Memory Policy — accounting and layer offload
Scope. How much memory a backend may use, when the refusal happens, and how the layer
offload plan is resolved. Both sections are decisions (E4, E5) whose per-ticket records are in
ARCHITECTURE-EXECUTION-PLAN.md §7.
AGENTS.md states each rule below as a one-line invariant and links here for the elaboration.
1. Account before allocating (E4)
- Memory is accounted before it is allocated, and the length contract lives in the
BufRef(E4).alloc_in_poolis fallible:weights (Backend::weights_bytes) + pooled + this allocation at its size classis checked against the backend's budget before the pool is asked for anything, and the refusal names the numbers (weights / pooled / request / budget, in MiB). The default budget is the backend's own answer — CUDA's current free bytes, or since #53 Metal'srecommendedMaxWorkingSetSize, with a quarter held back — resolved from an explicitallocplan::DeviceMemoryoutcome (Reported/QueryFailed/NoDevice) through the pureallocplan::budget_decision; a failed query is not a zero budget — it prints the real device error name once and charges weights only, while a measured zero still refuses (#122); CPU is unbounded unlessGraphAllocator::set_memory_budgetsets one.memory_report(backend)is the accounting surface:pool_bytes(what the pool holds — it only grows when the pool creates a buffer, so a recycled class buffer is not charged twice;Backend::pool_lenis the probe),live_bytes,peak_live_bytes,weights_bytes,budget,headroom_bytes(). The class ladder lives ingraph/allocplan.rs(class_size: powers of two to 16 KiB, then 16 KiB steps) and the pools allocate at it, so two shapes in one class share a buffer across a rebuild. Two rules follow, both latent bugs that rounding turned real: (a) a node's logical length isBufRef::len—fill_inputchecks the data against it and writes throughBackend::write_host_window, and every capture read is windowed (scheduler::window_of);write_hostkeeps the exact-length contract for the persistent KV regions and staging; (b) an input is host-filled before execution, so it must never take a buffer this build'ssweepreleased — all inputs are placed before the walk — and a liveness extension (in-place alias, D1 view) must move the buffer'sbuf_alivedeadline too (extend_through_views→extend_buffer_alive). S3 splits reservation from assignment: a released classed buffer goes to a reservation table (slots, keyed by(backend, class)), andalloc_class_in_pooltakes the smallest idle id of the class (deterministic — a LIFO list would move a rebuilt graph's slots around). A rebuild therefore re-maps:alloc_buffer/free_bufferare not called, CUDA'spool_gendoes not move (so its captured graphs survive), and the same topology gets the same slots.MemoryReportcarries the reservation's depth (idle_slots,reserved_classes), andCpuBackend::alloc_countis the CPU twin ofpool_genfor the gate.GraphCacheholds one graph perGraphParams(MRU,MAX_CACHED_GRAPHS), so a switch istry_reuse→alloc_graphwith no build and no fusion pass (stats()reports builds vs reuses). Cross-boundary staging is charged topool_bytesbut still allocated exact; backend-internal scratch is not in the report. Metal'sweights_bytescharges whatMpsStateregistered (#299) — the mmap-backedNoCopyweight slices and the per-weight copies alike — so the E4 gate subtracts the resident weights exactly as the E5autofit already charges them from the GGUF index. The term was the trait default0until #299, which let the gate admit the weight bytes more than the pool it protects had room for (on the measured Mac a 7B Q4_K_M's ~4.4 GiB, against a 38 339 MiB working set).
2. Layer offload is one plan read in three places (E5)
- Layer offload is one plan read in three places (E5).
graph/offload.rsownsOffloadPlan { gpu_layers, n_layers }: blocks0..gpu_layersrun on the device and the rest on the CPU, and tensors outside any block (embedding, final norm,lm_head) follow the device only when every block is offloaded (device_holds_unblocked— llama.cpp'sn_gpu_layers > n_layerconvention, stated once). The request is--gpu-layers N(CLI) orMINFER_GPU_LAYERS=N(environment; unset = every block a device can hold, the pre-E5 behaviour);OffloadRequest::planresolves it purely (CI-tested), and a spelling that is not a block count is a refused load, never a guess. The three readers must agree on the same number: (a) the loader registers a tensor on the device only whenOffloadPlan::allows_weight(name)says so, taking the block from the registry name (block_of:{ns}blk.{i}.…, including the fusedblk.{i}.attn_qkv/blk.{i}.ffn_guconcat copies); (b) the builder stampsCNode.layer(GraphBuilder::set_layer, once per block) and gates the device-only fused forms onlayer_gpu = gpu && il < gpu_layers; (c) the assignment pass (BackendScheduler::assign_backends→GraphAllocator::supports_for(op, dtype, layer)) never offers the device for a block past the plan.CParams.gpu_layerscarries it into the reuse identity, because the assignment is topology. The load verifies the plan against what was registered and drops to CPU-only with a printed reason when the offloaded blocks' weights are not usable there;offload_reportprints where the blocks landed with the measured device bytes. S2 (the automatic fit) adds theautospelling: the loader measures each block's weight bytes from the GGUF index (GgufTensorInfo::nbytes, before anything is loaded — the filter is the plan) and fits the largest prefix into the weight budget (fit_blocks, pure: a prefix, not a knapsack, because the plan is0..gpu_layersand a gap would put a CPU block between two device blocks for nothing), holding back a quarter for the KV arenas and the activation pool. The budget isMINFER_GPU_MEM=<MiB>if set, else three quarters of the device's own answer (CUDA's free bytes; Metal'srecommendedMaxWorkingSetSizesince #53) — the same default E4's feasibility gate uses, and the same resolver (models::device_memory()), so the fit and the gate talk about one number (weight_budget). Both device backends report a number on their platform; only a device with no state at all (the CPU, or a singleton that could not be created) fits nothing withoutMINFER_GPU_MEM, and a failed device query refusesautonaming the real error instead of reading it as "0 bytes free" (#122).
Decisions governing this document
This page is the current contract; the decisions behind it are frozen in the ADR corpus: