Memory Policy — accounting and layer offload

Scope. How much memory a backend may use, when the refusal happens, and how the layer offload plan is resolved. Both sections are decisions (E4, E5) whose per-ticket records are in ARCHITECTURE-EXECUTION-PLAN.md §7.

AGENTS.md states each rule below as a one-line invariant and links here for the elaboration.

1. Account before allocating (E4)

  1. Memory is accounted before it is allocated, and the length contract lives in the BufRef (E4). alloc_in_pool is fallible: weights (Backend::weights_bytes) + pooled + this allocation at its size class is checked against the backend's budget before the pool is asked for anything, and the refusal names the numbers (weights / pooled / request / budget, in MiB). The default budget is the backend's own answer — CUDA's current free bytes, or since #53 Metal's recommendedMaxWorkingSetSize, with a quarter held back — resolved from an explicit allocplan::DeviceMemory outcome (Reported / QueryFailed / NoDevice) through the pure allocplan::budget_decision; a failed query is not a zero budget — it prints the real device error name once and charges weights only, while a measured zero still refuses (#122); CPU is unbounded unless GraphAllocator::set_memory_budget sets one. memory_report(backend) is the accounting surface: pool_bytes (what the pool holds — it only grows when the pool creates a buffer, so a recycled class buffer is not charged twice; Backend::pool_len is the probe), live_bytes, peak_live_bytes, weights_bytes, budget, headroom_bytes(). The class ladder lives in graph/allocplan.rs (class_size: powers of two to 16 KiB, then 16 KiB steps) and the pools allocate at it, so two shapes in one class share a buffer across a rebuild. Two rules follow, both latent bugs that rounding turned real: (a) a node's logical length is BufRef::len — fill_input checks the data against it and writes through Backend::write_host_window, and every capture read is windowed (scheduler::window_of); write_host keeps the exact-length contract for the persistent KV regions and staging; (b) an input is host-filled before execution, so it must never take a buffer this build's sweep released — all inputs are placed before the walk — and a liveness extension (in-place alias, D1 view) must move the buffer's buf_alive deadline too (extend_through_views → extend_buffer_alive). S3 splits reservation from assignment: a released classed buffer goes to a reservation table (slots, keyed by (backend, class)), and alloc_class_in_pool takes the smallest idle id of the class (deterministic — a LIFO list would move a rebuilt graph's slots around). A rebuild therefore re-maps: alloc_buffer/free_buffer are not called, CUDA's pool_gen does not move (so its captured graphs survive), and the same topology gets the same slots. MemoryReport carries the reservation's depth (idle_slots, reserved_classes), and CpuBackend::alloc_count is the CPU twin of pool_gen for the gate. GraphCache holds one graph per GraphParams (MRU, MAX_CACHED_GRAPHS), so a switch is try_reuse → alloc_graph with no build and no fusion pass (stats() reports builds vs reuses). Cross-boundary staging is charged to pool_bytes but still allocated exact; backend-internal scratch is not in the report. Metal's weights_bytes charges what MpsState registered (#299) — the mmap-backed NoCopy weight slices and the per-weight copies alike — so the E4 gate subtracts the resident weights exactly as the E5 auto fit already charges them from the GGUF index. The term was the trait default 0 until #299, which let the gate admit the weight bytes more than the pool it protects had room for (on the measured Mac a 7B Q4_K_M's ~4.4 GiB, against a 38 339 MiB working set).

2. Layer offload is one plan read in three places (E5)

  1. Layer offload is one plan read in three places (E5). graph/offload.rs owns OffloadPlan { gpu_layers, n_layers }: blocks 0..gpu_layers run on the device and the rest on the CPU, and tensors outside any block (embedding, final norm, lm_head) follow the device only when every block is offloaded (device_holds_unblocked — llama.cpp's n_gpu_layers > n_layer convention, stated once). The request is --gpu-layers N (CLI) or MINFER_GPU_LAYERS=N (environment; unset = every block a device can hold, the pre-E5 behaviour); OffloadRequest::plan resolves it purely (CI-tested), and a spelling that is not a block count is a refused load, never a guess. The three readers must agree on the same number: (a) the loader registers a tensor on the device only when OffloadPlan::allows_weight(name) says so, taking the block from the registry name (block_of: {ns}blk.{i}.…, including the fused blk.{i}.attn_qkv / blk.{i}.ffn_gu concat copies); (b) the builder stamps CNode.layer (GraphBuilder::set_layer, once per block) and gates the device-only fused forms on layer_gpu = gpu && il < gpu_layers; (c) the assignment pass (BackendScheduler::assign_backends → GraphAllocator::supports_for(op, dtype, layer)) never offers the device for a block past the plan. CParams.gpu_layers carries it into the reuse identity, because the assignment is topology. The load verifies the plan against what was registered and drops to CPU-only with a printed reason when the offloaded blocks' weights are not usable there; offload_report prints where the blocks landed with the measured device bytes. S2 (the automatic fit) adds the auto spelling: the loader measures each block's weight bytes from the GGUF index (GgufTensorInfo::nbytes, before anything is loaded — the filter is the plan) and fits the largest prefix into the weight budget (fit_blocks, pure: a prefix, not a knapsack, because the plan is 0..gpu_layers and a gap would put a CPU block between two device blocks for nothing), holding back a quarter for the KV arenas and the activation pool. The budget is MINFER_GPU_MEM=<MiB> if set, else three quarters of the device's own answer (CUDA's free bytes; Metal's recommendedMaxWorkingSetSize since #53) — the same default E4's feasibility gate uses, and the same resolver (models::device_memory()), so the fit and the gate talk about one number (weight_budget). Both device backends report a number on their platform; only a device with no state at all (the CPU, or a singleton that could not be created) fits nothing without MINFER_GPU_MEM, and a failed device query refuses auto naming the real error instead of reading it as "0 bytes free" (#122).

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0015 — The offload auto fit takes a prefix, not a knapsack
  • ADR-0016 — A failed device-memory query is not a zero budget