minfer

A minimal local LLM inference engine built from scratch in Rust, modeled on llama.cpp's ggml_cgraph + backend scheduler, with zero ML framework dependencies.

  • Declarative compute graph — inference builds a ComputeGraph (pure IR), then assigns backends, fuses ops, allocates and executes via a scheduler. Params-only graph reuse (decode steps skip reconstruction).
  • Backends — CPU + Apple MPS (Metal); optional CUDA (feature-gated).
  • Models — Qwen2 / Qwen2.5 / Qwen3 (dense), GGUF v3 format.
  • Tooling — OpenAI-compatible server, CLI multi-turn conversation, an interactive web visualizer, and per-layer debug dumps.

Explore

Build & run

cargo build --release
./target/release/minfer <model.gguf> "hello"

Features

This page expands the feature list from the README with the full detail. Performance numbers refer to Qwen2.5-7B-Instruct q4_K_m prefill on GB10 (sm_121) unless noted.

Inference Core

Declarative compute graph

Inference builds a ComputeGraph (pure IR) then assigns backends, fuses ops, allocates and executes via a scheduler — inspired by llama.cpp's ggml_cgraph + backend scheduler. Graph reuse is params-only (decode steps skip reconstruction), backend assignment is per-op, and the whole design is documented in COMPUTE-GRAPH-DESIGN.md.

Interactive graph visualization (viz/)

A zero-dependency browser page for the compute graph. minfer viz <model> serves the page, live SSE inference, per-node tensor stats/heatmaps and logits top-5 in one process; --dump-graph-json / MINFER_TRACE export graphs and real traces. See the viz README for the full user guide.

GGUF loader

Parses GGUF v3 files (metadata + quantized tensors) with split multi-part support; weights are mmap'd and shared zero-copy with the GPU.

Self-contained BPE tokenizer

Loaded directly from GGUF metadata — no external dependency on tiktoken. tokenizer.ggml.pre selects the pre-tokenization rule (qwen2, alias deepseek-r1-qwen; qwen35), implemented as hand-written splitters because the regex crate has no lookahead; an unknown or missing rule, an empty merge table, a non-gpt2 model or an incomplete byte vocabulary refuses the load. Special tokens (GGUF type 3/4 table plus <|im_start|>/EOS fallbacks) match as single IDs before BPE, so special-token templates (DeepSeek-R1's <|User|>/<think>, etc.) tokenize exactly like llama.cpp, and an unmatched piece falls back byte by byte instead of silently emitting id 0. Token ids are gated byte-for-byte against transformers / llama.cpp over tests/fixtures/tokenizer/.

Backends

CPU — AVX2 / NEON+SDOT

All 8 quantized dot products as SIMD kernels (AVX2 on x86, NEON+SDOT via inline asm on Apple Silicon), plus a persistent row-parallel thread pool (-t/--threads). Qwen3-4B CPU decode runs ~52–58 tok/s on M4 Pro (vs 1.1 before the pool).

GPU — Metal (Apple Silicon)

Flash attention (single fused kernel for decode + prefill), simdgroup GEMM prefill for every quant type, SIMD-parallel RMSNorm, float4-vectorized kernels, a build-time precompiled .metallib (no per-run shader compile), and auto-selected f16 KV cache for 7B-class models. Tracked in METAL_OPTIMIZATIONS.md.

GPU — CUDA (NVIDIA, feature-gated --features cuda)

The performance headline of the project. The int8 tensor-core MMQ path is default-on in CUDA builds (opt-out per gate with "0"; MINFER_MMQ=0 restores the legacy f16 path):

  • Default prefill: ~3581 tok/s (7B q4_K_m @3314-token prompt) = 1.080× llama.cpp (llama-bench 3323.3 same shape) — from 441 tok/s when the path first landed, an 8.1× campaign documented step-by-step in CUDA_OPTIMIZATION.md (75-step history table).
  • Raw-nibble int8 mma.m16n8k32 GEMMs for q4_K and q6_K with producer-fused activation quantization (rms-norm/swiglu emit the transposed q8 plane directly, skipping intermediate writes), registration-time weight-expansion planes (W_exp / W_dsc) staged by cp.async, and flash attention with register-resident softmax (2.43× kernel).
  • CUDA Graph capture/replay for repeated identical-length prefills; decode uses the MMVQ weight-streaming path.
  • Memory/speed knobs: the weight-expansion planes cost ~3.3 GB device for ~+6% prefill; MINFER_MMQ_Q6K_EXP=0 / MINFER_MMQ_Q4K_DSC=0 return the memory.
  • Device adaptation (doc 105): the CUDA banner reports the resolved device tier — CUDA: device tier <name> (<provenance>, mmq <bool>) — from a cc-keyed table (GB10 measured; consumer GPUs adopted from llama.cpp; unknown → GENERIC). Dispatch gates (MMQ prefill availability, future batch caps) read the tier; on foreign devices the smem/VRAM feasibility checks self-degrade to slower-but-correct paths. MINFER_DEVICE_TIER=<key> forces a row for soak testing.

Model Support

Qwen2 / Qwen3 architectures

GQA attention, SwiGLU FFN, RoPE (Neox style), RMSNorm. Qwen3 adds the decoupled head dim and per-head Q/K RMSNorm (attn_q_norm/attn_k_norm, Op::QkNorm). Supported models include Qwen2.5 0.5B/7B, Qwen3 0.6B/4B, and DeepSeek-R1-Distill-Qwen-1.5B — see the matrix in AGENTS.md.

User-Facing

Model download

Auto-download from the Hugging Face Hub or the Ollama registry, with resume and cache-name resolution.

Multi-turn conversation CLI (--cnv)

Append-only KV + incremental chat-template rendering: each turn only prefills the new message delta while the whole conversation accumulates in the KV cache. In-session commands (/clear, /regen, …), automatic overflow handling, --session persistence. On overflow the dropped turn's KV rows are removed in place and the tail is re-based/re-roped (Phase C / C2), so the turn prefills its own delta instead of the retained history (measured 185 → 14 tokens per overflowing turn on the 0.5B probe); MINFER_NO_CONTEXT_SHIFT=1 restores the exact drop-and-re-render path. Plan: CLI-CONVERSATION-PLAN.md.

OpenAI-compatible HTTP server (serve)

/v1/chat/completions (streaming + non-streaming), /v1/models, /health, and /metrics — a Prometheus text snapshot of request counts, queue depth, live KV/arena occupancy and (under MINFER_OP_TIMING) per-op seconds; multi-slot with queued serial execution (MINFER_BATCH=0); a request that finds every slot busy on the batched path is refused with 503 (minfer_jobs_dropped_total +1, an SSE error frame when streaming) rather than queued — queueing is #150; a decode step whose forward fails answers every run in that batch with 500 server_error, clears their cached prefix and frees the slots (#151), so a deterministic failure cannot spin the worker on the same forward; a worker wedged inside a step ends itself after a counted STALL_STEP_LIMIT (64) consecutive no-progress steps, answering every live and queued request once with 500 server_error and publishing minfer_worker_stalled_total (#196) instead of spinning at 100% CPU; SIGINT/SIGTERM drain bounded by MINFER_DRAIN_MS. Plan: OPENAI-CHAT-API-PLAN.md, F8 record: ARCHITECTURE-EXECUTION-PLAN.md.

Chat templates (F7)

The model's own tokenizer.chat_template is rendered by minijinja plus a Python-str-method hook, so the published Qwen2.5/Qwen3 templates (including Qwen3's <think>-block split and re-emission) render as published. Reference renderings for every supported model live in tests/fixtures/chat/ (transformers 5.17.0, generated from each model's tokenizer_config.json), and a rendered prompt must equal them byte for byte. A template the engine cannot render is a loud refusal naming the construct and the template line — checked at load, so the CLI exits and the server refuses to start; the generic ChatML renderer applies only to a GGUF with no template at all. Design + accepted and refused construct sets: CHAT-TEMPLATE-AND-TOKENIZER-DESIGN.md.

Constrained decoding — grammar and JSON Schema (F2)

A GBNF-style grammar or a JSON Schema is compiled once per request into a pushdown automaton that masks the logits inside the one sampler pipeline, so decoding cannot leave the accepted language.

  • GBNF subset: rules, string literals, character classes with negation, ., grouping, alternation, */+/?, repetition ranges {m}/{m,}/{m,n}, and # comments.
  • JSON Schema subset: type (string or array), enum, const, object properties/required/additionalProperties, array items/prefixItems/minItems/maxItems, strings, integer bounds (inclusive and exclusive), numbers, booleans, null, anyOf/oneOf, and $defs + local $ref (recursive schemas work).
  • Anything outside the subset is a loud refusal naming the construct (CLI startup error or HTTP 400) — never a silent guess. The catalogues are in GRAMMAR-DESIGN.md.
  • Token advancement is byte-level correct: a token whose piece is one byte of a multi-byte character is handled, a token that a rule only partially accepts is rejected with its longest accepted prefix named, end-of-generation is legal only at a complete state, and "no token is allowed" stops with a printed reason instead of emitting an arbitrary token.
  • Surfaces: CLI --grammar/--grammar-str/--json-schema/--json-schema-str; the server's response_format (json_object / json_schema) and a grammar extension field. The mask is cached per automaton state and computed with a DFA-style transition memo (measured 5.4 ms per new state on a 151,936-token vocabulary).

Performance benchmark (bench)

minfer bench <model> runs llama-bench-style prefill (pp<P>) / decode (tg<T>) throughput tests on the active backend — mean ± stddev over reps after an untimed warmup, each rep from an empty KV context without a model reload — reported as a markdown/CSV/JSON table.

GGUF tooling — convert / quantize / split (F6)

minfer convert <hf-dir> out.gguf produces a GGUF v3 from a HuggingFace Qwen2 checkpoint with the metadata the strict tokenizer/template loader requires (tokenizer.ggml.model/pre, the 256 byte tokens, merges, special ids and tokenizer.chat_template); minfer quantize in.gguf out.gguf --type … re-encodes weights to q4_0/q4_1/q5_0/q5_1/q8_0 (byte-identical to llama-quantize on the same source) or f16/f32; minfer split in.gguf dir --max-size N writes split.no/split.count parts the loader merges back into one tensor index. Unsupported architectures, tensor names, dtypes and quant targets are refused by name. A converted f16 GGUF loads and runs (f16 weights are CPU-only). See GGUF-TOOLING.md.

Philosophy

No external ML framework — pure Rust; runtime deps are minimal (rand, regex, half, serde, serde_json, minijinja; axum/tokio only for the HTTP server). Attention, RMSNorm, RoPE, SiLU, Softmax and every quantized dot product are handwritten.

Build

Requirements: Rust (edition 2021, no ML-framework dependencies — runtime deps are minimal). A Nix flake devShell (nix develop) is available for a batteries-included dev environment.

Build commands

# CPU + Metal (macOS) — plain build, never touches nvcc
cargo build --release
./target/release/minfer <model.gguf> "hello"

# CUDA (NVIDIA GPU) — opt-in feature, requires the CUDA toolkit (nvcc)
cargo build --release --features cuda
# statically-linked cudart (no libcudart.so runtime dep)
cargo build --release --features cuda,cuda_static

# + per-node debug dumps (MINFER_DUMP_DIR)
cargo build --release --features debug_dump

GGUF tooling (F6) — no extra dependencies

minfer convert / quantize / split need no Python, no PyTorch and no gguf-py: safetensors is parsed as a length-prefixed JSON header plus raw bytes and tokenizer.json/config.json are plain JSON, both through the already-present serde_json. They are host-side file work — no GPU backend is initialized — and they build in every configuration (CPU, Metal, CUDA). The writer/encoders and their verification references are documented in GGUF-TOOLING.md.

Tests

cargo test --release                         # unit + integration, no model files needed
cargo fmt --all --check                      # formatting, with the pinned toolchain's rustfmt
scripts/real_model_gates.sh                  # the #[ignore]d real-model gate set (parallel on CPU)
PARALLEL=0 scripts/real_model_gates.sh       # CPU-only serial form
FEATURES=cuda scripts/real_model_gates.sh    # device gate set (serial, one GPU)
scripts/cuda_test.sh                         # the whole CUDA suite on a real GPU, serial

The #[ignore]d set needs the cached real models (the 0.5B for the default configuration, MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf for the f16-KV one) and writes temporary session files. scripts/real_model_gates.sh is the documented entry point: it runs serially when FEATURES includes cuda, which a device build requires (the CUDA state is a process-wide singleton, issue #64), and in parallel otherwise — on a CPU-only build that reason does not exist, and the parallel harness is where the server batching gate checks its own robustness. PARALLEL=1/0 overrides. Each platform's CI job compiles the test target as well as the crate: cargo test --release --no-run on the Linux CPU job (implicitly), cargo test --release --features cuda --no-run on the CUDA job and cargo test --release --no-run on build-macos (#303) — cargo build does not compile #[cfg(test)], and a macOS-only test module was invisible for eleven days because of it. Those three jobs are skipped when the change classifier finds no code change (ADR-0025): a docs-only change runs the doc gates instead, and the suite still runs whenever the suite ledger moves. Since the KV storage format became per engine (#99) the parallel harness no longer makes one gate size another gate's KV regions. The current counts live in TEST-BASELINES.md and its ledger scripts/test-baselines.toml, each record dated, box-labelled and command-labelled. Four fixes now make the parallel form trustworthy: the server batching gate (interleaved matched rounds, a median verdict), #158 (a work bound instead of absolute wall-clock deadlines), #160 (every while engine.busy() stepper in the server gates routed through one shared per-step WorkBound — progress plus a step budget, both counted, never a clock), and #196 (the production serve_loop itself carries the counted no-progress bound STALL_STEP_LIMIT = 64 consecutive steps, so the two gates that drive it are wedge-proof too, and a wedged server answers every live and queued request once and publishes minfer_worker_stalled_total instead of spinning). A wedge injected with MINFER_TEST_TICK=wedge now fails each of those gates in seconds instead of hanging the suite; =spin exercises the step-budget arm (the two serve_loop gates still have to be --skipped there: a step that keeps moving the work counter is what a counter cannot catch — see the #196 record). The wall-clock bounds that remain are named backstops, not verdicts: the two run_cli child-process ceilings (1800s for the real-model sessions, 60s for the no-model registry cases; MINFER_CLI_WATCHDOG_SECS overrides both) and the serve_loop feeder's FEEDER_POLL_BACKSTOP, now the last resort for a worker whose published metrics never settle rather than the only way out of a wedge. The rules behind all of this are GATE-CONTRACT.md §3/§4 and the full records are ARCHITECTURE-EXECUTION-PLAN.md §test-infrastructure.

Git hooks (git-hooks.nix)

The formatting gate is provided by git-hooks.nix (flake input) instead of a Cargo dev-dependency. Pinned in flake.lock; the rustfmt --check hook runs with the project's pinned toolchain (1.97.1).

  • The hook is installed when entering the devShell (nix develop installs .git/hooks/pre-commit + the .pre-commit-config.yaml symlink into the nix store). Committing outside the shell has no hook — those contributors should run cargo fmt --all -- --check manually first.
  • Format drift is rejected, not auto-fixed: run cargo fmt --all, re-stage the changed files, then commit again.
  • The same check runs sandboxed as a derivation: nix flake check (checks.pre-commit-check), or nix develop -c pre-commit run --all-files.
  • CI enforces it too (#211): the test-linux-cpu job runs cargo fmt --all --check before the suite, with the toolchain step installing the rustfmt component for the pinned 1.97.1 (dtolnay/rust-toolchain@1.97.1 + components: rustfmt — 1.97.1 ships no rustfmt by default). The pinned toolchain's rustfmt is used deliberately, never stable's: a formatter that disagrees with the toolchain is exactly how a "green CI, red hook" split appears.
  • In a nested worktree (.worktrees/<scope>, the documented parallel-work shape) only tracked files exist, so the ignored .pre-commit-config.yaml is absent and pre-commit run cannot work. The substitute is the CI command itself, after installing the component once: rustup component add rustfmt --toolchain 1.97.1 && cargo fmt --all --check. cargo fmt reads rust-toolchain.toml, so it uses the pinned rustfmt rather than whichever one stable happens to carry.

macOS / Metal notes

On macOS the Metal backend is built in automatically: build.rs compiles src/metal/kernels/ → a precompiled .metallib via /usr/bin/xcrun at build time (no per-run shader compile). If Xcode tools are unavailable the build still succeeds and the shader is compiled from source at first run.

CUDA build details

  • nvcc is located via PATH, CUDA_HOME/CUDA_PATH (e.g. /usr/local/cuda); with --features cuda a missing toolkit is a hard error (with a clear message), and plain builds never touch nvcc at all.
  • The host compiler is auto-detected: nvcc's default works when compatible; otherwise the first accepted GCC is pinned via -ccbin (e.g. a nix devShell putting GCC 15 first while CUDA 13 accepts ≤ GCC 13). Force one with MINFER_CUDA_CCBIN=/path/to/g++.
  • GPU architectures are auto-detected from what the toolkit accepts (SASS for sm_70…sm_121 as available, plus PTX for the highest and a backward-JIT compute_70/compute_72 PTX) — one binary covers older and newer GPUs. The candidate list includes sm_87/sm_88 (Jetson Orin) explicitly: the device-tier gate routes sm_87 to the int8 BT path, which is compiled out of any PTX below sm_80 — native SASS per target is mandatory because PTX JITs forward only (docs 105–106, plan §14 R9). The minimum is sm_70 (Volta): the kernels in src/cuda/kernels/*.cu use WMMA tensor cores (nvcuda::wmma), which require sm_70+, so Pascal (sm_61) is not a target. The Volta V100/Titan-V (sm_70/72) PTX is only emitted when nvcc supports it: CUDA 12.x does, CUDA 13 removed Volta, so keep Volta coverage by building with CUDA 12.8 (the only version supporting Volta + the Blackwell RTX 50 sm_120/121, which needs ≥ 12.8).
  • cudart linking (mirrors llama.cpp's GGML_STATIC): by default -lcudart is a shared link, so the binary NEEDEDs libcudart.so.N and needs the CUDA toolkit runtime present at runtime (an rpath to <cuda_home>/lib64 is baked in). Adding cuda_static links libcudart_static.a instead: the binary has no libcudart.so NEEDED dependency and only needs the NVIDIA driver (libcuda.so.1, dlopen'd lazily at runtime) + libstdc++ — deployable without a CUDA toolkit. The driver is never a link-time dependency in either mode.
  • In CUDA builds the int8 tensor-core MMQ prefill path is default-on (runtime gates MINFER_MMQ*, see the README's Performance section / CUDA_OPTIMIZATION.md).

Usage

Examples use the built binary (cargo build --release first — see the Build section of the project README); cargo run --release -- … works identically.

./target/release/minfer <model> [prompt] [OPTIONS]

<model> can be a local path, a download URI, or a cached model name:

FormatExample
Local file~/models/qwen2.gguf, ./model.gguf, /abs/model.gguf
Hugging Facehf:Qwen/Qwen2-0.5B-GGUF:qwen2-0.5b-q4_0.gguf (auto-download)
Ollamaollama:qwen2.5:0.5b (pull)
Cached model nameqwen2.5-0.5b-instruct-q4_0 (resolved from ~/.cache/minfer/models, see list)

If prompt is omitted, reads from stdin. Run minfer --help for the full option list; the subcommands are:

CommandPurpose
<model> [prompt] [OPTIONS]single-shot generation
serve [--port N] [--n-ctx N] [--n-slots N] <model>OpenAI-compatible HTTP server
info <model>print GGUF metadata + key tensors
download hf <repo> [quant] / download ollama <model>[:tag]fetch models
listlist locally cached models
viz [--port N] <model>self-contained viz server (default port 8081)
bench [-p N] [-n N] [-r N] [-o md|csv|json] <model>perf test: pp<P> prefill / tg<T> decode, mean ± stddev over reps
specverify [-p N] [-r N] [-o json|md] <model>D5-1a gate bench: batched verify cost C_T(nt) + per-token amortization at deep KV

Sampling and runtime options:

  • --temp <T> — sampling temperature (default 0.8; --greedy = greedy decoding)
  • --top-k <K> / --top-p <P> — top-K / nucleus sampling (defaults 40 / 0.95)
  • --repeat-penalty <N> — repeat penalty (default 1.1; 1.0 = off), plus --frequency-penalty / --presence-penalty
  • F3 sampler set (#48) — every default below leaves the pre-F3 chain unchanged:
    • --min-p <P> — drop tokens whose probability is below P * max (0 = off; 1.0 = argmax only)
    • --typical <P> — locally typical sampling (1.0 = off; 0 keeps the single most typical token)
    • --xtc-probability <P> / --xtc-threshold <T> — XTC: with probability P, exclude the top choices whose probability is at least T (T <= 0.5; 0 = off)
    • --dry-multiplier <N> — DRY (Don't Repeat Yourself) penalty strength (0 = off), with --dry-base <N> (default 1.75), --dry-allowed-length <N> (default 2), --dry-penalty-last-n <N> (default 64) and --dry-sequence-breakers <L> — restart sequences as token ids, e.g. --dry-sequence-breakers 198;13,2 (; between sequences, , between ids)
    • --mirostat <0|1|2> — mirostat off / v1 / v2, with --mirostat-tau <N> (target surprise in bits, default 5.0), --mirostat-eta <N> (learning rate, default 0.1) and --mirostat-m <N> (v1 estimator window, default 100). In mirostat mode the temperature is ignored (mirostat's mu truncation subsumes it); --temp 0 still means greedy. Mirostat cannot be combined with --spec-draft.
    • --logit-bias <L> — add to raw logits: ID:BIAS pairs separated by ,, repeatable, e.g. --logit-bias 15043:-2.0,198:1.5. A token id outside the vocabulary, or a bias outside [-100, 100], is refused at startup. A nonsensical value for any of these is refused at startup (exit 1), never silently ignored.
  • F2 constrained decoding (#47) — mutually exclusive; an unsupported construct is refused at startup (exit 1) with the offending token:
    • --grammar <FILE> / --grammar-str <GBNF> — constrain sampling to a GBNF grammar
    • --json-schema <FILE> / --json-schema-str <JSON> — constrain sampling to a JSON Schema (compiled to GBNF internally) The grammar is compiled once against the loaded vocabulary; the mask is applied inside the sampler, after the penalties/DRY and before the greedy shortcut, so --greedy respects it. --spec-draft cannot be combined with a grammar (a verify round samples several rows from one automaton state). The accepted GBNF and JSON-Schema subsets, and every construct that is refused, are catalogued in GRAMMAR-DESIGN.md; the honest narrowing to know about is that object properties are accepted in declaration order and oneOf is compiled as anyOf. Example:
    minfer model.gguf "Give me a person" --json-schema-str \
      '{"type":"object","properties":{"name":{"type":"string"},"age":{"type":"integer","minimum":0,
        "maximum":150}},"required":["name","age"],"additionalProperties":false}'
    # -> {"age": 25, "name": "John Doe"}
    
  • --stop <STR> — stop generation at this string (repeatable)
  • -n, --n-predict <N> — max tokens to generate (default 512)
  • --seed <N> — RNG seed for sampling
  • Speculative decoding (D5-R, ADR-0017) — a draft model proposes, the target verifies the proposals in one batched forward:
    • --spec-draft <model> — the draft model (any supported GGUF)
    • --spec-draft-n <N> — drafted tokens per round (default 2, so the verify batch is N + 1 rows)
    • --spec-draft-adaptive — pick d per round from the per-depth acceptance and cost EWMA, capped at 8 unless --spec-draft-n sets a smaller cap --spec-draft-n and --spec-draft-adaptive are no-ops without --spec-draft. Speculative decoding refuses mirostat and a grammar (both named above), because a verify round samples several rows from one shared RNG, or from one automaton state.
  • --n-ctx <N> — sizes the KV cache (clamped to the model's max context)
  • -t, --threads <N> — CPU worker threads
  • --gpu <N> — CUDA device index; unset auto-selects the highest compute capability. Ignored on CPU and Metal, where there is nothing to index.
  • --backend <name> — F4: fence the run to a set of backends. Repeatable, and both the comma-separated (--backend cpu,cuda) and the =-joined (--backend=cpu) forms work; the names cpu, metal and cuda are matched case-insensitively. cpu stays allowed whatever is named — it is the fallback — so --backend cpu is how a run is forced onto the CPU. Unset = every backend this build has; MINFER_BACKENDS is the environment form of the same surface, and the flag wins when both are set. It is extracted before subcommand dispatch, so serve, viz, bench and specverify honour it too. Three distinct refusals, each naming what it rejected: an unknown name (unknown backend 'nope'; known backends are: cpu, metal, cuda), a name this build does not contain (backend 'cuda' is known but not compiled into this build: the CUDA backend is compiled only with --features cuda), and a name this machine cannot use (backend 'cuda' is compiled in but not available on this machine: <why>).
  • --gpu-layers <N|auto> — E5: run the first N transformer blocks on the device and the rest on the CPU (0 = CPU only; unset/MINFER_GPU_LAYERS = every block the device can hold, the pre-E5 behaviour; auto = as many as the budget allows). The placement is printed at load (offload: 4 of 24 blocks on cuda, 20 on cpu; embed/output on cpu (32.0 MiB of device weights; --gpu-layers 4)), and the tensors outside the blocks — token_embd, the final norm and lm_head — stay on the CPU unless every block is offloaded.
  • MINFER_GPU_MEM <MiB> — the weight budget auto fits into; unset = three quarters of the device's free bytes (the same default the activation gate uses). auto also holds back a quarter of that budget for the KV arenas and the activation pool, and reports what it decided (offload: … auto: 5 of 24 blocks fit — weights budget 64 MiB, 16 MiB reserved for KV/activations; MINFER_GPU_MEM=64 MiB). Without a device that reports free memory (Metal today) auto needs MINFER_GPU_MEM.

Three flags describe the run instead of changing how it decodes:

  • --meta — dump the GGUF's metadata KV pairs and the key tensors (token_embd, the norms, lm_head) — the same two dumps minfer info <model> prints. Unlike info it does not stop there: the model still loads and generation proceeds, so it is the flag form for "show me the file, then answer".
  • --dump-graph <PATH> — export the compute graph the run just built (build → assign → fusion, with the real backend assignment and the same fusion environment toggles as the live path) as Graphviz DOT. It runs one prefill, writes the file, prints Graph DOT exported to <PATH> (<n> nodes) and exits before decoding.
  • --dump-graph-json <PATH> — the same graph as the JSON the viz/ web visualizer reads. It prints Graph JSON exported to <PATH> (<n> nodes, <kind>), where <kind> names the phase the graph was built for, and also exits before decoding.

Chat templates and tokenizer

Chat rendering uses the model's own tokenizer.chat_template from the GGUF. The published Qwen2.5/Qwen3 templates use Python string methods (content.split('</think>'), .lstrip('\n')) that minijinja does not provide natively; minfer supplies them through minijinja's unknown-method hook with CPython semantics, so those templates render as published (this is what makes Qwen3's think-block handling and tool-call formatting reach the model). Reference renderings — multi-turn, a system message, a generation prompt, a <think>-reasoning turn — are committed under tests/fixtures/chat/ with their provenance.

A template the engine cannot compile or render is an error, never a generic prompt:

Error: chat template error — unsupported template construct: unsupported Python str
method `splitlines` (template line 41); minfer refuses to fall back to a generic
ChatML prompt. Supported Python str methods: capitalize, count, endswith, find,
join, lower, lstrip, replace, rfind, rsplit, rstrip, split, startswith, strip,
title, upper.

The CLI exits before inference, serve/viz refuse to start (checked before the worker thread is spawned), and a per-request refusal is an HTTP 400. The generic ChatML renderer survives only for a GGUF that has no tokenizer.chat_template at all, and the startup path prints a one-line notice when that happens. --no-template still bypasses template rendering entirely.

The tokenizer is byte-level BPE and is equally strict about what it does not implement. tokenizer.ggml.pre selects the pre-tokenization rule — qwen2 (alias deepseek-r1-qwen) or qwen35 — and an unknown or missing value, a tokenizer.ggml.model that is not gpt2, an empty merge table, or a vocabulary missing any of the 256 byte tokens refuses the load with the offending value named. Token ids match the reference (transformers AutoTokenizer, or llama.cpp llama-tokenize on the same GGUF) byte for byte over the corpus committed in tests/fixtures/tokenizer/.

Multi-turn conversation

--cnv (docs/CLI-CONVERSATION-PLAN.md): append-only KV + incremental template rendering — each turn only prefills the new message delta, the whole conversation accumulates in the KV cache:

./target/release/minfer --cnv qwen2.5-0.5b-instruct-q4_0           # interactive REPL
./target/release/minfer --cnv -st qwen2.5-0.5b-instruct-q4_0 "hi"  # single turn

In-conversation commands: /exit /quit, /clear, /regen (regenerate the last reply), /help; EOF (Ctrl+D) exits.

Conversation options:

  • -st, --single-turn — run one turn, then exit
  • --system <STR> — system prompt
  • -mli, --multiline-input — submit input on an empty line
  • --color on|off|auto — color output (default auto = tty)
  • --session <FILE> (with --cnv) — save/load the conversation. FILE is the history as JSON; FILE.kv is a KV session companion written next to it (C5) that carries the rows those messages were rendered into, plus the host state they belong to. On start a matching companion is resumed and the history is not re-prefilled — the run prints resumed N message(s) and M KV row(s) … — 0 tokens prefilled; anything that does not match this run (another --n-ctx, another model's n_kv_embd, another MINFER_CACHE_TYPE, a history the user edited, an older file version) prints the reason and falls back to re-rendering the JSON, which is always correct if slower. On overflow the oldest turns are dropped automatically and generation continues: the dropped turn's KV rows are removed in place and the tail is re-based (Phase C / C2), so only the new turn's delta is prefilled; MINFER_NO_CONTEXT_SHIFT=1 forces the older, exact "drop the turns and re-prefill the rest" behaviour (an engine that cannot move rows — e.g. Metal, where it is Phase G — falls back to that path on its own and says so on stderr)

Qwen3-style <think>…</think> reasoning blocks are gray-highlighted (single-shot mode too, when stdout is a terminal or MINFER_COLOR=1).

OpenAI-compatible HTTP server

Continuous batching (the worker composes one decode batch across the active slots instead of one forward per slot — Phase E / E2) is on by default when the model's forwards run on CUDA or Metal, and off on CPU (E6; Metal joined in #44 part (b), 2026-10-06). The default was decided by measurement, not assumption: on this project's reference CPU batching measured slower than serving requests one at a time (0.49x on 7B Q4_K_M, 0.88x on 0.5B Q4_0, --n-slots 4), while on the GB10 it measured 1.97x faster (7B Q4_K_M, four identical prompts, equal work, --n-slots 4, default settings).

Both readings are dated, and both are historical. The CPU pair is the E2 record's step 3 (2026-09-17); the GPU 1.97x is E6's re-fetch with the default and no environment variable (2026-09-19, matching E2's 1.9x of 2026-09-18 within noise). All three predate the cross-request prefix sharing described below (C8a/C8b, 2026-09-21/22), so the CPU figure prices a per-request prefill that sharing has since made avoidable, and the reason the E2 record gave for it — concurrency forfeits the cross-request prefix reuse each slot otherwise keeps — no longer holds. They are history, not a current claim; the boxes, the workloads and the tables are in the E2 and E6 records of ARCHITECTURE-EXECUTION-PLAN.md §5.

--n-slots does not cap a request. The arena is divided among the slots as an initial, elastic share only: the server starts from n_ctx / n_slots per slot, but a request's real bound is the whole n_ctx. When a request needs more than its share, admission reclaims cells from idle slots (releasing their cached prefixes) and the allocator moves whatever runs are in the way, so a busy neighbour cannot block it (C7/C7b, the same plan §5). The serial path (MINFER_BATCH=0) is the exception: its per-slot graph region really is n_ctx / n_slots cells.

  • --slots-file <PATH> (batched engine only) — resume the server's slot contexts from PATH at startup and rewrite it after every completed request (C5 S2). A request whose prompt matches a restored slot's tokens prefills only its own delta; the file also carries the KV rows, so a restart no longer re-prefills the conversations that had finished. The startup line prints the size it will write per request (about 12 MiB for a 0.5B/512-row arena), because that cost is a decision. A snapshot from another --n-slots or --n-ctx (or another model / KV element type) is refused loudly and the server starts empty.
  • MINFER_BATCH=1 forces batching on (this is how to batch on CPU, for experiments or for a machine where your own measurement says it wins).
  • MINFER_BATCH=0 forces it off.
  • Any other value warns and uses the device default.
  • Metal joined the default only once it could take the node the batched path needs: until #44 part (b) (2026-10-06) supports_attn_span() was false there, so batching would have failed loudly instead of serving and Metal stayed opt-in. The explicit switch is no longer needed.
  • The server prints its choice at startup: [server] batching: on (device cuda; MINFER_BATCH=1 forces it on, =0 forces it off).

Two fuse-related switches are easy to confuse (D3):

  • MINFER_NO_FUSE_QKV=1 / MINFER_NO_FUSE_FFN=1 disable the corresponding decode fusion (the decoder builds the plain matmul/rope/store — or gate+up+ silu+mul — path instead).
  • MINFER_FFN_COMPOSITION=1 keeps the FFN fusion but builds it as the proven composition (concat matmul + gate/up windows + in-place SwiGLU) instead of the hand-written fused node. It is the reference the A/B is run against, and it is ignored with a warning on a backend without offset views (Metal until G5).

The KV cache type is its own switch (C4): MINFER_CACHE_TYPE=f32|f16|q8_0, strict — an unknown value fails the load on every device, f16 resolves to f32 on the CPU (no f16 KV kernel there) and q8_0 is refused where the attention kernel has no packed read — it is supported on the CPU, CUDA and Metal (Metal since #310). A packed q8_0 cache is 3.76× smaller and, since C4 S2, is read by a fused Q8_0 × Q8_0 K dot with V accumulated out of the cell; MINFER_NO_FUSED_Q8_KV=1 restores the older dequantize-into-a-scratch read for the A/B. Details, numbers and the named tolerance class: docs/BACKENDS.md, docs/ARCHITECTURE-EXECUTION-PLAN.md §5 (C4).

Two environment switches around the GPU are easy to get wrong:

  • MINFER_DISABLE_CUDA is checked for presence, not value: setting it to 0 disables CUDA (and therefore also turns the batching default off, since the model then runs on CPU). To force the CPU path deliberately use MINFER_DISABLE_CUDA=1; to use the GPU, leave it unset.
  • On CUDA, batched prefills stay per request by construction (fa_prefill tiles one query tile against one KV window — E1b), so --n-slots concurrency still pays one prefill per request there. The batched-decode win is unaffected.
  • Prefix reuse across slots is a share, not a copy, wherever the attention kernel gathers a kv_map (CPU, CUDA, and Metal since #362): admission points the arriving request at the donor slot's rows and no byte is copied. A device that cannot gather, and any device at all under MINFER_NO_KV_SHARE=1 (the A/B switch), falls back to C8a's row copy — either way the arriving request prefills only its own suffix.

Slot saturation

A request that arrives while every engine slot is busy is rejected loudly, not queued (the decision #121 pinned down). The worker answers

  • non-streaming: 503 Service Unavailable with {"error":{"code":503,"message":"no idle slot","type":"unavailable_error"}};
  • streaming: the event stream has already started, so the status line is 200 and the signal is a data: error frame carrying the same error object (followed by [DONE]) — not an empty stream a client could mistake for a generation that produced nothing.

minfer_jobs_dropped_total moves by exactly one per rejected request. The alternative — holding the request until a slot frees, which minfer_queue_depth would then measure and which the pre-E2 plan assumed (see OPENAI-CHAT-API-PLAN.md §"Slot Lifecycle") — is a feature request, #150; the serial path (MINFER_BATCH=0) still queues, because its single worker pulls one job at a time from the same channel.

Structured output (F2)

/v1/chat/completions accepts OpenAI's response_format and a llama.cpp-style grammar extension:

FieldEffect
"response_format": {"type": "text"} (or absent)no constraint
"response_format": {"type": "json_object"}any single JSON value
"response_format": {"type": "json_schema", "json_schema": {"name": "person", "schema": {…}}}the compiled schema (name and strict are accepted and ignored)
"grammar": "root ::= …"a GBNF grammar inline

grammar together with a non-text response_format is a 400 (one grammar per request, never a precedence rule), and so is any unsupported GBNF/schema construct — the schema is compiled on the handler side, before the request takes a slot, so the error is an HTTP 400 with the offending construct named, not a truncated generation:

curl -s localhost:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "messages": [{"role":"user","content":"Give me a person"}],
  "temperature": 0, "max_tokens": 40,
  "response_format": {"type":"json_schema","json_schema":{"name":"person","schema":
    {"type":"object","properties":{"name":{"type":"string"},"age":{"type":"integer","minimum":0,
     "maximum":150}},"required":["name","age"],"additionalProperties":false}}}}'
# {"choices":[{"message":{"content":"{\n  \"age\": 25,\n  \"name\": \"Qwen\"\n}"}}]}

A grammar is per request: the compiled automaton is shared by Arc and each request (serial or batch slot) carries its own position, so slots cannot perturb each other. A response body never ends mid-character: if a generation stops with a partial UTF-8 sequence pending, those bytes are dropped (with a printed note) rather than decoded to U+FFFD.

Metrics and observability (F8)

GET /metrics returns a Prometheus text snapshot (content type text/plain; version=0.0.4; charset=utf-8). It is served from its own router state, so a scrape needs neither the tokenizer nor the job channel and keeps answering while the model is busy; rendering only reads atomics, so scraping cannot perturb generation.

Family names and units (every family is minfer_-prefixed; *_bytes is bytes, the per-op family is seconds, everything else is a count):

  • Request lifecycle: minfer_requests_total, minfer_requests_completed_total, minfer_requests_rejected_total (refused before queueing: the server is draining, or the worker is gone), minfer_requests_in_flight (accepted and not yet finished — the drain surface), minfer_jobs_dropped_total (the worker could not place the job), minfer_worker_stalled_total (the worker's counted no-progress bound tripped: every live and queued request was answered 500 the worker stalled and the worker stopped — #196).
  • Depth: minfer_queue_depth (accepted - admitted: the channel backlog plus the worker's pending deque — the one number neither thread can see alone), minfer_worker_pending_jobs, minfer_requests_running (occupying an engine slot now).
  • Throughput: minfer_prompt_tokens_total, minfer_completion_tokens_total, and minfer_completion_tokens_per_second — generated tokens/s over a trailing 16-second window (a lifetime average would keep reporting a startup burst on an idle server). Tokens are counted where the response is produced, so the batched and serial paths agree; a client that disconnects before its answer is complete is not counted, because those tokens were never delivered.
  • Drain: minfer_draining (0/1), minfer_drain_abandoned_requests (still in flight when the deadline expired; 0 is a clean drain).
  • Allocator (E4 MemoryReport, for the backend the server runs on): minfer_memory_{weights,pool,live,peak_live,budget,headroom}_bytes, minfer_memory_idle_slots, minfer_memory_reserved_classes. On an unbounded backend (CPU/Metal with no budget) the budget/headroom families are omitted, not reported as 0.
  • KV arena: minfer_kv_{layers,rows,region_bytes}, minfer_kv_packed (1 for a packed Q8_0 cache), minfer_kv_{reserved,owned,shared,free}_cells, minfer_kv_{free_runs,sequences}, and the C3/C8b counters minfer_kv_{defrags,cells_moved,cows,cow_cells}_total.
  • Per-op timing (present only when the flag below is set): minfer_op_seconds_total{op="matmul"} and minfer_op_calls_total{op=…}. The interval is the scheduler's per-node dispatch, so it includes the backend's own prologue and excludes split-level syncs and cross-backend copies; only ops that actually ran appear.

Occupancy is a live reading: the worker republishes the allocator's numbers after every step. On the batched path that is the one shared arena; the serial path has one arena per slot, so it reports the arena of the slot that served the last request.

Flags:

  • MINFER_OP_TIMING — presence-checked (any value), off by default. Turns on the per-op timing above. Off, the scheduler never reads the clock and the timing family is absent from a scrape; on, the numbers reported change and the numbers computed do not (a greedy run is identical with and without it).
  • MINFER_DRAIN_MS — how long a graceful shutdown may take, in milliseconds (default 30000). A value that is not a whole number of milliseconds is reported and the default is used; 0 means "stop now". See below.
MINFER_OP_TIMING=1 ./target/release/minfer serve --n-ctx 4096 --n-slots 1 qwen2.5-0.5b-instruct-q4_0
curl -s http://127.0.0.1:8080/metrics

Graceful shutdown (F8)

On SIGINT or SIGTERM the server stops accepting new work and lets the requests already accepted finish, up to MINFER_DRAIN_MS:

  1. the listener is closed (a brand-new connection gets a connection error), and a request that arrives on an already-accepted connection gets 503 — it is counted in minfer_requests_rejected_total;
  2. in-flight responses are allowed to complete;
  3. at the deadline the server logs how many requests were still in flight, records that count in minfer_drain_abandoned_requests, and exits. It never waits on the worker indefinitely — an SSE client that never disconnects cannot keep the process alive.

With no signal the server runs forever exactly as before, and the --slots-file snapshot (written after every completed request) is unaffected.

./target/release/minfer serve --n-ctx 4096 --n-slots 1 qwen2.5-0.5b-instruct-q4_0
# POST /v1/chat/completions  (stream + non-stream)
# GET  /v1/models, GET /health
# GET  /metrics  (Prometheus text; see "Metrics and observability" above)

Performance testing (bench)

./target/release/minfer bench -r 3 <model>        # pp512 + tg128, markdown table
./target/release/minfer bench -p 3314 -n 128 <model>  # campaign-shape pp/tg

pp<P> ingests P prompt tokens (prefill-only, generate nothing); tg<T> prefills the context then decodes T tokens. -p 0 / -n 0 skip a test, -r sets the measured reps (1 untimed warmup each), -o csv|json emits the same fields machine-readable, --n-ctx only ever grows the auto KV sizing (P+T+16, clamped to the model's context length).

Verify-step gate bench (specverify)

./target/release/minfer specverify -p 512 -r 40 -o json <model>

Measures the batched verify-step cost C_T(nt) (nt = 1, 3, 5 by default) at a fixed deep KV depth and reports the per-token amortization nt·C_T(1)/C_T(nt) — the D5 speculative-decoding gate instrument (step doc 81). -p sets the depth, -r the timed reps (3 untimed warmups each), MINFER_SPECVERIFY_NTS=1,3,16 overrides the phase list, MINFER_SPECVERIFY_NOUT=1 forces prefill-style n_out=1. Exit code is 0 whenever the measurement completes; the PASS/FAIL verdict is in the JSON.

GGUF tooling — convert, quantize, split (F6)

# HuggingFace Qwen2 checkpoint -> GGUF (f16, or f32 for a lossless archive)
./target/release/minfer convert /path/to/Qwen2.5-0.5B-Instruct out.gguf --outtype f16
# quantize an existing single-file GGUF (q4_0, q4_1, q5_0, q5_1, q8_0, f16, f32)
./target/release/minfer quantize out.gguf out-q4_0.gguf --type q4_0
# split a single file into transport-sized parts (any type; tensor bytes copied verbatim)
./target/release/minfer split out-q4_0.gguf /tmp/parts --max-size 200M

convert reads config.json, tokenizer.json, tokenizer_config.json and model.safetensors (single file or an index.json shard map) and writes the GGUF v3 metadata the engine's strict loader requires: tokenizer.ggml.model = gpt2, tokenizer.ggml.pre = qwen2, the full 256-token byte vocabulary, the merge table, the special-token ids and tokenizer.chat_template. Supported architectures: Qwen2ForCausalLM only. 1-D tensors stay f32 under --outtype f16 (llama.cpp's rule); --outtype f32 is exact for bf16/f16 sources.

quantize re-encodes 2-D float weights; 1-D norms/biases keep their source type, and on a tied model a sub-8-bit target quantizes the shared token_embd.weight at q8_0 (both are printed). K-quants and I-quants are refused by name — minfer can read them but has no encoder that has been verified against llama.cpp.

split writes {stem}-NNNNN-of-MMMMM.gguf parts with split.no/split.count/split.tensors.count; the loader reads part 0 and merges every part into one tensor index. --split-max-size/--max-size accept bytes or a K/M/G suffix, and a tensor is never split across parts.

Full contract, supported/refused sets and verification references: GGUF-TOOLING.md.

Examples

# Local model
./target/release/minfer ~/models/qwen2-0.5b-q4_0.gguf "What is the capital of France?"

# Cached model by name (no full path needed)
./target/release/minfer qwen2.5-0.5b-instruct-q4_0 "Hello"

# Auto-download from Hugging Face + run (quant auto-detected, splits included)
./target/release/minfer hf:Qwen/Qwen2.5-0.5B-Instruct-GGUF:qwen2.5-0.5b-instruct-q4_0.gguf "Hello"

# Inspect GGUF metadata + key tensors
./target/release/minfer info qwen2.5-0.5b-instruct-q4_0

# List available GGUF files in a HF repo (without downloading)
./target/release/minfer download hf Qwen/Qwen2.5-0.5B-Instruct-GGUF

# Pull from Ollama and create a symlink
./target/release/minfer download ollama qwen2.5:0.5b

# List locally cached models
./target/release/minfer list

Sampler pipeline and invariants (moved from AGENTS.md)

sampler.rs (#48): one SamplerConfig drives one pipeline — logit bias → penalties (repeat / frequency / presence, last 64 tokens) → DRY → grammar mask (F2) → greedy shortcut (temp == 0) → top-k → typical → top-p → min-p → XTC → temperature or mirostat v1/v2, seeded StdRng. Every F3 knob defaults to a no-op, so the default path is bit-identical to the pre-#48 chain (pinned by test_default_pipeline_matches_the_pinned_pre_f3_sequence and, through the grammar-aware entry point, by test_default_pipeline_matches_the_pinned_pre_f2_sequence). SamplerConfig::validate refuses nonsensical values at CLI startup / HTTP 400 (never clamps silently), and logit-bias token ids are checked against the vocabulary. Mirostat's mu is caller-owned state (MirostatState: one per run / session / request / batch slot); speculative decoding refuses mirostat (--spec-draft), because a verify round samples several rows from one shared RNG. DRY sequence breakers are token-id sequences (--dry-sequence-breakers 198;13,2); llama.cpp's string form needs a tokenizer port (follow-up). CLI: --temp --greedy --top-k --top-p --repeat-penalty --frequency-penalty --presence-penalty --min-p --typical --xtc-probability --xtc-threshold --dry-multiplier --dry-base --dry-allowed-length --dry-penalty-last-n --dry-sequence-breakers --mirostat --mirostat-tau --mirostat-eta --mirostat-m --logit-bias --grammar --grammar-str --json-schema --json-schema-str -n --seed -t.

Constrained decoding (F2, #47). src/grammar.rs compiles a GBNF grammar (or a JSON Schema, through a generated GBNF) into one pushdown automaton — a flat program per rule, a set of {rule, pc} call stacks — and turns it into a per-state token bitset. The mask is applied inside sample_with_config_grammar, after DRY and before the greedy shortcut: every stage before it only shifts logits and every stage after it only removes candidates, so a forbidden token can never be chosen, and the mask consumes no RNG (mirostat/DRY are unperturbed). The compiled Arc<Grammar> is per request; the mutable GrammarState is per run, exactly like MirostatState. Token advancement is byte-level correct (a piece may be one byte of a multi-byte character); a token whose pending bytes can never complete to an accepted codepoint is rejected, EOG is legal only at an accepting state with no pending bytes, and no allowed token is a loud stop (SampleError::NoAllowedToken), never an arbitrary token. Unsupported GBNF or schema constructs are startup/400 refusals — never a silent guess; the accepted subset and every refusal are catalogued in docs/GRAMMAR-DESIGN.md. Server: response_format (json_object / json_schema) plus a grammar extension field; the two together are a 400. --spec-draft + a grammar is refused (a verify round samples several rows from one automaton state).

Decisions governing this document

This page is the current contract for the CLI surface; the decisions behind the flags and behaviours it documents are frozen in the ADR corpus:

  • ADR-0006 — The KV storage format is a per-engine gate, not a process-wide global
  • ADR-0014 — A KV session is a versioned, checksummed file — never a memory dump
  • ADR-0015 — The offload auto fit takes a prefix, not a knapsack
  • ADR-0017 — Speculative decoding refuses the features its identity contract cannot carry
  • ADR-0018 — The grammar mask is one stage inside the single sampler pipeline
  • ADR-0019 — A chat template that cannot be rendered refuses the load
  • ADR-0021 — bf16 is a round-to-nearest-even cast, and 1-D tensors stay f32

Support Matrix

Supported quantization formats and model architectures. This page is the expanded version of the two support sections formerly in the README.

Supported Quantization Formats

minfer supports GGUF v3 files with the following quantized weight types. The CPU backend quantizes activations on-the-fly (Q8_0 for the simple weight types, Q8_K — llama.cpp's format with precomputed per-subblock sums — for Q4_K/Q5_K/Q6_K); the GPU backends read f32 activations directly for their non-MMQ kernels, matching llama.cpp's Metal backend.

Supported

TypeBitsBlockCPUAVX2CUDA GPUMetal GPU
Q4_0418 B / 32 val✅✅✅✅
Q4_1420 B / 32 val✅❌✅✅¹
Q4_K4144 B / 256 val✅✅⁹✅✅¹
Q5_0522 B / 32 val✅✅✅✅¹
Q5_1524 B / 32 val✅❌✅✅¹
Q5_K5176 B / 256 val✅✅⁹✅✅¹
Q6_K6210 B / 256 val✅✅⁹✅✅¹
Q8_0834 B / 32 val✅✅✅✅¹
F16162 B / 1 val✅³✅³✅⁴✅⁵
BF16162 B / 1 val✅⁶—✅⁷✅⁸
F32324 B / 1 val✅—✅²✅²

¹ Metal prefill uses a simdgroup GEMM for every quant type (dispatched when nt ≥ 2 && (od ≥ 2048 || nt ≥ 9)); the scalar f32 multi kernels handle decode (nt==1) and tiny small-od batches. The compute-graph MetalBackend dispatches these kernels per op (quant_matmul_f32_on_gpu_buf), so every quant type above runs on the GPU. ² F32 weights are supported on both GPUs: 1-D norms/biases through the norm kernels and 2-D matmul weights through CUDA's launch_f32_f32_matmul and Metal's kernel_f32_f32_matmul (#317). Before #317 an f32 weight on Metal had no arm and silently ran the Q4_0 kernel. ³ F16 has no block: 2 B per element, so there is no integer dot to run. The CPU dot is vectorized — AVX2 uses F16C (_mm256_cvtph_ps) and aarch64 uses baseline NEON FCVTL (vcvt_f32_f16) — with an f64 scalar oracle/fallback (vec_ops::dot_f16_f32, f16_dot_path()), and the multi-token prefill decodes each weight row once and threads the row loop through the shared CPU pool (#141). The AVX2 column marks the hand-written x86 kernel; NEON is folded into CPU as in every other row. ⁴ CUDA decodes in-register (f16_f32_matmul_vec / _scalar, __half22float2) and the embedding gather has its own f16 kernel — the weights stay 2 B/element on the device, which is the point of the format. No MMQ route: MMQ streams quantized bytes and f16 is not one of its formats, so an f16 prefill runs the f32-activation kernel. Both supported architectures (Qwen2/Qwen2.5 and Qwen3) use it: #141 landed the registration branch and the graph type gate in the qwen2 loader/graph only, so until #167 an f16 Qwen3 model fell to the CPU on a CUDA build even though these kernels existed; the loaders now share one registration rule (models::weight_reg). The engine's f16 file contract is 2-D tensors f16 and 1-D norms/biases f32 (llama.cpp's rule; mat_mul_f16/the f16 embed decode have no f16-norm sibling) — what minfer convert --outtype f16 writes. ⁵ Metal registers the raw 2 B/element f16 weights and promotes in-register: kernel_f16_f32_matmul (src/metal/kernels/f16.metal) is the f32-activation matmul and kernel_get_rows_f16 the embedding gather, both selected by the TensorType::F16 arms of quant_matmul_f32_on_gpu_buf / embed_tokens_gpu (#164). The weights stay half width on the device — no registration-time f32 copy — and, like CUDA, an f16 prefill runs the f32-activation kernel, not a simdgroup GEMM. Both loaders admit the type, so weights_on_gpu's all-or-nothing check passes and the model is a Metal model; the file contract above means an f16 norm can never reach a d*2 kernel buffer. Measured on macbook (macOS 27.0.1, Apple M4 Pro) (2026-10-06) against the same file's CPU logits: max |Δlogit| 2.4e-3 on the 0.5B and 7.9e-3 on Qwen3-0.6B (bar 0.05), with an identical greedy continuation ([12095, 11, 323, 432] for Qwen2, [12095, 13, 576, 6722] for Qwen3). ⁶ BF16 weights (#142): the CPU decodes one row at a time (vec_ops::mat_mul_bf16, exact f32::from_bits(bits << 16), then the same vec_dot_f32 the f16 row path uses) and the embedding rows in Op::GetRows. The decode is a left shift, so there is no separate SIMD kernel to mark in the AVX2 column (the dot itself is the vectorized vec_dot_f32). minfer convert --outtype bf16 writes 2-D bf16 / 1-D f32 and is byte-identical to llama-quantize --pure <f32>.gguf … BF16 (docs/GGUF-TOOLING.md §4.1.1). ⁷ CUDA registers the raw 2 B/element bf16 words and promotes in-register — bf16_f32_matmul_vec / _scalar (the uint4 word load split by a bits << 16 shift, the f16 pair's exact sibling) and embed_rows_bf16 — selected by the TensorType::BF16 arms of matmul_f32_ptr_layout / embed_rows_on_gpu (#208; #141 is the f16 template). The decode is exact (f32::from_bits(bits << 16)), so unlike the quantized types there is no rounding at all; the weights stay half width on the device — no registration-time f32 copy — and, like f16, a bf16 prefill runs the f32-activation kernel, not the int8 MMQ GEMM (MMQ streams quantized bytes and bf16 is not one of its formats). Both loaders admit the type through the shared models::weight_reg::cuda_weight_reg rule, so weights_on_cuda's all-or-nothing check passes for both supported architectures and the graph's BF16 matmul / embed nodes are assigned Backend::CUDA; the 1-D side is the file contract above. bf16 does not fuse: cuda::concat_rows has no 2 B/element arm, so the attn_qkv / ffn_gu concat copies are not registered and the unfused matmul chain runs. Measured on a GB10 (2026-10-06, dgxspark): a 0.5B bf16 GGUF registers the device-weight figure recorded in §"f16 and bf16 weights" (the same number as its f16 twin, i.e. the 2 B/element claim is real), 169 bf16 matmul + 1 embed nodes on CUDA, device-vs-CPU max |Δlogit| 7.82e-5 absolute / 4.24e-6 relative (bar 0.01 / 1e-3) with an identical greedy continuation [12095, 13, 1084, 374]. ⁸ Metal registers the raw 2 B/element bf16 words and promotes in-register: kernel_bf16_f32_matmul (src/metal/kernels/bf16.metal) is the f32-activation matmul and kernel_get_rows_bf16 the embedding gather, both selected by the TensorType::BF16 arms of quant_matmul_f32_on_gpu_buf / embed_tokens_gpu (#208, the Metal half; the CUDA half is footnote 7). Its own kernel, not a dtype flag on the f16 one — bf16 and f16 are different 2 B/element layouts, so a shared kernel would branch per element in the hottest device kernel. The weights stay half width on the device — no registration-time f32 copy — and, like CUDA/f16, a bf16 prefill runs the f32-activation kernel, not a simdgroup GEMM. Both loaders' Metal arm (matches!(ttype, F32 | F16 | BF16)) admits the type, so weights_on_gpu's all-or-nothing check passes and the model is a Metal model; the file contract above covers the 1-D side. bf16 does not fuse (the fused device forms are CUDA-only). Measured on a Mac (2026-10-06, macbook (macOS 27.0.1, Apple M4 Pro)) against the same file's CPU logits: 169 bf16 matmul + 1 embed nodes all on Backend::METAL, the same device-weight figure, max |Δlogit| 1.889e-3 absolute / 1.025e-4 relative (bar 0.05 / 5e-3), with an identical greedy continuation [12095, 13, 1084, 374].

⁹ The K-quant dots (#56, 2026-10-08) have AVX2+FMA kernels (src/quants/avx2.rs) and AVX-512/VNNI variants (src/quants/avx512.rs), dispatched AVX-512 → AVX2 → scalar at runtime (MINFER_NO_AVX512=1 drops to AVX2, MINFER_NO_AVX2=1 to scalar — the x86 counterparts of MINFER_NO_NEON) and gated bitwise against the scalar reference (quants::avx2_correctness). The AVX2 column marks the hand-written x86 kernel; NEON is folded into CPU as in every other row.

CUDA notes: prefill (nt ≥ 16) runs the default int8 tensor-core MMQ path for the common quants (Q4_0/Q4_1/Q5_0/Q5_1/Q8_0/Q4_K via the f16-wmma GEMM, Q4_K/Q6_K via the raw-nibble int8 kernels — the promoted ~3581 tok/s path, see CUDA_OPTIMIZATION.md); the f32-activation kernels cover every type including Q5_1/Q5_K, and decode (nt==1) uses the dp4a MMVQ kernels (Q4_K/Q5_K/Q6_K with shape gates). Q5_K requires id % 32 == 0 (tail-masking granularity).

GPU grouping note: the old whole-layer layer_gpu path required all 7 weight matrices in a layer to share one quant group (all-Q4 or all-QK) and fell back to CPU otherwise. The compute-graph path (default) has no such restriction — backend assignment is per op, so mixed-group layers run fully on the GPU. Raw weights are not supported on GPU and select the CPU backend for those ops.

KV Cache Storage Type by Backend

MINFER_CACHE_TYPE picks the KV cache element type, which is a separate axis from the weight type above (graph/kvformat.rs is the single authority, and the answer for "can this backend read it" is the registry's reads_packed_kv). A value the backend has no kernel for is refused at load, never silently mapped to f32.

MINFER_CACHE_TYPECellCPUCUDAMetal
f32 (default)4 B/element, f32✅✅✅
f162 B/element in the f32-shaped region→ f32✅✅
q8_0packed Q8_0 blocks, 34 B per 32 elements, cell padded to whole f32 words✅ (C4 S1+S2)✅ (C4 S2b)✅ (C4 S2b, #310 enabled)

Notes:

  • f16 on the CPU resolves to f32 — the CPU has no f16 KV kernel, and an env var set for a GPU run must not break a CPU one.
  • The default on CUDA/Metal is the model's own auto policy (f16 when n_layers × n_kv_embd ≥ 8192, i.e. the 7B class, f32 for small models); the table's "default" row is the region shape, which f16 does not change.
  • Q8_0 is the packed one: 3.76× smaller than f32 and 1.88× smaller than the f16 auto policy. On CUDA a Q8_0 decode runs the layout-tagged split-K kernel together with the packed fused QKV epilogue (attn_bias_rope_store_q8_0, #144), and a prefill at head dim 128 runs the packed FA prefill — its f16-tile staging dequantizes each packed block, so the tensor-core route is offered for a packed cell too (#144: Qwen3-0.6B pp2048 564.5 → 8231.1 tok/s). Still off their tuned route, and stated: the verify band (1 < nt ≤ 16) takes the general layout-tagged kernel and the hybrid 4-warp decode dispatch is f16-typed. The dp4a packed K dot is not a follow-up — it landed in #186 (cc19b4f, 2026-09-27) as the packed decode route, and docs/ARCHITECTURE-EXECUTION-PLAN.md §C4 #186 records it DONE with its tolerance class re-measured. The general layout-tagged kernel remains the fallback for every packed path. A speculative session refuses a packed cache outright (its greedy identity contract rests on the batched split kernel). See docs/ARCHITECTURE-EXECUTION-PLAN.md §5 C4 #144 and docs/cuda_optimization_steps/107-c4-packed-q8-kv-cuda.md.
  • Metal (#310) — enabled. kernel_store_kv_q8_0 writes the same bytes as the CPU quantizer. Two read mechanisms cover every attention shape, selected by the pure crate::metal::packed_attn_route: mechanism A reads packed cells natively in the decode flash family (kernel_flash_attn_ext_q8_0 / _hd128_q8_0, nt == 1, hd ∈ {64,128}); mechanism B dequantizes the needed window into a transient f32 stage (kernel_dequant_kv_q8_0_to_f32) and runs the unchanged f32 prefill / windowed-flash family. The classic kernel_gqa_attn_q8_0 / _window_q8_0 / _map_q8_0 remain the fallback for a small/odd hd, an nt == 1 explicit window and any MINFER_NO_* opt-out. READS_PACKED_KV is now true, so MINFER_CACHE_TYPE=q8_0 loads and runs on Metal. The memory win is 3.76×; the measured speed (macbook (macOS 27.0.1, Apple M4 Pro), 2026-10-08, minfer bench -p 1024 -n 64 -r 3 --n-ctx 2048, 3 interleaved runs, medians; f16 baseline): Qwen3-0.6B pp1024 4844 → 4683 tok/s (0.967×) and tg64 193.5 → 176.6 (0.913×), Qwen2.5-0.5B pp1024 6175 → 6154 (0.997×) and tg64 295.7 → 245.5 (0.830×); region sizes 58 720 256 → 15 597 568 B and 6 291 456 → 1 671 168 B (both 3.76×). The stage is f32, not f16: f16 staging's second rounding was measured to amplify to 16.9 logit delta on Qwen3-0.6B over 8 decode steps, outside the inherited C4 class; f32 staging restores it (1.28 / 0.36) at parity speed. docs/METAL-BACKEND-DESIGN.md §4.4 records the mechanisms, the measurements and the gates.
  • A Q8_0 cell width must be a whole number of 32-element blocks (so n_kv_embd % 32 == 0, which every supported architecture satisfies); ensure_kv refuses anything else.
  • The KV write/move side (Backend::copy_cells for C3 compaction / C8a prefix copy / C8b S3 copy-on-write, and GraphAllocator::copy_kv_to_cpu for the C2 shift and C5 sessions) is implemented on all three backends since #44 part (b): Metal moves rows one at a time with MTLBlitCommandEncoder in the overlap-safe order and reads its regions back through the registry host_read hook. A physical shift of an f16 region refuses loudly and is pinned by a gate (#306: the host round trip has no dequantize → re-rope → requantize map, so the CLI re-renders the retained window instead); the per-engine kv_format is what makes a Metal session describe the width its region really uses.

Not Yet Supported

CategoryTypes
K-quantsQ2_K, Q3_K, Q8_K
I-quantsIQ1_S, IQ1_M, IQ2_XXS, IQ2_XS, IQ2_S, IQ3_XXS, IQ3_S, IQ4_NL, IQ4_XS
OtherQ1_0, TQ1_0, TQ2_0, MXFP4, NVFP4

Q5_K and Q5_1 are fully supported on CPU and both GPU backends — Q5_K_M models run at full GPU speed.

f16 and bf16 weights — the file contract and the measured device deltas

Moved here from AGENTS.md (its Support section keeps the one-line summary).

f16 weights (F6/#49; device half by #141 on CUDA, #164 on Metal): an f16 GGUF runs on CPU, CUDA and Metal for both supported architectures (Qwen2/Qwen2.5 and Qwen3). Both device loaders register the raw 2 B/element weights and the kernels promote in-register; the CUDA side shares one registration rule with the CPU (models::weight_reg, which also carries the q4_K W_dsc plane gate of #165), while the Metal side is the loader's matches!(ttype, F32 | F16 | BF16) registration arm (the BF16 arm is #208's Metal half). Op::MatMul decodes one f16 weight row at a time (vec_ops::mat_mul_f16) and Op::GetRows decodes f16 embedding rows; 1-D norms/biases stay f32 — the file contract every producer writes (minfer convert --outtype f16, llama-quantize … F16, minfer quantize --type f16), and the CUDA norm path checks the registered byte length so an f16 norm cannot reach a d*2 buffer (#169). CUDA registers the raw 2 B/element weights and converts in-register (f16_f32_matmul_vec / _scalar, embed_rows_f16); an f16 prefill does not enter the int8 MMQ GEMM — it runs the f32-activation kernel. Metal is the same shape: kernel_f16_f32_matmul + kernel_get_rows_f16 (src/metal/kernels/f16.metal) are its f32-activation matmul/embed, selected by the TensorType::F16 arms of quant_matmul_f32_on_gpu_buf / embed_tokens_gpu, and its prefill runs that same f32-activation kernel (no simdgroup GEMM). minfer quantize supports q4_0/q4_1/q5_0/q5_1/q8_0 and the K-quants q4_K/q5_K/q6_K (byte-identical to llama-quantize; the K-quant reference is llama-quantize --pure, and --type q4_K writes one uniform type, not the Q4_K_M mixture — #203), f16 and f32, and refuses every type without an encoder by name (#140); a K-quant row that is not a multiple of 256 is demoted the way llama.cpp's tensor_type_fallback does it; --type f16 keeps 1-D tensors f32, --type f32 writes every tensor f32. The gate inputs are recorded: tests/fixtures/f6-fixtures.json carries one entry per cached artifact content (path, bytes, sha256, the exact producer command, the producer identity — minfer commit, or llama.cpp commit + compiler + effective -ffp-contract — date and an absolute box label), scripts/check_f6_fixtures.py audits that manifest in CI, verifies the whole cache where it exists, and — since #345 — re-runs a recorded producer with --regenerate and re-records the content identity it produces (idempotently; a llama-quantize reference, or a minfer entry whose recorded commit is not the one that runs, is refused by name rather than re-recorded), and every F6 gate verifies the fixture it resolves (src/tooling/tests/f6_fixtures.rs) so a stale or replaced reference refuses the run by name and digest instead of being compared against — #205.

bf16 weights (#142): minfer convert --outtype bf16 writes 2-D bf16 / 1-D f32, f32→bf16 being round-to-nearest-even (ggml_compute_fp32_to_bf16, NaN forced quiet; general.file_type = 32 = LLAMA_FTYPE_MOSTLY_BF16). The CPU decodes one bf16 weight row at a time (vec_ops::mat_mul_bf16, exact f32::from_bits(bits << 16), then the same vec_dot_f32 the f16 row path uses) and Op::GetRows decodes bf16 embedding rows; TensorType::BF16 is the new type the loader maps GgmlType::BF16 to. CUDA registers bf16 since #208's CUDA half — models::weight_reg::cuda_weight_reg answers Raw for it (the f16 arm's sibling: 2 B/element, no f32 copy, clear_nb_bt_only), Op::MatMul takes the TensorType::BF16 arm of matmul_f32_ptr_layout (bf16_f32_matmul_vec when id % 8 == 0, else _scalar; both shift bits << 16 in-register, which is exact) and Op::GetRows the BF16 arm of embed_rows_on_gpu (embed_rows_bf16), and both loaders' CUDA branch admits the type through that one shared rule, so the all-or-nothing gate answers Cuda for both supported architectures. bf16 is not an MMQ format, so its prefill runs the f32-activation kernel; cuda::concat_rows has no 2 B/element arm, so the attn_qkv/ffn_gu concat copies are not registered and bf16 runs unfused. Measured on GB10 2026-10-06 (dgxspark, scripts/cuda_test.sh device gate f208_bf16_weights_run_on_the_cuda_device): 0.5B bf16 registers 942.4 MiB of device weights (the f16 twin's number), 169 bf16 matmul + 1 embed nodes all assigned Backend::CUDA, device-vs-CPU max |Δlogit| 7.82e-5 absolute / 4.24e-6 relative (bar 0.01 / 1e-3), greedy [12095, 13, 1084, 374] identical. Metal registers it too since #208's Metal half — kernel_bf16_f32_matmul + kernel_get_rows_bf16 (src/metal/kernels/bf16.metal, the pl_bf16_f32 / pl_get_rows_bf16 pipelines) are its f32-activation matmul/embed, selected by the TensorType::BF16 arms of quant_matmul_f32_on_gpu_buf / embed_tokens_gpu, and both loaders' Metal arm is now matches!(ttype, F32 | F16 | BF16), so weights_on_gpu passes and both supported architectures answer Device::Metal. Measured on a Mac (macbook (macOS 27.0.1, Apple M4 Pro), 2026-10-06, device gate f208_bf16_weights_run_on_the_metal_device): the 0.5B bf16 file's 169 bf16 matmul + 1 embed nodes are all assigned Backend::METAL, 942.4 MiB of device weights, device-vs-CPU max |Δlogit| 1.889e-3 absolute / 1.025e-4 relative (bar 0.05 / 5e-3), greedy [12095, 13, 1084, 374] identical. The writer is byte-identical to llama-quantize --pure <f32>.gguf … BF16, 290/290 tensor payloads (169 BF16 2-D, 121 F32 1-D; docs §4.1.1 — the reference is cast from the f32 conversion, because bf16→f16 is lossy below 2^-14 and the f16 file cannot carry those values). On the bf16-source 0.5B, bf16-vs-f16 CPU logits differ by max 2.29e-5 / 1.24e-6 relative (not bitwise — the f16 file's 123 024 subnormal weight values are the whole difference) with an identical greedy continuation [12095, 13, 1084, 374].

Operator Coverage by Backend

supports_op decides at graph-build time which backend runs each node (docs/ARCHITECTURE.md §5). This table is the contract, generated from the three implementations — keep it in step with them.

The table's concrete F32 Op rows are pinned by graph::op_matrix::support_table_matches_support_matrix_doc, which fails when a backend's supports_op disagrees (each backend column is checked wherever it is compiled in). Two kinds of row it cannot check, so they are prose plus their own gates: the composite rows (one line spelling several ops) and the capability nuances that are not an Op field — a partial View at offset 0 (the allocator backstops it; supports_op sees only the offset), the set-valued kv_map attention window (Device::gathers_attn_map), and a packed q8_0 KV region (BackendCaps::reads_packed_kv).

OperatorCPUMetalCUDA
Input, KvcacheLoad, View/Reshape/Permute✅✅✅
Add, Mul, Silu✅✅✅
RmsNorm, QkNorm✅✅✅
MatMul✅✅✅
GetRows (embedding, tail rows)✅✅✅
View with offset != 0 or a partial window (D1)✅❌✅ — Metal's kernels take a buffer and a length with no element offset, so it can express exact views only: a standing design limit, not a pending port (G5 landed the attention window and the KV cell store, not offset views — Op::View { offset, .. } => *offset == 0); it is what keeps the hand-written Op::FusedFFN on Metal (§D3). The allocator backstops the partial case, which supports_op cannot see
Attn✅✅✅
Attn with explicit_span (a window that starts at a non-zero cell, or several sequences in one batch)✅✅✅ — the one-range attn_span window is read on all three backends (Metal's kernel_gqa_attn_window_f32/_f16 landed in #44 part (a), and #44 part (b) gave Metal the matching write/move side so a batched and compacted multi-sequence run serves; CUDA's E1b instantiation is device-verified on GB10, including a bitwise batch-order-invariance gate). The set-valued kv_map window is read on all three backends too: CPU/CUDA always did, and Metal's sibling kernel_gqa_attn_map_f32/_f16 landed in #362 (Device::gathers_attn_map is now true for Metal). A packed q8_0 KV cache is read on Metal too since #310 (mechanisms A and B), alongside CPU and CUDA
KvcacheStore✅✅✅
SwiGLU (fused)✅✅✅
RoPE non-interleaved✅✅✅
RoPE interleaved✅✅❌
FusedQKV (decode)❌✅✅
FusedFFN (decode)❌✅✅
FusedQkvNorm (Qwen3 decode)❌✅❌
QkvBiasRopeStore (mixed-quant decode)❌❌✅
Scale, Softmax, BatchMatMul✅ / ✅ / ❌❌ / ❌ / ❌❌ / ❌ / ❌

Notes on the asymmetries — these are the rows where a model's decode path differs by platform:

  • FusedQkvNorm is Metal-only. Qwen3 decode on CUDA takes the unfused QkNorm path, which is numerically equivalent but issues more dispatches. Making CUDA fused is a Phase G / CUDA-verifiable ticket, not a correctness gap.
  • QkvBiasRopeStore is CUDA-only — a recorded decision (#52), not a gap. It is the mixed-quant decode epilogue: q/k/v use different quant types (so they cannot share FusedQKV's single concat weight), so three separate matmuls (no bias) feed one bias×3 + RoPE×2 + store×2 pass. CUDA fuses it (10 dispatches → 4 per layer, −6); Metal keeps the unfused chain and the graph builder never emits the node there (metal_backend.rs's false arm is a design statement). Porting would save 6 dispatches per mixed-quant layer — 84 per decode token on Qwen2.5-7B-Q4_K_M, the realistic case, whose 14 of 28 layers carry attn_v as Q6_K against Q4_K q/k — with no numerical difference (supports_op is a build-time gate and the unfused chain is the reference). The whole forward's host-encode is ~0.2 ms against a ~20 ms/token decode, so those 84 dispatches are a sub-1% slice of decode time; the A/B of the concat-class fusion that removes more dispatches (MINFER_NO_FUSE_QKV=1, −8 on the same 14 layers) sits within run-to-run variance on macbook (macOS 27.0.1, Apple M4 Pro) (2026-10-06, five interleaved bench -p 0 -n 128 -r 4 pairs: 48.06 vs 46.49 t/s means, individual pairs crossing zero), so a second kernel path and its bitwise gate are not earned by a ~1% ceiling. Provenance and the rejected alternative, stated so the numbers are not misread. The ~0.2 ms host-encode figure is MINFER_OP_PROFILE=1 on that 7B — per-op host-encode totals, not per-label counts — and the 10 → 4 / 84-per-token dispatch counts are the CUDA D3-8 ledger applied to Metal's dispatch table; Metal has no per-op profiler, so they were not re-counted on the device. And the port is cheap: Metal already has the class-1 attn_bias_rope_store kernel, so the refused work is mainly a three-pointer binding — the decision rests on the measured ceiling, not on the size of the change.
  • Interleaved RoPE is CPU-only. Both loaders hard-code NonInterleaved today, so no shipped model hits this; a family that needs interleaved RoPE needs a loader change plus a CUDA kernel.
  • Scale/Softmax are CPU-only and unused. Attention kernels fuse the softmax and carry the scale in AttnMeta, so no supported architecture emits either node.
  • BatchMatMul is deferred everywhere (single-output IR) and nothing emits it.

Supported Model Architectures

minfer currently supports two model architectures.

ArchitectureVariantsStatusDetection Key
Qwen2Qwen2, Qwen2.5, DeepSeek-R1-Distill-Qwen✅ Fully supportedgeneral.architecture = "qwen2"
Qwen3Qwen3 (dense: 0.6B–32B)✅ Fully supported (CPU + GPU)general.architecture = "qwen3"

Qwen3 support: dense architecture only (no MoE / hybrid-SWA / VL variants yet). The dense models reuse the Qwen2 graph with two deltas — the head dim is read from qwen3.attention.key_length (decoupled from n_embd / n_head) and Q/K go through a per-head RMSNorm (blk.{i}.attn_q_norm / attn_k_norm) before RoPE. See QWEN3-SUPPORT-PLAN.md for the design + verification record.

How Architecture Detection Works

minfer reads the general.architecture string from the GGUF metadata header. Only the exact values "qwen2" and "qwen3" (case-sensitive) are accepted. Any other value produces a clear error:

Unsupported architecture: 'llama'

The loader will not silently misinterpret a non-Qwen2 model — it fails immediately with a descriptive message. All model-agnostic components (BPE tokenizer, Jinja2 chat template renderer, samplers) are ready for additional architectures once the graph construction (build_graph) is added.

Hyperparameter Keys

The Qwen2 loader reads GGUF keys from both qwen2.* and llama.* prefixes. The llama.* fallback exists for compatibility with older GGUF converters that used the llama. prefix as a de-facto standard for Llama-family hyperparameters. This does not mean Llama architecture is supported.

Adding a New Architecture

See AGENTS.md for a step-by-step guide. In brief:

  1. Create src/models/<name>/ with mod.rs, graph.rs, loader.rs
  2. Add a match branch in src/models/mod.rs::load_model()
  3. Define HParams, LayerWeights, and implement the ModelDef trait (including build_graph(&self, params) -> ComputeGraph, which is deterministic in params — the graph-reuse invariant)

Architectures that share Qwen2's tensor naming convention (LLaMA, Mistral, Phi) should be relatively straightforward to port.

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0013 — The CPU quantizes activations to Q8_0; a device reads f32
  • ADR-0021 — bf16 is a round-to-nearest-even cast, and 1-D tensors stay f32
  • ADR-0005 — Metal becomes a first-class backend
  • ADR-0006 — The KV storage format is a per-engine gate, not a process-wide global
  • ADR-0020 — A quantized file is byte-identical to llama-quantize, or it is wrong
  • ADR-0032 — The packed-KV staging window is f32, not f16

minfer Architecture

A pure-Rust LLM inference engine written from scratch, inspired by llama.cpp, with zero ML framework dependencies. Inference runs through a declarative compute graph (builder → scheduler → per-backend kernels), modeled on llama.cpp's ggml_cgraph + backend scheduler. This document describes the overall design: module responsibilities, the compute graph pipeline, the CPU / Metal backend layering, quantization, adding a new model architecture, and adding a new backend. The pre-graph imperative forward is preserved at the end as an appendix (Appendix A).


1. Design Principles

  1. No ML framework — attention, RMSNorm, RoPE, SiLU, softmax are all handwritten. Only 5 external crates: rand, regex, half, serde/serde_json, minijinja.
  2. Declarative compute graph, not an imperative loop — the forward pass is built as a pure ComputeGraph (no side effects at build time), then assigned to backends, fused, allocated, and executed by the scheduler. This mirrors llama.cpp (ggml_cgraph + ggml_backend_sched) and enables graph reuse, per-op backend assignment, DOT export, and a clean path to new backends. Design and implementation record: docs/COMPUTE-GRAPH-DESIGN.md.
  3. Bytes-in / bytes-out tensors — weight tensors are raw &[u8]; SIMD dot-product kernels (AVX2 / Metal shaders) operate on byte slices matching the exact GGML quantized block layout (repr(C) in block.rs).
  4. Activations stay f32 — CPU matmuls quantize activations to Q8_0 on-the-fly; the Metal backend reads f32 activations directly for all weight types (matching llama.cpp's Metal backend).
  5. Backend assignment is a build-time decision, never a silent mid-execution fallback — supports_op decides which ops run where; cross-backend transfers happen at split boundaries; kernel-invariant violations abort (gpu_abort / Err), they never silently fall back to CPU.

2. Module Map

ModuleResponsibility
main.rsCLI, GGUF load, chat template, prefill → autoregressive generation loop, timing
graph/Compute graph core — IR, builder, scheduler, backends, reuse cache (see §4 table)
gguf.rsGGUF v3 parser (metadata KV + tensor table + data blob), multi-part (split) support, ggml_pad alignment
block.rs20+ quantized block types as repr(C) structs + fp16 conversions, matching ggml-common.h
quants.rs + src/quants/*.rsAVX2+FMA / NEON+SDOT dot-product kernels (Q4_0/Q4_1/Q5_0/Q5_1/Q8_0/K-quants × Q8_0/Q8_K) + f32→Q8_0/Q8_K quantization, scalar fallback. The module file is the decider; the parts are src/quants/{dot_q4_0,dot_q4_1,dot_q5,dot_q8_0,kquant,quantize_q8_0,quantize_q8_k,avx2,neon}.rs (#264, the source layout plan)
kernel.rs + src/kernel/*.rsQuantized matmul dispatch (Q4_0/Q4_1/Q5_0/Q5_1/Q4_K/Q5_K/Q6_K/Q8_0) over activations, CPU scalar fallback, the shared worker pool, and the shared embed_tokens row getter
vec_ops.rs + src/vec_ops/*.rsSIMD vector ops: RMSNorm, RoPE (Qwen2/Llama styles), softmax, SiLU, add/scale/mul, plus the f16/bf16 weight-row dots
tensor.rs4D Tensor (type/shape/strides/Vec<u8> data), ggml-compatible strides & byte sizing
sampler.rsRepeat/frequency/presence penalties → top-k → top-p → temperature, seeded StdRng
tokenizer.rsSelf-contained BPE tokenizer, loaded from GGUF metadata (no tiktoken)
template.rsGGUF chat_template rendering via minijinja + a Python-str-method hook (F7/#50), so Qwen3's think-block template renders; a template that cannot be rendered is a loud refusal, never a silent ChatML fallback (docs/CHAT-TEMPLATE-AND-TOKENIZER-DESIGN.md)
models/Architecture implementations. mod.rs has the ModelDef trait + factory dispatch
models/qwen2/Qwen2/Qwen2.5: mod.rs (model struct + trait impl), graph.rs (Qwen2Graph::build/forward), loader.rs (GGUF weights + hparams)
models/qwen3/Qwen3 (dense): same triple — decoupled head dim + per-head Q/K norm (qk_norm); its ChatML+<think> template renders since F7
conversation.rsMulti-turn session state (append-only KV) behind --cnv
server/OpenAI-compatible HTTP server (serve): axum + tokio, multi-slot, /v1/chat/completions streaming; viz.rs serves the viz page
src/metal/ + src/metal/kernels/*.metalApple MPS (Metal) backend: L1 runtime (MpsState/MetalDevice), L2 command-buffer encoding, L3 shaders (the legacy whole-layer layer_gpu is retained for tests). The split into src/metal/{runtime,encode,ops,policy}.rs + src/metal/kernels/ is the source layout plan (#265)
cuda.rs + src/cuda/kernels/*.cuNVIDIA CUDA device layer + launch layer + kernels (feature-gated --features cuda); executed through the graph via graph/cuda_backend.rs. The split is the source layout plan (#262, #263)
device_tier.rscc-keyed device tier table + selector (measured GB10 row, llama.cpp-adopted consumer rows, GENERIC fallback); resolved once at init, feeds the MMQ gate, smem feasibility and plane-VRAM budget checks (docs 105–106)
download/Hugging Face Hub + Ollama download, cached-name resolution, resume support
dump.rsPer-layer hidden-state debug dump (gated by --features debug_dump)
bench.rs / live.rs / trace.rs / spec_verify.rsllama-bench-style bench, live viz inference host, per-node trace export (MINFER_TRACE), D5 verify-step gate bench

The backend layers — runtime, launch, kernels, executors

Each device backend is organised along four layers, and only the last one is polymorphic:

layerCUDAMetalCPU
L1 device/runtimesrc/cuda/methods/{init,weights,stream,buffers,copy,events,capture}.rssrc/metal/runtime.rs— (std threads)
L2 launch/dispatchsrc/cuda/methods/<family>.rs + the extern "C" declarations each family owns (under src/cuda/methods.rs)src/metal/encode.rs + ops.rssrc/kernel/*.rs
L3 kernel sourcessrc/cuda/kernels/*.cu + common.cuhsrc/metal/kernels/*.metalsrc/quants/*.rs, src/vec_ops/*.rs
L4 graph executorsrc/graph/cuda_backend.rssrc/graph/metal_backend.rssrc/graph/cpu_backend.rs

(The CUDA paths in that table have landed (#262, #263); the Metal ones are still the target of Step 4 (#265, Mac-local). docs/SOURCE-LAYOUT-PLAN.md §8 has the file-by-file tree and each step updates this table as it lands.)

Two rules follow, and they are why the crate keeps one flat interface instead of a directory per layer:

  1. Device is the first axis, the layer is the second. Backend + registry.rs (src/graph/backend.rs) are the one device seam — llama.cpp keeps ggml-backend*.cpp/h flat beside its per-device directories for the same reason. L2 cannot leave L1: the CUDA launchers are inherent methods of CudaState, the Metal ones methods of MpsCommandBuffer; splitting them out would be a type refactor, not a file move.
  2. A common needs a second real implementation. A shared abstraction is added only when at least two backends implement it and at least two callers use it with the same semantics. The one candidate today is allocplan::DeviceMemory (CUDA answers it; the Metal half is #53). CPU's quants.rs/vec_ops.rs are deliberately not a device-private layer: they are the crate's numeric kernel library, consumed by graph/kvformat.rs and graph/cuda_backend.rs.

src/graph/ — the compute graph core

FileRole
mod.rsComputeGraph (topo-validated node list + inputs/outputs), CNode, DType, BufRef; re-exports the Backend handle from registry.rs
registry.rsThe backend registry (F4): the fixed id space (Backend::CPU/METAL/CUDA), the name surface and its three startup refusals, each entry's priority + capability record + pool/host-I/O hooks, and the --backend / MINFER_BACKENDS filter. Contract: docs/BACKEND-REGISTRY-DESIGN.md
ops.rsOp enum (full payload PartialEq), NodeMeta, AttnMode, FusedOp
builder.rsGraphBuilder — declarative construction (embedding/rms_norm/matmul/rope/attn/kvcache/…)
alloc.rsPer-backend liveness allocator + persistent per-layer KV regions + KvProvider
backend.rsBackend trait + KvProvider
cpu_backend.rsCPU execution (wraps kernel.rs + vec_ops.rs)
metal_backend.rsMetal execution (per-op MPS kernels; cfg(target_os = "macos"))
cuda_backend.rsCUDA execution (feature-gated): int8 MMQ prefill + MMVQ decode (default-on), split-KV attention, CUDA Graph capture/replay
scheduler.rsassign → split → execute (+ cross-backend copies at split boundaries)
fusion.rsPattern-matching fusion (SwiGLU), gated by backend supports_fused
cache.rsGraphCache — params-only deterministic graph reuse
params.rsGraphParams/CParams/GraphType — the reuse identity
dot.rsGraphviz DOT export (--dump-graph)
json.rsGraph JSON export for the viz page (--dump-graph-json, MINFER_TRACE)

3. Inference Pipeline

The top-level flow lives in main.rs. The whole engine is a single-pass prefill followed by an autoregressive decode loop; both call ModelDef::forward, which routes through the compute graph. Every forward is build → assign → fuse → alloc → execute (the subgraph below), but the built graph is cached per GraphParams and reused — only the first call with new params pays the build/assign/alloc cost.

flowchart TD
    A["CLI args: model, prompt, flags"] --> B{"resolve model<br/>download::resolve"}
    B -->|local path| C["load GGUF v3<br/>single or split parts (mmap)"]
    B -->|"hf:… / ollama:…"| D["auto-download → path"]
    B -->|cached name| C
    C --> E["parse metadata KV + tensor table"]
    E --> F["init GPU backend<br/>MPS / CUDA"]
    F --> G["load model<br/>dispatch on general.architecture"]
    G --> MODE{"mode?"}
    MODE -->|"--cnv"| CV["conversation REPL<br/>(run_conversation)"]
    MODE -->|"serve"| SV["OpenAI-compatible HTTP server<br/>multi-slot, streaming"]
    MODE -->|"viz"| VZ["viz server (page + live SSE)"]
    MODE -->|single shot| H["load BPE tokenizer from GGUF"]
    H --> J{"no-template?"}
    J -->|no| K["render chat template<br/>GGUF chat_template via minijinja<br/>(unrenderable → loud refusal;<br/>no template at all → ChatML)"]
    J -->|yes| L["raw prompt"]
    K --> M["tokenize prompt<br/>n_ctx = max(--n-ctx, prompt len)<br/>(sizes the KV regions once)"]
    L --> M
    M --> N["PREFILL<br/>graph forward, all prompt tokens at once"]
    N --> O["last-token logits (n_out = 1)"]
    O --> P{"DECODE loop<br/>while generated < n_predict"}
    P --> Q["sample next token<br/>penalties → top-k → top-p → temp"]
    Q --> R{"stop token or<br/>stop string match?"}
    R -->|yes| S["done"]
    R -->|no| T["append token, decode+print"]
    T --> U["graph forward, single token<br/>KV persists in the allocator"]
    U --> P

    subgraph GRAPH["every forward: build → assign → fuse → alloc → execute"]
        G1["GraphBuilder<br/>build_graph (pure IR)"] --> G2["assign backends<br/>priority Metal → CUDA → CPU"]
        G2 --> G3["fuse<br/>SwiGLU + decode fusions (gated)"]
        G3 --> G4["alloc<br/>liveness + persistent KV"]
        G4 --> G5["execute<br/>per split, cross-backend copies"]
    end
    N -.->|"GraphCache: params-only reuse"| GRAPH
    U -.->|"GraphCache: params-only reuse"| GRAPH

Generation parameters (defaults match llama.cpp): temp=0.8, top_k=40, top_p=0.95, repeat_penalty=1.1 (last 64 tokens), frequency_penalty=0.0, presence_penalty=0.0, seed=42, n_ctx=4096, n_predict=512. Sampling applies the three penalties in one pass, then top-k → top-p → temperature.

Timing is dual-caliber: Prefill: = prompt tokens / prefill wall time; Generated: = generated tokens / decode wall time (pure decode, matches llama-bench "Generation" caliber); Total: = blended.


4. Compute Graph Architecture

Inference = build a ComputeGraph (pure, side-effect free) → assign backends → fuse → allocate → execute. The graph is built once per distinct GraphParams and reused (decode steps reuse the same graph; the model's forward() routes through Qwen2Graph::forward).

4.1 The IR

  • ComputeGraph = topologically ordered nodes (builder appends sources before consumers), inputs, outputs, uid (for CUDA-Graph-style reuse).
  • CNode = op + src dependencies + out_shape/out_dtype + backend (None until assigned) + meta (weight names, rope/attn params).
  • Op carries full payloads (RmsNorm{eps}, MatMul{transpose_b}, KvcacheStore{layer} …) so graphs are structurally comparable.

4.2 Builder (GraphBuilder)

Per-architecture code calls builder methods (mirroring llama.cpp's llm_graph_context): embedding, rms_norm, qk_norm (Qwen3 per-head Q/K norm), matmul, get_rows, rope, silu, add, mul, swiglu, softmax, attn, kvcache_store/kvcache_load, plus the decode-fusion constructors fused_qkv/qkv_bias_rope_store/fused_ffn. Building is pure — no computation happens at build time.

4.3 Scheduler pipeline

assign_backends → fuse → alloc_graph → execute
  1. assign_backends — capability-driven: each node gets the highest-priority backend whose supports_op returns true (priority Metal → CUDA → CPU). Weight registration decides GPU feasibility.
  2. fuse — pattern matching (Mul(Silu(X),Y) → SwiGLU) gated per backend by supports_fused (no double-fusion with hand-written kernels). The plan's second rule (RoPE(Add(X,B)) → FusedBiasRope) was removed: no backend ever claimed the capability, and the bias+rope work ships as the builder's decode nodes (FusedQKV/QkvBiasRopeStore, whose attn_bias_rope_store kernel subsumes the bias+rope part). BatchMatMul stays deferred (single-output IR limitation, COMPUTE-GRAPH-DESIGN.md §5.4).
  3. alloc_graph — per-backend liveness allocator: buffers shared between nodes whose live ranges don't overlap; persistent per-layer KV regions survive rebuilds; in-place ops alias their input buffer (see §4.5).
  4. execute — per split (contiguous same-backend runs): sync the previous backend, copy split inputs across backends (copy_across, a host round trip via shared memory), run the nodes, then a final sync. Metal batches one MpsCommandBuffer per split, submitted at synchronize().

4.4 Reuse (GraphCache)

Params-only deterministic reuse (llama.cpp allow_reuse invariant): GraphParams = n_tokens / n_out (tail rows) / gtype / cparams (n_ctx, flash_attn, gpu, fuse_qkv, fuse_ffn, explicit_span) / weights_version deterministically determines the topology — equal params ⇒ identical graph. n_past is deliberately absent (it is execution data). CParams.gpu records backend participation so a backend toggle forces a rebuild. GraphCache owns the allocator (so the KV regions persist across rebuilds, e.g. the prefill→decode transition) and try_reuse compares params only; debug builds assert structural consistency (Op: PartialEq).

4.5 Core invariants (must not be violated)

  1. KV positions are data, not structure. KvcacheStore/Load carry only the layer index; write positions come from the positions input node. Topology never depends on n_past.
  2. Each layer owns TWO persistent KV regions (K and V), resolved by kv_pair(layer) (KvProvider). The store node's output buffer is the K region; backends write the V sibling via kv_pair.
  3. GGUF weight layout: metadata [in, out] (ne[0] fastest), memory [out][in] row-major → matmul od = shape[1], id = shape[0]. Activations: shape metadata [d, nt, 1, 1], memory token-major [nt][d]. I32 inputs are stored as f32::from_bits bit patterns (fill_input_i32).
  4. In-place ops (Silu, RoPE) alias their input buffer — the allocator maps the output to the input's BufRef (only when the input's sole consumer is this op AND it is on the same backend). Never host-copy a GPU-pending buffer: a host copy_in of a producer that is encoded but not submitted reads stale data (the Phase-3 KV-corruption bug). Cross-backend in-place inputs get a fresh buffer (the producer completed before the split boundary, so the copy is safe there).
  5. Execution follows build order (a valid topological order by construction) — guarantees a KV store executes before the attention that reads it. Nodes with no allocated buffer (dead, e.g. fusion orphans) are skipped.
  6. CPU vs GPU activation paths differ numerically: CPU matmuls are Q8_0×Q8_0 (activation-quantized), the GPU backends read f32 activations for all weight types (CUDA prefill additionally offers the default-on int8 MMQ path). Compare each path against its own reference, not against the other.

4.6 Per-layer computation (Qwen2) — as built by graph.rs

flowchart LR
    A["token_embd lookup (GetRows)"] --> B["hidden"]
    B --> C["RMSNorm attn_norm"]
    C --> D["WQ / WK / WV matmuls + bias"]
    D --> E["RoPE on Q and K (in-place alias)"]
    E --> F["kvcache_store: K/V → layer regions"]
    F --> G["GQA attention<br/>Q·K^T → softmax → ·V"]
    G --> H["WO matmul + bias"]
    H --> I["+ residual → hidden"]
    I --> J["RMSNorm ffn_norm"]
    J --> K["FFN gate + up matmuls"]
    K --> L["SiLU(gate) × up (fused SwiGLU)"]
    L --> M["FFN down matmul"]
    M --> N["+ residual → hidden"]
    N --> O["next layer / output_norm"]

GQA: each query head h maps to KV head hk = h / gqa. The KV head dimension is independent (n_kv_embd read from the K weight's ne[1]), so Qwen2.5-0.5B (n_embd=896, n_head=14, hd=64, n_kv_embd=128) strides correctly.

4.7 Prefill vs decode

Phasent (tokens)Notes
Prefill> 1graph type Prefill; KV store writes all positions, attention reads the full written prefix
Decode1graph type Decode; same topology as prefill modulo nt → the graph is rebuilt once (KV persists), then reused for every subsequent token

The graph path keeps K/V on the executing backend (attention and KV are on the same backend by construction), so there is no per-token KV drain.


5. Backend Layering

5.1 The Backend trait (graph/backend.rs)

Backend (the thing the IR stores on a node) is not this trait: since F4 it is a Copy handle — an id into src/graph/registry.rs — and the trait below is implemented by the pools (CPU, Metal, CUDA) and held by the registry entry. The name, id, priority and capability record live on the handle/entry, not on the trait (Backend::name() is registry.rs:105):

#![allow(unused)]
fn main() {
pub trait Backend: Send + Sync {
    fn kv_pair(&self, layer: usize) -> Option<(usize, usize)>;   // persistent per-layer K/V regions
    fn supports_op(&self, op: &Op, dtype: DType) -> bool;
    fn supports_fused(&self, fused: &FusedOp) -> bool;
    fn supports_attn_span(&self) -> bool { false }       // E1: explicit attention windows (CPU/CUDA/Metal)
    fn weights_bytes(&self) -> usize { 0 }               // E4: bytes the memory budget is charged for
    fn alloc_buffer(&mut self, size: usize) -> usize;    // backend's own pool
    fn free_buffer(&mut self, id: usize);
    fn pool_len(&self) -> usize;
    fn alloc_fresh(&mut self, size: usize) -> usize;     // bypasses the recycle free list (split-boundary staging)
    fn execute_node(&mut self, node: &CNode, in_bufs: &[usize],
                    out_buf: usize, kv_pair: Option<(usize, usize)>) -> Result<(), String>;
    fn copy_cells(&mut self, dst: BufRef, src: BufRef, dst_row: usize, src_row: usize,
                  rows: usize, elems_per_cell: usize) -> Result<(), String>;   // C3 compaction / C2 shift
    fn read_host(&self, id: usize) -> Option<&[f32]>;
    fn write_host(&mut self, id: usize, data: &[f32]) -> Result<(), String>;
    fn write_host_window(&mut self, id: usize, offset: usize, data: &[f32]) -> Result<(), String>;
    fn synchronize(&mut self);
    fn retire(&mut self) { self.synchronize(); }         // release device objects for a dropped pool
    // CUDA only: try to replay a captured graph for (uid, range); capture is
    // gated to decode-shaped graphs. Default impl returns false.
    #[cfg(feature = "cuda")]
    fn graph_replay(&mut self, uid: u64, range: (usize, usize), nt_hint: Option<usize>) -> bool;
}
}
  • CPU (cpu_backend.rs): Vec<f32> pool; executes via kernel.rs + vec_ops.rs; F32 weights use a plain f32 matmul, quantized weights use cpu_quant_matmul_f32 (Q8_0-activation path).
  • Metal (metal_backend.rs, macOS): shared-memory MTLBuffer pool; per-op dispatch to MpsState's kernels (rms_norm, quant_matmul_f32_on_gpu_buf, rope, silu, add, mul, swiglu, embed_tokens_gpu, store_kv, gqa_attn_f32); one command buffer per split. Weights resolve by name from MpsState's registry (weight_buf(name) -> (buffer, offset)).
  • CUDA (cuda_backend.rs, feature-gated --features cuda): wraps the cuda.rs device layer — per-op dispatch with int8 MMQ prefill + MMVQ decode (default-on; MINFER_MMQ=0 reverts; the MMQ gate resolves from the device_tier.rs device-tier table), split-KV attention, and CUDA Graph capture/replay keyed on the graph uid (graph_replay, decode-shaped only). Implementation record: docs/CUDA-BACKEND-DESIGN.md; per-step optimization history in docs/CUDA_OPTIMIZATION.md.

The allocator owns every backend pool (single source of truth); the scheduler orchestrates assignment, cross-backend copies, and sync.

5.2 Selection rules

Assignment asks each node in the registry's priority order — Metal 300, CUDA 200, CPU 100 (src/graph/registry.rs:72-75) — for the first backend that has a pool, is not fenced off, and reports the op supported; the CPU is tried last and is always allowed, so the answer is total (GraphAllocator::supports_for, src/graph/alloc.rs:403). The fence is --backend <name> / MINFER_BACKENDS=<csv>, and an unknown name, a name this build does not contain and a name this machine cannot use are three distinct startup refusals. The two orders, the name surface and those refusals are docs/BACKEND-REGISTRY-DESIGN.md's contract; the user-facing flags are docs/USAGE.md.

  • Metal: all graph weights must be GPU-registered (Qwen2Graph::weights_on_gpu mirrors the old per-layer check). MINFER_DISABLE_MPS=1 forces CPU.
  • CUDA: feature-gated (--features cuda), and every graph weight must be GPU-registered with a device present. A node whose block the E5 offload plan left on the CPU never reaches the device (supports_for answers CPU for it, so a partially offloaded model cannot run a block whose weights were never registered there). MINFER_DISABLE_CUDA=1 is presence-checked and forces the CPU.
  • CPU: always available; AVX2 dispatch via is_x86_feature_detected!("avx2") on x86, NEON+SDOT on aarch64, scalar fallback elsewhere (MINFER_NO_NEON=1 forces scalar; the K-quant dots also have an AVX-512/VNNI path since #56, gated by MINFER_NO_AVX512=1 and MINFER_NO_AVX2=1).

5.3 GPU safety

All Metal submits wait bounded (10 s) and check status; no early return past a threadgroup_barrier; device limits queried at runtime, never hardcoded; guard failures gpu_abort with actual values. In the graph architecture, kernel-invariant violations return Err from execute_node and must NOT be treated as a silent CPU fallback — backend assignment at build time decides where ops run; only genuine support limitations (e.g. Raw weights) select the CPU backend. See docs/GPU_SAFETY.md.


6. Quantization & Tensor Layout

  • Weight layout (block.rs): repr(C) blocks matching ggml-common.h. Q4_0/Q4_1/Q5_0/Q5_1/Q8_0 = 32-value blocks; Q4_K/Q5_K/Q6_K = 256-value super-blocks. Supported: Q4_0, Q4_1, Q8_0, Q4_K, Q6_K, Q5_0, Q5_1, Q5_K (CPU + Metal), F32/F16 norms & biases.
  • Activation quantization: CPU matmuls quantize f32 activations to Q8_0 on-the-fly (quants.rs); Metal reads f32 directly.
  • GGUF parsing (gguf.rs): ggml_pad(x, n) = (x + n - 1) & !(n - 1) alignment; tensor strides computed from type block size exactly like ggml; split multi-part models are merged into one tensor index.
flowchart LR
    A["GGUF file"] --> B["metadata KV<br/>hparams + tokenizer + template"]
    A --> C["tensor table<br/>name / type / shape / offset"]
    C --> D["quantized data blob"]
    D --> E["Tensor: type, shape, strides, Vec&lt;u8&gt;"]
    B --> F["HParams"]
    B --> G["Tokenizer"]
    B --> H["Chat template"]

7. KV Cache

The graph path owns the KV cache: two persistent regions per layer (K and V), sized n_kv_embd × n_ctx, allocated in the backend pool the layer runs on. kv_pair(layer) resolves them; the KV store node writes K/V at the positions carried by the positions input; attention reads the written prefix. The regions live inside the GraphCache's allocator and survive graph rebuilds (the prefill→decode transition). MINFER_CACHE_TYPE=f16 selects an f16 GPU cache where the kernels support it. The pre-graph cache.rs KVCache type — and the vestigial &mut KVCache argument of ModelDef::forward — was deleted in #252; these regions are the only KV store.


8. Adding a New Architecture

  1. Create src/models/<name>/ with mod.rs, graph.rs, loader.rs.
  2. Add a match branch in src/models/mod.rs::load_model() for the new general.architecture value.
  3. In loader.rs: define HParams (including n_kv_embd) and LayerWeights, parse them from GGUF metadata (both qwen2.* and llama.* prefixes are accepted).
  4. In graph.rs: implement build_graph(&self, params: &GraphParams) -> ComputeGraph deterministically in params (the reuse invariant), using GraphBuilder — mirror llama.cpp's llm_graph_context builder methods.
  5. In mod.rs: implement ModelDef (forward, build_graph, forward_graph, forward_graph_cached, as_any, special_tokens, n_layer/n_head_kv/n_embd_head/n_kv_embd/n_vocab, rope_style). models/qwen3/ is the worked example of a second architecture (decoupled head dim + per-head Q/K norm).
  6. If needed, add a chat template format in template.rs.

Architectures that share Qwen2's tensor naming convention (LLaMA, Mistral, Phi) are the easiest ports. The RopeStyle enum defines both Qwen2 (non-interleaved) and Llama (interleaved) pairings, but only the non-interleaved form is wired up: both loaders hard-code it (models/qwen2/loader.rs:155, models/qwen3/loader.rs:173) and the CUDA backend refuses the interleaved style (graph/cuda_backend.rs:1522). A family that needs interleaved RoPE requires a loader change plus a CUDA kernel — see docs/MODEL-SUPPORT-ROADMAP.md, "Cost model: what a port actually costs".


9. Model Download

download/mod.rs resolves hf:<repo>[:<quant>] and ollama:<model>[:<tag>] URIs, downloads via curl with size-checked resume, quant-matches single or split files case-insensitively, and stores them under ~/.cache/minfer/models. Cached filenames can be used directly as the model argument.


TopicLocation
Compute graph design + rewrite plan + implementation record (per-phase commits)docs/COMPUTE-GRAPH-DESIGN.md
llama.cpp compute-graph design analysis (ggml_cgraph / scheduler / reuse)docs/LLAMA-COMPUTE-GRAPH.md
Metal backend optimizations / gap analysis (primary tracking)docs/METAL_OPTIMIZATIONS.md
GPU safety conventions + auditdocs/GPU_SAFETY.md
CPU backend optimizationsdocs/CPU_OPTIMIZATIONS.md
CUDA optimization history + per-step recordsdocs/CUDA_OPTIMIZATION.md (+ docs/cuda_optimization_steps/)
Debug dump formatdocs/debug-dump.md
Historical bugs / debugging notesdocs/BUG-6-KV-CACHE-INDEXING.md, docs/QWEN2.5-*, docs/DEBUGGING-*

Appendix A — Legacy Imperative Architecture (removed in Phase 6)

Historical reference. The imperative per-layer forward (models/qwen2/forward.rs) was the engine's core until the compute graph replaced it (Phase 6); it is preserved here verbatim in structure so old notes, benchmarks, and kernel analyses remain interpretable. Do not treat this as the current design.

A.1 Design stance (then)

The engine ran a direct per-layer forward loop instead of a compute graph: "simpler, easier to trace, and the whole layer can be fused onto the GPU." The GPU fallback was per-layer and safe: a layer that could not run on the GPU (e.g. unsupported weight type) submitted partial GPU work, downloaded the hidden state, and continued on the CPU.

A.2 The old forward pass

ModelDef::forward(tokens, positions, kv) was implemented in models/qwen2/forward.rs as an imperative loop:

flowchart LR
    A["token_embd lookup"] --> B["hidden"]
    B --> C["RMSNorm attn_norm"]
    C --> D["WQ / WK / WV matmuls + bias"]
    D --> E["RoPE on Q and K"]
    E --> F["store K/V into KV cache"]
    F --> G["GQA attention<br/>Q·K^T → softmax → ·V"]
    G --> H["WO matmul + bias"]
    H --> I["+ residual → hidden"]
    I --> J["RMSNorm ffn_norm"]
    J --> K["FFN gate + up matmuls"]
    K --> L["SiLU(gate) × up"]
    L --> M["FFN down matmul"]
    M --> N["+ residual → hidden"]
    N --> O["next layer / output_norm"]

Decode-time optimizations: fused QKV (nt==1) via a concatenated blk.{il}.attn_qkv weight, fused bias+rope+store (attn_bias_rope_store), fused SwiGLU kernel, and the last layer computed only the tail n_out rows (an inp_out_ids-style partial-row optimization). The n_out tail-row optimization was subsequently carried into the graph path as the G3 work: the builder inserts GetRows(wo, tail_ids) (and the same for the residual) before the last layer's FFN, so the last FFN, the final norm and lm_head all run on n_out rows only (models/qwen2/graph.rs:64, :228-231; GraphParams.n_out is part of the reuse identity).

A.3 Old backend layering & fallback

flowchart TD
    A["forward nt tokens"] --> B{"embedding on GPU?"}
    B -->|yes| C["GPU embed lookup → buf_hidden"]
    B -->|no| D["CPU embed → upload hidden"]
    D --> E["upload positions"]
    C --> E
    E --> F{"per-layer: layer_gpu ok?"}
    F -->|"yes, all layers"| G["output_norm_gpu<br/>on GPU"]
    F -->|"no at layer i"| H["submit partial GPU work<br/>download hidden, sync KV to CPU"]
    H --> I["CPU loop from layer i"]
    G -->|"output on GPU"| J["download logits → return"]
    G -->|"output fell back"| I
    I --> K["output_norm + output matmul on CPU"]
    K --> L["return logits"]

Selection rules (then):

  • Metal: layer 0 must have all 7 weight matrices + norms registered on the GPU. Within a layer all 7 matrices must be all Q4 group (Q4_0/Q4_1) or all QK group (Q4_K/Q5_0/Q6_K); Q5_1/Q5_K use the f32 path and are exempt. MINFER_DISABLE_MPS=1 forces CPU.
  • CUDA (--features cuda): requires every layer's 7 matrices to be all Q4_0/Q4_1 or all Q4_K/Q6_K. Decode replays a captured CUDA Graph.
  • CPU: always available; AVX2 dispatch via is_x86_feature_detected!("avx2"), scalar fallback elsewhere (plus the AVX-512/VNNI K-quant dots, #56).

The GPU path skipped the per-token CPU→GPU KV drain (no sync_kv_to_cpu) because GPU-layer failure is deterministic by weight type — the sync only happened in the fallback branch.

A.4 Old KV cache

cache.rs provided an architecture-agnostic per-layer cache: k/v were pre-allocated Vec<f32> of max_size × dim, size tracked the current sequence length. store_multi wrote K/V for many positions at once (prefill); decode wrote one. The GPU maintained its own buffers and sync_kv_to_cpu copied them back only on the CPU-fallback path. MINFER_CACHE_TYPE=f16 selected an f16 GPU cache (opt-in).

A.5 Old "Adding a New Architecture"

  1. Create src/models/<name>/ with mod.rs, forward.rs, loader.rs.
  2. Add a match branch in src/models/mod.rs::load_model().
  3. Define HParams (including n_kv_embd) and LayerWeights in loader.rs.
  4. Implement the per-layer forward pass in forward.rs using kernel::, vec_ops::, cache::.
  5. Implement ModelDef in mod.rs.
  6. Add a chat template format in template.rs if needed.

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0001 — Inference runs through one declarative compute graph
  • ADR-0012 — Device is the first axis, the layer the second — and no premature common
  • ADR-0002 — Topology is a function of GraphParams alone, so positions cannot be structure
  • ADR-0007 — No ML frameworks: every operator is hand-written
  • ADR-0008 — GPU safety: bounded waits, no early return past a barrier, runtime device limits
  • ADR-0009 — A failure is an error, never a silent fallback
  • ADR-0013 — The CPU quantizes activations to Q8_0; a device reads f32

minfer Architecture Roadmap

Status: analysis + prioritized backlog (no code changed). Baseline: minfer HEAD = 293eb19 (2026-09-15, working tree clean). Method: direct reading of src/ (44,105 LOC) and docs/ (~15.6k LOC). Every claim about minfer carries a file:line anchor; claims taken from a document rather than code are marked "per docs".

Scope. Architecture-level work on the system layers: IR, scheduler, allocator, KV/sequence state, batching, backend abstraction, and model/quant coverage. Kernel micro-optimization is out of scope — CUDA_OPTIMIZATION.md and METAL_OPTIMIZATIONS.md cover it, and both campaigns are formally converged. Model-family selection is a separate document: docs/MODEL-SUPPORT-ROADMAP.md.


0. Verdict

minfer's kernel work is complete and competitive for the architectures it supports: eight quantized weight types on CPU, Metal and CUDA; int8 MMQ (quantized matrix-multiply on tensor cores) prefill and dp4a MMVQ (quantized matrix-vector multiply) decode on CUDA; simdgroup GEMM (general matrix-multiply) and flash attention on Metal; speculative decoding with adaptive draft depth. On the measured CUDA target, 7B Q4_K_M reaches ~3581 tok/s prefill and ~51.2 tok/s decode; on Metal, 7B decode is at parity and prefill reaches ~0.82–0.85× of the achievable ceiling, with the residual measured as not source-addressable.

The remaining work is at the system layer, and three items dominate it:

  1. No multi-sequence batching. The IR, the attention kernels, the KV (Key/Value) allocator and the server are all single-sequence. E1/E1b made the attention side sequence-aware and E2 landed sequence-addressable batches and a batching server worker; the CPU payoff measured negative while the GPU payoff was 1.9x, so the default follows the device (E6: batches on CUDA and, since #44 part (b) landed the one-range attn_span read and the matching write/move side on a Mac (2026-10-06), on Metal too; off on CPU).
  2. KV cache is a fixed per-layer buffer, not a sequence-addressable cell store: no sequence ids, no eviction/context shift, no defragmentation, no state save/restore, no quantized KV. (Phase C has since closed every one of those: C1/C2 the cell store and physical removal/shift, C3 the compaction — a pure planner, the counters, Backend::copy_cells on CPU and CUDA, the K re-rope a move needs while positions are cells, and a model-level continuation gate — C4 the packed Q8_0 cache (CPU, CUDA and Metal — the last in #310) and C5 the session container. What is still open is driving the compaction from the server's serving model, which reserves every slot once and packed, the CUDA packed-decode residual's stall attribution (#212 — #144/#186/#202 landed the fused epilogue, the dp4a dot and the FA prefill), and Metal's packed half (G5, #44; #310, enabled on Metal — mechanism A native packed decode + mechanism B f32 stage). See §2.4.)
  3. The server has no persistent context. Every request builds a fresh GraphCache (KV regions + device pool re-allocated, CUDA Graph capture re-warmed) and re-prefills the whole prompt. Fixed in B2/B3 for the prefix-matched case: a slot keeps its cache and reuses the rows its prompt already covers (measured ≈11× TTFT on a second turn).

Everything else — IR expressiveness (views, multi-output), memory placement policy (VRAM budget, layer offload), backend pluggability, additional model families, quantizer tooling — is real but secondary, and mostly enabled by fixing (1) and (2) first.

Ranked backlog (full list with numbering in §3)

§3 item(s)ItemLayerClassEffort
1–3Multi-sequence batch + continuous batchingL4+L5+L2XL3–6 w
1KV cache redesign (cells, seq ids, prefix reuse, shift/defrag)L4XL2–5 w
4Persistent server context (no per-request rebuild/realloc)L5+L3M3–5 d
7IR expressiveness: strided views/aliasing + multi-output nodesL1L1–2 w
8–9Memory placement policy: VRAM budget, layer offload, size-class allocator — E4 complete (accounting/gate, size-class pools, reserve/assign + the multi-graph cache) and E5 complete (explicit per-block offload and the auto fit); the Metal free-bytes query landed as #53's DeviceMemory answer, so auto and headroom_bytes() work on macOS too; only per-block tuning remainsL3+L6L1–2 w
12Backend registry (drop the hard-coded 3-way enum/match) — DONE (F4, 2026-09-24): the enum is a Copy/Hash/Ord handle over a fixed id space and the name-keyed table lives in graph/registry.rs; the nine match sites are gone. The set is still closed at compile time — this makes it data, not pluggable at runtimeL6M4–7 d
11CPU: AVX2/AVX-512 for the K-quant dots + weight repackingL7L1–2 w
10Chunked prefill (n_batch actually used) — E3, landed 2026-09-22L2+L5M3–5 d
15Grammar / JSON-schema constrained decoding — DONE (F2, 2026-09-24); the mask is host-side and per state (5.4 ms on a 151k vocabulary)L7M4–6 d
23Op × dtype × backend correctness matrix in CIL8M3–5 d

1. Current state

1.1 Scale

MetricValue
src/ Rust44,105 LOC across 48 files
Largest filesgraph/cuda_backend.rs 6,664 · src/cuda.rs 6,340 · src/metal/ 3,944 · models/qwen2/graph.rs 2,253 · gguf.rs 2,096
Compute-graph core (src/graph/)13,052 LOC; 3,391 LOC excluding the three backends
CUDA kernelssrc/cuda/kernels/*.cu (17 translation units + common.cuh, compiled by build.rs only under --features cuda; #263)
Metal shaderssrc/metal/kernels/ → minfer.metallib at build time, source-compile fallback
Tests5 integration files (tests/, ~3.1k LOC, four of them Metal-only) + inline #[cfg(test)]
Docs~182 Markdown files; 41 in docs/, 106 numbered CUDA campaign records

1.2 Landed surface

Compute graph (build → assign → fuse → alloc → execute) with params-only reuse (graph/cache.rs:47, graph/params.rs:51); per-op backend assignment (graph/scheduler.rs:60); liveness allocator with persistent per-layer KV regions (graph/alloc.rs:165, :385); eight quantized weight types (Q4_0/Q4_1/Q5_0/Q5_1/Q8_0/Q4_K/Q5_K/Q6_K) on CPU + Metal + CUDA; two model families (qwen2, qwen3 dense); GGUF v3 with split parts and Metal zero-copy mmap; self-contained BPE tokenizer; chat templates via minijinja; greedy/penalty/top-k/top-p/temperature sampling; speculative decoding with adaptive depth; multi-turn conversation CLI; OpenAI-compatible server (chat completions, streaming, multi-slot); a runtime graph introspection stack (MINFER_TRACE per-node capture, live SSE visualizer, pipeline view, DOT/JSON export).

1.3 Documentation hygiene

The index-level documents (README.md, ARCHITECTURE.md, CPU_OPTIMIZATIONS.md, QWEN3-SUPPORT-PLAN.md, OPENAI-CHAT-API-PLAN.md, MODEL-SUPPORT-ROADMAP.md) are the entry points AGENTS.md routes readers through, so a stale statement there costs more than one buried in a campaign record.

A pass on 2026-09-15 fixed these stale statements:

  • README.md and ARCHITECTURE.md — the removed BiasRope fusion in the two architecture diagrams, plus the fusion.rs row in the module table (the pass produces only SwiGLU; the decode fusions are built by the model code, not by FusionPass).
  • ARCHITECTURE.md — the "n_out tail-row optimization is not yet ported" claim (it landed as the G3 work), and the RopeStyle "already covers Qwen2 and Llama" claim (only the non-interleaved form is wired up).
  • CPU_OPTIMIZATIONS.md — the pre-compute-graph snapshot is now banner-marked as historical, with the still-open AVX2 gap pointed at §2.7.
  • QWEN3-SUPPORT-PLAN.md — the self-contradictory status line.
  • OPENAI-CHAT-API-PLAN.md — the "KV is f32" note (f16 is auto-selected for 7B-class models).
  • MODEL-SUPPORT-ROADMAP.md — the misleading status line.

Keep the invariant: when a change lands, update the index rows in the same commit. No outstanding doc-truth work item remains.


2. Layer-by-layer analysis

The layer names follow the vocabulary of docs/GLOSSARY.md, which classifies the campaign's terms into seven layers; §2.8 adds one cross-cutting layer for ops, safety, testing and observability.

Legend for Gap: 🔴 structural (blocks a class of use cases) · 🟠 material (measurable capability or performance loss) · 🟡 hygiene.

2.1 L1 — Graph IR and build

Today. ComputeGraph is a flat, topologically ordered vector of CNode (graph/mod.rs:83-115). Each node has exactly one output, a static [usize; 4] shape, and an Op carrying its full payload. Op::View { offset, shape }, Reshape, Permute exist (ops.rs:105-116) but both backends execute them as identity copies (cpu_backend.rs:384-388, cuda_backend.rs:433-438, metal_backend.rs likewise) — the allocator has no aliasing for views, so a view costs a full tensor copy. The builder appends sources before consumers, so node-id order is the execution order (scheduler.rs:214).

Gap. 🔴 Strided, zero-copy views and multi-output nodes were missing — closed by D1 (2026-09-19): views are zero-copy with allocator-known aliasing, and GraphBuilder::split_parts exposes a node's single output as independent parts, so what remains under this header is model work (item 17), not IR work. Without them the IR cannot express slicing, concatenation, or a copy between overlapping regions as compositions of primitive ops — which is why minfer carries four hand-written decode-specific fused ops instead: FusedQKV, QkvBiasRopeStore, FusedFFN, FusedQkvNorm (ops.rs:125-154). Each one costs changes in five places: the builder constructor (builder.rs:193-281), the allocator's special cases (alloc.rs:232-273), the scheduler's kv_pair resolution (scheduler.rs:261-271), every backend's supports_op + execute_node, and both models' build_graph. BatchMatMul is deferred for the same reason (COMPUTE-GRAPH-DESIGN.md §5.4).

Consequence for the model work. MODEL-SUPPORT-ROADMAP.md ranks new model families by how much of the existing graph they reuse. That ranking is accurate for parameter-isomorphic architectures (Llama/Mistral) but not for anything needing new structure: MoE (Mixture-of-Experts) needs mul_mat_id (3-D expert indexing), MLA (Multi-head Latent Attention) needs a different KV layout and a view-based latent split, and sliding-window attention needs a mask parameter threaded into attention. Each of those currently pays the five-site cost.

Recommendation. Add strided views with allocator-known aliasing, and let a node carry a small output list. Fusions can then be produced by the fusion pass from compositions rather than hand-written as IR variants, and the four decode fusions become compiler output instead of bespoke ops.


2.2 L2 — Scheduler and execution

Today. assign_backends (scheduler.rs:68-83) walks nodes in build order and assigns the highest-priority backend whose supports_op returns true. split_graph partitions into contiguous same-backend runs derived from that positional assignment (scheduler.rs:73-120); split inputs/outputs are the crossing edges. Execution then walks splits, calling sync_backend(prev) and copy_across for each cross edge (scheduler.rs:176-189), and finally sync_backend.

Gap. 🟠 Two distinct issues.

  1. No cost model. Splits are positional. A single unsupported op in the middle of an otherwise GPU graph produces three splits with two host round trips; the assignment never considers the cost of the resulting movement, or whether a neighbouring op could be moved so the op becomes supportable. In practice this is masked today because the supported models build a single GPU split on both GPU backends (CUDA-BACKEND-DESIGN.md:313-315), but it is exactly what breaks the first time an op is unsupported — e.g. interleaved RoPE on CUDA (cuda_backend.rs:1522), or FusedQkvNorm, which CUDA does not advertise (cuda_backend.rs:1905) while Metal does (metal_backend.rs:686).
  2. Synchronous, host-mediated cross-backend movement. — partly closed by F5 (#58, 2026-09-24): the boundary is now two registered phases. Phase A (copy_across → BackendEntry::copy_cross) enqueues the transfer — for a CUDA source a device→host cudaMemcpyAsync into a pinned slab plus a cudaEventRecord, where the pre-F5 code did copy_to_cpu (a full stream sync + a blocking cudaMemcpy) — and phase B (await_cross) waits on the recorded event at the documented synchronization point, once per staged input. The CPU's hooks are a synchronous host round trip and a no-op (it has no device memory); Metal runs the same two phases — a MTLBlitCommandEncoder copy into a shared staging buffer plus encodeSignalEvent on a MTLSharedEvent, waited once by a bounded waitUntilSignaledValue — ported by #137 and verified on a Mac (2026-10-05: blocking_host_copies 21 → 0, copies == waits, bitwise identical). Counters (graph/copystats.rs), the per-backend table, the enumeration of the synchronization points and the measured before/after are in docs/BACKEND-REGISTRY-DESIGN.md §11 and the plan's F5 record. What remains: nothing filed for the boundary copy. The wait is already deferred to the consumer's first use (#138), so a boundary with several staged inputs holds them in flight at once; true cross-split overlap has no target on the reachable macOS topology, and #300 recorded that negative result — a single device split feeding a host consumer whose first node reads the dominant staged tensor (docs/BACKEND-REGISTRY-DESIGN.md §11.6). Multi-device execution is still the place that would need it.

Also: the cross-boundary staging map is keyed by node id alone (alloc.rs:34), so a node consumed by two different foreign backends can only have one staging buffer; the scheduler compensates by filtering on backend (scheduler.rs:252-255), which is correct only while at most two backends are enabled. Fixed in A5: the map is keyed by (node, destination backend), one node can feed two foreign consumers, and the consumer-side filter is gone.

Recommendation. Add an assignment pass that (a) propagates support backwards from unsupported ops and (b) scores a candidate assignment by crossing count; and replace the host round trip with an async copy plus a recorded event when the source and destination are different devices (F5 landed the CUDA half of this; the Metal half and the actual consumer-side waiting that would let a copy overlap independent work are the remaining follow-ups). Neither is urgent today; both are prerequisites for multi-device execution and for any heterogeneous split.


2.3 L3 — Allocator and memory

Today. GraphAllocator::alloc_graph (alloc.rs:513-521) rebuilds the whole node→buffer mapping on every graph rebuild: it frees every previously live buffer back to the backend pool (:165-177), recomputes last_use over build order, and re-allocates. Buffer pools are per backend, not unified (alloc.rs:315-335), and allocation is by exact element count — both the Metal (metal_backend.rs:322-332) and CUDA (cuda_backend.rs:1333-1353) pools scan a free list for an exact byte-length match and otherwise allocate fresh. free_buffer never returns memory to the device (cuda_backend.rs:1355-1362).

Gap. 🟠 Three consequences.

  1. Per-rebuild teardown. Non-persistent buffers are freed and re-derived on every topology change. For a decode loop with a stable shape this happens once; for anything with a varying n_tokens (speculative decoding with adaptive depth, server requests of differing prompt lengths, future continuous batching) it happens every time, together with a full device re-allocation for any size not seen before.
  2. No size classes / no rounding. Because matching is exact, a workload touching k distinct activation shapes ends up with k sets of live buffers. The pool is a per-GraphCache high-water mark that never shrinks (documented as accepted debt, CUDA-BACKEND-DESIGN.md:447) — but that was reasoned about for a fixed-shape CLI run, not for a server whose prompt lengths vary per request. E4 S2 (2026-09-23) rounds every pooled activation to its class, so shapes inside one class share a buffer across rebuilds; E4 S3 then split reservation from assignment (a slots table per backend and class: a released buffer is re-assigned rather than returned to the pool, so a rebuild re-maps and CUDA's pool_gen — the capture generation — stops moving) and gave GraphCache one graph per GraphParams, so a switch is a re-map, not a build (§14 row 3, closed by E4 S3).
  3. No VRAM budget or feasibility check. cudaMemGetInfo is queried once at init and only printed (src/cuda/methods/init.rs:126); the sole consumer of free-memory information today is a valve guarding the optional f16 weight cache (src/cuda/methods.rs:57-62) — the activation/KV allocator has no accounting at all. Out-of-memory surfaces as a null pointer that fails at execute time (cuda_backend.rs:1345-1348). There is no "this graph will not fit, offload the last n layers" fallback because there is no layer-offload concept at all (§2.6).

Recommendation. Split allocation into a reserve phase (size the graph for the worst-case shape the session will use: n_ctx, n_batch) and an assign phase that re-maps nodes into reserved regions without touching the device. Round pool sizes to a size-class ladder (e.g. 16/64/256 KiB steps) so distinct shapes share. Add real memory accounting (bytes_reserved, bytes_live, bytes_device_free) and make the prefill path consult it before allocating, with CPU execution as an explicit fallback decided at build time (consistent with the "no silent fallback" rule — the fallback would be visible in the backend assignment).


2.4 L4 — KV cache and sequence state 🔴

Today. Each layer owns two persistent contiguous regions, K and V, created on first use by ensure_kv(layer, backend, size) (alloc.rs:385-393), sized n_kv_embd × n_ctx f32 (or f16 when the type flag is set). They live in the allocator inside GraphCache and survive graph rebuilds (graph/cache.rs:69). Positions are data, injected per step through the positions input node, so the topology never depends on n_past (ops.rs:94-99, COMPUTE-GRAPH-DESIGN.md §1.4).

That design is a genuine strength — it is the precondition for graph reuse — but it stops at the single-sequence append-only case. What is missing:

CapabilityStatus
Sequence ids / per-sequence views of one cache✔ C1 per-cell owner (SeqId); E1/E2 made several views of one cache real (reservations + explicit attn_span, device-aware batching default via E6); C7 increment 1 made the partition elastic. One cell still belongs to exactly one sequence — C8 (#41, closed 2026-09-22) turns owner into a set/refcount so a prefix can be shared, and C8b's span list + kv_map (#41) delivers it on all three backends — CPU and CUDA first, Metal once its kv_map kernel landed in #362
Partial removal / keep / copy between sequences◐ C2: physical removal of a row range; compaction inside one arena is C3 (done); sharing one cell range across sequences no longer needs owner[cell] to become a set — C8b S2 landed 2026-09-21 derives occupancy from per-sequence span lists and S3 landed 2026-09-22 adds the copy-on-write store rule, both gated on byte-equality; CUDA's gather (S4) is next
Defragmentation✔ C3 (2026-09-19): pure planner + KvArenaStats counters + Backend::copy_cells (CPU copy_within; CUDA kv_move_rows, one block walking rows ascending with a barrier, because overlapping device-to-device copies are undefined; Metal moves rows through MTLBlitCommandEncoder in the same order since #44 part (b), 2026-10-06) + host K re-rope while positions are cells + a mid-session compaction gate. The re-rope is gone (C6 landed): compaction moves rows verbatim and is bit-identical. C7 is the scheduled consumer — the server's dynamic-run trigger uses the planner to make a slot's reservation elastic
Sliding-window eviction◐ C2: physical shift + re-rope (exact mechanism; the retained rows keep the context they were written in, see below)
Recurrent / hybrid memory (state-space models)✗
Context shift (keep KV, shift positions)✔ C2: kv_rm/kv_shift plus the conversation's overflow shift — 185 → 14 prefilled tokens per overflowing turn on the 0.5B probe; MINFER_NO_CONTEXT_SHIFT=1 restores the exact re-render
State save/restore (session persistence)✔ C5 (2026-09-22): a versioned, checksummed container (graph/kvsession.rs) holds the shape, the backend, the KV element type, one K/V blob per layer and the arena's bookkeeping (owner table + run table + span lists); GraphAllocator::kv_save/kv_load stream it through the backends' host I/O. A load verifies the whole file before it applies anything, so a truncated/corrupted/foreign file is a no-op, and the real-model gate resumes a session that continues bitwise (max |Δlogit| = 0, CPU and CUDA). The element type rides in the header flags (FLAG_PACKED Q8_0 / FLAG_F16 f16 / 0 f32; the two bits are mutually exclusive and an unknown bit is refused), so C5 S3 (#130, 2026-09-25) closed the last gap — a CUDA f16 auto-policy session (--session or --slots-file) now restores instead of refusing its own snapshot, and the Qwen3-0.6B configuration's real-model set is 31/0. Resuming the CLI's --session and an E2 slot snapshot: #89 · #43
Prefix reuse across requests✔ B2/B3 — ≈11× TTFT on the second turn
Quantized KV◐ C4 S1 + S2a + S2b (2026-09-24, CPU and CUDA): MINFER_CACHE_TYPE=q8_0 stores packed Q8_0 cells — 3.76× smaller regions, measured on both cached models. S2a reads the packed blocks directly on the CPU (Q8_0 × Q8_0 K dot, V out of the cell: 1.16×/1.31× over S1's dequantizing read at ctx 512/2048) and makes a physical kv_rm/kv_shift work on a packed region; S2b adds the CUDA kernels (a KV_LAYOUT_F32/F16/Q8_0 tag with a byte-addressed kv4<LAYOUT> load, a packed store that uses the CPU's quantizer byte for byte) and flips the registry's reads_packed_kv. The win is memory — the figures in this sentence are the baseline 293eb19 measurement, kept as measured and superseded: decode against f16 was 1.13× slower on Qwen3-0.6B and 1.48× on the 0.5B (1.25× of it the packed load, the rest the stated no-fused-epilogue cut); at hd 128 the packed prefill was 15× slower because the f16-typed FA path was not offered for a packed cell. #144 items 1+3 (041de15, 2026-09-26), #186 (cc19b4f, 2026-09-27) and #202 (798fd32, 2026-09-27) landed the packed fused epilogue, the dp4a packed K dot and the packed FA prefill; the current A/B is the 2026-10-07 comment on #310, measured at 740e0ff on dgxspark (aarch64, GB10 sm_121): 1.23× decode / 1.24× prefill at hd 64 and 1.01× / 1.04× at hd 128 — q8_0 is still never faster than f16 on CUDA. Metal reads the packed cell since #310 (reads_packed_kv is true there; mechanism A reads the decode flash family natively and mechanism B stages the prefill/window window to an f32 buffer), so q8_0 is a CPU/CUDA/Metal feature; the one open CUDA lever is #212 (a stall-level attribution of the residual — #202's L1-request hypothesis is refuted). Format ownership is per engine since #99 (2026-09-25): the resolved format lives on the loaded model and reaches the graph through CParams::kv_format and the CPU kernels through GraphAllocator::set_kv_format; the CUDA device layout tag is still process-wide (#153)
KV memory growthfixed at first allocation, never resized
Position vs cell✔ C6 (landed 2026-09-20, 001b8cc): sequence-relative positions + an allocator-resolved cells input, so a cell move changes no rotation and compaction is bit-identical. The CUDA fused decode QKV family (Op::FusedQKV, Op::QkvBiasRopeStore) takes cells too, so the build-time gate that kept it on the unfused chain under explicit_span is now (cuda_on || !explicit_span) — Metal keeps the fused node (a CUDA-side gate, not a Metal limitation). A request placed in a non-zero-start slot exposed and fixed a server position bug (submit_on still added the run start). The two user-visible limits this design leaves open are C7 (one request may use the whole arena) and C8 (cross-sequence sharing), both now scheduled
Elastic per-slot KV partition (a long request vs n_slots)✔ C7 + C7b (2026-09-20): the engine sizes a slot from the request (prompt + max_tokens + 1, clamped to the arena); when that exceeds its share it reclaims idle runs for capacity and KvCache::set_cap returns a plan that moves whatever is in the way — in either direction, so a busy neighbour above the slot is no longer a wall (order_moves: upward top-down, downward bottom-up, upward first, with the owner table travelling in the same order). The HTTP bound moved from a slot's share to the whole arena, and CPU + CUDA copy_cells pin an overlapping upward move byte-for-byte. GB10: --n-slots 4 --n-ctx 8192 serves a 2054-token prompt (2048 → 2071 cells, three idle slots released) with a continuation byte-identical to --n-slots 1
Multi-sequence attention masks✔ E1 + E1b + E2 (device-aware default: see the batching row below): the allowed window is an explicit attn_span input resolved from the sequence's span list (C8b S1a/S1b, derived from cell ownership; multi-span layouts are refused until S2's kv_map), read by the CPU kernel and by CUDA's windowed kernel instantiations — device-verified on GB10 since 2026-09-18, and since 2026-09-19 swept over both KV dtypes (f16 and f32), which is what caught the f16 prefill mask fault (ARCHITECTURE-EXECUTION-PLAN.md §14 row 0). Metal reads the one-range attn_span too (#44 part (a), 2026-10-06 — kernel_gqa_attn_window_f32/_f16), so a batched multi-sequence run serves there; only the set-valued kv_map window stays CPU + CUDA (#310)

Gap. 🔴 This is the single largest structural gap, because it blocks four separate user-visible capabilities at once: multi-slot serving throughput, cross-request prompt caching, long-conversation context handling without re-prefill, and any state-space/hybrid model family.

Two concrete defects live here as well:

  • ensure_kv ignores the requested size after the first call (alloc.rs:385-393: the early if let Some(&pair) = self.kv.get(&layer) returns without comparing size). CParams.n_ctx is part of the reuse identity, so a session that changes n_ctx on the same GraphCache silently keeps the old region. The CPU backend then errors on out-of-range positions (cpu_backend.rs:170-172); the GPU backends do not check at all (below). Today no caller changes n_ctx on a live cache, so it is latent rather than live — but it is a landmine directly under the "grow the context" feature that a server wants.
  • GPU KvcacheStore had no bounds validation — closed. A3 moved the check to the backend-agnostic GraphAllocator::fill_input_i32 (check_positions_bound), the single point where positions/cells become graph data, so CPU, CUDA and Metal all refuse an out-of-range row before execution; the Metal gap-table G1 ticket #38 then added the arm-level guard for fill paths that bypass fill_input_i32 (MetalBackend::check_kv_store_rows, covering KvcacheStore, FusedQKV and FusedQkvNorm). A caller violating the documented contract now gets an Err naming the cell and the arena instead of an out-of-bounds device write.

What C1/C2 landed (2026-09-16). src/graph/kvcache.rs now owns a per-layer arena with an owner per cell, and GraphAllocator::kv_rm(start, len, &rope) removes a row range and re-bases the rows after it — a physical operation, so cell == pos survives and no backend needed a new kernel. The conversation's overflow path uses it instead of dropping turns and re-prefilling them. Its one approximation is inherent and recorded in ARCHITECTURE-EXECUTION-PLAN.md §5 (C2 record): the retained rows hold the values they were written with, so rows that attended to the dropped turns keep that influence — no shift that avoids re-prefilling can avoid it (llama.cpp's context shift behaves the same way). What is exact is the mechanism, and that is what the tests pin bitwise: a tail removal leaves the retained head byte-identical (so continuing from it matches a fresh prefill exactly), a middle removal copies V verbatim and re-ropes only K, and the whole operation keeps the identity cell mapping C1's scheduler gate requires.

Recommendation. Redesign the KV layer as a sequence-addressable cell store before adding batched attention, because batching without per-sequence KV addressing cannot be correct. Concretely: a KvCache owning a per-layer arena of n_ctx cells, each carrying the set of sequence ids that own it; the store node resolves (layer, seq_id) → cell index on the host and passes an index array to the kernel; the attention kernel receives an explicit per-query allowed-cell mask instead of deriving the bound from positions. E1 landed that last part on CPU: attn_span carries each query's [lo, hi) cell range (resolved from ownership + the query's position), and a representation as a range is complete because a sequence's cells are contiguous — a per-cell mask is what a hole-creating layout would need (C3/D1). Build one such cache with an optional window parameter and an optional recurrent state, rather than a family of per-variant caches. That single abstraction simultaneously delivers prefix reuse (cells already owned by the matching prefix), eviction (drop cell ownership rather than re-prefill), and defragmentation (a cell-copy op that a strided-view IR makes expressible).


2.5 L5 — Batching and serving 🔴

Today. Single sequence by default on CPU — batching exists and its default follows the device (E6: on for CUDA, off for CPU/Metal; MINFER_BATCH=0/1 forces it), slower on CPU (0.49x) and 1.9x faster on the GPU, where the E2 acceptance is met. GraphParams carries no sequence count at all: E2 deleted n_seqs, which A7 had kept as reserved for item 3, once item 3 landed and showed the count is data rather than topology (CParams.explicit_span carries the only topology decision it can force — see ARCHITECTURE-EXECUTION-PLAN.md §8). The other dead identity field, CParams.n_batch, was deleted in Phase A7; chunked prefill (item 10) will reintroduce it with its real semantics. Attention derives its causal bound from the per-token positions input (cpu_backend.rs:411-416, cuda_backend.rs:1100-1102), so two sequences in one batch would attend to each other.

The prefill is one graph covering the entire prompt (main.rs:862-890): n_ctx = max(--n-ctx, prompt_len), one forward with nt = prompt_len. The server rejects any prompt longer than the slot context (chat.rs:118-124); new_slots (slot.rs:33-36) divides n_ctx_total equally per slot. worker_loop drains the queue serially, one slot at a time (chat.rs:452-518), and every request starts from a fresh GraphCache (chat.rs:77). Default --n-slots is 1 (main.rs:211), so the default server is strictly serial with a cold KV per request.

Gap. 🔴 This is the largest capability difference between what minfer does today and what a serving workload needs, and it costs throughput on exactly the workload the server exists for. With --n-slots N minfer gets N independent serial sessions, not N-way batching: aggregate decode throughput stays at batch-1 tokens/s × (fraction of time the GPU is busy).

Recommendation. Work in dependency order: (1) KV cells → (2) IR seq_id and mask inputs → (3) attention kernels with explicit masks → (4) batch composition in the scheduler/worker → (5) n_batch chunking. Do not start with (4); it cannot be made correct first.

Two smaller but immediate items sit in this layer:

  • Persistent server context. chat.rs:491 discards the GraphCache per request. The comment justifies it as avoiding cross-request KV contamination (a real bug fixed in doc 97) — but the correct fix is sequence-aware KV invalidation (§2.4), not discarding the whole cache. As written, every request pays KV + pool re-allocation (≈235 MB for 7B at n_ctx=4096, f16) and invalidates the CUDA Graph capture warm-up (pool_gen bump, cuda_backend.rs:1342). Fixed in B2: the slot keeps its cache and a record of the tokens its rows hold; a request reuses the KV only when its prompt starts with exactly that sequence. Measured end-to-end (0.5B Q4_0, 219-token second turn): prefill 219 → 16 tokens, time-to-first-token ≈2.8 s → 0.25 s (≈11×), with the cold-slot turn unchanged.
  • No panic isolation. Fixed in A4. The guard that existed (guarded_forward, chat.rs:585-597) only covered forward_graph_cached; the speculative path calls both models' forwards directly (spec.rs) and the tokenizer, sampler and stop-string paths were bare. A panic in any of them unwound worker_loop, dropping the bounded job channel (capacity 64, server/mod.rs:66) — later requests got 503 — and the queued jobs' StreamEvent senders, ending their SSE (Server-Sent Events) stream after an empty [DONE] rather than an error (server/mod.rs:198-216). The whole per-job body now runs under run_job_isolated, which turns a panic into a logged 500 for that request and keeps the worker draining the queue.

2.6 L6 — Backend abstraction and device reach

Today. Backend is a Copy/Hash/Eq/Ord handle over a fixed id space (graph/registry.rs, F4) and a trait (graph/backend.rs:21-98). The ids are a KV-session file-format contract, and the name-keyed registry carries each backend's priority, capability matrix and pool hooks; consumers read it instead of matching. Before F4, adding a backend meant teaching nine #[cfg]-laden match sites about it: GraphAllocator::supports (alloc.rs:384-386), alloc_in_pool/alloc_fresh_in/free_in_pool (:819, :955, :959), sync_backend (alloc.rs:2400), copy_across (alloc.rs:2449), and the scheduler's execute match (scheduler.rs:274-292); the functions are registry-driven now.

GPU participation is decided by an all-or-nothing model-level gate: every weight must be registered on the GPU or the model runs on CPU (ARCHITECTURE.md §5.2; CUDA-BACKEND-DESIGN.md:251-255). The engine enumerates devices and honours --gpu N but uses exactly one (src/cuda/methods/init.rs:63-81); there is no tensor_split, no peer copies, and no remote-device backend.

Gap. 🟠 Two axes.

Pluggability — F4 (2026-09-24) landed the registry half: the backend set is data (a name-keyed table of entries registered at startup) rather than nine match sites, and a name that is unknown, not compiled into this build, or not usable on this machine is a loud startup refusal (BACKEND-REGISTRY-DESIGN.md). What is still closed at compile time is the set itself: there is no dlopen/plugin path, so a Vulkan backend (the only route to Windows/Linux AMD and Intel GPUs) or a remote/RPC backend is added by registering an entry — but it must still be compiled in. That is what remains of this axis.

Placement policy — a model that does not fit in VRAM could not run at all. The docs already record this as a real limitation on 8 GB devices (DEVICE-ADAPTATION-PLAN.md §9.2: "7B Q8_0 will not fit (7.2 GB weights alone)"). E5 (2026-09-23) added the missing granularity and the fit: CNode.layer (stamped by the model builders), an OffloadPlan (graph/offload.rs) that the loader's registration filter, the builder's device-only fused gates and the assignment pass all read, --gpu-layers N / MINFER_GPU_LAYERS=N, the startup report, and a verified mixed CPU+CUDA run on the 0.5B — then S2's auto: per-block weight bytes from the GGUF index, the pure fit_blocks prefix search against a weight budget (MINFER_GPU_MEM, else three quarters of the device's free bytes) with a quarter held back for the KV arenas and the activation pool. The registry/BackendId half is done (F4), and the free-bytes query for Metal landed as #53's DeviceMemory answer (recommendedMaxWorkingSetSize), so auto and headroom_bytes() work on macOS too.

Recommendation. (a) Introduce a BackendRegistry with register(Box<dyn Backend>), supports(op, dtype) -> Option<BackendId> and per-backend allocation/sync as trait methods; replace the enum with an opaque BackendId. Done in F4 (2026-09-24) — the shape that landed is a name-keyed Registry of BackendEntry values (priority + caps + the pool hooks, registered at startup) and a Backend handle over a fixed id space, with the assignment order made an explicit pinned number (docs/BACKEND-REGISTRY-DESIGN.md). It keeps the capability matrices in the trait's own modules (the trait methods forward to them), so the registry and the trait cannot disagree. (b) Add a layer-offload budget policy that consumes the memory accounting from §2.3.


2.7 L7 — Models, quantization, tokenizer, sampling

Architectures. Two (qwen2, qwen3 dense), dispatched on general.architecture (models/mod.rs:105-121). Every other family is absent: MoE, MLA, sliding-window/hybrid, state-space (Mamba/RWKV), multimodal, and every non-Qwen dense family (Llama, Mistral, Phi, GLM, InternLM). The per-family decision, its prerequisites and its cost live in MODEL-SUPPORT-ROADMAP.md; the caveat that matters here is §2.1 — its Tier 1 estimate holds only while a family needs no new IR structure. 🟠

Quantization coverage. Eight quant types; no Q2_K/Q3_K/Q8_K, no I-quants, no MXFP4/NVFP4 (SUPPORT-MATRIX.md §Not Yet Supported; CUDA_OPTIMIZATION.md §1.4 marks IQ/Q2/Q3 "not planned"). BF16 landed as a weight type: the writer and the CPU path in #142, the device halves (CUDA and Metal) in #208, 2026-10-06. The notable one still for reach is Q2_K/Q3_K (running large models on small machines). 🟠 but legitimately deprioritized.

Quantizer tooling — closed by F6 (#49, 2026-09-24), and its f16-weight gap closed by #141 (2026-09-25). A GGUF v3 writer (src/gguf_write.rs), byte-verified weight encoders (src/quantize.rs), a HuggingFace Qwen2 converter that satisfies the strict loader (src/convert.rs), and the convert / quantize / split subcommands (src/tooling.rs, docs/GGUF-TOOLING.md). A converted f16 model runs on the CPU and on CUDA — the CPU graph path gained f16 weight dispatch in F6 and #141 vectorized that dot (AVX2 F16C / NEON FCVTL) and moved the multi-token row loop onto the shared CPU pool (3.2 → 207 tok/s prefill on the 0.5B), while CUDA registers the raw 2 B/element weights and converts in-register. Both supported architectures now: #141's f16 registration and graph type gate landed in qwen2's loader/graph only, so #167 completed the qwen3 pair and gave the two loaders one shared registration rule (models::weight_reg, which also carries the q4_K W_dsc plane gate). Metal refuses f16 weights until #164, and a rewrite/split round-trip is bit-exact, and the HF conversion's f16 output is byte-identical to llama.cpp's converter on the same checkpoint, per tensor. What remains is scope F6 did not claim: the K-quant/I-quant encoders are still refused by name (#140) and bf16 output is refused (#142). 🟢 for conversion/quantization of the supported set.

CPU SIMD. AVX2 (Advanced Vector Extensions 2) dot kernels exist for Q4_0/Q8_0 (src/quants/dot_q4_0.rs:1-9, src/quants/dot_q8_0.rs:28), and since #56 (2026-10-08) the K-quants too: Q4_K/Q5_K/Q6_K × Q8_K have AVX2+FMA kernels (src/quants/avx2.rs) and AVX-512/VNNI variants (src/quants/avx512.rs), dispatched AVX-512 → AVX2 → scalar at runtime (MINFER_NO_AVX512=1 drops to AVX2, MINFER_NO_AVX2=1 to scalar — the x86 counterparts of MINFER_NO_NEON) and gated bitwise against the scalar reference by quants::avx2_correctness; Q4_1/Q5_0/Q5_1 remain scalar on x86 and NEON-only on aarch64. The support matrix is updated (SUPPORT-MATRIX.md, AVX2 column) and CPU_OPTIMIZATIONS.md §P1 is closed. What remains is weight repacking — the standard fix (repack the quant blocks at load time into SIMD-friendly interleaved layouts) that would expose the kernel gain end-to-end, since the current matmul re-reads each weight row per token; there is also still no aarch64 i8mm path and no AMX. 🟡 for x86 (the dots landed; the end-to-end win awaits repacking).

Tokenizer. Byte-level BPE only, loaded from GGUF metadata (tokenizer.rs). F7 (#50, 2026-09-24) made tokenizer.ggml.pre authoritative with two hand-written rule sets (qwen2, qwen35), replaced the silent unwrap_or(0) with a checked byte fallback, and made every other pre-tokenizer value (including a missing one), a non-gpt2 model, an empty merge table and an incomplete byte vocabulary a loud load refusal. What remains: no SentencePiece/unigram or WordPiece path, and no ignore_merges/multi-regex pre-tokenizers (llama3, default, deepseek-*, …) — those refuse by name, so the coverage gap is explicit rather than a wrong split. Also no NFC normalization (llama.cpp does not apply one for BPE either) and special-token matching is still a per-model hand-extension (the DeepSeek-R1 fix noted in AGENTS.md), which does not scale. 🟡 for the Qwen/Llama-3-shaped families this engine supports; 🟠 for anything SentencePiece-based. Follow-up: #132.

Chat templates. Landed in F7 (#50, 2026-09-24). minijinja 2.21.0's set_unknown_method_callback is the extension point, so Qwen3's chat_template renders (think-block extraction, tool-call formatting, enable_thinking) with no dependency change; the reference renderings are committed and byte-for-byte gated, and any template the engine cannot render is a loud refusal naming the construct and line — validated at load, so the CLI exits and the server refuses to start. What remains: a --chat-template <FILE> override (a user-supplied template for a GGUF that has none or a wrong one), strftime_now, and the Python str methods outside the implemented set (each refused loudly). Design + accepted/refused sets: docs/CHAT-TEMPLATE-AND-TOKENIZER-DESIGN.md. 🟢 for the supported models.

Sampling. One SamplerConfig pipeline in sampler.rs (#48, landed 2026-09-24): logit bias → penalties → DRY → greedy shortcut → top-k → typical → top-p → min-p → XTC → temperature or mirostat v1/v2, with every new knob defaulting to a no-op so the pre-F3 chain is bit-identical. Landed from the old gap list: min-p, typical, XTC, DRY, mirostat v1/v2, logit bias (and frequency/presence penalties from the OpenAI plan). Grammar / JSON-schema constrained decoding landed in F2 (2026-09-24): src/grammar.rs compiles a GBNF grammar (or a JSON Schema, via a generated GBNF) into a pushdown automaton and masks the logits inside that one pipeline, between DRY and the greedy shortcut; CLI --grammar/--json-schema, server response_format (json_object / json_schema) plus a grammar field. Design and the exact accepted subset: docs/GRAMMAR-DESIGN.md. What it deliberately does not do (and refuses loudly rather than guessing): pattern/minLength/maxLength, multipleOf, real-valued numeric bounds, object property permutations, and a device-side mask — the mask is host work before any backend runs. Still missing: top-n-sigma, adaptive-p, infill. 🟡

No LoRA (Low-Rank Adaptation) / adapter support. There is no adapter path at all. minfer's weights_version field in GraphParams (params.rs:88-98) exists precisely to break reuse on a weight swap, but nothing produces such a swap. 🟡


2.8 L8 — Ops coverage, safety, testing, observability

Op coverage. The Op vocabulary is deliberately small; six variants are parity-only stubs that no architecture emits (Scale, Softmax, View, Reshape, Permute, AttnMode::Mha — COMPUTE-GRAPH-DESIGN.md §1.3), and View/Reshape/Permute execute as copies. CUDA additionally refuses transpose_b matmul (cuda_backend.rs:1454-1459) and FusedQkvNorm (cuda_backend.rs:1905), and Metal refuses QkvBiasRopeStore (metal_backend.rs:700) — so the two GPU backends do not implement the same op set, and a model's decode path differs by platform. Neither is documented in SUPPORT-MATRIX.md. Fixed in A8: SUPPORT-MATRIX.md now carries an "Operator Coverage by Backend" table with the four asymmetric rows and their consequences.

Guard asymmetry. Both halves are closed on a Mac (2026-10-06). The decode-fusion debug_assert!s (#39, gap-table G2): FusedFFN / FusedQKV / FusedQkvNorm each return Err naming the node and the observed nt before the weight lookup, so a release build no longer passes an invalid shape through. The weightless-RMSNorm substitution (#40, gap-table G3): Op::RmsNorm / Op::QkNorm now call MetalBackend::norm_weight, which returns Err naming the node and the missing tensor for both None meanings (weight_name absent, or unregistered on the device) instead of running the weightless kernel where CUDA errors (cuda_backend.rs norm_weight). Neither violates the project's rule that "kernel-invariant violations return Err, never a silent fallback" (AGENTS.md §GPU Safety) any more.

Testing. Five integration files, four of them #![cfg(target_os = "macos")] Metal kernel isolation tests; CPU/CUDA correctness rests on inline unit tests plus real-model tests that skip when a GGUF is not cached (tests/conversation_cli.rs:1-12). CI is cargo build --release on macOS only — no test run, no CUDA build, no Linux build. Fixed in A2: CI now runs cargo test --release on Linux/CPU, compiles the CUDA backend in NVIDIA's devel image, and keeps the macOS build. 🟠 The remaining highest-value addition is a systematic op × dtype × backend correctness matrix (every op checked on every backend against a CPU reference); it would have caught several items in §4 automatically — that is ticket A1.

Observability. F8 landed 2026-09-24 (#51): GET /metrics (Prometheus text) exports live KV/arena occupancy, queue depth and the running/in-flight counts; per-op timing is available under MINFER_OP_TIMING (off by default, measured at the scheduler's per-node dispatch); and SIGINT/SIGTERM drains gracefully, bounded by MINFER_DRAIN_MS. Still missing: structured logging and levels (the server prints free-form startup lines), and worker-restart supervision (a panicking job is isolated per request by run_job_isolated, but a dead worker thread is not restarted). The trace/viz stack remains a development instrument, not an operations one. 🟡 for a research engine, 🟠 the moment it is deployed.


3. Recommendation backlog

Effort: S ≤ 2 d · M ≤ 1 w · L ≤ 2 w · XL > 2 w. Backlog items tracked on GitHub link their issue (the open list); an item whose work predates the tracker has no issue, and the plan is its record.

P0 — structural, unblocks a class of use cases

#ItemRefsEffort
1KV cache → sequence-addressable cell store (cells + seq-id sets; host-resolved (layer, seq) → cell indices; explicit per-query mask passed to attention). Prerequisite for everything in P0. — C1/C2/C3 landed (cell store, physical removal/shift, compaction), C6 landed 2026-09-20 (001b8cc: sequence-relative positions + an allocator-resolved cells input, so compaction is bit-identical) and C7 + C7b landed 2026-09-20 (elastic partition: a long request reclaims idle capacity and the planner moves runs in both directions, so a busy neighbour cannot block it; the continuation stays byte-identical to a one-slot engine). Remaining: C8 (#41, closed 2026-09-22), whose design splits it into C8a ✔ (shared prefill with copied rows; measured 36.6× cheaper than the re-prefill it replaces) and C8b (true sharing: a per-sequence span list plus an additive kv_map the attention read path gathers from — S1a/S1b landed 2026-09-21 (write path and read path both resolve through the list; nothing observable changed), S2 landed 2026-09-21 (admission shares a prefix in place — no bytes copied — and the CPU kernel gathers the map; block refcounts were not added, occupancy stays derived from the span lists), S3 landed 2026-09-22 (copy-on-write: a store inside a shared prefix rebases the sequence's run and shifts its own rows up, so a shared block is never written through), S4 landed 2026-09-22 (CUDA gathers the map in every attention path, FA prefill included, at 1.001x the span on decode and 1.1x on prefill; admission switches on Device::gathers_attn_map) and S5 landed 2026-09-22 (Metal refuses both window layouts loudly at assignment and at execution, so C8b is closed on CPU and CUDA — superseded: the G5 port and the Metal kv_map kernel of #362 lifted the refusal, so Device::gathers_attn_map is true for Metal and admission shares in place there too)§2.4XL
2IR seq_id + attention-mask inputs; attention kernels take an allowed-cell mask instead of deriving the bound from positions. — done (E1: span input + resolver + CPU kernel + two-sequence test; E1b: CUDA windowed kernels, compile-verified)§2.5L
3Batch composition + continuous batching in the scheduler and server worker; make n_seqs real (or delete it). — mechanism landed in E2: sequence-addressable batches, per-sequence reservations, the server composes one decode batch and batched prefills; n_seqs deleted (it turned out to be data, closing A7); opt-in (MINFER_BATCH=1): the payoff is device-dependent — 0.49x on CPU (7B Q4_K_M) but 1.9x on the GB10 GPU (measured 2026-09-18, after A0's "no device" verdict turned out to be an agent-sandbox artefact); the CPU refutation closed the ticket and the GPU result is the acceptance the closure predicted; the device-aware default landed in follow-up ticket E6 (2026-09-19: unset batches iff the model runs on CUDA, with MINFER_BATCH=0/1 to override — 1.97x on the 7B with no environment variable); see the plan's E2/E6 records§2.5XL
4Persistent server context: keep the GraphCache across requests, invalidate per sequence id; stop re-allocating KV and re-warming CUDA Graph capture per request. — done in B2/B3 (prefix-matched reuse, ≈11× TTFT on turn 2)§2.5M
5Fix ensure_kv size handling and add the missing pos < n_ctx guard on both GPU backends.§2.4S
6Worker panic isolation (catch_unwind + supervision + an error event instead of a silent empty stream). — done in A4§2.5S

P1 — material capability or performance

#ItemRefsEffort
7IR expressiveness: strided views with allocator-known aliasing, multi-output nodes; then re-express the four decode fusions as compositions. — DONE (D1, 2026-09-19), increments 1–3: exact views are zero-copy with allocator-known aliasing (CNode.view + liveness + no-op kernels), offset/partial windows work on CPU and CUDA (BufRef carries offset+len through the Backend trait; Metal is exact-only until G5), and the "multi-output node" is GraphBuilder::split_parts — one owning node plus one Op::View per part, so a producer's single output feeds several consumers with independently bindable tensors while the graph stays single-output. D2's windows are in production; the sketched Op::SplitParts was deliberately not added (a split op over one input is just views of that input). Item 17 (MoE) therefore no longer waits on an IR blocker — what it needs now is model work (expert weights, routing, the grouped GEMM), and item 16 (MLA) likewise§2.1L
8Allocator reserve/assign split + size-class rounding + real memory accounting + VRAM feasibility gate. — S1 landed 2026-09-22: the size-class ladder and a pure plan (graph/allocplan.rs), per-backend accounting (memory_report: weights/pool/live/peak + headroom_bytes), and a feasibility gate that refuses an over-budget graph before the pool is touched, naming its numbers. S2 landed 2026-09-23: the pools allocate at class_size(size) (two shapes in one class share a buffer across a rebuild), the node's logical length moved into BufRef (write_host_window, windowed capture reads, fills checked against BufRef::len), pool_bytes counts what the pool holds, all input buffers are placed before the walk (an input is filled before execution, so a released buffer would be clobbered by its previous owner's write), and a liveness extension now moves the buffer's buf_alive deadline — D1's view branch never ran at all. S3 landed 2026-09-23: reservation split from assignment (a slots table per backend/class; a released classed buffer is re-assigned, not returned to the backend), so a rebuild re-maps — CUDA's pool_gen stops moving and a captured graph survives it — and GraphCache holds one graph per GraphParams (a switch re-maps; a repeated chunked prefill builds nothing: 3 builds/5 reuses then 3/9) · #55§2.3L
9Layer offload policy on top of (8); needs a layer-granular assignment pass. — S1 landed 2026-09-23 (E5): graph/offload.rs (OffloadPlan + the pure resolver), CNode.layer + GraphBuilder::set_layer, GraphAllocator::supports_for refusing the device past the plan, per-block weight registration (OffloadPlan::allows_weight, the block parsed from the registry name), --gpu-layers/MINFER_GPU_LAYERS, CParams.gpu_layers in the reuse identity, the startup report, and the mixed-run gate (4/24 blocks on CUDA matching the all-CPU greedy tokens, one device split per offloaded block with cross-backend copies at the boundaries). S2 landed 2026-09-23: the auto request fits the largest block prefix into a weight budget (MINFER_GPU_MEM, else three quarters of the device's free bytes — the same default E4's gate uses) with a quarter held back for KV/activations; per-block bytes come from the GGUF index (measured before anything is registered, because the filter is the plan); the startup line explains the fit. Real-model gate: MINFER_GPU_MEM=64 → 5 of 24 blocks, 40.0 MiB on the device, greedy tokens matching the all-CPU run; uncapped → all 24 · #46§2.6L
10Chunked prefill: make n_batch real; cap activation memory and allow decode/prefill interleaving. — landed in E3 (2026-09-22): prefill_chunks + MINFER_N_BATCH (default 2048, a no-op for prompts that fit), remainder-last so the final forward carries the tail row, and the other slots take their decode step between chunks. Measured: 5 forwards / max nt 24 vs 1 / 98 for a 98-token prompt; logits bitwise on CPU and ≤ 0.218 (class 1.0) on CUDA; the interleaving A/B is 3 decode steps vs 0; the split costs one forward's fixed overhead per chunk (514 tokens in 4 forwards: CUDA 1.11x, CPU 1.005x; 98 tokens in 5: CUDA 2.0x). Not mixed prefill+decode batches yet · #45§2.5M
11CPU AVX2 (and AVX-512/VNNI where available) for the K-quant dots; then weight repacking.§2.7L
12Backend registry decoupling the enum from the nine match sites. — landed in F4 (2026-09-24): #57; the handle + the name-keyed table live in graph/registry.rs, the priority order is a pinned number, and the name surface (--backend / MINFER_BACKENDS) has three distinct loud startup refusals (docs/BACKEND-REGISTRY-DESIGN.md). The registered set itself stays compile-time (no dlopen)§2.6M
13Guard symmetry: Metal Err instead of debug_assert!/weightless fallback; CUDA gains FusedQkvNorm or SUPPORT-MATRIX.md gains a per-backend op column. — docs route done in A8; both Metal halves closed: the debug_assert! half in #39, the weightless-RMSNorm half in #40§2.8S
14Async cross-backend copy + events (needed for any heterogeneous split and for multi-device execution). — landed in F5 (2026-09-24): #58; the boundary's two registered phases (BackendEntry::copy_cross/await_cross), CUDA's cudaMemcpyAsync D2H + event + one cudaEventSynchronize at the documented point, the CPU's documented no-op, counter-gated (graph/copystats.rs: zero blocking boundary copies, one wait per copy) and bitwise-identical to the MINFER_SYNC_COPIES reference on a real split model. The deferred wait landed in #138 (2026-10-04): the boundary only enqueues, each staged input's single wait is issued at the consumer's first read (and drained after the last split), and the boundary's redundant sync_backend is gone for a backend that can order its close on its own stream (Backend::retire) — measured 21 → 0 full stream syncs over the 0.5B gate's 7 forwards and 2 copies in flight where the enqueue-then-wait order holds 1, still bitwise. Still open: Metal declines (unported, no Mac to verify) and true cross-split overlap (a later split's transfer beside an earlier split's kernels)§2.2M

P2 — coverage

#ItemRefsEffort
15Constrained decoding: GBNF-style grammar + JSON-schema → grammar. — landed in F2 (2026-09-24): #47, src/grammar.rs + the mask in src/sampler.rs; the design record is docs/GRAMMAR-DESIGN.md§2.7M
16Sampler set: min-p, typical, XTC, DRY, mirostat, logit bias. — landed in F3 (2026-09-24): #48§2.7M
17MoE support (see MODEL-SUPPORT-ROADMAP.md Tier 2 #1) — depends on item 7 for a clean implementation.§2.1, §2.7L
18Dense architecture port wave (see MODEL-SUPPORT-ROADMAP.md Tier 1) — parameter mapping plus the per-port items listed there.§2.7M–L
19Chat-template fidelity: replace or extend minijinja so Qwen3's template actually renders. — landed in F7 (2026-09-24): #50, the Python-str-method hook + the loud refusal; design record docs/CHAT-TEMPLATE-AND-TOKENIZER-DESIGN.md; remains: --chat-template <FILE> + strftime_now (#133)§2.7M
20Tokenizer generality: SentencePiece/unigram, per-model pre-tokenizer variants, data-driven special tokens. — partly landed in F7 (2026-09-24): #50 made tokenizer.ggml.pre authoritative (qwen2/qwen35, byte-for-byte gated, unknown values refused) and replaced the silent byte fallback; SentencePiece/unigram/WordPiece and the remaining pre-tokenizer rules are #132§2.7M
21Quantized KV (Q8_0 first) — after item 1. — landed in C4 S1 + S2a on the CPU and S2b on CUDA (fused read, quantize-aware shift, layout-tagged kernels); the packed fused epilogue, the dp4a packed dot and the packed FA prefill followed (#144 items 1+3, #186, #202), so what remains is the CUDA decode residual's stall attribution (#212) and Metal's packed half (#310, enabled)§2.4M
22Quantizer tooling: convert-hf-to-gguf + quantize + split.§2.7L

P3 — hygiene and operations

#ItemRefsEffort
23Op × dtype × backend matrix test. — done in A1§2.8M
24CI: run tests on macOS, add a Linux CPU job, add a CUDA build job. — done in A2§2.8S
25Metrics/observability: /metrics, KV occupancy, queue depth, per-op timing under a flag, graceful drain. — done in F8 (#51, 2026-09-24; the only member of the A-era batch that had no ticket). Structured logging/levels and worker supervision were not part of the ticket and remain open in §2.8§2.8M
26Remove the dead identity fields: delete CParams.n_batch; keep GraphParams.n_seqs marked reserved for item 3 (decision recorded in ARCHITECTURE-EXECUTION-PLAN.md §8). — done in A7; fully closed in E2, which deleted n_seqs too after item 3 showed it was redundant, not just unread§2.5S
27Re-key the cross-backend staging map by (node, dst_backend). — done in A5§2.2S
28CPU per-op allocations: cpu_backend.rs:157-158 clones the K/V sources on every store node and :195 allocates a Vec<&[f32]> per node. — closed in A6 as not worth doing: the allocation removal measured −1.2 % prefill / −1.8 % decode and was reverted (the loop is weight-streaming bound)§2.3S

4. Concrete defects found (verified in code)

Ordered by severity. Items 1–6 are behavioural; 7–12 are hygiene; 13–14 are behavioural defects found while executing the plan, already fixed.

  1. GPU KV store bounds check — closed. A3 moved the check to the backend-agnostic GraphAllocator::fill_input_i32 (check_positions_bound), so CPU, CUDA and Metal all refuse an out-of-range row before execution, and #38 added the arm-level Metal guard (MetalBackend::check_kv_store_rows) for the fill paths that bypass it — §2.4 states the same closure in the gap list.
  2. ensure_kv ignores a changed size (alloc.rs:995-1002). The KV region is frozen at first allocation while CParams.n_ctx remains part of the reuse identity, so a size change is neither honoured nor detected.
  3. Metal weakens kernel guards. Both halves are closed: the decode-fusion debug_assert!s (#39) return Err naming the node and the observed nt, and the silent weightless RMSNorm when a weight is missing (#40, MetalBackend::norm_weight) returns Err naming the node and the missing tensor instead. Neither contradicts docs/GPU_SAFETY.md and AGENTS.md's no-silent-fallback rule any more.
  4. Server worker had no panic isolation outside the forward call (src/server/chat.rs::guarded_forward_batch): a panic anywhere else unwound the worker, permanently degrading the server (503 for new jobs, empty 200/SSE for queued ones) with no log. Fixed in A4 — the whole per-job body is now guarded.
  5. Backend op-set asymmetry drives silent path changes. FusedQkvNorm is Metal-only (MetalBackend::supports_op) but absent from CUDA's supports_op, so Qwen3 decode takes the fused path on Metal and the unfused path on CUDA. QkvBiasRopeStore is the mirror case (MetalBackend::supports_op returns false for it). Neither is documented in SUPPORT-MATRIX.md. Documented in A8: SUPPORT-MATRIX.md now has an "Operator Coverage by Backend" table; the asymmetry is visible rather than silent.
  6. CUDA RoPE is non-interleaved only (cuda_backend.rs:1522), so any model needing the interleaved style splits every layer between CUDA and CPU, producing two host round trips per layer (§2.2); docs/ARCHITECTURE.md:455-459 now states that limitation rather than advertising both styles.
  7. Stale unreachable!("CUDA pool not implemented") in the non-CUDA arms — the string and any unreachable! are gone from alloc.rs. Fixed in #57 (cdf41b2).
  8. Dead fields in the reuse identity: CParams.n_batch and GraphParams.n_seqs are compared by params_match (cache.rs:95-101) but no builder reads them; every construction site hard-codes 1 / n_tokens. Fixed in A7, closed in E2: n_batch is deleted; n_seqs was kept for item 3 and then deleted by it (the sequence count is data — the topology decision lives in CParams.explicit_span), with sequence_count_is_data_not_topology pinning that a sequence-count change no longer rebuilds an otherwise identical graph.
  9. Single-entry cross-backend staging (GraphAllocator::cross_staging), mitigated by the consumer-side filter at scheduler::cross_input_ready. Fixed in A5 — keyed by (graph uid, node, dst_backend); two foreign consumers can now be served.
  10. read_host returns None on CUDA (cuda_backend.rs:2268), so the trait's host-read contract is backend-dependent; the allocator compensates with copy_to_host (alloc.rs:2137-2143).
  11. CUDA pool never releases device memory (cuda_backend.rs:2148-2156), documented as accepted debt but reasoned about for fixed-shape CLI runs; a varying-n_tokens workload accumulates one buffer set per distinct shape.
  12. Tests: no Linux/CUDA/CPU in CI, four of five integration files macOS-only (.github/workflows/ci.yml, tests/*.rs). Fixed in A2 for the CI half (Linux/CPU tests, CUDA compile, plus --no-run test compilation); the four macOS-only kernel-isolation files remain macOS-only by nature.
  13. Op::Softmax returned unnormalised values on the CPU backend (cpu_backend.rs): it called vec_soft_max_f32, which writes exp(x - max) and returns the sum, and discarded the return. No architecture emits a standalone Softmax node, so nothing exercised it. Found by the A1 op matrix and fixed: the arm now scales by 1/sum, with an op-matrix case pinning it.
  14. The CPU attention was not nt-invariant. Each token's scores were padded to the batch-wide nkv, and the softmax, the normalisation and the weighted sum all ran over nkv instead of the token's own causal window vl = pos+1. The padded entries are -inf → 0, so the arithmetic reads as equivalent — but the reduction length, and with it the rounding, depended on how many tokens shared the batch. Measured on 0.5B: a one-token decode step differed from a single-shot prefill by max|Δlogits| = 0.41, and the same token's K rows differed from layer 3 on (nt=6 vs nt=13). Found while testing prefix reuse for B2 and fixed: restricting the window to vl makes incremental prefill, the decode loop and a single-shot prefill bitwise identical, and removes the padding pass. It also matters beyond B2 — the speculative-decoding identity gates assume exactly this property.

5. Suggested execution order

The dependency structure is more informative than the priority table alone:

                       ┌──────────────────────────────────────────┐
                       │ 1. KV cell store (seq-addressable)       │
                       └───────────────┬──────────────────────────┘
                                       │
        ┌──────────────────────────────┼───────────────────────────────┐
        ▼                              ▼                               ▼
 2. IR seq_id + mask         4. persistent server context    21. quantized KV
        │                              │
        ▼                              │
 3. continuous batching  ◄──────────────┘   (4 alone already removes
        │                                    per-request realloc/re-warm)
        │
        ├──► 10. chunked prefill
        └──► 9. layer offload ◄── 8. allocator reserve + accounting

  7. IR views/multi-output ──► 17. MoE, 11. CPU SIMD (independent)
  12. backend registry      ──► 14. async copy ──► multi-device execution

Items 5, 6, 13, 26–28 are independent, small, and can land at any time — they are the ones worth doing first simply because they are cheap and they remove hazards that the larger work would otherwise have to work around.


Appendix — evidence index

AreaPrimary sources read
IR / builder / opssrc/graph/mod.rs, ops.rs, builder.rs
Schedulersrc/graph/scheduler.rs
Allocatorsrc/graph/alloc.rs
Reuse / paramssrc/graph/cache.rs, params.rs
Backend traitsrc/graph/backend.rs
Backend registry (F4)src/graph/registry.rs, docs/BACKEND-REGISTRY-DESIGN.md
CPU executionsrc/graph/cpu_backend.rs, src/kernel.rs, src/quants.rs
CUDA executionsrc/graph/cuda_backend.rs, src/cuda.rs, src/cuda/kernels/*.cu
Metal executionsrc/graph/metal_backend.rs, src/metal/, src/metal/kernels/
Modelssrc/models/mod.rs, models/qwen2/graph.rs, models/qwen3/graph.rs
Servingsrc/server/{mod,chat,slot}.rs, src/conversation.rs
Samplingsrc/sampler.rs, src/tokenizer.rs, src/template.rs
Speculativesrc/spec.rs
DocsAGENTS.md, README.md, docs/{ARCHITECTURE,COMPUTE-GRAPH-DESIGN,MODEL-SUPPORT-ROADMAP,SUPPORT-MATRIX,METAL_OPTIMIZATIONS,CUDA_OPTIMIZATION,CUDA-BACKEND-DESIGN,DEVICE-ADAPTATION-PLAN,OPENAI-CHAT-API-PLAN,SPECULATIVE-DECODING-PLAN}.md

minfer Architecture Execution Plan

Status: Phase A complete (9/9, 2026-09-16); Phase B complete (3/3, 2026-09-16); Phase C complete (8/8) — C1, C2, C3, C4 (quantized Q8_0 cache, CPU), C5 (session save/restore), C6 (logical positions), C7 (+C7b) and C8 (cross-sequence cell sharing) are done. C6 merged 2026-09-20 as 001b8cc; C7 landed 2026-09-20 (the partition is elastic, and growth moves runs in both directions, so a busy neighbour above the slot no longer blocks it); C8 split into C8a (shared prefill, duplicated rows: no IR change) and C8b (paged sharing: a block map and a gather in every attention kernel — S1a/S1b/S2/S3/S4/S5 landed, closed on all three backends: Metal's share path landed after the G5 port, in #362); C4 and C5 landed 2026-09-22 (C4's fused dots and the CUDA kernels are #87, Metal's packed read #310; the CLI/server surfaces C5 enables are #89). Phase D complete (3/3) (D1 done: views, multi-output via split_parts, D2, D3); Phase E complete (7/7) (E1, E1b, E2, E3, E4, E5, E6 all done); Phase F in progress (7/8) (F2, F3, F4, F5, F6, F7, F8 done; F1 waits for an x86 host); Phase G complete (7/7) — G1–G7 all landed on a Mac (the KV port G5 and the measurement bookend G7 last; #54 recorded the round's final baseline). Next: the open Linux-side tickets on dgxspark (aarch64, GB10 sm_121), beginning with #428's three fixture producers, and F1's weight-repacking increment, which still wants an x86 host. The order was deliberate: the Metal KV port (G5) came after the CUDA arena stopped changing shape (C7, C7b, C8), so those semantics were written into Metal once — it landed on a Mac 2026-10-06. Per-ticket evidence is in each phase's record and in the §14 open-risks table. Companion to: docs/ARCHITECTURE-ROADMAP.md (what is missing, why, and how it is ranked). This document is the how: phase-by-phase tickets with deliverables, acceptance criteria and dependencies. Baseline: HEAD = f32daa7 (2026-09-16); Phase A landed on architecture-phase-a (PR #1). This status was refreshed against master = 53b3ace (2026-10-06, the macOS/Metal round's final master); it is refreshed when a phase or round closes, not on every PR. Derived facts: the phase counters above, the next: sentence and the two commit ids are not hand-maintained prose — scripts/status.toml is the source of truth, and scripts/check_status.py --check (CI job check-docs) fails when this block disagrees with it, naming the file, the line and both values. Edit the source, not the counter. The suite counts are the sibling ledger scripts/test-baselines.toml, read by scripts/check_baselines.py: its --check-live compares the one CI-verifiable box against the test-linux-cpu log, while the CUDA / real-model / sanitizer rows are labelled recorded measurements because CI has no GPU. The per-ticket ✔ marks in the §11 diagram are not derivable from (done, total) and are deliberately out of scope for the checker.

Issue links. Work tracked on GitHub carries its issue link in its table row (prose sections carry it in the heading), and the record written when that work lands keeps the link. Tickets that predate the tracker — Phase A and B, C1–C3, C6/C7/C7b, C8a, D1–D3, E1/E1b/E2/E6 — have no issue, and their record here is the only one; the tickets below that do have one are exactly the open list in the tracker. Doc debt and test health are filed the same way: #62 (docs/USAGE.md staleness). The docs build was #63 and is closed: CI job check-docs builds the book with the same toolchain Docs deploys with and runs scripts/check_docs_links.py, so a renamed file now fails on the PR that renames it instead of rotting silently (it found four dead links the day it landed). Test health was #82 — the three #[ignore]d real-model tests that failed on master — and it is closed: two were writing into a directory nothing created, and the third asserted a model behaviour (the greedy 0.5B stopping on EOG within 16 tokens) instead of the engine's rule that ties need_insert_eot to the stream.

0. Decisions already taken

DecisionConsequence for this plan
Metal is out of scope this round — superseded: ADR-0005 reversed it on 2026-09-20 and Phase G is complete (7/7).Recorded as ADR-0003; no ticket here needs the old "defer to Phase G" marking any more.
Dead reuse-identity fields: option (a), then (c) — the decision is ADR-0002.A7 deleted CParams.n_batch; E2 then deleted GraphParams.n_seqs too — item 3 landed and showed the sequence count is data, not topology (A7 closed, rationale in §8).
Phase A (A0–A8) is complete (2026-09-16, PR #1); Phase B (B1–B3) is complete (2026-09-16, PR #1); Phase C's C1 and C2 are complete (2026-09-16, PR #2); E1 and its CUDA half (E1b) are complete (2026-09-17, PR #3) — E1b was compile-verified and SASS-checked then, and is device-verified since 2026-09-18 (see its record); E2 landed and is closed (2026-09-17, PR #4; re-measured on the GPU 2026-09-18): mechanism in, A7 closed by deleting n_seqs, acceptance refuted on CPU (0.49x) and met on GPU (1.9x) — the sign of the effect is a property of the device. The CUDA device is available from 2026-09-18 (A0 superseded); E1b's windowed attention is device-verified and its causal path is timing-neutral, and the A1 matrix's CUDA column now runs on hardware.The next work is C3 + D1 increment 3 (the explicit cell-copy op C3 needs, and multi-output nodes — also the MoE/MLA prerequisite), then C4/C5; E6 settled the batching default (device-aware: batches iff the model runs on CUDA or — since #44 part (b), 2026-10-06 — Metal, off on CPU, MINFER_BATCH=0/1 to force either way), and §14 row 0 closed the GPU-batching correctness blocker behind it (the f16 windowed FA prefill mask, fixed 2026-09-19); a CPU nt>1 decode kernel (F1 family) remains the only CPU route to the throughput claim; Phases D–G remain planned.

1. Standing rules

Every ticket is rejected if it breaks one of these. They restate the project's own invariants rather than inventing new ones.

  1. Params-only reuse. Anything that changes graph topology must join GraphParams / CParams and be compared in GraphCache::params_match (src/graph/cache.rs:57-64); n_past never enters the identity. Decision: ADR-0002.
  2. No silent fallback. A kernel-invariant violation returns Err from execute_node with the actual values; assignment is decided at build time. Decision: ADR-0009.
  3. Identity gate. Any change to a kernel or an execution path is A/B'd against the existing path. Bitwise-identical is the default bar; where that is impossible, the ticket must name the tolerance class and its cause (the project's existing example: nt ≤ 8 bitwise, nt = 9 tolerance-class). Decision: ADR-0010.
  4. GPU safety. Bounded waits, status checks, device limits queried at runtime, no hardcoded device constants. Decision: ADR-0008.
  5. Index docs move with the code. README.md, AGENTS.md, docs/SUPPORT-MATRIX.md and docs/ARCHITECTURE-ROADMAP.md are updated in the same commit as the change they describe.
  6. Deferred-Metal marking — retired. A ticket whose cross-backend design changed Metal had to add a Phase G line in the same commit. Metal is in scope since ADR-0005 (2026-09-20), so the rule has no subject.

2. Verification matrix (what dgxspark can prove)

dgxspark — the box every recorded measurement in this plan was taken on — is a DGX Spark (GB10), aarch64 Linux, CUDA 13.0 (/usr/local/cuda-13.0), with cached Qwen2.5-0.5B/7B/14B and Qwen3-0.6B GGUFs. Measured records name it absolutely, never as "this box" (gate contract rule 5).

BackendBuildRun/verifyNote
CPU (aarch64 NEON+SDOT)✅✅Primary correctness net here
CUDA (sm_121)✅✅ device since 2026-09-18GB10 (121.6 GiB, driver 580.178.04, CUDA 13.0). Device-gated tests are still local-only: CI has no GPU, so its CUDA job only compiles the harness (F-campaign note in §14 row 1)
Metal❌❌macOS-only code, not compilable here → Phase G
x86 AVX2 / AVX-512❌❌dgxspark cannot build/run x86. Item 11's dots were verified on an x86 host 2026-10-08 (F1 landing record in §9); weight repacking is open

A0 verdict (2026-09-16) — superseded (2026-09-18). It read "CUDA is compile-only in this environment", and for two days every CUDA ticket's acceptance was "compiles + reviewer-inspected". The reading was wrong: those probes ran under an agent file sandbox whose Landlock rules denied open() on /dev/nvidia* even though the nodes exist and are world-writable, so cuInit failed with err 304 for a reason unrelated to the driver. The device has been available since 2026-09-18, and CUDA work is now measured on it — E1b's windowed attention, E2's 1.9x, and the f16 windowed-prefill fix recorded in §14 row 0 all carry device evidence.

The design consequence worth keeping is about coverage, not capability: CI has no GPU (§14 row 1), so a device-gated assertion is a local, manual run. Where a guard can live in the backend-agnostic layer (the allocator) instead of cuda_backend.rs, put it there so CI still proves it — that reasoning is independent of whether a device happens to be present (applied in A3, and again by C3 below).

3. Phase A — instrument, then hazard removal

Why first: the roadmap's §5 calls items 5, 6, 13, 26–28 "cheap and independent". This plan pulls item 23 (op matrix) and item 24 (CI) to the front of that batch — they are the instrument that keeps Phases B–E honest, and one of them (the op/dtype/backend matrix) would have caught several of the roadmap §4 defects automatically.

IDItemTitleEffortStatus
A0—CUDA access spike on dgxsparkS✅ done — but the verdict is superseded (2026-09-18): the device is available; "unavailable" was an agent-sandbox artefact (§2)
A123Op × dtype × backend correctness matrixM✅ done — found + fixed an op defect
A224CI: test on Linux/CPU, build on CUDA, keep macOS buildS✅ done
A35KV bounds guard + ensure_kv size checkS✅ done
A46Server worker panic isolationS✅ done
A527Re-key cross-backend staging by (node, dst_backend)S✅ done
A628Remove CPU per-op allocationsS✅ measured — refuted, reverted
A726Dead identity fieldsS✅ done
A813Guard symmetry (docs half + CUDA FusedQkvNorm)S✅ done (docs route)

A0 — CUDA access spike — SUPERSEDED (2026-09-18): the device is available

  • Original verdict (2026-09-16): cargo build --release --features cuda succeeds (1m23s, targets sm_75…sm_121, PTX compute_121), but at runtime cudaGetDeviceCount returns err 304 and the engine logs CUDA: no CUDA devices found (cudaGetDeviceCount err 304, count 0) then CUDA: not available, using CPU fallback. The CPU path was unaffected (Qwen3-0.6B Q8_0: 120 tok/s prefill, 64.7 tok/s decode).
  • Correction (2026-09-18). The agent's execution sandbox is enough to produce every symptom A0 recorded, on a device that works. Under the harness's default file sandbox (Landlock) every open("/dev/nvidia*") returns EACCES even though the nodes are crw-rw-rw-, so cuInit fails with 304 and NVML prints Failed to initialize NVML: Unknown Error — the exact signature A0 recorded. With the sandbox widened (2026-09-18, after a reboot that also cleared a driver-upgrade state the maintainer had flagged): NVIDIA GB10, sm_121, 121.6 GiB, driver 580.178.04, CUDA 13.0, cuInit → CUDA_SUCCESS. Every A0-era probe was therefore run inside a sandbox that cannot reach a GPU even when the GPU is healthy, so those probes could not establish anything about the driver's state: "no device" was unsupported rather than merely pessimistic. (The maintainer recalls a driver upgrade without a reboot at the time; that account and the sandbox are both consistent with the record, and only the sandbox is reproducible today — which is why the correction is to re-run the verification, not to assume it always would have passed.) What is certain now: the CUDA half of every ticket between A0 and this session was compile-verified only, and in this session it is device-verified (see the E1b record's device section).
  • Consequence for the plan: tickets that were closed "compile-verified only" because of A0 are re-opened as verification, not as code: E1b's kernels, the A1 matrix's CUDA column, and the CUDA-side test suite all get their first execution on hardware below.

A1 — Op × dtype × backend matrix · item 23 · M — DONE

  • Files: new src/graph/op_matrix.rs, registered as #[cfg(test)] mod op_matrix;. It needs the crate's internals (the allocator, the backends), so it lives in the module tree rather than under tests/ — the crate is a binary target, so a file under tests/ is a separate crate that cannot name the binary's items at all (there is no lib target to link against). The test must therefore be an in-crate module.
  • Three tests:
    1. matrix_cases_match_their_reference — 17 cases (Add, Mul, Scale, Silu, SwiGLU, Softmax, RmsNorm, QkNorm, MatMul, GetRows, View/Reshape/Permute, RoPE, Attn, KvcacheStore/Load) run on every backend that claims the op, each compared against an analytic reference written in the test — never against another backend. Unavailable backends report SKIP (reason), never PASS: on dgxspark that is 17 CPU cells + 34 skips.
    2. support_table_matches_support_matrix_doc — supports_op for 23 op rows against the table published in SUPPORT-MATRIX.md, so the A8 doc and the code cannot drift. The CPU column is checked here; the Metal/CUDA columns check themselves wherever they are compiled in.
    3. every_op_has_a_matrix_decision — every Op variant is either covered or excused in EXCUSED. op_label matches with no wildcard arm, so adding an Op variant is a compile error until the matrix is updated.
  • Found a real defect, fixed here: the CPU Op::Softmax arm called vec_soft_max_f32 (which writes exp(x - max) and returns the sum) and discarded the return, so the op produced unnormalised output. Nothing caught it because no architecture emits a standalone Softmax node. The arm now scales by 1/sum; verified by neutering that line (the Softmax cell goes red) and restoring it.
  • Acceptance met: the matrix runs under cargo test; the A8 asymmetry table is now machine-checked; one real defect found and fixed; cargo test --release 162 passed / 0 failed.
  • Left for a GPU-verifiable run: the Metal and CUDA columns, and the four GPU-only fused ops (FusedQKV/FusedFFN/FusedQkvNorm/QkvBiasRopeStore) which are excused on a CPU-only box.
  • CUDA column executed on hardware (2026-09-18). With the sandbox corrected (A0) the CUDA column finally ran, and every cell it claimed passed: Add, Mul, Silu, SwiGLU, RmsNorm, QkNorm, MatMul, GetRows, View, Reshape, Permute, RoPE, Attn, KvcacheStore, KvcacheLoad (Scale/Softmax stay SKIP (Cuda does not claim …), which the support table asserts). Running it exposed four harness defects, all fixed here — the kind that only a device can show:
    1. the CUDA column silently depended on test order: CudaState::get() is None until something calls init(), which main does at startup and the model tests do through the loader, so a filtered run reported SKIP (no CUDA device) while a full-suite run used the device. The harness now initialises the state itself (in backend_claims and run_on).
    2. the synthetic weights were registered only on CPU (alloc.register_weight delegates to the CPU pool; in the product the loader registers each weight into CudaState), so every weight op failed with "weight not registered on CUDA". The harness now registers them into CudaState too.
    3. Attn (hd 2) and QkNorm (hd 2) used a head dim the CUDA kernels reject (must be a nonzero multiple of 4) — cells that could only ever pass on CPU. Both fixtures now use hd 4.
    4. the KvcacheStore/Load case wrote two of its four region rows and expected the rest to be zero: true for a fresh CPU pool, false for device memory. It now writes every cell, so the expectation is defined on every backend.
  • Acceptance met for the CUDA column too (2026-09-18): on GB10 (sm_121) the matrix is 17 CPU cells + 17 CUDA cells with no FAIL in either column.
  • Recorded limitation (2026-09-18): support_table_matches_support_matrix_doc compares the code against a mirror table hard-coded in Rust and only names SUPPORT-MATRIX.md in its failure message — it does not parse the markdown. A stale row label therefore survives a green suite (it did: the Attn row still said multi_seq after the field was renamed to explicit_span). Parsing the published table would make the drift impossible rather than merely visible; it is a hardening item for A1/A8, not a defect.

A2 — CI · item 24 · S — DONE (verified in CI)

  • Files: .github/workflows/ci.yml.

  • Deliverable: three jobs — test-linux-cpu (cargo test --release, the runtime net), build-linux-cuda (cargo build --release --features cuda in nvidia/cuda:12.8.0-devel-ubuntu22.04; no GPU on a hosted runner, so compile-only by necessity), and the unchanged macOS/Metal build.

  • Acceptance: the workflow parses and declares all three jobs; a broken commit now fails test-linux-cpu (unit tests run there, including the allocator/scheduler tests added in A3/A5).

  • Verified outcome (PR #1, run 35080755969):

    jobresulttime
    test-linux-cpu✅1m09s
    build-macos✅1m26s
    build-linux-cuda✅ (incl. cargo test --features cuda --no-run)4m36s
  • The x86_64 question is answered: the Linux/CPU suite passes. running 163 tests → 160 passed; 0 failed; 3 ignored, plus 3 passed; 6 ignored for the integration file. That is 2 fewer than the aarch64 dev box, and the difference is exactly the quants::neon_correctness module, gated #[cfg(all(test, target_arch = "aarch64"))] — no test is silently missing.

  • First CI run failed, and that was the point. Run 35080525201 came back build-linux-cuda ❌ at the dtolnay/rust-toolchain step: curl: command not found — the CUDA devel images are minimal, so rustup could not bootstrap and the CUDA build never started. Fixed by installing curl ca-certificates build-essential before the toolchain step (the last also supplies nvcc's host compiler). Only the fix commit's push turned all three green.

  • Container base moved (2026-09-30): the image tag above is now nvidia/cuda:12.8.2-devel-ubuntu24.04 — deliberately the same image the x64 release job builds in (.github/workflows/release-build.yml, linux-x64-cuda). The 22.04 tag was glibc 2.35 / gcc 11 against the artifact's 2.39 / gcc 13, so the compile check ran on a base no artifact was ever built on. The rest of this record stands as written.

A3 — KV bounds guard + ensure_kv size check · item 5 · S — DONE

  • Files: src/graph/alloc.rs (only — see the deviation note).
  • Deliverable, part 1: ensure_kv now records the element count and the backend each region was allocated for, and returns Err when a later graph asks for a different size or assigns the layer elsewhere. Before this, the early return silently handed back a wrongly sized region; the CPU backend then failed with a position error and the GPU backends wrote out of bounds.
  • Deliverable, part 2 (deviation from the ticket as written): the pos < n_ctx guard was not added to cuda_backend.rs. Because A0 found CUDA unverifiable here, the guard went into GraphAllocator::fill_input_i32 — the single point where positions become graph data. It is structural (an I32 input consumed by a KV-writing or attention op is bounded by the graph's n_ctx, taken from the kv_load node), so it covers all three backends including Metal without touching metal.rs, costs one O(nt) host scan, and is fully testable on CPU. token_ids is deliberately exempt (vocabularies exceed n_ctx).
  • Acceptance: kv_region_size_change_is_a_loud_error, position_beyond_n_ctx_is_rejected, token_ids_are_not_bounded_by_n_ctx (all fail before this change); cargo test --release 155 passed / 0 failed.
  • Defers to G: nothing — the Metal path is covered by the same allocator guard. Phase G keeps only the style asymmetry (debug_assert! vs Err). Superseded in part by gap-table G1 (#38, 2026-10-06): the allocator guard covers the fill_input_i32 path, and the Metal arm guard (MetalBackend::check_kv_store_rows) was added for the fill paths that bypass it.

A4 — Worker panic isolation · item 6 · S — DONE

  • Correction to the ticket as written: the worker did have a catch_unwind — guarded_forward (chat.rs:438-450) contains a panic inside forward_graph_cached. It is just too narrow: the speculative path calls both models' forwards directly (spec.rs:385/417/447/471/509), and the tokenizer, sampler, stop-string and streaming paths were unguarded too. One panic there unwound worker_loop, dropping job_rx (every later request is rejected) and the already-queued jobs' event senders (an empty 200 instead of an error).
  • Files: src/server/chat.rs — run_job_isolated + panic_message, called around the whole per-job body in worker_loop; the inner guarded_forward stays for the better message on the common case.
  • Deliverable: any panic in a job becomes a 500 on that request's stream, is logged, and the worker keeps draining the queue with the slot released.
  • Acceptance: isolated_job_turns_a_panic_into_an_error_event, isolated_job_forwards_a_normal_error, isolated_job_passes_success_through_silently; cargo test --release 158 passed / 0 failed.
  • Found while writing the test: panic_message(&payload) on a Box<dyn Any + Send> downcasts against the Box (which is itself Any) and always reports a non-string payload; &*payload is required. Verified against a throwaway rustc probe.
  • Not fixed (needs a device): a panic while the CUDA capture window holds the process-wide stream mutex poisons it, so every later request would panic too — now contained per job, but the server would still fail every request. Recorded for the CUDA-verifiable phase.

A5 — Staging map re-key · item 27 · S — DONE

  • Files: src/graph/alloc.rs (cross, copy_across, cross_buffer, plus a new write_pool helper that removes the duplicated four-arm write) and src/graph/scheduler.rs (the consumer-side filter).
  • Deliverable: cross is keyed by (NodeId, Backend). The scheduler's .filter(|cb| cb.backend == split.backend) is gone — the map itself can no longer offer a consumer the other backend's copy, and one node feeding two foreign backends now gets one staging buffer each instead of a single entry that only one of them could use.
  • Acceptance: staging_is_keyed_by_destination_backend (stages one node for CPU, asserts the CUDA consumer gets None, then stages CUDA and asserts the CPU entry survives); the existing scheduler tests still pass; cargo test --release 159 passed / 0 failed.
  • Honest limitation: a behavioural test needs a second usable backend, which dgxspark does not have (A0: CUDA unavailable; Metal not compiled). The test asserts the new keying contract directly through a #[cfg(test)] hook rather than through a real split boundary. Phase G should re-test it on a Mac.

A6 — CPU per-op allocations · item 28 · S — DONE: measured, refuted, reverted

  • Files (measured, not changed): src/graph/cpu_backend.rs — the per-node Vec<&[f32]> of resolved inputs, and the K/V source clone in the KV store arm.

  • What was tried: replace the per-node input Vec with a fixed [&[f32]; 4] array (every op the builders emit has ≤ 4 inputs) plus a heap spill for anything larger — i.e. zero allocation per node execution instead of one.

  • Measurement (Qwen2.5-0.5B Q4_0, CPU, 20 threads, bench -r 3, three interleaved before/after pairs, minfer-a6-before vs minfer-a6-after):

    pairbefore pp512after pp512before tg128after tg128
    176.08 ± 0.0374.79 ± 0.0340.30 ± 0.0939.60 ± 0.28
    276.05 ± 0.0175.33 ± 0.0240.25 ± 0.1639.07 ± 0.40
    376.14 ± 0.0375.39 ± 0.0540.00 ± 0.3139.68 ± 0.29

    The change is consistently slower (−1.2 % prefill, −1.8 % decode; every pair separated, and the intra-run error is ±0.05 or less). Reverting restores the baseline (76.10 / 76.06 pp512, 40.79 / 40.21 tg128), which confirms the cause.

  • Why it does not matter anyway: the decode loop is weight-streaming bound. At ~250 nodes/token the per-node Vec was one small allocation each, on the order of 0.1 % of the step — below what the harness can resolve. The K/V clone was left alone for the same reason: it is ~2 KB per layer and the buffers are needed for the borrow structure (the store arm needs disjoint access to four pool slots).

  • Outcome: no code change; the negative result is recorded so nobody re-opens it. Item 28 is closed as "not worth doing" rather than done.

A7 — Dead identity fields · item 26 · S — DONE, closed in E2

  • Decision taken: option (a) — delete CParams.n_batch, keep GraphParams.n_seqs marked reserved for item 3 (rationale in §8).
  • Files: src/graph/params.rs, src/graph/cache.rs:57-64, src/graph/json.rs:40, and the n_batch construction sites in src/models/*/graph.rs.
  • Deliverable: n_batch gone from CParams, from params_match and from the graph JSON export; n_seqs carries a comment naming the item that will read it.
  • Acceptance: no occurrence of n_batch remains in src/graph/; every field compared by params_match has at least one reader; cargo test green.
  • E2 closure (the second half of the ticket). Keeping n_seqs reserved was the right call only until item 3 landed. It did — and the field turned out not merely unread but redundant: the only topology decision a multi-sequence batch can force is the attention instantiation, and CParams.explicit_span, derived from the KV reservations (the authority on where each window starts), already carries it. E2 therefore deleted GraphParams.n_seqs, satisfying the acceptance by the "or deleted" branch. Measured justification (sequence_count_is_data_not_topology, models/qwen2/graph.rs): with the field present, a 2-sequence batch and a 1-sequence batch with the same n_tokens/n_out/gtype/explicit_span rebuilt the graph (uid 3 → 4); without it the same graph is reused and its logits are bitwise-identical to a fresh single-sequence forward. The test is permanent, so the claim cannot silently regress.

A8 — Guard symmetry · item 13 · S — DONE (docs route)

  • Files: docs/SUPPORT-MATRIX.md — a new "Operator Coverage by Backend" section.
  • Route taken: the ticket offered "CUDA gains FusedQkvNorm or the matrix records that it does not have it". A0 made CUDA unverifiable here, so writing a new CUDA kernel blind is strictly worse than documenting the asymmetry: the table now lists every op against CPU/Metal/CUDA, generated from the three supports_op implementations, with the four asymmetric rows called out and their consequences (Qwen3 decode is fused on Metal and unfused on CUDA; QkvBiasRopeStore is the mirror case; interleaved RoPE is CPU-only).
  • Acceptance: the roadmap §4 defect 5 asymmetry is now visible in the support matrix; A1's matrix must agree with this table.
  • Defers to G: the Metal half (debug_assert! → Err; the weightless RMSNorm fallback) and any decision to port FusedQkvNorm to CUDA. (The Metal half landed: debug_assert! → Err in #39 and the weightless RMSNorm fallback in #40 — both gap-table G2/G3. The FusedQkvNorm-to-CUDA question is #52.)

4. Phase B — persistent server context

Why second: it is independent of the KV redesign, and it is the first user-visible win — a chat client that resends the conversation stops re-prefilling it every turn.

IDItemTitleEffort
B14Reproduce the doc-97 contaminationS
B24Slot cache retention + prefix checkM
B34Measurement + docsS

B1 — Reproduce contamination — DONE: the documented mechanism does not reproduce

  • What the docs claimed. chat.rs reset the slot cache on every request because "re-prefilling a DIFFERENT prompt over the same regions leaves stale rows inside the new attention window" — the message of 39eceaa (doc 97), which fixed a cross-request KV bug and added the reset.
  • What was tested. reused_cache_across_prompts_matches_a_fresh_cache (src/models/qwen2/graph.rs) runs real prompts through one GraphCache and compares bitwise against a virgin cache, in three orderings:
    1. long prompt → short prompt (the stale-row case);
    2. short → long (the append case B2 wants);
    3. prefill → 3 decode steps → shorter prompt (the server's actual sequence, including the generated rows the comment is about). All three are bitwise equal.
  • Why. A prefill writes rows 0..nt and attention reads [0, max(pos)+1) = [0, nt) — i.e. only rows this request just wrote. Stale rows survive above the window, where nothing reads them. The original bug is consistent with the state recorded in the OpenAI plan's revision notes: the cache was process-global at the time, so two slots shared one set of regions. That is already fixed by per-slot caches; the reset was belt-and-braces whose stated mechanism does not hold on the current tree.
  • Residual uncertainty (why the reset stays until B2). This box can only run the CPU path. The fused GPU stores (FusedQKV, FusedQkvNorm, QkvBiasRopeStore) write K/V inside their own kernels and are unverified here (A0: CUDA compile-only; Metal not built). The chat.rs comment now records the finding and this caveat instead of the unverified mechanism.
  • Consequence for B2. Reuse is safe by construction if it is gated on an exact token-prefix match: every row read is then a row whose contents were verified to be the same tokens. That is what B2 implements.

B2 — Retention + prefix reuse — DONE

  • Files: src/server/slot.rs (Slot.cached_tokens), src/server/chat.rs (common_prefix_len, prefill_span, the suffix prefill, the per-token record, worker_loop no longer resets the cache).
  • Deliverable: the slot keeps its GraphCache and a record of the token sequence its KV rows hold; generate_seq reuses the cache only when the new prompt starts with exactly that sequence (else it prefills from position 0), and always feeds at least the last token because its logits seed the sampler.
  • Why it is safe by construction: the reuse gate is the verification — every row attention reads is a row common_prefix_len just proved to hold the same token. It is also bitwise: the CPU attention is now nt-invariant (item 14), so feeding a suffix at its own positions reproduces a single-shot prefill exactly.
  • Deliberate limits: the speculative path neither reuses nor records (a verify round writes rows past the committed tokens); a panicking job clears the record; MINFER_NO_PREFIX_REUSE=1 restores the pre-B2 behaviour for A/B.
  • Tests: common_prefix_len_finds_the_exact_match, prefill_span_always_feeds_the_last_token (pure), and prefix_reuse_matches_a_full_prefill (real model, bitwise).
  • Not covered here: the fused GPU stores are unverified on dgxspark, so the gate stays conservative; Phase G re-tests on Metal.

B3 — Measurement — DONE (interleaved A/B, same binary)

  • Setup: serve --n-ctx 1024 --n-slots 1 on Qwen2.5-0.5B Q4_0, one long system prompt (203 tokens) plus a second turn that appends an assistant reply and a new question (219 tokens), max_tokens=1 so the wall time is essentially time-to-first-token. Four alternating runs; after vs before differ only by MINFER_NO_PREFIX_REUSE=1.

    modeturnprompt tokfedreusedwall
    after120320302.66 s
    after2219162030.24 s
    before120320302.60 s
    before221921902.71 s
    after120320302.66 s
    after2219162030.26 s
    before120320302.88 s (turn 2)
  • Result: turn 2's prefill drops from 219 to 16 tokens and its time-to-first-token from ~2.7–2.9 s to ~0.24–0.26 s — ≈ 11× — while turn 1 (cold slot) is unchanged, as it must be. The per-request line [server] prefill fed N/M prompt tokens (R reused) is what the numbers come from, and it stays in the server so the behaviour is observable in production.

5. Phase C — KV cell store (item 1)

Five sub-steps, each keeping the tree green. C3 depends on Phase D.

IDTitleEffort
C1KvCache with cells, single implicit sequence (behaviour-preserving)L
C2seq_rm / seq_add: prefix truncation + context shift — DONEL
C3Defragmentation (cell copy) — needs D1M
C4Quantized KV (item 21) · #42M
C5State save/restore for session persistence · #43M

C1 — Cell store, one sequence — DONE

  • Landed: src/graph/kvcache.rs (KvCache, KvLayer, SEQ_MAIN/FREE, cells_for, is_identity, and the ownership/write bookkeeping C2 will use); the allocator's kv map is now the store, so ensure_kv validates through it (it gained the n_ctx argument) and kv_pair/copy_kv_to_cpu read it.
  • The gate has teeth: BackendScheduler::execute refuses any graph whose mapping is no longer the identity, because no backend consumes the resolved cell array yet — so a half-ported C2 fails loudly instead of writing the wrong row. a_non_identity_kv_mapping_is_refused proves it fires.
  • The resolver has a live consumer today: A3's position guard (fill_input_i32) now resolves through cells_for instead of re-deriving the bound from a node's shape, so when C2 changes the mapping the check keeps meaning the same thing without being touched.
  • Verified bitwise, as the ticket requires: cargo test --release 172 passed / 0 failed — including the real-model bitwise tests (reused_cache_across_prompts_matches_a_fresh_cache, prefix_reuse_matches_a_full_prefill, graph_logits_match_forward_real_model) — and a pre-C1 vs post-C1 binary A/B on Qwen2.5-0.5B Q4_0 greedy (-n 24) produced byte-identical generated text; only the timing lines differ.
  • Deliverable (as specified): per-layer cell arenas with an owner set; the KV store node resolves position → cell on the host and passes an index array to the kernel; attention still derives its bound from positions, so behaviour is unchanged.
  • Delivered in this pass: the arenas, the owner set and the host-side resolver. The index array reaching the kernel is deliberately not done — it is only needed once the mapping stops being the identity, and the scheduler gate refuses to execute in that state, so no backend can silently index the wrong row in the meantime. Handing the array across the Backend trait was planned as C2's first step; C2 landed without it because it removes rows physically and therefore never leaves the identity mapping (see the C2 record below) — so the hand-off stays available, not urgent.
  • Acceptance (as specified): bitwise-identical greedy output vs the pre-change binary on the full smoke set (0.6B Q8_0, 7B Q4_K_M, 14B Q4_K_M) at several context lengths; the graph topology is unchanged (no new topology-affecting params).
  • Acceptance met for: 0.5B Q4_0 byte-identical end to end, plus the in-tree bitwise model tests. Deferred at the time: the 0.6B/7B/14B smoke rows — minutes-long CPU jobs and C1 does not touch a kernel, so they were left to the C2 landing rather than claimed here. Discharged at C2 (2026-09-16): the pre-C2 (C1) binary and the C2 binary generate byte-identical text greedily (-n 16, -t 8, single-shot) on Qwen3-0.6B Q8_0, Qwen2.5-7B Q4_K_M and Qwen2.5-14B Q4_K_M; only the timing lines differ.
  • Defers to G: nothing — the Metal backend keeps the old regions until G.

C1 design (written before the code, 2026-09-16). The shape below is what C2 needs, so C1 builds it rather than a type with no consumer:

  1. KvCache (new src/graph/kvcache.rs) owns, per layer, the two persistent regions plus owner: Vec<SeqId> — one entry per cell, FREE for a cell no sequence holds. One implicit sequence for now, so every cell of a written row is owned by it.
  2. KvCache::cells_for(positions: &[usize]) -> Result<Vec<u32>, String> — the host-side resolver. Today it is the identity (cell == pos) and returns Err for a position outside the arena; C2 changes exactly this function to consult owner and the layer's window start.
  3. KvCache::is_identity() -> bool — true until C2 introduces a hole or a window. The scheduler consults it: while true the existing positions input reaches the backend exactly as today (so no kernel changes), and once false the resolved cell array must be passed instead. A backend that has not been ported returns Err instead of indexing the wrong row (standing rule 2). On dgxspark that means CPU first: CUDA was compile-verified only at the time (A0 — superseded 2026-09-18, the device is available) and Metal stays untouched (G), so C2 must either keep the mapping identity for them or refuse to run there.
  4. The index array travels as a graph input (kv_cells, I32), filled by the allocator from positions — positions stay data, so the topology and the params-only reuse identity are untouched.
  5. Tests: resolver edge cases (out-of-range → Err, identity while there are no holes), is_identity flipping exactly when C2 introduces one, and the existing model tests (B1, B2's prefix reuse, graph_logits_match_forward) staying bitwise — that is the refactor's acceptance gate.

The first thing to change is therefore the resolution, not the kernels: the Backend trait gains the resolved cell slice alongside kv_pair, and each backend uses it for row indexing while still using positions for the causal bound and RoPE.

C2 — Removal and shift — DONE

  • Deliverable: seq_rm (drop a range) and seq_add (shift positions), surfaced as graph-level operations; conversation.rs overflow switches from "drop the oldest turns + full re-render" to a context shift.
  • Acceptance: the conversation overflow test passes; a turn that overflows no longer re-prefills the whole conversation; the shift path is a named tolerance class with the reason recorded (positions change ⇒ RoPE inputs change ⇒ not bitwise).
  • Defers to G: n/a.

C2 record (2026-09-16)

What landed. GraphAllocator::kv_rm(start, len, &KvRope) removes KV rows [start, start + len) from every layer and re-bases the rows after them by -len, re-roping their K; kv_shift(drop, rope) is the start == 0 case (KvCache::after_rm / after_shift do the bookkeeping). rope_shift_kv now takes a signed delta so a rotation can be undone (the tests use that).

The removal is physical, which is the whole point: cell == pos survives, so is_identity() stays true, C1's scheduler gate stays satisfied, and no backend gains a kernel. The work is a host-side memmove plus one re-rope pass per layer, both through the existing copy_kv_to_cpu / write_host pair — which is also why CUDA/Metal need no change here (copy_kv_to_cpu has no Metal arm, so a Metal session falls back loudly rather than shifting wrongly; the port is G5).

Engine::kv_rm(start, len) is the conversation-facing hook (Err by default, so a mock or a non-shiftable backend is handled explicitly). Conversation plans the drop on a copy of the message list, resolves the dropped turn's token span from the chat template, and verifies both boundaries against the KV stream before touching anything; an unverifiable boundary or an engine that refuses falls back to the exact drop-and-re-render path, with the reason logged. MINFER_NO_CONTEXT_SHIFT=1 forces that path.

The retained rows keep the context they were computed with. This is the ticket's named tolerance class, and its cause is not rounding. Re-roping is exact to the rope tolerance (measured max|Δ| = 2.38e-7, relative 1.14e-7, by rope_shift_matches_roping_at_the_new_position), but a shifted row's value is whatever it was when it was written — it attended to the tokens that were later dropped. Measured on Qwen2.5-0.5B Q4_0 with the a/b/c probe (drop a, keep b, continue with c):

Layer0151223
retained-K `maxΔvs a freshb` prefill1.5e-52.557.80

Last-token logits after continuing with c: max|Δ| = 16.75 vs a fresh b+c prefill, and the argmax changes. This is inherent: recomputing the retained rows is exactly the re-prefill the shift exists to avoid. llama.cpp's context shift (llama_kv_cache_seq_rm + seq_add + llama_kv_cache_update) has the same property. What C2 guarantees instead is exactness of the mechanism, and that is what the tests assert bitwise:

  • removing a tail range leaves the retained head byte-identical, so continuing from it is bitwise equal to a fresh prefill of that head (kv_rm_is_exact_and_the_window_shift_is_a_named_tolerance_class);
  • removing a middle range copies V byte-for-byte, leaves [0, start) untouched, and moves K by exactly the re-rope (undone to < 1e-4 per layer);
  • removing everything written empties the arena, len == 0 is a no-op, and removing past the end is an Err, not a silent truncation.

Model-path A/B (discharging C1's deferral). The pre-C2 binary (HEAD at C1) and this build produce byte-identical generated text greedily (-n 16, -t 8, single-shot, CPU) on Qwen3-0.6B Q8_0, Qwen2.5-7B Q4_K_M and Qwen2.5-14B Q4_K_M — only the timing lines differ — which is the evidence that adding the removal path did not perturb the model path.

Conversation measurement. context_shift_real_model_measurement (0.5B Q4_0, n_ctx = 192, 12 turns with a system prompt) overflows on turns 8–11 and shifts on every one:

turn 8:  dropped 22 KV rows at 17 (2 messages), prefill 14 tokens instead of 185
turn 9:  dropped 23 KV rows at 17 (2 messages), prefill 14 tokens instead of 180
turn 10: dropped 23 KV rows at 17 (2 messages), prefill 14 tokens instead of 175
turn 11: dropped 23 KV rows at 17 (2 messages), prefill 14 tokens instead of 170

— a ~13× cut in prefill tokens per overflowing turn, start = 17 (the system prompt is retained and not even re-roped), no full re-render, and the answers stay correct on the shifted window (Stockholm,/Athens,/Warsaw,/ Lisbon, for Sweden/Greece/Poland/Portugal). prefill_tokens is the observable that makes "no longer re-prefills the whole conversation" checkable, so the ticket's acceptance is a test, not a claim.

Two pre-existing bugs the C2 tests found and fixed (both outside the ticket but blocking it, and both user-visible in --cnv):

  1. The first user_turn never prefilled pre-existing messages — it wrote only the new message's delta, so the --system prompt never reached the KV while still being rendered into every later delta. The first turn now prefills the whole canonical render (first_user_turn_prefills_the_system_prompt).
  2. A pending EOT was written before the overflow check, so a full context with a reply that had not reached EOG called forward at position == n_ctx and tripped the KV-region guard. The EOT is now deferred into the overflow path (it belongs at the end of the stream, so it commutes with the removal), and ContextFull keeps it pending for the next attempt.

Not done here. The logical mapping (a hole or a window with cell != pos, which is what would let a backend keep the old rows in place) is deliberately still refused by C1's scheduler gate — nothing consumes the resolved cell array yet, and the physical removal makes it unnecessary for the sliding-window case. C3's defragmentation, C4's quantized KV and C5's state save/restore are untouched. The server path is untouched too: a full slot is still reported as finish_reason: "length" and relies on B2's prefix-matched reuse, so a server-side shift policy (and the n_keep-style rule that protects a system prompt) belongs with E2's batching work.

C3 — Defragmentation

  • Deliverable: a cell-copy operation that compacts the arena; triggered when fragmentation exceeds a threshold.
  • Deps: D1 (a copy needs either a view or an explicit copy op).
  • Acceptance (as resolved, 2026-09-19): node-count and arena-utilisation counters before/after; the moved bytes are exact — V verbatim and K exactly rope_shift_kv(old, delta), both asserted — and a mid-session compaction leaves the continuation's greedy token intact (asserted; a wrong re-rope flips it). Bit-identical logits were not claimable before C6 (a cell move shifted every RoPE angle, so the tail moved by the amplified rounding of that shift — §14 row 9 carries the four probes that attributed it). C6 landed the construction that removes it: positions are sequence-relative and the allocator resolves cells, so a compaction now changes no angle and the continuation is bitwise (gated by a_compaction_between_steps_keeps_the_continuation and the inverted offset tests).
  • Follow-up: C6 does exactly this — sequence-relative positions plus an allocator-resolved cells input — which makes a cell move change nothing (a token's RoPE angle is its index within its sequence) and drops C3's K re-rope. When C6 lands, this ticket's acceptance tightens from the named amplified-rounding class back to bit-identical logits.

C3 design (written before the code, 2026-09-19)

The problem the arena has. KvCache::reserve_seq is first-fit over [0, n_ctx): it needs a contiguous free run of cap cells, while release_seq frees whatever run a sequence held. Two sequences that finish out of admission order therefore leave a hole in the middle, and the next admission can fail with no free run of N cells while the arena has far more than N free cells. That is fragmentation, and it is the one failure mode of E2's per-sequence reservations that no amount of capacity fixes.

Scope. C3 compacts downward: live runs are packed from cell 0 in ascending old-start order, each keeping its cap, so free space coalesces at the top of the arena. Compaction never changes a run's cap, never reallocates a KV region (the copy is within the same K/V buffer) and never changes a query's logical position — but it does change the cell a sequence's rows live in, so the caller that reserved the run must be told (the contract below).

Where the copy executes. The K/V regions are backend buffers (host memory on CPU, device memory on CUDA), so the copy is a backend operation: Backend::copy_cells(&mut self, dst: BufRef, src: BufRef, rows, elems_per_cell) — a new trait method, implemented by CPU and CUDA in this ticket and refused loudly by Metal (Phase G), which keeps standing rule 2: a backend that cannot move cells fails the request instead of skipping the compaction.

The contract is dst_row <= src_row (downward only) and the implementation must be correct for overlapping ranges, which the two obvious implementations are not:

  • CPU: copy_within (memmove semantics) is exactly right.
  • CUDA: device-to-device cudaMemcpyAsync is documented undefined for overlapping ranges, so the kernel is one block that walks rows in ascending order with a __syncthreads() between rows. Ascending is the safe direction when dst <= src (the row a write could clobber is already copied), and a row's elems_per_cell (<= n_kv_embd) parallelises across the block's threads. No staging buffer and no second pass.

The planner is pure. KvCache::compaction_plan(need) is a host-side function over the run table with no backend calls: it returns the moves (seq, from, to, rows) that compact the live runs, or — with Some(n) — the shortest prefix of that plan which opens a free run of n cells (None compacts fully). rows is the sequence's written length (last owned cell + 1 - start), not cap: unwritten cells hold no data worth copying, and moving them would shuffle another sequence's stale bytes. Because the planner is pure, its policy is unit-tested without a device — that is the half CI can prove.

The contract with the caller. Only whoever reserved a run caches its start: E2's server keeps one per slot (SlotState.start) and passes start + current_pos as the store position, so a move must be reported. GraphAllocator::kv_defrag(need) returns the applied moves plus before/after stats, and BatchEngine applies the new start to the slots it owns; everything else follows, because the graph's window (attn_span) is derived from seq_slot on every forward. A caller that ignores the report is not silently wrong in a hard-to-find way: its next store would write to the old cells and the span check would reject the query.

Trigger. GraphAllocator::kv_reserve_seq_with_defrag retries once through a compaction when first-fit fails, unless MINFER_NO_KV_DEFRAG (presence-checked) disables it — the A/B gate standing rule 3 requires: the same workload must produce identical output with and without the compaction, and the counters must show what it bought. The plain kv_reserve_seq stays pure, so a caller that keeps no run bookkeeping (the single-sequence path, the graph tests) is unaffected; the defrag-capable variant returns the moved runs precisely because a caller that does keep one (E2's server) has to follow them.

Counters (the ticket's other deliverable). KvArenaStats carries n_ctx, reserved/owned/free cells, the number of free runs (the "node count" C3 compacts), the largest free run, the live run count, and cumulative defrags/cells_moved. The acceptance test records them before and after; the same struct is what F8 (observability) will export.

Increments.

  1. Planner + counters + KvCache bookkeeping (apply_moves) with unit tests — pure host-side logic, so CI proves it.
  2. Backend::copy_cells (CPU + CUDA), GraphAllocator::kv_defrag, the reservation retry and the MINFER_NO_KV_DEFRAG gate.
  3. Integration: fragment the arena on purpose (--n-slots N, prompts of different lengths) so an admission needs the compaction, and assert the server's replies stay bit-identical to the serial path, plus a device run for the CUDA copy.

Deferred (recorded, not silently skipped). Compaction does not grow a sequence's run, does not evict, does not move a sequence up to make room for a larger one, and does not run inside a forward (it is a between-forwards, host-driven operation, so CUDA Graph capture is unaffected). Compacting the arena per layer separately is also out of scope: one plan moves every layer together, which is what keeps owner[] consistent across layers.

C3 record (2026-09-19) — increments 1 and 2

Landed: the planner, the counters and the bookkeeping (increment 1), and the copy path (increment 2) — Backend::copy_cells on CPU (copy_within, i.e. memmove) and on CUDA (a new kv_move_rows kernel: one block, ascending rows, a barrier between them, no staging buffer), Metal refusing loudly (G5), GraphAllocator::kv_defrag + kv_reserve_seq_with_defrag + the MINFER_NO_KV_DEFRAG gate, and the server applying the returned moves to its slot starts.

Evidence:

  • CPU end to end: kv_defrag_moves_the_bytes_and_opens_the_run fragments a real 16-cell arena (K/V regions allocated by a kvcache_store graph), asserts the 8-cell reservation is refused, compacts, then checks the bytes at the new cells, the counters (free_runs 2 -> 1, defrags/cells_moved), and that the refused reservation now fits — and repeats the fragmentation so the retry helper has to hand the moves back. CPU suite: 207 passed / 0 failed / 5 ignored.
  • Device: cuda_copy_cells_moves_overlapping_rows_down moves rows [1, 4) to [0, 3) — two of the three rows are read and overwritten — compares the whole buffer, and checks that the downward-only at the time (C7b later added the upward direction — see §5) contract is refused before a launch. CUDA suite: 255 passed / 0 failed / 5 ignored.
  • Coverage split, stated rather than implied: the allocator's CPU arm is proven end to end and the CUDA primitive is device-proven, but "the allocator drives the CUDA copy" is glue that only increment 3's integration run exercises.

Increment 3 (2026-09-19) added the part that makes a compaction correct rather than merely well-bookkept, and the model-level gate:

  • A compaction must re-rope the K rows it moves. positions are cells today, so RoPE angles are absolute: a K row copied from cell from to cell to still carries the angle of from, and the next decode step attends to it at the wrong relative angle. kv_defrag(need, rope) now re-ropes the moved K rows through the model's own rope_shift_kv — the C2 path, one host pass per moved run, with delta = from - to (C2's convention: rope_shift_kv(d) means new_pos = old_pos - d). V carries no rotation and moves verbatim.
  • The gate caught its own sign bug. a_compaction_between_steps_keeps_the_continuation (real 0.5B: prefill + one step at a non-zero start, release the holder below it, compact, then one more step) first failed on its argmax assertion because the delta was written as to - from; the allocator's unit test had passed anyway, because it computed its expectation with the same wrong sign. That is the whole argument for a behavioural gate next to a byte-level one.
  • The offset-alone effect, measured while gating. The control run differs from the compacted one only in the subject's cell offset, and that alone moves the logits by 2.6% relative (0.43 absolute) on the 0.5B before any compaction; the compaction's own contribution is the remaining ~0.2pp. A uniform RoPE shift is attention-neutral and V is unrotated, so an unidentified op is offset-sensitive. This is now §14 row 9 — it is why the C3 acceptance is a surviving greedy token plus byte-exact K/V rather than bit-identical logits, and it is the floor the compaction test prints.

Coverage: the planner, the counters and the bookkeeping are unit-tested; the compaction's bytes, ownership and counters are asserted end to end on CPU (kv_defrag_moves_the_bytes_and_opens_the_run, with the K re-rope computed independently through rope_shift_kv); the CUDA copy primitive is device-proven (cuda_copy_cells_moves_overlapping_rows_down); and the continuation gate runs the real model. Still open: driving the compaction from the server's dynamic runs (today every slot is reserved once at startup and packed, so the trigger cannot fire there), and the logical-positions change that would remove the re-rope and the offset sensitivity altogether.

C4 — Quantized KV cache, Q8_0 first · #42 — DONE (CPU: S1 + S2a; CUDA: S2b) 2026-09-24

Why. f32 (and the GPU's f16-in-an-f32-region) is all the cache could be, so context length was bounded by KV memory with no way to trade precision for it. Q8_0 is the first packed layout: one cell is ceil(n_kv_embd/32 * 34) bytes instead of 4 * n_kv_embd, which is a real footprint reduction — unlike f16, whose regions stay f32-shaped and buy bandwidth, not memory.

What S1 landed (src/graph/kvformat.rs is the new authority):

  • KvFormat { F32, F16, Q8_0 } with a strict MINFER_CACHE_TYPE parser and a device policy (resolve, pure so CI covers the matrix): an unknown value is refused on every device — CUDA used to read anything that was not f16 as f32, which is exactly the silent fallback the ticket forbids — and q8_0 is refused loudly on CUDA and Metal, whose attention kernels address f32/f16 rows (both since lifted: CUDA in C4 S2b below, Metal in #310). f16 on CPU stays the documented f32 (a process whose env is set for a GPU run must not fail a CPU one); the load path (models::load_model_ns) turns any refusal into a failed load with the reason printed.
  • Packed rows stay addressable by the C3/C8b machinery. A cell is rounded up to a whole number of f32 words, so elems / n_ctx is still one cell's width and copy_cells (the copy-on-write shift, the compaction) keeps moving rows verbatim with no change at all. KvcacheMeta::row_elems carries the cell width, ensure_kv sizes the region by it, and the node's shapes stay logical ([n_kv_embd, n_ctx]), which is what the store's K/V input means.
  • CPU store quantizes, CPU attention reads the packed blocks — S1 dequantized the window into a scratch; S2's fused read is below. The CUDA/Metal kernels are still issue #87.

Acceptance, as measured (CPU, cargo test --release, 2026-09-22):

  • Footprint: the real-model gate prints the persistent regions, f32 against q8_0 — 0.5B (n_kv_embd 128): 6 291 456 B → 1 671 168 B (3.76× smaller); Qwen3-0.6B Q8_0 (n_kv_embd 1024, head dim 128): 58 720 256 B → 15 597 568 B (3.76×).
  • Tolerance class (never bitwise — the store rounds every K/V cell): at the reference's argmax the logits move by ≤ 1.0 (measured 0.60 on the 0.5B, 0.55 on Qwen3), and over the whole 152k-way vector by ≤ 3.0 (measured 2.50, ≤ 8 % of the 37.8 spread). The greedy continuation agreement is reported, not asserted: the CLI diverges at the 5th token on a chat-templated 0.5B prompt, which is the format's honest cost, not a defect.
  • The format itself is pinned one level down: a_packed_kv_region_answers_like_the_f32_one_and_is_smaller asserts a stored cell is bitwise the Q8_0 quantizate of the row it was given (and that the f32 region holds the row verbatim), plus the 3× footprint and an attention comparison over three causal queries. The gate was mutation-checked: disabling the packed read path turns that comparison into max |Δ| = 3.2e38.
  • Refusals: MINFER_CACHE_TYPE=q8_0 on dgxspark (CUDA available) ends the load with minfer: MINFER_CACHE_TYPE=q8_0 is not supported on cuda yet: …, and a typo (banana) with … is not a KV cache type (f32, f16, q8_0); refusing rather than silently running with f32. ensure_kv backstops both: a packed width on a non-CPU backend, and a width Q8_0 cannot express, are Err where the region is sized. (Superseded: CUDA stops refusing it in C4 S2b below, Metal in #310.)

Coverage, stated rather than implied. The packed layout, the parser/policy matrix, the store-exactness check, the refusal and the region accounting are unit-tested and run in CI. The real-model gate is #[ignore]d (at this record's time it flipped a process-wide KV policy; since #99 the format is per engine and it no longer mutates shared state — see the #99 record below) and was run on both cached models. Not verified here: any packed region on a GPU (by design, S1 refuses it), and the fused dots' performance (S2).

C4 S2 — the fused read, the quantize-aware shift, and what is handed off (2026-09-23).

#87 named three items. The first two are CPU-only and land here; the GPU kernels are the follow-up at the end of this record.

  • The fused read (attn_heads_q8, src/graph/cpu_backend.rs). S1 dequantized the union of the batch's windows into a reusable f32 scratch — two allocations plus a write-then-read pass per attention node, per layer, per forward — and ran the unchanged f32 kernel over it. S2 reads the packed blocks directly: the K score is dot_q8_0_q8_0 between the stored K blocks and the Q8_0-quantized query row (llama.cpp's form for a quantized cache), and V accumulates out of the cell block by block (kvformat::accumulate_q8_0_row). No scratch, no second pass. MINFER_NO_FUSED_Q8_KV (presence-checked) restores S1's path — the A/B standing rule 3 asks for — and the S1 path stays as the fallback for a hd_kv narrower than one Q8_0 block.
  • The quantize-aware shift. kv_rm/kv_shift used to refuse a packed region because they re-rope K in f32 and a packed cell is not f32 rows. S2 keeps the property S1 already relied on — a cell is a whole number of words, so the survivors move verbatim, which is why copy_cells and the CoW/compaction paths needed no change at all — and then maps each survivor through kvformat::map_q8_0_cells: dequantize → re-rope → requantize. Stated honestly, the shifted K is quantize(rope(dequantize(quantize(row)))): C2's re-rope class composed with C4's packing class.
  • The store allocates once per node, not once per row (kvformat::pack_q8_0_cell_into): a 2048-token prefill packs ~98k cells across 24 layers, and the per-cell vec! was the packed store's one avoidable cost.

Acceptance, as measured (CPU aarch64 on the GB10 box, 20 threads, Qwen2.5-0.5B Q4_0, n_ctx 4096, minfer bench -r 3, greedy):

configtg128 (ctx 512)tg64 (ctx 2048)pp512pp2048
f32 KV57.9647.00182.94173.43
q8_0 fused (S2)59.7748.63182.52164.06
q8_0 S1 scratch51.5037.26187.54174.67
  • Decode is the measured effect: the fused read is 1.16× (ctx 512) and 1.31× (ctx 2048) faster than S1's dequantizing read, which is what the deferred item was for. Against f32 the packed cache went from 0.89× / 0.79× (a real regression, which is why "does not regress by more than a named factor" was the acceptance) to 1.03× (both contexts). A second run at ctx 2048 read 174.89/50.20 f32, 162.10/50.22 fused, 156.62/37.66 S1 — the same decode ratios; the f32 baseline itself moves 47.0–50.2 between runs, so only within-run ratios are quoted.
  • Prefill is not measurably affected: the three configs overlap inside their own stddev (3–6 tok/s).
  • Tolerance class: the real-model gate's whole-vector bound moved 3.0 → 4.0 because the fused read adds a new term — the query's own Q8_0 quantization. Measured on the same gate: 3.029 fused vs 2.505 S1 over the vector, and at the reference's argmax 0.646 vs 0.604; the decoding-relevant bound stays 1.0. MINFER_NO_FUSED_Q8_KV=1 reproduces the S1 column, so the delta is attributable to the term rather than to run-to-run noise.
  • The shift: after a 4-row shift of a 20-token prefill, the first decode step — the one whose input is the shifted context — is 2.466 from the f32 shift (0.064 at the argmax) on the fused path and 3.165 (1.074) on S1's; the greedy continuation agrees 4/8 and 1/8 respectively, reported, because a flipped token makes the two runs different sequences from that step on. (An earlier version of this gate compared the last step instead and read 13.9 for a difference that is 0.06 at the decision point — the measurement, not the code, was wrong.)
  • Gates, each mutation-checked, one level below the model: the_fused_q8_read_matches_the_dequantizing_reference (two KV heads whose K rows differ strongly, so a wrong head block cannot pass; using head 0's blocks for every head fails with max |Δ| 1.238 of a 1.313 spread — the shipped path measures 6.0e-5) and a_packed_physical_shift_moves_v_verbatim_and_requantizes_k (V verbatim compared as bits — a packed word is an f16 scale plus int8 quants, so as f32 it is frequently NaN and == on it is never true; K exactly the Q8_0 quantizate of the re-roped row; the tail zeroed; skipping the re-rope fails at row 0). The shift gate's first fixture filled only positions and left the builder's own cells input at zero, so all three rows went to cell 0 and the gate passed on zeros: that was found by mutating the re-rope and watching the mutant survive. The fixture now fills cells and asserts every row is non-zero.
  • Suites: CPU 282 passed / 0 failed / 15 ignored (was 280/0/15), green both with and without MINFER_NO_FUSED_Q8_KV.

Handed off, still on #87: the CUDA and Metal kernels — MINFER_CACHE_TYPE=q8_0 is still refused on both, and the refusals stay loud. CUDA is the one that matters for throughput, and it is a kernel project rather than a read-path change: kv_ld4<KV> and stride_kv address elements, so a packed cell needs byte-based addressing plus a block-dequantizing load in every attention kernel (three window modes each), a Q8_0 store, and the fused decode QKV epilogue's own store. Metal stays at G5 by the round's own decision. (Superseded: CUDA landed in C4 S2b below, Metal in #310.)

C4 S2b — the CUDA Q8_0 kernels: landed 2026-09-24 (#87). The subsection below is the plan of record as it was written before the kernels; the design decisions, the measured results and the honest scope are recorded after it, so the plan and the outcome can be read against each other.

The plan of record (2026-09-24, written before the first line of kernel code).

The CPU half of #87 landed (S2a). The device half is a kernel project, not a read-path change, so it lands as its own increment against the same issue. This is the map it starts from: where the CUDA backend touches a KV region, the design that covers those sites, and the dispatch cuts that keep the first increment bounded. It came from walking the three files below at a5f7961 + S2a, so the line numbers are the ones to read, not to trust.

What is there today. No CUDA code knows KvFormat. The layout is one process-wide bool (cuda.rs's KV_F16), and every kernel addresses a cell as row * nkt elements of a float*/__half* — there is no byte addressing anywhere, so a 34-byte-per-32-element cell cannot be expressed by any current kernel signature. set_kv_cache_type maps anything that is not exactly f16 to false, so flipping the load gate before the kernels exist would silently mean f32.

The design. A layout tag plus byte-addressed row accessors in cuda_kernels.cu:

  • KV_LAYOUT_F32 / KV_LAYOUT_F16 / KV_LAYOUT_Q8_0, kv_row(base, cell, row_bytes), and kv4<LAYOUT>(row, elem) -> float4 as the one load idiom: f32 is the old float4 load, f16 the old two-__half2 pair (both bit-identical to today's instantiations), and Q8_0 reads block elem/32's f16 scale plus four quants at 2 + elem%32. A 4-element group never straddles a block, because a KV head's base is hd-aligned and hd % 32 == 0 — the property ensure_kv's packed-width check already enforces;
  • kernels take const void* k/v + size_t row_bytes instead of a typed pointer plus stride_kv = nk * hd, templated on int LAYOUT rather than typename KV;
  • store_kv_q8_0: one thread per (row, 32-element block), with the CPU store's own quantizer (amax/127, f16 scale, round_ties_even), so both backends store the same bytes.

The access sites to convert (src/cuda_kernels.cu, plus the host rows below):

SiteLinesWhat it is
kv_ld4<KV> specialisations2935–2950the split body's two loads
attn_split_1w_body<KV,…>2957–3053 (3013, 3016)K+V, cell[j]-indexed, hd/4 dims per lane
gqa_attn_split_partial<KV,…>3055–3083decode (nt == 1)
gqa_attn_split_partial_bt<KV,…>3120–3137spec-verify (1 < nt ≤ 16)
gqa_attn_f32<CAUSAL,MAP>3456–3575 (3500/3510/3535/3544)the general kernel (nt > 16 prefill)
store_kv_f32 / store_kv_f162530–2569the store both dtypes use
attn_bias_rope_store_f322590–2667 (2649–2665)the fused decode epilogue's K/V store
gqa_attn_f32_f16kv, attn_split_h4w_body, fa_stage_kv_async, fa_prefill_f16kv2781–2900, 3273–3431, 4327–4372, 4374–4648the f16-specialised paths (see the cuts)
kv_move_rows8703–8720the compaction mover — a plain word copy, so a packed cell (a whole number of f32 words) needs no change; the host already passes the cell's word count
host launcherscuda.rs 349–897 (FFI), 4481–4755 (CudaState), 5393–5488 (store/epilogue)typename KV / f16_kv: bool become a layout tag
executorcuda_backend.rs 26/114/133 (the field), 1075–1109 (store), 1111–1262 (dispatch), 1510–1541 (copy_cells), 154–156 (test setter), + 41 kv_f16 test sitesthe format decides store, dispatch and the move stride
format plumbingcuda.rs 1456–1478, kvformat.rs 124–133 (supports), alloc.rs 867–939 (ensure_kv's packed refusal) and 2035–2059 (kv_element_format), models/{qwen2,qwen3}/loader.rs 462–464one authority: the device reads the format the loader already resolved

The dispatch cuts that keep the first increment bounded — build-time choices (no silent fallback), each stated when the format is enabled:

  • decode (nt == 1) takes the split-K path through the converted 1-warp body; its hybrid 4-warp dispatch (hd == 128, nkv >= 1921) is f16-typed and is skipped for Q8_0 (rpw_gate = 0);
  • 1 < nt ≤ 16 routes to gqa_attn_f32 instead of the batched split kernel, which exists for spec-verify's bitwise identity contract with sequential decode. A speculative session must therefore refuse a packed cache loudly (a draft keeps its own KV, and the two reduction schedules would no longer agree) — consistent with S2a, which already refuses a draft in the session container;
  • prefill (nt > 16) routes to gqa_attn_f32: the FA path (fa_prefill_f16kv, f16-typed shared-memory staging) is not offered for Q8_0 in this increment, so the prefill is correct but off its tuned path;
  • the fused decode epilogue (Op::FusedQKV / Op::QkvBiasRopeStore) is not built for a packed cache — the model builders' layer_gpu gate gains && !packed, and a Q8_0 decode takes the unfused bias/rope/store chain through the converted store kernel.

Acceptance (the issue's, made concrete): cuda_map_window_matches_the_span_over_the_same_rows (cuda_backend.rs:3678) extended to Q8_0 rows — it writes its own region contents and calls exec_ids directly, so it needs a packed cell encoder and a packed-word region size; a Q8_0 store→attention round trip (cuda_kv_f16_roundtrip_attn, :5150, is the pattern); copy_cells under the packed stride (cuda_f16_kv_cell_move_strides_by_row_bytes, :3946); the real-model gate a_packed_kv_cache_answers_like_the_f32_one run with the cuda feature on GB10 (since #123 that gate is CPU-forced on a CUDA build until these kernels land, so this acceptance includes re-pointing it at the device; CI has no GPU, and its MINFER_C4_MODEL arm covers a second model); and an A/B against f16 on the same model/context with the numbers recorded. The honest expectation, stated before the work: the win on the device is memory (3.76× less than f32, 1.88× less than f16); whether decode bandwidth also improves depends on the block-dequant instruction count, so the bar against f16 is "no worse than a named factor", not "faster".

S2b design decisions, frozen before the first line of kernel code (2026-09-24). The plan above is the map; these are the choices it leaves open, decided up front so the implementation cannot drift into a silent fallback:

  1. Layout tag and accessors. KV_LAYOUT_F32 = 0 / KV_LAYOUT_F16 = 1 / KV_LAYOUT_Q8_0 = 2 are #defines in src/cuda_kernels.cu, and 0/1/2 are a host contract: the Rust side stores the same codes (the registry's KvFormat discriminants) and every launcher takes the tag as an int. Two accessors are the only place a KV address is formed: kv_row(const void* base, int64_t cell, size_t row_bytes) (byte address of a cell) and kv4<LAYOUT>(const char* row, int elem) -> float4 (the single load idiom). F32 is the old float4 load and F16 the old two-__half2 pair, both bit-identical to today's kv_ld4<float> / kv_ld4<__half>; Q8_0 reads block elem/32's f16 scale plus four quants at 2 + elem%32. A 4-element group never straddles a block because a KV head's base is hd-aligned and hd % 32 == 0 — the property ensure_kv's packed-width check already enforces.
  2. No typed pointer survives. Every attention kernel that can see a packed row takes const void* k, const void* v plus size_t row_bytes, and templates on int LAYOUT instead of typename KV. The row_bytes == nk * hd * 4 value the f32 path passes is the same arithmetic as the old stride_kv * sizeof(KV), so the f32/f16 instruction streams are unchanged.
  3. Dispatch cuts, each stated when the format is enabled (build-time; a packed row never reaches a kernel that would address it as f32):
    • decode nt == 1 → the converted split-K 1-warp body (gqa_attn_split_partial), with rpw_gate = 0: the hybrid 4-warp dispatch is f16-typed (attn_split_h4w_body takes const __half*) and is not converted in this increment;
    • 1 < nt <= 16 → gqa_attn_f32. The batched split kernel (gqa_attn_split_partial_bt) exists for spec-verify's bitwise identity contract with sequential decode; rather than claim that contract without measuring it, a speculative session refuses a packed cache loudly (the draft already keeps its own KV, and C5 already refuses a draft session);
    • prefill nt > 16 → gqa_attn_f32. The FA path (fa_prefill_f16kv, f16-typed shared-memory staging and tensor-core QK^T) is not offered for Q8_0, so the prefill is correct but off its tuned path — measured and reported, not hidden;
    • the fused decode epilogue (Op::FusedQKV / Op::QkvBiasRopeStore) is not built for a packed cache: the model builders' layer_gpu gate gains && !packed, so a Q8_0 decode runs the unfused bias → rope → store chain through the converted store_kv_q8_0.
  4. The store is the CPU's quantizer, byte for byte. store_kv_q8_0 maps one thread to one (row, 32-element block), computes amax, d = amax/127 as an f16 (nearest-even), and round_ties_even for each quant — the same three steps quants::quantize_row_q8_0_into uses — and writes d then the 32 quants into the packed cell at 34-byte stride. Both backends therefore store the same bytes for the same f32 row, which is what makes the CPU/device Q8_0 comparison a layout check rather than a tolerance question.
  5. The gate flips only after the kernels exist, and through the registry. READS_PACKED_KV stays false until every dispatch cut above is in place; then it becomes true for CUDA and KvFormat::supports(Cuda) follows automatically (registry::reads_packed_kv is the one authority). CudaBackend gains an int layout field fed by the resolved KvFormat, so MINFER_CACHE_TYPE=q8_0 can never again mean "f32 with a packed region".
  6. Acceptance bar, named before measuring. The device win is memory: a Q8_0 cell is 34/128 bytes per element versus 4 (f32) and 2 (f16), i.e. 3.76x less than f32 and 1.88x less than f16. Whether speed also improves depends on the block-dequant instruction count in the load path, so the bar against f16 is stated as a factor, not as "faster": Q8_0 decode tokens/s must be no worse than 1/1.30 of the f16 decode rate on the same model and context (the same 1.30x already recorded for the CPU's fused Q8_0 read at ctx 2048), and the prefill is reported as its own number because it takes the untuned gqa_attn_f32 route.

Why it was not in the S2a increment. ~10 kernel sites, 3 launchers, ~15 host sites and 41 test sites, each needing an nvcc iteration and — for the gates — a serial device run. It was recorded here rather than half-wired. Metal's half stays at G5 — since landed as #310.

What actually landed (2026-09-24). The real counts are close to the estimate: 8 kernel sites, 3 launchers (launch_gqa_attn_split_q8_0, the layout-tagged launch_gqa_attn_f32, launch_store_kv_q8_0) plus two re-templated existing ones, ~25 host sites (cuda.rs's FFI + CudaState methods, cuda_backend.rs's store/attention/fused-epilogue dispatch, copy_cells, the registry entry), 4 new gates and 2 strengthened ones. The two files the plan expected not to change did not: kv_move_rows is a plain word copy (verified — the packed cell's word count is what the host already passes, and the Q8_0 move gate pins it), and the CPU path is untouched.

  • Layout and accessors, exactly as designed. KV_LAYOUT_F32/F16/Q8_0 are 0/1/2 on both sides of the FFI; kv_row + kv4<LAYOUT> are the one address/load idiom. The f32 and f16 kernels compile to their pre-C4 instructions (same loads, same cells, same order).
  • The gate flipped through the registry. BackendCaps::reads_packed_kv became true for CUDA and KvFormat::supports(Cuda) followed automatically; CudaBackend carries an int kv_layout fed by the resolved KvFormat, so the pre-S2b bool that mapped anything but f16 to f32 can no longer turn a packed region into f32 rows.
  • The store is the CPU's quantizer, byte for byte. cuda_q8_0_store_matches_the_cpu_quantizer asserts the device's packed words equal kvformat::pack_q8_0_cell's for two block widths, and that unwritten cells stay zero. Mutation: amax / 126 instead of / 127 fails it at nkt=64 row 0.
  • All three acceptance gates, mutation-checked. cuda_map_window_matches_the_span_over_the_same_rows now sweeps f32/f16/q8_0; cuda_kv_q8_0_roundtrip_attn (store → attention) and cuda_q8_0_kv_cell_move_strides_by_row_bytes (copy_cells under the packed stride) are the two new device gates. compute-sanitizer --tool memcheck over all three reports no memory error (the 2 pre-existing CUDA API errors it did report were root-caused in S2c below, #145). Mutations: dropping the Q8_0 block base in kv4 left the map-window gate green (every comparison there is between modes over the same bytes, so a value-level fault shifts both sides together — the F6 lesson) and made the round-trip and the real-model device arm fail (max |Δlogit| 2.48 → 26.34 of a 37.8 spread); halving the packed stride in copy_cells fails the move gate; flipping READS_PACKED_KV back to false makes the device arm refuse the region loudly. The map-window gate was strengthened: it now also compares one single-row Q8_0 window against the dequantized V cell (max |Δ| 1.71 under the block-base mutation), so it cannot pass on consistency alone.

The A/B against f16, measured on dgxspark (aarch64, GB10 sm_121) — the bar was named before the work as "Q8_0 decode no worse than 1/1.30 of f16" and it is not met against the fused f16 baseline:

Superseded 2026-10-07 — read the numbers as measured, not as current. This table is the pre-#144 S2b measurement (2026-09-24; its baseline was re-measured in the #144 round at master 85c712e, which reproduced it within a few percent), and three later changes moved it: #144 items 1+3 (041de15, 2026-09-26), #186 (cc19b4f, 2026-09-27, the dp4a packed K dot) and #202 (798fd32, 2026-09-27, the wide u16 block load). The current A/B is the 2026-10-07 comment on #310, measured at 740e0ff on dgxspark (aarch64, GB10 sm_121): 1.23x decode / 1.24x prefill at hd 64, and 1.01x / 1.04x at hd 128 — the 15.3x is gone, and q8_0 is still never faster than f16 on CUDA.

model / configf16 (pinned MINFER_CACHE_TYPE=f16)q8_0factor
Qwen2.5-0.5B q4_0, pp2048 @ n_ctx 40962693.54 tok/s2171.85 tok/s1.24x slower
Qwen2.5-0.5B q4_0, tg128 @ n_ctx 4096239.76 tok/s162.46 tok/s1.48x slower
…f16 with the fused epilogue cut (MINFER_NO_FUSE_QKV=1)203.58 tok/s—the Q8_0 kernel's own cost is 1.25x
Qwen3-0.6B Q8_0 (hd 128), pp20488604.56 tok/s562.76 tok/s15.3x slower
Qwen3-0.6B Q8_0 (hd 128), tg128137.85 tok/s122.21 tok/s1.13x slower

The f16 column is pinned, not auto. auto_device_format needs n_layers × n_kv_embd ≥ 8192 (AUTO_F16_MIN_KV_ELEMS, src/graph/kvformat.rs), so today's resolver answers f32 for the 0.5B (24 × 128 = 3072) and f16 for Qwen3-0.6B (28 × 1024 = 28672); both arms in this table were run with MINFER_CACHE_TYPE=f16 explicitly, and the header's older "(default policy)" wording does not describe the 0.5B arm on the current resolver.

Read honestly [as measured pre-#144; the clauses that no longer hold are marked inline]: decode on the 0.5B misses the named 1.30x because ~1.18x of the gap is the stated cut (no fused QKV epilogue for a packed cache — the unfused chain adds launches on a model whose decode is launch-bound) and the remaining 1.25x is the packed load itself (kv4<Q8_0> costs four int8 converts and four multiplies where f16 costs two __half2 converts; there is no dp4a packed dot in this increment — #186 landed it on 2026-09-27, cc19b4f]). On Qwen3-0.6B, where the fused epilogue is not on the f16 default path in the same way, decode is 1.13x — inside the bar (the superseding run at 740e0ff measures 1.01x). The prefill is a different story: at hd 128 the f16 path is fa_prefill_f16kv (8604 tok/s) and the packed path is the general kernel (563 tok/s), i.e. the packed cache is correct but off its tuned route by 15x — #144 item 3 then landed the packed FA prefill, so the superseding run measures 1.04x here]. The win is memory, as stated up front: measured on both models, the f32/f16 region is 6 291 456 B and the Q8_0 region 1 671 168 B — 3.76x smaller than f32 and, against f16's actual 2 B/element payload (3 145 728 B), 1.88x smaller.

Acceptance results.

  • Real model, device arm: a_packed_kv_cache_answers_like_the_f32_one runs both arms — the CPU one (kept, coverage on every build) and, on a CUDA build with a device, one that asserts device() == Cuda. 0.5B: f32 6 291 456 B vs q8_0 1 671 168 B (3.76x), CPU max |Δlogit| 3.03 / argmax 0.65, CUDA 2.48 / 0.60 of a 37.8 spread; the physical-shift arm: CPU 2.47 / 0.064, CUDA 3.34 / 0.42 — inside the bounds the CPU gate had already fixed.
  • Suites: CPU cargo test --release 432 / 0 / 28 unit + 10 / 0 / 6 integration (unchanged); CPU serial ignored 28 / 0; CUDA serial unit 490 / 0 / 31 (was 486/0/31: three new device gates + one new cuda.rs layout test; the packed gate now also runs on the device); CUDA serial ignored 0.5B 31 / 0; Qwen3-0.6B 30 / 1, the one failure the then-pre-existing #130 f16 session-container gap (closed in C5 S3: 31 / 0).
  • The Q8_0 session container works. a_session_resumed_from_disk_continues_bitwise with MINFER_CACHE_TYPE=q8_0 on the device prints live KV format: q8_0, saves 24 layers / 256 cells / 5 written / 1 696 464 B and continues bitwise (max |Δlogit| = 0). The container's FLAG_PACKED bit is what encodes it; f16 had no flag then — the gap #130 closed in C5 S3 (FLAG_F16), below.

Handed off after S2b: the packed fused decode epilogue, a dp4a packed K dot and the FA prefill on packed cells — taken up as #144 and recorded in the next subsection (items 1 and 3 landed; item 2 then landed as #186 below, DONE 2026-09-27). Metal's half stays at G5 — since landed as #310.

C4 — #144: the packed fused decode epilogue and the packed FA prefill · #144 — DONE (items 1 + 3) 2026-09-26

Why. C4 S2b's packed cache is correct and 3.76x smaller than f32, but three paths were off their tuned route: a Q8_0 decode ran the unfused bias/rope/store chain (the builders' layer_gpu gate carried && !packed), kv4<Q8_0> paid four int8 converts where f16 paid two __half2, and an hd-128 packed prefill could not enter fa_prefill_f16kv's f16 shared-memory staging, so it fell to the general layout-tagged kernel — 15.3x off on Qwen3-0.6B pp2048. This ticket took the first and third (a dp4a K dot is a numerics change with its own accuracy statement; it is filed separately).

What landed.

  1. The packed fused decode epilogue. attn_bias_rope_store_q8_0 keeps the f32/f16 epilogue's q section (bias + rope in place, one thread per pair) and re-maps K and V to one thread per (head, 32-element block): the thread computes all 32 K roped values itself (from the unroped row plus the pair partner d <-> d + hd/2) or the 32 V bias-added values, and hands them to q8_0_quantize_block — the same device quantizer store_kv_q8_0 now calls, factored out so the two cannot drift. Neither K nor V is written back (both buffers are dead in the fused classes) — which is also why the packed epilogue has no rope race. The builders dropped && !b.kv_is_packed() from the Qwen2 family's fuse_qkv gate. FusedQkvMeta / QkvBiasRopeStoreMeta / FusedQkvNormMeta gained a row_elems field (the packed cell width, KvFormat::row_elems(nkt)) and the allocator sizes the region from it: without that, the first packed fused graph died with the node declares 34 words per cell but the q8_0 layout packs one cell of 512 elements into 136.
  2. The packed FA prefill. fa_prefill_f16kv became fa_prefill_kv<CAUSAL, MAP, LAYOUT>: the staging is the only layout-dependent part (kv8_q8_0 dequantizes each packed 32-element block into the same f16 tile, one 16-byte smem store per 8 elements); the tensor-core QK^T, the fragment-resident softmax and P·V are untouched. gqa_attn_kv_prefill serves f16 and Q8_0 from one call and keeps the general layout-tagged kernel as the documented fallback; the launcher's failed-smem arm and MINFER_NO_FA_PREFILL=1 both reach it. The chokepoint bumps testfail::note_checked("cuda_fa_prefill_q8_0"), so the gate can prove the packed prefill took the FA route rather than silently falling back.

The A/B protocol and the numbers. All numbers: GB10 sm_121, cargo build --release --features cuda, minfer bench -p 2048 -n 128 --n-ctx 4096 -o json, MINFER_CACHE_TYPE pinned per arm, 5 interleaved rounds, medians, 2026-09-26. The bars were named before measuring (gate contract rules 3 and 5); the full protocol lives in CUDA_OPTIMIZATION.md and cuda_optimization_steps/107.

The re-measured baseline (master 85c712e, its own binary) reproduced the S2b table within a few percent: 0.5B q4_0 tg128 237.37 f16 / 161.49 q8_0, MINFER_NO_FUSE_QKV=1 f16 200.80 (a 1.182x cut — the ticket's 1.18x); Qwen3-0.6B Q8_0 pp2048 8323.24 f16 / 564.47 q8_0.

itembar (named first)measuredverdict
1 packed fused epilogue0.5B tg128 q8_0 ≥ 1.15x the unfused packed chain170.11 vs 161.63 same-binary = 1.052x (vs the baseline binary 161.49 = 1.053x); f16 gap 1.470x → 1.393xlanded, bar missed
3 packed FA prefillQwen3-0.6B pp2048 q8_0 ≥ 0.5x the f16 arm8231.05 vs 564.57 same-binary = 14.58x; vs the f16 arm 8540.83 = 0.964x, i.e. the 14.7x factor became 1.038xlanded
1/3 no-regression0.5B pp2048 q8_0 (hd 64: FA does not apply) flat2144.18 vs 2139.10 = 1.002xflat

Item 1's miss is a result, not a failure. The ticket's "~1.18x" was the fusion cut measured on the f16-weight arm; on the q4_0 packed arm the same binary's fused-vs-unfused A/B is 1.052x. The saving is the ~6 launches/layer the epilogue removes (~13 µs/layer at 24 layers ≈ the 0.32 ms/token the two arms differ by), so the launch-overhead component is what this cut actually buys; the residual 1.39x f16-to-packed gap is the packed load (kv4<Q8_0>'s four converts) and the 1-warp split-K decode body's rpw_gate = 0, which the dp4a item targets.

Verification. The two new device gates and the five named Q8_0 gates are green: cuda_q8_0_fused_epilogue_matches_the_cpu_quantizer (new: byte-exact K/V against the CPU quantizer at pos = 0, a value q check at pos = 7, plus V byte-exact), cuda_q8_0_fa_prefill_attention_parity (new: 100-token hd-128 prefill against the CPU attention over the same packed bytes, 5e-3 class, max err 4.8e-4, plus the note_checked observation arm), cuda_q8_0_store_matches_the_cpu_quantizer, cuda_kv_q8_0_roundtrip_attn, cuda_q8_0_kv_cell_move_strides_by_row_bytes, the Q8_0 arm of cuda_map_window_matches_the_span_over_the_same_rows, and cuda_fa_prefill_attention_parity (the f16 route, unchanged, max err 2.8e-4). Mutations: reverting the fused kernel's block offset (d = blk*32 + i → d = i) turns the byte arm red (it found exactly that bug during development); shifting kv8_q8_0's block base (elem >> 5 → elem >> 4) turns the FA parity red (max err 3.33); making the packed prefill skip the FA launch leaves the parity arm green (max err 5.8e-5) but turns the observation arm red — which is the whole reason it exists. MINFER_TEST_ISSUE162=1 drives all 124 audited launch sites (six new: three fa_prefill_kv layouts × modes and the packed epilogue) and is green.

Deliberately not taken. The dp4a packed K dot (item 2): a per-head-quantized query against the packed K accumulates in int, which is a numerics change needing its own accuracy statement and a re-measured real-model tolerance — filed as a follow-up rather than assumed. Qwen3's packed fused QKV (Op::FusedQkvNorm) also waits: that op has no CUDA kernel at all (it is Metal + CPU today), so the !packed gate there is not what keeps Qwen3's decode off the device; its 1.13x packed gap is unrelated to this ticket's cuts.

C4 — #186: the dp4a packed Q8_0 K dot · #186 — DONE 2026-09-27

Why. [#144]'s residual was 1.393x on Qwen2.5-0.5B q4_0 tg128 (170.11 q8_0 vs 236.97 f16), and [#144] names two components: the 1-warp split-K body's rpw_gate = 0 plus the kv4<KV_LAYOUT_Q8_0> load. Only the load is addressable by a dp4a K dot. The trap is that a whole-kernel q8_0-vs-f16 gap would mix the two: at hd 128 the f16 arm takes the 4-warp hybrid body while the packed arm cannot. hd 64 (the 0.5B) is the shape where both take the identical gqa_attn_split_partial<LAYOUT,true,false> on the same grid, so the delta there is the load alone.

The bar, named before measuring. The load-attributable share of the Q8_0 decode step ≥ 10%. Reasoning: only one of the residual's two components is addressable; a free load recovers at most share / 1.393; dp4a removes only the K side's converts and multiplies, adds a query quantize, and still pays the memory traffic; [#144]'s landed item 1 was 5.3% on this arm, the measured floor for "worth landing". Measured 20.3% (nsys, real decode, node tracing: the incumbent packed partial kernel is 67424 ns median per layer/token vs f16's 18528, = 1.182 ms of a 5.818 ms/token step). Bar cleared.

What landed. attn_split_1w_body<LAYOUT,CAUSAL,MAP,Q8DP4A> (the Q8DP4A default false, so the f32/f16 and verify instantiations are untouched) quantizes each lane's four query values against its 32-element block's amax (the eight lanes sharing a block reduce it with three shfl_xor), reads the K quants as int8 through kv4_q8_0_packed, accumulates __dp4a(qk, ki, 0) and scales by d_q * d_k once per block. V still dequantizes; only K changes. The launcher launch_gqa_attn_split_q8_0 takes an int dp4a and picks the instantiation; cuda::q8_kv_dp4a_enabled() resolves MINFER_NO_DP4A_Q8_KV=1 once per process (the same-binary A/B control), and the Rust call site bumps testfail::note_checked("cuda_q8_kv_dp4a") only when it launched the dp4a arm.

The numbers (GB10 sm_121, cargo build --release --features cuda, minfer bench -p 2048 -n 128 --n-ctx 4096 -o json, 5 interleaved matched rounds, medians, 2026-09-27):

arm0.5B q4_0 tg128Qwen3-0.6B Q8_0 tg1280.5B pp2048
incumbent (MINFER_NO_DP4A_Q8_KV=1)171.95123.982149.22
dp4a193.30 (1.124x)136.66 (1.102x)2150.37 (flat)

The 0.5B f16 arm in the same round is 240.03, so the packed/f16 residual went 1.396x → 1.242x; on Qwen3-0.6B the f16 arm is 136.73, i.e. the packed decode now matches f16. The isolated probe (same 1-warp geometry, load only) cuts the delta from +96% to +40% at hd 64, and ncu (collected as root, the module parameter unchanged) names the residual: dp4a removes 494984 instructions (−15.3%) and leaves the L1 load-sector count unchanged at 792904 — 1.71x f16's, while its L2 read sectors are 0.57x f16's — so the remaining 1.242x is the 34-byte block layout's L1 request count, not DRAM traffic and no longer the arithmetic.

Tolerance, re-measured not assumed. a_packed_kv_cache_answers_like_the_f32_one's CUDA arm reads max |Δlogit| 2.479504 of a 37.82 spread, at the argmax 0.552662, greedy 9/9 (incumbent arm: 2.479504 / 0.596050 / 9-of-9) — inside the ≤4.0 / ≤1.0 class. The max is unchanged because it is the nt = 512 prefill step's; the decode steps' deltas moved and nsys on the gate itself names gqa_attn_split_partial<(int)2,(bool)1,(bool)0,(bool)1> under the default and ...,(bool)0> under the control, so the class is not "both arms ran the old path".

Verification and refusal-to-over-claim. The dp4a arm was added to cuda_kv_q8_0_roundtrip_attn (two cells through an explicit span, an exactly Q8_0-representable query, plus the observation-counter arm); the first version used one cell and was vacuous — a mutated block base still passed, which is how the two-cell form was found. Mutation: elem >> 5 → elem >> 4 → red at max |Δ| = 0.35126442. Prefill/verify is deliberately not taken: the nt > 1 general kernel also reads kv4<Q8_0> and measures 1.32x off its f16 arm on the 0.5B prefill attention, but that is a different kernel (per token and head query, no shared block scale) and not the ticket's residual — named rather than assumed. Full record: cuda_optimization_steps/108.

C4 — #202: the packed Q8_0 KV cell's L1 request count · #202 — DONE 2026-09-27 (counter-only; throughput bar missed)

Why. [#186]'s ncu attribution named the residual: the packed decode arm issued 1.71x f16's L1 load sectors (~792904 vs 463008) with 0.57x its L2 sectors, because a 34-byte Q8_0 block's 4-element group at blk + 2 + (elem & 31) is 4-byte aligned only on odd blocks, so kv4<Q8_0> spends four s8 loads plus a scattered scale load per group. The ticket listed three candidates and said to prefer the one with no layout change if it clears the bar (a split plane or a 36-byte block would move the CPU store/read path, copy_cells, map_q8_0_cells, the FA-prefill staging and the C5 session version).

What landed. The no-layout candidate: 34k + 2 + 4m is always 2-byte aligned (34k is even for every k, and 4m is), so two unsigned short loads replace the four s8 loads for the same bytes. q8_0_load4<WIDE> / kv4<LAYOUT,WIDE> / kv4_q8_0_packed<WIDE> / a Q8WIDE template parameter on attn_split_1w_body, and a third arm in launch_gqa_attn_split_q8_0, selected once per process by cuda::q8_kv_wide_enabled() (MINFER_NO_Q8_KV_WIDE=1 is the same-binary control). Three new launch sites, driven by the #162 test and added to the audit fixture. No layout, CPU, copy-stride, session-format or map_q8_0_cells change.

The bar, named before measuring. (a) packed/f16 L1 load-sector ratio ≤ 1.30x; (b) same-binary tg128 ≥ +2% over the #186 dp4a baseline 193.30, pp2048 within 2%. Measured: (a) met with margin — 792904 → 455840, ratio 1.712x → 0.9845x, instructions −2.17%; (b) NOT met — 193.13 → 193.65 (+0.27%, 5 interleaved medians), the kernel itself only 1.4–1.7% faster (nsys 40320 → 39744 ns median).

Honest result: a partial refutation. The L1 request count moves exactly as the mechanism predicted and the decode step does not follow, so the 1.23x packed/f16 residual is not L1-request-bound — not L2 traffic (0.564x), not instruction count (1.10x), not load sectors (0.9845x). Landed counter-only, byte-identical (the nt > 1 general kernel keeps the byte form; only the dp4a decode arm opts in), with the latency hypothesis (per-block scale load + cvt + V dequant on a 1-warp block's critical path) recorded for the next attribution rather than claimed. Mutation: swapping the two u16 halves in q8_0_load4_wide turns cuda_kv_q8_0_roundtrip_attn red at max |Δ| = 2.5326836. Full record: cuda_optimization_steps/109.

Test infrastructure — #207: the CUDA count-row drift · #207 — DONE 2026-09-27

Why. The CUDA rows are live_check = false recorded measurements (CI has no GPU), but they count the same test binary plus the device-gated tests: a CPU-only ticket that adds a feature-independent test moves both. [#140]'s K-quant encoder tests (7 passed / 1 ignored) and [#142]'s bf16 writer + CPU-path tests (7 passed / 2 ignored) did exactly that and left the row reading 548 / 0 / 39 while a device run on the branch reads 562 / 0 / 42 and both CUDA real-model rows read 42 / 0; this was the second consecutive occurrence.

What landed. Mechanism (c), chosen over a pending field and a rule line: docs/status.toml's recorded rows may carry projection_key / projection_box / projection_base_passed, and scripts/check_status.py --check prints (never fails on) passed + (cpu_now − base) when the CPU twin has moved. The relation is inexact — the device-gated tests move independently — which is exactly why it is a hint and not a comparison; the point is that a stale recorded row cannot look current without a projection beside it. The unit/sanitizer rows project from cpu-unit (aarch64) and the two real-model rows from cpu-real-model (aarch64), each with that row's passed at this measurement as the base, so all four hints are silent today and fire the moment a CPU ticket adds tests. --selftest gained a moved/not-moved pair (12/12 cases).

Why. minfer bench on a CUDA build printed CUDA kernel launch error: 1 between its two loops — pre-existing on master, not an S2b regression — and compute-sanitizer --tool memcheck over the unit suite reported 36 CUDA API errors. CudaState::sync polls cudaGetLastError, which reports whatever an earlier call on the thread latched, so an API error from several operations ago was printed as a failure of the kernel that had just run.

Root cause: two origins, both unchecked return values.

  1. gemm_prefill_smem_init (1 of the 36; #145). The eager dynamic-smem opt-in looped over every gemm_f16_nt_kernel_t<tm, ks, af32> and asked cudaFuncSetAttribute(.., cudaFuncAttributeMaxDynamicSharedMemorySize, N) for a stale N: its formula assumed a 512-thread TM=256 launch and always added the AF32 mirror, so gemm_f16_nt_kernel_t<256,64,true> requested 131072 B against GB10/sm_121's cudaDevAttrMaxSharedMemoryPerBlockOptin of 101376 B. The rejected call's return value was never read, so the error latched. The same formula existed a second time, inline in launch_gemm_f16, and that copy dropped the AF32 mirror — the launcher declared 16384 B less than the kernel's own Bs/Cs offsets need at the default KS=32, an out-of-declaration shared-memory access that only worked because the block's smem happened to be carved where nothing else wrote.
  2. cudaGraphDestroy on a cudaGraphExec_t (26 of the 36; #128). CudaState::graph_destroy is the only destroy path for an exec handle and called the graph destructor. It returned cudaErrorInvalidValue, leaked the exec, and the latch surfaced at the next sync — the actual source of the bench line. The remaining 9 errors were cudaGetLastError observations of one of the two.

What landed.

  • One formula, read twice. gemm_dynamic_smem_bytes(tm, ks, af32) is the single source for the kernel's byte layout (As 2*TN*KS halves + Am 2*TN*KS floats [AF32] + Bs 2*TM*KS halves + Cs NW*256 floats, TN = 64, NW = blockDim.x/32 = 8 at every launch site); launch_gemm_f16 and gemm_prefill_smem_init both call it.
  • The eager opt-in is checked and deliberate. Every cudaFuncSetAttribute return value is read; a failure is named once, at init, with the function, attribute, requested bytes, device limit and cudaGetErrorName, and cleared there. A request above the queried device limit is skipped without calling it, with the reason printed — on GB10 that is gemm_f16_nt_kernel_t<256,64,true> (122880 B > 101376 B), which cannot launch on this device at all. CudaState::try_new reports the counts.
  • sync attributes honestly. latched_api_error_message names the observer (cudaGetLastError), the cudaGetErrorName symbol and the code, and states it is not attributed to a kernel; the error is counted (latched_api_error_count) and cleared — still visible, never dropped. debug_sync's label follows.
  • graph_destroy calls cudaGraphExecDestroy; the cudaGraph_t / cudaGraphExec_t distinction is stated at the extern and the call site, and a failure is named there and cleared.

#218 forward note (2026-09-29): the eager sweep this record describes is gone. #188 deleted gemm_prefill_smem_init's CudaState::try_new call site with no mention in any commit message or document, and #218 removed the orphaned function (and the checked/skipped introspection) instead of leaving it allow(dead_code)-annotated. The per-launch opt-in this record already describes — gemm_smem_optin → the shared minfer_smem_optin, reached by launch_gemm_f16 — is now the whole mechanism; the invariant it must uphold (the attribute is in force before a capture window opens, never set inside one) is upheld by the 3-run capture warmup (capture_warmup), cudaStreamCaptureModeThreadLocal and the per-instantiation cache, and gated by cuda_prefill_smem_optin_is_done_by_production, the control arm cuda_prefill_smem_optin_refusal_fails_the_prefill, the coverage arm cuda_prefill_smem_lazy_optin_admits_every_launchable_instantiation, and cuda_prefill_smem_optin_is_never_set_inside_a_capture_window — see the #218 record in §test-infrastructure. gemm_dynamic_smem_bytes remains the one formula; the "read twice" wording above is historical (the launcher reads it once).

#223 forward note (2026-09-29): the eager half is back, but not as the sweep this record describes. CudaState::try_new now calls gemm_prefill_smem_prewarm_one(tm, ks, af32) for every launchable combination — a thin dispatcher onto the same production gemm_smem_optin cache the launcher reads — so there is one opt-in mechanism and one cudaFuncSetAttribute site, not a second sweep with its own copy of the formula (the pre-#145 bug class). The placement is the argument: at try_new no CudaBackend (and so no capture window) can exist, so the attribute is set outside any window by construction. The checked/skipped counters and the startup banner are still gone; a successful pre-warm is silent, and each failure or deliberate over-limit skip is named per instantiation. The lazy path stays as defence in depth, and the #218 gates keep their claims by running their children against the documented control MINFER_NO_GEMM_PREWARM=1. On the real binary the loop costs ~2.2 ms (the fatbin's one-time module load), which prewarm_prefill() already paid — net new startup cost ≈ 0, hot path within ±1% — see the #223 record in §test-infrastructure.

Acceptance results (GB10, sm_121, CUDA 13.0, driver 580.178.04).

checkbeforeafter
compute-sanitizer --tool memcheck over the serial CUDA unit suite36 API errors (26 cudaGraphDestroy, 9 cudaGetLastError, 1 cudaFuncSetAttribute)0 errors
CUDA serial unit suite490 / 0 / 31495 / 0 / 31 (five new gates)
CUDA serial ignored, 0.5B config31 / 031 / 0
CUDA serial ignored, Qwen3-0.6B Q8_030 / 130 / 1 at this commit — only #130, closed in C5 S3 (31 / 0)
minfer bench -p 64 -n 8 -r 2, 0.5B Q4_K_Mprints CUDA kernel launch error: 1 between the two loopsno such line; the one skipped opt-in named instead
CPU cargo test --release432 / 0 / 28 unit + 10 / 0 / 6 integrationunchanged
CPU serial ignored28 / 028 / 0

The five new gates, all mutation-checked.

  • the_latched_error_message_never_blames_a_kernel (pure) — the message names cudaGetLastError and cudaErrorInvalidValue and does not contain "kernel launch". Mutation: the old message fails it.
  • the_gemm_smem_formula_matches_the_kernel_layout (pure) — pins gemm_dynamic_smem_bytes against the kernel's byte layout for all 12 (tm, ks, af32) combinations. This is the value-level arm the F6 lesson asks for: a mode-vs-mode comparison is blind to a shrunk formula (launcher and init shrink together, so a consistency check stays green). Mutation: dropping the AF32 term fails it.
  • cuda_prefill_smem_optin_covers_every_launchable_instantiation (device) — every >48 KiB request the device admits reads back opted in through cudaFuncGetAttributes().maxDynamicSharedSizeBytes; every over-limit one is skipped, never called. Mutation: continue-ing one admitted combination fails it.
  • cuda_graph_exec_destroy_leaves_no_latched_error (device) — capture → instantiate → destroy returns true and leaves the latch at 0. Mutation: cudaGraphDestroy fails it. The gate asserts the destroy call's own result (graph_destroy now returns bool), not just the latch: the failure is named and cleared inside graph_destroy, so a latch-only assertion passed with the wrong destructor — the first cut of this gate passed for the wrong reason and was fixed here.
  • cuda_sync_surfaces_a_latched_error_as_latched (device; env-gated behind MINFER_TEST_LATCH_ERROR=1 because it deliberately latches a real API error, which a compute-sanitizer run must not see) — sync reports the injected error once and clears it. Mutation: dropping the report fails it.

The >48 KB capture path stays covered by the existing prefill gates — prefill capture is ON by default (R3-B) and cuda_prefill_capture_defaults_on, cuda_prefill_capture_bit_parity_pp16_pp300, cuda_multisplit_capture_bit_parity and the real-prefill cuda_graph_generation_replay_parity_real_model are all green in the 495 / 0 / 31 run.

The deliberate-failure check. The pre-fix over-limit request was forced back with the sanitizer-clean skip bypassed, so cudaFuncSetAttribute really failed: the init printed cudaFuncSetAttribute(gemm_f16_nt_kernel_t<256,64,true>, cudaFuncAttributeMaxDynamicSharedMemorySize, 131072 B) failed: cudaErrorInvalidValue (1); device opt-in limit 101376 B, and cuda_prefill_smem_optin_covers_every_launchable_instantiation failed (failures = 1). Reverted; the reverted src/cuda_kernels.cu is byte-identical to the committed one (sha256sum, in the closing comment on #145).

Honest scope. The device limit — and therefore which instantiation is skipped — is measured on GB10/sm_121 only; the decision is a runtime query and the skip is printed, so another device skips a different subset. MINFER_GEMM_TM=256 + MINFER_GEMM_K64=1 + the f32-A path is the only combination that lands on the skipped instantiation, and it had no working opt-in before either. The remaining unchecked CUDA calls in the same file are enumerated in #147 rather than silently fixed here.

C4 — S2d: the remaining unchecked CUDA attribute / launch / destroy returns · #147 — DONE 2026-09-25

Why. S2c (#145) fixed the two latched-error origins compute-sanitizer could still see and left the CUDA unit suite at 0 API errors — but the audit it performed listed a whole family of call sites that still discard a return value which gates a later launch or allocation, the class that let #122, #128 and #145 hide in the first place. None was known to fail on sm_121, so this is latent hardening, not a live bug: the value is that the next driver, device or tile-config change cannot turn one of them into another phantom "kernel launch error".

The re-derived site list (verified against the tree, not the ticket). The S2c table named six rows; re-deriving them against the post-#145 source gives eight code sites, one of which the table understated (the wide-NT launcher already read its own return but reported nothing) and two of which were already fixed by #145 (the eager gemm_prefill_smem_init opt-in and graph_destroy's cudaGraphExecDestroy). What remained, and how each is closed:

sitewasnow
launch_mmq_raw_nt (kd <= 4 and kd > 4)cudaFuncSetAttribute's return discarded; the launch followed unconditionally; the void launcher reported nothing and the caller could not tellminfer_smem_optin reads and names the call, the following launch is refused (return 0), and the launcher's own <<<>>> error is read too (minfer_launch_ok); launch_mmq_raw_nt returns int and the Rust caller turns a 0 into an Err
launch_mmq_nt (MMQ_LAUNCH, one call per quant type)same, as a macrothe macro branches on the named opt-in, refuses, and returns the launch's own result; the launcher returns int → Err
launch_mmq_raw_nb_ntthe return was not read; a following cudaGetLastError() treated any latch as "smem/reg cap, return 0", discarding the code and its originthe opt-in's own return decides (attr:mmq_raw_nb), the launch's own error is named (launch:mmq_raw_nb), the 0-fallback contract is unchanged
launch_mmq_raw_nb_bt_nt / ..._q6k_ntsame post-hoc cudaGetLastError() idiom, attribution by positionthe same named opt-in + launch check (attr:…, launch:…), cleared at the site
launch_mmq_raw_wide_nt (kd <= 4 and kd > 4)the return was read, but a failure cleared the latch silently (no name, no value)routed through the shared named helper: site, instantiation, requested bytes, device limit, cudaGetErrorName
gemm_smem_optin + launch_gemm_f16the helper printed a failed opt-in but the caller launched anyway; launch_gemm_f16 returned void, so its own launch was unchecked (launch_gemm_f32a checked it via a post-hoc cudaGetLastError)gemm_smem_optin returns the admitted/refused answer (cached per instantiation), the launcher refuses on false, reads its own launch error and returns int; launch_gemm_f32a forwards it; prefill_gemm_f16_inner returns Err
graph_end_capture_to_execcudaGraphDestroy(graph)'s return discardedread, named with cudaGetErrorName, cleared at the site; the legacy graph_end_capture gets the same read

The mechanics (one place, no per-launcher copy). cuda_kernels.cu gained a small block of shared helpers — minfer_smem_optin (skip an over-limit request without calling, otherwise call once, name the function/attribute/bytes/device-limit/cudaGetErrorName, clear the latch, return false → refuse), minfer_launch_prelude (report a latch that predates the launch, so the post-launch read is a launch check and not attribution by position), minfer_launch_smem, and minfer_launch_ok (name the launch's own error and clear it) — plus minfer_site_fail_* introspection for the gates. minfer_smem_optin queries cudaDevAttrMaxSharedMemoryPerBlockOptin once and skips a request above it, which is what keeps compute-sanitizer clean (calling it would only observe cudaErrorInvalidValue).

Two signature changes, both followed through. launch_mmq_raw_nt and launch_mmq_nt went void → int, and launch_gemm_f16 / launch_gemm_f32a went void → int (1 = launched and accepted, 0 = refused). The Rust callers read the result and return Err — a failed launch is an error, never a silent pass over an unwritten output. The fallback launchers keep their documented 0 = clean fallback contract, so which kernel runs is unchanged. Nothing numeric or dispatch-related moved.

A measured correction to a documented belief. docs/CUDA-BACKEND-DESIGN.md said cudaFuncSetAttribute is illegal under cudaStreamCaptureModeGlobal and that a lazy opt-in inside a window fails. A probe on CUDA 13.0 / driver 580.178.04 / sm_121 (/tmp/fix147_attr_capture_probe.cu) shows the call now returns cudaSuccess inside an open Global capture window (both for the set value and a new one), so the eager init is defence-in-depth rather than the only working form; the launcher caches one answer per instantiation, so it is not re-asked on the hot path either way. Recorded in the design doc.

Acceptance results (GB10, sm_121, CUDA 13.0, driver 580.178.04).

checkbeforeafter
CUDA serial unit suite503 / 0 / 32508 / 0 / 32 (five new gates: three pure, two device/env-gated)
compute-sanitizer --tool memcheck over the serial CUDA unit suite0 API errors over 503 (349.46 s)0 errors over 508
CUDA serial ignored, 0.5B config32 / 032 / 0
CUDA serial ignored, Qwen3-0.6B Q8_032 / 032 / 0
CPU cargo test --release440 / 0 / 29 unit + 10 / 0 / 6 integrationunchanged
CPU serial ignored29 / 029 / 0
scripts/check_docs_links.py940 links / 184 files940 / 184

The five new gates, all mutation-checked.

  • the_graph_destroy_failure_message_names_the_matching_destructor (pure) — the Rust formatter names cudaGraphDestroy, the cudaGraph_t from cudaStreamEndCapture handle, the cudaGetErrorName symbol and the ticket, and must not name cudaGraphExecDestroy or "kernel launch" (a gate that only asserts "a message appeared" cannot see a wrong message). Mutation: naming the exec destructor fails it.
  • the_injection_matcher_matches_only_the_named_site (pure) — all, exact comma-separated tokens, whitespace-tolerant, and no substring rule in either direction (small must not arm attr:mmq_nt; attr:mmq_nt_extra must not either). Mutation: the substring form fails it.
  • cuda_issue147_attribute_sites_name_the_call_and_refuse_the_launch (device; env-gated behind MINFER_TEST_ISSUE147=1 because it makes a real CUDA call fail) — for each of the ten dynamic-smem tokens the injected over-limit cudaFuncSetAttribute is reported once, at that site, with the API, the attribute, the device limit, a request above it, the exact instantiation and cudaErrorInvalidValue; the launcher returns 0; take_last_error() is 0. Then a positive control (knob off) launches and leaves no latch.
  • cuda_issue147_launch_sites_name_the_call_and_refuse_the_launch (device; env-gated) — the same for the ten launch tokens (an over-limit dynamic-smem <<<>>>, which the launch call itself rejects with cudaErrorInvalidValue — probed before use, /tmp/fix147_launch_fail_probe2.cu — so the kernel never runs), plus a positive control.
  • cuda_issue147_graph_destroy_failure_is_named_and_the_exec_survives (device; env-gated) — the injection re-creates #145's bug at this site (cudaGraphDestroy(exec)), the site clears the latch, and the valid exec is still returned and still destroyable (a failed graph destroy leaks only the graph handle, so refusing the exec would be wrong).

Mutation evidence. Every hardening was reverted one at a time and the corresponding gate re-run; each reverted version failed, the file was restored, and the restored files were byte-identical (sha256sum, in the closing comment on #147). Round 1 (per-site attribute guards): A1 attr:mmq_raw_nt_kd4, A2 _kd8, A3 the mmq_nt macro, A4 attr:mmq_raw_nb, A5 attr:mmq_raw_nb_bt, A6 attr:mmq_raw_nb_bt_q6k, A7 attr:mmq_raw_wide_kd4, A8 _kd8 — all eight trivially failed the attribute gate at that site. Round 2: A9 the gemm opt-in guard, B the shared minfer_launch_ok check, C the gemm message naming a wrong instantiation (the "prints a wrong message" check), D minfer_smem_optin reporting but admitting the launch (the shape a message-only assertion would miss — the gate's ret == 0 catches it), E the Rust destroy read discarded, F the Rust formatter naming the wrong destructor, G the injection matcher weakened to a substring.

Honest scope. The over-limit skips and the device limit are measured on GB10/sm_121 only; the decision is a runtime query and every skip is printed, so another device skips a different subset. All deliberate-failure evidence is device-gated and driven by test-only env knobs, which invert a site's own answer rather than exercising a genuinely failing driver call in production — the injection makes the call fail for real (over-limit attribute / over-limit dynamic smem / wrong destructor), so the "latch cleared" half is exercised for real, but the production paths remain latent by construction. The MMQ launch check is shared by the mmq_nt and raw-NT launchers, so the per-launch check was mutated once (B) rather than per launcher; every launch token's own report was observed by the gate. Other void launchers in the file — 65 of them (launch_dequant_f16, launch_convert_f16, launch_gemm_qb_nt, every store/rope/attention/MVQ/MVQ-multi wrapper) — still do not read their own <<<>>> error and were left alone: a failure there is currently reported by the next hardened launch's prelude or by sync as a latched API error, and converting all of them is filed as #162 rather than smuggled in here. And the pre-#147 fallback semantics are preserved deliberately: launch_mmq_raw_nb_nt / _nb_bt_nt / _q6k_nt / _wide_nt still return 0 on any opt-in refusal, so the dispatch falls to the next kernel — the failure is named, not fatal, because a fallback is the designed behaviour and changing it would change which kernel runs.

Every <<<>>> reads its own launch error · #162 — DONE 2026-09-26

Why. The #147 audit enumerated the wider case: 104 <<<>>> sites in src/cuda_kernels.cu (the 65 wrappers in the ticket plus multi-site families and two wrappers the #147 list had not counted) enqueued a kernel and never read the launch's error. A launch that failed for real — an illegal grid/block shape, an out-of-resources configuration, a stale context — latched the error, and it surfaced at CudaState::sync or at the next hardened site's pre-launch check as a latched API error (#145's honest label) with no indication of which launch produced it. grep -c '<<<' on the tree is 122; two of those are the <<<>>> in prose comments, so the audited surface is 120 sites in 76 launch_* owners (the wrappers plus the static launch_gqa_attn_split_batched_kv helper).

What landed.

  • Every site reads its own error, through the shared helper. Each <<<>>> is preceded by minfer_launch_prelude(site, kernel) (which reports any pre-existing latch as not this launch's) and followed by a minfer_launch_ok / minfer_launch_ok_opt read that clears the latch it named. The site token starts with launch: and the report names the kernel instantiation and cudaGetErrorName — e.g. kernel launch gemm_f16_nt_kernel_t<128,32,true> failed: cudaErrorInvalidValue (1) — the launch is refused (#162/launch:gemm_f16_a32).
  • The decision per launcher is severity in the helper, not a signature change. minfer_launch_ok is required: it records a sticky failure that CudaBackend::execute_node — one Rust-side check (CudaState::take_launch_failure), not 104 signature changes and 104 Rust Err arms — drains on both arms and turns into an Err naming the site and the node, so the op never proceeds with a stale output (and the drain keeps a stale record from blaming the next node). minfer_launch_ok_opt names and clears without the sticky for a path with a documented fallback. Split of the 120 sites: 107 required → Err (67 launchers — the matmul, elementwise, rope/store-KV, embedding, dequant/convert, MMVQ, attention and split-attention families, plus the MMQ terminal launchers launch_mmq_raw_nt / launch_mmq_nt and the #147 "no later gate" ones) and 13 documented fallbacks (6 launchers — launch_fa_prefill_f16kv's three window modes fall back to the legacy attention kernel; launch_mmq_raw_nb_nt / _nb_bt_nt (kernels + k-split reduce) / _q6k_nt (kernels + k-split reduce) / _wide_nt are #147's clean fast-path fallbacks; launch_kv_move_rows returns non-zero to its Result caller). The choice per family is a comment at each family in cuda_kernels.cu.
  • The injection lever is shared geometry, so a site's coverage is data. minfer_launch_block(site, dim3|unsigned) replaces the block argument at every ordinary site; when MINFER_TEST_CALL_FAIL names the site the block becomes 4096 threads (over the 1024/block limit) and the launch itself returns cudaErrorInvalidValue for real, the kernel never runs — probed on GB10/sm_121. The dynamic-smem launchers keep #147's minfer_launch_smem lever. Adding a site's token is therefore a one-line string in the driver, not a bespoke mechanism.
  • A scripted audit, wired into CI. scripts/check_cuda_launch_returns.py parses the source (comments and string literals blanked, so a commented-out occurrence is not a site) and requires, for every <<<: an enclosing minfer_launch_prelude("<site>", …) before it, an minfer_launch_ok/_opt("<site>", …) after its statement naming the same token, a token that starts with launch:, and a real-failure lever (minfer_launch_block or minfer_launch_smem) in the launch geometry. It prints its result and exits 1 on any offending line. It runs in the check-docs job (python3 present; the CUDA container's is not guaranteed) together with its own --selftest and a --check-fixture against tests/fixtures/cuda_launch_sites.tsv, the committed site list (line, owner, token, kernel fragment) the device gate compares its driven set against.
  • Two k-split reduce sites are separate tokens (launch:mmq_raw_nb_bt_ksplit, launch:mmq_raw_nb_bt_q6k_ksplit). They share the launcher with the kernel site, whose _opt read returns 0 before the reduce; arming the kernel's token would leave the reduce site unreachable, so the reduce has its own prelude/read/token and the driver arms it alone.

Acceptance results (GB10/sm_121, CUDA 13.0, driver 580.178.04; serial device runs).

checkbeforeafter
scripts/check_cuda_launch_returns.py on src/cuda_kernels.cu104 unchecked sitesempty list: 120 / 120 sites read their own error and carry a lever
CUDA serial unit suite (scripts/cuda_test.sh)531 / 0 / 37 (master 09406ce; the recorded 526 predated #173's +3 and #98's +2)536 / 0 / 37 (+5 gates)
MINFER_TEST_ISSUE162=1 device gate: every audited site driven and named—5 / 0 tests; the coverage test's driven set equals the 118-token fixture
compute-sanitizer --tool memcheck over the serial CUDA unit suite0 API errors0 errors over 536
CUDA serial ignored, 0.5B config37 / 037 / 0
CUDA serial ignored, Qwen3-0.6B Q8_037 / 037 / 0

The sanitizer row's command is compute-sanitizer --tool memcheck --target-processes all target/release/deps/minfer-<hash> --test-threads=1 (wrapping the test binary, not cargo — the cargo/test build tree is not instrumented and --target-processes all around the whole bash scripts/cuda_test.sh chain stalls the harness; the binary form is the one that completes, in 350.17 s).

Mutation checks (each reverted; src/cuda_kernels.cu restored to sha256 da2e00fb79442fcd03bd3618301b4014cd835f832bbafa4153b4af4283dcdbfb byte-identically). Full transcripts in the closing comment on #162; the list: (a) deleting the read at one single-site family (launch_add_f32) fails the audit and the device coverage test (the armed site reports nothing); (b) the same at one switch case (launch:embed_rows__q4_k) and (c) at one branch of a templated family (launch:gqa_attn_split_f16kv__hybrid_causal); (d) making minfer_launch_read report but admit the launch (return true) is caught by the severity test's assert_eq!(rc, -1) on the fa-prefill fallback — a message-only assertion misses it; (e) the message naming a wrong instantiation is caught by the fixture's fragment check; (f) the sticky removed from minfer_launch_ok is caught by cuda_issue162_required_sites_set_the_sticky_opt_sites_do_not and by the node-level test; (g) execute_node's unconditional drain removed is caught by cuda_issue162_the_err_arm_also_drains_the_sticky (an f16 matmul whose Rust wrapper returns Err leaves the sticky pending, so a drain only on the Ok arm would blame the next node); (h) minfer_launch_block made a no-op is caught because no armed site fails and every armed set observes the empty set; (i) an _opt site that also sets the sticky is caught by the severity test's take_launch_failure().is_none(); (j) the audit's lever check disabled is caught by the selftest's "read but no lever" case; (k) the fixture's kernel fragment changed is caught by --check-fixture. An equivalent mutant is recorded too: deleting the trailing cudaGetLastError() in minfer_launch_read changes nothing (the read's own cudaGetLastError already resets the latch), and the gate correctly stays green — the trailing call is belt-and-braces, not the clear.

Honest scope. Nothing was failing on sm_121 before or after: the sanitizer was already 0, so the production paths remain latent and the evidence is that every site can be shown to refuse and name a real failing launch. The injection is a test-only knob (MINFER_TEST_CALL_FAIL + MINFER_TEST_ISSUE162=1, unset in every default, bench and sanitizer run) that drives the site's own geometry illegal; it does not exercise a genuine driver fault. The source audit is static: it proves the read and the lever are written, not that they run — a <<< inside a string literal or a macro the parser cannot resolve is reported, never silently accepted, but the parser resolves a site variable only through a plain = assignment in the enclosing function (the launch_site ternary of launch_gemm_f16 resolves to its first arm, which is why the driver arms gemm_f16_a32). The device gate's coverage assertion compares sets of site tokens; where two source sites share one token (the two mmq_raw_nb_bt kernel instantiations) the gate proves the token is reached, not that both branches were — the audit, not the gate, is what guarantees each source site has its own read. The compute-sanitizer and real-model rows are recorded measurements on dgxspark (CI has no GPU); the x86_64 CPU row is unaffected because no pure-Rust test was added.

C5 — Session save and restore · #43 — DONE 2026-09-22

Why. A session's KV rows, ownership and run table lived only in memory, so every restart re-prefilled the whole context.

What landed:

  • src/graph/kvsession.rs is the container. An 8-byte magic, a version, flags, the backend tag and the shape (n_layer, n_ctx, n_embd, row_elems), then one K/V blob per layer as raw little-endian pool words, then the bookkeeping, then an FNV-1a checksum. It streams both ways: a save never holds a second copy of a 6 MB–1 GB arena, and a load materializes one layer at a time.
  • KvCache::session_state / restore_session is the bookkeeping half — the owner table, n_used, the run table (SeqSlot: start, cap, shared prefix, written extent), each sequence's span list, identity and the C3/C8b counters — kept separate so it round-trips in unit tests without a device. restore_session validates before it applies: arena capacity, the owner table's length, every reservation and span inside the arena, and every live sequence carrying a span list.
  • GraphAllocator::kv_save / kv_load move the bytes through the backends' read_host/write_host (CUDA included) and enable the pool a file names if it does not exist yet, because a restore happens before the first graph — which is also how the CUDA run of the gate found the gap. The header records the KV element type (f32/f16/q8_0) in its flags word (0/FLAG_F16/FLAG_PACKED — C5 S3, #130), so a file written under one width cannot be resumed under another.
  • A failed load is a no-op. kv_load runs kvsession::verify — a full pass over the header, every layer's declared length, the bookkeeping, the checksum and end-of-file — before ensure_kv creates a single region. Every refusal path is asserted to leave no arena behind (kv_n_used(0).is_none()).
  • Truncation is caught by construction, not only by the checksum: every read is exact against a length the header fixes, so a short file fails wherever it stops with "truncated — the file ends inside …". A version bump, a bad magic, trailing bytes, an unknown flag, a header whose packed flag and cell width disagree, and a file whose backend / n_ctx / n_embd / element type does not describe this run are each refused with the reason.

Acceptance, as measured (2026-09-22):

  • Container, in CI: 9 unit tests — a round trip including the run table, a packed session, truncation, a version mismatch, a flipped byte (checksum), trailing bytes, a foreign file, a header whose flag and width disagree, and a writer short a layer.
  • Allocator, in CI: the region bytes and the run table survive a save and a load into a fresh allocator, the restored table resolves the same cells, and four refusal paths (another n_ctx, another row width, another backend, another element type) plus a truncated file each leave the allocator untouched.
  • Real model (0.5B q4_0, #[ignore]d, run on the CPU and on CUDA): prefill, save (24 layers / 256 cells / 5 written / 6 316 748 bytes), drop the cache, restore into a fresh one, then 8 greedy steps against the session that never left memory — max |Δlogit| = 0, i.e. bitwise, which is the strongest form of "the same continuation".
  • Handed off: resuming the CLI's --session (which still re-prefilled its history JSON) from this container, and an E2 slot-table snapshot for the server, are #89.

C5 S2 — the CLI resumes, the host state rides along (2026-09-24).

The container described the KV rows; the rows belong to a host state, and a restore that brought one back without the other would be a different session. S2 closes that gap for the CLI and records what the server still needs.

  • The container gained a host-state section (version 2). An opaque, length-prefixed blob after the bookkeeping and inside the checksum (KvSessionWriter::set_host, KvSessionReader::finish → KvSessionBody), so the two halves travel together or not at all; a version-1 file has no such section and is refused loudly, which is exactly what the caller's fallback path is for. kv_save_with_host/kv_load_with_host carry it through the allocator; kv_save/kv_load stay as the no-host wrappers every existing gate uses.
  • Conversation::snapshot / restore_snapshot is the host half: the message list, stream_tokens, current_pos, turn_pos, prev_tokens and need_insert_eot, as a versioned JSON blob (SNAPSHOT_VERSION), with validation before it is applied — a snapshot whose current_pos disagrees with the token mirror, or that does not fit this run's n_ctx, is refused rather than half-applied.
  • --session FILE (under --cnv) writes FILE.kv on exit and resumes it on start. Everything that could make a resume wrong falls back to re-rendering the JSON with the reason printed: another n_ctx/model/MINFER_CACHE_TYPE, a history the user edited between runs (the snapshot and the JSON must agree), an unknown snapshot version, a missing companion, an engine that cannot hand its KV to the host (a speculative session — the draft keeps its own KV — and the mocks), and any file the container itself refuses.

Acceptance, as measured (2026-09-24, Qwen2.5-0.5B Q4_0, CPU, --n-ctx 1024, 1530-character history, greedy, -n 8):

  • Continues alike: turn 2 in a new process resumed from the companion answers The answer to 2+2 is — byte-identical to the same turn in the process that never left memory, and to a run with the companion moved away (the re-seed path). So the resume is behaviour-preserving, not merely fast.
  • And prefills nothing at startup: the resumed run prints resumed 2 message(s) and 379 KV row(s) from …seed.json.kv (25 268 624 bytes) — 0 tokens prefilled. Wall clock to the identical continuation: 0.47 s resumed vs 2.36 s re-seeded (5.0×), the difference being the ~370-token history the JSON path re-renders.
  • Mismatches are loud: --n-ctx 512 against the 1024-cell companion prints KV session: the file describes a 1024-cell arena, this run has 512 (--n-ctx); re-seeding the history instead and continues correctly.
  • Gates, each mutation-checked: a_resumed_snapshot_prefills_nothing_and_continues_alike (the restored run issues the same engine calls, at the same positions, as the in-memory run — forgetting prev_tokens in restore_snapshot fails it), the_host_state_round_trips_and_is_covered_by_the_checksum (a flipped byte inside the host blob is refused; dropping the host from the writer fails it), a_version_1_file_is_refused_so_the_caller_can_re_seed, a_snapshot_that_contradicts_the_host_mirror_is_refused, and an_engine_without_a_kv_refuses_the_session_calls. Suites: CPU 287 passed / 0 failed / 15 ignored (was 282/0/15).
  • Not verified here: a CUDA or Metal session companion (the container is backend-tagged and the CPU path is what dgxspark measured; the CUDA half of C5's own gate was run at S1).

C5 S2b — the server's slot snapshot (2026-09-24).

server::batch keeps a per-slot table (seq, start, cap, cached_tokens) that admission rebuilt from scratch, so a restart dropped every slot's context even though the rows were recoverable in the same container.

The design decision, stated because the issue's wording admits two readings. A Run holds the response channel to a client that a restart has already disconnected, plus its RNG and the row's logits — none of it survives a process boundary, and "resuming" a stream nobody is reading would be a fiction. So the snapshot carries the context, not the in-flight request: each slot's reservation and the token sequence its rows hold. A request that arrives after the restart with the same prefix is admitted onto the restored rows and prefills only its own delta — B2's cross-request prefix reuse, with the rows coming from disk. That is the property the acceptance measures ("resume the slot, and get the same continuation").

  • SlotRow / SlotsSnapshot (server/batch.rs) are the table, as versioned JSON in the container's opaque host section — the same mechanism S2a added, so the two halves cannot drift apart and a version this build does not know is a refusal.
  • BatchEngine::save_slots(path) writes the table and the arena through kv_save_with_host; load_slots(path, model) loads it back, validates it against this run (another --n-slots/--n-ctx, another model, another KV element type, a table whose sequence ids or written extents disagree with the arena it rode in with), and installs the slots.
  • When it is written: after every completed request (finish). That is when a slot's context is stable and it is the last moment the rows are known-good — so a server killed without warning still resumes the conversations that had finished, which shutdown-time saving would not give. The cost is one arena write per completed request (--slots-file prints the size at startup; ~12 MiB for the 0.5B/512-row fixture), which is why the flag is opt-in.
  • --slots-file <PATH> (CLI) → server::run → the worker, which loads at startup and rewrites on completion, printing what it resumed or why it started empty. The serial (non-batched) path has no shared arena to snapshot and says so loudly instead of writing nothing quietly; the batched engine is the one with cross-request prefix reuse to restore.

Acceptance, as measured (2026-09-24, Qwen2.5-0.5B Q4_0, CPU, 512 rows over 2 slots):

  • Resumes without re-prefilling: the cold run feeds 5/5 prompt tokens; the restored engine feeds 1/5 (4 reused) for the same prompt, and its next turn — a continuation carrying the whole conversation — feeds 6/11 (5 reused). A re-render would have fed all 11.
  • Same continuation: the restored run's generated text and token count equal the run that never stopped (a_slot_snapshot_resumes_the_context_without_re_prefilling, #[ignore]d, on the cached model).
  • Refusals, naming both numbers: a 2-slot snapshot into a 1-slot engine (--n-slots), and a 512-row snapshot into a 1024-row server (--n-ctx).
  • Gate, mutation-checked: dropping the token mirror from load_slots makes the resumed prompt feed 5/5 again and fails the gate. Suites: CPU 288 passed / 0 failed / 16 ignored (was 287/0/15).
  • Sharing: kvformat.rs gains expect_for(model, n_ctx) / backend_of(device), and the CLI's --session engine now uses them too, so "what this file must match" has one definition.

C5 S3 — the container encodes the f16 element type · #130 — DONE 2026-09-25.

The header already carried the KV element type in its flags word, but only Q8_0 had a bit (FLAG_PACKED). An f16 region therefore wrote flags == 0 and read back as f32, so kv_load's header.format != live check refused the file its own writer had just produced — on any CUDA box whose model crosses the f16 auto-policy threshold (Qwen3-0.6B: 28 × 1024 = 28672 ≥ 8192), for --slots-file and --session alike. The real-model gate showed it as "the file was written with the f32 KV element type, this run uses f16".

The design decision: a flag bit, not a format field. FLAG_F16 = 1 << 1 is added, and flags_of / format_of_flags are each other's exact inverse. A format field would have had to either renumber FLAG_PACKED — changing the meaning of every Q8_0 file already on disk — or move the type elsewhere in the header, shifting the byte layout a version-2 reader is already parsing; both are the silent misread the container exists to prevent. A new bit keeps the flag word's layout (and every existing file's meaning), and the compatibility story follows for free: a pre-#130 build reading an f16 file sees an unknown bit and refuses it loudly at the unknown-flags check instead of decoding the region as f32, while a pre-#130 f32 file (flags == 0) still loads as f32. The two element-type bits are mutually exclusive — packed | f16 describes two incompatible cell layouts, so it is refused by name, never resolved by preferring one bit. No version bump: VERSION stays 2, because a bump is for a layout change and this is an additive flag whose older-reader behaviour is a loud refusal — exactly as FLAG_PACKED was when it landed.

What landed.

  • src/graph/kvsession.rs: FLAG_F16, KNOWN_FLAGS, the flags_of / format_of_flags pair, the writer encoding the format into the flag word, and the reader's unknown-bit + mutual-exclusion refusals before it decodes.
  • Six new gates, five in kvsession.rs (no device — the CI-covered half) and one in alloc.rs (kv_save → kv_load under an f16 policy): the f16 round trip, the flag-word sweep over all three formats (encode and decode), the both-bits refusal, the unknown-bit refusal, the legacy flags == 0 → f32 load, and the allocator-level f16 round trip.

Measured (GB10 sm_121, CUDA 13.0, serial).

checkbeforeafter
kvsession unit tests1116
CPU cargo test --release432 / 0 / 28 unit + 10 / 0 / 6 integration438 / 0 / 28 + 10 / 0 / 6
CPU serial ignored28 / 028 / 0
CUDA serial unit suite495 / 0 / 31501 / 0 / 31
CUDA serial ignored, 0.5B config31 / 031 / 0
CUDA serial ignored, Qwen3-0.6B Q8_030 / 131 / 0 — a_slot_snapshot_resumes_the_context_without_re_prefilling resumes its own snapshot

Mutation checks (each reverted byte-identically, sha256sum): breaking the encode (f16 → 0) fails the flag sweep, the container f16 round trip and the allocator f16 round trip (435/3); breaking the decode (dropping the FLAG_F16 branch) fails the same three; removing the mutual-exclusion check fails only the both-bits gate — and it fails on the message, because the width check would otherwise refuse the file for a different reason (that is the "passes for the wrong reason" hazard this gate's assertion closes); widening KNOWN_FLAGS to !0 fails only the unknown-bit gate; and encoding f32 as FLAG_F16 fails the legacy gate plus the two f32 gates. rustfmt --edition 2021 --check clean on both files (stable rustfmt 1.9.0 — the pinned 1.97.1 toolchain has no rustfmt component here; CI runs no fmt job).

C6 — Logical positions (positions ≠ cells)

Why. §14 row 9 closed the cell-offset sensitivity by measurement: a sequence's logits' tail changes when its run moves because positions is both the RoPE angle and the KV cell index, so a cell move shifts every rotation. The intervention experiment proved that RoPE is the only entry (offset_divergence_is_caused_by_the_rope_rounding_alone: injecting run A's 48 rope outputs into run B makes the logits bitwise identical), and the sweep showed the effect saturates at a ~1e-6 distributed perturbation (a_distributed_rope_perturbation_saturates_the_logits_tail). Two consequences drive this ticket: C3's compaction needs a K re-rope (and can never be bit-identical), and any future cross-slot sharing inherits the same coupling.

Semantics. positions becomes what a caller naturally has — the token's index within its sequence — and the allocator resolves, per token, the KV row it must be written to:

InputMeaningConsumed by
positionssequence-relative token indexRoPE (q and k), the causal bound
cells (new)the row KvCache resolves for (seq, position)KvcacheStore and the fused decode QKV family (Op::FusedQKV, Op::QkvBiasRopeStore)
attn_spanunchanged: the [lo, hi) cell rangeattention

The single-sequence case is unchanged by construction: its run starts at cell 0, so cells[t] == positions[t], and the classic path stays bitwise.

Invariants / gates.

  1. every existing single-sequence test stays bitwise (regression);
  2. batched forward == the same sequences run one at a time, same layout, bitwise (the existing gate, migrated to relative positions);
  3. a compaction between steps leaves the logits bitwise identical — the literal acceptance C3 could not claim before, and the reason this ticket exists;
  4. a backend that cannot resolve cells refuses the node (Err, no silent fallback); Metal is out of reach by construction because it already refuses multi-sequence attention (supports_attn_span() == false), so it never sees a non-zero run start and keeps positions == cells (recorded as G5, no Metal code change in this ticket);
  5. decode timing on GB10 does not regress (the fused QKV family is the only hot path this touches, ported in S3 with its own A/B).

Change surface.

  • IR: GraphBuilder::kvcache_store wires its row input to a new cells node (created like attn_span); the fused QKV family (FusedQKV, QkvBiasRopeStore, FusedQkvNorm) is gated off while explicit_span is set (S1) and takes cells in S3.
  • Allocator: fill_batch_inputs / fill_attn_inputs / fill_seq_ids stop treating positions as cells — own_range, kv_note_used and attn_span resolve through the run table — and fill cells. (History: E1's test-only fill_attn_inputs and the kv_note_used → own_prefix C1 remnant were deleted in #228; fill_batch_inputs is the one production fill entry point.)
  • KV store: KvCache::cells_for (today the identity) and attn_span (today requires position ∈ run and owner[position] == seq) become the resolver; check_positions_bound / check_attn_span check the cells form.
  • Models: forward_batch's explicit_span decision and positions stay as they are semantically; the store wiring is inside kvcache_store.
  • Server: positions = slot.start + current_pos becomes current_pos; SlotState.start drops to reporting/diagnostics.
  • Backends: the KvcacheStore arms need no change — they already write at the rows their third input names; only that input's producer changes. This is what makes S1 landable with the suite green on all three backends.
  • C3: kv_defrag(need, rope) loses the rope parameter, and the KvRope plumbing added for the re-rope goes away.

Staged plan (each step: local CPU suite, plus the CUDA suite when it touches CUDA, then a docs progress update, then a commit on feat/logical-positions).

StepContentGate
S0this design, the roadmap/status correctionsdocs build (check-docs)
S1cells wiring + allocator/kvcache resolver + model/server switch + fused QKV gated off under explicit_spaninvariants 1–2, 4; CPU+CUDA suites
S2compaction without a re-rope; the bit-identity gateinvariants 1–3
S1∪S2merged in execution: the re-rope removal is not optional once S1 lands — S1 makes the stored K rows' angles sequence-relative, so a compaction that still re-roped them by from - to would rotate them away from the correct angle. The first S1 run showed exactly that: a_compaction_between_steps_keeps_the_continuation failed with a flipped greedy token until kv_defrag stopped re-roping. Semantics switch and re-rope removal must therefore land in the same commit.as above
S3CUDA fused QKV family takes cells; the gate re-enabled for CUDA — landed, plus a server position bug the gate exposedinvariant 5 + fused-vs-unfused bitwise + capture regression
S4docs closure (roadmap §2.4, AGENTS rules, design docs)docs build (check-docs)
S5PR, CI four jobs green with zero annotations, rebase mergeCI

The docs build gate is the CI job check-docs (.github/workflows/ci.yml): mdbook build with the deployed toolchain plus scripts/check_docs_links.py. It did not exist when S0/S4 ran — the docs were checked by hand then — and it landed with #63; the rows above name it so the record and the gate agree.

C6 progress (2026-09-19)

S0 done (71b86da). S1 + S2 landed — the semantics switch and the re-rope removal went in as one commit, as the merged-step row above requires. Suites: CPU 213 passed / 0 failed / 5 ignored, CUDA 261 passed / 0 failed / 5 ignored.

What S1 changed (production):

  • GraphBuilder::cells_input; kvcache_store creates and consumes it, so callers no longer pass a row buffer (26 call sites migrated);
  • KvCache::attn_span resolves cell = run.start + position; GraphAllocator::kv_cells_for_seq fills cells, with the classic single-sequence identity fallback when no run is reserved;
  • fill_batch_inputs / fill_attn_inputs / fill_seq_ids use relative positions, and fill_attn_inputs only resolves cells when the graph has that input (a rope-only fixture must not need a KV arena). (History: fill_attn_inputs was deleted in #228, which moved its corpus onto fill_batch_inputs; the rope-only branch survives as the #[cfg(test)]-scoped GraphAllocator::fill_attn_inputs_without_cells.)
  • qwen2/qwen3 gate the fused QKV family off while explicit_span is set (S3 ports it for CUDA), and server/batch.rs passes relative positions;
  • kv_defrag no longer re-ropes and the KvRope plumbing is gone (S2).

What the tests became, i.e. the new gates: the minimal hand-built attention graph asserts bitwise invariance to the cell placement (it used to be "within the rotation's rounding"); the four offset-sensitivity tests invert into "a cell placement changes nothing" (logits, per-layer V, per-node q/k/attention); and the perturbation probe was deleted, because its premise — that the offset effect needs explaining — is gone.

S3 done — the CUDA fused decode QKV family takes cells:

  • attn_bias_rope_store_f32 gained a cells parameter: pos rotates q/k and row = cells[0] addresses the four KV store writes. The launcher, the FFI declaration and CudaState::attn_bias_rope_store carry it; the builder wires the shared cells input into fused_qkv (sources [x, pos, cells]) and qkv_bias_rope_store ([q, k, v, pos, cells]), so no model call site changed;
  • the qwen2 gate becomes (cuda_on || !explicit_span): CUDA keeps the fused chain under an explicit span, Metal keeps the pre-C6 gate (it has no explicit-span attention at all, G5). Qwen3 needs nothing: its fused family is Metal-only, so no CUDA arm existed to port. The hand fixtures in cuda.rs::d38_probe_tests pass their positions buffer as cells — identity by construction, which is the case they assert.

Suites: CPU 213 passed / 0 failed / 5 ignored, CUDA 261 passed / 0 failed / 5 ignored. End-to-end gate on GB10 (Qwen2.5-0.5B Q4_0, --n-slots 2 --n-ctx 1024, one short + one long request submitted together): cap = n_ctx/2 = 512, so the long request ran in slot 1 (server log) and, once the short one answered in a token, alone — a single-token (nt == 1) step in a run whose start is not 0, which is exactly the path the S1 gate had closed. Fused vs MINFER_NO_FUSE_QKV=1 produced byte-identical completions (sha1 36c55ed81464, 319/319 expected words): the cells port and the S1/S2 KvcacheStore path agree on a non-zero run start.

The gate also caught a real server bug, now fixed: BatchEngine::submit_on still built positions as start + i — the pre-C6 "positions are cells" semantics — so a request placed in a non-zero-start slot fed position 512 into a 512-cell run and kv_cells_for_seq rejected it ("past sequence 2's reserved run"), rejecting the job. Both the batch prefill path and current_pos were already relative; only this entry point had not been migrated, and no earlier test exercised a second slot's run (the CPU suite batches nothing by default, so every server test ran the single-sequence path). It now passes feed_from..nt.

Regression coverage: server_batch_matches_serial_and_is_faster pins each of four requests to the slot it would occupy through submit_on, so it drives runs with non-zero starts — it is #[ignore]d only because CI has no cached model, and it must be run locally for any change to positions or the run table: cargo test --release -- --ignored server_batch_matches_serial. It passes both on CPU (byte-equality of batched vs serial, 1.43x) and on CUDA (structural assertions, 1.28x).

One measurement note worth keeping: an end-to-end A/B of the fusion on the batched decode path cannot isolate this gate, because fuse_qkv requires nt == 1 — a step carrying four streams is unfused in both configurations. The measurement that matters is the single-stream step in a non-zero-start run above; a 4-slot throughput A/B (master, 7B Q4_K_M, 4 concurrent requests) showed the fused and unfused chains within noise (92.6 vs 92.7 tok/s best case), which is expected for a bandwidth-bound model and is why the port's value is correctness and parity, not throughput.

The two pre-C6 experiments that established the attribution above — the RoPE injection (EXP1) and the distributed perturbation sweep (EXP2) — are archived with their code and measured numbers in experiments/logical-positions/.

Not in this ticket: sharing one cell range across sequences (owner[cell] becomes a set/refcount — the real cross-slot prefix reuse, which this ticket makes possible), C4 (quantized KV) and C5 (state save/restore) — both should follow this, because each would otherwise multiply the addressing surface.

  • Explicitly not changed: C2's kv_rm/kv_shift re-rope stays: it changes positions, not cells, so the angles genuinely have to move. Only the compaction re-rope disappears.

C7 — Dynamic runs: one request may use the whole arena · follows C6 · M · #59

Why (user-visible). BatchEngine::new hands every slot a fixed cap = n_ctx_total / n_slots up front, and submit_on rejects anything whose prompt + generation exceeds that cap (prompt of N tokens exceeds slot context of M). With --n-slots 4 --n-ctx 8192, a single 6000-token request is refused while three slots sit idle and their rows are free: the arena is sized for the total, the partition is the limit. This is the most user-visible consequence of the fixed reservation and the reason C3's "server-side dynamic-run trigger" was recorded as a follow-up instead of being closed with C3.

What. Make the partition elastic instead of fixed:

  • a slot may grow past its initial cap when the admitted request needs it, via kv_reserve_seq_with_defrag(seq, need) — C3 already plans the move, copies the rows through Backend::copy_cells and returns the moves;
  • the shrinking slots are re-reserved around the growth, and every caller that cached a run start (E2's SlotState) applies the returned moves — the contract kv_defrag already documents. C6 is what makes this safe: a moved row keeps its sequence-relative position, so no other slot's logits change and nothing re-ropes;
  • the repartition happens at one serialization point (admission), so no in-flight step can observe a half-moved arena;
  • a genuinely full arena still rejects loudly with the existing message shape — one request never silently shrinks another's context.

Gates. (1) --n-slots 4 --n-ctx 8192 serves one request of ~8k tokens while the other slots are idle (the literal requirement); (2) after a grow, each idle slot's next request is bitwise identical to the same request on a fresh engine; (3) the 4-slot batched-vs-serial byte-equality test stays green (server_batch_matches_serial_and_is_faster, run locally with --ignored); (4) the CUDA batched-decode throughput stays within noise of the recorded numbers.

Not in this ticket: sharing (C8) — C7 only makes the partition elastic, two sequences still never look at the same cell.

C7 increment 1 (2026-09-20) — the partition is elastic

What landed. BatchEngine now sizes a slot from the request instead of from the startup partition: wanted_cells_from (pure, unit-tested) asks for prompt + max_tokens + 1 for a bounded request and prompt + slot cap + 1 for an unbounded one, clamped to the arena and never below the prompt. When that exceeds the slot's cap, ensure_slot_capacity reclaims the idle runs above it (they hold at most B2's prefix hint, which the slot re-prefills when it is next used), then re-reserves this slot's run at its existing start and re-owns its written prefix; the moves the KV store returns are applied to every slot's cached start (apply_run_moves, the contract C3 documented).

No new KV primitive was needed — kv_release_seq, kv_reserve_seq_with_defrag and kv_own_range compose into "grow in place" — and no backend changed, because the store already writes at the allocator-resolved cells (C6). A reservation that would land anywhere other than where the rows physically are is refused and the slot is left to re-prefill: the one outcome C6 exists to prevent (reading rows at an offset they do not have) cannot happen. The HTTP layer's request bound moved from a slot's share to the whole arena (AppState::n_ctx); the serial path (batching off) keeps its own slot-sized bound, because its graph region really is that size.

(Forward note, 2026-09-29: kv_own_range in the sentence above is GraphAllocator::kv_own_range, which #232 deleted as dead surface — it was a one-line forwarder to KvCache::own_range, the function production actually calls (through fill_batch_inputs → own_positions, and through kv_copy_prefix). The sentence records what this step landed with; it is not a live API.)

The boundary, and why it is no longer padded. tick forwards the committed token before advance can report length, so a generation that reaches its cap used to perform one more forward at the cell after its last token — a whole weight pass whose logits are discarded, plus a cell nothing reads. The reservation was therefore padded with one cell, and a reservation of exactly prompt + max_tokens rejected that batch (kv_cells_for_seq: position N is past sequence S's reserved run).

That padding is gone: advance now decides before committing whether another token can still be used (the request's token budget, and current_pos + 1 < cap), and ends the turn instead of committing one that could only be discarded. wanted_cells_from is exact — prompt + max_tokens, clamped to the arena — and the top-of-advance bound is the safety net it always was. The regression is a request with no token budget (max_tokens = -1) on a run sized to its prompt plus 48 cells: it must end with length and exactly 48 tokens, because a forward past the run is rejected by kv_cells_for_seq — before this change that request failed instead of finishing.

Gates. (1) wanted_cells_plans_from_the_request — the policy, no model needed; (2) a_long_request_may_use_the_whole_arena — a 302-token prompt on n_ctx = 366 with four slots (91 cells each) is served, and both admission paths (submit / prefill_group and submit_on, the one the server uses) produce a continuation byte-identical to a one-slot engine that needs no reclaim; (3) the four-slot batched-vs-serial byte-equality test still passes; (4) CPU and CUDA suites green; (5) on GB10, --n-slots 4 --n-ctx 8192 with a 2054-token prompt logs slot 0: capacity 2048 -> 2071 cells ... (released 3 idle slot(s) above) and answers the same bytes as --n-slots 1, which needs no reclaim at all. Gate (5) is the point of the ticket: C6 makes a moved row arithmetic-free, so where the partition puts a sequence cannot show up in its output.

C7b — both directions, so a busy neighbour is not a wall (2026-09-20). The increment-1 caveat is gone: Backend::copy_cells and CUDA's kv_move_rows kernel now move rows up as well as down (the kernel walks one row at a time with a barrier, descending when the run slides up), apply_moves accepts to > from and frees the vacated cells in either direction (guarded by the sequence's own stamp, so a plan can never clear another sequence's ownership), and KvCache::set_cap grows or shrinks a reservation and returns the plan — shrinking below a sequence's written rows is refused, and so is a growth the arena cannot hold with the other reservations. The allocator's kv_set_cap_with_defrag copies the rows and then renumbers, like kv_defrag.

The subtle half is order: wherever ranges overlap, a destination must never land on a row that has not been copied yet. kvcache::order_moves is the single rule — upward moves top-down, downward bottom-up, upward first — and both the data copy and the bookkeeping follow it (the owner table is a mirror of where the rows live, so they have to travel together). The unit test caught exactly this: applying the plan in its raw ascending order moved one sequence's stamps through cells another had already overwritten.

Gates. (1) growing_a_run_pushes_the_runs_above_it_up — a pure test of the plan, the upward move and the owner stamps; (2) a_resize_refuses_to_eat_rows_or_overcommit_the_arena; (3) the CPU and CUDA twins that used to pin "upward is refused" now pin the result of an overlapping upward move (a memmove's output) — the CUDA one on GB10; (4) on GB10 the CUDA twin pins the upward kernel path (the bytes a memmove would produce). Suites: CPU 216 passed, CUDA 264 passed (0 failed, 6 ignored each).

The cells bound is its own rule (2026-09-21) — #60. check_positions_bound classified any I32 input consumed by a KV-writing op as positions, and cells — which now feeds the same ops (C6) — was measured by that rule. It cannot false-reject (a cell always indexes the arena, which is exactly n_ctx cells), but the two bounds only coincide because a cell equals its position, which is the coincidence C8 removes. cells now has its own bound and its own message (input 'cells': cell N is past the M-cell arena), pinned by a test that also covers the no-arena case.

End-to-end (2026-09-20) — the engine scenario is verified. Four slots at --n-ctx 8192: a short request, then a long generation on another slot, then a 2148-token prompt admitted while that generation is still running. On GB10 the server logs slot 0: capacity 2048 -> 2157 cells for a request wanting 2157 (released 1 idle slot(s); 2 run(s) moved) — two runs moved, one of them the live generation above — and on CPU (forced with MINFER_BATCH=1, which makes the comparison deterministic) the same scenario moves one run. In both runs the live neighbour's continuation is byte-identical to the same request served alone: its rows travelled up under it and its answer did not change, which is exactly what C6 + C7b promise. The grown request itself is compared structurally (a shared 38-byte opening with its one-slot baseline), because it shares the batch with the live request and therefore takes the windowed (explicit_span) attention path while the baseline takes the causal one — a named tolerance class, not a bitwise property.

That gate was missed twice for an environmental reason worth recording: cargo test --release overwrites target/release/minfer with a CPU-only build, so the server logged batching: off (device cpu) and served the serial path while the script believed it was testing the engine. The script now asserts batching: on and the device line in the log before it measures anything — a gate that cannot check its own preconditions is not a gate.

C8 — Cross-sequence cell sharing (owner → set/refcount) · follows C7 · L · #41

Why. A prefix shared by several sequences (a system prompt on every slot; B2's prefix reuse, but across slots) is duplicated today: each slot stores its own copy and pays its own prefill, so N sharers cost N× the memory and N× the prefill for the same tokens. llama.cpp's unified cache shares one cell between sequences through a per-cell sequence set (seq[i] is a bitset, with seq_cp/seq_rm/seq_keep over it); this ticket is that capability, scoped to what minfer's server needs.

What. KvCache::owner[cell] becomes a set (bitset or refcount); attn_span and kv_cells_for_seq accept "owned by any of these sequences"; a new kv_seq_cp(src, dst, range) shares a prefix's cells with another sequence; kv_rm frees a cell only when its last owner drops it; and the store must not write into a shared prefix — the first divergent token copies the shared cell it would have overwritten into a private one (copy-on-write). The CoW rule is the delicate part: a store that wrote through a shared row would corrupt every sharer, so it is either implemented or refused, never ignored.

Gates. (1) two slots sharing a prefix produce continuations bitwise identical to the same two slots run with private copies (CPU first; CUDA may take a named tolerance class only if a kernel shape differs); (2) removing one sharing sequence leaves the other's logits bitwise; (3) a store that would write into a shared cell copies it or fails loudly; (4) KvArenaStats counts a shared cell once, not once per owner (memory-footprint measurement).

Depends on: C7 (elastic runs — shared prefix plus growth is what the server actually needs); C1's owner table and C6's resolver are the hooks.

C8 design (2026-09-21) — the read path, not the refcount, is the cost

Why. A prefix used by several sequences (a system prompt on every slot, a conversation resumed twice) is prefilled and stored once per sequence. B2 reuses a prefix within a slot; across slots there is no reuse at all. KvCache::owner[cell] holds a single SeqId, so a cell belongs to exactly one sequence.

The constraint that shapes everything. A sequence's rows are one contiguous run, and the read path depends on that: cells[t] = start + position, and attention reads a single [lo, hi) span from attn_span. The write path is already general — the allocator hands the backend a per-token cells vector, so a store may target any row (C6) — but a read cannot follow a sequence whose rows are not contiguous. Now let two sequences share a prefix and then diverge: only one of them can keep a contiguous layout (its tail sits right after the prefix); every other sharer's rows become [0, p) ∪ [private, ...) — two ranges. So true sharing is not "a refcount in the cell store"; it is a paged read path (a per-sequence block map and a gather in the attention kernels of each backend). That is the whole cost of this ticket, and it is why it was estimated L rather than M.

Two increments, because the cheap half is independent of the read path.

StepContentBenefitCost
C8aShared prefill, duplicated rows: a slot that matches another slot's prefix copies its K/V rows with copy_cells instead of re-running the forward, then diverges privately.Removes the N× prefill for a shared system prompt — the latency that is actually paid on the first turn. No IR change, no read-path work, works on every backend that already copies cells (CPU, CUDA; Metal behind G5 — all three since #362, see the S5 record's 2026-10-11 note).Memory stays N×: each slot holds its own copy.
C8bTrue sharing: owner[cell] becomes a refcount; the allocator resolves a per-sequence block map (64-cell blocks) into the per-token cells vector the write path already accepts; attention gains a block gather (supports_attn_span() is replaced by a block-map capability), CPU first, then CUDA; Metal stays behind its gate (G5) — landed on all three in #362.The blocks are stored once, so memory follows the sharing, on top of C8a's prefill win.Every kernel that reads a window needs the gather; attn_span is replaced (or supplemented) by the map; compaction and the arena counters must understand refcounts.

Invariants / gates.

  1. C8a: a prefix copied from another slot produces continuations byte-identical to the same prefix re-prefilled, on CPU and on CUDA (the same named-tolerance caveat as any device comparison).
  2. C8a: the copy must be measurably cheaper than the prefill it replaces (it is the point).
  3. C8b: two sharers are byte-identical to the same two sequences with private copies.
  4. C8b: kv_rm frees a block only when its last owner drops it; a store into a shared block copies it (copy-on-write) or fails loudly — never writes through.
  5. C8b: KvArenaStats counts a shared block once, and a compaction moves it once (not once per owner); the memory saving is measured.

C8a increment 1 (2026-09-21) — the copy primitive. GraphAllocator::kv_copy_prefix(src, dst, rows) copies the first rows written K/V rows of one sequence's run into another's (per layer, through Backend::copy_cells, which handles either direction — C7b) and hands dst their ownership via own_range. It refuses a missing run, a source with fewer written rows than requested, and a destination that reserved fewer cells than that, and it is a no-op for zero rows or a same-run copy. Unit coverage is the region-free half of those refusals; the data path and the written/capacity bounds are covered by the S3 gate, because they need real KV regions and the honest test for a copy is "the answer does not change".

The engine is not wired to it yet. S2 does that at admission: placement already prefers the slot with the best-matching prefix, but only among idle slots and only for the slot's own rows — what is missing is that the donor may be another slot (possibly busy, and its rows are stable while it generates).

C8a increment 2 (2026-09-21) — admission uses it. The reuse source is now every slot, not just the admitting slot's own rows: when another slot holds the longer match, its written rows are copied into this slot's run (the donor may be busy — its rows are stable for the duration of the copy) and only the suffix is prefilled. The reuse is clamped to leave one token to feed, exactly as the slot's own reuse is, because a forward with nothing to run produces no logits. A failed copy falls back to the slot's own cache and prefills the rest — never a wrong-row read.

The test-only counter prefix_rows_copied makes the copy observable, so the gate asserts both that it happened and that it changed nothing: the same prompt served via a copy and on a private single-slot run answers byte-identically (CPU, where the comparison is exact; a device comparison takes the usual named-tolerance caveat). C8a cost gate (2026-09-21, GB10). With a 1450-token prompt on --n-slots 2 --n-ctx 8192 and the donor held busy (a second request generating), the same prompt served through the copy costs 0.013 s against 0.467 s with prefix reuse switched off (MINFER_NO_PREFIX_REUSE=1) — 36.6× — and both answers are identical. Two scenario traps had to be fixed before the number meant anything, and both are easy to repeat: (a) if the sharing slot is idle, the engine simply places the request there and reuses its own cache (B2), so no copy happens at all — the donor has to be busy; (b) if the donor's own request needs more than its share of the arena, C7's growth reclaims the idle slot the measurement needs, leaving one slot and self-defeating the scenario. Each copy is reported for an operator as [server] slot N: copied R prefix row(s) from slot M.

Order. C8a first: it delivers the user-visible half with the machinery that already exists (C3's row copy + C7's moves) and its gate is byte-equality. C8b is then a read-path project, and it can be scheduled on its own evidence rather than on this ticket's estimate.

Not in this ticket: C4 (quantized KV) and C5 (state save/restore) — each multiplies the addressing surface this ticket touches, which is why C6 preceded C7 for the same reason.

C8b design (2026-09-21) — a block map for the read path

What sharing needs that C8a does not. C8a duplicated the rows so the destination could keep one contiguous run. Sharing for real means a block exists once, several sequences refer to it, and a sequence's rows become a list of spans rather than one run (cells[t] = span_of(t).start + offset). The write path needs nothing new — since C6 it takes an arbitrary per-token cells vector. The read path is the whole change: attn_span hands each query a single [lo, hi) range, while a sharing sequence's window is a set of ranges.

Data. KvCache gains (a) a per-cell refcount at block granularity (64 cells — small integers, and the unit that compaction moves and the stats count), and (b) a per-sequence span list covering [0, written). A sequence that shares nothing has a one-entry span list, which is the property that keeps this incremental: its code path stays today's, bitwise.

IR — additive, never a replacement. A graph whose sequences share blocks gets a kv_map input (per query, the sequence's spans) alongside attn_span; attention uses the map when the input is present and the span otherwise. attn_span's single range is load-bearing — the CAUSAL fast path and Metal's refusal (G5) both rest on it — so it is not generalized in place.

Allocator. kv_cells_for_seq and attn_span resolve through the span list. A store that would land in a shared block takes a private row for that token (kv_private_row_for(seq, t)): the allocator decides before the forward, so no backend learns that sharing exists and a shared block is never written through. That is the copy-on-write rule, and it is why the store path did not have to change.

Backends. CPU attention gains the span-list gather first (it is the reference), then CUDA's gqa_attn_* inner loop, whose KV walk becomes (block, offset); Metal stays behind G5 and refuses the map, as it already refuses the explicit span.

Why S2 does not split the way S1 did (2026-09-21, found while sizing it). S1 split cleanly because its two consumers already existed — the write resolver and attn_span — so each half was a live path with a bitwise gate. S2's halves are coupled instead. Sharing means one cell has several owners, and the store's ownership representation is per cell: Layer::owner is a Vec<SeqId>, written_rows scans it contiguously for owner[cell] == seq, and attn_span's written check compares it against the query's sequence. A shared prefix cannot be expressed without changing what "written by this sequence" means (block refcounts plus the span list as the read-side authority), and until the read path gathers over several spans a sharing sequence would resolve a wrong window. So the store half and the read half have to land together: staging them separately would add machinery nothing reads, which is the shape A7 deleted. Concretely S2 = block refcounts (64 cells), kv_seq_cp, refcount-aware release/kv_rm, and a compaction that moves a shared block once and renumbers every sharer's span list — plus the additive kv_map input, the CPU gather, and admission wired to kv_seq_cp, accepted by gate 2 below.

Gates.

  1. S1 changed nothing observable (S1a + S1b landed: both resolvers read the list, and the suites stayed bitwise); the refcounts arrive in S2 together with the readers that justify them.
  2. Two sharers answer byte-identically to the same two sequences with private copies (CPU; CUDA takes the usual named tolerance).
  3. kv_rm frees a block only when its last owner drops it, and a store into a shared block either takes a private row or fails loudly — never writes through.
  4. KvArenaStats counts a shared block once; a compaction moves it once and renumbers every sharer's span list; the memory saving is measured.

Why S1 is a used generalization, not dead bookkeeping. The first draft of this plan put the refcounts in S1 and the span list in S2. That is the shape A7 deleted: the campaign's own note says an identity field earns its place only if some topology decision reads it — "a future feature will need it" is not enough, and n_seqs was removed for exactly that reason. Refcounts are only read by the sharing that S2 introduces, so they move to S2 and S1 becomes the span list plus the two resolvers that read it (kv_cells_for_seq and attn_span). Those are live paths for every request, which is what makes S1's gate meaningful: a single-entry span list must reproduce today's answers bit for bit, and any slip shows up in the existing suites rather than in a field nobody consults.

C8b S1a landed (2026-09-21) — the write path resolves through spans. KvCache carries a per-sequence span list (position base, first cell, length), and GraphAllocator::kv_cells_for_seq resolves every store row through KvCache::cell_of instead of slot.start + position. Every path that changes a run's start/cap or drops the run republishes the list (refresh_spans: reserve, release, resize, and each relocation apply_moves performs), and a position outside the list is a loud error, not a fallback to the contiguous form — which is what makes a missed maintenance point visible rather than silent. Today the list always holds exactly one entry, so nothing observable changed: the full suites keep their counts (CPU 218 / CUDA 267, 0 failed) and every byte-equality gate still passes — the four-slot batched-vs-serial test, the C7 boundary case, and the C8a prefix copy. The unit test pins the resolution against start + position, including a real compaction move and a release.

S1b landed 2026-09-21 — the read path (attn_span) now resolves through the same list. A sequence's window is still one contiguous [lo, hi) range, which is all the current input layout can carry, so S1b resolves only the single-span case and refuses a multi-span sequence loudly (a window that is not one range needs S2's kv_map). With one span the answer is identical by construction: lo is the span's first cell and hi is min(span end, query cell + 1), which for rel < length is exactly the old start + rel + 1. The unit tests pin both halves — the window is compared against the old arithmetic for a run that does not start at cell 0 (where a cell == position slip could hide), and a hand-written two-span list proves the refusal. Full suites after S1b: CPU 221 passed / 0 failed / 7 ignored (two tests added; the S1a figures of 218 CPU / 267 CUDA were each one under the state they described — that state measures 219 / 268), CUDA 270 passed / 0 failed / 7 ignored, and the byte-equality gates (batched-vs-serial, C7 boundary, C8a prefix copy) unchanged.

C8b S2 landed 2026-09-21 — sequences share a prefix in place. A sequence's address space is now its span list: a SharedPrefix { cell, rows } it reads from a donor plus its private run, which holds positions [rows, rows + cap). KvCache::share_prefix establishes that by pointer — no bytes copied — and refuses loudly what it cannot express: an empty share, an unwritten donor range, a destination that already shares or has written rows of its own, itself, and a donor prefix that is not one contiguous cell range (the caller then falls back to C8a's copy). KvCache::attn_map resolves a query's window as KV_MAP_MAX_SPANS (cell, len) runs, and CParams.kv_map selects that layout for the window input instead of attn_span's single range: an input's size is topology, so the difference is fixed at build time, and a size that matches neither form is refused rather than guessed. The CPU kernel gathers the runs (decode_window + cpu_gqa_attn_runs, with cpu_gqa_attn kept as the one-range wrapper); CUDA's windowed arm refuses a window that is not 2 * nt (S4 ports the gather), and both models ask for a map only on CPU. Admission uses it: at C8a's reuse site kv_share_prefix replaces kv_copy_prefix where the device can gather, so the arena holds one copy of those bytes, and every other device keeps the copy. Gate 2 holds: sharing answers byte-identically to a private run, on the same gate that validated the copy (a_prefix_copied_from_another_slot_answers_identically), and server_batch_matches_serial_and_is_faster still passes with the sharing path live. Suites: CPU 228 / CUDA 277, 0 failed.

Two deliberate departures from the design above. (1) No block refcounts. Occupancy (reserve_seq, free_runs) and the written count are derived from the span lists: a released donor's rows stay taken for as long as a sharer's spans name them, they come back when the sharer drops them, and a compaction moves each run's own rows and renumbers every sharer's pointer. A per-block refcount would be a derived cache of exactly that union and could drift from it — the shape A7 deleted — and the four gates it was meant to serve hold without it. (2) The layout rides the window input's size rather than a new op field. A kv_map flag on Op::Attn would have touched every construction site across three backends, the models and the tests for no behavioural gain; the size is the topology here, and every backend that cannot gather refuses loudly instead of misreading pairs.

S3 then adds copy-on-write (kv_private_row_for) for a store that would land in a shared block; S4 ports the gather to CUDA's gqa_attn_* loop; S5 is Metal behind G5 — lifted later by the G5 port and Metal's kv_map kernel (#362); see the S5 record's 2026-10-11 note.

C8b S3 landed (2026-09-22) — a store inside a shared prefix copies the row first. KvCache::private_row_for(seq, t) plans a copy-on-write and apply_private_row books it, with GraphAllocator::kv_private_row_for driving the data half through Backend::copy_cells; both fill entry points (fill_batch_inputs, fill_attn_inputs) run it before the first cell is resolved, and the store resolver (kv_cells_for_seq) now refuses a position inside the share outright. (As of #228 there is one entry point — the production fill_batch_inputs — and the resolver's refusal is a belt-and-braces check it cannot reach with a shared position; the fill_attn_inputs half of that sentence was true when S3 landed.) A sharing sequence's run holds positions [shared.rows, shared.rows + cap), so a store at t < shared.rows gives the share up from t on: shared.rows drops to t and the rows the sequence already wrote shift up by d = old_base - t inside the same run. That arithmetic is what keeps the change small — the span list stays at two entries (the remaining share plus the run), the owner table mirrors the move with one copy_within, and written is unchanged: the sequence's readable positions are still [0, written), only private_written grows by d. KvArenaStats gains cows/cow_cells so a gate can see the mechanism run, and GraphAllocator::kv_cell_of is the read-side twin of the store resolver (a caller snapshotting a sharing sequence's rows cannot use the store one, which refuses those positions by design).

(Forward note, 2026-09-30: kv_cell_of is test-only as of #236 — #[cfg(test)] pub(crate), not allow(dead_code), because it never had a non-test caller. git log -S 'kv_cell_of' --all over src/ names only d43e716 (which introduced the forwarder with this doc), c9bbcbd (the test call) and 54f6de0 (the test-module extraction), so no production call site was ever deleted. The "caller snapshotting a sharing sequence's rows" the sentence above names is served two other ways in production: a sharer's rows are read as windows (KvCache::attn_map → the kv_map input, C8b S2/S4) and a whole-run snapshot goes through the C5 container (kv_save/kv_save_with_host, which stores the whole arena). The one consumer is the test-side observation instrument kv_rows_of (server::batch::tests), driving the C8b S3 gate a_store_inside_a_shared_prefix_takes_a_private_row; see the #236 record in §test-infrastructure. The sentence records what S3 landed with; it is not a live API.)

Why the run is rebased rather than given a fresh row per token. The design's kv_private_row_for (seq, t) says "a private row for that token", and the honest way to give it one is to move the run's base, not to hand out unrelated cells: a per-token fresh cell would put d one-length spans in the span list for a d-token divergence, while kv_map carries KV_MAP_MAX_SPANS = 4 runs — so any real divergence would fail the read path — and a second run per sequence is a data-model change (SeqSlot is one run) that every relocation path would have to learn. Growing the run downward instead of shifting ([start - d, …)) needs d free cells immediately below it, and C7's compaction packs free space upward, so that placement is usually blocked. The shift needs no arena space at all, only cap - private_written >= d, and copy_cells is overlap-safe (C7b), so it is one move inside one run.

The room is not assumed. The shift needs cap >= written - t, and the server's sizing guarantees it: a run is sized for the whole previous request (prompt + max_tokens, and generation stops at max_tokens), so it covers the shared rows as well as the private ones — cap >= rows + w >= written - t. A hand-sized run without that room gets a loud Err (gate 3's second arm), never a write-through.

Why the pre-pass is its own loop. The shift renumbers the run's cells, so a cell resolved before it would be stale; the resolver's refusal is what makes a missed maintenance point visible instead of silent. After the copy the share is [0, t) — still two spans, or one when t = 0 and the share disappears — so CParams.kv_map is unaffected: the builder saw a share at build time, and attn_map handles one entry as readily as two.

Gate 3 holds. Always-run coverage: three KvCache unit tests (the plan; its refusals — a run with no room, and a plan applied to a state it was not made from; the owner table and the spans; and a sequence that copies twice as it diverges earlier each time, ending with the share gone) plus one allocator test on real regions that refuses a shared position through the store resolver, drives the copy, checks K/V byte-for-byte at the moved cells, checks the donor's four rows are unchanged, and drives a second copy-on-write through the fill entry point (fill_batch_inputs since #228; fill_attn_inputs when this gate landed). The real-model gate (a_store_inside_a_shared_prefix_takes_a_private_row, ignored like the others) shares 16 rows between two slots and then serves slot 1 a prompt matching only their first three tokens: the request is served (the resolver would otherwise refuse it), cows > 0 proves the copy ran, the answer is byte-identical to the same prompt on an engine with nothing to share, and the donor's rows — 17 positions × K/V × 24 layers — are byte-identical after it. That last check is the one with teeth: disabling the pre-pass and the resolver's guard makes the store write through, and the gate fails on exactly that comparison ("the diverging request wrote through the shared prefix"). Gate 3's kv_rm half (occupied() derived from the span lists) landed with S2. Suites: CPU 233 / CUDA 282, 0 failed.

C8b S5 landed (2026-09-22) — Metal refuses both window layouts, and C8b closes. Metal derives every query's window from positions (the pre-E1 form) and has no cell-store read path, so both explicit layouts — attn_span's one [lo, hi) pair per query and kv_map's (cell, len) runs — are refused. Two layers say so: Backend::supports_attn_span (the trait default, now an explicit override on Metal) keeps such a node off the backend at assignment time, and the Op::Attn arm of execute_node returns a loud Err if one ever arrives — the alternative would be computing a causal window from positions and attending to the wrong rows, which is the silent-wrong this project refuses. Nothing else changes for Metal: the models ask for a map only where Device::gathers_attn_map() says a kernel reads one (CPU and CUDA then — Metal's share path landed later, see the 2026-10-11 note below), so on a Mac the server keeps C8a's copy path — and since Metal's copy_cells is also refused until G5, admission logs the failed copy and prefills, which is the same loud fallback it has always taken. S5's gate is the compile check (Metal is cfg(target_os = "macos"), so CI's build-macos is the only compile dgxspark cannot do) plus this record; the device claims stay #44's, on a Mac.

That closes C8b: S1a/S1b (the span list and its two resolvers), S2 (sharing in place + the CPU gather), S3 (copy-on-write), S4 (the CUDA gather in every kernel) and S5 (Metal's refusal). The ticket C8 itself is closed on CPU and CUDA — a prefix is stored once and shared, and a store into it never writes through — with Metal's share path moving to G5, where the cell store is ported.

Updated 2026-10-11 — Metal's share path landed, and C8 is closed on all three. The paragraphs above are the 2026-09-22 state. Two later rounds removed S5's premise: G5 ported the Metal cell store, the write/move side and the one-range attn_span read (2026-10-06, #44 parts (a)/(b)), and #362 added the set-valued kv_map read (kernel_gqa_attn_map_f32/_f16/_q8_0 in src/metal/kernels/attn_window.metal). Device::gathers_attn_map() therefore answers true for Metal, admission shares a prefix in place there instead of copying, and the "CPU, CUDA" sets above and the "stays behind its gate (G5)" line in the C8 design table are history, not the current contract. docs/SUPPORT-MATRIX.md's op row and docs/METAL-BACKEND-DESIGN.md §4.4 carry the current state.

C8b S4 landed (2026-09-22) — every CUDA attention path gathers the map. attn_span's single range became three window modes selected by the size of the node's window input (C8b S2's departure 2): positions (causal), one [lo, hi) pair per query (attn_span), or KV_MAP_MAX_SPANS (cell, len) runs per query (kv_map). Each mode is a separate template instantiation — MAP joins CAUSAL as a compile-time flag for the same reason (E1b: the causal kernels must compile to exactly their pre-E1 instructions) — and a size that matches neither is refused rather than mis-strided. The row walk resolves a linear window index through the run list (kv_cell), and the key insight that keeps every kernel's mask valid is that a map window is a prefix of the sequence's address space (every run before the query's own, plus its own row), so the existing index < limit form stays exact with limit = the run total. Covered: the split-K 1-warp decode body (the sharing slot's decode step), the batched split path (1 < nt <= 16), the 4-warp hybrid, the legacy per-(token, head) kernels (f16 and f32 KV), and FA prefill, whose staging loop now resolves each linear index through the runs of the tile's widest query (its list is a superset of the others' — the one-sequence-per-tile precondition the span path already relied on for its tile-wide min/max).

The first A/B found a real cost, and the fix was to stop resolving per row. The map window measured 1.513x the span at the 7B decode shape (nkv = 2048, nh = 28, nk = 4, hd = 128: 24.6 vs 37.2 µs/launch) — the run walk ran per staged row. A map's runs are long, so a four-row batch almost always sits inside one: resolving the batch's first row once (plus how many rows its run still holds) and adding for the rest brought it to 24.6 µs — 1.001x the span, with a straddling batch still walking. The same measurement on the prefill path (nt = 512, hd = 128, f16 KV) is 1.726 → 1.872 ms (1.1x) now that FA gathers too; before it, FA had to be skipped for a map, which the A/B priced at 170x (301 ms of legacy per-token attention) — that number is why the FA port is part of this ticket rather than a follow-up.

Admission switches on one authority. Device::gathers_attn_map() (CPU, CUDA, Metal since #362) is read by both the models' graph builders (which ask for the kv_map input) and the server's admission (which shares a prefix in place instead of copying it), so the share and the window layout cannot disagree; MINFER_NO_KV_SHARE=1 forces the copy as the A/B gate for the share itself.

Gate 2 holds, bitwise. Always-run: cuda_map_window_matches_the_span_over_the_same_rows compares a map window against the span over the same bytes for two KV dtypes across one/two/three runs, at both prefill- and decode-shaped batches (a single token at the window's end — the case a one-token-at-position-0 draft of this test could not reach, and the real-model gate caught it missing) — 72 cases, all bitwise. On GB10, the three real-model gates (C8a's copy, E2's batched-vs-serial, and S3's copy-on-write) pass with sharing live, on an f32-KV 0.5B and an f16-KV, hd = 128 Qwen3-0.6B — the two combinations that matter, since FA prefill and the half-width cell stride only exist on the second.

Two pre-existing bugs this ticket's gate found, both fixed here. (1) The server's decode position was off by one: advance incremented current_pos before the forward that writes the row, so the first token after a prefill was stored at nt + 1 and position nt was never written — the next step's attention read a stale row whose content came from the arena's history, which made answers depend on allocation (the gate's two identical share runs disagreed). The increment now happens where the row exists (tick, next to the mirror's push), with the assertion corrected to match. (2) f16 KV cell moves strode by f32 elements: copy_cells's elems_per_cell is a count of f32, while an f16 cache stores a row as nkt halves (nkt / 2 f32), so every moved row walked twice as far and landed in the wrong cell — a compaction or a copy-on-write on an f16 device silently corrupted the arena. CudaBackend::copy_cells now halves the stride for f16, pinned by cuda_f16_kv_cell_move_strides_by_row_bytes (always-run, and it fails without the fix). Neither bug could be seen by the existing gates: both compare server-against- server, and every CUDA gate ran on an f32-KV model.

Suites: CPU 233 / CUDA 284, 0 failed, 9 ignored.

S5 is Metal behind G5 — lifted later by the G5 port and Metal's kv_map kernel (#362); see the 2026-10-11 note at the end of this record.

Risks. (1) The attention inner loop changes on every backend — correctness and timing, so each backend gets its own A/B. (2) The map must stay additive, or the CAUSAL path and Metal regress. (3) CoW's per-token private rows must not fragment the arena beyond what C7's growth can absorb; the stats gate watches exactly that.

StepContentGate
S1Per-sequence span list + the resolvers (kv_cells_for_seq, attn_span) reading through it — single-entry in practice, so nothing observable changesno behaviour change: suites stay bitwise
S2Block-granular refcounts + kv_seq_cp (share a prefix) + the CPU attention gathergate 2 on CPU
S3Copy-on-write: kv_private_row_for at store resolution — landed 2026-09-22: the share shrinks to the first position this forward writes and the run's own rows shift up inside it (two spans, no arena space needed), the store resolver refuses a shared position, and a run without room is a loud Errgate 3 on CPU + the real-model donor-byte check
S4CUDA attention gather + device A/B — landed 2026-09-22: three window modes (causal / span / map) selected by the window input's size, the MAP flag threaded through every attention kernel including FA prefill, and the batch-level row resolution that makes the gather freegate 2 on GB10 (bitwise, f32 and f16 KV) + the decode/prefill A/B
S5Metal behind G5 — landed 2026-09-22: both explicit window layouts refused (supports_attn_span keeps them off the backend; the Op::Attn arm backstops with a loud Err), admission keeps C8a's copy, docs closed. Superseded 2026-09-22 → 2026-10-06/#362: the G5 port and the Metal kv_map kernel lifted it — gathers_attn_map is true for Metal and admission shares in place (see the record's 2026-10-11 note)compile check (CI build-macos) + G5 record

6. Phase D — IR expressiveness (item 7)

IDTitleEffort
D1Strided views with allocator-known aliasing + multi-output nodes — DONE (2026-09-19), increments 1–3: exact views are zero-copy and offset/partial windows work on CPU and CUDA (BufRef carries offset+len through the Backend trait; Metal is exact-only), and the "multi-output node" is GraphBuilder::split_parts — one owning node plus one Op::View per part, so the graph stays single-output and no backend changed. The design's sketched Op::SplitParts was not added: a split op over one input is just views of that input. MoE/MLA remain model work (producer ops, weights, routing)L
D2Re-express one decode fusion as a composition (proof) — DONE (2026-09-19): FusedFFN as concat MatMul + two partial windows + in-place SwiGLU; CUDA takes the composition, Metal keeps the node until G5; env gate MINFER_FFN_NODE=1 for the A/B; three models byte-identical, decode cost ≤0.8% (recorded)M
D3Decide the fate of the four hand-written fused ops — DONE (2026-09-19): FusedFFN keeps the default (the proven composition is 0.3–0.8% slower on CUDA and Metal needs the node until G5); the QKV family keeps (no composition proof); QkvBiasRopeStore is a fallback epilogue, not a redundancy. Policy rule + gate clarity recordedS
  • D1 acceptance: Op::View becomes zero-copy (a test asserts the allocator maps a view onto its parent's buffer); the allocator's liveness understands view_src; existing graphs unchanged and bitwise.

D1 design (written before the code, 2026-09-19)

What D1 is actually for. C3's own text allows either route ("a copy needs either a view or an explicit copy op"), so views are not on C3's critical path; what needs them is D2 (a decode fusion re-expressed as a composition: FusedQKV leaves q|k|v in one concat buffer, and the composition feeds attention from three windows of it), and what needs multi-output nodes is MoE/MLA (item 17). The diagram's "C3 needs D1" is therefore softer than it looks; the honest dependency is D2 → D1.

Today (measured in the code, not assumed). Op::View/Reshape/Permute are copies: the CPU arm is out.copy_from_slice(ins[0]), i.e. the IR has no aliasing at all. The one alias that exists is the in-place elementwise rule in GraphAllocator::alloc_graph (Silu/RoPE): when the input's sole consumer is the op and both are on the same backend, the op's output is the input's buffer (node_to_buf[op] = node_to_buf[src]) and the input's last_use is extended past it. BufRef is { backend, id } — no offset — and the Backend trait hands the backends plain ids (in_bufs: &[usize], out_buf: usize). An offset view therefore cannot be expressed without threading offsets through the trait, all three backends and their call sites (the A5 lesson: a trait-signature change silently skips the CUDA test call sites unless CI compiles them).

Three increments.

  1. Exact views — landed here. CNode gains view: Option<ViewAlias> where ViewAlias { src: NodeId, offset: usize } (the field carries offset from the start so the IR shape does not change again later), the builder marks View/Reshape/Permute as views of their single source, the allocator maps node_to_buf[view] = node_to_buf[parent] and extends the parent's liveness through the view_src chain, and the view kernels become no-ops (the alias is the output). The backend trait is untouched: an exact view is the same buffer id. Loud refusals: offset != 0 in this increment, a missing parent, a parent on another backend, or a window that would not fit inside the parent.
  2. Offset views. BufRef gains offset: usize, the trait passes BufRef, and each backend adds the offset where it resolves an id (CPU's split_at_mut plumbing, CUDA/Metal's ptr_of). This is what D2 needs.
  3. Multi-output nodes (CNode grows a list of outputs), which is what MoE and MLA need.

Acceptance for increment 1 (the D1 acceptance above, narrowed to what is landed): a view is zero-copy (a unit test asserts the allocator maps it onto the parent's buffer id), liveness understands view_src (the parent is not recycled while the view is live, and is freed after), and existing graphs are unchanged — the real-model bitwise tests and the op matrix are the evidence.

Increment 2 record (2026-09-19). BufRef now carries offset and len (the window, not just a buffer id), the Backend trait passes references instead of raw ids, and the allocator maps a view to parent.window(offset, len) after checking the window fits inside the parent:

  • Every backend applies the offset where it resolves a buffer: CPU slices each input/output window out of the pool (the split_at_mut plumbing is otherwise unchanged), CUDA adds offset * 4 bytes in a new ptr_of_ref and takes element counts from the window (in_bufs[k].len, which also corrected copy_d2d's size check), and the host paths are window-aware — copy_to_cpu returns exactly the window.
  • Metal is exact views only, and that is a capability boundary rather than an oversight: its kernels take a buffer and a length with no element offset, so a window would silently read the wrong bytes. supports_op(Op::View{offset}) is offset == 0 there, the allocator backstops the partial-window case (which supports_op cannot see, since the parent's length is not in the op), and SUPPORT-MATRIX.md gains the asymmetric row; G5 is where Metal would learn offsets. That path is compile-verified by CI's macOS job alone, and the job earned its keep on this change: it caught a BufRef passed to a Metal debug Display format, which no Linux build can see.
  • Evidence: a new op-matrix case, View offset (a partial window at offset 2 over an 8-element parent, expected [3,4,5,6]), passes on CPU and CUDA and is skipped on Metal; two allocator tests pin the mapping ((id, offset, len) == (parent, 2, 4)) and the new refusal boundary (a window past the parent); the device-gated CUDA test call sites were re-pointed through a test-only exec_ids shim that derives each reference's length from the pool (the A5 lesson: a trait change skips them unless the test harness is compiled); and the real-model bitwise gates are unchanged.

Increment 1 record (2026-09-19). Landed as designed:

  • CNode.view: Option<ViewAlias> (with offset present from the start), set automatically for View/Reshape/Permute in GraphBuilder::node — the one construction point, so hand-built graphs and the op matrix get it too.
  • GraphAllocator::alloc_graph maps such a node onto its parent's buffer and extends the parent's liveness through the view_src chain (extend_through_views), which is also now used by the in-place branch (an in-place op on a view writes through to the buffer it windows).
  • The three view kernels became no-ops on all backends, with the aliasing asserted instead of assumed: CPU debug_assert_eq! on the pointers, CUDA/Metal return Err rather than copying if the output is not the source's buffer (standing rule 2 — the previous arms performed a silent identity copy).
  • Refusals, all loud: offset != 0 (increment 2), a partial window (its element count differs from the parent's, which would make copy_to_cpu read the whole parent), a cross-backend view, a missing parent buffer.
  • Evidence: three new tests in graph::alloc::tests (a_view_aliases_its_parent_buffer_and_keeps_it_alive, views_that_need_more_than_exact_aliasing_are_refused, in_place_on_a_view_extends_the_parents_liveness); the op matrix's View/Reshape/Permute cells now exercise the alias on CPU and CUDA; and the real-model bitwise gates (B2, C2, prefix reuse) are unchanged: CPU 195 passed / 0 failed, --features cuda 241 / 0. No architecture emits these ops yet, so the payoff is D2's: it is the machinery a FusedQKV-as-composition needs, one increment short of the offsets it will actually use.

Risks. Aliasing changes who may write whose bytes: a view's consumer can now write into the parent's buffer (in-place ops on a view), so the existing "sole consumer + same backend" rule and the transitive liveness extension are both load-bearing, and chains (view of a view, in-place on a view) need their own tests. A backend that cannot express an alias must refuse loudly rather than copy silently (standing rule 2). And copy_to_cpu/fill_input on a view now read the parent's bytes — which is the point, but it is also why the op matrix's View/Reshape/Permute cells are the regression net for it.

  • D2 acceptance: FusedFFN re-expressed as MatMul + views + in-place SwiGLU; the hand-written node stays behind an env gate for A/B; bitwise identity; no decode regression beyond a recorded budget.

D1 increment 3 design (written before the code, 2026-09-19) — multi-part outputs via views

The ticket row above says "multi-output nodes". Before giving CNode a second output — which would touch the allocator, the scheduler, all three backends and the JSON/DOT/trace exporters, for a capability only MoE would exercise — consider what D1 increments 1–2 already bought: an output can be an offset view of another buffer (BufRef { offset, len }, CNode::view) and liveness follows views (extend_through_views). A node that produces several tensors therefore needs one output buffer, not several outputs: its parts are views over it, and consumers bind to those views exactly as they bind to any other producer's output.

Landing correction (2026-09-19, before the code): no new op is needed, and none was added. The sketch above proposed Op::SplitParts, but a split op taking one input is either an identity copy or — what it really is — a set of windows over its input, which Op::View already expresses. What the increment needs is the helper: GraphBuilder::split_parts(owner, sizes) -> Vec<NodeId> emits one Op::View per part (cumulative offsets, the owner's trailing dims, the owner's dtype), so:

  • a node whose one kernel writes several logical tensors keeps writing one output buffer and the parts are ordinary graph tensors — the graph stays single-output, no backend grows a second-output path, and the parts inherit the owner's liveness through the existing extend_through_views;
  • a (values, indices) pair needs no second output dtype: an indices part is i32 in f32 bit patterns (rule 4), which is what lets it drive Op::GetRows;
  • and the shape is already in production — D2's fused_ffn_composition is a concat matmul whose gate/up halves are consumed through two such windows.

The property that makes this the right shape: the graph stays single-output, so no allocator, scheduler or backend grows a second-output path that only MoE will ever use, and the parts are ordinary graph tensors (so MINFER_TRACE, the JSON export and DOT already show them).

D1 increment 3 record (2026-09-19)

Landed as GraphBuilder::split_parts (builder helper: no new Op, no kernel, no backend change — Op::View's arm is a debug assertion because the allocator maps a part onto the owner's buffer). Two tests, both CPU-runnable:

  • split_parts_maps_every_part_onto_the_owners_buffer — structural: a real producer (a concat matmul) split into [3, 3]; each part's BufRef is (owner.id, offset, len) with cumulative offsets, and the record is a view (CNode::view), which is what extends the owner's liveness.
  • split_parts_parts_feed_independent_consumers — functional, in the shape MoE routing needs: one owner buffer holding four value rows plus two indices, split into [8, 2], with the indices part driving Op::GetRows over the values part. The gathered rows are asserted against the Rust-side expectation, so the mechanism is proven end to end (one owner, two parts, two consumers, no copy).

What this closes and what it does not: the IR blocker is gone — a producer can feed several consumers with independently bindable tensors without a second output — and it is closed without touching the allocator, the scheduler or any backend. MoE and MLA still need their own work (expert weights, routing, the grouped GEMM; latent projections, per-head latents), which is model work rather than IR work, and is recorded as such instead of implied to be unblocked.

What increment 3 does not do — correcting "D1 unblocks MoE/MLA" a second time, after D1's own design already scaled it back: MoE still needs its architecture work (expert weights, routing, the grouped GEMM) and MLA needs its latent projections. Increment 3 removes the IR blocker and proves the mechanism with a kernel any model can use; it is not a MoE implementation.

D2 design (written before the code, 2026-09-19)

What FusedFFN is today (read from the code, not remembered). The model builders emit it only for a GPU decode step: nt == 1 && cparams.gpu && cparams.fuse_ffn && gu_concat_available(...) && nf <= 16384 (models/qwen2/graph.rs:243, qwen3 likewise). Its output shape is [2 * nf, nt] — a concatenated gate|up buffer — and its consumer (the down matmul) reads rows 0..nf, where the fused kernel leaves silu(gate) * up. CPU never sees the op: cpu_backend returns Err for Op::FusedFFN ("fusion not enabled for it"), and the GPU gate is what keeps it off the CPU path.

The composition, and why it is bitwise-comparable. With D1 increment 2 in place, the same graph is expressible as:

MatMul(x, gu_concat_weight) -> concat [2*nf, nt]
View { offset: 0,  shape: [nf, nt] } -> gate      (a partial window of concat)
View { offset: nf, shape: [nf, nt] } -> up        (the second window)
SwiGLU(gate, up) -> out                            (in place, into the gate window)

out occupies the same bytes the fused node leaves its result in (rows 0..nf of the concat), so the down matmul and everything downstream are unchanged and the two paths are comparable element by element. That is the whole point of doing D2 after D1: before offset views, the gate/up windows could only be copies.

What must change.

  1. A builder path that emits the composition (fused_ffn_composition), with the hand-written Op::FusedFFN kept behind an env gate for the A/B — the plan's wording, and the reason D3 exists: the node is not deleted until the composition is measured.
  2. The allocator's in-place rule must accept Op::SwiGLU writing into its first input's window. The two inputs are windows of the same pool buffer at different offsets, which the CPU's aliasing snapshot already handles (it clones each aliased input); the rule and a test for the two-input case are what is missing.
  3. A per-backend capability decision, and it is not uniform:
    • CPU never emits FFN fusion (cparams.gpu is required), so nothing changes there;
    • CUDA takes the composition (MatMul and SwiGLU are both long-supported ops, views landed in D1 increment 2);
    • Metal keeps the hand-written node: the composition's views are partial windows at a non-zero offset, which Metal's exact-view rule refuses until G5 (SUPPORT-MATRIX.md's D1 row). Falling back must be a build-time choice per backend, not a runtime copy.

Evidence plan (the increment this round did not reach). Bitwise A/B on the device with a real model: one binary, the env gate selecting node vs composition, comparing the token stream and the logits; a decode tokens/s measurement for the recorded budget (the composition adds one 2*nf f32 read+write per layer per token for the SwiGLU pass, and it may pick a different GEMM kernel than the fused special case — that is exactly what the budget is for); the real-model bitwise gates and the op matrix unchanged; cargo test --release --features cuda green.

Increment record (2026-09-19) — landed. GraphBuilder::fused_ffn_composition builds the composition (matmul_by_name over the concatenated gate|up weight → View{offset: 0} gate → View{offset: nf} up → Op::SwiGLU), and:

  • the allocator's in-place set accepts Op::SwiGLU only when its first input is a view — the fusion pass's own SwiGLU (CPU, non-view) keeps its separate buffer, so no existing graph moves;
  • the models pick per backend: CUDA takes the composition, Metal keeps the hand-written node (partial/offset windows are refused there until G5 — a build-time choice, never a runtime copy), and MINFER_FFN_NODE=1 forces the node anywhere, which is the A/B gate;
  • a CI-verifiable test (ffn_composition_swiglu_aliases_the_gate_window) pins the structure and, with no device, that the SwiGLU's buffer is the concat buffer at offset 0 — the same bytes the fused node writes.

Evidence (device, GB10, greedy, equal work). Same-mode repeat is the determinism control (identical), and the comparison strips the timing lines:

ModelHand-written nodeCompositionGenerated text
Qwen2.5-0.5B Q4_0, n=64516.3 tok/s513.6 tok/sbyte-identical
Qwen3-0.6B Q8_0, n=32228.9 tok/s228.2 tok/sbyte-identical
Qwen2.5-7B Q4_K_M, n=64199.7 tok/s198.2 tok/sbyte-identical — but vacuous, see below

Correction (2026-09-19, D3 round). The 7B row is vacuous: the fusion gate is nf <= 16384 and the 7B's nf is 18944, so that model never uses FusedFFN and MINFER_FFN_NODE=1 selected the same unfused path as the default — the two runs were identical by construction, not by evidence. The meaningful rows are the 0.5B (nf 4864) and Qwen3-0.6B, both of which do fuse. The 7B's 0.3%-faster reading is noise, and the recorded budget therefore rests on the 0.5B/Qwen3 numbers (0.3–0.8% slower). Anything measured on a model above the gate says nothing about D2.

So the acceptance is met: re-expressed as a composition, identical greedy output (the project's established proxy for bitwise), and a recorded budget of <=0.8% decode (measured 0.3–0.75% slower, consistently — the separate SwiGLU pass costs one 2*nf f32 read+write per layer per token, and it is not offset by the fused epilogue on this device). Suites: CPU 197 / 0, --features cuda 243 / 0.

Method note, because it cost a false alarm. The first A/B hashed stdout, which includes the Prefill:/Generated: timing lines that differ every run, and therefore reported DIFFER on all three models. The same-mode repeat caught it: if a mode does not reproduce itself, the comparison is measuring the harness. Any future A/B of this kind must strip those lines (or compare token ids) and keep the repeat control.

What D3 inherits. The composition is proven and slightly slower on CUDA, which is exactly the trade D3 decides: keep the hand-written node (and its device-specific kernels) or delete it for simplicity. It cannot be deleted outright regardless — Metal needs it until G5.

Risks. The budget is the honest risk: if the separate SwiGLU pass costs more than the fused epilogue saves on this device, D2's outcome is "re-expressed, measured slower, node stays" — which is a result, not a failure, and is why D3 is a decision ticket rather than a deletion. The second risk is the reverse of D1's: a composition that works everywhere invites deleting the hand-written node, and the CUDA-specific kernels (fused epilogue, MMQ tiling) exist for reasons the composition cannot express; D3 owns that call.

  • D3: with D2 proven, decide per fusion whether to keep the hand-written node (performance) or delete it (simplicity). MoE (item 17) is unblocked here.

D3 record (2026-09-19) — the fate of the four hand-written fusions

Decisions.

OpDecisionEvidence / why
FusedFFNkeep as the default; the composition stays behind MINFER_FFN_COMPOSITION=1D2 measured the composition 0.3–0.8% slower on CUDA on the models that actually fuse (0.5B nf 4864, Qwen3-0.6B; the 7B never fuses — nf 18944 > the 16384 gate — and its row is recorded as vacuous), and Metal requires the node until G5, since a partial window at a non-zero offset is refused there. Deleting it would cost performance on CUDA and break Metal.
FusedQKVkeep — no deletion decidedBoth GPU backends support it and it carries the bias+rope+store epilogue; no composition proof exists. A D2-style proof is expressible now (concat MatMul → three windows → in-place rope → store via Op::KvcacheStore) but unmeasured, so deleting it would be an unmeasured simplification that also has to hold on Metal.
FusedQkvNormkeep, same as FusedQKVAdds per-head Q/K RMSNorm inside the epilogue (Qwen3); a composition would need a norm on a window — expressible, still unproven.
QkvBiasRopeStorekeep — it is not a redundant fusionIt is the fallback epilogue for the mixed-quant case: three separate q/k/v matmuls without bias, plus one combined pass (bias×3 + rope×2 + store×2). Deleting it would slow that path, not simplify anything.

The policy rule (reusable). A hand-written fusion may be deleted only when all four hold: (1) a composition exists in the builder; (2) it is byte-identical on every backend that emits the fusion, not just the one measured; (3) its measured cost is within a recorded budget on those backends; (4) no backend needs the node for lack of a composition primitive. Otherwise the node stays and its A/B gate is documented. FusedFFN fails (3) on CUDA and (4) on Metal → kept. The QKV family fails (1) → kept.

Gate clarity (the follow-through). Two classes of environment gate exist and mean different things:

  • MINFER_NO_FUSE_QKV=1 / MINFER_NO_FUSE_FFN=1 — disable the fusion entirely (build the plain MatMul + rope/store, or gate+up+silu+mul, path). Unchanged.
  • MINFER_FFN_COMPOSITION=1 — keep fusing, but build the fusion as the composition (MatMul + gate/up windows + in-place SwiGLU) instead of the hand-written node. This is the D2 proof path and the A/B reference; it is refused with a warning on a backend without offset views (Metal until G5; CPU, which never fuses in the first place).

The choice is a pure function (models::ffn_composition) unit-tested across the matrix, so CI covers it without a GPU — the same shape as E6's batch_mode.

MoE prerequisite, honestly. Phase D's note says MoE is unblocked here; it is not. MoE and MLA need multi-output nodes — D1's third increment, which is not landed (D1 has delivered exact views and offset/partial windows). Nothing in D3 changes that.

Evidence run. With the new default, the node and the composition still produce byte-identical greedy text on the models that fuse: Qwen2.5-0.5B Q4_K_M (n=32) and Qwen3-0.6B Q8_0 (n=32), both IDENTICAL. Suites: CPU 198 passed / 0 failed / 5 ignored, --features cuda 244 / 0 / 5.

7. Phase E — batching, then memory policy

IDItemTitleEffort
E12IR seq_id + explicit attention masks (CPU) — DONEL
E1b2CUDA attention kernels read attn_span — DONE, device-verified (2026-09-18) (window test passes on GB10; causal-path timing unchanged)M
E23Batch composition + continuous batching — mechanism landed; CPU acceptance refuted and accepted, GPU acceptance MET (1.9x); opt-in at the time (MINFER_BATCH=1 — E6 later made the default device-aware); A7 closed by deleting n_seqs — ticket closedXL
E310Chunked prefill: make n_batch real · #45 — DONE (2026-09-22, see the record)M
E48Allocator reserve/assign split + size classes + memory accounting · #55 — DONE (2026-09-23): S1 the accounting + the size-class ladder + the feasibility gate; S2 the pools allocate at the class, the length contract moved to BufRef, two latent liveness bugs fixed; S3 split reservation (the slot table) from assignment — a rebuild re-maps without touching the pool, CUDA's pool_gen stops moving, and GraphCache holds one graph per GraphParams (a switch re-maps instead of rebuilding)L
E59Layer-offload budget (n_gpu_layers equivalent) · #46 — DONE (2026-09-23): S1 the layer-granular plan (--gpu-layers/MINFER_GPU_LAYERS), per-block weight registration and placement, the startup report, a verified mixed CPU+CUDA run; S2 the auto fit — per-block weight bytes from the GGUF index, the pure fit_blocks prefix search against a weight budget (MINFER_GPU_MEM, else three quarters of device free), a quarter held back for KV/activationsL
E63 (follow-up)Device-aware batching default — DONE (2026-09-19): MINFER_BATCH unset now batches iff the model's forwards run on CUDA, serial otherwise; =1/=0 force it either way; the decision is a pure unit-tested function. Refetched 1.97x on the 7B with no environment variable (see the record)S

E5 record, S2 (2026-09-23) — the offload plan as a fit, not a ceiling

Why. S1 shipped the placement machinery and a knob, but the knob was an explicit ceiling: the user had to know the model's block count and guess how many fit. The ticket's "budget knob" is the other half — compute the split from the device's memory, which is what a machine whose free memory is smaller than the model actually needs.

What S2 landed.

  • auto as a third request (MINFER_GPU_LAYERS=auto / --gpu-layers auto), parsed strictly with the rest (OffloadRequest::parse: a decimal count, auto, or unset; anything else is a refused load).
  • The byte table comes from the index. block_weight_bytes sums GgufTensorInfo::nbytes per block (block_of on the tensor name) before anything is loaded — the decision cannot wait for a measurement, because the registration filter is the plan. nbytes is now the shared arithmetic (the loader slices the part with the same method), so the fit and the loader cannot disagree about a tensor's size.
  • The fit is a pure prefix search (fit_blocks(budget, per_block, reserve)): the largest k with reserve + Σ per_block[0..k] ≤ budget. A prefix, not a knapsack — the plan is 0..gpu_layers, and a gap would put a CPU block between two device blocks for nothing; the walk stops at the first block that does not fit, which is a deliberate, documented conservatism.
  • The budget (weight_budget): MINFER_GPU_MEM=<MiB> when set, else three quarters of what the device reports free — the same default E4's feasibility gate uses, so the fit and the gate that later checks the activation pool are talking about one number. A quarter of the weight budget is then held back as reserve for the KV arenas and the activation pool: both are sized per graph at forward time, so no load-time fit can measure them; a prompt that needs more is refused by the E4 activation gate (loudly) instead of quietly swapping.
  • The report explains the fit (auto_source): offload: 5 of 24 blocks on cuda … (40.0 MiB of device weights; auto: 5 of 24 blocks fit — weights budget 64 MiB, 16 MiB reserved for KV/activations; MINFER_GPU_MEM=64 MiB) — the decision and the numbers behind it.
  • device_memory() (src/models/mod.rs, from CudaState::device_memory) is the one place that asks the device (CUDA: cudaMemGetInfo), and since #122 it returns an explicit allocplan::DeviceMemory (Reported / QueryFailed / NoDevice) rather than an Option<usize> where a failure and "no device" both looked like a small number. A failed query refuses the auto request with the real CUDA error name instead of fitting 0 blocks; Metal's wrapper reports no free-bytes number yet, so on macOS auto needs MINFER_GPU_MEM and otherwise fits nothing (the default and an explicit count still work there).

Acceptance, as measured (0.5B q4_0 on GB10):

  • an_auto_offload_plan_fits_the_budget (ignored, real model): with MINFER_GPU_MEM=64 the fit is a strict prefix — 5 of 24 blocks, 40.0 MiB of device weights registered, i.e. inside the 48 MiB the budget left for weights — the report names auto and the cap, and the model's four greedy steps match the all-CPU run; with no cap the same request selects all 24 blocks, so an auto default cannot silently under-offload a device that fits the model.
  • The pure matrix (4 tests, no device): the request spelling (auto case-insensitive, garbage refused with the alternatives named, plan() refusing auto so a caller cannot skip the fit), the prefix search (exact boundary, reserve eating the budget, empty/zero-size tables, the "a big block stops the walk even though smaller ones follow" conservatism), the budget (three quarters, explicit cap wins, no device → nothing fits, garbage refused) and the report text (device vs cap vs no budget).
  • Mutation-checked: making fit_blocks ignore the budget fails the auto gate on its first assertion (a 64 MiB budget must be a strict prefix, got gpu_layers: 24), and S1's gate covers the placement rule on the same code path (removing it fails the mixed run; forcing allows_weight true fails the device-bytes assertion).
  • Suites: CPU 280 passed / 0 failed / 15 ignored; CUDA (GB10, serial) 333 / 0 / 16; the #[ignore]d set serially 15 passed / 0 failed.

Honest scope. The fit measures raw tensor bytes; the auxiliary device copies the loader builds while loading (the fused attn_qkv/ffn_gu concats, the padded Q6_K layout, the q8_0 p32 split, the q4_K dsc pair) are not in the per-block table — the quarter held back and E4's loud activation gate are what keep the estimate honest, and an under-estimate ends as a refusal, never as a silent overcommit. The reserve is a fixed quarter rather than a function of n_ctx/n_batch (which are runtime choices): a long-context request on a tight budget may therefore be refused by the gate even though auto accepted the weights — the message names the numbers, and --gpu-layers/MINFER_GPU_MEM are the knobs. Metal has no free-bytes query yet, and the fit is per allocator, like every other budget.

E4 record, S3 (2026-09-23) — split reservation from assignment, then hold several graphs

Why. S1 and S2 left the ticket's two structural items. The allocator's walk freed the previous graph's buffers and re-allocated them, so even a same-shape rebuild went through the backend's pool — and CUDA's pool_gen, which moves on every allocate/free, is exactly what invalidates its captured graphs (graph_replay_step re-captures when the generation changed). Separately, GraphCache held one graph, so a server alternating a 1-wide and an N-wide decode step rebuilt on every step, and a chunked prefill paid E3's measured "one forward's fixed overhead per chunk" again on every repeat request.

What S3 landed.

  • Reserve, then assign. A released classed buffer no longer goes back to the backend: it goes to a reservation table (slots: HashMap<(Backend, class-elements), BTreeSet<pool-id>>), and alloc_class_in_pool takes the smallest idle id of the class, allocating a new buffer only when the class has nothing idle. buf_class marks which allocations are classed; staging (exact-sized) keeps the backend's own free list. A rebuild therefore re-maps: the pool is not touched, and the same topology gets the same slots. The smallest-id-first order is load-bearing and deterministic — liveness releases buffers in HashMap order, so the first (LIFO) version handed a rebuilt graph different ids each time, which the gate caught immediately (left: [(0, 2), (1, 1), (2, 0)] / right: [(0, 0), (1, 1), (2, 2)]).
  • GraphCache holds one graph per GraphParams (MRU first, MAX_CACHED_GRAPHS = 8). try_reuse matches any cached graph, makes it current and re-maps the allocator onto it (alloc_graph: liveness + slot assignment) — no build, no assign pass, no fusion pass. stats() reports (builds, reuses) for the gate; cached_graphs() is the depth.
  • Cross-backend staging is per graph: the map is keyed by (graph uid, node, backend) and survives a re-map. Node ids restart per graph, so the old (node, backend) key would have let a switch reuse another graph's buffer; and re-creating the entries per switch leaked, because staging is alloc_fresh (never recycled) — the CUDA suite found it as CUDA: OOM allocating 4 bytes after ~200 generate steps in the capture-parity gate. Entries of evicted graphs stay allocated (a handful of boundary-sized buffers, bounded by the cache's depth).
  • The reservation is visible: MemoryReport::{idle_slots, reserved_classes}, and CpuBackend::alloc_count is the CPU twin of CUDA's pool_gen.

Acceptance, as measured.

  • a_rebuild_remaps_instead_of_reallocating (CPU): the same graph and a neighbour in the same class (896x16 / 896x17, both class 16384) re-map — n_cpu_allocs and n_cpu_buffers unchanged, the node → slot mapping identical; a shape that leaves the class does reserve.
  • a_released_buffer_stays_reserved_and_idle: pool_bytes, live_bytes and idle_slots are exactly what they were after a rebuild, and the reservation reports at least one class.
  • a_rebuild_does_not_touch_the_device_pool (CUDA, GB10): pool_gen is unchanged across a same-shape rebuild and across a same-class different shape, and the mapping is identical — i.e. a captured graph survives a rebuild.
  • switching_between_cached_graphs_re_maps_instead_of_rebuilding (CPU): two shapes built, four switches, stats() == (2, 4), zero pool allocations during the switches, and each graph maps to the same slots every time it is switched in.
  • a_repeated_chunked_prefill_stops_rebuilding (real 0.5B, ignored): request 1 → 3 builds / 5 reuses; request 2 (same prompt, same chunk size) → 3 / 9, i.e. the second request builds nothing. Before S3 each chunk forward of the second request was a fresh build.
  • Mutation-checked: returning classed buffers to the backend instead of the reservation fails the re-map gate; making try_reuse match only the previous graph fails the switch gate.
  • Suites: CPU 276 passed / 0 failed / 14 ignored; CUDA (GB10, serial) 329 / 0 / 15; the #[ignore]d set serially 14 passed / 0 failed.

Honest scope. The reservation is per allocator (per GraphCache), not per process: two caches (two models, or the spec-decode pair) each hold their own slots, exactly as they always held their own pools. Staging entries of an evicted graph are not reclaimed (cross is a HashMap the allocator does not prune against the cache); the cost is a few boundary-sized buffers per evicted graph, and a prune keyed on the live uids is the obvious follow-up if the cache ever grows past a handful of graphs. Metal shares the code path but was compile-checked only (build-macos).

E5 record, S1 (2026-09-23) — a layer-granular offload plan

Why. Device participation was one all-or-nothing check over the whole model (Qwen2Model::device asked "is every weight registered on the GPU?"), so a model that does not fit in device memory could not run at all — the roadmap's own example being "7B Q8_0 will not fit (7.2 GB weights alone)". Two things were missing: granularity (nothing in the IR said which block a node belonged to, so nothing could decide "these blocks here, the rest there") and a knob to choose the split.

What S1 landed.

  • src/graph/offload.rs — OffloadPlan { gpu_layers, n_layers } and the pure resolver OffloadRequest::plan(env, n_layers, device_available) (the E6 batch_mode pattern: CI covers the matrix with no GPU). --gpu-layers N (CLI) and MINFER_GPU_LAYERS=N (environment) spell it; unset means "every block a device can hold" — the pre-E5 behaviour — and anything that is not a block count fails the load loudly. Blocks 0..gpu_layers run on the device; the tensors outside any block (the embedding, the final norm, lm_head) follow the device only when every block is offloaded, so a partial plan never has to fit the two largest tensors (llama.cpp's n_gpu_layers > n_layer convention, written down once).
  • The plan is one number read in three places, which must agree:
    1. the loader registers a tensor on the device only when OffloadPlan::allows_weight(name) says so — the block comes from the registry name (block_of, {ns}blk.{i}.…), including the fused blk.{i}.attn_qkv / blk.{i}.ffn_gu concat copies. Registering a non-offloaded block would spend exactly the device memory the plan exists to save;
    2. the builder stamps CNode.layer (GraphBuilder::set_layer, called once per block by both model builders) and gates the device-only fused forms on layer_gpu = gpu && il < gpu_layers;
    3. the assignment pass (BackendScheduler::assign_backends → GraphAllocator::supports_for(op, dtype, layer)) never offers the device for a block past the plan, and CParams.gpu_layers carries the plan into the reuse identity — the assignment is topology, so a different plan must rebuild.
  • Verification and reporting. After registering, the load asks device() whether the offloaded blocks can actually run there; if not, the plan drops to CPU-only with a printed reason (never a silent partial offload). offload_report() prints the startup line E5 asks for, with the device memory the offloaded weights measured.
  • A mixed plan refuses a KV session: a session file is one arena with one backend tag (C5), so kv_load refuses before reading the file. kv_save already refused mixed layers.

Acceptance, as measured.

  • A chosen split runs (a_partial_offload_runs_the_rest_on_the_cpu, 0.5B q4_0 on GB10): --gpu-layers 4 loads with 4 blocks on CUDA and 20 on the CPU, the startup line reads offload: 4 of 24 blocks on cuda, 20 on cpu; embed/output on cpu (32.0 MiB of device weights; --gpu-layers 4), and four greedy steps match the all-CPU run (CPU and device logits differ by design, rule 9, so tokens are the honest comparison). The device holds only the offloaded blocks: device_bytes >= Σ(blocks 0..4) and < Σ(all blocks).
  • The boundaries copy (the same gate): the built decode graph has 40 nodes on CUDA and 368 on the CPU, cut into 7 splits (3 device, 4 CPU), and both a device→CPU and a CPU→device boundary carry non-empty inputs — the scheduler's cross-backend copies are what make the mixed graph executable at all. (The split count is not one per block: a block's device nodes are contiguous with its neighbours' whenever the nodes between them are on the device too, so the claim asserted is the alternation plus the copies, not a count.)
  • Placement is a gate, not a hint (the_offload_plan_keeps_late_blocks_off_the_device, CUDA): with gpu_layers = 2/4, supports_for(...) answers Cuda for block 0, CPU for block 2, CPU for a node outside any block, and Cuda for an unblocked node only under a full plan.
  • The unit matrix (5 tests, no device): unset/empty/garbage spellings, clamping, the CLI form beating the environment, the "unblocked tensors need a full plan" rule, the weight-name filter (blk.3 yes, draft.blk.3.attn_qkv yes, blk.4 no, blk.0attn/token_embd no), and the report text.
  • The builder contract (set_layer_tags_the_nodes_created_after_it): nodes created after set_layer(Some(2)) carry Some(2), views inherit it, and the tag clears.
  • A mixed plan cannot resume a session (the refusal precedes the file read — the test's path does not exist).
  • Suites: CPU 273 passed / 0 failed / 13 ignored; CUDA (GB10, serial) 325 passed / 0 failed / 14 ignored; the #[ignore]d set serially 13 passed / 0 failed on the CPU build (the mixed-run gate skips there — no device — and passes on the device, above).
  • CLI, end to end (0.5B on GB10): --gpu-layers 4 prints the line above and generates; --gpu-layers 0 prints offload: cpu only — 0/24 blocks on the device (--gpu-layers 0); MINFER_GPU_LAYERS=6 prints the environment as the source; MINFER_GPU_LAYERS=banana fails the load with the reason instead of being ignored.
  • Mutation-checked: with the placement gate removed from supports_for, the mixed run fails (a non-offloaded block reaches the device and its weights are not there); with allows_weight forced true, the device-bytes assertion fails (the device would hold the whole model, which is what the plan exists to prevent).
  • Honest scope: dgxspark's device has ~128 GB, so "a model larger than device memory" cannot be staged here. The gate forces the split with the knob and asserts the device holds only the offloaded blocks — the knob is exactly how the constraint is expressed; the automatic fit is S2 below. Metal takes the same code path but is compile-checked only (build-macos).

What S1 does not do (it stays on #46): the automatic fit — consuming E4's memory accounting to put "as many blocks as the budget allows" on the device, which is what turns the knob into a policy (and what a device whose free memory is smaller than the model needs); per-block tensor_split-style tuning; and a per-layer map for the server's batching default (the server sees "the device participates", which is right for a mixed plan but does not distinguish it).

E4 record, S2 (2026-09-23) — the pools allocate at the class, and what that exposed

Why. S1 ended with the ladder in the plan and the accounting but exact pools: two shapes in one class still got a buffer each, so a rebuild with a slightly different nt grew the pool, and the ticket's "recycled buffers come from the same size classes (no silent growth)" was not claimed. The obstacle was that a pooled buffer is no longer the same thing as the node it serves: the pool buffer is class_size(n), while the node's real length is n, and every consumer that used a physical length had to be told.

What S2 landed.

  • The pool request is the class — alloc_class_in_pool asks for class_size(size). Because the free list matches lengths exactly and every buffer of a class has the same length, the second shape in a class now finds the first one's buffer. The persistent KV regions still come through alloc_exact_in_pool: a cell's width is a layout contract (row_elems), not a tuning knob.
  • The length contract lives in BufRef (new Backend::write_host_window(id, offset, data)): a fill is checked against the node's logical BufRef::len and written at the reference's offset, and write_host keeps the exact-length contract for the persistent regions and staging. get_buffer returns the reference's window, not the physical slice.
  • The capture reads are windowed (scheduler::window_of): the CPU readback, the staged Metal readback and the CUDA capture_enq/capture_drain path all fed data.len() to trace::analyze as n_total, so a class-rounded buffer would have reported 256 elements for a 1-element token_ids input and changed every viz statistic.
  • pool_bytes is what the pool holds (Backend::pool_len is the probe): it only grows when the pool creates a buffer, so a recycled class buffer is not charged again. S1 charged every request, which made the report grow on every rebuild (and the budget gate refuse graphs that fit). Staging buffers are now charged too (exact, but resident and live until the next rebuild frees them).
  • Two latent bugs, both found by the rounding and both fixed here:
    1. A liveness extension did not move the pool's deadline. buf_alive is written from last_use at allocation; an in-place alias and a D1 view extend last_use later, and sweep reads only buf_alive, so a buffer could be handed to another node while the alias still read it. extend_through_views now reports the nodes whose liveness grew and extend_buffer_alive bumps their buffers' deadlines. D1's view branch had this wrong twice over: it called the walk with from = the view and to = the view's own last use, so the first check ended the walk — the parent extension never ran at all; it now walks from the parent.
    2. An input could take a buffer the walk released. Inputs are host-filled before execution, so the previous owner's write (during execution) lands after the fill: seq_ids took the matmul K buffer in the 0.5B decode graph and read -8.47 where its cell index belonged. All input buffers are now placed before the walk, when the free list still holds only the previous graph's buffers.

Acceptance, as measured.

  • a_rebuild_inside_one_class_reuses_the_pool: two rebuilds inside one class (896×16 = 14336 and 896×17 = 15232, both class 16384) leave pool_bytes unchanged and add zero pool buffers; a shape that leaves the class (896×20) does reserve more. Mutation-checked against exact pools and against per-request pool_bytes accounting.
  • a_fill_must_match_the_nodes_logical_length: a 3-element input occupies one class, a 3-element fill is accepted, and 2- or 4-element fills are refused naming both numbers; the refused fills wrote nothing and get_buffer reads exactly 3 elements. Mutation-checked (drop the check).
  • an_input_never_takes_a_buffer_the_walk_released and a_view_keeps_its_parents_buffer_alive_through_later_consumers: two synthetic execute-and-compare graphs whose values change if the buffer is recycled early. Both mutation-checked (remove the input pre-pass; extend from the view again).
  • Suites: CPU 265 passed / 0 failed / 12 ignored, CUDA on GB10 (serial, scripts/cuda_test.sh) 316 passed / 0 failed / 13 ignored (+4 over S1's 312).
  • Real-model: the nine non-ignored real-model graph tests that the un-migrated rounding broke (0.5B q4_0, 24 layers) — graph_logits_match_forward_real_model, a_two_sequence_batch_matches_two_single_sequence_forwards, offset_sensitivity_is_narrowed_to_multi_query_attention and the rest — pass again, and the whole CPU suite is green a second time with MINFER_BATCH_TEST_MODEL pointed at Qwen3-0.6B-Q8_0.gguf (f16 KV, the second cache-width path). The #[ignore]d real-model set, run serially (cargo test --release --bin minfer -- --ignored --test-threads=1), is 12 passed / 0 failed — the same as on master; in parallel it is red, on master too, because the C4 packed-cache gate sets the process-wide KV format mid-run (issue #99, and now documented in AGENTS.md). MINFER_TRACE on the 0.5B reports n = 30 for the token_ids input (the prompt's length); with the readback window removed it reports the class's 256, which is the mutation evidence for the capture fix.

Findings filed while doing S2 (both pre-existing, neither is caused by the allocator): #98 — ComputeGraph::topo_order counts in-degree per source entry but decrements once per node, so a graph with a repeated source (add(x, x)) is rejected as a cycle, and alloc_graph calls it on every build; fixed 2026-09-26 (record below); #99 — the process-wide KV format above.

What S2 does not do (it stays on #55): the reserve/assign re-map (a reserved region a rebuild re-maps without touching the device — the literal "split reservation from assignment", which the later E4 S3 landed for every backend in the allocator itself, Metal included) and the multi-graph cache (§14 row 3, which is what removes E3's per-chunk rebuild). Cross-boundary staging is charged to pool_bytes but is still allocated at its exact length, and backend-internal scratch (Metal capture staging, CUDA positions scratch) is outside the report.

Graph record (#98, 2026-09-26) — a repeated source is one edge, not two

The defect. ComputeGraph::topo_order counted in-degree per source entry (for &s in &node.src { indeg[node.id] += 1 }) but released it once per node (if self.nodes[v].src.contains(&u) { indeg[v] -= 1 }). A consumer that lists the same predecessor twice therefore never reached in-degree 0 and the validator reported a false cycle — on the minimal add(x, x) graph, cycle detected: 1/2 nodes ordered — even though the DAG was legal.

Reachability. GraphBuilder::add/mul pass &[a, b] straight through with no dedup, so b.add(x, x) (2 * x written as an addition) is buildable; GraphAllocator::alloc_graph validates with topo_order()? on every build, so the graph could not be allocated at all (BackendScheduler only had a debug_assert!, so allocation was the hard failure). Found while writing E4 S2's input-buffer gate, which worked around it with two distinct inputs.

The fix. The in-degree pass now counts each distinct predecessor once (!node.src[..j].contains(&s)), so both passes share one notion of "u is a predecessor of v". A duplicate source is one edge read twice, not two edges. The release pass is unchanged, and duplicate sources stay legal by design — the issue's intent is that add(x, x) works, not that it is rejected.

The tests. topo_order_accepts_a_repeated_source (the order covers every node, source first) and a_repeated_source_allocates_and_executes_as_two_reads (end-to-end through alloc_graph + BackendScheduler::execute, asserting the value: add(x, x) = 2x and mul(x, x) = x² on concrete inputs, so a graph that dropped the second read fails instead of passing an is_ok()). The control arm, topo_order_detects_cycle, now asserts the exact message cycle detected: 0/2 nodes ordered: a fix that simply stopped detecting cycles fails it. The E4 S2 gate an_input_never_takes_a_buffer_the_walk_released keeps its two-input form on purpose (its property is input placement, not the duplicate-source path) and carries a comment naming #98 as the reason it was written that way.

Mutation evidence (rule 3). Reverting the in-degree pass to the per-entry form (dropping the !node.src[..j].contains(&s) guard) and running both new tests:

$ cargo test --release repeated_source
test graph::alloc::tests::a_repeated_source_allocates_and_executes_as_two_reads ... FAILED
test graph::tests::topo_order_accepts_a_repeated_source ... FAILED
thread '...a_repeated_source...' panicked at src/graph/alloc.rs:4123:37:
add(x, x) must allocate, got: cycle detected: 1/2 nodes ordered
thread '...topo_order_accepts_a_repeated_source' panicked at src/graph/mod.rs:337:36:
add(x, x) is acyclic: "cycle detected: 1/2 nodes ordered"
test result: FAILED. 0 passed; 2 failed; 0 ignored; 0 measured; 493 filtered out
exit=101

Restored byte-for-byte (diff -q clean).

Adjacent audit (the contains-vs-occurrence asymmetry). The only other src.contains( in src/graph/ is GraphAllocator's "is this input consumed by a KV-indexing node" existence test — a duplicate does not change existence. The liveness last_use pass takes a max over source entries, so a duplicate is idempotent. n_consumers does count per entry, but over-counting only makes the in-place rule's == 1 test stricter: it refuses an alias and keeps a private buffer, which is conservative and never a wrong read (a comment now says so next to the count). extend_through_views / extend_buffer_alive walk the single view.src chain and max a deadline, so no source list is involved. The fusion pass reads mul.src[0] / src[1]; a duplicate is arithmetically preserved because SwiGLU(gate, up) = silu(gate) * up is exactly the Mul(Silu(gate), up) it replaces. The scheduler pushes one input buffer per source entry (so add(x, x) really reads the buffer twice) and dedups cross-split inputs by contains (existence again). cache.rs's reuse identity compares src vectors element-wise, so duplicates are deterministic. No other defect found; no follow-up issue needed.

Counts (rule 4). cargo test --release on dgxspark (aarch64), 2026-09-26: unit 462 passed / 0 failed / 33 ignored (was 460; +2 for the two new tests), integration 10 / 0 / 6. The x86_64 (CI runner) row moves by the same +2 (458 → 460); test-linux-cpu's --check-live confirms it against its own log, and AGENTS.md and docs/status.toml carry both rows.

E4 record, S1 (2026-09-22) — account first, allocate second

Why. Roadmap §2.3 listed three consequences of an allocator that places and hands out memory in one step: an exact-match pool ends up with one set of buffers per shape, nothing can answer "will this fit?" before the device is touched (CUDA surfaced it as a null pointer at execute time), and peak memory was not a number anyone could read.

What S1 landed.

  • src/graph/allocplan.rs — the size-class ladder and a pure plan. class_size(elems): powers of two up to 16 KiB (4096 elements), then multiples of 16 KiB, floor 1 KiB; a class never wastes more than one 16 KiB step. AllocPlan::plan takes the (size, first use, last use) intervals and simulates the pool's own reuse rule (a buffer freed at step f may serve an interval whose first use is after f), reporting reserved_bytes, live_peak_bytes, buffers and reused.
  • Accounting — GraphAllocator::memory_report(backend) returns MemoryReport { weights_bytes, pool_bytes, live_bytes, peak_live_bytes, budget } plus headroom_bytes(). pool_bytes is the pool's high-water mark (the pools never return memory), live_bytes is what is handed out right now, and the peak is tracked as the build proceeds. weights_bytes comes from the backend (a new Backend trait method with a 0 default; CPU sums its registry, CUDA sums the device registry — Metal inherits the default; it does not track its device-resident weights, which are mmap-backed or per-weight copies).
  • The feasibility gate — alloc_in_pool is fallible and checks weights + pooled + this allocation (at its class size) against the backend's budget before the pool is asked for anything. The default budget is the backend's own answer: CUDA's current free bytes, or Metal's recommendedMaxWorkingSetSize since #53, with a quarter held back, resolved through the pure allocplan::budget_decision over an explicit allocplan::DeviceMemory outcome — a failed query is not a number (it falls back to weights-only accounting with the real device error named once; see the S4 record below). CPU is unbounded unless set_memory_budget sets one (tests). The refusal names the numbers: weights, pooled bytes, this request, budget, all in MiB.

Acceptance, as measured (in CI, no device needed):

  • a 4095-byte budget against a 4 KiB activation is refused with out of CPU memory: … MiB of weights + … MiB of pooled buffers + … MiB for this activation exceeds the … byte budget (… MiB), and the pool is untouched (n_cpu_buffers() == 0, pool_bytes == 0) — the gate runs before the first backend call;
  • the same graph fits at 1 MiB, and the report then shows 0 < live_bytes <= pool_bytes, peak_live_bytes >= live_bytes and headroom_bytes() < budget;
  • weights and activations are one comparison: a budget covering the registered weight but not the activation is still refused, and one byte more is accepted;
  • the ladder is mutation-checked in both directions (the roundings above, plus the pure plan tests: two shapes in one class share, overlapping lifetimes do not, the plan is order-independent) and the gate itself is mutation-checked (disabling the comparison fails a_graph_that_cannot_fit_is_refused_with_its_numbers).

What S1 does not do (it stays on #55):

  • The pools still allocate exact sizes, so two shapes in one class do not yet share — the ladder is the plan's and the accounting's view, and the ticket's "recycled buffers come from the same size classes (no silent growth)" is not claimed yet. Rounding the pools is not a one-line change: every consumer's length contract has to move to the owning BufRef at the same time (the host read/write checks, the trace/dump capture, the scheduler's staging copy and the CUDA copy_to_host all return a physical length today). That is the reserve/assign re-map, and it is the next increment.
  • The reserve/assign re-map itself (a reserved region a rebuild re-maps into without touching the device) and the multi-graph cache (which §14 row 3 points at, and which is what removes E3's per-chunk graph rebuild) are S2 as well.
  • The pools still never return memory to the host/device (documented, accepted debt in docs/CUDA-BACKEND-DESIGN.md); the accounting now makes it visible as pool_bytes vs live_bytes.

E4 record, S4 (#122, 2026-09-24) — a failed device query is not a zero budget

What was wrong. GraphAllocator::memory_budget derived CUDA's default budget as CudaState::device_free_bytes() / 4 * 3, and CudaState::device_memory discarded the cudaMemGetInfo return code:

#![allow(unused)]
fn main() {
let (mut free, mut total) = (0usize, 0usize);
unsafe { cudaMemGetInfo(&mut free, &mut total) };   // return code dropped
(free, total)
}

On failure free stayed 0, so the budget became Some(0) and the E4 feasibility gate refused every later device allocation with

out of Cuda memory: 553 MiB of weights + 0 MiB of pooled buffers + 0 MiB for this
activation exceeds the 0 byte budget (0 MiB); reduce --n-ctx/--n-batch, …

The refusal was correct given a 0 budget; the defect was that a failed query was indistinguishable from "the device is full", and the message blamed the budget instead of naming the cause. weights_bytes() shared the class in the other direction: it summed the registry through .map(…).unwrap_or(0), so a poisoned lock silently reported 0 weights — the fail-open twin, which under-charges the same comparison.

Root cause, and what actually latched the error. Instrumenting the call (eprintln! of the return code plus cudaGetErrorName/cudaGetErrorString) gave

rc=700 name=cudaErrorIllegalAddress desc=an illegal memory access was encountered free=0 total=0

cudaErrorIllegalAddress is a sticky error: once a kernel in the context performs an illegal access, every later CUDA call in the process — cudaMemGetInfo included — returns 700 until the process ends. The origin was then located with compute-sanitizer --tool memcheck over the failing two-test subset (conversation_real_model_smoke + cuda_map_window_costs_no_more_than_the_span_it_replaces): 4 227 errors, of which 28 were

Invalid __global__ read of size 4 bytes
    at void fa_prefill_f16kv<(bool)0, (bool)0>(…)+0xe70
    by thread (64,0,0) in block (0,0,0)
    Access to 0xf654445fff00 is out of bounds
    and is 249 bytes after the nearest allocation at 0xf654445ffe00 of size 8 bytes
    … Host Frame: CudaState::gqa_attn_f16kv
    … Host Frame: cuda_map_window_costs_no_more_than_the_span_it_replaces

fa_prefill_f16kv indexes q as nt token rows of nh * hd (q[t * nh * hd + h * hd + d], t < nt), but the timing gate's prefill A/B reused the decode phase's single-row qb (nh * hd elements) while calling the entry with nt = 512, so the kernel walked up to ~7 MB past the buffer. When those device pages happened to be mapped the read was silent garbage; when the heap layout left a hole unmapped it faulted and the context was gone for the rest of the process. Which of the two happened is a test-order artifact — the latch needs the conversation test(s) and the map-window test ahead of the next E4 allocation (subset matrix, same binary: map alone → 1 passed, no 700; map+packed → rc 700 count 0; ctx+map+packed → rc 700) — but the masking is not: any sticky or fatal CUDA error from any cause (a different kernel bug, a driver hiccup, an OOM in an earlier call) made cudaMemGetInfo fail and handed the user a message about a 0-byte budget.

So both readings of the bug are true, and both are fixed: (a) a production robustness bug — a failed query silently became a 0 budget with a misleading reason; (b) a test-isolation artifact — a fixture passed an under-sized q, and only the test ordering decided whether that became a visible fault. The production path sizes q as nt rows (the Attn node's input), so the kernel indexing itself is correct.

The fix.

  1. allocplan::DeviceMemory { Reported { free, total }, QueryFailed { code, name }, NoDevice } makes the query's outcome a type; CudaState::device_memory checks the return code and names it with a new cudaGetErrorName extern, and device_free_bytes() -> Option<usize> is None on failure.
  2. allocplan::budget_decision(explicit, &DeviceMemory) -> BudgetDecision { budget, note } is the pure mapping: an explicit set_memory_budget wins; a reported read keeps the pre-existing free / 4 * 3 byte for byte (a genuine free == 0 still refuses); a failed query falls back to weights-only accounting (Some(usize::MAX)) with a note naming cudaErrorIllegalAddress (700), printed once per process; no device state is unbounded and silent. The fallback lets the backend's own allocation be the authority — it reports the real error if the context is genuinely unusable — instead of turning a broken accounting query into a total outage.
  3. E5's fit shares the defect through device_free_bytes(), where a failure became Some(0) → weight_budget → 0 → fit_blocks(0, …) → 0 device blocks planned while the startup line said device free 0 MiB (three quarters of it, …). models::device_memory() now returns the three-way outcome, and weight_budget refuses auto with the real CUDA error (offering MINFER_GPU_MEM as the escape hatch) rather than planning around a non-measurement; auto_source can no longer render an unmeasured device free 0 MiB. "No device at all" keeps the documented Metal behaviour (fit nothing).
  4. weights_from_lock(lock, what, sum) recovers a poisoned registry (append-only, so the map behind the poison is still valid) and says so, instead of reporting 0 bytes.
  5. MemoryReport::budget_is_bounded() / headroom_bytes() treat the unbounded sentinel as unbounded, and kv_snapshot_from omits the minfer_memory_budget_bytes / headroom_bytes gauges for it, so the metrics surface never publishes a number that was not measured.
  6. The other memory queries whose return code was discarded in the same neighbourhood: the init banner now reports a failed read instead of printing 0 MB; the w16_cache memory-pressure valve became fail-closed (rc != 0 skips the optional cache, where rc == 0 && … used to fall through and allocate it); plane_budget_ok keeps the conservative "planes off" decision but prints the error name once instead of silently reading a failure as "not enough memory".
  7. The trigger is fixed at both ends: the fixture allocates its own nt-row qb2 (the timing assertions are untouched — this is a buffer-size fix, not a weakened gate), and the CUDA Op::Attn arm refuses a q input shorter than n_tokens * n_head * hd before the kernel can read past it (the BufRef::len contract of rule 13), so this class of mistake is a loud Err rather than a corrupted context.

Measured before / after (GB10, sm_121, CUDA 13.0, driver 580.178.04; all runs serial, --test-threads=1).

GateBefore (eeba0d0)After
Minimal repro (issue #122's 4 filters)3 passed / 1 failed, exceeds the 0 byte budget3 passed / 1 failed, the failure is the #87 q8_0-on-CUDA refusal
Full #[ignore]d serial set (CUDA)5 passed / 14 failed, 26 × exceeds the 0 byte budget20 passed / 2 failed, 0 × that message (17/2 before the #47 F2 rebase added three tests to the set)
compute-sanitizer --tool memcheck on the latching subset4 227 errors, 28+ Invalid __global__ faults11 errors, 0 kernel-memory faults (6 cudaGraphDestroy + 4 cudaGetLastError from graph_replay_step, 1 cudaFuncSetAttribute at init — filed as #128)
cargo test --release (CPU)~340 passed / 0 failed / 18 ignored + 3/0/6382 passed / 0 failed / 21 ignored + 3/0/6 on the rebased tip (346/0/18 before the F2 rebase; the +6 here are the new gates)
cargo test --release --bin minfer -- --ignored --test-threads=1 (CPU, no cuda feature)12 passed / 0 failed (a stale count in AGENTS.md)18 passed / 0 failed

The two residual failures were both #123's, and are fixed there. At this record's time a_packed_kv_cache_answers_like_the_f32_one was CPU-only by its own docstring and was refused on CUDA (it failed alone the same way before this change once the budget was healthy), and a_partial_offload_runs_the_rest_on_the_cpu was second-hand damage: the packed gate set the process-wide KV format to q8_0 and panicked at the #87 refusal before its set_kv_format(F32) restore ran, so the next test built a q8_0-sized CUDA KV region and was refused (issue #99's mechanism). Neither was a budget failure — the 0-byte-budget string was absent from the run. The test-hygiene record (#123) below fixes both: the packed gate is CPU-forced and its format restoration is panic-safe (twice over: the normal path and a Drop guard), so the serial CUDA set is 22 passed / 0 failed.

Acceptance, as measured (pure, CI-covered).

  • a_failed_device_query_is_not_a_zero_budget: a QueryFailed { code: 700, name: "cudaErrorIllegalAddress" } decision is not Some(0); the budget is unbounded, the note names the error and the code, and the note contains neither 0 MiB nor 0 byte budget.
  • a_reported_free_read_keeps_the_three_quarters_default: free = 4 MiB → 3 MiB, and the integer rounding 4 000 001 → 3 000 000, i.e. the happy path is the pre-#122 number.
  • a_measured_zero_free_read_is_still_a_zero_budget: a real free == 0 still refuses, with no excuse attached.
  • no_device_state_and_an_explicit_budget_are_unchanged: NoDevice is unbounded and silent; an explicit budget is taken as given on every outcome.
  • a_failed_device_query_refuses_an_auto_fit: E5's auto returns Err naming cudaErrorIllegalAddress (700), offers MINFER_GPU_MEM, and never quotes a fabricated 0 MiB; the explicit cap still plans.
  • the_weight_budget_prefers_the_explicit_cap / the_auto_source_names_what_the_fit_decided: the pre-existing matrix, adapted to the three-way outcome, plus "an unmeasured device is not device free 0 MiB".
  • a_poisoned_registry_does_not_report_zero_weights: a mutex poisoned by a panic is recovered and reports its real 18 bytes, not 0.
  • Mutation check (all three reverted before landing): making budget_decision return Some(0) on QueryFailed, making weight_budget return Ok(0), and making weights_from_lock return 0 on a poisoned lock each make their gate fail — 0 passed; 3 failed.

Honest scope.

  • Metal is not exercised: no Mac here. DeviceMemory::NoDevice is the Metal arm and keeps auto fitting nothing without MINFER_GPU_MEM; CI's build-macos job compiles the Metal backend, nothing more. MINFER_DISABLE_CUDA / no-device behaviour is unchanged and unit-tested through NoDevice.
  • The trigger (the under-sized fixture q) is a test defect. The device evidence that production is unaffected is structural — the graph builder sizes the Attn input as nt * n_head * hd and the real-model gates pass — not a device A/B of a fixed production buffer, because there was no production buffer to fix. The trigger's test-order dependence is measured (subset matrix + one sanitizer run), not inferred.
  • The compute-sanitizer numbers are a memcheck A/B, not a production measurement: the sanitizer changes allocation layout and timing, so the map-window timing assert fails under it (1.405x) for #123's reasons. The kernel-fault count (28+ → 0) is the part that matters.
  • The SIGTERM drain path and the rest of the F8 wiring are untouched.
  • The residual a_partial_offload failure is attributed to #123/#99 by mechanism and by the error string; it is not fixed here, per the ticket's "do not fix #123 here".
  • No Metal run, no full non-ignored CUDA suite run: the ticket's gates are the #[ignore]d serial set plus the CPU suites, all run as above.

Follow-ups. #123 (the two residual failures: the CPU-only packed gate and the load-sensitive timing margin), #87 (no CUDA q8_0 KV kernel), #128 (the API-level cudaErrorInvalidValue findings: cudaGraphDestroy on an exec handle, which leaks the exec, and the eager cudaFuncSetAttribute opt-in failing at init).

E4/E5 record, Metal half (#53, 2026-10-05) — Metal answers the device-memory question

Box macbook (macOS 27.0.1, Apple M4 Pro), Rust 1.97.1 (the repo pin), commands as stated.

What was wrong. allocplan::DeviceMemory had exactly one implementation — CudaState::device_memory() — so on macOS models::device_memory() answered NoDevice for every backend. Two consequences: E5's --gpu-layers auto fell through weight_budget(NoDevice) → Ok(0), fitted nothing, and planned every block on the CPU unless MINFER_GPU_MEM was set; and E4's memory_report(Metal) carried budget = None, so the accounting surface had no headroom_bytes() on a Mac. graph/alloc.rs::memory_budget also had a CUDA-only inline arm, so the E4 gate would not have seen a Metal budget even if one had been answered.

What landed.

  1. MpsState::device_memory() (src/metal/runtime.rs) — MTLDevice.recommendedMaxWorkingSetSize, answered through the same three-way outcome: a non-zero value → Reported { free, total } (both the same number, because Apple Silicon's unified-memory device exposes no separate total); 0 → QueryFailed (a device that gives no figure is not a measured zero, [#122]); the shared seam MINFER_TEST_CALL_FAIL=metal_device_memory → QueryFailed ([#171]). code is 0 — Metal has no numeric error here — and name carries the reason.
  2. models::device_memory() (src/models/mod.rs) — the one resolver now routes Metal first on macOS, then CUDA, then NoDevice; it is the same function E4's gate and both loaders' auto fits read, so the fit and the gate cannot disagree.
  3. GraphAllocator::memory_budget (src/graph/alloc.rs) — every non-CPU backend asks that resolver instead of the CUDA-only inline arm; CPU stays NoDevice (unbounded). The pure budget_decision / weight_budget and the DeviceMemory variants are unchanged.

The common-module decision the campaign left open is recorded in docs/SOURCE-LAYOUT-PLAN.md §1.3 rule 3: no new module — 2 implementations (CUDA, Metal), 3 same-semantics callers of the resolver (E4's gate + the two loaders), and the type/policy already live in the device-agnostic allocplan/offload.

Acceptance, as measured (box macbook (macOS 27.0.1, Apple M4 Pro), 2026-10-05).

  • recommendedMaxWorkingSetSize on this machine is 40 200 896 512 bytes (38 339 MiB), asserted equal to the value the gate reports.
  • ./target/release/minfer --gpu-layers auto <qwen2.5-0.5b q4_0> "hello" → auto: 24 of 24 blocks fit — weights budget 28754 MiB, 7188 MiB reserved for KV/activations; device free 38339 MiB (three quarters of it, the E4 default budget), and the model runs on Metal (a follow-up prompt generated tokens). Before this change the same command planned 0 blocks.
  • MINFER_GPU_MEM=1024 … --gpu-layers auto → weights budget 1024 MiB, 256 MiB reserved …; MINFER_GPU_MEM=1024 MiB — the explicit cap still wins and is named.
  • MINFER_TEST_CALL_FAIL=metal_device_memory … --gpu-layers auto → the load refuses with the device free-memory query failed with [testfail] … (code 0); set MINFER_GPU_MEM=<MiB> … — a QueryFailed, never a fabricated zero.
  • memory_report(Backend::METAL) carries budget = 3/4 × free and headroom_bytes() == budget on an empty pool (graph::alloc::tests::budget::metal_memory_report_carries_the_device_budget); a graph pinned to Metal with a 4095-byte explicit budget is refused before the pool is touched (pool_bytes == 0).

Mutation evidence (each reverted before landing; cargo test --release, box as above).

mutationgateresult
delete the testfail::guard("metal_device_memory") armmetal::tests::a_forced_device_memory_query_failure_is_not_zero0 passed; 1 failed — "a forced failure must be QueryFailed, got Reported { … }"
memory_budget returns NoDevice for every backend (the pre-change shape)graph::alloc::tests::budget::metal_memory_report_carries_the_device_budget0 passed; 1 failed — "Metal must now answer a budget"
Reported { free: 0, … } instead of the API valuemetal::tests::device_memory_reports_the_recommended_working_set0 passed; 1 failed — left: 0, right: 40200896512

Honest scope.

  • The macOS suite baseline is red and stays red: on this box cargo test --release --no-fail-fast is 483 passed / 20 failed / 38 ignored unit + 21/0/6 integration at 119b9da, and 485 passed / 21 failed / 38 ignored after this change. The three new unit gates all pass; the extra failure is the pre-existing order-dependent flake models::qwen2::graph::tail_tests::cuda_conversation_multiturn_reuse, which is part of the #255 red baseline (482/21 in one fail-fast baseline run before this change, 483/20 in a --no-fail-fast one; alone it passes 3/3 on master). No new failure is attributable to this change, but the suite is not green and this record does not claim it is.
  • Metal's weights_bytes stays the trait default 0: Metal's weights are mmap-backed NoCopy buffers or per-weight copies, not a registry the pool owns, so the E4 gate charges the device's own budget without a weights term there. CUDA and CPU keep their sums. This is the documented default ("a backend that does not track its weights is not charged"). The asymmetry it leaves is #299: the E5 auto fit does charge those weights (from the GGUF index) while the E4 gate does not, so on the measured Mac (recommendedMaxWorkingSetSize 40 200 896 512 B ≈ 38 339 MiB) a 7B Q4_K_M's ~4.4 GiB of mmap-backed weights can be admitted beyond what the pool has room for — bounded by the weight bytes and unable to corrupt anything, but the one place where the accounting rule's weights term is silently zero on a platform that has them.
  • No CUDA re-run: the allocator's CUDA arm now calls models::device_memory(), which on a CUDA build returns CudaState::device_memory() exactly as before; CI's build-linux-cuda compiles it, but the device behaviour is not re-measured here (no GPU).
  • The real-model set was not re-run as a set (the branch's #[ignore]d Mac baseline is already red per #255); the auto acceptance above is a single real-model CLI run, not the suite.

Test-hygiene record (#123, 2026-09-24) — the serial #[ignore]d set goes green on a CUDA build

What was wrong. The documented CUDA command cargo test --release --features cuda --bin minfer -- --ignored --test-threads=1 could never be green, for two reasons that have nothing to do with the code under test.

  1. A documented CPU-only gate ran anyway. a_packed_kv_cache_answers_like_the_f32_one says in its own docstring that a CUDA box refuses q8_0 by design (#87), and the default offload request let the device claim the model, so it failed alone with the ensure_kv refusal (KV region for layer 0 would live on Cuda, which has no kernel that reads a packed q8_0 region …).
  2. A load-sensitive timing margin. cuda_map_window_costs_no_more_than_the_span_it_replaces asserted p_map <= p_span * 1.25 from one 20-launch block per mode, span first and map second. Load arriving during the map block had nothing to absorb it: a loaded GB10 measured 1.267x (2.232 vs 1.761 ms), and a rerun of the same binary passed.

Collateral (this is #122's record, above). When the packed gate panicked at the #87 refusal it never reached its set_kv_format(F32) restore, so the process-wide KV format stayed q8_0 and the next test in the serial set — a_partial_offload_runs_the_rest_on_the_cpu — sized its CUDA KV region for the wrong format and was refused too (the mechanism of #99). The baseline at bac8440 was therefore 20 passed / 2 failed, both failures carrying the identical #87 string.

The fix.

  1. The packed gate is device-aware without losing coverage: it loads its model with OffloadRequest::Layers(0) (--gpu-layers 0, the established all-CPU configuration) and asserts model.device() == Device::Cpu. On a CUDA build the gate therefore executes the packed path CPU-forced instead of being skipped — the choice that keeps the most real coverage — and on a CPU build nothing changes. It prints the reason it is CPU-only, naming #87.
  2. A KvFormatGuard snapshots the process-wide format before the gate flips it and restores it on Drop, so a panic anywhere in the gate cannot leak q8_0 into the next test. The normal path's explicit restore stays; the guard is the panic-safe backstop. (Per-engine format is #99 and was deliberately not implemented here.)
  3. The timing gate interleaves the two modes' rounds (span, map, span, map, …) and asserts the median of the per-round map/span ratios — a matched pair per round, robust to up to rounds / 2 disturbed rounds. Decode: 9 rounds × 100 launches; prefill: 9 rounds × 50 launches (the old prefill form had no round structure at all), each launch group with a 3-launch warm-up. The threshold is unchanged at 1.25x — the statistic is the fix, not a wider margin — and the gate prints every per-round ratio and the medians it used.

Measured (GB10 sm_121, CUDA 13.0, driver 580.178.04; all runs serial, --test-threads=1).

GateBefore (bac8440)After
Packed gate alone (CUDA build)0 passed / 1 failed, the #87 refusal1 passed, [c4] CPU-only by construction …; max |Δlogit| 3.0289, region 3.76x smaller — identical to the CPU build
Full #[ignore]d serial set (CUDA, 0.5B)20 passed / 2 failed (both #87)22 passed / 0 failed, ten consecutive runs (5 + 5) of the same binary
Full #[ignore]d serial set (CUDA, Qwen3-0.6B config)—21 passed / 1 failed; the one failure was an unrelated C5 defect (a session could not encode an f16 element type), filed as #130 and closed 2026-09-25 (C5 S3: 31 / 0)
Timing gate, decode ratio1.001 / 1.001 / 1.006 / 1.001 / 1.001 (5 runs)1.001–1.004 (6 idle runs), 1.001–1.018 (6 loaded runs)
Timing gate, prefill ratio1.088 / 1.087 / 1.092 / 1.021 / 1.090 (5 runs)1.087–1.107 (6 idle), 1.079–1.145 (6 loaded)
cargo test --release (CPU)382 / 0 / 21 + 3 / 0 / 6382 / 0 / 21 + 3 / 0 / 6 (unchanged)
cargo test --release --bin minfer -- --ignored --test-threads=1 (CPU)green21 passed / 0 failed, with the packed gate executing

The pre-#123 prefill sample 1.021 is the old statistic's failure mode from the other side: a spike hit the span block, the ratio went down, and the gate would have passed for the wrong reason. Load for the "loaded" rows: 16 CPU spinners plus two concurrent processes running cuda_verify_attention_nt_invariance in a loop. Individual per-round ratios reached 8.60x under that load; the median absorbed them (worst prefill median 1.145, worst decode median 1.018).

Mutation checks (all reverted before landing).

  • Re-breaking the device-awareness (the pre-#123 load_model) makes the packed gate red again with the exact #87 refusal: 21 passed / 1 failed — and the collateral test stays green, because the guard now restores the format on the panic path.
  • Removing the guard as well restores the bac8440 baseline exactly: 20 passed / 2 failed, both with the #87 refusal. So both halves — the CPU-forced load and the guard — are load-bearing.
  • Doubling the map path's work in the timing gate's prefill closure trips the assert at 2.190x (182.2 vs 83.3 µs/launch), so the gate is live. Since the intrinsic ratio is ~1.09, a uniform map regression of ≥ ~15% crosses 1.25x and trips it.

Honest scope.

  • Metal is not exercised: no Mac here; CI's build-macos job compiles the Metal backend only. The packed gate's CPU-forcing is backend-agnostic, so it holds there too.
  • #99 and #87 remain open. #99 (per-engine KV format) was explicitly out of scope; the guard is a test-local mitigation, not the format-ownership fix. #87 (a device kernel that reads a packed q8_0 region) is why the gate is CPU-forced: when it lands, the gate's device() == Cpu assertion will fire and it should be re-pointed at the device.
  • The Qwen3-0.6B configuration's serial set was 21/1 for an unrelated pre-existing reason — the C5 container had no F16 flag, so a CUDA f16 session was refused on load (#130, closed 2026-09-25: FLAG_F16, C5 S3); that test failed alone with no #123 code in its path, so it was not a #123 regression.
  • The decode threshold could in principle be tighter than 1.25x (its intrinsic ratio is ~1.00). It is left at the pre-existing 1.25x on purpose, so the ticket removes flakiness without weakening the gate; it is not so wide that a real regression escapes (the 2.190x mutation trips, and the arithmetic above puts the detection floor at ~15%).

Follow-ups. #87 (device packed-q8_0 attention), #99 (per-engine KV format — landed 2026-09-25; its record below removes the global, and with it this record's item-2 KvFormatGuard), and #130 (the f16 KV session round trip, found while gating this ticket).

Test-infrastructure record (#99, 2026-09-25) — the KV format is per engine, and the ignored gate set stops lying in parallel

What was wrong. The #[ignore]d real-model gate set was red whenever the harness ran it in parallel. Measured on the CPU build at a756419: 19 passed / 9 failed parallel against 28 passed / 0 failed serially. Every failure had one shape —

KV region for layer 0 was allocated with 57600 elements but 15300 are requested
(n_ctx changed on a live GraphCache; the regions are persistent)

— with 57600 / 15300 = 3.765, exactly the Q8_0 packing ratio. One cause: models::qwen2::graph::tests::a_packed_kv_cache_answers_like_the_f32_one (the C4 packed gate, device-aware since #123) called kvformat::set_kv_format(KvFormat::Q8_0) for its measurement runs. The format was a process global read by GraphBuilder::new and CpuBackend::new, so while the gate ran any other test building a graph sized its KV nodes for the packed format while its own GraphCache — or the model it compared against — had been sized under another one; ensure_kv refused the mismatch, correctly (the regions are persistent). #123's Drop guard made the serial set deterministic by restoring the format on a panic, but it could not remove the hazard for a test running concurrently; the parallel red set was the reason the E4 S2 gate run was ambiguous (see the S2 record above).

The fix (the real one, not the guard). The format is now a property of the engine:

  • models::load_model_configured(gguf, ns, offload, cache_type) resolves MINFER_CACHE_TYPE once against device() and stamps the answer on the loaded model (ModelDef::kv_format / set_kv_format). load_model_with / load_model_ns are the environment-backed wrappers; the explicit argument is what a test passes instead of mutating the environment (which is process-global too).
  • CParams::kv_format carries it into the build and is part of the reuse identity, so a cached graph is never reused across formats; Qwen2Graph::build / Qwen3Graph::build call GraphBuilder::set_kv_format(params.cparams.kv_format), and GraphBuilder::new defaults to F32 instead of reading a global.
  • GraphAllocator::set_kv_format gives the CPU kernels the same answer (CpuBackend's field, which the registry's kv_format hook and kv_element_format read); forward_batch calls it once per forward with model.kv_format.
  • spec::SpecEngine::new takes the target's format as an argument instead of reading the global; the graph JSON exporter (graph/json.rs) takes it from the model.
  • graph/kvformat.rs's KV_FORMAT static, set_kv_format / kv_format and KvFormat::from_code are deleted, and the obsolete the_process_wide_format_can_be_redecided unit test with them (the CPU unit count therefore moves 438 → 437 passed, same 28 ignored).

The C4 gate now proves coexistence instead of mutating shared state. It loads two engines per arm — one resolved f32, one q8_0 — through load_model_configured, in a CPU arm and (on a CUDA build with a device) a device arm, and asserts each engine's kv_format(), its backend, the 3x-smaller region and the two logit tolerances (unchanged bounds). The KvFormatGuard is gone — nothing process-wide is left to restore. The device arm still sets the one process-wide tag #99 left in place (cuda::KV_LAYOUT, read by the device kernels themselves) under a small DeviceLayoutGuard that restores it on a panic. Both of those device-only parts are gone as of #153 (the guard deleted, the tag per engine — see the #153 record below); the sentences above describe the state at #99's landing.

The CUDA device half was deliberately not done at #99 and landed as #153. cuda.rs held the layout in a process-wide KV_LAYOUT that the launchers read directly (not through CudaBackend::kv_layout), so a per-engine device path needed the launchers and the captured-graph key threaded; a per-instance field alone would silently still have read the global. The #153 record below states what is now per-engine and what is not.

The entry point. scripts/real_model_gates.sh is the one command for the set: it defaults to --test-threads=1 (required on a device) and takes PARALLEL=1 (CPU-only parallel) and FEATURES=cuda. Documented in AGENTS.md rule 11 + the real-model-gates bullet and in docs/BUILD.md §Tests.

Measured (CPU build, dgxspark; --bin minfer for the gate set).

runbefore (a756419)after (rebased on 09ce9e7, #121 included)
cargo test --release --bin minfer -- --ignored (parallel)19 passed / 9 failed28 passed / 1 failed
cargo test --release --bin minfer -- --ignored --test-threads=128 passed / 0 failed29 passed / 0 failed
cargo test --release (unit + integration)438 / 0 / 28 + 10 / 0 / 6438 / 0 / 29 + 10 / 0 / 6
C4 gate alone ([c4] print)—cpu: f32 6 291 456 B vs q8_0 1 671 168 B (3.76x), max |Δlogit| 3.0289 of a 37.79 spread; shifted max |Δlogit| 2.4662

The unit count is unchanged because two opposite moves cancel: #99 deletes the obsolete the_process_wide_format_can_be_redecided test (the global it asserted is gone) and #121 adds one (38 -> 37 from #99, +1 from #121). The ignored set grew by #121's saturation gate, hence 28 -> 29 serial.

The single remaining parallel failure is not a KV failure and not a #99 regression: server::batch::tests::server_batch_matches_serial_and_is_faster asserts a wall-clock relation (t_serial > t_batch) from two sequential whole-workload measurements, so under a loaded parallel harness the first-measured phase absorbs the start-up wave — measured 15.60s batched vs 11.54s serial on the rebased tree (pre-rebase: 17.59 vs 9.94, then 18.39 vs 9.75), while the serial set passes it. All nine KV-region failures are gone. The load-sensitive assertion is the same class #123 fixed for the CUDA map-window gate and is filed as #154; until it is robust, the serial invocation is the documented entry point.

Mutation check (run before the #121 rebase; reverted byte-identically, sha256sum -c on all three files). Re-introducing the #99 shape — a process-global format read by GraphBuilder::new, flipped by the C4 gate per run — makes the parallel set 20 passed / 8 failed, every failure the KV region … was allocated with N elements but M are requested string with the 3.765x ratio. So the parallel gate set is what detects the interference, and the per-engine path is what removes it.

Honest scope.

  • Metal is not exercised (no Mac; CI's build-macos job compiles the backend only). The per-engine plumbing is backend-agnostic; Metal's kv_format hook still reads metal::kv_cache_is_f16, and its packed format is G5 on #44 either way.
  • The device (CUDA) parallel run is still not green and is not claimed. CudaState is a process-wide singleton (MMQ memo, captured graph execs, stream state — issue #64); cargo test --release --features cuda -- --ignored without --test-threads=1 remains the wrong command, and scripts/real_model_gates.sh keeps it serial. Measured serially on dgxspark (GB10 sm_121, CUDA 13.0): the unit suite is 501 passed / 0 failed / 32 ignored, and the ignored set is 32 passed / 0 failed in both the 0.5B (f32 KV) and the Qwen3-0.6B (f16 KV) configurations — unchanged by this increment except that one obsolete unit test is gone. (#153 removed the second reason — the device KV layout tag is per engine now — and re-measured the parallel form; see the #153 record below for the fresh counts.)
  • An explicit f16 cache type still lets the device layout follow set_kv_cache_type's auto policy (the pre-C4 split the loader comment records); the builder's f32/f16 region shapes are identical, so that is not a sizing hazard. A packed (q8_0) resolution is still restated explicitly on the device, as before. (#153 folded the auto policy into the engine's resolved format and deleted set_kv_cache_type; the loader no longer restates anything.)
  • The forward_graph (non-cached) path uses the process-global graph_cache(); two models with different formats driving that one cache would still be refused by ensure_kv. That is the pre-existing single-cache-per-process design, not the format global, and the server/CLI paths that matter pass a GraphCache explicitly.

Follow-ups. #154 (the load-sensitive batching timing gate); #153 (the CUDA per-graph layout) was its own follow-up and landed — its record follows.

#153 record — the CUDA KV layout is per engine, and the captured-graph identity carries it

What was wrong. #99 made the KV format per engine for the model, the graph builder, the allocator and the CPU kernels, but left the CUDA side on a process-wide static KV_LAYOUT: AtomicI32 in cuda.rs. CudaBackend snapshotted it into kv_layout at construction, yet the value came from the static (and entry()'s kv_format hook read the static directly), so a per-instance value would have been a half-wire: two engines loaded with different formats would still have run the last-loaded layout, and the loader had to restate the tag for q8_0 (models::load_model_configured). The honest-scope sentence in AGENTS.md rule 11 said exactly this and is the caveat this ticket retires.

The one authority is now the engine's resolved format. KvFormat::resolve takes the model dims and folds in the GPU's own auto policy (auto_device_format: f16 for the 7B class, f32 for small models; the CPU stays f32) — it used to live only in cuda::set_kv_cache_type, which is deleted along with kv_cache_layout / kv_cache_is_f16 / set_kv_cache_layout / set_kv_cache_f16. models::load_model_configured stamps the one answer on the engine, and:

  • CParams::kv_format carries it into the builder (region width) and the reuse identity, as before;
  • GraphAllocator::set_kv_format(format) stamps the CPU backend and, if it exists, the CUDA backend's tag (cuda::layout_of(format) → KV_LAYOUT_F32/F16/Q8_0);
  • GraphAllocator::enable_cuda builds a fresh CudaBackend::with_layout(...) from that stamp, so a backend created after the stamp cannot revert to a default;
  • the registry's kv_format hook answers from a.cuda().kv_format(), so a KV session's header element type is the engine's, and server::batch's row-width snapshot reads GraphAllocator::kv_format instead of the process static.

The kernels already took the tag as a launcher argument (gqa_attn_split, the prefill/verify attention, store_kv_f32/f16/q8_0, attn_bias_rope_store); what was missing was that the value came from a global. cuda::layout_of / format_of are the only bindings between KvFormat and the FFI codes, and the layout tag is exhaustive over the enum.

The captured-graph identity carries the layout. CapturedGraph gained a kv_layout field, and graph_replay_step's lookup refuses an exec whose recorded tag no longer matches the backend's — it is destroyed and the 3-run warmup restarts, exactly like a pool_gen change. set_kv_layout invalidates eagerly on a change as the first line of defence. The enforcement point is graph/cuda_backend.rs::graph_replay_step (the pool_gen / kv_layout comparison), with the identity recorded in close_capture_or_sync.

The device gate. models::qwen2::graph::tests::two_cuda_engines_with_different_kv_layouts_run_interleaved loads an f32 engine and a q8_0 engine before either runs, prefills and then decodes them interleaved on two live GraphCaches, and asserts four things: each cache's CudaBackend::kv_layout is the layout its engine named (and the two differ); the packed regions are ≥ 3x smaller; the packed engine's logits stay in the C4 class of the f32 engine's (at the argmax ≤ 1.0, tail ≤ 4.0); and each engine's interleaved logits are bitwise its own solo logits (isolation). Interleaved, not threaded: CudaState is a process-wide singleton and the capture path holds a process-wide stream lock, so two OS threads would serialize on that lock anyway — the form the ticket allows. A unit gate (cuda_graph_recaptures_on_kv_layout_change) pins the capture identity without the model, and alloc::tests::set_kv_format_stamps_the_cuda_layout_per_engine pins the stamp → backend path.

Measured (GB10 sm_121, CUDA 13.0, 2026-09-26).

  • scripts/cuda_test.sh: 539 passed / 0 failed / 38 ignored (was 536 / 0 / 37; #153 adds two device unit gates and one pure kvformat gate, and moves the two cuda::kv_dtype_tests to the KvFormat ↔ tag mapping they now assert).
  • FEATURES=cuda scripts/real_model_gates.sh: 38 passed / 0 failed in both the 0.5B (f32 KV) and the Qwen3-0.6B (f16 KV, hd 128) configurations (was 37 / 0; the new two-engine gate is #[ignore]d).
  • The two-engine gate alone prints [153] two live CUDA engines, interleaved 8 decode steps: tags f32=0 q8_0=2; regions f32 6291456 B vs q8_0 1671168 B (3.76x smaller); interleaved packed-vs-f32 max |Δlogit| = 2.479504 of a 37.821205 spread, at the argmax 0.59605026; interleaved-vs-solo drift 0 / 0.
  • Full cargo test --release --features cuda -- --ignored (integration targets included, which is what the ticket's acceptance line names): serial --test-threads=1 is green — the 38-test bin set plus the 6-test conversation_cli set, 0 failed. Parallel (default harness) is still red, with fresh counts and a fresh reason (below).
  • compute-sanitizer --tool memcheck --target-processes all over the CUDA unit suite: 0 API errors (539 / 0 / 38).

The parallel --ignored answer (restated, not a shrug). The parallel run remains the wrong command, and the reason is now measured rather than inherited: it is not the KV layout any more. Two runs of the same command:

  • run A: 33 passed / 5 failed, failures cuda_map_window_costs_no_more_than_the_span_it_replaces (a load-sensitive timing gate — its own per-round ratios ranged 0.63–2.20 with median 1.398x) and server_batch_matches_serial_and_is_faster (the known #154 timing gate), plus three process-global-state failures: conversation_real_model_smoke with cudaErrorStreamCaptureInvalidated (901) inside a capture window, async_cross_copies_never_block_and_stay_bitwise_identical comparing the process-wide cuda::stream_sync_count() (4160 async vs 728 sync across concurrent tests), and an_auto_offload_plan_fits_the_budget mutating the process-wide MINFER_GPU_MEM env;
  • run B: SIGSEGV (signal 11) after two unrelated failures — the failure set is not stable.

Every mechanism is the process-wide CudaState singleton (one stream, one capture window, one MMQ/pool state — issue #64) and the process-wide counters/env a few gates read, not the KV format: the layout is per engine now, and no failure names a KV region or a layout mismatch. scripts/cuda_test.sh and FEATURES=cuda scripts/real_model_gates.sh keep the device set serial.

Honest scope. The CUDA tag is per engine; Metal's kv_cache_is_f16 is still process-wide (its kernels read it, there is no Mac here to change it on — G5 for packed). The gate interleaves two engines in one process, which is the ticket's allowed form; it does not prove two threads can drive two CUDA engines concurrently, and the compute-sanitizer run is still the whole CUDA unit suite serially.

E3 record (2026-09-22) — a prefill in chunks, and what runs between them

Why. A prefill was one forward over the whole prompt. Two consequences: activation memory scaled with the prompt (GraphParams.n_tokens sizes every activation buffer), and a long prompt blocked every other slot's decode for its whole duration — the batch worker is single-threaded, so admit returning after the full prefill meant the other slots' tokens simply waited.

Rule. prefill_chunks(from, total, chunk) splits the fed suffix into spans of at most chunk tokens, in order, with the remainder last (so the final forward carries the tail row whose logits the request samples from); chunk == 0 is "off" — one span, the pre-E3 path — and a suffix that already fits is never split. MINFER_N_BATCH (default DEFAULT_PREFILL_CHUNK = 2048) sets it, and --n-batch-style plumbing is the same setter: BatchEngine::set_prefill_chunk. The server prints the outcome at startup, like the batching mode, because a default nobody can see is a default nobody can debug.

Interleaving is the point. Between chunks — never before the first or after the last — the prefill runs the same decode step serve_loop would have run next, if any other slot has a token waiting (has_pending_decode). That is safe by construction: the chunk's rows are already written, the slot has no Run yet, and finish keeps a completed slot's rows (B2 reuse) rather than compacting, so nothing moves under the in-flight prefill.

Why 2048 as the default: it is a no-op for every prompt that fits it — identical forwards, identical graph, identical timing — so the change cannot regress the common case, while a longer prompt gets bounded memory and interleaving. The cost of the split is real and measured, and it is one forward's fixed overhead per chunk (the weights re-stream and the graph re-fills: it is keyed on n_tokens, so a chunked prefill rebuilds once per chunk — §14 row 3; E4's multi-graph cache is what removes that):

promptforwardsCPUCUDA (GB10)
98 tokens, chunk 241 → 5484 → 485 ms (1.004x)29 → 57 ms (2.0x)
514 tokens, chunk 1711 → 42676 → 2691 ms (1.005x)107 → 119 ms (1.11x)

The small-chunk case is the worst one (five weight passes for 98 tokens); at the default, a 4096-token prompt is two forwards, i.e. one extra fixed cost on a prefill of that size.

Acceptance, as measured (cargo test --release --bin minfer -- --ignored, 98-token prompt, chunk 24):

  • Bounded: the chunked run issued 5 prefill forwards whose largest nt was 24; the unchunked run issued 1 forward of nt = 98. The bound is prefill_stats(), an observable, not a claim about buffers.
  • Equal: the prefill's tail-row logits are bitwise identical on CPU (max |Δ| = 0) and within the named cross-shape class on CUDA (0.218, class 1.0 — CUDA's prefill tiles by nt and quantizes activations to int8); the CPU continuation is equal byte for byte.
  • Interleaved (48×"buffalo " on slot 1 already decoding, a ~4× chunk prompt on slot 0): the chunked prefill ran 3 decode steps for the other slot and it emitted 3 bytes during the prefill call; with chunking off the same call ran 0 steps and the slot gained 0 bytes — the A/B is deterministic, not a timing race.
  • Both gates are mutation-checked: disabling the interleave fails the second (ticks 0, bytes 0), and making prefill_chunks never split fails the first (1 forwards for a 98-token prompt at chunk 24).

Coverage, stated rather than implied. The chunk plan and the env parser are pure and run in CI; the two real-model gates are #[ignore]d (they need the cached 0.5B) and were run on the CPU and on CUDA (GB10). Not in this increment: a --n-batch CLI flag (the setter and MINFER_N_BATCH are the surface), and true mixed prefill+decode batches — the decode steps between chunks are the same tick the worker already runs, one per chunk, which bounds the stall without yet sharing a weight pass between a chunk and a decode row.

E6 record (2026-09-19) — a device-aware default

Why it needed its own change. E2 measured the sign of the batching effect per device (CPU 0.49x, GPU 1.9x) and, correctly, left the default with the measured-better path on the box it could measure. With the GPU available that stopped being the honest default: a CUDA server was leaving a ~2x win on the table behind an environment variable almost nobody would set.

Rule. MINFER_BATCH unset -> batched iff the model's forwards run on CUDA or Metal; =1 -> batched (also the way to batch on CPU); =0 -> serial; any other value warns and falls back to the device default. Metal was excluded by construction, not caution: the batched path needs attn_span and Metal refused that node (supports_attn_span() was false there), so batching a Metal server would have failed loudly instead of serving.

Updated 2026-10-06 (#44 part (b), PR #316) — the rule now reads CUDA or Metal. The premise is gone: Metal reads the one-range attn_span (supports_attn_span() is true) and has the matching write/move side, so chat::batch_mode answers Batched for Device::Metal. The CPU row is unchanged (still Serial by default).

The decision has one authority. "The device participates" used to be recomputed inside each model's forward_batch (metal_available() && weights_on_gpu, CudaState::get().is_some() && weights_on_cuda). It is now Qwen2Graph::device / Qwen3Graph::device (an E6 addition), which the builder derives CParams.gpu from and the server derives its default from — so the two cannot disagree, and a partial weight registration (device present, weights not all registered) correctly keeps the server serial. ModelDef::device() exposes it as a Device::{Cpu, Metal, Cuda} with a CPU default.

The decision is pure and tested without a device, which matters because CI has no GPU: server::chat::batch_mode(requested, device) is a pure function and batch_mode_follows_the_device_and_honours_the_override covers the matrix (unset / "1" / "0" / invalid x cpu / metal / cuda). Every server now also prints the outcome at startup ([server] batching: on|off (device cuda|cpu|metal; ...)), so the default is observable rather than inferred.

Device verification (2026-09-19). CPU build: unset -> off (device cpu), =1 -> on, banana -> warning + off. CUDA build on the GB10: unset -> on (device cuda), =0 -> off. Acceptance re-measured with the default, no environment variable: 7B Q4_K_M, --n-slots 4, four identical prompts, max_tokens=16, equal work — 0.66 s batched (default) vs 1.30 s with MINFER_BATCH=0 = 1.97x, matching E2's 1.9x (1.314 vs 0.686) within noise. Suites: CPU 192 passed / 0 failed, --features cuda 238 / 0.

  • E1 acceptance: the mask is an explicit input, not a derivation from positions; a two-sequence test proves no cross-attention; single-sequence output is bitwise unchanged.
  • E2 acceptance: aggregate throughput at --n-slots 4 materially exceeds the serial baseline on a fixed workload; the n_seqs field is either real or deleted (closes A7 if it was kept). E2 design (written before the code, 2026-09-17). --n-slots N does not serve concurrently today: each Slot owns its own GraphCache and the worker is serial (server/slot.rs), so N only divides the context budget (n_ctx_slot = n_ctx / n_slots) — four slots buy nothing. E1 made the attention side sequence-aware; E2 has to make the bookkeeping and the serving loop sequence-aware:
  1. KvCache gains reservations. A sequence is a SeqId owning a contiguous cell run: reserve_seq(seq, cap) (first-fit over free cells, Err when the arena cannot fit it), release_seq(seq), and seq_range(seq) -> (start, cap) from the reservation rather than from a scan of owner. Ownership keeps its C1 meaning — it marks written cells — and becomes per-sequence: own_range(seq, from, to), replacing today's own_prefix(SEQ_MAIN, n) which would clobber a second sequence's cells.
  2. The single-sequence path is the cap == n_ctx special case. SEQ_MAIN takes the whole arena, so seq_range is (0, n_ctx) and attn_span still yields [0, pos + 1) — the existing model path stays bitwise, which is the refactor's acceptance gate.
  3. A batch is data. Batch { tokens, positions, seq_ids, n_out } with n_seqs = distinct ids; the graph is built per (n_tokens, n_seqs) (both already in GraphParams), and the allocator fills seq_ids/attn_span from the batch exactly as E1 does for one sequence. n_seqs becomes a real field (it now gates the multi_seq op flag and the graph identity) — A7's "real or deleted" question answered with real.
    • As implemented (outcome, 2026-09-17): the second half of that sentence did not survive contact. n_seqs was made real for one commit, then the flag it was said to gate moved to CParams.explicit_span (the reservation is the authority on the window, not the count), and a test showed the count now changed nothing about the topology while still forcing a rebuild. The field was deleted — A7's question answered with deleted, the other branch of the same acceptance clause. See §8 and the A7 ticket.
  4. The server batches decode steps. One shared cache for the batch; each active slot holds a reservation and its token stream; one forward per step carries every ready slot's next token (nt = ready slots, n_seqs = ready slots), so one weight pass serves all of them. Prefill stays per-slot in this increment (mixing prefill into a decode batch is E3's chunked prefill), and a decode batch is capped at 16 tokens so it rides the batched split-attention path on CUDA (fa_prefill's tile must not span two sequences — documented in E1b).
  5. Tests and measurement. Bitwise: a multi-sequence decode batch must equal the same sequences run one at a time on the CPU reference. Resolver: two reservations do not overlap, release_seq frees exactly its rows, Err when the arena is full. Serving: aggregate throughput at --n-slots 4 against the same workload run serially (the ticket's acceptance), reported as a ratio with the workload recorded.

Not in E2: mixed prefill+decode batches and chunked prefill (E3), moving a sequence's cells when the arena fragments (C3 needs D1; E2 reserves a slot's budget up front and fails loudly instead), layer offload (E5), Metal (G5), and CUDA runtime verification (deferred under A0; done 2026-09-18 for E1b and for C2's CUDA arm — see the sweep below).

E2 progress (2026-09-17). Step 1 of the design landed: KvCache holds per-sequence reservations (SeqSlot { start, cap }, reserve_seq first-fit, release_seq, seq_slot) and per-sequence ownership (own_range, replacing own_prefix's hard-coded sequence 0), and attn_span resolves a query's window from its sequence's reservation — with a written-row check (owner[pos] == seq), so a window can never include rows nobody wrote. A reservation is not ownership: reserving cells does not let a query attend to them, which is what keeps a batch's unwritten rows out of attention. after_rm/after_shift now refuse a multi-sequence cache explicitly (C2's shift is the single-sequence sliding window; a general move is C3).

The single-sequence path is the cap == n_ctx special case: own_prefix reserves the whole arena for SEQ_MAIN, so attn_span still returns [0, pos + 1) and the real-model bitwise tests are the refactor's gate — they did not move (cargo test --release 183 → 184 passed / 0 failed, the +1 being the new reservation test). The allocator exposes the E2-facing surface (kv_reserve_seq/kv_release_seq/kv_seq_slot/kv_own_range). (Since #228 the classic case is fill_batch_inputs's own reserve_seq(seq, n_ctx) when a group has no run, and own_prefix is test-only. Since #232, 2026-09-29: kv_own_range was deleted from that surface — a one-line forwarder with no production caller; KvCache::own_range is the production-used function, and the alloc tests drive it directly.)

E2 progress, step 2 (2026-09-17). The batch entry point landed: graph/batch.rs defines Batch { tokens, positions, seq_ids } with groups() (contiguous runs), out_rows() (one logits row per sequence for a batch; the last n_out rows for a single sequence) and check(), which refuses an interleaved batch — the CPU path could tolerate it, but CUDA stages a query tile against one window (E1b), so an interleaved batch would be right on CPU and wrong on CUDA; that is the silent divergence the project refuses. ModelDef::forward_batch (default: refuse) runs one forward over a batch, and forward_graph_cached is now its one-sequence case, which is why the classic path's tests did not move. GraphAllocator::fill_batch_inputs marks each sequence's written rows, fills seq_ids and resolves the span in one call, and kv_set_capacity lets a caller reserve before the first alloc_graph.

The flag had to change meaning. E2 makes a single sequence start at a non-zero cell (a slot's reserved run), and CUDA's causal instantiation derives nkv = positions[t] + 1 — correct only when the window starts at cell 0. So Op::Attn { multi_seq } became Op::Attn { explicit_span }, meaning “positions alone cannot bound this node”: more than one sequence or a window that does not start at 0. The model derives it from the KV reservations (kv_seq_slot(seq).start != 0), never from n_past, and it lives in CParams because it selects a kernel instantiation (two graphs per shape at most, exactly like the fusion flags). The classic path — one sequence, SEQ_MAIN, start 0 — keeps the causal instantiation, so E1b's SASS/performance argument still holds, and Metal refuses the flagged node as before (G5).

Evidence gate (a_two_sequence_batch_matches_two_single_sequence_forwards). Both sides use the same KV layout (sequence 7 at [0, la), sequence 9 at [la, la + lb)), so the only difference is one forward carrying two sequences versus two forwards carrying one each: the batched prefill of both prompts and the batched decode step are bitwise equal to the per-sequence runs on Qwen2.5-0.5B Q4_0, and the fixture is discriminating (the two sequences' argmaxes differ). cargo test --release 184 → 189 passed / 0 failed.

E2 progress, step 3 (2026-09-17): the server batches, and the measurement says not to default to it. server/batch.rs implements the design: one shared arena with one reservation per slot (BatchEngine, n_ctx_total rows split n_slots ways), a slot keeping its KV and its cached_tokens across requests so B2's prefix reuse still works, a submit path that pre-fills one request (reuse-aware admission: among idle slots it picks the one whose KV holds the longest prompt prefix), and a tick that carries every ready slot's next token in one forward_batch and then samples, commits and streams per slot. A speculative session keeps the per-slot caches and the run-to-completion loop (doc 94/97's identity contract is per request), and a panic in a batched forward fails the whole batch loudly, because the arenas may be half-written.

The measurement, end to end through the HTTP server with --n-slots 4 and max_tokens reached on every request (so both sides do equal work):

WorkloadSerialConcurrentRatio
Qwen2.5-0.5B Q4_0, prefix reuse on136 tok / 4.47 s143 tok / 5.08 s0.88x
Qwen2.5-0.5B Q4_0, MINFER_NO_PREFIX_REUSE=1136 tok / 6.36 s143 tok / 5.70 s1.12x
Qwen2.5-7B Q4_K_M, prefix reuse on32 tok / 16.50 s32 tok / 18.42 s0.49x

At the engine level (same slot layout on both sides, 4 requests × 16 tokens) the batched forward is 1.45x on the 0.5B and 1.00x on the 7B.

Verdict: the acceptance is not met, and the two causes are measurable. (1) The CPU decode kernels' nt > 1 path is not more efficient per token — the 7B's batched forward is exactly as fast as four serial forwards, so batching buys nothing there (the 0.5B, whose decode is launch/compute bound rather than weight bandwidth bound, gains 1.45x). (2) Concurrency forfeits B2's cross-request prefix reuse: each concurrent request needs its own KV home and starts cold, which on a model with a slow prefill (the 7B pays ~6 s for a 34-token chat prompt) dominates — hence 0.49x despite the neutral forward. Where decode is weight-bandwidth bound (a GPU) batching is the standard win, but nothing here can verify that (A0), so MINFER_BATCH=1 is opt-in and the default stays with the measured-better serial path, exactly as A6 was reverted on measurement.

E2 progress, step 4 (2026-09-17): batching the prefills too, and the trace that closes the diagnosis. BatchEngine::admit now places a group of arriving requests and combines their prefills into one forward when the group fits MAX_PREFILL_BATCH (per request otherwise, so the graph width does not churn; with a CUDA device the prefills stay per request, because fa_prefill's query tile must not span two sequences — E1b). MINFER_BATCH_TRACE=1 prints per-forward timings, which is how the numbers below were obtained.

On Qwen2.5-7B Q4_K_M, --n-slots 4, four identical prompts, max_tokens=4, equal work (16 tokens each side), reuse on:

serial:     7.65 s   (1 full prefill 4936 ms + 3 reused prefills ~140 ms + decode)
concurrent: 17.12 s  (batched prefill: 3 prompts, 129 tokens, 14788 ms + decode)
  • Prefill batching is exactly neutral on this CPU: 129 tokens in one forward cost 14788 ms = 3 x 4936 ms, three separate prefills to the millisecond. The prefill is compute-bound, so sharing one weight pass buys nothing here (it is the case a bandwidth-bound device would change).
  • The whole gap is B2's prefix reuse: the serial path re-feeds 1 token per request after the first (42 of 43 reused, ~140 ms), while four concurrent requests need four KV homes and each pays the full 4936 ms. Concurrency cannot reuse across slots without a cell copy — C3's operation, which needs D1.
  • Decode batching is neutral on this model: a 4-wide step costs 4.0x a single-token step (the engine measurement: 1.00x on the 7B, 1.45x on the 0.5B).

So --n-slots 4 cannot "materially exceed" the serial baseline on dgxspark: for identical prompts the serial path is ~3x cheaper on prefills that batching cannot recover, and for distinct prompts the two tie (neutral prefill batching + neutral decode batching). The remaining route to the acceptance on CPU is a nt > 1 decode kernel that actually exploits the shared weight read (the F1 family); on a bandwidth-bound device batching is the standard win, and the trace above is what a GPU re-measurement should compare.

Still to come in E2: nothing is left to build for the deliverable; what remains is the acceptance, which dgxspark cannot demonstrate (the step-4 trace above shows why, and MINFER_BATCH_TRACE=1 is the instrument for re-measuring elsewhere).

E2 progress, step 5 (2026-09-17): A7's second half — n_seqs deleted. The ticket's second acceptance clause is "the n_seqs field is either real or deleted (closes A7 if it was kept)". Step 1 had made it real (the multi_seq flag), step 3 moved that flag to CParams.explicit_span for a reason that has nothing to do with the count — a slot's reservation decides whether positions alone can bound the window — leaving n_seqs set on every path and read by none. Rather than "keep it reserved" a second time, the question was measured: a 2-sequence batch and a 1-sequence batch with the same n_tokens, n_out, gtype and explicit_span describe the same topology, and with the field in the identity they still rebuilt (uid 3 → 4, sequence_count_is_data_not_topology). The field is therefore deleted from GraphParams (and from params_match, the JSON export and every construction site); the builder's "more than one sequence involved" test now reads the batch. The test that exposed the rebuild is permanent and also asserts the reused graph's logits are bitwise-identical to a fresh single-sequence forward, so the deletion is pinned from both sides. A7 is now fully closed: both fields the A7 note called dead are gone. One incidental dead field went with it — BatchEngine's Run.finish, set to None and never read (the finish reason is a parameter of finish()), removed with the build warning it produced.

E2 progress, step 6 (2026-09-18): the acceptance is MET on the GPU. A0's "no device" verdict turned out to be an artefact of the agent sandbox, not the machine, so the re-measurement the closure note asked for was run on the GB10 — 7B Q4_K_M, --n-slots 4, four identical prompts, max_tokens=16, equal work on both sides (64 tokens each), CUDA confirmed active in both server logs:

ModeWall clockRatio
serial (default)1.314 s—
MINFER_BATCH=10.686 s1.9x

The trace explains where it comes from, and confirms the E1b constraint still holds on device: with a CUDA device the prefills stay per request (prefill_batch_ok() is false, because fa_prefill_f16kv tiles a query against one window) — 83 ms for the first 40-token prompt, 41 ms for each of the other three — so the whole win is the batched decode. That used to be an inference from the totals; MINFER_BATCH_TRACE=1 now also prints the decode step itself, and the re-run shows it directly: 15 steps of 4 sequences (plus one of 3 and one of 1) for the 64 tokens, at 6.5–6.7 ms/token, against ~19 ms/token for the serial path (1.23 s of decode for 64 tokens) — ~2.9x per decoded token, which the four full prefills the batched side pays turn into the 1.9x end to end. On CPU the same workload was 0.49x (step 3): the sign of the effect is a property of the device, exactly as the closure predicted.

Device verification sweep before merging (2026-09-18). Everything in PR #3 and PR #4 that had been left unverified for lack of a device was re-run on the GB10, and the ones that could not be asserted were made assertable:

ItemResult
E1b windowed kernels (cuda_two_sequences_do_not_cross_attend)passes on device, after fixing its hd = 2 fixture and read_host read
E1b causal vs windowed, same rowsnew gate, bitwise equal on device (closes the gap above)
E1b causal vs windowed, both KV dtypes (f32 and f16)cuda_windowed_attention_matches_causal_for_long_windows: 22/22 bitwise equal on device after fixing fa_prefill_f16kv's windowed row mask (§14 row 0) — the f16 prefill case (hd = 128, n = 34, start = 64) was the one the fix is about
Server batching, four different prompts, 7B Q4_K_M (f16 KV), --n-slots 4after the fix: batched and serial both 4/4 correct and identical (before it: batched = slot 0 right + slots 1-3 derailed, e.g. 0.1555555555555555555555)
E2 batched decode windows with a real modelbatch_order_does_not_change_a_sequences_logits passes bitwise on device
E2 engine acceptance (server_batch_matches_serial_and_is_faster, ignored)passes: 1.32x at 0.5B, all four continuations share a non-empty prefix
E2 server, 7B Q4_K_M, --n-slots 41.9x, 15 four-wide decode steps traced
Wider windowed batch--n-slots 8, 8 concurrent prompts: 11 eight-wide steps, ~1.2–1.9 ms/token, all eight answers distinct
Qwen3 (the other model E1/E2 touched) multi-sequence--n-slots 2 on Qwen3-0.6B Q8_0: 12 two-wide steps, both answers correct — the path had no test coverage on any backend before this
CUDA Graph capture under batchingno capture failure or self-disable in any run
C2's conversation path on device (adjacent: merged in PR #2, but it shares the KV cell store)context_shift_real_model_measurement passes on GPU: incremental prefills 30 then 14 tokens/turn, and the physical removal + re-rope shift takes 185 -> 14 prefill tokens with correct replies throughout
conversation_real_model_smoke (ignored, model-behaviour assertion)pre-existing red: fails identically on master + device at the same need_insert_eot assertion, so it is not from these PRs
dump_real_q4k_tensor / dump_real_q5k_tensor (ignored)pre-existing red debug dumps (they compare against llama.cpp artifacts); unrelated to these PRs and not used as gates
Full CUDA suite236 -> 237 passed / 0 failed / 5 ignored
Full CPU suite191 passed / 0 failed / 5 ignored

E2 closed (2026-09-17, maintainer decision; re-measured 2026-09-18). With the mechanism landed, A7 closed, and the throughput acceptance refuted with its two causes measured (steps 3–4), the maintainer accepted the refutation for the CPU and E2 was closed as mechanism landed, CPU throughput acceptance refuted — needs a bandwidth-bound device, with the serial path kept as the default and MINFER_BATCH=1 as the documented opt-in. Step 6 then supplied the bandwidth-bound device and the acceptance holds there (1.9x): the CPU result stands as the reason the default stays serial on CPU, and a device-aware default (enable batching automatically when a CUDA device participates) is now a supportable follow-up rather than a guess — it needs its own ticket, since it changes the server's default behaviour. This follows the A6 precedent (a ticket may close on a measured negative result, recorded so nobody re-opens it). The follow-on work the decision names is C3/D1 (cross-slot prefix reuse via a cell copy, which is what the 0.49x gap is actually made of) and C4/C5; a CPU nt > 1 decode kernel (the F1 family) is the only CPU-side route to the original throughput claim and is not part of E2. A GPU re-measurement should re-run the step-4 trace before concluding anything about batching on device — with a CUDA device the prefills stay per request (E1b), so the comparison there starts from a different baseline.

  • E4/E5 are what make a model that does not fit in VRAM runnable at all.

E1 design (written before the code, 2026-09-17). The bound a query may attend over is derived from positions today (cpu_backend.rs: let vl = pos[t] + 1, and the same derivation on device in CUDA). That is correct only while one sequence owns every written cell, and it is the one thing that makes E2 impossible: a batch holding two sequences would let each one attend to the other. E1 replaces the derivation with data:

  1. Two new inputs, not new topology. seq_ids (I32 [nt], one sequence id per query token) and attn_span (I32 [2*nt], the allowed cell range [lo, hi) per query). Both are Op::Input leaves filled per step by the allocator, exactly like positions — so GraphParams/CParams and the params-only reuse identity do not move.
  2. KvCache resolves the span. The cell store already owns owner[cell]; E1 adds seq_range(seq) -> Option<(start, len)> and attn_span(seq_ids, positions) -> Result<Vec<u32>, String>, which returns lo = start, hi = min(start + len, pos + 1). Ownership must be contiguous per sequence — that is the invariant the resolver asserts instead of silently producing a wrong bound, and the reason a range is enough.
  3. Op::Attn consumes it, and declares when positions would be wrong. Op::Attn { mode } becomes Op::Attn { mode, multi_seq }. The kernels always read the span (single code path, no "derive or read" branch); multi_seq exists so a backend that has not been ported refuses the op instead of falling back to the positions derivation — Metal (untouched, Phase G) is that backend. n_seqs > 1 is what sets the flag, giving GraphParams.n_seqs its first reader (E2 is what will make it a batch).
    • As built: the flag is named explicit_span (Op::Attn { mode, explicit_span }), and it is set from the batch and its KV reservations — "more than one sequence involved, or a window that does not start at cell 0" — never from GraphParams. E2 deleted n_seqs for exactly that reason (§8); this paragraph is the E1-era design that predicted otherwise.
  4. Bitwise for one sequence. With a single sequence, lo = 0 and hi = min(n_used, pos + 1) = pos + 1, so the CPU kernel's loop bounds, reduction length and accumulation order are unchanged — the existing real-model bitwise tests (graph_logits_match_forward_real_model, reused_cache_across_prompts_matches_a_fresh_cache, the C2 tests) are the gate, not a new tolerance class.
  5. Tests. The cell store resolves two sequences in one arena (and errors on non-contiguous ownership); an op-level two-sequence graph — one query per sequence, distinctive V rows — must reproduce the single-sequence result bitwise, which is what "no cross-attention" means; A1's op matrix gains the multi_seq cell so the Metal refusal is recorded rather than assumed.
  6. CUDA. Both attention kernels (gqa_attn_split and the non-split path) take the span; the I32 input reaches the device through the existing capture-safe positions_i32 conversion. This box had no device at the time (A0 — superseded 2026-09-18), so the CUDA half could only be compile-verified then — see the landing record below: it is deferred to E1b rather than changed blind, and the assignment gate keeps CUDA correct (single-sequence) in the meantime.

Not in E1: batch composition and continuous batching (E2 — nothing composes several sequences into one forward yet, so the model path fills seq_ids with SEQ_MAIN and stays a single-sequence caller); chunked prefill (E3); a full per-cell mask, which only a layout with holes needs (C3/D1 — a range covers every layout the engine can currently produce); Metal (G5).

E1 record (2026-09-17)

What landed. seq_ids and attn_span are IR inputs (Op::Input leaves, filled per step like positions); Op::Attn { mode } became Op::Attn { mode, multi_seq } with sources [q, kv, pos, span] — positions stays at index 2 because the Metal and CUDA arms read it there, and the span is the new fourth input. KvCache::seq_range / attn_span resolve each query's [lo, hi) from per-cell ownership (start) and its position (causal end), and GraphAllocator::fill_attn_inputs is the single call that records how far the forward writes, fills the ids and resolves the span, so the IR's ids and the kernel's window cannot drift apart. The CPU kernel now walks lo..hi instead of 0..pos[t] + 1. (E1's fill_attn_inputs was never called by production: it was deleted in #228 and its whole test corpus now drives fill_batch_inputs, the call forward_cached/forward_batch themselves make.)

The refusal has one definition. Backend::supports_attn_span (default false) plus graph::backend_takes decide assignment; GraphAllocator::supports and A1's op matrix both call it, so a backend that still derives its bound from positions is never handed a multi_seq node — and the op matrix gained an explicit asymmetric row for it.

Bitwise for one sequence, as specified. With one sequence lo = 0 and hi = min(n_used, pos + 1) = pos + 1, which is exactly the old derivation, so the loop bounds, reduction length and accumulation order are unchanged. Unit suite 179 → 183 passed / 0 failed; the real-model bitwise tests (graph_logits_match_forward_real_model, prefix_reuse_matches_a_full_prefill, reused_cache_across_prompts_matches_a_fresh_cache, the C2 tests) never moved, and the pre-E1 vs E1 binaries generate byte-identical greedy text on Qwen2.5-0.5B Q4_0, Qwen3-0.6B Q8_0, Qwen2.5-7B Q4_K_M and Qwen2.5-14B Q4_K_M (only timing lines differ).

No cross-attention (two_sequences_do_not_cross_attend). One arena, two sequences: sequence 0 owns row 0, sequence 1 owns row 2, queries at positions 0 and 2 with spans [0, 1) and [2, 3). The K/V values are chosen so a leak changes the answer: query 1 scores 1.0 against sequence 0's key, so a window that wrongly started at 0 would return [0.73, 0.27] instead of sequence 1's V = [0, 1]. The test asserts the exact window and the output, and the resolver has its own store-level tests (two_sequences_resolve_to_disjoint_windows, plus Err on non-contiguous ownership and on an empty window).

The CUDA half is deferred, and why. The ticket says "CPU + CUDA". CUDA's attention kernels still compute positions[t] + 1 (six of them: gqa_attn_f32_f16kv, gqa_attn_f32, the split partial/combine pairs, the batched variants and the flash-attention prefill path at cuda_kernels.cu:4194). A window with lo > 0 changes what the split-K chunking covers, so the port is not mechanical; and dgxspark had no device at the time (A0 — superseded 2026-09-18: E1b is now device-verified), which made it the one class of change that could not be verified then — a mistake would silently corrupt single-sequence GPU output that is known-good today. So: CUDA keeps its existing behavior, supports_attn_span() stays false for it (the trait default), the assignment gate refuses it a multi_seq node, and the port is ticket E1b. E1's acceptance is therefore met on CPU and not met on CUDA — recorded here rather than claimed, and the deferral was agreed with the maintainer on 2026-09-17 rather than taken unilaterally.

Not in E1: batch composition and continuous batching (E2 — nothing composes several sequences into one forward yet, so the model path fills seq_ids with SEQ_MAIN and stays a single-sequence caller); chunked prefill (E3); a full per-cell mask, which only a layout with holes needs (C3/D1 — a range covers every layout the engine can currently produce); Metal (G5); the CUDA port (E1b).

E1b record (2026-09-17) — CUDA's windowed attention

Shape. Every attention kernel is now template <bool CAUSAL>. CAUSAL (the existing behaviour) keeps nkv = positions[t] + 1 and a row base of 0; the windowed instantiation reads the [lo, hi) pair (bound[t], bound[nt + t]) and indexes rows as row0 + j. Row arithmetic is the only difference — loop trip counts, split chunking, merge order and the per-row op order are untouched, which is what preserves the doc-94 bitwise identity between the verify batch and sequential decode. Kernels: attn_split_1w_body, attn_split_h4w_body, gqa_attn_split_partial, _hybrid, _bt, gqa_attn_f32_f16kv, gqa_attn_f32 and fa_prefill_f16kv (whose tile-wide extent is max(hi)/min(lo) so a query tile must not span two sequences — E2's composition keeps a sequence contiguous). Op::Attn { multi_seq } picks the pointer (positions vs span) and the instantiation host-side, so a causal node never touches the span input; supports_attn_span() is now true for CUDA, and the op matrix's asymmetric row flips accordingly.

The performance constraint, with evidence. The requirement was "do not affect existing CUDA performance", and there was no device to measure it on (cuInit → 304, re-confirmed 2026-09-17: no seccomp, no container, the 580.178.04 module loaded, nodes present — all of it an artefact of the agent sandbox, see A0; the wall-clock half of this section was added 2026-09-18). So the evidence at the time was the generated code: both revisions compiled with the project's own nvcc flags (-O3, -gencode arch=compute_121,code=sm_121) and cuobjdump -sass compared per kernel:

  • instruction counts identical for every causal instantiation — gqa_attn_split_partial 448/416, _hybrid 1664, _bt 464/416, gqa_attn_f32_f16kv 1952, gqa_attn_f32 1520, fa_prefill_f16kv 2104 (the pre-E1b numbers);
  • opcode histograms identical for gqa_attn_split_partial[float] and gqa_attn_split_partial_hybrid[half]; the others differ only by LDG.E → LDG.E.CONSTANT (the bound parameter is const __restrict__, so the loads take the read-only path — a caching upgrade, not added work) and by MOV/CS2R/NOP register-allocation substitutions of equal count.

No kernel gained an instruction, none lost one. What is not verified is wall-clock time on a GPU: the windowed path is unreachable until E2 composes batches, and the first GPU session should re-check both the numbers and the timing (recorded in the status line).

Tests. cuda_two_sequences_do_not_cross_attend mirrors the CPU test at the backend level (device-gated: it compiles here and skips without a device), and the existing CUDA attention tests — including cuda_verify_attention_nt_invariance, the doc-94 identity — call the causal instantiation as before.

Device verification (2026-09-18) — the first execution on hardware. With the sandbox corrected (A0), E1b's windowed path was run on the GB10 it was written for, and it is correct and free:

  • The E1b test could not have passed anywhere. As written it used hd = 2, which the CUDA attention kernels reject (attention head dim 2 outside the kernel's supported range (multiple of 4, 1..=128)), and it read its result with the trait's read_host, which is None on CUDA by design (device memory cannot be borrowed). Both are fixed — hd = 4, copy_to_host — and the test now passes on the device: token 0 returns V(0), token 1 returns V(2), so a window that started at 0 (the pre-E1b derivation) would fail it. This is the window arithmetic itself, checked row by row.
  • The claim "no CPU behaviour change" holds: E1b's diff touches only src/cuda.rs, src/cuda_kernels.cu, src/graph/cuda_backend.rs and the op_matrix support-table test, and the CPU suite is green (191 passed / 0 failed / 5 ignored).
  • The claim "no causal-path performance change" now has wall-clock evidence next to the SASS instruction counts: pre-E1b (82f5109) vs this branch, 7B Q4_K_M, same prompt, greedy, on GB10 — prefill 529.8 → 530.3 tok/s, decode 51.2 → 51.1 tok/s, identical output text. The windowed instantiations cost nothing when they are not selected, exactly as the opcode histograms said.
  • The multi-sequence path is right with a real model. New gate batch_order_does_not_change_a_sequences_logits: two sequences in one batch (batch shape, KV layout, history and graph all held fixed) must reproduce each sequence's logits bitwise when their order in the batch is swapped. It passes on the device — and it is only bitwise because the window follows the sequence and its reservation, never the row index.
  • Still open on the timing side (2026-09-18): the bitwise equivalence is now asserted, but the windowed instantiation's cost at the same width was never measured — the GPU numbers in this record compare a 4-wide windowed step against a 1-wide causal step (6.5 vs ~19 ms/token), which mixes the width and the instantiation. A same-width A/B (windowed vs causal over identical rows, timed) is the remaining question, and it needs a device.
  • The causal-vs-windowed gap is closed (2026-09-18). The record previously left this open: the windowed test exercises the row arithmetic but never the causal pointer against it. cuda_causal_and_windowed_agree_on_the_same_rows now runs the same queries, K/V and rows through both instantiations — a window that starts at cell 0, so positions[t] + 1 and the explicit span name the same rows — and asserts the outputs are bitwise equal on the device. It passes on GB10, so the equivalence no longer rests on the SASS identity alone.
  • The f16 half of the windowed path was untested until 2026-09-19. That gate forced cb.kv_f16 = false (f32 KV), while the production default is f16 whenever n_layers * n_kv_embd >= 8192 (cuda::set_kv_cache_type): the 7B (14336) runs f16, the 0.5B (3072) f32. The sweep now loops over both dtypes and at once caught a real fault in fa_prefill_f16kv's windowed row mask (bound[t] is the window's lo, not its hi) — plan §14 row 0 has the root cause, the fix and the end-to-end server evidence. With the fix, 22/22 cases (11 shapes x 2 dtypes) are bitwise equal on device, including the hd = 128, n = 34, start = 64 prefill case the fault lived in.

E2 follow-up record (#121, 2026-09-25) — a rejected job is answered, not dropped

The defect. F8's minfer_jobs_dropped_total (#51) exposed it (see the F8 record's "defect found while gating this"): serve_loop calls admit on every pass, busy or not, and admit consumed the Job by value — so a job it could not place (no idle slot) had its mpsc::Sender<StreamEvent> dropped without an event. The handler read the closed channel as a completed answer: non-streaming collect_response returned Ok(("", "stop", 0)) → HTTP 200 with empty content; streaming sent an empty SSE stream followed by [DONE]. A client could not tell a dropped request from a model that produced nothing, and the request was silently lost.

Reproduced on dgxspark (CPU, cached 0.5B Q4_0, MINFER_BATCH=1 --n-slots 1, two concurrent max_tokens=200 requests, B sent ~0.4 s after A, scripts-free python client — see the PR):

beforeafter
A200, 976 chars, finish=length, 200 completion tokensunchanged
B, non-streaming200, 0 chars, finish=stop, completion_tokens=0503, {"error":{"code":503,"message":"no idle slot","type":"unavailable_error"}}
B, streaming200, SSE = role chunk + [DONE], no error frame200, SSE = role chunk + data: {"error":{"code":503,…}} + [DONE]
minfer_jobs_dropped_total11

Decision: reject loudly; queueing is a feature (#150). The batched worker admits on every pass instead of holding a backlog, so --n-slots N bounds concurrent requests and a request arriving while all N are busy is refused. Queueing it would turn admit's contract inside out (the caller would own a retry loop and a bound) for a behaviour nobody asked for, and the rejection is already the measured design: queue_depth is a near-zero transient, minfer_jobs_dropped_total is the honest signal, and this ticket's whole point is that the signal must reach the client. The serial path (MINFER_BATCH=0) already queues — its one worker pulls a job at a time from the same channel — so the two paths document their difference instead of one of them silently losing a request. #150 tracks making the batched path queue with a bound.

What landed. src/server/batch.rs: reject(job, e) sends StreamEvent::Err through the job's own sender before the sender drops and returns the error so the caller can still count the rejection; the three paths that give up on a job go through it — admit's no-idle-slot branch, admit's failed-group-prefill branch (the install loop is its last, infallible step, so nothing was installed) and submit_on's errors (invalid slot, busy slot, prompt over the slot context, failed prefill forward). serve_loop still counts the returned Errs, so F8's counter keeps its meaning. src/server/types.rs: ApiError gains Clone (the error is sent and returned); unavailable was already the 503 constructor (status: 503, error_type: "unavailable_error"), rendered by server::error_response (non-streaming) and server::to_event (streaming). A misplaced submit_on doc comment that sat above set_prefill_chunk was moved back to its function in passing.

Audit of every sender-drop path. worker_loop_serial's no idle slot and mirostat-refusal branches already sent their refusals (git log -S 'no idle slot' -- src/server/chat.rs puts both at Phase 2–6, before F8), and the no idle slot branch is unreachable in practice anyway: the serial loop pulls one job at a time, so every slot is idle when it looks. BatchEngine::fail sends; finish sends Finish; run_job_isolated turns a job panic into a 500. The handler's own job_tx.send failure and the drain refusal answer 503 directly, before a Job exists. Two residual shapes were found and not changed: (a) tick's failed forward returns Err before any per-run fail, so the affected runs keep needs_forward and serve_loop retries the same batch forever (a stuck sender, not a dropped one) — filed as #151; (b) a panic outside guarded_forward_batch in serve_loop would drop pending's senders, but every model-calling path is guarded.

Gates. a_rejected_job_answers_503_and_an_sse_error_frame (CI, no model) drives the real reject + collect_response + error_response + stream_response and asserts the 503 and the SSE error frame. a_job_rejected_for_want_of_a_slot_is_answered_with_503 (#[ignore], real model, one slot, two jobs queued before the loop starts) asserts one request is served (Finish, tokens > 0) and the other gets exactly one StreamEvent::Err (status 503, no idle slot) and nothing else, with jobs_dropped_total exactly 1. The old F8 round-2 gate kept both receivers dropped, so it counted the rejection but could not see the empty 200 — the new gate reads them.

Mutation checks. (a) reject replaced by drop(job) (the pre-fix silent drop): the CI gate fails with "a rejected job must not read as a completed empty answer: ("", "stop", 0)", the ignored gate with "a job's channel closed with 0 event(s) — the #121 silent drop is back (the handler would answer HTTP 200 with empty content)". (b) jobs_dropped_total.fetch_add(dropped, …) → fetch_add(0, …): the ignored gate fails "exactly one job could not be placed on the one slot: left: 0, right: 1"; the CI gate still passes, because the counter is not its property. Both reverted; sha256sum src/server/batch.rs equals the pre-mutation value, byte-identical.

Verification (2026-09-25).

CommandResult
cargo test --release (CPU)439 / 0 / 29 unit + 10 / 0 / 6 integration
cargo test --release --bin minfer -- --ignored --test-threads=1 (CPU)29 / 0 (baseline 28 / 0)
cargo test --release --features cuda -- --test-threads=1 (GB10 sm_121)502 / 0 / 32 unit + 10 / 0 / 6 integration (baseline 501 / 0 / 31)
… --ignored --test-threads=1, 0.5B f32 KV32 / 0 (baseline 31 / 0)
… MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf --ignored --test-threads=132 / 0 (baseline 31 / 0)
rustup run stable rustfmt --edition 2021 --check on the changed .rsclean (rustfmt 1.9.0-stable)
python3 scripts/check_docs_links.py940 relative links / 184 files (unchanged)

The counts move by exactly this ticket's two gates (+1 unit pass, +1 #[ignore]d real-model gate); the F8 serve-loop gate's own numbers (peak running 2, 1 dropped) are unchanged.

Docs. USAGE.md § Slot saturation (the decision, both status shapes, the counter); FEATURES.md and OPENAI-CHAT-API-PLAN.md § Slot Lifecycle state the batched/serial split (the plan's old "defer the task in the request queue" line was never implemented by the batched path and is now corrected); AGENTS.md's server/batch.rs bullet carries the contract and the new counts. ARCHITECTURE-ROADMAP.md is untouched: no roadmap gap closes here.

E2 follow-up record (#151, 2026-09-25) — a failed decode step answers its batch

The defect. The mirror of #121: there the sender was dropped without an event; here it was never dropped and never answered. BatchEngine::tick built one batch from every run with needs_forward set and called guarded_forward_batch(...)? — the ? returned before the per-row loop that takes each needs_forward and before advance/fail, and serve_loop only logged [server] step failed: …. So every run kept needs_forward = Some(tok), the next pass rebuilt the same batch and retried the same forward. A deterministic failure (a kernel-invariant violation, an E4 activation-budget refusal, an ensure_kv format mismatch, a panic caught by guarded_forward_batch) failed forever at 100% CPU; no client was ever told, in_flight stayed ≥ 1, so a graceful drain ran out its whole MINFER_DRAIN_MS deadline.

Reproduced on dgxspark (CPU, no model — a test double whose forward_batch panics, driven through the real guarded_forward_batch and the real serve_loop, with the handler's InFlight guard held until the stream ends and a 500 ms drain window):

beforeafter
serve_loop returnsnever (3 s window)yes
decode forwards attempted955 907 in 3 s (≈319 k/s — the spin)2 (1 prefill + 1 decode)
StreamEvent::Err to the client01 (status 500)
Finish00
in_flight after the drain deadline1 (stuck)0
running after the step10

Decision: answer every row of the failed batch, once, and do not retry. The forward is one weight pass, so the failure belongs to every row it carried — not to a slot, and not to a run that was not in it. tick already builds rows: Vec<usize> (the slot index per batch row) while collecting the pending tokens; that list is the membership test, so a run whose needs_forward was unset (already sampled) or that did not fit MAX_BATCH is untouched. Each affected run goes through the existing fail, which is exactly the shape #121's reject established per job: one StreamEvent::Err through the run's own sender, the run taken (slot freed) and cached_tokens cleared — a forward that failed part-way may have written rows the mirror does not describe, and reusing that prefix is the one thing a failure must not lead to. Clearing the slot is also what stops the retry: with needs_forward gone the next tick builds a different batch (or none). No retry is a choice, not an oversight: every reachable class here is deterministic and request-fatal, and a caught panic may have left the shared arena half-written, so trying again can only spin (the defect) or read a corrupted arena. The failure is answered for these runs only — nothing is latched on the server, so a later request builds a fresh batch. A transient failure would therefore still cost these particular requests their answer: that is the honest trade against the old infinite retry, and a genuinely transient class (if one ever appears) belongs in the backend as a retry around the device op. The error attribution is ApiError::server (500 server_error) — the step failed; 503 unavailable_error stays the saturation refusal of #121/#150.

What landed. src/server/batch.rs: tick's forward is a match; its Err arm calls fail_batch(&rows, &e) and then still returns Err(e), so serve_loop keeps its "step failed" signal while every affected run has been answered. fail_batch (rows empty ⇒ no-op) loops over rows calling fail, then prints one line naming how many runs it answered and that their slots were released. No behavior change on the success path, and no new test-only seam in the engine: the gate's injection is a ModelDef double, so the real guarded_forward_batch → ApiError::server path is what runs.

Gates. a_failed_decode_forward_answers_a_single_run_and_releases_its_slot and a_failed_decode_forward_answers_every_row_in_the_batch (both plain #[test], CI — no model on disk): the double's forward_batch panics and counts attempts; each gate asserts one StreamEvent::Err (500, server_error, non-empty message) and no Finish/Text per affected run, the sender closed and the slot free, cached_tokens cleared, and — after a second tick — that the attempt count did not grow (the bounded no-retry assertion, so a broken build fails instead of hanging). The multi-slot gate also installs a live run with needs_forward = None and asserts it survives with its prefix intact and its channel open. The real-model batched-vs-serial gates keep their counts (the new gates need no model).

Mutation checks. (a) the pre-fix early return (no fail_batch): both gates fail with "exactly one error, not zero and not a retry: left: 0, right: 1" (and the multi-slot one with "slot 0: exactly one error"). (b) answer the run but do not take it / clear it (the slot not released): both fail with "the run's sender is dropped: the slot is released" / "slot 0: released". Both reverted; sha256sum src/server/batch.rs equals the pre-mutation value, byte-identical (725827f6…).

Verification (2026-09-25).

CommandResult
cargo test --release (CPU)440 / 0 / 29 unit + 10 / 0 / 6 integration (baseline 438 / 0 / 29)
cargo test --release --bin minfer -- --ignored --test-threads=1 (CPU)29 / 0 (unchanged)
cargo test --release --features cuda -- --test-threads=1 (GB10 sm_121)503 / 0 / 32 unit + 10 / 0 / 6 integration (baseline 501 / 0 / 32)
… --ignored --test-threads=1, 0.5B f32 KV32 / 0 (unchanged)
… MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf --ignored --test-threads=132 / 0 (unchanged)
rustup run stable rustfmt --edition 2021 --check on the changed .rsclean (rustfmt 1.9.0-stable)
python3 scripts/check_docs_links.py940 relative links / 184 files (unchanged)

The CPU/CUDA unit counts move by exactly this ticket's two gates (+2). The #[ignore]d set does not move (the new gates are CI-covered). The baseline at ee4d0e1 measured 438 / 0 / 29 unit on CPU, one below the #121 record's literal 439 — the #94 count-drift class, recorded here rather than rewritten into another ticket's dated record.

Docs. AGENTS.md's server/batch.rs bullet carries the mirror case next to #121's; OPENAI-CHAT-API-PLAN.md § Slot Lifecycle step 5 and the Error Types table state the failed step (one 500 per affected run, no retry) and its note no longer claims the batched path defers; FEATURES.md's serve bullet names it. ARCHITECTURE-ROADMAP.md is untouched: no roadmap gap closes here — this is robustness inside an existing row (item 3, Batch composition + continuous batching), not a new capability. #154 (the wall-clock flake of server_batch_matches_serial_and_is_faster under the parallel harness) is left as-is on purpose; the ignored set is run serially.

Test-infrastructure record (#154, 2026-09-25) — the batching gate's throughput verdict is a median over interleaved rounds

The defect. server::batch::tests::server_batch_matches_serial_and_is_faster (an #[ignore]d real-model gate, src/server/batch.rs) measured the two whole workloads once, sequentially (batched then serial) and asserted the wall-clock relation t_serial > t_batch. Under the parallel --ignored harness the first-measured phase absorbs the start-up wave, so the verdict was a property of the load, not of the code. Measured at e1ac17f on dgxspark (CPU build, 0.5B q4_0, four requests, max_tokens = 16):

runbatchedserialratio
parallel --ignored (first)21.20s9.95s0.47x (fail; the set was 28 passed / 1 failed in 34.36s)
serial, same binary0.72s1.07s1.50x (pass; gate alone 2.68s)

The correctness half was never at fault — batched vs serial output (and the staggered-admission arm) already compared byte-for-byte on CPU and passed in both runs.

The fix (the #123 shape). The timed rounds are now interleaved — run_batched, then run_serial, repeated rounds times — so each ratio is a matched pair measured next to each other on the same machine state, and the assertion is on the median of the per-round serial/batched ratios (the location estimate that tolerates up to rounds / 2 disturbed rounds). This is the shape #123 gave cuda_map_window_costs_no_more_than_the_span_it_replaces. median and a factored assert_replies_match helper replace the two inline comparison loops; every per-round ratio, both time medians and the verdict median are printed, so a loaded box's result is auditable.

Statistic, threshold, cost. The threshold is unchanged at 1.0x ("must not be slower"); the statistic is what changed, and no margin was invented. rounds = 7 (env MINFER_BATCH_TEST_ROUNDS) is the smallest odd count whose median tolerates three disturbed rounds. The timed rounds use timing_tokens = 8 (env MINFER_BATCH_TEST_TIMING_TOKENS) while the correctness comparison stays a separate full-length (max_tokens = 16) pair, so a shorter timing workload cannot weaken the byte-equality. The gate alone is 8.92s on an idle box (2.68s before) and the whole parallel set ~41-44s (34.36s before). The ticket's cost note ("one pair is ~27s") describes the loaded harness, where a full-length pair reached ~31s; the idle pair is ~1.8s, which is what made 7 full-length rounds affordable-ish and 7 shorter ones clearly so.

Verification (2026-09-25, CPU build, dgxspark; --bin minfer for the gate set).

CommandResult
parallel --ignored, before28 passed / 1 failed — 21.20s batched vs 9.95s serial = 0.47x
serial --ignored, before29 passed / 0 failed
parallel --ignored, after run 129 / 0; per-round ratios [1.335, 1.804, 1.776, 1.434, 1.420, 1.416, 1.368]; median 1.420x
parallel --ignored, after run 229 / 0; ratios [1.350, 0.923, 1.640, 1.342, 1.404, 1.398, 1.379]; median 1.379x
parallel --ignored, after run 329 / 0; ratios [1.647, 1.360, 1.481, 1.472, 1.468, 1.396, 1.348]; median 1.468x
scripts/real_model_gates.sh (new default → parallel)29 / 0; ratios [1.204, 1.787, 1.837, 1.805, 1.709, 1.824, 1.818]; median 1.805x
parallel --ignored + 16 CPU spinnersthe set 28 / 1 — this gate passed at median 2.159x (ratios [1.099, 2.046, 2.142, 2.159, 2.182, 2.274, 2.286]); the one failure is a different, deadline-based gate, filed as #158
serial --ignored, after29 passed / 0 failed
mutation: the timed batched arm runs timing_tokens * 2fails at median 0.805x (ratios 0.784–0.842, every round < 1.0); reverted byte-identically, sha256sum src/server/batch.rs = 15b717e1…
cargo test --release (CPU)440 / 0 / 29 unit + 10 / 0 / 6 integration (unchanged — no new test)
cargo test --release --features cuda -- --test-threads=1 (GB10 sm_121)503 / 0 / 32 unit + 10 / 0 / 6 integration (unchanged)
… --ignored --test-threads=1, 0.5B f32 KV32 / 0
… MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf --ignored --test-threads=132 / 0
rustup run stable rustfmt --edition 2021 --check on the changed .rsclean (rustfmt 1.9.0-stable; the pinned 1.97.1 toolchain has no rustfmt component)
python3 scripts/check_docs_links.py940 relative links / 184 files (unchanged)

The wrapper default. scripts/real_model_gates.sh used to default to --test-threads=1 on every build, because the device needs it (CudaState is process-wide, issue #64). With #154 landed, the only CPU reason left is gone too, so the wrapper now defaults to serial when FEATURES includes cuda and to the parallel form otherwise (PARALLEL=1/0 overrides, with a warning if parallel is forced on a CUDA build). The parallel CPU form is where the #154 gate's robustness is actually exercised, so it is the default CPU command rather than an opt-in — the fix would otherwise be invisible to whoever runs the wrapper.

Mutation check. Doubling the timed batched arm's work (timing_tokens * 2, the natural "removes the batching win" injection for this statistic) makes the gate fail at median 0.805x with all seven rounds below 1.0 — the median is not blind. It was reverted byte-identically (sha256sum -c on 15b717e1…). The three plain parallel runs double as the robustness check from the other side: run 2's round 0 measured 0.923x (a load spike landing on one matched pair), and the median still returned 1.379x.

A gate that passes for the wrong reason — found, filed, not fixed here. The extreme-load run (16 extra CPU spinners) failed published_metrics_move_as_requests_are_served at "the long request finished inside the deadline" — it asserts a 180s wall-clock deadline on a 64-token generation, which is the same class of load-dependence as #154 but not the same statistic (an absolute deadline, not a ratio), and it is green in the plain parallel runs. Filed as #158 rather than widened here.

Docs. AGENTS.md's real-model-gates bullet no longer says the parallel CPU set is 28 / 1 or that the wrapper defaults to serial; docs/BUILD.md § Tests states the per-feature default and the new counts; the wrapper's own comments carry the same story. ARCHITECTURE-ROADMAP.md is untouched: this is robustness inside the existing test-infrastructure/batching rows (item 3, Batch composition + continuous batching), not a new capability — the same reason the #151 record gives. The stale "E2's acceptance … take materially less wall time" paragraph that sat on the serve_on helper (it described this gate, not the helper) moved onto the gate, where the new statistic is stated.

Test-infrastructure record (#158, 2026-09-25) — the F8 metrics gate bounds work, not wall-clock seconds

The defect. server::batch::tests::published_metrics_move_as_requests_are_served (an #[ignore]d real-model gate, src/server/batch.rs) drove its two requests with absolute wall-clock deadlines — a 120s loop for the warm 8-token request and a 180s loop for the long one — and asserted !engine.busy() when the loop ended. The deadline is not a property of the code: on this 20-core box, running the whole parallel --ignored set with 16 extra CPU spinners made the long request legitimately exceed 180s, and the gate panicked at src/server/batch.rs:3609 with "the long request finished inside the deadline". Measured before the fix: the set was 28 passed / 1 failed in 435.11s (460s wall with the spinners), green without them (29 / 0). This is the same class as #154 — a verdict that is a property of the load, not of the code — but an absolute deadline no statistic can absorb, unlike #154's ratio. #123 is the third member of the class (the CUDA map-window gate, which needed a justified margin rather than no margin).

The fix: a bound on progress. BatchEngine gained a monotone work_units counter, advanced once per row a decode forward wrote and once per token advance committed. The invariant that makes it a progress signal: a tick that leaves the engine busy must have moved it — if no forward ran, then some slot's advance returned Continue, and Continue commits exactly one token (every other advance outcome ends the run and takes it). The gate's two loops became two calls to a shared drive_by_work(engine, model, tok, rx, budget, what) helper that steps until idle, asserts the counter moved on every step that left the engine busy, and caps the step count. No wall-clock bound remains in the gate, so there is no bare literal to justify and no env knob to add: a slow box runs the same steps, only for longer, while a wedged engine trips the work assertion on the stalling step.

The step budget. step_budget(prompt, max_tokens) = 4 * (prompt + max_tokens + 8) (STEP_BUDGET_MARGIN = 4, STEP_BUDGET_SLACK = 8): one decode forward and one sample per answer token plus one prefill forward per chunk, times a deliberately loose margin. It is loose because the bound must catch an engine that cannot terminate, never one that is merely slow — a false negative hangs the suite, a false positive is the flaky gate this ticket removes. Measured on the 0.5B q4_0 (dgxspark, 2026-09-25): the warm request (8-token prompt, max_tokens = 4) took 4 steps against a budget of 80 (20x); the long one (120-token prompt, max_tokens = 64) took 64 against 768 (12x). The gate prints both counts on every run.

Verdict, before and after (CPU build, dgxspark, 16 extra CPU spinners on a 20-core machine).

CommandBeforeAfter
parallel --ignored + 16 CPU spinners28 / 1 in 435.11s — the gate panicked at batch.rs:3609 "the long request finished inside the deadline"29 / 0 in 424.17s (438s wall)
plain parallel --ignored29 / 029 / 0 in 36.60s
serial --ignored29 / 029 / 0 in 38.47s

Mutation check. Making tick return Ok(()) without forwarding or committing (the natural "wedge the engine" injection) makes the gate fail on step 1 of the warm request — "the warm request: the engine is wedged — step 1 left it busy without advancing the work counter (still 0)" — in 0.18s, where the old deadline would have waited the full 120s to report the same thing. Reverted byte-identically: sha256sum src/server/batch.rs back to 99a608e3….

The audit. Every real-model gate shape was checked for an absolute deadline used as its only failure signal (grep for Instant::now() + Duration::from_secs across src/ and tests/). The two #158 deadlines were the only ones; what remains is genuinely different and filed as #160:

  • src/server/batch.rs:2925 (serve_loop_publishes_the_queue_and_running_depth) — the feeder's 120s poll terminator is a redundant backstop: the verdict is the downstream peak_running > 0 and queue-arithmetic assertions, and the worker runs on the main thread, so the deadline cannot rescue a real wedge.
  • tests/conversation_cli.rs:41 — run_cli's 1800s child-process kill for the #[ignore]d real-model CLI sessions: a cross-process hang guard (a child exposes no in-process progress counter, and 1800s is ~10-100x the legitimate scalar-CPU runtime).
  • tests/backend_registry_cli.rs:27 — run_cli's fixed 60s kill: a cross-process hang guard, and not a real-model run (every case points at a nonexistent model path).
  • src/server/mod.rs:827 — a CI unit test's 10s ceiling on a 50ms bounded drain (200x margin), not a real-model gate.
  • Five while engine.busy() stepper loops in the other #[ignore]d server gates have no bound at all (a wedge hangs the suite rather than false-failing it); #158's helpers make hardening them mechanical, and #160 tracks it.

Verification (2026-09-25, dgxspark; --bin minfer for the gate set).

CommandResult
parallel --ignored + 16 CPU spinners, before28 / 1 in 435.11s (the gate panicked at batch.rs:3609)
parallel --ignored + 16 CPU spinners, after29 / 0 in 424.17s
plain parallel --ignored, after29 / 0 in 36.60s
serial --ignored (PARALLEL=0), after29 / 0 in 38.47s
mutation: tick returns without advancingfails on step 1 ("the engine is wedged … (still 0)"), 0.18s; reverted byte-identically (sha256sum = 99a608e3…)
cargo test --release (CPU)440 / 0 / 29 unit + 10 / 0 / 6 integration (unchanged — no new test)
cargo test --release --features cuda -- --test-threads=1 (GB10 sm_121)503 / 0 / 32 unit + 10 / 0 / 6 integration (unchanged)
… --ignored --test-threads=1, 0.5B f32 KV32 / 0
… MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf --ignored --test-threads=132 / 0
rustup run stable rustfmt --edition 2021 --check src/server/batch.rsclean (rustfmt 1.9.0-stable; the pinned 1.97.1 toolchain has no rustfmt component)
python3 scripts/check_docs_links.py940 relative links / 184 files (unchanged)

Docs. AGENTS.md's real-model-gates bullet and docs/BUILD.md § Tests carry #158 next to #154; this record is the per-ticket entry, cross-referencing #154 and #123 as the three members of the load-dependent-verdict class. ARCHITECTURE-ROADMAP.md is untouched: this is robustness inside the existing test-infrastructure/batching rows (item 3, Batch composition + continuous batching), not a new capability — the same reason the #154 and #151 records give.

Test-infrastructure record (#160, 2026-09-27) — the remaining server-gate steppers bound work, not wall-clock seconds

The defect. #158's audit left two shapes standing. The one this ticket was filed for: five #[ignore]d server gates drove the engine with while engine.busy() { engine.tick(…) } and neither a progress assertion nor a step budget, so a wedge inside BatchEngine::tick hung the suite forever instead of failing it (#158's own F8 metrics gate kept its work bound). The other shape: three absolute wall-clock bounds whose verdict does not depend on the clock. Both are review findings, not reproduced failures — the set was green before and after; the evidence below is injected wedges, not an observed hang.

What now bounds each site. Every stepper goes through one shared WorkBound (src/server/batch.rs), the single implementation of the invariant #158 introduced: a tick that leaves the engine busy must have advanced BatchEngine::work_units (if no forward ran, some slot's advance returned Continue, and Continue commits exactly one token), and the step count must stay inside step_budget(prompt, max_tokens). drive_by_work is now a thin wrapper over it.

site (gate)beforenow
run_batched (server_batch_matches_serial_and_is_faster, a_long_request_may_use_the_whole_arena)while !queue.is_empty() || engine.busy()WorkBound with step_budget(Σ prompts, n · cap)
run_serial (same two gates)while done.is_none()one WorkBound per request
serve_on (a_prefix_copied_from_another_slot_answers_identically, a_store_inside_a_shared_prefix_takes_a_private_row)while engine.busy()WorkBound
a_slot_snapshot_resumes_the_context_without_re_prefillingthree while X.busy() loopsthree WorkBounds (cold / resumed / delta), each with its own budget
a_chunked_prefill_answers_like_an_unchunked_onewhile engine.busy() in both armsWorkBound, the arm's own name in the message
a_repeated_chunked_prefill_stops_rebuildingwhile engine.busy() in both callsWorkBound
a_long_prefill_keeps_another_slot_decodingfor _ in 0..3 { tick } (already finite)the same three steps through WorkBound, so a wedge fails on the step that wedged it
published_metrics_move_as_requests_are_served (#158)drive_by_workunchanged behaviour, now on the shared WorkBound

step_cap folds an unbounded request (max_tokens < 0, which ends on its context bound) into the budget. No wall-clock number was added anywhere as a failure signal, and no bar was widened.

The wall-clock bounds that remain are named backstops, not verdicts.

  • serve_loop_publishes_the_queue_and_running_depth's feeder terminator is now FEEDER_POLL_BACKSTOP (120s) with its reasoning at the call site: the verdict is the downstream peak_running > 0 / queue arithmetic, and the worker runs on the test's own thread, so a real wedge hangs that join regardless — which is exactly why #158 left it and #160 classifies it a redundant backstop. peak_running is observed within milliseconds, so it is not load-sensitive.
  • tests/conversation_cli.rs's run_cli child-process ceiling (1800s for the #[ignore]d real-model sessions, 30s for the no-model cases) and tests/backend_registry_cli.rs's 60s are cross-process hang guards — a child exposes no in-process progress counter. Both are now env-overridable through MINFER_CLI_WATCHDOG_SECS (unset/unparsable/zero keeps the caller's number), with the reasoning recorded at each run_cli.
  • src/server/mod.rs's elapsed < Duration::from_secs(10) is untouched: a CI unit test's 200x ceiling over a 50ms bounded drain, not a real-model gate.

The seam (rule 3). MINFER_TEST_TICK (src/server/batch.rs::tick_seam, read once per process through a OnceLock, Off when unset or on any unrecognised value) injects the two arms of a bounded drive into the real tick: wedge returns before forwarding or committing anything (the work counter freezes), spin advances work_units without ever completing a run (the counter keeps moving along a path that cannot terminate). It is an environment switch — the mutation runs need no source revert, so git diff stayed clean throughout.

Mutation evidence — the wedge arm (rule 3). Command shape (CPU build, aarch64, dgxspark, 2026-09-27): MINFER_TEST_TICK=wedge cargo test --release --bin minfer -- --ignored --exact <gate> --nocapture. Every hardened gate fails on step 1 with "the engine is wedged — step 1 left it busy without advancing the work counter (still 0)" (src/server/batch.rs:3756):

gatewall clockwhat the message names
published_metrics_move_as_requests_are_served (#158 precedent)0.30sthe warm request
a_prefix_copied_from_another_slot_answers_identically0.21sthe slot-scoped request
a_store_inside_a_shared_prefix_takes_a_private_row0.29sthe slot-scoped request
a_long_request_may_use_the_whole_arena1.68sthe batched run
server_batch_matches_serial_and_is_faster0.30sthe batched run
a_slot_snapshot_resumes_the_context_without_re_prefilling0.22sthe cold run
a_chunked_prefill_answers_like_an_unchunked_one0.69sthe unchunked prefill
a_repeated_chunked_prefill_stops_rebuilding0.71sthe repeated chunked prefill
a_long_prefill_keeps_another_slot_decoding0.45sthe short run

The listed time is the whole process (start + 0.5B load + the failure); libtest's own per-test time is 0.15–1.61s. Every run's test result: line is FAILED. 0 passed; 1 failed — the assertion, not a timeout, ends it.

Mutation evidence — the step-budget arm. MINFER_TEST_TICK=spin over the whole #[ignore]d set with the two serve_loop gates skipped is 10 failed / 21 passed in 12.45s (CPU build, dgxspark, 2026-09-27; cargo test --release --bin minfer -- --ignored --nocapture --skip serve_loop_publishes_the_queue_and_running_depth --skip a_job_rejected_for_want_of_a_slot_is_answered_with_503). Each bounded stepper stops exactly one step past its budget, and the budget arm is the one that fires:

gatethe budget arm
published_metrics_move_as_requests_are_served81 steps exceeded the 80-step budget
a_prefix_copied_from_another_slot_answers_identically85 > 84
a_store_inside_a_shared_prefix_takes_a_private_row133 > 132
a_long_request_may_use_the_whole_arena1273 > 1272
server_batch_matches_serial_and_is_faster369 > 368
a_slot_snapshot_resumes_the_context_without_re_prefilling69 > 68
a_chunked_prefill_answers_like_an_unchunked_one457 > 456
a_repeated_chunked_prefill_stops_rebuilding441 > 440
a_long_prefill_keeps_another_slot_decodingnot the budget: the work counter does move under spin, and its priming loop is a fixed three steps, so the gate's own "slot 1 must have emitted something before the long prefill" assertion fires instead

The tenth failure is the_seam_fails_the_batch_forward_without_a_bespoke_mock: the #171 gate asserts a tick reaches the forward and reports the injected 500, and under any MINFER_TEST_TICK injection tick returns before the forward, so it fails on its own .expect_err. It is reported for completeness, not a bounded stepper.

The bound that could not be made to fail fast — honest scope. Two #[ignore]d gates step the engine through the production serve_loop on the test's own thread: serve_loop_publishes_the_queue_and_running_depth and a_job_rejected_for_want_of_a_slot_is_answered_with_503. A persistent wedge keeps the engine busy, so serve_loop keeps ticking and the test thread never returns — there is nowhere for a test assertion to run. Measured with MINFER_TEST_TICK=wedge … --exact under an outer timeout 30: the process is killed at exit 124 (it hangs), which is the shape #158's audit already recorded and #160's table classifies as a redundant backstop. Making those two wedge-proof needs a production liveness bound in serve_loop (a counted consecutive-no-progress limit) or the engine's work counter published to ServerMetrics so a spawned worker can be observed; either is a production change this ticket's scope fence excludes. Filed as #196.

Verification (2026-09-27, CPU build, dgxspark (aarch64); --bin minfer for the gate set).

CommandResult
cargo test --release465 passed / 0 failed / 33 ignored unit + 10 / 0 / 6 integration (unchanged — no test added or removed)
scripts/real_model_gates.sh (parallel, 0.5B)33 / 0 in 35.08s (42.9s wall)
PARALLEL=0 scripts/real_model_gates.sh (serial, 0.5B)33 / 0 in 36.67s (36.8s wall)
MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf PARALLEL=0 scripts/real_model_gates.sh33 / 0 in 46.39s (46.6s wall)
MINFER_TEST_TICK=wedge per gate (9 gates)each FAILED on step 1, 0.21–1.68s wall (table above)
MINFER_TEST_TICK=spin (set, minus the two serve_loop gates)10 failed / 21 passed in 12.45s (table above)
MINFER_TEST_TICK=wedge on the two serve_loop gates, outer timeout 30exit 124 — hangs, recorded as the limit above (#196)
rustup run stable rustfmt --edition 2021 --check on the changed .rsclean (rustfmt 1.9.0-stable; the pinned 1.97.1 toolchain has no rustfmt component)

Docs. docs/GATE-CONTRACT.md rule 4's closing "#160 tracks the remaining unbounded steppers" is replaced with the outcome (the steppers are bounded; the remaining watchdogs are named backstops; the seam is MINFER_TEST_TICK), and §3 documents the seam next to MINFER_TEST_CALL_FAIL. docs/BUILD.md § Tests records the same and the mutation lever. ARCHITECTURE-ROADMAP.md is untouched: this is robustness inside the existing test-infrastructure/batching rows, not a new capability — the same reason the #154, #151 and #158 records give.

Test-infrastructure record (#196, 2026-09-27) — serve_loop carries a counted no-progress bound, so a wedge answers its clients and ends

The defect. #160 made every while engine.busy() stepper in the #[ignore]d server gates fail fast under an injected wedge — except two, for a structural reason:

  • server::batch::tests::serve_loop_publishes_the_queue_and_running_depth
  • server::batch::tests::a_job_rejected_for_want_of_a_slot_is_answered_with_503

Both drive the engine through the production serve_loop on the test's own thread. A persistent wedge in BatchEngine::tick leaves the engine busy forever, serve_loop keeps ticking, and the test thread never returns — there is nowhere for an assertion to run. #160 measured MINFER_TEST_TICK=wedge … --exact <gate> under an outer timeout 30 as exit 124 (a hang); the feeder's FEEDER_POLL_BACKSTOP cannot rescue it, because the worker is on the same thread. In production the same shape is #151's client-visible failure one level up: the loop spins at 100% CPU, no client is ever told, in_flight stays ≥ 1, and a graceful drain burns its whole MINFER_DRAIN_MS deadline.

What landed — a production no-progress guard in serve_loop. src/server/batch.rs:

  • STALL_STEP_LIMIT: u64 = 64 — the number of consecutive steps that may leave the engine busy without advancing BatchEngine::work_units. The healthy maximum is 0: a tick that leaves the engine busy has either forwarded a decode row or committed a token through advance's Continue, and both increment the counter (#158's invariant; the gates' WorkBound asserts it per step, and this is the same invariant enforced in production). 64 is deliberately loose rather than tuned: a future engine change that legitimately defers work for a handful of steps is not mistaken for a stall, and 64 no-op steps cost well under a millisecond, so a real stall still ends immediately against the previous "spins forever". It is a count (steps), never a wall-clock number (rule 4).
  • On the 64th consecutive no-progress step the loop: counts minfer_worker_stalled_total; answers every live run exactly once with ApiError::server("the worker stalled") (500 server_error) through the existing fail machinery — new BatchEngine::fail_all, the fail_batch shape without a row list, since a wedged step may have built no batch; answers every queued-but-unadmitted job (the worker's pending deque plus the channel) with the same terminal error through #121's reject, so its sender never drops into the silent empty 200 — new reject_queued; zeroes worker_pending, publishes metrics, logs, and breaks the loop. The queued jobs leave the queue, so they are counted in jobs_admitted_total and jobs_dropped_total, keeping accepted - admitted a true queue depth.
  • The count is reset wherever the loop legitimately moves: on blocking_recv (an idle wait is not a spin), on any step that advanced work_units, and on #151's Err step (it answered its batch and released its slots; work_units does not move there because the increment sits after the successful forward). The last reset is deliberate: without it, a saturated server whose every forward fails deterministically — a state #151 already handles correctly, one 500 per batch — would accumulate no-progress steps and be declared stalled, turning a wrong-but-answering server into a stopped one.
  • Ending the loop drops job_rx, so a later request gets the existing 503 unavailable_error("server shutting down") (requests_rejected_total), not a hang.

The condition is observable next to the other AtomicU64s: ServerMetrics::worker_stalled_total and the minfer_worker_stalled_total family in the /metrics rendering. It is the guard's own signal — a count of stall events — and deliberately not #157's general terminal-error accounting.

Gates. Both #[ignore]d gates now assert worker_stalled_total == 0, the property the mutation breaks. serve_loop_publishes_the_queue_and_running_depth also prints the terminal error each client received, and its feeder leaves as soon as the counter moves, so the assertion runs at once instead of waiting out FEEDER_POLL_BACKSTOP (before that the mutated gate failed correctly but only after 120.19s). a_job_rejected_for_want_of_a_slot_is_answered_with_503 checks the counter after reading the channels, so a wedge surfaces as the error the client actually got. Two plain #[test]s cover the answer machinery in CI, with no model: the_stall_answers_every_live_run_exactly_once (three live runs and an idle slot: one 500 each, slot freed, prefix cleared, idle slot untouched) and the_stall_answers_every_queued_job_exactly_once (one job in the deque, two in the channel: one 500 each, senders closed).

Mutation evidence (rule 3). MINFER_TEST_TICK=wedge is an environment switch, so no source revert is involved and git diff stayed clean throughout. CPU build, dgxspark (aarch64), 2026-09-27, MINFER_TEST_TICK=wedge cargo test --release --bin minfer -- --ignored --exact <gate> --nocapture, no outer timeout:

gateexitwall (whole process)what failedthe client's answer
serve_loop_publishes_the_queue_and_running_depth1010.29s"the worker tripped its counted no-progress bound … left: 1, right: 0"two runs, one 500 the worker stalled each
a_job_rejected_for_want_of_a_slot_is_answered_with_5031010.18s"a rejected job is unavailable: the worker stalled" (left: 500, right: 503)the served run got 500 the worker stalled

Both logs also carry [server] the worker stalled: 64 consecutive steps left the engine busy without advancing its work counter (N run(s), 0 queued job(s) answered with 500); stopping the worker. A whole-set wedge run with no --skip is 21 passed / 12 failed in 12.04s (the nine bounded steppers, the two serve_loop gates, and the #171 seam gate that cannot reach its forward under any injection) — against #160's exit 124 for the same two gates under an outer timeout 30. The unmutated gates on the same commands are pass, exit 0 (0.5B: 3.09s / 0.21s; Qwen3-0.6B: 3.69s / 0.17s).

The spin arm still needs the two gates skipped — honest scope. MINFER_TEST_TICK=spin advances work_units without ever completing a run, so by construction a step that keeps "moving" cannot trip a count of no-progress steps; the two serve_loop gates would still spin (they have no step budget, and adding one is exactly the shape #196 offers as its second, rejected alternative). The #160 command is unchanged:

Command (2026-09-27, CPU build, dgxspark (aarch64))Result
MINFER_TEST_TICK=spin cargo test --release --bin minfer -- --ignored --skip serve_loop_publishes_the_queue_and_running_depth --skip a_job_rejected_for_want_of_a_slot_is_answered_with_50321 passed / 10 failed in 12.21s (the #160 budget arm, unchanged)

Verification (2026-09-27, CPU build, dgxspark (aarch64)).

CommandResult
cargo test --release467 passed / 0 failed / 33 ignored unit + 10 / 0 / 6 integration (baseline 465/0/33; the two CI stall gates)
the two gates --ignored --exact, unmutated, 0.5B1 / 0 each, 3.09s / 0.21s
the two gates --ignored --exact, unmutated, Qwen3-0.6B1 / 0 each, 3.69s / 0.17s
scripts/real_model_gates.sh (parallel, 0.5B)33 / 0 in 35.96s
PARALLEL=0 scripts/real_model_gates.sh (serial, 0.5B)33 / 0 in 38.80s
MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf scripts/real_model_gates.sh (parallel)33 / 0 in 42.48s
MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf PARALLEL=0 scripts/real_model_gates.sh (serial)33 / 0 in 46.95s
MINFER_TEST_TICK=wedge, whole #[ignore]d set, no --skip21 passed / 12 failed in 12.04s, no hang
MINFER_TEST_TICK=spin, two serve_loop gates skipped21 passed / 10 failed in 12.21s
rustup run stable rustfmt --edition 2021 --check src/server/batch.rs src/server/metrics.rsclean (rustfmt 1.9.0-stable; the pinned 1.97.1 toolchain has no rustfmt component)
python3 scripts/check_status.py --checkexit 0
python3 scripts/check_docs_links.py968 relative links / 186 markdown files, exit 0

Honest limits.

  • The bound is on the loop, not inside a single tick: a tick that never returns (or spins internally without returning to the loop) still has no in-process bound — only a cross-process watchdog can bound that, and none claims to here.
  • The guard stops the worker. Requests after the stall get 503 server shutting down until the process is restarted; that is the deliberate trade against spinning at 100% CPU with clients left hanging, and minfer_worker_stalled_total is the signal to alert on. A self-restart is out of scope.
  • STALL_STEP_LIMIT = 64 cannot be validated by observation on the healthy path (it never reaches 1); it is slack, chosen by the argument above, not a measured constant.
  • The spin arm above: a step that advances the work counter is not a stall by this guard's definition, so the two gates remain skipped in that mutation run. Filed as #198 — a production bound for the moving-but-non-terminating step, which must not false-positive a legitimately long generation.
  • The CUDA unit row of docs/status.toml (546, GB10 sm_121, 2026-09-27) was not re-measured: this ticket's two new CI tests are feature-independent, so a CUDA build would read +2; the row stays the dated device record and is not claimed to include them (the GPU rows were out of scope).
  • AGENTS.md no longer carries a per-module server/batch.rs bullet (the Docs lines of the #121 and #151 records predate that refactor); the server contract now lives in FEATURES.md, USAGE.md and OPENAI-CHAT-API-PLAN.md, which are the files updated here.

Docs. docs/GATE-CONTRACT.md rule 4's #160 paragraph now ends with the resolution (the serve_loop gates are wedge-proof; STALL_STEP_LIMIT; minfer_worker_stalled_total; FEEDER_POLL_BACKSTOP demoted to a last resort). docs/BUILD.md § Tests and scripts/real_model_gates.sh's mutation note record the same. Every /metrics enumeration (USAGE.md § Metrics and observability, FEATURES.md's serve bullet, OPENAI-CHAT-API-PLAN.md § Slot Lifecycle, the F8 table above) names the new counter. AGENTS.md's suite counts carry the +2 (467 aarch64 / 465 x86_64). ARCHITECTURE-ROADMAP.md is untouched: no roadmap gap closes here — this is robustness inside the existing test-infrastructure/batching rows, the same reason the #154, #151, #158 and #160 records give.

Test-infrastructure record (#218, 2026-09-29) — the prefill-GEMM dynamic-smem opt-in is lazy, per-instantiation, and gated

The finding. #188 deleted the production call site of gemm_prefill_smem_init — the eager dynamic-smem opt-in in CudaState::try_new — and said so nowhere: the commit message has no smem/shared/attribute/#145 token, the same PR's docs commit rewrote CUDA-BACKEND-DESIGN.md (+174/−74) without mentioning it, and the implementation commit touched six .rs files and no documentation. Two later dead-code-hygiene commits then annotated the orphan #[cfg_attr(not(test), allow(dead_code))]: f5a956a under the banner "55 items are reached only from #[cfg(test)] code", and 58379c0 left it because the #145 gate calls it. The instrument that should have surfaced an orphaned production path instead silenced the diagnostic, and the #145 gate cuda_prefill_smem_optin_covers_every_launchable_instantiation kept certifying a function production no longer called.

The decision (plan B). Do not reintroduce the eager init. Keep the lazy per-launch path (gemm_smem_optin, reached from launch_gemm_f16), remove the orphan rather than rename it, and make the invariant explicit and gated. The invariant: the >48 KiB attribute is in force before a capture window opens and is never set inside one. Three mechanisms uphold it — the 3-run capture warmup (capture_warmup, default 3), cudaStreamCaptureModeThreadLocal (#188's measured choice), and the per-instantiation cache — and all three are now written down at their sites (graph_replay_step's comment, gemm_smem_optin's comment, CUDA-BACKEND-DESIGN.md §2.4) so the next person who changes the capture mode or the warmup count reads why. The #145 gate was re-pointed at the production function and renamed cuda_prefill_smem_lazy_optin_admits_every_launchable_instantiation; the checked/skipped introspection and the gemm_prefill_smem_init symbol are gone, and the remaining introspection is a #[cfg(test)] extern "C" block, so no non-test build carries even a declaration for it (the allow(dead_code) pattern this ticket exists to end).

The bug the first gate run found. The gemm_smem_optin cache was not per-instantiation. It was template <typename K> and every gemm_f16_nt_kernel_t<TM,KS,AF32> shares one signature, so K deduced to one function-pointer type and static int state was one cache for the whole family. The re-pointed coverage gate found it on its first run: <64,64,true> (73728 B) answered "admitted" for <128,64,false> (57344 B), whose attribute was never set — the device still reported its 49152 B default for the 57344 B request. #218 made the cache genuinely per-instantiation (template <int TM, int KS, bool AF32>), which is what its comment always claimed. It is latent in production today (only one (tm, ks, af32) is used per process — both tile env knobs are read once — so nothing mixed families), but it is exactly the "unwritten, untested" property this ticket exists to end.

The gates.

gatevalue it assertsthe precondition that keeps it non-vacuous
cuda_prefill_smem_optin_is_done_by_productionafter a real prefill forward, the device's own cudaFuncGetAttributes().maxDynamicSharedSizeBytes read-back says <128,64,false> (57344 B) is opted init runs in a fresh process (src/cuda/test_child.rs) and asserts opted_in == 0 before the forward — the tile env and the attribute are both process-scoped, so "before" is only observable in a process no earlier launch has touched
cuda_prefill_smem_optin_refusal_fails_the_prefillthe control arm: MINFER_TEST_CALL_FAIL=attr:gemm_f16_f16 makes the production prefill refuse the launch and the site report name cudaFuncSetAttribute + cudaErrorInvalidValuea second child process, so the per-instantiation cache cannot have answered already; env-gated behind MINFER_TEST_ISSUE218=1 (it makes a real call fail, as #147/#162's gates are)
cuda_prefill_smem_optin_is_never_set_inside_a_capture_windowa >48 KiB prefill-shaped graph is captured (captured_count() == 1) and replays bitwise over 5 steps, and gemm_smem_optin_in_capture_count() == 0 — the opt-in ran before the window, never inside itfresh process; gemm_smem_need > 48 KiB asserted; captured_count() == 1 so an uncaptured path cannot pass; opted_in == 0 before; MINFER_TEST_CAPTURE_WARMUP=1 is the test-only seam that makes it fail
cuda_prefill_smem_lazy_optin_admits_every_launchable_instantiation (the re-pointed #145 gate)every launchable >48 KiB instantiation reads back opted in through the production gemm_smem_optin; every over-limit one is refused without a callit drives the production function (via gemm_prefill_smem_optin_one_for_test), not a mirror; the over-limit combination is a separate negative arm

Mutation evidence (rule 3), GB10 sm_121, 2026-09-29.

  1. gemm_smem_optin answers true without calling minfer_smem_optin (C++, one line). The coverage gate goes red at <64,64,true> ("cudaFuncGetAttributes reports its maxDynamicSharedSizeBytes below that"); the real-prefill gate's child goes red with minfer/cuda: kernel launch gemm_f16_nt_kernel_t<128,64,false> failed: cudaErrorInvalidValue (1) — the launch is refused (#162/launch:gemm_f16_f16) and then "cuda: prefill GEMM (f16): launch_gemm_f16 refused the launch … no kernel ran"; the capture gate's child goes red on the same refused launch. One mutation, three gates.
  2. MINFER_TEST_CAPTURE_WARMUP=1 (the documented seam, from the parent — the harness passes it through on purpose): the child's whole gate runs — the capture happens, replays bitwise, and opted_in == 1 — and only gemm_smem_optin_in_capture_count() trips: "assertion left == right failed: the smem opt-in must be performed before a capture window opens, never inside one — this is the load-bearing part of the design; left: 1, right: 0".
  3. prefill_gemm_f16_inner's if (launched == 0) return Err(..) arm removed (Rust): the control arm's expect_err goes red — "the injected attribute failure must refuse the >48 KiB prefill launch: ()".

Honest limit — now measured, and it still does not license removing the warmup. Mutation 2 is also the experiment the #218 issue said had never been run: it drives cudaFuncSetAttribute into an open same-thread cudaStreamCaptureModeThreadLocal window. On this runtime the call is tolerated — the attribute publishes, the capture completes, instantiates and replays bitwise-identically; only the counter moves. So on GB10 sm_121 / CUDA 13.0 / driver 580.178.04 an in-window opt-in would work. The design nevertheless keeps the call out of the window, because the historical failure is real on other toolkits, the 2026-09-25 probe measured the Global mode (not the adopted one), and the whole point of the #218 gate is to pin the property as an observed invariant rather than rely on driver behaviour. The counter gate is what makes that true regardless of what the driver tolerates.

Counts (rule 5). scripts/cuda_test.sh → 565 / 0 / 42, GB10 sm_121, 2026-09-29 (was 562 / 0 / 42; +3 device gates: the two fresh-process issue218_tests arms and the captured-graph arm). compute-sanitizer --tool memcheck --target-processes all <test binary> --test-threads=1 → 0 API errors over 565 / 0 / 42 (one aggregated ERROR SUMMARY for the whole process tree — the fresh-process children are followed too). The two real-model configurations were re-run and stay green: FEATURES=cuda scripts/real_model_gates.sh → 42 / 0 (0.5B config) and the same command with MINFER_BATCH_TEST_MODEL=~/.cache/minfer/models/hf/Qwen/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf → 42 / 0, GB10 sm_121, 2026-09-29.

Docs. CUDA-BACKEND-DESIGN.md §2.4 (the eager claim replaced by the lazy path, the three load-bearing mechanisms, the gate list, the measured in-window experiment); cuda_optimization_steps/02-wmma-f16-prefill-gemm-8m.md and this file's #145 record carry dated forward notes (history is not rewritten); inference_e2e_walkthrough/15-cuda-backend.md no longer says gemm_prefill_smem_init() runs eagerly; docs/GATE-CONTRACT.md rule 1 gains the reusable lesson from this ticket — a gate must exercise the production entry point, not a test-only helper that mirrors it — with the #145 gate as the instance. docs/status.toml and AGENTS.md carry the new CUDA counts.

Process lesson. A dead-code campaign treated "an item with no production caller that looks like production" as a fact to annotate rather than a question to ask. allow(dead_code) on a test-reachable production-looking item is a deferred question; the fix is to list such items for review. Future hygiene work should treat the annotation as a lead, not a resolution.

#223 forward note (2026-09-29): the "tested, not enforced" gap this record names is closed. CudaState::try_new now runs the eager pre-warm through the production gemm_smem_optin, at the earliest point in the process, where no capture window can exist yet; the lazy path and the three mechanisms this record lists remain as defence in depth. The four gates above keep their claims by running their children with MINFER_NO_GEMM_PREWARM=1 (the documented control and the "lazy path alone" arm); the runtime guarantee has its own fifth gate. See the #223 record below.

Test-infrastructure record (#223, 2026-09-29) — the eager prefill-GEMM smem pre-warm is back, as a runtime guarantee through the lazy entry

The gap #218 left. #218 made the prefill-GEMM dynamic-smem invariant explicit, gated and correctly cached, but production still upheld it only by emergent means: the 3-run capture warmup, cudaStreamCaptureModeThreadLocal, and the fact that the cache happens to be consulted on the first launch. The attribute was a tested property, not an enforced one — #188 had deleted the eager caller and nobody noticed.

The design (plan A). CudaState::try_new drives the production pre-warm once per process for every launchable (tm, ks, af32), at the earliest point in the process. The placement is the argument: try_new runs under CUDA.get_or_init, before the state is published and before any CudaBackend — the only thing that can hold a capture window — can exist, so "the attribute is set outside any window" holds by construction. Crucially it is not the old gemm_prefill_smem_init sweep: the Rust side enumerates the set and each entry goes through the new production gemm_prefill_smem_prewarm_one → gemm_smem_optin<TM,KS,AF32> — the same per-instantiation cache the launcher reads. One mechanism, one cudaFuncSetAttribute site, one copy of gemm_dynamic_smem_bytes. The C++ MINFER_GEMM_OPTIN_SET X-macro is now the single list of the launchable set (fatbin lookup, pre-warm and the #218 test seam all expand it). The lazy per-launch opt-in stays as defence in depth, and the three #218 mechanisms are demoted to defence in depth behind the pre-warm. A request above cudaDevAttrMaxSharedMemoryPerBlockOptin is skipped without calling the attribute (reason named); a failure is named per instantiation by minfer_smem_optin; the checked/skipped counters and the startup banner did not come back — a fully admitted pre-warm is silent. MINFER_NO_GEMM_PREWARM=1 is the documented control.

The fifth gate. issue223_tests::cuda_prefill_smem_prewarm_opts_in_every_launchable_instantiation_before_any_launch asserts, in a fresh process immediately after context creation and before any launch, that every launchable >48 KiB instantiation already reads back opted in. Non-vacuity: a second fresh process with MINFER_NO_GEMM_PREWARM=1 asserts the negation (the read-back is capable of answering 0), the child launches nothing before the check, and the assertion names the load-bearing gemm_f16_nt_kernel_t<128,64,false> (57344 B). This gate exists for a mutation the #218 arms cannot see: skip one (tm, ks, af32) in the pre-warm and the lazy path simply opts it in on first launch, so the coverage/counter arms stay green. The four #218 arms now run their fresh-process children with MINFER_NO_GEMM_PREWARM=1, which is where their pre-#223 opted_in == 0 preconditions are observable; their claims are unchanged (they are the lazy-path-alone arm) and they remain the cache-keying detector.

Mutation evidence (rule 3), GB10 sm_121, 2026-09-29. Drop (128, 64, 0) from GEMM_PREWARM_SET (the Rust production list; replaced by a duplicate so the array still type-checks) and run the new gate: the pre-warmed child fails with "immediately after context creation and before any kernel launch, gemm_f16_nt_kernel_t<128,64,false> (57344 B > 48 KiB) must already read as opted in … left: 0, right: 1". Under the same mutation cuda_prefill_smem_lazy_optin_admits_every_launchable_instantiation stays green (its child skipped the pre-warm, so the lazy path opts the dropped instantiation in on first launch) — the transcript that justifies the fifth gate. The over-limit skip is visible in every pre-warmed child's stderr: cudaFuncSetAttribute(gemm_f16_nt_kernel_t<256,64,true>, cudaFuncAttributeMaxDynamicSharedMemorySize, 122880 B) SKIPPED: the request exceeds cudaDevAttrMaxSharedMemoryPerBlockOptin (101376 B) ….

Performance, measured (rule 5) — the real binary disagrees with the #223 proxy by ~15×, and the net is still zero. Date/device/command: 2026-09-29, GB10 sm_121, CUDA 13.0, driver 580.178.04, MINFER_OP_TIMING=1 ./target/release/minfer bench -p 8 -n 1 -r 1, 25 fresh processes.

bar (named before measuring)proxy in #223measuredverdict
pre-warm loop's own duration≲ 0.2 ms (first attr call 152.9 µs)median 2249 µs (range 2126–2448, n=25)above the stated bar — it is the fatbin's one-time module load, not 152.9 µs; set-size independent (n=1 ≈ n=12 ≈ 2.2 ms). The range is warm-clock only — see the cold row below
pre-warm loop's own duration, first (cold / idle-clock) invocationsame one-time work, no separate bar named~14 526 µs (~6× the warm median; the three consecutive runs were 14526, 2157, 2364 µs)the 2126–2448 range is not the worst case: the module load is clock/state dependent, and its first cold run measured ~14.5 ms (#225)
minfer bench -p 2048 -n 128 tg128within ±1%ON 236.30 vs OFF 236.45 t/s (−0.06%)pass (7 interleaved matched rounds, medians, same binary)
minfer bench -p 2048 -n 128 pp2048within ±1%ON 2546.55 vs OFF 2544.34 t/s (+0.09%)pass
startup (model load → first token)net new ≈ 10 µsnet ≈ 0: the ~2.2 ms moves into the existing prewarm_prefill() module loadthe proxy's qualitative reading survives. The ≈ 0 net is coupled: it holds only while a later step pays that same module load — move prewarm_prefill() after the first launch, or remove it, and the pre-warm's loop becomes ~2.2 ms of net-new startup cost

The "the ~150 µs moves rather than appears" reading did survive on the real path, with a different magnitude. Controlled fresh-process probe on the same binary (the tiny gemm_f16_nt_kernel_t fixture, MINFER_MMQ 0/1 × MINFER_GEMM_K64 0/1): first prefill forward 2288–2314 µs with the pre-warm off vs 78–120 µs with it on; prewarm_prefill() (the r59 rider's minfer_prewarm_kernels, called at the end of Qwen2/Qwen3 weight registration) 4.3–4.6 ms off vs 2.2–2.4 ms on. So the ~2.2 ms is a one-time fatbin module load that the startup path already pays; the pre-warm only decides where — at try_new instead of at the end of registration. The end-to-end CLI phase measurement is consistent with that but cannot resolve the net to better than a few ms: paired cuda_ready (mid-context-creation) +2.53 ms, forward_ms +0.05 ms, model-load phase noise ±20 ms. The honest residual: if the module load ever stops being paid by prewarm_prefill() (e.g. that rider is removed), the pre-warm's loop becomes ~2.2 ms of net new startup cost, so the two are coupled and the next person to touch prewarm_prefill must know.

#225 forward note (2026-10-04): the table's 2126–2448 range is a warm-clock measurement of the one-time fatbin module load, not a bound. An independent three-run check under the same command (2026-09-29, GB10 sm_121) saw the first, cold / idle-clock invocation take 14 526 µs (~6× the warm median; the next two runs 2157 µs and 2364 µs), so the range above is warm-only and the load's magnitude is clock/state dependent. CUDA-BACKEND-DESIGN.md §2.4's cost table now carries both readings. The coupling is load-bearing: the ≈ 0 net holds only while a later step pays the same module load — today prewarm_prefill(), otherwise the first launch from the fatbin.

Counts (rule 5). scripts/cuda_test.sh → 566 / 0 / 42, GB10 sm_121, 2026-09-29 (was 565 / 0 / 42; +1 device gate, cuda::issue223_tests). compute-sanitizer --tool memcheck --target-processes all <test binary> --test-threads=1 → 0 API errors over 566 / 0 / 42. The two real-model configurations were re-run and stay green: FEATURES=cuda scripts/real_model_gates.sh → 42 / 0 (0.5B config) and with MINFER_BATCH_TEST_MODEL=~/.cache/minfer/models/hf/Qwen/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf → 42 / 0, GB10 sm_121, 2026-09-29.

Docs. CUDA-BACKEND-DESIGN.md §2.4 now states the restored design (eager pre-warm at context creation + lazy per-launch opt-in), demotes the three mechanisms to defence in depth, names the fifth gate, and carries the measured cost table; cuda_optimization_steps/02-wmma-f16-prefill-gemm-8m.md and this file's #145 and #218 records carry dated forward notes (history is not rewritten); inference_e2e_walkthrough/15-cuda-backend.md still describes only prewarm_prefill(), which is unchanged. docs/status.toml and AGENTS.md carry the new CUDA counts.

Test-infrastructure record (#228, 2026-09-30) — the E1 test corpus moves onto the production fill_batch_inputs

The finding (S1 of #227). GraphAllocator::fill_attn_inputs (alloc.rs:1724 on master b5a9d5f) was a production-shaped entry point with #[cfg_attr(not(test), allow(dead_code))] and a doc that claimed "the model path and hand-built graphs both use it". The model path does not: forward_batch and forward_cached — the latter literally forward_batch(&Batch::single(tokens, positions), …) (models/qwen2/graph.rs:458-465) — call fill_batch_inputs, a separate implementation of the same job. The coverage was therefore inverted: E2 had one test call site (qwen2::graph::tests:2068) while E1, which production never runs, had 15. A change to E2's COW ordering, reservation or ownership recording could leave the suite green — the "mirrored helper" hazard of GATE-CONTRACT §1.

Classification of the 15 E1 call sites (step 1 done before any rewriting). Every one asks the same question — can the scenario be a Batch? — and the answer is yes for 14; the exception is the rope-only fixture.

file:line (master)driving testgraph has cells?disposition
alloc/tests.rs:423kv_session_round_trips_the_rows_and_the_run_tableyesmoved → Batch::new(_, _, [SEQ, SEQ])
alloc/tests.rs:558an_f16_session_round_trips_through_save_and_loadyesmoved → Batch::new
alloc/tests.rs:794a_copy_on_write_moves_the_rows_and_never_writes_throughyesmoved → Batch::new(_, _, [2]) (not Batch::single: that names SEQ_MAIN, which holds no run here)
cpu_backend/tests.rs:377embedding_and_ropenoexception — see below
cpu_backend/tests.rs:436kvcache_store_load_and_attn_roundtripyesmoved → Batch::single
cpu_backend/tests.rs:507a_packed_kv_region_answers_like_the_f32_one_and_is_smalleryesmoved → Batch::single
cpu_backend/tests.rs:678a_packed_physical_shift_moves_v_verbatim_and_requantizes_kyesmoved → Batch::single; the redundant post-execute kv_note_used(nt) deleted with it
qwen2/graph/tail_tests.rs:203tail_reduction_matches_full_nt (run_keep)yesmoved → Batch::single
qwen2/graph/tail_tests.rs:403fused_qkv_matches_unfused_decodeyesmoved → Batch::single
qwen2/graph/tail_tests.rs:498fused_qkv_matches_unfused_decodeyesmoved → Batch::single
qwen2/graph/tests.rs:2907graph_logits_match_forward_real_model (run_prefill_decode)yesmoved → Batch::single
qwen2/graph/tests.rs:2955graph_logits_match_forward_real_model (decode step)yesmoved → Batch::single(&[next], &[nt])
qwen2/graph/tests.rs:3107graph_metal_layer0_isolation (macOS)yesmoved → Batch::single
qwen2/graph/tests.rs:3132graph_metal_layer0_isolation (macOS)yesmoved → Batch::single
qwen3/graph/tests.rs:406metal_prefill_determinism (macOS)yesmoved → Batch::single

The no-cells case exists — and it is a no-op. cpu_backend::tests::embedding_and_rope builds an embedding + RoPE graph with no kvcache_store, so it has none of seq_ids / cells / kv_map / attn_span. E1's has_cells guard made its call there a complete no-op, and E2 cannot express the shape at all: with no KV node there is no arena (n_ctx() == 0), and fill_batch_inputs's per-group implicit reservation is reserve_seq(seq, 0), which is refused. So this is the one site that keeps the shape, as #[cfg(test)]-scoped GraphAllocator::fill_attn_inputs_without_cells — the has_cells == false branch and nothing else, debug_asserted to reject a cells graph, with the doc naming embedding_and_rope. It is #[cfg(test)]-scoped, not allow(dead_code)-silenced, so a non-test build cannot see it.

What was deleted, and the honest disposition of each item that lost its only caller.

  • GraphAllocator::fill_attn_inputs — deleted with its header comment (it was not an entry point).
  • GraphAllocator::kv_note_used (alloc.rs:1240) — after the move its last caller was the redundant cpu_backend::tests:683 line; that line is gone, so it had nothing and was deleted (a dead item behind an allow is exactly the pattern Core Convention 5 exists to stop). Its doc claimed "the model calls this after a forward with max(positions)+1" — false; production records the extent through KvCache::own_positions.
  • KvCache::own_prefix (kvcache.rs:582) — kv_note_used was its last non-test caller, so it is now test-only; it keeps its #[cfg_attr(not(test), allow(dead_code))] and its doc now names the four kvcache::tests that drive it.
  • kv_cells_for_seq is not an orphan — the #228 premise (inherited from #227) was tested and falsified. The premise was that its only non-test callers are E1 and S3's kv_cell_of, so it is production-unreachable yet unannotated, "silent only because both its callers are annotated", and therefore missed by an annotation-grep census. The reachability chain says otherwise: models/qwen2/graph.rs:660 / qwen3/graph.rs:567 → fill_batch_inputs → fill_seq_ids → kv_cells_for_seq (the has("cells") branch), and every model graph has a cells input because kvcache_store always creates one (models/qwen2/graph.rs:217). It was verified at runtime, not by grep: an env-gated eprintln! at the top of kv_cells_for_seq plus cargo test --release models::qwen2::graph::tests::graph_logits_match_forward_real_model → 4 hits, test green (probe reverted). So it is production-reachable, has no annotation because rustc is right, and the annotation-grep/transitive-closure worry does not apply to this subgraph: kv_cells_for_seq needs no change. S1 and S3 are still one subgraph (joined through kv_cells_for_seq), but the closure direction is inverted, and that inversion is the result #227's census method needs. The genuinely dead chain is the smaller one: fill_attn_inputs → kv_note_used → own_prefix (E1 and kv_note_used deleted; own_prefix now test-only with its doc naming the tests). Its documented role is unchanged; only the "two fill entry points" comment was corrected.
  • Residual, not fixed (out of this ticket's scope, recorded): kv_cells_for_seq's classic branch (arena_stats().sequences == 0 → cell == position) is now reachable only from alloc/tests.rs; production always reserves before fill_seq_ids runs, because fill_batch_inputs's first loop covers every group.

Mutation evidence (rule 3), CPU aarch64, 2026-09-30. The production path must be the thing that goes red. Mutation: in fill_batch_inputs, own the wrong positions — own_positions(seq, positions.iter().map(|p| p + 1)) — so the ownership recording is off by one while everything still compiles (a plain "skip own_positions" mutation does not compile: the item would become dead and #![cfg_attr(not(test), deny(warnings))] fires, which is itself evidence the recording is load-bearing).

cargo test --release
test result: FAILED. 478 passed; 3 failed; 36 ignored; 0 measured; 0 filtered out
  graph::alloc::tests::kv_session_round_trips_the_rows_and_the_run_table
    src/graph/alloc/tests.rs:445: assertion `left == right` failed: positions 0..4 are written
      left: 5   right: 4
  graph::cpu_backend::tests::a_packed_physical_shift_moves_v_verbatim_and_requantizes_k
    src/graph/cpu_backend/tests.rs:722: assertion `left == right` failed: one row removed
      left: 3   right: 2
  models::qwen2::graph::tests::kv_rm_is_exact_and_the_window_shift_is_a_named_tolerance_class
    src/models/qwen2/graph/tests.rs:2676: assertion `left == right` failed: only A survives removing B
      left: 8   right: 7

Two of the three are moved E1 call sites; the third, kv_rm_is_exact_and_the_window_shift_is_a_named_tolerance_class, reaches fill_batch_inputs through the real production caller forward_graph_cached — the strongest form of the point, because the mutation is visible to production's own path and not only to test-driven fills. Mutation reverted; grep -n MUTATION src/graph/alloc.rs empty and git diff clean of it.

Counts (rule 5). No #[test] was added or removed, so the suite counts are unchanged and no CUDA row was re-measured: cargo test --release, box dgxspark (aarch64, GB10 sm_121), 2026-09-30 → 481 / 0 / 36 unit + 10 / 0 / 6 integration (the same as the recorded CPU row). cargo fmt --all --check clean; the non-test build warning-free with and without --features cuda (the one warning: line on the CUDA build is build.rs's pre-existing cargo:warning= target list, not a rustc diagnostic).

Limits. (1) The macOS-only arms (graph_metal_layer0_isolation, fused_qkv_matches_unfused_decode, metal_prefill_determinism) are type-checked only by the CI build-macos job here — no Mac. (2) The x86_64 CPU row cannot be computed locally; it is unchanged by construction (same test count) and the CI test-linux-cpu log is its source. (3) The no-cells helper's claim is weak by nature ("filling a graph with no KV input resolves nothing and touches no arena"); it is kept because the ticket asks for exactly that shape, not because it catches a bug.

Test-infrastructure record (#232, 2026-09-29) — GraphAllocator::kv_own_range deleted; the tests drive KvCache::own_range

The finding (S2 of #227). GraphAllocator::kv_own_range (alloc.rs:1645-1650 on master f4d3390) was a one-line forwarder — self.kv.own_range(seq, from, to) — carrying #[cfg_attr(not(test), allow(dead_code))] and a doc that claimed a production role: "Mark cells [from, to) as written by seq in every layer (E2's batched forwards write several sequences per step)." Production does not call it; its only callers were five sites in src/graph/alloc/tests.rs.

One implementation, not two. Unlike S1's fill_attn_inputs, the wrapper and production share one implementation: the inner, unannotated KvCache::own_range (kvcache.rs:536). It is production-reachable through two paths that do not go through the wrapper:

  • fill_batch_inputs (alloc.rs:1697 on master) → KvCache::own_positions (kvcache.rs:574) → own_range; production reaches fill_batch_inputs from models/qwen2/graph.rs:660 and models/qwen3/graph.rs:567 (forward_batch).
  • GraphAllocator::kv_copy_prefix (alloc.rs:1488) → own_range; its production caller is server/batch.rs:821.

(Correction to [#232]'s body and #227's S2 text: they place alloc.rs:1488 inside kv_private_row_for. It is not — kv_private_row_for (alloc.rs:1516) drives KvCache::private_row_for/apply_private_row and never calls own_range. Line 1488 is the tail of kv_copy_prefix. The conclusion — production does not call the wrapper — is unaffected, but the reachable path is the batch prefix-copy, not the copy-on-write.)

The disposition (b): delete, and move the call sites to the production spelling. The wrapper and its doc comment are deleted. All five call sites — in four test functions — now spell a.kv.own_range(...) / alloc.kv.own_range(...):

testcall sites (master)disposition
kv_session_round_trips_the_rows_and_the_run_table:433a.kv.own_range(SEQ, 0, 4)
an_f16_session_round_trips_through_save_and_load:570a.kv.own_range(SEQ, 0, 2)
kv_defrag_moves_the_bytes_and_opens_the_run:629alloc.kv.own_range(seq, …)
a_copy_on_write_moves_the_rows_and_never_writes_through:739, :742alloc.kv.own_range(1, 0, 4) / (2, 4, 6)

src/graph/alloc/tests.rs is a child module of the type's module (alloc.rs:2735 #[cfg(test)] mod tests;) and kv is a private field (alloc.rs:174), so the direct field access compiles as predicted — the fallback (#[cfg(test)] impl GraphAllocator inside alloc/tests.rs) was not needed. Every assertion is unchanged; the diff is the five spellings plus the deleted wrapper. The issue's "five tests" is five call sites in four test functions — no test was dropped or weakened.

Mutation evidence (rule 3), CPU aarch64, 2026-09-29. The mutation is in the production-used KvCache::own_range: an off-by-one in the written extent — Some(pos) => written = written.max(pos + 1) → written.max(pos).

cargo test --release graph::alloc::tests::
test result: FAILED. 34 passed; 3 failed; 0 ignored; 480 filtered out

---- graph::alloc::tests::a_copy_on_write_moves_the_rows_and_never_writes_through ----
panicked at src/graph/alloc/tests.rs:743:47:
called `Result::unwrap()` on an `Err` value:
  "share_prefix: sequence 1 has written 3 rows, 4 requested"

---- graph::alloc::tests::kv_session_round_trips_the_rows_and_the_run_table ----
panicked at src/graph/alloc/tests.rs:445:5:
assertion `left == right` failed: positions 0..4 are written
  left: 3   right: 4

---- graph::alloc::tests::kv_defrag_moves_the_bytes_and_opens_the_run ----
panicked at src/graph/alloc/tests.rs:662:5:
assertion `left == right` failed
  left: 3   right: 4

Three of the four moved tests go red — the required "at least one" with margin. The fourth (an_f16_session_round_trips_through_save_and_load) stays green under this mutation because its written is already raised to the asserted value by fill_batch_inputs's own_positions; only the owner-table half of its own_range call is load-bearing, and that half is what the copy-on-write test asserts. Mutation reverted; grep -rn MUTATION src/ empty and git diff clean of it (src/graph/kvcache.rs has no diff).

Counts (rule 5). No #[test] was added or removed, so no row moves and no CUDA row was re-measured: cargo test --release, box dgxspark (aarch64, GB10 sm_121), 2026-09-29 → 481 / 0 / 36 unit + 10 / 0 / 6 integration (the recorded CPU row). cargo fmt --all --check clean; python3 scripts/check_status.py --check exits 0; the non-test build warning-free with and without --features cuda (cargo build --release and cargo build --release --features cuda; the one warning: line on the CUDA build is build.rs's pre-existing cargo:warning= target list, not a rustc diagnostic).

Limits. (1) No CUDA device is used here: the --features cuda line is a compile, not a run; the CUDA unit row is unchanged by construction (no #[test] moved) and is not re-measured. (2) The x86_64 CPU row cannot be computed locally; it is unchanged by construction and the CI test-linux-cpu log is its source. (3) The moved call sites certify KvCache::own_range's owner-table and written-extent halves only as far as those tests assert them; the f16 session test does not depend on the mutated written arm.

Test-infrastructure record (#236, 2026-09-30) — GraphAllocator::kv_cell_of is test-only; production reads a sharer's rows as kv_map windows

The finding (S3 of #227). GraphAllocator::kv_cell_of (alloc.rs:1829 on master 691a9d2) was a one-line forwarder — self.kv.cell_of(seq, pos) — carrying #[cfg_attr(not(test), allow(dead_code))] and a doc that claimed a production role: "…this answers 'where would a reader look?', which is what a caller snapshotting a sharing sequence's rows needs." No such caller exists, and none was deleted: git log -S 'kv_cell_of' --all over src/ names exactly three commits — d43e716 (C8b S3 2/4, which introduced the forwarder and that doc), c9bbcbd (C8b S3 3/4, the test call) and 54f6de0 (the test-module extraction). Its only call site is server/batch/tests.rs:514 inside kv_rows_of (:502), the observation instrument of the C8b S3 gate a_store_inside_a_shared_prefix_takes_a_private_row (:385).

The consumer class the doc named is served two other ways — neither needs position→cell. (1) Reading a sharing sequence's rows for attention goes through windows: GraphAllocator::fill_seq_ids → KvCache::attn_map (kvcache.rs:803) → the kv_map input, which all three backends' attention kernels gather (C8b S2/S4; Metal's since #362). It returns runs, not per-position cells. (2) Whole-run snapshots go through the C5 container: BatchEngine::save_slots (server/batch.rs:442) calls kv_save_with_host (alloc.rs:2172), which stores the whole arena with the slot table as its host section — no position→cell mapping anywhere. C8b is closed (S1a–S5 landed 2026-09-21/22), so no pending slice was meant to consume the accessor. The kv field it forwards into is private (alloc.rs:174), which is why the tests cannot bypass the wrapper and call KvCache::cell_of directly from server::batch::tests; it is the forwarder, not KvCache::cell_of, that lacks a production caller.

The disposition (b): #[cfg(test)] pub(crate), in place. The annotation and the doc are replaced on the item itself (the pub(crate) fn is at alloc.rs:1840 after the change), matching #228's precedent (fill_attn_inputs_without_cells is #[cfg(test)]-scoped in place). The equivalent alternative — a #[cfg(test)] impl GraphAllocator in src/graph/alloc/tests.rs — was rejected: alloc/tests.rs is unit-test scaffolding for alloc, and this method's one consumer is in server::batch::tests, so putting the type's surface in a different module's test file would move the boundary without moving the caller. No call site changed — an inherent method resolves wherever its impl lives, and pub(crate) keeps the cross-module caller compiling. #[cfg(test)] is strictly stronger than allow(dead_code): the method does not exist in a non-test build, so a later cleanup cannot leave a production-looking orphan behind. KvCache::cell_of itself stays pub and production-used (the store resolver kv_cells_for_seq resolves through it); only the wrapper is gated.

Correction to an earlier record's wording. The #228 record lists "S3's kv_cell_of" among kv_cells_for_seq's non-test callers while it reports the premise it then falsifies. kv_cell_of forwards to KvCache::cell_of (kvcache.rs:512) and never called kv_cells_for_seq; the falsification that record reports — kv_cells_for_seq is production-reachable through fill_seq_ids — is unaffected, and as of this record kv_cell_of is itself test-only. The #228 text is left as written (it is a historical record); this is the dated forward correction for it.

Bar named before measuring. No test loses its assertion, and the driver stays sensitive: cargo test --release stays 481 / 0 / 36 unit + 10 / 0 / 6 integration, the ignored S3 gate is green on the unmutated tree, and a cell_of mutation must be visible to kv_rows_of.

Mutation evidence (rule 3), box dgxspark (aarch64, GB10 sm_121), 2026-09-30. Two mutations of the production-used KvCache::cell_of (kvcache.rs:512), each run as cargo test --release server::batch::tests::a_store_inside_a_shared_prefix_takes_a_private_row -- --ignored --nocapture, each reverted. Baseline (unmutated) run: 1 passed; 0 failed; 0 ignored; 516 filtered out.

(a) The span offset, cell + (pos - base) → + 1. The gate goes red — but at the answer arm, not through the instrument:

thread 'server::batch::tests::a_store_inside_a_shared_prefix_takes_a_private_row' panicked at
src/server/batch/tests.rs:492:5:
assertion `left == right` failed: the shared run must answer like the shape-matched copied one
  left: ".\nA. the a\nB."
 right: "\nA. Rome\nB. Naples"
test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 516 filtered out; finished in 1.50s

Every kv_rows_of-driven assertion passed under it, and that is structural, not luck: kv_rows_of resolves its cells through the same mutated cell_of the store uses, so a consistent offset cancels in a before/after snapshot (assert_eq!(kv_rows_of(donor), donor_before)) and in an A/B of two identical runs (assert_eq!(dst_a, dst_b)). The arm that caught it compares the sharing run's generated answer against the shape-matched copied run — generation reads through attn_map, which never calls cell_of, so the store's shifted cells surface as a wrong answer instead of a shifted snapshot. So the ticket's own example is evidence that the gate is sensitive, but not that the row-snapshot instrument is; it is recorded here rather than dropped, and the instrument is exercised by (b).

(b) The span cover, pos < base + len → pos + 1 < base + len — the last position of every span resolves to None. This one the instrument refuses itself:

thread 'server::batch::tests::a_store_inside_a_shared_prefix_takes_a_private_row' panicked at
src/server/batch/tests.rs:515:36:
no cell for sequence 2 position 2
test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 516 filtered out; finished in 0.60s

tests.rs:515 is kv_rows_of's own .unwrap_or_else(|| panic!("no cell for sequence {seq} position {p}")) — the failing assertion is the instrument, reached through kv_cell_of → KvCache::cell_of while it snapshots the diverging run's rows, before the answer arm can run. So at least one kv_rows_of-driven refusal fails under a mutation of the production cell_of, and the row snapshot is load-bearing. Both mutations reverted; git diff src/graph/kvcache.rs empty and grep -rn MUTATION src/ empty.

Counts (rule 5). No #[test] was added or removed, so no row moves and no CUDA row was re-measured: cargo test --release, box dgxspark (aarch64, GB10 sm_121), 2026-09-30 → 481 / 0 / 36 unit + 10 / 0 / 6 integration (the recorded CPU row). cargo fmt --all --check clean; python3 scripts/check_status.py --check exits 0; the non-test build warning-free with and without --features cuda (cargo build --release and cargo build --release --features cuda; the one warning: line on the CUDA build is build.rs's pre-existing cargo:warning= target list, not a rustc diagnostic).

Docs. AGENTS.md rule 1's last sentence no longer claims a production consumer (it now says test-only, names the kv_map/kv_save* production paths and the one consumer); the C8b S3 paragraph above carries a dated forward note; docs/COMPUTE-GRAPH-DESIGN.md was re-checked and never made the claim.

Limits. (1) No CUDA device is used here: the --features cuda line is a compile, not a run; the CUDA unit row is unchanged by construction and is not re-measured. (2) The x86_64 CPU row cannot be computed locally; it is unchanged by construction and the CI test-linux-cpu log is its source. (3) The S3 gate is #[ignore]d (it needs the cached 0.5B model), so the mutation transcripts come from -- --ignored, not from the default cargo test --release run. (4) The mutation that reaches the instrument is the span-cover one, not the span-offset one the ticket names; which arm each kills is stated above, and neither is presented as the other.

Test-infrastructure record (#238, 2026-10-01) — the bucket-C test-only wrappers become #[cfg(test)], in place

The ticket. #238 is T1 of the allow(dead_code) census. Its bucket C is "dead in the production --features cuda build, reached from test code, with at least one test caller outside the item's own module subtree". Those items cannot move into their module's tests.rs — that is bucket B / T2 (#239) — so the retirement is an explicit test-only scope at the item:

-    #[cfg_attr(not(test), allow(dead_code))]
-    pub fn foo(…)
+    /// Test-only (#238): driven by `…::tests::…`; `#[cfg(test)]` keeps it out of production builds.
+    #[cfg(test)]
+    pub(crate) fn foo(…)

#[cfg(test)] is strictly stronger than allow(dead_code): the item does not exist in a non-test build, so a later cleanup cannot leave a production-looking orphan behind (the #218 shape GATE-CONTRACT.md §1 asks about). The item stays where it is because its callers are in other modules: moving the surface into a different module's test file would relocate the boundary without relocating the caller — the same reasoning #236's record gives for kv_cell_of.

Reconciliation: the three tallies were 26, 71 and 56. Re-derived from the raw captures (minfer-allow-census/cuda.jsonl, tests_cuda.jsonl), taking every span of every dead_code diagnostic whose file starts with src/:

unitcount
distinct (file, line) dead sites, cargo check --release --features cuda314 (99 diagnostics)
… still dead under --tests --features cuda → bucket A (#242)142
… test-reachable (314 − 142) → 68 sites on B items + 71 on C items + 33 impl/mod header spans that rustc reports without an item identity172

At census item level (census.json, keyed by (file, line, name)) the cuda build has 260 dead items: A 121 (89 code / 32 shape), B 68 (58 / 10), C 71 (56 code / 15 shape). The three numbers the tracker carried are three different units, not three measurements:

  • 26 — the number of dead_code diagnostics whose primary span is a bucket-C item. The earlier analysis took each diagnostic's primary span only, so a grouped diagnostic collapsed a whole dead impl into one row. This record reproduces 26 exactly with that rule.
  • 71 — the item-level bucket-C count (cuda_C in census.json), i.e. 56 code + 15 shape.
  • 56 — the code-only C set, which is what T1 lists. tickets/T1.md carries exactly 56 item bullets (a naive `grep -c '^- `` reads 57: one Acceptance bullet also starts with a backtick).

Per-item verification, and what a re-read changed. Every item was read at its call site before it was touched; a name match is not a call site. The census's six recorded collision overrides (snapshot, plan, cuda_backend::new, submit, sample, source) all held. Six more were found here, and one of them changed a recorded call list without changing a bucket:

  • CudaBackend::kv_format vs GraphAllocator::kv_format / ModelDef::kv_format: the census's three recorded test callers for graph/cuda_backend.rs:314 are the other two (the registry hook is fn(&GraphAllocator), registry.rs:241). The real caller is graph/alloc/tests.rs:1526 (alloc.cuda().unwrap().kv_format()), so the item stays C.
  • optiming::record vs TimingSink::record: the census's optiming/tests.rs:142 caller is the sink's method; the free function's one caller is graph/scheduler/tests.rs:208.
  • GgufContext::get_val_bool vs GgufKv::get_val_bool: the four non-test matches (gguf.rs:770, :1752, :1849, main.rs:2398) are the inner type's method; the context's is test-only.
  • CpuBackend::pool_len vs BackendTrait::pool_len: GraphAllocator::pool_len_of calls the trait method; the inherent method's only caller is n_cpu_buffers.
  • Tensor::new vs Vec::new() / String::new() / OnceLock::new(): the census's ~200 "other callers" are all other types; the one real caller is graph/builder/tests.rs:6.
  • GraphAllocator::supports vs KvFormat::supports: kvformat::resolve calls the format's.

No item moved from C to A or C to B, but four of the 56 are not gated here, each with its reason written at the item (or, for the two (a) verdicts, in the tracker):

itemverdict
CudaState::matmul_f32_ptr (cuda.rs:3888)blocked: its caller CudaState::quant_matmul_f32_on_gpu (cuda.rs:5325) is bucket A and still compiled into production (part of #240's legacy wrapper layer). Gating the callee would leave the production CUDA build naming a function that does not exist. Deferred to #240/#242.
CpuBackend::pool_len (cpu_backend.rs:113)blocked: its only caller GraphAllocator::n_cpu_buffers (alloc.rs:2358) is bucket B (#239) and stays compiled, with its own allow, until T2 moves it. Deferred to #239.
ModelDef::as_any (models/mod.rs:211)(a) question: the doc says "downcast helper for the graph path's weight registration"; models::weight_reg never downcasts (its decision is a pure (ttype, geometry) predicate) and every as_any() call site under src/ is in a #[cfg(test)] module. A production-looking trait method whose documented caller does not exist — reported, not silenced.
ModelDef::offload (models/mod.rs:294)(a) question: the doc (and the qwen2/qwen3 module docs) says it "hands the plan to the graph builder", but the builder reads model.offload.plan (models/qwen2/graph.rs:556) and nothing outside tests calls the accessor. Same class, same disposition.

Both (a) items are defaulted/required trait methods, so #[cfg(test)] would also change the trait's public surface and (for as_any) both impls — a decision for #244's shape/API class rather than a mechanical retirement.

Stale docs corrected at the item. CudaState::stream_is_capturing and its FFI declaration cudaStreamIsCapturing both claimed a "registration path's refusal / inventory" consumer. No such caller exists: the production capture bookkeeping is the per-instance CudaBackend::capturing field, and the pair's only caller is the #188 probe in graph/cuda_backend/tests.rs. Their docs now say so.

Bar named before measuring. No #[test] is added or removed, and no assertion moves — the tests keep their own assertions and only the resolution of the item changes. The bar: the non-test build stays warning-free with and without --features cuda (the crate carries #![cfg_attr(not(test), deny(warnings))], so an exposed callee or a broken call site is a hard failure), the suite counts stay where the recorded rows put them, and one mutation per module group must be visible to the test that drives the scoped item.

Mutation evidence (rule 3), box dgxspark (aarch64, GB10 sm_121), 2026-10-01. One representative item per module group; each mutation applied, the named test run with --exact, then reverted and git status --porcelain <file> empty. All 16 went red:

group (item)mutationtest (all FAILED)first failing line
conversation (Conversation::start)prefill_tokens = toks.len() + 1conversation::tests::first_turn_full_render_and_eogtests.rs:209: left 22 / right 21
cuda (format_of)KV_LAYOUT_F16 => F32cuda::kv_dtype_tests::the_layout_tag_is_the_format_discriminantkv_dtype_tests.rs:14: left F32 / right F16
gguf (get_arr_n)get_ne() + 1gguf_write::tests::every_metadata_type_round_trips_through_the_parsertests.rs:58: left 3 / right 2
grammar (accepts)Err(_) => truegrammar::tests::gbnf_literals_classes_and_dottests.rs:83
graph/alloc (get_buffer)if true { return None }graph::alloc::tests::fill_and_read_inputtests.rs:356
graph/builder (swiglu)inputs &[up, gate]graph::builder::tests::swiglu_builder_and_metatests.rs:91: left [1, 0] / right [0, 1]
graph/cache (stats)builds + 1graph::cache::tests::switching_between_cached_graphs_re_maps_instead_of_rebuildingtests.rs:96: left (3, 0) / right (2, 0)
graph/copystats (delta)copies: self.copiesgraph::copystats::tests::the_two_phases_are_counted_separately_and_the_delta_is_exacttests.rs:28: left copies 5 / right 2
graph/cpu_backend (causal_span)(0, p) instead of (0, p + 1)graph::cuda_backend::tests::cuda_rope_kv_attn_roundtrip (CUDA)graph/cuda_backend/tests.rs:689
graph/cuda_backend (stream_sync_count)return 0graph::cuda_backend::tests::stream_sync_counts_are_per_backend_not_process_wide (CUDA)tests.rs:8317: left 0 / right 1
graph/kvcache (GraphAllocator::kv_clear_identity)no-opgraph::scheduler::tests::a_non_identity_kv_mapping_is_refusedtests.rs:85
graph/kvformat (pack_q8_0_cell)dst *= 2.0 after packinggraph::kvformat::tests::a_packed_cell_round_trips_within_the_q8_0_block_errortests.rs:172
graph/registry (is_unfiltered)all → anygraph::registry::tests::the_name_surface_fences_devices_and_keeps_cputests.rs:186
optiming (record)drop the recordgraph::scheduler::tests::a_concurrent_graph_load_cannot_move_a_private_sinktests.rs:242: left 256 / right 257
tensor (Tensor::from_data)zero the payloadgraph::op_matrix::matrix_cases_match_their_referenceop_matrix.rs:800
testfail (checked)`map_or(0,_1)`

The full suite run is the baseline these deltas are read against (below). The CUDA rows ran through the device build; the rest are CPU rows.

Counts (rule 5). No #[test] was added or removed, so no row moves and no CUDA row is re-measured: cargo test --release, box dgxspark (aarch64, GB10 sm_121), 2026-10-01 → 481 / 0 / 36 unit + 10 / 0 / 6 integration (the recorded CPU row, unchanged); bash scripts/cuda_test.sh, same box, 2026-10-01 → 566 / 0 / 42 (the recorded CUDA row, unchanged; the suite was run to confirm the device build still links the #[cfg(test)] items it now gates). cargo check --release and cargo check --release --features cuda both exit 0 with no rustc diagnostic (the one CUDA warning: line is build.rs's pre-existing cargo:warning= target list); cargo fmt --all --check clean; python3 scripts/check_status.py --check exits 0.

Docs. The #235 family issue gets a comment recording which of its four wrappers this ticket closed (kv_clear_identity, kv_shift, kv_n_used — kv_cells_for is bucket A and stays with #242); #238 gets the reconciliation, the four deferrals and the patch that brings its body in line; #239/#240/#242 keep their rows.

Limits. (1) The macOS-only modules (metal.rs, the Metal half of graph/metal_backend.rs) are not compiled on Linux, so their annotations are unchanged and unjudged — the macOS CI job is the only gate (src/models/mod.rs's as_any is reached from metal/mmap_align_test.rs, which is why it appears in this bucket at all). (2) The item-level classification still rests on textual test-caller resolution for the module of a caller (liveness is rustc's); the six new collisions above are the residue, each now verified by reading. (3) Tensor::new's only consumer is the f32_tensor helper in graph/builder/tests.rs, whose assertions read fields the helper sets itself (name, shape) — no mutation of new's body is observable through it, so the tensor group's mutation is on Tensor::from_data (same impl Tensor, same file), and that is a real gap in the helper's coverage, recorded rather than papered over. (4) --features debug_dump was not built by the census or here.

Test-infrastructure record (#239, 2026-10-01) — the same-module test-only items move into their module's tests.rs

The ticket. #239 is T2 of the allow(dead_code) census. Its bucket B is "dead in the production --features cuda build, reached from test code, and every test caller sits inside the item's own module subtree". That is what makes the retirement a move rather than #238's in-place gate: an inherent method goes into a #[cfg(test)] impl Type block in that module's tests.rs (an inherent impl may live in a child module — the cuda build proves it), a free function/const/FFI declaration goes in as a #[cfg(test)] item there. After the move nothing in a non-test build knows the item exists, which is strictly stronger than #[cfg(test)] pub(crate) in place: there is no production-file line left for a later cleanup to orphan.

Reconciliation: the ticket's "26" is not a B unit. Re-derived from the raw captures (minfer-allow-census/cuda.jsonl, tests_cuda.jsonl) with the same rule T1 used — every span of every dead_code diagnostic whose file starts with src/:

unitcount
distinct (file, line) dead sites, cargo check --release --features cuda314 (99 diagnostics)
… still dead under --tests --features cuda → bucket A (#242)142
… test-reachable (314 − 142) → 68 sites on B items + 71 on C items + 33 impl/mod header spans rustc reports without an item identity172

At census item level (census.json, keyed by (file, line, name)) the cuda build has 260 dead items: A 121 (89 code / 32 shape), B 68 (58 / 10), C 71 (56 / 15). The three numbers the tracker circulated are therefore:

  • 68 — the item-level bucket-B count (cuda_B), i.e. 58 code + 10 shape.
  • 58 — the code-only B set, which is T2's list. tickets/T2.md carries exactly 58 item bullets (a naive grep -c '^- \'reads 59 because one Method/Sentence line also matches);census.json` and the list agree name for name.
  • 26 — does not reproduce for B, and is not a third unit. Re-deriving it the only way it can be read (how many of the 99 diagnostics have their primary span on a bucket-B item) gives 45 (42 code + 3 shape), not 26. 26 reproduces exactly as the number of diagnostics whose primary span is a bucket-C item — which is T1's headline unit (#238). The ticket body's "26 (diagnostics whose primary span is a bucket item)" is a mis-attribution of T1's number to T2, and its body is patched to say so. (For completeness, the README's collapsed B = 28 reproduces as 27 diagnostics whose every src/ span is a B item; the one-diagnostic difference is a diagnostic that also carries an impl header span.)

Per-item verification, and what a re-read changed. Every item was read at its call site before it was touched. The census's collision overrides all held for these items (snapshot, plan, cuda_backend::new, submit, sample were checked again: the recorded cross-module matches really are other items — TimingSink::snapshot, OffloadRequest::plan, Vec::new, the Metal CommandBuffer::submit, server::metrics' own sample). Five items did not survive as a plain move:

itemverdict
KvCache::set_owner (kvcache.rs:322)(a) question: its doc said "C1 uses it from own_range", but own_range (kvcache.rs:517) writes l.owner[cell] = seq inline and never calls it. A helper whose documented production caller does not exist is the #218 shape — reported, not moved (the annotation stays), for #244.
OffloadPlan::all_on_device (offload.rs:44)(a) question: its doc said it is "what an unset request resolves to when a device is available", but OffloadRequest::plan (offload.rs:187) and resolve (offload.rs:232) build OffloadPlan { gpu_layers, n_layers } inline (they must — they clamp first). Same class, same disposition.
StreamScratch::slot (cuda.rs:1300)blocked: bucket B by the census (its only test caller is cuda::d35_probe_tests), but the bucket-A legacy wrappers upload_hidden / upload_positions / download_logits / get_positions_buf still call it and are still compiled, so it cannot leave the production file. Deferred to #240 / #242, with the reason at the item.
device_entry.rs's 4 items (DEVICE_ENTRY, DeviceHolder, DeviceEntry, enter)blocked: enter is still called by the bucket-A CudaState::layer_gpu (cuda.rs:7051), which is compiled in the cuda build. Deferred to #241, which deletes the guard.
CudaBackend::new (cuda_backend.rs:200)reclassified B→C, and still moved: the census calls it B because is_test_path only recognises tests.rs / *_tests.rs, but it also has a cross-module test caller, graph::op_matrix (a #[cfg(test)] mod). An inherent method's impl may live in a child module, so moving it into cuda_backend/tests.rs with pub(crate) keeps op_matrix compiling while production loses the constructor — verified by the --features cuda check and the device suite.

Three further stale doc claims were found on items that were moved (their own docs do not assert a current consumer, so they are not (a) items) and are recorded at the item and in the tracker: GraphAllocator::set_memory_budget's "a future offload policy uses it" (the E5 S2 policy landed and resolves its budget through MINFER_GPU_MEM + allocplan::weight_budget), KvCache::note_written's // E2 surface note (E2 records written extents through own_positions / fill_batch_inputs), and allocplan.rs's module doc ("the plan is checked against a per-backend budget before that loop runs" — the gate is alloc_in_pool, not AllocPlan::plan). The module doc keeps its wording and gains an (a) note rather than being quietly rewritten.

The moves: 51 items, 18 module groups. One item per line stayed free; the map is

module (tests.rs)moved itemsvisibility
conversationConversation::snapshotprivate
cuda (new src/cuda/tests.rs)13 test-only FFI declarations (cuda_test_latch_oversized_smem, minfer_site_fail_* ×8, minfer_site_hist_{len,site,name,msg}), latched_api_error_count, CudaState::take_last_errorpub(super) — the callers are sibling *_tests modules under cuda
device_tierIDENTITY_BATCH_BOUND, mmvq_batch_limit, mmvq_capprivate (the const is a standalone item, not a field/variant, so it can move; mmvq_cap was its only production-side reader)
grammarGrammarState::pending_bytesprivate
graph/allocset_memory_budget, n_cpu_buffers, n_mapped_buffersprivate
graph/allocplanAllocPlan, plan, live_peak, DeviceMemory::free_bytesprivate
graph/cachecached_graphsprivate
graph/cpu_backendweightprivate
graph/cuda_backendnew, exec_ids + elems (the #[cfg(test)] shim and its helper)pub(crate) on new (the op_matrix caller), private otherwise
graph/kvcacheown_prefix, note_written, after_shift, seq_range, private_writtenprivate
graph/modDType::sizeprivate
graph/registrynamesprivate
graph/schedulerwith_timingprivate
samplerapply_repetition_penalty, sample_with_penalties, sampleprivate (sample_with_penalties has no test caller of its own; it is reached only through sample, so it moved with it)
server/batchidle_slots, submit, prefill_stats, prefill_fed, interleaved_ticksprivate
server/metricscompletion_tokens_per_secondprivate
templatePyValue::to_jsonprivate
tokenizerPreTokenizer::gguf_name, Tokenizer::decodeprivate

The visibility rule: private when the only caller is that one tests.rs; pub(super) when the callers are several test files under the same module (cuda); pub(crate) when a #[cfg(test)] module elsewhere in the crate calls it (CudaBackend::new from graph::op_matrix). Each moved item carries a doc note naming the test that drives it (machine-checked against the named test's module and file) and the sentence "Test-only (#239)".

T1's pool_len deferral is closed. With GraphAllocator::n_cpu_buffers now in graph/alloc/tests.rs, CpuBackend::pool_len has no production caller left, so it became #[cfg(test)] pub(in crate::graph) — the narrowest spelling that reaches graph::alloc::tests (pub(crate) would also work; the narrower one is preferred). #238 gets a comment saying so. T1's other deferral, CudaState::matmul_f32_ptr, stays with #240/#242: its caller quant_matmul_f32_on_gpu is still compiled.

Bar named before measuring. No #[test] is added or removed and no assertion changes — the moved items are byte-identical bodies, so the claim each test makes is unchanged by construction. The bar: the non-test build stays warning-free with and without --features cuda (the crate carries #![cfg_attr(not(test), deny(warnings))]), the recorded suite counts do not move, and one mutation per module group must be visible to the test that drives the moved item.

Mutation evidence (rule 3), box dgxspark (aarch64, GB10 sm_121), 2026-10-01. One representative item per module group that actually moved; each mutation applied, the named test run, then reverted and the file verified byte-identical. All 18 went red:

group (item)mutationtest (all FAILED)first failing line
conversation (Conversation::snapshot)messages: Vec::new()conversation::tests::a_resumed_snapshot_prefills_nothing_and_continues_aliketests.rs:301
cuda (latched_api_error_count)return 0 (MINFER_TEST_LATCH_ERROR=1)cuda::issue145_tests::cuda_sync_surfaces_a_latched_error_as_latched (CUDA)issue145_tests.rs:230
device_tier (mmvq_cap)drop the .min(IDENTITY_BATCH_BOUND)device_tier::tests::caps_clamp_to_the_identity_boundtests.rs:85
grammar (pending_bytes)return 0grammar::tests::token_advancement_handles_partial_utf8tests.rs:514
graph/alloc (pool_len, through the moved n_cpu_buffers)return 0graph::alloc::tests::a_rebuild_remaps_instead_of_reallocatingtests.rs:1607
graph/allocplan (free_bytes)Some(*free + 1)graph::allocplan::tests::a_reported_free_read_keeps_the_three_quarters_defaulttests.rs:51
graph/cache (cached_graphs)return 0graph::cache::tests::switching_between_cached_graphs_re_maps_instead_of_rebuildingtests.rs:97
graph/cpu_backend (weight)return Nonegraph::cpu_backend::tests::f32_matmul_nt2_token_majortests.rs:249
graph/cuda_backend (new)return Nonegraph::cuda_backend::tests::cuda_pool_roundtrip (CUDA)tests.rs:108
graph/kvcache (seq_range)Ok(None)graph::kvcache::tests::two_sequences_resolve_to_disjoint_windowstests.rs:151
graph/mod (DType::size)F32 => 8graph::tests::dtype_sizetests.rs:101
graph/registry (names)return ["metal", "cpu", "cuda"]graph::registry::tests::names_resolve_and_unknown_names_are_refusedtests.rs:111
graph/scheduler (with_timing)force TimingMode::Globalgraph::scheduler::tests::a_concurrent_graph_load_cannot_move_a_private_sinktests.rs:230
sampler (apply_repetition_penalty)no-opsampler::tests::test_repeat_penalty_reduces_repeatedtests.rs:18
server/batch (idle_slots)return 0server::batch::tests::a_failed_decode_forward_answers_a_single_run_and_releases_its_slottests.rs:1955
server/metrics (completion_tokens_per_second)return 0.0server::metrics::tests::token_counters_and_the_trailing_ratetests.rs:277
template (PyValue::to_json)encode an int as a stringtemplate::tests::python_str_methods_match_cpythontests.rs:263
tokenizer (Tokenizer::decode)return String::new()tokenizer::tests::decode_bytes_reverses_byte_encodingtests.rs:65

The offload group has no row: its only bucket-B item, all_on_device, is an (a) finding and was not moved.

Two coverage gaps, recorded rather than papered over. (1) allocplan::live_peak's result is not observable at all: AllocPlan::plan computes peak in its main loop (that loop never subtracts, so peak == reserved_bytes on exit) and then peak.max(live_peak(…)) cannot exceed it — mutating live += class_bytes(classes[i]) to live += 0 left the_live_peak_is_not_the_reserved_total green. The group's mutation is therefore on a sibling in the same move, DeviceMemory::free_bytes. This is a real gap in live_peak's coverage and also a hint that the function is dead arithmetic — for #244 to decide. (2) pool_len's value is only ever compared with itself in a_rebuild_remaps_instead_of_reallocating, so a constant offset (+1) is invisible; the observable mutation is 0, which fails that test's buffers > 0 start assertion.

Counts (rule 5). No #[test] was added or removed, so no row moves: cargo test --release, box dgxspark (aarch64, GB10 sm_121), 2026-10-01 → 481 / 0 / 36 unit + 10 / 0 / 6 integration (the recorded CPU row); bash scripts/cuda_test.sh, same box, 2026-10-01 → 566 / 0 / 42 (the recorded CUDA row — run to confirm the device build still resolves the moved FFI declarations, the moved CudaBackend::new and the #[cfg(test)] extern "C" block the new cuda/tests.rs introduces). cargo check --release and cargo check --release --features cuda both exit 0 with no rustc diagnostic (the one CUDA warning: line is build.rs's pre-existing cargo:warning= target list); cargo fmt --all --check clean; check_status.py --check, check_docs_links.py and check_source_layout.py all clean (the new src/cuda/tests.rs is named by its #[cfg(test)] mod tests;).

Docs. #239 gets the reconciliation, the five non-moves and the mutation table as a comment and its body is patched (the 26 → 45/68/58 reconciliation); #238 gets the pool_len closure; #240/#241/#242 keep the two blocked groups; #244 gets the two (a) items and the three stale doc claims.

Limits. (1) The macOS-only modules (metal.rs, the Metal half of graph/metal_backend.rs) are not compiled on Linux and were not touched — the macOS CI job is their only gate. (2) The item-level B/C classification rests on textual resolution of a caller's module (liveness is rustc's); the one residue found here (CudaBackend::new, called from the #[cfg(test)] mod op_matrix) is exactly the limit, and it was handled by moving with pub(crate). (3) --features debug_dump was not built. (4) The two (a) verdicts are questions, not findings of fact about intent: in both cases the doc may simply be stale rather than a caller having been deleted, and either resolution (call it, or delete it, or say test-only) belongs to #244, not here.

Test-infrastructure record (#240/#241, 2026-10-01) — the dead legacy CudaState wrapper layer and the device_entry guard are deleted

The tickets. #240 is T3a of the allow(dead_code) census: delete the legacy CudaState wrapper layer — 23 methods and 18 backing fields that no build configuration calls. #241 is T3b: the legacy device_entry guard is unreachable and #185's remaining token should be retired. They landed as one PR because they are compile-coupled, not because they were convenient to batch.

Why one PR: the coupling is a compile error, not a preference. CudaState::layer_gpu is the only caller of has_weight, debug_sync, kv_ensure_layer, get_positions_buf, matmul_on_gpu, quant_matmul_q8 and quant_matmul_f32_on_gpu, and the only reader of fifteen buf_*/kv_* fields. Deleting them while layer_gpu stays is E0599. The reverse is also true: src/device_entry.rs opened with #![cfg_attr(not(feature = "cuda"), allow(dead_code))], so in the cuda build there is no allow — with layer_gpu gone, enter/DeviceEntry/DEVICE_ENTRY have no caller and deny(warnings) fails the build. A partial #240 that "skips layer_gpu" can therefore only delete the sixteen uncoupled leaf items and cannot close either of the deferrals T1/T2 left here; that split was rejected after the coupling was traced, and the user approved deleting the guard, whose only caller is itself dead.

Per-item disposition. Every item was read at its definition, its cross-module matches checked, and its history read (git log -S) before it was touched.

itemdisposition
stream_lock, clear_launch_failure, upload_hidden, download_hidden, upload_positions, init_kv_cache, get_kv_size, download_logits, graph_available, graph_end_capture, graph_launch, quant_matmul_f32_batch, quant_matmul_f32, output_norm_gpudeleted (uncoupled leaf; no caller in any build)
buf_logits (readers: download_logits/quant_matmul_f32/output_norm_gpu), decode_graph_exec (graph_available/graph_end_capture/graph_launch)deleted with their initializers
layer_gpu (bucket A; dead in every build)deleted (#241's item, absorbed by the coupling above)
has_weight, debug_sync, kv_ensure_layer, get_positions_buf, matmul_on_gpu, quant_matmul_q8, quant_matmul_f32_on_gpudeleted once layer_gpu went (it was their only caller)
buf_hidden, buf_bn, buf_bq, buf_bk, buf_bv, buf_ba, buf_bf, buf_bg, buf_q8_bn, buf_q8_ba, buf_positions, kv_k, kv_v, kv_sizedeleted with their initializers (only layer_gpu read them)
tierdeleted: never read from self; the name-level matches are device_tier::Selection's own tier in the two select arms (the collision the census note flags). Its doc said "Direct consumers arrive with the batch-cap activation (plan §14 R8)", which is the pattern the plan's own A7 note rejects — "a future feature will need it" is not enough. The effective gate production reads is tier_mmq; the R8 work re-adds the field from the selection it already computes.
cckept, reported: read by #[cfg(test)] CudaState::cc() (driven from graph/cuda_backend/tests.rs), so it is test-reachable and not a deletion. Its non-test-dead annotation is #243's tightening, the same shape T1/T2 handled for their items.
stream_wait_event / the FFI declaration cudaStreamWaitEventkept: the device→device staging copy it is the mechanism for. #138 landed 2026-10-04 (the F5 S2 record) without creating a reachable caller: the only pair that could express a device destination needs two device backends, copy_across early-returns on a same-backend pair, and Metal declines phase A. So the item is still dead in every compilable configuration and the census brief requires it be named.
src/device_entry.rs (DEVICE_ENTRY, DeviceHolder, DeviceEntry, enter) + its mod declaration and its test filedeleted: the guard exists for a caller that does not exist (#188 had already narrowed it to layer_gpu).
cuda_debug_enabled + the CUDA_DEBUG OnceLocknewly dead, identified here and deleted in the same commit: debug_sync was its only reader.

Two doc claims were read because a doc that names a production caller is the #218 lost-caller shape: clear_launch_failure's Err-arm claim (graph/cuda_backend.rs::execute_node drains take_launch_failure() on both arms instead — the doc was stale) and init_kv_cache's "must be called before the first forward pass" (superseded by the graph path, which owns KV through GraphAllocator::kv_pair). Both items are deleted; no lost caller was found, so nothing is reported as (a) for these two. No extern "C"/#[no_mangle] item is in the set and none of the names is a symbol in src/cuda_kernels.cu or build.rs; nothing memcpys a CudaState.

The two deferrals this closes. StreamScratch::slot (#239's deferral, pinned by the four bucket-A wrappers upload_hidden/upload_positions/download_logits/get_positions_buf) now lives in src/cuda/tests.rs as pub(super), the visibility the rest of that file uses for the sibling cuda::*_tests modules. CudaState::matmul_f32_ptr (#238's deferral, pinned by quant_matmul_f32_on_gpu) is now #[cfg(test)] pub(crate) — its callers are the device gates in graph::cuda_backend::tests, a #[cfg(test)] module in another file, so T1's in-place form (not T2's move) is the correct one.

The one intentional coverage decrease. src/device_entry/tests.rs held a single test, device_entry::tests::the_legacy_unbound_path_is_exclusive_across_threads_and_re_entrant_on_one, which asserted the guard's cross-thread exclusion and per-thread re-entrancy with mutation evidence. It is deleted with the guard, by decision (recorded on #241): the guard's only caller is dead, so there is no engine property left for the assertions to be about. This is the only place a test is removed in this ticket, and no other test loses an assertion. It is why the CPU rows move 481 → 480 (aarch64) and 479 → 478 (CI), and the CUDA rows 566 → 565.

Evidence form: the census re-run (this ticket's version of mutation evidence). There is no gate to mutate — the deleted code has no caller, so no test observes it, and "break the implementation and watch a gate go red" has nothing to break. For a deletion ticket the observable claim is instead "the dead set shrank by exactly what was removed, and nothing else became dead", and that is measured. Method, identical on both sides: strip.py in a scratch worktree (every allow(dead_code) replaced by a //STRIPPED comment, line numbers preserved), then cargo check --release --features cuda --message-format=json (the maximal configuration on Linux, because src/cuda.rs is #[cfg(feature = "cuda")] mod), counting every src/ span of every dead_code diagnostic. Liveness is rustc's, never a name grep.

treedead src/ sitesdead itemssrc/cuda.rs sites
e9bf4c5 — the durable census' recorded baseline (reproduced here from minfer-allow-census/cuda.jsonl)31426053
79c9837 — this branch's parent19115653
dc550a9 — after the uncoupled deletions17514037
2fbaf00 — the final tree1431097

The drop from the parent is 48 sites / 47 items, and the before/after item-name diff has an empty "new" side: zero newly-dead items. The 314 → 191 drop between the recorded baseline and this branch's parent is not this ticket: #238/#239 turned bucket-C items into #[cfg(test)] and moved bucket-B items into tests.rs, and an item that is not compiled into a non-test build leaves the dead set entirely. Truncation check: both captures re-run under RUSTFLAGS=--cap-lints=warn differ from the deny(warnings) captures by symmetric difference 0 (cpu and cuda), so deny(warnings) did not truncate the count.

The seven src/cuda.rs sites left are outside this ticket's list and are named, not silently left: stream_wait_event + its FFI declaration cudaStreamWaitEvent (#138), cc (#243), and cudaGetErrorString/cuda_error_string, minfer_site_hist_reset, stream_sync_count — dead before this ticket too (they are in the baseline set, not the diff), so they belong to #242/#243/#244's sweep, not to #240.

Bar named before measuring. No #[test] is added or removed except the guard's own, removed by decision (above), so no assertion outside it changes. The bar: the non-test build stays warning-free with and without --features cuda and without adding a single allow(dead_code); the suite rows move by exactly the one removed test (480/0/36 + 10/0/6 CPU, 565/0/42 CUDA); and the census dead-site count drops by at least the number of items deleted, with zero newly-dead sites.

Counts (rule 5), box dgxspark (aarch64, GB10 sm_121), 2026-10-01. cargo test --release → 480 / 0 / 36 unit + 10 / 0 / 6 integration. bash scripts/cuda_test.sh → 565 / 0 / 42. compute-sanitizer --tool memcheck --target-processes all <test binary> --test-threads=1 → 0 API errors over 565 / 0 / 42. cargo check --release, cargo check --release --features cuda and cargo check --release --tests --features cuda all exit 0; the only warning: line is build.rs's pre-existing cargo:warning= target list. cargo fmt --all --check clean; check_status.py --check, check_docs_links.py and check_source_layout.py clean (the deleted src/device_entry/tests.rs and the mod device_entry; line go together, so no .rs is left undeclared). The four rows that share the removed test are updated together: docs/status.toml's CPU rows (481 → 480, 479 → 478), the CUDA unit row (566 → 565) and the sanitizer row (566 → 565), each with its projection_base_passed; AGENTS.md carries the same numbers and the changelog sentence.

Docs. #240 is closed with the evidence comment; #241 carries the absorption and the deletion (not silent); #239 gets the slot closure, #238 the matmul_f32_ptr closure; #244 gets cc's annotation and the two stale-doc findings as decisions; #243 gets cc; #185 is closed with the guard's removal and what remains of its question. The live doc claims that named the removed knob (MINFER_CUDA_DEBUG in docs/CUDA-BACKEND-DESIGN.md, the walkthrough and the tutorial) and the #189 record's sentence about the guard covering layer_gpu are updated in place; the historical #185/#188 records keep their wording.

Limits. (1) The macOS-only modules were not touched — the build-macos CI job is their only gate. (2) The counted census is a cargo check: a "caller" is a compile-time reference, not a verified runtime exercise; the suites above are the independent runtime check. (3) --features debug_dump was not built, as in the census itself. (4) The real-model gate sets were not re-run here (they need the cached models and were unaffected — no test in that set was touched); their rows are unchanged and still dated 2026-09-27. (5) cc's tightening and the residual dead items are named above and left to #242/#243/#244.

Test-infrastructure record (#242, 2026-10-01) — the remaining no-caller items are deleted

The ticket. #242 is T3c of the allow(dead_code) census: bucket A — an item rustc reports dead in the maximal build (--release --features cuda) that is still dead with --tests --features cuda, i.e. nothing reaches it at all. It is the last of the census' code items after #238/#239 (buckets B/C, test-only) and #240/#241 (the legacy CUDA wrapper layer and the device_entry guard), and it closes #235's last row (kv_cells_for).

Reconciled numbers — the tracker was patched to this run. The ticket body carried the e9bf4c5 item list, so it was re-derived and the body rewritten:

treedead src/ sitesdead itemsbucket A codebucket A shape
e9bf4c5 — the durable census' recorded baseline3142608932
79c9837 — after #238/#239191156——
5147d2c — this branch's parent (after #240/#241)1431096515
8d3f54a — this branch's final tree7147315

Method, identical on both sides: strip.py in a scratch worktree (every allow(dead_code) replaced by a //STRIPPED comment, line numbers preserved), then cargo check --release --features cuda --message-format=json, counting every src/ span of every dead_code diagnostic — rustc groups a dead impl into one diagnostic, so a primary-span-only read undercounts ~2.6×. The --tests --features cuda capture is the A-vs-B/C discriminator, and the --cap-lints=warn twins prove the deny(warnings) capture was not truncated (symmetric difference 0 on both cpu and cuda). The branch deletes 62 code items; the counted drop is 72 sites / 62 items, exactly the number deleted.

Per-file disposition (62 census items; every item read at its definition, its cross-module name matches read rather than name-matched, and git log -S read where its doc claimed a caller).

fileitemswhat was deleted and why it was safely dead
src/cache.rs6KVCacheLayer::{store, get_k, get_v, clear, store_multi} and KVCache::clear — the graph path owns KV in the allocator; the type survives only as forward's vestigial &mut KVCache argument (its fields are the shape items below)
src/cuda.rs4the cudaGetErrorString FFI declaration + cuda_error_string (production names errors through cuda_error_name), the dead minfer_site_hist_reset FFI declaration (the .cu defines and calls its own), and the process-wide stream_sync_count with its STREAM_SYNCS static and the increment in CudaState::sync (#185 moved every gate to the backend's own counter)
src/gguf.rs22GgufType::type_name; GgufContext::{get_version, get_alignment, get_kv_type, get_val_u8…get_val_f64, get_val_str, get_val_data, get_n_tensors, find_tensor, get_tensor_offset, get_tensor_name, get_tensor_type, get_tensor_size, dump_metadata} (main.rs has its own dumper). The surviving accessors are production- or test-reached, so the impl GgufContext #[allow(dead_code)] went with them
src/graph/mod.rs2ComputeGraph::{node_mut, n_elements}
src/graph/alloc.rs2GraphAllocator::{kv_cells_for, reset_cross_stats}
src/graph/backend.rs1 (+3 impls)the Backend::name trait method and its cpu/cuda/metal impls — the live name surface is registry::Backend::name(self), the handle's own method
src/graph/params.rs1next_weights_version (the GraphParams::weights_version field stays in the reuse identity)
src/models/mod.rs1ModelDef::format_chat
src/models/qwen2/mod.rs1 (+1 impl)format_chatml and its format_chat impl
src/models/qwen3/mod.rs2 (+1 impl)format_chatml, the inherent Qwen3Model::n_layer, and the format_chat impl
src/tensor.rs20from_data_with_strides, nelements, nrows, ncols, data_mut, data_f32_mut, the eight data_q*/data_q*_mut byte accessors, get_f32, set_f32, copy_from, reshape; with them the impl Tensor allow went
src/server/batch/tests.rs—the test mock's format_chat (forced by the trait change; it was an unreachable!(), so no assertion changes)

Named keeps, not silent leftovers. cudaStreamWaitEvent + CudaState::stream_wait_event stay reserved for the F5 device→device staging copy, which #138 (landed 2026-10-04) found unreachable — see the F5 S2 record. ModelDef::forward_graph stays because its only caller is models::qwen2::graph::tests's #[cfg(target_os = "macos")] block (tests.rs:3021): the Linux census cannot compile it, cargo test on a Mac would not compile if the method were deleted, and CI's macOS job (cargo build --release) would not catch the break. That is this ticket's one platform-cfg blind spot in bucket A, found by reading the call site rather than by the oracle. The 15 bucket-A shape items (KVCacheLayer's fields, Vendor::{Amd,Mthreads, Apple}, AttnMode::Mha, SampledToken::logit, Slot::id, Tokenizer::{id_to_score,id_to_type}, HfCheckpoint::order) are #244's decisions; cc's annotation (2 of the 5 residual src/cuda.rs sites) is #243's.

No newly-dead items. The before/after diff keyed on (name, kind, container) has an empty "new" side: 0 newly-dead. The after-set is a strict subset of the before-set, so deletion unmasked no transitive dead code — the impl GgufContext and impl Tensor allow-removals were the two places that could have exposed some, and neither did.

Findings — every doc that named a caller was stale; no lost caller. GgufType::type_name ("only exercised by tests/debug tooling today") had no test caller, only the dead dump_metadata; ComputeGraph::node_mut ("fusion pass / debug tooling") is never called by the fusion pass; ComputeGraph::n_elements ("assertion / debug helper") by no assertion; kv_cells_for ("the resolver's C2 consumers are the backends") names a wrapper C2 never used — the backends read the resolved cells input through kv_cells_for_seq (#235); reset_cross_stats ("a gate that wants an absolute number") by no gate — every F5 gate reads cross_stats().delta(before); next_weights_version ("Phase 6 wires the model to bump it") was introduced by the Phase-4 commit and never gained a caller (git log -S), i.e. the A7 "a future feature will need it" pattern; cuda_error_string ("for diagnostics that want the prose form") by no diagnostic after #122; Backend::name ("part of the Backend API surface") by nothing — all .name() calls are on the registry handle; ModelDef::format_chat/format_chatml ("kept as the fallback implementation") by nothing — template.rs is the path; Qwen3Model::n_layer ("stays as a concrete-type helper") by no concrete-type caller; the impl Tensor "complete ggml_tensor interface (used by tests / debug tooling)" note and the impl GgufContext "public raw API surface (tests / debug tooling)" note likewise. None is the #218 lost-caller shape — no earlier cleanup removed a caller that this ticket then deleted — so nothing here is referred to #244 as a decision; the two bucket-B doc-vs-caller items #239 reported there (KvCache::set_owner, OffloadPlan::all_on_device) are outside this ticket's bucket and untouched.

Evidence form: the census re-run (this ticket's version of mutation evidence, as in #240). There is no gate to mutate — the deleted code has no caller, so no test observes it. The observable claim is instead "the dead set shrank by exactly what was removed and nothing else became dead", and that is measured on both sides (table above): sites 143 → 71, items 109 → 47, the item drop equal to the 62 deleted, and a before/after item diff whose "new" side is empty. The one test-adjacent edit is stream_sync_counts_are_per_backend_not_process_wide's mutation note, which named the deleted process-wide counter as its mutation; it now names the still-available mutation (bump a process-wide static in state_sync and read it) and the test's two assertions are unchanged.

Counts (rule 5), box dgxspark (aarch64, GB10 sm_121), 2026-10-01. cargo test --release → 480 / 0 / 36 unit + 10 / 0 / 6 integration. bash scripts/cuda_test.sh → 565 / 0 / 42. compute-sanitizer --tool memcheck --target-processes all <test binary> --test-threads=1 → 0 API errors over 565 / 0 / 42. No count row moves: no #[test] was added or removed, and every count above is the same number the #240/#241 record left. cargo check --release, cargo check --release --features cuda and cargo check --release --tests --features cuda all exit 0; the non-test builds are warning-free and the only warning: line is build.rs's pre-existing cargo:warning= target list. cargo fmt --all --check clean; check_status.py --check, check_docs_links.py and check_source_layout.py clean.

Bar named before measuring. No #[test] is added, removed or weakened: the bar is "no test loses an assertion; every deletion is a no-caller item", and the evidence form is the census before/after (the deleted code has no gate to mutate, so a mutation-checked gate would be a claim about nothing).

Limits. (1) The macOS-only modules and tests are not compiled here; ModelDef::forward_graph is the one bucket-A item that lives only behind target_os = "macos", and it was kept for that reason — build-macos compiles the library, not the macOS-gated tests, so nothing in CI would have caught a deletion of it. (2) --features debug_dump was not built; every bucket-A name was grepped against its gated code (src/dump.rs, src/quants.rs, src/main.rs) and none appears there. (3) The counted census is a cargo check: a "caller" is a compile-time reference, not a verified runtime exercise; the suites above are the independent runtime check. (4) The real-model gate sets were not re-run (no test in that set was touched); their rows remain dated 2026-09-27.

Test-infrastructure record (#243, 2026-10-01) — the config-gated dead-code annotations name their configuration

The ticket. #243 is T4 of the allow(dead_code) census: bucket D — an item rustc reports dead in the build without --features cuda and live with it. Its annotation must name that configuration (#[cfg_attr(not(feature = "cuda"), allow(dead_code))]), not blanket-silence the item everywhere. The same rule covers an item dead only in the non-test build (CudaState::cc, whose one reader is a #[cfg(test)] accessor), and the sentence this ticket adds to Core Convention 5 states it once.

Reconciled numbers — the tracker was patched to this run. The ticket body carried the e9bf4c5 item list, so the census was re-derived on 212e748 (after #242) and the body rewritten:

treedead src/ sitesdead itemsbucket D
e9bf4c5 — the durable census' recorded baseline31426022
79c9837 — after #238/#239191156—
5147d2c — after #240/#241143109—
212e748 — this branch's parent (after #242)714722
this branch's final tree714722

Method, identical on both sides: strip.py in a scratch worktree (every allow(dead_code) replaced by a //STRIPPED comment, line numbers preserved), then cargo check --release --message-format=json and cargo check --release --features cuda --message-format=json, taking every src/ span of every dead_code diagnostic (rustc groups a dead impl into one diagnostic, so a primary-span-only read undercounts ~2.6×). Bucket D is the set difference: items in the CPU capture that are not in the --features cuda capture. The --tests --features cuda capture is the dead-in-the-maximal-build discriminator; src/cuda.rs and src/graph/cuda_backend.rs do not exist without the feature, so their dead items are outside bucket D by construction.

Per-item disposition (all 22).

dispositionitems
tightened to #[cfg_attr(not(feature = "cuda"), allow(dead_code))]GraphAllocator::kv_format (src/graph/alloc.rs) — the only bucket-D item still bare
already precise, module-level form, no editDeviceKey, Provenance, TIERS, GENERIC, llama_key, Selected, select, select_forced, family_row, select_by_key — #[cfg_attr(not(feature = "cuda"), allow(dead_code))] mod device_tier; (src/main.rs); OffloadPlan::allows_weight — #[cfg_attr(not(any(target_os = "macos", feature = "cuda")), allow(dead_code))] pub mod offload; (src/graph/mod.rs); CudaWeightReg + cuda_weight_reg — #![cfg_attr(not(feature = "cuda"), allow(dead_code))] (src/models/weight_reg.rs); q4k_dsc_payload_bytes + q4k_dsc_payload_ok + q4k_dsc_plane_admitted — the same inner form (src/q4k_dsc.rs)
already precise, item-level form, no editDeviceMemory::{Reported.free, QueryFailed.code} — #[cfg_attr(not(any(feature = "cuda", test)), allow(dead_code))] (src/graph/allocplan.rs)
explicitly kept, naming #244Vendor, QClass, DeviceTier (src/device_tier.rs) — their bare item-level #[allow(dead_code)] is load-bearing for #244's never-constructed members in the cuda build (Vendor::{Amd,Mthreads,Apple}, QClass::Other, DeviceTier::{source,mmvq_batch_default,mmvq_batch_by_type}). Tightening the containing item to not(feature = "cuda") would expose exactly those members and break the warning-free cuda build; giving each its own note is #244's decision, so this ticket does not pre-empt it and a comment records the coupling on that issue.

CudaState::cc is not bucket D (it lives in the feature-gated src/cuda.rs): its only reader is the #[cfg(test)] CudaState::cc() accessor, so its honest form is #[cfg_attr(not(test), allow(dead_code))] — the configuration where it is unused is the test build's complement, not a feature. The #[cfg(test)] accessor and its graph/cuda_backend/tests.rs caller are untouched. The three remaining residual src/cuda.rs dead sites (cudaStreamWaitEvent + CudaState::stream_wait_event, the F5 device→device wait) are intentionally untouched, pending open #138.

The reverse direction does not exist here. Seven items are reported in the cuda capture but not the CPU one, all in src/device_tier.rs: Vendor::{Amd, Mthreads, Apple}, QClass::Other, DeviceTier::{source, mmvq_batch_default, mmvq_batch_by_type}. They are not "dead only with cuda": they are never constructed or read in either configuration. The CPU build reports the enclosing enum/struct as unused and never descends to its members, so the members never appear in its dead set. A #[cfg_attr(feature = "cuda", allow(dead_code))] on them would silence the cuda report of an item that is dead in both — the over-claim this ticket exists to remove. They are #244's shape items, and the census bucket definition has to be read as the difference between the two captures, not as "whichever capture names an item".

Evidence form: the stripped-oracle identity (this ticket's version of mutation evidence, as in #240/#242). There is no gate to mutate — the diff is two attributes and two comments, and an attribute-only change has no observable behaviour to break. The falsifiable claim is instead that tightening changes the diagnostics by exactly nothing, and that each cfg_attr names the configuration rustc proved dead:

capturebeforeaftersymmetric difference
cargo check --release (CPU)59 items / 80 sites / 38 diagnostics59 / 80 / 380
cargo check --release --features cuda47 / 71 / 2447 / 71 / 240
cargo check --release --tests --features cuda23 / 34 / 1623 / 34 / 160

Not one item was newly exposed and not one was hidden: the after-set is the before-set. (The :296 → :300 line move of kv_format is the four-line doc-comment addition; the item key is (file, name, kind).)

Non-truncation of the required captures, re-proved on this tree. #![cfg_attr(not(test), deny(warnings))] turns the lint into an error, so the build aborts after the lint pass; the census uses a twin build per configuration in which that one crate-level line is commented out (the cheap equivalent of the durable census' RUSTFLAGS=--cap-lints=warn, which would force an nvcc re-run for every target). Required vs twin item sets: CPU 59 = 59, cuda 47 = 47, symmetric difference 0 in both phases — so neither capture is truncated.

Per-item config split (the proof each cfg_attr is precise, not decorative).

itemcfg_attr coversdead there?dead in the other configuration?
GraphAllocator::kv_formatnot(feature = "cuda")CPU: yes--features cuda: no
CudaState::cc (field)not(test)non-test cuda: yes--tests --features cuda: no

Counts (rule 5), box dgxspark (aarch64, GB10 sm_121), 2026-10-01. cargo test --release → 480 / 0 / 36 unit + 10 / 0 / 6 integration. bash scripts/cuda_test.sh → 565 / 0 / 42. cargo check --release, cargo check --release --features cuda and cargo check --release --tests --features cuda all exit 0; the non-test builds are warning-free and the only warning: line is build.rs's pre-existing cargo:warning= target list. No count row moves: no #[test] was added or removed, so docs/status.toml and the AGENTS.md rows are unchanged. cargo fmt --all --check clean; check_status.py --check, check_docs_links.py and check_source_layout.py clean.

Bar named before measuring. No #[test] is added, removed or weakened, and the change is attributes plus comments only: the bar is "the stripped-oracle dead set is unchanged in each of the three configurations, and each tightened annotation is dead in exactly the configuration it names" — stated before the after-capture was taken and measured by the tables above.

Limits. (1) The macOS-only modules are not compiled here, but no bucket-D item is behind target_os = "macos" in the sense that would matter: --features cuda is the maximal Linux configuration and the CPU capture is the other side of the difference. (2) --features debug_dump was not built; neither the tightened item nor any bucket-D name appears in src/dump.rs. (3) The counted census is a cargo check: a "caller" is a compile-time reference, so kv_format's liveness with cuda rests on cuda_backend::entry's kv_format hook being compiled in that configuration (the suite above is the independent runtime check). (4) CudaState::cc's test build is --tests --features cuda, not the macOS test build; the #[cfg(test)] accessor is platform-independent.

Test-infrastructure record (#244, 2026-10-01) — the dead-code decisions get a verdict, and the unambiguous subset lands

The ticket. #244 is T5, the last ticket of the allow(dead_code) census: the items rustc still reports on the stripped oracle after T1–T4, which need a decision rather than a deletion. It is also the ticket the earlier records handed their findings to — #239's two lost callers (KvCache::set_owner, OffloadPlan::all_on_device), its vacuous live_peak gate and its Tensor::new coverage gap; #240/#241's stale-doc items; #242's ~12 stale-doc notes; and #243's three src/device_tier.rs container blankets.

Reconciled lineage — the oracle re-run on 0a81d47. Method identical on both sides: strip.py in a scratch worktree (every allow(dead_code) replaced by a //STRIPPED comment, line numbers preserved), then cargo check --release, cargo check --release --features cuda and cargo check --release --tests --features cuda, each with --message-format=json, taking every src/ span of every dead_code diagnostic (a primary-span-only read undercounts ~2.6× because rustc groups a dead impl into one diagnostic). The scratch worktree was pinned to the repo's 1.97.1 toolchain (RUSTUP_TOOLCHAIN is set to stable in this shell; the census unsets it).

treecuda dead sitescuda dead itemsCPU sites/itemstests-cuda sites/items
e9bf4c5 — the durable census' baseline314261——
79c9837 — after #238/#239191157——
5147d2c — after #240/#241143110——
212e748 — after #2427148——
0a81d47 — this branch's parent (after #243; re-run here)714880 / 6034 / 23
this branch's final tree573967 / 5122 / 15

A one-item parser bug the re-run exposed, fixed in the durable census.py. RopeStyle::Interleaved = 1 carries an explicit discriminant, and the item-name regex (VARIANT_RE) accepted only (/{/,/end after the identifier — so the variant's own row was dropped while its span still counted. The durable captures were re-parsed after the fix: e9bf4c5 is 261 cuda / 198 CPU items, not 260 / 197 — exactly one more in each non-test capture. The sites columns are span-based and unchanged, and the historical rows above are the recorded ones plus that one (the fix is deterministic and the variant is present in every tree). The cuda set is therefore 41 shape + 7 code items, not the 40 + 7 #242 reconciled, and the item this recovery adds — RopeStyle::Interleaved — is one of #244's own (its bare #[allow(dead_code)] is now #[cfg_attr(not(test), allow(dead_code))], the test that constructs it being graph::op_matrix).

The 7 code items are cudaStreamWaitEvent + CudaState::stream_wait_event (reserved for the F5 device→device staging copy — still unreachable after #138), KvCache::set_owner, OffloadPlan::all_on_device, ModelDef::{as_any, forward_graph, offload}.

Verdict table (all 48).

verdictitemslanded?
delete (internal, never read in any build, no API contract)SampledToken::logit; Slot::id; KVCacheLayer::{k,v,size,max_size,dim} + KVCache::layersyes
wire the lost caller (small, behaviour-preserving)OffloadPlan::all_on_device (called by resolve's unset-request arm)yes
keep + honest note, precise annotationDType::{F16,Q8_0}; Op::{Scale,Softmax,Reshape,Permute,BatchMatMul}; AttnMode::Mha; RopeStyle::Interleaved; TimingMode::{Off,Private}; OffloadRequest::AutoWithBudget; HfCheckpoint::order (deferred to #209); Tokenizer::{id_to_score,id_to_type}; Vendor::{Amd,Mthreads,Apple}; QClass::Other; DeviceTier::{source,mmvq_batch_default,mmvq_batch_by_type}; TurnOutcome::{text,stopped_by_eog,stopped_by_string}; Tokenizer::{special_tokens,im_start,im_end}; BackendCaps::{supports_op,supports_fused,supports_attn_span}; BackendEntry::name; ModelDef::{as_any,forward_graph,offload}yes (annotation + note; the membership questions are escalated)
keep, nothing to changeCudaState::cc (already not(test) from #243); cudaStreamWaitEvent + stream_wait_event (bare, dead in every build — #138 landed 2026-10-04 and did not create a reachable caller)n/a
escalate, do not landBackendCaps authority; ModelDef::as_any; ModelDef::offload; KvCache::set_owner; the deferred DType/Op/RopeStyle variant membership; removing ModelDef::forward's &mut KVCacheno

Two annotation-shape corrections ride along: Op::View's #[allow(dead_code)] was unnecessary (D1's GraphBuilder::split_parts constructs it, and the stripped oracle does not report it), so it was dropped; and the ModelDef trait-level blanket became four member-level annotations (not(test) for as_any/offload, any(not(test), not(target_os = "macos")) for forward_graph) so forward/kv_format/… stay checked.

The three device_tier.rs container blankets (#243's acceptance item). Vendor, QClass and DeviceTier lost their item-level #[allow(dead_code)]; each member rustc named in the cuda build now carries its own: Vendor::{Amd,Mthreads,Apple} bare (never constructed in either configuration — the CPU build reports the enclosing enum and never descends), QClass::Other and DeviceTier::{source,mmvq_batch_default,mmvq_batch_by_type} not(test) (constructed/read by device_tier::tests). The non-test cuda build stays warning-free, which is what #243 could not achieve without deciding these members first.

The vacuous gate (#239's finding), fixed rather than removed. AllocPlan::plan's in-loop live never decreased, so peak always equalled reserved_bytes and peak.max(live_peak(..)) could never observe live_peak; the_live_peak_is_not_the_reserved_total stayed green under a mutation of live_peak's arithmetic. The dead in-loop accumulation is gone; live_peak now takes its peak between the add and the remove pass (an interval whose first == last counts for its one step), and the test adds the shape where the two numbers genuinely differ — two different classes whose lifetimes never overlap (both reserved, one live at a time) — on top of the existing same-class-singletons and two-classes-overlapping cases. No assertion was removed or weakened; two were added to an existing test, so no count row moves. Mutation: live += class_bytes(classes[i]) → live += 0 fails the_live_peak_is_not_the_reserved_total at allocplan/tests.rs:171 (before this ticket the same mutation left it green). The [#244] module doc also replaces the stale sentence that claimed the plan is the feasibility gate: the gate is GraphAllocator::alloc_in_pool.

The Tensor::new coverage gap (#239's finding), closed. Its only consumer, f32_tensor in graph/builder/tests.rs, asserted neither the strides nor the allocation, so the body was unobserved. The helper now asserts strides[0..4], nbytes() and data.len() == nbytes() for the f32 shapes it is called with. Mutation: tensor.strides[1] = 0 fails builder_creates_topo_sorted_graph at the new assertion.

Census before/after (the evidence form for annotation-only work).

capturebefore (items / sites)after (items / sites)newly dead
cargo check --release (CPU)60 / 8051 / 670
cargo check --release --features cuda48 / 7139 / 570
cargo check --release --tests --features cuda23 / 3415 / 220

The cuda-set difference is exactly the nine items the ticket deleted or wired — the six KVCache fields, SampledToken::logit, Slot::id and OffloadPlan::all_on_device — and every other item is the same item at a different line. No item was newly exposed (a deletion can unmask a transitive callee, so this is checked, not assumed) and none was hidden. Item identity is (file, name, kind) (with the parser fix above); the line-keyed diff is misleading because the doc/annotation edits move lines.

Counts (rule 5), box dgxspark (aarch64, GB10 sm_121), 2026-10-01. cargo test --release → 480 / 0 / 36 unit + 10 / 0 / 6 integration. bash scripts/cuda_test.sh → 565 / 0 / 42. cargo check --release and cargo check --release --features cuda exit 0 with 0 diagnostics. The test build (which does not carry deny(warnings)) has an unchanged warning multiset: 79 diagnostics before and after, with no new, removed or re-counted message; the only difference is the source line of one pre-existing unused variable: pos in graph/builder/tests.rs, shifted by the assertions this ticket added above it. No count row moves: no #[test] was added, removed or weakened, so docs/status.toml and the AGENTS.md rows are unchanged. cargo fmt --all --check, check_status.py --check, check_docs_links.py and check_source_layout.py are clean.

Bar named before measuring. No item is load-bearing (the landed code changes are deletions and one constructor call whose result is identical), so the bar is the census identity plus the two test mutations: "the after-set is the before-set minus exactly the deleted/wired items", "the live_peak mutation is red", "the Tensor::new mutation is red" — stated before the after-capture and the two mutation runs, and measured by the tables above.

Escalated — reported on #244, not landed. (1) BackendCaps authority: the three supports_* fields are written by every entry() but read only by tests; the trait methods call the same module-level functions, so the answers cannot diverge, but the registry field is not the production read path. Options: make the trait/assignment read the caps (the design's intent, a real refactor), drop the three fields (small, removes the mirror), or keep them as the test-asserted mirror (what the code does). (2) ModelDef::as_any — no src/ downcast outside #[cfg(test)]; keep as the test harness' handle or delete it and restructure that harness. (3) ModelDef::offload — the builders read model.offload.plan directly; wire them through the trait method or delete it. (4) KvCache::set_owner — own_range clamps where set_owner errors, so wiring is not behaviour-preserving; wire behind the clamp, delete it with its loud check, or keep it test-only. (5) The deferred DType/Op/RopeStyle variant membership (keep the vocabulary or delete the variants). (6) Removing ModelDef::forward's legacy &mut KVCache parameter (the type is now an empty marker kept so the signature does not move). Each has a recommendation on the ticket; none changes behaviour, so none belongs in a dead-code cleanup silently.

#252 forward note (2026-10-02): escalation (6) landed. src/cache.rs, its mod cache; declaration, the &mut KVCache argument of ModelDef::{forward,forward_graph}, the KVCache::new call in main.rs and the tests that constructed one only for the signature are all gone; the parameter was never read by any path (the graph allocator owns KV). See the #252 record in §test-infrastructure.

Doc corrections landed. docs/BACKEND-REGISTRY-DESIGN.md §3 + §10 carry the BackendCaps correction above; src/models/{qwen2,qwen3}/mod.rs no longer claim ModelDef::offload() hands the plan to the builders (they read the field); src/graph/allocplan.rs's module doc no longer claims the plan is the feasibility gate. clear_launch_failure/init_kv_cache and the #242 stale-doc items were already corrected or deleted by #240/#242, and were re-checked here: no live doc still names them as callers. AGENTS.md Core Convention 5 gains one sentence: a deliberately retained deferred item says what would construct or read it.

Limits. (1) The macOS-only modules (metal.rs, the Metal half of graph/metal_backend.rs) are not compiled on Linux, so their annotations are outside this census — the macOS CI job is the only gate, and the ModelDef::forward_graph cfg is written to be true on the macOS test build, where its one caller lives. (2) --features debug_dump was not built; no #244 item appears in src/dump.rs. (3) The census is a cargo check: a "caller" is a compile-time reference, so all_on_device's new liveness rests on resolve being compiled in every build (the offload unit tests are the independent runtime check). (4) The device_tier.rs member annotations are asserted warning-free in the two non-test builds and the cuda test build, not on a macOS host (the module is CUDA-only). (5) The one-item parser fix above is applied to the durable census.py but is not in the repo: it does not change any site count, and it corrects item counts only.

Test-infrastructure record (#244 escalations, 2026-10-01) — the five decided verdicts land, and decision 6 becomes #252

The ticket. #244's six escalations each needed a decision rather than a dead-code deletion. Decision 5 (keep the deferred DType/Op/RopeStyle vocabulary, each member naming what would construct it) had already landed in #251; this branch lands 1–4 and 2, and decision 6 — retiring the legacy KVCache and ModelDef::forward's &mut KVCache parameter — is a trait-signature refactor, not dead-code cleanup, so it was filed separately as #252 with the verified inventory rather than smuggled into this ticket.

Verdict table.

#itemverdictwhy
1BackendCaps::{supports_op, supports_fused, supports_attn_span}delete the three fieldsevery entry() wrote them and production read none: the trait methods forward to the backend module's free function / constant, and assignment reads the trait. The registry keeps reads_packed_kv, the one field production reads (a format question answered without a &self, §8 of the registry design)
2ModelDef::as_anykeep, doc fixedits callers are #[cfg(test)] only (the harness deliberately holds Arc<dyn ModelDef>); no src/ production path downcasts, and there is no weight-registration consumer — the doc no longer claims one
3ModelDef::offloaddelete the trait method + both implsthe graph builders and the assignment pass read the concrete model.offload.plan field; the method's only readers were tests
4KvCache::set_ownerdelete it and its two test assertionsno production consumer, and it was not behaviour-equivalent to own_range (which truncates out-of-range silently where set_owner returned Err), so wiring it in would have changed behaviour — option (a) was rejected on that ground
5deferred DType/Op/AttnMode/RopeStyle vocabularyalready landed in #251verified only: every kept member names what would construct it
6ModelDef::forward's legacy &mut KVCacheescalated as #252a trait-signature refactor with its own acceptance criteria; the parameter is provably unread

#252 forward note (2026-10-02): the escalation landed as its own PR — decision 6 is no longer pending. The KVCache type, the mod cache; declaration and the &mut KVCache parameter are deleted, with zero behaviour change (every implementation ignored the argument, and both graph entry points already bound it as _kv). The #252 record in §test-infrastructure carries the evidence.

The caps≡trait gate was repointed, not deleted. registry_caps_match_the_backend_trait used to compare Backend::CPU.caps().supports_op (a registry field) against the trait method. With the fields gone, the same test now compares cpu_backend::supports_op / supports_fused / SUPPORTS_ATTN_SPAN against the trait method — i.e. the authority the field merely mirrored, so the property the gate was written for (the registry's answer and the trait's answer cannot disagree) is still asserted, and the two sides are now the forwarding pair rather than a field and the code it pointed at. names_resolve_and_unknown_names_are_refused had three caps().supports_* assertions in its "an unregistered handle claims nothing" block; they were repointed to the remaining field (reads_packed_kv), which is the only capability an unregistered handle can now be asked about. Mutation: making CpuBackend::supports_op diverge from cpu_backend::supports_op (return true unconditionally) fails the test — panicked at src/graph/registry/tests.rs:319: assertion left == right failed: Input F16, left: false, right: true; reverted.

The removed-assertion note (decision 4). Cache::set_owner's only callers were the two assertions in graph::kvcache::tests::own_prefix_and_release_round_trip: c.set_owner(0, 1, FREE).unwrap() + the owner[1] == FREE check (the release half) and assert!(c.set_owner(0, 9, SEQ_MAIN).is_err()) (the loud out-of-range check). Two assertions are removed with the method, which is the same discipline #240/#241 used for the device_entry guard's test: the deleted method's check has no production consumer, and the deliberate production contract is own_range's silent clamp, so keeping a test that pins the opposite (erroring) contract would pin behaviour the engine does not promise. No #[test] was added or removed; the test is renamed own_prefix_round_trip so its name does not overclaim, and it still covers own_prefix's owner table + n_used. The two ModelDef::offload deletions removed no assertion — the five test call sites (qwen2::graph::tests ×3, tooling::tests ×4 uses) now read model.offload.plan, and the one call inside a &dyn ModelDef closure takes the plan as an explicit argument because the plan surface is the concrete field.

Census before/after (same method as #244: strip.py in a scratch worktree at the pinned 1.97.1, three --message-format=json captures, every src/ span of every dead_code diagnostic). "sites" is the number of such spans; "items" the deduplicated (file, name, kind) set, the same two columns the previous records use.

capturebefore (items / sites)after (items / sites)newly dead
cargo check --release (CPU)51 / 6746 / 600
cargo check --release --features cuda39 / 5734 / 500
cargo check --release --tests --features cuda15 / 2215 / 220
cargo check --release --tests (CPU)7 / 97 / 90

The cuda set is 33 shape + 6 code → 30 shape + 4 code, and the difference is exactly the five deleted items: the three BackendCaps shape fields (supports_op, supports_fused, supports_attn_span) and the two code items (KvCache::set_owner, ModelDef::offload). The site column drops by 7 rather than 5 because two of the removed spans are container spans rustc had grouped into the now-gone diagnostics (pub struct BackendCaps { and impl KvCache {), which move or vanish with the grouping, not with an item. No item was newly dead (checked, not assumed: a deletion can unmask a transitive callee), and the cuda/CPU --cap-lints=warn captures have symmetric difference 0 in both the before and the after run, so neither required capture is a truncated lint pass.

Counts (rule 5), box dgxspark (aarch64, GB10 sm_121), 2026-10-01. cargo test --release → 480 / 0 / 36 unit + 10 / 0 / 6 integration; bash scripts/cuda_test.sh → 565 / 0 / 42 — all three numbers unchanged, because no #[test] was added or removed (the cap test was repointed, the two set_owner assertions were dropped from an existing test). cargo check --release and cargo check --release --features cuda exit 0 with 0 diagnostics (the non-test build's deny(warnings) gate). The test build's warning multiset is unchanged: cargo check --release --tests 68 warning diagnostics / 35 distinct messages before and after, cargo check --release --tests --features cuda 79 / 38 before and after, with no new, removed or re-counted message. No count row in docs/status.toml or AGENTS.md moves, so none was edited.

Bar named before measuring. The removed items are not load-bearing (three write-only fields and two methods with test-only callers; the landed code changes are deletions plus test reads of the same concrete field), so the bar is the census identity plus the one forwarding mutation: "the after-set is the before-set minus exactly the five deleted items, with zero newly dead", "the caps-divergence mutation is red", and "the suite counts are unchanged" — stated before the after-capture and the mutation run, and measured by the tables above. The mutation rule is applied where it bites (the repointed gate); for the pure deletions there is nothing to mutate, and what was verified instead is the census identity and the green suites.

Limits. (1) Metal is not compiled on Linux, so the metal_backend::entry() edit and the ModelDef::as_any/forward_graph annotations are asserted by the macOS CI job, not here. (2) The census is a cargo check: a "caller" is a compile-time reference, so the claim that the three deleted fields had no production reader rests on rustc plus the git grep of .caps() (the only remaining reads are reads_packed_kv). (3) --features debug_dump was not built; no #244 item appears in src/dump.rs. (4) The --tests --features cuda row is the maximal test build; the CPU --tests row is reported alongside it for completeness.

Test-infrastructure record (#256, 2026-10-02) — the 25 bare allow(dead_code) verdicts land, and the stripped oracle does not move

The ticket. #256 is the last Linux-side work on the allow(dead_code) population: the 18 sites whose bare allow is redundant because the item is live (delete the annotation), the 7 that are dead only outside cfg(test) (tighten to #[cfg_attr(not(test), allow(dead_code))]), and the 4 stale comments the deletions exposed. It lands as PR #257. It is the Layer-1 prerequisite #254 names for its shape ratchet ("resolve the 40 bare/other allow(dead_code) sites … so Layer 1's rule starts from a clean baseline"). The two macOS-only sites (src/metal.rs:971, :2029) stay with #255, which needs a Mac; the 9 bare-correct deferred sites (HfCheckpoint::order — #209 owns apply-or-drop — the two cuda.rs items #138 reserved and left dead, Vendor::{Amd,Mthreads,Apple}, tokenizer.rs::{id_to_score,id_to_type}, AttnMode::Mha) stay bare, because bare is the precise spelling for an item dead in every compilable configuration.

Method, identical on both sides. strip_dead_code.py in a scratch worktree (.worktrees/256-oracle): every allow(dead_code) — bare, cfg_attr and inner #![…] — replaced by a line-preserving //STRIPPED comment. Then four cargo check --release --message-format=json captures with RUSTFLAGS=--cap-lints=warn on the pinned 1.97.1 toolchain (RUSTUP_TOOLCHAIN is stable in this shell, so it is unset for every run): cpu, cuda (--features cuda), tests_cpu (--tests) and tests_cuda (--tests --features cuda). Sites are every src/ span of every dead_code diagnostic (a primary-span-only read undercounts ~2.6×); items are every member named in a diagnostic's message, keyed (file, name, kind) — "variants Scale, Softmax, … are never constructed" contributes five items, and a line move does not read as a change.

The baseline is the recorded one, byte-for-byte. The four f47e4c3 captures are diff-identical to the captures the 36-site table was built from (/home/yusiwen/minfer-36/captures/*.dead.json on the maintainer's box), so the before column is that measurement rather than a restatement of it.

capturebefore sites/itemsafter sites/itemsitem Δper-file site-count Δ
cargo check --release (CPU)60 / 4660 / 460none
cargo check --release --features cuda50 / 3450 / 340none
cargo check --release --tests22 / 1522 / 150none
cargo check --release --tests --features cuda22 / 1522 / 150none

Every diagnostic is unchanged, not only every dead_code one. The complete JSON diagnostic multiset (level, code, message) is equal before and after in all four configurations: 32 cpu (all dead_code), 17 cuda, 73 tests_cpu (61 warnings + 12 dead_code) and 86 tests_cuda (74 + 12). The stripped-annotation population falls 61 → 43 (36 bare + 25 cfg_attr → 11 bare + 32 cfg_attr).

A — the 18 deletions. Each allow was redundant: the item is live in every configuration that compiles it.

site on f47e4c3itemthe live reader
src/cuda.rs:551,553,555,557,559minfer_launch_fail_{pending,site,name,code,clear}via CudaState::take_launch_failure → graph/cuda_backend.rs:2137 execute_node (production, cuda)
src/cuda.rs:3329CudaState::take_launch_failurethe same execute_node read; the #162 gates (src/cuda/issue162_tests.rs:954…) are the other callers
src/cuda.rs:1346test_call_failure_requestedproduction graph_end_capture_to_exec (src/cuda.rs:3403, the destroy:graph_destroy injection)
src/cuda.rs:2965CudaState::get_or_grow~24 non-test impl CudaState call sites (MMQ / attention / f16 scratch), e.g. matmul_f32_ptr_layout ← cuda_backend.rs:1205/1346/1440
src/graph/alloc.rs:321GraphAllocator::enable_cudagraph/json.rs:94, models/qwen2/graph.rs:614, models/qwen3/graph.rs:534, and the cuda_backend::entry enable hook (:2034)
src/graph/alloc.rs:351GraphAllocator::cudathe registry hooks pool/host_read (cuda_backend.rs:2011/2017)
src/graph/backend.rs:139Backend::synchronizeGraphAllocator::sync_backend (alloc.rs:2356) ← scheduler.rs:221/417
src/graph/copystats.rs:63the impl CrossCopyStats blanketthe impl is empty in every non-test build (delta has carried #[cfg(test)] since #238)
src/graph/kvcache.rs:62KvLayer::ownerC2's resolver, and C1's own release_seq (kvcache.rs:363) ← alloc.rs:1623
src/graph/kvcache.rs:66KvLayer::n_usedC2's after_rm / kv_rm (kv_rm ← conversation.rs:189/258)
src/graph/kvcache.rs:281KvCache::iteralloc.rs's kv_rm (:1121) and kv_save_with_host
src/gguf.rs:1681MmapFile::_fileattribute line only — the field stays; rustc ignores _-prefixed fields, so the allow silenced nothing
src/models/qwen2/mod.rs:48Qwen2Model::n_layerthe concrete &Qwen2Model caller at models/qwen2/graph.rs:714
src/download/mod.rs:295HfSibling::sizedownload/mod.rs:274 .and_then(|s| s.size) (the l.size at :365 is OllamaLayer::size — a name collision, not this field)

Eight of the 18 are live only under --features cuda (enable_cuda, cuda, the five launch-fail FFI decls and take_launch_failure), which is why the CUDA capture and the CUDA warning-free build are load-bearing here: deleting them without the CUDA check would leave the CPU CI green while the CUDA deny(warnings) build broke — the trap this series exists to close.

B — the 7 tightenings, with the split that justifies the cfg. All seven items behave identically: dead in both non-test configurations, alive in both test configurations, so not(test) is exact and any other cfg (not(feature = "cuda")) or a bare allow would be wrong.

itemcpu (non-test)cuda (non-test)tests-cputests-cudaconstructed in test code by
DType::F16 (graph/mod.rs)deaddeadalivealivegraph/tests.rs (:102, :120)
DType::Q8_0 (graph/mod.rs)deaddeadalivealivegraph/tests.rs (:104, :122)
Op::Scale (graph/ops.rs)deaddeadalivealiveop_matrix.rs (:262, :555, :677), optiming/tests.rs
Op::Softmax (graph/ops.rs)deaddeadalivealiveop_matrix.rs (:573, :679), builder.rs's #[cfg(test)] fn softmax
Op::Reshape (graph/ops.rs)deaddeadalivealiveop_matrix.rs (:416, :625, :701, :886)
Op::Permute (graph/ops.rs)deaddeadalivealiveop_matrix.rs (:432, :633, :704, :895)
Op::BatchMatMul (graph/ops.rs)deaddeadalivealiveop_matrix.rs (:706, :937)

Each keeps the note naming what would construct it in production (a kernel taking an f16 activation, an IR node exposing the quantized buffer, a builder scaling in place, …); the not(test) cfg is the part that says when it is unused. The mutation direction that matters was checked: a wrong cfg is exactly what this split would expose, and the stripped after-captures keep all seven out of both test builds and in both non-test ones.

The copystats coupling, stated because the deletion is only conditionally safe. The blanket was on impl CrossCopyStats, whose one member (delta) already carries #[cfg(test)] (#238), so the impl is empty in the shippable build and an empty inherent impl draws no diagnostic. Deleting the blanket is therefore safe only while delta keeps its #[cfg(test)]: if someone un-gates delta, the non-test build must fail — which is the correct outcome, and the reason the blanket is deleted rather than narrowed (the member-level gate is already the precise spelling).

The gguf attribute-only note. src/gguf.rs MmapFile::_file is the one site where "delete the annotation" means the attribute line only: _file is the RAII handle that keeps the fd alive for the mapping's lifetime and must stay. rustc ignores _-prefixed fields for dead_code, so the attribute silenced nothing — the stripped capture reports no gguf.rs diagnostic at all, which is the baseline's confirmation rather than a reasoning claim.

Warning table — the un-stripped twin. The same four commands on both trees, each analysed from its own --message-format=json capture, with RUSTFLAGS=--cap-lints=warn so the crate's cfg_attr(not(test), deny(warnings)) cannot truncate the pass. Counting rustc compiler-message diagnostics (so a build script's cargo:warning= is not mixed in with them):

commandf47e4c3 diagnostics (dead_code / other)8c3f74bmultiset
cargo check --release0 (0 / 0)0 (0 / 0)equal (empty)
cargo check --release --features cuda0 (0 / 0)0 (0 / 0)equal (empty); the run's one warning: line is build.rs's pre-existing cargo:warning= target list, the same line #243 recorded
cargo check --release --tests68 (7 / 61)68 (7 / 61)equal
cargo check --release --tests --features cuda79 (5 / 74)79 (5 / 74)equal

Both production configurations are warning-free — no rustc diagnostic of any kind. The two --tests configurations are not warning-free on f47e4c3 either: 61/74 of their diagnostics are the test-code unused import / unused variable set and 7/5 more are dead_code on test helpers, the "separate, tracked cleanup" Core Convention 5 explicitly scopes out of the non-test gate. These counts differ from the stripped capture's 73/86 above only because stripping the annotations exposes 5/7 more dead items; the shared 61/74 non-dead_code warnings are the same set in both. (An un-stripped --tests --features cuda cargo run therefore prints 81 warning: lines: 79 diagnostics plus cargo's generated 79 warnings summary; --tests prints 69 = 68 + 1.) The acceptance property this ticket can honestly claim is therefore "no new warning and no new dead_code in any of the four configurations", which the multiset equality above proves directly; "the test build is warning-free" is not a property master has, and this record does not claim it.

Counts (rule 5), box dgxspark (aarch64, GB10 sm_121), 2026-10-02. cargo test --release → 480 / 0 / 36 unit + 10 / 0 / 6 integration. bash scripts/cuda_test.sh → 565 / 0 / 42. No #[test] was added, removed or weakened, so no count row moves and docs/status.toml / the AGENTS.md rows — including projection_base_passed — are untouched. cargo fmt --all --check clean; check_status.py --check, check_docs_links.py and check_source_layout.py clean.

Bar named before measuring. No #[test] is added, removed or weakened and the diff is 25 attribute lines plus 4 comments, so the bar is: the stripped-oracle (file, name, kind) item set and the per-file site-count histogram are unchanged in all four configurations, and the un-stripped warning multiset is unchanged in all four — stated before the after-captures and the twin builds were run, and measured by the two tables above.

Mutation evidence takes the annotation-shape form here. There is no behavioural gate to mutate: an attribute-only diff has no observable effect, and the falsifiable claims are (a) deleting a redundant allow changes no diagnostic once all annotations are stripped, and (b) tightening keeps the item dead in exactly the configuration the cfg_attr names. Claim (a) is measured by the before/after oracle identity (0 item Δ, 0 per-file site Δ, equal diagnostic multisets); claim (b) is mutation-checkable in the direction that matters and was checked — the after-captures' per-item split is dead/dead/alive/alive for all 7, which a wrong cfg (e.g. not(feature = "cuda"), or a bare allow) would have broken.

Limits. (1) macOS is not compiled here: src/metal.rs's two sites are #255's, the macOS half of graph/metal_backend.rs and the macOS branches of the loaders are invisible to all four captures, and this ticket touches none of them. (2) --features debug_dump was not built; no #256 site lives in src/dump.rs. (3) The oracle is a cargo check: "live" is a compile-time reference, not runtime reachability — Qwen2Model::n_layer is live precisely because a MINFER_GRAPH_DUMP branch references it, which is the right rule for a warning-free build and not a statement about hot paths. (4) cuda_static was not built (it changes linking only). (5) The cross-platform gate is CI's build-macos plus the three build jobs; the local aarch64 runs above are this box's.

Test-infrastructure record (#254, 2026-10-02) — the dead-code ratchet: an annotation-shape guard in check-docs, and a stripped oracle in the two building jobs

The ticket. The #218 → #227 → #228/#236/#238–#244 series established that an #[allow(dead_code)] is not a lint silence but a liveness root: rustc hides the annotated item and its whole call chain, so one stray annotation can hide a large dead subgraph while deny(warnings) stays quiet (the minimal case the census was built on: without the allow rustc reports both struct NeverBuilt is never constructed and fn annotated_root is never used; with it, neither). The series measured the real dead set by stripping every annotation and reading rustc's own dead_code diagnostics, but that census lived in a scratch script outside the repository. #254 productises it as two mechanical layers; it lands as PR #258. The workflow keeps its seven-job shape: Layer 1 rides in check-docs (no build), Layer 2 is the last step of test-linux-cpu (--config cpu) and build-linux-cuda (--config cuda) — the two jobs that already pay for their compile, so the strip pass is one more cargo check over a warm target/.

Layer 1 — the shape guard (scripts/check_dead_code_annotations.py). It enforces the annotation shape, not liveness, and its docstring says so — that is what keeps it honest. R1 a bare #[allow(dead_code)] is rejected unless it is one of the 11 sites that predate the ratchet (GRANDFATHERED_BARE, keyed path:item; a stale entry is a failure, so fixing a site forces its entry out — the two src/metal.rs entries must go when #255 resolves them). R2 a cfg_attr must name a cfg predicate: #[cfg_attr(dead_code, allow(dead_code))] names the lint, not a configuration. R3 the annotated item needs a reason marker in its own comment block — a ticket reference (#254), a named consumer/configuration (tokenizer::tests, not(feature = "cuda")), or one of the deferred-use phrases #244 established. R4 the rejection text points a test-only item at #[cfg(test)] / tests.rs, which needs no annotation. R3 is a presence test: the checker can see that a reason was written, never that it is true. --selftest pins 14 cases (13 shapes plus the stale-entry rule) and runs in check-docs.

Seven annotations had no reason under R3. Each gets a one-line doc-comment reason — no code and no attribute change — and all seven are reported here because the ticket asks for them:

siteitemthe reason added
src/conversation.rs:329TurnOutcome::textread only by conversation::tests; the streaming path hands each delta to the caller as it is produced
src/device_tier.rs:82QClass::Otherpoints at the enum note above: device_tier::tests constructs it; a caller classifying a non-K-quant weight type would in production
src/device_tier.rs:106DeviceTier::mmvq_batch_defaultread by device_tier::tests; the batch-cap activation (plan §14 R8) would read it
src/device_tier.rs:109DeviceTier::mmvq_batch_by_typeread by device_tier::tests; the batch-cap activation would read the per-class override
src/graph/ops.rs:22AttnMode::Mhawhat would construct it: a multi-head-attention builder (n_head_kv == n_head)
src/graph/ops.rs:136Op::Permutethe deferred note: a builder that wants a transposed view rather than a copied/reshaped buffer
src/graph/ops.rs:143Op::BatchMatMulthe FusedOp::BatchMatMul note: the single-output IR cannot express a batched matmul

(mmvq_batch_default's doc already mentioned “kernels”, which the marker test happens to accept; its reason was added anyway so the two fields of the pair read the same.)

Layer 2 — the oracle ratchet (scripts/check_dead_code_oracle.py + docs/dead-code-baseline.toml). It copies the tree to a scratch directory (the checked-out tree is never modified), replaces every allow(dead_code) — bare, cfg_attr, a combined list (dropping only dead_code) and the inner #![...] form — with a line-preserving marker, then runs cargo check --release [--features cuda] --message-format=json with RUSTFLAGS=--cap-lints=warn (never deny(warnings), which turns the lint into an error and can truncate the pass); a third configuration, --config macos, is that same plain cargo check --release run on a macOS host, where the #[cfg(target_os = "macos")] modules are compiled — elsewhere it is refused by name, because a run that compiles none of them would read green (#332). It takes every src/ span of every dead_code diagnostic — the primary span alone undercounts ~2.6×, the first census's error — and every (name, kind) the diagnostic's message names (fields a, b … is two field items, variants X, Y … two variant items), and compares that set with the manifest. An addition fails, printed with file:line, the diagnostic and a paste-ready entry; a removal is informational; a file move is a note. The strip cannot pass silently: an unrecognised spelling (a multi-line attribute, a clippy::dead_code path), a surviving dead_code in a code line, a zero-annotation strip, or a capture with no dead_code diagnostic at all is an exit-2 infrastructure error.

The scratch copy shares the calling tree's target/ by default, and that is the point: on a warm tree the cpu configuration costs 1.5 s (only the crate is rechecked) while the cuda one — whose non-test artifacts the job has not built — costs about two minutes over the job's already-warm dependency cache. Four CI steps, seconds-to-minutes inside jobs that compile the tree anyway.

Counts, both architectures. sites is every src/ span of every dead_code diagnostic; items is every (name, kind):

configurationdgxspark (aarch64, GB10 sm_121)--target x86_64-unknown-linux-gnu
cargo check --release (cpu)32 diagnostics / 60 sites / 46 items32 / 60 / 46
cargo check --release --features cuda17 / 50 / 3417 / 50 / 34

Both columns reproduce the #256 record's recorded cpu 60 / 46 and cuda 50 / 34 exactly, and the two architectures are identical at the (name, kind) level. The x86_64 column is a local cross-check the box can make because cargo check needs no x86 linker — the rustc invocation carries --target x86_64-unknown-linux-gnu and target_arch="x86_64" — but it is not a run on the runner. Since the sets agree, the manifest is the union of one set and carries no arch field; the authoritative architecture validation is the PR's own CI run, which passes no --target (the first run's oracle output is the confirmation this record's table is checked against).

Confirmed on the runner. The PR's first CI run is green 7/7 with zero code annotations (the single build-macos annotation is GitHub's platform notice about macOS runner capacity). On the x86_64 runner test-linux-cpu printed stripped 43 annotation(s) … dead_code diagnostics: 32 src/ spans (sites): 60 items: 46 … baseline [cpu]: 46 entry/entries — 0 addition(s), 0 removal(s); PASS in 17 s, and build-linux-cuda printed stripped 43 … 17 … 50 … 34 … 34 entry/entries — 0 addition(s), 0 removal(s); PASS in 2 m 47 s over its warm dependency cache. Those are the same sets as the aarch64 column, so no arch field was needed and the manifest needed no second seed; check-docs printed annotation shapes clean (11 grandfathered bare site(s)), the Layer-1 selftest's 14 cases pass and the oracle's strip, item and capture cases pass.

The manifest (docs/dead-code-baseline.toml). 45 [[cpu]] + 34 [[cuda]] + 42 [[macos]] entries, each name / kind / file / reason. The macOS section is a measured set since 2026-10-09, run on a Mac (#332's Mac half); the macOS set is the CPU set minus QueryFailed/Reported/allows_weight, which macOS makes live, so each [[macos]] entry reuses its [[cpu]] twin's reason (see the manifest header). The "unjudged" spelling still exists for a config with no recorded measurement yet (#332), and only a config whose spec sets unjudgeable may carry it: cpu/cuda are machine-checked in CI, so the marker there would silently disable a gate and is refused; macos is an on-demand Mac run and no longer carries it. The update rule is written into the file: an addition is a decision — make the item live, delete it, or add the entry in the same PR with a one-line reason — and the checker prints the entry to paste; a removal is informational. --print-toml regenerates the block (reusing the existing reasons), and --selftest pins the strip, the item derivation and the capture parser in check-docs, so regression in the strip is caught without cargo.

The blind spots, written into the manifest rather than discovered later. (1) macOS was judged on 2026-10-09 (#332's Mac half), so the old blind spot is closed. The run is the plain cargo check --release with every annotation stripped, on a Mac, where the #[cfg(target_os = "macos")] modules are compiled: it reports 42 items and no dead item under src/metal/ or src/graph/metal_backend* — the evidence #255's two dispositions rest on (the MpsState container blanket is deleted and matmul_on_gpu_buf is #[cfg(test)]). The macOS set is the CPU set minus QueryFailed/Reported/allows_weight, which macOS makes live, so the [[macos]] section mirrors the [[cpu]] reasons. The cross-platform half of the blind spot stays visible in the data: ModelDef::forward_graph is dead in the non-test check and its reason names the #[cfg(target_os = "macos")] test (models::qwen2::graph::tests::real_model::graph_metal_matches_cpu_logits) that keeps it live in the macOS test build. build-macos compiles the test target since #303 (cargo test --release --no-run); the oracle itself is still an on-demand Mac run, not a CI job. (2) The oracle is a cargo check: “live” is a compile-time reference, not runtime reachability. (3) --features debug_dump and cuda_static are not covered.

Fixture evidence — both layers can fail. Layer 1: appending a bare #[allow(dead_code)] fn layer1_bare_allow_probe() {} to a scratch copy makes check_dead_code_annotations.py exit 1 with two problems (the bare form; no reason); swapping it for a reason-less #[cfg_attr(not(test), allow(dead_code))] exits 1 with the “no reason” problem alone; removing the probe restores the clean run. Layer 2: a synthetic oracle_mutation_probe function added to a scratch copy makes the oracle exit 1 in both configurations, naming oracle_mutation_probe (fn) at src/vec_ops.rs:1349 and printing the manifest entry; pasting that entry into the scratch manifest's [[cpu]] array makes the same command pass (47 items, 47 entries — 0 addition(s), 0 removal(s)), which is the update path in one diff. Neither fixture is committed.

Bar named before measuring. Layer 1: the shape audit passes on the tree as it stands, and each fixture mutation exits non-zero. Layer 2: the stripped oracle reproduces the recorded cpu 60 / 46 and cuda 50 / 34 and the manifest compares equal (0 additions, 0 removals) in both configurations on both architectures. Both were stated before the after-runs and are met by the tables above.

Mutation evidence takes the fixture form here. There is no behavioural gate to break: an attribute-only diff plus two new checkers have no runtime effect, so the falsifiable claims are that Layer 1 rejects the two shapes it forbids and that Layer 2 fails on a newly hidden item. Both are the transcripts above, and the oracle's own strip is mutation-checked by construction — the oracle_mutation_probe run strips 44 annotations (43 + the probe) and the item set moves 46 → 47 exactly.

Counts (rule 5), box dgxspark (aarch64, GB10 sm_121), 2026-10-02. cargo test --release → 480 / 0 / 36 unit + 10 / 0 / 6 integration; bash scripts/cuda_test.sh → 565 / 0 / 42 (unit), integration 7 + 3 passed / 0 failed / 6 ignored. No #[test] was added, removed or weakened — the diff is two scripts, a TOML manifest, CI steps and seven doc comments — so no count row moves and docs/status.toml / the AGENTS.md rows are untouched. cargo fmt --all --check, check_status.py --check, check_docs_links.py and check_source_layout.py are clean.

Limits. (1) The x86_64 column is a cross-compile, not a run on the runner; the CI job is the confirmation. (2) The one-line reason strings are as good as the annotation notes they reuse: the oracle proves an item is dead, never that the reason for keeping it is still true. (3) A dead_code diagnostic whose only span is outside src/ is ignored by construction — this configuration compiles the crate, not tests/. (4) A new annotation that hides a large subgraph is caught the next time the oracle runs, not by Layer 1; that division of labour is the design.

Test-infrastructure record (#332, 2026-10-09) — the macOS dead-code set, judged on a Mac

The ticket. #332 is the macOS half of the dead-code ratchet: #254's stripped oracle carried only cpu (the test-linux-cpu job) and cuda (build-linux-cuda), so no oracle configuration compiled a macOS-only module and the manifest recorded a macos = "unjudged" marker instead of a verdict. The Linux half landed first (4188a72, PR [#424]): the checker gained a macos configuration (the plain cargo check --release, refused by name on a non-Mac host, covered by --selftest) and the schema accepted an unjudgeable config's marker. This record is the Mac-side half: the marker replaced by the measured set.

The run. Box macbook (macOS 27.0.1, Apple M4 Pro) (hostname macbookpro-ysw), 2026-10-09, python3 scripts/check_dead_code_oracle.py --config macos --print-toml: 40 annotations stripped, 29 dead_code diagnostics, 53 src/ spans (the sites), 42 items. Seeded into docs/dead-code-baseline.toml as [[macos]], the re-run reports 42 entry/entries — 0 addition(s), 0 removal(s); PASS.

What the Mac run found. The macOS set is exactly the CPU set minus three items, each live on macOS for a host reason rather than a feature one: Reported/QueryFailed (graph/allocplan.rs) are constructed by MpsState::device_memory — whose #[cfg_attr(not(any(feature = "cuda", test)), allow(dead_code))] is therefore redundant on macOS, as its doc comment already says — and allows_weight (graph/offload.rs) is called by both loaders under #[cfg(any(target_os = "macos", feature = "cuda"))], so the offload module's #[cfg_attr(not(any(target_os = "macos", feature = "cuda")), allow(dead_code))] applies no blanket there. Every one of the 42 is D4 (keep, with a reason): the set is the test-only/deferred vocabulary the CPU section had already judged, and each reason is platform-neutral, so the [[macos]] entries mirror them. The run reports no dead item under src/metal/ or src/graph/metal_backend*, which is the positive evidence #255's two dispositions rest on: the container-level impl MpsState blanket was correctly deleted (loaders, main.rs, graph/metal_backend.rs and get_or_grow/cmd_buffer read all ten members) and matmul_on_gpu_buf is #[cfg(test)].

ModelDef::forward_graph (the known cross-platform case). It appears in the macOS non-test set — its #[cfg_attr(any(not(test), not(target_os = "macos")), allow(dead_code))] (src/models/mod.rs) reduces to an allow in a non-test macOS build — and its reason names its one caller, the #[cfg(target_os = "macos")] test models::qwen2::graph::tests::real_model::graph_metal_matches_cpu_logits. In the macOS test build the same cfg applies no allow, so the method is live there; that is why the verdict is D4 (keep the annotation), not D1/D3. The three forward_graph reasons (the [[cpu]], [[cuda]] and [[macos]] rows) now spell the full real_model:: test path.

Properties. No source annotation changed (no D1/D2/D3 verdict), so no test count moves; the diff is the manifest, its header and this record. --config cpu and --config cuda are untouched, so the Linux CI jobs are unaffected. macos remains an on-demand run, not a CI job (build-macos compiles the test target only).

Test-infrastructure record (#252, 2026-10-02) — the legacy KVCache and ModelDef::forward's &mut KVCache parameter are deleted

The ticket. #252 is #244's decision 6, spun out because it is a trait-signature refactor rather than a dead-code cleanup: a dead-code ticket may delete an item, but it may not move a public trait's signature. #244 had already deleted KVCache's storage, leaving src/cache.rs a 31-line storage-free marker whose only job was to keep ModelDef::forward's kv: &mut KVCache argument spelled the same. The argument was provably unread: the graph path owns KV in the allocator's persistent regions, both graph entry points already bound the parameter as _kv (models/qwen2/graph.rs:435, models/qwen3/graph.rs:380), and every implementation only threaded it through. #252 lands as PR #259.

What was deleted, and the proof the parameter was dead. The inventory was re-grepped on the branch point (bef4bd9), not trusted from the ticket: git grep -n 'KVCache' src/ plus git grep -n '\.forward(' src/ names exactly the sites below, and nothing else.

whatsite beforechange
the marker typesrc/cache.rs (31 lines: pub struct KVCache;, KVCache::new)file deleted
its declarationsrc/main.rs:13 mod cache;deleted (with the file, so check_source_layout.py stays consistent)
the trait declarationsrc/models/mod.rs:207 (kv at :211)kv: &mut KVCache removed; doc comment updated
the default forward_graphsrc/models/mod.rs:264 (kv at :268)kv removed from the signature and the delegation
Qwen2 implssrc/models/qwen2/mod.rs:56/:97 (kv at :60/:101)both removed; use crate::cache::KVCache; gone
Qwen3 implssrc/models/qwen3/mod.rs:48/:89 (kv at :52/:93)both removed; use crate::cache::KVCache; gone
graph entry pointssrc/models/qwen2/graph.rs:435, src/models/qwen3/graph.rs:380_kv: &mut KVCache removed from forward; imports gone
CLI constructionsrc/main.rs:1227 (KVCache::new)deleted, with the then-unused n_kv_embd/n_layer locals
CLI forwardssrc/main.rs:1478, :1788argument dropped
test constructionsrc/models/qwen2/graph/tests.rs:2978/3019/3023KVCache::new calls and arguments dropped
test mocksrc/server/batch/tests.rs:1794 (FailingForward::forward)_kv parameter dropped

src/conversation.rs was verified rather than assumed: GraphEngine::forward (src/conversation.rs:43) is its own 3-argument method and the model calls go through forward_graph_cached (:237), neither of which ever took a KVCache — its diff is empty. src/graph/cache.rs (pub mod cache; at src/graph/mod.rs:29, the GraphCache) is untouched; only the other, legacy cache module went. Cargo.toml, src/graph/mod.rs and every GraphCache call site are unchanged.

Zero behaviour change — stated before the suites ran. No implementation read the argument (both graph entry points bound it _kv), KVCache had no fields since #244, and nothing else names the type. The edit is therefore signature-only: the graph every caller builds, the KV the allocator owns, the token stream and the logits are bit-identical. The suite counts are the check — no #[test] was added, removed or weakened, so the rows below are unchanged rather than re-baselined.

Warning table — four configurations, before and after. Same method as #256: four cargo check --release --message-format=json captures on the pinned 1.97.1 toolchain with RUSTFLAGS=--cap-lints=warn (the crate's cfg_attr(not(test), deny(warnings)) must not truncate the pass), counting rustc compiler-message diagnostics (a build script's cargo:warning= is not mixed in) and comparing the whole (level, code, message) multiset by sha256, not only counts:

commandbefore bef4bd9after (this PR)multiset
cargo check --release0 (0 dead_code / 0 other)0 (0 / 0)equal (empty) — sha e3b0c442… both
cargo check --release --features cuda0 (0 / 0)0 (0 / 0)equal (empty) — sha e3b0c442… both
cargo check --release --tests68 (7 / 61)68 (7 / 61)equal — sha cc1f8468… both
cargo check --release --tests --features cuda79 (5 / 74)79 (5 / 74)equal — sha 6b87d17c… both

Both non-test configurations are warning-free in the strict sense (the plain cargo build --release and cargo build --release --features cuda, without --cap-lints, both exit 0 under deny(warnings); the cuda run's one warning: line is build.rs's pre-existing cargo:warning= target list). The two --tests configurations carry #256's recorded 68/79 pre-existing test-code warnings, which Core Convention 5 explicitly scopes out. The honest claim is therefore "no new warning in any of the four configurations", and the sha equality is stronger than the count equality: not one message was added, removed or moved.

The ratchet's first real exercise — it said PASS, 0 additions, 0 removals. #254's stripped oracle (scripts/check_dead_code_oracle.py) ran in both configurations against docs/dead-code-baseline.toml:

[cpu]  stripped 43 annotation(s); dead_code diagnostics: 32  src/ spans (sites): 60  items: 46
       baseline [cpu]: 46 entry/entries — 0 addition(s), 0 removal(s); PASS
[cuda] stripped 43 annotation(s); dead_code diagnostics: 17  src/ spans (sites): 50  items: 34
       baseline [cuda]: 34 entry/entries — 0 addition(s), 0 removal(s); PASS

Those are exactly the recorded cpu 60/46 and cuda 50/34 that #254/#256 measured, and the manifest needed no edit — no entry added, none removed, no arch field. That the ratchet is silent here is itself the finding, and it is the expected one: the oracle strips allow(dead_code) annotations and compiles, and KVCache carried no annotation — it was live by reference, kept alive by the very signature this PR moves. The ratchet's subject is newly hidden code; this PR deletes a live-by-reference item, so its item set is unchanged. The manifest never listed KVCache (grep -n KVCache docs/dead-code-baseline.toml is empty), which is why there is no removal to report. The two runs also confirm the oracle survives a file deletion and a mod removal without an infrastructure error (the "every .rs is declared" half of check_source_layout.py is what would have caught a half-done deletion).

Counts (rule 5), box dgxspark (aarch64, GB10 sm_121), 2026-10-02. cargo test --release → 480 / 0 / 36 unit + 10 / 0 / 6 integration; bash scripts/cuda_test.sh → 565 / 0 / 42 unit + 7 + 3 passed / 0 failed / 6 ignored integration; PARALLEL=0 scripts/real_model_gates.sh (CPU) → 36 / 0. No #[test] was added, removed or weakened, so no count row moves and docs/status.toml, the AGENTS.md rows and projection_base_passed are untouched. cargo fmt --all --check, check_status.py --check, check_docs_links.py, check_source_layout.py and check_dead_code_annotations.py are clean.

Docs. AGENTS.md's layout block loses its cache.rs line; docs/ARCHITECTURE.md loses the module-table row and its "remains only as CLI plumbing" sentence; docs/COMPUTE-GRAPH-DESIGN.md §10's trait sketch loses kv and the other two src/cache.rs mentions are corrected; docs/METAL-BACKEND-DESIGN.md's legacy-surface paragraph is past-tensed; docs/CPU_OPTIMIZATIONS.md's P2/P4 "current state" lines stop citing the deleted file/parameter; docs/OPENAI-CHAT-API-PLAN.md's revision note and slot structure say the ignored argument was deleted; the walkthrough docs 01/03/09/13 carry dated forward notes (history is not rewritten); and the move of this record closes the two #252 forward notes in the #244 and #244-escalations records above.

Bar named before measuring. No #[test] is added, removed or weakened and the parameter was never read, so the bar is: the four warning multisets are unchanged, the stripped-oracle (site, item) set is unchanged in both configurations, and the three suite counts are unchanged — stated before the after-captures, the oracle runs and the suites, and measured by the tables above.

Mutation evidence takes the warning-multiset form here. A signature refactor has no runtime behaviour to mutate and no new gate to break — the falsifiable claim is "the edit is inert", and the sharpest inertness test available is that the compiler sees exactly the same message multiset in all four configurations, which cannot hold if the deletion moved a live path (it would surface as a new dead_code/unused diagnostic) or left a dangling reference (it would surface as a resolved error). The oracle's stripped captures are the second arm: if KVCache had been merely hidden rather than removed, or if the deleted parameter had a reader left behind, the (site, item) set would have moved; it did not, in either configuration. A behavioural mutation is not available because there is no behaviour in the diff, and inventing one would be the "passed for the wrong reason" failure docs/GATE-CONTRACT.md warns about.

Limits. (1) macOS is not compiled here: graph_metal_matches_cpu_logits is the one test that called forward_graph, and its call site is inside a #[cfg(target_os = "macos")] test block, so this box proves the edit compiles on Linux and CI's build-macos job (a type-check, not a test run) must prove the macOS half compiles. (2) The oracle is a cargo check: "live" is a compile-time reference, as #254 states. (3) --features debug_dump and cuda_static were not built; neither names the removed type. (4) The --tests warnings are the two configurations' pre-existing sets (68/79), not a claim that master's test build is warning-free.

8. Note — the dead identity fields (A7 rationale)

CParams.n_batch and GraphParams.n_seqs live in the two structs that define the graph-reuse identity (src/graph/params.rs), and GraphCache::params_match compares both (src/graph/cache.rs:57-64). A change in either forces a full graph rebuild.

They are dead because nothing reads them:

  • n_seqs is set to 1 at every construction site (models/qwen2/graph.rs:443, models/qwen3/graph.rs:382, graph/json.rs:31, …). No builder, kernel or allocator consults it.
  • n_batch is set to n_tokens at every site (models/qwen2/graph.rs:452, …) — it is a second copy of a field that already exists, so it can never carry independent information.

Both are therefore compared for equality but can never differ: they are inert.

Why it is worth a ticket. They read as evidence that the engine supports multi-sequence batches and prefill chunking, which it explicitly does not (docs/COMPUTE-GRAPH-DESIGN.md §1.3 lists multi-sequence batching as a non-goal). The concrete hazard is the next implementer: n_seqs sitting in the reuse identity looks load-bearing, so E2 might start by "just setting it", while every real decision point — attention masking, KV addressing, allocator liveness — has no notion of sequence ids. Deleting or annotating the fields forces that realisation at the point where it is cheap.

Options.

OptionWhat changesTrade-off
(a) Recommended — delete n_batch, keep n_seqs with an explicit "reserved for item 3" commentOne field goneE3 will reintroduce n_batch with its real meaning (a chunk size chosen by the scheduler, not n_tokens); keeping today's field would make the old meaning the default. n_seqs keeps its exact future name and semantics, so nothing is lost by keeping it.
(b) Keep both, comment as reservedNothingMinimal churn, but two inert fields remain in the identity and the comment is easy to ignore.
(c) Delete both, re-add in E2/E3Two fields goneCleanest now, but n_seqs must be re-added as part of E2 — and E2 is already the largest change in the plan.

Decision (2026-09-16): option (a). Delete n_batch in A7; keep n_seqs with a comment naming item 3 as its future reader.

Follow-up (E2, the rounding-out of this note). Option (a) assumed n_seqs would be made real by item 3. Item 3 landed and the assumption was wrong in an instructive way: the sequence count is data, exactly like n_past, and never needed to be in the identity at all. What item 3 did add to the identity is CParams.explicit_span — the one decision (which attention instantiation to build) that depends on how a batch is composed, derived from the KV reservations rather than from a count. So the field was deleted in E2 instead of being populated: option (c), chosen late, with the measurement that makes the case (sequence_count_is_data_not_topology). The general lesson this note now records: an identity field earns its place only if some topology decision reads it — "a future feature will need it" is not enough, because a redundant field costs a rebuild every time it changes shape with the topology unchanged.

9. Phase F — independent tracks

Can run in parallel with A–E by a different workstream.

IDItemTitleEffortBox needed
F111AVX2/AVX-512 dots for the K-quants + weight repacking · #56 — dots landed 2026-10-08 on an x86 host (bitwise gates; AVX-512/VNNI selected at runtime), weight repacking is the remaining incrementLx86
F215GBNF-style grammar + JSON-schema constrained decoding · #47 — DONE 2026-09-24 · follow-ups #125 (refused constructs), #126 (mask cost)Mdgxspark
F316Sampler set: min-p, typical, XTC, DRY, mirostat, logit bias · #48 — DONE 2026-09-24Mdgxspark
F412Backend registry (drop the compile-time enum) · #57 — DONE 2026-09-24 · the per-device KV-format capability #87 needs is now a used registry field (BackendCaps::reads_packed_kv), not a hardcoded CPU testMdgxspark
F514Async cross-backend copy + events · #58 — DONE 2026-09-24 · CUDA's boundary copy is an cudaMemcpyAsync D2H into a pinned slab plus an event, waited on once at a documented synchronization point; the CPU is a registered synchronous no-op and Metal declines (unported) · follow-ups #137 (Metal), #138 (deferred wait — DONE 2026-10-04, see the F5 S2 record)Mdgxspark (CUDA)
F622Quantizer tooling (convert-hf-to-gguf, quantize, split) · #49 — DONE 2026-09-24 · a GGUF v3 writer (gguf_write.rs), byte-exact weight encoders (quantize.rs), an HF converter that passes the strict loader (convert.rs), the three subcommands (tooling.rs), and a real download size check — follow-ups #140 (K-quant encoders), #141 (f16 on the device), #142 (bf16)Ldgxspark
F719/20Chat-template fidelity + tokenizer generality · #50 — DONE 2026-09-24 · follow-ups #132 (NFC + the remaining pre-tokenizer rules) and #133 (--chat-template, strftime_now)Mdgxspark
F825Metrics/observability (/metrics, KV occupancy, queue depth, per-op timing under a flag, graceful drain). Item 25 was the only member of the A-era batch (items 23/24/26/27/28 -> A1/A2/A7/A5/A6) with no ticket; it is independent of the critical path, hence this table · #51 — DONE 2026-09-24, both real-model gates device-verified on GB10 sm_121 2026-09-24; the serial #[ignore]d set it left red (#123) is green as of 2026-09-24 (22 passed / 0 failed)Mdgxspark

F1's dot half landed on an x86 host on 2026-10-08 (landing record below); the remaining increment — weight repacking into SIMD-friendly interleaved layouts — still needs that machine class and remains the largest single CPU win.

The hardware lane, stated where the sequencing is read (#335). F1's remaining increment (weight repacking, #56; the AVX2/AVX-512 dots themselves landed 2026-10-08) wants an x86 host — an entirely different machine class from the two boxes this project owns, dgxspark (aarch64, GB10 sm_121) and macbook (macOS 27.0.1, Apple M4 Pro); the x86_64 (CI runner) is GitHub's ephemeral runner and cannot host interactive work. It is therefore not the project's next: sentence, and docs/status.toml's next: no longer says so. What can start on the two boxes instead, per the §14 hardware row: on macbook the macOS round-2 tickets tracked in #333 (the Metal-only list and the order, with #310 since landed — its last capability gap closed; on dgxspark the open Linux-side tickets — #200 (the CUDA Op::FusedQkvNorm kernel), #212, #354 and #356. No schedule is attached to either lane; this is a statement about which machine each one needs.

F1 landing record — AVX2/AVX-512 K-quant dots (#56), 2026-10-08. On an x86_64 host (Intel i7-11700B, avx2/fma/avx512f/bw/dq/vl/vnni), src/quants/avx2.rs gained the Q4_K/Q5_K/Q6_K × Q8_K AVX2+FMA kernels and the new src/quants/avx512.rs VNNI variants; kquant.rs dispatches AVX-512/VNNI → AVX2+FMA → scalar at runtime (MINFER_NO_AVX512=1 drops to AVX2, MINFER_NO_AVX2=1 to scalar — the x86 counterpart of MINFER_NO_NEON). quants::avx2_correctness gates every kernel bitwise against its *_scalar reference (the int8×int8 products are exact in i32 and the per-superblock float arithmetic keeps the scalar order), and its #[ignore]d kquant_simd_dot_speedup harness records the dot-level ratios on a 4096-element row (median of five): AVX2 3.39× / 2.80× / 1.79× and AVX-512 3.97× / 3.23× / 2.51× scalar for Q4_K / Q5_K / Q6_K. End-to-end on the cached Qwen2.5-0.5B Q4_K_M (minfer bench, 16 threads, median of nine -r 3 samples at --n-ctx 656): prefill pp512 28.59 → 31.82 (AVX2) → 32.13 (AVX-512) tok/s, decode tg128 11.20 → 12.59 → 12.54 tok/s. The end-to-end gain is small because the CPU matmul re-reads each weight row per token (bandwidth-bound), which is exactly what the weight-repacking increment would fix; the scalar and NEON results are byte-unchanged and were not re-measured.

F3 — Sampler set (#48) — DONE 2026-09-24

What landed. src/sampler.rs gained a SamplerConfig (the pre-F3 knobs plus min_p, typical_p, xtc_probability/xtc_threshold, the DRY group, mirostat/mirostat_tau/mirostat_eta/mirostat_m, and logit_bias), a MirostatState { mu } that the caller owns (mirostat is the one sampler with cross-token state), and one pipeline entry point:

logit bias → penalties → DRY → (greedy shortcut) → top-k → typical → top-p →
min-p → XTC → temperature | mirostat v1/v2

The order is llama.cpp's common_sampler_init chain. Every filter is a pure function with its own boundary tests: apply_min_p (p = 0/1, ties, empty/one token, "the argmax always survives"), apply_typical (p = 1 off, p = 0 keeps one, dominated distributions, ties), apply_xtc (probability 0/threshold > 0.5 off — and no RNG draw, fewer than two candidates, the exact exclusion set), apply_dry (the reverse Z-algorithm plus the restart-sequence cap, an empty history, a window shorter than allowed_length, the exponential scale, the single-token-breaker exemption), sample_mirostat_v2 / sample_mirostat_v1 (the mu update direction, mu <= 0 never emptying the set, the documented degenerate s_hat = 1 rule), and apply_logit_bias (positive and negative).

Surface. --min-p --typical --xtc-probability --xtc-threshold --dry-multiplier --dry-base --dry-allowed-length --dry-penalty-last-n --dry-sequence-breakers --mirostat --mirostat-tau --mirostat-eta --mirostat-m --logit-bias on the CLI; the matching optional fields on the OpenAI request (min_p, typical_p, xtc_probability, xtc_threshold, dry_*, dry_sequence_breakers, mirostat, mirostat_tau, mirostat_eta, mirostat_m, logit_bias). SamplerConfig::validate refuses every nonsensical value at the boundary (CLI startup / HTTP 400), and validate_logit_bias rejects a token id outside the vocabulary once it is known — nothing is clamped or silently ignored.

Acceptance measurements. cargo test --release: the F3 gate default_config_is_bit_identical_to_the_old_path runs 64 steps through both sample_with_penalties and the new pipeline with one seed and asserts the token sequences are equal; default_pipeline_matches_the_pinned_pre_f3_sequence asserts the same 64 tokens as a sequence captured from master before the change ([5, 54, 21, 54, 105, …], seed 42). A mutation check (forcing min_p = 0.5 into the default config) fails the pinned gate, and reverting makes it pass — so the gate can fail. Full-suite counts are in the F3 record commit / PR (unit + integration lines).

Honest scope. (a) DRY sequence breakers are token-id sequences (--dry-sequence-breakers 198;13,2), not llama.cpp's strings: mapping a string breaker to the overlapping token sequences needs get_overlapping_token_sequences against the vocabulary, a tokenizer port left as a follow-up. (b) Speculative decoding (--spec-draft) carries the pure F3 filters but not mirostat — a verify round samples several rows from one shared RNG, so mu has no faithful home there; the CLI and the server refuse the combination loudly. (c) In mirostat mode the temperature is ignored (mirostat's mu truncation subsumes it, as in llama.cpp, which sets temp = 1.0); temp == 0 still wins as the greedy shortcut. (d) The conversation/KV session snapshot does not carry mu (it does not carry the RNG either): a resumed session restarts mu at 2 * tau, which affects sampling only, never the KV.

F2 — Grammar and JSON-schema constrained decoding (#47) — DONE 2026-09-24

What landed. src/grammar.rs (new) holds the whole engine, designed in GRAMMAR-DESIGN.md and implemented against that contract; src/sampler.rs gained the mask's position in the one pipeline.

GBNF subset. Rules, string literals (\n \r \t \\ \" \xNN \uNNNN), character classes with negation, ., grouping, alternation, * + ? {m} {m,} {m,n} (upper bound ≤ 1024), # comments. Refused loudly, each with the offending token: \d/\w/\p{...}, an empty or reversed class, a repetition whose upper bound is below its lower bound or above the cap, a missing root, a duplicate rule name, an undefined reference, and left recursion (detected while compiling every rule's epsilon closure, so it is a startup error rather than a generation-time hang).

JSON-Schema subset. type (string or array), enum, const, properties/required/additionalProperties (bool or schema), items, prefixItems, minItems/maxItems, string, integer bounds (inclusive and exclusive, ranges compiled by digit decomposition), number, boolean, null, anyOf/oneOf, $defs/definitions + local $ref (recursive schemas work). Refused loudly: pattern, format, minLength/maxLength, non-integer numeric bounds on number, multipleOf, propertyNames, patternProperties, dependent*, uniqueItems, contains, allOf, not, if/then/else, remote or nested $ref, items as an array (the draft-07 tuple form), and a required name that is not declared.

Engine. One flat program per rule (Cp/Class/Any/Split/Jump/Call/ Ret), a stack of {rule, pc} frames, a bounded set of nondeterministic stacks (64) with a stack-depth bound (256) that turns left recursion into an error, codepoint matching with a carried partial-UTF-8 buffer, EOG allowed only at an accepting state with no pending bytes, and an empty-piece token never allowed. The mask is a packed bitset per token id, cached per state (bounded at 64) and computed through a per-call DFA-style transition memo.

Pipeline. SamplerConfig.grammar: Option<Arc<Grammar>> is the compiled, immutable object; the mutable automaton state is per run (GrammarState), exactly like MirostatState. The one pipeline is now

logit bias → penalties → DRY → [GRAMMAR MASK] → greedy shortcut → top-k →
typical → top-p → min-p → XTC → temperature | mirostat v1/v2

The position is the argument: everything before the mask only shifts logits (finite additions cannot lift a masked -inf), and everything after it only removes candidates or reweights survivors, so a forbidden token can never win — including under --greedy, which is why the mask is before the shortcut. The mask consumes no RNG and writes nothing the other stages read, so mirostat's mu and DRY are unperturbed for the same token sequence. sample_with_config_grammar returns SampleError::{NoAllowedToken, Grammar}: an empty allowed set is a loud stop, never an arbitrary token, and a configured grammar with no run state is an error rather than a silent fallback.

Surfaces. CLI --grammar <FILE> / --grammar-str <GBNF> / --json-schema <FILE> / --json-schema-str <JSON> (mutually exclusive, compiled once after the vocabulary loads, refused at startup); --cnv resets the automaton per assistant turn; --spec-draft + a grammar is refused. Server response_format: {"type":"json_object"} and {"type":"json_schema", "json_schema":{"name":…,"schema":{…}}} plus the grammar GBNF extension field — grammar together with a non-text response_format is a 400, and so is an unsupported construct (compiled on the handler side, before a slot is taken). Batch slots and serial requests each build their own state.

Measured acceptance. Greedy, seed 42, 0.5B q4_0 (f32 KV) and Qwen3-0.6B Q8_0 (f16 KV), schema {name: string, age: integer 0..150}:

Gate0.5B q4_0Qwen3-0.6B Q8_0
schema run parses (serde_json::from_str){"age": 25, "name": "John"}{"age": 25, "name": "John Doe"}
response_format: json_object (server, 80 tok) parses{"field1": "value1", "field2": "value2"}—
empty case: max_tokens 0 → "" / "length"✅✅
empty case: root ::= "" → "" / "stop", EOG only✅✅
max-length (8 tok) → Grammar::accepts_prefix ✅, full parse ✗ (as asserted){"name": "John Doe{"name": "John",
mask cost per new state (151,936-token vocab)5.4 ms (16 states, 86 ms)5.4 ms
bitwise no-grammar pathpinned 64-token sequence, test_default_pipeline_matches_the_pinned_pre_f2_sequencesame

The two defects the gates found are recorded in GRAMMAR-DESIGN.md §9: a partial-UTF-8 token accepted where the automaton could never finish it (found by the real-model parse gate, reproduced over HTTP as "a\uFFFD" for {"grammar":"root ::= \"ab\""}), and the mask's first version costing 71.3 ms per state (13× the memoized cost). Four mutations were run against the new gates and each made one fail: an off-by-one in the sampler's mask index, a disabled partial-completion check, an ignored minItems, and a disabled EOG allowance.

Honest scope. Object properties are accepted in declaration order (a subset of the schema's language: extras may be interleaved only where no required property is pending); oneOf is compiled as anyOf; only integer-valued numeric bounds are compiled (number + bounds is refused); pattern/minLength/maxLength and multipleOf are refused. The mask is O(vocabulary) per new state — 5.4 ms on a 151k vocabulary, cached per state — so a long constrained generation still pays it once per unseen state. All acceptance is CPU-only: CUDA/Metal are not touched by this ticket (the mask is host-side, before any device work), so the CUDA build-linux-cuda job compiles the change but no device run exercises it. Two follow-up issues carry what was deliberately left out: #125 (the GBNF/JSON-Schema constructs the compiler refuses: pattern/minLength/ maxLength, real-valued numeric bounds and multipleOf, property permutations, an exact oneOf, allOf/not, \d/\w/\p{...}) and #126 (index the vocabulary by first codepoint to cut the residual mask cost).

F7 — Chat-template fidelity and tokenizer generality (#50) — DONE 2026-09-24

What landed. src/template.rs renders the model's own tokenizer.chat_template again, and src/tokenizer.rs makes tokenizer.ggml.pre authoritative and replaces the silent byte fallback. The accepted/refused sets, the loud refusal and the reference behind every gate are the design-first artifact CHAT-TEMPLATE-AND-TOKENIZER-DESIGN.md.

Templates. minijinja 2.21.0's own extension point (Environment::set_unknown_method_callback) closes the gap — no dependency change, and minijinja-contrib::pycompat was considered and rejected as a new dependency with different edge semantics. The implemented Python str methods are strip/lstrip/rstrip (a character set, as CPython), split/rsplit with maxsplit, startswith/endswith, replace, lower/upper/title/ capitalize, join, find/rfind/count, plus a raise_exception function for templates that refuse their own input. Qwen3's template therefore renders, including its <think>-block split and re-emission — the behaviour QWEN3-SUPPORT-PLAN.md §5 gotcha #9 recorded as lost. Everything else is refused loudly: the error names the construct and the template line and states that the engine will not substitute a generic prompt. validate() runs at load, so the CLI exits before inference and serve/viz refuse to start (before the worker thread is spawned); a per-request failure is an HTTP 400. The ChatML fallback survives for exactly one case — a GGUF with no tokenizer.chat_template at all, announced once on stderr — and callers that passed "" for a missing template (which rendered the empty prompt) are fixed. format_single/Conversation propagate the refusal as a turn error.

Tokenizer. tokenizer.ggml.pre selects the pre-tokenization rule: qwen2 (alias deepseek-r1-qwen) and qwen35, hand-written splitters ported from llama.cpp's unicode_regex_split_custom_* — hand-written because the Rust regex crate has no lookahead and \s+(?!\S) is load-bearing for whitespace runs. Every other value, including a missing key, refuses the load, as do a tokenizer.ggml.model that is not gpt2, an empty merge table, and a vocabulary missing any of the 256 byte tokens. The silent unwrap_or(0) is a checked byte fallback (one token per byte), and the "whole piece is in the vocabulary" BPE shortcut is gone: it produced a different split than the reference on Qwen3.5, where a vocabulary entry is not reachable through merges.

Accepted / refused, in one line each. Templates: all minijinja constructs, the Python str methods above, raise_exception, slices with a step (messages[::-1]); refused — any other method or an unparseable construct, each naming itself. Tokenizers: byte-level BPE with pre = qwen2/qwen35; refused — SentencePiece/unigram/WordPiece, ignore_merges and multi-regex pre-tokenizers (llama3, default, deepseek-*, falcon, starcoder, …), byte_encode = false, and a missing/unknown pre.

Measured acceptance.

GateEvidence
Template rendering, per supported modelmodel_templates_render_byte_for_byte: 4 models × 7 cases byte-for-byte against transformers 5.17.0 (tests/fixtures/chat/*.json)
The artifact the engine loadsgguf_template_renders_like_the_reference (#[ignore]d): the GGUF's own template renders the same bytes for 4 models × 7 cases, including multi-turn, system-message, generation-prompt and <think>-reasoning cases
Loud refusal4 unit gates; restoring the pre-F7 fallback (mutation M2) fails all four with unwrap_err() on an Ok, and reverting makes them pass
Pre-tokenizer rulespre_tokenizer_split_matches_the_reference: qwen2 + qwen35, 52 corpus entries each, == CPython regex on the model's own tokenizer.json pattern
Token idstoken_ids_match_the_reference (#[ignore]d): 52 entries × 5 cached models byte-for-byte — Qwen2.5-0.5B/7B/14B and Qwen3-0.6B against transformers, Qwen3.5-0.8B against llama.cpp llama-tokenize on the same GGUF; transformers and llama.cpp agree on all 52 entries where both exist
Byte fallbackbyte_fallback_emits_one_token_per_byte; mutation M5 (emit id 0 again) fails it [0,0] vs [1097,1098]
Mutation checksM1 py_trim ignores the char set → rendering gate red; M2 restore the ChatML fallback → refusal gates red; M3 is_number broken → split gate red; M4 unknown pre silently defaults → refusal gate red; M5 silent id 0 → fallback gate red; M6 merge order inverted → id gate red ([1519, 654, 78, …] vs [9707, 11, 1879, 0]). All six reverted
Full suitecargo test --release: 393 passed / 0 failed / 23 ignored + 3 / 0 / 6 (baseline 382/0/21 + 3/0/6; +11 unit gates, +2 #[ignore]d real-model gates)
Real-model serial setcargo test --release --bin minfer -- --ignored --test-threads=1: 23 passed / 0 failed on the CPU build (baseline 21/0) — the F2 grammar and F3 pinned-sampler gates included

Honest scope. (a) The tokenizer does not apply the HF normalizer (NFC): llama.cpp does not either for BPE, and every corpus entry is NFC-normalized; composed-vs-decomposed input can therefore differ from transformers, and a normalization table is a follow-up (#132). (b) Beyond qwen2/qwen35 no pre-tokenizer is implemented; each unsupported value is refused by name, and the classic GPT-2 rule was dropped from the design rather than shipped ungated (GPT-2's tokenizer.json has no Split pattern to reference). (c) SentencePiece / unigram / WordPiece models are out of scope — this is a byte-level BPE. (d) The template reference is the model's published tokenizer_config.json; the GGUF copies of Qwen2.5's and Qwen3's templates differ from it textually (a converter-escaped tool-call string; for Qwen3 an older but equivalent revision), so the real-model gate asserts the rendered bytes, and the fixture records both hashes. (e) Qwen3.5 (qwen35 arch) is not a supported architecture; only its tokenizer is used, as the qwen35 reference. (f) Templates that call strftime_now or use filters minijinja lacks are refused by name rather than supported. (g) The byte-for-byte and id gates need the cached GGUFs, so they are manual/local (#[ignore]d) evidence; the CI-runnable gates are the fixture- template rendering, the split fixtures, the str-method/semantics fixtures, the loud refusal and the mutation-checked unit gates. (h) Nothing here is device-adjacent: the CUDA and macOS CI jobs are compile-only for this change.

F8 — Metrics and observability (#51) — design (2026-09-24)

Scope. Four independent pieces, landed together because they share one registry: a Prometheus text /metrics endpoint, live KV/arena occupancy, queue depth, and per-op timing behind a flag. A fifth piece — graceful drain — is the risky one and is designed separately below.

Where the numbers come from. The HTTP side and the worker side are two threads with different reach: the handler owns the job sender and the tokenizer, the worker owns the model, the GraphCache and the BatchEngine. A model reading cannot happen on the handler thread (it would need a lock on the worker), and a queue reading cannot happen on the worker thread (the channel backlog is only visible from the sender side). So the design is one Arc<ServerMetrics> of relaxed atomics written by both sides and read by the renderer:

  • the handler writes requests_total, requests_rejected_total and in_flight (the drain surface);
  • the worker writes the queue/engine counters (accepted, admitted, running) and, after every step, a published snapshot of the allocator's occupancy.

minfer_queue_depth = accepted - admitted is the quantity neither side can see alone: it is everything between the handler's send and the engine's slot assignment (the channel backlog plus the worker's pending deque). It is monotone per request and converges to 0 when the worker catches up; a transient overshoot cannot happen because admitted only ever counts a job that accepted already counted.

KV occupancy is a live reading, not a startup snapshot. BatchEngine gains publish_metrics(), called from serve_loop after each tick and after each admission group. It copies GraphAllocator::memory_report(backend) (weights / pool / live / peak live / budget / reservation depth), the arena shape (kv_layer_count() x kv_n_ctx(), kv_region_bytes(), kv_is_packed()) and the C3 counters (kv_arena_stats()) into the atomics. memory_report is a handful of HashMap lookups, so publishing per step is cheaper than the step itself; the alternative — reporting only MemoryReport and leaving the arena shape to the startup log — was rejected because the plan's E4 record says "the same struct is what F8 will export", and a snapshot taken at startup is exactly what the ticket says it must not be.

Per-op timing is measured at one choke point, and the flag is off by default. BackendScheduler::execute already walks every node and dispatches it to its backend; the timer wraps that dispatch (Instant::now() only when the flag is on, one record() after). The alternative — instrumenting each of cpu_backend / metal_backend / cuda_backend separately — was rejected: it triplicates the code, still cannot separate kernel time from its prologue, and would leave the three backends' definitions of "an op" free to drift. The honest scope this buys is written down: the number is dispatch + execution of one node, and split-level syncs and cross-backend copies are not attributed.

Storage is a fixed [(name, AtomicU64 nanos, AtomicU64 calls)] table with an exhaustive op_name(&Op) -> &'static str, so a new Op variant is a compile error rather than a metric that silently disappears. MINFER_OP_TIMING is presence-checked (the repo's convention for opt-in flags, like MINFER_TRACE), resolved once into an atomic so the hot path is a relaxed load.

Graceful drain (the risky piece). Today server::run has no signal handling and the worker is a detached thread blocked on its job channel, and a naive join() can hang forever because in-flight HTTP handlers hold an Arc<AppState> clone and therefore keep job_tx open. The design is therefore bounded first, observable second, join third:

  1. The handler increments in_flight when a job is accepted, and its response guard decrements it when the response is finished (body sent, SSE stream closed, or the client gone). in_flight is thus "requests the HTTP side still owes a response to" — which is what axum::serve's graceful shutdown actually waits on. A client that disconnects mid-answer releases its count while the worker is still decoding, so step 4 adds a bounded wait for the worker itself; that is the case a single counter would otherwise lose.
  2. On SIGINT/SIGTERM, a task sets draining (so handlers refuse new work with 503 instead of queueing behind a shutdown) and signals axum::serve's graceful shutdown, which stops accepting connections.
  3. server::run then waits for the serve future up to MINFER_DRAIN_MS (default 30000). On timeout it reports the still-in-flight count to stderr and into minfer_drain_abandoned_requests, and returns — the process exits and the detached worker dies with it. It never joins unboundedly.
  4. Only if the graceful path completed inside the deadline does it wait (again bounded) for the worker, which by then has seen every sender drop.

The alternative — worker.join() after axum::serve returns — is the trap the ticket names: an SSE client that never disconnects keeps a handler clone alive, job_tx never drops, and the join blocks for as long as the client likes. The CLI default is untouched: with no signal the loop waits forever, exactly as before.

Acceptance. /metrics rendering is a pure function with boundary tests; /metrics is driven over a real HTTP connection on an ephemeral port; the occupancy/queue numbers are published and read back in a real-model engine run; MINFER_OP_TIMING off leaves the table empty and on accumulates (and the computation's output is unchanged either way); the drain is bounded and reports what was left. At least one new gate is mutation-checked.

What landed (2026-09-24). All five pieces, in three commits: the design note, feat(graph): F8 per-op timing behind MINFER_OP_TIMING, and feat(server): F8 /metrics, live KV/queue depth, bounded drain (#51).

Surface. GET /metrics (Prometheus text 0.0.4) on the OpenAI server, served from its own router state so a scrape needs neither the tokenizer nor the job channel and keeps answering while the model is busy. Metric families and units:

FamilyTypeUnitMeaning
minfer_requests_totalcounterrequestsaccepted and queued for the worker
minfer_requests_completed_totalcounterrequestsresponse finished (body sent, stream closed, client gone)
minfer_requests_rejected_totalcounterrequestsrefused before queueing (draining / worker gone)
minfer_requests_in_flightgaugerequestsaccepted, not yet finished — the drain surface
minfer_jobs_dropped_totalcounterjobsthe worker could not place the job in a slot
minfer_worker_stalled_totalcountereventsthe worker's counted no-progress bound tripped and ended the loop, answering every live and queued request 500 (added by #196)
minfer_queue_depthgaugejobsaccepted - admitted: channel backlog + the worker's pending deque
minfer_worker_pending_jobsgaugejobsthe worker's own deque right now
minfer_requests_runninggaugerequestsoccupying an engine slot right now
minfer_prompt_tokens_total, minfer_completion_tokens_totalcountertokensprompt / delivered completion tokens
minfer_completion_tokens_per_secondgaugetokens/sgenerated tokens/s over a trailing 16 s window
minfer_draininggauge0/1a shutdown was requested
minfer_drain_abandoned_requestsgaugerequestsstill in flight when the drain deadline expired
minfer_memory_{weights,pool,live,peak_live,budget,headroom}_bytesgaugebytesE4 MemoryReport; budget/headroom omitted when unbounded
minfer_memory_{idle_slots,reserved_classes}gaugecountthe E4 S3 reservation table's depth
minfer_kv_{layers,rows,region_bytes}gaugecount / count / bytesthe arena's shape
minfer_kv_packedgauge0/1packed Q8_0 cells vs f32/f16 words (C4)
minfer_kv_{reserved,owned,shared,free}_cells, minfer_kv_{free_runs,sequences}gaugecells / runsC3/C8b arena occupancy
minfer_kv_{defrags,cells_moved,cows,cow_cells}_totalcountercount / cellsC3 compaction and C8b copy-on-write
minfer_op_seconds_total{op=…}, minfer_op_calls_total{op=…}counterseconds / countper-op, present only when MINFER_OP_TIMING is set

minfer_completion_tokens_per_second is the one family the issue names that is not a plain counter: it is a trailing 16-second window (see the honest scope), so a scrape gets a throughput without a Prometheus server, and rate() on the counters remains available for anyone who wants a different window.

Flags. MINFER_OP_TIMING (presence-checked) turns on per-op timing. MINFER_DRAIN_MS (default 30000) bounds the graceful drain. Both are documented in docs/USAGE.md.

Landed as #120 on the branch feat/f8-metrics.

Drain. SIGINT/SIGTERM → set draining → axum::serve graceful shutdown → wait for the serve future at most MINFER_DRAIN_MS via bounded_drain → on a clean drain, a second bounded wait for the worker → otherwise log and record the still-in-flight count and return. No unbounded join, and with no signal the loop waits forever as before. The --slots-file snapshot is untouched (it is still written on every completed request inside BatchEngine::finish).

Two mechanisms stop new work, and it is worth naming which does what: the graceful shutdown closes the listener, so a brand-new connection after the signal gets a connection error; the draining check in the handler then returns 503 for a request that arrives on an already-accepted connection (a keep-alive client, or one accepted just before the signal). The 503 is therefore a backstop rather than the primary switch, and it is the one the automated gate drives directly.

Measured acceptance. cargo test --release: 340 passed / 0 failed / 18 ignored (unit) and 3 passed / 0 failed / 6 ignored (integration) — a delta of +27 passed, +2 ignored from the pre-F8 baseline (313/0/16 and 3/0/6, at 911ce2c). The 27 new non-ignored tests are the rendering boundary set, the HTTP round-trips, the drain bound, the in-flight guard, the deadline parse, the timing gates, and the three token-accounting gates (the trailing window's bucket arithmetic and decay, the counters, and the rendered family). The 2 new ignored tests are the real-model gates, and the whole ignored set run serially is 18 passed / 0 failed (cargo test --release --bin minfer -- --ignored --test-threads=1): published_metrics_move_as_requests_are_served (24 layers, 128 rows, 3 145 728 B of KV region, running 0 → 1 → 0) and serve_loop_publishes_the_queue_and_running_depth (the real serve loop; peak running 2, and the deterministic one-slot round dropping exactly 1 job). CI's CUDA step, cargo test --release --features cuda --no-run, is clean locally too (nvcc 13.0, no device needed).

The first version of the serve-loop gate was itself flaky and the failure was worth recording: it sampled for a peak every millisecond while serving 2-token answers that finish inside one sampling interval, so "the peak was 0" was true of the sampler, not the worker; and simply raising max_tokens with n_ctx = 128 made a request try to grow to the whole arena, which released the other slot's reservation and killed both requests. The gate now asks for 128-token answers over n_ctx = 1024, and only stops after it has actually seen running > 0. queue_depth is deliberately not peak-asserted: serve_loop calls admit on every iteration and a job is placed or rejected within that pass, so the non-zero window after a send is shorter than a sampler's interval — the arithmetic is gated purely and the loop's drain through accepted - queue_depth == n.

Issue #51's own acceptance, item by item. (i) "Metrics are exposed in a standard scrapeable format" — Prometheus text 0.0.4 with # HELP/# TYPE per family, asserted well-formed by a gate. (ii) "and are correct while slots are admitted, grown and released" — a real-model round drives exactly that transition: minfer_kv_sequences goes 2 → 1 → 1 → 2 while reserved_cells goes 256 → 184 and owned_cells 120 → 183, with the engine logging slot 0: capacity 128 -> 184 cells for a request wanting 184 (released 1 idle slot(s); 0 run(s) moved); minfer_requests_running is 0 → 1 → 0 across it. (iii) "A drain stops accepting new requests and finishes the running ones" — the handler refuses with 503 once draining is set, and the bounded drain waits for the in-flight set (manual SIGTERM run below: 1 in flight → drain complete). (iv) "Per-op timing is off by default and does not change results when on" — gated bitwise through the scheduler, plus a greedy CLI A/B. (v) tokens/s — the counter pair and the trailing-window rate, verified live on both response paths: non-streaming → prompt 32 / completion 32 / 2.000 tokens/s, streaming → 64 / 64 / 4.000 (32 and 64 generated tokens over the 16-second window).

End-to-end run (manual, on dgxspark, CPU, Qwen2.5-0.5B Q4_0, MINFER_BATCH=1). A live minfer serve on port 18099 with MINFER_OP_TIMING=1 and MINFER_DRAIN_MS=8000: before any request /metrics reports zeros and no timing family; one 16-token chat completion then shows minfer_requests_total 1, minfer_requests_completed_total 1, minfer_requests_in_flight 0, minfer_kv_layers 24, minfer_kv_rows 512, minfer_kv_region_bytes 12582912 (= 24 × 2 × 512 × 128 × 4, the f32 KV of the 24-layer model at --n-ctx 512), minfer_memory_weights_bytes 422782464, and a per-op breakdown of exactly the ops that ran (matmul 0.538 s / 2704 calls, attn 0.034 s / 384, rms_norm, rope, swiglu, add, get_rows, kvcache_store, kvcache_load; silu and softmax are absent because the kernels fuse them). A second request was started, SIGTERM sent mid-flight with MINFER_DRAIN_MS=8000: [server] SIGTERM: draining — refusing new requests; 1 request(s) in flight, up to 8000 ms to finish → drain complete: every in-flight request finished → worker stopped; exiting. Repeating with MINFER_DRAIN_MS=150 and a 400-token request produced the forced path: drain deadline (150 ms) reached with 1 request(s) still in flight; abandoning them and exiting, with the process gone shortly after. A greedy 12-token run of the same prompt with and without MINFER_OP_TIMING=1 produced the identical generated text (The capital of France is Paris. It is the largest city) — only the throughput lines differed — which is the flag's end-to-end "off changes nothing computed" claim.

Mutation checks (the gates can fail). (a) Rendering the queue-depth sample from the worker's pending gauge instead of accepted - admitted fails both_threads_write_into_one_registry and the HTTP metrics_endpoint_renders_over_http. (b) Mapping Op::MatMul to GetRows' index fails op_names_agree_with_op_index with "two variants map to index 9". Both were reverted; the shared test gate is poison-tolerant so a mutation produces one named failure rather than a cascade of PoisonErrors.

A defect found while gating this (pre-existing, not F8's). The new minfer_jobs_dropped_total counter made it visible: serve_loop calls admit every iteration, and admit consumes the Job, so a job rejected with "no idle slot" has its event sender dropped without an error event — the handler then answers 200 with empty content, which a client cannot tell from an empty generation. Reproduced on dgxspark (CPU, 0.5B Q4_0, MINFER_BATCH=1, --n-slots 1, two concurrent max_tokens=200 requests): A 200 with 1014 chars and finish=length; B 200 with 0 chars, finish=stop, completion_tokens=0, plus [server] job rejected: no idle slot in the log. It is out of F8's scope (it changes admit's contract and the error semantics of both serve paths), so it was written up as a follow-up: #121. Fixed 2026-09-25 by #121 — admit/submit_on now answer a job they cannot place through reject, so the same reproduction reads B 503 {"error":{"code":503,"message":"no idle slot",…}} (non-streaming) or an SSE error frame (streaming); the E2 follow-up record in §7 has the before/after transcript, the decision (reject loudly, not queue), the two new gates and their mutation checks.

Device verification (2026-09-24, NVIDIA GB10 sm_121, CUDA 13.0 / nvcc V13.0.88, driver 580.178.04). Every measurement above is CPU-only, so the two real-model gates were re-run on the GPU in an isolated worktree (.worktrees/f8gpu, branch verify/f8-gpu-gates, from master 3616570), leaving the CPU numbers untouched. Build: CUDA_HOME=/usr/local/cuda-13.0 PATH=/usr/local/cuda-13.0/bin:$PATH cargo build --release --features cuda → targets sm_75,sm_80,sm_86,sm_87,sm_88,sm_89,sm_90,sm_100,sm_103,sm_110,sm_120,sm_121; PTX compute_121 (build.rs's auto-detected list; cuobjdump --list-elf confirms a sm_121 cubin and --list-ptx a compute_121 PTX section in the binary). The device really participates: the load banner reads CUDA: using NVIDIA GB10 (SM 12.1, 124544 MB, 48 SMs) / CUDA: device tier DGX Spark GB10 (Measured, mmq true) / CUDA: GPU acceleration enabled, plus E5's offload: all 24 blocks + embed/output on cuda (403.2 MiB of device weights; default), where the control (MINFER_DISABLE_CUDA=1) reads CUDA: disabled by MINFER_DISABLE_CUDA and CUDA: not available, using CPU fallback (=0 disables too — presence-checked). A greedy 8-token CLI A/B on one prompt measured 1182.0 tok/s prefill / 208.5 tok/s decode on the device vs 194.4 / 59.2 with the device disabled (6.1x / 3.5x) — the difference the gate numbers themselves cannot show.

Both gates were run serially on the device (the CudaState singleton is process-wide, so the serial form is the only honest one) on the default 0.5B (f32 KV) and on MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf (f16 KV, hd 128 — the only combination that reaches FA prefill):

GateConfigurationResultNumbers it printed
serve_loop_publishes_the_queue_and_running_depth0.5B f32 KV, deviceokpeak running 2, 24 layers, 25 165 824 B region, 262 owned cells, 1 dropped
serve_loop_publishes_the_queue_and_running_depth0.5B f32 KV, MINFER_DISABLE_CUDA=1okpeak running 2, 24 layers, 25 165 824 B, 266 owned, 1 dropped
serve_loop_publishes_the_queue_and_running_depthQwen3-0.6B f16 KV, deviceokpeak running 2, 28 layers, 234 881 024 B, 262 owned, 1 dropped
published_metrics_move_as_requests_are_served0.5B f32 KV, deviceok24 layers, 128 rows, 3 145 728 B, owned 11, idle_slots 10, reserved_classes 5; 2→1→1→2, 256→184, 120→183
published_metrics_move_as_requests_are_served0.5B f32 KV, MINFER_DISABLE_CUDA=1oksame shape, reserved_classes 4
published_metrics_move_as_requests_are_servedQwen3-0.6B f16 KV, deviceok28 layers, 128 rows, 29 360 128 B, owned 11, idle_slots 17, reserved_classes 6; 2→1→1→2, 256→184, 120→183

The asserted numbers are device-independent by construction — model.n_layer(), n_ctx and the KV element width, plus the arena bookkeeping — so the device and the control agree on them, and that agreement is deliberately not offered as evidence. The device evidence is the in-test banner/offload line and the CLI A/B throughput above. Two non-asserted gauges did differ (owned_cells 262 vs 266 in the serve-loop gate, reserved_classes 5 vs 4 in the metrics gate), both incidental — the greedy stop point and the per-backend size-class ladder — not independent proof.

The F8 path that genuinely runs on the device is per-op timing, because BackendScheduler::execute dispatches CUDA through the same wrapper. On the device with MINFER_OP_TIMING=1 a serve plus one 8-token chat completion renders 11 op families / 26 lines, including the CUDA-only fused forms fused_qkv (72 calls) and fused_ffn (72), alongside matmul 316, rms_norm 196, add 192, attn 96, kvcache_load 96, rope 48, kvcache_store 24, swiglu 24 and get_rows 6; the same server with the flag unset renders 0 minfer_op_ lines, so the flag is off by default on the device too. The computed text is unchanged: a greedy CLI A/B (--greedy --seed 42 -n 12) produced The capital of France is Paris. It is the largest city byte-identically with the flag on and off (and across a repeated off run), and the same chat request served with and without the flag returned identical content.

The full #[ignore]d set on this CUDA build was red, and not because of F8. cargo test --release --features cuda --bin minfer -- --ignored --test-threads=1 gave 5 passed / 14 failed at 3616570 (deterministic across two runs), all 14 failing with the same refusal: the E4 default CUDA budget collapsed to 0 bytes once the conversation and map-window tests had run in one process — the cudaMemGetInfo return code was discarded, free stayed 0, and the gate then refused every later allocation — while /proc/meminfo sampled during the minimal reproduction showed 118.6–119.8 GB of 121 GB still available. That is fixed (#122): the same command at eeba0d0 + the S4 record below is now 17 passed / 2 failed (20 passed / 2 failed on the F2-rebased tip, which added three tests to the set), and the 0-byte-budget message appears 0 times. The two residual failures were both the packed-cache gate's #87 cause, tracked in #123: first a_packed_kv_cache_answers_like_the_f32_one itself (CPU-only by its own docstring, refused on CUDA), then a_partial_offload_runs_the_rest_on_the_cpu, because the packed gate sets the process-wide KV format to q8_0 and panics before restoring f32, so the next test's CUDA KV region is sized for the wrong format (the #99 mechanism, made visible by #87's refusal now being the first failure instead of the budget). The map-window gate's 1.25x timing margin (a loaded GB10 exceeded it once, 1.267x) is the third #123 item and passed in this run. All three are filed, and none reproduces with either F8 gate run alone, which is why the device claim below is scoped to those two gates plus the op-timing path. All three are fixed in #123 (its test-hygiene record lives in the E4 section, above): on the 0.5B configuration the serial set is 22 passed / 0 failed, ten consecutive runs of the same binary, and the timing gate's verdict no longer depends on the box being quiet — its median-of-ratios statistic stayed below 1.15x even with 16 CPU spinners and two concurrent CUDA attention loops. The Qwen3-0.6B configuration's set was 21/1, its one failure a separate C5 defect filed as #130 (closed 2026-09-25: the container's FLAG_F16 bit, C5 S3).

Honest scope. (a) Every measurement in the sections above this block was taken on a CPU-only build: that worktree was built with cargo build --release, without --features cuda, so the server's banner reads device cpu and the CUDA backend is not compiled in at all (the outer tree's CUDA binary was left untouched). This ticket does not need a device and the wiring is backend-agnostic — the timing hook is in the shared scheduler and kv_snapshot_from reads whatever backend the model reports — so the plan-level claim was made without one. The two real-model gates and the timing path are device-verified in the block above (2026-09-24, GB10); everything else here still carries no CUDA/Metal runtime claim. The CUDA compile path is covered locally with cargo test --release --features cuda --no-run (nvcc 13.0, no device needed) and by CI's build-linux-cuda; the macOS compile path is CI's build-macos only. (b) Per-op timing measures the scheduler's per-node dispatch — the one choke point CPU/Metal/CUDA share — so it includes the backend's prologue and excludes split-level syncs, cross-backend staging copies, allocator liveness and fill_input; there is no kernel-only timer. (c) There is no histogram and no latency accounting (no _bucket/_sum, no time-to-first-token): the counters and the trailing-window tokens/s the issue asks for are present, but a latency histogram would need a bucket policy the issue did not specify. The rate is a trailing 16-second window, not a lifetime average — a lifetime average is not a throughput, since an idle server would keep reporting its startup burst — and it is a dedicated one-writer bucket store (no lock, no allocation). Tokens are recorded where the response is produced, so the batched and serial paths are covered uniformly and nothing is plumbed through the engine; a client that disconnects before its answer is complete is not counted, because those tokens were never delivered. (d) The serial (non-batched) path has one GraphCache per slot, so /metrics reports the arena of the slot that served the last request rather than a sum; the batched path reports its single shared arena in full. (e) Queue depth is a subtraction of two relaxed counters, so a scrape can transiently read 0 while a job is in the channel; it is exact in the steady state and saturating, never negative. (f) The Metal path is compile-checked by CI's build-macos only (there is no Mac here); the CUDA path is compile-checked both by CI's build-linux-cuda and locally (see (a)), and the timing hook itself is exercised on CPU here and on the GB10 (the device block above), including the CUDA-only fused ops. (g) The signal path is not in the automated suite — only bounded_drain, the deadline parse and the draining 503 are. The end-to-end runs above were performed by hand on dgxspark and are recorded here as manual evidence, not as a CI gate; wiring a SIGTERM into a test would mean a subprocess and a port, which this ticket did not take on.

F4 — Backend registry (#57) — DONE 2026-09-24

What landed. src/graph/registry.rs (new) is the backend registry, and the compile-time enum Backend { CPU, Metal, Cuda } is gone. Backend is now a Copy/Hash/Eq/Ord handle over a fixed, configuration-independent id space (cpu = 0, metal = 1, cuda = 2) — the ids are a KV-session file-format contract, so they are appended, never renumbered. The design-first artifact is BACKEND-REGISTRY-DESIGN.md; the contract is summarized in COMPUTE-GRAPH-DESIGN.md §3.6, and the module map and the Extending → New backend recipe in AGENTS.md were rewritten to match.

The registry. One BackendEntry per backend, registered by its own module at startup, carrying the name, the assignment priority (Metal 300, CUDA 200, CPU 100 — the pre-F4 statement order inside supports_for, now an explicit number), the capability matrix (supports_op / supports_fused / supports_attn_span / reads_packed_kv) and the hooks that reach the pool (pool/pool_mut, host_read, kv_format, lazy enable, unavailable). Each backend's capability matrix moved to module-level functions that its impl Backend methods forward to, so the registry's answer and the trait's answer are one authority (asserted by a gate).

What stopped matching on the enum. The allocator's twelve dispatch sites (allocate / free / fresh, pool_len, weights_bytes, write_host_window, write_host, host read, synchronize, copy_cells, the KV element format, the session enable) are trait calls on the entry's pool hook, with the per-cfg fallback arms gone; supports_for walks registry().by_priority(); scheduler::execute is alloc.pool_mut(split.backend); the fusion wiring in graph/json.rs and both model graphs is alloc.fusion_backends() + alloc.fusion_backend_index() (the hand-built [cpu, metal?, cuda?] vector and its name() == "cuda" position lookup are gone from three call sites); kvsession's tag table is Backend::index() / from_index(); the JSON and DOT exporters, the op-matrix harness and kvformat::KvFormat::supports read the handle/registry. models::Device gained Device::backend(), the one bridge between the two id spaces (kvsession::backend_of now delegates to it).

The two orders. Identity (Backend::index) fixes the on-disk KV-session tag, the exported graph's backend index, Ord and the fusion backend list; priority (the numbers above) is the assignment preference. Nothing depends on HashMap iteration order: the table is a fixed-size array indexed by id, by_priority() sorts by (Reverse(priority), id) so no two entries can tie, and the filter is a [bool; 3]. Both orders and the numbers are pinned per configuration by a gate.

The name surface. --backend <name> (repeatable; comma-separated values accepted; extracted from argv before any subcommand dispatch, so serve, viz, bench and specverify all honour it) and MINFER_BACKENDS=<csv>; accepted names cpu, metal, cuda; trimmed, case-insensitive; the flag wins over the environment; unset is the pre-F4 behaviour. The request is a fence: it removes backends from participation and never adds one. cpu is always admitted — it is the universal fallback a graph must always be assignable to — so --backend cpu is the useful spelling (force the CPU for every device) and --backend cuda means "the device when it can take the node, the CPU otherwise". The fence is read from one filter in two places, which is what keeps them from disagreeing: Qwen2Graph::device / Qwen3Graph::device (so a fenced device never causes a device-only fused node to be built) and GraphAllocator::supports_for (so a node is never placed on a pool whose weights were never registered).

Refusals. Three classes, three messages, all before the model is even resolved:

unknown backend 'gpu2'; known backends are: cpu, metal, cuda
backend 'cuda' is known but not compiled into this build: the CUDA backend is compiled only with --features cuda
backend 'cuda' is compiled in but not available on this machine: no CUDA device is available, or CUDA is disabled (MINFER_DISABLE_CUDA)

Stage 1 (names, purely — CI-covered) runs at the top of main; stage 2 (availability) runs once the device layer is up and, on every path that will run a model (run/serve/viz in main, bench and specverify in their own run), before the model path is resolved, so a missing file or a bad GGUF cannot preempt it. Only a backend the request named is checked at stage 2: the default request means "whatever this build can use", so MINFER_DISABLE_CUDA=1 / MINFER_DISABLE_MPS=1 keep meaning "run on the CPU". That distinction is a behaviour-preservation gate of its own — the first cut of stage 2 checked every allowed backend and turned MINFER_DISABLE_CUDA=1 into a startup refusal, which the gate caught.

Feature gates. The registered set is the compile-time one (CUDA only under --features cuda, Metal only on macOS) and is pinned per configuration, but the names are unconditional: metal on Linux and cuda on a default build give the accurate "not compiled into this build" refusal rather than "unknown". A backend that is compiled out is never silently treated as absent.

#87. BackendCaps::reads_packed_kv is the per-device KV-format capability query, and it is used by this ticket's own code, not reserved: GraphAllocator::ensure_kv's packed-region refusal and KvFormat::supports both read it, replacing backend != Backend::CPU and matches!(device, Device::Cpu). The region-sizing gate and the C4 format gate now read one field; #87 is the work that flips CUDA's and Metal's value and adds their kernels. No follow-up issue is filed for it here.

Measured acceptance.

CommandResult
cargo test --release (CPU)399 passed / 0 failed / 23 ignored unit (baseline 393; +5 registry, +1 allocator fence) and 10 / 0 / 6 integration (baseline 3; the new tests/backend_registry_cli.rs is 7)
cargo test --release --features cuda --test backend_registry_cli7 passed / 0 failed (exercises the #[cfg(feature = "cuda")] branch of the disabled-device gate)
cargo test --release --bin minfer -- --ignored --test-threads=1 (CPU)23 passed / 0 failed (baseline 23)
CUDA_HOME=/usr/local/cuda-13.0 … cargo build --release --features cudaexit 0
cargo test --release --features cuda --bin minfer -- --ignored --test-threads=1 (GB10 sm_121, cached 0.5B)25 passed / 0 failed. The ticket's "22" predates F7's two reference gates: 0602eb2 has 26 #[ignore] attributes, 2 of them in the macOS-only src/metal.rs, i.e. 24 runnable here, and this ticket adds the 25th (the fence gate below)
cargo test --release --features cuda --bin minfer -- --test-threads=1 (GB10, whole unit suite)452 passed / 0 failed / 25 ignored
MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf cargo test --release --features cuda --bin minfer -- --ignored --test-threads=124 passed / 1 failed at this commit, the one failure server::batch::tests::a_slot_snapshot_resumes_the_context_without_re_prefilling with "the file was written with the f32 KV element type, this run uses f16" — the then-pre-existing #130 C5 defect, closed 2026-09-25 (C5 S3: 31 / 0), which the record above documents as this configuration's 21/1 (the count moved by the same +3: F7's two gates and this ticket's)
registry / CLI gatesnames_resolve_and_unknown_names_are_refused, the_registered_set_and_priority_order_are_pinned, the_name_surface_fences_devices_and_keeps_cpu, the_packed_kv_capability_is_the_registrys_answer, registry_caps_match_the_backend_trait, alloc::tests::a_fresh_allocator_inherits_the_runs_backend_filter, tests/backend_registry_cli.rs (7, incl. a_disabled_device_is_not_refused_unless_it_was_named)
rustfmt --edition 2021 --check on the 19 changed .rsclean (stable's rustfmt 1.9.0: the pinned 1.97.1 toolchain has no rustfmt component installed here — CI runs no fmt job)
python3 scripts/check_docs_links.py935 links resolve in 183 files (baseline 930 / 182)

Mutation checks (all reverted, all restored byte-identical).

MutationGate that failedObserved
(a) an unknown name falls back to the default backendnames_resolve_and_unknown_names_are_refused + the_name_surface_fences_devices_and_keeps_cpucalled 'Result::unwrap_err()' on an 'Ok' value: CPU, and Ok(BackendFilter { allowed: [true, false, false], … })
(b) swap the CUDA and CPU prioritiesthe_registered_set_and_priority_order_are_pinnedleft: [("cpu", 200), ("cuda", 100)] vs right: [("cuda", 200), ("cpu", 100)]
(b2) perturb one priority value (200 → 201), order unchangedsame gateleft: [("cuda", 201), ("cpu", 100)] vs right: [("cuda", 200), ("cpu", 100)]

(b2) exists because the expected numbers are deliberately literals: deriving them from the PRIORITY_* constants would have made the value half of the gate vacuous (only a reordering would have failed).

Manual evidence (device, GB10; not a CI gate). With the CUDA build, MINFER_DISABLE_CUDA=1 minfer --backend cuda <model> and the MINFER_BACKENDS spelling both print the stage-2 message; the same with serve, viz, bench and specverify; --backend metal on Linux prints the not-compiled message; --backend gpu2 prints the unknown-name message. The fence itself was checked end-to-end: minfer --backend cpu … "The capital of France is" and MINFER_DISABLE_CUDA=1 minfer … on the same CUDA build produce byte-identical stdout (prompt + 8 greedy tokens) apart from the timing lines — i.e. naming cpu selects exactly the path the pre-existing disable flag selects.

Honest scope. (a) Metal is compile-only: there is no Mac here, so the Metal entry's pool/host_read/kv_format/enable hooks, its unavailable probe and its refusal text are covered by CI's build-macos job and by rustfmt, and by nothing that runs. (b) Behaviour preservation rests on the existing suites plus the new order/name gates; the parts a suite would not notice are pinned explicitly — the identity order (the on-disk session tag, the exporters), the priority order and numbers, Ord, the Debug spelling, and the equality of the registry's capability matrix with the trait's. (c) One deliberate, unreachable behaviour delta: fill_input / write_pool name a disabled backend in an Err where the old code panicked with expect("… pool not enabled"); assignment can never select a backend whose pool is not enabled, so no suite path reaches it. (d) copy_kv_to_cpu keeps its pre-F4 CPU/CUDA-only shape (a CPU identity-debug helper) rather than widening to Metal through the new hook — widening is not this ticket's job. (e) One per-backend branch remains on purpose: the MINFER_TRACE/viz capture path, where each backend captures through its own mechanism (a borrowed read, a Metal blit, an async D2H into pinned CUDA staging) — that is machinery, not a capability. (f) The registry makes the backend set data, not pluggable: there is no dlopen path, so a Vulkan/remote backend is registered by editing the crate, not by dropping in a library (ROADMAP §2.6 keeps that distinction). (g) The CPU numbers above are the final tree (after the rustfmt pass and the stage-2 fix); the CUDA numbers are from the same final tree.

F5 — Async cross-backend copies and events (#58) — DONE 2026-09-24

What landed. The scheduler's split boundary is now two registered phases instead of one blocking host round trip. GraphAllocator::copy_across (phase A, enqueue) resolves the destination staging buffer, marks the entry pending, and calls the source backend's BackendEntry::copy_cross hook — the F4 rule applies: the allocator never asks "is this the CPU?", it asks the entry. GraphAllocator::await_cross (phase B, wait) calls the entry's await_cross, clears the pending flag and counts the wait. The scheduler walks Split::inputs once for the copies and once for the waits, before any node of the consuming split runs; the consumer resolves a staged input through the new GraphAllocator::cross_input, which refuses a still-pending entry with a loud Err naming the missing wait. The registry contract, the per-backend table and the enumerated synchronization points are in BACKEND-REGISTRY-DESIGN.md §11; the scheduler-side summary is COMPUTE-GRAPH-DESIGN.md §3.4.

The hot path, defined and measured. The ticket is about the split-boundary staging copies — the per-Split::inputs transfers execute makes when a value produced on one backend is consumed on another; they run on every forward, on the critical path of every partially offloaded decode step. The copies that are legitimately host-visible and therefore not in scope are enumerated in the registry doc §11.1 (logits readback, KV-session save, debug/trace dumps, weight/tokenizer load, fill_input) so "zero blocking boundary copies" is not read as "zero device→host copies anywhere". Before F5 a single CUDA→host boundary input cost two host stalls — a full stream synchronization inside copy_to_host and then a blocking cudaMemcpy D2H — which is why the counters below count both.

Instrumentation. graph/copystats.rs holds the per-allocator counters (copies, waits, blocking_host_copies, async_host_copies, event_syncs, stream_waits) and the MINFER_SYNC_COPIES=1 switch that restores the pre-F5 path as the bitwise reference (with set_sync_for_test as its programmatic form). Device-level twins: CudaBackend::blocking_readback_count() (actual blocking cudaMemcpy D2H calls) and cuda::stream_sync_count() (host stalls).

The device layer. src/cuda.rs gains the event primitives (record_event, wait_event, stream_wait_event, event_destroy) and the async transfer (copy_to_host_async, host_alloc/host_free for the pinned slabs) — every failure is a loud Err naming the cudaGetErrorName, never a silently missing synchronization. CudaBackend owns a small grow-on-demand pinned-slab pool and the pending-copy table; a slab is released when its event has been waited on, and Drop destroys the events and frees the slabs.

Measured acceptance (GB10 sm_121 unless noted).

CommandResult
cargo test --release (CPU)405 passed / 0 failed / 23 ignored unit (baseline 399; +2 copystats, +2 allocator, +2 scheduler) and 10 / 0 / 6 integration (unchanged)
cargo test --release --bin minfer -- --ignored --test-threads=1 (CPU)23 passed / 0 failed (unchanged — every F5 gate that needs a boundary is device-gated or in the unit suite)
cargo test --release --features cuda -- --test-threads=1 (GB10, whole unit suite)459 passed / 0 failed / 26 ignored (baseline 452 / 0 / 25; +6 CPU-visible tests, +1 new device test, +1 new #[ignore]d real-model gate)
cargo test --release --features cuda --bin minfer -- --ignored --test-threads=1 (0.5B)26 passed / 0 failed (baseline 25 / 0, +the F5 real-model gate)
MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf … --ignored --test-threads=125 passed / 1 failed at this commit — the single failure was server::batch::tests::a_slot_snapshot_resumes_the_context_without_re_prefilling with "the file was written with the f32 KV element type, this run uses f16", i.e. the then-pre-existing #130, not this ticket (closed 2026-09-25, C5 S3: 31 / 0). The F5 gate itself passes in that configuration
rustfmt +stable --edition 2021 --check on the 10 changed .rsclean (rustfmt 1.9.0-stable; the pinned 1.97.1 toolchain has no rustfmt component here, as F4 found — CI runs no fmt job)
python3 scripts/check_docs_links.py935 links resolve in 183 files (baseline 935 / 183)

The numbers the ticket asked for. The real-model gate (async_cross_copies_never_block_and_stay_bitwise_identical) runs the same 4-of-24-block offload graph 7 forwards (1 prefill + 6 decode) in both modes and compares every step's logits:

async (default)sync (MINFER_SYNC_COPIES=1, the pre-F5 path)
staging copies (copies)5656
boundary waits (waits)5656
blocking device→host copies07
async device→host copies70
event waits (host, the documented point)70
device-level blocking readbacks07
full stream syncs2734
max |Δlogit| vs the other mode0 (bitwise)—

So the before/after is: 7 → 0 blocking device→host copies and 34 → 27 full stream synchronizations (−7, exactly one per staged device→host copy) over 7 forwards, with byte-identical logits. The cheap device gate (cuda_backend::tests::a_split_graph_waits_once_per_staged_copy_and_stays_bitwise) reports the same shape on a 3-node CPU→CUDA→CPU graph: async copies=2, waits=2, blocking=0, async_host=1, event_syncs=1, readbacks=0, syncs=1; sync copies=2, waits=2, blocking=1, readbacks=1, syncs=2.

Mutation evidence (reverted; the tree was restored byte-identical). The boundary's phase-B loop was removed from execute:

graph::cuda_backend::tests::a_split_graph_waits_once_per_staged_copy_and_stays_bitwise
  panicked: called `Result::unwrap()` on an `Err` value: "staged cross-backend input 0
  for Cuda was read before its boundary wait: every copy_across owes one await_cross (F5, #58)"

models::qwen2::graph::tests::async_cross_copies_never_block_and_stay_bitwise_identical
  panicked: called `Result::unwrap()` on an `Err` value: "staged cross-backend input 3
  for Cuda was read before its boundary wait: every copy_across owes one await_cross (F5, #58)"

Which half of the missing-wait gate is deterministic, and which is not. The loud refusal above is deterministic: the pending flag left by the missing wait is turned into an Err by cross_input on the consumer path, on the first read. The counter half (all_copies_awaited(), i.e. copies == waits) is also deterministic. The bitwise comparison against the synchronous reference is the probabilistic half: a dropped event wait can also corrupt bytes, but whether it does depends on timing, so it is reported as evidence of equality between the two supported modes, never as the failure mode of the mutation. That is why the gate carries both.

Honest scope. (a) Metal is unported, deliberately: there is no Mac here and no macOS toolchain, so its copy_cross declines (Ok(false)) and the allocator's synchronous host round trip handles it exactly as before F5 — no half-written blit/event code that nothing can compile. A Metal source's copies therefore still count as blocking_host_copies; the port is #137. (b) True overlap did not land and is not claimed: the split loop is strictly sequential (enqueue, then wait), so there is no independent work for a copy to overlap with, and a host-side consumer must wait by definition. What the substrate buys today is that the transfers are enqueued back to back and the redundant per-copy stream syncs disappear — the 34 → 27 measurement above. Deferring a wait to the consumer's first use landed as #138 (2026-10-04; the F5 S2 record is below). (c) Only CUDA has an async device path, and only the CUDA→host direction (a device→device staging copy is unreachable — copy_across early-returns when source and destination backends match; CUDA→Metal on a macOS+CUDA build declines and stays synchronous, the pre-F5 behaviour). (d) The CPU is a synchronous no-op by construction — it has no device memory — and the gate that says so is scheduler::tests::the_cpu_path_never_enters_the_cross_copy_machinery: a CPU-only graph is one split, both modes are bitwise identical and every counter stays zero. (e) The latency claim is a count, not a timing: the per-copy full cudaStreamSynchronize is gone (measured as a sync count) and no blocking cudaMemcpy is issued; no wall-clock speedup is claimed, because at a boundary the consumer immediately needs the bytes and the dominant cost is unchanged. (f) The end-to-end counter assertion needs a real cross-backend boundary, so it runs on the CUDA device (#[ignore]d); CI (CPU-only) covers the allocator-level missing-wait invariant plus the graph-level refusal through execute with an injected pending entry (scheduler::tests::a_staged_boundary_input_cannot_be_consumed_before_its_wait), and the CPU no-op/bitwise gate. The mutation's CUDA failure output above is the record of the part CI cannot reach.

F5 S2 — the cross-backend wait is deferred to the consumer (#138) — DONE 2026-10-04

The ticket. #138 is F5's overlap follow-up: F5 enqueued one staging copy per Split::inputs entry and then waited on every one of them at the split entry, so no copy could overlap anything, and the boundary also paid a full cudaStreamSynchronize to retire a producer whose work the copies were already stream-ordered behind.

What landed.

  • src/graph/scheduler.rs — the boundary calls GraphAllocator::retire_backend(previous) instead of sync_backend(previous), its phase-B loop is gone, the node loop resolves each source through GraphAllocator::cross_input_ready, and drain_cross_pending() runs after the last split.
  • src/graph/backend.rs — a new Backend::retire, whose default body is synchronize (Metal's boundary is a submission, so it still blocks there) with the reason the CUDA override is safe.
  • src/graph/cuda_backend.rs — CudaBackend::retire closes the capture window and drops the MMQ memoization without the cudaStreamSynchronize; close_capture_or_sync(block) is that one difference; CudaBackend::cross_inflight_peak() is the device-side overlap metric.
  • src/graph/alloc.rs — cross_input_ready (the deferred wait), drain_cross_pending, retire_backend, and a copy_across that is idempotent while an entry is in flight.
  • src/graph/copystats.rs — CrossCopyStats::deferred_waits.

The reading of the ambiguous acceptance line, recorded here and in the issue comment. "A boundary with several staged inputs waits once, at the first actual use, not once per input at the split entry" is read as each staged entry's single wait is issued at its first use, not as one wait per boundary: F5's documented contract is one phase-B call per phase-A copy (BACKEND-REGISTRY-DESIGN.md §11.2), and the ticket's fourth acceptance line keeps GraphAllocator::cross_input's refusal of a pending entry, which a per-boundary collapse of the wait count would have to weaken. The counters show the deferral: deferred_waits == copies on the deferred path.

The finding the deferral exposed (fixed in the same commit). With the wait deferred, the same (graph, node, destination) can be staged by two boundaries of one execution while the first copy is still in flight — a state unreachable under F5, which always awaited before re-enqueuing. Re-issuing it duplicated the transfer and pushed a second CrossPending record into CudaBackend that the allocator's single pending key would never wait on: a leaked pinned slab and event. Measured on the 0.5B gate before the fix: 56 copies / 47 waits, 18 in-flight re-requests and 6 drains of 4. copy_across now returns early while its entry is pending (same staging buffer, one unchanged source node), so the contract holds again and copies means unique transfers per execution.

Measured acceptance (box dgxspark (aarch64, GB10 sm_121), 2026-10-04). The bar was named from the F5 record (34 → 27 stream syncs, 7 → 0 blocking D2H). On this box the pre-change tree at 6b6d94f measures 21 stream syncs for the same gate and mode (the ticket's 27 is the 2026-09-24/27 measurement), so the same gate's printed output is the before/after:

F5 (measured at 6b6d94f)#138
staging copies (copies)5647 (9 in-flight re-requests deduped)
waits (waits)5647
of which deferred to a consumer read035
blocking device→host copies00
device-level blocking readbacks00
full stream syncs210
max |Δlogit| vs MINFER_SYNC_COPIES=100

The cheap device gate (cuda_backend::tests::a_split_graph_waits_once_per_staged_copy_and_stays_bitwise, 3 nodes CPU → CUDA → CPU) reads: async copies=2 waits=2 deferred=2 blocking=0 async_host=1 event_syncs=1 readbacks=0 syncs=0 against the F5 reading syncs=1; sync mode copies=2 waits=2 blocking=1 readbacks=1 syncs=1 (F5: 2).

The new overlap gate (cuda_backend::tests::a_boundary_with_several_staged_inputs_defers_its_waits) makes the device split produce two values the CPU split consumes, with an independent CPU node between the boundary and the first staged read. Two deterministic device-side metrics:

  • in-flight copies — the pinned-slab high-water mark is 2 for the deferred boundary, while the F5 enqueue-then-wait discipline, driven by hand on the same buffers in the same test, cannot exceed 1 (the wait releases the slab before the next copy takes one). This is the "enqueuing copy N+1 while copy N is in flight" measurement.
  • host stalls — stream_syncs is 0 for the deferred boundary and 2 for the synchronous reference over the same two device→host copies; copies == waits == 3 and deferred_waits == 3, bitwise against the reference.

The gate also asserts the dedup directly: two copy_across calls for one pending entry count as one copy and owe one wait.

Mutation evidence (reverted; the tree was restored byte-identical). The node loop's resolver was replaced with a bare cross_input — i.e. the deferral removed, the F5 read path restored:

graph::cuda_backend::tests::a_split_graph_waits_once_per_staged_copy_and_stays_bitwise
  panicked: "staged cross-backend input 0 for Cuda was read before its boundary wait: every
  copy_across owes one await_cross (F5, #58)"
graph::cuda_backend::tests::a_boundary_with_several_staged_inputs_defers_its_waits
  panicked: "staged cross-backend input 0 for Cuda was read before its boundary wait: …"
graph::scheduler::tests::a_staged_boundary_input_is_waited_on_at_its_first_use
  panicked: "staged cross-backend input 1 for CPU was read before its boundary wait: …"
models::qwen2::graph::tests::async_cross_copies_never_block_and_stay_bitwise_identical
  panicked: "staged cross-backend input 3 for Cuda was read before its boundary wait: …"

Suite counts. cargo test --release 481 / 0 / 36 unit + 10 / 0 / 6 integration (baseline 480 / 0 / 36; graph::alloc::tests::the_drain_waits_on_a_staged_entry_nothing_read is the new feature-independent gate). scripts/cuda_test.sh 567 / 0 / 42 (baseline 565 / 0 / 42; the same alloc gate plus cuda_backend::tests::a_boundary_with_several_staged_inputs_defers_its_waits; no #[ignore]d test moved). FEATURES=cuda scripts/real_model_gates.sh 42 / 0 on the cached 0.5B and 42 / 0 with MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf. cargo build --release --features cuda and cargo test --release --features cuda --no-run are clean, and cargo fmt --all --check is clean on the pinned 1.97.1 toolchain.

Honest scope. (a) Metal is still unported (#137): its copy_cross declines, Backend::retire's default keeps its boundary blocking, and its copies still count as blocking — no half-written blit/event code. (Superseded by the F5 S3 record below, 2026-10-05, which ports it; the sentence is kept as the state this record measured.) (b) The cudaStreamWaitEvent device-consumer arm is not reached: the ticket's first bullet asks for it, but the only pair that could express a device destination would need two device backends, copy_across early-returns on a same-backend pair, and CUDA→Metal (macOS + CUDA) declines phase A — so a call site would be unreachable code, exactly what Core Convention 5 asks about rather than silences. The observable requirement of that bullet — "a device consumer does not block the host" — is met by the CPU→device direction, whose fill is already stream-ordered on the destination pool's own stream (§11.3). The two grandfathered bare sites therefore stay dead in every configuration, and this PR leaves GRANDFATHERED_BARE and docs/dead-code-baseline.toml unchanged; the stripped oracle confirms it below. (c) True cross-split overlap is still not claimed — the split loop remains sequential and #300 owns the measured overlap or the recorded negative result, delivered in the F5 S4 record below (the negative result; the wait is still deferred to where the data is read, not where it was produced).

The dead-code ratchet, measured. python3 scripts/check_dead_code_annotations.py reports the same 11 grandfathered bare sites (this PR adds no annotation and tightens none), and the stripped oracle run with RUSTFLAGS=--cap-lints=warn and RUSTUP_TOOLCHAIN unset reports 0 additions in both configurations — the (name, kind) set difference against docs/dead-code-baseline.toml is empty, so neither site became live and neither entry went stale.

F5 S3 — the async staging copy ported to Metal (#137) — DONE 2026-10-05

The ticket. #137 is F5's Metal half: the BackendEntry::copy_cross / await_cross hooks had a CUDA implementation and a Metal one that declined phase A (Ok(false)), so a Metal source kept the pre-F5 synchronous host round trip and its boundary copies counted as blocking_host_copies. The port needs a MTLBlitCommandEncoder copy plus an event, the mechanism chosen and justified, and the same counter / bitwise / loud-Err acceptance the CUDA side has.

What landed.

  • src/metal/encode.rs — MpsCommandBuffer::encode_blit_signal (a blit copy of the source window into the staging buffer plus encodeSignalEvent) and command_buffer().
  • src/metal/runtime.rs — MpsState::new_shared_event(); src/metal.rs — the MetalSharedEvent alias.
  • src/graph/metal_backend.rs — MetalBackend::cross_enqueue / cross_take, the copy_cross / await_cross hooks, cross_pending_len and sync_readback_count.
  • src/graph/scheduler.rs — phase A for a Metal→CPU boundary runs before retire_backend, so the blit lands in the producer split's own command buffer.
  • src/graph/alloc.rs — the cross_stats_mut / cross_buffer cfgs widened to target_os = "macos".

Mechanism choice. MTLSharedEvent, not MTLEvent and not a completion handler: a completion handler can only notify the host, and a plain MTLEvent has no host wait; MTLSharedEvent is the one primitive that both encodeSignalEvent / encodeWaitForEvent accept and that exposes a bounded host wait. Phase A enqueues (blit + signal) and never blocks; phase B's single wait is waitUntilSignaledValue:timeoutMS: with a 10 s bound, at the consumer's first read. The device-consumer arm (encodeWaitForEvent) is reserved and unreachable (see Honest scope).

The finding the port forced (fixed in the same commit). The first cut gave each staging copy its own command buffer, committed after retire. On this box the same mixed 0.5B graph then differed between async and MINFER_SYNC_COPIES=1 by max |Δlogit| 1.46, while each mode compared with itself was bitwise stable. Reading the source at enqueue and comparing it with the staging buffer at the wait showed the staged bytes were identical, so the divergence was the extra in-flight command buffer perturbing Metal kernel execution (the backend already has 20 red tests on this box — a red baseline makes the mode difference ambiguous). Encoding the blit into the producer split's own command buffer — copy_across before retire for a Metal→CPU boundary — restores the one-command-buffer rule and makes the comparison bitwise.

Measured acceptance (box macbook (macOS 27.0.1, Apple M4 Pro), 2026-10-05).

cargo test --release --bin minfer async_cross_copies_never_block_and_stay_bitwise_identical_on_metal -- --ignored --test-threads=1:

[f5-metal] async: copies=35 waits=35 deferred_waits=35 blocking_host_copies=0 async_host_copies=21 event_syncs=21 stream_waits=0 sync_readbacks=0
[f5-metal] sync : copies=35 waits=35 blocking_host_copies=21 async_host_copies=0 sync_readbacks=21
[f5-metal] 6 decode steps + 1 prefill over 4/24 Metal blocks: max |Δlogit| = 0; blocking host copies 21 -> 0, device readbacks 21 -> 0

The cheap device gate (graph::metal_backend::tests::staging::a_split_graph_waits_once_per_staged_copy_and_stays_bitwise, 3 nodes CPU → Metal → CPU) reads async copies=2 waits=2 deferred=2 blocking=0 async_host=1 event_syncs=1 sync_readbacks=0 against sync copies=2 waits=2 blocking=1 sync_readbacks=1.

Mutation evidence (each caught by a named gate).

mutationgateresult
drop a staged input's wait (read cross_input directly)graph::metal_backend::tests::staging::a_staged_entry_read_before_its_wait_is_refusedloud Err "… read before its boundary wait …"
re-request (graph, node, dst) while in flight…staging::a_re_request_while_in_flight_is_the_same_transferone copy, one wait, one record (cross_pending_len() == 1), no leak
make a blit/wait end in a failing status (MINFER_TEST_CALL_FAIL=metal_cross_copy)…staging::a_timed_out_cross_wait_is_a_loud_errErr naming the value waited for, the observed signaledValue, and the command buffer status

Suite counts. cargo test --release 490 / 20 / 39 unit (baseline 486 / 20 / 38; +4 Metal staging gates, +1 #[ignore]d real-model gate; the 20 failures are the box's pre-existing red baseline and are exactly the same set). scripts/real_model_gates.sh (the default parallel form on a CPU-only build) 37 / 2 on the cached 0.5B and 37 / 2 with MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf, against the baseline 37 / 1; the extra failure is server::batch::tests::prefill::a_long_prefill_keeps_another_slot_decoding, a pre-existing Metal path (the batch engine offloads all 24 blocks and hits the G5 copy_cells refusal / a CPU index bound), not a boundary copy. The new Metal gate passes in isolation — cargo test --release --bin minfer async_cross_copies_never_block_and_stay_bitwise_identical_on_metal -- --ignored --test-threads=1 → 1 passed — and one Qwen3 parallel run was 36 / 3 with the gate itself at Δ = 0.285: other concurrently-running tests submit to the same MpsState command queue without taking crate::metal::metal_test_lock() (the lock's own doc comment records that this drops kernel writes), so the gate can be corrupted by the shared harness. cargo build --release and cargo fmt --all --check are clean on the pinned 1.97.1 toolchain.

A cross-process CLI check is not usable on this box, and is recorded as such. minfer --gpu-layers 4 -n 8 --greedy <0.5B q4_0> "The capital of France is" is stable within one mode across repeats but produces different continuations in async and MINFER_SYNC_COPIES=1 runs, and neither matches the --backend cpu/--gpu-layers 0 reference ("The capital of France is Paris. It") — the pre-existing Metal graph-graph divergence the box's 20 red tests already record. The controlled same-process async-vs-sync comparison (both arms on the same loaded engine, mode toggled with copystats::set_sync_for_test) is bitwise and is the acceptance evidence.

Honest scope. (a) The device-consumer arm (encodeWaitForEvent) is reserved, not exercised: Metal→Metal early-returns and CUDA does not run on Apple Silicon, so no device→device pair exists on macOS. (b) Not verified on CUDA: the scheduler reorder is not #[cfg]-gated but only the Metal branch takes it (pb == BackendTag::METAL), so a CUDA build's order is unchanged; this box has no CUDA device and CI only compiles the CUDA harness. (c) The 20 red Metal tests on this box were not fixed and are not claimed to be. (d) The port adds no cross-split overlap — the split loop is still sequential; the wait is deferred to the consumer read, as on CUDA.

F5 S4 — true cross-split overlap: the recorded negative result (#300) — DONE 2026-10-09

The ticket. #300 is the follow-up the F5 S3 record names: after #137 the boundary blit rides the producer's command buffer, so the copy is asynchronous but serial with respect to the producer and the consumer. It asked for either a genuinely overlapping copy with an explicit dependency (a separate boundary command buffer that waits on a producer event) and evidence, or a dated, reproducible negative result with the mechanism.

The result is negative, and the mechanism is measured on macbook (macOS 27.0.1, Apple M4 Pro) at ad707c7 (2026-10-09), recorded in 9df405d. Two conclusions, both with numbers.

(1) The #137 divergence is a missing-dependency race, not a kernel perturbation. The F5 S3 record inferred from the then-red baseline that "the extra in-flight command buffer perturbed Metal kernel execution" (max |Δlogit| ≈ 1.46). Isolated, the pathological first cut — a standalone boundary command buffer submitted from copy_cross before retire submits the producer's buffer, with no dependency on it — is deterministic: on the real-model gate it reads max |Δlogit| = 26.718678 at step 0, the same value twice. The mechanism is the shared MpsState command queue's commit order: the standalone blit buffer is committed first, so it reads the producer's StorageModeShared source window before the producer's kernels wrote it, and the consumer waits on an already-signaled event holding the previous forward's bytes. The F5 S3 "staged bytes identical" comparison was stale-to-stale.

(2) A separate boundary buffer with an explicit dependency is correct — but buys no measurable cross-split overlap, so the production design stays. Committing the boundary blit on its own MTLCommandQueue with encodeWaitForEvent on an event the producer signals at its split end (then signalling its own event for the consumer) restores bitwise identity: the real-model gate reads max |Δlogit| = 0 over the prefill + 6-decode loop (twice), counters unchanged — copies=35 waits=35 deferred_waits=35 blocking_host_copies=0 async_host_copies=21 event_syncs=21 sync_readbacks=0 async against copies=35 waits=35 blocking_host_copies=21 sync_readbacks=21 sync — and the cheap 3-node gate passes (copies=2 waits=2 deferred=2 blocking=0 async_host=1 event_syncs=1 sync_readbacks=0). Cross-split overlap, however, needs a second split to overlap with, and the reachable macOS topology has none: the E5 mixed graph (4 of 24 blocks on the device) is a single CPU → Metal → CPU sequence, so there is exactly one device→host boundary per forward; its consumer is the host CPU split, whose first node reads the dominant staged tensor (the hidden state), so the copy is on the consumer's critical path; the other two staged tensors are read within the CPU split's first eight nodes (add at node 46, cells at 52, attn_span at 54), so a separate buffer could only overlap a blit of two tiny index/window tensors. A device→device pair (Metal→Metal) would have real overlap, and copy_across early-returns on it while no CUDA device exists on macOS. The boundary blit therefore keeps riding the producer's command buffer.

Reproduction. Both cuts are MINFER_300_*-gated temporary patches to MetalBackend::cross_enqueue (a standalone cmd_buffer() + submit() before retire; a second queue + encodeWaitForEvent / encodeSignalEvent), neither retained in the tree. Every arm runs the same one command:

cargo test --release --bin minfer async_cross_copies_never_block_and_stay_bitwise_identical_on_metal -- --ignored --test-threads=1 --nocapture

The production (blit in the producer's buffer) arm is max |Δlogit| = 0 with the counters above. The full write-up is docs/BACKEND-REGISTRY-DESIGN.md §11.6.

Suite counts. This record is documentation only — no production or test code changed. cargo test --release --no-fail-fast on macbook (macOS 27.0.1, Apple M4 Pro), 2026-10-09: 553 passed / 0 failed / 45 ignored unit + 21 / 0 / 6 integration (the macOS baseline of 2026-10-08, unchanged).

Honest scope. The Option A cut was validated in-process (bitwise) and its counters read; the overlap claim was not measured as a wall-clock win because the structural target is negligible — that is the point of the negative result, not a gap in it. The second queue and the explicit event dependency are not retained: keeping a second queue alive for a copy that cannot overlap the producer and overlaps almost nothing else would be complexity without a measured benefit.

F6 — Quantizer tooling: convert, quantize, split (#49) — DONE 2026-09-24

What landed. Four new modules and the subcommands that use them:

  • src/gguf_write.rs — the GGUF v3 writer. The reader in src/gguf.rs is the contract: magic/version/n_tensors/n_kv, the KV encoding for every one of the 13 GgufTypes (including arrays of strings and of the numeric types), the tensor index (name, n_dims, ne, type, offset), ggml_pad alignment before the data section and after every tensor, and the multi-part convention (split.no / split.count / split.tensors.count, {stem}-NNNNN-of-MMMMM.gguf). write_split assigns tensors to parts greedily by padded size, never splits a tensor, gives every part the full metadata (so each parses standalone), and preserves the global tensor order so the merged index equals the single-file index; a one-part assignment is written as a plain single file with no split.* keys at all. The writer validates shapes/names/payload sizes and refuses a wrong-length payload instead of shifting every later tensor.
  • src/quantize.rs — the weight encoders. q4_0, q4_1, q5_0, q5_1, q8_0 (plus the f16/f32 element casts), implemented as llama.cpp's quantize_row_*_ref (the CPU reference), with QuantTarget::parse refusing every other GGUF type by name ("known GGUF type but minfer has no weight encoder for it … writing it would emit wrong weights").
  • src/convert.rs — the HF → GGUF converter (safetensors parsed as a length-prefixed JSON header plus raw bytes; serde_json only, no Python and no ML framework) and QuantizePlan (re-encode an existing GGUF, with the llama.cpp "1-D tensors stay f32" rule and the tied-embedding → q8_0 policy). The metadata it writes is what the F7 strict tokenizer/template loader reads: tokenizer.ggml.model = gpt2, .pre = qwen2, the full vocab_size token array with llama.cpp's token types, merges, special ids, add_bos_token and tokenizer.chat_template. Unknown architecture, tensor name or dtype is a refusal.
  • src/tooling.rs — convert / quantize / split and the F6 gates.
  • One engine change the acceptance forced: f16 weights now run. The CPU graph path had no f16 weight dispatch at all (only f32 and the quants), so vec_ops::mat_mul_f16 decodes a weight row at a time and cpu_backend dispatches it for Op::MatMul, and Op::GetRows decodes f16 embedding rows. Without it "a converted model produces the same logits" was impossible for the f16 output the ticket names.
  • download::check_downloaded_size — the second acceptance line. http_download used to ignore the expected size it was passed; now a downloaded file whose length differs from the remote's is an error and is removed, so it cannot be mistaken for a cached complete file.

Measured acceptance (CPU; GB10 sm_121 for the CUDA row). The references and tolerances are stated in docs/GGUF-TOOLING.md §4.

CommandResult
cargo test --release (CPU)432 passed / 0 failed / 28 ignored unit (baseline 405 / 0 / 23; +27 F6 tests, +5 F6 real-model gates) and 10 / 0 / 6 integration (unchanged)
cargo test --release --bin minfer -- --ignored --test-threads=1 (CPU)28 passed / 0 failed (baseline 23; +5 F6 real-model gates)
cargo test --release --features cuda -- --test-threads=1 (GB10 sm_121)486 passed / 0 failed / 31 ignored (baseline 459 / 0 / 26)
cargo test --release --features cuda --bin minfer -- --ignored --test-threads=1 (0.5B)31 passed / 0 failed (baseline 26 / 0, +the 5 F6 gates)
… with MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf30 passed / 1 failed at this commit — the single failure is the then-pre-existing #130 (the file was written with the f32 KV element type, this run uses f16), as at the baseline; closed 2026-09-25 (C5 S3: 31 / 0)
the F6-produced q8_0 file on the deviceCUDA GATE: … is not printed; offload: all 24 blocks + embed/output on cuda (500.8 MiB of device weights), 1214.7 tok/s prefill, greedy Paris. — the same text as MINFER_DISABLE_CUDA=1. The f16 file on the same build prints weight 'token_embd.weight' (type F16) has no CUDA kernel or is not registered and running on CPU (#141)
HF → GGUF vs convert_hf_to_gguf.py (Qwen2.5-0.5B-Instruct, bf16 → f16)290/290 tensor payloads byte-identical (sha256 per tensor); metadata equivalent for every loader-read value (llama.cpp also writes the cosmetic general.size_label, and key order differs)
minfer vs llama.cpp f16 file, logits after a 4-token greedy continuationbitwise identical (assert_eq! over 151,936 logits) and identical greedy text [12095, 13, 1084, 374]
rewrite a cached GGUF with the writer291/291 tensors byte-identical, metadata key-for-key equal, logits bitwise identical
split the cached 0.5B into 4 partsmerged index exactly the single-file index (name/shape/type/nbytes, in order); logits bitwise identical; missing part / wrong split.no / filename-vs-split.count mismatch each fail the load
minfer quantize vs llama-quantize on the same f16 sourceq4_0, q4_1, q5_0, q5_1, q8_0 each 290/290 tensors byte-identical
f16 → q8_0 end-to-endgreedy continuation identical; max |Δlogit| 0.481 (mean 0.082) against max |logit| 18.43 — stated bound ≤ 1.0 absolute
download size gatepure test plus an end-to-end local HTTP server (correct 206 resume accepted; a server that ships the whole body for a range is rejected and the file removed)
rustfmt +stable --edition 2021 --check on the 8 changed/new .rsclean (rustfmt 1.9.0-stable; the pinned 1.97.1 toolchain has no rustfmt component here — CI runs no fmt job)
python3 scripts/check_docs_links.py940 links resolve in 184 files (baseline 936 / 183; the new doc + its registration)

The FPE detail worth recording. The first q4_0 encoder differed from llama-quantize in 12 of 64512 bytes on one tensor — every difference a single nibble off by one. The cause was not the algorithm but the compilation: llama.cpp's reference is compiled with -ffp-contract=fast, so x*id + 8.5f becomes an FMA on aarch64/x86, while Rust's two separate operations rounded twice. f32::mul_add reproduces it and the difference went to zero. Q8_0 uses a single multiply and matched without it. A quantizer that is "close" is a wrong file, which is why the gate is per-tensor byte equality.

Mutation evidence (reverted; the tree was restored byte-identical). Each new gate was broken in the way it guards and observed to fail:

Mutation (reverted after each run)Gate that failed (observed output)
writer pads between tensors to 16 while the index declares 32gguf_write::tests::tensor_index_offsets_match_the_parser_requirement — left: 880, right: 896: the file is 16 bytes short of the layout its own index declares
tensors assigned to split parts in reverse ordergguf_write::tests::split_writes_parts_the_reader_merges_back — left: 3, right: 1, part 0 holds t2
split.count written as 1 for a 3-part splitthe same test panics inside load_gguf_model: split count mismatch: filename implies 3 parts, split.count = 1
q4_0 encoder without mul_addf6_quantize_encoder_is_byte_identical_to_llamacpp — tensor blk.0.attn_k.weight payload differs (the 12 bytes)
check_downloaded_size returns Ok unconditionallya_wrong_size_file_is_refused_and_a_right_size_file_is_accepted and http_download_resumes_a_partial_file_and_size_checks_it both FAILED (the oversized / partial file is accepted)

The reverts were byte-identical (sha256sum of each mutated file before and after matches HEAD). Two of the gates had to be strengthened to catch their mutation, which is the point of the exercise: the first padding gate only checked the index the writer declares (the parser recomputes the same offsets and never reads past the last tensor, so a short inter-tensor pad parsed fine), so the test now also asserts file length == data_offset + size — for the single file and for each split part; and the first reverse-order mutation landed in validate_specs' loop rather than split_assignment's, so it proved nothing until the anchor was made specific. The table above is the re-run output against the committed tests.

Honest scope. (a) K-quants cannot be written. The encoders are the five legacy types; q4_K/q5_K/q6_K and every I-quant are refused by name (#140). Reading them is unchanged. (b) f16 weights are CPU-only. Superseded: #141 registered them on CUDA, #164 on Metal (2026-10-06), so the device row below is history, not a current limitation. As written at F6: the CPU path decodes them, but the Metal/CUDA weight registration accepts f32 and the supported quants only, so an f16 model runs the CPU path on a device build and that path is slow (measured ~3 tok/s prefill on the 0.5B); the device row therefore exercises a minfer quantize-produced q8_0 file (#141). (c) bf16 output was refused at F6 and landed in #142 (--outtype bf16, 2-D bf16 / 1-D f32, round-to-nearest-even, byte-identical per tensor to llama-quantize --pure <f32>.gguf … BF16), together with the CPU bf16 weight path; CUDA and Metal still do not register bf16 — both do since #208 (CUDA PR #321, Metal PR #323, 2026-10-06). (d) The exactness claims are named per step: f16/f32 copies and f16→f32, bf16→f32 are bit-exact; bf16→f16 is exact in the mantissa but not in the exponent range — it can overflow to inf, and below f16's smallest normal it rounds onto the subnormal grid (measured: 123 024 values on the 0.5B checkpoint, #142); f32→f16 is not exact and neither is f32→bf16 (RNE). (e) The HF reference is llama.cpp's converter output (byte-identical weights) plus, for logits/text, llama.cpp's own run — transformers was installed but not used as the logit reference, because the f16-vs-f16 byte comparison against llama.cpp's converter is the stronger claim. (f) The five real-model gates are #[ignore]d (they need the checkpoint and/or the cached 0.5B); CI covers the writer/encoder/converter/download unit and local-HTTP gates.

#205 (2026-10-07) — the fixture chain has a record, and the gates check it. The F6 gates compare against files under ~/.cache/minfer/f6-src/ that they do not produce, so a stale or replaced reference used to be compared against silently. docs/f6-fixtures.json now records one entry per artifact content identity — path, bytes, sha256 (or a sha256_prefix where the 2026-10-07 table truncated it), the exact producer command, the producer's identity (minfer commit, or llama.cpp commit + compiler + effective -ffp-contract + cflags), date and an absolute box label — and every path with two recorded contents carries a divergence_notes entry, which is how the two 2026-10-07 divergences became readable: the f16 source and its bf16 cast differ by minfer producer version (the dgxspark copy is the 2026-09-27 producer whose commit was never recorded; the Mac's is ab34a72, and bf16-from-f32 — whose f32 input is identical on both boxes — is byte-identical), while the ref/qwen2.5-0.5b-* references differ by the reference build's -ffp-contract (§4.2, #334). scripts/check_f6_fixtures.py audits the manifest's shape and the tree cross-check in CI (--check), verifies the whole cache where it exists (--verify), and carries its own tamper cases (--selftest); the F6 gates verify the fixture they resolve (src/tooling/tests/f6_fixtures.rs), so cargo test … --ignored refuses a tampered cache by name and digest. Measured on dgxspark (aarch64, GB10 sm_121), 2026-10-07: --verify → 36 recorded contents, the 21 present files all matching, the Mac's two f164/ entries reported absent; one flipped byte in a /tmp copy (qwen2.5-0.5b-q4_0.gguf, a7→a6) → exit 1 naming the file, the actual 9fe6b5e0… and both recorded digests (04634958… and the Mac's 51c2b000…); the same tamper under MINFER_F6_CACHE → the gate panics with the same message; an unrecorded .gguf inside the cache root → refused as unknown. Three new unit tests (485/0/40 on dgxspark) and a sha2 dev-dependency, so the production graph is unchanged. Still owed: the 11 truncated Mac digests (#342), the hf/ main revision pin, f164/minfer-f16.gguf's producer, and a regeneration command.

#334 (2026-10-07) — the byte-parity claim is build-dependent. Re-running the gate on macbook (macOS 27.0.1, Apple M4 Pro) made it red (blk.0.attn_k.weight payload differs; 168 of 290 q4_0 tensors, every difference a single data nibble, zero scale bytes) against a llama-quantize built by Apple clang without -ffp-contract=fast, while dgxspark's GCC-built reference still reproduced 290/290. The flag is the whole difference, measured on dgxspark (aarch64, GB10 sm_121), 2026-10-07, one variable at a time from llama.cpp HEAD 050dde50c: -O3 -DNDEBUG reproduces the cached ref/qwen2.5-0.5b-q4_0.gguf (04634958…) and adding -ffp-contract=off yields ea94611e…. The gate now detects which of the two builds produced the reference it compares against — quantize::FmaContract threads the contract through every encoder, and quantize_row_with(…, Off) reproduces that uncontracted reference 290/290 — so it refuses loudly naming the build instead of reporting a payload difference; a reference that matches neither variant still fails as a payload difference. docs/GGUF-TOOLING.md §4.2 carries the qualifier next to the 290/290 number, the two-box hash table and what #205 still owes. Mutation evidence on dgxspark: the -ffp-contract=off reference → the named-build refusal; one flipped payload byte → tensor blk.3.attn_output.weight payload differs (1 of 451584 bytes) with build unmatched; and a refactor slip that moved nmax's sign in make_qx_quants → q6_K blk.0.ffn_down.weight payload differs (2567682 of 3575040 bytes), green again 290/290 after the fix (the round-trip unit tests did not see it).

#349 (2026-10-07) — the K-quant half of that claim is compiler-dependent, and the gate's verdict now says which build it is looking at. #342 had already rebuilt the Mac's llama-quantize from c479922ac with -DCMAKE_C_FLAGS_RELEASE="-O3 -DNDEBUG -ffp-contract=fast" (the flag confirmed in the object) and measured: the five legacy quads pass 290/290, and every K-quant fails as tensor blk.0.ffn_down.weight payload differs (2285 of 2451456 bytes) … build unmatched — q5_K 97 of 2996224, q6_K 269 of 3575040 — i.e. the reference matches neither minfer's Fast nor its uncontracted Off model. git log 050dde50c..c479922ac -- ggml/src/ggml-quants.c touches only the q3_K/i-quant encoders, so the residual is Apple-clang-vs-GCC codegen of the search quantizers (§4.2.1), not the flag. Two defects followed: §4.2/§4.2.1 stated the K-quant 290/290/310/310 without naming the compiler, and build unmatched read like an encoder defect while being indistinguishable from a corrupted file. The fix records the decision in the manifest and in the gate. docs/f6-fixtures.json gained authoritative_reference: true on the one entry per llama-quantize path the claim is asserted against — for every ref/… path the dgxspark gcc 13.3.0 -ffp-contract=fast content, which §4.2.1 now names next to the number and states the claim is conditional on (compiler + -ffp-contract + llama.cpp revision). scripts/check_f6_fixtures.py enforces it as S6 (exactly one marked entry, on a llama-quantize record built with the flag; five new selftest cases, 26 total), and the Rust reader exposes reference_record + a pure ParityVerdict::classify that turns the three byte comparisons into five outcomes: reproduces; flag mismatch (the #334 refusal); not the recorded content (the resolver of #205 refuses first, by name and digest); a recorded foreign build (matches neither model but its digest is recorded → [f6 parity] SKIP … naming the file's build, the authoritative build, the §4.2.1 reason and this box's cc --version, so the Mac's K-quants stop being a permanently red ignored gate — the §14 row 6 failure mode); and the authoritative build, not reproduced, or an unrecorded file, which still fail. Measured on dgxspark (aarch64, GB10 sm_121), 2026-10-07: the eight targets stay 290/290 with build -ffp-contract=fast, recorded reference gcc 13.3.0 … on the pass line; a freshly built -ffp-contract=off llama-quantize (sha256 ea94611e…, the #334 digest) → the flag-mismatch refusal; one flipped byte in a /tmp copy → check_f6_fixtures.py --file exit 1 naming both recorded digests and the actual one, and the same copy under MINFER_F6_CACHE → the gate's resolver refusal before any compiler verdict; four flipped bytes at the end of the q4_K reference (a /tmp copy the record does not name) → tensor blk.23.ffn_up.weight payload differs (4 of 2996224 bytes) with the unattributable failure, not a compiler claim; and a fabricated third record marked non-authoritative plus MINFER_F6_CACHE → the SKIP line with exit 0. Mutations: the classifier's unrecorded arm returning RecordedForeignBuild → the unit test fails (left: RecordedForeignBuild, right: Unattributable) and the /tmp gate run wrongly skips as "a different compiler"; reference_record_in returning the default → the unit test fails on the empty record. Both restored from cp backups. Suites: 486 / 0 / 40 unit + 10 / 0 / 6 integration and the F6 #[ignore]d set 6 / 0, all on dgxspark. --regenerate landed in #345; still owed: the bare :NNN citation sweep (#336).

#354 (2026-10-07) — the fixture manifest is relocatable, so the skip verdict is a test. The F6 parity gate resolved its record through a compile-time path (concat!(env!("CARGO_MANIFEST_DIR"), …)), so #349's ParityVerdict::RecordedForeignBuild arm — a reference whose digest is recorded but is not the path's authoritative_reference, which the gate turns into the loud [f6 parity] SKIP — could only be exercised by editing a tracked file under a cp backup. MINFER_F6_MANIFEST is the record's own override, beside MINFER_F6_CACHE, honoured by the gate's reader (src/tooling/tests/f6_fixtures.rs) and by scripts/check_f6_fixtures.py (--manifest still wins over it; --regenerate refuses the variable outright, because re-recording into a manifest chosen by the environment is the accident the override must not create). Both halves honour it loudly: a value that is set but empty, or a record that is missing, unreadable or malformed, refuses by name — manifest_path() / load_manifest() in the reader and resolve_manifest() in the checker name the variable, the path and the problem — and neither falls back to docs/f6-fixtures.json. The reader's parse cache is keyed on the path rather than a OnceLock (a OnceLock freezes the first resolution, which is what made the verdict untestable in-process), and a reentrant process-wide guard serialises the override window against every reader. Two feature-independent tests land: the_manifest_override_drives_the_whole_resolver drives verify → reference_record → classify through a fabricated two-content 4-byte record and gets AuthoritativeNotReproduced, RecordedForeignBuild, Unattributable and the resolver's tampered-cache refusal, and a_broken_manifest_override_refuses_by_name_instead_of_falling_back covers the empty, missing and malformed override; the checker's --selftest gains four cases (43 → 47). Measured on dgxspark (aarch64, GB10 sm_121), 2026-10-07: 490 / 0 / 40 unit + 10 / 0 / 6 integration and the F6 #[ignore]d set 6 / 0; the recorded-foreign-build verdict driven end-to-end from a /tmp manifest against the real cache (the q4_0 reference's last four bytes flipped, its digest recorded as an Apple clang non-authoritative content) → [f6 parity] SKIP q4_0: the reference matches neither minfer's FmaContract::Fast model nor its uncontracted model (tensor blk.23.ffn_up.weight payload differs (4 of 2451456 bytes)) … with exit 0, where #349's evidence needed the tracked-manifest edit; and MINFER_F6_MANIFEST at a malformed or a missing file → exit 101 naming the path and the problem (… does not parse: key must be a string at line 1 column 3, … is unreadable: No such file or directory). Mutations: the reader ignoring the variable (always resolving the tracked record) → both new tests fail (the authoritative content must verify against the override: Err("F6 fixture cache: … is not in …/docs/f6-fixtures.json …") and a missing override must refuse, not fall back); a silent read_to_string fallback in load_manifest → a missing override fell back: 36 entries; the checker ignoring the variable → two --selftest cases fail. All restored from cp backups. Docs: docs/GGUF-TOOLING.md §4.2.2 names the override beside MINFER_F6_CACHE.

#344 (2026-10-07) — the anchor checker can now see a moved target. check_doc_line_anchors.py (#266) proves an anchor's file exists, its line is in range and its adjacent symbol is still in that file; a commit that inserts a line in an anchored file leaves every anchor below it resolving and wrong, and the checker stays green. That happened twice in one round — #329's dispatch change moved 7 anchors off by one (walkthrough docs 07 and 14), #299's weight accounting moved 69 across eight docs by 1–3 — and both were caught by a hand-written old→new line map. scripts/check_anchor_drift.py is that map as a gate: it diffs HEAD against the merge base (git diff -U0 <rev>...HEAD), reads the anchors of both revisions through the same Checker / ANCHOR / FROZEN objects (it imports the #266 script rather than restating them), pairs them on the doc line with its numbers normalised — PR #343's own ":NNN" -> ":N" proof, applied per line — and fails a pair that does not carry the mapped numbers, naming it doc:line → target:old (now new). A cited endpoint the range deleted has no image in the map, so it is reported as ambiguous for a human (--strict promotes it); an untouched target, an already re-pointed pair and an empty range are silent. It runs in check-docs on pull_request against origin/<base_ref> (hence the fetch-depth: 0 that check_status.py already required) plus its own --selftest (stale, re-pointed, rewritten, frozen and unchanged cases, each a real two-commit repository). Re-measured on dgxspark (aarch64, GB10 sm_121), 2026-10-07, with the recorded commands: over #299's range (--head eef052b 252932a) it named 73 anchors — the 69 PR #343 fixed and 4 it missed (docs/ARCHITECTURE-ROADMAP.md:98 and :452, docs/SOURCE-LAYOUT-PLAN.md:56, docs/inference_e2e_walkthrough/03-model-dispatch-weights.md:705, each verified by content) — and over #329's (--head 3bc896c 3bc896c^) the 7, with both ranges green at their re-pointed heads; the third live case, Cargo.toml:38-43 × 2 in docs/METAL-OBJC2-MIGRATION-PLAN.md (--head 506b26c 387fe91 → (now 47-52)), is resolved by adding that completed-migration record to FROZEN: the anchors cite the [lints.rust] unexpected_cfgs block commit 6a382a3 deleted, so they cannot be re-pointed (re-pointing is not even defined — no revision of Cargo.toml ever held that block at 38-43). The existing checker's verdicts are untouched: 1537 anchors, 297→304 frozen and 1014→1007 checked is the whole delta, and the new entry is reported UNUSED FREEZE (the ratchet's informational arm, not STALE FREEZE). Still owed: the bare :NNN continuations on those same lines are invisible to both checkers and belong to #336's sweep.

F6b — f16 weights on the device backends + a vectorized CPU f16 dot (#141) — DONE 2026-09-25

What landed. The third item F6 left open, plus the discovery that the fused concat registration was already fine but the type gate was not.

  • CUDA: f16_f32_matmul_vec / f16_f32_matmul_scalar in cuda_kernels.cu (the same NR0/NSG unit mapping and token-in-block loop as the f32 kernels; __half22float2 + FMA, accumulation in f32) and embed_rows_f16 (one thread per output element). matmul_f32_ptr_layout gained the F16 arm and embed_rows_on_gpu the f16 gather; the loader registers TensorType::F16 raw — deliberately its own branch, not folded into the quantized matches!, whose q4_K dsc-plane gate has no type check and would expand misinterpreted bytes from an f16 tensor into a plane no kernel reads (#165). Both launchers read their own launch return through the #147 helpers and return non-zero, so a failed launch is an Err at the call site instead of joining the 65 unchecked <<<>>> sites of #162.
  • The design choice was a native kernel over dequant-at-registration, and the reason is the memory trade the ticket asks to state: an f16 GGUF exists to halve the weight stream, and dequantizing to f32 on the device would give most of it back — 0.5B: 948 MiB of device weights as f16 against ~1.9 GiB as f32 (14 GiB → 28 GiB for a 7B). The measured offload line for the 0.5B f16 file is 942.4 MiB. f16 is not an MMQ format (MMQ streams quantized bytes and the f16-wmma GEMM is the MINFER_MMQ=0 fallback for the quantized types), so an f16 prefill runs the f32-activation kernel at every nt rather than the int8 GEMM.
  • Metal is a deliberate refusal, not a partial port. The loader does not register f16 there, so Qwen2Graph::weights_on_gpu fails its all-or-nothing check and an f16 GGUF prints the loader's "weights are not usable there — running on CPU" line. Registering a weight type no kernel can consume would make the device claim true while the op ran the wrong (or no) kernel, which is exactly what that gate exists to prevent; a Metal f16 matmul/embed kernel cannot be verified from dgxspark (no Mac; CI's build-macos compiles the crate and nothing runs it). Filed as #164. (#164 landed on a Mac 2026-10-06 — kernel_f16_f32_matmul + kernel_get_rows_f16 exist, both loaders register f16, and an f16 GGUF runs on Metal; see the G8 row.)
  • CPU: vec_ops::dot_f16_f32 (AVX2 F16C _mm256_cvtph_ps / aarch64 baseline NEON FCVTL vcvt_f32_f16, f64 scalar oracle) and vec_ops::decode_f16_row, with mat_mul_f16 decoding each weight row once for nt > 1 and folding every token into it, and the row loop going to the shared worker pool (kernel::par_for, the same MIN_PARALLEL_MACS threshold the quantized matmul uses). The SIMD loops run the same FMA tree in the same order as vec_dot_f32, so the f16 dot is bit-identical to vec_dot_f32 over the decoded row — which is what makes the row loop safe to reorder and to thread: n never changes a value, asserted. MINFER_NO_F16_ROWB=1 keeps F6's per-(row, token) shape as the A/B control.

Measured acceptance (GB10 CPU; GB10 sm_121 for the CUDA rows).

CommandResult
cargo test --release (CPU)445 passed / 0 failed / 30 ignored unit (baseline 440 / 0 / 29; +5 unit tests, +1 #[ignore]d device gate) and 10 / 0 / 6 integration (unchanged)
PARALLEL=0 scripts/real_model_gates.sh (CPU, serial)30 passed / 0 failed (baseline 29 / 0)
PARALLEL=1 scripts/real_model_gates.sh (CPU, parallel)30 passed / 0 failed (baseline 29 / 0)
CPU f16 prefill, before → after (34-token prompt, medians of 3 interleaved rounds per mode)3.2 → 207 tok/s (10.72 s → 0.16 s). The vectorized dot alone is 25.3 tok/s; the row blocking alone measured within noise; the pool is the rest. Decode 2.2 → ~10 tok/s
cargo test --release --features cuda -- --test-threads=1 (GB10 sm_121)513 passed / 0 failed / 33 ignored (baseline 508 / 0 / 32) + 10 / 0 / 6 integration
FEATURES=cuda scripts/real_model_gates.sh (0.5B config)33 passed / 0 failed (baseline 32 / 0)
MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf FEATURES=cuda scripts/real_model_gates.sh33 passed / 0 failed (baseline 32 / 0)
compute-sanitizer --tool memcheck over the serial CUDA unit suite0 API errors over 513 passed / 33 ignored (357.69 s)
The f16 file on the device (minfer convert --outtype f16 on the Qwen2.5-0.5B-Instruct checkpoint, 994,156,352 B)offload report "all 24 blocks + embed/output on cuda (942.4 MiB of device weights)"; 169 f16 matmul + 1 f16 embed nodes assigned Backend::CUDA; device vs CPU max |Δlogit| 7.34e-5 (mean 1.26e-5), 4.0e-6 relative, max |logit| 18.43; greedy [12095, 13, 1084, 374] identical on both
the same file under llama.cpp (--temp 0)Paris. — the same greedy continuation minfer gives on CPU and CUDA
rustup run stable rustfmt --edition 2021 --check on the 6 changed .rsclean (rustfmt 1.9.0-stable; the pinned 1.97.1 toolchain has no rustfmt component here — CI runs no fmt job)
python3 scripts/check_docs_links.py940 links resolve in 184 files (unchanged — no new file, only absolute issue URLs)
cargo check / cargo rustc --emit=obj for x86_64-unknown-linux-gnuthe AVX2+F16C path type-checks and codegens (the cross build reaches the link stage, where no cc cross-linker exists); running it is CI's ubuntu job

The gate asserts placement directly, not through a timing. f141_f16_weights_run_on_the_cuda_device (in src/tooling.rs, #[ignore]d): (1) Qwen2Graph::device() is Cuda; (2) the offload report says all 24 blocks + embed/output are on the device; (3) every F16 matmul / embedding node the scheduler assigns is Backend::CUDA, counted by walking the built graph's CNode.backend — so a registered-but-unsupported type that supports_op routes to the CPU fails here instead of quietly measuring the CPU; (4) then the logits against the same file's Layers(0) CPU run, at |Δ| ≤ 0.01 and ≤ 1e-3 relative (measured 7.34e-5 / 4.0e-6), greedy identical.

Mutation evidence (reverted; the files restored byte-identical).

MutationGate that failed (observed output)
the loader's f16 registration gated offf141_f16_weights_run_on_the_cuda_device — left: Cpu, right: Cuda at assertion (1)
the f16 arm removed from matmul_f32_ptr_layoutthe same gate, panicking on the loud Err("cuda: weight type F16 has no f32-activation matmul kernel …") — not a silent fallback
__half22float2 replaced with zeros for half of each 8-element chunk in f16_f32_matmul_vecthe same gate — greedy continuation differs between device and CPU, so a value-level fault cannot pass either
dot_f16_f32's SIMD match replaced by a direct scalar call (path reporting intact)vec_ops::tests::f16_dot_uses_the_vectorized_path — f16_dot_path() reports Neon but the SIMD dot never ran (0 -> 0)
the SIMD branch removed from decode_f16_row onlythe same test — …but the SIMD row decode never ran (1 -> 1)

The last two are why the vectorization gate does not read f16_dot_path() alone: the SIMD entry points (dot and row decode) bump a test-only thread-local counter, so a dispatch that reports the SIMD path while running the scalar dot still fails. Reading the path function by itself was the first version of this gate, and the mutation above passed it — a textbook "assertion on something the code under test clears".

Two findings worth recording.

  1. The gate had to load under a namespace, and the reason is a real hazard. The CUDA weight registry is process-global and name-keyed, and register_weight reuses a same-name+same-size device copy. The other real-model gates in the #[ignore]d set load the cached q4_k_m 0.5B under the default ns="", which shares 121 f32 norm/bias names and their byte sizes with the converted f16 file — but not their values. The f16 gate therefore failed only in the full serial set (device greedy [3110, 31139, 47, 34369] against the CPU's [12095, 13, 1084, 374]) and passed when run alone: the device arm was computing with the other file's norms. Loading both arms under ns="f141:" (the loader's own documented remedy for a second model's name-keyed entries) fixes it. This is the process-global hazard #64 describes, and the shape of the failure — a gate that passes alone and fails in the set — is worth remembering.
  2. The cached qwen2.5-0.5b-instruct-q4_k_m.gguf is not weight-identical to the Qwen/Qwen2.5-0.5B-Instruct HF checkpoint. Its f32 norms differ (blk.0.attn_norm.weight[0] = -0.046875 in the HF-derived f16 file — the exact bf16 value in the safetensors — against -0.082947 in the cached file), which is what turned finding 1 into a wrong-weights comparison rather than a last-bit one. No gate compares across the two files, so nothing is red; a future gate that does must not assume they are the same weights.

Honest scope. (a) Metal had no f16 weight kernels at the time of this record, so f16 was refused there and fell to the CPU loudly (#164) — that was a stated policy, not an untested implementation. (#164 landed on a Mac 2026-10-06: the kernels exist and f16 runs on Metal; see the G8 row.) (b) The device gate's tolerance is a backend tolerance: both paths compute f32 activations against f16 weights (an f16 weight has no integer form, so the CPU does not quantize its activations the way it does for the quantized types), so what remains is accumulation order plus the attention exp/softmax kernel — measured 4.0e-6 relative, bounded at 1e-3 with ~250x headroom. (c) The CPU f16 path is vectorized and pooled but not blocked over the K dimension, and the row blocking that is there measured neutral on this model; a K-tiled kernel that reuses a weight row across a token tile without re-reading it is not implemented. (d) x86_64 codegen is verified by cross-cargo, not by running the AVX2 kernel — CI's ubuntu job is the run. (e) The row-blocked and direct forms are asserted bit-identical on dgxspark; the argument that they must be (identical FMA tree and order) is also why the SIMD/scalar comparison uses a tolerance rather than bit equality. (f) The pre-existing finding filed with F6b is fixed in F6c (#165): the q4_K dsc plane was built for non-q4_K types with passing geometry.

F6c — the q4_K dsc plane is gated on q4_K, with a payload contract (#165) — DONE 2026-09-25

What landed. #165, found while porting #141 (F6b) and deliberately kept out of that branch (which is why f16 was not folded into the loader's quantized matches!).

  • The qwen2 loader built the r59 W_dsc f32-pair plane inside the else of if ttype == TensorType::Q6_K, so the registration gate was reached for every non-Q6_K type the enclosing matches! admitted (Q4_0/Q4_1/Q4_K/Q5_0/Q5_1/Q5_K/Q8_0), with no type check on the plane. Any of them whose geometry passed (id % 256 == 0, od even) had 144-byte q4_K super-blocks decoded out of its bytes and uploaded — a plane no kernel reads (the q4k_dsc map is keyed on the q4_K weight's device pointer and its only consumer is mmq_raw_nb_bt), at device-memory and host-CPU cost per tensor.
  • The fix is one pure rule with two load-bearing halves, src/q4k_dsc.rs::q4k_dsc_plane_admitted: ttype == TensorType::Q4_K and raw.len() == od * (id / 256) * 144 — q4_K's own block layout, an equality and not a lower bound. The qwen2 loader calls it; register_weight_q4k_dsc re-checks the payload before the budget query and before the host expansion, and expand_q4k_dsc itself returns None for a payload it cannot index, so a direct caller cannot bypass either. The rule is a free function in a non-cuda-gated module precisely so CI's CPU job runs its tests; the CUDA job only compile-checks the device modules.
  • Why both checks. q4_0's bytes/element equals q4_K's exactly (18/32 == 144/256), so the size check is blind to the difference — the type gate is the only thing that can refuse it. A q8_0 payload (34/32) is longer than the row arithmetic needs and was misread; the size check refuses it. A future type with a smaller ratio (a 2-bit K-quant: 84/256) is shorter and is refused instead of read past the tensor — the latent OOB #165 names. What the size check cannot do is tell a q4_K payload from another type's bytes of the same length; that is the type gate's job.

Measured acceptance (GB10, sm_121, CUDA 13.0, driver 580.178.04). The "before" numbers are the real pre-fix code path (the two gate halves reverted, everything else identical).

CriterionBeforeAfter
q4dsc_planes() after loading the cached 0.5B q4_0 (qwen2 arch)24 planes / 26 148 864 B0 / 0
q4dsc_planes() after loading /tmp/fix165/qwen2.5-0.5b-instruct-q8_0.gguf (a qwen2 q8_0 built by minfer quantize from the cached 0.5B q4_0)24 / 26 148 864 B0 / 0
q4dsc_planes() after loading Qwen3-0.6B-Q8_0.gguf — the model #165 names0 / 0: the qwen3 loader never had the call (honest scope below)0 / 0
the cached 0.5B q4_k_m (the q4_K positive control, end to end)—12 planes / 13 074 432 B, exactly the GGUF index's admissible q4_K set
cargo test --release --features cuda -- --test-threads=1—516 passed / 0 failed / 34 ignored (baseline 513 / 0 / 33)
FEATURES=cuda scripts/real_model_gates.sh (0.5B and Qwen3-0.6B-Q8_0 configs)—34 passed / 0 failed both (baseline 33 / 0)
compute-sanitizer --tool memcheck over the serial CUDA unit suite0 API errors over 514 passed / 2 failed / 34 ignored (the two new gates failing on the pre-fix path; 344.62 s)0 errors over 516 passed / 0 failed / 34 ignored (349.38 s)
cargo test --release (CPU)—447 passed / 0 failed / 30 ignored unit (baseline 445 / 0 / 30; +2 pure q4k_dsc tests) + 10 / 0 / 6 integration
PARALLEL=0 scripts/real_model_gates.sh (CPU)—30 passed / 0 failed, unchanged (the new #[ignore]d gate is cuda-gated)
rustup run stable rustfmt --edition 2021 --check on the changed .rs—clean (rustfmt 1.9.0-stable; the pinned 1.97.1 toolchain has no rustfmt component, CI runs no fmt job)
python3 scripts/check_docs_links.py—940 links resolve in 184 files, unchanged

Gates. Two pure tests in src/q4k_dsc.rs (CI's CPU job): the q8_0-length and wrong-type refusals with a q4_K positive control and a Q5_K wrong-type control, and the short/long/empty payload refusals. graph::cuda_backend::tests::cuda_q4dsc_plane_is_q4k_only (device): the q4_K control registers {name}__q4dsc{od}x{id} and the same q4dsc_planes() query the refusals use sees exactly that plane, while the q8_0-length and one-block-short payloads add nothing — a query blind to planes could not see the control either. The #[ignore]d cuda_real_model_registers_q4dsc_planes_only_for_q4k loads a real model and asserts the registered plane set equals the GGUF index's admissible q4_K set (0 for q4_0/q8_0, 12 for q4_k_m).

Mutations (reverted; every file restored byte-identical, sha256sum). (a) type gate removed (q4k_dsc_plane_admitted drops ttype == Q4_K): the pure wrong-type assertion fails ("the type gate must refuse a q8_0 weight regardless of its length") and the real-model gate fails on the 0.5B q4_0 at 24 planes / 26 148 864 B against 0 expected — the mutation reproduces #165 exactly. (b) size validation weakened (exact equality → payload_bytes >= want): the pure "one block long" assertion fails and the device gate fails at "a q8_0 payload must not register a __q4dsc plane" — so the gate really tests exactness, not a lower bound. (c) wrong plane name (__q4dsc → __q4dscX): the device gate fails at its positive control ("the q4_K payload must register f165q4k1024x3072__q4dsc1024x3072"), which is what proves the "nothing registered" arms observe the plane's real registry entry rather than passing vacuously.

Honest scope. (a) The model #165 names does not reproduce the defect: Qwen3-0.6B-Q8_0 is arch qwen3 and goes to src/models/qwen3/loader.rs, which has no register_weight_q4k_dsc call at all — its plane count is 0 before and after (measured). The ticket's "~22 MB on Qwen3-0.6B-Q8_0" is the qwen2 loader's arithmetic applied to the qwen3 model's ffn_down shape; the defect is qwen2-loader-only. The qwen2 q8_0 arm is measured on a q8_0 file built here with minfer quantize (from the cached q4_0 0.5B, since a K-quant source is not re-quantizable) and both qwen2 arms are named in the table. (b) The size check cannot tell a q4_K payload from another type's bytes of the same length — q4_0 shares q4_K's ratio exactly, which is why the type gate is not redundant; the pair is the contract, and only the pair is tested. (c) The #[ignore]d real-model gate assumes the default full offload plan (all blocks fit), which holds for every cached small model it runs on; a partial plan would register fewer planes than the GGUF index implies. (d) A separate loader divergence was filed here, and is fixed in F6d (#167): at the time the qwen3 loader lacked both the q4_K dsc plane this record gates and the f16 registration branch #141 gave qwen2, so a q4_K Qwen3 ran the in-kernel scalar dsc decode and an f16 Qwen3 model dropped to the CPU on CUDA (read from the loader then; device-verified in F6d, 2026-09-25).

F6d — the qwen3 loader shares qwen2's registration rule: f16 + the q4_K dsc plane (#167) — DONE 2026-09-25

What landed. #167, the loader divergence F6c filed ((d) above). Each loader carried its own copy of the #[cfg(feature = "cuda")] per-tensor registration block, and the copies had drifted twice:

  • the qwen3 copy had no register_weight_q4k_dsc call (r59), so a q4_K Qwen3 weight kept mmq_raw_nb_bt's in-kernel scalar dsc decode — correct, but the r59 prefill win was not available to Qwen3 even though the same kernel consumes both models;
  • it had no TensorType::F16 branch (#141), so an f16 Qwen3 weight was never registered and the all-or-nothing weights_on_cuda gate dropped the whole model to the CPU even on a build with the f16 kernels. Registering would not have been enough on its own: the qwen3 graph's own type gate (Qwen3Graph::weights_on_cuda → matmul_t_ok/embed_t_ok) was missing F16 as well.

The shared rule. src/models/weight_reg.rs now owns the dispatch. cuda_weight_reg is the pure decision — no CudaState, no environment (the r59 dispatch gates are passed in), so CI's CPU job runs its tests exactly like src/q4k_dsc.rs — and the cuda-gated register_cuda_weight carries it out. Both models/qwen2/loader.rs and models/qwen3/loader.rs call it for every tensor the E5 plan puts on the device, so the type coverage has one authority and cannot diverge by edit. It carries: the quantized set, the F16 raw branch, F32 (1-D norms/biases vs 2-D matmul weights), the Q6_K padded repack, the q8_0 p32 split plane, the q4_K W_dsc plane under q4k_dsc_plane_admitted, and the clear_mmq_nb_bt_only rule. The Metal per-tensor blocks were deliberately not folded in: Metal's admitted set is different (and, at that time, an f16 weight had to stay refused there — #164), so the extraction is CUDA-only and Metal's per-loader blocks are untouched (CPU/Metal behaviour unchanged was part of the acceptance). (#164 later added F16 to Metal's own per-loader arm, 2026-10-06; the CUDA extraction is still CUDA-only.) Qwen3Graph::weights_on_cuda gained TensorType::F16 in matmul_t_ok and embed_t_ok, matching qwen2's list; CudaState::q4dsc_plane_for was added so a gate can read the same pointer-keyed map mmq_raw_nb_bt reads.

Audit finding folded in. The extracted F32 arm used to test tensor.shape.len() == 2, but Tensor::shape is a [i64; 4], so the test was always false in both loaders and the NB-BT-only flag was never cleared for a 2-D f32 matmul weight. The shared rule takes the real rank (gguf_write::ggml_n_dims) instead. It can only disable the mode-2 skip-write MMQ optimization (never change a result), and no model in the gate set has a 2-D f32 matmul weight — pinned by only_a_2d_f32_weight_clears_the_nb_bt_flag.

Test models.

ModelProvenanceSize
/tmp/f167-work/qwen3-q4k.ggufHuggingFace unsloth/Qwen3-0.6B-GGUF Qwen3-0.6B-Q4_K_M.gguf (the official Qwen/Qwen3-0.6B-GGUF repo ships only Q8_0); a real Q4_K_M mix — 168 q4_K + 29 other quantized 2-D tensors396 705 472 B
/tmp/f167-work/qwen3-f16.ggufllama-quantize --allow-requantize <cached Qwen3-0.6B-Q8_0.gguf> … F16 — 2-D f16, 1-D f321 198 182 048 B
/tmp/f167-work/qwen3-q4_0.ggufminfer quantize --type q4_0 of the cached Q8_0 — the ratio-equal negative control (q4_0's bytes/element is exactly q4_K's)419 245 728 B
/tmp/f167-work/qwen3-f16-minfer-quantize.ggufminfer quantize --type f16 of the cached Q8_0 — not usable, see below1 198 050 976 B

minfer quantize --type f16 was tried first, as the ticket suggested, and does not work: it is a pure element cast, so it writes the 1-D norms as f16 too, and the engine's f16 contract requires 1-D f32 (the CPU RMSNorm reads Tensor::data_f32, which asserts F32). Loading that file panics at src/tensor.rs:277 — data_f32 on 'blk.0.attn_norm.weight' of type F16. The new gate asserts 1-D-stays-f32 (which is what caught it) and the f16 model was then produced with llama-quantize, which follows llama.cpp's "except 1d tensors" rule. The tooling gap is filed as #169.

Measured acceptance (GB10, sm_121, CUDA 13.0, driver 580.178.04).

CriterionResult
f167_qwen3_q4k_registers_the_dsc_plane_exactly — planes vs the GGUF index168 planes / 95 420 416 B, exactly the index's admissible set (168 q4_K, 29 other quantized 2-D); the kernel's own pointer-keyed lookup finds every one; the q8_0 and q4_0 negative arms register 0
f167_f16_qwen3_weights_run_on_the_cuda_devicedevice() == Cuda; 28/28 blocks + embed/output on the device (1137.0 MiB); 197 f16 matmul + 1 f16 embed node assigned Backend::CUDA; device-vs-CPU max |Δlogit| 8.92e-3 (mean 1.25e-3) / 4.46e-4 relative at max |logit| 19.99; greedy [12095, 13, 576, 6722] identical, asserted at ≤ 0.05 and ≤ 1e-3
cargo test --release --features cuda -- --test-threads=1521 passed / 0 failed / 36 ignored (baseline 516 / 0 / 34)
FEATURES=cuda scripts/real_model_gates.sh (0.5B and Qwen3-0.6B-Q8_0 configs)36 passed / 0 failed both (baseline 34 / 0)
compute-sanitizer --tool memcheck over the serial CUDA unit suite0 errors over 521 passed / 0 failed / 36 ignored (353.44 s)
cargo test --release (CPU)452 passed / 0 failed / 32 ignored unit (baseline 447 / 0 / 30; +5 pure weight_reg tests, +2 cuda-gated ignored) + 10 / 0 / 6 integration, unchanged
PARALLEL=0 scripts/real_model_gates.sh (CPU)32 passed / 0 failed (baseline 30 / 0; the two new #[ignore]d gates no-op and pass on a CPU build)
rustup run stable rustfmt --edition 2021 --check on the changed .rsclean (rustfmt 1.9.0-stable; the pinned toolchain has no rustfmt component, CI runs no fmt job)
python3 scripts/check_docs_links.py940 relative links in 184 files, unchanged

The device-vs-CPU spread is ~100× the 0.5B f16 gate's (7.34e-5 / 4.0e-6) because Qwen3 runs four norms per layer (attn_norm + per-head q_norm/k_norm + ffn_norm) and the CPU rms_norm (8-lane AVX2 FMA plus an f64 tail, then 1/sqrt) and the device rms_norm (warp-shuffle f32, then rsqrtf) differ in reduction order and reciprocal-sqrt form; the greedy continuation and the argmax agree, and a wrong f16 row would move the logits by O(1).

Gates. Five pure tests in src/models/weight_reg.rs (CI's CPU job runs them): the F16 arm; the q4_K dsc decision with each of its four gates refused independently (r59 gates off, odd od, the q4_0 type control, a longer q8_0 payload) plus a q4_K positive control; the quantized arm's pre-#167 dispatch (q6_K padded, q8_0 p32, every other quant raw + flag-clear, q4_K never clearing); the 1-D vs 2-D f32 rank rule; and a three-way variant discriminator. Two #[ignore]d real-model gates: the q4_K plane set (exact names/count against the GGUF index, the kernel's pointer map per weight, distinct non-null buffers, and the two negatives) and the f16 device path (placement asserted per node, then the numeric comparison). The plane gate calls CudaState::init() itself, so running it alone does not silently skip for lack of a device — a skip-shaped pass was found and removed during the mutation campaign.

Mutations (reverted; files restored byte-identical, sha256sum). (a) the F16 arm removed from cuda_weight_reg → the pure f16 test fails and the device gate fails at left: Cpu, right: Cuda; (b) F16 removed from Qwen3Graph::weights_on_cuda's matmul_t_ok → the same device-gate failure, so the graph type gate is separately load-bearing; (c) the q4_K admission bypassed (q4k_dsc = gates && od % 2 == 0) → the pure test fails and the real-model gate fails on its q4_0 negative arm (planes registered for a ratio-equal type); the q8_0 arm alone could not catch this, because the registry's own payload re-check refuses a longer payload — which is exactly why the q4_0 arm was added; (d) the plane forced off → the pure test fails and the real-model gate fails with 0 registered against 168 expected. Each mutation was re-run after the rustfmt pass and reverted with sha256sum -c OK.

Honest scope. (a) The prefill win the dsc plane buys is not measured here — the gate proves the plane exists, contains the right weights, and that the kernel's pointer-keyed lookup finds each; the r59 dsc prefill A/B on Qwen3 is not run. (b) The q4_K test model is a community quant (unsloth) because the official Qwen Qwen3-0.6B GGUF repo has only Q8_0; it is a real Q4_K_M mix, not a uniform q4_K, which is stronger for the type gate but means the per-tensor set is that file's. (c) The f16 test model is llama-quantize's F16 of the cached Q8_0, so its weights are a q8_0 re-quantization, not an HF bf16/f16 checkpoint; it preserves the 2-D-f16/1-D-f32 shape contract and is the same class of file #141 gated. (d) The f16 device gate's absolute bound (0.05) is looser than f141's (0.01) for the stated Qwen3 reason; the relative bound stays 1e-3. (e) An f16 Qwen3 file with f16 1-D norms is not loadable (CPU panic; on CUDA norm_weight has no type gate, so the kernel would read f16 bytes as f32) — not device-verified because the CPU panic is reached first; that tooling gap is #169. (f) minfer quantize --type f16 was left unchanged: fixing it is #169, out of this loader-focused ticket.

F6e — quantize --type f16 keeps 1-D tensors f32, and the CUDA norm weight type gate (#169) — DONE 2026-09-26

What landed. #169, the tooling gap F6d filed when its new "1-D stays f32" assertion rejected the model minfer quantize --type f16 had just written. QuantizePlan::plan treated the f16 target as a pure element cast (let keep = is_quant && (ne[1] <= 1 || !row_ok)), so every tensor — the 1-D norms/biases included — was converted to f16. The engine's f16 weight path requires 1-D f32: the CPU RMSNorm reads the weight through Tensor::data_f32, which asserts F32, and neither mat_mul_f16 nor the f16 embedding decode has an f16-norm sibling, so the tool could not run a file it had itself produced.

The predicate is now one arm per target, matching llama.cpp's tensor_allows_quantization (which returns tensor->type for ggml_n_dims < 2): a quant target keeps its 1-D or unaligned-row rule, f16 keeps 1-D (both minfer convert --outtype f16 and llama-quantize … F16 write those tensors f32, and that is the contract the engine reads), and f32 converts everything. The kept tensors are reported through the same preserved list; the CLI message now names the target's actual reason (1-D for f16, 1-D or row length not a multiple of N for a quant target) instead of printing a block-size clause that is meaningless at f16's block size 1.

The latent CUDA hazard the issue recorded is fixed and gated too. CudaBackend::norm_weight used to check only that the name was registered; it now compares the registered byte length (CudaState::weight_size, the same raw-length convention as has_weight_of_size) against the d*4 bytes the rms_norm kernel indexes and returns Err naming both lengths. Without it an f16 norm would be launched into a d*4-byte read out of a d*2 buffer — the "kernel-invariant violation is a refusal, never a silent wrong path" rule of docs/GPU_SAFETY.md, and the mutation below shows the read is a real invalid access, not a theoretical one.

Gates. (1) tooling::tests::quantize_f16_keeps_1d_f32_and_encodes_2d_f16 — CI-covered on a miniature f16-shaped source written by the real writer. Three arms (f16, f32, q8_0) assert the type of every tensor as the written file declares it, against a want computed from the source spec's rank (never from plan), and that a preserved 1-D tensor carries the source's own non-zero bytes; the f32 control differs in the property under test (its 2-D tensors must come back f32), so neither arm can carry the other. The preserved list is asserted against the source's 1-D set. (2) the ignored f6_quantize_end_to_end_stays_within_the_stated_bound gained an f16 arm on the real f16 file: 1-D f32 / 2-D f16 read back, the preserved count, and a bitwise f16→f16 logit equality (every source value is representable, so assert_eq! is the honest claim). (3) graph::cuda_backend::tests::cuda_norm_weight_size_is_part_of_the_invariant — the device gate: the same 64-element norm graph and d, two names both registered and both valid float4 dims, so the f32 arm (control) can only pass and the 2-byte-per-element arm can only fail through the length check; the refusal's text must name both lengths and "f16-norm".

Measured acceptance (GB10, sm_121, CUDA 13.0, driver 580.178.04, 2026-09-26). The source is the cached Qwen3-0.6B-Q8_0.gguf (639 446 688 B); the reference is llama-quantize --allow-requantize <src> /tmp/fix169/llama-f16.gguf F16.

CriterionResult
Before minfer quantize --type f16 (the bug)1 198 050 976 B; 113 1-D F16 (output_norm.weight + blk.*.{attn,ffn}_norm.weight + blk.*.attn_{q,k}_norm.weight) and 197 2-D F16; loading panics at src/tensor.rs:277 — data_f32 on 'blk.0.attn_norm.weight' of type F16
After minfer quantize --type f161 198 182 048 B; 113 1-D F32 + 197 2-D F16; sha256 6341ef3a7287cad1e42b5910cb3db4ee587ae2eff76f48f953fef63c00e26667; byte-identical (cmp) to the llama-quantize … F16 reference, same size
per-tensor type set vs llama-quantize … F16310 / 310 agree on (name → type), and every ne shape too
minfer quantize --type f32310 / 310 F32 (2 390 150 816 B), no preserved list
the f16 file on CPU (--backend cpu, greedy, 13-token prompt)runs; prefill 13 tokens in 0.10 s, 16 generated tokens in 1.11 s; text <think>\nOkay, the user asked for the capital of France. Let me think
the f16 file on CUDA (default backend)all 28 blocks + embed/output on cuda (1137.0 MiB of device weights); same 16-token greedy text; prefill in 0.09 s, 16 tokens in 0.22 s
CPU-vs-CUDA logits (MINFER_GRAPH_DUMP)argmax equal at prefill / decode_13 / decode_14; max |Δlogit| 0.0178 / 0.0169 / 0.0132 (6.0e-4 / 4.9e-4 / 4.8e-4 relative)
CPU cargo test --release457 passed / 0 failed / 33 ignored unit (baseline 456 / 0 / 33; +1 gate) + 10 / 0 / 6 integration
PARALLEL=0 scripts/real_model_gates.sh and the default parallel form33 / 0 each
cargo test --release --features cuda -- --test-threads=1526 passed / 0 failed / 37 ignored (baseline 524 / 0 / 37; +1 CI gate +1 device gate)
FEATURES=cuda scripts/real_model_gates.sh37 / 0 at the 0.5B config and 37 / 0 at the Qwen3-0.6B config
compute-sanitizer --tool memcheck over the CUDA unit suite0 errors over 526 / 0 / 37
rustup run stable rustfmt --edition 2021 --check on the changed .rsclean (rustfmt 1.9.0-stable; the pinned toolchain has no rustfmt component and CI runs no fmt job)
python3 scripts/check_docs_links.py957 relative links in 185 files

Mutations (reverted; sha256sum -c byte-identical). (a) QT::F16 => false in plan (the pre-fix behaviour, 1-D converted again) → quantize_f16_keeps_1d_f32_and_encodes_2d_f16 fails at the value assertion F16: tensor blk.0.attn_norm.weight came back f16, expected f32 (rank 1-D). (b) the CUDA length check forced off (if false && got != Some(want)) → the device gate fails (an f16 norm weight must be refused before the launch: (), i.e. the f16 arm executed), and under compute-sanitizer that run reports Invalid __global__ read of size 16 bytes ×12 / ERROR SUMMARY: 12 errors — the d*4 read out of a d*2 buffer, device-verified.

Honest scope. (a) The byte-identity with llama-quantize is a stronger result than the ticket asked for (per-tensor types) but it holds only for this source/target pair: an f16 source whose 1-D tensors were already f16 would be kept as f16 by both tools (llama.cpp returns tensor->type; neither converts a 1-D tensor to f32), so the "1-D is f32" contract is a property of the files the producers write, not of an arbitrary GGUF. (b) The miniature source in the CI gate is synthetic (zero-ish payloads, one architecture key) — the shape rule is exercised, not a real architecture's tensor mix; the real-file half is the #[ignore]d arm and the acceptance run above. (c) output.weight is absent from this tied Qwen3 model, so the tied-embedding policy is not exercised by the acceptance run (it is covered by the existing q4_0 byte-parity tests and unchanged here). (d) The CUDA norm gate is device-only (CI has no GPU); its CI-covered half is the weight_size arithmetic exercised through the pure path, and the device arm is run here on GB10. (e) The norm_weight change adds one size lookup per norm node at execute time; no timing was measured and none is claimed.

F6f — bf16 weights on the CUDA device, CUDA half (#208, 2026-10-06)

What landed. The device half of the second 2 B/element dtype, mirroring #141/#164 exactly: bf16_f32_matmul_vec / bf16_f32_matmul_scalar / embed_rows_bf16 in src/cuda/kernels/ops_misc.cu (plus the b2f helper in common.cuh), selected by the new TensorType::BF16 arms of CudaState::matmul_f32_ptr_layout / embed_rows_on_gpu; the raw registration arm in models::weight_reg::cuda_weight_reg; and the type gate in both Qwen2Graph::weights_on_cuda / Qwen3Graph::weights_on_cuda lists. The design decision — its own kernel, no flag on the f16 one — is argued in docs/CUDA-BACKEND-DESIGN.md §4.4: bf16 is f32's top 16 bits, so the promotion is a bits << 16 shift the vec kernel performs on one uint4 per 8 elements, materially cheaper than __half22float2, and a flag would put a per-element branch in the hottest device kernel. Registering raw keeps the 2 B/element claim: the 0.5B bf16 file registers 942.4 MiB of device weights, the same number its f16 twin reports (a dequantized copy would be ~1.9 GiB). bf16 never enters the int8 MMQ GEMM (not a quantized format) and it does not fuse: cuda::concat_rows has no 2 B/element arm, so blk.{i}.attn_qkv / blk.{i}.ffn_gu are not registered and the census is the unfused 169 matmul + 1 embed shape. Metal's arm is deliberately untouched (matches!(ttype, F32 | F16)), so a bf16 GGUF on a Metal build still drops to the CPU loudly; that half is the ticket's second PR.

Verification (rule 5 numbers). dgxspark (aarch64, GB10 sm_121), 2026-10-06, in the worktree:

CommandResult
cargo build --release --features cudaexit 0, warning-free (deny(warnings) off-test)
cargo build --releaseexit 0, warning-free
scripts/cuda_test.sh570 / 0 / 45 (was 567 / 0 / 42)
cargo test --release482 / 0 / 39 unit + 10 / 0 / 6 integration (was 481 / 0 / 36); the x86_64 CI runner reports 480 / 0 / 39 — the live check_status.py --check-live gate is what corrected the CI row from a derived 41 ignored to the measured 39
FEATURES=cuda scripts/real_model_gates.sh (0.5B)45 / 0 (was 42 / 0)
FEATURES=cuda MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf scripts/real_model_gates.sh45 / 0 (was 42 / 0)
PARALLEL=0 scripts/real_model_gates.sh (CPU)39 / 0 (was 36 / 0)
python3 scripts/check_{source_layout,dead_code_annotations,dead_code_oracle --config cpu,dead_code_oracle --config cuda,doc_line_anchors,docs_links,status --check}.pyexit 0 each (check_doc_line_anchors still prints its two pre-existing unused freezes)
cargo fmt --all --checkclean

The fixture is minfer convert --outtype bf16 ~/.cache/minfer/f6-src/hf/Qwen2.5-0.5B-Instruct /tmp/f6-work/f208/f208-bf16.gguf (994 156 352 bytes; minfer quantize has no bf16 encoder), and the gate takes it from $MINFER_F208_BF16_GGUF or converts the cached checkpoint itself.

The bar, named before measuring. Both arms read the same bf16 weights, so weight precision is not part of the device-vs-CPU difference at all — it is accumulation order plus the attention exp/softmax kernel (graph rule §9), the same class as #141's f16 measurement (7.34e-5 / 4.0e-6) and #142's CPU bf16-vs-f16 record (2.29e-5 / 1.24e-6). Stated bound: max |Δlogit| ≤ 0.01 and max relative ≤ 1e-3, greedy continuation identical. Measured: 7.82e-5 absolute / 4.24e-6 relative, greedy [12095, 13, 1084, 374] on both arms (the same continuation #142 recorded for the CPU).

Mutation evidence (rule 3), each reverted.

MutationGateResult
cuda_weight_reg's BF16 arm disabled (if false && ttype == TensorType::BF16)f208_bf16_weights_run_on_the_cuda_deviceFAILED at the placement assertion: left: Cpu, right: Cuda, with the loader's own CUDA GATE: weight 'f208:token_embd.weight' (type BF16) has no CUDA kernel or is not registered + "running on CPU" lines — the all-or-nothing gate really is what routes the model
the BF16 arm of matmul_f32_ptr_layout calls launch_f16_f32_matmul insteadcuda_bf16_matmul_matches_the_exact_shift_referenceFAILED on the first element: bf16 matmul [0,0] (od=10 id=64 nt=3): got 1.875 want 1 (an f16 misread of the same words)
both—reverted; sha256sum -c of the two files matches the pre-mutation hashes, and the gates are green again

Honest scope. (a) The device gates are #[ignore]d and device-only — CI has no GPU, so its only half is the compile check in build-linux-cuda plus the pure cuda_weight_reg unit test. (b) The Metal half is not in this PR: no Metal registration arm, no Metal kernel, no Metal claim — a bf16 GGUF on a Metal build still drops to the CPU loudly, by construction. (c) The real-model gate is the 0.5B (Qwen2) only; Qwen3-bf16 would need a second converted checkpoint (the weights_on_cuda arm is added to both architectures and is covered by the pure unit test, but no device run exercises a bf16 Qwen3 file). (d) bf16 is unfused and non-MMQ by construction, so this PR measures the f32-activation path only; no prefill-GEMM or fused-node bf16 census is claimed.

F6g — bf16 weights on the Metal device, Metal half (#208, 2026-10-06, PR #323)

What landed. The Metal half of the second 2 B/element dtype — the #321 CUDA half's twin and the #164 f16 pair's sibling. kernel_bf16_f32_matmul + kernel_get_rows_bf16 (src/metal/kernels/bf16.metal, listed in build.rs's SHADER_SOURCES) are selected by the new TensorType::BF16 arms of quant_matmul_f32_on_gpu_buf / embed_tokens_gpu through the pl_bf16_f32 / pl_get_rows_bf16 pipelines built in try_new, and both loaders' Metal arm now admits BF16 (matches!(ttype, F32 | F16 | BF16)), which is what makes Qwen2Graph::weights_on_gpu pass and the model a Metal model. The design decision — its own kernel, no dtype flag on the f16 one — mirrors the CUDA half: bf16 is f32's top 16 bits, so the promotion is an in-register as_type<float>(bits << 16) (exact, no rounding), and a shared kernel would branch per element in the hottest device kernel. Registering raw keeps the 2 B/element claim: the 0.5B bf16 file registers 942.4 MiB of device weights, the same number the CUDA and f16 twins report. bf16 does not fuse (the fused device forms are gated on the CUDA path), so the census is the unfused 169 matmul + 1 embed shape. The one documented limit is that a registered-but-unkerneled dtype would hit the dispatch ladder's _ fallback (the #317 trap), which is exactly why the registration arm and the kernel land together and the exactness gate is the proof. No Qwen3 real-model arm exists because minfer convert supports Qwen2ForCausalLM only, so there is no producer for a bf16 Qwen3 file; the Qwen3 loader arm is the same one-line registration change and the Metal device gate is the Qwen2 file.

Verification (rule 5 numbers). macbook (macOS 27.0.1, Apple M4 Pro), 2026-10-06, in the nested worktree:

CommandResult
cargo build --releaseexit 0, warning-free (deny(warnings) off-test)
cargo test --release531 / 0 / 43 unit + 21 / 0 / 6 integration (was 529 / 0 / 42 unit; the +2 passed are the kernel-exactness gates, the +1 ignored is the real-model gate)
PARALLEL=0 scripts/real_model_gates.sh42 / 1 (the +1 pass over the pre-change set is f208_bf16_weights_run_on_the_metal_device; the single residual is the #310 kv_sharing gate, not this change)
f208_bf16_weights_run_on_the_metal_device1 / 0 — 169 bf16 matmul + 1 embed nodes on Backend::METAL, 942.4 MiB, max
bf16_matmul_matches_the_exact_shift_reference / bf16_embed_gather_matches_the_reference2 / 0, bitwise against crate::block::bf16_to_f32
cargo fmt --all --check + the six check_*.py gatesclean / exit 0 each

The fixture is minfer convert --outtype bf16 ~/.cache/minfer/f6-src/hf/Qwen2.5-0.5B-Instruct /tmp/f6-work/f208/f208-bf16.gguf (948.1 MiB; minfer quantize has no bf16 encoder), taken from $MINFER_F208_BF16_GGUF or converted by the gate, and the gate skips loudly when neither exists.

The bar, named before measuring. Both arms read the same bf16 weights, so weight precision is not part of the device-vs-CPU difference at all; it is accumulation order plus the attention exp/softmax kernel (graph rule §9). The CPU bf16-vs-f16 record is 2.29e-5 / 1.24e-6, but Metal's scalar row lanes dominate — #164's f16 Metal gate recorded 2.4e-3 / 7.9e-3 against a 0.05 bar. Stated bound: max |Δlogit| ≤ 0.05 and max relative ≤ 5e-3, with the greedy continuation identical (the f16 Metal bar, not loosened). Measured: 1.889e-3 absolute / 1.025e-4 relative (max |logit| 18.43), greedy [12095, 13, 1084, 374] on both arms — the same continuation the CPU (#142) and CUDA (#208) records give.

Mutation evidence (rule 3), each reverted.

MutationGateResult
TensorType::BF16 dropped from the qwen2 Metal registration arm (`F32F16`)f208_bf16_weights_run_on_the_metal_device
the TensorType::BF16 matmul arm calls pl_f16_f32 insteadbf16_matmul_matches_the_exact_shift_referenceFAILED, exit 101: bf16 matmul [0,1] (od=64 id=96 nt=3): left: -4.375, right: 7.0 — the f16 misread of the same words
both—reverted with git checkout --; git status clean and the gates green again

Honest scope. (a) The device gates are #[ignore]d and device-only — CI's build-macos job compile-checks the kernels and the non-ignored metal_pipelines_compile gate compiles every pipeline, but the numeric device claims are the Mac runs above. (b) The real-model gate is the 0.5B (Qwen2) only, because minfer convert cannot produce a bf16 Qwen3 file; the Qwen3 loader arm is the same registration change and has no independent device run. (c) bf16 is unfused and non-MMQ by construction, so this PR measures the f32-activation path only; no bf16 prefill-GEMM or fused-node census is claimed. (d) No timing is claimed — this is a correctness/placement change.

10. Phase G — Metal alignment round (complete 2026-10-06; device claims need a Mac)

Metal is a first-class target — it is the default backend on macOS and a plain cargo build --release builds it — so this phase is scheduled, not deferred. Only the device half waits: CI's build-macos job already compile-checks any Metal change, and per the standing rule a Metal-only change is recorded as compile-verified or left behind a build-time gate when it cannot be run here.

Order and rationale. G1–G3 first: each is small, independent of the KV semantics, and removes a way Metal can be wrong (missing guard, debug_assert! on a release path, a silent weightless-RMSNorm fallback). Then G5 after C7/C8, deliberately: porting the cell store before the arena becomes elastic (C7) and shareable (C8) would mean writing the same semantics into Metal twice. G4/G7 follow G5; G6 landed on a Mac (2026-10-05) — its delta is Metal's DeviceMemory answer, because the allocator's reserve/assign split is backend-agnostic and Metal's pool already ran through it.

IDOriginWorkPosition
G1A3pos < n_ctx guard in Metal's KvcacheStore · #38landed on a Mac (2026-10-06): the allocator already bounds cells/positions on the fill_input_i32 path (docs/ARCHITECTURE-EXECUTION-PLAN.md §A3), so the delta is the arm that indexes the region — MetalBackend::check_kv_store_rows now bounds every Metal KV-write arm (KvcacheStore on cells; FusedQKV/FusedQkvNorm on positions) against the layer region's n_ctx and returns Err naming the cell and the arena before any dispatch — record in docs/METAL-BACKEND-DESIGN.md §4.9; gate metal_kvcache_store_refuses_a_cell_past_the_arena (PR #304)
G2A8debug_assert! → Err for FusedFFN/FusedQKV/FusedQkvNorm nt == 1 · #39landed on a Mac (2026-10-06): each decode-only arm now refuses nt != 1 with an Err naming the node and the observed nt before the weight lookup, so a release build no longer dispatches a shape the kernel does not handle — record in docs/METAL-BACKEND-DESIGN.md §4.9; gates metal_fused_ffn_refuses_nt_other_than_one / ..._qkv_... / ..._qkv_norm_... drive each arm through the real BackendScheduler::execute (PR #307)
G3A8Remove the silent weightless-RMSNorm fallback (metal_backend.rs:403-414, :457-468) · #40landed on a Mac (2026-10-06): Op::RmsNorm and Op::QkNorm now call MetalBackend::norm_weight, which returns Err naming the node and the missing tensor for both None meanings (weight_name absent, or a set name the device never registered) instead of falling through to the weightless kernel — no supported producer builds a weightless norm, so nothing legitimate is lost — record in docs/METAL-BACKEND-DESIGN.md §4.9; gates metal_rms_norm_refuses_a_weight_not_on_gpu / ..._weightless_node / metal_qk_norm_refuses_a_weight_not_on_gpu (PR #311)
G4A8CUDA/Metal op-set asymmetry: decide whether Metal gains QkvBiasRopeStore · #52landed on a Mac (2026-10-06): the decision is keep the refusal — the mixed-quant epilogue port would save 6 dispatches per mixed layer (10 → 4), 84/token on Qwen2.5-7B-Q4_K_M (14/28 layers carry attn_v as Q6_K against Q4_K q/k), with no numerical difference and a sub-1% time ceiling, so the asymmetry is recorded instead of closed; record in docs/SUPPORT-MATRIX.md (PR #318)
G5C1/C2/E1Port the cell store, the KV removal/shift and the explicit attention span to Metal (supports_attn_span() becomes true; today copy_kv_to_cpu has no Metal arm, so a Metal session re-renders instead of shifting, and a multi-sequence batch is refused outright) · #44landed on a Mac (2026-10-06), both halves. (a) read side — kernel_gqa_attn_window_f32/_f16 (src/metal/kernels/attn_window.metal) reads the one-range attn_span window for nt == 1 and nt > 1 at both KV widths, the Op::Attn arm selects the mode by the window input's size (causal / span / a refused map) and supports_attn_span() is true; the causal kernels are byte-untouched (PR #313). (b) write/move side (PR #316) — Backend::copy_cells moves rows one at a time with MTLBlitCommandEncoder in the overlap-safe order (both directions; f16 halves the stride), GraphAllocator::copy_kv_to_cpu reads through the registry host_read hook after flushing the split (so a Metal session shifts/saves instead of re-rendering), MetalBackend holds a per-engine kv_format (the process-wide KV_F16/kv_cache_is_f16/set_kv_cache_type are deleted) and batch_mode(Metal) follows the device. The packed q8_0 read is enabled since #310 (mechanism A native packed decode + mechanism B f32 stage) and a physical shift of an f16 region refuses loudly (#306, CUDA exposed too). Records: docs/METAL-BACKEND-DESIGN.md §4.4.2; gates metal_copy_cells_moves_overlapping_rows_in_both_directions / metal_f16_kv_cell_move_strides_by_row_bytes / metal_kv_format_is_per_engine / metal_copy_kv_to_cpu_reads_after_the_pending_split, the Metal arms of a_compaction_between_steps_keeps_the_continuation / kv_rm_is_exact_and_the_window_shift_is_a_named_tolerance_class / reused_cache_across_prompts_matches_a_fresh_cache and a_long_prefill_keeps_another_slot_decoding, plus metal_attn_span_*. The windowed kernel is a correctness path, not a speed one — a batched multi-sequence prefill runs it and its throughput is unmeasured (#315)
G6E4Adopt the reserve/assign allocator split in Metal's pool · #53landed on a Mac (2026-10-05): the allocator split was already backend-agnostic (E4 S3), so the delta is Metal's DeviceMemory answer (recommendedMaxWorkingSetSize) — record under the E4 record
G7METAL-OBJRe-run the Metal gap/parity measurements after G2–G3 (and again after G5), since each changes a kernel path · #54landed on a Mac (2026-10-06): re-measured the docs/METAL_OPTIMIZATIONS.md §0.1 rows at 6b95763 (macbook (macOS 27.0.1, Apple M4 Pro), minfer bench -p <P> -n 128 -r 3) — 0.5B Q4_0 decode 306.19 / prefill pp440 6249.60 t/s, 7B Q4_K_M decode 48.52 / prefill pp206 406.51, Qwen3-4B decode 74.11 / prefill pp241 729.83, 0.5B CPU decode 149.69 (the pre-round ~5.9 is not reproduced; see the §0.1 note); parity is green at the op level (metal_*_matches_cpu), the external oracle graph_metal_matches_llama_reference reproduces Qwen3's pinned greedy prefix, and support_table_matches_support_matrix_doc was extended with the Input and View-offset rows it had left unchecked (the model-level graph_metal_matches_cpu_logits was found degenerate in the graph era — #324); recorded the round's final macOS baseline (unit 531/0/43, integration 21/0/6, real-model 42/1 ×2 — the one residual is the deliberate #310 kv_map gap) in docs/METAL-BACKEND-DESIGN.md §7.4 (PR #325)
G8F6/#49f16 weight matmul + embedding kernels on Metal (an f16 GGUF runs on the device instead of falling to the CPU) · #164landed on a Mac (2026-10-06): kernel_f16_f32_matmul + kernel_get_rows_f16 (src/metal/kernels/f16.metal, listed in build.rs's SHADER_SOURCES) are the f32-activation matmul and the embedding gather, selected by the TensorType::F16 arms of quant_matmul_f32_on_gpu_buf / embed_tokens_gpu through the new pl_f16_f32 / pl_get_rows_f16 pipelines; both loaders' Metal branch registers raw 2 B/element f16 (`matches!(ttype, F32
G9#317f32-weight matmul on Metal — quant_matmul_f32_on_gpu_buf had no F32 arm, so an f32 weight hit the catch-all _ arm (kernel_q4_0_f32_matmul, which reads the f32 bytes as Q4_0 blocks and writes zeros); the order-dependent graph::op_matrix::matrix_cases_match_their_reference was the only reporter because its Metal column only ran when an earlier test had initialized MpsState · #317landed on a Mac (2026-10-06): kernel_f32_f32_matmul (src/metal/kernels/f32.metal, the pl_f32_f32 pipeline) is dispatched by the new TensorType::F32 arm (CUDA parity: launch_f32_f32_matmul), and op_matrix's Metal arm calls MpsState::init() like the CUDA arm — so the column no longer depends on test order — record in docs/METAL-BACKEND-DESIGN.md §4.4 and docs/SUPPORT-MATRIX.md footnote 2; gates metal_matmul_f32_matches_cpu and the device-present assertion in matrix_cases_match_their_reference (PR #320)
G10#314Chunked-prefill drift on Metal — the flash prefill (kernel_flash_attn_blk_*) partial last KV block read the last C rows (pos0 = nkv - C), overlapping the previous full block whenever nkv % C != 0 && nkv > C; the online softmax double-counted the overlap, so a chunked (or any >64-token) prefill drifted 1.445 (0.5B) / 0.382 (Qwen3-0.6B), judged against a CUDA-only bar · #314landed on a Mac (2026-10-06): the partial block reads its own [ic, ic + C) window (pos0 = ic) and kernel_kv_tail_pad pads from ic = (nkv / 64) * 64, so padded rows are masked instead of leading rows double-counted; a_chunked_prefill_answers_like_an_unchunked_one now carries the named Metal cross-shape class (0.1, measured 0.0087 / 0.0078) instead of treating Metal as CPU — record in docs/METAL-BACKEND-DESIGN.md §4.4.3; gate flash_prefill_matches_the_cpu_reference_at_every_kv_tail (nkv 72/96/98 red before at 0.164, ≤ 2e-4 after) (PR #322)
G11F6/#49bf16 weight matmul + embedding kernels on CUDA and Metal (a bf16 GGUF runs on the device instead of falling to the CPU) · #208Both halves landed. CUDA half on dgxspark (aarch64, GB10 sm_121) (2026-10-06): bf16_f32_matmul_vec / bf16_f32_matmul_scalar / embed_rows_bf16 (src/cuda/kernels/ops_misc.cu, with b2f in common.cuh) are the f32-activation matmul pair and the embedding gather, selected by the TensorType::BF16 arms of matmul_f32_ptr_layout / embed_rows_on_gpu; models::weight_reg::cuda_weight_reg gained the BF16 raw arm and both loaders' CUDA branch admits the type through that one shared rule. bf16 registers 2 B/element with no f32 copy and never enters the MMQ GEMM; cuda::concat_rows has no 2 B/element arm, so bf16 runs the unfused matmul chain. Metal half on macbook (macOS 27.0.1, Apple M4 Pro) (2026-10-06): kernel_bf16_f32_matmul + kernel_get_rows_bf16 (src/metal/kernels/bf16.metal, the pl_bf16_f32 / pl_get_rows_bf16 pipelines) are its f32-activation matmul/embed, selected by the TensorType::BF16 arms of quant_matmul_f32_on_gpu_buf / embed_tokens_gpu, and both loaders' Metal arm is `matches!(ttype, F32

G5 acceptance (on a Mac; the CPU/CUDA equivalents are the gates already in the suite): two sequences do not cross-attend, bitwise; a mid-session compaction is bit-identical; the C2 context shift matches CPU; and MINFER_BATCH unset may then batch on Metal, which is what E6's Metal exclusion is waiting for.

G5 (a)/(b) status (2026-10-06). The read side (a) satisfies the first acceptance clause: batch_order_does_not_change_a_sequences_logits (same shape, bitwise) is green on Metal, and the cross-shape batch gates (a_two_sequence_batch_matches_two_single_sequence_forwards, sequence_count_is_data_not_topology, prefix_reuse_matches_a_full_prefill) are green under the named cross-shape class Metal now carries (cross_shape_tolerance, 0.1; observed max|Δ| ≤ 0.0153), for the same reason CUDA has one — a batched forward runs the windowed kernel and a single-sequence one runs the causal flash/prefill kernel, and the GEMMs tile by nt. The write/move side (b) — compaction, the C2 shift, C5 sessions, packed q8_0, per-engine kv_format — landed in #316, so a_compaction_between_steps_keeps_the_continuation, kv_rm_is_exact_and_the_window_shift_is_a_named_tolerance_class, reused_cache_across_prompts_matches_a_fresh_cache and offset_sensitivity_is_narrowed_to_multi_query_attention are green in the full macOS run (recorded 2026-10-06 at 6b95763, docs/METAL-BACKEND-DESIGN.md §7.4). E6's Metal batching default therefore follows the device (chat::batch_mode answers Batched for Device::Metal). The two deliberately-owed pieces then were the set-valued kv_map read and packed q8_0 KV; both have since landed — the gathers_attn_map capability in #362 and the packed read in #310.

Entry condition: G1–G3 need only a machine that builds Metal (CI's build-macos); G5's device claims need a Mac. Exit condition: SUPPORT-MATRIX.md's per-backend op column matches supports_op on all three backends, with A1's matrix green.

11. Sequencing

Phase A  ├─ A0 ─ A1 ─┬─ A3 ─ A4 ─ A5 ─ A6 ─ A7 ─ A8 ──────────►  (A8 CUDA half)
         └─ A2 ──────┘
Phase B  ├─ B1 ─ B2 ─ B3                          (starts once A0/A1 exist)
Phase C  ├─ C1 ✔ ─ C2 ✔ ─────► C3 ✔ ─ C6 ✔ ─ C7 ✔ ─ C7b ✔ ─ C8a ✔ ─ C8b(S1a ✔ S1b ✔ S2 ✔ S3 ✔ S4 ✔ S5 ✔) ─ C4 ✔ (S1, CPU; S2 = #87) ─ C5 ✔   (C3 needed D1; the CUDA path first, per the 2026-09-20 decision)
Phase D  ├────────── D1 ─ D2 ─ D3 ──────────────►         (D unlocks MoE/MLA)
Phase E  ├──────────────────── E1 ✔ ─ E2 ✔ ─ E3 ✔ ─ E4 ─ E5        (E1b ✔ device-verified; E2 closed: CPU 0.49x, GPU 1.9x)
Phase F  └─ F2 F3 F4 F5 F6 F7 (parallel)        F1 = waits for an x86 host
Phase G  └─────────────────► G1 ─ G2 ─ G3 ─ G5 ─ G4 ─ G6 ─ G7   (after C8; G1–G3 compile-verified in CI; device claims need a Mac)

Critical path: A0 → A1 → C1 → C2 → E1 → E2 → C3 → C6 → C7 → C8 → G5. Deliberate exception to the roadmap's ordering: A1/A2 run before the hazard-removal tickets, because they are the instrument that proves those tickets and everything after them.

Hardware lanes (#335). F1 = waits for an x86 host in the diagram is a machine constraint, not a position in the order: F1 cannot run on either box this project owns. The two lanes that can start today are the macOS round-2 tickets tracked in #333 on macbook (macOS 27.0.1, Apple M4 Pro), and the open Linux-side tickets (cuda / tooling / docs) on dgxspark (aarch64, GB10 sm_121). §14 row 8 carries the same split and §9's F table states it next to F1. The x86_64 (CI runner) does not count as an x86 host for F1: it is GitHub's ephemeral runner, it cannot host an interactive session, and it is a CI job, not a box someone develops on — F1 waits for an x86 machine the maintainer can work on.

12. What "done" means per phase

PhaseDone when
AComplete 2026-09-16. cargo test green on Linux/CPU (aarch64 locally, x86_64 in CI); A1's matrix green (or every red row explained); A0's CUDA verdict recorded; each hazard ticket has a test that fails before and passes after; A6 is closed by measurement instead — a refuted hypothesis with numbers is a result, not a gap.
BComplete 2026-09-16. A multi-turn conversation prefills only the new turns (219 → 16 tokens, ≈11× TTFT); the contamination property is pinned by a bitwise test; numbers recorded interleaved with the same binary.
CCell store lands bitwise; shift is a documented tolerance class; a single request may use the whole arena while the other slots are idle, even with a busy run above it (C7 + C7b ✔), and one cell range can be shared across sequences, counted once (C8); quantized KV behind its gate; session save/restore round-trips.
DA view is provably zero-copy; one hand-written fusion is replaced by a composition, bitwise.
ETwo sequences can be batched without cross-attention; --n-slots 4 beats serial; n_batch chunks prefill; an over-VRAM model runs with layer offload.
GThree backends agree with the op matrix and the support table.

13. Phase A outcome (2026-09-16)

Branch: architecture-phase-a (PR #1). All nine tickets closed.

CheckResult
Unit suite (aarch64, dgxspark)162 passed / 0 failed / 3 ignored
Unit suite (x86_64, CI runner)running 163 → 160 passed / 0 failed / 3 ignored; the 2-test delta is quants::neon_correctness, gated cfg(all(test, target_arch = "aarch64"))
Op matrix17 cases × 3 backends (17 CPU cells, 34 explicit skips), 23 support rows, exhaustive Op enumeration
CIthree jobs green: 1m09s / 1m26s / 4m36s
rustfmt --check, pre-commitclean; the hook ran on every commit
mdbook buildpasses

Defects found and fixed during the phase — none of them were on the ticket list, which is the point of A1 and A2:

  • Op::Softmax returned unnormalised values on the CPU backend (found by A1's matrix; pinned by a case).
  • Four CUDA test call sites broken by A5's signature change — invisible to cargo build because it does not compile #[cfg(test)]; the CUDA CI job now runs cargo test --features cuda --no-run.
  • panic_message(&payload) downcast against the Box (itself Any) instead of the payload, so every panic reported a non-string payload.
  • CI could not bootstrap rustup in the CUDA container (no curl).

Deliberately not done, and where it is recorded:

  • GPU columns of the matrix and the four GPU-only fused ops → a GPU-verifiable run (A0: CUDA is compile-only here).
  • Metal — untouched by decision; its debt is Phase G.
  • A6 — no code change; the measurement is the deliverable.

Open for the maintainer: the three CI jobs carry a pre-existing Node.js 20 is deprecated annotation for actions/checkout@v4; bumping to v5 silences it.

14. Known gaps and open risks (2026-09-18)

Everything the campaign measured is recorded where it happened (the phase records above). This section exists because a risk that lives only in a chat message is not recorded at all: it lists what is known to be unfinished or unverified, in one place, so the next session does not have to rediscover it.

#Gap / riskKindHome / next step
0FIXED 2026-09-19, same day (found in the D3 round): the GPU-batched server corrupted concurrent responses — root cause, fix and evidence at the end of this cell. On GB10, MINFER_BATCH=1 with --n-slots N returns one correct reply and N-1 identical copies of 0.15555555555555 (N = 2, 3, 4 all reproduce; MINFER_BATCH=0 serial gives 4/4 correct). The engine-level gate server_batch_matches_serial_and_is_faster passes with all four replies sane, so the batched graph/engine path is fine and the fault is in the server path (or in a layout it produces). The server's per-row trace shows slot.start = n_ctx_total / n_slots, so slots are reserved as huge runs (e.g. 4096 rows each at --n-ctx 8192 --n-slots 2) and every failing slot is one with a large absolute KV offset — while slot 0 (offset 0) is always the correct one. The engine test uses a 512-row arena, so it never exercises those magnitudes: the server's layout, not the kernel tests, is what lacks coverage. Ruled out so far: prefix reuse misfiring (logs show 40/40 prompt tokens, 0 reused for every slot), admission placing two jobs on one slot (the taken map is per-slot and the traces show distinct seq/slot per row), and D2's new in-place SwiGLU (the corruption reproduces identically with MINFER_FFN_NODE=1, i.e. the hand-written node). Not yet pinned: whether the trigger is the magnitude of the offset, the span/kernel arithmetic at large row bases, or n_ctx itself — the follow-up experiment (vary n_ctx with --n-slots 2) was cut short by a harness wait bug, not by the code. Narrowed the same day (D3 round 2): it is not the layout magnitudes — --n-ctx 128 (slot start 64) reproduces it — and not the engine: server_batch_matches_serial_and_is_faster pointed at the 7B (MINFER_BATCH_TEST_MODEL) gives four sane replies at --n-slots 4, while the server with the same model and two slots gives one sane and one garbage. It is model-dependent: Qwen2.5-7B Q4_K_M reproduces, Qwen2.5-0.5B Q4_K_M does not. CUDA Graph capture and prefill capture are ruled out (MINFER_NO_CUDA_GRAPH=1 and MINFER_NO_PREFILL_CAPTURE=1 both reproduce). The server's and the engine test's remaining structural difference is staggered admission: the server admits requests as they arrive, so its step sequence mixes widths (the trace shows a 1-sequence step, then 2-sequence steps), while the engine test admits all four before its first tick and only ever runs 4-wide steps. That was the leading hypothesis, and it is now ruled out: the ignored engine test grew a stagger arm that reproduces the server's pattern faithfully (admit one request, tick, then admit the rest, so the step sequence mixes widths) and with the 7B at --n-slots 4 every staggered reply is as sane as its simultaneous counterpart (leading bytes shared: 55/41/66/33, same as the simultaneous run). So mixed step widths are not the trigger, and the engine test now pins that pattern either way. What remains between the server and the engine test: the inputs (the server renders the chat template, the test feeds raw encode(prompt); the sampling params come from the request), the engine configuration (n_ctx 2048 vs 512, n_slots 2 vs 4), and the fact that the server keeps running across waves. Root-caused to the windowed-attention kernel (same day, third round). Bisecting the inputs did it, and the answer is none of the candidates above: with the engine test made configurable (MINFER_BATCH_TEST_{SLOTS,CTX,TEMPLATED}) and, crucially, made to render the prompt exactly as the server does (template::render_messages with the GGUF template + generation prompt — model.format_chat is a different path and produced a 13-token prompt where the server's is 34, so the first bisect compared unequal inputs):
Config (7B Q4_K_M, GB10)Result
2 slots / 2048, 5-token raw promptsall four sane (batched and serial)
2 slots / 2048, 34-token templated promptsslot 0 sane; slot 1 garbage in both batched and serial
4 slots / 512, 34-token templated promptsslot 0 sane; slots 1-3 garbage, identically, in both batched and serial

So it is not the server, not batching, not the slot layout or the offset magnitude, and not the model per se — it is the prompt length combined with a non-zero start. Slot 0 (start 0) takes the causal attention instantiation and is always right; every slot with a non-zero start takes the windowed instantiation, and it produces garbage once the window is ~34 rows rather than ~5. That is consistent with everything measured: model-dependent because the two models select different kernel variants (hd 128/n_kv 4 vs hd 64/n_kv 2), unaffected by batching (the kernel is chosen by explicit_span = "start != 0", not by batch width), and invisible to E1b's device tests, which only ever exercised tiny windows (a synthetic hd = 4 fixture, and ~7-token sequences in the order-invariance gate).

RETRACTED (same day, fifth round): the windowed kernel is NOT at fault — the "minimal reproduction" above was a bug in its own harness. run was called twice (causal, then windowed) while it drew q/k/v from a single advancing LCG declared outside the closure, so the two calls compared different random data. The one assertion that kept passing, causal == V(row 0), is a single-key-softmax identity that holds whatever the data are — which is exactly why the failure looked like "the windowed path returns the wrong numbers". Two probes settled it:

  • A device printf in both instantiation entry points (gqa_attn_split_partial and gqa_attn_split_partial_bt) printed, for n = 1, start = 0: causal bound[0]=0 -> row0=0, nkv=1 and windowed bound[0]=0, bound[1]=1 -> row0=0, nkv=1 — identical, so the windowed path derived the right window from the right cells and the divergence had to be in the data itself.
  • With the LCG re-seeded from the shape inside run (never from start or explicit, which are precisely what the two calls differ in), the whole sweep is bitwise equal: 11 cases over (nh, nk, hd) = (1,1,4)/(1,1,128)/(2,2,64)/ (4,4,128), n = 1/2/16/34, start = 0/1/64/256. The test is un-#[ignore]d and is now the E1b/E2 gate it was meant to be (device run 2026-09-19: 11/11 bitwise equal, 249 other tests filtered out).

Consequences recorded here rather than silently dropped:

  • The "root-caused to the windowed-attention kernel" paragraph above is withdrawn, and with it the retraction of E1b's device evidence: cuda_causal_and_windowed_agree_on_the_same_rows was called degenerate; it is weak on its own (one-hot queries make the output score-insensitive) but nothing contradicts it, and its note now says so instead of blaming the kernel.
  • What the fourth round got right and keeps: the engine test did compare unequal inputs to the server's (model.format_chat's 13-token prompt vs the server's 34-token render), and it now renders the prompt exactly as the server does (chat_template_from_gguf + template::render_messages) — that part of the bisect stands and is an improvement to the test.
  • The server blocker itself is therefore unexplained again, and the earlier chain of eliminations must be re-read with that in mind: it is not prefix reuse, admission placement, D2's in-place SwiGLU, layout magnitude, CUDA Graph or prefill capture, staggered admission, model-specific kernel variants, or the windowed instantiation. The reproduction is re-run on the current build (below) to establish whether the symptom still exists at all.

Root cause found and fixed (same day, sixth round): the f16-KV FA prefill kernel's windowed row mask. The server blocker is real and still reproduced on the current build — 4 different prompts, --n-ctx 2048 --n-slots 4, greedy: the batched run returned slot 0 correct and slots 1-3 derailed (slot 3 literally 0.1555555555555555555555, the value in the original report), while MINFER_BATCH=0 returned 4/4 correct. Two facts broke it open:

  1. The dtype gate. cuda::set_kv_cache_type chooses f16 KV whenever n_layers * n_kv_embd >= 8192 — true for the 7B (28 x 512 = 14336), false for the 0.5B (24 x 128 = 3072). That is the recorded "model dependence". The new gate forced cb.kv_f16 = false, so no f16 windowed path was ever tested.
  2. The sweep extended to both dtypes went red immediately at exactly one point: kv_f16=true (4,4,128) n=34 start=64 -> DIVERGES, first=(512, ...). Index 512 is the first element of token 1's output — token 0 was bitwise correct, so the fault was per-row, not per-tile.

nt >= 2 && hd == 128 && !MINFER_NO_FA_PREFILL routes to fa_prefill_f16kv, whose per-row limit was qpos = bound[t]. With an explicit span bound[t] is the window's lo, not the causal upper bound, so every row kept only the lo column (win_lo <= col <= lo). Token 0 came out right because its window is {lo}. That is why the 7B's per-slot prefill (34-44 tokens, non-zero start) fed the rest of the network from corrupted hidden states: its KV and its logits were wrong from the first prefill, so every decoded token was garbage, while slot 0 (start 0, causal) stayed correct. The 0.5B never reproduced because it runs f32 KV, whose prefill kernel (gqa_attn_f32) had the windowed mask right.

The fix is the exclusive per-row limit qlim = CAUSAL ? bound[t] + 1 : bound[nt + t] (four mask sites). For the causal instantiation col <= bound[t] and col < bound[t] + 1 are the same integer comparison, so the pre-E1 instantiation keeps its codegen and output; only the windowed instantiation changes behaviour.

Evidence after the fix: the sweep is 22/22 bitwise equal (11 shapes x both KV dtypes, including the n=34, start=64, hd=128 prefill case); the server reproduction returns distinct, correct replies in both modes; the full CUDA suite passes on the device. The engine acceptance gate (server_batch_matches_serial_and_is_faster, 7B, 4 slots, --n-ctx 2048, templated prompts) passes too — and it passed before the fix as well, which is itself the lesson: its assertion is "batched text == serial text", and both paths ran the same faulty prefill, so a shared fault is invisible to a differential gate. Only a kernel-level sweep over the KV dtypes (plus the live server A/B with different prompts) could see this one. What this also says about the earlier rounds: the single broad structural fact they established still holds and was the useful half — the server exercises non-zero-start prefills that the engine test's 512-row single-sequence arena does not, which is why the gate now sweeps start = 64..256 and both dtypes.

Earlier text (kept for the record of how the diagnosis narrowed): a focused device test of the windowed instantiation with a long window at a non-zero start (34-64 rows) against the causal instantiation over the same rows — that should pin the exact kernel variant (the split/rows-per-warp bodies and the _bt variants are the candidates) and give a minimal reproduction, then the fix. The engine test as it stands is the end-to-end reproduction and now renders the server's prompt, so it fails the moment the bug is present and passes when it is fixed. Note also that this corrects the earlier "server-only, engine is fine" conclusion: that comparison used 5-token prompts on the engine side and 34-token ones on the server side. Impact: E6 made batching the default on CUDA, so this is the default behaviour on a GPU server today; E2's "GPU acceptance 1.9x" measurement inspected only request 0's text, so its timing stands but its correctness was never checked per request — that record is corrected here. Immediate mitigation: moot, and it would not have worked. Reverting E6's CUDA default (back to opt-in) was proposed while the cause was unknown; the f16 prefill fault lives in the per-slot prefill, which non-batched serving also performs (slots 1..n start at a non-zero KV cell), so MINFER_BATCH=0 was affected too — it merely produced plausible-looking text instead of derailed text. The default is left on and is correct as of the fix recorded below. | | 1 | CI has no GPU. The CUDA job only compiles the harness, so every device-gated test is a local, manual run — which is exactly how six device-only test bugs survived to 2026-09-18. | process | A self-hosted runner on this DGX Spark would put cargo test --features cuda into CI; nothing else does. Until then, anyone changing CUDA code must run it by hand and say so. | | 2 | The windowed instantiation's cost at equal width is unmeasured (cuda_causal_and_windowed_agree_on_the_same_rows proves equality, not speed). | measurement | E1b record; needs a device A/B over identical rows at one width. | | 3 | A varying batch width rebuilds the graph (GraphCache holds one graph at a time), so a server alternating 1-wide and N-wide decode steps re-allocates. | design | DONE (E4 S3, 2026-09-23): the cache holds one graph per GraphParams and a switch re-maps the allocator onto it (liveness + slots, no build); a repeated chunked prefill builds nothing the second time (a_repeated_chunked_prefill_stops_rebuilding: 3 builds/5 reuses, then 3/9). | | 4 | The op matrix's support table does not parse SUPPORT-MATRIX.md — it checks a Rust mirror, so a stale doc row stays green (it did, for multi_seq). | test hardening | A1/A8; recorded in the A1 record. | | 5 | Batching is opt-in even where it is measured faster. Closed by E6 (2026-09-19): the default follows the device, with =1/=0 as the override — amended 2026-10-06: #44 part (b) made Metal batched too, so it is CUDA and Metal on, CPU off. | product | done | | 6 | conversation_real_model_smoke and dump_real_q4k/q5k_tensor are red (ignored tests). Attribution done: the first fails identically on master + device, the others are pre-existing debug dumps. They are not gates, but a red ignored test is easy to mistake for noise. | pre-existing | Either fix their assertions/artifacts or mark them clearly in their doc comments; not caused by any PR in this campaign. | | 7 | Roadmap item 25 (metrics/observability) had no ticket — the only orphan from the A-era batch. | planning | F8, added with this section. | | 8 | The hardware lane: F1 (AVX2/AVX-512 K-quant dots, #56) needs an x86 host, which neither box this project owns is — dgxspark (aarch64, GB10 sm_121) and macbook (macOS 27.0.1, Apple M4 Pro) are the two machines, and the x86_64 (CI runner) is ephemeral and cannot host interactive work. Phase G's half of this row is closed (complete on a Mac 2026-10-06). | hardware | Sequencing §11 ("Hardware lanes") and §9's F table; what next: points at instead (#335): the macOS round-2 tickets tracked in #333 on macbook, and the open Linux-side tickets — #200 (CUDA Op::FusedQkvNorm), #212, #354, #356 — on dgxspark; no schedule is attached to either lane. F1 is the largest single CPU win and stays queued behind the x86 host. | | 9 | A sequence's logits' tail depended on its absolute arena offset — resolved by C6 (2026-09-19) (found while gating C3). The pre-C6 measurements stand and are what justified the fix: at cell 0 vs cell 8 the max |Δ| over the vocabulary was 2.6% relative with the greedy token unchanged and the run deterministic; the hand-built q/k/v → rope → store → attn graph is exact to ≤ 1.2e-7 at the model's own shape and equal for a 1-cell and an 8-cell offset; the arena layout is irrelevant (a split reservation is bit-identical per layer); positions had exactly four consumers per layer (96 = 24×4); a layer bisect put the entry at layer 0's attention output; and the rope-injection intervention proved the entry is RoPE alone (injecting run A's 48 rope outputs into run B made the logits bitwise identical), while a distributed ~1e-6 rope perturbation already saturates the tail (0.44 vs 0.43) with the greedy token stable from 1e-6 to 1e-2. C6 removed the coupling — positions are sequence-relative and the allocator resolves cells — so a cell move changes no angle: the offset tests now assert bitwise equality, C3's acceptance tightens from the named amplified-rounding class to bit-identical, and a compaction no longer re-ropes. Three method notes earned here: a zero from a perturbation probe means nothing without a loud control; a control validates the path, not the equivalence of the perturbation (a single element nudged by 1e-5 is not the offset's distributed 1.5e-5 — the earlier "refutation" was an over-read); and an intermediate buffer may only be read immediately after its own node runs (graph.outputs does not extend liveness). | measurement | Done by C6; the logical-positions design, its gates and the CUDA fused-op port (S3) are in §5. |

Test-infrastructure record (#171, 2026-09-26) — one home for the gate contract, and one failure-injection seam

The problem. The rules this campaign's gate work produced existed only as narrative inside two AGENTS.md bullets, interleaved with per-ticket history and the counts those tickets moved, so every agent re-derived them. Mutation checking in particular cost a bespoke mock per ticket (#147's MINFER_TEST_CALL_FAIL, #145's MINFER_TEST_LATCH_ERROR, #151's FailingForward, #167's switches), so "break it and watch the gate fail" was skipped.

What landed. The rules are stated once in GATE-CONTRACT.md — five of them (assert the value, not a relation between two code paths; a control arm that differs in the property under test; mutation evidence; work, not seconds; a device number with its provenance) — each with the precedent that produced it, and the two AGENTS.md bullets keep only the point-of-use command/default/counts plus a pointer. The failure-injection seam is src/testfail.rs: one presence-checked switch (MINFER_TEST_CALL_FAIL, the #147 name and semantics; exact-token comma list, all for every Rust site) with chokepoints at the batched forward (forward_batch), the allocator pool entry (alloc_in_pool), the backend execute entry (execute_node) and the weight registrar (register_weight), plus the device-side launch:*/attr:* sites of #147. The observation half is testfail::note_checked(site) / checked(site), bumped by the chokepoint itself, so a gate proves the path ran instead of reading the dispatch's own answer — the shape #141's f16 vectorization gate needed (F16_SIMD_PATH_CALLS). #147's matcher injection_names_site moved into testfail.rs, so there is one matcher and its exact-token tests now run in CI's CPU job instead of only on a --features cuda build.

Migration of the older knobs. MINFER_TEST_CALL_FAIL keeps its name and semantics (the contract the #147 gates depend on). #145's MINFER_TEST_LATCH_ERROR is left in place: it enables a deliberately-latching device gate, it does not select a chokepoint, so folding it into the token list would change what it means. #151's FailingForward mock stays in its CI gate (no model on a hosted runner), and the seam replaces it on the real path below. #167's gates use the pure registration rule and needed no switch.

Verification (2026-09-26, dgxspark: 20-core CPU, GB10 sm_121, CUDA 13.0).

CommandResult
cargo test --release (CPU)456 passed / 0 failed / 33 ignored unit + 10 / 0 / 6 integration
PARALLEL=0 scripts/real_model_gates.sh (CPU, serial)33 passed / 0 failed (37.59 s)
scripts/real_model_gates.sh (CPU, parallel)33 passed / 0 failed (33.98 s)
scripts/cuda_test.sh (GB10)524 passed / 0 failed / 37 ignored
FEATURES=cuda scripts/real_model_gates.sh (0.5B)37 passed / 0 failed
MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf FEATURES=cuda scripts/real_model_gates.sh37 passed / 0 failed
compute-sanitizer --tool memcheck over the serial CUDA unit suite0 API errors over 524 passed / 37 ignored (351.89 s)
cargo test --release --bin minfer -- --ignored --exact server::batch::tests::the_seam_fails_the_batch_forward_without_a_bespoke_mock --test-threads=1 --nocapture1 passed; prints [#171] forward_batch seam: one 500, slot released, no retry, 1 observed forward entries
MINFER_TEST_CALL_FAIL=forward_batch cargo test --release --bin minfer -- --ignored --exact server::batch::tests::server_batch_matches_serial_and_is_faster --test-threads=1FAILED — one environment variable reproduces #151's mutation and trips an existing gate ([testfail] deliberate panic injected at site 'forward_batch')
cargo test --release --bin minfer -- testfail::tests::the_seam_is_off_by_default with requested() ignoring the environmentFAILED at assertion failed: !requested("all") (reverted, sha256sum identical)
cargo test --release --bin minfer -- the_execute_chokepoint_is_observable with note_checked() a no-opFAILED left: 0, right: 2 (reverted, sha256sum identical)
rustup run stable rustfmt --edition 2021 --check on every changed .rsclean (rustfmt 1.9.0-stable)
python3 scripts/check_docs_links.py955 relative links resolve in 185 markdown files

The CUDA unit count moves 521 → 524 passed while gaining five tests, because #147's the_injection_matcher_matches_only_the_named_site was cuda-gated and moved to the always-compiled testfail module: −1 cuda-only, +4 always, +1 #[ignore]d real-model gate.

Honest scope. A script can require that mutation evidence is present; it cannot check that it is true. The seam removes the cost, not the discipline. The observation counter is thread-local, so a gate that runs the engine on another thread must read it there. The device-side matcher is necessarily a second implementation in C++ (a kernel cannot call into Rust); the two are documented as exact-token identical and the Rust half is the one unit-tested. The still-unbounded while engine.busy() steppers remain #160, and the count-consistency check that would enforce rule 5's provenance remains #94.

Test-infrastructure record (#173, 2026-09-26) — the op-timing gate reads its own sink, not a shared table

The problem. graph::scheduler::tests::op_timing_does_not_change_the_result_but_does_accumulate (F8/#51) took the process-global optiming::gate() mutex, forced the timing flag on, and asserted that a later quiet run left the process-global per-op table unchanged. The mutex serialized the flag flips only; it could not stop another test from executing a graph while the flag was on, so that test's records landed in the same two atomics between the gate's snapshots. Observed once on 2026-09-26 in a full parallel CPU suite (heavy device jobs also on the box): left: 2, right: 1 at the "a run with the flag off records nothing" assertion. A verdict that depends on what else happens to run beside it is not a gate — this is #173, the same isolation class #99 fixed for the KV format.

What landed. The accumulators moved into optiming::TimingSink — the same fixed [(nanos, calls)] table, now an instance rather than a static. The engine keeps one process-global sink (record / snapshot / reset delegate to it, so /metrics and every other consumer are unchanged), and BackendScheduler carries a TimingMode chosen at construction:

  • Global (default): the production policy — resolve MINFER_OP_TIMING once per execute and record into the process-global sink when it is on;
  • Off: never read the clock;
  • Private { sink: Arc<TimingSink>, enabled: bool }: record into a caller-owned sink, gated by a caller-owned flag — the isolation seam.

force(), GATE and gate() existed only to paper over the shared table and are deleted. The gate now asserts values from the sink its own scheduler wrote: exactly one silu and one add after the timing-on run (the small_graph is input → silu → add, and an input is never dispatched), and an empty sink after the timing-off run. A new control test, a_concurrent_graph_load_cannot_move_a_private_sink, runs 4 × 64 real executes into the shared global sink on four scoped threads while asserting eight private sinks stay at exactly one silu + one add each, then asserts the shared sink's total moved to 1 + 4 × 64 (the extra one through the free record). The load is bounded by work, every thread is joined by scope, and there are no sleeps or clocks (rule 4).

Why Private carries its own enabled instead of the recommended Private(Arc<TimingSink>) that always records: the "flag off records nothing" arm has to read the sink the scheduler would write, or the arm is refused by the mode rather than by the flag (rule 2). With an always-recording private sink, the off arm could only assert the global table — the shared state this ticket removes.

The stale-doc defect. The module doc claimed the property was "pinned by op_timing_off_by_default_leaves_the_table_empty"; no such test has ever existed. The doc now names the real ones — op_timing_flag_is_presence_checked_and_off_when_unset, off_mode_never_records, global_mode_follows_the_flag, and the scheduler gate — and the stale optiming::gate references in graph/copystats.rs are corrected.

Mutation evidence (rule 3). The failure-injection seam does not fit — the mutation ignores a flag, it does not fail a call — so this is the one-line implementation edit the rule allows. Dropping the enabled gate from TimingMode::resolve_for's Private arm (one line: enabled.then(|| sink.as_ref()) → Some(sink.as_ref())) is "the scheduler records while the flag is off":

$ cargo test --release --bin minfer -- \
    graph::scheduler::tests::op_timing_does_not_change_the_result_but_does_accumulate --exact
test graph::scheduler::tests::op_timing_does_not_change_the_result_but_does_accumulate ... FAILED

thread '...' panicked at src/graph/scheduler.rs:706:9:
a run with timing off records nothing: [OpTimingEntry { name: "add", calls: 1, nanos: 43584 },
  OpTimingEntry { name: "silu", calls: 1, nanos: 7056 }]

test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 492 filtered out; finished in 0.00s
error: test failed, to rerun pass `--bin minfer`
$ echo $?
101

Reverted byte-for-byte (diff -q clean; src/optiming.rs sha256 4f20689df7ea803ea703c96893725fd3914d7d9c2a911e4f864e9d92b9409d3f both sides).

Counts (rule 5). All on dgxspark (20-core aarch64), cargo test --release:

RunUnitIntegration
idle, run 1460 passed / 0 failed / 33 ignored10 / 0 / 6
idle, run 2460 passed / 0 failed / 33 ignored10 / 0 / 6
under 24 CPU spinners (20 cores)460 passed / 0 failed / 33 ignored (95.89 s)10 / 0 / 6
the two gate tests ×20 under the same load2 passed / 0 failed every iteration—

The delta is +3: optiming loses the_gate_switches_the_scheduler_path_on_and_off (it tested the deleted force) and gains off_mode_never_records, global_mode_follows_the_flag and private_mode_uses_its_own_sink_and_flag (−1, +3); the scheduler replaces one gate with the rewritten gate plus the concurrent-load control (+1). The x86_64 (CI runner) row moves by the same +3 (455 → 458), which test-linux-cpu's --check-live confirms against its own log; AGENTS.md and docs/status.toml carry both rows.

Audit of the other global-table readers (acceptance item 4). server::metrics::tests:: op_timing_family_renders_seconds_with_nanosecond_resolution builds MetricsSnapshot.ops by hand and never reads the process table, so it is outside the fault class. The two tests that render a live ServerMetrics::snapshot() (the_token_families_render, both_threads_write_into_one_registry) do read the global table incidentally, but assert only token/KV families and well-formedness — an op family a concurrent execution adds cannot change their verdict. No follow-up is needed: the only shared-table writer left in the suite is the #173 control test itself, which resets the sink before and after.

Honest scope. The new gate proves that a scheduler's verdict reads only its own sink, and the control test proves a concurrent global load cannot move it; it does not prove that any other future test will keep its hands off the process table. That is now a rule a reviewer checks, not a property CI can see. The removed force() also removes the suite's only way to flip the cached process flag — the production Global path is exercised by the pure resolve_for tests and the (unset) flag's rule, not by an in-process toggle; a real MINFER_OP_TIMING=1 end-to-end run is still the device/manual evidence recorded under F8.

CUDA test-infrastructure record (#185, 2026-09-26) — the parallel device suite's SIGSEGV is the global capture window racing an unlocked weight copy

The observation. A parallel cargo test --release --features cuda -- --ignored (--test-threads = 20 cores, 38 tests) on GB10 sm_121 / CUDA 13.0 / driver 580.178.04 either segfaulted or produced an unstable failure set. Reproduced in this session on the same box and binary (target/release/deps/minfer-d35a4f8f70a3b190 --ignored, 2026-09-26):

runsSIGSEGV (rc 139)completed runs' failuresfailure sets
6 (bare, with timeout 300)3 / 64, 3, 9 failedthree distinct sets
10 (under gdb, handle SIGSEGV stop print nopass)2 / 10——

Across the 16 runs the failing members moved over conversation::tests::*, server::batch::tests::*, models::qwen2::graph::tests::{a_packed_kv_cache…, an_auto_offload_plan…, async_cross_copies…, two_cuda_engines…} and graph::cuda_backend::tests::cuda_map_window… — five distinct sets once the ticket's own two runs are counted. Not one of them is a wrong value: every one is either a 901 capture invalidation, a process-global assertion, or a timing relation. ulimit -s is 8192 KiB (no guard-page frame in the backtraces); compute-sanitizer --tool memcheck is 0 errors on the serial suite, so a device memory error was not the candidate.

The backtrace. Two independent gdb crashes stopped at the same faulting instruction inside the driver:

Thread N "models::qwen2::" received signal SIGSEGV, Segmentation fault.
0x0000ffffab24da8c in ?? () from /lib/aarch64-linux-gnu/libcuda.so.1
#6  cuMemcpyHtoD_v2
#9  cudaMemcpy
#10 <minfer::cuda::CudaState>::register_weight
#11 minfer::models::weight_reg::register_cuda_weight
#12 minfer::models::qwen2::loader::load_tensor
#13 minfer::models::qwen2::loader::load
#14 minfer::models::load_model_configured
#15 tests::a_packed_kv_cache_answers_like_the_f32_one              (crash 1)
     tests::two_cuda_engines_with_different_kv_layouts_run_interleaved   (crash 2)

and, at that instant, a second thread inside the same driver:

#8  cuGraphInstantiateWithFlags
#10 cudaGraphInstantiate
#11 <minfer::cuda::CudaState>::graph_end_capture_to_exec
#12 <CudaBackend as minfer::graph::backend::Backend>::synchronize
#13 <minfer::graph::scheduler::BackendScheduler>::execute
#14 <minfer::models::qwen2::graph::Qwen2Graph>::forward_batch
#15 minfer::server::chat::guarded_forward_batch
#16 <minfer::server::batch::BatchEngine>::submit_on
#17 tests::a_long_prefill_keeps_another_slot_decoding       (crash 2's co-tenant)

One thread holds an open cudaStreamCaptureModeGlobal capture window (cudaStreamBeginCapture(stream, 1)) and is instantiating it; the other issues a plain, not stream-ordered, blocking cudaMemcpy (H2D) from the weight-registration path.

A third crash — captured once the two-chokepoint guard below was already in place — is a second unguarded entry rather than the same one, and it is why the guard has a third chokepoint:

Thread 5 "graph::cuda_bac" received signal SIGSEGV
#5  cuLaunchKernel
#8  __device_stub__gqa_attn_split_combine
#9  launch_gqa_attn_split_f32kv
#10 graph::cuda_backend::tests::cuda_map_window_costs_no_more_than_the_span_it_replaces

That gate drives CudaBackend/CudaState directly (no scheduler::execute), so no scheduler-level guard can see it; 3 of 8 gdb runs crashed there after the register-weight path was refused. It now takes the device token itself.

The determination: (a), the documented capture-stream race. CudaState::stream_lock() serializes the paths that take it, and graph_replay_step holds it across the capture window — but CudaState::register_weight takes no lock at all, and a Global-mode capture is invalidated by another thread's non-capturable driver call. The benign mode of that race is the cudaErrorStreamCaptureInvalidated (901) the run logs are full of; the recorded mode is the driver faulting instead of returning. (b) teardown/drop race: ruled out — neither crash has a Drop, cudaFree, cudaFreeHost, cudaGraphExecDestroy or graph_destroy frame; the faulting frame is a live forward cudaMemcpy and the co-tenant is cudaGraphInstantiate, not a destructor. (c) memory error the serial path only avoids by timing: ruled out as the root — the destination pointer is the one cudaMalloc returned two lines earlier in the same function, no minfer unsafe passes a wild pointer, and compute-sanitizer is clean serially. (iv) stack overflow: ruled out — the fault is inside libcuda several frames down, not at a guard page. What makes the crash bite the serial path too is not a timing accident but the sharing: any future two-thread device use (a second engine thread, a parallel harness) re-creates it.

What landed (the loud refusal #185 asks for). src/device_entry.rs owns a process-wide, re-entrant-per-thread token for "inside the CUDA device path". BackendScheduler::execute takes it for a whole execution when the graph has a CUDA split (the capture window's lifetime; the test is the split set, not the allocator's state, so a CPU-only graph on a CUDA-enabled allocator is not refused), and models::weight_reg::register_cuda_weight takes it for each registration (the copy that crashes). A second thread is refused with the operation it was doing, the operation already inside, the mechanism, the evidence and the remedy (--test-threads=1 / scripts/cuda_test.sh) — before any driver call. The module is feature-independent and pure on purpose, so the CPU CI job runs its test; the same reason models::weight_reg::cuda_weight_reg keeps its decision pure.

Verification that it refuses rather than crashes (rule 5 numbers). The same parallel --ignored command, GB10 sm_121, 2026-09-26, as the guard grew (before → 3 / 6 bare runs and 2 / 10 gdb runs SIGSEGV'd with no guard at all):

guardbare runsSIGSEGVgdb runsSIGSEGVnamed refusals per bare run
two chokepoints (execute + registration)61 / 683 / 8 (all cuLaunchKernel from cuda_map_window…)25, 23, 23, 25, 27, 0
three (landed, + the direct-driver gate)100 / 1060 / 624, 23, 23, 30, 23, 24, 24, 23, 23, 21

Every refusal is the device_entry message; no run reached the driver concurrently. That the intermediate state still crashed is the useful part: the guard converted the two backtraced mechanisms into refusals and exposed the third, which is a direct-driver test the scheduler cannot see. The serial configuration is untouched: scripts/cuda_test.sh, GB10 sm_121, 2026-09-26 → 544 / 0 / 38 (was 541 / 0 / 38; +3 tests: the guard's own test, the explicit-auto-budget test and the per-backend stream-sync gate), and the sanitizer stays at 0 errors over the same 544.

Mutation evidence (rule 3). Deleting the Some(holder) if holder.thread != me arm from device_entry::enter — i.e. letting every thread in — makes device_entry::tests::the_device_path_is_exclusive_across_threads_and_re_entrant_on_one fail at a second thread must be refused. The per-backend counter's gate is mutated by making CudaBackend::stream_sync_count return crate::cuda::stream_sync_count(), which fails stream_sync_counts_are_per_backend_not_process_wide on its second assertion.

What does not change, and where it is filed. The race is not fixed: the guard refuses the configuration instead of making it correct, and it is a chokepoint, not a structural exclusion — a caller that reaches CudaBackend::execute_node/synchronize/graph_replay_step, or CudaState::register_weight directly rather than through those two entry points, still runs unguarded (as does Drop for CudaBackend's frees, which the pre-existing stream_guard serializes unless the backend is mid-capture). Per-instance streams/capture contexts, or extending the capture discipline to every un-ordered device call (register_weight's cudaMalloc/cudaMemcpy, cudaMemGetInfo, cudaHostAlloc/cudaFreeHost, the Drop frees), is #188, with both backtraces attached. The S4 map-window timing gate's co-tenant sensitivity (median 1.398x once, in the parallel run; decode half 0.916x and the whole test green in the sibling run) is the #154 class, not the crash: #189.

Two test-hygiene defects, fixed here (they are why the failure set moved). async_cross_copies_never_block_and_stay_bitwise_identical read the process-wide cuda::stream_sync_count() (run A: the async arm counted 4160 stalls against the synchronous arm's 728 — a foreign test's syncs inside the delta). The count now lives on the CudaBackend (stream_syncs, bumped by the one state_sync helper) and both F5 gates read it there, exactly as blocking_readbacks and copystats' accumulators already did; the process-wide function stays for the "host stalls in this process" figure and its doc now says a gate must not read it. And an_auto_offload_plan_fits_the_budget mutated the process-global MINFER_GPU_MEM; it now uses OffloadRequest::AutoWithBudget(64) — an explicit argument, the repo's convention since #99/#153 — and its second arm uses the device's measured free bytes, so the gate neither reads nor mutates the environment.

CUDA concurrency record (#188, 2026-09-27) — the stream, the capture window and the activation scratch become per CudaBackend

The defect #188 names. #185 recorded the parallel #[ignore]d device suite's SIGSEGV as cudaStreamBeginCapture(stream, 1) — cudaStreamCaptureModeGlobal — racing a weight-registration cudaMemcpy that does not take the stream lock. Under Global semantics a driver call from another thread belongs to the capture window, so it either invalidates it (cudaErrorStreamCaptureInvalidated, 901: the benign mode, all over the run logs) or faults inside the driver (the recorded mode: cuMemcpyHtoD_v2 under register_weight while another thread sat in graph_end_capture_to_exec → cudaGraphInstantiate). The reason one window could even be "the" window was structural: static CUDA: OnceLock<Option<CudaState>> gave the process one stream, so two engines could not capture independently at all. #185's device_entry token made the configuration a loud refusal; it did not make it correct.

The instrument, built first (gate contract rule 5 needs a number, and the raw crash rate was too low to judge a fix). graph::cuda_backend::tests::capture_window_on_one_thread_survives_a_weight_registration_on_another opens a capture window on thread B, records a device→device copy into it, signals thread A, has A perform a weight registration while the window is open, then closes the window and checks both that cudaStreamEndCapture returned 0 (never 901; cuda::last_capture_end_code) and that the replayed graph produced the bytes the window recorded (the wrong value is seeded into dst first, so the value arm cannot pass on the setup — rule 1). Two env knobs make it the experiment: MINFER_PROBE_STREAM=context captures on CudaState's own (blocking) stream — the pre-#188 shared-stream model — and MINFER_PROBE_LEGACY_MEMCPY=1 uses the pre-#188 blocking cudaMemcpy; MINFER_CUDA_CAPTURE_MODE=0|1|2 picks relaxed/global/thread-local. GB10 sm_121, 2026-09-27, 5 process runs per cell (90 s watchdog):

capture streamregistrationmoderesult
context (blocking, shared)blocking cudaMemcpy (pre-#188)global (1, pre-#188)5/5 hang
context (blocking, shared)blocking cudaMemcpythread-local (2)5/5 hang
context (blocking, shared)blocking cudaMemcpyrelaxed (0)5/5 hang
instance (non-blocking)stream-ordered (landed)global (1)5/5 pass
instance (non-blocking)stream-orderedthread-local (2, adopted)5/5 pass
instance (non-blocking)stream-orderedrelaxed (0)5/5 end_code=901

The mode finding, measured rather than assumed. (i) The blocking copy is a hard deadlock against a capture window on a blocking stream — the null stream implicitly synchronizes with every blocking stream, including the one that cannot complete until the host closes it — in all three modes. So the mode is not what fixes the historical setup; a stream-ordered copy on a per-instance non-blocking stream is. (ii) With the structural fix in place, relaxed still returns 901 in 5/5 runs (the probe's exact forbidden outcome; cudaMalloc inside the window also fails), so relaxed is ruled out by measurement. (iii) global passes the instance cell but is the mode that lets a foreign thread's call belong to the capture — the class this ticket exists to remove — so the adopted mode is cudaStreamCaptureModeThreadLocal (2), which bounds invalidation to the capturing thread.

The structural change (both halves the ticket asks for). CudaBackend now owns a cudaStreamNonBlocking stream (CudaState::create_stream, #[188]; created in with_layout, destroyed in Drop after the pool/scratch/host/graph teardown). crate::cuda::bind_stream publishes it in a thread-local for the duration of every backend device operation, and CudaState::stream() answers with it — so the ~60 launch/copy/event/capture/replay/synchronize helpers keep their signatures and follow the instance: the launchers, the pinned H2D ring (write_input_async), D2D (copy_device_to_device), the F5 copy_cross/await_cross event path, cudaGraphLaunch/graph_begin_capture/graph_end_capture_to_exec, state_sync, copy_cells (kv_move_rows), alloc_buffer/free_buffer and Drop. Per stream too: the buf_* activation scratches became StreamScratch (a map keyed on the current stream), the MmqCache memo is keyed the same way (a hit records a scratch pointer, so a shared memo would hand one engine's plane to another), and the pinned staging ring is keyed on the stream. Weight registration (CudaState::register_weight) no longer issues a blocking cudaMemcpy at all: it queues the H2D copy on the context stream and waits on that stream, so it can never be recorded into a backend's window. Context-wide calls stay shared and are named as such: cudaMalloc/cudaFree, cudaMemGetInfo, cudaHostAlloc/cudaFreeHost, the weight registry and the derived planes. The process-wide stream_lock/stream_guard is gone — the capture window is the backend's own capturing field. Finally, the device tests that drive cb.state.* kernels directly (store_kv_q8_0, gqa_attn_split*, the q5_k/q6_k decode parities, the map-window A/B) now bind the backend's stream: a context-stream launch read back through an instance-stream sync was a race the shared stream used to hide, and the serial suite caught exactly two of them (cuda_capture_staging_order_and_fallback, cuda_verify_attention_nt_invariance) before the binding was added.

The negation: #185's guard is narrowed, and its test updated rather than deleted. device_entry's token no longer covers BackendScheduler::execute or weight_reg::register_cuda_weight — both are per-instance/stream-ordered now. It stays on the one path that still reaches CudaState's helpers without a backend to bind: CudaState::layer_gpu, the legacy per-layer path, which drives the context-keyed buf_* scratches (the graph path and main do not call it). The module's docs and its test now state that scope, and the test asserts the refusal messages name the narrowed mechanism (unbound) and the one remaining path (layer_gpu). The positive property that replaces the refusal is the concurrent gate below.

The concurrent gate. models::qwen2::graph::tests::two_cuda_engines_forward_concurrently_and_stay_bitwise_identical runs two CUDA engines (f32 and q8_0) on two OS threads, both caches alive across a barrier so the forwards genuinely overlap, and compares each engine's logits bitwise (max |Δ| == 0) against its own serial reference. It also asserts the two live backends hold different device stream pointers. The verdict is a value, not a timing, so the S4 map-window co-tenant gate (#189) does not decide it. GB10 sm_121, 2026-09-27: [188] two CUDA engines forwarding on two threads (8 decode steps each): streams 0xe8f738039e60 vs 0xe8f72c039de0; concurrent-vs-serial drift 0 / 0.

Before/after on the parallel #[ignore]d configuration (<test binary> --ignored, default parallel, GB10 sm_121, 2026-09-27). "Before" is master 440178b with #185's guard call sites removed — i.e. the crashing configuration the guard was landed to suppress. "After" is this PR (which no longer routes the graph path or registration through the guard at all). Bare, 6 runs each:

exit-okSIGSEGVruns with test failuresfailure set
before (guard removed)0 / 61 / 6 (rc 139, core dumped)6 / 6 (3, 5, 10, —, 5, 4 failures)unstable: 900/901 launch refusals across server::batch, conversation, a_session_resumed…
after (this PR)4 / 60 / 62 / 6 (1 each)cuda_map_window_costs… (#189 timing gate) and server_batch_matches_serial_and_is_faster (#154 timing gate)

Under gdb (gdb -batch -ex 'run --ignored' -ex 'thread apply all bt', 3 runs each): before 1 / 3 SIGSEGV (faulting inside libcuda.so.1), after 0 / 3. The "after" run also runs the configuration instead of refusing it: 39 tests pass in parallel, and the only failures left are the two known co-tenant timing gates — #189 (out of scope by the ticket) and #154's batch ratio.

Verification (rule 5 numbers). All GB10 sm_121, 2026-09-27. bash scripts/cuda_test.sh → 545 / 0 / 39 (was 544 / 0 / 38; +1 the probe, +1 the concurrent gate, which is #[ignore]d). cargo test --release (CPU, dgxspark) → 465 / 0 / 33 unit + 10 / 0 / 6 integration, unchanged — the renamed device_entry test keeps the count. compute-sanitizer --tool memcheck --target-processes all <test binary> --test-threads=1 → 0 API errors over the same 545 / 0 / 39. The real-model device set (FEATURES=cuda scripts/real_model_gates.sh) → 39 / 0 for the 0.5B config and 39 / 0 for the Qwen3-0.6B config (was 38 / 0; +1 the concurrent gate).

Mutation evidence (rule 3). (a) Mode: running the probe with MINFER_CUDA_CAPTURE_MODE=0 (relaxed) makes it fail 5/5 with end iteration 0: the capture window was invalidated (code 901 = cudaErrorStreamCaptureInvalidated) — the probe detects exactly the outcome the acceptance forbids, and it is why relaxed was not adopted. (b) Per-instance stream: replacing CudaBackend::with_layout's state.create_stream() with state.stream() (the pre-#188 shared stream) makes the concurrent gate die with SIGSEGV (signal 11) — the pre-#188 crash reproduced by one line — instead of reaching its two different device streams assertion.

What did not change / honest scope. The legacy layer_gpu path and the direct CudaState scratch helpers (upload_hidden, download_logits, …) still share the context stream and its scratch; they are #[allow(dead_code)] legacy surface (the graph path is the production path) and the guard keeps them to one thread at a time. Metal is untouched. The device suite's own default stays --test-threads=1 (scripts/cuda_test.sh): the two new gates are parallel-safe, but other device tests still mutate process-global state (MINFER_GPU_MEM history, timing gates) and the wrapper's serial default is not this ticket's to change.

Test-infrastructure record (#189, 2026-09-27) — the S4 map-window A/B is a paired sign test with a value arm

The defect. graph::cuda_backend::tests::cuda_map_window_costs_no_more_than_the_span_it_replaces — the S4 device half of #123's fix — was a pure stopwatch: it interleaved matched rounds of the span window (row0 + i) and the map window (row resolved through a run list) and asserted the median of 9 per-round ratios <= 1.25x. It failed once in the parallel #[ignore]d device run (GB10 sm_121, CUDA 13.0, driver 580.178.04, 2026-09-26) on the prefill phase:

[s4-ab] prefill nt=512 nkv=512 hd=128: span 89.9 / map 125.1 us/launch (median of 9 interleaved
rounds of 50); per-round ratios [0.632, 0.697, 0.989, 1.091, 1.398, 1.440, 1.466, 2.198, 6.695]
— median 1.398x
a map prefill costs 1.398x the span it replaces

Five of the nine matched pairs were above the bar — exactly the count a median of nine flips at. The decode half passed in the same run (median 0.916x) and the whole test passed in a sibling run. The 1.25x bar had been justified by #123 against "16 CPU spinners plus two concurrent CUDA attention loops" (median 1.079–1.145 prefill); the parallel #[ignore]d suite is a far heavier co-tenant (~38 device tests, several capturing CUDA graphs, one GPU). The verdict was about the machine, not the kernel — the #154 class, and not the #185 SIGSEGV that #188 fixed structurally.

The fix — the value arm first (rule 1). Before any timing the gate now asserts, on its own fixture (rows [512, 2560) carry distinct K/V; the map window's base cell is 512):

  • a one-row map window returns exactly that row's V, bit for bit against the host's input row — an absolute oracle, not a mode-vs-mode relation;
  • a two-run map window returns the span's bytes over the same rows, bit for bit — f32 KV at the decode shape and f16 KV at the FA-prefill shape. One run at cell 0 is indistinguishable to a resolver that reads (cell, len) as (lo, hi); two runs at a non-zero base are not;
  • testfail::note_checked("cuda_attn_map_window") — bumped in the launchers CudaState::gqa_attn_split / gqa_attn_kv_prefill — is 0 after a span call and 1 after a map call. The observation is counted rather than read from the dispatch's own report (the contract's "observation half"), and it is paired with the bitwise arms so "the map kernel ran" and "it resolved the span's rows" are separate facts.

The fix — a paired sign test (rule 4). The timing verdict is the count of matched pairs whose map arm is above 1.25x its own span arm; the gate refuses only at 7 of 9 (PAIRS = 9, fixed in advance), the one-sided sign test at alpha = 46/512 = 0.090. A minority of disturbed rounds can no longer decide the verdict, and neither can the bare majority the recorded run had; a doubled map cost moves all nine. The bar is unchanged at 1.25x and the timed fixture is the pre-#189 one (constant K/V, one run at cell 0), so the recorded margins stay comparable. Every per-round ratio and the refusal count are printed.

Verification (rule 5 numbers). All GB10 sm_121, CUDA 13.0, driver 580.178.04, 2026-09-27.

CommandResult
gate alone: cargo test --release --features cuda --bin minfer cuda_map_window_costs… -- --ignored --nocapture --test-threads=1decode 24.6 / 25.0 µs/launch, 0/9 refusals; prefill 82.8 / 91.1 µs/launch, 0/9
parallel #[ignore]d: cargo test --release --features cuda --bin minfer -- --ignored --nocapture, 6 runs6 × 39 passed / 0 failed / 0 ignored; gate refusals 0–3 per phase; worst single pair 5.09x
mutation: MINFER_S4_AB_MAP_REPS=2 + the gate-alone command0 passed / 1 failed, decode 51.7 vs 24.6 µs/launch = 9/9 above 1.25x
scripts/cuda_test.sh546 / 0 / 39 (was 545 / 0 / 39; +1 the pure statistic test)
FEATURES=cuda scripts/real_model_gates.sh (0.5B, then Qwen3-0.6B)39 / 0 and 39 / 0
compute-sanitizer --tool memcheck --target-processes all <test binary> --test-threads=10 API errors over 546 / 0 / 39
cargo test --release (CPU, dgxspark)465 / 0 / 33 unit + 10 / 0 / 6 integration, unchanged
python3 scripts/check_status.py --checkexit 0

Measured — the gate alone, idle (GB10 sm_121, CUDA 13.0, driver 580.178.04, 2026-09-27; cargo test --release --features cuda --bin minfer cuda_map_window_costs_no_more_than_the_span_it_replaces -- --ignored --nocapture --test-threads=1).

phasespan / map µs/launchper-round ratiosrefusals at 1.25xverdict
decode (nkv 2048, nh 28, nk 4, hd 128)24.6 / 25.0[0.871, 1.007, 1.011, 1.013, 1.014, 1.017, 1.044, 1.050, 1.051]0 / 9pass
prefill (nt 512, f16 KV)82.8 / 91.1[1.054, 1.089, 1.097, 1.099, 1.100, 1.100, 1.104, 1.104, 1.189]0 / 9pass

Measured — the parallel #[ignore]d configuration (the one that produced the failure), six runs (GB10 sm_121, CUDA 13.0, driver 580.178.04, 2026-09-27; cargo test --release --features cuda --bin minfer -- --ignored --nocapture, default parallel harness). The gate's refusal counts per run, and the whole set's result:

rundecode refusalsprefill refusalsworst per-round ratio seensuite
10 / 93 / 91.57 (prefill)39 passed / 0 failed / 0 ignored
21 / 93 / 92.07 (prefill)39 / 0 / 0
30 / 93 / 91.48 (prefill)39 / 0 / 0
41 / 92 / 95.09 (decode)39 / 0 / 0
50 / 90 / 91.15 (decode)39 / 0 / 0
62 / 93 / 94.09 (decode)39 / 0 / 0

The co-tenant moved up to 3 of 9 pairs above the bar (the recorded failure's 5 is inside the tolerance); a single disturbed pair reached 5.09x in run 4 and the sign test still returned green. The whole set was 6 / 6 green, including the #154 batching gate that had also failed once in this configuration. Honest reading: in these six runs the co-tenant was lighter than in the recorded one — no run's median exceeded 1.25x — so they show the gate is not decided by the co-tenant, not that the new statistic rescued a median-red run. That case is the recorded distribution itself, replayed by the pure test graph::cuda_backend::tests::the_s4_ab_statistic_absorbs_a_loaded_run_and_still_refuses_a_real_regression (5 of 9 refusals pass at 7; the old median of the same ratios is 1.398x and red).

Before/after on the parallel configuration. The #188 record's post-fix state had 2 of 6 parallel #[ignore]d runs with one failure each — this gate once and server_batch_matches_serial_and_is_faster (#154) once. With this ticket's statistic the same configuration is 6 / 6 green (39 / 0 each). Honest reading: in these six runs the co-tenant was lighter than in the recorded one — no run's median exceeded 1.25x — so the six runs prove the gate is not decided by the co-tenant, while the recorded distribution (5 of 9 above the bar, median 1.398x) is replayed by the new pure test graph::cuda_backend::tests::the_s4_ab_statistic_absorbs_a_loaded_run_and_still_refuses_a_real_regression, which asserts the sign test passes it and that the old median of the same ratios is red.

The stale #185 premise. The ticket text says the parallel --ignored device run is "refused loudly (src/device_entry.rs)" and asks for the guard to be removed for the test session. After #188 (master fe1b7cd) that guard covers only the legacy unbound CudaState::layer_gpu path — crate::device_entry::enter has no other call site in src/ (grep -rn 'device_entry::enter' src/ names only src/cuda.rs, plus a doc-comment mention in models/qwen2/graph.rs) — and BackendScheduler::execute and register_cuda_weight no longer take it. The parallel configuration was therefore run as-is, with no guard change; it is the same narrowing #188's record already measured.

Mutation evidence (rule 3). MINFER_S4_AB_MAP_REPS=2 (src/cuda.rs::s4_ab_map_reps) issues every map-mode attention launch twice, so the gate's timed map arm pays twice the work — the reproducible form of #123's map-work doubling, and an implementation seam rather than a test edit. Armed, the gate fails with 9/9 pairs above the bar:

[s4-ab] decode nkv=2048 nh=28 nk=4 hd=128: span 24.6 / map 51.7 us/launch (9 interleaved matched
pairs of 100); per-round ratios [1.706, 1.745, 2.052, 2.091, 2.095, 2.100, 2.102, 2.136, 2.358]
— 9/9 above 1.25x (sign test refuses at 7)
thread '…cuda_map_window_costs_no_more_than_the_span_it_replaces' panicked: the map window is above
1.25x the span in 9 of 9 matched pairs … FAILED

The mutation is an env switch, so the unmutated run is the same binary with the variable unset — there is no source mutation to revert, and git diff on the tree carries only the #189 change.

Docs. docs/CUDA-BACKEND-DESIGN.md gains §7.10 (the statistic, the bar, the value arm, the dated device tables); AGENTS.md's CUDA counts bullet and docs/status.toml move 545 / 39 → 546 / 39 for the pure statistic test; the gate's doc comment (which claimed "It no longer needs an otherwise quiet box" — the claim #189 refutes) now states the sign test and the value arm.

Test-infrastructure record (#261 step 0, 2026-10-04) — the source-layout plan and the four-layer convention

What landed. A docs-only PR (9174644, #269) that fixes the shape of the campaign splitting src/cuda.rs, src/cuda_kernels.cu, src/metal.rs and src/metal.metal: docs/SOURCE-LAYOUT-PLAN.md (new — the decision, the target tree, the per-backend file tables, the tooling changes, the documentation-anchor plan and the interaction table for the 18 affected open issues), its mdBook chapter, the four-layer convention in AGENTS.md (L1 runtime / L2 launch / L3 <backend>/kernels/ / L4 executors; device-first, one polymorphic seam), the backend-layer section in docs/ARCHITECTURE.md — including the rule that a shared common needs two real implementations, of which allocplan::DeviceMemory was the one candidate (its second real implementation, Metal, landed as #53 — and the answer was still no common module: plan §1.3 rule 3) — and a docs/BACKENDS.md footnote. Four of the six stale claims folded into #219 are fixed here; register_weight's prose and the src/cuda.rs:3471 banner naming the deleted CudaCommandBuffer ride with step 1 (#262), which is the step that moves that code.

Verification (rule 5 numbers). dgxspark (aarch64, GB10 sm_121), 2026-10-04, in the step's worktree:

CommandResult
python3 scripts/check_docs_links.py .989 relative links resolve in 189 markdown files, exit 0 (base 5386a1c: 985 / 188)
python3 scripts/check_status.py --checkprose agrees with docs/status.toml (7 phases, 7 count rows: 1 live-checkable, 6 recorded measurements); exit 0
CI run 371970273397 / 7 green (check-docs, test-linux-cpu, build-linux-cuda, build-macos, check-viz, check-pr-body, lint-workflows), zero code annotations — the single build-macos annotation is GitHub's runner-capacity notice

Mutation evidence (rule 3). Two mutations in the worktree, each reverted: a broken relative link at docs/BACKENDS.md:18 (./SOURCE-LAYOUT-PLAN-NOPE.md) → check_docs_links prints BROKEN LINK docs/BACKENDS.md:18: ./SOURCE-LAYOUT-PLAN-NOPE.md and 1 of 989 relative links do not resolve (in 189 markdown files), exit 1; a perturbed copy of the manifest (passed = 478 → 479, prose untouched) → check_status prints AGENTS.md:92: cpu-unit (x86_64 (CI runner)) passed: prose says '478', docs/status.toml says '479', exit 1.

Deliberately out of scope. No src/ file changed. The plan document's pre-split line ranges are rewritten by the step that moves the code; its two docs/CUDA-BACKEND-DESIGN.md line citations were replaced by section references in the follow-up that carries this record, because step 0 itself shifted them — the same class of rot this campaign exists to remove.

Test-infrastructure record (#267 step 6, file 1, 2026-10-04) — the long test files start moving: graph/cuda_backend/tests.rs becomes a parent plus 11 topic files

What landed. The crate's longest test file, src/graph/cuda_backend/tests.rs (8,603 lines; 89 fn definitions in all — 74 top-level items, which are 61 #[test] of which one is #[ignore]d plus 13 helpers, and 15 nested helpers: 12 inside test bodies and 3 in the test-only impl CudaBackend shim), is split by op family into src/graph/cuda_backend/tests/<topic>.rs, each named by a mod declaration in the now-106-line parent (#267, PR #275). It is a pure move: every top-level fn/const item is byte-identical apart from rustfmt, the 61 test leaf names are the same set, and nothing was added, removed, renamed, re-gated or relaxed. The topic file is staging (4 tests), pool (2), elementwise (4), matmul (6), mmvq (3), prefill (4), weights (6), kv (9), attention (4), attn_window (6) and capture (13); the shared fixtures (device, pool, assert_close), the use lines and the test-only impl CudaBackend shim stay in the parent, which every topic reaches through use super::*;. src/graph/cuda_backend.rs keeps its #[cfg(test)] mod tests; untouched — only the parent gained child mods, and the layout checker's declaration walk reaches tests/<topic>.rs through module_dir().

This is the shape the guard exists for. A test file no mod names is never compiled, so its tests silently stop running while cargo test and CI stay green. scripts/check_source_layout.py rule 2 is the only thing that catches it, and this is the first src/ tree in the repo where a test module is a directory; AGENTS.md's test-file paragraph records the convention in one sentence.

Verification (rule 5 numbers). dgxspark (aarch64, GB10 sm_121), 2026-10-04, in the step's worktree, with the pre-move baseline taken in that same worktree first:

CommandResult
scripts/cuda_test.sh567 / 0 / 42 (pre-move baseline in the same worktree: 567 / 0 / 42)
cargo test --release481 / 0 / 36 unit + 10 / 0 / 6 integration (baseline 481 / 0 / 36 + 10 / 0 / 6)
python3 scripts/check_source_layout.pysrc obeys the layout rules, exit 0
python3 scripts/check_doc_line_anchors.pyexit 0 (4 OUT-OF-RANGE anchors re-pointed at the moved tests)
python3 scripts/check_docs_links.py, scripts/check_status.py --checkexit 0 each
cargo fmt --all --checkclean
CI run on PR #2757 / 7 green, zero code annotations

Mutation evidence (rule 3). mod staging; → // mod staging; in the parent, then reverted. The layout checker exits 1 and names the orphan: src/graph/cuda_backend/tests/staging.rs: not reachable from src/main.rs — no \mod` declaration names it, so it is never compiled (tests in it would silently not run). The CUDA unit row drops **567 → 563** passed (the four stagingtests), 42 ignored unchanged — so the tests really did move out oftests.rs` into a file that only the declaration compiles.

Docs swept. AGENTS.md (the one-sentence convention, in the existing test-file paragraph), docs/SOURCE-LAYOUT-PLAN.md (§6.4 gains the tests.rs → tests/<topic>.rs mapping row; the §Status table's step-6 row), and the four live anchors in docs/cuda_tutorial/05-kernels-attention-host.md, which now name cuda_kv_f16_roundtrip_attn at cuda_backend/tests/kv.rs:1011 and cuda_graph_replay_bit_parity at cuda_backend/tests/capture.rs:274. The #238/#239 test-only annotations in src/graph/{cuda_backend,alloc,copystats,builder,cpu_backend}.rs, src/q4k_dsc.rs and src/server/batch/tests.rs are re-anchored to the deeper module path (…::tests::<topic>::<leaf>); src/cuda.rs is deliberately left alone because #262 owns it in a parallel worktree, and re-anchoring its four #238 paths is the one follow-up.

What the brief had wrong (the tree wins). The ticket and the step-6 blueprint both said the file has "88 fn items = 61 #[test] + 1 #[ignore]d test + 26 helpers". Measured on the pre-split file: 61 #[test], of which 2 are #[ignore]d (cuda_map_window_costs_no_more_than_the_span_it_replaces at #[ignore = "timing: needs a CUDA device"], and cuda_real_model_registers_q4dsc_planes_only_for_q4k) plus 13 top-level helpers = 74 top-level fn items; 12 further helper fns are nested inside test bodies and 3 inside the test-only impl CudaBackend shim, for 89 fn definitions in all. The "88" counted those nested helpers as items and double-counted the ignored test. The numbers that matter — 61 tests, 2 ignored — are unchanged, and docs/status.toml is not edited: the suite counts do not move.

Deliberately out of scope. The other eight files of #267 are untouched here, largest first next: src/models/qwen2/graph/tests.rs (3,399), src/server/batch/tests.rs (2,629), src/graph/alloc/tests.rs (1,833), src/tooling/tests.rs (1,669), src/sampler/tests.rs (1,232), src/conversation/tests.rs (1,156), src/graph/kvcache/tests.rs (1,134), and src/cuda/issue162_tests.rs (1,186) last, after #263 changes its launch-fixture column assertion.

Test-infrastructure record (#267 step 6, file 2, 2026-10-04) — models/qwen2/graph/tests.rs becomes a parent plus five topic files

What landed. src/models/qwen2/graph/tests.rs (3,399 lines, 21 #[test] — 7 of them #[ignore]d device/real-model gates — plus 6 top-level helpers) is split by topic into src/models/qwen2/graph/tests/{cuda_kv,offload_copy,kv_reuse,batching,real_model}.rs, each named by a mod declaration in the now-111-line parent (PR #277, part of #267). A pure move: every top-level fn/const item is byte-identical apart from rustfmt and the 21 test names are the same set. The topics are cuda_kv (4: the packed cache, two engines with different KV layouts, the concurrent bitwise gate, a session resumed from disk), offload_copy (3: E5 partial/auto offload and the F5 async cross copies), kv_reuse (4: cache/prefix reuse, compaction, physical kv_rm/kv_shift), batching (6: batch composition, offset sensitivity, sequence-count independence) and real_model (4: logits parity + the Metal gates). The 6 shared fixtures (cached_model_path, max_delta, cross_shape_tolerance, assert_across_shapes, argmax, compare) stay in the parent. src/models/qwen2/graph.rs keeps its #[cfg(test)] mod tests; untouched.

Verification (rule 5 numbers). dgxspark (aarch64, GB10 sm_121), 2026-10-04, in the worktree:

CommandResult
scripts/cuda_test.sh567 / 0 / 42
cargo test --release481 / 0 / 36 unit + 10 / 0 / 6 integration
python3 scripts/check_source_layout.pysrc obeys the layout rules, exit 0
python3 scripts/check_doc_line_anchors.pyexit 0 (the walkthrough anchor was OUT-OF-RANGE before the re-point)
cargo fmt --all --checkclean
CI on PR #2777 / 7 green

Mutation evidence (rule 3). mod batching; → // mod batching;: the layout checker exits 1 with src/models/qwen2/graph/tests/batching.rs: not reachable from src/main.rs …, and the CPU unit row drops 481 → 475 (the six batching tests), 36 ignored unchanged. Reverted.

Docs swept. The live models::qwen2::graph::tests::<leaf> prose gains its topic segment in AGENTS.md, docs/CUDA-BACKEND-DESIGN.md, and the #238/#244 annotations in src/graph/cuda_backend.rs, src/graph/alloc.rs and src/models/mod.rs; the walkthrough anchor in docs/inference_e2e_walkthrough/13-decode-loop-graph-reuse.md now names forward_cached_isolates_kv_between_caches at models/qwen2/graph/tests/batching.rs:742. The ARCHITECTURE-EXECUTION-PLAN.md records themselves are frozen and keep their pre-split anchors.

Deliberately out of scope. The remaining six files of #267: src/server/batch/tests.rs (2,629), src/graph/alloc/tests.rs (1,833), src/tooling/tests.rs (1,669), src/sampler/tests.rs (1,232), src/conversation/tests.rs (1,156), src/graph/kvcache/tests.rs (1,134), and src/cuda/issue162_tests.rs (1,186) last, after #263.

Test-infrastructure record (#267 step 6, file 3, 2026-10-04) — server/batch/tests.rs becomes a parent plus seven topic files

What landed. src/server/batch/tests.rs (2,630 lines, 21 #[test] — 13 of them #[ignore]d device/real-model gates — plus 19 helpers, two structs and three impl blocks) is split by topic into src/server/batch/tests/{kv_sharing,slots,prefill,batching,stall,http,metrics}.rs, each named by a mod declaration in the now-356-line parent (PR #278, part of #267). The topics are kv_sharing (4: a copied prefix, the copy-on-write store, the planned cells, the whole-arena request), slots (2: the table round trip and a resumed snapshot), prefill (5: chunk size, chunked-vs-unchunked, the interleaved ticks), batching (1: the batched-vs-serial verdict), stall (5: the injected FailingForward double and the one-answer-per-run gates), http (2: the 503 and the SSE error frame) and metrics (2: queue/running depth and the counter deltas). The parent keeps the shared fixtures (cached_model, Reply, sampling_params, run_batched, run_serial, the STEP_BUDGET_*/WorkBound cluster the helpers themselves call, and the impl BatchEngine accessors). src/server/batch.rs keeps its #[cfg(test)] mod tests; untouched.

The one non-comment edit. The file's only super::super:: reference — super::super::chat_template_from_gguf, which meant server::chat_template_from_gguf at the old module depth — becomes crate::server::chat_template_from_gguf. A topic file sits one module deeper, so the relative path would otherwise resolve to batch::; the absolute path names the same item. This is the only line of test text that differs from a pure move, and it is called out in the PR body too. (The super::serve_loop references are unaffected: the parent's use super::*; re-exports batch::serve_loop into the tests module.)

Verification (rule 5 numbers). dgxspark (aarch64, GB10 sm_121), 2026-10-04, in the worktree:

CommandResult
cargo test --release481 / 0 / 36 unit + 10 / 0 / 6 integration
scripts/cuda_test.sh567 / 0 / 42
python3 scripts/check_source_layout.pysrc obeys the layout rules, exit 0
python3 scripts/check_doc_line_anchors.py, check_docs_links.py, check_status.py --checkexit 0 each
cargo fmt --all --checkclean
CI on PR #2787 / 7 green

Mutation evidence (rule 3). mod stall; → // mod stall;: the layout checker exits 1 with src/server/batch/tests/stall.rs: not reachable from src/main.rs …, and the CPU unit row drops 481 → 477 passed and 36 → 35 ignored — the four running stall gates and the one #[ignore]d one, which is why the ignored count is the sharper half of this mutation. Reverted.

Docs swept. docs/SOURCE-LAYOUT-PLAN.md (§6.4 now carries the qwen2 and batch rows, so the three landed splits are all in the mapping table; the §Status step-6 row reads files 1–3) and this record. No live document anchors server/batch/tests.rs by line number, and the parent's own #239 annotations now name the child module (server::batch::tests::<topic>::<leaf>, five sites).

Deliberately out of scope. The remaining five files of #267: src/graph/alloc/tests.rs (1,833), src/tooling/tests.rs (1,669), src/sampler/tests.rs (1,232), src/conversation/tests.rs (1,156), src/graph/kvcache/tests.rs (1,134), and src/cuda/issue162_tests.rs (1,186) last, after #263.

Test-infrastructure record (#267 step 6, file 4, 2026-10-04) — graph/alloc/tests.rs becomes a parent plus six topic files

What landed. src/graph/alloc/tests.rs (1,833 lines, 42 #[test] — 3 of them #[ignore]d or #[cfg]-gated — plus 4 helpers and a test-only impl GraphAllocator) is split by allocator concern into src/graph/alloc/tests/{backend_fence,views,liveness,kv_arena,staging,budget}.rs, each named by a mod declaration in the now-72-line parent (PR #279, part of #267). The topics are backend_fence (4: F4's fence, the E5 offload plan and the per-engine KV format), views (8: D1 zero-copy windows, their liveness and split_parts), liveness (4: reuse along a chain, parallel chains, input fill and cycles), kv_arena (12: regions, C5 sessions, C3 defrag, C8b S3 copy-on-write and the C6/C7 cell bounds), staging (4: the F5 destination key, the pending copy and the #138 drain) and budget (10: E4/E4-S3 accounting, the length contract and rebuild re-mapping). The parent keeps chain and the impl GraphAllocator accessors; view_graph moves with views and tensor_f32 with budget, each used by that topic only. src/graph/alloc.rs keeps its #[cfg(test)] mod tests; untouched.

The non-comment edits. Ten super::super:: paths — DType, CNode, ops::NodeMeta and kvformat::KvFormat, all of which meant graph::… at the old module depth — become absolute crate::graph::… paths, because a topic file sits one module deeper and super::super would otherwise resolve to alloc::. (super::kv_defrag_enabled_from is unaffected: the parent's use super::*; re-exports it into the tests module.) These are the only test-text differences from a pure move.

Verification (rule 5 numbers). dgxspark (aarch64, GB10 sm_121), 2026-10-04, in the worktree:

CommandResult
cargo test --release481 / 0 / 36 unit + 10 / 0 / 6 integration
scripts/cuda_test.sh567 / 0 / 42
python3 scripts/check_source_layout.pysrc obeys the layout rules, exit 0
python3 scripts/check_doc_line_anchors.py, check_docs_links.py, check_status.py --checkexit 0 each
cargo fmt --all --checkclean
CI on PR #2797 / 7 green

Mutation evidence (rule 3). mod staging; → // mod staging;: the layout checker exits 1 with src/graph/alloc/tests/staging.rs: not reachable from src/main.rs …, and the CPU unit row drops 481 → 477 passed (the four staging gates), 36 ignored unchanged. Reverted.

Deliberately out of scope. The remaining four files of #267: src/tooling/tests.rs (1,669), src/sampler/tests.rs (1,232), src/conversation/tests.rs (1,156), src/graph/kvcache/tests.rs (1,134), and src/cuda/issue162_tests.rs (1,186) last, after #263.

Test-infrastructure record (#267 step 6, file 5, 2026-10-04) — tooling/tests.rs becomes a parent plus seven topic files

What landed. src/tooling/tests.rs (1,669 lines, 14 #[test] — 11 of them #[ignore]d real-model/device gates — and 14 helpers) is split by tooling concern into src/tooling/tests/{parse,f16_encode,f6_roundtrip,f141_device,f167_qwen3,quantize_bounds,bf16}.rs, each named by a mod declaration in the now-166-line parent (PR #280, part of #267). The topics are parse (2: the size parser and the split stem rule), f16_encode (1: the f16 writer's 1-D/2-D contract), f6_roundtrip (3: the llama.cpp rewrite, HF-conversion and split references), f141_device (1), f167_qwen3 (2), quantize_bounds (4: the end-to-end bounds and the byte-identical encoder) and bf16 (2). The parent keeps the shared fixtures (env_path, work_dir, cached_qwen05, PROMPT, logits_greedy, logits_greedy_on, logits_greedy_on_qwen3, assert_tensor_payloads_equal); miniature_f16_source_specs/tensor_of move with f16_encode and bf16_vs_f16_weight_value_diffs with bf16, each used by that topic only. src/tooling.rs keeps its #[cfg(test)] mod tests; untouched.

Verification (rule 5 numbers). dgxspark (aarch64, GB10 sm_121), 2026-10-04, in the worktree:

CommandResult
cargo test --release481 / 0 / 36 unit + 10 / 0 / 6 integration
scripts/cuda_test.sh567 / 0 / 42
python3 scripts/check_source_layout.pysrc obeys the layout rules, exit 0
python3 scripts/check_doc_line_anchors.py, check_docs_links.py, check_status.py --checkexit 0 each
cargo fmt --all --checkclean
CI on PR #2807 / 7 green

Mutation evidence (rule 3). mod parse; → // mod parse;: the layout checker exits 1 with src/tooling/tests/parse.rs: not reachable from src/main.rs …, and the CPU unit row drops 481 → 479 passed (the two parse gates), 36 ignored unchanged. (A topic whose gates are all #[ignore]d would move only the ignored count; parse was chosen because it moves the running half.) Reverted.

Deliberately out of scope. The remaining three files of #267 under 1,300 lines — src/sampler/tests.rs (1,232), src/conversation/tests.rs (1,156), src/graph/kvcache/tests.rs (1,134) — and src/cuda/issue162_tests.rs (1,186) last, after #263.

Test-infrastructure record (#267 step 6, files 6–8, 2026-10-04) — the four sub-1,300-line files land in one PR, a commit apiece

What landed. The three remaining files under the ticket's 1,300-line grouping threshold are split in one PR (#281, part of #267), one commit per file, each keeping its shared fixtures in tests.rs and declaring one mod <topic>; per topic file:

filebeforeaftertests
src/sampler/tests.rs1,232a 140-line parent + {greedy_topk,penalties,stops,minp_typical,xtc,dry,mirostat,bias_validate,defaults,grammar}.rs47
src/conversation/tests.rs1,156a 209-line parent + {turns,regen,spec,snapshot,overflow,real_model}.rs (the mock engine and 12 fixtures stay up)27
src/graph/kvcache/tests.rs1,134a 122-line parent + {cells,spans,sharing,resize,defrag}.rs33

All three are pure moves: every item is byte-identical apart from rustfmt and the test-name sets are unchanged. None of the three files contained a super::super:: reference, so no test-body path needed the absolute rewrite that #278 and #279 carried. The two live walkthrough anchors for conversation/tests.rs are re-pointed at conversation/tests/turns.rs:48 and conversation/tests/overflow.rs:22 in the same commit.

Verification (rule 5 numbers). dgxspark (aarch64, GB10 sm_121), 2026-10-04, in the worktree:

CommandResult
cargo test --release481 / 0 / 36 unit + 10 / 0 / 6 integration
scripts/cuda_test.sh567 / 0 / 42
python3 scripts/check_source_layout.pysrc obeys the layout rules, exit 0
python3 scripts/check_doc_line_anchors.py, check_docs_links.py, check_status.py --checkexit 0 each
cargo fmt --all --checkclean
CI on PR #2817 / 7 green

Mutation evidence (rule 3). One mutation per file, each reverted — the checker names the orphan and the CPU row drops: mod dry; in sampler/tests.rs → src/sampler/tests/dry.rs and 481 → 476; mod overflow; in conversation/tests.rs → src/conversation/tests/overflow.rs and 481 → 475; mod spans; in graph/kvcache/tests.rs → src/graph/kvcache/tests/spans.rs and 481 → 475.

What is left. src/cuda/issue162_tests.rs (1,186 lines) is the last file of #267 and is deliberately not split here: it carries the launch-fixture column assertion (assert_eq!(f.len(), 4)) that Step 2 (#263) changes, so it moves after #263 merges.

Test-infrastructure record (#267 step 6, file 9, 2026-10-04) — cuda/issue162_tests.rs becomes a parent plus four topic files

What landed. src/cuda/issue162_tests.rs (1,186 lines, 5 #[test] — all device + MINFER_TEST_ISSUE162=1 gated, none #[ignore]d — plus 18 non-test top-level items: 3 use, 5 consts, 2 structs, 3 impl blocks and 5 helper fns) is split by topic into src/cuda/issue162_tests/{sites,severity,control,node}.rs, each named by a mod declaration in the now-209-line parent (PR #283, part of #267). The topics are sites (1: the union driver that arms every audited <<< site and asserts armed set == observed set == the fixture), severity (1: a required site sets the sticky, a documented-fallback _opt site does not), control (1: the knob-off positive control that really computes 1 + 2) and node (2: a required launch failure fails a real Op::Add node with an Err naming launch:add_f32, and the f16-matmul Err arm still drains the sticky). Every use, const, fixture and helper stays in the parent — device, cstr, gate_enabled, fixture, Arm, Ctx and the run driver — so the #239 annotations in src/cuda/tests.rs, which name cuda::issue162_tests::run, keep resolving unchanged. A pure move: every top-level item is byte-identical apart from rustfmt and the 5 test names are the same set. src/cuda.rs keeps its #[cfg(test)] mod issue162_tests; untouched.

Verification (rule 5 numbers). dgxspark (aarch64, GB10 sm_121), 2026-10-04, in the worktree:

CommandResult
scripts/cuda_test.sh567 / 0 / 42
MINFER_TEST_ISSUE162=1 scripts/cuda_test.sh issue1625 / 0 before and after the move
cargo test --release481 / 0 / 36 unit + 10 / 0 / 6 integration
python3 scripts/check_source_layout.pysrc obeys the layout rules, exit 0
check_doc_line_anchors.py, check_docs_links.py, check_status.py --check, check_dead_code_annotations.pyexit 0 each
cargo fmt --all --checkclean
CI on PR #2837 / 7 green

Mutation evidence (rule 3). mod node; → // mod node;: the layout checker exits 1 with src/cuda/issue162_tests/node.rs: not reachable from src/main.rs — no \mod` declaration names it, so it is never compiled (tests in it would silently not run), the CUDA unit row drops **567 → 565** (the two nodegates) and this file's ownMINFER_TEST_ISSUE162=1 … issue162` gate drops 5 → 3. Reverted.

What the brief had wrong (the tree wins). The step-6 note deferred this file until after #263, because its launch-fixture assert_eq!(f.len(), 4) was expected to change. That change was measured unnecessary — the fixture keeps its four columns, its identity key being the ordered (owner, site, kernel-fragment) triple — so #263 touches neither this file nor the fixture and the split proceeds. No anchor re-point was needed either: assert_eq!(f.len(), 4) still lands on src/cuda/issue162_tests.rs:56 in the parent, so the live src/cuda/issue162_tests.rs:55 anchor in docs/SOURCE-LAYOUT-PLAN.md §4 Step 2 resolves unchanged. All nine files of #267 are now split.

Docs swept. docs/SOURCE-LAYOUT-PLAN.md (§6.4 gains the issue162_tests mapping row; the §Status step-6 row reads files 1–9) and docs/CUDA-BACKEND-DESIGN.md (the #162 runtime-gate paragraph names the four topic files). The ARCHITECTURE-EXECUTION-PLAN.md anchor into this file (line 6697) is inside this frozen record and keeps its pre-split text.

Test-infrastructure record (#261 step 1, 2026-10-04) — src/cuda.rs becomes src/cuda/{ffi_runtime,methods}.rs + 18 family files

What landed. The first code move of the source-layout campaign (#262, PR #276): src/cuda.rs loses its 4,546-line impl CudaState and its two extern "C" blocks and keeps the module doc, AttnWindow, pub struct CudaState (all 29 fields still private), the free items, the ten #[cfg(test)] mod declarations and the re-export lines — 6 594 → 1,014 lines. New: src/cuda/ffi_runtime.rs (the cudart/driver declarations, now pub(super), plus the test-only extern block), src/cuda/methods.rs (the 15 non-pub helpers whose callers land in a second family file, the 18 mod declarations and the two-level #[cfg(test)] pub(crate) use chain the launch tests resolve through use super::*), and the 18 src/cuda/methods/<family>.rs files (28–728 lines) holding the impl CudaState blocks and the 86 launch declarations (now pub(crate)), each with the family that uses it — no symbol has two consumer families.

The layout invariant. methods/<family>.rs are descendants of methods.rs, itself a child of cuda.rs, so a private field of CudaState and a private method in methods.rs are visible in every family file: 0 field-visibility edits and 0 pub(super). The 15 helpers that must live in the shared parent are the transitive closure of the cross-file call graph, not the family table's first guess: plane_budget_ok stays in methods/weights.rs (all four callers are there), while the eight no_*/fused_b_on predicates are cross-family and cannot sit in a sibling policy.rs without pub(super) — that is the one place the executed shape differs from the plan's §4 Step 1 bullet.

Verification (rule 5 numbers). dgxspark (aarch64, GB10 sm_121), 2026-10-04, in the step's worktree:

CommandResult
cargo build --release --features cudaexit 0
scripts/cuda_test.sh567 passed / 0 failed / 42 ignored — unchanged from 2f70fc6
cargo test --release481 / 0 / 36 unit + 10 / 0 / 6 integration — unchanged
cargo fmt --all --checkexit 0
python3 scripts/check_source_layout.py"src obeys the layout rules", exit 0
python3 scripts/check_dead_code_annotations.pyexit 0 (11 grandfathered bare sites)
python3 scripts/check_dead_code_oracle.py --config {cpu,cuda}0 additions, 0 removals; only the two moved file = lines in docs/dead-code-baseline.toml
python3 scripts/check_doc_line_anchors.py81 out-of-range → 0 (99 anchors re-pointed by the split's line map, 28 re-anchored to the symbol, 5 rebuilt by hand)
FEATURES=cuda scripts/real_model_gates.sh ×242 / 0 ×2, greedy output bitwise identical
CI run 372011328547 / 7 green, zero code annotations

Mutation evidence (rule 3). Two mutations in the worktree, each reverted: src/cuda/methods/orphan.rs → check_source_layout prints not reachable from src/main.rs — no mod declaration names it, so it is never compiled (tests in it would silently not run), exit 1; deleting mod buffers; from methods.rs → the same message for src/cuda/methods/buffers.rs, exit 1 (restored, exit 0).

Deliberately out of scope. No line of CUDA behaviour changed: the impl bodies, free items and declarations move verbatim, and the only textual edits are the visibility prefixes on the moved declarations, the generated use/mod preambles and the CudaCommandBuffer banner (rewritten to name graph/cuda_backend.rs, the deferred #219 item this step owned). The .cu file is untouched (Step 2, #263).

Test-infrastructure record (#264 stage A, 2026-10-04) — src/quants.rs becomes src/quants/*.rs

What landed. Stage A of the CPU split (#264, PR #290): src/quants.rs 1,340 → 61 lines (the module decider: mod declarations + the pub use list + the cross-module wiring) and nine part files — dot_q4_0.rs 71 · dot_q4_1.rs 44 · dot_q5.rs 94 · dot_q8_0.rs 64 · kquant.rs 185 · quantize_q8_0.rs 125 · quantize_q8_k.rs 114 · avx2.rs 81 · neon.rs 407 — plus the extracted src/quants/neon_correctness.rs 137.

The layout invariant. neon_kernels and neon_q8k are flattened into the one neon.rs: their item names do not collide (enabled/fp16/dot16/sdot_vec + the five dot_* kernels vs the three K-quant dots), and neon_q8k's super::neon_kernels:: prefixes become the same module's items. The move is otherwise verbatim: an item a sibling or the parent reaches becomes pub(super) — the same reachable set it had as a private item of quants, since every part file is exactly one level below — and the parent's pub use list keeps crate::quants::… resolving, which is what graph/kvformat.rs (two calls) and kernel.rs (the dot dispatch) compile against.

The rule widening (#274). Extracting the inline #[cfg(all(test, target_arch = "aarch64"))] mod neon_correctness removes the tree's only compound-cfg inline test module, and the same PR widens scripts/check_source_layout.py rule 1 from the literal #[cfg(test)] to the cfg predicate: all(test, …)/any(test, …) with or without further attributes are reported, not(test) is not (that module is the non-test build). Five selftest cases were added (12 total). Measured: the baseline tree's compound-cfg module exits 0 under master's checker and 1 under the widened one; after the extraction the widened checker exits 0 on the tree.

Verification (rule 5 numbers). dgxspark (aarch64, GB10 sm_121), 2026-10-04, in the step's worktree:

CommandResult
cargo test --release481 / 0 / 36 unit + 10 / 0 / 6 integration — identical to docs/status.toml
MINFER_NO_NEON=1 cargo test --release481 / 0 / 36 + 10 / 0 / 6 — the scalar arm, identical
cargo check --release --target x86_64-unknown-linux-gnuexit 0 (the AVX2/f16c bodies this aarch64 box never parses)
cargo fmt --all --checkexit 0
python3 scripts/check_source_layout.py (--selftest)"src obeys the layout rules", exit 0; 12 cases pass
python3 scripts/check_dead_code_annotations.pyexit 0 (11 grandfathered bare sites)
python3 scripts/check_dead_code_oracle.py --config cpu46 baseline entries, 0 additions, 0 removals (no file = field moved: the baseline names items only)
python3 scripts/check_doc_line_anchors.py17 out-of-range → 0; the 19 quants.rs:NNN anchors in live docs → 0
CI run 372148018447 / 7 green, zero code annotations

Mutation evidence (rule 3). Renaming src/quants/kquant.rs to a name no mod declares → check_source_layout exits 1 with not reachable from src/main.rs — no \mod` declaration names it, so it is never compiled (tests in it would silently not run). Deleting pub use dot_q4_0::dot_q4_0_q8_0;→cargo check --releaseexits 101 witherror[E0425]: cannot find function `dot_q4_0_q8_0` in module `crate::quants`atsrc/kernel.rs:148:63` (the re-export, not just the move, is load-bearing). Both reverted, exit 0.

What the brief had wrong (the tree wins). #264's body states the acceptance as CPU 480 / 0 / 36; docs/status.toml carries 481 (the two tests #138 added after that body was written) and the tree's rows are the ones this record uses. The plan's §4 Step 3 paragraph and its §10 row carried the same stale 480 and were corrected here.

Deliberately out of scope. No value, order or algorithm changed; the only textual edits are visibility prefixes, use/mod preambles, the flattened-module renames and the doc sweep. Stages B (src/vec_ops.rs) and C (src/kernel.rs) follow, one PR each; docs/CPU_OPTIMIZATIONS.md is frozen and keeps its quants.rs:NNN numbers.

Test-infrastructure record (#264 stage B, 2026-10-04) — src/vec_ops.rs becomes src/vec_ops/*.rs

What landed. Stage B of the CPU split (#264, PR #291): src/vec_ops.rs 1,344 → 52 lines (the module decider) and eight part files — vec.rs 394 · rms_norm.rs 154 · rope.rs 17 · softmax.rs 91 · silu.rs 102 · f16.rs 318 · bf16.rs 95 · neon.rs 165. The old file had 47 fn definitions and the new files have the same 47 (excluding the pre-existing tests.rs).

The layout invariant. mod neon_vec is promoted to the file neon.rs: its items keep pub(super), which is pub(in vec_ops) on both sides of the move, and vec_soft_max_inplace_f32's call becomes super::neon::vec_soft_max_f32_inplace. mod neon_f16 stays nested inside f16.rs because its super::F16_SIMD_PATH_CALLS is one module up either way. vec_exp_f32_avx2 and decode_{f16,bf16}_row widen to pub(super) (a sibling or the parent's wiring reaches them); the parent's pub use list keeps every crate::vec_ops::… path resolving, which is what graph/cpu_backend.rs (13 calls), models/*/loader.rs and graph/cuda_backend.rs's RopeStyle import compile against.

One deliberate API narrowing. dot_f16_f32, dot_f16_f32_scalar, f16_dot_path and F16DotPath have exactly one consumer, vec_ops::tests; a non-test pub use of an unused name is an unused_imports error under #![deny(warnings)], so their re-export is #[cfg(test)]. The items themselves stay pub in f16.rs. silu.rs imports vec_exp_f32_avx2 under #[cfg(target_arch = "x86_64")] for the same reason.

The baseline moved with the code. RopeStyle::Interleaved is the one docs/dead-code-baseline.toml entry naming this file; its two file = fields (the cpu and cuda arms) now read src/vec_ops/rope.rs. The oracle still reports 0 additions and 0 removals on the cpu arm.

Verification (rule 5 numbers). dgxspark (aarch64, GB10 sm_121), 2026-10-04, in the step's worktree:

CommandResult
cargo test --release481 / 0 / 36 unit + 10 / 0 / 6 integration — identical to docs/status.toml
MINFER_NO_NEON=1 cargo test --release481 / 0 / 36 + 10 / 0 / 6 — the scalar arm, identical
cargo check --release --target x86_64-unknown-linux-gnuexit 0 (the AVX2/f16c bodies and F16DotPath::Avx2)
cargo fmt --all --checkexit 0
python3 scripts/check_source_layout.py (--selftest)exit 0; 12 cases pass
python3 scripts/check_dead_code_annotations.pyexit 0 (11 grandfathered bare sites)
python3 scripts/check_dead_code_oracle.py --config cpu46 baseline entries, 0 additions, 0 removals
python3 scripts/check_doc_line_anchors.pythe 10 live vec_ops.rs:NNN anchors re-pointed, 0 out-of-range
CI run 372158108347 / 7 green, zero code annotations

Mutation evidence (rule 3). Renaming src/vec_ops/rms_norm.rs to a name no mod declares → check_source_layout exits 1 with not reachable from src/main.rs — no mod declaration names it, so it is never compiled (tests in it would silently not run). Deleting pub use rms_norm::{rms_norm_f32, rms_norm_fused_f32}; → cargo check --release exits 101 with error[E0425]: cannot find function rms_norm_fused_f32in modulecrate::vec_ops`` at src/graph/cpu_backend.rs:445:45 and the note that the function exists but is inaccessible. Both reverted, exit 0.

Deliberately out of scope. No value, order or algorithm changed. Stage C (src/kernel.rs) follows. docs/CPU_OPTIMIZATIONS.md is frozen and keeps its vec_ops.rs:NNN numbers.

Test-infrastructure record (#264 stage C, 2026-10-04) — src/kernel.rs becomes src/kernel/*.rs

What landed. Stage C closes the CPU trio (#264, PR #292): src/kernel.rs 667 → 22 lines (the module decider) and three part files — dispatch.rs 72 · pool.rs 313 · embed.rs 278. The old file had 10 fn definitions and the new files have the same 10 (excluding the pre-existing tests.rs).

The layout invariant. dispatch.rs submits through the pool, so MmJob (and its seven fields), ParForJob, PoolJob, Pool (and its five fields), mm_rows, chunk and get_pool become pub(super) — the same pub(in kernel) reach they had as private items of kernel — and dispatch.rs imports them explicitly (use super::pool::{chunk, get_pool, mm_rows, MmJob, PoolJob};) rather than through the parent, which keeps the dependency visible. Pool's gate field carries its hazard comment (the concurrent-submission use-after-free the lock prevents) into pool.rs verbatim. pool.rs also keeps the thread-locals and the use std::sync::… lines the one-file version had mid-file.

One re-export deliberately not carried. cpu_quant_matmul (the byte-in/byte-out worker entry) stays pub in dispatch.rs but is not re-exported at the kernel root: its one caller is cpu_quant_matmul_f32 in the same file, no crate::kernel::cpu_quant_matmul path is named anywhere in the tree, and a pub use of an unused name is an unused_imports error under #![deny(warnings)]. Every path that exists is preserved.

Verification (rule 5 numbers). dgxspark (aarch64, GB10 sm_121), 2026-10-04, in the step's worktree:

CommandResult
cargo test --release481 / 0 / 36 unit + 10 / 0 / 6 integration — identical to docs/status.toml
MINFER_NO_NEON=1 cargo test --release481 / 0 / 36 + 10 / 0 / 6 — the scalar arm, identical
cargo check --release --target x86_64-unknown-linux-gnuexit 0
cargo fmt --all --checkexit 0
python3 scripts/check_source_layout.py (--selftest)exit 0; 12 cases pass
python3 scripts/check_dead_code_annotations.pyexit 0 (11 grandfathered bare sites)
python3 scripts/check_dead_code_oracle.py --config cpu46 baseline entries, 0 additions, 0 removals
python3 scripts/check_doc_line_anchors.pythe 14 live kernel.rs:NNN anchors re-pointed, 0 out-of-range

Mutation evidence (rule 3). Renaming src/kernel/embed.rs to a name no mod declares → check_source_layout exits 1 with not reachable from src/main.rs — no mod declaration names it, so it is never compiled (tests in it would silently not run). Deleting pub use pool::{cpu_threads, par_for, set_cpu_threads}; → cargo check --release exits 101 with error[E0425]: cannot find function set_cpu_threadsin modulecrate::kernel`` at src/bench.rs:138:36 and the note function crate::kernel::pool::set_cpu_threads exists but is inaccessible. Both reverted, exit 0.

#264 is complete. src/quants.rs 1,340 → 61 + 9 parts (+ the extracted neon_correctness.rs 138), src/vec_ops.rs 1,344 → 52 + 8 parts, src/kernel.rs 667 → 22 + 3 parts, and the three module files are the deciders with the same public paths. The ISA axis is now the directory tree.

Test-infrastructure record (#261 step 2, 2026-10-04) — src/cuda_kernels.cu becomes src/cuda/kernels/ (backfilled 2026-10-05)

Why this entry is retroactive. Steps 0 (1c9e68d), 1 (PR #276), 3 (stages A–C) and 6 (files 1–9) each appended their dated record here; Step 2's lived only in docs/SOURCE-LAYOUT-PLAN.md §4 Step 2. Step 5 adds the missing entry so the campaign's own rule (§6.3: "every step: a dated entry in §test-infrastructure") holds for all seven, and so the module-count measurement has one home. The numbers below are the stage record's (plan §4 Step 2, one PR per group), not re-measured here.

What landed. Six PRs on dgxspark (aarch64, GB10 sm_121): G1 930e1e2 (#284) common.cuh 498 + guard.cu 305 · G2 286a5a4 (#285) attention_decode.cu 1,177 + attention_prefill.cu 511 · G3 f68a542 (#286) mmq_int8.cu 457 + mmq_raw.cu 645 + mmq_nb.cu 708 + mmq_bt_q6k.cu 481 · G4 c0ddf8b (#287) matmul_f32act.cu 719 + mmvq_aquant.cu 413 + mmvq_skipwrite.cu 648 + mmvq_q6k.cu 342 + mmvq_multi.cu 753 · G5 ee380ad (#288) ops_misc.cu 743 + ops_elementwise.cu 439 + kv_store.cu 435 + gemm_wmma.cu 936 + gemm_fused_dequant.cu 284 · G6 0cb21cb (#289) the 14-line remainder deleted. src/cuda_kernels.cu 10,215 → 0; src/cuda/kernels/ = common.cuh + 17 TUs (9,996 lines).

The layout invariant. Route (a) — the launcher lives in the TU that instantiates its kernel — is forced by the default toolchain: a templated __global__ launched or address-taken across TUs is a link error without -rdc=true (the four-pattern two-file probe in plan §5), and the plan deliberately does not add -rdc. Two launchers were the exception and were resolved by merging their kernel's home (attention_hybrid.cu into attention_decode.cu, the q6_K reducer next to the NB kernels). build.rs owns one explicit KERNEL_SOURCES/KERNEL_HEADERS list — the compile order, the launch audit's --source order and the fixture's row order — plus a new-file guard (a .cu in kernels/ absent from the list fails the build) and one rerun-if-changed per file. The five measured corrections the plan's draft table needed (guard.cu 305 not 402; five per-family pre-warm entries, not six; gemm_smem.cu merged into gemm_wmma.cu; the two-caller mmq_ksplit_reduce_kernel; 17 TUs, not 19) are recorded in plan §4 Step 2.

Verification (rule 5 numbers, the stage record's). Same at every stage: check_cuda_launch_returns.py 130 sites + --check-fixture exit 0; MINFER_TEST_ISSUE162=1 device gate green; CUDA unit 567 / 0 / 42; CPU 481 / 0 / 36 + integration 10 / 0 / 6; real-model 42 / 0 on both the 0.5B and the Qwen3-0.6B configuration, greedy output bitwise identical; 0 nvcc warnings. The module-load cost this step was predicted to change is re-measured in the Step 5 record below.

Deliberately out of scope. No -rdc=true, no -static-global-template-stub=false, no behaviour change; the pre-warm decomposition keeps the minfer_prewarm_kernels symbol the Rust side declares; and the cold/warm module-load comparison was left to Step 5 because it needs a quiet device and a fresh process per configuration.

Test-infrastructure record (#261 step 5, 2026-10-05) — the close-out: the module-load re-measure, the .cu prose sweep and the last two #219 claims

What landed. The Linux half of Step 5 on dgxspark (aarch64, GB10 sm_121), CUDA 13.0, driver 580.178.04, one GPG-signed branch, PR #293. Three concerns:

  1. The measurement the split made necessary. The fatbin is 17 modules since #263, so the recorded "~2.2 ms per module / ~14.5 ms cold" row (docs/CUDA-BACKEND-DESIGN.md §2.4, [#225]) was re-measured instead of extrapolated. The pre-split binary was rebuilt at 930e1e2^ (bc30152) in a second worktree and both were driven with the command of record MINFER_OP_TIMING=1 target/release/minfer <cached 0.5B q4_0> "hello", warm (fresh processes, back-to-back) and cold (the binary's and the model's pages evicted with posix_fadvise(POSIX_FADV_DONTNEED) first — an agent shell cannot drop_caches). nvidia-smi before the warm runs: SM clock 2,411 MHz, 0 % util, no other compute process.
  2. The prose sweep. Every live document that named the retired src/cuda_kernels.cu now names src/cuda/kernels/*.cu (or the specific TU): CUDA-BACKEND-DESIGN.md (the §2.1 layer table and the four current-state references), BUILD.md, the walkthrough's 15-cuda-backend.md and README.md, GPU_SAFETY.md, CUDA_OPTIMIZATION.md, LLAMA-CPP-MMQ-ANALYSIS.md (the r25-HEAD citation kept, the current path added), cuda_tutorial/{01,02,03,04,05,06,07,README,STYLE}.md, CUDA-TECH-PRIMER.md, DEVICE-ADAPTATION-PLAN.md, COMPUTE-GRAPH-DESIGN.md, ARCHITECTURE-ROADMAP.md, GATE-CONTRACT.md, README.md, AGENTS.md, .github/workflows/ci.yml, build.rs and the src/** comments. Two dead fallbacks for the deleted file were removed with it (build.rs::LEGACY_SOURCE and its kernel_translation_units()/rerun-if-changed arms; check_cuda_launch_returns.py::LEGACY_SOURCE and the docstring/--source text describing the "legacy remainder"). Frozen records (cuda_optimization_steps/*, QWEN2.5-*, DEBUGGING-*, KNOWN-CPU-ISSUES-*, CPU_OPTIMIZATIONS.md, PARAMETER_AUDIT.md, this file, experiments/cuda/*) keep their pre-split citations by policy (plan §6.2, where experiments/cuda/* is now named).
  3. The claims the earlier steps missed. docs/ARCHITECTURE.md still named src/cuda/impl/<family>.rs and src/cuda/{init,…}.rs (the inner module is methods; the files are under src/cuda/methods/) and docs/BACKENDS.md repeated the impl path; inference_e2e_walkthrough/15-cuda-backend.md carried the pre-#262/#267 sizes (src/cuda.rs 6,578 → 1,014, cuda_backend.rs + its tests re-counted, the .cu line count corrected to 9,996); and 17 already-wrong src/cuda.rs:NNN anchors in cuda_tutorial/{02,03}, ARCHITECTURE-ROADMAP.md and walkthrough/03 were re-pointed at the symbol or the owning file — the step-1 line map proved they named the extern-"C" declaration block on the old file too, so they were pre-existing wrong citations of the same class as #219, invisible to the checker because the numbers landed inside the 1,014-line remainder.

The measured row (the point of the split). Command as above; the MINFER_OP_TIMING=1 line is the prefill-GEMM smem pre-warm loop, i.e. the point at which the fatbin's module is finalized.

buildmodules the command forceswarm, fresh processcold (page-cache-evicted)
pre-split bc301521 (whole fatbin)2,300 µs median (2,203–2,468, n=12)18,478 / 20,723 µs
post-split shipped (4900298)1 of 17 (gemm_wmma.cu)350–590 µs4,288 / 6,109 / 6,983 µs
post-split, all 16 loadable TUs16 of 17 (temporary per-TU probe, reverted)≈2,370 µs total42,671 / 44,416 / 47,412 µs

Per-module warm (temporary probe inside minfer_prewarm_kernels, plus the try_new GEMM module): gemm_fused_dequant 40 · mmvq_q6k 57 · mmvq_aquant 57 · mmq_raw 70 · mmq_nb 73 · matmul_f32act 85 · mmq_int8 93 · mmvq_skipwrite 100 · kv_store 102 · mmvq_multi 111 · mmq_bt_q6k 118 · ops_misc 123 · attention_decode 223 · ops_elementwise 245 · attention_prefill 447 · gemm_wmma ≈350–590 µs — i.e. 40–450 µs (median ≈150 µs) per module, not the recorded 2.2 ms, which was the whole pre-split module. Cold, the same 16 registrations are 975–5,214 µs each: the extra is a per-registration fault of the fatbin's pages.

Verdict recorded in the docs. Acceptable, with the honest extra named: warm steady state is unchanged (2,300 → ≈2,370 µs, the shipped line reads lower only because it times 1/17 of the work), and a page-cache-cold start pays +25 ms once per process (≈19 → ≈45 ms, ≈1.8 % of this model's ~1.4 s cold start). The plan's named mitigation (minfer_prewarm_kernels trimmed to the modules a run needs) is deliberately not applied: it would move the cost into the first forward's lazy loads, not remove it. The old "cold / idle-clock" explanation of the 14.5 ms row is corrected — it is page-cache-cold, not the GPU clock (40 s of cooldown at a 208 MHz SM clock reads the warm figure; evicted pages reproduce the cold one). See docs/CUDA-BACKEND-DESIGN.md §2.4 and docs/SOURCE-LAYOUT-PLAN.md §5.1.

Verification (rule 5 numbers). 2026-10-05, in the step's worktree:

CommandResult
scripts/cuda_test.sh567 passed / 0 failed / 42 ignored — unchanged from docs/status.toml
cargo test --release481 / 0 / 36 unit + 10 / 0 / 6 integration (backend_registry_cli 7/0, conversation_cli 3/0/6) — unchanged
FEATURES=cuda scripts/real_model_gates.sh42 / 0 on the 0.5B and 42 / 0 with MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf
python3 scripts/check_dead_code_oracle.py --config {cpu,cuda}46 / 34 baseline entries — 0 additions, 0 removals
python3 scripts/check_source_layout.py / check_dead_code_annotations.pyexit 0 ("src obeys the layout rules"; 11 grandfathered bare sites)
python3 scripts/check_docs_links.py / check_status.py --check990 links resolve; prose agrees with docs/status.toml
python3 scripts/check_doc_line_anchors.py1 664 anchors · 0 bad, 3 unchanged UNUSED FREEZE notices, no stale exemption
python3 scripts/check_cuda_launch_returns.py --selftest / --check-fixture5 cases pass; 130 <<< sites
cargo fmt --all --check / cargo build --release --features cudaexit 0

Mutation evidence (rule 3). Two mutations, each reverted: appending `src/cuda.rs:99999` to docs/GPU_SAFETY.md → check_doc_line_anchors.py exits 1 with OUT-OF-RANGE docs/GPU_SAFETY.md:223: src/cuda.rs:99999 → src/cuda.rs — src/cuda.rs has 1014 lines; adding an unlisted src/cuda/kernels/zz_probe_mutation.cu → the build script panics at build.rs:59 with assertion left == right failed: build.rs KERNEL_SOURCES must list every .cu in src/cuda/kernels, in sorted order (an unlisted kernel file is never compiled) and exit 101 — the guard still bites after the LEGACY_SOURCE fallback was deleted. Both reverted; the gates re-run green.

#219 closed. The issue's two original claims and the four folded ones are all resolved: the walkthrough's register_weight section (and the same stale "blocking cudaMemcpy" sentence in cuda_tutorial/02) now documents the #188 stream-ordered cudaMemcpyAsync + cudaStreamSynchronize and the test-only register_weight_blocking_legacy probe; the CudaCommandBuffer banner was rewritten by Step 1 (src/cuda/methods/dispatch.rs:124); the ~4400 LOC, 120 sites and walkthrough-size figures were corrected by Step 0 and re-corrected here where #262/#267 invalidated them again.

Deliberately out of scope (the Mac round). #265 (Step 4: src/metal.rs/src/metal.metal → src/metal/{runtime,encode,ops,policy}.rs + src/metal/kernels/*.metal) and #53's Metal DeviceMemory answer — the second implementation the common rule waits on. src/main.rs gates the whole Metal module on target_os, so no Linux build or CI job can compile it. The common decision is therefore not taken: plan §1.3 records the pre-analysis (the type and the policy are already in the device-agnostic allocplan/offload; the device answer has one implementation, CudaState::device_memory() through the CUDA-only models::device_memory(), with no Backend trait hook), so creating a common module now would be the single-real-implementation case the rule forbids.

(Mac half, 2026-10-05: #53 landed Metal's DeviceMemory answer, so the decision this paragraph deferred is now taken — no new module; see plan §1.3 rule 3 and the E4/E5 Metal-half record above.)

Test-infrastructure record (#255, 2026-10-05) — the two macOS-only allow(dead_code) sites are judged, and the macOS test build is unblocked

The ticket. #255 is the Mac hand-off of the #238–#244 dead-code series: the two bare allow(dead_code) sites in src/metal.rs (macOS-only, so no Linux capture can compile them). It is [#260](https://github.com/yusiwen/minfer/issues/260)'s first item and a prerequisite of #265 (the Metal split).

Box. macbook (macOS 27.0.1, Apple M4 Pro), 2026-10-05, toolchain pinned at 1.97.1 (RUSTUP_TOOLCHAIN is stable in this shell, so it is unset for every run). Worktree .worktrees/255, branch 255-metal-dead-code.

Oracle method. allow(dead_code) is stripped line-preservingly across src/**.rs (43 attribute lines on a7fc07e; the issue's 61 is the f47e4c3 count — master has since lost 18). Then RUSTFLAGS=--cap-lints=warn cargo check --release --message-format=json is captured on the Mac; every src/ span of every dead_code diagnostic is read (60 spans). This is the first capture in the series that compiles src/metal.rs at all.

Verdict 1 — src/metal.rs:972 MpsCommandBuffer::matmul_on_gpu_buf → D3, #[cfg(test)]. The stripped macOS non-test build reports method 'matmul_on_gpu_buf' is never used; its only readers are src/metal/tests.rs:147/201/212. It stays in place under #[cfg(test)] (the T1 shape), because it is an inherent method on a lifetime-parameterised type whose body is a one-line forward to the live quant_matmul_f32_on_gpu_buf.

Verdict 2 — src/metal.rs:2029, the container-level blanket on impl MpsState → D1, delete. The stripped macOS non-test build reports no dead member of the impl: try_new/init (via main.rs), register_part/register_weight (both loaders), has_weight (both graph builders), cmd_buffer, weight_buf, new_f32_buffer (graph/metal_backend.rs), get/get_or_grow all have non-test readers. The comment's "layer-gpu reference path" claim was false since #240/#241 deleted that path, so the blanket and the comment both go.

Prerequisite found and fixed: the macOS test build did not compile. The ticketed acceptance asks for a warning-free cargo check --release --tests, but on clean master that command exits 101 with 16× E0599: no associated function or constant named 'Metal' found for struct 'graph::registry::Backend' — src/graph/metal_backend/tests.rs still spells the pre-F4 enum variant Tag::Metal where the F4 registry handle is Backend::METAL. CI's build-macos ran cargo build only, never --tests, so the breakage was invisible to CI and predates this ticket — #303 later made that job run cargo test --release --no-run. The byte-identical clean capture (clean_tests.log) proves it is not caused by the strip. The 16 sites are renamed and the test build then exits 0.

Step 6 of the issue, confirmed. src/models/mod.rs's #[cfg_attr(any(not(test), not(target_os = "macos")), allow(dead_code))] on as_any/forward_graph: the macOS test build emits no dead_code warning for either, so the annotation applies no allow there and the items are live exactly where the caller (models::qwen2::graph::tests) is compiled. The annotation is correct as written and is left alone.

Manifests. scripts/check_dead_code_annotations.py's GRANDFATHERED_BARE drops both src/metal.rs: keys (11 → 9 sites) and its [#255] comment is rewritten; the ratchet only turns one way. docs/dead-code-baseline.toml needs no row: the issue predicted a file = "src/metal.rs" entry, but the manifest only carries [[cpu]]/[[cuda]] rows and metal is not compiled on Linux, so the stripped Linux oracle cannot see either site — the tree is the authority here, not the issue text.

Verification (rule 5 numbers), macbook (macOS 27.0.1, Apple M4 Pro), 2026-10-05.

CommandResult
cargo build --release (unstripped, after the change)exit 0 — the deny(warnings) non-test gate is clean; target/release/build/minfer-*/out/minfer.metallib = 410 942 B, sha256 13af518ed447d71f10439a894380eb1bddcdb6e2d737a396e99cd6a823ef3523
RUSTFLAGS=--cap-lints=warn cargo check --release (stripped)exit 0; 60 dead_code src/ spans — matmul_on_gpu_buf present, zero impl MpsState members
RUSTFLAGS=--cap-lints=warn cargo check --release --tests (clean master)exit 101, 16× E0599 Tag::Metal
same, after the renameexit 0
cargo test --release483 passed / 20 failed / 38 ignored — see the note below
scripts/real_model_gates.sh (0.5B) and MINFER_BATCH_TEST_MODEL=…/Qwen3-0.6B-Q8_0.gguf37 / 1 each; the single failure is server::batch::tests::slots::a_slot_snapshot_resumes_the_context_without_re_prefilling ("KV session: layer 0 host read failed"), a pre-existing Metal KV-session gap
python3 scripts/check_dead_code_annotations.pyexit 0, 9 grandfathered bare sites
cargo fmt --all --checkexit 0

The macOS suite is red, and this record is the first honest measurement of it. With the compile fix in place the macOS-only tests run for the first time; 20 fail, and their causes are pre-existing, not this change: five graph::metal_backend::tests attention-input sites panic with attn: query 0 has span [0, 0) (the #228 migration that #231 says "has never been executed"), and the models::qwen2/qwen3 Metal comparisons diverge (graph_metal_matches_cpu_logits: CPU greedy token 220 vs GPU 353). They belong to #231/#260, not here, and this ticket adds no fix for them. #265 carries the same failing set before and after, which is what a pure move must show.

Test-infrastructure record (#303, 2026-10-06) — build-macos compiles the test target

The blind spot, and its window. cargo build never compiles #[cfg(test)], so a macOS-only test module is invisible to every Linux job and to build-macos, which ran cargo build only. The window was eleven days: from cdf41b2 (2026-09-24, F4 turned Backend into a registry handle and left 16 pre-F4 Tag::Metal spellings in src/graph/metal_backend/tests.rs) to 4add59f (2026-10-05, the #255 fix), no macOS test binary could compile — so no macOS test ran, and the suite's 21 failures were the first honest measurement of it (#298/#231).

The fix and the rule. build-macos now runs cargo test --release --no-run before its build check and keeps the cargo build step, so a #[cfg(target_os = "macos")]-gated test target is compiled on every push; the macOS-only assertions still need a Mac (the job is link-only and compiles for aarch64-apple-darwin). Cost: the added --no-run step measured 12.33 s warm on dgxspark (aarch64, GB10 sm_121) (2026-10-06) — it reuses the job's existing build. The rule this leaves behind: a platform-gated test target is compiled by neither cargo build nor another platform's CI job, so the platform's own job must pass --no-run — the same shape as the CUDA job's cargo test --release --features cuda --no-run.

Test-infrastructure record (#265, 2026-10-05) — the Metal split, in the three increments Addendum 2 asked for

The ticket. #265 is Step 4 of the source-layout campaign: split src/metal.rs (2 472 lines at a7fc07e) into src/metal/{runtime,encode,ops,policy}.rs and src/metal.metal (5 151 lines) into src/metal/kernels/*.metal + *.h. It is a Mac-local task — src/main.rs gates the module on target_os — and the ticket's Addendum 2 (2026-10-05) required three increments, each verified on the Mac, rather than one step. Box macbook (macOS 27.0.1, Apple M4 Pro), toolchain pinned at 1.97.1 (RUSTUP_TOOLCHAIN unset), worktree .worktrees/265 branched from the landed #255.

Increment 1 — build plumbing, shader set unchanged. build.rs gains an explicit SHADER_SOURCES list, concatenates it into $OUT_DIR/minfer.metal, compiles that into the one minfer.metallib, emits one cargo:rerun-if-changed per listed part, and runs check_shader_file_list() on every build (every .metal/.h under src/ must be listed; a listed file must exist). src/metal.rs's newLibraryWithSource fallback now include_str!s the same generated file. The generated source is byte-identical to src/metal.metal, and the metallib is unchanged at 410 942 B / sha256 13af518e… — expected, because no shader and no flag moved.

Increment 2 — one family per commit (16 commits), then src/metal.metal deleted. The final partition, verified by reassembling the ranges: 212 923 B, 5 151 lines, 61 kernel void kernel_* names, identical sets. src/metal/kernels/ is common.h 13 + dequantize.h 198 + mul_q4_0_q8_0 350 · mul_f32act_q4q5 504 · mul_f32act_kquant 492 · mul_mm 397 · mul_mm_kq 640 · get_rows 184 · norm_elementwise 219 · rope 37 · fa_parallel 129 · kv 115 · qkv_fused 283 · fa_split 262 · fa_decode 516 · fa_prefill 812. Every family was built and run (0.5B Metal) before its commit; the metallib hash moved to 7a4a7cd4… (same size — the functions are the same, the concatenation order is not). Three corrections against plan §4 Step 4's table: no quantize.h (the tree has no GPU-side quantize helper — the only quantize strings are comments); get_rows.metal includes the warm-up kernel (2 340–2 523); rope.metal is 37 lines because kernel_rope_f32 sits after the P1 parallel-attention section. The tree wins; the plan's table is corrected in place.

Increment 3 — the Rust split, 0 visibility edits. src/metal.rs 2 480 → 451 lines; the type definitions (MpsState, MpsStateInner, MpsCommandBuffer), the objc2 aliases, the free items and the private dispatch primitives (trace_op/set_params/barrier/dispatch_1d/dispatch_2d/dispatch_3d/ gemm_enabled/gemm_dispatch, plus the test-only matmul_on_gpu_buf and get_or_grow) stay in the parent, so the four children reach private fields and methods with no pub(super) — the same shape that kept CudaState in src/cuda.rs. metal/ops.rs 1 428 · metal/runtime.rs 475 · metal/encode.rs 114 · metal/policy.rs 66 (re-exported through pub use policy::{…} so crate::metal::kv_cache_is_f16 etc. still resolve). Two findings from doing it: get_or_grow is not internal to impl MpsState as the #255 note said — ops.rs calls it four times, so it moved to the parent; and matmul_on_gpu_buf had to stay in the parent (or move to tests.rs) because a private method of metal::ops is invisible to metal::tests.

A fourth compile entry point. The ticket names two (build.rs, the runtime fallback); the tree has four: tests/{flash_attn,flash_attn_blk,gqa_attn,gemm}_isolation.rs include_str! the shader at 9 sites and compile it themselves. They now read the same $OUT_DIR/minfer.metal, so all four consumers cover one file set by construction (env!("OUT_DIR") is available to integration tests). Leaving them on src/metal.metal is what first broke cargo test.

Gate counts, before (#255 baseline) → after. macbook (macOS 27.0.1, Apple M4 Pro), 2026-10-05:

Command#255 baseline#265
cargo build --releaseexit 0, metallib 410 942 B / 13af518e…exit 0, 410 942 B / 7a4a7cd4…
cargo test --release483 / 20 / 38483 / 20 / 38 (one run 482/21 — cuda_conversation_multiturn_reuse is flaky under the parallel harness; it passes 3/3 alone)
scripts/real_model_gates.sh (0.5B)37 / 137 / 1 (same a_slot_snapshot_resumes_the_context_without_re_prefilling)
scripts/real_model_gates.sh + Qwen3-0.6B37 / 137 / 1 (same)
check_doc_line_anchors.py1672 anchors, exit 01542 anchors, exit 0
check_source_layout.py / check_dead_code_annotations.pyexit 0 / 9 bareexit 0 / 9 bare
cargo fmt --all --checkexit 0exit 0

Runtime fallback verified, not assumed. With the build-time metal compile forced to fail (a bad flag), build.rs wrote the empty marker and the 0.5B model still loaded and ran through newLibraryWithSource; MINFER_METALLIB_FILE=/tmp/empty.metallib exercises the override-failure arm.

Mutation evidence. An unlisted src/zz_probe_mutation.metal (plus a build.rs touch, because cargo only reruns the script on a watched change) → build.rs:77 panics with build.rs SHADER_SOURCES must list every .metal/.h under src (a shader absent from the list is never compiled), exit 101.

Doc sweep. 130 live anchor occurrences re-pointed across 13 documents (line numbers dropped for the symbol anchor, the policy's preference) and the plain metal.rs/metal.metal mentions updated in 26 live files; README.md's mention too, even though the checker does not watch it. Frozen records (PARAMETER_AUDIT.md, ARCHITECTURE-EXECUTION-PLAN.md, QWEN2.5-*, KNOWN-CPU-ISSUES-*, cuda_optimization_steps/*) keep their pre-split anchors. The five cuda_tutorial anchors that cited src/metal.rs for CUDA symbols were re-pointed at the CUDA files (cuda/methods/elementwise.rs, cuda/methods.rs, graph/kvformat.rs) instead of the new Metal ones.

Source layout plan — the runtime, launch and kernel layers

Status: every step has landed (2026-10-05). Steps −1, 0, 1, 2, 3, 4 and 6 are on master with the merge SHAs in the table below; Step 5's Linux half (the MINFER_OP_TIMING module-load re-measure and the stale-prose sweep) landed with this document's own row, and its Mac half (#53's Metal DeviceMemory answer, the second implementation the common decision waits on) landed on macbook (macOS 27.0.1, Apple M4 Pro) — the decision itself is recorded in §1.3 rule 3. Step 4 (Metal, #265) landed from a Mac in three increments. This is the plan of record for splitting the four long backend files and for the naming convention the crate follows afterwards. Tickets: #261 (umbrella) + #262 cuda.rs · #263 cuda/kernels/ · #264 CPU · #265 Metal (Mac) · #266 anchor checker · #267 test files. Step −1 landed #138 and #225 first, as decided.

Status

stepticketstate
−1 #138 + #225#138, #225landed: #225 5386a1c (PR #268, 7/7 green; the cold first-run row is in §2.4); #138 cd39894 (PR #271, 7/7 green)
0 plan document + conventionsthis filelanded 9174644 (PR #269) + the step's own record 1c9e68d (PR #272): this document + SUMMARY.md + AGENTS.md + ARCHITECTURE.md + BACKENDS.md, plus 4 of #219's 6 stale claims
1 src/cuda.rs → src/cuda/*.rs#262landed d08e05e (PR #276, 7/7 green, zero code annotations): src/cuda.rs 6 594 → 1 014 lines, src/cuda/{ffi_runtime,methods}.rs + 18 src/cuda/methods/*.rs; counts unchanged (CUDA 567/0/42, CPU 481/0/36, integration 10/0/6); 0 visibility edits, 0 newly dead
2 src/cuda_kernels.cu → src/cuda/kernels/#263 (after #266)landed (PRs #284 G1 930e1e2, #285 G2 286a5a4, #286 G3 f68a542, #287 G4 c0ddf8b, #288 G5 ee380ad, #289 G6 0cb21cb): src/cuda_kernels.cu 10,215 lines → deleted; src/cuda/kernels/ = common.cuh (498 lines, 1 header) + 17 TUs (9,996 lines); audit 130 sites; counts unchanged (CUDA 567/0/42, CPU 481/0/36, integration 10/0/6, real-model 42/0 ×2); the per-stage record is at the end of §4 Step 2, and the re-measured module-load cost is in §5
3 CPU files#264landed: stage A quants cbee81f (PR #290) — src/quants.rs 1,340 → 61 lines + 9 part files + src/quants/neon_correctness.rs, scripts/check_source_layout.py rule 1 widened (#274); stage B vec_ops 395da59 (PR #291) — src/vec_ops.rs 1,344 → 52 lines + 8 part files; stage C kernel 4900298 (PR #292) — src/kernel.rs 667 → 22 lines + 3 part files; counts unchanged (481/0/36 + 10/0/6)
4 Metal (Mac-local)#265landed in three increments (2026-10-05, macbook (macOS 27.0.1, Apple M4 Pro)): src/metal.rs 2 480 → 451 lines + src/metal/{runtime,encode,ops,policy}.rs (475/114/1428/66) and src/metal.metal 5 151 lines → deleted, replaced by src/metal/kernels/ = 2 headers (common.h 13, dequantize.h 198) + 14 family .metal (largest fa_prefill.metal 812, §9 decision 2); build.rs compiles the concatenated parts into the one metallib and the runtime + the 4 tests/*_isolation.rs read the same generated $OUT_DIR/minfer.metal; 0 visibility edits; counts unchanged (macOS 483/20/38, real-model 37/1 ×2 — the same red baseline as #255, src/metal.metal is gone, so no anchor survives in the live docs); PR #295
5 close the loop—landed: Linux half (PR #293, 2026-10-05) — the MINFER_OP_TIMING module-load re-measure in §5 + the src/cuda_kernels.cu prose sweep + the last two #219 claims; Mac half (PR #296, 2026-10-05, macbook (macOS 27.0.1, Apple M4 Pro)) — #53's Metal DeviceMemory answer (§4 Step 5), and the common decision it unblocked is recorded in §1.3 rule 3: no new module
6 long test files#267files 1–9 landed: PR #275 424c64d (file 1), #277 66446e5 (2), #278 b439612 (3), #279 d3edb10 (4), #280 587c156 (5), #281 cd2c14a (6–8), #283 bc30152 (9): graph/cuda_backend/tests.rs 8,603 → a 106-line parent + 11 tests/<topic>.rs (61 tests); models/qwen2/graph/tests.rs 3,399 → a 111-line parent + 5 (21); server/batch/tests.rs 2,630 → a 356-line parent + 7 (21); graph/alloc/tests.rs 1,833 → a 72-line parent + 6 (42); tooling/tests.rs 1,669 → a 166-line parent + 7 (14); sampler/tests.rs 1,232 → a 140-line parent + 10 (47); conversation/tests.rs 1,156 → a 209-line parent + 6 (27); graph/kvcache/tests.rs 1,134 → a 122-line parent + 5 (33); cuda/issue162_tests.rs 1,186 → a 209-line parent + 4 tests/<topic>.rs (5 tests)

Each step appends its dated record here when it lands (gates run, counts, box label).

  1. Device is the first axis, the layer is the second axis inside each device. The crate keeps its per-device modules (cuda.rs, metal.rs, the CPU trio kernel.rs/quants.rs/vec_ops.rs) and each of them is split into its own inner axis. There is no top-level L1/, L2/, L3/ directory tree: the layers only become directories inside a device.

  2. The cross-device interface stays flat and singular. src/graph/backend.rs (Backend, KvProvider) plus src/graph/registry.rs remain the one device seam — exactly as llama.cpp keeps ggml-backend.cpp + ggml-backend-impl.h + ggml-backend-reg.cpp flat beside the per-device directories. No new trait is introduced by this plan.

  3. A common is only allowed to exist when it has a second real implementation. Interface eligibility = at least two real implementations and at least two callers using it with the same semantics. The one candidate is allocplan::DeviceMemory: CUDA answered it first and, since #53, Metal answers it too. This rule goes into docs/ARCHITECTURE.md.

    Step 5 pre-analysis (2026-10-05, on 4900298) — no common module yet, and the Linux half owes none. The candidate's three halves: the type (allocplan::DeviceMemory, src/graph/allocplan.rs:67) and the pure policy (budget_decision/weight_budget, allocplan.rs:124 + graph/offload.rs:310, with callers in graph/alloc.rs:889/908 and both loaders) are already device-agnostic and CI-tested; the device answer has exactly one implementation (CudaState::device_memory(), src/cuda/methods/accounting.rs:31, reached through the CUDA-only resolver models::device_memory() at src/models/mod.rs:74, 3 call sites), and no Backend trait hook for it exists. #53 adds the second answer behind that existing resolver, so nothing has to move to make it possible; whether the device-answer code then belongs in a common module stays open until it exists, and "no new module" is a live answer (the type and the policy are already in allocplan). Creating the abstraction now would be the single-real-implementation case this rule forbids, so it is deliberately not created here.

    Post-analysis (2026-10-05, on the #53 branch) — the second implementation landed, and the answer is still "no new module". #53 added MpsState::device_memory() (Metal's recommendedMaxWorkingSetSize, src/metal/runtime.rs), so allocplan::DeviceMemory now has the two real implementations the rule asks for. The eligible interface is not the enum alone, though: it is the whole device-answer path, and that path has three parts whose homes are already right. The type and the pure policy (budget_decision, weight_budget) stay in graph/allocplan.rs / graph/offload.rs — device-agnostic and CI-tested. Each device answer stays in its own device module, next to the state it queries (CudaState::device_memory(), src/cuda/methods/accounting.rs; MpsState::device_memory(), src/metal/runtime.rs). The routing is the four-line models::device_memory() (Metal > CUDA > CPU), and it is the two-caller seam the rule is about: graph/alloc.rs::memory_budget (E4's feasibility gate) and both loaders' auto fit (models/qwen2/loader.rs, models/qwen3/loader.rs) read it with the same semantics. A common module would therefore have to hold either the two device methods — moving them out of the modules that own the device state, for no caller's benefit — or the routing, a four-line match: a module with one function is the abstraction-for-one-caller case this rule exists to refuse. Counts at the decision: 2 implementations (CUDA, Metal) and 3 same-semantics callers (the E4 gate + the two loaders) of the resolver; 0 new modules.

  4. Each backend picks its own inner axis (this is what llama.cpp actually does — it is not uniform): CUDA = kernel family, Metal = layer (ggml-metal-device.* → ggml-metal-ops.cpp → kernels/), CPU = ISA (ggml-cpu/arch/{x86,arm,…}).

  5. Existing file paths are preserved. src/cuda.rs stays the module file and gains src/cuda/<part>.rs children (the layout already used by src/cuda/tests.rs); the same for metal.rs, quants.rs, vec_ops.rs, kernel.rs. This keeps all crate::… paths, the check_dead_code_annotations.py grandfather keys, the dead-code baseline file = fields and the documentation's file references valid.

  6. Policy is expressed as pure predicates next to the family they gate (llama.cpp's ggml_cuda_should_use_mmq / _mmvq / _mmf convention), not as a separate policy module and not by relocating the decision point. The predicates become pure functions with unit tests that need no device.

  7. src/cuda.rs keeps its name and its type; the family impl blocks become descendants of one intermediate parent (src/cuda/methods.rs — not impl.rs: impl is a Rust keyword, so mod impl; is a syntax error — expected identifier, found keyword 'impl', verified with a two-file rustc probe on 2026-10-04), so privacy does the work: private fields of CudaState (defined in cuda) and private helper methods (defined in cuda::methods) are visible in every family file. The split is therefore a pure move — 0 field-visibility edits, 0 pub(super) — with exactly one mechanical edit: the 86 extern "C" launch declarations get pub(crate) so the two launch test files keep resolving them through use super::*;.

Confirmed execution decisions (2026-10-04): land #138 and #225 first (Step −1); this document lands as a docs-only PR (Step 0); the seven tickets in §7.4 are filed now.

Non-goals: no behaviour change, no renaming of public items, no new abstraction layer, no Metal work on a non-Mac box.

2. The four layers today

LayerCUDAMetalCPU
L1 device/runtime — context, streams, memory, events, capture, resident weights, device querycuda.rs 1147–1930 + impl families A–Hmetal.rs (MpsState, MetalDevice, MpsCommandBuffer)— (std threads; kernel.rs's Pool is not a device layer)
L2 launch/dispatch — one thin host wrapper per opcuda.rs families I–R (~3,000 lines) + 86 extern "C" declarationsmetal.rs command-buffer encoding (~1,700 lines)kernel.rs + quants.rs + vec_ops.rs
L3 kernel sourcessrc/cuda/kernels/*.cu + common.cuhmetal.metalthe *_avx2 bodies and mod neon_* inside quants.rs/vec_ops.rs
L4 graph executor — Op → backend, buffers, capture replaygraph/cuda_backend.rsgraph/metal_backend.rsgraph/cpu_backend.rs

L2 is not a layer that can be moved away from L1: the CUDA launchers are inherent methods of CudaState, and the Metal ones are methods of MpsCommandBuffer. Splitting L1/L2 apart by directory would be a type refactor, not a file move — that is why the layer axis stays inside each device.

3. What each backend's second axis is

BackendSecond axisTarget shape
CUDAkernel familysrc/cuda/{ffi_runtime,policy}.rs + src/cuda/methods.rs + src/cuda/methods/<family>.rs (L2, Rust); src/cuda/kernels/*.cu + *.cuh (L3 + the C++ half of L2)
Metal (Mac round)layersrc/metal/{runtime,encode,ops,policy}.rs (L1/L2) + src/metal/kernels/*.metal + *.h (L3)
CPUISAsrc/quants/*.rs, src/vec_ops/*.rs, src/kernel/*.rs

The one rule both device backends share: kernel sources live in <backend>/kernels/. llama.cpp is not uniform here (its CUDA keeps *.cu/*.cuh flat in ggml-cuda/, with only vendors/ and template-instances/ as subdirectories, while its Metal puts shaders in ggml-metal/kernels/); this plan chooses the consistent form the request asked for, and states the rule once.

Two honest wrinkles:

  • A CUDA .cu holds the kernels and their host-side launchers (the C++ half of L2), because a launcher must live in the TU that instantiates its kernel (§5). src/cuda/kernels/ therefore means "the CUDA translation units", not "device code only"; the Rust half of L2 is src/cuda/methods/.
  • CPU is the exception: quants.rs and vec_ops.rs are not device-private layers — they are the crate's numeric kernel library (graph/kvformat.rs uses quants::quantize_row_q8_0_into, graph/cuda_backend.rs uses vec_ops::RopeStyle), so they stay where they are and are split by ISA, not into a kernels/ directory.

4. Steps

Step −1 — land the two colliding tickets first (decided 2026-10-04)

  • #138 (F5 late cross-backend wait). It edits copy_to_host and consumes stream_wait_event / cudaStreamWaitEvent — both in cuda.rs family G, and both are the only two grandfathered bare allow(dead_code) sites. Landing it first removes code the split would otherwise move and deletes two GRANDFATHERED_BARE keys plus one docs/dead-code-baseline.toml entry.
  • #225 (pre-warm cost table). It corrects the same measurement the .cu split will change (the fatbin's per-module one-time load), so the corrected record must exist before Step 2 re-measures it.
  • Acceptance: each ticket's own gates; master hard-synced afterwards.

Step 0 — documentation and hygiene (docs-only PR)

  • Land this document as docs/SOURCE-LAYOUT-PLAN.md; add it to the AGENTS.md docs index and update the AGENTS.md Layout block.
  • Fold the stale facts found while measuring into #219 (it already owns two of them): AGENTS.md:3 ~4400 LOC (production code is 55,529 lines), inference_e2e_walkthrough/15-cuda-backend.md:4/29/36 line counts, CUDA-BACKEND-DESIGN.md §"the gates" — 120 sites / 120 / 120 (the count at that revision; 130 on 6b6d94f). The banner naming the deleted CudaCommandBuffer was rewritten to src/cuda/methods/dispatch.rs:124 by Step 1 (#262), the step that moved it. If #219 is not widened, the rest become ticket 7 in §7.4.
  • Add the interface-eligibility rule (§1.3) and the layer definition (§2) to docs/ARCHITECTURE.md.
  • Acceptance: check-docs green (check_docs_links.py, check_status.py --check, book build). No counter in docs/status.toml changes (the suite counts do not move in this step).

Step 1 — src/cuda.rs → src/cuda/*.rs (landed)

  • src/cuda.rs (6 594 → 1 014 lines) keeps: the module doc, AttnWindow, pub struct CudaState (all 29 fields stay private), the free items (cuda_error_name, cstr_owned, CudaPtr, CudaDevicePropBuf, the memcpy/attribute consts, StreamScratch/StreamBinding/bind_stream, ModelLoadGuard, PinnedPool/PinnedBuf, CaptureStaging, MmqCache, layout_of/format_of, concat_rows, the KV-layout constants, gemm_prewarm_disabled, …), the ten #[cfg(test)] mod declarations, and the mod/use/#[cfg(test)] pub(crate) use lines that re-export the two moved FFI surfaces.
  • src/cuda/methods.rs (323 lines) — the 15 non-pub helpers whose callers land in a second family file, computed as the transitive closure of the cross-file call graph: context_stream, get_or_grow (called from seven families), mmq_quantize_transposed, mmq_quantize_native, decode_quantize_native, record_mmq_cache_native, prefill_gemm_f16, the four no_*_mmvq predicates, no_q80_p32, fused_b_on, no_w16cache, no_prefill_gemm, no_fa_prefill; plus the 18 mod <family>; declarations and the #[cfg(test)] pub(crate) use <family>::*; re-exports the two launch test files resolve through use super::*. plane_budget_ok stays in methods/weights.rs (every caller is there), so the split has 0 field-visibility edits and 0 pub(super).
  • 18 family files under src/cuda/methods/ (28–728 lines each): each holds its impl CudaState block and its own pub(crate) extern "C" launch declarations (the 86 declarations move with their family; no symbol is used by two families). The declarations are re-exported two levels (methods.rs → cuda.rs) under #[cfg(test)], because a single-level glob does not reach cuda::tests and an ungated one is an unused_import under deny(warnings); the families whose declarations a methods.rs helper calls are re-exported ungated.
  • src/cuda/ffi_runtime.rs (145 lines) — the cudart/driver FFI block (its declarations become pub(crate), the second mechanical widening) and the test-only extern block.
  • src/cuda/methods/policy.rs (113 lines) holds the MMQ gate family J (mmq_gate_on, mmq_enabled, mmq_active, cc, mmq_a_fuse_mode). The no_*/fused_b_on predicates are cross-family (dispatch, prefill_f16, attention), so they live in methods.rs with the other shared helpers: parking them in a sibling policy.rs is exactly what would have forced the pub(super) edits this step avoids.
  • Acceptance: CUDA unit 567 / 0 / 42; CPU 481 / 0 / 36 on dgxspark (aarch64, GB10 sm_121) (479 on the CI runner) + integration 10 / 0 / 6 — the rows #138 moved when it landed (565 → 567, 480 → 481 / 478 → 479); cargo fmt --all --check; check_source_layout.py; check_dead_code_annotations.py (its two src/cuda.rs: keys re-pointed to src/cuda/ffi_runtime.rs:cudaStreamWaitEvent and src/cuda/methods/events.rs:stream_wait_event); check_dead_code_oracle.py --config {cpu,cuda} with only the two moved file = lines in docs/dead-code-baseline.toml; real-model gates FEATURES=cuda scripts/real_model_gates.sh 42 / 0 ×2 with bitwise-identical greedy output; check_doc_line_anchors.py green (the split's per-line map is /home/yusiwen/minfer-split/step1/line-map.tsv, and the anchors whose old line was only a locator were re-anchored to the symbol).

Step 2 — src/cuda_kernels.cu → src/cuda/kernels/ (landed: 1 header + 17 TUs)

Route (a): launchers move with the kernels they launch (llama.cpp's CUDA shape). Two of the 77 launchers are the exception (blueprint: /home/yusiwen/minfer-split/step2/README.md §2) — they launch kernels that land in two different target files, so "the launcher moves to its kernel's file" needs the qualification: launch_gqa_attn_split_f16kv launches both gqa_attn_split_partial (decode) and gqa_attn_split_partial_hybrid (hybrid), resolved by merging attention_hybrid.cu into attention_decode.cu; launch_mmq_raw_nb_bt_nt launches mmq_raw_nb_bt_kernel (nb) and mmq_ksplit_reduce_kernel (bt_q6k), resolved by moving mmq_ksplit_reduce_kernel next to the NB kernels. No -rdc=true, no new nvcc flag (route (b), -static-global-template-stub=false, is the recorded fallback; see §5). Prerequisite: ticket 6 in §7.4 (the documentation-anchor checker) lands first or in parallel, because this step moves 163 line anchors in 23 documents.

Placement (decided 2026-10-04): src/cuda/kernels/, so that both device backends obey one rule — kernel sources live in <backend>/kernels/ (Metal gets src/metal/kernels/). The path src/cuda_kernels.cu disappears, so the 342 documentation mentions of it are swept in this step (they are being swept for anchors anyway).

Granularity: every file at or below ~800 lines, no kernel split across files. The previous draft stopped at "ten family TUs", which left attention (~1,600 lines) and mmq_prefill (~2,270) too large. The section inventory measured on 6b6d94f regroups into:

file (new)source sections (pre-split lines)≈ lines
kernels/common.cuhmacros (Q4B…WARP), warp_reduce_sum, h2f, get_scale_min_k4, Q8PB/MMQ_A_*, the KV layout + load idiom (2655–2793), declarations of the minfer_launch_*/minfer_smem_optin family~450
kernels/guard.cu5521–5922 — #147 gating + #162 sticky state + the minfer_launch_* definitions (one owner; external linkage)402
kernels/matmul_f32act.cu63–693 (+ its launchers)~750
kernels/mmvq_aquant.cu694–1094 — fused-producer A-quantize, transposed-A prepass, pad40 producer fusion401
kernels/mmvq_skipwrite.cu1095–1651 — P6 r52 mode-2 skip-write variants557
kernels/mmvq_q6k.cu1652–1942 — pipelined q6_K + dense split-plane291
kernels/ops_misc.cu1943–2375 — padded Q6_K matmul, row gather/embed, f32×f32, f16×f32433
kernels/ops_elementwise.cu2376–2654 — f32→Q8_0 quantize, RMSNorm, bias, add/mul/SiLU/SwiGLU, i32 decode, RoPE279
kernels/kv_store.cu2794–3084 + 10173–10215 — KV store, fused QKV epilogue (f16 + packed), arena row move334
kernels/attention_decode.cu3085–3677 — GQA f32, E1 window, kv_map, split-K, batched split593
kernels/attention_hybrid.cu3678–4047 — hybrid rpw (hd 128, f16 KV)370
kernels/attention_prefill.cu5136–5520 — FA-style prefill (staged KV)385
kernels/gemm_wmma.cu5923–6490 — dequant-to-f16 + wmma HGEMM568
kernels/gemm_smem.cu6491–6859 — prefill-GEMM dynamic smem formula + checked opt-ins369
kernels/gemm_fused_dequant.cu6860–7135 — 8p fused dequant-in-GEMM276
kernels/mmq_int8.cu7136–7540 — R1 int8 MMQ prefill GEMM405
kernels/mmq_raw.cu7541–8074 — P6 raw-byte MMQ534
kernels/mmq_nb.cu8075–8667 — raw-nibble NB + its A-layout transform593
kernels/mmq_bt_q6k.cu8668–9407 — r38 q6_K BT740
kernels/mmvq_multi.cu9408–10172 — multi-token MMVQ + doc103/doc104 decode arms765

(the extern "C" launcher block 4052–5135, 1,086 lines, contributes ~50–150 lines to each file above; that is why matmul_f32act and attention_* look slightly over their section size.)

  • minfer_prewarm_kernels (9188–9237) is decomposed into one extern "C" minfer_prewarm_<family>_kernels() per file plus a dispatcher that keeps the symbol name the Rust side declares at src/cuda/methods/prefill_mmq.rs:48.
  • Grouping into PRs (revised): 5–6 groups, not one family each — the file count grew from 10 to 20, so the natural batches are: (1) common.cuh + guard.cu (the infrastructure), (2) attention (3 files), (3) the MMQ prefill family (4 files), (4) the MMVQ decode family (4 files), (5) ops/KV/gemm (5 files), (6) matmul_f32act + leftovers. Each group is independently verifiable on the device.
  • Tooling in the same PRs:
    • check_cuda_launch_returns.py discovers src/cuda/kernels/*.cu as a list and audits one file per audit() call (_RESOLVE_LINES is a module global); it must also assert that no <<<>>> lives in a .cuh (the invariant that keeps the audit complete).
    • tests/fixtures/cuda_launch_sites.tsv keeps its four columns and is regenerated in the build list's order. The identity key is the ordered (owner, site, kernel-fragment) list, and only 2 of the 130 triples repeat, so four columns still identify every site; the --fixture writer emits no file column and src/cuda/issue162_tests.rs's assert_eq!(f.len(), 4) stays as it is. The one tooling change is --check-fixture's filter, len(w) == 4 → len(w) >= 4, so a future column cannot silently empty the fixture (a 5-column fixture under the old filter audited 0 of 130 sites — the pre-verified mutation in the blueprint's §5).
    • build.rs compiles the list, emits one .o per file, keeps libcuda_kernels.a, and adds one rerun-if-changed per .cu and per .cuh (a header edit that does not trigger a rebuild is the silent-stale hazard of this step). It also gains a new-file guard: every .cu/.cuh found in src/cuda/kernels/ must appear in the explicit list and every listed file must exist, so a new kernel file cannot silently not compile.
  • Acceptance per group: 130-site audit + --check-fixture; MINFER_TEST_ISSUE162=1 device gate; CUDA unit 565 / 0 / 42; real-model gates 42 / 0 ×2 with bitwise-identical greedy output; cold-start timing recorded (the fatbin module count changes — see §7 and #225).

Stage record (2026-10-04, dgxspark (aarch64, GB10 sm_121)). One PR per group. Every stage regenerates tests/fixtures/cuda_launch_sites.tsv in build.rs's order (still 130 rows, still four columns) and states the three numbers — the moved files, the remaining src/cuda_kernels.cu, the audit's site count — so "nothing lost" is checkable in each PR rather than only at the end. The non-mutating gates are the same at every stage: audit 130 / --check-fixture exit 0, CUDA unit 567/0/42 + integration 10/0/6, CPU 481/0/36 + 10/0/6, real-model 42/0 on both models, 0 nvcc warnings.

stagefiles (new)linescuda_kernels.cuPR
G1common.cuh 498 + guard.cu 3058039 475#284, 930e1e2
G2attention_decode.cu 1 178 + attention_prefill.cu 5121 6907 828#285
G3MMQ: mmq_int8.cu 457 + mmq_raw.cu 645 + mmq_nb.cu 708 + mmq_bt_q6k.cu 4362 2465 629—
G4MMVQ: matmul_f32act.cu 731 + mmvq_aquant.cu 413 + mmvq_skipwrite.cu 648 + mmvq_q6k.cu 342 + mmvq_multi.cu 7552 8892 803—
G5ops_misc.cu 743 + ops_elementwise.cu 439 + kv_store.cu 435 + gemm_wmma.cu 937 (incl. gemm_smem) + gemm_fused_dequant.cu 2842 83814—
G6the 14-line remainder deleted; src/cuda_kernels.cu retired0——

Three measured corrections to the tables above, applied as the stages land:

  • guard.cu is 305 lines, not 402. The §4 range 5521–5922 also covered launch_fa_prefill_kv (102 lines), but that launcher launches fa_prefill_kv, whose instantiations live in the FA-prefill section — route (a) puts it in attention_prefill.cu, so G1 takes only the #147/#162 state and its one-owner minfer_launch_* definitions.
  • minfer_prewarm_kernels needs five per-family registration functions, not six: mmq_nb, mmq_raw, mmq_bt_q6k, attention_prefill, attention_decode are the only translation units that own template __global__ instantiations the pre-warm address-takes (the other pre-warm entries are plain kernels and stay in the dispatcher, which is where the symbol src/cuda.rs declares it lives). The dispatcher itself moves with mmq_bt_q6k.cu in G3.
  • gemm_smem.cu merges into gemm_wmma.cu (937 lines, not 568 + 369): the MINFER_GEMM_OPTIN_SET table, gemm_f16_fn_for and both GEMM launchers address-take and launch gemm_f16_nt_kernel_t instantiations that only gemm_wmma.cu defines. Two files would be the cross-TU template shape that fails to link, so route (a) merges them — 18 family TUs, not the plan's 19.
  • The mmq_ksplit_reduce_kernel site appears twice. launch_mmq_raw_nb_bt_nt and launch_mmq_raw_nb_bt_q6k_nt both launch it (fixture rows 108 and 111), so moving the reducer next to the NB kernels leaves the q6_K BT launcher with a cross-TU launch. That is legal where route (a) forbids a cross-TU reference: the reducer is a plain __global__, and only a template instantiation fails without -rdc (blueprint §6 fact 2 vs §7.3).

Step 3 — CPU files

In-place split along the ISA axis that is already there (src/quants.rs mod neon_kernels / mod neon_q8k, src/vec_ops.rs mod neon_f16 / mod neon_vec, plus the inline #[cfg(target_arch = "x86_64")] *_avx2 bodies). No path changes: the three file names stay the module deciders. Three stages, one PR each, quants → vec_ops → kernel:

  • 3A src/quants.rs — dot_q4_0 / dot_q4_1 / dot_q5 / dot_q8_0 / kquant / quantize_q8_0 / quantize_q8_k / avx2 / neon (the two inline NEON modules, flattened into the one file), and the inline #[cfg(all(test, target_arch = "aarch64"))] mod neon_correctness promoted to src/quants/neon_correctness.rs — which is also the extraction that lets scripts/check_source_layout.py's first rule be widened to read the cfg predicate instead of the literal #[cfg(test)] (#274).
  • 3B src/vec_ops.rs — vec / rms_norm / rope / softmax / silu / f16 (with the inline mod neon_f16) / bf16 / neon (the promoted mod neon_vec).
  • 3C src/kernel.rs — dispatch / pool / embed.

Acceptance for each stage: CPU 481 / 0 / 36 + 10 / 0 / 6 (read from docs/status.toml; the plan's older text said 480, which predates #138's two tests), the dead-code (name, kind) set identical on aarch64 and x86_64, and check_source_layout.py / check_dead_code_annotations.py / cargo fmt --all --check green.

Stage A landed (PR #290): src/quants.rs 1,340 → 61 lines + dot_q4_0 71 · dot_q4_1 44 · dot_q5 94 · dot_q8_0 64 · kquant 185 · quantize_q8_0 125 · quantize_q8_k 114 · avx2 81 · neon 407, plus the extracted neon_correctness.rs 137. neon_kernels and neon_q8k are flattened into the one neon.rs (no item name collides), their cross-references lose the super::neon_kernels:: prefix. pub(super) replaces "private to quants" on the items a sibling reaches, so the reachable set is unchanged; the parent's pub use list is what keeps crate::quants::… (and graph/kvformat.rs's two calls) resolving. The old file had 48 fn definitions and the new files have the same 48 (excluding the pre-existing tests.rs); all 29 cpu-map.tsv items resolve in their mapped targets.

Stage B landed (PR #291): src/vec_ops.rs 1,344 → 52 lines + vec 394 · rms_norm 154 · rope 17 · softmax 91 · silu 102 · f16 318 (the inline mod neon_f16 stays nested, its super::F16_SIMD_PATH_CALLS unchanged) · bf16 95 · neon 165 (the promoted mod neon_vec — its items keep pub(super), the same pub(in vec_ops) reach they had as a nested module). 47 fn definitions on each side of the move. Three f16 re-exports (dot_f16_f32, dot_f16_f32_scalar, f16_dot_path, F16DotPath) are #[cfg(test)]: their only consumer is vec_ops::tests, and a non-test pub use of an unused name is an unused_imports error under #![deny(warnings)]. RopeStyle::Interleaved's docs/dead-code-baseline.toml file = field moves to src/vec_ops/rope.rs in the same PR.

Stage C landed (PR #292): src/kernel.rs 667 → 22 lines + dispatch 72 · pool 313 · embed 278. 10 fn definitions on each side of the move. Pool's fields and the MmJob/PoolJob/ParForJob types become pub(super) because dispatch.rs submits through them (the same pub(in kernel) reach they had as private items of kernel), and Pool's gate field carries its hazard comment into pool.rs. One re-export is deliberately not carried: cpu_quant_matmul keeps its pub in dispatch.rs but is not re-exported at the kernel root, because its only caller is cpu_quant_matmul_f32 in the same file and #![deny(warnings)] rejects a pub use of an unused name; no crate::kernel::cpu_quant_matmul path exists anywhere in the tree.

Step 4 — Metal (Mac round, after #255)

src/metal.rs and src/metal.metal are not compiled on Linux (src/main.rs:31-32 gates the module; a failed .metal compile only warns and writes an empty metallib marker), so this step is part of #260 and is verified by the Mac-local gates.

Shape: src/metal/{runtime,encode,ops,policy}.rs + src/metal/kernels/*.metal + *.h — the same rule as CUDA (<backend>/kernels/ = the shader sources), which is also llama.cpp's Metal shape (ggml-metal-device → ggml-metal-ops → kernels/, 22 .metal + common.h/dequantize.h/quantize.h).

Granularity: every .metal file at or below ~800 lines. The 5,151-line metal.metal regroups as:

file (new)source sections (pre-split lines)≈ lines
kernels/common.hshared macros/preamble~80
kernels/dequantize.h + quantize.h595–890 dequant helpers (shared by every GEMM) + the quantize helpers~330
kernels/mul_q4_0_q8_0.metal15–363 — Q4_0×Q8_0 + its prefill349
kernels/mul_f32act_q4q5.metal364–567, 1530–1847 — Q5_1, Q4_0 prefill, Q4_1/Q5_K matmul + prefill~500
kernels/mul_f32act_kquant.metal1848–2339 — Q4_K/Q6_K/Q8_0 matmul + prefill~490
kernels/mul_mm.metal568–1144 — Q4_0/Q4_1/Q8_0 simdgroup GEMM~580
kernels/mul_mm_kq.metal1145–1529 + 4897–5151 — Q5_0/Q5_1/Q6_K/Q4_K/Q5_K simdgroup GEMM~640
kernels/get_rows.metal2340–2508 — embedding lookups, all types169
kernels/norm_elementwise.metal2524–2742 — RMSNorm ×2, add, add-bias, mul, SiLU, SwiGLU~220
kernels/rope.metal2743–2746 + the RoPE kernels~60
kernels/fa_parallel.metal2747–2908 — P1 parallel prefill attention162
kernels/kv.metal2909–3023 — KV store + fused bias/rope/store epilogue115
kernels/qkv_fused.metal3024–3306 — fused decode QKV with per-head Q/K RMSNorm (Qwen3)283
kernels/fa_split.metal3307–3568 — KV-parallel split attention (decode)262
kernels/fa_decode.metal3569–4084 — flash attention decode516
kernels/fa_prefill.metal4085–4896 — flash attention prefill (812 lines — accepted as one unit, decision 2 in §9)812

build.rs compiles the parts into one metallib, and the runtime newLibraryWithSource fallback (src/metal/runtime.rs include_str!) needs the parts joined (concat!) or a thin umbrella source; both entry points must see the same set. Then #53 (reserve/assign + the DeviceMemory report) lands on top of the new layout.

Landed (2026-10-05, macbook (macOS 27.0.1, Apple M4 Pro)) — three increments as Addendum 2 asked. Increment 1 made build.rs's shader set an explicit SHADER_SOURCES list plus a check_shader_file_list() guard, concatenated it into $OUT_DIR/minfer.metal, and pointed the runtime include_str! at the same file (shader set unchanged; the metallib stayed byte-identical, 13af518e…). Increment 2 moved one family per commit (16 commits) and deleted src/metal.metal; the generated source is byte-for-byte the old one (212 923 B, 5 151 lines, 61 kernel void names, same set), and the metallib hash moved to 7a4a7cd4…. Increment 3 split src/metal.rs with 0 visibility edits — the type definitions and the private dispatch primitives stay in the parent module, so the children reach them without pub(super) (the CudaState-stays-in-cuda.rs shape) — and moved the test-only matmul_on_gpu_buf with them so src/metal/tests.rs still resolves it.

Three measured corrections to the table above, applied as the increments landed:

  • quantize.h is not created. The plan's dequantize.h + quantize.h row assumed GPU-side quantize helpers; the tree has none (the only quantize strings are comments — activations are quantized on the CPU). The pragmatic source of truth is the tree, so src/metal/kernels/ holds common.h + dequantize.h only; a quantize.h with no helper would be a file the guard must list and nothing reads.
  • get_rows.metal includes the warm-up kernel (2 340–2 523), so it is 184 lines, not 169.
  • rope.metal is 37 lines (the 2 743–2 746 banner is not adjacent to kernel_rope_f32, which is at 2 876–2 908 after the P1 parallel-attention section); fa_parallel.metal is 129 lines (2 747–2 875). The plan's "≈60 / 162" rows mixed the two.

A fourth compile entry point the ticket did not name: the four tests/*_isolation.rs integration tests include_str! the shader source and compile it themselves (9 sites). They now include_str! the same $OUT_DIR/minfer.metal, so all four consumers — build.rs, src/metal.rs's fallback and the integration tests — see one file set by construction.

Step 5 — close the loop

Re-measure the N-module cold start and update the §2.4 pre-warm table (#225); implement allocplan::DeviceMemory for Metal (#53, CUDA already answers it) so that "the device memory report" becomes the first interface with two real implementations; add the mechanical documentation-anchor check (§6) if it is not already landed.

Landed. The MINFER_OP_TIMING re-measure is §5.1 (Linux half, PR #293); the anchor checker is #266. The Metal DeviceMemory half landed on macbook (macOS 27.0.1, Apple M4 Pro) (PR #296) — the measured record (before/after counts, the red-baseline note, the mutation transcript) is the E4/E5 "Metal half" record in docs/ARCHITECTURE-EXECUTION-PLAN.md, and the common decision it unblocks is §1.3 rule 3 above: two implementations, three same-semantics callers, no new module.

Step 6 — the long test files (in scope, decided 2026-10-04)

The largest files in the crate are tests: src/graph/cuda_backend/tests.rs (8,384 lines, 143 tests), src/models/qwen2/graph/tests.rs (3,374), src/server/batch/tests.rs (2,629), src/cuda/issue162_tests.rs (1,186), src/graph/alloc/tests.rs (1,773), src/graph/kvcache/tests.rs (1,134), src/conversation/tests.rs (1,156), src/tooling/tests.rs (1,669), src/sampler/tests.rs (1,232), src/graph/cuda_backend/tests.rs … — this step splits them by op family / topic into <module>/tests/<topic>.rs (the same rule: every file named by a mod declaration, checked by scripts/check_source_layout.py).

  • Why last: src/graph/cuda_backend/tests.rs is the evidence base for Steps 1–2 (its 143 tests are what proves the moves), and every later PR's line references would churn if it moved first.
  • Order inside the step: graph/cuda_backend/tests.rs first (the largest), then the other >1,000-line test files, one PR each.
  • Acceptance: the test counts are identical (nothing added or removed — the same tests run from new files, which check_source_layout.py's rule 2 is precisely there to guarantee), plus the step-appropriate gates (CUDA unit 565/0/42 for the executor tests, CPU 480/0/36 + 10/0/6 for the rest).
  • Note: no size ratchet is added (decided 2026-10-04) — the ~800-line target in this document is guidance, enforced by review, not by a script.

5. Why no -rdc=true, and why route (a)

Measured on dgxspark (aarch64, GB10 sm_121), CUDA 13.0, 2026-10-03 (two-file probe, four cross-TU patterns, each compiled and run):

cross-TU usedefault nvcc-static-global-template-stub=false
plain __global__ launch✅ links and runs✅
plain __global__ address-taken + cudaFuncSetAttribute✅✅
templated __global__ launch❌ link error (hidden symbol … isn't defined, nvcc warning #20280-D)✅ runs
templated instance address-taken (the prewarm idiom)❌ link error✅ cudaSuccess

So the default toolchain forces "the launcher lives in the TU that instantiates the kernel" — which is route (a) and is also llama.cpp's CUDA shape. The evidence base for the split's other costs:

  • one nvcc invocation today, 13 -gencode targets (12 SASS + compute_121 PTX); serial compile 117.3 s / 114.1 s (two runs), 675 MB RSS, 35.8 MB object; nvcc --threads 0 16.4 s / 15.6 s;
  • the CUDA build is not bit-reproducible today (two identical serial runs differ by 16 bytes in the .text of two cubins) — so the split's evidence is runtime gates, not binary identity;
  • the recorded per-module fatbin load was ~2.2 ms, set-size independent (docs/CUDA-BACKEND-DESIGN.md §2.4's cost table), so ten modules were extrapolated at ~13–22 ms one-time. Step 5 measured it (2026-10-05) and the extrapolation was wrong — see below.

5.1 The re-measured module load (Step 5, 2026-10-05)

Box dgxspark (aarch64, GB10 sm_121), CUDA 13.0, driver 580.178.04, nvcc 13 -gencode targets (12 SASS + compute_121 PTX) × the 17 TUs; nvidia-smi before the run: SM clock 2 411 MHz idle (warm) / 208 MHz (cold), 0 % util, no other compute process. Two binaries, both built in this repository's worktrees: pre-split = bc30152 (930e1e2^, src/cuda_kernels.cu = one 10,215-line TU = 1 fatbin module) and post-split = 4900298 (17 TUs = 17 modules). Command of record (<0.5B> = the cached qwen2.5-0.5b-instruct-q4_0.gguf):

MINFER_OP_TIMING=1 target/release/minfer <0.5B> "hello"

which prints the prefill-GEMM smem pre-warm loop's own duration — the point at which the fatbin's module is finalized. fresh process = a new minfer invocation; cold = the binary's and the model's pages evicted with posix_fadvise(POSIX_FADV_DONTNEED) first (an agent shell cannot drop_caches), warm = back-to-back fresh processes with the pages resident.

buildmodule(s) the command forceswarm, fresh processcold (page-cache-evicted)
pre-split bc301521 (the whole fatbin)2 300 µs median (2 203–2 468, n=12)18 478 / 20 723 µs (2 runs)
post-split, shipped binary1 of 17 (gemm_wmma.cu)350–590 µs4 288 / 6 109 / 6 983 µs
post-split, all 16 loadable TUs16 of 17 (temporary per-TU probe)≈ 2 370 µs total (15 probes 1 851–2 023 µs + the GEMM module)42 671 / 44 416 / 47 412 µs

Per-module figures (post-split, warm, one live cudaFuncGetAttributes per TU, temporary probe inside minfer_prewarm_kernels — reverted before the PR): mmvq_q6k 57 · mmvq_aquant 57 · mmq_raw 70 · mmq_nb 73 · matmul_f32act 85 · mmq_int8 93 · mmvq_skipwrite 100 · kv_store 102 · mmvq_multi 111 · mmq_bt_q6k 118 · ops_misc 123 · ops_elementwise 245 · attention_decode 223 · gemm_fused_dequant 40 · attention_prefill 447 · gemm_wmma ≈ 350–590 µs (its line also carries the 12 attribute queries), i.e. 40–450 µs per module, ~150 µs median — not the recorded 2.2 ms, which was the whole pre-split module's cost, not a per-module constant.

Verdict: the split's cost is acceptable. In steady state the 17-module fatbin loads in the same ~2.3 ms as the pre-split single module, because the load tracks the code a module contains, not the module count; the MINFER_OP_TIMING line moves from 2.3 ms to 0.4 ms only because it now times one seventeenth of the work. The honest extra is cold: a page-cache-cold start pays ≈ +25 ms (≈ 19 ms → ≈ 45 ms), because every one of the 16 module registrations faults the fatbin's pages again. The cold per-module figures (same probe, evicted pages, 975–5 214 µs) are ~10× their warm values across the whole size range — including 975 µs for the 284-line gemm_fused_dequant.cu — so a per-registration overhead rides on top of the size-proportional part, and 16 registrations pay it 16 times. That is ≈ 1.8 % of the ~1.4 s cold-start wall time on this 0.5B model, once per process, and the mitigation the plan names (minfer_prewarm_kernels trimmed to the modules a run needs) would only move it into the first forward's lazy loads — which is why no trimming is applied. The correction also applies to §2.4's old explanation of the 14.5 ms cold row: it is page-cache-cold, not the GPU clock (a 40 s idle cooldown at a 208 MHz SM clock reads the warm 2.3 ms; the same binary with evicted pages reads 18.5–20.7 ms).

The full per-module transcripts and the two worktrees' build logs are in the Step 5 record in docs/ARCHITECTURE-EXECUTION-PLAN.md.

6. Documentation plan

6.1 What the split invalidates

measurecount
documents that mention one of the four paths (src/cuda.rs, src/cuda_kernels.cu, src/metal.rs, src/metal.metal), measured on 6b6d94f — the campaign's own documents (this plan, AGENTS.md, ARCHITECTURE.md, BACKENDS.md) have since added mentions126 (869 mentions)
documents carrying a line anchor into one of them (…:NNN), measured on 6b6d94f35 (401 anchors: cuda side 269, metal side 132)
the anchor hot spotsdocs/cuda_tutorial/* 180 (6 files), docs/LLAMA_METAL_E2E.md 50, docs/inference_e2e_walkthrough/14-metal-backend.md 35, docs/LLAMA-CPP-MMQ-ANALYSIS.md 18, docs/METAL-OBJC2-MIGRATION-PLAN.md 15, docs/inference_e2e_walkthrough/15-cuda-backend.md 10, docs/metal-inference-analysis.md 10, docs/ARCHITECTURE-EXECUTION-PLAN.md 11
documents that describe the layout and need rewriting, not sweeping18 (listed in §6.3)
machine-checked todaycheck_docs_links.py (relative link targets only — it cannot see path:NNN), check_status.py --check (AGENTS.md prose ↔ scripts/status.toml), build_book.sh (mdBook chapters from docs/SUMMARY.md)

6.2 Policy: live documents are edited, historical records are frozen

  • Live documents (the ones a maintainer reads to find code): edited in the step that moves the code, with anchors converted to symbol anchors (`prefill_mmq` (`src/cuda/kernels/mmq_*`)) wherever the line number was only a locator.
  • Historical records (docs/cuda_optimization_steps/*.md, docs/QWEN2.5-*.md, docs/DEBUGGING-*.md, docs/KNOWN-CPU-ISSUES-*.md, docs/PARAMETER_AUDIT.md's older tables, docs/ARCHITECTURE-EXECUTION-PLAN.md's per-ticket entries, and experiments/cuda/*.md — the probe run records, which quote the nvcc … ../../src/cuda_kernels.cu command as it was run) keep their text — they record a measurement taken against a revision, and rewriting them would falsify the record. They are resolved through the path mapping table this document keeps (§6.4).
  • The machine ledgers the step records name are of their day. They moved beside their checkers on 2026-10-09/10 (ADR-0023, ADR-0024) and now live at scripts/status.toml, scripts/test-baselines.toml, scripts/dead-code-baseline.toml and tests/fixtures/f6-fixtures.json. §6.1's "machine-checked today" row names the current reader; the step text above keeps the paths it was written with.
  • The checker must know about the freeze: scripts/check_doc_line_anchors.py (ticket 6) carries a frozen-file set (the GRANDFATHERED_BARE pattern), so a frozen record does not fail CI, and the set can only shrink. This is the one design constraint the frozen policy puts on ticket 6.

6.3 Per-step update table

StepLive documents edited (content)Mechanical sweep (paths + anchors)
0docs/SOURCE-LAYOUT-PLAN.md (new) + docs/SUMMARY.md (chapter entry) + AGENTS.md (Layout block, docs index, the CUDA/Metal bullets, the ~4400 LOC figure) + docs/ARCHITECTURE.md (module map + the layer/interface-eligibility convention) + docs/BACKENDS.md (the device-layer rows) + the stale-number list folded into #219none yet (no file has moved)
1 cuda.rsdocs/inference_e2e_walkthrough/15-cuda-backend.md, docs/cuda_tutorial/{02,04,05}.md (the Rust-side excerpts), docs/CUDA-BACKEND-DESIGN.md (§device layer), docs/DEVICE-ADAPTATION-PLAN.md, docs/COMPUTE-GRAPH-DESIGN.mdthe 269-anchor cuda-Rust half and the 269 mentions of cuda.rs across the live set
2 .cudocs/CUDA-BACKEND-DESIGN.md (§kernels), docs/cuda_tutorial/{03,04,05,06}.md, docs/LLAMA-CPP-MMQ-ANALYSIS.md, docs/CUDA-TECH-PRIMER.md, docs/CUDA_OPTIMIZATION.md, docs/GPU_SAFETY.md (the <<<>>>/opt-in rules), docs/BUILD.md (the nvcc file list)cuda_kernels.cu 342 mentions + its anchors, via the mapping table
3 CPUdocs/ARCHITECTURE.md, docs/inference_e2e_walkthrough/{10,11}.md, AGENTS.md (Layout) — docs/CPU_OPTIMIZATIONS.md is frozen (§6.2) and keeps its quants.rs/vec_ops.rs line numbersquants.rs/vec_ops.rs/kernel.rs mentions (19 quants.rs:NNN anchors in live docs, re-pointed in stage A; the frozen records resolve through §6.4)
4 Metaldocs/METAL-BACKEND-DESIGN.md, docs/METAL_OPTIMIZATIONS.md, docs/inference_e2e_walkthrough/14-metal-backend.md, docs/LLAMA_METAL_E2E.md, docs/METAL_OBJC2-MIGRATION-PLAN.md, docs/metal-inference-analysis.md, docs/multi-token-kernel-analysis.mdthe 132 metal anchors + 250 metal.rs/metal.metal mentions
6 testsAGENTS.md (the test-module convention paragraph), docs/GATE-CONTRACT.md if a gate's location is namedtest-file paths named in docs
every stepa dated entry in docs/ARCHITECTURE-EXECUTION-PLAN.md §test-infrastructure (the repo's per-ticket record) + the Status table of this document—

docs/status.toml is not edited by Steps 0–6: the suite counts do not move (code moves, tests move, no test is added or deleted). If a step ever changes a count, scripts/check_status.py --check must be updated in the same PR, and this document says so in that step's record.

6.4 The path mapping table (lives here; grows per step)

The frozen records resolve old paths through this table, and the live sweeps are generated from it:

oldnew
src/cuda_kernels.cu 20–50, 2655–2793, 5521–5922src/cuda/kernels/{common.cuh, guard.cu}
src/cuda_kernels.cu 4050–5135distributed: each launcher to its kernel's file
src/cuda_kernels.cu other rangesthe §4 Step 2 table (one row per new file)
src/cuda.rs 82–1128, 1141–1153, 1942–6495src/cuda/ffi_runtime.rs + src/cuda/methods.rs + src/cuda/methods/*.rs (the §8 tree; per-line map /home/yusiwen/minfer-split/step1/line-map.tsv)
src/metal.metal 1–13src/metal/kernels/common.h
src/metal.metal 577–593, 595–756, 1662–1679src/metal/kernels/dequantize.h
src/metal.metal 14–363src/metal/kernels/mul_q4_0_q8_0.metal
src/metal.metal 364–567, 1530–1661, 1680–1847src/metal/kernels/mul_f32act_q4q5.metal
src/metal.metal 1848–2339src/metal/kernels/mul_f32act_kquant.metal
src/metal.metal 568–576, 757–1144src/metal/kernels/mul_mm.metal
src/metal.metal 1145–1529, 4897–5151src/metal/kernels/mul_mm_kq.metal
src/metal.metal 2340–2523src/metal/kernels/get_rows.metal
src/metal.metal 2524–2742src/metal/kernels/norm_elementwise.metal
src/metal.metal 2743–2746, 2876–2908src/metal/kernels/rope.metal
src/metal.metal 2747–2875src/metal/kernels/fa_parallel.metal
src/metal.metal 2909–3023src/metal/kernels/kv.metal
src/metal.metal 3024–3306src/metal/kernels/qkv_fused.metal
src/metal.metal 3307–3568src/metal/kernels/fa_split.metal
src/metal.metal 3569–4084src/metal/kernels/fa_decode.metal
src/metal.metal 4085–4896src/metal/kernels/fa_prefill.metal
src/metal.rs 1–131, 193–312, 314–348, 392–493, 966–987, 2455–2473src/metal.rs (module doc, aliases, type definitions, dispatch primitives, matmul_on_gpu_buf, get_or_grow)
src/metal.rs 132–191src/metal/policy.rs
src/metal.rs 349–391, 1936–1985src/metal/encode.rs
src/metal.rs 495–965, 988–1935src/metal/ops.rs
src/metal.rs 1997–2454src/metal/runtime.rs
src/graph/cuda_backend/tests.rssrc/graph/cuda_backend/tests/{staging,pool,elementwise,matmul,mmvq,prefill,weights,kv,attention,attn_window,capture}.rs
src/models/qwen2/graph/tests.rssrc/models/qwen2/graph/tests/{cuda_kv,offload_copy,kv_reuse,batching,real_model}.rs
src/server/batch/tests.rssrc/server/batch/tests/{kv_sharing,slots,prefill,batching,stall,http,metrics}.rs
src/graph/alloc/tests.rssrc/graph/alloc/tests/{backend_fence,views,liveness,kv_arena,staging,budget}.rs
src/tooling/tests.rssrc/tooling/tests/{parse,f16_encode,f6_roundtrip,f141_device,f167_qwen3,quantize_bounds,bf16}.rs
src/sampler/tests.rssrc/sampler/tests/{greedy_topk,penalties,stops,minp_typical,xtc,dry,mirostat,bias_validate,defaults,grammar}.rs
src/conversation/tests.rssrc/conversation/tests/{turns,regen,spec,snapshot,overflow,real_model}.rs
src/graph/kvcache/tests.rssrc/graph/kvcache/tests/{cells,spans,sharing,resize,defrag}.rs
src/cuda/issue162_tests.rssrc/cuda/issue162_tests/{sites,severity,control,node}.rs
src/quants.rs 10–31, 292–355, 426–457src/quants/quantize_q8_0.rs
src/quants.rs 33–53, 96–140src/quants/dot_q4_0.rs
src/quants.rs 55–94src/quants/dot_q4_1.rs
src/quants.rs 142–202src/quants/dot_q8_0.rs
src/quants.rs 204–290src/quants/dot_q5.rs
src/quants.rs 357–424, 459–466src/quants/avx2.rs
src/quants.rs 471–663, 964–1198src/quants/neon.rs (the two flat NEON modules)
src/quants.rs 674–780src/quants/quantize_q8_k.rs
src/quants.rs 782–962src/quants/kquant.rs
src/quants.rs 1200–1340src/quants/neon_correctness.rs
src/vec_ops.rs 6–20src/vec_ops/rope.rs
src/vec_ops.rs 22–159, 349–569, 725–752src/vec_ops/vec.rs
src/vec_ops.rs 161–258src/vec_ops/silu.rs
src/vec_ops.rs 260–347src/vec_ops/softmax.rs
src/vec_ops.rs 571–721src/vec_ops/rms_norm.rs
src/vec_ops.rs 754–1069src/vec_ops/f16.rs
src/vec_ops.rs 1071–1163src/vec_ops/bf16.rs
src/vec_ops.rs 1165–1332src/vec_ops/neon.rs
src/kernel.rs 8–34, 319–358src/kernel/dispatch.rs
src/kernel.rs 36–317, 360–387src/kernel/pool.rs
src/kernel.rs 389–664src/kernel/embed.rs

6.5 The macOS hand-off (decided 2026-10-04: Step 4 is Mac-local)

Step 4 cannot be executed or verified on the Linux box, so this document must be sufficient alone for a macOS agent: §3's rule, §4 Step 4's file table, §5's cross-TU constraint, §6.3's doc sweep, and §10's verification row. The ticket (T5) and #260 both link here, and Step 4's record names the Mac box explicitly (gate-contract rule 5: an absolute box label, e.g. macbook (macOS 15.x, Apple M4)), never "this box".

7. Interaction with the open issues (as of 2026-10-03, 38 open)

Symbol-level scan of all 38 issue bodies against the identifiers defined in the files to be split (601 distinctive symbols; plus a direct src/<file>:NNN path scan). 18 issues reference affected code or files, or target code that this plan moves. (#150's worker_loop hit is the server's worker_loop_serial, not kernel.rs — counted as unaffected.)

7.1 Must be sequenced against this plan

IssueWhy it collidesAction
#138 F5 late cross-backend waitedits copy_to_host and consumes stream_wait_event / cudaStreamWaitEvent — both in cuda.rs family G, and both are the two grandfathered bare allow(dead_code) sitesland #138 first if it is next: it removes code the split would otherwise move and deletes two grandfather keys; otherwise keep it out of flight during Step 1
#219 two stale CUDA claimsowns walkthrough/15-cuda-backend.md §3.2.2 (register_weight) and CUDA-BACKEND-DESIGN.md — the same files Step 0 and Step 1 re-anchormerge Step 0's stale-fact list into #219 (one docs PR), or land Step 0 first and reference #219
#225 pre-warm cost tablethe split changes the fatbin module count, i.e. exactly what #225 records (2.2 ms per module, cold-run 14.5 ms)land #225's correction first (cheap), then re-measure in Step 2's first increment and cross-reference
#200 CUDA kernel for Op::FusedQkvNormadds a kernel and a launcher, in the attn_bias_rope_store* (family R) shapethe split has landed: add the kernel to src/cuda/kernels/kv_store.cu (or a new kernels/<family>.cu registered in build.rs's KERNEL_SOURCES) and the launcher to src/cuda/methods/kvstore.rs
#208 bf16 weights on deviceadds device kernels (+ metal.metal) and touches vec_ops::mat_mul_bf16same as #200 — device kernels go into src/cuda/kernels/
#212 packed Q8_0 residual attributionprofiles gqa_attn_f32 (attention family) with line-level referencesthe split has landed: gqa_attn_f32 is in src/cuda/kernels/attention_decode.cu, so its references re-anchor there (or to the symbol)
#164 Metal f16 matmul/embedding kernelsadds kernels to metal.metalMac round; do it after the Metal split (Step 4)
#255 two macOS-only dead-code annotationsits two targets are src/metal/ops.rs / src/metal/runtime.rs — line anchors the Metal split movesjudge them first (Mac), then split
#260 Mac round umbrellathe entry point for a Mac agent; it lists the Metal gaps and the orderupdate it with the Step 4 shape and the new #53 item
#53 Metal reserve/assign (+ the device-memory gap)the only issue that already owns the one interface this plan promotes; its pool code moves in Step 4Step 4 first, then #53; #53 supplies the second DeviceMemory implementation
#44 Metal KV cell store / explicit spanadds Metal kernels + copy_cells work in metal_backend.rs/metal.metalMac round, after Step 4
#56 AVX2/AVX-512 K-quant dots + repackingadds kernels to quants.rs — Step 3's target fileland after Step 3, or rebase onto src/quants/*.rs

7.2 Needs a body/anchor update only

#137 (async staging, copy_cross/await_cross + Metal), #135 (walkthrough/architecture stale Backend enum — same docs), #54 (re-run Metal gap measurements), #52 (mixed-quant QKV epilogue), #39 (debug_assert! in release on Metal), #231 (five macOS-only attention call sites, 8 src/…:NNN references).

7.3 Unaffected

#215, #209, #205, #204, #203, #198, #195, #179, #157, #150 (server-side worker_loop_serial, not kernel.rs), #133, #132, #126, #125, #118, #103, #62, #40, #38, and #229 (allocator dead-code bookkeeping only).

7.4 Tickets this plan files (decided 2026-10-04: filed now, labelled)

#TitleLabelsStep
1Umbrella: #261 [layout] split the runtime/launch/kernel layers per deviceenhancementall
2#262 [cuda] split src/cuda.rs into src/cuda/*.rs (pure move)enhancement1
3#263 [cuda] split src/cuda_kernels.cu into src/cuda/kernels/ (header + guard + 19 TUs)enhancement,test2
4#264 [cpu] split quants.rs / vec_ops.rs / kernel.rs along the ISA axisenhancement3
5#265 [metal] split metal.rs / metal.metal into src/metal/{runtime,encode,ops,policy}.rs + src/metal/kernels/enhancement4
6#266 [docs] mechanical check for src/<file>:NNN anchors + convert to symbol anchorsdocumentation,cibefore 2
7#267 [test] split the >1,000-line test files (cuda_backend/tests.rs first), counts identicaltest6

That is 7 tickets, all filed 2026-10-04; every sub-ticket carries Part of #261, and #261 carries the Step −1…6 checklist, the interaction table of §7.1, the target tree of §8 and the documentation plan of §6. The former "stale size/claim sweep" ticket was folded into #219 as a comment (decision 3, §9).

8. Resulting tree

src/kernel/, src/quants/, src/vec_ops/, src/metal/ and src/cuda/ already exist today — they hold only tests.rs (plus metal/mmap_align_test.rs and cuda/'s ten issue probes). The plan therefore does not create a new convention: the production parts simply join the directories that are already there. [S1]…[S4] name the step that produces each entry; a parenthesised line count is the file's length after the step that created it, and the .cu/.metal ranges refer to the pre-split file.

src/
├── main.rs                                    (unchanged)
├── cuda.rs                          [S1]  module cuda: doc + `AttnWindow` + `pub struct CudaState`
│                                             (fields stay private) + free items + `mod methods;`
│                                             `mod ffi_runtime;` + the `#[cfg(test)] pub(crate) use`
│                                             re-exports + the ten `#[cfg(test)] mod` declarations (1 014)
├── cuda/                                    (exists: 10 test files today, unchanged)
│   ├── methods.rs                   [S1]  the 15 cross-family helpers + the 18 `mod` declarations
│   │                                        + the `#[cfg(test)] pub(crate) use` re-exports (323)
│   ├── methods/
│   │   ├── accounting.rs            [S1]  B   42   weights_bytes / device_memory
│   │   ├── attention.rs             [S1]  P  516   gqa / split / batched / prefill
│   │   ├── buffers.rs               [S1]  E   28   cuda_malloc / cuda_free
│   │   ├── capture.rs               [S1]  H  137   CUDA-graph capture / replay
│   │   ├── copy.rs                  [S1]  F  182   H2D / async / D2H / pinned / D2D
│   │   ├── dispatch.rs              [S1]  I  440   matmul_f32_ptr* + the MMQ dispatch tree
│   │   ├── elementwise.rs           [S1]  O  195   norm / add / mul / silu / swiglu / rope
│   │   ├── events.rs                [S1]  G  183   events, async staging, sync, latch
│   │   ├── gpu_act.rs               [S1]  N  233   on-GPU quantize / gather / embed
│   │   ├── init.rs                  [S1]  A  331   device probe / tier / singleton
│   │   ├── kvstore.rs               [S1]  R  350   KV store + fused QKV epilogue
│   │   ├── mmq_quant.rs             [S1]  K  306   A-quantize + MmqCache
│   │   ├── mmvq.rs                  [S1]  Q  728   decode MMVQ + q8_0 p32 planes
│   │   ├── policy.rs                [S1]  J  113   MMQ gate predicates
│   │   ├── prefill_f16.rs           [S1]  M  307   f16 GEMM + w16 cache
│   │   ├── prefill_mmq.rs           [S1]  L  488   auto_ksplit, prefill_mmq
│   │   ├── stream.rs                [S1]  D   67   bound/context stream, create/destroy
│   │   └── weights.rs               [S1]  C  726   register_weight + q6k/q4k expansion
│   ├── ffi_runtime.rs               [S1]  cudart/driver FFI (`pub(crate)`) + the test-only extern
│   │                                        block (145)
│   └── kernels/                     [S2]  the CUDA translation units (kernels + their host
│       │                                   launchers; `<backend>/kernels/` is the one rule both
│       │                                   device backends share)
│       ├── common.cuh               [S2]  ~450: defines + device helpers + KV load idiom +
│       │                                   declarations of the #147/#162 helpers
│       ├── guard.cu                 [S2]  5521–5922  single owner of the #147/#162 state and of
│       │                                   the minfer_launch_* definitions (external linkage)
│       ├── matmul_f32act.cu         [S2]  63–693            ≈750 with launchers
│       ├── mmvq_aquant.cu           [S2]  694–1094          fused-producer A-quantize prepass
│       ├── mmvq_skipwrite.cu        [S2]  1095–1651         mode-2 skip-write variants
│       ├── mmvq_q6k.cu              [S2]  1652–1942         pipelined + dense split-plane q6_K
│       ├── ops_misc.cu              [S2]  1943–2375         padded Q6_K, gather/embed, f32×f32, f16×f32
│       ├── ops_elementwise.cu       [S2]  2376–2654         quantize f32→Q8_0, norm, bias, add/mul,
│       │                                                     SiLU/SwiGLU, i32 decode, RoPE
│       ├── kv_store.cu              [S2]  2794–3084 + 10173–10215  KV store, fused QKV epilogue,
│       │                                                     arena row move
│       ├── attention_decode.cu      [S2]  3085–3677         GQA, E1 window, kv_map, split-K, batched
│       ├── attention_hybrid.cu      [S2]  3678–4047         hybrid rpw (hd 128, f16 KV)
│       ├── attention_prefill.cu     [S2]  5136–5520         FA-style prefill (staged KV)
│       ├── gemm_wmma.cu             [S2]  5923–6490         dequant-to-f16 + wmma HGEMM
│       ├── gemm_smem.cu             [S2]  6491–6859         dynamic-smem formula + checked opt-ins
│       ├── gemm_fused_dequant.cu    [S2]  6860–7135         8p fused dequant-in-GEMM
│       ├── mmq_int8.cu              [S2]  7136–7540         R1 int8 MMQ prefill GEMM
│       ├── mmq_raw.cu               [S2]  7541–8074         P6 raw-byte MMQ
│       ├── mmq_nb.cu                [S2]  8075–8667         raw-nibble NB + A-layout transform
│       ├── mmq_bt_q6k.cu            [S2]  8668–9407         r38 q6_K BT
│       └── mmvq_multi.cu            [S2]  9408–10172        multi-token MMVQ + doc103/104
├── metal.rs                         [S4]  module metal: doc + free items + `pub use`
├── metal/                                   (exists: mmap_align_test.rs + tests.rs)
│   ├── runtime.rs                   [S4]  L1: MpsState / MetalDevice / library + pipeline cache
│   ├── encode.rs                    [S4]  L2: MpsCommandBuffer encoding
│   ├── ops.rs                       [S4]  L2: the op → encoding table
│   ├── policy.rs                    [S4]  pure predicates (MINFER_METAL_* / MINFER_* knobs)
│   └── kernels/                     [S4]  L3, ≤~800 lines each:
│       ├── common.h · dequantize.h          (no `quantize.h` — the plan's row
│       │                                    assumed GPU-side quantize helpers
│       │                                    and the tree has none; §4 Step 4)
│       ├── mul_q4_0_q8_0.metal · mul_f32act_q4q5.metal · mul_f32act_kquant.metal
│       ├── mul_mm.metal · mul_mm_kq.metal · get_rows.metal · norm_elementwise.metal · rope.metal
│       ├── kv.metal · qkv_fused.metal
│       ├── fa_parallel.metal · fa_split.metal · fa_decode.metal · fa_prefill.metal
│       ├── f16.metal (#164) · f32.metal (#317) · bf16.metal (#208) · attn_window.metal (#44a)
│                                            (the four added by the Metal round after S4)
│                                            (runtime fallback joins them with `concat!`)
├── kernel.rs                        [S3]  module kernel: `mod` + `pub use`
├── kernel/
│   ├── dispatch.rs                  [S3C] cpu_quant_matmul / cpu_quant_matmul_f32 (12–44, 322–390)
│   ├── pool.rs                      [S3C] Pool / par_for / set_cpu_threads (45–321, 364–390)
│   ├── embed.rs                     [S3C] embed_tokens (391–)
│   └── tests.rs                            (exists)
├── quants.rs                        [S3]  module quants: `pub use`
├── quants/
│   ├── dot_q4_0.rs · dot_q4_1.rs · dot_q5.rs · dot_q8_0.rs    [S3A]
│   ├── kquant.rs                    [S3]  Q4_K/Q5_K/Q6_K dots
│   ├── quantize_q8_0.rs · quantize_q8_k.rs                    [S3A]
│   ├── neon.rs                      [S3A] was `mod neon_kernels` / `mod neon_q8k`, flattened
│   ├── avx2.rs                      [S3A] was the inline `*_avx2` bodies
│   ├── neon_correctness.rs          [S3A] was the inline `#[cfg(all(test, aarch64))] mod
│   └── tests.rs                            (exists)
├── vec_ops.rs                       [S3]  module vec_ops: `pub use`
├── vec_ops/
│   ├── vec.rs · rms_norm.rs · rope.rs · softmax.rs · silu.rs   [S3B]
│   ├── f16.rs                       [S3B] the f16 dot/matmul + the nested `mod neon_f16`
│   ├── bf16.rs                      [S3B] bf16 row decode + matmul
│   ├── neon.rs                      [S3B] was `mod neon_vec`
│   └── tests.rs                            (exists)
├── graph/                                   (unchanged: `backend.rs` + `registry.rs` stay the one
│                                             device seam; `*_backend.rs` stay the executors)
└── … (all other modules unchanged)

Non-src/ changes that ride along: build.rs (CUDA/Metal file lists + one rerun-if-changed per .cu/.cuh/.metal/.h + one .o per .cu, still one libcuda_kernels.a; Metal parts → one metallib; plus the new-file guard that a file in kernels/ cannot silently be absent from the list), scripts/check_cuda_launch_returns.py (directory discovery + the "no <<<>>> in a .cuh" assertion) + tests/fixtures/cuda_launch_sites.tsv (regenerated in build-list order, still 130 rows, still four columns), scripts/check_doc_line_anchors.py (new, ticket 6), docs/SOURCE-LAYOUT-PLAN.md (this file) + AGENTS.md Layout block + docs/ARCHITECTURE.md (layer definition + interface-eligibility rule) + the path/anchor sweeps in docs/BACKENDS.md, docs/CUDA-BACKEND-DESIGN.md, docs/inference_e2e_walkthrough/{07,15}*.md, docs/GPU_SAFETY.md, docs/BUILD.md, docs/ARCHITECTURE-ROADMAP.md, docs/CUDA_OPTIMIZATION.md, docs/LLAMA-CPP-MMQ-ANALYSIS.md, docs/ARCHITECTURE-EXECUTION-PLAN.md.

9. Decisions (all taken — nothing open)

#Decision
1No file-size ratchet. The ~800-line target is guidance enforced by review; no script.
2kernels/fa_prefill.metal (812 lines) is accepted — the prefill flash-attention kernel plus its helpers is one unit; llama.cpp splits FA instantiations (template-instances/), not the body, so there is no better boundary here.
3The stale-fact sweep is folded into #219; no separate documentation ticket.
4Step 4 is Mac-local, and this document must be sufficient alone for a macOS agent (§6.5): the file table, the two compile entry points, the verification commands, and an absolute Mac box label in the record.
5The long test files are in scope as Step 6 (split by op family/topic, counts identical).
6The frozen set is approved (§6.2): docs/cuda_optimization_steps/*, docs/QWEN2.5-*.md, docs/DEBUGGING-*.md, docs/KNOWN-CPU-ISSUES-*.md; it may only shrink, and ticket 6's checker carries it with a reason per file.
7docs/cuda_tutorial/* is live (180 of the cuda anchors live there): each example is re-pointed in Steps 1–2, not frozen.

10. Verification matrix

StepCommandsExpected
0 docsscripts/check_docs_links.py, scripts/check_status.py --check, scripts/build_book.shgreen
1 cuda.rscargo build --release --features cuda, scripts/cuda_test.sh, cargo test --release, cargo fmt --all --check, check_source_layout.py, check_dead_code_{annotations,oracle}.py, FEATURES=cuda scripts/real_model_gates.sh ×2565/0/42; 480/0/36 + 10/0/6; 42/0 ×2; all checkers green
2 .cuStep 1's list plus check_cuda_launch_returns.py (+--selftest, --check-fixture), MINFER_TEST_ISSUE162=1 device gate, MINFER_OP_TIMING=1 cold-start record130 sites; 42/0 ×2 bitwise; module-load cost recorded
3 CPUCPU suites (cargo test --release, plus MINFER_NO_NEON=1) + the two-arch dead-code set comparison481/0/36, 10/0/6, sets identical
4 Metalon a Mac: cargo build --release (non-empty metallib), real-model gates, #255's two judgmentsrecorded on the Mac box
5 closeLinux: #225's table re-measured (§5.1); Mac: #53's DeviceMemory for Metalthe measured row in CUDA-BACKEND-DESIGN.md §2.4 / plan §5.1; the common decision resolved on the Mac half (§1.3 rule 3: 2 implementations, 3 callers, no new module)
6 teststhe step-appropriate suite (CUDA 565/0/42 for the executor tests, CPU 480/0/36 + 10/0/6) + check_source_layout.pycounts identical; every new test file named by a mod
allscripts/check_docs_links.py, scripts/check_status.py --check, scripts/build_book.sh, scripts/check_doc_line_anchors.py (once ticket 6 lands)green; no stale anchor

Suite counts are read, not remembered. The numbers in the table above are the pre-#138 rows; the current rows are docs/status.toml (after #138: CUDA 567 / 0 / 42, CPU 481 / 0 / 36 on dgxspark (aarch64, GB10 sm_121) and 479 on the CI runner, integration 10 / 0 / 6). Every step must read the rows from that file rather than quote this document, and a step that moves a row updates AGENTS.md and docs/status.toml in its own PR.

Standing rule for every step: a layout PR moves code and nothing else; anything else it notices is filed, not fixed inside it.

The layering convention and the test-module rule (moved from AGENTS.md)

The backend layers: each device backend is organised as L1 runtime, L2 launch/dispatch, L3 kernel sources and L4 graph executor, and only L4 is polymorphic — graph/backend.rs + graph/registry.rs stay the single device seam, so the directory tree is device-first and the layer is the second axis inside each device (<backend>/kernels/ is where kernel sources live). A shared common is added only when two backends implement it and two callers use it (allocplan::DeviceMemory is the one candidate today). CPU's quants.rs/vec_ops.rs are deliberately not a device-private layer: they are the crate's numeric kernel library, shared with graph/kvformat.rs and graph/cuda_backend.rs. The four long backend files are split along this convention — the three CUDA files, the CPU trio and the two Metal files are done (src/cuda/, src/cuda/kernels/, src/quants/, src/vec_ops/, src/kernel/, src/metal/, src/metal/kernels/). Plan and target tree: docs/SOURCE-LAYOUT-PLAN.md (#261); the file list above is updated by each step as the code moves.

src/graph/: mod.rs ComputeGraph/CNode · ops.rs Op + NodeMeta · builder.rs GraphBuilder · scheduler.rs assign → split → execute · backend.rs + cpu_backend.rs/metal_backend.rs/cuda_backend.rs executors · registry.rs backend registry (F4; F5's copy_cross/await_cross) · alloc.rs liveness allocator + persistent KV regions (E4) · allocplan.rs size-class ladder + pure plan + DeviceMemory/budget_decision · offload.rs layer offload plan (E5) · kvcache.rs cell store, removal/shift/compaction, span list, prefix sharing (C1–C3, C8b) · kvformat.rs KV format + MINFER_CACHE_TYPE gate (C4) · kvsession.rs versioned KV session container (C5) · cache.rs/params.rs params-only graph reuse · fusion.rs SwiGLU fusion · copystats.rs split-boundary counters · batch.rs batch composition · dot.rs/json.rs exporters.

Unit tests live beside their module as <module>/tests.rs, declared #[cfg(test)] mod tests;, so a non-test build does not parse them (e.g. src/graph/alloc/tests.rs for src/graph/alloc.rs). src/graph/op_matrix.rs was already this pattern. Note this buys build hygiene, not speed: a warm cargo check --release --features cuda measured 1.43–1.45 s with the tests inline vs 1.52–1.57 s extracted. Both halves are enforced: scripts/check_source_layout.py (CI check-docs) rejects an inline #[cfg(…test…)] mod … { (the cfg predicate is read, so a compound #[cfg(all(test, …))] is caught too), and rejects a src/**.rs that no mod declaration names — the latter is the quiet one, because an undeclared file is never compiled and the tests inside it would silently not run. A long test module is split further into <module>/tests/<topic>.rs, each topic declared by a mod <topic>; in its tests.rs (src/graph/cuda_backend/tests/ is the first, #267); the same declaration walk reaches those files, so the orphan rule guards the split as well.

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0012 — Device is the first axis, the layer the second — and no premature common
  • ADR-0005 — Metal becomes a first-class backend
  • ADR-0023 — Each machine ledger lives beside its checker, one per prose target

Architecture Decision Records — index

An ADR records a decision, the alternatives it beat, and what it cost — and then stops changing. This directory answers "why is it this way?", a question the design docs cannot answer because they describe the current contract, which is allowed to move.

The boundary rule

This is what keeps the corpus from becoming a fourth copy of the documentation. Each kind of fact has exactly one home:

ContentOne homeMutable?
The decision, the alternatives considered, the accepted consequencesdocs/adr/NNNN-*.mdNo — a change is a new ADR plus Superseded by on the old one
The current contract: capabilities, file formats, semanticsthe design doc (docs/KV-CACHE-DESIGN.md, docs/COMPUTE-GRAPH-DESIGN.md, …), which links Decision: ADR-NNNNYes
Measurements, suite counts, per-ticket historyscripts/test-baselines.toml, docs/TEST-BASELINES.md, docs/ARCHITECTURE-EXECUTION-PLAN.mdYes
The plan's phase counters, next: sentence and baseline commitsscripts/status.toml, docs/ARCHITECTURE-EXECUTION-PLAN.mdYes
Unresolvable path citations in this corpus, each with a reasonscripts/adr-citations.toml, beside its checkerYes
A machine-read record: fixture provenance, an accepted-annotation baselinebeside its readers, never in the book tree — tests/fixtures/f6-fixtures.json, scripts/dead-code-baseline.toml (ADR-0024)Yes

What a map lists. A design document ends with ## Decisions governing this document — its map. It lists the ADRs whose decision that document's current contract depends on (states it, or requires it); naming another document's decision is a cross-reference, which belongs in the prose and not in the map, so a document may legitimately mention an ADR its map omits (ADR-0028).

Three kinds of page carry no map, and that is not an omission — each is a place a fact about a decision lives, not a document whose contract depends on one:

  • a record page: the plan and its per-ticket history (docs/ARCHITECTURE-EXECUTION-PLAN.md), the optimization records (docs/CUDA_OPTIMIZATION.md, docs/METAL_OPTIMIZATIONS.md), and a plan's own root-cause notes (docs/QWEN3-SUPPORT-PLAN.md);
  • a ledger page: docs/TEST-BASELINES.md, whose numbers are machine-checked against scripts/test-baselines.toml;
  • an upstream comparison: docs/LLAMA-COMPUTE-GRAPH.md, which is not this repository's contract.

And a map may omit an ADR that names the page, in two cases — both are references about a document rather than a dependency of it:

  • a correction reference: ADR-0022 names the documents whose citations it fixes (ADR-0029 records the class);
  • a contrast reference: ADR-0037 names docs/CUDA-BACKEND-DESIGN.md §4.3 as a different mechanism, not as a dependency.

Three rules keep it that way. All three are about kind, not about banning a form:

  • A capability is a dated consequence, never a present-tense claim. "At the Date: above, #310 enabled X" is allowed; "X is now true" is not — that is the sentence that rots. The authority for what is true today is the design doc and docs/SUPPORT-MATRIX.md, which are mutable.
  • A citation is dated — a path, a symbol and a count, not only a capability. An ADR may say "the ledgers then lived under docs/"; it may not say "the ledgers live under docs/". Cite the kind of home (the machine ledger, the design doc) or date the reference. An unresolvable path citation is either pinned with a reason in scripts/adr-citations.toml or it fails check_adr.py; telling history from staleness is that ledger's whole purpose (ADR-0027).
  • A number may be evidence, never a baseline. A count that is part of an argument — the rejected alternative's failures, a named boundary — is frozen with the ADR and belongs in it. A current suite baseline belongs in docs/TEST-BASELINES.md and its ledger scripts/test-baselines.toml, which are machine-checked. An ADR that quotes today's suite result is a bug.

Three consequences worth stating outright:

  • Superseding is additive. Correcting a decision means writing a new ADR and setting the old one's Status to Superseded by ADR-NNNN. The old text is never edited — the record of what we believed, and why, is the useful part.
  • Corrections are additive too. A defect in a frozen ADR's text is corrected by a new ADR carrying - Corrects: ADR-NNNN, and the corrected ADR's row below names the corrector — both ends, enforced by check_adr.py. The old text is never edited. See ADR-0022 (three citations at once) and ADR-0026 (a path a follow-through had moved).
  • Numbers are citations. They are dense from 0001, never reused, never renumbered. A number is assigned once, in the order decisions are established when their ADR is written, and the Date: field carries the date the decision was actually taken. Because most of this corpus is backfilled from existing records, a decision discovered later keeps its true (possibly earlier) Date: and takes the next free number — so the sequence is approximately chronological, and Date: is always the authority. That is the cost of keeping numbers stable; stability is worth more here than a perfect chronology.

scripts/check_adr.py (CI job check-docs) enforces the mechanically checkable half: filename and heading agree, dense numbering, one of four Status values, a YYYY-MM-DD date, the required sections, every ADR listed below, Superseded by written from both ends, a Corrects: target that exists, is lower-numbered and is named on its own index row, and every path citation either resolving or pinned in scripts/adr-citations.toml.

Index

Ordered by Date:, because a number is a citation and a date is the chronology. Backfilled ADRs are numbered in the order they were written, so a lower number does not imply an earlier decision — see the numbering rule above. Within one date, the order is by number.

#DateDecisionStatus
00072026-06-24No ML frameworks: every operator is hand-writtenAccepted (citation corrected by ADR-0024)
00132026-06-24The CPU quantizes activations to Q8_0; a device reads f32Accepted
00082026-08-02GPU safety: bounded waits, no early return past a barrier, runtime device limitsAccepted
00012026-08-21Inference runs through one declarative compute graphAccepted (citation corrected by ADR-0022)
00022026-08-21Topology is a function of GraphParams alone, so positions cannot be structureAccepted
00092026-08-21A failure is an error, never a silent fallbackAccepted
00102026-08-28The identity gate: bitwise by default, a named tolerance class otherwiseAccepted (citation corrected by ADR-0022)
00032026-09-16Metal is out of scope for this roundSuperseded by ADR-0005
00042026-09-19The batching default follows the deviceAccepted
00052026-09-20Metal becomes a first-class backendAccepted (supersedes ADR-0003; corrected by ADR-0022)
00142026-09-22A KV session is a versioned, checksummed file — never a memory dumpAccepted
00152026-09-23The offload auto fit takes a prefix, not a knapsackAccepted
00112026-09-24Backend ids are a file-format contract: appended, never renumberedAccepted
00162026-09-24A failed device-memory query is not a zero budgetAccepted
00172026-09-24Speculative decoding refuses the features its identity contract cannot carryAccepted
00182026-09-24The grammar mask is one stage inside the single sampler pipelineAccepted
00192026-09-24A chat template that cannot be rendered refuses the loadAccepted
00062026-09-25The KV storage format is a per-engine gate, not a process-wide globalAccepted
00202026-09-27A quantized file is byte-identical to llama-quantize, or it is wrongAccepted (citation corrected by ADR-0024)
00212026-09-27bf16 is a round-to-nearest-even cast, and 1-D tensors stay f32Accepted
00122026-10-04Device is the first axis, the layer the second — and no premature commonAccepted
00222026-10-09A defect in a frozen ADR is corrected by a new ADR, not by editing itAccepted (corrected by ADR-0023 and ADR-0027)
00232026-10-09Each machine ledger lives beside its checker, one per prose targetAccepted
00242026-10-09A machine-read record is not book contentAccepted
00252026-10-09A docs-only change runs the docs gate, not the compilersAccepted (corrected by ADR-0026)
00262026-10-10The classifier carries no docs-path exception, and its list is ratchetedAccepted
00272026-10-10A citation is dated too — paths, symbols and counts, not only capabilitiesAccepted (corrects ADR-0022)
00282026-10-10A document's ADR map lists what its contract depends onAccepted (corrected by ADR-0029)
00292026-10-10An ADR's prose count is dated too, and its References are map evidenceAccepted (corrects ADR-0028)
00302026-10-10Launch severity lives in the helper, not in 120 call sitesAccepted
00312026-10-10Capture runs in thread-local mode, because Global lets a foreign thread's call join the windowAccepted
00322026-10-10The packed-KV staging window is f32, not f16Accepted
00332026-10-10The CUDA pool recycles exact byte lengths, never frees, and reports OOM as an errorAccepted
00342026-10-10The smem opt-in is an eager pre-warm by construction, with the lazy path as defence in depthAccepted
00352026-10-10The q4_K dsc plane is admitted by two gates, and the payload test is equalityAccepted
00362026-10-10bf16 weights get their own device kernels, not a dtype flag on the f16 onesAccepted
00372026-10-10A Metal weight dtype a kernel cannot consume is refused, never run as a wrong kernelAccepted
00382026-10-10The per-tensor registration dispatch is one shared rule, not a copy per loaderAccepted

The template

# NNNN. One-line decision title

- Status: Accepted | Proposed | Rejected | Superseded by ADR-NNNN
- Date: YYYY-MM-DD
- Issues: #NN, #NN            (optional)
- Supersedes: ADR-NNNN        (required when this ADR replaces one)
- Corrects: ADR-NNNN          (when this ADR corrects an earlier one's text; see ADR-0022)

## Context

What the situation was, and what made it a decision rather than a default. Cite the record.

## Decision

What was decided, in one or two sentences, then the concrete facts (symbols, files, flags).

## Alternatives considered

Each alternative and the reason it lost — a measurement, a constraint, a cost. If the record holds
no rejected alternative, say so explicitly rather than inventing one.

## Consequences

What this makes easy, what it makes expensive, and which costs were knowingly accepted.

## References

The design doc or plan section that holds the current contract, and the issues/PRs involved.

Model Support Roadmap — Which Model Families to Support Next

Status: planning (no code written for the architectures proposed below). Recorded 2026-08 after adding DeepSeek-R1-Distill-Qwen support; serves as the decision reference for the next model-family work.

Companion document. docs/ARCHITECTURE-ROADMAP.md covers the system layers (IR, scheduler, allocator, KV cache, batching, backend abstraction). The Prerequisites column below references its backlog item numbers (§3 item N); several Tier 2/3 entries are gated on that work.

Current Coverage

minfer supports the Qwen2/Qwen2.5 graph (general.architecture = "qwen2", including DeepSeek-R1-Distill-Qwen) and the Qwen3 dense graph ("qwen3", 0.6B–32B). All dense models share one graph family; the only attention-level deltas are Qwen3's decoupled head dim and per-head Q/K RMSNorm (Op::QkNorm). Backend kernels are architecture-agnostic and reused as-is — but only 2 of the 8 quantized CPU dot products have AVX2 kernels (docs/ARCHITECTURE-ROADMAP.md §2.7), so "reused as-is" is not the same as "equally fast everywhere".

Selection Criteria

Ranked by (a) ecosystem weight — HF download trends as of 2026 put Qwen, DeepSeek, Llama, GLM, Gemma in the first tier — and (b) how much of the existing Qwen2/Qwen3 graph the family reuses.

Criterion (b) only dominates the cost estimate for parameter-isomorphic families. For anything needing new IR structure, the reuse is small and the cost is dominated by the touch sites listed below — see the next section before reading the effort column in the tier tables.

Cost model: what a port actually costs

Parameter-isomorphic (the cheap case)

The bulk of the work is model logic: add models/<name>/{mod,loader,graph}.rs and dispatch in models/mod.rs::load_model(). The shared matmul / RMSNorm / attention kernels and the backend scheduling are untouched. The Tier 1 effort estimates below assume exactly this.

Four items are not covered by that estimate and apply to every port:

  1. Decode-fusion weight registration. The decode path fuses QKV (blk.{i}.attn_qkv) and gate+up (blk.{i}.ffn_gu) only when the loader has registered the concatenated tensors — once for Metal and once for CUDA (models/qwen2/loader.rs:425, :488; models/qwen3/loader.rs:426, :468) — plus the matching qkv_concat_available / gu_concat_available predicates (models/qwen2/graph.rs:288-326). Skip them and the port is still correct, but decode fusion silently disables: no error, just a slower decode path.
  2. RoPE style. rope_style is hard-coded to NonInterleaved in both loaders (models/qwen2/loader.rs:154, models/qwen3/loader.rs:172). RopeStyle::Interleaved exists and the CPU implements it (graph/cpu_backend.rs:491-492), but CUDA refuses it (graph/cuda_backend.rs:1324) and no loader ever selects it. A family needing interleaved RoPE (Llama 1/2, Mistral 7B v0.1) needs a loader change and a CUDA kernel, or must be converted to the non-interleaved convention.
  3. RoPE scaling. Only *.rope.freq_base and *.rope.frequency_scale are read (models/qwen2/loader.rs:146-149). The long-context scaling types (rope.scaling.type = llama3, i.e. the low/high-frequency factors in Llama 3.1+) are not implemented, so a Llama 3.x port is correct at short context and wrong past the original training window until this lands.
  4. Chat template. minijinja 2.21 exposes no str methods natively; since F7 (#50) minfer installs an unknown-method hook implementing the Python str methods with CPython semantics, so a port whose template uses them renders as published. A template using a construct outside the implemented set is refused loudly at load (docs/CHAT-TEMPLATE-AND-TOKENIZER-DESIGN.md), never silently replaced by a generic prompt.

New structure (the expensive case)

Anything that needs a new operator shape pays a five-site change (docs/ARCHITECTURE-ROADMAP.md §2.1): the builder constructor, the allocator's special case, the scheduler's kv_pair resolution, every backend's supports_op + execute_node, and both models' build_graph. The existing decode fusions (FusedQKV, QkvBiasRopeStore, FusedFFN, FusedQkvNorm) are four worked examples of this tax.

The counterpart fix is ARCHITECTURE-ROADMAP.md §3 item 7 (strided views with allocator-known aliasing, plus multi-output nodes): with it, the new-structure cases below can be expressed as compositions and derived by the fusion pass instead of being hand-written per backend.

Tier 1: Nearly Isomorphic to Qwen2 (parameter mapping + the prerequisites below)

ArchitectureDelta vs Qwen2PrerequisitesEffort
Llama 3.1/3.2/3.3/3.4no attention bias (optional in the IR), RoPE variant, SwiGLU gate/up orderRoPE scaling (cost model #3) — short context works without it, long context does not; concat-weight registration (#1)smallest + rope work
Mistral 7Bsame as Llama 3 (no bias)Interleaved RoPE for v0.1 (cost model #2): loader + CUDA kernel; or convert to the non-interleaved conventionsmall–medium
InternLM2~noneconcat-weight registration (#1) onlysmallest
Phi-3/Phi-4qkv bias, RoPE variant, no normqkv bias is already supported (models/qwen2/loader.rs:403); verify the RoPE variantsmall
GLM-4-9Bdense, minor attention detailsnone known beyond #1small
Gemma 2GeGLU, shared QKV layer, alternating SWAGelu op (absent from Op, graph/ops.rs — GeGLU needs it) and a sliding-window mask in the attention kernels (adjacent to ARCHITECTURE-ROADMAP.md §3 item 2, but a distinct parameter — a window bound rather than a cell set); GeGLU itself is a five-site change unless §3 item 7 lands firstmedium

Tier 2: New Operators Needed (see prerequisites before starting)

1. Qwen3-MoE (30B-A3B / 32B / 235B-A22B) — highest structural value

  • The Qwen3 dense graph already contains all attention logic (QkNorm, decoupled head dim) — fully reused.
  • Three additions: the router (ffn_gate_inp linear + top-k softmax), 3-D expert weight indexing ([n_embd, n_ff_exp, n_expert] layout), and a moe_ffn operator.
  • Prerequisite: ARCHITECTURE-ROADMAP.md §3 item 7. The IR has no mul_mat_id equivalent and no multi-output node, so 3-D expert indexing is not expressible today; moe_ffn would have to be a bespoke op at the five-site cost. With item 7 landed, the three additions above are compositions.
  • Watch out for expert_weights_scale (1.0 for 30B-A3B, 0.5 for 235B-A22B).
  • 30B-A3B activates only 3B params; Q4_K_M is ~17–18 GB — runs on M4 Pro, which makes it the natural first structural target once item 7 lands.
  • Reference: llama.cpp build_moe_ffn (src/models/qwen3moe.cpp).

2. DeepSeek-V3/R1 (MLA + MoE)

  • MLA is a KV-cache revolution: per token it stores only the compressed latent (kv_lora_rank 512 + RoPE 64), not n_kv_heads × hd — a

    10× KV footprint reduction at long context.

  • Needs the wq_a / wq_b / wkv_a_mqa / wkv_b weight chain, a new KV cache shape, and a matching attention kernel.
  • Prerequisites: §3 items 1, 2, 7 — the KV layout change is exactly the case the current fixed per-layer K/V regions cannot express (ARCHITECTURE-ROADMAP.md §2.4), and the latent split needs views.
  • V3 Q4_K_M ~20 GB — marginal on M4 Pro; R1's inference popularity makes it high value.

3. Qwen3-Next (hybrid SWA + MoE) — after the above

  • Adds a sliding-window mask on top of the Qwen3 graph; the flash-attention kernels need mask support (the same prerequisite as Gemma 2, so the two share one piece of work).
  • Prerequisites: the SWA mask (see Tier 1, Gemma 2), plus MoE (§3 item 17) and §3 item 7.

Tier 3: Large Architectural Deltas (new kernels, high cost)

  • Gemma 3 / Qwen3-VL: multimodal (vision encoder) — minfer has no multimodal framework; highest cost.
  • Llama 4 Scout: MoE + interleaved attention; open weights but special training.
  • RWKV / Mamba / Jamba: SSM recurrence, entirely different kernels, and a KV/memory model that is neither the current per-layer regions nor the cell store proposed in ARCHITECTURE-ROADMAP.md §2.4.

Suggested Order

Two tracks, because the gates are the whole point: the unblocked ports can start today, and the structural ones should not be scheduled before their prerequisites.

Unblocked today (Tier 1):
  1. Llama 3.x dense          ← widest ecosystem coverage; add RoPE scaling
  2. GLM-4-9B                 ← easy; no known prerequisite
  3. InternLM2 · Phi-3/4      ← easy; same wave as 1–2

Gated on ARCHITECTURE-ROADMAP work:
  4. Qwen3-MoE (30B-A3B)      ← §3 item 7 (IR views / multi-output)
  5. DeepSeek-V3 / R1 (MLA)   ← §3 items 1 + 2 + 7
  6. Gemma 2 · Qwen3-Next     ← sliding-window mask (+ §3 item 17 for Qwen3-Next)

The single highest-leverage system-layer item for this document is ARCHITECTURE-ROADMAP.md §3 item 7: it unlocks Qwen3-MoE, is a prerequisite for MLA, and removes the five-site tax from every future structural port. The second is the sliding-window mask, which Gemma 2 and Qwen3-Next share.

Qwen2 / Qwen2.5 Architecture Support

Status page for minfer's support of the Qwen2 architecture (Qwen2 and Qwen2.5 dense models). Qwen2 is the engine's oldest and most-verified model family — the reference implementation for "adding an architecture" in docs/ARCHITECTURE.md §5, and the model every optimization campaign was measured on. The Qwen3 counterpart is docs/QWEN3-SUPPORT-PLAN.md.

1. What the architecture is

A Qwen2 decoder layer (as built by src/models/qwen2/graph.rs, walkthrough diagram in docs/ARCHITECTURE.md §4.6):

h ─ RMSNorm ─ WQ/WK/WV matmuls + bias ─ RoPE(Q, K) ─ KV store
  ─ GQA attention (Q·Kᵀ·scale → softmax → ·V) ─ WO matmul ─ + residual
  ─ RMSNorm ─ gate/up matmuls ─ SiLU(gate)·up ─ down matmul ─ + residual

Distinguishing properties a loader/graph must get right:

PropertyQwen2 valueWhere in minfer
NormalizationRMSNorm, pre-norm, f_norm_rms_eps from GGUF (qwen2.attention.layer_norm_rms_epsilon, loader.rs:143)HParams.f_norm_rms_eps (loader.rs:19)
FFN activationSwiGLU (silu(gate) · up), no biasfused SwiGLU node
AttentionGQA — n_head query heads share n_head_kv K/V headsattn_heads GQA mapping (walkthrough 11 §3.2.4)
QKV biasespresent on WQ/WK/WV (none on WO, none in FFN)QKVBiasRopeStore decode fusion needs them
Output-head biasoptional output.bias tensorQwen2Model.output_b: Option<Tensor> (mod.rs:19, consumed at graph.rs:267-271)
RoPENonInterleaved pairs (i, i+hd/2) — GGUF's NEOX conventionhardcoded RopeStyle::NonInterleaved (loader.rs:154)
Attention scale1/√n_embd_headHParams::attention_scale() (loader.rs:36-38)
KV head dimmay differ from query head dim (n_kv_embd decoupled)HParams.n_kv_embd — Qwen2.5-0.5B: kv_dim=128 vs n_embd=896 (loader.rs:25-28)

2. Where the implementation lives

src/models/qwen2/
├── mod.rs     # Qwen2Model + ModelDef impl (forward / build_graph / forward_graph)
├── loader.rs  # GGUF tensor loader + HParams (standard qwen2.* metadata keys)
└── graph.rs   # the compute-graph builder — deterministic in GraphParams
  • Dispatch: models/mod.rs::load_model() selects Qwen2 by general.architecture == "qwen2" (Qwen3 has its own module).
  • HParams fields (loader.rs:11-29): n_embd, n_head, n_head_kv, n_layer, n_ff, n_vocab, max_seq_len, f_norm_rms_eps, rope_freq_base / rope_freq_scale (read from qwen2.rope.freq_base, falling back to llama.rope.* keys — loader.rs:146-150), eos_token_id, im_end_token_id, rope_style, n_kv_embd.
  • Per-layer weights (LayerWeights): attn_norm, wq/wk/wv (+ biases), wo, ffn_norm, ffn_gate/ffn_up/ffn_down — all as the GGUF lays them out; weights are never dequantized as a whole (walkthrough 03/10).

3. Graph-shape specifics (Qwen2-flavored)

  • G3 tail-row optimization — when n_out < nt, the graph reduces the hidden states to the last n_out rows before the final layer's FFN, so the tail FFN, final RMSNorm and lm_head all run on n_out rows only (graph.rs:61-64, 214-222). With tied embeddings (Qwen2.5-0.5B: output is token_embd), the output matmul is a n_vocab × n_embd Q4_0/Q8_0 matmul — the single most expensive decode op, which is why the tail cut matters (walkthrough 05 §3, 10 §3.2).
  • Decode fusions — Op::FusedQKV (concat WQ/WK/WV matmul + 3 biases + 2 RoPEs + KV store in one node) exists because Qwen2 has QKV biases; Op::FusedFFN (gate+up concat + SwiGLU). Both are env-revertable (MINFER_NO_FUSE_QKV=1 / MINFER_NO_FUSE_FFN=1) and bit-identical when fused; backend support is per-supports_fused (see docs/BACKENDS.md §3).
  • RoPE style is model-level: NonInterleaved is the Qwen2 family convention. Mixing it with the Llama-style Interleaved layout produces plausible-but-wrong scores with nothing crashing — it was debugging suspect #1 during Qwen2.5 bring-up (docs/DEBUGGING-PLAN.md H1).

4. Verified models

Per docs/SUPPORT-MATRIX.md and the AGENTS verified list (CPU + graph-GPU backends; greedy output matches llama.cpp where noted):

ModelQuants verifiedNotes
Qwen2.5-0.5BQ4_0, Q4_K_M, Q5_K_M24 layers · n_embd=896 · 14 query heads / 2 KV heads (GQA 7:1) · hd=64, n_kv_embd=128 · tied embeddings · KV cost 24 KB/token across all layers (walkthrough 11 §2.2)
Qwen2.5-7BQ4_K_Mthe 7B decode-bandwidth reference (28 GB→4.4 GB quantization argument, walkthrough 10 §2.1)
Qwen2.5-1.5Bhistoricalhad a dedicated debugging era — docs/QWEN2.5-1.5B-BUGS.md, docs/QWEN2.5-DEBUGGING-NOTES.md (predates the graph refactor; the kernel fixes it drove — Metal attention hd=128 overflow, Q4_K interleaved scale/min layout, Q6_K embedding-scale indexing — are folded into src/metal/kernels//quants.rs and regression-tested)
DeepSeek-R1-Distill-Qwen-1.5Bworksneeds the tokenizer special-token match (same Qwen2 architecture)

Any GGUF with general.architecture = "qwen2" that stays inside the supported quant set (docs/SUPPORT-MATRIX.md) should load; unverified sizes are untested rather than known-broken.

5. Tokenizer and chat

  • Byte-level BPE, self-contained from GGUF metadata (src/tokenizer.rs).
  • Stop tokens: SpecialTokens { eos, im_end } (src/models/mod.rs:88-91) — Qwen2's <|im_end|> id is read from GGUF (HParams.im_end_token_id) and stops generation alongside EOS.
  • Chat template: the GGUF's Jinja template is rendered with minijinja (src/template.rs); on missing/invalid templates minfer falls back to multi-message ChatML (template.rs:10-14, :119-130) — which is Qwen2.5's native format, so the fallback is semantically safe here. System prompt default: "You are a helpful assistant.".

6. Context, KV sizing, and backend notes

  • max_seq_len comes from GGUF (qwen2.context_length); the KV cache is sized by --n-ctx, not by the model's training context — 0.5B costs 24 KB of KV per position (all 24 layers), so --n-ctx 4096 ≈ 96 MB of persistent regions (walkthrough 07).
  • KV positions are data: the graph is identical for prefill, decode, and multi-turn continuation; positions[i] < n_ctx is enforced with a hard Err in the KV-store arm (walkthrough 11 §3.2.3).
  • Quant support per backend: docs/SUPPORT-MATRIX.md (CPU activations are Q8_0, Q8_K for K-quant weights; GPU reads f32; CUDA prefill int8 MMQ).
  • CPU-vs-GPU logits differ by design; each path is compared against its own llama.cpp reference.

7. Troubleshooting checklist

When a Qwen2-family model produces incoherent output, check in this order (every item here is a real bug class from the Qwen2.5 bring-up):

  1. RoPE style — NonInterleaved must be in effect (§3); wrong style keeps magnitudes plausible.
  2. Quant block layouts — Q4_K scales/mins are interleaved within the 12-byte field (docs/QWEN2.5-DEBUGGING-NOTES.md Bug 4); Q6_K embedding scale indexing was its own bug (Bug 3/5). Both are covered by cargo test quants:: parity tests now.
  3. Tokenizer special tokens — chat-mode garbage on distilled models is usually an EOS/im_end mismatch, not a kernel bug.
  4. Head-dimension limits on GPU — hd=128 models need the widened Metal attention registers (the oc[32] fix, Bug 1); guard failures abort with values per docs/GPU_SAFETY.md.
  5. --n-ctx — a context larger than the KV allocation is rejected by the store arm with Err, not truncated silently.

8. Adding a Qwen2-like architecture

Qwen2 is the worked example for new architectures: mirror src/models/qwen2/{mod,loader,graph}.rs with your HParams + LayerWeights, wire tensor names, set the correct RopeStyle and attention_scale, register in models/mod.rs::load_model(), and keep build_graph deterministic in GraphParams (the reuse invariant). Details: docs/ARCHITECTURE.md §5, docs/COMPUTE-GRAPH-DESIGN.md.

Qwen3 Support Plan (dense architecture)

STATUS: IMPLEMENTED (2026-08-23). Dense Qwen3 is fully supported on CPU and Metal GPU. See §6 Implementation Record below; the rest of this document remains the design record the implementation followed.

Status at the time of writing — historical; superseded by the STATUS banner above. Qwen3 was not supported then: load_model() (src/models/mod.rs) only dispatched on "qwen2", and a general.architecture = "qwen3" GGUF failed with Unsupported architecture: 'qwen3'.

This document analyzes the Qwen3 dense architecture against minfer's existing Qwen2 support (which already implements all the primitives Qwen3 needs) and gives a phased implementation plan. Reference: llama.cpp's upstream Qwen3 implementation.

Reference model: Qwen3-0.6B-Instruct-Q8_0.gguf (Qwen/Qwen3-0.6B-GGUF, Q8_0, file_type 7, quantization_version 2).


1. What the model is

KeyValueminfer impact
architectureqwen3 (dense — no MoE, no sliding window)new dispatch branch
block_count28—
embedding_length1024n_embd
head_count / head_count_kv16 / 8 (GQA)n_head / n_head_kv
key_length / value_length128 / 128n_embd_head — decoupled from n_embd/n_head = 64
feed_forward_length3072n_ff (SwiGLU)
rope.freq_base1 000 000freq_base (NeoX / NonInterleaved RoPE)
layer_norm_rms_epsilon1e-6f_norm_rms_eps
context_length40960max_seq_len / n_ctx
tokenizergpt2 model, pre = "qwen2"minfer's BPE tokenizer works unchanged
eos / bos / pad151645 (`<im_end
chat_templateQwen3 ChatML + <think> tagsrenders since F7 (#50, 2026-09-24): the Python str methods (.split()/.lstrip()/.rstrip()/.strip()) are provided through minijinja's unknown-method hook, so the model's own template is used and the ChatML fallback is gone (see §5 gotcha #9)
vocab151936 (÷32 ✓ for Q8_0 blocks)—
weightsall Q8_0 (q/k/v/wo/gate/up/down) + f32 normsQ8_0 fully supported (CPU + Metal)

Tensor inventory per layer (from the GGUF):

blk.{i}.attn_norm.weight    f32   [1024]
blk.{i}.attn_q.weight       q8_0  [1024, 2048]      # 16 × 128
blk.{i}.attn_k.weight       q8_0  [1024, 1024]      #  8 × 128
blk.{i}.attn_v.weight       q8_0  [1024, 1024]
blk.{i}.attn_output.weight  q8_0  [2048, 1024]
blk.{i}.attn_q_norm.weight  f32   [128]   # NEW vs Qwen2 — per-head Q RMSNorm
blk.{i}.attn_k_norm.weight  f32   [128]   # NEW vs Qwen2 — per-head K RMSNorm
blk.{i}.ffn_norm.weight     f32   [1024]
blk.{i}.ffn_gate/up/down    q8_0  [1024,3072] / [1024,3072] / [3072,1024]

Gotcha: both the minfer info listing and the original metadata dump truncate the tensor list, so attn_q_norm / attn_k_norm do not appear in them. They are in the file (56 occurrences, verified with strings model.gguf | grep attn_q_norm). Do not "fix" the file or skip the tensors — the loader must load them or the model produces garbage.


2. Architecture deltas vs Qwen2.5 (the entire difference)

Everything else — RMSNorm pre-norm, SwiGLU FFN with ffn_norm, no biases, GQA attention, NeoX (non-interleaved) RoPE, tied lm_head — is byte-identical in structure to Qwen2.5 and already implemented in minfer. Only two things differ:

2.1 Decoupled head dimension (n_embd_head = 128, not n_embd/n_head = 64)

Qwen3 decouples head_dim from hidden_size / n_head (HF config head_dim). Consequences:

  • attn_q projects to n_head × 128 = 2048, attn_k/attn_v to n_head_kv × 128 = 1024, attn_output maps 2048 → 1024.
  • RoPE rotates the full 128-dim head (llama.cpp: GGML_ASSERT(n_embd_head == n_rot)), freq_base = 1e6.
  • QK scale is 1/sqrt(128).
  • KV cache row is 1024 floats wide (n_kv_embd = 1024, hd_kv = 128).

llama.cpp derives this generically in llm_load_hparams (src/llama-model.cpp:1219-1226):

hparams.n_embd_head_k_full = hparams.n_embd / hparams.n_head();
ml.get_key(LLM_KV_ATTENTION_KEY_LENGTH, hparams.n_embd_head_k_full, false);   // "qwen3.attention.key_length" = 128
hparams.n_embd_head_v_full = hparams.n_embd / hparams.n_head();
ml.get_key(LLM_KV_ATTENTION_VALUE_LENGTH, hparams.n_embd_head_v_full, false); // "qwen3.attention.value_length" = 128
hparams.n_rot_full = hparams.n_embd_head_k_full;                              // no "qwen3.rope.dimension_count" → n_rot = 128

minfer's Qwen2 HParams::n_embd_head() computes n_embd / n_head and the loader only overrides n_kv_embd from the K weight — for Qwen3 we must also carry an explicit n_embd_head read from qwen3.attention.key_length (and value_length, same value here).

2.2 Per-head Q/K RMSNorm (attn_q_norm, attn_k_norm)

New for Qwen3: after the Q and K projections (no biases), each head's hd elements are RMS-normalized and scaled by a per-head weight of length hd, before RoPE:

Qcur = Wq @ x                        # [nt, 2048]
Qcur = reshape [128, 16, nt] → rms_norm(eps=1e-6) per (head, token) → × attn_q_norm[128]
Qcur = rope(Qcur, n_rot=128)         # NeoX

llama.cpp reference (src/models/qwen3.cpp graph):

  • load_arch_tensors creates attn_q_norm / attn_k_norm with shape {n_embd_head_k} (qwen3.cpp:33-34).
  • graph: build_qkv(...) reshapes Q/K/V to 3D [n_embd_head, n_head, n_tokens] (src/llama-graph.cpp:build_qkv), then build_norm(Qcur, attn_q_norm, NULL, LLM_NORM_RMS, il) — i.e. ggml_rms_norm over ne[0] = 128 per (head, token), then multiply by the weight — then ggml_rope_ext with n_rot = 128 (qwen3.cpp:78-100).

In minfer's flat token-major activation layout this is a contiguous [nt·n_head][hd] matrix, so it is exactly the existing RMSNorm kernel with d = hd, n = nt·n_head, weight length hd — no new math, just a new op that encodes the row grouping (n_heads instead of 1).

The rest of qwen3.cpp (FFN, output) is a standard rms_norm → swiglu(gate,up) → down → add + output_norm → lm_head, identical to Qwen2.


3. Implementation plan

Phased, correctness first; each phase leaves the tree compiling and the model run-able.

Phase A — models/qwen3 scaffolding (mirror models/qwen2)

  1. src/models/qwen3/mod.rs
    • Qwen3Model { hparams, tok_embd, output_norm, output, output_b, layers } and impl ModelDef — copy qwen2/mod.rs (forward routes through the graph; special_tokens from hparams; accessors use explicit n_embd_head).
    • tensor_names: copy + add attn_q_norm(i) = "blk.{i}.attn_q_norm.weight", attn_k_norm(i) = "blk.{i}.attn_k_norm.weight".
  2. src/models/qwen3/loader.rs
    • HParams: same fields as Qwen2 plus explicit n_embd_head: i64 (from qwen3.attention.key_length, fallback n_embd/n_head); n_embd_head() returns the stored value; attention_scale() = 1/sqrt(n_embd_head).
    • Read keys with the qwen3. prefix (qwen3.embedding_length, qwen3.attention.head_count[_kv], qwen3.block_count, qwen3.feed_forward_length, qwen3.context_length, qwen3.attention.layer_norm_rms_epsilon, qwen3.rope.freq_base, qwen3.rope.frequency_scale → default 1.0).
    • LayerWeights: add q_norm: Option<Tensor>, k_norm: Option<Tensor>.
    • Load blk.{i}.attn_q_norm.weight / attn_k_norm.weight (f32, [128]); register on Metal/CUDA like other weights (f32 registration already exists in load_tensor).
    • Resolve n_kv_embd from blk.0.attn_k.weight ne[1] (= 1024) and call set_kv_cache_type(n_layer, 1024) after that (28×1024 = 28672 ≥ 8192 → f16 KV auto-pick; using the naive default 8×64 = 512 would still give f16, but be correct).
    • Keep the QKV concat (blk.{i}.attn_qkv) and FFN concat (blk.{i}.ffn_gu) GPU registration from qwen2 — FFN fusion is reused as-is; QKV fusion is gated off in Phase C.
  3. src/models/mod.rs — add pub mod qwen3; and dispatch "qwen3" => qwen3::loader::load(model).

Phase B — new graph op Op::QkNorm { hd, nh }

Per-head RMSNorm is not expressible with the existing RmsNorm (it normalizes the whole row). Add a dedicated op; the kernels themselves are the existing norm kernels with a different row grouping:

  1. src/graph/ops.rs — add variant QkNorm { hd: usize, nh: usize } (eps comes from RmsNorm-style meta; reuse NodeMeta::Norm { weight_name, bias_name }); add to NodeMeta mapping (out_shape = in_shape).
  2. src/graph/builder.rs — pub fn qk_norm(&mut self, x, weight_name, hd, nh, eps) -> NodeId.
  3. src/graph/cpu_backend.rs — execute: the input buffer is [nt · nh · hd] floats and the norm rows are contiguous (t·(nh·hd) + h·hd), so run the existing loop with d = hd, n = nt·nh, using vec_ops::rms_norm_fused_f32(hd, dst, row, w.data_f32(), eps) (weight length hd). Add a supports_op arm (CPU: always true).
  4. src/graph/metal_backend.rs — QkNorm arm → cb.rms_norm_256(x, w, w_off, y, hd, nt·nh, eps, 0) (fall back to rms_norm if the 256-thread kernel is disabled); supports_op arm. The MPS kernel takes (d, n) — no shader change needed. Add to supports_op gates so backend assignment works.
  5. Fusion pass (src/graph/fusion.rs): QkNorm is not a fusion target — ensure the matcher leaves it untouched (default no-op is fine).

Math to match (llama.cpp): y = x · rsqrt(mean(x²) + eps) · w with eps = f_norm_rms_eps = 1e-6.

Phase C — graph wiring (src/models/qwen3/graph.rs)

Copy qwen2/graph.rs and change only the attention spine:

#![allow(unused)]
fn main() {
let hd  = hp.n_embd_head as usize;            // 128 (NOT n_embd/n_head = 64)
let nkt = hp.n_kv_embd as usize;              // 1024
let hd_kv = nkt / nk;                         // 128
let attn_scale = hp.attention_scale();        // 1/sqrt(128)

// per layer, unfused path (decode AND prefill):
let q = b.matmul(normed, l.wq, None);
let k = b.matmul(normed, l.wk, None);
let v = b.matmul(normed, l.wv, None);
let q = b.qk_norm(q, "blk.{i}.attn_q_norm.weight", hd, nh, eps);
let k = b.qk_norm(k, "blk.{i}.attn_k_norm.weight", hd, nk, eps);
let q = b.rope(q, inp_pos, RopeStyle::NonInterleaved, RoPEMeta { freq_base: 1e6, freq_scale: 1.0, n_head: nh, hd });
let k = b.rope(k, inp_pos, ..., RoPEMeta { n_head: nk, hd });
b.kvcache_store(i, k, v, inp_pos, n_ctx);
let kv = b.kvcache_load(i, nkt, n_ctx, nk);
let attn = b.attn(q, kv, inp_pos, Gqa, AttnMeta { n_head: nh, n_head_kv: nk, hd, hd_kv, nkt, scale });
}
  • Do NOT use the fused decode QKV path (FusedQKV) for Qwen3 initially: the attn_bias_rope_store Metal kernel does bias+rope+store in one pass and cannot express the per-head norm between projection and rope. Gate fuse_qkv = false for Qwen3 (the env toggle MINFER_NO_FUSE_QKV already exists for A/B; make it the default for this arch until Phase E). The unfused 3-matmul path is correct on both backends.
  • FFN unchanged: gate/up/down + ffn_norm; the fused FFN decode path (blk.{i}.ffn_gu concat, nf = 3072 ≤ 16384) is reused as-is.
  • G3 tail-row reduction (get_rows on the last layer + logits), KV-in-allocator persistent regions, params-only reuse, weights_on_gpu, register_graph_weights (add q_norm/k_norm to the lists) — all carried over unchanged.

Phase D — verification

  1. cargo build --release; minfer info <model> parses; add a strings-level check that attn_q_norm/attn_k_norm load (loader logs layer count).
  2. Self-consistency: prefill + decode produce a sensible greedy continuation; prefill logits are deterministic across runs.
  3. CPU vs Metal: with MINFER_GRAPH_DUMP=/tmp/d, compare prefill/decode logits — greedy tokens must agree, logits may differ ~1e1 (CPU quantizes activations to Q8_0, Metal reads f32 — known path difference, see AGENTS.md §9).
  4. Cross-check vs llama.cpp (the authoritative reference for this arch): run the same Q8_0 GGUF in llama-cli (or llama-server) with temp 0, same prompt, and compare greedy token sequences — they must match; compare raw logits (--logits-all / server logprobs) to the minfer Metal graph path (f32 activations on both sides): expect ~1e-2 max abs diff (reduction-order noise), which also validates the per-head norm math.
  5. Tokenizer/template smoke: minfer <model> "hello" without --no-template — the GGUF chat template renders via minijinja (since F7 [#50] the Python str methods are provided, tools is none like transformers passes, and enable_thinking stays undefined → falsy; an unrenderable template is now a loud refusal, not a ChatML fallback).
  6. Add hermetic tests in qwen3/graph.rs modeled on the qwen2 tests (graph_logits_match_forward_real_model, forward_cached_isolates_kv_between_caches) using the cached Qwen3 GGUF; update the model matrix in AGENTS.md / README.md.

Phase E — follow-ups (explicitly out of scope for the initial port)

  • Fused QKV decode with qk-norm: extend the Metal attn_bias_rope_store kernel (or add attn_qk_norm_rope_store) so nt==1 decode gets the single concat matmul + one-pass norm+rope+store back for Qwen3. Until then Qwen3 decode pays 3 matmuls (same as Qwen2 prefill path).
  • Thinking-tag output: Qwen3 (instruct) emits <think>…</think> blocks; optionally strip/collect them like llama.cpp's --reasoning-format — a CLI/formatting concern, not an engine change.
  • Other Qwen3 family members (different arch IDs in llama.cpp): MoE (LLM_ARCH_QWEN3MOE — expert tensors, router), hybrid SWA/MLA (LLM_ARCH_QWEN3NEXT — sliding-window + per-layer dense flags), VL (QWEN3VL*), reranker/embedding (pooling + cls_out), Qwen3.5 (LLM_ARCH_QWEN35*). Each is a separate project; the dense port above already covers Qwen3-0.6B/1.7B/4B/8B/14B/32B.

4. File touch list

FileChange
src/models/mod.rsdispatch "qwen3"
src/models/qwen3/mod.rsnew — Qwen3Model + ModelDef + tensor_names
src/models/qwen3/loader.rsnew — HParams (explicit n_embd_head), LayerWeights + q/k_norm, tensor load + GPU registration
src/models/qwen3/graph.rsnew — Qwen3Graph::build/forward/forward_cached (per-head norm spine, fuse_qkv off)
src/graph/ops.rsnew Op::QkNorm { hd, nh }
src/graph/builder.rsnew qk_norm() builder method
src/graph/cpu_backend.rsQkNorm exec arm (d=hd, n=nt·nh) + supports_op
src/graph/metal_backend.rsQkNorm exec arm (rms_norm_256) + supports_op
docs/ (this plan), AGENTS.md, README.mdmodel matrix / docs

5. Risks & gotchas

  1. Head dim trap: if n_embd_head falls back to n_embd/n_head = 64, the Q/K projections are misinterpreted, RoPE rotates the wrong dims, the attention scale is wrong, and the KV row stride mismatches — silent garbage. Read qwen3.attention.key_length and assert n_kv_embd == n_head_kv · key_length (1024 = 8 × 128) and ne[1](attn_q) == n_head · key_length.
  2. Truncated tensor listing: minfer info and naive metadata dumps hide attn_q_norm/attn_k_norm; the loader must load them by name regardless.
  3. Fused decode QKV cannot express qk-norm — keep fuse_qkv off for Qwen3 or decode logits will silently skip the norm (bitwise different from llama).
  4. RoPE: freq_base = 1e6, freq_scale = 1.0, n_rot = hd = 128, NeoX/NonInterleaved — minfer's cpu_rope/Metal rope_f32 already rotate the full head dim given hd, so only the hd value changes.
  5. KV f16 auto-pick uses n_layer × n_kv_embd; pass the real 1024.
  6. ChatML template — fixed in F7 (#50, 2026-09-24): it used to fall back to ChatML because the Qwen3 chat_template uses Python string-method syntax that minijinja 2.21.0 cannot run; gotcha #9 below records the root cause, and the F7 record (docs/CHAT-TEMPLATE-AND-TOKENIZER-DESIGN.md) records the fix — the Python str methods are supplied through minijinja's own extension point, so the template's think-block extraction, tool-call formatting and enable_thinking handling are live, and an unrenderable template is a loud refusal instead of a silent fallback.
  7. No biases anywhere in Qwen3 — bq/bk/bv/output_b stay None; the loader's optional-bias paths already handle that.
  8. Context 40960: n_ctx defaults from qwen3.context_length; KV region sizing is n_kv_embd × n_ctx = 1024 × 40960 ≈ 40 MB/layer → ~1.1 GB f16 — fine, but keep the f16 KV path (auto-selected).
  9. Qwen3 chat_template is NOT rendered by minijinja (2026-08-27) — FIXED 2026-09-24 in F7 (#50). The historical record: the engine logged chat template rendering failed (unknown method: string has no method named split), falling back to ChatML, because the template (tokenizer.chat_template from the GGUF) is written with Python string-method syntax (message.content.split('</think>'), .lstrip('\n'), .rstrip('\n') at template lines 35-36) and minijinja 2.21.0 strings are Rust strings exposing no str methods. The consequence was that think-block extraction, tool-call formatting, multi-step-tool collapse and enable_thinking were lost and the ChatML fallback was fed back verbatim. The fix (F7): template.rs now installs Environment::set_unknown_method_callback and implements the Python str methods with CPython semantics (a character set for strip/lstrip/ rstrip, split/rsplit with maxsplit, startswith/endswith, replace, case, join, find/rfind/count), plus raise_exception. The model's own template now renders; the acceptance is byte-for-byte against the reference renderings committed under tests/fixtures/chat/ (Qwen2.5-0.5B, Qwen2.5-7B, Qwen2.5-14B, Qwen3-0.6B). Any template the engine still cannot render is a loud error naming the construct and the template line — the silent ChatML fallback is gone.

6.4 Known follow-ups (unchanged from §3 Phase E)

  • Fused QKV decode with qk-norm — DONE 2026-08-27 (Op::FusedQkvNorm + the no-bias kernel_attn_rope_store). We added a new fused decode op that concatenates Wq/Wk/Wv into one matmul and applies the per-head Q/K RMSNorm + no-bias RoPE + KV store in place on the concat buffer (a single op replacing 3 matmul + 2 qk_norm + 2 rope + 2 store). The Qwen2 attn_bias_rope_store path (biases, no per-head norm) is untouched.
  • Qwen3 MoE / hybrid-SWA / VL / reranker variants (LLM_ARCH_QWEN3MOE, QWEN3NEXT, QWEN3VL*) — separate architectures, out of scope here.
  • Optional <think>-block stripping at the CLI layer — the Qwen3 template now renders (F7, gotcha #9), so the template itself decides how a replayed assistant turn's <think> block is re-emitted; CLI-side stripping is no longer needed and was never implemented.

minfer Inference E2E Walkthrough — one run, stage by stage

This series follows one real inference run through the minfer engine, from the moment you type a prompt to the moment generated text streams out — one document per pipeline stage, each with the actual code, the data shapes at that point, and the reasoning behind every design choice.

Who this is for. You can read Rust, but you have never built an LLM inference engine. Every concept (token, embedding, KV cache, attention, quantization, sampling) is defined where it first appears. If you want the compressed version first, read docs/ARCHITECTURE.md; this series is the long version.

Version skew. Line numbers were verified against commit e7fa0da (2026-09-11). Functions move; the file + function name is the stable address, the line number is a convenience.

The run in one paragraph

Numbers in parentheses are chapter numbers of this series — each one below is a link to its chapter; the master table lists them with descriptions.

You type ./target/release/minfer model.gguf "Hello". minfer resolves the model name to a GGUF file, memory-maps it, and parses its metadata and quantized weight tensors (01–02). Metadata dispatches the file to a model implementation (Qwen2/Qwen3) whose weights register into the compute-graph allocator and, when eligible, into the GPU backend (03). Your prompt is rendered through the model's chat template and tokenized into integer ids (04). For those ids the engine builds a declarative compute graph — a pure data structure describing every math op of the transformer — assigns each node to a backend, fuses op patterns, and allocates buffers by liveness with persistent per-layer KV regions (05–08). The prefill forward executes the graph once over all prompt tokens (09): quantized matmuls on CPU (10), RoPE / RMSNorm / GQA attention over the fresh KV (11), producing the last-token logits — a score per vocabulary entry. The sampler turns those scores into one next token (12). From then on the decode loop repeats with a single token per step, reusing the cached graph and the KV accumulated so far (13). On macOS the same graph runs on Metal (14); with --features cuda it runs on NVIDIA GPUs with int8 MMQ prefill and CUDA Graph replay (15).

flowchart LR
    subgraph ACT1["Act 1 — startup & load"]
        D01["01 CLI + model resolution"] --> D02["02 GGUF load (mmap)"]
        D02 --> D03["03 model dispatch + weights"]
        D03 --> D04["04 tokenizer + template"]
    end
    subgraph ACT2["Act 2 — life of the graph"]
        D04 --> D05["05 graph build (IR)"]
        D05 --> D06["06 assign + fusion"]
        D06 --> D07["07 allocator + KV regions"]
        D07 --> D08["08 scheduler + execute"]
    end
    subgraph ACT3["Act 3 — inside one forward"]
        D08 --> D09["09 prefill path"]
        D09 --> D10["10 CPU matmul kernels"]
        D10 --> D11["11 attention + vec ops + KV"]
        D11 --> D12["12 sampler"]
    end
    subgraph ACT4["Act 4 — decode & backends"]
        D12 --> D13["13 decode loop + graph reuse"]
        D13 -.->|next token| D09
        D14["14 Metal backend"] -.->|replaces 10/11| D09
        D15["15 CUDA backend"] -.->|replaces 10/11| D09
    end

Master table — stage → document → code

StageDocEntry code (verified e7fa0da)What happens
CLI parse + model resolution01main.rs (main, subcommand dispatch, GenParams), download/mod.rs::resolveFlags → defaults; local path / hf: / ollama: / cached name → a GGUF path; mode branches (--cnv, serve, viz)
GGUF load02gguf.rs::load_gguf_model, MmapFileParse GGUF v3 (header, metadata KV, tensor table), mmap the data blob zero-copy, merge multi-part splits
Model dispatch + weights03models/mod.rs::load_model, models/qwen2/loader.rs, MpsState::init / CudaState::init_with_gpu, GraphAllocator::register_weightgeneral.architecture → ModelDef; hparams + weight tensors; GPU init; weights registered by name into allocator/backend registries
Tokenizer + template04tokenizer.rs::Tokenizer::load/encode, template.rs::render_templateBPE from GGUF metadata; chat template via minijinja (ChatML fallback); text → token ids
Graph build05graph/mod.rs (IR), graph/builder.rs, models/qwen2/graph.rs::buildPure-IR ComputeGraph: one node per math op, per-layer topology, decode fusions, n_out tail rows; topology = f(GraphParams)
Assign + fusion06graph/scheduler.rs::assign_backends, graph/fusion.rs::runEvery node → the best backend whose supports_op says yes (Metal → CUDA → CPU); the Mul∘Silu→SwiGLU rewrite
Allocate07graph/alloc.rs, graph/cache.rsLiveness-based buffer sharing, in-place aliasing, fill_input_i32, two persistent KV regions per layer; GraphCache owns it all so KV survives rebuilds
Execute08graph/scheduler.rs::split_graph/executeContiguous same-backend splits; cross-backend copies at boundaries; execution in build order; one Metal command buffer per split; CUDA replay hook
Prefill forward09main.rs (prefill block), models/mod.rs::forward → forward_graph_cachedAll prompt tokens through the graph; positions drive KV writes and causal masking; last-token logits out; timing calibers
CPU matmul kernels10kernel.rs, quants.rs, block.rsQuantized weight × on-the-fly Q8_0 activations; AVX2 / NEON+SDOT dots; persistent thread pool; repr(C) block layouts
Attention + vec ops + KV11vec_ops.rs, graph/cpu_backend.rsRMSNorm, RoPE (two styles), softmax, SiLU; GQA attention; kvcache_store/load — positions are data, not structure
Sampler12sampler.rs::sample_with_penaltiesPenalties (repeat/frequency/presence, last-64 window) → top-k → top-p → temperature → seeded sample; stop-string byte matching
Decode loop + reuse13main.rs (decode loop), graph/cache.rs::try_reuse, conversation.rsOne token per step: only input data changes; first decode step rebuilds the graph, KV survives; multi-turn append-only prefill
Metal backend14src/metal/ (MpsState), graph/metal_backend.rs, src/metal/kernels/Same graph on Apple GPU: zero-copy weight buffers, per-op shaders, one command buffer per split, fused decode kernels, flash attention
CUDA backend15graph/cuda_backend.rs, cuda.rs, cuda/kernels/*.cuSame graph on NVIDIA: resident weights, int8 MMQ prefill / MMVQ decode, split-KV attention, CUDA Graph capture/replay, pinned async staging

Reading orders

  • Timeline (default): 01 → 15 in order; dashed arrows in the diagram show where 14/15 swap in for the CPU kernels and where 13 loops back to 09.
  • "I only care about the graph": 05 → 06 → 07 → 08 → 13.
  • "I only care about GPU": 08 (splits) → 14 → 15, with docs/CUDA_OPTIMIZATION.md + docs/cuda_optimization_steps/ for the kernel campaign history and docs/METAL_OPTIMIZATIONS.md for Metal's.
  • "What does X mean?": every doc's §2 defines its terms; the GLOSSARY is the backstop.

Conventions

  • Docs follow STYLE.md (structure, voice, forensics protocol).
  • Cross-refs use repo-relative links; env-var behavior is stated where it matters and collected per doc in §4 (Observe & verify).
  • The series mirrors docs/cuda_optimization_steps/ in form: numbered standalone docs + this index + the shared writing contract.

← Start · 01 — CLI args and model resolution →

01 · CLI args and model resolution

Stage: user types the command → this stage → GGUF file opens (doc 02). Here the shell hands minfer a list of strings; this stage turns that list into (a) typed generation parameters, (b) a chosen run mode, and (c) one concrete filesystem path to a GGUF model file — the only thing the next stage needs. Code: src/main.rs (main, GenParams, print_usage), src/download/mod.rs (resolve, resolve_cached_name, download_hf, download_ollama, list_local).

1. Background — where this stage sits

You type:

./target/release/minfer qwen2.5-0.5b-instruct-q4_0 "Why is the sky blue?" --temp 0.7

and, roughly a second later, text starts streaming out. Everything that happens before the model file is even opened is this stage. The binary starts with three raw ingredients: the process arguments (the strings after the program name), the environment variables, and whatever is in the local model cache. It must end the stage holding exactly the things every later stage takes for granted:

  1. GenParams — the generation parameters as a typed struct: how many tokens to generate, how creative the sampling should be, how big the context is. ("Sampling" means the step where the engine picks the next token from the model's scores; doc 12 does the math, this doc only collects the knobs.)
  2. A run mode — single-shot generation, a multi-turn conversation (--cnv), an OpenAI-compatible HTTP server (serve), or a visualization server (viz). The mode decides what happens after the model loads.
  3. model_path — a path to a real GGUF file on disk.
  4. A prompt — the text to feed the model, taken from the command line, or from stdin, or (for the server modes) not yet existing at all.

The model file is a GGUF file — "GGUF" (GPT-Generated Unified Format) is the single-file container llama.cpp popularized: one file holds the model's metadata (layer counts, tokenizer, chat template) and all of its weights in a compressed, quantized form ("quantized" means each weight is stored in a few bytes instead of a full 4-byte float, so a 0.5-billion-parameter model fits in ~350 MB instead of ~2 GB; doc 02 reads the format, doc 10 reads the compressed numbers). Doc 02 is what opens that file; this stage's whole job is to guarantee it has a valid path to hand over.

The reason a whole stage exists for this is that "the model" is not a path — it is a reference. minfer accepts four spellings of the same idea: a local path, a Hugging Face repo (hf:Qwen/Qwen2.5-0.5B-Instruct-GGUF:q4_0), an Ollama model (ollama:qwen2.5:0.5b), or a bare cached name (qwen2.5-0.5b-instruct-q4_0). The next stage (the GGUF parser, the model dispatcher, the tokenizer) should not know any of that. Resolution is a small name-service layer in front of them: it translates any of the four forms into one PathBuf, downloading if it must, and returns an error with actionable hints when it cannot.

What would break without it? If the GGUF loader had to handle hf: URIs, every later stage would drag network code, cache-layout knowledge, and retry logic around with it. If the flags were parsed ad hoc in each mode, --temp 0.7 would mean different things in single-shot and conversation mode. This stage is also where minfer's promise of comparable behavior lives: the generation defaults are copied from llama.cpp on purpose, so that the same command produces the same distribution of output as the reference implementation — which is what makes the benchmark comparisons throughout this series meaningful.

2. Principle — how it works and why

2.1 Two independent decisions

main() makes two decisions that are easy to confuse because they both read the same argument list:

  • Parse — classify each argument as a flag (starts with -), a flag value, or a positional argument. Positionals assemble into [subcommand?] <model> [prompt…]. There is no argument-parsing crate: the parser is a hand-rolled while loop with a match, because minfer keeps its dependency list tiny (5 crates for the core engine — see Cargo.toml; no clap).
  • Resolve — turn positional[0] into a filesystem path. This happens after parsing and before anything loads, in one call: download::resolve(&model_ref).
flowchart TD
    A["argv strings"] --> B{"arg starts with '-'?"}
    B -->|yes| C["flag → typed value in GenParams / mode vars<br/>(unknown flag → usage + exit 1)"]
    B -->|no| D["positional: subcommand? model? prompt words"]
    C --> B
    D --> B
    B --> E{"positional[0]"}
    E -->|bench / specverify| F["separate parser, exit before global parse"]
    E -->|download / list / info| G["do it, return — no model load"]
    E -->|serve / viz| H["set mode flag, drop the token, keep the model arg"]
    E -->|anything else| I["single-shot / --cnv inference"]
    I --> J["download::resolve(model_ref)"]
    J -->|local path| K["PathBuf (must exist)"]
    J -->|hf: / ollama:| L["download → cached path"]
    J -->|bare name| M["search ~/.cache/minfer/models for *.gguf"]
    K & L & M --> N["model_path — hand this to doc 02"]

The order of the checks inside resolve matters. A path-like reference (starts with /, ., or ~) is taken literally and must exist — no downloads, no name search. The hf: and ollama: prefixes are next. Only if the argument matched none of those forms does minfer try it as a relative path, and then as a cached model name. Trying the cache first would mean a file called model.gguf in the current directory loses to a cached file with the same name; trying the path first makes local files always win, which matches the reader's intuition: the most specific thing you wrote wins.

2.2 What each parameter controls (intuition only)

Sampling — choosing the next token from the model's output scores — is a weighted lottery over the vocabulary (a "token" is one piece of text, a word or word-fragment, represented as an integer id). GenParams collects the lottery's rules. The full math is doc 12; here is what each knob means:

ParameterDefaultWhat it controls, in one sentence
temp0.8Sharpness of the lottery: low values make the best token win almost always, high values flatten the odds. 0 = greedy (always the argmax).
top_k40Before the lottery, keep only the 40 highest-scoring tokens.
top_p0.95Keep the smallest prefix of tokens whose probabilities sum to ≥ 0.95 ("nucleus" sampling).
repeat_penalty1.1Divide the score of any token seen in the last 64 tokens by 1.1 — a mild "stop repeating yourself".
frequency_penalty0.0Penalize each token in proportion to how many times it appeared in the window (0 = off).
presence_penalty0.0Penalize every token that appeared at all in the window (0 = off).
seed42The random number generator's starting point: same seed + same flags + same model ⇒ same output.
n_predict512Hard cap on generated tokens per run.
n_ctx4096Context size — how many tokens the engine makes room for (see §2.4).
stop_strings—Stop generating when this byte string appears in the output (repeatable).

2.3 Why copy llama.cpp's defaults instead of inventing

Three reasons, in order of importance.

Comparability. minfer's development loop constantly compares itself to llama.cpp — greedy output verified token-for-token, throughput measured against llama-bench (see docs/PERF-QWEN3-4B-VS-LLAMACPP.md, and docs 14/15 in this series). If temp or repeat_penalty defaulted to something else, every "minfer matches llama.cpp" claim would need an asterisk listing different knobs. Copying the defaults means identical commands, comparable results: minfer -n 256 --seed 1 model.gguf "prompt" and the equivalent llama.cpp invocation draw from the same sampling distribution, so differences in output are attributable to the engine (kernels, backends), not to the lottery rules. ARCHITECTURE.md §3 states this as policy: "Generation parameters (defaults match llama.cpp)".

Less invention risk. Each default above is a tuned compromise — repeat_penalty 1.1, for instance, is strong enough that small models stop looping ("the the the") yet weak enough that code generation, which wants repeated spaces and braces, still works. Re-tuning that by hand is a research project with no payoff for the reader.

Convention transfer. Anyone who has used llama.cpp, Ollama, or OpenAI's API recognizes these names and value ranges (--temp, --top-k, --top-p, --frequency-penalty). Familiar flags mean the CLI needs no tutorial.

Two deliberate deviations, both visible in the code comments:

  • n_predict: 512 is a finite cap, so a command-line run always ends — convenient for benchmarking and for tests that pipe stdin and wait for the process to exit (llama.cpp's CLI generates until the context is full).
  • seed: 42 is fixed, while llama.cpp defaults to a random seed. A fixed default makes every run with the same flags bit-identical — which is what lets the conversation tests in tests/conversation_cli.rs pipe scripted stdin and assert on the output. Determinism by default is worth more to a from-scratch engine than lottery variety; --seed is one flag away when you want variety.

2.4 What --n-ctx means here (vs what it will mean later)

At this stage, n_ctx is just an integer sitting in a struct — the transformer has not run, so nothing has consumed it yet. Its meaning arrives in two steps:

  • Later (docs 05/07): the compute graph's per-layer KV cache — the model's notepad of intermediate attention states, one K (key) and one V (value) vector pair per token per layer (docs 07 and 11 explain fully) — is allocated with room for n_ctx tokens. n_ctx therefore decides a memory number: for Qwen3-4B (36 layers, KV width 1024, f32) each token of headroom costs 36 × 2 × 1024 × 4 B = 288 KB, so the default 4096 reserves ≈ 1.2 GB while the model's maximum (40960 tokens) would reserve ≈ 11.8 GB — the "12 GB+" the code comment warns about (src/main.rs:744-748, citing docs/PERF-QWEN3-4B-VS-LLAMACPP.md §2).
  • And it is clamped twice, defensively. In main.rs the effective context is params.n_ctx.max(input_ids.len()) — a long prompt must never overflow the notepad — and the model's forward pass clamps again with n_ctx.min(max_seq_len) where max_seq_len comes from the GGUF metadata (src/models/qwen2/graph.rs:398). The CLI flag requests; the model's own context length caps.

So: at the command line --n-ctx is "how much room to reserve"; after doc 07 it is "the size of the persistent KV regions the allocator carved out once, for both prefill and decode". The clamp chain (max with prompt length → min with model max) is the contract that keeps that single number consistent for the whole run.

2.5 Modes: where the same parsed data goes

The parsed state fans out into mode-specific structs, and each mode takes a different subset. In run_conversation (src/main.rs:2008), the conversation gets a ConversationSpec (template, special tokens, seed, n_ctx, system prompt) plus a TurnParams (all the sampling knobs plus stop_strings) — every field is copied out of GenParams, so the conversation REPL and the single-shot loop sample identically given the same flags. The server instead receives n_ctx as the total across all slots (each slot's conversation gets a slice of it — usage text, src/main.rs:128-133), because the server holds several independent KV sets at once.

3. Implementation

3.1 Data in / data out

In: std::env::args() — a Vec<String>, plus environment reads (HOME, MINFER_MODEL_DIR, MINFER_DISABLE_MPS, …).

Out (by the end of the stage, all as owned values in main's stack frame):

ValueTypeProduced by
paramsGenParams (typed defaults + flag overrides)parse loop
mode flagsconv_mode, server_mode, viz_mode, meta_flag, …parse loop + subcommand match
model_pathString — a path that exists on diskdownload::resolve
promptStringpositional join, or one stdin line
KV sizingparams.n_ctx (and the model's n_kv_embd)loaded model + params (the graph's persistent KV regions; the KVCache this row used to name was deleted in #252)

The stage boundary is gguf::load_gguf_model — the first call doc 02 covers. Everything above it in main() is this stage.

3.2 Key code

The parameter struct and its llama.cpp defaults

src/main.rs:50-78 — ten fields, each default annotated with its origin:

#![allow(unused)]
fn main() {
struct GenParams {
    n_predict: usize,
    temp: f32,
    top_k: usize,
    top_p: f32,
    repeat_penalty: f32,
    frequency_penalty: f32,
    presence_penalty: f32,
    seed: u64,
    n_ctx: usize,
    stop_strings: Vec<String>,
}

impl Default for GenParams {
    fn default() -> Self {
        Self {
            n_predict: 512,
            temp: 0.8, // llama.cpp default (sampling, not greedy)
            top_k: 40,
            top_p: 0.95,            // llama.cpp default
            repeat_penalty: 1.1,    // 1.0 = disabled; mild penalty reduces repetition
            frequency_penalty: 0.0, // llama.cpp default (0.0 = disabled)
            presence_penalty: 0.0,  // llama.cpp default (0.0 = disabled)
            seed: 42,
            n_ctx: 4096,
            stop_strings: Vec::new(),
        }
    }
}
}

Read the comments as the design record: the four sampling values that llama.cpp tunes (temp, top_p, repeat_penalty, and the two penalties' off-state) are copied verbatim, and the two "engine convenience" values (n_predict, seed) are minfer's own choice for the reasons in §2.3. stop_strings starts empty because a default stop string would silently truncate someone's output — it is opt-in per run (--stop "USER:", repeatable).

Subcommand dispatch: bench and specverify leave early

src/main.rs:145-162 — the very first thing main does:

fn main() {
    let raw_args: Vec<String> = std::env::args().collect();
    let prog = raw_args[0].clone();

    // `bench` subcommand: parsed separately (its -p/-n/-r/-o flags are
    // bench-local and must not collide with the global inference options,
    // which reject unknown flags below).
    if raw_args.get(1).map_or(false, |s| s == "bench") {
        let code = bench::run(&prog, &raw_args[2..]);
        std::process::exit(code);
    }

    // `specverify` subcommand: D5-1a verify-step cost micro-bench (its -p/-r/-o
    // flags are bench-local too, so it parses separately like `bench`).
    if raw_args.get(1).map_or(false, |s| s == "specverify") {
        let code = spec_verify::run(&prog, &raw_args[2..]);
        std::process::exit(code);
    }

Why the separate parse? The global parser (below) rejects unknown flags with an error — a safety feature. But bench legitimately uses -p / -n with different meanings (bench -p 512 = 512 prompt tokens for the prefill test, per src/bench.rs:38-43) that collide with nothing in the inference flag set yet would still be rejected there, and --n-ctx means something different again (bench sizes KV as pp+tg+16 unless overridden, src/bench.rs:244). Giving bench the raw tail &raw_args[2..] and its own hand-rolled loop (src/bench.rs:62: "Parse (hand-rolled, bench-local flags)") keeps each flag space self-consistent: bench flags stay llama-bench-shaped, inference flags stay inference-shaped, and neither parser needs conditional meanings. std:: process::exit(code) then means the two paths never share state.

The parse loop: flags, values, and the unknown-flag rejection

The loop is ~220 lines of one match (src/main.rs:194-420); the shapes worth seeing are the value-taking flag, the seed/greedy pair, and the fallback:

#![allow(unused)]
fn main() {
        match a {
            "-h" | "--help" => {
                print_usage(&prog);
                std::process::exit(0);
            }
            // ... (every flag arm looks like one of the three below)
            "--greedy" => {
                params.temp = 0.0;          // boolean-style: a sugar flag
                i += 1;
            }
            "--seed" => {
                if let Some(v) = next_val(a) {          // value-taking flag
                    params.seed = v.parse().unwrap_or_else(|_| {
                        parse_err = Some(format!("invalid --seed '{v}'"));
                        0
                    });
                }
                i += 2;
            }
            _ => {
                if a.starts_with('-') && a.len() > 1 {
                    // Unknown option — reject instead of treating as model path.
                    print_usage(&prog);
                    eprintln!("Error: unknown option '{a}'");
                    std::process::exit(1);
                }
                positional.push(raw_args[i].clone());   // a word → positional
                i += 1;
            }
        }
}

(src/main.rs:204-419, condensed.) Three details carry the design:

  • next_val (src/main.rs:517) is a closure that peeks at raw_args[i+1] and records parse_err = Some("missing value for …") if it is absent — the error is remembered and reported after the whole parse, so the user sees one clean message, not an early exit mid-list.
  • A bad value degrades, never crashes: on a failed .parse() the arm sets parse_err but still writes a placeholder (e.g. seed 0), so the rest of the parse continues over well-typed data.
  • Unknown flags are rejected loudly. This is the other half of the bench design: because the global parser refuses anything starting with - that it does not know, a typo like --tem 0.7 cannot silently become the model path. (Note the a.len() > 1 guard — a lone - is treated as a positional, e.g. a file named -.)

Subcommands that stay: download, list, info, serve, viz

src/main.rs:433-524 (condensed; positional[0] decides):

#![allow(unused)]
fn main() {
    match positional[0].as_str() {
        "download" => { /* build the hf:/ollama: URI, call download::resolve,
                           print "Model downloaded: <path>", return;       */ }
        "list"     => { download::list_local()?;  return;  }   // cache listing
        "info"     => { /* resolve + load_gguf_model + dump metadata/tensors,
                           return — no generation                       */ }
        "viz"  => { viz_mode = true;   positional.remove(0); } // fall through
        "serve" => { server_mode = true; positional.remove(0); } // fall through
        _ => {} // fall through to model inference
    }
    let model_path = &positional[0];
}

The split is by side effects: download, list, info finish and return — they never load the model into the graph. serve and viz only set a mode flag and remove their own token, so positional[0] becomes the model and the normal load path continues — the model must be loaded before a server can serve it. The branches themselves are far down the same main (server: src/main.rs:663-686, viz: 688-696, conversation: 698-713), each taking the already-loaded model + tokenizer. This is why a wrong flag in serve mode is still caught by the same global parser — only bench/specverify opted out.

Resolution call site in main

src/main.rs:538-555 — right before the prompt logic, and before any load:

#![allow(unused)]
fn main() {
    // Resolve paths, hf:/ollama: URIs, and cached model names.
    let is_uri = model_path.starts_with("hf:")
        || model_path.starts_with("ollama:")
        || (!model_path.starts_with('/')
            && !model_path.starts_with('.')
            && !model_path.starts_with('~'));
    let model_path = match download::resolve(model_path) {
        Ok(p) => {
            if is_uri {
                eprintln!("Model ready: {}", p.display());
            }
            p.to_string_lossy().to_string()
        }
        Err(e) => {
            eprintln!("Error: {}", e);
            std::process::exit(1);
        }
    };
}

The is_uri predicate re-derives "was this a non-path reference?" so the progress line prints only when resolution might have done work — a local path loads silently, while hf:/ollama:/bare names get a Model ready: /home/you/.cache/minfer/models/… line confirming where the reference landed. Note the same predicate is true for a bare relative filename like model.gguf (no leading //./~), so it gets the "Model ready" line too — harmless, and consistent: anything that went through the name-lookup path announces its result.

resolve: the four sources

src/download/mod.rs:21-53 — the whole dispatcher is 33 lines:

#![allow(unused)]
fn main() {
pub fn resolve(uri: &str) -> Result<PathBuf, String> {
    let cache_dir = default_cache_dir();

    if uri.starts_with('/') || uri.starts_with('.') || uri.starts_with('~') {
        // Local path
        let p = if uri.starts_with('~') {
            let home = std::env::var("HOME").map_err(|e| format!("HOME not set: {}", e))?;
            PathBuf::from(home).join(&uri[2..])
        } else {
            PathBuf::from(uri)
        };
        if p.exists() {
            return Ok(p);
        }
        return Err(format!("File not found: {}", p.display()));
    }

    if let Some(repo) = uri.strip_prefix("hf:") {
        return download_hf(repo, &cache_dir);
    }
    if let Some(model) = uri.strip_prefix("ollama:") {
        return download_ollama(model, &cache_dir);
    }

    // Treat as local path fallback
    let p = PathBuf::from(uri);
    if p.exists() {
        return Ok(p);
    }

    // Bare model name → resolve against the local cache (e.g. `minfer qwen2.5-0.5b-instruct-q4_0`)
    resolve_cached_name(uri, &cache_dir)
}
}

Source (a) local path: ~ is expanded by hand (minfer has no shellexpand crate) and existence is the only check — note it is exists(), not is_file(), so a directory passes here and fails later with a better error (§3.2 "Error UX"). Sources (b) and (c) delegate to the downloaders. Source (d) is the fallback chain: relative path first, then cache lookup. The cache dir itself honors an env override (src/download/mod.rs:7-13): MINFER_MODEL_DIR if set, else ~/.cache/minfer/models.

Cached names: exact, then prefix, with split collapsing

src/download/mod.rs:57-105 — the ergonomics feature:

#![allow(unused)]
fn main() {
fn resolve_cached_name(name: &str, cache_dir: &Path) -> Result<PathBuf, String> {
    let mut paths = Vec::new();
    collect_gguf_paths(cache_dir, &mut paths);        // recursive *.gguf walk

    let mut exact = Vec::new();
    let mut prefix = Vec::new();
    for p in &paths {
        let fname = /* file_name() as String */;
        if fname == name {
            exact.push(p.clone());
        } else if fname.starts_with(name) {
            prefix.push(p.clone());
        }
    }
    let candidates = if !exact.is_empty() { exact } else { prefix };

    match candidates.len() {
        1 => Ok(candidates[0].clone()),
        0 => Err(format!(
            "Model '{}' not found. Use `minfer list` to see cached models, or pass a path, hf:<repo>[:file], or ollama:<model>[:tag].",
            name
        )),
        _ => {
            // If every candidate is a part of ONE split model, resolve to part 0
            // (its split.count drives the loader, which finds the rest).
            /* ... gguf::split_file_info() on every candidate; if all share one
               prefix, return the part-0 path ... */
            Err(format!(
                "Ambiguous model name '{}':\n  {}",
                name,
                candidates.iter().map(|p| p.display().to_string()).collect::<Vec<_>>().join("\n  ")
            ))
        }
    }
}
}

Three ordered outcomes: exactly one candidate (exact matches preferred over prefix matches — qwen2.5-0.5b prefix-matches every Qwen2.5-0.5B quant file, exact matches only the whole name), none (an error whose text tells you the two escape hatches: minfer list, or a URI), and many — where the split-aware rescue applies. A 7B model downloaded from HF lands as …-00001-of-00002.gguf / …-00002-of-00002.gguf (gguf.rs: 1958 parses that pattern into prefix/index/count). Without the rescue, the bare name would always be "ambiguous" for split models; with it, the parts of one model collapse to part 0 — the correct entry point, because doc 02's loader reads part 0's split.count metadata and finds the siblings itself. Only genuinely different models remain ambiguous, and the error lists every candidate path so you can copy-paste one.

The HF downloader: API listing, quant matching, size-checked resume

src/download/mod.rs:210-290 (condensed to the decision loop):

#![allow(unused)]
fn main() {
    let hf_dir = cache_dir.join("hf").join(&repo);          // hf/<org>/<repo>/
    // GET https://huggingface.co/api/models/<repo> → JSON "siblings" list
    let gguf_files: Vec<&HfSibling> = api_resp.siblings.iter()
        .filter(|s| s.rfilename.ends_with(".gguf")).collect();
    let parts = match_model(&filenames, file.as_deref())?;  // pick quant group

    for name in &parts {
        let file_path = hf_dir.join(name);
        let download_url = format!("https://huggingface.co/{}/resolve/main/{}", repo, name);
        // Expected size: prefer the HF API `size`; fall back to a HEAD request
        // (many repos omit `size`), so a complete cached file is skipped, not
        // re-fetched.
        let size = gguf_files.iter().find(|s| &s.rfilename == name).and_then(|s| s.size)
            .or_else(|| head_content_length(&download_url));
        // Skip only when the file exists AND its size matches the remote one —
        // a partial/interrupted download must be resumed, not skipped.
        let complete = file_path.exists()
            && size.map_or(false, |s| {
                file_path.metadata().map(|m| m.len() == s).unwrap_or(false)
            });
        if complete {
            eprintln!("Already cached: {}", file_path.display());
            continue;
        }
        http_download(&download_url, &file_path, size)?;
    }
}

The interesting decision is the idempotency check. "File exists" is not enough — a Ctrl-C'd download leaves a truncated file that would parse as garbage in doc 02. So "complete" means exists and byte-length equals the remote size (from the API listing, or a HEAD request when the repo omits it). The transfer itself is a curl subprocess (src/download/mod.rs:449-473) with -C - — curl's own resume flag, which appends from the current file length — plus -L for HF's redirect chain and --progress-bar. Delegating to curl buys HTTP/2, TLS, retries, and resume without an HTTP-client crate; it is the same trade as the Ollama path (src/download/mod.rs:309-395), which shells out to ollama pull and then symlinks the largest blob from Ollama's own store into ~/.cache/minfer/models/ollama/<model>/model.gguf so both sources share one cache layout.

Quant selection (match_model, src/download/mod.rs:136-205) groups the repo's files by split prefix, matches the requested quant case-insensitively against the group base name (…-q4_k_m), expands an exact part filename to its whole group, and errors with the available choices when the request is missing or ambiguous. Unit tests cover all of those branches (src/download/tests.rs:14-89).

Cache layout and minfer list

~/.cache/minfer/models/          (or $MINFER_MODEL_DIR)
├── hf/
│   └── Qwen/Qwen2.5-0.5B-Instruct-GGUF/
│       └── qwen2.5-0.5b-instruct-q4_0.gguf     (or -0000X-of-0000Y parts)
└── ollama/
    └── qwen2.5:0.5b → model.gguf (symlink into ~/.ollama/models/blobs)

list_local (src/download/mod.rs:524) prints this tree with human sizes ("412.3 MB"), grouped under Hugging Face: and Ollama: headers — and because resolution reads exactly these files, every name it prints is directly usable as the model argument. That is the whole point of source (d): the listing is the autocomplete.

Error UX: directory, missing, unparsable

Resolution can reject a path ("File not found") but it cannot detect every failure — the decisive check is the first read, which is the GGUF loader. src/main.rs:579-620 turns that failure into three distinct diagnoses (condensed):

#![allow(unused)]
fn main() {
    let gguf_model = match gguf::load_gguf_model(std::path::Path::new(&model_path)) {
        None => {
            let p = std::path::Path::new(&model_path);
            if p.is_dir() {
                // A directory is almost always a cache dir with several .gguf
                // candidates — list them instead of a bare "parse GGUF" panic.
                /* read_dir → collect names ending in .gguf, sort */
                eprintln!("Error: {model_path} is a directory — minfer needs a .gguf file path");
                /* + "candidates:" lines, each a copy-pasteable
                   `minfer viz <dir>/<file>` command, or `minfer list` hint */
            } else if !p.exists() {
                eprintln!("Error: file not found: {model_path}");
                eprintln!("       run `minfer list` to see cached models");
            } else {
                eprintln!(
                    "Error: failed to parse GGUF: {model_path} (not a valid GGUF or corrupt)"
                );
            }
            std::process::exit(1);
        }
}

Why the directory case gets its own branch: resolve deliberately accepts any existing path, so a slipped argument — minfer ./models "hi", or the cache directory itself — survives resolution and surfaces here. The branch assumes the common case (a cache directory) and answers with the files you probably meant, each rendered as a ready-to-run command, instead of a bare parse panic. Missing-vs-unparsable keeps the two failure families apart: "you typed a wrong name" (fixable with list) vs "the file is there but not a GGUF" (wrong download, truncated file, corrupt split part) — the latter is exactly the truncated-download scenario the size check in §3.2 exists to prevent.

Where the prompt comes from

src/main.rs:556-575 — after resolution, before loading:

#![allow(unused)]
fn main() {
    // Conversation mode: the positional prompt is the FIRST user turn; stdin
    // is read interactively by the loop (never consume it here).
    let first_prompt = if conv_mode && positional.len() > 1 {
        Some(positional[1..].join(" "))
    } else {
        None
    };
    let prompt = if positional.len() > 1 {
        positional[1..].join(" ")          // every word after the model = prompt
    } else if conv_mode || server_mode || viz_mode {
        // serve / viz / --cnv take no positional prompt and never consume
        // stdin for it. (Bare `minfer viz model` would otherwise hang on
        // read_line until you press Enter — the model+grep host already shows
        // startup, so don't block on a prompt these modes don't use.)
        String::new()
    } else {
        let mut input = String::new();
        std::io::stdin().read_line(&mut input).unwrap_or(0);
        input.trim().to_string()           // single-shot: one line from stdin
    };
}

Because parsing collects all non-flag words, quoting is optional: minfer m.gguf why is the sky blue joins into one prompt. The stdin fallback makes the classic Unix pipe work (echo "hi" | minfer m.gguf), while the server-like modes explicitly take an empty prompt — they would otherwise hang waiting on a stdin that nobody will write. Conversation mode splits the difference: an optional positional becomes the first turn, and further turns come from the REPL's own reader (read_user_input, src/main.rs:1133), never from this one-shot read.

--n-ctx at the moment it starts to matter

The first consumer of params.n_ctx after load (src/main.rs:744-757):

#![allow(unused)]
fn main() {
    // n_ctx (--n-ctx, default 4096) sizes the graph KV regions — NOT the
    // model's max_seq_len, which would allocate 12 GB+ and pay a first-submit
    // Metal tax (docs/PERF-QWEN3-4B-VS-LLAMACPP.md §2). It never shrinks below
    // the prompt length, and the model's forward clamps it to max_seq_len.
    // Computed ONCE so prefill and decode size the same KV regions.
    let ctx = params.n_ctx.max(input_ids.len());
    /* ... */
    let logits = model.forward(&input_ids, &positions, &mut kv_cache, 1, ctx);
}

and the second clamp inside the model (src/models/qwen2/graph.rs:398):

#![allow(unused)]
fn main() {
        let n_ctx = n_ctx.min(model.hparams.max_seq_len as usize);
}

The comment is the invariant in words: one number, computed once, clamped from below by the prompt and from above by the model's own context length, used for both the prefill forward and every decode step — because the KV regions are allocated once and reused (docs 07/13). A single-shot run and its own decode loop must agree, or the second forward would rebuild a graph with different KV geometry mid-generation.

3.3 Design choices (why this shape and not another)

Hand-rolled parsing instead of clap. The core engine pins itself to 5 crates (rand, regex, half, serde/serde_json, minijinja — Cargo.toml). A parser crate would be the least painful dependency, but the flag set is small enough (~25 arms) that the loop is shorter than the clap derive boilerplate, and total control over the error text is worth real money in a teaching tool: every error above ends in a hint (minfer list, candidates, usage). The cost — hand-maintaining i += 2 bookkeeping and parse_err plumbing — is contained in one function.

Separate parsers for bench/specverify, shared parse for everything else. The alternative designs each fail: (1) teach the global parser the bench flags — then -p means two things depending on position, and the unknown-flag safety net needs exceptions; (2) make bench a separate binary — duplicating model loading and build wiring for ~500 lines of code. The chosen shape — bench::run(&prog, &raw_args[2..]) with its own loop and its own exit — gives each flag space exactly one meaning, keeps llama-bench-style flag letters (-p, -n, -r, -o) for people comparing tools, and shares all the loading code through the library modules.

Resolve cached names at all. The alternative is "always pass a full path", which pushes cache-layout knowledge onto every user and every shell alias. With name resolution, minfer list is self-documenting: what it prints, you can run. The prefix matching (typing qwen2.5-0.5b instead of the full 43-character filename) and the split-part collapse are the two places where the cache's shape (quant variants, multi-part files) would otherwise leak into the command line. The cost is one ambiguity class — prefix matches can hit several files — handled by refusing to guess and listing candidates instead.

curl subprocess instead of an HTTP crate. Resume (-C -), redirects (-L), progress, and TLS come free, at the price of requiring curl on PATH and giving up programmatic retry loops. For a download that happens once per model, the dependency saving wins — the same reasoning as the ollama pull delegation, which also inherits Ollama's own manifest/digest handling instead of reimplementing it.

exists() instead of is_file() in resolve. Strictness here would duplicate the "what is wrong" logic before the file is ever read; instead, resolution stays a one-line existence check and the richer diagnostics live where the file is opened, where the three-way directory/missing/corrupt branch (§3.2) has the path in hand. The invariant to preserve: one layer owns each diagnosis — resolve owns "not found / not downloaded / not unambiguous", the loader's error branch owns "found but unusable".

Defaults copied, deviations explicit. Every llama.cpp-matching default carries a comment saying so; the two deviations (n_predict 512, seed 42) are chosen for determinism and finite runs. This keeps the "compare against llama.cpp" discipline cheap forever: no benchmark needs to document a sampling delta, and the conversation/pipe tests can assert on output only because the seed is stable.

3.4 Pitfalls & invariants

  • n_ctx is consumed once, consistently. ctx is computed once (max(n_ctx, prompt_len)) and passed to both prefill and every decode forward; the model clamps it to max_seq_len internally. Letting any call site pass a different value would change KV geometry mid-run — the graph-reuse identity (doc 13) treats KV size as a rebuild trigger, so a mismatched n_ctx would silently reallocate and discard accumulated KV.
  • --cnv rejects --no-template up front (src/main.rs:530-536), before resolution: conversation history is rendered through the chat template, so a template-less conversation would corrupt the KV with un-formatted turns. Failing at parse time (not first-turn time) keeps the error message adjacent to the mistake.
  • Unknown flags die loudly; flag words never become the model path. The starts_with('-') rejection (src/main.rs:410-415) protects the positional contract. A flag missing its value is remembered in parse_err and reported after the full parse — one error, at the end, with usage.
  • Downloads must be size-checked before they are skipped. Existence alone would treat a Ctrl-C'd partial file as cached (the truncated file becomes doc 02's "corrupt" error, and worse, is not resumed because the code believed it was complete). The exists && len == remote rule (src/download/mod.rs:276-281) is what makes repeated hf: runs idempotent.
  • Ambiguity must refuse to guess. Both ambiguity sites — cached-name prefix matches and repo quant matches — enumerate candidates in the error rather than picking one. Guessing wrong here downloads or runs the wrong model silently, which is the worst failure class a resolver can have.
  • Split models enter through part 0. Both resolution paths (cached-name collapse, match_model group expansion) guarantee the returned path is …-00001-of-0000N.gguf, because gguf::resolve_splits (gguf.rs:1790) refuses to act as the entry when handed a later part. The resolver and the loader co-own that convention.
  • Server-like modes never consume stdin. serve/viz/--cnv read prompts from requests or the REPL; the single-shot stdin read (src/main.rs:572-574) would otherwise block startup on an empty pipe — the exact hang the code comment describes for minfer viz model.

4. Observe & verify

  • minfer --help — prints the usage block (print_usage, src/main.rs:80-143): the four MODEL forms and every option with its default, which is this stage's contract in readable form.
  • minfer list — runs list_local: the cache tree with sizes; every printed name is directly runnable as the model argument. With an empty cache it says so and points at minfer download.
  • minfer info <model> — resolves the model argument through the same download::resolve (including hf:/cached names, printing Model ready: …), then dumps metadata and key tensors without running inference — a dry run of stages 01+02.
  • --meta — with single-shot inference, replaces the one-line GGUF summary with the full metadata dump and still continues to generate (src/main.rs:630-652).
  • minfer download hf Qwen/Qwen2.5-0.5B-Instruct-GGUF q4_0 — shows the whole download path: API listing, quant match, curl progress bar, and on a second run the Already cached: … line proving the size check passed.
  • Unit tests — src/download/tests.rs:14-89 covers quant matching (single/split, case-insensitivity, ambiguity, exact-filename expansion) and src/gguf/tests.rs:6 covers split-filename parsing — the two pieces of logic the resolver leans on.
  • Integration tests — tests/conversation_cli.rs spawns the real binary with piped stdin: it asserts --cnv --no-template errors before any model load, that invalid --color values are rejected, that --help lists the conversation flags, and (against a cached model) that a piped prompt works — the determinism that makes this possible is seed: 42.

What you should see for the doc-02 handoff on a real run:

$ ./target/release/minfer qwen2.5-0.5b-instruct-q4_0 "Why is the sky blue?" -n 8
Loading model: /home/you/.cache/minfer/models/hf/Qwen/Qwen2.5-0.5B-Instruct-GGUF/qwen2.5-0.5b-instruct-q4_0.gguf ...
File: 407676704 bytes (388.8 MB) in 1 part(s)
GGUF: 31 KV, 168 tensors
Model loaded.
Vocabulary: 151936 tokens
Prompt: 23 tokens
...   ← doc 02 takes over at "Loading model"

5. Cross-references

  • docs/ARCHITECTURE.md §3 — the pipeline map this stage opens (CLI → resolve → load → mode branches); §9 — the download module's one-paragraph summary.
  • docs/USAGE.md — the full CLI reference this doc only samples: every flag of every subcommand.
  • docs/CLI-CONVERSATION-PLAN.md — the --cnv design: REPL, turn params, color/session flags parsed here and consumed by run_conversation.
  • docs/OPENAI-CHAT-API-PLAN.md — the serve design: how --n-ctx/--n-slots divide the context across concurrent requests. The division is an initial, elastic partition — a request that needs more than its share reclaims idle capacity from the other slots (ARCHITECTURE-EXECUTION-PLAN.md §5, C7/C7b).
  • docs/PERF-QWEN3-4B-VS-LLAMACPP.md §2 — why --n-ctx sizes KV instead of the model maximum (the 12 GB arithmetic and the Metal first-submit tax).
  • 02 — GGUF load — the next stage: what happens when the path this doc produced is finally opened (header, metadata KV, tensor table, mmap).
  • 04 — Tokenizer + template — where the prompt string collected here is rendered and turned into token ids.
  • 07 — Allocator + KV regions — where n_ctx becomes bytes and two persistent regions per layer.
  • 12 — Sampler — the math behind the GenParams sampling defaults this stage collects.

← Index · 02 — GGUF load →

02 · GGUF load — from file bytes to a model handle

Stage: model path resolved (01) → GGUF load: parse header + metadata, mmap the data blob → model dispatch + weight registration (03). Code: src/gguf.rs — load_gguf_model (line 1998), GgufContext::init_from_reader (line 1021), MmapFile (line 1866), ggml_pad (line 328); call site src/main.rs line 579. Block layouts: src/block.rs.

1. Background — where this stage sits

Doc 01 ended with a plain string: a filesystem path to a .gguf file (resolved from a local path, a hf:/ollama: URI, or a cached model name). This doc is where that path stops being a name and starts being data. When main.rs runs gguf::load_gguf_model(...) (line 579), the engine has exactly one asset: a file on disk, typically hundreds of megabytes to several gigabytes. When the call returns, the engine holds a GgufModel handle: the model's hyperparameters and tokenizer sit in a parsed key-value table, every weight tensor is catalogued by name/type/shape, and the raw weight bytes are reachable in memory without having copied a single one of them.

First, the vocabulary, because everything else builds on these five words:

  • GGUF (GPT-Generated Unified Format) is llama.cpp's single-file model container. One file holds everything needed to run the model: the weight tensors, the hyperparameters (layer count, head count, …), the tokenizer's vocabulary, and the chat template. Version 3 is the current revision; minfer reads exactly that (GGUF_VERSION = 3, gguf.rs line 10).
  • A tensor is an n-dimensional array of numbers. A transformer's weights are a few hundred of them: embeddings, per-layer projection matrices, norm vectors. In this engine a tensor is just named bytes — the GGUF file stores each one under a name like blk.0.attn_q.weight with a shape and a quantization type.
  • Quantization is compressing those numbers: instead of storing each weight as a 4-byte f32, store it in fewer bits (4.5 bits per weight for Q4_0) by sharing one scale factor across a small block of values. The details are §2.4 and doc 10; for this doc you only need the consequence: a tensor's byte size is not "element count × 4" and the parser must know each type's block arithmetic to slice tensors correctly.
  • Metadata KV (key-value) pairs are the file's self-description: strings, numbers, and arrays such as general.architecture = "qwen2" or tokenizer.ggml.tokens = [151,936 strings]. Think of them as a JSON document embedded at the front of a binary file.
  • mmap (memory mapping) asks the operating system to make a file's bytes appear at an address in the process's address space. Nothing is read up front: the first touch of any page triggers an OS fault that pulls that 4 KiB page from disk. The process reads file bytes through an ordinary slice, and the OS page cache (shared memory of disk pages, reused by all processes) does the caching once, for everyone.

Why does an engine written from scratch need its own GGUF parser? Because there are no ML-framework dependencies to lean on — minfer's whole gguf.rs is a faithful Rust port of llama.cpp's gguf.cpp, down to the gguf.cpp line numbers in the comments. The parser is deliberately dumb but strict: it reads the container faithfully, checks every length and offset, and hands clean data to the next stage. It does not interpret the model — deciding that general.architecture = "qwen2" means "build a Qwen2 transformer" is doc 03's job.

What would break without this stage? Everything downstream. Doc 03 needs the metadata to pick an architecture and the tensor table to find blk.{i}.attn_q.weight by name. Doc 04 needs the tokenizer arrays. Docs 07–08 need tensor strides (byte distances between rows) to be ggml-compatible so kernels index memory correctly. And doc 14's zero-copy Metal path needs the weight bytes to live in an mmap region — which is why this doc, not doc 03, is where the mmap is created. The load stage is also the last line of defense against a corrupt file: a truncated download or a misaligned write is caught here, with a message, instead of producing garbage logits at step 200.

2. Principle — how it works and why

2.1 The file is four regions, read in three

A GGUF v3 file is a strictly ordered sequence of regions:

byte 0 ┌──────────────────────────────┐
       │ magic "GGUF" (4 B)           │  ┐
       │ version  u32 (4 B)           │  │
       │ n_tensors i64 (8 B)          │  ├─ header: 24 bytes, always
       │ n_kv      i64 (8 B)          │  ┘
       ├──────────────────────────────┤
       │ metadata KV pairs (n_kv of)  │   key string, type tag, value
       │  · scalars: u32/i32/f32/…    │   arrays: elem type + count + data
       │  · strings: u64 len + bytes  │   (the tokenizer vocab lives here:
       │  · arrays of the above       │    151,936 strings ≈ megabytes)
       ├──────────────────────────────┤
       │ tensor info table (n_tensors)│   name, n_dims, shape[nd], ggml
       │                              │   type, offset (from data start)
       ├──────────────────────────────┤
       │ (padding to `general.alignment`, default 32 B)
       ├──────────────────────────────┤
       │ tensor data blob             │   quantized blocks, back to back,
       │  tensor 0 | tensor 1 | …     │   each padded to the alignment
       └──────────────────────────────┘   (offsets are absolute within it)

The parser consumes the first three regions with a cursor (GgufReader) and stops. The fourth region — 99% of the file — is never read by the parser at all. We measured this on a real file (Qwen2.5-0.5B-Instruct Q4_0, 428,730,208 bytes): the header is 24 bytes, the 26 KV pairs end at byte 5,931,189 (the tokenizer vocabulary and BPE merges are the bulk of it), the 291-entry tensor table ends at byte 5,947,741, three alignment pad bytes follow, and the data section starts at byte 5,947,744. That is 1.39% of the file parsed and 98.61% left untouched — untouched not skipped-and-copied, but never read, because the data region becomes an mmap that the OS pages in lazily, on demand, forever after.

This is the first design decision worth internalizing: loading a model must not mean reading it. Loading model: … → Model loaded. prints in tens of milliseconds precisely because the only O(file-size) work is the kernel setting up the mapping (a few syscall-rounds), not transferring gigabytes.

2.2 Alignment: why every region is padded, and to what

The header, KV block, and tensor table are all variable-length — strings and arrays make their ends land on arbitrary byte offsets. But the data section must start at an address that is friendly to consumers: SIMD loads want 16–64 byte alignment, GPU buffer offsets want more. GGUF solves this with one rule: the data section starts at the next multiple of general.alignment (default 32), and every tensor's data is padded up to a multiple of that alignment too, so tensor N starts right where tensor N−1 ended.

The padding operation is ggml's GGML_PAD macro, ported verbatim (gguf.rs lines 327–332):

#![allow(unused)]
fn main() {
#[inline]
pub fn ggml_pad(x: usize, n: usize) -> usize {
    // ((x) + (n) - 1) & ~((n) - 1)
    // Assumes n is power of 2
    (x + n - 1) & !(n - 1)
}
}

One line of bit tricks: add n−1 to force a carry past the next boundary, then mask off the low bits so only multiples of n survive. For n = 32 the mask is !31 = …11100000. Worked example: the tensor table of the 0.5B file ends at byte 5,947,741; (5,947,741 + 31) & !31 = 5,947,744, so 3 pad bytes are inserted. Because the trick requires a power of two, the parser rejects any general.alignment that isn't one (line 1488) rather than silently rounding wrong.

Why does the parser care at all — couldn't each consumer align itself? Because the offsets are stored in the file: each tensor's offset field is relative to the (aligned) start of the data section, and tensor bytes must be found at exactly that offset. Three things go wrong if padding is ignored or miscomputed:

  1. Wrong slices. Offsets would drift by the pad bytes; every tensor after the first misalignment would be cut from shifted bytes — a garbled model that produces fluent nonsense or crashes in a kernel far from the cause. The parser defends itself: it re-derives each tensor's expected offset as a running sum of padded sizes and rejects the file if the stored offset disagrees (lines 1692–1713).
  2. Slow or faulting vector loads. The CPU kernels read quantized blocks as 16/32-byte SIMD chunks; a base pointer off by 3 bytes splits every chunk across cache lines (x86 tolerates this slowly; some aarch64 NEON loads fault on misalignment). Doc 10's kernels assume ggml's alignment exactly.
  3. GPU mapping breaks. Metal's zero-copy path wraps the mmap in a GPU buffer, and newBufferWithBytesNoCopy requires a page-aligned base address (16,384 bytes on Apple Silicon). minfer satisfies this by mapping the whole file from offset 0: the mapping base — hence part.data.as_ptr() — is page-aligned by construction, and every tensor is then addressed as an offset into that one buffer (§3.2 l). Per-tensor 32-byte alignment keeps those offsets sane for SIMD loads; the page alignment comes from the mapping itself.

The alignment is "per-region" in one more sense: it is applied twice, once before the data section (after the table) and once after each tensor inside it. Both applications use the same ggml_pad, and the running-sum check verifies both. A writer that forgot interior padding would fail the parse rather than poison memory.

2.3 Quantized sizing: 32 values in 18 bytes

Quantization is the reason this 630M-parameter model (its own general.size_label metadata) fits in 409 MiB of Q4_0/Q8_0 bytes instead of the ~2.3 GiB its f32 form would take. The scheme (fully dissected in doc 10) is block quantization: group values into blocks, store one shared f16 scale per block, and store each value as a small integer relative to that scale. Two block shapes exist:

  • 32-value blocks ("Q4_0 family"): Q4_0, Q4_1, Q5_0, Q5_1, Q8_0.
  • 256-value super-blocks ("K-quants"): Q4_K, Q5_K, Q6_K, … — 256 values per super-block, subdivided into 8–16 sub-blocks that each get their own quantized scale.

The canonical example is Q4_0. Its 32 values are stored in 18 bytes — an f16 (2-byte) scale plus 16 bytes holding 32 4-bit nibbles:

BlockQ4_0 (18 bytes)                      value ≈ d × (q − 8)
┌─────────────┬───────────────────────────────┐
│ d : f16 (2B)│ qs : 16 bytes = 32 × 4 bits   │
└─────────────┴───────────────────────────────┘
  18 × 8 / 32 = 4.5 bits per weight

The layouts are repr(C) structs in block.rs mirroring llama.cpp's ggml-common.h — BlockQ4_0 { d: Fp16, qs: [u8; 16] } (lines 51–56) — with compile-time asserts that the Rust sizes equal the C sizes (lines 190–204). The type table in gguf.rs is what the parser actually consults (type_size(), lines 200–246). Every supported type, with its block arithmetic:

Typevalues/block (blck_size)bytes/block (type_size)bits/weightanatomy
F321432.0raw
F161216.0raw
Q4_032184.5f16 d + 16 B nibbles
Q4_132205.0f16 d + f16 m + 16 B nibbles
Q5_032225.5f16 d + 4 B high bits + 16 B nibbles
Q5_132246.0f16 d + f16 m + 4 B high + 16 B nibbles
Q8_032348.5f16 d + 32 i8
Q4_K2561444.5f16 d + f16 dmin + 12 B scales + 128 B nibbles
Q5_K2561765.5f16 d + f16 dmin + 12 B scales + 32 B high + 128 B nibbles
Q6_K2562106.5625128 B low + 64 B high + 16 i8 scales + f16 d

(source: gguf.rs lines 200–283; struct comments in block.rs lines 16–27, 135–167. Q8_K at 290 bytes is the activation format — it appears in files rarely and is used at runtime on CPU; doc 10 covers it.)

Note the two f16 fields in a Q4_K block are the super-block scale and min; the 12-byte scales field packs eight 6-bit sub-block scales and eight 6-bit mins (the unpack_q4k_scales bit-shuffling in block.rs lines 32–44 decodes it). That two-level structure is the whole trick of K-quants: fine scales capture local variance, the super-scale keeps them honest.

The byte-size formula follows directly from the table: blocks per tensor × bytes per block. ggml_nbytes (gguf.rs lines 828–849) implements it via the strides (next section); in its simplest one-dimensional reading it is n_elements / blck_size × type_size. Real numbers for the 0.5B model's largest tensor, token_embd.weight (shape [896, 151936] = 136,134,656 elements, on disk as Q4_0):

LayoutbytesMiB
F32544,538,624519.3
Q8_0144,643,072137.9
Q4_0 (as on disk)76,575,74473.0
Q4_K76,575,74473.0
Q6_K111,672,960106.5

Q4_0 and Q4_K land on the same 4.5 bits/weight — the K-quant's extra scale structure buys quality at equal size, which is why Q4_K_M variants are the popular download. And note the file itself uses different types for different tensors: in this very file output.weight is Q8_0 (137.9 MiB — the output projection is quality-sensitive) while token_embd.weight is Q4_0 (73.0 MiB), and the norm vectors are F32 (896 × 4 = 3,584 bytes each). The parser never assumes a uniform type; it reads each tensor's type tag and computes sizes per tensor.

2.4 Strides: the shape's byte-geometry, computed once in the parser

A stride is the number of bytes you skip to move one step along a dimension. GGML (and therefore GGUF) stores shapes as ne[4] (elements per dimension, ne[0] fastest-varying — for weight matrices ne[0] is the input dim, ne[1] the output dim) and strides as nb[4] (bytes per step). Strides are what make a kernel able to walk a tensor without knowing what a "Q4_K super-block" is: a row is nb[1] bytes away from the previous row, period.

The parser computes strides the instant it learns shape + type (gguf.rs lines 1650–1655):

#![allow(unused)]
fn main() {
// calculate byte offsets (gguf.cpp lines 728-732)
info.nb[0] = type_size;
info.nb[1] = info.nb[0] * (info.ne[0] / blck_size) as usize;
for j in 2..GGML_MAX_DIMS {
    info.nb[j] = info.nb[j - 1] * info.ne[j - 1] as usize;
}
}

Read it as three claims:

  • nb[0] = type_size — the atom of dimension 0 is one block, not one element. For F32 that's 4 bytes per step; for Q4_0 that's 18 bytes per step (one block of 32 values).
  • nb[1] = nb[0] × ne[0]/blck_size — one row contains ne[0] values, which is ne[0]/blck_size blocks, so a row is that many block-sizes long. For a Q4_0 row of ne[0] = 896: 896/32 = 28 blocks × 18 B = 504 bytes per row.
  • Every higher stride multiplies: nb[2] = nb[1] × ne[1] (one full plane), nb[3] = nb[2] × ne[2].

ggml_nbytes then totals it: dimension 0 contributes ne[0]/blck_size blocks' worth (ne[0] × nb[0] / blck_size), each higher dimension contributes (ne[i] − 1) × nb[i] (the last step doesn't add bytes). The division-aware form is exactly what keeps the 4.5-bit types honest — you cannot compute this size without the block size, which is why the parser demands ne[0] % blck_size == 0 (line 1627) and fails otherwise (a Q4_0 tensor with 895 elements per row is unrepresentable, not merely odd).

Two facts make this stage's stride work load-bearing. First, the same formula is reimplemented at tensor-creation time (tensor.rs lines 142–149, loader.rs lines 192–197), so the GGUF-parsed strides and the in-memory Tensor.strides agree by construction — kernels can be written against one geometry. Second, the geometry is ggml's, byte for byte, which is what allows minfer to consume llama.cpp-quantized files and to be compared against llama.cpp numerically (docs record the greedy-output matches).

2.5 mmap: borrowing the file instead of owning a copy

The obvious loader reads the file into a Vec<u8> (read()-into-heap) and hands out slices of that vector. minfer instead calls mmap and hands out slices of the mapping. The difference is not academic:

  1. Load latency. read() copies every byte before "Model loaded." can print; mmap copies no bytes at all. The kernel just installs a page table entry: microseconds, independent of file size.
  2. RAM footprint vs file size. With read(), RSS (resident memory, the pages actually in RAM) immediately includes the whole 4 GB. With mmap, RSS grows only as pages are touched — the embedding table and the layers you actually use — and the OS can evict cold pages under pressure because it knows they're backed by a file it can re-read.
  3. Page cache sharing. The OS keeps one copy of the file's pages in its page cache. mmap'd readers attach to it: a second minfer process on the same model adds no new copies, and download's freshly written file is already warm.
  4. Zero-copy GPU mapping (the payoff that dominates later docs). Because the weight bytes live in one stable, page-aligned mapping for the life of the process, Metal can wrap that same memory in a MTLBuffer with newBufferWithBytesNoCopy — the GPU reads the model file through the page cache with no upload step at all (doc 14). That is only possible because the data was never copied into a heap allocation; the mapping is the storage.

The cost of mmap is honesty about lifetime: the mapping must outlive every slice borrowed from it. minfer's answer is the simplest correct one — the MmapFile is deliberately leaked (Box::leak, line 1999) so its backing memory lives until process exit, and every tensor slice is &'static [u8]. For a process whose whole purpose is to run one model, that is not a leak in any meaningful sense; it is a pool that is freed by exit. (The alternatives — Arc<MmapFile> reference counting, or self-referential structs — buy nothing here and cost real complexity; §3.3.)

2.6 Multi-part files: many GGUFs, one model

Models above a few GB are published as splits — name-00001-of-00002.gguf, …-00002-of-00002.gguf — because of file-size limits on hosting platforms. Each part is a complete, valid GGUF file: own header, own (small) metadata, own tensor table, own data blob. The metadata in part 0 carries split.no = 0, split.count = 2, split.tensors.count = 339 (verified in a real 7B Q4_K_M part 0). Loading = mmap and parse each part in order, then concatenate the tensor catalogs: the merged index maps every tensor name to "the part that lists it", and a lookup slices from that part's own mapping. Doc 03's loader builds exactly that map (loader.rs lines 329–343); this doc's load_gguf_model produces the parts: Vec<GgufPart> it iterates.

The merge rule mirrors llama.cpp: entry = part 0. Part 0's metadata is the model metadata (architecture, tokenizer, template — the 7B part 0 has 29 KV pairs where a non-split file of the same family has 26, the extras being the split.* trio), and part 0 must be the file you point minfer at (split.no != 0 is rejected, line 2020). Tensors are distributed round-robin-ish by the quantizer; minfer never assumes where, it just looks the name up in whichever part claims it.

2.7 What the parse hands to doc 03

The product of this stage is one struct, and its shape is the contract with everything after:

GgufModel
└── parts: Vec<GgufPart>            one per file; [0] is the entry
    ├── ctx: GgufContext            parsed regions (header + KV + tensor table)
    │   ├── version, alignment      format facts
    │   ├── kv: Vec<GgufKv>         metadata: keys → typed values/arrays
    │   └── info: Vec<GgufTensorInfo>  name, ne[4], nb[4], type, offset
    └── data: &'static [u8]         the mmap'd file bytes (whole file)

Who reads what next:

Metadata key(s)ConsumerDoc
general.architecturemodels/mod.rs::load_model (line 97) → "qwen2" / "qwen3" dispatch03
qwen2.block_count, qwen2.embedding_length, qwen2.attention.head_count(_kv), qwen2.feed_forward_length, qwen2.context_length, qwen2.attention.layer_norm_rms_epsilon, qwen2.rope.freq_base, qwen2.rope.frequency_scale (+ llama.* aliases)models/qwen2/loader.rs::hparams_from_gguf (lines 118–150) → HParams03
tensor table (info)loader's merged tensor map → Tensors → GPU registration03
tokenizer.ggml.tokens/scores/token_type/merges/bos_token_id/eos_token_idtokenizer.rs::Tokenizer::load (lines 82–149)04
tokenizer.chat_templatemain.rs::get_chat_template (lines 1495–1504)04
split.no, split.countload_gguf_model itself (lines 2002–2025)02
general.alignmentinit_from_reader (lines 1481–1486)02

(general.quantization_version is present in files — value 2 on disk — but minfer doesn't consult it; the ggml type tag per tensor is the operative fact.)

3. Implementation

3.1 Data in / data out

In: a &Path — one .gguf file, which is either a whole model or part 00001 of a split. Nothing else: no config, no sidecar files.

Out: Option<GgufModel> (None ⇒ error already printed, main.rs exits). The two structs (gguf.rs lines 602–623, 1945–1954):

#![allow(unused)]
fn main() {
pub struct GgufTensorInfo {
    pub name: String,
    pub ne: [i64; GGML_MAX_DIMS],   // number of elements per dimension
    pub nb: [usize; GGML_MAX_DIMS], // stride in bytes per dimension
    pub type_: GgmlType,
    pub offset: u64, // offset from start of data section
}

pub struct GgufContext {
    pub version: u32,
    pub kv: Vec<GgufKv>,
    pub info: Vec<GgufTensorInfo>,
    pub alignment: usize,
    pub offset: usize, // offset of data section from beginning of file
    pub size: usize,   // size of data section in bytes
}

pub struct GgufPart {
    pub ctx: GgufContext,
    /// `'static` slice of the (leaked, process-lifetime) mmap of the part file.
    pub data: &'static [u8],
}

pub struct GgufModel {
    pub parts: Vec<GgufPart>,
}
}

A GgufKv (line 337) is { key, is_array, type_: GgufType, data: Vec<u8>, data_string: Vec<String> } — scalars live in data as little-endian raw bytes, strings in data_string, so typed getters decode on demand.

The tensor bytes themselves are not yet wrapped as tensors — that is doc 03. What this stage guarantees: for any ti in ctx.info, the bytes data[ctx.offset + ti.offset .. + ggml_nbytes(ti)] are the tensor's full contents, contiguous, and the address data.as_ptr() is page-aligned (it is mmap's return value).

3.2 Key code

(a) The format constants (gguf.rs lines 9–19). Everything downstream of this point is these numbers plus the type table:

#![allow(unused)]
fn main() {
const GGUF_MAGIC: [u8; 4] = [b'G', b'G', b'U', b'F'];
const GGUF_VERSION: u32 = 3;
const GGUF_DEFAULT_ALIGNMENT: usize = 32;
const GGUF_KEY_GENERAL_ALIGNMENT: &str = "general.alignment";

const GGUF_MAX_STRING_LENGTH: u64 = 1024 * 1024 * 1024;
const GGUF_MAX_ARRAY_ELEMENTS: u64 = 1024 * 1024 * 1024;

// Note: GGML_MAX_DIMS and GGML_MAX_NAME from ggml.h
const GGML_MAX_DIMS: usize = 4;
const GGML_MAX_NAME: usize = 64;
}

The two MAX_ caps are not decoration: every length read from the file is checked against them before allocating (read_string, read_vec), so a corrupt header cannot make the parser malloc 2⁶⁴ bytes and die. The parser is, among other things, a hostile-input parser — model files come from the internet.

(b) The type table's arithmetic half (gguf.rs lines 209–229). The comments are the spec — each entry is sizeof(struct) spelled out:

#![allow(unused)]
fn main() {
// sizeof(block_q4_0) = sizeof(ggml_half) + QK4_0/2 = 2 + 16 = 18
GgmlType::Q4_0 => 18,
// sizeof(block_q4_1) = sizeof(ggml_half)*2 + QK4_1/2 = 2 + 2 + 16 = 20
GgmlType::Q4_1 => 20,
// sizeof(block_q5_0) = sizeof(ggml_half) + QK5_0/2 + QK5_0/8 = 2 + 16 + 4 = 22
GgmlType::Q5_0 => 22,
// sizeof(block_q8_0) = sizeof(ggml_half) + QK8_0 = 2 + 32 = 34
GgmlType::Q8_0 => 34,
// QK_K=256, super-block types — sizes from type_traits
GgmlType::Q4_K => 144,
GgmlType::Q5_K => 176,
GgmlType::Q6_K => 210,
GgmlType::Q8_K => 290,
}

The sibling method blck_size() (lines 250–283) returns values-per-block (32 for the first family, 256 for the K family, 1 for F32/F16). type_size and blck_size together are the complete sizing algebra — everything in §2.3/§2.4 derives from them. (The full enum spans all 42 ggml type discriminants, lines 109–152, including the unsupported IQ/quaternion families — their entries exist so the parser can recognize and reject files that need them.)

(c) The reader's primitives. All typed reads go through one function, read_val::<T> (lines 905–918): bounds-check against nbytes_remain, copy size_of::<T>() bytes, then ptr::read_unaligned them into a T — unaligned because the cursor has no alignment guarantee (strings of odd length precede most scalars). GGUF is little-endian; x86-64 and aarch64 are little-endian hosts, so the byte copy is the decode, and the version field's own check (below) is the early tripwire for byte-swapped files. Strings are u64 length + bytes (read_string, lines 938–962): the length is checked against the 1 GiB cap and the remaining file size before any allocation, so a corrupt header cannot make the parser malloc itself to death.

(d) Header + version guards (gguf.rs lines 1071–1096). After the 4-byte magic comparison (lines 1033–1065, which prints the four offending characters it found), the version is screened:

#![allow(unused)]
fn main() {
if let Some(version) = gr.read_val::<u32>() {
    ctx.version = version;
    if ctx.version == 0 {
        eprintln!("GGUF: bad GGUF version: {}", ctx.version);
        ok = false;
    }
    // endianness check (gguf.cpp lines 490-500)
    if ok && (ctx.version & 0x0000FFFF) == 0x00000000 {
        eprintln!("GGUF: failed to load model: this GGUF file version {} is extremely large, is there a mismatch between the host and model endianness?", ctx.version);
        ok = false;
    }
    if ok && ctx.version == 1 {
        eprintln!(
            "GGUF: GGUFv1 is no longer supported, please use a more up-to-date version"
        );
        ok = false;
    }
    if ok && ctx.version > GGUF_VERSION {
        eprintln!("GGUF: this GGUF file is version {} but this software only supports up to version {}", ctx.version, GGUF_VERSION);
        ok = false;
    }
}
}

The endianness heuristic is clever: if the file were written big-endian, the u32 version 3 (0x00000003 little-endian) would read back as 0x03000000, whose low 16 bits are zero — an "impossible" version number, reported with the endianness hint instead of a mystery failure downstream.

(e) The KV loop's type dispatch (gguf.rs lines 1162–1190). Each pair is key-string → type tag → (if array) element type + count → value:

#![allow(unused)]
fn main() {
let mut type_: GgufType;
let mut is_array: bool = false;
let mut n: u64 = 1;

match gr.read_gguf_type() {
    Some(t) => type_ = t,
    None => { ok = false; break; }
}

if type_ == GgufType::Array {
    is_array = true;
    match gr.read_gguf_type() {
        Some(t) => type_ = t,   // element type of the array
        None => { ok = false; break; }
    }
    match gr.read_val::<u64>() {
        Some(v) => n = v,       // element count
        None => { ok = false; break; }
    }
}
}

The 13 type tags (GgufType, lines 25–39) are u8…f64 plus string and array; the match type_ block that follows (lines 1197–1471) reads each, arrays element-wise. Duplicate keys are rejected (lines 1151–1157) — llama.cpp relies on unique keys and so does every consumer.

(f) Per-tensor type checks and stride computation (gguf.rs lines 1623–1655). This is where §2.3 and §2.4 become code, per tensor:

#![allow(unused)]
fn main() {
let type_size = info.type_.type_size();
let blck_size = info.type_.blck_size();

// check that row size is divisible by block size
if blck_size == 0 || info.ne[0] % blck_size != 0 {
    eprintln!("GGUF: tensor '{}' of type {} ({}) has {} elements per row, not a multiple of block size ({})",
        info.name, type_val, info.type_.type_name(), info.ne[0], blck_size);
    ok = false;
    break;
}

// check that size in bytes is representable
let nelements = ggml_nelements(&info.ne);
if ok && (nelements / blck_size) as u64 > (usize::MAX / type_size) as u64 {
    eprintln!(
        "GGUF: tensor '{}' with shape ({}, {}, {}, {}) has a size in bytes > {}",
        info.name, info.ne[0], info.ne[1], info.ne[2], info.ne[3], usize::MAX
    );
    ok = false;
    break;
}

// calculate byte offsets (gguf.cpp lines 728-732)
info.nb[0] = type_size;
info.nb[1] = info.nb[0] * (info.ne[0] / blck_size) as usize;
for j in 2..GGML_MAX_DIMS {
    info.nb[j] = info.nb[j - 1] * info.ne[j - 1] as usize;
}
}

Before this: the shape itself is validated (n_dims ≤ 4, lines 1548–1562; negative dims rejected; a product-of-dims overflow check, lines 1585–1598). After it: the tensor's data-section-relative offset is read as u64 (lines 1661–1668) — not computed, read, because the writer chose the layout.

(g) Alignment, contiguity, total size (gguf.rs lines 1679–1714). The parse's final act is to pin down the data section and prove the tensor table consistent with it:

#![allow(unused)]
fn main() {
// align to data section (gguf.cpp lines 751-756)
if n_tensors > 0 {
    let aligned_offset = ggml_pad(gr.tell() as usize, ctx.alignment);
    if !gr.seek(aligned_offset as u64) {
        eprintln!("GGUF: failed to seek to beginning of data section");
        return None;
    }
}

// store data section offset (gguf.cpp line 759)
ctx.offset = gr.tell() as usize;

// compute total data section size (gguf.cpp lines 762-782)
{
    ctx.size = 0;
    for i in 0..ctx.info.len() {
        let ti = &ctx.info[i];
        if ti.offset != ctx.size as u64 {
            eprintln!(
                "GGUF: tensor '{}' has offset {}, expected {}",
                ti.name, ti.offset, ctx.size
            );
            eprintln!("GGUF: failed to read tensor data");
            return None;
        }
        let padded_size = ggml_pad(ggml_nbytes(ti), ctx.alignment);
        if usize::MAX - ctx.size < padded_size {
            eprintln!(
                "GGUF: tensor '{}' size overflow, cannot accumulate size {} + {}",
                ti.name, ctx.size, padded_size
            );
            return None;
        }
        ctx.size += padded_size;
    }
}
}

The running-sum check is the quiet hero: it recomputes where each tensor must start (previous end, padded) and compares with the file's claim. Our real-file walk shows it passing exactly: output.weight (Q8_0) is 151,936 × 896 / 32 blocks × 34 bytes = 144,643,072 bytes, sits at offset 0, and indeed token_embd.weight (Q4_0) follows at offset 144,643,072; blk.0.attn_norm.weight follows the 73 MiB embedding at 221,218,816 = 144,643,072 + 76,575,744. No drift, no gaps: quantized arithmetic and padding compose perfectly.

(h) The mmap (gguf.rs lines 1861–1890 and 1892–1927). No memmap2 crate — raw libc, declared once:

#![allow(unused)]
fn main() {
/// Read-only mmap of a GGUF file part (zero-dependency: raw mmap/munmap via
/// the system libc, which Rust links by default). The file pages are shared
/// with the CPU and (via newBufferWithBytesNoCopy) the GPU instead of being
/// copied — llama's `llama_mmap` / `newBufferWithBytesNoCopy` equivalent
/// (ggml-metal-device.m:1668). MAP_PRIVATE (no writes happen), PROT_READ.
pub struct MmapFile {
    ptr: *mut u8,
    len: usize,
    #[allow(dead_code)]
    _file: std::fs::File, // keeps the fd alive for the mapping's lifetime
}

// Generic POSIX mmap (Linux + macOS; the syscall ABI is identical on both).
#[cfg(unix)]
const PROT_READ: i32 = 0x1;
#[cfg(unix)]
const MAP_PRIVATE: i32 = 0x0002;

#[cfg(unix)]
extern "C" {
    fn mmap(addr: *mut std::ffi::c_void, len: usize, prot: i32,
            flags: i32, fd: i32, offset: i64) -> *mut std::ffi::c_void;
    fn munmap(addr: *mut std::ffi::c_void, len: usize) -> i32;
}
}

and the map call:

#![allow(unused)]
fn main() {
pub fn map(path: &std::path::Path) -> Option<Self> {
    #[cfg(unix)]
    {
        use std::os::unix::io::AsRawFd;
        let file = std::fs::File::open(path).ok()?;
        let len = file.metadata().ok()?.len() as usize;
        let ptr = unsafe { mmap(std::ptr::null_mut(), len, PROT_READ,
                                MAP_PRIVATE, file.as_raw_fd(), 0) };
        // MAP_FAILED = (void*)-1
        if ptr as isize == -1 {
            return None;
        }
        // (The GPU-side warm-up read happens in src/metal/ register_part —
        // the first GPU access to file-backed pages costs ~44 ms of one-time
        // page/TLB setup per process, METAL_OPTIMIZATIONS #39. A CPU-side
        // madvise/touch does NOT fix it — the cost is the GPU's own access.)
        Some(MmapFile { ptr: ptr as *mut u8, len, _file: file })
    }
}
}

Three details deserve their sentence each. _file is kept (though unused) so the file descriptor cannot close — on some systems closing the fd is allowed while mapped, but holding it makes the lifetime story trivially airtight. MAP_FAILED is (void*)-1, not NULL, hence the as isize == -1 check rather than a null check. And the comment records a measured fact from docs/METAL_OPTIMIZATIONS.md (#39): the first GPU touch of file-backed pages costs ~44 ms of page/TLB setup, which Metal's register_part pays once at load, on purpose, outside the timed region. as_slice() (lines 1930–1932) turns the pointer into &[u8]; Drop (lines 1935–1942) calls munmap — though with the leak below, it effectively never runs.

(i) The entry point: parts, leak, merge (gguf.rs lines 1998–2014):

#![allow(unused)]
fn main() {
pub fn load_gguf_model(path: &std::path::Path) -> Option<GgufModel> {
    let mmap0 = Box::leak(Box::new(MmapFile::map(path)?));
    let data0: &'static [u8] = mmap0.as_slice();
    let ctx0 = GgufContext::init_from_data(data0)?;
    let split_count = ctx0
        .get_key_val_i64("split.count")
        .map(|v| v as usize)
        .unwrap_or(1);

    if split_count <= 1 {
        return Some(GgufModel {
            parts: vec![GgufPart {
                ctx: ctx0,
                data: data0,
            }],
        });
    }
}

Box::leak is the 'static trick from §2.5: MmapFile::map returns an owned value; leaking it converts ownership into a process-lifetime guarantee, which is exactly the borrow lifetime every downstream tensor slice needs. Then either the fast path (no split: parse and return) or the split path — checks first (split.no must be 0, lines 2016–2025; filename-derived part list must match split.count, lines 2027–2034), then parse each part in order:

#![allow(unused)]
fn main() {
let mut parts = Vec::with_capacity(split_count);
for (i, p) in part_paths.iter().enumerate() {
    let mmap = Box::leak(Box::new(MmapFile::map(p)?));
    let data: &'static [u8] = mmap.as_slice();
    let ctx = GgufContext::init_from_data(data)?;
    let no = ctx
        .get_key_val_i64("split.no")
        .map(|v| v as usize)
        .unwrap_or(0);
    if no != i {
        eprintln!("GGUF: split {p:?} has split.no={no}, expected {i}");
        return None;
    }
    parts.push(GgufPart { ctx, data });
}
Some(GgufModel { parts })
}

Note what is not here: no merging of KV maps, no re-allocation of tensor data, no concatenation of bytes. Each part keeps its own ctx and its own mmap slice; the "one tensor index" is built lazily by doc 03's loader from the parts (tensor_map over all part.ctx.info, loader.rs lines 331–337). The parse's job is to make that trivial, not to do it.

(j) Split filename parsing (gguf.rs lines 1958–1974) — split_file_info recognizes the name-0000X-of-0000Y.gguf pattern (exactly 5 digits, 1 ≤ X ≤ Y) and resolve_splits (lines 1979–1992) rebuilds the full ordered path list from part 00001, rejecting non-first parts:

#![allow(unused)]
fn main() {
pub fn split_file_info(name: &str) -> Option<(String, usize, usize)> {
    let stem = name.strip_suffix(".gguf")?;
    let dash = stem.rfind("-of-")?;
    let idx_part = &stem[..dash];
    let count_str = &stem[dash + 4..];
    let idx_dash = idx_part.rfind('-')?;
    let prefix = &idx_part[..idx_dash];
    let idx_num = &idx_part[idx_dash + 1..];
    if idx_num.len() != 5 || count_str.len() != 5 {
        return None;
    }
    let idx: usize = idx_num.parse().ok()?;
    let count: usize = count_str.parse().ok()?;
    if idx == 0 || count == 0 || idx > count {
        return None;
    }
    Some((prefix.to_string(), idx - 1, count))
}
}

Both functions have unit tests right below (lines 2058–2096) covering the 2-part, 3-part, non-split, and invalid-index cases.

(k) The call site's accounting (main.rs lines 622–635). This is the "File: … bytes … in N part(s)" line you see at startup, and the part-0 convention made explicit:

#![allow(unused)]
fn main() {
let n_parts = gguf_model.parts.len();
let total_bytes: usize = gguf_model.parts.iter().map(|p| p.data.len()).sum();
println!(
    "File: {} bytes ({:.1} MB) in {n_parts} part(s)",
    total_bytes,
    total_bytes as f64 / 1_048_576.0
);

let ctx = &gguf_model.parts[0].ctx;
if meta_flag {
    dump_gguf_metadata(ctx);
} else {
    println!("GGUF: {} KV, {} tensors", ctx.kv.len(), ctx.info.len());
}
}

For the 7B split on disk here: 3,993,201,344 + 689,872,288 = 4,683,073,632 bytes → File: 4683073632 bytes (4466.1 MB) in 2 part(s). Note the sizes summed are the mmap lengths — the whole files, data blob included — not the data sections; this line is honest about disk footprint, while ctx.size (from §3.2 g) is the tensor-data footprint.

(l) The handoff, one step further (models/qwen2/loader.rs lines 182–199, excerpted; doc 03 owns the full story). When the next stage wants a weight, it computes the global offset and slices the mmap:

#![allow(unused)]
fn main() {
let off = ctx.offset + ti.offset as usize;
// Use GGML type for byte-size calculation — always correct regardless of TensorType mapping
let ts = ti.type_.type_size();
let bs = ti.type_.blck_size() as usize;
let n = (shape[0] * shape[1] * shape[2] * shape[3]) as usize;
let nbytes = (n / bs) * ts;
// Borrow the tensor bytes straight from the mmap'd part file (zero-copy —
// the file pages are shared with the CPU and GPU instead of a per-tensor copy).
let src = &raw[off..off + nbytes];

let mut strides = [0usize; 4];
strides[0] = ts;
strides[1] = strides[0] * (shape[0] / bs as i64) as usize;
for j in 2..4 {
    strides[j] = strides[j - 1] * shape[j - 1] as usize;
}

let mut tensor = Tensor::from_data_borrowed_with_strides(ttype, &shape, &strides, src);
}

The stride recomputation (loader lines 192–197) is deliberately identical to the parser's (§3.2 f) — same formula, same result, and a cross-check: if the two ever disagreed, tensors would be sliced with one geometry and walked with another. Inside Tensor (tensor.rs lines 201–214), the slice becomes Cow::Borrowed:

#![allow(unused)]
fn main() {
/// Create a weight tensor as a Borrowed slice of the mmap'd GGUF file
/// (zero-copy load — the file pages are shared with the CPU/GPU instead of
/// being copied per tensor). The slice must be 'static: the gguf loader
/// leaks the Mmap for the process lifetime.
pub fn from_data_borrowed_with_strides(
    ttype: TensorType,
    shape: &[i64; 4],
    strides: &[usize; 4],
    data: &'static [u8],
) -> Self {
    Tensor {
        ttype, shape: *shape, strides: *strides,
        data: std::borrow::Cow::Borrowed(data),
        name: String::new(),
    }
}
}

Cow (clone-on-write) is Rust's "borrowed until someone needs to modify it" enum: scratch/activation tensors are Cow::Owned(Vec<u8>), weights are Cow::Borrowed(&'static [u8]) — same type, two ownership stories, and a clone() of a weight tensor copies nothing (the t.clone() at model call sites is free for borrowed data; graph/cpu_backend.rs line 37 notes this).

And the final link in the zero-copy chain, doc 14's anchor (src/metal/ lines 2330–2344):

#![allow(unused)]
fn main() {
let page = 16384; // macOS page size on Apple Silicon
let base = data.as_ptr() as usize;
debug_assert!(base % page == 0, "mmap'd GGUF part not page-aligned");
let buf = unsafe {
    self.inner
        .device
        .newBufferWithBytesNoCopy_length_options_deallocator(
            NonNull::new(data.as_ptr() as *const std::ffi::c_void as *mut c_void)
                .unwrap(),
            (data.len() as u64) as usize,
            MTLResourceOptions::StorageModeShared,
            None,
        )
        .unwrap()
};
}

The GPU buffer is the file's pages — data here is exactly the &'static [u8] this doc's mmap produced, registered per part before any weight is wrapped as (buffer, offset) (loader lines 322–327). mmap's page-aligned return value is what makes newBufferWithBytesNoCopy legal; a heap Vec could never qualify.

3.3 Design choices (why this shape and not another)

mmap vs read()-into-heap. §2.5 gave the four wins (latency, RSS tolerance, page-cache sharing, GPU mapping). The honest costs: a mapping is address-space (irrelevant on 64-bit), pages can fault mid-inference if the file is truncated underneath you (llama.cpp documents this failure mode — minfer's download layer size-checks resumes for the same reason, per docs/ARCHITECTURE.md §9), and lifetime needs the leak discipline below. On net, for multi-GB read-only files on machines that also want GPU access to the same bytes, there is no contest — llama.cpp made the identical call (llama_mmap), and minfer is explicitly in that lineage.

Leak-for-'static vs Arc<MmapFile>. The tensors that borrow the mapping sit inside a ModelDef behind Box<dyn ModelDef>, get cloned, and get registered by name in two or three registries (CPU tensors, Metal buffers, CUDA device copies). Threading an Arc through all of them (or a self-referential struct crate) would infect every signature with a lifetime story that never varies in practice: the model lives for the whole process. Box::leak states that invariant once, at the only place that could violate it. The Drop impl still exists and is correct — it runs if a non-leaked mapping (e.g. a future tool) is dropped — but the load path never triggers it.

Raw bytes now, decode never-at-load. The tempting alternative is "dequantize everything to f32 at load" — one clean uniform representation, simple kernels. Why minfer doesn't: (1) it would quadruple memory (73.0 MiB → 519.3 MiB for one 0.5B-model tensor, from §2.3's table) and add a full file-size pass to load time; (2) the GPU backends want the quantized bytes — Metal kernels dequantize in registers and CUDA's MMQ path (int8 matrix-multiply quantized, doc 15) multiplies in quantized space directly, so decoding at load would force an f32 path that is both slower and numerically a different model; (3) the CPU kernels' speed comes precisely from SIMT/SIMD-friendly block layout (doc 10), not from pre-decoded rows. The engine's actual decode budget is spent where it pays: on activations, per matmul, at runtime ("Activations stay f32 … CPU quantizes to Q8_0 on the fly", docs/ARCHITECTURE.md §1.4).

Parser strictness as a feature. Every length is capped, every offset is cross-checked, duplicates are rejected, alignment must be a power of two, version > 3 fails. A more permissive parser would "work" until the first corrupt file — and then mis-slice weights and hand the failure to a matmul kernel hundreds of milliseconds later, ten layers deep in a forward pass. The error messages (with the tensor name, the expected and found offsets) turn a hex-editor session into a one-line diagnosis. The cost is a few branches per metadata element; metadata is 1.39% of the file.

Parse-then-borrow, not parse-and-copy. GgufContext deliberately does not hold tensor data (the gguf.cpp-parity comment says so, lines 621–622: "data … handled by the caller"). The parse produces coordinates; the mmap is the territory. That separation is what lets the same GgufModel serve the CPU path (slice → Cow::Borrowed), the Metal path (newBufferWithBytesNoCopy on the same pointer), and the CUDA path (device copies made once at registration, from the same slices) without the parser knowing any backend exists.

Why strides are computed at parse time, not in the kernel. GGUF stores ne[] (shape) but derives nb[] (strides) — they are not on disk. Computing them at parse time (rather than at kernel time) means every consumer agrees on geometry before any kernel runs, and the divisibility check (ne[0] % blck_size == 0) fails at load with the tensor's name, instead of as a mis-decoded value inside a dot product. ggml computes identical strides in ggml_new_tensor; minfer mirrors it in two places (parser + tensor factory) on purpose — the redundancy is an assertion.

3.4 Pitfalls & invariants

  • Alignment must be a power of two — ggml_pad's masking is only correct then, and the parser enforces it (lines 1488–1491) instead of trusting the file. A general.alignment of 48 would otherwise round down for some offsets and corrupt every subsequent slice.
  • Tensors are contiguous, and the parse proves it. Each stored offset must equal the running padded sum (lines 1694–1703). A file with gaps (or with padding computed at a different width) is rejected here — not at first inference.
  • The mapping must outlive every slice. The Box::leak at lines 1999 and 2038 is load-bearing: drop or unmap early and every weight tensor in the process is a use-after-free. Corollary: model loading is one-way — there is no unload-and-load-another within a process (the server keeps one model per slot, per docs/ARCHITECTURE.md §2).
  • ne[0] must be a multiple of the block size (line 1627). This is why a model whose vocabulary isn't divisible by 32 cannot be exported as Q8_0 without padding — docs/QWEN3-SUPPORT-PLAN.md records Qwen3's vocab 151,936 passing exactly this check ("151936 (÷32 ✓ for Q8_0 blocks)").
  • Version guards are endianness-aware: version & 0xFFFF == 0 catches byte-swapped files before any misread scalar can do damage (lines 1079–1082); v1 is refused outright, v4+ refused with a "file is newer than software" message (lines 1083–1092).
  • The data blob is never parsed, only addressed. If any future code "just reads one tensor" during load by scanning bytes instead of seeking to ctx.offset + ti.offset, it breaks the lazy-paging model (and on a cold cache, load time). The invariant: the parser reads bytes [0, data_start) and nothing beyond.
  • Part 0 is the entry, always. Loading …-00002-of-00002.gguf fails by design (resolve_splits rejects non-first parts, line 1982; split.no re-checked per part, lines 2041–2048), and the filename pattern must agree with split.count (lines 2028–2034). Doc 01's resolver can hand over any part of a split from a cache listing — this stage is where the wrong one is caught.
  • Name and key uniqueness: duplicate metadata keys (lines 1151–1157) and duplicate tensor names (lines 1531–1540) abort the parse — consumers do linear scans or hash maps keyed by name, and "first match wins" would make file ordering semantic. It isn't.

4. Observe & verify

  • minfer info <model> (main.rs line 492) runs this exact stage and dumps its output instead of continuing: you get the full metadata KV dump (dump_gguf_metadata, main.rs lines 1396–1449 — every key with its type and value, arrays itemized) followed by dump_key_tensors' name/type/shape table. On the 0.5B Q4_0 file: n_kv=26, n_tensors=291, general.architecture = "qwen2", qwen2.block_count = 24, output.weight q8_0 [896,151936], token_embd.weight q4_0 [896,151936].
  • The normal startup lines are this stage's stdout: Loading model: …, File: N bytes (X MB) in P part(s) (mmap lengths summed, lines 622–628), GGUF: 26 KV, 291 tensors (line 634), then Model loaded. after doc 03. On a split you'll see in 2 part(s); the number printed is the sum of the parts.
  • GgufContext::dump_metadata (gguf.rs lines 1721–1788) is a self-contained debug printer of the same data (kept #[allow(dead_code)] for tooling/tests).
  • Unit tests: cargo test gguf runs the split-pattern tests (split_file_info_parses_pattern at line 2059, resolve_splits_builds_all_parts at line 2078) — the filename grammar and part ordering, no model file needed.
  • The failure side is observable too: point minfer at a non-GGUF file and you get GGUF: invalid magic characters: '…', expected 'GGUF' then Error: failed to parse GGUF: … (main.rs lines 614–619); at a directory you get a candidate list (lines 582–610); at part 2 of a split, the split.no error.
  • What you can't see directly (by design): the mmap. top/Activity Monitor show RSS climbing during the first prefill as tensor pages fault in, not during Loading model — that is §2.5's point made visible. The one mmap-related timing you may notice was paid deliberately: Metal's part warm-up (~44 ms page/TLB setup, METAL_OPTIMIZATIONS.md #39) happens at registration, outside the Total: timing.
  • MINFER_TRACE / --dump-graph / debug_dump are graph-stage tools; they say nothing about this stage. The graph's weight names (visible in traces) are this stage's tensor names passed through untouched.

5. Cross-references

  • docs/ARCHITECTURE.md §6 — the quantization/tensor-layout summary this doc expands (ggml_pad, block sizes, split merging); §3 places the stage in the pipeline.
  • 01 — CLI args and model resolution — where the path came from (download/cache resolution, size-checked resume — the reason truncation-induced mmap faults are rare).
  • 03 — Model dispatch and weights — the direct consumer: general.architecture dispatch, HParams from qwen2.* keys, the merged tensor map, Tensor creation, GPU registration.
  • 04 — Tokenizer and template — the other consumer of the KV map (tokenizer.ggml.*, tokenizer.chat_template).
  • 10 — CPU matmul kernels — the block layouts catalogued here, used in anger (dot products on 18-byte Q4_0 blocks).
  • 14 — Metal backend — the payoff of mmap: newBufferWithBytesNoCopy, (buffer, offset) weight wrapping, part warm-up.
  • 15 — CUDA backend — the third consumer of the same slices (registration-time device copies; Q6_K's padded-224 variant shows why per-type layout knowledge matters end to end).
  • docs/GLOSSARY.md — backstop definitions for every term used here.
  • llama.cpp provenance: gguf.rs comments cite gguf.cpp/ggml.c line numbers throughout (e.g. lines 334, 600, 851) — the parser is a readable diff against upstream.

← 01 — CLI args and model resolution · Index · 03 — Model dispatch and weights →

03 · Model dispatch and weights

Stage: GGUF parsed (02) → this stage: the file becomes a runnable model, with weights wired into every backend → tokenizer + template (04). Code: models/mod.rs::load_model (dispatch), models/qwen2/loader.rs::load and models/qwen3/loader.rs::load (the two implementations), main.rs:637-657 (GPU init + legacy KV cache), tensor.rs (Cow<'static,[u8]> weight bytes), graph/alloc.rs::register_weight, src/metal/::register_weight / cuda.rs::register_weight (per-backend registries), models/qwen2/graph.rs::weights_on_gpu (the GPU participation gate).


1. Background — where this stage sits

Doc 02 left us holding a GgufModel. It is a container, not yet a model: a parsed metadata table (key/value pairs), a tensor index (name, quant type, shape, offset for every tensor), and the raw data blob of the file memory-mapped into the process. Nothing in that container knows what a transformer is. If you asked it "how many layers does this model have?", it can only answer "there is a metadata key somewhere that might say". The bytes are all there, but they have no meaning yet.

This stage gives them meaning. Two things happen, and they happen in this order. First, the engine initializes its GPU backends — Metal on macOS, CUDA when the binary was built with --features cuda and an NVIDIA GPU is present. Second, the engine looks at one metadata string, general.architecture, and dispatches the whole file to the model implementation that matches it: "qwen2" goes to the Qwen2/Qwen2.5 loader, "qwen3" to the Qwen3 loader. The loader then reads the hyperparameters (the dimension numbers of the architecture — layer count, head counts, embedding width; these are the settings that were chosen before training and are never learned) and gathers every weight tensor by name. A weight is one block of the model's learned numbers: the big matrices that projections multiply by, the small gain vectors of the normalization layers, and the optional bias vectors. "Learned" means fixed by training; inference never changes them.

The output of this stage is a single object behind the ModelDef trait — for example Qwen2Model — holding hyperparameters plus every weight tensor, with its bytes still sitting in the mmap'd file. Alongside it, each available backend has been told about the weights it cares about. Those registrations are what make the next stages work: doc 05 will build a compute graph whose nodes reference weights by name, and docs 07–08 will allocate buffers and execute ops that fetch their weight through the registries this stage fills.

Why is the ordering "GPU init first, then dispatch" and not the other way round? Because loading is not a read-only operation. The loaders register weight tensors into the GPU registries while they walk the tensor list, so the registry must already exist. Get this wrong in one specific way — CUDA initialized lazily in the middle of a load — and you get a half-registered model whose backend gate flips between the first and the second half of the tensor list. The loader defends against exactly that (we will see the guard in §3.2).

What would break without this stage? Almost everything downstream, and in ways that are loud rather than subtle. The graph builder needs the hyperparameters to know how many layers to emit and what shapes the ops have. The allocator needs n_kv_embd (the per-layer key/value width) to size the persistent KV regions — a KV cache stores, per layer, the key and value vectors of every token generated so far (doc 11 covers the mechanism). The samplers need the end-of-sequence token ids, which are metadata this stage parses. And the backends need the name→weight registries: a matmul whose weight is not registered fails at execute time with weight '...' not registered. In short, this stage turns a parsed file into a machine the rest of the pipeline can drive. One thing to keep in mind from here on: nothing is dequantized at load — the weights stay exactly the bytes the file shipped, 4-bit nibbles and block scales and all. §2.3 explains why that is a feature, not laziness.

2. Principle — how it works and why

2.1 The cast of characters

The stage has five players. It is worth naming them once, because the rest of the doc is just their handshakes.

  1. GgufModel (doc 02) — metadata key/values, the tensor index, and &'static [u8] slices into the mmap'd file parts.
  2. load_model (models/mod.rs) — reads general.architecture and picks the implementation. This is the dispatch.
  3. The per-architecture loader (models/qwen2/loader.rs, models/qwen3/loader.rs) — parses hyperparameters, gathers tensors by name, registers them with the backends, and assembles the model struct.
  4. ModelDef (models/mod.rs) — the architecture-agnostic interface the rest of the engine talks to. Downstream code never says "Qwen2"; it says "whatever model is loaded, give me n_layer(), build me a graph, format this chat".
  5. The backend registries — CpuBackend (a name→tensor map), MpsState (a name→(Metal buffer, offset) map), CudaState (a name→(device pointer, size) map). "Registry" here just means a hash map from a weight's GGUF name to whatever the backend needs in order to use it.
 GgufModel ──"general.architecture"──► load_model ──► qwen2::loader::load
                                                          │
                           ┌──────────────────────────────┴───────────────┐
                           ▼                                              ▼
                HParams (dims, token ids)               Tensor per weight (bytes
                           │                            borrowed from the mmap)
                           ▼                                              │
                 Qwen2Model : ModelDef ────t.clone()────► GraphAllocator  │
                           │                     (CPU: name → Tensor)    │
                           │                                              │
       GPU init happens BEFORE load:                                      │
         MpsState::init() → name → (MTLBuffer, offset) ───────────────────┤
         CudaState::init   → name → (device ptr, size) ───────────────────┘
                                          (graph nodes reference weights
                                           by name only, not by pointer)

2.2 Dispatch on a metadata string

GGUF writes the architecture family into a top-level metadata key. minfer reads it once and matches it against the implementations it ships:

  • "qwen2" → Qwen2 / Qwen2.5 (the 0.5B, 1.5B, 7B… checkpoints all report qwen2; Qwen2.5 differs from Qwen2 only in training, not in tensor layout).
  • "qwen3" → Qwen3 dense (same overall wiring plus two twists we cover in §3.2: a decoupled head dimension and per-head Q/K norms).

The alternative — "try each loader until one succeeds" — is strictly worse, and the reasons are worth spelling out because they shape the whole design.

Determinism. A string match is a total function: one key, one answer. A trial-parse loop depends on what each loader happens to tolerate, and both of minfer's loaders deliberately accept llama.* metadata keys as a fallback (some fine-tunes re-label their metadata). Two tolerant loaders plus a try-loop is a recipe for loading a Qwen3 file as "some kind of qwen2" — a parse success that is semantically wrong and would corrupt attention shapes.

One clear error. When the string does not match anything, the user gets Unsupported architecture: 'llama' — the actual offending string — and the run stops before any GPU work happens. A try-loop's failure mode is instead "N loaders each printed a different complaint", or worse, a quiet success.

Load is side-effectful. This is the mechanical clincher. Loading registers weights into GPU registries, and on CUDA a registration is a one-way upload into device memory. A "try and reject" loader would leave half a model uploaded with no clean way to un-register (see §3.4: CUDA deliberately never frees stale weight buffers). Dispatching first, then loading exactly once, keeps the side effects all-or-nothing too.

2.3 Weights stay raw bytes

The single most important data decision of this stage: a loaded weight tensor is a view, not a copy. The Tensor type stores its payload as std::borrow::Cow<'static, [u8]> — a Rust "clone-on-write" enum that here is always the Borrowed variant, i.e. a plain (pointer, length) slice pointing into the mmap'd GGUF file. The 'static lifetime works because doc 02's loader Box::leaks each mmap for the process lifetime (gguf.rs:1804-1810).

Three consequences follow, and each one answers a "why not the obvious alternative":

Why not dequantize to f32 at load? Because every consumer wants the bytes as they are. The CPU matmul kernels (doc 10) consume packed 4-bit nibbles directly — they dequantize a block on the fly inside the dot product, one 32-value block at a time, and never materialize the full f32 tensor. The GPU kernels do the same in shaders (Metal) or stream raw quantized bytes (CUDA's int8 MMQ path). Meanwhile the memory math is brutal for the alternative: a Q4_0 block is 18 bytes for 32 values (2-byte fp16 scale + 16 bytes of nibbles, block.rs:53-56), which is ≈ 0.56 bytes per value. A 0.5-billion-parameter checkpoint is then ≈ 0.28 GB of file bytes; dequantized to f32 it would be 2 GB — 7× more, copied at load time, for zero benefit.

Why does cloning a Tensor not copy bytes? Because Cow::Borrowed clones as pointer + length. That is what makes it affordable for the graph path to re-register all weights on every graph (re)build (t.clone() at qwen2/graph.rs:354) — the clone duplicates a small struct and a name string, not gigabytes. The one caveat — registries still guard against re-registration when a caller hands them owned bytes — has a measured war story attached, told with excerpt 10.

Why do GPU backends get their own representation? Because "the weight" means something different per backend: on CPU it is the mmap bytes; on Metal it is a byte range inside a shared-memory buffer; on CUDA it is a pointer into device memory. The registry abstracts exactly that, and §2.4 walks each one.

2.4 What "the weight is on the GPU" means, per backend

The phrase "weights on GPU" hides three quite different mechanisms. Getting them straight explains everything the loader does.

CPU — registration is free. CpuBackend keeps HashMap<String, Tensor>. Registering inserts the tensor struct. The bytes were already in the process (they are the mmap pages, faulted in on first touch), so the "registration" moves no data at all. On CPU, "the weight is registered" means only "the name resolves to a byte range".

Metal — the GPU reads the same physical pages. On Apple Silicon, CPU and GPU share one physical memory ("unified memory"). Metal exposes buffers that both sides can address (StorageModeShared). The trick is that minfer does not copy each weight into such a buffer. Before any weight is registered, the loader hands each mmap'd file part to MpsState::register_part, which wraps the whole mmap in one Metal buffer via newBufferWithBytesNoCopy — "no copy" is the API's name and its contract. Each individual weight is then registered as (buffer, byte offset) into that one buffer. The GPU reads the file's pages directly; there is no GPU-side allocation and no memcpy, ever. The one cost is a first-touch one: the very first GPU access to file-backed pages pays ~44 ms of page/TLB setup, which the loader deliberately triggers once at load time, outside the timed inference window (src/metal/runtime.rs).

CUDA — one upload, resident forever. NVIDIA GPUs have discrete memory (device memory, VRAM) that the CPU cannot address; bytes must be copied across the PCIe bus ("H2D", host-to-device). CudaState::register_weight does cudaMalloc for the tensor's size, one cudaMemcpy H2D, and stores (device pointer, size) under the name. That copy happens exactly once, at load. From then on the weight is resident: every decode step reads it from device memory at GPU bandwidth instead of re-uploading. This is the Phase 7 thesis in one sentence — a graph backend is only fast if the weights are already addressable on the executing device, so registration is a load-time job, not a per-step one. For scale: the 7B Q4_K_M model is ~4.4 GB of weights; the alternative (per-step host staging) is exactly what the old imperative path did for activations, and the CUDA campaign measured such host round-trips at "~6 PCIe round trips × 24 layers ≈ 144 DMA operations per decode step, 2–7 ms" (docs/CUDA_OPTIMIZATION.md). Resident weights delete that entire class of cost.

2.5 The gate: GPU participation is all-or-nothing

Per §2.4, a backend can only execute an op if the op's weight lives where the op runs. The graph's backend assignment is per-op (doc 06), but the weights constrain it globally, so before building anything the graph path asks: "is every weight this model will use registered — and kernel-supported — on this backend?" Two functions do this:

  • Qwen2Graph::weights_on_gpu (Metal): every weight name must be present in MpsState's registry (models/qwen2/graph.rs:636-677).
  • Qwen2Graph::weights_on_cuda (CUDA): every weight must be registered and of a type a kernel exists for — e.g. the embedding gather supports every registered type except Q4_1 (models/qwen2/graph.rs:695-756).

The result — metal_on || cuda_on — is stored in CParams.gpu, which is part of the reuse identity: the fingerprint that decides whether a cached graph can be reused (doc 13). Flip any env toggle or unplug the eGPU and the next forward rebuilds the graph rather than executing stale assignments.

Why all-or-nothing rather than "put what fits on the GPU, layer by layer"? The old imperative engine had a per-layer fallback, and it is preserved in docs/ARCHITECTURE.md Appendix A.3 as a cautionary diagram: the moment one layer failed its GPU check, the hidden state had to cross back to host memory, the KV cache had to be synced to CPU, and the rest of the layers ran on CPU — per token. In the graph path that cost is even sharper: KV regions live on the backend that executes attention (by construction), so one CPU-resident layer would force the whole layer's KV traffic across the bus every step. The one-buffer-at-a-split-boundary design (doc 08) exists precisely so cross-backend traffic happens a handful of times per forward, not per op. All-or-nothing is how the design keeps that promise: either the backend can host everything the graph reads, or it does not participate at all.

2.6 Two architectures, one interface

Qwen2 and Qwen3 differ in exactly two load-time-visible ways, and both exist to keep the graph builder simple.

First, the head dimension. In Qwen2, the per-head size is derivable: n_embd_head = n_embd / n_head. Qwen3 broke that identity — the 0.6B model has n_embd / n_head = 64 but its keys are 128-wide — so the loader reads qwen3.attention.key_length explicitly and asserts the K weight's actual output width agrees (qwen3/loader.rs:311-326). Trust the bytes, not the derived formula.

Second, per-head Q/K RMSNorm: Qwen3 normalizes each head's query and key vectors before RoPE, with small learned gain vectors (q_norm, k_norm). The loader stores them; the graph builder has a dedicated qk_norm op (graph/builder.rs:99-119) that consumes them. The loader's job is recognizing that these tensors exist and must not be lost — the minfer info listing truncates names, but they are in the file. Everything else — the loader shape, the registration calls, the fused-QKV and fused-FFN concat weights — is deliberately mirrored between the two loaders, so a new architecture is a copy-and-edit job (§5).

3. Implementation

3.1 Data in / data out

In: the GgufModel from doc 02. Concretely, per part: ctx.kv (metadata key/values), ctx.info (tensor index; each entry has name, type_, ne[4] shape, and offset — offset from the start of the part's data section), and part.data: &'static [u8] (the mmap'd bytes). A weight's file position is ctx.offset + ti.offset, where ctx.offset is where the data section starts in the file (gguf.rs:603-619).

Out: three things.

  1. Box<dyn ModelDef> — the polymorphic model object (Qwen2Model / Qwen3Model): HParams + tok_embd + output_norm + output + optional output_b + one LayerWeights per layer.
  2. Populated backend registries: CPU always; Metal's (buffer, offset) entries when MPS initialized; CUDA's device copies when a device exists.
  3. The KV element-type decision (f16 vs f32), set once from the model's dimensions before any forward runs.

Shapes to internalize now (they recur in every later doc): a GGUF weight matrix is stored with shape metadata [in, out] (ne[0] = input dim, fastest-varying) while memory is row-major [out][in] — so wq of a 0.5B model is 896×896 and wk is 128×896 as bytes even though its logical projection is 896 → 128. The embedding table token_embd.weight is [n_vocab, n_embd]: one row per vocabulary entry, each row the vector that token id looks up to. Norm weights (attn_norm, ffn_norm, output_norm) are 1-D f32 gain vectors of length n_embd (a gain is just a learned per-feature multiplier applied after normalizing); biases (bq, bk, bv, attn_output has none, output.bias optional) are 1-D f32 too. Quantized matmul weights are one of Q4_0/Q4_1/Q5_0/Q5_1/Q8_0 (32-value blocks) or Q4_K/Q5_K/Q6_K (256-value super-blocks) — quantization stores values in fewer bits, grouped into blocks that share a scale factor.

3.2 Key code

Excerpt 1 — the startup order in main.rs: GPU init, dispatch, legacy KV cache. (src/main.rs:637-657; the KV block below is historical — see the forward note under the annotations)

#![allow(unused)]
fn main() {
    // === GPU backends ===
    #[cfg(target_os = "macos")]
    metal::MpsState::init();
    #[cfg(feature = "cuda")]
    cuda::CudaState::init_with_gpu(gpu);
    // On CPU/Metal builds `--gpu` is a no-op: it is parsed but unused.
    #[cfg(not(feature = "cuda"))]
    let _ = gpu;

    // === Load model (dispatches on general.architecture) ===
    let model = models::load_model(&gguf_model).expect("load model");
    ...
    // === KV Cache ===
    let n_kv_embd = model.n_kv_embd();
    let n_layer = model.n_layer();
    let mut kv_cache = cache::KVCache::new(n_layer, n_kv_embd, params.n_ctx);
}

Forward note (#252, 2026-10-02): the // === KV Cache === block above is gone. #244 had already deleted the KVCacheLayer storage; #252 then deleted the empty marker src/cache.rs, its mod cache; declaration and the vestigial &mut KVCache parameter of ModelDef::forward and forward_graph, so main.rs loads the model and the tokenizer and nothing else. There is no per-load KV allocation to narrate at all — the graph allocator's persistent regions are the only KV store (#244, #252).

Annotations: MpsState::init() is a OnceLock singleton init — inside, it honors MINFER_DISABLE_MPS by returning None, so "disabled" and "no device" are the same state downstream (src/metal/runtime.rs). CUDA likewise honors MINFER_DISABLE_CUDA and takes the --gpu N index here. The load_model call is where this entire doc's work happens — note .expect: an unsupported architecture is fatal, by design (§2.2). The final three lines in the excerpt allocated the legacy KV cache, whose only remaining job (after #244 deleted its storage) was to satisfy the ModelDef::forward signature's &mut KVCache parameter — a parameter no path read, and the claim that it cost ≈ 100 MB of zeroed memory stopped being true with #244. #252 deleted the argument and the type rather than keep a dead allocation alive for the API shape.

Excerpt 2 — the dispatch itself. (src/models/mod.rs:95-112)

#![allow(unused)]
fn main() {
pub fn load_model(model: &GgufModel) -> Option<Box<dyn ModelDef>> {
    let ctx = &model.parts[0].ctx;
    let arch = ctx.get_key_val_str("general.architecture")?;
    match arch.as_str() {
        "qwen2" => {
            let m = qwen2::loader::load(model)?;
            Some(Box::new(m))
        }
        "qwen3" => {
            let m = qwen3::loader::load(model)?;
            Some(Box::new(m))
        }
        other => {
            eprintln!("Unsupported architecture: '{}'", other);
            None
        }
    }
}
}

Annotations: part 0 is authoritative for metadata even in a multi-part split (doc 02 merged the tensor index across parts; metadata comes from the first). ? on the string lookup means a file without the key is "no model", not a panic — the caller reports it. The error branch prints the offending string, which is what makes a mistyped or future architecture diagnosable in one glance.

Excerpt 3 — the interface everything downstream codes against. (src/models/mod.rs:22-45, :77-85 — abridged)

#![allow(unused)]
fn main() {
pub trait ModelDef: Send + Sync {
    fn forward(&self, tokens: &[u32], positions: &[usize],
               n_out: usize, n_ctx: usize) -> Vec<f32>;
    /// Downcast helper for the graph path's weight registration.
    fn as_any(&self) -> &dyn std::any::Any;
    /// Build the declarative compute graph for one forward step (Phase 5).
    /// Topology is a deterministic function of `params` (reuse invariant).
    fn build_graph(&self, _params: &GraphParams) -> ComputeGraph { ... }
    /// Graph-based forward with a caller-provided cache and explicit context
    /// size (server / multi-slot path).
    fn forward_graph_cached(&self, tokens: &[u32], positions: &[usize],
                            n_out: usize, n_ctx: usize, cache: &mut GraphCache)
                            -> Vec<f32> { ... }
    fn special_tokens(&self) -> SpecialTokens;
    fn n_layer(&self) -> usize;
    fn n_head_kv(&self) -> usize;
    fn n_embd_head(&self) -> usize;
    fn n_kv_embd(&self) -> usize;
    fn n_vocab(&self) -> usize;
    fn rope_style(&self) -> RopeStyle;
}
}

This trait looks wide for an interface with two implementations, and that is the point: each method exists because a downstream stage needs it and must not know which architecture it is talking to.

MethodWho consumes it, and for what
forward, forward_graph_cachedthe CLI loop (doc 09) and the server's per-slot path — both just "run a forward"; the default forward_graph routes to the graph
build_graphdoc 05: the pure-IR graph builder; topology is a function of GraphParams only
n_layerloader-loop sizing here, the graph builder's per-layer loop (doc 05), the legacy KV cache above
n_head_kv, n_embd_head, n_kv_embdGQA (grouped-query attention: fewer K/V heads than query heads) head mapping and strides (doc 11), and the KV region width n_kv_embd × n_ctx (doc 07)
n_vocablogits width — the sampler's input size (doc 12)
special_tokensthe sampler's stop condition: main.rs fetches eos/im_end ids once and checks every sampled token against them (main.rs:839,900,1020)
rope_styledoc 11: RoPE (rotary positional encoding) has two layout styles — Qwen's non-interleaved vs Llama's interleaved — and the vec-op must be told which
as_anylets graph code downcast to the concrete model when it needs specifics
Send + Syncthe HTTP server shares the model across threads (Arc<dyn ModelDef>)

Excerpt 4 — hyperparameter parsing with the dual metadata prefix. (src/models/qwen2/loader.rs:117-156, abridged)

#![allow(unused)]
fn main() {
    // Try qwen2 prefix first, fall back to llama/generic
    let n_embd = get_i64(ctx, "qwen2.embedding_length")
        .or_else(|| get_i64(ctx, "llama.embedding_length"))?;
    let n_head = get_i64(ctx, "qwen2.attention.head_count")
        .or_else(|| get_i64(ctx, "llama.attention.head_count"))?;
    let n_head_kv = get_i64(ctx, "qwen2.attention.head_count_kv")
        .or_else(|| get_i64(ctx, "llama.attention.head_count_kv"))
        .unwrap_or(n_head);                       // no GQA ⇒ KV heads = Q heads
    let n_layer =
        get_i64(ctx, "qwen2.block_count").or_else(|| get_i64(ctx, "llama.block_count"))?;
    ...
        f_norm_rms_eps: get_f32(ctx, "qwen2.attention.layer_norm_rms_epsilon")
            .or_else(|| get_f32(ctx, "llama.attention.layer_norm_rms_epsilon"))
            .unwrap_or(1e-6),
        rope_freq_base: /* ... llama.* fallback ... */ .unwrap_or(10000.0),
        rope_style: RopeStyle::NonInterleaved,
        n_kv_embd: n_head_kv * (n_embd / n_head), // default, updated from K weight below
}

Annotations: every dimension is a fallback chain — the architecture's own prefix first, then the LLaMA-family prefix that several fine-tunes use. The .unwrap_or defaults are also data: n_head_kv defaulting to n_head means "no grouped-query attention"; rms_eps = 1e-6 and freq_base = 10000.0 are the values llama.cpp would use. n_vocab is not read from a metadata key at all but counted from the tokenizer.ggml.tokens array — the tokenizer data is the ground truth (it arrives next stage). And n_kv_embd starts as the naive product, purely so the struct is initialized; the loader immediately overwrites it from the K weight's real shape.

Excerpt 5 — a weight tensor is born as a borrowed slice, then handed to the GPU registries. (src/models/qwen2/loader.rs:176-220, abridged; the CUDA branch at :221-285 is discussed in the annotations)

#![allow(unused)]
fn main() {
fn load_tensor(ctx: &GgufContext, raw: &'static [u8], ti: &GgufTensorInfo) -> Tensor {
    let ttype = TensorType::from_ggml_type(ti.type_);
    ...
    let off = ctx.offset + ti.offset as usize;
    let ts = ti.type_.type_size();      // bytes per block
    let bs = ti.type_.blck_size() as usize; // values per block
    let n = (shape[0] * shape[1] * shape[2] * shape[3]) as usize;
    let nbytes = (n / bs) * ts;
    // Borrow the tensor bytes straight from the mmap'd part file (zero-copy —
    // the file pages are shared with the CPU and GPU instead of a per-tensor copy).
    let src = &raw[off..off + nbytes];
    ...
    let mut tensor = Tensor::from_data_borrowed_with_strides(ttype, &shape, &strides, src);

    // Register weight tensors with GPU backends.
    #[cfg(target_os = "macos")]
    if let Some(mps) = crate::metal::MpsState::get() {
        if matches!(ttype,
            TensorType::Q4_0 | TensorType::Q4_1 | TensorType::Q4_K
          | TensorType::Q5_0 | TensorType::Q5_1 | TensorType::Q5_K
          | TensorType::Q6_K | TensorType::Q8_0)
        {
            mps.register_weight(&ti.name, tensor.data());
        } else if ttype == TensorType::F32 {
            mps.register_weight(&ti.name, tensor.data());
        }
    }
    // #[cfg(feature = "cuda")] branch: same shape, more work — see below.
    tensor
}
}

Annotations: byte size comes from the GGML type's block size, not from a TensorType guess — the comment notes this is "always correct regardless of TensorType mapping". src is a sub-slice of the mmap: constructing the tensor did one range check and zero copies, and the registration is woven into the same walk rather than done as a second pass over the model. On Metal, essentially everything quantized plus f32 gets registered — the shader kernels handle the block formats natively. The CUDA branch is pickier and does more work at registration: Q6_K weights are repacked into padded 224-byte slots (raw blocks are 210 bytes, which forces byte-granular GPU loads — padding restores 16-byte-aligned vector loads), a f32-pair plane may be precomputed for the Q4_K kernel, and unsupported types clear a fast matmul-mode flag because a mode that assumed certain weight layouts would otherwise read garbage. The point for this doc: registration is where per-backend representation is decided — bytes for CPU, (buffer, offset) for Metal, device allocation (+optional repack) for CUDA.

Excerpt 6 — two load-time side decisions: the KV element type, and the true KV width. (src/models/qwen2/loader.rs:310-317 and :495-498; Qwen3's guarded version at src/models/qwen3/loader.rs:311-331)

#![allow(unused)]
fn main() {
    // KV cache element type (GPU path): auto-select f16 for the 7B class (KV
    // bandwidth-bound decode) unless MINFER_CACHE_TYPE overrides. Must run
    // before the first forward (kv_cache_is_f16 reads the OnceLock).
    #[cfg(target_os = "macos")]
    crate::metal::set_kv_cache_type(hparams.n_layer as usize, hparams.n_kv_embd as usize);
    // 8b: CUDA side shares the same policy and MINFER_CACHE_TYPE override.
    #[cfg(feature = "cuda")]
    crate::cuda::set_kv_cache_type(hparams.n_layer as usize, hparams.n_kv_embd as usize);
    ... // (later, after the per-layer weights are loaded:)

    // Override n_kv_embd from layer 0 K weight's actual output dimension
    if let Some((_, ti)) = tensor_map.get(&tn::attn_k(0)) {
        hparams.n_kv_embd = ti.ne[1];
    }
}
#![allow(unused)]
fn main() {
    // qwen3/loader.rs — same override, resolved BEFORE the KV type pick, plus
    // an assert that is only sound for Qwen3 (Qwen2's `n_kv_embd` may
    // legitimately differ from `n_head_kv × n_embd_head`, so it cannot assert):
    if let Some((_, ti)) = tensor_map.get(&tn::attn_k(0)) {
        hparams.n_kv_embd = ti.ne[1];
        // sanity: kv dim must equal n_head_kv * n_embd_head (catches a wrong
        // key_length fallback before it silently corrupts attention)
        assert_eq!(
            hparams.n_kv_embd, hparams.n_head_kv * hparams.n_embd_head, ...)
    }
}

Annotations: the K projection is a real matrix sitting in the file — on the 0.5B it is [896 → 128] — so its output width is the KV width, whatever the head-count metadata might imply; the loader trusts it over any derived value. The Qwen3 loader reads it before the KV type pick (its comment says why: the f16 auto-select multiplies n_layers × n_kv_embd), and its assert turns a wrong key_length fallback into a load-time crash instead of silently corrupting attention. The policy set_kv_cache_type implements (src/metal/policy.rs): if MINFER_CACHE_TYPE says f16/f32, obey (since C4 the value is parsed strictly on every device — an unknown spelling, or q8_0 on a backend whose attention kernel has no packed read, fails the load instead of quietly running f32); otherwise auto-select — f16 (half precision: 2 bytes per value instead of 4) when n_layers × n_kv_embd ≥ 8192, i.e. models big enough that decode is KV-bandwidth-bound (measured −1 ms/token on the 7B at 2K context), f32 for small models where f16 measured ~3% slower.

Excerpt 7 — the merged tensor index and name lookup. (src/models/qwen2/loader.rs:329-343)

#![allow(unused)]
fn main() {
    // Merged tensor index across all split parts (llama.cpp weights_map): each
    // tensor lives in the part that lists it, read from that part's own data.
    let mut tensor_map =
        std::collections::HashMap::<String, (usize, &GgufTensorInfo)>::new();
    for (pi, part) in model.parts.iter().enumerate() {
        for ti in &part.ctx.info {
            tensor_map.insert(ti.name.clone(), (pi, ti));
        }
    }
    let load_one = |n: &str| -> Option<Tensor> {
        tensor_map.get(n).map(|(pi, ti)| {
            let part = &model.parts[*pi];
            load_tensor(&part.ctx, &part.data, ti)
        })
    };
}

The loader then reads weights by canonical name: token_embd.weight, output_norm.weight, output.weight, blk.{i}.attn_norm.weight, blk.{i}.attn_q.weight, … — the names come from a small tensor_names module (models/qwen2/mod.rs:124-166), so a naming convention change is a one-file edit. Two notable lookups: output falls back to the embedding table when absent (load_one(tn::OUTPUT).unwrap_or_else(|| tok_embd.clone()) — weight tying: small models reuse the embedding table as the final projection instead of shipping a second matrix), and every per-layer tensor is Option because Qwen3 has no biases while Qwen2.5-7B does.

Excerpt 8 — Metal's zero-copy registry. (src/metal/, annotated condensation of register_part :2320-2372 and register_weight :2374-2424)

#![allow(unused)]
fn main() {
    pub fn register_part(&self, data: &'static [u8]) {
        let page = 16384; // macOS page size on Apple Silicon
        let base = data.as_ptr() as usize;
        debug_assert!(base % page == 0, "mmap'd GGUF part not page-aligned");
        let buf = unsafe {
            self.inner.device
                .newBufferWithBytesNoCopy_length_options_deallocator(
                    ptr, data.len(), MTLResourceOptions::StorageModeShared, None)
                .unwrap()
        };
        self.inner.mmap_parts.lock().unwrap().push((base, data.len(), buf.clone()));
        // GPU-side warm-up (#39): the FIRST GPU access to file-backed pages
        // costs ~44 ms of one-time page/TLB setup → do it here, at load.
    }

    pub fn register_weight(&self, name: &str, data: &[u8]) {
        let force_copy = std::env::var("MINFER_WEIGHT_COPY").map_or(false, |v| v == "1");
        // Zero-copy path: the weight is a slice of a registered mmap'd part
        // → (part buffer, offset). The GPU reads the mapped file pages
        // directly — no CPU→GPU memcpy, no GPU-side allocation.
        let entry = if !force_copy {
            parts.iter().find(|(base, len, _)| ptr >= *base && ptr + data.len() <= base + len)
                .map(|(base, _, buf)| (buf.clone(), (ptr - base) as u64))
        } else { None };
        let (buf, off) = match entry {
            Some(e) => e,
            None => { /* copy into a fresh shared buffer (offset 0) */ }
        };
        self.inner.weights.lock().unwrap().insert(name.to_string(), (buf, off));
    }
}

Annotations: the loader calls register_part for every mmap'd part before registering any weight (qwen2/loader.rs:319-327) — the ordering is load- bearing, because register_weight locates its zero-copy entry by finding the part that contains the pointer. StorageModeShared on Apple Silicon means one physical allocation both CPU and GPU address; "zero-copy" is literal. The copy fallback exists for the two cases where bytes are not file pages: the fused attn_qkv/ffn_gu concat weights (built in RAM at load, qwen2/loader.rs:391-426,446-491) and MINFER_WEIGHT_COPY=1, an A/B switch that makes the cost of the zero-copy path measurable.

Excerpt 9 — CUDA's upload-once registry. (src/cuda/methods/weights.rs:18-97, abridged)

#![allow(unused)]
fn main() {
    pub fn register_weight(&self, name: &str, data: &[u8]) {
        if data.is_empty() { return; }
        {
            let w = self.weights.lock().unwrap();
            if let Some((_, size)) = w.get(name) {
                if *size == data.len() {
                    // Device weights are immutable: same name + size ⇒ the
                    // same GGUF tensor ... Reuse the existing device copy
                    // instead of leaking one buffer per load.
                    return;
                }
                // Different size ...: replace the entry. The stale buffer is
                // deliberately NOT freed — a live captured graph may still
                // reference it; ...
            }
        }
        let mut ptr: *mut std::ffi::c_void = std::ptr::null_mut();
        let err = unsafe { cudaMalloc(&mut ptr, data.len()) };
        ...
        let err = unsafe {
            cudaMemcpy(ptr, data.as_ptr() as *const c_void, data.len(),
                       CUDA_MEMCPY_HOST_TO_DEVICE)
        };
        ...
        self.weights.lock().unwrap().insert(name.to_string(), (CudaPtr(ptr), data.len()));
    }
}

Annotations: three details repay attention. (1) The dedup check makes re-registration a no-op — graph rebuilds and unit tests that reload a model must not each leak another full-weight-set upload (~4.4 GB on the 7B). (2) The refusal to free stale buffers is deliberate, not sloppy: a captured CUDA Graph (doc 15) holds raw device pointers; freeing under it would be use-after-free. (3) A plain registration clears any stale "padded" flag for that name, so a second model reusing a tensor name with a non-Q6_K type cannot be dispatched through the padded-224 kernel on a raw-210 buffer (a Phase 8 review finding).

Excerpt 10 — CPU registration: the cheapest one. (src/graph/cpu_backend.rs:34-47)

#![allow(unused)]
fn main() {
    /// Register a weight tensor by name (Phase 6 wires this from the model).
    pub fn register_weight(&mut self, name: &str, t: Tensor) {
        // Skip re-registration of an already-known weight: Tensor carries its
        // bytes as Cow::Owned, so the `t.clone()` at the model call sites
        // deep-copies the full weight set (~4.4 GB on 7B) on EVERY graph
        // (re)build — measured as a ~635 ms pure-CPU stall at the
        // prefill→decode graph switch (no CUDA calls, no kernels). Model
        // weights are immutable after load (weights_version guards any future
        // change), so a same-name registration always carries the same data.
        if self.weights.contains_key(name) {
            return;
        }
        self.weights.insert(name.to_string(), t);
    }
}

The comment is the whole lesson: with borrowed bytes, even the unguarded insert is cheap; the guard exists because one call path produced owned clones. The graph allocator simply forwards to it (graph/alloc.rs:135-138: self.cpu.register_weight(name, t)), which is why the allocator's registration costs nothing on CPU.

Excerpt 11 — the graph's own registration pass, at first build. (src/models/qwen2/graph.rs:344-378, abridged)

#![allow(unused)]
fn main() {
    /// Register every weight the graph references on the allocator's backend.
    pub(crate) fn register_graph_weights(model: &Qwen2Model, alloc: &mut GraphAllocator) {
        for t in [&model.tok_embd, &model.output_norm, &model.output, &model.output_b] {
            if let Some(t) = t {
                let name = t.name.clone();
                alloc.register_weight(&name, t.clone());
            }
        }
        for l in &model.layers {
            for t in [&l.attn_norm, &l.wq, &l.bq, &l.wk, &l.bk, &l.wv, &l.bv,
                      &l.wo, &l.ffn_norm, &l.ffn_gate, &l.ffn_up, &l.ffn_down] {
                if let Some(t) = t {
                    let name = t.name.clone();
                    alloc.register_weight(&name, t.clone());
                }
            }
        }
    }
}

Annotations: this runs inside forward_cached on graph build only (guarded by the reuse check), and the t.clone() is the cheap borrowed-clone of §2.3. Twelve entries per layer is the Qwen2 inventory: 7 matmul weights (wq, wk, wv, wo, ffn_gate, ffn_up, ffn_down), 2 norm gains, and 3–4 biases. The list is deliberately spelled out — not derived by reflection — so the compiler catches field renames in both the registration and the gate (excerpt 12), which must enumerate the same weights.

Excerpt 12 — the Metal participation gate. (src/models/qwen2/graph.rs:634-677, names list abridged)

#![allow(unused)]
fn main() {
    /// Every weight the graph reads must be GPU-registered for the Metal path.
    #[cfg(target_os = "macos")]
    fn weights_on_gpu(model: &Qwen2Model) -> bool {
        let names: Vec<String> = {
            let mut v = Vec::new();
            for t in [&model.tok_embd, &model.output_norm, &model.output, &model.output_b] {
                if let Some(t) = t { v.push(t.name.clone()); }
            }
            for l in &model.layers {
                for t in [&l.attn_norm, &l.wq, /* ... all 12 per layer ... */ &l.ffn_down] {
                    if let Some(t) = t { v.push(t.name.clone()); }
                }
            }
            v
        };
        let Some(mps) = crate::metal::MpsState::get() else { return false; };
        names.iter().all(|n| mps.has_weight(n))
    }
}

And where the verdict lands (src/models/qwen2/graph.rs:431-457,462-470):

#![allow(unused)]
fn main() {
        #[cfg(target_os = "macos")]
        let metal_on =
            crate::graph::metal_backend::metal_available() && Self::weights_on_gpu(model);
        ...
        #[cfg(feature = "cuda")]
        let cuda_on = crate::cuda::CudaState::get().is_some() && Self::weights_on_cuda(model);
        ...
        let params = GraphParams {
            n_tokens: nt, n_out,
            gtype: if nt == 1 { GraphType::Decode } else { GraphType::Prefill },
            cparams: CParams {
                n_ctx, flash_attn: false,
                gpu: metal_on || cuda_on,          // ← participation recorded
                fuse_qkv: nt == 1 && (metal_on || cuda_on) && !env("MINFER_NO_FUSE_QKV"),
                fuse_ffn: nt == 1 && (metal_on || cuda_on) && !env("MINFER_NO_FUSE_FFN"),
            },
            weights_version: 1,
        };
        if !cache.try_reuse(&params) { /* build → register → assign → fuse → alloc */ }
}

Annotations: gpu is a single flag in the reuse identity, so toggling the environment forces a rebuild (doc 13). The CUDA gate (weights_on_cuda, :689-750) is stricter than Metal's: each matmul weight must be registered and match a kernel type (has_weight_of_size compares the byte length so padded Q6_K registrations still match by raw size), and the embedding is checked separately because its gather kernel supports one fewer type (Q4_1). On failure it names the first offending tensor instead of returning a bare false — the difference between a debugging session and a support ticket.

3.3 Design choices (why this shape and not another)

Dispatch on a string, once, before any side effect. Covered in §2.2; the one-line summary: deterministic, one honest error message, and compatible with the fact that loading mutates GPU state.

A wide trait instead of a narrow one. The tempting alternative is a minimal trait (forward + a getter or two) with the rest downcast via as_any. That pushes every consumer into arch-specific code. The chosen shape inverts it: the trait declares everything the pipeline needs (dims for graph shapes, build_graph for doc 05, special_tokens for doc 12, rope_style for doc 11, n_vocab for the sampler width), each architecture implements them once, and downstream code stays architecture-blind. The cost is some #[allow(dead_code)] ceremony on methods only reached through Box<dyn ModelDef> — noted in the trait's own comment (models/mod.rs:17-21) — which is a fair price for the type safety.

The IR references weights by name, not by pointer. Graph nodes carry NodeMeta::MatMul { weight_name, weight_ttype, in_dim, out_dim } (graph/builder.rs:127-145); backends resolve the name through their registry at execute time. The alternatives: embedding raw byte pointers in the IR (couples the pure graph to mmap lifetimes and makes the CUDA representation impossible), or embedding Tensors (makes graph comparison — the reuse identity — expensive). Names are cheap, comparable, and each backend maps them to its own representation. The same choice is what makes fusion possible: the fused attn_qkv weight is just another name, so a fused node differs from three unfused ones only in metadata.

Zero-copy on Metal, upload-once on CUDA, free on CPU. Three honest answers to "where can this hardware read bytes from?" — not three implementations of one idea. What they share is the invariant: after load, no backend ever moves weight bytes again during inference.

No dequantization at load. §2.3's arithmetic: 0.56 B/value vs 4 B/value, and both CPU and GPU kernels are built to consume the packed forms directly. The exceptions prove the rule — every repack that does happen (CUDA's padded Q6_K slots, the Q4_K descriptor plane, the optional f16 dequant cache, the fused concat weights) exists because a specific kernel measured faster on a different layout, is gated on model size or env flag, and is documented with its byte math at the registration site.

All-or-nothing GPU participation, recorded in the reuse identity. §2.5. The alternative (per-layer fallback) is the old engine's design, and its cost — 144 DMA operations per decode step in the worst case — is on record in docs/CUDA_OPTIMIZATION.md.

The legacy KVCache is gone (#252). #244 deleted the type's storage (it was never read), and #252 then deleted the empty marker src/cache.rs, its mod cache; declaration, the ModelDef::forward/forward_graph &mut KVCache parameter, the KVCache::new call in main.rs and the tests that constructed one only to satisfy the signature. The graph path's real KV lives in the allocator's persistent per-layer regions (docs 07–08) and always did; the "vestigial 100 MB allocation" this paragraph used to call the cheaper mess stopped existing with #244's storage deletion, so keeping the parameter bought nothing.

3.4 Pitfalls & invariants

  • Registration order on Metal: register_part for every mmap part before any register_weight. The zero-copy lookup finds weights by pointer containment in a registered part; weights registered first would silently take the copy path. Page alignment of the mmap base is a debug_assert, not a hope (src/metal/runtime.rs).
  • CUDA init must complete before the first registration. The loaders call CudaState::init() up front and hold a model-load guard for the whole load (qwen2/loader.rs:295-306). The recorded failure mode: lazy init mid-load flips the backend gate between tensors, producing a graph that mixes CPU/CUDA assignment against persistent KV regions that were sized for one of them.
  • Weights are immutable after load. Both CPU and CUDA registries skip or dedup same-name re-registration; weights_version in GraphParams is the escape hatch if that ever changes. The bug behind the CPU guard cost a measured 635 ms per prefill→decode switch.
  • CUDA device buffers are never freed on replace. A captured graph may reference them; the leak is bounded by distinct (architecture, tensor) shapes ever loaded (src/cuda/methods/weights.rs:24-37).
  • The gate and the registration must enumerate the same weights. Loader registers; weights_on_gpu/weights_on_cuda check the same field list spelled out twice. That duplication is intentional — a new weight field fails the compile in both places until acknowledged.
  • Loader registers ⊋ gate accepts (on CUDA). Some types are registered for the legacy path but have no graph kernel; the gate's type check is what keeps those on CPU. The embedding's Q4_1 exclusion is the standing example (qwen2/graph.rs:685-693).
  • The KV element-type decision is write-once. set_kv_cache_type initializes a OnceLock; it must run before the first forward, which is why the loaders do it mid-load — and why the Qwen3 loader resolves n_kv_embd before calling it.
  • positions[i] < n_ctx is a caller obligation. The KV regions are sized n_kv_embd × n_ctx once; forward_cached asserts it loudly (qwen2/graph.rs:421-426) rather than corrupting a region.

4. Observe & verify

  • ./target/release/minfer info <model> — dumps the tensor table (names, quant types, shapes) and metadata KV, so you can see exactly the names the loader will look up and the general.architecture value dispatch matches.
  • Startup log, the stage's own narration: File: … bytes in N part(s) (doc 02), then MPS: GPU acceleration enabled or MPS: disabled by MINFER_DISABLE_MPS / CUDA: GPU acceleration enabled, then Loaded: N layers, Model loaded., and Vocabulary: N tokens (the n_vocab this stage counted).
  • MINFER_DISABLE_MPS=1 — forces CPU on macOS; the log flips to MPS: disabled by MINFER_DISABLE_MPS and the graph's backend colors (next bullet) go all-CPU. MINFER_DISABLE_CUDA=1 is its CUDA twin.
  • MINFER_WEIGHT_COPY=1 / MINFER_CACHE_TYPE=f16|f32|q8_0 (q8_0 is read by CPU, CUDA and Metal — the last since #310, C4) — the first makes Metal copy each weight into a fresh buffer instead of wrapping the mmap pages (A/B the zero-copy path); the second pins the KV element type instead of the size-based auto-select.
  • --dump-graph out.dot / --dump-graph-json (or MINFER_TRACE=/tmp/t.json) — exports the built graph with real backend assignment; nodes whose matmuls reference registered weights show their assigned backend, which is the visible outcome of this stage's gate. minfer viz renders the same in a browser.
  • MINFER_NO_FUSE_QKV=1 / MINFER_NO_FUSE_FFN=1 — skips building the fused concat weights at load too, so their memory cost disappears from your process footprint; a way to feel the difference between "raw GGUF bytes" and "registration-time derived copies".
  • Tests: cargo test covers the weight registry round-trip (cpu_backend tests register and look up by name), and the Metal/CUDA graph suites run the same tiny graphs on GPU and CPU asserting bit-identical output — the end-to-end proof that registration made weights reachable on each backend.

5. Cross-references

  • docs/ARCHITECTURE.md §2 (module map), §3 (pipeline position of this stage), §5 (backend layering + selection rules), §8 (the add-a-new-architecture checklist that mirrors this doc).
  • docs/COMPUTE-GRAPH-DESIGN.md — Phase 5/6 record: how the imperative forward became build_graph + registries, and why nodes carry names.
  • docs/CUDA-BACKEND-DESIGN.md — Phase 7 design and §2 inventory of the CUDA weight registry; the resident-weights thesis this doc leans on.
  • docs/CUDA_OPTIMIZATION.md (+ docs/cuda_optimization_steps/) — the measured cost of per-step host round-trips that resident weights delete; Q6_K padded registration details (7e②).
  • docs/METAL_OPTIMIZATIONS.md — the mmap-part zero-copy design and the #39 first-touch warm-up; KV f16 auto-select measurements (§0/§2.5).
  • docs/QWEN3-SUPPORT-PLAN.md §2 — the decoupled head dim and per-head Q/K norm rationale.
  • docs/PERF-QWEN3-4B-VS-LLAMACPP.md — why n_ctx (not the model's max context) sizes the KV regions.
  • Neighbors: 02 (what the GgufModel container is), 04 (tokenizer + template — the next consumers of metadata), 05 (the graph that finally reads these weights), 07 (the allocator that owns registration and KV regions), 14/15 (the Metal and CUDA backends whose registries were filled here).

← 02 — GGUF load · Index · 04 — Tokenizer and chat template →

04 · Tokenizer + chat template

Stage: model dispatch + weights (03) → tokenizer + template → graph build (05). The weights are registered; nothing has run yet. This stage turns the string you typed into the integer token list that every later document operates on — the raw material of the whole rest of the series. Code: src/tokenizer.rs (Tokenizer::load :353, encode :619, decode_bytes :661), src/template.rs (render_template :603, render_messages :494), src/main.rs (template block :1404-1439, get_chat_template :2449, decode loop :1749-1768) — lines verified at commit 15fa45c.

1. Background — where this stage sits

Doc 03 left the engine with the GGUF memory-mapped and parsed, the architecture dispatched (Qwen2 or Qwen3), and every weight tensor registered in the compute-graph allocator — possibly on the GPU. But not one byte of your prompt has been touched: you typed "What is 2+2?", and the model, so far, has no idea it exists.

This stage closes that gap, and it has two halves. First, the chat template: a string, stored in the GGUF metadata, that describes how a conversation is spelled out in the exact marker format the model was trained on — your raw prompt is wrapped in that format before anything else happens. Second, the tokenizer: the code that turns that wrapped text into a list of integers.

Those integers are called token ids. A token is the model's unit of text — a short chunk such as "What", " is", or a single character — and each distinct chunk the model knows has a number. The model cannot read characters at all. Its very first layer is a lookup table (the embedding matrix) that maps the integer 3838 to a vector of floats; integer 3837 maps to a different vector. Feed it raw characters and there is simply no table entry — nothing downstream can run. That is why this stage gates the entire series: the token list is the input to graph build (05), the prefill forward (09), the sampler (12), and the decode loop (13).

The template half matters just as much, and it is the more surprising one. A chat model was not trained to continue arbitrary text; it was trained to answer when it sees a very specific arrangement of marker strings like <|im_start|>user. Get that arrangement wrong and a perfectly good model produces garbage — it will happily continue your sentence instead of answering it. The markers are not decoration; they are the protocol.

2. Principle — how it works and why

2.1 The stage in one picture

 prompt: "What is 2+2?"
     │
     ▼
 get_chat_template()          reads GGUF metadata key "tokenizer.chat_template"
     │                        (missing, or --no-template → use the raw prompt)
     ▼
 template::render_template()  minijinja (+ a Python-`str`-method hook) renders
     │                        with add_generation_prompt=true; a template it cannot
     │                        render is a LOUD error naming the construct (F7/#50)
     ▼
 "<|im_start|>system\nYou are a helpful assistant.<|im_end|>\n
  <|im_start|>user\nWhat is 2+2?<|im_end|>\n
  <|im_start|>assistant\n"
     │
     ▼
 Tokenizer::encode()
     ├─ 1. special-token scan   whole strings like <|im_end|> become one id
     ├─ 2. pre-tokenize          the `tokenizer.ggml.pre` rule (qwen2 / qwen35)
     │                          splits into word / number / punctuation pieces
     ├─ 3. byte-encode          map every raw byte to a printable unicode char
     └─ 4. greedy BPE merges    merge the adjacent pair with the lowest rank
     ▼
 Vec<u32>  [151644, 3838, 374, 220, 17, 10, 17, 30, 151645, 151648, 198]
     │
     └──► doc 05: graph build consumes these ids (positions 0,1,2,…)

2.2 Why does the model need a template at all?

Start with what a language model fundamentally does: given a sequence of tokens, predict a probability for every token in the vocabulary of what comes next. A base model (trained only on raw documents) uses this to continue text: prompt it with "The capital of France is" and it predicts " Paris" — or equally " known", because continuing documents is its whole job.

A chat model is a base model that went through a second training phase (usually called instruction tuning or alignment). The training data in that phase was conversations, serialized in a fixed format:

<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
What is 2+2?<|im_end|>
<|im_start|>assistant

Those <|im_start|> / <|im_end|> strings are special tokens — vocabulary entries that were reserved during training and shown to the model millions of times as turn boundaries. The model's chat behavior lives entirely in that format: during alignment training, every example was markers, a user turn, <|im_start|>assistant, then an answer — so the model learned the conditional distribution "text that follows <|im_start|>assistant": answers, not continuations.

Now the punchline about the last line. The rendered prompt ends with <|im_start|>assistant\n — an empty assistant turn opener. This is what the flag add_generation_prompt controls, and it is not optional. With the opener, the model's next-token distribution is "the first token of an assistant answer", and it says "2+2 equals 4...". Without it, the prompt ends inside the user turn, and the model keeps writing the user turn — more question text, or a stray <|im_end|>. It will not answer.

So the template is not a display nicety; it selects which distribution the model samples from. minfer always passes add_generation_prompt=true on the CLI path (src/main.rs), because a one-shot prompt is by definition a "generate the assistant's next turn" request.

One more piece: bos and eos. bos (beginning of sequence) is a token some model families expect at the very start of every input; eos (end of sequence) is the token the model was trained to emit when it is done talking. The template context exposes bos_token as a variable (src/template.rs); whether a bos marker appears is the template string's choice, not the engine's. For eos, minfer does not rely on the template — the GGUF metadata carries tokenizer.ggml.eos_token_id and <|im_end|>'s id directly (§2.5), and the decode loop treats them as stop signals.

2.3 The tokenizer: byte-level BPE, end to end

BPE (Byte Pair Encoding) is the algorithm that decided which chunks of text become tokens. It was run once, months before you ever run inference, on a huge training corpus:

  1. Start with every single byte as its own token (256 of them).
  2. Count which pair of adjacent tokens occurs most often in the corpus; merge that pair into a new token; repeat. Each merge gets a merge rank — the order number in which it was learned. Rank 0 was learned first, i.e. it was the most frequent pair in the whole corpus.
  3. Stop when the vocabulary reaches its target size — for Qwen models, 151,936 entries (docs/QWEN3-SUPPORT-PLAN.md:37).

The training-time result ships inside the model file: the GGUF metadata carries the token strings (tokenizer.ggml.tokens), the merge list in rank order (tokenizer.ggml.merges), and the special-token types. Encoding is simply replaying those merges on your text.

A first example — byte-exact, because it comes straight from minfer's test suite, whose expected ids were cross-checked against llama.cpp (src/tokenizer.rs):

text:    <|User|>What is 2+2?<|Assistant|><think>\n
ids:     151644 3838 374 220 17 10 17 30 151645 151648 198
idstored token textwhat a human seeshow it was produced
151644<|User|>(role marker)special token, matched whole, before BPE
3838WhatWhatpre-token piece whose characters merge into the stored entry
374Ġis␣isregex piece " is"; already a single vocab entry
220Ġ␣regex \s+ piece — the lone space before a digit
1722regex \p{N} piece — one digit only
10++regex punctuation piece
151645<|Assistant|>(role marker)special token
151648<think>(reasoning marker)special token
198Ċnewlineregex \s*[\r\n]+ piece

Two oddities in that table are the pre-tokenizer at work. The rule named by tokenizer.ggml.pre (qwen2 for Qwen2.5 and Qwen3; F7/#50) splits text into word pieces, single digits, and punctuation runs before any merging happens — merges can never cross a piece boundary. Digits are matched one at a time (\p{N} matches exactly one in the qwen2 rule), which is why 2+2 costs four tokens and why models are famously weak at long arithmetic: every digit is a separate concept. And a space before a digit attaches to nothing (the word rule only glues a leading space to letters), so it becomes a bare Ġ — token 220.

Now the greedy merge loop itself. minfer splits the piece into characters and repeatedly merges the adjacent pair with the lowest rank. A toy illustration (invented ranks): if the piece is [m][i][n][f][e][r] and ("i","n") has rank 88 while every other adjacent pair ranks higher, [i][n] fuses first; the scan then repeats on the shorter list until no adjacent pair is in the merge table, and each surviving piece is looked up in the vocab. The real ranks live in the GGUF; the real loop is excerpted in §3.2.3.

Why lowest rank first, and why does that give good tokenizations? Because rank order is frequency order from training. The first merges ever learned were the most common byte pairs; later merges built on earlier ones. Replaying lowest-rank-first reconstructs the same segmentation the vocabulary was built for, so your text is cut exactly the way the model saw text cut during training. The practical payoff is compression: common words were merged thousands of merges ago and exist as single tokens, so "the" costs one position instead of three. That matters downstream because every cost in this engine scales with token count — prefill matmuls, KV cache size (each token reserves a K and V row per layer), and decode latency per generated token.

Why byte-level? Because the alphabet is bytes, not characters. Before merging, every raw byte 0–255 is mapped to a printable unicode character (build_byte_to_unicode, src/tokenizer.rs; printable ASCII and most Latin-1 map to themselves, the rest get chars from code point 256 upward — space becomes Ġ, newline Ċ), and every vocab entry is stored in that mapped form. The consequence: any byte string round-trips — Chinese, emoji, binary junk — and decode (§2.4) can always invert the mapping exactly. A character-level tokenizer cannot make that promise: a character the vocabulary never saw has no representation at all.

2.4 Decode: ids back to bytes, and why bytes and not a String

Generation runs the same table backwards. decode_bytes concatenates id_to_token[id] for each id — producing the mapped-form text, e.g. Ġis — then maps every character back to its raw byte via the reverse table (unicode_to_byte). Out come raw bytes, exactly the bytes that were encoded.

Why insist on bytes rather than a Rust String? Because a multi-byte UTF-8 character can be split across two tokens. Consider the Chinese character 中 (U+4E2D), whose UTF-8 encoding is the three bytes E4 B8 AD; minfer's test (src/tokenizer.rs) contains a token whose mapped text is 中 — the mapped forms of exactly those three bytes. If the model emits the first two bytes of the character in one token and the third in the next, a per-token String::from_utf8_lossy conversion would stamp a � (U+FFFD replacement character) into your output stream permanently — the bytes were already thrown away. decode_bytes never attempts the conversion: it emits raw bytes, so E4 B8 + AD reassembles perfectly wherever they land. The tests pin this: decode_bytes_keeps_multibyte_bytes (src/tokenizer.rs).

2.5 Special tokens: ids with a job

A special token is a vocabulary entry that is not a piece of human text but a control signal: <|im_start|>, <|im_end|>, <think>, <|User|>, and so on. In the GGUF they are flagged by tokenizer.ggml.token_type — the values the code checks are 3 (control) and 4 (user-defined) (src/tokenizer.rs).

They get special treatment at both ends of the pipeline:

  • Encode: a special token must survive as one id — the BPE machinery would otherwise shred <|im_end|> into ordinary character pieces. minfer scans for special-token strings before running BPE on each segment (src/tokenizer.rs), matching the earliest position first and the longest string at a given position. This is not cosmetic: DeepSeek-R1-style markers <|User|> use fullwidth unicode bars that the pre-tokenizer rule would happily split apart; the regression test at :533 keeps them intact.
  • Decode/generate: the ids of eos and <|im_end|> are handed to the generation loop as stop sentinels — when the sampler produces one, the engine stops instead of appending it. They are also fed into the sampler's penalty window (src/main.rs, doc 12). The ids come from ModelDef::special_tokens() (doc 03), sourced from GGUF metadata: tokenizer.ggml.eos_token_id, plus a lookup of <|im_end|> that falls back to the eos id (src/models/qwen2/loader.rs:130-131). One vocabulary, two directions, and a set of reserved ids that act as the protocol's punctuation.

3. Implementation

3.1 Data in / data out

Input data — GGUF metadata (parsed in doc 02; the tokenizer reads it via GgufContext, src/tokenizer.rs):

GGUF keyTypeLands in
tokenizer.ggml.tokensstring array (~151,936 entries for Qwen)id_to_token: Vec<String>, inverted into vocab: HashMap<String,u32>
tokenizer.ggml.scoresf32 arrayid_to_score (loaded for llama.cpp parity, unused)
tokenizer.ggml.token_typei32 array (1 normal, 3 control, 4 user-defined)id_to_type, drives the special-token table
tokenizer.ggml.mergesstring array "A B" per merge, in rank ordermerges: HashMap<(String,String), usize> — pair → rank
tokenizer.ggml.prestring (qwen2, qwen35, …)pre: PreTokenizer — which rule split() applies (F7/#50). An unknown or missing value refuses the whole load
tokenizer.ggml.bos_token_id / eos_token_idu32bos_token / eos_token
tokenizer.chat_templateone long stringpassed to minijinja verbatim

Note what is not here: no tokenizer model file, no external vocabulary. The vocab ships inside the GGUF because the model file already had to describe its own output layer (output.weight is [n_embd, 151936] — the vocabulary size is baked into the weight shape), so the conversion tool writes the matching token table alongside it.

The flow is: &str prompt + metadata → rendered String → Vec<u32> → ctx = max(--n-ctx, ids.len()) (src/main.rs), which sizes the persistent KV regions once for the whole run → forward(&ids, positions 0..n) (doc 05+). During generation the direction reverses: one sampled id per step → decode_bytes → raw bytes → stdout/SSE. The template's token cost is real memory: the rendered wrapper becomes part of the prompt, and the prompt length feeds ctx — a template that bloats the prompt bloats the KV allocation.

3.2 Key code

3.2.1 Template selection and rendering (CLI path)

The whole template stage in main.rs is deliberately small — read the template out of metadata, render, encode:

#![allow(unused)]
fn main() {
// src/main.rs — the whole CLI template stage (F7/#50)
// === Chat template (need tokenizer for bos_token text) ===
let processed = if no_template {
    prompt.clone()
} else if let Some(tmpl) = get_chat_template(&gguf_model.parts[0].data) {
    let bos_text = tokenizer
        .id_to_token
        .get(tokenizer.bos_token as usize)
        .map(|s| s.as_str())
        .unwrap_or("");
    // An unrenderable template refuses the run here, before inference —
    // `validate` renders a canary conversation through it and a failure
    // names the construct and the template line.
    if let Err(e) = template::validate(&tmpl) {
        eprintln!("Error: {}", e.message());
        std::process::exit(1);
    }
    match template::render_template(&tmpl, &prompt, true, bos_text) {
        Ok(p) => p,
        Err(e) => {
            eprintln!("Error: {}", e.message());
            std::process::exit(1);
        }
    }
} else {
    eprintln!(
        "Notice: this GGUF has no tokenizer.chat_template; using the generic ChatML renderer"
    );
    prompt.clone()
};
let input_ids = tokenizer.encode(&processed);
}

Three branches, in priority order: --no-template bypasses everything and tokenizes the raw prompt (useful for base models and for comparing token counts); otherwise get_chat_template pulls tokenizer.chat_template from the GGUF metadata bytes (a tiny re-parse of metadata only — src/main.rs); with no template key at all, the raw prompt is used as-is. The literal true argument to render_template is add_generation_prompt — §2.2 explained why it must always be on for a one-shot prompt. If the tokenizer produces zero ids, the run aborts: an empty token list would leave the graph builder with no tokens to embed.

The renderer wraps minijinja, a small Jinja-compatible template engine, plus a Python str-method hook (F7/#50). That hook is what lets the published Qwen2.5/Qwen3 templates run at all: they use Python string syntax (message.content.split('</think>').lstrip('\n')), and minijinja strings expose no methods, so the engine registers Environment::set_unknown_method_callback and implements the methods with CPython semantics (split, lstrip/rstrip/strip with a character set, replace, startswith/endswith, join, find, …):

#![allow(unused)]
fn main() {
// src/template.rs — the environment every render goes through
fn environment() -> Environment<'static> {
    let mut env = Environment::new();
    env.set_unknown_method_callback(unknown_method);
    env.add_function("raise_exception", |msg: String| -> Result<Value, Error> {
        Err(Error::new(ErrorKind::InvalidOperation, msg))
    });
    env
}
}

The context exposes exactly what real chat templates expect: messages, add_generation_prompt, bos_token, and tools as none (transformers' default), so {% if tools %} branches take the no-tools path.

A template that cannot be compiled (add_template fails) or cannot be rendered (render_messages returns Err(TemplateError)) is a refusal, not a fallback. The refusal names the construct and the template line:

chat template error — unsupported template construct: unsupported Python str
method `splitlines` (template line 41); minfer refuses to fall back to a generic
ChatML prompt. Supported Python str methods: capitalize, count, endswith, find,
join, lower, lstrip, replace, rfind, rsplit, rstrip, split, startswith, strip,
title, upper

The CLI prints that and exits; serve/viz run template::validate once at startup and refuse to start; a per-request failure is an HTTP 400; a conversation turn reports it. The hand-written ChatML renderer (fallback_chatml_messages, src/template.rs) survives for exactly one case — a GGUF with no tokenizer.chat_template key at all, where there is no model format to lose (that is the branch that prints the Notice: line above).

ChatML is the marker convention Qwen models are trained on and the de-facto lingua franca of chat templates, which is why it is a reasonable default for a GGUF that carries no template. It is deliberately not used as a fallback when a template exists: the F7 design record (docs/CHAT-TEMPLATE-AND-TOKENIZER-DESIGN.md) shows what that silently cost — Qwen3's think-block extraction and tool-call formatting, fed back verbatim.

3.2.2 Loading the tokenizer from GGUF metadata

Tokenizer::load (fallible since F7/#50: it returns Result<Self, String> and refuses a tokenizer it cannot reproduce byte for byte) walks the metadata key-value list once per data kind. The token strings become the id_to_token vector and an inverted vocab map (:114-118); special-token ids and types fill special_tokens (:135-143) plus bos_token / eos_token / im_end (:145-149). BPE's data structure is built at :120-133: for each tokenizer.ggml.merges entry — a string like "Ġ t", two space-separated halves — the code splits on the first space and inserts merges[(first, second)] = i, where i is the array index. That index is the merge rank, because converters write merges in the order they were learned. Splitting on the first space is enough because each half is one byte-encoded string with no literal spaces in it (spaces were mapped to Ġ precisely so they could never appear inside a half).

Special tokens need one more data structure, and its comment explains the invariant:

#![allow(unused)]
fn main() {
// src/tokenizer.rs, 172-179 (the <|im_start|>/<|im_end|>/eos
// fallback inserts between, described in the text below)
// Merge GGUF special tokens (type 3/4) with hardcoded fallbacks, then
// group by first char with longest-first ordering inside each group
// (an earliest-position, longest-match scan needs both).
let mut special_by_first: HashMap<char, Vec<(String, u32)>> = HashMap::new();
for (pat, id) in merged {
    let first = pat.chars().next().unwrap_or('\0');
    special_by_first.entry(first).or_default().push((pat, id));
}
for group in special_by_first.values_mut() {
    group.sort_by(|a, b| b.0.len().cmp(&a.0.len()));
}
}

The skipped middle starts from special_tokens.clone() and defensively inserts <|im_start|>, <|im_end|>, and the eos token only when the GGUF did not already provide them (contains_key guards): some converted models mark their specials as ordinary type-1 tokens, so minfer hardcodes the ChatML markers as fallbacks, and real metadata always wins. The resulting special_by_first index buckets patterns by their first character; the encode scan (next) will jump straight to the bucket for the character it is looking at instead of testing every pattern against every position.

3.2.3 Encode: specials first, then regex, then merges

The top-level encode is a loop over "segments": text up to the next special token goes through BPE, the special token becomes a single id, repeat (src/tokenizer.rs, doc comment at :273-279):

#![allow(unused)]
fn main() {
// src/tokenizer.rs (fn head at :280-283, final `result` at :309-310)
loop {
    // Find the earliest position where any special token starts.
    let mut earliest: Option<(usize, u32, usize)> = None; // (byte_pos, id, byte_len)
    'scan: for (ci, ch) in remaining.char_indices() {
        if let Some(group) = self.special_by_first.get(&ch) {
            let rest = &remaining[ci..];
            for (pat, id) in group {
                if rest.starts_with(pat.as_str()) {
                    earliest = Some((ci, *id, pat.len()));
                    break 'scan; // group is longest-first; earliest char wins
                }
            }
        }
    }

    if let Some((pos, id, len)) = earliest {
        // Encode text before the special token
        if pos > 0 {
            result.extend(self.encode_bpe(&remaining[..pos]));
        }
        result.push(id);
        remaining = &remaining[pos + len..];
    } else {
        // No more special tokens, encode the rest
        result.extend(self.encode_bpe(remaining));
        break;
    }
}
}

The double ordering matters: the outer scan takes the first character that starts any special token ("earliest position wins"); within one position, the bucket is sorted longest-first, so the first starts_with hit is the longest match ("<think▁begin|> beats <think>"). The dedicated tests special_token_earliest_position_wins and longest_special_token_wins_at_same_position pin both rules.

Inside a segment, encode_bpe (src/tokenizer.rs) applies the pre-tokenization rule named by tokenizer.ggml.pre (F7/#50) over the text — for qwen2, which Qwen2.5 and Qwen3 both select:

(?:'[sS]|'[tT]|'[rR][eE]|'[vV][eE]|'[mM]|'[lL][lL]|'[dD])|[^\r\n\p{L}\p{N}]?\p{L}+|\p{N}| ?[^\s\p{L}\p{N}]+[\r\n]*|\s*[\r\n]+|\s+(?!\S)|\s+

The alternation, read left to right: contractions ('s, 't, 're…), an optional leading space glued to a run of letters, a single digit, an optional leading space glued to punctuation, line breaks, and any other whitespace. It is implemented as a hand-written scan rather than a regex pattern because the Rust regex crate has no lookahead and the (?!\S) on the whitespace alternative is load-bearing for duplicated spaces; PreTokenizer::split returns byte slices that tile the input exactly, pinned byte for byte against CPython regex on the model's own tokenizer.json pattern by the committed split fixtures. Each piece becomes one piece here: byte_encode (the per-byte character mapping from §2.3) maps its bytes to printable chars, and the piece goes to bpe_encode. These piece boundaries are sacred — merges never cross them — which is why " is" and "2" never fuse into one token no matter what the merge table says.

Then the merge loop itself:

#![allow(unused)]
fn main() {
// src/tokenizer.rs (whole-piece shortcut at :218-221, lookup tail at :251-254)
loop {
    // Find the best merge (lowest rank)
    let mut best_rank: Option<usize> = None;
    let mut best_idx: Option<usize> = None;

    for i in 0..word.len().saturating_sub(1) {
        let pair = (word[i].clone(), word[i + 1].clone());
        if let Some(&rank) = self.merges.get(&pair) {
            if best_rank.is_none() || rank < best_rank.unwrap() {
                best_rank = Some(rank);
                best_idx = Some(i);
            }
        }
    }

    if best_idx.is_none() {
        break;
    }

    // Merge at best_idx
    let idx = best_idx.unwrap();
    let merged = format!("{}{}", word[idx], word[idx + 1]);
    word.splice(idx..=idx + 1, std::iter::once(merged));
}
}

Read it as: loop {scan every adjacent pair, keep the lowest-rank one, splice it}, until no pair is in the merge table. There is deliberately no "whole piece is already a vocab entry" shortcut (F7/#50 removed one): a vocabulary entry is not necessarily reachable through merges — Qwen3.5 holds a Devanagari cluster as one entry with no rank producing it, and every reference splits it — so the loop always starts from single characters, exactly like HF tokenizers and llama.cpp. After it, each surviving piece is looked up in the vocab, and a piece that is not there falls back one token per byte (byte_fallback), which Tokenizer::load guarantees is possible by refusing a vocabulary that lacks any of the 256 byte tokens. The old tail mapped such a piece to id 0 (unwrap_or(0)) — see §3.4 for why that was a bug. Complexity is O(pieces²) per word with tiny constants; tokenization runs once per prompt, so it is not a hot path.

3.2.4 Decode: ids → bytes → streamed text

#![allow(unused)]
fn main() {
// src/tokenizer.rs
pub fn decode_bytes(&self, ids: &[u32]) -> Vec<u8> {
    let mut encoded = String::new();
    for &id in ids {
        if (id as usize) < self.id_to_token.len() {
            let token = &self.id_to_token[id as usize];
            encoded.push_str(token);
        }
    }

    // Reverse byte-level encoding
    let mut result = Vec::new();
    for c in encoded.chars() {
        if let Some(&b) = self.unicode_to_byte.get(&c) {
            result.push(b);
        } else {
            // Fallback: encode the char as UTF-8
            let mut buf = [0u8; 4];
            let s = c.encode_utf8(&mut buf);
            result.extend_from_slice(s.as_bytes());
        }
    }

    result
}
}

Two quiet robustness choices: out-of-range ids are skipped — a corrupted sample cannot panic the stream (tested at :466-471) — and characters not in the reverse byte map (vocab entries holding genuine unicode text rather than byte-mapped forms) pass through as their UTF-8 bytes (:336-341).

The streaming holdback, used by the server and conversation paths, is complete_utf8_prefix_len (src/tokenizer.rs). Its doc comment says it "mirrors llama.cpp's format_incomplete_utf8 holdback", and the mechanism is just the UTF-8 length grammar walked once: the lead byte of a sequence determines its length (1 for ASCII, 2–4 for multi-byte, judged by the 0xC0/0xE0/0xF0 masks on the top bits), so the function scans forward until the next character would run past the end of the buffer, and returns the offset where the complete prefix ends — everything from there onward waits for the next token. The callers wire it into per-step emission: server/chat.rs:162 appends newly decoded bytes to a full accumulator and server/chat.rs:176 flushes up to emitted + complete_utf8_prefix_len(&full[emitted..]); conversation.rs:548/564 does the same for the REPL. (The CLI decode loop instead writes raw bytes straight to stdout and lets the terminal assemble the glyph; complete_utf8_prefix_len_holds_incomplete_trailing at :474 pins the grammar.)

And the consumer end — how the ids the sampler produces meet this decoder in the CLI decode loop (doc 13 walks the loop; here, only the tokenizer-relevant lines). The sampled id first runs the stop-sentinel check, then is recorded for the penalty window (generated.push / prev_tokens.push, :903-907):

#![allow(unused)]
fn main() {
// src/main.rs, 909-918
if is_stop_token(sampled.token_id, &special) {
    break;
}

// Stop-string detection on the FULL byte stream before emitting.
full.extend_from_slice(&tokenizer.decode_bytes(&[sampled.token_id]));
if let Some(cut) = sampler::match_stop_suffix(&full, &stop_refs) {
    full.truncate(cut);
    if cut > emitted {
        hi.feed(&full[emitted..]);
        emitted = full.len();
    }
    break;
}
}

special came from model.special_tokens() (src/main.rs), and is_stop_token (src/main.rs) is the two-line sentinel check id == special.eos || Some(id) == special.im_end — the SpecialTokens struct itself is two fields (eos, im_end: Option<u32>, src/models/mod.rs). Note the byte-accumulation pattern: decoded bytes go into full before emission (the rest of the loop, :919-922, flushes newly completed bytes), so a stop string (--stop "Let me think") that straddles two tokens is caught in the accumulated stream — the same byte-first philosophy as §2.4. The tokens handed to prev_tokens also feed the sampler's repeat-penalty window (doc 12), so prompt and generated token ids influence penalties from the very first step (src/main.rs).

3.3 Design choices (why this shape and not another)

Why a chat template at all — why not just tokenize the prompt? Because tokenization is the wrong layer for chat behavior. The model was aligned on a marker protocol; the only way to reach its "answer mode" is to reproduce that protocol byte-for-byte at the input (§2.2). This gets chat behavior out of the same weights by changing only the prompt text, where separate per-mode models or hidden role channels would multiply the model. The cost: template rendering is a compatibility surface (§3.4's minijinja gotcha), and template correctness is invisible until the model answers wrong — hence the debug dump (§4) and the --no-template bypass.

Why self-contained BPE instead of a tokenizer crate? The obvious alternative is a dependency — tokenizers (HuggingFace) or tiktoken — bringing a large dependency tree and its own version skew. minfer's constraint is zero ML framework deps (ARCHITECTURE.md §1: five crates total, minijinja the newest). The decisive fact is that the tokenizer's data already ships in the GGUF — tokens, scores, types, merges, special-token flags — so a crate would mostly re-read the same tables and add a second source of truth; the algorithm itself is ~80 lines (§3.2.3). The honest trade-off is exactness risk: BPE implementations differ in pre-tokenization details, and a divergence silently changes every id. minfer buys that risk down with llama.cpp-parity tests — special_tokens_match_as_single_ids_before_bpe hardcodes ids copied from llama.cpp's tokenizer as the expected output (src/tokenizer.rs), so any divergence fails CI rather than shipping as subtly different model behavior.

Why greedy lowest-rank merging? Because the merge table is a frequency-ordered construction history, and lowest-rank-first replay is the inverse operation (§2.3): it reproduces the segmentation the model was trained on, with no search — one linear scan per merge round. The alternatives are worse on both axes: longest-first or highest-rank-first produce segmentations the model never saw (the pieces would still be valid tokens, but the embedding each maps to was trained on different contexts — quality quietly degrades), and optimal segmentation search (minimize token count, e.g. Viterbi) costs orders of magnitude more. The compression payoff is concrete: single-token " is" costs one position instead of three — one fewer row through every attention head and one fewer KV row per layer, multiplied by every layer.

Why add_generation_prompt=true (and hard-coded on the CLI path)? The rendered prompt must end with the empty assistant-turn opener (<|im_start|>assistant\n), or the model's next-token distribution is "continue whatever turn is open" — usually a continuation of the user's own text (§2.2). It is hard-coded true on the CLI path (src/main.rs) because a one-shot CLI prompt is definitionally a "start the assistant's turn" request. The multi-turn paths pass it explicitly too — and only on the final render: format_single (src/template.rs), the incremental renderer behind --cnv, renders the recorded past with false and only the new state with true, then diffs the two strings so the KV cache is appended with just the delta — a mid-history render must not append an opener, or the KV would contain an assistant header that never led to an answer.

Why byte-level (and why decode in bytes)? Byte-level BPE gives total input coverage (§2.3) and byte decode gives lossless streaming (§2.4). A string-oriented decoder would corrupt output precisely in the cases that matter — emoji, CJK — and permanently: from_utf8_lossy cannot be undone. The design keeps lossy conversion strictly at the presentation edge (the doc comment on decode, src/tokenizer.rs, says streaming paths must use decode_bytes), never in the data path.

Why match special tokens before BPE, with earliest-then-longest rules? Special tokens are protocol punctuation; letting BPE see them destroys their meaning (and with R1-style fullwidth markers, the regex pieces can never recombine into the special string — the test comment at src/tokenizer.rs says exactly this). Earliest-position-wins matches how a human reads: the leftmost marker is the next structural event. Longest-at-position-wins disambiguates prefixes (<think vs <think▁begin|>); any other priority would be arbitrary.

3.4 Pitfalls & invariants

The minijinja 2.21 gotcha — fixed in F7, and it is why the hook exists. minfer uses minijinja = "2" with default-features = false. In minijinja 2.21 template strings are Rust strings and expose no str methods, so Qwen3's shipped chat_template (Python method syntax: message.content.split('</think>'), .lstrip('\n')) used to fail at render time (unknown method: string has no method named split) and silently fall back to ChatML — losing think-block extraction, tool-call formatting and enable_thinking (docs/QWEN3-SUPPORT-PLAN.md §5 gotcha #9 keeps the historical record). F7 (#50) installs minijinja's unknown-method callback and implements those methods with CPython semantics, so the model's own template runs. The new failure mode is the opposite of the old one: if a template needs a construct the hook does not implement you get Error: chat template error — unsupported template construct: … and no inference — never a generic prompt that silently changes model behaviour.

Specials must never reach BPE. The whole-string scan happens before encode_bpe, and the regression test exists because the R1 template broke otherwise. Invariant: a new special-token source must join merged/special_by_first before encode runs.

Unknown pieces used to map to id 0 — now they cannot. bpe_encode's tail maps an unknown piece to one token per byte, and Tokenizer::load refuses a vocabulary that lacks any of the 256 byte tokens (F7/#50). Before that, a vocab/merges inconsistency silently yielded the vocabulary's first entry and one word of output was consistently garbage. The empty-encode guard in main.rs still catches the louder "nothing encoded at all" failure.

Byte-decode invariant: no lossy conversion in the streaming path. The lossy decode (src/tokenizer.rs) exists for tests only — its doc comment says so explicitly. Streaming paths must pair decode_bytes with complete_utf8_prefix_len, or multi-byte characters split across tokens become permanent U+FFFD in the transcript.

Template and conversation modes are coupled. --cnv refuses --no-template (src/main.rs) because the conversation session's append-only KV scheme requires template rendering to compute what the next turn appends.

Template output feeds KV sizing. ctx = max(n_ctx, prompt_len) (src/main.rs): the rendered prompt's token count participates in sizing the persistent KV regions (doc 07). A runaway template (e.g. one that duplicates history) does not just slow prefill — it changes the allocation.

4. Observe & verify

  • The two printed counts. Every CLI run prints Vocabulary: 151936 tokens (src/main.rs — the number for Qwen-family models) and then Prompt: {} tokens (printed right after tokenization). Run the same prompt with and without --no-template: the difference is exactly the boilerplate the template added (system turn, role markers, the assistant opener).
  • See the rendered prompt. Build with --features debug_dump and set MINFER_DUMP_DIR: crate::dump::maybe_dump_text("minfer_dump_prompt", …) (src/main.rs) writes the post-template, pre-tokenization string — the literal <|im_start|>…<|im_end|>…<|im_start|>assistant text of §2.2. Format reference: docs/debug-dump.md.
  • Unit tests are the fastest oracle — and the llama.cpp cross-check. cargo test tokenizer:: covers the byte round-trip (decode_bytes_reverses_byte_encoding, the CJK decode_bytes_keeps_multibyte_bytes), the holdback grammar (complete_utf8_prefix_len_holds_incomplete_trailing), and all three special-token rules — including the id list copied from llama.cpp, and the real-model token_ids_match_the_reference gate (5 cached models × 52 corpus entries) that fails on any id shift. cargo test template:: covers rendering byte-for-byte against the committed transformers references, the loud refusal, and the incremental format_single diff semantics (format_single_diffs_only_new_user_message renders the real Qwen ChatML template shape, src/template.rs).
  • Per-token text in traces and the loud refusal. With MINFER_TRACE=<path>, the decode loop attaches each sampled token's decoded text to the trace (crate::trace::set_token, src/main.rs) — handy for spotting byte-fallback pieces from §3.4. A template failure prints Error: chat template error — … on stderr and stops before inference: the model's own template is not renderable and the engine refused to substitute a different prompt. The only line that means ChatML replaced a template is Notice: this GGUF has no tokenizer.chat_template; using the generic ChatML renderer.

5. Cross-references

  • 02 — GGUF load: where tokenizer.ggml.* and tokenizer.chat_template come from — this stage is a metadata consumer.
  • 03 — Model dispatch and weights: provides ModelDef::special_tokens() (eos, im_end) and the weights this stage's ids will drive.
  • 05 — Graph build: the IR and the builder: consumes the token list; GraphParams and the KV sizing that ctx = max(n_ctx, prompt_len) feeds.
  • 09 — Prefill forward path: the ids become the embedding lookup rows with positions 0..len.
  • 12 — Sampler: stop sentinels and the penalty window that receives prompt + generated ids; stop-string byte matching from §3.2.4.
  • 13 — Decode loop + graph reuse: the loop whose per-token decode_bytes + holdback streaming this doc set up; also the multi-turn path where format_single renders only the appended turn.
  • docs/QWEN3-SUPPORT-PLAN.md §5 #9: the full minijinja 2.21 record — the exact template lines that failed and the fallback consequences, kept as the history of the bug F7 fixed.
  • docs/CHAT-TEMPLATE-AND-TOKENIZER-DESIGN.md: the F7 contract — accepted and refused template constructs, the loud refusal, the pre-tokenizer rules, and the reference behind every gate.
  • docs/CHAT-TEMPLATE-AND-TOKENIZER-DESIGN.md: the F7 contract — accepted and refused template constructs, the loud refusal, the pre-tokenizer rules, and the reference behind every gate.
  • docs/OPENAI-CHAT-API-PLAN.md and docs/CLI-CONVERSATION-PLAN.md: the server-side template handling (render_messages, tools) and the incremental-render design (format_single) behind multi-turn sessions.
  • docs/USAGE.md: every flag this stage reads (--no-template, --stop, --cnv). docs/GLOSSARY.md: backstop definitions (BPE, merge rank, ChatML, bos/eos).
  • ARCHITECTURE.md §3: the flowchart rows this doc expands (template render → tokenize → prefill).

← 03 — Model dispatch and weights · Index · 05 — Graph build: the IR and the builder →

05 · Graph build — the IR and the builder

Stage: tokenizer + template (04) → graph build (IR) → backend assignment + fusion (06). By this point the prompt is a list of integer token ids and every weight is registered by name in the allocator — but nothing has been computed yet. This stage writes the entire transformer forward pass as one pure data structure: a declarative compute graph that later stages assign to hardware, fuse, allocate, and execute. Code: src/graph/mod.rs (CNode :156, ComputeGraph :201, topo_order :238), src/graph/ops.rs (Op :47, NodeMeta :196), src/graph/builder.rs (GraphBuilder :16, method family :264–711), src/models/qwen2/graph.rs (Qwen2Graph::build :39, forward_cached :450), src/graph/params.rs (GraphParams :88), src/graph/cache.rs (try_reuse :69) — lines verified at commit 15fa45c.

1. Background — where this stage sits

After doc 04 the engine holds two things: a Vec<u32> of token ids — one integer per piece of your prompt, as produced by the byte-pair-encoding tokenizer — and a loaded model whose weight tensors (the quantized matrices from the GGUF file) are registered by name in the graph allocator. A forward pass must now turn those ids into logits: one floating-point score per vocabulary entry, from which the sampler will pick the next token. That takes hundreds of math operations — normalizations, matrix multiplies, rotations, attention — arranged in a very specific order, repeated for each of the model's transformer layers.

The engine does not run those operations as a hand-written loop, at least not any more. Instead it first describes the whole computation as a data structure called a compute graph — a directed acyclic graph where each vertex ("node") is one math operation and each edge means "this node's output is that node's input". The adjective declarative is the point: building the graph states what should be computed, never how or when. The "how" (which CPU instructions, which GPU kernel) is decided later, per node, by the backend scheduler (doc 06). The "when" is the node order itself (doc 08).

Why go through this indirection? Four concrete payoffs, each of which gets a full section in §3.3:

  1. Reuse. Decoding is a loop: one forward pass per generated token, often hundreds of them. If the graph is a pure function of a small parameter struct (GraphParams), identical parameters mean an identical graph — so the engine builds it once and replays it for every decode step, refreshing only the input data (the new token id, the new position).
  2. Global decisions. "Which backend runs this node?" and "which buffers can share memory?" need the complete picture of the computation before anything runs — and a data structure you can walk is exactly that picture.
  3. Optimization as rewriting. Pattern-optimizations (fusing silu∘mul into one operation) become a small graph-rewrite pass instead of scattered special cases inside a forward loop.
  4. Observability. The graph can be printed (--dump-graph exports Graphviz DOT; --dump-graph-json exports JSON for the web visualizer), so you can literally see one forward pass.

There was an older design: models/qwen2/forward.rs, an imperative per-layer loop that computed as it went and hard-coded its GPU fallbacks. It is gone — deleted in Phase 6 of the graph refactor (commit 6af12a4) only after the graph path reproduced its logits bit-identically, prefill and decode (docs/COMPUTE-GRAPH-DESIGN.md §17, Phases 5–6 and 8). Its shape survives as a historical record in docs/ARCHITECTURE.md Appendix A.

So this document is the heart of the engine: the transformer forward pass, written as a data structure. Doc 06 walks it and gives every node a backend; doc 07 gives every node a buffer; doc 08 finally executes it.

2. Principle — how it works and why

2.1 The vocabulary: graph, node, edge, IR

Some terms we will use constantly:

  • Compute graph — the data structure describing one forward pass. Each node is one operation ("multiply these two matrices"), each edge is a data dependency ("node 13's output feeds node 14"). No cycles are allowed: information flows one way, input to output.
  • IR — intermediate representation, a compiler word for "a program expressed as data, sitting between the source and the machine". minfer's IR is ComputeGraph: a Rust struct you can traverse, compare, rewrite, and export. llama.cpp has the same idea (ggml_cgraph); minfer mirrors it.
  • Node (CNode) — one operation plus everything the rest of the engine needs to know about it: which operation (an Op enum value with full parameters), which other nodes it reads (src, the edges), the shape and element type of its output, and (only later, doc 06) which backend it runs on.
  • Topological order — an ordering of nodes where every node appears after all of its inputs. The builder produces this by construction: it only lets you reference nodes that already exist. "Append sources before consumers" is one sentence, and it buys an enormous amount: no sort is ever needed, and executing nodes in list order (doc 08) is automatically correct.
  • Backend — one compute device and its kernels: CPU (AVX2/NEON SIMD), Metal (Apple GPU), Cuda (NVIDIA GPU). Each node gets assigned one, but that happens after this stage.
  • Buffer — a chunk of memory holding one node's output; a BufRef is just a handle (which backend's pool, which slot). Allocated in doc 07; the graph only records shapes.

2.2 The three data structures

The whole IR lives in src/graph/mod.rs (273 lines including tests). You have met the CNode fields in §2.1; what ComputeGraph itself adds is the node list plus two id sets — inputs (the leaves filled with external data each step) and outputs (the logits) — and a uid, a monotonic id assigned when the graph is cached, which the CUDA backend keys its replay cache on. A BufRef { backend, id } is just a handle for "buffer id in that backend's pool" — created only in doc 07.

2.3 Operations carry full payloads

The Op enum is not just a tag like "MatMul". Every variant carries the parameters that make it this operation and no other: RmsNorm { eps }, MatMul { transpose_b }, RoPE { style }, KvcacheStore { layer }, FusedQKV { layer }. Node-level extras (which weight tensor, which bias, frequency base) ride along in a NodeMeta enum. Two consequences:

  • Two graphs can be compared structurally — node by node, payload by payload. Op and NodeMeta both derive PartialEq for exactly this, and a debug/test check (GraphCache::verify_structural) asserts that rebuilding a graph with identical parameters yields an identical node sequence — the tripwire guarding the reuse invariant of §2.7.
  • The dumped graph is self-describing: the DOT export prints RmsNorm { eps: 1e-6 } and KvcacheStore { layer: 23 }, so you can read a forward pass off the page.

The full op list is wider than what Qwen2 emits (Softmax, View, Permute, Scale… exist for ggml parity), but the live graph uses a compact set — see the census in §2.5.

2.4 The builder: an append-only factory

GraphBuilder (496 lines) is the only way to construct a graph. Its methods come in two kinds:

  • input(name, shape, dtype) — declare a leaf node whose data arrives from outside each step. It records the id in graph.inputs and nothing else.
  • The operation family — embedding, rms_norm, qk_norm (Qwen3), matmul, get_rows, rope, silu, add, mul, swiglu, softmax, attn, kvcache_store/kvcache_load, plus the decode-fusion constructors fused_qkv / qkv_bias_rope_store / fused_ffn / fused_qkv_norm. Each one computes the output shape from its inputs' shapes (so shapes propagate through the graph automatically), wraps the parameters into an Op + NodeMeta, appends the node, and returns its id.

The crucial property is stated in the module's first lines: "The builder is pure: it only appends nodes to the graph, it never computes." Building a graph allocates a few Vecs and does integer shape arithmetic — no matmul runs, no GPU is touched, no buffer exists yet. Purity is what makes the reuse invariant (§2.7) even statable.

One naming convention to keep straight, because it recurs in every excerpt: GGUF stores weight metadata as [in, out] (input dim first) but memory layout is [out][in] row-major, and activations are token-major [nt][d] (nt = token count, d = features) — so matmul's output shape is [w.shape[1], nt]: output dim, then token count.

2.5 The main event: one forward pass, node by node

Qwen2Graph::build (src/models/qwen2/graph.rs:39) is where the forward pass is actually written down. The real thing below is from the Qwen2.5-0.5B model (24 layers, hidden 896, 14 query heads of dim 64, 2 KV heads — more on that below, FFN width 4864, vocabulary 151936), prefilled with a 30-token prompt. Node ids are the real ones from a --dump-graph export.

token_ids(0)  positions(1)  tail_ids(2)      <- input leaves (filled per step)
      \          |              \______ (consumed at the LAST layer only)
       v        v
    embed(3) GetRows             h = one token_embd row per id, [896, nt]
       |
  == layer 0 ==================================================================
   rms_norm(4) <- h                                  "attn_norm", eps 1e-6
     |-- matmul(5) blk.0.attn_q     [896 -> 896]   Q = Wq @ normed (+ bias)
     |-- matmul(6) blk.0.attn_k     [896 -> 128]   K = Wk @ normed (+ bias)
     +-- matmul(7) blk.0.attn_v     [896 -> 128]   V = Wv @ normed (+ bias)
   rope(8)  <- n5, positions                        rotate Q by position
   rope(9)  <- n6, positions                        rotate K by position
   kv_store.0(10) <- n9(K), n7(V), positions        write into layer-0 KV region
   kv_load.0(11)                                    view of that region (no edges in)
   attn(12) <- n8(Q), n11(KV), positions            GQA attention, scale 1/sqrt(64)
   matmul(13) blk.0.attn_output    [896 -> 896]     project heads back
   add(14) <- embed(3), n13                          <-- FIRST residual add
   rms_norm(15)                                      "ffn_norm"
   matmul(16) blk.0.ffn_gate      [896 -> 4864]
   matmul(17) blk.0.ffn_up        [896 -> 4864]
   silu(18)  <- n16                                  (orphaned by fusion, see §2.6)
   SwiGLU(19) <- n16, n17                            silu(gate) * up (fused form)
   matmul(20) blk.0.ffn_down      [4864 -> 896]
   add(21) <- n14, n20                               <-- SECOND residual add
  == layers 1..22 identical (18 nodes each) ==================================
  == layer 23 (the last): as above, plus two GetRows nodes (428, 429) =========
   rms_norm(438)                                     final norm
   matmul(439) output (lm_head)   [896 -> 151936] -> OUTPUT logits

Reading it top to bottom is reading the forward pass. What a beginner should take away is the shape of a transformer layer, which the IR makes unmistakable:

  • An attention block: normalize, then three matrix multiplies produce Q (query), K (key), V (value) — attention's three roles; the positions are mixed in by RoPE (Rotary Position Embedding, which rotates Q and K so that attention scores depend on token distance); the fresh K and V are written into the layer's KV cache region (persistent memory holding every past token's keys and values — the thing that makes decoding cheap, doc 11); then one attention operation reads Q plus the whole KV region, and one final matrix multiply projects the result back to the hidden width.
  • An FFN block (feed-forward network): normalize, two wide matrix multiplies (gate, up), the SwiGLU activation — silu(gate) * up, where SiLU is the smooth gate function x·sigmoid(x) — then a narrow matrix multiply back down.
  • Two residual adds — h = h + attention(...) and h = h + ffn(...). Each block computes a correction to the running hidden state instead of replacing it; the add nodes are the only places the hidden state is rewritten.

Nodes 4–21 are layer 0 (0–2 are the input leaves, 3 the embedding); layers 1 through 22 repeat the same 18-node pattern (layer 23, with the two extra get_rows of §2.8, spans nodes 418–437); node 438 is the final norm and 439 the lm_head. The full prefill graph is 440 nodes (measured: --dump-graph on 0.5B, 30-token prompt, n_out = 1). The complete census:

Op kindCountWhere
MatMul1697 per layer (q,k,v,wo,gate,up,down) × 24 + lm_head
RmsNorm492 per layer × 24 + 1 final
RoPE482 per layer (q, k)
Add482 residual adds per layer
SwiGLU241 per layer (rewritten from silu+mul)
Silu24orphans left behind by that rewrite
KvcacheStore / KvcacheLoad24 / 241 each per layer
Attn241 per layer
GetRows3embed + 2 tail-row selects
Input3token_ids, positions, tail_ids

A note on GQA — grouped-query attention, the Qwen2 trick visible in the shapes above. Full multi-head attention stores a key and a value per query head: with 14 heads of dim 64, that is 14 × 64 = 896 numbers per token per layer, per K and V. Qwen2.5-0.5B instead shares each key/value among a group of 7 query heads: only n_head_kv = 2 KV heads exist, so K and V are 128-wide (n_kv_embd = 128), and the KV cache shrinks 7×. The graph expresses this in plain shapes — the wk/wv matmuls output 128, not 896 — and the attn node's metadata carries the mapping (n_head, n_head_kv, head dims, the KV row stride nkt, the scale 1/sqrt(64)).

2.6 Prefill vs decode: same skeleton, two fusion classes

A forward pass is built for exactly one n_tokens:

  • Prefill (nt > 1): all prompt tokens flow through together; the KV store writes a block of positions and attention reads the fresh prefix.
  • Decode (nt == 1): one token at a time — same 18-node skeleton, the shapes just narrow to [.., 1].

Decode is where minfer buys speed with two build-time fusions — extra builder constructors that emit fewer, bigger nodes. They are decided inside build itself (gated on nt == 1 and a GPU backend), not by a later pass:

  • FusedQKV (the concat class): instead of 3 matmul + 3 bias-add + 2 rope + 2 store = 10 kernel dispatches per layer, one concat matmul against a pre-concatenated blk.{i}.attn_qkv weight, then one fused kernel applying the three biases, roping Q and K, and storing K/V — 2 dispatches. Attention reads Q from offset 0 of the concat buffer.
  • FusedFFN: instead of gate matmul + up matmul + silu + mul = 4 dispatches, one concat matmul against blk.{i}.ffn_gu plus one in-place SwiGLU pass — 2 dispatches. Measured on the 0.5B: decode ~269 → ~299 tok/s (+~11%) for QKV, ~303 → ~312–331 (+~3%) for FFN (COMPUTE-GRAPH-DESIGN.md §17, Phases 10–11).

Both fusions are gated by measurement, not ideology: FusedFFN is built only when nf ≤ 16384, because on the 7B model (nf = 18944, so the concat matmul's output is 37 888 wide) the single wide matmul measured slower than two narrow ones — 42.5 vs 46.7 tok/s — and the gate turns it off there. Mixed-quantizer layers (e.g. Q6_K attention-V among Q4_K q/k) cannot share a concat weight at all; for those, a second class (qkv_bias_rope_store) keeps three separate matmuls and fuses only the epilogue. Qwen3 has its own variant, FusedQkvNorm, which additionally folds the per-head Q/K RMSNorm that Qwen3 requires: 3 matmul + 2 qk_norm + 2 rope + 2 store → 2.

Also note the SwiGLU/Silu rows in the census: on the unfused path the builder still emits silu then mul, and a later pass (doc 06) rewrites the mul into SwiGLU in place, orphaning the silu node. The orphan stays in the node list — the scheduler skips nodes without buffers — which is why both counts are 24. In the DOT dump you can literally see the stale edge n16 → silu(18) alongside the fused n16 → SwiGLU(19).

2.7 The reuse invariant: topology = f(GraphParams)

Here is the sentence the whole design stands on: the graph topology is a deterministic function of GraphParams — equal parameters produce an identical graph, node for node. GraphParams (src/graph/params.rs:88-98) contains exactly n_tokens, n_out (tail rows, §2.8), gtype (prefill or decode), cparams (context size, flash-attn flag, GPU participation, the two fusion toggles), and weights_version.

Conspicuously absent: n_past — how many tokens are already in the KV cache. That number changes on every decode step, and it is deliberately not part of the structure. Where does the position live instead? In an input node: the positions leaf (node 1) is filled with fresh values before every execution, and the KV store/load and RoPE nodes read it as data. Positions are data, not structure — the one invariant everything else protects (ops.rs:3-6 states it; §3.3 Q5 asks "what breaks without it").

Given that invariant, reuse is a six-field comparison: GraphCache::try_reuse compares the new GraphParams against the cached one; equal means the graph — and the allocator, with its persistent KV regions — are reused as-is, and only the input data is refreshed. During a 500-token generation the prefill graph is built once, the decode graph is built once (the first decode step, because nt changed 30 → 1), and the remaining ~500 steps replay it.

2.8 The n_out tail: shrink the graph to the rows you sample

One forward pass produces logits for some tokens, but generation only ever samples the last token (single sequence). Computing lm_head — the final [896 → 151936] matmul — for all 30 prompt tokens costs 30 × 896 × 151936 ≈ 4.1 × 10⁹ multiply-accumulates; for the last token only, ≈ 1.4 × 10⁸. minfer, like llama.cpp's inp_out_ids, draws the line further up: after the last layer's attention output projection, two GetRows nodes select the tail n_out rows — of the attention output and of the residual — so the last FFN block, both final residual adds, the final norm, and lm_head all run on n_out rows instead of nt. That is nodes 428/429 in the dump; with n_out = 1 the entire output stack is one row. The row indices are the tail_ids input node (node 2) — again data, not structure — filled with [nt-n_out .. nt) before execution. Measured impact on 0.5B prefill: ~3900–4000 tok/s, ~+55% over the full-nt graph (plan §17, deviation 16).

One subtlety worth admiring: tail_ids is declared at the head of the graph, next to the other inputs, though its consumers sit at the very end. The code comment explains why (graph.rs:55-60): a mid-graph input node would split execution into extra backend boundary segments, each costing full-stream syncs and host round trips on the GPU path. Node order is not semantics — only the edges are — so declaration position is free to optimize execution, not readability.

3. Implementation

3.1 Data in / data out

In (from doc 04 and doc 03):

  • tokens: &[u32] — the tokenized prompt (prefill) or the single generated token (decode). Becomes the token_ids input node, [nt, 1, 1, 1], I32.
  • positions: &[usize] — the KV slot of every token: 0..len for prefill, the running position per decode step. Becomes the positions input node, same shape, I32.
  • The loaded model: hparams (layer count, head counts, FFN width, eps, RoPE base/scale) and the weight tensors — which the builder references by name only (MatMulMeta.weight_name = "blk.7.attn_q.weight"), not by pointer. The graph stays pure data; the allocator resolves names to buffers at execution time.
  • GraphParams — the full set of structure-deciding knobs (§2.7).

Out:

  • A ComputeGraph: 440 nodes (0.5B prefill with n_out = 1), inputs = [0, 1, 2], outputs = [439] (the lm_head matmul). No buffers, no backends, no numbers — nothing computed.

Two layout conventions the shapes encode (from docs/ARCHITECTURE.md §4.5): weights are metadata [in, out] but memory row-major [out][in] — so a matmul's output dim is w.shape[1]; activations are shape [d, nt, 1, 1] with token-major memory [nt][d]; and integer inputs (I32) are stored bit-exactly inside f32 buffers as f32::from_bits patterns (fill_input_i32, doc 07) so one buffer pool serves both dtypes.

A sense of scale for the persistent parts the graph merely declares: each layer's K region is [n_kv_embd=128, n_ctx=4096] f32 = 2 MB; K + V per layer = 4 MB; 24 layers ≈ 101 MB of KV the allocator must keep alive across every rebuild (doc 07 owns that machinery).

3.2 Key code

The node and the graph (src/graph/mod.rs:82-115)

#![allow(unused)]
fn main() {
/// Single compute node.
#[derive(Debug, Clone)]
pub struct CNode {
    pub id: NodeId,
    pub name: String,
    pub op: Op,                 // operation + full payload
    pub src: Vec<NodeId>,       // input dependencies (the edges)
    pub out_shape: [usize; 4],  // output shape [d, nt, 1, 1] convention
    pub out_dtype: DType,       // output element type
    /// Backend assigned by the scheduler (Phase 4); None = undecided.
    pub backend: Option<Backend>,
    pub meta: NodeMeta,         // weight names, rope/attn params, ...
}

/// Compute graph: topologically ordered node sequence + input/output sets.
#[derive(Debug, Clone, Default)]
pub struct ComputeGraph {
    pub nodes: Vec<CNode>,
    pub inputs: Vec<NodeId>,    // input nodes that need external filling
    pub outputs: Vec<NodeId>,   // output nodes (logits, etc.)
    /// Graph identifier for reuse detection (CUDA Graph caching etc.).
    /// Populated by `GraphCache::replace_graph` (monotonic per process);
    /// a reused graph keeps its uid.
    pub uid: u64,
}
}

Note what is absent: no buffers, no data pointers, no backend — decisions for later stages. The IR records only structure and shape; the backend: Option<Backend> slot exists so doc 06 can fill it in place, and at this stage it is None everywhere.

Topological order is not maintained — it is implied. The topo_order() doc comment (src/graph/mod.rs:151-153) says it outright: "The builder appends sources before consumers, so nodes is already topologically ordered; this validates the invariant and returns a stable order (used by the allocator)." Because a node can only reference ids the builder has already returned, nodes is sorted the moment it is built. topo_order() still runs a full Kahn pass — but its real job is validation (cycle detection, dangling ids), and its result is deliberately not used as the execution order (a G3 bug story told in doc 07).

Operations with payloads (src/graph/ops.rs:36-101, abridged)

#![allow(unused)]
fn main() {
/// Operator type. Implements full `PartialEq` (payloads included) so debug
/// builds can verify graph-rebuild structural identity; the production graph
/// reuse decision is params-only (see docs/COMPUTE-GRAPH-DESIGN.md §6).
#[derive(Debug, Clone, PartialEq)]
pub enum Op {
    /// Leaf input node (token ids, positions, KV idx, ...). Filled externally
    /// each step via `GraphAllocator::fill_input`; never part of the topology.
    Input,

    Add,
    Mul,
    Silu,
    // ... Scale(f32), Softmax { dim } — ggml-parity vocabulary, no live
    //     architecture emits them today ...
    RmsNorm {
        eps: f32,
    },
    QkNorm {
        hd: usize,
        nh: usize,
        eps: f32,
    },
    MatMul {
        transpose_b: bool,
    },
    GetRows,
    RoPE {
        style: RopeStyle,
    },
    Attn {
        mode: AttnMode,
    },
    // ---- KV cache (persistent external buffer; positions are data) ----
    KvcacheStore {
        layer: usize,
    },
    KvcacheLoad {
        layer: usize,
    },
    // ... fused: SwiGLU, FusedQKV{layer},
    //     QkvBiasRopeStore{layer}, FusedFFN, FusedQkvNorm{layer}
}

Two things to notice. First, the KV ops carry only the layer index — no position, no length. The module doc at ops.rs:3-6 states the invariant: positions are data, injected via the positions input node, so the graph topology never depends on n_past. Second, the PartialEq derive is annotated with its purpose: structural identity checks for the reuse invariant.

Weights and other per-node parameters travel in NodeMeta, a concrete enum (ops.rs:160-178) rather than the plan's original Box<dyn Any> — a recorded deviation that buys PartialEq, no downcast panics, and Clone nodes. The workhorse is MatMulMeta { weight_name, bias_name, weight_ttype, in_dim, out_dim }: the graph references the weight by name ("blk.7.attn_q.weight"), and the backend resolves that name to a registered buffer at execution time. That indirection is what keeps the IR pure data — and it is why doc 03 registered every weight under exactly these GGUF names.

The builder's only real primitives are node(name, op, src, out_shape, out_dtype, meta) — which appends a CNode with id = graph.nodes.len() and backend: None, and returns that id — and input(name, shape, dtype), which creates an Op::Input leaf with no sources, records its id in graph.inputs, and is otherwise an ordinary node (its doc comment already names the payoff: "so n_past/positions changes never force a graph rebuild"). Everything else in the 496-line file is sugar over node.

The builder in practice (src/graph/builder.rs:127-145)

A representative convenience method, matmul:

#![allow(unused)]
fn main() {
    pub fn matmul(&mut self, x: NodeId, w: &Tensor, bias: Option<&Tensor>) -> NodeId {
        let out = w.shape[1] as usize;   // GGUF: out dim = shape[1]
        let nt = self.graph.nodes[x].out_shape[1];
        let name = format!("matmul_{}", w.name);
        self.node(
            &name,
            Op::MatMul { transpose_b: false },
            &[x],
            [out, nt, 1, 1],             // output width, then token count
            DType::F32,
            NodeMeta::MatMul(MatMulMeta {
                weight_name: w.name.clone(),
                bias_name: bias.map(|b| b.name.clone()),
                // ... weight_ttype, in_dim = w.shape[0], out_dim = w.shape[1]
            }),
        )
    }
}

Shapes propagate automatically: the output width comes from the weight, the token count from the input. The bias is not a separate node — it is a name in the metadata, and each backend's matmul kernel applies it inline (a small taste of how aggressively this IR avoids node spam). transpose_b is false everywhere in the live graph: minfer stores weights row-major [out][in] and never needs the transposed form; the flag exists for ggml vocabulary parity.

And the KV pair (src/graph/builder.rs:368-389, store side — trimmed):

#![allow(unused)]
fn main() {
    /// Write this step's K/V into the layer's persistent KV region at the
    /// positions carried by `pos`. `n_ctx` sizes the persistent region.
    pub fn kvcache_store(
        &mut self,
        layer: usize,
        k: NodeId,
        v: NodeId,
        pos: NodeId,
        n_ctx: usize,
    ) -> NodeId {
        let n_embd = self.graph.nodes[k].out_shape[0];
        // shape mirrors the persistent region so the allocator can size it
        self.node(&format!("kv_store.{layer}"), Op::KvcacheStore { layer },
                  &[k, v, pos], [n_embd, n_ctx, 1, 1],
                  DType::F32, NodeMeta::Kvcache(/* ... */))
    }
}

The store node reads three sources — the rotated K, the raw V, and the positions input — and its declared output shape [n_kv_embd, n_ctx] is the allocator's sizing contract for the persistent region. kvcache_load is even simpler: no sources at all (it is a view of the region, which is why kv_load nodes have no incoming edges in the DOT dump), carrying {layer} again and nothing positional.

The main event, part 1: inputs and the QKV branch (src/models/qwen2/graph.rs:53-130)

#![allow(unused)]
fn main() {
        let inp_ids = b.input("token_ids", [nt, 1, 1, 1], crate::graph::DType::I32);
        let inp_pos = b.input("positions", [nt, 1, 1, 1], crate::graph::DType::I32);
        // G3 tail-row reduction input, declared at the graph HEAD (not beside
        // its consumers at the last layer): an input node mid-graph splits the
        // forward into extra CPU/CUDA boundaries (2 full-stream syncs + host
        // round-trip copies per step on the split path). R3-A1,
        // docs/CUDA_OPTIMIZATION.md Part III. Node order is not semantics —
        // the consumers below just reference the handle.
        let tail_ids = (params.n_out < nt).then(|| {
            b.input(
                "tail_ids",
                [params.n_out, 1, 1, 1],
                crate::graph::DType::I32,
            )
        });

        let mut h = b.embedding(inp_ids, model.tok_embd.as_ref().unwrap());
}

Then, inside the for (il, l) in model.layers.iter().enumerate() loop, after rms_norm, the Q/K/V projection has three build-time classes, selected by a gate that reads only nt, GraphParams, and the layer's own tensors:

#![allow(unused)]
fn main() {
            let fuse_qkv = nt == 1
                && params.cparams.gpu
                && params.cparams.fuse_qkv
                && l.bq.is_some()
                && l.bk.is_some()
                && l.bv.is_some();
            let (q, kv) = if fuse_qkv && Self::qkv_concat_available(&l.wq, &l.wk, &l.wv) {
                let qkv = b.fused_qkv(
                    normed,
                    inp_pos,
                    il,
                    FusedQkvMeta {
                        qkv_weight: format!("blk.{il}.attn_qkv"),
                        // ... biases, dims, rope params, kv_elems ...
                    },
                );
                // q lives at concat offset 0 (rows 0..nqt); K/V went into the
                // persistent regions via the fused store — read them back.
                let kv = b.kvcache_load(il, nkt, n_ctx, nk);
                (qkv, kv)
            } else if /* mixed-quant class: 3 matmuls + qkv_bias_rope_store */ {
                // ...
            } else {
                let q = b.matmul(normed, l.wq.as_ref().unwrap(), l.bq.as_ref());
                let k = b.matmul(normed, l.wk.as_ref().unwrap(), l.bk.as_ref());
                let v = b.matmul(normed, l.wv.as_ref().unwrap(), l.bv.as_ref());
                let q = b.rope(q, inp_pos, hp.rope_style, RoPEMeta { /* ... */ });
                let k = b.rope(k, inp_pos, hp.rope_style, RoPEMeta { /* ... */ });
                b.kvcache_store(il, k, v, inp_pos, n_ctx);
                let kv = b.kvcache_load(il, nkt, n_ctx, nk);
                (q, kv)
            };
}

This is where "the graph is data" pays its first visible dividend: the decision about decode fusion is an ordinary if in a pure function, and each branch emits a different topology. A different GraphParams (nt = 1 vs 30, or MINFER_NO_FUSE_QKV=1 in the environment) yields a different graph — and because the deciding values all live in GraphParams, the difference is exactly reproducible.

The main event, part 2: attention, tail rows, FFN (src/models/qwen2/graph.rs:201-269, abridged)

#![allow(unused)]
fn main() {
            // attention
            let attn_out = b.attn(
                q,
                kv,
                inp_pos,
                mode,
                AttnMeta {
                    layer: il,
                    n_head: nh,
                    n_head_kv: nk,
                    hd,
                    hd_kv,
                    nkt,             // KV row stride = n_kv_embd
                    scale: attn_scale,   // 1/sqrt(hd), loader.rs:36-38
                },
            );

            // output projection + residual
            let wo = b.matmul(attn_out, l.wo.as_ref().unwrap(), None);
            let is_last = il == model.layers.len() - 1;
            if is_last && params.n_out < nt {
                let tail_ids = tail_ids.expect("tail_ids input declared when n_out < nt");
                let cur_tail = b.get_rows(wo, tail_ids, [ne, params.n_out, 1, 1]);
                let res_tail = b.get_rows(residual, tail_ids, [ne, params.n_out, 1, 1]);
                h = b.add(res_tail, cur_tail);
            } else {
                h = b.add(residual, wo);
            }
}
#![allow(unused)]
fn main() {
            let fuse_gu = nt == 1
                && params.cparams.gpu
                && params.cparams.fuse_ffn
                && Self::gu_concat_available(&l.ffn_gate, &l.ffn_up)
                && nf <= 16384;      // 7B: concat matmul measured slower
            // ... rms_norm(ffn_norm) ...
            let ffn_out = if fuse_gu {
                let gu = b.fused_ffn(normed, FusedFfnMeta { /* ... */ });
                b.matmul(gu, l.ffn_down.as_ref().unwrap(), None)
            } else {
                let gate = b.matmul(normed, l.ffn_gate.as_ref().unwrap(), None);
                let up = b.matmul(normed, l.ffn_up.as_ref().unwrap(), None);
                let g = b.silu(gate);
                let sw = b.mul(g, up);
                b.matmul(sw, l.ffn_down.as_ref().unwrap(), None)
            };
            h = b.add(residual, ffn_out);
}

After the loop, two lines finish the pass: rms_norm(h, output_norm) then matmul against model.output (lm_head), marked with b.output(logits), and b.build() hands back the finished ComputeGraph. Two details deserve a pause: the attention node takes the positions input as a third source — attention needs it for causal masking (token t may attend to positions 0..=pos[t]; builder comment at builder.rs:325-327) — and on the last layer the residual is also narrowed by a second get_rows, because both sides of the add must shrink or the shapes would disagree.

The reuse check (src/graph/cache.rs:47-64)

#![allow(unused)]
fn main() {
    /// Params-only reuse check. On success the previously stored graph is
    /// reused without rebuilding (caller then refreshes input data).
    pub fn try_reuse(&mut self, params: &GraphParams) -> bool {
        match (&self.prev_params, &self.graph) {
            (Some(prev), Some(_)) if Self::params_match(prev, params) => {
                self.prev_params = Some(params.clone());
                true
            }
            _ => false,
        }
    }

    fn params_match(a: &GraphParams, b: &GraphParams) -> bool {
        a.n_tokens == b.n_tokens
            && a.n_out == b.n_out
            && a.gtype == b.gtype
            && a.cparams == b.cparams
            && a.weights_version == b.weights_version
    }
}

No node walk, no comparison of 440 nodes — six field comparisons, because §2.7's invariant makes them sufficient. The caller (Qwen2Graph::forward_cached, models/qwen2/graph.rs:456-472) shows the whole production loop in one glance: try_reuse; if it fails, build → register weights → assign_backends (doc 06) → FusionPass (doc 06) → alloc_graph (doc 07) → store in the cache with a fresh uid; then execute (doc 08) and refresh input data. Notice that CParams.fuse_qkv/fuse_ffn are read from the environment at params construction (MINFER_NO_FUSE_QKV=1 etc.), so the A/B toggles force rebuilds through the front door — a test (fuse_flags_are_part_of_the_reuse_identity) guards exactly that.

3.3 Design choices (why this shape and not another)

Q1: Why a declarative graph instead of the obvious imperative loop? The loop is genuinely simpler to write — the old forward.rs did exactly that, and it worked. Four things it could not do well:

  1. Decode-step reuse. With a graph, one token of generation = fill 2–3 input buffers + walk the node list; the expensive part (build → assign → fuse → allocate) runs once per distinct GraphParams.
  2. Backend assignment before execution. Deciding "this node runs on Metal" requires the global picture before running anything, so splits and cross-backend copies can be planned (doc 06/08). A loop that computes as it goes must decide mid-flight — which is how the old design ended up with per-layer fallback heuristics and surprise host round trips (ARCHITECTURE.md Appendix A.3).
  3. Fusion as a rewrite pass. Mul(Silu(x), y) → SwiGLU is a few lines of pattern matching over the IR (fusion.rs:39-87), applied wherever a backend has the kernel. In an imperative loop the same optimization is hand-woven into the control flow of every model's forward code.
  4. Observability. --dump-graph / --dump-graph-json are ~40 lines each because the forward pass is data (dot.rs:13-55). The old loop had to be reverse-engineered from breakpoints.

The cost is real but bounded: 273 + 322 + 496 lines of IR/builder, and every model expresses its forward as builder calls (Qwen2's build function is ~240 of its 2244 lines). In exchange, the old forward.rs was deleted outright once the graph path matched it bit-for-bit (Phase 6, plan §17: "after deletion full suite 78 pass; CLI default-path output consistent with the old implementation"). Nothing regression-tested was lost.

Q2: Why must building be pure / side-effect-free? Because reuse requires determinism: try_reuse compares six parameters and then skips building entirely — safe only if the skipped build would have produced the same graph. Purity is what makes "same params ⇒ same graph" even meaningful, and the debug structural check enforces it (rebuild and compare, node for node). Purity also makes building cheap and repeatable: the dump path builds a second graph just to export it (json.rs::build_runtime_graph), no GPU context is touched, and there is no hidden state to invalidate. The moment someone reaches into the allocator or the GPU during build, the invariant becomes unverifiable.

Q3: Why does Op carry full payloads? So that graphs are comparable and self-describing. Comparison is concrete: GraphCache::verify_structural walks two graphs asserting op == op per node, and the unit test op_partial_eq_compares_payloads (mod.rs:255-264) pins the semantics — RmsNorm{eps: 1e-5} ≠ RmsNorm{eps: 1e-6}, KvcacheLoad{layer: 2} ≠ {layer: 3}. Without payloads in the equality, two graphs differing only in an eps or a layer index would look identical and the tripwire would be blind. Self-description is the DOT dump of §2.5 — KvcacheStore { layer: 23 } printed on the node — and the same payloads drive execution: the scheduler resolves the layer's persistent K/V regions from KvcacheStore{layer}, and backends read AttnMeta.scale/nkt directly.

Q4: Why are KV store/load explicit graph nodes at all? The KV cache is "just memory" — the old design hid it inside a cache object. Making store and load nodes buys three things. First, ordering for free: execution follows build order, and the store node is built before the attention that loads — so "this step's K/V are written before attention reads them" is list order, guaranteed by the executor rather than by discipline (doc 08; ARCHITECTURE.md §4.5.5). Second, the allocator sees the truth: the store node's declared shape [n_kv_embd, n_ctx] sizes the persistent region, and the load node's absence of edges marks the region as alive — which is how KV regions survive graph rebuilds while ordinary buffers are recycled by liveness (doc 07). Third, backends resolve locality: the layer index lets any backend find the sibling K/V regions (kv_pair(layer)) — attention and its KV stay on the same backend by construction, with no per-token KV drain.

Q5: What breaks if topology secretly depends on n_past? Everything in §2.7, concretely. Suppose the build baked the cache length into node shapes — say attention reading a KV view of [n_kv_embd, n_past + nt]:

  • params_match (which does not compare n_past) would call two different graphs "equal" — step 5 would silently reuse step 4's graph, reading the wrong rows. And if you instead added n_past to GraphParams, try_reuse would fail on every decode step: the graph rebuilds every token, the allocator recomputes its liveness mapping every token, and the uid churns, defeating the CUDA Graph replay cache keyed on it — decode throughput collapses to build-and-allocate speed.
  • The debug structural check would catch the mismatch (equal params, different topology ⇒ panic) — but only in test builds; production takes the silent path.

This is why the invariant is repeated in the three places a reader will trip over it — the ops.rs module doc, the params.rs module doc ("n_past deliberately absent: it is execution data"), the builder's input doc comment — and asserted by the unit test kv_nodes_carry_layer_only (builder.rs:470-484), the smallest expression of "positions are data".

3.4 Pitfalls & invariants

  1. KV positions are data, not structure. KvcacheStore/Load carry only {layer}; positions arrive via the positions input node. Topology must never branch on n_past (Q5). Guarded by kv_nodes_carry_layer_only.
  2. Topological by construction. A node can only reference ids already returned by the builder; topo_order() validates rather than sorts. Execution order = build order (doc 08) — which is also why a KV store built before its attention is guaranteed to run before it.
  3. Node order is not semantics — the edges are. tail_ids is declared at the graph head but consumed at the last layer (graph.rs:55-60), purely to keep input nodes from fragmenting execution into extra backend boundary segments. When reading build, follow src, not position.
  4. Fused and unfused are both real graphs, and they must agree. The fusions are gated by GraphParams fields, so both topologies are first-class, A/B-able via environment variables, and bit-identical in output (asserted by fused_qkv_matches_unfused_decode). One recorded lesson (plan §17.26): when comparing fused vs unfused, the unfused path must still run the FusionPass — an unfused graph executed without it computes silu+mul as two kernels, differing from the single SwiGLU kernel by ~1e-6 float noise that amplifies at large magnitudes.
  5. Fusion orphans stay in the node list. After Mul→SwiGLU rewrites, the old Silu node has no consumers and gets no buffer; the executor skips bufferless nodes (scheduler.rs:228-233: "dead nodes … are skipped, not executed"). Do not panic at a node count that exceeds the number of "real" operations.
  6. Attention's output shape comes from metadata, not from Q. With FusedQKV, Q is a slice of a wider concat buffer, so attn sizes its output from AttnMeta (builder.rs:607). Shapes in this IR are declared facts, not derived ones.
  7. GPU participation and fusion toggles are part of the reuse identity (CParams.gpu, fuse_qkv, fuse_ffn): they change the topology, so they must (and do) change the reuse decision — a test locks this in after it was nearly broken (cache.rs:135-140).
  8. Weights are referenced by name, registered elsewhere (doc 03): a MatMulMeta.weight_name with no registered weight is an execution-time error — the IR cannot check it, and purity forbids it from trying.

4. Observe & verify

  • --dump-graph <path> — exports the forward pass as Graphviz DOT (nodes colored by assigned backend, inputs/outputs as double circles) and exits; the run prints the node count. The §2.5 census came from MINFER_DISABLE_MPS=1 ./target/release/minfer <model> "Hello" --dump-graph /tmp/g.dot (440 nodes) and the same with --no-template and a one-token prompt (437, the decode graph).
  • --dump-graph-json <path> — the same graph as JSON for viz/ (web visualizer); the JSON also carries the GraphParams that produced it.
  • minfer viz <model> — live graph page with per-node data over SSE, built by the same build_runtime_graph helper.
  • MINFER_TRACE=<path> — per-node real-data trace during a run; the bridge from this doc's structure to doc 08's execution.
  • MINFER_GRAPH_DUMP=<dir> — dumps per-node logits/KV outputs of the live graph (any build) for offline comparison.
  • Tests (all in the current tree): topo_order_validates_chain / topo_order_detects_cycle / op_partial_eq_compares_payloads (graph/mod.rs), builder_creates_topo_sorted_graph / kv_nodes_carry_layer_only (graph/builder.rs), the reuse quartet reuse_requires_equal_params / fuse_flags_are_part_of_the_reuse_identity / allocator_survives_rebuild / structural_check_detects_different_graph (graph/cache.rs), and the model-level graph_logits_match_forward_real_model / tail_reduction_matches_full_nt / fused_qkv_matches_unfused_decode (models/qwen2/graph.rs). Historical acceptance bar: graph logits max diff 0.000e0 vs the old imperative path, prefill and decode (plan §17, Phases 5 and 8).

5. Cross-references

  • 04 — Tokenizer and chat template: produces the token ids this stage wraps into the token_ids input node.
  • 06 — Backend assignment and fusion: walks this graph and fills every backend: None; runs the FusionPass whose Mul∘Silu → SwiGLU rewrite orphaned the silu nodes in the census.
  • 07 — Allocator, liveness, KV regions: turns shapes into buffers, shares memory by liveness, and owns the persistent per-layer KV regions this stage merely declared.
  • 08 — Scheduler + execute: executes nodes in build order — the guarantee behind "store before the attention that reads it".
  • 09 — Prefill forward path and 13 — Decode loop + graph reuse: the two callers of forward_cached — where GraphParams comes from and how the params-only reuse behaves across hundreds of steps.
  • 11 — Attention + vec ops + KV: the kernels behind RmsNorm, RoPE, Attn, and the KV store/load — including how positions drive masking and cache writes.
  • 03 — Model dispatch and weights: registered every weight under the names this stage's metadata references.
  • docs/ARCHITECTURE.md §4 — the verified design summary this doc expands (§4.6 has the mermaid layer diagram); Appendix A preserves the imperative design this replaced.
  • docs/COMPUTE-GRAPH-DESIGN.md §3–6 — the design record (IR, builder, fusion rules, reuse) and §17 — the phase-by-phase implementation log with the measured fusion/tail numbers quoted above and the full deviation list.
  • docs/LLAMA-COMPUTE-GRAPH.md: the llama.cpp equivalent (ggml_cgraph, llm_graph_context, inp_out_ids) minfer mirrors. viz/README.md: the visualizer data format; docs/GLOSSARY.md: backstop definitions (GQA, RoPE, SwiGLU, liveness…).

← 04 — Tokenizer and chat template · Index · 06 — Backend assignment and fusion →

06 · Assign + fusion — every node picks an engine, then patterns fold

Stage: 05 — Graph build: the IR and the builder → assign backends + fuse op patterns → 07 — Memory allocation: liveness and the KV regions (row 6 of the README master table: the graph leaves its "pure IR" form here and becomes executable, but nothing has been allocated or run yet). Code: src/graph/scheduler.rs::assign_backends (line 60) · src/graph/alloc.rs::supports (line 140) · src/graph/backend.rs (the Backend trait) · cpu_backend.rs / metal_backend.rs / cuda_backend.rs (supports_op, supports_fused) · src/graph/fusion.rs::run (line 23) · call site src/models/qwen2/graph.rs::forward_cached (line 472, mirrored in qwen3/graph.rs) · src/graph/params.rs (CParams).

1. Background — where this stage sits

Doc 05 left us with a ComputeGraph: a plain data structure that lists every math operation of the transformer as a node, with edges pointing from each node to the nodes that produce its inputs. Nothing has been computed; nothing has been allocated; the graph is a description, not a program. In compiler language this description is called an IR (intermediate representation) — a neutral, engine-independent way of writing down "what should be computed" so that several different engines can later agree on how to compute it.

This stage performs the two transformations that turn that description into something an engine can actually run:

  1. Backend assignment. Every node gets a backend — one of the engine implementations that can execute a node: CPU, Metal (macOS GPU), or Cuda (--features cuda, NVIDIA GPU). A backend is more than a pile of math code: it owns its own buffer pool (its own memory for intermediate results), its own kernels (the small, hand-tuned functions that actually crunch numbers — a matmul kernel, a RoPE kernel, …), and the machinery to schedule that work on its device. Assignment answers one question per node: who computes this?
  2. Fusion. A rewrite pass walks the graph and looks for small, fixed patterns of nodes that some backend knows how to compute in a single kernel — a fused op. When both the pattern matches and the node's assigned backend says "I have a kernel for that", the pass replaces the pattern with one fused node. This stage ships one pattern: Mul(Silu(x), y) folds into a single SwiGLU node. (The plan's second pattern, RoPE(Add(x, b)) → FusedBiasRope, was removed once it became clear no backend would claim the capability — the bias+rope work ships as the builder-built decode nodes; see §2.3.)

Both transformations happen once per graph build, not once per token. The engine builds the graph on the first forward (and on any rebuild), runs assign → fusion immediately after, and then stores the finished graph in the GraphCache. Every decode step afterwards reuses that exact graph and only refreshes its input data (doc 13). So the cost of assignment and fusion is paid a handful of times per run, while their benefit — cheaper execution — is paid off on every single forward.

Why does this stage exist at all? Two reasons, one per transformation.

Without backend assignment, the engine would have to pick an engine while running — and the tempting version of that ("this kernel failed on the GPU, let's quietly redo it on the CPU") is exactly what minfer forbids. Silent mid-run fallback makes results backend-dependent in surprising ways, hides kernel bugs, and makes execution non-deterministic. minfer's convention (documented in docs/GPU_SAFETY.md) is the opposite: the decision is made once, at build time, and if a kernel invariant is violated at run time the run aborts with an error rather than falling back.

Without fusion, the graph would execute every tiny operation as its own kernel launch. Each launch has a fixed cost (on a GPU it means encoding and submitting work; on the CPU it means a thread-pool dispatch), and each intermediate result is written to memory and read back. Elementwise ops like SiLU and Mul are dominated by exactly those memory round trips — the arithmetic itself is trivial. Fusing the two-launch, three-memory-trip pattern into one kernel removes a full write + read of the intermediate buffer and one launch, per layer, per forward. §2.3 does that arithmetic with real byte counts.

One piece of vocabulary before we dive in: a dispatch (or launch) is the act of handing one kernel invocation to an engine — on the CPU, waking the thread pool for one operation; on Metal, recording the kernel into a command buffer (a batched to-do list for the GPU; doc 08 covers submission). Dispatches are individually cheap but never free, and decode (one token per forward) multiplies their cost by every layer of the model.

2. Principle — how it works and why

2.1 Capability-driven assignment: ask each engine, in priority order

The policy that assigns backends is one line long (we'll read it in §3.2):

for every unassigned node: the first backend — trying Metal, then CUDA, then CPU — whose supports_op(op, dtype) answers "yes" gets the node; if nobody answers, the node goes to CPU.

supports_op is a capability query: a pure yes/no function from (operation, data type) to bool, implemented by each backend about itself. The scheduler never consults a master table of "which ops run where"; it just asks the engines in order. The priority order is the order of the checks inside the allocator's supports function — Metal is asked first because, when available, it is the fastest engine; CUDA second; CPU last, and CPU is also the landing pad (or(Some(CPU))) when nothing else is enabled.

dtype (data type) matters here: nodes carry out_dtype (f32 for activations; i32 inputs are stored as f32 bit patterns — a doc 05 fact). A backend that only has f32 kernels answers "no" for anything else, and the query result reflects what that backend can actually execute, not what would be nice to execute.

Two facts make GPU assignment possible at all, and both are decided before this stage runs:

  • Weights must already live on the GPU. A matmul node assigned to Metal is useless if the weight tensor it needs is only on the host. The model layer therefore gates GPU participation all-or-nothing: metal_on = metal_available() && Self::weights_on_gpu(model) — either every matmul weight is registered on the GPU registry, or the whole graph stays on CPU (src/models/qwen2/graph.rs:431-433; the CUDA arm is the same shape at line 435). This is why doc 03 (weight registration) is a precondition of this stage.
  • The user can opt out. MINFER_DISABLE_MPS=1 makes metal_available() false (src/metal/runtime.rs), so the priority question falls straight through to CPU.

And one design rule makes the assignment trustworthy: it is a build-time decision, full stop. Three concrete consequences, all visible in the code:

  1. No mid-run fallback, ever. Once a node is assigned, its buffer lives in that backend's pool and its kernel is that backend's kernel. If a kernel invariant is violated at execution (a shape mismatch, an unregistered weight, a device-limit problem), execute_node returns Err and the run aborts — docs/GPU_SAFETY.md rule 1: "Kernel-invariant violations return Err from execute_node — never a silent CPU fallback." The CPU backend even has an explicit arm that turns "I have no kernel for this fused op" into a loud error (cpu_backend.rs:433-441, quoted in §3.4). A quiet fallback would mask the bug the guard exists to catch.
  2. Splits are deterministic. The scheduler partitions the graph into splits — maximal runs of consecutive nodes on the same backend (doc 08) — and copies data across backends only at split boundaries. Because every node's backend is fixed before execution, the split map and the copy set are pure functions of the graph. Nothing about them can wobble run to run.
  3. Participation is recorded in the reuse identity. CParams.gpu stores "did a GPU take part in this graph?" Because graph reuse is params-only (equal GraphParams ⇒ rebuild skipped), the flag makes a backend toggle — MPS becoming available mid-session, or the env var flipping between runs — force a rebuild instead of silently reusing a graph whose backends are wrong. The comment on the field says exactly this (params.rs:19-21, quoted in §3.2).

2.2 What fusion is, and why elementwise fusion is a bandwidth story

Fusion means replacing a fixed pattern of small operations with one operation that computes the same result. The pattern this stage actually folds today is the FFN activation of every Qwen2/Qwen3 layer:

gate = matmul(normed, W_gate)        # [nt, nf]
up   = matmul(normed, W_up)          # [nt, nf]
g    = silu(gate)                    # [nt, nf]   SiLU(x) = x · sigmoid(x)
sw   = g * up                        # [nt, nf]   (Op::Mul)
out  = matmul(sw, W_down)            # [nt, d]

The silu-then-mul pair is the pattern Mul(Silu(x), y); the fused form is one node Op::SwiGLU(gate, up) — the same acronym the model cards use for this activation (SwiGLU = SiLU-gated GLU).

Why is this worth a pass? Look at what the two kernels do to memory. Call A the size of one intermediate buffer, A = nt · nf · 4 bytes in f32. Elementwise kernels read their inputs and write their output in full; they do ~one multiply per 4 bytes moved. That ratio (work per byte) is called arithmetic intensity, and it is so low that the memory traffic — not the math — sets the runtime. This is what "bandwidth-bound" means: you could make the arithmetic ten times faster and barely notice.

Count the traffic of the unfused pair, per layer, per forward:

kernelreadswrites
silugate: Asilu-out: A
mulsilu-out: A, up: Aout: A
total3A2A

The fused swiglu kernel reads gate + up and writes out once:

kernelreadswrites
swiglu2AA
total2AA

Fusion removes one full write plus one full read of the intermediate buffer = 2A bytes per layer, and one dispatch. Now put real numbers on it (Qwen2.5-0.5B: hidden 896, intermediate nf = 4864, 24 layers):

  • Prefill (say a 512-token prompt, nt = 512): A = 512 × 4864 × 4 B ≈ 9.96 MB, so 2A ≈ 19.9 MB saved per layer → ≈ 478 MB of memory traffic avoided in one prefill, plus 24 fewer dispatches. On a GPU moving tens of GB/s for scattered elementwise work, that is real time.
  • Decode (nt = 1): A = 19.4 KB, so 2A ≈ 39 KB per layer → under 1 MB across the model — bandwidth is irrelevant here. The decode win is the dispatch: every Metal pass costs a host-side encode (MINFER_OP_PROFILE=1 prints this cost per op; metal_backend.rs:253-271), and at one token per forward there is nothing else to hide it behind.

That decode asymmetry is why minfer has two fusion mechanisms, and keeping them apart is the clearest way to understand this stage:

  • The FusionPass (this doc) folds general patterns — the SwiGLU pair — on any graph, prefill or decode, CPU or GPU. It is small, safe, and always runs.
  • The builder's decode fusions (doc 05 introduced them; this doc explains the division of labor) go much further but only for nt == 1 on GPU: Op::FusedQKV replaces 3 matmul + 3 bias + 2 rope + 2 KV-store dispatches with 2 dispatches (10 → 2, measured ~+11% decode throughput on 0.5B), and Op::FusedFFN replaces 2 matmul + silu + mul with 2 dispatches (4 → 2, ~+3%; docs/COMPUTE-GRAPH-DESIGN.md §17 rows G4/G5). Those are built as single nodes because they need special kernels with unusual shapes (a concat weight blk.{i}.attn_qkv, an in-place offset swiglu), not because a pattern matcher found them.

2.3 The rewrite, node by node

The SwiGLU rewrite on one FFN fragment. Note what the pass does not do: it does not delete the old silu node — it merely stops referencing it. The node becomes an orphan (no consumers), and the next stage (the allocator, doc 07) gives orphans no buffer, so the executor skips them (§3.4). The pass itself stays a pure op-substitution:

BEFORE (silu and mul are two nodes, two intermediate buffers)

  normed ──► MatMul(ffn_gate) ──► Silu ──┐
                                         ├──► Mul ──► MatMul(ffn_down)
  normed ──► MatMul(ffn_up) ─────────────┘

AFTER (silu folded into SwiGLU; the Silu node is orphaned)

  normed ──► MatMul(ffn_gate) ───────────┐
                                         ├──► SwiGLU ──► MatMul(ffn_down)
  normed ──► MatMul(ffn_up) ─────────────┘

  Silu out: [nt, nf] buf written+read     Silu: still in the node list, but
  Mul  out: [nt, nf] buf written          orphaned → no buffer → never runs

The matcher's exact logic (fusion.rs:48-79, read in full in §3.2):

  1. Find a node whose op is Op::Mul with exactly two sources.
  2. Check whether either source is an Op::Silu (the pattern may arrive as Mul(Silu(x), y) or Mul(y, Silu(x)) — both are handled; the other operand becomes up).
  3. Ask the assigned backend of the Mul node: supports_fused(FusedOp::SwiGLU)?
  4. If yes: rewrite the Mul node in place — op becomes Op::SwiGLU, sources become [gate, up]. Shapes are unchanged (SwiGLU's output shape equals the Mul's), so nothing downstream moves.

The plan's second pattern, RoPE(Add(x, b), pos) → Op::FusedBiasRope(base, bias, pos), is gone. It was implemented and gated on supports_fused(FusedOp::BiasRope), but no backend ever claimed that capability (CPU, Metal and CUDA each list only SwiGLU), so the rewrite was dormant by construction and no Op::FusedBiasRope node could exist. Rather than keep an unreachable arm, the rule, the op and the capability tag were removed. The bias+rope work it described is covered by the builder-built decode nodes: Metal/CUDA attn_bias_rope_store folds 3 biases + 2 ropes + 2 KV stores into one kernel — a strictly more aggressive version of the same idea. The invariant that matters is the gating (fusion can never invent an op its backend did not claim), not this particular pattern: a future backend with a standalone bias+rope kernel would re-add the rule together with its capability tag.

For contrast, the builder-built decode fusion (not this pass) collapses ten nodes into two:

DECODE (nt == 1, GPU, fuse_qkv on) — built by GraphBuilder, not by FusionPass

BEFORE:  MatMul(wq) ─ Add(bq) ─ RoPE ─┐
         MatMul(wk) ─ Add(bk) ─ RoPE ─┼─ KvcacheStore   ⇒  10 dispatches / layer
         MatMul(wv) ─ Add(bv) ────────┘
AFTER:   FusedQKV(concat matmul W_qkv) ── attn_bias_rope_store   ⇒  2 dispatches / layer

2.4 The gating story: who may fuse what

The pass never decides alone. Every rewrite is gated by the target backend's own answer to supports_fused(&FusedOp), and the three capability tables currently read:

backendsupports_fused claimsconsequence for the pass
CPUSwiGLUsilu+mul folds on CPU too — one node, executed as one single-pass vector kernel (§3.4)
MetalSwiGLUsilu+mul folds (the kernel is swiglu_f32)
CUDASwiGLUsilu+mul folds

Why does the decoder-side FusedQKV/FusedFFN not appear in this table, even though they are fused ops? Because they are not produced by this pass at all. The division of labor:

  • The builder decides topology: when CParams.fuse_qkv / fuse_ffn are on (decode, GPU, concat weights available, nf ≤ 16384 for FFN), it emits Op::FusedQKV / Op::FusedFFN nodes directly (models/qwen2/graph.rs:93-105 and 236-262). When they are off it emits the decomposed chain — and the FFN branch deliberately builds silu + mul rather than a swiglu node, with the in-code comment "built as silu+mul so the fusion pass folds it" (line 228). One builder, two shapes, one downstream pass.
  • The FusionPass decides local rewrites on whatever graph it is handed — model-built or test-built — using only pattern + capability.

This split is what makes double fusion impossible by construction. Double fusion would mean fusing an already-fused node again — e.g. wrapping FusedQKV in another pattern. It cannot happen here, for three independent reasons:

  1. The pass's patterns match only decomposed ops (Mul, Silu). Fused nodes (SwiGLU, FusedQKV, FusedFFN, …) are not triggers, so a rewritten graph matches nothing the second time around — the pass is idempotent.
  2. The pass runs exactly once per build, at a single call site right after assign_backends (§3.2). There is no loop that could re-run it on its own output.
  3. Fused nodes produced by the builder contain no sub-nodes to fold — FusedQKV is one node whose "insides" live inside a Metal/CUDA kernel, invisible to a graph pattern matcher.

And the reason gating matters at all: a fusion applied to a backend without the kernel would produce an op no engine can execute — the graph would abort at run time (loudly, per the convention above, but still uselessly). Gating at the source means the graph only ever contains ops its assigned backend claimed.

2.5 The env toggles: fusion as a first-class A/B experiment

Both builder fusions are controlled by environment variables read at graph-build time (models/qwen2/graph.rs:468-473):

MINFER_NO_FUSE_QKV=1  →  CParams.fuse_qkv = false  →  builder emits the decomposed QKV chain
MINFER_NO_FUSE_FFN=1  →  CParams.fuse_ffn = false  →  builder emits matmul + silu + mul + matmul

fuse_qkv/fuse_ffn are fields of CParams, which is part of GraphParams, which is the entire input to graph reuse. Flip either env var and three things follow, mechanically:

  1. The reuse check try_reuse fails (params differ) → the graph is rebuilt with the new topology. Toggling fusion can never leave a stale fused graph in place.
  2. The rebuilt graph has different nodes (fused vs decomposed) — the two graphs are a valid A/B pair, and the experiment is cheap: no process restart, no re-loading of weights.
  3. The comparison is bit-identical, not approximately equal. Fused kernels were written to produce exactly the same bits as the decomposed chain; the G4/G5 records measured fused-vs-unfused logit differences of 0.000 on 0.5B and 7B (COMPUTE-GRAPH-DESIGN.md §17 G4/G5), and the test fused_qkv_matches_unfused_decode asserts it (§4).

That bit-identity has one famous footnote — the ~1e-6 noise lesson (plan doc deviation 26): during bring-up, a test ran the unfused graph without the FusionPass, so silu and mul executed as two separate kernels; the result differed from the fused path by ~1e-6 relative noise (two float roundings vs one, amplified across large intermediate values) and the bit-identity assertion "failed" for a reason that had nothing to do with the fused kernels. The resolution is a rule, not a workaround: the real forward path always runs FusionPass — the only difference between the A and B sides is which nodes the builder emitted, never whether the pass ran. An unfused graph that skipped the pass would not be "more primitive"; it would be wrong as a reference.

3. Implementation

3.1 Data in / data out

In: the freshly built ComputeGraph (doc 05) — every node carries op (with full payloads), src (input node ids), out_shape/out_dtype, and backend: Option<Backend> where None means "undecided" (Backend is a re-export of the registry handle — src/graph/mod.rs:105). Plus the registry (src/graph/registry.rs): which backends this build contains, their names, capability records and priority order — and, per allocator, which of them are enabled (Metal on macOS when available and all weights registered; CUDA when the feature is on and a device exists).

Out: the same graph object, mutated in place:

  • every node has backend: Some(…),
  • some Mul nodes are now Op::SwiGLU nodes with re-routed src = [gate, up],
  • nothing else: no node is added or removed, no shape changes, build order is untouched. The orphaned Silu nodes are still in the list — they only die at allocation time.

The subsequent stages consume exactly this: alloc_graph (doc 07) uses the backends to pick pools and gives orphans nothing; split_graph/execute (doc 08) uses the backends to cut splits and dispatch.

Where the stage physically runs: inside the model's forward_cached — the first place a forward needs a graph. The sequence there is try_reuse → build → register weights → enable backends → assign_backends → FusionPass → alloc_graph → replace_graph (src/models/qwen2/graph.rs:611-650). Note the pipeline lives in model code, not the scheduler: the scheduler provides assign_backends, but the model orchestrates, because only the model knows whether its weights made it onto the GPU.

3.2 Key code

The whole assignment policy — src/graph/scheduler.rs:68-83:

#![allow(unused)]
fn main() {
/// Assign every node to the best backend that supports it (capability
/// driven via the allocator's backend registry).
pub fn assign_backends(&self, graph: &mut ComputeGraph, alloc: &GraphAllocator) {
    for node in &mut graph.nodes {
        if node.backend.is_some() {
            continue; // keep explicit assignments
        }
        node.backend = alloc
            .supports_for(&node.op, node.out_dtype, node.layer)
            .or(Some(BackendTag::CPU));
    }
}
}

Three lines of policy, each deliberate. The loop is order-independent — each node is asked on its own, so the result cannot depend on graph traversal order. The continue keeps any explicit assignment (tests use it; the production builder assigns none). The .or(Some(CPU)) fallback makes the assignment total: every node leaves with a backend even if supports_for returned None. That cannot hide an unsupported op — the CPU backend's execute_node has an explicit Err arm for ops it has no kernel for (§3.4) — it just guarantees the abort names the right node instead of crashing on an Option unwrap.

Where the priority actually lives — src/graph/alloc.rs:403-433, trimmed:

#![allow(unused)]
fn main() {
/// Which backend may take a node of `(op, dtype)` that belongs to `layer` — the
/// assignment rule with the E5 offload policy and F4's backend fence applied.
pub fn supports_for(&self, op: &Op, dtype: DType, layer: Option<usize>) -> Option<Backend> {
    let eligible = |b: &dyn BackendTrait| super::backend_takes(b, op, dtype);
    let device_ok = /* E5: is this node's block on the device? */;
    if device_ok {
        // F4: the registry's **priority** order, not a statement order here.
        for &backend in super::registry::registry().by_priority() {
            if backend == Backend::CPU {
                continue; // the fallback, tried below even when device_ok is false
            }
            if !self.filter.allows(backend) {
                continue; // `--backend` / `MINFER_BACKENDS` fenced it off
            }
            if let Some(pool) = self.pool(backend) {
                if eligible(pool) {
                    return Some(backend);
                }
            }
        }
    }
    if self.filter.allows(Backend::CPU) && eligible(&self.cpu) {
        return Some(Backend::CPU);
    }
    None
}
}

The order is data, not statement order. Backend is a Copy handle into a fixed id space (src/graph/registry.rs:85-87: CPU = 0, METAL = 1, CUDA = 2) and the registry holds the entries in priority order — Metal 300, CUDA 200, CPU 100 (PRIORITY_METAL / PRIORITY_CUDA / PRIORITY_CPU, src/graph/registry.rs:72-75). by_priority() is that list, so the assignment preference is one table rather than the order of if blocks in this function.

enable_metal() / enable_cuda() still populate the allocator's per-backend pools (model code calls them only when its own weight gates passed, §2.1); a backend with no pool is skipped, and self.filter is the F4 fence. A backend this build does not contain is not in the registry at all, so the old #[cfg]-gated questions are gone — but the order and the answer are identical, which is what the registry's own pinning test (graph::registry::tests::the_registered_set_and_priority_order_are_pinned) asserts. The contract — the two orders, the name surface and the three startup refusals — is docs/BACKEND-REGISTRY-DESIGN.md; this doc only shows how the assignment reads it.

The contract every backend implements — src/graph/backend.rs:21-191. The full method list is docs/ARCHITECTURE.md §5.1, and the contract around it is docs/BACKEND-REGISTRY-DESIGN.md:

#![allow(unused)]
fn main() {
pub trait Backend: Send + Sync {
    /// Op support by (op, dtype). `supports_fused` gates the fusion pass
    /// (Phase 4) so fused IR nodes are only produced when a kernel exists.
    fn supports_op(&self, op: &Op, dtype: DType) -> bool;
    fn supports_fused(&self, fused: &FusedOp) -> bool;
    fn supports_attn_span(&self) -> bool { false }   // E1: explicit attention windows
    fn weights_bytes(&self) -> usize { 0 }           // E4: what the memory budget is charged
    // ... buffer pool, copy_cells, execute_node, host read/write, synchronize, retire
}
}

The trait has no name(): since F4 a backend's identity, id, priority and names live on the registry's handle and entry (src/graph/registry.rs), which is what lets a backend be requested by name at all.

A backend's self-description, CPU — src/graph/cpu_backend.rs:73-105:

#![allow(unused)]
fn main() {
fn supports_op(&self, op: &Op, dtype: DType) -> bool {
    if dtype != DType::F32 {
        return false;
    }
    matches!(
        op,
        Op::Input
            | Op::Add
            | Op::Mul
            | Op::Scale(_)
            | Op::Silu
            | Op::Softmax { .. }
            | Op::RmsNorm { .. }
            | Op::QkNorm { .. }
            | Op::MatMul { .. }
            | Op::GetRows
            | Op::RoPE { .. }
            | Op::Attn { .. }
            | Op::KvcacheStore { .. }
            | Op::KvcacheLoad { .. }
            | Op::SwiGLU
            | Op::View { .. }
            | Op::Reshape { .. }
            | Op::Permute { .. }
    )
}

fn supports_fused(&self, fused: &FusedOp) -> bool {
    // The SwiGLU rewrite IS applied to CPU nodes (CPU is first in the
    // fusion pass's backend list); `Op::SwiGLU` below executes it as a
    // single pass. bias+rope and batch-matmul are not fused (batch QKV
    // quantize-sharing is a Phase 5+ win).
    matches!(fused, FusedOp::SwiGLU)
}
}

The CPU claims the whole per-layer vocabulary — everything is F32-only, and the matches! list is the CPU's capability table. Note Op::SwiGLU is claimed but FusedQKV/FusedFFN are not: decode fusions are GPU-only (§2.3). The supports_fused answer is true for SwiGLU, so the pass does fold on CPU; what "fused" means there is a single-pass kernel, shown next.

What a "fused" SwiGLU means on CPU — src/graph/cpu_backend.rs:373-379:

#![allow(unused)]
fn main() {
Op::SwiGLU => {
    // silu(gate) * up, one pass: bit-identical to the old
    // vec_silu_f32 + vec_mul_f32 pair (same formula and
    // per-element order) without its full-size temp buffer.
    crate::vec_ops::vec_swiglu_f32(out.len(), out, ins[0], ins[1]);
    Ok(())
}
}

One node, one pass (vec_ops::vec_swiglu_f32), no intermediate buffer. The earlier form ran vec_silu_f32 then vec_mul_f32 through a full-size scratch copy (out.to_vec()); the fused version computes silu(gate[i]) * up[i] in one loop and is bit-identical to that pair (same formula, same per-element order — pinned by vec_ops::tests::swiglu_matches_silu_then_mul).

Metal's table, with the decode fusions and the CUDA-only epilogue negative — src/graph/metal_backend.rs:294-319:

#![allow(unused)]
fn main() {
fn supports_op(&self, op: &Op, dtype: DType) -> bool {
    match op {
        Op::Input => true,
        Op::Add | Op::Mul | Op::Silu | Op::RmsNorm { .. } | Op::QkNorm { .. } | Op::SwiGLU => {
            dtype == DType::F32
        }
        Op::MatMul { .. } => {
            matches!(dtype, DType::F32) // activations are f32; weight type in meta
        }
        Op::GetRows | Op::RoPE { .. } | Op::Attn { .. } => dtype == DType::F32,
        Op::KvcacheStore { .. } | Op::KvcacheLoad { .. } => dtype == DType::F32,
        Op::FusedQKV { .. } | Op::FusedQkvNorm { .. } | Op::FusedFFN => dtype == DType::F32,
        Op::View { .. } | Op::Reshape { .. } | Op::Permute { .. } => true,
        Op::Scale(_) | Op::Softmax { .. } | Op::BatchMatMul => false,
        // Mixed-quant decode QKV epilogue (D3-8 class 2) is CUDA-only; on
        // Metal the graph builder never emits it (qkv_epilogue_ok = false
        // without `--features cuda`), so it is never assigned here.
        Op::QkvBiasRopeStore { .. } => false,
    }
}

fn supports_fused(&self, fused: &FusedOp) -> bool {
    // swiglu_f32 is the only fusion-pass kernel. The bias+rope+store
    // capability is a build-time fused node (FusedQKV/FusedQkvNorm), not a
    // FusionPass target, so it is not advertised here.
    matches!(fused, FusedOp::SwiGLU)
}
}

The negative arms are design statements, not gaps: Op::QkvBiasRopeStore => false documents that the mixed-quant decode epilogue belongs to CUDA only, and Op::BatchMatMul => false is a deferred vocabulary entry. Metal does claim the builder's decode fusions (FusedQKV, FusedQkvNorm — the Qwen3 per-head-norm variant — and FusedFFN), which is what makes the builder's fuse_qkv gate safe.

CUDA's table differs where its kernels differ — src/graph/cuda_backend.rs:supports_op, trimmed to the interesting arms:

#![allow(unused)]
fn main() {
/// Capability matrix (docs/CUDA-BACKEND-DESIGN.md §4.3): the full per-layer
/// chain runs on CUDA, including the embedding/tail gather (7e③) and the
/// decode fusions FusedQKV/QkvBiasRopeStore/FusedFFN. Scale, Softmax,
/// BatchMatMul and the Qwen3-only FusedQkvNorm have no kernels and stay on
/// the CPU backend; weight-quant eligibility is the model-level
/// all-weights-registered gate. RoPE is gated to the neox
/// (non-interleaved) layout — the only style the supported architectures
/// emit.
fn supports_op(&self, op: &Op, dtype: DType) -> bool {
    if dtype != DType::F32 {
        return false;
    }
    match op {
        Op::Input | Op::Add | Op::Mul | Op::Silu | Op::SwiGLU | /* … */
        Op::MatMul { .. } | Op::Attn { .. }
        | Op::KvcacheStore { .. } | Op::KvcacheLoad { .. } | /* … */
        Op::GetRows
        | Op::FusedQKV { .. }
        | Op::QkvBiasRopeStore { .. }
        | Op::FusedFFN => true,
        Op::RoPE { style } => matches!(style, RopeStyle::NonInterleaved),
        _ => false,
    }
}
}

The instructive line is RoPE: capability can be payload-conditional — CUDA rotates only the NonInterleaved style, so a hypothetical interleaved-RoPE node would silently route to CPU instead of producing wrong numbers. This is supports_op earning its keep as a per-node query rather than a per-backend yes/no.

The fusion pass, part 1 — dispatch — src/graph/fusion.rs:20-35:

#![allow(unused)]
fn main() {
/// Run all supported fusions over the graph. Returns the number of nodes
/// rewritten. `backend_for` returns the backend a node is assigned to
/// (used to gate fusions per backend capability).
pub fn run(
    &self,
    graph: &mut ComputeGraph,
    backends: &[&dyn Backend],
    backend_of: &dyn Fn(&ComputeGraph, usize) -> Option<usize>, // node id -> backend index
) -> usize {
    let mut n = 0;
    n += self.fuse_swiglu(graph, backends, backend_of);
    n
}
}

The pass takes the backends as trait objects plus a closure mapping node id → index into that slice. The indirection exists because backends are stored as Option fields on the allocator in cfg-conditional order; the closure is built at the call site where that layout is known. The None case of the closure means "unassigned" — and unassigned nodes are never fused (the gate treats None as false, §2.4's promise again).

The fusion pass, part 2 — the SwiGLU matcher — src/graph/fusion.rs:48-79 (inside fuse_swiglu):

#![allow(unused)]
fn main() {
for id in 0..n {
    if !matches!(graph.node(id).op, Op::Mul) {
        continue;
    }
    let mul = graph.node(id);
    if mul.src.len() != 2 {
        continue;
    }
    let (s, y) = (mul.src[0], mul.src[1]);
    let is_silu = |x: usize| matches!(graph.node(x).op, Op::Silu);
    let (silu_in, gate, up) = if is_silu(s) {
        (graph.node(s).src[0], graph.node(s).src[0], y)
    } else if is_silu(y) {
        (graph.node(y).src[0], graph.node(y).src[0], s)
    } else {
        continue;
    };
    let _ = silu_in;
    // gate the fusion on the mul node's backend capability
    let ok = match backend_of(graph, id) {
        Some(bi) => backends[bi].supports_fused(&FusedOp::SwiGLU),
        None => false,
    };
    if !ok {
        continue;
    }
    // replace Mul with SwiGLU(gate, up)
    new_ops[id] = Some(Op::SwiGLU);
    if let Some(node) = graph.nodes.get_mut(id) {
        node.src = vec![gate, up];
    }
    replaced += 1;
}
}

Mechanics worth noticing: the rewrite is collected into a new_ops staging vector and applied after the scan, so the matcher never walks a half-mutated graph. The rewrite keeps the Mul's node id and the Mul's backend — consumers of the Mul (the down matmul) keep pointing at the same id and need no edits. The new src = [gate, up] re-routes dataflow: gate is the Silu node's input (skipping the dead Silu — the silu_in binding holds the same value, which is why the source discards it with let _ =), up is the other operand. Shapes never change because SwiGLU's output shape equals the Mul's by definition.

The second matcher is gone. The old fuse_bias_rope (RoPE(Add(x, b)) → Op::FusedBiasRope) and its FusedOp::BiasRope capability tag were removed (§2.3): no backend claimed the capability, so the matcher was unreachable code. run() now composes exactly one rewrite.

The call site that wires it all together — src/models/qwen2/graph.rs:611-650, trimmed:

#![allow(unused)]
fn main() {
if !cache.try_reuse(&params) {
    let mut graph = Self::build(model, &params);
    let sched = BackendScheduler::new();
    {
        let alloc = cache.alloc();
        Self::register_graph_weights(model, alloc);
        #[cfg(target_os = "macos")]
        if metal_on { alloc.enable_metal(); }
        #[cfg(feature = "cuda")]
        if cuda_on { alloc.enable_cuda(); }
        sched.assign_backends(&mut graph, alloc);
        // fusion pass gated per node's assigned backend.
        //
        // F4: the backend list and the node → index map come from the allocator's
        // registry view (the enabled entries in identity order), replacing the
        // hand-built vector and its `name() == "cuda"` position lookup.
        let backends: Vec<&dyn Backend> = alloc.fusion_backends();
        FusionPass::new().run(&mut graph, &backends, &|g, id| {
            g.node(id)
                .backend
                .and_then(|b| alloc.fusion_backend_index(b))
        });
        alloc.alloc_graph(&graph).unwrap();
    }
    cache.replace_graph(graph, params);
}
}

This is the stage's place in the pipeline, verbatim: reuse check → build → enable → assign → fuse → allocate → store. The whole block is skipped when try_reuse succeeds — assignment and fusion run once per GraphParams, which is why they can afford to be thorough. (Qwen3Graph mirrors this block at src/models/qwen3/graph.rs:521-548.)

The reuse identity that makes toggles work — src/graph/params.rs:17-35:

#![allow(unused)]
fn main() {
/// Runtime parameters that affect graph construction.
///
/// `gpu` records whether the GPU backend participates — the backend assignment
/// is part of the built graph, so a change (e.g. MPS init between runs) must
/// force a rebuild.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct CParams {
    pub n_ctx: usize,
    pub flash_attn: bool,
    pub gpu: bool,
    /// G4 decode QKV fusion enabled (part of the topology: toggling
    /// `MINFER_NO_FUSE_QKV` must force a rebuild).
    pub fuse_qkv: bool,
    /// G5 decode FFN gate+up fusion enabled (part of the topology: toggling
    /// `MINFER_NO_FUSE_FFN` must force a rebuild). Decoupled from `fuse_qkv`
    /// so A/B-ing one fusion does not flip the other.
    pub fuse_ffn: bool,
}
}

and the comparison that consumes it — src/graph/cache.rs:47-64:

#![allow(unused)]
fn main() {
pub fn try_reuse(&mut self, params: &GraphParams) -> bool {
    match (&self.prev_params, &self.graph) {
        (Some(prev), Some(_)) if Self::params_match(prev, params) => {
            self.prev_params = Some(params.clone());
            true
        }
        _ => false,
    }
}

fn params_match(a: &GraphParams, b: &GraphParams) -> bool {
    a.n_tokens == b.n_tokens
        && a.n_out == b.n_out
        && a.gtype == b.gtype
        && a.cparams == b.cparams
        && a.weights_version == b.weights_version
}
}

a.cparams == b.cparams is derived PartialEq over all six fields — gpu, fuse_qkv, fuse_ffn included. That one derived impl is why every toggle in this doc is a rebuild instead of a silent stale-graph reuse. The dedicated test fuse_flags_are_part_of_the_reuse_identity (cache/tests.rs:36) flips each flag and asserts non-reuse.

Where fusion orphans die — src/graph/scheduler.rs:225-233 (doc 08's executor, forward pointer):

#![allow(unused)]
fn main() {
if node.is_input() {
    continue; // data pre-filled by the allocator
}
// dead nodes (no consumers, not outputs) get no buffer — the
// fusion pass can orphan them (e.g. silu folded into SwiGLU);
// they are skipped, not executed
let Some(br) = alloc.node_buffer(id) else {
    continue;
};
}

The comment is the contract between this stage and the next two: the fusion pass is allowed to leave garbage nodes in the graph because the allocator (doc 07) refuses to give dead nodes buffers and the executor (doc 08) skips bufferless nodes. Cleaning them out of the node list would mean renumbering every node id — the riskier operation by far.

3.3 Design choices (why this shape and not another)

Q1: Why is supports_op a per-backend query instead of a global capability table?

The obvious alternative is a central table — SUPPORTS: [(Op, Backend, bool); N] — that the scheduler reads. It looks simpler and it rots faster: every new backend must be threaded into the table, every new op must be threaded into the table, and the two edits happen in different files with no compiler help linking them. The trait inverts that: a capability table is replaced by a type. Adding the CUDA backend meant implementing Backend (including supports_op/supports_fused) and writing its registry entry — the scheduler's assignment code and supports_for have not changed for it at all since F4 (the loop reads the registry's priority list; the CUDA enable call is cfg-gated model-side). The compiler enforces completeness: a backend that forgets to answer for an op answers false via its own match exhaustiveness (and the enum's missing-arm error), never "table out of date". It also lets capability be stateful in a legitimate way — Metal's yes depends on which kernels its GPU supports at runtime, CUDA's RoPE answer depends on the payload style — which a static table cannot express.

Q2: Why does fusion run after assignment?

Because gating needs an answer to "fused for whom?", and only assignment provides it. The pass's gate is literally backend_of(mul_node) → supports_fused(...). Fuse before assignment and the pass would have to guess a backend (fusing against Metal and CUDA "just in case") or fuse unconditionally (producing a fused node its eventual owner cannot execute — aborting the run later, or worse, tempting someone to add a fallback). Assignment-first means each node's rewrite is decided by the engine that will actually run it; the graph can never contain a fused op its owner didn't claim. There is a bonus: the order is safe in the other direction too, because every fused op the pass can produce (Op::SwiGLU) is itself in all three backends' supports_op — so had assignment somehow run again after fusion, it would be a no-op. The pipeline relies on the ordering, not on that coincidence.

Q3: Why is the fusion pass a graph→graph rewrite instead of fusing inside the kernels?

Four reasons, in descending order of weight:

  1. Kernels stay single-purpose. swiglu_f32 exists because one Metal kernel computes it; the CPU instead runs two vec passes inside one node. Both decisions are backend-internal and can change (the plan doc calls for a single-pass CPU kernel) without touching the IR, the scheduler, or any other backend. Fusion-in-kernel would instead mean every backend re-implements pattern detection over raw buffers.
  2. The IR stays inspectable. After this stage the graph still prints as nodes with names, ops and backends — which is what --dump-graph-json, the DOT export, the P2 trace and the live viz server show (§4). Debugging "why is decode slow" becomes reading a graph, not decoding backtrace soup.
  3. Verification becomes A/B on graphs. Fused vs unfused is two GraphParams and one env var, executed through the same scheduler, compared bit-for-bit (§2.5). If fusion happened opaquely inside kernels, the "unfused" reference would be unreachable — there would be nothing to compare against.
  4. Downstream stages see the truth. The allocator sizes and shares buffers for the fused graph (no buffer for the dead Silu), the splitter draws splits around the fused node, and the executor dispatches exactly what was planned. A kernel-internal fusion would make all three plan against a graph that lies.

Q4: Why are fused ops (and gpu) part of the reuse identity?

Reuse exists so decode steps don't rebuild the graph; it is safe only if equal params guarantee an identical graph. Fusion decisions are topology: a fused decode graph and an unfused one are different graphs by any structural measure. Since try_reuse compares params only (deliberately — node-by-node structural comparison was rejected as unnecessary given deterministic building, COMPUTE-GRAPH-DESIGN.md §6), the fusion switches must ride along in those params or the cache could hand back a graph that contradicts the current configuration: you'd set MINFER_NO_FUSE_QKV=1 for your A/B run and the engine would quietly reuse the fused graph — the experiment would show "no difference" and teach you nothing. Determinism and A/B-ability are the same requirement viewed from two sides: identical params ⇒ identical fusion decisions ⇒ identical topology ⇒ reuse is sound and toggles are observable.

Two smaller choices worth naming:

  • Why is priority hardcoded as if-block order instead of a sorted backend list? — F4 reversed this. The ordered list is now the registry's priority field (PRIORITY_METAL 300 / PRIORITY_CUDA 200 / PRIORITY_CPU 100, src/graph/registry.rs:72-75), for the reason the old argument missed: once Backend is a handle rather than an enum variant, the order has more than one reader — the assignment pass, the fusion pass's node → index map, and the exporters — so one table is what keeps them from disagreeing. The if-chain was readable; it was also the beginning of three copies. The contract is docs/BACKEND-REGISTRY-DESIGN.md.
  • Why does the pass keep orphans instead of removing them? Removal renumbers node ids, which invalidates every stored src, every output id, the trace bookkeeping — all to save a skipped iteration at execution. Leaving them costs one continue per orphan (§3.2, last excerpt) and zero risk.

3.4 Pitfalls & invariants

  • Assignment is total and final. Every node exits this stage with a backend, and no stage after this may change it. Execution errors are Err + abort: the CPU backend's arm for ops it cannot run says so in its error text — cpu_backend.rs:433-441:

    #![allow(unused)]
    fn main() {
    Op::BatchMatMul
    | Op::FusedQKV { .. }
    | Op::QkvBiasRopeStore { .. }
    | Op::FusedQkvNorm { .. }
    | Op::FusedFFN => Err(format!(
        "op {:?} unsupported on CPU (fusion not enabled for it)",
        node.op
    )),
    }

    ("fusion not enabled for it" — i.e. if you ever see this error, a fused node reached a backend that never claimed it: a gating bug, not a runtime condition. docs/GPU_SAFETY.md rule 1 is the convention; this arm is the CPU-side enforcement.)

  • The fusion pass must run — on every path, including the "unfused" one. The deviation-26 lesson (§2.5): a graph that skips FusionPass is not a valid reference for bit-identity experiments, because two-kernel silu+mul differs from one-kernel swiglu by ~1e-6 float noise. That is why the JSON export path (json.rs::build_runtime_graph, lines 50-99) re-runs assign + FusionPass — so a dumped graph matches what actually executed — and why the test helper in fused_qkv_matches_unfused_decode runs the pass on both sides.

  • Double-fusion is impossible by construction (§2.4): patterns match only decomposed ops, the pass is idempotent, it runs once per build, and builder-fused nodes have no graph-visible insides. If you add a third fusion pattern, preserve all four properties.

  • The pass has exactly one rule. RoPE(Add) → FusedBiasRope was removed (§2.3) because no backend ever claimed the capability. If a future backend gains a standalone bias+rope kernel, re-add the rule together with its FusedOp tag and a test that pins the negative case — the gating design is the invariant, and the removed rule is the worked example of what not to advertise.

  • Fusion never changes shapes or node count. Orphans stay, ids stay, shapes stay. Anything that does change the node set (decode fusions) happens in the builder, where shapes are computed with full context (FusedQKV's [nqt+2·nkt, 1] output, FusedFFN's [2·nf, 1]). Keep that boundary: pattern pass = local substitution, builder = structural change.

  • GPU feasibility is a precondition, not a per-node property. supports_op(Metal) says nothing about whether the weights are resident; that is the model-level all-or-nothing gate (weights_on_gpu, qwen2/graph.rs:813) feeding metal_on, feeding enable_metal(). A backend enabled without its weights would abort at the first matmul with "weight not registered" — loud, but avoidable.

  • In-place aliasing interacts with fusion only indirectly (full story in docs 07/08): the fused SwiGLU reads gate and up as separate inputs on GPU, while Op::Silu and Op::RoPE are the ops with in-place aliasing rules. Fusion removes Silu nodes (orphaning them), which slightly reduces the number of aliased buffers — a quiet benefit for the allocator.

4. Observe & verify

  • MINFER_GRAPH_TRACE=1 prints the split map and a per-op/per-backend node census at execute time (scheduler.rs:127-144), e.g. op SwiGLU backend Metal x24 — the direct output of this stage: 24 folded nodes, one per layer, on their assigned backend.
  • MINFER_NO_FUSE_QKV=1 / MINFER_NO_FUSE_FFN=1 switch the builder's decode fusions off for one run; combined with the trace (or the decode throughput line) they are the A/B switch, and the rebuild they force is the mechanism §2.5 describes.
  • --dump-graph-json <file> / --dump-graph <dot> rebuild the graph through json.rs::build_runtime_graph — explicitly "build → assign → FusionPass" (json.rs:50-68) — so the exported node list shows the fused ops ("swiglu", "fused_qkv") and each node's backend, i.e. a picture of this stage's output. MINFER_TRACE / the viz server show the same graph live, with per-node data.
  • MINFER_OP_PROFILE=1 (Metal) prints host-encode time per op after the run (metal_backend.rs:36-39, 224-242) — the dispatch cost that fusion removes, made visible.
  • MINFER_DISABLE_MPS=1 forces metal_on = false and shows the priority chain collapsing to CPU (everything in the trace census becomes backend CPU).
  • Tests (all cargo test): fusion.rs — swiglu_fusion_applies_when_backend_supports (rewrite happens, src == [gate, up]), swiglu_fusion_skipped_when_backend_does_not_support (unassigned node ⇒ no fusion); vec_ops::tests::swiglu_matches_silu_then_mul (the CPU single-pass kernel is bit-identical to the old silu+mul pair); cache.rs::fuse_flags_are_part_of_the_reuse_identity (flag flip ⇒ no reuse); qwen2/graph.rs::fused_qkv_matches_unfused_decode (fused nodes present iff gated on, logits bit-identical, plus per-layer output comparison); the tail-reduction test asserts the SwiGLU node exists and consumes the gate/up matmuls (qwen2/graph/tail_tests.rs:111-146) — the pass is exercised on every graph-level test run, not just the fusion unit tests.

5. Cross-references

  • docs/ARCHITECTURE.md §4.3 (the assign → fuse → alloc → execute pipeline), §4.5 (invariants 4–5: aliasing, dead nodes), §5.1–5.3 (Backend trait, selection rules, GPU safety) — the condensed version of this doc.
  • docs/COMPUTE-GRAPH-DESIGN.md §5 (fusion rules incl. the deferred BatchMatMul), §7 (Metal per-op mapping — where Op::SwiGLU's kernel comes from), §17 rows G4/G5 (decode fusion measurements: 10→2 and 4→2 dispatches, +11%/+3% on 0.5B, fused-vs-unfused diff 0.000) and deviation 26 (the ~1e-6 FusionPass-must-run lesson).
  • docs/GPU_SAFETY.md — the Err-not-fallback convention this stage's build-time assignment exists to honor.
  • Doc 05 (builder: where decomposed vs fused topologies come from) · doc 07 (allocator: what happens to fusion orphans) · doc 08 (scheduler: splits and cross-backend copies that assignment makes deterministic) · doc 13 (decode loop: why assign+fuse run once, not per token) · doc 14/15 (Metal/CUDA: the kernels behind the capability tables).

← 05 — Graph build: the IR and the builder · Index · 07 — Memory allocation: liveness and the KV regions →

07 · The allocator: liveness, buffer reuse, and the KV cache

Stage: docs 05–06 built a compute graph and gave every node a backend → this stage decides WHERE every intermediate tensor lives (and owns the KV cache) → doc 08's scheduler then executes the graph through those buffers. Code: src/graph/alloc.rs (GraphAllocator::alloc_graph, fill_input_i32, ensure_kv, copy_across), src/graph/cache.rs (GraphCache), src/graph/backend.rs (Backend trait: pool + alloc_fresh, KvProvider), src/models/qwen2/graph.rs::forward_cached (input filling call site).

1. Background — where this stage sits

By the end of doc 06, the engine holds a compute graph — a pure data structure that lists every math operation of one transformer forward pass as CNodes ("compute nodes"): a RmsNorm node, three projection MatMul nodes, a RoPE node, a fused SwiGLU node, and so on. Each node already knows which backend will run it (CPU, Metal, or CUDA) and how big its output is (out_shape, e.g. [896, 1, 1, 1] for one token's hidden state on Qwen2.5-0.5B). What no node has is a place to put its result. The nodes are pure description: "the silu of node 17". Silu of what memory?

This stage answers that. The allocator (GraphAllocator in src/graph/alloc.rs) walks the graph once and hands every node a buffer — a region of memory inside a backend's pool, addressed by a small handle (BufRef { backend, id }). It also does two jobs that are easy to overlook but are just as load-bearing:

  • it owns the KV cache: the per-layer scratchpad that attention reads and writes (defined properly in §2.5), and
  • it fills the input buffers: the token ids and positions your prompt was turned into in doc 04 are written into pool buffers here, before any kernel runs.

Why not just give every node its own fresh buffer and be done? Arithmetic makes the naive version ugly fast. The 0.5B decode graph has 437 nodes (recorded in docs/COMPUTE-GRAPH-DESIGN.md Phase 8), so naive allocation means 437 separate memory regions per backend — and for the GPU that is 437 driver buffer objects to create, register, and keep alive. Worse, the two persistent things (KV regions) must be sized once and survive; a throwaway allocate-per-node scheme has no place to put them.

The saving observation is old and simple: a buffer's contents only matter between the moment they are written and the moment they are last read. Node 17 (hello again, silu) writes its output; some later node reads it once; after that the memory is dead weight that the next operation could reuse. Bookkeep those windows — the live ranges — and buffers can be shared by operations that never overlap in time. This is exactly what llama.cpp's ggml_gallocr does, and alloc.rs says so in its first line: *"Mirrors llama.cpp's ggml_gallocr: buffers are shared between nodes whose live ranges do not overlap" (src/graph/alloc.rs:1-10, the module doc).

But sharing memory is also where the two most instructive bugs of this codebase happened: one where a copy read data the GPU had not produced yet (the whole KV cache silently became zeros), and one where liveness was computed in a different order than execution, so a buffer was recycled while its reader was still waiting (logits off by 21.79). Both bugs, and the invariants that now prevent them, are told in §3.4 — because "why is reuse safe now" is the single best question you can ask about this stage.

2. Principle — how it works and why

2.1 Buffers, pools, and handles

Three terms, defined once and used everywhere after:

  • A buffer is a contiguous region of memory holding size f32 numbers (4 bytes each). On the CPU backend a buffer is literally a Vec<f32> CpuBackend (src/graph/cpu_backend.rs:20-23, the buffers pool); on Metal it is an MTLBuffer the CPU and GPU can both see; on CUDA it is device memory.
  • A pool is the backend's list of all its buffers, plus a free list of ids that are currently unused. Allocating means "find me a buffer of this size" — from the free list if one fits, otherwise create one.
  • A handle (BufRef) is just { backend, id } — which pool, which slot. Nobody outside the backend ever touches the memory through the id directly; the backend resolves id → &mut [f32] (read_host/write_host) or passes the id to its own kernels.

One deliberate simplification shapes everything: every pool buffer is f32-typed. The allocator counts sizes in f32 elements (Backend::alloc_buffer "allocate / release a buffer of size f32 elements", supports_fused (backend.rs:25), Metal sizes buffers as size * 4 bytes in alloc_buffer (metal_backend.rs:897), and weights keep their quantized bytes elsewhere (registered by name in doc 03). One dtype means one allocator, one copy path, one set of host-access functions — and, as §2.6 shows, even integers ride along as f32 bit patterns.

The last pool property to internalize: ids are stable. Once buffer #7 exists, it is buffer #7 until the whole graph is torn down; a recycled id keeps its size; nothing ever moves. That stability is what lets a GPU record raw pointers into a captured kernel launch (CUDA Graph replay, doc 15) and replay them later — the backend trait says it outright: "Implementations must keep captured pointers stable (pool ids never move memory)" execute_node (backend.rs:80).

2.2 Liveness: when a buffer's contents are precious

A node's output is live from the moment the node executes (first write) until the moment its last consumer has executed (last read). That window is its live range. Liveness analysis is just computing everyone's window.

The windows come from the graph structure, not from a clock. If node h is consumed by nodes at execution positions 40 and 240, h's live range is [40, 240] — its buffer must hold h's value through position 240 and not a step longer. Two nodes whose ranges never overlap can safely share one buffer: whoever comes first writes, its readers finish, and only then does the second writer overwrite. A chain of ops shares beautifully:

exec position:   0     1      2      3      4      5
node:          input  silu   add    silu   add    output
live range:    [0,5]  [1,2]  [2,3]  [3,4]  [4,5]  [5,∞)

buffers in use at any moment: 2 (input + "current value")
naive:                        6 buffers

Every result only feeds the next op, so one scratch buffer ping-pongs with the input buffer. Contrast a parallel shape — a2 = add(silu(a0), a0) and b2 = add(silu(b0), b0) computed independently — where both branches are live simultaneously and must get separate buffers. minfer's unit tests assert exactly these two behaviors: liveness_reuses_buffers_along_chain (fewer buffers than nodes) and parallel_chains_do_not_share (alloc/tests/liveness.rs:24).

Two bookkeeping rules extend the basic window, and both exist because of execution-order realities rather than graph theory:

  1. Graph outputs live forever (well, until the scheduler copies them out). Logits are read after the whole graph ran, so their live range ends at order.len() — one past the last node.
  2. Graph inputs live forever too. This is the subtler one, and it is a recorded bug fix (deviation 23, §3.4): inputs are filled on the host before execution starts, so a buffer that liveness would normally recycle for another input would get clobbered by the later fill. Inputs get the same "live to the end" treatment as outputs alloc_graph (alloc.rs:513-561, the pinning loop).

2.3 The alloc_graph walk: intervals, sweep, free list

alloc_graph runs once per graph (re)build, after doc 06's assign + fusion passes. The walk is a single forward pass with a running clock:

  1. Tear down the previous graph's liveness buffers — alloc_graph (alloc.rs:513-519): every id tracked in buf_alive goes back to its pool's free list; the node_to_buf map is cleared. Persistent regions (§2.5) are not in buf_alive, so they sail through untouched.
  2. Order check: call topo_order() — but only to validate that the graph is acyclic. The order actually used is plain build order, 0..n_nodes (alloc.rs:536-537, inside alloc_graph: the acyclicity check and the build-order vector). §3.4 explains why this one line is the tombstone of the G3 bug.
  3. Compute windows: exec[id] = i gives each node its execution position; then one pass over all nodes raises last_use[src] to the latest consumer position. Finally outputs and inputs are pinned to order.len() alloc_graph (alloc.rs:513-561, the last_use pass).
  4. Count consumers per node, n_consumers (alloc.rs:563-577) — the safety input for in-place aliasing (§2.4).
  5. Walk nodes in order (alloc.rs:597-786, the main walk). For each node, first sweep: free every tracked buffer whose last_use < i — its readers have all been positioned earlier, so its contents are officially dead GraphAllocator::sweep (alloc.rs:1256-1267). Then decide where this node's output lives:
    • KV store/load nodes → the layer's persistent K region (§2.5);
    • fused QKV nodes → their persistent regions plus an ordinary output buffer for the concatenated q|k|v result;
    • Silu / RoPE / QkvBiasRopeStore → try to alias the input buffer in place (§2.4);
    • everything else → alloc_in_pool(backend, size), which asks the pool for a recycled buffer of exactly that size or creates a new one, then records (backend, id) → last_use in buf_alive.
  6. Dead nodes get no buffer at all: last_use[id] > i is false when a node has no consumers (fusion orphans the Silu inside a fused SwiGLU), so no allocation happens, and the scheduler skips bufferless nodes at node_buffer (scheduler.rs:321-326, the bufferless-node skip).

The pool side of step 5 is where reuse actually happens in alloc_buffer (cpu_backend.rs:276-301):

#![allow(unused)]
fn main() {
fn alloc_buffer(&mut self, size: usize) -> usize {
    if let Some(idx) = self
        .free
        .iter()
        .position(|&id| self.buffers[id].len() == size)
    {
        let id = self.free.swap_remove(idx);
        self.buffers[id].fill(0.0);   // recycled: zero it, so stale data can't leak
        return id;
    }
    self.buffers.push(vec![0.0f32; size]);
    self.buffers.len() - 1
}
}

Note the exact-size match: recycling only takes a free buffer whose length equals the request. That keeps the bookkeeping trivial (a buffer's size never changes) at the cost of occasionally missing a "big enough" free buffer — a deliberate trade: first-fit-with-growth would save a few allocations but complicates every downstream size assertion.

Since E4 S2 the size the allocator asks for is the node's size class, not its element count (graph/allocplan.rs::class_size: powers of two up to 16 KiB, then 16 KiB steps). The backend still matches exactly, but every buffer of a class has the same length, so the second shape in a class now finds the first one's buffer and a rebuild with a slightly different n_tokens stops growing the pool. The node keeps its real length in its BufRef (offset + len), and every consumer slices to that window — a pool buffer is routinely longer than the node it serves, which is why fill_input checks the data against BufRef::len and writes through write_host_window, and why the MINFER_TRACE/viz capture windows its readback.

How much does all this save? A hand tally of the 0.5B prefill graph (24 layers, per-layer node list in docs/ARCHITECTURE.md §4.6) with a 440-token prompt makes it concrete. Every hidden-width buffer holds 896 × 440 × 4 B ≈ 1.58 MB; every FFN-width buffer 4864 × 440 × 4 B ≈ 8.6 MB. Naive per-node allocation would put ≈55 MB of activation buffers per layer × 24 layers ≈ 1.3 GB of live-at-build-time buffers on the heap. With liveness, the peak simultaneous set is roughly six hidden-width buffers plus three FFN-width ones — around 35–40 MB, about 30× less, and the pool only ever holds as many buffers as the peak demanded. For decode (nt = 1) the byte savings are small (a hidden buffer is 3.5 KB), but the buffer count still drops from 437 to a couple dozen — which is what matters for GPU buffer objects.

2.4 In-place aliasing: Silu and RoPE write into their input

An alias means two nodes share the same buffer on purpose, not by recycling accident. Silu (the sigmoid-linear unit activation) and RoPE (rotary position embedding, which rotates pairs of numbers by an angle derived from the token's position) are elementwise transforms with a special property: their output has the same shape as their input, and their input has no other reason to keep its old value if nothing else reads it. So instead of

q_buf ──(read)──▶ rope kernel ──(write)──▶ q_rope_buf   [2 buffers, 2 passes over memory]

the allocator maps the rope node's output to the input's buffer:

q_buf ──▶ rope kernel reads and overwrites q_buf in place   [1 buffer]

(alloc.rs:736-775, the in-place alias arm) implements this, guarded by exactly two conditions — the input's sole consumer is this op (n_consumers[src] == 1, from step 4 above) and the input lives on the same backend. Both guards are load bearing. If another node also reads the input, overwriting it destroys data that reader still needs. If the input is on another backend, "just use the input's buffer" would mean running your kernel against memory in another device's pool — and the cross-backend case gets its own treatment (§2.7), which is precisely where the Phase-3 bug lived (§3.4).

When aliasing applies, the aliased input's live range is extended to cover the aliasing op's consumers — extend_through_views (alloc.rs:763-768), the in-place alias extension — the buffer now carries two logical tensors' worth of deadlines, and liveness must respect the later one.

Why bother? Three reasons, in decreasing order of "wow":

  1. Correctness on GPU. This is the surprising one. On Metal, one split's kernels are encoded into a command buffer and only submitted at the split boundary (doc 08/14). If the allocator instead made rope read a host-side copy of its input, that copy would read a GPU buffer whose producing kernel is still queued, not run — stale data. Aliasing keeps the read/write inside the same command buffer in kernel order, which is always coherent. This is ARCHITECTURE.md invariant 4's "never host-copy a GPU-pending buffer" rule, and it was learned the hard way (§3.4).
  2. Memory traffic. Each avoided alias-copy is a full pass over the activation. The prefill graph runs two RoPEs per layer (Q and K) whose inputs have sole consumers, so aliasing skips (896 × 440 + 128 × 440) × 4 B ≈ 1.8 MB of copy per layer — about 43 MB of pure memcpy per 440-token prefill across 24 layers. (The FFN silu copy is skipped too on the fused path, but there the whole Silu node is folded into SwiGLU, so it is fusion's win, not aliasing's.)
  3. Parity with llama.cpp, which executes rope and silu in place for the same reasons (alloc.rs:736-775, the same in-place alias arm).

The model-side code cooperates with the rule. In the mixed-quant QKV decode path, the builder deliberately wires attention to the epilogue node so that the q matmul's buffer has exactly one consumer and can alias — the epilogue is qkv_bias_rope_store (models/qwen2/graph.rs:169-181: "Attention is wired to the epilogue node so q's matmul buffer has exactly one consumer (in-place alias rule, §5)").

2.5 The KV cache as persistent regions

Time for the term this doc has been promising. A KV cache is the transformer's memory of tokens it has already processed: for each layer and each past token, the attention mechanism's K (key) and V (value) vectors (what those are is doc 11's business; here they are just tensors named K and V). Autoregressive generation works by appending each new token's K/V to this notepad and letting attention read the whole accumulated prefix — that is why decode is cheap per token. KV cache preview (doc 01's phrase) means deciding how big that notepad is before anything is written.

minfer's allocator owns it as persistent regions: each layer gets two buffers, K and V, each sized n_kv_embd × n_ctx f32 elements, allocated the first time any node of that layer mentions the layer and then never freed and never recycled ensure_kv (alloc.rs:995). n_kv_embd is the KV width — 128 for Qwen2.5-0.5B (2 KV heads × head-dim 64), 1024 for Qwen3-4B — and n_ctx is the context budget from the CLI (--n-ctx, default 4096). The store/load node shapes carry the size kvcache_store (builder.rs:655) builds the store node with shape [n_embd, n_ctx, 1, 1], "shape mirrors the persistent region so the allocator can size it").

The byte arithmetic you should carry around:

one layer  : 2 regions × n_kv_embd × n_ctx × 4 B
0.5B       : 2 × 128  × 4096 × 4 B =  4.2 MB/layer  × 24 layers ≈ 100 MB
Qwen3-4B   : 2 × 1024 × 4096 × 4 B = 33.6 MB/layer  × 36 layers ≈ 1.2 GB
             ...at n_ctx = 40960 (10× the tokens):              ≈ 12.1 GB  (!)

That 12.1 GB is not hypothetical — it is the recorded lesson of docs/PERF-QWEN3-4B-VS-LLAMACPP.md §2, retold from the sizing side in §3.3 below.

Why "persistent, never recycled"? Two lifetimes matter, and both are longer than one graph execution:

  1. Across forward calls. The whole point of a KV cache is to survive between steps: token 50's attention must read tokens 0–49's K/V, written during previous forward calls. A liveness-recycled buffer would be overwritten by the very next matmul.
  2. Across graph rebuilds. Prefill (many tokens) and decode (one token) have different GraphParams, so the prefill→decode transition rebuilds the graph — new nodes, new node_to_buf mapping (doc 13). The allocator object, though, is the same object (that is the GraphCache design, §2.8), so ensure_kv finds the existing pair and returns it untouched. Zero copies: the KV the prefill just wrote is exactly where decode's attention will read it. Deviation 14 records this as the analogue of llama.cpp's KV living in the memory context rather than the graph's buffer set.

Mechanically, the K region does double duty as the store node's output buffer — node_to_buf (alloc.rs:680-681: "the node's buffer = the K region") — and the CPU executor enforces that contract supports_op (cpu_backend.rs:129): "KV store out buffer must be the K region"). The V region is a sibling the kernel reaches through the kv_pair handle (§3.2, excerpt 8). The load node executes as a no-op — it is a view of the K region execute_node (cpu_backend.rs:308).

2.6 Filling inputs: why f32 buffers, and the I32 bit-pattern ride

The graph declares three inputs for a prefill — token_ids (models/qwen2/graph.rs:57-84) [nt,1,1,1], positions [nt,1,1,1], and (when the tail-row optimization is active) tail_ids — all typed DType::I32 in the IR. Yet every pool buffer is f32 (§2.1). The bridge is fill_input_i32 fill_input_i32 (alloc.rs:1903): each u32 is packaged as f32::from_bits(v) — a pure bit reinterpretation, not a numeric conversion — and written into the input node's buffer via the backend's write_host. On the consumer side the kernels run the inverse, x.to_bits(), recovering the exact integer:

consumercode
embedding row gather (CPU)ins[0][t].to_bits() → token id execute_node (cpu_backend.rs:308)
generic get_rows (CPU)ins[1][t].to_bits() as usize execute_node (cpu_backend.rs:308)
RoPE positions (CPU)ins[1][t].to_bits() as usize execute_node (cpu_backend.rs:308)
attention positions (CPU)ins[2][t].to_bits() as usize execute_node (cpu_backend.rs:308)
CUDA kernelsdevice-side __float_as_int in one pass gather_rows_f32 (ops_misc.cu:145)

Why this trick at all? Because of the uniform-pool decision. The alternatives were a second, integer-typed pool per backend (double the allocator state, double the copy paths, and a special-case alloc_buffer(size, dtype) in every backend) or converting integers to their float values (which is exact only for small integers and lossy in surprising ways). Riding the bits keeps one pool and is lossless. The safety envelope recorded in the code — "exact for |v| < 2^24" (cpu_backend.rs:2-4, the module doc) — is generous headroom: 2²⁴ = 16,777,216, and real data sits far inside it — the largest vocabulary here is 151,936 token ids, and contexts top out in the tens of thousands (the biggest n_ctx in the perf tables, 65,536, is clamped to the model's 40,960-token max_seq_len before it reaches the allocator). Within that envelope the pattern is robust even if some stage ever treated the contents as a float value instead of bits.

The CUDA note is worth savoring because it shows the constraint pushing back: the decode kernels need raw int32, but converting on the host would need a sync (and would break CUDA Graph replay, doc 15). So a tiny device kernel f32_bits_to_i32 reinterprets the bits on the GPU, "fully device-side, so the per-layer path needs no host sync (and stays CUDA-Graph-replayable)" f32_bits_to_i32 (ops_elementwise.cu:249).

Filling happens at a strict moment: after alloc_graph, before the scheduler runs forward_batch (models/qwen2/graph.rs:465): cache.current() → three fill_input_i32 calls). That ordering is exactly why inputs must be pinned out of the recycling pool (§2.2 rule 2) — the fills would otherwise fight each other over a shared buffer before any node had executed (§3.4, bug 2b).

2.7 Cross-backend staging and alloc_fresh

When doc 06's assignment puts a producer on Metal and its consumer on CPU, the scheduler inserts a split boundary: sync the previous backend, then copy the consumer's inputs across execute (scheduler.rs:266-275, the phase-A staging enqueue). The copy lands in the allocator's copy_across, which routes through copy_to_cpu (a host round trip — Metal/CUDA buffers here are CPU-visible, so this is a plain memcpy) and then write_host into a buffer on the destination pool copy_across (alloc.rs:2449), the staging copy).

That destination staging buffer must be fresh — alloc_fresh_in (alloc.rs:955-957) — never drawn from the recycle free list. The trait comment is the design record (backend.rs:56-62):

#![allow(unused)]
fn main() {
/// Allocate a buffer that bypasses the recycle free list. Split-boundary
/// staging needs this: at execute time the free list holds ids whose
/// physical contents are still referenced by node_to_buf and get
/// read/written later in the same execute — recycling one would clobber
/// in-flight data. Fresh buffers enter the normal free list on
/// free_buffer (at graph rebuild), where liveness recycling is safe.
}

Unpacking that: during the build loop, the free list is safe to draw from because the sweep clock (§2.3 step 5) advances monotonically through liveness order — anything freed is dead from that position onward, and every later allocation is also later in execution. But a split boundary is an out-of-band allocation: it happens at execution position P, with the free list frozen in whatever state the build left it. The list can still contain a buffer whose last reader sits at position ≥ P (nothing after it in the build happened to want that size), and a staging write at P would clobber data that execution has not consumed yet. Fresh allocation sidesteps the whole question by never consulting the list.

Staging buffers are one-per-(node, destination backend) per graph: the first execute allocates, every later execute of the reused graph just rewrites the same buffer (alloc.rs:119-121, the cross field's "no per-step allocation" note). They are freed at the next rebuild (alloc.rs:121-124, where E4 S3 keeps the staging entries) — the "at graph rebuild" moment the trait comment mentions, where returning them to the normal free list is safe because the next build's monotonic sweep re-establishes the invariant from scratch.

2.8 GraphCache: the allocator outlives the graph

The final principle is an ownership decision that makes §2.5 possible. GraphCache (cache.rs:38-45) is a tiny struct: the cached graphs, the allocator, and the params each was built for. Reuse is decided by try_reuse (cache.rs:63-73), which compares GraphParams only — n_tokens, n_out, gtype, cparams (including n_ctx, the GPU flag, and the fusion toggles), and weights_version. Equal params ⇒ the topology is deterministic ⇒ reuse the graph and just refresh input data (§2.6). Mismatched params ⇒ the caller builds a new graph and — keeping the allocator — replace_graph (cache.rs:103-109) swaps it in.

That is the entire reason the KV regions survive: the regions live inside the allocator, the allocator lives inside the cache, and rebuilds replace only the graph. The unit test allocator_survives_rebuild pins this contract with a planted persistent region allocator_survives_rebuild (cache/tests.rs:159).

3. Implementation

3.1 Data in / data out

In:

  • A ComputeGraph fresh from doc 06: nodes with backend: Some(_) assigned, fusion applied (some nodes orphaned, some replaced by FusedQKV/SwiGLU style ops), shapes and dtypes final.
  • CParams.n_ctx — riding inside GraphParams — which sizes every KV region (§2.5).
  • Registered weights, already inside the backends' registries (doc 03) — the allocator's CPU pool is the same object weight registration went through GraphAllocator (alloc.rs:107) delegates to self.cpu.register_weight).
  • Host data for inputs: &[u32] token ids, positions, tail ids forward_batch (models/qwen2/graph.rs:465).

Out:

  • node_to_buf: HashMap<NodeId, BufRef> — the answer to "where does node N's output live". The scheduler consumes it for every node of every split node_buffer (scheduler.rs:324-326, the per-node read).
  • kv: HashMap<layer, [BufRef; 2]> + persistent: Vec<PersistentBuf> — the KV regions with stable names like "kv.7.k" / "kv.7.v", exposed to backends through the trait KvProvider (backend.rs:12-19), alloc_graph (alloc.rs:513).
  • cross: HashMap<NodeId, BufRef> — split-boundary staging copies, filled lazily during the first execute and rewritten on later ones.
  • Filled input buffers, ready before the scheduler's first node.

Shapes at this stage (Qwen2.5-0.5B, decode step, CPU path): inputs [1,1,1,1] f32-carried I32; hidden-width buffers [896,1,1,1] = 3.5 KB; FFN-width [4864,1,1,1] = 19.5 KB; KV regions [128, 4096, 1, 1] = 2.1 MB each, two per layer, 24 layers ≈ 100 MB total. All f32.

3.2 Key code

Excerpt 1 — the fields of GraphAllocator (src/graph/alloc.rs:106-188) — the struct and its field map below reappear in the walk; the comments record the ownership rules.

#![allow(unused)]
fn main() {
pub struct GraphAllocator {
    cpu: CpuBackend,
    #[cfg(target_os = "macos")]
    metal: Option<super::metal_backend::MetalBackend>,
    #[cfg(feature = "cuda")]
    cuda: Option<super::cuda_backend::CudaBackend>,
    node_to_buf: HashMap<NodeId, BufRef>,
    /// Cross-backend copies for the CURRENT graph (split-boundary staging):
    /// node → buffer on the consuming split's backend. NOT part of the node's
    /// canonical assignment — node_to_buf must stay re-executable (a remap
    /// would break the next execute of a reused graph, whose producing split
    /// would find its buffer on another backend). The same staging buffer is
    /// rewritten on every execute (no per-step allocation).
    cross: HashMap<NodeId, BufRef>,
    /// (backend, pool id) → last exec index it stays alive until
    buf_alive: HashMap<(Backend, usize), usize>,
    /// per-layer KV persistent regions: [k, v]
    kv: HashMap<usize, [BufRef; 2]>,
    /// All persistent regions (never freed).
    pub persistent: Vec<PersistentBuf>,
}
}

Note the deliberate separation of node_to_buf (canonical, re-executable) from cross (staging). A naive design would move a node's buffer to the consuming backend — which would break execute #2 of a reused graph, when the producing split needs its buffer back where it was.

Excerpt 2 — liveness in build order, with inputs and outputs pinned (alloc.rs:526-561). This is the code that bug G3 rewrote; the comment is the tombstone.

#![allow(unused)]
fn main() {
// The scheduler executes nodes in BUILD order (node id order — the
// builder appends sources before consumers), so liveness must use the
// same order: topo_order() can reorder srcless nodes (kv_load) ahead,
// which would let a later consumer's buffer reuse clobber an input the
// scheduler has not yet read (G3 tail get_rows regression). Validate
// acyclicity, but keep build order.
graph.topo_order()?;
let order: Vec<NodeId> = (0..graph.n_nodes()).collect();
let n = graph.n_nodes();

let mut exec = vec![0usize; n];
for (i, &id) in order.iter().enumerate() {
    exec[id] = i;
}
let mut last_use = exec.clone();
for (i, &id) in order.iter().enumerate() {
    for &s in &graph.node(id).src {
        if last_use[s] < i {
            last_use[s] = i;
        }
    }
}
for &o in &graph.outputs {
    last_use[o] = order.len();
}
// Inputs are filled on the host BEFORE execution starts, so every
// input buffer is live at fill time; liveness (which tracks execution
// order) must never reuse an input's buffer for another input — the
// later fill would clobber the earlier one. Treat inputs like outputs.
for &i in &graph.inputs {
    last_use[i] = order.len();
}
}

last_use starts as each node's own position, so a node with no consumers (a fusion orphan) has last_use == exec and will get no buffer; a node read by many consumers ends at the latest reader.

Excerpt 3 — the main walk: sweep, then per-node decision (alloc.rs:597-784, the main walk; the KV arm at 669-682, the generic arm at 777-784 is three lines of "alloc if alive").

#![allow(unused)]
fn main() {
for (i, &id) in order.iter().enumerate() {
    self.sweep(i);
    let node = graph.node(id);
    let backend = node.backend.unwrap_or(Backend::CPU);
    match node.op {
        Op::KvcacheStore { layer } | Op::KvcacheLoad { layer } => {
            let pair = self.ensure_kv(layer, backend, node.n_elements());
            // the node's buffer = the K region
            self.node_to_buf.insert(id, pair[0]);
        }
        Op::FusedQKV { layer } => {
            // fused decode QKV: also needs the layer's persistent KV
            // regions (the kernel stores K/V), but its output is a
            // normal concat buffer (q|k|v), not the K region.
            let kv_elems = match &node.meta {
                NodeMeta::FusedQkv(m) => m.kv_elems,
                _ => node.n_elements(),
            };
            self.ensure_kv(layer, backend, kv_elems);
            if last_use[id] > i { /* … ordinary buffer for the concat … */ }
        }
}

GraphAllocator::sweep (alloc.rs:1256-1267) collects every buf_alive entry whose deadline passed (al < i), removes it, and hands the id to free_in_pool — which pushes it onto the backend's free list. Nothing is deallocated; "free" here means "return to the recycling pool", which is why the next alloc_in_pool of the same size is a zero-cost reuse (plus one zero-fill on CPU).

Excerpt 4 — the in-place alias arm (alloc.rs:736-775). The two guards and the live-range extension, exactly as argued in §2.4.

#![allow(unused)]
fn main() {
// In-place elementwise transforms: alias the input buffer
// (llama.cpp executes rope/silu in place). Same-backend
// aliasing avoids a host-side copy between a pending GPU
// producer and this kernel — it reads/writes the buffer the
// producer wrote, in kernel order. Cross-backend inputs get
// a fresh buffer: the producer completed before the split
// boundary, so the backend's host copy is safe there.
if last_use[id] > i {
    let in_ref =
        self.node_to_buf.get(&node.src[0]).copied().ok_or_else(|| {
            format!("in-place op src buffer missing (node {id})")
        })?;
    // alias only when the input's sole consumer is this op
    // (in-place overwrites the input) AND it is on the same
    // backend
    if in_ref.backend == backend && n_consumers[node.src[0]] == 1 {
        self.node_to_buf.insert(id, in_ref);
        // the aliased input must stay alive through this
        // node's consumers
        last_use[node.src[0]] = last_use[node.src[0]].max(last_use[id]);
    } else {
        let size = node.n_elements();
        let pid = self.alloc_in_pool(backend, size);
        self.buf_alive.insert((backend, pid), last_use[id]);
        self.node_to_buf.insert(id, BufRef { backend, id: pid });
    }
}
}

The else branch matters as much as the if: a cross-backend or multi-consumer input silently falls back to a normal buffer. Aliasing is an optimization with strict preconditions, never an assumption.

Excerpt 5 — persistent region creation ensure_kv (alloc.rs:995).

#![allow(unused)]
fn main() {
/// Per-layer KV persistent regions (K and V), created on first use on the
/// layer's assigned backend.
fn ensure_kv(&mut self, layer: usize, backend: Backend, size: usize) -> [BufRef; 2] {
    if let Some(&pair) = self.kv.get(&layer) {
        return pair;
    }
    let k = self.alloc_persistent(&format!("kv.{layer}.k"), backend, size);
    let v = self.alloc_persistent(&format!("kv.{layer}.v"), backend, size);
    self.kv.insert(layer, [k, v]);
    [k, v]
}

/// Allocate a persistent (never-freed) region on a backend.
pub fn alloc_persistent(&mut self, name: &str, backend: Backend, size: usize) -> BufRef {
    let id = self.alloc_in_pool(backend, size);
    self.persistent.push(PersistentBuf {
        name: name.to_string(),
        backend,
        id,
    });
    BufRef { backend, id }
}
}

Two quiet details: the pair is created once per layer per process — the if let Some early-return is what makes rebuilds zero-copy (§2.5) — and alloc_persistent never touches buf_alive, so no sweep can ever free it. The region is also sized on first use only: if a later graph asked for a different size, it would silently get the old buffer — one reason n_ctx must stay consistent across a run (§3.3, question 3).

Excerpt 6 — I32 input filling fill_input_i32 (alloc.rs:1903) plus the routing tail of fill_input_impl, alloc.rs:2077).

#![allow(unused)]
fn main() {
/// Fill an I32 input (token ids / positions). Stored as `f32::from_bits`
/// patterns — exact for |v| < 2^24.
pub fn fill_input_i32(
    &mut self,
    graph: &ComputeGraph,
    name: &str,
    data: &[u32],
) -> Result<(), String> {
    let bits: Vec<f32> = data.iter().map(|&v| f32::from_bits(v)).collect();
    self.fill_input_impl(graph, name, &bits)
}
}
#![allow(unused)]
fn main() {
let id = graph.inputs.iter().copied()
    .find(|&i| graph.node(i).name == name)
    .ok_or_else(|| format!("no input node named '{name}'"))?;
let br = self.node_buffer(id)
    .ok_or_else(|| format!("input '{name}' has no buffer (not allocated)"))?;
// F4: the pool is looked up through the registry handle, not matched on a variant.
let (backend, id, offset) = (br.backend, br.id, br.offset);
match self.pool_mut(backend) {
    Some(pool) => pool.write_host_window(id, offset, data),
    None => Err(format!("{} is not usable on this allocator: {}", backend.name(), …)),
}
}

Inputs are found by name, not position — the graph is rebuilt between prefill and decode, so node ids may shift, but the names "token_ids" / "positions" / "tail_ids" are stable API.

Excerpt 7 — the copy that must be fresh alloc_graph (alloc.rs:513), the tail of copy_across).

#![allow(unused)]
fn main() {
// `copy_across_blocking` (the host leg), `alloc.rs:2546-2557`:
let data = self
    .copy_to_cpu(node_id)
    .ok_or_else(|| format!("node {node_id} host read failed"))?;
self.write_cross_staging(dst, &data)

// `write_cross_staging` (`alloc.rs:2630`) → `write_pool` (`alloc.rs:2810`): one
// registry lookup through the handle, never a `match` on a backend variant.
match self.pool_mut(dst.backend) {
    Some(pool) => pool.write_host(dst.id, data),
    None => Err(…),
}
}

The staging BufRef is minted by cross_staging (alloc.rs:2499-2532), which is also where the entry is inserted into self.cross — the destination buffer is allocated fresh there (alloc_fresh_in), so the copy never lands in recycled memory.

copy_to_cpu is safe here — this code only runs at a split boundary, i.e. after the producing split was synchronized (§2.7). The same host copy performed inside a split, against an unsubmitted command buffer, is the Phase-3 bug (§3.4).

Excerpt 8 — how backends receive the KV regions KvProvider (backend.rs:12-19) and the scheduler's resolution, kv_pair (scheduler.rs:358-370).

#![allow(unused)]
fn main() {
pub trait KvProvider {
    /// (k_buf_id, v_buf_id) of a layer's persistent regions on this pool.
    fn kv_pair(&self, layer: usize) -> Option<(usize, usize)>;
}
}
#![allow(unused)]
fn main() {
let kv_pair = match &node.op {
    Op::KvcacheStore { layer } => alloc.kv_pair(*layer),
    Op::FusedQKV { layer } => alloc.kv_pair(*layer),
    Op::QkvBiasRopeStore { layer } => alloc.kv_pair(*layer),
    Op::FusedQkvNorm { layer } => alloc.kv_pair(*layer),
    Op::Attn { .. } => match &node.meta {
        NodeMeta::Attn(m) => alloc.kv_pair(m.layer),
        _ => None,
    },
    _ => None,
};
}

execute_node takes kv_pair: Option<(usize, usize)> alongside the ordinary input ids execute_node (backend.rs:80) — the K/V regions are not the node's src inputs; they are process-lifetime siblings only KV-aware ops know about. The CPU store kernel shows the split-brain clearly execute_node (cpu_backend.rs:308-359): K is written through out_buf (which the allocator guaranteed is the K region), V through the sibling id, both reached with split_at_mut for disjoint mutable borrows, and positions decoded from the I32 input with to_bits (cpu_backend.rs:332-337) — with a hard error if a position exceeds n_ctx (cpu_backend.rs:358-359), never a silent overflow.

Excerpt 9 — the pool's two remaining flavors, free_buffer (cpu_backend.rs:291-306) and alloc_fresh: alloc_buffer was already shown in §2.3, so this is just its siblings.

#![allow(unused)]
fn main() {
fn free_buffer(&mut self, id: usize) {
    if !self.free.contains(&id) { self.free.push(id); }   // "free" = recycle
}
fn alloc_fresh(&mut self, size: usize) -> usize {
    // never recycled from the free list (see Backend::alloc_fresh)
    self.buffers.push(vec![0.0f32; size]);
    self.buffers.len() - 1
}
}

(Metal's pool is the same shape with MTLBuffer lengths in bytes, alloc_buffer (metal_backend.rs:897), except recycled buffers are not re-zeroed — kernels fully overwrite their outputs, and the driver zero-fills only new allocations.)

Excerpt 10 — GraphCache (cache.rs:38-114): params-only reuse, allocator kept (the struct, try_reuse, and replace_graph).

#![allow(unused)]
fn main() {
pub fn try_reuse(&mut self, params: &GraphParams) -> bool {
    match (&self.prev_params, &self.graph) {
        (Some(prev), Some(_)) if Self::params_match(prev, params) => {
            self.prev_params = Some(params.clone());
            true
        }
        _ => false,
    }
}

fn params_match(a: &GraphParams, b: &GraphParams) -> bool {
    a.n_tokens == b.n_tokens
        && a.n_out == b.n_out
        && a.gtype == b.gtype
        && a.cparams == b.cparams
        && a.weights_version == b.weights_version
}

/// Store a freshly built graph. The allocator is kept (KV regions persist);
/// its liveness mapping is recomputed by the caller via `alloc_graph`.
pub fn replace_graph(&mut self, mut graph: ComputeGraph, params: GraphParams) {
    graph.uid = NEXT_GRAPH_UID.fetch_add(1, Ordering::Relaxed);
    self.graph = Some(graph);
    self.prev_params = Some(params);
}
}

Note what is absent from params_match: n_past (how many tokens are already in the KV cache). Positions are data, not structure — ARCHITECTURE.md invariant 1 — which is the precondition for the whole reuse scheme: if topology depended on n_past, every decode step would rebuild.

3.3 Design choices (why this shape and not another)

Why does the allocator own the backend pools? Why is the scheduler a pure orchestrator? (COMPUTE-GRAPH-DESIGN.md deviation 11.) Three forces point the same way. One id space: buffer ids appear in node_to_buf, in split input lists, in kernel launches, and in captured CUDA Graphs; if two components each held a pool, every id would need a "whose?" qualifier and every bug a suspect. One lifetime: pools must live exactly as long as the cached graph (rebuilding them per step would re-create hundreds of GPU buffer objects per token); the cache owns the graph, so the cache owns the pools, through the allocator. One registration path: weights land in the same CPU pool object (register_weight), which is how "does this node's weight live on the GPU?" becomes a simple registry query during assignment (doc 06). The scheduler keeps only orchestration logic — assign, split, copy, run — and borrows the backends mutably through the registry pool hook — pool_mut (scheduler.rs:380-398) at execution time.

Why is buffer reuse safe here when it broke twice? Because each bug was a missing precondition, not a flaw in liveness itself, and the fixes wrote the preconditions into the code:

  1. Reuse is only sound if liveness is computed in the same order the executor runs the nodes. The G3 bug computed it in Kahn topological order while execution used build order (§3.4). Now both are build order — literally 0..n_nodes — so "dead after position i" means the same thing to both components.
  2. Reuse must respect who fills memory outside the node walk. Inputs are host-filled before execution; outputs are read after. Both classes are pinned to order.len() and never recycled (§2.2).
  3. Reuse must respect who allocates outside the build loop. Split-boundary staging bypasses the free list via alloc_fresh (§2.7).
  4. In-place sharing (aliasing) is a stronger claim than reuse — two live tensors, one buffer — so it carries its own extra guards: sole consumer, same backend (§2.4).

The general lesson: sharing memory is safe exactly when every writer's schedule is known and every reader is accounted for in one order. minfer now has that schedule (build order) and that accounting (last_use + pins + fresh staging), enforced by code, comments, and the unit tests of §4.

Why size KV by n_ctx and not the model's max_seq_len? The regions are allocated once and their size is n_kv_embd × n_ctx — so n_ctx is the single biggest memory decision in the process, and for Qwen3-4B the wrong answer was 12.1 GB: the single-shot CLI used to pass max_seq_len = 40960 straight through, giving 36 layers × 2 regions × 40960 × 1024 × 4 B = 12.1 GB of Metal shared buffers for a 10-token prompt. The damage was not resident memory (peak RSS was identical, ~2.1 GB, at 4096 and 40960) but the Metal driver's one-time first-submit setup, which scales with total buffer bytes: 289 ms at n_ctx 40960 vs 106 ms at 4096 — a 3× tax on the first token (docs/PERF-QWEN3-4B-VS-LLAMACPP.md §2). The fix put the choice in the CLI's hands (--n-ctx, default 4096; doc 01 covered that side) and clamped it: main.rs computes ctx as the larger of params.n_ctx and input_ids (main.rs:1451-1453) — a long prompt must never overflow the notepad — and the model clamps again with Qwen2Graph::forward (src/models/qwen2/graph.rs:438-445), which applies n_ctx.min(max_seq_len). One more consistency requirement hides here: because ensure_kv sizes on first use only (excerpt 5), prefill and decode must pass the same n_ctx so the regions created during prefill are correctly sized for every decode step — the comment "Computed ONCE so prefill and decode size the same KV regions" main.rs:1452 pins that.

Why are inputs f32 buffers at all? Because the pool is uniform and the two numeric paths agree on f32 as the interchange format: GPU backends read f32 activations directly (convention #1 in AGENTS.md; CPU quantizes activations to Q8_0 at the matmul, inside the kernel), so f32 is already the lingua franca of every buffer. Integer inputs ride as bit patterns (§2.6). The alternative — per-dtype pools — would multiply allocator state, copy paths, and backend code for the sake of two [nt]-element integer buffers per graph; the bit-pattern trick costs one from_bits/to_bits pair per element and one explanatory comment.

3.4 Pitfalls & invariants

Bug 1 — the Phase-3 KV-corruption bug (never host-copy a GPU-pending buffer). After doc 06's assignment, a GPU-resident layer's RoPE input sometimes needed a copy: the original allocator materialized cross-backend and in-place inputs through a host copy_in. On Metal, though, one split's kernels are encoded into an MpsCommandBuffer as they execute — and only submitted at the split boundary capture_split (metal_backend.rs:173), execute_node (metal_backend.rs:925). A host copy enqueued mid-split therefore read the buffer's old contents: freshly allocated Metal memory, i.e. zeros. The copy captured zeros, RoPE dutifully rotated them, KvcacheStore wrote them into the layer's persistent region — and the whole KV region was zeros, so every attention read garbage and the output was unintelligible. The recorded fix (COMPUTE-GRAPH-DESIGN.md deviation 18; ARCHITECTURE.md invariant 4): same-backend in-place ops alias their input (the read and the write happen inside the same command buffer, in kernel order, so coherence is guaranteed by the GPU's own queue), and cross-backend inputs get a fresh buffer — safe, because the producer split was synchronized at the boundary before any copy runs (excerpt 7). After the fix, all 437 nodes of the 0.5B graph ran correct on a single command buffer. The invariant, verbatim from the architecture doc: "Never host-copy a GPU-pending buffer: a host copy_in of a producer that is encoded but not submitted reads stale data."

Bug 2 — the G3 liveness-order bug (liveness must follow execution order). The allocator originally computed liveness over topo_order() — a Kahn topological sort — while the scheduler executes in build order. Both are valid topological orders, but they are not the same order: Kahn's queue front-loads every source-less node (inputs, kv_load) and can reorder two independent nodes relative to each other. In the G3 tail-shrink graph, that reordering made the allocator believe the residual stream h was dead earlier than execution would prove — so when the attention node was allocated (after h's supposed last use, in Kahn order), the sweep had already recycled h's buffer to it. Execution then ran in build order: the attention node wrote its output into what was still h's buffer, and the later get_rows(h) read the attention output instead of the residual — logits off by 21.79 (COMPUTE-GRAPH-DESIGN.md deviation 22). The fix is excerpt 2: call topo_order()? purely to reject cycles, then compute liveness over 0..n_nodes — the order the scheduler actually runs. (Small forensics note: the doc comment on topo_order still says "used by the allocator" (graph/mod.rs:237) — a stale leftover; GraphAllocator (alloc.rs:107) is authoritative.)

Bug 2b — input buffers are never freed. Same fix series, complementary rule (deviation 23): inputs are host-filled before execution, but liveness only tracks consumers during execution — so token_ids (last consumer: the embedding, position 1) looked dead long before positions was filled, the two inputs' buffers were reconciled into one, and the later fill clobbered the earlier (recorded as "the embedding_and_rope regression: token_ids overwritten by positions"). Every embedding then gathered garbage rows. Fix: last_use[i] = order.len() for all inputs (excerpt 2's final loop) — the cost is a few dozen bytes pinned per graph; the benefit is that fill order no longer matters.

The remaining invariants, in one list (each traceable to a §2 section):

  • Aliasing requires sole-consumer and same-backend; everything else allocates normally (§2.4).
  • Inputs and outputs are pinned to the end of the execution; persistent regions are outside buf_alive entirely (§2.2, §2.5).
  • Split-boundary staging always allocates fresh; it rejoins the free list only at rebuild (§2.7).
  • KV positions are data: the region is sized n_kv_embd × n_ctx, and a position ≥ n_ctx is a loud error, not an overflow (cpu_backend.rs:359), plus the pre-flight assert maxp < n_ctx in register_graph_weights (models/qwen2/graph.rs:393).
  • Dead nodes get no buffer and the scheduler skips them — so adding an op the fusion pass orphans cannot corrupt memory, it just does nothing — skipped where node_buffer (scheduler.rs:321-326, the bufferless-node skip) reads None.

4. Observe & verify

  • cargo test — the allocator's own unit tests alloc_graph (src/graph/alloc.rs:513): liveness_reuses_buffers_along_chain and parallel_chains_do_not_share assert the two liveness behaviors of §2.2; kv_regions_two_per_layer asserts store and load share the K region, V is a sibling, and exactly two persistent regions exist for one layer; cycle_graph_allocation_fails proves the acyclicity check is live. allocator_survives_rebuild (src/graph/cache/tests.rs:159) pins "persistent regions survive rebuilds". Filter with cargo test liveness / cargo test kv_regions.
  • MINFER_TRACE=/tmp/t.json ./target/release/minfer model.gguf "Hello" — records per-node real data for the viz page; input nodes appear host-filled, and you can watch a buffer's contents change across the nodes that share it.
  • MINFER_GRAPH_DUMP=/tmp/d … — dumps logits and the KV regions after a run; the KV dump is exactly the persistent regions of §2.5, so zeros there would reproduce bug 1's symptom.
  • --dump-graph / --dump-graph-json — re-runs build → assign → fusion and exports the 437-node graph with backend colors; the node ids it shows are the build order the allocator's liveness uses.
  • MINFER_NO_FUSE_QKV=1 / MINFER_NO_FUSE_FFN=1 — flips the fusion toggles, which changes cparams, which forces a graph rebuild with the same allocator — a hands-on way to watch KV regions survive a rebuild (§2.8) while the node/buffer mapping is recomputed.
  • Greedy equivalence checks — the recorded acceptance for both bug fixes: --temp 0 output identical pre/post fix, and fused-vs-unfused decode logits diff 0.000 (COMPUTE-GRAPH-DESIGN.md deviations 18, 22-24, 25-26).

5. Cross-references

  • docs/ARCHITECTURE.md §4.3 (pipeline position), §4.4 (GraphCache), §4.5 (the invariants this doc expanded, esp. 1, 2, 4, 5), §7 (KV cache summary) — the compressed version of this stage.
  • docs/COMPUTE-GRAPH-DESIGN.md §3.3 (original allocator design — note where the implementation diverged: per-backend pools, build-order liveness, two regions instead of one [K|V] block), §17 deviations 11 (pool ownership), 14 (allocator survives rebuilds), 18 (aliasing fix), 20 (two regions per layer), 22-23 (the G3 liveness fixes).
  • docs/PERF-QWEN3-4B-VS-LLAMACPP.md §2 — the 12 GB n_ctx lesson with its measurements.
  • 05 — Graph build (IR) — where the node list, input names, and KV node shapes come from.
  • 06 — Backend assignment and fusion — upstream: decides which backend each buffer must be allocated on, and creates the orphan/fused node shapes the allocator must handle.
  • 08 — The scheduler: splits, copies, execution — downstream: consumes node_to_buf/cross/kv_pair, runs the split boundaries whose staging rules this doc motivated.
  • 11 — Attention + vec ops + KV — what K and V actually mean and how attention reads the written prefix.
  • 13 — Decode loop + graph reuse — the prefill→decode rebuild that the persistent regions are designed to survive.

← 06 — Backend assignment and fusion · Index · 08 — The scheduler: splits, copies, execution →

08 · The scheduler: splits, synchronization, and execute

Stage: graph allocated (07 — Memory allocation: liveness and the KV regions) → this stage: the graph finally does work → the forward produces logits (09 — Prefill: the first forward).

Code: src/graph/scheduler.rs — BackendScheduler::split_graph (L73) and execute (L123); the Backend trait contract in src/graph/backend.rs; the split-boundary helpers in src/graph/alloc.rs (sync_backend L542, copy_across L575); the three execute_node implementations in src/graph/cpu_backend.rs, src/graph/metal_backend.rs, src/graph/cuda_backend.rs. All line numbers verified against commit e7fa0da (the current HEAD).

1. Background — where this stage sits

Everything so far in Act 2 of this walkthrough has been bookkeeping. Doc 05 built a compute graph: a list of nodes, one per math operation of the transformer, each describing what to compute but computing nothing. Doc 06 assigned every node to a backend — CPU, Metal, or CUDA — by asking each backend "can you do this op?" in priority order. Doc 07 allocated buffers: every node's output has a home, memory is shared between nodes whose lifetimes don't overlap, and each layer owns two persistent KV (KV = key/value, the attention memory of the model) regions. What exists as this stage begins is: a ComputeGraph whose nodes are in build order (the order the builder appended them, sources before consumers); a GraphAllocator owning one buffer pool per backend, with a node → buffer map covering every live node; and the graph's input buffers already filled with this step's data (token ids, positions — doc 07). Nothing has been computed yet.

This stage is where the math happens. The scheduler walks the node list, hands each node to the backend it was assigned to, and — the new, subtle part — makes sure that when a node reads its inputs, those inputs actually contain the values their producers wrote. On the CPU that is almost trivial: a function computes, returns, the result is in memory. On a GPU it is not trivial, because a GPU is an asynchronous device: asking it to do work and getting the result are two separated events in time. Most of this document is about managing that gap.

Two terms we will use constantly. Scheduling is deciding where work runs (which backend — doc 06) and in what order, with what synchronization (this document). A split is minfer's unit of scheduling: a maximal contiguous run of nodes all assigned to the same backend. Without this stage's synchronization the model would not merely be slow — it would be wrong: attention would read KV rows before they were written, and the logits buffer would be read while the GPU was still filling it. The scheduler's whole job is "never read a value before its producer finished writing it", across three very different devices.

2. Principle — how it works and why

2.1 Execution is a walk over a list that is already in the right order

The scheduler does not run a sort before executing. It walks graph.nodes[0..n] in index order — build order — and executes each node on its assigned backend. Why is that correct? Because of how doc 05's builder works: every time the builder creates an operation node, the node's sources already exist as earlier nodes, so every consumer sits after all of its producers. An ordering with that property has a name: a topological order of the graph (a DAG — directed acyclic graph, a dependency network with no cycles — is the data structure here).

minfer keeps the formal sort only as a checker: execute opens with debug_assert!(graph.topo_order().is_ok(), ...) (scheduler.rs L124–125; a Kahn sort, src/graph/mod.rs L154). In release builds nothing is re-sorted — the walk is the order. §3.3 returns to this "one order everywhere" decision with the bug that proved it.

Build order buys two guarantees for free:

  1. A KV store node always executes before the attention node that reads the KV it wrote. kvcache_store/kvcache_load being explicit nodes (doc 05) makes this an ordinary data-dependency instead of hidden control flow: attention's src list contains the load node, the load node sits after the store node, so the walk order enforces write-before-read with no special case anywhere.
  2. The allocator's liveness analysis uses the same order. Doc 07 computes "when is this buffer's last read?" over node ids 0..n. If the executor and the allocator disagreed about the order, the allocator would recycle a buffer while execution still needed it. That disagreement actually happened — the G3 bug, logits off by 21.79 — and the fix was to make both sides use build order (§3.3, §3.4).

2.2 Three backends, two execution styles

The three backends differ fundamentally in when work happens relative to the call that requests it:

  • CPU — execute_node computes immediately and returns with the result already in the buffer. Synchronous: the call's return means the work is done.
  • Metal (Apple GPU) and CUDA (NVIDIA GPU) — execute_node only records work and returns immediately. Asynchronous: the call's return means "the work has been enqueued", nothing more.

For Metal, "recording" means appending a kernel launch to a command buffer: a list of GPU commands the CPU builds up in memory, which the GPU executes only after the CPU explicitly submits it (minfer keeps one command buffer per split). For CUDA, recording means launching a kernel on a stream: an ordered queue of GPU work; kernels on one stream run in launch order, and the launch call returns long before the kernel finishes. The third word is synchronize: wait until every piece of work previously enqueued on this backend has finished — the mirror image of "submit" (hand the batch to the GPU).

Why be asynchronous at all? Because encoding is cheap and GPU execution is long: while the GPU chews through kernel #37, the CPU can already be encoding kernel #38. Synchronizing after every node would idle one side or the other at each step. Batching a whole split into one submission keeps both busy: a 0.5B decode step costs ~3.9 ms end-to-end (~256 t/s, the G1–G3 figures in docs/METAL_OPTIMIZATIONS.md L55) — one submit per split pays the CPU↔GPU round-trip cost once per step, not once per node.

Asynchrony creates exactly one hazard, and the scheduler exists to police it: a host read of a buffer the GPU has not finished writing returns stale data. Every place the scheduler touches GPU-written memory is therefore placed after a synchronize.

2.3 Splits: contiguous runs are submission units

split_graph scans the node list in build order and cuts it every time the assigned backend changes. Each piece is a Split:

#![allow(unused)]
fn main() {
// src/graph/scheduler.rs L22-32
pub struct Split {
    pub backend: BackendTag,
    /// Node id range [start, end) in graph.nodes.
    pub node_range: (usize, usize),
    /// Nodes whose source values live on another backend (copied in).
    pub inputs: Vec<NodeId>,
    /// Nodes consumed by a later split on another backend (copied out).
    pub outputs: Vec<NodeId>,
}
}

A small ASCII example — three nodes assigned CPU → Metal → CPU:

nodes:     0(input x, CPU)   1(silu, Metal)   2(add s+x, CPU)
split 0:   [0..1) CPU          outputs: [0]  (x is read by the Metal split)
split 1:   [1..2) Metal        inputs:  [0]  outputs: [1]
split 2:   [2..3) CPU          inputs:  [1, 0]
boundary 0→1: sync CPU (no-op) · copy x  CPU→Metal
boundary 1→2: sync METAL (submit+wait) · copy silu-out Metal→CPU · (x already CPU: no copy)
end:          sync CPU (no-op)

(That is the split_on_backend_change unit test, src/graph/scheduler.rs L487–503: it asserts splits[1].inputs == vec![0], splits[2].inputs == vec![1, 0].)

Why contiguous runs, rather than "all nodes of backend X, wherever they sit"? Both answers are about count: each boundary costs a synchronize plus copies, and contiguity makes the boundary count equal the number of backend alternations along the list — the minimum possible for a given assignment. A split is also the natural submission unit for an async backend: one command buffer encodes end − start kernel launches and pays one submit + one wait at the boundary (per-node units would multiply round trips). §3.3 revisits this choice.

Consequences for the two trivial shapes: an all-CPU graph is one split (no backend change → no boundaries → no copies, no syncs — prev_backend never differs), and an all-Metal graph is likewise one split. In practice minfer runs are overwhelmingly single-split: Metal and CUDA are enabled all-or-nothing per model — a GPU backend runs only when every graph weight is registered on it (Qwen2Graph::weights_on_gpu; "same all-or-nothing rule as Metal", src/graph/cuda_backend.rs L1296–1297). Mixed multi-split graphs are real but rare — they occur when one backend rejects an op the other supports (CUDA accepts only the non-interleaved RoPE style, L1298) — and the code still handles them matter-of-factly (scheduler.rs L243–258 analyzes "x goes to a Metal silu split and a CPU add split").

2.4 The split protocol

execute is a loop over splits, and each iteration follows the same four-beat protocol:

  1. Sync the previous backend (alloc.sync_backend(pb)). Everything the previous split enqueued is now finished: its output buffers hold final values and are safe to read from the host.
  2. Copy this split's cross-backend inputs (alloc.copy_across(inp, split.backend) for each inp in split.inputs). Each copy is a host round trip: read the producer's buffer to host memory, then write host memory into a staging buffer (scratch memory holding a value in transit) in the consumer backend's pool.
  3. Run the split's nodes in order via the backend's execute_node. CPU computes inline; Metal encodes into the split's command buffer; CUDA launches onto its stream.
  4. Flush — at the next boundary (or the very end), sync this backend and read back any observability captures queued during the split.

After the loop, the same protocol runs once more with no copies: sync the last backend, so that when execute returns Ok(()) every buffer in the graph — including the logits — holds its final value and is safe to read. That is the guarantee doc 09's prefill code relies on when it reads the logits right after execute.

One definition borrowed from GPU_SAFETY explains why beat 1 must precede beat 2: a barrier is an explicit ordering guarantee making writes from one piece of GPU work visible to work that follows. Across backends, minfer's barrier is the split-boundary synchronize (submit + bounded wait); inside one Metal encoder, the backend inserts memoryBarrierWithScope between dispatches, because Metal does not promise cross-dispatch write visibility on its own (docs/GPU_SAFETY.md §3). The scheduler never thinks about the second kind; the backend does.

2.5 The error contract: abort, never silently fall back

Every execute_node returns Result<(), String>. When a kernel's preconditions do not hold — a dimension mismatch, an unregistered weight, a KV position beyond capacity — the backend returns Err(...), the scheduler propagates it with ?, and the whole run aborts. It never quietly re-runs the node on the CPU. The reason is numerical, not aesthetic:

  • CPU matmuls quantize activations to Q8_0 on the fly (a per-32-value integer block format with one scale, doc 10); the GPU backends read activations as f32. The two paths produce different numbers by design (AGENTS.md rule 9, ARCHITECTURE.md §4.5 invariant 6). A mid-run fallback would splice Q8_0-rounded numbers into an otherwise f32 stream — every downstream value subtly wrong, with nothing to notice but a degraded, hard-to-trace output.
  • Backend placement is a build-time decision (doc 06, via supports_op); execution-time failures are invariant violations — bugs in the assignment or the kernel — and bugs should be loud. docs/GPU_SAFETY.md §2.3 and CUDA rule 1 (L203) state this as a hard rule. The call site makes "loud" literal: sched.execute(graph, alloc).unwrap(); (src/models/qwen2/graph.rs L549–550) turns an Err into a panic, and the server wraps the forward so an unexpected abort becomes a 500 instead of killing the worker thread (src/server/chat.rs L234–247).

3. Implementation

3.1 Data in / data out

In:

  • graph: &ComputeGraph — nodes in build order, each with op, src (ids of producer nodes), out_shape/out_dtype, backend: Option<BackendTag> (set by doc 06's assign_backends), and meta (weight names, RoPE/attention parameters).
  • alloc: &mut GraphAllocator — owns the per-backend buffer pools; maps every live node to a buffer (node_to_buf); holds the two persistent KV regions per layer (resolved by kv_pair(layer) → (k_id, v_id)); holds the cross staging map for cross-backend copies. Input buffers are already host-filled (fill_input_i32, doc 07).

Out:

  • Every allocated node buffer contains its computed value; the KV regions contain this step's K/V rows at the positions given by the positions input; the graph's output buffers (logits) are readable on the host.
  • The Result verdict: Ok(()) means "all splits ran and all backends synced"; Err(msg) means an invariant failed and the run is aborted.

Shapes worth keeping concrete (Qwen2.5-0.5B: n_embd = 896, vocab 151936, f32 = 4 B): a hidden-state buffer for one decode token is 896 × 4 = 3,584 B (~3.5 KB); for a 512-token prefill it is ~1.8 MB; the logits buffer is 151936 × 4 B ≈ 608 KB (the size docs/cuda_optimization_steps/07-r3-small-model-overhead.md L79 works with). Cross-backend copies move buffers of exactly these sizes — which is why "how many splits exist" is a performance question, not just a correctness one.

3.2 Key code

The split walk — boundaries first

The heart of execute is one loop over the splits, with the boundary protocol in its first lines:

#![allow(unused)]
fn main() {
// src/graph/scheduler.rs L176-189
for split in &splits {
    if let Some(pb) = prev_backend {
        if pb != split.backend {
            // 1. flush the previous backend's async work
            alloc.sync_backend(pb);
            // 1b. staged Metal/CUDA captures are valid now — read back
            flush_metal_captures(graph, alloc, &mut staged, trace_on, live_on);
            flush_cuda_captures(graph, alloc, &mut cuda_caps, trace_on, live_on);
            // 2. copy this split's inputs across backends
            for &inp in &split.inputs {
                alloc.copy_across(inp, split.backend)?;
            }
        }
    }
}

Beat by beat: sync_backend dispatches on the backend tag (CPU is a no-op — src/graph/alloc.rs L542–562; Metal submits the pending command buffer; CUDA closes any capture window and stream-syncs). Only after that may the captures be read (their staging writes just landed) and copy_across run (it reads the producer's buffer host-side — guaranteed final now). If the previous split was on the same backend, none of this runs: same-pool buffers need no copies and no sync.

The CUDA replay hook — skip the whole split

Before the node loop, one backend-specific fast path: a captured CUDA Graph (a recording of every kernel launch of this split, replayable as a single launch) can replace the entire node walk:

#![allow(unused)]
fn main() {
// src/graph/scheduler.rs L194-213 (cfg attributes abridged)
let replayed = if capture {
    false
} else {
    match split.backend {
        BackendTag::Cuda => {
            let c = alloc.cuda_mut().ok_or("CUDA backend not enabled")?;
            c.graph_replay(graph.uid, split.node_range, graph.capture_nt_hint())
        }
        _ => false,
    }
};
if replayed {
    // the captured launch covers every node of this split (the
    // boundary sync of the NEXT split still closes it out)
    prev_backend = Some(split.backend);
    continue;
}
}

true = "I handled this split": the scheduler continues past all of the split's nodes, and the replayed work — stream-ordered like any other launch — is drained by the next boundary's sync. The replay is keyed on (graph.uid, split.node_range): uid is the graph's identity for reuse (src/graph/mod.rs L111–115), so step 50 replays the recording made at step 3. And capture (trace/viz per-node readback) forces replayed = false: a readback inside a recorded capture would be recorded into the graph and corrupt it — the same rule as GPU_SAFETY CUDA rule 2. The capture mechanics live in graph_replay_step (src/graph/cuda_backend.rs L184–244): executions 1–2 of a split run direct (warmup), the 3rd records, later ones replay; the nt_hint gate (capture_nt_hint, src/graph/mod.rs L126–132) restricts capture to decode-shaped graphs. Docs 14/15 cover the rest; this hook is all the scheduler sees.

The node loop — three skip rules, then dispatch

#![allow(unused)]
fn main() {
// src/graph/scheduler.rs L225-241 (input capture under trace omitted)
if node.is_input() {
    continue; // data pre-filled by the allocator
}
// dead nodes (no consumers, not outputs) get no buffer — the
// fusion pass can orphan them (e.g. silu folded into SwiGLU);
// they are skipped, not executed
let Some(br) = alloc.node_buffer(id) else {
    continue;
};
if br.backend != split.backend {
    let op_full = format!("{:?}", node.op);
    let op = op_full.split(['(', '{']).next().unwrap_or("?");
    return Err(format!(
        "node {id} ({op}) buffer on {:?} but executing split is {:?} (assignment/alloc mismatch)",
        br.backend, split.backend
    ));
}
}

Three ways a node can be not executed, in order. First, it is an input — no math; the allocator filled its buffer before execute was called (doc 07's fill_input_i32); under trace/viz capture its (host-written, therefore current) data is recorded at the top of the loop, before the continue. Second, it has no buffer — the allocator only maps nodes with consumers or output status; when doc 06's fusion pass folded Mul(Silu(x), y) into SwiGLU, the original silu node may survive in the list as a dead orphan, and it is skipped, not executed — build order stays intact and the dead weight costs one map lookup. Third, it is an error: a buffer on a different backend than the split it landed in means assignment and allocation disagree — a bug, and it gets a descriptive Err naming the node and both backends (split_graph derives splits from node.backend, so this "cannot happen"; if it ever does, it fails loudly instead of writing a buffer with the wrong kernel).

Input resolution: which buffer is "the" input?

For each of the node's src producers, the scheduler picks the buffer the kernel should read:

#![allow(unused)]
fn main() {
// src/graph/scheduler.rs L242-258
let mut in_bufs = Vec::with_capacity(node.src.len());
for &s in &node.src {
    // a cross-backend staging copy (split boundary) takes
    // precedence only when it was made FOR this split's backend:
    // a node feeding two different backends leaves one stale
    // cross-buffer (e.g. x goes to a Metal silu split and a CPU
    // add split — cross_buffer(x) ends up Metal), which must not
    // be read by the CPU consumer. Otherwise fall back to the
    // node's canonical buffer (already on the split's backend if
    // no copy was needed for it).
    let sbr = alloc
        .cross_buffer(s)
        .filter(|cb| cb.backend == split.backend)
        .or_else(|| alloc.node_buffer(s))
        .ok_or_else(|| format!("node {s} has no allocated buffer"))?;
    in_bufs.push(sbr.id);
}
}

The rule in one sentence: prefer a cross-backend staging copy if one was made for this split's backend, otherwise read the producer's own buffer. The subtlety (the comment's scenario): copy_across keeps one staging buffer per (node, destination) pair, so a producer feeding two consumer backends leaves behind a staging buffer for one of them; a consumer on the other backend must not read it. The .filter makes the stale entry invisible to the wrong split.

Resolving kv_pair, then dispatching

KV ops need the layer's two persistent regions (doc 07). The scheduler resolves them before borrowing the backend mutably, because the backend — not the scheduler — knows which sibling buffer an op writes:

#![allow(unused)]
fn main() {
// src/graph/scheduler.rs L259-292 (cfg attributes abridged)
// resolve the layer's KV region pair BEFORE the mutable backend
// borrow (the backend needs it for KV store / attention)
let kv_pair = match &node.op {
    Op::KvcacheStore { layer } => alloc.kv_pair(*layer),
    Op::FusedQKV { layer } => alloc.kv_pair(*layer),
    Op::QkvBiasRopeStore { layer } => alloc.kv_pair(*layer),
    Op::FusedQkvNorm { layer } => alloc.kv_pair(*layer),
    Op::Attn { .. } => match &node.meta {
        NodeMeta::Attn(m) => alloc.kv_pair(m.layer),
        _ => None,
    },
    _ => None,
};
// NOTE: execution follows node id order (build order), which is
// the graph's topological order by construction.
match split.backend {
    BackendTag::CPU => alloc
        .cpu_mut()
        .execute_node(node, &in_bufs, br.id, kv_pair)?,
    BackendTag::Metal => {
        let m = alloc.metal_mut().ok_or("Metal backend not enabled")?;
        m.execute_node(node, &in_bufs, br.id, kv_pair)?;
    }
    /* ... Cuda arm: the same call on alloc.cuda_mut() ... */
}
}

This is the Backend trait contract in use — every backend implements the same four-argument call, and ? propagates any kernel-invariant failure straight out of execute (§2.5). The trait itself:

#![allow(unused)]
fn main() {
// src/graph/backend.rs L49-70
fn execute_node(
    &mut self,
    node: &CNode,
    in_bufs: &[usize],
    out_buf: usize,
    kv_pair: Option<(usize, usize)>,
) -> Result<(), String>;
/* ... read_host / write_host ... */
/// Wait for async work to complete (CPU: no-op; Metal: submit the pending
/// command buffer). Called between splits and after the last split.
/// A backend that captured a CUDA Graph window for the current split must
/// close it here (capture records launches without executing them).
fn synchronize(&mut self);
}

in_bufs[i] is the pool-local id of the buffer for node.src[i]; out_buf is the pool-local output id; kv_pair is Some((k_id, v_id)) only for the KV-touching ops listed above (attention uses it to find the V region while iterating the K view — doc 11). Note what the contract does not contain: no shapes-in, values-out signature — buffers are identified, not passed, and the data never moves for a normal op.

The tail after the loop is the same boundary protocol one last time, minus the copies (src/graph/scheduler.rs L346–353): sync the last prev_backend, flush the staged captures, return Ok(()). Everything enqueued by the whole graph is done; the logits sit on a host-readable buffer; the caller (doc 09) can read them immediately.

What a copy across backends actually is

#![allow(unused)]
fn main() {
// src/graph/alloc.rs L575-582 + L608-611 (doc comment trimmed)
pub fn copy_across(&mut self, node_id: NodeId, dst_backend: Backend) -> Result<(), String> {
    let br = self
        .node_buffer(node_id)
        .ok_or_else(|| format!("node {node_id} has no buffer"))?;
    if br.backend == dst_backend {
        return Ok(());
    }
    /* one staging buffer per (node, dst) pair is reused when present, then: */
    let data = self
        .copy_to_cpu(node_id)
        .ok_or_else(|| format!("node {node_id} host read failed"))?;
    let new_id = self.alloc_fresh_in(dst_backend, data.len());
}

The function's own doc comment (L564–574) names the mechanism: "host round trip through read_host/write_host (shared-memory GPU buffers make this a plain memcpy both ways)", with the copy landing in the cross staging map — the node's canonical buffer is left untouched, and consumers resolve it via cross_buffer (the mechanism the previous excerpt showed). So: device → host (a plain host-pointer read for Metal's shared-memory buffers, a staged cudaMemcpy-style read for CUDA) → host → device (write_host into a fresh staging buffer). The early Ok(()) for same-backend pairs is why split.inputs can be copied without pre-filtering: in the §2.3 example, copy_across(node 0, CPU) is a no-op even though node 0 appears in split 2's inputs. Staging buffers are allocated fresh (bypassing the free-list recycle) because a recycled buffer could still be referenced by nodes that execute later in the same run (Backend::alloc_fresh doc, backend.rs L36–42).

CPU execution: one representative arm

CpuBackend::execute_node (src/graph/cpu_backend.rs L133) is a big match on the node's op, dispatching to the handwritten kernels of kernel.rs / vec_ops.rs. Before the match, two pieces of plumbing make every arm safe: the KV store is handled first (it needs mutable access to two regions at once, L143–177), and aliased in-place inputs (liveness reuse may map out_buf onto an input — doc 07) are snapshotted so the output can be carved out of the pool with split_at_mut (L179–201). Then the arms; the one that matters most for the rest of this series is MatMul:

#![allow(unused)]
fn main() {
// src/graph/cpu_backend.rs L265-299 (bias loop abridged)
Op::MatMul { .. } => {
    let meta = match &node.meta {
        NodeMeta::MatMul(m) => m,
        other => return Err(format!("matmul node missing MatMulMeta: {other:?}")),
    };
    let w = self
        .weights
        .get(&meta.weight_name)
        .ok_or_else(|| format!("weight '{}' not registered", meta.weight_name))?;
    // llama.cpp/GGUF convention: metadata [in, out], memory [out][in]
    let od = w.shape[1] as usize; // output dim
    let id = w.shape[0] as usize; // input dim
    let nt = node.out_shape[1];
    if w.ttype == crate::tensor::TensorType::F32 {
        // plain f32 matmul: out[t*od+o] = dot(w[o], x[t])
        crate::vec_ops::mat_mul_f32(od, nt, id, out, w.data_f32(), ins[0]);
    } else {
        // quantized weight × f32 activations (Q8_0-quantized on the fly)
        kernel::cpu_quant_matmul_f32(w, ins[0], out, od, id, nt);
    }
    if let Some(bname) = &meta.bias_name { /* per-row bias add */ }
    Ok(())
}
}

Everything about the CPU path the later docs (10, 11) zoom into is visible here: weights resolve by name out of the backend's own registry, the GGUF [in, out] metadata convention fixes which shape index is which, and the f32-vs-quantized fork is exactly where cpu_quant_matmul_f32 (doc 10's subject — quantized weight rows × on-the-fly Q8_0 activation blocks) takes over. The elementwise and normalization ops are the same shape but simpler — Op::RmsNorm { eps } loops rows and calls crate::vec_ops::rms_norm_fused_f32(d, dst, row, w.data_f32(), *eps) (L223–241); Op::Silu is one vec_silu_f32 call. The two Err arms right here show the error contract on CPU too: a missing MatMulMeta or an unregistered weight aborts rather than improvising. (The KV store arm — the code that writes the KV regions using kv_pair and the positions input — is excerpted in doc 11, where it belongs.)

Metal execution: encode now, compute later

The Metal execute_node is structurally identical — same match, same arms — but every arm ends in a command-buffer encode, not a computation:

#![allow(unused)]
fn main() {
// src/graph/metal_backend.rs L323-341 (abridged)
let cb = self.cb();          // the split's one command buffer
/* ... */
match &node.op {
    Op::Input => Ok(()),
    Op::Silu => {
        if in_bufs[0] != out_buf {
            self.copy_in(out_buf, in_bufs[0]);
        }
        let n = self.pool[out_buf].length() as usize / 4;
        cb.silu_f32(self.buf(out_buf), n);   // enqueue, don't compute
        Ok(())
    }
    /* ...add, mul, rms_norm, matmul, attn, ... all encode into cb... */
}

self.cb() (L146–157) lazily creates the split's single MpsCommandBuffer on the first op and hands it to every subsequent op — that is the "one command buffer per split" rule from AGENTS.md (rule 8) living in code. Each cb.* call appends a kernel dispatch; nothing runs on the GPU yet. The actual submission happens in synchronize (L1023–1024), which just calls submit_pending (L159–186): take the box, cb.submit(), clear the pointer.

submit() (src/metal/ L1935–1978) is where the GPU_SAFETY rules get enforced. It commits the command buffer, then waits on a dispatch semaphore with a 10-second timeout and checks the final MTLCommandBufferStatus: Completed → Ok(()); any other status → Err with the recent dispatch labels attached; timeout → Err("Metal command buffer timed out after 10s (GPU hang)"). The history is a real machine freeze (GPU_SAFETY §1, 2026-08-02): the old submit waited forever on the semaphore and never checked status, so one GPU fault hung the whole process — and, because Metal clients share the GPU, threatened the machine. Now a hang costs at most 10 seconds and produces a diagnosable error (the last 16 dispatch labels, recorded when MINFER_TRACE is on, identify the faulting kernel family). Doc 14 shows the full submit path.

What synchronize actually guarantees

Because the whole async story funnels into this one call, it is worth stating its guarantee precisely. After alloc.sync_backend(b) returns for backend b:

  1. Every node of every split already run on b has finished executing. Metal: its pending command buffer was committed and its completion handler fired with status Completed (or the wait timed out → error). CUDA: the stream was synchronized (and any open capture window was closed by instantiating + launching the recorded work). CPU: nothing was ever pending.
  2. Therefore every buffer written by b holds its final value — safe to read from the host (trace readbacks, copy_to_cpu), and safe for another backend to read via a host round trip: copy_across reads host-side precisely at this point in the protocol. Note the division of labor: synchronize does not copy across backends, and copy_across does not synchronize — the pairing is the scheduler's, and the invariant "never host-copy a GPU-pending buffer" (ARCHITECTURE.md §4.5 #4, from a real Phase-3 KV-corruption bug) is enforced by that ordering.
  3. It says nothing about backends never synchronized — which is why the scheduler syncs on every boundary and once after the last split, and why a CPU-only build "never calls it" (the trait doc's note, backend.rs L62–64): a synchronous backend has no gap to close.

3.3 Design choices (why this shape and not another)

Why execute in build order instead of re-sorting topologically? Four reasons stack up:

  1. It is already a topological order — the builder appends sources before consumers (doc 05). Re-sorting would compute, per execution, a permutation the builder already guaranteed.
  2. One order everywhere is a correctness feature, not a convenience. The allocator's liveness must predict exactly which buffers are still needed at each step of execution (doc 07). This is not hypothetical: the G3 bug was exactly that disagreement — the Kahn order moved source-less nodes (like kv_load) earlier, liveness concluded a buffer was dead before the build-order executor had read it, attention reused the residual buffer, and the tail get_rows read the attention output instead of the residual — logits off by 21.79 (docs/COMPUTE-GRAPH-DESIGN.md §22, fixed in commit 8febf4c). The fix was not "sort better"; it was "everyone uses build order".
  3. Store-before-attention for free. With explicit KV nodes (doc 05) plus build order, the write-before-read property of KV needs no scheduling logic at all — it falls out of list order.
  4. It matches ggml. llama.cpp executes ggml_cgraph.nodes[0..n] in order; the scheduler's module doc (L1–15) says so explicitly: "matching ggml, which executes nodes[0..n] in order".

Why is a host round trip acceptable for cross-backend copies? Because the case is rare, the buffers are small, and the alternative is a lot of plumbing. Splits are rare (§2.3: all-or-nothing GPU gates make single splits the norm; typically zero to two cross-boundary copies per forward), and the values that cross are small — a decode-step hidden state for 0.5B is 3.5 KB, and even a 512-token prefill's hidden state is ~1.8 MB, one memcpy each way on Metal where GPU buffers live in shared host-visible memory (the copy_across doc: "shared-memory GPU buffers make this a plain memcpy both ways"). The alternative — device-to-device peer-to-peer copies — would need each backend to expose cross-device primitives (and CPU↔GPU are not peer devices at all; the host is the intermediary), plus allocator plumbing to track peer mappings, for a code path that runs zero or a handful of times per forward. The one real cost of the round trip — the sync it requires — is a cost the correctness protocol already pays: nothing may read a GPU buffer before its split syncs.

Why must execute_node take &mut self? Two independent reasons, both visible in the code:

  1. The backend mutates its own state to do the work. The CPU backend writes results into its buffer pool (split_at_mut over self.buffers, src/graph/cpu_backend.rs L191–201) and mutates the pool in the KV-store arm (L160–166). The backend trait doc says this outright (backend.rs L3–5: "execute_node takes &mut self (the CPU backend mutates its own pool)").
  2. Asynchronous backends mutate encoding state per call. Metal appends to the split's command buffer and lazily creates it (self.cb() mutates cb_ptr, metal_backend.rs L146–157); CUDA tracks capture windows and launch bookkeeping. "Encode one op" is a stateful operation on the backend, not a pure function on the node.

The &mut self is also a safety choice: it is a compile error to execute a node while something else borrows the backend's pool — which is why the scheduler resolves kv_pair from the allocator before borrowing the backend mutably (scheduler.rs L259–261).

Why are splits contiguous runs rather than per-backend op groups? Because the boundary costs (sync + copies) scale with the number of boundaries, not with the number of nodes moved — contiguity minimizes boundaries for a given assignment (§2.3) and keeps split_graph a single O(n) scan plus an O(n·src) pass to derive cross-split edges: no graph reordering, no risk of changing execution order.

Why derive cross-split inputs/outputs from src edges instead of annotating them at assignment time? Because assignment (doc 06) and the fusion rewrites happen before this stage and both can change which nodes feed which; deriving from the final src lists at split time keeps the boundary set consistent with the graph actually being executed — which is what lets the split_on_backend_change test (L487–503) assert the exact input/output vectors.

3.4 Pitfalls & invariants

  • One order everywhere (invariant 5). Execution, liveness, and the "buffer may be freed" logic all use build order; topo_order() only proves acyclicity. Origin: the G3 regression — 21.79 of logit drift traced to an executor/allocator order mismatch, fixed in 8febf4c together with "inputs are never freed" (COMPUTE-GRAPH-DESIGN §23: two inputs whose liveness-shared buffer got refilled by the later input's fill, clobbering the first — the embedding_and_rope case).
  • Never host-copy a GPU-pending buffer (invariant 4's corollary). A host read of a buffer whose producer kernel is encoded-but-not-submitted reads stale bytes. Every host-side read in the split protocol sits after a boundary sync. Origin: the Phase-3 KV-corruption bug (ARCHITECTURE.md §4.5 #4). The trace/viz capture code is arranged around the same trap: CPU outputs are read immediately (already final), Metal outputs are blitted to staging and read after the split's submit, CUDA outputs queue a stream-ordered async D2H drained at the boundary (scheduler.rs L293–330).
  • Dead fusion orphans are skipped, not executed. A node without a buffer would crash a naive pool[buf_id] — the let Some(br) = alloc.node_buffer(id) else { continue } is a correctness requirement, not a nicety.
  • Stale cross-buffers. One staging buffer per (node, dst backend) means a producer feeding two backends leaves an entry that is wrong for the other backend; the .filter(|cb| cb.backend == split.backend) at input resolution (L252–256) is the guard. Removing it reads valid-looking but wrong-split data — the kind of bug that survives small graphs and fails at multi-backend ones.
  • Kernel-invariant violations abort (invariant 6's enforcement). Silent CPU fallback would quantize activations to Q8_0 mid-stream and corrupt the numerics (§2.5). The error path is: backend Err → scheduler ? → forward_graph_cached .unwrap() → abort (or HTTP 500 in the server). Assignment is build-time; runtime surprises are bugs.
  • Bounded GPU waits, always. Metal submits wait ≤10 s and check status (GPU_SAFETY §2.1); CUDA syncs poll cudaGetLastError plus stream state (GPU_SAFETY CUDA rule 4). The scheduler never blocks on a GPU without a timeout in the path beneath it.
  • No host readbacks inside a CUDA capture window. The scheduler disables per-node capture (and therefore replay) whenever trace/viz capture is on (L190–193); a readback inside a recorded launch sequence corrupts the capture — GPU_SAFETY CUDA rule 2, learned from the 7e② "faster but wrong" incident. (The capture vectors metal_srcs/staged/cuda_caps are declared unconditionally with empty non-macOS/non-CUDA flush stubs, L167–175, L444–456, so no #[cfg] maze forks the main loop.)

4. Observe & verify

  • MINFER_GRAPH_TRACE (any value) — at the top of execute (scheduler.rs L127–145) prints one line per split ([graph] split 0: Metal nodes 0-412) plus a per-op/per-backend node-count table: how many splits your model produced and which ops went where.
  • MINFER_TRACE=<path> — per-node real-data trace for the web visualizer: after each node executes, its output buffer is analyzed (min/max/mean/abs-mean stats + downsampled values) and recorded per step (src/trace.rs; the hook is record_node_data, scheduler.rs L369–389). KV-region nodes are skipped (up to n_embd × n_ctx per layer — huge). Load the file at viz/index.html to watch the graph compute, node by node; details in viz/README.md (no doc 17 exists; that README is the reference).
  • Live viz — minfer viz <model> serves the same graph page with a live SSE feed: crate::live::enabled() (src/live.rs L51–56) is checked once per execute() and flips the same capture flag, so nodes light up as each split executes. Again: viz/README.md.
  • MINFER_OP_PROFILE — Metal-side: per-op host encode times plus per-submit GPU wait times (first submit prints a full table; metal_backend.rs L224–242). Shows the encode/execute split of §2.2 in real microseconds.
  • Tests — src/graph/scheduler.rs L458–519 covers this stage end to end: assign_all_cpu_and_single_split (all-CPU ⇒ exactly one split, no cross edges), split_on_backend_change (CPU→Metal→CPU ⇒ three splits with the derived inputs/outputs of §2.3), and execute_single_backend_graph (computes silu(x) + x through the full allocate→fill→execute path). The backend suites pin cross-backend copies (copy_across_cpu_to_cuda_and_back, cuda_backend.rs L1477) and replay parity (cuda_backend.rs L6159).
  • MINFER_GRAPH_DUMP=<dir> — dumps the logits and layer-0 KV after execute, for CPU-vs-GPU comparison of what this stage produced (src/models/qwen2/graph.rs L552).

5. Cross-references

  • docs/ARCHITECTURE.md §4.3 (the assign→fuse→alloc→execute pipeline), §4.5 (invariants 4–6), §5.1 (the Backend trait list), §5.3 (GPU safety mapping).
  • docs/GPU_SAFETY.md — the hard rules this stage enforces: bounded submit + status check (§2.1), Err-not-fallback (§2.3), no sync inside a capture window (CUDA rule 2), stream order as the async contract (CUDA rule 5).
  • docs/COMPUTE-GRAPH-DESIGN.md §16/§22/§23 (L948–955) — the G3 tail-row optimization and the two liveness-vs-order bugs (21.79 logit drift; input-buffer clobber) that made build order the one order.
  • 05 — Graph build: why the node list is topological by construction, and why KV ops are explicit nodes.
  • 06 — Assign + fusion: where node.backend comes from and which fused ops exist (the source of dead orphans).
  • 07 — Allocator: buffer pools, liveness sharing, the allocator side of the cross staging map, KV regions.
  • 09 — Prefill: the caller that reads logits immediately after execute returns — the consumer of this doc's guarantee.
  • 10 — CPU matmul kernels and 11 — Attention + vec ops + KV: inside cpu_quant_matmul_f32, and the KV-store/attention arms that the kv_pair contract feeds.
  • 14 — Metal backend / 15 — CUDA backend: the full story of the one-command-buffer-per-split discipline and CUDA Graph capture/replay that this doc only hooks into.
  • viz/README.md: the trace format, the live SSE view, and how to load a run into the graph visualizer.

← 07 — Memory allocation: liveness and the KV regions · Index · 09 — Prefill: the first forward →

09 · Prefill: the first forward

Stage: the graph machinery is ready — build, assign, fuse, allocate, execute (docs 05–08) → this stage: the CLI actually calls it, for the whole prompt at once, and gets logits back → the math inside the graph (10 — CPU matmul kernels, 11 — Attention + vec ops + KV) and the first sample (12 — Sampler). Code: src/main.rs prefill block (L1441–1485, timing prints L1480–1485 and L1848–1865), the call chain src/models/mod.rs::forward (L222) → src/models/qwen2/mod.rs::forward (L55) → src/models/qwen2/graph.rs::forward (L432) → forward_cached (L450), src/graph/params.rs (GraphParams L88), src/trace.rs::begin_phase (L70) — lines verified at commit 15fa45c.

1. Background — where this stage sits

Docs 01–04 got your typed prompt turned into a list of integer token ids. Docs 05–08 built the machinery that can run a transformer forward pass — a forward pass is one trip through the whole network: embeddings, then every transformer layer (attention + feed-forward), then a final projection — without computing anything. The machinery is a compute graph: a data structure that lists every math operation as a node, knows which backend (CPU, Metal, CUDA) runs each node, where each result lives in memory, and in what order the nodes execute. What it does not have yet is data to chew on and a caller. This document is that caller: the ~30 lines of main.rs that say "here are the prompt's token ids, please run the graph once over all of them, and give me the answer."

The answer has a name: logits. The last operation of the network — the lm_head matrix multiply — produces one floating-point score per entry of the model's vocabulary (the list of all tokens the model knows; about 151,936 for the Qwen models). A score is not a probability yet — it is an unnormalized "how plausible does each next token look" number, bigger = more plausible. The sampler (doc 12) turns these scores into one chosen token. For a 512-token prompt the raw scores occupy 151,936 × 4 bytes ≈ 608 KB — roughly half a megabyte of "opinions" per forward.

This specific forward is called prefill: the model reads all prompt tokens in one graph execution, computing every token's hidden state and — importantly — writing every token's K/V (key/value, the attention memory) into the KV cache, the per-layer notepad of past attention states. Contrast that with decode, every later step of generation: one new token per forward, reusing the cached graph. The prefill/decode split is the two-act structure of the whole engine, and most of this document is about the decisions the boundary forces: why all prompt tokens go through together, why only the last token's logits are wanted, why the context size is fixed once for both phases, and why the logits come back as an owned Vec<f32> rather than a borrowed slice.

What would break without this stage working as it does? Three things. If the prompt went through token-by-token, the engine would stream every weight byte from memory once per token instead of once per prompt — prefill would be hundreds of times more memory traffic for the same arithmetic. If earlier positions computed full logits too, the biggest matrix in the model would run 512 times more work than needed. And if prefill and decode disagreed about the KV region size, the cache the prefill just filled would have to be copied — or silently misplaced — before the first decode step could read it. Each of these is a design decision with real arithmetic behind it; §2 and §3.3 walk them.

One more orientation point. The single-shot CLI path this doc follows is the simplest caller: one sequence, one user prompt, one sample per step. The conversation REPL (--cnv) and the HTTP server (serve) call the very same forward with their own bookkeeping (append-only KV, per-slot caches); doc 13 returns to them. Everything here — the prefill block, the params it builds, the logits it receives — is the shared spine of all three.

2. Principle — how it works and why

2.1 Four words, defined once: forward, prefill, decode, context

  • Forward pass — one execution of the graph: tokens in, hidden states transformed layer by layer, logits out. minfer runs one forward per GraphParams shape; the graph itself is cached and replayed (docs 05/13).
  • Prefill — the first forward of a run, feeding all prompt tokens at once (n_tokens = prompt length). It computes two things at once: the hidden state of every prompt position (needed so the last position's prediction is informed by all of them), and the K/V entries for every prompt position (the KV cache's initial content).
  • Decode — every forward after that: one token in (the one just sampled), whose K/V appends to the cache, and whose logits pick the next token. Same graph skeleton, n_tokens = 1.
  • Context — how many token slots the KV cache has room for (n_ctx, the CLI flag --n-ctx, default 4096). It is a memory reservation, not a speed knob: it decides how long a conversation can grow before the notepad is full.

2.2 Why a positions vector exists at all

The prefill block starts with one unassuming line:

#![allow(unused)]
fn main() {
let positions: Vec<usize> = (0..input_ids.len()).collect();
}

Prompt token i gets position i. Why carry that in a vector — can't the code just know "the i-th token of the prompt is at position i"? It can't, and the reason is the central invariant of the graph design (docs 05/07): KV positions are data, not structure. The graph's topology never depends on where in the cache a token lands. Instead, every execution fills a positions input buffer, and three different kernels read it as ordinary data:

  • the kvcache_store nodes use it to decide which KV rows to write (position p writes row p of the layer's K and V regions);
  • the attn node uses it for causal masking — token t may only attend to positions 0..=pos[t] (the builder's own words: "pos carries the per-token write positions (I32 input), needed for causal masking (vl = pos[t]+1)", builder.rs:325-327);
  • the rope nodes use it as the rotation angle's input (position determines the angle — doc 11).

This indirection is what lets one graph topology serve prefill and decode: prefill fills positions with 0, 1, 2, …, nt-1; the decode loop fills it with one running number (current_pos, starting at the prompt length, main.rs:840, incrementing per step, main.rs:940). If positions were instead baked into the topology (say, attention shaped to read exactly n_past + nt cache rows), the graph would have to be rebuilt on every decode step, because n_past changes every step — and the params-only reuse scheme of doc 05 §2.7 would collapse. The positions vector is the price of reusability, paid in a few bytes of input data.

There is a second, subtler payoff: decoupling slot from loop iteration means the slot need not equal "how many tokens came before in this call". The conversation path (doc 13) re-prefills a new user turn at positions that continue from the previous turn; the server (one GraphCache per slot) gives each slot its own position range. Same graph, different data.

2.3 One forward for the whole prompt — the parallelism argument

The obvious naive design is: run the graph once per prompt token, feeding it tokens 0..=i each time, i.e. token-by-token prefill. minfer (like llama.cpp) does the opposite: one forward with n_tokens = nt. The reason is arithmetic about what a matmul is.

Every weight matrix in the model is read from memory during any forward — that part is unavoidable; the weights are the model. The question is how much useful work each byte of weight buys. In a decode-shaped forward (nt = 1), each weight byte participates in exactly one multiply-accumulate for the one token being processed: the forward is dominated by reading the weights, not by math. In a prefill-shaped forward (nt = 512), the same weight byte is reused for all 512 columns of the activation matrix — 512 multiply-accumulates per weight element. Same memory traffic, 512× the arithmetic. In roofline language: decode is memory-bandwidth-bound (the bottleneck is streaming the weights), prefill is compute-bound (the bottleneck is the ALUs/SIMD lanes, which are now saturated). That single ratio — arithmetic per weight byte scales with nt — is why prefill throughput is measured in thousands of tokens/second and decode in hundreds, on the same hardware, and docs 10/14 build their kernel strategies directly on this split.

Token-by-token prefill would throw that away: 512 sequential decode-shaped forwards stream the weights 512 times. Concretely for Qwen2.5-0.5B (weights total ≈ 0.5 × 10⁹ elements, mostly in quantized matmul weights): one batched 512-token prefill reads each weight element once and does ≈ 5 × 10¹¹ multiply-accumulates; the token-by-token version does the same math but pays 512× the weight traffic, and then still could not beat the batched version even with infinite bandwidth, because each of the 512 walks also repeats the per-node dispatch overhead of a 440-node graph (doc 05's census) 512 times.

Two more dividends of batching, beyond bandwidth:

  1. The KV cache is filled in one pass. The kvcache_store nodes write all nt positions contiguously (0..nt) in the same execution that computes attention over them. Token-by-token would interleave 512 store passes with 512 attention reads, with no benefit — attention for token i needs exactly the rows 0..=i, which are all available in the batched version the moment each layer's store node runs (build order guarantees store-before-attention, doc 08).
  2. One build/assign/fuse/allocate for the prompt. Each distinct GraphParams pays the graph-build pipeline once (doc 05 §2.7). One prefill forward = one build; the decode graph that follows is a second build (because n_tokens changed), and then hundreds of decode steps replay it for free.

The costs of batching are real but bounded: activation buffers scale with nt (a hidden-width buffer on 0.5B is 896 × nt × 4 B ≈ 1.8 MB at nt = 512, doc 07 §2.3), and attention grows quadratically (every token scores against every earlier token, O(nt²) per layer — doc 11). Both are why real engines chunk very long prompts into batches; minfer's CLI keeps the simple one-shot shape and relies on n_ctx clamping (§2.5) to keep the problem bounded.

2.4 Why only the last token's logits (n_out = 1)

The prefill forward's fourth argument is n_out, and the CLI passes 1 (main.rs:757). This is the tail-row optimization whose topology doc 05 §2.8 built; here is the why.

A transformer is causal: token t's hidden state is computed only from tokens 0..=t. A consequence that surprises people the first time: the forward pass could, in principle, produce a "next-token prediction" for every prompt position — position 17's logits predict token 18, and so on. But during prefill we already know what token 18 is: it is sitting in the prompt (positions 18..nt-1 are the prompt itself). Training uses those predictions as the learning signal; inference does not need them. The only position whose continuation is genuinely unknown is the last one. So the CLI requests n_out = 1: one row of logits, for one token, the last.

What does n_out = 1 change? Doc 05 §2.8 drew the line not at lm_head but one layer earlier: after the last layer's attention output projection, two get_rows nodes select the tail n_out rows of the attention output and of the residual. Everything after that — the last FFN block, the final residual add, the final RMSNorm, and the lm_head — runs on n_out rows instead of nt. The row indices themselves arrive via a third input node, tail_ids, filled with [nt-n_out .. nt) before execution — again data, not structure. The saving is the largest single-matrix win in the graph: for a 30-token 0.5B prompt, full-nt lm_head costs 30 × 896 × 151936 ≈ 4.1 × 10⁹ multiply-accumulates; the last-row-only version costs 1.4 × 10⁸ — 30× less on the widest matrix in the model (doc 05 §2.8 measured the whole-graph effect at ~+55% prefill throughput).

Worth pausing on why n_out = 1 is safe here and not a general-purpose setting: it is a property of the single-sequence CLI. The engine keeps n_out as a GraphParams field precisely because other callers want more — a trainer computing loss over all positions would pass n_out = nt (and the graph would keep every row), and the graph's own fallback (forward_cached, graph.rs:613-626) handles the n_out == nt case with the same code path. The CLI's choice is the degenerate, cheapest corner of a general mechanism.

2.5 Sizing the context once for both phases

Right before the forward call, main.rs computes:

#![allow(unused)]
fn main() {
let ctx = params.n_ctx.max(input_ids.len());
}

and passes that single number to every forward of the run — prefill at main.rs:757, each decode step at main.rs:932. The comment above it (L744–748) records the three facts that make this the right shape:

  1. n_ctx sizes the graph's KV regions — the allocator carves each layer's K and V regions as n_kv_embd × n_ctx f32 elements, once (doc 07 §2.5). It is not the model's max_seq_len: using that unconditionally was the 12 GB lesson of docs/PERF-QWEN3-4B-VS-LLAMACPP.md §2 (a 10-token prompt reserving 12.1 GB of KV and paying a 3× first-token Metal submit tax).
  2. It never shrinks below the prompt length (max): the prompt itself needs nt KV slots, positions 0..nt, and the decode loop continues at current_pos = input_ids.len() (main.rs:840) — if ctx < prompt len, the very first decode position would overflow the regions.
  3. The model clamps again to max_seq_len (n_ctx.min(max_seq_len), graph.rs:392): a prompt longer than the model was trained for is capped at the model's own limit, the same way llama.cpp clamps. The two clamps chain: CLI requests → max(prompt) floors it → min(max_seq_len) caps it. One number survives.

The byte arithmetic for that number (per token of context headroom; the regions are f32):

per token, per layer : 2 regions (K + V) × n_kv_embd × 4 B
Qwen2.5-0.5B         : 2 × 128  × 4 B = 1 KB/layer × 24 layers =  24 KB/token
  at n_ctx 4096      : 24 KB × 4096 ≈ 96 MB total (doc 07 §2.5's ≈100 MB)
Qwen3-4B             : 2 × 1024 × 4 B = 8 KB/layer × 36 layers = 288 KB/token
  at n_ctx 4096      : 288 KB × 4096 ≈ 1.2 GB   (doc 01 §2.4; doc 07 §2.5)
  at max_seq_len 40960 (the old bug):          ≈ 12.1 GB  (!)

Now the once part, which is the design question hiding in the comment's last line ("Computed ONCE so prefill and decode size the same KV regions"). Suppose prefill passed max(4096, prompt=512) = 4096 but the decode loop recomputed with some other value. Two failure modes, both bad:

  • A larger decode n_ctx would be silently ignored. ensure_kv sizes each region on first use only and returns the existing pair thereafter (doc 07, excerpt 5: "The region is also sized on first use only: if a later graph asked for a different size, it would silently get the old buffer"). The decode graph would declare bigger regions (kvcache_store shape [n_kv_embd, n_ctx]) but write into the small ones — an out-of-bounds write waiting for a long generation.
  • A smaller decode n_ctx would trip the pre-flight assert — forward_cached asserts max(positions) < n_ctx (graph.rs:415-420) — after the prefill already filled positions 0..512. Loud, but a crash the design made unnecessary.

And there is a third reason that is about work, not safety: n_ctx lives inside CParams, which is part of the reuse identity (params_match, doc 07 §2.8). Any change forces a graph rebuild. During a rebuild the graph is replaced but the allocator — and with it the KV regions and their freshly written prefill contents — is kept (doc 07 §2.5). So identical n_ctx buys the best possible outcome: the prefill→decode transition rebuilds the graph (because n_tokens changed 512 → 1), the regions survive untouched, and decode's very first attention reads exactly the rows prefill wrote. Zero copies, zero refills.

2.6 What comes back: an owned Vec<f32>, moved not copied

forward returns Vec<f32> — the logits of the last n_out tokens, concatenated in token order. For n_out = 1 that is exactly one row of n_vocab f32 values: 151,936 × 4 B ≈ 608 KB on 0.5B/Qwen3. (The code comment says "607 KB"; 151,936 × 4 = 607,744 bytes — same number, rounding choice.)

Where the bytes travel is worth tracing, because it explains the owned return type. Inside forward_cached, after execute returns, the logits live in the graph's output buffer — a pool buffer owned by the allocator (doc 07). The engine copies them out exactly once:

#![allow(unused)]
fn main() {
let logits = alloc.copy_to_cpu(graph.outputs[0]).expect("logits buffer");
// R3-A2: the buffer is always exactly n_out*nv (G3-reduced, or
// n_out == nt) — skip the redundant full-logits clone.
if logits.len() == n_out * nv { logits } else { logits[..n_out * nv].to_vec() }
}

(graph.rs:617-625.) That one copy is unavoidable: the pool buffer will be overwritten by the next forward's lm_head (the allocator pins it only for the duration of one execute, doc 07 §2.2), so the caller must own its own copy before the next step. What the design avoids is every copy after that: copy_to_cpu already returns a Vec sized exactly n_out × nv (the else branch — the slice-and-to_vec() second copy — is dead on the G3-reduced path; the comment records its removal as R3-A2), and from there the value travels by move, Rust's zero-byte ownership transfer:

#![allow(unused)]
fn main() {
let logits = model.forward(&input_ids, &positions, &mut kv_cache, 1, ctx); // owned Vec
let last_logits: Vec<f32> = logits;      // move, main.rs:758
...
let mut logits = last_logits;            // move into the decode binding, main.rs:833
...
let sampled = sampler::sample_with_penalties(&mut logits, ...); // borrow, mutate in place
...
logits = model.forward(&[sampled.token_id], &[current_pos], ...); // move again, main.rs:932
}

The decode loop's comment (main.rs:924-925) states the stakes: forward() returns n_out*nv logits (n_out=1 for single-token decode, exactly n_vocab), so move the Vec in place instead of copying 607 KB/token. At ~300 decode tokens/second, an extra 608 KB copy per token would be ~180 MB/s of pure memcpy — a measurable tax on a loop that is already memory-bound. The borrowed-slice alternative (forward returning &[f32] into the pool buffer) is not viable for a second, harder reason: the borrow would pin the allocator for as long as the caller holds the logits, but the very next statement needs &mut access to run the next forward — the borrow checker forbids the loop outright. And even ignoring the checker, the pointed-to buffer is recycled by the next execute; a held slice would read garbage (the next token's logits, or whatever the pool put there). Owned-at-the- boundary is the minimal-copy, borrow-checker-friendly shape: exactly one copy per forward (pool → caller), zero after that. Doc 13 picks this thread up from the decode side.

3. Implementation

3.1 Data in / data out

In (what the prefill block holds when it calls forward):

  • input_ids: Vec<u32> — the tokenized prompt from doc 04; becomes the token_ids input node, shape [nt, 1, 1, 1], typed I32 (carried as f32 bit patterns, doc 07 §2.6).
  • positions: Vec<usize> — 0..nt (§2.2); becomes the positions input node, same shape.
  • n_out = 1 — the tail-row count (§2.4).
  • ctx = max(params.n_ctx, nt) — the KV region width for the whole run (§2.5); params.n_ctx defaults to 4096 (main.rs:74).
  • kv_cache — gone (#252). Until #252 the call passed a legacy KVCache created in main.rs; the graph path always ignored it (the callee bound it as _kv) and the graph owns KV in its persistent regions. #252 deleted the argument, the type and that construction (doc 03).
  • The model itself: hparams and weight tensors, registered by name in the allocator's registries (doc 03).

Out:

  • Vec<f32> of n_out × n_vocab = 1 × 151936 logits — the last prompt token's unnormalized scores over the vocabulary (§2.6). On the graph path this is one copy_to_cpu out of the output buffer, owned by the caller.
  • A filled KV cache — the second, invisible output: every layer's K and V regions now hold rows 0..nt (written by the kvcache_store nodes), which every subsequent decode step reads. Nothing is returned for it; the regions live in the GraphCache's allocator and simply persist (doc 07).
  • A wall-clock prefill_time covering the whole call — including the one-time graph build, backend assignment, allocation, and (on Metal) the first-submit setup cost (§3.2, timing calibers).

3.2 Key code

The prefill block (src/main.rs:737-769)

#![allow(unused)]
fn main() {
    // === Prefill ===
    let infer_start = Instant::now();
    let positions: Vec<usize> = (0..input_ids.len()).collect();
    // forward() computes logits for only the LAST n_out tokens (n_out=1 here:
    // single sequence, only the final token is sampled). llama.cpp does the same
    // via ggml_get_rows(inp_out_ids) at the last layer, shrinking the lm_head
    // to n_outputs rows — saves the full-nt output GEMM + logits download.
    // n_ctx (--n-ctx, default 4096) sizes the graph KV regions — NOT the
    // model's max_seq_len, which would allocate 12 GB+ and pay a first-submit
    // Metal tax (docs/PERF-QWEN3-4B-VS-LLAMACPP.md §2). It never shrinks below
    // the prompt length, and the model's forward clamps it to max_seq_len.
    // Computed ONCE so prefill and decode size the same KV regions.
    let ctx = params.n_ctx.max(input_ids.len());
    // P2 trace (MINFER_TRACE=<path>): the scheduler records per-node data; the
    // CLI marks phase boundaries and attaches tokens/logits for the page.
    let trace_on = crate::trace::enabled();
    if trace_on {
        crate::trace::begin_phase("prefill");
    }
    let logits = model.forward(&input_ids, &positions, &mut kv_cache, 1, ctx);
    let last_logits: Vec<f32> = logits;
    if trace_on {
        crate::trace::attach_step(&last_logits);
    }

    let prefill_time = infer_start.elapsed();
    println!(
        "Prefill: {} tokens in {:.2}s ({:.1} tok/s)",
        input_ids.len(),
        prefill_time.as_secs_f64(),
        input_ids.len() as f64 / prefill_time.as_secs_f64()
    );
}

Segment by segment: L738 starts the prefill stopwatch (after tokenization — tokenize cost is not in the prefill number); L739 builds the positions vector (§2.2); L740–748 is the comment that documents n_out (and its llama.cpp analogue inp_out_ids), the n_ctx sizing policy, and the compute-once rule; L749 the double-clamped context (§2.5); L753–756 opens the prefill trace phase (only when MINFER_TRACE is set — one env read per run, hoisted); L757 is the stage itself — one call, all prompt tokens, n_out = 1; L758–761 renames the result and (for the trace page) attaches the logits' top-5 to the prefill step; L763–769 prints the prefill caliber.

The call chain — four hops, one delegation each

main.rs:757 calls through the architecture-agnostic trait object (Box<dyn ModelDef>), so the static type knows nothing about Qwen. Each hop adds exactly one concern:

main.rs:757   model.forward(&input_ids, &positions, &mut kv_cache, 1, ctx)
  │           trait method (models/mod.rs:26-33): the architecture-agnostic
  │           signature; doc comment: "the legacy `kv` arg is ignored" on the
  ▼           graph path; callers must guarantee positions[i] < n_ctx
qwen2/mod.rs:33-42   impl ModelDef for Qwen2Model
  │           one line: graph::Qwen2Graph::forward(self, tokens, positions, kv, n_out, n_ctx)
  ▼           (qwen3/mod.rs mirrors this identically for Qwen3)
qwen2/graph.rs:390-401   Qwen2Graph::forward
  │           n_ctx = n_ctx.min(model.hparams.max_seq_len)  ← 2nd clamp
  │           locks the process-global GraphCache (graph_cache())
  ▼           delegates to forward_cached — the CLI wrapper; server code
              calls forward_cached directly with a slot-scoped cache
qwen2/graph.rs:409-632   forward_cached — the real work (below)

The two trait-level aliases on the way are worth one glance (models/mod.rs:262-298): forward_graph and forward_graph_cached expose the same two implementations at trait level — forward_graph_cached is the server/multi-slot entry point (it takes an explicit &mut GraphCache instead of using the process-global one). The CLI goes through plain forward and lands in the same forward_cached; there is one forward-pass implementation, not one per caller.

forward_cached: the build → fill → execute → read pipeline

(src/models/qwen2/graph.rs:409-476, 520-550, 617-625 — abridged)

#![allow(unused)]
fn main() {
    pub fn forward_cached(model: &Qwen2Model, tokens: &[u32], positions: &[usize],
                          n_out: usize, n_ctx: usize, cache: &mut GraphCache) -> Vec<f32> {
        let nt = tokens.len();
        debug_assert!(n_out <= nt);
        // Out-of-range positions would write past the KV regions (which are
        // sized n_kv_embd * n_ctx): fail loudly instead of corrupting memory.
        if let Some(&maxp) = positions.iter().max() {
            assert!(maxp < n_ctx, "position {maxp} exceeds n_ctx {n_ctx} ...");
        }
        /* metal_on / cuda_on: device present AND every weight registered —
           all-or-nothing GPU participation (docs 03/14/15) */
        let params = GraphParams {
            n_tokens: nt,
            n_out,
            gtype: if nt == 1 { GraphType::Decode } else { GraphType::Prefill },
            cparams: CParams { n_ctx, flash_attn: false,
                               gpu: metal_on || cuda_on,
                               fuse_qkv: nt == 1 && (metal_on || cuda_on) && ...,
                               fuse_ffn: nt == 1 && (metal_on || cuda_on) && ... },
            weights_version: 1,
        };
        if !cache.try_reuse(&params) {
            /* build → register weights → assign_backends → FusionPass →
               alloc_graph → cache.replace_graph (docs 05-08) */
        }
        let (graph, alloc) = cache.current().unwrap();
        // refresh input data (positions/ids are data, not topology)
        alloc.fill_input_i32(graph, "token_ids", &ids).unwrap();
        alloc.fill_input_i32(graph, "positions", &pos).unwrap();
        if graph.inputs.iter().any(|&i| graph.node(i).name == "tail_ids") {
            let tail = ((nt - n_out)..nt).collect::<Vec<u32>>();
            alloc.fill_input_i32(graph, "tail_ids", &tail).unwrap();
        }
        let sched = BackendScheduler::new();
        sched.execute(graph, alloc).unwrap();
        /* MINFER_GRAPH_DUMP block omitted (logits + KV dumps for debugging) */
        let logits = alloc.copy_to_cpu(graph.outputs[0]).expect("logits buffer");
        if logits.len() == n_out * nv { logits } else { logits[..n_out * nv].to_vec() }
    }
}

This is the whole "call side of the graph machinery" in one function, and it is worth reading as five beats:

  1. Pre-flight assert (L415–420): every position must be < n_ctx — the loud version of the KV-overflow check that the store kernel also enforces (doc 07 §3.2, excerpt 8).
  2. GraphParams construction (L438–470): every field derived from the call's arguments plus device availability. This is the prefill graph's birth certificate — for a 512-token CPU prompt: n_tokens = 512, n_out = 1, gtype = Prefill, cparams = { n_ctx: 4096, flash_attn: false, gpu: false, fuse_qkv: false, fuse_ffn: false }. Note the fusion flags are nt == 1 && gpu: the decode fusions of doc 05 §2.6 are off during prefill by construction — they pay off at nt = 1 only, and their being params-derived is what keeps fused and unfused graphs reproducible (and A/B-able via MINFER_NO_FUSE_QKV=1).
  3. Reuse-or-build (L472–518): try_reuse compares the six fields against the cached graph (doc 05 §2.7); on mismatch — which for the CLI happens exactly twice, at the prefill call and at the first decode call — the full build → assign → fuse → allocate pipeline runs, and replace_graph swaps the graph in keeping the allocator (hence the KV regions).
  4. Fill inputs, execute (L520–550): the three input buffers are overwritten with this step's data (by name — node ids shift between rebuilds, names do not), then the scheduler walks the nodes (doc 08). execute returning Ok(()) is the guarantee that every buffer — logits included — holds its final value (doc 08 §2.4).
  5. Read the output (L617–625): one copy_to_cpu of the output buffer, returned as an owned Vec (§2.6).

GraphParams → topology: what each field changes

Doc 05 established the mapping; here is the prefill-relevant digest, with the field's home in params.rs:

Field (params.rs)Prefill value (CLI)Topology effect (doc 05 §)
n_tokensprompt lengthevery activation shape's nt; decode-fusion gates read nt == 1
n_out1tail get_rows pair after the last attention projection (§2.8); the whole output stack runs on 1 row
gtypePrefillpart of the reuse identity; with nt it names the graph class (§2.6)
cparams.n_ctxmax(--n-ctx, prompt)KV store/load node shape [n_kv_embd, n_ctx] → region size (doc 07 §2.5)
cparams.gpufalse (CPU) / truebackend-assignment eligibility — Cuda/Metal claim nodes only when every weight is registered
cparams.fuse_qkv / fuse_ffnfalse (nt > 1)decode-only fused topologies; prefill builds the plain matmul+rope+store skeleton
weights_version1invalidates reuse if weights change (future LoRA/reload hook)

(src/graph/params.rs:10-63; GraphType L10–15, CParams L22–48, GraphParams L50–63. The module doc's first line is the invariant: these are "the ONLY inputs to graph reuse".)

The trace phase boundary (src/trace.rs:70-80, CLI at main.rs:753-761)

#![allow(unused)]
fn main() {
/// CLI: mark the start of a phase (prefill / decode). Repeated calls within the
/// same phase are no-ops; a kind change starts a new phase.
pub fn begin_phase(kind: &str) {
    let mut t = trace().lock().unwrap();
    match t.phases.last() {
        Some(p) if p.kind == kind => {}
        _ => t.phases.push(Phase { kind: kind.into(), steps: Vec::new(), graph: None }),
    }
}
}

begin_phase("prefill") opens a Phase record; the scheduler then appends one Step per execute() call (begin_step, called from the scheduler — prefill is one step, each decode forward one more), and attach_step(&last_logits) (L759–761, trace.rs:130-145) staples the prefill logits' top-5 (token_id, probability) pairs onto that step. The JSON export (trace.rs::finish) embeds the prefill-phase graph, so the viz page can show the whole prefill execution node by node. Note the phase naming lives in the CLI, not the engine: the engine has GraphType, the trace has human-facing labels.

The timing calibers (src/main.rs:763-769, 837, 1000-1017)

Three stopwatches, three printed calibers — and the differences between them are the point:

#![allow(unused)]
fn main() {
    let prefill_time = infer_start.elapsed();          // L763 (started L738)
    println!("Prefill: {} tokens in {:.2}s ({:.1} tok/s)", ...);   // prompt / prefill wall

    let gen_start = Instant::now();                    // L837, "pure-decode start"
    ...
    let gen_time = gen_start.elapsed();                // L1000
    let total_time = infer_start.elapsed();            // L1001
    // Pure-decode rate (generated tokens / decode time) — matches llama.cpp's
    // "Generation:" caliber. The "Total:" line below keeps the previous blended
    // caliber (prompt+generated / prefill+decode) for comparison.
    println!("Generated: {} tokens in {:.2}s ({:.1} tok/s)", ...); // L1006-1011
    println!("Total:     {} tokens in {:.2}s ({:.1} tok/s)", ...); // L1012-1017
}
  • Prefill = prompt tokens ÷ wall time of the single prefill forward. It includes the one-time costs of the run: graph build, backend assignment, allocation, and on Metal the first-submit setup (the 3× first-token tax of doc 07 §3.3). Long prompts amortize it; a 1-token prompt measures mostly setup.
  • Generated = generated tokens ÷ decode-only wall time (gen_start L837 → L1000). This is the llama.cpp/llama-bench "Generation" caliber: steady-state speed, the number people mean by "tok/s". It includes the sampler and streaming write per token — which is why MINFER_TIMING exists to split it further (§4).
  • Total = (prompt + generated) ÷ (prefill + decode) — the blended caliber llama.cpp prints as its "Total" line; dominated by whichever phase has more tokens.

Why split the calibers at all, rather than one honest number? Because the two phases are bottlenecked by different resources, and one blended number hides both. Prefill is compute-bound (§2.3): every weight byte is reused for nt multiply-accumulates, so throughput scales with ALU/SIMD throughput and kernel efficiency — GPUs love it, and that is where int8 MMQ prefill (doc 15) pays. Decode is memory-bound: each token streams the entire weight set from memory to do only 2 × n_params flops with it, so throughput tracks memory bandwidth, not arithmetic. A hardware or kernel change (say, faster dot products) moves prefill a lot and decode barely; a memory-side change (quantization, bandwidth) moves decode and barely touches prefill. Reporting them separately is what makes such measurements interpretable — and minfer bench (§4) adopts llama-bench's pp/tg test split for exactly this reason (bench.rs:1-12: "pp<P>: prefill-only … tg<T>: prefill P context tokens (untimed setup), then time the decode").

3.3 Design choices (why this shape and not another)

Q1: Why process ALL prompt tokens in ONE forward instead of token-by-token? §2.3 gave the bandwidth argument; the full tally has four legs:

  1. Amortized weight traffic. One batched forward reads every weight byte once and reuses it for all nt tokens' math (each weight element feeds nt multiply-accumulates instead of 1). Token-by-token multiplies the dominant cost of decode-shaped work by nt for zero extra information.
  2. One graph build. The build → assign → fuse → allocate pipeline runs once per distinct GraphParams; token-by-token would run it nt times (or force one graph to serve growing nt, which violates the topology-=-f(params) reuse identity).
  3. KV filled in one pass. The kvcache_store nodes write positions 0..nt contiguously in the same execution; build order guarantees each layer's store precedes its attention (doc 08 §2.1). The batched attention is also the natural shape for the causal-mask kernel: token t reads rows 0..=pos[t] of a region that is already there.
  4. Parallelism inside the kernels. Wide matmuls ([out, nt] outputs, doc 05 §2.4's shape convention) give SIMD lanes and GPU threadgroups nt columns of independent work — the difference between a GEMV (one token: bandwidth-bound) and a GEMM (many tokens: compute-bound). This is precisely why prefill on CUDA uses a different kernel family (int8 MMQ) than decode (MMVQ), doc 15.

The alternative was rejected on measurement, not taste: llama.cpp made the same choice, and minfer's own fusion A/B numbers (doc 05 §2.6) show how even decode-side batching decisions are gated by measurements — prefill batching is the same discipline applied where the win is 100×, not 10%.

Q2: Why only the last token's logits? §2.4 gave the causality argument; the design-shaped summary: for an autoregressive generator, every prompt position except the last has a known continuation (it is in the prompt), so computing logits for them buys nothing the CLI can use. The n_out mechanism generalizes (any tail count; n_out = nt restores full logits and is the code's own fallback path, graph.rs:613-626), so choosing 1 is a caller policy, not an engine limitation — the engine still offers every row to callers who want them (trainers, scoring tools). The placement of the cut after the last attention projection (not at lm_head) is doc 05's §2.8 refinement: everything downstream of the tail select — last FFN, final norm, lm_head — shrinks with it, and the measured effect was ~+55% prefill throughput on 0.5B.

Q3: Why is n_ctx sized once for both phases? §2.5 gave the failure modes; the principle underneath: the KV regions are process-lifetime state, allocated on first use inside an allocator that outlives every graph (doc 07 §2.8). Their size is therefore a run-level decision, not a phase-level one — and the only run-level facts available are the CLI flag and the prompt length. Sizing per phase would either silently reuse the first phase's regions (too small → corruption on later positions) or force a region reallocation (a full KV copy, or a loss of everything prefill just wrote — the cache is the run's memory). One number, computed once from max(CLI, prompt) and capped by the model, is the smallest contract that makes "prefill fills, decode appends" work with zero copies.

Q4: Why does forward return owned Vec<f32> instead of a borrowed slice? §2.6 gave the mechanics; the ownership-shaped summary:

  1. The pool buffer's lifetime is shorter than the caller's need. The logits buffer is pinned for one execute (doc 07 §2.2); the next forward overwrites it. A borrowed return would hand the caller a reference into memory whose contents are dead by the time the loop iterates.
  2. The borrow checker agrees. Holding &[f32] into the allocator's pool while calling forward again (which needs &mut GraphCache) is a compile error — the loop shape requires ownership at the boundary.
  3. One copy is the minimum anyway. The data must leave the pool before the next execute; copy_to_cpu produces the owned Vec at exactly n_out × nv size (R3-A2 removed the second, slice-to_vec copy that a non-shrunk logits buffer would have needed). After that, every handoff — prefill binding → decode binding → sampler borrow → next forward's move — transfers 8 bytes of pointer/len/cap, never the 608 KB payload. The recorded alternative, copying 607 KB/token at ~300 tok/s, is ~180 MB/s of pure memcpy inserted into the engine's most memory-sensitive loop.

3.4 Pitfalls & invariants

  • Positions must satisfy positions[i] < n_ctx — always. The preflight assert (graph.rs:415-420) and the store kernel's hard error (cpu_backend.rs:169-171, doc 07) both enforce it, because a bad position is an out-of-bounds write into a persistent region. The CLI satisfies it structurally: positions are 0..nt and ctx ≥ nt by the max clamp.
  • n_ctx is a run-level constant. Passing a different n_ctx to decode than to prefill breaks the run (§2.5, Q3): either silently (regions sized on first use) or loudly (the preflight assert). This is why the comment says "computed ONCE" — and why main.rs:932 passes the same ctx binding, not a recomputation.
  • The legacy kv_cache argument is gone (#252). It was dead on this path from the start — nothing read it in forward_cached — and #252 deleted it (with the KVCache type and the main.rs construction) once the storage had already gone in #244. forward_cached's signature never took it; the server's GraphCache/forward_batch surface is unchanged.
  • The prefill timer includes one-time setup. Comparing Prefill: lines across runs of different prompt lengths (or across builds with different fusion env toggles) mixes steady-state prefill cost with build/first-submit cost. For steady-state numbers use minfer bench (warmup + reps, bench.rs:292), not the CLI's single-shot print.
  • n_out is part of the reuse identity. Changing it rebuilds the graph (it changes the tail topology). The CLI never does mid-run; the server could, and the params comparison handles it (params_match, doc 07 §2.8) — but a caller that flips n_out per step would rebuild per step.
  • Prefill and decode graphs are two graphs, one cache. The transition costs exactly one rebuild (first decode step); the KV regions and the allocator survive it (doc 07 §2.5). If decode's first step ever seems to "lose" prefill's cache, the bug is in region identity (layer index, backend, dtype) — not in the params scheme, which is what the allocator_survives_rebuild test pins (doc 07 §4).

4. Observe & verify

  • The three printed calibers — any run prints Prefill: … tok/s, Generated: … tok/s, Total: … tok/s (§3.2). On 0.5B CPU expect prefill in the thousands of tok/s and decode in the hundreds — the compute-bound vs memory-bound split of §2.3, visible in two numbers.
  • MINFER_TIMING=1 — decomposes the decode side one level further (main.rs:863-868): per token it times the sampler call (t_samp, around sample_with_penalties, L884–898) and the forward call (t_fwd, L931–939 — "CPU encode + GPU exec + logits download" per the comment), then prints one line, e.g. [MINFER_TIMING] over 512 tokens: sample 0.02 ms/tok (1.2%), forward 3.30 ms/tok (98.8%) (L949–953). One paragraph is all it needs: it exists to answer "is my decode time inference or sampling/streaming?" — on every model in the support matrix the answer is overwhelmingly forward, which is why the kernel docs (10/14/15) own the optimization story.
  • MINFER_TRACE=/tmp/t.json — the trace records a prefill phase (opened by begin_phase, §3.2) with one step per execute; the step carries per-node buffer stats for all 440 nodes and the prefill logits' top-5 (attach_step). Load at viz/index.html or minfer viz.
  • MINFER_GRAPH_DUMP=/tmp/d — writes logits_prefill.f32 (n_out × n_vocab lef32 values — the exact Vec this doc is about) plus per-node and per-layer KV dumps (graph.rs:554-611), so CPU-vs-GPU logits comparisons start from this stage's output.
  • --dump-graph / --dump-graph-json — rebuilds and exports the prefill graph (the same GraphParams the runtime used, doc 05 §4): the n_out tail rows are visible as the two get_rows nodes before the last FFN, and the node count differs from the decode graph's (440 vs 437 on 0.5B).
  • minfer bench — the measurement-grade version of the calibers: pp<P> (prefill-only) and tg<T> (prefill untimed + decode timed) rows with llama-bench-style mean/stdev over reps (bench.rs:1-12, greedy sampling for reproducibility). Use it instead of single-shot prints for any number you intend to compare.
  • Tests — graph_logits_match_forward_real_model (graph path vs the old imperative path, the acceptance bar for this whole pipeline), tail_reduction_matches_full_nt (the n_out mechanism of §2.4: reduced and full-logits graphs agree on the tail rows), and the reuse quartet in graph/cache.rs (params-only identity) — all cited with locations in doc 05 §4.

5. Cross-references

← 08 — The scheduler: splits, copies, execution · Index · 10 — CPU matmul: quantized weights × Q8_0 activations →

10 · CPU matmul: quantized weights × Q8_0 activations

Stage: prefill invoked (doc 09) → this stage: inside a CPU MatMul node → attention and vec ops (doc 11). Code: src/kernel.rs (cpu_quant_matmul_f32, mm_rows, the thread pool, embed_tokens), src/quants.rs (quantize_row_q8_0, quantize_row_q8_k_buf, the dot_* kernels), src/block.rs (block layouts).

1. Background — where this stage sits

Doc 09 followed the prefill call from main.rs into the graph machinery. Doc 08 ended at the moment the scheduler walks a split and calls execute_node for every node. For a MatMul node on the CPU backend, that call bottoms out in the code this document walks: the quantized matrix multiplication that produces essentially every number a transformer computes.

First, the vocabulary, because the rest of the series assumes it:

  • A matrix multiply (matmul) takes a matrix of weights W and a matrix of activations X and produces an output where element (t, o) is the dot product of output row o of W with token row t of X. A dot product multiplies matching elements and adds them up: Σ W[o][i] · X[t][i]. In a transformer, roughly 169 of the 440 nodes of a 0.5B prefill graph are matmuls (doc 05's census) — they dominate both compute time and model size.
  • Weights are the learned matrices loaded from the GGUF file (doc 03). In a 7B model they total ~4.4 GB. Storing them as plain 32-bit floats would need ~14 GB; storing them quantized — fewer bits per number, organized in small groups with a shared scale — is what makes local inference practical at all.
  • Quantization (for this doc) means: take a group of 32 floating-point values, find the largest absolute value in the group, store that as one shared half-precision scale (d), and store the 32 values as small integers relative to that scale. Dequantizing gives value ≈ q · d — the integers times the scale. The error is bounded by half a scale step per value, which for a well-chosen group size is small enough that llama.cpp ships this design at 7B+ scale with usable quality.

minfer's CPU matmul follows llama.cpp's core trick, which this document makes explicit:

The weights stay quantized on disk and in memory — they are never dequantized as a whole. Activations are quantized to int8 on the fly, one block at a time, and the inner loop is an int8×int8 dot product with per-block scales folded in.

Why that is the winning design is §2. The implementation walk is §3: the dispatch (mm_rows), the activation quantizers (quantize_row_q8_0 and the K-quant quantize_row_q8_k_buf), the AVX2 kernel line by line, its scalar reference, the NEON/SDOT counterpart, the persistent thread pool, and the repr(C) block layouts that make all of it possible. Doc 14 and doc 15 reuse the same math on the GPU with different execution models.

2. Principle — why quantized weights × Q8_0 activations

2.1 The bandwidth argument (why not dequantize once at load?)

The obvious alternative to on-the-fly work: at load time, dequantize every weight to f32 and run a plain float matmul. The numbers kill it.

During decode (doc 13), the engine produces one token per forward pass. Every matmul must read its entire weight matrix to produce that one token's outputs — with one token there is nothing to amortize over. The 7B model's ~4.4 GB of q4_K_M weights therefore stream from RAM per token. On a memory subsystem moving ~60–100 GB/s, that alone sets a ceiling of roughly 15–25 tokens/s — and that ceiling is the decode physics on CPU (the same argument appears on the GPU side, where doc 15 calls it "weight streaming").

Now compare the two storage options:

Storage7B weight bytesStream time @ ~80 GB/sDecode ceiling
f32 weights~28 GB (7B × 4 B)~350 ms/token~3 tok/s
q4_K_M weights (4.5–5.5 bits/weight)~4.4 GB~55 ms/token~18 tok/s

The same arithmetic per byte-of-precision: Q4_0 spends 18 B / 32 values = 0.5625 B per weight — 7.1× less traffic per multiply than the 4 bytes an f32 weight costs. Quantization is not a quality knob here — it is the difference between usable and unusable. Dequantizing 4.4 GB into 28 GB of f32 at load would also cost the load itself seconds and 6× the RSS.

2.2 Why the activations also go int8 (the Q8_0 trick)

The weights being 4-bit is only half the story. The inner loop must multiply weight elements by activation elements. If activations stayed f32, every inner-loop iteration would need a convert (int→float) plus a float multiply-add — and the CPU's fastest small-integer machinery goes unused.

So minfer (following llama.cpp) quantizes the activations to Q8_0 — per 32-value block, one f16 scale + 32 int8 values — right before the matmul (§3.2). The inner loop then becomes an int8×int8 dot product, for which x86 has vpmaddubsw/vpmaddwd (the AVX2 path below) and ARM has SDOT (the NEON path): single instructions that multiply and add 8–16 integer pairs each. The per-block scales are folded in once per block, outside the integer loop, as one float multiply-add.

The accuracy story: activations are quantized per 32 values with their own exact scale, so the input side carries no cross-tensor approximation drift; weights are 4-bit but their scales were chosen at model-conversion time by the quantizer with the whole tensor in view. The repo's verification records show the resulting CPU logits matching llama.cpp bit-for-bit — the format is the same, the kernel math is the same, and the accumulated error across 28 layers lands identically.

2.3 The dispatch in one picture

For a matmul node with weight type T, output od rows, input dim id, nt tokens:

x: [nt][id] f32 (token-major, from the allocator's f32 pool — doc 07)
        │
        ▼  quantize_row_q8_0_buf   (K-quant weights → quantize_row_q8_k_buf instead)
x_q: [nt][id] int8 blocks + f16 scales (34 B per 32 values, or 306 B per 256)
        │
        ▼  cpu_quant_matmul → thread pool → mm_rows (one row per worker)
w row o: [id] quantized blocks (18/22/24/34/144/176/210 B per 32/256 values)
        │
        ▼  dot_<T>_q8_0(wrow, xrow)   ← the inner loop, per (output, token)
out: [nt][od] f32

The block sizes come from block.rs and are the contract between the GGUF file (doc 02), the loader (doc 03), and the kernels:

#![allow(unused)]
fn main() {
// src/block.rs:18-27
pub const Q4B: usize = 18;  // sizeof(block_q4_0)
pub const Q8B: usize = 34;  // sizeof(block_q8_0)
pub const Q4KB: usize = 144; // sizeof(block_q4_k)
pub const Q6KB: usize = 210; // sizeof(block_q6_k)
...
pub const Q8KB: usize = 2 + 256 + 16 + 32; // 306  (block_q8_k)
}

Sanity check with a real number: token_embd of Qwen2.5-0.5B in Q4_0 with id = 896 has 896/32 = 28 blocks per row, so one weight row is 28 × 18 = 504 bytes — and doc 02's byte-size formula (n / blck_size) × type_size lands on exactly that.

2.4 Two activation formats, one pairing rule

The entry-point branch (src/kernel/dispatch.rs:9-31, below) is not cosmetic — each weight family pairs with a specific activation format, fixed by the kernels' inner loops:

Weight familyWeight blockPaired activation formatActivation block
Q4_0 / Q4_1 / Q5_0 / Q5_1 / Q8_032 values (18/20/22/24/34 B)Q8_032 values, 34 B: f16 d + 32 × i8
Q4_K / Q5_K / Q6_K256-value super-blocks (144/176/210 B)Q8_K256 values, 306 B: f16 d + 256 × i8 + 16 reserved + 16 × i16 bsums

The Q8_K pairing exists because the K-quant kernels unpack one 256-value weight super-block — with its 8 sub-block scales and mins — and want the matching 256-value activation super-block with its own group sums (bsums, used by the K-quant dot's correction term) adjacent. The layout comment is pinned in the source:

// src/quants/quantize_q8_k.rs:8-18 — activation q8_K block: d(f16) + qs[256 i8] +
// bsums[16 i16] = 306 bytes (crate::block::Q8KB).

(Note a subtlety doc 02 flagged: block.rs's BlockQ8_K struct — a f32 d ggml-style layout, block.rs:173-177 — is not the same 306-byte arrangement the activation quantizer writes. The activation-side format is defined by the quantizer + dot kernel pair, and that is the only contract the matmul path relies on.)

3. Implementation

3.1 Data in / data out

ItemShape / layoutWhere it comes from
Activations x[nt][id] f32, token-majorthe allocator's f32 pool (doc 07), filled by the previous node's output
Quantized activations x_q[nt] rows of id/32 × 34 B (Q8_0) or id/256 × 306 B (Q8_K)allocated per call by cpu_quant_matmul_f32
Weights w[od][id] row-major quantized bytes, exactly as the GGUF has themTensor.data — a borrowed slice of the mmap'd file (docs 02/03); row stride = blocks_per_row × block_size
Output out[nt][od] f32, token-majorthe node's output buffer in the pool (doc 07)

Note what is absent: any f32 copy of the weights, anywhere. The GGUF bytes flow untouched from file page cache → Tensor.data → the dot kernel.

One call, in real numbers. Prefill of 100 tokens through Qwen2.5-0.5B's blk.0.attn_q (an [896, 896] Q4_0 weight, so od = id = 896, 28 blocks/row) with --threads 8:

StepWorkBytes touched
quantize activationsquantize_row_q8_0_buf on 100 × 896 f32read 358 KB, write 100 × 28 × 34 = 95 KB
submitMmJob under job, bump gen, pool wakes~64 B
8 workers × 112 rows eachmm_rows rows 0..896weights: 896 × 28 × 18 = 451 KB; activations re-read per row: 112 × 95 KB from cache
dot kernels896 × 100 = 89,600 calls to dot_q4_0_q8_028 block-iterations each → 2.5M block dots
outputout[t][o] = Σwrite 100 × 896 × 4 = 358 KB f32

Every number here comes straight from the formulas of §2.3/§3.2 — this table is the sanity check to run in your head whenever a matmul result looks wrong: block counts (id/32), row strides (blocks × block_size), and output size (nt × od × 4) must all be whole and consistent.

3.2 Key code

The entry point: quantize activations, then dispatch. cpu_quant_matmul_f32 is what the CPU backend's MatMul arm calls (doc 08):

#![allow(unused)]
fn main() {
// src/kernel/dispatch.rs:9-31
pub fn cpu_quant_matmul_f32(w: &Tensor, x: &[f32], out: &mut [f32],
                            od: usize, id: usize, nt: usize) {
    match w.ttype {
        TensorType::Q4_K | TensorType::Q5_K | TensorType::Q6_K => {
            let n_super = id / 256;
            let mut qb = vec![0u8; nt * n_super * Q8KB];
            crate::quants::quantize_row_q8_k_buf(x, nt, id, &mut qb);   // → Q8_K
            cpu_quant_matmul(w, &qb, out, od, id, nt)
        }
        _ => {
            let nbe = id / 32;
            let mut qb = vec![0u8; nt * nbe * Q8B];
            crate::quants::quantize_row_q8_0_buf(x, nt, id, &mut qb);   // → Q8_0
            cpu_quant_matmul(w, &qb, out, od, id, nt)
        }
    }
}
}

The two activation formats of §2.4 are chosen here, once, so no call site can ever pair the wrong kernel with the wrong activation bytes.

The row kernel: mm_rows. One function handles every weight type; the type only changes the byte arithmetic and which dot_* gets called:

#![allow(unused)]
fn main() {
// src/kernel/pool.rs:98-120 (Q4_0 arm; the other 7 arms are the same shape)
unsafe fn mm_rows(job: &MmJob, r0: usize, r1: usize) {
    let od = job.od;          // outputs (weight rows)
    let id = job.id;          // input dim
    let nt = job.nt;          // tokens
    match job.ttype {
        TensorType::Q4_0 => {
            let nb = id / 32;          // blocks per row
            let ws = nb * Q4B;         // weight row stride: 28 × 18 B for 0.5B
            let rowb = nb * Q8B;       // activation row stride: 28 × 34 B
            for o in r0..r1 {          // this worker's rows
                let wrow = std::slice::from_raw_parts(job.w.add(o * ws), ws);
                for t in 0..nt {
                    let xrow = std::slice::from_raw_parts(job.x.add(t * rowb), rowb);
                    *job.out.add(t * od + o) = crate::quants::dot_q4_0_q8_0(wrow, xrow);
                }
            }
        }
        ...
}

Read the loop order carefully — it is the parallelism story: the outer loop is over output rows o, and each row is written by exactly one worker. That is why the comment at the top of the function can promise the multi-threaded result is bit-identical to single-threaded: no output element is touched twice, so there is no float summation order to disagree about.

The other arms only change the stride constants and the kernel name: Q8_0 has ws == rowb (both 34 B per block); the K-quant arms compute strides in 256-value units (nk = id/256, weight row nk × 144 / 176 / 210 B, activation row nk × 306 B) and call dot_q4_k_q8_k / dot_q5_k_q8_k / dot_q6_k_q8_k; Q5_0/Q5_1 add their high-bit planes (22/24 B blocks). Also note job.w.add(o * ws): the weight "matrix" is a byte pointer plus a stride — the loader never materialized an f32 matrix (doc 03), and the kernel never asks for one.

One AVX2 kernel, line by line. dot_q4_0_q8_0 first picks its engine (all the dot_* wrappers share this shape):

#![allow(unused)]
fn main() {
// src/quants/dot_q4_0.rs:9-25 (abridged to the dispatch)
pub fn dot_q4_0_q8_0(q4: &[u8], q8: &[u8]) -> f32 {
    let nb = q8.len() / Q8B;                       // 32-value blocks
    #[cfg(target_arch = "x86_64")]
    { if avx2_enabled() { return unsafe { dot_q4_0_q8_0_avx2(q4, q8, nb) }; } }
    #[cfg(target_arch = "aarch64")]
    { if neon_enabled() { return unsafe { dot_q4_0_q8_0_neon(q4, q8, nb) }; } }
    dot_q4_0_q8_0_scalar(q4, q8, nb)               // portable fallback
}
}

Runtime detection (is_x86_feature_detected!, wrapped in avx2_enabled()) is why one binary ships everywhere and picks AVX2 only where it exists; MINFER_NO_AVX2=1 forces the portable fallback for A/B (MINFER_NO_NEON=1 on aarch64). The K-quant dots (dot_q4_k_q8_k/dot_q5_k_q8_k/dot_q6_k_q8_k) follow the same shape with a third arm — an AVX-512/VNNI variant selected first, then AVX2, then scalar (#56) — so the same wrapper text applies to all ten dot_* entry points; doc 06's supports_op story is about ops, this one is about instructions, and both follow the same capability-query philosophy.

The AVX2 kernel processes one 32-value block per iteration with two 256-bit registers:

#![allow(unused)]
fn main() {
// src/quants/dot_q4_0.rs:29-52 (core loop, comments mine)
unsafe fn dot_q4_0_q8_0_avx2(x: &[u8], y: &[u8], nb: usize) -> f32 {
    let mut acc = _mm256_setzero_ps();
    for ib in 0..nb {
        let xp = x.as_ptr().add(ib * Q4B);   // weight block: 2 B scale + 16 B nibbles
        let yp = y.as_ptr().add(ib * Q8B);   // activation block: 2 B scale + 32 × i8
        let xd = f16_to_f32_bits(*xp.cast::<u16>());  // weight block scale
        let yd = f16_to_f32_bits(*yp.cast::<u16>());  // activation block scale
        let d = _mm256_set1_ps(xd * yd);     // scales fold in ONCE per block
        let tmp = _mm_loadu_si128(xp.add(2) as *const __m128i);   // 16 packed nibbles
        let bytes = _mm256_set_m128i(_mm_srli_epi16(tmp, 4), tmp); // high nibbles → hi lane
        let mut qx = _mm256_and_si256(bytes, _mm256_set1_epi8(0xF)); // keep 4 bits
        qx = _mm256_sub_epi8(qx, _mm256_set1_epi8(8));  // unsigned 0..15 → signed −8..7
        let qy = _mm256_loadu_si256(yp.add(2) as *const __m256i);   // 32 × i8
        let ax = _mm256_sign_epi8(qx, qx);   // |qx|          (unsigned-ify)
        let sy = _mm256_sign_epi8(qy, qx);   // qy × sign(qx) (move w's sign onto acts)
        let dot = _mm256_maddubs_epi16(ax, sy);  // 16 u8×i8 → 8 × i16 pairwise dots
        let q = _mm256_cvtepi32_ps(_mm256_madd_epi16(_mm256_set1_epi16(1), dot));
        acc = _mm256_fmadd_ps(d, q, acc);    // acc += (xd·yd) · q
    }
    hsum_float_8(acc)                        // horizontal add of the 8 lanes
}
}

The moves worth understanding:

  1. Nibble unpacking without a lookup table. A 4-bit weight is stored as two nibbles per byte (low = element 2i, high = element 2i+1). One _mm256_and_si256(0xF) recovers the low halves; _mm_srli_epi16(4) on a 128-bit half plus _mm256_set_m128i re-packs the high halves into the upper lane — 32 centered nibbles in one register, no per-byte work.
  2. The sign trick. maddubs multiplies unsigned × signed bytes. Weights are centered to −8..7 (signed), so the kernel takes their absolute value (sign_epi8(qx, qx)) and instead flips the activations' signs by the weights' signs (sign_epi8(qy, qx)). Same product, and now the fast unsigned×signed instruction applies.
  3. Scales outside the integer loop. d = xd·yd is one scalar per block; the integer chain (maddubs → madd → widen) produces the exact integer dot for the block, and one fmadd folds it into the float accumulator. The integer math is exact; only the block quantization itself approximates.
  4. The horizontal sum is its own small art — one lane value out of eight:
#![allow(unused)]
fn main() {
// src/quants/avx2.rs:76-81
unsafe fn hsum_float_8(x: __m256) -> f32 {
    let x128 = _mm_add_ps(_mm256_extractf128_ps(x, 1), _mm256_castps256_ps128(x));
    let x128 = _mm_add_ps(x128, _mm_movehl_ps(x128, x128));
    _mm_cvtss_f32(_mm_add_ss(x128, _mm_movehdup_ps(x128)))
}
}

The scalar kernel is the reference semantics. Before admiring the SIMD, read the portable version — it is the mathematical definition every fast path must reproduce:

#![allow(unused)]
fn main() {
// src/quants/dot_q4_0.rs:54-71
fn dot_q4_0_q8_0_scalar(x: &[u8], y: &[u8], nb: usize) -> f32 {
    let mut s = 0.0f32;
    for ib in 0..nb {
        let xb = &x[ib * Q4B..];
        let yb = &y[ib * Q8B..];
        let dx = block::fp16_to_f32(u16::from_le_bytes([xb[0], xb[1]])); // weight scale
        let dy = block::fp16_to_f32(u16::from_le_bytes([yb[0], yb[1]])); // activation scale
        let mut si = 0i32;
        for j in 0..16 {
            let v0 = (xb[2 + j] & 0x0F) as i8 - 8;        // low nibble, centered
            let v1 = (xb[2 + j] >> 4) as i8 - 8;          // high nibble, centered
            si += (v0 as i32) * (yb[2 + j] as i8 as i32);
            si += (v1 as i32) * (yb[2 + j + 16] as i8 as i32);
        }
        s += si as f32 * dx * dy;                          // exact int dot × scales
    }
    s
}
}

This is also the honest worked example. Take dx = dy = 1.0 for clarity and a weight byte xb[2] = 0x21 (binary 0010_0001): low nibble 1 − 8 = −7, high nibble 2 − 8 = −6 — two dequantized weights −7·dx and −6·dx from one byte. The i32 accumulator never rounds: the only approximation in the whole kernel happened when the values were quantized.

The Q8_0 kernel is the same skeleton, minus unpacking. With both sides already int8, the loop shrinks to load-scale-multiply-accumulate:

#![allow(unused)]
fn main() {
// src/quants/dot_q8_0.rs:28-48 (core, abridged)
unsafe fn dot_q8_0_q8_0_avx2(x: &[u8], y: &[u8], nb: usize) -> f32 {
    let mut acc = _mm256_setzero_ps();
    for ib in 0..nb {
        let xd = f16_to_f32_bits(*x.as_ptr().add(ib * Q8B).cast::<u16>());
        let yd = f16_to_f32_bits(*y.as_ptr().add(ib * Q8B).cast::<u16>());
        let d = _mm256_set1_ps(xd * yd);            // the two block scales
        let qx = _mm256_loadu_si256(x.as_ptr().add(ib * Q8B + 2).cast::<__m256i>());
        let qy = _mm256_loadu_si256(y.as_ptr().add(ib * Q8B + 2).cast::<__m256i>());
        let ax = _mm256_sign_epi8(qx, qx);          // |qx|, signs moved onto qy
        let sy = _mm256_sign_epi8(qy, qx);
        let dot = _mm256_maddubs_epi16(ax, sy);     // 16 u8×i8 → 8 × i16 dots
        let q = _mm256_cvtepi32_ps(_mm256_madd_epi16(_mm256_set1_epi16(1), dot));
        acc = _mm256_fmadd_ps(d, q, acc);
    }
    hsum_float_8(acc)
}
}

Why the sign_epi8 dance again when Q8_0 values are already signed? Because maddubs insists on unsigned × signed, and the weight side must be the unsigned one — so the same abs-and-transfer trick from the Q4_0 kernel reappears. Once you have seen it twice, every minfer dot kernel reads as a variation on one theme: unpack to centered int8 → exact integer dots → one float fold per block.

This kernel is also the entire lm_head story: the output projection matmul (vocab ≈ 151k rows of Q8_0 for Qwen3) is this loop 151,936 times per token — doc 05's n_out tail-row optimization exists precisely to cut how many of those rows run.

The NEON counterpart: one instruction, 16 MACs. On aarch64 (Apple Silicon, doc 14's home turf) the equivalent of the unpack-and-multiply chain is a single instruction — SDOT, issued through inline assembly because Rust's std::arch exposed no stable intrinsic at the time:

#![allow(unused)]
fn main() {
// src/quants/neon.rs:44-54
#[target_feature(enable = "dotprod")]
pub(super) unsafe fn sdot_vec(acc: int32x4_t, a: int8x16_t, b: int8x16_t) -> int32x4_t {
    std::arch::asm!(
        "sdot {acc:v}.4s, {a:v}.16b, {b:v}.16b",   // acc += 4-way i8 dot of 16-byte lanes
        acc = inout(vreg) acc, a = in(vreg) a, b = in(vreg) b,
        options(nomem, nostack),
    );
    acc
}
}

Each sdot takes two 16-byte int8 vectors and adds sixteen multiply-accumulates into four i32 lanes. The NEON dot_q4_0_q8_0 (src/quants/neon.rs:58) unpacks nibbles with NEON shuffles and drives sdot_vec per 16 bytes — the aarch64 answer to maddubs. MINFER_NO_NEON=1 disables the whole NEON layer (neon_enabled()) and drops to scalar, which is how the optimization campaign A/Bs the SIMD paths; on x86 the counterparts are MINFER_NO_AVX2=1 (whole quants AVX2 layer) and MINFER_NO_AVX512=1 (only the AVX-512/VNNI K-quant dots, falling back to AVX2).

The activation quantizers. The Q8_0 one, entry first:

#![allow(unused)]
fn main() {
// src/quants/quantize_q8_0.rs:85-93 (entry; _buf variant at :95 writes into a caller buffer)
pub fn quantize_row_q8_0(x: &[f32]) -> Vec<u8> {
    let k = x.len();
    debug_assert!(k % 32 == 0);
    let nb = k / 32;
    let mut y = vec![0u8; nb * Q8B];
    quantize_row_q8_0_to(x, &mut y);
    y
}
}

Per 32-value block: find amax = max|x|, store d = amax / 127 as f16 (2 bytes), store each x[i]/d rounded to nearest i8 (32 bytes) → 34 bytes. The AVX2 variant (quantize_avx2, src/quants/avx2.rs:7) is worth reading because it shows the max-reduce and the rounding in registers:

#![allow(unused)]
fn main() {
// src/quants/avx2.rs:7-27 (core, abridged)
unsafe fn quantize_avx2(x: &[f32], y: &mut [u8], k: usize) {
    for i in 0..nb {
        let v0..v3 = /* four 8-float loads: the 32-value block */;
        let ma = _mm256_max_ps(
            _mm256_max_ps(_mm256_andnot_ps(sb, v0), _mm256_andnot_ps(sb, v1)),
            _mm256_max_ps(_mm256_andnot_ps(sb, v2), _mm256_andnot_ps(sb, v3)));
        // sb = -0.0; andnot(sb, v) = |v| — abs without an extra op
        let ms = /* horizontal max of ma */;
        let d = ms / 127.0f32;
        y[yo]..y[yo+1] = f16(d).to_le_bytes();       // the block scale
        let id = if ms != 0.0 { 127.0f32 / ms } else { 0.0f32 };
        let i0 = _mm256_cvtps_epi32(_mm256_round_ps(
            _mm256_mul_ps(v0, _mm256_set1_ps(id)), _MM_ROUND_NEAREST as i32));
        ...
}

Two details to notice: the absolute value is free (_mm256_andnot_ps with -0.0 clears the sign bit), and the zero-block guard (ms != 0.0 → inverse 0) keeps an all-zero block from producing NaNs — a whole layer of 0.0/0.0 if skipped. The debug_assert!(k % 32 == 0) is where doc 02's alignment story pays off — every activation row is a whole number of blocks, so the quantizer never sees a partial block.

The K-quant activations quantizer (quantize_row_q8_k_buf, src/quants/quantize_q8_k.rs:25) fills the 306-byte Q8_K blocks of §2.4. Its NEON worker (:75) shows every field:

#![allow(unused)]
fn main() {
// src/quants/quantize_q8_k.rs:75-114 (core, abridged)
unsafe fn quantize_row_q8_k_buf_neon(row: &[f32], out: &mut [u8]) {
    for s in 0..n_super {
        let blk = &row[s * 256..(s + 1) * 256];
        let o = s * crate::block::Q8KB;                 // 306-byte stride
        let mut amax = 0.0f32;
        for g in 0..64 { /* amax over the 256 values, exact max reduction */ }
        let d = amax / 127.0f32;
        out[o]..out[o+1] = f16(d).to_le_bytes();        // ① block scale, f16
        for g in 0..16 {                                // 16 groups of 16 values
            let q8 = saturating_round(blk[g*16..g*16+16] * (1/d));   // ② int8 quants
            vst1_s8(out.as_mut_ptr().add(o + 2 + base) ..., q8);     //    at o+2..o+258
            bsums[g] = exact int sum of the SATURATED q8;            // ③ group sums
        }
        out[o + 258 + g] = 0;                           // ④ 16 reserved bytes, zeroed
        out[o + 274..o + 306] = bsums;                  //    16 × i16 group sums
    }
}
}

Three details carry weight: ② uses saturating narrowing (clamping to [−128, 127] exactly like the scalar .clamp()), ③ computes the bsums from the saturated values so the integer group sums are exact — the K-quant dot kernels use them for their correction term and any drift there would break parity — and ④ keeps the reserved field zeroed so the region reads deterministically. The AVX2/scalar paths write byte-identical layouts, which is what lets one kernel consume activations from any build.

The payoff: a K-quant dot kernel, walked. All of §2.4's structure (super-scales, 6-bit sub-scales, mins, bsums) exists to serve this loop — dot_q4_k_q8_k_scalar (src/quants/kquant.rs:30-69), the reference every K-quant fast path must match:

#![allow(unused)]
fn main() {
// src/quants/kquant.rs:30-69 (scalar, abridged but complete in structure)
fn dot_q4_k_q8_k_scalar(q4: &[u8], q8k: &[u8]) -> f32 {
    for i in 0..n_super {
        let d    = w_scale(i) * a_scale(i);          // super-scale × activation scale
        let dmin = w_dmin(i)  * a_scale(i);
        let (scales, mins) = block::unpack_q4k_scales(&q4b[4..16]);  // 12 B → 8+8 6-bit values
        // ① the MIN term: Σ mins[s] × (sum of quants in sub-block s)
        let mut mterm = 0i32;
        for s in 0..8 {
            let b0 = bsums(2 * s);  let b1 = bsums(2 * s + 1);   // from Q8_K's ③ above
            mterm += mins[s] as i32 * (b0 + b1);
        }
        sumf -= dmin * mterm as f32;
        // ② the VALUE term: nibble × int8 dots per sub-block, weighted by scales
        for j in 0..4 {
            let mut s_lo = 0i32;
            let mut s_hi = 0i32;
            for l in 0..32 {
                s_lo += (q4b[q4off + l] & 0x0F) as i32 * (q8b[q8off + l] as i8 as i32);
                s_hi += (q4b[q4off + l] >> 4) as i32 * (q8b[q8off + 32 + l] as i8 as i32);
            }
            sumi1 += s_lo * scales[2 * j] as i32;
            sumi2 += s_hi * scales[2 * j + 1] as i32;
        }
        sumf += d * (sumi1 + sumi2) as f32;
    }
    sumf
}
}

The Q4_K value formula is value = q · d_s − min_s per sub-block s (a non-centered 4-bit scheme: instead of subtracting 8 like Q4_0, each sub-block carries its own min). Unrolling the math shows why the kernel has two terms:

Σ_values (q·d_s − min_s)·a        per sub-block
= d_s · Σ(q·a)  −  min_s · Σa     ← the min term is a sum over ACTIVATIONS only

Σ(q·a) is the integer dot ②; Σa per sub-block is exactly what Q8_K's bsums carry — precomputed once at quantization time (doc 10's §3.2 ③) so the kernel never re-touches the activation bytes for the min correction ①. That is the entire reason the Q8_K format exists, and the reason doc 02's byte-size table and this kernel must agree on where bsums live (offset 274 in the 306-byte block — visible in both quantize_row_q8_k_buf_neon and dot_q4_k_q8_k_scalar).

The persistent thread pool. Decode runs ~250 matmuls per token (doc 13), each tiny (one token row through od rows of weights). Spawning threads per matmul measured ~170 µs (kernel.rs's own comment) — against a per-token budget of a few milliseconds, that alone would be the bottleneck:

// src/kernel/pool.rs:242-249 (the dispatch every worker wakes for)
let job = *pool.job.lock().unwrap();
match job {
    PoolJob::MatMul(m) => {
        let (r0, r1) = chunk(pool.n + 1, my_idx, m.od);  // even row split
        unsafe { mm_rows(&m, r0, r1) };
    }
    PoolJob::ParFor(p) => { ... }                        // generic parallel-for
}
pool.done.fetch_add(1, Ordering::SeqCst);

The design (kernel.rs comment block, lines ~274–287):

  • Workers are spawned once, lazily, by get_pool (OnceLock, process lifetime) and spin on an atomic gen counter — 8000 spin iterations, then yield_now, so idle workers cost nothing but wake in microseconds.
  • The submitting thread publishes a MmJob under job, bumps gen, and waits on done == n+1 (the main thread participates as the last worker — no wasted idle main).
  • A gate Mutex serializes submissions: the comment (src/kernel/pool.rs:212-218) records the real hazard — two concurrent callers (parallel tests, the multi-slot server) would clobber job and share done, letting one caller return before its range was computed while workers still read its stack-local context. That is a use-after-free, and the fix is one lock around submit→wait.
  • chunk(parts, idx, total) splits od rows evenly; each row belongs to exactly one worker → bit-identical results at any thread count (§2.3's promise).
  • Worker count: set_cpu_threads (CLI --threads, src/kernel/pool.rs:21) must run before the first matmul because the pool is spawned lazily on first use; cpu_threads() (:28) reads the CPU_THREADS atomic where 0 means auto-detect.
  • The same pool also serves par_for (src/kernel/pool.rs:290) — which is exactly what attention reuses for per-head parallelism in doc 11.

The embedding "matmul". embed_tokens (src/kernel/embed.rs:5) is not a matmul at all: for each token id it walks the quantized token_embd row block-by-block, dequantizing (scale × nibble for Q4_0/Q4_1, scale × byte for Q8_0) into the output row:

#![allow(unused)]
fn main() {
// src/kernel/embed.rs:5-26 (Q4_0/Q8_0/Q4_1 arm, abridged)
pub fn embed_tokens(ids: &[u32], t: &crate::tensor::Tensor, out: &mut [f32], ne: usize) {
    for (ti, &id) in ids.iter().enumerate() {
        let idx = id as usize;                       // the token's row in the table
        for b in 0..nbp {                            // blocks in one embedding row
            let off = (idx * nbp + b) * bb;
            let d = fp16_to_f32(...);                // block scale
            ...                                      // dequantize 32 values → out row
        }
    }
}
}

This is the GetRows node of doc 05 made concrete — "the embedding table is a lookup" — and it exists here because CPU and GPU backends share the same row getter, keeping embeddings byte-for-byte identical across backends.

3.3 Design choices (why this shape and not another)

Why not dequantize weights to f32 at load? §2.1's arithmetic: 6× the RAM, 7× the per-multiply bandwidth, seconds of extra load time — and the GPU backends want the quantized bytes anyway (doc 14's shaders consume the same block layouts). The quantized bytes are not an intermediate representation; they are the storage format.

Why Q8_0 (int8) for activations, not f32 or int4? The inner loop is the hot code; int8×int8 has dedicated silicon on both ISAs (§2.2). f32 activations would forfeit maddubs/SDOT; int4 activations would double the quantization error exactly where values are least pre-calibrated (activations change every call; weights were tuned at conversion time). llama.cpp's CPU path makes the same choice, and matching it is what makes the bit-parity tests possible at all.

Why do K-quant weights get a different activation format (Q8_K)? The K-quant kernels' inner loop processes 256 values per super-block and needs the activation side in matching 256-value groups with exact group sums (bsums) for its correction term. A 34-byte Q8_0 stream would force the kernel to re-group and re-sum activations per call; the 306-byte Q8_K block carries everything precomputed (§2.4). The pairing is enforced in one place (cpu_quant_matmul_f32) so it cannot drift.

Why per-row thread ownership instead of splitting the inner reduction? Splitting a dot product across threads needs a reduction (adding partial sums in some order), and float addition is not associative — results would drift with thread count. One row per worker makes --threads a pure performance knob with zero numerical effect, which the whole verification methodology (greedy token equality across machines, doc 12's seeded gates) quietly depends on.

Why is a Vec<u8> allocated for quantized activations on every matmul call? It is sized per call (nt × row_bytes) because nt changes between prefill and decode. Doc 07's allocator owns the node buffers; this scratch lives one call deep and is invisible to the graph. (CPU_OPTIMIZATIONS.md's "batched QKV shared quantization" record — the prefill gain in COMPUTE-GRAPH-DESIGN §12 — exists precisely because three sibling matmuls each re-quantizing the same normed row was visible in the profile.)

Why are Q2_K / Q3_K / I-quants not supported? Each extra format costs a hand-written AVX2 + NEON + scalar kernel triplet (×2 for the activation-format pairing) plus parity tests. The supported set — Q4_0/Q4_1/Q5_0/Q5_1/Q8_0 (32-value) and Q4_K/Q5_K/Q6_K (256-value) — covers every GGUF quant llama.cpp recommends for quality at 4–8 bits; Q2_K/Q3_K trade real quality for size, and the I-quants' lookup-table design resists this kernel shape. docs/SUPPORT-MATRIX.md is the authoritative list.

3.4 Pitfalls & invariants

  • Block-size divisibility is a format contract, not a hint. id % 32 == 0 (and id % 256 == 0 for K-quants) is guaranteed by GGUF conversion tools and asserted in the quantizer; a model that broke it would silently misindex without the assert.
  • The transposed-output bug that decode hid (src/vec_ops/vec.rs:366-370, the mat_mul_f32 comment): an earlier version wrote C[row*n + col] — an [m,n] output — while every caller wanted [nt][od]. For decode (nt == 1) the two layouts coincide, so no decode-only test caught it; prefill output was transposed. Lesson the repo kept: test with nt > 1 or the token-major convention will bite.
  • Row ownership = bit-identical parallelism. Never "optimize" the pool into splitting a row's reduction; that trades away the determinism the verification gates rely on (§3.3).
  • The gate lock is load-bearing (src/kernel/pool.rs:212-218): removing it works in single-threaded tests and corrupts memory the first time two threads submit concurrently (the server's multi-slot path).
  • K-quant weights need Q8_K activations, 32-value weights need Q8_0 — crossing the pairing (e.g. feeding Q8_0 blocks to dot_q4_k_q8_k) misindexes the super-block scales. The cpu_quant_matmul_f32 branch exists to make the pairing unstateable from the call site.
  • The activation-Q8_K layout is kernel-pair-defined — 306 bytes as written by quantize_row_q8_k_buf (src/quants/quantize_q8_k.rs:20-25 comment), not the BlockQ8_K struct layout (block.rs:173); doc 02 flagged the same nuance on the on-disk side. When touching either side, re-verify the quantizer→kernel byte contract together.
  • Scales fold once per block, integers stay exact — any refactor that converts intermediate integer dots to float mid-block changes the numerics and breaks parity with llama.cpp.

4. Observe & verify

  • cargo test quants:: / cargo test kernel:: — per-kernel parity tests: every SIMD path is checked against the scalar reference you read in §3.2 (and, for the formats llama.cpp also ships, against llama.cpp-produced reference values).
  • cargo test --release greedy token gates — end-to-end: doc 12's seeded/greedy gates (-n 32 --greedy --seed 42) would shift the instant any kernel changed one ulp of accumulated behavior.
  • MINFER_TIMING=1 ./target/release/minfer <model> "hi" — splits per-token wall time into sampling vs forward (doc 09 §4); on CPU, forward is these kernels.
  • --threads N — the pool's worker count; output must be identical for any N (that property is itself tested).
  • MINFER_NO_NEON=1 (aarch64) — forces scalar, the A/B lever for the NEON layer.
  • MINFER_NO_AVX2=1 (x86) — forces the whole quants AVX2 layer scalar; MINFER_NO_AVX512=1 drops only the AVX-512/VNNI K-quant dots to AVX2.

5. Cross-references

  • 02 — GGUF load §2.3/§3.2 — where the block layouts and byte-size formulas come from (and the on-disk Q8_K nuance this doc's §2.4 completes).
  • 03 — Model dispatch and weights — why Tensor.data is the raw GGUF bytes.
  • 08 — The scheduler §3.2 — the MatMul dispatch arm that lands here.
  • 11 — Attention, vec ops, and the KV cache — the non-matmul half of the layer; reuses this doc's par_for pool for per-head parallelism.
  • 13 — The decode loop — why decode is bandwidth-bound (250 matmuls/token, each streaming weights).
  • 14 / 15 — the same math with GPU execution models (f32 activations on Metal; int8 MMQ tensor-core GEMM on CUDA).
  • docs/CPU_OPTIMIZATIONS.md, docs/SUPPORT-MATRIX.md — the optimization history and the authoritative quant-format matrix.

← 09 — Prefill: the first forward · Index · 11 — Attention, vec ops, and the KV cache →

11 · Attention, vec ops, and the KV cache

Stage: CPU matmul kernels (doc 10) → this stage: every non-matmul op of the layer → the sampler (doc 12). Code: src/vec_ops.rs (rms_norm_f32, vec_soft_max_f32, vec_silu_f32, mat_mul_f32), src/graph/cpu_backend.rs (cpu_rope:469, the KvcacheStore/KvcacheLoad/Attn arms, cpu_gqa_attn:503, attn_heads:572), src/graph/builder.rs:323 (the attn node contract).

1. Background — where this stage sits

Doc 05 built the layer as a graph; docs 08–10 executed its matmuls. What is left is everything a transformer does between and around those matmuls — and one op that is unlike any other in the network:

  • RMSNorm — rescale each token's vector so its root-mean-square is 1, then apply learned per-dimension gains. Runs twice per layer, plus a final one before the output matmul.
  • RoPE — rotate the query and key vectors by an angle that depends on the token's position. This is how a transformer without recurrence knows word order.
  • Attention — the op the whole architecture orbits: every token looks at other tokens and mixes their information. It is the only op whose inputs reach across positions, which is why the engine needs a KV cache, a causal mask, and positions-as-data (docs 05/07/09 set those up; this doc shows the mechanics).
  • SiLU and the elementwise glue — add (residuals), mul, scale, softmax: short vector passes that doc 06's fusion pass already minimized.

For a beginner, the mental model of one decoder layer is:

h ──► RMSNorm ──► [W_q W_k W_v matmuls] ──► RoPE(q, k) ──► store K,V into the cache
                                                                │
              ┌─────────────────────────────────────────────────┘
              ▼
   attention: each query token reads the cached K/V of its allowed prefix
              ──► [W_o matmul] ──► (+ residual) ──► RMSNorm ──► FFN ──► (+ residual) ──► h'

This document walks the four mechanisms in order: RMSNorm (§3.2.1), RoPE (§3.2.2), the KV store/load pair (§3.2.3), and attention (§3.2.4), each with the actual code and a worked numeric example. Everything here is the CPU implementation; the GPU backends run the same math with different execution models (docs 14/15).

One modern-model wrinkle to know up front: Qwen3 makes the query head dimension (hd) and the key/value head dimension (hd_kv) decoupled — they can differ per model (doc 03's loader). Every function in this doc carries both (hd and hd_kv parameters) and cpu_gqa_attn refuses (Err) any configuration where hd < hd_kv, because a query head shorter than its key cannot be dotted against it. Qwen2/Qwen2.5 models simply have hd == hd_kv.

2. Principle — attention from zero

2.0 The residual stream: why the norm sits before each block

One piece of architecture context explains half the ops in this doc. A decoder layer does not compute h' = ffn(attn(h)) — it computes:

h = h + attn_out( rms_norm(h) )      // attention sub-block, pre-norm
h = h + ffn_out( rms_norm(h) )      // FFN sub-block, pre-norm

This is called pre-norm (the norm comes before the attention/FFN instead of after), and the running h that both sub-blocks read and add onto is called the residual stream: a highway where each layer's contribution is an update, not a replacement. Two consequences worth knowing:

  • If any sub-block is useless for a given token, the layer can output it with (near-)zero weight and the token's information passes through untouched — this is why very deep models train stably at all.
  • The two + operations are Op::Add nodes (doc 05), the plainest vector op in the engine — and they are why doc 06's fusion pass cares so much about elementwise buffer reuse: in the 0.5B graph the residual adds are among the most frequent nodes.

2.1 What attention computes

A transformer represents each token as a vector (the hidden state, dimension d = n_embd). Attention lets token t build a new representation out of the other tokens' vectors. Mechanically, each token's vector is projected three ways — by three of doc 10's matmuls:

  • a query vector q — "what am I looking for?" (W_q),
  • a key vector k — "what do I contain?" (W_k), and
  • a value vector v — "what do I hand over if someone attends to me?" (W_v).

The projections are split into heads — n_head independent attention channels of dimension hd each (n_embd = n_head × hd; Qwen2.5-0.5B: 14 × 64 = 896). Heads attend in parallel and are concatenated before the output projection W_o.

Then, for query token t in head h:

  1. Score every earlier-or-equal token s (s ≤ t): score[t][s] = q_t · k_s × scale, with scale = 1/√hd. The dot product measures similarity between what t seeks and what s offers; the scale keeps the numbers in a range where softmax behaves (a raw dot over hd elements grows like hd, and softmax over huge numbers degenerates).
  2. Mask the future: s > t is forbidden (a token must not see tokens that come after it — that is the causal property that makes one pass over the prompt meaningful). Implemented not by a mask matrix but by length: s only ranges over [0, pos[t]] (§3.2.4).
  3. Softmax the scores into weights that sum to 1.
  4. Output the weighted sum of the values: out_t = Σ_s weight[t][s] · v_s.

The matmuls are doc 10's job. This doc is steps 1–4 plus the ops that condition the inputs (RMSNorm, RoPE).

2.2 Why the KV cache exists

Consider decode, step 100 (doc 13): one query token needs scores against the keys of tokens 0..99. The keys and values of tokens 0..98 were already computed on earlier steps — recomputing them would mean re-running W_k/W_v over the entire history every step, making step 100 cost as much as prefill of 100 tokens. Instead, the engine stores every token's K and V rows as they are produced and attention reads them back.

That is the kvcache_store / kvcache_load node pair from doc 05, backed by doc 07's persistent per-layer regions: each layer owns two buffers (K and V) of n_kv_embd × n_ctx floats, and the store node writes this step's rows at the positions the positions input carries. "KV positions are data" — the write offset comes from an input, not from any internal counter — is what lets the same graph serve prefill (write 100 positions), decode (write 1 position), and multi-turn continuation (write positions 100..120 — doc 13's conversation path).

The cost arithmetic, concretely for Qwen2.5-0.5B (24 layers, n_kv_embd = 128): one position costs 128 floats × 4 B × 2 (K and V) × 24 layers = 24 KB — the number doc 09 quoted. For Qwen3-4B (36 layers, n_kv_embd = 1024): 1024 × 4 × 2 × 36 = 288 KB per position, i.e. 33.6 MB per layer at n_ctx = 4096 (doc 07's number). GQA (§2.4) is what keeps that from being much larger.

2.3 RoPE: position without position embeddings

Dot products are permutation-invariant: shuffle the input tokens and every dot product stays the same. A transformer needs some way to know that "dog bites man" differs from "man bites dog". Older models added a fixed position vector to each token's embedding; modern LLM-family models use rotary position embeddings (RoPE): rotate each q and k vector by an angle proportional to its position, in consecutive 2-dimensional subspaces.

The property that makes it work: for a query at position m and a key at position n rotated in the same 2-D subspace, their dot product depends only on m − n — the relative distance — because rotating both vectors by their own angles leaves a net rotation of m − n between them. Attention scores then encode "how far apart are these tokens", which generalizes to positions never seen in training far better than absolute embeddings.

The angles: in subspace i (of hd/2), the angular speed is freq_i = 1 / base^(2i/hd) — low subspaces rotate fast (fine, local position distinctions), high subspaces rotate slowly (coarse, long-range distinctions). base is freq_base from the model's hyperparameters (e.g. 1,000,000 for Qwen2.5 — doc 03's HParams). This is also why one long-context trick is simply scaling those frequencies (freq_scale in the RoPE meta — NTK/RoPE-scaling territory).

2.4 GQA: many query heads share few key/value heads

Attention runs in parallel heads: n_head independent query/attention channels, each with its own smaller dimension hd. Classic multi-head attention gives every query head its own K/V heads, so the KV cache costs n_head × hd per token per region.

Grouped-query attention (GQA) shares: n_head_kv K/V heads serve all n_head query heads, with query head h reading kv head hk = h / (n_head / n_head_kv). Qwen2.5-0.5B: 14 query heads, 2 kv heads → each kv head serves 7 query heads → the KV cache is 7× smaller than classic MHA with negligible quality loss (the models are trained that way). The code and the cache arithmetic both inherit this: the K/V regions are n_kv_embd = n_head_kv × hd_kv wide (128 floats for 0.5B), and doc 07's per-position byte cost follows from it.

The attention scale is defined per model next to the head geometry it depends on:

#![allow(unused)]
fn main() {
// src/models/qwen2/loader.rs:36-38 (Qwen3 has its own at loader.rs:48)
pub fn attention_scale(&self) -> f32 {
    1.0 / (self.n_embd_head() as f32).sqrt()
}
}

and flows into the graph as attn_scale (models/qwen2/graph.rs:93), landing in every AttnMeta (graph.rs:200-207). Doc 05 covered the meta's plumbing; here it is the c.scale of §3.2.4.

2.5 What changes on the GPU (preview)

Docs 14/15 run the same formulas with different execution models; three deltas to keep in mind so nothing here surprises you later:

  • Storage: Metal/CUDA may keep the K/V regions in f16 (MINFER_CACHE_TYPE=f16), halving the bandwidth of §2.2's arithmetic. The CPU, CUDA and Metal (the last since #310) store f32 by default and offer a packed Q8_0 cache instead (MINFER_CACHE_TYPE=q8_0, C4): a cell is a Q8_0 block row — ceil(n_kv_embd/32 * 34) bytes instead of 4 * n_kv_embd, 3.76x smaller measured — with the store quantizing and the attention read dequantizing the window it is about to use (§3.3).
  • Shape: the GPU attention kernels tile the (query, key) matrix and apply the softmax online — max and sum accumulate block-by-block instead of one full pass — the "flash attention" trick; the CPU path computes full rows because everything already fits in cache.
  • Parallelism axis: CPU splits by head (§3.3); CUDA additionally splits the KV dimension across blocks and reduces (split-KV), because a GPU has thousands of threads and only 14–40 heads to give them.

The invariants survive all three deltas: positions are data, windows come from pos[t]+1, and store-before-attention still holds.

3. Implementation

3.1 Data in / data out

OpReadsWrites
RmsNorm[nt][d] f32 + d gains[nt][d] f32 (aliases its input — doc 07)
RoPEq or k [nt][n_head×hd], positions [nt] (I32-as-f32)same buffer, rotated in place (aliases — doc 07)
KvcacheStorek [nt][nkt], v [nt][nkt], positionsthe layer's persistent K/V regions at those positions
KvcacheLoadnothing (a view of the K region)nothing
Attnq [nt][n_head·hd], the K/V regions, positions[nt][n_head·hd] f32
Softmax / SiLU / Add / Mulelementwise f32elementwise f32

nkt is the K/V row width (n_kv_embd); positions ride as I32-in-f32 bit patterns (doc 07's fill_input_i32), and every arm decodes them with .to_bits() as usize.

3.2 Key code

3.2.1 RMSNorm — src/vec_ops/rms_norm.rs:8

The formula: y = x / sqrt(mean(x²) + eps) × weight. The scalar fallback shows every piece:

#![allow(unused)]
fn main() {
// src/vec_ops/rms_norm.rs:8-31 (scalar path; AVX2 path at :35)
pub fn rms_norm_f32(n: usize, y: &mut [f32], x: &[f32], eps: f32) {
    let mut sum_sq = 0.0f64;
    for i in 0..n {
        sum_sq += (x[i] as f64) * (x[i] as f64);       // ① sum of squares, f64
    }
    let mean = (sum_sq / n as f64) as f32;
    let scale = 1.0 / (mean + eps).sqrt();             // ② 1/rms(+eps)
    if y.as_ptr() != x.as_ptr() {
        vec_cpy_f32(n, y, x);                          // ③ skip the copy when aliased
    }
    vec_scale_f32(n, y, scale);                        // ④ normalize
    ...                                                // ⑤ multiply by gains (weight)
}
}

Design notes a beginner should keep:

  • No mean subtraction. Classic LayerNorm subtracts the mean before normalizing; RMSNorm skips it (mean of squares directly). One less pass over the vector, and empirically it works as well — llama.cpp's models all use it, and minfer matches them op-for-op. eps comes from the model hyperparameters (f_norm_rms_eps, doc 03) and just keeps sqrt away from zero for an all-but-zero vector.
  • The accumulation is f64 (①) — 896-wide sums in f32 would lose real precision; the AVX2 path keeps the same f64 accumulator semantics so both paths agree.
  • Line ③ is doc 07's aliasing made visible: the fusion/allocator pass maps RMSNorm's output onto its input's buffer where legal, and the copy self-suppresses by pointer comparison. rms_norm_fused_f32 (src/vec_ops/rms_norm.rs:78) is the variant that also multiplies the gains in one pass.

Worked example: x = [1, 2, 3, 4], eps ≈ 0. mean(x²) = (1+4+9+16)/4 = 7.5; scale = 1/√7.5 ≈ 0.365; normalized x ≈ [0.365, 0.730, 1.095, 1.461], then element-wise multiplied by the learned gains.

3.2.2 RoPE — the node arm (cpu_backend.rs:343) and cpu_rope (:469)

The RoPE node arm shows the two conventions this series keeps meeting — positions decoded from f32 bits, and in-place execution:

#![allow(unused)]
fn main() {
// src/graph/cpu_backend.rs:343-356 (abridged)
Op::RoPE { style } => {
    let meta = match &node.meta { NodeMeta::Rope(m) => m, ... };
    let nh = meta.n_head;
    let hd = meta.hd;
    let nt = node.out_shape[1];
    // positions are I32 bit patterns in ins[1]
    let pos: Vec<usize> = (0..nt).map(|t| ins[1][t].to_bits() as usize).collect();
    out.copy_from_slice(ins[0]);                 // free when aliased (doc 07)
    cpu_rope(out, &pos, nh, hd, meta.freq_base, meta.freq_scale, *style);
    Ok(())
}
}

cpu_rope then does the rotation of §2.3:

#![allow(unused)]
fn main() {
// src/graph/cpu_backend.rs:469-501 (core loop, abridged)
pub(crate) fn cpu_rope(x: &mut [f32], pos: &[usize], nh: usize, hd: usize,
                       freq_base: f32, freq_scale: f32, style: RopeStyle) {
    let half = hd / 2;
    let mut freqs = [0.0f32; 128];
    for i in 0..half {
        freqs[i] = freq_scale / freq_base.powf((2 * i) as f32 / hd as f32);  // ① angular speeds
    }
    for t in 0..pos.len() {
        let p = pos[t] as f32;
        for h in 0..nh {
            let b = t * nh * hd + h * hd;
            for i in 0..half {
                let th = p * freqs[i];                       // ② this subspace's angle
                let (sn, cs) = th.sin_cos();
                let (i0, i1) = match style {
                    RopeStyle::NonInterleaved => (b + i, b + i + half),  // ③ Qwen2 layout
                    RopeStyle::Interleaved   => (b + 2*i, b + 2*i + 1),  //    Llama layout
                };
                let (x0, x1) = (x[i0], x[i1]);
                x[i0] = x0 * cs - x1 * sn;                   // ④ 2-D rotation
                x[i1] = x0 * sn + x1 * cs;
            }
        }
    }
}
}
  • ①: the angular speeds — subspace 0 spins fastest, the last subspace slowest (§2.3). The table holds hd/2 entries; the fixed [128; ...] bound covers every supported head dim.
  • ②: the angle is position × speed — position enters only here, from the pos input.
  • ③: the two layout styles are a pure memory convention — which two slots form a rotating pair. Qwen2/Qwen3 put the pairs at (i, i+half) (all "first halves" contiguous — GGUF's NEOX convention); Llama interleaves (2i, 2i+1). Same rotation math, different addresses; ModelDef::rope_style() (doc 03) picks per model, and getting it wrong silently scrambles position information (it is one of the first things the Qwen2.5 bring-up docs checked — see docs/DEBUGGING-PLAN.md H1).
  • ④: the 2-D rotation, straight from §2.3.

Worked example: hd = 4 (half = 2), base = 10⁴ → freqs = [1.0, 10⁻¹]. Token at pos = 2, pair 0: angle 2×1.0 = 2 rad → (x0,x1) ← (x0·cos2 − x1·sin2, x0·sin2 + x1·cos2). Pair 1 rotates 10× slower — exactly the "fine vs coarse" structure of §2.3.

3.2.3 The KV store / load pair — cpu_backend.rs:143 and :381

The store arm is where "KV positions are data" becomes memory writes:

#![allow(unused)]
fn main() {
// src/graph/cpu_backend.rs:143-176 (core, abridged)
if let Op::KvcacheStore { layer } = &node.op {
    let (k_id, v_id) = kv_pair.ok_or_else(...)?;          // doc 07's persistent regions
    if k_id != out_buf { return Err("KV store out buffer must be the K region".into()); }
    let nkt = node.out_shape[0];                           // K/V row width (n_kv_embd)
    let n_ctx = node.out_shape[1];
    let nt = self.buffers[in_bufs[0]].len() / nkt;
    let pos: Vec<usize> = self.buffers[in_bufs[2]].iter()
                              .map(|b| b.to_bits() as usize).collect();  // I32-as-f32
    let (k_dst, v_dst): (&mut [f32], &mut [f32]) = /* split_at_mut over the two regions */;
    for t in 0..nt {
        let p = pos[t];
        if p >= n_ctx {
            return Err(format!("KV store position {p} >= n_ctx {n_ctx}"));  // ① hard error
        }
        let ks = p * nkt;                                  // ② THE write offset
        k_dst[ks..ks + nkt].copy_from_slice(&k_src[t * nkt..(t + 1) * nkt]);
        v_dst[ks..ks + nkt].copy_from_slice(&v_src[t * nkt..(t + 1) * nkt]);
    }
}
}

Two things to internalize:

  • ② is the entire KV cache write: a memcpy of one K row and one V row to position × row_width inside the layer's persistent regions. No scaling, no math — the K/V rows arriving on in_bufs[0..1] were already computed by the W_k/W_v matmuls and rotated by RoPE.
  • ① is the doc 08 error contract in miniature: a position out of range is a topology/contract violation, so it returns Err and aborts the run — it never silently clamps and continues, because a clamped write would corrupt a different token's cache row and you would see it a thousand tokens later as subtly wrong text.
  • The load arm is one line (cpu_backend.rs:381): Op::KvcacheLoad { .. } => Ok(()). The load node's output buffer is the K region (doc 07 mapped it directly), so "loading" is a bookkeeping view — no data moves. The scheduler's build-order guarantee (doc 08) is what makes the view safe: this layer's store always executes before this layer's attention.

3.2.4 GQA attention — the Attn arm (cpu_backend.rs:388) and attn_heads (:572)

The node arm resolves the regions and computes one derived quantity — the current KV length:

#![allow(unused)]
fn main() {
// src/graph/cpu_backend.rs:388-428 (core, abridged)
Op::Attn { .. } => {
    let meta = ...;                                        // AttnMeta: n_head, hd, nkt, scale, layer
    let (k_id, v_id) = kv_pair.ok_or_else(...)?;           // this layer's K and V regions
    let k_slice /*, v_slice */ = /* the two persistent regions, borrowed */;
    let n_ctx = k_slice.len() / nkt;
    let nkv = (0..nt).map(|t| ins[2][t].to_bits() as usize + 1).max()   // ① max position + 1
                     .unwrap_or(0).min(n_ctx);
    let pos: Vec<usize> = (0..nt).map(|t| ins[2][t].to_bits() as usize).collect();
    cpu_gqa_attn(ins[0], k_slice, v_slice, &pos, nt, nkv,
                 meta.n_head, meta.n_head_kv, meta.hd, meta.hd_kv, nkt, out, meta.scale)?;
}
}

① is the causal mask in data form: the attention window is max(positions) + 1 — during prefill of 100 tokens that is 100; during decode at position 100 it is 101. There is no mask tensor anywhere; the window is the mask, derived from the same positions input the store used. (builder.rs:323-325 documents the contract: vl = pos[t]+1.)

cpu_gqa_attn (:503) validates the head geometry (hd < hd_kv → Err — the Qwen3 decoupled-dims guard from §1), then farms head ranges to the same thread pool as the matmuls (kernel::par_for, doc 10), each worker with a private scores buffer. Heads never reduce against each other → bit-identical at any thread count, the same invariant as doc 10's row ownership. Single-threaded fallback when --threads 1 or n_head < 2.

The per-head worker does the math of §2.1:

#![allow(unused)]
fn main() {
// src/graph/cpu_backend.rs:572-625 (core loop, abridged; runs per head range via par_for)
let gqa = c.nh / c.hk;                                     // e.g. 14/2 = 7
for h in h0..h1 {                                          // this worker's query heads
    let hk = h / gqa;                                      // ① my shared kv head
    for t in 0..c.nt {
        let vl = (*c.pos.add(t) + 1).min(c.nkv);           // ② causal window for token t
        for kv in 0..vl {
            let s = vec_dot_f32(c.hd_kv, &q_row, &k_row(kv)) * c.scale;   // ③ score
            scrs[kv] = s; if s > mx { mx = s; }            //    (max tracked on the fly)
        }
        for kv in vl..c.nkv { scrs[kv] = f32::NEG_INFINITY; }             // ④ mask = −∞
        let sm = vec_soft_max_inplace_f32(c.nkv, &mut scrs, mx);          // ⑤ softmax
        vec_scale_f32(c.nkv, &mut scrs, (1.0 / sm) as f32);
        out_row.fill(0.0);
        for kv in 0..c.nkv {                                              // ⑥ weighted V sum
            vec_muladd_f32(c.hd_kv, out_row, &v_row(kv), scrs[kv]);
        }
    }
}
}

Walk it against §2.1:

  • ①: GQA's only appearance in the code — one integer division mapping query head → kv head (0.5B: heads 0–6 read kv head 0, heads 7–13 read kv head 1).
  • ②③: scores are q·k × scale over the causal window vl = pos[t]+1.
  • ④: tokens beyond the window get -INF scores rather than a shorter loop — see ⑥ for why that is safe and cheap.
  • ⑤: softmax with max subtraction (§3.2.5). Weights now sum to 1.
  • ⑥: the output is the weighted sum of V rows. The loop runs over all nkv rows, including masked ones — their softmax weight is exactly exp(−∞ − mx) = 0.0, and 0.0 × v = 0.0 contributes nothing. Unwritten region bytes would be zeros even if read, so there is no uninitialized-memory hazard either.

Worked example (one query head, hd_kv = 4, GQA 1:1 for simplicity). Query at pos = 1 (the second token); cache holds K rows k0 = [1,0,0,0], k1 = [0,1,0,0], V rows v0 = [1,2,0,0], v1 = [3,4,0,0]; q = [1,1,0,0], scale = 1/√4 = 0.5; suppose the KV region is sized for 3 (nkv = 3) so there is a masked third slot:

vl = pos+1 = 2                          (token 1 may see tokens 0 and 1)
score0 = q·k0 × 0.5 = 1 × 0.5 = 0.5
score1 = q·k1 × 0.5 = 1 × 0.5 = 0.5
score2 = −∞                             (④ masked: beyond the causal window)
softmax([0.5, 0.5, −∞]) = [0.5, 0.5, 0.0]
out = 0.5·v0 + 0.5·v1 + 0.0·v2 = [2.0, 3.0, 0.0, 0.0]

A query aligned equally with both cached keys blends both values. Change the query to [1,0,0,0] and score0 wins (0.5 vs 0.0 after scale → softmax ≈ [0.73, 0.27, 0.0]) and the output tilts toward v0 — that tilting is, mechanically, all "attention" is.

The same example one step later (decode). Suppose step 3 now emits token at pos = 2 with q2 = [0,1,1,0]; the store arm appends k2 = [1,1,0,0], v2 = [5,0,0,0] at offset 2 × nkt (② in §3.2.3); the attention arm re-derives nkv = max(2)+1 = 3. Nothing about the graph changed — the positions input grew by one element, and the window grew with it:

token 0 sees:  vl = 1  →  [w0]                     (still blind to 1, 2)
token 1 sees:  vl = 2  →  [w0, w1]
token 2 sees:  vl = 3  →  [w0, w1, w2]
score2·k0 = 0×0.5 = 0.0 · k1 = 1×0.5 = 0.5 · k2 = 1×0.5 = 0.5
softmax([0.0, 0.5, 0.5]) ≈ [0.27, 0.37, 0.37]
out2 ≈ 0.27·v0 + 0.37·v1 + 0.37·v2

That is decode in miniature: one new K/V row per step, windows growing monotonically, and every past token's cached rows read again without recomputation — the entire reason §2.2's cache exists. (Doc 13 shows the loop that drives it and doc 07 the regions that hold it.)

3.2.5 Softmax and SiLU — src/vec_ops/softmax.rs:8, src/vec_ops/silu.rs:8

#![allow(unused)]
fn main() {
// src/vec_ops/softmax.rs:8-32 (scalar path; caller supplies the max)
pub fn vec_soft_max_f32(n: usize, y: &mut [f32], x: &[f32], max: f32) -> f64 {
    let mut sum = 0.0f64;
    for i in 0..n {
        let val = (x[i] - max).exp();      // ① subtract max BEFORE exp
        y[i] = val;
        sum += val as f64;                 // ② f64 sum, returned for normalization
    }
    sum
}
}

① is the numerical-stability trick the whole series keeps meeting: softmax is invariant under subtracting any constant, but e^1000 overflows f32 while e^(1000−1000) = 1 does not. So every caller finds the max first — the standalone Softmax node arm (cpu_backend.rs:357-368) shows the full ceremony (scan for max → copy → softmax → caller divides by the returned sum), while attention's hot path uses the in-place variant (vec_soft_max_inplace_f32, :276) since its scores buffer is scratch anyway. The f64 sum ② then normalizes without precision loss.

SiLU is one formula — silu(x) = x / (1 + e^(−x)) (vec_silu_f32, :158, scalar shown):

#![allow(unused)]
fn main() {
for i in 0..n { y[i] = x[i] / (1.0 + (-x[i]).exp()); }
}

It is the FFN's activation (doc 05's silu(gate) × up); doc 06 fused it into SwiGLU, whose CPU execution (two passes over vec_silu_f32 then vec_mul_f32) you saw in doc 06 §3. The add/mul/scale/muladd helpers are the same pattern — a short SIMD-able loop each — and attn_heads ⑥ uses vec_muladd_f32 for the weighted-V accumulation. The f32 weights path (mat_mul_f32, src/vec_ops/vec.rs:371) closes the loop back to doc 10: norm biases and F32 tensors skip quantization entirely and use this plain dot-product matmul.

3.3 Design choices (why this shape and not another)

Why is attention the op everything else serves? Every matmul is per-token: token t's matmul output depends only on token t's input row. Attention is the only op whose output depends on other tokens — that is where the model's ability to relate words to each other lives, and it is the only reason the engine needs positions, a causal window, and a cross-call cache. Remove attention and the remaining stack is just per-token transforms that a single forward could do in any order.

Why length-as-mask instead of a mask matrix? llama.cpp-style engines could build an [nt, nt] additive mask; minfer derives the window from positions (nkv = max(pos)+1, vl = pos[t]+1). That is cheaper (no mask buffer, no per-score mask add), it composes with multi-turn continuation for free (a resumed sequence's positions just continue — doc 13), and it keeps the "positions are data" invariant doing double duty. The −∞ tail (④) is the one concession, and it exists so the softmax normalize step can treat scores as one fixed-length buffer.

Why store K/V at input-carried positions instead of an internal n_past counter? A counter inside the store would make the graph's behavior depend on hidden execution state — the second decode step would write somewhere else than the first, with identical topology. With positions as data, the same rebuilt graph is correct for prefill, for decode, and for resuming a saved conversation (doc 13's conversation.rs just continues the position sequence); doc 07's reuse invariant stays intact.

Why compute scores in f32 (vec_dot_f32) instead of the int8 trick? Doc 10's Q8_0 machinery quantizes matmul activations; attention scores are computed once per (query, key) pair, and quantizing q/k per score would cost more than it saves — the dot is over hd_kv (≤128) elements, not n_embd. The K/V storage is where bandwidth matters (hence the GPU f16 cache, below), not the score math on CPU.

Why is the KV cache f32 on CPU while GPUs offer f16? The CPU path never re-quantizes attention inputs; f16 K/V on CPU would add a conversion per score for negligible bandwidth win at these sizes. The GPU backends, where bandwidth per token is the budget, do offer the f16 cache (MINFER_CACHE_TYPE=f16, docs 03/14). The CPU's own trade is different and is about footprint, not bandwidth: MINFER_CACHE_TYPE=q8_0 (C4) stores packed Q8_0 cells, 3.76x smaller regions on both cached models, and f16 is refused there because the path has no f16 KV kernel. Same invariant — regions are persistent and position-addressed — different storage per backend.

Why parallelize over heads instead of over tokens? Heads are perfectly independent (① is the only cross-head coupling, and it is read-only), so the split has zero communication; splitting over tokens would share each head's scrs buffer across workers. It also composes with decode: at nt=1 the token axis is empty, but 14 heads still parallelize the dot products.

Why refuse hd < hd_kv instead of supporting it? A query head shorter than its key cannot be dotted against it without a padding convention; no supported model needs it (Qwen3's decoupling goes the other way where relevant), and a silent pad would hide model-definition mistakes. Fail loudly (doc 08's error contract).

3.4 Pitfalls & invariants

  • positions[i] < n_ctx is a caller obligation — the store arm enforces it with Err (① in §3.2.3); doc 09 showed the CLI's double clamp making that hold. Silent clamping would be a data-corruption bug.
  • Store must execute before the attention that reads the region — guaranteed only by build-order execution (doc 08's invariant 5); a "smart" reordering that hoisted attention above store would read stale rows. This is why kvcache_store/kvcache_load are explicit nodes rather than hidden side effects.
  • RoPE styles are model-level, not per-call — mixing NonInterleaved and Interleaved on one model silently produces wrong scores with plausible magnitudes; it was suspect #1 in the Qwen2.5 bring-up (docs/DEBUGGING-PLAN.md H1) precisely because nothing crashes.
  • The attention window derives from positions, never from a counter — if you find yourself reaching for n_past inside a kernel, you are breaking the reuse invariant (docs 05/07).
  • Softmax needs the max first — skipping the subtraction works on toy inputs and overflows on real logits (score magnitudes grow with context); the AVX2 path preserves the same subtract-then-exp order.
  • Masked V rows are multiplied by exactly 0.0 — any change to the −∞ masking that produced NaN (−∞ × 0 patterns) would poison the output; exp(−∞) → 0.0 is the behavior the tail loop relies on.
  • bsums-style exactness discipline applies here too — the Q8_K bsums lesson of doc 10 (sums computed from the saturated integers) has its mirror in attention: 1/sm normalization happens once, outside the V accumulation, so every worker's output matches the single-threaded order bit for bit.

4. Observe & verify

  • cargo test graph:: — the CPU attention round-trip tests (KV store → GQA attention against hand-computed references) and the decode/prefill KV-persistence tests live with the graph suite; attn_parallel_realdata_correctness checks multi-threaded attention against a real-dump reference.
  • cargo test vec_ops:: — per-op parity: every AVX2/NEON path vs the scalar reference you read above.
  • MINFER_TRACE=<path> ./target/release/minfer <model> "hi" — per-node traces (doc 08 §4); the attention nodes' output stats land in the viz pipeline view (viz/README.md).
  • MINFER_DUMP_DIR (debug-dump builds, --features debug_dump) — per-layer hidden-state dumps; comparing layer-by-layer against llama.cpp dumps is how the RoPE-style and attention-scale issues in docs/DEBUGGING-*.md were cornered.
  • The KV arithmetic (24 KB/token for 0.5B, 288 KB/token for Qwen3-4B — §2.2) is directly observable: run a long context and watch RSS grow by the KV region totals doc 07 computed.

5. Cross-references

  • 05 — Graph build §3 — where the RmsNorm/RoPE/Kvcache*/Attn nodes and their metas come from.
  • 07 — Allocator §3.2 — the persistent K/V regions and the aliasing this doc's in-place ops rely on.
  • 08 — Scheduler — the build-order guarantee that orders store before attention.
  • 10 — CPU matmul — the W_q/W_k/W_v/W_o projections around these ops; the shared thread pool; the Q8_K bsums exactness lesson.
  • 12 — Sampler — the next stage: what happens to the logits the last layer produces.
  • 13 — Decode loop — the cache in action: one position written per step, nkv growing by one.
  • 14/15 — the same attention on GPU (flash/split variants on Metal; split-KV on CUDA).
  • docs/COMPUTE-GRAPH-DESIGN.md §17 deviations 4/5 — the KV layout and pos-input decisions this doc's mechanics implement.

← 10 — CPU matmul: quantized weights × Q8_0 activations · Index · 12 — The sampler: from logits to a token →

12 · Sampler — from logits to one chosen token

Stage: last-token logits out of the forward (docs 09–11) → this stage → one chosen token id, which doc 13 feeds back into the graph for the next step. This is the bridge between math and language: the model's entire output for one step is a list of 151,936 floating-point scores; the sampler turns that list into a single integer — the next piece of text. Code: src/sampler.rs (apply_penalties :71, recent_window :108, match_stop_suffix :120, apply_top_k :141, apply_top_p :165, sample_temperature :242, sample_with_config :1181) and its call site, the decode loop in src/main.rs :1553–1791 (GenParams defaults :73–140, seeded RNG :1566, sampling call :1729–1744, stop gates :1749–1768, is_stop_token :1868) — lines verified at commit 15fa45c.

1. Background — where this stage sits

Doc 11 ended with attention producing each layer's hidden states, and doc 09 ended with the final projection: for the last token of the sequence, the engine computes one score per vocabulary entry. Those scores are called logits. A logit is a raw, unnormalized score — a plain f32 that says "how much the model likes this token as the next one". Higher means more likely, but the numbers are not probabilities: they can be negative, they do not sum to anything in particular, and their absolute scale is arbitrary (it depends on the model, the quantization, even the layer's final RMSNorm gain). For the Qwen family the vocabulary has 151,936 entries, so what this stage receives is a Vec<f32> with 151,936 elements — 151,936 × 4 bytes = 607,744 bytes ≈ 0.6 MB of scores (the 607 KB the code comment at src/main.rs:924-925 refers to).

The sampler's whole job is to turn that list into one u32 token id. That sounds trivially small next to the transformer's billions of multiplies, and in CPU time it is — microseconds against milliseconds. But it is where the model's character lives. The same weights produce a careful, repetitive assistant or a creative, rambling one depending entirely on how this stage picks. Every knob doc 01 collected — --temp, --top-k, --top-p, --repeat-penalty, --frequency-penalty, --presence-penalty, --seed, --greedy — is consumed here and nowhere else.

Three quiet facts shape the design:

  1. The model only ever proposes; the sampler disposes. The forward pass computes a score for every token in the vocabulary, including absurd ones. Sampling decides which of those voices gets heard.
  2. Sampling needs memory of the past. The repeat/frequency/presence penalties need to know which tokens appeared recently. That means the decode loop keeps a sliding window of the last 64 token ids — seeded from the prompt's tail — and passes it in alongside the logits (src/main.rs:842-847).
  3. "Stop generating" is a sampling-adjacent concern. After a token is chosen, three gates decide whether generation ends: the token is the end-of-text sentinel (eos or <|im_end|>, doc 04), the generated byte stream now ends with a user-supplied stop string, or the n_predict cap is reached. The first two live in the same loop, a few lines below the sampling call, and one of them (match_stop_suffix) is implemented in sampler.rs.

What would break without this stage? Literally everything downstream: the decode loop has nothing to append, the tokenizer has nothing to decode, and the terminal stays silent. But the subtler failure is a wrong sampler: the engine's claims of matching llama.cpp rest on producing the same output for the same parameters, and the sampler is half of that contract (the other half is bit-close kernels). That is why minfer copies llama.cpp's default values, its penalty semantics, and its stage order, and why the whole thing is seeded and reproducible by default.

2. Principle — how it works and why

2.1 The pipeline in one picture

The sampler is a short chain of filters followed by one random draw. Each stage takes the logits buffer, modifies it in place, and hands the same buffer to the next stage:

 raw logits   Vec<f32>, one entry per vocab token (151,936 for Qwen)
              ← the output of the last forward (docs 09–11)
     │
     ▼
 apply_penalties      repeat ×/÷ rule + frequency −count·f + presence −p
                      one pass over the last-64 token window
     │                (temp == 0? → greedy argmax here and skip the rest)
     ▼
 apply_top_k          keep the k=40 highest logits, mask the rest to −∞
     │
     ▼
 apply_top_p          nucleus: keep the smallest set of tokens whose softmax
     │                probability sums ≥ p (0.95); mask the rest to −∞
     ▼
 sample_temperature   logits × 1/temp (0.8) → softmax → multinomial draw
     │                ← one uniform random number from StdRng(seed = 42)
     ▼
 SampledToken { token_id: u32, logit: f32 }
     │
     ▼
 stop gates           is_stop_token(eos/<|im_end|>)?  match_stop_suffix()?
     │                generated.len() < n_predict?
     ▼
 decode_bytes → stdout  ·  doc 13 feeds token_id back into the graph

Everything before the last box is deterministic arithmetic on the logits. The single random element is one uniform number per generated token. Fix that number stream (the seed) and the entire run becomes reproducible.

2.2 From logits to a probability distribution: softmax

A probability distribution over the vocabulary is a list of numbers, one per token, that are all ≥ 0 and sum to exactly 1 — think of a pie cut into 151,936 slices, one slice per token, sized by "how likely is this token next". Raw logits are not that (they can be negative and don't sum to 1), so the engine converts them with softmax, the standard score→probability converter:

softmax(z)_i = exp(z_i) / Σ_j exp(z_j)

In words: raise e (≈ 2.718, a mathematical constant) to the power of each logit — this makes every score positive while preserving order (bigger logit ⇒ bigger exp) — then divide each by the total, so everything sums to

  1. The exponential's superpower is contrast amplification: a logit difference of 2 becomes a probability ratio of e² ≈ 7.4. Small score gaps turn into lopsided probabilities.

One implementation detail you will see in the code: before exponentiating, every logit has the maximum subtracted (exp(v - max)). This changes nothing mathematically — the max appears in numerator and denominator and cancels — but it prevents overflow. exp(30) is already ~10¹³ and logits can reach the tens; exp(large) in f32 overflows to infinity. After the shift the biggest exponent is exactly exp(0) = 1, which cannot overflow. You will see this "subtract the max" pattern in every softmax in this engine, including attention's (doc 11).

2.3 Why sample at all — and what "greedy" means

The softmax hands you a full probability distribution. The simplest policy is: always take the highest-probability token. That is called greedy decoding (or argmax sampling — "argmax" means "the index of the maximum"). It is fully deterministic and it is the right default for benchmarking and correctness gates — minfer's own bench uses it (src/bench.rs:24,384), and the CUDA optimization campaign's second gate is "greedy byte-for-byte identity" before/after a kernel change (docs/cuda_optimization_steps/77-verification-methodology.md §2.2).

So why not always greedy? Because text generated purely by "what is most likely next" has a failure mode every LLM user has seen: loops. "The cat sat on the mat. The cat sat on the mat." A greedy decoder, once it steers into a rut, has no way out — the most likely continuation of a repetition is more repetition. Multinomial sampling is the alternative: treat the probability distribution as a weighted lottery and draw one token at random, with each token's win chance equal to its probability. Now a 0.7-likely token wins ~70% of the time — and, crucially, a 0.2-likely token sometimes wins, which is exactly the escape hatch that breaks loops and gives the model its range of phrasing. The trade is coherence for diversity; temperature (§2.6) is the dial between the two, and the dial's zero stop is greedy itself — --greedy in the CLI is literally --temp 0 (src/main.rs:312-314), and sample_temperature returns the argmax when temp < 1e-6 (src/sampler.rs:230-232). Greedy is not a different algorithm here; it is a degenerate temperature.

minfer's default is llama.cpp's: temp = 0.8 — mildly random. The rest of the pipeline (penalties, top-k, top-p) exists to make that randomness tasteful: fair over plausible candidates, rigged against degenerate ones.

2.4 Stage 1 — penalties: teaching the model not to repeat itself

The first filter looks at the last 64 generated-or-prompted tokens (the window; llama.cpp calls the setting repeat_last_n and 64 is its default) and pushes down the logits of anything that appears in that window. minfer implements all three penalties in one pass (apply_penalties, src/sampler.rs:47-79), matching llama.cpp's llama_sampler_init_penalties semantics, which the doc comment spells out:

for each distinct token t in the window, with count(t) occurrences:
    logits[t] -= count(t) · frequency_penalty     (if frequency ≠ 0)
    logits[t] -= presence_penalty                 (once, if count(t) > 0)
    logits[t] = logits[t] ≤ 0 ? logits[t] · repeat_penalty
                               : logits[t] / repeat_penalty

Three separate ideas share the pass:

  • Repeat penalty (default 1.1): a token seen in the window has its positive logit divided by 1.1 — or its negative logit multiplied by 1.1. Both branches make the token strictly less likely; §3.3 explains why one parameter covering both signs forces the ÷/× split.
  • Frequency penalty (default 0, used by the OpenAI-compatible server): subtract count × 0.something per occurrence — the more a token already appeared, the harder it is pushed.
  • Presence penalty (default 0): subtract a flat amount once if the token appeared at all — count-blind, it just discourages re-raising any recent topic.

A worked example. Suppose the window's last 64 tokens contain "the" three times and "a" once, and the model's raw logits for five candidates are:

tokenraw logitin window?after repeat 1.1
the8.53×8.5 / 1.1 = 7.727
cat7.9no7.900 (untouched)
a4.01×4.0 / 1.1 = 3.636
dog−2.0no−2.0 (untouched)
said−0.5no−0.5 (untouched)

Before the penalty, the (8.5) beats cat (7.9) and greedy would emit the — possibly the third the in a row. After the penalty, cat wins (7.900 > 7.727) and the loop is broken. Note what did not happen: dog and said, which are not in the window, were not touched; the penalty never invents new preferences, it only taxes recency. (If the window had contained dog, its negative logit would become −2.0 × 1.1 = −2.2 — pushed further down, not up; §3.3 covers why.)

With the OpenAI-style penalties on (frequency = 0.5, presence = 0.3), the order inside one pass matters and the code applies subtraction first, then the repeat rule — llama.cpp's order. For the (count 3): 8.5 − 3 × 0.5 − 0.3 = 6.7, then ÷ 1.1 = 6.09. For a (count 1): 4.0 − 0 − 0.3 = 3.7, then ÷ 1.1 = 3.36.

The counting itself is a HashMap<u32, u32> built by one loop over the window, then one mutation per distinct token — at most 64 hash entries and 64 logit writes per generated token, i.e. microseconds against the milliseconds the forward pass costs.

2.5 Stage 2 and 3 — top-k and top-p: pruning the long tail

After penalties, the distribution still spans the whole vocabulary. Most of those 151,936 candidates are nonsense — misspellings, lone bytes from the middle of a CJK character, whitespace runs. Individually each is unlikely, but collectively the long tail is where sampling goes to produce garbage, and two standard filters chop it.

Top-k (default k = 40) keeps only the 40 highest logits and masks every other entry to −∞ (negative infinity — the float value that loses every comparison and maps to probability 0 in softmax). After this, the lottery has at most 40 tickets, and they are by construction the strongest ones. That is the entire point: kill the tail of absurd tokens in one stroke, regardless of how flat or peaked the distribution currently is.

Top-p (default p = 0.95), also called nucleus sampling, is adaptive where top-k is fixed. It computes the softmax probabilities, sorts them from most to least likely, and keeps the smallest set of top tokens whose probabilities sum to at least p — the "nucleus" of the distribution. On a confident step (one token at 0.98), the nucleus is that one token. On a genuinely uncertain step (ten tokens at ~0.1 each), the nucleus stretches to ten. The cutoff tracks the shape of the distribution instead of a fixed count, which is why it complements top-k rather than replacing it.

A worked example, small enough to check by hand. Five tokens with softmax probabilities 0.40, 0.30, 0.15, 0.10, 0.05, and p = 0.8:

ranktokenprobabilityrunning sumverdict
1A0.400.400.40 ≤ 0.8 → keep going
2B0.300.700.70 ≤ 0.8 → keep going
3C0.150.850.85 > 0.8 → stop; C is kept
4D0.10—masked to −∞
5E0.05—masked to −∞

The nucleus is {A, B, C}: the smallest prefix whose sum (0.85) reaches past 0.8. D and E — 15% of the probability mass between them — are now unreachable. The token that crosses the threshold is kept, so the nucleus is never smaller than the first token that gets you to p.

One subtlety of minfer's implementation matters enough that the code documents it (src/sampler.rs:142-145): top-p does not rewrite the kept entries with their probabilities. It computes probabilities only internally, to decide who stays; then it masks the losers' raw logits to −∞ and leaves the winners' raw logits untouched. Why that distinction is load-bearing is §3.3's third design question — short version: the final softmax has not happened yet, and it still needs the real logits.

2.6 Stage 4 — temperature, the final softmax, and the draw

Everything so far selected candidates. The last stage weights them and picks one. Three sub-steps, all in sample_temperature (src/sampler.rs:229-278):

  1. Divide every surviving logit by temp. Temperature is a dial on the softmax's contrast. Dividing by a small number stretches the logit gaps, making the distribution sharper; dividing by a large number compresses them, flattening it. Watch the same three logits (4.0, 3.0, 2.0) go through softmax at different temperatures:

    tempscaled logitsprobabilitiescharacter
    0.58.0, 6.0, 4.00.867, 0.117, 0.016sharp — best token nearly always wins
    0.85.0, 3.75, 2.50.731, 0.209, 0.060default — mildly random
    1.04.0, 3.0, 2.00.665, 0.245, 0.090the raw distribution
    2.02.0, 1.5, 1.00.507, 0.307, 0.186flat — long shots get real odds

    As temp → 0 the probabilities collapse onto the argmax (the 1/t multiplier grows without bound, so the top logit's exp wins by an infinite margin): temperature 0 is greedy. As temp → ∞ the distribution approaches uniform. Everything the earlier stages decided survives this step, because dividing by a positive constant never changes which logits are bigger — only how much bigger.

  2. Softmax. The real one this time — subtract-max, exp, normalize — exactly §2.2. Two implementation details: the running sum accumulates in f64 (152k exp values summed in f32 would lose low-order bits), and entries already masked to −∞ are skipped rather than exponentiated — exp(−∞) is exactly +0.0, so skipping is bit-identical and saves ~150k transcendental calls per token when top-k/top-p have pruned hard (src/sampler.rs:239-242).

  3. The multinomial draw. Multinomial sampling means: pick one token with probability equal to its weight. Picture the unit interval [0, 1) carved into segments whose widths are the probabilities, then drop a uniform random dart into [0, 1) and see which segment it lands in:

    probs:   B=0.731          C=0.209      D=0.060
    [0 ────────────────┬──────────────┬────────┬ 1)
               dart r = 0.85 ────────────────┘ lands in D's segment
    

    The code walks the tokens in index order accumulating a running total and returns the first token whose cumulative sum reaches the dart (src/sampler.rs:260-273). The dart comes from rng.gen() — one uniform f32 in [0, 1) per generated token — from a seeded random number generator: StdRng::seed_from_u64(params.seed), default seed 42 (src/main.rs:845). A seed is the starting state of a pseudo-random generator: same seed in, same sequence of "random" numbers out — which is why the same command produces the same text even while sampling (§3.3's fourth design question covers why minfer wants that).

2.7 Stopping: when does generation end?

The sampler produces a token; the loop decides whether to keep going. Three gates, checked in this order each step (src/main.rs:883-918):

  1. Stop tokens. is_stop_token(id, &special) — true when the sampled id equals the model's eos (end of sequence) token or <|im_end|> (doc 04: both come from GGUF metadata via ModelDef::special_tokens()). When the model emits its own "I'm done" marker, the loop breaks and the marker is not appended or printed. This is the normal ending of an answer.
  2. Stop strings. The user can pass --stop "some text" (repeatable). Each step appends the new token's decoded bytes to a buffer full, and match_stop_suffix(&full, &stop_refs) asks: does the accumulated stream now end with any stop string? Matching bytes — and matching the whole stream, not just the newest token — is what makes a stop string split across token boundaries work: "world" might arrive as wo + rld, and neither token is the string, but after the second one the byte suffix matches (doc 04 explains why the engine deals in raw bytes at all). The code excerpt in §3.2 shows the truncation and the already-streamed-bytes behavior on a match.
  3. n_predict. The hard cap from doc 01 (-n, default 512) — the while generated.len() < params.n_predict loop condition.

The match_stop_suffix choice of "suffix" is deliberate economy: a stop string can only complete on the step that emits its last byte, and at that moment it ends the buffer — so checking only the buffer's tail each step finds every possible match exactly once, in O(len(stop)) time.

3. Implementation

3.1 Data in / data out

In (all in main's frame, handed to sampler::sample_with_config):

ValueType / shapeFrom
logits&mut Vec<f32>, exactly n_vocab entries (151,936 for Qwen ≈ 0.6 MB)the last forward (docs 09–11); prefill for the first step, one-token decode for every later step
params.tempf32, default 0.8 (0 = greedy)GenParams (src/main.rs:73)
params.top_kusize, default 40:68
params.top_pf32, default 0.95:69
params.repeat_penaltyf32, default 1.1:70
params.frequency_penalty / presence_penaltyf32, defaults 0.0:71-72
prev_tokens&[u32], ≤ 64 entriessliding window: prompt tail (recent_window(&input_ids, 64), :847) + every generated token so far (:903-907)
rng&mut StdRng, seeded once at :845seed_from_u64(params.seed), default 42

Out: a SampledToken { token_id: u32, logit: f32 } (src/sampler.rs:7-13). Only token_id drives the loop; logit is result metadata (a sampled probability after the final softmax, actually — the field name is historical).

The buffer contract is the quiet star of the API: the sampler mutates the caller's logits buffer in place and returns only the tiny result struct — no full-vocab copy crosses the sampler boundary, and no second buffer needs to stay alive per token. Why that shape matters is design question 5 in §3.3.

After the call (still inside one loop iteration, src/main.rs:900-922): stop-token gate → generated.push(id) + window update → decode to bytes into full → stop-string gate → stream newly finished bytes → next forward.

3.2 Key code

The whole pipeline in one function

src/sampler.rs:280-307 — the complete sampler is eleven lines of calls; the order is the design:

#![allow(unused)]
fn main() {
// src/sampler.rs:283-307
pub fn sample_with_penalties<R: Rng>(
    logits: &mut [f32],
    temp: f32,
    top_k: usize,
    top_p: f32,
    repeat_penalty: f32,
    frequency_penalty: f32,
    presence_penalty: f32,
    prev_tokens: &[u32],
    rng: &mut R,
) -> SampledToken {
    apply_penalties(
        logits, prev_tokens, repeat_penalty, frequency_penalty, presence_penalty,
    );
    if temp < 1e-6 {
        return sample_greedy(logits);
    }
    apply_top_k(logits, top_k);
    apply_top_p(logits, top_p);
    sample_temperature(logits, temp, rng)
}
}

Read the greedy early-return carefully: in greedy mode the penalties still apply (a repeated token can lose the argmax to a non-repeated one — that is the whole point of the default repeat penalty, and the unit test test_sample_pipeline_greedy_applies_penalty pins it), but top-k and top-p are skipped — correctly, because masking logits to −∞ can never remove the maximum, so those stages cannot change greedy's choice. Fewer operations, identical result.

Penalties: one pass, one HashMap, llama.cpp semantics

src/sampler.rs:47-79 (doc comment :31-46 spells out the semantics). First the counting — note the early-out when every penalty is at its "off" value:

#![allow(unused)]
fn main() {
// src/sampler.rs:54-60
    if (repeat - 1.0).abs() < 1e-6 && frequency.abs() < 1e-6 && presence.abs() < 1e-6 {
        return;
    }
    let mut counts: HashMap<u32, u32> = HashMap::new();
    for &t in prev_tokens {
        *counts.entry(t).or_insert(0) += 1;
    }
}

The HashMap gives two things at once: distinct tokens (iterate the map, not the window — a token repeated 10 times is taxed once per its count, not 10 times) and the counts for frequency penalty. Then the combined mutation per distinct token:

#![allow(unused)]
fn main() {
// src/sampler.rs:61-78 (loop body)
    for (&t, &c) in &counts {
        let idx = t as usize;
        if idx >= logits.len() {
            continue;                       // out-of-range window token: skip, never panic
        }
        let v = logits[idx];
        let mut nv = v;
        if frequency.abs() >= 1e-6 || presence.abs() >= 1e-6 {
            nv -= c as f32 * frequency;     // −count·f  (scales with occurrences)
            if c > 0 {
                nv -= presence;             // −p        (once, count-blind)
            }
        }
        if (repeat - 1.0).abs() >= 1e-6 && repeat >= 1.0 {
            nv = if nv <= 0.0 { nv * repeat } else { nv / repeat };
        }                                   //           (the asymmetric rule)
        logits[idx] = nv;
    }
}

Three guards worth noticing: the idx >= logits.len() skip (a corrupt or out-of-model window token must not panic the stream — the bench path feeds arbitrary seed tokens); repeat >= 1.0 (a repeat penalty below 1.0 would reward repeats — minfer ignores it rather than implementing the opposite of the feature's name); and the subtraction-before-division order, which the test test_freq_presence_then_repeat_penalty pins with the comment "llama.cpp order".

The window, and how the loop keeps it sliding

src/sampler.rs:92-98 is trivial — the interesting part is the call-site choreography in main.rs:

#![allow(unused)]
fn main() {
// src/sampler.rs:95-98
pub fn recent_window(tokens: &[u32], last_n: usize) -> Vec<u32> {
    let start = tokens.len().saturating_sub(last_n);
    tokens[start..].to_vec()
}

// src/main.rs:842-847 — seeded once, window seeded from the PROMPT tail
    let mut rng = rand::rngs::StdRng::seed_from_u64(params.seed);
    const REPEAT_LAST_N: usize = 64;
    let mut prev_tokens = sampler::recent_window(&input_ids, REPEAT_LAST_N);

// src/main.rs:903-907 — every sampled token enters the window
        generated.push(sampled.token_id);
        prev_tokens.push(sampled.token_id);
        if prev_tokens.len() > REPEAT_LAST_N {
            prev_tokens.drain(0..prev_tokens.len() - REPEAT_LAST_N);
        }
}

The comment at :842-844 carries the design point: the window is initialized from the prompt's last 64 tokens "so the first generated tokens are penalized too". Without that, step 1 of the generation would have an empty window and a prompt that says "translate: the the the" could immediately sample the. The window is a fixed-capacity sliding window: push, then drain from the front to 64 — the prompt tail ages out as generation proceeds, and by token 65 of output the penalties look only at generated text.

Top-k: O(n) threshold, mask in place

src/sampler.rs:128-140:

#![allow(unused)]
fn main() {
// src/sampler.rs:128-140
pub fn apply_top_k(logits: &mut [f32], k: usize) {
    if k == 0 || k >= logits.len() {
        return;
    }
    let mut sorted = logits.to_vec();
    sorted.select_nth_unstable_by(k - 1, |a, b| b.total_cmp(a));
    let threshold = sorted[k - 1];
    for v in logits.iter_mut() {
        if *v < threshold {
            *v = f32::NEG_INFINITY;
        }
    }
}
}

The trick is select_nth_unstable_by: Rust's partial-selection primitive that partitions a slice so the element that would be k-th in sorted order lands at index k−1, in O(n) average time instead of the O(n log n) of a full sort (for n = 151,936: roughly 300k comparisons vs ~2.6M). It runs on a copy — select_nth_unstable_by reorders whatever slice it is given, which would destroy the index↔token mapping in the caller's buffer; the original is only read to extract the threshold and masked, never moved (the doc comment :121-127 records exactly this). After the mask, ties are kept: *v < threshold is strict, so tokens exactly at the threshold survive — top-k keeps at least k candidates.

The k == 0 guard doubles as the off switch: --top-k 0 disables filtering entirely and hands the full vocabulary to top-p (which then takes its full-array path, below).

Top-p: nucleus over survivors, mask raw logits

src/sampler.rs:152-201, the function whose doc comment (:142-151) is the design record. First the fast path — after top-k, at most k entries are finite:

#![allow(unused)]
fn main() {
// src/sampler.rs:156-168
    let survivors: Vec<(usize, f32)> = logits
        .iter()
        .enumerate()
        .filter(|(_, &v)| v > f32::NEG_INFINITY)
        .map(|(i, &v)| (i, v))
        .collect();
    if survivors.is_empty() {
        return;
    }
    if survivors.len() > 1024 {
        apply_top_p_full(logits, p);
        return;
    }
}

Then the nucleus decision, computed on the survivors alone:

#![allow(unused)]
fn main() {
// src/sampler.rs:173-200 (core)
    let max_val = survivors.iter().fold(f32::NEG_INFINITY, |a, &(_, v)| a.max(v));
    let sum: f64 = survivors
        .iter()
        .map(|&(_, v)| ((v - max_val) as f64).exp())
        .sum();
    let mut cand: Vec<(usize, f32)> = survivors
        .iter()
        .map(|&(i, v)| (i, (v - max_val).exp() / sum as f32))
        .collect();

    cand.sort_by(|a, b| b.1.partial_cmp(&a.1).unwrap_or(std::cmp::Ordering::Equal));

    let mut cumulative = 0.0f32;
    let mut keep = cand.len();
    for (i, &(_, prob)) in cand.iter().enumerate() {
        cumulative += prob;
        if cumulative > p {
            keep = i + 1;               // the token that crosses the line is kept
            break;
        }
    }
    for &(idx, _) in &cand[keep..] {
        logits[idx] = f32::NEG_INFINITY;  // mask the LOGIT, keep the winners' logits raw
    }
}

Why is summing only the survivors legal? Because the excluded entries are already −∞, and exp(−∞) = +0.0 exactly in IEEE floats — adding exact zeros to an f64 running sum changes nothing, so the survivor-only sum is bit-identical to the full-array sum. That is the claim in the comment at :170-172, and it is what makes the ≤1024-candidate fast path a pure optimization rather than an approximation. When top-k is disabled (or the distribution is pathologically flat), more than 1024 entries survive and the code falls back to apply_top_p_full (:206-226) — the original softmax-over-everything + full sort, kept so --top-k 0 never silently changes semantics.

The final loop is the part §2.5 promised: losers get −∞ in logit space; the winners' entries stay untouched raw logits for the temperature stage.

Temperature + softmax + the draw

src/sampler.rs:229-278, all three sub-steps in one function:

#![allow(unused)]
fn main() {
// src/sampler.rs:229-256 (core)
pub fn sample_temperature<R: Rng>(logits: &mut [f32], temp: f32, rng: &mut R) -> SampledToken {
    if temp < 1e-6 {
        return sample_greedy(logits);          // greedy IS temp 0
    }

    let inv_temp = 1.0 / temp;
    for v in logits.iter_mut() {
        *v *= inv_temp;                        // (−∞)·finite stays −∞
    }

    // Softmax. Masked (-INF) logits map to exp(-INF)=0 and contribute nothing
    // to the running max or sum; skipping the exp() call for them avoids
    // ~n_vocab transcendental evaluations per token while staying bit-identical
    // (exp(-INF) == +0.0 exactly).
    let max_val = logits.iter().fold(f32::NEG_INFINITY, |a, &b| a.max(b));
    let mut sum = 0.0f64;
    for v in logits.iter_mut() {
        if *v > f32::NEG_INFINITY {
            *v = (*v - max_val).exp();
        } else {
            *v = 0.0;
        }
        sum += *v as f64;
    }
}

Then normalization and the draw (:253-277): divide by the sum in place, draw one uniform f32 via rng.gen(), and walk the buffer accumulating:

#![allow(unused)]
fn main() {
// src/sampler.rs:260-277
    let r: f32 = rng.gen();
    let mut cumulative = 0.0f32;
    for (i, &v) in logits.iter().enumerate() {
        if v <= 0.0 {
            continue;                          // pruned entries: zero probability
        }
        cumulative += v;
        if r <= cumulative {
            return SampledToken { token_id: i as u32, logit: v };
        }
    }
    SampledToken { token_id: (logits.len() - 1) as u32, logit: logits[logits.len() - 1] }
}

The v <= 0.0 skip is why the walk is cheap after pruning (≤ 40 real entries after top-k) — and the trailing fallback return is the float-safety net: if rounding leaves the cumulative total at 0.9999… while the dart drew 0.99995, the last finite entry wins instead of nobody. Note the buffer is destroyed by this function (it now holds probabilities, not logits) — which is fine, because the caller never reads it again; the next forward returns a fresh Vec (§3.1).

Stop tokens and stop strings at the call site

The gates run immediately after sampling, before anything is appended:

#![allow(unused)]
fn main() {
// src/main.rs:883-918 (condensed to the sampler-adjacent lines)
    while generated.len() < params.n_predict {
        let sampled = sampler::sample_with_penalties(
            &mut logits,
            params.temp, params.top_k, params.top_p,
            params.repeat_penalty, params.frequency_penalty, params.presence_penalty,
            &prev_tokens,
            &mut rng,
        );

        if is_stop_token(sampled.token_id, &special) {
            break;                                  // eos / <|im_end|>: silent stop
        }
        generated.push(sampled.token_id);
        prev_tokens.push(sampled.token_id);
        /* … window drain … */

        // Stop-string detection on the FULL byte stream before emitting.
        full.extend_from_slice(&tokenizer.decode_bytes(&[sampled.token_id]));
        if let Some(cut) = sampler::match_stop_suffix(&full, &stop_refs) {
            full.truncate(cut);                     // drop the stop string itself
            if cut > emitted {
                hi.feed(&full[emitted..]);          // flush bytes not yet streamed
                emitted = full.len();
            }
            break;
        }
        if emitted < full.len() {
            hi.feed(&full[emitted..]);              // stream newly finished bytes
            emitted = full.len();
        }
}

is_stop_token is the two-line sentinel (src/main.rs:1020-1022):

#![allow(unused)]
fn main() {
fn is_stop_token(id: u32, special: &models::SpecialTokens) -> bool {
    id == special.eos || Some(id) == special.im_end
}
}

and match_stop_suffix is the byte-level tail check (src/sampler.rs:107-119):

#![allow(unused)]
fn main() {
// src/sampler.rs:107-118
pub fn match_stop_suffix(buf: &[u8], stops: &[&[u8]]) -> Option<usize> {
    let mut best: Option<usize> = None;
    for s in stops {
        if s.is_empty() || s.len() > buf.len() {
            continue;                               // empty stops ignored; too-long can't match
        }
        if &buf[buf.len() - s.len()..] == *s {
            let start = buf.len() - s.len();
            best = Some(best.map_or(start, |b| b.min(start)));
        }                                           // several matches: earliest start wins
    }
    best
}
}

"Earliest start wins" means the longest applicable stop string decides where to cut (a later start truncates less text). The bytes already flushed to the terminal before the match completed are not un-printed — that is the llama.cpp antiprompt behavior the comment at src/main.rs:849-853 describes, and it is why full (everything generated) and emitted (everything streamed) are tracked as separate cursors.

3.3 Design choices (why this shape and not another)

1. Why sample at all instead of always taking the max? §2.3 explained the failure mode: greedy steps are locally optimal, but locally optimal steps compound into globally degenerate sequences — the model repeating "the the the" does nothing wrong per step, it has simply fallen into a self-reinforcing basin where the most likely token after a repetition is the same repetition. Multinomial sampling injects exactly enough noise to escape those basins while staying biased toward good tokens (a 0.73-likely token still wins 73% of the time); temperature is the coherence-vs-diversity dial and greedy its zero point (--greedy, src/main.rs:312-314). The composition minfer chose matters as much as the mechanism: penalties apply in both modes (a repeated token can lose the argmax), so even greedy output is loop-resistant — the default 1.1 penalty does real work in every verification run in this repo.

2. Why penalize repeats in logit space, with the asymmetric ÷/× rule — and why a 64-token window? Work backwards from what the downstream consumer needs. Softmax only cares about logit differences; "make token t less likely" always means "move t's logit down relative to everyone else". The penalties therefore edit logits directly, where one small, composable step does the whole job — no re-normalization, no separate probability pass. The multiplicative form (÷1.1 / ×1.1) rather than subtracting a constant is a scale choice: logit magnitudes vary across models, quants, and positions, and a ratio taxes a confident repeat (logit 12) harder than a hesitant one (logit 2) in exactly the proportion that matters — while one fixed subtraction would be negligible for the first and absurd for the second. The asymmetry is then forced by the sign: dividing a negative logit by 1.1 would move it toward zero, i.e. make an unlikely repeated token more likely — the opposite of the feature. Multiplying instead pushes it further from the pack. Either way the rule means "this token got less likely", which is the only contract softmax cares about. (An additive design could also be made sign-aware, but it would need a scale-calibrated constant per model; the ratio needs none, which is why llama.cpp ships one default that works everywhere.)

The window is 64 for a bent-cost curve: a degenerate loop repeats within a handful of tokens, and phrase-level echoes ("as I said above") live within tens — 64 covers both, so anything the model is stuck on gets taxed every step. Meanwhile the topical words a document genuinely needs — names, terms of art — recur over spans of hundreds of tokens; a window of 64 has already forgotten them by the time they are legitimately needed again. Longer windows start punishing correctness (a code generator needs its brackets and whitespace back after 70 tokens), shorter ones let loops survive past the horizon. 64 is llama.cpp's tuned compromise, and minfer copies it instead of re-tuning (doc 01 §2.3's "less invention risk").

3. Why is top-p implemented as logit masking rather than truncating the probability vector? The comment at src/sampler.rs:142-145 states the rule — "sets excluded tokens' raw logits to −∞ (does NOT overwrite logits with probabilities, so the final temperature softmax stays correct)" — and the reason is composition. The pipeline applies top-p before temperature, and temperature divides logits by 1/temp and re-softmaxes. If top-p had overwritten the survivors' entries with their temp-1 probabilities, the temperature stage would then scale probabilities (multiply by 1/temp and exponentiate them), producing a different distribution than "nucleus-filter, then temperature" — e.g. flattening p = 0.9 → p⁰·⁸-style distortions instead of a clean renormalized subset. Keeping raw logits makes the two stages commute cleanly: filtering is a set operation (remove, don't rewrite), temperature is a shape operation on whatever set remains. It also matches llama.cpp, whose sampler chain passes logits through link by link with each link mutating in place — top-p sees untempered logits, temp runs after — so parity of semantics, not just of defaults, is preserved. And there is a numerical dividend: masking −∞ survivors is exact (exp(−∞) = +0.0), so the fast path that softmaxes only the ≤1024 (in practice ≤40) survivors is bit-identical to the full-vocab computation (:147-152, :170-172) — no epsilon drift between code paths.

4. Why a fixed seed (42) by default instead of entropy? Because in this repo, determinism is infrastructure. Doc 01 §2.3 records the decision; the consumers are everywhere: the conversation tests (tests/conversation_cli.rs) pipe scripted stdin and assert on the output — impossible if every run rolls new dice; the CUDA optimization campaign's verification gate runs -n 32 --greedy --seed 42 before and after every kernel change and requires byte-identical token streams (docs/cuda_optimization_steps/77-verification-methodology.md §2.2) — and when the greedy stream is the gate, you want the seeded sampling path audited by the same machinery (the campaign's gate chain explicitly pairs greedy identity with the rp=1.0 seeded stream, §2.3/§2.4 there); and a support report that says "same model, same flags, same output" turns a heisenbug into a reproduction recipe. Entropy by default would buy lottery variety nobody asked for and cost every one of those properties. One seed value is one flag away (--seed N), so users who want fresh rolls per run can have them — opt-in, not default. Note the subtlety that makes seeded sampling a usable gate: the single StdRng is constructed once before the loop (:845), so the whole generation consumes one deterministic stream of draws; identical model + flags + seed ⇒ identical tokens even at temp 0.8, because minfer's forwards are themselves deterministic. (One honest limit of the gate: GPU kernels must also be deterministic for the equality to hold — which is exactly why the campaign pairs greedy identity with the parity tests rather than trusting either alone.)

5. Why does the sampler mutate the logits buffer in place? Because the obvious alternative — let probs = sampler::sample(logits.clone(), …) or returning a new Vec — buys a full-vocab copy per token for zero benefit. The arithmetic: 151,936 × 4 B = 607,744 B ≈ 0.6 MB per generated token, which is precisely the copy the call-site comment at src/main.rs:924-925 brags about not making ("move the Vec in place instead of copying 607 KB/token") for the forward boundary — the sampler boundary gets the same treatment. The ownership dance that makes it work: logits starts as the prefill output; each iteration passes &mut logits into the sampler (which consumes and destroys it — after sample_temperature it holds probabilities); then logits = model.forward(&[sampled.token_id], …) (:932) rebinds the variable to the fresh Vec the forward allocated. One buffer alive at a time, no .clone(), no caller-visible aliasing (the sampler takes &mut [f32], so the borrow checker guarantees nobody reads the half-transformed buffer mid-pipeline). The one internal copy that remains — apply_top_k's to_vec() for the selection pass — is a deliberate, short-lived temp: select_nth_unstable_by reorders what it sorts, and reordering the caller's buffer would scramble the index→token mapping (:121-127). Copy-then-select keeps the mutation mask-only at the cost of one temp that dies at the end of the function.

6. Why this stage order — penalties → top-k → top-p → temperature? Because it is llama.cpp's order, and each neighbor-pair has a reason. Penalties first, so a taxed token can also be pruned by k/p (and so greedy benefits). Top-k before top-p because it is the cheap O(n) pre-filter that guarantees top-p's softmax only ever handles ≤40 survivors — the reverse order would compute a full-vocab softmax for nothing. Temperature last, because it must shape the final set (§3 above), and because applying it earlier would change the probabilities top-p uses to draw the nucleus. The greedy shortcut (:301-303) sits where it does because penalties must run even when the stochastic stages are skipped.

3.4 Pitfalls & invariants

  • −∞ is the "removed" sentinel, end to end. Every pruning stage masks with f32::NEG_INFINITY, and every downstream consumer is written against that contract: softmax skips −∞ (exp(−∞) = +0.0), top-p's survivor scan filters v > f32::NEG_INFINITY, the draw skips v <= 0.0. A new stage that writes 0.0 instead of −∞ would corrupt top-p's survivor count; one that writes any finite sentinel would sneak into the softmax.
  • Top-p's survivor fast path is exact, not approximate — but only because the excluded entries contribute exact zeros. The >1024 fallback (apply_top_p_full) exists so --top-k 0 (or a bizarrely flat distribution) never falls off the bit-exactness claim; it is the full-vocab softmax + full sort, deliberately preserved (:203-205).
  • The penalty pass silently ignores a repeat penalty < 1.0 and out-of-range window tokens (:62-65, :74). Neither is an error path: a sub-1.0 penalty would encourage repetition, and a bad id must not panic a generation stream.
  • The order inside apply_penalties is load-bearing: frequency/ presence subtraction then the repeat ÷/× — the reverse order gives different numbers (and the test test_freq_presence_then_repeat_penalty pins llama.cpp's order with a worked case: 4 − 1 − 1 = 2, then ÷2 = 1).
  • The window spans prompt and generation (:842-847 + :903-907). A new code path that resets prev_tokens at generation start re-opens the "first generated token repeats the prompt" hole.
  • Greedy still pays the penalty pass (:301-303): any change that moves the greedy return before apply_penalties changes the output of every --greedy run with a nonzero repeat penalty — and every verification gate in this repo that uses -n 32 --greedy --seed 42.
  • Stop strings are bytes, not strings (match_stop_suffix(&[u8]…)): matching happens on the accumulated byte stream because tokens can split multi-byte characters (doc 04 §2.4). Converting to String mid-stream would make split-CJK stop strings unmatchable and lossy.
  • Already-emitted bytes stay emitted (:911-918): the stop-string truncation rewinds full, never the terminal. Anything else would be a lie about what the user saw.
  • The draw's fallback return is unreachable in exact math but reachable in floats: cumulative sums of probabilities can land at 0.99999…, so a dart in the gap returns the last finite entry rather than nothing (:274-277). Removing that arm invites a rare-but-real "no token" panic.
  • The sampler destroys its input buffer (it holds probabilities after sample_temperature). The caller-side invariant that makes this safe: logits is reassigned from the next forward before it is ever read again (src/main.rs:932). Reading logits between sample and reassign would read probabilities and call them logits.

4. Observe & verify

  • MINFER_TIMING=1 — decomposes per-token wall time into t_samp (the sampling call, :896-898) vs the forward. Expect sampling in the microseconds: the whole pipeline is a 64-entry hash pass, one O(n) selection, a ≤40-entry softmax, and one draw.
  • MINFER_TRACE=<path> — the decode loop attaches each sampled token's id and decoded text to the trace (crate::trace::set_token, src/main.rs:926-930); the viz page shows the token stream the sampler produced.
  • The reproducibility experiment — run the same command twice, e.g. ./target/release/minfer qwen2.5-0.5b-instruct-q4_0 "Tell me a story" -n 24: both runs print identical text (seed 42); add --seed 7 to either and the text differs while the prompt handling stays identical. That pair of observations is the seeded-sampler contract, live.
  • Greedy mode — --greedy (or --temp 0) runs the penalty pass + argmax; bench subcommand output is produced this way (src/bench.rs:24) so throughput numbers measure the engine, not the lottery.
  • Stop strings — --stop with a string that straddles tokens, e.g. --stop "the mat" against a prompt that will produce it; the output ends before the stop text, and with a multi-byte stop string you can watch the byte-level match fire only when the final byte arrives.
  • Unit tests — cargo test sampler:: covers every stage with small hand-checkable numbers: greedy picks max; positive logit ÷ penalty and negative × penalty; frequency scales with count while presence is once-only; freq+presence precede repeat; defaults are a no-op; top-k masks below threshold; top-p keeps only the nucleus and does not overwrite the surviving raw logit; two identical seeds give identical tokens; stop-suffix basics, longest-wins, and the CJK E4 B8 AD-split-across-tokens case.

5. Cross-references

← 11 — Attention, vec ops, and the KV cache · Index · 13 — The decode loop and graph reuse →

13 · The decode loop and graph reuse

Stage: the sampler picks the first token (doc 12) → this stage: the autoregressive decode loop — one token per forward, with the compute graph rebuilt once and then reused for every step → the same loop runs unchanged on the GPU backends (docs 14–15, which swap out the kernels inside the loop, never the loop itself). Code: src/main.rs (decode loop (main.rs:1727-1791, the while generated.len() < params.n_predict loop), the repeat window REPEAT_LAST_N (conversation.rs:426), loop setup main.rs:1553-1566 (the decode-loop setup), is_stop_token main.rs:1868-1870), the cached forward src/models/qwen2/graph.rs::forward_cached (src/models/qwen2/graph.rs:456-472), the reuse decision src/graph/cache.rs::try_reuse try_reuse (cache.rs:69-85), the params src/graph/params.rs (GraphParams params.rs:88-97), the allocator's persistent KV regions src/graph/alloc.rs (alloc_graph alloc.rs:513, ensure_kv alloc.rs:995), the per-step execution BackendScheduler::execute (scheduler.rs:137-473), and the multi-turn session src/conversation.rs (user_turn conversation.rs:757-764, generate_assistant_with_logits conversation.rs:1271-1278) — lines verified at commit 15fa45c.

1. Background — where this stage sits

By the end of doc 12, one full forward pass has happened. The prompt was rendered, tokenized, and pushed through the compute graph in a single prefill step (doc 09): every layer's math ran over all nt prompt tokens at once, each layer's K and V projections were written into the graph allocator's persistent KV regions (doc 07), and the graph returned the logits — one score per vocabulary entry — for the last prompt token only. The sampler then turned those 151,936 scores into exactly one token id. What the process now holds is a curious half-finished sentence: the prompt tokens, which the model has read, and one new token, which the model has written but not yet read.

This document is about the loop that finishes the sentence. Language models generate text autoregressively — auto ("self") + regressive ("moving backward"): each new token is produced by feeding the model everything so far, including the tokens the model itself just wrote. So the engine enters a loop:

  1. take the token sampled last step (for the first step, the last prompt token's logits were already produced by prefill),
  2. run one forward pass whose only input is that single token, at the next free slot in the KV cache,
  3. sample one token from the resulting logits,
  4. append it, print it, and check whether generation should stop.

Step 2 is a full transformer forward — 24 layers, every weight matrix — but for one token instead of nt. That single-token forward is called a decode step, and the loop that repeats it is the decode loop. In src/main.rs it is a plain Rust while loop (main.rs:1727-1791, the decode loop); everything dramatic about LLM inference — chatbots streaming words one at a time — is this loop running hundreds of times per second.

Two facts make this loop interesting enough for its own document.

Fact one: a decode step is cheap in a very specific way. Every matmul still reads its entire weight matrix (there is only one token's activation to multiply against each row, so nothing about the weights can be skipped) — that makes decode memory-bound: its speed is set by how fast RAM streams weights to the compute units, not by how fast they multiply (doc 10 §2.1 walks the arithmetic). But everything else shrinks: the matmul arithmetic drops from nt token-rows to 1, and attention (doc 11) drops from the quadratic all-pairs work of prefill to a single query scanning the KV window built so far. A decode step costs roughly "one prefill token's worth of matmul work plus one attention scan of nkv cached keys" — no more.

Fact two: the graph is rebuilt exactly once, then reused for free. Recall from docs 05–08 what one forward costs besides the math: build a few hundred CNode structs describing the topology, walk them to assign each a backend, run the fusion pass, and run the liveness allocator to hand every node a buffer. Doing all of that per decode step would be pure overhead — the topology does not change from step to step. minfer avoids it with a design rule that shapes the whole codebase: the graph topology is a deterministic function of GraphParams, and n_past (how much KV is cached) is not one of those params — it is input data. Equal params ⇒ identical graph ⇒ the cached graph, its backend assignment, its buffers, and above all its KV regions can be reused as-is; the only thing that changes per step is a handful of bytes written into the input nodes. The reuse decision itself is a six-field struct comparison in GraphCache::try_reuse (cache.rs:69-85) — no graph traversal, no node-by-node diff.

That rule is why the KV regions must live outside the graph, in an allocator owned by the cache (doc 07's contract). The prefill graph (built for nt tokens) and the decode graph (built for 1 token) are different graphs — different buffer shapes, a different gtype, decode-only fused ops on the GPU — so switching from prefill to decode does force one rebuild. The rebuild throws away the node list and the liveness mapping, but the allocator that owns the KV regions survives, and the regions are addressed by position, not by step. All the prompt's context is still there, and the very first decode step reads it.

The loop also decides when to stop. Three gates live in (or beside) the loop: the sampled token is the end-of-text sentinel is_stop_token (main.rs:1749-1751), the generated byte stream now ends with a user-supplied stop string match_stop_suffix (main.rs:1761-1768), or the -n token cap is reached (the while condition itself, the loop head (main.rs:1727). Doc 12 introduced these; here we see them as the loop's control flow.

Finally, there is a longer-timescale version of the same trick: a multi-turn chat session (--cnv, src/conversation.rs) keeps one token stream in the KV across whole conversations. Each user turn prefills only the delta — the few tokens of the new message — at the next free position, and the decode loop resumes on the same cached graph. A ten-turn conversation never re-prefills the previous nine turns; it pays O(delta) per turn instead of O(full history). Same invariant, larger loop: positions are data, the KV is append-only, and the graph just keeps getting reused.

What would break without this stage? Nothing downstream exists: no second token, no streaming, no chat. And without graph reuse, every one of the hundreds of decode steps would pay the build → assign → fuse → allocate tax again — the plan document estimated the wasted work as "recomputes topology every step + CPU scratch reallocation" and made parameter-deterministic reuse a headline goal of the rewrite (docs/COMPUTE-GRAPH-DESIGN.md §1, "Graph reuse" row). The rest of this document walks the loop one excerpt at a time: the per-step data flow (§2, §3.1), the code of the loop and the reuse machinery (§3.2), why it is built this way (§3.3), and the traps the design has to defuse (§3.4).

2. Principle — how it works and why

2.1 The loop in one picture

Here is one run of minfer model.gguf "Hello" with -n 4, annotated with which graph the engine used at each moment:

 prompt "Hello" ──tokenize──▶ [ids: nt tokens]

 ┌────────────────────────────────────────────────────────────────────┐
 │ PREFILL (one call, doc 09)                                         │
 │   graph: built for n_tokens = nt        (GraphType::Prefill)       │
 │   inputs: token_ids = ids, positions = [0, 1, .., nt-1]            │
 │   KV regions after: nt slots filled per layer                      │
 │   output: logits for the LAST token          ← 607 KB of f32      │
 └────────────────────────────────────────────────────────────────────┘
                    │ sample → t₁                     (doc 12)
                    ▼
 ┌────────────────────────────────────────────────────────────────────┐
 │ DECODE step 1                                                      │
 │   REBUILD here (n_tokens nt→1, gtype Prefill→Decode) — exactly     │
 │   once per run; KV regions + buffer pools survive the rebuild      │
 │   inputs: token_ids = [t₁], positions = [nt]                       │
 │   KV: layer ℓ writes t₁'s K/V at slot nt, attention reads slots    │
 │        0..=nt (nkv = nt+1 — the window grows by one)               │
 │   output: logits for t₂'s sampling                                 │
 └────────────────────────────────────────────────────────────────────┘
                    │ sample → t₂
                    ▼
 ┌────────────────────────────────────────────────────────────────────┐
 │ DECODE step 2, 3, …   SAME graph, zero rebuilds                    │
 │   try_reuse(params) → true every step                              │
 │   inputs refreshed: token_ids = [tᵢ], positions = [nt+i-1]         │
 │   one scheduler walk per step (one split walk per token)           │
 └────────────────────────────────────────────────────────────────────┘
                    │ sample → t₄
                    ▼
        stop gate: EOS? stop string? generated == n_predict?

Two structural things to notice before any code:

  • The loop consumes logits that already exist. Each iteration samples first, from the logits the previous iteration's forward produced (or from prefill's logits on iteration one), and only then runs the forward that produces logits for the next iteration. The forward is the last statement of the body — ModelDef::forward (main.rs:1782), not the first.
  • The rebuild happens inside the first forward, invisibly. The loop calls model.forward(...) on every step; the cache check (try_reuse) is inside forward_cached. The loop never knows whether a given call rebuilt or reused — from the loop's point of view every step is "fill two small inputs, execute, get logits".

2.2 Why a decode step is cheap relative to prefill

The run header prints two rates for a reason — prefill throughput and decode throughput differ by an order of magnitude, and it is worth being precise about where the savings come from. Take Qwen2.5-0.5B as the worked example (d_model = 896, 24 layers, 14 query heads / 2 KV heads × head-dim 64, so nkt = 128 KV channels per layer, n_vocab = 151,936; doc 10 uses the same model).

Budget 1 — weight streaming (unchanged, and dominant). Every matmul of every layer reads its whole weight matrix per forward, whether it multiplies it by 128 token rows or by 1. That is ~0.3 GB of quantized weights streamed from RAM per decode step, so on a ~60–100 GB/s memory system the floor is a few milliseconds per token no matter what the loop does — doc 10 §2.1 calls this "the decode physics". Graph reuse cannot shrink this budget; it only removes the overhead around it.

Budget 2 — matmul arithmetic (shrinks by nt). The multiply-add work of a matmul scales with the number of token rows. Prefill of a 128-token prompt does 128× the arithmetic of a decode step for the identical weights. This is why prefill shows high tok/s (the cost is amortized over many tokens) while decode shows the raw weight-streaming rate: the plan document records ~100 tok/s CPU decode for 0.5B (docs/COMPUTE-GRAPH-DESIGN.md §12, first row) — the loop is running as fast as RAM allows.

Budget 3 — attention (shrinks from quadratic to one scan). Prefill attention is causal all-pairs work over nt queries (doc 11); decode has one query, whose scan a data window bounds: the allocator resolves it from the KV cell store and hands it to the kernel as the attn_span / kv_map input (E1, C8b). The CPU path decodes that window — decode_window (cpu_backend.rs:683):

#![allow(unused)]
fn main() {
                // E1/C8b S2: the allowed cells are an explicit input, not a bound
                // derived from `positions` — that is what lets a batch hold several
                // sequences, and (S2) what lets one query's window be a list of runs.
                // The allocator resolved it from the cell store; here it is only
                // validated and decoded.
                let (runs, off) = decode_window(ins[3], nt, n_ctx)?;
}

No field of the graph knows or cares how large the cache has grown; the kernel reads exactly the cells the window names out of a region allocated at the full n_ctx from the start — the whole reason a fixed graph can serve a growing cache (§2.4). The cost grows linearly with context: at nkv = 1024 a decode step reads ≈ 24 MB of K/V (2 × 24 layers × 128 × 1024 × 4 B), and ~96 MB at n_ctx = 4096. The position-derived nkv = max_pos + 1 predates E1 (PR #3, 2026-09-17) and now lives only on Metal (metal_backend.rs:1588-1589).

What reuse removes. Build + assign + fuse + allocate are per-graph costs, not per-token math. Reuse means a decode step pays only: two small input fills (a 4-byte token id and a 4-byte position), one scheduler walk over the cached graph, and the math. The plan document's expected-gains table puts it plainly: decode's gain is "skips topology rebuild and CPU scratch allocation" (docs/COMPUTE-GRAPH-DESIGN.md §12).

2.3 One rebuild, then none: the params identity

Every forward call constructs a GraphParams value and asks the cache whether it may reuse. Trace the three calls of one run:

Calln_tokensn_outgtypecparams.gpufuse_qkv/fuse_ffnCache verdict
prefill (nt=23)231Prefillfalse (CPU build)false / falsemiss (empty cache) → build
decode step 1 (nt=1)11Decodefalsefalse / falsemiss (params differ) → rebuild
decode step 2, 3, … (nt=1)11Decodefalsefalse / falsehit → reuse

The interesting row is the second. The prefill and decode graphs are genuinely different programs, which is why the rebuild is honest work and not a technicality:

  • every buffer's shape is [nt, …] — activations are 23 rows vs 1 row (doc 07's layouts), so the whole liveness mapping must be redone;
  • the prefill graph contains the G3 tail reduction — an extra tail_ids input node and two GetRows nodes that cut the last layer's FFN and the lm_head down to the n_out output rows (src/models/qwen2/graph.rs:78-84, src/models/qwen2/graph.rs:256-257). In decode n_out == nt == 1, so those nodes don't exist at all;
  • on a GPU build, the decode fusions apply only when nt == 1 (src/models/qwen2/graph.rs:131-137), and the same gate in the CParams construction (src/models/qwen2/graph.rs:584-590): fuse_qkv. Op::FusedQKV merges 3 matmuls + 3 biases
    • 2 RoPEs + 2 KV stores into one kernel. Different node set ⇒ different graph.

And then the third row is the payoff: step 2's params and step 3's params are equal — same n_tokens, same n_out, same gtype, same cparams — so every remaining decode step of the run takes the reuse branch and does zero structural work. Within one generation run there is exactly one rebuild (prefill → first decode) and zero after it.

It is worth being explicit about what does not appear in the comparison: the token ids, the positions, and n_past. Those are execution data, injected into input nodes after the reuse check (src/models/qwen2/graph.rs:655-678). The plan document records this as the design's founding correction: an earlier sketch encoded n_past into the KV-cache op, which "causes the decode topology to change at every step and structurally breaks graph reuse" (docs/COMPUTE-GRAPH-DESIGN.md, revision note 1). Moving the position from the graph's structure into its data is what makes row 3 of the table possible at all.

2.4 Why a params-only comparison is enough to decide reuse

The claim that carries the whole design: if the params are equal, the topology is equal — so comparing six scalar/enum fields is a sound substitute for comparing graphs. For that to be sound, each field must be genuinely load-bearing — it must be able to change the node sequence. It is worth checking each one GraphParams (params.rs:88-97), defined below in §3.2):

  • n_tokens — every activation buffer's shape and loop trip counts derive from it; also selects the per-layer QKV build path (nt == 1 enables the decode fusions, src/models/qwen2/graph.rs:131-137).
  • n_seqs — deleted in E2 (A7 closure). A batch's sequence count is data, like n_past: the one topology decision it can force — explicit_span, the explicit attention window — lives in cparams and is derived from the KV reservations. Carrying the count here bought nothing and cost a rebuild whenever batching changed shape with the topology unchanged.
  • n_out — decides whether the G3 tail_ids input + tail GetRows nodes exist (n_out < nt) and how many rows the output buffer has.
  • gtype — Decode vs Prefill; today it is redundant with n_tokens == 1, but it names the intent (llama.cpp's llm_graph_params carries the same distinction) and keeps the door open for graph shapes that are not a function of token count alone.
  • cparams — the runtime knobs that reach into topology: n_ctx (sizes the KV regions and the RoPE/attention metadata), flash_attn (selects the attention node's mode), gpu (backend assignment is part of the built graph — a GPU that initialized between two calls must force a rebuild (params.rs:33, the gpu field), and the fusion gates fuse_qkv / fuse_ffn (the A/B env toggles must reliably force a rebuild, hence they live inside cparams and inside the equality check — the test fuse_flags_are_part_of_the_reuse_identity pins this, verify_structural (cache.rs:134).
  • weights_version — the future LoRA/reload hook: bumped whenever weights change, invalidating every cached graph (params.rs:96-97, the weights_version field).

If topology were not a pure function of these, a params hit would reuse a graph that was subtly wrong for the new inputs — the worst kind of bug, because it looks like a numerics problem. The defense is a debug-build-only second line of checks: GraphCache::verify_structural verify_structural (cache.rs:134-145) compares two graphs built from equal params node by node (op, shape, dependencies) and is wired into tests, so a non-deterministic builder — the one way params-equality could lie — is caught in CI rather than in production. In release builds the six-field comparison is all there is, and it is enough because the builder is deterministic by contract (src/models/qwen2/graph.rs:37-38: "deterministic in params — the reuse invariant").

This is a deliberate about-face from an earlier plan sketch that built the new graph first and compared node sequences ("build then compare") — which can never skip the rebuild, defeating the purpose. The plan document records the correction: "no graph build, no node-sequence comparison; the graph topology is a deterministic function of the parameters — the same invariant as llama.cpp allow_reuse()" (docs/COMPUTE-GRAPH-DESIGN.md §6).

2.5 What survives a rebuild vs what is recomputed

The rebuild between prefill and decode (and between turns in a chat session) redraws a sharp line. On the "recomputed" side, everything that describes the graph; on the "survives" side, everything that holds data:

Survives the rebuildRecomputed on rebuild
The GraphAllocator itself (it lives inside GraphCache, cache.rs:38-45)The node list (Self::build, src/models/qwen2/graph.rs:611)
Registered weights (registered once by name; register_weight, alloc.rs:374)Backend assignment (assign_backends, src/models/qwen2/graph.rs:628)
The per-layer KV regions — kv.{ℓ}.k / kv.{ℓ}.v, allocated once at full n_ctx size and never freed (ensure_kv, alloc.rs:995-1002; alloc_graph explicitly frees only liveness buffers, alloc_graph (alloc.rs:513-519))The fusion pass (FusionPass::run, src/models/qwen2/graph.rs:636-640)
Backend buffer pools (freed liveness buffers return to their pool; the memory is recycled, not released)The node→buffer mapping (alloc_graph clears node_to_buf, alloc.rs:519)
The monotonic graph uid of the reused graph (a rebuilt graph gets a fresh uid — which is exactly what invalidates a stale CUDA Graph capture, replace_graph (cache.rs:106-114) + §3.4)Cross-backend staging buffers (keyed by (graph uid, node, backend) and surviving a rebuild and a re-map — E4 S3; alloc_graph clears only the in-flight cross_pending set, alloc_graph (alloc.rs:513-528))

The first row is the one that matters most: the allocator is a field of the cache, not a local of the forward function, so a rebuild cannot take the KV cache with it. cache.rs's module comment states the contract in two sentences (cache.rs:9-12, the "survives graph rebuilds" module comment): the allocator "lives inside the cache and survives graph rebuilds: the persistent KV regions are exactly the KV cache, so a prefill→decode transition (different n_tokens/gtype ⇒ rebuild) must not lose them. Only the node/buffer mapping is recomputed on rebuild." Doc 07 covers the region mechanics; this document cares about the consequence — the decode step after a rebuild finds the prompt's context already in place, at the same addresses, and simply appends.

2.6 Multi-turn conversation: append-only KV, delta prefill

A chat session (--cnv) is the decode loop wearing a longer timeline. The session keeps a host-side mirror of the KV contents — stream_tokens, a plain Vec<u32> — plus a write cursor current_pos that always equals its length (conversation.rs:399-401, the stream_tokens/current_pos fields); the real-model test asserts this invariant (conversation/tests/real_model.rs:90-91, the real-model assert). The KV region itself is position-addressed (slot p of layer ℓ's K region holds the K vector of the token at position p), which makes the whole session strategy possible:

  • Append-only. Turn 2 does not rewrite turn 1's slots. user_turn renders only the delta — the new user message wrapped in the template's turn separator — tokenizes it, and prefills it at positions current_pos .. current_pos + delta.len() (conversation.rs:949-952, the delta prefill). The engine then decodes the assistant reply with the same loop as before. Each turn costs O(delta), never O(history).
  • Same graph, more rebuilds. Each turn boundary flips n_tokens from 1 back to delta_len (Prefill) and then to 1 again (Decode) — a rebuild per flip. That is two cheap rebuilds per turn against one very expensive re-prefill of the whole history; the KV regions and buffer pools survive every flip (§2.5).
  • Rollback without erasing. /regen rewinds current_pos to turn_pos (the start of the last turn's delta) and regenerates (conversation.rs:965-977, the regen_turn rollback). The rolled-back slots in the KV region are now stale but never read: attention reads only the window the cursor bounds. Regeneration simply overwrites those slots as it appends. (A full /clear or a template mismatch falls back to rehydrate_full — reset the cache, re-render everything, re-prefill once, rehydrate_full (conversation.rs:518).)
  • Seams kept consistent. If a turn ended without an end-of-turn token, the next turn first inserts the EOT token into the KV at the cursor (conversation.rs:771-776, the EOT insert) so the region keeps matching what the chat template's canonical render would have produced — the module doc calls this the §5.4 KV-consistency invariant (conversation.rs:16-19, the module's KV-consistency note).

2.7 What stops generation, and where that is decided

Three gates, three different owners, all on the sampler's output before or instead of the next forward (§3.2 walks the code):

  1. End-of-generation token. The sampled id equals the model's eos or <|im_end|> (is_stop_token, main.rs:1868-1870; the ids come from special_tokens() (main.rs:1560). Decided in the loop is_stop_token (main.rs:1749-1751); the token is not appended and not fed to the graph — the run just ends. (Conversation mode deliberately does the opposite: it writes the EOG token into the KV before breaking, to keep the region matching the template's canonical next-turn render — (conversation.rs:1330-1343, the EOG write to the KV).)
  2. Stop strings (--stop). Byte-level suffix match over the entire generated byte stream, so a stop string split across token boundaries is still caught by match_stop_suffix (main.rs:1761-1768; doc 12 §3 owns the matcher). The matching bytes are truncated out of the kept text; already-printed bytes stay in the terminal, llama.cpp-style.
  3. The -n cap. The while condition generated.len() < params.n_predict the loop head (main.rs:1727), default 512 in GenParams::default (main.rs:110-112). In conversation mode a fourth gate joins: the write cursor reaching n_ctx stops cleanly instead of overflowing the KV regions (conversation.rs:1314-1318, the cursor guard).

Note the asymmetry between the gates: EOS and stop strings break before the forward at the bottom of the body, so no forward is wasted on a token nobody reads. The -n cap, checked only at the loop head, lets the final iteration's forward run — one discarded forward per full-length run, the price of a simple loop condition (§3.4).

3. Implementation

3.1 Data in / data out

Per decode step, the engine touches a remarkably small amount of changing data. Everything else — weights, graph, KV regions, buffers — was set up once and is only read:

DataType / shapeWhere it comes fromWhere it goes
token_ids inputI32, shape [1, 1, 1, 1] (one token)step 1: the token sampled from prefill's logits; every later step: the previous iteration's sampled.token_idfill_input_i32 writes it into the input node's buffer (src/models/qwen2/graph.rs:657)
positions inputI32, shape [1, 1, 1, 1]current_pos (main.rs:1561) — starts at input_ids.len(), incremented once per step at current_pos (main.rs:1790)same, src/models/qwen2/graph.rs:659; the KV store uses it as the write slot, attention as the last readable slot
K/V regions (per layer)f32, [nkt][n_ctx] (128 × 4096 = 2 MiB per region for Qwen2.5-0.5B at --n-ctx 4096)allocated once at first use (ensure_kv, alloc.rs:995-1002); contents: prefill's prompt + every generated token so farstep ℓ's store writes slot position; step ℓ+1's attention reads slots 0..=nkv-1
logits outputf32, [n_vocab] = 151,936 × 4 B ≈ 607 KBthe graph's output buffer (graph.outputs[0], copied to host at src/models/qwen2/graph.rs:752-764)moved into the loop's logits variable (main.rs:1782-1783, the logits = model.forward(...) assignment) for the next sample
prev_tokens windowVec<u32>, ≤ 64 idsprompt tail + generated tokens — recent_window (main.rs:1568, 1752-1757)the sampler's repeat/frequency/presence penalties (doc 12)
generatedVec<u32>pushed per step (main.rs:1752, the generated.push(...) call)stop checks, final stats; decode_bytes streams it to stdout

Two details of this table deserve unpacking.

The I32 bit-pattern trick. The allocator's buffers are f32 pools — every node reads and writes f32 slices, whatever its logical type. Token ids and positions are integers. Rather than special-case integer buffers, fill_input_i32 stores each u32 bit pattern reinterpreted as an f32 value fill_input_i32 (alloc.rs:1903), and the kernels that consume these inputs (attention, KV store) convert back with f32::to_bits() as usize (cpu_backend.rs:334-338, the store's position decode; :919-920, attention's). This is exact for values below 2²⁴ — vocabulary ids and positions never come close — and it keeps one uniform buffer format across the whole graph (AGENTS.md "Compute Graph" rule 4). f32::from_bits(v) does no rounding at all; it is a transmute, not an arithmetic conversion, which is why a round trip through the buffer is lossless.

One forward call per step, five hidden phases. The loop's single model.forward(...) — ModelDef::forward (main.rs:1782) — expands inside forward_cached into: build params (six fields) → try_reuse (usually a hit) → fill 2–3 small inputs → one scheduler.execute walk → copy the output buffer back. On a rebuild step the same call additionally runs build → register → assign → fuse → alloc. The loop code has no idea any of that exists — which is the point of the ModelDef::forward facade ModelDef::forward (models/mod.rs:222).

One name in the call needs a sentence for honesty: forward's signature used to take a &mut KVCache (main.rs passed a kv_cache created at load). On the graph path that legacy type was always ignored — the callee bound it as _kv and forward_cached never took it — so #252 deleted the parameter, the KVCache type and the construction rather than keep a dead API shape. The real KV lives in the allocator's persistent regions (doc 07), and the call now reads forward(tokens, positions, n_out, n_ctx).

3.2 Key code

The loop's setup (src/main.rs:1553-1568, the decode-loop locals)

Before the loop starts, a handful of locals are established — each one is a hand the loop plays with on every iteration:

#![allow(unused)]
fn main() {
    let mut logits = last_logits;
    if trace_on {
        crate::trace::begin_phase("decode");
    }
    let gen_start = Instant::now(); // pure-decode start (llama "Generation" caliber)
    let mut generated: Vec<u32> = Vec::new();
    let special = model.special_tokens();
    let mut current_pos = input_ids.len();

    // Seeded RNG for reproducible sampling; recent-token window for the
    // penalties (llama.cpp repeat_last_n default = 64), seeded with the prompt
    // tail so the first generated tokens are penalized too.
    let mut rng = rand::rngs::StdRng::seed_from_u64(params.seed);
    const REPEAT_LAST_N: usize = 64;
    let mut prev_tokens = sampler::recent_window(&input_ids, REPEAT_LAST_N);
}
  • logits starts as prefill's last-token logits — the loop's first iteration samples from them without running a forward first. This is the delayed-forward shape of §2.1 made concrete.
  • current_pos = input_ids.len(): the first generated token will be written at KV slot nt — the slot after the prompt. The KV regions were sized for n_ctx slots, and ctx = params.n_ctx.max(input_ids.len()) (main.rs:1452-1453, the ctx computation) guaranteed at prefill time that the prompt fits; from here the cursor only ever increments.
  • special carries the model's stop-token ids (eos, and <|im_end|> when the model defines one) — gate 1 of §2.7.
  • rng is seeded once (--seed, default 42): the whole run's sampling is a deterministic function of the seed, which is what makes minfer's llama.cpp-comparison claims testable.
  • prev_tokens seeds the 64-token penalty window with the prompt's tail, so even the first sampled token is subject to repeat penalties — doc 12's recent_window helper.

The stop-string machinery follows immediately — stop_refs (main.rs:1570-1582): user --stop strings are copied to byte vectors (stop_bytes/stop_refs), and two cursors track output — full, every generated byte ever, and emitted, how many of those bytes have already been flushed to stdout. The pair exists because a stop string may straddle token boundaries: earlier tokens were already printed before anyone could know a stop string was forming.

The decode loop, first half — sample and stop (src/main.rs:1727-1772, through the stop-string check)

#![allow(unused)]
fn main() {
    while generated.len() < params.n_predict {
        t0 = std::time::Instant::now();
        let sampled = match sampler::sample_with_config_grammar(
            &mut logits,
            &sampler_cfg,
            &prev_tokens,
            &mut mirostat,
            &mut grammar_state,
            &mut rng,
        ) {
            Ok(s) => s,
            Err(e) => {
                // The loud stop the grammar contract requires: the emitted text
                // is a valid prefix, nothing illegal is appended.
                eprintln!("\n[grammar] {e}");
                break;
            }
        };
        if timing {
            t_samp += t0.elapsed().as_secs_f64();
        }

        if is_stop_token(sampled.token_id, &special) {
            break;
        }
        generated.push(sampled.token_id);
        token_trace(current_pos, sampled.token_id);
        prev_tokens.push(sampled.token_id);
        if prev_tokens.len() > REPEAT_LAST_N {
            prev_tokens.drain(0..prev_tokens.len() - REPEAT_LAST_N);
        }

        // Stop-string detection on the FULL byte stream before emitting.
        full.extend_from_slice(&tokenizer.decode_bytes(&[sampled.token_id]));
        if let Some(cut) = sampler::match_stop_suffix(&full, &stop_refs) {
            full.truncate(cut);
            if cut > emitted {
                hi.feed(&full[emitted..]);
                emitted = full.len();
            }
            break;
        }
        if emitted < full.len() {
            hi.feed(&full[emitted..]);
            emitted = full.len();
        }
}

Reading it as a state machine, one iteration:

  1. Gate: the cap. The while condition is gate 3: a full quota exits without touching anything.
  2. Sample. sample_with_config_grammar (doc 12) consumes and mutates logits (hence &mut): the config's penalties and DRY, then the grammar mask, then top-k → top-p → temperature (or mirostat) → one seeded draw. Out comes one token id; the loop never inspects logits itself, and the 607 KB of scores are noise to everyone but the sampler. A SampleError (no token the grammar allows) is the loud stop the excerpt handles: the loop prints it and breaks, never emitting an arbitrary token.
  3. Gate: EOS. is_stop_token (main.rs:1868-1870) checks eos/im_end; a hit breaks before the token is pushed — the sentinel is not text.
  4. Append. The token enters generated and the 64-entry sliding penalty window, drained from the front.
  5. Gate: stop strings — the tail of the excerpt: the token's bytes join full, and match_stop_suffix (main.rs:1760-1768) tests whether it now ends with a stop string; on a hit the text is truncated before the match, else the new bytes are flushed to stdout (main.rs:1769-1772).

Only when all three gates pass does the iteration continue to the forward — the loop never runs the transformer for a token it has already decided to discard.

The decode loop, second half — forward and advance (src/main.rs:1774-1791, the forward call and cursor advance)

#![allow(unused)]
fn main() {
        // forward() returns n_out*nv logits (n_out=1 for single-token decode,
        // exactly n_vocab), so move the Vec in place instead of copying 607 KB/token.
        if trace_on {
            let text = String::from_utf8_lossy(&tokenizer.decode_bytes(&[sampled.token_id]))
                .into_owned();
            crate::trace::set_token(sampled.token_id, &text);
        }
        t1 = std::time::Instant::now();
        logits = model.forward(&[sampled.token_id], &[current_pos], 1, ctx);
        if trace_on {
            crate::trace::attach_step(&logits);
        }
        if timing {
            t_fwd += t1.elapsed().as_secs_f64();
            n_tok += 1;
        }
        current_pos += 1;
    }
}

The heart of the whole document is one line:

#![allow(unused)]
fn main() {
logits = model.forward(&[sampled.token_id], &[current_pos], 1, ctx);
}

Four arguments, and every one of them is either constant across the whole loop or a single value that changes: the input is a one-element slice containing last iteration's sampled token; the position is a one-element slice containing the cursor; n_out is 1 (we need logits for exactly one token); ctx is the constant that sized the KV regions. Nothing here says "decode" or "step 37 of 500" — the loop is identical whether it is the first or the four-hundredth step, and the graph machinery underneath treats it as just another forward whose params happen to be unchanged. The MINFER_TIMING arrows around it (t1, t_fwd) split the per-token wall clock into sampling time and forward time — §4 uses this to show that the loop overhead is microseconds against the forward's milliseconds.

The trace_on block records the token and the logits for the viz trace (§4); note it is hoisted out of the loop as a bool — trace_on (main.rs:1456-1457) so the steady-state loop does one env-var read fewer per step. And current_pos += 1 is the KV cursor's whole life: prefill filled slots 0..nt, this line claims slot nt, nt+1, … one per step, until either a stop gate fires or (in conversation mode) a guard stops it before the regions overflow.

is_stop_token itself is the smallest function in the pipeline is_stop_token (main.rs:1868-1870):

#![allow(unused)]
fn main() {
fn is_stop_token(id: u32, special: &models::SpecialTokens) -> bool {
    id == special.eos || Some(id) == special.im_end
}
}

im_end is an Option because not every model defines <|im_end|> (Qwen2.5 and Qwen3 chat models do). Two ids, one boolean — but this function is the model's only "voice": everything else about stopping is user policy (stop strings, -n), while this is the model saying "I'm done".

Inside the forward: building the params — GraphParams (src/models/qwen2/graph.rs:541-597)

The loop's forward call lands in forward_cached, which first expresses "what kind of graph does this step need?" as a plain data value:

#![allow(unused)]
fn main() {
        let params = GraphParams {
            n_tokens: nt,
            n_out,
            gtype: if nt == 1 {
                GraphType::Decode
            } else {
                GraphType::Prefill
            },
            cparams: CParams {
                n_ctx,
                flash_attn: false,
                explicit_span,
                kv_map,
                gpu: metal_on || cuda_on,
                gpu_layers: if metal_on || cuda_on {
                    model.offload.plan.gpu_layers
                } else {
                    0
                },
                fuse_qkv: nt == 1
                    && (metal_on || cuda_on)
                    && (cuda_on || !explicit_span)
                    && !std::env::var("MINFER_NO_FUSE_QKV").map_or(false, |v| v == "1"),
                fuse_ffn: nt == 1
                    && (metal_on || cuda_on)
                    && !std::env::var("MINFER_NO_FUSE_FFN").map_or(false, |v| v == "1"),
                kv_format: model.kv_format,
            },
            weights_version: 1,
        };
}

Everything the reuse decision will ever need is assembled here, before the cache is consulted (comments elided): gtype is derived from nt; the new explicit_span / kv_map pair selects the attention window layout — both are topology, both derived from the KV reservations, never from n_past — and gpu (params.rs:33) records whether a GPU backend participates, a param because backend assignment is baked into the graph, so "Metal became available between two calls" must force a rebuild. gpu_layers carries the E5 offload plan (its assignment is topology too), and kv_format (params.rs:61-67, C4) sizes every KV cell, so a graph built for one format never serves another. The fusion gates are the subtlest part: runtime env vars (MINFER_NO_FUSE_QKV=1 / MINFER_NO_FUSE_FFN=1) change which Ops the graph contains, so they live in CParams — flipping one mid-run fails the comparison and rebuilds with the new state. That is how an environment variable safely joins a build cache, and the two gates are separate so that A/B-ing one cannot flip the other.

Before this struct is built, forward_cached has already done two quiet checks worth noting (src/models/qwen2/graph.rs:513-527): it asserts every position is below n_ctx — out-of-range positions would write past the KV regions, so the failure is a loud panic, not silent corruption (src/models/qwen2/graph.rs:513-520) — and it probes GPU availability (metal_on / cuda_on), which feeds the gpu field above.

The reuse decision itself GraphCache (src/graph/cache.rs:38-64)

With params in hand, the cache is asked one question (src/models/qwen2/graph.rs:605): if !cache.try_reuse(&params).expect("re-map onto a cached graph") { … }. Here is the whole machinery:

#![allow(unused)]
fn main() {
    pub fn try_reuse(&mut self, params: &GraphParams) -> Result<bool, String> {
        let Some(pos) = self
            .graphs
            .iter()
            .position(|(p, _)| Self::params_match(p, params))
        else {
            return Ok(false);
        };
        let (_, graph) = self.graphs.remove(pos);
        …(the `Err` carries the failed re-map back to the caller)…
        self.alloc.alloc_graph(&graph)?;
        self.graphs.insert(0, (params.clone(), graph));
        self.reuses += 1;
        Ok(true)
    }
}

That is the entire cache: an MRU Vec of at most MAX_CACHED_GRAPHS = 8 (GraphParams, ComputeGraph) pairs, one per distinct shape (cache.rs:36-45) — a session alternates a handful, so they all fit. Six field comparisons (params_match, cache.rs:95-101) replace rebuilding a few hundred nodes, re-assigning backends, re-running fusion, and re-walking liveness. A miss means no cached graph has these params; a hit moves that graph to the front and re-maps the allocator onto it, which is the one step that can fail — hence the Result, which the caller .expects (loud panic).

The GraphParams and CParams types behind the comparison CParams (src/graph/params.rs:30-63) and GraphType (params.rs:19) carry exactly the fields argued for in §2.4:

#![allow(unused)]
fn main() {
pub struct GraphParams {
    pub n_tokens: usize,
    /// Number of output (tail) rows: the last layer's FFN + lm_head run on the
    /// last `n_out` rows only (llama `inp_out_ids`). Part of the topology —
    /// a change forces a rebuild.
    pub n_out: usize,
    pub gtype: GraphType,
    pub cparams: CParams,
    /// Bumped by the model whenever weights change (LoRA switch, reload).
    pub weights_version: u64,
}
}

and inside CParams: n_ctx, flash_attn, gpu, fuse_qkv, fuse_ffn (src/graph/params.rs) — each documented there with the reason it belongs in the identity. The module's opening comment is the invariant in one breath (params.rs:1-7, the module comment): these are "the ONLY inputs to graph reuse … n_past (KV position) is deliberately absent: it is execution data."

The rebuild branch forward_cached (src/models/qwen2/graph.rs:456, 599-645)

When the comparison fails, the five-phase pipeline of docs 05–08 runs, and its result is stored back into the same cache:

#![allow(unused)]
fn main() {
        if !cache
            .try_reuse(&params)
            .expect("re-map onto a cached graph")
        {
            let rebuild_t0 = std::time::Instant::now();
            let trace_rb = std::env::var("MINFER_REBUILD_TRACE").map_or(false, |v| v == "1");
            let mut graph = Self::build(model, &params);
            let sched = BackendScheduler::new();
            {
                let alloc = cache.alloc();
                Self::register_graph_weights(model, alloc);
                …(the `#[cfg]` enable_metal / enable_cuda blocks, doc 07)…
                alloc.set_offload_plan(Some(model.offload.plan));
                sched.assign_backends(&mut graph, alloc);
}

…(the middle of the block is the fusion pass's backend list, which comes from the allocator's registry view rather than a hand-built Vec — src/models/qwen2/graph.rs:629-635)…

#![allow(unused)]
fn main() {
                let backends: Vec<&dyn Backend> = alloc.fusion_backends();
                FusionPass::new().run(&mut graph, &backends, &|g, id| {
                    g.node(id)
                        .backend
                        .and_then(|b| alloc.fusion_backend_index(b))
                });
                alloc.alloc_graph(&graph).unwrap();
            }
            cache.replace_graph(graph, params);
        }

        let (graph, alloc) = cache.current().unwrap();
}

Three details turn this from "a rebuild" into "a rebuild that preserves the session", and the .expect("re-map onto a cached graph") above makes a failed re-map a loud panic, never a silent rebuild:

  • cache.alloc() — the allocator is borrowed from the cache; register_graph_weights re-registers weights by name, idempotently (doc 03/07), and set_offload_plan (E5) puts the block plan in force before assignment. alloc_graph frees the old liveness mapping but keeps the persistent KV regions (§2.5, below).
  • The pipeline is the full one — assign, fusion (backend-gated per node), allocate — because a rebuilt graph must be indistinguishable from a fresh-process graph; determinism (§2.4) is what makes that sound. MINFER_REBUILD_TRACE=1 prints rebuild_t0's build+assign+alloc ms.
  • replace_graph (cache.rs:106-114) stores the new node list and params, and stamps the graph with a fresh monotonic uid — replace_graph (cache.rs:106-107) — the CUDA Graph cache keys captures by uid, so a new topology naturally invalidates the old capture while a reused graph keeps its uid and its replay (doc 15).

Then, on both paths (rebuilt or reused), the inputs are refreshed:

#![allow(unused)]
fn main() {
        // refresh input data (positions/ids are data, not topology)
        let ids: Vec<u32> = tokens.to_vec();
        alloc.fill_input_i32(graph, "token_ids", &ids).unwrap();
        let pos: Vec<u32> = positions.iter().map(|&p| p as u32).collect();
        alloc.fill_input_i32(graph, "positions", &pos).unwrap();
        …(E1/E2: `fill_batch_inputs` resolves every query's window)…
        alloc
            .fill_batch_inputs(graph, batch)
            .unwrap_or_else(|e| panic!("batch inputs: {e}"));
        // G3: the last-layer tail-row reduction reads `tail_ids` (filled when
        // the graph was built with n_out < nt, i.e. prefill)
        if graph
            .inputs
            .iter()
            .any(|&i| graph.node(i).name == "tail_ids")
        {
            alloc.fill_input_i32(graph, "tail_ids", &out_rows).unwrap();
        }
}

This is "positions are data" in executable form: the same code runs for a 23-token prefill (23 positions) and for decode step 400 (one position, value 422) — the graph is never told which situation it is in; it reads the buffers. fill_batch_inputs (src/models/qwen2/graph.rs:667-669) resolves each sequence's window first; the tail_ids fill is conditional and takes the batch's out_rows (src/models/qwen2/graph.rs:677) — that input only exists in prefill graphs (n_out < nt), so its presence is data too.

Finally sched.execute(graph, alloc) runs one scheduler walk (src/models/qwen2/graph.rs:692), and the output buffer is copied back as the returned logits (src/models/qwen2/graph.rs:752-764). With G3 active the output buffer already holds exactly n_out × n_vocab values, so the return is either the buffer itself or a truncated copy — never a full-nt logits matrix (doc 09 covered the prefill-side benefit; in decode n_out == nt == 1, so the buffer is one row regardless).

Why the KV survives: the allocator's two kinds of memory (alloc_graph (src/graph/alloc.rs:513-519), alloc_persistent (alloc.rs:1080-1089))

The claim everywhere above is that a rebuild "keeps the KV". The mechanism is that the allocator distinguishes two kinds of buffers, and only one kind is freed on rebuild. First, alloc_graph — which runs on every rebuild — clears only the liveness-managed mappings:

#![allow(unused)]
fn main() {
    /// Runs on every graph (re)build: previous liveness buffers are released
    /// back to their pools, while **persistent regions (KV cache) survive** —
    /// they are the KV cache and must persist across prefill→decode rebuilds.
    pub fn alloc_graph(&mut self, graph: &ComputeGraph) -> Result<(), String> {
        let prev: Vec<(Backend, usize)> = self.buf_alive.keys().copied().collect();
        for (b, id) in prev {
            self.free_in_pool(b, id);
        }
        self.buf_alive.clear();
        self.node_to_buf.clear();
        // F5: a rebuild starts a fresh execution — no staging copy can be in
        // flight across it (the boundary either completed or failed loudly).
        self.cross_pending.clear();
}

Note the free_in_pool call: freed buffers return to their backend's pool rather than being deallocated — the prefill graph's big activation buffers are immediately available to serve the decode graph's smaller ones, which is why a rebuild does not reallocate GPU memory or fragment the pools (doc 07's pool mechanics). cross_pending.clear() only forgets the in-flight copies: the staging buffers themselves are keyed by (graph uid, node, backend) and survive both a rebuild and a re-map — re-creating them per switch would leak, as staging is alloc_fresh, which never recycles from the free list (alloc.rs:520-528). The KV regions are never in buf_alive: ensure_kv creates them on first use into a separate persistent list nothing touches:

#![allow(unused)]
fn main() {
    fn ensure_kv(…, n_embd: usize, row_elems: usize, n_ctx: usize) -> Result<[BufRef; 2], String> {
        …(the C4 packed-width checks)…
        let elems = row_elems * n_ctx;
        if let Some(region) = self.kv.get(layer) {
            …(elems / backend / packed checks)…
            return Ok([region.k, region.v]);
        }
        let k = self.alloc_persistent(&format!("kv.{layer}.k"), backend, elems);
        let v = self.alloc_persistent(&format!("kv.{layer}.v"), backend, elems);
        self.kv
            .insert(layer, k, v, n_embd, row_elems, n_ctx, packed);
        Ok([k, v])
    }

    /// Allocate a persistent (never-freed) region on a backend.
    pub fn alloc_persistent(&mut self, name: &str, backend: Backend, size: usize) -> BufRef {
        let id = self.alloc_exact_in_pool(backend, size);
        self.persistent.push(PersistentBuf {
            name: name.to_string(),
            backend,
            id,
        });
        BufRef::own(backend, id, size)
    }
}

ensure_kv is the "allocate once, then look up forever" pattern: the first graph that touches layer ℓ's KV creates both regions at the full n_ctx-sized extent (elems = row_elems × n_ctx — 2 MiB per region for 0.5B f32 at --n-ctx 4096; a packed Q8_0 region counts packed words, not f32 elements), and every later graph — including the decode graph of every subsequent step — just gets the same BufRefs back (alloc.rs:1045-1071, the existing-region early return), allocated exact into the never-freed persistent list by alloc_persistent (alloc.rs:1080-1089), during alloc_graph's walk of KvcacheStore/KvcacheLoad (alloc.rs:670-681: the store node's buffer is the K region; V is its sibling) — n_ctx is a CParams field because kv_elems: nkt * n_ctx (src/models/qwen2/graph.rs:158) fixes its extent.

This is the exact contract doc 07 promised and the decode loop depends on: the KV cache is not a structure the graph owns — it is two never-freed buffers per layer that the graph borrows by position. The graph can be thrown away and rebuilt at every step boundary without touching a single cached K or V value.

One split walk per token (src/graph/scheduler.rs:219-285, the per-split boundary and walk)

Reuse also means the execution machinery runs the same cached plan every step. execute split_graph (scheduler.rs:87) splits the graph into contiguous same-backend ranges once per call, then walks them. The per-split boundary handling is where cross-backend sync and copies happen — and on a CPU-only run there is exactly one split, so none of it fires:

#![allow(unused)]
fn main() {
        for split in &splits {
            if let Some(pb) = prev_backend {
                if pb != split.backend {
                    // 1. retire the previous backend's async work (#138)
                    alloc.retire_backend(pb);
                    // 1b. staged Metal/CUDA captures are valid now — read back
                    flush_metal_captures(graph, alloc, &mut staged, trace_on, live_on);
                    flush_cuda_captures(graph, alloc, &mut cuda_caps, trace_on, live_on);
                    // 2. copy this split's inputs across backends
                    for &inp in &split.inputs {
                        alloc.copy_across(graph.uid, inp, split.backend)?;
                    }
                }
            }
}

and inside each split, the per-node walk (scheduler.rs:305-326, the node loop):

#![allow(unused)]
fn main() {
            for id in split.node_range.0..split.node_range.1 {
                let node = graph.node(id);
                if capture && node.is_input() {
                    // inputs are host-filled before execute — no pending GPU
                    // work, so reading them here is always current
                    if let Some(br) = alloc.node_buffer(id) {
                        if let Some(d) = read_host_buffer(alloc, br.backend, br.id)
                            .and_then(|d| window_of(br, d))
                        {
                            record_node_data(node, d, trace_on, live_on);
                        }
                    }
                }
                if node.is_input() {
                    continue; // data pre-filled by the allocator
                }
                // dead nodes (no consumers, not outputs) get no buffer — the
                // fusion pass can orphan them (e.g. silu folded into SwiGLU);
                // they are skipped, not executed
                let Some(br) = alloc.node_buffer(id) else {
                    continue;
                };
}

The comment at scheduler.rs:184-186 (the one-step-per-execute() note) is the decode-relevant summary: capture happens "one step per execute() (prefill = 1 step, each decode forward = 1)". So "one split walk per token" is literal: each decode step re-walks the split list and executes every node in build order — the kernels run fresh every step (new data!), but the plan (which nodes, which backends, which buffers, which splits) is the cached graph's. copy_across only enqueues (F5/#138): the single wait is issued at the consumer's first read, and the capture read above is windowed (window_of, scheduler.rs:483). The KV ops resolve their layer's regions by index — kv_pair (scheduler.rs:360-370) — then dispatch to the backend's execute_node (scheduler.rs:398); a GPU build replays a captured split by uid (doc 15 owns that).

The one structural difference: prefill's G3 tail — tail_ids (src/models/qwen2/graph.rs:252-258)

It is worth seeing the actual node-level difference that forces the prefill→decode rebuild — the G3 tail-row reduction, which exists only in prefill graphs:

#![allow(unused)]
fn main() {
            let wo = b.matmul(attn_out, l.wo.as_ref().unwrap(), None);
            let is_last = il == model.layers.len() - 1;
            // G3: reduce to the tail n_out rows BEFORE the last layer's FFN
            // (llama `ggml_get_rows(cur/inpSA, inp_out_ids)` at
            // qwen2.cpp:106-108) — ffn_norm, gate/up/down, swiglu, both
            // residuals and lm_head all run on n_out rows only. The tail_ids
            // input itself is declared at the graph head (see there).
            if is_last && params.n_out < nt {
                let tail_ids = tail_ids.expect("tail_ids input declared when n_out < nt");
                let cur_tail = b.get_rows(wo, tail_ids, [ne, params.n_out, 1, 1]);
                let res_tail = b.get_rows(residual, tail_ids, [ne, params.n_out, 1, 1]);
                h = b.add(res_tail, cur_tail);
            } else {
                h = b.add(residual, wo);
            }
}

Prefill computes attention and projections for all nt tokens (it must — every token's K/V goes into the cache), but only the last token's logits are wanted. So after the last layer's attention output projection, two GetRows nodes gather just the n_out tail rows (tail_ids is filled with nt-1, …, nt-n_out at src/models/qwen2/graph.rs:677), and the remaining FFN + residual + lm_head run on n_out rows instead of nt. With n_out = 1 and a 2000-token prompt, that saves a 2000-row lm_head GEMM — a 151,936 column × 2000 row output — per run. In decode, n_out == nt == 1, the if is false, and the graph simply has no tail nodes at all. Same builder, one branch, two different programs — the honest reason the params comparison must fail across the prefill→decode boundary.

Multi-turn conversation: the delta prefill (src/conversation.rs:949-952, the delta prefill call)

The conversation engine (GraphEngine, conversation.rs:113-117) wraps the same forward_graph_cached with a session-private cache — turn 9 of a chat reuses turn 8's decode graph, not just this turn's. When the user submits a new message, user_turn prefills only the delta:

#![allow(unused)]
fn main() {
        // 5. Record the user message, set the rollback point, prefill the delta.
        self.messages
            .push(("user".to_string(), Some(input.to_string())));
        self.turn_pos = self.current_pos;
        let positions: Vec<usize> =
            (self.current_pos..self.current_pos + delta_toks.len()).collect();
        let logits = engine.forward(&delta_toks, &positions, 1);
        self.stream_tokens.extend_from_slice(&delta_toks);
        self.current_pos += delta_toks.len();
        // The penalty window is reseeded **after** the delta is appended (llama.cpp feeds prompt tokens into the sampler window).
        self.prev_tokens = sampler::recent_window(&self.stream_tokens, REPEAT_LAST_N);
}

Compare with doc 09's first prefill: positions 0..nt there, positions current_pos..current_pos + delta here — the same graph builder, the same forward, just a later window of slots. turn_pos records the delta's start so /regen can rewind to it (§2.6). After this call the cache holds a prefill graph for delta_len tokens; the loop's first decode forward rebuilds to a 1-token graph — two rebuilds per turn, both cheap, both KV-preserving.

The conversation decode loop — generate_assistant_with_logits (conversation.rs:1271-1389) is the same sample → gates → forward shape, with two additions the single-shot CLI doesn't need. First, a context guard before sampling (conversation.rs:1314-1318, the cursor guard): if current_pos >= n_ctx, the turn ends "cleanly" — reported as hit_n_predict — because writing at slot n_ctx would overflow the regions (the single-shot path instead relies on forward_cached's position assert, src/models/qwen2/graph.rs:513-520). Second, the EOG token is written into the KV before breaking:

#![allow(unused)]
fn main() {
            if self.is_eog(sampled.token_id) {
                // The EOG must be written to the KV: the canonical render carries the EOG marker after the assistant message
                // (§5.4), and llama.cpp also decodes the EOG before stopping. Not writing it would break the invariant.
                stopped_by_eog = true;
                self.prev_tokens.push(sampled.token_id);
                if self.prev_tokens.len() > REPEAT_LAST_N {
                    self.prev_tokens
                        .drain(0..self.prev_tokens.len() - REPEAT_LAST_N);
                }
                self.stream_tokens.push(sampled.token_id);
                let _ = engine.forward(&[sampled.token_id], &[self.current_pos], 1);
                self.current_pos += 1;
                break;
            }
}

The single-shot loop treats EOS as pure control flow — break, drop the token. The conversation loop cannot: the session's stream_tokens mirror must equal what the chat template would render for the next turn, and a canonical render contains the end marker after each assistant message. So the EOG token is pushed, decoded through the graph at the cursor (its K/V land in the regions), and only then does the turn end. Two stop policies, one principle each: the CLI stops the text; the session keeps the KV mirror canonical (conversation.rs:16-19 — the module's §5.4 note — documents the invariant this preserves).

3.3 Design choices (why this shape and not another)

Params-only comparison, not build-then-compare. The original plan sketch built the new graph and diffed node sequences against the cached one (docs/COMPUTE-GRAPH-DESIGN.md §6 records the correction). That design can never skip the build — the expensive half of the rebuild — so it saves nothing. The shipped design inverts the burden of proof: the builder is required to be deterministic in GraphParams (each architecture's build doc-comment carries the invariant, src/models/qwen2/graph.rs:37-38), and once that contract holds, six field comparisons are a complete decision procedure. The failure mode of the contract (a non-deterministic builder) is covered in debug builds by verify_structural (cache.rs:134-145), which re-derives the graph and compares node-by-node — a test-only safety net that keeps the production path comparison-free.

Rebuild once, don't parameterize the graph. An obvious alternative to "two graphs, one rebuild": a single graph shape that serves both prefill and decode — e.g. always build for [n_ctx] token rows and mask off the unused ones, or mutate the n_out tail in place between phases. Both lose. A [n_ctx]-wide graph would run every activation at the full context width (4096 rows of scratch instead of 1 for decode) and forfeit the G3 tail optimization. In-place mutation of shapes would make the graph stateful — the exact property the reuse design is trying to eliminate — and would need its own invalidation logic, i.e. a second, ad-hoc cache protocol. One params-derived rebuild per phase boundary is measured in microseconds and keeps every graph immutable and self-describing.

The allocator lives in the cache, not in the forward call. The GraphCache struct holds graph and alloc side by side GraphCache (cache.rs:38-45) precisely so that one can be replaced without the other. If the allocator were created per forward (or per rebuild), every rebuild would zero the KV cache and chat would forget itself at every turn boundary; if it were a global keyed by nothing, two concurrent sessions (server mode, OPENAI-CHAT-API-PLAN.md) would share and corrupt each other's regions. Hence the three ownership tiers that exist in the tree: the CLI's plain mode uses a process-wide static cache (graph_cache(), src/models/qwen2/graph.rs:1020-1025 — one run, one session); the conversation engine holds a session-private cache (conversation.rs:113-117, the GraphEngine cache field); the server hands each slot its own cache via forward_graph_cached (models/mod.rs:288-298). Same mechanism, scoped ownership.

Positions as data — the founding decision. The plan's first revision note tells the story of the alternative (docs/COMPUTE-GRAPH-DESIGN.md revision note 1): an early design encoded n_past into the KV-cache op, meaning the decode topology changed every step and reuse was structurally impossible. The fix was to make the KV a persistent external buffer addressed by position, with the position injected through an input node — the design llama.cpp's allow_reuse also relies on. Everything else in this document (params-only comparison, KV-survives-rebuild, append-only sessions) is downstream of that one move. It also explains the odd-looking window derivation in §2.2: there is no "cache length" field anywhere in the engine, because the query window is data: it is decoded at execution time (decode_window, cpu_backend.rs:683).

Sample-first loop shape. The loop samples from logits that already exist and runs its forward at the end of the body. The alternative — forward first, sample after — would need a dummy forward before the loop (what would it compute?) or a special first iteration. Sampling first also gives the stop gates a natural home before the forward: a stop decision costs zero forward time. The one asymmetry (§3.4) is the discarded final forward when the run ends via -n rather than via a stop gate.

KV regions sized up front, not grown per step. ensure_kv allocates nkt × n_ctx f32 on first touch — the full context window — even though prefill may use 30 slots. Growing per step (realloc at slot 129, 257, …) would mean periodic huge copies and, on GPU, buffer re-creation; sizing once at n_ctx means a step's KV write is a plain copy_from_slice into existing memory (cpu_backend.rs:377-378, the KV copy_from_slice). The cost is bounded by --n-ctx, which the CLI deliberately decouples from the model's max_seq_len (main.rs:1448-1451 cites the multi-GB over-allocation and first-submit Metal tax this avoids — docs/PERF-QWEN3-4B-VS-LLAMACPP.md §2). The cursor-bounded window (§2.2) is what makes the slack harmless: unread region contents beyond it are never touched by attention.

Two stop policies for EOG, not one. §3.2's conversation excerpt shows EOG being written to the KV before breaking, while the CLI's loop drops it. The difference is not inconsistency but the invariant each loop must keep. The CLI's obligation ends at the printed text; the session's obligation extends to a KV mirror that the next turn will rely on — and the next turn's canonical template render contains the end-of-message marker. Breaking the mirror would surface later as subtly wrong context (the model would effectively "see" a conversation missing its turn boundaries). llama.cpp makes the same choice — the comment that cites it is at conversation.rs:1331-1332.

3.4 Pitfalls & invariants

1. A position past n_ctx is corruption, so every layer of the stack guards it. The cursor is incremented per step with no loop-level ceiling in single-shot mode (only -n bounds the run), so the defense is layered: forward_cached asserts max(positions) < n_ctx before anything runs (src/models/qwen2/graph.rs:513-520 — "fail loudly instead of corrupting memory"); the CPU KV-store kernel bounds-checks each slot and returns Err — never a silent clamp (cpu_backend.rs:358-360, the p >= n_ctx refusal); the conversation loop checks the cursor before sampling and stops the turn cleanly (conversation.rs:1314-1318, the cursor guard). Three guards, one rule: an out-of-range KV write would overwrite another position's cached vector — the resulting nonsense would look like a model-quality bug, which is why it must crash instead.

2. In-place ops and GPU-pending buffers: never host-copy a KV region's producer mid-flight. RoPE and the fused KV-store ops alias their input buffers in place (sole-consumer rule, doc 07); on a GPU backend the "buffer" may be un-submitted device memory. Host-copying it at the wrong moment was the Phase-3 KV-corruption bug (recorded in docs/COMPUTE-GRAPH-DESIGN.md and AGENTS.md rule 5): the copy read stale bytes and wrote them back over freshly stored K/V. The scheduler's split boundaries (the boundary's own close-out wait before any cross-backend read, sync_backend (scheduler.rs:466-471) are the only sanctioned sync points — which is also why the decode loop's reuse discipline matters: no code path between steps touches buffers out-of-band.

3. Liveness and fills must follow build order. alloc_graph computes buffer lifetimes in node-id (build) order, not topological order, because topo_order() may reorder src-less nodes (like KV loads) ahead of nodes the scheduler reads first — the G3 tail regression (alloc.rs:530-536, the build-order comment). Related decode-side rule: input buffers are treated as live for the whole step — last_use (alloc.rs:555-561) — so that a liveness reuse can never clobber token_ids after it was filled but before its consumer ran. On a rebuild, node_to_buf is cleared and re-derived — the input fill in forward_cached happens after alloc_graph has produced the new mapping (src/models/qwen2/graph.rs:641 before src/models/qwen2/graph.rs:655-659), so fills always land in the buffers the scheduler will read.

4. Stale-but-unread KV after rollback. /regen rewinds the cursor without erasing the region (conversation.rs:965-977, the regen_turn rollback), so slots past the cursor hold abandoned tokens. This is safe only because attention reads a window that ends at the cursor: the span input names the cells per head (cpu_backend.rs:1111-1140, the per-head vl window) and the store overwrites slot p on the next append. The invariant "the cursor is the truth; region contents past it are garbage" is what makes rollback O(1) — but it means nothing may read a region by extent (only by position), or it would see the garbage.

5. Fusion toggles are part of the identity — A/B tests must keep the FusionPass. fuse_qkv/fuse_ffn live inside CParams (src/models/qwen2/graph.rs:584-590), so MINFER_NO_FUSE_QKV=1 at any point forces a rebuild with the unfused node set — reliable A/B. The flip side: the fused and unfused graphs must be bit-identical in output, and the unfused path must still run the FusionPass (AGENTS.md rule 7) so that the only difference between the two runs is the QKV fusion itself. The test fuse_flags_are_part_of_the_reuse_identity (cache/tests.rs:36) pins the identity half.

6. Graph uid discipline. A reused graph keeps its uid; a rebuilt one gets the next monotonic value replace_graph (cache.rs:106-114). The CUDA backend's replay cache is keyed by uid — graph_replay (scheduler.rs:286-304, doc 15) — so the pairing is load-bearing: same uid ⇒ same topology ⇒ replay is valid; new uid ⇒ the old capture is orphaned (and simply never requested again). Breaking this — say, reusing a uid across a rebuild — would replay the wrong graph's captured launches.

7. The discarded final forward. When a run ends by reaching -n, the last iteration's forward has already run at the bottom of the body and its logits are dropped by the loop condition. One wasted forward per full-length run (a few milliseconds) buys a loop with a single trivial exit condition. The stop gates that can save it (EOS, stop strings) do, since they break before the forward — so the cost only applies to runs that end by quota, and those are exactly the runs where the model would have kept going anyway.

8. n_ctx is computed once, before prefill, for both phases. ctx = params.n_ctx.max(input_ids.len()) (main.rs:1452-1453, the ctx computation) with the comment "Computed ONCE so prefill and decode size the same KV regions" (main.rs:1452, the "Computed ONCE" comment). If prefill and the first decode step disagreed on ctx, they would disagree on CParams, the params comparison would fail, and — worse — ensure_kv would have sized regions from the prefill graph that the decode graph considers too small. One variable, computed early, keeps the whole run inside one region geometry.

9. The field that looked like cheap insurance — and was not. GraphParams used to carry n_seqs, "so that adding batching later cannot silently reuse a single-sequence graph for a two-sequence call". E2 added the batching and found the reasoning backwards: what a multi-sequence batch changes about the topology is which attention instantiation is built, and that is cparams.explicit_span, decided by the KV reservations. The count itself is data. Keeping it meant a 2-sequence and a 1-sequence batch of the same shape described the same graph and still rebuilt it — measured (uid 3 → 4) by sequence_count_is_data_not_topology in models/qwen2/graph.rs, which now pins the reuse instead. An identity field earns its place only if some topology decision reads it.

4. Observe & verify

The decode loop and the reuse machinery are unusually observable — most of the evidence is on stderr of any plain run.

  • The run header itself. Prefill: N tokens in … vs Generated: M tokens in … (tok/s) (main.rs:1479-1483, 1854-1858, the two rate headers) — two deliberately different calibers (pure-decode rate vs blended rate, main.rs:1851-1853, the pure-decode-vs-blended caliber comment). The gap between the two numbers is this document's §2.2 argument, measured.
  • MINFER_TIMING=1. Splits the per-token wall clock into sampling vs forward milliseconds — t_samp/t_fwd (main.rs:1584-1587, 1800-1804). Expect sampling in the tens of microseconds and forward in the milliseconds — the loop's own overhead is the difference between the two.
  • MINFER_GRAPH_TRACE=1. The scheduler prints the split layout and a per-op/backend census once per execute call (scheduler.rs:164-182, the MINFER_GRAPH_TRACE print) — i.e. once per token on stderr. A CPU-only run shows a single split; a Metal run shows the decode graph's split boundary and the fused-op census. Watching it repeat N times for N tokens is the loop made visible.
  • MINFER_TRACE=<path>. Records a per-node data trace; the CLI marks the prefill/decode phase boundary — begin_phase (main.rs:1459, :1556), tags each decode step with its sampled token (set_token, main.rs:1690-1693), and attaches both the prefill graph and a separate decode graph JSON at the end — export_graph_json (main.rs:1826-1844). Opening the trace in viz/ shows the two topologies side by side — the G3 tail nodes present in one and absent in the other.
  • --dump-graph <dot> / --dump-graph-json <path>. Exits after exporting the graph the runtime would use — labeled "decode" when the prompt is one token and "prefill" otherwise (main.rs:1521-1525, 1545-1548, the kind label). Compare a prefill dump against a decode dump (feed a 1-token prompt) to see exactly which nodes the rebuild adds or removes.
  • MINFER_GRAPH_DUMP=<dir>. Writes logits_decode.f32 plus every layer's kv{ℓ}_decode.f32 per step, overwritten in place (src/models/qwen2/graph.rs:694-750). Each file is the full persistent region (fixed 2 MiB for 0.5B at --n-ctx 4096), so across steps you watch the valid prefix of kv0_decode.f32 grow — 512 B of new K values per step (128 f32) — which is the append-only KV made tangible; the per-step logits dumps let you diff GPU vs CPU decode runs layer by layer.
  • Unit tests. src/graph/cache.rs covers the reuse contract directly: reuse_requires_equal_params (cache/tests.rs:126 — n_tokens / gtype / weights_version changes all force rebuilds), allocator_survives_rebuild (cache/tests.rs:159 — a persistent region registered before a rebuild is still there after), fuse_flags_are_part_of_the_reuse_identity (cache/tests.rs:36), and structural_check_detects_different_graph (cache/tests.rs:176). On the model side, forward_cached_isolates_kv_between_caches forward_cached_isolates_kv_between_caches (src/models/qwen2/graph/tests/batching.rs:742) proves two caches don't share KV. The conversation state machine is tested without a model via a mock engine: second_turn_appends_only_delta (conversation/tests/turns.rs:48) asserts the delta prefill is smaller than the first turn's, and context_fill_during_decode_stops_cleanly (conversation/tests/overflow.rs:22) pins the cursor-vs-n_ctx guard.
  • bench -n N <model>. Runs decode for a fixed N and reports the pure-decode rate in md/csv/json — the reproducible version of the Generated: line, and the tool the plan document's §12 expectations were written against.

A five-minute experiment that exercises the whole document: run once with MINFER_GRAPH_TRACE=1 and a two-word prompt, and count the split-layout lines — prefill once, then one per generated token, all with identical layouts (same graph reused). Then rerun with --stop set to a word early in the output and watch the trace stop mid-stream without a final forward.

5. Cross-references

← 12 — The sampler: from logits to a token · Index · 14 — The Metal backend →

14 · The Metal backend

Stage: decode loop running (docs 09–13) → this stage: the same graph, executed on the Apple GPU instead of the CPU kernels of docs 10/11 → the CUDA backend (doc 15). Code: src/graph/metal_backend.rs (MetalBackend, the Backend trait implementation — supports_op at :265, execute_node at :843, synchronize at :1750), src/metal/ (MpsState device layer, MpsCommandBuffer, submit at :1935), src/metal/kernels/ (the Metal Shading Language kernels).

1. Background — where this stage sits

Docs 05 through 08 built a declarative compute graph — one node per math op of the transformer — assigned every node to a backend, and allocated buffers by liveness. Docs 09 through 11 then followed the CPU execution of that graph: quantized matmuls with AVX2/NEON dot kernels (doc 10), RMSNorm, RoPE, GQA attention over the KV cache (doc 11). Doc 13 showed the decode loop re-running that graph one token at a time. This document is the parallel universe: the identical graph, node for node, executed on an Apple GPU through Metal. Nothing about the graph changes — the builder emits the same nodes, the scheduler walks the same splits — only the thing that runs each node changes. Doc 14 and doc 15 (NVIDIA) are the two chapters of "the same math, a different executor".

Two words this document leans on constantly, defined before anything else:

  • A backend is an executor: a component that owns a pool of buffers and knows how to run every operation of the graph on one piece of hardware. In minfer a backend is whatever implements the Backend trait from src/graph/backend.rs — the CPU backend of doc 08, the Metal backend of this doc, the CUDA backend of doc 15.
  • A kernel is one program that runs on the accelerator: for Metal, a function written in Metal Shading Language (MSL — Apple's C-derived GPU language, the code in src/metal/kernels/) that thousands of GPU threads execute in parallel. "The backend dispatches ops to kernels": MetalBackend::execute_node is the dispatch, kernel_rms_norm_f32 is a kernel.

So the division of labor is: src/graph/metal_backend.rs (~2080 lines) is the graph-facing half — it takes one CNode at a time and decides which kernel to encode. src/metal/ (~3940 lines) is the device-facing half — it owns the Metal device object, the compiled kernels, the weight registry, and the command buffers. src/metal/kernels/ (~5150 lines) is the hardware-facing half — the actual shader programs. One name to un-confuse early: the code calls its device layer MpsState ("MPS" = Metal Performance Shaders, Apple's own math library). That is a historical name — minfer uses none of Apple's MPS library; every kernel is handwritten, following llama.cpp's Metal kernels, because the project has zero ML-framework dependencies by design.

Why does a GPU backend exist at all, and why does macOS get one first? Apple Silicon is a unified-memory machine: the CPU and the GPU physically share one pool of RAM and one memory controller. A traditional discrete GPU needs every weight and every activation explicitly copied over PCIe (doc 15's pinned-staging story); on an M-series Mac the GPU can be handed a pointer to memory the CPU already filled. That removes the classic tax of GPU inference — the copy — and leaves the GPU's real advantage: thousands of arithmetic lanes that can chew through a quantized matmul while the CPU's handful of cores (doc 10's thread pool) stream the same bytes far more slowly. The measured stakes, from docs/METAL_OPTIMIZATIONS.md §0.1: Qwen2.5-0.5B decodes at ~306 tokens/s on Metal versus the CPU's 149.7 tokens/s on the same machine (the older ~5.9 figure is not reproduced — a throttled box, and §0.1 says so explicitly); the 7B model runs at ~19.3 ms/token — llama.cpp parity.

What would break without this stage? Nothing crashes — the engine would simply always take the CPU path. But the design must answer a harder question: how do you plug a second executor into a graph that was built for the CPU without introducing hangs, silent wrong answers, or GPU faults that freeze the whole machine? That last failure mode is real: docs/GPU_SAFETY.md opens with an incident where a GPU hang froze an entire M4 Pro (WindowServer included, forced shutdown). So the Metal backend is built around four contracts, all verified in this doc: ops are only assigned to it when it can run them (supports_op), every weight it needs is registered before the graph builds (the all-weights-registered gate), its async GPU work is flushed exactly at split boundaries (synchronize), and every wait on the GPU is bounded (submit).

The walk: §2 covers the principles — the trait contract, unified memory, the f32-activations decision, the one-command-buffer-per-split rule, the dispatch matrix, and when Metal actually wins. §3 walks the code: the trait implementation, the device layer, zero-copy weight registration, the gate, the bounded submit, and one full shader. §4 shows how to watch it run; §5 points onward.

2. Principle — how it works and why

2.1 The contract: what a backend must provide

The Backend trait (src/graph/backend.rs) is the seam between the graph machinery (docs 05–08, which know nothing about GPUs) and the hardware executors. Seven methods, each with one job:

#![allow(unused)]
fn main() {
// src/graph/backend.rs:21-45 (contract core; read_host/write_host/synchronize at :59-70)
pub trait Backend: Send + Sync {
    fn name(&self) -> &str;

    /// Op support by (op, dtype). `supports_fused` gates the fusion pass
    /// (Phase 4) so fused IR nodes are only produced when a kernel exists.
    fn supports_op(&self, op: &Op, dtype: DType) -> bool;
    fn supports_fused(&self, fused: &FusedOp) -> bool;

    /// Buffer pool: allocate / release a buffer of `size` f32 elements.
    fn alloc_buffer(&mut self, size: usize) -> usize;
    fn free_buffer(&mut self, id: usize);

    /// Allocate a buffer that bypasses the recycle free list. Split-boundary
    /// staging needs this: at execute time the free list holds ids whose
    /// physical contents are still referenced by node_to_buf and get
    /// read/written later in the same execute — recycling one would clobber
    /// in-flight data. Fresh buffers enter the normal free list on
    /// free_buffer (at graph rebuild), where liveness recycling is safe.
    fn alloc_fresh(&mut self, size: usize) -> usize;
}

Read the method list as a small operating system for one accelerator: capability queries (supports_op, supports_fused — "can you run this op at all?") are asked at build time, before any execution; memory management (alloc_buffer/free_buffer/alloc_fresh) is asked by the allocator of doc 07, once per graph build; execution (execute_node) is asked by the scheduler of doc 08, once per node per forward; and data movement (read_host/write_host — host↔buffer access for filling inputs and reading logits) plus synchronize (flush async work) are asked at split boundaries. The CPU backend implements all of these trivially — its buffers are Vec<f32>s, execute_node just calls the function, synchronize is a no-op. The Metal backend implements the same signatures with an asynchronous accelerator behind them, which is where every interesting design decision in this chapter comes from.

Two contract details deserve a second look because the Metal backend leans on them hard. First, execute_node receives buffer ids into the backend's own pool, not raw pointers, and takes a kv_pair argument — the (k_id, v_id) ids of the layer's two persistent KV regions — because KV ops (store, attention) need to touch regions the allocator keeps alive across the whole session. Second, the trait doc for execute_node warns that the output buffer may alias an input buffer (liveness reuse, doc 07) — the backend must handle in-place execution. On the CPU that is automatic (same memory); on Metal it becomes a visible piece of code you will meet in §3.2 (copy_in).

2.2 Unified memory changes the copy math

On a discrete GPU the data story is: allocate device memory, copy weights in once, copy activations back and forth every step over PCIe. On Apple Silicon, every buffer can be one allocation both sides can see. minfer leans on this in three places:

  1. Weights are never copied at all. The GGUF file is memory-mapped (doc 02); at load time minfer wraps the mapped bytes in a Metal buffer with newBufferWithBytesNoCopy (§3.2.7) and registers each weight as (buffer, byte-offset) into that one wrapper. The GPU reads the model directly out of the file page cache. A 7B model's ~4.4 GB of weights (doc 10's number) occupy the same physical pages for the CPU and the GPU.
  2. The activation pool is shared-memory. new_f32_buffer allocates with MTLResourceOptions::StorageModeShared — host and GPU address the same bytes. Filling an input is a plain memcpy into the buffer; reading logits is a plain read out of it. read_host/write_host are therefore views, not transfers.
  3. Cross-backend copies are host round trips. When the scheduler must move a value from a Metal split to a CPU split (doc 08), copy_across does read_host → write_host — two memcpys, no DMA engine, no staging queue. src/graph/alloc.rs:565 documents exactly this: "host round trip through read_host/write_host (shared-memory GPU buffers make this a plain memcpy both ways)".

The consequence for the reader of docs 14 vs 15: this chapter has almost no data movement code, because on this hardware there is almost no data movement. The CUDA chapter is, in large part, the story of managing exactly the copies that unified memory makes unnecessary.

2.3 Why f32 activations on the GPU when the CPU quantizes to Q8_0

Doc 10 established the CPU's inner-loop trick: activations are quantized to Q8_0 (per 32 values, one f16 scale + 32 int8 values) right before each matmul, so the inner loop becomes an int8×int8 dot product that the CPU has dedicated silicon for (vpmaddubsw on x86, SDOT on ARM). The Metal path does none of that: its shaders read the activation buffers as plain f32. Three reasons, in decreasing order of importance:

  • There is no int8 hardware to exploit. MSL exposes no fast int8×int8 matrix instruction to shaders comparable to SDOT — Apple's integer muscle lives in the Neural Engine, which Metal Shading Language cannot reach. An int8 activation path on Metal would add a quantize kernel before and a dequantize inside the shader, and the inner loop would still be float math converted from ints. llama.cpp's Metal backend makes the same choice (minfer's METAL_OPTIMIZATIONS.md #4 records adopting it: "Q4_0 → f32 activations (aligned with llama Metal, removed Q8_0 quantize) — decode +5-10 %" — removing the quantize step made it faster).
  • Decode bandwidth is bound by weights, not activations. During decode (doc 13) one token's activation row is read once per matmul, while the entire weight matrix is streamed per matmul. Arithmetic for one layer of Qwen2.5-0.5B (attn_q, an [896, 896] Q4_0 weight): weights = 28 blocks × 18 B × 896 rows ≈ 451 KB; activations = 1 token × 896 × 4 B ≈ 3.5 KB — the weights are 99.2 % of the traffic. Quantizing activations to 1 byte would shrink the 3.5 KB to 0.9 KB: a ~0.6 % saving on a stream that is memory-bound. The weights stay 4-bit (§3.2: the same GGUF bytes as the CPU reads); only the activation side differs.
  • The two paths are allowed to disagree. minfer's rule (AGENTS.md, core rule 9): "CPU quantizes activations to Q8_0, GPU reads f32 — CPU-vs-GPU logits differ by design; compare each path against its own reference." The CPU path is verified bit-for-bit against llama.cpp (doc 10); the Metal path is verified against its own f32 reference within a small tolerance (§4's tests use max diff < 1e-3). Greedy token outputs still match in practice across backends, but the guarantee lives per path, and any test that quietly compared CPU logits to Metal logits would be testing the wrong invariant.

2.4 One command buffer per split — the encode/submit rhythm

A command buffer is Metal's unit of work submission: you encode into it a list of kernel launches (a compute command encoder wraps them), then commit it to the GPU, which executes the launches in order. Encoding is cheap CPU-side bookkeeping; the expensive part is the commit — the driver round trip. That asymmetry shapes minfer's execution model:

  • execute_node never talks to the GPU. It encodes one kernel launch (or a few) into the split's command buffer and returns immediately. A whole decode forward — ~50 nodes for one layer loop — costs the CPU only the encode time.
  • The command buffer lives for exactly one split (doc 08: a maximal run of consecutive nodes assigned to the same backend). At the split's end the scheduler calls synchronize, which submits — commits the buffer and waits for the GPU to finish.

Why per split and not per op, or per whole graph? Per op would pay the commit round trip dozens of times per forward for no benefit — nothing between two ops in one split needs the GPU's results. Per whole graph would be wrong the moment the graph has a CPU node: the CPU node downstream needs its input now, so the GPU work feeding it must be flushed at the boundary — and the split boundaries are precisely where the scheduler already stops and syncs (scheduler.rs:177-188, shown in §3.2). On a fully-GPU-eligible model the graph is one Metal split, so one submit per forward; a mixed graph pays one submit per Metal split. The metal_backend.rs module header states the rule: "One MpsCommandBuffer is kept per split and submitted by synchronize() (called at split boundaries), so ops within a split share a single GPU submission — the plan's §15 'split shares one command buffer' rule."

One more subtlety hides inside the encoder: within a split, kernels run back-to-back on the GPU, and Metal does not guarantee that one kernel's writes are visible to the next kernel without an explicit barrier. Every dispatch helper in src/metal/ therefore ends with enc.memoryBarrierWithScope(MTLBarrierScope::Buffers) (src/metal/). This is not paranoia — METAL_OPTIMIZATIONS.md #28 records the bug it fixed: on 1.5B/7B prefill, the RMSNorm output buffer bn was reused as the input of the next op, and "a kernel that reads a buffer written by a preceding dispatch can race with that dispatch's last threadgroups, intermittently corrupting the tail rows" — last-2-token garbage, roughly 10-30 % wrong tokens on the first token, fixed by inserting the barrier (and a threadgroup_barrier in the GEMM kernels), after which 24/24 dump comparisons became deterministic.

2.5 The dispatch matrix: which op runs which kernel

supports_op (metal_backend.rs:675-702) is the capability table, and it is nearly a yes for everything the Qwen2/Qwen3 builders emit: elementwise ops (Add/Mul/Silu/SwiGLU), norms (RmsNorm, Qwen3's per-head QkNorm), MatMul in every supported quant type, GetRows (embedding gather and the tail-row gather), RoPE, KvcacheStore/KvcacheLoad, Attn, and the decode fusions FusedQKV/FusedFFN/FusedQkvNorm. It refuses Scale, Softmax and BatchMatMul — none of which the model builders emit in the graph path (attention is one fused node that does its own softmax; the attention scale rides in AttnMeta; BatchMatMul is a deferred vocabulary entry) — and it refuses QkvBiasRopeStore, a CUDA-only fusion (doc 15). Refusal here is not an error: it simply makes assign_backends (doc 06) hand such a node to the CPU, creating a split boundary.

Within the ops Metal does run, the dispatch is a decision tree of kernel families — this is where docs/METAL_OPTIMIZATIONS.md's campaign lives, so here is the map (the doc has the measurements):

  • Matmul, three tiers (quant_matmul_f32_on_gpu_buf, src/metal/ops.rs): a simdgroup GEMM kernel (64×32 output tile per threadgroup, 8 KB of threadgroup scratch) when prefill-shaped (nt ≥ 2 and big enough), a _multi kernel (one threadgroup per 2 output rows, handles all nt at once) when multi-token but small, and a single-token kernel for decode (nt == 1). Every supported quant type (Q4_0 through Q6_K) has all three tiers where it matters; the q4_K decode kernel was ported to llama.cpp's layout in #27 and took 7B decode from ~51 to ~19.3 ms/token.
  • Attention, five variants (the Op::Attn arm, §3.2.5): decode (nt == 1) picks the flash kernel (llama's kernel_flash_attn_ext port, online softmax, hd 64 or 128) or the two-pass split KV-parallel kernel, else the classic kernel; prefill picks the prefill-flash port, else the 3-pass parallel implementation (scores → masked softmax → output, fully barrier-free), else classic. The CPU (doc 11) computes full attention rows because everything fits in cache; the GPU variants exist because a 32-lane simdgroup cannot hold a row and the fix is to tile and accumulate online — the "flash attention" trick doc 11 §2.5 previewed.
  • RMSNorm, two widths: the 32-thread single-simdgroup kernel and the 256-thread multi-simdgroup one (default, ~2× faster per dispatch — §3.2.8 walks the shader).
  • KV cache, three element widths: MINFER_CACHE_TYPE=f16 stores K/V as half floats (auto-on for the 7B class, src/metal/policy.rs), halving attention memory bandwidth (#13 measured f16 decode 1.60 → 0.95 s on a long-context case). The region allocation stays f32-shaped — the IR dtype is F32 — and the f16 kernels use the first half of the bytes as half storage; the win is bandwidth, not footprint. The third is a packed q8_0 cache (#310): crate::metal::packed_attn_route picks mechanism A, a native packed read in the decode flash family, or mechanism B, which dequantizes the needed window into a transient arena-addressed f32 stage and runs the unchanged f32 prefill / windowed-flash family.
  • Decode fusions (docs 05/06 introduced them): FusedQKV is one concat matmul (blk.{i}.attn_qkv, built at load with concat_rows) + one attn_bias_rope_store kernel replacing 3 matmuls + 3 biases + 2 RoPEs + 2 stores (10 dispatches → 2); FusedFFN is one concat matmul (blk.{i}.ffn_gu) + an in-place swiglu (4 → 2), gated nf ≤ 16384 because on the 7B the concat matmul measured slower than two separate ones; FusedQkvNorm (Qwen3) adds the per-head Q/K RMSNorms between the matmul and the rope+store.

2.6 When Metal wins — and by how much

The physics from doc 10 carries over unchanged: decode is weight streaming (one token ⇒ nothing to amortize; every matmul reads its whole weight matrix), prefill is compute-bound (many tokens share each weight byte). Metal wins both races on an M-series machine, for different reasons:

  • Decode on unified memory: the ~4.4 GB of 7B weights stream from the same LPDDR the CPU would use, but thousands of GPU lanes keep the memory controller saturated in a way 8–12 CPU cores cannot. METAL_OPTIMIZATIONS.md §0.1: 7B Q4_K_M decode ~49 t/s on the graph path (steady ~19.3 ms/token after #27 — 4.4 GB / 19.3 ms ≈ 228 GB/s of effective weight streaming), Qwen3-4B ~74.1 t/s, 0.5B ~306 t/s — versus the CPU's 149.7 t/s on the 0.5B (§0.1's re-measurement; the pre-round ~5.9 is not reproduced, and doc 10's table's ~18 tok/s 7B ceiling is bandwidth arithmetic, not a measurement).
  • Prefill on the GPU: doc 10's quantized int8 kernels are genuinely fast, but a GEMM-shaped workload is exactly what simdgroup matrix hardware is for; the graph path reached ~3900–4000 t/s prefill on the 0.5B (+~55 % over minfer's own pre-graph path, llama-parity on Qwen3-4B per docs/PERF-QWEN3-4B-VS-LLAMACPP.md), with the remaining gap on 7B prefill ~−10 % (METAL_OPTIMIZATIONS.md §0.1 reading).

When does Metal not win? The doc records the counterexamples as guardrails: the 7B FFN fusion is disabled (nf ≤ 16384 gate) because the concat matmul measured slower on the 7B's decode scalar kernel; the 0.5B f16 KV cache measured ~3 % slower (dispatch-latency-bound, so auto-f16 is reserved for the 7B class); and the very first GPU touch of file-backed pages costs ~44 ms of one-time page/TLB setup — which is why register_part runs a warmup read at load (§3.2.7). "GPU faster" is a per-op, per-shape, measured claim here, never an assumption.

2.7 GPU participation is a build-time decision — priority and the gate

Three rules decide whether any node runs on Metal, all resolved before the graph exists:

  1. Priority order Metal → CUDA → CPU. GraphAllocator::supports (alloc.rs:384-385) asks backends in that fixed order and returns the first yes. If Metal is enabled, it wins every op it supports; CUDA (when compiled in) gets the leftovers it can run; CPU mops up the rest. Doc 15's backend joins the same queue.
  2. Metal is enabled only if the whole model made it to the GPU. The gate (models/qwen2/graph.rs:636, §3.2.6) requires every weight the graph reads — embeddings, all 28+ layer weights and biases, norms, lm_head — to be registered in the Metal registry (has_weight). One missing tensor ⇒ metal_on = false ⇒ the entire graph runs on CPU. This is the all-weights-registered gate: all-or-nothing participation, never a partial run where some layers are on the GPU and some on the CPU (which would be correct-but-mysterious; a split boundary per layer would also hammer the sync path).
  3. Participation is part of graph identity. The gate's outcome is recorded in CParams.gpu (graph/params.rs:27), and GraphParams is the only thing graph reuse compares (doc 13). So "was the GPU on" is baked into the cached graph: a change (MPS unavailable, weights not registered) changes CParams.gpu, fails try_reuse, and forces a rebuild with the new assignment — rather than silently reusing a graph whose backend assignments no longer hold.

And the failure posture, which the repo treats as a hard rule: kernel-invariant violations return Err from execute_node — never a silent CPU fallback. If a weight is somehow not registered when a matmul executes, the arm returns Err("weight '...' not on GPU") and the run aborts with that message; if an op reaches the Metal arm that supports_op rejected, the arm returns Err too (metal_backend.rs:1092-1100). The reasoning (GPU_SAFETY.md §2.3): a silent fallback masks a broken invariant — the model would keep producing text while the user has no idea half the graph silently ran somewhere else, or that a shape assumption was violated. Loud failure at the offending node is debuggable; quiet degradation is not. (Frontier guards that sit below the graph layer — dispatch-time checks like gpu_abort for device-limit overruns — print the actual values and exit, for the same reason.)

3. Implementation

3.1 Data in / data out

The data story at this stage is the doc 07 allocator's, seen from the GPU side. Everything the backend touches is an id into its own pool of shared-memory MTLBuffers:

ItemShape / layoutWhere it comes from, where it goes
Pool buffersize f32 elements = size × 4 bytes, StorageModeSharedalloc_buffer/alloc_fresh → MpsState::new_f32_buffer; host fills inputs via write_host, reads outputs via read_host
Weightsraw GGUF quantized bytes, [out][in] row-major, registered by name as (MetalBuffer, u64 offset)register_weight at load (zero-copy into the mmap'd part wrapper); looked up per node via state.weight_buf(name)
Activations x[nt][id] f32, token-major — never quantized on the GPU pathprevious node's output buffer; the matmul shader reads them directly
Positions[nt] u32 values stored as f32::from_bits bit patterns (doc 07's fill_input_i32)input buffer; kernels decode with float_to_int semantics, host helpers read them as u32
KV regionsper layer two persistent buffers [n_kv_embd, n_ctx] (f32 elements; f16 mode uses the first half of the bytes)alloc_persistent at first touch (doc 07); written by KvcacheStore/fused kernels at positions[t], read by attention
Logits[n_out][n_vocab] f32, token-majorthe graph's output buffer; copy_to_cpu at the end of the forward (doc 09's extraction)

Two entries deserve a beginner's double-take. Positions as f32 bit patterns: the graph's IR has one dtype (DType::F32) for buffers, so integer token ids and positions are smuggled through as raw bit patterns and converted back with .to_bits() wherever the backend needs the integer (positions_max at metal_backend.rs:232-237 reads them as u32 directly, which is the same bytes). The f16 KV region: the graph shapes the region as [n_kv_embd, n_ctx] F32 no matter what, so the allocation is unchanged; when kv_cache_is_f16() is true the kernels (kernel_store_kv_f16, kernel_gqa_attn_f16, the f16 flash variants) treat that memory as half* — 2 bytes per element in the first half of the allocation. The saving is bandwidth during attention (half the bytes per read), which is what the decode loop is bound by.

And one negative entry, because it is the headline difference from doc 10: there is no Q8_0 activation scratch anywhere on the GPU path. The CPU quantizes activations per matmul call (doc 10 §3.2); the Metal path's activations stay f32 from embedding lookup to logits.

3.2 Key code

3.2.1 The backend object: pool, command buffer, and the 'static trick

MetalBackend (metal_backend.rs:104) is a thin shell over the device singleton plus its own pool:

#![allow(unused)]
fn main() {
// src/graph/metal_backend.rs:85-106
pub struct MetalBackend {
    state: &'static MpsState,
    /// f32-element pool: id → shared MTLBuffer (size * 4 bytes)
    pool: Vec<crate::metal::MetalBuffer>,
    free: Vec<usize>,
    /// P2/P3 capture staging: host-readable buffers written by per-split blits.
    /// Only allocated while a trace/live capture is armed.
    staging: Vec<crate::metal::MetalBuffer>,
    free_staging: Vec<usize>,
    /// Pending command buffer for the current split. Stored as a leaked box
    /// pointer (null = none) because MpsCommandBuffer is !Send/!Sync; all
    /// access happens sequentially through &self/&mut self methods on the
    /// scheduler thread, so the raw pointer is contained.
    cb_ptr: *mut crate::metal::MpsCommandBuffer<'static>,
}

// Safety: every field is either owned (pool/free), a 'static reference
// (MpsState is a Sync singleton), or the command-buffer pointer which is only
// dereferenced inside &self/&mut self methods (sequential, single-threaded).
unsafe impl Send for MetalBackend {}
unsafe impl Sync for MetalBackend {}
}

Three fields carry the design. state is a &'static MpsState — the device layer is a process-wide singleton (static MPS: OnceLock<Option<MpsState>>, src/metal/), created once by MpsState::init() from main.rs:639 before the model loads; every backend instance just borrows it. pool/free implement the Backend buffer-pool contract: pool[id] is a shared MTLBuffer, and free is the recycle list the allocator's liveness analysis drives (doc 07). cb_ptr is the oddest: the in-flight command buffer is stored as a raw pointer to a leaked Box, because the objc2 command-buffer type is !Send/!Sync and Rust's borrow checker would otherwise forbid the pattern create the buffer on first op → keep appending on later ops → submit at the split end while also letting cb() hand out a &'static mut that does not borrow self (so encode methods can still touch the pool). The comment block is the safety argument: all access is sequential through &mut self on the single scheduler thread. This is also why unsafe impl Send/Sync appears — Metal objects are thread-safe in reality, but objc2 cannot prove it, so the type asserts it with the reasoning written out (src/metal/ says the same for MpsState).

The command buffer is created lazily on the first op of a split:

#![allow(unused)]
fn main() {
// src/graph/metal_backend.rs:175-186
    /// The current split's command buffer (created on first op of a split).
    /// The box is leaked, so the returned reference is 'static and does not
    /// borrow `self` — callers can freely touch the pool afterwards.
    fn cb(&mut self) -> &'static mut crate::metal::MpsCommandBuffer<'static> {
        if self.cb_ptr.is_null() {
            let cb = Box::new(self.state.cmd_buffer());
            self.cb_ptr = Box::into_raw(cb);
        }
        // SAFETY: cb_ptr is null or points to a live box created here; all
        // callers hold &mut self, so no concurrent mutation.
        unsafe { &mut *self.cb_ptr }
    }
}

MpsState::cmd_buffer() (src/metal/runtime.rs) asks the device's command queue for a fresh MTLCommandBuffer and immediately opens a compute command encoder on it — so from the first execute_node of the split, every dispatch lands in one encoder. The Drop impl (metal_backend.rs:274-287) flushes any never-submitted buffer so a dropped backend cannot leak an open encoder.

3.2.2 Capability table: supports_op and supports_fused

#![allow(unused)]
fn main() {
// src/graph/metal_backend.rs:294-319 (abridged to the shape; full match in tree)
    fn supports_op(&self, op: &Op, dtype: DType) -> bool {
        match op {
            Op::Input => true,
            Op::Add | Op::Mul | Op::Silu | Op::RmsNorm { .. } | Op::QkNorm { .. } | Op::SwiGLU => {
                dtype == DType::F32
            }
            Op::MatMul { .. } => {
                matches!(dtype, DType::F32) // activations are f32; weight type in meta
            }
            Op::GetRows | Op::RoPE { .. } | Op::Attn { .. } => dtype == DType::F32,
            Op::KvcacheStore { .. } | Op::KvcacheLoad { .. } => dtype == DType::F32,
            Op::FusedQKV { .. } | Op::FusedQkvNorm { .. } | Op::FusedFFN => dtype == DType::F32,
            Op::View { .. } | Op::Reshape { .. } | Op::Permute { .. } => true,
            Op::Scale(_) | Op::Softmax { .. } | Op::BatchMatMul => false,
            // Mixed-quant decode QKV epilogue (D3-8 class 2) is CUDA-only; on
            // Metal the graph builder never emits it (qkv_epilogue_ok = false
            // without `--features cuda`), so it is never assigned here.
            Op::QkvBiasRopeStore { .. } => false,
        }
    }

    fn supports_fused(&self, fused: &FusedOp) -> bool {
        // swiglu_f32 is the only fusion-pass kernel. The bias+rope+store
        // capability is a build-time fused node (FusedQKV/FusedQkvNorm), not a
        // FusionPass target, so it is not advertised here.
        matches!(fused, FusedOp::SwiGLU)
    }
}

Two things to notice. The dtype checks are almost tautological — the graph's node outputs are DType::F32 everywhere — but the weight type is not part of this signature; it lives in NodeMeta::MatMul.weight_ttype and is checked at dispatch time against the shader families that exist (§3.2.4). And the two false arms are informative negatives: QkvBiasRopeStore is refused with a comment explaining that the builder never emits it on macOS — the support table and the builder's emission rules are kept in lockstep by that comment, and doc 15's backend claims the same op because CUDA can fuse it. supports_fused similarly gates the fusion pass (doc 06): only SwiGLU is claimed (its kernel exists, and the fusion pass has no other rule), so a Mul(Silu(x), y) chain gets rewritten while everything else is left alone.

3.2.3 The buffer pool: recycle vs fresh

#![allow(unused)]
fn main() {
// src/graph/metal_backend.rs:321-343
    fn alloc_buffer(&mut self, size: usize) -> usize {
        if let Some(idx) = self
            .free
            .iter()
            .position(|&id| self.pool[id].length() as usize == size * 4)
        {
            return self.free.swap_remove(idx);
        }
        self.pool.push(self.state.new_f32_buffer(size));
        self.pool.len() - 1
    }

    fn free_buffer(&mut self, id: usize) {
        if !self.free.contains(&id) {
            self.free.push(id);
        }
    }

    fn alloc_fresh(&mut self, size: usize) -> usize {
        // never recycled from the free list (see Backend::alloc_fresh)
        self.pool.push(self.state.new_f32_buffer(size));
        self.pool.len() - 1
    }
}

The recycle logic is an exact-size free-list search: reuse any dead buffer whose byte length matches size × 4, else allocate a new shared buffer. Since E4 S2 the allocator passes the node's size class as size (so buffers of one class are interchangeable), and a host write may address a window of a longer buffer (write_host_window); write_host keeps the exact-length contract for the persistent KV regions and staging. The reason alloc_fresh exists is written in the trait (§2.1): split-boundary staging buffers must not be recycled mid-execute, because the free list at that moment holds ids whose contents are still in flight (encoded kernels will write them later this same execute). Fresh buffers join the normal free list only when the whole graph is freed at rebuild. Note what is not here: no reference counting, no GPU-side allocator — the allocator of doc 07 does the liveness math and calls these methods; the backend just owns the bytes. Underneath, new_f32_buffer (src/metal/runtime.rs) is one call: device.newBufferWithLength_options(size * 4, StorageModeShared) — the unified-memory allocation that makes read_host/write_host plain views.

3.2.4 execute_node: the dispatch arms

execute_node (metal_backend.rs:925-1967) is a 782-line match &node.op (the match opens at :941), and every arm follows the same rhythm: resolve metadata → look up weights by name → encode one or two kernel launches into cb → Ok(()). Nothing waits; the submit happens at the split boundary. The MatMul arm is the cleanest example:

#![allow(unused)]
fn main() {
// src/graph/metal_backend.rs:561-590
            Op::MatMul { .. } => {
                let meta = match &node.meta {
                    NodeMeta::MatMul(m) => m,
                    other => return Err(format!("matmul node missing MatMulMeta: {other:?}")),
                };

                let (wb, w_off) = self
                    .state
                    .weight_buf(&meta.weight_name)
                    .ok_or_else(|| format!("weight '{}' not on GPU", meta.weight_name))?;
                let nt = node.out_shape[1];
                cb.quant_matmul_f32_on_gpu_buf(
                    &wb,
                    w_off,
                    meta.weight_ttype,
                    self.buf(in_bufs[0]),
                    0,
                    self.buf(out_buf),
                    meta.out_dim,
                    meta.in_dim,
                    nt,
                );
                if let Some(bname) = &meta.bias_name {
                    let (bb, b_off) = self
                        .state
                        .weight_buf(bname)
                        .ok_or_else(|| format!("bias '{}' not on GPU", bname))?;
                    cb.add_bias_f32(self.buf(out_buf), &bb, b_off, meta.out_dim, nt, 0);
                }
                Ok(())
            }
}

Every line here teaches the backend's shape. The weight arrives as (buffer, byte offset) from the registry — the offset is the zero-copy trick of §3.2.7, letting one buffer hold the whole mmap'd model with per-tensor offsets. The failure mode is the gate leaking: if a weight somehow is not registered, the arm returns Err("weight '...' not on GPU") — the run stops with the tensor's name, it does not fall back (§2.7). And the bias is a second encode into the same command buffer: matmul then add_bias_f32, two kernels back-to-back with the encoder's memory barrier between them, both paid only when the split submits.

The kernel side, quant_matmul_f32_on_gpu_buf (src/metal/ops.rs), is the three-tier dispatch of §2.5. Its head shows the tier decision and a GPU-safety guard:

#![allow(unused)]
fn main() {
// src/metal/ops.rs (Q8_0 arm of the dispatch; guard + GEMM tier selection)
        if matches!(
            ttype,
            TensorType::Q4_K | TensorType::Q5_K | TensorType::Q6_K
        ) && id % 256 != 0
        {
            gpu_abort(&format!(
                "matmul input dim id={id} is not 256-aligned for {ttype:?} (K-quant kernels use K/256 super-block floor)"
            ));
        }
        match ttype {
            TensorType::Q8_0 => {
                if nt >= 2 && (od >= 2048 || nt >= 9) && Self::gemm_enabled() {
                    self.gemm_dispatch(
                        &self.state.pl_q8_0_mm_f32,
                        wb,
                        w_off,
                        x,
                        x_off,
                        out,
                        od,
                        id,
                        nt,
                    );
                } else {
                    self.enc.setComputePipelineState(
                        &**(if nt > 1 {
                            &self.state.pl_q8_0_f32_multi
                        } else {
                            &self.state.pl_q8_0_f32
                        }),
                    );
}

The guard is audit finding M1 from GPU_SAFETY.md §3: the K-quant shaders compute super-block counts as K/256 (integer floor), so a non-256-aligned id would silently drop the tail elements — wrong numbers, no fault. Rather than risk it, gpu_abort (src/metal/) prints the actual misaligned value plus the hint "(force CPU with MINFER_DISABLE_MPS=1)" and exits. The tier rule nt >= 2 && (od >= 2048 || nt >= 9) sends prefill-shaped work to the simdgroup GEMM (a 64×32 output tile per threadgroup amortizes the weight load across 2048 MACs) and everything else to the _multi/single kernels. Each per-type arm repeats this skeleton — buffer 0 = weights (with offset), buffer 1 = activations, buffer 2 = output, bytes 3 = [od, id, nt] params — with per-kernel threadgroup shapes tuned in the optimization campaign (#27's q4_K port is the headline: METAL_OPTIMIZATIONS.md records 7B attn_q going 70 → 265 GB/s effective).

The attention arm (metal_backend.rs:684-836) is the largest, and its opening is the safety story:

#![allow(unused)]
fn main() {
// src/graph/metal_backend.rs:689-709 (guards; dispatch decision follows)
                // GPU safety (H1): kernel_gqa_attn strides KV by nk*hd
                if meta.nkt != meta.n_head_kv * meta.hd {
                    return Err(format!(
                        "Metal attention: nkt={} != n_head_kv*hd={} (kernel_gqa_attn strides KV by nk*hd)",
                        meta.nkt, meta.n_head_kv * meta.hd
                    ));
                }
                if meta.hd != meta.hd_kv {
                    return Err(format!(
                        "Metal attention: hd={} != hd_kv={} (kernel_gqa_attn uses query head dim)",
                        meta.hd, meta.hd_kv
                    ));
                }
                let (k_id, v_id) = kv_pair
                    .ok_or_else(|| format!("KV regions for layer {} not allocated", meta.layer))?;
                let nt = node.out_shape[1];
                // G1: dispatch the fast attention kernels (flash / split /
                // parallel) exactly like the legacy layer_gpu path. The fast
                // paths are gated to the isolation-tested shapes (hd 64/128);
                // anything else falls back to the classic kernel.
}

These are audit findings H1 (GPU_SAFETY.md §3): the attention kernels stride the KV cache by nk*hd, so a model where the K/V row width nkt was built from a different head dim (hd_kv ≠ hd, possible on other Qwen2.5 family models) would read out of bounds — a GPU fault, the class of bug that froze the machine once. Both checks are Errs with the actual numbers, per the invariant-violation rule. Then the dispatch decision tree of §2.5: decode (nt == 1) tries flash_attn_enabled(hd) → gqa_attn_flash with a chunk count from attention_chunks (one chunk per 32 KV rows, capped at 16, MINFER_ATTN_CHUNKS override — metal_backend.rs:220-227); else the split two-pass kernel for hd 64/128 (unless MINFER_NO_SPLIT_ATTN=1); else classic. Prefill tries attn_flash_prefill, then the 3-pass attn_parallel_prefill (unless MINFER_NO_MATMUL_ATTN=1), then classic. Note the vocabulary: "falls back to the classic kernel" here means another GPU kernel in the same family, not the CPU — the backend never hands a node back to the CPU mid-flight.

3.2.5 The fused decode ops: fewer dispatches per token

The decode-fusion arms show why the graph's fused nodes exist. Op::FusedQKV (metal_backend.rs:1727-1816) is two encodes for what unfused would be ten:

#![allow(unused)]
fn main() {
// src/graph/metal_backend.rs:894-917 (head of the FusedQKV arm)
            Op::FusedQKV { layer } => {
                let meta = match &node.meta {
                    NodeMeta::FusedQkv(m) => m,
                    other => return Err(format!("fused_qkv node missing FusedQkvMeta: {other:?}")),
                };
                let (wb, w_off) = self
                    .state
                    .weight_buf(&meta.qkv_weight)
                    .ok_or_else(|| format!("qkv weight '{}' not on GPU", meta.qkv_weight))?;
                let nt = node.out_shape[1];
                debug_assert!(nt == 1, "FusedQKV is decode (nt==1) only, got nt={nt}");
                let od_total = meta.nqt + 2 * meta.nkt;
                // 1) concat matmul: x × [wq|wk|wv] → q|k|v concat buffer
                cb.quant_matmul_f32_on_gpu_buf(
                    &wb,
                    w_off,
                    meta.weight_ttype,
                    self.buf(in_bufs[0]),
                    0,
                    self.buf(out_buf),
                    od_total,
                    meta.in_dim,
                    nt,
                );
}

The weight blk.{i}.attn_qkv was built at load time by concat_rows (src/metal/) — a byte-level row-major concatenation of the Q4_0-or-whatever blocks of wq|wk|wv, possible only when all three share a type and input dim. One matmul over the concat produces the q|k|v packed buffer; then step 2 (not shown) encodes attn_bias_rope_store — one kernel that adds the three biases, applies RoPE to q and k, and scatters k and v into the layer's KV regions at the current position. FusedFFN (:738-787) mirrors it for gate/up + an in-place swiglu (the shader reads gate rows 0..nf and up rows nf..2*nf of the same buffer and writes back into the gate rows — the aliasing the Backend contract warned about). FusedQkvNorm (:863-985, Qwen3) inserts two in-place per-head RMSNorms between the concat matmul and a no-bias rope+store, reusing the rms_norm_256 kernel with byte offsets into the concat buffer. Each fused arm ends the same way as the unfused ones: encode, Ok(()), no waiting — the win is fewer dispatches per decode step (10 → 2), which METAL_OPTIMIZATIONS.md §0.1 credits for part of the 0.5B decode gain.

3.2.6 The device layer: one device, ~70 compiled pipelines, one queue

MpsState::try_new (src/metal/runtime.rs) is where the GPU is actually acquired, and its opening decides participation:

#![allow(unused)]
fn main() {
// src/metal/runtime.rs
    pub fn try_new() -> Option<Self> {
        if std::env::var("MINFER_DISABLE_MPS").is_ok() {
            eprintln!("MPS: disabled by MINFER_DISABLE_MPS");
            return None;
        }
        // dummy for non-macOS — never called due to cfg
        #[cfg(not(target_os = "macos"))]
        return None;

        #[cfg(target_os = "macos")]
        {
            let device = MTLCreateSystemDefaultDevice()?;
}

MINFER_DISABLE_MPS=1 returns None before anything GPU-ish happens — this is the documented force-CPU switch (AGENTS.md build table). Then: grab the default Metal device; optionally start an Xcode GPU capture (MINFER_METAL_CAPTURE=1, src/metal/runtime.rs:73-79); load the shader library — the build precompiles src/metal/kernels/ into a metallib at build time (llama.cpp-style, src/metal/runtime.rs:38-43), falling back to a ~0.3-1 s runtime source compile when the toolchain was missing, with MINFER_METALLIB_FILE as a runtime override for A/B-ing compiler flags. Then comes the part that looks like boilerplate and is actually the capability table made real: ~70 get_pl("kernel_...") calls (one MetalComputePipelineState per shader entry point — 3 matmul tiers × 7 quant types, 9 get_rows variants, 2 rms_norms, elementwise ops, 5 attention families × f32/f16, store_kv × 2, the fused store kernels, warmup) — and any missing kernel name makes try_new return None here, so a shader typo degrades to CPU at startup, loudly printed ("MPS: no function '...'"), rather than faulting later. The regression test metal_pipelines_compile (metal/tests.rs:18) exists because of exactly that failure mode: "a duplicate/missing kernel or a Metal compile error makes MpsState::init fall back to CPU silently, which looks like a 'GPU throttling' slowdown" — the Q5_0 incident of 2026-08-06.

Finally the state is built with the device, the runtime-queried limits, and the single command queue:

#![allow(unused)]
fn main() {
// src/metal/runtime.rs (fields that matter; the full struct init spans :2166-2240)
            let dummy_buf = device
                .newBufferWithLength_options((1) as usize, MTLResourceOptions::StorageModeShared)
                .unwrap();
            let m = MpsStateInner {
                device: device.clone(),
                max_threadgroup_memory: device.maxThreadgroupMemoryLength() as u64,
                queue: device.newCommandQueue().unwrap(),
                pl_q4_0_f32,
                pl_q4_0_f32_multi,
                pl_q4_0_mm_f32,
                ...
}

max_threadgroup_memory is the GPU_SAFETY.md §4 rule in code: "device-specific thresholds … MUST be queried at runtime — never hardcoded." Every dispatch that uses threadgroup scratch compares its need against this cached value first (the GEMM's 8 KB check at src/metal/ is the worked example — §3.3). One queue serializes all of the engine's GPU work; try_new ends by printing the line every minfer-on-Mac user knows: MPS: using Metal on Apple M4 Pro (unified: yes) (:2241-2249).

3.2.7 Zero-copy weights: wrapping the mmap in a Metal buffer

The registration chain starts in the loader (doc 03). load_tensor (models/qwen2/loader.rs:177-184) hands every weight tensor's raw bytes to the Metal registry as it is parsed:

#![allow(unused)]
fn main() {
// src/models/qwen2/loader.rs:202-220
    // Register weight tensors with GPU backends.
    #[cfg(target_os = "macos")]
    if let Some(mps) = crate::metal::MpsState::get() {
        if matches!(
            ttype,
            TensorType::Q4_0
                | TensorType::Q4_1
                | TensorType::Q4_K
                | TensorType::Q5_0
                | TensorType::Q5_1
                | TensorType::Q5_K
                | TensorType::Q6_K
                | TensorType::Q8_0
        ) {
            mps.register_weight(&ti.name, tensor.data());
        } else if ttype == TensorType::F32 {
            mps.register_weight(&ti.name, tensor.data());
        }
    }
}

Registration happens per tensor at parse time, before any graph exists — which is what makes the §3.2.8 gate meaningful later. Before the tensors, though, the loader registers the parts — the mmap'd GGUF blobs themselves (loader.rs:319-327, one mps.register_part(part.data) per file part, multi-part included):

#![allow(unused)]
fn main() {
// src/metal/runtime.rs (register_part core; warmup dispatch follows)
    pub fn register_part(&self, data: &'static [u8]) {
        #[cfg(not(target_os = "macos"))]
        {
            let _ = data;
        }
        #[cfg(target_os = "macos")]
        {
            if data.is_empty() {
                return;
            }
            let page = 16384; // macOS page size on Apple Silicon
            let base = data.as_ptr() as usize;
            debug_assert!(base % page == 0, "mmap'd GGUF part not page-aligned");
            let buf = unsafe {
                self.inner
                    .device
                    .newBufferWithBytesNoCopy_length_options_deallocator(
                        NonNull::new(data.as_ptr() as *const std::ffi::c_void as *mut c_void)
                            .unwrap(),
                        (data.len() as u64) as usize,
                        MTLResourceOptions::StorageModeShared,
                        None,
                    )
                    .unwrap()
            };
}

newBufferWithBytesNoCopy is the whole zero-copy story in one call: Metal wraps an existing virtual-address range (the mmap of doc 02) as a GPU-visible buffer, with no copy — the hardware requirement is a page-aligned base (16 KB on Apple Silicon), which mmap guarantees and the debug_assert pins. The part is stored as (base_ptr, len, buffer); then register_weight (src/metal/runtime.rs) resolves each weight by pointer-range containment: find the part whose range contains the weight's data.as_ptr(), and record (part_buffer, ptr - base) as the weight's (buffer, offset) — the offset that execute_node passes to setBuffer_offset_atIndex (§3.2.4). The fallback path (MINFER_WEIGHT_COPY=1, or a weight that somehow falls outside every part) copies into a fresh per-weight buffer at offset 0, so correctness never depends on the zero-copy path landing.

register_part does one more thing worth understanding — the warmup. Right after wrapping the part it encodes a trivial kernel_warmup_read over the whole buffer and submits it. The comment (src/metal/runtime.rs) records the measurement: "the FIRST GPU access to file-backed (mmap) pages costs ~44 ms of one-time page/TLB setup. Doing a dummy full-buffer read HERE (at model load, outside the CLI's Total timing) moves that cost out of the first prefill." Without it, the first prompt of every session pays a hidden ~44 ms + the GPU's own page-in; with it, load time absorbs the cost and benchmark numbers are equally warm (llama.cpp does the same thing — ggml-metal-device.m).

The same loader file also pre-builds the fusion weights: when wq/wk/wv share a quant type, concat_rows (src/metal/) byte-concatenates them and registers the result as blk.{i}.attn_qkv (src/models/qwen2/loader.rs:500-501); likewise ffn_gu (src/models/qwen2/loader.rs:567-568). This is why the decode fusions of §3.2.5 can look up a single weight — the concat exists in the registry before the graph builder ever runs, but the graph only uses it when GraphParams says fuse (the params-only reuse rule of doc 13 keeps fused and unfused builds distinct).

3.2.8 The all-weights-registered gate and CParams.gpu

At forward time, before the graph is built or reused, the model asks a yes/no question (models/qwen2/graph.rs:427-439):

#![allow(unused)]
fn main() {
// src/models/qwen2/graph.rs:427-439
        // GPU availability is part of the reuse identity (backend assignment
        // lives in the built graph, not in the params' other fields). Uses
        // attributes rather than `cfg!()` so the `metal_backend` path is not
        // resolved on non-macOS builds (the module does not exist there).
        #[cfg(target_os = "macos")]
        let metal_on =
            crate::graph::metal_backend::metal_available() && Self::weights_on_gpu(model);
        #[cfg(not(target_os = "macos"))]
        let metal_on = false;
        // CUDA participation (Phase 7): requires a usable device AND every
        // matmul weight registered on the CUDA registry in a kernel-supported
        // type (all-or-nothing; 7e③ moved the embedding gather on device, so
        // tok_embd is gated like every other weight).
        #[cfg(feature = "cuda")]
        let cuda_on = crate::cuda::CudaState::get().is_some() && Self::weights_on_cuda(model);
        #[cfg(not(feature = "cuda"))]
        let cuda_on = false;
}

metal_available() (metal_backend.rs:2098-2100) is just MpsState::get().is_some() — device present, not disabled. weights_on_gpu (:721-777) is the gate itself: it enumerates every tensor name the graph will read (embedding, output norm, lm_head, output bias, then per layer the norm, wq/bq/wk/bk/wv/bv/wo, the FFN norm, ffn_gate/ffn_up/ffn_down) and requires all of them in the Metal registry:

#![allow(unused)]
fn main() {
// src/models/qwen2/graph.rs:671-677 (tail of weights_on_gpu)
        #[cfg(target_os = "macos")]
        {
            let Some(mps) = crate::metal::MpsState::get() else {
                return false;
            };
            names.iter().all(|n| mps.has_weight(n))
        }
}

One false ⇒ metal_on = false ⇒ the whole model on CPU. Why all-or-nothing instead of per-layer participation? Three reasons: a mixed graph would put a split boundary at every layer (§2.4's sync-per-split cost, multiplied); the layer loop's intermediate buffers would need cross-backend copies of GB-scale activations; and — the decisive one — the reuse identity would become fragile, since CParams.gpu could no longer describe "the GPU runs this graph". The gate's name list and register_graph_weights' list (:339-372, the CPU-side registration) are maintained as mirror images, so "the CPU has it" and "the GPU has it" stay in sync by construction.

The result flows into CParams (:447-468): gpu: metal_on || cuda_on, plus fuse_qkv/fuse_ffn gated on nt == 1 && (metal_on || cuda_on) — on the GPU the fusions are always shape-eligible at decode, on CPU they are not built (the CPU prefers the batched-quantization shape of doc 10). From there, if try_reuse fails, the rebuild path enables the backend and assigns:

#![allow(unused)]
fn main() {
// src/models/qwen2/graph.rs:478-492 (rebuild: register → enable → assign)
        if !cache.try_reuse(&params) {
            let mut graph = Self::build(model, &params);
            let sched = BackendScheduler::new();
            {
                let alloc = cache.alloc();
                Self::register_graph_weights(model, alloc);
                #[cfg(target_os = "macos")]
                if metal_on {
                    alloc.enable_metal();
                }
                #[cfg(feature = "cuda")]
                if cuda_on {
                    alloc.enable_cuda();
                }
                sched.assign_backends(&mut graph, alloc);
}

enable_metal (alloc.rs:260-272) constructs the MetalBackend (which succeeds only if MPS is live), and from this moment assign_backends's alloc.supports(...) (§2.7's priority order) hands nodes to Metal. When the gate passed, that is every node of the Qwen2/Qwen3 graph — the resulting graph is one all-Metal split, and CParams.gpu=true means the cached graph keeps that assignment until the params change.

3.2.9 One command buffer per split, flushed by the scheduler

Now assemble §2.4's rhythm from both sides. The scheduler's execute (scheduler.rs:123-354) walks splits; at every backend change it flushes and copies:

#![allow(unused)]
fn main() {
// src/graph/scheduler.rs:176-189
        for split in &splits {
            if let Some(pb) = prev_backend {
                if pb != split.backend {
                    // 1. flush the previous backend's async work
                    alloc.sync_backend(pb);
                    // 1b. staged Metal/CUDA captures are valid now — read back
                    flush_metal_captures(graph, alloc, &mut staged, trace_on, live_on);
                    flush_cuda_captures(graph, alloc, &mut cuda_caps, trace_on, live_on);
                    // 2. copy this split's inputs across backends
                    for &inp in &split.inputs {
                        alloc.copy_across(inp, split.backend)?;
                    }
                }
            }
}

Step 1 is the command buffer's submit: sync_backend (alloc.rs:2400-2406) routes to MetalBackend::synchronize, which is one line — self.submit_pending() (metal_backend.rs:2078-2080). Step 2 is the cross-backend copy of §2.2 (host round trip through the shared buffers, into a fresh staging buffer so the producer's own buffer is untouched for the graph's re-executability). On a fully-Metal graph there is one split, so this if never fires mid-graph — but the final sync after the loop (:377-381) always does, which is where the decode forward's single submit lands. If any node's buffer turned out to live on a different backend than the split executing it, the scheduler returns a hard Err ("assignment/alloc mismatch", :263-270) — the same no-silent-fallback posture, one level up.

The submit itself (metal_backend.rs:189-215) takes the leaked box back, calls cb.submit(), and — under MINFER_OP_PROFILE=1 — accumulates the GPU wait time that §4's profile table prints. The Drop impl does the same flush best-effort (let _ = cb.submit()) so a backend dropped mid-split cannot leak an unterminated encoder.

3.2.10 submit(): the bounded wait

MpsCommandBuffer::submit (src/metal/ops.rs) is the GPU-safety centerpiece — the fix for the incident that motivated docs/GPU_SAFETY.md:

#![allow(unused)]
fn main() {
// src/metal/encode.rs (core; encoder-end at :1936-1938)
        // dispatch_semaphore_t is already a reference-counted opaque pointer.
        let sem = unsafe { dispatch_semaphore_create(0) };
        let sem_val = sem as usize;

        let blk = RcBlock::new(
            move |_cb: NonNull<ProtocolObject<dyn MTLCommandBuffer>>| unsafe {
                dispatch_semaphore_signal(sem_val as *mut c_void);
            },
        );
        unsafe {
            self.cmd_buf.addCompletedHandler(RcBlock::into_raw(blk));
        }
        self.cmd_buf.commit();

        // Bounded wait (10 s). If the GPU hangs (hardware fault), the completion
        // handler never fires and we bail out instead of blocking forever.
        let timeout = unsafe { dispatch_time(0, 10_000_000_000i64) }; // 10 s from now
        let rc = unsafe { dispatch_semaphore_wait(sem, timeout) };
        unsafe {
            dispatch_release(sem);
        }

        if rc == 0 {
            // Command buffer finished (possibly with an error status).
            match self.cmd_buf.status() {
                MTLCommandBufferStatus::Completed => Ok(()),
                st => Err(format!(
                    "Metal command buffer status={st:?}. recent dispatches: {}",
                    self.recent_trace()
                )),
            }
        } else {
            // Timed out: the GPU did not complete the work.
            Err(format!(
                "Metal command buffer timed out after 10s (GPU hang). recent dispatches: {}",
                self.recent_trace()
            ))
        }
    }
}

Walk it as three defenses. (1) A completion handler on a semaphore: addCompletedHandler registers an Objective-C block that fires when the GPU finishes (or fails) the command buffer; the host sleeps on a dispatch_semaphore with a 10-second deadline — dispatch_semaphore_wait returns non-zero on timeout instead of blocking forever. Before this hardening (GPU_SAFETY.md §2.1) the code waited DISPATCH_TIME_FOREVER and never checked status, so "a single GPU fault would block minfer forever (and, since Metal clients share the GPU, could stall WindowServer → whole-machine freeze)". (2) A status check: even when the semaphore fires, the buffer's MTLCommandBufferStatus is verified — Completed means success; Error means a kernel faulted. (3) A diagnosis trail: both failure paths append recent_trace() — the ring of the last 16 dispatch labels recorded by trace_op when MINFER_TRACE=1 (src/metal/) — so the error message names the kernel family that was last encoded. All three exist because the alternative (an unkillable hang, or an error with no clue) is disproportionately expensive on shared-GPU macOS. Note the return type: Result<(), String>, which submit_pending turns into .expect(...) — a submit failure panics the run rather than pretending it succeeded, the same fail-loud rule as execute_node's Errs.

3.2.11 Host read/write: views, plus the readback that feeds the sampler

#![allow(unused)]
fn main() {
// src/graph/metal_backend.rs:1104-1127
    fn read_host(&self, id: usize) -> Option<&[f32]> {
        let buf = self.pool.get(id)?;
        let len = (buf.length() as usize) / 4;
        Some(unsafe { std::slice::from_raw_parts(buf.contents().as_ptr() as *const f32, len) })
    }

    fn write_host(&mut self, id: usize, data: &[f32]) -> Result<(), String> {
        let buf = self.pool.get(id).ok_or_else(|| format!("no buffer {id}"))?;
        let len = (buf.length() as usize) / 4;
        if len != data.len() {
            return Err(format!(
                "buffer {id}: expected {len} elements, got {}",
                data.len()
            ));
        }
        unsafe {
            std::ptr::copy_nonoverlapping(
                data.as_ptr(),
                buf.contents().as_ptr() as *mut f32,
                data.len(),
            );
        }
        Ok(())
    }
}

This is §2.2's promise made concrete: read_host reinterprets the Metal buffer's contents() pointer as a Rust &[f32] — no copy, no sync — and write_host is memcpy with a length check. The two call patterns that matter: inputs are written here before the split's nodes encode (the scheduler skips Op::Input nodes precisely because "data pre-filled by the allocator", scheduler.rs:225-227), and outputs are read after the split's sync — the logits path ends with alloc.copy_to_cpu(graph.outputs[0]) (models/qwen2/graph.rs:624), which lands here once the final submit's completion handler has fired. Reading a buffer whose producing kernels are still encoded but not submitted would be the classic race; the one-command-buffer-per-split discipline plus the bounded submit is what makes these plain views safe.

There is one more readback path, used only by the trace/viz tooling (MINFER_TRACE, the viz server): capture_split (metal_backend.rs:173-186) encodes a blit pass — encode_captures (src/metal/encode.rs), a GPU→GPU copy into per-split staging buffers appended after all the split's kernels — so the capture reads this step's data without forcing a per-node flush. The scheduler queues (node, staging) pairs during the split (scheduler.rs:306-344) and drains them in flush_metal_captures right after the boundary sync (:394-417). It is a nice illustration of the submit model: even debug tooling has to work with the async pipeline, by scheduling its reads into the same command buffer.

3.2.12 One shader, walked: kernel_rms_norm_f32

Every graph node ends in something like this — a Metal Shading Language function in src/metal/kernels/. RMSNorm (doc 11 §3.2.1) is the best first shader: short, and it shows every GPU-programming concept this backend uses:

// src/metal/kernels/norm_elementwise.metal — one threadgroup per row
kernel void kernel_rms_norm_f32(
    device const float * x       [[buffer(0)]],
    device const float * w       [[buffer(1)]],
    device       float * y       [[buffer(2)]],
    constant    int    & d       [[buffer(3)]],
    constant    float  & eps     [[buffer(4)]],
    uint3 tgpig [[threadgroup_position_in_grid]],
    uint3 tpitg [[thread_position_in_threadgroup]],
    uint3 ntg   [[threads_per_threadgroup]]
) {
    int row = tgpig.x;
    int d4 = d / 4;

    device const float4 * x4 = (device const float4 *)(x + row * d);

    float ss = 0.0f;
    for (int i = tpitg.x; i < d4; i += 32) {
        ss += dot(x4[i], x4[i]);
    }
    int rem = d - d4 * 4;
    if (tpitg.x == 0) {
        device const float * x_tail = x + row * d + d4 * 4;
        for (int i = 0; i < rem; i++) ss += x_tail[i] * x_tail[i];
    }
    ss = simd_sum(ss);

    float scale = 1.0f / sqrt(ss / (float)d + eps);

    device float4 * y4 = (device float4 *)(y + row * d);
    device const float4 * w4 = (device const float4 *)w;
    for (int i = tpitg.x; i < d4; i += 32) {
        y4[i] = x4[i] * scale * w4[i];
    }
    if (tpitg.x == 0) {
        device const float * x_tail = x + row * d + d4 * 4;
        device       float * y_tail = y + row * d + d4 * 4;
        device const float * w_tail = w + d4 * 4;
        for (int i = 0; i < rem; i++) y_tail[i] = x_tail[i] * scale * w_tail[i];
    }
}

Beginner's decoder ring, line by line:

  • kernel void marks an entry point the CPU can launch (the get_pl("kernel_rms_norm_f32") of §3.2.6 compiles it into a pipeline). The [[buffer(N)]] attributes are the argument slots — they match the setBuffer_offset_atIndex(_, _, N) and setBytes_length_atIndex(_, _, N) calls you saw in src/metal/: buffer 0/1/2 are the device pointers for x, the norm gains w, and the output y; buffers 3/4 are small by-value scalars (d, eps) passed via setBytes. This is the whole CPU↔shader calling convention: pointers into shared memory plus a few ints.
  • The built-in coordinates replace the CPU's loop indices. A GPU launch is a grid of threadgroups (here: one threadgroup per row — rms_norm dispatches dispatch_2d(n, 1, 32, 1) with n = row count, src/metal/ops.rs), each holding 32 threads (one simdgroup — the hardware unit of 32 lanes that execute in lockstep). tgpig.x is "which row am I", tpitg.x is "which of my 32 lanes am I". Compare doc 10's mm_rows: there the loop variable was a row owned by a worker thread; here it is a row owned by 32 lanes.
  • The strided loop for (i = tpitg.x; i < d4; i += 32) splits the row's d/4 float4-vectors across the 32 lanes — lane 0 takes vectors 0, 32, 64…, lane 1 takes 1, 33, …. Each dot(x4[i], x4[i]) is a 4-wide multiply-add per lane: with d = 896 that is 224 vectors, so 7 iterations per lane. The float4 cast is free vectorization — the compiler emits 128-bit loads against the row-major layout the doc 07 allocator laid down.
  • simd_sum(ss) is the cross-lane reduction: one instruction that adds the 32 lanes' partial sums and broadcasts the total. On the CPU this was hsum_float_8's shuffle dance (doc 10 §3.2); here the hardware does it, in lockstep, with no explicit synchronization — this is the "same math, different execution model" of doc 11 §2.5 in one line.
  • The tail loop (rem = d - d4*4, if (tpitg.x == 0)) handles d not divisible by 4 in scalar — lane 0 alone sweeps the last rem elements. A beginner-relevant detail: which lanes do work is decided by lane index, not by data, so all 32 lanes still reach the simd_sum together — the barrier-safety rule of §3.4 in miniature.
  • The second pass reuses the same strided pattern to write y = x · scale · w. Note the two passes over x (sum of squares, then normalize): on the CPU, doc 11 could do the same because the row fits in cache; a flash-style online variant exists for the cases where it does not — here the row is small enough that two passes through L1/L2 are cheaper than saving state across the reduction.

The 256-thread sibling kernel_rms_norm_f32_256 (src/metal/kernels/norm_elementwise.metal, dispatched by default — rms_norm_256_enabled(), ~2× faster per METAL_OPTIMIZATIONS.md #16) shows the other reduction tool, and the repo's most important GPU-safety rule in its natural habitat:

// src/metal/kernels/norm_elementwise.metal — two-barrier handshake across 8 simdgroups
    ss = simd_sum(ss);

    threadgroup_barrier(mem_flags::mem_threadgroup);
    if (tiisg == 0) {
        shmem[sgitg] = ss;
    }
    threadgroup_barrier(mem_flags::mem_threadgroup);

    ss = shmem[tiisg];
    ss = simd_sum(ss);

With 256 threads (8 simdgroups), one simd_sum is no longer enough: each simdgroup reduces its own 32 lanes, then lane 0 of each simdgroup publishes its subtotal into threadgroup shared memory (shmem, declared [[threadgroup(0)]], 8 floats). The two threadgroup_barriers are the handshake: the first guarantees shmem is zeroed/ready, the second guarantees every simdgroup's store landed before anyone reads. A threadgroup_barrier is a lane-count-wide rendezvous — every thread must arrive; if any lane took an early return, the rest would wait forever and the GPU would deadlock (the machine-freeze class from GPU_SAFETY.md §2.2). That is why the rule is "no early return past a threadgroup_barrier": guards must be expressed as predication (lanes run the code with harmless values, then skip their stores) rather than exits — exactly how this kernel and the attention kernels do it.

3.3 Design choices (why this shape and not another)

Why handwritten kernels instead of Apple's MPS/MPSGraph? Because minfer has zero ML-framework dependencies and needs llama.cpp-parity numerics. Apple's MPS is a closed library: fixed quant formats, no visibility into summation order, and no way to guarantee the byte-identical greedy outputs the verification gates rely on. Writing kernels in MSL (mostly as transcriptions of llama.cpp's, which src/metal/kernels/ says openly — "faithful llama.cpp transcription") keeps the GGUF block layouts as the single contract between disk, CPU, and GPU, the same way doc 10's CPU kernels do. The cost is real — every quant needs Metal kernels for its tiers, and the optimization campaign is hand-measured (METAL_OPTIMIZATIONS.md) — but the control is what makes the correctness claims checkable.

Why the objc2 ecosystem — how does Rust call Metal without bindgen? Metal is an Objective-C API; Rust reaches it by sending Objective-C messages. minfer's device layer speaks that protocol through objc2-metal (plus block2 for the completion-handler block and objc2-foundation for strings), migrated from the legacy metal/objc 0.2 crates on 2026-08-25 (docs/METAL_OBJC-ECOSYSTEM.md records the whole story). The practical differences you can see in this doc's code: Retained<ProtocolObject<dyn MTLBuffer>> is a typed, ref-counted handle (RAII — drop releases, versus the old manual msg_send! retain/release); method calls are real Rust signatures checked at compile time (the old msg_send! with a typo'd selector failed at runtime); and RcBlock (§3.2.10's completion handler) wraps the Objective-C block safely. The old ecosystem was frozen at objc 0.2.7 (2019) and carried a future-rustc breakage that minfer had to vendor-patch; the migration removed the vendored patch, the metal crate, and that whole risk class. What did not change: the architecture — MpsState as the singleton, MpsCommandBuffer as the encoder wrapper — survived the migration mostly untouched, which is the best argument that the layering was right.

Why name-keyed weight registration? register_weight(name, bytes) stores (buffer, offset) under the tensor's GGUF name, and execute_node looks weights up from NodeMeta's name strings (meta.weight_name), never from a Tensor. The alternative — handing the backend a reference to the tensor — would couple the backend's lifetime to the model object and break on graph rebuilds (doc 13: the graph is rebuilt while the model, and its registered weights, stay put). With names, registration is a one-time load event and every rebuild re-resolves by name; the all-weights-registered gate is then literally a contains_key sweep over the same map (§3.2.8).

Why is Metal first in the priority queue? supports (alloc.rs:141-158) checks Metal, then CUDA, then CPU — so on a Mac with both compiled in, Metal wins every op it can run and CUDA gets only what Metal refuses (today: nothing on the standard graph, since supports_op covers it all; CUDA's unique claims are its own fusions like QkvBiasRopeStore). The order is a policy statement: on Apple Silicon, Metal is the native, zero-copy path — CUDA there runs through a translation layer with real copy costs. On a datacenter box the CUDA backend's supports_op is broader (int8 MMQ prefill, doc 15), so "first yes wins" naturally routes to whichever backend is both present and most capable per op.

Why a fresh allocation for cross-split staging (alloc_fresh) instead of reusing the free list? The trait comment (§2.1) is the record of a real bug class: at execute time, free-list ids still back node_to_buf entries that later nodes in the same execute will read or write; recycling one for a staging copy would clobber in-flight data. The staging buffer is born outside the recycle economy and joins it only at graph-rebuild time, when liveness recycling is safe again (doc 07's allocator owns that transition).

Why is the f16 KV cache auto-selected by model size? set_kv_cache_type (src/metal/policy.rs) turns f16 on when n_layers × n_kv_embd ≥ 8192 — the 7B class — and keeps f32 for the 0.5B class. The asymmetry is measured, not aesthetic: on the 7B, attention streams the whole KV per decode step, so halving those bytes is a direct win (−~1 ms/token at 2K ctx; #13 measured 1.60 → 0.95 s on a long-context case); on the 0.5B the decode is dispatch-latency-bound, the f16 conversion kernels cost more than the bandwidth saves, and f16 measured ~3 % slower (the comment cites the decided-not record). MINFER_CACHE_TYPE=f16|f32 overrides either way.

Why does execute_node return Err while some dispatch guards gpu_abort? Two layers, two severities. Err is for graph-level invariant violations — wrong meta, missing weight, unsupported op — where the scheduler can propagate a clean message naming the node and abort the run. gpu_abort is for dispatch-level hazards — a shape that would make a shader index out of bounds or overrun threadgroup memory — where the safest thing is to print the actual offending numbers and exit before the GPU ever sees the launch (GPU_SAFETY.md §2.3/§4). Both exist because the third option, silently continuing on another path, hides the bug while the model keeps talking.

3.4 Pitfalls & invariants

  • Never host-copy a GPU-pending buffer. Shared memory makes every pool byte look readable at all times — but between execute_node and the split's submit, a buffer's future contents are still in flight. The Phase-3 KV-corruption bug (recorded in AGENTS.md core rule 5) came from exactly this: an in-place op's input was host-copied while the GPU had pending writes to it. The invariants that make the current code safe: inputs are host-written before their split encodes; outputs are read only after synchronize; and in-place GPU ops snapshot their input inside the backend when aliasing is not already guaranteed:
#![allow(unused)]
fn main() {
// src/graph/metal_backend.rs:239-251
    fn copy_in(&self, dst: usize, src: usize) {
        // in-place-ish ops (silu/rope) may alias; snapshot to dst first
        let src_buf = self.buf(src);
        let dst_buf = self.buf(dst);
        let n = (src_buf.length().min(dst_buf.length()) / 4) as usize;
        unsafe {
            std::ptr::copy_nonoverlapping(
                src_buf.contents().as_ptr() as *const f32,
                dst_buf.contents().as_ptr() as *mut f32,
                n,
            );
        }
    }
}

The Silu and RoPE arms call this when in_bufs[0] != out_buf (e.g. metal_backend.rs:363-370) — the allocator promised the alias is safe for the same backend (doc 07's sole-consumer rule), and this memcpy makes it byte-identical to executing in place, host-side or GPU-side.

  • One memory barrier between every pair of dispatches. Metal orders kernels but does not make writes visible to the next kernel for free; the barrier() at the end of dispatch_1d/2d/3d (src/metal/, 405, 421) is mandatory glue. The 2026-08-19 incident (METAL_OPTIMIZATIONS.md #28, GPU_SAFETY.md §3 post-audit finding): without it, the reused bn buffer raced between RMSNorm's writes and the next op's reads, producing first-token nondeterminism on 1.5B/7B (~10-30 % wrong tokens) that vanished and reappeared with system load. Corollary recorded there: when a threadgroup-memory buffer is reused for a different purpose at a different loop stage, a threadgroup_barrier must separate the last read from the first write — the GEMM temp_str fix.

  • No early return past a threadgroup_barrier (§3.2.12's ending). The original sin: kernel_gqa_attn_f32 had if (h >= nh) return; before a barrier; when nh % nk != 0, some simdgroups exited while others waited at the barrier — "GPU permanent deadlock = machine freeze" (GPU_SAFETY.md §2.2). The fix pattern — invalid lanes run the full loop on a dummy index and skip only the final store via a valid_head flag — is now the review rule for every new kernel, and the flash kernels' discipline (GPU_SAFETY.md §4b: mask computed inline per lane, break-only control flow on lane-independent conditions, shuffles instead of barrier-protected shared arrays) is the same rule applied to a harder kernel.

  • Device limits are queried, then compared — never guessed. The worked example is the GEMM's threadgroup-scratch check:

#![allow(unused)]
fn main() {
// src/metal/
        if 8192 > self.state.max_threadgroup_memory {
            gpu_abort(&format!(
                "GEMM needs 8192 B threadgroup memory, device max is {} B",
                self.state.max_threadgroup_memory
            ));
        }
}

max_threadgroup_memory was captured from the device at init (§3.2.6); the guard compares the kernel's real need (4 KB + 2 KB + reused 8 KB staging, per the comment at src/metal/) and aborts with both numbers. The rule exists because the first draft hardcoded a guessed 32 KB for the M4 Pro (GPU_SAFETY.md §4) — a number that would be silently wrong on the next chip.

  • Quant-shape assumptions get refused, not absorbed. The K-quant id % 256 guard (§3.2.4, audit M1) exists because the failure mode is wrong numbers, no crash — the worst kind. The same audit table accepts genuinely low-risk gaps with reasoning (L1: matmul pointers computed past the buffer for out-of-range rows but reads guarded; L2: store_kv trusts the host to keep positions < capacity, which doc 07's region sizing guarantees). Every accepted risk is written down with its mitigation — the audit doc is the invariant ledger.

  • The registry and the graph must agree on shapes. FusedQkvNorm's per-head norms assume the concat layout (q at byte 0, k at nqt*4 via off_k — metal_backend.rs:1880-1881); attn_bias_rope_store assumes the same packing; the f16 KV kernels assume the region's first half is theirs (§3.1). These cross-file layout contracts are exactly where doc 10's lesson applies: they are tested with real-scale isolation tests (metal_attn_kv_real_scale, metal_store_real_dims, … — the #[cfg(test)] module at metal_backend.rs:1141), not just small shapes, because the transposed-output bug of doc 10 hid from nt == 1 tests.

  • A shader typo must fail at startup, not at token 500. §3.2.6's pipeline table makes every kernel name a startup-checked fact; the metal_pipelines_compile test pins it in CI. The Q5_0 incident (2026-08-06) — a duplicate symbol that silently downgraded the process to CPU and "looked like GPU throttling" — is the reason this is treated as an invariant rather than an annoyance.

4. Observe & verify

  • Startup line: a Metal-enabled run prints MPS: using Metal on <device> (unified: yes) then MPS: GPU acceleration enabled (src/metal/runtime.rs, 2262). Its absence — or MPS: not available, using CPU fallback — is the first thing to check when numbers look CPU-shaped; MINFER_DISABLE_MPS=1 produces the explicit MPS: disabled by MINFER_DISABLE_MPS.
  • MINFER_GRAPH_TRACE=1 prints the split table (scheduler.rs:127-144): one line per split (split 0: Metal nodes 0-440) plus an op×backend census — the quickest way to see the all-Metal split a passing gate produces, or the CPU stragglers when the gate failed.
  • MINFER_OP_PROFILE=1 turns on the backend's built-in profiler (metal_backend.rs:26-83, 224-242): after the first submit it prints a host-encode-per-op table (top 20 by time), then one line per submit with the GPU wait — prefill shows one big submit, decode shows one line per token. This is the tool that measured the 32-vs-256-thread RMSNorm dispatch cost (#16).
  • MINFER_TRACE=<dir> records per-node real-data traces (doc 08's staged Metal capture path — blits at split end, read after sync); the same env var arms submit()'s dispatch-label ring, so a Metal error/timeout message names the last 16 kernels encoded.
  • A/B levers for every §2.5 decision: MINFER_NO_FLASH=1 (decode flash → split), MINFER_NO_PREFILL_FLASH=1, MINFER_NO_MATMUL_ATTN=1 (parallel prefill → classic), MINFER_NO_SPLIT_ATTN=1, MINFER_NO_RMS_256=1, MINFER_GEMM=0 (GEMM tier off), MINFER_CACHE_TYPE=f16|f32|q8_0 (KV width; all three backends read it, Metal since #310, C4), MINFER_ATTN_CHUNKS=N, MINFER_NO_FUSE_QKV=1 / MINFER_NO_FUSE_FFN=1 (fusion off — changes CParams, forces a rebuild). Each is a one-env-var kernel-family A/B, the same levers the optimization campaign used.
  • MINFER_METAL_CAPTURE=1 starts an Xcode GPU capture at device init (src/metal/runtime.rs) — open the .gpu capture in Xcode to see every encoded dispatch of a run; MINFER_METALLIB_FILE=<path> swaps the precompiled shader library at runtime; MINFER_WEIGHT_COPY=1 forces the copied-weight path to isolate zero-copy registration bugs.
  • Tests (macOS, cargo test): metal_pipelines_compile (metal/tests.rs:18) fails if any kernel is missing or the library does not compile — the anti-silent-CPU-fallback guard; the metal_backend.rs test module carries per-op correctness gates against host-computed references, e.g. metal_matmul_q8_matches_cpu builds a graph, runs it through the real scheduler with every node forced to Metal, and asserts max diff < 1e-3 against a manual Q8×f32 reference (:1775-1782), plus cross-backend copy, split alternation, KV attention, decode-step, and real-scale (d=896) variants; the greedy end-to-end gates of doc 12 close the loop (byte-identical greedy output across the optimization campaign is the standard the records claim).
  • The honest caveat: on a CPU-only build these tests print MPS unavailable; skipping and pass — Metal coverage exists only where Metal does. A no-op GPU test suite is one of the failure modes docs/GPU_SAFETY.md's recurrence playbook is written for.

5. Cross-references

  • 06 — Assign and fusion — how supports_op/supports_fused are consumed at build time; the priority walk of §2.7 starts there.
  • 07 — Allocator, liveness, and KV regions — who owns the buffers this backend pools, the aliasing rule copy_in implements, and the persistent KV regions the KV ops write.
  • 08 — Scheduler and execute §3.2 — the split loop, cross-backend copies, and the "one Metal command buffer per split" rule this doc implements from the backend side.
  • 10 — CPU matmul kernels — the execution model this doc replaces: Q8_0 activations vs f32 (§2.3), thread-pool rows vs simdgroup lanes, and the bandwidth physics both share.
  • 11 — Attention, vec ops, and the KV cache §2.5 — the GPU preview this doc expands (flash vs full-row softmax, f16 KV); §3.2's CPU RMSNorm is this doc's shader in scalar form.
  • 13 — The decode loop and graph reuse — why CParams.gpu is part of the reuse identity and when the gated graph gets rebuilt.
  • 15 — The CUDA backend — the same Backend trait where every unified-memory shortcut of §2.2 becomes an explicit copy, plus int8 MMQ and CUDA Graph replay.
  • docs/GPU_SAFETY.md — the hard rules and the incident report behind §3.2.10, §3.2.12, and §3.4; read before touching Metal/CUDA code.
  • docs/METAL_OPTIMIZATIONS.md — the kernel campaign behind §2.5/§2.6 (§0.1's table maps every graph op to its kernel, with measurements).
  • docs/METAL_OBJC-ECOSYSTEM.md — why objc2, what the 2026-08-25 migration changed, and the nix/Xcode toolchain gotchas (§3.2.6's metallib story).
  • docs/metal-inference-analysis.md, docs/PERF-QWEN3-4B-VS-LLAMACPP.md — deeper per-kernel analyses and the Qwen3-4B llama-parity A/B.

← 13 — The decode loop and graph reuse · Index · 15 — The CUDA backend →

15 · The CUDA backend

Stage: the Metal contrast (doc 14) → this stage: the same compute graph on NVIDIA GPUs → end of the series (Index). Code: src/graph/cuda_backend.rs (CudaBackend, execute_node_inner, graph_replay_step), src/cuda.rs (CudaState singleton, register_weight, matmul_f32_ptr_layout, prefill_mmq, gqa_attn_split), src/cuda/kernels/*.cu (the CUDA C++ kernels — common.cuh + 17 translation units, 9,996 lines, since #263), build.rs (the nvcc build chain).

1. Background — where this stage sits

Docs 08 and 13 left the inference loop in a particular shape: the scheduler walks each backend-uniform split of the compute graph and calls execute_node for every node, and the decode loop repeats that walk once per generated token, reusing the cached graph. Doc 10 showed what a MatMul node becomes on the CPU; doc 14 showed the same nodes executing on Apple's GPU via Metal. This document completes the trilogy: the same graph, the same allocation rules, the same safety contract — executed on an NVIDIA GPU.

The CUDA backend is the third backend of the engine and the only one that is opt-in at build time. A plain cargo build --release never touches the CUDA toolchain at all; adding --features cuda makes build.rs locate nvcc, compile each src/cuda/kernels/*.cu into an object file, archive them into libcuda_kernels.a, and link it in (docs/BUILD.md records the full recipe, including which GPU architectures get machine code baked in). That opt-in flag is why the file src/graph/cuda_backend.rs — the subject of most of this document — is wrapped in #[cfg(feature = "cuda")] and simply does not exist in other builds.

Three source files cooperate, and keeping their roles separate makes the rest of the document easy to follow:

  • src/cuda.rs (1,014 lines of Rust) plus the 18 src/cuda/methods/*.rs family files — the device layer. A process-wide singleton, CudaState, owns the CUDA context pieces: the device, one command stream, the weight registry (name → device pointer), the KV-cache regions, pinned-host staging pools, and one host-side wrapper function per kernel. It talks to the CUDA runtime through hand-written extern "C" declarations (there is no -lcuda crate and no bindgen — the FFI surface is explicit and auditable).
  • src/graph/cuda_backend.rs (2,351 lines + 8,511 lines of tests in a 106-line parent and 11 tests/<topic>.rs files) — the graph backend. It implements the same Backend trait as the CPU and Metal backends (doc 08): supports_op, a device buffer pool, execute_node, host read/write, synchronize. Its job is translation, not math: turn a CNode into one or two kernel launches on the shared stream, and enforce the kernel invariants loudly (Err) when they are violated.
  • src/cuda/kernels/*.cu — the kernels themselves, CUDA C++ compiled by nvcc, one translation unit per kernel family (common.cuh holds the shared macros/helpers). Every function the backend calls is a launch_* wrapper in cuda.rs that eventually reaches a __global__ kernel here.

Everything upstream of this doc is unchanged by the backend swap: the graph was built by the model code (doc 05), backends were assigned at build time by supports_op priority Metal → CUDA → CPU (doc 06), buffers were placed by the liveness allocator with two persistent KV regions per layer (doc 07). The CUDA backend executes whatever it was assigned; it never rewrites the graph and never falls back to the CPU mid-run.

What this document adds beyond the Metal story is three CUDA-specific mechanisms, each answering a question the other backends never had to ask:

  1. Prefill runs an int8 tensor-core GEMM ("MMQ") — quantized weight bytes feed NVIDIA's mma integer matrix instruction directly, llama.cpp-style, instead of dequantizing first (§2.2–§2.3, §3.2.5–§3.2.6).
  2. Decode attention is "split-KV" — the KV dimension is cut into 32 stripes processed by 32× more warps than heads alone would occupy, then reduced (§2.4, §3.2.7).
  3. Decode steps are recorded as CUDA Graphs and replayed — one launch per split instead of hundreds of kernel launches (§2.5, §3.2.8).

The implementation arc itself is worth knowing up front, because the code carries its history in comments: the backend was built in five planned sub-phases (7a skeleton → 7b per-op parity → 7c model wiring → 7d graph replay → 7e polish, all recorded in docs/CUDA-BACKEND-DESIGN.md), and then a long measurement-driven optimization campaign (Phase 8 and the r-series, recorded in docs/CUDA_OPTIMIZATION.md plus one document per step in docs/cuda_optimization_steps/) turned a working backend that ran 7B prefill at 30.7 tok/s into the current default path at ~3,581 tok/s — 1.080× llama.cpp on the same hardware and model. §2.3 and §3.3 cite that record where a design decision needs its measurement.

2. Principle — how it works and why

2.1 Five GPU words you need before the code makes sense

NVIDIA's execution model can be compressed into five terms, and every kernel excerpt in §3 uses them:

  • A thread is the unit of work — it computes, say, four elements of one dot product. A kernel launch creates a grid of threads organized into blocks; the hardware schedules whole blocks onto SMs (Streaming Multiprocessors, the GPU's ~100-odd cores).
  • Threads within a block can share a small shared memory scratchpad and synchronize with barriers. Threads are grouped 32 at a time into a warp, which executes one instruction across its lanes in lockstep — a warp reduction (__shfl_xor_sync) sums 32 lanes in 5 steps, which is why the attention kernel in §3.2.7 maps one warp to one KV stripe.
  • Occupancy is how many warps an SM can keep resident at once. Latency (a few hundred cycles for a VRAM read) hides only behind other warps' work, so "more resident warps" is the default medicine — until shared memory or registers per block cap it. Several optimization records in this doc are occupancy stories (r40's third resident block, D3-4's dual-kernel dispatch).
  • Coalescing: adjacent threads should read adjacent addresses so one memory transaction serves the whole warp. The Q6_K weight repack (§3.2.2) exists purely to make this possible.
  • The stream is the GPU's in-order work queue. Everything the backend enqueues — copies and kernels — runs in issue order on one stream, which is precisely what makes the pinned-staging async fills (§3.2.9) race-free and what CUDA Graph capture records.

One more word: a tensor core is a piece of silicon inside each SM that executes small matrix-multiply-accumulate instructions (the mma family) at many times the rate of ordinary arithmetic. The f16 flavor (wmma) is the floor for minfer's CUDA kernels — sm_70 (Volta) is the minimum supported architecture because of it (docs/BUILD.md) — and the int8 flavor (mma.m16n8k32, sm_80+) is what prefill uses (§2.3).

2.2 Two forwards, two physics: why prefill and decode run different matmul kernels

The single most important fact about GPU inference is that prefill and decode are bound by different resources, and the backend dispatches on that difference. Recall the two phases (docs 09 and 13):

  • Prefill pushes all prompt tokens through the graph at once. Each matmul is now a true GEMM (General Matrix-Multiply): nt token rows stream through the same weight matrix, so the weights are read once while the arithmetic grows with nt. With enough tokens, the GPU runs out of math to do before it runs out of bytes to fetch — compute-bound.
  • Decode produces one token per forward. Every matmul must still read its entire weight matrix to produce that one token's outputs — there is nothing to amortize over. Weight-streaming-bound: the speed of light is VRAM bandwidth, not FLOPs.

The engine's own numbers make both limits visible. The 7B Q4_K_M model's weights total ~4.4 GB; measured decode is ~51.2 tok/s (README.md performance table), which is ~225 GB/s of sustained weight streaming — exactly the "200–225 GB/s class" the dispatch comments in cuda.rs report for the decode kernels. That also shows why quantization matters just as much on a GPU as on a CPU (doc 10 §2.1): the same model in f32 weights would move 4× the bytes per token and stream 4× slower, and would not fit next to the KV cache on most GPUs anyway (14B Q4_K_M already occupies ~14.1 GB of device memory). Decode reads every weight byte per token from VRAM — "weight streaming" — so bits-per-weight converts 1:1 into tokens-per-second.

Prefill is the opposite trade. Once nt is large enough, the weight bytes are a one-time cost and the tensor cores want maximum math throughput. That is why matmul_f32_ptr_layout — the single dispatch every CUDA matmul flows through — opens with a token-count gate:

#![allow(unused)]
fn main() {
// src/cuda/methods/weights.rs:487-509 (dispatch gate; abridged comment)
if (nt >= 9 || small_m_gemm)
    && id % 32 == 0
    && !Self::no_prefill_gemm()
    && matches!(ttype, TensorType::Q4_0 | TensorType::Q4_1
        | TensorType::Q5_0 | TensorType::Q5_1 | TensorType::Q8_0
        | TensorType::Q4_K | TensorType::Q5_K | TensorType::Q6_K)
{
    if self.mmq_active() {
        // doc 92 resolution: auto-ksplit is the DEFAULT for every route
        // into the BT GEMM (ksplit_req = -1 lets prefill_mmq decide).
        return self.prefill_mmq(wptr, ttype, x, out, od, id, nt, padded_q6k, ksplit_req);
    }
    return self.prefill_gemm_f16(wptr, ttype, x, out, od, id, nt, padded_q6k);
}
}

nt >= 9 (Step 82 lowered it from 16) sends prefill-shaped batches to one of two tiled GEMMs; the nt == 1 decode shapes fall through to per-type kernels that keep the f32 activations and optimize for weight streaming. There is even a middle band: nt 2–8 runs multi-token MMVQ — MMVQ with an in-block token loop — because a batch that small still cannot fill GEMM tiles, but re-running the whole per-token path per token wastes weight reads (Step 82, docs/CUDA_OPTIMIZATION.md §0 row 82). The small_m_gemm term in the gate is the doc-91 experiment knob (MINFER_SMALL_M_GEMM=1) that routes that middle band into the BT path instead — measured ~1.7× slower than multi-MMVQ at small nt, so it stays off by default.

So the answer to "why does prefill use int8 MMQ while decode uses a different path?" is not taste — it is which resource is scarce in each phase:

PhaseShapesScarce resourceKernel familyWhy
Prefillnt ≥ 9math throughputint8 MMQ tensor-core GEMM (or f16 wmma GEMM)one tiled GEMM; tensor cores do the MACs; weights stream once
Small batchnt 2–8weight bytes, but too few rows for tilesmulti-token MMVQ (dp4a, token loop in-block)weights-once like a GEMM, launch-lean like decode
Decodent == 1VRAM bandwidthper-type MMVQ (dp4a, one row per block) + f32-activation kernelsmaximize bytes/s; integer dots are free alongside the stream

2.3 MMQ, explained from zero: int8 tensor cores eat quantized weights

MMQ (llama.cpp's name for "matrix-matrix quantized") is the trick the CPU doc 10 introduced in scalar form — keep weights quantized, quantize activations to int8 on the fly, do an integer dot — lifted onto NVIDIA's int8 tensor-core instruction. The pieces:

  • The weights are never dequantized. Their raw block bytes (block_q4_K's nibbles, sub-scales, mins...) are staged into shared memory as-is and decoded in registers right where the mma instruction needs them.
  • The activations are quantized once per matmul call into padded int8 blocks — nt × (id/32) × 40 bytes (each 32-value block takes 40 bytes: 32 int8 values plus scale/sum fields the kernels consume). A small prepass kernel does this before the GEMM.
  • The mma.m16n8k32.s8 instruction multiplies an 16×32 int8 tile by a 32×8 int8 tile and accumulates into int32 — exact integer math, exactly like the AVX2 vpmaddubsw chain of doc 10, but 8× wider and issued by tensor-core silicon.
  • The float scales fold in afterwards: each accumulated integer tile gets multiplied by the matching weight/activation block scales once, outside the integer loop, before being added to the f32 output accumulator.

This is llama.cpp's structure, transplanted: docs/LLAMA-CPP-MMQ-ANALYSIS.md dissects the reference kernel instruction-by-instruction, and minfer's kernel grew against it round by round. The campaign record (docs/CUDA_OPTIMIZATION.md §0) shows the arc — the first parity-clean int8 MMQ measured 441 tok/s of 7B prefill; the promoted default path measures ~3,581 tok/s, 8.1× — with each lever measured in isolation:

Campaign step (docs/cuda_optimization_steps/)LeverEffect
R1 (step 08)first int8 MMQ GEMM, opt-inparity-clean but 8× off llama — the gap was unprofiled
r34 (step 37)quantize+transpose activations in the prepass (llama's quantize_mmq_q8_1 design) so the GEMM's A-staging is a bulk copy+9.72% whole prefill
r41 (step ~22 of r-series)q6_K "B-expand" plane widened to uint4 group loads — 32 per-byte loads per thread-k had become the stallq6_K kernel −61.5% time
r52RMSNorm/SwiGLU producers fuse the activation quantize (mode 2 skips the f32 output write entirely)+5.45%
r59/r60q4_K scale-pair plane + promotion: the verified gate set flips default-on+11.1%; final 1.080× vs llama.cpp

Two details of that table deserve a beginner's pause. First, the planes (W_exp, W_dsc): for q4_K/q6_K the kernels can either decode the packed sub-scales inside the hot loop or read precomputed f32 scale pairs staged at load time — trading ~3 GB of extra VRAM (the default path peaks at ~9.5 GB for 7B, vs ~20.5 GB for the legacy f16 escape, which is heavier, not lighter — CUDA_OPTIMIZATION.md §1.1) for removing a stall from the inner loop. Second, promotion (r60) is the campaign's exit ritual: a lever proves itself behind an opt-in env gate, the A/B record accumulates, and only then does the default flip — with "0" opt-outs kept so every step stays A/B-able forever.

And why does decode not want MMQ? At nt == 1 the GEMM degenerates to a matrix-vector product: there is one output row, so the 16-row M-tile of mma.m16n8k32 is 15/16 wasted, and the binding resource is weight bytes anyway. The decode path instead quantizes the single activation row to the same padded int8 format and runs MMVQ ("matrix-vector quantized"): one 256-thread block per weight row, integer dp4a (4-way int8 dot) over the row's blocks, one launch table per type. Measured: +74–77% on 7B shapes for q4_K (8e②), and the K-quant arms carry measured shape gates for when the integer path loses to the simpler f32-activation kernel (§3.2.5 shows the exact gates).

2.4 Split-KV decode attention: thousands of threads, few heads

Attention on the CPU (doc 11) parallelizes over heads. During prefill that is fine — the grid also spans tokens. But decode has one query token, so a per-head kernel for Qwen2.5-7B occupies 28 warps... on a GPU that fits hundreds. The rest of the machine idles while each warp serially walks the whole KV history.

Split-KV (the "split-K" family, flash-decoding style) adds a second parallel axis: cut the KV sequence dimension into ATTN_SPLITS = 32 stripes and launch one warp per (stripe, head) pair. Each warp computes a partial attention over its stripe only, and a second small kernel merges the 32 partials. Two pieces of math make the merge correct, and both are worth internalizing because they appear verbatim in the kernel excerpt of §3.2.7:

  1. Online softmax (the "flash attention" trick, already previewed in doc 11 §2.5). Instead of a first pass to find the row max, each new score s = q·k·scale updates a running max m and rescales everything seen so far: with corr = exp(m_old − m_new), the running sum S ← S·corr + exp(s − m_new) and the running output accumulator oc ← oc·corr + exp(s − m_new)·v. Nothing overflows, and one pass suffices.
  2. Split merging. A stripe's partial is exactly (m_sp, S_sp, oc_sp). The global max is gmx = max_sp m_sp; each partial's weight is w_sp = exp(m_sp − gmx); then S = Σ w_sp·S_sp and out = Σ w_sp·oc_sp / S. The rescaling makes the 32 partial sums combine into precisely the softmax the single-warp version would have computed — up to float addition order, which is a semantic difference here: merging 32 partials reorders the float sum, so split-KV output is not bitwise-identical to a one-stripe scan (the D1 record measured the drift at ~1e-9 magnitude). That is why ATTN_SPLITS is frozen at 32 rather than tuned per context length: the split grid is baked into captured CUDA Graphs (§2.5), and changing it would change numerics, not just speed.

The payoff was immediate (8d: 7B decode 10.1 → 13.7 tok/s, +36%), and the campaign kept refining the body: staging K and V rows into registers before the serial softmax chain (D2: kernel 34.1 → 19.4 µs/launch), a dim-parallel rewrite (R4: +10–15%), and a dual-kernel dispatch that swaps in a 4-warp body only when the per-warp row count is high enough to amortize it (D3-4: the same structure ran slower at short context — rows-per-warp pathology — so both kernels launch and each self-gates on the device-side context length, keeping CUDA-Graph replay valid).

2.5 CUDA Graph capture/replay: one launch per split

A decode step executes ~150–300 nodes — hundreds of kernel launches on the stream. Each launch carries a few microseconds of host-side submission overhead, and at decode's ~50 tok/s pace the overhead is a measurable tax (7d measured +18% decode on 0.5B from removing it).

A CUDA Graph is a recorded, replayable bundle of stream work. Capture the sequence of launches once, and every later execution of the identical sequence becomes a single cudaGraphLaunch. That is a huge if — "identical" means identical kernel parameters, which includes identical device pointers. The whole capture machinery in cuda_backend.rs exists to keep that promise:

  • The graph's node buffers live in the backend's device pool and are reused across decode steps (doc 07's allocator never frees while the graph is alive), so every kernel keeps the same address arguments step after step.
  • Anything address-affecting bumps a pool_gen counter, and captured graphs keyed to an older generation are destroyed and re-captured (§3.2.8 shows the invalidation branch).
  • No host decisions inside the window: the causal attention bound is read from the device-side positions buffer inside the kernel (the Attn arm's comment calls this out as a precondition for replay), and input data is H2D-copied into the stable staging addresses before the split, so replay reads fresh data through fixed pointers.
  • Capture uses llama.cpp's warmup protocol: executions 1 and 2 of a split run direct launches (warming up per-kernel state), the 3rd opens a capture window, and every execution after that replays. One-shot graphs — a CLI prefill that runs once — never reach 3 runs and never pay capture cost. Repeated identical-length prefills (the server's slot scenario) do capture, by default since R3-B (MINFER_NO_PREFILL_CAPTURE=1 opts out).

The replay hook lives in the scheduler (doc 08's split loop), not in the backend alone: before executing a CUDA split the scheduler asks graph_replay(uid, node_range, nt_hint); a true answer means the captured graph covers the whole node loop and the loop is skipped. Tracing/viz capture forcibly disables replay — per-node host readbacks inside a capture window are exactly the "host decisions inside the window" that break it.

2.6 What the GPU doesn't change

It is worth stating the invariants explicitly, because the series has spent ten documents building them and the CUDA backend preserves every one:

  • KV positions are data, not structure — the attention kernel derives its causal bound from the device positions buffer at run time.
  • Weights are the GGUF bytes — registered to the device once at load; execution never host-copies a weight (llama.cpp's "ops follow their weights" rule, CUDA-BACKEND-DESIGN.md §3).
  • Backend assignment is a build-time decision — supports_op + the all-weights gate decide placement before the first forward; a mid-run invariant violation is an Err, never a silent CPU detour (§3.4).
  • CPU-vs-GPU logits differ by design — the GPU path uses f32 activations (int8 only inside the MMQ GEMM), so parity gates compare each path against its own reference plus greedy-token equality, never cross-path bitwise.

One thing does change on the GPU: the KV cache element type. The device KV regions may be f16 — halving attention's read bandwidth — auto-selected for models where KV streaming dominates decode (n_layers × n_kv_embd ≥ 8192, so 7B and up; MINFER_CACHE_TYPE overrides), mirroring the Metal policy (doc 14). Store and attention come in matching f32/f16 kernel pairs; doc 11's invariants (positions are data, windows are pos[t]+1, store-before-attention) are untouched.

3. Implementation

3.1 Data in / data out

The backend's whole world is device pointers. Every buffer the allocator hands it (doc 07) is an id into its device pool; every weight is a name resolved to a device address; the KV regions are pool slots that simply never rejoin the free list.

ItemLayoutWhere it lives
Node buffers (activations)[nt][d] f32, token-major — 4 bytes/elem device buffersCudaBackend.pool: Vec<CudaBuf> (cudaMalloc'd; free-list recycled)
Weightsraw quantized bytes exactly as the GGUF has them ([out][in] row-major, block layouts of doc 02); Q6_K repacked to 224-byte block slots; f32 norms/biases as-isCudaState.weights: HashMap<String, (CudaPtr, usize)> — registered once at load
MMQ activation scratch[nt][id/32] × 40 B padded int8 blocks (+ transposed variant for the raw kernels)CudaState grow-on-demand scratch buffers (buf_q8_prefill, buf_q8_decode)
MMQ precomputed planesq6_K expanded-B / scale planes, q4_K f32 scale-pair planes (~3 GB total on 7B)keyed by weight device pointer in CudaState maps
KV regionsper layer, [n_ctx][nkt] f32 or f16 (auto-selected, §2.6)the pool's persistent regions (kv_pair(layer), doc 07)
Positions[nt] I32 — arriving as f32::from_bits bit patterns (doc 07), decoded on device to raw int32 by a tiny kernelpos_scratch, memoized per execution window
Host→device input fillsRust &[f32] → pinned staging slot → cudaMemcpyAsync H2D8 × 2 MiB cudaHostAlloc ring (staging)
Device→host readbacklogits (and viz/trace node dumps) D2H through a pinned readback bufferreadback / CaptureStaging (128 MB ceiling)

Two asymmetries against the CPU backend are worth noticing. First, the CPU matmul quantizes activations per call into a Vec<u8> (doc 10); the CUDA backend keeps them in persistent device scratch that grows on demand and is memoized per execution window — because a cudaMalloc in the hot path would sync the device and because decode re-quantizes the same producer output every step. Second, read_host returns None unconditionally: a staged D2H transfer cannot hand back a borrowed slice, so host reads go through copy_to_host (the allocator's copy_to_cpu CUDA arm) instead.

3.2 Key code

3.2.1 The build chain: what --features cuda (and cuda_static) actually do

build.rs only runs the CUDA section when the cargo feature is set — plain builds never touch nvcc. With the feature requested, CUDA is required: a missing toolkit is a hard error rather than a silent CPU-only binary, because src/cuda.rs declares launch_* symbols that only the compiled kernel archive can satisfy.

#![allow(unused)]
fn main() {
// build.rs:182-241 (abridged)
let nvcc = match find_nvcc() {
    Some(n) => n,
    None => panic!(
        "CUDA feature requested but nvcc was not found — install the CUDA \
         toolkit or point CUDA_HOME at its root (e.g. /usr/local/cuda)"
    ),
};
...
// nvcc inherits the first cc/g++ on PATH as its host compiler and
// hard-fails when that is newer than the toolkit supports. Keep nvcc's
// own default whenever it works ...; only pin -ccbin when the default is
// rejected.
let (ccbin, ccbin_label) = match std::env::var("MINFER_CUDA_CCBIN") {
    Ok(v) if !v.is_empty() => (Some(v), format!("MINFER_CUDA_CCBIN={v}")),
    _ => match detect_host_compiler(&nvcc, &out_dir, &include_flag) {
        Some(HostCompiler::Default) => (None, "nvcc default".to_string()),
        Some(HostCompiler::Pinned(c)) => {
            println!("cargo:warning=CUDA: nvcc default host compiler rejected, \
                      pinning -ccbin {c}");
            (Some(c.clone()), format!("-ccbin {c}"))
        }
        None => panic!(...),
    },
};
let archs = detect_archs(&nvcc, &out_dir, &include_flag, ccbin.as_deref());
...
let cudart_static = std::env::var_os("CARGO_FEATURE_CUDA_STATIC").is_some();
}

Read the three decisions:

  • Host compiler pinning (-ccbin): nvcc compiles host C++ too, and it rejects host compilers newer than the toolkit supports (a nix devShell putting GCC 15 first breaks a CUDA 13 that accepts ≤ GCC 13). The build probes nvcc's default first and only pins an older GCC when that probe fails; MINFER_CUDA_CCBIN forces one explicitly.
  • GPU arch coverage: detect_archs tries sm_70…sm_121 and keeps what the toolkit accepts — SASS (native machine code) per supported architecture plus PTX for the highest, so one binary covers older and newer GPUs (older ones JIT the PTX forward). The floor is sm_70/Volta because the f16 wmma kernels require tensor cores.
  • cuda_static: the plain feature links cudart shared (the binary needs libcudart.so.N and a CUDA toolkit runtime at run time, with an rpath baked in). --features cuda,cuda_static links libcudart_static.a instead: the binary has no libcudart.so dependency — it needs only the NVIDIA driver libcuda.so.1 (never a link-time dependency; dlopen'd at run time, with a preload_driver helper that resolves it from well-known paths because nix shells bypass /etc/ld.so.cache) plus libstdc++. That is the deployment story: build on a toolkit machine, copy the binary to a driver-only machine.

3.2.2 Device init and the weight registry

CudaState::try_new probes devices, honors --gpu N (falling back to auto-selecting the highest compute capability), creates the one shared stream, queries the device properties at run time — SM count, compute capability, free/total memory, name — and prints the banner you see at startup (CUDA: using ... (SM 12.1, ... MB, ... SMs)). Two details are load-bearing beyond the boilerplate: the compute capability is stored as an integer (major*100 + minor) and feeds the device-tier resolution (src/device_tier.rs, docs 105–106) — the resolved tier's MMQ availability plus a dynamic-smem feasibility check gate the int8 MMQ path (mmq_active(); the fallback rule for unknown devices is still cc >= 800, i.e. sm_80+, because mma.m16n8k32 exists only from Ampere on); and gemm_prefill_smem_init() does not run any more: since #218 the prefill GEMM opts into

48 KB dynamic shared memory lazily, on an instantiation's first launch (gemm_smem_optin, reached through launch_gemm_f16), and caches the answer per instantiation. What keeps that attribute out of a stream-capture window is the 3-run capture warmup (capture_warmup) plus cudaStreamCaptureModeThreadLocal; the reasoning and the gates that pin it are in docs/CUDA-BACKEND-DESIGN.md §2.4.

Weights register through register_weight — a cudaMalloc plus one stream-ordered H2D copy of the raw GGUF bytes. Since #188 that copy is cudaMemcpyAsync on the context stream followed by a cudaStreamSynchronize, deliberately not the legacy-null-stream blocking cudaMemcpy: the blocking form is not a stream operation at all, so while another thread holds a capture window open on a different stream it participates in the legacy default stream's implicit global synchronization — exactly the call that invalidated the capture before #188.

#![allow(unused)]
fn main() {
// src/cuda/methods/weights.rs:39-87 (core of register_weight)
let mut ptr: *mut std::ffi::c_void = std::ptr::null_mut();
let err = unsafe { cudaMalloc(&mut ptr, data.len()) };
if err != 0 || ptr.is_null() {
    eprintln!("CUDA: failed to allocate {} bytes for '{}'", data.len(), name);
    return;
}
let err = unsafe {
    // Issue #188: stream-ordered, not the legacy-null-stream blocking
    // `cudaMemcpy`. ... Queuing the copy on the context's own stream
    // and waiting on that stream keeps the registration a bounded,
    // stream-scoped operation.
    cudaMemcpyAsync(
        ptr,
        data.as_ptr() as *const std::ffi::c_void,
        data.len(),
        CUDA_MEMCPY_HOST_TO_DEVICE,
        self.context_stream(),
    )
};
if err == 0 {
    let serr = unsafe { cudaStreamSynchronize(self.context_stream()) };
    // ... a non-zero `serr` names the error and frees the buffer
}
if err != 0 {
    eprintln!("CUDA: failed to copy '{}' to device", name);
    unsafe { cudaFree(ptr) };
    return;
}
// a plain (unpadded) registration must clear any stale padded flag
// for the same name ...
self.padded_weights.lock().unwrap().remove(name);
self.weights.lock().unwrap()
    .insert(name.to_string(), (CudaPtr(ptr), data.len()));
}

The pre-#188 blocking form survives only as the test-only register_weight_blocking_legacy probe (weights.rs:99-131), which the capture acceptance probe needs in order to measure the capture mode against the historical setup.

The comment lines this excerpt elides contain two ownership rules that matter: a re-registration with the same name and size reuses the existing device copy (unit tests reload the same file; weights are immutable, so there is nothing to refresh), while a different size replaces the entry and deliberately leaks the stale buffer — a live captured CUDA Graph may still reference the old address, and the leak is bounded by the number of distinct tensor shapes ever loaded. Device memory ownership is serious enough that this one rule got its own bullet in docs/GPU_SAFETY.md.

Special layouts register through special entries. The important one is Q6_K: the raw 210-byte block stride forces 1-byte-per-instruction weight loads that capped 7B decode near ~38 GB/s, so the loader repacks each block into a 224-byte slot (a multiple of 16) and the kernels stream with aligned uint4 loads:

#![allow(unused)]
fn main() {
// src/models/qwen2/loader.rs:234-244
if ttype == TensorType::Q6_K {
    // 7e②: register Q6_K in the padded 224-byte block layout so
    // the matmul kernel can use aligned uint4 weight loads
    // (the raw 210-byte stride forces 1-byte-per-instruction
    // reads and caps 7B decode near ~38 GB/s).
    cuda.register_weight_q6k_padded(
        &ti.name,
        tensor.data(),
        tensor.shape[1] as usize,
        tensor.shape[0] as usize,
    );
}
}

That one repack was a 3.1× decode win (8.4 → 26.4 tok/s on 7B — the 7e② record), and it made has_weight_of_size match padded entries by their original raw length so the participation gate (below) stays honest.

Also at load, the loaders flip global policy switches that must be settled before the first forward: set_kv_cache_type auto-selects the f16 KV policy from model dims, and a non-q4_K/q6_K quantized weight (or any 2-D f32 matmul weight) clears the nb_bt_only flag, which degrades the skip-write fused producers of §3.2.6 to a safe mode — the flag exists so a mixed-quant model can never feed a dead buffer to a GEMM that reads f32 activations.

3.2.3 Build-time assignment: supports_op and the all-weights gate

The CUDA backend claims almost the whole per-layer chain — everything except the ops it has no kernel for (a standalone Softmax node; the Scale node) and RoPE in a layout it does not implement:

#![allow(unused)]
fn main() {
// src/graph/cuda_backend.rs:1259-1300 (abridged to the decision arms)
fn supports_op(&self, op: &Op, dtype: DType) -> bool {
    if dtype != DType::F32 {
        return false;
    }
    match op {
        Op::Input
        | Op::Add | Op::Mul | Op::Silu | Op::SwiGLU
        | Op::RmsNorm { .. } | Op::QkNorm { .. }
        | Op::MatMul { .. } | Op::Attn { .. }
        | Op::KvcacheStore { .. } | Op::KvcacheLoad { .. }
        | Op::View { .. } | Op::Reshape { .. } | Op::Permute { .. }
        | Op::GetRows                     // 7e③: embed + tail gather on device
        | Op::FusedQKV { .. }             // D3-8: decode QKV fusion (G4 port)
        | Op::QkvBiasRopeStore { .. }     // D3-8: mixed-quant QKV epilogue
        | Op::FusedFFN => true,           // 7e⑤: decode FFN fusion
        // MatMul ttype gating happens at the model level (weights must all
        // be registered on CUDA — same all-or-nothing rule as Metal).
        Op::RoPE { style } => matches!(style, RopeStyle::NonInterleaved),
        _ => false,
    }
}
}

The allocator's supports() walks priority Metal → CUDA → CPU, so on a CUDA build every claimed node lands in a single CUDA split per graph (the 7e③ record: embedding and the G3 tail gather moved on device precisely to eliminate the last cross-backend copies — a prefill/decode graph is now one CUDA split, no host round trips at all).

But per-op support is not sufficient: a graph with one weight missing from the device registry would interleave CPU and CUDA execution against persistent KV regions. So the model wiring applies an all-or-nothing gate before the graph is even built:

#![allow(unused)]
fn main() {
// src/models/qwen2/graph.rs:436-457 (abridged)
// CUDA participation (Phase 7): requires a usable device AND every
// matmul weight registered on the CUDA registry in a kernel-supported
// type (all-or-nothing; 7e③ moved the embedding gather on device, so
// tok_embd is gated like every other weight).
#[cfg(feature = "cuda")]
let cuda_on = crate::cuda::CudaState::get().is_some() && Self::weights_on_cuda(model);
...
cparams: CParams {
    n_ctx,
    flash_attn: false,
    gpu: metal_on || cuda_on,
    ...
},
}

weights_on_cuda (qwen2/graph.rs:881) walks every weight the graph reads — embedding, output head, every layer's norms/biases/matmul weights — and requires each to be (a) registered and (b) a type with a matching kernel. The type check is per role: matmul weights admit all eight quant types plus f32, while the embedding gather lacks a Q4_1 kernel, so a Q4_1 tok_embd keeps the whole model on CPU. When the gate fails it names the first offending tensor in the log rather than emitting a generic complaint. The boolean result is recorded in CParams.gpu, which makes GPU participation part of the reuse identity (doc 13): flipping CUDA on or off deterministically changes the built graph, so a stale reuse can never mix assignments.

3.2.4 execute_node: one match, many kernels

The scheduler calls execute_node(node, in_bufs, out_buf, kv_pair) per node. The CUDA implementation is a single match over the op enum; each arm resolves device pointers (failing loudly if any is missing), checks the kernel's structural invariants, and launches. The wrapper adds one capture-window duty:

#![allow(unused)]
fn main() {
// src/graph/cuda_backend.rs:1354-1375
fn execute_node(
    &mut self,
    node: &CNode,
    in_bufs: &[usize],
    out_buf: usize,
    kv_pair: Option<(usize, usize)>,
) -> Result<(), String> {
    match self.execute_node_inner(node, in_bufs, out_buf, kv_pair) {
        Ok(()) => Ok(()),
        Err(e) => {
            // A node error during an open capture window dooms the window:
            // the scheduler propagates before the boundary sync, so nothing
            // would close it — later input fills would be RECORDED into the
            // window and the eventual close would cache a multi-step graph
            // (double KV commit on every replay). Abort the window loudly.
            if self.capturing.is_some() {
                self.abort_capture(&e);
            }
            Err(e)
        }
    }
}
}

A tour of the arms, in execution order per layer, tells you what actually runs on the GPU:

  • Input / KvcacheLoad — no kernel. Inputs were H2D-filled by the allocator before the split (§3.2.9); a KV load is a view of the persistent region (out_buf is the region).
  • View/Reshape/Permute — a device-to-device copy (layout-only nodes keep the CPU backend's identity-copy semantics).
  • GetRows — the embedding gather (dequantize-on-gather, type-dispatched on device, ids read from the I32-as-f32 buffer) or the generic f32 tail gather for the n_out reduction (doc 05).
  • RmsNorm/QkNorm — the float4 RMSNorm kernel; decode producers fuse an int8 quantize epilogue (the n == 1 branch at the bottom of the arm) so the following MMVQ group skips its standalone quantize launch.
  • MatMul — the decision tree of §3.2.5.
  • RoPE — neox-style rotate in place (a D2D copy first if the allocator did not alias input and output); positions decoded from the f32-bits buffer by positions_i32.
  • KvcacheStore — scatter k/v rows into the persistent regions at the positions; f32 or f16 destination per the KV policy.
  • Attn — the split-KV decode kernel (nt == 1) or the prefill attention kernel, §3.2.7.
  • FusedQKV / QkvBiasRopeStore / FusedFFN — the decode fusions (doc 06): one concat matmul + one fused bias/rope/store (or offset-swiglu) launch, decode-only (nt != 1 returns Err — the kernels are not shaped for prefill), replacing 3 matmuls + bias×3 + rope×2 + store×2 launches with two.
  • The catch-all arm (line 1156) returns Err("cuda: op ... has no kernel (stays on the CPU backend per supports_op)") — a deferral that can only fire if supports_op and the dispatcher drift apart, which is exactly the loud-abort behavior the safety contract wants.

3.2.5 The matmul decision tree

Both prefill and decode funnels converge on CudaState::matmul_f32_ptr_layout(weight_ptr, ttype, x, out, od, id, nt, padded_q6k). Its opening gate you saw in §2.2 (nt >= 9 → MMQ or f16 GEMM). What falls through is the decode-side dispatch, and reading one arm teaches you the shape of all of them:

#![allow(unused)]
fn main() {
// src/cuda/methods/dispatch.rs:312 (Q4_K arm; abridged comment)
TensorType::Q4_K => {
    // 8e-reversal: decode (nt == 1) runs the MMVQ structure
    // (dp4a over q8 activations, one row per 256-thread block) —
    // +74–77% at 7B shapes (bench8e2); id >= 2048 gate (below
    // that it is launch-latency noise), id % 32 == 0 for the
    // sub-block tail granularity. Prefill keeps the f32 kernel.
    if nt == 1 && id >= 2048 && id % 32 == 0 {
        self.q4_k_decode_mmvq(wptr, x, out, od, id, nt);
        Ok(())
    } else if nt >= 2 && nt <= 8 && id % 32 == 0 {
        // Step 82: multi-token MMVQ (in-block token loop).
        self.q4_k_decode_mmvq_multi(wptr, x, out, od, id, nt);
        Ok(())
    } else {
        launch!(launch_q4_k_f32_matmul)
    }
}
}

Three things to take from this arm:

  1. The shape gates are measured, not guessed. Below id = 2048 the kernel-count win is lost in launch latency; the q5_K/q6_K arms carry an even sharper measured crossover (od*id ≥ 24M for q5_K, lowered to 4M for q6_K's attn_v class) — under the crossover the simpler coalesced f32-activation kernel wins because the MMVQ byte loads are uncoalesced and 1–2 units per thread expose their latency (the arm comments cite the on-device micro-bench numbers).
  2. The activation quantize is shared machinery. q4_k_decode_mmvq calls decode_quantize_native, which consults the MmqCache first — when the preceding fused RMSNorm/SwiGLU already wrote the padded int8 plane for this exact source buffer, the standalone quantize launch is skipped entirely (D3-5 1a: standalone quantize launches dropped 78% at 14B).
  3. Every arm ends in a launch or an Ok(()) — the dispatch never returns "unsupported" for a type supports_op admitted; the weights_on_cuda gate guaranteed the type exists here.

The full tree, in one table:

ConditionPath
nt ≥ 9, quant type, MMQ active (sm_80+, MINFER_MMQ on)prefill_mmq — int8 tensor-core GEMM (§3.2.6)
nt ≥ 9, quant type, MMQ offprefill_gemm_f16 — dequant weights to an f16 scratch, f16 wmma GEMM (the pre-campaign path; also what MINFER_MMQ=0 escapes to)
nt == 1, per-type shape gate passesMMVQ dp4a kernel (q4_K id ≥ 2048; q5_K od·id ≥ 24M; q6_K od·id ≥ 4M)
nt 2–8multi-token MMVQ (weights-once token loop)
everything elseper-type f32-activation kernels (the original Phase-7 family; every supported type has one)

3.2.6 Inside prefill_mmq: the int8 GEMM and its fallback ladder

prefill_mmq maps the quant type to a small integer id, pre-checks the activation scratch for OOM (so a failure surfaces before any launch), picks the Q6_K block stride (224 padded vs 210 raw), and then walks a ladder of increasingly general kernels — each specialized variant tries first and each failure falls through cleanly:

#![allow(unused)]
fn main() {
// src/cuda/methods/dispatch.rs:289-364 (the q4_K raw-byte branch, abridged)
// P6: raw-byte staging variant (q4_K, whole super-blocks only).
// Same quantized activations; the GEMM stages RAW weight bytes via
// cp.async and dequants in registers (docs/CUDA_OPTIMIZATION.md).
if type_id == 5 && Self::mmq_gate_on("MINFER_MMQ_RAW") && (id / 32) % 8 == 0 {
    ...
    // P6 r34: relocate the A-side layout transform out of the mma
    // kernel into a quantize-transpose prepass (llama.cpp's design).
    ...
    let (qa8g, sdag) = self.mmq_quantize_transposed(
        x as *const f32, id as i32, nt as i32, nchunk, ntb, stream,
    );
    // r59: the q4_K W_dsc f32-pair plane (null on miss ->
    // the DSC=false in-kernel scalar decode instantiation).
    let w_dsc = self.q4k_dsc.lock().unwrap()
        .get(&(wptr as usize)).map(|cp| cp.0)
        .unwrap_or(std::ptr::null_mut());
    nb_ok = qa8g != 0 && sdag != 0
        && launch_mmq_raw_nb_bt_nt(
            type_id, wptr as *const u8, w_dsc as *const u8,
            qa8g as *const u8, sdag as *const u8,
            out as *mut f32, nt as i32, od as i32, id as i32,
            nchunk, stream, kd,
        ) == 1;
    ...
}
}

The vocabulary, translated: A is the activation matrix, B the quantized weight matrix; "raw NB-BT" is the fastest kernel family (raw weight bytes staged by cp.async — the hardware copy engine — into shared memory, nibbles unpacked in registers, mma per k-chunk); the transposed prepass (mmq_quantize_transposed) emits the int8 activations already laid out the way the GEMM's staging loop wants them, removing the layout transform from the hot kernel (+9.72% whole prefill); and the W_dsc plane is the load-time precomputed f32 scale-pair table that replaces in-kernel sub-scale decoding (+11.1%). Each plane lookup is null on miss, and the launcher returns 0 when its preconditions (shared-memory caps, geometry gates) fail — which is how the ladder degrades to the generic launch_mmq_nt fallback at the bottom:

#![allow(unused)]
fn main() {
// src/cuda/methods/prefill_mmq.rs:466 (generic tail)
unsafe {
    let q8 = self.mmq_quantize_native(x as *const f32, id as i32, nt as i32, stream);
    if q8 == 0 {
        return Err("cuda: prefill MMQ q8 scratch OOM".to_string());
    }
    launch_mmq_nt(
        type_id,
        wptr as *const u8,
        q8 as *const u8,
        out as *mut f32,
        nt as i32,
        od as i32,
        id as i32,
        block_stride,
        stream,
    );
}
Ok(())
}

Note the one path that is not a clean fallback: when the A-quantize helper returns 0 because the r52 skip-write guard refused to re-quantize a dead buffer, prefill_mmq returns Err (§3.4). Perf fallbacks are safe — correctness fallbacks are not, and the code keeps the distinction visible.

3.2.7 Split-KV attention, the kernel

The host side (gqa_attn_split, src/cuda/methods/attention.rs:397) computes the partial-row stride pstr = (4 + hd + 3) & !3 (running max, running sum, then the hd-wide output accumulator, rounded to a 16-byte boundary for the float4 writes), grows the partials scratch once (a fixed [32][nh][pstr] slab — nh/hd are graph constants, so it never grows inside a capture window), and launches two kernels on the stream: the partial kernel and the combine.

The partial kernel's grid is (ATTN_SPLITS, n_head) — 32 stripes × heads — with one warp (32 threads) per block. Each lane owns 4 consecutive head dimensions (a 128-dim head fills all 32 lanes × 4 = 128 dims exactly; the dispatch rejects hd > 128 or hd % 4 != 0 before launch). The core loop is the online softmax of §2.4:

#![allow(unused)]
fn main() {
// src/cuda/kernels/attention_decode.cu:185 (attn_split_1w_body core loop)
for (int base = lo; base < hi; base += 4) {
    int nr = min(4, hi - base); // warp-uniform
    // D2: stage BOTH K and V for the whole 4-row window before the first
    // softmax step. All 8 row loads then issue back-to-back and their
    // latency overlaps the serial chain; ...
    float4 k4[4], v4[4];
    pragma unroll
    for (int j = 0; j < 4; j++) {
        k4[j] = (live && j < nr)
            ? kv_ld4<KV>(k + (size_t)(base + j) * stride_kv + hk * hd + d0)
            : make_float4(0.0f, 0.0f, 0.0f, 0.0f);
        v4[j] = (live && j < nr)
            ? kv_ld4<KV>(v + (size_t)(base + j) * stride_kv + hk * hd + d0)
            : make_float4(0.0f, 0.0f, 0.0f, 0.0f);
    }
    pragma unroll
    for (int j = 0; j < 4; j++) {
        if (j >= nr) break; // warp-uniform: all lanes exit together
        // Full-row dot: this lane's 4-dim partial, then a warp reduction
        // so every lane holds the row's complete dot (uniform softmax).
        float d = q4.x * k4[j].x + q4.y * k4[j].y
                + q4.z * k4[j].z + q4.w * k4[j].w;
        pragma unroll
        for (int off = 16; off > 0; off >>= 1)
            d += __shfl_xor_sync(0xFFFFFFFF, d, off);
        float s = d * scale;
        float nmx = fmaxf(mx, s);
        float corr = expf(mx - nmx);
        float e = expf(s - nmx);
        S = S * corr + e;
        mx = nmx;
        if (live) {
            float4 vv = v4[j];
            oc.x = oc.x * corr + e * vv.x;
            oc.y = oc.y * corr + e * vv.y;
            oc.z = oc.z * corr + e * vv.z;
            oc.w = oc.w * corr + e * vv.w;
        }
    }
}
}

Walk it as a beginner: each warp iteration covers 4 KV rows. The warp first stages 4 K rows and 4 V rows into registers (memory latency overlaps the following serial math — the D2 record measured −42% on cold DRAM from exactly this scheduling). Then per row: every lane multiplies its 4 query dims by its 4 key dims, __shfl_xor_sync butterflies the 32 partials into a full row dot (5 shuffle rounds), the online-softmax rescale updates mx, S, and the 4-dim output accumulator oc. Each lane owns distinct output dims, so there is no cross-lane reduction at the end — lane i simply writes its 4 accumulator dims into the stripe's partial. Stripes beyond the KV end write an empty partial (mx = -INF, S = 0), which the combine weights to zero.

The combine kernel is 20 lines and re-derives the whole answer:

#![allow(unused)]
fn main() {
// src/cuda/kernels/attention_decode.cu:360
__global__ void gqa_attn_split_combine(
    const float* __restrict__ partial,
    float* __restrict__ o,
    int nh, int hd, int pstr
) {
    int h = blockIdx.y;
    int i = threadIdx.x; // hd threads
    if (i >= hd) return;
    float gmx = -INFINITY;
    for (int sp = 0; sp < ATTN_SPLITS; sp++)
        gmx = fmaxf(gmx, partial[((size_t)sp * nh + h) * pstr]);
    float S = 0.0f, acc = 0.0f;
    for (int sp = 0; sp < ATTN_SPLITS; sp++) {
        const float* p = partial + ((size_t)sp * nh + h) * pstr;
        float w = expf(p[0] - gmx);
        S += p[1] * w;
        acc += p[4 + i] * w;
    }
    o[h * hd + i] = (S > 0.0f) ? acc / S : 0.0f;
}
}

One thread per output dim; 32 reads of the partials slab; the §2.4 merge math verbatim. (For f16-KV, hd==128 models at long context, a second 4-warp body exists and both kernels launch with static grids, each self-gating on the device-side context length — the D3-4 record's answer to a dispatch that must stay replay-safe.)

Two design consequences to keep: the grid is static — no kernel reads n_past on the host to size the launch, so capturing this launch in a CUDA Graph and replaying it at any later context length is safe (the kernel re-derives everything from positions[0] per replay); and the partials scratch is size-stable, so its address never churns under a captured graph.

3.2.8 CUDA Graph capture/replay, the state machine

graph_replay_step is called by the scheduler before each CUDA split executes (scheduler.rs:201) and returns whether replay covered the whole node loop. Three states live in the backend: graph_execs (the captured, instantiated graphs keyed by (uid, range)), graph_runs (the warmup counter per key), and capturing (an open capture window, closed by synchronize):

#![allow(unused)]
fn main() {
// src/graph/cuda_backend.rs:199-250 (abridged)
let key = (uid, range);
if let Some(pos) = self.graph_execs
    .iter().position(|g| g.uid == uid && g.range == range)
{
    if self.graph_execs[pos].pool_gen != self.pool_gen {
        // pool churned since capture — pointers may differ, re-capture
        let g = self.graph_execs.remove(pos);
        self.graph_runs.remove(&key);
        self.state.graph_destroy(g.exec);
    } else {
        let exec = self.graph_execs[pos].exec;
        // a plain stream launch — serialized like any other stream op
        let _sg = self.stream_guard();
        if self.state.graph_launch_exec(exec) {
            return true;
        }
        eprintln!("CUDA: graph replay launch failed; graphs disabled for this session");
        self.graphs_mode = GraphMode::Disabled;
        return false;
    }
}
let runs = self.graph_runs.entry(key).or_insert(0);
*runs += 1;
// 8g①: capture decode-shaped graphs by default ... Prefill-shaped graphs
// (nt > 1) capture only with the prefill_capture gate: 8g② made
// that a deliberate opt-in after an audit caught unvalidated
// capture; R3-B (2026-08-31) flips the default ON — the 3-run
// protocol bounds the cost ...
if *runs >= 3
    && self.capturing.is_none()
    && nt_hint.map_or(true, |nt| nt == 1 || self.prefill_capture)
{
    // Hold the process-wide stream lock across the capture window:
    // any other backend's stream work would otherwise be recorded
    // into this graph (capture is per-stream, not per-thread).
    let guard = self.state.stream_lock().lock().unwrap();
    if self.state.graph_begin_capture() {
        self.capturing = Some(key);
        self.stream_guard = Some(guard);
    } else {
        drop(guard);
        eprintln!("CUDA: stream capture unavailable; graphs disabled for this session");
        self.graphs_mode = GraphMode::Disabled;
    }
}
false
}

Every line answers a "what could go wrong":

  • Pool churn (pool_gen mismatch): any buffer allocation can move pointers; the exec keyed to the old generation is destroyed and the next executions re-warm and re-capture. This is the pointer-stability promise of §2.5 enforced mechanically.
  • Replay failure: logs, disables graphs for the session, returns false — the split then executes by direct launch, still correct.
  • The 3-run protocol: executions 1–2 direct-launch (llama.cpp's warmup ×2), the 3rd opens the window. nt_hint (the first MatMul node's token count, graph/mod.rs:127) distinguishes decode-shaped from prefill-shaped graphs so the prefill capture gate applies only where it was validated.
  • The process-wide stream lock: stream capture is per-stream; while one backend holds an open window, every other backend's stream work (the CPU-side fills of a mixed split boundary, another model slot's launches) must block instead of being recorded into the graph. The capturing backend holds the mutex across its window and its own enqueues skip re-locking.

Closing the window happens in synchronize → close_capture_or_sync (cuda_backend.rs:268): end capture, instantiate, launch once so the step still produces output, cache the exec, and only then release the stream lock. If instantiate/launch fails, the recorded launches never executed — this step's outputs are undefined, and the code says so out loud (graphs disabled for the session; rerun with MINFER_NO_CUDA_GRAPH=1), instead of pretending the step succeeded.

One subtlety ties the whole section together: input staging. Replay reads whatever is in the recorded input addresses at launch time, so the allocator's fill_input H2D copies happen before graph_replay is called for the split — fresh token data lands at the same addresses every step, which is the invariant that makes one captured graph serve every decode step.

3.2.9 The buffer pool, pinned staging, and the sync contract

The pool mirrors Metal's (doc 14): a Vec<CudaBuf> plus a byte-length-matched free list. Two details are CUDA-specific:

#![allow(unused)]
fn main() {
// src/graph/cuda_backend.rs:1307-1327 (alloc_buffer)
fn alloc_buffer(&mut self, size: usize) -> usize {
    let _sg = self.stream_guard(); // cudaMalloc syncs the device
    let bytes = size * 4;
    if let Some(pos) = self.free.iter()
        .position(|&id| self.pool[id].bytes == bytes)
    {
        let id = self.free.remove(pos);
        self.pool_gen += 1;
        return id;
    }
    // On OOM, cuda_malloc logs and returns null; the null buffer fails
    // cleanly (Err) at execute time via ptr_of — do NOT panic here: the
    // backend may be holding the process-wide stream lock, and panicking
    // under a mutex poisons it for every other user.
    let ptr = <crate::cuda::CudaState>::cuda_malloc(bytes);
    self.pool.push(CudaBuf { ptr, bytes });
    self.pool_gen += 1;
    self.pool.len() - 1
}
}

Both the free-list reuse and the fresh cudaMalloc bump pool_gen — the replay-invalidations trigger of §3.2.8. free_buffer recycles (never cudaFrees) so persistent KV regions survive rebuilds, and teardown is Drop's job. And the OOM comment is worth a second read: allocating under the stream lock means panicking would poison the mutex for every other backend user, so OOM is carried as a null pointer and converted to Err at first use.

Host transfers avoid the pageable-memory penalty with pinned staging. Input fills (write_host) copy into a slot of a lazily-allocated ring of 8 × 2 MiB cudaHostAlloc buffers and enqueue cudaMemcpyAsync — returning before the copy lands, which is race-free because everything consumer-side runs later on the same stream (GPU_SAFETY rule 5). Logits readback goes through a pinned readback buffer for the same reason in reverse: a blocking cudaMemcpy into pageable memory bounces through a driver-internal pinned buffer (R3-A2). Viz/trace node dumps queue async D2H copies into a 128 MB CaptureStaging arena and drain with one sync at the split boundary — replacing what would otherwise be a per-node full-stream sync.

synchronize (called by the scheduler at split boundaries) closes the capture window if one is open — otherwise it plain-syncs — and also clears the two execution-window memos (the MMQ A-quantize cache and the positions i32 conversion), because at a boundary the pool reuses buffer ids for different data and a stale memo would alias old content onto a new node. The sync itself is bounded and checked, per GPU_SAFETY:

#![allow(unused)]
fn main() {
// src/cuda/methods/weights.rs:125-134
pub fn sync(&self) {
    let err = unsafe { cudaGetLastError() };
    if err != 0 {
        eprintln!("CUDA kernel launch error: {}", err);
    }
    let err = unsafe { cudaStreamSynchronize(self.stream()) };
    if err != 0 {
        eprintln!("CUDA stream sync error: {}", err);
    }
}
}

Launch errors are checked here, at sync points, rather than after every launch (rule 4) — the stream serializes everything, so one checked sync after a batch of launches observes all of their errors.

3.3 Design choices (why this shape and not another)

Why two layers (a CudaState singleton wrapped by a CudaBackend)? cuda.rs predates the graph (it began as a direct-inference device layer) and CUDA-BACKEND-DESIGN.md §2.3 made the call explicit: wrap, do not rewrite. The singleton owns everything device-global (the stream, the weight registry, the KV regions, staging pools) and is shared by tests and the legacy surface; the backend owns everything graph-shaped (the buffer pool, capture state machine, per-node dispatch). The benefit shows at the seams: the model loader talks only to CudaState (register weights), the scheduler talks only to the Backend trait, and neither sees the other.

Why hand-written kernels instead of cuBLAS? cuBLAS has no quantized-weight GEMM entry point that consumes llama.cpp block layouts, so the quantized matmuls — the whole point of the engine — would need manual dequantization into f16/f32 scratch anyway (the pre-8m path did exactly that, at 30.7 tok/s). The campaign's answer was to write the quantized GEMMs directly against the mma instruction, and the plan doc lists "cuBLAS paths" under deliberately-skipped llama.cpp machinery. The payoff is that weights are never materialized in f32 — the 8p f16 cache, the one exception, costs +8.6 GB on 7B and was itself made obsolete by MMQ (the MMQ gate skips the warm pass, src/cuda/methods/prefill_f16.rs:89).

Why is capture keyed on (uid, range, pool_gen) instead of llama.cpp's node-props snapshot? llama.cpp memcmps per-node properties and can update a captured exec in place; minfer's topology is already deterministic in GraphParams (doc 13), so the only thing that can invalidate a capture is a device-address change — exactly what pool_gen tracks. The simpler key buys a much smaller state machine, at the cost of a full re-capture (warmup ×2 again) whenever allocation churns; decode steps don't churn, so the common case pays nothing.

Why capture on the 3rd run, and why is prefill capture on by default? Warmup exists because per-kernel state (dynamic shared-memory opt-ins, module loading) must be settled before capture — capturing a first-call launch would record a slow path forever. The 3-run protocol also makes capture self-funding: one-shot graphs never pay for it, and repeated identical-length prefills (the server slot scenario, where the same nt prefills over and over) amortize capture across many replays. The 8g② default-off interlude is instructive: an audit found unvalidated prefill capture, the default flipped off until the bit-parity harness (pp16/pp300) proved replay at real prefill scale, and R3-B flipped it back on — the promote-then-default pattern the optimization campaign uses everywhere.

Why f16 KV on the GPU (and f32 on CPU)? Decode attention streams the KV history per token; at 7B that is 28 layers × 512 kv-dims × 2 (K and V) × 2 bytes (f16) × context — halving it is an ~11% whole-decode win (8b). The CPU path stays f32 because its attention is compute/latency-carried, not KV-bandwidth-bound, and f32 keeps the reference path simplest (its own option is the packed Q8_0 cache of C4, which buys footprint rather than bandwidth). The policy boundary is measured, not aesthetic: n_layers × n_kv_embd ≥ 8192.

Why a 224-byte padded Q6_K layout (§3.2.2) and precomputed scale planes instead of faster kernels? Both are memory-layout answers to a memory-bound problem: the padded stride makes weight loads coalesced, the planes (W_exp/W_dsc, ~3 GB) move scale decoding out of the inner loop. Both carry opt-outs (MINFER_MMQ_Q6K_EXP=0 saves 1.52 GB for −5%; r54's record quantified the exact trade), because VRAM is the one resource you cannot grow at run time.

Why does the norm arm require a weight when Metal's degrades? Metal's rms_norm silently runs weightless if the gains are not registered on the backend — a debuggable-but-wrong result. CUDA's norm_weight returns Err("cuda: weight '...' not registered") instead (cuda_backend.rs:1220-1246, comment citing docs/GPU_SAFETY.md). The difference is deliberate: the all-or-nothing participation gate makes a missing CUDA weight unreachable through normal wiring, so reaching that error means an invariant is already broken — and broken invariants abort.

Why is the backend feature-gated at all, when Metal is unconditional? Metal ships with macOS; CUDA needs a toolkit at build time and a driver at run time that not every machine has. The opt-in flag keeps plain builds nvcc-free (build.rs never probes for it), keeps CI green on GPU-less machines, and pairs with cuda_static for the toolkit-free deployment story of §3.2.1.

3.4 Pitfalls & invariants

  • Never sync inside a capture window. A cudaStreamSynchronize inside the window corrupts the capture — the 7e② record's "faster but wrong" incident was exactly this: a temporary debug sync inside the matmul dispatch produced garbage only when graphs were enabled. Rule 2 of docs/GPU_SAFETY.md exists because of it; the general lesson (a perf win that changes load counts and breaks only one configuration is a bug) is recorded alongside.
  • Invariant violations return Err — with the values. The Attn arm rejects nkt != n_head_kv*hd, mismatched query/KV head dims, hd > 128, hd % 4 != 0, and non-divisible GQA head counts, each message carrying the actual numbers; the MatMul arm rejects non-multiple-of-32 quant dims and buffer-size mismatches; FusedQKV/FusedFFN reject nt != 1; RoPE rejects any non-neox style; a transposed-B matmul is refused outright. The catch-all arm turns any un-kernelled op into an error naming the op. None of these fall back to CPU — placement was decided at build time, so a guard failure means the run aborts with the blocking node's name.
  • The dead-write refusal (r52's skip-write mode). In default mode 2, a fused RMSNorm/SwiGLU producer writes only the int8 quantize plane and skips the f32 output. If any GEMM then tries to quantize that f32 source again, mmq_quantize_native refuses (the dead_write flag in the MmqCache) and prefill_mmq returns Err("... mode-2 dead-write A refused ...") — a silent garbage read was the alternative, and the record says the guard "turns any window violation into a loud error instead of reading the unwritten buffer". The nb_bt_only flag (§3.2.2) degrades mode 2 to mode 1 at load time when the model's quant mix makes skip-write unsound at all.
  • Positions and ids cross the host boundary as bit patterns. The allocator fills I32 inputs as f32::from_bits (doc 07); the GPU never converts them on the host — positions_i32 runs a tiny device kernel (bits_to_i32) once per execution window (memoized), and the embed kernel reads token ids with __float_as_int (the 7e③ record warns that __float2int_rn would read the denormal float and always yield 0).
  • Memo lifecycles are execution-window-bounded. The MmqCache, the positions memo, and the capture staging all key on or clear at synchronize — pool buffer ids are recycled across executions, so a memo that leaked past a boundary would alias different data under the same id. The positions_i32 scratch growth even bumps pool_gen, because the freed scratch pointer may be embedded in captured graphs.
  • Device memory is not host memory. On GB10, plain memcpy of a device pointer SIGSEGVs (GPU_SAFETY rule 3); every D2H goes through copy_to_host/pinned staging, and read_host returns None so no caller can even attempt a borrowed device read.
  • Test with the fusion passes and both graph modes. The decode fusions (FusedQKV/FusedFFN) are part of the reuse identity; the unfused path must run the same FusionPass to be comparable (graph rule 7), and graph replay A/B (MINFER_NO_CUDA_GRAPH=1) is the standard harness for anything that touches stream state — the 7d verification matrix runs replay-vs-direct bit-parity as its first gate.

4. Observe & verify

  • Startup banner — CUDA: using <device> (SM 12.1, ... MB, ... SMs) + CUDA: GPU acceleration enabled (or not available, using CPU fallback) tells you in one line which device was picked and whether the graph will run on it at all. --gpu N pins the device; MINFER_DISABLE_CUDA=1 forces CPU.
  • The dispatch labels — MINFER_MMQ_RAW_NB_DEBUG=1 prints which MMQ kernel variant is active per launch class, including why a fast path was skipped (fallback! vs the deliberate exp=off — the r53 lesson that a fallback-correct fast path needs a visible label, since parity tests cannot see which path ran).
  • Escape hatches, one per mechanism (all default to the promoted path; each restores the pre-optimization behavior so A/Bs stay reproducible): MINFER_MMQ=0 (f16 prefill GEMM), MINFER_NO_PREFILL_GEMM=1 (legacy per-type kernels), MINFER_NO_KQ_MMVQ=1 (K-quant decode back to f32 kernels), MINFER_NO_DECODE_A_FUSE=1 (standalone decode quantize), MINFER_NO_W16CACHE=1 (per-call f16 scratch), MINFER_MMQ_Q6K_EXP=0/MINFER_MMQ_Q4K_DSC=0 (drop the precomputed planes), MINFER_NO_FUSE_QKV=1/MINFER_NO_FUSE_FFN=1 (decode fusions), MINFER_CACHE_TYPE=f32|f16|q8_0 (KV element type; q8_0 is the packed cache the CPU, CUDA and Metal kernels read, C4 — it is now 1.04× f16's time on a packed hd = 128 prefill post-#144), MINFER_NO_PINNED_READBACK=1 (pageable readback).
  • CUDA Graph state — MINFER_NO_CUDA_GRAPH=1 forces direct launches (the A/B lever for replay); MINFER_NO_PREFILL_CAPTURE=1 disables prefill capture only. Note that MINFER_TRACE/viz capture and MINFER_GRAPH_DUMP node dumps silently disable replay while active (per-node host readbacks are illegal in a capture window) — and mode-2 fused producers degrade to mode 1 when any dump/trace reader is on, so instrumented runs are not perf runs.
  • Device-side debugging — MINFER_CUDA_DEBUG used to turn on per-node labeled syncs (debug_sync) that reported launch/sync errors with a layer tag. That knob lived on the legacy layer_gpu surface and was deleted with it (#240/#241); the corresponding graph-path tool today is MINFER_OP_TIMING=1 with CudaState::sync() as the drain point.
  • Kernel selection in one command — MINFER_NO_CUDA_GRAPH=1 MINFER_TIMING=1 ./target/release/minfer <model> "hi" gives clean per-forward timing (doc 09's caliber), and bench [-p N] [-n N] measures prefill/decode tok/s the way the campaign's tables do. nsys/ncu profiles are what the optimization records cite per kernel.
  • Tests (device-gated: they skip when no GPU answers, so CI without CUDA stays green) — cargo test --features cuda cuda_ covers the whole backend: per-kernel parity (cuda_elementwise_parity, cuda_matmul_parity, cuda_kquant_matmul_parity, the per-type MMVQ parity tests), attention round trips (cuda_attn_split_decode_parity, f16-KV variants), the fusions (cuda_fused_qkv_epilogue_bitwise, cuda_fused_ffn_parity), embedding/gather (cuda_embed_getrows_parity), MMQ prefill bit-parity (cuda_prefill_mmq_parity, cuda_q6k_dsc_dense_byte_exact), and the graph machinery (cuda_graph_replay_bit_parity, cuda_graph_recaptures_on_pool_gen_change, cuda_prefill_capture_bit_parity_pp16_pp300, cuda_capture_abort_on_error). End-to-end: greedy output must equal the CPU path's greedy output at temp 0 — the Phase-7 acceptance gate for every supported model.

5. Cross-references

  • 08 — The scheduler — the split walk that calls execute_node and the replay hook this doc's §3.2.8 plugs into.
  • 10 — CPU matmul kernels — the scalar/AVX2 origin of the quantize-and-integer-dot scheme MMQ lifts onto tensor cores; §2's block-layout vocabulary is reused here without redefinition.
  • 11 — Attention, vec ops, and the KV cache — the attention math the split-KV kernel parallelizes (§2.5 previews the GPU deltas this doc implements).
  • 13 — The decode loop — why decode reuses the graph and its buffers, the precondition for CUDA Graph replay.
  • 14 — The Metal backend — the sibling backend: same Backend trait, same pool/free-list shape, same f16-KV policy; Metal dispatches one command buffer per split where CUDA captures one graph per split, and Metal's norm arm degrades silently where CUDA's returns Err.
  • docs/CUDA-BACKEND-DESIGN.md — the backend's design + implementation record: §3's llama.cpp reference map, §4 the design, §5 phases 7a–7e with per-item A/B numbers (7e②'s 3.1× decode, 7d's +18% replay).
  • docs/CUDA_OPTIMIZATION.md + docs/cuda_optimization_steps/ — the optimization campaign: §0's master history table (Phase 8 rows 8b–8q, the R-series, P5, the r28–r60 MMQ rounds, the D-series decode sessions), one standalone record per step; §1.1 the current perf/memory state.
  • docs/CUDA-TECH-PRIMER.md — the technology primer behind §2.1's five words: warps, occupancy, cp.async, mma, CUDA Graphs, each with its campaign context.
  • docs/LLAMA-CPP-MMQ-ANALYSIS.md — the reference MMQ kernel dissected (dispatch, tiling, numerics, SASS census) plus minfer's round-by-round contrast; the source for §2.3's claims.
  • docs/BUILD.md — the build reference this doc's §3.2.1 summarizes: nvcc/host-compiler detection, MINFER_CUDA_CCBIN, arch coverage (sm_70…sm_121, CUDA 12.8-vs-13 Volta note), cuda_static linking.
  • docs/SUPPORT-MATRIX.md — the per-quant support table incl. the CUDA notes (which path runs at which nt, the Q5_K id % 32 gate, the F32 caveats).
  • docs/GPU_SAFETY.md §"CUDA" — the seven hard rules §3.4 paraphrases (capture windows vs syncs, launch-error policy, weight-registry ownership, runtime device limits).
  • docs/GLOSSARY.md — the campaign glossary if a term (NB-BT, KDR, W_dsc, rpw...) from the records above needs its formula.

← 14 — The Metal backend · Index →

minfer Compute Graph Design

How minfer turns one forward pass into a declarative graph and executes it across CPU, Metal and CUDA. This is the authoritative design and implementation record for src/graph/: the IR, the builder, the allocator, the scheduler, the fusion rules, graph reuse, and the per-backend execution mapping.

Status. Landed. Every mechanism described here is implemented in the tree; the phase ledger and the deviations from the original plan are in §17. Baseline: HEAD = 5471680 (2026-09-13); the working tree additionally carries the post-baseline fixes listed at the end of §17.3.

Provenance. This file was docs/GRAPH-REFACTOR-PLAN.md, written before the rewrite as a plan. The section skeleton is preserved; the body has been rewritten in the present tense against the current code, and the plan-time "current state" descriptions have been replaced by the landed implementation. One historical caveat: the commit hashes in the plan-era phase table no longer resolve (the repository history was rewritten); the resolvable equivalents are noted in §17.

Related documents. The end-to-end walkthrough (docs/inference_e2e_walkthrough/05-graph-builder-ir.md … 08-scheduler-execute.md) narrates the same machinery line by line for a first-time reader; this document is the design of record and does not repeat that narrative. Backend-specific optimization history lives in docs/CUDA-BACKEND-DESIGN.md, docs/CUDA_OPTIMIZATION.md, docs/cuda_optimization_steps/ and docs/METAL_OPTIMIZATIONS.md.


1. Design Goals

1.1 The problem

The pre-rewrite engine computed a forward pass imperatively, and that shape had four structural costs: no reuse across decode steps, fusion hard-wired and invisible, a backend chosen inside the loop (so a support limitation could silently change the execution path mid-run), and ~620 lines of per-architecture loop to rewrite for a new model.

Those costs, and the decision they forced, are ADR-0001; this page assumes it and states what was built instead.

1.2 Goals and outcome

GoalLanded outcomeEvidence
Multi-backend dispatch decided before executionPer-node assignment at build time via GraphAllocator::supports and each backend's supports_op/supports_fused; the scheduler partitions the assigned graph into splits§3.4, §3.5
Fusion as a first-class IR citizenFusionPass rewrites Mul(Silu(x), y) → SwiGLU; decode additionally builds FusedQKV, QkvBiasRopeStore, FusedFFN, FusedQkvNorm as single nodes§5
Graph reuse with no rebuildGraphCache compares GraphParams only; a decode step reuses graph + allocator + KV regions and refreshes input buffers§6
Extensible architecturesModelDef::build_graph per architecture; Qwen2 and Qwen3 each own a graph.rs; the imperative forward.rs was deleted§10, §13
No silent fallbackexecute_node returns Result<(), String>; kernel-invariant violations abort with the actual values, matching docs/GPU_SAFETY.md§3.5, §7.4
CUDA as a real backend, not a stubCudaBackend wraps the existing cuda.rs device layer and preserves CUDA Graph capture/replay keyed by graph uid§9

1.3 Non-goals

  • Multi-sequence batching on CPU. The IR, the allocator and the attention kernels are sequence-aware (E1/E1b/E2: seq_ids, attn_span, per-sequence KV reservations, Batch), and the server composes a batch — but the default follows the device (E6: batches on CUDA and Metal, since #44 part (b); off on CPU — MINFER_BATCH=0/1 forces either way), because batching measures 0.49x the serial path on CPU while it is 1.9x on the GB10 (E2's acceptance is device-dependent; see ARCHITECTURE-EXECUTION-PLAN.md §7 and the E6 record).
  • A generic ggml operator set. Scale, Softmax, View, Reshape, Permute, AttnMode::Mha and FusedOp::BatchMatMul are present in the vocabulary but no supported architecture emits them; they are kept for parity and future use. Op::FusedBiasRope and its fusion rule were removed outright (the capability was never claimed — see §5.2).
  • Cross-vendor graph transpilation. Each backend implements its own execute_node; there is no lowering pass.

1.4 Core invariants

These are the rules the rest of the document elaborates. Breaking one is a bug, not a tuning choice.

  1. KV positions are data, not structure. KvcacheStore/KvcacheLoad carry only the layer index; the write row arrives through the cells input node (C6: positions is the token's index within its sequence and drives RoPE and the causal bound, while the allocator resolves cells — the arena row — from the run table; the two coincide only while a run starts at cell 0). The topology never depends on n_past — this is the precondition for decode reuse (llama.cpp allow_reuse behaves the same).
  2. Topology is a deterministic function of GraphParams. Equal params ⇒ identical node sequence ⇒ the graph may be reused without rebuilding. The debug build asserts this structurally.
  3. Weights are named, not copied. A node references its weight by name (MatMulMeta.weight_name); the backend resolves the name in its registry at execution time.
  4. The allocator is the single owner of buffers. Backends own pools; the scheduler never allocates.
  5. Execution follows build order. The builder appends sources before consumers, so node-id order is a valid topological order; the allocator's liveness uses the same order.
  6. Fusion never duplicates an existing fused kernel. A rewrite is applied only when the target backend reports supports_fused.
  7. Errors are errors. A backend that cannot execute a node returns Err; it never falls back to CPU mid-run.

2. Overall Architecture

2.1 Pipeline

 model.build_graph(&GraphParams)                models/<arch>/graph.rs
        │  ComputeGraph (pure IR, no execution)
        ▼
 scheduler.assign_backends(&graph, &alloc)      capability-driven, per node
        │
        ▼
 FusionPass::run(&graph, backends, backend_of)  SwiGLU rewrite (see §5)
        │
        ▼
 alloc.alloc_graph(&graph)                      liveness + persistent KV regions
        │
        ▼
 scheduler.execute(&graph, &alloc)              split → sync → copy → execute
        │
        ├── CpuBackend    (kernel.rs / vec_ops.rs)
        ├── MetalBackend  (src/metal/ + src/metal/kernels/)
        └── CudaBackend   (cuda.rs + cuda/kernels/*.cu)

GraphCache owns the graph, the allocator and the last GraphParams; a decode step that reuses the cached graph skips the first three stages and only refills inputs.

2.2 Module map

ModuleRole
graph/mod.rsComputeGraph, CNode, DType, Backend, BufRef, PersistentBuf, topo_order
graph/ops.rsOp, NodeMeta and the per-op metadata structs, AttnMode, FusedOp
graph/builder.rsGraphBuilder — the declarative construction API
graph/params.rsGraphType, CParams, GraphParams (the reuse identity)
graph/cache.rsGraphCache — params-only reuse, graph uid, allocator lifetime
graph/backend.rsBackend trait, KvProvider
graph/registry.rsthe backend registry (F4): the Backend handle, the name-keyed entries (priority + caps + pool/host-read/kv-format/enable hooks), the --backend/MINFER_BACKENDS fence — see BACKEND-REGISTRY-DESIGN.md
graph/alloc.rsGraphAllocator — liveness, node→buffer map, KV regions, cross-backend staging
graph/scheduler.rsBackendScheduler — assign_backends, split_graph, execute
graph/fusion.rsFusionPass — pattern-matching rewrite
graph/cpu_backend.rsCPU executor over kernel.rs / vec_ops.rs
graph/metal_backend.rsMetal executor over src/metal/ per-op methods
graph/cuda_backend.rsCUDA executor over cuda.rs, including CUDA Graph capture/replay
graph/dot.rsGraphviz DOT export
graph/json.rsJSON export for the interactive visualizer (viz/)

2.3 Scheduler policies

The original plan stated three policies; all three are landed, with two nuances.

  1. Attention and its KV regions share a backend. A layer's two KV regions are created on the backend that first uses the layer (GraphAllocator::ensure_kv), and the layer's attention node resolves the same pair through kv_pair(layer). Keeping them together avoids a per-step O(n_kv) host round trip.
  2. A GPU split shares one submission. Metal accumulates kernels into one command buffer for the split and submits at the split boundary (sync_backend). CUDA can go further: a decode split may be replayed as a single captured CUDA Graph launch (§9.4). Nuance: the whole-layer layer_gpu() fast path the plan proposed keeping was not kept. Per-op execution through MetalBackend/CudaBackend is the only inference path; the legacy cuda.rs::layer_gpu survives as dead code, and the plan's "whole-layer fast path" fallback is unnecessary because assignment is decided at build time.
  3. Fusion is capability-driven, not forced. FusionPass consults supports_fused per node, and the decode fusions are gated in CParams so they can be A/B-tested. Nuance: only the SwiGLU rewrite is actually accepted by a backend today; see §5.2.

2.4 Execution order and the correctness contract

BackendScheduler::execute walks the graph split by split, and within a split node id by node id. This is the same contract ggml uses (nodes[0..n_nodes]): the builder guarantees that a producer precedes its consumers, so a KV store runs before the attention that reads the KV view.

The allocator's liveness analysis deliberately uses the same order rather than the Kahn order returned by topo_order(). topo_order() may move source-less nodes (e.g. kv_load) ahead of nodes built before them; liveness computed on that order can consider an input dead while the scheduler has not read it yet, and reuse its buffer — the G3 regression, recorded as deviation 22 in §17.


3. Core Data Structures

3.1 The IR (graph/mod.rs, graph/ops.rs)

#![allow(unused)]
fn main() {
pub type NodeId = usize;

pub enum DType { F32, F16, I32, Q8_0 }   // activations are F32; F16/I32/Q8_0 for inputs

// F4: the backend handle — a fixed id space (cpu = 0, metal = 1, cuda = 2), not an
// enum. The ids are a KV-session file-format contract; `registry.rs` maps names to
// them and carries each backend's priority, capabilities and pool hooks.
pub struct Backend(u16);
impl Backend {
    pub const CPU: Backend = Backend(0);
    pub const METAL: Backend = Backend(1);
    pub const CUDA: Backend = Backend(2);
}

pub struct BufRef { pub backend: Backend, pub id: usize }   // id inside that backend's pool
pub struct PersistentBuf { pub name: String, pub backend: Backend, pub id: usize }

pub struct CNode {
    pub id: NodeId,
    pub name: String,
    pub op: Op,
    pub src: Vec<NodeId>,
    pub out_shape: [usize; 4],
    pub out_dtype: DType,
    pub backend: Option<Backend>,     // decided by the scheduler
    pub meta: NodeMeta,
}

pub struct ComputeGraph {
    pub nodes: Vec<CNode>,
    pub inputs: Vec<NodeId>,
    pub outputs: Vec<NodeId>,
    pub uid: u64,                     // CUDA Graph cache key; assigned by GraphCache
}
}

Shapes follow the llama.cpp convention: activations are feature-major [d, nt, 1, 1] (feature fastest, tokens in dim 1), and a weight tensor carries GGUF metadata [in, out] while its memory is [out][in] row-major. MatMulMeta therefore records in_dim = shape[0], out_dim = shape[1].

DType::size() is defined for all four variants, but only F32 activations and I32 inputs are constructed by the supported models; F16 KV storage is a backend-internal concern, not an IR dtype.

The Op vocabulary

#![allow(unused)]
fn main() {
pub enum Op {
    Input,                                     // leaf, host-filled every step
    Add, Mul, Scale(f32), Silu,                // element-wise
    Softmax { dim: usize },                    // reduction (vocabulary only)
    RmsNorm { eps: f32 },                      // normalization
    QkNorm { hd: usize, nh: usize, eps: f32 }, // per-head RMSNorm (Qwen3 q/k norm)
    MatMul { transpose_b: bool },              // linear algebra
    GetRows,                                   // embedding lookup / tail-row selection
    RoPE { style: RopeStyle },                 // positional encoding
    Attn { mode: AttnMode, explicit_span: bool },  // attention (softmax fused inside the kernel);
                                               //   `explicit_span` = `positions` cannot bound the
                                               //   node (several sequences, or a window that does
                                               //   not start at cell 0), so only a backend that
                                               //   reads the explicit span may take it
    KvcacheStore { layer: usize },             // persistent KV write; the row comes from `cells`
    KvcacheLoad  { layer: usize },             // view of the persistent KV region
    View { offset: usize, shape: [usize; 4] },
    Reshape { shape: [usize; 4] },
    Permute { dims: [usize; 4] },
    SwiGLU,                                    // FusionPass output
    BatchMatMul,                               // planned, not emitted (single-output IR)
    FusedQKV { layer: usize },                 // decode: concat matmul + bias/rope/store
    QkvBiasRopeStore { layer: usize },         // decode mixed-quant: 3 matmuls + one epilogue
    FusedFFN,                                  // decode: gate|up concat matmul + in-place swiglu
    FusedQkvNorm { layer: usize },             // decode Qwen3: concat matmul + qk_norm + rope/store
}
}

Op derives a full PartialEq (payloads included). The production reuse decision does not use it — it compares GraphParams only — but the debug structural check compares op payloads, shapes and dependencies, so a payload that silently changed is caught.

The variants marked "vocabulary only" (Scale, Softmax, View, Reshape, Permute, BatchMatMul, AttnMode::Mha) carry #[allow(dead_code)] and no supported architecture emits them (§5.5). The KV rule from invariant 1 is visible directly in the payloads: KvcacheStore/KvcacheLoad carry the layer index and nothing else.

Node metadata

meta uses a concrete enum rather than the plan's Box<dyn Any + Send + Sync> — it is PartialEq (needed by the structural check), cannot panic on downcast, and keeps CNode Clone:

#![allow(unused)]
fn main() {
pub enum NodeMeta {
    None,
    MatMul(MatMulMeta), Norm(NormMeta), Rope(RoPEMeta), Attn(AttnMeta),
    Kvcache(KvcacheMeta), Embed(EmbedMeta),
    FusedQkv(FusedQkvMeta), QkvBiasRopeStore(QkvBiasRopeStoreMeta),
    FusedFfn(FusedFfnMeta), FusedQkvNorm(FusedQkvNormMeta),
}
}
MetaCarriesConsumed by
MatMulMetaweight_name, bias_name, weight_ttype, in_dim, out_dimbackends pick the kernel by weight_ttype without holding the Tensor
NormMetaoptional weight/bias names (shared by RmsNorm and QkNorm)CPU / Metal / CUDA
RoPEMetafreq_base, freq_scale, n_head, hdrope kernel
AttnMetalayer, n_head, n_head_kv, hd, hd_kv, nkt (KV row stride), scaleattention; layer resolves kv_pair. The allowed cells are data (attn_span, E1), not a field
KvcacheMetan_embd, n_head_kvKV region sizing / attention strides
EmbedMetavocab_size, weight_name, weight_ttypeembedding lookup
FusedQkvMetaconcat weight name, three bias names, in_dim, nqt, nkt, hd, nh, nk, rope params, kv_elemsFusedQKV
QkvBiasRopeStoreMetathree bias names, nqt, nkt, hd, rope params, kv_elemsQkvBiasRopeStore
FusedFfnMetagu_weight, weight_ttype, in_dim, nfFusedFFN
FusedQkvNormMetaconcat weight, q_norm_name, k_norm_name, dims, rope params, kv_elems, epsFusedQkvNorm

ComputeGraph::topo_order() is a Kahn sort used for validation; capture_nt_hint() (CUDA builds) returns the first MatMul node's output row count for the capture gate.

3.2 Graph builder (graph/builder.rs)

GraphBuilder is an append-only factory: it assigns ids, records shapes/dtypes/meta and wires src edges. It never computes and never allocates.

#![allow(unused)]
fn main() {
impl GraphBuilder {
    pub fn new() -> Self;
    pub fn node(&mut self, name: &str, op: Op, src: &[NodeId],
                out_shape: [usize; 4], out_dtype: DType, meta: NodeMeta) -> NodeId;

    pub fn input(&mut self, name: &str, shape: [usize; 4], dtype: DType) -> NodeId;

    // shape-aware convenience constructors
    pub fn embedding(&mut self, ids: NodeId, weight: &Tensor) -> NodeId;
    pub fn rms_norm(&mut self, x: NodeId, weight: Option<&Tensor>, eps: f32) -> NodeId;
    pub fn qk_norm(&mut self, x: NodeId, weight: Option<&Tensor>,
                   hd: usize, nh: usize, eps: f32) -> NodeId;
    pub fn matmul(&mut self, x: NodeId, w: &Tensor, bias: Option<&Tensor>) -> NodeId;
    pub fn matmul_by_name(&mut self, x: NodeId, weight_name: &str, ttype: TensorType,
                          out_dim: usize, in_dim: usize) -> NodeId;
    pub fn get_rows(&mut self, x: NodeId, ids: NodeId, out_shape: [usize; 4]) -> NodeId;
    pub fn rope(&mut self, x: NodeId, pos: NodeId, style: RopeStyle, meta: RoPEMeta) -> NodeId;
    pub fn silu(&mut self, x: NodeId) -> NodeId;
    pub fn add(&mut self, a: NodeId, b: NodeId) -> NodeId;
    pub fn mul(&mut self, a: NodeId, b: NodeId) -> NodeId;
    pub fn softmax(&mut self, x: NodeId, dim: usize) -> NodeId;
    pub fn attn(&mut self, q: NodeId, kv: NodeId, pos: NodeId,
                mode: AttnMode, meta: AttnMeta) -> NodeId;  // src = [q, kv, pos, span]
    pub fn swiglu(&mut self, gate: NodeId, up: NodeId) -> NodeId;
    pub fn kvcache_store(&mut self, layer: usize, k: NodeId, v: NodeId,
                         pos: NodeId, n_ctx: usize) -> NodeId;
    pub fn kvcache_load(&mut self, layer: usize, n_embd: usize,
                        n_ctx: usize, n_head_kv: usize) -> NodeId;
    pub fn output(&mut self, node: NodeId);
    pub fn build(self) -> ComputeGraph;

    // decode fused constructors (§5.3)
    pub fn fused_qkv(&mut self, x: NodeId, pos: NodeId, layer: usize, meta: FusedQkvMeta) -> NodeId;
    pub fn qkv_bias_rope_store(&mut self, q: NodeId, k: NodeId, v: NodeId, pos: NodeId,
                               layer: usize, meta: QkvBiasRopeStoreMeta) -> NodeId;
    pub fn fused_ffn(&mut self, x: NodeId, meta: FusedFfnMeta) -> NodeId;
    pub fn fused_qkv_norm(&mut self, x: NodeId, pos: NodeId, layer: usize,
                          meta: FusedQkvNormMeta) -> NodeId;
}
}

Two shape rules are worth calling out because the fused nodes depend on them:

  • attn() takes its output shape from AttnMeta (n_head * hd) rather than from the q input, because a fused QKV node's q handle actually points at a larger q|k|v concat buffer.
  • fused_qkv / fused_qkv_norm output [nqt + 2*nkt, nt]; fused_ffn outputs [2*nf, nt]; the downstream matmul reads the rows it needs (0..nqt q for attention, 0..nf for the FFN down projection).

3.3 Allocator (graph/alloc.rs)

#![allow(unused)]
fn main() {
pub struct GraphAllocator {
    cpu: CpuBackend,
    metal: Option<MetalBackend>,      // macOS
    cuda: Option<CudaBackend>,        // feature = "cuda"
    node_to_buf: HashMap<NodeId, BufRef>,
    cross: HashMap<NodeId, BufRef>,   // split-boundary staging for the current graph
    buf_alive: HashMap<(Backend, usize), usize>,
    kv: HashMap<usize, [BufRef; 2]>,  // per-layer [K, V] persistent regions
    pub persistent: Vec<PersistentBuf>,
}
}

alloc_graph runs on every build/rebuild and performs, in order:

  1. Release the previous graph's liveness buffers into their pools (buf_alive and node_to_buf are cleared; the cross staging buffers are freed and re-materialized on the next execute). Persistent regions are not touched — they are the KV cache.
  2. Validate acyclicity with topo_order(), then keep build order for everything else.
  3. Compute liveness. last_use[node] starts at its exec index and is extended to the largest index of any consumer. graph.outputs get last_use = order.len(). Inputs get the same treatment: they are host-filled before execution, so reusing an input's buffer for another input would let the later fill clobber the earlier one (deviation 23).
  4. Count consumers per node, for in-place alias safety.
  5. Walk in build order, calling sweep(i) to return expired buffers, then allocate:
    • KvcacheStore/KvcacheLoad bind the node buffer to the layer's K region; ensure_kv allocates the never-freed [K, V] pair on first use for that layer/backend;
    • FusedQKV / FusedQkvNorm / QkvBiasRopeStore also ensure_kv (their kernels store K/V, with the size taken from meta.kv_elems) but keep a normal concat/epilogue output buffer;
    • Silu, RoPE and QkvBiasRopeStore alias their input buffer when the input's sole consumer is this node and both are on the same backend; otherwise they get a fresh buffer;
    • every other node gets a pooled buffer sized by n_elements().

The in-place alias rule is what makes kernel-order execution correct without a host copy: a pending GPU producer and this kernel read/write the same physical buffer. A cross-backend input is never aliased (the producer already completed at the split boundary, so the staging copy is safe).

Host access and transfer helpers:

#![allow(unused)]
fn main() {
pub fn fill_input(&mut self, graph, name: &str, data: &[f32]) -> Result<(), String>;
pub fn fill_input_i32(&mut self, graph, name: &str, data: &[u32]) -> Result<(), String>;
pub fn get_buffer(&self, graph, id) -> Option<&[f32]>;          // CPU pools only
pub fn copy_to_cpu(&mut self, id) -> Option<Vec<f32>>;          // any backend
pub fn copy_kv_to_cpu(&mut self, layer) -> Option<(Vec<f32>, Vec<f32>)>;
pub fn sync_backend(&mut self, backend: Backend);
pub fn copy_across(&mut self, node_id, dst_backend) -> Result<(), String>;
pub fn cross_buffer(&self, node_id) -> Option<BufRef>;
pub fn alloc_persistent(&mut self, name, backend, size) -> BufRef;
}

I32 inputs (token ids, positions, tail ids) are stored as f32::from_bits bit patterns — exact for |v| < 2^24, which covers every supported model's vocab and context. CPU and Metal read them back with to_bits(); CUDA converts the buffer to a real i32 plane on device, memoized per execution window.

Cross-backend staging is deliberately not an overwrite of the node's canonical buffer. copy_across allocates a staging buffer on the consumer's backend (once per graph, then rewritten on each execute) and records it in cross; consumers on the producer's own backend keep reading the canonical buffer. This is what keeps a reused graph re-executable.

GraphAllocator implements KvProvider, so kv_pair(layer) is the only way any executor reaches the persistent regions — the V half in particular has no dedicated node and is reachable only through this pair.

3.4 Scheduler (graph/scheduler.rs)

#![allow(unused)]
fn main() {
pub struct Split {
    pub backend: Backend,
    pub node_range: (usize, usize),   // [start, end) in graph.nodes
    pub inputs: Vec<NodeId>,          // produced by another backend, copied in
    pub outputs: Vec<NodeId>,         // consumed by another backend
}

impl BackendScheduler {
    pub fn assign_backends(&self, graph: &mut ComputeGraph, alloc: &GraphAllocator);
    pub fn split_graph(&self, graph: &ComputeGraph) -> Vec<Split>;
    pub fn execute(&self, graph: &ComputeGraph, alloc: &mut GraphAllocator) -> Result<(), String>;
}
}
  • assign_backends asks alloc.supports(op, dtype) for each unassigned node and falls back to CPU via .or(Some(Backend::CPU)). supports consults Metal first (macOS), then CUDA (feature = "cuda"), then CPU. Nodes with an explicit backend keep it.
  • split_graph scans nodes in order and cuts whenever the assigned backend changes, then derives cross-split edges: a src living in another split becomes an input of this split and an output of the producer's split. The result is a contiguous partition — there is no split-merging pass.
  • execute retires the previous backend, copies the new split's inputs across, then runs its nodes. Cross-backend inputs are resolved through cross_buffer(src) filtered to the executing backend, falling back to the canonical node_buffer(src); this filtering matters when one value feeds two backends and only one staging copy exists.
  • The boundary is two phases (F5, #58; the wait is deferred since #138). At a backend change, execute first enqueues one staging copy per entry of Split::inputs (GraphAllocator::copy_across — a no-op if that entry is already in flight, because it is the same transfer into the same staging buffer), then runs the consuming split. Each staged entry's single wait (GraphAllocator::await_cross) is issued at the consumer's first read of that buffer (GraphAllocator::cross_input_ready); whatever nothing downstream reads is drained once after the last split. The transfer itself is the source backend's registered hook (BackendEntry::copy_cross), so a device source can make it an cudaMemcpyAsync plus a recorded event instead of the pre-F5 blocking stream-sync + cudaMemcpy host round trip. The wait is the backend's await_cross: cudaEventSynchronize for a device→host copy (host memory is about to be read — the invariant that makes it necessary), a no-op for a CPU source (there is no device transfer to wait for), and cudaStreamWaitEvent for a device consumer. One wait per copy is the contract; the checked GraphAllocator::cross_input refuses a staged entry whose wait has not been issued rather than consuming a transfer that may still be in flight. The boundary's own sync_backend is gone for a backend that can order its close on its own stream (Backend::retire), since the copies are stream-ordered behind it. Counters, the not-hot-site list and the per-backend table are in BACKEND-REGISTRY-DESIGN.md §11.
  • KV resolution happens before the backend borrow: KvcacheStore, FusedQKV, QkvBiasRopeStore, FusedQkvNorm and Attn (via AttnMeta.layer) all call alloc.kv_pair(layer) and pass it to execute_node.
  • Dead nodes are skipped. Fusion can orphan a node (e.g. the Silu folded into SwiGLU); such a node has no buffer and is not executed.
  • Split mismatch is an error, not a fallback: if a node's buffer is on a different backend than the executing split, execute returns an Err naming the node and both backends.

Diagnostics: MINFER_GRAPH_TRACE=1 prints the split list and a per-op/per-backend node count.

The scheduler is also where the optional per-node data capture for the visualizer hooks in (§8.3) and where CUDA Graph replay short-circuits a split's node loop (§9.4).

3.5 Backend trait (graph/backend.rs)

#![allow(unused)]
fn main() {
pub trait KvProvider {
    fn kv_pair(&self, layer: usize) -> Option<(usize, usize)>;   // (k_buf_id, v_buf_id)
}

pub trait Backend: Send + Sync {
    fn name(&self) -> &str;                                       // diagnostics
    fn supports_op(&self, op: &Op, dtype: DType) -> bool;
    fn supports_fused(&self, fused: &FusedOp) -> bool;

    fn alloc_buffer(&mut self, size: usize) -> usize;             // from the recycle free list
    fn free_buffer(&mut self, id: usize);
    fn alloc_fresh(&mut self, size: usize) -> usize;              // bypasses the free list
    fn pool_len(&self) -> usize;                                  // E4: did the pool grow?

    fn execute_node(&mut self, node: &CNode, in_bufs: &[BufRef], out_buf: BufRef,
                    kv_pair: Option<(usize, usize)>) -> Result<(), String>;

    fn read_host(&self, id: usize) -> Option<&[f32]>;
    fn write_host(&mut self, id: usize, data: &[f32]) -> Result<(), String>;   // exact length
    fn write_host_window(&mut self, id: usize, offset: usize, data: &[f32])    // E4 S2: a window of
        -> Result<(), String>;                                                 // a class-sized buffer
    fn synchronize(&mut self);

    #[cfg(feature = "cuda")]
    fn graph_replay(&mut self, uid: u64, range: (usize, usize),
                    nt_hint: Option<usize>) -> bool { false }
}
}

Three deviations from the plan's sketch are worth stating explicitly:

  • execute_node takes &mut self (the pool is mutated) and returns a Result (invariant 7).
  • kv_pair is passed in rather than looked up: the backend cannot see the allocator's KV map, and the scheduler already resolves it.
  • alloc_fresh exists because a split-boundary staging buffer must not be recycled during the same execute: the free list can hold ids whose physical contents are still referenced by node_to_buf and read later in the same pass. Fresh buffers re-enter the free list on the next alloc_graph.

read_host is a borrowed view, which CUDA cannot provide; CUDA returns None and exposes the real device→host path through the allocator's copy_to_cpu (§9.5).

graph_replay is CUDA-only. Returning false may still have armed or entered capture mode as a side effect (backend-internal warmup bookkeeping); the captured window is closed at the split's synchronize. The default implementation is a no-op, so CPU and Metal drop the method entirely.

3.6 Backend registry (graph/registry.rs)

The set of backends is a registry, not a compile-time enum. Backend is a Copy/Hash/Eq/Ord handle over a fixed id space (cpu = 0, metal = 1, cuda = 2) — the ids are a KV-session file-format contract, so they are appended, never renumbered. Each backend module registers one BackendEntry at startup: its name, its assignment priority, its capability matrix (supports_op / supports_fused / supports_attn_span / reads_packed_kv) and the hooks that reach its pool (pool/pool_mut, host_read, kv_format, lazy enable, unavailable) plus the split boundary's two phases (F5: copy_cross / await_cross, §3.4). The capability matrices are module-level functions and the trait methods forward to them, so the registry's answer and the trait's answer cannot diverge.

Three things follow for this document:

  • supports_op is offered in priority order (Metal 300, CUDA 200, CPU 100) read from the registry, not in a hardcoded chain inside GraphAllocator::supports_for. The order is a pinned number because the assignment is topology (§6.1). Nothing here may depend on HashMap iteration order.
  • the allocator's dispatch (allocation, free, sync, copy_cells, host I/O, the KV element format, the session enable) is a trait call on the entry's pool hook — the twelve match backend { … } sites are gone, together with their #[cfg] fallback arms.
  • a fence by name (--backend / MINFER_BACKENDS) removes a backend from participation, read by the graph builders' Device and by supports_for from one filter; and a name that is unknown, not compiled into this build, or not usable on this machine is a loud startup refusal that names which of the three it is.

The full contract — the two orders, the determinism rule, the name surface, the exact refusal messages and the per-configuration feature-gate table — is BACKEND-REGISTRY-DESIGN.md.

3.7 Parameters and the reuse cache (graph/params.rs, graph/cache.rs)

#![allow(unused)]
fn main() {
pub enum GraphType { Decode, Prefill }

pub struct CParams {
    pub n_ctx: usize,
    pub flash_attn: bool,
    pub gpu: bool,        // whether a GPU backend participates (assignment is part of topology)
    pub fuse_qkv: bool,   // G4 decode QKV fusion (MINFER_NO_FUSE_QKV=1 disables)
    pub fuse_ffn: bool,   // G5 decode FFN fusion (MINFER_NO_FUSE_FFN=1 disables)
}

pub struct GraphParams {
    pub n_tokens: usize,
    pub n_out: usize,          // tail rows (G3): part of topology, not just execution
    pub gtype: GraphType,
    pub cparams: CParams,
    pub weights_version: u64,  // reserved for weight reload / LoRA switch
}

pub struct GraphCache {
    graph: Option<ComputeGraph>,
    alloc: GraphAllocator,     // outlives rebuilds: holds the KV regions
    prev_params: Option<GraphParams>,
}
}

GraphCache::try_reuse compares params_match — n_tokens, n_out, gtype, cparams (all of it, including gpu/fuse_qkv/fuse_ffn/explicit_span) and weights_version. It never inspects the node sequence. A batch's sequence count is deliberately absent: it is data (A7/E2), and a 1-sequence and a 2-sequence batch of the same shape share one graph. replace_graph assigns a fresh monotonic uid for CUDA Graph caching and leaves the allocator in place. See §6 for the full reuse story.

weights_version is currently a constant 1 at both model call sites: there is no LoRA or weight reload path yet, and no weights_version() trait method. The field and the next_weights_version() counter exist so a future reload can break reuse by bumping it.


4. Model Graph Construction

Both supported architectures build the graph in models/<arch>/graph.rs through GraphBuilder. The skeleton is shared; the differences are all in the projections and norms.

4.1 Qwen2 — prefill and decode

Qwen2Graph::build emits, in order:

  1. Inputs — token_ids [nt] I32, positions [nt] I32, and (only when n_out < nt) tail_ids [n_out] I32. tail_ids is declared at the graph head, not next to its consumer: an input node in the middle of the graph would split a GPU run into extra CPU/CUDA boundaries, each with a stream sync and host round trip (R3-A1, docs/CUDA_OPTIMIZATION.md).
  2. Embedding — b.embedding(token_ids, tok_embd).
  3. Per layer il:
    • residual = h; normed = rms_norm(h, attn_norm, eps);
    • Q/K/V projections — one of three shapes (§4.2);
    • attn_out = attn(q, kv, positions, mode, AttnMeta{ layer: il, ... }), where mode is Flash when cparams.flash_attn else Gqa;
    • wo = matmul(attn_out, wo);
    • residual add, with the G3 tail reduction on the last layer (§4.3);
    • residual = h; normed = rms_norm(h, ffn_norm, eps);
    • FFN — fused or unfused (§4.2);
    • residual add.
  4. Output — rms_norm(h, output_norm) → matmul(normed, output, output_b) → b.output(logits).

For nt > 1 the QKV and FFN blocks are always the unfused chains, because every fusion gate requires nt == 1. Prefill therefore runs the FusionPass-rewritten SwiGLU in the FFN and nothing else.

4.2 The decode fusion branches

The Q/K/V construction is where topology differs materially between prefill and decode. For decode (nt == 1) with a GPU participating and CParams.fuse_qkv on, and provided all three biases exist:

ClassConditionGraph
Concat (G4/D3-8 class 1)qkv_concat_available(wq, wk, wv) — same quant type, same input dim, block-aligned, so the loader registered blk.{i}.attn_qkvone fused_qkv(normed, positions, il, ...) node (the builder also wires cells), then kvcache_load for attention
Mixed quant (D3-8 class 2)concat unavailable, CUDA presentthree bias-less matmuls + one qkv_bias_rope_store(q, k, v, positions, il, ...) epilogue (plus cells from the builder); attention is wired to the epilogue node so q's matmul has exactly one consumer and can be aliased in place
Unfusedany gate off (or a Metal-only mixed-quant layer)matmul×3 (+bias) → rope×2 → kvcache_store → kvcache_load

Class 2 is CUDA-only: without the cuda feature qkv_epilogue_ok is false and those layers keep the unfused chain on Metal, bitwise-neutral against the pre-D3-8 behaviour. The FFN branch: for decode with CParams.fuse_ffn, gu_concat_available(ffn_gate, ffn_up) and nf <= 16384, the builder emits one fused_ffn(normed, ...) followed by the down matmul; otherwise it emits matmul(gate), matmul(up), silu, mul and lets FusionPass fold the last two.

Two gates are measured decisions, not correctness ones:

  • nf <= 16384. On the 7B class the concat matmul (od = 2*nf ≈ 37888) is slower than two separate matmuls with the decode nt == 1 kernel, while on 0.5B the fusion is worth ~3%.
  • concat_rows_feasible on the CUDA Qwen2 path. The probe is metadata-only because rebuilding the concat rows during graph construction cost ~920 ms per decode graph build (a ~1.9 GB probe); the loader has already built the concatenated weight, so the builder only needs to confirm it is feasible.

Both gates live in the builder, and CParams.fuse_qkv / fuse_ffn keep the A/B honest across graph reuse.

4.3 G3 — the n_out tail-row reduction

Prefill computes logits for the last n_out tokens only (the CLI uses n_out = 1). Rather than computing all nt rows and slicing, the builder inserts, after the last layer's wo:

cur_tail = get_rows(wo,       tail_ids, [n_embd, n_out])
res_tail = get_rows(residual, tail_ids, [n_embd, n_out])
h        = add(res_tail, cur_tail)

so the last layer's ffn_norm, gate/up/down, the residual adds and the lm_head matmul all run on n_out rows. This mirrors llama.cpp's ggml_get_rows(cur, inp_out_ids) at the final layer. When n_out == nt (decode), no tail_ids input and no reduction node are built. GraphParams.n_out participates in the reuse identity precisely because it changes the topology.

The logits buffer is exactly n_out * n_vocab either way (G3-reduced, or n_out == nt), so the extraction path does not clone a full-vocabulary row set.

4.4 Qwen3 differences

Qwen3 shares the skeleton and adds a per-head RMSNorm on Q and K (attn_q_norm / attn_k_norm) and has no attention biases. Its head dim is decoupled from n_embd / n_head (128 vs 64 on 0.6B), which drives the Q/K/V/wo widths, the RoPE dims, the attention scale and the KV row stride.

  • Unfused: qk_norm(q, q_norm, hd, nh, eps) and qk_norm(k, k_norm, hd, nk, eps) after the three bias-less matmuls and before the rope. QkNorm treats the flat token-major buffer as a contiguous [nt*nh, hd] matrix and reuses the RMSNorm kernels with d = hd, n = nt*nh.
  • Decode fused: fused_qkv_norm(normed, positions, il, ...) — one concat matmul, then per-head q/k norm, then a no-bias rope + K/V store pass. It exists as a separate op because the Qwen2 attn_bias_rope_store kernel cannot express the per-head norm; the Qwen2 path is untouched. This fusion is Metal-only (the Qwen3 concat probe has no CUDA arm and fuse_qkv is gated on metal_on), so CUDA Qwen3 runs the unfused qk_norm chain.
  • No QkvBiasRopeStore class exists for Qwen3 (there are no attention biases to fold).
  • Qwen3 uses the same G3 tail reduction and the same FusedFFN gate.

4.5 Graph inputs

InputShapeDTypePresent when
token_ids[nt, 1, 1, 1]I32always
positions[nt, 1, 1, 1]I32always (sequence-relative: RoPE + the causal mask)
cells[nt, 1, 1, 1]I32when the graph writes or resolves KV (KvcacheStore, FusedQKV, QkvBiasRopeStore): the arena row per token (C6)
tail_ids[n_out, 1, 1, 1]I32n_out < nt (prefill with the G3 reduction)

The prefill / decode call sites fill tail_ids with [(nt - n_out) .. nt); decode leaves it absent. Positions in the graph are always data: no node payload contains n_past, and the KV region is allocated as [n_embd, n_ctx] with only its written prefix read at execution time.


5. Operator Fusion

minfer fuses in two distinct ways, and the distinction matters when reading the IR.

5.1 Two mechanisms

MechanismWhenExamplesGated by
Build-time fused nodeThe builder knows the pattern statically at graph constructionFusedQKV, QkvBiasRopeStore, FusedFFN, FusedQkvNormCParams.fuse_qkv / fuse_ffn (+ weight availability, quant class, nf cap)
FusionPass rewriteA peephole pass over the built graph, per node's assigned backendSwiGLU (the only rule)backend supports_fused

Build-time fusion is preferred where the pattern spans weight registration (concat weights) or where the fused node needs extra metadata (layer index, biases, rope parameters). The peephole pass covers patterns that are cheap to recognize and safe to leave unfused — the unfused form is still correct, just slower.

5.2 FusionPass rules

#![allow(unused)]
fn main() {
impl FusionPass {
    pub fn run(&self, graph: &mut ComputeGraph, backends: &[&dyn Backend],
               backend_of: &dyn Fn(&ComputeGraph, usize) -> Option<usize>) -> usize;
}
}
  • SwiGLU: Mul(Silu(x), y) → SwiGLU(x, y); recognized with Silu on either side. Gate: the mul node's backend reports supports_fused(FusedOp::SwiGLU) — true for CPU, Metal and CUDA.
  • FusedBiasRope (removed): the plan's RoPE(Add(x, b), pos) → FusedBiasRope(x, b, pos) rule was implemented but no backend ever advertised FusedOp::BiasRope, so the rewrite was unreachable and Op::FusedBiasRope could never be constructed. Rule, op and capability tag were removed; FusedOp now has a single variant (SwiGLU). The bias+rope work is covered by the build-time fused nodes instead (Metal/CUDA attn_bias_rope_store, which also does the KV store — a strictly stronger fusion).

The pass rewrites op and re-points src; orphaned producers (the old Silu) become dead nodes and are skipped by the scheduler and the allocator. It returns the rewrite count, which the tests use.

One CPU nuance: CpuBackend::supports_fused accepts SwiGLU, and its Op::SwiGLU execution is a single pass (vec_ops::vec_swiglu_f32, dst[i] = silu(gate[i]) * up[i]). It is bit-identical to the vec_silu_f32 + vec_mul_f32 pair it replaced (same formula, same per-element order) but does not allocate the full-size intermediate buffer that pair needed.

Because fusion rewrites the IR, an unfused-vs-fused comparison must run FusionPass on the unfused side too; otherwise silu + mul execute as two kernels and differ from the single swiglu kernel by float noise (~1e-6, amplified at large values). The real forward path always runs the pass.

5.3 Build-time fused ops

OpShape of the winWhere it executes
FusedQKVreplaces 3 matmul + 3 bias + 2 rope + 2 store (10 dispatches) with one concat matmul + one fused bias/rope/store kernel (2)Metal and CUDA
QkvBiasRopeStoremixed-quant layers that cannot share a concat matmul: the 3 matmuls stay, but 3 bias + 2 rope + 2 store become 1 epilogue passCUDA only (D3-8 class 2)
FusedFFNreplaces 2 matmul + silu + mul (4) with one concat matmul + one in-place swiglu (2); gate nf <= 16384Metal and CUDA
FusedQkvNormQwen3's concat matmul + per-head q/k norm + no-bias rope/storeMetal only

All four are part of the topology, so their enable flags are part of CParams and therefore of the reuse identity. Fused vs unfused output is bit-identical on the same backend (verified by fused_qkv_matches_unfused_decode and its Qwen3 counterpart), and the fused kernels are decode-only: nt == 1 is a debug_assert plus a shape gate in each backend.

5.4 Deferred: BatchMatMul

The plan proposed folding three sibling matmuls sharing one activation into one BatchMatMul node, to quantize the activation once on the CPU path. It is not implemented: the IR gives every node exactly one output buffer, and BatchMatMul is inherently multi-output. Expressing it needs either a multi-output node kind or a concat-plus-view encoding, and both are larger IR changes than the measured CPU prefill gain justifies today. The Op::BatchMatMul variant remains in the IR with a comment recording the reason; the corresponding FusedOp::BatchMatMul capability tag was removed (no rule ever probed it).

5.5 Operator vocabulary not emitted

Scale, Softmax, View, Reshape, Permute, AttnMode::Mha and BatchMatMul are represented in the IR for ggml parity but no supported architecture builds them: attention kernels fuse their own softmax, the Qwen2/Qwen3 graphs need no view/reshape node, and the bias+rope rewrite was removed (§5.2). Scale and Softmax are additionally unsupported by Metal and CUDA at execution time.

5.6 Toggles

Env varEffectReuse impact
MINFER_NO_FUSE_QKV=1disables the decode QKV fusion (CParams.fuse_qkv = false)forces a rebuild, by design
MINFER_NO_FUSE_FFN=1disables the decode FFN fusion (CParams.fuse_ffn = false)forces a rebuild
MINFER_DISABLE_MPS=1no Metal participation (CParams.gpu / assignment change)forces a rebuild
MINFER_GRAPH_TRACE=1prints splits + per-op/backend countsnone (diagnostic only)
MINFER_TRACE, MINFER_GRAPH_DUMPper-node data capture / dumps (§8)none

5.7 Measured effect

ChangeModelEffect
G4 FusedQKVQwen2.5-0.5B Q4_0 decode~269 → ~299 tok/s (+11% at KV 440); logits bit-identical
G4 FusedQKVQwen2.5-7B Q4_K_M decodeflat (Q4_K GEMM-bound)
G5 FusedFFNQwen2.5-0.5B Q4_0 decode~303 → ~312–331 tok/s (+3%)
G5 FusedFFNQwen2.5-7B Q4_K_M decodegate off (nf > 16384), unchanged

Numbers are from the phase ledger (§17); docs/CPU_OPTIMIZATIONS.md, docs/METAL_OPTIMIZATIONS.md and docs/CUDA_OPTIMIZATION.md hold the full measurement context.


6. Graph Reuse Mechanism

6.1 The invariant

The graph topology is a deterministic function of GraphParams. Equal params ⇒ identical topology ⇒ the cached graph and its buffers can be reused as-is; only input data is refreshed.

n_past (the KV position) is deliberately absent from GraphParams: it enters through the positions input node. This is the same invariant as llama.cpp's llm_graph_params::allow_reuse.

6.2 Reuse flow

#![allow(unused)]
fn main() {
// GraphCache::try_reuse — params-only
match (&self.prev_params, &self.graph) {
    (Some(prev), Some(_)) if Self::params_match(prev, params) => {
        self.prev_params = Some(params.clone());
        true
    }
    _ => false,
}

// GraphCache::replace_graph — keep the allocator (KV lives there), assign a fresh uid
pub fn replace_graph(&mut self, mut graph: ComputeGraph, params: GraphParams) {
    graph.uid = NEXT_GRAPH_UID.fetch_add(1, Ordering::Relaxed);
    self.graph = Some(graph);
    self.prev_params = Some(params);
}
}

A caller that gets false from try_reuse builds a new graph, assigns backends, runs FusionPass, calls alloc_graph, and stores it. The allocator object is never replaced, so the persistent KV regions survive a prefill→decode rebuild — the transition that changes n_tokens/gtype and therefore necessarily rebuilds.

6.3 What forces a rebuild

ParameterWhy it changes topology
n_tokensmatmul output shapes, KV store shape, n_out < nt decision
n_outdecides whether the G3 tail reduction nodes exist
gtypeDecode vs Prefill, and the decode-only fused branches
cparams.n_ctxsizes the persistent KV regions
cparams.flash_attnselects AttnMode::Flash vs Gqa
cparams.gpubackend assignment is a build-time decision
cparams.fuse_qkv / fuse_ffnfused nodes are topology
weights_versionreserved for weight reload / LoRA switch (constant 1 today)

6.4 Debug structural check

GraphCache::verify_structural(graph) is compiled in debug test builds only. It compares two graphs built from equal params node by node (op with full payloads, out_shape, src) and returns false on any difference. cache.rs has a test asserting that a graph with an extra node fails the check — the guard against a non-deterministic builder silently corrupting reuse.

cache.rs also pins the fusion flags into the reuse identity with a test (fuse_flags_are_part_of_the_reuse_identity): flipping fuse_qkv or fuse_ffn in either direction must force a rebuild, so no one can drop the fields from CParams and silently break the A/B envs.

6.5 uid

uid is monotonic per process, starts at 1, and is assigned only by replace_graph; a reused graph keeps its uid. CUDA uses it as the CUDA Graph cache key together with the node range and the token hint (ComputeGraph::capture_nt_hint returns the first matmul's output row count, None for graphs without matmuls).


7. CPU and Metal Backend Mapping

The full Metal backend design (device/kernel layer, dispatch, command buffers, model wiring, safety) is docs/METAL-BACKEND-DESIGN.md; this section is the condensed mapping of the graph-facing surface.

7.1 CPU (graph/cpu_backend.rs)

CpuBackend is { buffers: Vec<Vec<f32>>, free: Vec<usize>, weights: HashMap<String, Tensor> }. alloc_buffer reuses the first free id whose length matches exactly and zero-fills it; alloc_fresh always appends; free_buffer only returns the id to the list. Re-registering an existing weight is skipped, and that skip is load-bearing: Tensor owns its bytes, so the model call sites' clone() would deep-copy ~4.4 GB on 7B at every graph rebuild (~635 ms measured).

supports_op gates on dtype == DType::F32 for every op, including Input. I32 input nodes are therefore not "supported" by CPU in the capability sense and land on CPU only through the scheduler's or(Some(Backend::CPU)) default. supports_fused accepts only SwiGLU.

OpCPU execution
Inputno-op (host-filled)
Add / Mulvec_ops::vec_add_f32 / vec_mul_f32
Siluvec_ops::vec_silu_f32 in place
Scalecopy + vec_scale_f32
RmsNorm / QkNormper-row rms_norm_fused_f32 (or rms_norm_f32); QkNorm uses d = hd
MatMulF32 weight → vec_ops::mat_mul_f32; quantized → kernel::cpu_quant_matmul_f32 (Q8_0 activations on the fly); optional bias added scalar
GetRowsembedding path (kernel::embed_tokens) or generic gather
RoPElocal cpu_rope (copy then transform)
Softmaxdims 0/1 only; other dims are an Err
SwiGLUvec_ops::vec_swiglu_f32, one pass (dst[i] = silu(gate[i]) * up[i])
KvcacheStoreper-token copy into the K and V regions; pos >= n_ctx is an Err
KvcacheLoadno-op — the node buffer is the K region
Attncpu_gqa_attn, parallelized over heads via kernel::par_for
View/Reshape/Permuteidentity copy
any fused op it does not supportErr("op ... unsupported on CPU ...")

Aliased inputs (Silu, RoPE) are snapshotted into a local Vec before the pool is split, so the kernel reads the producer's values even though output and input are the same buffer. A recorded regression (8a②) is that an F32-weight matmul with nt > 1 must produce token-major [nt][od] output; the earlier [od][nt] layout was silently wrong for prefill only.

7.2 Metal (graph/metal_backend.rs)

MetalBackend holds { state: &'static MpsState, pool, free, staging, free_staging, cb_ptr }. Its pool buffers are shared-mode f32 MTLBuffers, so read_host/write_host are direct memory views. staging/free_staging are capture-only buffers used by capture_split. The backend holds no weight map: every op resolves state.weight_buf(name) -> (MetalBuffer, byte offset) — weights are registered in MpsState as zero-copy newBufferWithBytesNoCopy slices over the mmap'd GGUF parts (MINFER_WEIGHT_COPY=1 forces a copy). One MpsCommandBuffer per split is submitted by synchronize(); Drop flushes a pending buffer.

supports_op accepts Input at any dtype and F32 for the element-wise, norm, matmul, rope, attention, KV, GetRows and decode-fused ops; View/Reshape/Permute are accepted. supports_fused is matches!(fused, FusedOp::SwiGLU) — the only fusion rule that exists (§5.2).

Graph opMpsState method / kernel
Add / Muladd_f32 / mul_f32
Silusilu_f32 (in place)
RmsNormrms_norm_256 when rms_norm_256_enabled() (G2; MINFER_NO_RMS_256 opt-out), else rms_norm
QkNormsame kernels with d = hd, n = len/hd
MatMulquant_matmul_f32_on_gpu_buf; per-ttype *_f32_matmul (nt==1) / *_multi (nt>1) / GEMM (`nt >= 2 && (od >= 2048
GetRowsembed_tokens_gpu (per-quant kernels) or get_rows_f32 for the G3 gather
RoPErope_f32 (copy_in first when not aliased)
SwiGLUswiglu_f32
AttnG1 dispatch (below)
KvcacheStoretwo store_kv calls (K then V); f32 or f16 KV selected by the engine's per-instance kv_format (MINFER_CACHE_TYPE / auto policy)
KvcacheLoadno-op (view of the K region)
FusedQKVconcat quant_matmul_f32_on_gpu_buf + attn_bias_rope_store (3 bias + q/k rope + K/V store in one pass)
FusedFFNconcat matmul + swiglu_f32_off (in-place on the concat buffer, gate at offset 0, up at nf)
FusedQkvNormconcat matmul + two in-place per-head rms_norm[_256] (q at byte offset 0, k at nqt*4) + attn_rope_store
Scale / Softmax / BatchMatMulErr (no kernel)
QkvBiasRopeStoreErr (CUDA-only epilogue)

The *_off kernel variants (swiglu_f32_off, rope_f32(off), rms_norm(off_x, off_y), add_bias_f32(off), store_kv(off), attn_bias_rope_store(bias offsets)) exist so a fused node can operate on a section of one physical buffer without an extra copy.

Attention dispatch (G1)

Pre-dispatch guards return Err before anything is encoded: nkt == n_head_kv * hd, hd == hd_kv, and the layer's kv_pair must exist.

For decode (nt == 1), first match wins:

  1. flash_attn_enabled(hd) (MINFER_NO_FLASH != "1" && hd ∈ {64,128}) → gqa_attn_flash, with the hd128 kernel variants.
  2. else hd ∈ {64,128} and MINFER_NO_SPLIT_ATTN != "1" → gqa_attn_split_f32 (partial + combine).
  3. else → gqa_attn_f32 (classic).

For prefill (nt > 1) with hd ∈ {64,128}:

  1. prefill_flash_enabled(hd) (MINFER_NO_PREFILL_FLASH != "1") → attn_flash_prefill.
  2. else matmul_attn_enabled() (MINFER_NO_MATMUL_ATTN != "1") → attn_parallel_prefill.
  3. else → gqa_attn_f32(..., nt).

Other head dims always use the classic kernel. The KV-parallel chunk count is MINFER_ATTN_CHUNKS or ((max_pos + 1 + 31) / 32).clamp(1, 16). This mirrors the pre-graph path exactly, which is why G1 was bitwise-neutral; the fast paths live in kernels that were already isolated-tested.

7.3 In-place execution and the aliasing rule

Silu and RoPE execute in place whenever the allocator aliased their input (sole consumer, same backend). FusedFFN's swiglu runs in place on the concat buffer; QkvBiasRopeStore runs in place on q's buffer. The backends assume the allocator did the safety analysis — execute_node does not re-check consumer counts. Metal's copy_in snapshots only the non-aliased case; CPU snapshots every aliased input because its kernels are written as slice operations.

The hard rule that came out of Phase 3: never host-copy a buffer with pending GPU work. A per-node host readback inside a split with an open command buffer reads stale data. This surfaced as an all-zero KV region and garbled output before the in-place aliasing rule was introduced.

7.4 Error contract

Kernel-invariant violations (unsupported shape, device-limit guard failure, missing weight) return Err(String) naming the node and the actual values, and abort the run. There is no "delete the node and fall back to CPU" path: assign_backends already decided the backend before execution, and on Metal the submit itself panics with the status when a bounded 10 s wait fails. This is the graph-level expression of docs/GPU_SAFETY.md; metal::gpu_abort remains the device-configuration escape hatch that refuses to risk a GPU fault and exits.


8. Graph Export and Debugging

8.1 DOT (graph/dot.rs)

ComputeGraph::dump_dot(w) writes a Graphviz digraph: one node per CNode labelled with its name and op, filled by backend (Metal light blue, CPU light yellow, CUDA light green, unassigned white), edges from src, and double-circle markers for inputs and outputs.

./target/release/minfer --dump-graph /tmp/g.dot <model.gguf> "hello"
dot -Tpng /tmp/g.dot -o /tmp/g.png

The export path does not reuse the runtime cache: main.rs rebuilds the graph via json::build_runtime_graph (build_graph → enable GPU → assign_backends → FusionPass), so the DOT and JSON exports contain the fusion and backend assignment the live run used. Both flags exit immediately after writing (before decode).

8.2 JSON for the visualizer (graph/json.rs)

export_graph_json / ComputeGraph::export_json emits {format:"minfer-graph", version:1, model, kind:"prefill"|"decode", inputs, outputs, nodes:[{id,name,op,detail,shape,dtype,backend,src,meta}]}. op_name, op_detail and meta_json cover every Op and NodeMeta variant, including the fused ones. The schema reserves per-node stats / values fields for trace data, so the page treats their absence as "no data for this node in this step".

./target/release/minfer --dump-graph-json /tmp/g.json <model.gguf> "hello"
./target/release/minfer viz <model.gguf>          # page + live SSE (default port 8081)

8.3 Per-node data capture

With MINFER_TRACE=<path> the scheduler records, after each node executes, the node's output stats and a downsampled value sample; viz's live mode uses the same machinery over SSE. The capture is backend-aware:

  • CPU outputs are read directly from the pool.
  • Metal outputs are blitted into staging after all of the split's kernels and read back once the split's command buffer is submitted — one submit per split, never a per-node GPU flush.
  • CUDA outputs are queued as stream-ordered D2H copies into pinned capture staging and drained with a single sync at the split boundary; tensors above the staging ceiling fall back to a per-node synchronous copy.
  • Input nodes are host-filled, so they are read directly without a sync; KvcacheLoad has no execution at all and is skipped.
  • KV regions are skipped (a full region per layer would dominate the trace), so the visualizer shows them as "no data".

CUDA Graph replay is disabled while capture is on: a host readback inside a capture window would corrupt the recorded graph.

The trace JSON is {format:"minfer-trace", version:1, model, prompt, phases:[{kind, graph, steps: [{token, text, logits_top, nodes:[{id, dtype, stats, values, stride, n}]}]}]}. viz/README.md documents the page, the SSE endpoints (GET /viz/graph, GET /viz/events, POST /viz/run) and the MINFER_VIZ_DIR sample directory.

8.4 Diagnostics and env vars

Env var / flagEffect
--dump-graph <path>Graphviz DOT export, then exit
--dump-graph-json <path>JSON export for viz/, then exit
MINFER_TRACE=<path>per-node real-data trace (JSON)
MINFER_GRAPH_TRACE=1split list + per-op/backend node counts on stderr
MINFER_GRAPH_DUMP=<dir>logits / per-layer KV dumps (read in the model graph modules)
MINFER_REBUILD_TRACElogs each graph rebuild (Qwen2)
MINFER_DISABLE_MPS=1force CPU participation off for Metal
MINFER_NO_FUSE_QKV / MINFER_NO_FUSE_FFNdisable the decode fusions (A/B)
MINFER_VIZ_DIRdirectory for viz samples (default viz)
MINFER_BENCH_WARMUP_MSbench warmup budget

Resolved (2026-09): export/trace fusion gate. The export/trace paths used to gate fuse_qkv on metal_on only while the Qwen2 runtime gates it on metal_on || cuda_on, so a CUDA-only run exported a graph without the FusedQKV nodes the runtime built — diverging node ids. Both paths (and the model's own CParams construction) now share graph::json::preview_fuse_flags(nt, metal_on, cuda_on), and viz/README.md no longer claims QKV fusion is Metal-only. The gate is pinned by graph::json::tests::preview_fuse_flags_include_cuda_only_runs.


9. CUDA Backend (design level)

The kernel-level optimization campaign — MMQ (int8 quantized GEMM), MMVQ (quantized matrix-vector), FA (flash attention) tiling, weight-plane prepasses, occupancy work and the per-step measurements — is documented in docs/CUDA-BACKEND-DESIGN.md, docs/CUDA_OPTIMIZATION.md and docs/cuda_optimization_steps/. This section covers only what the graph design depends on.

9.1 Structure

CudaBackend (src/graph/cuda_backend.rs, feature = "cuda") wraps the device layer in src/cuda.rs (CudaState, kernel launchers, weight registry, CUDA Graph API). It implements the same Backend trait as CPU and Metal: its own device buffer pool, execute_node dispatch, host read/write, and synchronize. CudaState is a process-wide OnceLock singleton reached through CudaState::get(); CudaBackend::new() returns None when no device is present (or MINFER_DISABLE_CUDA), and the allocator then declines to enable CUDA. The legacy whole-layer cuda.rs::layer_gpu path is no longer driven by inference; the graph backend is the only path.

Device work is serialized through CudaState::stream_lock(), held around every enqueue except while this backend itself owns an open capture window.

9.2 Eligibility

supports_op is f32-activation only (dtype != DType::F32 ⇒ false). Unconditionally supported: Input, Add, Mul, Silu, SwiGLU, RmsNorm, QkNorm, MatMul, Attn, KvcacheStore, KvcacheLoad, View, Reshape, Permute, GetRows, FusedQKV, QkvBiasRopeStore, FusedFFN. RoPE is supported for RopeStyle::NonInterleaved only. FusedQkvNorm is not supported (Qwen3's fused path is Metal-only). supports_fused accepts only FusedOp::SwiGLU.

Weight-quant eligibility is a model-level all-or-nothing gate, not a per-op one: Qwen2Graph::weights_on_cuda / Qwen3Graph::weights_on_cuda requires every matmul weight to be one of Q4_0/Q4_1/Q5_0/Q5_1/Q8_0/Q4_K/Q5_K/Q6_K/F32 and to be registered on the device; on any failure CUDA is not enabled for that model and CParams.gpu records the decision. Kernel-invariant checks that cannot be made at build time (head dims, alignment, nt == 1 for fused nodes, rope style) run inside execute_node and return Err — never a silent CPU fallback.

9.3 Execution dispatch

execute_node maps each op to a CUDA path selected by shape and weight type. Representative mapping (the full table is docs/CUDA-BACKEND-DESIGN.md §4.4 and the per-round records):

OpCUDA path
GetRowsembedding kernels per quant type, or gather_rows_f32 for the G3 tail reduction
Add / Mul / Silu / View-familydevice kernels (identity D2D copy for views)
RmsNorm / QkNormfloat4 rms_norm; the MMQ producer-fused variant when a following GEMM can consume a pre-quantized activation plane
SwiGLUswiglu_f32, or the producer-fused swiglu_quant when the consumer is an MMQ GEMM
MatMulprefill: int8 MMQ (MINFER_MMQ, default on) or f16 wmma GEMM; decode/small-batch: per-quant MMVQ (*_decode_mmvq, _multi for nt 2..8) or f32 kernels
Attngqa_attn_split (decode split-KV), gqa_attn_split_batched (2..16 tokens, per-position bitwise-equal), gqa_attn_f32/_f16kv (prefill; FA tiled prefill for hd 128)
KvcacheStorestore_kv_f32 / store_kv_f16; the output buffer must be the K region
FusedQKVconcat matmul + attn_bias_rope_store (decode, neox, even hd)
QkvBiasRopeStorein-place q + attn_bias_rope_store (mixed-quant class 2)
FusedFFNconcat matmul + in-place offset swiglu (decode)
any op with no kernelErr("cuda: op ... has no kernel ...")

Two implementation details matter to the graph contract: the positions buffer is converted to a real i32 plane on device and memoized per execution window (the causal bound is derived device-side, so no host scalar crosses — a precondition for capture), and the MMQ activation-quantize memo is invalidated at every non-matmul node.

9.4 CUDA Graph capture/replay

graph_replay(uid, range, nt_hint) is the graph-level hook the original plan asked for: a previously captured decode split replays as one captured CUDA Graph launch, and the scheduler skips that split's node loop (continue). Key properties:

  • The cache key is (uid, node range), with capture_nt_hint gating capture to decode-shaped graphs so a prefill graph is not captured accidentally. Prefill capture is default ON since R3-B (MINFER_NO_PREFILL_CAPTURE=1 opts out; MINFER_CAPTURE_PREFILL=1 is accepted but redundant). MINFER_NO_CUDA_GRAPH=1 disables capture entirely.
  • Executions 1 and 2 of a key run direct launches (warmup); on the 3rd the backend takes the stream lock and begins capture. The window is closed by synchronize, which ends capture, instantiates and launches once so the step still produces output; later executions replay.
  • Pointers inside a captured window must stay stable. Pool ids never move memory, but any allocation bumps pool_gen; a captured exec whose pool_gen differs is destroyed and re-captured.
  • Capture is disabled under MINFER_TRACE / live viz, because per-node host readbacks inside the window are illegal. abort_capture ends a window without launching when a node returns Err, and disables graphs for the session (the recorded launches never executed, so that split's outputs are invalid).
  • Replay is refused while this backend already has an open capture window.

9.5 Buffer pool, staging and readback

alloc_buffer matches exact byte sizes in a free list and bumps pool_gen (both on reuse and on cudaMalloc); free_buffer only recycles — persistent KV regions survive rebuilds and the pool keeps device memory. alloc_fresh always allocates and bypasses the free list, which is what the allocator uses for split-boundary staging. OOM is not a panic (a null buffer fails later with a real error) because the backend may be holding the stream lock.

read_host returns None for CUDA: a staged device→host copy cannot return a borrowed slice. The real path is copy_to_host, called by the allocator's copy_to_cpu, which syncs and reads back through a grow-on-demand pinned buffer (MINFER_NO_PINNED_READBACK=1 reverts to a pageable copy). Async H2D input fills use a small pinned ring. The trace path uses a dedicated pinned CaptureStaging (ceiling 128 MiB) so the scheduler can enqueue one stream-ordered D2H per node and drain them with a single sync at the split boundary.


10. ModelDef Trait

#![allow(unused)]
fn main() {
pub trait ModelDef: Send + Sync {
    fn forward(&self, tokens: &[u32], positions: &[usize],
               n_out: usize, n_ctx: usize) -> Vec<f32>;
    fn as_any(&self) -> &dyn std::any::Any;

    fn build_graph(&self, params: &GraphParams) -> ComputeGraph;            // default: unimplemented!
    fn forward_graph(&self, tokens, positions, n_out, n_ctx) -> Vec<f32>;   // default: forward()
    fn forward_graph_cached(&self, tokens, positions, n_out, n_ctx,
                            cache: &mut GraphCache) -> Vec<f32>;            // default: unimplemented!

    fn special_tokens(&self) -> SpecialTokens;
    fn n_layer(&self) -> usize;
    fn n_head_kv(&self) -> usize;
    fn n_embd_head(&self) -> usize;
    fn n_kv_embd(&self) -> usize;
    fn n_vocab(&self) -> usize;
    fn rope_style(&self) -> RopeStyle;
}
}
  • forward is retained as the single-shot convenience entry and both models route it to the graph path: QwenXGraph::forward clamps n_ctx to max_seq_len, locks the process-global graph_cache(), and calls forward_cached. Its pre-#252 legacy kv: &mut KVCache argument was ignored (the KV lives in the graph allocator) and was deleted with the type.
  • build_graph is the immutable-topology constructor; each architecture implements it in models/<arch>/graph.rs.
  • forward_graph_cached is the real primitive used by the CLI, the server, the conversation engine and bench. Callers that need an isolated KV (a server slot, a draft model) must own their GraphCache; the process-global one exists only for the single-shot CLI path.
  • There is no weights_version() method: the value is the GraphParams.weights_version field, set to a constant 1 today.

Dispatch is models/mod.rs::load_model_ns, matching general.architecture (qwen2, qwen3) and passing a weight-registry namespace. The namespace matters when two models are loaded in one process (the D5-R draft): the registry is process-global and name-keyed, and a second model without a prefix would collide and silently drop the primary model to CPU.

ArchitectureGraph moduleHighlights
Qwen2 / Qwen2.5models/qwen2/graph.rsbiases; FusedQKV (concat or CUDA mixed-quant), FusedFFN
Qwen3 (dense)models/qwen3/graph.rsno biases; per-head q/k norm (QkNorm, FusedQkvNorm on Metal), FusedFFN

11. Runtime Wiring

11.1 The reuse kernel

Every path funnels through QwenXGraph::forward_cached:

params = GraphParams { n_tokens, n_out, gtype, cparams, weights_version: 1 }
if !cache.try_reuse(&params) {
    graph = model.build_graph(&params)
    register_graph_weights(...)          // CPU/Metal/CUDA registration
    alloc.enable_metal() / enable_cuda() // per the same gates recorded in cparams.gpu
    scheduler.assign_backends(&mut graph, alloc)
    FusionPass::run(...)
    alloc.alloc_graph(&graph)
    cache.replace_graph(graph, params)
}
let (graph, alloc) = cache.current().unwrap()
alloc.fill_input_i32(graph, "token_ids", tokens)
alloc.fill_input_i32(graph, "positions", positions)
alloc.fill_input_i32(graph, "tail_ids", tail)   // only when n_out < nt
scheduler.execute(graph, alloc)?
logits = alloc.copy_to_cpu(graph.outputs[0])

cparams.gpu is metal_on || cuda_on, and fuse_qkv/fuse_ffn are nt == 1 && gpu minus their env opt-outs, so the params carry the same decisions the builder will make.

11.2 CLI

The CLI computes ctx = params.n_ctx.max(prompt_len) once so prefill and decode size the same KV regions, prefills with model.forward(&input_ids, ..., 1, ctx), then decodes one token per step with model.forward(&[tok], &[pos], ..., 1, ctx). The prefill→decode transition changes n_tokens and gtype, so it rebuilds once; the KV regions survive because the allocator lives in the cache (deviation 14). Decode steps then reuse the same graph indefinitely.

11.3 Server and conversation

  • server/slot.rs gives every slot its own GraphCache and n_ctx_total / n_slots context; a slot's cache is reset per request so a request cannot observe another request's KV.
  • server/chat.rs routes all inference through a catch_unwind-guarded forward_graph_cached call, so a backend panic becomes a 500 instead of a dead worker.
  • conversation.rs owns a GraphEngine { model, cache, n_ctx } for multi-turn sessions; the append- only KV is exactly the persistent regions in that cache.

11.4 Speculative decoding (D5-R)

Speculative decoding is a graph-reuse consumer built on the same primitive:

  • the draft model runs a single-token forward_graph_cached chain against its own GraphCache;
  • the target verifies a draft block of d tokens with one forward_graph_cached call at nt = d + 1, which builds or reuses a Decode-shaped graph for that token count;
  • specverify and the greedy-identity tests compare the spec output against sequential decode byte-for-byte.

Because n_tokens changes as the draft length changes, the target graph rebuilds when the draft depth changes; adaptive depth therefore pins the identity boundary deliberately (docs/SPECULATIVE-DECODING-PLAN.md).

11.5 bench

bench uses a local GraphCache: it warms up the prompt forward for a time budget, then repeats prompt forwards with identical params (the graph is reused; the KV is simply rewritten from position 0), and measures decode by one untimed prefill followed by timed single-token forwards.


12. Measured Results

Representative numbers, all from the phase ledger (§17) or the referenced documents. They are environment-dependent (M4 Pro for Metal, GB10 for CUDA) and greedy-decoded unless noted; treat them as evidence that the design works, not as a benchmark suite.

ChangeModel / pathResult
Graph vs imperative forwardQwen2.5-0.5B Q4_0, prefill + decodelogits max diff 0.000 (bit-identical), KV carried across steps
G1–G3 (attention dispatch, rms_norm_256, tail rows)0.5B decode KV206, Metal~122 → ~256 tok/s (2.1×, ≈ the old path)
G1–G30.5B prefill pp440, Metal~2530–2620 → ~3900–4000 tok/s (+55%)
G1–G37B Q4_K_M decode KV206 / prefill pp206~32.5 → ~49 tok/s; prefill ~217 tok/s (−10% vs old; attention not the bottleneck)
G4 FusedQKV0.5B decode KV440 / KV550~269 → ~299 tok/s (+11%) / 265 → 295; 7B flat
G5 FusedFFN0.5B decode~303 → ~312–331 tok/s (+3%); 7B gate off
CUDA Phase 7 (7e②)7B Q4_K_M decode8.4 → 26.4 tok/s (kernel vectorization)
CUDA Phase 8 / R / MMQ (r56, R4)7B @2K prefill / decode~3212 tok/s (≈1.035× llama.cpp) / ~43–45 tok/s — details in docs/CUDA_OPTIMIZATION.md

Test surface: 83 #[test] across src/graph/*.rs (CUDA backend 40, Metal 14, CPU 6, allocator 5, cache 4, mod 4, builder/scheduler 3 each, fusion 2, dot/json 1 each; ops.rs, params.rs, backend.rs have none directly), plus 2 in vec_ops for the CPU SwiGLU kernel. The model graph modules add their own end-to-end tests. Not all 83 compile in a single configuration because the Metal and CUDA backends are platform/feature-gated.


13. Module Inventory

The plan's file-change manifest, replaced by the landed inventory.

FileLinesRole
src/graph/mod.rs273IR types, topo_order, capture_nt_hint
src/graph/ops.rs319Op, NodeMeta, metadata structs
src/graph/builder.rs496GraphBuilder
src/graph/params.rs71GraphType, CParams, GraphParams
src/graph/cache.rs235GraphCache, uid, structural check
src/graph/backend.rs95Backend, KvProvider
src/graph/alloc.rs795liveness allocator, KV regions, staging
src/graph/scheduler.rs519assign / split / execute, capture hook
src/graph/fusion.rs151FusionPass
src/graph/cpu_backend.rs916CPU executor
src/graph/metal_backend.rs2081Metal executor
src/graph/cuda_backend.rs6662CUDA executor + graph capture
src/graph/dot.rs80DOT export
src/graph/json.rs357JSON export, preview_fuse_flags
src/models/qwen2/graph.rs2253Qwen2 build + weights + tests
src/models/qwen3/graph.rs1132Qwen3 build + weights + tests

src/models/qwen2/forward.rs was deleted in Phase 6; the imperative path no longer exists. src/metal/ and src/cuda.rs remain the per-op kernel/device layers that the graph backends wrap; src/cuda/kernels/*.cu holds the CUDA kernels (common.cuh + 17 translation units since #263). src/cache.rs (the legacy KVCache the graph path ignored) was deleted in #252 — the allocator owns KV.


14. Implementation Timeline

The plan's build order, with landed status:

PhaseContentStatus
1IR + buffer infrastructure: mod.rs, ops.rs, builder.rs, alloc.rs✅
2CPU backend: backend.rs + cpu_backend.rs (+ minimal executor)✅
3Metal fine-grained backend: metal_backend.rs + cross-backend scheduling✅
4Scheduling + fusion + debugging: scheduler.rs, fusion.rs, dot.rs, cache.rs, params.rs✅
5Qwen2 graph construction: models/qwen2/graph.rs✅
6Wiring + cleanup: ModelDef routes to the graph path, forward.rs deleted✅
7CUDA backend wrapping cuda.rs, preserving CUDA Graph✅
8Verification: old/new logits, 7B GPU, --dump-graph✅
9Metal wiring optimization G1/G2/G3 + allocator liveness fixes✅
10G4 decode QKV fusion (Op::FusedQKV)✅
11G5 decode FFN fusion (Op::FusedFFN)✅
12+Post-plan workstreams (Qwen3, CUDA phases, viz, spec decode, bench) — see §17.2✅

15. Design Decisions, Invariants and Risks

15.1 Standing decisions

These were deviations from the original plan that are now deliberate design:

  1. NodeMeta is a concrete enum, not Box<dyn Any + Send + Sync>: PartialEq for the structural check, no downcast panic, CNode: Clone. The cost is that a new metadata kind edits the enum.
  2. Op::Input is a node kind, so inputs are visible in the IR and in ComputeGraph::inputs.
  3. The allocator is the single owner of every buffer. Backends own pools; the scheduler is a pure orchestrator (assign_backends queries allocator capabilities, execution goes through alloc.cpu_mut() / metal_mut() / cuda_mut()).
  4. In-place ops alias their input when it is their sole consumer on the same backend. This is what makes kernel-order execution correct without host copies.
  5. Execution follows build order, and so does liveness.
  6. Inputs are never freed by liveness, exactly like outputs.
  7. Fusion flags are part of the reuse identity. CParams carries gpu, fuse_qkv, fuse_ffn; toggling any of them forces a rebuild, which is what makes the env A/Bs valid.
  8. Fused decode nodes are built, not pattern-matched. The builder knows the layer, the biases and the concat weights; the peephole pass only handles the backend-generic SwiGLU rewrite.

15.2 Risks and status

RiskStatus
KV position coupling (the original plan's fatal flaw)Resolved: positions are input data; no topology depends on n_past
Reuse decision too loose / missing payloadsResolved: params-only comparison plus a debug structural check; fusion flags pinned by a test
Per-node host round trips across backendsResolved: buffer-id interface; cross-backend copies only at split boundaries
Lost n_out optimizationResolved: G3 GetRows tail reduction; n_out is in the reuse identity
Metal per-op dispatch overheadResolved: one command buffer per split; G1 fast-path dispatch is bitwise-neutral and ~2.1× on 0.5B decode
Cross-backend KV copyResolved: KV regions are created on, and resolved from, the layer's backend
Double fusionResolved: supports_fused gates the pass; build-time fused ops bypass it entirely
GPU safety rules brokenResolved: execute_node returns Err; docs/GPU_SAFETY.md rules are enforced in the guards
CUDA functionality regressionResolved: CUDA Graph capture/replay preserved and default-on for decode
BatchMatMul fusionOpen by design: deferred (single-output IR); recorded in §5.4
FusedBiasRope ruleRemoved: no backend ever claimed the capability; the op, rule and tag are gone (§5.2)
Dump/trace fusion gate differs from runtime on CUDAResolved: shared preview_fuse_flags gate + test (§8.4)

16. References

Internal

TopicWhere
IR / builder / scheduler / allocator walkthroughdocs/inference_e2e_walkthrough/05-graph-builder-ir.md … 08-scheduler-execute.md
Architecture overview, adding an architecturedocs/ARCHITECTURE.md
Backend overview (CPU/Metal/CUDA, feature gates)docs/BACKENDS.md
CUDA backend design + Phase 7 recorddocs/CUDA-BACKEND-DESIGN.md
CUDA optimization history + env-gate referencedocs/CUDA_OPTIMIZATION.md, docs/cuda_optimization_steps/
Metal optimization plansdocs/METAL_OPTIMIZATIONS.md
CPU optimization recorddocs/CPU_OPTIMIZATIONS.md
GPU safety rulesdocs/GPU_SAFETY.md
Speculative decoding (D5)docs/SPECULATIVE-DECODING-PLAN.md
Graph visualizer (JSON, trace, live SSE)viz/README.md
Debug dumpsdocs/debug-dump.md
Qwen2 / Qwen3 supportdocs/QWEN2-SUPPORT.md, docs/QWEN3-SUPPORT-PLAN.md

llama.cpp analogues

ConceptReference
Graph IRggml/src/ggml-impl.h — ggml_cgraph
Reuse decision (params only, no n_past)src/llama-graph.h — llm_graph_params::allow_reuse
Graph construction contextsrc/llama-graph.h — llm_graph_context
Result container / reusesrc/llama-graph.h — llm_graph_result, can_reuse()
Backend split graphggml/src/ggml-backend.cpp — split_graph()
Per-split executionggml/src/ggml-backend.cpp — compute_splits
Tail rowssrc/models/llama.cpp — ggml_get_rows(cur, inp_out_ids)
In-place rope/silu, swiglu splitggml op semantics; minfer's *_off kernels mirror them

17. Implementation Record

17.1 Phase ledger

The G-phase rows below are the plan-era record. Their original commit hashes no longer resolve because the repository history was rewritten; the resolvable documentation commits are 8d7cb38 (G1–G3), 96404fb (G4), ec922f1 (G5). Post-plan work is in §17.2.

PhaseContentStatusEvidence
1IR infrastructure (mod.rs, ops.rs, builder.rs, alloc.rs)✅unit tests: topo sort / cycle detection, op payload equality, chained liveness reuse, parallel-chain isolation, input fill, persistent KV sharing
2CPU backend (backend.rs, cpu_backend.rs, minimal executor)✅graph vs manual computation (matmul+bias+silu+scale), rms_norm, embedding+rope, KV store/load + GQA attention
3Metal backend + cross-backend scheduling✅15 tests: element-wise / norm / Q4_0+Q8_0 matmul vs reference, cross-backend copies, multi-split alternation, KV + decode, layer-0 full-Metal vs CPU bit-identical; CLI GPU output fluent
4Scheduler + fusion + diagnostics (scheduler.rs, fusion.rs, dot.rs, cache.rs, params.rs)✅fusion apply/gate, DOT format, cache reuse semantics, split boundaries
5Qwen2 graph construction✅real model logits: graph vs imperative max diff 0.000 (prefill + decode, KV across steps)
6Wiring + cleanup: forward.rs deleted, ModelDef on the graph path, --graph a compatibility no-op✅full suite green after deletion; CLI output consistent
7CUDA backend wrapping cuda.rs (7a–7e: skeleton, per-op dispatch, graph integration + staging, CUDA Graph, kernel/path optimizations)✅GB10: 7B Q4_K_M decode 8.4 → 26.4 tok/s; test suites pass; see docs/CUDA-BACKEND-DESIGN.md
8Verification: old/new logits, 7B GPU, --dump-graph export, graph path default✅0.5B Q4_0 bit-identical; 7B Q4_K_M GPU fluent (~42 tok/s at the time); 437-node DOT
9 (G1–G3)Metal attention dispatch, rms_norm_256, n_out tail rows; two allocator liveness fixes✅0.5B decode 2.1×; prefill +55%; 7B decode ≈ old; tail-reduction per-node comparison 0.000
10 (G4)Decode QKV fusion Op::FusedQKV (concat matmul + bias/rope/store), flags in CParams, env revert✅0.5B decode +11%; fused vs unfused logits 0.000 (0.5B + 7B); fused_qkv_matches_unfused_decode
11 (G5)Decode FFN fusion Op::FusedFFN (gate+up concat + in-place swiglu), nf <= 16384 gate✅0.5B decode +3%; 7B unchanged (gate off); fused vs unfused logits 0.000

17.2 Post-plan workstreams

After G5 (2026-08-22) the graph path gained 12 workstreams (101 commits through HEAD 5471680):

WorkstreamGraph impactRepresentative commits
Chat API / server / conversation plumbingexplicit n_ctx, per-slot GraphCache, conversation engine cachef8f1124, 55296f2, 5866d36, 99f2188, 98f05bb, 921bb4c, b6c9473
Qwen3 dense on the graph pathsecond architecture: QkNorm, FusedQkvNorm, Metal fixes; KV sized by --n-ctx283c7d6, 94d57ac, d5b8023, 6756d42
Viz live + CLI subcommandsscheduler capture/live events, JSON export, viz/serve subcommandseac8180, c95df6a, 163cb6c, 5b22353, 2a70e35, 071b5f6, ed5fc9b
Metal objc2 migrationbackend internals re-homed; no topology change6a382a3, be3df55, ee9b65b, 94d57ac
CUDA Phase 7a–7ecuda_backend.rs, per-op dispatch, staging, CUDA Graph, FusedFFN port, async H2Ddfa3516, ad1512d, 8fb88f8, 4fcd0d8, 7123adb, 92ad586, 8d2ee6f, 78d410a, f54f721, 082095c
CPU graph correctnesstail_ids, buffer-pool race, cross-buffer handling, token-major F32 matmul, fusion flags in the reuse identity2ed3eb1, b849601, 058ea97
CUDA Phase 8 campaign (8b–8p)f16 KV, prefill GEMMs, split-K decode attention, wmma/FA prefill, MMVQ; graph-side 8o removes ~1.6 s per-rebuild CPU stallsf7b0036, 69a27c5, a5af60f, b959ec9, eb24054, b7b8e73, 1298cb2, 1d28235, acca28f, ba3f317, cb66fca, 65b686c, 2992f57
CUDA R1–R4 + capture defaultsint8 MMQ, MMVQ weight streaming, pinned D2H, prefill capture default-on, R3-A1 tail_ids at the graph head, decode split attention40e97c9, 6df3245, 761e236, a213c89, 70f57db, 86ca78c, 4d8c666, f45c241
MMQ campaign r34–r60no topology change; weight-plane precompute + producer-fused A-quantize; gates default-on87a75a3, cf1ed4b, 910d967, 83fee77, 4cf7c74, 36a481f, 57edcf6
D3-x CUDA decodedecode attention dispatch, D3-5/D3-7, G4 FusedQKV ported to CUDA (class 1 concat + class 2 mixed-quant epilogue)22336b2, 3230b2b, 92b0712, b3b6077, 3857633, a448a4a, 9f419f9
bench subcommandharness around forward_graph_cached (pp/tg, warmup budget)646f22e, cdc1e21, 62b9b90
Speculative decoding D5 → D5-Rdraft nt=1 chain + target verify at nt=d+1; specverify; adaptive depth; conversation/server integration; draft weight namespacing3fd0880, a6b7cf3, 90ef8ea, 8d7ce50, c3d4bb1, 0fe132f, 39eceaa, c26c114

17.3 Deviations from the plan

Items 1–26 are the plan-era deviations (kept, lightly updated). Items 27+ are post-plan changes that a reader of the original plan would not find there.

  1. NodeMeta is a concrete enum, not Box<dyn Any + Send + Sync>.
  2. Op::Input leaf node; ComputeGraph::inputs records the ids.
  3. rope() carries style, and kvcache_store/load carry n_ctx so the persistent region has a fixed length; only the written prefix is read at execution time.
  4. (Superseded by 20) The first KV layout was one contiguous [K | V] block per layer.
  5. attn() takes a pos input: the causal mask needs pos[t] + 1, matching llama.cpp's idx-dependent mask.
  6. Metadata expansion beyond the plan's sketch: EmbedMeta.weight_name, RoPEMeta.n_head/hd, AttnMeta.nkt/scale/layer, plus the four fused metas.
  7. Backend::execute_node takes &mut self (the pool is mutated), and the allocator is the single owner of the pool (deviation 11).
  8. I32 inputs are stored as f32::from_bits bit patterns (exact for |v| < 2^24), filled through fill_input_i32.
  9. The CPU backend supports F32-weight matmul through vec_ops::mat_mul_f32; the quantized path quantizes activations to Q8_0 on the fly.
  10. BatchMatMul fusion deferred: the single-output IR cannot express a multi-output node.
  11. Allocator single ownership (final form): the scheduler is a pure orchestrator; execution goes through alloc.cpu_mut()/metal_mut()/cuda_mut().
  12. GraphParams is the entire reuse basis; n_past is explicitly absent; weights_version is reserved (constant 1 today).
  13. Weight layout is the GGUF layout: metadata [in, out], memory [out][in] row-major, so od = shape[1], id = shape[0].
  14. The KV persistent region survives graph rebuilds: GraphCache owns the allocator, so a prefill→decode rebuild swaps only the graph.
  15. The scheduler executes in build order, which guarantees a KV store precedes the attention that reads it; fusion-orphaned nodes are skipped.
  16. G3 tail rows implemented: GetRows(wo/residual, tail_ids) after the last layer's wo; forward.rs's old "compute all, slice the tail" no longer exists.
  17. forward.rs deleted in Phase 6; embed_tokens moved to kernel.rs; --graph is a compatibility no-op.
  18. In-place op buffer aliasing (the key Phase-3 fix): Silu/RoPE alias their input when it is the sole consumer on the same backend; this fixed the all-zero-KV corruption.
  19. Backend configuration is part of the reuse identity (CParams.gpu): MPS/CUDA initialization changes force a rebuild.
  20. Each layer's KV is two independent persistent regions (kv.{layer}.k, kv.{layer}.v), exposed through kv_pair(layer); the plan's contiguous [K | V] was replaced.
  21. Metal G1/G2 wiring: attention dispatches by nt/hd to flash/split/parallel/classic with the same gating and env vars as the old path; RmsNorm selects rms_norm_256 when enabled.
  22. Liveness must match the scheduler's execution order (G3 fix): using topo_order() for liveness let a later node reuse an input the scheduler had not read yet (logits off by 21.79). Liveness now uses build order; topo_order() only validates acyclicity.
  23. Input buffers are never freed (G3 fix): inputs are host-filled before execution, so two inputs sharing a liveness block let the later fill clobber the earlier one.
  24. G1–G3 measured (M4 Pro, greedy): 0.5B decode KV206 ~122 → ~256 tok/s; prefill pp440 ~3900–4000 tok/s (+55%); 7B decode ~49 tok/s; 7B prefill ~217 tok/s (−10%). Greedy output per token identical to pre-G1.
  25. FusedQKV (G4): decode builds one node (concat matmul + attn_bias_rope_store); attn() takes its output shape from AttnMeta, not from the (larger) fused q buffer; CParams.fuse_qkv is in the reuse identity; CPU never builds it.
  26. FusedFFN (G5): decode builds concat gate+up matmul + in-place offset swiglu; the down matmul reads rows 0..nf; gated nf <= 16384 because the 7B concat measured slower. Debugging lesson: an unfused comparison must run FusionPass, or two-kernel silu+mul vs one swiglu kernel differ by ~1e-6.

Post-plan additions:

  1. Qwen3 on the graph path: Op::QkNorm (per-head RMSNorm) and Op::FusedQkvNorm (concat matmul + qk-norm + no-bias rope/store). The fused variant is Metal-only; CUDA Qwen3 runs the unfused qk_norm chain.
  2. D3-8 class 2 (Op::QkvBiasRopeStore): mixed-quant decode layers (e.g. Q6_K attn_v among Q4_K q/k) run three separate matmuls plus one bias/rope/store epilogue, CUDA-only. On Metal those layers keep the unfused chain, bitwise-neutral.
  3. G4 FusedQKV ported to CUDA (class 1 concat) — Op::FusedQKV is no longer Metal-only.
  4. CUDA Graph replay: Backend::graph_replay + ComputeGraph::capture_nt_hint, keyed by (uid, node range), invalidated by pool_gen; two direct-launch warmups, capture on the third; prefill capture default-on. The scheduler disables replay under trace/live capture.
  5. alloc_fresh added to the Backend trait for split-boundary staging that must not be recycled mid-execute.
  6. R3-A1: tail_ids is declared at the graph head, not beside its consumer, to avoid extra CPU/CUDA split boundaries. R3-A2: the logits buffer is exactly n_out * n_vocab, so the per-step full-logits clone was dropped.
  7. weights_version is static 1: there is no weights_version() trait method and no LoRA path yet; the counter exists but is unused.
  8. Export/debug surface: --dump-graph, --dump-graph-json, MINFER_TRACE (P2) and the viz live path (P3) are graph consumers added after the plan.
  9. Runtime consumers of forward_graph_cached: the server's per-slot caches, the conversation engine and the D5-R draft/verify loops.
  10. Export/trace fusion gate (fixed): the export and trace paths used to gate fuse_qkv on Metal only while the Qwen2 runtime used metal_on || cuda_on, so CUDA previews lacked FusedQKV and node ids diverged from live events. Both call sites now share graph::json::preview_fuse_flags, pinned by a unit test (§8.4).

Post-baseline additions (working-tree changes made while writing this document, after HEAD 5471680):

  1. CPU single-pass SwiGLU: vec_ops::vec_swiglu_f32 replaces the vec_silu_f32 + vec_mul_f32 + full-size temporary pair in CpuBackend::execute_node. Bit-identical to the pair (same formula and per-element order); pinned by vec_ops::tests::swiglu_matches_silu_then_mul.
  2. Dormant FusedBiasRope removed: Op::FusedBiasRope, FusionPass::fuse_bias_rope, the FusedOp::BiasRope tag and the never-probed FusedOp::BatchMatMul/FusedOp::QKVBiasRopeStore tags are gone; FusedOp now has a single variant. Metal's supports_fused no longer advertises a tag whose op it cannot execute (§5.2/§5.4).
  3. Shared preview gate: graph::json::preview_fuse_flags(nt, metal_on, cuda_on) is the single source for the --dump-graph* / MINFER_TRACE preview's fusion flags (item 36).
  4. The compile-time backend enum became an opaque handle (F4, #57): pub enum Backend { CPU, Metal, Cuda } is now pub struct Backend(u16) over a fixed, configuration-independent id space, with the name-keyed registry in graph/registry.rs (§3.6, the design record is BACKEND-REGISTRY-DESIGN.md). Every consumer that matched on the enum now reads the registry: the allocator's twelve dispatch sites (pool, sync, copy_cells, host I/O, KV format, session enable), the scheduler's execute, the fusion wiring in graph/json.rs + both model graphs, the JSON/DOT exporters, the KV-session tag table and the op-matrix harness. The two orders are preserved and pinned: identity (id order — the session tag, the exporters, Ord) and priority (Metal > CUDA > CPU, the assignment preference). One per-backend branch deliberately remains: the MINFER_TRACE/viz capture path dispatches CPU / Metal / CUDA because each backend captures through its own mechanism (a borrowed read, a blit into Metal staging, an async D2H into the pinned CUDA staging), which is not a capability.

17.4 Test surface and baseline

83 inline #[test] functions live in src/graph/*.rs (distribution in §12), plus 2 in vec_ops for the CPU SwiGLU kernel. The CUDA backend dominates because capture/replay needs bit-parity tests (cuda_graph_replay_bit_parity, cuda_prefill_shaped_graph_never_captures, cuda_multisplit_capture_bit_parity, cuda_graph_recaptures_on_pool_gen_change, and the prefill capture parity harness). The model graph modules add end-to-end tests that compare graph and reference logits.

Two plan-era baseline notes are now historical: the attn_parallel_realdata_correctness test needs a real dump under /tmp/dp3 (environment-dependent, a pre-existing failure), and the "46 pass / 1 fail" figure dates from Phase 1 — later rows of the ledger record larger suites (80/81 for G3–G5, 144+130 for the CUDA phases). Run cargo test --release for the current number; a bare cargo test --release does not compile the CUDA backend unless --features cuda is passed.

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0001 — Inference runs through one declarative compute graph
  • ADR-0002 — Topology is a function of GraphParams alone, so positions cannot be structure
  • ADR-0009 — A failure is an error, never a silent fallback
  • ADR-0010 — The identity gate: bitwise by default, a named tolerance class otherwise
  • ADR-0013 — The CPU quantizes activations to Q8_0; a device reads f32
  • ADR-0004 — The batching default follows the device
  • ADR-0006 — The KV storage format is a per-engine gate, not a process-wide global

KV Cache Design — cells, format, sessions

Scope. The contract for the KV cache: how a run resolves its cells, what the cell store guarantees, what MINFER_CACHE_TYPE is allowed to do, and what a session file contains. The graph-side invariants are in COMPUTE-GRAPH-DESIGN.md; the per-ticket history (C1–C8, #99, #130, #153, #306, #310) is in ARCHITECTURE-EXECUTION-PLAN.md §5.

AGENTS.md states each rule below as a one-line invariant and links here for the elaboration.

1. Positions and cells

  1. KV positions are data, not structure — topology never depends on n_past (precondition for decode reuse). So is the allowed window: attn_span (E1) carries each query's [lo, hi) cell range, resolved by KvCache from its per-sequence span list (position base, first cell, length) — not start + position. A read path refuses a multi-span sequence loudly; a set-valued window (shared prefix + private run) goes through the kv_map input instead (C8b S2/S4, gathered by all three backends — Metal's since #362). A backend expresses its window in one of three modes, selected by the size of the window input: causal positions, one [lo, hi) pair per query, or KV_MAP_MAX_SPANS (cell, len) runs. Op::Attn { explicit_span } marks a node positions cannot bound (more than one sequence, or a window not starting at cell 0); only a backend with supports_attn_span() (CPU, CUDA and — since #44 part (a) landed on a Mac 2026-10-06 — Metal, via kernel_gqa_attn_window_f32/_f16 in src/metal/kernels/attn_window.metal) may take it. Metal reads both explicit layouts — the one-range attn_span since #44 part (a) and the set-valued kv_map since #362, through the sibling kernel_gqa_attn_map_f32/_f16 — and a packed q8_0 KV cache is read on Metal since #310 (mechanism A's native packed decode plus mechanism B's f32 stage, crate::metal::packed_attn_route). The write/move side landed in #44 part (b) on a Mac 2026-10-06: copy_cells (C3 compaction), the copy_kv_to_cpu Metal arm (C2 shift / C5 sessions, f32-only) and the per-engine kv_format. graph/batch.rs composes such batches; forward_cached is its one-sequence case. A store may never write into a prefix a sequence reads in place: production has one fill entry point, GraphAllocator::fill_batch_inputs, and it runs the copy-on-write (GraphAllocator::kv_private_row_for, C8b S3) for every group before the first cell is resolved — the share shrinks to the first position that forward writes and the run's own rows shift up inside it. The kv_cells_for_seq refusal of a position still inside the share is a belt-and-braces invariant check on the store resolver, not a second entry point: production cannot reach it with a shared position, because fill_batch_inputs copied first (E1's test-only fill_attn_inputs and its C1 remnant kv_note_used/own_prefix were deleted in #228). GraphAllocator::kv_cell_of is that same read side, but test-only (#236): production reads a sharing sequence's rows as windows (the kv_map input above) and snapshots the whole arena (kv_save*, C5), and the one consumer is the C8b S3 gate's kv_rows_of in server::batch::tests. The windowed Metal path now has fast prefill families (kernel_flash_attn_window_blk_* for the one-range attn_span, #359; kernel_flash_attn_window_map_* for the set-valued kv_map, #369 — both in src/metal/kernels/fa_window.metal) selected for nt > 1 at hd ∈ {64,128}, alongside the #44/#362 correctness families (which stay byte-untouched and serve the other shapes), so a batched multi-sequence prefill is no longer the slow path — #315 measured the correctness families, #359 and #369 the fast ones.

2. The cell store: ownership, removal, compaction, sharing

  1. Each layer owns two persistent KV regions (K/V) via kv_pair(layer); they survive rebuilds (the allocator lives in GraphCache). kvcache.rs tracks the owner of every cell (C1), can drop a row position in place (C2), can compact the arena (C3), lets a sequence read another's prefix in place (C8b S2) and copies a row out of that prefix the moment a store would land on it (C8b S3). C6 splits position from cell: positions is a token's index within its sequence (what RoPE rotates by); the allocator resolves cells — the row to write — from the run's span list. They coincide only while a run starts at cell 0, which is why the single-sequence path is bitwise unchanged. The fused decode QKV family (Op::FusedQKV, Op::QkvBiasRopeStore) consumes cells too — positions ropes, cells stores ((cuda_on || !explicit_span); Metal keeps the pre-C6 gate, G5) — and whoever feeds a batch must pass sequence-relative positions, never run.start + pos. A physical kv_rm/kv_shift (C2) changes positions, so it still re-ropes; a compaction (C3) changes only cells, so rows move verbatim (Backend::copy_cells: CPU copy_within; CUDA's kv_move_rows — one row at a time with a barrier, ascending when the run slides down and descending when it slides up (C7b), no staging buffer, because overlapping device-to-device cudaMemcpy is undefined; Metal's arm walks one row at a time via MTLBlitCommandEncoder in the same order, since #44 part (b) 2026-10-06) and kv_defrag(need) takes no rope. Whoever copies must use kvcache::order_moves — upward moves top-down, downward bottom-up, upward first — because a destination may never land on a row that has not been copied yet. Whoever caches a run's start (E2's server keeps one per slot) must apply the moves kv_defrag / kv_reserve_seq_with_defrag return.

3. The storage format is a gate, never a guess

  1. The KV storage format is a gate, never a guess (C4), and it is per engine (#99, CUDA device half #153). MINFER_CACHE_TYPE is parsed strictly into KvFormat { f32, f16, q8_0 } (graph/kvformat.rs, the single authority) and resolved once per load against the device in models::load_model_configured; the resolver folds in the GPU's own auto policy (kvformat::auto_device_format: f16 for the 7B class, f32 for small models, never on the CPU), so one answer drives both the region width and the kernel. The answer is stored on the loaded engine (ModelDef::kv_format/set_kv_format) and reaches the graph through CParams::kv_format (part of the reuse identity: the builder stamps each KV node's KvcacheMeta::row_elems from it) and the CPU, CUDA and Metal kernels through GraphAllocator::set_kv_format. There is deliberately no process-wide kvformat global (the decision, and the measured failure that forced it, are in ADR-0006), and no process-wide CUDA layout tag either: cuda::layout_of/format_of are the only KvFormat ↔ FFI-tag binding, and the tag is part of the captured-graph identity — cuda_backend.rs::graph_replay_step refuses an exec whose recorded kv_layout moved, exactly like a pool_gen change. An unknown value fails the load on every device; q8_0 fails on a backend whose attention kernel has no packed read — the registry's BackendCaps::reads_packed_kv, true for CPU, CUDA and Metal (#310 enabled it on Metal via mechanisms A and B). A Q8_0 cell is a whole number of Q8_0 blocks rounded up to whole f32 words, so elems / n_ctx is one cell's width and every copy_cells move stays verbatim; KvcacheMeta::row_elems carries that width while the node's shape stays logical, and ensure_kv refuses a packed width on a backend without the capability. The store quantizes with one quantizer on both backends (store_kv_q8_0: amax/127, f16 scale, round-ties-even), and attention reads the packed blocks directly — on the CPU the K score is dot_q8_0_q8_0 against the Q8_0-quantized query row and V accumulates out of the cell (kvformat::accumulate_q8_0_row; MINFER_NO_FUSED_Q8_KV restores S1's dequantize-into-scratch A/B) — while on CUDA the layout-tagged kv4<KV_LAYOUT_Q8_0> dequantizes each 4-element group. A physical kv_rm/kv_shift works on a packed region: survivors move verbatim and each is mapped through kvformat::map_q8_0_cells (dequantize → re-rope → requantize). Honest scope: a physical shift on an f16 region refuses loudly, and that refusal is pinned by a gate (graph::alloc::tests::kv_shift::an_f16_region_refuses_the_physical_shift_and_the_other_formats_take_it, with f32 and packed q8_0 controls on the same fixture) (#306 — the host round trip has no hunk→f32→hunk map; CUDA is exposed too), so the CLI --cnv overflow falls back to re-rendering the retained window (measured dgxspark (aarch64, GB10 sm_121), 2026-10-07, 7B Q4_K_M MINFER_CACHE_TYPE=f16 ./target/release/minfer --cnv --n-ctx 512 -n 8: 499–505 tokens re-prefilled per overflow in 0.30–0.35 s, ≈ +0.09 s over the f32 shift's 22-token / 0.26 s delta, and cheaper than the f32 fallback's own 0.72–0.76 s), and a speculative session refuses a packed cache (spec::SpecEngine::new) because the batched split kernel's bitwise identity with sequential decode is what its greedy contract rests on.

4. A session is a file with a header, never a memory dump

  1. A KV session is a file with a header, never a memory dump (C5). graph/kvsession.rs writes a versioned, checksummed container: shape, backend, KV element type (the header's flags word: FLAG_PACKED for Q8_0, FLAG_F16 for f16, 0 for f32; the two element-type bits are mutually exclusive and an unknown bit is refused loudly, so a pre-#130 flags == 0 file still loads as f32 and an older build refuses a newer f16 file instead of decoding it as f32) — one K/V blob per layer as pool words, then the owner table + run table + span lists + written extents, then — version 2, C5 S2 — an opaque, length-prefixed host-state blob inside the checksum: the KV rows belong to a host state, so the container carries both or neither and a version-1 file is refused loudly. GraphAllocator::kv_save/kv_load stream it through the backends' host I/O and enable the pool a file names if the graph has not been built yet. kv_load verifies the whole file before it applies anything, so a truncated, corrupted or foreign file is a no-op, and the header must describe the run that loads it (backend, n_ctx, n_embd, element type). KvCache::restore_session validates the bookkeeping — arena capacity, owner-table length, every reservation and span inside the arena, every live sequence carrying a span list — before applying. The real-model gate asserts the resumed session continues bitwise (max |Δlogit| = 0). CLI (--session FILE under --cnv): keeps the JSON history and writes FILE.kv beside it on exit; a matching companion is resumed with no prefill (the host state — messages, stream_tokens, current_pos, turn_pos, prev_tokens, need_insert_eot — rides in the container as a versioned JSON blob), and everything else (another --n-ctx/model/MINFER_CACHE_TYPE, an edited history, an unknown snapshot version, an engine that cannot hand its KV to the host) prints the reason and re-renders. --slots-file <PATH> is the server twin (S2b): it snapshots the batched engine's slot table after every completed request and resumes it at startup, so a request whose prompt matches a restored slot prefills only its delta; the in-flight request is deliberately not in the snapshot. A mixed offload plan has KV on two backends, so kv_load refuses a session (one arena per file).

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0002 — Topology is a function of GraphParams alone, so positions cannot be structure
  • ADR-0006 — The KV storage format is a per-engine gate, not a process-wide global
  • ADR-0014 — A KV session is a versioned, checksummed file — never a memory dump
  • ADR-0005 — Metal becomes a first-class backend
  • ADR-0011 — Backend ids are a file-format contract: appended, never renumbered

Memory Policy — accounting and layer offload

Scope. How much memory a backend may use, when the refusal happens, and how the layer offload plan is resolved. Both sections are decisions (E4, E5) whose per-ticket records are in ARCHITECTURE-EXECUTION-PLAN.md §7.

AGENTS.md states each rule below as a one-line invariant and links here for the elaboration.

1. Account before allocating (E4)

  1. Memory is accounted before it is allocated, and the length contract lives in the BufRef (E4). alloc_in_pool is fallible: weights (Backend::weights_bytes) + pooled + this allocation at its size class is checked against the backend's budget before the pool is asked for anything, and the refusal names the numbers (weights / pooled / request / budget, in MiB). The default budget is the backend's own answer — CUDA's current free bytes, or since #53 Metal's recommendedMaxWorkingSetSize, with a quarter held back — resolved from an explicit allocplan::DeviceMemory outcome (Reported / QueryFailed / NoDevice) through the pure allocplan::budget_decision; a failed query is not a zero budget — it prints the real device error name once and charges weights only, while a measured zero still refuses (#122); CPU is unbounded unless GraphAllocator::set_memory_budget sets one. memory_report(backend) is the accounting surface: pool_bytes (what the pool holds — it only grows when the pool creates a buffer, so a recycled class buffer is not charged twice; Backend::pool_len is the probe), live_bytes, peak_live_bytes, weights_bytes, budget, headroom_bytes(). The class ladder lives in graph/allocplan.rs (class_size: powers of two to 16 KiB, then 16 KiB steps) and the pools allocate at it, so two shapes in one class share a buffer across a rebuild. Two rules follow, both latent bugs that rounding turned real: (a) a node's logical length is BufRef::len — fill_input checks the data against it and writes through Backend::write_host_window, and every capture read is windowed (scheduler::window_of); write_host keeps the exact-length contract for the persistent KV regions and staging; (b) an input is host-filled before execution, so it must never take a buffer this build's sweep released — all inputs are placed before the walk — and a liveness extension (in-place alias, D1 view) must move the buffer's buf_alive deadline too (extend_through_views → extend_buffer_alive). S3 splits reservation from assignment: a released classed buffer goes to a reservation table (slots, keyed by (backend, class)), and alloc_class_in_pool takes the smallest idle id of the class (deterministic — a LIFO list would move a rebuilt graph's slots around). A rebuild therefore re-maps: alloc_buffer/free_buffer are not called, CUDA's pool_gen does not move (so its captured graphs survive), and the same topology gets the same slots. MemoryReport carries the reservation's depth (idle_slots, reserved_classes), and CpuBackend::alloc_count is the CPU twin of pool_gen for the gate. GraphCache holds one graph per GraphParams (MRU, MAX_CACHED_GRAPHS), so a switch is try_reuse → alloc_graph with no build and no fusion pass (stats() reports builds vs reuses). Cross-boundary staging is charged to pool_bytes but still allocated exact; backend-internal scratch is not in the report. Metal's weights_bytes charges what MpsState registered (#299) — the mmap-backed NoCopy weight slices and the per-weight copies alike — so the E4 gate subtracts the resident weights exactly as the E5 auto fit already charges them from the GGUF index. The term was the trait default 0 until #299, which let the gate admit the weight bytes more than the pool it protects had room for (on the measured Mac a 7B Q4_K_M's ~4.4 GiB, against a 38 339 MiB working set).

2. Layer offload is one plan read in three places (E5)

  1. Layer offload is one plan read in three places (E5). graph/offload.rs owns OffloadPlan { gpu_layers, n_layers }: blocks 0..gpu_layers run on the device and the rest on the CPU, and tensors outside any block (embedding, final norm, lm_head) follow the device only when every block is offloaded (device_holds_unblocked — llama.cpp's n_gpu_layers > n_layer convention, stated once). The request is --gpu-layers N (CLI) or MINFER_GPU_LAYERS=N (environment; unset = every block a device can hold, the pre-E5 behaviour); OffloadRequest::plan resolves it purely (CI-tested), and a spelling that is not a block count is a refused load, never a guess. The three readers must agree on the same number: (a) the loader registers a tensor on the device only when OffloadPlan::allows_weight(name) says so, taking the block from the registry name (block_of: {ns}blk.{i}.…, including the fused blk.{i}.attn_qkv / blk.{i}.ffn_gu concat copies); (b) the builder stamps CNode.layer (GraphBuilder::set_layer, once per block) and gates the device-only fused forms on layer_gpu = gpu && il < gpu_layers; (c) the assignment pass (BackendScheduler::assign_backends → GraphAllocator::supports_for(op, dtype, layer)) never offers the device for a block past the plan. CParams.gpu_layers carries it into the reuse identity, because the assignment is topology. The load verifies the plan against what was registered and drops to CPU-only with a printed reason when the offloaded blocks' weights are not usable there; offload_report prints where the blocks landed with the measured device bytes. S2 (the automatic fit) adds the auto spelling: the loader measures each block's weight bytes from the GGUF index (GgufTensorInfo::nbytes, before anything is loaded — the filter is the plan) and fits the largest prefix into the weight budget (fit_blocks, pure: a prefix, not a knapsack, because the plan is 0..gpu_layers and a gap would put a CPU block between two device blocks for nothing), holding back a quarter for the KV arenas and the activation pool. The budget is MINFER_GPU_MEM=<MiB> if set, else three quarters of the device's own answer (CUDA's free bytes; Metal's recommendedMaxWorkingSetSize since #53) — the same default E4's feasibility gate uses, and the same resolver (models::device_memory()), so the fit and the gate talk about one number (weight_budget). Both device backends report a number on their platform; only a device with no state at all (the CPU, or a singleton that could not be created) fits nothing without MINFER_GPU_MEM, and a failed device query refuses auto naming the real error instead of reading it as "0 bytes free" (#122).

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0015 — The offload auto fit takes a prefix, not a knapsack
  • ADR-0016 — A failed device-memory query is not a zero budget

Backends — the three execution engines

minfer runs one declarative compute graph (build → assign → fuse → allocate → execute, docs/COMPUTE-GRAPH-DESIGN.md) on three interchangeable backends:

CPUMetalCUDA
Executorsrc/graph/cpu_backend.rssrc/graph/metal_backend.rssrc/graph/cuda_backend.rs
Device layer— (std threads)src/metal/ + src/metal/kernels/ shaderssrc/cuda.rs + src/cuda/kernels/*.cu
PlatformanymacOS (Apple GPU)NVIDIA, opt-in --features cuda
Assign prioritylast (always answers)first on macOSsecond, when built in
Activationsquantized to Q8_0 (Q8_K for K-quant weights)read as f32f32; int8 MMQ for prefill
KV cachef32 regions, or packed Q8_0 (MINFER_CACHE_TYPE=q8_0 — C4, 3.76× smaller)f32, f16 or packed Q8_0 (C4 S2b, #310)f32, f16 or packed Q8_0 (C4 S2b)
Async modelsynchronousone MpsCommandBuffer per splitstream + CUDA Graph capture/replay
Deep diveswalkthrough 10, 11, CPU optimizationswalkthrough 14, Metal optimizationswalkthrough 15, backend plan, campaign

The src/<backend>/kernels/ paths and the CUDA src/cuda/methods/*.rs families are the source layout plan (#261): the CUDA half has landed (src/cuda/kernels/*.cu, src/cuda/methods/, src/quants/, src/vec_ops/, src/kernel/), while the Metal split (src/metal/kernels/*.metal) and #53's DeviceMemory answer are the Mac round. The Metal shaders are still the single src/metal/kernels/.

This page is the overview: what the backend contract is, how nodes land on a backend, and where the three differ. The linked pages carry the per-backend detail.

1. The contract: the Backend trait

Every backend implements one trait (src/graph/backend.rs:21-98). The scheduler knows nothing about Metal or CUDA specifics — it only talks to this surface:

MethodMeaning
supports_op(op, dtype)capability query, asked per node at build time
supports_fused(fused)gates the fusion pass — fused IR nodes are only produced where a kernel exists (fusion.rs:72)
alloc_buffer / free_bufferthe backend's own buffer pool, sized in f32 elements
alloc_freshsame, but bypasses the recycle free list — split-boundary staging needs ids whose physical contents are still referenced later in the same execute (backend.rs:36-45)
execute_node(node, in_bufs, out_buf, kv_pair)run one node; kv_pair carries the layer's persistent (K, V) region ids for KV ops; the output may alias an input (in-place ops)
read_host / write_hosthost access to a pool buffer — direct slices on CPU, staged transfers on GPU
synchronizewait for async work: CPU no-op; Metal submits the pending command buffer; CUDA closes a capture window if one is open
graph_replay (CUDA only)replay a previously captured CUDA Graph for a split (backend.rs:94-97, feature-gated; default no-op)

Two supporting traits/rules ride along:

  • KvProvider::kv_pair(layer) — each layer owns two persistent KV regions (K and V) that live in the backend's pool and survive graph rebuilds (backend.rs:12-19; allocator detail in walkthrough 07).
  • The allocator is the single owner. Backends own pools, but buffers are only created through the allocator's liveness pass; input buffers are never freed, and in-place aliasing is decided at allocation time, not by kernels.

2. How a node gets its backend

Assignment happens once, at build time — never mid-run:

  1. GraphAllocator::supports(op, dtype) walks the registered backends highest priority first: Metal → CUDA → CPU (src/graph/alloc.rs:140-153). The first backend whose supports_op answers true wins that node. The CPU backend supports the full op set, so it always terminates the walk.
  2. GPU participation is gated on weights: a GPU backend only claims ops once all weight tensors are registered on it; the gate fails → the run aborts with the actual values, never a silent CPU fallback (docs/GPU_SAFETY.md).
  3. Whether any GPU took part is recorded in CParams.gpu, which is part of the graph-reuse identity — a run that switches between GPU and CPU gets a different GraphParams and therefore a rebuilt graph (walkthrough 13).

Because assignment is per node, one forward pass can mix backends. The scheduler cuts the node list into splits — maximal runs of the same backend — and at every split boundary it synchronizes the previous backend and copies split inputs across (copy_across, a host round trip through shared memory; src/graph/scheduler.rs:176). Metal encodes one MpsCommandBuffer per split and submits it at synchronize. Mechanically this is walkthrough 08 §3.

The error contract at execution time mirrors the build-time gate: kernel-invariant violations return Err from execute_node — never a silent fallback to another backend (e.g. a KV-store position ≥ n_ctx, an attention head geometry mismatch, a device-limit shortfall). docs/GPU_SAFETY.md is the binding rules page for the GPU backends (bounded submit waits, no early return past a threadgroup_barrier, device limits queried at runtime).

3. What differs between the three

CPU — deterministic, zero-setup, bit-identical

  • Quantized weights straight from the GGUF mmap; activations quantized to Q8_0 (32-value blocks) or Q8_K (306-byte super-blocks) per matmul — the int8×int8 dot kernels are AVX2 / NEON+SDOT with scalar fallbacks (walkthrough 10).
  • Thread parallelism is ownership-based (one matmul row / one attention head per worker), so results are bit-identical for any --threads value — the property the greedy-token verification gates rely on.
  • KV regions are f32; scores are computed in f32.

Metal — zero-copy on unified memory

  • Weights register with newBufferWithBytesNoCopy: the GGUF bytes are the Metal buffer, no copy, on Apple Silicon's unified memory (walkthrough 14).
  • Activations stay f32; a per-op dispatch matrix picks handwritten shaders (3 matmul tiers, 5 attention variants, rms_norm 2 widths) and decode fusions.
  • Optional f16 KV regions halve attention bandwidth (MINFER_CACHE_TYPE=f16).

CPU KV cache types — MINFER_CACHE_TYPE

f32 (default) · f16 (resolves to f32 here: this path has no f16 KV kernel) · q8_0 (C4: packed Q8_0 cells, ceil(n_kv_embd/32 * 34) bytes per cell instead of 4 * n_kv_embd, so the regions are 3.76× smaller; the store quantizes and the attention reads the packed blocks directly — the K score is a Q8_0 × Q8_0 dot against the quantized query, V accumulates out of the cell, and S1's dequantize-into-a-scratch pass is gone). An unknown value is refused on every device, and a backend without a packed-read kernel refuses q8_0 loudly rather than run f32 — the answer is the registry's reads_packed_kv, which is **true for the CPU (C4 S1/S2a), CUDA (C4 S2b) and Metal (#310 enabled Metal's packed read, whose attention window rides #44; the three C4 items left on #87 — the packed fused epilogue, the packed FA prefill and the dp4a packed dot — landed (#144 items 1+3, #186 item 2; the CUDA residual is #212). A physical context shift (kv_rm/kv_shift) works on a packed region: the survivors move verbatim and are re-rope/re-quantized one row at a time. MINFER_NO_FUSED_Q8_KV=1 restores the S1 read path (the A/B of standing rule 3; measured 1.16× at ctx 512 and 1.31× at ctx 2048 in the fused read's favour). The tolerance class against f32 is named, never bitwise: see the C4 record in docs/ARCHITECTURE-EXECUTION-PLAN.md §5.

CUDA — opt-in, campaign-tuned

  • Built only with --features cuda (plain builds never touch nvcc); --features cuda,cuda_static links cudart statically for deployment (docs/BUILD.md).
  • Prefill runs int8 MMQ tensor-core GEMMs; decode runs weight-streaming MMVQ kernels; attention is split-KV with a combine pass — the whole arc is the CUDA optimization campaign (r5–r60, D1–D4-4).
  • Captures decode-shaped splits as CUDA Graphs and replays them (graph_replay; MINFER_NO_CUDA_GRAPH=1 to disable) — the only backend with a capture/replay protocol on the trait.

Fusion capability differs per backend

The fusion pass consults supports_fused before producing fused IR nodes, so the same model builds a different graph per backend: Metal accepts SwiGLU (metal_backend.rs:705-709); CUDA accepts SwiGLU (cuda_backend.rs:1303-1305) and carries its own fused decode kernels from the campaign (see the CUDA docs for the inventory). The decode fusions (Op::FusedQKV, Op::FusedFFN) are env-revertable (MINFER_NO_FUSE_QKV=1 / MINFER_NO_FUSE_FFN=1) and are part of the reuse identity. Fused vs unfused is bit-identical; when comparing, the unfused path must still run the FusionPass.

4. Support matrix and forcing a backend

  • Quant formats per backend (including the Metal prefill GEMM dispatch window and CUDA MMQ notes): docs/SUPPORT-MATRIX.md. Short version: Q4_0/Q4_1/Q5_0/Q5_1/Q8_0/Q4_K/Q5_K/Q6_K everywhere; Q2_K/Q3_K/I-quants nowhere.
  • MINFER_DISABLE_MPS=1 — force the CPU backend on macOS.
  • MINFER_NO_NEON=1 (aarch64) — drop the CPU NEON layer to scalar (A/B lever).
  • MINFER_NO_AVX2=1 (x86) — drop the whole CPU quants AVX2 layer to scalar (A/B lever); MINFER_NO_AVX512=1 drops just the AVX-512/VNNI K-quant dots to AVX2.
  • MINFER_NO_CUDA_GRAPH=1, MINFER_CACHE_TYPE=f32|f16|q8_0 (q8_0 = packed, on CPU + CUDA + Metal since #310, C4), MINFER_NO_FUSED_Q8_KV=1 (keep S1's dequantizing read of a packed cache, for the A/B), MINFER_NO_FUSE_QKV / MINFER_NO_FUSE_FFN — per-backend behavior levers.
  • Which backend to expect: the startup banner and MINFER_TRACE / MINFER_GRAPH_TRACE show per-node assignments (walkthrough 08 §4).

Numerics across backends are not identical by design — CPU quantizes activations, GPUs read f32 (and CUDA prefill quantizes differently still), so CPU-vs-GPU logits differ; every path is verified against its own reference (greedy output equality with llama.cpp where noted in the support matrix).

5. Reading order

  1. walkthrough 08 — the scheduler — splits, copies, execution.
  2. Backend episodes of the walkthrough: 10 / 11 (CPU), 14 (Metal), 15 (CUDA).
  3. docs/GPU_SAFETY.md — the hard rules before touching GPU code.
  4. Per-backend history: docs/CPU_OPTIMIZATIONS.md, docs/METAL_OPTIMIZATIONS.md, docs/CUDA-BACKEND-DESIGN.md + docs/CUDA_OPTIMIZATION.md.

F4 — Backend registry (design record)

Ticket: #57 ("[F4] Backend registry instead of the compile-time enum"). This document is the design-first artifact: it is committed before the implementation and states the registry's shape, the name surface, the ordering rule and the exact refusals. Gaps and the backlog item live in docs/ARCHITECTURE-ROADMAP.md; the ticket record lives in docs/ARCHITECTURE-EXECUTION-PLAN.md.

1. Why the enum was the problem

Backend was enum Backend { CPU, Metal, Cuda } in src/graph/mod.rs, and every consumer matched on it. GraphAllocator alone dispatched twelve operations through match backend { … }; BackendScheduler::execute matched to pick the executing pool and again for the MINFER_TRACE capture path; and the fusion wiring built a Vec<&dyn Backend> by hand and found a node's backend with a position(|b| b.name() == "cuda") lookup that existed only because the match could not express "whichever device is present". Adding a backend therefore meant editing the allocator, the scheduler, the fusion wiring, the exporters, the KV-session tag table and the op matrix; and none of it told a user anything — there was no way to ask for a backend by name or to learn that a requested one did not exist.

That argument is a decision, and it is frozen in ADR-0001 (one device seam, build-time assignment) and ADR-0011 (the id space and the name surface). This page keeps the shape that replaced it — the handle, the registered set, the ordering rule and the refusals.

2. The handle: Backend(u16)

Backend stays a cheap, Copy, Hash, Eq, Ord value; it stops being an enum. It is an opaque index into a fixed id space:

#![allow(unused)]
fn main() {
pub struct Backend(u16);

impl Backend {
    pub const CPU:   Backend = Backend(0);
    pub const METAL: Backend = Backend(1);
    pub const CUDA:  Backend = Backend(2);
}
}

The id space is compile-time and configuration-independent: cpu = 0, metal = 1, cuda = 2 on every build, whether or not the backend is compiled in. That is deliberate — the id is already on disk and in exports:

  • the KV-session backend tag (graph/kvsession.rs, tag_of / backend_of_tag) is a u32 in a versioned file, so the numbering is a file-format contract;
  • graph/json.rs / graph/dot.rs name and index the backend in exported graph documents;
  • Ord is derived from the id, and the pre-F4 derived Ord on the enum was declaration order — CPU < Metal < Cuda.

Backend::index(), Backend::from_index() and Backend::name() are the only ways to leave and re-enter the id space; kvsession's tag functions and the fusion index are now one-line calls to them instead of three-arm matches.

Debug is implemented by hand so diagnostics print exactly what the enum printed (CPU, Metal, Cuda) — no log line or error message changes shape.

2.1 Two orders, stated once

The registry has two orders and they are not the same; conflating them is the failure this document exists to prevent.

orderwhat it fixesauthority
identity (Backend::index())the on-disk KV-session tag, the exported graph's backend index, Ord, the fusion pass's backend listBackend::CPU/METAL/CUDA
priority (assignment preference)which backend supports_for offers firstBackendEntry::priority, descending

Before F4 the priority order was implicit in the statement order of GraphAllocator::supports_for: Metal, then CUDA, then CPU. It is preserved exactly, and it is now a number on each entry — Metal 300, CUDA 200, CPU 100 — so the ordering gate can pin it and a future accidental reordering fails a test instead of silently changing where a graph's nodes land.

Determinism rule. The assignment pass fixes the graph's topology, which is part of GraphParams' reuse identity (standing rule 3). Nothing in the registry may depend on HashMap iteration order:

  • the registry table is a fixed-size array indexed by the handle id — not a HashMap;
  • Registry::by_priority() sorts by (Reverse(priority), index), so two entries can never tie and the result is a total order derived from the pinned numbers;
  • BackendFilter is a [bool; N_BACKENDS] indexed by id, never a set that is iterated;
  • the fusion-pass backend list is built by walking the identity order, so the index it hands the pass is Backend::index() on every run.

3. The registry table

One BackendEntry per backend, built once at startup (registry(), a OnceLock, is forced by main before anything else runs):

#![allow(unused)]
fn main() {
pub struct BackendCaps {
    pub reads_packed_kv: bool,          // the #87 seam, §8 — the only field
}

pub struct BackendEntry {
    pub handle: Backend,
    pub name: &'static str,             // "cpu" | "metal" | "cuda"
    pub priority: u16,
    pub caps: BackendCaps,
    pub pool:     fn(&GraphAllocator) -> Option<&dyn BackendTrait>,
    pub pool_mut: fn(&mut GraphAllocator) -> Option<&mut dyn BackendTrait>,
    pub host_read: fn(&GraphAllocator, usize) -> Option<Vec<f32>>,
    // F5 (#58): the split boundary's two phases (§11).
    pub copy_cross:  fn(&mut GraphAllocator, u64, NodeId, Backend) -> Result<bool, String>,
    pub await_cross: fn(&mut GraphAllocator, u64, NodeId, Backend) -> Result<(), String>,
    pub kv_format: fn(&GraphAllocator) -> KvFormat,
    pub enable:    fn(&mut GraphAllocator) -> bool,
    pub unavailable: fn() -> Option<&'static str>,
}
}

Each backend module owns its entry and its hooks (cpu_backend::entry(), metal_backend::entry(), cuda_backend::entry()), and Registry::build() calls their register() — that call site is the only place a backend is introduced.

The capability authority is the module-level item, stated once (#244, 2026-10-01). Each backend defines its op×dtype matrix and its fusion matrix as module-level free functions (cpu_backend::supports_op, cuda_backend::supports_fused, …) and its attn_span answer as a module-level constant (cpu_backend::SUPPORTS_ATTN_SPAN, …). The impl Backend for X methods are one-line forwards to them, and the assignment pass reads the trait method (graph::backend_takes → dyn Backend::supports_op). The registry carries no copy of those three answers. The supports_op / supports_fused / supports_attn_span fields that used to sit here were written by every entry() and read only by registry::tests — a mirror with no production reader, which is why #244 deleted them (option (b) of that ticket's escalation). The gate that used to compare the field against the trait now compares the module-level function against the trait (registry::tests::registry_caps_match_the_backend_trait), i.e. the authority the field merely mirrored; putting the two on one line is no longer possible even in principle.

BackendCaps therefore carries exactly one field, reads_packed_kv, and that one is read in production because the question is asked about a format, not an engine: GraphAllocator::ensure_kv's packed-region refusal and KvFormat::supports call registry::reads_packed_kv(backend), where no &self is available (§8). It is the one capability the registry itself carries.

pool / pool_mut are why "sync" and "copy" are not a match any more: the allocator's dispatch helpers ask the entry for the pool and then call the Backend trait method (synchronize, copy_cells, write_host, …) on it. The two backends that have a native host-read path different from the trait's borrowed read_host (CUDA's stream-ordered copy_to_host) say so in host_read, which is why copy_to_cpu and the CUDA KV debug read all collapse to one call without a match.

enable is the lazy "a session names this backend, so bring its pool up" hook; unavailable is the runtime probe (device present, not disabled) used by the availability refusal in §5.

copy_cross / await_cross are F5's addition (§11): the split boundary's cross-backend staging copy, split into enqueue and wait so a backend with device memory can make the transfer asynchronous and name exactly where the consumer waits. They follow the same rule as the rest of the entry — the allocator never asks "is this the CPU?", it asks the entry.

4. What no longer matches on Backend

sitebeforeafter
alloc.rs pool ops (12)match backend { Backend::CPU => …, Metal => …, Cuda => … }(entry.pool_mut)(self) → trait method
alloc.rs supports_forhardcoded Metal-then-CUDA-then-CPU if let chainregistry().by_priority() + entry.caps
alloc.rs kv_load enablethree-arm match with per-cfg fallbacksentry.enable + unavailable
alloc.rs kv_element_formatmatch with per-cfg fallbacksentry.kv_format (takes the allocator: the answer is the engine's resolved format, per #99/#153 — the CPU field and the CUDA backend's kv_layout — not a process global)
alloc.rs copy_cells_in_poolmatch + per-cfg stringsentry.pool_mut + registry-aware refusal
scheduler.rs executematch split.backend { … }alloc.pool_mut(split.backend)
scheduler.rs read_host_buffermatchentry.host_read
json.rs / qwen2,3/graph.rs fusion wiringhand-built Vec + index matchalloc.fusion_backends() + Backend::index()
json.rs / dot.rs namingmatch on the enumBackend::name() / entry
kvsession.rs tagsmatch 0/1/2Backend::index() / from_index()
op_matrix.rs (test matrix)match tag { … }equality on the handle + registry caps
ensure_kv packed check, KvFormat::supportsbackend != Backend::CPU / matches!(device, Device::Cpu)entry.caps.reads_packed_kv

Everything the enum used to decide is now either a registry field or a trait call behind the entry's pool hook.

5. The name surface

Accepted names: cpu, metal, cuda. Comparison is case-insensitive and whitespace is trimmed ( Metal is metal); the canonical spelling is lower-case and is what every message uses.

Two spellings request a set of backends:

  • --backend <name> on the CLI (repeatable, and comma-separated values are accepted: --backend cpu,metal). It is extracted from argv before any subcommand dispatch, so serve, viz, run and bench all honour it.
  • MINFER_BACKENDS=<csv> in the environment, for callers that cannot pass a flag. The flag wins when both are present.

The request is a fence: it removes backends from participation. It never adds one, and it is not an offload policy (--gpu-layers / MINFER_GPU_LAYERS stay the authority for how many blocks the device holds). cpu is always allowed, whether or not it is named: it is the universal fallback and a graph must always be assignable, so --backend cuda means "the device if it can take the node, the CPU otherwise", and --backend cpu is the useful spelling — force the CPU exactly as MINFER_DISABLE_MPS=1 does, but for every device.

Unset (the default) means "every backend this build has and this machine can use" — the pre-F4 behaviour, byte for byte. The pre-existing MINFER_DISABLE_MPS and MINFER_DISABLE_CUDA fences keep working unchanged, both presence-checked: they are read by the device layer (MpsState, CudaState) and a fenced backend is simply never available, whatever the name surface says.

Where the fence is enforced:

  1. Device participation — Qwen2Graph::device / Qwen3Graph::device (the single authority CParams.gpu and the server's batching default read) return Cpu for a fenced device, so the builder never emits a device-only fused node (FusedQKV, QkvBiasNorm, FusedFFN, QkvBiasRopeStore) that the CPU cannot execute.
  2. Assignment — GraphAllocator::supports_for skips a fenced backend, so a node is never placed on a pool whose weights were never registered there.

Both read one process-wide filter installed at startup (registry::active_filter()), which is what makes them unable to disagree.

6. Refusals — three classes, three messages, always loud

A name is resolved in two stages, because the compile-time answer and the runtime answer are different questions and a backend that is compiled out must not be silently treated as absent.

Stage 1 — names, before anything else in main. Purely a function of the name and the compile-time registry (no device is touched, so it is covered by CI). Two failures:

Error: unknown backend 'gpu2'; known backends are: cpu, metal, cuda
Error: backend 'cuda' is known but not compiled into this build: the CUDA
       backend is compiled only with --features cuda

The first is "no such backend name" — always an error, on every build. The second is "the name exists, this binary does not contain it", and it names which of the two situations it is. A metal name on Linux and a cuda name on a default build both land here, with the reason spelled out (Metal is compiled only on macOS (target_os = "macos")).

Stage 2 — availability, once the device layer is up. Still startup (before the model is loaded), but now the device layer can be asked. A compiled-in backend that this machine cannot use is a third, different message:

Error: backend 'cuda' is compiled in but not available on this machine: no CUDA
       device is available, or CUDA is disabled (MINFER_DISABLE_CUDA)

A cuda name on a --features cuda build with MINFER_DISABLE_CUDA=1, or with no device, lands here — not in the "not compiled" bucket, and never in a silent fallback to the CPU.

Only a named backend is checked at stage 2. The filter therefore carries two facts per backend — allowed (may this run use it?) and requested (did the request name it?) — because the default request means "whatever this build can use". Preserving that distinction is what keeps the pre-existing MINFER_DISABLE_CUDA / MINFER_DISABLE_MPS flags meaning "run on the CPU" rather than turning them into "refuse to start": with no --backend / MINFER_BACKENDS, nothing was named, so stage 2 has nothing to check and the run proceeds on the CPU exactly as it did before F4. Naming the fenced backend (--backend cuda with MINFER_DISABLE_CUDA=1) is a refusal, because the user asked for it.

Stage 2 runs before the model path is resolved on every path that will run a model (run/serve/viz in main, and bench/specverify in their own run), so an unrelated failure — a missing file, a bad GGUF — cannot preempt it.

Both stages exit non-zero and print nothing else about backends. The reverse direction is also loud: no code path may drop a named backend and keep going.

7. Feature gates

configurationmetal entrycuda entrymetal namecuda name
default Linuxabsentabsentnot compilednot compiled
--features cudaabsentpresentnot compiledresolves
macOSpresentabsentresolvesnot compiled
macOS + --features cudapresentpresentresolvesresolves

The registry's registered set is the compile-time one: cuda_backend::register is #[cfg(feature = "cuda")], metal_backend::register is #[cfg(target_os = "macos")]. The names are not gated: Backend::name() and the known-name list are unconditional, so an unregistered name still resolves to a handle and still produces the accurate "not compiled into this build" refusal instead of "unknown backend". The ordering gate pins the registered set and the priority order per configuration (all four rows above), so a change to either is visible in CI on every configuration.

8. The seam for a later per-format capability query (#87)

#87 wants the registry to answer "can this backend read a packed q8_0 KV region?" instead of a hardcoded CPU-only test. The registry carries exactly that field — BackendCaps::reads_packed_kv — and it is used by this ticket's own code, not reserved for the future: GraphAllocator::ensure_kv's packed-region refusal and KvFormat::supports both read it, replacing backend != Backend::CPU and matches!(device, Device::Cpu) respectively. That is why it is a field and not dead abstraction: there is one authority for the answer, and #87 is the work that flips CUDA's and Metal's value to true (and adds their kernels). No other per-format query is added here.

9. What the registry is not

  • Not a plugin/dlopen system: the set of backends is fixed at compile time. The registry makes the set data instead of control flow; it does not make it extensible at runtime.
  • Not a device-selection policy: --gpu, --gpu-layers and MINFER_GPU_LAYERS keep their meanings.
  • Not a weight registry: GraphAllocator::register_weight (CPU) and CudaState::register_weight (CUDA) are unchanged; the "all weights registered" gate is per architecture and stays where it is.
  • Not a place to move Device: models::Device stays the coarse "the device participates" fact that the server's batching default reads, and gains a Device::backend() mapping so the two id spaces have one bridge.

10. Acceptance and the gates

  • Behaviour preservation. The identity and priority orders, the supports_op / supports_fused / supports_attn_span answers, the sync/copy arms and the memory accounting are unchanged; the existing suites are the measurement, and the order/name gates pin the parts a suite would not notice.
  • Gates
    1. registry::tests::names_resolve_and_unknown_names_are_refused — the known names resolve to the pinned handles, an unknown name and a compiled-out name are distinct loud errors, and diagnostics keep the pre-F4 spelling (pure; CI covers it).
    2. registry::tests::the_registered_set_and_priority_order_are_pinned — the registered set, the priority order and the priority numbers, per configuration, plus Ord and a fresh allocator's deterministic answer.
    3. registry::tests::the_name_surface_fences_devices_and_keeps_cpu — the fence surface (comma/repeat spellings, cpu always admitted, the flag winning over the environment, both refusals).
    4. registry::tests::the_packed_kv_capability_is_the_registrys_answer and registry::tests::registry_caps_match_the_backend_trait — the #87 seam is the field both C4 gates read; and the capability answer is one authority: each Backend trait method forwards to its backend module's own free function / constant, and the assignment pass reads the trait. #244 deleted the three mirrored BackendCaps fields, so the registry holds no second copy to disagree with (§3).
    5. alloc::tests::a_fresh_allocator_inherits_the_runs_backend_filter — the fence reaches the assignment pass through the same active filter.
    6. tests/backend_registry_cli.rs — the process level: an unknown name exits non-zero naming the accepted set, a compiled-out name gives the other message, a known name passes the gate (the failure moves on to the model), bench honours the flag, --help documents it, and an unnamed device disabled by the pre-existing flags is still "run on the CPU".
    7. alloc::tests::cuda_the_backend_fence_moves_assignment_off_a_usable_device — #[ignore]d (needs a device): the fence moves assignment off an enabled, usable device without tearing it down.
  • Mutation evidence. (a) making an unknown name fall back to the default backend fails gate 1; (b) swapping two priorities fails gate 2; (b2) perturbing one priority value without reordering also fails gate 2, which is why the expected numbers are literals. All three are reverted; the observed failure output is recorded in the ticket record.

11. F5 — the async staging copy and its synchronization points (#58)

Ticket: #58 ("[F5] Async cross-backend copies and events"). The implementation record (measurements, mutation evidence, honest scope) lives in docs/ARCHITECTURE-EXECUTION-PLAN.md §F5; the scheduler-side contract is also summarized in docs/COMPUTE-GRAPH-DESIGN.md §3.4. This section is the registry contract those records refer to.

11.1 What the hot path is, exactly

BackendScheduler::execute partitions a graph into contiguous same-backend splits (assign_backends → split_graph). When a value produced by one split is consumed by a split on another backend, the boundary must move it; the scheduler does that through GraphAllocator::copy_across once per entry of Split::inputs. Those copies are the hot path this ticket is about — they run on every forward, on the critical path of every decode step of a partially offloaded model.

Everything else that reads device memory back to the host is legitimately host-visible and out of scope, and is enumerated here so "zero blocking copies" is not read as "zero device→host copies anywhere":

sitewhy it blocks on purpose
GraphAllocator::copy_to_cpu on the logits / output paththe run's answer; the caller is about to read it
GraphAllocator::copy_kv_to_cpu (KV session save, --session)a file write; no overlap to exploit
Copy / debug dumps, MINFER_GRAPH_DUMP, doc dumpsdiagnostics
MINFER_TRACE / viz capture (CudaBackend::copy_to_host fallback, CaptureStaging)already batched asynchronously; the per-node fallback is for tensors above the staging ceiling
weight / tokenizer loading, fill_input, write_hosthost→device or host-only; no device leg to wait on

Before F5 a single CUDA→host boundary input cost two host stalls: a full cudaStreamSynchronize plus a blocking cudaMemcpy D2H inside CudaBackend::copy_to_host. The counters below are how that became a number (graph::copystats, per allocator, read by the gates).

11.2 The two phases

copy_across (phase A) resolves/allocates the destination staging buffer, marks the entry pending, increments copies, and calls the source backend's copy_cross. A second request for the same (graph, node, destination) while that entry is still pending is the same transfer — one staging buffer, one unchanged source node — so it is a no-op and is not counted again (#138). Under the F5 boundary that state was unreachable (the first copy was always awaited before the second was enqueued); the deferral below is what makes it reachable, and re-issuing it would duplicate the transfer and leave the first copy's pinned slab and event behind, because the allocator clears one pending key per entry.

await_cross (phase B) calls the source backend's await_cross, clears the pending flag and increments waits. One phase-B call per phase-A call is the contract, and GraphAllocator::cross_input — the checked accessor — refuses a still-pending entry with a loud Err naming the missing wait, so dropping a wait can never be bought with a silent read of in-flight data. The scheduler reads a consumer input through GraphAllocator::cross_input_ready, which issues a pending entry's wait at that first use and then calls cross_input; whatever nothing downstream reads is drained once, after the last split (GraphAllocator::drain_cross_pending). The wait is still issued exactly once per copy, at the latest safe point rather than at the boundary (#138).

Phase A returning Ok(true) means "this backend issued the transfer"; Ok(false) means "declined — use the synchronous host round trip". Declining is not a silent CPU fallback: the allocator does perform the pair, just synchronously, and the boundary counters record it as a blocking copy.

Mechanism choice on Metal (#137). Three primitives were candidates. A command-buffer completion handler can only notify the host — it is not waitable, so it cannot express phase B at all. A plain MTLEvent is device-scoped and exposes no host wait. MTLSharedEvent is the one primitive that both MTLCommandBuffer::encodeSignalEvent / encodeWaitForEvent accept and that exposes a bounded host wait (waitUntilSignaledValue:timeoutMS:); it is therefore what the port uses, and it is also the reserved device-side mechanism. The two waits are enumerated in §11.3. The blit is encoded into the producer split's own command buffer rather than a dedicated one; §11.5 records why that is a correctness requirement, not tidiness.

backendphase A (copy_cross)phase B (await_cross)
cpusynchronous host round trip — the CPU has no device memory, so there is no transfer to make asynchronous. The device leg of CPU→device is the destination pool's own stream-ordered write_host (pinned + cudaMemcpyAsync, 7e⑥), which never blocked the host either.documented no-op — nothing was enqueued that needs waiting for; the destination device orders its own fill on its stream. Still counted, so the one-wait-per-copy contract is backend-independent.
cudadevice→host: cudaMemcpyAsync D2H into a pinned slab + cudaEventRecord, both stream-ordered after the producing kernels. Any other destination declines: a device→device staging copy (unreachable — copy_across early-returns when the source and destination backends match), CUDA→Metal on a macOS+CUDA build.device→host: cudaEventSynchronize — the host block, and the only one the async path takes for that copy — then the bytes are published into the staging buffer. A device consumer uses cudaStreamWaitEvent (no host block).
metaldevice→host: a MTLBlitCommandEncoder copy of the source StorageModeShared MTLBuffer into a fresh shared staging buffer, plus MTLCommandBuffer::encodeSignalEvent on a fresh MTLSharedEvent. Both are encoded into the producer split's own command buffer (one command buffer per split — see §11.5). Any other destination declines: a Metal→Metal staging copy is unreachable (copy_across early-returns when the source and destination backends match) and CUDA does not run on Apple Silicon, so no device→device pair exists on macOS.device→host: MTLSharedEvent::waitUntilSignaledValue:timeoutMS: with a 10 s bound — the host block, and the only one the async path takes for that copy — then the bytes are published into the staging buffer. A timeout is a loud Err naming the value waited for, the observed signaledValue and the command buffer's real status() / error(); never an unbounded block (MetalBackend::cross_take). A device consumer would use MTLCommandBuffer::encodeWaitForEvent (no host block); it is reserved, because the device→device pair is unreachable on macOS.

11.3 The synchronization points, enumerated

Every execution, in order:

  1. The boundary retire (BackendScheduler::execute step 1) — GraphAllocator::retire_backend(previous). Not a copy: it retires the previous backend's kernels (and, for CUDA, closes an open graph-capture window and clears the MMQ memoization). Since #138 it calls Backend::retire, whose default body is synchronize — a backend whose boundary work is a submission (Metal) still blocks here — while CUDA overrides it: its close is already stream-ordered with the copies of step 2 (both run on the backend's own stream), so the cudaStreamSynchronize that synchronize adds orders nothing new and is not taken. sync_backend (the blocking form) remains the after-the-last-split flush.
  2. Phase A, per staged input — copy_across. A CUDA→host copy enqueues cudaMemcpyAsync + cudaEventRecord; a CUDA→device copy would use cudaMemcpyAsync D2D; the CPU does its host memcpy; Metal encodes a MTLBlitCommandEncoder copy plus encodeSignalEvent into the producer split's own command buffer (it runs before that split's retire, so the copy is ordered with the kernels that wrote its source — §11.5). No host wait here, and all of the boundary's inputs are enqueued before any of them is waited on.
  3. Phase B, per staged input — await_cross, issued at the consumer's first read of that staging buffer (or, for an entry nothing reads, by the end-of-execution drain):
    • device→host, CUDA: cudaEventSynchronize on the event recorded in step 2. Invariant that makes it necessary: the D2H destination is host memory, and the host is about to read it; without the wait the consumer reads bytes the DMA may not have written yet. Deferring it to the read is what lets the other copies of the same boundary stay in flight.
    • device→host, Metal: MTLSharedEvent::waitUntilSignaledValue:timeoutMS: with a 10 s bound, then a host read of the shared staging buffer. Invariant: the event is signaled at the end of the producer split's command buffer, and the host's StorageModeShared read is only ordered after that signal; the wait is deferred to the consumer's read exactly as CUDA's is. A timeout is a loud Err naming the real status. The bound and the loudness are the GPU-safety contract (docs/GPU_SAFETY.md §2.4).
    • host→device: no host wait; the fill is ordered on the consuming backend's own stream ahead of the kernels that read it. Invariant: stream order — the copy and the first consumer share one stream, so no host synchronization is needed and none is taken. This is the device-consumer direction the ticket's first bullet names, and it is already host-free by construction.
    • device→device (no backend implements it today): cudaStreamWaitEvent on the consuming stream, or MTLCommandBuffer::encodeWaitForEvent with the signalling MTLSharedEvent, would be the mechanism; the host never blocks. It stays reserved: copy_across early-returns when source and destination match, and no other device destination is reachable on a macOS build (CUDA does not run on Apple Silicon), so there is no pair to exercise. See the census note in the execution plan.
    • The contract invariant, independent of backend: one phase-B wait per phase-A copy, enforced by cross_input's refusal.
  4. The next boundary's retire, or the final flush after the last split — sync_backend, which for CUDA is also where an open capture window closes.

The gated evidence for 2–3 is copystats::CrossCopyStats (per allocator: copies, waits, deferred_waits, blocking_host_copies, async_host_copies, event_syncs, stream_waits), CudaBackend::blocking_readback_count() (a device-level count of blocking cudaMemcpy D2H calls), CudaBackend::stream_sync_count() (host stalls; per backend since #185, the process-wide counter deleted in #242), CudaBackend::cross_inflight_peak() (how many staging copies were enqueued but not yet waited on at once — the device-side overlap metric of #138) and, on Metal, MetalBackend::sync_readback_count() (host reads of a StorageModeShared pool buffer — there is no blocking-copy API to count, so this is the device-level analogue; the async path never moves it). MINFER_SYNC_COPIES=1 restores the pre-F5 synchronous path as the bitwise reference; copystats::set_sync_for_test is its programmatic form.

11.4 Overlap: what F5 delivers, and what the deferred wait adds

F5 delivered the async substrate + documented waits: no blocking copy on the boundary path, one explicit wait per staged input, and one measured reduction in host stalls (the per-copy stream synchronizations inside copy_to_host are gone).

#138 makes the wait late. The boundary only enqueues; each staged entry's single wait is issued at the consumer's first read of it, and the end-of-execution drain covers an entry nothing reads. Two things follow, both measured on the 0.5B mixed-offload gate:

  • the boundary's own cudaStreamSynchronize is gone (it retired the producer before copies that are already stream-ordered behind it — one full stream sync per CUDA→CPU boundary, 21 → 0 over the gate's 7 forwards), and
  • the boundary's copies stay in flight while the consumer works, so a boundary with several staged inputs holds several transfers at once (2 where the F5 enqueue-then-wait order cannot exceed 1).

Resolved by #300: there is no cross-split overlap to gain on the reachable macOS topology, and §11.6 records the measured negative result — the boundary blit stays on the producer's command buffer, which is both correct (it reads the source in the submission that wrote it) and not slower than the explicit-dependency alternative. The split loop remains strictly sequential, and a host-side consumer must wait by definition. What the deferral buys is that the wait happens where the data is needed, not where it was produced.

11.5 The Metal port (#137): why the blit shares the split's command buffer

MetalBackend keeps one MpsCommandBuffer per split (compute-graph rule 8); the producer submits and bounded-waits it in Backend::retire. The first cut of this port gave each staging copy its own command buffer, committed from copy_cross after the retire. That is a second submission overlapping the next split's, and on the measurement box it changed the kernels' results: the same mixed 0.5B offload graph run twice with the copy mode toggled differed by max |Δlogit| ≈ 1.4, while each mode compared with itself was bitwise stable. The staged bytes were identical (the source was read at enqueue and compared with the staging buffer at the wait), so the divergence was the extra in-flight command buffer perturbing Metal kernel execution — a pre-existing fragility of the backend, not the copy.

Encoding the blit into the producer split's own command buffer fixes it and restores the one-command-buffer invariant: copy_across for a Metal→CPU boundary now runs before BackendScheduler::execute calls retire_backend, so MetalBackend::cross_enqueue appends an MTLBlitCommandEncoder pass (and encodeSignalEvent) to the still-open split buffer, and the split's own submission carries the copy behind the kernels that wrote its source. A source with no open split buffer gets a standalone buffer that cross_enqueue submits itself; the synchronous reference (MINFER_SYNC_COPIES=1) keeps the post-retire order, because its host read must see the producer's finished output.

Measured on macbook (macOS 27.0.1, Apple M4 Pro) — the F5 S3 record in docs/ARCHITECTURE-EXECUTION-PLAN.md.

11.6 The #300 result: the serialization is a missing dependency, and there is no overlap to gain

Ticket: #300 ("true cross-split overlap after #137"). It asked for either a measured overlap or a recorded negative result beside the §11.5 note. The result is a negative one, with the mechanism measured on macbook (macOS 27.0.1, Apple M4 Pro) at ad707c7 (2026-10-09), recorded in 9df405d.

The #137 divergence is a missing-dependency race, not a kernel perturbation. §11.5 inferred from the then-red baseline that "the extra in-flight command buffer perturbed Metal kernel execution". Isolated, the mechanism is simpler and deterministic. The pathological first cut is a standalone boundary command buffer submitted from copy_cross before retire submits the producer's split buffer, with no dependency on that buffer. On the shared MpsState command queue the standalone buffer is committed first, so the blit reads the producer's StorageModeShared source window before the producer's kernels wrote it, and the consumer's wait on an already-signaled event returns the previous forward's bytes. On models::qwen2::graph::tests::offload_copy::async_cross_copies_never_block_and_stay_bitwise_identical_on_metal that cut reads max |Δlogit| = 26.718678 at step 0, the same value on two runs (the §11.5 "staged bytes identical" check was stale-to-stale, which is why it looked consistent).

Option A — a separate boundary buffer with an explicit dependency — is correct. Committing the boundary blit on its own MTLCommandQueue, with encodeWaitForEvent on an event the producer signals at its split end before the blit and its own event for the consumer, restores bitwise identity: the real-model gate reads max |Δlogit| = 0 over the prefill + 6-decode loop (twice) with the counters unchanged — copies=35 waits=35 deferred_waits=35 blocking_host_copies=0 async_host_copies=21 event_syncs=21 sync_readbacks=0 async against copies=35 waits=35 blocking_host_copies=21 sync_readbacks=21 sync — and the cheap 3-node gate passes (copies=2 waits=2 deferred=2 blocking=0 async_host=1 event_syncs=1 sync_readbacks=0). So the perturbation is avoidable: it is a dependency the shared queue did not supply, not a fragility of Metal kernel execution.

But Option A buys no measurable cross-split overlap, so the production design stays. Cross-split overlap needs a second split whose work can run beside the copy. On the reachable macOS topology there is none: the E5 mixed graph (4 of 24 blocks on the device) is a single CPU → Metal → CPU sequence, so there is exactly one device→host boundary per forward; that boundary's consumer is the host CPU split, whose first node directly reads the dominant staged tensor (the hidden state), so the copy is on the consumer's critical path by definition; and the boundary's other two staged tensors are read within the CPU split's first eight nodes (add at node 46, cells at 52, attn_span at 54 in the decode graph), so a separate buffer could only overlap a blit of two tiny index/window tensors with a handful of host-side node setups. A general non-blocking producer retire cannot be made safe without reordering the next device submission after the boundary blit, which reintroduces exactly the serialization it would remove. A device→device pair (Metal→Metal) would be the one topology with real overlap, and copy_across early-returns on it while no CUDA device exists on macOS — it stays reserved (§11.3). The boundary blit therefore keeps riding the producer's command buffer; it is bitwise, and on this topology it is not slower than Option A.

Reproduction. Both cuts are MINFER_300_*-gated temporary patches to MetalBackend::cross_enqueue (a standalone cmd_buffer() + submit() before retire for the pathological cut; a second queue + encodeWaitForEvent / encodeSignalEvent for Option A) — neither is in the tree. The acceptance command is the same in every arm:

cargo test --release --bin minfer async_cross_copies_never_block_and_stay_bitwise_identical_on_metal -- --ignored --test-threads=1 --nocapture

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0001 — Inference runs through one declarative compute graph
  • ADR-0011 — Backend ids are a file-format contract: appended, never renumbered
  • ADR-0006 — The KV storage format is a per-engine gate, not a process-wide global
  • ADR-0008 — GPU safety: bounded waits, no early return past a barrier, runtime device limits
  • ADR-0009 — A failure is an error, never a silent fallback

GPU Safety (Metal + CUDA Backends)

Scope: preventing GPU faults/hangs from freezing the machine, and the review discipline that produced these guards. The Metal sections are the original incident-driven rules; the CUDA section at the end translates them for the graph CUDA backend. See METAL_OPTIMIZATIONS.md for performance work and AGENTS.md for the project overview.


1. Incident: M4 Pro GPU hang (2026-08-02)

A full diagnostic report lives at ~/macbook-gpu-hang-report-2026-08-02.md. Summary:

  • The GPU (AGXG16X = M4 Pro) hardware hung at ~15:37:25. WindowServer's compositing thread froze in mtl_submit → IOGPU → AGXG16X (40 s zero progress), then every Metal client (incl. new minfer instances stuck in MTLCreateSystemDefaultDevice) blocked behind WindowServer. The machine required a forced shutdown.
  • minfer was the only active GPU workload at the time (llama-cli / Ollama were idle). The exact faulting kernel was not identified (snapshots only show the post-freeze state).
  • No deterministic OOB/deadlock was found for Qwen2-0.5B, but the review found three structural amplifiers and several latent landmines for other models.

2. Safety fixes applied (2026-08-02)

2.1 submit() hardening — no more infinite block

MpsCommandBuffer::submit() previously waited DISPATCH_TIME_FOREVER on a semaphore and never checked MTLCommandBufferStatus. A single GPU fault would block minfer forever (and, since Metal clients share the GPU, could stall WindowServer → whole-machine freeze).

Now (src/metal/):

  • Bounded wait: dispatch_semaphore_wait(sem, dispatch_time(NOW, 10s)).
  • On completion, checks MTLCommandBufferStatus — non-Completed reports an error instead of silently continuing.
  • On timeout, reports "GPU hang".
  • MINFER_TRACE=1 records the last 16 dispatch op labels (rms_norm/matmul/gqa_attn/store_kv/swiglu/add/bias/rope/ embed); an error/timeout prints the trace so the faulting kernel family can be identified. Trace recording is env-gated (zero overhead when off).
  • submit() now returns Result<(), String>; all callers print + exit (or fall back to CPU for embed_tokens_gpu).

2.2 gqa_attn barrier deadlock — never return before a barrier

kernel_gqa_attn_f32/f16 used if (h >= nh) return; before the threadgroup_barrier. When nh % nk != 0, some simdgroups exit early while others wait on the barrier → GPU permanent deadlock = machine freeze.

Fix (src/metal/kernels/, both kernels): no early return. Invalid heads (h0 >= nh) run the full loop with a dummy head index (h = 0, keeps pointers in-bounds) so all simdgroups reach every barrier, then skip the output write via a valid_head flag.

2.3 Runtime guards — fall back instead of risking a fault

The legacy whole-layer layer_gpu / output_norm_gpu entry points (deleted with the imperative forward, Phase 6) returned false (CPU fallback) when the kernels' assumptions did not hold; the graph path reports the same invariants as Err from MetalBackend::execute_node (never a silent CPU fallback). The checked assumptions:

  • nh % nk != 0 — attention barrier participation (see 2.2).
  • hd > 256 — the float acc[256] private array would overflow. Note: the threadgroup memory limit is device-specific and must be queried, not assumed — see §4. This hardcoded 256 is the array size, which is a fixed kernel declaration; the threadgroup-smem check belongs in the dispatch (§4).
  • ne/nqt/nkt/nf % 32 != 0 — quantized-matmul block alignment.
  • ne % 32 != 0 in the output matmul.

2.4 The cross-backend staging copy (F5/#137) — bounded wait, loud status

The split boundary's Metal→host staging copy is a MTLBlitCommandEncoder copy plus a MTLSharedEvent signal, encoded into the producer split's own command buffer (docs/BACKEND-REGISTRY-DESIGN.md §11.2/§11.5). The consumer's wait:

  • is bounded — MTLSharedEvent::waitUntilSignaledValue:timeoutMS: with a 10 s bound (MetalBackend::cross_take), the same order as submit()'s. A GPU that never signals the event is an Err, never an unbounded host block.
  • is loud — a timeout reports the value waited for, the observed signaledValue, and the command buffer's real status() / error(). There is no silent CPU fallback: BackendScheduler::execute propagates the Err and the run stops.
  • happens once per copy, at the consumer's first read — the blit is encoded (and the split submitted) without a per-copy host wait; the one wait is the documented synchronization point.
  • stays inside the split's command buffer — the blit is encoded into the producer's open buffer before retire submits it. A separate boundary command buffer overlapped the producer and changed the kernels' results (measured on macbook (macOS 27.0.1, Apple M4 Pro)); the F5 S3 record in docs/ARCHITECTURE-EXECUTION-PLAN.md carries it.

3. Audit findings (2026-08-02) — status

Review of all 30 Metal kernels for the same failure classes (barrier deadlock, OOB, fixed arrays, dimension assumptions, infinite waits).

IDFindingRiskStatus
H1Attention assumes hd == hd_kv — kernel uses stride_kv = nk*hd but the KV cache row is nkt = nk*hd_kv; OOB reads if they differ (other Qwen2.5 models may have hd_kv != hd).High (fault)GUARDED 2026-08-03 — layer_gpu aborts (nkt != nk*hd → gpu_abort) instead of risking misaligned KV reads. Qwen2.5 0.5B/1.5B have hd == hd_kv, so this never fires on supported models
H2Attention threadgroup smem 2*32*hd*4 may exceed the device limit for large hd.High (dispatch failure)GUARDED 2026-08-02 — layer_gpu queries device.max_threadgroup_memory_length() at init (cached) and gpu_aborts when 2*32*hd*4 exceeds it (see §4)
M1Q4_K/Q5_K/Q6_K kernels assume K % 256 == 0 (nbe = K/256 floor); non-aligned K (e.g. 896) → missed elements (wrong).MediumGUARDED 2026-08-03 — quant_matmul_f32_on_gpu_buf aborts when id % 256 != 0 for Q4_K/Q5_K/Q6_K (GEMM and scalar paths both use K/256)
M2kernel_get_rows_q4_0 (embed) computes (token_id*nb+b)*Q4B with no token_id < vocab check. Sampler guarantees valid ids (low risk), but no defense.MediumGUARDED 2026-08-03 — embed_tokens_gpu aborts if any token_id >= vocab (host-side)
L1Matmul kernels compute ax pointers past the buffer for OOB rows; reads guarded by if (r0+N < p[0]) — pointer arithmetic only, no fault.LowAccept
L2store_kv has no in-kernel position bound; host kv_ensure_layer keeps positions < capacity.LowAccept
H3 (CUDA, C4 S2b)A packed Q8_0 kv4<LAYOUT> load assumes the 4-element group sits inside one 32-element block: it forms elem/32 and elem%32. A hd that is not a multiple of 32, or a KV head base that is not hd-aligned, would read a neighbouring block's scale/quants — wrong values (and, at the row end, past the cell).High (silent wrong values)GUARDED — KvFormat::Q8_0.check_width refuses a packed width that is not a non-zero multiple of 32 in ensure_kv, hd % 32 == 0 holds for every supported architecture, and the head base is always hd-aligned. cuda_kv_q8_0_roundtrip_attn and the Q8_0 arm of cuda_map_window_matches_the_span_over_the_same_rows (which compares a single-row window against the dequantized cell) fail loudly on a wrong block base; compute-sanitizer --tool memcheck over both reports no memory error
H4 (CUDA, C4 S2b)A KV cell is addressed by bytes (kv_row(base, cell, row_bytes)), and row_bytes comes from the backend's layout, while the region is sized from the builder's KvFormat in ensure_kv. Two policies that disagree would stride a packed region as f32 rows (the pre-S2b bool did exactly that for anything not f16).High (silent corruption)GUARDED — one tag (KV_LAYOUT_F32/F16/Q8_0 = the KvFormat discriminants) reaches the kernels, set_kv_cache_type maps q8_0 to KV_LAYOUT_Q8_0 rather than to false, models::load_model_ns re-states the layout from the resolved format, and the KvFormat::supports / reads_packed_kv gate refuses the load when the backend cannot read it

Verified clean: simple elementwise kernels (add/mul/silu/bias/swiglu/rope all have tid < n guards), Q4_0/Q4_1/Q5_0/Q5_1/Q8_0 matmuls (K reads bounded), GEMM smem/bc_out (within 8192 B), Q5_K qh/qs reads (within the 176 B block).

Post-audit finding (2026-08-19, fixed) — cross-kernel and within-kernel write visibility (not a deadlock, but a correctness race):

  • No memoryBarrier between dispatches in the single prefill compute encoder: dispatches are ordered but write-visibility across them is NOT guaranteed by Metal. bn reused as RMSNorm output / QKV input / WO output / ffn_down output raced → last-2-token garbage. Fixed with memoryBarrierWithScope after every dispatch_* (src/metal/), matching llama.cpp. Rule for future code: any buffer written by one dispatch and read by the next in the same encoder needs the barrier; do not rely on "it's serialized".
  • GEMM partial-tile temp_str overlaps sa/sb: after the K-loop, a fast simdgroup could overwrite sa/sb while a slow simdgroup still read them — threadgroup_barrier is required BEFORE the temp_str stores (all 8 mm kernels). Audit rule: when a threadgroup-memory buffer is REUSED for a different purpose at a different loop stage, there must be a threadgroup_barrier between the last read and the first write of the reuse.

4. Device metrics: query at runtime, never guess (2026-08-02 rule)

Rule: device-specific thresholds (threadgroup memory, max threads per threadgroup, alignment limits, etc.) MUST be queried at runtime via the metal crate's MTLDevice properties — never hardcoded from a guessed or remembered value.

Background: the initial H2 estimate used a guessed "32 KB threadgroup-memory limit" for the M4 Pro. The correct approach (what llama.cpp does, ggml-metal-device.m:851 + ggml-metal-ops.cpp:2367) is:

#![allow(unused)]
fn main() {
// src/metal/ — query the real limit and guard against it:
let shmem = 2 * 32 * hd * 4; // Bc * hd * 2 * sizeof(f32)
let max = self.inner.device.max_threadgroup_memory_length(); // real device value
if shmem > max {
    // CPU fallback
    return false;
}
}

The metal crate exposes DeviceRef::max_threadgroup_memory_length() and max_threads_per_threadgroup for exactly this purpose. Guards that hardcode a magic number should be re-examined: prefer a query, and document the queried value when a hard limit (like a kernel's fixed array size) genuinely exists.

4a. Split-attention and float4 kernel guards (2026-08-03)

  • kernel_gqa_attn_partial_f32 (KV-parallel split, pass 1) preserves the classic kernel's barrier discipline: every simdgroup reaches every threadgroup_barrier (empty KV chunks produce mx=-INF/S=0/acc=0 — they skip the tile loop together, so no barrier divergence). The final acc reduction MUST be a uniform d loop (all 32 lanes step the same d together) — a per-lane loop makes simd_sum reduce mismatched acc components (divergent reduction bug caught by the isolation test).
  • kernel_gqa_attn_combine_f32 (pass 2) is pure elementwise: no shared memory, no barriers. Guards: t<nt, h<nh, and m==-INFINITY → write zeros (avoids exp(-INF - -INF) = NaN).
  • New layer_gpu guard: hd % 4 == 0 (gpu_abort) — the float4 vectorized acc requires it. Existing hd <= 256 guard covers the acc4[64] array (64 float4s = 256 floats).
  • Shared-mutable-state change lesson (KV growth): a typo that cloned the K buffer into old_v during KV-cache growth polluted the V cache (Q4_K_M garbage). The split-vs-classic A/B did NOT catch it — both paths share the same corrupted KV. Any change to shared mutable GPU state (KV cache, buffer growth) must be checked against a known-good reference output, not just an A/B of two code paths over the same state.

4b. Flash-attention kernels (kernel_flash_attn_ext_f32/_f16, 2026-08-14)

The llama-port decode attention kernel (NSG=1, one 32-lane simdgroup per (t, h, chunk) threadgroup). Deadlock/race discipline:

  • Mask is computed inline per lane ((ic+NE*tx+ty < nkv) ? 0 : -MINF_MAXHALF), never via a shared sm[] array. llama's sm[tiisg] write → sm[NE*tx+ty] read is a cross-lane threadgroup access with NO barrier (works only by NSG=1 lockstep) — a race this kernel removes on purpose.
  • All control flow is break-only: if (ic >= nkv) break depends on lane-independent values, so all 32 lanes exit together. No continue, no per-lane early returns. Every lane reaches both threadgroup_barriers per chunk. Out-of-range KV reads are clamped to nkv-1 (in-bounds, value masked to ~0 via exp(-MINF_MAXHALF)).
  • Shuffle reductions are intra-simdgroup (simd_shuffle_down(8,4,2,1) + simd_shuffle(·, NL*ty) broadcast) — no threadgroup barrier inside the reduce. The route-to-lane-0/16 pattern keeps the DK4=16-lane reduction pure within each NE group.
  • Fixed-shape guard on the host: flash_attn_enabled(hd) gates dispatch on hd == 64 (DK/DV are hardcoded); layer_gpu falls back to the split path for any other hd (a support limitation, not a silent safety degradation).
  • shared ss[] handoffs (QK^T → softmax → PV) are the only cross-lane threadgroup accesses and are protected by the two threadgroup_barriers per chunk.
  • Partials {M, S, O[hd]} reuse the combine kernel's format; the combine's m==-INFINITY → zeros guard also covers empty flash chunks (iwg with iwg*C >= nkv, which break on the first iteration and write an empty partial).

5. Recurrence playbook

  1. On a GPU fault/hang, submit() now reports the dispatch trace (MINFER_TRACE=1 for labels). Reproduce with one app (kill Ollama / llama-cli / screen-capture tools), bounded -n generations, no long loops.
  2. Bisect with MINFER_GEMM=0, MINFER_CACHE_TYPE=f32, and git stash of the metal changes.
  3. If the machine freezes again: SSH in (enable Remote Login in System Settings) and run sudo spindump -n minfer + sudo spindump; check /Library/Logs/DiagnosticReports/ for new spin/ips artifacts.
  4. Only after 2+ recurrences in 1–2 weeks: run Apple Diagnostics (hold D at boot) and file a GPU hang report via Feedback Assistant.

CUDA (Phase 7, aarch64 GB10)

Hard rules (mirror the Metal section; enforced in graph/cuda_backend.rs + cuda.rs):

  1. Kernel-invariant violations return Err from execute_node — never a silent CPU fallback. Backend assignment at build time (supports_op + registered-weight gates) decides placement; a mid-execution guard failure aborts the run with the blocking node's name.
  2. No sync inside an active capture window. The 7d CUDA Graph capture wraps decode splits; any cudaStreamSynchronize (e.g. a debug print that reads device data) inside the window corrupts the capture — the 7e② "faster but wrong" incident was exactly this (a temp wrapper's internal sync produced garbage only with graphs ON). Debug reads must go through close_capture_or_sync paths.
  3. Device memory is not host-readable via plain memcpy on GB10 — host probes that copy_from_slice a device pointer SIGSEGV (__memcpy_sve). All D2H goes through cudaMemcpy staging (copy_to_host).
  4. A latched error is never blamed on the kernel that just ran. CudaState::sync() polls cudaGetLastError + cudaStreamSynchronize. The first reports whatever an earlier call on the thread latched — it is not evidence about the kernel — so its message names the observer and the cudaGetErrorName symbol and says explicitly that it is not attributed to a kernel, and the error is counted (latched_api_error_count) and cleared, never dropped. A persistent latched error is a call-site bug: find the call that discarded its return value and fix it there, do not add a louder sync. Issue #145: a rejected cudaFuncSetAttribute in the eager prefill-GEMM smem opt-in and a cudaGraphDestroy called on a cudaGraphExec_t both latched cudaErrorInvalidValue, and minfer bench printed it as "CUDA kernel launch error: 1". Its follow-on #147 closed the remaining sites of the same class (below) and added the MINFER_TEST_CALL_FAIL / MINFER_TEST_ISSUE147 gates that prove each one refuses loudly; #162 removed the class entirely — every <<<>>> in src/cuda/kernels/*.cu reads its own error, so a latched error at sync() is by construction an error no launch site read.
  5. A return value that gates a later launch or allocation is read where the call is made.
    • cudaFuncSetAttribute (the dynamic-smem opt-in) returns cudaErrorInvalidValue for a request above cudaDevAttrMaxSharedMemoryPerBlockOptin — a request already known to exceed the queried device limit must be skipped with the reason printed, never called just to observe the error (compute-sanitizer counts every such call, and the launch cannot succeed anyway). Any other failure must be named once (function, attribute, requested bytes, device limit, cudaGetErrorName), cleared at the call site, and the following launch refused — a launch over an un-opted-in dynamic smem cannot succeed, so it must not be issued into a checked error. Every opt-in in src/cuda/kernels/*.cu goes through minfer_smem_optin.
    • A <<<>>> has no return value, so its error is read with an immediately-following cudaGetLastError — that is a launch check, not attribution by position, provided nothing else runs in between: minfer_launch_prelude clears (and reports) a latch that predates the launch first, so the post-launch read can only be the launch's. Every launch goes through minfer_launch_ok or minfer_launch_ok_opt (#162), and the choice is the op's severity: minfer_launch_ok is a required launch, and its failure is recorded as a sticky that execute_node (via CudaState::take_launch_failure) turns into an Err naming the site, so the op never proceeds on a stale output; _opt is for a path with a documented fallback (the MMQ fast paths, the fa-prefill smem fallback, the int-returning launchers whose Rust caller decides) and only names and clears. The site token starts with launch: and the message names the kernel instantiation and cudaGetErrorName. scripts/check_cuda_launch_returns.py audits the source in CI, and a real-failure lever (minfer_launch_block — an illegal block geometry — or minfer_launch_smem) is part of every site's geometry, so the MINFER_TEST_ISSUE162=1 gates can drive each one into a real cudaErrorInvalidValue.
    • A cudaGraph_t (from cudaStreamEndCapture) is destroyed with cudaGraphDestroy, a cudaGraphExec_t (from cudaGraphInstantiate) with cudaGraphExecDestroy — mixing them returns cudaErrorInvalidValue and leaks the handle. Both returns are read; a failed graph destroy leaks the graph but leaves the exec valid, so it is named and cleared, and the exec is still returned.
    • The test knobs: MINFER_TEST_CALL_FAIL=<comma-separated site tokens|all> makes a hardened site perform its real call with a value that fails (an over-limit attribute / dynamic-smem request, an illegal block geometry, or cudaGraphDestroy on the exec), MINFER_TEST_ISSUE147=1 / MINFER_TEST_ISSUE162=1 enable the gates that drive it. All are unset in every default, bench and compute-sanitizer run; the gates name the site, the instantiation, the requested value and the error, and assert the latch is gone.
  6. Same-stream ordering is the correctness contract for async fills: the 7e⑥ pinned-staging write_input_async queues cudaMemcpyAsync on the stream and returns before the copy lands — safe because every consumer kernel runs later on the same stream. The staging ring syncs once when it wraps (>8 fills without a sync); never hand a slot back before that.
  7. The same contract, host-side, for async readbacks (F5): an cudaMemcpyAsync D2H into a pinned slab is stream-ordered but the host is not — the bytes are undefined until the event recorded after it has been waited on (cudaEventSynchronize). So (a) a pinned slab may only be reused or freed after its event wait, (b) nothing on the host may read the destination before it, and (c) the failure of the wait is a loud Err (never a fallback), because the alternative is reading bytes the DMA may not have written. CudaBackend::take_cross is the only place that wait happens, and GraphAllocator::cross_input refuses a staging entry whose wait has not been issued.
  8. Weight registry ownership: register_weight is name-keyed (same name+size reuses the device copy; different size replaces and deliberately leaks the stale buffer — bounded, a live captured graph may still reference it). Padded Q6_K repacks register through register_weight_q6k_padded and dispatch on is_weight_padded.
  9. Device limits are queried at runtime (SM count, compute capability, cudaMemGetInfo, and the per-function cudaFuncGetAttributes / per-device cudaDevAttrMaxSharedMemoryPerBlockOptin the smem opt-in reads); no hardcoded SM/arch assumptions beyond the compiled targets list.
  10. A registration-time expansion validates its payload against the layout it is about to index. A plane built by decoding a weight's raw bytes (the q6_K W_exp/_dsc, the q4_K _dsc, the q8_0 p32) carries a host-side row arithmetic — block count × block size — that the raw tensor only guarantees if its type matches. The gate is therefore two checks, and both are load-bearing: the type the plane is for, and the exact payload length that type's ratio implies (src/q4k_dsc.rs, called by the loader and re-checked inside register_weight_q4k_dsc). Issue #165: the q4_K W_dsc gate sat in the else of the Q6_K branch, so it ran for every non-Q6_K type in the loader's matches!; a q4_0 weight has exactly q4_K's bytes/element ratio, so only the type check can refuse it, and a q8_0 weight (longer rows) was misread into a plane no kernel reads — the size check is what a future smaller-ratio type (a 2-bit K-quant) hits, and it must refuse rather than read past the tensor. Exact equality, never >=: a lower bound admits the q8_0 length. The check cannot tell a q4_K payload from another type's bytes of the same length — that is the type check's job, and the pair is what is tested.

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0008 — GPU safety: bounded waits, no early return past a barrier, runtime device limits
  • ADR-0009 — A failure is an error, never a silent fallback

minfer Metal Backend Design

How minfer runs the compute graph on an Apple GPU: the MetalBackend graph executor, the src/metal/ device/kernel layer, the src/metal/kernels/ shaders, the command-buffer rhythm, and the memory and safety rules that hold it together.

Status. Landed. The backend arrived with compute-graph Phase 3, and the G1–G6 wiring passes plus the objc2 migration brought it to parity with the pre-graph imperative path. Every mechanism described here is implemented in the tree. Baseline: HEAD = 62a0a3e (2026-09-14).

Provenance. New document. It mirrors the layout of docs/CUDA-BACKEND-DESIGN.md and is the design of record for the Metal backend; the optimization campaign and its measurements stay in docs/METAL_OPTIMIZATIONS.md.

Related records. docs/METAL_OPTIMIZATIONS.md is the optimization history and current-state ledger (including §0.1, the graph-path integration status), docs/METAL_OBJC-ECOSYSTEM.md and docs/METAL-OBJC2-MIGRATION-PLAN.md cover the objc2 crate migration, docs/LLAMA_METAL_E2E.md is the llama.cpp Metal reference, docs/GPU_SAFETY.md holds the hard safety rules (Metal sections), and docs/inference_e2e_walkthrough/14-metal-backend.md narrates the backend for a first-time reader. The graph contract this backend implements is docs/COMPUTE-GRAPH-DESIGN.md §3.5/§7.


1. Design Goals and Outcome

1.1 Goal

Implement MetalBackend (src/graph/metal_backend.rs) as the macOS backend of the compute graph by wiring the existing per-op kernels on MpsState (src/metal/) — no new kernels required — so that on Apple Silicon the whole per-layer chain runs on the GPU through the standard build → assign → fuse → alloc → execute pipeline, with the same correctness contract as CPU: backend placement is decided at build time, kernel-invariant violations return Err, and there is never a silent mid-run fallback.

1.2 Outcome

GoalLanded outcomeEvidence
MetalBackend implements the Backend trait over MpsStateFull trait: shared-memory buffer pool, weight_buf offsets, per-op dispatch, direct host views, split-scoped command buffer§2, §4.1–§4.2
One command buffer per split, submitted at boundariescb() creates it lazily, synchronize() submits it; Drop flushes a pending one§2.4, §4.8
Per-node placement decided at build timesupports_op + the model-level weights_on_gpu all-or-nothing gate feed CParams.gpu; the scheduler splits the graph§4.3, §4.6
Zero-copy weightsGGUF parts are wrapped with newBufferWithBytesNoCopy; weights are (buffer, byte offset) pairs; a warm-up read moves the ~44 ms first-touch page cost to model load§2.3
Attention dispatch matches the old pathG1 wires nt==1 and nt>1 to the flash/split/parallel/classic kernels with the same gates and env vars§4.4, §4.6
Decode fusions on MetalG4 FusedQKV, G5 FusedFFN, G6 FusedQkvNorm (Qwen3) are built as single nodes§4.4, §4.6
Same correctness gates as CPUPer-op parity tests, cross-backend copy tests, model-level CPU-vs-Metal logits / greedy equality, kernel isolation tests§7
PerformanceGraph path at or above the old imperative path, re-measured 2026-10-06 at 6b95763 on macbook (macOS 27.0.1, Apple M4 Pro) (minfer bench -p <P> -n 128 -r 3): 0.5B Q4_0 decode 306.19 ± 1.00 tok/s / prefill pp440 6249.60 ± 10.52 tok/s; 7B Q4_K_M decode 48.52 ± 0.18 tok/s / prefill pp206 406.51 ± 0.83 tok/s; Qwen3-4B Q4_K_M decode 74.11 ± 0.19 tok/s (llama-Metal 79.7). Metal correctness: the external oracle graph_metal_matches_llama_reference reproduces the pinned greedy prefix, the model-level graph_metal_matches_cpu_logits compares a Layers(0) CPU engine against the full-plan Metal engine (restored by #324), and the per-op metal_*_matches_cpu gates are greenMETAL_OPTIMIZATIONS.md §0.1

1.3 Non-goals

  • A whole-layer layer_gpu fast path. The pre-graph imperative path was removed; src/metal/ is the per-op device/kernel layer only (the name survives solely in the legacy CUDA code).
  • f16/bf16 activations. Graph activations are f32; the f16 story is the KV cache (MINFER_CACHE_TYPE) and the flash-attention f16-KV kernel variants.
  • New kernels for the graph. The graph path is wiring plus the G2 rms_norm_256 selection and the decode *_off kernel variants that already existed.
  • Multi-GPU / device selection. One integrated GPU per Mac.
  • Training or fine-tuning.
TopicWhere
Optimization history, current state, graph-path status, env gatesdocs/METAL_OPTIMIZATIONS.md
objc 0.2 → objc2 crate migration (phases, gotchas, checklist)docs/METAL-OBJC2-MIGRATION-PLAN.md, docs/METAL_OBJC-ECOSYSTEM.md
llama.cpp Metal (MPS) end-to-end path (reference baseline)docs/LLAMA_METAL_E2E.md
GPU safety rules (barriers, bounded submit, capture windows)docs/GPU_SAFETY.md
Graph contract (IR, allocator, scheduler, backend trait)docs/COMPUTE-GRAPH-DESIGN.md
Beginner narrative of this backenddocs/inference_e2e_walkthrough/14-metal-backend.md
Kernel-level analysesdocs/metal-inference-analysis.md, docs/multi-token-kernel-analysis.md
Qwen3-4B vs llama.cpp (Metal)docs/PERF-QWEN3-4B-VS-LLAMACPP.md

2. Architecture at a Glance

2.1 Three layers

LayerFileRole
Graph executorsrc/graph/metal_backend.rsImplements Backend: shared-memory buffer pool, name→buffer-offset weight resolution, per-op dispatch, split-scoped MpsCommandBuffer, capture staging, error contract
Device/kernel layersrc/metal/MpsState singleton: device/queue/library init, zero-copy weight registry, MpsCommandBuffer (encoder, barriers, submit), and one Rust method per op/kernel
Shaderssrc/metal/kernels/The Metal Shading Language kernels (norms, matmul tiers, attention variants, elementwise, KV store, get_rows, fused epilogues)
Build chainbuild.rsRuntime shader compilation (default) or a precompiled metallib (MINFER_METALLIB_FILE/_PATH); the objc2 framework links

The split follows CUDA: src/metal/ is the only place that touches Objective-C/Metal APIs, and metal_backend.rs is the only place that knows about graph nodes.

2.2 The MetalBackend surface

#![allow(unused)]
fn main() {
pub struct MetalBackend {
    state: &'static MpsState,                 // process-wide MPS singleton
    pool: Vec<MetalBuffer>,                   // f32-element pool: id -> shared MTLBuffer
    free: Vec<usize>,                         // exact-byte-length free list
    staging: Vec<MetalBuffer>,                // trace/viz capture staging (blit targets)
    free_staging: Vec<usize>,
    cb_ptr: *mut MpsCommandBuffer<'static>,   // pending command buffer (null = none)
}
}

Pool buffers are StorageModeShared MTLBuffers — host and GPU see the same memory — so read_host/write_host are direct memory views and a cross-backend copy is a host round trip, not a staged transfer. The command buffer is stored as a leaked box pointer because MpsCommandBuffer is !Send/!Sync; all access happens sequentially through &self/&mut self, and the struct is unsafe impl Send/Sync on that basis.

2.3 Weight residency and registry

  • Zero-copy parts. MpsState::register_part wraps an mmap'd GGUF part with newBufferWithBytesNoCopy (StorageModeShared). The base must be page-aligned (16 KiB on Apple Silicon); mmap guarantees it, and the code debug_assert!s it. #39 audited the remaining debug_assert!s and kept this one deliberately: it is an OS contract on a stdlib call (mmap's guarantee), not a shape or kernel invariant, and register_part returns () so there is no Result for the value to travel in — the one asymmetry the release-build rule allows.
  • Per-weight offsets. register_weight(name, data) records (part buffer, byte offset); the executor resolves any weight through state.weight_buf(name) -> Option<(MetalBuffer, u64)> and passes the offset to the kernel. MINFER_WEIGHT_COPY=1 forces a copy per weight for A/B.
  • Load-time warm-up. The first GPU access to file-backed (mmap) pages costs a one-time page/TLB setup (~44 ms measured); a dummy full-buffer read at model load moves that cost out of the first prefill (METAL_OPTIMIZATIONS.md §0 Done #39).
  • KV element type (per engine, #44 part (b)). models::load_model_configured resolves MINFER_CACHE_TYPE once (kvformat::auto_device_format picks f16 when n_layers × n_kv_embd >= 8192 — the 7B class, measured ~−1 ms/token at 2K context — and f32 otherwise; f16 measured ~3% slower on the 0.5B, dispatch-latency-bound). The answer is stored on MetalBackend as kv_format, stamped through GraphAllocator::set_kv_format, and passed as an explicit f16 argument to every store/attention/fused decode. There is no process-wide Metal tag any more (ADR-0005, ADR-0006): two engines with different dims hold different layouts in one process. A third value, q8_0, is the packed cache the CPU, CUDA and Metal kernels read; Metal's packed path is enabled (#310): READS_PACKED_KV is true, two read mechanisms cover every shape (mechanism A native packed decode, mechanism B an f32 staging window for the fast prefill/window families), and the classic packed kernels stay the fallback — MINFER_CACHE_TYPE=q8_0 loads and runs. f16 remains the device default.

2.4 Command buffers and submission

One MpsCommandBuffer is kept for the current split:

  • cb() creates it on the first op of a split (a leaked box, so the returned reference is 'static and callers can still touch the pool).
  • Every dispatch helper ends with memoryBarrierWithScope(MTLBarrierScope::Buffers), so kernels in one command buffer see each other's writes.
  • synchronize() submits it; the scheduler calls that at split boundaries and once at the end.
  • Drop submits a pending buffer so an unterminated encoder can never be left behind.
  • MpsCommandBuffer::submit() commits with a dispatch-semaphore completion handler (avoiding the ~20 ms scheduler wakeup of waitUntilCompleted), waits with a bound and checks the command buffer status. On a failure it returns Err; the backend expects it, and device-configuration problems go through gpu_abort instead.

2.5 Legacy surface

The pre-graph whole-layer layer_gpu path was removed when the graph became the default: there is no layer_gpu function in src/metal/ (the name survives only in the legacy CUDA code), and src/graph/metal_backend.rs is the only live consumer of the device layer. What remains is a #[allow(dead_code)] block in src/metal/ holding the old-forward scaffolding and a few methods kept for tests (e.g. matmul_on_gpu_buf); the loaders, the graph backend and the kernel tests are the live callers. (The legacy KVCache type in src/cache.rs was likewise unused by the graph path; KV lives in the allocator's persistent regions, and the type — with the ModelDef::forward argument that kept it alive — was deleted in #252.)


3. llama.cpp Metal Reference Map

docs/LLAMA_METAL_E2E.md documents llama.cpp's Metal path end to end. What minfer borrowed, and what it deliberately does not do:

llama.cpp Metal conceptminfer analogStatus
Backend interface + scheduler splitsGraph Backend trait + split_graph; per-node execute_nodeBorrowed, reshaped
Multi-command-buffer scheme with status trackingOne command buffer per split; submit at boundariesSimplified
Per-op encoder + explicit memory barriersOne compute encoder per split, memoryBarrierWithScope(Buffers) after every dispatchBorrowed
MUL_MAT three-way kernel selection (matrix/tile/vector)quant_matmul_f32_on_gpu_buf tiers: simdgroup GEMM, _multi, single-tokenBorrowed in spirit, own kernels
FLASH_ATTN_EXT variant selection (DK/DV, f16 KV)gqa_attn_flash (nt==1) and attn_flash_prefill (nt>1) with hd ∈ {64,128} guardsBorrowed in spirit, own port
Fusion rules in the Metal backendminfer fuses at the graph level: FusionPass SwiGLU + build-time FusedQKV/FusedFFN/FusedQkvNormDiverged (graph-level fusion)
Unified-memory weight buffers (mmap or copies)newBufferWithBytesNoCopy over mmap'd GGUF parts, per-weight (buffer, offset)Borrowed
KV cache element-type policyModelDef::kv_format auto f16 for the 7B class, per engineBorrowed
MPS MPSGraph / higher-level MPS APIs—Not used (all kernels are hand-written MSL)
Multi-GPU / device selection, command-buffer concurrency tuning—Out of scope

4. Design

4.1 MetalBackend lifecycle and state

new() returns None when MPS is unavailable or MINFER_DISABLE_MPS is set (both handled inside MpsState::try_new); otherwise it stores the 'static singleton reference and starts with an empty pool and a null command-buffer pointer. MpsState::init() runs once at model load.

Pool rules:

  • Exact byte-length reuse only — alloc_buffer matches size * 4 against the free list, else allocates a StorageModeShared f32 MTLBuffer.
  • free_buffer never releases — the id goes back to the free list so persistent KV regions survive rebuilds; the MTLBuffer stays alive for the process.
  • alloc_fresh always allocates — split-boundary staging must not be recycled out from under in-flight node buffers.
  • No generation counter. Unlike CUDA there is no captured-exec cache, so no pointer invalidation machinery is needed.

Capture staging (staging / free_staging) exists only for trace/viz: staging_alloc reuses an exact-length staging buffer or allocates one, capture_split(src_ids) encodes blits at the end of the split's command buffer (after all kernels, so the staging holds this step's output) and returns ids valid only after the next submit, read_staging(id) reads them back, and release_staging_all() returns them to the free list after that split's readback.

In-place aliasing is the allocator's decision (sole consumer + same backend). The executor calls copy_in(dst, src) only when in_bufs[0] != out_buf — the non-aliased case — so an in-place kernel runs directly on out_buf, which for an aliased node is the producer's buffer.

MINFER_OP_PROFILE=1 accumulates host encode time per op label and per-submit GPU wait; the first submit prints a top-20 table, later submits print one line each; zero overhead when unset. Drop submits any pending command buffer (never leaving an unterminated encoder) and prints the profile.

4.2 Backend trait mapping

Trait methodMetal implementation
name()"metal"
supports_op(op, dtype)§4.3
supports_fused(fused)matches!(fused, FusedOp::SwiGLU) — the only variant in the enum
alloc_buffer / free_buffer / alloc_fresh§4.1
execute_node(node, in_bufs, out_buf, kv_pair)opens the split's command buffer once, then dispatches; guards return Err
read_host(id) / write_host(id, data)direct &[f32] views over shared memory (no staging, no copy)
synchronize()submit_pending()
graph_replay(..)not overridden (the trait default returns false) — Metal has no capture/replay path

Alongside the trait, MetalBackend exposes the capture helpers (capture_split, read_staging, release_staging_all) that the scheduler's trace/viz path calls, and the module exposes metal_available().

4.3 Eligibility

supports_op is:

OpMetal
Inputyes (any dtype)
Add, Mul, Silu, RmsNorm, QkNorm, SwiGLUF32
MatMulF32 activation (the weight type rides in MatMulMeta.weight_ttype)
GetRows, RoPE, Attn, KvcacheStore, KvcacheLoadF32
FusedQKV, FusedQkvNorm, FusedFFNF32
View, Reshape, Permuteyes (identity copy)
Scale, Softmax, BatchMatMulno (vocabulary only)
QkvBiasRopeStoreno — the mixed-quant decode epilogue is CUDA-only; on macOS the builder never emits it (the #52 recorded decision, with the dispatch arithmetic, is in docs/SUPPORT-MATRIX.md)

Two layers of checks are deliberately elsewhere:

  1. Weight-type eligibility is a model-level all-or-nothing gate (weights_on_gpu, §4.6): either every graph-referenced weight is registered on the GPU, or the model runs entirely on CPU.
  2. Shape invariants are enforced in execute_node and return Err: attention requires nkt == n_head_kv * hd (the kernel strides KV by nk*hd) and hd == hd_kv (it uses the query head dim); the KV pair must exist; a missing weight is an error; the fast attention paths are limited to hd ∈ {64, 128} and fall back to the classic kernel otherwise.

One Metal/CUDA divergence worth stating: Op::RoPE carries the style into the kernel (rope_style 0 = non-interleaved/Qwen2, 1 = interleaved/LLaMA), so Metal supports both styles; CUDA's supports_op gates RoPE to NonInterleaved only. All supported models are non-interleaved.

4.4 Execution dispatch

execute_node opens the split's command buffer once, then dispatches:

OpMetal path
Inputno-op (host-filled)
Silucopy_in if not aliased, then silu_f32 in place
Add / Muladd_f32 / mul_f32
RmsNormweight from NormMeta; rms_norm_256 when rms_norm_256_enabled(), else rms_norm. A missing gain (no NormMeta, no weight_name, or a name the device never registered) is a loud Err — never the weightless kernel (#40)
QkNormsame kernels with d = hd, n = len/hd over the flat [nt*nh, hd] rows
MatMulquant_matmul_f32_on_gpu_buf (tier below) + optional add_bias_f32
GetRows + Embed metaembed_tokens_gpu (per-weight-type row gather + dequant)
GetRows + no metaget_rows_f32 (the G3 tail-row selection)
RoPEcopy_in if not aliased, then rope_f32 with the node's rope_style
SwiGLUswiglu_f32
KvcacheStorekv_pair required; two store_kv calls (K then V); f32 or f16 by the engine's per-instance kv_format
KvcacheLoadno-op — the output buffer is the persistent K region
Attn§4.4.1
View / Reshape / Permutecopy_in when the output differs, else no-op
FusedQKVconcat matmul (blk.{i}.attn_qkv) + attn_bias_rope_store; refuses nt != 1 with Err
FusedFFNconcat matmul (blk.{i}.ffn_gu, od = 2*nf) + in-place swiglu_f32_off; refuses nt != 1 with Err
FusedQkvNormconcat matmul + two in-place per-head rms_norm[_256] (q at offset 0, k at byte offset nqt*4) + attn_rope_store; refuses nt != 1 with Err
Scale / Softmax / BatchMatMulErr("op ... unsupported on Metal (Phase 3)")
QkvBiasRopeStoreErr("op ... unsupported on Metal (CUDA-only)") — reaching it is a scheduling invariant violation

MatMul tiers (quant_matmul_f32_on_gpu_buf): a simdgroup GEMM kernel (64×32 tile, 128 threads, 8 KiB threadgroup scratch) when nt >= 2 && (od >= 2048 || nt >= 9) && gemm_enabled(); otherwise the _multi kernel for nt > 1 (one threadgroup per two output rows); otherwise the single-token kernel. Every supported quant type has the tiers that matter. Guards: K-quant id % 256 != 0 aborts via gpu_abort; the GEMM checks the threadgroup-memory request against the device limit queried at init. MINFER_GEMM=0 disables the GEMM tier for A/B. An unregistered weight dtype is refused with Err (#329) — the catch-all _ arm is a guard, not a fallback (see the f32 paragraph below).

f16 weights run on the device (#164, landed on a Mac 2026-10-06). The tiers above exist for the quantized types; f16 and f32 have their own single-token arms. The loader registers TensorType::F16 raw (2 B/element — no registration-time f32 copy) via the Metal branch's matches!(ttype, F32 | F16 | BF16), and quant_matmul_f32_on_gpu_buf's TensorType::F16 arm dispatches kernel_f16_f32_matmul (src/metal/kernels/f16.metal), a f32-activation matmul that promotes each half weight in-register over NR0*NSG = 8 output rows per 64-thread threadgroup (grid (ceil(od/8), 1), the token loop inside so a prefill re-streams a weight row once per threadgroup). The embedding gather is the sibling kernel_get_rows_f16, selected by embed_tokens_gpu's F16 arm (one element per thread, nb = ne). Both are listed in build.rs's SHADER_SOURCES and their pipelines (pl_f16_f32, pl_get_rows_f16) are built in try_new, so Qwen2Graph::weights_on_gpu passes and an f16 GGUF is a Metal model. Like CUDA, an f16 prefill runs this f32-activation kernel, not a simdgroup GEMM; 1-D norms/biases stay f32 (the file contract), so an f16 norm can never reach a kernel. Before #164 the type was refused here — registering a weight type a kernel cannot consume would make the device claim true while the op silently ran the wrong (or no) kernel, exactly what the registration gate exists to prevent. A second 2 B/element dtype (bf16, #208) is the subject of the paragraph after the f32 one. Per docs/SUPPORT-MATRIX.md, f16 is on both device columns.

f32 weights run on the device too (#317, landed on a Mac 2026-10-06). The loader's matches!(ttype, F32 | F16 | BF16) branch registers an f32 2-D weight raw (4 B/element), and quant_matmul_f32_on_gpu_buf's TensorType::F32 arm dispatches kernel_f32_f32_matmul (src/metal/kernels/f32.metal, the pl_f32_f32 pipeline) — the f32 twin of the f16 kernel and the peer of CUDA's launch_f32_f32_matmul. Before #317 an f32 weight had no arm and hit the catch-all _ arm (the Q4_0 kernel): kernel_q4_0_f32_matmul reads the f32 bytes as Q4_0 blocks (the first two bytes of 1.0f32 are 0x0000, an f16 scale of 0) and writes zeros. The gap was invisible because the reporting test, graph::op_matrix::matrix_cases_match_their_reference, only ran its Metal column when some earlier test in the process had already initialized MpsState; its Metal arm now calls MpsState::init() explicitly, exactly as its CUDA arm calls CudaState::init(), so the column no longer depends on test order. That silent-wrong-kernel fallback is the registration-gate failure the f16 paragraph above describes, and docs/SUPPORT-MATRIX.md's footnote 2 is updated to match. Two scope notes: no model in the gate set carries a 2-D f32 weight, so this path is covered by the synthetic metal_matmul_f32_matches_cpu gate and the op-matrix case only; and the catch-all _ arm that silently ran the Q4_0 kernel is gone — since #329 it returns Err naming the node, the observed dtype and the kernel that would have run (pl_q4_0_f32 / _multi), so the next unregistered dtype aborts instead of repeating #317's silent zero. That refusal is driven through the production dispatch by graph::metal_backend::tests::metal_matmul_refuses_an_unkerneled_weight_dtype, whose control arm builds the same graph with an F32 weight and asserts it still computes, so the dtype is the only difference (rule 2).

bf16 weights run on the device too (#208, landed on a Mac 2026-10-06). The second 2 B/element dtype — the Metal half of the ticket whose CUDA half is PR #321. kernel_bf16_f32_matmul + kernel_get_rows_bf16 (src/metal/kernels/bf16.metal, the pl_bf16_f32 / pl_get_rows_bf16 pipelines built in try_new and listed in build.rs's SHADER_SOURCES) are dispatched by the TensorType::BF16 arms of quant_matmul_f32_on_gpu_buf / embed_tokens_gpu — the f16 pair's geometry with an in-register as_type<float>(bits << 16) promotion (the device twin of crate::block::bf16_to_f32, exact for every value including NaNs). Its own kernel, not a dtype flag on the f16 one — the same decision #208's CUDA half made: bf16 and f16 are different 2 B/element layouts, so a shared kernel would branch per element in the hottest device kernel. Both loaders' Metal arm (matches!(ttype, F32 | F16 | BF16)) registers it raw (2 B/element, no f32 copy), so weights_on_gpu passes and both architectures are Metal models; 1-D norms/biases stay f32. The kernel-exactness gates bf16_matmul_matches_the_exact_shift_reference / bf16_embed_gather_matches_the_reference assert bitwise against crate::block::bf16_to_f32 (a wrong kernel — the f16 or f32 one — is red), and the ignored real-model gate f208_bf16_weights_run_on_the_metal_device measures 169 bf16 matmul + 1 embed nodes all on Backend::METAL, 942.4 MiB of device weights, max |Δlogit| 1.889e-3 absolute / 1.025e-4 relative (bar 0.05 / 5e-3) with an identical greedy continuation. Per docs/SUPPORT-MATRIX.md, bf16 is now on both device columns.

Aliasing. Only Silu, RoPE and the view ops call copy_in(dst, src), and only when the allocator did not alias them; an aliased node runs its in-place kernel directly on out_buf.

4.4.1 Attention dispatch

Pre-dispatch guards return Err: nkt == n_head_kv * hd (the classic kernel strides KV by nk*hd), hd == hd_kv (it uses the query head dim), and the layer's KV pair must exist.

  • Decode (nt == 1): flash_attn_enabled(hd) → gqa_attn_flash (chunked, partials merged by the shared combine kernel); else hd ∈ {64,128} and MINFER_NO_SPLIT_ATTN != "1" → gqa_attn_split_f32 (two-pass KV-parallel); else the classic gqa_attn_f32.
  • Prefill (nt > 1) with hd ∈ {64,128}: prefill_flash_enabled(hd) → attn_flash_prefill (the llama flash_attn_ext_blk port, with the tail-pad kernel for a partial last KV block); else matmul_attn_enabled() → attn_parallel_prefill (3-pass scores → masked softmax → output); else the classic gqa_attn_f32.
  • Other head dims always take the classic kernel.

Window modes (E1, issue #44 part (a), landed on a Mac 2026-10-06). The dispatch above is the causal path: every kernel derives token t's window from positions (nkv = positions[t] + 1, or a host max_pos + 1 for prefill). When the node is Op::Attn { explicit_span: true } — more than one sequence in a batch, or a run that does not start at cell 0 — the arm selects the mode from the size of the window input (topology, fixed at build time), mirroring CUDA's arm:

  • in_bufs[3].len == 2 * nt → the attn_span layout (one [lo, hi) pair per query). A prefill (nt > 1) at hd ∈ {64,128} with MINFER_NO_WINDOW_FLASH unset takes the fast windowed family kernel_flash_attn_window_blk_{f32,f16} / _hd128_{f32,f16} (src/metal/kernels/fa_window.metal, issue #359): a copy of the causal kernel_flash_attn_blk_* tile structure (Q=8 × C=64 simdgroup GEMM, inline online softmax, the kernel_kv_tail_pad tail) whose mask reads each query's explicit [window[t], window[nt+t]) instead of the causal [0, positions[t] + 1). The threadgroup processes the launch's global [lo_min, hi_max) union; the host advances K/V by lo_min for the tail pad and passes lo_min for the mask, so a windowed prefill does the causal tile work with extra blocks masked out. Every other shape — nt == 1 decode, any hd outside {64,128}, the opt-out — keeps the correctness kernel gqa_attn_window_f32 / _f16 (src/metal/kernels/attn_window.metal), which reads K/V at the run's cells [lo, hi) with the classic kernel's Bc = 32 tiling (the CPU gather cpu_gqa_attn_runs is the structural reference). f32/f16 is selected from the engine's KV format for both families; the fast family is the one measured below, the correctness family stays its reference.
  • in_bufs[3].len == nt * KV_MAP_MAX_SPANS * 2 → the kv_map layout (a list of (cell, len) runs per query, C8b S2/S4). A prefill (nt > 1) at hd ∈ {64,128} with MINFER_NO_WINDOW_FLASH unset takes the fast map family kernel_flash_attn_window_map_{f32,f16} / _hd128_{f32,f16} (also in src/metal/kernels/fa_window.metal, issue #369): the same Q=8 × C=64 tile over the launch's global [lo_min, hi_max) union, with the mask a run-membership walk (fwin_map_has, ≤ KV_MAP_MAX_SPANS runs) instead of [lo, hi) — the arithmetic CUDA does in attn_map_nkv / kv_cell. Every other shape — nt == 1, any hd outside {64,128}, the opt-out — keeps the #362 correctness kernels gqa_attn_map_f32 / _f16 (also in src/metal/kernels/attn_window.metal), which resolve each flat window row to a cell by walking the ≤ 4 runs instead of the window's lo + ki. The correctness one-range and map kernels are a separate family on purpose: their instruction streams are the #315 measured contract, so their code path is byte-untouched. A sharing sequence's window (a shared prefix plus a private run) is exactly this shape; the input's size selects the layout.
  • anything else → a loud Err naming the accepted sizes.

supports_attn_span() is now true (SUPPORTS_ATTN_SPAN) and Device::gathers_attn_map is now true for Metal, so both explicit layouts are read on the device. The causal paths (flash / split / parallel-prefill / classic) are byte-untouched: the windowed families are used only for an explicit window, so a single-sequence causal forward keeps its previous numbers.

Packed q8_0 KV is enabled (#310). Metal reads a packed region: READS_PACKED_KV = true (the registry's reads_packed_kv, which GraphAllocator::ensure_kv and KvFormat::Q8_0::supports read), so MINFER_CACHE_TYPE=q8_0 loads and runs on Metal and a C5 session file carries FLAG_PACKED. The store (kernel_store_kv_q8_0, byte-identical to the CPU quantizer) and two read mechanisms cover every attention shape, selected by the pure crate::metal::packed_attn_route:

  • Mechanism A — native packed reads in the decode flash family. kernel_flash_attn_ext_q8_0 and kernel_flash_attn_ext_hd128_q8_0 (fa_decode.metal) are the f16 kernels' twins — the same threadgroup layout, chunk loop, barriers, MINF_MAXHALF masking and partial buffer, so the shared kernel_gqa_attn_combine_f32 merges them unchanged. Only the per-lane K/V read changes: one 34-byte block's four dequantized elements via dequant_q8_0_kv4 (dequantize.h, four scalar byte loads — a block begins at row*row_bytes + 34*b, and 34*b is 2 mod 4 for odd b, so no vector load is alignment-safe). row_bytes is the one added argument (buffer 11), so buffers 0..10 keep their indices. nt == 1, hd ∈ {64,128}, causal.
  • Mechanism B — an f32 staging window for every other packed path. kernel_dequant_kv_q8_0_to_f32 (kv.metal) dequantizes the needed window of packed cells into a transient, arena-addressed f32 scratch buffer (row i = cell i), then the unchanged f32 prefill (kernel_flash_attn_blk_*) or windowed-flash (kernel_flash_attn_window_{blk,map}_*) family runs against it with f16 = false. Because the stage is arena-addressed, the windowed kernels keep their absolute lo_min addressing and the tail-pad kernel its relative offset — no kernel in those families is edited. nt > 1 causal prefill and every nt > 1 explicit window.
  • Anything neither covers — a small/odd hd, an nt == 1 explicit window, any MINFER_NO_* opt-out — keeps the classic packed kernels kernel_gqa_attn_q8_0 / kernel_gqa_attn_window_q8_0 / kernel_gqa_attn_map_q8_0, still reachable and still the correctness fallback.

MINFER_PACKED_KV_STAGE=1 routes a shape mechanism A would take through mechanism B, so the decode A/B is measurable; the differential gate (metal_packed_decode_stage_matches_the_native_read) drives that switch through a #[cfg(test)] thread-local rather than the environment.

Why f32 staging, not f16. Staging to f16 would round every already-Q8_0-dequantized cell element a second time. The extra rounding is small per element, but on Qwen3-0.6B it is amplified through 8 decode steps: the interleaved f16-stage run measured max |Δlogit| 16.9 (at the argmax 9.46) against the f32 engine, over the same prompt where the classic f32-dequant packed path measures 1.29 / 0.43 and the f16 cache measures 0.011. The prefill-step delta is tiny (0.19) but compounds; the mechanism is the decode trajectory's sensitivity to the specific error realization, not a kernel fault. f32 staging removes the second rounding and restores the classic / CUDA class (docs/SUPPORT-MATRIX.md): the ignored real-model gate measures 1.28 / 0.36 on Qwen3-0.6B and 2.37 / 0.46 on Qwen2.5-0.5B, both inside the inherited ≤ 4.0 tail / ≤ 1.0 argmax bars — at parity speed (below).

Measured (issue #310, macbook (macOS 27.0.1, Apple M4 Pro), hostname macbookpro-ysw, 2026-10-08). Size and speed with the capability enabled and the two mechanisms in place, MINFER_CACHE_TYPE=f16 vs q8_0, 3 interleaved runs (minfer bench -p 1024 -n 64 -r 3 --n-ctx 2048, medians):

modelKV regions f32 vs q8_0pp1024 f16 → q8_0tg64 f16 → q8_0
Qwen3-0.6B Q8_0 (hd 128)58 720 256 B → 15 597 568 B (3.76×)4 844 → 4 683 tok/s (0.967×)193.5 → 176.6 tok/s (0.913×)
Qwen2.5-0.5B Q4_0 (hd 64)6 291 456 B → 1 671 168 B (3.76×)6 175 → 6 154 tok/s (0.997×)295.7 → 245.5 tok/s (0.830×)

The memory win is 3.76×. Prefill (mechanism B) lands within noise of f16 (0.997× / 0.967×); decode (mechanism A) is 1.20× / 1.10× slower than f16 — the native packed kernel issues five 2-byte loads per lane against the f16 kernel's one vectorized half4, so at these short, dispatch-bound contexts it is LSU-bound rather than bandwidth-bound and the 34 B/32-element read does not pay off. That is inside the CUDA precedent after #144 (1.01–1.24×). MINFER_PACKED_KV_STAGE=1 on the 0.5B tg64 shape (mechanism B decode: dequantize, then the f32 flash) measures 238.8 tok/s, ~2.7% below mechanism A's 245.5 — the stage cost is small and the mechanism-A decode is the slower one. The gate metal_packed_decode_stage_matches_the_native_read pins A vs B at max |Δ| 1.7e-8 (f32 reduction order). This is the enabled behaviour, not a refusal.

Measured (issue [#315], macbook (macOS 27.0.1, Apple M4 Pro), hostname macbookpro-ysw, 2026-10-07). Bar named before the run (docs/GATE-CONTRACT.md rules 3 and 5): the windowed arm's tokens/s must reach >= 0.8x the causal prefill's at the same total token count, on the median of interleaved matched rounds (rule 4). It was not met on either model. Three runs each of

cargo test --release --bin minfer -- --ignored --nocapture a_windowed_prefill_is_not_materially_slower
MINFER_315_MODEL=~/.cache/minfer/models/hf/Qwen/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf \
  cargo test --release --bin minfer -- --ignored --nocapture a_windowed_prefill_is_not_materially_slower

(a 5-round interleaved harness, #[cfg(target_os = "macos")] #[ignore]d, in src/models/qwen2/graph/tests/batching.rs; it prints the model, n_ctx, the node/attention-node counts and the kernel each arm took). Both models use n_ctx = 1024, n_total = 512; the second exercises the f16 windowed kernel and hd = 128, the first the f32 kernel at hd = 64.

Model (KV, hd)causal single seq (attn_flash_prefill)windowed two-seq batchwindowed one-seq @ cell 128
Qwen2.5-0.5B Q4_0 (f32, 64)6387 ± 10 tok/s3439 ± 52410 ± 3
Qwen3-0.6B Q8_0 (f16, 128)5258 ± 17 tok/s838 ± 1505 ± 1
Pair (median of the 3 runs)0.5BQwen3-0.6B
primary — one 512-token causal sequence vs a two-sequence batch of the same 512 tokens0.539x (0.537 / 0.541 / 0.539)0.159x (0.159 / 0.160 / 0.159)
shape-matched — one causal sequence vs the same sequence behind a 128-cell holder, same nt/n_out/tokens, only the kernel differs0.376x (0.375 / 0.377 / 0.376)0.096x (0.096 / 0.097 / 0.095)

So the windowed kernel runs at ~0.54x / ~0.16x the causal prefill's tokens/s for the two-sequence batch (~1.9x / ~6.3x slower) and ~0.38x / ~0.10x at the same shape (~2.7x / ~10.4x slower); the f16/hd = 128 instantiation is the worse of the two.

The causal kernel and code path are untouched (the round adds only the harness); the gap is the simple windowed kernel's serial per-(query, KV-head) walk over the run, which the tuned attn_flash_prefill tiles. A windowed fast path — tiling the run list the way fa_prefill.metal tiles a contiguous window, without disturbing the causal kernels (the #137 lesson) — was therefore warranted and is filed as #359; the correctness kernel stays the reference.

Measured (issue [#359], macbook (macOS 27.0.1, Apple M4 Pro), hostname macbookpro-ysw, 2026-10-07). Same harness, bar and n_ctx/n_total as the #315 run above (median of 5 interleaved rounds, three runs each). The fast family is now selected for every windowed prefill arm; the causal arm and its kernel are unchanged, and both windowed arms print kernel_gqa_attn_window_* because the harness's kernel label is hard-coded to the correctness family (the harness is the yardstick and was re-run unchanged).

Pair (median of the 3 runs)0.5B (before → after)Qwen3-0.6B (before → after)
primary — one 512-token causal sequence vs a two-sequence batch of the same 512 tokens0.540x → 0.968x0.158x → 0.918x
shape-matched — one causal sequence vs the same sequence behind a 128-cell holder, same nt/n_out/tokens, only the kernel differs0.379x → 0.996x0.093x → 0.995x

Both arms clear the >= 0.8x bar on both models (0.5B primary 0.968 / 0.968 / 0.971, shape 0.995 / 0.996 / 0.997; Qwen3-0.6B primary 0.918 / 0.917 / 0.920, shape 0.995 / 0.997 / 0.995). The residual gap in each primary pair is the two-sequence batch's own per-forward overhead, not the attention kernel: the windowed-global launch runs the same tile count the causal prefill does. With MINFER_NO_WINDOW_FLASH=1 both pairs fall back to the #315 numbers above, which is the A/B control for the fast family.

Measured (issue [#369], macbook (macOS 27.0.1, Apple M4 Pro), hostname macbookpro-ysw, 2026-10-07). The same harness gained a fourth arm for the set-valued kv_map layout: a donor owns the first N_TOTAL/2 positions and a subject reads them in place while prefilling its own N_TOTAL tokens after them, so its window is a shared run plus a private run. First-fit puts the subject's run at cell N_TOTAL/2, adjacent to the donor, so the map's global union is one contiguous range and the fast map kernel pays the same tile cost the one-range fast kernel does (the run walk is the only extra work). Bar and protocol are the #315/#359 ones (median of 5 interleaved rounds, three runs each). Before: the #362 correctness map kernel; after: the #369 fast map kernel.

Map arm (median of the 3 runs)0.5B (before → after)Qwen3-0.6B (before → after)
map — one 512-token causal sequence vs one shared-prefix kv_map subject, same N_TOTAL tokens0.229x → 0.943x0.046x → 0.914x

Both reach the >= 0.8x bar (0.5B 0.942 / 0.947 / 0.943; Qwen3-0.6B 0.912 / 0.915 / 0.914). MINFER_NO_WINDOW_FLASH=1 restores the before numbers on both layouts (map 0.229x, one-range primary 0.539x / shape 0.378x), the A/B control the win is attributed through. The map arm's residual gap to the causal prefill is the adjacent shared run (the subject's attention window is N_TOTAL/2 rows wider than a causal query at the same position), not the run walk: the union is contiguous, so fwin_map_has runs over the same tiles the one-range kernel would.

Measured: the map's remaining shape limit. A window whose runs are far apart pays the gap between them — the fast map kernel tiles the global [lo_min, hi_max) union, while the correctness kernel walks just the runs; a map whose shared prefix and private run are non-adjacent would therefore do the union's tile work, not the runs'. The harness's first-fit layout (the production server's) is adjacent and does not pay it; a kv_defrag-spread layout could. This is the recorded tradeoff, not a silent one.

Gate: metal::tests::window_map_flash_matches_the_cpu_reference drives the fast map kernel directly at hd 64 / hd 128 (f32 and f16), with a shared run plus a growing private run, both adjacent (gap = 0) and gapped (gap = 8, so the run walk is load-bearing), nt = 197 and a non-zero lo_min. Bars named before measuring: a one-cell map returns the named cell's V row bitwise (0 on every case), the two-run map matches the CPU reference to <= 0.01 (f32) / <= 0.05 (f16) — measured 1.0e-5 to 2.0e-5 — and a wrong shared base changes the output. The #362 map gates (metal_map_single_cell_matches_the_v_row, metal_map_matches_the_span_and_a_wrong_base_differs) and the #44/#359 gates are unchanged and green.

Gate: metal::tests::window_flash_matches_the_cpu_reference drives the fast kernel directly at hd 64 (f32 and f16) and hd 128 (f32 and f16), with lo_min ∈ {0, 64, 96}, nt = 197 (a non-multiple of 8, so the Q-tile tail padding runs) and nkv = 197 (a partial 64-row tail, so kernel_kv_tail_pad runs). Its bars were named before measuring: a one-cell window returns the named cell's V row bitwise (max|Δ| = 0 on every case), the growing window matches the CPU reference to <= 0.01 (f32) / <= 0.05 (f16) — measured 1.5e-4 / 2.7e-4 — and a window shifted one cell changes the output (the rule-2 control). The #44 correctness gates (metal_attn_span_matches_cpu, metal_attn_span_multi_key_matches_cpu, metal_attn_span_nonzero_start, metal_map_single_cell_matches_the_v_row, metal_map_matches_the_span_and_a_wrong_base_differs) are unchanged and green.

4.4.2 KV write/move side (issue #44 part (b), landed on a Mac 2026-10-06)

  • C3 row move — Backend::copy_cells. CUDA-rejecting on both directions (dst_row above or below src_row, overlapping): the arm opens one MTLBlitCommandEncoder in the current command buffer and copies row by row in the overlap-safe order — ascending when the run slides down, descending when it slides up (dst_row <= src_row), the mirror of CUDA's kv_move_rows. A bulk blit is not an option: Apple documents an overlapping same-buffer copy as undefined, and a separate submission would overlap the producer split (#137). f16 halves elems_per_cell (half[cell * nkt] is nkt / 2 f32 words apart); a packed Q8_0 cell is already whole f32 words including padding and is passed through unchanged (the caller passes region.elems / n_ctx, which ensure_kv sized to KvFormat::Q8_0.row_elems), so the move is a plain whole-word copy that keeps every packed block and its padding verbatim (#310). Gates: metal_copy_cells_moves_overlapping_rows_in_both_directions, metal_f16_kv_cell_move_strides_by_row_bytes.
  • C2 shift / C5 sessions — GraphAllocator::copy_kv_to_cpu. The CPU || CUDA hardcode is gone; the read goes through the registry host_read hook for every backend, so a Metal session shifts (kv_rm/kv_shift) and saves (kv_save*) instead of re-rendering. The read is ordered after the split's submission (copy_kv_to_cpu takes &mut self and calls sync_backend first — MetalBackend::read_host takes &self and does not submit, so a pending buffer would be read stale, the #301 shape). Gate metal_copy_kv_to_cpu_reads_after_the_pending_split leaves a real store dispatch un-submitted and reads through copy_kv_to_cpu: without the flush the read is the region's zeros. The C2 re-rope stays f32-only: an f16 region has no host map (the raw halves would be rotated as f32), so kv_rm — and its start == 0 spelling kv_shift — refuses it loudly, naming #306; CUDA is exposed too and is not fixed here. That refusal is the recorded end state, pinned by graph::alloc::tests::kv_shift::an_f16_region_refuses_the_physical_shift_and_the_other_formats_take_it (a genuine set_kv_format(F16) region, with f32 and packed q8_0 controls that do shift on the same fixture), and the caller re-renders the retained window instead — measured on dgxspark (aarch64, GB10 sm_121), 2026-10-07, 7B Q4_K_M --cnv overflow at --n-ctx 512: 499–505 tokens / 0.30–0.35 s re-prefilled per overflow, against the shift's 22-token / 0.26 s delta.
  • Per-engine kv_format. MetalBackend now carries its own KvFormat, stamped from GraphAllocator::set_kv_format (and read from the allocator stamp on enable_metal), and every attention/store dispatch takes it as an explicit f16 argument instead of reading a process-wide tag. The old metal::KV_F16 OnceLock, kv_cache_is_f16, set_kv_cache_type and the #[cfg(test)] override are deleted: two engines in one process now hold their own layouts and a C5 session header is described under the width its region really uses (the registry kv_format hook reads the allocator's stamp so it answers before the pool is enabled — the CUDA hook's shape). Gate metal_kv_format_is_per_engine.
  • Packed q8_0 is enabled since #310 (READS_PACKED_KV = true); §4.4 records the two mechanisms and the measurements, and docs/SUPPORT-MATRIX.md carries the format row.

The decode chunk count is MINFER_ATTN_CHUNKS or ((max_pos + 1 + 31) / 32).clamp(1, 16) — one chunk per 32 KV rows, capped at 16. nkv for the prefill kernels and the chunk count come from a host read of the positions buffer; that is safe because positions are host-written input data, never GPU-computed. The read is bounded by the node's logical length (positions_max(positions, nt), BufRef::len), not the class-rounded pool buffer: a recycled positions buffer keeps a prior graph's tail, and scanning it derived a stale max_pos that made a causal prefill attend past its own window — the reused_cache_across_prompts_matches_a_fresh_cache failure fixed in #44 part (b). (CUDA instead derives the bound on device so nothing host-side enters a captured graph; Metal has no replay to protect, so the host read is free.)

One measured order-dependence. a_compaction_between_steps_keeps_the_continuation compares two paths (a prefill whose KV is compacted mid-session against one that is not); the two take different kernels, and under the full-suite Metal state the drift is ~0.0058 while it is 0 when the gate runs alone — the same class as matrix_cases_match_their_reference. The gate therefore keeps its behavioural assertion (the greedy token) rather than a tight numeric one, and the number is recorded here so a future full-suite run does not read it as new.

4.4.3 Prefill flash tail-pad overlap (issue #314, landed on a Mac 2026-10-06)

The flash prefill kernel (kernel_flash_attn_blk_f32/_f16, incl. the hd=128 variants) tiles KV into C = 64 blocks. For a partial last block (nkv % C != 0) it read the last C rows (pos0 = nkv - C) from the [2][64][nkt] tail pad. That window overlaps the previous full block whenever nkv % C != 0 && nkv > C (e.g. nkv = 72: block 0 covers rows 0–63, the "partial" block covers 8–71), so the online softmax counted the overlapped rows twice — a real attention error (measured 0.05–0.16 on a synthetic reference), which surfaced as a 1.445 (Qwen2.5-0.5B) / 0.382 (Qwen3-0.6B) chunked-vs-unchunked logit drift because every chunk > 64 tokens hit it, and as a wrong answer for any >64-token prefill. The fix: the partial block reads its own [ic, ic + C) window (pos0 = ic) and kernel_kv_tail_pad pads from ic = (nkv / 64) * 64, so the rows past nkv are zero+masked (kpos0 <= qpos) instead of the leading rows being double-counted. CPU is bitwise across chunk shapes; after the fix Metal's residual drift is the ordinary cross-shape accumulation class (named 0.1, measured 0.0087 on the 0.5B and 0.0078 on Qwen3-0.6B), which the a_chunked_prefill_answers_like_an_unchunked_one gate now carries alongside CUDA. Why no gate caught it for so long: the only cross-shape Metal gate was graph_metal_matches_cpu_logits, whose prompt is 33 tokens — below C = 64, so it never reached a partial block, and the drift was invisible by construction. The isolating experiment named the kernel rather than an accumulation-order effect: MINFER_NO_PREFILL_FLASH=1 (the 3-pass parallel attention) drops the drift from 1.445 to 0.007, and a chunk sweep (1.71 / 1.61 / 1.68 / 1.45 for different chunk counts) is not proportional to the number of chunks, which rules out accumulation order and leaves the flash-prefill kernel as the entry. Gates: flash_prefill_matches_the_cpu_reference_at_every_kv_tail (nkv 72/96/98 red before at 0.164, all ≤ 2e-4 after) and the real-model chunked gate.

4.5 Allocator and scheduler integration

  • Assignment priority is Metal → CUDA → CPU; enable_metal() mirrors enable_cuda().
  • KV regions are created by ensure_kv on the layer's assigned backend, so a Metal-assigned layer's K/V live in the Metal pool and kv_pair resolves to pool ids the executor passes to the kernels.
  • Cross-backend values are staged by the allocator (alloc_fresh on the consumer's backend) and copied through host memory — with shared MTLBuffers that is a plain copy_nonoverlapping.
  • The scheduler calls sync_backend (→ MetalBackend::synchronize) at every backend change and after the last split, which is exactly when the split's single command buffer is submitted.
  • Because Metal supports the whole Qwen2/Qwen3 op set (including the embedding and tail gathers), a normal forward is a single Metal split; CPU splits appear only for an op Metal refuses (Scale/Softmax) or in synthetic/mixed graphs — and the tests deliberately exercise the multi-split alternation.

4.6 Model wiring

metal_on = metal_available() && weights_on_gpu(model)   // #[cfg(target_os = "macos")]
CParams.gpu = metal_on || cuda_on
  • weights_on_gpu is registration-only: it builds the exact list of weight names the graph reads (tok_embd, output_norm, output, output_b, and per layer attn_norm, wq, [bq Qwen2], wk, [bk], wv, [bv], wo, ffn_norm, ffn_gate, ffn_up, ffn_down, plus Qwen3's q_norm/k_norm) and requires mps.has_weight(name) for every one. It is all-or-nothing: one missing name and the model runs entirely on CPU. Unlike CUDA's weights_on_cuda there is no type whitelist here and no diagnostic print — a Metal gate failure is silent.
  • Type support is enforced at loader registration: load_ti registers a tensor only when its type is Q4_0/Q4_1/Q4_K/Q5_0/Q5_1/Q5_K/Q6_K/Q8_0, or F32/F16/BF16 (the last two since #164 and #208). An unsupported type is simply never registered, so the gate fails and the model falls back to CPU. Names are namespaced (mps.register_part per mmap'd part; {ns}{tensor} for the registry) so a second model cannot collide with the first.
  • Concat weights are built once at load with metal::concat_rows and registered: Qwen2 and Qwen3 both register blk.{i}.attn_qkv (all three of wq/wk/wv present, same type/dim, block-aligned) and blk.{i}.ffn_gu, the latter only when nf <= 16384 && MINFER_NO_FUSE_FFN != "1" (the 7B concat would otherwise hold ~2 GiB of weights no node reads).
  • Fused nodes per model: Qwen2 builds FusedQKV for decode QKV; Qwen3 builds FusedQkvNorm (per-head Q/K norm, which the Qwen2 bias+rope+store kernel cannot express); both build FusedFFN when nf <= 16384. The mixed-quant QkvBiasRopeStore class is CUDA-only, so those layers keep the unfused chain on macOS. Gates: CParams.fuse_qkv / fuse_ffn (Qwen3's fuse_qkv is metal_on-only), so no CUDA presence enables a Metal-incompatible fusion.
  • FusionPass gets [cpu, metal, maybe cuda] and a backend_of that maps CPU→0, Metal→1 and CUDA→the index found by name; the SwiGLU rewrite is gated by supports_fused(SwiGLU) and applies to the unfused FFN path (when FusedFFN is built, silu+mul are inside the fused kernel).

4.7 Memory, residency and staging

  • Unified memory changes the copy math. Activations are StorageModeShared MTLBuffers, so host reads/writes are direct views and copy_across is a host round trip; there is no pinned staging ring and no pool_gen.
  • Weights are zero-copy over the mmap'd GGUF parts, with the ~44 ms first-touch page cost paid at load by a dummy warm-up read (METAL_OPTIMIZATIONS.md §0 Done #39). MINFER_WEIGHT_COPY=1 forces per-weight copies.
  • E4 charges the registered weights (#299). The registry (MpsStateInner::weights) stores each weight's own byte extent, and MpsState::weights_bytes sums it (recovering a poisoned lock rather than reporting 0, the CUDA twin); MetalBackend::weights_bytes forwards it, so weights + pooled + request > budget is the one comparison on Metal too. Until #299 the trait-default 0 let the E4 gate ignore the resident weights the E5 auto fit already charged from the GGUF index — it could admit weights against a budget that never counted them. The recorded Mac measurement is in the plan's #299 entry. The gate is graph::alloc::tests::budget::metal_registered_weights_are_charged_in_the_budget_gate; with weights_bytes forced back to 0 its budget refusal does not happen and the gate fails at the unwrap_err. The per-weight extent is tensor.data().len(): 2 B/element for f16/bf16, the block size for a quant — the same bytes CUDA's registry sums.
  • KV element type. The persistent regions stay f32-shaped in the IR, but the engine's per-instance kv_format picks f16 for the 7B class (KV-bandwidth-bound; measured ~−1 ms/token at 2K) and f32 for small models (f16 measured ~3% slower there); MINFER_CACHE_TYPE overrides. The arm is GraphAllocator::set_kv_format (per engine, #44 part (b)).
  • Capture staging exists only while trace/live capture is armed; per-split blits write node outputs into host-readable staging at the end of the command buffer, read back after submit, then released.
  • Device limits are queried once at init (maxThreadgroupMemoryLength) and referenced by dispatch-time guards; the remaining hardcoded numbers are kernel-declared array sizes, documented in GPU_SAFETY.md §3.
  • Profiling. MINFER_OP_PROFILE=1 reports host-encode time per op and per-submit GPU wait, which is how the decode dispatch-cost story in METAL_OPTIMIZATIONS.md was measured.

4.8 Command buffers, submission and trace capture

One MpsCommandBuffer per split, submitted at boundaries, is the whole execution model:

  • Every dispatch helper ends with memoryBarrierWithScope(MTLBarrierScope::Buffers). Metal does not guarantee write visibility between dispatches in one compute encoder; the missing barrier caused intermittent last-2-token corruption on 1.5B/7B prefill before the 2026-08-19 fix.
  • submit() commits with a dispatch-semaphore completion handler (avoiding the ~20 ms scheduler wakeup of waitUntilCompleted), waits with a 10 s bound, and requires MTLCommandBufferStatus::Completed; otherwise it returns Err carrying the recent dispatch trace. The backend expects the submit result; device-configuration problems go through gpu_abort (print the actual limits and exit) rather than degrading silently.
  • Trace/viz: the scheduler pushes each non-KV Metal node's output buffer, and at the split end capture_split encodes the blits; after sync_backend submits, flush_metal_captures reads every staging buffer and records its stats. KV regions are skipped (a full region per layer would dominate the trace).
  • No graph replay. Metal has no CUDA-Graph equivalent here; graph_replay is the trait default. MINFER_METAL_CAPTURE=1 instead starts an MTLCaptureManager GPU trace at init for Xcode, and MINFER_TRACE=1 records a 16-deep per-dispatch label ring used to diagnose a GPU fault (note: MINFER_TRACE also names the graph trace path in the CLI, which is a separate mechanism).

4.9 GPU safety (Metal edition)

docs/GPU_SAFETY.md holds the rules and the incident history; the backend implements them as:

  1. Bounded submit with a status check — never DISPATCH_TIME_FOREVER, never a silent non-Completed status.
  2. No early return past a threadgroup_barrier — the attention kernels were rewritten to run a dummy head instead of returning early, then skip the output write via a valid_head flag.
  3. Device limits queried at runtime — threadgroup memory and thread limits come from the device at init; guards compare against the queried values.
  4. Barriers between dispatches — a buffer written by one dispatch and read by the next needs the explicit scope barrier; a reused threadgroup-memory buffer needs a threadgroup_barrier between the last read and the first write.
  5. Err from execute_node, never a CPU fallback — missing weight, bad shapes, missing KV regions, unsupported op.
    • KV-store bound (#38, gap-table G1). Every Metal arm that writes the persistent K/V region — KvcacheStore (row = cells), FusedQKV and FusedQkvNorm (row = positions) — reads the small StorageModeShared index buffer back and returns Err naming the offending cell and the region's n_ctx before any dispatch (MetalBackend::check_kv_store_rows). A kernel-side range check would be a silent no-write, which this rule forbids. The allocator bounds the same input on the fill_input_i32 path (GraphAllocator::check_positions_bound), so the arm guard closes the fill paths that bypass it. Gate: metal_kvcache_store_refuses_a_cell_past_the_arena. Scope: the check compares against the layer region's n_ctx, not the allocator's arena max when a per-layer n_ctx is smaller; and the FusedQKV/FusedQkvNorm arms share the helper but have no dispatch gate of their own. Record: the plan's #38 entry.
    • Decode-fusion shape guards (#39, gap-table G2). The three decode-only fused arms — FusedFFN, FusedQKV, FusedQkvNorm — refuse a non-decode shape (nt != 1) with an Err naming the node and the observed nt, and the check runs before the weight lookup (shape validation is weight-independent and cheaper, so a bad shape on a weightless node reports the shape, not the missing weight; this matches CUDA's FusedQKV/QkvBiasRopeStore order). Before #39 the arms asserted this with debug_assert!, which a release build compiles out — the asymmetry CUDA never had. Gates: metal_fused_ffn_refuses_nt_other_than_one, metal_fused_qkv_refuses_nt_other_than_one, metal_fused_qkv_norm_refuses_nt_other_than_one (src/graph/metal_backend/tests/fusion_shape.rs). Record: the plan's #39 entry.
    • Norm-weight guards (#40, gap-table G3). Op::RmsNorm / Op::QkNorm no longer fall through to the weightless rms_norm kernel when the gain cannot be resolved. Both None meanings — NormMeta::weight_name absent, or a set name the device never registered — are a missing gain, and the old path produced plausible-looking output from a wrong computation, the failure mode this section forbids. Both arms call MetalBackend::norm_weight, which returns Err naming the node and the missing tensor (the CUDA twin is CudaBackend::norm_weight); no supported producer (models/qwen2, models/qwen3) builds a weightless norm, so there is no legitimate path to preserve. Gates: metal_rms_norm_refuses_a_weight_not_on_gpu, metal_rms_norm_refuses_a_weightless_node, metal_qk_norm_refuses_a_weight_not_on_gpu (src/graph/metal_backend/tests/norm_weight.rs). Scope: the CPU arm keeps its None => rms_norm_f32 fall-through (src/graph/cpu_backend.rs) — untouched, since #40 is a Metal ticket — so CPU remains exposed to the same weightless computation. Record: the plan's #40 entry.
    • KV-store row count (#305). Op::KvcacheStore derives nt from the K input's logical length (BufRef::len), not self.pool[id].length(): the pool allocates at the E4 S2 size class, so the physical length over-counts nt whenever nkt * nt is not itself a class size, and the store would read cells past the filled prefix into the class's uninitialised tail and write those garbage rows into an arena other nodes read. The other self.pool[..].length() uses were audited and are harmless: Silu/Add/Mul/QkNorm over-process their own output padding, and FusedQKV/FusedQkvNorm slice to the rounded length but read only p[0] for the concat. Record: the plan's #305 entry.
  6. gpu_abort for configurations the GPU path cannot run — dimension misalignment, device-limit overruns, kernel-array overflow: print the actual values and exit.
  7. Recurrence playbook — reproduce with one app and a bounded -n; bisect with MINFER_GEMM=0 and MINFER_CACHE_TYPE=f32; on a freeze, spindump over SSH and check the diagnostic reports.

Accepted and documented (not fixed): the audit's L1/L2 landmines, and the fact that a kernel-level fault is not always attributable to a single dispatch (hence the MINFER_TRACE ring).

5. Implementation Phases

5.1 Graph-backend phases

Metal arrived as Phase 3 of the compute-graph rewrite and was then wired to parity with the pre-graph path through the G-series:

PhaseContentStatus
3MetalBackend per-op adapter (metal_backend.rs) + cross-backend scheduling✅
4–6scheduler assign/split/execute, FusionPass, DOT/cache; Qwen2 graph build; imperative forward.rs deleted✅
G1Attn dispatch mirrors the old path (flash / split / parallel / classic) with the same gates✅ 4e105ce
G2RmsNorm selects rms_norm_256 when enabled (~2× faster per dispatch)✅ 4e105ce
G3n_out tail-row GetRows + two allocator liveness fixes✅ d81af71 (docs 8d7cb38)
G4Decode Op::FusedQKV (concat matmul + attn_bias_rope_store)✅ bd28047 (docs 96404fb)
G5Decode Op::FusedFFN (concat gate+up + in-place swiglu), gated nf <= 16384✅ 1dee1b5 (docs ec922f1)
G6Qwen3 Op::QkNorm + Op::FusedQkvNorm (per-head norm + no-bias rope/store)✅ 283c7d6, 94d57ac, d5b8023
objc2 migrationmetal 0.28 / block + vendor/block patch → objc2-metal / block2 / objc2-foundation, Phases 0–6✅ 9e238bd, 6a382a3, be3df55, ee9b65b

The graph-era details live in docs/COMPUTE-GRAPH-DESIGN.md §17 (Phases 1–11, deviations 18–26) and METAL_OPTIMIZATIONS.md §0.1 / §4.3.

5.2 Optimization and cold-start record

Hash caveat. The commit hashes printed in docs/METAL_OPTIMIZATIONS.md predate a repository history rewrite and no longer resolve. The hashes below are subject-matched equivalents from the current history; each was verified with git cat-file -t. Cite these, and the METAL_OPTIMIZATIONS.md section, together.

WorkstreamWhat it changedDoc referenceRepresentative commits
Initial Metal portDevice/queue/encoder skeleton, the full kernel set, RMSNorm/GQA simdgroup parallelism, SwiGLU fusion§3.1, §5.1ac2cb3d, e7df395, 6811da3, 2a76d0c, b0819e7, 2f484f8
Correctness foundationRoPE freq_scale, output_b, softmax max, dynamic hd, Q5_K formula, GQA simd_max partial-tile divergence, first isolation suites§3.1 (#1–#5)df2da9f, c34b7d8, 87fec18
GPU-safety hardeningBounded submit() + status check, dispatch-label trace ring, no early return past a barrier, runtime guards, autorelease retain fix§3.1 (#6–#7), GPU_SAFETY.md §1–§4aef6cde7, 5f5a42d
Decode fusions + split attention (old path)Fused QKV/FFN decode matmuls, 2-pass KV-parallel split attention, float4 acc, adaptive chunks, KV geometric growth§3.3e5db3aa, 39c2ba9, eb0f812, e031282
GEMM (simdgroup) workllama kernel_mul_mm port (64×32 tile, 4 simdgroups), per-quant GEMMs, hot-loop unroll, ik-loop simdgroup_barrier, the partial-tile race + memoryBarrier fix§3.4 (#11/#12/#28/#29/#30)c34b7d8, d83bd25, 5202548, e7be3d3, 69a47d4, f52628c, e997b99
Decode matmul layout portsq6_K stride-2/float4 (72→209 GB/s), q4_K stride-4/sc16 (7B decode ~51→~19.3 ms/token)§3.3 (#27)36c9b01, b59c8c8
Flash-attention portsDecode flash_attn_ext_vec hd=64/hd=128, prefill flash_attn_blk hd=64/hd=128, tail-pad kernel§3.3/§3.4 (#22/#24–#26), §5.5c1177f7, b56b5dd, ac5e1ea, a78620c
Parallel prefill attention + RMSNorm-2563-pass barrier-free prefill attention; 256-thread RMSNorm; chunk cap 16; drop the KV→CPU sync§3.4/§3.5 (#16/#17)35a9659, 89aa1fa
KV f16store_kv_f16 + _f16 attention kernels; auto-select f16 for the 7B class§3.5 (#13/#37)0aad968, ff60ed7
n_out tail-row reductionFinal norm + lm_head on output rows; last-layer FFN + both residuals on the tail rows§3.7 (#32/#34)59a97f9, 660f59e
Cold-start axisEmbedded precompiled metallib, GGUF mmap + zero-copy weights, load-time warm-up read§4.2b7cb5bc, 1032ca9, a5709e7
Embedding coverage + GEMM thresholdGPU get_rows for all 8 quant types; adaptive GEMM dispatch `nt≥2 && (od≥2048nt≥9)`
Prefill-GEMM investigationGrid-shape probe, exact-shape replay, 7B decomposition, structural-equivalence audit, gap acceptance§3.6 (§4.3.1–§4.3.10)6205f7e, 9a201b5, c900127, 89e2468
Compute-graph IR + allocatorPhases 1–6 of the rewrite (IR/builder/allocator, CPU backend, Metal backend, scheduler/fusion/cache, Qwen2 graph build)COMPUTE-GRAPH-DESIGN.md §17.1a163a07, 308cc74, be35b1b, 8091b61, 941d34f, e54070c

Numbers like #27 are local to a METAL_OPTIMIZATIONS.md §0 table — always qualify them, because the Done / To-do / Decided tables reuse the same numbers.


6. Risks and Open Questions

#Risk / questionStatus
1Cross-dispatch write visibility (2026-08-19 incident)Fixed: scope barrier after every dispatch; rule recorded in GPU_SAFETY.md
2Early return past a threadgroup_barrier in attentionFixed: no early returns; invalid heads run the loop and skip the store
3Device-limit guessesFixed: limits queried at init, never hardcoded
4Concurrent MPS access under parallel tests flips fused/unfused greedy tokensMitigated for tests: metal_test_lock() + MpsState::init(); production runs one scheduler thread
5Q8_0 multi-token matmul race (missing trailing threadgroup_barrier)Fixed; pinned by metal_prefill_determinism
6Metal gate failure is silent (no diagnostic like CUDA's CUDA GATE:)Open (diagnostics): a missing/unregistered weight drops the model to CPU with only the init line to explain it
7Whole-layer layer_gpu reference path still referenced in comments/testsOpen (cleanup): the function is gone from src/metal/; the #[allow(dead_code)] block and comments remain
87B prefill ~10% behind the old pathAccepted: GEMM-bound; the GEMM transfers fully, and the prefill-GEMM investigation closed as "decided not to change" (METAL_OPTIMIZATIONS.md §3.6)
9Warm-up / cold-start costsTracked: mmap first-touch solved by the load-time warm-up; remaining cold-start to-dos in METAL_OPTIMIZATIONS.md §4.2

7. Verification

7.1 Device test suite

src/graph/metal_backend.rs carries 14 graph-backend tests, all serialized by metal_test_lock() and skipped with MPS unavailable; skipping when Metal is absent:

  • Elementwise / norm: metal_elementwise_matches_cpu (silu+add, bit-for-bit), metal_rmsnorm_matches_cpu, metal_rmsnorm_real_scale (d=896, nt=8).
  • Cross-backend: metal_cross_backend_copy, metal_cross_backend_copy_large, metal_embed_then_rmsnorm_cross_backend, metal_multi_split_alternation.
  • Matmul: metal_matmul_q8_matches_cpu (real wk shape), metal_matmul_q4_matches_reference (layer-0 wq, [896, 896], vs a manual Q4_0×f32 reference).
  • KV + attention: metal_attn_kv_matches_cpu (bit-exact), metal_attn_kv_real_scale (nh=14, nk=2, hd=64, nkt=128, nt=30), metal_attn_decode_step (nt=1 with 30 stored rows), metal_store_after_gpu_op (KV store whose K comes from a GPU op), metal_store_real_dims (n_ctx=32768).

Model-level Metal tests live with the models and skip on non-macOS: graph_metal_matches_cpu_logits, graph_metal_layer0_isolation, graph_metal_real_wk_matmul (Qwen2), fused_qkv_matches_unfused_decode (Qwen2, with the metal_test_lock rationale recorded), fused_qkv_norm_matches_unfused_decode (Qwen3), graph_metal_matches_llama_reference (Qwen3; pins the first 9 tokens of a 60-token byte-identical llama-Metal run) and metal_prefill_determinism (Qwen3; the Q8_0 multi-token matmul race regression).

7.2 Kernel isolation suites

tests/ carries four macOS-only integration binaries (#![cfg(target_os = "macos")], so they do not exist on Linux). Each builds its own deterministic inputs and embeds a scalar CPU reference — none needs an external dump:

FileTestsCoverage
tests/flash_attn_isolation.rsflash_attn_ext_isolation (hd 64 + hd 128), flash_attn_matches_splitdecode flash vs a scalar online-softmax reference and vs the split path; partial/empty KV chunks, nt 1–2, nkv up to 4097
tests/flash_attn_blk_isolation.rsflash_attn_blk_isolationprefill blk port (hd 64/128, NSG=4), partial last KV block via the tail-pad kernel, nt up to 200, GQA heads, f32 + f16 KV, vs classic
tests/gqa_attn_isolation.rsgqa_attn_isolation, gqa_attn_split_isolation, gqa_attn_split_timingclassic + split attention vs a scalar reference, including the nkv % 32 != 0 divergent-simd_max case
tests/gemm_isolation.rsgemm_isolation, qkv_row_concat_layout, non_q4_0_gemm_isolation, get_rows_q4_k_isolation, get_rows_multi_type_isolationGEMM determinism + correctness vs scalar, concat layout, per-type row-gather

7.3 Acceptance gates

  • Bit-exactness where the math is identical: elementwise and KV/attention round trips are checked bit-for-bit; matmul and norm are checked within float tolerance.
  • CPU-vs-Metal logits: #324 restored the model-level graph_metal_matches_cpu_logits. The pre-#324 body compared ModelDef::forward with ModelDef::forward_graph, but in the graph era both route through Qwen2Graph::forward under the same device decision, so its max |Δ| = 0 compared the Metal graph with itself — a vacuous gate. It now loads the same cached 0.5B Q4_0 twice — a Layers(0) CPU engine and an explicit full-plan Metal engine — drives the same greedy continuation through forward_graph_cached, asserts the two built graphs genuinely differ in backend assignment (metal arm 248 METAL / 0 CPU nodes; cpu arm 0 METAL / 440 CPU nodes), pins the greedy continuation to a literal, and compares the final-step logits. Bar named before measuring: 1.0 absolute (the q8_0 weight-quantisation class of docs/GGUF-TOOLING.md, 3× the observed) and 5e-2 relative. Measured on macbook (macOS 27.0.1, Apple M4 Pro), 2026-10-07, cargo test --release --bin minfer -- --nocapture graph_metal_matches_cpu_logits: max |Δlogit| 0.347 absolute / 1.57e-2 relative against max |logit| 22.04, greedy [12095, 11, 323, 432] identical. The class is quantization, not accumulation order: the CPU quantizes activations to Q8_0 while Metal reads f32 (rule 9), which is why the bar is looser than the f16 #164 gate's 0.05. The row is still backed more widely by the per-op metal_*_matches_cpu gates and the external oracle graph_metal_matches_llama_reference.
  • Fused-vs-unfused decode must be bit-identical, with the unfused side running the FusionPass.
  • llama-reference oracle: Qwen3's first 9 greedy tokens are pinned against llama-Metal.
  • Determinism: metal_prefill_determinism and the isolation suites' repeat-run checks.

Suite baselines: METAL_OPTIMIZATIONS.md / KNOWN-CPU-ISSUES-2026-08-29.md record the single-threaded and parallel runs (e.g. 152 passed / 3 ignored single-threaded main-bin, plus the integration binaries green on device). Run cargo test --release on macOS for the device suites; on Linux the Metal-only tests self-skip and the isolation binaries are empty.

7.4 The macOS suite baseline (#298, recorded by #54 at the round's end)

The current machine-checked baseline is not on this page. It is docs/TEST-BASELINES.md ("macOS unit", box macbook (macOS 27.0.1, Apple M4 Pro), 2026-10-08) backed by the macos-unit row of scripts/test-baselines.toml, together with that page's macOS real-model row. Edit that ledger, never a counter. A later Mac gate run diffs against it: anything new is a regression.

The round-era record is in the plan. The round's final master was 6b95763; the round's start was 97823e4, where the same suite had 21 failures, and the enumeration that followed — the three root causes they decomposed into, the two order-dependent flakes, and the eleven-day build-macos blind spot that hid the macOS test target from cdf41b2 to 4add59f — is the plan's record: G7's row (#54) for the baseline, the #298/#231 suite enumeration for the failures, and the #303 record for the blind spot.

The purpose is unchanged: a Mac gate run compares against a dated baseline, so a green suite cannot hide a new failure. Only the baseline's home moved.


8. Out of Scope / Future

  • Not planned: MPSGraph / higher-level MPS APIs, multi-GPU, f16 activations, training. (f16 weights landed on Metal in #164 and bf16 in #208 — see §4.4 and docs/SUPPORT-MATRIX.md; it is the activation dtype that stays f32.)
  • Closed by measurement (METAL_OPTIMIZATIONS.md §3.6/§4): the prefill GEMM gap (params match llama's; decided not to change), and the flash/split/parallel attention lineup.
  • Remaining research (METAL_OPTIMIZATIONS.md §4.1/§4.2): cold-start items and the residual 7B-prefill gap; see that document's roadmap section rather than duplicating it here.
  • Cleanup candidates: the #[allow(dead_code)] legacy block/comments in src/metal/, and adding a Metal gate diagnostic to match CUDA's CUDA GATE: line (§6 #6/#7).

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0001 — Inference runs through one declarative compute graph
  • ADR-0002 — Topology is a function of GraphParams alone, so positions cannot be structure
  • ADR-0005 — Metal becomes a first-class backend
  • ADR-0006 — The KV storage format is a per-engine gate, not a process-wide global
  • ADR-0008 — GPU safety: bounded waits, no early return past a barrier, runtime device limits
  • ADR-0009 — A failure is an error, never a silent fallback
  • ADR-0010 — The identity gate: bitwise by default, a named tolerance class otherwise
  • ADR-0013 — The CPU quantizes activations to Q8_0; a device reads f32
  • ADR-0014 — A KV session is a versioned, checksummed file — never a memory dump
  • ADR-0021 — bf16 is a round-to-nearest-even cast, and 1-D tensors stay f32
  • ADR-0012 — Device is the first axis, the layer the second — and no premature common
  • ADR-0015 — The offload auto fit takes a prefix, not a knapsack
  • ADR-0016 — A failed device-memory query is not a zero budget
  • ADR-0032 — The packed-KV staging window is f32, not f16
  • ADR-0036 — bf16 weights get their own device kernels, not a dtype flag on the f16 ones
  • ADR-0037 — A Metal weight dtype a kernel cannot consume is refused, never run as a wrong kernel

minfer CUDA Backend Design

How minfer runs the compute graph on an NVIDIA GPU: the CudaBackend graph executor, the cuda.rs device layer, the kernel families it dispatches, the CUDA Graph capture/replay machine, and the memory and safety rules that hold it together.

Status. Landed. Phase 7 (7a–7e) shipped the backend, and the later campaigns (Phase 8, R/P5, the MMQ line, the decode campaign D1–D4, D5-R speculative decoding) built on it. Every mechanism described here is implemented in the tree. Baseline: HEAD = 916a7d2 (2026-09-14).

Provenance. This file was docs/CUDA-BACKEND-PLAN.md, written before Phase 7 as a plan. The section skeleton is preserved; the body has been rewritten in the present tense against the current code, and the plan-time "current state" / "placeholder" / v1-matrix text has been replaced by the landed design. Plan-era commit hashes that no longer resolve are noted where cited.

Related records. This document is the backend design (what it is and how it fits the graph). The optimization campaign — every measured lever, accepted or reverted — lives in docs/CUDA_OPTIMIZATION.md (live state + §0 history table + Appendix A env-gate reference) and its expansion layer docs/cuda_optimization_steps/ (one document per step). docs/CUDA-TECH-PRIMER.md explains the GPU techniques, docs/GPU_SAFETY.md holds the hard safety rules (CUDA section), and docs/inference_e2e_walkthrough/15-cuda-backend.md narrates the backend for a first-time reader. The graph contract this backend implements is docs/COMPUTE-GRAPH-DESIGN.md §3.5/§9.


1. Design Goals and Outcome

1.1 Goal

Implement CudaBackend (src/graph/cuda_backend.rs) as the third backend of the compute graph by wrapping the existing src/cuda.rs device layer — not stubbing it — so that on a CUDA machine the whole per-layer chain runs on the GPU through the standard build → assign → fuse → alloc → execute pipeline, with the same correctness contract as CPU and Metal: backend placement is decided at build time, kernel-invariant violations return Err, and there is never a silent mid-run fallback.

1.2 Outcome

GoalLanded outcomeEvidence
CudaBackend wraps cuda.rs in the Backend traitcuda_backend.rs implements the full trait (pool, name→ptr weights, per-op dispatch, host transfers, graph replay); the device layer keeps the registry, streams and kernel launchers§2, §4.1–§4.2
Per-op dispatch, not a whole-layer callEvery Op the model builders emit has a CUDA arm; the legacy cuda.rs::layer_gpu is no longer driven§4.4
Per-node placement decided at build timesupports_op + the model-level weights_on_cuda all-or-nothing gate feed CParams.gpu; the scheduler splits the graph§4.3, §4.6
Weights resident from load, no execution-time host copiesThe loaders register every graph-referenced tensor by name at model load; execute_node resolves pointers from the registry§2.3, §4.6
CUDA Graph capture/replayDecode splits capture after a 2-run warmup and replay as one launch; repeated identical-nt prefills capture too (default on since R3-B); keyed by (uid, node range) with pool-generation invalidation§4.8
Same correctness gates as MetalPer-op parity tests, per-quant matmul parity, capture bit-parity, greedy-text equality vs CPU; CPU-vs-GPU logits differ by design (f32 activations vs Q8_0) and are compared with tolerance classes§7
PerformancePost-campaign default: 7B Q4_K_M whole-prefill ~3581 tok/s (1.080× llama.cpp same-window) and decode 51.2 tok/s tg128 (1.074×); the full ledger is CUDA_OPTIMIZATION.md §1§5, CUDA_OPTIMIZATION.md §1.1

1.3 Non-goals

  • Multi-GPU / tensor split (llama.cpp's split machinery) — single-GPU engine.
  • FP16/BF16 activations and cuBLAS/cublasLt. Prefill uses f16 weight tiles and int8 tensor-core MMQ, and the KV cache can be f16, but activations stay f32 and the matmul path is minfer's own.
  • IQ*/Q2_K/Q3_K kernels — not implemented and not planned without a model that needs them.
  • Windows and a self-hosted CUDA CI runner (deferred; device-gated tests skip gracefully).
  • graph_optimize-style node reordering — the graph topology is the builder's decision.
TopicWhere
Optimization history, current state, env-gate referencedocs/CUDA_OPTIMIZATION.md
Per-step records (every lever, accepted/reverted/closed)docs/cuda_optimization_steps/
GPU technique primerdocs/CUDA-TECH-PRIMER.md
Safety rules (capture windows, D2H staging, weight ownership)docs/GPU_SAFETY.md (CUDA section)
Graph contract (IR, allocator, scheduler, backend trait)docs/COMPUTE-GRAPH-DESIGN.md
Beginner narrative of this backenddocs/inference_e2e_walkthrough/15-cuda-backend.md
MMQ analysis, speculative decodingdocs/LLAMA-CPP-MMQ-ANALYSIS.md, docs/SPECULATIVE-DECODING-PLAN.md

2. Architecture at a Glance

2.1 Three layers

LayerFileRole
Graph executorsrc/graph/cuda_backend.rsImplements Backend: device buffer pool, per-op dispatch, positions conversion, CUDA Graph state machine, trace staging, error contract
Device layersrc/cuda.rsCudaState context: device probe, weight registry, the per-instance stream binding (bind_stream/stream/create_stream), extern "C" kernel launchers, CUDA Graph API, pinned/staging memory, per-stream activation scratches, MMQ caches and gate reads
Device tier tablesrc/device_tier.rscc-keyed tier rows (measured GB10 + llama.cpp-adopted consumer rows + GENERIC) resolved once at init; feeds the MMQ gate, smem feasibility and plane-VRAM budget checks. Design + status: DEVICE-ADAPTATION-PLAN.md, docs 105–106
Kernelssrc/cuda/kernels/*.cu + common.cuhThe __global__ kernels (quantized matmul families, attention, norms, elementwise, KV store, embedding gather, quantize planes) — 17 nvcc translation units since #263
Build chainbuild.rsOpt-in --features cuda, nvcc/-ccbin probe, per-arch SASS/PTX (incl. native sm_121), cudart link + rpath

The split is deliberate: cuda.rs is the only place that touches the CUDA runtime API, and cuda_backend.rs is the only place that knows about graph nodes.

2.2 The CudaBackend surface

#![allow(unused)]
fn main() {
pub struct CudaBackend {
    state: &'static crate::cuda::CudaState,   // process-wide device context
    stream: *mut c_void,                      // THIS instance's non-blocking stream (#188)
    kv_layout: i32,                           // KV layout tag for this instance
    pool: Vec<CudaBuf>,                       // id -> { ptr, bytes }
    free: Vec<usize>,                         // byte-length-matched free list
    pool_gen: u64,                            // bumped on every pool allocation
    pos_scratch: *mut c_void,                 // device i32 positions plane
    pos_scratch_bytes: usize,
    pos_memo: Option<(usize, u64)>,           // one-execution-window conversion memo
    graph_execs: Vec<CapturedGraph>,          // instantiated graphs
    graph_runs: HashMap<(u64, (usize, usize)), u32>, // warmup counters
    capturing: Option<(u64, (usize, usize))>, // open capture window (on `stream`)
    graphs_mode: GraphMode,                   // Enabled / Disabled
    prefill_capture: bool,                    // default ON (MINFER_NO_PREFILL_CAPTURE opts out)
    cap: crate::cuda::CaptureStaging,         // viz/trace async D2H staging
}
}

new() returns None when CudaState::get() fails (no device, or MINFER_DISABLE_CUDA=1, both handled inside the device layer), which is how GraphAllocator::enable_cuda declines. The struct is unsafe impl Send/Sync: raw device pointers are dereferenced only by the GPU, and mutation happens only through &mut self (mirroring MetalBackend).

2.3 Weight residency and registry

Every graph-referenced tensor is uploaded once at model load and referenced by name afterwards:

  • CudaState::register_weight(name, data) allocates and H2D-copies; the same name and size is a no-op (reloading the same GGUF must not leak a second device model), while a different size replaces the entry and deliberately leaks the stale buffer, because a live captured graph may still reference it. The leak is bounded by the number of distinct (architecture, tensor) shapes.
  • has_weight_of_size(name, raw_len) is the size-aware gate used by the model wiring; it matches repacked weights by their original raw length, so a foreign-architecture entry reads as "not registered" and that model stays on CPU.
  • get_weight_ptr(name) is the only lookup the executor uses.
  • Quant-specific registrations: Q6_K uses a 224-byte-padded repack (register_weight_q6k_padded, enabling 16-byte uint4 loads); Q8_0 optionally registers a p32 split plane (register_weight_q80_p32); Q6_K/Q4_K optionally build pre-expanded / pre-decoded planes (register_weight_q6k_exp, register_weight_q6k_dsc, register_weight_q4k_dsc, and the D4-4 {name}__dpl dense split plane). Plane maps are keyed by the weight's device pointer and looked up inside prefill_mmq.
  • The q4_K W_dsc plane is admitted for q4_K and only q4_K (src/q4k_dsc.rs, q4k_dsc_plane_admitted — issue #165). The rule has two halves: the type gate (TensorType::Q4_K is the one type mmq_raw_nb_bt dispatches the dsc template for, so the loader admits no other type) and the payload gate (raw.len() must be exactly od * (id / 256) * 144, q4_K's own block layout — equality, not a lower bound). register_weight_q4k_dsc re-checks the payload before the budget query and before expand_q4k_dsc, and expand_q4k_dsc itself returns None for a payload it cannot index, so a direct caller cannot bypass either. Why both: a q4_0 payload has exactly q4_K's bytes/element ratio (18/32 == 144/256), so the size check cannot refuse it; a q8_0 payload (34/32) is longer and would be misread as 144-byte q4_K super-blocks; and a future type with a smaller ratio (a 2-bit K-quant: 84/256) is shorter than the row arithmetic needs, so the size check is what refuses it instead of reading past the tensor. The check cannot tell a q4_K payload from another type's bytes of the same length — that is the type gate's job.
  • The per-tensor registration dispatch is one shared rule (src/models/weight_reg.rs, issue #167). Both loaders call register_cuda_weight for every tensor the E5 plan puts on the device; it carries the whole contract: the quantized-type matches! set, the F16 raw branch, the F32 (1-D norms/biases vs 2-D matmul weights) branch, the Q6_K padded repack, the q8_0 p32 plane, the q4_K W_dsc plane under q4k_dsc_plane_admitted, and the clear_mmq_nb_bt_only rule. The decision (cuda_weight_reg) is pure — no CudaState, no environment; the r59 dispatch gates are passed in — so CI's CPU job runs its tests, exactly like src/q4k_dsc.rs. Both loaders previously carried a copy of this block and the copies had drifted twice: the qwen3 copy had neither the f16 branch (#141) nor the q4_K dsc call (r59/#165), so an f16 Qwen3 fell to the CPU and a q4_K Qwen3 kept the in-kernel scalar dsc decode. The graph-side type gate is per architecture and must list the same types (Qwen3Graph::weights_on_cuda gained F16 in #167).
  • A per-weight f16 dequant cache (w16_cache) is enabled by the loader only when quantized matmul weights exceed 2 GiB and MMQ is off; MINFER_NO_W16CACHE=1 reverts.
  • ModelLoadGuard (reentrant, process-wide) serializes loader registration so two models with same-named tensors cannot interleave, and real-model tests hold it across their forwards.

The graph-side CPU registry is separate: register_graph_weights registers into the allocator's CPU backend, which is what a CPU split of a mixed graph executes from; the CUDA registry is filled by the loaders at model-load time.

2.4 Streams, capture windows and the per-instance discipline

Everything the device path issues — kernels, cudaMemcpyAsync staging, events, capture/replay, synchronize — runs on a stream owned by the CudaBackend instance, not on a process-wide one. That is the #188 change:

  • CudaState stays the process-wide context: the device, the name-keyed weight registry, the derived weight planes (q6k_exp/q6k_dsc/q4k_dsc/w16_cache) and the device_memory/host_alloc queries are genuinely context-scoped and shared.

  • CudaState::create_stream() returns a fresh cudaStreamNonBlocking stream. CudaBackend::with_layout creates one per instance (no stream, no backend) and Drop destroys it after the pool, the scratches and the captured graphs.

  • Every CudaBackend device operation starts with let _bound = self.bind(). bind_stream publishes the instance's stream in a thread-local; CudaState::stream() answers with it, so the ~60 launch/copy/event helpers in cuda.rs keep their signatures and still follow the instance. The stream consumers, named:

    ConsumerHow it gets its stream
    kernel launchers (cuda.rs, the MMQ/attention/fused families)self.stream() → the bound instance stream
    H2D input fill (write_input_async), D2D (copy_device_to_device), D2H staging (copy_to_host_async)self.stream(); the pinned staging ring is keyed on the stream too
    events (record_event, stream_wait_event) and the F5 copy_cross/await_cross hooksself.stream() via enqueue_cross_host/take_cross, both of which bind
    cudaGraphLaunch in graph_replay_step, graph_begin_capture, graph_end_capture_to_execself.stream() while the backend's own window is open
    synchronize / state_syncself.stream() — one stream sync, counted per backend
    copy_cells (kv_move_rows) and alloc_buffer/free_bufferthe bound stream (the cudaFree/cudaMalloc themselves are context-wide)
    Drop for CudaBackendbinds, then frees pool/scratch/host/graph, then destroys its stream
    weight registration (CudaState::register_weight)the context stream, never a backend's: an H2D copy queued on the context stream + cudaStreamSynchronize(context), so it is stream-ordered and cannot be recorded into anybody's capture window
    context-wide, stays sharedcudaMalloc/cudaFree, cudaMemGetInfo (device_memory), cudaHostAlloc/cudaFreeHost (host_alloc/host_free), the weight registry and the derived planes, cudaGetLastError
  • Activation scratch is per stream. The buf_hidden/buf_q8_prefill/buf_qa8_t/ buf_q8_decode/buf_attn_partial/… slots are now StreamScratch, a map keyed on the current stream, and the MmqCache memo is keyed the same way (a hit records a scratch pointer, so a shared memo would hand one engine's plane to another). The staging ring (write_input_async) is keyed on the stream for the same reason. Everything unbound — the legacy layer path, direct CudaState tests — keys on the context stream, which is why the #185 guard is narrowed to that path (§2.5).

Capture mode. Windows are opened with cudaStreamCaptureModeThreadLocal (graph_begin_capture), not cudaStreamCaptureModeGlobal. Under Global another thread's capture-unsafe driver call belongs to the window: it either invalidates the capture (cudaErrorStreamCaptureInvalidated, 901) or faults inside the driver. The recorded SIGSEGV (#185) was exactly that — one thread at cuMemcpyHtoD_v2 under register_weight while another was at cuGraphInstantiateWithFlags under graph_end_capture_to_exec. Thread-local mode scopes invalidation to the capturing thread, so a registration on another thread is benign. MINFER_CUDA_CAPTURE_MODE=0|1|2 overrides the mode (relaxed / global / thread-local) — it exists for the #188 probe's measurement, not for production.

The #188 probe is the instrument. graph::cuda_backend::tests:: capture_window_on_one_thread_survives_a_weight_registration_on_another opens a capture window on one thread and issues a weight-registration copy on another while the window is open, then closes it and checks both that cudaStreamEndCapture returned 0 (never 901) and that the replayed graph produced the bytes it recorded. It has two env knobs so the mode can be judged rather than assumed: MINFER_PROBE_STREAM=context captures on the context (blocking) stream — the pre-#188 shared-stream model — and MINFER_PROBE_LEGACY_MEMCPY=1 issues the registration with the pre-#188 blocking cudaMemcpy. Measured on GB10 sm_121, 5 process runs per cell (90 s watchdog; crash = the probe's own assertion failed), 2026-09-27:

capture streamregistrationmoderesult
context (blocking), sharedblocking cudaMemcpy (pre-#188)global (1, pre-#188)5/5 hang
context (blocking), sharedblocking cudaMemcpythread-local (2)5/5 hang
context (blocking), sharedblocking cudaMemcpyrelaxed (0)5/5 hang
instance (non-blocking)stream-ordered (this PR)global (1)5/5 pass
instance (non-blocking)stream-orderedthread-local (2, adopted)5/5 pass
instance (non-blocking)stream-orderedrelaxed (0)5/5 end_code=901

Three measured readings, none assumed:

  1. The blocking copy is a hard deadlock, independent of the mode. A blocking cudaMemcpy is issued on the legacy null stream, which implicitly synchronizes with every blocking stream — including the one holding the open capture window, which by construction cannot complete until the host closes it. All three modes hang 5/5. So the mode is not the fix for the historical setup; a stream-ordered copy on a non-blocking instance stream is.
  2. Relaxed is ruled out by direct measurement. With the structural fix in place, relaxed still returns cudaErrorStreamCaptureInvalidated (901) — the exact code the acceptance forbids — in 5/5 runs (cudaMalloc inside the window also fails, CUDA: failed to allocate 16384 bytes). It is not adopted.
  3. Thread-local is the mode. With it, the probe passes 5/5 in both the pre-#188 shared-stream cell (the deadlock aside) and the instance cell; global also passes the instance cell but is the mode that lets a foreign thread's driver call belong to the capture, which is the class this ticket exists to remove.

Two readings: the mode is what makes the historical shared-stream setup safe (global is not viable; thread-local and relaxed both are, and thread-local keeps the capturing thread's own mistakes fatal, so it is the one adopted), and the structural change makes the mode irrelevant by removing the sharing. The concurrent device gate (models::qwen2::graph::tests::cuda_kv::two_cuda_engines_forward_concurrently_and_stay_bitwise_identical) is the positive half: two engines on two threads, two distinct streams, bitwise equal to their serial references.

The prefill-GEMM smem opt-in (#145, #147; lazy and gated since #218). A gemm_f16_nt_kernel_t instantiation whose dynamic shared memory exceeds the 48 KiB default must be opted in with cudaFuncSetAttribute(.., cudaFuncAttributeMaxDynamicSharedMemorySize, N) — a launch over an un-opted-in dynamic smem is rejected with cudaErrorInvalidValue and cannot succeed. The number requested is the kernel's own byte layout — `As 2TNKS halves + Am 2TNKS floats (AF32 only)

  • Bs 2TMKS halves + Cs NW*256 floats, TN = 64, NW = blockDim.x/32= 8 — andgemm_dynamic_smem_bytes(tm, ks, af32) is the single source the launcher (launch_gemm_f16) reads; before #145 an eager sweep carried a stale copy of it while the launcher carried a copy that dropped the AF32 mirror. A request that exceeds the device's own cudaDevAttrMaxSharedMemoryPerBlockOptinis **skipped with the reason printed**, instead of called: the call could only returncudaErrorInvalidValue(whichcompute-sanitizercounts) and the instantiation cannot launch on that device at all. On GB10/sm_121 (limit 101376 B) that is exactly one combination,gemm_f16_nt_kernel_t<256,64,true>` at 122880 B.

The design is eager pre-warm at context creation + lazy per-launch opt-in (#223 restored the eager half). #188 deleted gemm_prefill_smem_init's CudaState::try_new call site with no mention in its commit message or its docs commit; two later "dead code hygiene" commits annotated the orphan #[cfg_attr(not(test), allow(dead_code))] instead of asking why a production-looking init had no production caller; #218 removed the function and its checked/skipped introspection, leaving the invariant tested but not enforced by production. #223 put the runtime guarantee back at the same site: CudaState::try_new calls the production entry gemm_prefill_smem_prewarm_one(tm, ks, af32) once per process for every launchable combination (the MINFER_GEMM_OPTIN_SET X-macro in src/cuda/kernels/common.cuh, shared with the fatbin lookup and the test seam). The placement is the argument: try_new runs under CUDA.get_or_init, before the state is published, before any CudaBackend exists, and therefore before the per-instance stream graph_begin_capture needs — so "the attribute is set outside any capture window" holds by construction, not by inference from the warmup count or the capture mode. It is once per process, not once per backend. The entry drives the same gemm_smem_optin<TM,KS,AF32> the launcher reads, so the pre-warm and the lazy path share one per-instantiation cache and one cudaFuncSetAttribute site; a cache-keying regression therefore cannot hide behind the pre-warm (it would leave the pre-warmed instantiations un-opted-in, which the gates read back from the device). A successful pre-warm prints nothing; a failure or a deliberate over-limit skip is named per instantiation, and MINFER_NO_GEMM_PREWARM=1 is the documented control that skips the loop (same-binary A/B and the "lazy path alone" gate arms).

The lazy per-launch opt-in stays as defence in depth: gemm_smem_optin (called from the launcher's GEMM_ONE) invokes the shared minfer_smem_optin helper on an instantiation's first launch and caches the answer in a function-local static per instantiation (a test injection is never cached). The helper names the site, the instantiation, the attribute, the requested bytes, the queried device limit and cudaGetErrorName, clears the latch, and the launcher does not launch when the answer is false; its own <<<>>> error is read too (minfer_launch_ok) and returned as 0, which prefill_gemm_f16_inner turns into an Err. The af32 wrapper launch_gemm_f32a returns the same result. A request above cudaDevAttrMaxSharedMemoryPerBlockOptin is skipped without calling the attribute on both paths, with the reason named. The same treatment covers every MMQ launcher (launch_mmq_nt/launch_mmq_raw_nt/launch_mmq_raw_nb_nt/launch_mmq_raw_nb_bt_nt/ launch_mmq_raw_nb_bt_q6k_nt/launch_mmq_raw_wide_nt); the terminal two return 0 → Err at the Rust caller, the fallback ones keep their documented 0 = clean fallback contract. The operator's signal stays the first-launch site report: a refused or skipped instantiation prints once, at the launch site (minfer_smem_optin). The removed checked/skipped counters did not come back, and a fully admitted pre-warm is silent.

Why the invariant holds — the pre-warm by construction, then three defence-in-depth mechanisms. The historical claim was that cudaFuncSetAttribute is illegal inside a capture window and poisons the context (error 700). Nothing in the repository establishes whether that holds for the adopted mode (see the honest limit below), so the design does not rely on the call being legal in a window; it relies on the call never happening in one. The eager pre-warm makes that true by construction (above). The three emergent mechanisms remain, now as the lazy path's fallback — called out at their sites (graph_replay_step in src/graph/cuda_backend.rs, gemm_smem_optin in src/cuda/kernels/gemm_wmma.cu); any change to one is a design change, not a tuning knob:

  1. The 3-run capture warmup (capture_warmup, default 3): capture opens only from the third run of a (uid, range) key, and prefill-shaped graphs (nt > 1) capture by default since R3-B. An instantiation's first launch — the one that calls cudaFuncSetAttribute — therefore always runs uncaptured.
  2. cudaStreamCaptureModeThreadLocal (#188's measured choice): the window belongs to the capturing thread, so a foreign thread's driver call cannot join it or invalidate it. Changing the mode must not silently change (1)'s guarantee.
  3. The per-instantiation cache: the in-window launch re-reads the cached answer instead of asking the driver again, so the >48 KiB launch inside the window never calls the attribute. The cache is instantiated on the (tm, ks, af32) template parameters, not on a deduced K: every gemm_f16_nt_kernel_t shares one signature, and the pre-#218 template <typename K> gave the whole family one static (the #218 coverage gate found <64,64,true> answering for <128,64,false>, whose attribute had never been set). #223's pre-warm drives this same function, so the cache is exercised for every instantiation at context creation — the regression can no longer hide behind a sweep that bypassed the cache.

The #218 gates pin that as observed behaviour. cuda_prefill_smem_optin_is_done_by_production (a real prefill forward in a fresh process; asserts the device's own opted_in read-back), cuda_prefill_smem_optin_refusal_fails_the_prefill (the control arm: MINFER_TEST_CALL_FAIL=attr:gemm_f16_f16 makes the production prefill refuse the launch and name the site), cuda_prefill_smem_optin_is_never_set_inside_a_capture_window (a >48 KiB prefill captures, replays bitwise, and gemm_smem_optin_in_capture_count() == 0), and cuda_prefill_smem_lazy_optin_admits_every_launchable_instantiation (every launchable >48 KiB instantiation reads back opted in through the production function). The capture_warmup test seam (MINFER_TEST_CAPTURE_WARMUP=1) is the mutation lever for the counter.

#223 adds the runtime guarantee's own gate. issue223_tests::cuda_prefill_smem_prewarm_opts_in_every_launchable_instantiation_before_any_launch runs in a fresh process and asserts, immediately after CudaState::init() and before any kernel launch, that every launchable >48 KiB instantiation already reads back opted in; a second fresh process with MINFER_NO_GEMM_PREWARM=1 asserts the negation, so the read-back is not always 1. It is the detector for the mutation the #218 gates cannot see — a pre-warm that skips one (tm, ks, af32) (the lazy path simply opts it in on first launch, so the coverage/counter arms stay green). The four #218 arms run their fresh-process children with MINFER_NO_GEMM_PREWARM=1, the documented control, so their claims stay the lazy path's and the pre-warm cannot make them vacuous.

Cost. The pre-warm loop's own duration is the fatbin's one-time module load — which the process pays before the first kernel from it can run either way — so with MINFER_OP_TIMING=1 the line reads ~2.3 ms warm, and the pre-split/synthetic 152.9 µs proxy was not the cost. The full measurement is in docs/SOURCE-LAYOUT-PLAN.md §5.1: the pre-split and post-split per-module transcripts (warm and page-cache-cold — the correction that the ~6× cold factor is page-cache, not the GPU clock), the 17-module split's effect on it, and the MINFER_OP_TIMING reading that only looks smaller because it now times one seventeenth of the work. The Step 5 record in docs/ARCHITECTURE-EXECUTION-PLAN.md holds the ticket-level entry.

The net effect is still the proxy's conclusion: the module load moves rather than appears — but only because something later would pay it anyway. prewarm_prefill() (the r59 rider's minfer_prewarm_kernels, called at the end of Qwen2/Qwen3 weight registration, before the first forward) already pushes that fatbin load into the startup path; with the pre-warm on it costs ~2.3 ms, with the pre-warm off ~4.5 ms — the same ~2.2 ms, moved earlier. The coupling is explicit and load-bearing: the ≈ 0 net holds only while a later step pays that same module load — today prewarm_prefill(), otherwise the first launch from the fatbin. Move that rider after the first launch, or remove it, and the pre-warm's loop becomes ~2.2 ms of net-new startup cost of the same clock-dependent magnitude. Controlled probe on the real binary (fresh process, MINFER_MMQ=0/1 × MINFER_GEMM_K64=0/1): the first prefill forward is 2288–2314 µs with the pre-warm off and 78–120 µs with it on; prewarm_prefill() is 4.3–4.6 ms off vs 2.2–2.4 ms on. The hot path is untouched: minfer bench -p 2048 -n 128, same binary, 7 interleaved matched rounds, medians — pp2048 2546.55 vs 2544.34 t/s (+0.09%), tg128 236.30 vs 236.45 t/s (−0.06%), bar ±1%. The full transcript is in the #223 record.

The same-thread ThreadLocal in-window case — measured (2026-09-29). The open question this section used to carry — is cudaFuncSetAttribute legal inside a same-thread cudaStreamCaptureModeThreadLocal window? — is now measured: the MINFER_TEST_CAPTURE_WARMUP=1 mutation arm drives the opt-in into the first, captured run, and on GB10 sm_121 / CUDA 13.0 / driver 580.178.04 the call is tolerated — the attribute publishes, the

48 KiB graph still captures, instantiates and replays bitwise-identically, and only gemm_smem_optin_in_capture_count() moves. (The 2026-09-25 probe below measured the Global mode instead; this one is the adopted mode.) The design nevertheless keeps the call out of the window: that behaviour is not contractual across toolkits, and the counter gate is what makes "never set inside a window" an observed property rather than a driver assumption. The full transcript is in the #218 record.

Measured correction (2026-09-25, CUDA 13.0 / driver 580.178.04 / sm_121). A probe (/tmp/fix147_attr_capture_probe.cu) shows the historical claim above no longer holds verbatim on this runtime: cudaFuncSetAttribute returns cudaSuccess when called inside an open cudaStreamCaptureModeGlobal window, both for the already-set value and for a new one. The eager sweep that was in tree then has since been removed by #218 (it was already dead code — #188 had dropped its caller), and because the launcher caches one answer per instantiation it does not re-ask inside a window either.

#223 forward note (2026-09-29): the eager half is back, as a pre-warm through the lazy entry, not as the old sweep: CudaState::try_new calls gemm_prefill_smem_prewarm_one for every launchable instantiation, i.e. the same gemm_smem_optin cache the launcher reads. Nothing in the measured driver behaviour above changes; the point of the placement is to stop depending on it (the attribute is set before any window can exist, by construction). On the real binary the loop's measured cost is ~2.2 ms (the fatbin's one-time module load), which prewarm_prefill() already paid during registration — net new startup cost ≈ 0, hot path unchanged. See §2.4 and the #223 record in ARCHITECTURE-EXECUTION-PLAN.md.

2.5 Legacy surface

The pre-graph imperative path still compiles but is not driven by inference: layer_gpu, output_norm_gpu, init_kv_cache + the kv_k/kv_v/kv_size slots, the buf_* persistent-slot pool with get_or_grow/upload_*, the host-staged quant_matmul_* wrappers, and the old single-slot capture API. Each carries a scoped #[allow(dead_code)] with a reason; the release build is warning-free. The graph path owns KV through the allocator's persistent regions, so init_kv_cache is bypassed entirely.


3. llama.cpp Reference Map

llama.cpp's CUDA backend (ca3d5a3e1 at the time of the port) was the reference for the device layer, the replay state machine and the kernel strategy. What was borrowed and what was deliberately not:

llama.cpp conceptminfer analogStatus
Backend interface (graph_compute, synchronize, async tensor set/get, supports_op)Backend trait: per-node execute_node instead of a whole graphBorrowed, reshaped
Weights placed in device buffers at load; ops follow their weights; never host-copied during executionregister_weight at load + name→ptr lookup in execute_nodeBorrowed
Device buffer pool, alloc/free hot pathCudaBackend pool with a byte-length free list; alloc_fresh for split stagingBorrowed
Enqueue during compute; synchronize only at scheduler boundariesexecute_node launches on the one stream; synchronize() at split boundariesBorrowed
CUDA Graph: warm up twice, capture the third run, replay keyed per graph, invalidate on pointer change(uid, node range) key, 2-run warmup, pool_gen invalidation, launch-once at closeBorrowed
Capture disqualifiers (host syncs, arch floor, env off-switch)No syncs/readbacks inside a window; MINFER_NO_CUDA_GRAPH=1Borrowed
Replay requires stable buffer addresses across stepsGraphCache keeps the allocator (and pool pointers) alive across decode stepsBorrowed
int8 tensor-core quantized GEMM (MMQ)The r-series MMQ line: pad40 transposed A planes, fused quantize producers, raw-byte NB-BT kernelsBorrowed in spirit, own kernels
Flash-attention tilingdecode split-KV attention, batched verify attention, FA-style tiled prefill attentionBorrowed in spirit, own kernels
Fusion pass patterns (ggml_cuda_try_fuse)minfer's FusionPass + build-time fused nodes; CUDA implements SwiGLU, FusedQKV, QkvBiasRopeStore, FusedFFNDiverged (graph-level fusion)
Multi-GPU split, NCCL, VMM pool, cuBLAS paths, graph_optimize reordering—Deliberately skipped
Abort-on-capture-failurelogs, disables graphs for the session, continues with direct launchesDiverged (chosen)

The one llama.cpp idea still on the table as a step function is the q8_1 GEMM-prologue fusion (see CUDA_OPTIMIZATION.md §1.4); everything else in the campaign is closed or sub-bar.


4. Design

4.1 CudaBackend lifecycle and state

with_layout() resolves the device singleton, creates the instance's own non-blocking stream (issue #188; a device that cannot give one gives no backend), snapshots the engine's KV element type (crate::cuda::layout_of), reads the two graph gates (MINFER_NO_CUDA_GRAPH=1 → GraphMode::Disabled; prefill_capture defaults ON unless MINFER_NO_PREFILL_CAPTURE=1), and starts with an empty pool. Drop binds the stream, frees every pool pointer, the positions scratch, the cross-backend staging and every captured graph exec, then destroys the stream.

State groups:

GroupFieldsLifecycle
Device + KV policystate, kv_f16fixed at construction
Poolpool, free, pool_gengrows on demand; free_buffer only recycles; alloc_fresh bypasses the list for split staging; pool_gen bumps on every allocation
Positionspos_scratch, pos_scratch_bytes, pos_memogrown on demand; the scratch pointer is embedded in captured execs, so growth bumps pool_gen to force re-capture
Capturegraph_execs, graph_runs, capturing, graphs_mode, prefill_capture (all on the instance's own stream)see §4.8
Tracecappinned async D2H staging, see §4.7

Pool rules worth restating because they carry correctness weight:

  • Exact byte-length reuse only — alloc_buffer scans free for pool[id].bytes == size * 4.
  • free_buffer never frees — a persistent KV region must survive rebuilds, and the pool keeps device memory for the next graph. Only Drop returns memory to the driver.
  • alloc_fresh exists for split-boundary staging: ids in the free list are still referenced by node_to_buf and physically live during the execute that follows.
  • OOM is not a panic. cuda_malloc logs and returns null; the null buffer fails cleanly at execute time (ptr_of). Panicking is forbidden because it would poison the shared scratch maps and the device-entry token (the legacy path) for every other user.

4.2 Backend trait mapping

Trait methodCUDA implementation
name()"cuda"
supports_op(op, dtype)§4.3 table
supports_fused(fused)matches!(fused, FusedOp::SwiGLU) — the only variant in the enum
alloc_buffer / free_buffer / alloc_fresh§4.1 rules
execute_node(node, in_bufs, out_buf, kv_pair)wraps execute_node_inner; on Err with an open capture window it calls abort_capture first (§4.8)
read_host(id)None — a staged D2H cannot return a borrowed &[f32] from &self; the allocator's copy_to_cpu arm calls copy_to_host instead
write_host(id, data)size-checked H2D; small inputs go through the pinned async ring (§4.7)
synchronize()clears the MMQ cache + positions memo, then close_capture_or_sync()
graph_replay(uid, range, nt_hint)§4.8 state machine

4.3 Eligibility

supports_op is f32-activation only (dtype != DType::F32 ⇒ false) and answers:

  • Unconditionally supported: Input, Add, Mul, Silu, SwiGLU, RmsNorm, QkNorm, MatMul, Attn, KvcacheStore, KvcacheLoad, View, Reshape, Permute, GetRows, FusedQKV, QkvBiasRopeStore, FusedFFN.
  • Conditional: RoPE only for RopeStyle::NonInterleaved (the neox layout; the only style the supported architectures emit).
  • Everything else (Scale, Softmax, BatchMatMul, FusedQkvNorm) stays on CPU.

Two layers of checks are deliberately not in supports_op:

  1. Weight/quant eligibility is a model-level, all-or-nothing gate (weights_on_cuda, §4.6). A layer whose weights are not all registered in a kernel-supported type keeps the whole graph on CPU rather than creating a partial-GPU split. The whitelist for matmul weights is Q4_0/Q4_1/Q5_0/Q5_1/Q8_0/Q4_K/Q5_K/Q6_K/F32; the embedding (tok_embd) additionally has its own type list because it is gathered, not multiplied.
  2. Shape and feature invariants are enforced in execute_node and return Err — never a silent fallback. Guards include: transpose_b unsupported; quantized matmul id % 32 == 0; RmsNorm/ QkNorm dim a nonzero multiple of 4 (the float4 kernel); attention nkt == n_head_kv * hd, hd == hd_kv, hd a multiple of 4 in 1..=128, n_head % n_head_kv == 0; the fused decode nodes nt == 1; RoPE neox; KvcacheStore's output buffer must be the K region; a missing declared norm weight is an error.

4.4 Execution dispatch

execute_node_inner opens with one stream-lock acquisition and a cache rule:

The MMQ A-quantize memo is valid only across consecutive MatMul/FusedFFN nodes; every other node kind clears it (clear_mmq_cache). FusedFFN is in the preserve set because its input is the FFN-norm output whose pre-quantized plane the fused producer just recorded.

OpCUDA pathPicking conditions
Input, KvcacheLoadno kernelhost-filled / the output is the persistent K region
View/Reshape/Permutecopy_d2d identity—
GetRows + Embed metaembed_rows_on_gpu (per weight type, incl. the padded Q6_K layout and, since #141, a dedicated f16 gather)weight registered; type via the model gate
GetRows + no metagather_rows_f32_on_gputhe G3 tail-row gather
Add / Muladd_f32 / mul_f32input element counts must match
Silucopy_d2d if not aliased, then silu_f32 in placein-place alias rule
SwiGLUproducer-fused swiglu_quant_nw (mode 2) → swiglu_quant → swiglu_f32rows >= 16 && dim % 256 == 0 plus the MMQ gate set and MINFER_MMQ_A_FUSE mode; plane OOM degrades mode 2 → 1 → unfused
RmsNormprefill producer-fused rms_norm_quant_nw/rms_norm_quant → decode rms_norm_quant_on_gpu → rms_normprefill n >= 16 && d % 256 == 0; decode n == 1 && d % 32 == 0 && !MINFER_NO_DECODE_A_FUSE
QkNormrms_norm with d = hd over the flat [nt*nh, hd] viewhd % 4 == 0, nonzero, divides the element count
MatMulmatmul_f32_ptr_layout + optional add_bias_f32see the family table below
RoPEcopy_d2d if not aliased, then rope_f32neox; hd even
KvcacheStorestore_kv_f32 / store_kv_f16, or store_kv_q8_0 for a packed cache (K then V), per kv_layoutout_buf == k_id; nt = elems(K_in)/nkt; the rows are the cells input (C6) — device data, not re-validated against n_ctx here (the allocator's kv_cells_for_seq and fill_input_i32 own that). The packed store maps one thread to one (row, 32-element block) and uses the CPU's quantizer, so both backends write the same bytes
Attnf32/f16: nt == 1 → gqa_attn_split; 1 < nt <= 16 → gqa_attn_split_batched; nt > 16 → gqa_attn_f16kv (FA prefill when hd == 128 && !MINFER_NO_FA_PREFILL, else legacy) or gqa_attn_f32. q8_0 (C4 S2b): nt == 1 → gqa_attn_split_q8_0 (the same 1-warp split-K body, rpw_gate = 0 — the hybrid 4-warp body is f16-typed); every nt > 1 → gqa_attn_f32 (the batched split kernel's bitwise-identity purpose is not claimed for Q8_0, and fa_prefill_f16kv is f16-typed shared-memory staging)the attention guards of §4.3; the batched verify path is bitwise-equal per position. A packed cache must therefore refuse --spec-draft (spec::SpecEngine::new), and its prefill is correct but off the tuned FA route
Attn windowed (explicit_span)the same entry points, instantiated with CAUSAL = false; bound carries [lo, hi) pairs (bound[t] = lo, bound[nt + t] = hi) instead of positions, and every per-row limit must come from hi. All three window modes are instantiated per layoutcuda_windowed_attention_matches_causal_for_long_windows sweeps (nh, nk, hd) × n × start × both KV dtypes; cuda_map_window_matches_the_span_over_the_same_rows sweeps all three layouts (f32/f16/q8_0) over the same rows. fa_prefill_f16kv used bound[t] (the window's lo) as the causal limit until 2026-09-19, which made every non-zero-start prefill attend to a single row — see ARCHITECTURE-EXECUTION-PLAN.md §14 row 0
FusedFFNconcat matmul_f32_ptr_layout + in-place swiglu_quant_off/swiglu_f32_offnt == 1; offset fuse when n % 32 == 0
FusedQKVconcat matmul over [wq|wk|wv] + attn_bias_rope_store (sources [x, positions, cells])nt == 1, neox, even hd; concat weight + 3 biases registered; KV pair present. positions[0] ropes q/k, cells[0] addresses the four KV writes (C6), so the node is valid for a run that does not start at cell 0
QkvBiasRopeStorecopy_d2d for q + attn_bias_rope_store over three separate matmul outputs (sources [q, k, v, positions, cells])nt == 1, neox, even hd; 3 biases registered; same positions/cells split
anything elseErr("cuda: op ... has no kernel ...")—

MatMul family selection (matmul_f32_ptr_layout in cuda.rs):

  • Prefill GEMM when (nt >= 9 || MINFER_SMALL_M_GEMM=1) && id % 32 == 0 && !MINFER_NO_PREFILL_GEMM and the type is quantized: the promoted int8 MMQ path (MINFER_MMQ, default on) — pad40 pre-transposed A planes, fused producers, raw-byte NB-BT tensor-core kernels for q4_K/q6_K; else the f16 wmma GEMM (or the persistent f16 weight cache / MINFER_FUSED_B dequant-in-GEMM).
  • Decode / small batch: per-type MMVQ (dp4a over q8_0 activations) for nt == 1 with shape gates, _multi variants for nt 2..8, and f32-activation kernels otherwise. Q4_1/Q5_0/Q5_1 have f32 kernels only; F32 weights use f32_f32_matmul_vec/_scalar.
  • f16 weights (#141): f16_f32_matmul_vec when id % 8 == 0, else f16_f32_matmul_scalar — for every nt, because f16 is not an MMQ format (MMQ streams quantized bytes) and the f16-wmma path is the MINFER_MMQ=0 fallback for the quantized types, not an f16-weight kernel. The weight bytes stay 2 B/element on the device: there is no registration-time dequant to f32, so the memory the f16 file exists to save is actually saved (a 0.5B f16 GGUF registers 942.4 MiB of device weights; ~1.9 GiB if it were dequantized). The kernel converts in-register with __half22float2 and FMA's against the f32 activations, so the accumulation is f32 like every other CUDA matmul. Both launchers read their own launch return through the #147 helpers and return non-zero → Err, rather than joining the unchecked <<<>>> sites of #162. The f16 embedding gather is embed_rows_f16 (one thread per output element) — without it the all-or-nothing weights_on_cuda check would drop a converted f16 GGUF to the CPU over its token_embd alone.
  • bf16 weights (#208, the CUDA half): the exact sibling of the f16 pair — bf16_f32_matmul_vec when id % 8 == 0, else bf16_f32_matmul_scalar, and embed_rows_bf16 for the embedding. The promotion is cheaper than f16's: bf16 is f32's top 16 bits, so the vec kernel loads 8 elements as one uint4 and shifts each word (b2f(bits) == __uint_as_float(bits << 16), exact — no rounding at all, unlike the quantized types), while the scalar kernel shifts one word per element. The launchers are their own (launch_bf16_f32_matmul / launch_embed_rows_bf16), each reading its own launch return through the #147 helpers, so a failed launch is an Err at the call site. The same two design decisions as f16 apply and for the same reasons: the weights stay 2 B/element on the device (the f16 bullet above carries the measured figure — the f16 twin's number — where a dequantized copy would be ~1.9 GiB), and a bf16 prefill does not enter the int8 MMQ GEMM quantized types only; MMQ streams quantized bytes and bf16 is not a format). bf16 is registered by the shared models::weight_reg::cuda_weight_reg rule, so both architectures admit it in one place; cuda::concat_rows has no 2 B/element arm, so the attn_qkv / ffn_gu concat copies are not registered and the bf16 graph runs the unfused matmul chain. The exactness gate is cuda_backend::tests::weights::cuda_bf16_matmul_matches_the_exact_shift_reference (both launcher arms, bitwise against f32::from_bits(bits << 16)) and cuda_backend::tests::weights::cuda_bf16_embed_gather_matches_the_exact_shift_reference; the placement/real-model gate is f208_bf16_weights_run_on_the_cuda_device (169 bf16 matmul + 1 embed nodes all on CUDA; device vs CPU max |Δlogit| 7.82e-5 / 4.24e-6 relative, bar 0.01 / 1e-3, greedy identical). The Metal twin is #208's other half and landed separately (PR #323), which widened the Metal registration arm to matches!(ttype, F32 | F16 | BF16) — see METAL-BACKEND-DESIGN.md §4.4.
  • Bias is applied by add_bias_f32 after the GEMM; its last argument is the row count nt (the kernel maps one block row per token).

4.5 Allocator and scheduler integration

  • GraphAllocator::supports priority is Metal → CUDA → CPU; enable_cuda() mirrors enable_metal(). On a CUDA-only host the practical effect is CUDA first.
  • KV compaction (C3) is not a node op — it is the Backend::copy_cells trait method, called by GraphAllocator::kv_defrag between forwards. CUDA implements it with kv_move_rows: one block, rows walked ascending when dst_row <= src_row and descending otherwise (C7b — a compaction slides a run down into the gap below it, and growing a run can move one up), with a __syncthreads() between rows, because the ranges may overlap in either direction and device-to-device cudaMemcpyAsync is documented undefined for overlap. No staging buffer, no second pass. The launcher returns non-zero on a contract violation and the Rust side turns that into an Err, so the allocator fails the compaction before it renumbers any run. It runs on the backend's own stream, so it is ordered after the previous forward's kernels.
  • KV regions are created by ensure_kv on the layer's assigned backend, so with CUDA assignment the per-layer K/V regions live in the CUDA pool and KvProvider::kv_pair returns pool ids that execute_node resolves to device pointers. init_kv_cache is bypassed.
  • Cross-backend values go through the allocator's staging path: copy_to_cpu (pinned D2H) + write_host (pinned H2D) into a buffer allocated with alloc_fresh on the consumer's backend.
  • The scheduler syncs at split boundaries (sync_backend → CudaState::sync()), which is also where a pending capture window closes. In practice the post-7e③ graph is a single CUDA split for Qwen2/Qwen3 (embed gather and tail gather are on device), so there are no per-step cross-backend copies at all; splits appear only in synthetic or intentionally mixed graphs.

4.6 Model wiring

Both architectures use the same shape (Qwen2 shown; Qwen3 mirrors it):

cuda_on  = CudaState::get().is_some() && weights_on_cuda(model)     // #[cfg(feature = "cuda")]
metal_on = metal_available() && weights_on_gpu(model)
CParams.gpu = metal_on || cuda_on
  • weights_on_cuda is the all-or-nothing gate: every graph-referenced tensor must be registered (has_weight_of_size) and, for matmul/embedding tensors, of a kernel-supported type. On failure it prints CUDA GATE: weight '<name>' (type <t>) has no CUDA kernel or is not registered and the model runs entirely on CPU. Since 7e③ tok_embd is gated like every other weight (the embedding gather is on device); its own type list is F32/Q4_0/Q8_0/Q4_K/Q5_0/Q5_1/Q6_K/Q5_K.
  • Registration happens in the loader, not in register_graph_weights (which fills the CPU registry only): cuda.register_weight per tensor, register_weight_q6k_padded for Q6_K, register_weight_q80_p32 for the q8_0 split plane, and the _exp/_dsc/__dpl planes under their gates. The loaders also build the fused concat weights with cuda::concat_rows and register them: blk.{i}.attn_qkv (Qwen2 only — the Qwen3 fused-QKV path is Metal-only, so CUDA registers no Qwen3 attn_qkv) and blk.{i}.ffn_gu (both models, gated on nf <= 16384).
  • Concat availability: Qwen2's qkv_concat_available / gu_concat_available use the metadata-only cuda::concat_rows_feasible probe on the CUDA arm, because rebuilding the concat bytes during graph construction measured a ~920 ms stall per decode graph build. Qwen3's gu_concat_available uses the eager cuda::concat_rows (its concat is built once per layer at load either way); Qwen3's qkv_concat_available has no CUDA arm.
  • Fused nodes per model: Qwen2 builds FusedQKV (concat class) or QkvBiasRopeStore (mixed-quant class, CUDA-only) for decode QKV, and FusedFFN when nf <= 16384. Qwen3 builds FusedFFN but not its FusedQkvNorm on CUDA: that fused path is Metal-only (the Qwen3 concat probe has no CUDA arm), so CUDA Qwen3 runs the unfused qk_norm chain.
  • FusionPass receives the enabled backends in its Vec<&dyn Backend> (F4: GraphAllocator::fusion_backends) and its node → index map is GraphAllocator::fusion_backend_index, so the SwiGLU rewrite is gated by CUDA's own supports_fused and the two cannot drift apart. The pre-F4 hand-built vector and its b.name() == "cuda" position lookup are gone.
  • The dump/debug tags in forward_cached are backend-agnostic (MINFER_GRAPH_DUMP; MINFER_REBUILD_TRACE=1, Qwen2 only). MINFER_CUDA_DEBUG was a device-layer trace on the legacy layer_gpu surface; #240/#241 deleted that surface, so the knob and its per-node syncs (debug_sync) are gone — the graph path's CudaState::sync() is the remaining drain point.

4.7 Memory, residency and staging

Residency. Weights are uploaded at load and never copied during execution; only activations, positions and logits cross PCIe (tiny for decode). Quant-specific registrations trade device memory for speed: the Q6_K padded repack (224-byte slots), the q8_0 p32 split plane (≈ +94% of that tensor's bytes), the q6_K W_exp dense plane (~1.52 GB on 7B), the q4_K/q6_K W_dsc f32-pair planes (~1.46 GB), and the D4-4 q6_K dense split plane. Each has an opt-out gate (§5, env-gate reference in CUDA_OPTIMIZATION.md Appendix A); the promoted default spends ~3.27 GB to reach the 1.080× prefill path.

KV layout. The persistent K/V regions keep their f32 IR shape, but the store/attention kernels run one of three layouts, tagged by crate::cuda::KV_LAYOUT_F32/F16/Q8_0 — the same 0/1/2 codes KvFormat uses, and a host contract the kernels are templated on (int LAYOUT in the src/cuda/kernels/*.cu KV/attention TUs):

  • KV_LAYOUT_F32 — one f32 per element;
  • KV_LAYOUT_F16 — one f16 per element in the first half of the f32-shaped region (kvformat::auto_device_format selects it when n_layers × n_kv_embd >= 8192, the 7B class, and MINFER_CACHE_TYPE=f16 overrides). Its snapshots are restorable since #130: the session container records the type in its header flags (FLAG_F16, C5 S3), so an auto-f16 model's --slots-file / --session companion is no longer refused on load;
  • KV_LAYOUT_Q8_0 — packed 34-byte Q8_0 blocks, one cell rounded up to whole f32 words (MINFER_CACHE_TYPE=q8_0, C4 S2b).

Every KV address is formed in bytes: kv_row(base, cell, row_bytes) names a cell and kv4<LAYOUT>(row, elem) -> float4 is the one load idiom — the old float4 load for f32, the old two-__half2 pair for f16 (both bit-identical to the pre-C4 instantiations), and for Q8_0 the f16 scale plus four quants of block elem/32. A 4-element group never straddles a block because a KV head's base is hd-aligned and hd % 32 == 0 (ensure_kv's packed-width check). The kernels take const void* k/v plus size_t row_bytes, and the launchers take the layout as an int.

CudaBackend holds that int in its kv_layout field, and kv_row_bytes(nkt) derives the stride (nkt*4, nkt*2, or KvFormat::Q8_0.row_bytes(nkt)). The packed store is store_kv_q8_0, whose quantizer is the CPU's step for step (amax/127, f16 scale, round-ties-even), so both backends store the same bytes. Before C4 S2b this was a bool that mapped anything not exactly f16 to f32 — which would have addressed a packed region as f32 rows, the silent corruption the layout tag exists to make impossible.

Per-engine scope (#99, completed by #153). #99 made the KV format per engine for the model, the graph builder (CParams::kv_format), the allocator and the CPU kernels; #153 finished the device half. There is no cuda::KV_LAYOUT static any more: the engine's resolved KvFormat — including the GPU auto policy, which kvformat::resolve now folds in from the model dims — reaches CudaBackend through GraphAllocator::set_kv_format (and enable_cuda builds a fresh backend from the same stamp), so CudaBackend::kv_layout is the only source the dispatch reads. cuda::layout_of / format_of are the one binding between KvFormat and the FFI tag. The launchers already took the tag as an argument; what was process-wide was the value. Two engines in one process therefore run their own layouts, and the captured-graph key carries the tag (below), so an exec instantiated for one layout cannot replay for another. models::load_model_configured no longer restates anything.

All three backends keep the format per engine. Metal joined in #44 part (b); the decision and the deleted process-wide symbols are in ADR-0005 and ADR-0006, and the Metal detail is docs/METAL-BACKEND-DESIGN.md.

Host transfers.

  • H2D fills go through a lazy ring of 8 × 2 MiB pinned slots (write_input_async): the Rust slice is copied into a slot and the cudaMemcpyAsync is queued; same-stream ordering guarantees the consumer kernels see the data. A ring wrap retires in-flight copies with one stream sync; oversized inputs or a failed cudaHostAlloc fall back to a blocking copy.
  • D2H readbacks use a single grow-on-demand pinned buffer (PinnedBuf), pre-grown to 4 MiB at first prefill because the lazy cudaHostAlloc showed up as a 0.78 ms tail malloc at the logits readback. MINFER_NO_PINNED_READBACK=1 reverts to pageable copies.
  • Device memory is not host-readable by plain memcpy on GB10 — a host probe that dereferences a device pointer faults. All D2H goes through cudaMemcpy staging (copy_to_host).
  • Trace/viz uses CaptureStaging: one pinned buffer (128 MiB ceiling) into which each captured node's output is queued as a stream-ordered async D2H right after its launch, drained with a single sync at the split boundary. Node outputs above the ceiling fall back to the per-node synchronous copy.

Caches. The MMQ A-quantize memo (consecutive-window rule above), the positions→i32 memo (keyed by (input buffer id, pool_gen), cleared in synchronize), and the optional persistent f16 weight cache (w16_cache, enabled when quantized matmul weights ≥ 2 GiB and MMQ is off) are all execution-window caches with explicit invalidation points rather than long-lived state.

4.8 CUDA Graph capture and replay

graph_replay(uid, range, nt_hint) is called by the scheduler once per split, before its node loop. The state machine:

  1. graphs_mode != Enabled → direct launches.
  2. An open capture window of our own → direct launches (a nested replay would be CUDA-invalid; the graph is single-split today so this is unreachable).
  3. A stored exec for (uid, range) with a matching pool_gen and kv_layout → cudaGraphLaunch; on launch failure, disable graphs for the session and fall back.
  4. A stored exec with a different pool_gen (pointer layout may differ) or a different kv_layout (the recorded kernels were instantiated for the old tag — store_kv_f16 vs store_kv_q8_0, the layout-tagged attention) → destroy it, drop the warmup counter, re-warm. #153 added the kv_layout term; CudaBackend::set_kv_layout also invalidates eagerly when a stamp moves, so the lookup check is the second line of defence.
  5. Warmup: executions 1 and 2 of a key run direct launches (llama.cpp warms up twice); a one-shot prefill never reaches capture.
  6. On the third execution, if nt_hint.map_or(true, |nt| nt == 1 || prefill_capture), the backend opens a capture window on its own stream, in cudaStreamCaptureModeThreadLocal (§2.4) — no process-wide lock, because no other backend shares this stream. The gate means decode-shaped graphs always capture; prefill-shaped graphs capture only when prefill_capture is on (default ON since R3-B; MINFER_NO_PREFILL_CAPTURE=1 opts out).
  7. The window closes at the split's synchronize → close_capture_or_sync: end capture, instantiate, launch once so the step still produces output, cache the exec at the current pool_gen and kv_layout, then sync. A failure destroys the exec, logs loudly, and disables graphs for the session.

Replay correctness rests on stable addresses: pool ids never move memory, copy_across rewrites the same staging buffers each step, and the positions scratch pointer is embedded in captured execs — so growing it bumps pool_gen and forces re-capture.

Two interactions are part of the contract:

  • Trace/viz disables replay. The scheduler skips graph_replay entirely while MINFER_TRACE or live viz capture is active, because per-node host readbacks inside a capture window are illegal.
  • A node error inside the window aborts it (abort_capture): end capture without launching, destroy the exec, disable graphs, sync. Later steps run direct-launch with graphs disabled; the aborted step's outputs were never produced and are consumed as-is — there is no poisoned-error mechanism, and the code says so explicitly.

4.9 GPU safety (CUDA edition)

The hard rules live in docs/GPU_SAFETY.md (CUDA section); this is how the backend implements them:

  1. Errors are errors. Guards return Err naming the node and the blocking values; the scheduler aborts. There is no mid-run CPU fallback.
  2. No sync inside an active capture window. The 7e② incident — a temporary sync wrapper that produced garbage only with graphs on — is the recorded reason. Debug reads go through the boundary.
  3. D2H always through staging (GB10 device memory is not host-readable by plain memcpy).
  4. A latched error is never blamed on the kernel that just ran. sync() polls cudaGetLastError + cudaStreamSynchronize; the first reports whatever an earlier call on the thread latched, so its message names the observer and the cudaGetErrorName symbol and says it is not attributed to a kernel (the error is counted and cleared, not dropped). The two origins behind the old phantom "kernel launch error: 1" were a rejected cudaFuncSetAttribute and cudaGraphDestroy called on a cudaGraphExec_t — both fixed at their call sites by #145, and the remaining unchecked sites of the same class (every MMQ dynamic-smem opt-in and launch, the prefill-GEMM launcher's own launch, and graph_end_capture_to_exec's cudaGraphDestroy) by #147. #162 then removed the class entirely: every <<<>>> in src/cuda/kernels/*.cu reads its own error, so a latched error at sync() is by construction an error no site read (a non-launch API call), never an unattributed launch.
  5. A return value that gates a later launch is read where the call is made. Every dynamic-smem opt-in goes through minfer_smem_optin: an over-limit request is skipped with the reason, any other failure is named and cleared at the call site and the launch is refused (a launch over an un-opted-in dynamic smem cannot succeed). Every launch goes through minfer_launch_ok, whose immediately-following cudaGetLastError is a launch check because minfer_launch_prelude cleared (and reported) any latch that predates the launch. The sync poll is the backstop for a missed site, not the place to diagnose one; #147's issue147_tests gates inject a real failure at every one of the sites and assert the named report, the refusal and a clean latch. #162 states the per-op severity in the helper (the plan's #162 record holds the site census): minfer_launch_ok is required — it records a sticky failure that CudaBackend::execute_node turns into an Err naming the site (one Rust-side check, not one per launcher, so the op never proceeds on a stale output) — while minfer_launch_ok_opt only names and clears for a path with a documented fallback (the MMQ fast paths, the fa-prefill smem fallback, and the int-returning launchers whose Rust caller already decides). Both levers are data: minfer_launch_block (an illegal block geometry) and minfer_launch_smem (an over-limit dynamic smem request) make the real launch fail.
  6. Same-stream ordering is the async-fill contract; the pinned ring syncs on wrap and never hands a slot back early.
  7. Weight-registry ownership: name+size reuse, different-size replace with a deliberate, bounded leak (a live captured graph may still reference the old buffer).
  8. Device limits are queried at runtime (SM count, compute capability, free memory); the only hardcoded shape knowledge is the compiled target list and the documented kernel invariants.

5. Implementation Phases

5.1 Phase 7 (7a–7e)

PhaseContentStatus
7aSkeleton + wiring: CudaBackend struct/pool/trait impl, allocator arms, enable_cuda, supports() priority; execute_node handles only Input✅
7bPer-op execution + parity tests: full dispatch; un-allow the used device-layer methods; launch error checks✅
7cModel wiring + E2E: CUDA gate + weights_on_cuda, FusionPass backend index, legacy KV pre-alloc removal✅
7dCUDA Graph capture/replay: uid population, device-derived attention bound, capture state machine, replay hook in the trait/scheduler✅
7e①CPU-path residual diagnosed as a path-identity artifact (cross-backend f32 reduction order), not a bug; gates switched to greedy equality on CUDA builds✅
7e②Vectorized q4_K/q6_K kernels + Q6_K padded repack: 7B decode 8.4 → 26.4 tok/s (3.1×)✅
7e③Embed + generic GetRows on device: prefill/decode become a single CUDA split (no cross-backend copies)✅
7e④F32×F32 matmul kernels; F32-weight models participate in CUDA✅
7e⑤FusedFFN on CUDA (concat_rows + swiglu_f32_off); CParams.fuse_ffn decoupled from the QKV gate; 0.5B +10%, Qwen3-0.6B +4%✅
7e⑥Async H2D input fill through pinned staging; Q8_0 prefill GEMM shape-gated (8c)✅
7e⑦Docs + cleanup: per-item #[allow(dead_code)] with reasons, release build warning-free✅

5.2 After Phase 7

The optimization campaign is indexed in docs/CUDA_OPTIMIZATION.md §0 and expanded in docs/cuda_optimization_steps/. For the design record, its eras and what each one changed in the backend or its graph integration:

Era / workstreamBackend impactStep docsRepresentative commits
Phase 8 foundations (8m–8p, 8e)wmma f16 prefill GEMM, FA-style tiled prefill attention, decode-start stall elimination, persistent f16 weight cache, decode MMVQ02–06ba3f317, cdc6599, cb66fca, 65b686c, 2992f57, b7b8e73, 1298cb2, 1d28235
Phase 8 correctness/coverage (8a–8q)KV f16, shaped Q8_0 GEMM, split-K attention, Q5_K/Q5_1/Q5_0 kernels, the F32-matmul latent bug, llama baseline78, 79f7b0036, 69a27c5, a5af60f, b959ec9, acca28f, 9f419f9
R1–R4 + P5R1 int8 MMQ, R2 MMVQ weight streaming, R3-A1 single-split prefill (tail_ids at the graph head), R3-A2 pinned D2H, R3-B prefill capture default ON, R4 decode split attention rewrite07–1140e97c9, 6df3245, 029a9a4, a213c89, 761e236, 70f57db, 86ca78c
q4_K MMQ line (r5–r37)Staging-shape search → raw-byte NB kernels → quantize-transpose prepass; the int8 tensor-core prefill path12–40d440d16, 774a116, 0957a08, bfe6bba, 851a896, ba977bf
q6_K + FA + promotion (r38–r60)q6_K BT kernel, FAP2 register softmax, shared-A dedup, fused producers, W_exp/W_dsc planes; the verified gate set promoted default-on (1.080×)41–6475aabb9, d38744d, 87a75a3, cf1ed4b, 910d967, 83fee77, 4cf7c74, 36a481f, 57edcf6
Decode campaign D1–D4-4Split-K decode attention, hybrid rpw dispatch, fused decode A-quantize, D4-2 correctness fix, D4-4 q6_K dense split plane65–76a5af60f, 22336b2, 3230b2b, b31084c, ffce151
D5 → D5-R speculative decodingNew forward_graph_cached consumer: draft nt=1 chain + target verify at nt=d+1; verify attention bitwise-equal per position; adaptive depth80–104a6b7cf3, c3d4bb1, 0fe132f, 5471680, b3e5dab

The retired CUDA-FOLLOWUP-PLAN.md was consolidated into step docs 78/79 (commit ace6242); its residual open items are listed in CUDA_OPTIMIZATION.md §1.4.


6. Risks and Open Questions

The plan's original risk table, with its resolution:

#RiskStatus
1Host-scalar nk baked into captured graphs → stale attention windowResolved: the causal bound is derived from the device positions buffer inside the attention kernel; no host scalar crosses, and the v1 positions readback never existed in the graph path
2Sync/readback inside a capture window corrupts captureResolved: replay is skipped under trace/viz; abort_capture handles node errors; the 7e② incident is recorded in GPU_SAFETY.md
3store_kv layout vs allocator KV region layout mismatchResolved: KV roundtrip tests (cuda_rope_kv_attn_roundtrip, cuda_kv_f16_roundtrip_attn)
4Legacy default-stream implicit sync is load-bearingResolved: explicit sync() before every D2H
5Pool never shrinks → VRAM high-water markAccepted: same policy as CPU/Metal; Drop frees; documented
6Q5_K/F32-weight models silently fall back to CPUMostly resolved: Q5_K/Q5_1/Q5_0 and F32 matmul kernels landed; the gate remains all-or-nothing and logs the failing tensor
7RoPE kernel is neox-onlyAccepted: guard + Err; all supported models are neox
8No CUDA CIOpen, deferred (8h②): device-gated tests skip gracefully; GB10 is the reference bench
9F32-activation CUDA vs Q8_0-activation CPU logits differ by designAccepted: gates use greedy-text equality + tolerance classes, never cross-path bitwise

Newer, still-standing items:

ItemStatus
End-of-capture failure leaves that step's outputs undefined (loudly logged, graphs disabled)Accepted: effectively unreachable (no host syncs/readbacks in the window); there is deliberately no re-execution path
Weight-registry different-size replace leaks the stale bufferAccepted, bounded by distinct (arch, tensor) shapes; freeing would break live captured graphs
Planes cost ~3.27 GB on 7B for the promoted prefill pathDocumented, per-plane opt-out gates available
Open performance leads (q8_1 GEMM-prologue fusion, rms_nw roofline, small-od re-tile, fused ffn_gu, FA deep-opt)Tracked in CUDA_OPTIMIZATION.md §1.4 and the step-doc index; all sub-bar or step-function

7. Verification

7.1 Device test suite

src/graph/cuda_backend.rs carries the device suite (39 test functions); the device layer's own gates live in src/cuda.rs. Both are device-gated: tests skip when no CUDA device is present. Categories and the invariants they pin:

  • Capture / replay — cuda_graph_replay_bit_parity, cuda_prefill_shaped_graph_never_captures, cuda_multisplit_capture_bit_parity, cuda_prefill_capture_defaults_on, cuda_prefill_capture_bit_parity_pp16_pp300, cuda_capture_abort_on_error, cuda_graph_recaptures_on_pool_gen_change, cuda_graph_generation_replay_parity_real_model, cuda_capture_staging_order_and_fallback; the device-layer gate cuda_graph_exec_destroy_leaves_no_latched_error (#145: destroying the cudaGraphExec_t must not latch an API error) and the env-gated cuda_sync_surfaces_a_latched_error_as_latched.
  • Pool / allocator / host transfer — cuda_pool_roundtrip, cuda_pinned_readback_roundtrip, copy_across_cpu_to_cuda_and_back, kv_persistent_regions_survive_realloc, cuda_scheduler_chain, cuda_row_marginal_bench (env-gated bench).
  • Per-op parity — elementwise (cuda_elementwise_parity), norms (cuda_norm_parity), matmuls (cuda_matmul_parity, cuda_kquant_matmul_parity, cuda_q5_matmul_parity, and the per-type MMVQ parity tests for q4_K/q6_K/q5_K), attention (cuda_attn_split_decode_parity, cuda_verify_attention_nt_invariance, cuda_rope_kv_attn_roundtrip, cuda_kv_f16_roundtrip_attn), embedding gather (cuda_embed_getrows_parity), fused FFN (cuda_fused_ffn_parity).
  • KV cell movement (C3/C7b) — cuda_copy_cells_moves_overlapping_rows_in_both_directions (rows [1, 4) -> [0, 3) and then [0, 3) -> [1, 4), i.e. two of three rows are read and overwritten in each direction; the whole buffer is compared against the bytes a copy through a temporary would produce. The kernel walks the rows in the order the overlap requires — ascending when the run slides down, descending when it slides up — and kvcache::order_moves fixes that order across several runs).
  • Prefill GEMM / MMQ / FA — cuda_prefill_mmq_parity, cuda_prefill_f16_gemm_parity, cuda_q4_0_prefill_q8_0_gemm_parity, cuda_fa_prefill_attention_parity (note: causal, single-sequence, start = 0 — the windowed FA mask is covered by cuda_windowed_attention_matches_causal_for_long_windows, which is what caught the fa_prefill_f16kv window-limit fault on 2026-09-19), the byte-exact plane tests (cuda_q6k_exp_dense_byte_exact, cuda_q6k_dsc_dense_byte_exact, cuda_q4k_dsc_dense_byte_exact), cuda_prefill_fused_b_bitparity, and cuda_multi_token_matmul_bitwise (one nt=3 forward bitwise-equal to three nt=1 forwards). The lazy >48 KiB opt-in (#218) has its own gates: the_gemm_smem_formula_matches_the_kernel_layout (pins gemm_dynamic_smem_bytes against the kernel's byte layout — the value-level arm, so a silent shrink is seen); cuda_prefill_smem_optin_is_done_by_production (a real prefill forward makes the device read back opted_in == 1 for gemm_f16_nt_kernel_t<128,64,false>, in a fresh process so the "not opted in before" precondition is observable); the control arm cuda_prefill_smem_optin_refusal_fails_the_prefill (MINFER_TEST_CALL_FAIL=attr:gemm_f16_f16 refuses the launch, and the site report names the call and cudaErrorInvalidValue); cuda_prefill_smem_lazy_optin_admits_every_launchable_instantiation (every launchable >48 KiB instantiation reads back opted in through production's own gemm_smem_optin, and every over-limit one is refused without a call); and cuda_prefill_smem_optin_is_never_set_inside_a_capture_window (a >48 KiB prefill captures, replays bitwise, and gemm_smem_optin_in_capture_count() == 0 proves the opt-in ran before the window opened). Plus the_latched_error_message_never_blames_a_kernel. #147's issue147_tests module adds the_graph_destroy_failure_message_names_the_matching_destructor (pure; the injection matcher itself moved to testfail::tests::the_matcher_is_exact_and_comma_separated in #171) plus the three cuda_issue147_* deliberate-failure gates (device; env-gated behind MINFER_TEST_ISSUE147=1, which arms MINFER_TEST_CALL_FAIL per site): every dynamic-smem site names its failed opt-in and refuses the launch, every launch site names its failed <<<>>> and refuses, and the destroy site names a failed cudaGraphDestroy and leaves no latch.

Model-level CUDA coverage: cuda_conversation_multiturn_reuse (Qwen2) asserts that an incremental multi-turn session reusing the decode graph and appended KV produces the same turn-2 text as a fresh conversation that re-prefills. graph_logits_match_forward_real_model's comparison helper switches to greedy-token equality when a CUDA device is present (the 7e① path-identity artifact). The Metal-only tests (fused_qkv_matches_unfused_decode, fused_qkv_norm_matches_unfused_decode, graph_metal_*) skip on a Linux CUDA build.

7.2 Acceptance gates

PhaseGate
7aalloc/copy roundtrip + copy_across; plain (non-CUDA) build untouched, zero nvcc
7ball per-op parity tests; 0.5B Q4_0 full model: CUDA greedy text == CPU greedy text
7cE2E table (0.5B / 0.6B / 7B) with throughput; disable-env negatives; graph reuse across decode steps
7dreplay bit-parity; re-capture on pool_gen change; long-generation parity; graphs-off A/B identical greedy text
7e+per-lever A/B, recorded in CUDA_OPTIMIZATION.md and the step docs

The campaign's verification methodology — the five-gate chain (parity ×3, greedy-32 identity, interleaved A/B medians with a +1.5% whole-prefill bar, the device suite, the ncu/nsys/SASS protocol) and its transferable lessons — is docs/cuda_optimization_steps/77-verification-methodology.md. Read it before running any A/B on this engine. Two standing measurement rules: never quote llama-bench long-context throughput as an attention target without an ncu byte-count or a llama-cli recall cross-check, and an end-to-end max|Δlogits| ≈ 0.38/0.39 is the inherent class of any accumulation-order change — argmax + greedy divergence + A/B are the operative gates.

Suite baselines in the campaign records: the Phase-7e entries say "144/0 (CUDA parallel + single), 130/0 plain"; the decode campaign's later rows reach 187/0/3. The non-CUDA baseline on the current tree is 145 passed / 0 failed / 3 ignored. Run cargo test --release for the plain suite and cargo test --release --features cuda on a device.

7.2a Profiling on dgxspark (ncu / nsys)

The reusable recipe, so the next session does not re-derive it (recorded 2026-09-27, GB10 sm_121, CUDA 13.0, driver 580.178.04):

  • ncu is installed but not on PATH. The binary is /usr/local/cuda-13.0/bin/ncu (2025.3.1). which ncu finding nothing means the directory is not on PATH, not that the tool is missing — use the absolute path.
  • A normal user cannot collect counters here — RmProfilingAdminOnly: 1 makes ncu fail with ERR_NVGPUCTRPERM. sudo collects (passwordless on dgxspark); no module parameter change and no driver reload is needed. Record "collected as root; module parameter unchanged".
  • sudo changes HOME to /root, so the model cache under /home/yusiwen/.cache/minfer/models is invisible to a sudo-launched binary. Pass absolute model paths (MINFER_BATCH_TEST_MODEL=/home/yusiwen/.cache/..., or the path on the command line) or use sudo -E. A "model not found" under sudo is this, not a missing file.
  • Build as the normal user first, then attach ncu to an already-built binary. ncu does not write repository files, so target/ stays yusiwen-owned; if a sudo run does leave a root-owned file, chown it back before the next cargo build.
  • Some metrics are n/a on GB10/sm_121 — dram__bytes.sum among them (the integrated-memory / DGX Spark form exposes no classic DRAM counters). Confirm a metric name exists with ncu --query-metrics first, prefer the SM / instruction / L1 / L2 families, and treat a metric you actually collected as the only evidence; n/a is not a number.
  • Collect targeted, not whole-run. Counter replay is slow: filter with -k regex:<kernel> and bound it with --launch-count, rather than replaying a full bench.
  • A -k regex must match ncu's base kernel name, not the demangled signature. ncu lists the available kernels by base name (gqa_attn_split_partial, no template arguments), so a regex that includes < — e.g. -k 'regex:gqa_attn_split_partial<' — matches nothing and ncu prints an "Available Kernels" list instead of collecting (No kernels were profiled). Anchor it (-k 'regex:gqa_attn_split_partial$') so sibling kernels (..._combine, ..._bt) are excluded rather than eating into --launch-count. Recorded by #202, which lost one collection to this.

nsys stays the cheap, always-available instrument for per-kernel durations (nsys profile --trace=cuda --cuda-graph-trace=node ... — without node, kernels launched from a replayed CUDA graph are traced as one graph and never appear individually, which silently hides the whole decode path).

7.3 Recorded verifications — #145, #147, #141, #165, #167

These five are per-ticket verification records, and their durable home is the plan's ticket section (docs/ARCHITECTURE-EXECUTION-PLAN.md, one section per ticket). What their gates assert is stated where it belongs, not here: the failure-injection seam in docs/GATE-CONTRACT.md (#171), the launch-return rules in §4.9, and the q4_K dsc admission in §2.3.

The old §7.4–§7.7 and §7.9 headings were folded into this one; §7.8 and §7.10 keep their numbers because their records are not wholly duplicated — they carry the norm-weight invariant and the #189 value arm respectively.

7.8 The norm weight type is part of the rms_norm invariant

norm_weight(node, elems) resolves CudaState::weight_size(name) (the registered raw length, the same convention has_weight_of_size uses for a padded Q6_K plane) and requires exactly elems * 4, returning Err that names the registered length, the length the kernel reads, the node and "f16-norm". Both callers pass the dim the kernel uses (node.out_shape[0] for Op::RmsNorm, *hd for Op::QkNorm), so the check is the kernel's own geometry, not a name heuristic. It exists because the CUDA rms_norm kernel indexes the weight as d f32 elements regardless, so an f16 (2 B/element) norm weight — which minfer quantize --type f16 used to write for every 1-D tensor — would be read past its end.

The gate is device-only (CI has no GPU) and its arithmetic has no CI-covered pure twin: it is one usize comparison against elems * 4. It does not make an f16-norm model loadable on CUDA — registration still admits an f16 1-D weight, and the refusal happens at execute time with the node named (the Err-not-fallback rule of docs/GPU_SAFETY.md). The verification record (the mutation that made the f16 arm execute, the compute-sanitizer output, the suite counts and the #169 end-to-end acceptance run) is in the plan's #169 section.


7.10 The S4 map-window A/B: a value arm and a paired sign test

The defect. The gate was a pure stopwatch: it interleaved matched rounds of the span window and the map window and asserted the median of 9 per-round ratios <= 1.25x, which flips once five pairs are disturbed — and one parallel #[ignore]d device run had exactly five above the bar. The verdict was about how loaded the box was, not about the kernel. The failing transcript, the co-tenant reading and the bar's #123 provenance are in the plan's #189 section.

The fix, two halves.

  1. A value arm first (gate contract rule 1). Before any timing the gate asserts
    • a one-row map window at cell 512 returns exactly that row's V — an absolute value computed on the host, not a relation between the two modes (softmax over one key is exactly 1.0, so the kernel is the identity on V);
    • a two-run map window at cell 512 returns the span's bytes over the same rows, bit for bit — f32 KV at the decode shape (nt = 1), f16 KV at the FA-prefill shape (nt = 512). One run is indistinguishable to a resolver that reads (cell, len) as (lo, hi); two runs, at a non-zero base, are not;
    • the map instantiation actually ran, by counted observation rather than the dispatch's own report: testfail::note_checked("cuda_attn_map_window") is bumped in the launcher (CudaState::gqa_attn_split / gqa_attn_kv_prefill, src/cuda.rs), and the gate resets it, runs one span call (counter stays 0) and one map call (counter 1). The bitwise arms say what was computed; the counter says the map path was the thing that computed it.
  2. A paired sign test (gate contract rule 4). The timing verdict is the count of matched pairs whose map arm is above 1.25x its own span arm, and the gate refuses only at 7 of 9 — the one-sided sign test at alpha = 46/512 = 0.090. A load spike that disturbs a minority (or even a bare majority) of pairs cannot decide it; a doubled map cost moves all 9 and does. The bar is unchanged at 1.25x and the timed fixture is the pre-#189 one, so the recorded margins stay comparable; the round count is fixed in advance (PAIRS = 9) and every per-round ratio and the refusal count are printed. The reproducible mutation is MINFER_S4_AB_MAP_REPS=2 (src/cuda.rs::s4_ab_map_reps), which issues every map-mode attention launch twice — an implementation mutation, not a test edit.

Honest scope. Every number is a local GB10 measurement; CI has no GPU, so its CUDA job only compiles the harness. The two-run value arms are the gate's own fixtures, not the graph path — the graph-path bitwise coverage stays with cuda_map_window_matches_the_span_over_the_same_rows (which sweeps f32/f16/q8_0 over one/two/three runs and both batch shapes). The MINFER_S4_AB_MAP_REPS seam is wired to the two launches this gate drives (gqa_attn_split, gqa_attn_kv_prefill), not to the batched split path. The bar 1.25x is inherited unchanged from #123; this ticket changed the statistic and added the value arm, and did not widen it. The measured tables — the idle gate and six parallel runs, with every per-round ratio — the mutation output, the suite counts and the superseded src/device_entry.rs guard note are in the plan's #189 section.

8. Out of Scope / Future

  • Not planned (revisit with a concrete need): cuBLAS/cublasLt, VMM pool, multi-GPU + peer copies, graph_optimize-style node reordering, Windows, self-hosted CUDA CI, IQ/Q2/Q3 quants.
  • Open leads, all sub-bar or step-function (details in CUDA_OPTIMIZATION.md §1.4): the q8_1 GEMM-prologue fusion (the identified step change); the rms_nw roofline (+0.5–1%); a wave re-tile for small-od classes (+0.3–0.8%, needs ≤85 registers); a fused ffn_gu concat (needs the G5 nf <= 16384 gate re-measured); FA deep-opt only with a numerics-order-preserving structure.
  • Inherited Phase-8 ledger: the shape-dependent q6_K MMVQ idle-tail rule (not started), the CUDA CI runner (deferred), and the /tmp/minfer_phase7/ ledger cleanup (awaiting a decision).
  • Landed since the original plan and therefore no longer in this list: prefill CUDA-Graph capture, FA-style prefill attention, fused decode nodes on CUDA, pinned host buffers, f16 KV, Q5_K/Q5_1/Q5_0 and F32 kernels.

Running on a real device — stream ownership and launch checks (moved from AGENTS.md)

scripts/cuda_test.sh (cargo test --release --features cuda -- --test-threads=1) is the only way to run the device-gated tests — CI has no GPU and its CUDA job only compiles the harness — and because the device state is a process-wide singleton (CudaState) it must run serially (#64); a local run must rebuild the CLI with the feature (a plain cargo test --release overwrites target/release/minfer with a CPU-only build and silently measures the CPU), and MINFER_DISABLE_CUDA is presence-checked (=0 disables CUDA). The per-instance stream, capture-window and scratch design, the #188 probe and its measured table, the launch-return audit (scripts/check_cuda_launch_returns.py) and the injection lever (MINFER_TEST_ISSUE162=1, which drives every audited launch return through a real failing launch) are §2.4; the retired src/device_entry.rs guard is the #240/#241 record in docs/ARCHITECTURE-EXECUTION-PLAN.md.

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0001 — Inference runs through one declarative compute graph
  • ADR-0002 — Topology is a function of GraphParams alone, so positions cannot be structure
  • ADR-0006 — The KV storage format is a per-engine gate, not a process-wide global
  • ADR-0008 — GPU safety: bounded waits, no early return past a barrier, runtime device limits
  • ADR-0009 — A failure is an error, never a silent fallback
  • ADR-0010 — The identity gate: bitwise by default, a named tolerance class otherwise
  • ADR-0013 — The CPU quantizes activations to Q8_0; a device reads f32
  • ADR-0014 — A KV session is a versioned, checksummed file — never a memory dump
  • ADR-0021 — bf16 is a round-to-nearest-even cast, and 1-D tensors stay f32
  • ADR-0030 — Launch severity lives in the helper, not in 120 call sites
  • ADR-0031 — Capture runs in thread-local mode, because Global lets a foreign thread's call join the window
  • ADR-0033 — The CUDA pool recycles exact byte lengths, never frees, and reports OOM as an error
  • ADR-0034 — The smem opt-in is an eager pre-warm by construction, with the lazy path as defence in depth
  • ADR-0035 — The q4_K dsc plane is admitted by two gates, and the payload test is equality
  • ADR-0036 — bf16 weights get their own device kernels, not a dtype flag on the f16 ones
  • ADR-0038 — The per-tensor registration dispatch is one shared rule, not a copy per loader

CPU Inference Path Analysis & Optimization Results

This document analyzes minfer's CPU inference performance, compares with llama.cpp, and documents optimization attempts and their outcomes.

Status and scope (updated 2026-10-08). This file is a layered record. The sections from the top down to §Conclusion are a 2026-07-01 snapshot taken before the compute-graph rewrite: their src/models/qwen2/forward.rs:NNN line references point at a file that was deleted in Phase 6, and their "minfer is single-threaded" premise no longer holds. The current numbers are in §"NEON + Threading Overhaul" (2026-08) below and in docs/PERF-QWEN3-4B-VS-LLAMACPP.md §3.

Four of the four "Remaining Optimization Opportunities" from that snapshot have since shipped: P4 multi-threading (persistent CPU thread pool, src/kernel.rs), P2 f16 KV cache (MINFER_CACHE_TYPE, auto-selected for 7B-class models), flash attention (on both GPU backends; the CPU path keeps the multi-pass form), and P1 — the AVX2/AVX-512 dot products for the K-quant types (#56, landed 2026-10-08; see §P1 below). The remaining piece of that lane is weight repacking, tracked in docs/ARCHITECTURE-ROADMAP.md §2.7 and docs/SUPPORT-MATRIX.md.

Current State

Baseline performance: 27.1 tok/s decode on Qwen2-0.5B (AVX2, i7-1260P). llama.cpp comparison: ~60-80 tok/s on the same model. Performance gap: ~2.5-3×.

Key finding: The codebase is already well-optimized. All major SIMD optimizations identified in initial analysis were already present in the codebase. Optimization attempts caused performance regressions due to function call overhead and interference with compiler auto-vectorization.


Existing Optimizations (Already Present)

The forward pass already uses SIMD-optimized operations throughout:

1. RMSNorm with Fused Weight Multiply

Location: src/models/qwen2/forward.rs:242-250

#![allow(unused)]
fn main() {
fn rms_norm(x: &[f32], eps: f32, out: &mut [f32], n: usize, d: usize, w: Option<&[f32]>) {
    for t in 0..n {
        let row = &x[t * d..(t + 1) * d];
        let dst = &mut out[t * d..(t + 1) * d];
        match w {
            Some(w) => crate::vec_ops::rms_norm_fused_f32(d, dst, row, w, eps),
            None => crate::vec_ops::rms_norm_f32(d, dst, row, eps),
        }
    }
}
}

Uses rms_norm_fused_f32 from vec_ops.rs:511-533 with AVX2 sum-of-squares computation using _mm256_fmadd_ps. Fuses RMSNorm + weight multiply in a single pass.

2. Vectorized Residual Add

Location: src/models/qwen2/forward.rs:110-115

#![allow(unused)]
fn main() {
unsafe {
    crate::vec_ops::vec_add_f32(hidden.len(),
        std::slice::from_raw_parts_mut(hidden.as_mut_ptr(), hidden.len()),
        std::slice::from_raw_parts(hidden.as_ptr(), hidden.len()),
        &bn);
}
}

Uses vec_add_f32 with AVX2 _mm256_add_ps (8-wide). Called twice per layer (48 times total for 24-layer model).

3. Vectorized Gate × Up Multiply

Location: src/models/qwen2/forward.rs:127-130

#![allow(unused)]
fn main() {
crate::vec_ops::vec_mul_f32(len,
    std::slice::from_raw_parts_mut(bg.as_mut_ptr(), len),
    std::slice::from_raw_parts(bg.as_ptr(), len),
    &bf);
}

Uses vec_mul_f32 with AVX2 _mm256_mul_ps. Called once per layer (24 times).

4. RoPE Sin/Cos Cache

Location: src/models/qwen2/forward.rs:257-275

#![allow(unused)]
fn main() {
let mut sin_cache = vec![0.0f32; half];
let mut cos_cache = vec![0.0f32; half];
for t in 0..pos.len() {
    let p = pos[t] as f32;
    for i in 0..half {
        let th = p * freqs[i];
        let (sn, cs) = th.sin_cos();
        sin_cache[i] = sn;
        cos_cache[i] = cs;
    }
    // Use cached values for all heads
    for h in 0..nh { ... }
}
}

Precomputes sin/cos once per position (64 values), shared across all 32 heads. Reduces sin_cos() calls from 2048 to 64 per token per layer (32× reduction).

5. SiLU Activation

Location: src/models/qwen2/forward.rs:120

#![allow(unused)]
fn main() {
crate::vec_ops::vec_silu_f32(len, &mut bg);
}

Uses vec_silu_f32 with AVX2 implementation.


Optimization Attempts & Results

Attempt 1: Activation Quantization Reuse

Goal: Quantize activation once, reuse Q8_0 buffer for Q/K/V matmuls.

Implementation: Added quant_matmul_f32_batch_prequantized and quantize_activation_f32 to kernel.rs. Pre-allocated Q8_0 buffer outside layer loop.

Result: 28.0 → 26.0 tok/s (-7% regression)

Why it failed:

  • Added function call overhead for small batch sizes (nt=1 during decode)
  • Extra memory traffic from separate quantization step
  • Compiler couldn't optimize the separated code as well
  • The original code's per-matmul quantization allows better inlining and register allocation

Attempt 2: Online Softmax

Goal: Merge 4-pass attention into single-pass online softmax.

Implementation: Replaced multi-pass softmax with running max/sum algorithm (ref: arxiv 2112.05682).

Result: 27.1 → 27.4 tok/s (+1% marginal gain, then regressed to 25.1 tok/s after further changes)

Why it failed:

  • For small sequence lengths (typical in decode), the overhead of maintaining running max/sum outweighs the benefit
  • Compiler auto-vectorization of the simple multi-pass version is very effective
  • The attention loop is already memory-bound, not compute-bound

Attempt 3: RoPE Cache Stack Allocation

Goal: Avoid heap allocation in RoPE by using stack arrays.

Implementation: Changed vec![0.0f32; half] to [0.0f32; 64].

Result: 27.1 → 26.7 tok/s (-1.5% regression)

Why it failed:

  • Stack allocation of 64 f32s (256 bytes) is not significantly faster than heap allocation for this size
  • May have interfered with compiler's register allocation strategy
  • The heap allocation happens once per layer (24 times), not per token

Attempt 4: Attention Head Parallelization (Multi-Threading)

Goal: Parallelize attention heads across CPU cores using std::thread.

Implementation: Added gqa_attn_parallel function that splits attention heads across multiple threads. Each thread processes a subset of heads independently, with its own scores buffer to avoid contention. Used std::mem::transmute to bypass Rust's Send trait check on closures capturing raw pointers (safe because each thread accesses non-overlapping output regions).

Two variants tested:

  1. Always parallel: Parallelize for both decode (nt=1) and prefill (nt>1)
  2. Prefill-only: Only parallelize when nt>1, keep original single-threaded path for decode

Results:

VariantPrefill (42 tokens)Decode
Baseline (single-thread)29.2 tok/s28.0 tok/s
Always parallel25.3 tok/s (-13%)25.3 tok/s (-10%)
Prefill-only parallel27.1 tok/s (-7%)26.3 tok/s (-6%)

Why it failed:

  • Thread creation overhead: Each gqa_attn call spawns 16 threads (one per head). With 24 layers, that's 384 thread create/join cycles per forward pass. Thread creation on Linux costs ~10-20μs each.
  • Small workload per thread: For decode (nt=1), each head processes only hd=64 elements per KV position. The computation per thread is too small to amortize thread startup cost.
  • Memory allocation per thread: Each thread allocates its own vec![0.0f32; nkv] scores buffer, adding heap allocation overhead.
  • Code bloat: Adding the parallel path increases binary size, reducing instruction cache hit rate for the hot decode path.
  • llama.cpp uses a thread pool: Unlike our per-call thread creation, llama.cpp uses a persistent thread pool (ggml_threadpool) where threads are created once and reused across all operations. This eliminates the creation overhead.

What would be needed to make this work:

  1. Implement a persistent thread pool (create threads once at startup)
  2. Use work-stealing or task queue pattern instead of per-call spawn
  3. Consider using rayon crate which provides an efficient work-stealing thread pool with minimal overhead
  4. Only beneficial for larger models (7B+) or longer sequences where the attention computation per head is substantial

Attempt 5: Crossbeam Work-Stealing & usize Pointer Trick

Goal: Research crossbeam's work-stealing deque and find a way to bypass Rust's Send/Sync checks on raw pointers in multi-threaded closures.

Key Discovery: usize Pointer Trick

Rust's type system prevents passing raw pointers (*const f32, *mut f32) across thread boundaries because they don't implement Send. Even wrapping them in structs with unsafe impl Send fails because the compiler recursively checks struct fields.

Solution: Convert raw pointers to usize before the closure, then reconstruct inside:

#![allow(unused)]
fn main() {
// Before closure (main thread)
let ptr = slice.as_ptr() as usize;  // usize is Send

std::thread::scope(|s| {
    s.spawn(move || {
        // Inside closure (worker thread)
        let slice = unsafe { std::slice::from_raw_parts(ptr as *const f32, len) };
        // Use slice...
    });
});
}

This works because:

  • usize is just an integer type that implements Send
  • The compiler doesn't see raw pointers in the closure capture
  • The pointer values remain valid because std::thread::scope ensures all threads join before the scope exits (pointees stay alive)

Crossbeam Research Findings:

Tested both std::thread::scope and crossbeam::scope:

  • Both require Send on closures passed to spawn()
  • Both allow the outer closure to be non-Send
  • crossbeam::deque provides work-stealing deques but doesn't solve the fundamental pointer issue
  • The usize trick works with both std and crossbeam

Implementation:

Applied the usize pointer trick to gqa_attn with conditional parallelization:

  • Only parallelize when nt > 8 (large prefill)
  • Keep original inline single-threaded code for decode (nt=1)
  • Each thread gets its own score buffer to avoid contention

Results:

VariantPrefill (35 tokens)Decode
Baseline (single-thread)29.1 tok/s25.4 tok/s
usize trick (nt>8 parallel)27.2 tok/s (-6.5%)24.4 tok/s (-3.9%)

Why it still failed:

  • Thread creation overhead dominates: Even with conditional parallelization, creating 4-16 threads per attention call (24 layers) adds 240-960μs overhead per forward pass
  • Small workload per head: For 0.5B model with hd=64, each head processes only 64 elements per KV position. Even with 35-token prefill, the work per thread is ~2240 operations, completing in ~10-20μs—not enough to amortize 10-50μs thread creation cost
  • Compiler optimization interference: The presence of parallel code paths may prevent the compiler from fully optimizing the single-threaded path (instruction cache pressure, register allocation changes)
  • Memory allocation overhead: Each thread allocates vec![0.0f32; nkv] score buffer

Crossbeam Work-Stealing Deque Analysis:

crossbeam::deque::Worker<T> + Stealer<T> pattern:

  • Main thread creates Worker, pushes task indices
  • Worker threads get Stealer handles, steal tasks when idle
  • Good for load balancing when task sizes vary

However, for attention head parallelization:

  • All heads have identical work (same hd, same nkv)
  • No load imbalance to address
  • Work-stealing overhead > benefit for uniform tasks
  • Still requires persistent thread pool to avoid creation overhead

Conclusion:

The usize pointer trick successfully bypasses Rust's Send/Sync checks, enabling raw pointer sharing across threads. However, for small models (0.5B) and typical sequence lengths, the thread creation overhead dominates. Multi-threading would only benefit:

  • Larger models (7B+) where attention computation per head is substantial
  • Very long sequences (1000+ tokens) where work per head amortizes thread costs
  • With a persistent thread pool (like llama.cpp's ggml_threadpool) to eliminate creation overhead

For this 0.5B model, the compiler's auto-vectorization and single-threaded execution remain more efficient than multi-threading.


Lessons Learned

1. Compiler Auto-Vectorization is Highly Effective

For small models (0.5B) and short sequences (decode phase with nt=1), the Rust compiler's auto-vectorization is extremely effective. Simple loops with clear patterns are automatically vectorized to AVX2.

Manual SIMD function calls add overhead:

  • Function call overhead (even with #[inline])
  • Memory traffic from explicit loads/stores
  • Interference with compiler's optimization passes

2. Function Call Overhead Matters

For operations on small arrays (e.g., 128 elements in attention), the function call overhead can exceed the computation time. The existing code already uses SIMD functions where the array size justifies the call overhead.

3. Memory Bandwidth is the Bottleneck

During decode (nt=1), the inference is memory-bound, not compute-bound. The dominant cost is loading weights from memory, not computation. Optimizations that add memory traffic (separate quantization, intermediate buffers) hurt performance.

4. Don't Fix What Isn't Broken

The initial analysis incorrectly identified "scalar hotspots" as bottlenecks. In reality, the code already had SIMD optimizations. The perceived gap between minfer and llama.cpp is due to:

  • llama.cpp's more aggressive optimizations (flash attention, tiled GEMM)
  • Better memory layout (transposed K cache)
  • Multi-threading (minfer is single-threaded)
  • Model size differences (llama.cpp benchmarks often use larger models where amortization is better)

Remaining Optimization Opportunities

Superseded in part (2026-10-08). P2, P3 and P4 below have shipped since this section was written, and P1 — the K-quant AVX2/AVX-512 dots — landed 2026-10-08 (#56); only its weight-repacking half remains. See the status banner at the top of this document.

P1: K-quant AVX2 / AVX-512 Dot Products — DONE 2026-10-08 (#56)

Impact: Huge for Q4_K/Q5_K/Q6_K models (all matmul operations). Difficulty: High (requires SIMD bit manipulation).

Current state (closed): Q4_K/Q5_K/Q6_K × Q8_K have AVX2+FMA kernels in src/quants/avx2.rs and AVX-512/VNNI variants in src/quants/avx512.rs, dispatched AVX-512 → AVX2 → scalar at runtime (MINFER_NO_AVX512=1 drops to AVX2, MINFER_NO_AVX2=1 to scalar — the x86 counterparts of MINFER_NO_NEON). quants::avx2_correctness gates each kernel bitwise against its *_scalar reference, and its #[ignore]d kquant_simd_dot_speedup harness records the dot ratios on a 4096-element row: AVX2 3.39× / 2.80× / 1.79×, AVX-512 3.97× / 3.23× / 2.51× scalar for Q4_K / Q5_K / Q6_K (Intel i7-11700B, 2026-10-08). minfer bench on the cached Qwen2.5-0.5B Q4_K_M (16 threads, median of nine -r 3 samples, --n-ctx 656) moves prefill pp512 28.59 → 32.13 tok/s and decode tg128 11.20 → 12.54 tok/s; the modest end-to-end delta is the weight-row re-read per token, which the remaining weight-repacking increment addresses.

llama.cpp approach: ggml/src/ggml-cpu/arch/x86/quants.c uses the denibble()/maddubs AVX2 shape and the VNNI dpbusd AVX-512 shape; minfer's kernels follow the same structure and keep the scalar float order so the result is bitwise, not merely close.

P2: f16 KV Cache

Impact: 2× memory bandwidth reduction during attention. Difficulty: Medium.

Current state: the KV store is the graph allocator's persistent regions, whose width is the per-engine KvFormat (graph/kvformat.rs; f16 on the GPU backends, packed q8_0 on CPU + CUDA + Metal (#310), f32 default — MINFER_CACHE_TYPE). The pre-graph src/cache.rs Vec<f32> this line used to name was deleted in #252.

llama.cpp approach: Default is GGML_TYPE_F16, reducing memory by half.

Recommendation: Add Vec<half::f16> option. The half crate is already a dependency. Attention dot product would need f16→f32 widening.

P3: Flash Attention

Impact: 10-15% for long sequences, grows with context length. Difficulty: High.

Current state: Multi-pass softmax with strided KV access.

llama.cpp approach: Tiled flash attention with online softmax (ggml/src/ggml-cpu/ops.cpp:8347-8871).

Recommendation: Low priority for decode (short sequences). Only beneficial for prefill or long context. The current implementation is already efficient for typical use cases.

P4: Multi-Threading

Impact: Linear speedup with core count (expected +50-70% on 8-core CPUs). Difficulty: Medium.

Current state: no parallel libraries (rayon, crossbeam) in Cargo.toml, so a forward is one thread. The &mut KVCache this line used to blame was a dead parameter, not a constraint: it was never read and was deleted in #252, and the KV store the graph really uses lives inside the allocator (so top-level parallelization is bounded by the allocator's ownership, not by a caller-held cache handle).

llama.cpp approach: OpenMP parallelism across layers and attention heads.

Parallelization Opportunities

1. Attention Heads Parallel (Highest Priority)

Location: forward.rs:287 in gqa_attn function

#![allow(unused)]
fn main() {
for h in 0..nh {  // ← Can be parallelized!
    let hk = h / gqa;
    for t in 0..nt {
        // Each head writes to different output slice: out[os..os + hd]
        // Reads same Q, K, V caches (read-only)
    }
}
}

Why it can be parallel:

  • Each head writes to independent memory region (out[h*hd .. (h+1)*hd])
  • All heads read same Q/K/V caches (read-only access)
  • Qwen2-0.5B has 16 heads, can use 16 threads

Expected gain: 4-6× speedup on attention portion (30-40% overall)

2. Q/K/V Matmul Batch Parallel

Location: forward.rs:95-99

#![allow(unused)]
fn main() {
crate::kernel::quant_matmul_f32_batch(&mut [
    (l.wq.as_ref().unwrap(), &mut bq, nqt),  // Independent output
    (l.wk.as_ref().unwrap(), &mut bk, nkt),  // Independent output
    (l.wv.as_ref().unwrap(), &mut bv, nkt),  // Independent output
], &bn, ne, nt);
}

Why it can be parallel:

  • Three matmuls write to different buffers (bq, bk, bv)
  • All read same input bn (read-only)
  • Currently sequential, can run 3 threads in parallel

Expected gain: 2-3× speedup on matmul portion (10-15% overall)

3. FFN Gate/Up Parallel

Location: forward.rs:118-121

#![allow(unused)]
fn main() {
crate::kernel::quant_matmul_f32_batch(&mut [
    (l.ffn_gate.as_ref().unwrap(), &mut bg, nf),  // Independent
    (l.ffn_up.as_ref().unwrap(),   &mut bf, nf),  // Independent
], &ffn_in, ne, nt);
}

Why it can be parallel: Same as Q/K/V, two independent matmuls

Expected gain: 2× speedup on FFN matmuls (5-10% overall)

What Cannot Be Parallelized

1. Layer Loop (Sequential Dependency)

#![allow(unused)]
fn main() {
for il in 0..model.n_layer() {  // ← Must be sequential
    // layer[il] needs hidden output from layer[il-1]
    // Cannot parallelize different layers
}
}

2. KV Cache Writes

#![allow(unused)]
fn main() {
kv_cache.layers[il].store_multi(positions, &bk, &bv);
}

If multiple threads write to KV cache simultaneously, need locks or atomic operations, which would negate parallelization benefits.

Implementation Recommendations

Option 1: Rayon (Simplest)

# Cargo.toml
rayon = "1.10"
#![allow(unused)]
fn main() {
// forward.rs - modify gqa_attn
use rayon::prelude::*;

fn gqa_attn(...) {
    (0..nh).into_par_iter().for_each(|h| {
        // Each head's computation
        for t in 0..nt { ... }
    });
}
}

Pros: Minimal changes, automatic thread pool management Cons: Need to ensure output buffers don't conflict

Option 2: Crossbeam (More Control)

crossbeam = "0.8"
#![allow(unused)]
fn main() {
// Manual thread control
crossbeam::scope(|s| {
    for h in 0..nh {
        s.spawn(|_| {
            // Each head's computation
        });
    }
});
}

Pros: Fine-grained control Cons: Manual thread lifecycle management

Option 3: Pre-allocated Thread Pool (Highest Performance)

#![allow(unused)]
fn main() {
use std::sync::Arc;
use std::thread;

struct ThreadPool {
    workers: Vec<thread::JoinHandle<()>>,
    // ...
}

// Initialize once in main.rs, reuse across entire inference
}

Pros: Avoids thread creation overhead Cons: Complex implementation

Expected Performance Gains

Based on current 27.1 tok/s baseline:

OptimizationExpected GainNew Performance
Attention heads parallel (16 heads → 8 threads)+30-40%35-38 tok/s
+ Q/K/V parallel+10-15%38-43 tok/s
+ FFN parallel+5-10%40-47 tok/s
Total+50-70%40-46 tok/s

This would close the gap with llama.cpp from 2.5-3× to 1.5-2×.

Recommendation: Start with Rayon for attention heads parallelization. This is the easiest change with the highest impact. Then add Q/K/V and FFN parallelization if needed.


Bottleneck Analysis (Updated)

Current Bottlenecks (in order of impact)

  1. Memory bandwidth (70% of time): Loading Q4_0 weights from memory.

    • Mitigated by: Q4_0 quantization (4× compression vs f32)
    • Further optimization: Q4_K AVX2 (P1), better prefetching
  2. Attention computation (15% of time): KV cache access + softmax.

    • Current implementation is already efficient for short sequences
    • Further optimization: Flash attention (P3), f16 KV cache (P2)
  3. Element-wise operations (10% of time): RMSNorm, SiLU, residual add.

    • Already optimized with SIMD
    • Further optimization: marginal gains only
  4. Quantization (5% of time): f32→Q8_0 conversion.

    • Already optimized in quants.rs
    • Further optimization: activation reuse (but this caused regressions)

Comparison with llama.cpp

Why llama.cpp is 2.5-3× faster

  1. Multi-threading: llama.cpp uses OpenMP for parallel execution across layers and attention heads. minfer is single-threaded. This alone accounts for ~2× difference on modern CPUs with 8+ cores.

  2. Flash attention: Tiled flash attention with online softmax reduces memory traffic and improves cache utilization. minfer uses multi-pass softmax.

  3. Transposed K cache: llama.cpp stores K cache transposed (K_f32[dk][kv]) so the KV dimension is contiguous for SIMD. minfer uses position-major layout with strided access.

  4. Better matmul kernels: llama.cpp uses tinyBLAS with register blocking and microkernels for matmul. minfer uses simple nested loops.

  5. More aggressive quantization: llama.cpp supports Q4_K, Q6_K, Q8_0 for weights and KV cache. minfer primarily uses Q4_0.

What minfer does well

  1. Clean Rust code: No unsafe code in most places, strong type safety.
  2. Simplicity: Easy to understand and modify.
  3. Good baseline optimizations: RMSNorm fusion, RoPE cache, vectorized element-wise ops are already present.
  4. Fast for small models: 27.1 tok/s on 0.5B model is reasonable for single-threaded inference.

Recommendations

Short-term (1-2 weeks)

  1. Implement Q4_K AVX2 dot product (P1)

    • Highest impact for Q4_K models
    • Reference: llama.cpp quants.c:1900-2076
    • Expected gain: 2-3× for Q4_K matmul operations
  2. Add multi-threading (P4)

    • Use rayon or crossbeam for parallel layer execution
    • Expected gain: 2-4× on 8-core CPUs

Medium-term (1 month)

  1. Add f16 KV cache (P2)

    • Reduces memory bandwidth by 2×
    • Expected gain: 10-20% for attention-bound workloads
  2. Implement flash attention (P3)

    • Only beneficial for long sequences
    • Expected gain: 10-15% for prefill, minimal for decode

Long-term (2+ months)

  1. Tiled GEMM with register blocking

    • Replace simple nested loops with microkernel approach
    • Reference: llama.cpp tinyBLAS (sgemm.cpp)
    • Expected gain: 20-30% for matmul operations
  2. AVX-512 support

    • 16-wide f32 operations (vs 8-wide AVX2)
    • Requires runtime detection and fallback
    • Expected gain: 1.5-2× for compute-bound operations

Conclusion

Superseded (2026-09-15). This conclusion predates the 2026-08 NEON/threading work and the compute-graph rewrite: minfer is no longer single-threaded, and the CPU path described as "production-ready for single-threaded use cases" has since gained a persistent thread pool and Q8_K activation quantization. Read §"NEON + Threading Overhaul" for the current state.

The minfer codebase is already well-optimized for single-threaded CPU inference on small models. The initial analysis incorrectly identified "scalar hotspots" as bottlenecks, but these were already optimized with SIMD operations.

All optimization attempts (activation quantization reuse, online softmax, RoPE cache stack allocation) caused performance regressions due to:

  • Function call overhead
  • Interference with compiler auto-vectorization
  • Extra memory traffic

The performance gap with llama.cpp (~2.5-3×) is primarily due to:

  1. Multi-threading (llama.cpp uses OpenMP)
  2. Flash attention with tiled execution
  3. Better memory layout (transposed K cache)
  4. More aggressive quantization support

Future optimizations should focus on:

  • Q4_K AVX2 dot product (highest impact for Q4_K models)
  • Multi-threading (easiest way to close the gap)
  • f16 KV cache (reduces memory bandwidth)

The code is production-ready for small models and single-threaded use cases. Further optimizations require significant engineering effort and are only justified for larger models or multi-core servers.


NEON + Threading Overhaul (2026-08, Qwen3-4B Q4_K_M on M4 Pro)

The CPU path went from 1.1 tok/s decode (single-threaded scalar) to ~52–58 tok/s at -t 8 (~80–90 % of llama.cpp's 63–69), and prefill from ~1.1 to ~70–75 tok/s. Full analysis: docs/PERF-QWEN3-4B-VS-LLAMACPP.md §3.

What was added

  1. aarch64 NEON/SDOT dot kernels (src/quants.rs) — all 8 quantized dot products (Q4_0/Q4_1/Q8_0/Q5_0/Q5_1/Q5_K/Q4_K/Q6_K) use the ARMv8.2 sdot instruction (16 MACs/instr) via stable inline asm (vdotq_s32 is unstable in std::arch). Bit-exact with the scalar kernels (exact int32 accumulation, per-block float ops in identical order); MINFER_NO_NEON=1 reverts. The x86 counterpart since #56 is MINFER_NO_AVX2=1 (the whole quants AVX2 layer) with MINFER_NO_AVX512=1 dropping just its AVX-512/VNNI layer to AVX2.
  2. Q8_K activations for K-quant matmuls (src/block.rs Q8KB, src/quants.rs quantize + dots, src/kernel.rs dispatch) — llama.cpp's activation format: 256-element blocks with precomputed int16 per-subblock sums, so dots never re-reduce the activation and use one scale per 256 elements. Kernels restructured to llama's shape (scales applied to the SDOT vectors via vmlaq_n_s32 before the single horizontal reduce).
  3. Persistent CPU thread pool + row-parallel matmuls (src/kernel.rs) — per-call thread::scope spawns cost ~170 µs (measured), so a persistent pool (atomic generation handoff, spin-then-yield workers, main thread participates as the last worker) dispatches each matmul in ~1–3 µs. Output is bit-identical to single-threaded (each row computed by one worker). Generic par_for extends it to other ops (attention over heads).
  4. NEON vector ops (src/vec_ops.rs) — vec_dot/muladd/scale/add and an in-place softmax (polynomial exp); the attention/norm path was fully scalar on ARM before.
  5. CLI -t/--threads <N> (default = macOS P-core count via hw.perflevel0.logicalcpu; measured best 8 for the 4B; E-cores hurt, matching llama).

Pitfalls hit

  • vld1q_s16 loads 8 lanes only — loading 16 bsums needs two loads; vpaddq_s16(a, b) yields 8 lanes (a's 4 pairs then b's 4 pairs), and vget_high_s16 on an 8-lane array read undefined stack bytes. Both caused silent garbage; caught by NEON-vs-scalar unit tests on random data.
  • int16 vmlal_s8 accumulation overflows for Q8_0 (127×127×4); SDOT's int32 accumulation avoids the whole class.
  • Workers spinning between matmuls show up as ~50 % idle in profilers — that is the main thread's serial work (quantize, dispatch, attention), not a kernel problem; the attention-parallelization and NEON quantize closed most of it.

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0013 — The CPU quantizes activations to Q8_0; a device reads f32
  • ADR-0007 — No ML frameworks: every operator is hand-written

Metal Backend Optimizations

Goal: match llama.cpp (Apple M4 Pro) performance. Current gap (2026-10-07 same-model, same-parameter A/B against llama.cpp c479922ac): decode is ~0.88-0.99× llama (7B Q4_K_M ≈ parity — 48.33 vs 48.59 t/s; 0.5B Q4_K_M 0.88×; Qwen3-0.6B Q8_0 0.89×) and long prefill is ~0.86-0.93× (0.5B Q4_K_M 0.86×; 7B Q4_K_M 0.93×; Qwen3-0.6B Q8_0 0.86×). The full protocol and every raw number are in §0.2; measured at minfer 97cc0d7 on macbook (macOS 27.0.1, Apple M4 Pro). The 2026-08-17 record (llama.cpp 88b47a755) is superseded and kept as history in §1.1. The re-measured minfer-only rows (throughput on the cached models + Metal-vs-CPU parity) live in §0.1, taken 2026-10-06 at 6b95763. See §1 Current state.

⚠️ The §0 progress table is the single source of truth for tracking; §1-§6 are the detailed explanations behind it. Update §0 first before changing code.

⚠️ Commit hashes in this document predate a repository history rewrite and do not resolve. Measured on 2026-10-10: 44 of the 47 hash-like tokens here fail git cat-file -t, and the three that pass are not this repository's commits anyway (llama.cpp's own c479922ac and 88b47a755, and an AppleClang version). They are kept as history, not as citations. Subject-matched equivalents from the current history — each verified with git cat-file -t — are listed in METAL-BACKEND-DESIGN.md §5.2; cite the equivalent and that section together, and treat a bare hash in this file as a pre-rewrite label rather than a locator.

📌 Compute-graph era note (2026-08-21+, Phase 6): inference now runs through the declarative compute graph (src/graph/, MetalBackend), which dispatches these kernels per op instead of the old whole-layer layer_gpu. Every kernel/optimization below remains valid — the gap is in wiring, not kernels: see §0.1 Graph-path (MetalBackend) integration status. The old-path measurements in §1 are pre-graph references; graph-path numbers live in §0.1.

📖 Backend design of record: METAL-BACKEND-DESIGN.md (MetalBackend lifecycle, dispatch, command buffers, model wiring, safety). This document stays the optimization ledger; commit hashes printed below predate a history rewrite and no longer resolve — see the design doc's §5.2 for subject-matched equivalents.


§0 Progress Overview (single tracking source)

Legend: [x] done (with commit) · [ ] to-do · [—] decided not to change. Commits are from git log; a few early items are marked "TODO-trace" (long history, to be filled in later).

§0.1 Graph-path (MetalBackend) integration status (2026-08-21)

The graph's MetalBackend (src/graph/metal_backend.rs) dispatches one op per call into a split-scoped MpsCommandBuffer. Which of the optimizations below already transfer:

Graph opKernel used by MetalBackendStatus
MatMulquant_matmul_f32_on_gpu_buf → per-type GEMM / _multi / ported q4_K kernel✅ full transfer (#11/#12/#27/#29/#30/#40 apply)
GetRows (embed)embed_tokens_gpu → get_rows kernels✅ full transfer (#33/#38 apply)
GetRows (tail n_out)get_rows_f32 kernel (dispatch_2d gather, G3)✅ G3 — last layer runs on the tail n_out rows (#32/#34 semantics, plan §5.5)
KvcacheStorestore_kv✅
RoPE / Silu / Add / Mul / SwiGLUrope_f32 / silu_f32 / add_f32 / mul_f32 / swiglu_f32✅
RmsNormrms_norm_256 (G2)✅ G2 — rms_norm_256_enabled() (#16, ~2× faster decode dispatch cost)
Attn (prefill + decode)flash / split / parallel / classic, gated like the old path (G1)✅ G1 — nt==1 → flash_attn_enabled(hd) → gqa_attn_flash (chunked) else gqa_attn_split_f32 else classic; nt>1 → hd 64/128 + prefill_flash_enabled(hd) → attn_flash_prefill else matmul_attn_enabled() → attn_parallel_prefill else classic (#9/#10/#13/#17/#22/#24/#25/#26)
FusedQKV (decode QKV, G4)concat matmul (blk.{i}.attn_qkv) + attn_bias_rope_store✅ G4 — nt==1 builds one fused node replacing 3 matmul + 3 bias + 2 rope + 2 store (10 dispatches → 2); q at concat offset 0 feeds attention; MINFER_NO_FUSE_QKV=1 reverts
FusedFFN (decode gate+up, G5)concat matmul (blk.{i}.ffn_gu) + in-place swiglu_f32_off✅ G5 — nt==1 builds one fused node replacing 2 matmul + silu + mul (4 → 2 dispatches); down reads rows 0..nf; gated nf ≤ 16384 (7B Q4_K concat matmul slower); MINFER_NO_FUSE_FFN=1 reverts
Fused decode QKV / bias+rope+store (attn_bias_rope_store)— (the deleted layer_gpu fused kernels)⬜ superseded: the graph's Op::FusedQKV covers the unmixed case (row above) and the mixed-quant epilogue is CUDA-only by decision (#52); this row described the pre-graph command path and is kept only as history
FusedQkvNorm (Qwen3 decode, per-head Q/K RMSNorm)concat matmul (blk.{i}.attn_qkv) + rms_norm_256 (q/k in place) + attn_rope_store✅ Qwen3 Op::FusedQkvNorm (2026-08-27) — decode fuses 3 matmul + 2 qk_norm + 2 rope + 2 store → 1 concat matmul + 2 per-head norm + 1 no-bias rope+store; Qwen2's attn_bias_rope_store path untouched
n_out tail-row optimization (#32/#34)tail GetRows + reduced last layer (G3)✅ G3 — full-nt FFN/lm_head work dropped (prefill ↑~1.5× on 0.5B)

Graph-path measured numbers (M4 Pro, --temp 0 greedy; old-path figures from §1/§1.6 — same models; the graph-path column was re-measured 2026-10-06 at 6b95763 on macbook (macOS 27.0.1, Apple M4 Pro), minfer bench -p <P> -n 128 -r 3 <model> with the Metal default except the CPU row (MINFER_DISABLE_MPS=1 -p 0). The pre-round figures it replaces were taken before the Metal round; the Δ column is against the old path, not against the pre-round figure):

ScenarioOld path (layer_gpu)Graph path (2026-10-06)Δ
0.5B Q4_0 decode, KV ~200+~279 t/s (§1.1, 128 tok)306.19 ± 1.00 t/s (tg128, pp440, n_ctx 584)≈ parity→+~10 % over old
0.5B Q4_0 prefill pp390–440~2530–2620 t/s (§1.1 pp430)6249.60 ± 10.52 t/s (pp440, n_ctx 584)+~140 % over old
7B Q4_K_M decode, KV ~200+~48 t/s steady48.52 ± 0.18 t/s (tg128, pp206, n_ctx 350)≈ parity
7B Q4_K_M prefill pp206~240 t/s (§1.6 pp252)406.51 ± 0.83 t/s (pp206, n_ctx 350)+~70 % over old
0.5B CPU decode~5.9 t/s149.69 ± 1.79 t/s (tg128, pp0, n_ctx 144)the old ~5.9 t/s is not reproduced — see the note below
Qwen3-4B Q4_K_M decode—74.11 ± 0.19 t/s (tg128, pp241, n_ctx 385)llama-Metal 79.7 → ≈ parity
Qwen3-4B Q4_K_M prefill steady-state—729.83 ± 0.20 t/s (pp241)llama ~900 → ≈ 0.81×
Qwen3-4B first-request prefill (241 tok)—~330 ms (241 / 729.83)was 552 ms

Parity (re-measured 2026-10-06). graph_metal_matches_llama_reference (Qwen3-0.6B) reproduces the pinned llama-Metal greedy prefix [12095, 13, 576, 6722, 315, 9625, 374, 1083, 279] — the one model-level oracle that compares Metal against an external reference. The per-op Metal-vs-CPU gates (metal_*_matches_cpu in src/graph/metal_backend/tests.rs) and the four kernel isolation suites are green in the full macOS run (532 / 0 / 43 unit, 21 / 0 / 6 integration; the +1 is #329's dispatch-refusal gate). The Metal-vs-CPU logits row cannot be taken from graph_metal_matches_cpu_logits: in the graph era both ModelDef::forward and forward_graph route through Qwen2Graph::forward (src/models/qwen2/mod.rs:56/:96), so its max |Δ| = 0 compares the Metal graph with itself — the test's "CPU reference" comment is stale. Filed as #324; no parity regression is implied. The f16/bf16 device-gate max |Δlogit| numbers are in docs/SUPPORT-MATRIX.md (their model files are not cached for a fresh run here).

The old ~5.9 t/s CPU row is not reproduced at 6b95763: the same 0.5B Q4_0 file decodes at 149.69 ± 1.79 t/s on the CPU with -p 0. A 25× gap is not measurement noise, so the old figure describes a different configuration (it is quoted from the pre-graph §1 table and may predate the CPU kernel work); it is left visible in the Old path column but must not be read as the current CPU number.

Reading: G1 (attention dispatch: flash/split/parallel) closed the decode KV-growth regression (−44 %/−32 % → parity; 0.5B 2.1×, 7B ~1.5×); G3 (n_out tail-row reduction) cut the full-nt last-layer FFN + lm_head work so 0.5B prefill now exceeds the old path. Remaining graph-path gap: 7B prefill is now well ahead of the old path too; the residual llama gap (attention is not the bottleneck there — GEMM dominates and transfers fully) is the 2026-08-17 figure above and was not re-measured. Greedy outputs are byte-identical to the pre-G1 graph path (flash/split/parallel and tail-reduced paths all verified).

Windowed prefill (#315 / #359 / #369). The explicit prefill — the batched multi-sequence path — first ran the correctness kernel families kernel_gqa_attn_window_f32/_f16 (one-range attn_span) and kernel_gqa_attn_map_f32/_f16 (set-valued kv_map) where a single-sequence prefill keeps the tuned attn_flash_prefill (src/metal/kernels/attn_window.metal). Measured on this Mac (2026-10-07, macbook (macOS 27.0.1, Apple M4 Pro), 512 tokens, bar 0.8x named before the run) the one-range path reached only 0.539x the causal prefill's tokens/s for a two-sequence batch on the 0.5B Q4_0 (f32, hd 64) and 0.159x on Qwen3-0.6B Q8_0 (f16, hd 128); at the same shape 0.376x (~2.7x slower) and 0.096x (~10.4x slower), and the map path (a shared prefix read in place) 0.229x / 0.046x. [#359] added the fast one-range family kernel_flash_attn_window_blk_* (src/metal/kernels/fa_window.metal, a copy of the causal attn_flash_prefill tile with an explicit [lo, hi) mask, selected for nt > 1 at hd ∈ {64,128}) and re-measured 0.968x / 0.918x primary and 0.996x / 0.995x shape-matched; [#369] added the fast map family kernel_flash_attn_window_map_* (the same tile with a run-membership mask, fwin_map_has) and measured the map arm at 0.943x / 0.914x — all above the 0.8x bar, with MINFER_NO_WINDOW_FLASH=1 restoring the before numbers on both layouts as the A/B control. A map whose runs are far apart pays the gap between them (the kernel tiles the global union); the first-fit shared prefix the server emits is adjacent and does not. Protocol, node/kernel counts and the bar are in METAL-BACKEND-DESIGN.md §4.4.1.

Graph-path TODO (wire into MetalBackend): ① ✅ G1 attention dispatch; ② ✅ G2 rms_norm_256; ③ ✅ G3 n_out tail-row GetRows; ④ ✅ G4 fused decode QKV + bias+rope+store; ⑤ ✅ fused decode FFN gate+up (Op::FusedFFN); ⑥ ✅ G6 Qwen3 fused decode QKV with per-head Q/K RMSNorm (Op::FusedQkvNorm, 2026-08-27) — decode 3 matmul + 2 qk_norm + 2 rope + 2 store → 1 concat matmul

  • 2 per-head norm + 1 no-bias rope+store. G1–G6 are wiring (kernels already isolated-tested); G3 additionally fixed two allocator liveness bugs (liveness used topo_order while the scheduler executes in build order; input buffers were freed and reused before host fills finished). The FFN fusion is gated on nf ≤ 16384 — the Q4_K concat matmul (od ≈ 37888) is slower than two separate matmuls on the 7B decode scalar kernel, so only the 0.5B class fuses FFN.

§0.2 Refreshed llama.cpp A/B (2026-10-07)

The §1.1 (2026-08-14/08-17) A/B against llama.cpp 88b47a755 was re-measured on the Mac against a HEAD llama.cpp build, c479922ac (build 11458, AppleClang 21.0.0.21000334, CMake Release -O3 -DNDEBUG; CMAKE_C_FLAGS_RELEASE additionally -ffp-contract=fast), at minfer 97cc0d7. Build: cmake --build ~/git/reading/llama.cpp/build-fpc --target llama-bench llama-cli -j. This is outcome 1 of #331: the pair is refreshed, not merely re-worded.

Protocol (gate-contract rules 3 and 5). Same GGUF file for both engines, from /Volumes/WD_BLACK/models/hf/Qwen/; same model, same prompt length and context for both. minfer:

./target/release/minfer bench -p P -n 128 -r 3 <model>                       # Metal default
MINFER_CACHE_TYPE=f16 ./target/release/minfer bench -p 430 -n 128 -r 3 <Q8_0 model>

llama.cpp (the 2026-08-17 pair used minfer --greedy vs llama-bench -b 512 -t 8; kept here — minfer bench is the current harness: pp times the prefill alone, and tg runs P prefill tokens and then times only the n_gen decode steps (the prefill happens before the timer in src/bench.rs's decode_once), which is llama-bench's tg caliber — the direct analogue of the old Generated: pure-decode number, not a prefill-inclusive average; both harnesses do a warmup pass before the measured reps and report a mean over -r 3):

~/git/reading/llama.cpp/build-fpc/bin/llama-bench -m <model> -p P -n 128 -r 3 -b 512 -t 8

Ambient state. macbook (macOS 27.0.1, Apple M4 Pro), hostname macbookpro-ysw, AC power (battery 80 %, not charging), no recorded thermal or performance warning; the box was otherwise loaded (Microsoft Edge + opencode/agent active, 15-min load average ≈ 7). Both engines ran back-to-back in one session, so the ratio is the robust quantity; each engine's absolute t/s carries the shared ambient load. Each cell is the harness mean ± stddev over the 3 measured reps.

Qwen2.5-0.5B-Instruct Q4_K_M (qwen2.5-0.5b-instruct-q4_k_m.gguf):

Testllama.cpp c479922ac (t/s)minfer 97cc0d7 (t/s)minfer / llama
pp430 (long)6781.16 ± 280.145827.26 ± 51.130.86×
tg128 after pp430298.52 ± 13.44262.68 ± 3.200.88×
pp30 (short)2844.87 ± 100.802053.06 ± 8.420.72×
tg128 after pp30286.04 ± 7.41277.09 ± 1.670.97×

Qwen2.5-7B-Instruct Q4_K_M (split GGUF, entry part …-00001-of-00002.gguf):

Testllama.cpp c479922ac (t/s)minfer 97cc0d7 (t/s)minfer / llama
pp495 (long)469.79 ± 1.89438.61 ± 1.580.93×
tg128 after pp49548.59 ± 1.6048.33 ± 0.310.99×
pp30 (short)337.77 ± 0.30285.93 ± 0.800.85×
tg128 after pp3048.76 ± 0.2647.95 ± 0.770.98×

Qwen3-0.6B Q8_0 (f16 KV cache both sides; minfer MINFER_CACHE_TYPE=f16, llama-bench default -ctk f16 -ctv f16):

Testllama.cpp c479922ac (t/s)minfer 97cc0d7 (t/s)minfer / llama
pp430 (long)5645.83 ± 160.284881.25 ± 10.440.86×
tg128 after pp430235.97 ± 5.40210.67 ± 0.830.89×

For the Q4_K_M rows minfer's auto KV rule keeps the 0.5B class on f32 and the 7B class on f16 while llama-bench is f16 for both; this is stated, not corrected (the 2026-08-17 record had the same asymmetry).

Reading. The 2026-08-17 "long prefill ~2.6-2.7×" no longer holds: long prefill is now 0.86-0.93× (1.07-1.16× slower) and decode is 0.88-0.99× (7B at parity). Short-prompt pp30 stays dispatch/overhead-bound on both engines and is not a clean attention lever (0.72-0.85×).

§0.3 Small-batch prefill re-measurement (#40, 2026-10-09)

#40 closed with "the nt∈[2,8] small-od gap vs llama's ext kernel (0.5B pp4 141 vs 583) remains open — low value (2-8 token prompts), deferred": a 0.24× A/B taken 2026-08-21 against a then-HEAD llama.cpp. That figure is stale, and a reader of §0 could not tell the deferral from an oversight. Re-measured 2026-10-09 at minfer 270b6b3 on macbook (macOS 27.0.1, Apple M4 Pro), hostname macbookpro-ysw, AC power (~79 %), the box otherwise loaded (Edge + agent; 15-min load ≈ 7-8), against the §0.2 llama.cpp reference c479922ac (build 11458, AppleClang 21.0.0.21000334, -O3 -DNDEBUG -ffp-contract=fast).

Protocol. Same GGUF for both engines, from /Volumes/WD_BLACK/models/hf/Qwen/Qwen2.5-0.5B-Instruct-GGUF/: qwen2.5-0.5b-instruct-q4_k_m.gguf (the §0.2 model) and …q4_0.gguf. §0 #40 recorded the pp point ("pp4") but not the tg, so the §0.2 harness is used whole. Each engine ran its own contiguous sweep (minfer first, then llama) from a settled box, -r 5:

./target/release/minfer bench -p P -n 128 -r 5 <model>                      # Metal default
~/git/reading/llama.cpp/build-fpc/bin/llama-bench -m <model> -p P -n 128 -r 5 -b 512 -t 8
ppminfer k_mllama k_mratiominfer q4_0llama q4_0ratio
2230.5352.50.65279.1501.10.56
4306.1917.70.33533.2821.70.65
8446.81658.10.27884.31679.00.53
302258.22136.71.062432.82703.30.90

Reading. The recorded 0.24× is stale: minfer's pp4 roughly doubled since #40 and the ratio at pp2-8 is now ~0.3-0.65×. Repeats spread that band — the tiny-pp timing is a few ms of per-call setup, so the session's ambient dominates (llama's own pp30 moved 2845 → 2137 between §0.2 and this run; each engine's pp4 ratio across four sessions fell in 0.33-0.52 for k_m and 0.52-0.65 for q4_0). At pp30 the engines are within ~±15 %. The gap is real and lives exactly where #40 left it: the nt∈[2,8] small-od matmuls still take the serialized _multi kernel where llama runs kernel_mul_mv_ext (the batched matvec), because the GEMM gate is nt >= 2 && (od >= 2048 || nt >= 9) (src/metal/ops.rs:40), so a 0.5B's attn_qkv (od 1152) stays on the multi path through nt 8; the GEMM does kick in at nt >= 9, which is why the ratio improves through pp12-16.

Decision — kept deferred (2026-10-09). The fix is a mul_mv_ext-style batched-matvec kernel family (comparable surface to the q4_K decode port of #27) that speeds up only 2-8 token prompts, a window real chat/system prompts never occupy: the engines are already at 0.86-0.93× at pp30-495 (§0.2). The cost buys a few milliseconds of one-time prompt processing for single-word-style inputs and does not justify the kernel surface. Revisit trigger: a workload that actually runs 2-8 token prompts at volume (single-token completions, an embedding/similarity path, or speculative-draft prefill), or a batched-matvec kernel landing anyway for a larger reason. The work is filed as #431 so the finding is schedulable and searchable; the trigger above still decides when it is worth starting. Until then this is a decision, not an oversight.

✅ Done

#ItemMeasured effectCommit (pre-rewrite)
1Metal backend foundation + 4 correctness fixes (RoPE freq_scale / output_b / softmax max / stack array)Qwen2-0.5B 130→334 t/s2473981 / 26f0e4d (early; TODO-trace)
2Q5_K formula (unsigned) + qh index fixQ5_K_M CPU/GPU output correctpre-3f23560 (TODO-trace)
3Q5_1 / Q5_K Metal matmul kernelsQ5_K_M fully GPU26f0e4d / later (TODO-trace)
4Q4_0 → f32 activations (aligned with llama Metal, removed Q8_0 quantize)decode +5-10 %ba51f68
5GQA attention simd_max divergence fix (partial tiles)long prefill output correct28d4ba2
6GPU-hang safety hardening (bounded wait + dispatch trace + barrier guards)deadlock → error-exitbff73db
7Metal cb/encoder autorelease retain fixbg-thread cb no longer assertsb1256d5
8Fused QKV + FFN gate/up matmuls (decode, nt==1)decode ~5 % (24 % was a GPU-state artifact, corrected in 26b145b)6f0c847
9KV-parallel split attention (decode, 2-pass)200-token 1.56→1.06 s (~32 %)b3d4c7a
10Attention float4 acc + adaptive chunks + KV geometric growthextra ~15 % + long-context fix66f4290
11simdgroup GEMM: non-Q4_0 quants (Q8_0/Q5_0/Q5_1/Q4_K/Q5_K/Q6_K)K_M prefill 300→650 t/sc9f865c / 2c03bd1
12Q4_1 simdgroup GEMMevery quant has a GEMM5b914f0
13f16 split attention (MINFER_CACHE_TYPE=f16)f16 decode 1.60→0.95 s387d612
14float4 elementwise + parallel RoPE (P6/P7)200-token ~0.88→~0.80 sddd3eb0
15CPU sampler speedup (top_k O(n) / top_p sorts only survivors)sampling ~12.6-14.8→~5.5-6.5 ms/token (2×)192378d
16256-thread RMSNorm + per-kernel profile + KV-growth fixes (chunk cap 32→16, drop sync_kv_to_cpu)~3 % decode + long-context ~0.25 ms/tokena7f21e4
17Parallel prefill attention (3-pass, barrier-free)pp430 212→144 ms (~32 %); 7B 944→832 msb2c97fd
18Generated: pure-decode caliber + dual-caliber bench.shmeasurement credibility fixdc66d0d
19Q4_K AVX2 test-reference fix (implementation was already correct)29 bin tests green266ffb7
20Split-GGUF (7B multi-part) support7B loads and runscbba68c / 34eaf10
21Same-model, same-parameter A/B benchmark docgap baseline established09d27ae
22Flash attention port (llama kernel_flash_attn_ext_vec, NSG=1 fixed DK=DV=64)decode GPU 0.25-1.0 ms/token faster; wall ~10 % (0.5B, byte-identical)2e0c8b3
23Prefill gap root-cause (2026-08-14): GEMM is NOT the prefill leversee to-do #4 status—
24Prefill flash attention port (llama kernel_flash_attn_ext_blk, legacy simdgroup-matrix, NSG=4 fixed DK=DV=64)prefill GPU ~110→~93 ms (~16 %); f32+f16 byte-identical to classic5974eb1
25hd=128 (7B) prefill flash port (2026-08-15, §3.4 done)7B pp310 prefill GPU 1042→~949 ms (~9 %), f32/f16 byte-identical, fixes 7B f16-cache garbage§3.4
26hd=128 (7B) decode flash port (2026-08-17, §3.3 extend)7B decode steady GPU ~51.1 ms vs split ~51.6 ms (pp205), f32+f16+split byte-identical§3.3
27q4_K decode matmul layout port (7B) (2026-08-17, to-do #7, §3.3)7B q4_K dims 70→265 GB/s (attn_q), 18→146 (attn_k), 74→243 GB/s (ffn_g/u); 7B decode steady GPU ~51 → ~19.3 ms/token (~2.6×, now ≈ llama's 50.5 t/s); 7B/0.5B byte-identicalto-do #7
28GEMM partial-tile race + missing Metal memoryBarrier (2026-08-19, §3.6): (a) all 8 simdgroup mm kernels lacked the threadgroup_barrier BEFORE the partial-tile temp_str stores — temp_str overlaps sa/sb, so a fast simdgroup overwrites sa/sb while a slow one still reads them → intermittently corrupted last-2-token logits (partial x-tile only); (b) the single prefill encoder had NO memoryBarrier between dispatches → RMSNorm write raced QKV read of the reused bn buffer (last-2 token slots, huge stale values)1.5B/7B first-token nondeterminism (~10-30 % wrong tokens, dump-localized) → 24/24 deterministic, output matches CPU byte-for-byte7253a3b
29mm-kernel hot-loop #pragma unroll (2026-08-19, §3.4): llama FOR_UNROLLs the staging/ik/load/mac loops; minfer's 8 mm kernels had none → added the 6 unroll points (llama-parity set)7B pp495 1438.8 → 1355.6 ms (~5.8 %), 0.5B ~6.9 %, 1.5B ~2.4 %; byte-identical + 24/24 determinismf3a499d
30ik-loop threadgroup_barrier → simdgroup_barrier(mem_none) (2026-08-19, §3.4 follow-up): .air diff showed the pre-unroll corruption was a rolled-loop compiler artifact, not a memory need; with the unroll in place llama's exact barrier form is now safe7B pp495 min 1387.4 → 1370.6 ms (~1.2 %); byte-identical (1.5B×24 / 7B×8 / 0.5B×3); removes the last structural mm-kernel difference vs llama0e756f3
31Phase-0 7B prefill decomposition (2026-08-20, §3.6): exact 7B MUL_MAT graph mapped (197 GEMMs, 7.000 TFLOP, wk/down/output = q6_K on the q4_k_m; CORRECTS the earlier "all q4_K except ffn_down=q6_K" assumption); llama GPU-busy measured by host timestamps (CB1 83 ms + CB0 964 ms ≈ 1043 ms @ pp495 = 6.71 TF clean window); every remaining factor refuted via an exact-shape replay harness (kernels A/B in-batch 6.21 vs 6.26 TF, grid/smem/buffer-mode/pooling/barriers free, interleave + 2-CB split hurt, weight data no effect, concurrent dispatch no benefit with the per-dispatch barrier)exact-shape replay (minfer kernel, real 7B shapes, one CB) = ~1126 ms = 6.20 TF, converging with llama under comparable system load; engine GEMM-only ~1240 ms (residual ~90-115 ms engine-vs-replay, unattributable); gap vs llama stays ~1.25× (consistent with §3.4)bd89eab
32lm_head / final-norm output-rows-only (2026-08-21, §3.7): llama computes the final norm + last-layer FFN + lm_head on n_outputs rows only (ggml_get_rows(cur, inp_out_ids) at qwen2.cpp:106-108; graph dump shows output GEMM out=[152064 1] and last-layer gate/down [.. 1]); minfer computed [152064×495]. §3.6's "7.000 TFLOP" was WRONG (assumed output N=495; correct llama total ≈ 6.26 TFLOP — the kernel-level "gap" was mostly this over-count). Fix: forward()/output_norm_gpu/CUDA all take n_out; final rms_norm + lm_head run on the tail n_out rows (n_out=1), logits buffer/download shrink 301 MB → 608 KB7B pp495 GPU ~1354 → ~1255 ms (stable; pre-change noisy 1354-2754), download ~150 ms → ~0.1 ms (wall −~150 ms); 0.5B GPU output byte-identical pre/post; 1.5B/7B greedy generation correct43989da
33GPU Q4_K embedding (get_rows) (2026-08-21): minfer's kernel_get_rows_q4_0 was Q4_0-only, so the 7B/1.5B (Q4_K token_embd.weight) fell back to CPU scalar dequant + 7 MB upload_hidden per prefill (O(nt) wall: ~5-15 ms @ 495 tok, ~100-200 ms @ 8K). Added kernel_get_rows_q4_k (reuses the validated dequant_q4_k_16, llama kernel_get_rows_q-equivalent) + pipeline + type routing in embed_tokens_gpu7B/1.5B embedding now on GPU; greedy output byte-identical pre/post (same seed); get_rows_q4_k_isolation test bit-exact vs CPU; no prefill GPU regressionuncommitted
34Last-layer FFN output-rows-only (2026-08-21, §3.7 follow-up): llama reduces cur + inpSA to n_out rows BEFORE the last layer's FFN (ggml_get_rows(cur, inp_out_ids) + get_rows(inpSA, inp_out_ids) at qwen2.cpp:106-108) — the last layer's ffn_norm, gate/up/down, swiglu and BOTH residuals run on 1 row. minfer ran layer-27's entire FFN on all nt. Fix: layer_gpu now takes n_out/is_last; the wo matmul stays on all nt (llama build_attn precedes the reduction), then the wo-residual, ffn_norm (byte-offset read of the hidden tail), gate/up/down (dispatched with nt=n_out via the new x_off matmul param), swiglu and the final residual all run on the tail n_out rows (add_f32_off); CPU path mirrors. minfer's total graph work drops 6.46 → ≈6.26 TFLOP — now exactly llama's total7B pp499 GPU ~1278 → ~1234 ms (~44 ms, ~3.4 %) — the doc's ~40 ms estimate confirmed; byte-identical (7B/0.5B GPU greedy + 0.5B CPU fallback, same seeds); decode unchanged (nt==1 path untouched); 33/34 bin tests green (1 pre-existing env-dependent failure: attn_parallel_realdata_correctness needs /tmp/dp3 dumps)b6ecbd3
35Precompiled metallib (2026-08-21, §4.2 to-do #1): build.rs compiles src/metal/kernels/ → embedded .metallib (xcrun metal -O3 + metallib, llama's exact flags; clang module cache redirected via -fmodules-cache-path since the default cache dir is unwritable under the build sandbox) → loaded with newLibraryWithData; empty-marker fallback to newLibraryWithSource when the toolchain is absent. Numerics verified byte-identical to the runtime source compile (7B @615/499 + 0.5B greedy, all 4 isolation suites, 33/34 bin tests) — a first apparent divergence was an A/B reference prompt-mixup, not a compiler difference; -O0 metallib lacks kernel_q4_1_f32_matmul (falls back to CPU), so -O3 only. Runtime hook MINFER_METALLIB_FILE=<path> loads an external metallib without rebuilds0.5B process wall ~1.32 → ~1.09 s (warm driver cache; the first-ever run benefits most — no per-process compile), prefill/decode perf unchanged (equal within noise); shader errors now caught at build timed09b8db
36GGUF mmap + zero-copy weights (2026-08-21, §4.2 to-do #2): std::fs::read (4.4 GB) + per-tensor extend_from_slice copy + GPU new_buffer+memcpy were THREE full-weight copies (~8.8 GB RAM + 4.4 GB GPU). Now: each part is mmap'd (MAP_PRIVATE, zero-dep raw mmap/munmap FFI, leaked for the process) and Tensor.data is a Cow<'static,[u8]> Borrowed slice of it (zero per-tensor copy; the output = tok_embd.clone() weight-tying fallback is now a shallow clone too); the Metal backend wraps each part with ONE page-aligned newBufferWithBytesNoCopy (register_part) and registers weights as (buffer, byte offset) into it — llama's exact design (ggml_metal_buffer_map page-aligns; newBufferWithBytesNoCopy requires a page-aligned base — per-weight NoCopy at 32-aligned bases was tried first and reads SHIFTED data on the GPU, hence the offset design). MINFER_WEIGHT_COPY=1 forces the old copy path for A/B7B load+prefill wall ~4.3 → ~2.7 s (warm; first-run 7.2 → 3.3 s); peak RSS 20.9 → 4.7 GB (~4.4×) — the weights are file-backed pages shared with CPU/GPU; 7B/0.5B greedy byte-identical + CPU fallback identical; prefill/decode perf unchanged (equal within noise); 34/35 bin tests green (1 pre-existing env-dependent)d51e8b8
37f16 KV auto-default (2026-08-21): set_kv_cache_type(n_layers, n_kv_embd) at model load auto-selects the GPU KV element type — f16 for the 7B class (n_layers×n_kv_embd ≥ 8192: KV bandwidth-bound decode), f32 for small models; MINFER_CACHE_TYPE=f16/f32 overrides. llama always defaults F16 (llama-context.cpp:3539); minfer had kept f32 default because 0.5B f16 measured ~3% slower (§0 decided-not #8) — the auto rule applies f16 exactly where it wins and keeps the 0.5B class on f327B @2K ctx steady decode f16 ≈ 20.15 ms vs f32 21.13 ms/token (~1 ms, ~5 %); 7B greedy byte-identical across auto/f32/f16 (16 tokens); 0.5B untouched (auto→f32); 34/35 bin tests greenthis commit 9099f79
38GPU get_rows for the remaining embedding types (2026-08-21): llama's kernel_get_rows_q covers every quant; minfer had Q4_0+Q4_K only (#33), so the 0.5B Q5_0/Q5_1/Q8_0/Q6_K-embedding models (q4_k_m embd=Q5_0, q5_k_m embd=Q5_1, q5_0/q8_0/q6_k models) fell back to CPU scalar dequant + upload_hidden per prefill. Added MSL templates kernel_get_rows_q32 (Q4_1/Q5_0/Q5_1/Q8_0, one thread per 32-elem block, reuses the validated dequant_*_16 helpers) + kernel_get_rows_q256 (Q6_K/Q5_K, one thread per 16-elem group, same structure as the Q4_K kernel) + pipeline routing/guards (ne%32 / ne%256) in embed_tokens_gputests/gemm_isolation.rs::get_rows_multi_type_isolation — all 6 new kernels bit-exact vs CPU (rel 0); end-to-end q5_0/q5_k_m/q8_0 GPU == CPU greedy (same seed); 0.5B q4_k_m (embd Q5_0) now embeds on GPU (was CPU fallback); 1.5B Q4_K / 7B Q4_K unchanged; 34/35 bin tests green. Note: the 0.5B q6_k model shows a PRE-EXISTING GPU-vs-CPU greedy divergence (reproduced on the pre-#38 binary — not from this change; its 896-dim Q6_K embd is ne%256≠0 → CPU fallback) — flagged for later investigationthis commit
39GPU warm-up read + merged embed (2026-08-21): (a) the mmap loader (#36) REGRESSED cold-start prefill — the FIRST GPU access to file-backed (mmap) pages costs ~44 ms of one-time page/TLB setup per process (0.5B pp1 wall 20 → 56 ms vs the copy path; a CPU-side madvise/touch does NOT fix it — the cost is the GPU's own access). Fix: register_part now dispatches a dummy kernel_warmup_read over the whole part buffer at model load (outside the CLI's Total timing; llama-bench numbers are equally warm). (b) the embed is now dispatched into the MAIN command buffer (llama builds ggml_get_rows into the main graph — one submit instead of two).0.5B pp1 wall 56 → 14 ms, pp31 prefill ~430 → ~950 t/s (llama 2686, gap 6× → 2.4×); 7B pp31 ~190 → ~247 t/s (llama 328, gap 1.7× → 1.3×), pp499 GPU 1215 → 1142 ms (~6 %), first-decode −~50 ms; cost: +~0.2 s 7B load (the read is amortized into load), 0.5B/7B outputs byte-identical, 34/35 bin tests greenthis commit
40Small-batch prefill matmul threshold (2026-08-21): minfer dispatched the simdgroup GEMMs only for nt ≥ 16 (chosen on the 0.5B in P0); nt∈[9,15] fell to the _multi kernels which serialize t INSIDE the threadgroup — measured 7B pp12: 16.6 t/s vs llama 130 (~7.8×), 0.5B pp12 ~3.9×. Fix: adaptive rule `nt ≥ 2 && (od ≥ 2048nt ≥ 9)— GEMM for all nt ≥ 9 (llamane11_mm_min=8→ MM for ne11>8) AND for nt∈[2,8] with large od (7B class: od≥3584 — GEMM ≈ llama'skernel_mul_mv_ext` there, measured pp4 7B 34 vs llama 61); small-od 0.5B matmuls keep the multi at [2,8] (measured better). The nt∈[2,8] small-od gap vs llama's ext kernel was re-measured 2026-10-09 and kept deferred by decision — the 2026-08-21 "0.5B pp4 141 vs 583" (0.24×) is stale; fresh numbers and rationale in §0.3; filed as #431

Completed optimizations in detail (every done / decided-not item): §3.1 Correctness fixes · §3.2 CPU sampler · §3.3 Decode optimizations · §3.4 Prefill optimizations · §3.5 KV / long context · §3.6 Prefill GEMM gap investigation (resolved)

🔜 To-do (required path to match llama.cpp)

Principle (2026-08-12): we do NOT accept the current state — whatever llama.cpp can achieve, minfer must too. The former "accept the architecture floor" verdict is revoked; §4 is the only action path.

#ItemGoalStatus
1GPU trace (minfer + llama): Performance Limiters + per-kernelper-phase bottleneck + per-op durations for both sidesDONE 2026-08-13 (§3.3): per-kernel + limiter comparison. Early "1.6-3.9×" numbers were trace-semantics artifacts (fused-vs-separate, mixed od); the clean isolation A/B (llama test-backend-ops perf) shows q5_0/q8_0 at parity and q6_K ffn_down 3× slower — fixed (this row)
2decode matmul per-call executionq6_K 72→209 GB/s (llama 217); decode 4.27→3.72 ms/tok (~13%)q6_K DONE 2026-08-13 (§3.3): ported llama's stride-2/float4 kernel layout; byte-identical + tests green. q5_0/q8_0 already at parity. q4_K next if a K_M model uses it
3flash attention portdecode attention 42.8→~4-6 µs/layer (~7-10×)DONE 2026-08-14 (§3.3): ported kernel_flash_attn_ext_vec (NSG=1, DK=DV=64/NE=2/C=32) as kernel_flash_attn_ext_f32/_f16; KV-layout check PASSED (minfer [nkv][nk*hd] == llama physical layout — no cache rework). Isolation-verified (tests/flash_attn_isolation.rs: cos vs CPU >0.999 for nkv 1..4097 incl. partial/empty chunks; flash-vs-split cos=1.0 through the shared combine), A/B byte-identical (0.5B Q4_K_M f32 + Q4_0 f16 + 7B Q4_K_M), decode GPU 0.25-1.0 ms/token faster (interleaved MINFER_TIMING), wall ~10 %; no long-context regression. Gate: MINFER_NO_FLASH=1 reverts to split
4prefill flash attention port (llama kernel_flash_attn_ext_blk, legacy simdgroup_matrix) + prefill GEMM/small efficiencyprefill 2.3-2.8× → ~1.5× (135 → ~90 ms); GEMM/small 89→~44 ms secondaryGEMMs RULED OUT 2026-08-14 (§3.6). Grid-shape probe (3.5-5.4 variance) + barrier/store experiments rule out the GEMM kernels (mem_none ≈ mem_threadgroup ~2-3 % and RACES in minfer). Real pp325 decomposition (0.5B Q4_K_M, MINFER_SKIP_ATTN): attention 46 ms (34 %), everything-else 89 ms. llama pp320 = 47.7 ms total with attention only ~3 ms (6803 vs 6373 t/s -fa on/off). llama's prefill attention is kernel_flash_attn_ext_blk = legacy simdgroup-matrix (has_simdgroup_mm, NOT the M5 tensor API) — single fused kernel vs minfer's 3-pass. PORT DONE 2026-08-14 (§3.4): kernel_flash_attn_blk_f32/_f16 (fixed-shape NSG=4, Q=8, C=64, DK=DV=64, 7168 B shmem, inline causal mask, kernel_kv_tail_pad for the partial last block) + host attn_flash_prefill. Isolation-verified (tests/flash_attn_blk_isolation.rs: cos vs CPU >0.999 across 16 nt/nkv configs incl. partial blocks + GQA, f32+f16, deterministic), A/B byte-identical to the classic gqa_attn_f32 at every layer (f32 AND f16 cache — maxabs 0.0), interleaved MINFER_TIMING prefill GPU ~110→~93 ms (~16 %), all 34 bin + 9 isolation tests pass. Bonus: FIXES the f16-cache prefill 3-pass bug (the 3-pass kernel_attn_scores/kernel_attn_output read the f16 KV cache as float* → garbage "!!!!!!"; the f16 blk kernel reads half K/V correctly). Default for hd==64 (0.5B/1.5B) and — since the 2026-08-15 hd=128 port — for hd==128 (7B) too. Gate: MINFER_NO_PREFILL_FLASH=1 reverts to 3-pass. Non-attention 89 vs ~44 ms remains a secondary structural gap — 7B direct per-kernel GEMM A/B 2026-08-18 (§3.6): minfer GEMMs at 87-94 % of llama (parity). Phase 0 prefill decomposition 2026-08-18 (§3.6): GEMMs are 76 % (0.5B) / 88 % (7B) of prefill; small kernels only 10 % / 4 % → #1 fusion ceiling low. Phase X 2026-08-18 (§3.6): §4.3.4's parity was a test-backend-ops measurement artifact — llama real prefill GEMMs ≈ 6.9 TFLOPS (467 t/s pp466) vs minfer ≈ 5.2 (≈1.33×). Concurrency, fusion, small kernels, dequant type, ik-loop barrier, sb-staging vectorization, and geometry ALL measured/verified — none explains the 1.33×; the gap is not addressable from minfer source (compiler-level only). FINAL 2026-08-20 (§3.6): exact 7B MUL_MAT graph mapped (197 GEMMs, 7.000 TFLOP; wk/down/output = q6_K — CORRECTS the earlier assumption), llama GPU-busy by host timestamps ≈ 1043 ms @ pp495 (6.71 TF, clean window), exact-shape replay (real 7B shapes, minfer kernels, one CB) ≈ 1126 ms = 6.20 TF. Isolated AND in-batch kernel A/B equal (6.21 vs 6.26 TF); grid/smem/buffer-mode/pooling/weight-data/barriers all free; GEMM interleave + 2-CB split hurt; concurrent dispatch no benefit with the engine's required per-dispatch barrier. Residual engine-vs-replay ~90-115 ms unattributable (GPU-side scheduling, below static visibility). Under comparable system load (avg 3-4) llama-bench pp495 degrades to 1250-1760 ms and converges with the replay. CLOSED: attention port DONE, GEMM/small gap confirmed NOT source-addressable at every tested level — only the tensor-API GEMM (mpp::tensor_ops, llama disables on M4) remains as a beat-llama research direction, not a parity fix. 2026-08-21 CORRECTION (§3.7 / #32): the '7.000 TFLOP / 6.71 TF' figures were WRONG — the output GEMM is N=1 (not N=495), so llama ≈ 6.26 TFLOP ≈ 6.0 TF ≈ the replay. The measured 'gap' was mostly minfer's full-nt lm_head over-count (≈539 GFLOP + 301 MB download), now FIXED (output-rows-only): 7B pp495 GPU ~1354 → ~1255 ms + download −~150 ms. Residual ≈1.2× remains not source-addressable (kernel/graph level).
57B same-model A/B + per-step regression check (0.5B is the research model; 7B is the user-facing one)7B decode/prefill gap vs llama quantified; no 7B regression from each stepBASELINE 2026-08-14 (§1.6): 7B Q4_K_M pp252: prefill ~240 t/s (52 % of llama 461), decode ~18.8 t/s (37 % of llama 50.5), steady GPU 50.1-51.3 ms/token. 0.5B sanity: pp252 2010 t/s (33 %), tg32 243 t/s (83 %) — no regression. CURRENT 2026-08-21: decode at parity (≈50 t/s, #27); prefill pp495 ~1255 ms vs llama ~1043-1064 ms (~82 %, §3.6/§3.7) after the lm_head output-rows-only fix (#32: GPU −~100 ms + download −~150 ms). Regression checks run per step; all green
6decode small-elementwise efficiency—CLOSED 2026-08-13: trace shows small-op parity (1.2-2.0 vs 1.3-1.9 µs) — the old 4× claim was subtractive noise (§3.3)
7q4_K decode matmul layout port (7B) — the next decode lever7B decode 37 % → closer to llama (steady GPU ~50 → target ~30 ms/token)DONE 2026-08-17 (§3.3): 7B K_M decode matmuls are Q4_K-dominated (attn_q/k/output + ffn_gate/up all Q4_K; Q6_K only output/ffn_down/attn_v). Ported llama's kernel_mul_mv_q4_K_f32_impl stride-4/float4 layout into kernel_q4_k_f32_matmul (TG(32, nsg=2) dispatch, sc16/kmask nibble unpack — the scale/min high/low nibble interleave reproduces llama's get_scale_min_k4 exactly, verified against llama's dequantizer). Steps: ① isolation probe at 7B dims — old kernel 70/18/74 GB/s (attn_q 3584/3584, attn_k 3584/512, ffn_g/u 18944/3584) → new 265/146/243 GB/s ② kernel port ③ 7B steady-decode A/B: ~49.7 → ~19.3 ms/token GPU (~2.6×), 17-19 → 45-49 t/s (≈ llama's 50.5), git-stash A/B byte-identical ④ 0.5B (no q4_K weights — trivially identical) + 1.5B (q4_K decode + multi path) regression green; all 34 bin tests pass

Note: §3.3's per-kernel table is superseded by the clean isolation A/B (§3.3): the trace mixed fused-vs-separate and different od per kernel name. The reliable decode gap = q6_K ffn_down kernel saturation (now fixed).

❌ Decided not to change (has measured or llama-source evidence)

#ItemEvidence
12D simdgroup_matrix (mpp tensor) GEMM portllama disables tensor GEMM on M4 Pro (PARAMETER_AUDIT A) — not llama's advantage
2bf16 / f16 intermediate activationsCore convention #1: llama Metal reads f32 activations (only KV → f16)
3non-blocking multi-cbminfer encode already hidden (0.13 ms); MINFER_SPLIT_CB measured linear regression
4parallel command buffers (A1)measured regression (1.67/1.08/1.43 s vs serial 0.93 s), reverted
5dispatch fusion (store_kv_both / residual_rms_norm)measured no gain (1.79 vs 1.74 s), reverted
6nt==1 matmul rewrite (full-block matvec)measured at ~200 GB/s bandwidth floor, no gain
7Q6_K / Q4_K dequant vectorizationQ5_0 full vectorization only +2.6 % — same low-value class
8f16 KV cache as default0.5B measured ~3 % slower (dispatch-latency-bound) → REPLACED by auto-select (#37): f16 for the 7B class (KV bandwidth-bound), f32 for small models — the 0.5B-side of this decision still stands
9MTLDispatchTypeConcurrent encoder (2026-08-20, §3.6)helps only WITHOUT the per-dispatch barrier (replay 1126 vs 1143 ms); with the engine's required memoryBarrier it is noise (FULL 1331-1370 vs serial 1336) — reverted
10GEMM-interleave / 2-CB split (2026-08-20, §3.6)dummy attn/norm dispatches between GEMMs: 845 vs 840 ms (hurts); llama-style split CB1+CB0: 1195-1210 vs 1126 ms (hurts)
11Prefill GEMM kernel experiments (2026-08-14, §3.6)grid-shape probe (3.5-5.4 TF variance) not the lever; mem_none ik-loop barrier + vectorized sb-store both RACY in minfer (deterministic corruption) — the mem_threadgroup barrier is a genuine correctness requirement

1. Current state

1.1 Historical same-model, same-parameter A/B (2026-08-14 record, llama.cpp 88b47a755)

Historical record — superseded. These numbers are the 2026-08-14/08-17 A/B against llama.cpp 88b47a755; they were re-measured 2026-10-07 against c479922ac in §0.2. Kept as history.

minfer --greedy (pure decode, llama "Generation" caliber); llama.cpp llama-bench -b 512 -t 8 (pure eval). Model Qwen2.5-0.5B-Instruct. Prefill numbers updated after the prefill flash port (5974eb1): minfer pp430-435 uses the kernel_flash_attn_blk_f32 path (hd==64).

Q4_K_M (qwen2.5-0.5b-instruct-q4_k_m.gguf):

Testllama.cppminferGap
prefill 30 tok2720 t/s~580 t/s (pp35, flash; short-prompt fixed overhead dominates)—
prefill 430 tok6909 t/s~2530-2620 t/s (pp435, flash)2.6-2.7×
decode 128 tok (pure GPU)293-299 t/s~218 t/s (4.47 ms/tok steady)1.3-1.4×
decode, default sampling247 t/s~197 t/s1.25×

Q4_0 (qwen2.5-0.5b-instruct-q4_0.gguf):

Testllama.cppminferGap
prefill 30 tok2610 t/s~950 t/s (pp31, #39; was ~430)~2.4×
prefill 430 tok7449 t/s~2770 t/s (pp435, flash)2.7×
decode 128 tok314-339 t/s~279 t/s (3.90 ms/tok steady)1.1-1.2×

Reading:

  • Decode is now 72-88 % of llama (pure GPU 1.1-1.4×, default sampling 1.25×) — driven by rms_norm_256, the chunk-cap/sync fixes, and the per-kernel non-matmul profile (was 1.47× before 2026-08-10).
  • Prefill (long) improved from 2.8-3.6× to ~2.6-2.7× after the prefill flash port (§3.4): GPU ~164 → ~144 ms at pp435 (~12 %; pp294 ~16 %). The residual gap is the non-attention 89 → ~44 ms structural difference (GEMMs + small kernels under-occupied, §3.6), NOT attention (now ~3-4 ms, llama-like).
  • Short prefill (pp30): dominated by per-dispatch fixed overhead (~950 t/s at pp31 after #39 — the mmap first-access TLB cost is now absorbed at load; was ~430 t/s), so pp30 is no longer a meaningful attention lever; llama's pp30 is similarly launch-bound (0.5B 2686, 7B 328 t/s — minfer now 2.4×/1.3×).

Post-#28-32 (2026-08-21): the §1.1 0.5B numbers remain the reference — the GEMM-partial-tile barrier fix (#28) was a correctness fix (no regression), the unroll/barrier changes (#29/#30) improved 0.5B prefill ~7-9 % on top, and the lm_head output-rows-only fix (#32, §3.7) is a prefill/lm_head change that does not alter the 0.5B decode reference here. No re-measurement has invalidated any §1.1 figure. The user-facing 7B state is the current focus — see §1.6.

1.2 Per-token GPU decomposition (decode, nt==1, Q4_K_M 0.5B)

Categoryminfer GPUllama GPUEvidence
matmul (QKV/O/GU/down/output, ~97 kernels)~3.0 ms (~130 GB/s)~3.0 ms (source+params identical)minfer measured / llama inferred
attention (split 2 kernels)0.54 ms~0.15-0.2 ms (flash vec 1 kernel)minfer measured (skip-ATTN) / llama inferred
small elementwise (norm/bias/rope/store/add/swiglu, ~300)~0.5 ms~0.1-0.3 msminfer measured / llama inferred
base infra (encode+submit+download)encode 0.13 + download 0.02-0.03~0.3-0.5 (incl. multi-cb encode)minfer measured / llama inferred
Total~4.35-4.55 ms/token GPU~3.1-3.3 ms GPU / 3.51 wallinterleaved A/B

SUPERSEDED 2026-08-13 by the per-kernel trace (§3.3): the llama-side numbers here were inferred (no per-op timing existed). The trace shows the real picture differs: matmuls are NOT "zero gap" (minfer 1.6-3.9× slower per call), and small elementwise is NOT "4× slower" (parity). The Total is right; the category split was not.

1.3 Whole-pipeline comparison (decode token, nt==1)

Stageminferllama.cppCPU/GPU
Samplingsampler.rs top_k/top_p/temp/repeat-penalty (O(n) + candidate list)llama-sampler.cpp candidate chain (partial_sort)CPU
Dispatch encodeMpsCommandBuffer set_buffer/set_params ×N, single-threadedsame; multi-cb threads hide encodeCPU
GPU executionsingle cb serial ~483 dispatches (Q4_K_M)single/multi cb, ~490-530 dispatchesGPU
EmbeddingQ4_0: kernel_get_rows_q4_0; others: CPU embed + uploadggml_get_rows → MetalGPU (or CPU upload)
KV storekernel_store_kv_f32/_f16 (2 dispatches)kernel_cpy_f32_f16 ×2 (K,V)GPU
Attention2 kernels/layer (partial + combine)1 kernel/layer (flash vec)GPU
Logits readbackcopy_from_gpu 607 KBMetal buffer readGPU→CPU

1.4 Per-layer kernel sequence (Q4_K_M 0.5B, nt==1)

minfer — 20/layer (fused QKV OFF: Q5_0/Q5_0/Q8_0 mixed types, cannot concat):

RMSNorm → Wq/Wk/Wv 3×matmul → 3×add_bias → 2×RoPE → 2×KV store → attention split(partial+combine) → Wo matmul → residual → RMSNorm → fused gate+up matmul → SwiGLU → Ffn_down matmul → residual

×24 = 480 + output_norm 3 = 483. Q4_0 model (all Q4_0): fused QKV + BSR active → 12/layer ×24 + output 3 + GPU embed 1 = 292.

llama.cpp — 17/layer (flash_attn on):

RMSNorm → Wq/Wk/Wv 3×matmul → 2×RoPE → 2×KV store(f32→f16) → flash attention(1 dispatch) → Wo → residual → RMSNorm → gate+up 2×matmul → SwiGLU → Ffn_down → residual

×24 = 408 + output 3 + embed 1 ≈ 412 base; graph 822 nodes → ~490-530 dispatches (f16 cast/cont/reshape are non-no-op nodes).

1.5 Early performance milestones (Qwen2-0.5B, historical)

PhaseOptimizationDecode (short)Cumulative
BaselineGPU + 4 correctness fixes130 tok/s1.0×
+2Flash Attention + float4151 tok/s1.2×
+3SIMD-parallel attention196 tok/s1.5×
+4SIMD-parallel RMSNorm334 tok/s2.6×
+5SwiGLU fusion312-334 tok/s2.5×

(Early numbers used the blended caliber; since 2026-08-06 Generated: is pure decode.)

1.6 7B state (2026-08-21 record, user-facing model, Q4_K_M)

Historical record. The llama-comparison gaps below are the 2026-08-21 standing; the refreshed 2026-10-07 A/B is in §0.2.

The 7B (qwen2.5-7b-instruct-q4_k_m, split GGUF) is the user-facing model; this is the current standing after the decode q4_K port (#27), the correctness/unroll/ barrier work (#28-30), the Phase-0 prefill decomposition (#31, §3.6), the lm_head output-rows-only fix (#32, §3.7), and the GPU Q4_K embedding port (#33).

Metricminferllama.cppGap
decode (steady GPU)~19.3 ms/token (45-49 t/s)~50.5 t/s (19.8 ms/tok)≈95 % (parity)
prefill pp495 GPU~1234 ms @ pp499 (was ~1278 pre-#34; ≈5.07 TF @ ≈6.26 TFLOP)~1043-1064 ms (≈6.0 TF @ 6.26 TFLOP)~86 % (~1.16-1.18×)
prefill pp495 logits download~0.1 ms (608 KB, was ~150 ms/301 MB)blit, 608 KBparity
exact-shape replay (pure GEMM, one CB)~1126 ms (6.20 TF)—converges with llama under comparable load
  • Decode is at parity (essentially closed by #27's q4_K decode matmul port).
  • Prefill: the big lm_head over-count (minfer computed [152064×495] logits, llama computes [152064×1] after get_rows(inp_out_ids)) is FIXED (#32) — GPU ~1354 → ~1255 ms, download ~150 → ~0.1 ms. The corrected FLOP accounting (llama ≈ 6.26 TFLOP, NOT 7.0) shows llama ≈ 6.0 TF ≈ the replay's 6.2 TF.
  • The residual ~1.2× prefill gap is kernel/graph-level, below static visibility (GPU backend machine-code scheduling / execution environment) and accepted (§3.6). The only remaining research direction is the tensor-API GEMM (mpp::tensor_ops) — llama disables it on M4 by default, so it is a beat-llama option, not a parity fix. The ~40 ms follow-up (last-layer FFN on output rows only, §3.7) is now DONE (#34) — minfer's total graph work (≈6.26 TFLOP) exactly matches llama's, and the pp495 gap closed 1.2× → ~1.16-1.18×.

2. Gap analysis: verified vs inferred

Core finding (2026-08-06 #5, verbatim): "Structural" is an inference, not a proven architectural inferiority. §2.1 is what is strictly VERIFIED; §2.2 is the inference after elimination. The decisive per-kernel measurement was completed 2026-08-13 (§3.3) — see its per-kernel table, which refuted several of the §2.2-§2.4 inferences.

2.1 Strictly verified

  1. matmul kernel source is line-for-line identical (nt==1 mul_vec_q_n_f32_impl / block_q*_dot_y translations).
  2. dispatch count is comparable (~436 vs ~490-530).
  3. dispatch params match (llama ggml-metal-impl.h N_R0/N_SG vs minfer, 2026-08-06 #6).
  4. dispatch count is nearly identical (~484 vs ~490-530); the gap is per-dispatch GPU execution time (10.3 µs vs 6.2 µs).
  5. prefill GEMM ceiling ~5.4 TFLOPs/s (prefill_gemm_throughput_profile, 2026-08-11 A1); llama ~7 TFLOPs/s effective.

2.2 Closed hypotheses (2026-08-06 #6)

  1. Dispatch params — DISPROVEN (see 2.1.3).
  2. Attention is the main decode lever — DISPROVEN: llama -fa on/off = 3.64 vs 3.88 ms (only ~0.25 ms); even flash OFF llama beats minfer.
  3. Multi-cb is a lever — DISPROVEN: minfer encode is only 0.13 ms; MINFER_SPLIT_CB=N regresses linearly (0.67 → 0.93/1.23/1.62 s).
  4. CPU side (encode/sampler) — DISPROVEN: encode 0.13 ms; sampler fixed (2×).
  5. Q5_0 scalar dequant is the matmul bottleneck — DISPROVEN: full vectorization only +2.6 %; ~130 GB/s is nt==1 small-grid structural latency.
  6. Matmul is the prefill bottleneck (before the attention fix) — 2026-08-11 proved classic attention ~100 ms (48 %) was; matmuls became the main remainder after the fix.

2.3 Per-kernel non-matmul profile (2026-08-10, P0)

src/metal/::tests::non_matmul_bandwidth_profile (batched-cb, median of 3) — a single dispatch is dominated by the ~165 µs cb launch+sync floor, so batch dozens and take the median:

Kernelµs/dispatchNotes
rms_norm 32t (1 simdgroup)13.87× elementwise — latency-bound
rms_norm 256t (8 simdgroups, P1)3.7~3.7× faster, bit-identical
add_f32 / add_bias / swiglu / rope / store_kv~1.6-2.3256-thread elementwise baseline
attn_bias_rope_store (BSR)3.1
attention split pair (partial+combine, nkv=430)44.3dominant non-matmul kernel
attention classic (single-pass, nkv=430)3528× worse than split — confirms the split design

Findings: ① the attention split pair is the dominant non-matmul kernel (44 µs/layer); only a faithful flash port can cut it. ② the small elementwise tail (~300 kernels × 2-3 µs) is structurally cheap per-dispatch latency; rms_norm was the one exception and is now fixed.

2.4 Final gap report (2026-08-06, precise decomposition) — SUPERSEDED

Componentminfer GPUllama GPUgap
matmuls~3.0 ms (bandwidth-bound, source+params identical)~3.0 ms~0
non-matmul (attention + small + serialization)~1.2 ms~0.3 ms (inferred)~0.9 ms

The structural gap is 100 % in non-matmul — minfer's ~340 small kernels run at ~4× llama's efficiency. (Honest uncertainty: minfer's 1.2 ms is reliable; the attention-vs-small split has ±0.2 ms noise; llama's 0.3 ms is inferred.)

SUPERSEDED 2026-08-13 by §3.3's per-kernel trace. The "~0 gap matmuls" and "small ~4×" were subtraction artifacts. Real per-kernel data: matmuls ARE the gap (1.6-3.9× per call) and small-op is at parity. This section is kept as the historical pinpoint that motivated the trace.

2.5 KV-growth component (partially fixed 2026-08-10)

minfer's average decode grows with context (5.05 → 6.7 ms at -n 64→512), llama's stays flat. Two addressable causes were fixed: attention chunk cap 32→16 (avoid over-parallelization) + removal of sync_kv_to_cpu on the pure-GPU paths (an O(nkv)/token copy). Interleaved A/B: 4.65-4.76 → 4.50-4.55 ms/token (~0.2-0.25 ms). The remainder is sub-linear attention KV-read, which llama amortizes natively via its f16 cache + flash.


3. Completed optimizations in detail

3.1 Correctness fixes (Metal backend foundation)

4 early bugs (all affected output correctness, Qwen2-0.5B):

  1. RoPE freq_scale not applied (src/metal/ rope_f32 got the param + forward.rs passes hp.rope_freq_scale).
  2. output_b not applied (output_norm_gpu adds the bias).
  3. softmax max initialization (-INFINITY instead of 0, prevents NaN).
  4. attention stack array hardcoded (hd dimension made dynamic).

Q5_K formula + qh index fix (2026-07-31, affects CPU + Metal):

  • Formula: Q5_0-style signed dl*(u-16)-ml was wrong → llama's unsigned dl*u-ml.
  • qh high-bit index: qh[sub*4+pos/8] bit pos%8 was wrong → qh[pos] bit sub.
  • Fixed in quants.rs / kernel.rs / forward.rs (embed).

GQA attention simd_max divergence fix (2026-08-01, 28d4ba2): in a partial KV tile (nkv % 32 != 0), out-of-range lanes exited the loop early → simd_max(dot) ran across divergent lanes with stale registers → corrupted online-softmax running max → repetition loops. Fix: uniform iteration count + valid mask (invalid lanes dot=-INF, e=0). Result: prefill logits cos 0.83→0.999. Regression test tests/gqa_attn_isolation.rs.

GPU-hang safety hardening (2026-08-03, bff73db):

  1. submit() bounded 10 s wait + MTLCommandBufferStatus check + MINFER_TRACE dispatch trace (GPU fault errors out instead of freezing the machine).
  2. attention kernels never return early before the barrier (prevents nh % nk != 0 deadlock).
  3. layer_gpu/output_norm_gpu runtime guards: nh % nk == 0, hd ≤ 256, id % 32 == 0; error-exit (gpu_abort) on violation.

3.2 CPU sampler (2026-08-06)

Root cause: sampler.rs ran a full-vocab O(n·log n) sort per token + 607 KB copy (top_k) + a full sort of 151,936 (usize,f32) tuples (top_p). llama.cpp uses a candidate-list chain (std::partial_sort O(n·log k) → later samplers operate only on the ≤k survivors).

Fix:

  • top_k → select_nth_unstable_by (O(n), on a copy to preserve the index→token mapping).
  • top_p → softmax + sort only the ≤k survivors (falls back to the full-array path when >1024 survive).
  • temp → skip exp() for masked (-INF) logits.
  • main.rs moves the logits Vec instead of logits_all[..].to_vec() (607 KB/token).

Measured: -n 128/256/512 all ~2.0×; default sampling ~12.6-14.8 → ~5.5-6.5 ms/token; fixed-seed output byte-identical (7 sampler tests pass).

3.3 Decode optimizations (GPU)

ItemEffectCommit (pre-rewrite)
Fused QKV + FFN gate/up (nt==1 single matmul/group)~5 % decode6f0c847
KV-parallel split attention (2-pass online-softmax)~32 % decodeb3d4c7a
float4 acc + adaptive chunks + KV geometric growthextra ~15 % + long context66f4290
f16 split attention (partial _f16)f16 1.60→0.95 s387d612
float4 elementwise + parallel RoPE (P6/P7)~2-3 %ddd3eb0
256-thread RMSNorm (llama kernel_rms_norm_fuse_impl port)~3-4 %a7f21e4
Fused bias+RoPE+KV-store (BSR, 7→1 kernel, nt==1)~5 %5c106dd

Split-attention design (b3d4c7a):

  • Pass 1 kernel_gqa_attn_partial_f32/_f16: grid (nt, nk, n_chunks); each TG computes an online-softmax partial (mx, S, acc) for its KV chunk, same tile/barrier/valid-head structure as the classic kernel.
  • Pass 2 kernel_gqa_attn_combine_f32: grid (nt, nh) merges the partials (pure elementwise, no shared mem/barriers).
  • n_chunks = clamp((max_pos+1+31)/32, 1, 16) (lowered from /16..32 on 2026-08-10; MINFER_ATTN_CHUNKS overrides). Correctness is invariant to chunk count.

Fused QKV essentials (6f0c847; the "24 %" figure was a GPU-state artifact, corrected to ~5 % in 26b145b): Wq/Wk/Wv and ffn_gate/up are row-major-concatenated at load (concat_rows → blk.{i}.attn_qkv / ffn_gu) when types + input dim match. nt==1 runs ONE matmul/group; rope/store/swiglu read sections via set_buffer byte offsets. MINFER_NO_FUSE_QKV=1 A/B byte-identical; gemm_isolation.rs::qkv_row_concat_layout locks the layout.

Measurement-trap notes (avoid repeating):

  • Single-dispatch isolation timing is unreliable (~165 µs cb launch+sync floor) — batch dozens and take the median.
  • Each kernel's first timed run has a cold-start/GPU-clock-ramp artifact (~4×) — warm first, measure twice.
  • Sustained benchmarking thermally throttles the M4 Pro (extreme: all configs ~1.3 s) — interleave configs, take min/median.

q6_K decode matmul layout port (2026-08-13, §3.3) — the decode gap was matmul-dominated, not attention/small-op (per-kernel trace §3.3). Clean isolation A/B (minfer matmul_bandwidth_profile vs llama test-backend-ops perf -b MTL0 -o MUL_MAT, same od/id/nt) showed q6_K ffn_down (896/4864) 72 → 217 GB/s (3.0× gap) — minfer's stride-64 super-block loop left only 19/64 TG threads busy (~30 % util) with scalar inner loops. Fix: ported llama's kernel_mul_mv_q6_K_f32_impl stride-2 + float4 layout (TG(32, nsg=2) dispatch):

  • q6_K isolation 72 → 209 GB/s (llama 217); decode steady GPU 4.27 → 3.72 ms/token (~13 %); wall ~203 → ~220 t/s (Generated: pure-decode).
  • byte-identical (git-stash A/B), all tests green. q5_0/q8_0 already at parity.

q4_K decode matmul layout port (2026-08-17, to-do #7, §3.3) — 7B K_M decode is Q4_K-dominated (attn_q/k/output + ffn_gate/up all q4_K; Q6_K only output/ffn_down/attn_v) and carried the same gap q6_K had. Rewrote kernel_q4_k_f32_matmul as a faithful kernel_mul_mv_q4_K_f32_impl transcription: stride-4 super-block loop, float4 acc over the kmask nibble unpack, sc16 scale unpack (verified byte-exact vs llama's get_scale_min_k4, ggml-quants.c dequantize_row_q4_K), TG(32, nsg=2), grid od/4:

  • 7B isolation at decode dims: attn_q (3584/3584) 70 → 265 GB/s, attn_k (3584/512) 18 → 146 GB/s, ffn_gate/up (18944/3584) 74 → 243 GB/s.
  • 7B steady decode GPU ~49.7 → ~19.3 ms/token (~2.6×); wall 17-19 → 45-49 t/s (llama 50.5). 0.5B/1.5B regression green, all 34 bin tests pass.

flash decode attention port (2026-08-14 hd=64 + 2026-08-17 hd=128, §3.3) — isolation confirmed attention at 42.8 µs/layer (partial+combine) vs llama's ~4-6 µs (flash vec) = ~7-10×, ~0.9 ms of the 3.72 ms decode step (~24 %). Root cause: llama's flash uses dot(float4,float4) + simd_shuffle_down reductions with no threadgroup barrier within a KV tile; minfer's split had 2 threadgroup_barrier per 32-row tile. Ported llama's kernel_flash_attn_ext_vec as kernel_flash_attn_ext_f32/_f16 (NSG=1, DK=DV=64, NE=2, C=32, NL=16, fixed-shape for Qwen2, writes the same {M,S,O} partials so the shared combine merges unchanged). KV-layout check PASSED (minfer [nkv][nk*hd] == llama physical layout — no cache rework). GPU-safety deviations: inline per-lane partial-chunk mask (llama's cross-lane sm[] write is a race), break-only control flow so all 32 lanes reach both barriers, clamped reads + -MINF masking instead of a pad buffer. Verified: tests/flash_attn_isolation.rs (cos vs CPU >0.999 for nkv 1..4097, cos=1.0 vs split through the shared combine), A/B byte-identical (0.5B Q4_K_M f32 + Q4_0 f16 + 7B Q4_K_M), decode GPU ~0.3-1.0 ms/token faster, wall ~10 %. hd=128 decode flash (7B, 2026-08-17): separate kernel_flash_attn_ext_hd128_f32/_f16 (NE=1, DK4=DV4=32, full simd_sum(mqk) per cc-iteration); 7B pp205 steady decode GPU flash ~51.1 vs split ~51.6 ms/token (weight-read-bound; attention was NOT the 7B decode bottleneck — the q4_K port below was).

3.4 Prefill optimizations

simdgroup GEMM (P0/P1, 2026-08-01): faithful llama legacy kernel_mul_mm port (64×32 tile, 4 simdgroups × 32 threads, Q4_0 dequant staged into sa, f32 activations into sb). P0's initial version had 3 bugs (B-staging unclamped rows, store transpose direction, barrier must be mem_threadgroup) → after P1 fixes: +11 % at 30 tok, +34 % at 70 tok. Dispatched for nt ≥ 16; MINFER_GEMM=0 falls back to f32 multi. Isolation test gemm_isolation.rs (nt=12/30/32/33).

Non-Q4_0 GEMMs (2026-08-03, c9f865c/2c03bd1/5b914f0): one simdgroup GEMM per quant — Q8_0/Q5_0/Q5_1 (32-elem blocks) + Q4_K/Q5_K/Q6_K (256-elem super-blocks). K_M prefill 300→650 t/s; 1.5B Q4_K_M 48→442 t/s (~9×). 8 KB threadgroup-memory guard. non_q4_0_gemm_isolation verifies.

Parallel prefill attention (2026-08-11, b2c97fd): classic kernel_gqa_attn_f32 is latency-bound at prefill (grid (nt,nk) sequential KV loop, ~24K barriers, ~100 ms = 48 % of prefill, ~25× llama's attention). Replaced by a 3-pass barrier-free design:

  1. kernel_attn_scores: one 256-thread TG per (t,h) row; each thread computes one score.
  2. kernel_softmax_attn: masked softmax over the kv axis.
  3. kernel_attn_output: softmax·V sum.

GQA via per-head hk = h/gqa (the broadcast-GEMM idea was tried and abandoned — a 2D GEMM can't produce the per-head 3D scores tensor). ⚠️ threadgroup-memory bug: the softmax's shmem[tiisg] writes 32 floats but only 8 were allocated (OOB corrupted adjacent memory → NaN rows) — fixed to 32×4=128 B; rms_norm_256 had the same latent bug.

Measured: pp430 classic 212 → 144 ms (attention 100→30 ms); pp30 44→40 ms; 7B pp230 944→832 ms (attention 169→57 ms); 7B decode unchanged. 34 bin + 6 isolation tests pass, end-to-end byte-identical.

Prefill flash attention port (2026-08-14, §3.4, 5974eb1) — llama's prefill attention is a SINGLE fused kernel_flash_attn_ext_blk (simdgroup_matrix, NSG=4, Q=8, C=64) vs minfer's 3-pass (scores + softmax + output); measured minfer attention 46 ms of the 135 ms pp325 (34 %) vs llama's ~3 ms. Ported kernel_flash_attn_blk_f32/_f16 (fixed-shape NSG=4, DK=DV=64, 7168 B shmem, inline causal mask, kernel_kv_tail_pad for the partial last KV block). Isolation-verified (tests/flash_attn_blk_isolation.rs: cos vs CPU

0.999 across 16 nt/nkv configs incl. partial blocks + GQA, f32+f16, deterministic), A/B byte-identical to the classic gqa_attn_f32 at every layer (f32 AND f16 cache — maxabs 0.0), interleaved MINFER_TIMING prefill GPU ~110 → ~93 ms (~16 %). Bonus: fixes the f16-cache prefill 3-pass bug (the 3-pass kernels read the f16 KV cache as float* → garbage "!!!!!!"; the f16 blk kernel reads half K/V correctly). Default for hd==64 (0.5B/1.5B). Gate: MINFER_NO_PREFILL_FLASH=1.

hd=128 (7B) prefill flash port (2026-08-15, §3.4, 89ac39a) — 7B pp310 prefill GPU 1042 → ~949 ms (~9 %), f32/f16 byte-identical, fixes 7B f16-cache garbage. Default for hd==128 (7B) too.

GEMM partial-tile race + missing Metal memoryBarrier fix (2026-08-19, §3.4, 7253a3b) — (a) all 8 simdgroup mm kernels lacked the threadgroup_barrier BEFORE the partial-tile temp_str stores (temp_str overlaps sa/sb, so a fast simdgroup overwrites them while a slow one still reads → intermittently corrupted last-2-token logits, partial x-tile only); (b) the single prefill encoder had NO memoryBarrier between dispatches → RMSNorm write raced the QKV read of the reused bn buffer (last-2 token slots, huge stale values). Result: 1.5B/7B first-token nondeterminism (~10-30 % wrong tokens) → 24/24 deterministic, output matches CPU byte-for-byte.

mm-kernel hot-loop #pragma unroll (2026-08-19, §3.4, f3a499d) — llama FOR_UNROLLs the staging/ik/load/mac loops; minfer's 8 mm kernels had none. Added the 6 unroll points (llama-parity set): 7B pp495 1438.8 → 1355.6 ms (~5.8 %), 0.5B ~6.9 %, 1.5B ~2.4 %; byte-identical + 24/24 determinism.

ik-loop threadgroup_barrier → simdgroup_barrier(mem_none) (2026-08-19, §3.4 follow-up, 0e756f3) — the pre-unroll corruption was a rolled-loop compiler artifact, not a memory-visibility need (.air diff: the barrier swap changes ONLY the barrier instruction — zero IR scheduling difference). With the unroll in place llama's exact barrier form is now safe: 7B pp495 min 1387.4 → 1370.6 ms (~1.2 %), byte-identical (1.5B×24 / 7B×8 / 0.5B×3). Removes the last structural mm-kernel difference vs llama.

3.5 KV / long context

  • KV geometric growth (66f4290): kv_ensure_layer grows ×2 instead of reallocating + copying the whole old KV every token (0.5 ms@KV140 → 4.2 ms@KV2510 → 0.13 ms). ⚠️ an old_v clone typo polluted the V cache (Q4_K_M garbage) — the A/B didn't catch it (both paths share the corrupted KV); found against a known-good reference.
  • f16 KV auto-default (387d612 + bff73db + #37): MINFER_CACHE_TYPE=f16/f32 overrides; unset → set_kv_cache_type at load picks f16 for the 7B class (n_layers×n_kv_embd ≥ 8192) and f32 for small models (0.5B measured ~3 % slower on f16 — dispatch-latency-bound). kernel_store_kv_f16 + the _f16 flash / partial kernels serve the f16 path.
  • Split-GGUF (cbba68c/34eaf10): multi-part models (7B -0000X-of-0000Y), merged tensor index, 7B verified.

3.6 Prefill GEMM gap investigation — RESOLVED (2026-08-14 → 08-20, decided not to change)

The prefill GEMM throughput gap was investigated exhaustively and is now closed as "not source-addressable" (the mm kernels are proven identical to llama's). Record of the investigation (was §4.3.1-§4.3.10):

StepFindingVerdict
§4.3.1 (08-14): grid-shape probe + kernel experimentsGEMM efficiency varies 3.5→5.4 TF purely by grid shape; mem_none barrier + vectorized sb-store either no-gain or RACY in minfer (deterministic corruption)GEMM kernels ruled out as the prefill lever
§4.3.4 (08-18): 7B direct per-kernel GEMM A/Bminfer GEMMs at 87-94 % of llama (test-backend-ops)GEMMs at parity (isolation)
§4.3.5 (08-18): Phase 0 subtractive decompositionGEMMs = 76 % (0.5B) / 88 % (7B) of prefill; small kernels only 10 % / 4 %#1 fusion ceiling low
§4.3.6 (08-18): real-chain mechanism§4.3.4's parity was a test-backend-ops artifact — llama real prefill GEMMs ≈ 6.9 TF vs minfer ≈ 5.2 (≈1.33×); concurrency/fusion/small-kernels/dequant/barrier/vectorization/geometry all measured — none explains 1.33×gap not addressable from minfer source
§4.3.7 (08-19): partial-tile race + missing memoryBarrierFIXED (→ §3.4)correctness
§4.3.8 (08-19): re-baseline + last cheap leversMSL -O3 no benefit; concurrent dispatch ~0.5 % (noise); dependency-aware barrier ≈ barrier-alwaysgap accepted
§4.3.9 (08-19): #pragma unroll + ik-loop barrier~5.8 % + ~1.2 % recoveries (→ §3.4); structural-equivalence audit: source/IR/.air/smem/dispatch/runtime-compile ALL identicallast structural mm-kernel difference gone
§4.3.10 (08-20): Phase-0 7B decomposition + exact-shape replay7B MUL_MAT graph = 197 GEMMs / 7.000 TFLOP (wk/down/output q6_K — CORRECTS earlier assumption); llama GPU-busy ≈ 1043 ms @ pp495 (6.71 TF, host timestamps); exact-shape replay (real shapes, one CB) ≈ 1126 ms = 6.20 TF; isolated + in-batch kernel A/B EQUAL (6.21 vs 6.26 TF); grid/smem/buffer-mode/pooling/weight-data/barriers free; GEMM-interleave + 2-CB split HURT; concurrent dispatch no benefit with the per-dispatch barrier; engine-vs-replay residual ~90-115 ms unattributable (GPU-side scheduling, below static visibility)CLOSED: gap confirmed not source-addressable at every tested level

Residual status (pre-#32, 2026-08-20): 7B pp495 minfer FULL ~1324-1336 ms vs llama ~1043-1064 ms (~73 %, ≈1.25-1.3×), converging under comparable system load (llama-bench degrades to 1250-1760 ms at load avg 3-4). Updated 2026-08-21 by #32 (§3.7): after the lm_head output-rows-only fix minfer is ~1255 ms GPU (≈82 %, ≈1.2×) with download ~0.1 ms — see §1.6 / §3.7. The only remaining research direction is the tensor-API GEMM (mpp::tensor_ops) — llama disables it on M4 by default, so it is a beat-llama option, not a parity fix (decided-not, §0).

⚠️ 2026-08-21 CORRECTION — the §3.6 FLOP accounting was WRONG. The "7.000 TFLOP" total assumed the output (lm_head) GEMM runs on all N=495 tokens. The graph dump actually shows out=[152064 1 1 1] for the output and the last layer's gate/up/down on N=1 — llama's ggml_get_rows(cur, inp_out_ids) (qwen2.cpp:106-108) reduces to n_outputs rows after the last attention, so the last layer's FFN + final norm + lm_head all run on N=1. Correct llama total ≈ 6.26 TFLOP → llama efficiency ≈ 6.0 TF (NOT 6.71), essentially equal to the replay's 6.2 TF. minfer was computing the full-nt output GEMM [152064×495] (≈539 GFLOP waste) + final norm on all rows + a 301 MB logits download — a large fraction of the measured "gap" was this over-count, not the kernels. Fixed 2026-08-21 (§3.7 / changelog #32).

3.7 lm_head / final-norm output-rows-only (2026-08-21, changelog #32)

Discovery (from the LLAMA_METAL_E2E.md reference): llama's graph reduces the hidden state to n_outputs rows right after the last layer's attention (ggml_get_rows(cur, inp_out_ids) + get_rows(inpSA, inp_out_ids), src/models/qwen2.cpp:106-108), so the last layer's FFN, the final norm and the lm_head all run on 1 row for a single-sequence prefill (graph dump: output GEMM out=[152064 1 1 1], last-layer gate/down [.. 1 ..]). minfer computed the output projection over all nt tokens ([nv×nt]) and downloaded all of it.

Fix: forward() / output_norm_gpu (Metal + CUDA) now take n_out (number of output rows = the LAST n_out tokens, single-sequence row-major [nt][ne] hidden). The final rms_norm + output GEMM + bias + logits buffer + download all operate on n_out rows (n_out=1 for the minfer CLI). Host: src/models/qwen2/forward.rs (GPU output_norm_gpu(…, n_out, …), CUDA path, CPU fallback slices hidden[(nt-n_out)*ne..]), src/metal/runtime.rs + rms_norm(.., off) byte-offset, src/main.rs passes n_out=1.

Measured (7B Q4_K_M pp495, interleaved A/B same window):

  • GPU submit-wait: pre ~1354 (noisy 1354-2754) → post 1253-1271 ms (stable)
  • logits download: ~150 ms (301 MB) → ~0.1 ms (608 KB)
  • 0.5B GPU generation byte-identical pre/post (same seed, same text); 1.5B / 7B greedy generation correct.
  • Corrected efficiency: llama ≈ 6.0 TF @ 6.26 TFLOP; minfer post-#32 ≈ 6.46 TFLOP @ ~1.255 s ≈ 5.15 TF — the remaining ~1.2× was the §3.6 kernel/graph residual plus the last layer's FFN still on all nt (llama reduces before it) ≈ ~40 ms follow-up potential.

Follow-up — DONE 2026-08-21 (#34): the last layer's FFN now runs on the output rows only, exactly matching llama's get_rows(inp_out_ids) reduction (qwen2.cpp:106-108 reduces cur + inpSA BEFORE the last layer's FFN). Fix in layer_gpu (src/metal/): the wo matmul stays on all nt (llama build_attn precedes the reduction), then the wo-residual, ffn_norm (byte-offset read of the hidden tail), gate/up/down (dispatched with nt=n_out via the new x_off matmul param) and the final residual all run on the tail n_out rows (add_f32_off); the CPU path in forward.rs mirrors (gate/up/down on n_out rows, residuals on the hidden tail slice).

  • Measured (7B Q4_K_M pp499, interleaved A/B same window): GPU submit-wait ~1278 → ~1234 ms (~44 ms, ~3.4 %) — the ~40 ms estimate confirmed.
  • minfer's total graph work drops 6.46 → ≈6.26 TFLOP — now exactly llama's total; efficiency ≈ 5.07 TF @ 6.26 TFLOP (llama ≈ 6.0 TF) → gap ~1.16-1.18×.
  • Correctness: 7B/0.5B GPU greedy byte-identical pre/post (same seeds) + 0.5B CPU fallback identical; decode untouched (nt==1 ⇒ ffn_nt==nt, the reduced path never triggers); 33/34 bin tests green (attn_parallel_realdata_correctness fails on missing /tmp/dp3 dumps — pre-existing environment dependency).

4. Roadmap status — all items resolved (2026-08-20)

The 2026-08-12 principle — "we do not accept the current state; whatever llama.cpp can achieve, minfer must too" — drove the work below. Every roadmap item is now resolved (done, or decided-not-to-change with measured / source evidence). The detailed records live in §3 ("Completed optimizations in detail", incl. the §3.6 investigation record) — see the §3 links in §0's Progress Overview; decided-not items are in §0's "Decided not to change" table. This chapter holds the remaining research + operational reference, plus the cold-start to-dos added 2026-08-21 (§4.2) (a separate axis from the steady-state gap — not part of the §0 "match llama.cpp" table).

4.1 Remaining research (not a parity fix)

The only direction left to beat llama on prefill is the tensor-API GEMM (mpp::tensor_ops, kernel_mul_mm_id) — llama disables it on M4 Pro by default (PARAMETER_AUDIT A), so it is a research exploration, not a parity requirement. No other prefill path remains open (§3.6 / §0 decided #1).

2026-08-21 (#32, §3.7): the last-layer FFN output-rows-only follow-up (llama's get_rows(inp_out_ids) shrinks the graph before layer-27's FFN, closing the minfer-vs-llama prefill work difference 6.46 vs 6.26 TFLOP) is now DONE (#34) — no prefill work difference remains; minfer's total equals llama's.

2026-08-21 (#41, DEFERRED): quantized KV cache (llama #27390/#27438 — Q4_0/Q4_1/Q5_0/Q5_1/Q8_0 KV: prefill dequantizes to f16 scratch via kernel_flash_attn_ext_kv_f16 when ne01≥32; decode reads quantized blocks directly in the flash vec kernel). Analyzed against llama source (ops.cpp:2805- 2831, ggml-src/metal/kernels//7864-7885) and deferred: the shipped models cap at max_seq=32768 — f16 KV (#37) needs only 1.9 GB at 32K (within the 8.7 GB mmap-era RSS), so Q8_0 KV (0.96 GB) has no validation target on the current models. Revisit when a >32K-context model is the user-facing one.

4.2 Cold-start optimization to-dos (2026-08-21)

Observed: 7B run twice in a row → the second run ≈ 2× faster (Total 1.46 s → 0.77 s). Note the CLI's Total timing starts AFTER model load (main.rs:1442 infer_start), so the 2× is in the inference phases (prefill + decode), not the load. Root causes:

FactorEffectEvidence
GPU weight-buffer cold state (run 1)the 5.2 GB weights are freshly CPU-memcpy'd into Shared buffers at load; the GPU's first read hits cold MMU/TLB + page residency → slower prefill + decode on run 1decode 32.6 t/s (~84 GB/s) run 1 vs 46.4 t/s (~119 GB/s) run 2
GPU clock rampfirst GPU burst after idle starts below max clocksecondary
Model-load wall (not in Total, but real wall time)4.4 GB model load (mmap, gguf.rs:1711-1738) + Metal shader source compile (newLibraryWithSource, src/metal/ops.rs) — run 2 mitigated by the OS page cache + the Metal driver's on-disk shader cacherun 1 load visibly slow, run 2 ~free

Warm steady-state = run 2's numbers (pp30 ~0.17 s, decode ~46 t/s). Not a bug.

#ItemCurrentApproach (llama parity)ExpectedBlocker / Risk
1Precompiled metallibevery process compiles src/metal/kernels/ from source via newLibraryWithSource → DONE 2026-08-21 (#35)build-time metal compiler → embed a .metallib → load with newLibraryWithDataremove the per-invocation shader compile — 0.5B process wall 1.32 → 1.09 s (warm; the first-ever run benefits most)the standalone Metal toolchain (xcrun metal) IS installed; build.rs falls back to source compile when absent — no numerics risk (byte-identical, verified)
2GGUF mmap + zero-copy weight buffersstd::fs::read the whole file into a Vec, then copy each weight into a Shared GPU buffer (2 passes) → DONE 2026-08-21 (#36)mmap the GGUF; newBufferWithBytesNoCopy over the mapped data (llama ggml-metal-device.m:1668)remove the 4.4 GB copy pass + the 4.4 GB intermediate Vec; the GPU reads the mapped file pages directly — 7B load wall 4.3→2.7 s, peak RSS 20.9→4.7 GBnewBufferWithBytesNoCopy requires a page-aligned base → ONE buffer per mmap'd part + per-weight (buffer, offset), exactly llama's design. Cold-start regression (fixed #39): the first GPU access to file-backed pages cost ~44 ms per process (short prompts) — absorbed at load by a dummy GPU warm-up read (kernel_warmup_read)
3Persistent / server mode (note)one-shot CLI: every invocation re-loads the modelkeep the process alive and reuse the loaded model (llama-server pattern)removes reload for repeated calls — the definitive fix for the observed run-1/run-2 gapout of the current CLI scope

Item 1 removes a genuinely avoidable per-invocation cost; item 2 is mostly a memory + one-copy-pass win (the 4.4 GB disk read is unavoidable either way). Item 3 is the only way to make every invocation fast, at the cost of a daemon. Together they target the cold-start axis only — steady-state inference is unaffected (already at the §1.6 numbers).

4.3 Graph-path (MetalBackend) integration to-dos (2026-08-21+, compute-graph era)

The compute graph is now the inference path (Phase 6). Kernel-level parity work is complete (§0); the remaining gap is wiring already-tested kernels into MetalBackend — see §0.1 for the status table and measured numbers.

#ItemStatusApproach / ResultCommit (pre-rewrite)
G1Attention: wire the fast kernels✅ doneper-op Attn dispatch: nt==1 → flash_attn_enabled(hd) → gqa_attn_flash (chunked MINFER_ATTN_CHUNKS) else gqa_attn_split_f32 else classic; nt>1 → hd 64/128 + prefill_flash_enabled(hd) → attn_flash_prefill else matmul_attn_enabled() → attn_parallel_prefill else classic. Gated like the old path (MINFER_NO_FLASH/MINFER_NO_SPLIT_ATTN/MINFER_NO_PREFILL_FLASH/MINFER_NO_MATMUL_ATTN). Decode KV-growth closed: 0.5B KV206 ~122 → ~256 t/s (2.1×), 7B ~32.5 → ~49 t/s; greedy byte-identical32b0d03
G2RMSNorm: 256-thread kernel✅ doneRmsNorm dispatches rms_norm_256 when rms_norm_256_enabled() (#16), falls back to the 32-thread kernel32b0d03
G3n_out tail-row optimization✅ doneGetRows(wo/residual, tail_ids) after the last layer's wo; the last FFN + output_norm + lm_head run on n_out rows (llama inp_out_ids, plan §5.5). GraphParams.n_out joins the reuse identity; decode (nt==1) builds no reduction. 0.5B prefill pp440 ~3900–4000 t/s (+~55 % over the old path); full-nt vs reduced graphs bit-identical through the whole tail block. Also fixed two allocator liveness bugs surfaced by the reduced graph: (a) liveness used topo_order() while the scheduler executes in build order — a later consumer's buffer reuse could clobber a still-alive input (attn reused the residual h before get_rows(h) read it); (b) input buffers were freed and reused before the host-side fills finished (two inputs sharing a buffer clobbered each other).8febf4c
G4Fused decode QKV / bias+rope+store✅ donedecode (nt==1) builds one Op::FusedQKV node: a single concat matmul (blk.{i}.attn_qkv, loader-registered `wqwk
G5Fused decode FFN gate+up✅ donedecode (nt==1) builds one Op::FusedFFN node: a single concat matmul (blk.{i}.ffn_gu, loader-registered `ffn_gateffn_up) into a gate|up concat buffer + one in-place swiglu_f32_offpass (silu(gate rows 0..nf) × up rows nf..2·nf, llamaggml_swiglu_split); the down matmul reads rows 0..nf. 4 dispatches → 2 per layer. **Gated nf ≤ 16384** (7B Q4_K concat matmul od≈37888 measured slower than separate matmuls on the decode scalar kernel; 0.5B gains). MINFER_NO_FUSE_FFN=1` reverts. 0.5B decode ~303 → ~312-331 t/s (+~3 %); fused-vs-unfused decode logits max diff 0.0 (0.5B fused, 7B gated-off)

4.4 Reference: GPU profiling tooling (done, operational)

  • xctrace: /usr/bin/xctrace is a broken stub ("tool not found"); the real binary is /Applications/Xcode.app/Contents/Developer/usr/bin/xctrace.
  • Record in Instruments: Metal System Trace, Counter Set → Performance Limiters, Enable Shader Timeline on, Deferred.
  • scripts/export_trace.sh <trace> [run]: exports per-kernel durations (metal-shader-profiler-intervals), the limiter profile (gpu-counter-value), and per-forward intervals (metal-gpu-intervals); TRACE_PROC=<proc> filters per-forward intervals.
  • Interpretation: percentage counters (Limiter/Utilization/Occupancy) are ratios; bandwidth counters (L1/LLC Read Bandwidth) are cumulative — ignore magnitude.

4.5 Process (was §4.5 backfill)

After each item completes: update the §0 progress table (check, fill in the commit, update measured effect) → record the implementation + verification in the relevant §3 section → update the §1 gap numbers.


5. Historical appendix (reference only)

⚠️ This appendix is a historical record, for reference only — NOT the basis for future plans. Its architecture conclusions, metric calibers, and kernel structures may have been superseded by §1-§4. Future plans follow §0 + §4 only.

5.1 Early phase summary (Qwen2-0.5B, 130→334 t/s, 2026-07-27~08-01)

PhaseContentKey points
14 correctness bug fixesRoPE freq_scale, output_b, softmax max, stack array (see §3.1)
2Flash Attention (online softmax) + float4 vectorizationKV-parallel chunks + running max/sum fix
3SIMD-parallel attention (vec kernel)32-lane simd_dot, threadgroup barrier sync
4SIMD-parallel RMSNormmulti-simdgroup reduction + threadgroup buffer
5SwiGLU fusionsilu+mul single kernel, saves 1 dispatch/layer

5.2 Early gap-analysis records

  • KV-growth 2.2× (2026-08-01): f32 KV full re-read vs llama f16 (fix in §2.5).
  • Per-dispatch encode ~24µs vs llama ~7µs (2026-08-01 conclusion) → REVOKED 2026-08-03: encode measured at only ~1 ms/step; decode is GPU-execution-bound.
  • Q4_0 dual dispatch (quantize+matmul) (2026-08-01) → fixed: f32 activations path (§3.1 #4).

5.3 Tested-and-rejected ideas (with commits)

IdeaResultCommit (pre-rewrite)
Parallel command buffers (A1)regressed (encode already hidden), revertedb1256d5
nt==1 matmul full-block matvec rewriteat bandwidth floor, not integrated—
store_kv_both / residual_rms_norm fusionno gain, reverted—
naive 1-kernel attention (classic single-pass)4.80 vs split 4.15 ms—
f16 KV as default~3 % slower (0.5B)387d612 (kept opt-in)
Q5_0 full vectorizationonly +2.6 % matmul—
broadcast-GEMM for prefill attention2D GEMM can't produce per-head 3D scores—

5.4 Stale data tables (do not cite)

The following come from 2026-08-01/03 early measurements with different calibers (blended t/s vs pure decode; old llama baseline) — for historical cross-check only:

  • "Q4_K_M/Q5_K_M prefill 7.3×" (35-token table) — before the non-Q4_0 GEMMs existed; now filled in.
  • "decode short ~187 / long ~86 t/s" (KV-growth table) — before split attention.
  • "pure decode 2.0× / 3.2×" early gap — now 1.1-1.4× (§1.1).

5.5 flash attention: llama.cpp vs minfer — detailed kernel comparison (reference for §3.3)

Line-by-line structural comparison behind the §3.3 ~7-10× attention gap (minfer split attention 42.8 µs/layer vs llama flash ~4-6 µs/layer at nkv=430). Source: llama kernel_flash_attn_ext_vec (ggml-src/metal/kernels/), llama dispatch (ggml-metal-ops.cpp:2959), minfer kernel_gqa_attn_partial_f32 (src/metal/kernels/kv.metal), _f16 (:3127), kernel_gqa_attn_combine_f32 (:3243).

A. Overall design

llama kernel_flash_attn_ext_vecminfer partial + combine
kernels/layer1 (nwg cross-TG reduce built in)2 (partial + combine)
per-TG scope1 query × 1 head, whole KV loop inside the kernel1 query × 1 KV-head × 1 chunk
KV parallelism32 workgroups each sweep a KV slice (stride NWG*NSG*C), online-softmax partials → temp buffer + a 2nd reduce kerneln_chunks chunks, partials → combine kernel
threads/layer (0.5B decode)grid (1, 14, 32), 32 threads/TG → 14×32 = 448grid (1, 2, n_chunks), 32×gqa=224 threads/TG → 448×n_chunks
KV readhalf4/float4 direct from global (f16 cache), no explicit shmem stagingexplicit KV-tile stage into threadgroup shmem per 32-row tile

Key: the cross-TG reduction idea is identical on both sides (online softmax partials + a combine pass). All the difference is inside a single threadgroup.

B. Inside one threadgroup — the core gap

llama (dk64 template: NE=2, NL=16, C=32, 32 threads = 2 rows × 16 cols):

QK^T:  for cc in 0..C/NE:                 // 16 cache columns per simdgroup
         for ii in 0..DK4/NL:             // 16 float4 dots
           mqk[cc] += dot(pk4[..], pq4[..])
         reduce = simd_shuffle_down ×5 + simd_shuffle   // ← NO threadgroup barrier
  • QK^T accumulates in registers mqk[]; the reduction is simd_shuffle_down (16 lanes of ONE simdgroup merge via pairwise shuffles — no threadgroup barrier needed).
  • softmax update: M = simd_max, S = S*ms + simd_sum(vs) (same simd reduces).
  • PV: lo[ii] += float4(pv4)*float4(sst) register accumulate + simd_shuffle_down reduce.
  • only 2 simdgroup_barrier per tile (QK→softmax, softmax→PV), and simdgroup-level (much cheaper than threadgroup-level).
  • KV read straight as half4 from global; relies on HW cache + barrier-free pipelining.

minfer (partial kernel, 224 threads/TG = 7 simdgroups × 32):

per 32-row tile:
  1. all threads stage the KV tile into threadgroup shmem (k_tile4/v_tile4)
  2. threadgroup_barrier                          // ① make data visible
  3. each lane does one row's QK^T:
       dot = 0; for d4: qv = qhead4[d4]*kj4[d4]; dot += qv.x+qv.y+qv.z+qv.w
       batch_mx = simd_max(dot); online-softmax (corr, acc4 *= corr)
  4. threadgroup_barrier                          // ② before reusing shmem

Gap sources:

  • After every QK^T reduce and before every PV accumulate llama only needs simd-level shuffles; minfer must threadgroup_barrier — its 32 rows are processed serially by the same simdgroup's lanes (for j0 loop, tile_sz=32) while different lanes of the 224-thread TG handle different heads (gqa=7). Result: 2 global barriers per tile, 28/layer at nkv=430; each barrier stalls the whole TG to the slowest lane.
  • llama's 16-column reduce stays inside one simdgroup (32 threads) — zero cross-simdgroup sync.
  • minfer's acc4 is register-accumulated but serialized by the barrier sequence; llama's mqk/lo register accumulation pipelines through the shuffle chain.

C. Why the gap is ~7-10× and not smaller

Of minfer's 42.8 µs/layer (nkv=430), the 28 threadgroup barriers + per-tile shmem stage-in/out dominate: a barrier makes all 224 threads wait, while each 32×64 tile does only ~2048 MACs — swamped by sync overhead. llama's simd reductions drop sync to near zero, and half4 reads with no shmem staging keep the memory pipeline unbroken.

D. Port (option C) — actual change surface

For Qwen2.5-0.5B (hd=64, nh=14, nk=2, gqa=7, nt==1 decode):

  1. KV-layout check (prerequisite) — DONE 2026-08-14, PASS: llama's K/V physical layout IS [nkv][nk*hd] (all KV heads packed in dim0 of the cache tensor, token stride nk*hd), and the flash kernel reads it with nb11 = nk*hd*elem (token), nb12 = hd*elem (head), ns10 = nk*hd — the same indexing minfer already uses (k + ki*stride_kv + hk*hd, stride_kv = nk*hd). No cache rework needed; the host passes the stride args mirroring minfer's layout (nb10 = elem, nb11 = nk*hd*elem, nb12 = hd*elem, ns10 = nk*hd). Evidence chain in §3.3.
  2. Single-kernel rewrite — DONE 2026-08-14: fixed DK=DV=64, NE=2, C=32, NSG=1 (0.5B/7B decode dims), dropping llama's function-constant system (hardcoded constants; n_chunks is the host-tunable grid depth instead of llama's fixed NWG). Shmem sq4 + ss + so4 (no sm[] — mask inlined), QK^T/PV via float4 + simd_shuffle_down(8,4,2,1) + simd_shuffle(·, NL*ty) broadcast (the reduce routes the full-head sum to lanes 0/16).
  3. Cross-TG reduce — DONE: reuse minfer's existing combine kernel — the flash kernel writes the same {M, S, O[hd]} partials (strided C=32-block chunking, interchangeable with the split's contiguous chunking).
  4. isolation tests — DONE: tests/flash_attn_isolation.rs (scalar CPU ref, multi-nkv incl. partial/empty chunks, nt 1-2, f32+f16; flash-vs-split A/B through the shared combine) + end-to-end byte-identical A/B (MINFER_NO_FLASH=1).
  5. prefill untouched: nt>1 keeps the current 3-pass parallel attention.

Risk points:

  • KV layout mismatch was the most likely failure — CLEARED (2026-08-14, see D-1 above); the layout is compatible as-is.
  • Function constants → hardcoded constants is NOT a line-by-line translation; shmem offsets must be re-derived (sgitg*SH terms are all 0 for NSG=1, which simplifies).
  • simd_shuffle_down lane participation (NE=2 uses only the NE>1/NE>2 branches; simd_shuffle_down(mqk[cc], 16) folds 16 cols onto lane 0) needs careful index alignment.

5.6 Glossary — variables and functions used in this document

This table covers only this document's variable/function symbols. Concept-level terms (quant formats, kernel techniques, methodology) are classified corpus-wide in GLOSSARY.md; the ggml layout symbols below are also collected there (L2).

Common (both kernels, GGUF/model dims):

SymbolMeaning
n_embd / nemodel hidden size (embedding dim), e.g. 896 (0.5B), 3584 (7B)
n_head / nhnumber of query heads, e.g. 14 (0.5B), 28 (7B)
n_kv_embdKV cache head dim (hd_kv), e.g. 128 (0.5B); may differ from n_embd
hd / hd_kvattention head dim: n_embd/n_head (64 for 0.5B) and KV head dim
gqagroup size = nh/nk (e.g. 7 for 0.5B) — GQA heads share a KV head
nknumber of KV heads (n_head_kv), e.g. 2 (0.5B)
nkvnumber of KV cache positions (context length so far)
nkttotal KV cache capacity (max positions)
ntnumber of tokens in the batch (prefill nt>1, decode nt==1)
nqtnumber of query tokens (prefill)
odoutput dim of a matmul (rows of the weight)
idinput dim of a matmul (cols of the weight)
nfFFN intermediate dim (gate/up/down)
positionsper-token KV position array; nkv = positions[t] + 1
max_poslargest position used so far
n_chunkssplit-attention chunk count (MINFER_ATTN_CHUNKS, adaptive)
BcKV tile size in rows (32) used by the minfer attention kernels
n_layers / n_layertransformer layer count (24 for 0.5B, 28 for 7B)
Q4_0/Q4_K/Q5_0/Q5_K/Q6_K/Q8_0GGUF weight quant types (see §Quantization in AGENTS.md)

llama ggml tensor/strided-layout args (used in flash/matmul kernels):

SymbolMeaning
ne00..ne33tensor dimensions: ne0x=dim0(dims of x-th src), ne1x=dim1, ne2x=dim2, ne3x=dim3
nb10..nb33byte stride of each dim for src1 (nb10=elem stride in bytes, nb11=row/token stride, etc.)
ns10 / ns20nb11/nb10 and nb21/nb20 — element count per head/row/token (used as the flash KV inner-loop stride)
ne11KV cache length dim (nkv) in flash-attn args
ne12/ne13KV head dim2/dim3 (GQA heads, batch)
nwgnumber of workgroups (32 for flash vec, each sweeps a KV slice)
nsgsimdgroups per threadgroup (flash vec: 1; blk: 4-8)
NWG/NSGfunction-constant copies of nwg/nsg inside the kernel
NEcolumns per simdgroup in a flash tile (dk64: 2)
NLlanes per column = NW/NE (dk64: 16), NW = 32 (simd width)
Cflash tile columns (32), SH = 4*C shared memory per simdgroup
DK/DVflash key/value head dims (dk64/dv64 templates; Qwen2.5-0.5B: 64/64)
DK4/DV4DK/4, DV/4 (float4 element count)
PK/PVPAD2(DK,128)/PAD2(DV,128) — padded head dim for shmem
NL/NEsee above (flash vec tile geometry)

Metal thread variables:

SymbolMeaning
tgpigthreadgroup position in grid (.x/.y/.z = the 3 grid dims)
tiisgthread index in simdgroup (0..31)
sgitgsimdgroup index in threadgroup (0..nsg-1)
simd_max/simd_sumSIMD (warp) reduction builtins
simd_shuffle_downSIMD shuffle reduce (lanes exchange values pairwise)
threadgroup_barrierTG-wide memory + execution barrier (all simdgroups)
simdgroup_barriersimdgroup-wide barrier (cheaper, single warp)
float4/half44-wide vector types (128-bit / 64-bit) used for SIMD loads

minfer kernels (src/metal/kernels/ unless noted):

FunctionMeaning
kernel_gqa_attn_f32/f16classic single-pass attention (grid (nt,nk), sequential KV loop) — prefill-past, superseded
kernel_gqa_attn_partial_f32/_f16split attention pass 1: online-softmax partials (grid (nt, nk, n_chunks))
kernel_gqa_attn_combine_f32split attention pass 2: merge partials (grid (nt, nh))
kernel_attn_scores / kernel_softmax_attn / kernel_attn_output3-pass parallel prefill attention (nt>1, barrier-free)
kernel_store_kv_f32/f16KV cache store (f32 or f16 cache)
kernel_q4_0_mm_f32Q4_0 prefill GEMM (simdgroup, nt≥16)
kernel_q6_k_f32_matmulq6_K matmul (stride-2/float4 layout, decode; was the 3× gap, fixed §3.3)
kernel_mul_mv_q6_K_f32_implllama's q6_K kernel whose layout was ported
kernel_rms_norm_fuse_implllama's fused rms_norm kernel pattern (256-thread port kernel_rms_norm_f32_256)
kernel_flash_attn_ext_vec / _blkllama flash kernels: vec = decode (nb<20, simd-shuffle), blk = prefill (simdgroup_matrix)
kernel_cpy_f32_f16copy f32→f16 (used for f16 KV cache)
kernel_get_rows_q4_0Q4_0 embedding gather

minfer env vars:

VarMeaning
MINFER_ATTN_CHUNKSoverride split-attention chunk count
MINFER_CACHE_TYPEf16/f32 KV cache override; unset → auto (7B class f16, small f32, #37). q8_0 (C4, packed) is read on Metal since #310 via mechanisms A/B; the classic kernel_gqa_attn_q8_0 family remains the fallback
MINFER_GEMM0 = disable the Q4_0/non-Q4_0 prefill simdgroup GEMMs
MINFER_NO_FUSE_QKV1 = disable fused QKV/FFN-gu decode matmuls
MINFER_SPLIT_CBN = split the decode into N command buffers
MINFER_TIMING1 = per-category decode GPU timing split
MINFER_TRACE1 = record per-dispatch labels (GPU hang debug)

CUDA Inference Path — Optimization History and Current State

STATUS (2026-09-06, post-r60): history-organized reference. This document was restructured from a part-based roadmap into a history-ordered record: §0 is the master history table (the outline — every landed, reverted, or measured lever with its commit and perf delta), §1 is the current state, §2 is one chapter per table row, and §3 holds the appendices (env-gate reference, methodology, and the pre-Phase-7 legacy history). Single-sourced implementation records: docs/CUDA-BACKEND-DESIGN.md (Phase 7a–7e) and the per-step documents of §2 (Phase 8: docs 01–11 plus the supplementary records 78–79 — the former CUDA-FOLLOWUP-PLAN.md was consolidated into them and retired on 2026-09-10); the per-round MMQ redesign records are mirrored in docs/LLAMA-CPP-MMQ-ANALYSIS.md §11. Default env = the verified 1.080×-vs-llama path; MINFER_MMQ=0 = the legacy f16 escape (§1, Appendix A).

§0 Master history table — the complete optimization record

One row per optimization step: every lever that landed, every lever that was reverted, and the measurement-only rounds that directed the campaign. Chapters in §2 follow this table row by row.

Reading conventions.

  • Perf column: whole-prefill tok/s for the 7B q4_k_m model on DGX Spark GB10 unless noted. The absolute anchor moves with the prompt length used by each session (2630/2659 → 3325/3314/3354 tokens ≈ "pp2K/3314-eq") and with machine state (co-tenant load); every row's numbers are interleaved same-binary A/B medians within one session window — cross-session absolutes are not comparable (the r59b lesson). "—" = not measured at whole prefill (kernel-level or decode-level metric instead).
  • vs-llama column: whole-prefill vs llama-bench at the 3325-eq (later pp3314) anchor — 3324.42 tok/s (clean machine) and 3323.29 (r59b window). The 8m–8p/P5 rows' early multipliers are vs the 3401 @2K llama-bench figure instead. Before r37 the campaign ran on the opt-in MMQ path, so no whole-prefill vs-llama was recorded at the 3325-eq anchor for those rows ("—"); the default f16 path sat at 1.43× (2340–2370 vs 3401) from P5 until the q6_K/FA lines moved it.
  • Commit(s): code commit first, record/docs commit second where both exist. Twelve hashes quoted in the older session records are pre-amend duplicates that no longer resolve (e.g. the r59 record's feb37de); this table cites their reachable twins (same subject — see the r59 chapter for the one case worth naming).
  • Cross-reference: the r28→r59 MMQ-redesign rounds are mirrored, round for round, in docs/LLAMA-CPP-MMQ-ANALYSIS.md §11.8–§11.37 (11.8 = r28 Phase-2 outcome, 11.26 = r47, 11.34 = r55, 11.36 = r58, 11.37 = r59).
  • r26–r27 do not appear in the record — the P6 round numbering skips from r25 to r28 (the intermediate commits are the MMQ-analysis §11 design/doc work, kept in docs/LLAMA-CPP-MMQ-ANALYSIS.md).
StepLeverCommit(s)Perf before→after (anchor)Δvs-llamaStatusOne-line lesson
Phase 7CUDA backend: raw FFI device layer + graph backend (7a–7e)0dc2a54baseline 7B @2K 30.7 tok/s—~110×LANDEDresident weights + per-op dispatch ended the Part-IV ping-pong
8m/8m②tiled wmma f16 prefill GEMM + cp.async tile stagingba3f317, cdc659930.7 → 294 → 1204 tok/s @2K39×~2.8×LANDEDone tensor-core GEMM over all 8 types replaces per-token weight re-streaming
8nFA-style tiled prefill attentioncb66fca176 → 8.5 ms/layer @2K20×—LANDEDonline softmax + register O accumulator; 256 B P stride avoids a score-clobber race
8odecode-start CPU stalls killed65b686cfirst decode step 724 → 35 ms20×—LANDEDCow::Owned clone + eager concat probe cost ~1.6 s per graph rebuild
8ppersistent f16 weight cache + fused dequant-in-GEMM2992f57 (+b9e7a91 docs)7B @2K prefill → ~1400–1500~4.7× vs 8m~2.3×LANDEDdequant once per weight at load (≥2 GB gate); exposed a latent Q5_0 misaligned load
8bKV f16 on CUDA (store_kv_f16 + f16-KV attention mirror)f7b00367B @2K decode +~11%+11%—LANDEDauto-f16 at n_layers×n_kv_embd ≥ 8192 (MINFER_CACHE_TYPE override); f32 accumulation
8cprefill Q8_0-activation GEMM, shape-gated69a27c50.5B @3.6K 1005 → 1246 tok/s+24%—LANDEDthe shape gate is load-bearing: −63% at 7B ffn_down (weight-bound shapes stream slower)
8dsplit-K flash-decoding decode attentiona5af60f7B @2K decode 10.1 → 13.7 tok/s+36%—LANDED28 warps → an 8-way KV-split scan; superseded by R4's dim-parallel rewrite
8fQ5_K + Q5_1 f32-activation kernelsb959ec90.5B q5_k_m admitted to CUDA (was CPU wholesale)——LANDEDthe all-or-nothing gate needs a kernel for EVERY matmul weight type
8e/8e②decode MMVQ (dp4a, per-type kernels, shape gate)b7b8e73, 1298cb2, 1d282357B decode +37% (q4_K), then q6_K/q5_K+37%—LANDEDinteger dp4a dots + llama.cpp MMVQ_PARAMETERS_GB10 launch table
8lllama.cpp parity benchmark (the decode + prefill gap sheet)acca28fdecode 1.15–1.76×, prefill 18–110×——MEAS-ONLYfound the Q5_K registration gap (51.6 → 246.3 tok/s, 4.8×); the sheet ranked 8m–8p / R1 / the r-campaign
8qQ5_0 CUDA enablement + u16 qh loads9f419f90.5B q4_k_m 148.7 → ~1200 prefill / 56.9 → ~306 decode——LANDEDa 22-byte block's qh word is not 4-byte aligned — two u16 loads; the CPU fallback eliminated
R3-A1single-split prefill (tail_ids input at graph head)029a9a44 splits → 1 per prefill forward——LANDEDa mid-graph declared input forced 2 extra full-stream syncs per forward
R3-A2pinned D2H logits readback, no redundant clonea213c89parity-to-slightly-ahead under load——LANDEDpageable readback paid a driver-internal pinned bounce
R3-Bprefill capture defaults ON (3-run protocol)761e236repeated identical-nt prefills capture automatically——LANDEDone-shot CLI prefill never reaches 3 runs and pays nothing
R1int8 MMQ prefill GEMM, opt-in MINFER_MMQ=140e97c9155 (co-tenant) / 412 quiet vs f16 630–880 / 1460; 441 in the r7–r8 window——LANDED (opt-in; superseded by raw line)parity-clean but ~2.9 TMAC/s vs llama ~24 — the 8× gap was unprofiled
R2MMVQ weight-streaming rework (per-thread sub-pairs, uint4)6df3245tg128 42.2 → 45.1; @2K 36.7 → 38.8+6.9% / +5.7%1.05× / 1.16×LANDEDper-sub-block nibble re-reads doubled load instructions; L1 hid the bytes, not the issue stream
R4decode split-attention dim-parallel rewrite70f57db@2K 39.2 → 43.2–45.1; tg128 45.1 → 47.5+10–15%~0–4% / aheadLANDEDLOCAL-memory float4 oc[32] accumulator = ~80 MB/layer of local traffic
P5·0FA prefill P·V on tensor cores86ca78c10.06 → 4.24 ms/layer; @2K prefill +15%+15%2.37×LANDEDwmma for P·V halves the FA pass
P5·1elementwise vectorization (store_kv, convert)d713e6e1435 → 1493+4%2.28×LANDED1-elem kernels left 15/16 of every transaction unused
P5·2128-wide GEMM tiles (TM=128)725e3071493 → 2267+30%1.50×LANDEDhalves B-panel L2 re-reads and barriers per FLOP; fb[1] offsets +16 elements, not rows
P5·3FA softmax on all 8 warps + padded smem rowsfc07c042267 → 2365–2371+4–5%1.44×LANDED256 B rows ≡ 0 mod 32 banks = 8-way ldmatrix conflicts; +8-half stride fixes
P5·4GEMM k-step KS=641365c821464 vs 2345−38%—REVERTED56 KB footprint halves resident blocks — depth vs occupancy inverted
P5·negTM=256 GEMM tilea189837−3%−3%—REVERTEDwider tile, same wall
P5·negin-kernel f32→f16 A staging (AF32 mirror)b254c22 (WIP a3b0dcd, 69c3933, aa40ed3)−8% end-to-end; parity hole open−8%—REVERTEDthe convert pass is cheaper off the hot path (r23 later quantified it at 6%)
r5–r6MMQ re-rank + structural rewrite spec1e0673f, 491eb5cKD=4 re-negative (427 vs 438); 4-warp 32×32 tile 399 + parity hole; dequant pass is at LOAD, not in the wall——MEAS-ONLY + REVERTEDdepth (KD=8) beats occupancy for the word-staging kernel; spec: raw-byte smem, dequant at mma time
r7–r8raw-byte MMQ kernel + 128-token wide tile; FA KV L2-prefetch probed440d16, d9d626a, a41eac0, ef9d5b4, 87bade0raw 472 vs 441 (+7%), quantize 129→74 ms; FA probe null (2319–2345 vs 2345)+7%—LANDED (raw) + REVERTED (probe)wide KD=8 first measured 2124 = phantom (silent smem-cap failure) — guard the cap; GQA already keeps KV L2-hot
r9llama.cpp MMQ reference decoded; shape axis closed84831d2narrow cp.async KD=8 481 = local optimum (6-shape matrix)——MEAS-ONLY~0.018 inst/MAC/thread vs our 0.133 — the ratio, not tile shape, is their speed
r10reference inner-loop decomposition ported025a69f462–468 vs 470 (flat, ~6.4 TMAC/s)~0—REVERTED (edits lost to a post-checkout hook; measurements valid)same decomposition as llama still 5× slower — residual is ILP depth/ldsm/tile
r11ILP-chain reading verified (tile ne = I·J/32)83e5580, 6f29e65———MEAS-ONLYthe I·J/64 ne was the AMD MFMA branch; NVIDIA needs 4 C regs per m16n8k32
r1216-chain warp tile + ldmatrix774a116wide-16 KD=4 1020–1058 vs narrow 441–481 (~2.3×); KD=8 973–9952.3×—LANDEDaccumulator depth wall passed; 4-warp 32×32 and 48B A-padding both negative
x-tile256-token wide-MMQ blockf061cb8~942 vs 1035–1058; kernel +27%−9%—REVERTEDA re-reads (327 MB) already ≈2× B re-reads (152 MB) — trading B for A is strictly negative
j-tile128tok × 256od A-reuse (outer jh loop)4993804~1067 (+2.6%, bar 1150 missed)+2.6%—REVERTEDL2-byte savings do not convert to time at 1 block/SM — latency-bound, not byte-bound
cp.async-dbcp.async double-buffered raw staging784786d~1034 (neutral; baseline band 1035–1050)~0—REVERTEDkernel is L2-throughput-bound, not MLP-starved; re-timing identical bytes wins nothing
r13ncu counter forensics vs llama.cpp + FULL/MINIMAL staging fixes5ca037dFULL 865–871 (−16%); MINIMAL +0.3% (noise)——MEAS-ONLYthe 3× gap is the per-MAC warp-instruction stream (10.14 vs 6.06 M/GMAC), not bytes
r14B-fragments via one ldmatrix.x4 + widened scale loadsc64cd991225 @KD=4 / 1273 @KD=8; kernel 3.632 → 2.378 ms+18.5% / +23–30%—LANDEDslot-major 48B qb8 tiling is conflict-free; fewer/wider smem ops, zero staging-ALU growth
r15f32-accum s8 mma probe (dead) + rank-1 term2 rescaleb999e9a1295 @KD=8; inst 9.45 → 8.69 M/GMAC+1.9%—LANDEDf32-accumulate integer mma does not exist (ptxas probe); merged-chunk rescale is mathematically invalid
r16narrow kernel gets the rank-1 fold151fa97480 vs 447–473+1.7% (noise-band)—LANDED3-site port of r15; narrow is not the perf path
r17wide warp remap 32od × 64tok9d09a81inst −5.8%; wall +0.9/+1.0% (noise)~0—REVERTEDpure per-MAC instruction cuts pay ~0 wall while SM% sits at ~30 — stall-bound
r18load-time B pre-expansion (bulk-copy staging)0a26b35KD=8 +0.9% (noise); KD=4 −19% median~0—REVERTED+5.8 GB for a −4.5% kernel-time that does not reach the wall; EB/SB machinery preserved for future L2 experiments
r19weight L2 residency (__ldg, persisting window)072dd9a__ldg +2% (noise); L2WIN −50%~0 / −50%—REVERTEDweight tiles already re-read from L2; a persisting carveout starves C stores/activations/KV
r20split-phase A staging5ac89171230.4 → 1317.7 @KD=4; 1275.8 → 1319.9 @KD=8+7.1% / +3.5%—LANDEDthe gap carrier is long_scoreboard (97% of named-stall excess) in the LDG→STS chains; freed stalls re-saturate on lg_throttle
r21coalesced block-linear A staging3c009ccsectors −28.6%, lg_throttle −90%, wall −2.3%/−0.7%−2%—REVERTEDstall mass is conserved: sectors/queue are not the binder, warp-instruction count is
r22qa8 XOR swizzle (+ d/ssum fold negative)5b40058KD=8 1329.6 vs 1311.8 (+1.4%, 3/3); op_ld 16.86 M → 0+1.4%—LANDED (fold REVERTED)precompute all 8 A-frag offsets once — per-ldmatrix address ALU eats freed wavefronts; d/ssum loads are L1 hits
r23f16-path full-graph wall decomposition + FA_TKV=32 occupancy lifte8c348dlift: occ 16.7 → 32.68%, kernel −6.7%, wall −0.3%−0.3%1.43× (f16 path)MEAS-ONLY + REVERTEDFA's 2.5×/layer gap is structural (llama keeps 128-wide KV tiles); GEMM is 74% of the f16 wall
r24scheduling ladder: tile-order swizzle + persistent blocksd90b3e9swizzle −2.3/−4.7%; persistent −3.3%negative—REVERTEDdefault B-hot x-fastest order is best; no wave-quantization tail exists to remove
r25SASS opcode-class census + kd-unroll attempt8658f1binst −9.7% (surplus halved to +46.6k/tile); wall +0.37/+0.49%~0—MEAS-ONLY (census is the deliverable) + REVERTED100% of the +25% inst surplus is support instructions (int ALU 77%) — but it is wall-inert at 1 block/SM
r28Direction-A raw-nibble NB kernel (2 blocks/SM)0957a08 (design 2f783a3)1375.2 → 1410.4 @KD=8; 45,056 B smem, 123 regs+2.56%—LANDEDoccupancy was the binding resource; unsigned-nibble + rank-1 rescale is the B-frag contract
r29NB kd-loop unrollbfe6bba1387.9 → 1426.8; int ALU −25%, inst −6.5%+2.80%—LANDEDr25's wall-inert int-ALU cut becomes real at 2 blocks/SM — occupancy unlocks instruction cuts
r30SWAR word-granular B-nibble unpack0071b31+0.54% (noise); SASS byte-identical~0—REVERTEDr29's unroll already induced the exact CSE — check SASS before writing the lever
r31q-major sda scale-read repack851a896 (+76d495a docs)1424.10 → 1439.40; LDS.64 32→0, LDS.128 16→32+1.07%—LANDED (sub-bar)naive q-major repack was 2-way conflicted — region-split layout is conflict-free
r32finite-lever sweep (staging hoist, epilogue widening)153d28cepilogue +0.46% (structurally capped ~0.4%)~0—REVERTEDstaging addressing already hoisted by ptxas; run-once epilogue cannot clear a bar
r33hybrid inner-loop port (llama j0/n order)697ef04median −0.25%; SASS byte-identical~0—REVERTED (hypothesis falsified)ptxas already schedules the 8-mma + rescale loop optimally — source reorders are SASS no-ops
r34quantize-transpose prepass (A-side layout transform out of the kernel)ba977bf1364.2 → 1496.8 @3354 tok; prepass 0.908×; 103 regs+9.72%—LANDEDthe residual was layout-transformation locality (llama's quantize_mmq_q8_1 design), not instruction composition
r35sda d/ssum scale pre-decode6112db3−0.46%; SHF 64→0 but wall flat~0—REVERTEDdecode ALU hides in the IMMA shadow — removing int/fp ops that fill idle slots frees nothing
r36A-frag wavefront economics (H1/H2 endpoint)f44fc441.76× shared wavefronts/IMMA but 1.85× wavefronts/s at equal IMMA rate——MEAS-ONLY (H1 refuted)the MIO pipe is not scarce; llama's edge is A-fragment reuse (0.125 vs 0.5 LDSM/IMMA), a tiling property
r37post-parity whole-prefill attributionea234f11521 tok/s; q6_K GEMM 1094.7 ms = 51.2% of wall at 6.38×/GMAC—2.15×MEAS-ONLYMMQ made q4_K fast and left q6_K on a slower-than-f16 path — the next lever is a different kernel
r38q6_K BT-style raw-byte mma kernel (KSPLIT=2, KDR=4)75aabb91518.4 → 1561.9; q6_K 368.9 → 221.8 µs/GMAC (1.66×)+2.87%2.13×LANDEDq6_K is 16 sub-blocks of 16 (not 8×32) → two m16n8k16 with separate dsc; KDR=8 regressed to 1097.8
r39q6_K KDR=2 double-buffer (A+B pipelined)f2b9e541568.7 → 1777.5; attn_v kernel −19.7%+13.3%1.87×LANDEDdoubling every plane at KDR=4 = the 1-block/SM trap; KDR=2 hits the same 29,696 B with real overlap
r403rd resident block via __launch_bounds__(256,3)65ecef71784.0 → 2015.6; kernel −23%+13.0%1.65×LANDEDthe 0-spill gate is disproven-immaterial: 80 regs + 4 B spill beats 87 regs at 2 blocks
r41q6_K B-expand widened to uint4 groupsb891e1b (+aa82e8f docs)1979.9 → 2605.2; kernel 1.70 → 0.654 ms (−61.5%)+30.7%1.27×LANDED32 per-byte ql/qh LDGs per thread-kt = the L1TEX scoreboard (85.5% → 33.6%)
r42q6_K stage-wide dsc scale reada1421e6−0.19%; L1TEX traffic down, stall share unchanged (33.6%)~01.27×REVERTEDcutting dsc bytes does not cut dsc latency — the stall is at the I2F consumer
r43PC-sampling attribution + pre-expand-B (parity FAIL)b7fa305attribution: B-expand recomb 45% + A-sts 28% + dsc I2F 26%; W_exp byte-correct but diff 448—1.27×MEAS-ONLY + REVERTEDattribute stalls to the consuming instruction; byte-correct data at wrong offsets = stride mismatch next door
r44W_exp stride mismatch root-caused; dense-index fix6d02017parity green; kernel −10.9% but wall −0.42%~01.27×REVERTEDdense W_exp indexed with the padded raw-W stride = the paradox; removing recomb only transforms the latency
r45cp.async the q6_K A-side staging9825ffdkernel −10.2%, longsb −18%, wall −0.34%~01.27×REVERTEDafter r41 the q6_K GEMM is no longer the bottleneck — a faster kernel that is not the wall does not reach it
r46 (FAP1)FA audit + FA_TKV 64→32 + S/P row paddinga186f51FA kernel 5.16 → 4.58 ms (−11%); wall +0.27%~01.27×REVERTEDFA was already wmma + online-softmax — occupancy-starved and S/P-conflicted; but not wall-critical yet
r47converged-regime wall decomposition (r37 table stale)11e36402585–2623 tok/s; q6_K 1094.7 → 196.4 ms; FA = #1 residual 5.72× (125.8 ms, 10.2%)—1.27×MEAS-ONLYq4_K 1.06× and q6_K 1.13× both at parity — recommend FAP2 register-resident softmax (2× → −4.9% wall)
r48 (FAP2)register-resident softmax in fa_prefill_f16kvd38744d (+7e2ee62 docs)2603.5 → 2749.9; FA kernel 5.16 → 2.12 ms (2.43×)+5.6%1.21×LANDEDsoftmax on the QK^T accumulator fragments; P built in-register as the P·V A-operand; K col_major vs V row_major is the trap
r49A-quantize prepass shared-A dedup (window memoization)87a75a32734.1 → 2797.5; prepass 193 → 110 launches, 118.4 → 83.9 ms+2.32%1.18×LANDEDq/k/v and gate/up share one A — the prepass was re-quantizing it per GEMM; cache keyed on (src ptr, nt, id), cleared at any non-MatMul node
r50FA_TKV 32→16 occupancy trial9128468−0.5% / −0.01% (3 blocks/SM reached); greedy identity lost~01.18×REVERTEDoccupancy gain cancelled by doubled per-tile sync/softmax overhead — FA_TKV reduction is a dead lever that also breaks byte-identity
r51producer-fused A-quantize, mode 1 (rms/swiglu emit pad40_t)cf1ed4b (+bf0c986 docs)2803.4 → 2856.4; prepass 110 → 28 launches, 83.0 → 10.1 ms+1.89%1.16×LANDEDthe quantize input is L2-hot in the producer; register-resident swiglu quantize was rejected (uncoalesced f32 stores)
r52fused-producer phase 2: skip-write mode (MINFER_MMQ_A_FUSE=2)910d967 (+fb659f7 docs)2855.7 → 3011.3; fused producers 151.9 → 86.5 ms+5.45%1.09×LANDEDf32 output is provably dead under the window enumeration; dead-write backstop turns violations into loud errors
r53q6_K bundle: pre-expanded-B dense W_exp + cp.async B staging83fee77 (+4907d9f docs)3024.7 → 3176.9; ffn_down kernel −20.5%; +1.52 GB device+5.03%1.05×LANDEDr44 (removes WORK) + r45 (removes WAIT) are individually wall-neutral and compose — the basket thesis
r54MINFER_MMQ_Q6K_EXP opt-out of the W_exp plane3252e96 (+b860b7e docs)default 3181 (unchanged) / EXP=0 3020.7; 7636 vs 6182 MiB−5.04% for 1.52 GB1.04×LANDED (gate)memory-for-speed knob with a three-way liveness label (exp=off vs fallback!)
r55fused-swiglu roofline audit + one-shot prefill CUDA-Graph decision83c3c67swiglu at 242 GB/s = 89% roofline (cap +0.74%); capture ≤ +0.1% + capture-illegal malloc—1.05×MEAS-ONLY (both documented skips; campaign CONVERGED)bound the roofline before coding — no implementation of this kernel can clear the bar
r56q6_K A-side bundle: A cp.async + W_dsc f32 plane4cf7c74 (+29084de docs)3138.6 → 3212.5; ffn_down −5.9%, attn_v −4.2% kernel; +363 MB+2.35%1.035×LANDEDpost-r53 the A-side wait and dsc consumer became the wall — r45's mechanism finally lands in a bundle
r57FA KV staging double-buffer (FA_TQ=48)c3268ccgreedy-32 identity lost at token 19 (×2 attempts)—1.035×REVERTEDFA_TQ is a tile size too — the r50 rounding caveat applies to any FA tile change
r58q4_K BT structural spec + cp.async-db2 transplant093ae41 (code reverted)2819.3 vs 3227.6 (−12.6%); spec: gate/up −37% potential, ceil-waves 2.6%−12.6%1.03×REVERTEDa mechanism whose COST depends on the granularity of what it replaces is amortization-bound (the r45 mirror)
r59q4_K W_dsc f32-pair plane + pre-warm/pre-grow riders36a481f (+15c04ba docs)recorded +26.2% (2843.2 → 3588.8, co-tenant) — superseded by r59b; bt kernel busy −30.9%; +1456 MB+11.1% (true)~aheadLANDED (Δ corrected)gate/up (−37%) and q/o (−18%) carried it, not ffn_down — the r58 "staging scales with kt" reading confounded decode cost with A-plane DRAM traffic
r59bclean-machine re-measure + baseline-poisoning correction074ca94definitive 3590.8 (HEAD) vs 3232.0 (fresh baseline rebuild) = 1.080× llama-bench 3323.29 @pp3314+11.1%1.080× (ahead)MEAS-ONLY (correction)the r59 "baseline" binary was the stale r58-delta build (−12.5%) — anchor every A/B baseline behaviorally in the same window
r60PROMOTION: the verified MMQ gate set flips DEFAULT-ON57edcf6 (+7029ee4 docs)default ≈3578–3599 (~3581); MINFER_MMQ=0 legacy f16 ~2226–2353-class; planes +3.27 GB—1.080×LANDEDpromotion = default-on with "0" opt-outs (r54 pattern); the bisect caught the mode-2 multiturn break → NB-BT-only guard
D1decode @1641 KV attribution: gqa_attn_split_partial is 100% of the KV-scaling wall (34.1 µs/launch, 76.5% long_scoreboard); ATTN_SPLITS sweep = dead endmeasurement-only (/tmp/d1/)tg128 49.3 (KV~1) / 47.2 (@1641) vs llama 49.41 (tg128)—0.956× (tg128)MEASUREDprobe-verified: staging-depth changes bitwise-safe (ndiff=0); ATTN_SPLITS changes reorder the float sum
D2explicit K+V register staging in gqa_attn_split_partial (4-row window staged before the softmax chain); cp.async smem pipe + pair-lookahead measured worseD2 commit (this row)decode @1641 47.2 → 48.2 (+2.0%); kernel 34.1 → 19.4 µs/launch; tg128 49.4 flat—0.975× (tg@1641)LANDEDbitwise-identical end-to-end (greedy-32/256 byte-identical); probe −42%, nsys −43%, wall +2.0% agree
D3b-1bdown-q6K pipelined MMVQ q6_k_q8_mmvq_v2_pf (npair>256: both serial units' weight+q8 loads issue up front)f1825b57B decode tg128 48.03→49.47 (+3.0% SEP), @1641 46.66→47.94 (+2.7% SEP); 14B tg128 22.80→22.90 (+0.44%), @3254 21.01→21.06 (+0.24% SEP)+2.7–3.0% (7B decode)0.970× (7B tg@1641, this window)LANDEDbitwise-identical (114/114 dump memcmp, greedy-256 byte-identical, suite 169/0/3); for npair>256 the second serial unit's exposed load latency WAS the 198.9-vs-220 GB/s gap — D4-2 CORRECTION: the 7B numbers are void (the dispatch dropped units 512..591 on npair-592 rows; the "gain" was mostly the missing work — see the D4-2 chapter); 14B numbers stand
D3b-1aattn_v-q6K off the padded-f32 kernel: (a) MMVQ routing via a lowered od*id>=24M gate — NOT bitwise (MMVQ quantizes activations to q8, different accumulation semantics); (b) NSG 2→1 row→warp re-map — bitwise-green but kernel 36.4→39.9 µs (2× warps = 2× y re-read L2 traffic)reverted (both routes)14B tg128 −0.74%, @3254 −0.33%——REVERTEDthe padded kernel is not warp-starved; y re-read traffic scales 1:1 with warp count — rows-per-warp is the only bitwise-free knob and 2 is already the sweet spot
D3b-1coutput-head dynamic block size (npair 160 → 160-thread blocks, warp-count-bounded mmvq_block_reduce)reverted (patch /tmp/d3/patch_1c.py)7B @1641 +0.26% (SEP); 14B tg128 +0.04%, @3254 +0.09%——REVERTEDGB10's 1536-thread/SM limit: 9 blocks×160 live threads ≈ 6×256 allocated (960 live) — the idle-thread win does not exist at 14B shapes
D3b-2short-KV combine skip (single-split path for nkv ≤ threshold)not implemented———ANALYSIS-NEGATIVEsingle-split ≠ 32-split partial+combine bitwise for ANY nkv>1 (the merge reorders the float sum — D1's split-count evidence: ndiff 3.6e-3 of outputs, max|Δ|~3e-9); the split grid is frozen by CUDA-graph replay capture; the bitwise-safe residual (combine early-out of empty splits — exact +0.0 terms) is ≤ ~15 µs/step, below every bar
D3a4-warp fattn-vec-style split-attention rewrite (gqa_attn_split_partial_h4w, hd=128: 128 threads, K/V streamed, Q in registers, 8-lane subgroups, 32-row windows/warp, smem LSE merge; grid unchanged, replay-safe)reverted (patch /tmp/d3/d3a_kernel_patch.diff, findings /tmp/d3/D3A_FINDINGS.md)kernel 14B @3254 73.4 → 68.4 µs (−6.9%) but 7B @1641 21.1 → 34.7 µs (+64%); wall 14B @3254 −0.95% (noise), 7B @1641 −3.4% (real), tg128 noise-level——REVERTEDrows-per-warp pathology: rpw = ceil(ceil(nkv/32)/4) = 26/13/1 at @3254/@1641/tg128 — the 32-row window idles 59–75% of lanes below rpw≈16 and per-block fixed costs amortize over rpw; kernel −6.9% at the best shape is only ~+0.5% wall (attention = 7.4% of the step), under the +1.5% bar and the ±2% A/B noise; numerics fully green (probe ≤1.3e-7 vs CPU, argmax hard-gated, greedy 0/10 diverged) — the session's durable output is the tolerance-gate calibration: end-to-end max|Δlogits| is 0.38/0.39 (14B/7B) for ANY accumulation-order change, so the D3-1 ≤1e-3 logits gate is unsatisfiable; argmax+greedy+A/B are the operative gate set
D3-4 L1hybrid rpw dispatch: dual-kernel self-gating split attention (f16 KV, hd==128) — 4-warp h4w kernel when rpw = ceil(ceil(nkv/32)/4) ≥ 16 (nkv ≥ 1921), incumbent 32-thread kernel below; BOTH launch per layer with static grids, each re-reads positions[0] per replay, exactly one is live per nkv (nkv-uniform branch → replay-safe)22336b214B @3254 split 72.1 → 62.1 µs (−13.9%) + 1.5 µs dud launch; wall 14B @3254 21.20 → 21.33 (+0.61%, SEP), tg128 22.96 → 22.94, 7B tg128 50.28 → 50.20, @1641 48.78 → 48.68 (guards hold)+0.61% (14B @3254)0.944× / 0.877× (14B tg128/@3254 vs llama 24.31/24.32)LANDED1-warp path bitwise (7B @1845 dump: all gated files identical; node{3,5,8}_prefill diffs = pre-calibrated slot aliasing); h4w tolerance class (max|Δlogits| 0.309, argmax identical margin 0.716, 1 greedy flip at the regime entry = 1/256 < 2%, temp-0.8 controls identical); suite 169/0/3 incl. the hd=128/n_ctx-4200 parity shape sweeping the rpw 15/16 boundary; an in-kernel 1-warp-fallback form was REJECTED pre-commit: inside 128-thread blocks the 1-warp body caps at 12 working warps/SM (1536/128) = +78% kernel at 7B @1641 (35.4 vs 19.8 µs nsys) — geometry, not math
D3-4 L2window-level K/V prefetch pipelining in the h4w body (K-pass software pipeline +2 uint4, 4-deep V-bulk ring +8 uint4; issue-point-only → bitwise vs h4w by construction)reverted (patch /tmp/d3/patch_l2.py)14B @3254 h4w 62.1 → 66.5 µs (+7%)−7% kernel—REVERTEDthe kernel is bytes+tail-bound at 79% of the 48.9 µs floor, not chain-bound enough: funding the pipeline buffers needs __launch_bounds__ minBlocks 8→4 (64→128 regs) → occupancy 32→16 warps/SM and 3.33→4.44 waves — the occupancy/wave-tail tax outweighs the shorter load chains; the D2 4-row-scale lesson (issue-point moves are free) does NOT transplant to window scale under a 64-reg budget
D3-4 findingspre-existing behaviors calibrated this session: (a) at long prompts (≥2.8K tokens) MINFER_GRAPH_DUMP PREFILL-phase files (all kv*_prefill, logits_prefill, prefill nodes) are non-deterministic pre-vs-pre (wholesale, garbage-magnitude — aliased dump reads); decode-phase dumps stay deterministic; (b) CLI prompts longer than the n_ctx default 4096 leave zero generation headroom (position N exceeds n_ctx N panic, graph.rs:367); bench unaffectedmeasurement-only———RECORDEDdump gates at long prompts must anchor pre-vs-pre at the EXACT shape and gate only decode-phase files; long-prompt CLI greedy needs prompt + n ≤ 4096 until n_ctx sizing is fixed
D3-5 1afused-producer decode A-quantize: rms_norm_quant_pad40 (rms body + the standalone per-block quantize body, both verbatim) and swiglu_quant_pad40 write the pad40 q8 plane beside their f32 output; decode matmuls consult decode_quantize_native (MmqCache, native form) and skip the standalone quantize launch on a hit; FusedFFN joins the cache-clear preserve set; MINFER_NO_DECODE_A_FUSE=1 opt-out3230b2bnsys 14B @3254 (NO_CUDA_GRAPH): standalone quantize_q8_0_pad40 4448 → 964 launches (−78%; the ~50/step remainder = the attn_o class), total kernels 29402 → 25918, sub-2µs 15171 → 11723; wall 14B tg128 23.05 → 23.35 (+1.30% SEP), @3254 21.36 → 21.68 (+1.50% SEP), 7B tg128 49.86 → 50.69 (+1.66% SEP), @1641 48.47 → 49.16 (+1.43% SEP)+1.30% (14B tg128)0.961× / 0.891× (14B tg128/@3254 vs llama 24.31/24.32); 7B 1.026× / 0.995×LANDEDq8 bytes bit-identical by construction (max is exact for any association; rintf/clamp elementwise — the epilogue IS the standalone body) and probe-verified bitwise (rms+swiglu q8 buffers, f32 producer outputs, MMVQ outputs through the cache-hit path); suite 170/0/3; 14B −n 4 dump gate: logits both phases + all KV byte-identical, the 3 node-dump diffs are same-binary pool-slot aliasing reproduced pre-vs-pre AND post-vs-post; greedy −n 256 token streams byte-identical both models; fused epilogues add ~0.2–0.25 µs/launch (swiglu 2.06 → 2.31 µs), priced into the wall
D3-5 1boutput-head od-split / 512-thread re-map (lm_head q6_K od 152064, id 5120, npair 160)not implemented———ANALYSIS-NEGATIVEall three candidate mechanisms are measured or computed dead at this shape: (a) idle-thread removal (96 of 256 idle at npair 160) = D3b-1c, measured neutral; (b) rows-in-flight: 288 resident rows either way (6×256-thread blocks/SM vs 3×512 dual-row), and D3b-1c's 9×160 = 432-row form was ALSO neutral — occupancy is not the limiter; (c) block-scheduling rate: the head sustains 47.6 blocks/µs while ffn_gu demonstrates 76/µs — not the limiter. A dual-row 512-thread form is bitwise-capable (per-row half-block reduce with the same 8-warp tree) but carries no mechanism → not built per the "measured mechanism, don't guess" rule
D3-5 1cffn_down-q6K 512-thread single-unit variant (id 13824 → npair 432)not implemented———ANALYSIS-NEGATIVENOT bitwise vs the landed v2_pf: in the 256-thread form thread t accumulates fma(u_t) then += fma(u_{t+256}) into ONE float acc before the block reduce; at 512 threads those units live in different threads and their sum happens in the reduce tree (16-warp cross-warp serial order) — a different float sum. A bitwise emulation (smem pair-exchange so thread t still sums u_t+u_{t+256} first) adds a barrier for zero resident-parallelism gain (5120 rows = 18 waves either way), and the exposure mechanism 1c targets was already fixed by v2_pf's up-front load issue (D3b-1b)
D3-6 2aGQA q-head batching in the decode split attention (grid (ATTN_SPLITS, n_kv_heads), 32*gqa threads, warp w = q head hk*gqa+w, full split stripe per warp, shared h4w_warp_windows window pass, per-warp 8/16-butterfly epilogue; same rpw ≥ 16 dispatch slot, static grids → replay-safe)reverted (patch /tmp/d3/patch_2a.py)nsys 14B @3254: h4w 63.76 → batched 64.79 µs (+1.6%); ncu lts__t_sectors: 2,121,671 (4.63× analytic 1×) → 472,583 (1.03×) — traffic ÷4.5 with time flat−1.6% kernel—REVERTEDmechanism-nailed: the 5× L2 re-read is fully HIDDEN under the latency roofline in the live regime (D1's latency-bound attribution stands; D3-4's L2-composition re-attribution revised) — ncu's −28% appears only serialized/cold; ALL gates were green first: parity ≤1e-4 at gqa 5/7 incl. exact nkv 2808 + outlier calibration (shapes kept as permanent h4w coverage), dump argmax HARD gate (margins 2.187/0.557), greedy byte-identical with repeat-penalty 1.0 both models, default-penalty flips = 1/256 sampler knife-edges (14B step-63 raw top-2 probgap 0.0167; 7B step-8 penalized rank-6 winner), temp-0.8 controls identical, suite 170/0/3; new gate rule: attribute greedy flips to sampler vs kernel via the penalty-free stream + per-step logits_top trace
D3-7 2battn_v-q6K decode MMVQ routing: Q6_K decode dispatch gate lowered od*id >= 24M → >= 4M (the only affected shape in the supported set is the 14B attn_v, od 1024 × id 5120 = 5.24M, 11 layers; GGUF census; 7B attn_v od 512 × id 3584 = 1.8M stays padded-f32)this commitnsys 14B @3254: attn_v kernel 33.16 → 24.32 µs (−26.7%, ~177 GB/s — short of the 220 class, as the 8e small-shape crossover data warned for od≈1024 but still −26.7%); ×11 layers ≈ 97 µs/step; quantize launch count UNCHANGED (attn_v joins attn_o's MmqCache hit — same src buffer + id)+0.42% (14B @3254, 8-pair median)see D3-7LANDEDTolerance-gated per the D3a package: logits_decode max|Δ| 0.254 (calibrated 0.39-class), logits_prefill byte-identical, argmax HARD gate green (margin 1.915), kv1+ decode-side f16-boundary drift (kv0 bitwise — earlier onset than D3-6's sub-ULP class, expected for input-quantization noise); penalty-free (rp=1.0) greedy streams byte-identical both models (the D3-6 clean kernel gate); default-penalty flips = 1 knife-edge event/256 steps (5/5 seeds, coherent text, no degeneracy); temp-0.8 controls: 7B identical, 14B reorders (sampled reordering expected at 0.22-logit drift); wall: @3254 21.605 → 21.72 (+0.42%, 7/7 clean pairs positive, sign-test p≈0.008; strict SEP missed by 0.05% — min-new-excl-outlier 21.66 vs max-base 21.67 — medians carried per the D3-5 outlier precedent); tg128 clean-window +0.26% (sub-bar; the extension window was co-tenant-contaminated post-side)
D3-7 2crms/elementwise-launch consolidation, two bitwise sub-levers: (i) rms_norm_quant_pad40 wide-block geometry (launch 32 → 128 threads; the reduction keeps lanes 0..31 exactly — same element→lane map, serial per-lane chains, warp_reduce_sum tree; scale broadcasts via smem; write/quantize loops are element/per-32-block independent so their wider mapping cannot move a bit; reduce loop #pragma unroll 8 deepens load pipelining), (ii) positions_i32 one-execution-window memo (every Rope/KvcacheStore/Attn node re-converted the same positions buffer: 240 launches/step at 14B; key (buf id, pool_gen), cleared in synchronize next to the MmqCache clear; capture-safe: only the first consumer's conversion is recorded and replay re-executes it)this commitnsys 14B @3254 decode census: rms_norm_quant_pad40 9.43 → 5.66 µs (−40%), 94.6–96/step; f32_bits_to_i32 239.6 → 1.2 launches/step; wall-effective ≈ −0.62 ms/step (rms −0.348 + bits −0.275)+1.76% (14B @3254 cumulative with 2b, SEP)see D3-7LANDEDBitwise end-to-end: dump gate (both sides under MINFER_NO_KQ_MMVQ=1) 109/114 files byte-identical, the 5 diffs = the documented slot-aliasing class; logits both phases + all KV byte-identical; 7B greedy streams byte-identical 5/5 seeds + temp-0.8 control; suite green. GATE GOTCHA recorded: MINFER_NO_KQ_MMVQ=1 also reverts the Q5_K decode arm (pre-existing), so a 2b-off control must set it on BOTH sides — a one-sided control shows a fake 0.22-logit drift from the Q5_K f32-activation fallback
D3-8FusedQKV decode fusion ported to CUDA (Stage-3 Tier A; the G4 Metal fusion): (1) attn_bias_rope_store_f32 kernel + launch_attn_bias_rope_store — one launch replaces the per-layer add_bias×3 + rope×2 + store_kv×2 chain, pointer-form (serves concat sections q=base/k=base+nqt/v=base+2nkt AND three separate buffers), positions read device-side (positions[0]) so the launch is capture-safe, math verbatim add_bias_f32+rope_f32+store_kv_f32/f16; (2) class 1 (wq|wk|wv same quant type): Op::FusedQKV — one concat matmul (blk.{i}.attn_qkv loader-registered wq|wk|wv rows) + the fused epilogue (MMVQ is per-row, dispatch on (ttype,id,nt) only → concat bitwise-equal to 3 separate matmuls, probe-proven); (3) class 2 (mixed quant, e.g. Q6_K attn_v among Q4_K q/k — 24/48 layers at 14B, 14/28 at 7B): new Op::QkvBiasRopeStore — the three SEPARATE matmuls (bias-free) + one epilogue launch (CUDA-only; Metal keeps the unfused chain for these layers), builder wires attention to the epilogue node so q's matmul buffer has exactly one consumer and the §5 in-place alias applies; gated by nt==1 && gpu && fuse_qkv (part of the reuse identity — MINFER_NO_FUSE_QKV=1 reverts both classes for A/B)this commitnsys 14B @3254: total launches −4968/trace (−22.5%); per decode step −310 (add_bias −144, rope −96, store_kv −96, fused +48, mmvq −48 (24 concat layers 3→1), quantize +24) ≈ the D3-7 §2-listed 0.45 ms/step qkv-chain item+1.63% (14B @3254; tg128 +3.11%; 7B +1.23%/+1.05%; isolation A/B post-vs-post NO_FUSE_QKV: +2.28%/+3.15%/+1.00%/+1.03% — all 3/3 pairs clean-separated)see D3-8LANDEDBitwise: probe tests (epilogue vs the 7-launch chain bitwise on q/k/v sections + KV rows, f32+f16 KV, both pointer forms, 14B+7B shapes; concat matmul vs 3 separate bitwise) + dump gate logits_prefill/decode + ALL kv*.f32 byte-identical both models (98/98 + 58/58; the informational node* dumps are a documented instrument limitation — recycled pool slots, binary-layout-dependent) + greedy −n 256 byte-identical 5/5 seeds × both models + rp=1.0 + MINFER_NO_FUSE_QKV=1 control + temp-0.8 controls (the only diffs are the perf-banner tok/s numbers); prefill DOT graph byte-identical (prefill untouched); suite 172/0/3 (D3-7's 170 + 2 probes); prefill graph topology unchanged (nt>1 gate) so prefill perf untouched (pp3254 1833 t/s pre-vs-post)
D4-2Decode-GEMM Tier B session (design /tmp/d4/D4_DESIGN.md): (B0) correctness — v2_pf dispatch bounded to npair ≤ 512 (see the D4-2 chapter; D3b-1b's 7B gains were mostly dropped units); (A) llama L2-prefetch port closed PRE-BUILD: prefetch distance 2·bpi requires bpr > 32 blocks (QI4_K=16/VDR=2, QI6_K=8/VDR=1 → bpi = 4·nwarps) — at 14B only ffn_down (bpr 54) qualifies, and our kernels map one 64-elem unit per thread over 256 threads → exactly ONE K-loop iteration at every decode shape (npair 80/216/160; v2_pf's 432 are unrolled u0/u1): there is no "2 iterations ahead" to prefetch, and where llama's prefetch does fire its ffn_down-q4K runs 224.1 GB/s vs our 228.6; (B1) __launch_bounds__(256,6) on v2_pf: 48→40 regs + 40 B stack spill, probe tg128 +0.25% / @3254 −0.75% → killed; (B1c) v2-loop at npair 432 via MINFER_Q6K_PF=0: v2_pf wins/ties (the 5-block × 2-unit-MLP form beats 6-block × 1-unit) → default kept, env kept as opt-out; (B2) 160-thread v2 right-size (= D3b-1c repeat, re-measured with per-kernel isolation): bitwise 98/98 but nsys lm_head +1.51%, attn_v +3.04% → killedb31084c (B0) + docs commitper-kernel nsys deltas above; walls ≈ 0 as expected for a fix-only treecorrectness fix; perf-neutral—see D4-2Dump gates 107/107 (14B pre-vs-post) + 98/98 (B2 bitwise check); greedy byte-identical 14B pre-vs-post; 7B fixed-vs-v1 first-step logits at the v1-vs-v2 rounding class (max|Δ| 0.254, argmax same) vs 4.72-4.79 pre-bug; suite 173/0/3
D4-3Attention-structure attempt 2 (probe /tmp/d4/probe_attn2.cu, 690-row sweep, 0 skips): llama-fattn-geometry split-attention kernel vec_attn over pb/T/R/minb/STG (load-scheduling axis); NO-GO per the pre-registered bar — best 41.07 µs kernel-total @14B (bar ≤~32; 1.64× vs current 67.4) and 16.22 @7B@1641 (bar ≤~10.35; 1.53×) → projected wall +1.6–1.8% < the +2% bar → no integration. Headline: the D4-1 llama attention target (14.02 µs/layer @14B/@3254) is a llama-bench artifact — ncu: the bench decode fattn-vec (grid (1,2,40)) loads a constant 5,427,200 B ≈ one 128-row KV iteration per block = 256 of 3255 rows covered, byte-identical at KV 1024/2474/3255, while llama-cli's decode (grid (1,7,40)) loads 53.2/142.7 MB scaling with context (mid-prompt recall A/B confirms). Honest llama full-context decode attention ≈ 2.0–2.1 TB/s ≈ 1.7 ms/step @14B — minfer's 3.32 ms is ~1.9× off, not 4.9×; the honest @3254 gap is ~10% wall (~3.5% attention)docs commitsweep table in the D4-3 chapterline closed (measurement-corrected)—see D4-3probe gate 0.05 abs w/ adversarial outliers, CPU ref in double; SASS-level LDG counts + recall A/B + reductio (9.6 TB/s impossible) all consistent
D4-4Final decode-kernel session (three levers, /tmp/d4/probe_l1_dpl.cu + /tmp/d4/probe_l3_fuse.cu): (L1) dense split-plane (dpl) q6_K decode MMVQ LANDED — the padded 256-elem/224B row layout streams 14 dead bytes per super-block (215/256 useful = 84%); dpl repacks to [ql: nbe×128][qh: nbe×64][sc: nbe×16][d: nbe×2] = 210B content/row at row stride (nbe·210+15)&~15 (16B-aligned uint4, zero pad sectors), same per-unit values + accumulation order → bitwise; sibling plane (+2.0 GB 14B, +0.9 GB 7B) under MINFER_Q6K_DPL ("0" opt-out), only id % 256 == 0 shapes, padded plane retained (prefill MMQ block_stride 224 + W_exp/W_dsc derivation + dequant/embed fallback). Probe: ffn_down 176.5→212.9 GB/s content (−17.1%), lm_head 208.3→250.0 (−16.7%); the group-split (gs) probe variant (−7.9/−10.7%, tolerance) dominated → dropped. (L2) PDL on the decode chain NO-GO in-situ: standalone probe green (graph capture+instantiate with cudaLaunchAttributeProgrammaticStreamSerialization OK on driver 580.173.02, 200 replays stable, PDL-graph vs plain-eager bitwise; but a compute-bound chain probe ran +2.8% slower — co-residency tax warning), full integration (PSS launch attribute + cudaGridDependencySynchronize() on 13 decode-chain kernels, MINFER_PDL gate) passed all bitwise gates, then the same-binary env-flip isolation A/B read 14B tg128 −2.6%/−1.8%, 7B ≈ 0%, @3254 within noise → below the +0.3% bar → reverted; mechanism: PSS early-launch co-residency taxes the compute-tail kernels (attention h4w, lm_head) more than the ~2 µs/launch graph-gap pool it recovers. (L3) fused gate+up+SwiGLU+q8 (Form B) NO-GO: 32-row-block fused q4_K kernel (grid nf/32 = 432 blocks, 64 serial per-row dots each, in-kernel silu + quantize_pad40_block) measured +28.2% vs the incumbent gu-matmul + swiglu_quant pair (412.1 vs 321.4 µs at 27648×5120; 193.2 vs 247.8 GB/s content) — grid = 1.5 waves at 6 blocks/SM (wave quantization) + exposed per-row latency; Form A (fused gu+swiglu f32, separate quantize) saves ~2–5 µs/step by arithmetic = sub-bar, not probedcode commit + docs commitwall deltas this row (3× interleaved same-window A/B, medians of 3)14B tg128 23.28→24.57 (+5.53%) / @3254 22.00→22.94 (+4.27%); 7B tg128 47.55→51.20 (+7.68%) / @1641 46.44→50.18 (+8.05%) — +5.53%/+4.27% (14B)vs-llama: 14B tg128 1.018×, @3254 0.950×; 7B tg128 1.074×, @1641 1.052× (llama 24.14/24.14/47.65/47.69)LANDED (L1); L2/L3 closed with mechanismL1: dpl kernels bitwise vs padded (probe + unit test both forms: od 512/id 8960 pf-form, od 4096/id 1024 loop-form); 14B −n 1 dumps 107 identical + 7 node{N} diffs (= the documented D4-2 pool-slot aliasing class), 7B 72 + 2; greedy rp=1.0 byte-identical both models; suite 174/0/3 (incl. the new dpl bitwise test; one earlier full-suite run flaked 2 pool/parity tests under the sglang co-tenant window — both pass in isolation and the rerun is green)
D5-0Speculative-decoding cost model (gate, no engine change): measured per-token costs (7B q4_k_m CUDA 54.3 tok/s; 0.5B q4_0 CUDA 342.2 / CPU 73.3; q4_k_m 365.0; q5_k_m 321.8) + real greedy acceptance via llama.cpp speculative-simple (0.5B-on-7B: aggregate 58.8/42.6/25.0% at d=2/4/8 → conditional p ≈ 0.68–0.70, stable across prose/code and draft quant) → break-even requires p* = 0.73/0.81/0.90 at d=2/4/8 → conditional go at d=2 only, gate = minfer measured nt=3 verify-batch amortization ≥ 2.5× (D4-1 anchor 2.7× at nt=4; interpolation 2.28× vs tile-step 2.7× disagree — D5-1 re-ordered primitive-first to measure); projected 1.04–1.05× at the anchor, ceiling ~1.2×; CPU-draft cross-device dead (1.35×)docs commitbattery: 3× interleaved -p 0 -n 128 medians; acceptance n=256 greedy, prose+code——MEAS-ONLY (gate open)measure the gate variable with someone else's binary; the cross-device fallback died by measurement not argument; interpolation is not measurement — nt=3 lands either side of the 2.5× line and only C_T(3) arbitrates
D5-1aThe gate measured end-to-end — FAILED, D5 closed: new minfer specverify instrument (src/spec_verify.rs; fixed-depth protocol, warmups absorb the graph rebuild, two-pass drift check) drives the generic forward_graph_cached at nt>1/n_out=nt; 7B q4_k_m CUDA @KV512: C_T(1)=18.33 ms, C_T(3)=106.0 ms → per-token amortization 0.52× vs ≥2.5× (needed C_T(3) ≤ 22.1 ms). Full curve: nt=2–8 costs a flat ~35 ms/token (batched path re-streams weights per row — 1.9× worse per token than the nt=1 MMVQ path; zero amortization anywhere), the real tile-regime step sits at M≥16 (nt=16 = 56.7 ms total, 3.1× the weight floor; nt=64 = 64.9 ms) — unreachable for verify (nt=d+1 ≤ 9), and even padding nt=3→16 caps at 0.97×. Five probes eliminated alternatives (graph-launch asymmetry, KV depth, FA kernel, n_out path, drift). The D4-1 2.7×@nt=4 anchor was a kernel micro-bench that never existed at graph levelthis commit + docs addendummedians of 15–40 reps × 2 passes, ±2% spread; probes + llama-cli battery in doc 81gate FAIL 4.8× (0.52× vs 2.5×); external check: llama-cli -md same pair lands 0.99–1.00× (doc 81 §4.1) — VOID, see doc 81 §4.3 errata: the battery never engaged the draft (missing --spec-type); corrected = 1.53–2.43×—see D5-1aLANDED (instrument) · D5 CLOSED per the pre-registered stop rule
82small-M dispatch fix: multi-token MMVQ (K-quants, nt 2–8, in-block token loop) + token-looped legacy kernels (grid.y=nt removed) + GEMM gate 16→9this commit7B batched decode nt=3 105.9 → 29.4 ms (3.60×), nt=8 279.3 → 48.6 ms (5.75×); marginal 34.4 → 4.3 ms/token; nt=1 paths bitwise-unchanged (tg128/pp512 clean)——LANDEDthe pre-registered 2.5× amortization bar was mis-derived (marginal ≈ nt×(attention+compute), not ε) — measured 1.87×; weight traffic is nt-independent now, D5 stays closed (C_T(3)=29.4 > 22.1); small models gain ~1.0× only (per-layer weights already L2-buffered)
D5-R ①speculative loop LANDED (reopen per doc 81 §4.3 + doc 82): src/spec.rs SpecEngine — d×nt==1 draft forwards + one nt=d+1 verify through forward_graph_cached, lazy accept loop (unit-tested), full-accept draft-KV repair, CLI --spec-draft/--spec-draft-n; two-model process fixes: namespaced GPU weight registries (load_model_ns), nb_bt_only global-mix semantics (q4_0 draft degrades mode-2→mode-1)this commit + doc 8314B+0.5B q4_0 d=2 greedy n=128: 34.0/40.6 tok/s = 1.34×/1.58× vs serial 25.5/25.5; 7B 1.18×; suite 179 greenG1 re-scoped by measurement: batched-verify vs decode kernels differ ~0.01–0.05 logits → near-tie flaps only (first-flap margin 0.043); d=0 fallback == non-spec to one exact-tie flap—LANDED"single-model-per-process" was load-bearing in three places (registry, dispatch flags, prewarm); all-or-nothing CUDA failure is silent CPU — watch tok/s, not errors; greedy equivalence across kernel paths is a numerics statement, not a logic statement
D5-R ②same-window dual-engine battery: 3 interleaved reps × prose/code × 4 cells (minfer off/d2 × llama base/d2), doc 81 §4.3 protocolthis commit + doc 84minfer 1.33×/1.59× (34.0/40.6) vs llama 1.64×/2.08× (38.8/49.0) in one window; minfer = 88%/83% of llama's absolute spec speed; gate ≥1.2× PASS——LANDEDsame-window interleaving beats rep count (llama's spec cell moved 13% between windows, base <1%); the whole gap to llama is the verify row marginal (minfer 8.8 vs llama 2.5 ms/row) — closing it prices at 1.62×; acceptance prose 51.6% (near-tie dilution) / code 73.1%

Footnotes.

  1. r59 correction (visible in-table). The r59 record originally reported +26.2% interleaved under a co-tenant and attributed −12% to a "co-tenant tax". r59b proved the r59 baseline binary was itself the stale r58-delta build (~−12.5% deficient), making the true clean delta +11.1% (3232.0 → 3590.8). The +26.2% and the co-tenant-tax claim are superseded; the corrected value is what the table's Δ column carries.
  2. Measurement contexts. Rows r12–r25 interleaved on a box that drifted session to session (−9% to +38% vs neighbors) — only intra-session deltas are meaningful. r59's interleaved series ran with a 46 GB sglang co-tenant (later shown irrelevant: idle residency taxes nothing, r59b §1). r55's baseline sanity was 3144.4–3151.4 under a live co-tenant vs the quiet-box 3181.
  3. Hash substitutions. The older records quote HEAD/revert anchors that were later amended away (e.g. r22's "HEAD 9819410", r18's "HEAD 1e0dded", r55's "HEAD dd6d842", r58's "HEAD 2105b08", r59's "commit feb37de"). Each has a reachable twin with an identical subject; the table cites the twins. The P5 record's session range start b8568cd does not resolve to any commit — the P5 code commits are 86ca78c, d713e6e, 725e307, fc07c04, 1365c82. llama.cpp-side hashes (ca3d5a3e1 bench build) are upstream identifiers, not minfer commits.
  4. Campaign arc: R1 MMQ 441 tok/s (first parity-clean MMQ measurement, r7–r8 window) → 3590.8 tok/s (r59b definitive) = 8.1×.

§1 Current state (post-D4-4, 2026-09-09)

1.1 Performance summary (DGX Spark GB10, 7B q4_k_m unless noted)

llama.cpp reference: llama-bench @ ca3d5a3e1 (upstream build), 8 threads, -ngl 99.

Config (all default env unless noted)7B q4_k_m whole-prefillMemory (peak, per-PID)Note
Default (promoted MMQ gate set, mode 2)~3581 (r59b definitive median; r60 A/B window 3578.0–3598.7)9484 MiB= the verified 1.080× path
MINFER_MMQ_Q6K_EXP=0 MINFER_MMQ_Q4K_DSC=0 (planes off)−5%-class (r54: −5.04%; r59-class)6217 MiB (planes cost +3.27 GB)fine-grained opt-out
MINFER_MMQ=0 (legacy f16 w16-cache escape)~2353 clean-class (~2226 in the r60 window)~20.5 GBthe escape is ~11 GB HEAVIER, not lighter
llama-bench pp3314 (r59b window)3323.29 ± 3.08—minfer 3590.8 / 3323.29 = 1.080×

Decode (nt==1) is untouched by the MMQ campaign (r60 evidence: no MINFER_MMQ* read on the nt==1 path; decode -n 16 --greedy byte-identical). Final decode-campaign state (D1→D4-4, 2026-09, all measured on the B0-fixed engine — the D4-2 correction found the v2_pf dispatch dropping units on npair-592 rows and re-anchored every 7B claim):

modeltg128long-ctxvs-llama (same-window)
7B q4_k_m51.20@1641 50.181.074× / 1.052× (ahead)
14B (48L)24.57@3254 22.941.018× tg128 / 0.950× @3254

Small models (pre-MMQ-campaign numbers, Part-I record): 0.6B q8_0 prefill @2K 4792 (llama 23909), decode tg128 ~195 (290); 0.5B q4_0 prefill ~3020 (30550), decode ~257 (453).

The per-session narratives (D2/D3-4/D3-5/D3-7/D3-8 updates), the D4-2/D4-3 correction chain, and the standing measurement rules live in the §0 rows D1–D4-4 and step docs 65–76. Two rules worth surfacing: never quote llama-bench long-ctx tg rates as attention targets without an ncu byte-count or llama-cli recall cross-check (D4-3: the bench decode fattn-vec covers only ~8% of KV rows; honest llama decode attention ≈ 1.7 ms/step, minfer ~1.9× off), and an end-to-end max|Δlogits| ≈ 0.38/0.39 is the inherent class of ANY accumulation-order change — argmax + greedy-divergence + A/B are the operative gates (D3a calibration).

1.2 Wall decomposition (converged regime, r55/r58/r59-era records)

Production nsys at nt=3314, kernel busy ≈ 984 ms (r58 census, pre-r59):

SliceSharevs-llamaState
q4_K BT GEMM (mmq_raw_nb_bt)63.2% (622.4 ms; r59: −30.9% → 526.8 ms)1.06×closed absent a llama.cpp-style q8_1 GEMM-prologue rewrite
q6_K BT GEMM (mmq_raw_nb_bt_q6k)15.4% (125.5–148.8 ms)~1.1× kernel-sideclosed (r53/r56: both staging planes precomputed, all stagings async)
fused A-producers (swiglu+quant, rms+quant)8.8% (swiglu 64.0 ms)swiglu at 89% of the 273 GB/s DRAM rooflineswiglu closed (cap +0.74%); rms at 56% roofline = last incremental lead (ideal +0.98%)
FA prefill attention5.3% (51.8 ms)2.43× taken in r48 (5.16 → 2.12 ms)closed for tile levers (r46/r50/r57)
elementwise/rope/kv/store~6%—bandwidth-bound
standalone wo quantize + mmvq tail~2%—tail section runs the MMVQ/native path
host/launch gaps~1%—one-time stalls 0.6% + recurring gaps ~0.1% + tail malloc 0.1%

Campaign verdict (r55, re-validated by r59/r60): no identified lever ≥ +1.5% remains within the current architecture; the next meaningful step is the step-function q8_1 GEMM-prologue fusion, not incremental optimization.

1.3 What the engine does now (dispatch shape)

All in src/cuda/kernels/*.cu + src/cuda.rs, dispatched by src/graph/cuda_backend.rs:

  • Weights resident at load: every matmul weight uploaded once and registered by name; q6_K optionally 224-byte-padded (7e②). Registration additionally builds the q6_K W_exp (1.52 GB) and q4_K/q6_K W_dsc (363.2 MB + 1456 MB) planes when their gates are on (Appendix A).
  • Prefill (nt ≥ 16): the promoted MMQ path — quantize_q8_0_pad40_t pre-transposed A planes (r34), fused producers (r51/r52), NB-BT raw-byte int8-tensor-core GEMMs for q4_K (mmq_raw_nb_bt_kernel, 2 blocks/SM) and q6_K (mmq_raw_nb_bt_q6k_kernel, KSPLIT=2, 3 blocks/SM, cp.async A/B/dsc staging), FA-style tiled prefill attention (8n, FAP2 register softmax). Fallbacks compile in and are byte-identical (EXP=false, DSC=false, mode-1 producers, generic mmq_nt/f16 arms).
  • Decode (nt == 1): per-type MMVQ over q8_0 activations (8e/8e② + R2 v2), fused bias+rope+store, whole-step CUDA-graph capture/replay (7d); pinned D2H logits readback (R3-A2). Repeated identical-nt prefills capture after the 3-run protocol (R3-B); one-shot prefills never capture (r55: measured no-win + capture-illegal mid-window malloc).
  • Memory etiquette (shared box): no raw allocation probes; check free -g before suite runs (the suite transiently reserves up to ~100 GB of the overcommitted pool); benches stay at single-process 7B scale while sglang serves.

1.4 Remaining roadmap (post-campaign)

  • D5 speculative decoding — CLOSED 2026-09-10, REOPENED as D5-R 2026-09-12 (docs 81 §4.3, 82, 83, 84): the original closure's bar was mis-derived and its llama-cli anchor was a measurement artifact (--spec-type silently defaults to none); doc 82's small-M dispatch fix made the verify amortization 2.14×. D5-R stage ① landed the loop (1.34×/1.58× at 14B d=2), stage ② priced the gap to llama (verify row marginal 8.8 vs 2.5 ms/row → 1.62× recoverable). Current plan: SPECULATIVE-DECODING-PLAN.md (the old plan is an appendix there).
  • Not planned (revisit with a concrete need): cuBLAS/cublasLt (closed as 8k — 8m's wmma GEMM covered the f16 path), VMM pool, multi-GPU, node reordering, Windows, IQ/Q2/Q3 quants.
  • Open leads, all sub-bar or step-function: q8_1 GEMM-prologue fusion (the step change); rms_nw roofline (+0.5–1%); wave re-tile for small-od classes (+0.3–0.8%, needs ≤85 regs); fused ffn_gu concat (needs the G5 nf≤16384 gate re-measured); FA deep-opt only with numerics-order-preserving structure (r50/r57 caveat).
  • Open Phase-8 ledger items (inherited from the retired CUDA-FOLLOWUP-PLAN.md; records in step docs 78/79): 8e② follow-up — llama.cpp's shape-dependent halve_iters idle-tail rule, not started; 8h② — self-hosted CUDA CI runner, DEFERRED (needs standing runner infrastructure); 8h③ — the Phase-7 /tmp/minfer_phase7/ ledger cleanup, awaiting user decision.
  • Closed Phase-8 ledger item: 8a① macOS Metal regression run (fuse_ffn decoupling + the MINFER_NO_FUSE_FFN A/B on 0.5B + 7B) — was BLOCKED on hardware; DONE 2026-09-10 on an Apple M4 Pro, all three checks green (pre/post greedy byte-identity, fused-vs-unfused byte-identity, 0.5B decode graph still emitting 24 × fused_ffn on Metal). Record: doc 78 §3.2.

§2 Step documents — one doc per history row

The full per-step chapters (process narrative, principle explanations, real code excerpts, verification gates, lessons — previously inlined here) now live in docs/cuda_optimization_steps/ as 109 standalone documents (01–109, including the Phase-8 supplementary records 78–79 and the verification-methodology capstone 77). The §0 master table above remains the one-row-per-step index; the tables below link each row to its step document. Appendix B points at the cross-cutting methodology.

Part I · Era A — Phase 7/8 foundations (rows 1–6 + 78–79)

Part II · Era B — R and P5 sessions (rows 7–13)

Part III · Era C — the P6 q4_K MMQ line, r5–r37 (rows 14–51)

#doc
12r5–r6 re-ranking + structural rewrite spec (MEAS-ONLY + REVERTED)
13r7–r8 raw-byte MMQ kernel + wide tile + FA probe (LANDED (raw) + REVERTED (probe))
14r9 llama.cpp MMQ reference decode; shape axis closed (MEAS-ONLY)
15r10–r11 — reference inner-loop decomposition port; ILP reading verification (REVERTED / MEAS-ONLY)
16r12 — 16-chain warp tile + ldmatrix (LANDED)
17x-tile / j-tile / cp.async-db — the staging-shape family, closed (REVERTED / closed)
18r13 — counter forensics against llama.cpp (MEAS-ONLY, closed)
19r14 — B fragments via ldmatrix + widened scale reads (LANDED)
20r15 — f32-accumulate mma probe (dead end) + rank-1 term2 rescale (LANDED)
21r16 — Narrow kernel gets the rank-1 fold (LANDED)
22r17 — Wide warp remap 32od × 64tok (REVERTED)
23r18: Load-time B pre-expansion — staging becomes a bulk copy (W_exp's debut, REVERTED)
24r19: Weight L2 residency — __ldg imperceptible, persisting window catastrophic (REVERTED)
25r20: Split-phase A staging — attribute first, shoot second: the first hit (LANDED)
26r21 — Coalesced block-linear A staging: stall-mass conservation (REVERTED)
27r22 — qa8 XOR swizzle (LANDED); d/ssum fold reverted separately
28r23 — f16-path whole-graph wall decomposition + FA_TKV occupancy raise (MEAS-ONLY + REVERTED)
29r24 — The scheduling-structure ladder (REVERTED; the +1.5% whole-prefill landing bar calibrated here)
30r25 — SASS opcode census; the unroll is wall-inert (MEAS-ONLY + REVERTED)
31r28 — Direction-A raw-nibble NB kernel, 2 blocks/SM (LANDED)
32r29 — NB kd-loop unroll: 2 blocks/SM lets integer-ALU pruning move the wall clock for the first time (LANDED)
33r30 — SWAR unpack: the compiler already did it (REVERTED)
34r31 — q-major sda scale-read repack: a sub-bar positive gain caught by conflict analysis (LANDED)
35r32 — The finite lever sweep: two regions fenced off (REVERTED)
36r33 — Hybrid inner-loop port: SASS fully identical, hypothesis falsified (REVERTED)
37r34 — The quantize-transpose prepass: layout-transform locality (LANDED, +9.72%)
38r35 — sda scale predecode: a total SASS win, a wall-clock tie (REVERTED)
39r36 — A-frag wavefront economics: H1 falsified (MEAS-ONLY, no code change)
40r37 — Post-parity whole-prefill attribution: the wall clock re-decomposed (MEAS-ONLY, no code change)

Part IV · Era D — q6_K, FA, prepass, promotion, r38–r60 (rows 52–75)

#doc
41r38 — q6_K BT-style raw-byte mma kernel (LANDED)
42r39 — q6_K KDR=2 double-buffer (LANDED)
43r40 — __launch_bounds__(256,3) third resident block (LANDED)
44r41 — q6_K B-expand widened to uint4 group loads (LANDED)
45r42 — q6_K stage-wide dsc scale reads (REVERTED)
46r43 — PC-sampling attribution + pre-expand-B parity FAIL (MEAS-ONLY + REVERTED)
47r44 — W_exp stride mismatch root cause: fix goes parity all-green but wall-neutral (REVERTED)
48r45 — cp.async for the q6_K A-side staging: mechanism confirmed, wall-neutral (REVERTED)
49r46 (FAP1) — FA audit + occupancy/bank-conflict levers: kernel −11% but wall-neutral (REVERTED)
50r47 — converged-era whole-wall re-decomposition (MEAS-ONLY)
51r48 (FAP2) — register-resident softmax: deleting the S/P smem round trip outright (LANDED)
52r49 — A-quantize prepass shared-A dedup: consecutive-window memoization (LANDED)
53r50 — FA_TKV 32→16 occupancy experiment (REVERTED)
54r51 — producer-fused A-quantize mode 1 (LANDED)
55r52 — skip-write mode 2: skipping the f32 intermediate write-out (LANDED)
56r53 — q6_K bundle: W_exp pre-expansion plane + cp.async B staging (LANDED)
57r54 — MINFER_MMQ_Q6K_EXP: an exit valve for the 1.52 GB W_exp plane (LANDED)
58r55 — swiglu roofline audit + one-shot prefill CUDA-Graph: both closed on the record (CLOSED)
59r56 — q6_K A-side bundle: A cp.async + W_dsc f32 plane (LANDED)
60r57 — FA KV staging double buffering (REVERTED)
61r58 — q4_K BT spec + cp.async-db2 transplant (REVERTED)
62r59 — q4_K W_dsc plane + riders (LANDED, Δ corrected by r59b)
63r59b — clean re-measurement + baseline-contamination correction (measurement round)
64r60 — the coronation: flipping the verified gate set to default-on (PROMOTION, LANDED)

Part V · The decode campaign, D1→D4-4 (rows 76–87)

#doc
65D1 decode attribution: split-attention staging depth is the only wall that grows with KV (measurement round, CLOSED)
66D2: explicit K+V register staging (LANDED, +2.0% @1641) and the cp.async negative result
67D3: 14B decode attribution (D3-1) + the D3b bitwise MMVQ triple (1b LANDED; 1a/1c REVERTED)
68D3a — the 4-warp fattn-vec-style split-attention rewrite (REVERTED) + tolerance-gate calibration
69D3-4 — L1 hybrid rpw dispatch (LANDED) + L2 window-prefetch pipelining (REVERTED) + long-prompt dump calibration
70D3-5 — decode-alignment plan Stage 1: fused-producer decode A-quantize (LANDED) + negative analysis of the output-head/ffn_down geometry levers (1b/1c)
71D3-6 — GQA q-head batched attention: all gates green, still reverted — the 5× L2 re-read was not the residual (REVERTED)
72D3-7 — attn_v-q6K MMVQ routing (2b) + rms wide-block / positions memo (2c): the Stage-2 closing ledger (LANDED ×2)
73D3-8 — G4 FusedQKV ported to CUDA: both layer classes covered, 14B short-KV breaks through parity (LANDED)
74D4-2 — B0 latent correctness fix + all bitwise occupancy/prefetch axes closed (LANDED)
75D4-3 — attention structure rewrite attempt 2 NO-GO + D4-1's llama target was a llama-bench artifact (CLOSED)
76D4-4 — the endgame kernel session: dpl dense split-plane q6_K decode MMVQ lands (bitwise); PDL and fused-FFN closed with mechanism (LANDED)
80D5-0 — speculative-decoding cost model: measured baseline, acceptance, and the d=2 gate (MEAS-ONLY)
81D5-1a — the verify-batch gate measured end-to-end: no amortization at any nt, D5 closed (LANDED · gate FAIL)
82small-M dispatch fix — multi-token MMVQ + token-looped legacy kernels: the batching invariant restored, D5 verdict unchanged (LANDED)
83D5-R stage 1 — speculative decode loop: two-model process fixes, accept-loop unit tests, greedy-identity investigation (LANDED)
84D5-R stage 2 — same-window dual-engine battery vs llama.cpp: 1.33×/1.59× vs 1.64×/2.08×, gap = verify row marginal (LANDED)
D5-R ③verify marginal priced with an nsys per-kernel ledger (specverify per-nt runs, exact forward spans)
D5-R ④aattention for the verify shapes: fa_prefill gate nt≥64 → nt≥2 (one line; hd==128, kill switch, rc fallback kept; nt=1 keeps split-KV)
D5-R ④bmulti-MMVQ nt 9–16: token-groups-of-8 (parity 85.0 vs 84.9 — group re-streams weights from DRAM) and acc[16] single-pass (111 ms, register spill) both measured; doc-82 GEMM boundary at nt ≥ 9 stands, kernels keep the group structure (bitwise at nt ≤ 8)
D5-R ⑤final dual-engine battery + capture A/B: the doc-85 capture prize was already banked by R3-B prefill capture (48.08 captured vs 50.15 eager)
D5-R ⑤+row-marginal localization (no ncu): cold-L2 real-kernel bench + chain nsys + ablation
D5-R ⑤++doc-89 menu item 1 (R-rows-per-block) implemented → measured → reverted
D5-R ⑤+++mma path audit: BT kernel already tensor-core; small-M floor = block starvation (40 blocks < SMs); conditional double-buffer shipped
D5-R ⑤++++K-split shipped for both BT kernels (grid.z + deterministic reduce), gated; flip condition C_T(9)≤55 NOT met (72.9)
D5-R ⑤+++++auto-ksplit enabled on the DEFAULT path (user decision, doc 92 §3b)
D5-R follow-updraft-quant sweep (mixed knob: q4_k_m code +3.0%/prose −2%) + greedy identity test: spec output ≠ sequential — flips originate in batched verify attention/softmax, nt-invariance campaign proposed
D5-R ⑥++++++greedy identity ACHIEVED: spec output byte-identical to sequential (4-prompt battery) — batched verify attention (bitwise position-invariant, nkv<1921) + spec penalty window capped to repeat_last_n; cost ≈1% C_T(3); d=8 door CROSSES with q4_k_m draft (code 76.5% ≥ 0.755: 46.6 tok/s vs d=2 43.6, 96% of llama; prose stays d=2)
D5-R ⑦adaptive draft depth (--spec-draft-adaptive): per-round d from beta-smoothed per-depth acceptance + min-of-4 verify/draft costs, 10% switch hysteresis, unobserved depths inherit the depth-1 rate (emergent exploration); identity boundary discovered and pinned — verify nt ≤ 8 (single+multi MMVQ) bitwise vs decode, nt=9 (BT-MMQ) lm_head is tolerance-class while its KV stays bitwise (48-layer dump proof) → adaptive cap d=7
D5-R ⑦P0nt=9 verify profiled (Phase 0, pre-registered stop-gate): nsys — BT-MMQ kernels = 84% of GPU time (attention ~1%, ksplit reduce 1.3%, quantize 1.7%); ncu (sudo; ERR_NVGPUCTRPERM workaround) — both BT kernels at ~15-24% SM / ~15-26% memory throughput, smem scoreboard stall = 40–59% of warp cycles → the <30% stop-gate says PROCEED; Phase 1 = cp.async double-buffered staging (bitwise-preserving), EV narrow (prefill lever / identity-relaxed d=8 — adaptive d≈3.5 already beats d=8 identity-safely) → re-priced separately below 战役 97
D5-R ⑦✦speculative decoding wired into ALL frontends: --cnv --spec-draft (Engine trait hooks + SpecAwareEngine + a spec sibling decode loop mirroring the plain one token-for-token) and serve --spec-draft (per-slot draft engines, seed carry across rounds, mid-batch stop/EOG termination); pre-existing server bug fixed in passing — slot GraphCache reuse across different-prompt requests leaked stale KV rows into the new attention window (the plain path was contaminated too; identical back-to-back requests hid it) → per-request slot-cache + draft reset
D5-R ⑦P196 Phase 1 lever (cp.async double-buffered staging for the q4_K BT kernel) implemented in full — A/B planes + one-group-per-tile (r56 pattern) + the dbuf regime extended to the K-split path — and measured NULL on GB10: q4_K BT 188.5/146.1 µs vs baseline 187.6/142.4, dbuf on/off within noise of each other, C_T ladder and pp512 unchanged; bitwise-preservation gate passed (pre == post == dbuf-off, 4/4 prompts). The 40–59% Short-Scoreboard stall is compute-side smem dependency (ldmatrix→mma operand chains), not tile staging — doc 92's staging attribution corrected. Remaining lever class reorders fp32 accumulation → tolerance-class → excluded on the identity-claimed path → patch reverted, campaign closed; C_T(9) ≈ 73 ms stands as the identity-safe floor
D5-R ⑦P1b99: fast-verify P0/P1 — doc 98's tolerance-class blanket corrected (int mma is exact; only the per-kd fp32 fold carries order → the bitwise-safe set is wider), then four interventions measured: cp.async staging NULL, fragment prefetch NULL, non-volatile mma NULL, B-plane XOR swizzle kept (−4.3% instance / e2e noise; the 40% excessive shared wavefronts eliminated). pc-sampling fixes the root-cause chain: the stall is L1TEX (global) latency at a register-file-capped 16 warps/SM — not smem, not staging, not scheduling. MINFER_FAST_VERIFY not built: tolerance-class levers act on scheduling and are re-priced to single-digit %; weight repack (−27% staging sectors) is the only remaining priced lever (~7–10%)
D5-R ⑦P2100: the last priced lever (32-B-aligned qs plane for the BT B staging, +0.89x q4_K memory) implemented and measured timing-NULL under interleaved A/B (141.9 vs 141.6 ms kernel totals; sequential "−12.4%" was clock-ramp drift). Discovery: dgxspark's short benchmarks carry a ±7–12% SM-clock-ramp band (208 MHz idle → 3 GHz) — sequential comparisons across docs 92–99 all carry it; only interleaved A/B is valid. Traffic levers don't convert in a latency-bound regime → repack family closed; doc 99's swizzle stays on its mechanism evidence with honest error bars
D5-R ⑦P3101: doc 100's method rule made executable — specverify/bench warm on a time budget (default 2000 ms, MINFER_SPECVERIFY_WARMUP_MS / MINFER_BENCH_WARMUP_MS); headline table re-measured in one hot session: pp512 2083.0 ± 7.2 (was ±25), e2e prose 25.3→36.6 (d2) / 35.9 (adaptive), code 25.3→44.8 (adaptive, +7.7% over best static); doc-95 structure reproduces, doc-95 absolutes confirmed as drift-band artifacts
D5-R ⑦P4102: draft-scale sweep — bigger drafts lose (Qwen3-0.6B Q8_0: 39.7 vs incumbent 44.8 code-adaptive; 1.7B: 34.2; acceptance bounded by target predictability at 76.5% same-family vs 63-67% cross-family) → default draft unchanged. Flushed a latent bug: qwen3 loader registered Q6_K padded weights under the raw GGUF name, silently replacing the target's registry entry under the namespaced draft load and dropping BOTH models to CPU — fixed to reg_name; post-EOS token-text divergence downgraded to a warning (cross-family drafts legitimate; identity re-proven 4/4 vs fresh sequential)
⑧ legacy-quant decode103: the 8e MMVQ structure extended to q4_0/q8_0 decode (nt=1 + multi, NEW CODE ONLY — every landed K-quant kernel/arm byte-identical; size-floored gates id≥2048, q4_0 multi only above the 8c territory id>8192; MINFER_NO_Q40_MMVQ/MINFER_NO_Q80_MMVQ fallbacks). 7B q4_0 tg128 48.1→56.8 (+18.1%), 7B q8_0 26.7→28.8 (+7.9%), 7B q8_0 spec e2e 26.8→68.0 tok/s (+154%) — q8_0's verify batch previously ran the f32-activation kernel and re-read f32 activations per token. Same-file llama.cpp: q4_0 decode 1.045× FASTER, q8_0 closed 0.87×→0.94×; pp512 unchanged (prefill gap 0.21–0.24× is pre-existing, out of scope)
⑧ q8_0 p32 planes104: nsys shows 97% of q8_0 decode GPU time inside the MMVQ kernels; cold-L2 microbench isolates the 34B-stride 2B-load scatter (~2x q4_0's wavefront cost/byte) → p32 split planes (uint4 x2 payload + dense d, traffic unchanged, raw registration untouched, byte-equal outputs, MINFER_NO_Q80_P32 fallback). 7B Q8_0 tg128 → 31.9-32.1 tok/s = llama.cpp parity (0.87x→0.94x→1.00x across docs 103-104); spec e2e 79.0 tok/s = 2.47x sequential; both engines now bounded by the same ~253-266 GB/s DRAM ceiling — decode gains for any quant now need a higher streaming ceiling, not better kernels
T1 device tier tables105 (T-series, plan: DEVICE-ADAPTATION-PLAN.md): cc-keyed tier table + selector (device_tier.rs, pure data) — GB10 row Measured (docs 94-104), consumer rows Adopted from llama.cpp (1200 Blackwell q4_K→5/q5_K→6/q6_K→7, 870 Orin K-quants→1, 890/860 generic-8, 750 Turing mmq=false per ruling #4), GENERIC fallback (mmq = cc≥800); MINFER_DEVICE_TIER forced-key override for soak tests; mmq_active() = resolved tier flag; batch caps tabled + unit-tested but unwired (R8: no destination for the vacated nt range until small-nt BT / field A/B). Encoding fix en passant: runtime cc is 1201 (major*100+minor), stale "1210" comments corrected; table keys use llama.cpp encoding via llama_key() conversion
T2 query formula gates106 (T-series T2): (1) doc-92 auto-ksplit extracted to auto_ksplit(), block target = max(256, 2*SM) — GB10 (48 SMs) keeps the calibrated 256 exactly, larger SM counts scale; (2) BT smem feasibility: the tile demand is a single-source C formula exposed as cuda_mmq_smem_bytes(), init folds demand > per-block-optin → tier_mmq=false (f16 GEMM serves) — R2 fixed (externs queried device 0 unconditionally, now the current device); (3) plane_budget_ok() (free > extra +25%) gates every optional plane upload (p32 pair +100%, q6_K W_exp/W_dsc, q4_K W_dsc) — Orin-Nano-class 8 GB devices self-disable planes, raw paths serve. Tile candidate search + batch-cap activation deferred per plan §6.4/R8 (Orin Nano A/B first); T3 stays later/independent
C4 #144 packed Q8_0107: the packed cache's two missing tuned routes. (1) Packed fused decode epilogue — attn_bias_rope_store_q8_0 gives one thread a whole (head, 32-element K block) and one V block, quantizing with the same q8_0_quantize_block store_kv_q8_0 uses; the Qwen2 builders dropped && !packed, and the fused metas gained row_elems so the allocator sizes a packed region the way the store node does. (2) Packed FA prefill — fa_prefill_f16kv became fa_prefill_kv<CAUSAL,MAP,LAYOUT>; only the staging is layout-dependent (kv8_q8_0 dequantizes each packed block into the same f16 tile), the tensor-core QK^T/softmax/P·V are untouched, and the general layout-tagged kernel stays the documented fallback. (3) A dp4a packed K dot was deliberately NOT taken (numerics change, own accuracy statement) and is filed separately
C4 #186 dp4a packed K dot108: the packed Q8_0 decode K dot accumulates in int (__dp4a) against a per-(head, block)-quantized query and scales by d_q*d_k once per block — K is never converted to float (V still is); MINFER_NO_DP4A_Q8_KV=1 is the same-binary control; the nt > 1 prefill/verify kernel is deliberately not taken (per token and head query, no shared block scale)
C4 #202 packed KV L1 loads109: the packed cell's four s8 K/V quant loads become two u16 loads (34k + 2 + 4m is always 2-byte aligned even when only odd blocks are 4-byte aligned) — no layout, CPU, copy-stride, session-format or map_q8_0_cells change; MINFER_NO_Q8_KV_WIDE=1 is the control

Methodology

§3 Appendices

Appendix A — Env-gate reference (post-r60 semantics)

Promoted gates (r60): absent / any non-"0" value = ON (the verified best path); explicit "0" = opt-out (the pre-r60 default behavior). All reads single-sourced in CudaState::mmq_gate_on.

GateDefault"0" effectIntroduced
MINFER_MMQonwhole MMQ path off → legacy f16 w16-cache prefill (~20.5 GB, ~2353-class)R1 (opt-in), r60 (default-on)
MINFER_MMQ_RAWonraw-byte kernels off → generic mmq_nt armsr7–r8, r60
MINFER_MMQ_RAW_NBonNB kernels off → wide raw kernelr28, r60
MINFER_MMQ_A_TRANSPOSEonpre-transposed prepass off → native quantize + NB kernelr34, r60
MINFER_MMQ_Q6K_NBonq6_K BT kernel off → generic q6_K pathr38, r60
MINFER_MMQ_A_FUSEabsent = mode 2 (skip-write)off (mode 1 = "1": fused producers write plane AND f32; "2" = skip-write override)r51/r52, r60
MINFER_MMQ_Q6K_EXPonq6_K W_exp plane not built (−1.52 GB, ~−5% prefill) → EXP=false r41 pathr54
MINFER_MMQ_Q4K_DSConq4_K/q6_K W_dsc planes not built (−1.46 GB) → in-kernel decoder59
MINFER_Q6K_DPLondense split-plane q6_K decode planes not built (−2.0 GB 14B / −0.9 GB 7B) → padded-224B MMVQ path (bitwise)D4-4

Overrides / debug (opt-in "1"): MINFER_MMQ_RAW_KD (default 8), MINFER_MMQ_RAW_WIDE (wide 128×128 kernel), MINFER_MMQ_RAW_NB_DEBUG (prints the dispatch label — B=W_exp-cp.async / DSC=f32-plane / exp=off / fallback! — the liveness instrument from r53/r54).

Legacy f16-path A/B gates (unchanged): MINFER_NO_PREFILL_GEMM, MINFER_NO_W16CACHE, MINFER_NO_FA_PREFILL, MINFER_NO_CUDA_GRAPH, MINFER_NO_PREFILL_CAPTURE (MINFER_CAPTURE_PREFILL=1 accepted, redundant), MINFER_NO_PINNED_READBACK, MINFER_MMVQ_V1, MINFER_NO_KQ_MMVQ, MINFER_GEMM_TM (64), MINFER_GEMM_K64, MINFER_FUSED_B.

Dispatch guards (unchanged by r60, now protect the default path): MMQ entry nt >= 16 && id % 32 == 0 && !no_prefill_gemm; NB-BT (id / 32) % 8 == 0; plane registration id % 256 == 0 (+ od % 2 == 0 for dsc); fused producers rows >= 16 && dim % 256 == 0; mode-2 auto-degrade under MINFER_GRAPH_DUMP / MINFER_DUMP_DIR / MINFER_TRACE / viz capture; the r60 nb_bt_only flag degrades mode 2 → mode 1 on mixed-quant models.

Post-r60 gates (D3–D5-R and the spec/bench harness) — added after the promotion round, so they are not in the table above:

GateDefaultEffect
MINFER_NO_DECODE_A_FUSEoff1 skips the decode fused-producer A-quantize, restoring the standalone quantize launch (D3-5)
MINFER_NO_Q40_MMVQ / MINFER_NO_Q80_MMVQoff1 forces Q4_0/Q8_0 decode off the MMVQ path (f32 kernels)
MINFER_NO_Q80_P32off1 reverts the q8_0 p32 split planes and their dispatch (doc 104)
MINFER_Q6K_PFon0 disables the q6_K prefetching MMVQ form (D4-2)
MINFER_SMALL_M_GEMMoff1 routes nt 2..8 into the mma BT path (doc 91; measured ~1.7× worse)
MINFER_MMQ_KSPLIT_TARGETmax(256, 2×SM)target resident-block count for the auto-K-split (doc 92; doc 106 SM-parameterized — GB10 keeps 256; an explicit value overrides the formula)
MINFER_DEVICE_TIERoffllama.cpp-style tier key forced onto the selector, e.g. 870 (Orin) / 750 (Turing, MMQ off) / -1 (GENERIC) — soak-test override for the doc-105 device tier tables; the banner logs it as FORCED

Instrument / harness (not backend gates): the specverify instrument reads MINFER_SPECVERIFY_WARMUP_MS / _NTS / _NOUT, spec round tracing uses MINFER_SPEC_DEBUG, bench takes its steady-clock warmup budget from MINFER_BENCH_WARMUP_MS, and MINFER_BENCH_ROW_MARGINAL gates the cold-L2 row-marginal device bench test.

Appendix B — Verification methodology (summary)

The full capstone — the five-gate chain (① parity ×3 via cuda_prefill_mmq / cuda_prefill / cuda_fa_prefill_attention_parity, ② greedy-32 token identity, ③ interleaved A/B medians with the +1.5% whole-prefill bar, ④ the device suite, ⑤ the ncu/nsys/SASS protocol incl. the GB10 metric gaps and the sudo-LD_LIBRARY_PATH gotcha) and the campaign's transferable lessons (baseline anchoring, liveness labels, tile-size vs greedy-identity, the pipeline-value formula, roofline-before-coding, occupancy-before-instructions, SASS-first, stall-mass conservation, layout-transformation locality, mechanism composition, consumer attribution, expiring wall decompositions, phantom results, shared-box memory etiquette) — lives in cuda_optimization_steps/77-verification-methodology.md. Read it before running any A/B on this engine.

Appendix C — Part IV legacy: the pre-Phase-7 roadmap (2026-08-29) and where it ended

Everything below described the deleted imperative path (layer_gpu, forward.rs) on an RTX 4080 Laptop with Qwen2-0.5B Q4_0: prefill 40, decode 20 tok/s (CPU 18/15). Kept for the record; outcomes annotated.

Root cause as diagnosed then: per-op CPU↔GPU ping pong.

CPU path → quantize f32→Q8_0 (CPU) → cudaMemcpy H2D → CUDA kernel → sync → cudaMemcpy D2H → CPU path

~6 PCIe round trips × 24 layers ≈ 144 DMA operations per decode step, 2–7 ms of pure overhead. (Correct for that path; Phase 7's resident-weight graph backend eliminated it structurally.)

Original P0–P5 and actual outcomes:

ItemClaim thenOutcome
P0 full-layer GPU offloadadd Q4_1/Q8_0 kernels, kill 144 DMAs, 3–4×Absorbed by Phase 7 graph backend (weights resident, per-op dispatch, split syncs)
P1 GPU-side activation quantizeGPU q8_0 kernel unused, 1.2×Landed as 8c with a measure-first gate (nt>1 && id≤8192); the q8_0 path LOSES 63% at 7B ffn_down (weight-bound) — a blind wire would have regressed
P2 fused GQA on GPUwire gqa_attn_f32, 1.5×Landed in Phase 7a/7e (gqa_attn_f32_f16kv); prefill attention replaced by FA tiling (8n: 20×); decode attention is 0.12 ms/token at 2K — no longer material
P3 cuBLAS for output projectioncublasSgemm "leverages tensor cores", 2×Closed as 8k (not planned). Two errors: cublasSgemm is FP32 SGEMM — tensor cores require cublasGemmEx with f16/int8; and the need disappeared once 8m's custom wmma GEMM covered large matmuls
P4 tiled quantized matmulllama.cpp MMQ "shared-memory tiling with Stream-K decomposition", 1.5×First judged negative, then REVERSED same-day (8e): the real design is integer __dp4a dots over q8_0 activations with a per-CC launch table — ported in-tree as the decode MMVQ win; "Stream-K" was never part of llama.cpp's MMQ. The prefill int8 version became R1
P5 CUDA graph for launch overheadcapture decode, 1.2×Landed as Phase 7d decode capture/replay (one ~57 µs graph launch per token) + opt-in prefill capture (8g②, default since R3-B). The original "2,000+ launches (…× ~14 heads)" miscounted — heads don't multiply launches; the true figure is ~95 nodes/layer × 28 layers ≈ 2.7K, same order

Implementation order as drawn then:

P0 (layer_gpu) ─→ P1 (GPU quantize) ─→ P2 (GPU attention)
                                      ↘
                                       P3 (cuBLAS) ─→ P4 (tiled MMQ) ─→ P5 (CUDA Graph)

All six landed in some form by 2026-08-31 — none via its original mechanism except P0's idea. Measured budgets recorded at the Part-III era: f16 wmma GEMM ~35 TFLOPS (llama.cpp int8 MMQ ≈ 52 equivalent); MMVQ weight streaming 130–147 GB/s effective vs the 252.7 GB/s read-only probe (93% of the 273 GB/s theoretical); llama.cpp ~197 GB/s on the same decode shape.

CUDA Optimization Step Documents — Master Index

This directory is the expansion layer of docs/CUDA_OPTIMIZATION.md: it writes every step of the twelve sessions, the two Phase-8 batch records (78–79), and ~60 optimization levers from 2026-08-29 → 2026-09-09 as a standalone, readable document — background, the GPU principle (arguing with arithmetic), real code before/after, verification methods, results, and lessons.

Reading guidance: only care about the current state → read docs/CUDA_OPTIMIZATION.md §1 (current status) and the doc 77 methodology at the end of this directory; want to understand why a mechanism is the way it is → find that step in the table below; want to follow the whole campaign → read in number order — the story is continuous.

Status legend: 🟢 LANDED · 🔵 MEAS-ONLY · 🔴 REVERTED (with the veto mechanism) · ⚪ CLOSED (no code, or analytical conclusions)

The writing contract is in STYLE.md.


Part I · Foundations (Era A: Phase 7/8, 2026-08-29 → 08-30)

#doctopicstatus
01phase7-cuda-backendCUDA backend: raw-FFI device layer + graph backend (the 30.7 tok/s starting point)🟢
02wmma-f16-prefill-gemm-8mwmma f16 tiled GEMM: 30.7→1204 tok/s (39×)🟢
03fa-tiled-prefill-attention-8nFA-style tiled prefill attention: 176→8.5 ms/layer (20×)🟢
04decode-start-stall-8odecode start stall: killing the 635 ms heavyweight clone (724→35 ms first step)🟢
05persistent-f16-cache-8president f16 weight cache + dequant folded into the GEMM (→~1400 tok/s)🟢
06decode-mmvq-8edecode MMVQ (dp4a × q8_0): +37% on q4_K🟢
78phase8-correctness-batchthe 8a review batch (11 fixes) + the F32-matmul latent bug + 8h①/8i tests🟢
79phase8-coverage-batchKV f16, shaped Q8_0 GEMM, split-K attention, Q5_K/Q5_1/Q5_0 kernels, the 8l llama baseline🟢/🔵

Part II · The R-and-P5 sessions (Era B: R/P5, 2026-08-31 → 09-01)

#doctopicstatus
07r3-small-model-overheadsmall-model per-token overhead: prefill single-split🟢
08r1-int8-mmq-prefill-gemmint8 MMQ prefill GEMM (opt-in): the parity-first strategy🟢
09r2-mmvq-weight-streamingMMVQ weight-streaming rework: tg128 +6.9%🟢
10r4-split-attention-dim-parallelsplit-attention dim-parallel rewrite: the bulk of the @2K gap🟢
11p5-gemm-tiles-fa-rewriteP5: TM=128 big tiles + FA rewrite (2.37×→1.43×)🟢

Part III · The q4_K campaign (Era C: r5–r37, 2026-08-31 → 09-05)

#doctopicstatus
12r5-r6-rerank-rewrite-specre-ranking + the structural rewrite spec⚪
13r7-r8-raw-byte-kernelraw-byte kernel + wide tile + FA probe🟢
14r9-llama-mmq-referencellama.cpp MMQ reference decode; the shape axis closed⚪
15r10-r11-inner-loop-portreference inner-loop decomposition port; ILP verification🟢
16r12-warp-tile-ldmatrix16-chain warp tile + ldmatrix🟢
17staging-shape-familythe x-tile / j-tile / cp.async-db staging shape family⚪
18r13-counter-forensicscounter forensics vs llama.cpp⚪
19r14-b-fragments-ldmatrixldmatrix B fragments + widened scale reads🟢
20r15-f32-acc-mma-rank1f32-accumulate mma probe + rank-1 term2 rescale🟢
21r16-narrow-kernel-rank1-foldthe narrow kernel gains the rank-1 fold🟢
22r17-wide-warp-remapwide warp remap 32od×64tok🔴
23r18-load-time-b-preexpansionload-time B pre-expansion (W_exp's debut)🔴
24r19-weight-l2-residencyweight L2 residency🔴
25r20-split-phase-a-stagingsplit-phase A staging🟢
26r21-coalesced-block-linear-ablock-linear coalesced A staging🔴
27r22-qa8-xor-swizzleqa8 XOR swizzle (the d/ssum fold reverted separately)🟢
28r23-f16-wall-decompositionf16-path wall decomposition + FA_TKV lift🔴
29r24-scheduling-ladderthe scheduling-structure ladder (the +1.5% bar calibrated)🔴
30r25-sass-opcode-censusthe SASS opcode census; the unroll wall's inertia⚪
31r28-nb-kernel-2blocksthe Direction-A raw-nibble NB kernel, 2 blocks/SM🟢
32r29-nb-kd-loop-unrollthe NB kd-loop unroll🟢
33r30-swar-unpackSWAR unpack: the compiler already did it⚪
34r31-qmajor-sda-repackthe q-major sda scale-read rearrangement🟢
35r32-finite-lever-sweepthe finite-lever sweep: both regions bounded⚪
36r33-hybrid-inner-loopthe hybrid inner-loop port: falsified by SASS identity🔴
37r34-quantize-transpose-prepassthe quantize-transpose prepass (+9.72% of layout-transform locality)🟢
38r35-scale-predecodescale pre-decode🔴
39r36-a-frag-wavefrontA-fragment wavefront economics: H1 falsified⚪
40r37-post-parity-attributionpost-parity whole-prefill attribution⚪

Part IV · q6_K-FA and the promotion (Era D: r38–r60, 2026-09-05 → 09-06)

#doctopicstatus
41r38-q6k-bt-rawbyte-mmathe q6_K BT-style raw-byte mma kernel (+2.9%)🟢
42r39-q6k-kdr2-double-bufferq6_K KDR=2 double buffer (+13.3%)🟢
43r40-third-resident-blockthe launch_bounds(256,3) third resident block (+13.0%)🟢
44r41-q6k-bexpand-uint4q6_K B-expand uint4 widen (+30.7%)🟢
45r42-stage-wide-dsc-readstage-wide dsc scale reads🔴
46r43-pc-sampling-attributionPC-sampling attribution; the pre-expand-B parity FAIL⚪
47r44-wexp-stride-mismatchthe W_exp stride-mismatch root cause; the fix wall-neutral🔴
48r45-cpasync-q6k-a-stagingcp.async q6_K A-side staging🔴
49r46-fap1-fa-auditFAP1: the FA audit + occupancy/conflict levers (−11% kernel but wall-neutral)🔴
50r47-converged-wall-decompositionthe converged-regime wall decomposition⚪
51r48-fap2-register-softmaxFAP2: register-resident softmax (FA 2.43×, prefill +5.6%)🟢
52r49-a-quantize-shared-dedupthe A-quantize prepass's shared-A dedup🟢
53r50-fa-tkv-16FA_TKV 32→16🔴
54r51-producer-fused-a-quantizeproducer-fused A quantize (mode 1)🟢
55r52-skip-write-mode2skip-write mode 2 (MINFER_MMQ_A_FUSE=2)🟢
56r53-q6k-wexp-cpasync-bundlethe q6_K bundle: W_exp + cp.async B staging (prefill 1.05×)🟢
57r54-q6k-exp-optoutthe MINFER_MMQ_Q6K_EXP opt-out (−5.04% for 1.52 GB back)🟢
58r55-swiglu-roofline-prefill-graphthe swiglu roofline + prefill CUDA-Graph (both recorded on the books and skipped)⚪
59r56-q6k-a-cpasync-wdscthe q6_K A-side bundle: A cp.async + the W_dsc plane🟢
60r57-fa-kv-staging-dbFA KV staging double buffer🔴
61r58-q4k-bt-cpasync-transplantthe q4_K BT spec + cp.async-db2 transplant (−12.6%, the pipeline value formula)🔴
62r59-q4k-wdsc-planethe q4_K W_dsc plane + riders🟢
63r59b-clean-remeasurethe clean re-measurement + the baseline-pollution correction⚪
64r60-promotion-default-onthe promotion: the verified gate set flipped default-on (the 1.080× path)🟢

Part V · The decode campaign (§2D: D1–D4-4, 2026-09-07 → 09-09)

#doctopicstatus
65d1-decode-attributionD1 attribution: the split-attention staging depth is the only kernel that grows with KV⚪
66d2-kv-register-stagingD2: explicit K+V register staging (+2.0% @1641) + two negative results🟢
67d3-14b-attribution-bitwise-mmvqD3-1 14B attribution + the D3b bitwise MMVQ trio (1a/1b/1c)🟢/🔴
68d3a-fattn-rewrite-rpwD3a: the 4-warp fattn rewrite reverted (rpw pathology) + the tolerance gate package calibrated🔴
69d3-4-hybrid-rpw-dispatchD3-4: the hybrid rpw dual-kernel dispatch landed; the L2 prefetch reverted🟢
70d3-5-fused-producer-a-quantizeD3-5: fused-producer decode A quantize (quantize launches −78%)🟢
71d3-6-gqa-batching-revertedD3-6: GQA batching — all gates green, still reverted; the 5× L2 re-read was not the residual🔴
72d3-7-attnv-mmvq-rmsD3-7: attn_v MMVQ routing + the rms wide-block/positions memo🟢
73d3-8-fusedqkv-portD3-8: FusedQKV ported to CUDA (both layer classes covered, short KV breaks through)🟢
74d4-2-b0-correctness-fixD4-2: the B0 latent correctness fix (7B dropped 13.5% of down-proj) + all bitwise axes closed🟢
75d4-3-attention-attempt2-artifactD4-3: attention attempt 2 NO-GO + the llama-bench artifact correction🔴
76d4-4-dpl-q6k-finalD4-4: dpl dense split-plane q6_K (+5.5/+4.3%, +7.7/+8.1%); PDL/fused-FFN closed🟢
80d5-0-cost-modelD5-0: the speculative-decoding gate — measured costs + acceptance p≈0.68–0.70; conditional go at d=2, gate = nt=3 verify amortization ≥ 2.5×📏
81d5-1a-verify-gate-measuredD5-1a: the gate measured end-to-end (specverify instrument) — C_T(3)=106 ms, per-token amortization 0.52× vs ≥2.5× required; nt=2–8 batched path costs a flat ~35 ms/token (no regime anywhere), real tile step only at M≥16 → D5 CLOSED by the pre-registered stop rule; external check: llama-cli -md (same pair) lands at 0.99–1.00×🔴
82small-m-multi-token-mmvqsmall-M dispatch fix: multi-token MMVQ + token-looped legacy kernels — 7B nt=3 105.9→29.4 ms (3.60×), nt=8 5.75×, marginal 34.4→4.3 ms/token; bitwise batched-vs-serial on all 8 quants; pre-registered 2.5× bar missed at 1.87× (cost-model error recorded); D5 verdict unchanged🟢

Part V-B · D5-R — speculative decoding reopened and closed (2026-09-12)

The doc 81 §4.3 errata voided the original closure's external anchor, doc 82 made the verify amortization real (2.14× at nt=3), and the plan was rewritten (SPECULATIVE-DECODING-PLAN.md) with the old plan kept as an appendix. Six records:

#docwhatverdict
83d5-r-stage1-spec-loopthe greedy d=2 loop (--spec-draft): second GraphCache, lazy accept loop (unit-tested), namespaced weight registries + nb_bt_only global-mix fix — the three single-model assumptions a second model breaks; 14B d=2 = 1.34×/1.58×🟢
84d5-r-stage2-dual-engine-batterysame-window dual-engine protocol (3 reps × prose/code × 4 cells): minfer 1.33×/1.59× vs llama 1.64×/2.08× — the whole gap = verify row marginal (8.8 vs 2.5 ms/row)🟢
85d5-r-stage3-verify-marginal-ledgernsys per-kernel ledger: nt=3 marginal 17.6 ms = attention nt 2–63 hole 9.0 (legacy per-(token,head) kernel vs the 0.8 ms split path) + matmul 8.1 + elt 1.9 + idle 0.7; nt=9 = dispatch cliff onto padded GEMM📏
86d5-r-stage4a-attention-verify-shapesone gate: fa_prefill nt≥64 → nt≥2 — C_T(3) 56.9→48.8, C_T(9) 101.3→86.4; e2e 1.42×/1.68× (code ≥ llama's same-window 1.64×); ledger projection validated ~5%🟢
87d5-r-stage4b-multi-mmvq-nt16-closedmulti-MMVQ nt 9–16: groups-of-8 = parity (weights re-streamed per group), acc[16] = register spill (111 ms) → doc-82 GEMM boundary stands; d=8 retired (0.71× prose projected, 0.63× measured in doc 88)🔴
88d5-r-stage5-final-batteryfinal battery: minfer d=2 1.42×/1.68× (35.7/42.5 tok/s) = 95%/88% of llama's absolute speed; capture prize verified already banked (R3-B); D5-R closes🟢
89d5-r-row-marginal-localizationthe absolute-gap leader localized without ncu: cold-L2 real-kernel bench + chain nsys + ablation — q4_K 2.4 / q6_K 1.0 / norm-quant 0.7 ms/row; ~half the matmul term = block-per-row activation re-read; fix menu priced (R-rows-per-block, small-M mma, chain hygiene)📏
90d5-r-rrows-per-block-closedmenu item 1 implemented → measured → reverted: ~0 at d=2 (chain keeps act rows L2-hot; block-parallel latency hiding dominates utilization); nt≥6 flatten recorded; small-M mma re-confirmed as the only lever of size🔴
91mma-path-block-starvationBT GEMM already mma.m16n8k32; small-M floor root-caused to block starvation (ntb=1 → 40 blocks); conditional double-buffer shipped (bitwise-safe, prefill guarded); K-split designed as the fix that revives d=8🔬
92ksplit-shipped-flip-resolved-noK-split (grid.z + deterministic reduce) shipped for both BT kernels behind the gate; C_T(9) 86.4→72.9 but the ≤55 flip condition failed — multi-MMVQ stays production; residual = per-tile staging serialization; nt 9..64 auto-ksplit → §3b: enabled on the default path by user decision (default C_T(9) 73.0; d=8 still acceptance-bound)✅
93draft-quant-and-greedy-identitydraft-quant swap is a mixed knob (±3 pts acceptance, opposite signs per cell); greedy identity test FAILS — spec ≠ sequential, flips traced to batched verify attention/softmax; nt-invariance campaign proposed with the identity test as acceptance criterion🔬
94greedy-identity-and-d8-crossinggreedy identity achieved (spec = sequential byte-for-byte; attention nt-invariance + penalty-window cap); d=8 crosses on code with q4_k_m (46.6 tok/s, 96% of llama); prose stays d=2✅
95adaptive-draft-depthadaptive per-round draft depth (beta acceptance + min-window costs + 10% hysteresis + optimism for unseen depths); identity boundary pinned — verify nt ≤ 8 bitwise (multi-MMVQ), nt=9 BT lm_head tolerance-class → adaptive capped at d=7; all four gates pass; adaptive BEATS the best static on both code cells (46.5/48.4 vs 43.9/46.6 tok/s) — the buggy static sweep had never measured d=3..7✅
96nt9-profile-phase0nt=9 verify profiled (stop-gate measured): BT-MMQ = 84% of GPU time, attention ~1%; ncu: both BT kernels at ~20% of both roofs, smem scoreboard stalls = 40–59% of warp cycles → doc 92's staging serialization confirmed dominant; gate verdict PROCEED; cp.async double-buffer staged as the bitwise-preserving Phase-1 lever (narrow EV: prefill lever / identity-relaxed d=8, not spec throughput) — Phase 1 re-priced separately, below 战役 97✅ P0
97spec-conversation-serverspeculative decoding in --cnv and serve (Engine trait hooks + SpecAwareEngine + sibling spec loops mirroring the plain decode token-for-token); server position/termination contract (seed carry, mid-batch stop/EOG ends the turn); pre-existing plain-server bug fixed: cross-request slot GraphCache reuse leaked stale KV (identical requests hid it) → per-request cache reset✅
98bt-cpasync-null96 Phase 1 measured: cp.async double-buffered staging brought to the q4_K BT kernel (full r56 treatment, dbuf extended to ksplit) → NULL on GB10 (kernel µs / C_T / pp512 all baseline-within-noise; dbuf on/off identical) — the BT stall is compute-side (ldmatrix→mma chains), not staging; remaining levers are tolerance-class → patch reverted per doc-90 discipline, D5-R closed at its identity-safe ceiling⚫
99fastverify-p0-p1fast-verify P0/P1: doc 98's tolerance-class claim corrected (int mma exact → wider bitwise-safe set); fragment prefetch / non-volatile mma / cp.async all measured NULL; B-plane XOR swizzle landed (shared wavefront excess 40%→0, −4.3% kernel instance, bitwise 4/4); pc-sampling re-attributes the stall to L1TEX latency × 16-warp occupancy ceiling → knob not built, EV re-priced down✅
100qs-plane-driftthe last priced lever (aligned qs plane) measured timing-NULL under interleaved A/B; sequential "−12.4%" was clock-ramp drift (±7–12% band, 208 MHz idle → 3 GHz) — sequential before/after runs on dgxspark are invalid instruments; repack family closed, D5-R fully closed⚫
101steady-state-methoddoc 100's rule made executable: time-budget warmups in specverify/bench; headline table re-measured tight (pp512 2083 ± 7.2; adaptive 35.9 prose / 44.8 code; d8 collapse reproduced) — doc-95 absolutes confirmed as drift artifacts, structure exact✅
102draft-scale-sweepbigger drafts lose (acceptance bounded by the target, not draft capacity) — default draft stays 0.5B Q4_K_M; flushed + fixed a latent qwen3-loader namespaced-registration bug (Q6_K padded weights under raw name → both models to CPU); post-EOS token-text gate relaxed to warning, cross-family identity 4/4✅ fix + null
103q40-q80-mmvqq4_0/q8_0 decode joins the MMVQ family (8e structure, NEW CODE ONLY — every landed K-quant kernel/arm untouched, size-floored gates + MINFER_NO_Q40_MMVQ/MINFER_NO_Q80_MMVQ fallbacks): 7B q4_0 tg128 48.1→56.8 (+18.1%), 7B q8_0 26.7→28.8 (+7.9%), 7B q8_0 spec e2e 26.8→68.0 tok/s (+154%); same-file llama.cpp comparison — q4_0 decode 1.045× faster than llama.cpp, q8_0 closed 0.87×→0.94×; pp512 unchanged ✓, suite 187/0/3, q8_0 identity 4/4✅
104q80-p32-split-planensys located doc 103's remaining q8_0 gap inside the kernel (97% of decode GPU time; 16 scattered 2B weight loads per block at a 34B lane stride = ~2x q4_0's L1TEX wavefront cost per byte) → p32 split planes (payload 32B/block 16B-aligned for uint4 x2 + dense 2B d plane, traffic unchanged, raw registration untouched, byte-equal outputs): 7B Q8_0 tg128 28.5→31.9-32.1 (+12.3% over f32) = parity with llama.cpp (0.87x→0.94x→1.00x), spec e2e 69.1→79.0 tok/s = 2.47x sequential; q4_0-p32 and 36B-pad variants measured and rejected✅
105device-tier-tablesT-series T1: cc-keyed device tier table + selector (device_tier.rs, pure data, offline-tested) — GB10 Measured, consumer rows Adopted from llama.cpp (Blackwell K-quant caps 5/6/7, Orin K-quants→1, Turing mmq=false ruling #4), GENERIC fallback; MINFER_DEVICE_TIER override; mmq gate = resolved tier flag; batch caps tabled but unwired (R8). Encoding fix: runtime cc is 1201, table keys llama-encoded via conversion✅
106query-formula-gatesT-series T2: auto-ksplit SM-count parameterization (target max(256,2*SM), GB10-invariant), BT smem feasibility gate (single-source C formula + cuda_mmq_smem_bytes(); R2 fixed — externs now query the selected device, not device 0), plane VRAM budget gate on all optional planes (p32/W_exp/W_dsc) for 8 GB unified-memory devices. Tile candidates + cap activation + T3 deferred per plan✅
107c4-packed-q8-kv-cudaC4 #144: the packed Q8_0 cache's two missing tuned routes — the packed fused decode epilogue (attn_bias_rope_store_q8_0, one thread per (head, 32-element K block) and per V block, the store's own quantizer) and the packed FA prefill (fa_prefill_kv<CAUSAL,MAP,LAYOUT> dequantizes each block into the same f16 tile; general kernel stays the fallback). Qwen3-0.6B pp2048 564.5 → 8231.1 tok/s (14.58x; packed/f16 14.7x → 1.038x); 0.5B tg128 q8_0 161.5 → 170.1 (1.052x, the pre-registered 1.15x bar missed — the ticket's 1.18x was the f16-weight arm's cut). Two new device gates, three mutations, 124 launch sites🟢
108c4-dp4a-packed-q8-kv-cuda#186: the dp4a packed Q8_0 K dot — the decode split-K body accumulates int (__dp4a) against a per-(head, block)-quantized query and scales by d_q*d_k once per block; K is never converted to float (V still is). Bar named first (load-attributable share ≥ 10%), measured 20.3% on the 0.5B's hd-64 decode where both layouts share one 1-warp geometry; landed at 0.5B tg128 171.95 → 193.30 (1.124x, packed/f16 1.396x → 1.242x) and Qwen3-0.6B tg128 123.98 → 136.66 (parity with its f16 136.73); pp2048 flat. Real-model class re-measured (tail 2.479504, argmax 0.5527, greedy 9/9); prefill/verify left alone by measurement🟢
109c4-packed-q8-kv-l1-request#202: the packed Q8_0 KV cell's L1 request count, no layout change — the four s8 K/V quant loads per 4 elements become two u16 loads (34k + 2 + 4m is always 2-byte aligned even when only odd blocks are 4-byte aligned). The counter moves exactly as the ticket's mechanism predicted — packed/f16 L1 load-sector ratio 1.712x → 0.9845x (792 904 → 455 840), instructions −2.17% — but the pre-registered +2% tg128 bar was NOT cleared (+0.27%: 193.13 → 193.65, medians of 5 interleaved rounds) and the kernel is only 1.4-1.7% faster (nsys). A partial refutation: the 1.23x packed/f16 decode residual is not L1-request-bound (not L2: 0.56x, not instructions: 1.10x, not sectors: 0.98x). Landed counter-only, byte-identical, no layout/session/CPU change; the latency hypothesis is filed for the next attribution🟢 counter-only

Part VI · Methodology

#doctopic
77verification-methodologythe verification system in full: the gate chain, the GB10 tool protocol, the master library of transferable rules

State at the campaign's close (2026-09-12, after D5-R)

tg128@long KVdevice memory
Qwen2.5-7B Q4_K_M1.074× vs llama.cpp1.052× @1.6K~10.4 GB
Qwen2.5-14B Q4_K_M1.018×0.950× @3.3K~14.1 GB

Prefill: 7B pp3314 ~3581 tok/s (1.080×); 14B pp3254 ~1830 tok/s (1.12×). Speculative decoding (D5-R, closed): 14B+0.5B q4_0 d=2 = 1.42×/1.68× (prose/code, 35.7/42.5 tok/s) = 95%/88% of llama.cpp's absolute speculative speed in the same window; d=8 measured 0.63× (retired). Open leads: ncu on the small-M MMQ gap (counter permissions), nt-invariant accumulation (prose acceptance + exact greedy identity).

77 · Verification Methodology & Transferable Lessons (the campaign's master gate system)

Result: 12 sessions and ~60 optimization levers, every landing/veto adjudicated by the same gate chain — 3 parity tests + greedy byte-for-byte identity + interleaved A/B medians + the suite + the ncu/nsys evidence protocol. Commit: no repo change (this doc expands docs/CUDA_OPTIMIZATION.md Appendix B). Date: 2026-08-30 → 2026-09-09.

1. Background — why a fixed gate set is needed

This campaign had one recurring scenario: an optimization runs fast and correct in isolated kernel tests, but end to end it is either a phantom gain (the fast path is never actually taken) or correctness erosion (a floating-point summation order changed; the parity simulator cannot see it, but the sampling chain can amplify it into divergence).

The matrix: kernel-level correct ≠ end-to-end correct; fast ≠ truly fast (baseline drift, co-tenant interference, silent fallbacks). Without a fixed gate chain, every lever would invent its verification on the spot, and a method invented on the spot is precisely blind to the defect class you most need it for.

So from R1 onward, every lever shares the same gate chain, each gate defending one defect class. This doc first covers the gate chain itself, then the tool protocol (the GB10 specifics of ncu/nsys/SASS), and finally collects the transferable rules the campaign sedimented.

2. Principle — the gate chain and the defect classes it defends

2.1 Gate 1: the parity trio (numeric correctness)

Three independent calls before every landing:

testformdefends
cuda_prefill_mmq1/0, 8 quant types × 8 shapes swept against the host referencethe quant kernels' numeric path
cuda_prefill7/0the prefill graph end to end
cuda_fa_prefill_attention_parity1/0the attention kernel

The 1e-3 tolerance is informative: a nibble-layout error (misaligned unpacking) shows up as a ~1e0-magnitude deviation, while legitimate f32 rounding noise is only ~1e-5. So the 1e-3 tolerance is not "lowering the bar" — it is a discriminative window that separates bug classes from noise classes.

2.2 Gate 2: greedy byte-for-byte identity (end-to-end correctness)

-n 32 --greedy --seed 42 against prompt2k on the pre-change binary; the token stream must be byte-for-byte identical.

It defends against graph-level/memory-level corruption the parity fixtures miss — two real cases: r52's rms kernel out-of-bounds write (OOB), r58's smem buffer-1 cross-write. Such defects may happen to be invisible in kernel-output tests, but once any layer is polluted the whole generation sequence necessarily diverges.

The FA exception (important): a tile-size change necessarily changes the grouping order of floating-point accumulation (the r50/r57 lesson); even with parity all green, greedy will diverge at some step. For changes of this "naturally breaks byte-for-byte identity" class, the gate is swapped for a whole calibrated tolerance package (see §2D D3a/D3-6: kernel-level vs CPU ≤1e-4 on realistic outlier data, the argmax hard gate, the rp=1.0 greedy identity stream + sampler knife-edge attribution).

2.3 Gate 3: interleaved A/B medians (performance truth)

3×/5× same-window pairing, warmup, alternating order; headline numbers require distribution separation (min-new > max-base); the whole-prefill landing bar is +1.5% (relative to the re-measured baseline, calibrated at r24).

The alternating order is the key design: the GPU is shared (sglang is co-tenanted on dgxspark); running one side back-to-back disguises window drift as a trend. Paired medians cancel the co-tenant noise.

2.4 Gate 4: the suite and co-tenant flakes

The suite grew from 166/0/3 at the campaign's start to 174/0/3 (gate 1's byte-level tests were progressively promoted to permanent tests). Tests that flake occasionally under a co-tenanted window are adjudicated with an isolated --exact rerun — a flaky test is either proven isolated or fixed, never silently retried to green.

2.5 Gate 5: the ncu/nsys/SASS protocol (GB10 specifics)

This GB10 (DGX Spark, sm_121) has several unavoidable tool pitfalls:

  • ncu must go through sudo -n env LD_LIBRARY_PATH=...: plain sudo strips environment variables, and ncu silently profiles the legacy path (the r56 lesson — you think you are measuring the new kernel but are measuring the old one);
  • GB10/GB20B has no dram__*/launch__grid_size/shared-sector counters — bandwidth rooflines can only be derived from lts__t_sectors_aperture_device (L2 sector count × 32 B) plus byte counts (established at r55);
  • ncu serializes replay; nsys is the wall-clock authority: ncu's per-kernel times are distorted under serialization and are used only for structural metrics like occupancy/sectors/residency;
  • PC-sampling's attribution rule: --page source attributes a stall to the consumer instruction waiting on it, not the instruction producing the bytes (r20/r43) — do not read the causality backwards;
  • SASS first: read cuobjdump -sass before writing any lever — cp.async in SASS is LDGSTS.E.BYPASS.128 (grep LDGSTS to verify the compiler actually emitted it, r45); ptxas -Xptxas -v reports the register/spill/occupancy budget (r40 used it to confirm the 3rd resident block).

3. Implementation — the transferable rules that sedimented out

Every rule below was paid for with real campaign losses; the source is in parentheses.

Baseline anchoring (r59b): every A/B baseline must be behaviorally anchored in the same window — re-measure a known binary with a historical record, or rebuild the baseline commit from a worktree. Idle co-tenancy = clean equivalence; without anchoring you may not attribute performance fluctuation to the co-tenant tax. One of r59's +Δs was proven by r59b to be a polluted baseline; the clean re-measurement corrected the number.

Liveness checks (r53/r54): an optimization with "a fallback as safety net" must have a counter/label for "did the fast path actually go live" — neither parity nor greedy can see a fast path that is never taken. Distinguish an intentional fallback (exp=off) from an accidental one (fallback!).

Tile size vs greedy identity (r50/r57): for kernels sensitive to accumulation order, strict byte-for-byte identity is satisfiable only for changes that "preserve the accumulation order" — once the tile size changes, ULP regrouping necessarily happens, and green parity is not enough. This rule directly spawned D3a's calibrated tolerance package.

The pipeline value formula (r58, the mirror of r45): a staging mechanism's wall-clock value = what it removes − its granularity cost. Replacing expensive work (q6_K r39/r53/r56) → +13/+5/+2.35%; replacing cheap copies (q4_K r58) → −12.6%. "Not dead, waiting for its scenario" is valid only while the mechanism's cost model holds.

Roofline before code (r55): derive the byte-traffic lower bound first (when there is no dram__*, use sector count × 32 B); if even a perfect kernel cannot pass the bar, skip the implementation. D4-2's Lever A (the llama L2 prefetch port) was this rule's pure-inference veto — closed without writing a single line of code.

Buy occupancy first, then save instructions (r13→r25→r28/r29): at 1 block/SM an instruction surplus is real but wall-clock inert; buy the occupancy first and the same instruction reduction starts paying (+2.6/+2.8%). At 3 blocks/SM, 4 B of spill is irrelevant (r40).

The compiler already did it (r30/r32/r33): read the SASS before writing a lever — if ptxas has already scheduled a source-level rearrangement, changing it is a SASS-level no-op. r33's "line-by-line port" of the llama inner loop died exactly here: the SASS was completely identical, so of course the performance was too.

Stall mass conservation (r20/r21): fix one bottleneck and the stall moves to the next (latency → lg_throttle → wait) — strike in stages in that order; do not expect one fatal blow.

Layout-transform locality (r34): promote the layout transform to a prepass (or fold it straight into the producer) instead of adapting inside every tile consumer — moving the transform out of the kernel itself is worth +9.72%.

Mechanisms compound across the wall (r53/r56): two levers individually wall-clock-neutral (one reducing WORK, one reducing WAIT) compose almost additively once the first one unblocks the bottleneck.

Attribute to the consumer (r42/r43): the stall counter tells you the resource; PC-sampling tells you the instruction waiting on it — cut the latency where it is exposed, not where the bytes move.

Wall decompositions expire (r37→r47): re-attribute the whole wall after each line converges; the hidden tax (q6_K prepass's +31.6 ms) must be weighed against its gain. D4-1 overturning D3-8's "matmul aggregation 2.9 ms" attribution is this rule's decode edition.

Phantom results (r8, P5·3, r52a): silent fallback / attribute-set failure / OOM masking all produce "fast and wrong" timings and fake errors — guard trips must report loudly, and whether the fast path fired must have evidence.

Shared-machine etiquette (2026-08-31): a kernel OOM under pool exhaustion kills someone else's workload (it happened once); bare allocation probes are forbidden; check free -g before the suite; while sglang is serving, run only single-process 7B-scale benches.

4. Verification

This doc is the verification system itself and has no independent object to verify. The way it is "verified": across the 12 sessions, not one REVERTED decision was ever proven afterwards to have been the wrong revert, and the single defect that slipped through (D4-2 B0: 7B losing 13.5% of its down-proj compute) is precisely this system's blind-spot case — when both sides of the A/B share the same bug, the comparison is bit-identical. The rules added from it: cross-binary comparison must use the -n 1 first-step dump (token cascades pollute all subsequent KV), and D3-6's sampler knife-edge attribution gate (the --repeat-penalty 1.0 identity stream = the clean kernel-numerics gate).

5. Results

The gate chain's final form evolved with the campaign:

  • R1 (2026-08-31): parity ×3 + greedy-32 + interleaved A/B + suite 166 — the four-piece set finalized;
  • r24 (2026-09-01): the +1.5% whole-prefill landing bar calibrated;
  • r50/r57 (2026-09-05): the applicability boundary of byte-for-byte identity was delimited, spawning the calibrated tolerance package;
  • r59b (2026-09-06): the baseline-anchoring rule established;
  • D3a (2026-09-07): the decode tolerance gate package calibrated (the argmax hard gate, sampler knife-edge attribution, kernel-level ≤1e-4 on outlier data);
  • D4-2 (2026-09-09): the -n 1 first-step dump rule + closing the cross-binary comparison blind spot.

6. Lessons

  1. The gate chain's value is the matrix of defect classes it defends — every gate can state "what I defend" in one sentence; a gate that cannot is ritual, not verification.
  2. All isolated "fast and correct" evidence is a suspect — liveness, baseline, co-tenancy, fallback: only after these four suspects are eliminated one by one is a headline allowed.
  3. Magnitude discrimination matters more than tolerance values: the 1e0 vs 1e-3 vs 1e-5 deviation magnitudes directly identify the bug class.
  4. A revert is not a failure: the campaign's 12 negative-result levers are each archived with their mechanism, and later sessions (e.g. D4-1's design doc) reused these veto evidences directly, saving at least three rounds of duplicate construction.
  5. On a shared machine no performance number is "naked" — every conclusion is a difference after paired anchoring, not an absolute value.

← 76 · Index →

01 · Phase 7 — CUDA backend: raw FFI device layer + graph backend 7a–7e (LANDED)

Result: the first working NVIDIA GPU inference path; 7B q4_k_m @2K prefill 30.7 tok/s (the Era-A baseline anchor — ~110× behind llama-bench at the time, but from here on every optimization step had ground to stand on). Commit: 0dc2a54 (raw-CUDA-FFI device layer; graph backend 7a–7e landed phase by phase per docs/CUDA-BACKEND-DESIGN.md, with the 7e series closing out on 2026-08-29). Date: 2026-07-15 (device-layer commit) / 2026-08-30 (Era A record).

1. Background — where things stood

Before this step, minfer's GPU story existed only on macOS: the Metal backend (metal.rs + metal.metal) ran on the same compute-graph architecture (Phase 6 had already deleted Qwen2's imperative forward, so all inference went through build graph → assign backend → allocate → execute). On x86-64 Linux there was no GPU path at all — every token came out of the CPU's AVX2 kernels.

There was also an earlier failure on record (kept in CUDA_OPTIMIZATION.md Appendix C, the "Part IV" history): a CUDA attempt made without the compute graph, shaped so that every op call moved weights/activations back and forth between CPU and GPU. That experiment's outcome was distilled into the Part-IV diagnosis: per-op H2D/D2H round trips are fatal at the 7B scale — the GPU idles between ops waiting for transfers, the transfer cost eats the entire speedup, and in the end the whole path was abandoned, leaving only a problem list.

The goal of Phase 7 was therefore explicit: not "port a few kernels to CUDA", but make CUDA the compute-graph architecture's third backend — peer to CPU and Metal, behind the same Backend trait and the same scheduler. Two preconditions make that possible, and they are also the thesis of this document:

  1. Resident weights: every matmul weight is uploaded to device memory once at model load and registered by name in a device-side registry; no forward ever touches host memory for weights again.
  2. Graph-backend dispatch: the entire prefill/decode graph executes as one CUDA split; activations cross PCIe/unified memory only at split boundaries.

Without this step, everything that follows (8m's wmma GEMM, 8n's FA attention, R1's MMQ, the whole 39×/8.1× campaign) has no footing: 7B on pure CPU is a single-digit token rate, and levers like wmma, cp.async and tensor cores only mean something in an architecture where weights are already resident and dispatch has already eliminated the round trips.

The hardware target is the DGX Spark (GB10, sm_121, Blackwell family, unified memory architecture). That is also why every number in the chapters that follow comes from the GB10 — it is the fixed battlefield of this campaign.

2. Principle — the GPU mechanism

Why per-op round trips are fatal. Let the byte counts speak. The 7B q4_k_m weights total about 4.4 GB. In the Part-IV shape, every matmul op re-streams the weights on every forward (from the host, or from a one-shot device-side staging buffer). A prefill of 2048 tokens × 28 layers × 7 large matmuls per layer, with non-resident weights, means the device either pulls the data across PCIe over and over or shuffles it around on-chip over and over — the GB10 DRAM roofline is 273 GB/s (a number later measured precisely in the r55 roofline audit), and 4.4 GB × a few tens of re-transmissions puts transfer time in the tens of seconds. The measured 30.7 tok/s means a 2048-token prefill takes 66.7 seconds — the true price of the "transfer-dominated" shape, and the reason it was abandoned.

What residency + graph dispatch remove. With weights resident, a forward's cross-device traffic shrinks to activations only: 2048 tokens × 3584 dims × f32 ≈ 29 MB per operator-level buffer, and most buffers never leave VRAM within their device-side lifetime (the allocator's liveness reuse).

The dispatch-side accounting splits in two. Prefill: the graph dispatches dozens of nodes as a single CUDA split, and cross-device copies happen only at split boundaries (ideally zero — after 7e③ both prefill and decode are a single split). Decode (nt==1): a few hundred kernels per step, each Rust→C launch costing ~2–5 µs of host time, which accumulates to milliseconds per step — a non-negligible share of a 20+ tok/s target. That is why 7d introduced CUDA Graph capture/replay: every launch inside one execution window is natively recorded into a cudaGraphExec_t, and each subsequent step replays the whole step with a single cudaGraphLaunch. This is the CUDA counterpart of the Metal side's command-buffer batch submission.

Host→device filling has to be async too. Decode writes small inputs (token ids, positions, …) into device buffers every step; a cudaMemcpy from pageable memory introduces driver-internal sync points. 7e⑥ turned input filling into pinned-staged async: host data first lands in a pinned-memory ring staging area (STAGING_SLOTS slots), then goes to the device via cudaMemcpyAsync — same-stream ordering makes this naturally lock-free against the kernels that consume it, and a sync is needed only when the ring runs out of slots. write_host (excerpt 3 in §3.2) is that layer.

Why raw FFI. This project's hard line is zero ML-framework dependencies; the CUDA runtime API actually used here amounts to about a dozen functions (cudaMalloc/cudaMemcpyAsync/cudaStreamCreate/cudaStreamBeginCapture/cudaGraphLaunch/cudaGetDeviceProperties…), which can be declared directly with extern "C" — no bindgen, no third-party crate. build.rs compiles src/cuda_kernels.cu with nvcc into a static library linked into the binary; when nvcc is absent the whole feature degrades cleanly to "does not compile", never touching CUDA at all. The initial kernel set was 12 kernels (q4_0 matmul, rms_norm, rope, silu, swiglu, gqa_attn, …), Q4_0-only, with every other quantization type transparently falling back to CPU — the coverage was deliberately narrow: first make "a whole layer runs on the GPU" true.

3. Implementation

3.1 Design choices (why this shape and not another)

  • The CudaState singleton mirrors MpsState: the division of responsibilities — weight registry, buffer pool, KV management — directly reuses the shape already proven by the Metal backend; in the graph architecture the two backends are "the same thing on different devices".

  • Phased landing (7a–7e), each phase independently testable, with different acceptance gates:

    PhaseContentAcceptance
    7aSkeleton: CudaBackend struct/buffer pool/trait impl; alloc.rs wiring (enable_cuda + supports()); execute_node handles only InputPool alloc/write/read round trip, copy_across both directions, KV persistent regions survive re-runs
    7bFull per-op dispatch + error checkingPer-op parity (bit-identical / tolerance, two tiers) + whole-layer chain + whole-graph logits + greedy equality
    7cModel wiring (qwen2/qwen3 cuda gate, weights_on_cuda, FusionPass)GB10 E2E greedy == CPU for three models; negative paths fall back cleanly to CPU
    7dCUDA Graph capture/replay (the dispatch tax from §2)Replay bit-identical, recapture triggers, 200-token KV growth, injected-failure fallback
    7ePolish + independently landable perf items (①–⑥)Each item A/B'd on its own

    A bad phase never contaminates the ones before it; this is also how the "wrap, do NOT stub" decision materialized — the cuda.rs device layer was kept as-is and the graph backend merely wraps it.

  • At the initial commit 0dc2a54 the code still hung off kernel.rs dispatch + a forward.rs whole-layer GPU path; the 7a–7e graph backend later replaced that temporary path.

  • The error model follows the GPU Safety conventions: kernel-invariant violations always return Err from execute_node and abort via the scheduler — never a silent fallback to CPU. CUDA's failure model is friendlier than Metal's (it does not freeze the whole machine): a fault surfaces on the next API call, so every launch is followed by a cudaGetLastError check.

  • At the initial commit 0dc2a54 the code still hung off kernel.rs dispatch + the forward.rs whole-layer GPU path (Q4_0-only, other quants transparently falling back to CPU); the 7a–7e graph backend subsequently replaced that temporary path, while the device layer itself (cuda.rs/cuda_kernels.cu) was preserved in full — which is where the plan's "wrap, do NOT stub" decision came from.

3.2 Key code

What the raw FFI layer looks like (excerpted from src/cuda.rs as introduced by 0dc2a54; the hand-written repr(C) cudaDeviceProp exists because the CUDA headers cannot be depended on):

#![allow(unused)]
fn main() {
// ─── FFI declarations for CUDA runtime API ────────────────────
extern "C" {
    fn cudaSetDevice(device: i32) -> i32;
    fn cudaFree(ptr: *mut std::ffi::c_void) -> i32;
    fn cudaMalloc(ptr: *mut *mut std::ffi::c_void, size: usize) -> i32;
    fn cudaMemcpy(dst: *mut std::ffi::c_void, src: *const std::ffi::c_void,
                  count: usize, kind: i32) -> i32;
    fn cudaMemcpyAsync(dst: *mut std::ffi::c_void, src: *const std::ffi::c_void,
                       count: usize, kind: i32, stream: *mut std::ffi::c_void) -> i32;
    fn cudaStreamCreate(stream: *mut std::ffi::c_void) -> i32;
    fn cudaStreamSynchronize(stream: *mut std::ffi::c_void) -> i32;
    fn cudaGetDeviceCount(count: *mut i32) -> i32;
    fn cudaGetDeviceProperties(prop: *mut cudaDeviceProp, device: i32) -> i32;
    // …kernel launch wrappers (launch_q4_0_q8_0_matmul and 11 more — 12 total) in another extern "C" block
}
}

The weight registry (current tree src/cuda.rs:1446) — note that the "resident" semantics land here as idempotent registration:

#![allow(unused)]
fn main() {
pub fn register_weight(&self, name: &str, data: &[u8]) {
    if data.is_empty() { return; }
    {
        let w = self.weights.lock().unwrap();
        if let Some((_, size)) = w.get(name) {
            if *size == data.len() {
                // Device weights are immutable: same name + size ⇒ the
                // same GGUF tensor. Reuse the existing device copy instead
                // of leaking one buffer per load.
                return;
            }
            // Different size: replace the entry. The stale buffer is
            // deliberately NOT freed — a live captured graph may still
            // reference it; the leak is bounded by distinct (arch, tensor).
        }
    }
    let mut ptr: *mut std::ffi::c_void = std::ptr::null_mut();
    let err = unsafe { cudaMalloc(&mut ptr, data.len()) };
    // …cudaMemcpy H2D upload, stored into the registry by name
}
}

The capture/replay entry layer (7d, src/cuda.rs:2297; on begin failure it clears the error and returns false so the caller falls back to direct launches instead of continuing on a poisoned stream):

#![allow(unused)]
fn main() {
    pub fn graph_begin_capture(&self) -> bool {
        let stream = self.stream();
        let err = unsafe { cudaStreamBeginCapture(stream, 1) };
        if err != 0 {
            unsafe { cudaGetLastError(); }   // clear the error — never leave the stream poisoned
            false
        } else {
            true
        }
    }
    // graph_end_capture: cudaStreamEndCapture → cudaGraphInstantiate →
    // stored into decode_graph_exec; each later step replays with one cudaGraphLaunch.
}

Two passages from the Backend trait implementation that best show the architectural constraints (current tree src/graph/cuda_backend.rs:1354): when execute_node fails inside an open capture window, the window must be aborted loudly (otherwise later input fills get recorded into a multi-step mega-graph and every replay double-commits KV); and write_host uses 7e⑥'s pinned-staged async fill:

#![allow(unused)]
fn main() {
fn execute_node(&mut self, node: &CNode, in_bufs: &[usize], out_buf: usize,
                kv_pair: Option<(usize, usize)>) -> Result<(), String> {
    match self.execute_node_inner(node, in_bufs, out_buf, kv_pair) {
        Ok(()) => Ok(()),
        Err(e) => {
            // A node error during an open capture window dooms the window:
            // nothing would close it — later input fills would be RECORDED
            // into the window and the eventual close would cache a
            // multi-step graph (double KV commit on every replay).
            if self.capturing.is_some() {
                self.abort_capture(&e);
            }
            Err(e)
        }
    }
}

fn write_host(&mut self, id: usize, data: &[f32]) -> Result<(), String> {
    // …size guard…
    // 7e⑥: pinned-staged async fill (same-stream ordering makes this
    // race-free with the kernels that read the input).
    let src = unsafe { std::slice::from_raw_parts(data.as_ptr() as *const u8, bytes) };
    self.state.write_input_async(src, dst);
    Ok(())
}
}

3.3 Pitfalls

The parallel test suite was the touchstone of this phase; it flushed out several traps a single-threaded path would never meet:

  • Capture is per-stream, not per-thread. While one backend holds an open capture window, any other thread's enqueue onto the shared stream (fills, copies, launches) gets recorded into that graph. The fix is a process-level stream_lock: the capturer holds the lock for the entire window, and every other stream touchpoint takes it per operation.
  • Weight-registry leak. The parallel suite loads the same GGUF repeatedly; the early implementation cudaMalloced a fresh copy of the weights on every load — 100+ OOM aborts. The fix is the "same name + size ⇒ reuse" rule above; on a size mismatch the entry is replaced but the stale buffer is deliberately NOT freed (a live captured graph may still reference it).
  • ModelLoadGuard: loading another architecture in parallel flips the CUDA gate between two forwards, producing a mixed state of "CPU-allocated persistent KV regions + a CUDA-executed split". The loader holds a re-entrant lock until registration completes.
  • No legacy cudaMemcpy inside a capture window: copy_device_to_device was switched to cudaMemcpyAsync. A mine that a graph-free engine would never step on, yet it sat one step away.
  • The 7e② lesson: a sync added temporarily in the matmul dispatch silently corrupts the capture window (garbage output whenever the graph is on) — the rule was hardened into: never sync inside execute_node. The same round produced a "faster but wrong" red flag: a v-selector that read only ¼ of the y values was actually faster — proof that any change to load counts must be suspected of breaking correctness first.
  • The 7e③ quantization-convention trap: minfer's Q4_0 stores round(v/d) + 8, so the embed kernel must subtract the 8 back. A version without the -8 passed the entire suite (the coverage happened to include no q4_0 model) and produced garbage only on q4_0 models — once a shared kernel changes, E2E must be re-run per quantization type.

4. Verification

The gates, set per phase — each defends against one class of regression:

  • 7a: buffer-pool alloc/write/read round trip, copy_across CPU↔CUDA in both directions, KV persistent regions surviving a re-run of alloc_graph — defends the lowest-level addressing/lifetime errors.
  • 7b: per-op parity — elementwise/rms ops bit-identical against vec_ops; each quant type's matmul within tolerance against the CPU Q8_0-activation reference and bit-identical against itself; RoPE / KV store-load round trip (including n_past growth) / GQA attention matching the CPU implementation; whole-layer chain tests + 0.5B Q4_0 whole-graph CUDA-vs-CPU logits + greedy text identical word for word — defends against "the GPU computed a different answer".
  • 7c: GB10 E2E over three models (0.5B Q4_0 / 0.6B Q8_0 / 7B Q4_K_M), greedy == CPU greedy; prefill→decode→multi-turn session checks KV growth and graph reuse; negative paths (MINFER_DISABLE_CUDA=1, no device, quant outside the gate) → clean CPU fallback — defends against wiring and gating mistakes.
  • 7d: replay decode logits bit-identical against direct launches; a pool_gen change triggers recapture; a 200-token generation run (KV growth + the nk fix); MINFER_NO_CUDA_GRAPH=1 A/B; injected capture failure → session fallback still correct — defends against the state leaks specific to the capture/replay layer.

5. Results

  • The first working CUDA path: 7B q4_k_m @2K prefill 30.7 tok/s. Against llama-bench's contemporaneous ~3401 tok/s @2K that is ~110× behind — a gap deliberately kept as an honest baseline: it measures exactly the "structure in place, kernels not yet optimized" starting point.
  • The 7e series close-out (each item A/B'd independently): 7e① judged the graph-vs-forward 0.449/0.525 diff a path-identity artifact (after Phase 6, forward also goes through the graph, so the "reference" was itself the CUDA graph — cross-backend f32 reduction-order noise, not a defect); 7e② vectorized q4_K/q6_K decode kernels 8.4 → 26.4 tok/s (3.1×); 7e③ moved embed/GetRows onto the device, making prefill/decode a single CUDA split (zero cross-backend copies); 7e④ F32×F32 matmul and 7e⑤ FusedFFN completed the dispatch table; 7e⑥ pinned-staged async input filling (§2).
  • The structural deliverables (what matters most to every later chapter): resident weights, by-name registration, a single split, CUDA-graph replay — on this skeleton, 8m swapped in a new GEMM kernel and turned 30.7 into 294 and then 1204.

6. Lessons

  1. Resident weights + graph-backend dispatch are the precondition for every later optimization; per-op H2D/D2H on the hot path is fatal at the 7B scale (the Part-IV lesson — this step killed it at the architecture level).
  2. A capture window is a process-level mutually exclusive resource: the per-stream semantics plus "never sync inside execute_node" must be written down as hard rules.
  3. The weight registry must be idempotent (same name+size reuses, replacement does not free) — only a parallel test suite can flush out this class of trap.
  4. Raw FFI is enough: a dozen runtime functions + kernel wrappers, no bindgen and no third-party crate — the dependency red line and engineering cost can both be had.

← Index · 02 →

02 · 8m/8m② — Tiled wmma f16 prefill GEMM (LANDED)

Result: 7B q4_k_m @2K prefill 30.7 → 294 (8m) → 1204 tok/s (8m②), 39× for the whole row; GEMM kernel 31 → 35 TFLOPS. Against llama-bench's 3401 @2K, the gap converged from ~110× to ~2.8×. Commit: ba3f317 (8m: tiled wmma f16 GEMM), cdc6599 (8m②: cp.async tile staging). Date: 2026-08-30 (both commits the same day, 50 minutes apart).

1. Background — where things stood

After Phase 7 raised the CUDA backend's skeleton (the previous chapter), the engine had its first working path, but 7B @2K prefill was only 30.7 tok/s. The bottleneck diagnosis was already clear at that point: prefill was running kernels shaped for decode. The decode kernels (the per-op matmuls that predate MMVQ) have a grid shape of grid.y = nt — one row of blocks per token, and every block re-reads the whole weight matrix (or its slice of output rows) from VRAM. At nt==1 that is a fair price; at nt=2048 the weight stream is re-read 2048 times.

Arithmetic shows how severe the mismatch is: 7B q4_k_m weights ~4.4 GB, GB10 DRAM roofline 273 GB/s — ideally one pass of the weights over DRAM takes ~16 ms. But 30.7 tok/s means the 2048-token prefill runs 66.7 s — the extra tens of seconds are almost entirely "the same bytes re-entering L2/SM over and over".

Two roads were open for the fix at the time:

  1. Keep the quantized format and build an int8 tensor-core GEMM (the llama.cpp MMQ route) — numerical correctness demands a q8 activation pipeline and per-type dedicated kernels: a large engineering effort;
  2. Dequantize the weights to f16 first and let one standard f16 tensor-core GEMM swallow the entire prefill — dequant logic is isolated per type into separate small kernels, the GEMM body is a single one, and the wmma API is directly usable.

8m chose route 2. The reason is the leverage structure: route 2 covers all 8 quantization types with one GEMM, breaks through the wall first, and establishes the f16 baseline; route 1 (R1's int8 MMQ) was later built as its own effort and became the new default at r60 — but that is a later story, and its existence in no way negates the f16 baseline's value: MINFER_MMQ=0 remains the escape hatch to this day.

2. Principle — the GPU mechanism

2.1 wmma 16×16×16 fragments

NVIDIA's warp-level matrix API (nvcuda::wmma) wraps one 16×16×16 multiply-accumulate as a warp-cooperative operation. Three objects:

  • fragment<matrix_a, 16,16,16, __half, row_major>: a 16×16 sub-block of A, held in pieces by the warp's 32 lanes (which lane holds which element is opaque — the root of a pitfall hit later);
  • fragment<matrix_b, 16,16,16, __half, col_major/row_major>: a 16×16 sub-block of B; the layout flag decides the addressing interpretation when loading from shared memory;
  • fragment<accumulator, 16,16,16, float>: the 16×16 result block of C, accumulated in f32.

load_matrix_sync loads a fragment from shared memory (with a leading-dimension stride), mma_sync(acc, a, b, acc) performs the multiply-accumulate (underneath are HMMA tensor-core instructions: f16 inputs, f32 accumulation), and store_matrix_sync writes back. f16 inputs give each tensor-core instruction several times the throughput of the contemporary f32 path, while f32 accumulation preserves numerical accuracy — that is why "dequant to f16" can harvest the hardware dividend.

2.2 Tile geometry: why 64×64

The 8m② baseline shape: output tile TN=64 (nt direction) × TM=64 (od direction), k-step width KS=32, 256 threads = 8 warps. Each warp owns one output sub-block of 32 rows (nt) × TM/4 columns (od), i.e. 2×(TM/64) pairs of 16×16 accumulator fragments.

Shared memory budget (KS=32, TM=64): A panel 2×64×32 halves (double buffer) = 8 KB, B panel 2×64×32 halves = 8 KB, C scratch 8 warps × 256 f32 = 8 KB, 24 KB total — inside the 48 KB static limit, so multiple blocks can be resident per SM. Two opposite directions were later nailed down by measurement (the launcher comments keep the record): KS=64 grows dynamic smem to 56 KB and halves resident blocks, −38% (depth traded for inverted occupancy); TM=256 is wider but hits the same wall, −3%.

The weight-traffic arithmetic is the whole point of this step. The grid is arranged as blockIdx.x = nt tile, blockIdx.y = od tile, with blocks consecutive in blockIdx.x sharing the same od-tile's B panel (64 rows × id columns f16, ~0.5 MB f16 for 7B's ffn_gu): this panel streams from DRAM into L2 once, and all nt/64 nt-tile blocks then hit it from L2. Compared with the decode-shaped kernels' "re-read all weights per token", weight DRAM traffic drops from nt× to ~1×.

The overall FLOP ledger: one 7B @2K prefill's matmul total ≈ 2 × 6.9e9 × 2048 ≈ 2.8×10¹³ FLOP. At 31 TFLOPS (8m) the GEMM takes ~0.9 s — the difference against the 2048/294 ≈ 7.0 s wall at 294 tok/s is made up of attention (still 176 ms/layer then, the next chapter's protagonist), the dequant pass, and elementwise work.

2.3 The cost and payoff of dequant-to-f16-then-GEMM

Payoff: one GEMM covers 8 types; all type differences are isolated into per-type dequant kernels (each a plain "read block → compute f16 → write row" loop); tensor cores at full strength.

The cost: one matmul call passes the data three times — quantized weights read in (~4.4 GB for the whole 7B model), f16 written out (~13.6 GB), and the GEMM reads the f16 back (~13.6 GB). In the 8m shape the dequant re-runs on every call (the scratch buffer buf_f16_w is rewritten per call) — 8p's record prices this pass at 288 ms/call (7B). This directly spawned 8p's two successors: a persistent f16 cache dequantized once at load time (gated at ≥2 GB, W16_ENABLE_BYTES), and the more memory-frugal dequant-in-GEMM fused kernel (MINFER_FUSED_B=1). In the 8m era this cost was knowingly paid: it is still far smaller than the 2048× weight re-read it replaces.

2.4 8m②'s cp.async

In the synchronous-staging shape, every warp must wait for the global→shared round trip to complete before it can start computing at every 32-k step (31 TFLOPS stalls right there). cp.async.cg.shared.global (sm_80+) turns 16 B chunk copies into async operations: the main loop issues the fetch for tile k+32 while computing tile k, and commit_group/wait_group 1 maintains the double-buffer rhythm of "one group in flight, one group ready" → 35 TFLOPS.

3. Implementation

3.1 Design choices (why this shape and not another)

  • One dequant kernel per type + one shared GEMM (dispatch on type_id 0–7) rather than 8 copies of the GEMM. The GEMM only ever sees __half* B and never knows the source format.
  • Gate: nt >= 16 takes this path, and id % 32 == 0 — the latter is both a requirement of the block math (all GGUF types use 32-element base blocks) and guarantees the 16 B alignment of the uint4 tile loads; real models' ids (3584/5120/13824…) satisfy it naturally.
  • The 64×64×32 tile (see §2.2's arithmetic and the two REVERTED counterexamples).
  • C lands through shared memory: accumulator fragments store_matrix_sync into the per-warp Cs, then a lane loop writes global with the nt/od tail masks — the tail masking lives in exactly one place.
  • The dynamic smem opt-in mechanism: TM=128/KS=64 instances need 56 KB, over the 48 KB static limit, so a cudaFuncSetAttribute path is mandatory (this became one of §3.3's pitfalls).

3.2 Key code

Type isolation on the dequant side. A simple-type sample (src/cuda_kernels.cu, with the 8m section-header comment excerpted alongside):

// ─── 8m: prefill dequant-to-f16 + wmma HGEMM ────────────────────────────
// Prefill (nt >= 16) routes quantized matmuls through ONE tiled
// tensor-core GEMM instead of the decode-shaped kernels whose
// grid.y = nt re-streamed the whole weight matrix once per token
// (7B q4_k_m @2K: 30.7 tok/s vs llama.cpp MMQ 3401). Weights are
// dequantized to f16 once per call into a scratch buffer, activations
// converted to f16, then C[nt, od] = A[nt, id] · B[od, id]^T via
// 16x16x16 wmma with f32 accumulation. Gated on id % 32 == 0.

__global__ void dequant_q4_0_f16(
    const uint8_t* __restrict__ w, __half* __restrict__ out, int od, int id
) {
    int nb = id / 32;
    long long g = (long long)blockIdx.x * blockDim.x + threadIdx.x;
    if (g >= (long long)od * nb) return;
    int row = (int)(g / nb);
    const uint8_t* blk = w + g * 18;          // Q4_0 block = 2B scale + 16B nibbles
    float d = h2f(*reinterpret_cast<const uint16_t*>(blk));
    const uint8_t* q = blk + 2;
    __half* o = out + (long long)row * id + (int)(g % nb) * 32;
    // minfer Q4_0 stores round(v/d) + 8 (same -8 offset as the matmuls).
    #pragma unroll
    for (int i = 0; i < 16; i++) {
        o[i]      = __float2half(d * (float(q[i] & 0x0F) - 8.0f));
        o[i + 16] = __float2half(d * (float(q[i] >> 4) - 8.0f));
    }
}

All K-quant complexity hides inside the dequant — a q6_K 16-element sub-block sample (the block_stride parameter handles both the 7e② 224B padded and the original 210B registration layouts, something the GEMM never needs to know):

__global__ void dequant_q6_k_f16(
    const uint8_t* __restrict__ w, __half* __restrict__ out,
    int od, int id, int block_stride
) {
    int nsub = id / 16; // 16-element units, 16 per 256 super-block
    long long g = (long long)blockIdx.x * blockDim.x + threadIdx.x;
    if (g >= (long long)od * nsub) return;
    int row = (int)(g / nsub), s = (int)(g % nsub);
    int sp = s / 16, sub = s % 16;
    const uint8_t* blk = w + ((long long)row * (id / 256) + sp) * block_stride;
    float d = h2f(*reinterpret_cast<const uint16_t*>(blk + 208));
    const uint8_t* ql = blk;              // 128B low nibbles
    const uint8_t* qh = blk + 128;        // 64B high 2 bits
    const int8_t* sc = (const int8_t*)(blk + 192);  // 16 per-block scales
    // …ql/qh/sc interleaved addressing (n/tt/gq decomposition); o written at the in-row sp*256 + … offset
    float dsc = d * float(sc[sc_idx]);
    #pragma unroll
    for (int r = 0; r < 16; r++) {
        int nib = (tt < 2) ? (ql[ql_off + r] & 0x0F) : (ql[ql_off + r] >> 4);
        int q2 = (qh[qh_off + r] >> (tt * 2)) & 3;
        o[r] = __float2half(dsc * float((nib | (q2 << 4)) - 32));
    }
}

The GEMM body's skeleton (gemm_f16_nt_kernel_t, current tree; at writing time the tree already contains the P5 series' TM=128 default and the AF32 variant — the 8m② baseline is TM=64/KS=32):

template <int TM, int KS, bool AF32 = false>
__global__ void gemm_f16_nt_kernel_t(
    const __half* __restrict__ A, const __half* __restrict__ B,
    float* __restrict__ C, int nt, int od, int id
) {
    using namespace nvcuda;
    constexpr int TN = 64;
    constexpr int ODC = TM / 64;  // od 16-col fragments per warp row-half
    // …dynamic smem layout: As(2×TN×KS) + Bs(2×TM×KS) + Cs(NW×256 f32)
    const int NW = blockDim.x >> 5;   // warps: 8 for TM<=128, 16 for TM=256
    int warp = tid >> 5;
    int wm = warp >> 1;               // od chunk of this warp
    int wn = warp & 1;                // nt sub-tile: 2 x 32 rows
    // blockIdx.x = nt tile, blockIdx.y = od tile: consecutive blocks share
    // the same od-tile's B panel (64 rows x id f16, ~0.5MB) in L2, so the
    // f16 weight matrix streams from DRAM ~once instead of nt/64 times.
    int m0 = blockIdx.y * TM;
    int n0 = blockIdx.x * TN;

    wmma::fragment<wmma::matrix_a, 16, 16, 16, __half, wmma::row_major> fa[4];
    wmma::fragment<wmma::matrix_b, 16, 16, 16, __half, wmma::col_major> fb[2];
    wmma::fragment<wmma::accumulator, 16, 16, 16, float> fc[2][ODC];
    // …fill_fragment(fc, 0)

The wmma inner loop — note the v1 bug lesson carved in place in the comments:

        // fa[n-half][k-half]; fb[k-half] per od chunk. Both k halves of each
        // 32-slice must accumulate (the v1 bug: only the first 16 k's were
        // multiplied); fb's k offset is +16 ELEMENTS (one k-half), not +16
        // rows.
#pragma unroll
        for (int kh = 0; kh < KHC; kh++) {
            wmma::load_matrix_sync(fa[0], &As[buf*TN*KS + wn*32*KS + kh*32], KS);
            wmma::load_matrix_sync(fa[1], &As[buf*TN*KS + (wn*32+16)*KS + kh*32], KS);
            wmma::load_matrix_sync(fa[2], &As[buf*TN*KS + wn*32*KS + kh*32 + 16], KS);
            wmma::load_matrix_sync(fa[3], &As[buf*TN*KS + (wn*32+16)*KS + kh*32 + 16], KS);
#pragma unroll
            for (int oc = 0; oc < ODC; oc++) {
                wmma::load_matrix_sync(fb[0], &Bs[buf*TM*KS + (ob+oc*16)*KS + kh*32], KS);
                wmma::load_matrix_sync(fb[1], &Bs[buf*TM*KS + (ob+oc*16)*KS + kh*32 + 16], KS);
                wmma::mma_sync(fc[0][oc], fa[0], fb[0], fc[0][oc]);
                wmma::mma_sync(fc[1][oc], fa[1], fb[0], fc[1][oc]);
                wmma::mma_sync(fc[0][oc], fa[2], fb[1], fc[0][oc]);
                wmma::mma_sync(fc[1][oc], fa[3], fb[1], fc[1][oc]);
            }
        }
        __syncthreads();

Structure: the warp sub-block of 32 rows × TM columns decomposes into 2 n-halves (16 rows each) × 2 k-halves × ODC od chunks; fa[0..4] is loaded once and reused by all od chunks — A-fragment reuse is the first dividend the tile shape pays.

8m②'s cp.async double-buffer staging:

// 8m②: cp.async global→shared staging (sm_80+). The synchronous load
// stalled every warp on the L2 round trip each 32-k step (~31 TFLOPS
// measured); async copies overlap the k+32 tile fetch with the k compute.
__device__ __forceinline__ void gemm_cp16(__half* smem_dst, const __half* gsrc, bool full) {
    unsigned d = (unsigned)__cvta_generic_to_shared(smem_dst);
    int sz = full ? 16 : 0; // src-size 0 => zero-fill the 16B chunk
    asm volatile("cp.async.cg.shared.global [%0], [%1], 16, %2;\n" ::"r"(d),
                 "l"(gsrc), "r"(sz));
}
__device__ __forceinline__ void gemm_cp_commit() { asm volatile("cp.async.commit_group;\n"); }
__device__ __forceinline__ void gemm_cp_wait1()  { asm volatile("cp.async.wait_group 1;\n"); }

// P4: stage the A (TN rows) and B (TM rows) k-tiles [k0, k0+KS) into the
// double-buffered dynamic-smem regions. …
template <int TM, int KS, int TN, bool AF32 = false>
__device__ __forceinline__ void gemm_stage_ab(/* … */) {
    for (int c = tid; c < TN * KS / 8; c += blockDim.x) {
        int r = (c * 8) / KS, d = (c * 8) % KS;
        gemm_cp16(As + bbuf*TN*KS + r*KS + d, A + (long long)n*id + k0 + d,
                  n < nt && k0 + d < id);      // zero-fill out-of-bounds chunks (src-size 0)
    }
    for (int c = tid; c < TM * KS / 8; c += blockDim.x) { /* same for the B panel */ }
}

The rhythm in the main loop (issue async for the next tile, wait ready for the current tile, one __syncthreads aligns the whole block):

    for (int k = 0; k < id; k += KS, buf ^= 1) {
        if (k + KS < id) {
            gemm_stage_ab<TM, KS, TN>(A, B, As, Bs, Am, buf ^ 1, n0, m0, k + KS, …);
            gemm_cp_commit();
            gemm_cp_wait1();   // wait until the CURRENT tile landed (one group in flight)
        } else {
            gemm_cp_wait0();
        }
        __syncthreads();
        // …the wmma inner loop consumes buf…

Out-of-bounds handling hides in gemm_cp16's full parameter: a cp.async with src-size 0 is a hardware zero-fill — no branches writing zeros needed for any nt/od/k tail.

The launcher and the timeline of tile parameters (launch_gemm_f16; the comments double as tombstones for three later-REVERTED directions):

void launch_gemm_f16(const __half* a, const __half* b, float* c,
                     int nt, int od, int id, cudaStream_t stream, bool af32) {
    // P2: od-tile width (128 default = halved B re-reads; MINFER_GEMM_TM=64
    // reverts to the 8m② baseline for A/B).
    static int tm = -1;
    if (tm < 0) { /* MINFER_GEMM_TM: 64/128/256, default 128 */ }
    // KS = staged k-width per tile. KS=64 halves the barriers per FLOP but
    // measured -38% (56KB dynamic smem halves resident blocks on GB10);
    // KS=32 (8m2 baseline) stays the default. MINFER_GEMM_K64=1 re-tries 64.
    static int ks = -1;
    if (ks < 0) { ks = getenv("MINFER_GEMM_K64") ? 64 : 32; }
    const size_t dyn_smem = (size_t)(2*64*ks + 2*tm*ks) * 2 + 8 * 256 * 4;
    // …GEMM_LAUNCH: grid((nt+63)/64, (od+TM_-1)/TM_), 256 threads, dyn_smem

The Rust-side routing (the non-fused branch of cuda.rs's prefill_gemm_f16_inner; the current tree already contains 8p's persistent cache w16_get — in the 8m era every call went straight through get_or_grow(&self.buf_f16_w) + launch_dequant_f16):

#![allow(unused)]
fn main() {
        let w16 = match self.w16_get(wptr, type_id, od, id, block_stride) {
            Some(p) => p, // persistent copy, dequant already done (8p)
            None => {
                let w16 = Self::get_or_grow(&self.buf_f16_w, od * id * 2);
                unsafe {
                    launch_dequant_f16(type_id, wptr as *const u8, w16, …, stream);
                }
                w16
            }
        };
        // …then launch_gemm_f16(x16, w16, out, nt, od, id, stream, false)
}

3.3 Pitfalls

  • The wmma fragment addressing unit is elements, not rows. The v1 kernel multiplied only the first 16-k half of each 32-k slice; fb's second k-half offset was written as +16 rows instead of +16 elements. The fragment's interpretation of layout is completely opaque, and this class of error shows up numerically as "results systematically too small / misaligned" — the parity test catches it immediately.
  • Setting the >48 KB dynamic smem attribute silently fails during graph capture and poisons the first captured launch. The fix: gemm_prefill_smem_init() calls cudaFuncSetAttribute eagerly on all pre-registered GEMM instances at stream creation, never leaving it to a running capture window.

    #218 forward note (2026-09-29): the eager sweep described just above no longer exists. #188 deleted its CudaState::try_new call site without a note, and #218 removed the orphaned function outright rather than re-annotating it as dead. The opt-in is the lazy per-instantiation gemm_smem_optin reached through launch_gemm_f16; the invariant — the attribute is in force before a capture window opens and is never set inside one — is upheld by the 3-run capture warmup (capture_warmup), cudaStreamCaptureModeThreadLocal, and the per-instantiation cache. It is gated by cuda_prefill_smem_optin_is_done_by_production, the control arm cuda_prefill_smem_optin_refusal_fails_the_prefill (MINFER_TEST_CALL_FAIL=attr:gemm_f16_f16), the coverage arm cuda_prefill_smem_lazy_optin_admits_every_launchable_instantiation, and cuda_prefill_smem_optin_is_never_set_inside_a_capture_window. See docs/CUDA-BACKEND-DESIGN.md §2.4 and the #218 record in docs/ARCHITECTURE-EXECUTION-PLAN.md. #223 forward note (2026-09-29): the eager half is restored, but not as the sweep above. CudaState::try_new calls gemm_prefill_smem_prewarm_one once per process for every launchable instantiation — a dispatcher onto the same production gemm_smem_optin cache the launcher uses, so there is still one opt-in mechanism and one copy of gemm_dynamic_smem_bytes (the pre-#145 bug class was two copies). The placement is the guarantee: at try_new no CudaBackend — and so no capture window — can exist. The lazy path stays as defence in depth; MINFER_NO_GEMM_PREWARM=1 is the control. Measured on the real binary (GB10 sm_121, 2026-09-29): the loop is ~2.2 ms — the fatbin's one-time module load, which prewarm_prefill() already paid at the end of registration — so the net startup cost is ≈ 0 and minfer bench -p 2048 -n 128 is unchanged (+0.09% pp2048, −0.06% tg128, bar ±1%). The fifth gate is cuda_prefill_smem_prewarm_opts_in_every_launchable_instantiation_before_any_launch; see the #223 record in docs/ARCHITECTURE-EXECUTION-PLAN.md.

  • Q5_0's 22-byte blocks are naturally 2-byte aligned: a u32 load at blk+2 on an even block is cudaErrorMisalignedAddress (716). The mine was planted in 8m and only detonated in 8p's bitparity test; today's code assembles qh from two u16s, with the comment written right in the dequant kernel.
  • Templates cannot have C linkage: the extern "C" block must open and close around the templated GEMM (paired comment markers exist in the .cu), otherwise nvcc name-mangling conflicts arise.
  • cp.async chunk alignment: id % 8 == 0 guarantees a 16 B chunk never straddles a boundary (naturally covered by the outer id % 32 gate).
  • The three REVERTED boundaries that later nailed down the tile shape (KS=64 −38%, TM=256 −3%, in-kernel f32→f16 A conversion −8%) all remain in the launcher as comments — every axis of the tile was tried at a more "aggressive" value, and 64×64×32 is this kernel's local optimum on GB10.

4. Verification

  • cuda_prefill_f16_gemm_parity (Rust-side test, current tree src/graph/cuda_backend.rs:4457): for each quantization type it generates random but valid block bytes (small d/dmin so the f16 scratch never overflows), and the reference is computed on the Rust side as an f32 matmul over the dequant of the exact same bytes — it tests kernel-vs-reference parity, independent of quantization quality. Tail shapes od=70, nt=33 (id fixed at 256, since real tensors' id is always %32==0), all 8 types covered one by one, ending with a real 7B Q4_K shape check (skipped when no dump is available, keeping the suite hermetic). Defends against dequant addressing errors and fragment assembly errors.
  • E2E greedy equality: logits comparison of CUDA graph output vs CPU graph output + greedy text identical token by token (the standard gate established in 7b/7c). Defends against "each kernel is right, the assembly is wrong".
  • The bitparity test (introduced in the 8p era, feeding back into this step): the dequantized f16 is compared bit-for-bit against the CPU reference — it detonated §3.3's Q5_0 alignment mine. Defends against silent errors that happen to stay inside tolerance.
  • Per-kernel-change standalone nvcc A/B compile verification has been the convention since 7e②; both of this step's commits followed it.

5. Results

  • 8m (ba3f317): 7B @2K prefill 30.7 → 294 tok/s (9.6×) — the full dividend of replacing per-token weight re-reads with one tiled wmma GEMM.
  • 8m② (cdc6599): cp.async double buffering took the GEMM kernel from 31 → 35 TFLOPS; the whole-prefill A/B inside that commit's window was 1082 → 1204 tok/s (the climb from 294 to 1082 came mostly from the same-day 8n attention fix — see the next chapter — and the decode start fix 8o; the row number is the window value after 8m② landed).
  • The whole row (master table row 2): 30.7 → 294 → 1204 tok/s, 39×; vs llama-bench 3401 @2K ≈ 2.8×.
  • Later: 8p's load-time f16 weight cache pushed the whole row to ~1400–1500 tok/s; P5·2 (TM=128) +1.5×; R1 MMQ finally became the default at r60 — the f16 GEMM was demoted from workhorse to the MINFER_MMQ=0 escape hatch, but the entire P5-era GEMM optimization stack (TM=128, KS=32, cp.async staging) was trained on this f16 path.

6. Lessons

  1. The first-principles problem of prefill GEMM is "the weight stream crosses DRAM only once": tiling plus letting consecutive blocks share one B panel turns nt re-reads into ~1 — this single structural fix was worth 9.6×.
  2. Dequant-to-f16 is the highest-leverage shape for the starting phase: 8 quantization types isolated into 8 plain dequant kernels, one tensor-core GEMM eating all prefill; the cost (the 288 ms/call dequant pass) was later clawed back step by step via the load-time cache / dequant-in-GEMM.
  3. The wmma fragment addressing unit is elements, not rows — the +16 elements vs +16 rows mistake is the fragment API's #1 trap.
  4. Kernels needing >48 KB smem must opt in before capture; setting the attribute inside a running capture window silently fails and poisons the first captured launch.

← 01 · Index · 03 →

03 · 8n — FA-style tiled prefill attention (LANDED)

Result: 7B q4_k_m @2K prefill attention 176 → 8.5 ms/layer (20×); K traffic ~132 GB → ~0.8 GB/layer. This step broke through the entire post-8m prefill wall — the 1082 tok/s baseline in the 8m② commit window owes most of itself to this. Commit: cb66fca. Date: 2026-08-30 (same day as 8m: landed 30 minutes after 8m and 20 minutes before 8m②).

1. Background — where things stood

Once 8m swapped the prefill GEMM for tiled wmma, the wall's composition flipped instantly: attention became the overwhelming majority. The old kernel gqa_attn_f32_f16kv measured 176 ms/layer at 7B @2K — 76% of the entire 2K prefill wall (28 layers × 176 ms ≈ 4.9 s, against a whole wall of ~7.0 s at 294 tok/s). The GEMM was already running at 31 TFLOPS; continuing to optimize it would have been the wrong next target — the 76% had to fall first.

The old kernel's disease is the same one decode attention later cured (R4): one block per (token, head). One block per q-head per token, and that block must read K in full — every byte of K is re-read nt × (q heads per kv head) times, ~132 GB/layer at 7B @2K, all burned on the L2/DRAM transport side. Meanwhile the hd-wide output accumulator float4 oc[32] occupies 128 registers and triggers spills — the same disease as the "LOCAL-memory accumulator = ~80 MB/layer of local traffic" entry in R4's later decode table.

The fix held no suspense: FlashAttention had already established "tile the q dimension + online softmax" as the standard shape. minfer's KV cache has stored f16 since the 7e series, so Q/K/V can feed tensor cores directly for QK^T. 8n's job was to land that shape on CUDA: one block per (64-token q tile, head), K/V entering shared memory tile by tile, S = Q·K^T on wmma, softmax in online form, the O accumulator resident in registers.

2. Principle — the GPU mechanism

2.1 Online softmax

Naive attention must finish computing the entire row of scores before softmax (it needs the full-row max and full-row sum), which means the S matrix materializes at full size. Online softmax makes it incremental: KV is processed tile by tile, and each row keeps three pieces of state — running max m, running sum l, output accumulator O. When processing tile k:

  1. Find the new max within the tile: m_new = max(m_old, max(S_tile));
  2. Rescale factor alpha = exp(m_old − m_new); multiply the old O by alpha and the old l by alpha;
  3. In-tile probabilities p = exp(S − m_new), accumulated into O and l.

After all tiles are processed, O / l is the correct softmax-weighted output. Mathematically an identity transformation; the price is one O rescale per tile — in exchange, O and S both need only tile-sized storage and K/V is consumed as a stream.

2.2 The byte ledger: 132 GB → 0.8 GB

Old shape: every (token, head) block reads all of K/V → ~132 GB per layer. For scale: reading the full 2048×2048 K matrix (hd=128, f16) once is 2048×2048×128×2 B ≈ 1 GB — the old kernel's re-read factor is exactly the nt × heads order of magnitude, with V doubling it again. Tiled shape: a block's 64 tokens share the same K/V tiles (reused 64 times once in shared memory), so K/V's effective read volume is diluted by the q-tile width; with blocks for the same kv-head hitting L2 against each other, the measured figure lands at ~0.8 GB/layer. That 20× transport difference is the main source of 176 → 8.5 ms — at this scale attention is a bandwidth war just like the GEMM.

2.3 Register O and the thread geometry

Inside a block, 256 threads = 64 rows × 4 quadrants: each thread owns the f32 accumulator acc[32] for one row's 32 dims (a quarter of hd=128). Keeping O in registers pays twice: the alpha rescale is a pure register operation (no reading O back from shared memory, scaling, writing back), and every V read in P·V is an intra-warp broadcast. The code comment, verbatim: Keeping O in registers (instead of shared) makes every P·V V-read a warp-wide broadcast and the alpha reads conflict-free.

QK^T uses wmma: 8 warps split the 64×64 S matrix by wm = warp>>1 (4 q 16-blocks) × wk = warp&1 (2 kv 32-blocks), and each warp runs mma_sync over the hd loop (f16 inputs, f32 accumulation). P·V is still a scalar FMA loop in this step (P stored f16, V f16, acc f32) — moving P·V onto tensor cores was a separate later step, P5·0 (10.06 → 4.24 ms/layer).

2.4 f16 probs and the 256 B stride

The post-softmax probabilities P are stored f16, directly aliasing the score matrix Sf's memory — scores are dead data once softmax has consumed them. This alias is where a 64×64 buffer is saved from the 97 KB smem budget, and it is also this step's only correctness mine (see §3.3): P's row stride must be 256 B (FA_PSTR = FA_TKV*2 halves), so row r's probabilities overlap only the first half of row r's own scores.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Grid (nt/64 q tiles, heads), 256 threads, dynamic smem ~65 KB (actual allocation 66,304 B: Qs/Ks/Vs at 64×128×2 B = 16 KB each, Sf 64×64×4 B = 16 KB, m/l/alpha 768 B). Over the static limit, so it follows 8m's procedure with the cudaFuncSetAttribute opt-in. (The layout table in the kernel's header comment says "~97 KB" and lists a row O [64*hd] f32 — a leftover from the draft period when O still lived in shared memory: the actual code moved O into registers and the launcher's allocation expression has no O. Small drift of this kind between comment and code is itself worth recording.)
  • The Q tile lands as f16 with the scale folded into Q: q * scale is computed once at load, so QK^T's inner loop no longer multiplies by scale every step — what is saved is scalar work inside the wmma loop; the price is that Q's f16 rounding happens early (numerically absorbed into the 5e-3 tolerance gate).
  • S/P alias: saves a 64×64 buffer; correctness is secured by the program-order argument for the 256 B stride (§3.3).
  • GQA slicing: q head h reads only the K/V slice of kv head h/gqa (stride_kv = nk*hd), not all of K — 7B is 28:4, and each kv head's K/V is shared among its 7 q heads' tile blocks.
  • kv_end = positions[last_t] + 1: KV positions are data, not structure (graph rule §1) — the kernel reads the position of the q tile's last token from the positions input, and every KV row beyond that bound is skipped.
  • Gated on hd == 128 (the Qwen2.5/2.5-7B shape); other head dims take the old path.
  • Masked positions are handled inside the data flow: the score comparison kv_g <= qpos; a fully-masked tile writes 0 probabilities and keeps the softmax state untouched.

3.2 Key code

All excerpts below come from the original fa_prefill_f16kv introduced by cb66fca (in the current tree this kernel has since evolved through the P5·0/P5·3/FAP2 series into the wmma P·V + register-softmax version, but this chapter's tile geometry, online-softmax state machine, and race argument survive unchanged to this day).

The smem layout and tile constants:

// Shared layout (dynamic, ~97 KB — opt-in via cudaFuncSetAttribute):
//   Qs [64*hd] f16   q tile (scale folded in, f16 for the tensor-core QK^T)
//   Ks [64*hd] f16   K tile           Vs [64*hd] f16  V tile
//   S  [64*64]       f32 scores, aliased as f16 probs after the row softmax
//   m/l/alpha [64] f32 per-row online-softmax state
#define FA_TQ 64
#define FA_TKV 64
#define FA_PSTR (FA_TKV * 2) // probs row stride in halves (256B): probs row r
                             // aliases only Sf row r's first half, already read
                             // by the same thread — no cross-thread race
#define FA_HQ 32 // hd/4 dims per accumulator thread (kernel is gated to hd == 128)

__global__ void fa_prefill_f16kv(
    const float* __restrict__ q, const __half* __restrict__ k,
    const __half* __restrict__ v, float* __restrict__ o,
    const int* __restrict__ positions,
    int nh, int nk, int hd, float scale, int nt
) {
    extern __shared__ __align__(256) uint8_t smem[];
    __half* Qs = reinterpret_cast<__half*>(smem);
    __half* Ks = Qs + FA_TQ * hd;
    __half* Vs = Ks + FA_TKV * hd;
    float* Sf = reinterpret_cast<float*>(Vs + FA_TKV * hd);
    __half* Pf = reinterpret_cast<__half*>(Sf); // alias: probs after softmax
    float* msh = reinterpret_cast<float*>(Sf + FA_TQ * FA_TKV);
    float* lsh = msh + FA_TQ;
    float* alpha = lsh + FA_TQ;

Q tile load (scale folded in + tail zeroing) and K/V tile staging (16B/lane, out-of-bounds zero-fill):

    // load q tile (scale folded in) as f16
    for (int i = tid; i < FA_TQ * hd; i += 256) {
        int r = i / hd, d = i % hd;
        int t = tq0 + r;
        float qv = (t < nt) ? q[(size_t)t * ne_q + h * hd + d] * scale : 0.0f;
        Qs[i] = __float2half(qv);
    }
    if (tid < FA_TQ) {           // per-row online-softmax state init
        msh[tid] = -INFINITY;
        lsh[tid] = 0.0f;
    }
    …
    // stage K/V tile (16B per lane; rows beyond kv_end zero-filled)
    const uint4 z4 = make_uint4(0, 0, 0, 0);
    for (int i = tid * 8; i < FA_TKV * hd; i += 2048) {
        int r = i / hd, d = i % hd;
        int p = kt + r;
        if (p < kv_end) {
            kk4 = *reinterpret_cast<const uint4*>(&k[(size_t)p * stride_kv + hk * hd + d]);
            vv4 = *reinterpret_cast<const uint4*>(&v[(size_t)p * stride_kv + hk * hd + d]);
        } else {
            kk4 = z4;  vv4 = z4;   // zero-fill out-of-bounds KV rows → S=0, backstopped again by the mask logic
        }
        *reinterpret_cast<uint4*>(&Ks[i]) = kk4;
        *reinterpret_cast<uint4*>(&Vs[i]) = vv4;
    }
    __syncthreads();

Register O and thread ownership:

    // Per-thread output accumulator: thread owns (row, quadrant) with
    // row = tid & 63, quadrant = tid >> 6 (FA_HQ dims each). Keeping O in
    // registers (instead of shared) makes every P·V V-read a warp-wide
    // broadcast and the alpha reads conflict-free.
    float acc[FA_HQ];
#pragma unroll
    for (int dd = 0; dd < FA_HQ; dd++) acc[dd] = 0.0f;
    __syncthreads();

    const int last_t = min(nt - 1, tq0 + FA_TQ - 1);
    const int kv_end = positions[last_t] + 1;   // KV positions are data
    const int arow = tid & (FA_TQ - 1);
    const int aquad = tid >> 6; // 0..3

QK^T on the tensor core (each warp computes one 16×32 sub-block of S):

        using namespace nvcuda;
        int warp = tid >> 5;       // 0..7
        int wm = warp >> 1;        // q 16-block: 4
        int wk = warp & 1;         // kv 32-block: 2
        wmma::fragment<wmma::matrix_a, 16, 16, 16, __half, wmma::row_major> fa;
        wmma::fragment<wmma::matrix_b, 16, 16, 16, __half, wmma::col_major> fb[2];
        wmma::fragment<wmma::accumulator, 16, 16, 16, float> fc[2];
        for (int d = 0; d < hd; d += 16) {
            wmma::load_matrix_sync(fa, &Qs[wm * 16 * hd + d], hd);
            wmma::load_matrix_sync(fb[0], &Ks[wk * 32 * hd + d], hd);
            wmma::load_matrix_sync(fb[1], &Ks[(wk * 32 + 16) * hd + d], hd);
            wmma::mma_sync(fc[0], fa, fb[0], fc[0]);
            wmma::mma_sync(fc[1], fa, fb[1], fc[1]);
        }
        wmma::store_matrix_sync(&Sf[wm*16*FA_TKV + wk*32],      fc[0], FA_TKV, wmma::mem_row_major);
        wmma::store_matrix_sync(&Sf[wm*16*FA_TKV + wk*32 + 16], fc[1], FA_TKV, wmma::mem_row_major);

The online softmax in full — including the verbatim race comment (this chapter's lesson carrier):

        // online softmax per row (thread = row): probs land in Pf (f16)
        if (tid < FA_TQ) {
            int r = tid;
            int qpos = (tq0 + r < nt) ? positions[tq0 + r] : -1;
            float m_old = msh[r], m_new = m_old;
            for (int kk = 0; kk < FA_TKV; kk++) {
                int kv_g = kt + kk;
                if (kv_g <= qpos && kv_g < kv_end) {
                    float s = Sf[r * FA_TKV + kk];
                    if (s > m_new) m_new = s;
                }
            }
            float a = 1.0f;
            // Pf rows use a 256B stride: row r's probs overlap ONLY Sf row r's
            // first half, which this same thread has already read (each read
            // precedes its clobbering write in program order). A 128B stride
            // would race: probs for row r land on scores of rows 2r/2r+1 that
            // other softmax threads have not read yet.
            if (m_new == -INFINITY) {
                // nothing valid in this tile: keep state, zero probs
                for (int kk = 0; kk < FA_TKV; kk++) Pf[r * FA_PSTR + kk] = __float2half(0.0f);
            } else {
                a = (m_old == -INFINITY) ? 0.0f : __expf(m_old - m_new);
                float sum = 0.0f;
                for (int kk = 0; kk < FA_TKV; kk++) {
                    int kv_g = kt + kk;
                    float p = 0.0f;
                    if (kv_g <= qpos && kv_g < kv_end)
                        p = __expf(Sf[r * FA_TKV + kk] - m_new);
                    Pf[r * FA_PSTR + kk] = __float2half(p);
                    sum += p;
                }
                lsh[r] = lsh[r] * a + sum;
            }
            alpha[r] = a;  msh[r] = m_new;
        }
        __syncthreads();

Rescale + P·V + epilogue:

        // rescale the accumulator by alpha, then add P · V
        float ar = alpha[arow];
#pragma unroll
        for (int dd = 0; dd < FA_HQ; dd++) acc[dd] *= ar;
        for (int kk = 0; kk < FA_TKV; kk++) {
            float p = __half2float(Pf[arow * FA_PSTR + kk]);
            if (p != 0.0f) {
                const __half* vrow = &Vs[kk * hd + aquad * FA_HQ];
#pragma unroll
                for (int dd = 0; dd < FA_HQ; dd++) acc[dd] += p * __half2float(vrow[dd]);
            }
        }
        __syncthreads();
    }

    // write out: acc / l — rows with l == 0 stay 0 (fully masked)
    if (tq0 + arow < nt) {
        float inv = (lsh[arow] > 0.0f) ? 1.0f / lsh[arow] : 0.0f;
        float* orow = &o[(size_t)(tq0 + arow) * ne_q + h * hd + aquad * FA_HQ];
#pragma unroll
        for (int dd = 0; dd < FA_HQ; dd++) orow[dd] = acc[dd] * inv;
    }

3.3 Pitfalls

  • The score-clobber race (this step's #1 pitfall). After Pf aliases Sf, P's writes run concurrently with reads of not-yet-consumed S. Under the naive 128 B stride, row r's probabilities land on the first half of Sf rows 2r/2r+1 — rows belonging to other softmax threads whose reads have not happened yet in program order: a race. The 256 B stride (FA_PSTR = FA_TKV*2) confines row r's probabilities to the first half of row r's own scores, which the same thread has already finished reading — every read precedes its clobbering write in the same thread's program order, so no race exists. This bug was not found by the graph parity test: the end-to-end logits comparison can pass under lucky scheduling; it was the standalone harness (a test rig that drives the kernel repeatedly, independent of graph execution, and compares outputs) that exposed the cross-thread race class.
  • The numeric path of a fully-masked tile: when m_new == -INFINITY (no valid KV position in the tile) the old state must be kept and probabilities written as 0 — taking the normal branch would let __expf(-INF − -INF) = NaN propagate down through l/O.
  • The 97 KB dynamic smem opt-in: when the attribute set fails, call cudaGetLastError() first to clear the error before returning — leaving it set would poison the subsequent stream (the same lesson as 8m's capture poisoning).
  • P's f16 rounding: S stays f32 throughout and the softmax output converts to f16 — this is where the 5e-3 tolerance gate comes from (P·V reads f16 back), and it is also the numerical contract that had to be preserved when P5·0 later moved P·V onto tensor cores.

4. Verification

  • The fa_prefill_f16kv parity test (src/graph/cuda_backend.rs:4306): seeded pseudo-random q/k/v, run through the real graph nodes (kvcache_store lands the KV, then the attention node executes), reference = cpu_gqa_attn computed on the same f16-rounded K/V, gate assert_close(..., 5e-3). Defends against: kernel numeric errors, mask errors, GQA slicing errors.
  • The standalone harness: a driver independent of graph execution that runs the kernel repeatedly and compares outputs. Defends against: scheduling-dependent cross-thread races — the class the graph parity test cannot catch (this step's race is exactly what it caught).
  • E2E greedy equality + whole-prefill timing: defends against assembly errors and confirms the wall-clock gain.
  • Masked / n_past-grown positions: the parity test's positions sequence covers non-zero starting points. Defends against: KV position handling errors (the data-fied positions of graph rule §1).

5. Results

  • Kernel level: 176 → 8.5 ms/layer (7B @2K, 20×); K traffic ~132 GB → ~0.8 GB/layer.
  • Wall-clock level: attention's share of the wall fell from 76% to single digits. Cross-check: at 8m's landing, 294 tok/s ⇒ whole wall 6.97 s, of which attention 4.93 s; after the fix the wall ≈ 2.04 + 0.24 ≈ 2.3 s ⇒ an estimate of ~900 tok/s. The 8m② window's measured baseline was 1082, and after cp.async 1204 tok/s (against llama-bench 3401 @2K, ~2.8×) — the gap between estimate and measurement belongs to machine state and small same-window fixes; the orders of magnitude agree.
  • Later evolution (each has its own chapter; not expanded here): P5·0 moved P·V onto tensor cores (10.06 → 4.24 ms/layer); P5·3 added padding to kill ldmatrix bank conflicts; r48 (FAP2) moved softmax wholesale into registers (5.16 → 2.12 ms/layer); r50/r57 nailed down the boundaries for changing FA tile sizes. The skeleton 8n built — the 64-token q tile, the online-softmax state machine, GQA kv slicing, register O — all continues in today's FA kernel.

6. Lessons

  1. Prefill attention's root disease is the same as decode's: per-(token,head) K/V re-reads — tiling the q dimension is the only correct cure, and the gain comes from the byte ledger (÷64 reuse + L2), not from smarter FLOPs.
  2. Every byte saved by aliasing needs a program-order argument attached: the P/S alias's correctness depends on the stride making each "clobber" land only on data the same thread has already read; 128 B's "save half" is a race.
  3. The standalone harness and graph parity are complementary gates: a scheduling-dependent race can hide behind lucky scheduling in end-to-end comparisons; only controlled repeated runs expose it reliably.
  4. Moving the O accumulator into registers was a two-step walk: 8n first made O register-resident (the alpha rescale with zero smem round trips), while softmax itself only entered registers at FAP2 — change one data path at a time, each step measurable.

← 02 · Index · 04 →

04 · 8o — Killing the CPU stall at decode start (LANDED)

Result: first decode step 724 → 35 ms (~20×); steady-state decode rate unchanged. Commit: 65b686c ("perf(graph): 8o — kill ~1.6 s one-time decode-start CPU stalls (4.4 GB weight re-clone + 1.9 GB concat probe per rebuild)"). Date: 2026-08-30.

1. Background — where things stood

Phase 7 (row 1) had just raised the CUDA backend: resident weight registry, per-op dispatch, CUDA Graph capture/replay. 8m/8m② (row 2) and 8n (row 3) had pushed prefill from 30.7 tok/s to 1204 tok/s and attention from 176 ms/layer down to 8.5 ms/layer. The prefill line finally looked respectable — naturally, decode was next.

But the moment decode was touched, it showed: the first decode step waits 724 ms before emitting the first token, while steady-state decode steps are far faster. This is not a kernel problem — in nsys the GPU has no kernel queued during that window — it is the host side doing two things at graph switch that it should never have done.

Background mechanism: minfer's inference runs through a "declarative compute graph" — built deterministically from GraphParams (including n_tokens); different parameters mean a rebuild and a new cached graph. prefill (nt = prompt length) and decode (nt == 1) differ in parameters, so the prefill→decode switch necessarily triggers one decode graph rebuild. That rebuild should have been purely structural (node list + buffer planning), but at the time it:

  1. Re-registered every weight on the CPU backend. The model-side call sites did t.clone() per weight before handing it to register_weight, and Tensor's byte payload is a Cow::Owned — clone() is a full deep copy. All of 7B q4_k_m's matmul weights are ~4.4 GB, memcpy'd on pure CPU at every graph (re)build, measured as a ~635 ms pure-CPU stall (no CUDA calls, no kernels — just memcpy).
  2. Answered a feasibility question with a "build it and throw it away" probe. When building the decode graph, the engine must decide whether the 28 ffn gate/up weight pairs can be concatenated and registered as blk.{i}.ffn_gu (the precondition of the G5 FFN fusion); the CUDA branch of the time called the eager concat_rows — to obtain a can-it-be-done answer of type Option<Vec<u8>>, it actually re-concatenated ~1.9 GB (7B's 28 gate/up pairs), then used only the metadata and dropped the bytes outright. Measured ~920 ms.

Together, ~1.6 s of one-time startup stall (the commit title's framing). For interactive use, this 1.6 s lands on a user who has already finished waiting for prefill, and it reads as "the first word is especially slow"; for bench it pollutes the first-step timing. 8o's job was to turn both segments into zero cost.

2. Principle — the GPU mechanism

This step's "mechanism" is not on the GPU but in the host↔GPU pipeline relationship: during the encode phase the GPU is idle. CUDA Graph capture only happens after the graph rebuild, weight resolution, and launch arguments are all ready; while the host rebuilds the graph the GPU has nothing to do. So every 1 ms of needless host work becomes first-step latency 1:1.

The byte arithmetic of the two stalls (7B q4_k_m, matmul weights ~4.4 GB):

  • The weight deep copy: a Vec<u8> memcpy is one read + one write ≈ 8.8 GB of DRAM traffic. Single-threaded memcpy runs at ~7 GB/s, so a 4.4 GB copy ≈ 630 ms — matching the measured ~635 ms. A pure host memory-bandwidth problem, completely unrelated to the GPU.
  • The eager concat probe: a freshly allocated ~1.9 GB output buffer (page-faulted on first touch) + ~1.9 GB read of inputs + ~1.9 GB write of output ≈ 3.8 GB of traffic, measured ~920 ms. Yet the question it answers — "are these tensors' types, block structures, and row byte counts compatible, and can they be concatenated row-wise" — requires only reading each tensor's ttype and shape, a few hundred bytes in total.

The key insight: the Clone semantics of Cow<'static, [u8]>. The #[derive(Clone)] implementation for Cow is: cloning Borrowed(b) copies only the reference (O(1)), while cloning Owned(v) deep-copies the whole Vec (O(n)). minfer's weight tensors took the Owned path at the time, so a casual t.clone() in the model code was a GB-scale memcpy on GB-scale objects like weights. The type system raises no error, the tests don't fail — it is just slow.

Equally key is the idempotence argument: weights are immutable after load (reuse consistency is guarded by GraphParams.weights_version; any future weight change bumps the version first and forces a rebuild), so a duplicate same-name registration necessarily carries identical data — skipping it is safe. That is what makes the contains_key early-exit lock legitimate.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Do not eliminate the rebuild — only eliminate the duplicated labor inside it. That different prefill/decode parameters trigger a rebuild is orthogonal to the reuse design (the same invariant as llama.cpp's allow_reuse); touching it is risky, and "rebuild = structural work" is the right shape. So the fix turns the two big pieces of work inside the rebuild into no-ops, rather than adding a rebuild cache.
  • Idempotent registration rather than making clone cheap by reference. The alternative was switching Tensor's weights to Borrowed (mmap zero-copy) so clone becomes cheap — but that touches the loader's ownership structure, a wide blast radius. The early-exit lock lands in one line: the first registration proceeds as before, and later same-name registrations skip.
  • The feasibility probe and the real concatenation must share "the same precondition". The real concatenation happens once, in the loader (concatenated and registered as blk.{i}.ffn_gu at load); graph building merely asks "can it be concatenated". The two paths' decisions must match branch by branch, or you get the silent mismatch of "the probe says yes, the loader concatenated a different layout". So the new function mirrors concat_rows' precondition checks item by item, and this is written into the doc comment as a maintenance contract.
  • No speculative generalization: the probe is simply concat_rows_feasible — no traits, no callbacks; the two precondition lists exist explicitly side by side, locked to each other by comments. Small, visible duplication drifts less than an abstraction.

3.2 Key code

First site: the registration idempotence lock in src/graph/cpu_backend.rs (10 lines added):

#![allow(unused)]
fn main() {
/// Register a weight tensor by name (Phase 6 wires this from the model).
pub fn register_weight(&mut self, name: &str, t: Tensor) {
    // Skip re-registration of an already-known weight: Tensor carries its
    // bytes as Cow::Owned, so the `t.clone()` at the model call sites
    // deep-copies the full weight set (~4.4 GB on 7B) on EVERY graph
    // (re)build — measured as a ~635 ms pure-CPU stall at the
    // prefill→decode graph switch (no CUDA calls, no kernels). Model
    // weights are immutable after load (weights_version guards any future
    // change), so a same-name registration always carries the same data.
    if self.weights.contains_key(name) {
        return;
    }
    self.weights.insert(name.to_string(), t);
}
}

The Tensor definition (current tree src/tensor.rs) explains why the call site's t.clone() is so expensive:

#![allow(unused)]
fn main() {
#[derive(Clone)]
pub struct Tensor {
    pub ttype: TensorType,
    pub shape: [i64; 4],      // ne[0..3]: number of elements per dimension
    pub strides: [usize; 4],  // nb[0..3]: stride in bytes per dimension
    pub data: std::borrow::Cow<'static, [u8]>,  // ← when Owned, clone() = full memcpy
    pub name: String,
}
}

Second site: the metadata-only feasibility probe added to src/cuda.rs (the core of the 34 added lines):

#![allow(unused)]
fn main() {
pub fn concat_rows_feasible(tensors: &[&Tensor]) -> bool {
    if tensors.len() < 2 {
        return false;
    }
    let tt = tensors[0].ttype;
    if tensors.iter().any(|t| t.ttype != tt) {
        return false;
    }
    let bq = quant_block_q(tt);          // elements per block
    let bb = quant_block_bytes(tt);      // bytes per block
    if bb == 0 {
        return false;
    }
    let ne0 = tensors[0].shape[0] as usize;
    if tensors.iter().any(|t| t.shape[0] != ne0 as i64) {
        return false;
    }
    if ne0 % bq != 0 {
        return false;
    }
    let row = (ne0 / bq) * bb;           // compressed bytes per row
    tensors
        .iter()
        .all(|t| t.data().len() == row * (t.shape[1] as usize))
}
}

Compared clause by clause against the eager concat_rows' preconditions: ≥2 tensors, identical quantization type, identical ne0 divisible by the block size, and each tensor's actual byte length equal to rows × row bytes. The final data-length check subsumes the eager version's trailing out.len() != rows * row defense — so both functions give the same can-it-be-done answer for the same input.

Third site: the call site in src/models/qwen2/graph.rs, before → after:

#![allow(unused)]
fn main() {
// before — concatenate 1.9 GB for one bool and throw it away:
crate::cuda::concat_rows(&[fg, fu]).is_some()
// after — reads metadata only:
crate::cuda::concat_rows_feasible(&[fg, fu])
}

3.3 Pitfalls

  • The pitfall is language semantics itself: Cow's derived Clone is an O(n) deep copy on the Owned variant. Whoever writes t.clone() is thinking "copy a handle" and gets a 4.4 GB memcpy. No compile-time or runtime signal whatsoever — only a timer can catch it.
  • The probe had side effects: concat_rows(&[fg, fu]).is_some() looks perfectly harmless but actually allocates a GB-scale buffer. The lesson: when you only want a bool, do not call a function that returns data — no matter how convenient.
  • The sync risk of two precondition lists: concat_rows_feasible and concat_rows are two hand-written checks, and the comment states explicitly "must mirror concat_rows' preconditions; the data-length check covers the trailing total-length defense". This is a deliberate trade — a comment lock instead of an abstraction merge; the price is that any future change to concat_rows must be mirrored here.

4. Verification

This step is a startup-path/host-side change that touches no numeric path, and the source record sets up no dedicated bitwise-dump gate for it. Its gates are behavioral-equivalence gates:

  • Steady-state decode rates unchanged (row 4 records "all decode rates unchanged") — defends against the "the work we skipped was actually needed" class of regression: if the skipped registration or probe had had a needed side effect, the decode step itself would change.
  • The equivalence argument: the registration early-exit reuses the very Tensor registered by the same load (ownership unchanged, no copies); the probe changes only "whether bytes get constructed", never the returned verdict (the branch-by-branch mirror of §3.2).
  • The end-to-end metric "first decode step 724 → 35 ms" is itself the most direct verification: it is the observable surface of this very bug.

5. Results

  • End to end: first decode step 724 → 35 ms (~20×), steady-state decode rates unchanged (row 4's Δ records 20×).
  • Component level (segmented timing): weight re-clone ~635 ms → 0; eager concat probe ~920 ms → 0; the combined ~1.6 s one-time startup stall disappears (the commit title's framing). A note on framing: the 635/920 ms figures are re-assembly numbers from segment-timing the rebuild path, while 724 → 35 ms is the end-to-end metric over the "first decode step" window — the two are not additive over the same measurement window, and row 4's conclusion defers to the latter.
  • No new kernels introduced, no numeric path changed; the comparison target is the engine's own baseline (before/after, same machine, same model).

6. Lessons

  1. Cow::Owned turns clone() into a deep copy — every clone() on a startup path deserves an audit in bytes (row 4's own words: audit clones on startup paths).
  2. Never answer a can-it-be-done question with a "build the result and throw it away" function — a feasibility check should read metadata only.
  3. Weights are immutable after load (weights_version guards changes), so registration can be made idempotent; idempotent registration turns the rebuild's duplicate registrations into no-ops wholesale.
  4. The end-to-end "first decode step" time is the sentinel metric for startup-path regressions — when it moves, always check the host side first, not the kernels.

← 03 · Index · 05 →

05 · 8p — Persistent f16 weight cache + fused dequant-in-GEMM (LANDED)

Result: 7B q4_k_m @2K prefill 1201 → ~1400–1500 tok/s (+17–24%); decode unchanged (40.8 vs 40.6 A/B). The price is 2 B/element (~8.6 GB on 7B). Commit: 2992f57 (docs record b9e7a91). Date: 2026-08-31.

1. Background — where things stood

8m/8m② (row 2) unified the prefill GEMM into a 64×64 wmma f16 tensor-core kernel, taking 7B @2K from 30.7 → 294 → 1204 tok/s. But this path carries a hidden tax: wmma consumes an f16 B operand, while the GGUF weights are quantized (q4_K 0.5 B/weight, q8_0 1 B/weight, …). So 8m's two-pass GEMM re-dequantizes W into an f16 scratch on every matmul call, and the GEMM then reads that scratch.

How expensive is "every call"? Measured at 7B q4_k_m @2K: 288 ms/call — dequantizing the 4.4 GB of quantized weights (read 4.4 GB + write ~8.8 GB f16) plus scheduling for all 28 layers × 7 matrices. The 1204 tok/s prefill forward itself is no longer slow, but every forward first pays this fixed 288 ms tax. It was also one of the main components of the then ~2.8× gap to llama.cpp on prefill.

Dequantization has an obvious asymmetry: weights are immutable after load (reuse consistency guarded by weights_version), yet the two-pass GEMM re-derives their values on every forward. The only defense of "dequant per call" is saving VRAM — holding an f16 copy on 7B costs +2 B/element ≈ +8.6 GB. 8p was built around this time-vs-memory triangle:

  1. Persistent f16 cache (the default path): each weight is dequantized once at load, the f16 copy stays resident in VRAM, and the GEMM reads it directly — the per-call 288 ms goes to zero.
  2. Fused dequant-in-GEMM (the alternate path, MINFER_FUSED_B=1): no cache is built; inside the GEMM, B tiles are dequantized in registers from the raw quantized bytes — zero extra VRAM, but slower at large nt (§2), kept as the fallback for memory-constrained settings.
  3. A size gate (W16_ENABLE_BYTES = 2 GB): the cache is enabled only when total quantized matmul weight bytes ≥ 2 GB (7B q4_k_m's 4.4 GB passes; the 0.5–1.5B test fixtures do not and keep their pre-8p footprint — the reason is the suite OOM pitfall in §3.3).

This combination still serves today: the "f16 w16-cache" under §1.1's MINFER_MMQ=0 legacy f16 escape path (~2353 tok/s, ~20.5 GB — about 11 GB heavier than the MMQ default path) is exactly the skeleton this step built.

2. Principle — the GPU mechanism

Let W be quantized weights (q4_K-class, 0.5 B/element), id×od elements, f16 copy 2 B/element. The relative per-forward costs of the three shapes:

  • Dequant per call (pre-8p): the dequant pass reads 4.4 GB of quantized bytes + writes 8.8 GB of f16 scratch, which the GEMM then reads — the dequant pass alone is one full-weight DRAM sweep, 288 ms/call. The scratch's write traffic (8.8 GB) is generated for nothing, and every GEMM B-panel re-read hits this scratch.
  • Persistent f16 cache (8p default): the dequant's read+write each happen once (at load); from then on every forward's B panel reads only the f16 copy. The full 288 ms per call is saved, in exchange for +8.6 GB resident. The GEMM's byte footprint is unchanged (the same f16 as before); what is saved is the dequant ALU plus the wholesale in/out of the intermediate scratch.
  • Fused dequant-in-GEMM (MINFER_FUSED_B=1): B tiles are staged into smem as quantized bytes (q4_K 0.5 B/element, a quarter of f16) and dequantized to f16 in registers before entering wmma. The cheapest in bytes and zero VRAM delta, but the dequant ALU becomes part of the kernel's inner loop. What matters is the re-read count: with 64×64 tiles, the same B panel is re-read by every m0 (a tile along nt) — at nt=2048 that is 32 passes, i.e. 32 rounds of dequant ALU in the fused shape; the f16 cache shape's "re-read" is only a byte re-read (with L2 backing it), no ALU. This is why the commit measured fused as slower than the cp.async f16 GEMM at large nt ("every nt tile re-dequantizes the B panel"), while in the nt==1 decode context fused is the one with the byte advantage (decode later went exactly the quantized-byte-stream + MMVQ/MMQ route — see 06).

Why the cache key can be a device pointer: register_weight reuses the same device copy for same-name same-size registrations and never frees on replace — so a registered weight's device pointer is stable for the process lifetime and the keys of w16_cache: HashMap<usize /*wptr*/, (CudaPtr, usize)> never dangle. Another cash-in of the "weights are immutable" invariant (04's registration idempotence lock leans on the same invariant).

The arithmetic of the alignment problem (Q5_0's latent crash, the protagonist of §3.3): a Q5_0 block is 22 bytes (2 B f16 d + 4 B qh bit plane + 16 B nibble qs). Block g sits at base + 22g; the u32 load of qh at blk+2 is 4-byte aligned iff 22g + 2 ≡ 0 (mod 4), i.e. g is odd. Every even block is misaligned — CUDA's scalar loads require natural alignment, violation is cudaErrorMisalignedAddress (error code 716) and the kernel fails outright. Q5_1's block is 24 bytes and 24g + 4 ≡ 0 (mod 4) always holds, so q5_1's u32 loads were always safe — alignment is a function of "stride × g + offset", not of provenance.

3. Implementation

3.1 Design choices (why this shape and not another)

  • The cache is the default and fused is the escape hatch, not the reverse. The target scenario (7B @2K prefill) is large nt, long-running inference — the dequant cost amortizes over countless forwards and one-time materialization wins outright; the VRAM price is acceptable on GB10's unified memory pool. The fused kernel's value is in memory-constrained settings, kept as the "zero VRAM delta" alternative.
  • The gate lives in the loader's warm pass (W16_ENABLE_BYTES = 2 << 30), not in runtime adaptation: the total bytes of quantized matmul weights are known the moment the model loads, making the check O(1); and "should we occupy 8.6 GB more" is fundamentally a model-level decision that should not fluctuate with load.
  • Dequantization happens at load (warm_w16), not lazily at the first GEMM: the load path already registers weight by weight, so warming in passing lets the very first forward run at full speed; the price is one extra (one-time) segment of load time.
  • Two env vars, one direction each: MINFER_NO_W16CACHE=1 turns the cache off and falls back to per-call scratch; MINFER_FUSED_B=1 turns fused on. Both A/B and regression have valves.
  • The guard added from R1 on: the int8 MMQ prefill GEMM (R1, later made default via r60) streams raw quantized bytes directly, so the f16 cache is dead weight while it is active — the loader's warm condition gains !cuda.mmq_active() (visible in the current tree), and MMQ models no longer pay 8.6 GB for nothing.

3.2 Key code

First site: the loader-side warm gate (src/models/qwen2/loader.rs; the Qwen3 loader carries the same 50-line change):

#![allow(unused)]
fn main() {
let warm_bytes: usize = warm
    .iter()
    .filter_map(|(t, _)| t.as_ref())
    .map(|t| t.data.len())              // total bytes of quantized matmul weights
    .sum();
// R1: with the int8 MMQ prefill GEMM active the f16 cache would be
// dead weight (MMQ streams raw quantized bytes) — skip the warm pass.
let warmable = warm_bytes >= crate::cuda::W16_ENABLE_BYTES && !cuda.mmq_active();
if warmable {
    cuda.enable_w16_cache();
}
for (t, name) in &warm {
    if warmable {
        if let Some(t) = t {
            cuda.warm_w16(name, t);     // dequantize once at load, f16 stays resident
        }
    }
}
}

The constant comment on W16_ENABLE_BYTES states the 7B/fixture boundary outright (src/cuda.rs):

#![allow(unused)]
fn main() {
/// 8p: warm the f16 weight cache only for models whose quantized matmul
/// weights total at least this much (7B q4_k_m = 4.4 GB warms; the 0.5-1.5B
/// test fixtures do not, keeping their footprint at pre-8p levels).
pub const W16_ENABLE_BYTES: usize = 2 << 30;
}

Second site: the alignment fix forced out by the bitparity test — dequant_q5_0_f16 (src/cuda_kernels.cu, post-fix form):

__global__ void dequant_q5_0_f16(
    const uint8_t* __restrict__ w, __half* __restrict__ out, int od, int id
) {
    int nb = id / 32;
    long long g = (long long)blockIdx.x * blockDim.x + threadIdx.x;
    if (g >= (long long)od * nb) return;
    int row = (int)(g / nb);
    const uint8_t* blk = w + g * 22;          // 22-byte block, stride not a multiple of 4
    float d = h2f(*reinterpret_cast<const uint16_t*>(blk));
    // 22-byte blocks are only 2-byte aligned: assemble qh from two u16
    // loads — a u32 load at blk+2 misaligns for even g
    // (cudaErrorMisalignedAddress 716; latent until 8p's bitparity test
    // exercised Q5_0 prefill GEMM for the first time).
    uint32_t qh = (uint32_t)*reinterpret_cast<const uint16_t*>(blk + 2)
                | ((uint32_t)*reinterpret_cast<const uint16_t*>(blk + 4) << 16);
    const uint8_t* qs = blk + 6;
    __half* o = out + (long long)row * id + (int)(g % nb) * 32;
    #pragma unroll
    for (int j = 0; j < 16; j++) {
        float lo = float(qs[j] & 0x0F) + 16.0f * float((qh >> j) & 1) - 16.0f;
        float hi = float(qs[j] >> 4) + 16.0f * float((qh >> (j + 16)) & 1) - 16.0f;
        o[j] = __float2half(d * lo);
        o[j + 16] = __float2half(d * hi);
    }
}

Two u16 loads (addresses 22g+2 and 22g+4 are both even, so 2-byte alignment holds) reassemble the original u32 — the value is identical and alignment is restored. The same fix lands on the fused path's block-level assembly helper bqa_q5_0 (bqa_q5_1 was unified into the two-u16 form while at it, even though the 24-byte block's u32 was safe already):

__device__ __forceinline__ void bqa_q5_0(
    const uint8_t* w, int row, int id, int e0, __half* dst
) {
    const uint8_t* blk = w + (long long)row * ((id >> 5) * 22)
                             + (long long)(e0 >> 5) * 22;
    float d = h2f(*reinterpret_cast<const uint16_t*>(blk));
    // 22-byte blocks are only 2-byte aligned: assemble qh from two u16
    // loads (a plain u32 load at blk+2 misaligns for even block indices —
    // cudaErrorMisalignedAddress, caught by cuda_prefill_fused_b_bitparity).
    uint32_t qh = (uint32_t)*reinterpret_cast<const uint16_t*>(blk + 2)
                | ((uint32_t)*reinterpret_cast<const uint16_t*>(blk + 4) << 16);
    int b = e0 & 31;
    #pragma unroll
    for (int l = 0; l < 8; l++) {          // 8 elements from nibble + high bit plane
        int e = b + l;
        uint8_t byte = blk[6 + (e & 15)];
        float nib = (e < 16) ? float(byte & 0x0F) : float(byte >> 4);
        float v = nib + 16.0f * float((qh >> e) & 1) - 16.0f;
        dst[l] = __float2half(d * v);
    }
}

Third site: the fused GEMM body's skeleton (gemm_qb_nt_kernel; 8m's 64×64 wmma structure unchanged, the B panel's staging switched to "quantized bytes into smem + register dequantization"):

__global__ void gemm_qb_nt_kernel(
    const __half* __restrict__ A, const uint8_t* __restrict__ W,
    float* __restrict__ C, int nt, int od, int id,
    int type_id, int q6_stride
) {
    __shared__ __half As[2][64 * 32];
    __shared__ __half Bs[2][64 * 32];        // dequantized f16 B tile (double buffer)
    ...
    int buf = 0;
    gemm_qb_load_tile(A, W, As[0], Bs[0], n0, m0, 0, nt, od, id, type_id, q6_stride);
    __syncthreads();
    for (int k = 0; k < id; k += 32, buf ^= 1) {
        if (k + 32 < id)
            gemm_qb_load_tile(A, W, As[buf ^ 1], Bs[buf ^ 1], n0, m0, k + 32,
                              nt, od, id, type_id, q6_stride);
        // same as 8m: fa[4] × fb[2] wmma pipeline, fc[2] accumulation
        wmma::load_matrix_sync(fb[0], &Bs[buf][wm * 16 * 32], 32);
        ...
        wmma::mma_sync(fc[0], fa[0], fb[0], fc[0]);
        wmma::mma_sync(fc[1], fa[1], fb[0], fc[1]);
        __syncthreads();
    }

gemm_qb_load_tile is the type-dispatched B assembly (one bqa_* branch per quant type, 8 total); its dequant math/rounding is bit-aligned with the standalone dequant kernel — the precondition for fused passing the bitparity gate.

3.3 Pitfalls

  • The latent alignment crash (this chapter's main pitfall): dequant_q5_0_f16's u32 load at blk+2 is aligned only for odd blocks (§2's arithmetic). It "lived happily" because the CUDA prefill GEMM had never been run on Q5_0 weights before — any Q5_0 model on the prefill GEMM would crash deterministically (cudaErrorMisalignedAddress 716). What caught it was not a Q5_0 user but 8p's own newly written bitparity test (full enumeration of 8 types × 2 super-block configurations). The fix is two-u16 load assembly, synchronized across three sites (the dequant kernel + bqa_q5_0/bqa_q5_1).
  • A performance feature pushed the test suite off a cliff (memory, not numerics): warming blindly for the small fixtures below 7B would cost each fixture +1–2 GB — and the suite keeps multiple loaded models co-resident in one oversubscribed CUDA pool, so later-registered models would probabilistically OOM while uploading weights. Hence the W16_ENABLE_BYTES = 2 GB gate: "a perf feature can break tests via footprint rather than via numerics". The gate keeps small models at their exact pre-8p footprint.
  • Bit-level equivalence means "the same rounding path", not "the same math": fused's dequantization must run in exactly the same order as the standalone dequant kernel — the same __float2half rounding, the same wmma accumulation — or the bitparity gate means nothing. The test even lays out each block's benign d (and min-type values) explicitly as f16 bytes, so random bytes cannot produce NaN/Inf that would disturb the bit comparison.
  • A pointer as the cache key = inheriting an ownership invariant: w16_cache keying on device pointers presumes that register_weight reuses the same device copy for same-name same-size and does not free on replace. That presumption is written in the field comment — whoever changes the registration semantics in the future hits the comment first.

4. Verification

  • cuda_prefill_fused_b_bitparity (new, src/graph/cuda_backend.rs): fused vs the legacy two-pass must be bit-identical across 8 quant types × 2 super-block configurations. Defends against: drift between fused's register dequantization and the standalone dequant kernel's rounding path (any "equivalent rewrite" of nibble order or scaling order gets caught). This is the test that snagged the Q5_0 alignment crash.
  • The legacy path's reference frame: the two-pass path itself is verified by the existing cuda_prefill_f16_gemm_parity (against a reference implementation) — one end of the bitparity chain must be pinned first.
  • Decode A/B: 40.8 vs 40.6 — defends against "prefill optimization hurting decode" (the cache changes the weight-resolution path; decode's kernel dispatch should be untouched).
  • Dual env valves: MINFER_NO_W16CACHE=1 (fall back to per-call scratch) and MINFER_FUSED_B=1 (the fused alternate) make both alternative paths independently re-testable.

5. Results

  • Prefill (7B q4_k_m @2K): 1201 → ~1400–1495 tok/s (+17–24%; row 5 records ~1400–1500, with the Δ column "~4.7× vs 8m" — anchored on 8m's landing at 294 tok/s, 1400/294 ≈ 4.7×; before 8m②'s cp.async staging the baseline was 294).
  • vs llama: ~2.3× (the §0 reading convention: the early 8m–8p rows' multiples use the llama-bench 3401 @2K figure as denominator).
  • Component level: the 288 ms per-forward dequant pass → 0 (paid once at load).
  • Decode: unchanged (40.8 vs 40.6 A/B) — the cache serves only the prefill GEMM.
  • VRAM: +2 B/element, ~8.6 GB on 7B q4_k_m; the fused alternate (MINFER_FUSED_B=1) adds nothing but is slower at large nt, positioned as the memory-constrained escape hatch.
  • Later evolution: once R1's int8 MMQ path became the default, the f16 cache retreated to being the skeleton of the MINFER_MMQ=0 escape path (§1.1: ~2353 tok/s, ~20.5 GB) — the infrastructure this step built outlived the entire MMQ campaign.

6. Lessons

  1. One-time materialization of immutable weights beats per-call recomputation (row 5's own words: load-time materialization beats per-call dequant) — the test is always "will this value change", never "is computing it once expensive".
  2. The bitparity test pays for itself: the full 8-type × 2-configuration enumeration caught, on merge day, a latent bug that would have crashed every Q5_0 prefill (row 5's own words: a bitparity test pays for itself immediately).
  3. Alignment of strided access is the arithmetic of stride × g + offset (mod width) — a u32 load on 22-byte blocks is aligned only for odd blocks; the fix is assembling from narrower naturally-aligned loads.
  4. Performance features need "a size gate + an exit valve": footprint side effects reach the test suite as probabilistic OOM, which is harder to attribute than numeric errors.

← 04 · decode-start CPU stalls (8o) · Index · 06 · decode MMVQ (8e)

06 · 8e/8e② — decode MMVQ: dp4a integer dot products + the llama.cpp launch table (LANDED)

Result: 7B decode +37% (q4_K); kernel-level +74–77% per matmul (194–207 vs 112–117 GB/s, L2-defeated microbenchmark); q6_K/q5_K followed + an od·id ≥ 24M shape gate. Commit: b7b8e73 (8e: the q4_K MMVQ reversal), 1298cb2 (8e follow-up: q6_K/q5_K kernels), 1d28235 (the shape gate into graph dispatch). Date: 2026-08-30.

1. Background — where things stood

8o (chapter 04) removed the one-time decode-start stall, but decode's steady state still ran on the f32-activation row-wise kernel. This path has an earlier dark history: in the P4 draft era before Phase 7 (docs/CUDA_OPTIMIZATION.md Part IV), a "tiled quantized matmul" had already been tried once, the verdict then being negative return, with a seemingly authoritative explanation — "116 GB/s = the platform's streaming bandwidth limit" — so the direction was buried, and llama.cpp's corresponding design was misread as "shared-memory tiling with Stream-K decomposition".

8e began with doubt about that buried verdict, and the evidence for the doubt took only one control experiment: an empty kernel that computes nothing and only reads weights reaches 252.7 GB/s on GB10 — 93% of the theoretical 273 GB/s. So 116 GB/s could not possibly be a platform limit; it had to be the kernel's own geometry.

The disease was immediately clear. The old f32-activation kernel used a 2 warps × 4 rows layout, each lane serially processing an entire 144-byte block — on the 7B ffn_down shape that is only ~28K threads in flight across the whole grid. GB10's LPDDR latency hides behind concurrency: 28K threads cannot spread enough outstanding loads, and the kernel actually ran at ~46% of platform bandwidth.

llama.cpp's decode dot-product path mul_mat_vec_q (MMVQ) is a different geometry, and a parameterized one: its MMVQ_PARAMETERS_* gives a launch-parameter table per GPU generation — the GB10 entry is 8 warps, one output row per block. Porting that geometry as-is was 8e's reversal: within a single day, the old "negative return + platform limit" verdict was overturned and 7B decode gained +37%. Row 6's lesson follows from it: port the launch-table parameters, not just the math — half of the original mistake lived outside the math.

2. Principle — the GPU mechanism

dp4a (DOT4-accumulate, an integer instruction since sm_61+): __dp4a(a, b, c) completes 4 pairs of int8 products and accumulates into a 32-bit integer in one instruction — `a0·b0 + a1·b1

  • a2·b2 + a3·b3 + c`. For quantized dot products it is a natural building block: one uint32 register holds 4 bytes (4 int8 activations, or 4 unpacked 4-bit nibble components), and one dp4a does 4 multiply-adds. The integer pipeline's throughput far exceeds an f32 chain, and it keeps weight traffic at the quantized byte width (q4_K 0.5 B/weight) — the price is that activations must become integers too.

The q8_0 activation quantization path (decode side): at nt==1 the activation row x (f32, id elements) is quantized into q8_0 in 32-element blocks before entering the matmul: per block d = max|x|/127, q_e = rintf(x_e/d) clamped to [-128, 127]. This matches the CPU path's established convention (§Core Conventions: all CPU matmuls consume Q8_0 activations) — CUDA decode moving from f32 activations to q8 activations is essentially pulling minfer back into llama.cpp's isomorphic design. minfer uses the pad40 layout: each block is 40 bytes = [f16 d][2B pad][32B int8 payload] (plus a 4 B sum of quantized values at the block tail, used by the prefill MMQ's min-term correction; the MMVQ kernel reads only d and the payload and never sees that word). The pad's job is to land the int8 payload on 4-byte alignment — dp4a's uint32 reads require it.

The min term also has a non-obvious trick. A Q4_K dequantized value = d·s·nib − dm·m (s/m are the sub-block scale/min, d/dm the super-block's), so:

W·x = Σ_sub-blocks [ d·s·Σ_e(nib_e·x_e) − dm·m·Σ_e(x_e) ]

The second term needs not the weight sum but the activation sum Σx. In the kernel it is accumulated per sub-block with one __dp4a(0x01010101, xa, sx) — four ones times four activation bytes is an incremental update of Σx, done in one line, with no separate bsums pass.

The MMVQ geometry (what the launch table contains): llama.cpp's mul_mat_vec_q for GB10 is 8 warps (256 threads) per block, one block computing one output row, with that row's (block, 32-element sub-block) units round-robined across the 256 threads and reduced by a block-wide reduction at the end. The concurrency arithmetic: 7B ffn_down (od 3584, one row per block) = 3584 blocks × 256 threads ≈ 917K threads in flight, versus the old kernel's ~28K — 32× the concurrency, and LPDDR latency is finally covered by outstanding loads. This is exactly what "launch-table parameters are part of the design" means: the same math spread as 8 warps × one-row-per-block versus 2 warps × four-rows-per-block — the former +74–77%, the latter stuck at 46% of platform bandwidth.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Copy llama.cpp's structure; invent nothing. 8e's first commit title says "re-examined" — what was reversed was our own old verdict, and the method was porting mul_mat_vec_q's block geometry, unit mapping, and reduce structure as-is, aligning only the K-quant block-format details (q4_K's scale/min layout, q6_K's 16-element units) to GGUF semantics.
  • One kernel per type: one each for q4_K / q6_K / q5_K (1298cb2), sharing the block shape and reduce while each owning its own nibble/byte unpacking code — the K-quant super-block formats differ enough that forcing an abstraction would sacrifice readability.
  • The shape gate lives in dispatch (1d28235): MMVQ does not win everywhere. Measured: at od·id < ~24M elements it actually loses (od 512 → 4.5× slower, 896 → 3.0×, 2048×4864 → 1.66×) — with too few rows each thread gets only 1–2 units and q5/q6's uncoalesced byte loads are exposed; large shapes win (7B ffn_down 3584×18944 → 1.5× faster, lm_head 152064×3584 → 1.4×). The q4_K gate uses id ≥ 2048 && id % 32 == 0 (below id 2048 the margin shrinks to launch-latency noise), the q5_K/q6_K gate uses od·id ≥ 24M. Tensors outside the gate keep the padded-f32 kernel — small tensors stay on the safe path, the "dispatch gated by shape" recorded in row 6.
  • The q8 scratch's capture constraint: buf_q8_decode is sized by id, constant within a graph, grown during warmup and never inside a capture window — CUDA Graph replay must contain no allocation events; this is 7d's existing capture/replay discipline applied in this step.

3.2 Key code

The q4_K MMVQ kernel (src/cuda_kernels.cu, q4_k_q8_mmvq; some unpacking details elided in the middle):

__global__ void __launch_bounds__(256) q4_k_q8_mmvq(   // 8 warps = launch table
    const uint8_t* __restrict__ weights, const uint8_t* __restrict__ acts8,
    float* __restrict__ output, int od, int id, int nt
) {
    const int row = blockIdx.x;                        // one block = one row
    const int nbe = (id + 255) / 256;
    const int row_stride = nbe * Q4KB;
    const int nsub = (id + 31) / 32;
    const uint8_t* x8row = acts8 + (size_t)t * nsub * Q8PB;

    float acc = 0.0f;
    for (int u = threadIdx.x; u < nsub; u += 256) {    // unit round-robin
        const int blk_i = u >> 3, sub = u & 7;
        const uint8_t* blk = weights + (size_t)row * row_stride + blk_i * Q4KB;
        const float d  = h2f(*reinterpret_cast<const uint16_t*>(blk));
        const float dm = h2f(*reinterpret_cast<const uint16_t*>(blk + 2));
        uint8_t s8, m8;
        get_scale_min_k4(sub, blk + 4, &s8, &m8);
        const uint32_t* qw = reinterpret_cast<const uint32_t*>(blk + 16 + (sub >> 1) * 32);
        const bool lo = (sub & 1) == 0;                // even subs take the low nibble, odd the high
        const uint8_t* x8 = x8row + (size_t)u * Q8PB;  // pad40 q8 activation block
        const float d8 = h2f(*reinterpret_cast<const uint16_t*>(x8));
        const uint32_t* xw = reinterpret_cast<const uint32_t*>(x8 + 4);
        int dot = 0, sx = 0;
        #pragma unroll
        for (int v = 0; v < 8; v++) {                  // 32 bytes = 8×uint32
            const uint32_t w = qw[v];
            const int n = lo ? (int)(w & 0x0F0F0F0F) : (int)((w >> 4) & 0x0F0F0F0F);
            const int xa = (int)xw[v];                 // 4 activation int8s
            dot = __dp4a(n, xa, dot);                  // 4 multiply-adds/instruction
            sx  = __dp4a(0x01010101, xa, sx);          // Σx, for the min term
        }
        acc += d8 * ((float)s8 * (float)d * (float)dot
                    - (float)m8 * (float)dm * (float)sx);
    }
    // …shfl_xor 8-warp tree reduction, thread 0 writes output[t*od + row]…
}

The dispatch gate (src/cuda.rs, the Q5_K branch; the comment carries the crossover data measured at the time):

#![allow(unused)]
fn main() {
// 8e follow-up: decode (nt == 1) joins the MMVQ structure
// (dp4a over q8 activations, one row per 256-thread block).
// Shape gate measured on-device (dbg micro-bench, padded f32
// vs mmvq): od*id < ~24M elements loses (od 512 → 4.5x
// slower, 896 → 3.0x, 2048x4864 → 1.66x) because 1-2 units
// per thread expose the uncoalesced q5/q6 byte loads; large
// shapes win (7B ffn_down 3584x18944 → 1.5x faster, lm_head
// 152064x3584 → 1.4x). MINFER_NO_KQ_MMVQ=1 forces f32.
if nt == 1 && od * id >= 24_000_000 && !Self::no_kq_mmvq() {
    self.q5_k_decode_mmvq(wptr, x, out, od, id, nt);
    Ok(())
} else {
    launch!(launch_q5_k_f32_matmul)
}
}

The q8_0 activation quantization body (per-block quantization in the pad40 layout; D3-5 later moved this body verbatim into the rms/swiglu producers for fusion — the quantization math is exactly 8e's standalone kernel):

__device__ __forceinline__ void quantize_pad40_block(
    const float* __restrict__ src, uint8_t* __restrict__ dst
) {
    float4 sv[8];                                       // 32 floats = 8×float4
    #pragma unroll
    for (int v = 0; v < 8; v++)
        sv[v] = *reinterpret_cast<const float4*>(src + 4 * v);
    float am = 0.0f;                                    // max|x| within the block
    for (int v = 0; v < 8; v++)
        am = fmaxf(am, fmaxf(fmaxf(fabsf(sv[v].x), fabsf(sv[v].y)),
                             fmaxf(fabsf(sv[v].z), fabsf(sv[v].w))));
    float d = am / 127.0f;
    float di = (d != 0.0f) ? 1.0f / d : 0.0f;
    *reinterpret_cast<__half*>(dst) = __float2half(d);  // [f16 d]
    uint32_t packed[8];
    for (int v = 0; v < 8; v++) {                       // rintf + clamp → int8
        const float* e = &sv[v].x;
        uint32_t p = 0;
        for (int j = 0; j < 4; j++) {
            int q = int(rintf(e[j] * di));
            q = max(-128, min(127, q));
            p |= (uint32_t)(uint8_t)(int8_t)q << (8 * j);
        }
        packed[v] = p;
    }
    for (int v = 0; v < 8; v++)                         // payload lands at offset 4:
        *reinterpret_cast<uint32_t*>(dst + 4 + 4 * v) = packed[v];  // 4B aligned
    *reinterpret_cast<uint32_t*>(dst + 36) = /* Σq, for MMQ; not read by MMVQ */;
}

3.3 Pitfalls

  • "Platform limit" was the wrong attribution. The original 8e's 116 GB/s was recorded as the platform's streaming limit, fossilizing a kernel-geometry problem into a physical constant. One data point — a read-only empty kernel at 252.7 GB/s (93% of theoretical bandwidth) — dismantled it. The lesson, mechanized: for any "XX GB/s = the limit" conclusion, run a ceiling probe first.
  • A test generator's block-layout bug impersonating a kernel bug (1298cb2). When q6_K followed, tests failed for a while and the investigation suspected the kernel had wrong high bits ("wrong high bits"); the final finding: the kernel draft was correct — it was the test's generator writing d at offset 0 — a Q6_K block is ql[128] + qh[64] + sc[16], 208 bytes, and only then the 2-byte d (block 210 B). The generator did not lay out data per the real layout, so the bit comparisons went all red, naturally.
  • The small-shape reversal. MMVQ's round-robin assumes "enough units per thread to amortize overhead"; with few rows (od 512) each thread is left with 1–2 units and the latency of q5/q6's uncoalesced byte loads is exposed directly — 4.5× slower than padded-f32. The shape gate (24M / id≥2048) is not conservative decoration; it is the measured crossover.
  • Allocation inside a capture window. If the q8 scratch grows during capture it breaks replay; handled by the "grow fully in warmup, constant within the graph" discipline (§3.1).

4. Verification

  • The L2-defeated microbenchmark (bench8e2): 194–207 GB/s vs 112–117 GB/s on the 7B shapes — L2 warmth deliberately flushed before measuring, defending against fake bandwidth from "reading L2, not DRAM".
  • Dispatch equivalence (1d28235): in-gate shapes through graph dispatch are bit-identical to direct kernel invocation (bit-exact 0.0000) — defends against "the shapes entering through the gate actually running a different kernel/parameters than the directly-invoked shapes that were verified"; out-of-gate sub-gate shapes (attn_v, 0.5B ffn_down) verify the kernel itself by direct invocation.
  • The full CUDA suite: 158 passed, 0 failed (as of 1d28235).
  • MINFER_NO_KQ_MMVQ=1 forces the f32 path back — the standing valve for A/B and regression. (A semantics note: MMVQ vs the padded-f32 kernel is not a bitwise relationship — q8 activation quantization vs reading f32 directly are different numerical semantics classes; what it aligns with is the CPU path's Q8_0 activation convention. The later D-series campaign built the mature gate set of argmax/greedy/A/B for exactly this kind of "quantization semantics switch" — see the decode chapters after 06.)

5. Results

  • Row 6: 7B decode +37% (q4_K); q6_K/q5_K followed and landed; the vs-llama column reads "—" (no whole-prefill comparison was recorded at the time; decode-vs-llama became an accounting item only in the D-series era).
  • Kernel level: +74–77% per matmul (194–207 vs 112–117 GB/s, L2-defeated); the old f32-activation kernel's 46% of bandwidth → the MMVQ structure pushed the 7B shapes to ~200+ GB/s.
  • The shape gate: od·id ≥ 24M (q5_K/q6_K) / id ≥ 2048 (q4_K) — small tensors outside the gate keep padded-f32. D3-7 2b later lowered the q6_K gate to 4M for the attn_v class (a story for another chapter).
  • The long tail: this dp4a MMVQ skeleton became the foundation of all later decode work — R2's weight-streaming rework (tg128 +6.9%), D3b-1b's pipelining, D4-4's dense split-plane, D3-5's fused-producer A quantization, and D3-8's FusedQKV concat equivalence argument all stack on this structure. After D4-4, 7B decode tg128 is 51.2 (1.074× vs llama) — the starting point was this step's +37%.

6. Lessons

  1. Port the launch-table parameters, not just the math (row 6's own words): block geometry, warps, and unit mapping are half the design — with the same dp4a math, 2 warps × four-rows-per-block stalls at 46% of bandwidth while 8 warps × one-row-per-block gains +74–77%.
  2. Run a read-only ceiling probe before believing any "bandwidth = platform limit" conclusion — the 116 GB/s "limit" was dismantled in one stroke by a 252.7 GB/s empty kernel.
  3. For a latency-bound kernel, fix concurrency first (28K → 917K threads in flight), then talk bytes — 32× concurrency is not tuned into existence; it is rearranged into existence by geometry.
  4. When a K-quant test goes red, check the generator's block layout first (d at offset 208, not 0) before suspecting the kernel — fixture bugs and kernel bugs are fixed in completely different ways.

← 05 · persistent f16 weight cache (8p) · Index · 07 →

78 · Phase-8 correctness & engineering-debt batch — 8a review fixes, 8h①, 8i (LANDED, ledger closed)

Result: 11 Phase-8 review findings fixed (961f696); the F32-weight E2E exercise (8a②) caught a latent CPU bug — vec_ops::mat_mul_f32 wrote token-TRANSPOSED output for nt > 1; rebuild-gate, multi-split-capture and multiturn-reuse tests added; device suite grew 147 → 158 across the batch. The last open item (8a① macOS Metal regression run) was hardware-blocked at the time and closed on 2026-09-10 on an Apple M4 Pro — §3.2. Commits: 961f696 (review batch), 789e64b→b849601 (batch-1 + 8g①), 60e9cc1 (8h① + 8i). Date: 2026-08-29 (8a① closed 2026-09-10).

1. Background — where things stood

Phase 7 had just landed (the CUDA graph backend, 7a–7e — see doc 01). Phase 8 opened deliberately with debt instead of performance: an independent review of 4fcd0d8..5cbb4ca surfaced 11 findings, and the 7e leftovers still had open verification items. The ordering rationale is structural: every later optimization session gates its landings on the device suite and on greedy token-identity checks, so any hole here silently weakens every gate downstream. A correctness batch is cheap; a perf campaign built on a leaking suite is not.

Where things stall without this step: the capture-window lifecycle had error paths that could abort mid-window, several kernels lacked weight-READ row guards, and there was no end-to-end F32-weight model coverage at all (the F32 kernels from 7e④ had parity fixtures but never ran a real model) — exactly the kind of gap that hides a real bug forever.

2. Principle — why these fixes take this shape

This is a batch document, so there is no single mechanism; the common threads:

  • Lifecycle pairing: every CUDA Graph capture window needs abort, replay, and error paths that agree with each other (capture-window abort on execute_node errors; the replay-vs-open-window guard).
  • Guard the reads, not just the writes: kernels that read weight rows (q4_K/q6_K raw + padded, f32_vec) got row guards — a malicious/corrupt shape must fail loudly, not read out of bounds.
  • Bookkeeping must track reality: pos_scratch pool_gen bump, ring-wrap reset race, stale padded_weights flag, pinned-ring alloc logging — all cases where cached state outlived the condition that justified it.
  • Env flips must be part of the identity: MINFER_NO_FUSE_FFN / MINFER_NO_FUSE_QKV have to flip CParams, otherwise an A/B comparison silently reuses the wrong graph (the "A/B footgun" class — 8a③ closes it permanently with a test).

3. Implementation

3.1 The 8a review batch (961f696) — 11 findings, all fixed

From the independent Phase-8 review of 4fcd0d8..5cbb4ca:

  1. Capture-window abort on execute_node errors (no half-open windows).
  2. Replay-vs-open-window guard (a replay must not interleave with an open capture).
  3. Kernel weight-READ row guards (q4_k/q6_k raw + padded, f32_vec).
  4. pos_scratch pool_gen bump (stale positions buffer across pool generations).
  5. Ring-wrap reset race in the pinned ring.
  6. Stale padded_weights flag.
  7. Pinned-ring alloc logging (diagnosability).
  8. Metal ffn_gu loader gate (~1.99 GiB concat on 7B — memory-only, no perf intent).
  9. A stray 7 MB trace file removed from the tree.
  10. Dump/trace/viz fuse_ffn gating aligned with the engine's actual gate.

Suites re-verified after the batch: cuda 147/0, plain 133/0, fmt clean; 7B and 0.5B end-to-end greedy output coherent.

3.2 Batch-1 (789e64b→b849601): 8a③ + 8g① + 8a②

  • 8a③ (rebuild-gate unit test) — DONE. fuse_flags_are_part_of_the_reuse_identity asserts that flipping MINFER_NO_FUSE_FFN (and MINFER_NO_FUSE_QKV) changes CParams, i.e. GraphCache::try_reuse returns false in both directions.

  • 8g① (decode-only capture gate) — DONE, shipped with this batch: ComputeGraph::capture_nt_hint() + a decode-only capture gate in graph_replay_step. Before it, the scheduler ran the 3-run capture protocol for EVERY CUDA split, so a repeated identical-nt prefill (the server/slot scenario) would silently start capturing a ~437-node graph. cuda_prefill_shaped_graph_never_captures runs a prefill-shaped (nt=8) graph 4× and asserts zero captures. (Productization followed in 8g②/R3-B — doc 07.)

  • 8a② (F32-weight GGUF E2E) — DONE, and it caught a real CPU bug. 7e④'s F32×F32 kernels had parity coverage but no end-to-end model (none cached). A byte-exact GGUF→F32 converter (numpy) produced test models, and qwen2.5-0.5B-F32 produced GARBAGE on CPU while the same weights kept as Q4_0 ran fine. Root cause: vec_ops::mat_mul_f32 wrote token-TRANSPOSED output for nt > 1 — decode (nt==1) was accidentally correct, which is why nothing had ever caught it. After the fix, both F32 models (0.5B + qwen3-0.6B) produce identical greedy text on CPU and CUDA. Regression test: f32_matmul_nt2_token_major.

  • 8a① (macOS Metal regression run) — CLOSED 2026-09-10 (Apple M4 Pro, Metal 4, cargo build --release, rustc 1.92.0). The qwen2 FFN-fusion gate was decoupled from fuse_qkv to CParams.fuse_ffn in 7e⑤ (mirroring Qwen3's existing intent), and 961f696 additionally gated the Metal ffn_gu loader registration on the same condition (nf ≤ 16384 + MINFER_NO_FUSE_FFN) — the two edits were Metal-relevant but had never run on a Mac. Verification, all three checks green:

    1. Post-vs-pre-change greedy identity — the pre-change tree (78d410a, = f54f721~1: FFN fusion keyed off fuse_qkv, ffn_gu registered unconditionally) was built in a worktree and A/B'd against the current tree, same prompt/seed/--greedy: byte-identical on 0.5B q4_k_m ×2 prompts, 0.5B q4_0 ×1, 7B q4_k_m ×2 (64 tokens each), all on Metal.
    2. MINFER_NO_FUSE_FFN A/B — fused vs unfused (unfused still running the FusionPass, per the AGENTS.md comparison rule) is byte-identical on the same 5 model/prompt pairs, so the decoupled gate does not change numerics. This also closes the pre-change footgun: before 7e⑤ the toggle was read inline in the build gate but was not part of the reuse identity (fuse_flags_are_part_of_the_reuse_identity now asserts it).
    3. 0.5B still fuses on Metal — MINFER_TRACE decode graph: 24 × fused_ffn nodes with weight: blk.{i}.ffn_gu, nf: 4864, backend: metal; the 7B decode graph has 0 × fused_ffn and 28 × swiglu (nf = 18944 > 16384, gate closed, exactly as intended). The fused run succeeding at all proves the loader registration happened — execution aborts if any weight the graph reads is unregistered.
    • Doc correction: the nf quoted for the 0.5B gate check was 2944; the actual blk.0.ffn_gate.weight out-dim is 4864 ([896, 4864], well under the 16384 gate). Conclusion unchanged.
    • Suites on the same Metal machine: cargo test --release 169 passed / 0 failed / 11 ignored = 155 bin + 14 integration (conversation_cli 3, gemm_isolation 5, flash_attn_isolation 2, flash_attn_blk_isolation 1, gqa_attn_isolation 3 — the four Metal isolation suites ran on device and are green). The 11 ignored are the env-dependent helpers (real-data dumps, throughput profiling, and the conversation_cli model-dependent cases), unchanged from before this check. cargo fmt --check clean.

    The interim state (2026-08-29 → 2026-09-10) stayed on the open ledger rather than being silently dropped; this entry is that ledger item being closed.

3.3 8h① — stale docs marked SUPERSEDED (60e9cc1)

CUDA_OPTIMIZATION.md / CUDA_PROBLEMS.md (the pre-Phase-7 records) got SUPERSEDED banners pointing at the current plans, with the absorbed ideas named: cuBLAS → 8k evaluation, MMQ tiling → 8e (later reversed into the MMVQ win), GPU quantize → 8c. (2026-09 consolidation note: both legacy docs have since been retired — CUDA_PROBLEMS.md's surviving conclusions live in CUDA_OPTIMIZATION.md Appendix C.)

3.4 8h② / 8h③ — deferred, deliberately

  • 8h② (optional CUDA CI runner) — DEFERRED: device-gated tests skip gracefully today; a self-hosted GB10 runner would keep the suite honest on every commit, but requires standing runner infrastructure (a wired-up machine + runner registration) — not achievable from a dev session. The 158-test device suite runs green locally.
  • 8h③ (temp files): the Phase-7 ledger (/tmp/minfer_phase7/TEMPS.md) is closed; cleanup awaits the user's decision (no auto-delete policy).

3.5 8i — graph integration test debts (60e9cc1)

  1. Multi-split capture — cuda_multisplit_capture_bit_parity: a CUDA → CPU (Softmax) → CUDA graph yields two CUDA splits; both capture (captured_count == 2) and replay bit-identical to direct launches.
  2. Multi-turn conversation — cuda_conversation_multiturn_reuse (q4_0 0.5B, device): turn-2 incremental (append-only KV + reused decode graph) vs turn-2 rehydrated from history (fresh graphs + full re-prefill) produce IDENTICAL greedy text. ConversationSpec now derives Clone; note that device tests must call CudaState::init() themselves (get() only reads the singleton).
  3. Slot loop — covered by the same test: both paths run the GraphCache prefill→decode alternation the OpenAI server slot uses; a dedicated axum-level test remains out of scope (needs a live HTTP harness).

4. Verification

  • Device suite progression across the batch: cuda 147/0 → 155/0 → 158/0 (the later steps in doc 79 added their own tests); plain suite 133/0; fmt clean.
  • fuse_flags_are_part_of_the_reuse_identity — defends the A/B workflow (an env flip that doesn't change the reuse identity would compare the same graph twice).
  • cuda_prefill_shaped_graph_never_captures — defends the "prefill never captures" invariant that 8c later relies on for capture-safety.
  • f32_matmul_nt2_token_major — pins the fixed nt>1 F32 matmul layout.
  • cuda_multisplit_capture_bit_parity / cuda_conversation_multiturn_reuse — integration-level bit-parity and KV-continuity guards.
  • 7B/0.5B E2E greedy coherence after every commit in the batch.

5. Results

A correctness-only tree: no perf deltas intended or claimed. The concrete outcomes are the 11 review fixes, the fixed nt>1 F32 CPU matmul, three new permanent tests, and the completed Phase-8 ledger (8a③/8g①/8a②/8h①/8i DONE; 8h② deferred; 8h③ user-decision; 8a① closed 2026-09-10, §3.2). The batch also unblocked the rest of Phase 8: 8c's capture-safety argument ("prefill never captures since 8g①") and the 8f model-coverage work both stand on this foundation.

6. Lessons

  1. Run the correctness batch FIRST — every later gate (parity, greedy identity, suite) inherits its strength from this layer.
  2. "Add the missing E2E coverage" is a test that can fail: the F32 model exercise caught a latent, year-class CPU bug that decode-only usage had hidden forever (nt>1 transpose).
  3. Env toggles must join the reuse identity — otherwise A/B comparisons silently compare identical graphs; assert it, don't remember it.
  4. Hardware-blocked verification stays on the open ledger — never silently dropped, and never assumed to be free once the hardware appears. 8a① sat blocked for 12 days and, on being re-run (2026-09-10), was green on all three checks — but that verdict had to be produced, not inferred from "the change looked memory-only".

← 77 · Index · 79 →

79 · Phase-8 coverage & first-measurement batch — 8b KV f16, 8c shaped Q8_0 GEMM, 8d split-K attention, 8f Q5_K kernels, 8l llama parity baseline, 8q Q5_0 (LANDED)

Result: KV f16 +~11% 7B @2K decode (8b); shaped Q8_0-activation prefill GEMM +24% 0.5B prefill with a −63% negative that proved the shape gate load-bearing (8c); split-K flash-decoding +36% 7B @2K decode (8d, later superseded by R4); Q5_K/Q5_1 kernels completed the CUDA model gate (8f); the 8l parity benchmark set the llama.cpp baseline (decode 1.15–1.76×, prefill 18–110× gaps) that directed 8m–8p and the whole MMQ campaign; 8q enabled Q5_0 models (0.5B q4_k_m: 148.7 → ~1200 prefill / 56.9 → ~306 decode tok/s). Commits: f7b0036 (8b), 69a27c5 (8c), a5af60f (8d), b959ec9 (8f), acca28f (8l), 9f419f9 (8q, squashed). Dates: 2026-08-29 → 2026-09-08.

1. Background — where things stood

After the correctness batch (doc 78), Phase 8 chased two things in parallel: decode/prefill performance headroom, and model coverage (a weight type the CUDA gate rejects means a whole model silently falls back to CPU). This doc records the six items that neither produced a master-table "campaign" of their own nor fit the perf-step docs 02–06 — the coverage work and the first honest measurements against llama.cpp. The 8l benchmark deserves special weight: its decode and prefill gap tables are the origin of the 8m–8p prefill line, the R-series, and the r5–r60 MMQ campaign.

2. Principle — the GPU mechanism (per item, in brief)

  • 8b — KV f16 halves the attention/KV byte traffic. K/V rows are read once per decode step per head; storing them as f16 instead of f32 halves those bytes while the accumulation stays f32. The win grows with context (more KV rows to stream).
  • 8c — activation quantization only pays when the matmul is activation-heavy. q8_0×int8 dots trade per-element dequant for a smaller byte stream; at id ≤ 8192 (attn/qkv/o shapes) the A-side savings dominate; at 7B ffn_down (id=18944) the shape is weight-bound and the q8_0 kernel streams weight bytes SLOWER than the f32 kernel — hence a shape gate.
  • 8d — split-K fills the machine. The incumbent kernel ran one warp per (token, head): 28 warps total at 7B @2K (the GPU is idle), streaming 2K KV rows serially at ~3 GB/s effective. Splitting the KV scan into SPLITS=8 chunks scanned in parallel (fixed grid → capture-safe) raises warps-in-flight ~8×; pass 2 merges (mx, S, oc) partials.
  • 8f — the all-or-nothing gate needs EVERY matmul weight type. The CUDA participation gate admits a model only if every matmul weight has a kernel; the 0.5B "q5_k_m" file actually stores Q5_1 + Q6_K + Q8_0, so Q5_1 kernels were required alongside Q5_K.
  • 8l — measurement, not mechanism. Cross-benchmarking against llama.cpp on the same GGUF files turns "we feel slower" into per-shape gap numbers, which is what ranks the follow-up levers.
  • 8q — the CUDA alignment contract is not advisory. A reinterpret_cast<const uint32_t*> load of a word at byte offset 2 of a 22-byte block is only 2-byte aligned for even block indices — on GB10 unified memory it faults nondeterministically depending on page-mapping state (small-shape parity tests passed; the real 93.6 MB tok_embd faulted).

3. Implementation

3.1 8b — KV f16 on CUDA (f7b0036)

store_kv_f16 + gqa_attn_f32_f16kv — an exact structural mirror of the f32 attention kernel: same online softmax, same warp reductions, same guards; the only delta is half4 → float4 K/V loads with f32 accumulation. Policy mirrors Metal: MINFER_CACHE_TYPE=f16|f32 override, auto f16 when n_layers × n_kv_embd ≥ 8192 (the 7B class), decided at model load and cached per CudaBackend instance. Caveat on record: MINFER_GRAPH_DUMP reads KV regions as f32, so it is incompatible with f16 KV (debug dump path only).

3.2 8c — prefill Q8_0-activation GEMM, SHAPED (69a27c5)

Measure-first verdict (standalone nvcc A/B, quantization included): +38–44% at id ≤ 8192 (0.5B shapes, 7B attn/qkv/o — activation-heavy), +4.7% at 7B ffn_gu (weight-bound), −63% at 7B ffn_down (id=18944). A blind wire of the 7e⑥ idea would have REGRESSED 7B-class Q4_0 ffn_down by ~60%. Wired only the winning region: nt > 1 && id ≤ 8192 routes Q4_0 prefill through quantize_q8_0 + q4_0_q8_0_matmul (grow-on-demand scratch; capture-safe because prefill never captures since 8g①, doc 78). E2E: 0.5B prefill @3.6K tokens 1005 → 1246 tok/s (+24%), greedy text unchanged.

3.3 8d — split-K flash-decoding decode attention (a5af60f)

nsys capture was fixed first (the 7e② "no kernel data" issue was the report workflow, not the config): nsys profile --trace=cuda + nsys stats / sqlite over CUPTI_ACTIVITY_KIND_KERNEL, decode step isolated by gap segmentation. Attribution at 7B @2K decode: gqa_attn 48.3% of the step, q4_k matmul 36.4%, q6_k 13.9%. Fix: split-K flash-decoding — pass 1 scans SPLITS=8 KV chunks in parallel (FIXED grid × nh blocks, ranges derived from device-side positions so CUDA Graph capture stays valid), pass 2 merges (mx, S, oc) partials. Size- stable state scratch, grown during warmup only; templated over the KV element type (f16 + f32 layouts); dispatched at nt == 1 only. Superseded later by R4's dim-parallel lane rewrite (doc 10), which removed the local-memory accumulator 8d still paid for.

3.4 8f — Q5_K + Q5_1 kernels (b959ec9)

Both f32-activation matmuls mirror the Q4_0/Q4_K structures; Q5_K decodes the transposed qh (bit s of byte l) and deinterleaved qs chunks, with sub-level tail masking for partial last super-blocks (0.5B id = 896 = 3.5×256; dispatch requires id % 32 == 0). embed_rows_q5_1/_q5_k cover the q5_1 token embedding. Gates updated in BOTH qwen2 and qwen3 weights_on_cuda. Q5_0 was deferred at the time (no cached model needed it) — until 8q.

3.5 8l — the llama.cpp parity benchmark (acca28f)

Cross-benchmark vs llama.cpp ca3d5a3e1 (build 10665), same GGUF files, GB10, -ngl 99 -t 8, llama-bench -r 3 (FA 0/1 matrix) + llama-cli cross-check; minfer side 3 reps --greedy.

Found + fixed the Q5_K registration gap first: the CUDA whitelist in models/qwen2/loader.rs (added in 7c) was missing TensorType::Q5_K (Metal's list had it). Every Q5_K matmul silently ran on CPU with per-token GPU↔CPU copies — 0.5B q5_k_m decoded at 51.6 tok/s behind CUDA GATE: ... spam. A one-line fix → 246.3 tok/s (4.8×).

modelllama fa0llama fa1minfergap (fa1)
0.5B q4_0417.6453.5258.01.76×
0.5B q5_k_m311.0394.6246.31.60×
0.6B q8_0273.8290.3197.81.47×
7B q4_k_m46.447.141.21.15×

7B @2K context: llama 44.9 (llama-cli) / ~45.1 (bench) vs minfer 31.3 → 1.43×. Depth penalty: llama −5% vs minfer −24% — the attention/KV path loses ~3 ms/token at 2K. Prefill gap: 110× (7B q4_k_m: 3401 vs 30.7 tok/s), 69× (0.6B q8_0), ~18–30× (0.5B). Root cause: minfer's quantized prefill reused the decode-shaped kernels with grid.y = nt — every token block re-streamed the full weight matrix (7B: ≈4.4 GB × 1920 tok ÷ 62 s ≈ 135 GB/s of pure redundant traffic). Decode-gap attribution: both engines bandwidth-limited at 7B (llama ≈221 GB/s ≈ 81% of the 273 GB/s peak, minfer ≈193 ≈ 71%); the residual 15% + the small-model 1.5–1.8× are per-token overhead (graph replay is worth +24% on 0.5B: 208 → 258 with MINFER_NO_CUDA_GRAPH=1) plus llama's FA-style single-pass decode attention vs minfer's multi-pass scores kernel. The ranked follow-ups became the later campaigns: ① prefill tiled int8 GEMM (→ R1 + 8m–8p + r5–r60), ② FA-style decode attention (→ 8d/R4/D3a), ③ per-token CPU overhead audit (→ R3).

3.6 8q — Q5_0 CUDA enablement + the misaligned-load fix (9f419f9)

0.5B q4_k_m GGUFs carry Q5_0 weights (token_embd, per-layer attn_q/k/v/o, ffn_gate/up; attn_v q8_0, ffn_down q6_K, output q8_0), but the CUDA participation whitelist (models/{qwen2,qwen3}/graph.rs) predated Q5_0, so the whole model silently fell back to CPU (148.7 tok/s prefill) behind per-token CUDA GATE: a matmul weight has an unsupported type spam.

Enablement (type-symmetric, no dispatch changes): gate whitelists extended with Q5_0 via matmul_t_ok/embed_t_ok helpers + a diagnostic that names the offending weight, type, and reason; cuda.rs gains the embed_rows_on_gpu arm (type_id 6) and a launch_q5_0_f32_matmul dispatch arm; cuda_kernels.cu gains embed_rows_q5_0 and q5_0_f32_matmul (f32-activation legacy structure, warp-per-4-rows). Prefill GEMM/MMQ reuse the existing type-agnostic path. The fault that mattered: the first E2E run died in a sticky cudaErrorMisalignedAddress (716) cascade — allocs/launches/syncs failing from the first prefill matmul onward, then a null-device-pointer panic at decode. Root cause: the Q5_0 block is 22 bytes, so the qh word at block offset 2 is NOT 4-byte aligned for even block indices. Fix: both kernels load qh as two 2-byte-aligned uint16_t loads; no shared kernel touched. (Same root-cause family as 8p's latent dequant_q5_0_f16 fix — doc 05.)

4. Verification

  • 8b: f16 roundtrip parity vs a half-rounded-KV reference (1e-4, kernel isolated from quantization noise); 7B 96-token greedy text identical f16 vs f32. cuda 154/0 after 8d.
  • 8c: cuda_q4_0_prefill_q8_0_gemm_parity (nt>1 vs the kernel's exact math; nt=1 f32 path vs hand dequant); cuda_matmul_parity's q4_0 arm mirrors activation quantization; greedy text unchanged.
  • 8d: standalone A/B nkv 440 +49% / 2000 +64% / 8000 +80% (maxdiff 1e-8); parity vs cpu_gqa_attn 1e-4 (empty + partial splits, both KV layouts); 7B E2E @2K greedy identical; @440 neutral (3 pairs within noise).
  • 8f: parity test (q5_1 id 64; q5_K id 896 tail, decode-formula weights over real unpack_q4k_scales, 5e-3); 0.5B q5_k_m E2E greedy identical to CPU; cuda 155/0.
  • 8l: llama-bench -r 3 × FA 0/1 + llama-cli cross-check on the llama side; minfer 3 reps --greedy; same GGUF files both sides.
  • 8q: cuda_q5_0_realshape_isolation (new device test — embed shape-bisect + legacy/f16/MMQ/decode matmuls at the model's exact shapes, with a real state.sync() after each step because Backend::synchronize does not wait on the stream outside capture windows); full suite 173; E2E clean 3/3 runs.

5. Results

itemheadline number
8b KV f167B @2K decode +~11% (swap-order pairs 10.3/10.2 vs 9.4/9.0 tok/s); win grows with context
8c shaped Q8_0 GEMM0.5B prefill @3.6K 1005 → 1246 tok/s (+24%); −63% negative at 7B ffn_down kept out by the gate
8d split-K attention7B E2E @2K decode 10.1 → 13.7 tok/s (+36%); @440 neutral
8f Q5_K/Q5_10.5B q5_k_m admitted to CUDA (was CPU wholesale), greedy identical
8l parity baselinedecode 1.15–1.76×, prefill 18–110× vs llama — the campaign's target sheet
8q Q5_00.5B q4_k_m prefill 148.7 → ~1200 tok/s, decode 56.9 → ~306 tok/s (CPU fallback eliminated)

6. Lessons

  1. Measure first, then wire (8c): the shape gate is load-bearing — the same kernel is +38–44% at one shape and −63% at another.
  2. Coverage is performance (8f/8q): a missing weight-type kernel silently costs a whole model its GPU; the "CUDA GATE" spam is the tell.
  3. Benchmark the competitor before optimizing (8l): per-shape gap numbers rank levers better than any profile of your own code.
  4. Respect the alignment contract (8q): uint32_t loads on non-4-aligned offsets are illegal and fault nondeterministically on unified memory — small-shape tests passing proves nothing; test the REAL shapes.
  5. Attention structure evolves in generations (8d → R4 → D3a/D3-4): 8d's split-K removed the warp-count bottleneck but kept a local-memory accumulator; each rewrite removed the bottleneck the previous one created.

← 78 · Index

07 · R3 — Small-model per-token overhead: single-split prefill (LANDED)

Result: of 0.5B decode's ~4.0 ms/token, ~2.4 ms of non-GPU overhead was located and dismantled item by item — prefill collapsed from 4 splits into a single CUDA split (A1), the D2H logits readback moved to a self-held pinned buffer and the per-step clone was cut (A2), prefill capture flipped default-on (B); greedy output bit-identical throughout. Bench recorded under ~96% co-tenant load: the pinned path parity to slightly ahead. Commit: 029a9a4 (A1) · a213c89 (A2) · 761e236 (B). Date: 2026-08-31.

1. Background — where things stood

The state at the close of Era A (Phase 7/8): 7B @2K prefill reached ~1204 via 8m's wmma GEMM and was pushed to ~1400–1500 tok/s by 8p's resident f16 weight cache; on the decode side, 8e's MMVQ (dp4a × q8_0) gave q4_K +37%, and 8o cut decode-start's 635 ms heavy clone to 35 ms. The big-model numbers were moving, but the 0.5B small model exposed another wall: the GPU accounts for only a minor share of each token's wall time. The MINFER_GRAPH_TRACE + DOT dump profile shows 0.5B decode at ~4.0 ms/token, of which the GPU floor is only ~1.6 ms — the remaining ~2.4 ms is CPU/synchronization overhead. 1000/4.0 ≈ 250 tok/s, matching Part-I's recorded 0.5B decode ~257 tok/s (llama 453) exactly: no matter how much faster the kernels get, this overhead caps them in place.

Small models are a magnifying glass for this wall, for a direct reason: GPU kernel time scales with model size, while the per-step fixed costs — split-boundary syncs, host↔device round trips, the driver bounce of pageable readbacks, redundant clones, the launch structure — do not scale. The smaller the model, the larger the fixed costs' share; on 0.5B it is 60% of the wall. The same fixed costs also surface on large models with short prompts (small nt, thin GPU work), so this is not a "small-model-only" corner case — it is the graph-execution architecture's bill.

Why fix it now: first, the graph architecture's core claims (declarative build → single-pass execute, CUDA Graph capture/replay) must hold at every scale, and a 4-split prefill on 0.5B plainly violates the "single-pass" promise; second, R3-B (prefill capture default-on) presupposes that prefill is one continuous split — with a host fill point in the middle of the split, capture can only wrap the body segment. Without this step, all later "whole-graph replay" dividends (launch-overhead elimination in server scenarios) have no footing.

The trace gave three targets, ordered "structural → constant → policy":

  • A1: the G3 tail-row reduction's input tail_ids was declared mid-graph (right before its consumers in the last layer), slicing every prefill forward into 4 splits (inputs | body | tail_ids | tail) — each forward pays 2 extra full-stream syncs + host round-trip copies;
  • A2: every decode step's logits readback goes through a blocking cudaMemcpy into a PAGEABLE Vec (the driver first bounces into its own pinned bounce buffer), and forward_graph then .to_vec()-clones logits that are already exactly n_out*nv;
  • B: 8g② had made prefill capture a deliberate opt-in (default off), so the server's repeated same-shape prefills got no benefit.

2. Principle — the GPU mechanism

Why a mid-graph input = a split. The scheduler executes nodes in build order (topological order); when it meets an input node that needs host filling, it must first let all queued async work on the stream complete (full-stream sync), copy the data from host into that input buffer, and only then continue encoding later nodes — this "stop, sync, fill, continue" point is a CPU/CUDA boundary. When inputs are declared at the graph head, filling happens before execution begins, none at all; declared between two CUDA ops, every such input adds one more cut to the forward. The pre-R3 prefill graph was:

[inputs | body | tail_ids | tail section]   ← 4 splits
          ↑ 2 extra full-stream syncs + host round-trip copies per forward

The crux: node order is not semantics — tail_ids's consumers merely reference the handle; declaring it at the graph head and consuming it at the tail leaves the dataflow completely unchanged while removing the mid-execution fill point from scheduling. That is what "input declaration position is graph topology" means: declaration position decides split boundaries even when semantics are equivalent.

The cost arithmetic: 0.5B's GPU floor is ~1.6 ms/step, and each of the two full-stream syncs drains the entire pipeline — the CPU wakes up, does bookkeeping, fills data, re-encodes; that chain is ms-scale on a machine with a busy co-tenant. The extra syncs are the same order of magnitude as the entire GPU step — the largest single source of the 2.4 ms overhead on small models (on the prefill side).

Why a pageable D2H readback is slow. For a blocking cudaMemcpy into pageable (plain malloc) memory, the driver has no stable bus address available, so it must first use its internal pinned bounce buffer as intermediary: device → driver-internal pinned buffer → copy into the caller's pageable destination — two DMA hops plus a possible staging allocation each time. 7e⑥ already paid this tuition on the H2D direction (write_input_async) and switched to a self-held pinned ring; the D2H direction never got the same treatment — the logits readback is precisely the D2H that runs every decode step. After switching to a cudaHostAlloc self-held buffer, the destination address is DMA-reachable, one hop direct, and reading out afterwards is just an ordinary CPU memory copy. The clone is cut in passing: the graph output is already exactly n_out*nv (reduced by G3, or n_out==nt), so the 608 KB (151936 vocab × f32) .to_vec() per step is pure waste.

Why capture has a 3-run protocol. CUDA Graph's benefit model: a one-time cost (capture + instantiate, ms-scale) buys the per-launch CPU cost down to zero (one cudaGraphLaunch for the whole graph). A 437-node prefill graph's per-launch CPU cost is substantial, but for a CLI prefill that runs only once, capture is a pure loss. The 3-run protocol: capture only when the same (uid, nt) graph appears a 3rd time — one-shot calls never reach 3 and pay nothing; repeated same-shape prefills like a server slot start netting a profit after the 3rd. A1's single split is the precondition: a capture window must contain no host fill point, otherwise it can only wrap the body segment (exactly the crippled form before 8g②).

3. Implementation

3.1 Design choices (why this shape and not another)

A1 — move the input to the head rather than eliminating the input. tail_ids carries G3's tail-row reduction (get_rows the n_out rows before the last layer's FFN so ffn/lm_head compute only the output rows — llama's ggml_get_rows(cur/inpSA, inp_out_ids) semantics); the data dependency itself must stay. The chosen form is .then(|| b.input(...)) under the params.n_out < nt condition: the condition is identical to before, and when n_out == nt the input node simply does not exist — the graph's determinism is decided by params (part of the reuse identity), and the declaration-position change does not touch that invariant. The decode graph never had tail_ids (nt==1 does not trigger the tail-row reduction), so the decode graph changed not at all.

A2 — a blocking memcpy into self-held pinned, not a switch to async. The readback point's semantics are "I want the result now": the caller has already sync()ed the stream. The gain comes from skipping the driver's bounce, not from overlap — so the shape is the simplest blocking memcpy + pinned destination, not an async chain with events/callbacks. The buffer is grow-on-demand: the first read allocates dst.len().max(4 MiB) (4 MiB of headroom prevents small size jitter from churning alloc/free), and later only growth triggers reallocation; if cudaHostAlloc fails, silently fall back to the pageable path and warn once. MINFER_NO_PINNED_READBACK=1 keeps the A/B switch. The clone cut is defensive: forward_graph returns the readback buffer directly only when its length already equals n_out*nv — "always true", but a shape check is kept rather than betting on structure.

B — flip the default rather than add a new mechanism. 8g②'s verification assets (the pp16/pp300 bit-parity harness) already existed; all R3-B did was flip the switch's default: MINFER_CAPTURE_PREFILL=1 is redundant but still accepted, MINFER_NO_PREFILL_CAPTURE=1 restores the old default. No new protocol, no change to the 3-run threshold.

3.2 Key code

A1: the input declaration moves from mid-graph to the graph head (src/models/qwen2/graph.rs, commit 029a9a4). Before — the input declared right beside its consumers, inserted before the last layer's FFN:

#![allow(unused)]
fn main() {
let wo = b.matmul(attn_out, l.wo.as_ref().unwrap(), None);
let is_last = il == model.layers.len() - 1;
if is_last && params.n_out < nt {
    // G3: reduce to the tail n_out rows BEFORE the last layer's FFN
    // (llama `ggml_get_rows(cur/inpSA, inp_out_ids)` at
    // qwen2.cpp:106-108) — ffn_norm, gate/up/down, swiglu, both
    // residuals and lm_head all run on n_out rows only.
    let tail_ids = b.input(                 // ← the input node appears mid-graph:
        "tail_ids",                         //    every CUDA op before it
        [params.n_out, 1, 1, 1],            //    gets cut by this boundary
        crate::graph::DType::I32,
    );
    let cur_tail = b.get_rows(wo, tail_ids, [ne, params.n_out, 1, 1]);
    let res_tail = b.get_rows(residual, tail_ids, [ne, params.n_out, 1, 1]);
    h = b.add(res_tail, cur_tail);
}

After — the same input declared at the graph head beside token_ids/positions, condition unchanged:

#![allow(unused)]
fn main() {
let inp_ids = b.input("token_ids", [nt, 1, 1, 1], crate::graph::DType::I32);
let inp_pos = b.input("positions", [nt, 1, 1, 1], crate::graph::DType::I32);
// G3 tail-row reduction input, declared at the graph HEAD (not beside
// its consumers at the last layer): an input node mid-graph splits the
// forward into extra CPU/CUDA boundaries (2 full-stream syncs + host
// round-trip copies per step on the split path). R3-A1.
// Node order is not semantics — the consumers below just reference the handle.
let tail_ids = (params.n_out < nt).then(|| {          // condition verbatim-identical to before
    b.input(
        "tail_ids",
        [params.n_out, 1, 1, 1],
        crate::graph::DType::I32,
    )
});
}

The consumption site changes by one line — unwrap the handle from the Option; the reduction logic untouched:

#![allow(unused)]
fn main() {
if is_last && params.n_out < nt {
    let tail_ids = tail_ids.expect("tail_ids input declared when n_out < nt");
    let cur_tail = b.get_rows(wo, tail_ids, [ne, params.n_out, 1, 1]);
    let res_tail = b.get_rows(residual, tail_ids, [ne, params.n_out, 1, 1]);
    h = b.add(res_tail, cur_tail);
}

A2: the pinned D2H readback (src/cuda.rs, commit a213c89). Self-held buffer + grow-on-demand + failure fallback; the core path:

#![allow(unused)]
fn main() {
pub fn copy_from_device_pinned(&self, src: *const std::ffi::c_void, dst: &mut [u8]) {
    static FALLBACK_WARNED: std::sync::atomic::AtomicBool =
        std::sync::atomic::AtomicBool::new(false);
    if std::env::var("MINFER_NO_PINNED_READBACK").as_deref() == Ok("1") {
        self.copy_from_device(src, dst);            // A/B switch: the old pageable path
        return;
    }
    // headroom so small size changes don't churn the allocation
    let need = dst.len().max(4 * 1024 * 1024);      // ≥4 MiB, guards against small-size jitter
    let mut guard = self.readback.lock().unwrap();
    if guard.as_ref().map_or(true, |b| b.bytes < need) {
        if let Some(old) = guard.take() {
            drop(old); // cudaFreeHost
        }
        let mut p: *mut std::ffi::c_void = std::ptr::null_mut();
        let err = unsafe { cudaHostAlloc(&mut p, need, 0) };
        if err != 0 {                               // alloc failed → pageable fallback,
            drop(guard);                            // warn only once
            if !FALLBACK_WARNED.swap(true, std::sync::atomic::Ordering::Relaxed) {
                eprintln!(
                    "CUDA: pinned readback alloc failed (err {err}); pageable D2H fallback"
                );
            }
            self.copy_from_device(src, dst);
            return;
        }
        *guard = Some(PinnedBuf { ptr: p as *mut u8, bytes: need });
    }
    let buf = guard.as_mut().unwrap();
    unsafe {
        cudaMemcpy(                                 // destination is pinned: one DMA hop,
            buf.ptr as *mut std::ffi::c_void,       // no driver-internal bounce
            src,
            dst.len(),
            CUDA_MEMCPY_DEVICE_TO_HOST,
        );
        std::ptr::copy_nonoverlapping(buf.ptr, dst.as_mut_ptr(), dst.len());
    }                                               // reading out is just an ordinary CPU memcpy
}
}

The backend-side call site (copy_to_host in src/graph/cuda_backend.rs); the clone cut pairs with it:

#![allow(unused)]
fn main() {
// R3-A2: read through the pinned staging buffer (pageable-memcpy
// bounce removed); MINFER_NO_PINNED_READBACK=1 reverts.
self.state.copy_from_device_pinned(b.ptr, dst);
}

B: the capture default flip (src/graph/cuda_backend.rs, commit 761e236):

#![allow(unused)]
fn main() {
-        let prefill_capture = std::env::var("MINFER_CAPTURE_PREFILL").as_deref() == Ok("1");
+        let prefill_capture = std::env::var("MINFER_NO_PREFILL_CAPTURE").as_deref() != Ok("1");
}

The 3-run protocol's gate itself is unchanged; only prefill_capture's initial value changes:

#![allow(unused)]
fn main() {
if *runs >= 3
    && self.capturing.is_none()
    && nt_hint.map_or(true, |nt| nt == 1 || self.prefill_capture)
}

3.3 Pitfalls

  • Input nodes serve readability and pay in scheduling. G3 declared tail_ids beside its consumers at the time — it reads nicely in source, but an input node's declaration position is scheduling metadata. This class of pitfall has no compile-time signal whatsoever: the graph still topo-validates, execution is still correct — just slow.
  • Env vars are process-global, and the suite runs tests in parallel. R3-B changed 8g①'s "prefill never captures" negative test to drive the opt-out through set_prefill_capture_for_test(false) — had the test kept setting the env var, it would cross-contaminate other configurations inside the parallel test processes.
  • cudaHostAlloc can fail. The pinned pool is not an infinite resource; the failure path must silently fall back to pageable and warn only once (the FALLBACK_WARNED static bit), otherwise every decode step sprays a stderr line.
  • The measurement window was occupied by a co-tenant. This session's bench ran throughout under ~96% sglang utilization — absolute numbers are incomparable; only interleaved A/B on the same binary, gate on/gate off, is meaningful. This was an early rehearsal of r59b's later "absolute values across windows are incomparable" lesson.

4. Verification

  • The split trace (A1): the prefill forward's split count fell from 4 to 2 and CPU/CUDA boundaries from 2 to 1 (the commit's own words "prefill split trace 4 -> 2", "2 -> 1 boundaries"); the master table condenses it to "4 splits → 1 per prefill forward". What this gate proves is the structural claim itself — the split boundaries really disappeared.
  • Greedy bit-identity (verified separately for A1/A2/B, 0.5B q4_0, 48 tokens): moving the input, changing the readback path, and flipping the capture default must not change a single bit of output. Defends against "structural refactoring casually changing the math".
  • cuda_pinned_readback_roundtrip (A2): a 5.6 MB round trip, deliberately larger than the 4 MiB initial buffer — forcing out the grow-on-demand path. Defends the buffer growth logic and copy-out correctness.
  • The pp16/pp300 bit-parity harness (B, 8g②'s legacy): capture/replay vs direct launch compared bit-for-bit, covering pp16 and pp300 (~437 nodes, real prefill scale). Defends against "capture semantics missing some node class".
  • The 8g① negative test redirected (B): under the opt-out the prefill graph never captures even after 3+ runs; a new default-on test proves the flipped default from the other side.
  • Suite: 161 (A1) → 162 (A2) → 163 (B) all green, run bounded at 8 threads (sglang was serving on the shared box). Defends against cross-module regressions.
  • Interleaved A/B (A2): same binary, MINFER_NO_PINNED_READBACK toggled on/off alternately, 0.5B decode, ~96% co-tenant — the pinned path parity to slightly ahead. Defends against "taking a single sample as a conclusion in a contended environment".

5. Results

Three structural settlements, all LANDED:

  • A1: the prefill forward merged from 4 splits into a single CUDA split (decode already was one); each forward saves 2 full-stream syncs + host round-trip copies. It is the direct precondition of R3-B's whole-graph capture.
  • A2: every decode step's logits readback (608 KB at 0.5B/7B-class vocab) skips the driver-internal pinned bounce, and the same-size per-step clone is cut. The master table's numeric verdict: parity-to-slightly-ahead under load — this session had no quiet window, and the honest record is "not worse under contention, structural waste deterministically eliminated".
  • B: repeated same-shape prefills (server/slot scenarios) automatically capture/replay from the 3rd occurrence; one-shot CLI prefills never reach 3 and pay nothing. (r55 later measured whole-prefill capture on 7B at ≤ +0.1% with a capture-illegal malloc mid-window — prefill capture stays default-on, but its value scenario is repeated prefill, consistent with the judgment that followed.)

The small models' absolute level (Part-I records, pre-MMQ-campaign): 0.5B q4_0 decode ~257 tok/s (llama 453), prefill ~3020 (llama 30550). R3 located and partially removed the self-inflicted fixed overhead; the remaining small-model gap is structural in launch (the number of kernels per step and the launch chain) — territory of the later decode campaign (the D series), outside this step's scope.

6. Lessons

  1. Input declaration position is graph topology: an input declared mid-graph cuts execution into multiple segments — wherever the consumers are, inputs belong at the graph head; node order is not semantics, but it is scheduling.
  2. Trace first, then assume where the overhead is: of 4.0 ms/token, 2.4 ms is not GPU work — without the trace and DOT profiles, none of the three targets would have been found.
  3. D2H and H2D are a symmetric tax: the pageable readback's driver bounce is the same money as the H2D side; and after reading back, do not clone a buffer that is already exact.
  4. Land structural-by-necessity changes even when the window shows no big win: A1 alone has no pretty tok/s number, but it is the switch for B's whole-graph capture.

← 06 · decode MMVQ (8e) · Index · 08 →

08 · R1 — int8 MMQ prefill GEMM (opt-in): the parity-first strategy (LANDED)

Result: a self-built int8 tensor-core MMQ kernel (64×64×256 tile) behind MINFER_MMQ=1, parity fully green (8 types × 8 shapes, max diff < 1e-3; greedy 7B token-identical to the f16 path); performance 155 (co-tenant) / 412 (quiet) / 441 (r7–r8 window re-measure) tok/s vs the f16 path's 630–880 / 1460 — parity clean but ~8× slow (vs llama ~24 TMAC/s), a gap unattributable in that window (ncu refused by the device). It is the scaffold of the r7+ raw-line campaign: the 441 → 3590.8 = 8.1× campaign arc starts here. Commit: 40e97c9. Date: 2026-08-31.

1. Background — where things stood

After R3 located the small models' fixed overhead, prefill's main contradiction returned to the GEMM itself. At that point the engine's prefill weight path was 8p's resident f16 cache: all quantized weights are dequantized to f16 once at load and laid flat in VRAM (the ≥2 GB gate), and the GEMM reads the f16 panels on wmma. This path was working on 7B — by P5's end, @2K 2340–2370 tok/s, 1.43× vs llama-bench 3401 — but it carried two structural costs:

  1. The weight traffic starts out multiplied. An f16 panel costs 64 B per 32 elements; raw q4_0 blocks are 18 B, q8_0 34 B, K-quants 18–34 B per 32-k block — resident f16 makes weight DRAM traffic 2–4× the raw bytes, and the panel is the operand the GEMM re-reads over and over.
  2. Dequant is a one-time tax, but memory is a permanent tax. The 2–4 GB of extra resident plane buys only "reads fast", while for llama.cpp this plane does not exist at all — its MMQ (matrix-multiplication with quantized weights) feeds int8 straight to the tensor cores: weights stay raw nibbles, activations are quantized to q8, int8×int8 accumulates exactly in int32, and block scales are corrected per block outside the mma.

That llama.cpp's prefill is fast — this is a core link of it. For minfer to close the 1.43× gap there is no way around standing this pipeline up. But at the time llama.cpp's MMQ internals had not yet been dissected the way they are today (that dissection only later became docs/LLAMA-CPP-MMQ-ANALYSIS.md), so R1's shape was: implement llama's MMQ math in a self-built kernel on minfer's own 8p tile skeleton — not a line-by-line port. This step's strategic significance outweighs its performance significance: nail down the correctness of "int8 pipeline + per-block rescale" behind an opt-in switch first, so the later campaign (the raw-byte line from r7 on) can chase speed on a parity-green foundation.

Where things would stick without this step: the f16 path's weight-traffic floor locks the GEMM's byte budget; without a parity gate, any int8 attempt faces two unknowns at once — "is it right" and "is it fast" — and events soon proved the first version was destined to be slow.

2. Principle — the GPU mechanism

2.1 The int8 tensor-core mma (s8s8s32)

R1's workhorse instruction is mma.sync.aligned.m16n8k32.row.col.s32.s8.s8.s32: one warp-collective call completes an M=16 × N=8 × K=32 matrix-multiply fragment, A/B operands int8, accumulator int32. Three key properties:

  • Exactness. Integer multiply-add has no rounding — Σ w·q within one 32-k block is exact and associative in int32 (nibble grid |w|≤127 × activations |q|≤127; the in-block magnitude is far within range). This means all floating-point work reduces to the block-scale correction multiplies, and the parity tolerance only needs to cover the ordering differences of those few f32 multiply-adds, not the dot product itself. That is the mathematical basis for daring to set the later parity gate at 1e-3.
  • Warp collectivity: there is profit only at M≥16. mma.m16n8k32's fragment layout is fixed: A's 16×32 is spread across the 32 lanes with row = groupID (lane/4) and k = the quad; C is 4 int32 registers per lane. M=16 is the instruction's hard size — run a decode-shaped nt=1 problem on it and 15/16 of the rows are padding, with the per-instruction cost undiminished. This exactly explains the engine's two-tier dispatch: decode takes MMVQ (dp4a per-thread dot products, the scheme 8e landed), and only prefill (large nt) takes MMQ's mma pipeline. (The D series later quantified this as "MMQ-at-M=1 = 0.14-wave collapse", with nt=4 the crossover where BT-MMQ pays — but the geometric reason was already written into the kernel comments at R1 time.)
  • The throughput gear. On Ampere-class and later SMs, the int8 tensor pipe's per-cycle MAC count is on the order of 4× f16's; stacked on "weights stay raw bytes", MMQ earns both the compute gear and the traffic at once.

R1's tile family is 64×64×256: block tile 64 tokens × 64 od rows (MMQ_BI=64, MMQ_BJ=64), staging 8 32-k chunks at once along k (MMQ_KD=8, i.e. 256-k depth, llama.cpp ITER_K style); warp tile 32 tokens × 16 rows, 8 warps (wm = warp>>1 covering 4×16 rows, wn = warp&1 covering 2×32 tokens) exactly tiling 64×64. The shared-memory ledger:

qa:  2 × KD × BI×WS int  = 2×8×64×9 ×4B = 36,864 B   ← A fragments (WS=9: 8 data + 1 pad)
qb:  same                  = 36,864 B                ← B fragments
ssa: 2 × KD × BI int      =  4,096 B                ← activation block int sums (for the rank-1 term)
sda/sds/sds1/sdm: 4 × 2×KD×BJ f32 = 16,384 B        ← block scales (q6_K dual sub-scales)
─────────────────────────────────────────────────
Total ≈ 94,208 B ≈ 94 KB → 1 block/SM on sm_121 (GB10)

2.2 The q8_1-style activation pipeline

llama.cpp's MMQ quantizes activations into q8_1: each 32-element block stores 32 int8s + f16 scale d + one int32 in-block integer sum s. Why a sum beyond the dot product? Because min-carrying weight types (q4_1/q5_1/K-quant) have values w_i = ds·nib_i − dmin, so:

Σ w_i·q_i = ds·Σ nib_i·q_i − dmin·Σ q_i
            └── acc computed by int mma ──┘   └── dmin × activation block sum (rank-1 term)

The activation block sum the rank-1 term needs must be produced in the same pass as quantization (summing in passing during quantization is the cheapest). R1's q8 pipeline is an isomorphic implementation of llama's q8_1: activations are quantized once per launch into pad40 blocks (40 B per 32-element block), and the 4 slack bytes written at offset 36 are exactly the block's int sum — during staging that one word travels into shared memory together with the quants (the kernel's ssa[]), zero extra reads outside the mma. This was also the design prototype later contrasted with llama's quantize_mmq_q8_1 in the r34 quantize-transpose prepass; llama's own step of folding this quantization into the GEMM prologue (the q8_1 pipeline) remained on record as a "step-function next" all the way to the campaign's close.

K-quant nibbles stay UNSIGNED (llama's unpack_scales trick): the grid is 0..15 or 0..31, the integer part is non-negative, and the scale/offset pair (d·s, −dmin) is applied per 32-k sub-block; q6_K's scales live on 16-element sub-blocks and one 32-k chunk spans two → each chunk runs two m16n8k16s with separate int accumulators, multiplied by (d·sc0, d·sc1) respectively.

2.3 Why this shape was "destined to be slow first but worth making correct first"

The bandwidth ledger: A side 40 B per chunk per token (pad40 q8), B side 16–34 B per chunk per row of raw weight bytes, against the f16 cache path's 64 B/row — weight traffic cut 2–4× outright. But R1's staging is synchronous (cp.async cannot carry quantized bytes — quantization must happen before the move), so each chunk's load latency is exposed directly on the critical path — the kernel comments recorded the numbers: staging one 32-k chunk at a time gives 2.5 TMAC/s, and batching 8 chunks amortizes it 8×. This "staging depth vs exposed latency" contradiction is the target the whole later raw line (r12/r14/r20/r34) would shoot at; R1 stood the target up rather than breaking through it on the spot.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Self-built tile, ported math. Not a line-by-line transplant of llama.cpp: the tile skeleton reuses 8p's wmma GEMM structure (8 warps, consecutive blocks sharing one od-tile weight panel's L2 locality), and the mma fragment layout strictly follows the PTX ISA documentation (C's get_i/get_j mapping uses the set llama.cpp has production-verified). The reasoning: porting the math contract (per-block rescale, unsigned nibbles, q6_K dual accumulators) gets parity; porting the code would drag llama's launch/stream-k structure in with it, and that part was not yet dissected at the time.
  • Weights stay RAW. No dequant, no w16 cache — the B-side staging unpacks nibbles straight into shared memory. This is the fundamental fork from the 8p route, and the source of the traffic advantage.
  • Types as templates. Three axes, <int TYPE, int KSPLIT, bool HAS_OFF>: TYPE 0–7 covers all 8 supported quant types; KSPLIT=1 (one m16n8k32 per chunk, types 0–6) vs KSPLIT=2 (two m16n8k16s + separate accumulators, q6_K only); HAS_OFF = min-carrying types take the rank-1 term.
  • KD=8 vs KD=4: ~94 KB deep staging (1 block/SM) measured faster than KD=4 (2 blocks/SM) — under co-tenant load, depth beat occupancy (a conclusion r5–r6 would flip once more; see later).
  • Landed as opt-in. Enabled only with MINFER_MMQ=1; the default stays f16 — in this state MMQ is ~3.5× slower, and parity-first does not mean a default switch. mma.m16n8k32 does not exist on sm_75, so compile-time __CUDA_ARCH__ >= 800 falls straight back to the f16 path.
  • q6_K does not use a per-16 loop. The naive cut of mma per 16-element sub-block is 4× slower end to end — q6_K carries 7B's ffn_down + lm_head, far too much of the wall. k32 staging + dual m16n8k16 preserves full throughput.

3.2 Key code

The design contract (src/cuda_kernels.cu, commit 40e97c9, excerpt):

// ─── R1: int8 MMQ prefill GEMM ────────────────────────────────────────
// llama.cpp's MMQ math structure on minfer's 8p tile skeleton. Activations
// are quantized to q8_0 once per launch (pad40 blocks; the kernel writes the
// per-block int sum into the 4 slack bytes at offset 36), weights stay RAW —
// no f16 dequant pass, no w16 cache. A tiled mma.m16n8k32 (s8) GEMM
// accumulates one 32-k int chunk per step; the int C fragment is rescaled
// per (token, row, k-block) with the block-scale products and, for
// min-carrying types, a rank-1 offset term (weight min × activation block
// sum) — exactly llama.cpp's per-block correction, so results sit within
// f32 rounding of the CPU q8_0-activation dot path.
//
// K-quants keep their nibbles UNSIGNED in the int GEMM and carry the min
// term separately (llama.cpp's unpack_scales trick — the nibble grid is
// 0..15/0..31, so the "integer part" is non-negative and the scale/offset
// pair (d·s, −dmin·m) is applied per 32-k sub-block). q6_K's scales live on
// 16-element sub-blocks: each 32-k chunk spans two of them, so the chunk
// runs as TWO m16n8k16 mmas (low/high k halves) with separate int
// accumulators, rescaled by (d·sc0, d·sc1).

The int8 mma inlines — note the accumulator C folded in place ("+r"(d[i])), 4 int32s per lane:

__device__ __forceinline__ void mmq_mma_k32(int* d, const int* a, const int* b) {
    asm volatile(
        "mma.sync.aligned.m16n8k32.row.col.s32.s8.s8.s32 "
        "{%0,%1,%2,%3}, {%4,%5,%6,%7}, {%8,%9}, {%0,%1,%2,%3};\n"
        : "+r"(d[0]), "+r"(d[1]), "+r"(d[2]), "+r"(d[3])
        : "r"(a[0]), "r"(a[1]), "r"(a[2]), "r"(a[3]), "r"(b[0]), "r"(b[1]));
}

__device__ __forceinline__ void mmq_mma_k16(int* d, const int* a, int b) {
    asm volatile(
        "mma.sync.aligned.m16n8k16.row.col.s32.s8.s8.s32 "
        "{%0,%1,%2,%3}, {%4,%5}, {%6}, {%0,%1,%2,%3};\n"
        : "+r"(d[0]), "+r"(d[1]), "+r"(d[2]), "+r"(d[3])
        : "r"(a[0]), "r"(a[1]), "r"(b));
}

The main loop skeleton (mmq_nt_kernel, excerpt) — smem layout, double buffering, fragment assembly, rescale:

template <int TYPE, int KSPLIT, bool HAS_OFF>
__global__ void __launch_bounds__(256) mmq_nt_kernel(
    const uint8_t* __restrict__ W, const uint8_t* __restrict__ q8x,
    float* __restrict__ C, int nt, int od, int id, int bstride
) {
#if __CUDA_ARCH__ >= 800   // mma.m16n8k32 (s8) needs sm_80+; sm_75 keeps f16
    extern __shared__ int mmq_sh[];
    int* qa  = mmq_sh;                               // [2][KD][BI*WS]
    int* qb  = qa + 2 * MMQ_KD * (MMQ_BI * MMQ_WS);  // [2][KD][BJ*WS]
    int* ssa = qb + 2 * MMQ_KD * (MMQ_BJ * MMQ_WS);  // [2][KD][BI] activation block sums
    float* sda  = reinterpret_cast<float*>(ssa + 2 * MMQ_KD * MMQ_BI);
    float* sds  = sda + 2 * MMQ_KD * MMQ_BI;         // weight scales [2][KD][BJ]
    float* sds1 = sds  + 2 * MMQ_KD * MMQ_BI;        // q6_K's second 16-sub
    float* sdm  = sds1 + 2 * MMQ_KD * MMQ_BI;        // min (the rank-1 term)

    float sum[16] = {0.0f};   // [nh][h][l]: 2 B-frags × 2 A-frags × 4 C regs
    // ... double buffer: stage block 0 → inside the loop, stage kt+1 before computing kt ...
    for (int kt = 0; kt < nktile; ++kt, buf ^= 1) {
        if (kt + 1 < nktile) MMQ_STAGE_TILE(kt + 1, buf ^ 1)
        for (int kd = 0; kd < MMQ_KD; kd++) {
            // A fragments: 2×16-token halves; B fragments: 2×8-row halves; 4 mmas per chunk
            int a[2][4], b[2][2];
            int clow[2][2][4], chigh[2][2][4];
            #pragma unroll
            for (int h = 0; h < 2; h++) {
                const int r0 = (i0w + h * 16 + (lane >> 2)) * MMQ_WS + (lane & 3);
                const int r1 = (i0w + h * 16 + 8 + (lane >> 2)) * MMQ_WS + (lane & 3);
                a[h][0] = qat[r0]; a[h][1] = qat[r1];
                a[h][2] = qat[r0 + 4]; a[h][3] = qat[r1 + 4];
            }
            #pragma unroll
            for (int nh = 0; nh < 2; nh++)
                #pragma unroll
                for (int h = 0; h < 2; h++) {
                    if constexpr (KSPLIT == 1) {
                        mmq_mma_k32(clow[nh][h], a[h], b[nh]);
                    } else {                          // q6_K: low/high k halves split into two
                        mmq_mma_k16(clow[nh][h], a[h], b[nh][0]);
                        mmq_mma_k16(chigh[nh][h], a[h] + 2, b[nh][1]);
                    }
                }
            // per (token, row, k-block) rescale: value = ds·int (+ dm·sa), all × da
            const float da_q[4] = { sdat[i0w + lane / 4], sdat[i0w + 8 + lane / 4],
                                    sdat[i0w + 16 + lane / 4], sdat[i0w + 24 + lane / 4] };
            // ... after the dsv/dmv reads: sum[idx] += da·(ds·acc + dm·sa) (when HAS_OFF)

q6_K's unsigned-nibble B fragment assembly — the 6-bit grid 0..63 subtracts 32 per byte into the signed domain; __vsubss4 does the 4-byte SIMD subtract in one instruction:

uint32_t nib = (g < 2) ? (QL & 0x0F0F0F0Fu) : ((QL >> 4) & 0x0F0F0F0Fu);
uint32_t hi  = ((QH >> (2 * g)) & 0x03030303u) << 4;
qb[r * MMQ_WS + w] = __vsubss4((int)(nib | hi), 0x20202020);  // −32/byte

The parity test's reference frame (current tree src/graph/cuda_backend.rs::cuda_prefill_mmq_parity — this gate lives on today, still the hard gate for every MMQ change). The test comment's own words give the tolerance rationale:

#![allow(unused)]
fn main() {
    // dot math (the structure llama.cpp's MMQ implements): int8×int8 dots
    // are exact on both sides and the block scales are f16→f32 on both
    // sides; only accumulation order differs, so 1e-3 absolute leaves
    // orders of magnitude of headroom over f32 rounding while still failing
    // loudly on any fragment-layout or unpacking mistake. All 8 types ×
    // {odd tile edges, 2 super-blocks}; q6_K in both registered layouts.
}

The reference implementation is the CPU-side per-32-block q8_0-activation dot product:

#![allow(unused)]
fn main() {
        // reference: CPU q8_0-activation dot math, per 32-block:
        //   out += da · (ds · Σ w_i·q_i + dm · Σ q_i)
        // q6_K carries 16-element sub-scales → two halves per 32-block.
        for b in 0..nb {
            let blk = &x[t * id + b * 32..t * id + b * 32 + 32];
            let am = blk.iter().fold(0f32, |m, v| m.max(v.abs()));
            let d = am / 127.0;
            da[b] = half::f16::from_f32(d).to_f32(); // f16 rounding, as the GPU kernel stores it
            let di = if d != 0.0 { 1.0 / d } else { 0.0 };
            for (i, v) in blk.iter().enumerate() {
                let qi = (*v * di).round_ties_even();
                qi = qi.clamp(-128.0, 127.0);
                q[b * 32 + i] = qi as i32;
                sa[b] += qi as i64;                  // block sum: the rank-1 term's reference
            }
        }
}

3.3 Pitfalls

  • Synchronous staging's exposed latency. Quantized bytes cannot ride cp.async (it moves but does not unpack), so the first version synchronized once per chunk, exposing global-load latency in full on the mma critical path — 2.5 TMAC/s. Batching 8 chunks (KD=8) amortized it 8× to reach usable. This pitfall defined the shape of the entire campaign that followed.
  • The q6_K per-16 loop is a 4× trap. Cutting the chunk into k16 along "the scales live on 16-element sub-blocks" doubles the mma count and halves the staging rhythm; end to end it is 4× slower. The correct shape is k32 staging + one m16n8k16 each for the low/high k halves, separate accumulators, each paired with its own (d·sc0, d·sc1).
  • The K-quant signed domain. Nibbles stay unsigned and min goes through the rank-1 term, rather than pre-converting nibbles to signed before the mma — the former keeps (d·s, −dmin) in the rescale to be applied per sub-block, and q6_K needs only one __vsubss4 SIMD subtract to move the 6-bit grid into the signed domain. This "B-frag contract" was later reused verbatim in r28's NB kernel.
  • The profiling channel was sealed. On GB10 in that window ncu reported ERR_NVGPUCTRPERM (device counter permission), so the ~8× performance gap had no counter evidence at all — only TMAC/s estimates. r13's counter forensics had to wait for the channel to be repaired; this is also one of the premises that made the "parity first, speed later" strategy sound: while the gap is unattributed, the only certainly-correct asset is parity itself.

4. Verification

  • cuda_prefill_mmq_parity: an 8-type × 8-shape sweep (odd tile edges, 2 super-blocks, q6_K in both registered layouts all covered), max diff < 1e-3. Defends against fragment-layout errors (misaligned lane→(row,col) mapping) and nibble-unpacking errors — the integer dots are exact on both sides, so 1e-3's entire margin belongs to the f32 rescale's ordering differences, and any "real error" fails loudly instead of hiding in noise.
  • Greedy 7B ≡ the f16 path, token for token: end-to-end comparison against the in-service f16 path. Defends against regressions in the integration surfaces beyond the kernel — dispatch, epilogue, type registration.
  • Suite all green: the existing CPU/Metal/f16-CUDA paths unbroken.
  • Interleaved A/B measurement: same binary, MINFER_MMQ=1/0 alternated — 155 vs 630–880 under co-tenant, 412 vs 1460 on a quiet machine; the r7–r8 window re-measured the same code at 441. Defends against "mistaking the co-tenant tax for a code tax" — this gate later matured into r59b's "baseline behavior anchoring" rule.

5. Results

Kernel level and wall-clock level (7B q4_k_m, DGX Spark GB10):

  • Performance: the MMQ path at 155 tok/s (sglang ~96% co-tenant) / 412 (quiet) / 441 (r7–r8 window re-measure); the f16 path in the same windows 630–880 / 1460. I.e. the opt-in state is ~3.5× slower, ~2.9 TMAC/s vs llama ~24 on the quiet baseline — a ~8× gap, unattributable in that window (ncu ERR_NVGPUCTRPERM). Fixed overhead ~0.6 ms per prefill (launch + quantization).
  • Shape conclusions: KD=8 (1 block/SM deep staging) beat KD=4 (2 blocks/SM) under load; the per-chunk latency exposure of synchronous staging was the largest single item (2.5 TMAC/s → KD=8 amortizes 8×).
  • The landing: MINFER_MMQ=1 opt-in, the default stays f16 (a 3.5× regression cannot be the default). The master table's status reads LANDED (opt-in; superseded by raw line) — from r7 on, the raw-byte line rebuilt the kernel on this scaffold (raw-byte smem, ldmatrix, 2 blocks/SM, the quantize-transpose prepass … until r60 flipped it default-on).
  • What it bought: the parity gate has been the hard gate for every MMQ change since R1 — none of the dozens of levers after r7 needed to reinvent a correctness standard; the campaign arc of master-table footnote 4, 441 → 3590.8 = 8.1×, starts at this row's 441. The slow R1 is the only step in the entire campaign where "performance did not matter", because what it delivered was not speed but a foundation on which speed could be chased with confidence.

6. Lessons

  1. A parity-clean opt-in lands even when slow: with the correctness scaffold in place first, the speed campaign has a stable right/wrong standard — the raw line's 8.1× walked on this gate the whole way.
  2. Re-measure the same code in a quiet window before concluding: 155 (co-tenant) → 412 (quiet) → 441 (re-measure); judge the environment before judging the code.
  3. A kernel you cannot profile can only be estimated: with the counter channel sealed, TMAC/s is the only clue; the counter forensics after the channel was repaired (r13) changed the direction of the entire campaign.
  4. Port the math contract, not the code: implementing llama's per-block rescale + unsigned-nibble contract on one's own tile skeleton gets parity; a line-by-line port drags the un-dissected structures in with it.

← 07 · R3 small-model overhead · Index · 09 →

09 · R2 — MMVQ weight-streaming rework (LANDED)

Result: 7B q4_k_m decode: tg128 42.2 → 45.1 tok/s (+6.9%), @2K 36.7 → 38.8 (+5.7%); the gap to llama.cpp (47.1 / 44.9) narrowed by about half in both cases. Commit: 6df3245. Date: 2026-08-31.

1. Background — where things stood

At the end of Era A the engine's decode (nt==1) ran on the MMVQ path 8e landed (see chapter 06): the activation vector is quantized to q8_0 before every matmul, weights stay in native GGUF format without dequant, the kernel does 4-byte-packed integer dot products with __dp4a, and the launch parameters copy llama.cpp's MMVQ_PARAMETERS table. That step lifted 7B q4_K decode by +37%, but measured against llama.cpp in the same window it still trailed:

  • tg128 (generation at a 128-token context): 42.2 vs 47.1 tok/s;
  • @2K (a 2048-token context): 36.7 vs 44.9 tok/s.

The gap is ~10% (short context) to ~18% (long context). The @2K portion comes mainly from split-attention (R4's story — chapter 10), while the tg128 gap falls almost entirely on the matmul kernel itself: each decode step's wall time is basically "stream the ~4.5 GB of weights through VRAM once", with attention and elementwise a mere rounding error.

The record's calibration of the 8e kernel: effective weight-stream rate ~147 GB/s versus llama's same-class kernel at ~197 GB/s — the record summarized it as "8e runs at only about 60% of llama's effective stream rate" (the per-kernel normalization basis for the two numbers was not preserved in the summary; the ratio itself defers to the session record). GB10's DRAM roofline is 273 GB/s (calibrated by r55's measurement), meaning minfer's decode kernels during their active period consume barely half of DRAM peak — far from the bandwidth ceiling, showing the bottleneck is not "the bytes cannot move" but "how the moving instructions are organized".

Where things would stick without this step: decode is the metric users feel most (generation speed), and its theoretical ceiling is decided solely by the weight-stream rate. As long as the MMVQ kernels' stream efficiency stays at ~150 GB/s, every later decode-side optimization (R4's attention, the D series' quantization folding) is pressed down by this foundation. R2 chose to raise the matmul kernels' stream efficiency before touching attention, because the attribution was clear: the tg128 gap is 100% matmul, and the matmul gap is structural (load instructions per byte) — not something parameter tuning can fix.

2. Principle — the GPU mechanism

Weight-streaming-bound decode. At nt==1 every matmul is a matrix-vector product y = W·x: producing one output row requires reading that row of the weights once and dotting it with x, which resides in registers/L1. Weight bytes are used exactly once per token — there is no k-fold inner-product reuse (that is the prefill GEMM's dividend), so decode's ideal speed is simply:

ideal tok/s = effective DRAM bandwidth ÷ weight bytes per token

7B q4_K_M's weights are ~4.5 GB; DRAM roofline 273 GB/s gives an ideal ceiling on the order of ~60 tok/s. Both engines measure far below it (llama 47.1, minfer 42.2), and the difference between them can only come from how efficiently each kernel streams weight bytes into the pipeline — how many instructions and memory transactions each byte costs.

Effective GB/s is the quantity defined on exactly this basis: weight bytes streamed per token × tok/s. It translates "tok/s" into the kernel's physical workload, freeing the minfer/llama comparison from implementation differences between model layers.

The v1 (8e) kernel's two wastes. A q4_K 256-element super-block contains 8 32-element sub-blocks; adjacent sub pairs (even/odd) share the same 32 B nibble chunk — the even sub takes the low 4 bits, the odd sub the high 4. v1's thread mapping was "one sub-block per thread":

  • Each sub's thread reads the whole 32 B chunk in via 8 4-byte loads, using only half the nibbles; the sibling sub's other thread reads the same 32 B again — every weight byte is touched by load instructions twice per row;
  • The second read hits L1 (the same sector), so DRAM bytes do not double, but the load-instruction count per byte doubles. The kernel at that point is issue/latency-bound, not bandwidth-bound; the instruction stream is stuffed with redundant loads and long scoreboard latency cannot be effectively spread;
  • q6_K is worse: the raw layout's block stride is 210 B (not a multiple of 4), so ql/qh can only be read 2 bytes at a time (llama.cpp get_int_b2 style) — one 16 B ql piece takes 8 loads.

v2's structure. The thread mapping changes from "one sub per thread" to "one sub-PAIR (64 elements) per thread": the pair's two subs share the nibble bytes, the chunk is read once via 2 16-byte uint4 loads and serves both subs' half-nibbles — exactly 1 load instruction per weight byte per row (with the uint4 widening, per-pair weight loads drop from 16 4B loads to 2 16B loads). On the q6_K side, 7e② had already padded the block stride to 224 B (14×16), making every ql/qh piece in a block 16 B aligned, so uint4 loads are legal. q5_K's qh plane is a 32 B shared by all 8 subs, reusable via bit indexing alone.

In one sentence: the bytes were not wasted (L1 caught them) — what was wasted was the instruction stream; v2 cuts the instruction stream back.

3. Implementation

3.1 Design choices (why this shape and not another)

One v2 kernel each for the three types (q4_K/q5_K/q6_K), coexisting with v1, selected by a dispatch gate:

  • The gate condition is id % 256 == 0 (whole super-blocks). v2's pair mapping assumes id divisible by 256 — shapes with a partial tail (e.g. id=2176 = 8.5 super-blocks) do not map and must stay on v1. This is precisely the root of the later pitfall (§3.3).
  • q6_K's v2 additionally requires the padded 224 B stride: the uint4 alignment was bought by the padded layout; the raw 210 B layout keeps v1.
  • MINFER_MMVQ_V1=1 forces v1 back — preserving the ability to A/B measure and regression-compare (a habit running through the whole campaign: every default switch keeps a "0"/"V1" opt-out).
  • The grid shape is unchanged (grid(od, nt), one 256-thread block per output row): v2 doubles each thread's work and halves thread coverage, but the reduce structure (warp shuffle + block tree) and llama's launch table are untouched — change the mapping, not the scheduling, so the A/B attributes to exactly one thing.

3.2 Key code

Before — v1 q4_K: one sub per thread, the whole chunk read with only half used (current tree src/cuda_kernels.cu, q4_k_q8_mmvq):

for (int u = threadIdx.x; u < nsub; u += 256) {   // u = 32-element sub-block
    const int blk_i = u >> 3, sub = u & 7;
    const uint8_t* blk = weights + (size_t)row * row_stride + blk_i * Q4KB;
    uint8_t s8, m8;
    get_scale_min_k4(sub, blk + 4, &s8, &m8);
    // sub-block nibbles: chunk (sub>>1) of 32B, lo nibbles for even sub,
    // hi for odd; element l of the sub-block ↔ byte l.
    const uint32_t* qw = reinterpret_cast<const uint32_t*>(blk + 16 + (sub >> 1) * 32);
    const bool lo = (sub & 1) == 0;
    ...
    #pragma unroll
    for (int v = 0; v < 8; v++) {                 // 8×4B = the whole 32B chunk
        const uint32_t w = qw[v];
        const int n = lo ? (int)(w & 0x0F0F0F0F) : (int)((w >> 4) & 0x0F0F0F0F);
        const int xa = (int)xw[v];
        dot = __dp4a(n, xa, dot);
        sx  = __dp4a(0x01010101, xa, sx);
    }
    acc += d8 * ((float)s8 * (float)d * (float)dot
               - (float)m8 * (float)dm * (float)sx);
}

Each sub thread issues 8 4B loads; the same chunk is read again verbatim by the sibling sub's thread — a 32 B chunk consumes 16 load instructions per row in total.

After — v2 q4_K: one sub-PAIR per thread, the chunk read once for both subs (commit 6df3245, also in the current tree, q4_k_q8_mmvq_v2):

for (int u = threadIdx.x; u < npair; u += 256) {  // u = 64-element sub-PAIR
    const int kbx = u >> 2, c = u & 3;
    const uint8_t* blk = weights + (size_t)row * row_stride + (size_t)kbx * Q4KB;
    const int s0 = 2 * c, s1 = 2 * c + 1;
    get_scale_min_k4(s0, blk + 4, &s8a, &m8a);
    get_scale_min_k4(s1, blk + 4, &s8b, &m8b);
    // one 32B nibble chunk: lo nibbles = sub s0's 32 elements, hi = sub s1's
    // (16B-aligned: 144·kbx + 16 + 32·c ≡ 0 mod 16)
    const uint4 w0 = *reinterpret_cast<const uint4*>(blk + 16 + c * 32);
    const uint4 w1 = *reinterpret_cast<const uint4*>(blk + 16 + c * 32 + 16);
    const uint32_t ws[8] = {w0.x, w0.y, w0.z, w0.w, w1.x, w1.y, w1.z, w1.w};
    ...
    #pragma unroll
    for (int v = 0; v < 8; v++) {
        const uint32_t wv = ws[v];
        const int xa_v = (int)xa[v], xb_v = (int)xb[v];
        dota = __dp4a((int)(wv & 0x0F0F0F0F), xa_v, dota);        // sub s0
        sxa  = __dp4a(0x01010101, xa_v, sxa);
        dotb = __dp4a((int)((wv >> 4) & 0x0F0F0F0F), xb_v, dotb); // sub s1
        sxb  = __dp4a(0x01010101, xb_v, sxb);
    }
    acc += d8a * ((float)s8a * d * (float)dota - (float)m8a * dm * (float)sxa)
         + d8b * ((float)s8b * d * (float)dotb - (float)m8b * dm * (float)sxb);
}

The same 32 B chunk: 2 16B loads, two __dp4a accumulators advancing in parallel. The low nibbles feed s0's dot directly, (wv >> 4) & 0x0F0F0F0F feeds s1 — two views of one datum, zero repeated loads. The alignment comment is a hard guarantee: 144·kbx + 16 + 32·c is always a multiple of 16 against q4_K's 144 B block header.

q6_K v2: the uint4 bought by the padded stride (q6_k_q8_mmvq_v2, abridged):

// v1 mapping with s = 2*pair + half: chunk = s>>3 = pair>>2,
// g = (s>>1)&3 = pair&3, is = s&1 = half (the pair's two subs share
// chunk/g; only the 16-byte is-half differs)
const int chunk = pair >> 2, g = pair & 3;
// padded 224B stride ⇒ every ql/qh piece is 16B aligned
const uint4 qla = *reinterpret_cast<const uint4*>(blk + chunk * 64 + (g & 1) * 32);
const uint4 qlb = *reinterpret_cast<const uint4*>(blk + chunk * 64 + (g & 1) * 32 + 16);
const uint4 qha = *reinterpret_cast<const uint4*>(blk + 128 + chunk * 32);
const uint4 qhb = *reinterpret_cast<const uint4*>(blk + 128 + chunk * 32 + 16);

Compare the v1 kernel's header-comment confession (the comment before the current tree's q6_k_q8_mmvq): "q6_K block strides are 210B raw / 224B padded (7e② repack) — both even but not 4-aligned, so the weight side reads 2-byte halves (llama.cpp get_int_b2 style)" — one 16 B ql piece costs 8 2B loads; v2 does the same mapping with 1 uint4.

q5_K v2: the qh plane shared via bit indexing (q5_k_q8_mmvq_v2, abridged):

// the qh plane is 32 bytes SHARED by all 8 sub-blocks (byte l holds
// one high bit per sub for element l) — every chunk reads the same
// bytes, only the bit index (s0/s1) differs
const uint4 h0 = *reinterpret_cast<const uint4*>(blk + 16);
const uint4 h1 = *reinterpret_cast<const uint4*>(blk + 16 + 16);
...
const uint32_t hia = (((qhv >> s0) & 0x01010101u) << 4);
const uint32_t hib = (((qhv >> s1) & 0x01010101u) << 4);
dota = __dp4a((int)((wv & 0x0F0F0F0F) | hia), xa_v, dota);

qh's byte layout is "byte l stores one high bit per sub for each of the 8 subs" — the same 32 B serves all chunks, only the bit index changes with the sub. Shift + mask, 3 ALU instructions, reinsert the high bit into the nibble, replacing v1's wrong shape of reading qh by chunk offset.

The dispatch gate (src/cuda.rs, the q4_K arm; shared as mmvq_v2):

#![allow(unused)]
fn main() {
fn mmvq_v2(id: usize) -> bool {
    id % 256 == 0 && !std::env::var("MINFER_MMVQ_V1").map_or(false, |v| v == "1")
}
}

The three type arms each pick between the v1/v2 launchers by this gate (the q6_K arm adds a blk_stride_padded condition).

3.3 Pitfalls

The gate condition hid the bugs for two weeks. The first v2 passed the first suite round carrying two bugs: q6_K's nibble-group computed wrong (written as g = pair >> 1, correct is g = pair & 3), and q5_K's qh offset computed per chunk (correct is the shared 32 B indexed by bit). Why it went uncaught: the parity/greedy shapes of the time used the non-square id=2176, and 2176 % 256 = 128 ≠ 0 — the dispatch gate sent it back to v1. The v2 code was never executed at all; the all-green tests were a fake green. Only after the id=2560 shape (10 whole super-blocks, takes v2) was added did the engine-level greedy check immediately emit garbled tokens, and the two bugs surfaced. The lesson was later written into the step-document rules: test shapes must cover every dispatch arm's gate condition, not just the kernel math.

The remaining details went relatively smoothly: v1/v2 accumulate in a different order (within a pair, the two subs each dp4a first, then add), but each output row's summation-tree structure is unchanged, and greedy output is identical (v1 ≡ v2 token-for-token) — no need to touch the tolerance gate.

4. Verification

  • Parity shape expansion: the id=2560 shape was added to the parity sweep, specifically targeting the id % 256 == 0 gate arm — defending against exactly §3.3's "gate routes the code around the test" fake green.
  • Engine-level greedy comparison: the whole engine runs greedy generation, v1 ≡ v2 token by token — a gate one level above kernel unit tests, catching mapping/dispatch errors kernel-level tests miss (both of this step's bugs were caught by it).
  • Full suite 164 passing: the regression safety net, confirming the change did not ripple into other types and paths.
  • Same-binary interleaved A/B: tg128 and @2K each measured under both a quiet-GPU window and sglang contention, medians taken — defending against machine-state drift reading noise as gain (+6.9% is the quiet window; the contention window kept a +5–8% relative gain, same direction).

5. Results

Metric (7B q4_k_m)before (8e)after (R2)Δllama.cpp same window
tg128 decode42.245.1+6.9%47.1 (gap ~10% → ~4%)
@2K decode36.738.8+5.7%44.9 (gap ~18% → ~14%)

Under the contention window (sglang on the same machine) the gain held at +5–8% relative. The remaining 14% gap at @2K is mostly split-attention (handled by the next step, R4); the tg128-side matmul gap narrowed to ~4%. At the kernel level, load instructions per weight byte per row halved (re-reads eliminated) and q6_K's ql/qh went from 8×2B to 1×16B, lifting the decode kernels' effective stream rate after v2 shipped — the record defers to whole-step tok/s and the two-window consistency, and does not list a separate kernel GB/s after value.

This commit touches 6 files: src/cuda_kernels.cu +208 lines (the three v2 kernels + launcher), src/cuda.rs dispatch wiring +131 lines, src/graph/cuda_backend.rs a 520-line refactor (v2 gate wiring), the rest documentation.

6. Lessons

  1. Test shapes must walk every dispatch arm's gate condition — id=2176 happened to land in the v1 gate, letting two v2 bugs pass all-green; adding one id=2560 shape was worth more than ten more lines of unit tests.
  2. When the bottleneck is the instruction stream rather than the bytes, first check "how many load instructions per byte": L1 will hide re-read bytes, but it cannot hide redundant instructions in the issue stream.
  3. Data-path granularity (how much weight one thread covers) is a first-class design axis for MMVQ-class kernels, of the same order as launch-table tuning; llama's parameter table gives the scheduling, but the bytes→threads mapping — the alignment and sharing relationships — you must derive yourself.
  4. Alignment is bought: q6_K's uint4-ization depends on 7e②'s 224 B padded layout; the dispatch gate must check both premises — "id divisible" and "stride already padded".

← 08 · R1 int8 MMQ prefill GEMM · Index · 10 →

10 · R4 — decode split-attention dim-parallel rewrite (LANDED)

Result: the gqa_attn_split_partial kernel 148 → 79 µs/layer (nsys); 7B @2K decode 39.2 → 43.2–45.1 tok/s (gap to llama's 44.9: 14% → ~4%); tg128 45.1 → 47.5–47.6 (overtaking llama's 47.1). Commit: 70f57db. Date: 2026-09-01.

1. Background — where things stood

After R2 raised the decode matmuls' weight-stream efficiency (chapter 09), @2K decode stalled at 38.8 tok/s, still ~14% behind llama's 44.9. The nsys attribution was very clean: the gap sat almost entirely on the split-attention kernel gqa_attn_split_partial — about 150 µs per layer, 28 layers ≈ 4.2 ms, one sixth of the ~25.8 ms per step at 7B @2K. And tg128 (short context) had already reached 45.1 after R2, showing the matmuls were not the problem: the attention kernel's time grows linearly with KV length — the @2K gap is exactly it.

This kernel is the flash-decoding path introduced by 8d (Phase 8's decode attention revision): at nt==1 each query head has a single query row yet must attend over the whole context's K/V — 28 heads' parallelism alone cannot fill the GPU, so the KV axis is cut into several splits, each split independently computes an online-softmax partial sum, and a combine kernel merges them. 8d's version fixed the split count at 8: 7B is 28 head × 8 split = 224 single-warp blocks.

The nsys readings of the disease (the record's own numbers):

  • ~150 µs per layer against a single 4.3 MB K+V read (= 2 × nkv 2048 × 4 kv-heads × hd 128 × 2 B f16) works out to an effective stream rate of 28 GB/s — under 11% of DRAM peak;
  • the ACC accumulator is a runtime-indexed float4 oc[32] living in LOCAL memory (~80 MB/layer of re-read/rewrite traffic);
  • lanes walk rows with 4-byte loads, 64 scattered sectors per row, 12.5% sector efficiency;
  • the whole kernel is only 224 single-warp blocks, occupancy naturally poor.

Worse, the directional evidence: sweeping the split count up from 8, the kernel gets monotonically worse — 148/172/419/609 µs @ SPLITS 8/16/32/64. Doubling parallelism makes it slower, meaning the bottleneck is not "too few blocks" but the per-warp geometry itself being broken: more splits just copy the same broken geometry more times, and LOCAL memory traffic floods L1. That is why this step does not "raise SPLITS" — it rewrites the whole dims→lane mapping.

2. Principle — the GPU mechanism

2.1 The structure of split-KV decode attention (why it must be two-stage)

Decode attention computes o = softmax(q·Kᵀ·scale)·V, with nkv the current context length. At nt==1, without splitting, the parallel units are only the nh heads — 28 warps on 7B, not enough to fill GB10 (48 warp slots per SM). Split-KV (flash-decoding) cuts the KV dimension into ATTN_SPLITS pieces:

  • The partial kernel: each (split, head) block runs online softmax over its own row range [lo, hi), maintaining the triple (mx, S, oc): mx is the max score seen so far, S the sum of exp weights, oc the weighted V accumulator. Each incoming row: nmx = max(mx, s), corr = exp(mx − nmx), S = S·corr + e, oc = oc·corr + e·v — numerically stable, and the full score matrix never needs to materialize; (mx, S, oc) is written into the partial buffer.
  • The combine kernel: first take gmx = max(mx_sp) over all splits, then re-weight and sum each split's S and oc by w_sp = exp(mx_sp − gmx), and o = acc / S. Mathematically equivalent to one softmax over the whole context — the merge weights exp(mx_sp − gmx) exactly correct the scale differences caused by each split's differing local maxima.

2.2 nkv independence: the hard precondition of graph replay

Decode's whole-step compute graph is CUDA-Graph captured and replayed repeatedly (7d). Grid dimensions are frozen into the graph at capture and replay cannot change them; meanwhile nkv grows every token. So split-attention's grid must be a static shape like dim3(ATTN_SPLITS, n_head), and nkv can only be read at kernel runtime as data:

const int nkv0 = positions[0] + 1;                    // read from device memory at runtime
const int chunk0 = (nkv0 + ATTN_SPLITS - 1) / ATTN_SPLITS;

When the context is short, most splits get no rows (lo ≥ nkv) — their loop body never runs, and they naturally write partial sums of mx = −INF, S = 0; on the combine side w = expf(−INF − gmx) = 0, contributing exactly zero. Idle splits need no special handling at all — that is why "static grid + data-side nkv" satisfies both replay safety and numerical correctness. This structure was set by 8d and R4 keeps it unchanged — what R4 changes is the geometry inside each warp.

2.3 The old geometry's three sins

8d's partial kernel assigns each lane rows (2 rows per lane); each lane computes the full 128-dim dot product itself and maintains a full 128-dim accumulator:

  1. A LOCAL memory accumulator. float4 oc[32]'s bound hd4 = hd/4 is a kernel parameter (a runtime value), so the compiler cannot fold it into registers → 512 B of local memory spill per lane. Online softmax rescales once per batch of rows: all 32 float4s read out of local, multiplied by corr, written back. nsys works it out to ~80 MB/layer of local traffic — an accumulator that should be registers, faking memory transfers row by row.
  2. Scattered row traversal. Each lane walks its row with 4-byte (f16 __half2) loads: a 256 B row should be 8 sectors, measured at 64 sectors touched per row (12.5% efficiency) — each DRAM transaction uses 1/8 of its bytes.
  3. A parallelism mismatch. 224 single-warp blocks; and "raising SPLITS" only gives each block fewer rows, making the fixed costs (q load, partial write-out) harder to amortize and multiplying the local-traffic copies — the monotonic SPLITS-sweep degradation is this mechanism's direct symptom.

2.4 The new geometry: lane ↔ dim binding

R4 flips the mapping: each lane permanently owns 4 dims (d0 = lane_id·4; at hd=128 the 32 lanes exactly cover all dims), and rows are traversed cooperatively by the whole warp:

  • The accumulator oc is one float4 register — zero spill, and the rescale is 4 FMAs;
  • Row access via kv_ld4: a 16-byte load where 32 lanes each take the row's 4 consecutive dims → one coalesced access covers the whole row, sector efficiency restored;
  • A single row's dot = this lane's 4-dim partial + a 5-step __shfl_xor butterfly → every lane holds the row's complete dot product, and the softmax state is naturally warp-uniform. (In the old design uniformity was free — each lane owned whole rows; the new design trades one butterfly per row for all of the local traffic — a trade that always wins.)
  • Rows are processed 4 at a time: the 4 rows' K loads are issued in batch first, and the serial softmax chain (max → exp → rescale) latency is covered by the later rows' loads.

The arithmetic comparison: old kernel 4.3 MB / 150 µs = 28 GB/s; the new kernel at 79 µs → ~54 GB/s, doubled — the remaining gap is a compute/latency mix, no longer a memory-geometry error.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Lane-owns-dims rather than lane-owns-rows: shrinking the accumulator from O(hd) to O(1) per lane is the root-cause fix; handing row traversal to warp cooperation incidentally makes the access pattern coalesced.
  • The dispatch gate hd % 4 == 0 && hd <= 128 (the kernel comment's own words: "enforced by the dispatch"): the hd=128 7B/14B take the new kernel; shapes not meeting it keep the old path — with hd > 128 the dims→lane mapping leaves many idle lanes, not worth it.
  • ATTN_SPLITS 8 → 32: under the new geometry each split's fixed cost is no longer amplified by local traffic, and more splits buy parallelism from 224 → 896 blocks; tg128 (nkv=128, only 4 rows per split) benefits too. The static grid stays untouched — replay safety.
  • The partial layout follows the geometry: the old layout was "lane 0 writes out all hd4 float4s"; the new layout is dim-sliced — each live lane writes its own oc to dst + 4 + d0, no cross-lane reduction of any kind needed. The combine side reads by index p[4 + i] (i = threadIdx.x, hd threads).
  • Idle lanes are not culled: lanes with d0 ≥ hd keep participating in the butterfly with q4 = 0 (zero contribution) — one warp-uniform fast path, no branches introduced.

3.2 Key code

Before — the 8d kernel: rows to lanes, LOCAL accumulator, per-item rescale (commit 70f57db's deleted side, excerpted verbatim):

const float4* q4 = reinterpret_cast<const float4*>(q + h * hd);
int hd4 = hd / 4;                       // a runtime value → dynamic indexing
float mx = -INFINITY, S = 0.0f;
float4 oc[32];                          // 512 B/lane → LOCAL memory
#pragma unroll
for (int i = 0; i < hd4; i++) oc[i] = make_float4(0, 0, 0, 0);

for (int base = lo; base < hi; base += 64) {   // 2 rows per lane, 64 rows per batch
    ...
    float nmx = fmaxf(mx, bmx);
    float corr = expf(mx - nmx);
    ...
    #pragma unroll
    for (int i = 0; i < hd4; i++) {      // each rescale: 32 items of local read+write
        oc[i].x *= corr; oc[i].y *= corr; oc[i].z *= corr; oc[i].w *= corr;
    }
    S *= corr;
    if (kv0 < hi) {                      // in-row 4-byte load walking (scattered)
        const KV* vrow = v + (size_t)kv0 * stride_kv + hk * hd;
        #pragma unroll
        for (int i = 0; i < hd4; i++) {
            float2 a = kv_ld2<KV>(vrow + i * 4);
            float2 b = kv_ld2<KV>(vrow + i * 4 + 2);
            oc[i].x += e0 * a.x; oc[i].y += e0 * a.y;
            oc[i].z += e0 * b.x; oc[i].w += e0 * b.y;
        }
    }
    ...
}
#pragma unroll                            // at the end: cross-lane reduction, 32 items × 4 dims
for (int i = 0; i < hd4; i++) {
    oc[i].x = warp_reduce_sum(oc[i].x); ...
}

After — the R4 kernel: dims to lanes, register accumulator, a 4-row window (commit 70f57db's added side; in the current tree this body was later refactored by D2 into attn_split_1w_body with the V loads hoisted into the window too — see chapter 66):

// Each lane owns 4 consecutive dims (hd % 4 == 0 and hd <= 128 are
// enforced by the dispatch); lanes with d0 >= hd are idle but keep
// participating in the warp reductions (zero contribution).
int d0 = lane_id * 4;
bool live = d0 < hd;
const float4 q4 = live ? *reinterpret_cast<const float4*>(q + h * hd + d0)
                       : make_float4(0.0f, 0.0f, 0.0f, 0.0f);

float mx = -INFINITY, S = 0.0f;
float4 oc = make_float4(0.0f, 0.0f, 0.0f, 0.0f);   // one register float4

for (int base = lo; base < hi; base += 4) {
    int nr = min(4, hi - base); // warp-uniform
    // Stage K for the whole batch first; the V addresses are already
    // known, so the compiler hoists those loads above the softmax chain.
    float4 k4[4];
    #pragma unroll
    for (int j = 0; j < 4; j++) {
        k4[j] = (live && j < nr)
            ? kv_ld4<KV>(k + (size_t)(base + j) * stride_kv + hk * hd + d0)
            : make_float4(0.0f, 0.0f, 0.0f, 0.0f);
    }
    #pragma unroll
    for (int j = 0; j < 4; j++) {
        if (j >= nr) break; // warp-uniform: all lanes exit together
        // Full-row dot: this lane's 4-dim partial, then a warp reduction
        // so every lane holds the row's complete dot (uniform softmax).
        float d = q4.x * k4[j].x + q4.y * k4[j].y
                + q4.z * k4[j].z + q4.w * k4[j].w;
        #pragma unroll
        for (int off = 16; off > 0; off >>= 1)
            d += __shfl_xor_sync(0xFFFFFFFF, d, off);
        float s = d * scale;
        float nmx = fmaxf(mx, s);
        float corr = expf(mx - nmx);
        float e = expf(s - nmx);
        S = S * corr + e;
        mx = nmx;
        if (live) {
            float4 v4 = kv_ld4<KV>(v + (size_t)(base + j) * stride_kv + hk * hd + d0);
            oc.x = oc.x * corr + e * v4.x;   // rescale+accumulate done in one FMA
            oc.y = oc.y * corr + e * v4.y;
            oc.z = oc.z * corr + e * v4.z;
            oc.w = oc.w * corr + e * v4.w;
        }
    }
}

The partial write-out: from "lane 0 writes all" to dim-sliced (before → after):

// after: no cross-lane reduction; lane 0 writes (mx, S), each live lane writes its own 4 dims
float* dst = partial + ((size_t)sp * nh + h) * pstr;
if (lane_id == 0) {
    dst[0] = mx;
    dst[1] = S;
}
if (live) {
    *reinterpret_cast<float4*>(dst + 4 + d0) = oc; // 16B-aligned via pstr
}

The combine kernel (current tree, identical to the R4 version except 8 → ATTN_SPLITS constant-folding):

__global__ void gqa_attn_split_combine(
    const float* __restrict__ partial, float* __restrict__ o,
    int nh, int hd, int pstr
) {
    int h = blockIdx.y;
    int i = threadIdx.x; // hd threads
    if (i >= hd) return;
    float gmx = -INFINITY;
    for (int sp = 0; sp < ATTN_SPLITS; sp++)
        gmx = fmaxf(gmx, partial[((size_t)sp * nh + h) * pstr]);
    float S = 0.0f, acc = 0.0f;
    for (int sp = 0; sp < ATTN_SPLITS; sp++) {
        const float* p = partial + ((size_t)sp * nh + h) * pstr;
        float w = expf(p[0] - gmx);
        S += p[1] * w;
        acc += p[4 + i] * w;             // dim-sliced layout: thread i reads dim i
    }
    o[h * hd + i] = (S > 0.0f) ? acc / S : 0.0f;
}

nkv data-side + static grid (the current tree's launcher, src/cuda_kernels.cu):

#define ATTN_SPLITS 32
...
gqa_attn_split_partial<__half><<<dim3(ATTN_SPLITS, n_head), 32, 0, stream>>>(...);

The grid (32, n_head) is as static at replay as it was at capture; nkv is read device-side from positions[0] + 1, each split takes chunk = ceil(nkv/32), lo = sp·chunk, hi = min(nkv, lo+chunk) (the first three lines of attn_split_1w_body).

3.3 Pitfalls

  • The partial layout and combine must change in lockstep. When the write side switched from "lane 0 writes hd4 float4s" to dim-sliced, the combine's read index changed from o4[i] (a float4 array) to p[4 + i] (a scalar index) — changing either side alone is a silent data misalignment; the 16 B alignment comment on pstr ("16B-aligned via pstr") is the precondition for the float4 direct write.
  • Bigger SPLITS is not better. Under the new geometry, SPLITS=64 measured 44.7 tok/s, worse than 32's 45.1 — once each split's row count halves, the fixed costs (the q load, two partial round trips, combine's 32-step loop) start eating the gain. 32 is the measured stationary point for this shape family, not a theoretically derived value.
  • break must be warp-uniform. The window loop's if (j >= nr) break relies on nr = min(4, hi - base) being identical across the whole warp — exit on a per-lane condition and the butterfly's __shfl_xor_sync will not line up its participants. The tail rows at the end of a batch are deliberately handled in warp-uniform form.
  • Know the numeric-order change. The dot product went from "each lane serially over 32 chunks" to "a 4-dim partial + butterfly tree" — the floating-point summation order changed. This is not a bitwise change; correctness is secured by the parity sweep (covering SPLITS=32's chunk-boundary shapes), not by byte-for-byte comparison.

4. Verification

  • The parity sweep extended to SPLITS=32 chunk boundaries: nkv takes integer multiples of the chunk length and their neighborhoods — defends against split-cutting off-by-one errors and empty-split handling mistakes (the idle split's mx=-INF, S=0 path is only truly exercised by boundary shapes).
  • The parity comparison: the attention output's numerical tolerance check against a reference implementation (the CPU path) — defends against sum-order errors introduced by the geometry rearrangement being waved through as "precision noise".
  • The full suite: full regression — defends against the hd gate mis-dispatching other shapes (hd≠128).
  • Same-binary interleaved A/B: medians over the @2K and tg128 windows — the kernel-level 2× gain must reproduce on the wall clock, and the direction must agree across both context lengths.

5. Results

Metric (7B q4_k_m)before (8d)after (R4)llama.cpp same window
split-attention kernel148 µs/layer79 µs/layer—
Effective K+V stream rate28 GB/s~54 GB/s (4.19 MB / 79 µs)—
@2K decode39.243.2–45.144.9 (gap 14% → ~4%)
tg128 decode45.147.5–47.647.1 (overtaken)

SPLITS sensitivity (the record's numbers): old kernel 8/16/32/64 → 148/172/419/609 µs (monotonic degradation); new kernel SPLITS=64 → 44.7 tok/s vs 32 → 45.1 (no further gain; the stationary point is 32). The two steps R2+R4 together took 7B decode from 42.2/36.7 (tg128/@2K) to 47.5/43.2–45.1 — short context overtakes llama, @2K enters the ~4% gap zone, and the decode-side chase of llama.cpp was essentially complete here (the later D series extended it to long context).

6. Lessons

  1. A runtime-indexed array = LOCAL memory: an accumulator like float4 oc[32] that "looks like registers" goes to local the moment its bound is a runtime value — 512 B per lane, re-read and rewritten row by row; an accumulator's dimension must be compile-time foldable into registers.
  2. The correct fix for insufficient parallelism is changing the geometry, not adding copies: the monotonic SPLITS-sweep degradation was already warning "the per-warp geometry is broken"; raising the split count just copies the disease.
  3. The lane↔dim binding fixes three things at once: a register accumulator, coalesced row access, and a butterfly reduction trading for warp-uniform softmax — the access pattern is a function of the mapping; the same math under a new mapping doubled the kernel.
  4. A decode kernel's grid must be a compile-time constant; anything that changes per step (nkv) goes through device-side data (positions[0]) plus idle units writing neutral values — the structural constraint for every decode kernel in the CUDA Graph replay era.

← 09 · R2 MMVQ weight-streaming · Index · 11 →

11 · P5 — prefill gap session: TM=128 big tiles + FA rewrite (LANDED, with three REVERTED probes)

Result: 7B @2K prefill 1435 → 2340–2370 tok/s (net +64%); the gap to llama.cpp (3401 @2K) narrows 2.37× → 1.43×. Commit: 86ca78c (P5·0) → d713e6e (P5·1) → 725e307 (P5·2) → fc07c04 (P5·3) landed; 1365c82 (KS=64), a189837 (TM=256), b254c22 (AF32) — three negative results reverted. Date: 2026-09-01 (single-day session). Record note: the session-range start b8568cd cited by the P5 chapter does not resolve to any commit (explicitly stated in §0 footnote 3); all 7 hashes listed above were verified reachable in practice, and the code excerpts key off them.

1. Background — where things stood

After 8p (resident f16 weight cache + dequant folded into the GEMM), 7B @2K prefill sat at ~1435 tok/s while llama.cpp measured 3401 on the same machine — a 2.37× gap. Decode had already approached parity through 8e and R2; prefill had become the most glaring shortcoming across the entire product line.

R1 (int8 MMQ, doc 08), started on August 31, had proven the quantized route could reach parity but was still slow at the time (441 tok/s, and not yet profiled); 8m's wmma f16 GEMM was the serving prefill engine. P5's choice: break the f16 route wide open first — it was both the fastest path available today and the baseline the coming MMQ campaign would measure itself against. By attribution, the prefill wall was made of three pieces: GEMM (the biggest), FA-style tiled attention, and elementwise/convert odds and ends. The session fired six probes in a single day, each gated by the full suite plus interleaved A/B on the same binary (full-output diff, never prompt-echo grep), under the rule "whatever clears the bar lands; negative results revert on the spot with the mechanism preserved."

Where things stall without this step: 2.37× is not a single bug — it is the superposition of three structural wastes: tile geometry (B-panel re-reads), tensor-core coverage (FA's P·V was still on the scalar path), and single-element elementwise kernels. Fixing any one of them alone does not reach parity, so the day was essentially about turning over every block of the prefill wall and measuring it.

2. Principle — the GPU mechanism

2.1 Why a larger N-tile raises tensor-core utilization

The 8m GEMM's output tile is TN × TM = 64 × TM (TN=64 nt rows, TM od columns). Every k-step must move the B panel (a TM-row × KS-k-element weight slice) from L2 into smem, and that B is then reused by all 64 A rows inside the tile. Going TM 64 → 128 means the same B bytes serve twice the mma work:

  • B-panel re-reads per FLOP halve. Each weight-matrix row is read once per k-step and produces 2× the FLOP → the L2 bandwidth needed to sustain the same TFLOPS halves;
  • Barriers per FLOP halve. Double-buffered staging takes one __syncthreads per k-step; with the tile doubled, each sync amortizes over 2× the mma;
  • Accumulator fragments per warp go 2 → 4 (fc[2][ODC], ODC = TM/64): the fragment-load-to-mma ratio gets healthier and tensor-core issue between sync points goes deeper.

The grid organization amplifies this: blockIdx.y is the od tile, so neighboring blocks share the same od-tile's B panel (code comment: 64 rows × id f16 ≈ 0.5 MB) — one weight panel in L2 feeds several blocks, and the f16 weight matrix streams from DRAM roughly once. Kernel time 455 → 302 ms (−34%) is the direct reading of that mechanism.

Running the ledger for one k-step (KS=32, f16 = 2 B/element):

staged into smem: A tile   64 × 32 × 2 =  4 KB
                  B panel  TM × 32 × 2 =  4 KB (TM=64) / 8 KB (TM=128)
FLOP produced:    64 × TM × 32 × 2   = 262K (TM=64) / 524K (TM=128)
smem bytes / KFLOP: ≈ 30.7 (TM=64) → ≈ 23.0 (TM=128)

The bytes on both the A and B sides are amortized over 2× the output elements — per-FLOP smem traffic drops ~25%, barriers per FLOP halve; and the od-tile count halving means the B panel's number of passes over DRAM halves too. The tensor core itself did not get faster; what changed is the fraction of time spent feeding it.

2.2 What the FA rewrite changed

The FA prefill attention is the online-softmax tiled kernel landed by 8n (fa_prefill_f16kv, Q tile 64 rows × KV tile FA_TKV columns). P5 took two cuts at it:

  1. P5·0 — put P·V on tensor cores. QKᵀ was already wmma, but the product of the score matrix P with V still went through scalar FMAs — half the FA kernel's arithmetic had no tensor core. Changed to wmma::mma_sync: P (f16, matrix_a) × V (matrix_b, row_major) accumulating into 16×16 f32 accumulators, 8 accumulator fragments spread over hd=128 per 16-row block. 10.06 → 4.24 ms/layer — the FA channel halved.
  2. P5·3 — parallelize softmax + pad smem rows. (a) The online softmax originally ran as "one 64-row-deep serial chain per warp" pressing on only 2 of the 8 warps; changed to warp-per-row (8 warps × 8 rows, shuffle reduction, −INF seeding) — the serial chain became 8-way parallel; (b) at hd=128 the smem row width is exactly 256 B ≡ 0 (mod 32 banks), so all 8 rows of every ldmatrix land in the same bank group → 8-way conflict; spacing rows +8 halves (272 B) shifts each row by 4 banks. The two cuts together took 4.25 → 1.92 ms/layer.

3. Implementation

3.1 Design choices (why this shape and not another)

  • TM/KS/AF32 all template-parameterized (gemm_f16_nt_kernel_t<TM, KS, AF32>): tile size is the "depth vs occupancy" trade-off axis (the three negative results in §3.3 are all views of it), and parameterization makes every probe a compile-time specialization rather than a runtime branch — a negative result only needs the launch selector changed; TM=64 keeps a MINFER_GEMM_TM=64 opt-out.
  • FA's P matrix enters P·V directly as fragments: at P5·0 P landed in smem and was ldmatrix'd from there; later r48 (doc 51) moved softmax into registers and converts P in-register into the matrix_a fragment — the current-tree code excerpted here is the post-evolution shape, but the tensor-core P·V structure itself has not changed since P5·0.
  • Padded stride is the constant sstr = hd + 8: the 272 B row spacing that fixes the bank conflict is independent of tile size, so it is hard-coded as a kernel constant (still in the current tree).

3.2 Key code

The TM-template GEMM: tile organization and the L2-sharing comment (comment above + opening of gemm_f16_nt_kernel_t in the current tree):

// C[nt, od] = A[nt, id] · B[od, id]^T. 64 x TM output tiles (TM = 64
// baseline, 128 halves the B-panel re-reads through L2 and the per-k-step
// barrier count), k-step 32, double-buffered shared staging, 8 warps (each
// owns 32 nt rows x TM/4 od cols as 2 x TM/64 f32 fragment pairs). f32
// accumulation.
template <int TM, int KS, bool AF32 = false>
__global__ void gemm_f16_nt_kernel_t(...) {
    constexpr int TN = 64;
    constexpr int ODC = TM / 64;  // od 16-col fragments per warp row-half
    ...
    // blockIdx.x = nt tile, blockIdx.y = od tile: consecutive blocks share
    // the same od-tile's B panel (64 rows x id f16, ~0.5MB) in L2, so the
    // f16 weight matrix streams from DRAM ~once instead of nt/64 times.
    int m0 = blockIdx.y * TM;
    int n0 = blockIdx.x * TN;

The inner loop: the 4 mma of fa[4]×fb[2], plus the lesson comment about the fb[1] offset (current tree, introduced by 725e307):

// fa[n-half][k-half]; fb[k-half] per od chunk. Both k halves of each
// 32-slice must accumulate (the v1 bug: only the first 16 k's were
// multiplied); fb's k offset is +16 ELEMENTS (one k-half), not +16
// rows.
#pragma unroll
for (int kh = 0; kh < KHC; kh++) {
    wmma::load_matrix_sync(fa[0], &As[buf * TN * KS + wn * 32 * KS + kh * 32], KS);
    wmma::load_matrix_sync(fa[1], &As[buf * TN * KS + (wn * 32 + 16) * KS + kh * 32], KS);
    wmma::load_matrix_sync(fa[2], &As[buf * TN * KS + wn * 32 * KS + kh * 32 + 16], KS);
    wmma::load_matrix_sync(fa[3], &As[buf * TN * KS + (wn * 32 + 16) * KS + kh * 32 + 16], KS);
    #pragma unroll
    for (int oc = 0; oc < ODC; oc++) {
        wmma::load_matrix_sync(fb[0], &Bs[buf * TM * KS + (ob + oc * 16) * KS + kh * 32], KS);
        wmma::load_matrix_sync(fb[1], &Bs[buf * TM * KS + (ob + oc * 16) * KS + kh * 32 + 16], KS);
        wmma::mma_sync(fc[0][oc], fa[0], fb[0], fc[0][oc]);
        wmma::mma_sync(fc[1][oc], fa[1], fb[0], fc[1][oc]);
        wmma::mma_sync(fc[0][oc], fa[2], fb[1], fc[0][oc]);
        wmma::mma_sync(fc[1][oc], fa[3], fb[1], fc[1][oc]);
    }
}

At TM=128, ODC=2: between consecutive barriers each of the 8 warps issues 2×2×2 = 8 mma (4 at TM=64) — the sync-overhead-to-tensor-core-work ratio halves directly.

FA's padded row stride: the constant behind the bank-conflict fix (opening of fa_prefill_f16kv in the current tree, introduced by fc07c04 and in use ever since):

extern __shared__ __align__(256) uint8_t smem[];
// Padded smem row stride: hd=128 halves = 256B ≡ 0 mod 32 banks makes
// every wmma ldmatrix row land on the same bank group (8-way conflict
// per load). +8 halves (272B) shifts each row by 4 banks.
const int sstr = hd + 8;
__half* Qs = reinterpret_cast<__half*>(smem);
__half* Ks = Qs + FA_TQ * sstr;
__half* Vs = Ks + FA_TKV * sstr;

FA's tensor-core P·V (current-tree shape; the P5·0 prototype went through a smem relay, r48 moved it to in-register construction):

// Build the P@V A-operand (f16 matrix_a) from the scaled fragments IN
// PLACE. matrix_a m16n16k16 row_major and the f32 accumulator use the
// SAME (row,col) layout, so pa.x[i] == fc[cc].x[i] element-wise. Ks in
// QK^T is col_major; V in P@V is row_major (both validated standalone).
wmma::fragment<wmma::matrix_a, 16, 16, 16, __half, wmma::row_major> pa[FA_TKV / 16];
...
// acc = acc*alpha + P · V. V (B) is row_major from Vs.
#pragma unroll
for (int kk0 = 0; kk0 < FA_TKV; kk0 += 16) {
    #pragma unroll
    for (int ob = 0; ob < 8; ob++) {
        wmma::fragment<wmma::matrix_b, 16, 16, 16, __half, wmma::row_major> vb;
        wmma::load_matrix_sync(vb, &Vs[kk0 * sstr + ob * 16], sstr);
        wmma::mma_sync(acc[ob], pa[kk0 / 16], vb, acc[ob]);
    }
}

P5·1's two elementwise kernels (the current tree is already the vectorized form) — store_kv_f16 does 4 dims per lane (one float4 load + two half2 stores, scalar fallback for the tail rows):

int t = blockIdx.x;
int j = (blockIdx.y * blockDim.x + threadIdx.x) * 4;   // 4 dims per lane
if (t >= nt || j >= nkt) return;
int p = positions[t];
if (j + 3 < nkt) {
    float4 v = *reinterpret_cast<const float4*>(src + (size_t)t * nkt + j);
    __half2* d = reinterpret_cast<__half2*>(dst + (size_t)p * nkt + j);
    d[0] = __floats2half2_rn(v.x, v.y);
    d[1] = __floats2half2_rn(v.z, v.w);
} else { /* tail: scalar loop */ }

convert_f32_f16_kernel does 8 elements per lane (the kernel comment carries its own ledger; the "P1" label in the comment is a stale historical numbering — the P5 record itself also notes the "8p" label collision, see the §5 table).

// P1: 8 elements per thread (2x float4 -> 4x half2) instead of one
// scalar element — 8x fewer transactions on the same traffic.
long long base = ((long long)blockIdx.x * blockDim.x + threadIdx.x) * 8;
if (base + 7 < n) {
    float4 a = *reinterpret_cast<const float4*>(x + base);
    float4 b = *reinterpret_cast<const float4*>(x + base + 4);
    __half2* o = reinterpret_cast<__half2*>(out + base);
    o[0] = __floats2half2_rn(a.x, a.y); ...
}

The single-element kernels' waste is purely at the transaction level: an f32 element is 4 B, so a 32 B sector uses 1/8 of itself; vectorization divides the transaction count over the same traffic by 4–8 — together these two kernels only lifted +4%, because they were never at the wall's center of gravity.

The survival note for the AF32 mirror mechanism (current tree, b254c22's mechanism kept along with its own cause of death):

// AF32 A staging, mirror scheme: cp.async the F32 k-tile into a smem
// mirror (16B = 4 f32 chunks; async again — the v1 synchronous global
// loads stalled every k-tile and measured -8%), then convert
// smem->smem f32->f16 right before compute. Requires id % 8 == 0.

3.3 Pitfalls

  • The fb[1] offset: +16 elements, not +16 rows. The first version after widening TM fetched the second k-half's B fragment as "+16 rows" — the GEMM result was flat-out wrong. B-fragment slices in ldmatrix/wmma advance along the k (column) direction, so fb[1]'s offset is k-half = +16 elements; cuda_prefill_f16_gemm_parity caught it on the spot. The comment still sits above the inner loop today (see the excerpt above).
  • Silent fallback on smem overflow. P5·3's double-buffered working set exceeded the 99 KB/block cap, cudaFuncSetAttribute failed → the kernel silently fell back to the legacy path and the whole machine dropped to 313 tok/s — a performance collapse hidden inside "no error reported". Fix: the fallback prints a warning; the padded layout shipped single-buffered (69 KB). "Resource limits must fail loudly" became a campaign rule from then on (same class as the r8 phantom and r58's fake OOM).
  • __syncthreads does not order cp.async. When the double buffer was cut, cp.async's commit_group/wait_group 0 were removed along with it, leaving a bare __syncthreads to wait on staging — async copies are not constrained by the barrier, and parity came out at 0.28. Lesson: __syncthreads only orders ordinary memory accesses; cp.async must commit/wait group.
  • KS=64 (P5·4, REVERTED): the idea of halving per-FLOP barriers again was beaten back by occupancy — TM=128+KS=64 needs 56 KB smem (current-tree comment verbatim), resident blocks halved, 1464 vs 2345 tok/s (−38%). The mechanism is kept (MINFER_GEMM_K64=1 can re-measure it), KS=32 stays default.
  • TM=256 (P5·neg, REVERTED): a wider tile hit the same wall, −3%; the session went ahead and parameterized the thread count anyway (NW: TM≤128 → 8 warps, TM=256 → 16 warps), and that part of the mechanism survives in the current tree.
  • In-kernel f32→f16 A staging (AF32, REVERTED): the goal was folding the separate convert pass into GEMM staging; end-to-end −8%, and while the standalone kernel version had proven correct, the integrated version's parity never closed — abandoned (WIP chain a3b0dcd/69c3933/aa40ed3). r23 later measured the convert pass at only 6% of the f16 wall — the upper bound never supported this direction; the lesson "measure the upper bound before building" was booked here.

4. Verification

  • cuda_prefill_f16_gemm_parity: numeric parity of GEMM output against the reference implementation — specifically defends against tile-reshuffle / wmma-slicing bugs (it caught the fb[1] bug).
  • Full suite every step + interleaved A/B on the same binary (full-output diff): defends against "changed A, broke B" and window-drift noise; the P5 chapter states outright "never prompt-echo grep".
  • Parity check on cp.async ordering: the 0.28 output difference is the direct signal of the removed wait_group — here the parity gate defends asynchronous memory ordering, not math.
  • Explicit warning on the smem cap: after the fix, the fallback path prints the actual smem requirement and the failure reason — prevents a repeat of the "performance silently collapsed" scenario.

5. Results

StepCommitMetricbefore → afterVerdict
P5·0 FA P·V on wmma86ca78cFA ms/layer; @2K prefill10.06 → 4.24; +15%🟢
P5·1 elementwise vectorizationd713e6e@2K prefill1435 → 1493 (+4%)🟢
P5·2 TM=128 big tile725e307@2K prefill; GEMM kernel1493 → 2267 (+30%); 455 → 302 ms🟢
P5·3 all-warp softmax + padded rowsfc07c04@2K prefill; FA kernel2267 → 2365–2371 (+4–5%); 4.25 → 1.92 ms/layer🟢
P5·4 KS=641365c82@2K prefill1464 vs 2345 (−38%)🔴 reverted
P5·neg TM=256a189837@2K prefill−3%🔴 reverted
P5·neg AF32 in-kernel convertb254c22@2K prefill−8%, parity hole never closed🔴 reverted

Net effect: 1435 → 2340–2370 tok/s (+64%), vs llama 2.37× → 1.43×. Comparison basis: llama.cpp 3401 @2K (ca3d5a3e1 bench build). After P5·3, fa_prefill_f16kv runs 1.92 ms/layer vs llama's 0.79 — FA is still part of the residual gap, left to the later FAP line (docs 49/51).

Post-P5 nsys budget (per 2K prefill): GEMM ~597 ms (effective ~46 TFLOPS; llama 455 ms), FA ~54 ms, convert ~56 ms, swiglu ~51 ms — the 1.3× kernel gap to llama on GEMM plus the misc items together make up the remaining 1.43×. That budget table then became Era C's (the MMQ campaign's) starting coordinates.

6. Lessons

  1. Silent fallbacks are correctness-grade hazards: resource limits (smem caps etc.) must fail loudly and print the actual value — "performance suddenly collapsed but nothing errored" always means the fallback path is running.
  2. __syncthreads does not order cp.async: ordering of async copies can only be established by commit_group/wait_group; the barrier is transparent to them.
  3. Tile widening's benefit boundary sits at "smem doubles → resident blocks halve": TM 64→128 gave +30%, KS 32→64 gave −38% — the depth-vs-occupancy mutual exclusion is a trade-off this campaign rediscovers repeatedly (r38/r39/r40 learned it again on MMQ).
  4. Measure the upper bound before writing the kernel: AF32's −8% and r23's 6% upper bound show that measuring the target pass's wall-clock share before building can veto an entire route up front.

← 10 · Index · 12 →

12 · r5–r6 re-ranking + structural rewrite spec (MEAS-ONLY + REVERTED)

Result: both queued levers re-measured negative — KD=4 re-test 427 vs 438 tok/s, and the 4-warp 32×32 warp tile 399 tok/s with a parity hole (zero cells in mmq_w80) — all reverted; the wall decomposition was also corrected (the 352 ms q4_K dequant pass runs at LOAD time and is not inside the prefill wall at all). r6 then landed the execution spec: raw-byte smem staging + dequant at mma time — it became the execution contract for the 30+ rounds of work that followed in Era C. Commit: 1e0673f, 491eb5c (both docs commits; the code under test (KD=4 switch, 4-warp tile, dequant vectorization) was reverted right after measurement and never became its own commit). Date: 2026-09-01.

1. Background — where things stood

As P5 wrapped up (around 2026-09-01), minfer's CUDA engine stood here: the default f16 path, pushed by the TM=128 big tile (P5·2) and the FA softmax balancing (P5·3), had reached 2340–2370 tok/s @2K — 1.44×/1.43× against llama-bench's 3401 @2K; the decode line (R2 MMVQ weight- streaming rework, R4 split-attention dim-parallel rewrite) had already pushed decode to a seesaw against llama.cpp. Prefill was the remaining main gap, and prefill's next direction had to be a binary choice: keep pressing the f16 path's wall, or pull the opt-in int8 MMQ line (R1, 40e97c9) up from 441 tok/s.

R1's position was awkward. It was parity-clean and structurally correct (int8 tensor-core mma + q8_0 activations), but its throughput was only ~6.1 TMAC/s: the f16 GEMM itself ran about 24 TMAC/s and llama.cpp's MMQ about 30. That is, R1 trailed the f16 GEMM by 4× and llama.cpp's MMQ by 5× — and when R1 landed, that 8×-scale gap had never been attributed (the master table's R1 row, verbatim: "the 8× gap was unprofiled"). Every overt lever on the r1–r4 incremental line (MMQ landing, MMVQ streaming rework, split-attention rewrite) had been harvested; the MMQ line stalled on the question "what is the next lever?"

Two candidate levers were queued at the time, both "obvious" changes to the R1 kernel shape:

  • KD=4: drop the number of 32-k chunks resident in each double-buffer half from 8 to 4, halving the smem footprint and raising occupancy from 1 block/SM to 2 — the textbook occupancy lever;
  • 4-warp 32×32 warp tile: change the warp tile from 32(i)×16(j) to 32×32 so each fragment word feeds 2× the mma, halving smem traffic — the textbook data reuse lever.

A third "obvious" lever was even more tempting: the 352 ms q4_K dequant pass. In the one- shot CLI prefill nsys timeline of the time there was a ~352 ms w16 fill that looked like forty percent of the 890 ms total — vectorizing it seemed like free money.

What r5–r6 did was knock down each of these three "obvious" ideas one by one, then write the only road left standing into an executable spec. The cost of not doing this is concrete: without re-measuring KD=4 and the 4-warp tile, the team would keep investing on a wrong assumption ("occupancy/reuse is the gap's source"); without correcting the dequant pass's attribution, the team would optimize a pass that isn't inside the wall; and without a spec, a structural rewrite of that size would be attempted with no parity gate — which is exactly how r5's 4-warp tile broke parity on the spot.

2. Principle — the GPU mechanism

Depth vs occupancy: why KD=4 was slower. R1's word-staging kernel keeps this cadence per k-tile: [dequant ALU + smem expansion + __syncthreads] → mma section. The staging section is a block-wide barrier-synchronized serial section — all 8 warps must finish their unpack before anyone crosses the barrier. At KD=8 you pay that cadence once per 256-k; KD=4 doubles the cadence count, so the fixed overhead amortized per weight byte (barriers, addressing, scale reads) rises. Going from 1 to 2 blocks/SM does let another block's compute section overlap with this block's staging section, but for this kernel the overlap gain cannot buy back the doubled cadence — the master table's verdict is "staging-depth amortization dominates". 427 vs 438: the occupancy lever lost by 2.5%. The transferable conclusion: the depth-vs-occupancy trade- off is a property of the kernel class — the heavier the staging section (more ALU, denser barriers), the more valuable depth amortization is; once r7's raw-byte kernel cut the staging section down to pure cp.async, the trade-off point would move (the r6 spec wrote this explicitly as "raw staging shifts the tradeoff", requiring both KD depths to be re-measured).

4-warp 32×32: why the reuse lever broke on correctness. Changing the warp tile from 32×16 to 32×32 has two theoretical gains: each fragment word (one smem read) feeds 2× the mma, and the B-side fragment count stays fixed while coverage doubles, halving smem traffic. The cost: 256 threads must be cut into a 4×2 sub-block grid, staging partitioning changes from a single loop to loopified, and B fragments go from 2 to 4 — the whole fragment mapping had to be re-derived by hand. The result: 399 tok/s (9% slower than 438) and a mmq_w80 parity failure: some C cells got zero contribution (zero cells) — not numeric drift, but a hand-written mapping that missed some cells' contribution paths entirely. What the reuse lever saved on "more mma per word" was outweighed by the index ALU and scheduling losses the mapping complexity introduced; and the parity hole showed the complexity had already passed what could be hand-verified at the time.

Wall decomposition: what inside the 890 ms is movable. r5 re-read the one-shot CLI prefill with nsys and split the wall into: GEMM ~600 ms + convert 56 + fa 54 + swiglu 51 + add 19 + host gaps. Three secondary conclusions:

  • swiglu re-measured at ~257 GB/s = bandwidth peak — already at roofline, no implementation headroom left;
  • convert (the f32→f16 activation conversion) at 56 ms is an inherent tax of the f16 path;
  • the 352 ms dequant pass runs at LOAD time (the w16 warm phase one-shot dequantizes q4_K into the f16 weight cache) and is not inside the prefill wall. Vectorizing it measured a null delta — the scalar version was already warp-coalesced, and it runs on the cold-start path, invisible inside the wall.

GEMM is therefore the only material lever: everything else is either one-shot (dequant), at roofline (swiglu), or small (fa/add). The source record puts it as "GEMM ~85% is the only material lever" — meaning that after host gaps and incompressible items, nearly all actionable time is GEMM.

The dequant-at-mma-time principle. The spec's core insight: the nibble unpack's int ALU does not have to finish during staging — in the gaps between tensor-core mma.m16n8k32 issue slots, the INT32 pipeline is largely idle. Resident the B side in smem as raw bytes (the 144 B Q4KB super-block: d[2] + dmin[2] + scales[12] + qs[128], nibbles unexpanded, un-centered), and staging degenerates into pure cp.async copies; move the unpack into the mma loop, one shift/mask per 32-bit word, __vsubss4 for centering — the ALU then overlaps tensor-core issue instead of serializing ahead of the barrier.

New concepts in this doc: pad40 = the 40 B q8_0 32-element block layout (d 2B + qs 32B @4 + ssum 4B @36, padded to 40 overall); Q4KB super-block = q4_K's 256-k weight block (144 B); KDR = the number of 32-k chunks resident per double-buffer half (KD=8 means 256-k); word staging = R1's approach — expanding nibbles into int8/words in smem during staging.

3. Implementation

3.1 Design choices (why this shape and not another)

r5's two negative results marked the end of the road called "keep looking for staging knobs on the R1 shape": occupancy (KD=4) and reuse (32×32) had both been tried and both lost. There was still no profiler (ncu would not work on this device until r13), so the gap could not be located with counters; what could be established was that the common bottleneck was in the per-32-k-chunk inner loop overhead (the rescale FMA chain, scale smem reads, sync cadence) — a property of R1's staging structure itself, not of any single knob. So r6's choice: stop looking for knobs and write an execution spec for a structural rewrite, pinning the entire next phase (later r7 through r34) to a design.

The spec's seven points, each with an explicit "why":

  1. smem holds RAW bytes only: A side 40 B pad40 per (token, 32-k chunk) (2×uint4 qs + uint2 d/ssum) — the pad40 layout is itself a raw format, so fragment words can be read straight out of it; B side 144 B Q4KB per (row, 256-k super-block) (9×uint4). At KD=8 the footprint is A 2×8×64×40 = 41 KB + B 2×64×144 = 18.4 KB + scales ≈ 62 KB → 1 block/SM; KD=4 ≈ 31 KB → 2–3 blocks/SM. Both depths re-measured — r5 had already shown this trade-off moves.
  2. staging = pure cp.async 16B chunks (A: 2×16 + 8 B; B: 9×16), commit_group per (kt, buf); zero dequant ALU on the staging path.
  3. In-register dequantization inside the mma loop: the math matches mmq_stage_b TYPE==5 (get_scale_min_k4 per (row, sub), nibble unpack, __vsubss4), so the ALU overlaps tensor- core issue.
  4. Warp tile stays 32(i)×16(j) × 8 warps on 64×64 — the 4-warp 32×32 variant measured slower and broke parity; the spec says outright "do not retry".
  5. Land as mmq_raw_nt_kernel<TYPE> beside mmq_nt_kernel, entering prefill_mmq behind the MINFER_MMQ_RAW=1 gate — R1 stays untouched, always A/B-able and revertible; the first cut covers q4_K only (79% of MMQ time), with q6_K's KSPLIT=2 structure left for later.
  6. Three-stage parity gate: cuda_prefill_mmq_parity gains a raw-mode arm (same host reference) → 7B greedy token identity vs the f16 path → quiet-window A/B (vs R1 MMQ and vs the default f16). Parity before performance, and the order is non-negotiable — r5's parity hole is the cautionary tale for getting this order backwards.
  7. The quantize pass (129 ms) is follow-up work once the GEMM wins: it is limited by convert bandwidth (~87–102 GB/s) and the serial fmaxf chain's latency; the approach is tree-reduce amax + register-packed stores, target ~60–70 ms.

The spec also wrote down two tiers of expected payoff: at GEMM = f16-parity (≥24 TMAC/s), the MMQ path (quantize ~130 + GEMM ~600) drops the 56 ms convert → ~2670 tok/s; at llama-parity (GEMM ~480 ms) → wall ~634 ms → ~3250 tok/s (conversion per the spec's anchors at the time).

3.2 Key code

The spec verbatim (git show 491eb5c -- docs/CUDA_OPTIMIZATION.md, design body excerpted):

### MMQ structural rewrite — execution spec (P6 r6, for next session)

Goal: mmq GEMM 6.1 TMAC/s (23 ms per ffn_gu call) -> >=24 (f16-GEMM
parity) or ~30 (llama.cpp parity). ...

Design (llama.cpp mmq structure; q4_K first = 79% of MMQ time):
1. smem holds RAW bytes only: A per (token, 32-k chunk) = 40 B pad40
   (2x uint4 qs + uint2 d/ssum); B per (row, 256-k super-block) =
   144 B Q4KB (9x uint4). KD = 8 chunks (256-k) per double buffer.
   Footprint: A 2x8x64x40 = 41 KB + B 2x64x144 = 18.4 KB + scales
   ~ 62 KB -> 1 block/SM at KD=8; KD=4 -> ~31 KB -> 2-3 blocks/SM
   (re-measure both; r5 showed depth beats occupancy for the word-
   staging kernel, raw staging shifts the tradeoff).
2. Staging = pure cp.async 16B chunks (A: 2x16 + 8 B; B: 9x16),
   commit_group per (kt, buf); NO dequant ALU in the staging path.
3. The mma loop dequants IN REGISTERS from smem raw bytes (same math
   as mmq_stage_b TYPE==5: get_scale_min_k4 per (row, sub), nibble
   unpack, __vsubss4) so the ALU overlaps tensor-core issue instead of
   serializing before __syncthreads.
4. Warp tile stays 32(i)x16(j) x 8 warps on 64x64 (the 4-warp 32x32
   variant measured slower AND broke parity — do not retry).
5. Land as mmq_raw_nt_kernel<TYPE> beside mmq_nt_kernel; gate
   MINFER_MMQ_RAW=1 through prefill_mmq so R1 stays intact; q4_K only
   in the first cut (q6_K KSPLIT=2 later).
6. Parity: extend cuda_prefill_mmq_parity with a raw-mode arm (same
   host reference), then the 7B greedy-token-identity check vs the f16
   path, then quiet-window A/B vs R1 MMQ and vs the default f16 path.
7. Quantize pass (129 ms) is follow-up work once the GEMM wins: ...
   tree-reduce amax + register-packed stores, target ~60-70 ms.

The R1 word-staging that the spec replaced (the q4_K branch of mmq_stage_b in the current tree src/cuda_kernels.cu; after r15 dmin is precomputed as ds/dm pairs and the nibble words stay unsigned — the same branch at r6's time did __vsubss4 centering in staging; the mechanism is identical: dequant ALU and expanded words both live in the staging section, serialized ahead of the barrier):

} else if constexpr (TYPE == 5) {   // q4_K: value = d·s·nib − dmin·m (nib unsigned)
    int sb = c >> 3, s = c & 7;
    const uint8_t* blk = W + (size_t)j * ((nb32 >> 3) * 144) + (size_t)sb * 144;
    #pragma unroll
    for (int w = 4 * half; w < 4 * half + 4; w++) {
        uint32_t N = *(const uint32_t*)(blk + 16 + (s >> 1) * 32 + 4 * w);
        uint32_t nib = (s & 1) ? ((N >> 4) & 0x0F0F0F0Fu) : (N & 0x0F0F0F0Fu);
        qb[r * MMQ_WS + w] = (int)nib;          // ← nibble expanded into int words resident in smem
    }
    if (half == 0) {
        uint8_t sc, m;
        get_scale_min_k4(s, blk + 4, &sc, &m);  // ← scale decode also sits in the staging section
        ds[r] = h2f(*(const uint16_t*)blk) * (float)sc;
        dm[r] = -(h2f(*(const uint16_t*)(blk + 2)) * (float)m);
    }

Segment by segment: each (row, chunk) unit reads a 144 B Q4KB block from global memory, does 4 words of nibble shift/mask, writes 4 expanded words into smem, and then the half==0 thread decodes the scale — this whole set of ALU and smem writes happens before __syncthreads, pure up-front serial cost ahead of the mma section. The raw-byte scheme compresses this into "a 9×16 B cp.async copy", deferring all dequantization to registers inside the mma loop.

3.3 Pitfalls

  • The parity hole's shape was "zero cells": the 4-warp tile's mmq_w80 failure was not numeric deviation (the max-diff-82.9 kind of partial-sum misalignment) but entire output cells never being written — the hand-written fragment mapping had coverage holes. The lesson from this class of error went into spec point 6: structural changes must pass a parity gate first, and the parity gate must be designed to expose "zero contributions" (per-cell comparison against the same host reference).
  • The attribution trap in wall decomposition: in the one-shot CLI prefill's nsys timeline, the LOAD-phase w16 fill (352 ms) and the prefill wall (890 ms) sit on the same trace; not slicing by phase makes it look like an in-wall cost. r5's commit message first wrote "~352ms of the 890ms wall"; the spec commit 13 minutes later corrected it to "runs at LOAD time, NOT inside the Prefill wall" — the correction was grounded in the vectorized dequant's measured null delta.
  • Negative results must not be extrapolated: KD=4 was negative for the word-staging kernel, but the spec did not write it up as a universal conclusion — it wrote "raw staging shifts the tradeoff" — and r7's raw kernel re-measuring KD indeed gave a different answer (KD=8 472 vs KD=4 440; depth still wins, but the meaning of the margin had changed).

4. Verification

  • cuda_prefill_mmq_parity sweep: every quantization sub-shape compared per-cell against the same host reference — the 4-warp tile's mmq_w80 zero cells were caught by exactly this gate (defends: fragment-mapping coverage holes / hand-written indexing bugs).
  • suite 169/0: all green after the reverts (defends: the reverts themselves introducing a regression).
  • nsys phase slicing: the one-shot CLI prefill's w16-fill split measurement, used to attribute the dequant pass to LOAD (defends: prioritizing an out-of-wall cost as if it were in-wall).
  • Default-path isolation: everything was measured under the MINFER_MMQ=1 opt-in; the default f16 path was untouched the whole time (2320–2370 tok/s unchanged) (defends: measurement changes polluting the production path).

5. Results

All measurement/revert, no landed code:

  • KD=4: 427 vs 438 tok/s — R1-era conclusion holds; occupancy 1→2 blocks/SM cannot buy back the staging-depth amortization loss.
  • 4-warp 32×32 tile: 399 tok/s (9% slower than 438) + mmq_w80 parity zero cells — reverted, and marked "do not retry" in the spec.
  • Vectorized q4_K dequant: null delta — reverted (it is not inside the wall; the scalar version was already warp-coalesced).
  • The r6 spec: became Era C's execution contract. r7–r8 landed the raw kernel per points 1–5 (472 tok/s), point 7 was cashed in early by r7–r8 (quantize 129 → 74 ms), and point 3's "dequant at mma time" plus point 6's parity gate carried through all subsequent MMQ kernel work.
  • Veto mechanism: the retry conditions for KD=4/4-warp are that the staging structure changes first (raw-byte conversion zeroes the staging section's ALU, reshuffling the depth/occupancy trade-off); the precondition for retrying dequant vectorization is that it returns to inside the wall (i.e. abandoning the w16 cache path) — neither happened later, so the verdicts stand.

6. Lessons

  1. Re-reading a profile must be done against the current code and phase slicing — the 352 ms dequant pass was never inside the wall, and one mis-attributed number nearly defined the whole optimization direction.
  2. Write an execution spec with parity gates before any large rewrite; the spec must state what NOT to retry (the 4-warp tile's "do not retry" saved every later re-argument cost).
  3. Depth vs occupancy has no universal answer — it is a function of the staging section's weight; when the staging structure changes, re-measure instead of inheriting the old conclusion.
  4. Parity gate before performance gate: zero-cell-class structural errors should be caught before any meaningful timing measurement happens.

← 11 · Index · 13 →

13 · r7–r8 raw-byte MMQ kernel + wide tile + FA probe (LANDED (raw) + REVERTED (probe))

Result: mmq_raw_nt_kernel (MINFER_MMQ_RAW=1, q4_K) landed: 472 vs 441 tok/s (+7%), GEMM kernel 20.2 → 23.0 ms same-window comparison, quantize pass 129 → 74 ms; the wide tile's first 2124 tok/s was a phantom — KD=8 needs 135 KB dynamic smem, over the ~99 KB opt-in cap, and both attr-set and launch failed silently with the GEMM writing nothing; after adding the cap guard the honest matrix has narrow KD=8 472 as the optimum. The FA KV L2-prefetch probe was null (2319–2345 vs 2345), reverted. Commit: d440d16 (raw kernel), d9d626a (quantize rebuild), a41eac0 (wide tile wip), ef9d5b4 (phantom fix + guard), 87bade0 (FA probe record, probe code not kept). Date: 2026-09-01.

1. Background — where things stood

The r6 spec had just been written (step 12) and Era C's execution contract was in place: smem holds raw bytes only, staging is pure cp.async, dequant moves into the mma loop, the MINFER_MMQ_RAW gate keeps R1 intact, and the three-stage parity gate goes first. This step is the spec's first round of execution, plus three adjacent items carried along:

  • The raw kernel itself (spec points 1–5): replace R1's word-staging (dequant ALU + expanded words resident in smem, serialized ahead of the barrier) with raw-byte staging. R1's 441 tok/s trailed the f16 GEMM path's 2318 tok/s by 4.9×, and the spec judged the common bottleneck to be the per-32-k-chunk inner-loop overhead — raw conversion is the first cut at compressing the inner loop's up-front cost.
  • The quantize pass rebuild (spec point 7 done early): quantize_q8_0_pad40 runs ~129 ms per 7B @2K prefill (380 launches, ~87 GB/s). In the current MMQ wall (4.7 s) that is only ~2.7%, so it was destined to measure flat — but it was explicitly positioned as groundwork: once the raw kernel lands, the MMQ wall shrinks several-fold and quantize's share surfaces. Fixing it first avoids conflating two variables later.
  • Wide tile + FA probe: the wide tile is the other face of the spec's KD/footprint discussion (a 128-token block halves B traffic); the FA probe was a free check of whether the f16 wall's FA item (1.86 ms/layer, ~6%) still had cheap gains — fa_stage_kv_async is single-buffered, and if the stage stall comes from DRAM latency, prefetching the next KV tile early should be able to hide it.

The campaign's position did not change qualitatively after this step — best 472 is still 4.9× from 2318 — but elimination advanced a great deal: the three hypotheses of B-traffic reduction, doubled reuse, and KV prefetch all went out with numbers attached, and the remaining hypothesis converged to "inner-loop instruction economy", leading directly to r9's reference decode.

2. Principle — the GPU mechanism

The raw-byte staging byte arithmetic. For each (row, 32-k chunk), R1's word-staging must: read 144 B of Q4KB, do 4 shift/masks, write 4 expanded int words into smem, decode the scale — then the whole block crosses the barrier. The raw scheme's ledger is completely different:

  • A side (activations, pad40): 40 B per (token, chunk), and the pad40 layout is itself a raw format — d(2B) + qs(32B @4) + ssum(4B @36); word w of a fragment is exactly the raw int8 lane group k∈[4w, 4w+3), with not a single byte of conversion needed. Five 8 B cp.async.ca complete one chunk's transfer.
  • B side (weights, Q4KB): 144 B per (row, 256-k super-block) = nine 16 B cp.async. Nibbles unexpanded, un-centered, resident in smem as-is.
  • The only "compute" done in the staging section is the per-(row, chunk) scale prefetch — a synchronous global read of 3–6 B per unit, written into R1's sds/sdm layout. This is deliberate: keeping the C-fragment rescale code line-for-line identical to R1 leaves fragment mapping as the only parity risk surface.

smem footprint: A 2×8×64×40 = 41 KB + B 2×64×144 = 18.4 KB + scales ≈ 62 KB at KD=8 → 1 block/SM; KD=4 ≈ 31 KB → 2–3 blocks/SM. r5's lesson (depth vs occupancy is a function of staging weight) applies here: once the staging section's ALU is zeroed, both depths are worth re-measuring.

Dequant at mma time. Inside the mma loop, one shift/mask per 32-bit word unpacks the nibbles. In the gaps between mma.m16n8k32 issue slots the INT32 pipeline is idle, so the unpack ALU overlaps tensor-core issue instead of serializing. The fragment mapping is deliberately identical to R1's — the int accumulator and accumulation order are bit-for-bit unchanged — so the parity gate only has to verify "did the data movement move the right bytes", not the math again.

The wide tile's theoretical gain and its hidden premise. Growing the block from 64×64 to 128×64: grid.y halves and per-output-cell B re-reads halve; meanwhile each fragment word feeds 2× the mma. This lever's implicit premise is that the B re-reads actually happen at DRAM — as we'll see, that premise is wrong.

The FA probe's principle. fa_prefill_f16kv's KV staging is single-buffered: stage(kt) → compute(kt) → stage(kt+1) in series; if the stage stall comes from DRAM latency, issuing an L2 prefetch for the next block early could hide it inside compute. But Qwen2.5-7B is GQA (7 q-heads sharing 1 kv-head): the same K/V rows are read concurrently by 7 attention blocks, and after the first read brings a row into L2 the remaining 6 blocks all hit in L2 — the tile is already L2-resident, and the stage stall was never DRAM-bound.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Whole-super-block constraint: raw staging moves B in whole 256-k blocks (KDR=8 is exactly one super-block), with a launcher-side guard nb32 % 8 == 0 — when the data doesn't form whole blocks the raw path is refused and we fall back to R1. The constraint buys zero addressing branches on the B staging side.
  • Scale prefetch into sds/sdm, rescale code untouched: the C-fragment rescale is code that R1's parity has repeatedly validated; reusing it means the raw kernel's parity verification focuses on staging and the fragment mapping — minimizing "new code volume" is a direct application of the r5 parity-hole lesson.
  • The wide tile as a clone of narrow, not a rewrite: mmq_raw_wide_nt_kernel is cloned from the parity-proven narrow kernel, changing only the block size (128 tokens), the warp mapping (4×2 subs of 32×32), the B fragment count (2→4), and sum[16]→[32]. The clone strategy makes mapping errors diffable hunk by hunk — in hindsight this was exactly the key to localizing the phantom.
  • Why quantize's tree-reduce amax is bitwise-safe: max is exact for any association order (float max has no rounding), so replacing the serial 32-deep fmaxf chain with a 4-way tree leaves amax bit-identical; the rintf quantization pass is untouched line for line → output bit-identical. The performance gain (~87 GB/s → higher, 129 → 74 ms) comes entirely from the shorter latency chain and wider stores; numerics untouched.

3.2 Key code

The raw kernel: staging and the double-buffer main loop (git show d440d16 -- src/cuda_kernels.cu, excerpt):

// ─── P6: raw-byte MMQ (llama.cpp structure) — q4_K first ─────────────
// smem holds RAW quant bytes; staging is pure cp.async; dequant happens
// in registers inside the mma loop. Only the per-(row, chunk) scales are
// pre-computed at staging (shared by the whole warp; the C-fragment
// rescale needs all 8 fragment rows, and computing them per lane would
// be 4x redundant work in the hot loop).
//   A: pad40 chunk = d(2B) qs(32B @4) ssum(4B @36) — consumed as-is
//      (word w of the fragment = the raw int8 lane group k 4w..4w+3).
//   B: Q4KB=144B super-block per row per 256-k; nibbles unpacked at mma.
// Requires whole super-blocks: nb32 % 8 == 0 (launcher guard).
__device__ __forceinline__ void mmq_cp8(uint8_t* smem_dst, const uint8_t* gsrc,
                                        bool full) {
    unsigned d = (unsigned)__cvta_generic_to_shared(smem_dst);
    int sz = full ? 8 : 0; // src-size 0 => zero-fill the 8B chunk
    asm volatile("cp.async.ca.shared.global [%0], [%1], 8, %2;\n" ::"r"(d),
                 "l"(gsrc), "r"(sz));
}

#define RAW_STAGE(kt, b)                                                       \
    do {                                                                       \
        /* A: (token, chunk) pad40 chunks as 5x 8B cp.async */                 \
        for (int x = threadIdx.x; x < MMQ_BI * KDR * 5; x += blockDim.x) {     \
            int kd = x / (MMQ_BI * 5);                                         \
            int rem = x % (MMQ_BI * 5);                                        \
            int r = rem / 5, u = rem % 5;                                      \
            int c = (kt) * KDR + kd;                                           \
            int tok = i0 + r;                                                  \
            const uint8_t* src =                                               \
                q8x + ((size_t)tok * nb32 + c) * 40 + u * 8;                   \
            uint8_t* dst = qa8 + ((size_t)(b) * KDR + kd) * MMQ_BI * 40        \
                         + (size_t)r * 40 + u * 8;                             \
            mmq_cp8(dst, src, tok < nt && c < nchunk);                         \
        }                                                                      \
        /* B: whole 144B super-block per row as 9x 16B cp.async */             \
        int sb = ((kt) * KDR) >> 3;                                            \
        for (int x = threadIdx.x; x < MMQ_BI * 9; x += blockDim.x) {           \
            int r = x / 9, u = x % 9;                                          \
            int j = j0 + r;                                                    \
            const uint8_t* src =                                               \
                W + (size_t)j * ((size_t)nsb * 144) + (size_t)sb * 144         \
                  + u * 16;                                                    \
            uint8_t* dst = qb8 + ((size_t)(b) * MMQ_BI + r) * 144 + u * 16;    \
            gemm_cp16((__half*)(void*)dst, (const __half*)(const void*)src,    \
                      j < od && sb < nsb);                                     \
        }                                                                      \
        /* per-(row, chunk) scales from global (a few B per unit, L2-hot) */   \
        ...  get_scale_min_k4(sg, blk + 4, &sc, &m);                           \
             sds[...] = h2f(*(const uint16_t*)blk) * (float)sc;                \
             sdm[...] = -(h2f(*(const uint16_t*)(blk + 2)) * (float)m);        \
    } while (0)

    RAW_STAGE(0, 0);
    gemm_cp_commit();
    int buf = 0;
    for (int kt = 0; kt < nktile; ++kt, buf ^= 1) {
        if (kt + 1 < nktile) RAW_STAGE(kt + 1, buf ^ 1);  // prefetch the next k-tile
        gemm_cp_commit();
        gemm_cp_wait1();          // at most one prefetch group in flight
        __syncthreads();          // scales (plain stores) + landed bytes visible
        for (int kd = 0; kd < KDR; kd++) { /* mma loop: unpack nibbles in registers */ }
    }

Against R1: the staging section shrinks from "read 144 B + 4 shift/masks + write expanded words + decode scale" to "five 8 B + nine 16 B cp.async + a 3–6 B scale read"; the dequant ALU moves entirely into registers in the mma loop, overlapping tensor-core issue.

The quantize pass rebuild (current tree from src/cuda_kernels.cu:692, matching d9d626a's change; the kernel comment itself records the bitwise justification):

__global__ void quantize_q8_0_pad40(
    const float* __restrict__ x, uint8_t* __restrict__ y, int dim, int nt
) {
    ...
    // P6: tree-reduced amax (the serial fmaxf chain was latency-bound)
    // and 16B loads / 4B register-packed stores. Math is bit-identical:
    // max is exact for any association, the rintf pass is unchanged.
    float4 sv[8];
    #pragma unroll
    for (int v = 0; v < 8; v++)
        sv[v] = *reinterpret_cast<const float4*>(src + 4 * v);   // 16B load
    float am = 0.0f;
    #pragma unroll
    for (int v = 0; v < 8; v++)                                   // 4-way tree amax
        am = fmaxf(am, fmaxf(fmaxf(fabsf(sv[v].x), fabsf(sv[v].y)),
                             fmaxf(fabsf(sv[v].z), fabsf(sv[v].w))));
    ...
    uint32_t packed[8];
    #pragma unroll
    for (int v = 0; v < 8; v++) {
        const float* e = &sv[v].x;
        uint32_t p = 0;
        #pragma unroll
        for (int j = 0; j < 4; j++) {
            int q = int(rintf(e[j] * di));                        // rintf as-is
            q = max(-128, min(127, q));
            p |= (uint32_t)(uint8_t)(int8_t)q << (8 * j);         // register packing
            s += q;
        }
        packed[v] = p;
    }
    #pragma unroll
    for (int v = 0; v < 8; v++)
        *reinterpret_cast<uint32_t*>(dst + 4 + 4 * v) = packed[v]; // 4B store ×8

The wide tile's warp mapping (git show a41eac0, narrow's wm = warp >> 1, wn = warp & 1 replaced by 4×2 subs):

const int wm = warp >> 2, wn = warp & 3;   // wide: 4x2 subs of 32x32
const int i0w = wn * 32;                   // each warp covers 32 od rows
const int j0w = wm * 32;                   // each warp covers 32 token columns
...
float sum[32] = {0.0f};   // [nh][h][l]: 4 B-frags x 2 A-frags x 4 C regs

The smem-cap guard (the phantom's fix) (git show ef9d5b4 -- src/cuda_kernels.cu, the original r8-era shape; the current tree's launcher has since been rewritten by r14's 16-chain layout with a recomputed smem budget, but the "check attr-set, check launch, return 0 on refusal" structure survives to this day):

extern "C" int launch_mmq_raw_wide_nt(...) {          // void → int: refusals reportable
    (void)type_id;
    // KD=8 needs 2*8*128*40 + 36.9KB + 16KB = 135KB — over the ~99KB
    // opt-in cap; an over-cap request fails SILENTLY (attr + launch both
    // ignored) and the GEMM silently writes nothing. Guard it: KD=4
    // (86KB) is the only feasible wide depth. Returns 0 when refused so
    // the caller can fall back to the narrow raw kernel.
    if (kd > 4) return 0;
    dim3 grid((nt + 127) / 128, (od + 63) / 64);
    const int smem = 2 * 4 * MMQ_WBI * 40 + 2 * MMQ_WBI * 144
                   + 2 * 2 * 4 * MMQ_WBI * 4;
    cudaError_t e = cudaFuncSetAttribute(
        reinterpret_cast<const void*>(&mmq_raw_wide_nt_kernel<4>),
        cudaFuncAttributeMaxDynamicSharedMemorySize, smem);
    if (e != cudaSuccess) {
        cudaGetLastError();
        return 0;                                     // attr failed: refuse + clear error
    }
    mmq_raw_wide_nt_kernel<4><<<grid, 256, smem, stream>>>(w, q8, c, nt, od, id);
    e = cudaGetLastError();
    if (e != cudaSuccess) {
        fprintf(stderr, "minfer/cuda: mmq raw wide launch failed: %s\n",
                cudaGetErrorString(e));
        return 0;                                     // launch failed: refuse + report
    }
    return 1;
}

The Rust side correspondingly turned "launch succeeded" into an explicit protocol (git show ef9d5b4 -- src/cuda.rs): wide_ok = launch_mmq_raw_wide_nt(...) == 1; if !wide_ok { /* fall back to narrow raw */ }.

3.3 Pitfalls

  • The phantom 2124: a double silent failure. The wide tile's first measurement was 2124–2124 tok/s — 4.5× faster than narrow raw's 472 and nearly matching the f16 GEMM path's 2318 — numbers so good they were suspicious. The truth: KD=8 needs 135 KB of dynamic smem, over the ~99 KB opt-in cap; cudaFuncSetAttribute returned an error code that nobody checked, and the kernel launch then failed on the smem oversize too, its error swallowed by the subsequent error-clearing convention — the GEMM did not write a single byte. The 2124 tok/s wall clock was "the sound of no GEMMs running" (verbatim). The mmq_w4k parity failure the wide tile hit back at r7 (max diff 82.9, "partial-sum signature") was a fragment of the same root cause — over-cap garbage, not a mapping bug: at KD=4 the same mapping went parity all-green, proving the mapping itself correct.
  • The semantic chain of silent failure: cudaFuncSetAttribute unchecked → the launch error swallowed by the cudaGetLastError() clearing convention → the Rust side receives "success" → the A/B numbers enter the record. Each link was "reasonable" on its own; combined they produce a phantom that can pollute decisions. The fixed protocol: the C side explicitly returns 1/0, attr and launch are each checked, refusals print to stderr; the Rust side falls back to narrow on a 0 return.
  • Hypothesis-cycling without ncu: ncu was not yet available on this device (GB10) — it only worked from r13 — so phantom triage had to run on parity diffs and conservation reasoning. The r8 record states the risk explicitly: "further MMQ work is hypothesis-cycling" — which directly pushed r9 toward reading llama.cpp's source.

4. Verification

  • cuda_prefill_mmq_parity sweep through the raw path: the env-routed q4_K sub-case runs the raw kernel and compares per-cell against the same host reference (defends: staging/mapping moving the wrong bytes).
  • 7B greedy token identity vs the f16 path: the whole generation token-for-token identical (defends: accumulation-order drift amplifying into behavioral differences over long sequences).
  • suite 169/0 (defends: revert/guard changes introducing a regression).
  • Guard refusal protocol: over-cap requests return 0 + Rust falls back to narrow (defends: phantom recurrence — numbers must be able to prove the kernel actually ran).
  • Wide tile KD=4 parity all-green: the evidence decoupling r7's w4k failure from mapping correctness (defends: discarding a correct mapping design by mistaking environmental garbage for a mapping bug).

5. Results

Raw kernel (LANDED, MINFER_MMQ_RAW=1 opt-in):

  • 7B @2K CLI prefill: 472 vs 441 tok/s (+7%), at KD=8; KD=4 = 440 (depth still wins after staging ALU zeroed).
  • GEMM kernel: 20.2 vs 23.0 ms; quantize pass: 129 → 74 ms (tree amax + packed stores, bit-identical).

Wide tile (honest matrix, all parity-clean, interleaved in the same window):

Shapetok/s
narrow cp.async KD=8472
wide KD=4 (86 KB, the only feasible wide depth)428
narrow KD=4427
R1 word-stage KD=8441

Both of the wide tile's theoretical gains were falsified: halving B traffic bought no time (L2 absorbs the B re-reads — B's DRAM traffic was never binding); feeding 2× mma per fragment word didn't pay either. The constraint-reranking conclusion: the common bottleneck is the per-32-k-chunk inner-loop overhead (the rescale FMA chain, scale smem reads, sync cadence) — no staging scheme can fix it.

FA KV L2-prefetch probe (REVERTED, 87bade0): null (2319–2345 vs the 2345 baseline). Veto mechanism: GQA's 7 q-heads/kv-head sharing keeps KV tiles naturally L2-hot, so the stage stall is not DRAM-bound; fa_prefill_f16kv stays at its P5 state (1.86 ms/layer, ~6% of the prefill wall). Retry conditions: only when the KV tiles' L2 residency is broken (e.g. much longer contexts or less q-head sharing) does prefetch have latency left to hide.

MMQ state (as of r8): best 472 is still 4.9× from the f16 GEMM path's 2318; the default f16 path untouched (2320–2370). The next step's information gap is not another staging knob but "what does a fast MMQ look like" — leading to r9's reference decode.

6. Lessons

  1. For a suspiciously fast wall-clock number, first prove the kernel actually ran: check both attr-set and launch, report failures explicitly; guarding the smem cap is the launcher's duty, not the kernel's.
  2. B-traffic reduction is a dead lever on L2-resident re-reads — ask at which level of the memory hierarchy the re-reads happen before paying a tile-shape price for them.
  3. "No speculative changes": a null probe reverts, even if the change looks harmless; harmless ≠ beneficial, and a negative change left in the tree only pollutes later attribution.
  4. Do groundwork before it gets hot: quantize was only 2.7% of the 4.7 s MMQ wall, but the moment the raw kernel landed it became a visible share — and in the r34 era it grew into one of the main levers.

← 12 · Index · 14 →

14 · r9 llama.cpp MMQ reference decode; shape axis closed (MEAS-ONLY)

Result: the reference source (the mmq-config-ampere.cuh family) decode complete: 256 threads, occupancy 1 (targeted), SRAM tile I=128 × J≤128, ITER_K=256, synchronous staging (no cp.async), float/int accumulator sum[64], 16 mma.m16n8k32 per k01 sub-iteration — an instruction model of ~0.018 inst/MAC/thread against our 0.133. All six shape points measured (all parity-clean): narrow cp.async KD=8 481 is the local optimum, and the shape axis closes with evidence; the residual lever points at llama's pre-arranged mma-fragment layout. Commit: 84831d2 (wip: sync-wide variant + shape matrix; that variant was later overwritten by r12's 16-chain rewrite and is not kept in the current tree). Date: 2026-09-01.

1. Background — where things stood

The MMQ battle state when r8 closed: the raw-byte kernel had pushed R1's 441 to 472, but that was still 4.9× from the f16 GEMM path's 2318; the wide tile and B-traffic hypotheses were out; the FA probe was null. The r8 record wrote down this phase's real predicament: ncu was unavailable on this device (GB10), and "further MMQ work is hypothesis-cycling" — with no counters, every lever could only be blind-tested through the parity gate + interleaved A/B, one hypothesis at a time, extremely costly and unable to localize "which instruction class is slow".

The only reliable high-density information source at hand was the opponent's own source code: llama.cpp's MMQ runs at ~30 TMAC/s on the same chip (GB10) — f16-GEMM class. Every one of its design decisions — how big a tile, whether staging uses cp.async, where scales live, how fragments are arranged — is written in ggml-cuda/mmq*.cuh. r9's choice: stop blind-testing and read the reference source as a profile, producing an instruction-level model of "why it is fast", then verify one separable hypothesis with a controlled experiment (sync-wide).

Meanwhile there was a methodology gap to close: the shape axis had only two measured points at the time (narrow 64×64, wide 128×64), and "which shape is optimal" and "why is it fast" are two independent questions. Without sweeping the shape axis first, any within-point optimization could be working on the wrong peak. r9 did both.

2. Principle — the GPU mechanism

This section is r9's decode output (verified at source-line level, later formalized by the companion analysis docs/LLAMA-CPP-MMQ-ANALYSIS.md §1–§6; line numbers below come from that document's verification system). Layer by layer:

2.1 Configuration and tile geometry

q4_K's instance on NVIDIA (Turing+) is mul_mat_q<GGML_TYPE_Q4_K, 128, false>, with the config-table entry (mmq-config-ampere.cuh:172; on GB10 the Blackwell table falls through to the Ampere table):

CASE(GGML_TYPE_Q4_K, 256, 1, 128, 128, GGML_CUDA_MMQ_SRAM_LAYOUT_Q8_1,
     MMQ_ITER_K, true, false)
→ nthreads=256, occupancy=1, I=128 (od), J=128 (tokens),
  sram_layout=Q8_1, K_vram=MMQ_ITER_K=256, stream_k=true
  • 256 threads (8 warps), occupancy 1 is "targeted" — it is not a failure to fit more blocks; the design deliberately budgets smem and registers for 1 block/SM, trading tile size for per-block resource headroom.

  • Tile 128×128: each block covers 128 od rows × 128 tokens; the warp mapping is 4 od-groups × 2 warps per group, rows_per_warp = 32, ntx = 2. Each warp covers 64 token columns (J=128 split across a warp pair) and 32 od rows per warp.

  • Accumulator: sum[J*I/(nwarps*warp_size)] = 128·128/256 = sum[64] (int32) — 64 accumulator slots per thread, i.e. 16 tile<16,8,int> C fragments (ne = 4).

  • Barrier cadence: 4 barriers per 256-k iteration (load_tiles → stage first half of y → barrier → vec_dot(k00=0) → barrier → stage second half of y → barrier → vec_dot(k00=32) → barrier) = 2 per 128-k — the same cadence as our KD=8.

  • stream-k and fixup: stream_k=true in the config means the host launch (launch_mul_mat_q, mmq.cuh:1393-1473) chooses between an xy-tiled grid and a stream-k grid by nsm (SM count): when the wave tail doesn't form whole blocks, some blocks take an extra tile and write partial sums to global, and a separate fixup pass (~34 µs) merges the cross-block fp partial sums in order. This is the scheduling tax llama pays for "occupancy=1 but big tiles and few waves" — the same territory we later probed in r24 (persistent blocks); llama's answer was "pay the fixup", our measurement was "there is no wave tail to flatten".

2.2 The q8_1 activation pipeline: a once-per-GEMM standalone quantization kernel

Activations (f32) are quantized into block_q8_1_mmq by a standalone kernel, quantize_mmq_q8_1 (quantize.cu:458, inside the host-side ggml_cuda_mul_mat_q, once per GEMM call, not per block), before the kernel launches:

  • block_q8_1_mmq = a 128-element block (QK8_1_MMQ = 4·QK8_1): a leading 16 B scale union (d4[4] / ds4[4] / d2s6[8]) + int8_t qs[128], sizeof = 144 B. The layout comment says it outright: 128 values chunked, transposed, each block padded by 16 B with the pad reused as the block scale and partial sum ("d/ssum in pad bytes"). Q4_K/Q5_K use the DS4 layout: half2 ds4[4], one 16-bit scale + one 16-bit partial sum per 32 values (d0,s0,d1,s1,…).
  • Inside the kernel the activation tile tile_y has row stride MMQ_TILE_Y_K = 36 ints = 144 B — the smem row layout equals the global block_q8_1_mmq layout: [scale || qs], scale read as (half2*)y, the qs plane at y+4.
  • Design implication: the activation operator finishes all "dequant preparation" (scale and ssum in place) before entering the kernel, so the in-kernel activation operand needs zero dequantization; dsB.y (the partial sum) is a free input for the rank-1 correction term.

2.3 Weight staging: the raw-nibble smem layout

load_tiles_q4_K (mmq-load-tiles.cuh:703-812) stages the 144 B raw Q4KB weights into tile_x, keeping raw nibbles (one 0..15 nibble per byte), unexpanded into signed int8 and un-centered:

// mmq-load-tiles.cuh:736-737 — the 0x0F mask isolates nibble by nibble, stored per byte
x_qs[i*sram_stride + 16*(txi/8) + txi%8 + 0] = (qs0 >> 0) & 0x0F0F0F0F;
x_qs[i*sram_stride + 16*(txi/8) + txi%8 + 8] = (qs0 >> 4) & 0x0F0F0F0F;
  • dmin is not subtracted at staging — it is folded into the scale: x_dm[...] = (bxi->dm · make_half2(1.0f, -1.0f)) · make_half2(sc8[l], m8[l]) (:772-777), storing (d·sc, −dmin·m) in a half2. No __vsubss4, no I2F chain anywhere in staging.
  • smem row stride: sram_stride = 2·MMQ_TILE_NE_K + 2·MMQ_TILE_NE_K/QI8_1 + 4 = 64 + 8 + 4 = 76 ints (304 B); the "+4" is a 16 B pad that rotates the bank phase of adjacent rows (defending against ldmatrix column conflicts); the nibble plane occupies the first 64 ints and the half2 scale plane follows at offset 64.
  • Staging is synchronous: plain global→(register)→smem stores ordered by __syncthreads; the classic MMQ path has no cp.async / TMA pipeline — latency hiding rests entirely on resident warps (at occupancy=1 that is 8 warps × 32 lanes of ILP plus intra-tile reuse).

2.4 The compute loop and numerics: raw nibbles into mma, dmin as a rank-1 fold

The compute loop is the generic ggml_cuda_mmq_vec_dot_q8_1_q8_1_mma (mmq-vec-dot.cuh:369-442):

  • vec_dot is called once per 32-k chunk and contains 64 mma: j0 8 steps × k01 4 steps × n 2 steps (mma.m16n8k32.row.col.s32.s8.s8.s32, int32 accumulate). Because the sum index doesn't depend on k01, every sum slot is accumulated 4×.
  • A fragment: ldmatrix.m8n8.x4 loads straight off the raw-nibble rows (load_ldmatrix(A[n], x_qs + ...·sram_stride + k0, sram_stride)) — ldsm consumes the nibble bytes directly, no pre-expansion needed.
  • B fragment: plain LDS (source comment verbatim: "faster than load_ldmatrix").
  • The fragment-reuse economics (the most productive ledger entry of the r9 decode): each vec_dot loads 8 A fragments (tile_A A[ntx][4], loaded once before the j0 loop and reused across the whole loop) and 32 B fragments (one per (j0, k01), reused across the 2 n iterations); the 64 mma consume them. Per MAC: llama issues ~0.125 LDSM per IMMA (r36's accounting), and the A-fragment reads amortize another 2× over the 32 od rows a warp covers. By contrast our kernel paid ~0.5 shared reads per IMMA — and this ratio difference is a property of tiling (how much a warp covers, how many mma a fragment feeds), not of staging or numerics. r9 accordingly put "raise fragment reuse" on the residual-lever list.
  • Per-chunk rescale (fp32): sum[i] += dmA.x·dsB.x·C.x + dmA.y·dsB.y, where dmA.x = d·sc (the weight-quantization product), dmA.y = −dmin·m, dsB.x = d (activation scale), dsB.y = ssum (activation partial sum). The dmin correction is a rank-1 term on the (token, od-col) plane, applied at accumulate time — that is the cost structure of "mma eats unsigned nibbles, centering deferred to rescale": staging saves ALU in exchange for two FMULs per chunk.
  • get_scale_min_k4 semantics (ggml-quants.c:880-887): q4_K's 12 B packed scales decode into per-32-value-sub-block (d_scale, dmin) — six 6-bit codes + six 4-bit combination codes; a value v ∈ [0,15] dequantizes to d·s·v − dmin·m.

The q6_K specialization — a different scale flow from q4_K (noted at r9, cashed in by Era D's q6_K line): q6_K has its own sram layout (stride also 76 ints, different composition: 64 + 1 + 4 + 7), its own load_tiles_q6_K and vec_dot_q6_K_q8_1_mma. Two key differences: the nibbles are fully centered at staging (__vsubss4(ql | qh, 0x20202020), subtracting 32 per byte into signed int8 ∈ [−32,31] — there is no dmin term to defer); and the scales are split into two levels (one float d per row + one int8 sc per 16-value sub-block), so the rescale is two-stage: tmp = (C0·scA0 + C1·scA1)·dB accumulated per j0 step, finally sum += tmp·dA. One MMQ framework, two numeric flows — which explains why q6_K's mma kernel could not be "casually" derived from q4_K's, and why (r38) it indeed needed a standalone BT kernel.

2.5 The instruction model: 0.018 vs 0.133

Converting the structure above into warp instructions per MAC: llama's MMQ issues about 0.018 instructions per thread per MAC, our raw kernel about 0.133 — a 7× gap. That is the origin of their ~30 TMAC/s against our ~6–8 TMAC/s (r9's verdict: "the ratio, not tile shape, is their speed"). The ratio's composition is worth taking apart, because it names which instruction classes "support" rather than "compute":

  • The compute core is full-price: one int8 mma multiply-add per MAC — this part is bit-identical between the two (r13/r20 later confirmed with counters that IMMA/ FFMA bytes were at parity). The gap is never in the mma itself.
  • The gap is entirely in the support stream: fragment address arithmetic, scale smem reads and I2F/FMUL, nibble-unpack shift/mask, barriers and loop overhead. llama compresses it with four mechanisms:
    1. Fragment reuse — 8 A-frags reused across 64 mma, 32 B-frags each reused 2×, pushing per-MAC load instructions to the minimum;
    2. The raw-nibble plane halves smem bytes — the expanded-byte form (our then qb8 per-k int8) is ~2× the raw form; on small tiles that is the difference between 1 and 2 blocks/SM;
    3. The half2 scale short path — rescale goes through 16-bit half2 rather than a full fp32 I2F→FMUL chain, only a handful of instructions per chunk;
    4. Tight indexing — the 32 od rows × 64 token columns warp shape keeps A reads, B reads, and address ALU all at amortized lows.
  • The costs it accepts are part of the model too: mma consuming unsigned nibbles (accumulator carries an unsigned dot product), the dmin correction deferred into a rank-1 rescale term, stream-k's fixup pass, and the tile-size/ubatch coupling.

The mechanisms sustaining the low ratio were later confirmed item by item by the r13/r25 ncu censuses (per-GMAC warp instructions 6.06 vs 10.14 M, the delta all support instructions) — but the direction was already set at r9: the optimization target changed from "find a faster shape" to "push per-MAC support instructions down and expose the ILP chains".

3. Implementation

3.1 Design choices (why this shape and not another)

  • Line-level verification, not impressions: every conclusion is pinned to source line numbers (config table :172, sum size mmq.cuh:903, vec_dot main loop mmq-vec-dot.cuh:414-440, nibble mask mmq-load-tiles.cuh:736-737, scale fold :772-777, the ne formula mma.cuh:226-227, C-fragment lane mapping mma.cuh:245,262). This discipline later crystallized into the companion analysis docs/LLAMA-CPP-MMQ-ANALYSIS.md.
  • Readings must be branch-aware: two earlier misreadings were corrected by the source — (1) "C fragment ne = I·J/64 = 2 regs" is the AMD MFMA branch's formula (mma.cuh:107-108); NVIDIA Turing+ is ne = I·J/32 = 4 regs (formally fixed at r11); (2) the scale word had been counted as 1 int, but the DS4 layout is really 4 ints (half2 ds4[4]) — which decides that tile_y's row stride is 144 B and not smaller, and in turn the bank phase and ldmatrix behavior.
  • Shape-semantics clarification: the campaign profile's "nt-512" is not a dispatch threshold — it is llama.cpp's ubatch size (-p 2600 → 512-token ubatch → each mul_mat_q sees 4 J=128 token tiles); J=128 is chosen coupled to the 512-token ubatch (mul_mat_q_switch_J takes the largest J that fits). That explains why "copy the tile shape" is not necessarily meaningful for minfer — our call shapes and tile-selection constraints differ.
  • Deriving the instruction model: dividing tile geometry (per-warp mma count, fragment loads, scale operations) by per-warp MACs gives the 0.018 vs 0.133 ratio — an ncu-independent "upstream quantity of speed" derivable from source alone.

3.2 Key code

The decode produced one testable hypothesis: can "staging style (sync vs cp.async) + footprint (single vs double buffer → occupancy)" alone explain the gap? r9 ported synchronous staging onto our wide kernel as the control: single buffer, 54 KB, 2 blocks/SM, plain loads + one __syncthreads per tile (git show 84831d2 -- src/cuda_kernels.cu, excerpt):

// Single-buffer sync-staged layout — KD=8 totals 54,272B so TWO blocks
// fit per SM (16 resident warps hide the staging latency):
//   qa8   [KDR][128] x 32B  chunk qs only (d/ssum in sda_q)
//   sda_q [KDR][128] x 8B   (d f16 | ssum i16) packed
//   qb8   [64][144]         B super-blocks (64 od-rows per block tile)
//   sds   [KDR][64] f32, sdm likewise
uint8_t* qa8 = mmq_raw_sh;
uint32_t* sda_q = reinterpret_cast<uint32_t*>(qa8 + KDR * MMQ_WBI * 32);
uint8_t* qb8 = reinterpret_cast<uint8_t*>(sda_q + KDR * MMQ_WBI * 2);
float* sds = reinterpret_cast<float*>(qb8 + 64 * 144);
float* sdm = sds + KDR * 64;
...
#define RAW_STAGE(kt)                                                          \
    do {                                                                       \
        /* llama.cpp-style synchronous staging, single buffer: plain global    \
         * -> smem loads, one syncthreads orders them. At 128x64 tiles the     \
         * smem is small enough for 2 blocks/SM - latency hiding comes from    \
         * occupancy, not prefetch depth. */                                   \
        for (int x = threadIdx.x; x < MMQ_WBI * KDR * 8; x += blockDim.x) {    \
            int u = x % 8, r = (x / 8) % MMQ_WBI, kd = x / (8 * MMQ_WBI);      \
            int tok = i0 + r, c = (kt) * KDR + kd;                             \
            unsigned v = 0;                                                    \
            if (tok < nt && c < nchunk)                                        \
                v = *(const unsigned*)(q8x + ((size_t)tok * nb32 + c) * 40     \
                                       + 4 + u * 4);                           \
            *(unsigned*)(qa8 + ((size_t)kd * MMQ_WBI + r) * 32 + u * 4) = v;   \
        }                                                                      \
        ...

(In the double-buffer version, the corresponding __syncthreads() comment also changed from "scales + landed bytes visible" to "single-buffer stage visible to all warps".) The experiment held the variables: tile shape and the raw-byte mma math unchanged, only the staging/footprint/occupancy axis swapped — if it came significantly closer to narrow cp.async, the gap was on the staging axis; if it still lost, the gap was in the instruction model.

3.3 Pitfalls

  • sync-wide still lost (462–466 vs 481): porting llama's staging philosophy did not pay. That is not an experimental failure — it is precisely the hypothesis test's negative output: the gap is not on the staging axis.
  • The compact variant's serialization trap: compact in the shape matrix (cp.async 128×64, A side compressed to 32 B qs-only chunks, sda_q packed scales) measured only 410 — the synchronous load of sda_q serialized the A-side reads. The bytes the compressed layout saved could not buy back the lost load parallelism.
  • Calibrating the "16 mma/chunk" reading: the r9 record's "16 mma per warp per chunk" counts per k01 sub-iteration (8 j0 × 2 n = 16); the full vec_dot (32-k) is 64 (× the 4 k01 steps). The companion analysis §6 makes this multiple explicit — when reading a reference kernel the "chunk" boundary must be pinned first, or the instruction model will be off by 4×.
  • The shelf life of wip code: 84831d2 is a wip commit; the sync-wide variant was overwritten at r12 by the 16-chain rewrite and not kept; this doc's code excerpts were verified against that commit's diff.

4. Verification

  • sync-wide parity green: cuda_prefill_mmq_parity fully passing (defends: the single-buffer + plain-load rework moving the wrong bytes — it has none of the raw path's double-buffer semantics to lean on).
  • The six-shape matrix all parity-clean (@2K, MINFER_MMQ=1, interleaved within one session) (defends: correctness differences contaminating cross-shape comparisons; also machine-state drift — all points measured inside one window).
  • narrow cp.async held at 481 as the control: every new shape re-tests the anchor (defends: baseline drift turning "the new shape is worse" into an artifact).

5. Results

The six-shape matrix (GB10 @2K, all parity-clean):

#Shapetok/sNotes
1narrow cp.async 64×64 KD=8481local optimum (r8's raw kernel)
2sync wide 128×64 KD=8464llama-style sync staging, single buffer 54 KB, 2 blocks/SM
3compact cp.async 128×64 KD=8410sda_q sync load serialization
4wide cp.async 128×64 KD=4428r8's honest wide-tile value
5narrow cp.async 64×64 KD=4427
6R1 word-stage 64×64 KD=8441the starting point
  • The shape axis closes with evidence: the six points cover the main axes of {narrow, wide, compact} × {KD=4, KD=8} × {cp.async, sync}, and the optimum is the starting point itself (481). The B-DRAM-halving theory is dead (L2 absorbs the re-reads); wider tiles pay more in staging/sync than they save.
  • The staging philosophy does not port: llama is still fastest with synchronous staging + occupancy=1 because its per-MAC instruction stream is 1/7 of ours — it doesn't need prefetch to hide latency, its mma density is high enough. Our kernel at 0.133 inst/MAC gets patched by every staging scheme for a problem that is really an over-dense instruction stream.
  • Adopted (feeding later steps):
    1. The instruction-model lens — the optimization target switches from "shape/staging" to "instructions per MAC and ILP depth". Direct products: r12's 16-chain warp tile + ldmatrix (wide-16 KD=4 1020–1058 vs 441–481, ~2.3×) and r14's ldmatrix B-fragments (+18.5% / +23–30%).
    2. "The pre-arranged mma-fragment layout can be materialized at load/quantize time" — the residual lever r9 named. Mechanism: the fragment ldmatrix consumes has an exact word→lane arrangement; re-arranging it live inside the kernel, per block, costs exactly the address/convert ALU r9 wanted to eliminate. Moving that rearrangement to a cold path (weight load time, or the activation-quantize prepass) does it once, and the hot loop is left with only ldmatrix itself. The A-side version was cashed in at r34 (quantize-transpose prepass, +9.72%, the implementation comment explicitly cites llama's quantize_mmq_q8_1 design); the B-side "pre-arranged fragments at load time" was cashed in at r12/r14 via ldmatrix + the slot-major smem layout.
    3. The q8_1-style "once-per-GEMM standalone quantize prepass + self-describing block layout" — became the structural prototype of minfer's pad40 prepass lineage (r34 → r51/r52).
  • Vetoed/closed: synchronous staging (464 < 481, no longer a main axis); the compact A layout (410); further peak-hunting on the shape axis (the matrix is closed; the local optimum is the current global optimum).
  • The r10–r11 follow-up control (the reference inner-loop decomposition ported and still flat) squeezed the residual further to the three items ILP depth / ldmatrix / tile — that is step 15's content; r9's contribution was fixing the search space: "instruction-level structure is the suspect class".

6. Lessons

  1. Copy the instruction model, not the shape: the 0.018 vs 0.133 ratio is the speed source; transplanting the opponent's tile/staging shape onto a kernel with a different instruction stream measures 464 vs 481.
  2. Sweep the axis first, then work within a point: a six-point shape matrix eliminates a whole family of hypotheses at once — far cheaper, and far more honest, than five rounds of incremental optimization on one point.
  3. When the profiler is unavailable, the reference source is the highest-density source of fact: a line-verified instruction model can replace counters in localizing the search space (and was later confirmed by the ncu census).
  4. Pin the counting boundaries before reading a kernel (the chunk/k01/vec_dot multiples, the NVIDIA vs AMD ne formula branches, the scale word's int count) — off by one multiple in the reference reading, off by an order of magnitude in the conclusion.

← 13 · Index · 15 →

15 · r10–r11 — reference inner-loop decomposition port; ILP reading verification (REVERTED / MEAS-ONLY)

Result: porting llama.cpp's q8_1×q8_1 generic decomposition verbatim into our MMQ kernel: parity green, but 462–468 vs narrow's 470 tok/s (@2K) — the same decomposition still runs ~6.4 TMAC/s (llama ~30): the gap is not in the math. r10 pinned the residual to exactly three items — ILP depth / ldmatrix / tile — which became r12's execution recipe. r11 verified the tile ne = I·J/32 (4 registers per m16n8k32 C fragment), establishing the "16 independent mma chains" reading. Commit: 025a69f (r10, docs-only), 83e5580 + 6f29e65 (r11, docs-only). All three are docs commits — the kernel code changes were swallowed by a post-checkout hook revert after measurement and never became code commits. Date: 2026-09-01.

1. Background — where things stood

By the end of r9 the q4_K MMQ campaign stood at the "shape axis exhausted" node. R1 (40e97c9) had built the int8 tensor-core prefill GEMM — parity-clean but only ~6.1 TMAC/s; r7–r8's raw-byte kernel (the d440d16 family) swapped staging to bare bytes and unpacked nibbles in registers, reaching 472 tok/s; after r8's wide tile (128-token block) landed, r9 finished the whole shape matrix: narrow cp.async KD=8 481 (local optimum) > sync-wide 128×64 464 > compact 410 > wide KD=4 428 > narrow KD=4 427 > R1 441. All six shapes parity-clean: under the current inner-loop structure, tile/staging shape is no longer a lever.

Two details from this trajectory belong in the record before r10. One: R1's first measurement was 155 (co-tenant) / 412 (quiet), and 441 was only calibrated in the r7–r8 window — the same opt-in path swings 2.7× across machine states, the direct reason the campaign later made "same-window interleaved A/B" a hard rule. Two: r8's wide tile's first 2124 tok/s was a phantom of silent smem over-cap (attr-set and launch both failing silently, the GEMM writing nothing); after the guard the honest number was 428 — "suspiciously fast" readings should be checked against resource caps, a reflex r10's measurements would need again. The six-point matrix in one sentence: every shape and staging variant sits in the narrow 410–481 band while llama runs ~30 TMAC/s — the difference is structure, not parameters.

r9's real output was a completed read of llama.cpp's MMQ reference implementation (the mmq-config-ampere.cuh Q4_K branch): 256 threads, targeted occupancy 1, SRAM tile I=128 od rows × J≤128 tokens, ITER_K=256, synchronous staging, 16 mma.m16n8k32 per warp per 32-k chunk. The instruction model reconciles to llama ~0.018 inst/MAC/thread vs our 0.133 — a 7× gap. That, not a magical tile shape, is the source of their ~30 TMAC/s.

r6's campaign target was MMQ from 6.1 TMAC/s to ≥24 (f16 parity) or ~30 (llama parity): f16 parity deletes the 56 ms convert pass (whole-prefill ~2670 tok/s), llama parity gives ~3250. We were still 5× from the f16 default path (same window 2320–2370) and every shape point had been tried — without eliminating the "different math decomposition" hypothesis first, all subsequent kernel-side work would rest on a wrong foundation.

r10's problem statement is thus concrete: is our kernel doing "the same math" as llama's? The q4_K dot product can be organized in different orders — when nibbles unpack, when dmin folds in, when scales multiply, whether the B side stays per-k signed int8 — and every organization changes the instruction stream. r9 reconciled "instructions per MAC" but not "what each instruction is". r10 ported llama's inner-loop decomposition item by item, made both sides identical in math organization, and measured the remaining gap.

This step also has a special archival situation: r10's kernel edits were reverted by the repo's post-checkout hook after that day's measurement, and the code never became a commit. What survives is only docs commit 025a69f, carrying the full redo recipe and all measurement numbers. r11's two commits are likewise docs-only: first declaring r10's 16-chain claim unverified by code (83e5580), then verifying it against llama.cpp's mma.cuh (6f29e65). Both of this doc's "results" are measurement records, not code records.

2. Principle — the GPU mechanism

mma chains and ILP. mma.sync.aligned.m16n8k32 (int8, s32 accumulator) has fixed instruction latency. In a warp, if the next mma's accumulator depends on the previous one's result the two mma serialize; if 16 mma each write their own accumulator (16 independent chains), the warp issues all 16 back-to-back before returning to consume the first result — the issue-window depth is the chain count. That is what "chain count ≈ ILP depth" means.

Comparing the chain structures arithmetically. llama: SRAM tile 128 od × ≤128 tokens, 256 threads, 16 m16n8k32 per warp per 32-k chunk, backed by 128 accumulator registers (int C fragments + float sum). Ours (r8 wide shape): 8 warps × 32×32 per-warp tile, 2×4 = 8 chains per chunk, ~64 accumulator registers. Half the chains means an issue window half as deep; and occupancy is 1 block/SM in both kernels — no other warp exists to fill the mma latency holes, so chain depth IS throughput. llama dares to spend 128 accumulator registers precisely because it targets occupancy 1 and its register budget (255/thread at 256 threads) cannot be exhausted.

C-fragment register count: ne = I·J/32. Each m16n8k32's C operand is 4 int registers per thread (16×8 output / 32 lanes = 4). The constant r11 had to verify: is llama.cpp's tile<16,8,int>::ne 2 or 4? If ne = I·J/64 = 2, one tile::mma() issues two hardware mma, 16 calls could be 32 chains or another structure, and r10's reading is void; if ne = 4, the 16 chains are confirmed and r12's "16-chain warp tile" plan stands. Verdict: I·J/64 is the AMD MFMA branch (one MFMA covers a larger tile, fewer registers per thread); the NVIDIA Turing+ branch is I·J/32, 16×8 → 4.

ldmatrix vs per-lane loads. The A fragment (m16n8k32's A operand, 4 ints = 16 bytes per thread) can come from shared memory in one ldmatrix.m8n8.x4 (1 LDSM instruction, hardware-distributed in mma layout), or each lane can compute its own addresses and issue 4 LDS.32 (8 load instructions per fragment pair). llama uses LDSM; we were on per-lane loads. Not just instruction count — LDSM's address generation reuses one set of matrix coordinates, so ALU overhead is smaller too.

Breaking that cost down. In the per-lane form each lane first computes "where my 4 ints are": row from lane >> 2, column from lane & 3, times the chunk's row pitch — 4 integer ALU plus 4 narrow loads per lane per fragment, and the narrow loads can still hit shared-memory bank conflicts. LDSM replaces the whole set with one instruction: the lane supplies a base address (distributed within the warp by convention), the hardware does the rest. r10 listed this as the second residual item; r12 cashed it in, and r14 later applied the same weapon to the B fragments.

Why the decomposition itself matters. The q4_K super-block dot product has several equivalent forms: unpack timing (all at staging vs item-by-item at mma), dmin fold timing (once per (row, chunk) vs repeatedly per accumulator), B-side form (keep nibbles vs pre-convert to signed int8). Different organizations, different instruction mixes. Only after r10 aligned all these axes with llama and performance did not move an inch did the elimination carry force.

Spelling out llama's decomposition. q4_K's scale is two-stage: the super-block scale d (f16) and the per-32-k-chunk correction dmin. llama has staging unpack the nibbles into per-k signed int8 and pre-fold the dmin term in one pass, so the compute side does only two FMAs per accumulator — one for the weight scale dsv, one for the correction dmv. Our raw-byte kernel of the time (excerpt B's shape) left the nibble unpack at mma time item by item, mixing SHF/AND-class integer ALU into the compute-side stream. The MAC totals of the two organizations are exactly equal; what differs is the composition of the support instructions — r10's question: if that composition is also made identical, does the 5× gap survive?

Why the 462–468 band. In throughput terms: the r9 window's 481 tok/s corresponds to MMQ ~6.4 TMAC/s, and r10's 462–468 lands in the same band — the port changed no throughput number. Meanwhile llama's same decomposition runs ~30 TMAC/s: the same math, 5× the machine time. This "flat" is more informative than any +x%: it crosses the whole "math" line off the residual list.

3. Implementation

3.1 Design choices (why this shape and not another)

r10 changed only the math organization, no shape or staging parameter: the host is the sync-wide kernel r9 had just measured (128×64 tile, single buffer 54 KB, 2 blocks/SM, plain loads + one __syncthreads per tile) — the shape matrix's point closest to narrow (462–466 vs 481), so the post-change delta is directly comparable. Textbook controlled variables: change the tile to llama's 128×128 at the same time and any gain is unattributable.

A second reason for sync-wide: it is the only point in r9's matrix that is "structurally close to llama (synchronous staging, no cp.async pipeline) yet not performance-competitive". r9's "copy the instruction model, not the shape" is purest here — staging was already synchronized, so the remaining differences concentrate in the inner loop, exactly what r10 wanted to isolate.

The ported decomposition (aligned item-by-item with llama's vec_dot_q4_K_q8_1): 1. nibble unpack + dmin fold move into staging, once per (row, chunk) — the compute side never sees a nibble; 2. the compute side only applies pre-loaded scales (two FMAs per accumulator, weight-side scale registers pre-loaded); 3. the B side stays per-k signed int8 (unpacked signed bytes go straight into the mma).

Point 3 in full: q4_K's nibbles are unsigned 4-bit, and feeding them straight into the s8 mma reads the high nibble with the wrong sign; llama's decomposition splits the two nibbles into separate signed byte streams at staging (low nibble & 0x0F, high nibble >> 4), the dmin correction carrying the unsigned-offset compensation. Our old kernel did the same work at mma time (excerpt B's (sg & 1) ? ((n0 >> 4) & 0x0F0F0F0Fu) : (n0 & 0x0F0F0F0Fu)) — redone for every fragment of every chunk; after the port the three-line expression leaves the compute loop and appears once in staging.

3.2 Key code

⚠️ Code survival note: the sync-wide kernel r10 actually edited is no longer in the current tree (hook revert, no code commit). The excerpts come from two verifiable locations: (a) the current tree's mmq_raw_nt_kernel inner loop — the narrow sibling kernel that survived r12 (plus a rank-1 fold in r16), same structural lineage as the r10 era, shown as the "pre-port" chain-structure class; (b) llama.cpp's mma.cuh — r11's verification target. r10's decomposition itself is reconstructed from context per the redo recipe; r12's commit (doc 16) is its true landed form.

Excerpt A · the mma chain unit (current tree, mmq_mma_k32) — the hardware contract r11 verified is, in our code, this one wrapper: one instruction, 4 C registers accumulated in place:

// src/cuda_kernels.cu (current tree)
__device__ __forceinline__ void mmq_mma_k32(int* d, const int* a, const int* b) {
    asm volatile(
        "mma.sync.aligned.m16n8k32.row.col.s32.s8.s8.s32 "
        "{%0,%1,%2,%3}, {%4,%5,%6,%7}, {%8,%9}, {%0,%1,%2,%3};\n"
        : "+r"(d[0]), "+r"(d[1]), "+r"(d[2]), "+r"(d[3])   // 4 C registers = I·J/32
        : "r"(a[0]), "r"(a[1]), "r"(a[2]), "r"(a[3]), "r"(b[0]), "r"(b[1]));
}

Each call consumes one m16n8k32: 16×8×32 MAC / 32 lanes. Chain count = the number of calls per chunk whose accumulators do not overlap.

Excerpt B · the surviving representative of the pre-port chain structure (current tree narrow kernel inner loop) — note the A fragments are per-lane groups of 4 narrow loads (the "8 LDS.32 per fragment pair" class), accumulators grouped by (B fragment × A fragment) into 4 int[4], 4 mma per chunk:

// src/cuda_kernels.cu (current tree, mmq_raw_nt_kernel after r16; same structural class as the r10 era)
for (int kd = 0; kd < KDR; kd++) {
    ...
    // A fragments: int8 lane words straight out of the raw chunk.
    int a[2][4], b[2][2];
    int clow[2][2][4], chigh[2][2][4];
    #pragma unroll
    for (int h = 0; h < 2; h++) {
        const int r0 = i0w + h * 16 + (lane >> 2);
        const uint8_t* p0 = qat + (size_t)r0 * 40 + 4;
        a[h][0] = *(const int*)(p0 + 4 * (lane & 3));        // per-lane narrow loads:
        a[h][1] = *(const int*)(p1 + 4 * (lane & 3));        // 4 LDS.32 per A fragment
        a[h][2] = *(const int*)(p0 + 4 * ((lane & 3) + 4));  // (llama's counterpart is
        a[h][3] = *(const int*)(p1 + 4 * ((lane & 3) + 4));  //  1 ldmatrix.x4)
    }
    // B fragments: unpack the raw nibbles in registers.
    #pragma unroll
    for (int nh = 0; nh < 2; nh++) {
        uint32_t n0 = *(const uint32_t*)(rb8 + ...);         // nibbles unpacked at mma time
        b[nh][0] = (int)((sg & 1) ? ((n0 >> 4) & 0x0F0F0F0Fu)
                                  : (n0 & 0x0F0F0F0Fu));
    }
    #pragma unroll
    for (int nh = 0; nh < 2; nh++)
        #pragma unroll
        for (int h = 0; h < 2; h++)
            mmq_mma_k32(clow[nh][h], a[h], b[nh]);           // 4 independent chains/chunk
    ...
}

r10's port replaced this segment's B side: unpack and dmin fold move into staging (once per (row, chunk)); the compute-side b[] becomes the staged signed int8 directly, scales applied as two pre-loaded FMAs. The reconstructed shape (rebuilt per the redo recipe, not surviving code): staging produces int8 b8[MMQ_BI][32] (per-k signed) plus the pre-folded dsv/dmv, and the compute loop body shrinks to "4 mma + 2×8 FMA". Its only difference from excerpt B is the instruction mix — exactly the variable r10 isolated.

Excerpt B′ · the half of the port left untouched: the two-term rescale (current tree narrow kernel) — however the decomposition is organized, every accumulator must multiply the weight scale (dsv term) and the dmin correction (dmv term); the port moved the B side into staging but kept this scale-application chain, aligned item-by-item with llama's counterpart:

// src/cuda_kernels.cu (current tree, post-r16 state; same structure in the r10 era)
// rescale: identical math/layout to the R1 kernel; A-side
// d/ssum come straight from the raw chunk.
float da_q[4];
int sa_q[4];
#pragma unroll
for (int t4 = 0; t4 < 4; t4++) {
    const uint8_t* at = qat + (size_t)(i0w + (lane >> 2) + t4 * 8) * 40;
    da_q[t4] = h2f(*(const uint16_t*)at);            // A-side super-block scale d (f16)
    sa_q[t4] = (int)*(const uint32_t*)(at + 36);     // A-side ssum (i32)
}
const float dma[4] = { da_q[0] * (float)sa_q[0], ... };

The A-side d/ssum reads straight from the bare chunk (bytes 0–1 of the 40 B/chunk layout are d, bytes 36–39 are ssum) — same lineage as llama's q8_1 A side, and the part the port need not touch: r10 aligned only the decomposition's organization; the scale semantics were already identical on both sides.

Excerpt C · r11's verification target (llama.cpp mma.cuh) — the two ne branches; I·J/64 is the AMD one:

// llama.cpp ggml/src/ggml-cuda/mma.cuh (reference repo, the evidence lines r11 read)
#if defined(AMD_MFMA_AVAILABLE)
        static constexpr int ne = I * J / 64;      // ← the branch r10 first misread (MFMA)
        T x[ne] = {0};
...
#if defined(VOLTA_MMA_AVAILABLE)
        static constexpr int ne = I * J / WARP_SIZE; // NVIDIA half-precision branch: 16·8/32 = 4
        half2 x[ne] = {{0.0f, 0.0f}};

The int8 tile's NVIDIA branch (tile<I,J,int>) is ne = I·J/32, likewise 4 for tile<16,8> — strictly matching excerpt A's 4 C registers. 16 tile::mma() calls = 16 hardware mma = 16 independent chains. r10's reading stands.

3.3 Pitfalls

  • The post-checkout hook revert swallowed the working tree. The day's kernel edits were restored by the hook after measurement completed; the code entered no commit. The salvage was writing the redo recipe and all numbers into docs commit 025a69f. The lesson is procedural: the first action when a measurement session ends is to commit (even docs-only) — code can be redone, numbers cannot.
  • Measurement validity must be argued separately. The hook revert happened after measurement, and the interleaved A/B numbers were recorded into the commit message on the spot — so "the code is gone" does not mean "the numbers are suspect". But it is a question that must be answered explicitly: had the revert come before measurement, this doc would be void. The master table's Status column states exactly this: "REVERTED (edits lost to a post-checkout hook; measurements valid)".
  • The ne reading: wrong first, verified later. r10's record justified "16 chains" via the reference's tile<16,8,int> — but the ne = I·J/64 = 2 it read contradicted the 4-register C fragment. r11's first commit (83e5580) publicly suspended the contradiction instead of glossing over it; the second (6f29e65) located the #if defined(AMD_MFMA_AVAILABLE) branch and resolved it. A compile-time-constant branch selection decided whether the entire r12 plan stood.

4. Verification

  • Parity gate (greedy token identity): the port changes only instruction organization, not the numeric path, so output must be token-for-token identical — defends against "optimization changed the math". r10: parity green.
  • Suite (169/0): whole-model regression — defends against the port breaking non-q4_K paths (the decomposition entered only the sync-wide kernel, but the suite runs everything).
  • Same-window interleaved A/B: 462–468 vs narrow's 470 interleaved within one session window, medians taken — defends against cross-session machine drift (forerunner of the r59b lesson, already consciously practiced here).
  • Default-path isolation: the MINFER_MMQ=1 opt-in gate keeps the f16 default path untouched — defends against experiments polluting production.
  • r11's verification is static, but still verification: the ne reading runs no performance gate; it reconciles against the hardware ISA contract — excerpt A's {%0,%1,%2,%3} 4-register C operand is our code-side evidence, the mma.cuh branch the reference-side evidence, and both must yield the same number (4). Defends against deriving a plan from a wrong hardware model.

A note on this doc's special status: a normal step verifies surviving code (byte-exact dumps, re-runnable greedy streams); this doc verifies a measurement record — the numbers and redo recipe inside three docs commits. The code cannot be re-verified (the hook restored it), which is exactly why every number went into the commit message: the record is the archive. This doc's code has no surviving verification target: all three commits are docs-only; excerpts A/B's line references were checked against the current tree and excerpt C against the reference repo; r10's own code shape rests only on the redo recipe's text (flagged in §3.2).

5. Results

r10 (decomposition port, REVERTED / measurement valid): sync-wide + llama decomposition = 462–468 tok/s @2K, control narrow 470 — delta inside the noise band, ~6.4 TMAC/s on both sides, while llama's same decomposition runs ~30 TMAC/s. Veto mechanism: this was a controlled experiment whose "no difference" outcome is itself the conclusion — the math decomposition is eliminated and the residual converges to three items: 1. ILP depth — llama 16 independent mma chains + 128 accumulator registers, ours 8 chains / ~64 registers; 2. ldmatrix A staging — 1 LDSM vs 8 LDS.32 per fragment pair; 3. tile 128×128 (llama) vs 128×64 (ours).

The revert itself carried no numeric cost (the code never survived anyway); the "under what future conditions is a retry worthwhile" answer is that r12 executed the redo recipe directly — no second r10 needed.

r11 (MEAS-ONLY): no performance numbers. The output is two definitive conclusions: ne = I·J/32 = 4 (the NVIDIA Turing+ branch), and the 16-chain reading stands. Of r10's three residuals, (1) and (2) both depend on this reading — r11 is r12's prerequisite proof, not an optional footnote.

Campaign-level significance: r9 eliminated the shape axis, r10 the math axis; the one suspect class their eliminations converge on — instruction-level structure (ILP depth, load width) — is exactly r12's target. r12's subsequent 2.3× is that reasoning chain cashed in.

Matching the redo recipe against r12's execution item by item shows the complete "handoff" the r10 record produced:

r10's residualr12's execution
ILP depth: 8 chains / ~64 accumulator registers16-chain warp tile: clow[8][2][4] + sum[64] = 128 accumulator-class registers
ldmatrix A staging: 8 LDS.32 per fragment pairA fragments loaded per group by ldmatrix.m8n8.x4 (8 groups covering the 128-token rows)
Tile 128×128block tile 128 tokens × 128 od (the host had been 128×64)

The third item especially shows r10/r11's value: without r11 pinning ne = 4, "16 chains" is an unverified mental calculation, and r12 would not have dared fill the register budget (128 accumulators) to the brim.

6. Lessons

  1. When the same math runs 5× slower, the difference is not in the math: align the decompositions with one controlled experiment first (even if the outcome is "flat") to converge the suspect class to instruction-level structure — far cheaper than intuition-driven kernel tinkering.
  2. A negative result's only vehicle is the record: code can be swallowed by a hook, but interleaved A/B numbers written into a docs commit survive; commit the record first when a measurement session ends.
  3. Read ISA details against the correct backend branch: I·J/64 (AMD MFMA) and I·J/32 (NVIDIA) differ by 2×, and that decides whether "16 chains" — the whole next-step plan — stands; publicly suspending the contradiction and then verifying beats glossing forward.

← 14 · Index · 16 →

16 · r12 — 16-chain warp tile + ldmatrix (LANDED)

Result: after the q4_K MMQ wide kernel's warp-tile rearrangement, 7B @2K same-window interleaved measurement gives wide-16 KD=4 1020–1058 vs narrow 441–481 tok/s ≈ 2.3× (KD=8 973–995; vs the pre-rewrite wide ~719 = 1.44×) — the "accumulator-depth wall" r10 located is flattened in one stroke, the largest structural landing of P6 before r34. Commit: 774a116. Date: 2026-09-02.

1. Background — where things stood

The r9–r11 elimination ladder had narrowed the q4_K MMQ suspects to instruction-level structure. r9 measured the entire shape matrix: narrow cp.async KD=8 481 was the local optimum, all six tile/staging shapes parity-clean — the shape axis closed — and read out llama.cpp's instruction model, ~0.018 inst/MAC/thread vs our 0.133. r10 ported llama's inner-loop math decomposition verbatim: 462–468 vs narrow's 470, motionless — the math decomposition was eliminated and the residual locked to three items: ILP chain depth (llama 16 independent mma chains + 128 accumulator registers; ours 8 chains / ~64 registers), ldmatrix A staging (1 LDSM vs 8 per-lane LDS.32 per fragment), and tile 128×128. r11 verified against llama.cpp's mma.cuh that ne = I·J/32 (NVIDIA Turing+ branch; I·J/64 is the AMD MFMA branch), the "16 chains" reading stood, and r12's plan had its prerequisite proof.

At this point MMQ was stuck at 441–481 tok/s (opt-in MINFER_MMQ=1) against the same-window f16 default path's 2284 — a 5× gap. r6's target (≥24 TMAC/s for f16 parity, ~30 for llama parity) meant prefill could go from ~2340 to ~2670/3250. MMQ was the only part of the prefill wall still beyond 5×.

The direct host of the rewrite was r8's wide kernel mmq_raw_wide_nt_kernel: 128-token block, 8 warps × 32×32 per-warp tile, best cross-machine-state result ~719 tok/s (the source of the 1.44× comparison in r12's results). Its margin over narrow (481) came from halving B traffic — but r8 had already established "L2 absorbs the B re-reads", so the wide tile's gains stopped there; the shared bottleneck is the per-chunk inner-loop overhead (r8's words). r12 cut on that bottleneck instead of moving the tile again.

One process note worth recording: r10's code died to the post-checkout hook, and in r12's session the hook was absent (the nix shell had no rusty-hook binary); commits explicitly bypassed it plus a make-up fmt-equivalence check. The 16-chain plan could not afford to be lost twice.

r12 executed r10's redo recipe but touched only the first two items: double the chain depth + ldmatrix for the A fragments; the B-side staging format and the two-term rescale were not changed by a single character. The cut was deliberately narrow — which is exactly why the 2.3× number could later be attributed cleanly.

2. Principle — the GPU mechanism

The warp tile's geometry. The new block tile is 128 tokens × 128 od with 8 warps. Each warp claims one private 16-od-row slice × the entire 128-token tile:

  • A fragments (token axis): 128 tok ÷ 16 tok per m16n8k32 fragment = 8 A fragments;
  • B fragments (od axis): 16 od ÷ 8 rows per fragment = 2 B fragments;
  • per chunk (32-k): 8 × 2 = 16 mma.

The key is "who shares whom": A fragments are reused by all 16 chains in the warp (the token axis is the warp's common axis), while B fragments are warp-private (the od axis is the warp's private axis). Against the old shape — 2 A fragments × 4 B fragments = 8 chains per warp (clow[4][2][4]) — the new shape moves the entire token axis into a single warp, doubling the chain count.

The accumulator register ledger. 16 chains cost registers: clow[8][2][4] = 64 int C-fragment registers + sum[64] = 64 float accumulators, 128 accumulator-class registers live simultaneously — exactly the llama depth r10/r11 recorded. This budget is only payable at 1 block/SM: 256 threads/block against a 64K register/SM file ≈ a hard cap of ~256 per thread, and 128 accumulators plus operands and indices barely fit. That is why llama's config says "target occupancy 1" — not a flaw, a budget choice: when no second block exists to fill the mma latency holes, chain depth IS the issue window. r5–r6 had already measured that occupancy 1→2 does not rescue this kernel (depth inverted against occupancy); r12 maxed out the depth instead.

ldmatrix. ldmatrix.sync.aligned.m8n8.x4.shared.b16 is one instruction fetching four 8×8 b16 matrices (512 B) from shared memory, distributed to the 32 lanes' 4 registers in the standard mma A-operand layout — exactly one m16n8k32 A fragment's worth. The per-lane alternative has each lane compute 4 addresses and issue 4 narrow loads (r10's accounting: 1 LDSM vs 8 LDS.32 per fragment). What is saved is more than instruction count: the address ALU disappears too, and the load granularity goes from 4 B to 16 B/lane.

KDR=4 restage-skip. B's raw bytes are organized in 256-k super-blocks; at KDR=4 two adjacent k-tiles fall inside one super-block, so the expanded qb8 can survive across the k-tile barrier — restaging happens only when ((kt·KDR) & 7) == 0, halving B-side movement. Free bandwidth, but not the main line: the main line is chain depth.

Why KD=4 overtakes KD=8. Intuition says doubling kd depth should amortize per-MAC staging overhead, but the smem ledger says otherwise: at the 128×128 tile, KD=8's dynamic smem is markedly thicker, and r8's phantom lesson (wide KD=8 once hit the ~99 KB cap) means every depth point must pass the launcher's explicit cap check. KD=4 with restage-skip already halves B movement, and depth's small amortization loses to the tighter smem budget and shallower residency — measured KD=4 (1020–1058) > KD=8 (973–995). This is also restage-skip's value: it lets the shallow depth stop paying full price for B movement for the first time.

The issue-window arithmetic. 1 block/SM, 8 warps/SM (4 schedulers × 2 warps each): the schedulers spend most cycles waiting on mma results. The old kernel issues 8 mma per thread per chunk; the new one 16 — each scheduler pass can accumulate twice the pending mma. No second block's warp exists to fill the holes (r5–r6 measured occupancy 1→2 not helping), so chain depth is the only source of issue window. This is the concrete form of "depth inverted against occupancy" in this kernel: r5–r6 stopped at KD depth; r12 pushed the same principle onto the accumulator chains.

A coverage reconciliation. Verifying the tile is self-consistent from another angle: per chunk per warp MAC = 16 mma × 16×8×32 = 65,536, and the warp's private slice is 16 od × 128 tok × 32 k = 65,536 — exact. Block level: 8 warps × 16 od = 128 od ✓, every warp covers all 128 tokens ✓. This reconciliation is not decoration: 16-chain rearrangements most easily produce coverage holes where some output rows are computed twice and others missed (the r5–r6 4-warp 32×32 hand mapping leaked exactly this way — zero cells, parity blew up immediately). r12 passing parity on the first try says the (A fragment g, B fragment nh) → (token, od) mapping is a complete bijection.

The cost accounting for leaving B alone. r12 deliberately kept B fragments on "per-lane narrow reads + unpack nibbles at mma time": 2 B fragments per chunk × 2 32-bit reads each + unpack shift/mask. That overhead is not small, but it is a fixed quantity per chunk, not multiplied by the chain count — whereas the A side's chain structure multiplies onto every mma. Double the chains first, clean up B later; in the reverse order (both sides at once) the 2.3× could not be attributed. r14 validated this ordering two weeks later: switching B fragments to ldmatrix took another +18.5%, showing B did still have meat — but that was the next cut.

Why ~2× was the expectation. 1 block/SM, 8 warps/SM: old kernel 8 mma per thread per chunk with a shallow issue window; new kernel 16. Doubling chain depth lets each scheduler's mma issue port fit twice the work between dependency stalls. The measured 2.3× slightly exceeds expectation, indicating LDSM's load efficiency also contributed.

3. Implementation

3.1 Design choices (why this shape and not another)

  • The warp takes "16 od × all 128 tok" instead of a 2D sub-block: putting the whole token axis in the warp is the source of chain depth (8 A fragments); cutting od into 8 slices is the source of the warp count. The reverse (4-warp 32×32 sub-blocks) cannot raise the chain count — r12 tried it on the spot, negative.
  • The B side untouched: r10's lesson is one variable at a time. B's unpack still happens in registers at mma time (ldmatrix B fragments only at r14, pre-expansion tried at r18), so the 2.3× can only come from the A side's two changes.
  • KDR=4, not 8: KD=8's wide shape has larger smem, closer to the 99 KB cap (r8's phantom lesson), while restage-skip stops KD=4's B movement from being penalized — measured KD=4 (1020–1058) > KD=8 (973–995).
  • Scales loaded per group for the accumulators: da_q/sa_q changed to per-token-group reads to preserve accumulator registers (see excerpt B's comment) — 128 accumulators already fill the register file, so the rescale side's intermediates must yield.

3.2 Key code

Excerpt A · the inner loop before/after in the r12 diff (git show 774a116 -- src/cuda_kernels.cu, segment by segment) — the whole secret of 8-chains-to-16 is in these 20 lines:

// ---- BEFORE (r8/r9 shape): 2 A frags x 4 B frags = 8 chains, A frags per-lane narrow loads ----
int a[2][4], b[4][2];
int clow[4][2][4], chigh[4][2][4];
#pragma unroll
for (int h = 0; h < 2; h++) {                    // 2 A fragments
    const int r0 = i0w + h * 16 + (lane >> 2);
    const uint8_t* p0 = qat + (size_t)r0 * 32;
    a[h][0] = *(const int*)(p0 + 4 * (lane & 3)); // each lane computes its own address:
    a[h][1] = *(const int*)(p1 + 4 * (lane & 3)); // 4 LDS.32 per fragment
    a[h][2] = *(const int*)(p0 + 4 * ((lane & 3) + 4));
    a[h][3] = *(const int*)(p1 + 4 * ((lane & 3) + 4));
}
...
for (int nh = 0; nh < 4; nh++)                    // 4 B frags x 2 A frags
    for (int h = 0; h < 2; h++)
        mmq_mma_k32(clow[nh][h], a[h], b[nh]);    // 8 chains

// ---- AFTER (r12): 8 A frags x 2 B frags = 16 chains, A frags via ldmatrix ----
// A fragments: 8 independent 16-token groups tile the full
// 128-token row, one ldmatrix.x4 per group (16 rows x 32B:
// lanes 0-7 -> rows 0-7 byte 0, 8-15 -> rows 8-15 byte 0,
// 16-23 -> rows 0-7 byte 16, 24-31 -> rows 8-15 byte 16 —
// the standard m16n8k32 A-fragment distribution).
int a[8][4], b[2][2];
int clow[8][2][4];
#pragma unroll
for (int g = 0; g < 8; g++) {                     // 8 A fragments, 1 LDSM each
    const uint8_t* p = qat
        + (size_t)(g * 16 + (lane & 7) + ((lane >> 3) & 1) * 8) * 32
        + ((lane >> 4) & 1) * 16;
    unsigned r0_, r1_, r2_, r3_;
    asm volatile(
        "ldmatrix.sync.aligned.m8n8.x4.shared.b16 "
        "{%0,%1,%2,%3}, [%4];\n"
        : "=r"(r0_), "=r"(r1_), "=r"(r2_), "=r"(r3_)
        : "r"((unsigned)__cvta_generic_to_shared(p)));
    a[g][0] = (int)r0_; a[g][1] = (int)r1_;
    a[g][2] = (int)r2_; a[g][3] = (int)r3_;
}
// B fragments: 2 minitiles of the warp's private 16 od-rows;
// unpack the raw nibbles in registers (staging format unchanged).
...
// 16 independent mma chains per thread per chunk, all C
// fragments live simultaneously (llama.cpp accumulator depth).
#pragma unroll
for (int g = 0; g < 8; g++)                       // 8 A fragments
    #pragma unroll
    for (int nh = 0; nh < 2; nh++)                // x 2 B fragments
        mmq_mma_k32(clow[g][nh], a[g], b[nh]);    // 16 chains, C fragments live throughout

Three correspondences to note: b[4][2] → b[2][2] (B fragments 4→2, the warp owns 16 rows outright); clow[4][2][4] → clow[8][2][4] (accumulators organized by A fragment); and the comment's "llama.cpp accumulator depth" directly cites r10's finding.

Excerpt A′ · the block tile itself widening (the r12 diff's staging-layout header) — od 64 → 128 is the spatial precondition for doubling chain depth (only then can each warp hold the full token row × its private 16 rows):

// BEFORE (r8/r9 wide): od tile 64
-    //   sds   [KDR][64] f32, sdm likewise
-    float* sdm = sds + KDR * 64;
// AFTER (r12): od tile 128
+    //   sds   [KDR][128] f32, sdm likewise
+    float* sdm = sds + KDR * MMQ_WBJ;        // MMQ_WBJ: 64 → 128
...
+    float sum[64] = {0.0f};   // [g][nh][l]: 8 A-frags x 2 B-frags x 4 C regs

MMQ_WBI (tokens, 128) unchanged, MMQ_WBJ (od) doubled: the smem scale plane doubles with it, buying each warp a private od slice instead of a shared one. The sum[64] comment is the ledger itself — 8 × 2 × 4.

Excerpt B · the float accumulators and register yielding (r12 diff) — sum[64] is the other half of chain depth; the rescale side switches to per-group loads:

float sum[64] = {0.0f};   // [g][nh][l]: 8 A-frags x 2 B-frags x 4 C regs
...
// rescale: identical math/layout to the R1 kernel; A-side
// d/ssum come straight from the raw chunk. da/sa load per
// token-group to keep registers for the accumulators.
float dsv[2][8], dmv[2][8];   // BEFORE: float dsv[4][8], dmv[4][8]

The two-term rescale (dsv weight scale, dmv dmin term) keeps math identical to the r7–r8 kernel — the diff only shrinks the arrays' first dimension from 4 to 2 (B fragment count halved).

Excerpt B′ · the accumulator fold epilogue (r12 diff) — the line folding the int C fragments back into float sum; the index changes from (nh, h) to (g, nh), the multiplication structure untouched:

// BEFORE:
sum[idx] += da * dsv[nh][jj] * (float)clow[nh][h][l];
// AFTER:
sum[idx] += da * dsv[nh][jj] * (float)clow[g][nh][l];

dsv[nh][jj] is the warp-private 16 rows' weight scale (jj walks the od columns), da the A-side token scale — per chunk, 64 int C values fold into 64 float accumulators through the two scale terms. This chain is not one of the 16 mma chains (it reads already-finished C fragments), so doubling chain depth does not disturb it; its register footprint (dsv/dmv, 32 floats) does compete with the accumulators for budget — the origin of §3.1's "scales loaded per group".

Excerpt C · KDR=4 restage-skip (introduced by r12; still at lines 5983–5986 of the current tree) — the expanded B bytes survive across k-tiles:

/* At KDR=4 two consecutive k-tiles share the super-block and the
 * expanded qb8 persists across the kt barrier - restage only when
 * this k-tile starts a new super-block. */
if (((kt) * KDR & 7) == 0) {
    // ... B staging: global read of qb8 raw bytes + register unpack written to smem
}

(At r12 the B fragments still unpacked nibbles at mma time; the "register unpack written to smem" above evolved at r14 into pre-expansion + the slot-major 48B layout — see below.)

Excerpt D · this kernel's direct descendants (current tree, mmq_raw_wide_nt_kernel) — r12's skeleton lives as-is, with B fragments also on ldmatrix (r14) and addresses switched to precomputed XOR-swizzle offsets (r22):

// src/cuda_kernels.cu (current tree): r14's B-fragment ldmatrix + r22's G[g] swizzle
int a[8][4], b[2][2];
int clow[8][2][4];
#pragma unroll
for (int g = 0; g < 8; g++) {
    const uint8_t* p = qat + G[g];          // r22: XOR-swizzled granule offset
    unsigned r0_, r1_, r2_, r3_;
    asm volatile("ldmatrix.sync.aligned.m8n8.x4.shared.b16 "
                 "{%0,%1,%2,%3}, [%4];\n" ...);
}
// B fragments: ONE ldmatrix.x4 serves both 8-od-row
// minitiles (matrices 0/1 = od-rows 0-7 at k-halves 0/1,
// matrices 2/3 = od-rows 8-15). reg_i of lane L = matrix_i row
// L/4, bytes (L%4)*4 — the exact mma.m16n8k32 B-operand
// distribution the plain LDS pattern produced.
{
    const uint8_t* rb8 = qb8 + (size_t)sg * (MMQ_WBJ * MMQ_WBQ)
        + (size_t)(j0w + (lane >> 4) * 8 + (lane & 7)) * MMQ_WBQ
        + (size_t)((lane >> 3) & 1) * 16;
    asm volatile("ldmatrix.sync.aligned.m8n8.x4.shared.b16 " ...);  // 1 LDSM
}
#pragma unroll
for (int g = 0; g < 8; g++)
    #pragma unroll
    for (int nh = 0; nh < 2; nh++)
        mmq_mma_k32(clow[g][nh], a[g], b[nh]);   // 16 chains, unchanged to this day

r12 is not the endpoint but the skeleton: r14 (B-fragment ldmatrix, another +18.5%), r20 (split-phase A staging), and r22 (qa8 XOR swizzle) all grew on this 16-chain shape; the mmq_raw_nb_bt kernel promoted today (r28+) inherits the same "chain depth × register budget" ledger.

3.3 Pitfalls

  • The 4-warp 32×32 variant is negative: chain depth does not double automatically with warp count — 4 warps × 32×32 sub-blocks still leaves 8 chains per warp, and it fragments the token axis so A fragments lose reuse. Rejected by measurement.
  • 48B padding of A rows measured flat: the A-side row padding for ldmatrix bank phase measured flat (48B on the B side is r14's business; not worth it on A at the time) — reverted; changes are not kept for "looking tidier".
  • Registers maxed out: 128 accumulators + 32 A-fragment registers exhaust the budget, so da_q/sa_q must be re-read per group rather than resident throughout — the hidden tax paid for chain depth.
  • smem cap checked point by point: doubling the od tile doubles the scale plane; KD=8 totals 98,304 B and KD=4 73,728 B — both inside the ~99 KB opt-in cap, with the launcher explicitly checking attr/launch results (r7's phantom lesson institutionalized: over-cap must be refused loudly, never silently fall back).
  • Hook bypass: rusty-hook was missing from this session's nix shell and the post-checkout hook was bypassed (r10's code had just been swallowed by it); a make-up fmt-equivalence check + suite 169/0 were run at commit time. The process lesson is the mirror of r10's: a hook can swallow code (r10) or be absent itself (r12) — both states need a manual-verification backstop.

4. Verification

  • Parity gate (cuda_prefill_mmq): old and new kernels element-wise identical — defends against the 16-chain rearrangement changing summation order (each chain's k order is unchanged, so bitwise was the theory; measured green). The mechanism deserves spelling out: the 16 chains each accumulate their own (A fragment × B fragment) pairs, and the summation order at block level is isomorphic to the old kernel's — only issue timing changed, not the numeric path. So the parity expectation is identity, not tolerance; a 1e-3-scale drift would mean the rearrangement accidentally changed the reduction structure and must be investigated as a bug.
  • Suite 169/0: whole-model regression — defends against the wide branch working only for q4_K and breaking other types.
  • Greedy token identity: byte-compared token stream against the default path — defends against any end-to-end drift. This layer backstops the kernel parity: parity covers a single kernel, greedy covers the whole forward.
  • Same-window interleaved A/B: wide-16 vs narrow interleaved in one session window — defends against cross-session machine drift (the campaign's hard rule).
  • Default-path isolation: four env gates (MINFER_MMQ=1 MINFER_MMQ_RAW=1 MINFER_MMQ_RAW_WIDE=1 MINFER_MMQ_RAW_KD=4 to reach the new path); the f16 default path unaffected.

5. Results

ComparisonNumber (7B @2K, same-window interleaved)Ratio
wide-16 KD=41020–1058 tok/svs narrow 441–481 ≈ 2.3×
wide-16 KD=8973–995 tok/sslightly below KD=4
vs pre-rewrite wide~719 tok/s (cross machine states)1.44×
vs same-window f16 default2284 tok/sstill ~2.2× behind — below the promotion threshold

parity green, suite 169/0, greedy token identity. Best config: MINFER_MMQ=1 MINFER_MMQ_RAW=1 MINFER_MMQ_RAW_WIDE=1 MINFER_MMQ_RAW_KD=4.

Why not a promotion. The 2.3× is against our own old kernel; against the f16 default GEMM path (same window 2284) it is still 2.2× behind, and the f16 path was the default production behavior. Flipping a path that is still 2.2× slower to default would be pure regression — so r12's landing semantics are "the biggest step on the MMQ line", not "a step for the engine's default behavior". MMQ's default-on had to wait for r60 (after post-parity and the q6_K/FA/prepass lines converged). The commit message's "not promotion material yet" means exactly this, and it pre-records the next lever (load-time B-side fragment pre-formatting).

Campaign coordinates: this is the largest structural landing of P6 before r34 (quantize-transpose prepass, +9.72%); of r10's redo recipe (16 chains + ldmatrix + 128×128 tile), two items were executed here and the third (128×128 tile) was achieved along with this shape's 128×128 block tile. The commit also records the next step: load-time B-side fragment pre-formatting — the line that became r14's B-fragment ldmatrix / r18's pre-expansion experiments. The residual 2.2× to f16 was handed to r13's counter forensics (per-MAC warp instruction stream 10.14 vs 6.06 M).

The basis of the numbers. Every ratio in this doc names its comparison target, because the machine state drifted across r12's two days (master table footnote 2: within the r12–r25 window, adjacent sessions on dgxspark drifted −9% to +38%): 2.3× is wide-16 vs narrow inside one narrow same-window band (1020–1058 vs 441–481, mutually comparable); 1.44× is cross-machine-state against the pre-rewrite wide's ~719 (weak control, corroborating only); the KD=8 vs KD=4 ordering (973–995 vs 1020–1058) is likewise a narrow-band reading. Copy the basis along with the number, or 2.3× and 1.44× will be mistaken for two measurements of one quantity.

6. Lessons

  1. Write the hardware's required ILP explicitly — the compiler cannot invent chains the source never declared: 8 chains and 16 chains are both "correct" at the PTX level, but only the latter fills the mma issue ports; chain depth is a source-level contract, not a compiler optimization knob.
  2. Chain depth's price is registers, and registers' budget is occupancy: 128 accumulators are only payable at 1 block/SM — do the ledger first, then choose the shape; occupancy and depth are inverted in this kernel class (r5–r6 measured it long before).
  3. One-variable slicing makes ratios attributable: with the B side and rescale untouched by a single character, the 2.3× lands cleanly on "chain depth + A-fragment LDSM", and r14's follow-on stacking had a stable baseline.
  4. Ratios must carry their basis: 2.3× (in-band vs narrow) and 1.44× (cross-state vs old wide) are both correct; mixing them misleads — draw the interleaved window's boundary first, then the ratio means something.

← 15 · Index · 17 →

17 · x-tile / j-tile / cp.async-db — the staging-shape family, closed (REVERTED / closed)

Result: on top of r12's landed state (baseline band 1035–1058 tok/s), three staging-shape experiments were vetoed in a row: x-tile (256 tok × 64 od) ~942 (−9%, kernel +27%); j-tile (128 tok × 256 od, A reuse) ~1067 (+2.6%, bar 1150 not met); cp.async-db (double-buffered async staging) ~1034 (neutral, L2 SOL 75.7 → 78.2%). After three angles measured flat, the verdict: the staging-order axis is closed; the remaining levers are per-MAC L2 bytes and SM-side instruction efficiency — not traffic shape. Commit: f061cb8 / 4993804 / 784786d — all three docs-only negative-result records; the code snapshots lived in /tmp (j-tile noted as /tmp/cuda_kernels_jh256.cu) and no longer exist. Date: 2026-09-02 (all three same day).

⚠️ Code evidence note: none of this doc's three changes has a verifiable code commit. The excerpts below come in two kinds: those marked "current tree" are genuinely surviving code — the r12 kernel shape (774a116, the common host of all three experiments) and its descendants today; fragments marked "reconstructed" are rebuilt from the record's text, not surviving code.

1. Background — where things stood

r12 had landed the 16-chain warp tile that same morning: MMQ jumped from 441–481 to 1020–1058 tok/s (≈2.3×). But the same-window f16 default path was at 2284 — still 2.2× away, and r6's parity target (≥24 TMAC/s) unmet. r10/r11 had narrowed the suspects from "math" to "instruction-level structure", and r12 executed two of the three items (chain depth, A-fragment LDSM); the next natural hypothesis was the third structural class: traffic shape — the A-vs-B re-read ratio in L2, staging depth and order.

The hypothesis had concrete numbers behind it. At the 128×128 baseline tile, the A side's activation re-reads already outweighed B's weight re-reads: A streams 4,480 B per token over all K (40 B/chunk × 112 chunks, id=3584), re-moved once for every y-block (od-direction tile); the weight side touches 2,016 B per row per visit (144 B/super-block × 14 super-blocks). Cumulative over the whole prefill: A re-reads 327 MB vs B re-reads 152 MB — A is already 2× B. Intuition says that is unbalanced: why not tune the tile shape toward "read less A"?

And r12's fresh profile handed the hypothesis a knife: the 128×128 baseline's q-proj kernel already ran Memory SOL at 75.7% — near the ceiling; meanwhile the j-tile experiment's pre-run measured issue 0.22 and active warps per scheduler 2.00. Together these two readings are this doc's drama: the memory subsystem is nearly saturated (high SOL) but the SM's issue rate is very low (low issue) — byte bottleneck or latency bottleneck? There was no criterion yet. The three experiments are a differential diagnosis of exactly this question, one from each direction.

So three shapes were measured in one day, each moving one axis: x-tile (widen the token dimension to 256, indirectly squeezing od tile to 64), j-tile (merge od to 256 so A moves once and is reused by both), and cp.async-db (keep the tile, swap synchronous staging for a cp.async double-buffered async pipeline). The three represent the two directions of "make A cheaper" and one direction of "make the waiting disappear".

2. Principle — the GPU mechanism

Why x-tile's ledger is negative. The block tile goes 128 tok × 128 od → 256 tok × 64 od: doubling the token dimension halves the x-block count; halving od doubles the y-block count. Two traffic streams follow:

  • A re-reads ∝ y-block count: every block stages the whole A tile. y-blocks double → A traffic doubles: 327 MB → ~654 MB (+327 MB);
  • B re-reads ∝ x-block count: each block touches its own 64/128 weight rows. x-blocks halve → B traffic halves: 152 MB → ~76 MB (−72 MB).

Paying +327 MB of A to buy −72 MB of B, when A is already the 2× majority — strictly negative. The reverse, 128 tok × 256 od (widen both to keep the product), is geometrically feasible but 512 threads/block cannot stay resident at REG 156–166 — the register wall seals that road. The ncu evidence matches the ledger: q-proj kernel 4.584 vs 3.619 ms (+27%), Memory SOL 73.6% vs 75.7%, Compute 22.1% vs 22.9% — utilization barely moved; the same work took 27% more time.

j-tile: bytes saved, time not. 128 tok × 256 od + an outer jh loop: A stages once per k-tile and is reused by both 128-od halves — A re-reads measured halved (−163 MB L2). Geometrically, merging block od 128 → 256 halves the y-block count and A's move count with it; inside the block, an outer loop computes od half 0 first, then half 1, the two halves sharing the same staged A and accumulator structure (sum[2][64], two copies). The price is one extra round of B-stage + __syncthreads per k-tile. The kernel's state: 1 block/SM, 2.00 active warps per scheduler, issue 0.22 — a classic latency bottleneck: too few warps, so the saved L2 bytes have no queue pressure to relieve, while the added stage+barrier round lands directly on the critical path. Result ~1067 (+2.6%), far below the 1150 bar (ncu side: q-proj kernel 3.62 → 3.55 ms, only −2%; Memory SOL 75.7 → 74.0% — it went down, from waiting one more barrier round). One engineering trap besides: the shared qb8's restage-skip is unsound in this shape — the previous round's stage always holds the other half's rows; partial-slot expansion (KDR/2 pairs) pushes per-block B bytes back to 2× — blocked from both ends.

cp.async-db: prefetch works, but the bottleneck isn't prefetch. Two-level staging: raw bytes go out one k-tile early via cp.async, expansion happens smem→smem. L2 SOL rose 75.7% → 78.2% — async movement does work, waiting does shrink. But the kernel is L2-throughput-bound (bytes/s at the ceiling), not MLP-starved (insufficient concurrency): sending the same bytes to L2 earlier leaves the throughput ceiling unchanged, wall clock neutral (~1034, baseline band 1035–1050).

The mechanism behind the family verdict. Together the three experiments covered traffic shape's three degrees of freedom: od split ratio (x-tile, negative), A reuse (j-tile, bytes saved but not time), async depth (cp.async-db, prefetch works but no bottleneck to solve). All three roads lead to the same reading: at 1 block/SM this kernel is bound by L2 throughput and the latency structure — rearranging the same bytes produces no time.

How to read the counters: a two-metric cheat sheet. Memory SOL (Speed of Light) is ncu's percentage of measured traffic over the hardware's theoretical peak — 75.7% means the memory subsystem is doing useful work three quarters of the time; high does not mean "more traffic still helps", because the bottleneck may be rotating elsewhere. issue (issued warp instructions per cycle per scheduler) is the utilization of the SM's issue ports — 0.22 means the four schedulers issue nothing most cycles, warps waiting on something (memory, barriers, dependency chains). Reading "high SOL + low issue" together: the SM is waiting on memory, but memory is also nearly full — adding more bytes (x-tile's A, j-tile's second stage round) makes both sides worse together; making the same bytes arrive earlier (cp.async-db) improves waiting, not throughput. That is the differential logic the three experiments' data combine into.

3. Implementation

3.1 Design choices (why this shape and not another)

  • All three experiments share the r12 host: everything sits on mmq_raw_wide_nt_kernel (KD=4, baseline band 1035–1058), one geometry/scheduling parameter changed at a time — so the negative results corroborate each other instead of polluting each other.
  • x-tile chose 256×64 over 128×256: the latter's 512 threads cannot stay resident at REG 156–166 (the record rejects it explicitly); 256×64 keeps 256 threads and the per-warp structure (warp = 16 od rows × 128-token half) — the minimal testable variant.
  • j-tile's bar pre-calibrated at 1150 (≈ baseline +9%): A halving was a profiled fact; the bar answers "is the structural complication worth it" — +2.6% misses the bar, withdraw, no feelings involved.
  • cp.async-db leaves the tile alone: it isolates the single "async depth" variable, same mechanism as the narrow kernel's existing cp.async double buffer (excerpt C).

3.2 Key code

Excerpt A · the host's staging shape (current tree, r12's synchronous-staging comment kept verbatim) — the part all three experiments touch:

// src/cuda_kernels.cu (current tree, the RAW_STAGE macro header of mmq_raw_wide_nt_kernel)
#define RAW_STAGE(kt)                                                          \
    do {                                                                       \
        /* llama.cpp-style synchronous staging, single buffer: plain global    \
         * -> smem loads, one syncthreads orders them. */                      \
        {  /* from r20 this evolves into split-phase: fire a batch of          \
             * independent LDGs into registers first, then STS them —         \
             * a later story; at r12 it was LDG->STS interleaved one by one */\
        ...
    } while (0)

Excerpt B · the geometry the experiments changed (current tree launcher, real code) — the two divisors of grid are exactly what x-tile/j-tile modified:

// src/cuda_kernels.cu (current tree, launch_mmq_raw_wide_nt)
// 16-chain layout: 128-token x 128-od block tile. ...
dim3 grid((nt + 127) / 128, (od + 127) / 128);   // ← x-tile: (nt+255)/256, (od+63)/64
                                                 // ← j-tile: (od+255)/256 + in-block jh loop
if (kd <= 4) {
    const int smem = 4 * MMQ_WBI * 32 + 4 * MMQ_WBI * 8
                   + 8 * MMQ_WBJ * MMQ_WBQ + 2 * 4 * MMQ_WBJ * 4;
    cudaError_t e = cudaFuncSetAttribute(..., cudaFuncAttributeMaxDynamicSharedMemorySize, smem);
    if (e != cudaSuccess) { cudaGetLastError(); return 0; }   // over-smem: explicit refusal
    mmq_raw_wide_nt_kernel<4><<<grid, 256, smem, stream>>>(w, q8, c, nt, od, id);
}

With the token dimension pulled to 256, x-tile's KD=8 smem request goes over the cap — the launcher refuses explicitly per the table above and falls back to the narrow path (record verbatim: "KD=8 refused at the launcher -> narrow fallback"). This is the rule instituted after r7's phantom actually working: resource caps must fail loudly.

Excerpt C · the mechanism cp.async-db wanted to port (current tree narrow kernel, real code) — the narrow kernel has had a double-buffered cp.async pipeline since r7–r8; the db experiment carried it into the wide kernel:

// src/cuda_kernels.cu (current tree, narrow kernel main loop — the db experiment's mechanism blueprint)
__device__ __forceinline__ void gemm_cp_commit() { asm volatile("cp.async.commit_group;\n"); }
__device__ __forceinline__ void gemm_cp_wait1()  { asm volatile("cp.async.wait_group 1;\n"); }

RAW_STAGE(0, 0);
gemm_cp_commit();
int buf = 0;
for (int kt = 0; kt < nktile; ++kt, buf ^= 1) {
    if (kt + 1 < nktile) RAW_STAGE(kt + 1, buf ^ 1);  // prefetch the next k-tile
    gemm_cp_commit();
    gemm_cp_wait1();          // at most the prefetch group pending
    __syncthreads();          // scales (plain stores) + landed bytes visible
    ...
}

Reconstructed fragments (flagged: not surviving code, rebuilt from the record) — the three variants' shape differences:

// [reconstructed] x-tile: block constants become 256 tok x 64 od, warp structure unchanged
// #define MMQ_WBI 256
// #define MMQ_WBJ 64
//   warp = 16 od-rows x 128-token half; KD=8 refused at the launcher by the smem cap → narrow fallback
//   result: A traffic ∝ y-blocks doubled (+327 MB), B traffic ∝ x-blocks halved (−72 MB) → ~942 tok/s

// [reconstructed] j-tile: od 256, outer jh loop lets two 128-od halves reuse the same staged A
// for (int jh = 0; jh < 2; jh++) {
//     /* A staging once per k-tile (hoisted out of jh); sum[2][64], two accumulator copies */
//     /* B re-expanded per (kt, jh) — the second stage + barrier round lands on the critical path */
// }
//   result: A re-reads −163 MB, wall clock ~1067 (+2.6%) < bar 1150

// [reconstructed] cp.async-db: wide kernel switched to double buffer — cp.async raw bytes one k-tile early, expansion smem→smem
//   result: L2 SOL 75.7 → 78.2%, wall clock ~1034 = neutral (baseline band 1035–1050)

3.3 Pitfalls

  • x-tile's KD=8 over-cap did not become a phantom: the launcher refused explicitly and fell back to narrow — r7's silent-failure lesson (the 2124 phantom) institutionalized and cashed in for the first time, with the fallback path covered by the parity gate.
  • j-tile's restage-skip is unsound: when the shared qb8 is reused across jh, "the previous round's stage holds the other half's rows" voids the skip condition; partial-slot expansion pushes B bytes back to 2× — reuse and skip are mutually exclusive in this shape.
  • cp.async-db "reverted before landing": the neutral result was withdrawn on the spot, never even entering the tree behind an opt-in gate — the standard treatment of negative results: into the docs, leaving no live complexity behind.
  • All three lived in /tmp: experiment code written, measured, and deleted the same day; only docs remain in the repo. r10's hook incident and this /tmp convention are two faces of one process rule: the only vehicle for a negative result is the record.

4. Verification

  • Parity gate: all three variants parity green (x-tile both paths — the main path and the narrow fallback — pass) — defends against geometry rearrangement changing outputs.
  • Same-window interleaved A/B: x-tile 927/953/945 interleaved three times vs the baseline band 1035–1058; j-tile median 1067 vs 1035–1050; cp.async-db ~1034 — all interleaved in the same window, defending against machine drift.
  • ncu evidence: x-tile's duration/Memory-SOL/Compute triple proves "same utilization, longer"; j-tile's issue 0.22 / 2.00 warps-per-scheduler proves the latency bottleneck; cp.async-db's L2 SOL 78.2% proves prefetch worked.
  • Default-path isolation: experiments sat behind opt-in env gates; the f16 default and the narrow raw path unaffected (narrow ~476 unchanged in the same window).
  • The verification gap on the code side must be recorded honestly: the three experiments' code existed only in that day's /tmp working tree, with no re-verifiable commit — this doc's "verification" reproduces the measurement evidence chain (parity + A/B + ncu), not the code itself. That x-tile's narrow fallback path was covered by the parity gate matters especially: the fallback was not "untested", it was a second, tested path.

5. Results

ExperimentShapeWall clock (7B @2K, same-window interleaved)ΔKey evidenceDisposition
x-tile256 tok × 64 od~942 (927/953/945)−9%kernel +27%, SOL flat; A +327 MB vs B −72 MBREVERTED
j-tile128 tok × 256 od + jh loop~1067+2.6% (bar 1150 unmet)A re-reads −163 MB; issue 0.22, 2.00 warps/schedREVERTED
cp.async-dbdouble-buffered async staging~1034neutral (band 1035–1050)L2 SOL 75.7 → 78.2%REVERTED

Veto mechanisms (each worth archiving): x-tile lost on the ledger (A is already the 2× majority, and it added to A); j-tile lost on bottleneck mismatch (under a latency bottleneck, byte savings don't cash); cp.async-db lost on bottleneck conservation (the L2 throughput ceiling doesn't move, so sending the same bytes earlier buys nothing).

Retry conditions (STYLE rule: a REVERTED doc must state under what future conditions a retry is worthwhile). None of this family's three vetoes is a permanent judgment — they are judgments under the kernel's current state; change the state and the ledger changes:

  • x-tile: when A re-reads are no longer the 2× majority (e.g. after r34's quantize-transpose prepass turns A into a one-shot pre-transposed plane, or a shape where A resides in L2), the "+A buys −B" directionality inverts and the wide token dimension is worth re-measuring;
  • j-tile: when the kernel escapes the 1-block/SM latency bottleneck (occupancy raised enough to cover the stage+barrier rounds, like the NB kernel's 2 blocks/SM after r28), the −163 MB of A bytes get a chance to cash into time;
  • cp.async-db: when the kernel is genuinely MLP-starved (issue even lower, SOL also low, staging latency exposed on the critical path) — r53/r56 found exactly that state on the q6_K BT kernel, which is when the cp.async bundle finally landed (+5.03%/+2.35%).

In other words: what this family closes is "rearranging traffic on the 128×128 16-chain kernel", not traffic optimization itself. The same mechanisms revived once the bottleneck class changed — the most important hidden thread between this doc and the Era D docs that follow.

The family-closing verdict: after all three degrees of freedom (od split, A reuse, async depth) were measured, the staging-order axis formally closed. The next lever had to change "bytes per MAC" or "SM-side instruction efficiency" — and the immediately following r13 (counter forensics: 10.14 vs llama's 6.06 M warp instructions per GMAC) and r14 (B-fragment ldmatrix, +18.5%) walked exactly those two directions.

6. Lessons

  1. Compute the re-read ledger before tuning the tile: A side 4,480 B/token × y-block count, B side 2,016 B/row × x-block count — the 327 vs 152 MB imbalance sits right there, and "symmetrically widen the token dimension" necessarily adds to the majority.
  2. Byte savings only cash under a byte bottleneck: the 1-block/SM latency bottleneck (issue 0.22) ate j-tile's −163 MB; the L2 throughput ceiling ate cp.async-db's prefetch — confirm the bottleneck class before choosing the lever.
  3. Negative results need a complete evidence chain too: each of the three carries parity + interleaved A/B + an ncu triple; only then does a family verdict stand. Code in /tmp, records in docs — this campaign's fixed archival form for negative results.

← 16 · Index · 18 →

18 · r13 — counter forensics against llama.cpp (MEAS-ONLY, closed)

Result: the first working ncu session split "we are ~2.6× slower than llama" down to hardware-counter level — the per-GMAC warp instruction stream 10.14 M vs 6.06 M (1.67×) is the gap's carrier, while IMMA tensor work equals exactly 2×MAC on both sides (mma structure irrelevant); two patches purpose-built to fix "bytes/conflicts/store efficiency" (FULL −16%, MINIMAL +0.3% noise) disproved all three hypotheses at once by counter- evidence. Commit: 5ca037d (docs-only; the two kernel patches were restored after measurement as cmp-verified, code not committed). Date: 2026-09-03.

1. Background — where things stood

r12 had landed the 16-chain warp tile + ldmatrix A fragments, jumping the wide kernel from 441–481 to 1020–1058 tok/s (~2.3×), the largest step of P6's structural rewrite. But on the vs-llama ruler: the same q-proj-class GEMM runs 107.7 µs/GMAC on minfer wide KD=4 versus 41.1 µs/GMAC on llama.cpp — the remaining gap is still ~2.6×, and nobody could say where it "lives" in the hardware.

The candidate explanations were a long list, each intuitively plausible: is our L2 read traffic larger? Too many shared-memory bank conflicts? Low store sector efficiency on the C matrix? A staging barrier structure that is too synchronous? r9–r10 had already given a "paper" answer — llama's instruction model is ~0.018 inst/MAC/thread versus our 0.133 — but that was a ratio derived from source structure, not read off the hardware. Paper reasoning indicates direction but carries no evidentiary force: it cannot distinguish "more instructions but hidden by latency" from "more instructions directly lengthening execution".

An earlier step had measured and closed the whole staging-shape family: x-tile (256-token wide block) −9% reverted, j-tile (A reuse) +2.6% below bar reverted, cp.async double buffer neutral reverted. All three were "traffic shape" levers, all ineffective. The empty result was itself a signal — if bytes and access shape are not the binding resource, the only remaining candidate on the binding-resource list is the SM-side per-MAC instruction stream. But "only remaining" is not "confirmed".

That is r13's positioning: attribute the 2.6× gap with first-hand ncu counters, then disprove with two patches designed as mutually exclusive explanations. If "byte theory" is right, the FULL patch that cuts L1 requests by 55% should win; if it loses, byte theory is out. Without this step, every lever after r14 would be guessing. One engineering constraint came attached: this was the campaign's first working ncu session — getting ncu running on this GB10 machine and finding the right counter names was itself a pitfalls lesson (see §3.3).

2. Principle — the GPU mechanism

Why normalize per-GMAC. The two binaries differ in tile shape, kernel count, and launch shape; comparing absolute counts is meaningless. Dividing every counter by the GEMM's total MACs (normalized to "per 10⁹ multiply-accumulates") yields hardware resources consumed per unit of math work — a measure of the kernel's "instructional structure", decoupled from problem size. Both engines do the same amount of math (the IMMA counts prove this below), so the entire difference is structural.

duration's linear law is this case's instrument of judgment. The execution time of the SM-side warp instruction stream is approximately

duration ≈ warp_inst / (issue_rate × SM × scheduler × clock)

The converted slope measured in this session, on both engines and all data points, is the same order (~0.10–0.15 warp-inst/ns, the record's basis), i.e. duration and instruction count are nearly collinear on the scatter plot. llama's issue efficiency is 0.41 vs our 0.25 — so the 1.67× instruction difference multiplied by the issue difference explains the ~2.6× time difference exactly. Conversely: if instruction counts were close with a 2.6× time gap, the answer would live in the stalls; if the IMMA counts differed by 1.67×, the answer would live in mma structure. The counter matrix turns these three hypotheses into three decidable readings.

GB10's counter availability (ncu device id GB20B) sets the forensics checklist.

  • No dram__* metric family — the DRAM-level byte/SOL view is unreadable on this device; the deepest memory view is the L2 sector counter lts__t_sectors_aperture_device.
  • int8 tensor instructions do not enter the hmma (fp16 mma) counters — it reads 0 for an int8 kernel; the IMMA sub-pipeline counters must be read instead.
  • Available and used here: warp instruction execution counts, IMMA tensor operations, shared load/store bank-conflict counts, L2 read/write sectors, L1 request counts, kernel duration.

The FULL/MINIMAL counter-evidence design. The two patches map onto "byte theory"'s three sub-hypotheses: FULL = merged A staging (cuts L1 requests/L2 reads) + per-superblock register scale staging + qb8 row pitch 256→272 B (kills bank conflicts) + float2-ized C stores (raises store sector efficiency); MINIMAL keeps only the last two. If any of the three hypotheses were the binding resource, MINIMAL should show at least visible gains and FULL should be clearly positive. Both measured nothing — all three hypotheses out simultaneously.

3. Implementation

3.1 Design choices (why this shape and not another)

Why the q-proj-class GEMM as the alignment target. It is a regularly-shaped near-square GEMM, both engines run it through the same int8 mma path, and its share of the prefill wall is stable — a natural "same problem" control. The llama-side alignment leans on llama- bench's chunking behavior: -p 2600 is cut into a nt=512 ubatch, so llama's q-proj- class launch runs as grid (48,1,1) at 0.24–0.29 ms each — a real launch ncu can sample directly, comparable in size to ours, not a synthetic micro-benchmark.

Why two patches rather than one. One patch mixing four sub-changes means a win cannot be attributed and a loss cannot be blamed. FULL/MINIMAL's set difference (merged staging, register scales) alone carries "byte theory"'s strong hypothesis, and the shared part (272 B row pitch, float2 C stores) carries the "conflicts/store efficiency" hypothesis. The two outcomes together (FULL slower, MINIMAL flat) carry more information than any single mixed patch.

Why restore immediately after measuring. Neither patch reached the bar and both mechanisms were already disproven; keeping them would only pollute the next round's A/B "baseline" (a lesson re-validated bloodily at r59b — a stale baseline nurtured a fake +26.2% reading for two weeks). Restoration was confirmed byte-for-byte against HEAD with cmp, and the variant code was archived in /tmp for later reference.

3.2 Key code

r13 sampled the wide kernel's compute loop as it then stood. That state has no standalone commit (restored after measurement); the only citable real-code source is the deletion side of r14's commit c64cd99. First, the B-fragment read pattern the counters measured — 4 scalar LDS.32 per (warp, chunk):

// git show c64cd99 deletion side (the r13-era tree; r14 replaced this segment with ldmatrix.x4)
// B fragments: 2 minitiles of the warp's private 16 od-rows;
// the staged bytes are already per-k int8 in element order
// (slot sg of the row), so both words are plain smem loads.
#pragma unroll
for (int nh = 0; nh < 2; nh++) {
    const int jr = j0w + nh * 8 + (lane >> 2);          // this minitile's od row
    const uint8_t* rb8 = qb8 + (size_t)jr * 256 + sg * 32;
    b[nh][0] = *(const int*)(rb8 + 4 * (lane & 3));     // LDS.32
    b[nh][1] = *(const int*)(rb8 + 16 + 4 * (lane & 3));// LDS.32
}

The scale and A-side d/ssum reads are likewise narrow scalar streams — the per-chunk shared- read instruction count was r14's target (record basis: scales 8 LDS.32 per minitile → 2 LDS.128; sda_q 16 LDS.32 per chunk → 8 LDS.64):

// also c64cd99's deletion side: the od-column scales fully enumerated, 8 elements each way
// (after ptxas DCE, 8 LDS.32 actually execute per lane per nh)
float dsv[2][8], dmv[2][8];
#pragma unroll
for (int nh = 0; nh < 2; nh++)
    #pragma unroll
    for (int jj = 0; jj < 8; jj++) {
        dsv[nh][jj] = sdst[j0w + nh * 8 + jj];   // sds: one f32 per row
        dmv[nh][jj] = sdmt[j0w + nh * 8 + jj];   // sdm: separate plane
    }
...
// A-side d/ssum: two scalar reads per (chunk, g); the token pair (t, t+8) sits 16B apart
#pragma unroll
for (int t4 = 0; t4 < 2; t4++) {
    const unsigned pk = *(const unsigned*)(sda_q
        + (size_t)kd * MMQ_WBI * 2
          + (g * 16 + (lane >> 2) + t4 * 8) * 2);
    da_q[t4] = h2f((unsigned short)(pk & 0xFFFF));
    sa_q[t4] = (int)(short)(pk >> 16);
}

Honestly flagged: the FULL/MINIMAL patch implementations existed only in the r13 session's working tree, restored after measurement and never committed (5ca037d is docs- only), so per this directory's conventions no code excerpt can be provided here; their composition is the four-item list in §2 (merged A staging, per-superblock register scale staging, qb8 row pitch 256→272 B, float2 C stores), with numbers in §5.

3.3 Pitfalls

  1. ncu's launch form. The invocation that finally worked on dgxspark is sudo -n env LD_LIBRARY_PATH=... ncu ... — sudo strips the user environment, and the CUDA toolchain's shared-library paths must be carried through explicitly with env, or ncu's injection fails.
  2. The entire dram__* family is absent. GB20B has no DRAM counters; a session hunting "decisive DRAM-traffic evidence" comes back empty-handed. Forensics must land on lts__t_sectors_aperture_device (L2 sectors, split by aperture).
  3. int8 mma is not in the hmma counters. hmma reads 0 for both kernels — unfamiliarity would misread it as "the tensor cores are idle". int8 mma volume lives under the IMMA sub-pipeline counters.
  4. The 66 MB/GMAC L2 write-sector mystery. The minfer side showed ~66 MB/GMAC of L2 write sectors with no corresponding writer findable in the source; the float2 C-store patch (which explicitly changes the store shape) had zero effect on it. Suspended unexplained at the time, later classified as a counter artifact / denominator effect. Lesson: before concluding, verify an unattributed counter reading with a patch that "should change it if fixed".

4. Verification

  • Parity gate (once per patch): both FULL and MINIMAL passed bitwise parity before being allowed into A/B — this session is a counter-evidence design; had a patch computed wrongly, "slower" would no longer point at a mechanism conclusion (defends: misreading an implementation bug as a mechanism veto).
  • Restoration check: after measurement the kernel was cmp'd byte-for-byte against HEAD and confirmed identical (defends: a half-restored state polluting every later A/B's baseline).
  • suite 169/0: full regression green after restoration (defends: collateral damage from the restoration).
  • Counter cross-consistency: IMMA equals 2×MAC on both sides (proving the two control kernels really do the same math, legitimizing the normalization); FULL's inst +33% → duration +35% lands on §2's linear law (proving the instruction→time chain is trustworthy, not a sampling flutter).
  • llama-side launch authenticity: the sample target is llama-bench's real nt=512 ubatch launch (grid (48,1,1), 0.24–0.29 ms), not a self-made benchmark (defends: synthetic and real launches being structurally incomparable).

5. Results

The forensics table (q-proj class, per-GMAC normalized):

Counterminfer wide KD=4llama.cppRatio
warp instructions10.14 M6.06 M1.67×
IMMA tensor ops2.0–2.05 G (= 2×MAC)2.0–2.05 G (= 2×MAC)parity
LDS bank conflicts (read+write)940.3 K + 499.0 K6.5 + 0~100×+ (see below)
L2 read sectors22.0 MB13.7 MB1.6×
L2 write sectors66.1 MB1.45 MB45× (no source, later judged artifact)
duration / GMAC107.7 µs41.1 µs2.6×

The two patches' measurements (interleaved A/B in the same session, baseline 1023–1043 tok/s):

PatchContentCounter changeResult
FULLmerged A staging + register scale staging + qb8 272B pitch + float2 C storesL1 requests −55%, L2 reads −14%, but inst 10.14 → 13.48 M/GMAC (+33%)865–871 tok/s, −16%
MINIMALonly qb8 272B pitch + float2 C storesthe corresponding conflict/store metrics improved1041–1046 tok/s, +0.3% (noise)

Attribution conclusion: L2 bytes, store sector efficiency, bank conflicts — the three hypotheses disproven simultaneously (all three fixed at once, wall clock zero effect, FULL even slower from instruction bloat). The gap's carrier is the per-MAC warp instruction stream × issue efficiency: 1.67× instruction difference × issue 0.25 vs 0.41 ≈ the 2.6× time difference, closing against the linear law. r13 thereby pointed the campaign's next lever at "cut shared-read instruction count in the compute loop without adding any staging ALU or global traffic" — exactly the design input for r14 (B fragments via ldmatrix

  • widened scale reads, +18.5%/+23–30%).

Veto mechanism (this step is MEAS-ONLY): the step's "product" is the attribution conclusion, not code; the two patches were restored (cmp-verified) for missing the bar and being mechanically disproven, with the command-level evidence chain recorded in 5ca037d. The future retry condition for FULL-class "byte theory" levers: when the kernel leaves the stall-bound regime (r25's correction: the instruction stream is the first-order predictor, but only under issue-stall constraints), or when L2 bandwidth itself becomes the bottleneck.

6. Lessons

  1. The per-MAC instruction stream is this kernel class's (GB10, stall-bound regime) first- order predictor — bytes, sector efficiency, and bank conflicts can all be fixed simultaneously while the wall clock does not move.
  2. Attribution needs a counter-evidence patch: a patch that "removes hypothesis X" measuring zero effect beats ten pages of reasoning; the preconditions are a parity-clean patch and one that genuinely hits the hypothesis's target (FULL's L1 requests −55% proved that).
  3. Establish which counters the device exposes before designing the forensics: GB20B has no dram__*, and int8 mma sits in IMMA not hmma — get the checklist wrong and the session is wasted.
  4. Flag unattributed counter readings (the 66 MB/GMAC L2 write sectors) instead of concluding, and only trust them after a "fixing the write shape should change it" patch verifies.

← 17 · Index · 19 →

19 · r14 — B fragments via ldmatrix + widened scale reads (LANDED)

Result: wide kernel KD=4 1036 → 1225 tok/s (+18.5%), KD=8 1273 (+23–30%); the same kernel's ncu duration 3.632 → 2.378 ms (−34.5%), shared-load instructions −64%, LDS conflicts −47%, warp instructions −6.8%; the per-GMAC gap to llama narrows from 2.6× to 70.4 vs 41.1 µs/GMAC (1.7×). Commit: c64cd99. Date: 2026-09-03.

1. Background — where things stood

r13's counter forensics had just classified the case: the gap's carrier is the per-MAC warp instruction stream (10.14 M vs 6.06 M per GMAC, 1.67×), and the three hypotheses — bytes, bank conflicts, store efficiency — were killed simultaneously by the two counter- evidence patches. The conclusion landed on the compute loop's shared-read instruction count — so the next cut belongs on shared reads, and it must satisfy the two constraints r13 drew: add no staging ALU whatsoever (r18 would later prove staging-ALU-class substitutions are wall-clock ineffective at SM% ~30), and add no new global traffic (x-tile's −9% already demonstrated the price of moving bytes around).

Inventorying the wide kernel's (r12's landed 16-chain tile, by now the MMQ path's performance carrier) compute-loop shared-read streams: A fragments moved to ldmatrix.m8n8.x4 at r12 (8 per chunk), but B fragments were still 4 scalar LDS.32; the od-column scales were two streams of 8 LDS.32 each (the post-DCE per-lane execution count; record basis: 8 LDS.32 per minitile); A-side d/ssum was 2 LDS.32 per (chunk, g) (16 per chunk). llama.cpp's B fragments are fed by ldmatrix — precisely the not-yet-ported half of what r9 flagged at the time as "their pre-arranged mma-fragment B layout, producible at weight-load time".

r14's target list was therefore explicit: change B fragments from "move bytes, then read them one scalar LDS at a time" to ldmatrix fragments; widen the scale and d/ssum reads. The constraints were equally explicit: all three changes live on the read side of the compute loop, and the staging side changes only address arithmetic (which slot an element is written to), adding no expansion multiply-adds — i.e. "fewer/wider smem ops with zero staging-ALU growth". The narrow kernel is not the performance path and was left alone (r16 later confirmed "narrow is not the perf path" formally).

2. Principle — the GPU mechanism

What ldmatrix is. ldmatrix.sync.aligned.m8n8.x4.shared.b16 is a warp-level shared→register matrix move: one instruction loads four 8×8 b16 matrices, with the 32 lanes each supplying one row address (lanes 0–7 → matrix 0's 8 rows, 8–15 → matrix 1, and so on); each lane receives 4 32-bit registers whose contents land in the mma operand's fragment distribution. Against scalar LDS: one LDS.32 moves 4 bytes with addresses computed per lane; ldmatrix.x4 moves 512 bytes in one instruction and the addresses occupy only the 32 lanes' addressing slots. The difference is not bytes (the total is identical) — it is MIO queue entries and issue slots, which r13 had just proven to be the binding resource.

Why one ldmatrix.x4 can replace 4 LDS.32. mma.m16n8k32's B operand is 16 rows (two 8-row minitiles) × 32 k of int8, each lane holding 2 32-bit words (8 bytes). The old code used 2 LDS.32 per minitile, 4 for both; the data those 4 move is exactly four 8×8 b16 matrices — matrices 0/1 = od rows 0–7's k first/second halves, matrices 2/3 = od rows 8–15's. As long as the bytes in smem sit in the row-address distribution ldmatrix expects, one x4 performs the identical data movement with a register distribution exactly matching what the mma needs (r14's session verified this equivalence standalone — see §4).

The 48 B slot stride's bank arithmetic. smem bank conflicts are decided by the 4-byte- granularity bank phase. With row addresses laid out by slot, row r's starting bank = (stride/4 × r) mod 32. The old layout's 32 B stride gives 8r mod 32 = 0: all 8 rows slam the same phase (an 8-way conflict — the same disease seen on the A side at r12); the new 48 B stride gives 12r mod 32, and r = 0..7 yields

{0, 12, 24, 4, 16, 28, 8, 20}  — 8 distinct phases, zero conflicts

48 = 3×16 preserves 16 B alignment (a ldmatrix row-address requirement), at the price of each slot growing 32 B → 48 B (32 B of content + 16 B pad), the qb8 plane growing 32,768 → 49,152 B (+16 KB, KDR-independent — qb8 holds one super-block). KD=8's block total goes 81,920 → 98,304 B, still 1 block/SM and nearly filling the ~99 KB opt-in cap — which explains why this kernel never had plane-widening room again.

The same arithmetic for widening the scale / d:ssum reads. The C fragment's epilogue consumes only the od column pair (j, j+1) and token pair (t, t+8) per lane:

  • sds packs each row's two scales (d | dmin·m) into a float2, so adjacent rows join into one float4 and one LDS.128 serves a minitile — lanes 0–3's four float4 addresses sit 16 B apart covering a contiguous 64 B, conflict-free.
  • sda_q is retilted to [KDR][16 g][8 q] as uint2: the two packed u32 of the token pair (t, t+8) a C fragment needs sit adjacent, read by one LDS.64 (the old code: two LDS.32 with addresses 16 B apart).

Together, per (warp, chunk) the shared-read instruction count drops from ~30 (the scalar streams of 4 LDS.32 + 16 LDS.32 + 16 LDS.32, plus 8 A-side ldmatrix) to 9 ldmatrix + 10 wide LDS — the origin of the next ncu's −6.8% warp-inst and −64% shared-load against r13's instruction-stream reading. The key asymmetry: r17/r25 later proved that pure "cut support instructions" is wall-clock ineffective at 1 block/SM, yet r14 immediately gained +18.5% — because r14 cut shared-read entries in the MIO queue, hitting precisely the LDG→STS→LDS latency chain r13's forensics identified, not the SM-side int ALU.

3. Implementation

3.1 Design choices (why this shape and not another)

slot-major (sg-major), not row-major. One B-fragment ldmatrix addresses a single chunk (sg)'s 16 rows × 2 16 B halves — laying slots out as [8 sg][128 od-row][48B] puts those 16×48 B inside one sg's contiguous 6,144 B region, leaving every per-lane address term but sg loop-invariant (only + sg * (WBJ*WBQ) moves per chunk). The old row-major [128 row][8 sg][32B] scattered the 16 rows across the same row's 8 sg segments — irrelevant for scalar reads, but ldmatrix's 32 row addresses would have to cross segments. Why not pick some other width than 48 B: 44/48/52 all give distinct phases, but only multiples of 16 keep ldmatrix row alignment — 48 is the distinct-phase solution nearest 32 among 16 B-aligned widths.

sds packed as (d | dmin·m) rather than two planes. Two planes mean any widened read has to stitch across planes — two LDS on different planes can never merge into one. Packed as float2, the width follows naturally from the epilogue's consumption granularity (column pairs): one float4 = row j's and row j+1's float2 each.

sda_q as uint2, not uint4. The first cut used uint4 tiling and wrote straight past the sda_q plane into the neighboring qb8 (see §3.3) — a consumer's C fragment needs exactly one pair of tokens, so uint2 is the consumption granularity's exact width; uint4 spends smem budget feeding a nonexistent fourth consumer.

3.2 Key code

The excerpts below all come from the current tree's mmq_raw_wide_nt_kernel in src/cuda_kernels.cu (r14's changes survive verbatim after the adjacent r20/r22/r15 changes; line numbers measured against the current tree).

(a) smem layout and MMQ_WBQ (current tree 5861–5888) — the 48 B slot's definition and the four planes' arrangement:

#define MMQ_WBI 128
#define MMQ_WBJ 128
#define MMQ_WBQ 48  // padded per-(sg,row) qb8 slot: 16B-aligned, 12r mod 32
...
    //   qb8   [8][128][48]      B sub-blocks pre-expanded to per-k int8,
    //                           SLOT-MAJOR (sg-major); 48B row stride puts
    //                           every ldmatrix row on a distinct bank phase
    //                           (raw nibbles 0..15; 128 od-rows per tile)
    //   sds   [KDR][128] float2 (d | dmin*m): one float4 load serves the
    //                           (j, j+1) od-col pair per minitile
    uint8_t* qa8 = mmq_raw_sh;
    uint32_t* sda_q = reinterpret_cast<uint32_t*>(qa8 + KDR * MMQ_WBI * 32);
    uint8_t* qb8 = reinterpret_cast<uint8_t*>(sda_q + KDR * MMQ_WBI * 2);
    float2* sds = reinterpret_cast<float2*>(qb8 + 8 * MMQ_WBJ * MMQ_WBQ);

The old version (c64cd99's deletion side) was qb8 [128][8][32] + separate sds/sdm f32 planes + a KD=8 total of 81,920 B — the new layout pushes the block to 98,304 B (+16 KB, still 1 block/SM).

(b) the staging side's slot-major expansion (current tree 5986–6015) — the write side changes only address arithmetic, no ALU growth; the uint4 group's low/high nibbles each write two 16 B granules, the low half (pair p's sub-block 2p) and high half (2p+1) in two adjacent sg rows:

        if (((kt) * KDR & 7) == 0) {                     // at KDR=4, reused across kt
        for (int x = threadIdx.x; x < MMQ_WBJ * 4; x += blockDim.x) {
            const int r = x >> 2, p = x & 3;
            const int j = j0 + r, sb = ((kt) * KDR) >> 3;
            uint4 v0 = make_uint4(0,0,0,0), v1 = make_uint4(0,0,0,0);
            if (j < od && sb < nsb) {
                const uint8_t* src = W + (size_t)j * ((size_t)nsb * 144)
                                  + (size_t)sb * 144 + 16 + p * 32;
                v0 = *(const uint4*)(src);               // 32B qs read as-is
                v1 = *(const uint4*)(src + 16);
            }
            const unsigned M = 0x0F0F0F0Fu;
            uint8_t* dst = qb8 + (size_t)(p * 2) * (MMQ_WBJ * MMQ_WBQ)
                         + (size_t)r * MMQ_WBQ;          // slot-major 48B row pitch
            *(uint4*)(dst)      = make_uint4(v0.x & M, v0.y & M,
                                             v0.z & M, v0.w & M);
            *(uint4*)(dst + 16) = make_uint4(v1.x & M, v1.y & M,
                                             v1.z & M, v1.w & M);
            uint8_t* dst1 = dst + MMQ_WBJ * MMQ_WBQ;     // high-nibble neighbor row
            *(uint4*)(dst1)     = make_uint4((v0.x >> 4) & M, ...);
            *(uint4*)(dst1 + 16) = make_uint4((v1.x >> 4) & M, ...);
        }
        }

(c) the compute side: B fragments before/after. Before (c64cd99's deletion side), 4 scalar LDS.32 per (warp, chunk):

// B fragments: 2 minitiles of the warp's private 16 od-rows;
// the staged bytes are already per-k int8 in element order
// (slot sg of the row), so both words are plain smem loads.
#pragma unroll
for (int nh = 0; nh < 2; nh++) {
    const int jr = j0w + nh * 8 + (lane >> 2);
    const uint8_t* rb8 = qb8 + (size_t)jr * 256 + sg * 32;
    b[nh][0] = *(const int*)(rb8 + 4 * (lane & 3));      // LDS.32 ×2/minitile
    b[nh][1] = *(const int*)(rb8 + 16 + 4 * (lane & 3));
}

After (current tree 6085–6104), one ldmatrix.x4, the per-lane address loop-invariant except for sg:

            // B fragments: ONE ldmatrix.x4 serves both 8-od-row
            // minitiles (matrices 0/1 = od-rows 0-7 at k-halves 0/1,
            // matrices 2/3 = od-rows 8-15). reg_i of lane L = matrix_i row
            // L/4, bytes (L%4)*4 — the exact mma.m16n8k32 B-operand
            // distribution the plain LDS pattern produced. Per-lane address
            // parts are loop-invariant; only the sg term moves per chunk.
            {
                const uint8_t* rb8 = qb8
                    + (size_t)sg * (MMQ_WBJ * MMQ_WBQ)
                    + (size_t)(j0w + (lane >> 4) * 8 + (lane & 7)) * MMQ_WBQ
                    + (size_t)((lane >> 3) & 1) * 16;
                unsigned b0_, b1_, b2_, b3_;
                asm volatile(
                    "ldmatrix.sync.aligned.m8n8.x4.shared.b16 "
                    "{%0,%1,%2,%3}, [%4];\n"
                    : "=r"(b0_), "=r"(b1_), "=r"(b2_), "=r"(b3_)
                    : "r"((unsigned)__cvta_generic_to_shared(rb8)));
                b[0][0] = (int)b0_; b[0][1] = (int)b1_;
                b[1][0] = (int)b2_; b[1][1] = (int)b3_;
            }

The 32 lanes' row-address distribution: lanes 0–7 → od rows j0w+0..7's first 16 B (k first half), lanes 8–15 → the same 8 rows' second 16 B (k second half), lanes 16–23/24–31 → od rows 8–15's first/second halves. The four 8×8 matrices land in b0_..b3_, matching the old 4-LDS register layout one for one — the precondition for a bitwise "relayout-only" change.

(d) the epilogue read side: sds float4 and sda_q uint2 (current tree 6124–6142):

            // od-col scales: one float4 per minitile serves the (j, j+1)
            // column pair the C fragment consumes (float2-packed at staging).
            float dsv[2][2], dmv[2][2];
            #pragma unroll
            for (int nh = 0; nh < 2; nh++) {
                const float4 sc4 = *(const float4*)(sds
                    + (size_t)kd * MMQ_WBJ + j0w + nh * 8 + (lane & 3) * 2);
                dsv[nh][0] = sc4.x; dsv[nh][1] = sc4.z;   // rows j and j+1's d
                dmv[nh][0] = sc4.y; dmv[nh][1] = sc4.w;   // rows j and j+1's dmin*m
            }
            #pragma unroll
            for (int g = 0; g < 8; g++) {
                float da_q[2];
                int sa_q[2];
                // token pair (t, t+8) in one LDS.64 (uint2 tiling)
                const uint2 pk2 = *(const uint2*)(sda_q
                    + (size_t)kd * MMQ_WBI * 2 + g * 16 + (lane >> 2) * 2);
                da_q[0] = h2f((unsigned short)(pk2.x & 0xFFFF));
                sa_q[0] = (int)(short)(pk2.x >> 16);
                da_q[1] = h2f((unsigned short)(pk2.y & 0xFFFF));
                sa_q[1] = (int)(short)(pk2.y >> 16);

The before side was dsv[2][8]/dmv[2][8] full enumeration + the t4 twin scalar reads (the deletion-side shape in §3.2(c); record basis: scales 8 LDS.32 per minitile → 2 LDS.128, sda_q 16 LDS.32 per chunk → 8 LDS.64). The launcher changed only the smem-size expression (8 * MMQ_WBJ * MMQ_WBQ replacing MMQ_WBJ * 256, current tree 7195–7221); both KD settings remain inside the ~99 KB cap.

3.3 Pitfalls

  1. The first cut's uint4 tiling overflowed and polluted qb8. Retilting the sda_q plane as uint4 wrote past the plane's KDR·1024 B boundary and smashed the head of the neighboring qb8 — the symptom was not off-by-one results but whole regions of data invalidated. Localization came by bisect: two earlier bugs were fixed first (a dropped j0w term, a word-offset misalignment), parity stayed red, and only then did the plane overflow surface. Switching to the consumption-exact uint2 passed on the first try.
  2. The verification order for relayout-class changes. ldmatrix's register distribution must equal the old scalar reads' one-for-one for "change the layout, not the semantics" to hold — r14's discipline was standalone verification first (confirming the ldmatrix distribution == the mma B-operand distribution), then wiring into the kernel; once in, parity went green on the first build.
  3. The smem budget is a hard ceiling. After +16 KB, KD=8 sits at 98,304 B of ~99 KB — this kernel never had room to add a plane again (r17's A-side padding attempt and everything after had to subtract within this budget).

4. Verification

  • Standalone distribution verification: ldmatrix.x4's lane→register distribution was verified offline first to equal the mma.m16n8k32 B-operand distribution (defends: layout right but registers misassigned — the result "looks like a fixed permutation off", easily misdiagnosed as a scale error).
  • Parity gate (both KD=4 and KD=8 pass): the change touches only byte placement and read patterns, not math, so bitwise parity is required (defends: semantic drift introduced by the relayout).
  • greedy-32 identity: the 32-token greedy sequence byte-identical (defends: summation- order or sampling-chain drift invisible to parity's short samples).
  • suite 166/0/3: full regression (defends: the launcher's smem-size change breaking other paths).
  • narrow kernel control: narrow untouched, its readings noise — proving the gains really come from this wide-kernel change rather than machine-state drift (defends: polluted A/B attribution).
  • The ncu evidence chain: warp-inst −6.8%, shared-load inst −64%, LDS conflicts −47%, duration −34.5% — four readings interlocking; the instruction stream was cut, the conflicts cleared, and the time cashed per r13's linear law (defends: wall-clock gains from mechanisms unrelated to this one).

5. Results

MetricbeforeafterΔ
wide KD=4 whole-machine wall1036 tok/s1225 tok/s+18.5%
wide KD=8 whole-machine wall~1000-class1273 tok/s+23–30%
Kernel duration (ncu, same session)3.632 ms2.378 ms−34.5%
shared-load instructions——−64%
LDS bank conflicts——−47%
warp instructions——−6.8%
per-GMAC duration107.7 µs (r13)70.4 µsvs llama 41.1 (1.7×)

narrow control noise; parity green at both depths; suite 166/0/3; greedy-32 identity. Against r13's prediction: the instruction stream (especially shared-read entries) was cut by more than half and the wall narrowed to a 1.7× ratio per the linear law — r13's "instruction-stream-bound" conclusion cashed in its predictive power on its first formal application. r15 (rank-1 rescale, +1.9%) and r20 (split-phase A staging, +7.1%/+3.5%) then kept stacking on this kernel; the A-side d/ssum epilogue was subsequently rewritten by r15 (the rank-1 fold after current-tree 6143 is outside this doc's scope).

6. Lessons

  1. In the stall-bound regime, "fewer/wider smem ops + zero staging-ALU growth" is the lever class that cashes into wall clock — r13 classified it, r14 cashed it, and r17/r25 proved that outside this regime (SM-side int-ALU class) the same effort buys ~0%.
  2. A relayout-only change's correctness is designed, not tested: prove the new layout's register distribution standalone-equal to the mma operand distribution first; only then does parity qualify as a "green on first build" gate rather than a debugging tool.
  3. A widened read's width is decided by consumption granularity: the C fragment consumes column/row pairs, so float4 and uint2 are exactly enough; wider than consumption (the first cut's uint4) only overflows the plane, smashes the neighbor, and buys nothing.
  4. Compute the bank phases before touching a layout: the 48 B stride's 12r mod 32 eight distinct phases are the sufficient condition for zero conflicts — one line of arithmetic like this can veto a doomed 8-way-conflict scheme before any code is written.

← 18 · Index · 20 →

20 · r15 — f32-accumulate mma probe (dead end) + rank-1 term2 rescale (LANDED)

Result: wide-kernel MMQ KD=8 1271 → 1295 tok/s (+1.9%, 3/3 consistent); ncu per-GMAC warp instructions 9.45 → 8.69 M (−8.1%), q-proj kernel 2.378 → 2.262 ms (−4.9%). At the same time a seemingly tempting ISA route was closed: integer mma has no f32-accumulator form (proven by a ptxas probe). Commit: b999e9a. Date: 2026-09-03.

1. Background — where things stood

On 2026-09-03 the q4_K MMQ campaign (Era C) reached round 15. The campaign target had not changed since r6: lift the MMQ GEMM from 6.1 TMAC/s to ≥24 (f16-path equivalent) or ~30 (llama.cpp equivalent), turning the whole quantized prefill path around. The previous rounds' chain of results:

  • r12 (16-chain warp tile + ldmatrix A fragments): punctured the ILP-depth wall, wide kernel KD=4 reaching 1020–1058 tok/s, ~2.3× faster than narrow (441–481) — the largest single-step structural gain to that point.
  • x-tile / j-tile / cp.async-db (the staging-shape family): all flat or reverted; the verdict was that the staging-order axis is closed and "the only remaining levers are per-MAC instruction efficiency".
  • r13 (ncu counter forensics): the first hard evidence — per-GMAC warp instruction count is this kernel class's first-order predictor. minfer 10.14 M/GMAC vs llama.cpp's 6.06 M (1.67×), with duration tracking instruction count linearly at ~0.10–0.15 warp-inst/ns across all data points of both engines. It also eliminated three "suspects": L2 bytes, store efficiency, bank conflicts — all three fixed simultaneously, the wall unmoved.
  • r14 (B fragments on a single ldmatrix.x4 + widened scale reads): down to 9.45 M/GMAC, wide kernel KD=4 1225 / KD=8 1273 tok/s, kernel 3.632 → 2.378 ms. A big win for the "fewer/wider smem ops" lever class.

The epilogue itself has a genealogy worth telling: the two-term rescale's math passed unchanged from the R1 word-level kernel (mmq_nt_kernel, 2026-08-31) all the way into the raw kernel — the comment above the wide kernel's epilogue at r15 time reads verbatim "rescale: identical math/layout to the R1 kernel". That is, this epilogue is the oldest living code in the campaign: in the R1 era it was the price of correctness; r12/r14's structural rewrites bypassed it twice (touching the tile, touching fragment loads) but nobody touched its instruction composition. r13's ledger pushed it onto the stage: llama's 6.06 M/GMAC contains no such mass of per-value floating-point operations — llama uses float accumulators (floating-point mma) and never needs an I2F boundary at all, while we must pay the translation tax for an s32 accumulator.

At r15's opening the instruction-stream ledger read: llama 6.06, us 9.45, a 3.39 M/GMAC difference. r14's record named the next lever: mma...f32.s8.s8.f32 — an f32-accumulating integer mma. The idea's appeal is the entire reason the rescale epilogue exists: the integer mma's accumulator is s32, so every C value must pass through I2F (int→float conversion) before multiplying two scale terms. If the accumulator were f32 to begin with, that int→float boundary disappears entirely and the epilogue's instruction stream could collapse by a large block.

r15's first move was therefore an ISA probe: does this instruction even exist? The answer is no (see §2), and one ptxas call falsified it — part of this step's methodological value: before designing around an instruction, spend a few minutes verifying the instruction exists.

With the probe dead, the same math spawned two alternatives: merging the integer dot products of 2 chunks before one I2F (vetoed algebraically, see §2.3), and the ultimately landed rank-1 fold — it does not reduce I2F, but cuts the floating-point multiplication stream in term2 by 3/4.

2. Principle — the GPU mechanism

2.1 Where the two-term rescale comes from

MMQ's (int8 tensor-core quantized GEMM) numeric structure: A (activations) quantized to q8, each 32-element block carrying scale da and the in-block integer sum sa (the q8_1 format's d/s fields, computed for free by the quantize prepass); B (weights, q4_K) with each 32-sub-block having effective scale dsv = d·sc and negative bias term dmv = −dmin·m (expanded into float2 at staging).

The s8 mma (mma.m16n8k32, B side fed unsigned nibble values v ∈ [0,15] directly — exactly representable in s8) computes the integer dot product:

S_ij = Σ_k a_int[i,k] · v[j,k]

while the true product is b = d·sc·v − dmin·m, so every C value needs two correction terms (the "two-term rescale"):

true(i,j) = Σ_chunks [ da_i · dsv_j · S_ij   ← term1: the main term
                     + dmv_j · (da_i · sa_i) ] ← term2: the correction term

term1 multiplies the mma's integer result (I2F mandatory); term2 multiplies the A block sums. The key observation is term2's structure:

term2(i,j) = dmv_j · (da_i · sa_i) = (row vector da·sa) ⊗ (column vector dmv)

a rank-1 matrix (outer product) on the (token, od-col) plane. The row-side factor da_i·sa_i is identical for all columns of a row; the col-side factor dmv_j is identical for all rows of a column.

By contrast term1 is not rank-1: S_ij in da_i · dsv_j · S_ij is the mma's per-(i,j) output — a full-rank matrix. term2 is foldable precisely because it contains no mma result, being only the product of two per-row/per-column coefficient sets. "Ask the correction term's rank first, then decide how to evaluate it" is this step's reusable criterion.

Where the two coefficient sets originate in the data path (both computed once at quantization, expanded at staging; the GEMM hot loop only multiplies):

  • col-side dsv/dmv: expanded per sub-block into float2 at B staging (around src/cuda_kernels.cu:6024 in the current tree):
float d = h2f(*(const uint16_t*)blk);        // super-block scale d (f16)
float dmin = h2f(*(const uint16_t*)(blk + 2)); // super-block dmin (f16)
// …sc/m are the sub-block's 6-bit indices (get_scale_min_k4 table semantics):
dv = d * (float)sc;                          // effective scale  → dsv
mv = -(dmin * (float)m);                     // negative bias    → dmv
sds[(size_t)kd * MMQ_WBJ + r] = make_float2(dv, mv);
  • row-side da/sa: the A quantize prepass computes the in-block integer sum for free while writing each q8 block (around src/cuda_kernels.cu:5301: da[r] = d; sa[r] = s;) — the q8_1 format's s field. The GEMM kernel just reads it from staging, with no extra pass.

2.2 Where the old code wasted

Before the fold, term2 was evaluated in full for every C value:

sum[idx] += da * dmv[nh][l & 1] * sa;   // old: 2 FMUL per C value

The wide kernel has 64 C values per thread per chunk (8 A-frags × 2 B-frags × 4 C regs), but the row side's da·sa takes only 2 distinct values (the token pair each A-frag covers). That is, 48 of the 64 FMULs repeat the same product — the commit message's count "64 FMUL(da·dmv)/chunk → 16 FMUL/chunk" is exactly this ledger: fold the row-side product to once per row (8 g × 2 rows = 16), leaving each C value a single dma·dmv, which the compiler issues as an FFMA (fused multiply-add into the sum accumulator).

Per-chunk, per-thread op ledger for the term2 part (term1 and the I2F are unchanged, hence excluded from the delta):

Old (r14 shape)New (r15 fold)Δ
FMUL (row side da·sa or dma)64 (one da·dmv per C value, then × sa)16 (one dma = da·sa per row)−48
FMUL/FMA (col side × dmv into sum)64 FMUL + 64 FADD64 FFMAfused, no standalone FADD
I2F (sa to float)repeated per value (CSE-able)exactly 16 (inside dma)structured

The three components of sum[idx] += dma[l >> 1] * dmv[nh][l & 1] — reading dma, reading dmv, FFMA into sum — were all going to happen in the epilogue anyway; the fold saves only those 48 redundant FMULs.

2.3 The two rejected alternatives (falsified on paper, before implementation)

Route A: f32-accumulating integer mma. If mma...f32.s8.s8.f32 existed, the I2F boundary vanishes. Probe result: ptxas (CUDA 13.0, PTX 8.8/9.0) reports "Unexpected instruction types" for every target sm_80 through sm_121; the control group's s32-accumulator form (.s32.s8.s8.s32) assembles clean across the board. The conclusion is not a syntax problem but an ISA fact: integer mma comes only with an s32 accumulator.

Could one fall back to a floating-point mma with f16/tf32/fp8 operands and f32 accumulation to emulate it? No — for the product to pass the parity gate (1e-3), the operands must carry the q8·scale product bit-exactly. int8's 8-bit magnitude times an f16 scale's 11 significant bits needs ~18 significant bits in the product; f16 has only 11, tf32 only 11, fp8 only 3–4. Every operand format loses precision before the multiply, and the parity gate fails.

Route B: 2-chunk integer merge. Add the integer dot products of two adjacent 32-k chunks in s32 first, then do one I2F, hoping to halve the I2F count. Overflow-wise fully feasible (per chunk |S| ≤ 32·127·15 ≈ 2^16, two merged chunks are far from 2^31 — the meaning of the record's "regardless of the 2^21 bound": range was never the obstacle). But algebraically void: the rescale coefficients differ per chunk (get_scale_min_k4(c&7) yields per-sub-block d/dmin, and the A side's d/ssum also varies per block); the merged S = S₀+S₁ can only be multiplied by one pair of coefficients, and the information needed to split back into two chunks is already lost. The commit message's conclusion: under any pairing, the error is O(1e2) — 5 orders of magnitude worse than the 1e-3 parity gate.

Why llama.cpp doesn't pay this tax: the key reading from r9's reference decode — llama's MMQ uses float sum accumulators (16 mma chains, each warp accumulating floating-point sums directly per chunk). It performs the same two-term rescale but has no I2F boundary: the mma emits floats and the rescale is a pure FFMA chain. Our instruction-count chase keeps running into this structural difference: the accumulator bit-width/bandwidth the integer mma saves must be bought back in the epilogue with I2F+rescale. r15's probe formally archived "can this tax be waived" as: no.

2.4 Why the fold is numerically safe

The new code's only numeric change is one multiplication's association order:

old: (da · dmv) · sa     new: (da · sa) · dmv

Three reals, two floating-point multiplies, only the parentheses exchanged — the difference is at ~ulp(|sum|) scale (the commit message's verbatim "numerics within ~ulp(|sum|)"), two orders of magnitude of margin against the 1e-3 gate. Note this is not "harmless enough to skip verification": floating-point multiply is commutative but not associative, and such changes still pass every gate (see §4) — it is just that the error class here is foreseeable and explainable.

2.5 Why this magnitude clears the +1.5% bar

r13's 1:1 tracking law (duration ≈ instruction count / 0.10–0.15 warp-inst/ns) predicts: an instruction stream −8.1% should buy back kernel time of the same order. The kernel measured −4.9% — sub-linear, because SpeedOfLight Compute (SM) is only 31.5%: the kernel is stall-bound, issue is not the bottleneck, and only part of the removed instructions actually shortened the critical path. This is a boundary fact delivered alongside the step: in this regime, ALU-class cuts' payout decays from 1:1 (r17 would push this rule to its extreme).

3. Implementation

3.1 Design choices (why this shape and not another)

  • Probe first. Verifying the f32-accumulate mma costs one raw-PTX file + one ptxas call; writing the kernel first and discovering the instruction doesn't exist wastes a full implement+verify round. In hindsight the commit message's structure — probe verdict in the first paragraph, the landing after — is the record form of "falsify first, design second".
  • The veto happened on paper. The 2-chunk merge needed no code to reject: per-chunk-varying coefficients is a fact visible to the eye in the source (get_scale_min_k4(c&7) called per sub-block inside the unrolled loop), and one algebraic argument closes it.
  • The landed choice is "the cheapest cut within the same instruction-stream family". The original target (eliminating the I2F boundary) died at the ISA level, but the motive behind it — shrinking the epilogue's floating-point op stream — still stood. The rank-1 fold touches no I2F, no term1, no layout; it only re-orders term2's multiplication tree — the smallest-surface, cleanest-numeric cut available.
  • Wide kernel only. The wide kernel (mmq_raw_wide_nt_kernel) is the performance path (1225–1273 vs narrow's 441–481); narrow stays as the control, ported one round later (r16, next doc).

3.2 Key code

The change is 10 lines of the wide kernel's epilogue in src/cuda_kernels.cu (the b999e9a diff, +8/−2). Before/after:

// ── BEFORE (r14 state): term2 evaluated in full for every C value ──────────────
#pragma unroll
for (int nh = 0; nh < 2; nh++)
    #pragma unroll
    for (int l = 0; l < 4; l++) {
        const float da = da_q[l >> 1];
        const float sa = (float)sa_q[l >> 1];        // ← repeated per value (CSE-able)
        const int idx = (g * 2 + nh) * 4 + l;
        sum[idx] += da * dsv[nh][l & 1] * (float)clow[g][nh][l];  // term1
        sum[idx] += da * dmv[nh][l & 1] * sa;        // ← the source of 64 FMUL/chunk
    }

// ── AFTER (r15): the row-side product folded to once per row ─────────────────
// r15: the dmv correction term is rank-1 in (token, od-col) —
// the row-side product da*sa is shared by the od-col pair of
// each C fragment, so fold it once per row (16 FMUL/chunk)
// instead of once per C value (64 FMUL/chunk). The dsv term
// and the per-chunk scale application are unchanged.
const float dma[2] = { da_q[0] * (float)sa_q[0],
                       da_q[1] * (float)sa_q[1] };   // ← once per g (16/chunk)
#pragma unroll
for (int nh = 0; nh < 2; nh++)
    #pragma unroll
    for (int l = 0; l < 4; l++) {
        const float da = da_q[l >> 1];
        const int idx = (g * 2 + nh) * 4 + l;
        sum[idx] += da * dsv[nh][l & 1] * (float)clow[g][nh][l];  // term1 unchanged
        sum[idx] += dma[l >> 1] * dmv[nh][l & 1];    // ← one FFMA per C value
    }

The two subscripts deserve a read: dma[l >> 1] indexes by row (l=0,1 share row 0; l=2,3 share row 1 — in m16n8k32's C-fragment layout, a thread's 2×2 accumulators share rows within a column pair), and dmv[nh][l & 1] indexes by column pair. The fold's entire gain comes from "sharing within a row", so the row/col indexing must not be swapped (see §3.3).

Its location in the current tree: src/cuda_kernels.cu:6136 (the wide kernel mmq_raw_wide_nt_kernel's per-chunk epilogue, with da/sa read from the uint2-packed staging in one LDS.64):

// token pair (t, t+8) in one LDS.64 (uint2 tiling)
const uint2 pk2 = *(const uint2*)(sda_q
    + (size_t)kd * MMQ_WBI * 2 + g * 16 + (lane >> 2) * 2);
da_q[0] = h2f((unsigned short)(pk2.x & 0xFFFF));
sa_q[0] = (int)(short)(pk2.x >> 16);
da_q[1] = h2f((unsigned short)(pk2.y & 0xFFFF));
sa_q[1] = (int)(short)(pk2.y >> 16);
const float dma[2] = { da_q[0] * (float)sa_q[0],
                       da_q[1] * (float)sa_q[1] };

sa itself comes from the A quantize prepass's free computation (the q8_1 block's s field, around src/cuda_kernels.cu:5301's da[r] = d; sa[r] = s;) — term2's row-side data was in place at quantization time; the GEMM side merely consumes it.

3.3 Pitfalls

  • Row/col index confusion is the only high-risk point. dma by l >> 1, dmv by l & 1 — written swapped, it compiles clean, term1 stays correct, and only term2 is wrong; the output deviation drowns in the scale magnitude but is findable, at the cost of localization time. m16n8k32's C-fragment layout (4 C values per thread = 2 rows × 2 columns) is this fold's spatial precondition — draw the layout clearly before touching the epilogue.
  • Don't bet on the compiler's CSE. The old code wrote (float)sa_q[l >> 1] inside the value loop; ptxas could theoretically hoist it. r15's approach makes the sharing explicit — the dma array is constructed outside the loop. Lesson: redundancy left for the compiler to eliminate is "maybe saved"; redundancy eliminated structurally is "definitely saved".
  • Probes should use raw PTX, not inline asm. Inline asm's error paths tangle in constraint parsing; feeding a raw .ptx file straight to ptxas yields a pure instruction-type verdict ("Unexpected instruction types"), and swapping the accumulator type in the same file runs the control group — the difference between two calls is the conclusion.

4. Verification

  • Parity gate (KD=4 + KD=8): against the CPU reference implementation at 1e-3 tolerance — defends against term2 indexing errors and re-association drift beyond expectation; the ulp-scale error class argued in §2.4 is backed by this gate.
  • Suite 166/0/3: the full test suite (166 pass / 0 fail / 3 skipped) — defends against collateral damage beyond the GEMM.
  • greedy-32 token identity: 32 greedy-decoded tokens compared token-for-token against the default path — defends against "numbers pass but semantics drift".
  • ncu counter A/B (q-proj GEMM: inst/GMAC + duration + SM%) — defends against "the wall clock moved but nobody knows why": this step captured −8.1% inst and −4.9% duration together, and the ratio is itself evidence of the stall-bound regime.
  • Interleaved 3× same-window A/B (KD=8 3/3 consistent) — defends against machine drift polluting the delta (r59b's lesson, by now internalized as procedure).

5. Results

Metricbefore (r14)after (r15)Δ
Wide kernel KD=8 whole-prefill1271 tok/s1295 tok/s+1.9% (3/3 consistent)
Wide kernel KD=4~1226 tok/s~1233 tok/s+0.5% (noise band)
Narrow kernel control—stable—
ncu warp-inst / GMAC (q-proj)9.45 M8.69 M−8.1%
Kernel duration2.378 ms2.262 ms−4.9%
SpeedOfLight Compute (SM)—31.5%stall-bound evidence

The instruction stream's campaign trajectory: r13 baseline 10.14 M/GMAC → r14 9.45 → r15 8.69 (llama.cpp 6.06; the control is llama-bench @ ca3d5a3e1's same-shape q-proj kernel). This round also drew a floor: the rescale epilogue is down to ~346 ops/chunk, and the I2F stream is ISA-irreducible — as long as the s32 accumulator is the only form, one I2F per C value is the fixed tax of integer GEMM.

Reading these numbers in the campaign's coordinates at the time: wide KD=8's 1295 tok/s was still ~1.8× from the same-session f16 default path (2284–2370), and the per-GMAC duration gap measured at r14 was 70.4 vs llama's 41.1 µs/GMAC; r15 cut 0.76 M out of the 3.39 M instruction gap, picking the per-MAC instruction-efficiency main line's low-hanging fruit clean. The next lever after r15's ALU cuts was named as warp-tile shape (2× od-rows per warp) or prefetch/stall structural work — the former is r17's story (reverted), the latter cashed in at r20 (+7.1%).

6. Lessons

  1. Build the ISA boundary with a raw-PTX probe, then design around it — an instruction that doesn't exist (or that ptxas rejects) should never get the chance to grow into a kernel.
  2. Look for rank-1 (outer-product) structure in rescale/correction terms — a per-value-evaluated correction term that separates by row/column pays FMULs per row instead of per value.
  3. Floating-point re-association is an ulp-scale numeric change: it can land, but it needs an explicit argument + the full parity gate; "just moving parentheses" is not a reason to skip.
  4. When SM% is low, ALU cuts pay out sub-linearly (−8.1% inst → −4.9% duration): the instruction stream is still the correct lever, but diminishing returns have begun — this signal foreshadows r17's lesson.

← 19 · Index · 21 →

21 · r16 — Narrow kernel gets the rank-1 fold (LANDED)

Result: narrow-kernel MMQ KD=8 480.2/480.9 vs baseline 447.0/472.6 tok/s (+1.7%/+1.8%, inside the noise band) — this step's value is not the tok/s but re-aligning the instruction-stream structure of the two raw kernels: the narrow kernel serves as the wide kernel's control and fallback path, and cannot keep living with an old epilogue. Commit: 151fa97. Date: 2026-09-03 (same-morning follow-up, 26 minutes after r15 landed).

1. Background — where things stood

r15 landed the rank-1 term2 fold in the wide kernel (mmq_raw_wide_nt_kernel): KD=8 1271 → 1295 tok/s, instruction stream 9.45 → 8.69 M/GMAC. The boundary decision at landing was to touch only the wide kernel — it is the performance path (1295 vs the narrow kernel's ~480, 2.7×), and the narrow kernel was kept as the control for that round's A/B.

But that left a structural problem. At this point the narrow kernel mmq_raw_nt_kernel sat at the intersection of three roles:

  1. Control: every later wide-kernel experiment's "invariant" is vouched for by it — if the control itself carries an epilogue with a different instruction-stream structure, the comparability of the "control stable" conclusion quietly leaks.
  2. Fallback path: the wide kernel at KD=8 needs 135 KB of dynamic smem; above the ~99 KB opt-in cap the launcher refuses it and falls back to the narrow kernel (the guard established in r7–r8). A slower fallback is acceptable; structural drift is not — it is another implementation of the same math.
  3. Template for variants: later kernel experiments (r17's warp remap, the more distant NB variants) all copy fragments from these two implementations. If the source epilogues disagree, everything copied needs a per-copy audit.

So r16's question is pure hygiene: port r15's fold to the narrow kernel so the two raw kernels' term2 evaluation becomes consistent again. The budget was "one-shot follow-up" — r15 landed at 10:20 in the morning, r16 was committed at 10:46, a 12-line change (+10/−2).

One more layer of on-the-ground reality about the "fallback path" role: the launcher carries an smem-cap guard for the wide kernel (the r7–r8 lesson — the wide kernel's first version at KD=8 needed 135 KB of dynamic smem and failed silently above the ~99 KB opt-in cap, producing ghost numbers; since then the launcher guards explicitly and falls back to the narrow kernel). The post-r12-rewrite wide kernel at KD=8 is 98,304 B (single-buffer synchronous staging, 1 block/SM; KD=4 is 73,728 B) — just under the line — but the guard is kept as a safety net, and any smem-budget fallback executes the narrow kernel. The fallback path's instruction structure is therefore not "historical legacy" but "active insurance" — one of the real motives for this round's port.

The config entry point too: the recommended config at the time was MINFER_MMQ=1 MINFER_MMQ_RAW=1 MINFER_MMQ_RAW_WIDE=1 MINFER_MMQ_RAW_KD=4 (recorded in r12) — MINFER_MMQ_RAW_WIDE=0 drops back to the narrow kernel. The narrow kernel is the only "structurally equivalent, smaller shape" alternative execution path in this gate set.

The expected gain was noise-scale to begin with: the narrow kernel's absolute level (~447–481) is 2.7× away from the campaign's main battlefield (wide kernel → llama.cpp's 41.1 µs/GMAC), and ±2% of session noise swamps any single-step narrow-kernel gain. The master table's verdict on this row is honest: "+1.7% (noise-band); narrow is not the perf path".

2. Principle — the GPU mechanism

2.1 The fold itself (a three-sentence refresher)

In q4_K's two-term rescale, term2 = dmv_j · (da_i · sa_i) — the column factor dmv (the per-od-col negative-bias term) and the row factor da·sa (the per-token-block A scale × in-block integer sum) form a rank-1 outer product on the (token, od-col) plane. Per-C-value evaluation re-multiplies the row factor over and over; folding = compute dma = da·(float)sa once per token row, reducing each C value to one FFMA. The numeric-class change is a single re-association of one multiplication (~ulp(|sum|)), already argued and gated in r15.

2.2 The narrow kernel's geometry: why the fold's accounting differs

First a glance at the battlefield as it stood on r16's morning (all figures are token rates measured interleaved within their own sessions; not comparable across sessions):

Kernel/configtok/s (@2K-class anchor)Source
R1 word-level MMQ (first parity-clean version)441r7–r8 window
Narrow raw kernel KD=8 (local optimum after r9's shape matrix)481r9
Narrow raw kernel (baseline band on r16's day)447–481r15/r16
Wide 16-chain kernel KD=4 / KD=8 (after r14)1225–1233 / 1273–1295r14/r15
Same-session f16 default path2284–2370r8–r12 period

The narrow kernel's position in this table determines the round's nature: it is a maintained path, not an advanced path.

The fold's benefit depends on the ratio of per-thread C values to distinct rows — a ratio set by the warp tile, which differs between the two kernels:

Narrow mmq_raw_nt_kernelWide mmq_raw_wide_nt_kernel
block tile64 tok × 64 od (MMQ_BI/BJ = 64)128 tok × 128 od (MMQ_WBI/WBJ = 128)
warp tile32 tok × 16 od (8 warps split as 4 od slots × 2 tok slots)128 tok × 16 od (each warp owns the full token axis)
per-thread C values/chunk16 (sum[16] = [nh][h][l])64 (sum[64] = [g][nh][l])
distinct token rows per thread4 (lane>>2 + t4*8, t4 ∈ 0..3)2 (uint2-packed token pairs)
folded FMUL/chunk4 (dma[4])16 (8 g × dma[2])

The narrow kernel's 16 per-thread C values land on 4 token rows; the row formula is plain in the current tree (src/cuda_kernels.cu:5801):

const uint8_t* at = qat + (size_t)(i0w + (lane >> 2) + t4 * 8) * 40;
da_q[t4] = h2f(*(const uint16_t*)at);            // f16 scale @ +0
sa_q[t4] = (int)*(const uint32_t*)(at + 36);     // in-block integer sum @ +36

That is, each thread covers the four rows lane>>2 + {0,8,16,24} (t4 ∈ 0..3), and the 16 C values map back onto exactly these four rows via h*2 + (l>>1) — after folding, dma[4] holds one entry per row, no more and no less. The fold compresses it from "multiply per value" to "multiply per row"; the commit message records the accounting: 8 I2F + 16 FMUL → 4 I2F + 12 FMUL per chunk. Half of the I2F cut comes from "structured sharing": the old code wrote (float)sa_q[h*2+(l>>1)] inside the value loop, and ptxas kept 8 copies after unrolling; the new code's dma[4] array is built explicitly outside the loop, exactly 4 copies — no longer relying on the compiler's CSE behavior.

2.3 Why a noise-scale gain is still worth landing

The narrow kernel's +1.7% falls inside the session noise band (the two same-round baseline runs, 447.0 and 472.6, differ by 5.7% — itself a confession of the noise width). But this is not a "no-op change":

  • Control validity: from here on, the "narrow control stable" conclusion of every wide-kernel experiment round is clean again — the two kernels' term2 evaluation structures agree, and the control's sensitivity to instruction-stream-class changes returns to the same baseline. This is not an abstract worry: r17 (the next round) was about to lean on the narrow kernel as its control, and r15/r16 are precisely the process of making that control meaningful.
  • Proof of pattern portability: r15's fold landed under two very different layouts (the wide kernel's uint2-packed staging vs the narrow kernel's direct 40 B raw-chunk reads) — this paved the way for extending it to the q6_K kernel family later (Era D's r38+ series; in the current tree dma[l >> 1] * dmv[...] appears in the NB/BT kernels).
  • Epilogue convergence signal: with both raw kernels' term2 folded, the question "how much epilogue is left to shave" has a common answer (the rescale floor ~346 ops/chunk recorded in r15) — the narrow kernel is no longer the exception "still carrying the old tail".

For completeness, the numeric intuition behind the fold's correctness (the same argument as r15 §2, here in the narrow-kernel version): for the same (i, j) the old code computed term2 as (da_i · dmv_j) · sa_i, the new one as (da_i · sa_i) · dmv_j — three reals, two floating multiplies, only parentheses swapped. Under IEEE 754 multiplication is commutative but not associative, so the difference bound is ~ulp(|sum|); two orders of magnitude of headroom against the 1e-3 parity gate. r15 already gated on this argument, and r16 reuses the same numeric class.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Port, don't re-derive. The math was argued in r15; the narrow kernel's C-fragment row/column mapping copies the same l>>1 / l&1 conventions; the only thing to recompute is the row count (4 vs 2). The diff is therefore just two hunks: add the dma[4] precompute, rewrite one term2 line.
  • Comments carry provenance. The ported comment keeps the "r15" tag (current tree src/cuda_kernels.cu:5805) — it records the fold's originating round, not the porting round; a reader who follows the tag finds the full numeric argument without re-deriving it in both kernels.
  • term1 and the per-chunk scale application untouched. Minimal blast radius: the only permitted numeric change is the one re-association already argued in r15.
  • The R1 word-level kernel (mmq_nt_kernel) is explicitly not ported. It is the legacy comparison path (reached only with MINFER_MMQ_RAW=0); investing in it has negative value — which makes explicit half of this doc's lesson: "keep sibling kernels instruction-compatible when a fold lands — or drop the sibling". The two raw kernels chose "keep"; the R1 kernel chose "drop" (left living with its old shape until retirement).
  • Why fold term2 and not term1 (the question review will ask): in term1 = da · dsv · clow the clow is mma's per-(i,j) output — full rank, nothing to fold; term2 is the only outer-product term built purely from per-row/per-column coefficients. r15 already drew the fold's applicability boundary; the port does not reinvent it.

3.2 Key code

The complete diff of 151fa97 (src/cuda_kernels.cu, +10/−2, two hunks):

// ── HUNK 1: dma[4] precompute (r15 pattern + the narrow kernel's 4-row geometry) ────────
             da_q[t4] = h2f(*(const uint16_t*)at);
             sa_q[t4] = (int)*(const uint32_t*)(at + 36);   // direct 40B raw-chunk read
         }
+        // r15: the dmv correction term is rank-1 in (token, od-col) —
+        // the row-side product da*sa is shared by the od-col pair of
+        // each C fragment, so fold it once per row (4 FMUL/chunk)
+        // instead of once per C value (8 FMUL/chunk). The dsv term and
+        // the per-chunk scale application are unchanged.
+        const float dma[4] = { da_q[0] * (float)sa_q[0],
+                               da_q[1] * (float)sa_q[1],
+                               da_q[2] * (float)sa_q[2],
+                               da_q[3] * (float)sa_q[3] };
         float dsv[2][8], dmv[2][8];

// ── HUNK 2: term2 rewrite ────────────────────────────────────────
                     for (int l = 0; l < 4; l++) {
                         const float da = da_q[h * 2 + (l >> 1)];
-                        const float sa = (float)sa_q[h * 2 + (l >> 1)];
                         const int jj = (lane & 3) * 2 + (l & 1);
                         const int idx = nh * 8 + h * 4 + l;
                         sum[idx] += da * dsv[nh][jj] * (float)clow[nh][h][l];
-                        sum[idx] += da * dmv[nh][jj] * sa;
+                        sum[idx] += dma[h * 2 + (l >> 1)] * dmv[nh][jj];
                     }

Two layout differences vs the wide kernel are worth comparing side by side (this is the "port ≠ copy" part):

  • Row indexing: the wide kernel has 2 rows per g, dma[l >> 1]; the narrow kernel's warp tile interleaves two A-frags (the h dimension) with 4 rows per thread, so it is dma[h * 2 + (l >> 1)]. The row sets differ (wide: uint2-packed token pairs; narrow: the four rows lane>>2 + t4*8), so the fold point must be re-derived from each kernel's C-fragment row mapping.
  • A-side data source: the narrow kernel reads da (f16 @ +0) and sa (uint32 @ +36) directly from the 40 B raw q8 block; the wide kernel reads two token pairs in a single LDS.64 from uint2-packed staging. The fold logic is transparent to both — it depends only on the rank-1 structure, not on the layout.

Current-tree location: src/cuda_kernels.cu:5805–5833 (the narrow kernel's per-chunk epilogue). The A-side sa has the same origin as in r15: a free byproduct of the A-quantize prepass (the q8_1 block's s field).

The third shape in the same file (the control): the R1 word-level kernel mmq_nt_kernel's epilogue is still the pre-fold form (src/cuda_kernels.cu:5620):

float dsv[2][8], dmv[2][8];
// …
if (HAS_OFF) sum[idx] += da * dmv[nh][jj] * sa;   // old form, legacy path kept intentionally

After this round the file genuinely carries three shapes — the two raw kernels (new) and the R1 word-level kernel (old). When reading the code, judge which form to align to by "which kernel is the active path", not by "which spelling is newer".

3.3 Pitfalls

  • Row indexing was the port's only real minefield. The two kernels' dma[...] index expressions look almost identical (l >> 1 vs h * 2 + (l >> 1)); a pure copy would move the wide kernel's 2-row indexing into the 4-row narrow kernel — it compiles, term1 stays correct, and term2 is wrong for exactly half the rows. The parity gate catches this class of error, but locating it costs far more than the two minutes spent drawing each kernel's row map before porting.
  • Don't "fix" the old kernel in passing. After the port the file contains three term2 shapes: the two raw kernels' dma[...] * dmv[...] (new) and the R1 word-level kernel's da * dmv[nh][jj] * sa (old, src/cuda_kernels.cu:5620, behind the HAS_OFF gate). The old form is a deliberately kept legacy path — a reader should not mistake it for a missed bug.
  • Noise-band measurement takes patience. Measured as a single round, the narrow kernel's +1.7% (the single pair 447 → 480) reads as a real gain; only interleaved 2× plus triangulation against r15's narrow-control series (473/472/472) puts it back in the noise band. The honesty of the conclusion depends on the density of controls, not on the sign of the delta.

4. Verification

Each gate defends one class of regression (STYLE's one-sentence-per-gate convention):

  • Parity gate ×3 configs (cuda_prefill_mmq): default (MMQ off — defends against "touched the raw kernel and broke the default path"), MINFER_MMQ=1 MINFER_MMQ_RAW=1 (KD=8, the narrow kernel's main config), and MINFER_MMQ_RAW_KD=4 (KD=4). Both depths must pass: KDR changes the restage-loop structure outside the epilogue, so the fold must hold at both cadences; HAS_OFF-class branch differences can also surface at only one depth.
  • Suite 166/0/3: the full test suite (166 pass / 0 fail / 3 skipped) — defends against collateral damage outside the GEMM.
  • Interleaved 2× A/B vs the HEAD binary: 480.2/480.9 vs 447.0/472.6 — defends against machine drift polluting the delta (the r59b lesson, internalized as procedure: only same-window pairs count).
  • Noise-frame triangulation: r15's same-round narrow-control series (473/472/472) serves as the noise-band reference, placing +1.7%/+1.8% back inside the band — defending against "reading noise as gain".
  • Deliberately no ncu: the narrow kernel is not the performance path; the profiling budget goes to the wide kernel. This is a measurement discipline — tool time follows the wall, not the change — but the flip side is that the change's conclusions lack instruction-level evidence, so the master table honestly records "no ncu (narrow is not the perf path)", making the evidence boundary explicit.

5. Results

MetricbeforeafterVerdict
Narrow kernel KD=8 (interleaved 2× median)447.0 / 472.6480.2 / 480.9+1.7%/+1.8%, inside the noise band (r15 control 473/472/472)
Parity—KD=8 + KD=4 all green (3 configs)✅
Suite—166/0/3✅
Wide kernel—untouched (1295 tok/s held)✅
Per-chunk epilogue ops (commit accounting)8 I2F + 16 FMUL4 I2F + 12 FMULfold delivered

Master table row 32's record: "480 vs 447–473 | +1.7% (noise-band) | LANDED | 3-site port of r15; narrow is not the perf path". The campaign-level accounting: the narrow-vs-wide gap stays 2.7×, and this round changed no battlefield number; what changed is that the two raw kernels' term2 evaluation structures agree again, and the fold pattern's portability was proven on a second layout (direct 40 B raw-chunk reads) — the pattern was later reused across Era D's q6_K kernel family.

A note on the "3 sites" wording in the record: the diff itself is a single epilogue in two hunks (add dma[4], drop the per-value sa conversion, rewrite the term2 line); "3-site" refers to the three code locations the port touched. Counted either way, this is a line-scale follow-up after r15; the master table gives it its own row because of the status change (the control kernel became trustworthy again), not a performance change.

6. Lessons

  1. When a fold lands, either bring the sibling kernels' instruction structure into alignment or explicitly retire the sibling — a control kernel carrying an old epilogue quietly devalues every later round's "control stable".
  2. Noise-scale gains can still be worth landing when the lever is hygiene (control validity, pattern portability) rather than speed; the criterion is "what do later measurements depend on", not "how many points did this earn".
  3. When porting a numeric fold, re-derive the indices from the target kernel's C-fragment row mapping — do not copy by textual similarity; the same math yields different index expressions under different warp tiles, and the failure symptom (term2 half wrong) is inconspicuous.

← 20 · r15 f32-accumulate probe + rank-1 fold · Index · 22 · r17 wide warp remap →

22 · r17 — Wide warp remap 32od × 64tok (REVERTED)

Result: instruction stream 8.68 → 8.18 M inst/GMAC (−5.8%, LDSM.x4 72 → 48/block-chunk), but wall clock KD=4 1250 vs 1239 / KD=8 1301 vs 1288 (+0.9%/+1.0%, noise band) — SM% 31.5 → 29.75 and SM Active Cycles +9.2%: the drop in issue efficiency ate the entire instruction reduction. Third independent confirmation: in the SM% ~30 stall-bound regime, pure per-MAC instruction reduction buys ~0 wall clock. Commit: 9d09a81 (record only; kernel code reverted to 708067d and cmp-verified, the remap variant lived only in /tmp and was never committed). Date: 2026-09-03.

1. Background — where things stood

r15 had landed the rank-1 fold that very morning (wide kernel 1295 tok/s @KD=8, 8.69 M inst/GMAC), and its record pointed the next lever at two directions: warp-tile shape (2× od-rows per warp) or prefetch/stall-structure work. r17 chose the former — the next stop on the "per-MAC instruction efficiency" main line that r13 (counter forensics) and r15 had both pointed to.

The chain of reasons for picking warp shape runs like this:

  1. r9 had noted a shape difference while decoding the llama.cpp reference: llama's Q4_K MMQ config is 256 threads, tile I=128 × J≤128, 16 mma chains per warp per chunk — the chain count matches ours (post-r12), but the warp's split differs: llama gives each warp 32 od-rows × 64 tokens, ours is 16 od-rows × 128 tokens.
  2. r13's ledger: ours 10.14 → 8.69 M inst/GMAC after r15, llama 6.06 M — still 2.6 M apart. Within the instruction stream, the A/B fragments' smem reads (LDSM) are one of the biggest items, and the LDSM count is set directly by warp shape (how many chains reuse each fragment).
  3. r14's precedent: routing B fragments through a single ldmatrix.x4 (fewer smem ops) had bought +18.5–30% — "fewer/wider smem ops" was a proven paying lever class at the time, and warp remapping is another entrance into that same lever class.

Hypothesis: change the warp shape to llama's 32od × 64tok; the A-fragment group count drops from 8 to 4 (each warp covers only 64 tokens) so A-side LDSM halves; B-side rises to 2 ldmatrix.x4 covering 32 od-rows; total LDSM per block-chunk falls from 72 to 48 (−33%), and by r13's 1:1 tracking law this should buy back a few percent of wall clock. The session's pre-registered bar: KD=8 ≥ 1350 tok/s (baseline 1288–1295, i.e. requiring +4% or more).

Two background readings in the execution environment were, in hindsight, both warnings: the wide kernel at KD=8 has 98,304 B of smem (single-buffer synchronous staging) — 1 block/SM, with latency hiding resting entirely on resident-warp depth rather than prefetch; and the issue-efficiency gap vs llama had been measured in r20's matched-nt comparison: 0.26 vs 0.42 issue/sched (same occupancy) — 3 of every 4 issue slots waiting. The instruction stream was a "thin yet fat" problem, but the binding constraint was the waiting.

In hindsight, the hypothesis erred by applying r13's tracking law outside its domain of validity — r13 itself had written the qualifier in its lessons: "per-MAC instruction count is the first-order predictor but see r25: only while issue is stall-bound" (r25 had not happened yet, but the SM% 31.5 reading already had the premise sitting right there). r17's value lies in nailing down that boundary with one clean full-remap experiment.

2. Principle — the GPU mechanism

2.1 The ledger of the two warp shapes

Before (r12–r16 shape)After (llama-shape remap)
warp owns128 tok × 16 od64 tok × 32 od
warp grid8 warps each in one 16-od slot (j0w = warp*16)wn=warp&1 → i0w=wn*64, wm=warp>>1 → j0w=wm*32
chain grid (per chunk per thread)8 A-frags × 2 B-frags = 16 chains4 A-frags × 4 B-frags = 16 chains
sum[] budget16 chains × 4 C regs = 64 floats (unchanged)64 floats (unchanged)
LDSM per thread per chunkA 8 + B 1 = 9A 4 + B 2 = 6
LDSM.x4 per block-chunk8 warps × 9 = 728 warps × 6 = 48

The unchanged chain count is the clever part of this remap: 4×4 and 8×2 are both 16 independent mma chains, so the register budget (sum[64]), staging, qb8/sds smem layout, and launcher all stay untouched — a pure index rearrangement.

2.2 The ignored side: B-fragment amortization halves

The LDSM ledger saves "read counts", but the reuse rate changes:

Reuse structure (per chunk, per warp)Before 8×2 gridAfter 4×4 grid
chains reusing each B fragment84 (amortization halves)
chains reusing each A fragment24 (amortization doubles)
block-level A-side LDSM8 warps × 8 = 648 warps × 4 = 32
block-level B-side LDSM8 warps × 1 = 88 warps × 2 = 16 (and the same 32-od span is read redundantly once each by the two warps wn=0/1)
total7248
  • Before: each B fragment (1 LDSM.x4) is reused by 8 A chains; each A fragment by 2 B chains.
  • After: each B fragment is reused by only 4 A chains; each A fragment by 4 B chains.

More hidden is the block-level redundancy: after the remap the same 32-od span is ldmatrix'd once each by two warps (wn=0/1) (B-side LDSM per block rises from 8 to 16), while A-side drops from 64 to 32 — the total does go 72 → 48, but the saving is on the better-amortized side (A) and the addition on the worse-amortized side (B). The commit message's mechanism sentence "the per-warp B-side went 2x-shared" says exactly this: B fragments' per-read amortization fell from 8 chains to 4.

2.3 Why −5.8% instructions cannot buy back wall clock: the stall-bound accounting

r13's tracking law (duration ≈ instruction count / 0.10–0.15 warp-inst/ns) carries an implicit premise: issue is the constraint. The kernel's true state at this point is SpeedOfLight Compute (SM) ≈ 31.5% — only three-tenths of each SM's issue slots are doing work, the rest are waiting (r20 later localized it: waiting on the A-staging LDG→STS chain's long_scoreboard, a ~600-cycle full-batch stall per 4-deep batch). In a regime where issue is far from saturated:

  • Cutting instructions that do not stand on the critical path → wall clock does not move (the critical path is stall, not issue width);
  • Worse: the remap makes the stall structure worse — B-side amortization halves and two warps redundantly read the same B span, so the waits on the MIO/smem dependency chain thicken.

Three ncu readings piece together the full causal chain: instructions −5.8% (the stream got thinner), SM% 31.5 → 29.75 (issue efficiency actually dropped), SM Active Cycles 4.61 → 5.04 M (+9.2%) (the same work, but the SM spent more active cycles waiting) → duration −0.3% (net effect ≈ 0). This is not "measurement noise masking the gain" — it is the gain being mechanically cancelled.

2.4 Why this veto has general value

This is the third independent confirmation of the same rule, and the three lever classes all differed: x-tile/j-tile/cp.async-db cut traffic shape (L2 bytes), r17 cut instruction count — two different dimensions of "getting thinner", both zeroing out at SM% ~30. The rule therefore upgraded from "some class of change doesn't work" to "this regime's constraint is latency/stall structure, not resource consumption volume". The rule's positive form was later cashed by r20: without changing a byte of traffic or removing a single instruction, only re-ordering the LDG→STS dependency timing, +7.1%.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Pure index remap, not one staging line touched. Staging buffers, qb8 slot layout, sds scale, launcher all stay — the experiment answers only "how much is the warp-shape variable alone worth", mixing in no other degrees of freedom. This keeps the veto conclusion clean: what was measured is the shape's own causality.
  • The unchanged sum[64] budget is both a constraint and a moat. The 16-chain × 4-reg accumulator structure is the ILP depth validated in r12; a remap that changed the chain count (say 4×4 → 2×8) would have introduced a second variable. Keeping 16 chains keeps the experiment single-variable.
  • Parity-first build discipline: full index rearrangement fails most easily in the C fragment's lane→(row, col) mapping; this experiment's first build came out parity-all-green (KD=4 + KD=8 + default), showing that m16n8k32's fragment layout math holds under both shapes — itself a validation of the layout understanding.
  • Variant kept in /tmp, not committed. After the revert the remap code had no commit value (the mechanism was vetoed), but the complete recipe went into the commit message (the five wn/wm/i0w/j0w expressions + h<4/nh<4 + 2×ldmatrix.x4) so it can be rebuilt from the record if ever needed. This is "even a veto must be reproducible" handling.

3.2 Key code

Before shape (the r17-era tree, 151fa97's mmq_raw_wide_nt_kernel; the warp mapping of the time was one 16-od slot per warp, j0w = warp*16, over all 128 tokens). A fragments: 8 groups, one ldmatrix.x4 each —

int clow[8][2][4];
#pragma unroll
for (int g = 0; g < 8; g++) {                 // 8 groups of 16 tokens = 128 tok
    const uint8_t* p = qat
        + (size_t)(g * 16 + (lane & 7) + ((lane >> 3) & 1) * 8) * 32
        + ((lane >> 4) & 1) * 16;
    unsigned r0_, r1_, r2_, r3_;
    asm volatile(
        "ldmatrix.sync.aligned.m8n8.x4.shared.b16 "
        "{%0,%1,%2,%3}, [%4];\n"
        : "=r"(r0_), "=r"(r1_), "=r"(r2_), "=r"(r3_)
        : "r"((unsigned)__cvta_generic_to_shared(p)));
    a[g][0] = (int)r0_; a[g][1] = (int)r1_;
    a[g][2] = (int)r2_; a[g][3] = (int)r3_;
}

B fragments: one ldmatrix.x4 serves both 8-od-row minitiles at once (matrices 0/1 = the two k halves of od rows 0–7, matrices 2/3 = od rows 8–15); the register distribution is exactly mma.m16n8k32's B-operand layout —

{
    const uint8_t* rb8 = qb8
        + (size_t)sg * (MMQ_WBJ * MMQ_WBQ)
        + (size_t)(j0w + (lane >> 4) * 8 + (lane & 7)) * MMQ_WBQ
        + (size_t)((lane >> 3) & 1) * 16;
    unsigned b0_, b1_, b2_, b3_;
    asm volatile(
        "ldmatrix.sync.aligned.m8n8.x4.shared.b16 "
        "{%0,%1,%2,%3}, [%4];\n"
        : "=r"(b0_), "=r"(b1_), "=r"(b2_), "=r"(b3_)
        : "r"((unsigned)__cvta_generic_to_shared(rb8)));
    b[0][0] = (int)b0_; b[0][1] = (int)b1_;
    b[1][0] = (int)b2_; b[1][1] = (int)b3_;
}
#pragma unroll
for (int g = 0; g < 8; g++)                   // 8×2 = 16 chains
    #pragma unroll
    for (int nh = 0; nh < 2; nh++)
        mmq_mma_k32(clow[g][nh], a[g], b[nh]);

After shape (never committed; rebuilt from the 9d09a81 record): the warp mapping becomes wn=warp&1 → i0w=wn*64, wm=warp>>1 → j0w=wm*32; the A-fragment loop g<8 → h<4 (64 tokens per warp = 4 groups); the B fragment single ldmatrix.x4 → two (covering 32 od-rows, nh<4); the mma chain grid 8×2 → 4×4 (still 16 chains, sum[64] unchanged); staging/qb8/sds/launcher untouched line for line.

Forensics note: per STYLE rule 0, this doc's git budget was spent on locating the before-shape code; the after-shape code exists in no commit (the variant lived only at /tmp/cuda_kernels_r16remap_variant.cu, gone after the machine rebooted), so its shape is given per the commit message's item-by-item description — no "pseudocode" excerpt is dressed up as a real source.

3.3 Pitfalls

  • First build was already parity-all-green — a "pit not stepped in" worth recording: full index remaps fail most easily in the lane→(row,col) fragment mapping (the same minefield as the r15/r16 fold series); under the 4×4 grid both the A-fragment and B-fragment lane distributions changed and both were derived correctly. The precondition: writing out on paper first what each of m16n8k32's registers should hold.
  • A pre-registered bar is what gives the veto teeth. Had the bar been set after seeing 1250/1301, +0.9% could easily have been defended into "the direction is right, keep it". This round's bar (KD=8 ≥ 1350) was registered before measurement, and the noise-band reading triggered the revert directly.
  • Reverts must be cmp-verified: after reverting the kernel to 708067d, a byte-for-byte comparison against the pre-experiment state (the commit message's own words: "cmp-verified") — preventing a "revert" from becoming "yet another unverified change".

4. Verification

  • Parity gate (KD=4 + KD=8 + default, three configs): after the full remap the output still matches the CPU reference — defends against lane-mapping errors; passed on the first build.
  • Interleaved 3× A/B vs the HEAD binary: KD=4 1250 vs 1239, KD=8 1301 vs 1288 — defends against machine drift; the readings landed in the noise band (+0.9%/+1.0%), and against the pre-registered bar ≥1350 the verdict is a loss.
  • ncu counter A/B (q-proj KD=4): inst/GMAC, LDSM counts, duration, SM%, SM Active Cycles — this round ncu was not "a nice-to-have" but the veto instrument: the wall clock only says "no effect"; the counters explain "why no effect" (the −5.8% instruction cut was cancelled by the issue-efficiency drop).

5. Results

Metricbefore (r16 tree)after (remap)Verdict
KD=4 whole-prefill12391250+0.9%, noise band
KD=8 whole-prefill12881301+1.0%, noise band; bar ≥1350 missed
ncu inst / GMAC8.68 M8.18 M−5.8% (the stream really got thinner)
LDSM.x4 / block-chunk7248−33%
ncu duration2.262 ms2.256 ms−0.3% (the wall does not move)
SpeedOfLight Compute (SM)31.5%29.75%issue efficiency dropped instead
SM Active Cycles4.61 M5.04 M+9.2% (more waiting)

Veto mechanism (why it was reverted, and what the evidence chain is):

  1. The wall clock +0.9%/+1.0% falls in the noise band, 4 percentage points short of the pre-registered bar (KD=8 ≥ 1350, requiring +4%+);
  2. Mechanically this is not "not yet tuned": the −5.8% instruction cut was fully cashed in the issue stream, but SM% fell 1.75 points instead and Active Cycles rose +9.2% — B-fragment amortization halved plus dual-warp redundant reads of the same B span made the stall structure worse, cancelling the thinner stream;
  3. This is the third independent confirmation that "at SM% ~30, per-MAC instruction cuts buy ~0 wall clock" (the first two: the traffic-shape cuts of x-tile/j-tile, and the staging family); the rule upgraded to a regime-level constraint;
  4. Retry conditions: only after the stall structure is fixed and SM% rises materially can the −5.8% instruction cut cash out at ~1:1 — and r20 (split-phase A staging, +7.1%) is precisely the correct fix that repaired the stall structure first; after r28 the NB kernel re-defined the occupancy regime at 2 blocks/SM, the wide kernel was no longer the performance path, and this shape experiment was never re-run under the new regime.

Revert execution: the kernel was restored to 708067d (r16's fold kept) and cmp-verified; the remap variant lived only in /tmp (never committed), with the recipe fully recorded in 9d09a81's commit message.

The second half of this rule's story (later verification within the same campaign, showing the "retry conditions" are not lip service): r25's SASS census found that 100% of the stream's excess was supporting instructions (int ALU 77%), and the wall clock stayed lazy after cuts — the same regime conclusion as r17; while r29, after the NB kernel landed at 2 blocks/SM, measured int ALU −25% / inst −6.5% buying a real +2.80% wall clock — once occupancy rose (the SM%-degraded premise removed), instruction cuts' payout recovered to near 1:1. r17's veto is therefore "wrong timing" rather than "wrong direction forever": in the 1 block/SM stall-bound regime it is zero; in the 2 blocks/SM regime it is real.

6. Lessons

  1. An instruction cut's payout is premised on SM%: when issue is far from saturated (~30%), cuts land where they don't stand on the critical path and the wall clock does not budge — read SpeedOfLight first, then pick the lever class.
  2. "Getting thinner" has two dimensions (bytes, instructions), and both zero out in the stall-bound regime; the third axis that can move the wall clock is the stall structure itself (dependency timing, staging phases) — r20's +7.1% was the first cash-out of that axis.
  3. Shape is not a separable degree of freedom: llama's warp shape is embedded in its own staging depth, ITER_K, and accumulator structure; porting the shape without the surroundings yields not "llama's shape" but a "mismatch".
  4. A veto experiment's entire value is in the mechanism readings: the wall clock says "useless"; ncu explains "why useless and when it is worth another try" — otherwise the revert is just a waste.

← 21 · r16 narrow-kernel rank-1 fold · Index · 23 · r18 load-time B pre-expansion →

23 · r18: Load-time B pre-expansion — staging becomes a bulk copy (W_exp's debut, REVERTED)

Result: parity green on first build; wall clock KD=8 +0.9% (noise band), KD=4 −19% (median, variance out of control); kernel duration KD=8 −4.5%, KD=4 +2.1%; warp instruction count only −0.3%; the price +5.8 GB of VRAM. The ≥1350 tok/s bar decisively missed → reverted; the EB/SB machinery preserved as a prior asset for later L2-residency / pre-expansion experiments. Commit: 0a26b35 (record commit; the code left no separate commit with the revert — the originally cited HEAD 1e0dded is an anchor that was swallowed by an amend and no longer resolves, see master-table footnote 3). Date: 2026-09-03.

1. Background — where things stood

2026-09-03, the P6 q4_K MMQ campaign at mid-game. The wide kernel (mmq_raw_wide_nt_kernel ) had just taken two structural landings in a row: r12's 16-chain warp tile + ldmatrix (441 → 1020–1058, +2.3×), r14's B-fragment ldmatrix-ization + widened scale reads (→ 1225 @KD=4 / 1273 @KD=8), and r15's rank-1 term2 rescale added another +1.9% (KD=8 1295). But the front was beginning to show fatigue: r16's port of the rank-1 fold into the narrow kernel bought only a noise-level +1.7%; r17's pure index remap measured +0.9%/+1.0%, landed in the noise band and was reverted.

The campaign's provisional bar stood at ≥1350 tok/s, and the best config was stuck at 1295 — a gap of only ~4%, yet three consecutive steps harvested only noise. More important than the numbers was the qualitative judgment: r13 had concluded (per-MAC warp instruction count is the first predictor, 10.14 vs 6.06 M/GMAC), but the IMMA counts were dead even on both sides — what was slow was not the tensor cores' work, but the supporting instructions and the latency structure. r17's revert tightened the line further: with SM% stuck at ~30, pure instruction removal buys no wall clock. The bottleneck was classified as "latency/stall structure" — something was making the 8 resident warps wait.

r18's hypothesis came from there (labeled in the record as the Phase-7 task-1 hypothesis): the stall structure's suspect was the B-side staging round trip. Each super-block's staging is two-phase — first move the 144 B q4_K raw super-block (16 B header + 128 B packed nibbles) into smem, then use ALU to unpack/expand it in-loop into qb8 [8][128][48] (one int8 per k). This "raw moved in → ALU expand" segment happens between compute phases, and at 1 block/SM with 8 resident warps there is no other warp to fill the gap.

There was also an earlier seed: r9, while decoding the llama.cpp reference implementation, had written "the remaining lever is llama's pre-arranged mma-fragment B layout, which can be produced at weight-load time" — r18 was the first attempt to cash that sentence. (The day's box absolute values ran ~9% below the r15/r17 sessions; this doc takes only same-session interleaved A/B relative values.)

2. Principle — the GPU mechanism

First quantify what "pre-expansion" is replacing. The wide kernel's thread block covers 128 od rows × 128 tokens; B-side staging works at super-block granularity (256 k, 8 sub-blocks of 32 k); the per-(super-block, od-row) cost:

PhaseBytes/ALU
read raw144 B (128 B qs nibbles + 16 B header d/dmin/scale codewords)
ALU unpackeach 32 B word group split into high/low nibbles, & 0x0F0F0F0F mask
write smem256 B of payload into a 384 B padded slot (48 B/slot × 8; the excess is ldmatrix alignment padding)
scale sideone header parse per chunk + get_scale_min_k4 table lookup → (d·sc, −dmin·m) float2

The key arithmetic is how much of the instruction stream this ALU occupies — r18's ncu gave the after-the-fact answer: the entire expansion is ~0.1% of the per-block warp stream (total warp instructions differ by only −0.3% before/after). Hoping to speed things up by "cutting the staging ALU" runs into a lever ceiling of 0.1% itself; the real bet was on dependency structure: the two-phase raw→qb8 staging strings a smem dependency chain through "move → expand → compute", and at 1 block/SM this chain has nowhere to hide.

The pre-expansion plan reshapes the bet into another form — generate two planes once at load time:

  • EB plane (expanded B): per-k int8, layout [sb][od][8][32], 256 B per (super-block, od-row), with slot bytes kept as element-ordered raw nibbles 0..15 — byte-for-byte equal to the old in-loop expansion's output. The core layout trick: one k-tile's 128 rows × 8 slots assemble exactly into one contiguous 32 KB global span, so staging degenerates into 2048 × 16 B LDG/STS pure bulk copies with zero unpack ALU.
  • SB plane (scales of B): the per-(chunk, od-row) (d·sc, −dmin·m) float2 pairs, [chunk][od] layout, replacing the per-chunk header parse + table lookup.

The cost is VRAM: EB ≈ 1 B per weight element (raw q4_K is 0.5 B/element), and with SB roughly od·id + 8·od·id/32 bytes — measured on 7B q4_k_m as +5.8 GB (free VRAM 26.1 → ~20 GB of 130.6 total, verified per-tensor actual values). And the staging bytes actually grow: the bulk copy reads 256 B/row (EB's 8×32 B) while the ALU path reads only 144 B/row (the raw super-block) — B-side global read bytes ~2×. This is a trade of "bytes for ALU + for dependency depth": r13/r17's evidence points at latency (the KD=8 kernel's −4.5% confirms the latency gain exists), while KD=4's failure shows the price on the other side of the trade (see §5).

3. Implementation

Forensics note: r18's code vanished with the revert; the complete variant survives only in record commit 0a26b35's narration and in /tmp artifacts (/tmp/patch_p7_t1.py etc., lost on machine reboot). The before excerpt is taken from the current tree, i.e. the original path it replaced; the after is reconstructed from the record's layout and byte-for-byte contract narration.

3.1 Design choices (why this shape and not another)

Expand at load time, not in the kernel. The expansion is a per-weight invariant, while the in-loop version re-expands every super-block once per block that touches the weight — moving it to load time replaces hot-path repetitive labor with a one-time, off-hot-path cost, the same family as 8p's persistent f16 cache.

Two planes, not one mixed plane. Nibble expansion and scale decoding have different lifecycles (qb8 changes per super-block — at KDR=4 two consecutive k-tiles share one super-block and can skip the re-move; sds changes per k-tile). Splitting the layouts lets EB achieve the "one k-tile = one contiguous 32 KB" bulk-copy form; SB's [chunk][od] aligns directly with the staging loop's enumeration order.

Byte-for-byte contract alignment, not layout reshuffling. EB's slot bytes are defined as element-ordered raw nibbles 0..15 — byte-for-byte identical to the old in-loop expansion's output. Getting the layout change to "smem end-state byte-identical" first leaves parity with only one variable, copy correctness — and indeed it was green on the first build, with none of r14's bisect debugging.

Registry + fallback, not a hard switch. register_weight_q4k_expanded registers the EB/SB buffers into the new mmq_expanded registry keyed by the raw wptr; the wide kernel launch requires the planes to exist and falls back to raw-W narrow-kernel staging if missing; the raw tensor itself stays registered and other paths are unaffected.

3.2 Key code

Before — the in-loop B expansion r18 replaces (current tree src/cuda_kernels.cu, the B section of the wide kernel's RAW_STAGE macro; since r18 this section had its A side rewritten again by r20/r22, but the B-side core is unchanged):

/* B: expand ONE 256-k super-block to per-k int8 AT STAGING - raw
 * nibble values 0..15 ... At KDR=4 two consecutive k-tiles share the
 * super-block: restage only when this k-tile starts a new super-block. */
if (((kt) * KDR & 7) == 0) {                       // restage-skip guard
for (int x = threadIdx.x; x < MMQ_WBJ * 4; x += blockDim.x) {
    const int r = x >> 2, p = x & 3;               // od-row, nibble-pair
    const int j = j0 + r, sb = ((kt) * KDR) >> 3;
    uint4 v0 = make_uint4(0,0,0,0), v1 = make_uint4(0,0,0,0);
    if (j < od && sb < nsb) {
        const uint8_t* src = W + (size_t)j * ((size_t)nsb * 144)
                          + (size_t)sb * 144 + 16 + p * 32;   // raw qs
        v0 = *(const uint4*)(src);
        v1 = *(const uint4*)(src + 16);
    }
    const unsigned M = 0x0F0F0F0Fu;                // ← exactly what r18 eliminates
    uint8_t* dst = qb8 + (size_t)(p * 2) * (MMQ_WBJ * MMQ_WBQ)
                 + (size_t)r * MMQ_WBQ;            //   mask unpack + reordered writes
    *(uint4*)(dst)      = make_uint4(v0.x & M, v0.y & M,     // 384 B padded
                                     v0.z & M, v0.w & M);    //   smem slot
    /* ... the high-nibble half likewise written to dst1 = dst + MMQ_WBJ*MMQ_WBQ ... */
} }

r18's after (reconstructed per the record): the staging B section no longer reads W and no longer has the & M unpack; it does pure 16 B copies from the EB plane at k-tile base addresses — one super-block is exactly 128 od-rows × 8 slots × 32 B = 32 KB contiguous, each thread enumerating one of the 2048 16 B moves, LDG.128 → STS.128, zero ALU.

The scale section likewise — this per-chunk parse:

for (int x = threadIdx.x; x < MMQ_WBJ * KDR; x += blockDim.x) {
    int r = x % MMQ_WBJ, kd = x / MMQ_WBJ;
    int j = j0 + r, c = (kt) * KDR + kd;
    float dv = 0.0f, mv = 0.0f;
    if (j < od && c < nchunk) {
        const uint8_t* blk = W + (size_t)j * ((size_t)nsb * 144)
                          + (size_t)(c >> 3) * 144;
        float d    = h2f(*(const uint16_t*)blk);            // f16 d
        float dmin = h2f(*(const uint16_t*)(blk + 2));      // f16 dmin
        uint8_t sc, m;
        get_scale_min_k4(c & 7, blk + 4, &sc, &m);          // 6-bit codeword table lookup
        dv = d * (float)sc;  mv = -(dmin * (float)m);
    }
    sds[(size_t)kd * MMQ_WBJ + r] = make_float2(dv, mv);
}

is replaced by a direct read of the SB plane's float2 — (d·sc, −dmin·m) is already a load-time product.

Load side (reconstructed): one expand_q4k_kernel device-side launch + stream sync per tensor, producing EB/SB and registering them; at wide-kernel launch the raw wptr is looked up to decide between the pre-expanded path and the fallback.

3.3 Pitfalls

  • No layout pit was stepped in — because the contract was nailed down first. r14 had once silently rewritten qb8 due to a uint4 tiling overrun (caught only by bisect); r18's EB was defined as "byte-for-byte equal to the old expansion output", locking the layout variable away in advance, and parity was green on the first build.
  • A hidden test-coverage trap was plugged in advance: the expanded buffers were also registered into the parity fixture; otherwise the fallback logic would have let the wide kernel quietly bypass the new path in tests — what was measured would have been the fallback branch.
  • KD=4's variance is itself a signal: three interleaved runs produced an outlier like 816.2 (median −19%, variance out of control), meaning not a stable slowdown but hitting some resource boundary (the record attributes it to the staging-byte excess, see §5 reading 2).

4. Verification

  • 8-shape parity sweep (KD=4 + KD=8 + default + narrow-kernel control) green on first build — defends against the layout/stride-error class of "runs but computes wrongly" regression, and incidentally confirms the fallback branch works.
  • Parity fixture registers the expanded buffers — defends against the false green of "fallback logic means the new path was never tested".
  • 3 same-session interleaved A/B runs (baseline binary vs r18 binary) + narrow-kernel control — defends against cross-session machine drift and global environment noise (both narrow-kernel pairs are in the noise band, showing no environment-level discontinuity; the day's box ran ~9% low, making interleaved relative values the only trustworthy source).
  • ncu single-launch profiling (q-proj, launch 1, nt 2630, od=id=3584, grid (21,28), per-GMAC = 33.78e9) — separates the two opposite-direction effects of "latency gain" and "byte cost", grounding the wall-clock conclusion in kernel-level mechanism.

5. Results

Wall clock (7B @2630 tok, same-session 3× interleaved medians):

Configbaseliner18 expanded-BΔ
wide KD=8 (default)1170.9 / 1164.7 / 1155.5929.8 / 1181.7 / 1175.0+0.9% (noise band)
wide KD=41145.7 / 1141.3 / 1130.61071.6 / 816.2 / 907.5−19% median, high variance
narrow kernel (control)463.3 / 429.9444.5 / 441.5noise

Kernel level (ncu, q-proj):

MetricKD=4 baselineKD=4 r18KD=8 baselineKD=8 r18
warp instructions8.68 M8.65 M (−0.3%)8.49 M8.46 M (−0.3%)
duration2.288 ms2.335 ms (+2.1%)2.155 ms2.059 ms (−4.5%)
Compute (SM)31.1%30.4%32.4%33.7%
Memory Throughput44.2%36.5%46.2%43.0%

Three readings (summarizing the record's own text):

  1. The staging-ALU cut is instruction-neutral: the expansion ALU was only ever ~0.1% of the warp stream — r15/r17's "cutting instructions at SM% 30 is useless" rule thereby extends to staging instructions.
  2. The latency gain is real but small and config-dependent: KD=8's −4.5% kernel duration translates to only +0.9% wall clock (the GEMM is only one of ~200 launches, and only the q4_K matmul benefits); KD=4 instead went +2.1% at the kernel level — the bulk copy reads 256 B/row while the ALU path reads 144 B (~2× staging bytes), and on KD=4 the excess outweighed the removed latency.
  3. The fourth independent confirmation that the stall structure is not in the B expansion: SM% does not move (30–34), and removing the B-expansion work did not shrink the stall — the latency binding the kernel is elsewhere (in hindsight, exactly the A-side staging, r20's subject).

Veto mechanism: the bar ≥1350 was decisively missed, and at every level the reason is structural rather than measurement fluctuation — no gain at the instruction level (ALU share 0.1%), a reversal at the kernel level on KD=4 (the bytes-for-latency trade nets a loss at shallow staging depth), the wall clock diluted by "one launch among many + a single weight type benefits", plus the permanent +5.8 GB VRAM price. Reverted to the then-HEAD (cmp-verified).

Under what future conditions a retry is worthwhile: only when the EB plane gains a consumer that is not staging, changing the trade's denominator — the record's explicit candidate is "L2-residency experiments on pre-expanded weights" (this is exactly why the EB/SB machinery was fully preserved). That prophecy was later partially cashed: the pre-expansion route revived on q6_K as the W_exp plane and landed (from r43 on), but hit a dense-plane stride mismatch along the way — a story for later, see docs 46/47.

6. Lessons

  1. Staging-ALU cuts extend r17's rule: with SM% stuck at ~30, cutting instructions (staging instructions included) buys no wall clock — first ask what percentage of the instruction stream the operation occupies (this time: 0.1%).
  2. "Eliminating work" at a byte cost is a trade, not a gain: −0.1% ALU for ~2× staging bytes nets −19% at shallow staging depth (KD=4).
  3. The arithmetic of kernel-level wins being diluted by the wall must be done in advance: −4.5% kernel → +0.9% wall; the dilution factor = the kernel's share of the wall × the share of weight types that benefit; when the bar is set at wall-clock level, compute the lever's global ceiling first.
  4. A reverted machine can be a valuable asset: the EB/SB materialization machinery was fully preserved, becoming the direct precursor of the later L2 experiments and even the q6_K W_exp route — a revert decision and asset preservation are not in conflict.

← 22-r17-wide-warp-remap · Index · 24-r19-weight-l2-residency →

24 · r19: Weight L2 residency — __ldg imperceptible, persisting window catastrophic (REVERTED)

Result: both levers built in the same round. The __ldg read-only path: neutral (KD=4 median +2.0%, KD=8 sunk in noise) — the weight tiles were already being re-read from L2, and L2 is not the constraint (SOL 37–46%). The MINFER_MMQ_L2WIN=1 persisting window (hitRatio 1.0): −50%, 6/6 consistent — the 12.8–34 MB of resident rows per weight squeeze the C stores (37 MB per GEMM), activations and KV out of the normal L2, and every re-mark churns the carveout. The bar ≥ +5% decisively missed; both levers reverted. Commit: 072dd9a (record commit; the code left no separate commit with the revert — the originally cited revert anchor HEAD 384b3d9 is an unresolvable pre-amend twin, see the master-table footnote-3 pattern). Date: 2026-09-03.

1. Background — where things stood

r19 shares r18's day and session, is item 2 of the Phase-7 tasks (task-2), and steps directly on r18's autopsy conclusion. r18 had just proven: cutting the B-expansion ALU buys nothing, with SM% motionlessly stuck at 30–34, "the latency binding the kernel is elsewhere". But r18's revert record left one door open — the EB/SB materialization machinery was fully preserved, "as the precursor for any future L2-residency experiment". r19 puts that L2 direction directly on the table: without waiting for the pre-expanded planes, ask on the raw weight byte span itself — when the same weight bytes are re-read over and over, do they actually come from DRAM or L2? Can they be forced to stay resident in L2?

The problem's magnitude deserves computing first. The wide kernel's tile is 128 tokens × 128 od; the B (weight) panel is blocked by od column, and every x-block (token direction) must re-read the same B panel completely once more. At nt=2630 the x-block count = ⌈2630/128⌉ = 21 — the same weight bytes are re-read 21× within a single launch; adding the cross-launch re-reads within a layer, the total repeated access to weight bytes is substantial.

The campaign history had in fact probed this question sideways twice:

  • r7–r8's x-tile lesson said "B-traffic reduction is a dead lever in the face of L2-resident re-reads" (L2 absorbed the B re-reads);
  • r13's counter forensics measured our per-MAC L2 read traffic at only 1.6× llama.cpp's — far below the 1.67× instruction-count gap.

But both are indirect evidence. r19 wanted direct intervention: use CUDA's two levels of L2 control (the read-only path hint, the persisting access policy window) to try turning "weight bytes resident in L2" from accident into policy.

The session environment remained harsh: box absolute values ran ~9% below the r15/r17 sessions, and the day drifted violently (the KD=8 baseline's three interleaved runs gave 780.0 / 1061.7 / 1252.4 — a ±25% spread). This sets r19's metrological situation: only signals whose effect size exceeds the ±25% noise band are judgeable. In hindsight this constraint mattered enormously — it is what made "−50%" the only judgeable signal in the whole round.

2. Principle — the GPU mechanism

The two levers' mechanisms are entirely different; audit them separately.

2.1 Lever one: __ldg (the read-only data path)

__ldg routes a load through the non-coherent read-only path, and on that basis the compiler can emit LDG.E.CI-class instructions and cache the data on the path optimized for read-only data. It matters for performance only under two premises:

  1. The data really is read-only (r19's targets — the weight bytes within the wide kernel's lifetime — satisfy this);
  2. The compiler didn't already know it.

The latter is key: the wide kernel's weight pointers are declared const __restrict__, and nvcc for such pointers already tends to emit non-coherent loads — if it already does, __ldg is a pure no-op. Also, it only affects cache-path selection and changes no byte counts: if the bottleneck is not the source of the weight bytes (DRAM vs L2) at all, the hint has nothing to push on.

2.2 Lever two: the L2 persisting window (cudaAccessPolicyWindow)

This is CUDA's L2 residency mechanism. Set an access policy window on a stream:

Parameterr19's valueSemantics
base_ptr / num_bytesthe raw q4_K weight byte span (12.8–34 MB scale)the address range the window covers
hitRatio1.0accesses inside the window are marked hitProp with this probability
hitPropPersistinghit lines can only be evicted by other persisting accesses or an explicit reset
missPropStreamingunmarked accesses take the streaming (low-priority residency) path

Alongside, cudaLimitPersistingL2CacheSize (the persisting carveout) must be raised to its maximum — r19 does this once at per-process init; the window is set before each wide-kernel launch.

2.3 What hitRatio 1.0 does, in arithmetic

The three L2 working sets of a single GEMM:

  • Weights (the object the window marks): 12.8–34 MB per weight; hitRatio 1.0 = every line request gets the residency mark;
  • C output: nt × od × 4 B; at nt=2630, od=3584 that is ≈ 37 MB/GEMM of write traffic, on the Normal path;
  • A activations and KV traffic: smaller than C, but competing on the Normal side all the same.

The carveout is raised to the driver-allowed ceiling (persistingL2CacheMaxSize), while the resident set the window demands is the same size as the carveout or larger — the result has two layers.

Layer one: capacity hijacking. A fixed slice of L2 is monopolized by weight lines; the C stores' 37 MB, activations, KV can only circle in the remaining normal L2, and the miss rate climbs.

Layer two (more hidden): re-mark churn. When the window-marked weight set exceeds the carveout, newly accessed weight lines must evict old persisting lines to become persisting themselves — "residency" degenerates into a high-frequency mark/evict cycle, every line shuttling in and out of the carveout, with L2 control-path metadata overhead stacked on top. The mechanism itself is deterministic — it deterministically does the wrong thing.

2.4 The premise both levers shared and should have checked first

Is this kernel's L2 actually tight? ncu's Speed-of-Light reading had been sitting there all along: L2 Cache Throughput 37–46% SOL. A resource below half saturation cannot be the source of a −50% or +5% no matter how its residency policy is optimized. Half of r19's value is measuring this veto; the other half is nailing the "read SOL first, then act" ordering into the methodology.

3. Implementation

Forensics note: r19's code vanished with the revert (the variant survives in artifacts like /tmp/patch_p7_t2.py and /tmp/cuda_kernels_p7t2.cu, which do not travel with the repo). This section's code excerpts are the target load sites these two levers aimed at on the current tree (src/cuda_kernels.cu wide kernel RAW_STAGE macro); the levers themselves are reconstructed from the record's narration — their form is only a few lines of API calls, which the narration reproduces exactly.

3.1 Design choices (why this shape and not another)

The read-only hint hits only the two weight-side load classes. Weights are read-only within the kernel's lifetime; the activation side is not — __ldg's targets are limited to: (a) the B-staging super-block uint4 bulk reads (the mainstay of the 144 B raw rows), (b) the scale-header uint16 reads (the d/dmin f16 headers). This is parity-neutral by construction (a cache-path hint changes no values).

The window targets the raw weight span, set before every launch. The set being re-read 21× is the raw W byte span itself, so the window should aim at it; cudaAccessPolicyWindow is a stream (stream) attribute, not a process attribute, so the set point is before each wide-kernel launch — which incidentally guarantees only the MMQ wide kernel's launches are affected and no other work on the stream is hit by mistake. The carveout raise is a per-process cudaDeviceSetLimit.

hitRatio 1.0 first is hypothesis testing, not engineering tuning. r19's question structure is "does weight residency in L2 help at all" — the cleanest experiment pushes the hypothesis to the extreme (the whole weight set resident), and only if significantly positive comes back to selective variants; if significantly negative, the whole family closes. In hindsight, the +5% bar combined with the ±25% noise band means: intermediate variants (hitRatio ~0.25) are simply unjudgeable in this noise environment; the extreme test is actually the only metrologically meaningful first step.

The env-var gate MINFER_MMQ_L2WIN=1. The window mechanism has global side effects (the carveout is a process-level resource), so it must default off and be enabled explicitly — which also lets A/B toggle via env var on the same binary.

3.2 Key code

Lever one's aiming point — the weight-side loads in the wide kernel's staging (current tree src/cuda_kernels.cu; r19 adds __ldg to these two read classes):

/* B-staging superblock bulk read (__ldg target a): */
if (j < od && sb < nsb) {
    const uint8_t* src =
        W + (size_t)j * ((size_t)nsb * 144)      // weight row stride 144 B/superblock
          + (size_t)sb * 144 + 16 + p * 32;      // +16 skips the f16 d/dmin header
    v0 = *(const uint4*)(src);                   // → __ldg(const uint4*)
    v1 = *(const uint4*)(src + 16);
}

/* scale-header read (__ldg target b): */
float d    = h2f(*(const uint16_t*)blk);         // → __ldg(const uint16_t*)
float dmin = h2f(*(const uint16_t*)(blk + 2));

Lever two itself (reconstructed per the record; the whole logic is these few lines):

// once per process (at init):
cudaDeviceSetLimit(cudaLimitPersistingL2CacheSize,
                   persistingL2CacheMaxSize);
// before each wide-kernel launch (when MINFER_MMQ_L2WIN=1):
cudaStreamAttrValue attr = {};
attr.accessPolicyWindow.base_ptr  = /* base address of this weight's raw byte span */;
attr.accessPolicyWindow.num_bytes = /* od * nsb * 144 (12.8–34 MB scale) */;
attr.accessPolicyWindow.hitRatio  = 1.0f;        // ← the disaster parameter, proven after the fact
attr.accessPolicyWindow.hitProp   = cudaAccessPropertyPersisting;
attr.accessPolicyWindow.missProp  = cudaAccessPropertyStreaming;
cudaStreamSetAttribute(stream, cudaStreamAttributeAccessPolicyWindow,
                       &attr);

3.3 Pitfalls

  • Parity passed in one shot, on both sides — cache hints and access policy windows change no values; window-on and window-off at both KD=4/KD=8 depths all green. r19 has no correctness story, only a performance story.
  • The noise band itself is this doc's biggest pit: the KD=8 baseline's three interleaved runs 780.0 / 1061.7 / 1252.4 make a median meaningless on its own; __ldg's +2% (KD=4) is below the judgeable threshold in this environment. Lesson: a lever whose effect size is smaller than the session noise band cannot be judged even if measured — estimate effect sizes first, then order the experiments.
  • "The mechanism works" and "the mechanism helps" are two different things: the window delivered a 6/6-consistent −50% — extremely deterministic. Without the same-session three-arm interleave (baseline / __ldg / __ldg+L2WIN), a deterministic large negative effect could have been misread as an environment problem.

4. Verification

  • Parity two-way gate (window on/off × KD=4/KD=8) — defends against "a cache-policy change quietly moved the values": an access policy window should not affect bits in theory; that must be proven by measurement.
  • Same-session three-arm 3× interleaved A/B (baseline / __ldg / __ldg+L2WIN) — defends against machine drift and single-arm hallucination; the ±25% noise band makes any effect smaller than it unjudgeable, and makes the −50% conclusion unusually solid.
  • 6/6 consistency check (−50% reproduced in all three repeats of both the KD=4 and KD=8 groups) — a large negative effect's credibility comes not from its magnitude but from reproduction consistency.
  • ncu SOL reading (L2 Cache Throughput 37–46%) — provides the mechanism explanation for __ldg's neutrality: the resource is unsaturated, so optimizing its residency is moot.

5. Results

Wall clock (7B @2630 tok, same-session 3× interleaved; that day's box noise was extreme, the KD=8 baseline spanning ±25%):

Configbaseline (1e0dded state)r19 __ldgr19 __ldg + L2WIN
wide KD=8780.0 / 1061.7 / 1252.4700.7 / 1278.6 / 1282.3453.6 / 535.4 / 548.2 (−50%)
wide KD=41195.8 / 1144.4 / 1240.01201.9 / 1219.7 / 1229.9 (+2.0% median)557.6 / 575.4 / 577.2 (−50%)

Verdict: bar ≥ +5% (relative to the task-1 state); the only signal in the whole round exceeding the noise band is −50%. __ldg neutral (+2.0% KD=4 median, no consistent KD=8 gain); the persisting window catastrophic and 6/6 consistent. Both levers reverted to the then-HEAD (cmp-verified).

Veto mechanism (each lever stands independently):

  1. __ldg: the weight tiles were already being re-read from L2 (r13: per-MAC L2 read traffic only 1.6× llama's), and L2 throughput 37–46% SOL says it is not the constraint; moreover under const __restrict__ nvcc most likely already emitted non-coherent loads — a "read-only hint" is literally a no-op once the compiler already holds the information.
  2. L2WIN: the mechanism works completely (deterministic 6/6), but with the hitRatio 1.0 + weight set ≥ carveout parameter combination the direction is negative — the resident set eats the carveout, the C stores' 37 MB/activations/KV are squeezed into the remaining normal L2, and every re-mark churns the carveout.

Under what future conditions a retry is worthwhile: the record names three untested cheap variants (the budget was exhausted at the time) —

  • (a) a selective window at hitRatio ~0.25 (resident only some lines, leaving L2 for normal traffic);
  • (b) windowing only small weights (q/k/v-type weights whose sets ≤ carveout, churn-free objects);
  • (c) per-layer cudaCtxResetPersistingL2Cache (preventing cross-layer mark accumulation).

The shared higher-level premise: first find a domain where L2 really is the constraint (shapes/kernels with SOL near saturation); otherwise every L2-residency optimization is optimizing a bottleneck that does not exist. r18's preserved EB/SB machinery and this doc's named variant (b) are a natural pair — "pre-expanded plane + selective residency" remains an unclosed road.

6. Lessons

  1. A "read-only hint" is a no-op once the compiler already knows: const __restrict__ pointers most likely already take non-coherent loads — before adding the hint, confirm what information it changes.
  2. Read SOL before ordering levers: the L2 37–46% reading could have killed this direction before any code was written — optimizing an unsaturated resource is a zero-sum game of luck.
  3. A persisting window needs selectivity: hitRatio 1.0 applied to a working set ≥ carveout turns a shared cache into a churning private region — the mechanism works deterministically, including deterministically biting back.
  4. A lever whose effect size is below the noise band is unjudgeable: ±25% session noise makes a +2% "micro-win" meaningless — experiment ordering should do the effect-size arithmetic first and put the judgeable extreme hypotheses up front.

← 23-r18-load-time-b-preexpansion · Index · 25-r20-split-phase-a-staging →

25 · r20: Split-phase A staging — attribute first, shoot second: the first hit (LANDED)

Result: 7B whole-prefill KD=4 1230.4 → 1317.7 (+7.1%), KD=8 1275.8 → 1319.9 (+3.5%), 6/6 reproducible; ncu (matched nt=512 q-proj): long_scoreboard 6.22 → 2.92 (−53%), kernel duration 632.4 → 555 µs (−12.2%), warp instructions −16%; the freed stall share migrated to lg_throttle 0.33 → 2.38. Although the ≥1350 bar was not reached, it landed on the strength of the complete mechanism — the turning point where the campaign moved from "guessing levers" into "attribute first" mode. Commit: 5ac8917 (code + record in the same commit). Date: 2026-09-03.

1. Background — where things stood

2026-09-03 afternoon, the third lever round of the same day. The morning's r18 (B pre-expansion) and r19 (L2 residency) had been reverted one after the other, plus the previous day's r17 (warp remap); the campaign was already holding four independent null results, which together pointed at a judgment that kept being repeated but never named: SM% stuck at 30–34, the kernel is stall-bound, "the latency binding the kernel is elsewhere". But where "elsewhere" is, nobody knew.

r13 had given a coarse localization five rounds earlier ("duration tracks instruction count 1:1"), but its stall table proved afterwards to be doubly distorted — it used the pre-r14 old kernel, and compared misaligned launch points, nt=2630 vs llama's nt=512. The attribution infrastructure itself was untrustworthy — that is the structural cause of the three nulls in a row before this: every lever was a shot fired on an uncalibrated map.

This round's starting question was asked differently. llama.cpp's mul_mat_q vs minfer's mmq_raw_wide_nt_kernel at exactly the same occupancy (2.00 active warps/cyc/sched, 8 warps / 4 schedulers) have a per-scheduler issue rate of 0.42 vs 0.26. Same occupancy, same tensor work (IMMA counts exactly equal), yet a 1.6× issue-rate gap — meaning the gap is not in resource allocation but in what the warps spend their cycles on.

That question can be answered directly with warp-state sampling: the sampler periodically snapshots each resident warp's PC and stall reason, turning "where did the cycles go" into a readable table. r20's session flow thereby rewrote the campaign methodology: do matched-shape attribution first, calibrate the stall structure clearly, then decide what to move. The matching conditions tightened to the same layer's q-proj GEMM, nt=512, id=od=3584, taking the launch points on both sides and comparing them one by one (llama mul_mat_q<12,128,0> grid (48,1,1) + fixup (48,4,1); minfer grid (4,28) = 112 blocks ≈ 2.33 waves). Falsify first, localize second, fix third — that order itself is one of this doc's main outputs.

2. Principle — the GPU mechanism

2.1 First establish how to read the attribution

ncu's PC-sampling (warp-state sampling) records the state of resident warps sampled per SM cycle: if a warp cannot issue because some instruction's dependency is not ready, the sample lands on that instruction's PC, with the stall reason attached. How to read the common reasons:

Stall reasonMeaning
long_scoreboardwaiting on a register produced by a long-latency instruction (global load)
short_scoreboardsame, but the producer is a short-latency instruction (smem, shared values)
waitfixed-latency instructions' (most ALU) waiting slots
barrierwaiting for other warps at bar.sync
not_selectedeligible to issue but the scheduler picked another warp (a good sign — there is slack)
lg_throttleLSU input queue/issue port saturated (request rate overloaded)
math_pipe_throttlecompute pipeline backpressure

The critical reading (this doc's core lesson): the sample lands on the "consumer" instruction — the one waiting on the dependency — not on the "producer". When an LDG is followed by the STS depending on it, the warp stalls on the STS, with reason long_scoreboard. Reading the table as "the store is slow" gets it entirely wrong; the store is only where the load's latency surfaces.

2.2 The matched nt=512 comparison table

Per-issue-active warp ratios; llama taken from launch 0 of its 8 launches (launch 6 verified consistent):

Metricminfer <KDR=4>minfer <KDR=8>llama <12,128>
duration q-proj (µs)632.4609.6263.6
issue /cyc/sched0.16–0.260.200.42
eligible /cyc0.22–0.280.280.64
warps active /cyc2.002.002.00
warp inst / tile (k)499488356
IMMA mma-inst1,605,6321,605,6321,605,632
long_scoreboard6.225.761.15
wait0.630.570.58
barrier0.260.170.19
not_selected0.410.400.50
lg_throttle0.330.600.09
LDS bank conflicts3,211,2643,211,26442

This table rules out four candidate explanations in one stroke:

  • Barrier structure: 4 barriers/256 k on both sides — the "their staging is less synchronized" hypothesis is false at the source level;
  • Tensor work: IMMA 1,605,632 identical across all three;
  • Occupancy: 2.00 warps/cyc the same;
  • Launch shape: the stream-k vs 3-wave difference measured as a wash; llama's fixup +34 µs counted in does not change the conclusion.

The only remaining significant difference is long_scoreboard: 6.22 vs 1.15, 97% of the total named-stall difference — minfer's warps spend ~86% of their resident cycles on global load latency, llama only ~24%.

2.3 Using PC sampling to pin long_scoreboard onto an instruction

The top stall site holds 16.7% of all stalls: the first STS in the A-qs staging loop. The mechanism taken apart:

The old staging loop interleaves "load → store"; after the compiler unrolls 4-deep into batches, each batch = 4 LDG + 4 dependent STS. The warp issues 4 LDGs, and the immediately following STS_0 must wait for LDG_0's registers to return — the A-activation read is an 8-sector scatter pattern (40 B per token row: d(2B) | qs(32B) | ssum(4B), 8 lanes each reading one 4 B word, sector efficiency 62.5%), one round trip ~600 cycles. So every batch pays one full memory latency, and batches serialize against each other.

Compute the per-thread batching: at KD=4 the A-qs side has 16 load words (128·KDR·8 / 256 thr = KDR·4) → 4 batches of depth 4; each batch sleeps ~600 cycles with an in-flight depth of only 4. The aggregate in-flight volume of 8 warps/SM cannot cover a 600-cycle latency, and the issue rate collapses to 0.16–0.26. Each warp's MLP (memory-level parallelism, the number of in-flight memory requests) depth is clamped to 4 by the dependency chain — that is the microscopic composition of long_scoreboard 6.22.

2.4 The fix's shape is uniquely determined by the attribution result

Since what binds is "dependency scheduling" rather than bytes, instruction count, or barriers, the minimal fix is to split the phases: first issue all staging global loads into a register array (one k-tile's worth of deep, independent LDGs), then write smem in one go. Addresses, traffic, instruction counts unchanged item by item — the only change is the dependency schedule: the warp now sleeps once per k-tile, at the first STS (by then it has issued 16 in-flight loads, MLP depth 4×), instead of once per batch. This is the cleanest class of diff in experiment design: any delta can only be attributed to the scheduling itself.

And one expectation that must be stated upfront: this class of fix moves the constraint; it does not eliminate it. Once the latency is hidden, the constraint that ranked second surfaces — measured here it migrated to lg_throttle (LSU queue saturation, 0.33 → 2.38): staging latency and request rate are serially co-bound. r21 will prove this conservation with a "fix the request side only" experiment.

3. Implementation

3.1 Design choices (why this shape and not another)

Why not cp.async / double buffering? The family history tried both: cp.async double buffering (784786d) measured neutral — but that experiment's interpretation needs correcting today: it moved the same bytes and kept the same "move while computing" interleaved structure, measuring the "prefetch depth" variable; the attribution pointed at dependency scheduling. The split-phase plan introduces no new transfer mechanism, adds no smem, adds no barriers — it is the only candidate that maps one-to-one onto the attribution result.

Why a register array? The old code actually had 4-deep unrolling too — the problem is that within each unrolled batch loads and stores interleave, and batches serialize against each other on the first STS. The register array av[KDR*4] splits "issuing loads" and "consuming loads" into two independent loops; the compiler naturally issues the load loop as a whole and consumes it as a whole in the store loop; _Pragma("unroll") guarantees the enumeration is not folded back into the interleaved form.

Why the A side only? The attribution's top site (16.7%) is in the A-qs loop; B-side staging at KDR=4 has the restage-skip guard (two k-tiles share one super-block), so its diluted stall share is an order of magnitude lower. Fix where the attribution points — this also lets the +7.1% gain be booked cleanly to the A side.

Why nt=512 for the matched capture? Both before and after use the same launch point, the same shape comparison (q-proj, id=od=3584), guaranteeing the longsb/duration before/after difference is not an artifact of shape drift — r13's lesson (nt=2630 vs llama nt=512 double distortion) converts directly into this round's experimental discipline.

3.2 Key code

First the layout context of the objects being operated on (current tree src/cuda_kernels.cu wide-kernel smem comment; the two A-side blocks are r20's targets):

//   qa8   [KDR][128] x 32B  chunk qs only (d/ssum in sda_q) ...
//   sda_q [KDR][128] x 8B   (d f16 | ssum i16) packed, uint2-tiling ...

Before — the interleaved LDG→STS chain (git show 5ac8917 deletion side, the RAW_STAGE macro's A section):

/* Old: loads and stores interleave — every 4-deep batch pays one full
 * memory latency at its first STS (PC sampling: that STS alone held
 * 16.7% of all warp stalls). */
for (int x = threadIdx.x; x < MMQ_WBI * KDR * 8; x += blockDim.x) {
    int u = x % 8, r = (x / 8) % MMQ_WBI, kd = x / (8 * MMQ_WBI);
    int tok = i0 + r, c = (kt) * KDR + kd;
    unsigned v = 0;
    if (tok < nt && c < nchunk)
        v = *(const unsigned*)(q8x + ((size_t)tok * nb32 + c) * 40
                               + 4 + u * 4);                     // LDG (scattered 4B words)
    *(unsigned*)(qa8 + ((size_t)kd * MMQ_WBI + r) * 32 + u * 4) = v; // STS (dependent)
}
for (int x = threadIdx.x; x < MMQ_WBI * KDR; x += blockDim.x) {   // d/ssum isomorphic
    int r = x % MMQ_WBI, kd = x / MMQ_WBI;
    ...
    d16 = *(const unsigned short*)src;               // LDG.16
    ss   = (unsigned)(short)*(const int*)(src + 36); // LDG.32
    *(unsigned*)(sda_q + ...) = d16 | (ss << 16);    // STS (dependent)
}

After — split-phase (5ac8917 addition side; the current tree keeps it to this day, r22 only rewrote the store addresses' swizzle):

/* r20: split-phase A staging. The old interleaved LDG->STS chains
 * stalled the warp at the first store of every 4-deep unroll batch
 * (PC-sampled: the leading STS held 16.7% of all warp stalls = one
 * full memory latency per batch, ~4 batches per kt). Issue ALL
 * global loads into registers first - one deep independent LDG batch
 * per warp per kt - then store to smem. Identical addresses, traffic
 * and instruction count; only the dependency schedule changes. */
{
    unsigned av[KDR * 4];          /* qs words: 128*KDR*8 / 256 thr */
    unsigned short dv[KDR / 2];    /* d f16 words: 128*KDR / 256 */
    unsigned sv[KDR / 2];          /* ssum words */
    _Pragma("unroll")
    for (int i = 0; i < KDR * 4; ++i) {                 // phase 1: issue loads only
        const int x = threadIdx.x + i * 256;
        const int u = x & 7, r = (x >> 3) & (MMQ_WBI - 1),
                  kd = x >> 10;    /* x/(8*MMQ_WBI) */
        const int tok = i0 + r, c = (kt) * KDR + kd;
        unsigned v = 0;
        if (tok < nt && c < nchunk)
            v = *(const unsigned*)(q8x
                + ((size_t)tok * nb32 + c) * 40 + 4 + u * 4);
        av[i] = v;                                      // independent batch of depth KDR*4
    }
    /* ... dv/sv isomorphic load loop ... */
    _Pragma("unroll")
    for (int i = 0; i < KDR * 4; ++i) {                 // phase 2: stores only
        const int x = threadIdx.x + i * 256;
        const int u = x & 7, r = (x >> 3) & (MMQ_WBI - 1),
                  kd = x >> 10;
        *(unsigned*)(qa8 + ((size_t)kd * MMQ_WBI + r) * 32 + u * 4)
            = av[i];                                    // consume the already-arrived registers
    }
    /* ... dv|ss merged-word store loop ... */
}

The address enumeration between the two loops is identical item by item (the same linearization x = threadIdx.x + i*256), the stores are still the same 4 B/8 B writes — anything outside the diff (B side, compute section, barrier count) untouched.

3.3 Pitfalls

  • The attribution table's reading trap: PC sampling records the stall on the STS, and the first instinct is "optimize the store" — e.g. widen the store, merge writes. Exactly backwards: the store is the victim; the producer chain (interleaved scheduling + scattered loads) is the pathology. This doc's fix touches not one store instruction's form.
  • r13's old stall table is not reusable: pre-r14 kernel + nt=2630 vs nt=512 launch-point misalignment, doubly distorted. Attribution must be redone on the post-fix kernel, matched shape — an old map is more dangerous than no map.
  • Register pressure is this plan's natural risk: av[16] + dv[2] + sv[2] at KD=4 costs roughly a dozen extra registers; the kernel already runs at 1 block/SM occupancy, and register spills would eat the gain directly. The post-landing ncu showed no spill growth (warp-inst −16% is a net reduction), so the risk was covered by measurement.
  • Hunt down where the "freed stalls" went: after longsb −53%, without looking at lg_throttle it is easy to misread it as "53% of the same problem remains"; in fact the constraint migrated — the next round's lever (r21's merged uint4) was derived from exactly that reading.

4. Verification

  • parity green KD=4 + KD=8 — split-phase only reorders independent smem writes; the end state is byte-identical; the parity gate confirms the "schedule change" did not quietly become a "value change".
  • Greedy token stream byte-identical — the end-to-end behavior gate, defending against a kernel-level diff amplifying over multi-step inference.
  • Suite 166/0/3 — the cross-shape/cross-model regression matrix, defending the long-tail shapes outside the matched shape (short prompts, nt boundaries) against enumeration-difference breakage.
  • 6/6 interleaved A/B reproducible — +7.1%/+3.5% must hold repeatedly on a box drifting ±9% that day to be bookable (the same measurement discipline as r18/r19).
  • Matched nt=512 before/after ncu (longsb 6.22→2.92, duration 632.4→555 µs, warp-inst −16%, lg_throttle 0.33→2.38) — the mechanism gate: the named stall must fall as predicted, and the freed share must have a findable destination; a mismatch in direction or magnitude means the attribution was wrong.
  • The exclusion list (barrier density, IMMA count, occupancy, launch shape, memory bytes all measured flat) — the attribution's control group: without these "exclusions", "long_scoreboard is the carrier" is only a correlation.

5. Results

Wall clock (7B whole-prefill, same-session interleaved, 6/6 reproducible):

ConfigbeforeafterΔ
wide KD=41230.41317.7+7.1%
wide KD=81275.81319.9+3.5%

Kernel level (matched nt=512 q-proj): duration 632.4 → 555 µs (−12.2%); warp-inst −16%; long_scoreboard 6.22 → 2.92 (−53%; the ~86% resident share on global-load latency falls sharply); lg_throttle 0.33 → 2.38 (the constraint migrated to the request-rate side). Against llama's same-shape 263.6 µs (+34 µs fixup): the per-kernel gap converges from 2.4× to ~2.0×.

The verdict on the bar: the session bar ≥1350 was not reached (1317.7 < 1350), but this step landed. Unlike r18/r19's "miss the bar, revert", r20's evidence structure is complete:

  1. The named stall was cut in half as predicted;
  2. 6/6 reproducible;
  3. The causal chain between mechanism and gain is closed;
  4. The residual constraint (lg_throttle/request rate) was already picked up by the next experiment (r21).

The bar is a heuristic, not a mechanism — when the mechanism evidence is complete, the bar yields. This is also the prelude to the campaign's later "+1.5% relative bar" replacing the absolute bar (formally calibrated in r24).

Follow-ups directly connected to this doc: the lg_throttle reading directly spawned r21 (merged block-linear A staging, targeting sector efficiency and request count — reverted, proving stall-mass conservation); the split-phase structure itself became the wide kernel's permanent shape; later r22 (swizzle) and r34 (prepass-ization) are both built on it.

6. Lessons

  1. Attribute first, shoot second: the r17/r18/r19 triple-null were all guesses on uncalibrated maps; one matched-shape PC-sampling session pinned 97% of the named-stall difference onto a single site, and the fix hit in one shot for +7.1%. An attribution session costs far less than a round of blind trials.
  2. PC-sampling names the consumer instruction: the sites in a stall table answer "who is waiting", never "who is slow" — reading the LDG's latency as the STS's fault sends you to optimize the wrong instruction.
  3. A schedule-only diff is the cleanest experiment: when addresses, traffic, and instruction counts are unchanged item by item, any delta can only come from the dependency schedule — variable isolation needs no statistics, only construction.
  4. A fix moves the constraint without eliminating it: the longsb → lg_throttle migration shows staging latency and request rate are serially co-bound; before celebrating, bring back the "next binder" reading (r21 immediately proved the other side of this conservation).

← 24-r19-weight-l2-residency · Index · 26-r21-coalesced-block-linear-a →

26 · r21 — Coalesced block-linear A staging: stall-mass conservation (REVERTED)

Result: A-side fetch sector count 79.35 → 56.67 M (−28.6%), lg_throttle 1.42 → 0.14 (−90%) — but wait +110%, short_scoreboard +126%, warp instruction count +11.5%, wall clock −2.3% / −0.7% (negative), reverted. Commit: 3c009cc (record/docs commit). Date: 2026-09-03.

Code provenance note: r21's code changes were fully reverted and have no directly reachable code commit — 3c009cc contains only the 64-line record for docs/CUDA_OPTIMIZATION.md. Per the forensics protocol, this doc's "before" code comes from the current tree (the split-phase staging shape r20 landed; the store addresses on today's tree carry r22's XOR swizzle beyond the r21-era baseline, noted inline), and "after" is given as record narration + addressing arithmetic — no code is fabricated.

1. Background — where things stood

This is the middle of the q4_K MMQ campaign (Era C). r20 (the previous doc) had just taken the campaign's first clear scheduling-level victory with split-phase A staging: tearing apart the serial chain where "the first STS of every 4-deep unrolled batch stalls on the global load latency", issuing all global loads into the register array first and then writing smem in one go — addresses, traffic, instruction counts all unchanged, only the dependency schedule changed. KD=4 1230.4 → 1317.7 (+7.1%), KD=8 1275.8 → 1319.9 (+3.5%), 6/6 reproducible, and it landed even without reaching the ≥1350 bar.

But r20's ncu evidence also handed over two "not dead yet" leads:

  1. The A fetch is still 8-sector scattered. The A-side source layout is per-token pad40 blocks (d(2B) | qs(32B) | ssum(4B), pitch 40B), and staging enumerates "one 4B word per thread"; one warp step covers 4 non-adjacent 32B segments; measured sector efficiency only 62.5% (the top stall site holds 16.7% of all stalls, each 4-deep batch eating about one ~600-cycle memory latency).
  2. The freed stalls immediately squeezed in somewhere else: long_scoreboard 6.22 → 2.92 (−53%), yet lg_throttle surged from 0.33 to 2.38. r20's lesson's own words: "staging latency and request count are serially co-bound" — fix one, and the other becomes the new binder.

r21's hypothesis is therefore very direct: if the binder is request quality (sector efficiency) and request count (L2TEX queue depth), then change the A fetch to coalesced block-linear access — the warp steps through contiguous per-token regions, moving one granule per 16B uint4, with d/ssum folded into the same pass along the way. In theory sector efficiency → 100%, load instruction count ÷4, and lg_throttle's queue pressure should fall with it. This is a frontal assault on item (1) of r20's leftover list.

The overall coordinates of the moment: the wide-tile raw kernel at ~1320 tok/s (KD=8), llama-bench at the same anchor 3401, bar ≥1350. Every percentage point of the campaign had to be attribution-driven — r21 is the decisive experiment for "is r20's remaining gap really a sector/request problem".

2. Principle — the GPU mechanism

2.1 The source layout's geometry ledger

The q8_0-quantized activation plane q8x is stored per token, one pad40 block per token per 32-k chunk:

token t, chunk c:  [d: 2B][qs: 32B @+4][ssum: 4B @+36]   block pitch = 40B

The r20-shape staging enumeration (see the §3.1 code): x = tid + i*256, u = x & 7 (the u-th 4B word within the block), r = (x>>3) & 127 (token row within the tile). That is, each thread fetches 4B; 32 threads in a warp = one (r, u=0..7) 32B qs row × 4 adjacent r's — the global address is (tok*nb32 + c)*40 + 4 + u*4, and rows of adjacent r's are 40B apart. One warp step issues 4 32B transactions landing on 4 different cache lines, and each 40B block's qs(4..35) + ssum(36..39) straddles two 32B sectors, with the d field's header in between — the combined measured efficiency is 62.5%.

2.2 r21's coalescing plan and theoretical gains

r21 flips the enumeration: the warp steps through contiguous per-token regions, each lane fetching one 16B uint4 at a time. Each (token, chunk)'s qs plane is 32B = two 16B granules; one warp step = 32 lanes × 16B = 512B contiguously covering 4 blocks' qs; d(2B) and ssum(4B) are folded into smem in the same store pass. Three theoretical gains:

  • Sector efficiency → ~100%: the 16B granularity divides the qs plane exactly, no more half-empty sectors from the 40B pitch;
  • Load instruction count ÷4: one uint4 covers 4 4B words, so the instruction count for the same bytes drops linearly (lg_throttle is essentially the L2TEX queue full — queue slots are counted per request, so requests drop 4×);
  • d/ssum folding: in the old shape d/ssum was a separate round of scattered small loads; after coalescing that second round disappears.

On paper this is a change where "every line should win". r21's verdict value lies precisely here: it won every line, yet lost the wall clock — that is exactly the mechanism this doc leaves behind.

2.3 Stall-mass conservation

A warp issues at most one instruction per cycle; the kernel's wall clock is decided by "whichever resource stalls the warp first". Moving stalls from long_scoreboard (data latency) to lg_throttle (queue full) was r20; r21 wanted to move stalls away from lg_throttle, premised on the freed issue slots not being immediately refilled by the next binder. Coalesced staging saves load instructions and requests, but to compute where each 16B granule goes, the enumeration itself pays new address ALU and branches — if these added instructions plus the reshuffled, shallower load batches (deep LDG batches chopped up) refill the issue slots, the wall clock does not move or even regresses. r21's measurement was the latter ending.

3. Implementation

3.1 before: r20's split-phase staging (current tree, real code)

Below is the wide-tile kernel (mmq_raw family) A-side staging as it stands today. The r20-introduced "all LDGs into registers first, then unified STS" structure is preserved as-is (the three sections av[KDR*4] / dv[] / sv[]); the only difference from the r21-era baseline is that qa8's store addresses got r22's XOR swizzle (noted in the comment) — r21 at the time faced the un-swizzled flat address qa8 + R*32 + u*4:

/* src/cuda_kernels.cu — wide-tile raw MMQ kernel, RAW_STAGE macro A side (today's tree) */
{
    unsigned av[KDR * 4];          /* qs words: 128*KDR*8 / 256 thr */
    unsigned short dv[KDR / 2];    /* d f16 words: 128*KDR / 256 */
    unsigned sv[KDR / 2];          /* ssum words */
    _Pragma("unroll")
    for (int i = 0; i < KDR * 4; ++i) {              /* phase 1: all qs loads */
        const int x = threadIdx.x + i * 256;
        const int u = x & 7, r = (x >> 3) & (MMQ_WBI - 1),
                  kd = x >> 10;    /* x/(8*MMQ_WBI), 8*128 = 1024 */
        const int tok = i0 + r, c = (kt) * KDR + kd;
        unsigned v = 0;
        if (tok < nt && c < nchunk)
            v = *(const unsigned*)(q8x
                + ((size_t)tok * nb32 + c) * 40 + 4 + u * 4);   /* 4B scatter */
        av[i] = v;                                     /* accumulate into registers first */
    }
    /* …d/ssum likewise loaded into dv[]/sv[] first (the other phase-1 batch)… */
    _Pragma("unroll")
    for (int i = 0; i < KDR * 4; ++i) {              /* phase 2: unified STS */
        const int x = threadIdx.x + i * 256;
        const int u = x & 7, r = (x >> 3) & (MMQ_WBI - 1),
                  kd = x >> 10;
        const int R = kd * MMQ_WBI + r;
        *(unsigned*)(qa8 + (size_t)(R & ~3) * 32
            + (size_t)(((((R & 3) << 1) + (u >> 2))      /* r22 swizzle */
                        ^ ((R >> 2) & 7)) << 4)
            + (size_t)(u & 3) * 4) = av[i];
    }
    /* …sda_q's d|ssum packed store… */
}

Key point: addresses, traffic, and instruction counts are word-for-word identical to r20's landing (except the store-side swizzle); the dependency schedule is "one batch of deep LDG → one batch of STS".

3.2 after: r21's block-linear enumeration (record narration)

r21's shape (reverted, no code survives; its structure restored per the record):

  • Enumeration granularity 4B → 16B: each lane handles one 16B uint4, the warp step covering contiguous per-token qs regions (one warp step = 512B = 4 blocks' qs planes);
  • d/ssum folded into the same pass: when the qs loads land and store, d and ssum are written along the way, eliminating the separate d/ssum scattered small round;
  • Split-phase kept: still the "load first, store second" phase structure, with only the enumeration and addresses changed.

Three parity fixes (quoted as-is from the record; all are the index accounts easiest to get wrong when rewriting an enumeration):

  1. Each lane's slice count must be recomputed at the new granularity (at 16B granularity, 256 threads × K slices must exactly tile the 128×KDR×32B smem mirror);
  2. The granule index is derived from lane, not from the old 4B word's x decomposition;
  3. The global token index must add back i0 (row number within the tile ≠ global token number) — this bug is exposed only by the nt=256 sweep shape (at small shapes i0=0 happens to be harmless), a textbook case of "single-shape testing missing an addressing bug".

Alongside came a byte-exact CPU simulator: on the CPU, replay both enumerations' smem write sequences byte by byte, confirming the old and new shapes produce exactly the same smem mirror — verifying "did the layout rewrite change the data" off the GPU.

3.3 Pitfalls

  • The i0 loss only surfaced at nt=256: the single-point nt (the main measured shapes like 2630/2659) happens to start from tile 0, so i0=0 and the lost global token base never fires; only the sweep shape reveals it. The lesson, distilled: addressing-class changes must pass a shape sweep.
  • The index accounting of enumeration rewrites: once granularity changes, the "thread → (kd, r, u)" three-dimensional decomposition must be fully re-derived; two of the three parity fixes were this class of decomposition error.
  • ptxas will not preserve batch depth for you: the old 4-deep unroll gave each warp one coherent deep LDG batch; the coalesced enumeration has wider single loads, but the reshuffled batch structure and boundary checks made the issue shape worse (see §5's wait/short_scoreboard rebound).

4. Verification

  • CPU byte-exact simulator: defends against "the layout rewrite quietly changed the smem mirror" — old and new enumerations compared byte by byte, confirmed equivalent.
  • Shape sweep parity (nt=256 included): defends against i0/slice-count-class addressing bugs that fire only at specific shapes (this doc's fix #3 is what it caught).
  • ncu counter pair (sectors / lg_throttle / wait / short_scoreboard / warp-inst): defends against "looking at only one wall-clock number" — all of this doc's mechanism evidence comes from this pair.
  • Same-window A/B interleaved measurement: defends against cross-session drift; the wall clock −2.3%/−0.7% is the median difference of adjacent same-binary pairs.

5. Results

Metricbefore (r20 shape)r21 coalesced shapeΔ
l1tex sectors (A-side fetch)79.35 M56.67 M−28.6%
lg_throttle (stall/issue-active)1.420.14−90%
mio / math throttlepresentgone—
wait (barrier-class stalls)——+110%
short_scoreboard (smem dependency)——+126%
total warp instructions——+11.5%
kernel duration——+10.9%
wall clock (same-window A/B)——−2.3% / −0.7%

Veto mechanism: all three "should-win" accounts cashed in (sectors, request count, and queue stalls all improved massively), yet the wall clock is negative. The freed issue slots did not become useful work; they were refilled by the reshuffled, shallower load batches (wait/short_scoreboard rebound) and the enumeration's own added instructions (warp-inst +11.5%) — stall-mass conservation: sector efficiency and request counts are not the binder; the warp instruction stream is (r25 later corrected "instruction count" to "issue/occupancy", but the direction agrees: the MIO/byte side is not this phase's ceiling). The code was fully reverted.

Under what future conditions a retry is worthwhile: only when the instruction stream already matches the opponent's and the issue slots genuinely have slack can "fewer, wider requests" convert into wall clock. The campaign actually took another road — r34 moved the entire A-side layout transform out of the kernel (quantize-transpose prepass), so the staging problem was eliminated rather than optimized; r51/r52 went further and had the producers emit the already-quantized plane directly. Looking back at r21, it tried to optimize within the premise of "per-token 40B scattered source" — and the premise itself was later changed.

6. Lessons

  1. Stall-mass conservation: fixing the current binder only makes room for the next one — predicting (and measuring) "who picks up the freed slots" is what completes an attribution.
  2. Counters winning ≠ the wall clock winning: sector/request/queue-class metrics improving while the wall clock regresses means they are not on the critical path; measure the next stall before writing the next lever.
  3. Deep load batches are an asset: when re-enumerating, the old shape's implicit "one deep independent LDG batch per warp" gets torn apart unintentionally — the issue shape is an explicit design object, not a byproduct.
  4. Addressing rewrites must pass a shape sweep: bugs like a lost i0 base are symptom-free on main measured shapes where i0=0; the byte-exact simulator and the shape sweep are two gates, both indispensable.

← 25-r20-split-phase-a-staging · Index · 27-r22-qa8-xor-swizzle →

27 · r22 — qa8 XOR swizzle (LANDED); d/ssum fold reverted separately

Result: shared-load conflict counter op_ld 16,859,136 → 0; KD=8 1311.8 → 1329.6 tok/s (+1.4%, 3/3 pairs positive); KD=4 −0.5% (noise band). The d/ssum stream fold attempted in the same round: both enumeration variants −19–20%, reverted. Commit: 5b40058 ("cuda(mmqr): r22 — qa8 XOR swizzle kills the A-side LDSM 2-way (op_ld 16.86M -> 0, KD=8 +1.4%); d/ssum stream fold NEGATIVE (reverted)"). Date: 2026-09-03.

1. Background — where things stood

r21 (the previous doc) had just delivered an expensive verdict: after changing the A fetch to coalesced block-linear, sectors −28.6% and lg_throttle −90%, yet the wall clock instead −2.3%/−0.7% — stall-mass conservation; neither the byte side nor the queue side is the binder. The natural next question: on the real binder, the "instruction stream/issue", is there a lever that actually deletes work (rather than reshuffling it)?

r22's Step 0 was pure attribution: on the q-proj, nt=2630, KD=4 shape, break the shared-access conflict counters apart. GB10 has no dedicated ldmatrix conflict metric (ldmatrix's bank conflicts fold into the op_ld/shared-load wavefront counts), but the ledger balances:

  • op_ld conflicts 16,859,136 = the A-side qa8 ldmatrix's 2-way conflicts — 4.74 M ldmatrix.x4 × 4 phases each, each 2-way phase paying one extra wavefront, summing exactly onto op_ld;
  • op_st 6.5 M is the qb8 staging store + sds conflicts — untouched this round.

Mechanically this comes from the qa8 smem plane's 32B row stride (§2 expands). The key judgment: conflicts are pure wavefront waste — the same bytes, the same instruction count, each ldmatrix paying double the MIO wavefronts; fixing it is "deleting work" rather than "moving work", fundamentally different from r21's reshuffle. And r21's lesson already pointed the way: the fix must not add address ALU per load — which determined r22's final shape as "each thread precomputes all 8 offsets once".

2. Principle — the GPU mechanism

2.1 The mechanical definition of a bank conflict

Shared memory is split into 32 banks, each 4B wide; an address lands in bank (addr/4) mod 32. In one memory instruction, if the 32 lanes' addresses hit mutually distinct banks, one wavefront completes; two lanes hitting the same bank is a 2-way conflict, and that phase splits into 2 serial wavefronts — time ×2, bytes unchanged. ldmatrix is a warp-level instruction: 32 lanes each supply a row address, .x4 takes one 8×8 b16 matrix per phase across 4 phases, conflicts counted per phase.

2.2 Why the qa8 plane is 2-way

The A-side smem plane is qa8[KDR][128 tokens] × 32B (the qs plane of each token per 32-k chunk). ldmatrix takes A fragments per m16n8k32's standard distribution: lanes 0–7 supply matrix 0's 8 rows, lanes 8–15 supply matrix 1… rows taken by adjacent phases are 4 rows apart. And the row stride is 32B: 4 rows = 128B = 32 banks × 4B, wrapping exactly back to bank phase zero. So row addresses "4 rows apart" all collide on the same bank phase — 2-way per phase, one extra wavefront per phase per ldmatrix.x4, cumulatively 16.86 M extra wavefronts.

No swizzle: address delta between row r and row r+4 = 4 × 32B = 128B ≡ 0 (mod 32 banks×4B)
           → the 8 row addresses within an ldmatrix phase collide pairwise → 2-way

With swizzle: the 8 16B granules within a 128B super-row are placed by XOR permutation
           gr(row, h) = ((row&3)*2 + h) ^ ((row>>2)&7)
           4 rows apart ⇔ (row>>2)&7 differs (bit 2 flips) ⇒ 8 rows → 8 distinct phases

A concrete comparison (the 4 rows of super-row group 0 plus group 1's first row; granule number = in-row half ×2 + in-group row ×1, then the XOR mask): without swizzle, granule 0 of rows 0/4/8/12 all sit at byte offsets {0, 128, 256, 384} — all ≡ 0 (mod 128B), same bank phase; with swizzle they are scattered by masks 0/1/2/3 into 4 different granule slots, and the phase's 8 lanes each land on a distinct phase.

2.4 Why these 16.86 M wavefronts are worth fixing

r20/r21's stall chain already showed: this kernel's issue slots are a scarce resource, and every MIO-side wavefront has to compete with LDG/STS issue. One 2-way-conflicted ldmatrix phase = the same bytes crossing the LSU pipe twice — 16.86 M extra wavefronts are pure issue-slot rent. It also satisfies two "worth fixing" criteria: (a) an exact zero is reachable (the conflict counter can be verified to 0, not "reduced a bit"); (b) the fix can amortize the ALU cost outside the loop (r21's lesson absorbed head-on).

2.3 The XOR swizzle: putting 8 granules into 8 phases

r22 makes every 4 consecutive 32B rows form a 128B super-row, with the 8 16B granules inside the super-row placed by XOR permutation:

gr(row, h) = ((row&3)*2 + h) ^ ((row>>2)&7)
  row   : the low 2 bits of the tile-global row number R (which row within the super-row)
  h     : which 16B granule within the row (0/1)
  row>>2: the super-row group number, which sets the XOR mask

Two row numbers 4 apart always differ in (row>>2)&7 (bit 2 flips) → their granule placements are scrambled by different masks → ldmatrix's 8 row addresses per phase scatter into 8 distinct bank phases → conflicts to zero, with zero smem growth (only a permutation inside the 128B). The store side and the ldmatrix side use the same mapping: whichever slot the data lands in is the slot it is read from — the two ends just have to agree.

3. Implementation

3.1 Design choices: verify standalone first, then integrate "computed once"

Two deliberate orderings:

  1. Standalone first: first write an independent test kernel to verify the swizzle mapping — 131,072 LDSM instructions' conflicts fell from 524,288 to 0, and the A fragments' lane→(row, word) distribution matches the old layout exactly (distribution unchanged is what makes bitwise parity possible). Only after the mechanism measured clean did we touch the real kernel.
  2. Offset precompute: r21's lesson (added address ALU eats the gain) directly determined the shape — the 8 ldmatrix addresses are lane-invariant per thread (they only shift with the kd base), so G[0..7] is computed once outside the loop; inside the loop each ldmatrix keeps only one IADD, fewer than the baseline's ~3; registers actually fell below baseline (129/145 vs 141/149, zero spill).

3.2 Key code

Store side (staging writes place bytes by the same mapping; today's tree's form, with the r20 split-phase structure kept and only qa8's store address changed — the comment is r22's on-site record):

/* src/cuda_kernels.cu — wide-tile RAW_STAGE macro, qa8 store (today's tree) */
/* r22: r20's split-phase A staging is kept verbatim; only the qa8
 * store address gained the XOR swizzle. The d/ssum stream fold
 * (single 9-word-per-chunk pass) was tried and REVERTED: the old
 * scattered d/ssum loads are L1 hits (the qs pass of the same
 * chunks has the lines resident), so folding saves little sector
 * traffic while the flat enumeration costs ALU + branchy batches
 * (-19% wall, see docs r22). */
const int R = kd * MMQ_WBI + r;
*(unsigned*)(qa8 + (size_t)(R & ~3) * 32                 /* 4-row group base */
    + (size_t)(((((R & 3) << 1) + (u >> 2))              /* granule number within the group */
                ^ ((R >> 2) & 7)) << 4)                  /* XOR permutation × 16B */
    + (size_t)(u & 3) * 4) = av[i];                      /* 4B word within the group */

Load side (each thread precomputes G[0..7] once; one IADD per ldmatrix inside the loop):

/* src/cuda_kernels.cu — wide-tile kernel, before the main loop (today's tree) */
// r22: precomputed swizzled A-frag byte offsets. The 8 ldmatrix
// addresses per chunk are lane-invariant except for the kd base:
// addr = qat + G[g], G[g] = g*512 + (lane&12)*32
//      + (((lane&3)*2 + (lane>>4&1)) ^ (lane>>2&3) ^ ((g&1)*4)) << 4.
// One IADD per ldmatrix (below baseline's 3), and the granule XOR
// gives every ldmatrix phase 8 distinct bank phases (the 32B row
// stride is 2-way conflicted).
const unsigned l12m = (unsigned)(lane & 12) * 32;
const unsigned grc = (unsigned)(((lane & 3) << 1) + ((lane >> 4) & 1)
                         ^ ((lane >> 2) & 3)) << 4;
unsigned G[8];
#pragma unroll
for (int g = 0; g < 8; g++)
    G[g] = (unsigned)g * 512 + l12m + ((g & 1) ? (grc ^ 64u) : grc);
/* consumption point: one ldmatrix.x4 per 16-token group, address = qat + G[g] */
for (int g = 0; g < 8; g++) {
    const uint8_t* p = qat + G[g];          /* the only per-instruction cost: one IADD */
    unsigned r0_, r1_, r2_, r3_;
    asm volatile(
        "ldmatrix.sync.aligned.m8n8.x4.shared.b16 "
        "{%0,%1,%2,%3}, [%4];\n"
        : "=r"(r0_), "=r"(r1_), "=r"(r2_), "=r"(r3_)
        : "r"((unsigned)__cvta_generic_to_shared(p)));
    a[g][0] = (int)r0_; a[g][1] = (int)r1_;
    a[g][2] = (int)r2_; a[g][3] = (int)r3_;
}

The read-side formula and the write-side formula are two expansions of the same mapping, item by item:

Read-side termWrite-side counterpartMeaning
g * 512— (carried by the qat base via kd)base of the g-th 16-token A-frag group: 16 rows × 32B = 512B per group
(lane & 12) * 32(R & ~3) * 32byte base of the 4-row subgroup the lane sits in (row stride 32B)
(lane & 3) * 2(R & 3) << 1in-group row number (low 2 bits) × 2: selects one of the row's two 16B granules
(lane >> 4) & 1u >> 2the row's 16B half h (byte 0 vs byte 16)
(lane >> 2) & 3 | (g & 1) * 4(R >> 2) & 7the XOR mask: the in-group row number's high 2 bits ∥ group parity (carry bit)

Two key invariants: (1) G[g] is a loop invariant per thread — once the lane is fixed, its (row, half) in the m16n8k32 A-fragment distribution is fixed, and the 8 groups are only base shifts, so the whole table is computed once before the main loop; (2) the kd base does not enter G[g] — the consumption point carries kd via qat = qa8 + kd * MMQ_WBI * 32, and G[g] describes only the in-group geometry; that is the precise meaning of "the 8 ldmatrix addresses are lane-invariant, shifting only with the kd base".

3.3 Pitfalls: the d/ssum fold is r21's echo

In the same round r22 tried Lever 2: folding d/ssum's separate scattered small round into the qs's big enumeration ("one pass sweeps a chunk's 9 words"). Both enumeration variants were −19–20% wall clock, three mechanisms stacked:

  1. What is saved are L1-hit duplicate requests: the so-called "scattered" d/ssum loads land on the same cache line the qs big pass just fetched — an L1 hit; the sector traffic merging saves was nearly free to begin with;
  2. The flat enumeration's cost: mod-9 address ALU + branchy store batches, tearing apart r20's carefully kept "one deep LDG batch per warp";
  3. Register blowout: ptxas hit 255 regs + 112 B spill in the folded form.

This is r21's "stall-mass conservation" replayed at micro scale: a request-count "saving" paid for in ALU and batch depth always loses. The fold was reverted, the swizzle kept — one positive and one negative in the same commit, booked separately.

4. Verification

  • Standalone conflict test: 131,072 LDSM instructions' conflicts 524,288 → 0, defending against "mapping derived wrong, swizzle wasted";
  • A-fragment distribution consistency (standalone cross-check): defends against "the swizzle changed the lane→data mapping" — distribution unchanged is the precondition for bitwise parity;
  • Parity (KD=4 / KD=8 all green) + greedy-32 token identity (vs the default binary, only the timing lines differ): defends against values being polluted by the layout change;
  • ncu pair (op_ld → 0, op_st unchanged): proves the fixed counter is exactly the attributed one and the conflicts were not pushed onto the store side;
  • smem size check (73,728 / 98,304 B unchanged): defends against the "zero growth" promise being quietly broken by alignment;
  • 3× same-window interleaved A/B (vs the pre-change binary, 3/3 pairs positive): defends against cross-session drift.

5. Results

MetricbeforeafterΔ
op_ld (shared-load conflict wavefronts)16,859,1360zeroed
op_st (qb8/sds store conflicts)6.5 M6.5 Muntouched (not this round's target)
registers / spill141 / 149129 / 145below baseline
smem73,728 / 98,304 Bsamezero growth
wall clock KD=8 (default path)1311.81329.6+1.4% (3/3)
wall clock KD=4——−0.5% (noise band)

LANDED (default path positive + mechanism closed loop: the attributed counter exactly zeroed). The combined ≥1350 bar was not reached — booked under the stop conditions as a "split result": swizzle kept, fold reverted.

In campaign coordinates, this +1.4% is the close of the r20/r21/r22 "staging attribution series" — r20 proved scheduling can win (+7.1%), r21 proved the byte/request side is not the binder (reshuffling actually lost), r22 proved the MIO-conflict side can be zeroed and cashed into wall clock. Together the three calibrate "the space left to squeeze on the A side" down to the instruction stream itself, directly pushing r24/r25 toward scheduling structure (negative) and the SASS census (the issue/occupancy verdict), with r28's occupancy rewrite then taking over. A +1.4% lever's campaign value lies mainly in what question it closed, not in the number itself.

The d/ssum fold's veto mechanism: what it saves are L1-hit duplicate small requests (nearly free), and the price is mod-9 ALU, branchy batches, torn-up deep LDG batches, and register blowout (255 regs + 112 B spill). Retry conditions: only worth another look when d/ssum becomes a true main-memory scatter (L1 no longer hits) and the enumeration adds no per-word ALU; the later r34 (quantize-transpose prepass) changed the A-side layout at the root, and this line closed naturally.

6. Lessons

  1. Generate addresses once per thread, not once per load: a lane-invariant offset table (G[g]) compresses the swizzle's ALU cost to one IADD per ldmatrix — the key for a "reshuffle"-class change to gain rather than lose.
  2. Prefer levers with an exact zero reachable: conflict-class counters terminate at exactly 0 (op_ld → 0 is verifiable), a far cleaner evidence chain than reshuffle-class levers' "reduced a bit".
  3. Attribute first; make the ledger balance to death: 16.86 M conflicts = 4.74 M LDSM.x4 × 4 phases × 2-way extra wavefronts — make the counter equal the mechanism before touching anything.
  4. Book positive and negative results separately within one commit: r22 = the LANDED swizzle + the REVERTED fold; reporting them as one "round" misleads both reuse and revert.

← 26-r21-coalesced-block-linear-a · Index · 28-r23-f16-wall-decomposition →

28 · r23 — f16-path whole-graph wall decomposition + FA_TKV occupancy raise (MEAS-ONLY + REVERTED)

Result: default f16 path (2659-token prefill, quiet window 2285 tok/s) whole-graph decomposition: GEMM 74% (gate+up 37.8% @ mem SOL 85% / down 27.5% / q+o 7.5%), FA 7.2%, convert 6.0%, swiglu 5.6% — llama's split gives the byte-width gap as 59.5 vs 39.5 µs/GMAC. The FA_TKV 64→32 raise: occupancy 16.7 → 32.68%, kernel −6.7%, but wall clock −0.3% < +3% bar, reverted. Commit: e8c348d (docs/record commit; the FA_TKV raise's code was reverted, no reachable code commit — this doc's FA code comes from the current tree, i.e. the shape after r48's FAP2 landed, noted inline). Date: 2026-09-03.

1. Background — where things stood

r21/r22 had probed the MMQ wide kernel's A side to the bottom in two consecutive rounds: reshuffling lost (stall-mass conservation), deleting conflicts won (+1.4%), KD=8 stuck at ~1330 tok/s. But one increasingly awkward fact: MMQ is still opt-in to this day (MINFER_MMQ=1) — what users run by default is the f16 path (8p's resident f16 weight cache + dequant-in-GEMM), ~2285–2370 tok/s (quiet window 2285), vs llama-bench at the same anchor 1.43×.

The campaign had invested eleven rounds (r12–r22) in the MMQ kernel, yet "how much can MMQ actually win" had always been an extrapolation rather than a measured wall-clock decomposition. r23 decided to pause adding levers and first do two things:

  1. A complete whole-graph wall decomposition of the default f16 path — nsys per-launch bucketing + ncu SOL cross-checks, charging every millisecond to an arithmetic class. This answers "how much wall clock is the MMQ campaign's premise (the byte-width gap) worth", and also "what other big non-MMQ items remain on the f16 path".
  2. Raise FA's occupancy in passing. In the decomposition FA is 7.2% with occupancy only 16.7% (the FA_TKV=64 set in the 8n era pushes smem past a single SM's one-block headroom). A one-line define, FA_TKV 64→32, halves the KV tile's smem and doubles the block count — a cheap pilot testing whether "occupancy for wall clock" holds on FA.

This round's nature is measurement first: the decomposition itself (MEAS-ONLY) is the main deliverable, and the FA_TKV raise is a piggybacked experiment.

2. Principle — the GPU mechanism

2.1 Arithmeticizing the campaign premise

llama's MMQ and our f16 path do the same MACs; the difference is how many bytes flow per weight:

f16 path:  B side 16 bit per weight (f16)
q4_K MMQ:  B side ~4.5 bit per weight (raw nibble stream + per-super-block scale/dmin)
           + A side 8 bit per token per k (q8) — same class on both sides

Measured split (quiet window, same anchor): llama's prefill GEMMs total 597 ms (39.5 µs/GMAC); our f16 path GEMMs 59.5 µs/GMAC + 73 ms of f32→f16 convert pass (llama pays ~0, its quantization happens in-kernel). 59.5 / 39.5 ≈ 1.5×/GMAC — not the 16/4.5 = 3.5× theoretical extreme, because MMQ also pays nibble expansion, scale rescaling, and int8 mma's extra supporting instructions; but at gate+up's 85% mem SOL, the B-stream byte width is the source of the 1.5×. This is the entire MMQ campaign's founding premise, pinned to wall-clock numbers for the first time. Corollary: if the MMQ GEMMs reach llama's GEMM level, the f16 default path should reach ~2900 tok/s — the target of every MMQ round after r23.

2.2 The gains and costs of halving FA_TKV

8n's FA kernel caches FA_TKV × hd K/V tiles per block; the smem bulk is KV staging (the precise figure from r46's later audit: 69.38 KB total → 1 block/SM → occupancy 16.64%, ~2 warps/scheduler, nowhere for latency to hide). TKV 64→32 halves the KV-side bytes (~43.8 KB) → 2 blocks/SM → occupancy 32.68%.

But the other side of the gain is the per-tile fixed cost paid twice as often: the total KV column count is unchanged, so halving the tile means the k loop iterates ×2, each iteration paying the fixed overhead of staging + barriers + the softmax rescale chain. What occupancy buys is latency hiding; what the fixed cost eats is the actual gain — r23's result (kernel −6.7%, wall −0.3%) is the record of these two nearly cancelling.

2.3 The hidden mine of shrinking the tile: lane masks

The wmma accumulator fragments' lane→(row, col) mapping is derived from the tile geometry. After TKV is halved, each warp's accumulator fragments go from 4 to 2, and lanes 16–31's in-tile column numbers no longer coincide with the old geometry — the validity mask must contain both the local bound (c0/c1 < FA_TKV) and the global bound (kt + c0 < kv_end); missing either, the tail tile feeds garbage scores from out-of-range columns into the softmax. r23's integration was caught mid-way by the verification system with exactly one such real bug (the record's own words: "TKV=32 lanes 16–31 must be masked by c0/c1 < FA_TKV, not just kt+c0 < kv_end").

3. Implementation

3.1 The decomposition method: per-launch bucketing + SOL cross-checks

  • Full nsys trace: the 2659-token prefill's default f16 path, per-launch timing bucketed by arithmetic class (gate/up, down, q/o, k/v, FA, convert, swiglu, add/rms/rope, host gap);
  • ncu SOL: spot-check a representative kernel per class against its throughput ceiling (mem SOL, SM busy), defending against the misreading "nsys bucketed correctly, but that kernel itself runs unsaturated";
  • Co-tenant pollution cleaning: outlier launches caused by co-tenant load within the window are removed; all conclusions re-verified on the quiet window (2285 tok/s reference).

Graph-structure discovery (an unexpected harvest before the decomposition): the prefill is 27 full layers + one nt=1 TAIL — q6_K's lm_head (1.45 TMAC = 9.6% of the whole model's MACs) runs only once in the tail, already "tail-priced"; the graph contains no [nt, vocab] logits GEMM in the wall clock. This overturns the intuition "lm_head is a prefill bulk item" and means prefill optimization only needs to watch the 27-layer body.

3.2 The FA_TKV raise: one define and one mask

The raise itself is #define FA_TKV 64 → 32 (all downstream sizes symbolic, no other changes); the real work was the mask fix (§2.3). The raise's code was reverted and survives in no commit; today's tree's FA kernel is already the post-r48-FAP2 shape, but the "local + global double predicate" mask structure survives verbatim and can serve as a living specimen of this fix class:

/* src/cuda_kernels.cu — fa_prefill_f16kv, today's tree (post-r48 FAP2 shape) */
float mnew0 = -INFINITY, mnew1 = -INFINITY;
#pragma unroll
for (int q = 0; q < FA_TKV / 16 * 4; q++) {
    /* valid = causal (kv <= query pos) AND within the stored KV range
     * (rows >= kv_end are zero-staged and must NOT contribute). */
    bool v0 = (gcol[q] <= qpos0) && (gcol[q] < kv_end);   /* global bound */
    bool v1 = (gcol[q] <= qpos1) && (gcol[q] < kv_end);
    if (v0) mnew0 = fmaxf(mnew0, sm[q]);
    if (v1) mnew1 = fmaxf(mnew1, sm1_[q]);
}

Today's tree tile geometry (FA_TKV 32 is the resident value; r50's attempt at 32→16 was rejected, and r57 states "stays 32"):

/* src/cuda_kernels.cu:3936 — today's tree FA tile width */
#define FA_TKV 32

Worth spelling out the causality: r23/r46's two independent "64→32 raises" were both reverted for missing the wall-clock bar, yet today's tree has TKV=32 — it entered and solidified as the new kernel's geometry in r48's FAP2 rewrite (register-resident softmax, deleting the S/P smem round trip entirely, 69.38 → 34.82 KB). The occupancy wall was ultimately climbed not by shrinking the tile but by deleting another block of smem; r23's pilot supplied the evidence that "2 blocks/SM is worth wanting", and r48 supplied the path that pays no fixed-cost tax.

3.3 Pitfalls

  • The lane mask bug: see §2.3 — shrink the tile geometry and validity masks derived from the old geometry quietly fail; the global bound does not stop lanes whose "local column is out of range but global column still legal".
  • Co-tenant outliers: a decomposition done on a polluted window skews every bucket's share; outlier removal + quiet-window re-verification are hard steps (after r59b this became the whole campaign's standard protocol).
  • The misleading nature of "FA is 7.2%": a small share does not mean a small lever — when r47 later re-decomposed in the converged regime, FA was already the 10.2% #1 structural residual. Shares travel with the baseline (decompositions have a shelf life).

4. Verification

  • Bucket-conservation check: the sum of the arithmetic classes' launch times ≈ nsys wall clock (GPU-busy basis), defending against bucket omissions or double counting;
  • ncu SOL cross-check: each class's representative kernel's mem/SM utilization reconciled against the physical expectation of "should it saturate" (gate/up 85% mem SOL = reasonable saturation; down's 20%/MAC asymmetry is an anomaly, booked as such);
  • Quiet-window re-verification: all shares re-measured on a co-tenant-free window;
  • The FA_TKV raise's full gate set: parity + greedy identity + interleaved A/B — it was exactly this gate set that caught §2.3's mask bug mid-integration (a real bug stopped by the gates, not by the naked eye).

5. Results

5.1 f16-path whole-graph decomposition (2659-token prefill, quiet window 2285 tok/s)

Arithmetic classshare of wallNotes
gate+up GEMM37.8%mem SOL 85% (saturated)
down GEMM27.5%20%/MAC slower than gate/up — cause of the asymmetry unknown, booked
q / o GEMM7.5%—
k / v GEMM1.4%—
FA attention7.2%occupancy 16.7% (1 block/SM)
convert f32→f166.0%llama pays ~0 (quantization in-kernel)
swiglu5.6%DRAM peak
add / rms / rope~5%—
host gaps0.8%—

GEMMs total 74% — the f16 path's wall is GEMM. q6_K layers carry no per-MAC penalty on the f16 kernel (the w16 resident cache smooths out the type difference), i.e. the f16 path is insensitive to weight type; all of MMQ's gain space comes from byte width. Against llama's split (GEMM 597 ms @ 39.5 µs/GMAC vs our 59.5 µs/GMAC + 73 ms convert): the prefill gap = the GEMM byte-width gap; the raw q4_K B stream (4.5 bit/w) at 85% SOL overwhelms f16 (16 bit/w) at ~1.5×/MAC — the campaign premise holds, and yields the quantitative target: MMQ GEMMs to llama's level → ~2900 tok/s.

5.2 The FA_TKV 64→32 raise (REVERTED)

Metricbefore (TKV=64)after (TKV=32)Δ
occupancy16.7%32.68%1 → 2 blocks/SM
FA kernel——−6.7%
whole-prefill wall clock——−0.3% (bar +3%)

Veto mechanism: occupancy doubled and the kernel got 6.7% faster, yet the wall moved only 0.3% — the fixed cost of the k loop's doubled iteration count ate the latency-hiding gain; and FA held only 7.2% of the wall at the time, so a −6.7% kernel amortizes to noise level across the whole graph. Reverted (but the mask-fix knowledge gained en route was booked). FA's 2.5×/layer gap to llama is structural: llama's FA keeps 128-wide KV tiles — shrinking the tile is not the answer to that gap.

Retry conditions: tile/occupancy-class levers are worth touching again only when (a) FA dominates the wall clock and (b) the per-tile fixed cost has been cut (r48's register softmax is exactly (b)). Two re-tests confirmed this verdict: r46 (TKV 64→32 + S/P padding, −11% kernel / +0.27% wall) and r50 (TKV 32→16, the occupancy gain offset by sync overhead + breaking bitwise identity); r48 FAP2 reached 2 blocks/SM by "deleting the S/P smem" and landed +5.6% — same wall, different road.

6. Lessons

  1. Byte-width arithmetic can adjudicate the campaign premise before any kernel exists: the ledger of 59.5 vs 39.5 µs/GMAC + 4.5 vs 16 bit/w is worth more than any single-point kernel experiment done first.
  2. Decompositions have a shelf life: re-run after each convergence (r47 therefore re-judged FA from a 7.2% "small item" to the #1 structural residual); hidden taxes (like convert's 6.0%) must appear in the same table as gains.
  3. Shrink the tile geometry and the lane masks must be re-derived: local bound + global bound are two predicates; missing either is silent numeric pollution of the tail tile (r46/r50 inherited this lesson).
  4. Occupancy is not a free lunch: the fixed cost of ×2 iterations offsets latency hiding; cut the fixed cost first (r48), then talk tile size.

← 27-r22-qa8-xor-swizzle · Index · 29-r24-scheduling-ladder →

29 · r24 — The scheduling-structure ladder (REVERTED; the +1.5% whole-prefill landing bar calibrated here)

Result: all three scheduling-structure rungs negative — tile-order swizzle: A-hot transposed −2.3%, G-grouped −2.4% (default x-fastest/B-hot 1370.0 optimal); persistent blocks −3.3% (the non-persistent u-loop also −3.3%, persistent occ=1/occ=2 ≈ identical). All reverted. This round's real durable output is calibration: the whole-prefill landing bar changed from an absolute ≥1350 tok/s to relative ≥ +1.5% (vs the same-window re-measured baseline) — used ever since by r25/r28/r31/r46/r55/D3a. Commit: d90b3e9 (docs-only record; r24's experiment code never became a code commit, see below). Date: 2026-09-04.

Forensics note (STYLE rule 0): r24's ladder code (the swizzle bijection + persistent grid) is an in-session "implement → measure → revert, never committed" artifact — d90b3e9 contains only the 64-line docs/CUDA_OPTIMIZATION.md record, and a full git-history search finds no r24 code commit. This doc's code excerpts are therefore taken from the current tree throughout (Grep + Read bounded ranges): the current-tree code shows what the ladder acted on (the wide kernel's grid geometry and tile walk), while the ladder itself is presented as clearly-labeled narrative pseudocode.

1. Background — where things stood

By the time r23 closed, the q4_K MMQ line's shape was: mmq_raw_wide_nt_kernel (128 tok × 128 od wide tile, 8 warps × 32×32 warp tile, r14's ldmatrix B fragments, r20's split-phase A staging, r22's qa8 XOR swizzle), with the same-window baseline at KD=8 in the ~1330 tok/s band (r20 landed 1317.7→1319.9, r22 1329.6). r13–r23, eleven rounds, turned over both axes — "how many instructions are issued per MAC" and "where the stalls land": r17 proved pure per-MAC instruction cuts pay ~0 wall clock at SM% ~30, r21 proved stall-mass conservation, r23 decomposed the whole f16-path wall and confirmed GEMM ~85% is the only lever.

One last untried structural family remained: scheduling — touching neither instructions, nor bytes, nor dependencies, only "which blocks run when". It has three natural rungs:

  1. tile-order swizzle: change the block-index→tile mapping (a bijection on the grid) so adjacent blocks share different operands;
  2. persistent blocks: one block resident per SM, a software loop walking the tile list, eliminating the wave-quantization tail;
  3. full stream-k: split along K + a fixup pass write-back (llama's plan; r13 had already measured the launch-shape-level equivalent: wash + fixup +34 µs).

Meanwhile a metrological problem exploded this round: the box drifted faster. The KD=8 re-measured reading was ~1385–1391, while the recorded band was ~1330 (r20/r22's landing numbers). This is not progress — master-table footnote 2 states it plainly: the r12–r25 rounds interleaved on boxes drifting −9% to +38%; only same-window deltas mean anything. But this directly sentenced the then-current working bar: the campaign had been using absolute ≥1350 tok/s as the landing line (r18 "Bar ≥1350 decisively missed", r20 "landed despite missing the ≥1350 bar", r22 "the combined ≥1350 bar not reached") — with the baseline itself already at 1385+, a zero-work binary could "pass the line". The absolute bar's premise (a stable box) had collapsed, the bar had to be recalibrated, and r24 happened to be the round with "no result to protect" — the most suitable round for doing exactly this.

2. Principle — the GPU mechanism

2.1 Why the default order is B-hot: the grid geometry's arithmetic

The wide-tile kernel's grid is two-dimensional: blockIdx.x is the token tile (i0) and blockIdx.y is the od tile (j0) — CUDA unrolls x fastest, so consecutive blocks on the same od column share the same B (weight) panel, while each block's A (activation) token tile is distinct. That is "default x-fastest / B-hot":

grid = (ceil(nt/128), ceil(od/128))
  blockIdx.x = token tile  (x, fastest)  → consecutive blocks: same j0, new i0
  blockIdx.y = od tile     (y, slow)     → the weight panel changes only when y changes

Instantiating the layer-0 q-proj GEMM (nt=512, od=id=3584): grid (4, 28) = 112 blocks ≈ 2.33 waves (at ~48 scheduling slots per SM, 112/48 = 2.33; llama's same GEMM uses a stream-k grid (48,1,1) + fixup (48,4,1)). Each od column has 4 blocks taking turns re-reading the same B (per kt phase a qb8 panel of 128 od × 8 chunks × 48 B ≈ 49 KB, 258 KB over the full K range) — small panel, short reuse window, naturally absorbed by L2; A, meanwhile, is a fresh token tile per block, the streaming side. The x-tile round (row 26) already measured the two streams' sizes: A re-read 327 MB vs B re-read 152 MB, A at 2.1:1.

2.2 Why the swizzle can only lose

Tile-order swizzle changes only the temporal locality window; the total re-read volume is a schedule-invariant quantity (how many bytes each tile must read is decided by the tile geometry). Three candidates:

  • A-hot transposed (x/y swapped): consecutive blocks share A tiles. But A is the dominant stream at 2.1× — making the dominant stream "shared-resident" stretches its reuse window into conflict with the other waves, while B was already free (L2-resident); there is no gain to trade. Measured −2.3%.
  • G-grouped: a third bijection (group-clustered traversal), likewise creating no new operand sharing; measured −2.4% (the master-table row records −2.3/−4.7%, a cross-window basis).
  • Default B-hot 1370.0: this window's optimum. Mechanism: the weight panel is small (the key difference between MMQ and an ordinary GEMM — q4_K weights are 4.5 bit/weight, so the 128-od B panel is far smaller than the 128-token A panel), the L2 residency cost ≈ 0, and the exclusive A stream stays one-step-one-tile, naturally pipelined.

The structural reason behind the conclusion: at 1 block/SM and ~2 warps/sched occupancy, the kernel is latency-bound, not byte-bound (r21's "stall-mass conservation" already falsified the byte side once). Scheduling can change only L2 temporal locality, and both streams' temporal locality under the default order is already no bottleneck — there is nothing separable.

2.3 Why persistent blocks have no tail to eliminate

A persistent grid = launch exactly num_sms resident blocks, each software-looping over a strided tile list. It pays off only under two premises: (a) a wave-quantization tail exists — the fractional part of ceil(waves) leaves the last batch of SMs idling; (b) the per-block fixed cost (smem zeroing, attr setup) is a significant share. Here 112 blocks ≈ 2.33 waves, but the GPU's launch is pipelined, not lockstep waves — as soon as a block of the previous grid exits, the next grid's block fills in immediately; the 0.33-wave "tail" was never idle. With no tail, persistent merely replaces the overlap the hardware block scheduler does for free with a software loop doing it itself, plus the loop's address arithmetic. Measured: the non-persistent u-loop alone is −3.3% (merely changing one-launch-one-tile into looping over tiles), and persistent occ=1 / occ=2 ≈ identical — no difference even along the occupancy dimension.

2.4 Why stream-k (rung 3) was vetoed before any code

Full stream-k splits along K across blocks, and the partial sums need a fixup pass re-ordering floating-point accumulation (llama pays +34 µs for this). This campaign's hard gate is greedy-32 byte-for-byte identity — a floating-point re-order necessarily flips some argmax blade at some step. So rung 3 is not "never got around to it" but vetoed a priori by the gate: r24's bijection rungs were deliberately chosen as bit-identical shapes; stream-k is the only one mathematically guaranteed to break identity, not worth spending implementation cost to hit a known wall.

3. Implementation

3.1 Design choices (why this shape and not another)

  • A bijection, not a reordered loop: rung 1 uses the env var MINFER_MMQ_RAW_SCHED to select the block index → tile mapping function. The bijection guarantees each tile is computed exactly once and the output memory image is bit-for-bit identical — correctness holds by construction, and measurement only needs to pass the performance gates. Far cleaner than adding a "schedule mode" branch inside the kernel: zero change to the hot loop.
  • Rung 2 split into two variables: first measure the non-persistent u-loop (changing the tile walk into a loop, still one-block-one-tile semantics), then persistent (num_sms resident + strided list). −3.3% appeared at the first step, and adding persistence changed nothing — the two variables cleanly isolated: the cost comes from loop-ification itself, not from persistence.
  • Register-identity check: the persistent rewrite's ptxas register allocation matches the original's (register-identical per ptxas), ruling out the confound "the negative came from register pressure" — the −3.3% is the pure scheduling-structure cost.
  • Using r24 to calibrate the bar: precisely because all three rungs were expected negative, this round was pure measurement-framework practice — same-window interleaved A/B, distribution-separation reading (later hardened by doc 77's methodology into "min-new > max-base"). See §5.2.

3.2 Key code (current tree; what the ladder acted on)

The ladder acts on the wide kernel's grid geometry and tile walk. The current tree's launch_mmq_raw_wide_nt grid mapping and smem budget (src/cuda_kernels.cu:7200-7205):

    // 16-chain layout: 128-token x 128-od block tile. r14: qb8 slot-major
    // 48B stride (ldmatrix-for-B) + packed scales. KD=8 totals 98,304B and
    // KD=4 73,728B — both inside the ~99KB opt-in cap, 1 block/SM. The
    // attr/launch results are checked: an over-cap request used to fail
    // SILENTLY (r7 phantom 2124).
    dim3 grid((nt + 127) / 128, (od + 127) / 128);

The kernel-side tile walk (src/cuda_kernels.cu:5892-5902) — x is the token tile, y the od tile, exactly the mapping the swizzle wanted to rearrange:

    // Each warp owns a private 16-od-row slice and reads the FULL 128-token
    // tile: B fragments become warp-exclusive, A fragments are warp-shared.
    const int i0 = blockIdx.x * MMQ_WBI;      // token tile  (x, fastest)
    const int j0 = blockIdx.y * MMQ_WBJ;      // od tile     (y, slow)  = B-hot
    const int j0w = warp * 16;
    const int nb32 = id >> 5;
    const int nchunk = nb32;
    const int nsb = nb32 >> 3;
    const int nktile = (nchunk + KDR - 1) / KDR;

    float sum[64] = {0.0f};   // [g][nh][l]: 8 A-frags x 2 B-frags x 4 C regs

The ladder's own shape (narrative pseudocode, never committed — d90b3e9 is a docs-only commit):

# Rung 1 — tile-order swizzle (a bijection on the grid, selected by MINFER_MMQ_RAW_SCHED)
sched = env("MINFER_MMQ_RAW_SCHED")            # default "x" (B-hot)
(x, y)  = (blockIdx.x, blockIdx.y)
match sched:
    "x"    -> tile = (x, y)                    # default: B-hot, measured optimal 1370.0
    "a"    -> tile = (y, x)                    # A-hot transposed, −2.3%
    "g"    -> tile = group_order(x, y)         # G-grouped, −2.4%
# bijection ⇒ each tile exactly once ⇒ output bit-identical

# Rung 2 — persistent blocks
for t in tiles(stride = gridDim, offset = blockIdx):   # u-loop, −3.3%
    compute_tile(t)                                    # occ=1 / occ=2 ≈ identical
# ptxas: register allocation identical to the original

# Rung 3 — full stream-k: not implemented. An fp accumulation re-order ⇒ greedy identity necessarily breaks.

3.3 Pitfalls

  • The absolute anchor failed silently: nobody "broke" anything — the box drifting faster turned the ≥1350 absolute bar into something passable at zero cost. The lesson: an anchor must live in the same frame as the reading (same-window relative quantities); otherwise the bar loses its force without anyone noticing.
  • "struct family untried" ≠ "struct family promising": the scheduling family has no separable resource on a latency-bound, 1 block/SM kernel — working through §2's arithmetic (the A/B re-read ratio, wave count, tail existence) before acting predicts it; half of r24's value is turning that prediction into a measured record.
  • The u-loop has no free lunch: even without persistence, merely changing "one block one tile" into "a block looping over tiles" is −3.3% — the hardware scheduler's work should not be replicated in software.

4. Verification

  • bit-identical (rung 1): the bijection guarantees by construction that the output memory image is bit-for-bit identical; greedy-32 identity therefore holds automatically (still run once as a regression confirmation). It defends against "a schedule change quietly skipping/recomputing some tile".
  • Interleaved A/B medians (uniform across the three rungs): same window, warmed up, alternating order, read against the same re-measured baseline. It defends against co-tenancy/thermal drift disguising the box's drift as a trend — this round's box drift is precisely why this reading protocol exists.
  • Register identity (rung 2): ptxas output comparison, ruling out the register-pressure confound.
  • A priori gate veto (rung 3): the greedy identity gate sentenced stream-k before any code was written — the cheapest use of a verification gate is "before building".

5. Results

5.1 The three rungs' numbers (same-window interleaved, vs the re-measured baseline)

rungshaperesultverdict
1swizzle MINFER_MMQ_RAW_SCHEDdefault B-hot 1370.0 optimal; A-hot −2.3%; G-grouped −2.4%REVERTED
2persistent blocksu-loop −3.3%; persistent occ=1/occ=2 ≈ identicalREVERTED
3full stream-knot implemented — vetoed a priori by the greedy identity gateCLOSED

The box's re-measured band: KD=8 ~1385–1391 (recorded band ~1330). Veto mechanism: the scheduling family has no separable resource on a kernel with "1 block/SM + latency-bound + L2 already absorbing B reuse"; retry conditions — (a) a real quantization tail appears in the grid (large-tile shapes with tile count < SM count); (b) the per-block fixed cost is visible relative to per-tile work (very short K); (c) the operand re-read ratio inverts to B-dominant (only then does the A-hot order have a theoretical gain); (d) stream-k only if the numeric gate is replaced by a tolerance regime (never happened in this campaign).

5.2 This round's real output: calibrating the +1.5% relative bar

Why can a reverted round calibrate a bar? Because calibration needs exactly a "disinterested" measurement-framework rehearsal, and r24 is exactly that:

  1. The absolute bar's death certificate: the baseline at 1385+ already passed ≥1350 by itself — the absolute anchor's premise (a stable box) was falsified. The bar had to become a same-window relative quantity.
  2. Resolution demonstrated: this round's three rungs' −2.3% / −2.4% / −3.3% were all unambiguously judged negative under same-window interleaved A/B — the protocol has clean resolution for ≥2% effects, with the noise band clearly narrower than that magnitude. The landing line is set at +1.5%: above the noise band (readings reliable), below typical structural levers' magnitude (a real lever can pass: r28 +2.56%, r29 +2.80%), and high enough to block "sub-noise positives with an unproven mechanism".
  3. The reading discipline, distilled: same-window pairing, alternating order, distribution separation (min-new > max-base), later hardened by doc 77's methodology into §2.3 — "the whole-prefill landing bar is +1.5% (relative to the re-measured baseline, calibrated in r24)".

Every round since reads inside this frame: r25's instruction cut +0.37/+0.49% → below the bar, reverted; r28's Phase-2 gate 4 wrote in black and white "≥ +1.5% over the re-measured baseline (r24 convention)"; r31's +1.07% was below the bar but mechanism-confirmed by ncu and landed after explicit recording as real-but-sub-bar; r46's FA +0.27% → reverted; r55 even used a roofline argument to derive "even a perfect kernel could not pass the bar" and skipped implementation; D3a listed it alongside the ±2% A/B noise as veto grounds. In one sentence: r24's levers all died, but they bought the common yardstick used by the 30+ rounds since.

6. Lessons

  1. A bar must live in the same frame as its reading: an absolute anchor holds only while the box is stable; once the box drifts it becomes a ritual — relative (same-window paired) quantities are the only cross-session comparable thing.
  2. A "failed" round's durable output need not be code: baseline re-measurement + resolution demonstration + reading discipline are calibrations only a disinterested round can fix accurately.
  3. A scheduling-structure lever's premise is "separable temporal locality or a quantization tail exists" — compute the re-read ratio, wave count, and launch pipelining first; most swizzle/persistent attempts can be predicted negative before any code.
  4. A verification gate's cheapest use is before implementation: stream-k's floating-point re-order being incompatible with greedy identity is an a priori fact — not one line of code should be written for it.

← 28 · r23 f16 wall decomposition · Index · 30 · r25 SASS opcode census →

30 · r25 — SASS opcode census; the unroll is wall-inert (MEAS-ONLY + REVERTED)

Result: the census attributes the +25.2% instruction surplus 100% to supporting instruction classes — integer ALU +69,465/tile (77%), fp32 rescale FMUL +15.2k, conversions +14.3k; IMMA and FFMA are equal item-for-item with llama (114,688 FFMA/tile on both sides). The accompanying kd-unroll fix cut integer ALU −38% and total instructions −9.7% (surplus 89.8k → 46.6k/tile, halved), but the wall clock moved only +0.37/+0.49% — below the +1.5% bar calibrated in r24, reverted; the census itself is this round's deliverable. Paradigm verdict: the kernel is issue/occupancy-bound (98 KB smem → 1 block/SM → ~2 warps/sched → latency unhidden), not instruction-count-bound. Commit: 8658f1b (docs-only record; the unroll attempt's code, like r24's, never became a code commit). Date: 2026-09-04.

Forensics note (STYLE rule 0): r25 is a "measure + attempt + revert" round; 8658f1b contains only the 87-line docs/CUDA_OPTIMIZATION.md record; the census numbers were reproduced verbatim in docs/LLAMA-CPP-MMQ-ANALYSIS.md §7 (this doc quotes that table directly and notes the source). Code excerpts come from the current tree: the wide kernel's 1-block/SM budget comment, and the NB kernel's kd loop as landed by r29 (the same unroll lever's final home, for comparison).

1. Background — where things stood

r24 had just closed the scheduling-structure family and recalibrated the landing bar to relative +1.5%. At this point the campaign had two layers of evidence for "why it is slow": r13's counter forensics said the 3× gap's carrier is the per-MAC warp instruction stream (10.14 vs 6.06 M/GMAC, not bytes), and r20's stall table localized that stream's latency exposure to long_scoreboard (6.22 vs 1.15, 97% of the named-stall surplus). But r13's table had one defect, named when r20 re-reviewed it: the two sides' shapes mismatch (the pre-r14 kernel ran nt-2630, llama ran nt-512) — "double-distorted". Attribution had reached the stall, but not yet the instructions themselves: which SASS classes make up the surplus? Math (mma)? Movement (LDG/STS/LDSM)? Or glue (address arithmetic, predicates, loop control)?

This question decides the next lever's life or death: if the surplus is in mma, it is a decomposition problem; if in movement, a layout problem (r14/r20/r22 had already taken three rounds); if in supporting instructions, it is a "glue per MAC" problem — cuttable, but r17 had already rehearsed "cut it and the wall doesn't move" once. r25 answers both questions at once: the surplus is real (+25.2%, precise to the class), and at the current occupancy cutting it is wall-inert — and explains why. That explanation (issue/occupancy-bound) directly spawned r28's Direction-A design.

2. Principle — the GPU mechanism

2.1 How the census was done: the per-opcode-class counting chain

The tool is ncu's thread-granularity per-opcode metrics (the sass_thread_inst_executed_op_* family, pred_on counts): each SASS instruction class is counted at thread granularity; divide by 32 to get warp instructions (the /32 conversion was cross-verified against the warp counters). Steps:

  1. Match the GEMM: both sides run layer-0 q-proj, nt=512, id=od=3584. On minfer's side mmq_raw_wide_nt_kernel grid (4,28) = 112 128×128 tiles; on llama's side mul_mat_q<12,128,0> grid (48,1,1) + fixup (48,4,1) (stream-k, equivalent to the same GEMM's 112 tiles).
  2. Sum by class, normalize by tile: each class's total divided by the tile count = per-tile warp instruction count. The ledger reconciles: ours smsp__inst_executed.sum 49,900,928 → 445,544 warp-inst/tile; theirs 39,844,864 → 355,758; surplus +89,786/tile (+25.2%).
  3. The reconciliation gate: the gap between the classes' sum and smsp__inst_executed.sum converges to ~1.4% — the anchor of the census's credibility, defending against "class definitions under- or double-counting".

2.2 The census table (reproduced verbatim from docs/LLAMA-CPP-MMQ-ANALYSIS.md §7)

SASS class (ncu opcode metric)ours/tiletheirs/tiledelta/tileratioreading
integer ALU (IADD3/IMAD/LEA/SHF/SEL/ISETP/LOP3)113,55244,087+69,4652.58×← 77% of the surplus
FP32 FMUL (rescale)72,59257,344+15,2481.27×dequant-rescale
conversion (I2FP/F2I)72,60058,254+14,3461.25×int-mma→fp32
misc (NOP/CS2R)13,3365,851+7,4852.28×loop/init
control-flow (BRA/isync)3,5841,160+2,4243.09×loop control
uniform datapath (UR)1,80855+1,75333×uniform regs
FP32 FFMA (rescale/accum)114,688114,688+01.00×exactly equal
bit (LOP3/PRMT/SHF)8456−4480.02×(theirs higher)
fp16 HADD2/HFMA path15,23236,400−21,1680.42×(theirs higher)
memory (LDG/STS/LDS/LDSM)31,81636,836−5,0200.86×(theirs higher)

Accompanying per-tile memory-family counts: global_ld ours 6,944 / theirs 6,384; LDSM ours 8,064 / theirs 1,792 (4.5×); shared_ld ours 8,960 / theirs 19,346 (ours is lower instead); shared_st ours 5,376 / theirs 7,730.

2.3 Three structural readings

  • The compute side is exactly MAC-bound: IMMA (r20: 1,605,632 equal on both sides) and FFMA (114,688/tile equal on both sides) match item for item — the mma work, decomposition, and accumulation depth are all aligned; compute was never the gap.
  • The movement side is cheaper on ours: shared_ld below llama's, LDSM 4.5× theirs — r14's ldmatrix B path is exactly "fewer, wider smem ops". "The surplus comes from moving bytes" is rejected.
  • The surplus = 100% supporting instructions: integer ALU (address/predicate/loop arithmetic) +69.5k holds 77%, and fp32 rescale (FMUL + I2FP, ~29.6k combined) holds most of the rest. (An honest annotation attached: llama's fp16 classes are higher, and the "llama uses half2 scales" reading is still marked interpreted rather than source-verified in the analysis doc's §Corrections — the census records counts, not facts about the opponent's source.)

2.4 The mechanism of "wall-inertness": issue arithmetic at 1 block/SM

Wall-inert = the instruction stream genuinely shortens and the wall clock does not budge. It has a precise mechanism, with every number in r20/r25's counters:

  • The wide kernel KD=8's smem budget is 98,304 B → 1 block/SM (inside the ~99KB opt-in cap, but doubling would breach it); 256 threads per block = 8 warps, 4 schedulers → ~2 resident warps per scheduler.
  • r20's table: warps_active 2.00 vs 2.00 (equal residency), but issue/cyc/sched 0.20–0.26 vs 0.42 and eligible 0.28 vs 0.64 — llama squeezes 2× the issue out of the same 2 resident warps. Where the difference lies: our 2 warps spend most of their resident cycles stalled on long_scoreboard (2.92, post-r20; llama 1.15) — when both warps wait on the same load's return, the scheduler holds no eligible warp and the issue slots idle.
  • So cutting instructions cannot change the wall clock: the wall is decided by latency-chain length ÷ available overlap (warp count), not by total instruction count. Issue slots are only 25–26% used (≪100%), showing "instructions too dense" is not the constraint at all — the constraint is too few resident warps. ∂wall/∂inst ≈ 0.

The verdict yields a falsifiable prediction: the same instruction cut should start paying once occupancy is bought (2 blocks/SM → ~4 warps/sched). r29 cashed that prediction (+2.80%), and r28's job was to buy the occupancy.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Census first, lever second: attribute the instruction stream first, then decide what to cut — r17's lesson (blind per-MAC instruction cuts paid 0 wall clock) upgraded into procedure: this time there is a per-class ledger before the cut.
  • The matched-shape protocol: the root cause of r13's table's distortion was the two sides' different nt; r25 made "same layer, same nt, both binaries, tile counts aligned (112 = 112)" the census's precondition gate.
  • Lever choice: the kd-unroll: within the integer ALU's composition, each chunk's select/base/ bounds (is_hi, smem base, bound compares) are loop-invariant yet recomputed every iteration — #pragma unroll folds them into compile-time constants, the cheapest targeted cut, and it changes no accumulation order (bit-identical by construction).

3.2 Key code

The tested lever itself is one #pragma unroll line on the wide kernel's kd loop (attempted then reverted, never committed). Its final home in the campaign is r29 — the same lever line landed on the NB kernel. The current tree's mmq_raw_nb_kernel kd loop (src/cuda_kernels.cu:6333-6341; the current-tree form includes r29's unroll; it was absent when r28 landed):

    for (int kt = 0; kt < nktile; ++kt) {
        if (kt > 0) RAW_STAGE_NB(kt);
        __syncthreads();

        #pragma unroll                    // ← exactly the line r29 landed: the kd loop unrolled,
        for (int kd = 0; kd < KDR; kd++) { //   is_hi/smem base/bounds all become compile-time constants
            const int c = kt * KDR + kd;
            if (c >= nchunk) break;
            const int sg = c & 7;
            const uint8_t* qat = qa8 + (size_t)kd * MMQ_NBI * 32;

The wide kernel the census targeted and its 1-block/SM budget (src/cuda_kernels.cu:7200-7202):

    // 16-chain layout: 128-token x 128-od block tile. r14: qb8 slot-major
    // 48B stride (ldmatrix-for-B) + packed scales. KD=8 totals 98,304B and
    // KD=4 73,728B — both inside the ~99KB opt-in cap, 1 block/SM.

The post-unroll-fix reconciliation arithmetic (the derivation of the halved surplus):

int-ALU class cut    = 0.38 × 113,552        ≈ 43,150 warp-inst/tile
total instruction cut = 43,150 / 445,544      = −9.7%            (matches measurement)
total surplus          = 89,786 − 43,150       ≈ +46.6k/tile      ("surplus halved")
int-ALU surplus        = 69,465 − 43,150       ≈ +26.3k/tile      (still the largest class)
wall clock             = +0.37 / +0.49%        <  +1.5% bar       ⇒ REVERTED

3.3 Pitfalls

  • The shape-mismatch re-accounting trap: r13's table's nt-2630 vs nt-512 distorted the per-tile comparison; any "per-tile normalized" census must first align tile counts (112 = 112), otherwise the surplus number itself is an artifact.
  • ncu is a structural authority only: on GB10 ncu serializes replays and per-kernel times are distorted (doc 77's methodology §2.5) — the census reads counts (inst/occupancy classes); all wall-clock verdicts go to interleaved A/B measurement; the two evidence sets are never mixed.
  • "Below the bar" ≠ "measured for nothing": without r24's freshly calibrated relative bar, +0.37/+0.49% would have been read as "a small positive in the noise" and wrongly landed. The bar's value cashed in for the first time here: vetoing a real but inconsequential improvement.

4. Verification

  • Reconciliation gate: the per-class sum vs smsp__inst_executed.sum differs by ~1.4%; the thread→warp /32 conversion cross-verified against the warp counters — defends against class-definition under- or double-counting.
  • Matching gate: same layer-0 q-proj, same nt=512, both binaries, tile counts aligned — defends against r13-style shape artifacts.
  • Interleaved A/B measurement: the unroll fix +0.37/+0.49% (read against r24's relative bar) — defends against the intuition-swap of "counters better = wall better".
  • Identity gate: the unroll changes no accumulation order; parity/greedy still ran as usual (the attempt itself was reverted; the identity gate's meaning is confirming "the negative result was not bought with numeric damage").

5. Results

  • The census (the deliverable): surplus +89,786 warp-inst/tile (+25.2%), 100% supporting instruction classes (int ALU 77%); IMMA/FFMA equal item-for-item on both sides; the movement side lower on ours. Per-tile wall 211 µs vs 113 µs (+34 µs fixup) — a ≈1.4× occupancy-bound residual (211 / (113+34) ≈ 1.44).
  • The unroll attempt (reverted): int ALU −38%, total instructions −9.7% (surplus 89.8k → 46.6k/tile), wall clock +0.37/+0.49% — below the +1.5% bar.
  • Veto mechanism: at 1 block/SM and ~2 warps/sched, instruction cuts are wall-inert — the root cause of the idling issue slots is too few resident warps, not too many instructions. Retry conditions: occupancy ≥ 2 blocks/SM. That condition was satisfied by r28's Direction-A kernel, and r29 immediately cashed +2.80% on the same lever (int ALU −25%, total instructions −6.5%).
  • The paradigm verdict (this round's most important output): the kernel is issue/occupancy-bound, not instruction-count-bound. r28's design doc §11.3 quotes this verdict directly for a reverse bet: "since a −38% integer cut cannot move the wall clock by even 0.5%, an instruction increase of a few percentage points should also be wall-inert — provided the occupancy lever actually fires".

6. Lessons

  1. Attribute the instruction stream before cutting it — the per-class ledger (with its reconciliation gate) turns "where to cut" from a guess into a reading; after cutting, also check whether the cut class stands on the critical path.
  2. Issue arithmetic can predict wall-inertness: issue 0.25 vs 0.42 @ warps_active 2.00 = 2.00 says the issue slots are not the constraint; look at occupancy before deciding whether to save instructions.
  3. Occupancy and instruction cuts are ordered, not independent: buy occupancy first (r28) and only then do instruction cuts become visible (r29) — the same lever, reordered, went from +0.4% to +2.8%.
  4. A MEAS-ONLY round's deliverable can be a ledger: the census table was repeatedly cited afterwards by r29 (the NB kernel's re-census) and §8/§9's design comparisons; its compound interest exceeds most landed patches.

← 29 · r24 scheduling-structure ladder · Index · 31 · r28 raw-nibble NB kernel →

31 · r28 — Direction-A raw-nibble NB kernel, 2 blocks/SM (LANDED)

Result: the companion kernel mmq_raw_nb_kernel (64 tok × 128 od, KD=8-native, smem 45,056 B → 2 blocks/SM; the wide kernel is 98,304 B → 1): whole-prefill median 1375.2 → 1410.4 (+2.56%, 5/5 positive; the earlier batch +2.43%, 3/3), clearing the +1.5% bar set by r24; parity 1.5e-5..9.9e-5 (pure f32 rounding); greedy-32 byte-identical; ncu: sm__warps_active.avg.per_cycle_active 16.17 ≈ 4.04 warps/sched (2 blocks/SM confirmed), long_scoreboard 2.92 → 2.03, issue_active 25 → 37.36%; 123 regs / 0 spill. r25's verdict inverted: occupancy is the bound resource — and it can be bought back with smem. Commit: 0957a08 (design doc 2f783a3). Date: 2026-09-04.

Code provenance note (STYLE rule 0): the kernel/launcher/dispatch excerpts come from the current tree (Grep + bounded Reads). The current tree = the r28 landed form + two later small changes, both flagged in place below: ① the #pragma unroll on the kd loop (landed in r29); ② sda_q repacked to one uint32 per token (r31, smem 45,056 → 43,008 B). The r28-era dispatch gate is taken from a narrow-range diff of git show 0957a08 -- src/cuda.rs.

1. Background — where things stood

r24 had closed the scheduling-structure family and r25's census delivered the paradigm verdict: the wide kernel is issue/occupancy-bound — 98,304 B smem → 1 block/SM → 8 warps/SM → ~2 warps/sched; with both resident warps stalled together on long_scoreboard the issue slots spin empty (issue 0.25 vs llama 0.42 at warps_active 2.00 on both sides). That verdict in turn pushed a never-touched lever onto the stage: occupancy itself can be bought back — paid for in smem.

The largest slice of those 98,304 B is exactly the B (weight) plane: qb8's expanded per-k int8 storage (1 byte per nibble, slot-major 48 B slots) takes 8 × 128 × 48 = 49,152 B — exactly half. llama's MMQ, by contrast, keeps the weights as raw nibbles (2 nibbles/byte, x_qs[...] = (qs0>>0) & 0x0F0F0F0F, unpacked only at use time). Direction A's bet took shape from this: swap the B smem plane for a raw-packed qs plane, accept the small ALU cost of unpacking inside the inner loop, and squeeze the per-block budget to ≤ ~49.5 KB → 2 blocks/SM → ~4 warps/sched. r25 had already proven "−38% instruction cut, <0.5% wall-clock", so "+a few percent of instructions" should also be wall-inert provided occupancy lands — both ends are functions of occupancy; this bet is placed on occupancy.

The danger in this step was numerics: nibble-layout mapping is the breeding ground of the r13-era "82.896 max-diff" parity pattern (wrong nibble position, wrong sign extension, wrong dmin fold — any one of them produces garbage diffs at the 1e0 scale). The design doc (2f783a3, §11.5 risk table #1) listed "wrong B-fragment nibble layout" as the top risk and mandated standalone verification before integration.

2. Principle — the GPU mechanism

2.1 The occupancy arithmetic: why 45,056 B is exactly 2 blocks/SM

On consumer-class parts like GB10 the per-SM shared-memory budget is on the order of ~99-100 KB (r7's measured "~99KB opt-in cap" is the same ballpark; the design doc §11.1 rule of thumb: per-block ≤ 49.5 KB ⇒ 2 blocks/SM). Candidate geometries (KD=8, KDR=8):

Geometry (T×O)QA8SDAQB_expSDSexp totalQB_rawraw totalblocks/SM (exp / raw)
64×12816,3844,09649,1528,19277,82416,38445,0561 / 2
128×6432,7688,19224,5764,09669,6328,19253,2481 / 1
64×6416,3844,09624,5764,09649,1528,19232,7682 / 2–3

The wide kernel 128×128 + QB_exp = 98,304 B → 1 block. 64×128 + raw is the only shape satisfying both constraints at once: (a) 2 blocks/SM; (b) the od tile stays 128. The latter is dictated by the A/B re-read arithmetic (x-tile row 26: A 327 MB vs B 152 MB = 2.1:1) — A re-reads scale with od/O, B re-reads with nt/T; the dominant stream is A, so O cannot shrink: 64×64 saves more smem but doubles od/O, driving the dominant stream 327 → 654 MB — strictly worse; 64×128 merely sacrifices the smaller stream (B re-reads 152 → 304 MB, already absorbed by L2).

2.2 Why raw nibble halves the plane, and where the cost lands

  • Halving: qb8 expanded = 1 byte per nibble (48 B slot/row); a raw qs plane = 128 B per od row holding 256 nibbles (2/byte) → B plane 49,152 → 16,384 B. This is the inverse trade of llama's scheme: it unpacks at stage time (x_qs masking), we mask at use time (& 0x0F) — only that way does the smem saving materialize.
  • The cost (quantified with the r25 census, §11.3): the census proved minfer was already using the cheaper ldmatrix B path (LDSM 8,064/tile vs theirs 1,792). Option 2 moves B back to plain LDS + unpack: ~30-60 extra ALU per (warp, chunk); the hedge is that the A side drops from 8 LDSM per chunk to 4 (T 64 → 4 token groups). Net instructions +3-6% — an order of magnitude smaller than the −38% cut r25 proved wall-inert, contingent on the occupancy lever firing.
  • Warp shape (§11.2): 8 warps × 16 od-rows = 128 od; each warp exclusively owns 16 od rows and reads all 64 tokens; mma m16n8k32 maps m=token, n=od; 8 independent mma chains per 32-k chunk (4 A-frags × 2 B-frags), sum[32] (down from 64); register estimate ~110-130 (avoiding r22 Lever-2's 255-spill cliff).

2.3 The numerics contract: unsigned nibble + rank-1 two-term rescale

mma consumes the unsigned 0..15 nibble directly as the int8 B operand (the high nibble is zero ⇒ always positive); the integer accumulator holds C_int = Σ_k nib(k)·act(k); the per-chunk fp32 fold is the exact r15 two-term form — sum += da·dsv·C_int + dma·dmv, where dsv = d·sc and dmv = −dmin·m. This is precisely the dequant form of d·s·nib − dmin·m. A centered fold of the shape (nib − m) is forbidden anywhere: q4_K's dmin offset is multiplied by the sub-block scale (−dmin·m); a fixed subtraction is a mathematical error — the root of the 82.896-diff pattern. The fp32 write-back epilogue stores C[i·od + j] = sum[...] directly; no f16 anywhere on the mma→store path.

3. Implementation

3.1 Design choices (why this shape and not another)

  • A companion kernel, not a wide-kernel replacement: NB is gated by MINFER_MMQ_RAW_NB=1; the wide kernel stays byte-identical and remains the default raw path — A/B and rollback are free, the risk ceiling is "zero".
  • KD=8-native: the raw qs plane encodes the full 256-k super-block; KD=4 is meaningless for it → for KD≠8 the launcher cleanly returns 0 (fall back to wide/narrow), never ambiguous.
  • Standalone verification before integration: the B-fragment nibble mapping was derived from the wide kernel's validated ldmatrix path, then given a standalone byte-equivalence test (8 sgs × 32 lanes × 4 regs, 0 mismatches) — the top risk was cleared before integration.
  • Keep the proven machinery: r20 split-phase A staging, the r22 XOR swizzle, and the r15 two-term rescale are kept as-is — NB changes only "kernel shape + B representation", stacking no new variables.
  • Pre-registered kill criteria (§11.7): parity fails three fixes in a row, or occupancy arrives (ncu reads ~4 warps/sched) but wall-clock does not move (falsify and close Direction A), or KD=8 register spill proves incurable — any hit pulls the plug; no "let's try again" gray zone.

3.2 Key code

Kernel signature and smem map (src/cuda_kernels.cu:6199-6233; the Q-major sda_q layout in the comment is the post-r31 form; at r28 it was uint2-per-token, 4,096 B):

__global__ void __launch_bounds__(256) mmq_raw_nb_kernel(
    const uint8_t* __restrict__ W, const uint8_t* __restrict__ q8x,
    float* __restrict__ C, int nt, int od, int id
) {
    extern __shared__ uint8_t mmq_nb_sh[];
    // Single-buffer sync-staged, KD=8 totals 43,008 B -> 2 blocks/SM:
    //   qa8     [KDR][64][32]   chunk q8 planes (r22 XOR swizzle, r20 split)
    //   sda_q   [KDR][64]         (d f16 | ssum i16) packed, one uint32 per
    //                             token, Q-MAJOR: ...
    //   qb_raw  [128][128]        raw GGUF qs plane (2 nibbles/byte, full
    //                             super-block per od-row)
    //   sds     [KDR][128] float2 (d | dmin*m): r15 rank-1 rescale terms
    uint8_t* qa8 = mmq_nb_sh;
    uint32_t* sda_q = reinterpret_cast<uint32_t*>(qa8 + KDR * MMQ_NBI * 32);
    uint8_t* qb_raw = reinterpret_cast<uint8_t*>(sda_q + KDR * MMQ_NBI);
    float2* sds = reinterpret_cast<float2*>(qb_raw + MMQ_NBJ * 128);

A-side staging: r20 split-phase + r22 XOR swizzle kept as-is (src/cuda_kernels.cu, inside RAW_STAGE_NB; excerpt of the LDG batch and the swizzled store):

            unsigned av[KDR * 2];   /* 8 words x 64 tok x KDR / 256 thr */
            _Pragma("unroll")
            for (int i = 0; i < KDR * 2; ++i) {          /* LDGs issued first, batched */
                const int x = threadIdx.x + i * 256;
                const int u = x & 7, r = (x >> 3) & (MMQ_NBI - 1),
                          kd = x / (8 * MMQ_NBI);        /* KDR=8, 8*NBI=512 */
                const int tok = i0 + r, c = (kt) * KDR + kd;
                unsigned v = 0;
                if (tok < nt && c < nchunk)
                    v = *(const unsigned*)(q8x
                        + ((size_t)tok * nb32 + c) * 40 + 4 + u * 4);
                av[i] = v;
            }
            ...
                const int R = kd * MMQ_NBI + r;           /* r22 XOR swizzle store */
                *(unsigned*)(qa8 + (size_t)(R & ~3) * 32
                    + (size_t)(((((R & 3) << 1) + (u >> 2))
                                ^ ((R >> 2) & 7)) << 4)
                    + (size_t)(u & 3) * 4) = av[i];

B-side staging: the raw qs plane as a pure bulk copy (zero staging ALU) + the SDS two-term form (src/cuda_kernels.cu, excerpt):

        /* ---- B: bulk raw qs super-block copy (r18-style, no staging ALU) */
        for (int off = threadIdx.x; off < MMQ_NBJ * 8; off += blockDim.x) {
            const int jj = off >> 3, c8 = off & 7;
            const int j = j0 + jj;
            uint4 v = make_uint4(0, 0, 0, 0);
            if (j < od && sb < nsb)
                v = *(const uint4*)(W + (size_t)j * ((size_t)nsb * 144)
                    + (size_t)sb * 144 + 16 + (size_t)c8 * 16);
            *(uint4*)(qb_raw + (size_t)jj * 128 + (size_t)c8 * 16) = v;
        }
        /* ---- B: SDS per-(chunk, od-row) rank-1 rescale terms ---- */
        ...
                dv = d * (float)sc;            /* d·sc    */
                mv = -(dmin * (float)m);       /* −dmin·m: the two-term form, never folded */
            sds[(size_t)kd * MMQ_NBJ + r] = make_float2(dv, mv);

The core: the inner-loop raw-nibble B-fragment unpack (src/cuda_kernels.cu:6358-6374 — where all of this step's risk and all of its payoff live):

            // B fragments: raw-nibble in-loop unpack (validated == wide kernel
            // ldmatrix). reg0 = qs[(sg>>1)*32 + (l&3)*4 + 0..3],
            //            reg1 = qs[(sg>>1)*32 + 16 + (l&3)*4 + 0..3].
            {
                const int p = sg >> 1, is_hi = sg & 1, lm3 = lane & 3;
                const unsigned M = 0x0F0F0F0Fu;
                #pragma unroll
                for (int nh = 0; nh < 2; nh++) {
                    const int jj = j0w + nh * 8 + (lane >> 2);
                    const uint8_t* qs = qb_raw + (size_t)jj * 128;
                    const uint32_t* q0 = (const uint32_t*)(qs + p * 32 + lm3 * 4);
                    const uint32_t* q1 = (const uint32_t*)(qs + p * 32 + 16 + lm3 * 4);
                    uint32_t v0 = *q0, v1 = *q1;
                    b[nh][0] = (int)(is_hi ? ((v0 >> 4) & M) : (v0 & M));
                    b[nh][1] = (int)(is_hi ? ((v1 >> 4) & M) : (v1 & M));
                }
            }

Key points: sg & 1 picks the high/low nibble (one 32-bit word packs 8 nibbles, half low and half high); after masking there is no sign extension — the unsigned 0..15 value is used as int8, and centering is left entirely to SDS's two-term rescale. is_hi/p/lm3 are all loop-invariant within the kd loop; r29's unroll (the outer #pragma unroll on the next line, current-tree line 6337) folds them into compile-time constants.

8 independent mma chains + the two-term rescale fold (src/cuda_kernels.cu, excerpt):

            #pragma unroll
            for (int g = 0; g < 4; g++)          /* 4 A-frags x 2 B-frags, */
                #pragma unroll                    /* all C fragments live at once */
                for (int nh = 0; nh < 2; nh++)
                    mmq_mma_k32(clow[g][nh], a[g], b[nh]);
            ...
                for (int nh = 0; nh < 2; nh++)
                    #pragma unroll
                    for (int l = 0; l < 4; l++) {
                        const float da = da_q[l >> 1];
                        const int idx = (g * 2 + nh) * 4 + l;
                        sum[idx] += da * dsv[nh][l & 1] * (float)clow[g][nh][l];
                        sum[idx] += dma[l >> 1] * dmv[nh][l & 1];
                    }   /* sum += da·dsv·C_int + dma·dmv — the r15 two-term form, verbatim */

fp32 epilogue (src/cuda_kernels.cu:6434-6444):

    for (int g = 0; g < 4; g++)
        for (int nh = 0; nh < 2; nh++)
            for (int l = 0; l < 4; l++) {
                const int i = i0 + g * 16 + (l >> 1) * 8 + (lane >> 2);
                const int j = j0 + j0w + nh * 8 + (lane & 3) * 2 + (l & 1);
                if (i < nt && j < od)
                    C[(size_t)i * od + j] = sum[(g * 2 + nh) * 4 + l];
            }

Launcher: smem arithmetic + every guard "cleanly returns 0" (src/cuda_kernels.cu:7164-7193; 43,008 is the post-r31 number; at r28 it was 45,056 = sda_q 4,096 B):

extern "C" int launch_mmq_raw_nb_nt(
    int type_id, const uint8_t* w, const uint8_t* q8, float* c,
    int nt, int od, int id, cudaStream_t stream, int kd
) {
    (void)type_id;
    // 64-token x 128-od block tile, KD=8 native. Raw qs plane + single-buffer
    // staging = 43,008 B (r31 q-major sda repack shrinks sda_q 4,096 ->
    // 2,048 B) >>> 2 blocks/SM on GB10. KD!=8 is inapplicable to the
    // raw-nibble variant (the qs plane encodes a FULL 256-k super-block), so
    // clean-fallback (return 0) to the wide kernel. smem/reg guards return 0
    // on any cap failure (never silently launch over cap).
    if (kd != 8) return 0;
    const int smem = 8 * MMQ_NBI * 32   // qa8
                   + 8 * MMQ_NBI * 4    // sda_q (one uint32 per token)
                   + MMQ_NBJ * 128      // qb_raw
                   + 8 * MMQ_NBJ * 8;   // sds (float2 = 8B)
    dim3 grid((nt + MMQ_NBI - 1) / MMQ_NBI, (od + MMQ_NBJ - 1) / MMQ_NBJ);
    cudaFuncSetAttribute(reinterpret_cast<const void*>(&mmq_raw_nb_kernel<8>),
                         cudaFuncAttributeMaxDynamicSharedMemorySize, smem);
    cudaError_t e = cudaGetLastError();
    if (e != cudaSuccess) { cudaGetLastError(); return 0; }
    mmq_raw_nb_kernel<8><<<grid, 256, smem, stream>>>(w, q8, c, nt, od, id);
    e = cudaGetLastError();
    if (e != cudaSuccess) { fprintf(stderr, ...); return 0; }
    return 1;
}

Dispatch gate: r28's before/after (narrow range of git show 0957a08 -- src/cuda.rs):

#![allow(unused)]
fn main() {
// BEFORE: the RAW path had only the wide/narrow tiers
let wide_ok = wide && launch_mmq_raw_wide_nt(...) == 1;
if !wide_ok { launch_mmq_raw_nt(...); }

// AFTER: NB is inserted at the front; on failure each tier falls back cleanly (return 0 = did not run, not an error)
let nb = std::env::var("MINFER_MMQ_RAW_NB").as_deref() == Ok("1");
let nb_debug = std::env::var("MINFER_MMQ_RAW_NB_DEBUG").as_deref() == Ok("1");
// Direction-A NB raw-nibble kernel is KD=8-native; it activates
// only under the full MMQ gate set (MINFER_MMQ=1 + MINFER_MMQ_RAW=1 here).
// launcher returns 0 on KD!=8 or smem/reg cap failure
// -> clean fallback to the wide/narrow raw path below.
let nb_ok = nb && launch_mmq_raw_nb_nt(...) == 1;
if nb_ok && nb_debug { eprintln!("minfer/cuda: mmq raw NB kernel active (KD=8)"); }
if !nb_ok { /* wide → narrow, unchanged */ }
}

(This segment has since evolved further in the current tree: r34's A-transpose prepass path now sits at the front, with plain NB as its fallback arm — src/cuda.rs:3352-3380; the NB-BT variant carries the default path.)

3.3 Pitfalls

  • The top risk was dismantled up front: the nibble-layout mapping was not "tested in passing" inside the kernel — the mapping was derived from the wide kernel's validated ldmatrix B path, given a standalone byte-equivalence check (full comparison over 8 sgs × 32 lanes × 4 regs, 0 mismatches), and only then integrated. r13's 82.896 pattern (layout error → 1e0 garbage diff) was stopped before integration.
  • The register cliff: the inner-loop B-unpack temps + sum[32] risked pushing ptxas to 255 regs + spill (r22 Lever-2's failure mode). Landed at 123 regs / 0 spill — occupancy was not bitten back by registers.
  • The semantics of the KD=4 arm: the raw qs plane = full 256-k super-block; KD=4 does not apply; the parity matrix's KD=4 arm falls back cleanly to the wide kernel. This must be written down, otherwise someone will misread "the KD=4 numbers = wide kernel" as an NB result.
  • The ghost of silent failure lives on: r7's phantom-2124 lesson is baked into the launcher's shape — the attr result is checked, over-cap never launches silently, a returned 0 triggers explicit fallback, plus the MINFER_MMQ_RAW_NB_DEBUG=1 liveness label ("NB kernel active").

4. Verification

Six Phase-2 gates (scheduled by the design doc §11.6, all green) plus one up-front gate:

  • Standalone B-unpack byte equivalence (before integration): full comparison against the wide kernel's ldmatrix B fragments (8 sgs × 32 lanes × 4 regs, 0 mismatches) — defends against nibble-layout mapping errors (the source of the 82.896 pattern).
  • ptxas register gate: -Xptxas -v reports 123 regs / 0 spill — defends against register spill silently killing the 2-blocks/SM goal itself.
  • Parity: NB-active cuda_prefill_mmq max diff 1.5e-5..9.9e-5 — pure f32 rounding magnitude; a layout bug would be ~1e0 (the 82.896 pattern), so the magnitude directly identifies the defect class. The KD=4 arm falls back cleanly and parity is green.
  • greedy-32 identity: byte-identical to the default f16 path — defends against graph-level/numerics-level breakage.
  • Perf (gate 4, the first formal application of the r24 bar): interleaved 3× median, NB 1410.4 vs re-measured baseline 1375.2 = +2.56% (5/5 positive) ≥ +1.5%; the headline number requires distribution separation.
  • ncu occupancy gate (gate 5): sm__warps_active.avg.per_cycle_active 16.17 = ~4.04 warps/sched → 2 blocks/SM; long_scoreboard 2.92 → 2.03 (closing on llama's 1.15); issue_active 25 → 37.36% — the bet's mechanism demonstrably happened. This is the pre-registered kill criterion used in reverse: had occupancy arrived while wall-clock stayed flat, the entire Direction-A line would close.
  • Suite 166/0/3 — defends against regressions.

5. Results

  • Wall-clock: 1375.2 → 1410.4 (+2.56%, 5/5 positive; the earlier batch +2.43%, 3/3) — the first clean crossing of the +1.5% landing bar set by r24.
  • Kernel level: smem 98,304 → 45,056 B; ~2 → ~4.04 warps/sched; long_scoreboard 2.92 → 2.03; issue_active 25 → 37.36%. 123 regs / 0 spill. The two kernels coexist; the wide kernel stays the default, NB is gate-selected.
  • The verdict inverted: r25 said "cutting instructions does not move the wall"; r28 says "the lever that moves the wall is occupancy, the purchase price is smem (the B representation), and the change is +3-6% instructions — and that instruction increment, exactly as the mirror image of r25's verdict predicted, is wall-inert." Methodology doc #77 fixed this arc into a transferable rule: "Buy occupancy, then cut instructions (r13→r25→r28/r29): at 1 block/SM the instruction surplus is real but wall-inert; buy occupancy first, and the same instruction cuts start paying out (+2.6/+2.8%)."
  • The arc closes (foreshadowing r29): the same kd-unroll lever, +0.37/+0.49% at 1 block/SM in r25 (reverted), +2.80% on NB's 2 blocks/SM in r29 (landed) — occupancy unlocks instruction cuts, not the other way around.
  • The NB kernel has been the chassis of the q4_K line ever since: r29 (unroll) and r31 (sda repack) land directly on top of it; r34's A-transpose prepass and the r58/r59 BT variants evolve along its geometry (in the current tree the default path is carried by NB-BT, with plain NB as the fallback arm).

6. Lessons

  1. Occupancy and instruction cuts are an ordering relationship: at 1 block/SM, buy occupancy first; any surgery on the instruction stream waits until occupancy is up — in the wrong order, a good lever measures as noise.
  2. smem is the occupancy currency on consumer-class parts: squeeze the budget from 98 KB to ≤49.5 KB and blocks/SM doubles; what you squeeze is decided by the A/B re-read ratio (2.1:1) — shrink the token dimension, keep the od dimension, sacrifice the smaller stream.
  3. Dismantle layout-mapping-class risk up front with standalone byte equivalence: derive the new mapping from a validated path, run the full comparison (0 mismatches), then integrate — an order of magnitude cheaper than "debugging a 1e0 diff" inside end-to-end parity.
  4. Pre-registered kill criteria let even a null result wind the line down: "occupancy arrives but the wall does not move ⇒ close Direction A" was written before the build, so the measured negative did not trigger wobbly retries.

← 30 · r25 SASS opcode census · Index · 32 · r29 NB kd-loop unroll →

32 · r29 — NB kd-loop unroll: 2 blocks/SM lets integer-ALU pruning move the wall clock for the first time (LANDED)

Result: 7B q4_K whole-prefill 1387.9 → 1426.8 tok/s (+2.80%, interleaved 5/5 positive) — a one-line #pragma unroll change: integer ALU −25% (49.3M → 36.8M warp instructions), total inst −6.5%, 123 regs / 0 spill, the 2-blocks/SM occupancy preserved intact. r25 measured the same pruning at 1 block/SM: +0.37/+0.49%, wall-clock inert. Commit: bfe6bba. Date: 2026-09-04.

1. Background — where things stood

Era C's target had been fixed at r6: lift the q4_K MMQ GEMM from 6.1 TMAC/s to ≥24 (f16-path parity) and ideally ~30 (llama.cpp parity). r25 ran a SASS opcode-class census that reconciled minfer's per- tile warp instruction stream against llama.cpp's: 445,544 vs 355,758 total, a +25.2% surplus — and that surplus is 100% supporting instructions: integer ALU +69.5k/tile (77% of it), fp32 rescale FMUL +15.2k, conversions +14.3k, while IMMA and FFMA are exactly equal on both sides (114,688 FFMA/tile on each). In other words: not one extra instruction computes a MAC; everything extra is "the overhead of carrying the MACs".

At the time, r25 also tried the kd-unroll in passing: integer ALU −38%, total inst −9.7% (the surplus halved), but the wall clock moved only +0.37/+0.49% — below the +1.5% landing bar — and it was reverted. The conclusion then was: the wide kernel's 98 KB smem admits only 1 block/SM, ~2 warps per scheduler, latency completely unhidden — the kernel is issue/occupancy-bound, not instruction- count-bound. The instruction stream was pruned, but the issue slots were never waiting for it.

r28 inverted that verdict: the one occupancy lever r13–r25 had never touched was smem itself. The new Direction-A kernel mmq_raw_nb_kernel squeezes the B-side smem to raw-packed (a 2-nibbles/byte qs plane), accepts a small in-loop B-unpack cost, and drops smem to 45,056 B → 2 blocks/SM. Result +2.56% (1375.2 → 1410.4), with ncu confirming warps_active 16.17 ≈ 4.04 warps/scheduler, long_scoreboard 2.92 → 2.03, issue_active 25 → 37.36%. Occupancy had been bought.

The NB kernel's own pedigree is worth recording: its B-fragment nibble layout was not guessed — r28 derived it from the wide kernel's validated ldmatrix path and ran a standalone byte-equivalence check before integration (all 8 sgs × 32 lanes × 4 registers, 0 mismatches); at landing it is double-gated (MINFER_MMQ_RAW_NB=1 and kd==8, with a clean fallback to the wide kernel for K dims that are not multiples of 8), the wide kernel remains the default raw path, and the two are byte- equivalent. That is, r29's measurement subject is a kernel already wrapped in two independent verifications — instruction pruning no longer needs to worry about layout correctness. That is the precondition that let it "land in one line".

Landing-bar context: the wall bar stays at +1.5%, judged on 7B whole-prefill wall clock. The NB kernel serves the KD=8 shape of q4_K MMQ, so a kernel-level gain must be amplified by its share of the wall to clear the bar — r28's +2.56% showed the amplification factor is sufficient, and same- order kernel-level improvements were worth continuing to mine.

So r29's question became natural: once occupancy is fixed, is r25's "wall-inert" instruction scissors alive again? The census was re-run on the NB kernel first — and the answer still had the original shape: integer ALU 2.85× llama per MAC (2.142 vs 0.751 e-3/MAC), FFMA still level on both sides; stalls concentrate on the shared path (mio_throttle 18.89% + long_scoreboard 16.80%). The same class of surplus, the same lever — it only needed to be verified once.

2. Principle — the GPU mechanism

The NB kernel's K-dimension main loop runs KDR=8 32-token chunks per k-tile (kd = 0..7). Each chunk's iteration body does five things, and four of them have addresses that depend on the runtime value of kd:

  1. A fragments: one ldmatrix.x4 for each of the 4 16-token groups, base qat = qa8 + kd*64*32;
  2. B unpack: read 2 raw 32-bit words from qb_raw; p = sg>>1 (sg = c&7) selects which pair of 16-k groups, is_hi = sg&1 selects low or high nibble;
  3. 8 independent mma chains (4 A-frags × 2 B-frags);
  4. sds scale read: a float4 from sds + kd*128 + ...;
  5. sda scale read: from sda_q + kd*64 + ... (sda/sds are the two sides' A/B scale data planes — A side per-token d|ssum, B side per-row d·sc|−dmin·m, fully defined in r31 — here you only need to know they are shared data read on every iteration).

Without unrolling, all these offsets are runtime integer arithmetic: sg = c&7, p = sg>>1, is_hi = sg&1, the multiply-add chain of the kd * constant bases — every LDS/LDSM address hangs off the end of this integer dependency chain, and every step on the chain is a source of the 77% surplus from r25's census. #pragma unroll lays the 8 iterations out statically, and each chunk's is_hi selection, smem bases, and boundary checks all become compile-time constants: the address arithmetic collapses into immediate offsets and the integer chain disappears; at the same time ptxas can see through, for the first time, that "adjacent chunk pairs read the same pair of raw words" (even kd pairs share the same p), and CSEs the raw-word load across chunks — which is exactly the SWAR equivalent that r30 goes on to test.

Why does the wall clock move this time? The occupancy ledger is unchanged: 256 threads × 2 blocks = 16 warps/SM = 4 warps per scheduler. In the r25 era there were only ~2 warps per scheduler, idle issue slots were the norm, and instruction count was not the bottleneck; now, under 4-warp issue pressure, the supporting instructions themselves start competing for issue slots — every instruction pruned lets a real IMMA/FFMA issue one step earlier. Occupancy is the precondition for instruction pruning; the order cannot be reversed. The unroll also adds no register pressure (the compiler merely folds constants; it does not need more live registers), so 123 regs / 0 spill is preserved intact — that is the guarantee that occupancy is not bitten back.

Dissect one chunk's iteration body into an instruction list and the location of the surplus is obvious. Per chunk per lane: 4 ldmatrix.x4 on the A side, 2 LDS.32 on the B side, 8 independent mmas (the full 4 A-frag × 2 B-frag combination), two LDS.128 for sds (float4, the nh=0/1 od rows), four LDS.64 for sda, plus the rescale FFMA/FMUL pairs. The MAC portion (mma + rescale) is level with llama class by class — the census's own words were "114,688 FFMA/tile, equal on both sides" — and the extra integer ALU is all on the address side: each iteration recomputes the kd-dependent offset for each of the four bases qat/qb_raw/sds/sda, plus the bit ops and boundary checks around sg/p/is_hi. After unrolling, this arithmetic collapses into immediates and disappears from the SASS instruction stream outright.

The issue-slot accounting closes too: the int-ALU class is 77% of the surplus, about ~19% of the whole instruction stream; cutting 25% of it ≈ ~4.8% of the whole stream, the same order as the measured total inst −6.5% (including the knock-on effect of address folding). Under 4-warp-per- scheduler issue pressure, that ~5% release of issue slots lands precisely in the gaps of the IMMA- dense segments — matching the +2.80% wall-clock move.

3. Implementation

3.1 Design choices (the candidate ladder: three vetted, two shot down, one landed)

After the census re-run, r29 arranged the candidate levers into a ladder, every one verified before acting:

  • (a) PRMT nibble extraction — REFUTED. A standalone sm_120 micro-benchmark showed PRMT still needs shift+mask alongside it to extract a nibble and cannot beat the existing SHF+LOP3 combination; the census also showed the kernel was already emitting 0 PRMT. The instruction that looks better on paper does not exist on the hardware.
  • (b) Software pipelining of the B raw-word load — NEUTRAL, reverted. Moving the B raw-word load one stage ahead into double buffering measured wall-neutral, and registers inflated 123 → 177. The occupancy ledger could be computed before the measurement: GB10 has 64K registers per SM; 123 regs × 512 threads (2 blocks) ≈ 63K, already nearly full; 177 regs × 512 ≈ 91K, and the second block would inevitably be squeezed out — even with a neutral wall clock, the register ledger alone is enough to veto. Risking occupancy for a neutral gain: voted down on both counts.
  • (c) LDSM A-fragments — already in place. The r14/r22 legacy; the NB kernel was born with it.
  • LANDED: the kd-loop #pragma unroll — r25's scissors, a one-line change; the census is its source of legitimacy, the zero register increment is its safety margin.

The essential reason for choosing it: it is the only lever on the ladder whose "mechanism is census- confirmed and whose cost is structurally zero". The other two were either falsified by the hardware or carried an occupancy side effect.

3.2 Key code

The change itself is one line (git show bfe6bba -- src/cuda_kernels.cu: 1 insertion). Below is the landed form (current tree src/cuda_kernels.cu), annotated with which quantities become constants because of the unroll:

// src/cuda_kernels.cu — mmq_raw_nb_kernel main K loop (current tree, post-r29)
for (int kt = 0; kt < nktile; ++kt) {
    if (kt > 0) RAW_STAGE_NB(kt);
    __syncthreads();

    #pragma unroll                      // ← r29: this line only (bfe6bba)
    for (int kd = 0; kd < KDR; kd++) {
        const int c = kt * KDR + kd;
        if (c >= nchunk) break;         // boundary guard kept; predicated per iteration after unroll
        const int sg = c & 7;           // ← compile-time constant (kd statically known)
        const uint8_t* qat = qa8 + (size_t)kd * MMQ_NBI * 32;  // ← folded to an immediate offset
        ...

The part that really eats integer ALU is inside the iteration body. After unrolling, the B unpack's p/is_hi and both load bases, the A LDSM G[g] offsets, and the sds/sda kd* offsets are all fixed at compile time:

        // B fragments: raw-nibble in-loop unpack (after unroll, p/is_hi are constants,
        // and the two chunk tiers (2p, 2p+1) read the same word pair → ptxas cross-chunk CSE, see r30)
        {
            const int p = sg >> 1, is_hi = sg & 1, lm3 = lane & 3;
            const unsigned M = 0x0F0F0F0Fu;
            #pragma unroll
            for (int nh = 0; nh < 2; nh++) {
                const int jj = j0w + nh * 8 + (lane >> 2);
                const uint8_t* qs = qb_raw + (size_t)jj * 128;
                const uint32_t* q0 = (const uint32_t*)(qs + p * 32 + lm3 * 4);
                const uint32_t* q1 = (const uint32_t*)(qs + p * 32 + 16 + lm3 * 4);
                uint32_t v0 = *q0, v1 = *q1;
                b[nh][0] = (int)(is_hi ? ((v0 >> 4) & M) : (v0 & M));
                b[nh][1] = (int)(is_hi ? ((v1 >> 4) & M) : (v1 & M));
            }
        }

3.3 Pitfalls

  • The un-warmed ~1206 outlier: the first batch of interleaved measurements produced one outlier at ~1206 tok/s, traced to a GPU power-state artifact under the un-warmed harness (clock ramp not settled); it vanished once warmup was added. From then on the campaign's A/B interleaved protocol always includes warmup — check the measurement environment before suspecting the code.
  • unroll coexisting with break: the loop body contains if (c >= nchunk) break, and #pragma unroll still takes effect (the 8 iterations are laid out statically and the guard is predicated per iteration); the boundary does not need to be rewritten for a "constant trip count" — do not complicate tail-block logic for unrollability's sake.
  • Do the software-pipelining register accounting first: candidate (b)'s threat of 177 regs to 2 blocks/SM should have been visible before measuring; a "neutral + register-inflating" lever is a liability in an occupancy-sensitive kernel.
  • The cost structure of a "one-line change": the diff is 1 line (+#pragma unroll), but the forensics took three rounds — the census re-run, the PRMT micro-benchmark, the software-pipelining trial. This class of lever's real cost is in the burden of proof, not the edit; an unroll without the proof is a gamble, not an optimization.

4. Verification

  • Parity (NB-active) 1/0: cuda_prefill_mmq cross-check passes — defends against the B-fragment mapping being broken by the unroll's reordering (a layout error is ~1e0 scale, f32 rounding is 1e-5 scale; one test tells them apart).
  • greedy-32 byte-identical: the greedy 32-token output is byte-identical — defends against any change in floating-point summation order (the unroll only touches address arithmetic and must not touch summation order).
  • ptxas ledger: 123 regs / 0 spill, warps_active 3.94 — defends against register inflation knocking 2 blocks/SM back to 1, which would destroy this lever's entire precondition.
  • Suite 166/0/3: the campaign-wide full regression gate (the same standard for the r28/r30/r31 steps). The unroll touches only one function in the NB kernel; the suite gates the whole engine's behavior surface — defends against "local optimization, global regression".
  • Interleaved 5-pair (with warmup): +2.80%, 5/5 positive; adjacent pairs +1.88% — defends against single-point noise and machine drift. 5/5 positive has probability 2⁻⁵ ≈ 3% under the "no real difference" null hypothesis — sign-test-level directional evidence, not just a nice-looking mean.

5. Results

Metricr28 baselineafter r29
whole-prefill (interleaved 5-pair median)1387.91426.8 (+2.80%)
integer ALU (warp instructions)49.3M36.8M (−25%)
total inst—−6.5%
regs / spill123 / 0123 / 0 (unchanged)
warps_active (per scheduler)~4.043.94 (2 blocks/SM kept)

The control group is the same lever on r25's wide kernel (1 block/SM): int ALU −38%, total inst −9.7%, wall-clock only +0.37/+0.49%. The same scissors, once occupancy moved 1 → 2 blocks/SM, went from inert to +2.80% — r25's census conclusion ("wall-inert ≠ the class does not matter") and r28's occupancy conclusion converge here: the two levers are not independent items but ordered moves.

One more output that does not enter the comparison table but did enter the follow-up agenda: the cross-chunk CSE that unrolling lets ptxas perform (the B raw words are read-once — loaded once per kd-pair, shared by the lo/hi uses) is clearly visible in the SASS. That became r30's test subject — whether the SWAR word-granular unpack proposal still had headroom (answer: none; see the next doc).

Baseline-convention note: this step's 1387.9 → 1426.8 holds only within the same session window; r31's control baseline reads 1424.10 rather than 1426.8, because machine state drifts across sessions (the master-table reading convention: all A/B numbers are measured interleaved within the same window; absolute values must not be subtracted across windows).

6. Lessons

  1. Occupancy unlocks instruction pruning, not the other way around: before the issue slots are contended, supporting instructions are free; buy occupancy first, then cut instructions — in the wrong order, both come up empty.
  2. The census is the lever's source of legitimacy: reconcile first to confirm which class is over-represented, then operate on that class — what separates r29 from r25 is not better pruning but better occupancy.
  3. Vet the candidate ladder in layers: the micro-benchmark falsified PRMT (the hardware shortcut does not exist), the register accounting vetoed software pipelining (an occupancy side effect), the census supported the unroll — every rejection has a concrete mechanism, not guesswork.
  4. Interleaved A/B must include warmup: the un-warmed outlier was a power-state artifact, not a code signal.

← 31-r28-nb-kernel-2blocks · Index · 33-r30-swar-unpack →

33 · r30 — SWAR unpack: the compiler already did it (REVERTED)

Result: the word-granular SWAR B-nibble unpack measured +1.09% (un-warmed) / +0.54% (warmed) — noise level, below the +1.5% bar; SASS reconciliation +2 SHF/+3 LOP3, everything else identical class by class = behaviorally equivalent machine code. Reverted (cmp-verified = HEAD). r29's kd-unroll had already induced ptxas to perform exactly this CSE — "read the SASS before writing the lever" thereby became this campaign's standard up-front gate. Commit: 0071b31 (record commit, docs-only — the code experiment never entered the tree). Date: 2026-09-04.

1. Background — where things stood

After r29 took +2.80% with a single #pragma unroll, the NB kernel's integer-ALU surplus versus llama still stood at ~2.1×/MAC (1.598 vs 0.751 e-3/MAC — what remained after r29 cut 25%, from a starting 2.85×) — only the top of the instruction-stream mountain had been shaved off. In the MMQ analysis doc's (§11) task list sat a candidate recorded since Direction A was chartered — Task 2: replace the per-chunk raw-nibble unpack with a word-granular SWAR — and its paper ledger was tempting: "one 32-bit word holds 8 nibbles; read once, ~3 bit ops produce both lo and hi copies; LDS count cut by more than half".

Before occupancy was fixed (the r25 era) this candidate was never prioritized: under 1 block/SM both the mio/longsb stalls and the instruction surplus hid inside idle issue slots. Now occupancy had doubled and issue slots were contended, so "issue half as many shared loads" once again looked like real money. r30's plan was therefore the textbook three steps: standalone byte equivalence → kernel integration → interleaved A/B.

But this step added one new gate to the campaign: SASS-first — before touching the integration, cuobjdump -sass the r29 kernel and take the B-raw path apart. That gate is r30's methodological legacy to the rest of the campaign: r32 (staging addressing already hoisted out of the loop by ptxas) and r33 (hybrid inner loop whose ported SASS is byte-identical) both reused the same veto pattern.

One more time skew needs spelling out: Task 2's original ledger in the §11 task list was written against the r28-era kernel shape — the 8 chunks each loaded their own words, 32 LDS.32 per lane per k-tile; word-granular SWAR reads once and produces two copies, saving more than half on paper (the proposal's own words: "4× fewer LDS"). Once r29 landed, that ledger's implicit premise — "the chunks' loads are independent of each other" — had already been eliminated by the compiler. The lever itself did not change; what changed is the thing it was meant to optimize. All of r30's work was quantifying this skew.

2. Principle — the GPU mechanism

SWAR (SIMD Within A Register): use whole-word integer bit ops to process several sub-fields of a word in parallel — here, read one 32-bit raw word (8 4-bit nibbles) and use shift+mask to produce the low-half and high-half nibbles simultaneously, instead of each chunk loading and extracting on its own.

First the paper ledger. The NB kernel has KDR=8 chunks per k-tile; each chunk per lane unpacks 2 B-fragments (the nh=0/1 od rows) × 2 32-bit words:

  • Per-chunk independent loads (naive ledger): 8 kd × 2 nh × 2 words = 32 LDS.32 / lane / k-tile.
  • The key structural fact: chunks 2p and 2p+1 unpack from the same pair of words (same p = sg>>1; only is_hi = sg&1 differs — the low nibble goes to one chunk, the high nibble to the other). There are only 16 unique words.

Word-granular SWAR's selling point is compressing those 32 loads into 16 "read-once, produce-two". But r29's kd-unroll happened to lay the 8 iterations out statically, letting ptxas see through the "adjacent chunks share a word" fact for the first time — it had already performed this CSE. SASS evidence (SM121, cuobjdump -sass):

  • The B-raw path is already 16 × LDS.32 = read-once: each unique word is loaded exactly once;
  • The lo use is a plain LOP3 v&M; the hi use is SHF.R.U32.HI + LOP3 — instruction-for-instruction isomorphic to handwritten SWAR's "read once, shift+mask to produce lo/hi".

In other words, the mechanism the SWAR proposal promised (halved loads + minimal bit ops) is already in the binary. A source-level SWAR rewrite would, at best, make ptxas re-discover the same schedule (SASS unchanged); at worst it would add explicit carry variables and reordered instructions, pushing register pressure and issue count up. It cannot beat "what the compiler already emits" unless ptxas's schedule happens to be suboptimal. r30's value was turning that sentence into a measured fact and hardening it into a rule.

The extraction arithmetic is priced identically on both sides, which is also why the SASS reconciliation was doomed to show only a ±few-instruction residual. Per unique word: the lo use costs 1 LOP3 (v & M), the hi use 1 SHF + 1 LOP3 ((v >> 4) & M); the two unique words per kd-pair total 2×LOP3 + 2×(SHF+LOP3), and the handwritten SWAR version is exactly the same. Load side: both versions issue 16 LDS.32 (8 unique word-pairs × 2 words). Bank behavior is unchanged too — same address set, same conflict distribution. The only degree of freedom between the two versions is instruction scheduling order, and scheduling is precisely ptxas's job.

Why the CSE only appeared after r29: CSE requires the compiler to see both uses of the same word within one visible scope. In the per-chunk loop the two uses belong to def-use chains of different iterations, and ptxas does not merge across iterations; once unrolling spreads the 8 iterations into one basic block, the redundancy is directly exposed and the merge is routine dataflow analysis. Put differently, SWAR's benefit was always a free byproduct of r29's unroll — the proposal simply did not realize it had already landed.

Why shared-instruction count is worth chasing separately in this kernel: LDS/LDSM go through the MIO queue, serialized with the LSU issue slots and the shared-memory bank ports; the post-r29 stall profile (mio_throttle 18.89%) shows the MIO side genuinely backing up. But "which class of shared instruction is backing up" must be attributed class by class — the residual list r30 left at revert time (A-frag LDSM, sda/sds scale reads, staging indexing, epilogue) is the output of exactly that attribution. After the CSE, the B-unpack class is neither the largest family nor a further-compressible one, and its MIO-side suspicion is hereby cleared.

3. Implementation

3.1 Design choices (gate order: verify the map first, then the machine code)

The experiment advanced through three gates; a failed earlier gate stops entry into the next:

  1. Standalone byte equivalence (/tmp/minfer_nb/b_swar_validate.cu): the word-granular variant was cross-checked against the validated ldmatrix B-fragment reference, sweeping all 8 sg × 32 lanes × 4 registers — 0 mismatches. This gate defends against layout errors introduced by "re-deriving the lane/word→fragment byte mapping" (r28's top risk was exactly this).
  2. SASS-first: disassemble the r29 kernel's SASS and count the B-raw path's instruction classes. This gate runs before integration — §2's conclusion comes from here. Strictly speaking, once this gate passed the experiment's fate was sealed: to win, source-level SWAR's SASS would have to issue fewer instructions than "the CSE the compiler already did", which is mechanically impossible. Operationally, the SASS gate is a reproducible procedure: for an SM121 target, cuobjdump -sass the NB kernel, count by opcode class (LDS/LDS.64/LDS.128/LDSM/IMMA/SHF/LOP3…), and reconcile against the expected list — here 16 LDS.32 (B-raw read-once), 64 IMMA, a minimal SHF/LOP3 set. Only a mismatched list leaves room for a source-level lever; when every line matches, the experiment can be judged a loss before integration.
  3. Faithful integration measurement: since the SASS had already ruled, the measurement was still run — in the "faithful b_hi-carry" form (even kd reads the raw words and produces/stages lo/hi; odd kd reuses them), keeping extraction and consumption semantics point-for-point identical, eliminating any claim that "what was measured was not the proposal itself".

3.2 Key code

The SWAR code never entered the tree (cmp-verified after the revert), so there is no commit to cite; as the contrast, what fell back to the current tree is the B-unpack it tried to replace (src/cuda_kernels.cu, the post-r29 unrolled form — note each chunk branches only on is_hi, and the word pair is shared within a kd-pair):

// src/cuda_kernels.cu — mmq_raw_nb_kernel, B fragments (current tree)
{
    const int p = sg >> 1, is_hi = sg & 1, lm3 = lane & 3;
    const unsigned M = 0x0F0F0F0Fu;
    #pragma unroll
    for (int nh = 0; nh < 2; nh++) {
        const int jj = j0w + nh * 8 + (lane >> 2);
        const uint8_t* qs = qb_raw + (size_t)jj * 128;
        const uint32_t* q0 = (const uint32_t*)(qs + p * 32 + lm3 * 4);
        const uint32_t* q1 = (const uint32_t*)(qs + p * 32 + 16 + lm3 * 4);
        uint32_t v0 = *q0, v1 = *q1;                      // ← CSE'd within the kd-pair
        b[nh][0] = (int)(is_hi ? ((v0 >> 4) & M) : (v0 & M));  // LOP3 / SHF+LOP3
        b[nh][1] = (int)(is_hi ? ((v1 >> 4) & M) : (v1 & M));
    }
}

The SWAR proposal's source shape (illustrative, not repository code — reconstructed from the b_hi-carry scheme recorded in §11.10): turn "each chunk reads its own" into explicit cross-chunk word carrying —

// illustrative (never in the tree): even chunks read the words and produce lo/hi, odd chunks reuse b_hi directly
uint32_t v0 = *q0, v1 = *q1;          // once per kd-pair only
uint32_t b_lo0 = v0 & M, b_hi0 = (v0 >> 4) & M;   // carried to the next kd
b[nh][0] = is_hi ? b_hi0 : b_lo0;     // consumption site unchanged

— which is exactly the shape ptxas had already generated after r29's unroll; writing it out explicitly only extends the carry variables' live ranges.

3.3 Pitfalls

  • Gate 1 is necessary but not sufficient: byte equivalence verifies that the mapping is correct (the fragment bytes land where they should); it has no say on "is it faster". r30's lesson is not "SWAR was wrong" but "correctness verification ≠ benefit verification" — the SASS sits between them.
  • The paper LDS ledger's implicit premise: "save half the LDS" assumes each load happens independently; after r29's unroll that premise was already dead. A lever's benefit model must be bound to the current compilation artifact, not to the kernel shape as it stood when the proposal was written.
  • Why the faithful form matters: the pre-revert measurement used a b_hi-carry semantically identical to the proposal, leaving no footing for the "you measured something else" objection — which is what gives the revert its full force.
  • The measurement gate was not skipped after the SASS gate ruled: "the SASS says equivalent" and "the wall clock says no difference" are two independent pieces of evidence, and this campaign wants both: the SASS reconciliation proves the machine code is behaviorally equivalent, the A/B measurement proves the wall clock really does not move. A revert that skips the measurement gate leaves an open case in the "it was actually 0.1% better" scenario.

4. Verification

  • Standalone byte equivalence: full sweep of 8 sg × 32 lanes × 4 regs, 0 mismatches — defends against mapping/layout errors (the r28-class risk; unrelated to benefit).
  • SASS opcode reconciliation (cuobjdump -sass, SM121): LDS/LDS.64/LDS.128/ LDSM = 16/32/16/32, identical item by item; 64 IMMA identical; the only deltas +2 SHF, +3 LOP3; 113 regs / 0 spill — proves the two versions are behaviorally equivalent machine code, so any measured difference can only be noise. Incidentally: the SWAR version's 113 regs is fewer than the incumbent's 123 — registers were never this step's constraint axis, the instruction stream was; "fewer registers" cannot rescue a SASS-equivalent lever.
  • Parity + greedy-32: the integrated build's cross-check and greedy output all green — confirms "what was measured was the equivalent".
  • Interleaved A/B (4-pair, alternating order, round-4 regression): un-warmed +1.09% / warmed +0.54% — both inside the noise band, below the +1.5% bar. Alternating order (AB-BA rotation) controls for bias in the measurement order itself (clock-ramp/thermal-drift directionality); "round-4 regression" means the sequence regressed by the fourth round, further showing the signal had no stable direction — a real +2.80% (r29) is monotonically identifiable from the first pair.
  • Revert gate: cmp-verified — after the revert the working tree is byte-identical to HEAD, ruling out an incomplete revert.

The four gates' division of labor, collapsed into one table:

GateWhat it defends againstr30 outcome
Byte equivalenceMapping/layout errors (the r28-class risk)0 mismatch, pass
SASS reconciliationA do-nothing lever isomorphic to the compilation artifact+2 SHF/+3 LOP3, ruled out
Integration correctness (parity + greedy)Semantic drift introduced by integrationall green
Interleaved A/BPhantom wall-clock gains+0.54%/+1.09%, noise

The ruling was made by the second gate; the last two turned it into "a revert backed by measurement" rather than "an armchair abandonment".

5. Results (REVERTED: the veto mechanism)

Measured: +1.09% (un-warmed) / +0.54% (warmed, alternating-order 4-pair median) — noise level; the SASS reconciliation proves behaviorally equivalent machine code. After the revert = HEAD.

Veto mechanism: when the SASS shows the compiler already emits the instruction stream a proposal promises (here: read-once 16 × LDS.32 + minimal SHF/LOP3), a source-level equivalent rewrite can only add overhead or tread water — there is no "better source shape" left for the compiler to discover, because the endpoint is already occupied. Retry conditions (any one being met makes a retry worthwhile; the common thread is that the trigger is observable in the SASS):

  1. A toolchain upgrade changes ptxas's scheduling strategy — more than 16 LDS.32 reappear on the B-raw path in the SASS (the CSE disappears);
  2. A kernel-shape change (unroll removed, kd-pair structure rearranged) makes "same word, double use" invisible again;
  3. A genuinely non-isomorphic alternative path appears (e.g. an ldmatrix B side) and changes the reconciliation baseline.

Even then the working order remains: re-read the SASS first, then decide whether to rewrite the source — the criterion is the compilation artifact, not the proposal document.

Residual attribution (this step's real output): at the §11.10 revert, the NB kernel's remaining surplus classes were pinned down — the int-ALU surplus of 1.598 e-3/MAC and the shared-path stalls (mio_throttle 15.4%, long_scoreboard 19.0%) come from A-frag LDSM, sda/sds scale reads, staging index arithmetic, and the epilogue — and not from byte-vs-word unpacking. The next item on that list became r31's sda scale-read repack.

6. Lessons

  1. Read the SASS before writing the lever: a source-level equivalent rewrite that is isomorphic to machine code the compiler already emits can only add overhead (the r30 pattern — reproduced by r32 and r33 in succession; three same-pattern vetoes).
  2. Byte equivalence verifies the map, not the gold: 0 mismatches says "runs correctly"; between it and "runs fast" stands the compilation artifact.
  3. Bind the benefit model to the current compilation artifact: when the paper ledger's implicit premise (independent loads) was destroyed by the previous landed step (r29's unroll), the lever had to be re-valued — it could not be advanced on the old books.
  4. A REVERTED step's residual attribution is the next step's signpost: r30's class list fed directly into r31.

← 32-r29-nb-kd-loop-unroll · Index · 34-r31-qmajor-sda-repack →

34 · r31 — q-major sda scale-read repack: a sub-bar positive gain caught by conflict analysis (LANDED)

Result: 7B q4_K whole-prefill 1424.10 → 1439.40 tok/s (+1.07%, median of 45 samples; range +0.49 ~ +2.38) — below the +1.5% bar, but the mechanism is ncu-confirmed (long_scoreboard 24.61% → 21.46%, mio_throttle −0.92 pp), register-neutral (111 → 109 regs / 0 spill), smem 45,056 → 43,008 B — landed as a sub-bar positive gain. SASS: LDS.64 32 → 0, LDS.128 16 → 32 (scale path 48 → 32 conflict-free LDS/k-tile). The first, naive q-major layout had a 2-way bank conflict and reached only +0.57%; conflict analysis caught it and produced the group-region split. Commit: 851a896 (+ 76d495a docs). Date: 2026-09-04.

1. Background — where things stood

r30's revert left behind a pinned residual-attribution list: the NB kernel's integer-ALU surplus versus llama (1.598 e-3/MAC) and the shared-path stalls (mio_throttle 15.4%, long_scoreboard 19.0%) come from four classes — A-frag LDSM, sda/sds scale reads, staging index arithmetic, epilogue. The B-unpack class had been proven to sit at the compiler's floor (the r30 pattern), so the next item on the list was the sda scale reads.

sda is the A side's per-token scale data: each 32-token chunk's q8 quantization block carries an f16 d (the dequant scale) and an i16 ssum (the sum of the block's q8 values), packed together into one 8-byte pair for consumption by the two-term rescale after the mma. The role of ssum deserves a sentence: a q8 block's dot product splits into two terms, "the q8-value contribution + the block-sum contribution"; the latter is multiplied by the A-side scale (dma = da * sa, sa being ssum) and combined with the B-side min term (dmv = −dmin·m) to form the correction term of r15's two-term rank-1 fold. That is why each token's (d, ssum) must be available as a pair at the rescale point — they are the inputs to two scalar corrections of the same mma accumulator value, and any reordering on the read side must preserve that pairing. The r28/r29-era layout was one uint2 (8 B) per token; the read side issued 4 LDS.64 per chunk (one per token group) — 32 LDS.64/lane per k-tile across the 8 chunks — the largest family of shared instructions on the scale path.

The SASS-first gate established at r30 "passed" a lever for the first time here instead of vetoing it: cuobjdump -sass showed ptxas had not merged those 4 LDS.64 (each kd issues its own at a 0x40 stride) — the opposite of r30's B-unpack (where the compiler had done everything); here was a genuine compiler blind spot, so the lever had footing. The motivation was twofold: replace 32 narrow loads with 16 wide ones (saving MIO issue slots), and land every load on a conflict-free bank distribution (compressing the stall class).

The other piece of context is landing-bar politics: after r29 the wall bar is +1.5%. sda reads are only a few percent of the kernel's instruction stream, so this lever would most likely miss the bar — whether to land a measured sub-bar result became the question this step had to answer head-on.

2. Principle — the GPU mechanism

sda/sds distinction: sda = the A side's per-token packed (d f16 | ssum i16) plane; sds = the B side's per-(chunk, od-row) (d·sc | −dmin·m) float2 (the B-side term of r15's two-term rank-1 rescale). This step touches only sda — sds reads are already at maximum width (LDS.128, float4), at their own floor.

Old layout (the r28 original): sda_q holds one uint2 per token, word index kd*128 + g*16 + q*2 + half (g = token group 0..3, q = within-group pair 0..7, half = the 8-token half-group 0..1). The read side issues one LDS.64 per (kd, g) reading 2 adjacent words. The raw material for wide loads was all there: the same lane's 4 groups' words sit 16 B apart — ptxas failed to see it.

First, pin down the old layout's "charge sheet" precisely: its LDS.64 is not conflicted. Within one phase of an LDS.64, the 16 lanes present only 4 unique 8-B words (q = lane>>2; the 4 lanes of a group broadcast), 4 broadcasts per phase, zero conflicts. The problem is purely instruction count — each load moves only 8 B, 32 of them per k-tile; the wide-load merging opportunity (the same lane's 16 B adjacency across groups) was always there, but the per-g read order gave ptxas no foothold. So this step's benefit model is "instruction-count halving first, conflict removal second" — which determines the direction of the region-split derivation below.

Naive q-major, first version (the rejected shape): lay out the 4 groups' words each lane needs, [q][g0..g3], contiguously — 32 B per lane, 32 B stride within a warp. Issuing LDS.128, one phase (8 lanes) reads 8 16-B words at byte offsets 0, 32, 64, …, 224:

bank(word start) = (byte_offset/4) mod 32
  → 0, 8, 16, 24, 0, 8, 16, 24
  → 4 bank quads each hit by 2 words = 2-way conflict

A 32 B stride is exactly 1/4 of the bank space (128 B); 8 words stomp 4 quads twice each — every LDS.128 splits into two waves, and the wide-load benefit is cut in half. Measured: only +0.57%; the conflict analysis explains why.

Group-region split (the landed shape): the word index becomes kd*64 + rg*32 + q*4 + gsel*2 + half, where rg = g/2 (the 4 groups split into two 32-word regions) and gsel = g&1. Within a region, q*4 gives a 16 B stride and gsel*2 + half gives the 4 word slots inside a 16 B word-group. A lane reads its full per-chunk (d|ssum) set with two LDS.128s (s0 = groups 0,1; s1 = groups 2,3): each instruction has the warp read 8 unique 16-B words at byte offsets 0, 16, 32, …, 112 → banks 0-3, 4-7, …, 28-31 — each of the 32 banks exactly once, zero conflicts. The write side is just as clean — one staging warp decomposed by lane (g = lane>>4, half = (lane>>3)&1, q = lane&7) produces this bank sequence for its 32 4-B writes:

lane  0.. 7 → gsel=0, half=0 → word slot q*4   → banks 0,4,8,12,16,20,24,28
lane  8..15 → gsel=0, half=1 → word slot q*4+1 → banks 1,5,9,13,17,21,25,29
lane 16..23 → gsel=1, half=0 → word slot q*4+2 → banks 2,6,10,14,18,22,26,30
lane 24..31 → gsel=1, half=1 → word slot q*4+3 → banks 3,7,11,15,19,23,27,31

— 32 banks hit exactly once each; the whole warp's writes complete in a single instruction, single transaction. Both the read and write sides are conflict-free.

Semantic invariant: only the storage order moves; the math does not. Every word's value is unchanged (d | ssum<<16), the consumption sites are unchanged (still used per (g, half) in r15's two-term fold), and the r22 qa8 swizzle is untouched — so the correctness gate can demand bit-identical results.

The occupancy side effect is positive: sda_q shrinks from 4,096 B to 2,048 B, total smem 45,056 → 43,008 B; 2 blocks/SM is kept with a thicker margin; registers 111 → 109 (the per-g address arithmetic becomes one uint4 read + constant selection). The smem ledger can be re-verified item by item: qa8 8×64×32 = 16,384 B, sda_q 2,048 B, qb_raw 128×128 = 16,384 B, sds 8×128×8 = 8,192 B — total 43,008 B, matching the kernel comment and the launcher's formula; substituting the old sda_q's 4,096 B back in gives exactly the pre-change 45,056 B.

3. Implementation

3.1 Design choices (two layouts, one conflict analysis)

  • Why naive q-major was tried first: [q][g0..g3] contiguous at 32 B/lane is the most intuitive arrangement of "one lane's data together", and the write-side indexing is simplest. It fails at the warp dimension: layout correctness is per-lane, bank behavior is per-warp — the intuitive arrangement buries the warp conflict inside its 32 B stride.
  • Conflict analysis before integration: the bank re-check was done only after the naive version measured +0.57% (positive but suspiciously weak), which found the 2-way conflict. Run in the other order — analyze first, measure second — the lesson would have saved one integration.
  • The region-split derivation direction: the target read order is "an LDS.128's 8 unique words spread across all 32 banks"; working backwards, the word layout must have a 16 B intra-warp stride; 16 B × 8 = 128 B = one region; 4 groups do not fit → split into two regions (rg = g/2), each LDS.128 handling two groups.
  • Two LDS.128s, not one: each lane needs 8 words per chunk (4 groups × 2 halves) = 32 B — exactly two LDS.128s; the naive version also issued two. The only difference is the intra-warp stride (32 B → 2-way conflict, 16 B → zero conflict). Identical instruction count, different address pattern — the entire wall-clock difference comes from the conflict waves. This is the minimal specimen of "layout is performance": the instruction counters see no difference; the bank analysis and the wall clock do.
  • sds untouched: it is already at the LDS.128 floor; touching it would be a lever with no mechanism (the r30 pattern applied preemptively).

3.2 Key code

Write side (inside the staging macro, src/cuda_kernels.cu current tree; in git show 851a896 changed from uint2/token-pair tiling to region-split):

// before r31 (a − line of 851a896): uint2 per token, g-major
// *(unsigned*)(sda_q + ((size_t)kd * MMQ_NBI + (r >> 4) * 8 + (r & 7)) * 2
//              + ((r >> 3) & 1)) = dv[i] | (sv[i] << 16);

// after r31 (current tree): region-split, the warp's 32 writes cover all 32 banks once each
for (int i = 0; i < 2; ++i) {
    const int x = threadIdx.x + i * 256;
    const int r = x & (MMQ_NBI - 1), kd = x / MMQ_NBI;
    const int g = r >> 4, t = r & 15, q = t & 7, half = t >> 3;
    const int rg = g >> 1, gsel = g & 1;
    /* conflict-free: region=g/2 block, q*16B stride, gsel*8B */
    *(unsigned*)(sda_q + (size_t)kd * MMQ_NBI
                  + rg * 32 + q * 4 + gsel * 2 + half) =
        dv[i] | (sv[i] << 16);          // value unchanged: d f16 | ssum i16
}

Read side (two LDS.128s per chunk replace four LDS.64s; current tree):

// before r31: one LDS.64 per (kd, g), 32 per k-tile
// const uint2 pk2 = *(const uint2*)(sda_q
//     + (size_t)kd * MMQ_NBI * 2 + g * 16 + (lane >> 2) * 2);

// after r31: s0 = groups 0,1; s1 = groups 2,3 — 16 B warp stride, zero conflicts
const uint32_t* sda_blk = sda_q + (size_t)kd * MMQ_NBI
                          + (size_t)(lane >> 2) * 4;   // q*16B stride
const uint4 s0 = *(const uint4*)(sda_blk);       // rg=0: g0h0,g0h1,g1h0,g1h1
const uint4 s1 = *(const uint4*)(sda_blk + 32);  // rg=1: g2, g3, each half
#pragma unroll
for (int g = 0; g < 4; g++) {
    float da_q[2]; int sa_q[2];
    const unsigned w0 = g == 0 ? s0.x : (g == 1 ? s0.z : (g == 2 ? s1.x : s1.z));
    const unsigned w1 = g == 0 ? s0.y : (g == 1 ? s0.w : (g == 2 ? s1.y : s1.w));
    da_q[0] = h2f((unsigned short)(w0 & 0xFFFF));    // consumption identical point-for-point to the old layout
    sa_q[0] = (int)(short)(w0 >> 16);
    da_q[1] = h2f((unsigned short)(w1 & 0xFFFF));
    sa_q[1] = (int)(short)(w1 >> 16);
    ...
}

3.3 Pitfalls

  • Layout-correct ≠ warp-correct: every lane of the naive q-major got the right data; what was broken was the warp-level bank distribution of the 32 B stride. A shared-memory layout review must do both layers: the per-lane value mapping + the per-warp bank trace of the accesses.
  • Positive but suspiciously weak = a mechanism problem signal: a number like +0.57% ("right direction, limping magnitude") deserves the question "why"; only after the conflict analysis answered it did region-split reach +1.07%. Without asking, a real lever gets sold at half price.
  • Handling a flaky suite: this round's suite once came out 164/2 flaky; it was recorded only after a rerun came back all green (166/0/3) — a flaky run is rerun-confirmed, neither counted nor ignored outright.
  • Shrink one region, fix every derived pointer: after sda_q went from uint2/token to uint32/token, the qb_raw base formula had to change from sda_q + KDR * MMQ_NBI * 2 to * 1 in lockstep, and the launcher's smem formula from * 8 → * 4 — miss any one of them and the staging writes overrun/corrupt the adjacent plane (off by exactly one region; parity will certainly explode, but this class of error is best caught in diff review, not waiting for the cross-check). "Shrinking one smem region" is the classic three-site coupled change.
  • uint4 alignment is a property the layout gives you: the new read side is a uint4 load and needs 16 B natural alignment; the region split's q*4-word offset (= q×16 B) provides exactly that. The old layout needed only 8 B alignment; switching to uint4 reads directly on the old indices would send half the accesses across 16 B boundaries — alignment constraints belong in the layout design, not left to runtime.
  • ptxas's blind spots are selective: at r30 it had already done the B-unpack CSE; at r31 it did not merge the sda loads (the 0x40 stride was right there) — both "the compiler already did it" and "it didn't" must be verified point by point in SASS, never extrapolated by intuition.

4. Verification

  • Parity (NB-active) 1/0: the strongest gate for a layout reorder — any (g, q, half) misalignment is ~1e0 scale while f32 rounding is 1e-5 scale; one test tells them apart.
  • greedy-32 byte-identical: the greedy 32-token output is byte-identical — confirms "same values, same consumption sites"; the rescale math and summation order were untouched.
  • SASS reconciliation: LDS.64 32 → 0, LDS.128 16 → 32 (scale path 48 → 32 per k-tile, all conflict-free layout) — the mechanism cashes out at the compilation-artifact level.
  • ncu: long_scoreboard 24.61% → 21.46%, mio_throttle 16.53% → 15.61% (−0.92 pp) — the stall class really was compressed, and no new stall class appeared.
  • Resource ledger: 109 regs / 0 spill (−2), smem 43,008 B — occupancy stays 2 blocks/SM with no regression.
  • Suite 166/0/3 (including one 164/2 flaky rerun-confirmed).
  • 45-sample large-N A/B: a small effect (+1.07%) is unresolvable at the 5-pair protocol's scale (r29's protocol size); only a 45-sample median + range (+0.49 ~ +2.38) could lift the signal out of the noise band.

5. Results

Metricr29 baselineafter r31
whole-prefill (45-sample median)1424.101439.40 (+1.07%)
sda reads (SASS, per k-tile/lane)32 × LDS.6416 × LDS.128 (0 LDS.64)
long_scoreboard24.61%21.46%
mio_throttle16.53%15.61%
regs / spill111 / 0109 / 0
smem45,056 B43,008 B (sda_q 4,096 → 2,048)

+1.07% is below the +1.5% bar; the ruling to land it rests on: (1) the mechanism is doubly confirmed by ncu and SASS — it genuinely eliminated an entire stall-contributing class (32 narrow scale reads) rather than being coincidental positive noise; (2) a zero-regression surface — registers, smem, and occupancy are all neutral or better; (3) the direction is stackable — the scale-read class belongs to the same family as the later r35 (sds predecode, REVERTED), and this step's floor is that step's starting point. The ruling is recorded in the master table's "LANDED (sub-bar)" status: a sub-bar but mechanism-confirmed positive gain may be kept when it "compresses some stall class with no regression" — the bar-decision itself must be recorded explicitly, not left for posterity to excavate.

Two post-hoc notes. First, the 43,008 B smem thickened the 2-blocks/SM margin, but the NB kernel was never pushed to 3 blocks afterwards — the q6_K line's r40 later proved the 3rd resident block is bought with __launch_bounds__ register trade-offs, not smem subtraction; the NB 2-block equilibrium held until the campaign's end. Second, an effect of +1.07% magnitude established the 45-sample median protocol — every later sub-bar candidate (r32/r33 etc.) used the "large sample + mechanism confirmation" double referee, with the 5-pair protocol reserved for candidates above +2%.

6. Lessons

  1. A shared layout passes two reviews: the per-lane value mapping (parity's job) + the per-warp bank trace (conflict analysis's job) — naive q-major lost at the second review, and only a suspiciously weak measurement exposed it.
  2. The SASS-first gate is bidirectional: at r30 it vetoed a lever the compiler had already exhausted; at r31 it passed a real lever inside a compiler blind spot — the gate's value is turning "compiler behavior" from guesswork into evidence.
  3. Landing sub-bar requires the trio: mechanism confirmation (ncu/SASS) + zero regression (regs/smem/occupancy) + an explicitly recorded bar decision; missing any one, a sub-bar positive gain should be reverted.
  4. Small effects need large samples: a wall-clock effect of ~+1% is unresolvable inside the 5-pair protocol's noise band; a 45-sample median + range is the right measuring instrument for this class of lever.

← 33-r30-swar-unpack · Index · 35-r32-finite-lever-sweep →

35 · r32 — The finite lever sweep: two regions fenced off (REVERTED)

Result: a SASS region census of the post-r31 NB kernel (2,617 instructions total) fenced off both of the integer-ALU surplus's remaining "cuttable candidates" — the staging's kt-independent addressing was already hoisted into the prolog by ptxas (source-level lever dead, no measurement needed), and the epilogue write-back widening measured +0.46% (noise) with a structurally capped dynamic share of ~0.4%. Both reverted, cmp-verified = HEAD. The NB kernel's integer-ALU surplus thereby reached the compiler floor. Commit: 153d28c (docs-only record commit; both code experiments were completed, measured, and reverted in the local tree — never landed as code commits). Date: 2026-09-04.

1. Background — where things stood

After r28 swapped in the raw-nibble NB kernel (2 blocks/SM, +2.56%), this line ate two positive gains in a row: r29's kd-loop unroll (+2.80%, integer ALU −25%, total instructions −6.5%) and r31's q-major sda scale-read repack (+1.07%, 1424.10 → 1439.40 tok/s, longsb 24.61 → 21.46%). One side conclusion of r29 deserves singling out: the purely instruction-count cuts that were "wall-inert" in the r25 era started paying out at 2 blocks/SM — occupancy unlocks instruction cuts, not the other way around. That kept "keep hunting cuttable instructions" reasonable on September 4.

But the levers were visibly thinning. r31 itself was a sub-bar landing below the +1.5% bar (the bar was calibrated by r24's scheduling-ladder experiment: wall-clock changes below it are indistinguishable from interleaved A/B noise) — mechanism real, magnitude already brushing the top of the noise band. The earlier r30 was a sharper warning: the SWAR word-granular unpack measured +0.54% (noise), and the SASS comparison showed r29's unroll had long since induced ptxas to CSE every raw word — writing in source a version "the compiler already generates" can only add overhead.

By the time r32 started, the books left by r30/r31 read: integer-ALU surplus 1.598 e-3/MAC, about 2.1× llama's (0.751 e-3), attributed three ways — A-frag LDSM (claimed irreducible), staging index math, epilogue. r32's problem statement was deliberately modest: is there any source-level cuttable component left in this surplus? If no, the "instruction count" axis closes as a whole and the next round of hypotheses (r33's loop-organization theory) gets a clean start. The method was a finite lever sweep: a SASS region census to apportion the kernel's instructions, then one decisive experiment per remaining candidate — falsify what can be falsified, cap what can be capped.

The archival situation matches doc 15's r10: git show 153d28c --stat contains only docs/CUDA_OPTIMIZATION.md +49 lines and docs/LLAMA-CPP-MMQ-ANALYSIS.md +32 lines — the code changes were reverted the same day after measurement, so the numbers and SASS evidence exist only in the record commit. Every piece of "current-tree code" cited here is a before form that survived the revert.

The master table's row 46 gives the verdict: "staging addressing already hoisted by ptxas; run-once epilogue cannot clear a bar" — two regions, two kinds of death (proven dead by SASS vs capped by measurement), demonstrating both forms of "fencing off".

2. Principle — the GPU mechanism

The region census — apportioning 2,617 instructions. The toolchain is r25's SASS census (cuobjdump -sass), but classified by execution region instead of opcode — r25 answered "which instruction types are over-represented", r32 answers "where do the extra instructions live". The post-r31 mmq_raw_nb_kernel<8>'s 2,617 instructions split into four segments (the four regions total 2,585; 32 strays):

RegionCountExecution frequencyContents
prolog + stage0500once per blockindex base materialization, first k-tile's staging (incl. stage0's staging share)
per-kt in-loop staging455once per k-tileA-side LDG batch + swizzle STS + sda repack (~21% — the "staging = 21% of kernel instructions" cited later in r34)
per-kt compute-kd1,515once per k-tileldmatrix A-frags + B unpack + 64 IMMA + rescale (the hot path)
run-once epilogue115once per block32 scalar STG write-backs + guards

Statically staging is 455/2617 ≈ 17%; adding stage0's staging share gives the recorded ~21% — the largest nominally "possibly cuttable" block on the books.

Region → lever mapping. The census's value is turning "where else can we cut" into a finite list: prolog/stage0 is the product of hoisting, "optimized" by definition; compute-kd is the hot path just harvested by r28–r31 (ldmatrix A-frags = the claimed-irreducible A-frag LDSM; B unpack = proven compiler-CSE'd at r30; 64 IMMA = the campaign's reason to exist; rescale = r15's fp semantics contract); staging's index math and the epilogue are the only two regions that "look untouched" — hence r32's two candidates.

Swizzle/repack — the staging region's two protagonists. The index math r32 examined is not casual: r22's XOR swizzle ((((R&3)<<1 + (u>>2)) ^ ((R>>2)&7)) << 4) lands each 8×8 ldmatrix tile's 8 rows on mutually conflict-free bank groups — a conflicted LDSM would turn "read A-frags" into a serialization hotspot (r22 once precomputed all 8 tile offsets to zero the address ALU per ldmatrix); r31's q-major region-split repack converges each warp's per-chunk scale reads into two LDS.128s (8 16-B words, 16 B stride, zero conflicts). The staging region's instruction count is the price paid for conflict-free consumption — r32 proves that price cannot be cut further at source level (ptxas has hoisted everything hoistable), and r34 will prove it can be moved wholesale.

Dynamic-share arithmetic for run-once regions. Static count is not dynamic share: the epilogue runs once per block, staging + compute once per k-tile. For the 7B GEMM with hidden = 3584: nchunk = id/32 = 112, KDR = 8 → nktile = 14. The dynamic share is

115 / (115 + (455 + 1515) × 14) = 115 / 27,695 ≈ 0.42%

That is where the record's "epilogue ~0.4%" comes from (prolog/stage0 likewise). The implication is hard: deleting a run-once region entirely has a theoretical ceiling below 0.5% — it can never touch the +1.5% bar. The epilogue experiment knew this ceiling from the start — it was measured to get one clean data point and to verify the premise "ptxas really does not vectorize scalar STGs".

ptxas's hoist mechanism: uniform registers. The evidence chain is SASS-level. The A-side staging address has two parts: a kt-independent term (the A-token base (i0+r)*nb32) and a kt-dependent term (the k-tile stride). ptxas materializes the former once in the prolog and loads the latter into a uniform register (UR — a per-warp bank of uniform scalar registers, present since Volta, carrying loop invariants identical across the warp, outside the per-thread register file); the in-loop STS target reads STS [R57+UR11+0x400..]: R57 is the prolog-computed base, UR11 carries the only per-kt term, 0x400 is a constant offset. The prolog's and the loop body's STS targets are byte-for-byte isomorphic — the SASS definition of "the compiler already did it". The av loads likewise merge into 3 base registers + immediate offsets (4 + kd*40 — in the native pad40 chunk qs starts at byte 4, with a 40 B kd stride). The remaining per-kt addressing is intrinsic (the k-tile is genuinely varying), not a hoistable index chain.

The minimal SASS forensics workflow. r32's forensics loop, recorded verbatim: cuobjdump -sass <binary> exports the target kernel's disassembly → split the instruction stream into regions by MMQ.-prefixed labels or register usage (the prolog's signature is "index math executed once"; the loop body is delimited by BRA/labels) → per region, count instructions and inspect operand sources (UR* uniform registers = hoisted loop invariants). The whole "staging is hoisted" ruling took under half an hour from export to reading — replacing a pointless round of rewrite + compile + measure.

Why store widening stops at float2. The mma m16n8k32 C-fragment lane map fixes the row/col coordinates of the 8 output points each thread holds: for each (g, nh) combination the thread has one point at row iA and one at iA+8, with adjacent columns (j, j+1) (l&1 is the column low bit; see §3.2). So per combination a thread can assemble one pair of adjacent-column 8 B float2s (one pair per row) — STG.64 available; the two float2s are a full row apart, no contiguous 16 B pair, STG.128 not available. Alignment is not the problem: j0w + nh*8 is a multiple of 8 and (lane&3)*2 is even, so the pair start is always an even column and 8 B alignment holds. 8 B is this fragment geometry's physical ceiling, and boundary blocks (od/nt not divisible by 128/64) must keep a scalar tail. As for "why doesn't ptxas do it automatically": vectorizing scalar stores requires proving the two STGs' addresses are adjacent, aligned, and side-effect-free in between — the induction across l iterations is not free for ptxas, and here it chose conservatism.

3. Implementation

3.1 Design choices (why this shape and not another)

The two candidates were vetted in different orders. Staging (455 instructions, ~21%) is the biggest block on the books, but r30's lesson is forensics before action — it took the pure SASS-forensics route, concluded source-level no-op, and wrote no code at all. The epilogue was measured even though its ~0.4% dynamic cap was known: it is the cleanest vehicle for "halving the STG count", one experiment answering two questions at once (does ptxas really not vectorize? how much does the widened guard ALU cost?) — both answers matter for every later kernel, and the 0.4% ceiling made it a low-risk probe.

The widening experiment's shape was float2 interior + scalar tail: divisible interior tiles take the STG.64 path, boundary blocks keep the scalar path — trading one runtime branch (dual-path) for halving the interior's store count. That shape choice is itself one of r32's propositions under test: in a run-once region, is a branch-for-width trade worth it?

3.2 Key code

First the staging side — the complete object the SASS proved "ptxas has already hoisted". RAW_STAGE_NB's A side has three phases (r20 split-phase: LDG batch → scale reads → STS write-back):

// src/cuda_kernels.cu:6242-6253 (RAW_STAGE_NB's LDG batch — the before form, still alive today)
_Pragma("unroll")
for (int i = 0; i < KDR * 2; ++i) {
    const int x = threadIdx.x + i * 256;
    const int u = x & 7, r = (x >> 3) & (MMQ_NBI - 1),
              kd = x / (8 * MMQ_NBI);   /* KDR=8, 8*NBI = 512 */
    const int tok = i0 + r, c = (kt) * KDR + kd;
    unsigned v = 0;
    if (tok < nt && c < nchunk)
        v = *(const unsigned*)(q8x                     // ← native pad40: 40 B per (token, chunk)
            + ((size_t)tok * nb32 + c) * 40 + 4 + u * 4); // base (i0+r)*nb32 + kt term + 4 + u*4
    av[i] = v;                                          //   — SASS: 3 base registers + immediate offsets
}

In source, every address explicitly contains kt; the SASS ruling is that the base term is materialized in the prolog and the kt stride term goes into UR11. Next the STS write-back — where r32 had hoped to find a lever on this index chain:

// src/cuda_kernels.cu:6268-6277 (RAW_STAGE_NB's qa8 write-back — A-side XOR swizzle)
_Pragma("unroll")
for (int i = 0; i < KDR * 2; ++i) {
    const int x = threadIdx.x + i * 256;
    const int u = x & 7, r = (x >> 3) & (MMQ_NBI - 1),
              kd = x / (8 * MMQ_NBI);
    const int R = kd * MMQ_NBI + r;
    *(unsigned*)(qa8 + (size_t)(R & ~3) * 32                       // ← group base: the kt-independent part
        + (size_t)(((((R & 3) << 1) + (u >> 2))                    // ← XOR swizzle index math
                    ^ ((R >> 2) & 7)) << 4)
        + (size_t)(u & 3) * 4) = av[i];
}

Intuitively R contains both kd (per-kt) and r (kt-independent) — an index chain recomputed every iteration. The SASS ruling: R's kt-independent component is materialized in the prolog, the per-kt component goes into UR11, and the XOR/shift parts survive instruction-for-instruction but with hoisted operands — everything hoistable has been hoisted; what remains is intrinsic per-kt addressing, and a source-level staging rewrite has no actionable lever.

The sda-side repack write (r31's region-split formula) belongs to the census's staging region too:

// src/cuda_kernels.cu:6278-6288 (RAW_STAGE_NB's sda write — r31 q-major region split)
_Pragma("unroll")
for (int i = 0; i < 2; ++i) {
    const int x = threadIdx.x + i * 256;
    const int r = x & (MMQ_NBI - 1), kd = x / MMQ_NBI;
    const int g = r >> 4, t = r & 15, q = t & 7, half = t >> 3;
    const int rg = g >> 1, gsel = g & 1;
    /* conflict-free: region=g/2 block, q*16B stride, gsel*8B */
    *(unsigned*)(sda_q + (size_t)kd * MMQ_NBI                // ← kd term: per-kt (goes into UR)
                  + rg * 32 + q * 4 + gsel * 2 + half) =     // ← region-split: kt-independent (hoistable to prolog)
        dv[i] | (sv[i] << 16);                               //   f16 d | i16 ssum packed into one u32
}

(This packing formula is reused byte-for-byte in r34's prepass — see doc 37.)

The epilogue's before form still lives in the current tree (the widening was reverted) — mmq_raw_nb_kernel's write-back, 4 g × 2 nh × 4 l = 32 (i,j) points, one scalar STG.E per point:

// src/cuda_kernels.cu:6434-6444 (mmq_raw_nb_kernel epilogue — the before of r32's widening experiment)
#pragma unroll
for (int g = 0; g < 4; g++)
    #pragma unroll
    for (int nh = 0; nh < 2; nh++)
        #pragma unroll
        for (int l = 0; l < 4; l++) {
            const int i = i0 + g * 16 + (l >> 1) * 8 + (lane >> 2);   // row: l-high bits + lane-high bits
            const int j = j0 + j0w + nh * 8 + (lane & 3) * 2 + (l & 1); // col: (lane&3)*2 + l-low bit
            if (i < nt && j < od)
                C[(size_t)i * od + j] = sum[(g * 2 + nh) * 4 + l];    // ← 32 scalar STG.E
        }

Read the column map: (lane & 3) * 2 + (l & 1) — the same lane's l=0/1 points are adjacent columns (stride 1), so each pair merges into an 8 B float2; l=2/3 differ in row by 8 ((l>>1)*8), another pair in the other row. The widened version (never committed) added a divisibility branch for the interior: SASS result 32 STG.E → 16 STG.E.64 + 24 STG.E — stores really did halve, but the dual-path guard pushed static integer ALU up (IMAD 55 → 74, LEA.HI.X 8 → 24), at ptxas 111 regs / 0 spill.

3.3 Pitfalls

  • "Static instruction count down" is not the objective function. In the widening experiment int ALU actually rose (the dual-path guard), and even had it fallen, run-once instructions are irrelevant to the wall — r25's wall-inert conclusion was only ever inverted for the per-kt hot path × 2 blocks/SM combination (r29); the epilogue is not on the hot path.
  • A store widening's ceiling is decided by the fragment lane map, not by desire. STG.128 needs the thread to hold 16 contiguous output bytes, and the mma C-fragment's row distribution (iA/iA+8) excludes that from the start. Draw the lane map before setting the widening target and you save an entire experimental round.
  • A census's region split must fix execution frequency first. The same static instruction carries a dynamic weight differing by nktile (=14) between a "once per block" and a "once per kt" region — classifying by opcode (r25) cannot see this; classifying by region can.
  • The forensics duty when experiment code is never committed. Same as r10: post-hoc re-inspection of the widened source is impossible; what can be re-inspected is only the SASS numbers in the record commit and the before form in the current tree. A docs-only commit must describe "what changed" well enough that a reader can mentally reconstruct it.

4. Verification

  • SASS region census (cuobjdump): defends against fake levers — see what the compiler already generates before acting; the r30 pattern made institutional.
  • SASS comparison (widening experiment): 32 STG.E → 16 STG.E.64 + 24 STG.E plus the IMAD/LEA counts — confirms the change touches only the write-back region and quantifies the guard's ALU cost.
  • Parity 1/0 + greedy-32 byte-identity: defends against "the widened stores corrupting output / boundary blocks written wrong" — the dual-path branch is exactly where boundary mistakes hide.
  • Interleaved 4-pair A/B (baseline 1441.5 → 1448.15, alternated within one window): defends against co-tenant drift reading +0.46% of noise as signal.
  • cmp-verified = HEAD revert check: confirms the experiment tree is byte-identical to HEAD — the negative result carries no residue.

5. Results

The staging lever: dead at the source level, no measurement needed. SASS proves the A-token base is materialized in the prolog, the per-kt term lives in a uniform register, and the av loads merge into 3 bases + immediate offsets — a source-level staging rewrite is a no-op. This is the second instance of r30's "compiler already did it" pattern, this time confirmed by SASS rather than source intent.

The epilogue lever: measured, capped, reverted. The SASS store count halved as predicted (32 → 16×64-bit + 24 scalar), the guard ALU rose (IMAD 55 → 74, LEA.HI.X 8 → 24); the interleaved 4-pair measurement gave +0.46% (1441.5 → 1448.15) — inside the noise band, and the ~0.4% structural ceiling (the §2 dynamic-share arithmetic) means it can never clear the bar. Reverted, cmp-verified = HEAD.

The total ledger. The NB kernel's remaining integer-ALU surplus = A-frag LDSM (irreducible) + intrinsic sda/sds scale decode + fp rescale (FFMA/FMUL/I2FP already at parity/deficit) — there is no source-level integer-ALU cut left.

The lever's reincarnation. The "staging instructions cannot be cut" death sentence is valid only for source-level rewrites: the next day, r34 moved the staging's entire transform component out of the kernel (a quantize prepass pre-transpose), and that 21% of index math vanished wholesale in the bt kernel — the bottleneck was not cut away, it was relocated. r32's census (the two numbers: staging 21%, epilogue ~0.4%) is exactly the map r34 used to aim that cut (doc 37).

Veto mechanism and retry conditions: a run-once region's dynamic share = static count / (per-kt count × nktile); in a GEMM kernel with nktile ≫ 1 this quotient is always under 1%. Only when the write-back itself becomes a per-tile hot path (e.g. f16 written directly to C, or a tile-geometry change making the epilogue scale with kt) is it worth revisiting this lever. The staging side has exactly one retry condition: a shape that changes the staging's instruction composition (not its count) — precisely the direction r34 picked up.

6. Lessons

  1. Compute the dynamic-share quotient before writing code: a run-once region's benefit ceiling = static count / (per-kt count × nktile); a region with a quotient < 1% does not deserve an experiment.
  2. Forensic ptxas's hoisting ability in SASS, not in source intuition: the uniform register is how it expresses "loop invariant"; an index chain that looks recomputed every iteration may already be split into prolog + UR.
  3. A store widening's width ceiling is written in the fragment lane map: confirm the thread's output contiguity first, then set the STG.64/128 target.
  4. A census's split dimension decides what it can see: splitting by opcode (r25) finds instruction-type surplus; splitting by region (r32) finds frequency structure — the same 2,617 instructions, but only the second split exposes "21% in staging, 0.4% in epilogue".
  5. A closing sweep's value is turning open questions into answered ones: r32 spent a day fencing off two levers so r33's hypothesis could be tested single-variable — negative results queue into the record, and only then is the road clean for positive ones.

← 34-r31-qmajor-sda-repack · Index · 36-r33-hybrid-inner-loop →

36 · r33 — Hybrid inner-loop port: SASS fully identical, hypothesis falsified (REVERTED)

Result: porting the shape of llama.cpp's j0-outer/n-inner inner loop into mmq_raw_nb_kernel (only the loop enumeration order changes — no layout, math, or shell changes): every emission-relevant gate untouched — SASS byte-identical to r31 (64 IMMA in the same order, 109 regs, same LDS/LDSM counts), parity 1/0, greedy byte-identical, interleaved 4-pair median −0.25%. The "remaining 1.15×/GMAC residual = loop- organization-induced SASS codegen" hypothesis is falsified; the census integer-ALU 1.66 e-3/MAC (llama 0.751) did not converge toward llama, and the line closes here. Commit: 697ef04 (docs-only record commit; the ported code was completed, measured, and reverted in the local tree — never landed as a code commit). Date: 2026-09-04.

1. Background — where things stood

r32 (doc 35) sealed off the "instruction count" axis entirely: staging addressing already hoisted by ptxas (source-level lever dead), epilogue structurally capped, and the remaining integer-ALU surplus = A-frag LDSM + intrinsic sda/sds decode + fp rescale, all "compiler floor". On the books, the NB kernel's per-GMAC warp instruction stream still ran 1.15× above llama.cpp — a number with its own shrinking history: at r13's counter forensics the gap was 10.14 vs 6.06 M/GMAC (1.67×); r28's NB kernel (2 blocks/SM) and r29's unroll squeezed it to 1.15×. Every conceivable, source-level cuttable instruction had been tried or proven not to exist.

One surviving explanatory framework remained — the campaign's last open "soft" hypothesis: is this 1.15×/GMAC residual purely loop organization — the SASS codegen difference that the loop's nesting shape induces? ptxas's scheduler takes as input not just the instruction set but the loop's nesting shape — perhaps llama's inner loop enumerates the mmas and rescales in some particular order that lets ptxas build a tighter pipeline (different issue gaps, different register-liveness peaks, different stall landings); perhaps we only need to write the loop in llama's shape and the machine code will "converge" toward it.

The hypothesis has a history worth recording. r10 (doc 15) ported llama's math decomposition — when nibbles unpack, when dmin folds in, when scales multiply — and still measured 462–468 vs 470 tok/s: the same decomposition, still 5× slower; the residual was not in the math organization. r30 issued a method warning: the SWAR unpack was written before anyone noticed the SASS-level ptxas had long CSE'd it — look at the machine code before acting. r33 tests the next organization level down: no math change, only the issue order of the mmas and rescales (the loop's enumeration shape). If this level is falsified too, the "codegen" class of hypotheses has exactly one exit left — changing the instruction composition itself (r33 explicitly records it as a scope caveat, handed to r34).

Archival situation identical to r32: git show 697ef04 --stat contains only docs/CUDA_OPTIMIZATION.md +69 lines and docs/LLAMA-CPP-MMQ-ANALYSIS.md +47 lines — the ported code was measured and reverted; the record commit is the only carrier.

2. Principle — the GPU mechanism

Two loop shapes, one DAG. The computation inside a 32-k chunk is fixed: 4 A-frags (activations, each 16 tokens × 32k) × 2 B-frags (weights, each 8 od × 32k) = 8 mma.m16n8k32, and each mma's accumulator is consumed by the same chunk's two-term rank-1 rescale. Draw the chunk's 8 mmas + 8 rescale groups as a dependency graph: the nodes are fully determined by the chunk geometry, and the only edges are "mma → its own rescale" — both sources enumerate the same graph:

  • minfer shape (g-outer × nh-inner): load 4 A-frags, load 2 B-frags, then for g { for nh { mma(A[g], B[nh]) } } — A-frag loads are amortized across the nh loop, B-frag loads across the g loop.
  • llama shape (j0-outer / k01 / n-inner): for j0 { load B; for k01 { for n { mma(B, A[n]); rescale } } } — the B-frag (weight) load is hoisted outside the n loop, one weight fragment reused by all n token-minitiles (llama's B-reuse-across-n).

Fragment-load counts, mma counts, and rescale application points are identical — the only difference is "which load is written at which nesting level in the source". Once fully unrolled, source order is merely an enumeration order of the DAG, not a scheduling constraint: #pragma unroll flattens the loop bodies into straight-line code, and ptxas's scheduler reshuffles the same DAG freely. r33's entire bet was that "the same DAG, enumerated differently, might schedule differently".

The hypothesis's steelman — what source order can change in theory. Stated fairly, it is not baseless: before unrolling, source order genuinely determines (a) the issue timing of loads relative to mmas (stall landings), (b) each fragment's register liveness interval (pressure peaks), (c) which load already sits outside the inner loop at the source level (approximate manual hoisting). All three are real constraints in unrolled-less source — the bet is that they still constrain ptxas after unrolling. r33 is the controlled experiment for that bet.

The hypothesis earns serious treatment because it makes two observable predictions: (1) if source order affects codegen, the ported SASS must change (instruction order, register allocation, or sync points — at least one); (2) if the SASS is unchanged, the wall-clock difference must be zero. The predictions are mutually exclusive and both cheap to test. r33's design puts both on the table and lets the SASS diff referee.

In method lineage, r33 is the last step of the "SASS forensics" three-step ladder: r30's first SASS-first (look at what the compiler already generates before writing a lever), r32's region census (where instructions live, how often they execute), r33 promoting identity itself to a criterion (byte-identical = hypothesis dead). All three steps share one tool (cuobjdump -sass) — getting cheaper and more lethal each step.

Geometry conversion: where the 8 mmas come from. llama's warp tile is 32 od × 64 tokens: rows_per_warp / tile_C::I = 2 M-minitiles (ntx = 2), with j0 stepping by ntx × tile_C::J. Our warp tile is 16 od × 64 tokens (NB kernel: 8 warps × 16 od = MMQ_NBJ 128) — under the same m16n8k32 / tile<16,8,int> lane map, the port folds into j0 = 2 od-groups × n = 4 token-minitiles = 8 mmas per 32-k chunk — the same 8, only the enumeration order changed. The 64 IMMA in static SASS = the KDR=8 kd expansion × 8 mmas per chunk — "64 IMMA in the same order" refers to exactly this batch's issue order.

A note on why the hypothesis tempts: llama's tile amortizes the weight B-frag load across all token-minitiles, which looks like a free locality win — the next day r36 quantified with wavefront counting that llama's fragment-load rate is 0.125 LDSM/IMMA (ours 0.5), so the gap really is in fragment reuse. But that is a tiling property (32 od rows/warp vs 16), not a loop-order property — r33 proves with SASS identity that under the same tile, enumeration order changes no load count.

k01 degenerates → parity by construction. llama's k01 loop steps by QI8_1 (the 8-k sub-blocks inside one q8_1 block) because their one m16n8k32 consumes 32-k while one q8_1 tile holds several sub-blocks. In our geometry k01 is degenerate: one m16n8k32 covers exactly the whole 32-k chunk (nchunk = id/32), so the k01 loop has a single iteration. The per-32-k-chunk rescale boundaries therefore do not move, and the fp accumulation order never changes — numerical equivalence is not a measured gamble but constructed. It also means any SASS difference could come only from scheduling, never from numerical reordering — a single-variable experiment, cleanly rare.

SASS identity = the definition of falsification. This experiment has a logical shortcut: identical SASS ⇒ identical cycle behavior ⇒ wall-clock difference necessarily 0. So the correct experimental order is diff the SASS first, then decide whether to run performance — byte-identical machine code cannot produce a different wall clock; running A/B merely adds a formal number for the record. r33 is the campaign's first use of "SASS identity" as an independent gate that can terminate an experiment early.

3. Implementation

3.1 Design choices (why a "hybrid" port)

"Hybrid" means: loop order only, shell fully kept. The port keeps the kernel signature, the 64×128 block geometry, KD=8, all smem layouts (qa8 / the r31 q-major sda repack / qb-raw / sds), r20 split-phase staging, the r22 XOR swizzle, the launcher and MINFER_MMQ_RAW_NB gate, r15's two-term fp32 rank-1 rescale semantics, and the fp32 write-back. The only change is the compute loop's enumeration shape. This makes any SASS difference attributable solely to loop order — attributional singularity matters more than "porting more like llama".

The fragment mapping table is validated before compiling. The easiest mistake in an enumeration-order port is not scheduling but miscopying one line of the (g, nh, l) → (i, j) output-unit mapping — parity would catch it, but how it catches (a few ulps vs large offsets) wastes half a day. The port first tabulated, for the new enumeration, which C accumulator each mma writes and which output grid point each accumulator maps to, then checked them one by one against the original: fragment maps validated, 0 mismatches. Only after that does the k01-degenerate constructive parity argument truly close.

Not porting the operand orientation is deliberate. In llama's mma, A = weights, B = activations (their activation fragment goes through load_generic, a trivial LDS — the source comment's own words: "faster than load_ldmatrix"; only the weight fragment earns ldmatrix). minfer is A = activations (ldmatrix), B = weights (raw-nibble register unpack). Flipping the orientation too would change the instruction composition (eliminating the A-frag LDSM class) — no longer a "loop organization" experiment. That half-step is explicitly recorded as a scope caveat, left for r34's narrow slice.

Write the equivalence proof first, then the code. The k01-degeneracy argument was written before the implementation: parity is guaranteed by construction, and the experiment's output space has only two points — "SASS same / different".

3.2 Key code

minfer's original compute loop (survived the revert in the current tree; each chunk's 8 independent mma chains, with A-frag loads and B-frag unpack):

// src/cuda_kernels.cu:6344-6388 (mmq_raw_nb_kernel — g-outer × nh-inner, the before of r33)
// A fragments: 4 independent 16-token groups (T=64), r22 G[].
int a[4][4], b[2][2];
#pragma unroll
for (int g = 0; g < 4; g++) {                       // ← A (activations): loaded once via ldmatrix
    const uint8_t* p = qat + G[g];
    unsigned r0_, r1_, r2_, r3_;
    asm volatile(
        "ldmatrix.sync.aligned.m8n8.x4.shared.b16 "
        "{%0,%1,%2,%3}, [%4];\n"
        : "=r"(r0_), "=r"(r1_), "=r"(r2_), "=r"(r3_)
        : "r"((unsigned)__cvta_generic_to_shared(p)));
    a[g][0] = (int)r0_; a[g][1] = (int)r1_;
    a[g][2] = (int)r2_; a[g][3] = (int)r3_;
}
// B fragments: raw-nibble in-loop unpack (weight-side register unpack, 2 frags per chunk)
{
    const int p = sg >> 1, is_hi = sg & 1, lm3 = lane & 3;
    const unsigned M = 0x0F0F0F0Fu;
    #pragma unroll
    for (int nh = 0; nh < 2; nh++) {
        const int jj = j0w + nh * 8 + (lane >> 2);
        const uint8_t* qs = qb_raw + (size_t)jj * 128;
        const uint32_t* q0 = (const uint32_t*)(qs + p * 32 + lm3 * 4);
        const uint32_t* q1 = (const uint32_t*)(qs + p * 32 + 16 + lm3 * 4);
        uint32_t v0 = *q0, v1 = *q1;
        b[nh][0] = (int)(is_hi ? ((v0 >> 4) & M) : (v0 & M));
        b[nh][1] = (int)(is_hi ? ((v1 >> 4) & M) : (v1 & M));
    }
}
…
// 8 independent mma chains per thread per chunk (4 A-frags x 2 B-frags),
// all C fragments live simultaneously.
#pragma unroll
for (int g = 0; g < 4; g++)                         // ← outer loop over A frags
    #pragma unroll
    for (int nh = 0; nh < 2; nh++)                  // ← inner loop over B frags
        mmq_mma_k32(clow[g][nh], a[g], b[nh]);      // 8 mmas, then the same chunk's rescale

llama's reference shape (the source the port was checked against, mmq-vec-dot.cuh:408-437, current upstream tree):

// llama.cpp ggml/src/ggml-cuda/mmq-vec-dot.cuh:408-437 (the q8_1×q8_1 mma branch)
#pragma unroll
for (int j0 = 0; j0 < J; j0 += ntx*tile_C::J) {         // ← outer loop over od-groups (weights)
#pragma unroll
    for (int k01 = 0; k01 < MMQ_TILE_NE_K; k01 += QI8_1) {
        tile_B   B;
        float2 dsB[tile_C::ne/2];
        load_generic(B, y_qs + j0*MMQ_TILE_Y_K + k01,
                     MMQ_TILE_Y_K);                     // ← B frag hoisted outside the n loop (trivial LDS)
#pragma unroll
        for (int n = 0; n < ntx; ++n) {                 // ← inner loop over token-minitiles
            tile_C C;
            mma(C, A[n][k01/QI8_1], B);
#pragma unroll
            for (int l = 0; l < tile_C::ne; ++l) {
                sum[(j0/tile_C::J + n)*tile_C::ne + l] +=
                    dmA[n][l/2][k01/QI8_1].x*dsB[l%2].x*C.x[l];
                sum[(j0/tile_C::J + n)*tile_C::ne + l] +=
                    dmA[n][l/2][k01/QI8_1].y*dsB[l%2].y;   // ← two-term rank-1 rescale
            }
        }
    }
}

The other big shell piece kept as-is is the rescale section — r15's two-term rank-1 fold (the d*sc main term + the -dmin*m rank-1 term) and r31's sda reads (two LDS.128s per warp) survive unchanged in the port, because they define the fp-accumulation numerical contract:

// src/cuda_kernels.cu:6399-6427 (mmq_raw_nb_kernel's sda reads + rescale — kept shell, excerpt)
// r31: Q-major sda repack — one uint32 per token; the group-region
// split (g/2 region block, q*16B stride) makes each warp LDS.128
// read 8 unique 16B words at 16B stride = bank-conflict-free.
const uint32_t* sda_blk = sda_q + (size_t)kd * MMQ_NBI
                          + (size_t)(lane >> 2) * 4;
const uint4 s0 = *(const uint4*)(sda_blk);
const uint4 s1 = *(const uint4*)(sda_blk + 32);
#pragma unroll
for (int g = 0; g < 4; g++) {
    …
    da_q[0] = h2f((unsigned short)(w0 & 0xFFFF));      // d (f16 half-word)
    sa_q[0] = (int)(short)(w0 >> 16);                  // ssum (i16 half-word)
    …
    #pragma unroll
    for (int nh = 0; nh < 2; nh++)
        #pragma unroll
        for (int l = 0; l < 4; l++) {
            …
            sum[idx] += da * dsv[nh][l & 1] * (float)clow[g][nh][l];  // main term d*sc
            sum[idx] += dma[l >> 1] * dmv[nh][l & 1];                 // rank-1 term -dmin*m (r15)
        }
}

The port (never committed) replaced minfer's for g { for nh } with the for j0 (2 od-groups) { for n (4 token-minitiles) } above: the B-frag load hoisted outside the n level, the A-frag load still outermost (once per chunk, unchanged), the rescale fold copied verbatim in the sum[…] += … two-term form. All load counts, mma counts, and rescale application points map one-to-one to the before (the §3.1 mapping-table check targets exactly this). The two enumerations side by side:

before (minfer):     load a[0..4]; load b[0..2];
                     for g in 0..4 { for nh in 0..2 { mma(a[g], b[nh]); rescale(g,nh) } }

ported (llama shape): load a[0..4];                     // ← still outermost, once per chunk
                      for j0 in 0..2 {                  // od-groups (weights)
                          load b[j0];                   // ← hoisted outside the n loop
                          for n in 0..4 { mma(a[n], b[j0]); rescale(j0,n) }   // token-minitiles
                      }

Both are 8 mmas, 8 rescale groups, 6 fragment loads per chunk — isomorphic DAG, different enumeration order. ("Hybrid" gets its name here: llama's loop shape + all of minfer's shell.)

3.3 Pitfalls

  • A SASS diff must be same-version, same-flags. Byte-identity is meaningful only under the same ptxas and the same compile options; compiling each side separately and diffing cuobjdump -sass means any flag drift manufactures fake differences.
  • Mapping-table errors and scheduling differences are two diseases — do not use one medicine. The biggest risk in an enumeration-order port is a miscopied (g,nh,l)→(i,j) mapping (a parity disease), not scheduling (a perf disease); only the 0-mismatch mapping check gives the SASS comparison its "pure scheduling" interpretive authority.
  • Do not read "the census did not move" as "the experiment was botched". After the port the integer-ALU census read 1.66 e-3/MAC, unmoved — that is not a sign the port failed; it is positive evidence that "source order does not affect instruction composition". The census measures composition; r33 manipulated order.
  • Take the logical shortcut first. Running the 4-pair A/B before looking at the SASS wastes a round of machine time; the SASS identity check alone condemns the experiment in ~10 minutes. Fixing "diff first, then measure" as an order is this experiment's real methodological output.

4. Verification

  • SASS byte-identity gate (the experiment's main gate): cuobjdump -sass comparing the port against the r31 baseline — defends against both the reverse illusion "thought we changed scheduling, actually didn't" and the false negative "thought we didn't change, actually did". Result: byte-identical — 64 IMMA in the same order, same LDS/LDSM counts.
  • Fragment mapping-table check (0 mismatches): defends against the subtlest output-unit misalignment an enumeration-order port can produce — the precondition for the SASS comparison's "pure scheduling" authority.
  • ptxas resource audit: mmq_raw_nb_kernel<8> 109 regs / 0 spill, smem 43,008 B, 2 blocks/SM — all identical to r31, ruling out the bypass "loop reordering changed register allocation".
  • Parity 1/0 + greedy-32 byte-identity: confirms the k01-degenerate constructive-equivalence argument holds on real hardware.
  • Interleaved 4-pair A/B: −0.73% / −0.59% / +1.30% / −0.10%, median −0.25% — completes the record for the "wall clock unchanged" ruling.
  • Instruction census: integer-ALU 1.66 e-3/MAC (llama 0.751), unmoved before and after the port — quantifies the "no convergence toward llama".

The gates' execution order is itself part of the conclusion: SASS diff (~10 minutes) → resource audit (as before) → parity/greedy (constructive confirmation) → A/B (the formal number) → census (quantifying non-convergence) — the cheaper the gate, the earlier it runs; the first gate delivered the verdict and the remaining four were purely record-keeping.

5. Results

Hypothesis falsified, line closed. With the loop shape swapped to llama's j0-outer/n-inner, ptxas produced SASS byte-identical to r31: the compiler schedules the unrolled 8-mma + rescale stream the same way whether the source enumerates g-outer/nh-inner or j0-outer/n-inner. §2's three steelman points (issue timing, liveness intervals, manual hoisting) all evaporate after unrolling — they were always ptxas's scheduling degrees of freedom, not source-level constraints. Byte-identical SASS physically cannot give a different wall clock; the measured median −0.25% (inside the noise band) is just the footnote on that logical necessity. Ported code reverted, cmp-verified = HEAD.

The verdict, stated fully (the §11.13 wording): the 1.15×/GMAC residual is not loop-organization codegen; it is an inherent difference in instruction composition (A-frag LDSM consumption + intrinsic sda/sds scale decode + staging index math — the regions r32 attributed) plus nvcc/ptxas's scheduling of the whole kernel, and source-level reordering touches neither. To change ptxas's output, the DAG itself must change.

The verdict's later footnotes: the next day, r36 added a mechanism footnote to the residual with wavefront counting — the MIO pipe is not scarce, and llama's real advantage is the A-frag reuse rate (0.125 vs 0.5 LDSM/IMMA, a tiling property) — refining but not overturning r33's conclusion; and r34 picked the "change the composition" direction out of the scope caveat and landed +9.72% (doc 37). The door r33 closed, r34 dismantled around the frame and walked through.

The experiment's cost-benefit. The entire cost of this falsification: one local loop rewrite (after the mapping check), two compiles + a SASS diff, one 4-pair A/B — bought the closure of the entire "loop organization" axis, plus a cheaply re-runnable gate (after any future ptxas version change, one SASS diff suffices). Against the routes it eliminated (continuing trial and error along loop shape, each round a full port + full verification), this is one of the campaign's best cost-performance "negative results".

Veto mechanism and retry conditions: this hypothesis's retry conditions are written precisely in the scope caveat — only ports that change instruction composition can move the wall: flip the orientation to A=weights (eliminating the per-tile activation A-frag LDSM, weights moving to trivial LDS), or move the A-side layout transform out of the kernel. The former was out of reach within the then-current budget; the latter was realized by r34 as a narrow slice. If a future ptxas version changes behavior, the re-check costs one SASS diff — the cheap re-inspection entrance this experiment left behind.

6. Lessons

  1. A loop-organization hypothesis can be falsified by SASS identity before any performance measurement: byte-identical machine code is definitional evidence of "hypothesis dead" — diff the SASS first, then decide whether to spend machine time.
  2. After #pragma unroll, source order is only an enumeration order of the DAG: issue timing, register liveness, and load hoisting are real constraints at the source level and scheduling degrees of freedom after unrolling — to change ptxas's output, change the DAG first.
  3. Split a large residual into separately falsifiable sub-hypotheses: "codegen difference" hides three components — composition, order, scheduling; r32 sealed composition, r33 sealed order, and the remaining scheduling component can only be bypassed by changing the problem itself (r34).
  4. The mapping table precedes the compiler: in a loop-order experiment, the output-unit mapping check (0 mismatches) is the precondition for the SASS comparison to mean anything.
  5. The scope caveat is a negative result's will: the sentence in r33's record about "the deeper direction, and why it was not taken" became r34's work order directly — writing down clearly what was not done has more long-term value than one more measured number.

← 35-r32-finite-lever-sweep · Index · 37-r34-quantize-transpose-prepass →

37 · r34 — The quantize-transpose prepass: layout-transform locality (LANDED, +9.72%)

Result: the A-side layout transform (the r22 XOR swizzle + r31 q-major sda repack index math) moves out of the mma kernel's per-tile staging into a per-GEMM quantize prepass — llama.cpp's quantize_mmq_q8_1 design. New quantize_q8_0_pad40_t (bit-identical, transposed layout, zero-filled pad tokens)

  • mmq_raw_nb_bt_kernel (A staging degenerates to bulk LDG→STS with no per-element index math): 1364.2 → 1496.8 tok/s (+9.72%, interleaved 4-pair all positive, no interval overlap), prepass 0.405 vs 0.446 ms (0.908× — smaller, not bigger), bt kernel 103 regs / 0 spill (NB 109 — the direct evidence that the staging index math left the kernel). P6's largest single-mechanism gain since r28. Commit: ba977bf (src/cuda.rs +135, src/cuda_kernels.cu +309). Date: 2026-09-05.

1. Background — where things stood

After two consecutive negative results (r32/r33, docs 35/36), the line stood at a thoroughly cleaned-up crossroads. r32 had sealed the "instruction count" axis: staging addressing already hoisted by ptxas, epilogue structurally capped, remaining integer-ALU surplus = A-frag LDSM + intrinsic sda/sds decode + fp rescale, all "compiler floor". r33 had sealed the "loop organization" axis: source-level reordering produces byte-identical SASS. Together the two verdicts point at the only exit — to change ptxas's output, the DAG itself must change. r33's scope caveat wrote that "way of changing" as a will: flip the mma operand orientation to A=weights (eliminating the per-tile activation A-frag LDSM), or move the A-side layout transform out of the kernel. The former was out of budget; r34 took the latter — a llama-faithful narrow slice.

The problem itself was restated in §11.14 as layout-transform locality. Look at what the minfer NB kernel's staging does: for every (64-token block, od-tile column) combination it reads the activations' qs/d/ssum out of the native token-major 40 B chunks, then writes them into the mma-consumption smem layout via the per-element index math of the r22 XOR swizzle and the r31 q-major repack. And the grid is (nt/64, od/128) — the same 64-token A tile gets restaged once per od-tile column: 28 times per buffer for the od=3584 projections (q/o/down), and 148 times for od=18944 gate/up. The transform's index math scales linearly with the restage count, while the transform itself is od-independent — pure repeated labor.

llama.cpp never had this problem: its quantize-side quantize_mmq_q8_1 pre-transposes the activations into exactly the layout the mma kernel consumes, inside the prepass (both the transpose and the intra-block shuffle happen on the write side), so the kernel's A staging is a near-trivial copy. Its planes are y_qs (qs) / y_dm (packed scale half-words) — one-to-one with minfer's new yqs/ysda planes; the activation-side trivial LDS r33 quoted ("faster than load_ldmatrix") reads exactly this pre-transposed plane. r32's census priced the contrast: staging = 21% of kernel instructions (455 per-kt) — r34 relocates the "transform" share of that 21%.

One more anchor: whole-prefill was still on the ~1440 tok/s plateau (the r31 window), while the campaign target (set at r6) was f16 parity ~24 TMAC/s / llama parity ~30. r34 is among the last levers of the "q4_K kernel itself" — a week later r37's attribution would show q6_K is the new wall (51.2% of wall), but that is another line's (Era D's) business.

2. Principle — the GPU mechanism

Layout-transform locality — splitting the two costs. Each byte of A-side staging costs two things:

  • Copy: moving bytes from global to smem. No ALU, pure LSU, and repeated reads are absorbed by cache — paid again per restage, but cheap.
  • Transform: computing where each byte lands in smem. The XOR swizzle's shifts/xors/masks, the q-major repack's region indices — pure ALU, register-occupying, scaling linearly with the restage count.

Before r34 the two frequencies were bound together: transform count = copy count = restage count (the od/128 column count). r34 moves the transform into the prepass so it is paid once per buffer; the copy stays in the kernel but degenerates into a uint4 bulk copy:

restages (per A buffer) = grid.y = ceil(od / 128)
  od = 3584 (q / o / down) → 28 times        od = 18944 (gate/up) → 148 times
transform cost: old = once × grid.y; new = once × 1     ← the multiplication r34 removes
copy cost: both sides = once × grid.y (old: a 455-instruction mixed stream; new: pure uint4 copies)

Transform count ÷28 for the od=3584 projections, and the in-kernel index math disappears wholesale. As for re-reading the same plane across od-tile columns: it is the same class as the B side's weight re-reads (r19 proved that class is absorbed by L2 and creates no DRAM traffic) — not a new cost.

The plane layout's arithmetic. The planes the new prepass emits are the exact global mirrors of the smem layout: one swizzled 2048 B qs block per (64-token block, chunk) ([ntb][nchunk][2048]), one 256 B packed d|ssum block ([ntb][nchunk][256]). Against the old path: native pad40 is 40 B per (token, chunk) (4 B d + 4 B ssum + 32 B qs), so each block-chunk read 64×40 = 2560 B and wrote smem 2048 + 256 = 2304 B; the new plane is directly 2304 B per block-chunk — d and ssum packed into one u32 (f16 d + i16 ssum) halves the scale bytes from 8 → 4 B/token/chunk. At nt=3354, id=3584: the planes are ≈ 11.6 MiB (qs) + 1.45 MiB (sda) ≈ 13 MiB — a one-time per-GEMM activation-buffer cost.

The bulk copy's issue arithmetic. Each k-tile's A-segment copy: qs side KDR × NBI × 32 / 16 = 8×64×32/16 = 1024 uint4s (16 KB), sda side 8×64×4/16 = 128 uint4s (2 KB); 256 threads issue 4-5 LDG.128 → STS.128 each, with contiguous per-thread addresses = perfectly coalesced. Against the old path's same segment: 455 mixed instructions with XOR/shift/mask. The two call sites RAW_STAGE_NB_BT(0) / if (kt > 0) RAW_STAGE_NB_BT(kt) are identical to the NB kernel's — r20 split-phase double buffering as before (prefetch the next tile while computing the current); only the stage macro's internals change.

Why the prepass does not grow — it shrinks. The quantize body itself (8×float4 reads, the amax tree, rintf/clamp, the ssum sum) is output-layout-independent — bit-identical in both versions. Only the write side changes: 8 u32 qs writes + 1 u32 sda write per thread, using the same addressing formula as the old smem STS (the same swizzle routine) with the target switched from smem to global. Packing additionally halves the written scale bytes. Measured 0.405 vs 0.446 ms (0.908×) — the transpose carries no tax; it nets a small saving.

The register mechanism — why regs drop. The per-kt staging index math (the swizzle's XOR/shifts, the repack's region indices) is a real source of register pressure: their live ranges span the LDG batch and the STS write-back. With the transform out of the kernel, this set of intermediates vanishes — 109 → 103 regs, 0 spill kept, smem 43,008 B unchanged → 2 blocks/SM kept. Occupancy unmoved while per-tile instructions shrink: a clean positive combination.

The consumption-side contract — what layout ldmatrix demands. Why must the old staging XOR-swizzle, and why must the new plane replicate it byte for byte? Because the consumer's ldmatrix.sync.aligned.m8n8.x4 reads 8-row × 16 B tiles from smem, and if the 8 rows' addresses crowd into the same bank group, one ldmatrix tears into multiple serial replays. r22's swizzle (the inter-row XOR (R>>2)&7 term) exists precisely to stagger adjacent tiles' rows across banks; the kernel-side G[4] precompute (r22: each tile's start computed once) then moves the address ALU from every ldmatrix into the prolog. r34 does not touch one hair of this contract — it only moves "who arranges the bytes into this layout" from the kernel's per-tile staging to the prepass's once-per-buffer. The layout is unchanged, so the consumer's read pattern need not change; the byte-exactness validator guards exactly this contract.

Consumer-side invariance is the proposition's boundary. A-frags are still read from smem by ldmatrix — and ldmatrix does not consume an arbitrary layout but the m8n8 bank-friendly row-staggered layout, which is why the swizzle existed (r22). r34 changes neither the consumption layout nor the B weight staging, the SDS fold, the mma loop, or the fp32 write-back. The only thing that changes is "how the qa8/sda_q tiles are produced" — guaranteeing any wall-clock change attributes to the staging alone (the "eliminate LDSM" half-step of the r33 caveat was not done).

3. Implementation

3.1 Design choices (why this shape and not another)

A new kernel; the old one untouched. The minimal-intrusion shape for swapping the A-side supply is a copied mmq_raw_nb_bt_kernel (bt = bulk-transposed) with mmq_raw_nb_kernel kept alive as-is. More than conservatism: the old kernel becomes the A/B integrity control — its SASS in the new binary must be byte-identical to the old binary's, proving the A/B difference can only come from the bt kernel itself. Forensic check: git show ba977bf --stat shows src/cuda_kernels.cu +309 lines; quantize_q8_0_pad40_t, the MMQ_A_* constants, mmq_raw_nb_bt_kernel, and RAW_STAGE_NB_BT were all introduced by that commit; the current tree's A segment is still its introduced form (the later r59 DSC plane touched only the B-side sds staging).

Gate shape: MINFER_MMQ_A_TRANSPOSE=1, paired with MINFER_MMQ_RAW_NB=1 (current-tree dispatch: if nb && at && kd == 8). The opt-in experimental gate makes rollback free; this gate later carried a building — r49 added shared-A dedup on top, r51/r52 folded quantization into producers, r59 added the W_dsc plane, and r60 flipped the default on with the whole gate set. One detail: the native pad40 quantize (quantize_q8_0_pad40) was not deleted — after r34 it fires only when the bt path is unavailable (the dispatch comment's own words: "the native q8 buffer is only filled on the (rare) bb-bt fallback below").

Plane layout = mirror of the smem layout is the design's load-bearing wall: only when the global plane's byte order matches the smem target's byte order exactly can staging degenerate into an index-math-free bulk copy. To that end the prepass's write side embeds the old STS's swizzle routine directly (byte-for-byte replication; see §3.3).

Pad-token zero fill: blocks whose tail has fewer than 64 tokens emit all-zero pad rows (d=0, ssum=0, qs=0) — the transposed plane is thus independent of buffer-reuse history (deterministic), and the write-back's i < nt guard guarantees pad rows never land in C. "Zero fill instead of skipping writes" also keeps the bulk copy length constant (KDR*NBI*32 bytes exactly), so the kernel needs no length special-casing for tail blocks.

The launcher keeps the native quantize's shape (a 1-D grid covering all (pad-token, chunk) pairs):

// src/cuda_kernels.cu:3536-3545 (launch_quantize_q8_0_pad40_t — introduced in r34, unchanged)
void launch_quantize_q8_0_pad40_t(
    const float* x, uint8_t* yqs, uint8_t* ysda,
    int dim, int nt, int nchunk, int ntb, cudaStream_t stream
) {
    long long total = (long long)ntb * MMQ_A_BLK * nchunk;
    int block = 256;
    long long grid = (total + block - 1) / block;
    if (grid > 2147483647LL) grid = 2147483647LL;
    quantize_q8_0_pad40_t<<<(int)grid, block, 0, stream>>>(x, yqs, ysda, dim, nt, nchunk, ntb);
}

At nt=3354, id=3584: total = 53 × 64 × 112 = 379,904 threads → 1,484 256-thread blocks, one launch covering the whole plane; each thread writes 9 u32s (8 qs + 1 sda) — exactly the native quantize's write granularity, only the target address formula differs.

3.2 Key code

The prepass kernel (src/cuda_kernels.cu, introduced in r34; current-tree version — after r51, the producer-fused path reuses exactly this quantize body):

// src/cuda_kernels.cu:753-810 (quantize_q8_0_pad40_t — excerpt)
__global__ void quantize_q8_0_pad40_t(
    const float* __restrict__ x,
    uint8_t* __restrict__ yqs,    // [ntb][nchunk][2048] ← the swizzled qs plane
    uint8_t* __restrict__ ysda,   // [ntb][nchunk][256]  ← the packed d|ssum plane
    int dim, int nt, int nchunk, int ntb
) {
    int tid = blockIdx.x * blockDim.x + threadIdx.x;
    int total_pad = ntb * MMQ_A_BLK * nchunk;      // one thread per (pad-token, chunk)
    if (tid >= total_pad) return;
    int t = tid / nchunk;
    int b = tid % nchunk;
    const int r = t & (MMQ_A_BLK - 1);  // local token 0..63
    const int tb = t >> 6;              // 64-token block index
    // …quantize body: 8×float4 reads + amax tree + rintf/clamp + ssum (bit-identical to
    //   quantize_q8_0_pad40; pad tokens (t >= nt) emit all zeros, keeping the plane deterministic)…
    // Swizzled qs write — byte-identical to the NB kernel's old smem staging
    //   (qa8 + (R&~3)*32 + ((((R&3)<<1 + (u>>2)) ^ ((R>>2)&7)) << 4) + (u&3)*4).
    const int t4 = r & 3, grp = r & ~3;
    const int xswz = (r >> 2) & 7;
    size_t qbase = ((size_t)tb * nchunk + b) * MMQ_A_QASZ + grp * 32;
    #pragma unroll
    for (int u = 0; u < 8; u++) {
        const int off = (((t4 * 2 + (u >> 2)) ^ xswz) << 4) + (u & 3) * 4;  // ← the same swizzle
        *reinterpret_cast<uint32_t*>(yqs + qbase + off) = packed[u];         //   just written into global
    }
    // Packed d|ssum (r31 Q-major region split of the old sda_q).
    const int g = r >> 4, t15 = r & 15, q = t15 & 7, half = t15 >> 3;
    const int rg = g >> 1, gsel = g & 1;
    size_t sbase = ((size_t)tb * nchunk + b) * MMQ_A_SDASZ
                   + (rg * 32 + q * 4 + gsel * 2 + half) * 4;
    __half dh = __float2half(d);
    uint16_t dbits = *reinterpret_cast<uint16_t*>(&dh);
    *reinterpret_cast<uint32_t*>(ysda + sbase) =
        (uint32_t)dbits | ((uint32_t)(uint16_t)ssum << 16);   // f16 d + i16 ssum = 4 B
}

The bt kernel's header comment (r34's original, src/cuda_kernels.cu:6448-6455) states the design's boundary in one breath:

// --- P6 r34: NB kernel with the A-side layout transform relocated into a
// quantize-transpose prepass (MINFER_MMQ_A_TRANSPOSE=1). Byte-identical
// compute to mmq_raw_nb_kernel (same qa8/sda_q smem content, same ldmatrix
// fragment reads, same rescale) — only the A STAGING differs: the qs plane and
// the packed d|ssum are emitted PRE-TRANSPOSED (quantize_q8_0_pad40_t) so the
// per-(kt, warp) A reads become contiguous bulk LDG->STS with no per-element
// index math (the r22 XOR swizzle and the r31 q-major sda repack are baked into
// the prepass layout). The B (weight) + SDS staging is unchanged.

After — the bt kernel's A staging (the A segment of RAW_STAGE_NB_BT; introduced in r34, the A path unchanged to this day; the current tree's DSC/B-side sds is r59's later story):

// src/cuda_kernels.cu:6486-6498 (RAW_STAGE_NB_BT — A side)
/* ---- A: bulk LDG->STS of the pre-transposed qa8/sda (no math) ----*/
{
    const size_t qbase = ((size_t)blockIdx.x * nchunk + (size_t)(kt) * KDR) * MMQ_A_QASZ;
    for (int off = threadIdx.x; off < (KDR * MMQ_NBI * 32) / 16;        \
         off += blockDim.x)                                             \
        ((uint4*)(qa8))[off] = ((const uint4*)(qa8g + qbase))[off];     // ← pure uint4 copy
    const size_t sbase = ((size_t)blockIdx.x * nchunk + (size_t)(kt) * KDR) * MMQ_A_SDASZ;
    for (int off = threadIdx.x; off < (KDR * MMQ_NBI * 4) / 16;         \
         off += blockDim.x)                                             \
        ((uint4*)(sda_q))[off] = ((const uint4*)(sdag + sbase))[off];   // ← same, no index math
}

Before — the NB kernel's A staging (RAW_STAGE_NB, LDG batch + swizzle STS, excerpting the STS write-back; the full contrast is doc 35 §3.2):

// src/cuda_kernels.cu:6268-6277 (RAW_STAGE_NB — the index math every tile paid before r34)
_Pragma("unroll")
for (int i = 0; i < KDR * 2; ++i) {
    const int x = threadIdx.x + i * 256;
    const int u = x & 7, r = (x >> 3) & (MMQ_NBI - 1),
              kd = x / (8 * MMQ_NBI);
    const int R = kd * MMQ_NBI + r;
    *(unsigned*)(qa8 + (size_t)(R & ~3) * 32                       // ← these three addressing lines
        + (size_t)(((((R & 3) << 1) + (u >> 2))                    //   all vanish in the
                    ^ ((R >> 2) & 7)) << 4)                        //   bt kernel
        + (size_t)(u & 3) * 4) = av[i];
}

The Rust-side dispatch (src/cuda.rs; the comments are r34's original, the cache logic is r49/r52's later story):

#![allow(unused)]
fn main() {
// src/cuda.rs:3288-3307 (prefill_mmq's routing branch, excerpt)
// P6 r34: relocate the A-side layout transform out of the mma
// kernel into a quantize-transpose prepass (llama.cpp's design).
// Under MINFER_MMQ_A_TRANSPOSE=1 the activations are emitted
// PRE-TRANSPOSED (quantize_q8_0_pad40_t) so the NB kernel's A
// staging is a bulk LDG->STS; …
let mut nb_ok = false;
if nb && at && kd == 8 {
    let nchunk = (id / 32) as i32;
    let ntb = ((nt as i64 + 63) / 64) as i32;
    // r49: cache-backed transposed A-quantize prepass (reuses
    // the previous same-A matmul's qa8g/sdag when consecutive).
    let (qa8g, sdag) = self.mmq_quantize_transposed(
        x as *const f32, id as i32, nt as i32, nchunk, ntb, stream,
    );
    // …(launch mmq_raw_nb_bt_kernel; null planes fall back to the NB kernel)…
}
}

3.3 Pitfalls

  • The swizzle routine must be replicated byte for byte, with the formula written into the comment. The qs write side's comment embeds the full formula — because its correctness criterion is not "looks right" but "byte-identical to the old smem staging's product". A byte-exactness validator was written for exactly this: 9 shapes (including nt values that are not multiples of 64) comparing the planes against the old path's smem content, 0 mismatches. The non-multiple-of-64 nt values are deliberate: the tail blocks' all-zero pad rows, the grp/tb out-of-range index behavior, and the bt kernel's write-back guard are only truly exercised on tail blocks.
  • The ncu census failed this round. ncu injection failed for both mmq_raw_nb_kernel and mmq_raw_nb_bt_kernel ("Unknown Error on device 0" — a platform/toolchain limitation), so the instruction census was unobtainable. The landing was not blocked: the mechanism is corroborated by three independent pieces of evidence — the register drop (109→103, the static evidence that the staging index math left the kernel), the prepass timing (0.908×), and the wall clock (+9.72%).
  • "Transpose" is easy to read as a cost. Intuition says writing an extra transposed plane should be taxed; in fact 0.908×. The lesson: do not estimate cost from a name — the transform's ALU did not grow (the same swizzle formula), and the scale packing even saved 256 B/block-chunk of writes.
  • The discipline of not deleting the old kernel. mmq_raw_nb_kernel stays as the control and fallback path (serving normally when the planes are missing or the gate is off), which gives the A/B integrity gate a checkable anchor. The bt kernel's B/SDS segments are verbatim identical to the NB kernel's — maintaining two copies is the price of duplication, exchanged for the comparability of "the difference is only in the A segment": any B-side drift in a SASS diff is immediately suspicious.

4. Verification

  • byte-exactness validator (9 shapes, including non-multiples-of-64 nt): defends against "one XOR of the replicated swizzle written wrong" — the transposed planes must be byte-identical to the old smem staging, the foundation of the entire "supply swap only" proposition. The 0 mismatches cover every (tb, chunk) index corner: multi-block, tail-block pad rows, and the 64-token alignment boundaries.
  • plain NB kernel SASS byte-identity (across the old and new binaries): A/B integrity — proves the control was untouched and the performance difference can only come from the bt kernel.
  • ptxas resource audit: bt 103 regs / 0 spill, smem 43,008 B unchanged — defends against an occupancy regression (2 blocks/SM must hold) and doubles as static mechanism evidence.
  • cuda_prefill_mmq parity 1/0 + greedy-32 token identity (matching the f16 default path): defends against the layout relocation changing any numerics.
  • Prepass timing contrast (0.405 vs 0.446 ms): defends against "the transpose hiding its tax in the prepass" — the wall-clock gain must be proven not traded away on the quantize side.
  • Suite 166/0/3: defends against cross-shape regressions.

In the campaign's gate numbering (the ba977bf commit message's own words): Gates 1-6,8 green — byte-exactness, parity, greedy identity, suite, interleaved A/B, and the resource audit all green; the only absentee is the ncu census (platform injection failure), substituted by the three alternative pieces of evidence. This "gate numbers + absence declaration" combination later became the docs-commit standard signature.

  • Interleaved 4-pair A/B (7B q4_k_m @3354-token prefill, same window): 1364.2 → 1496.8, every pair positive with no interval overlap — defends against co-tenant/drift reading the signal as noise.

5. Results

Metricbeforeafter
whole-prefill (7B q4_k_m @3354 tok, 4-pair median)1364.2 tok/s1496.8 tok/s (+9.72%)
quantize prepass0.446 ms (native pad40)0.405 ms (0.908×)
mma kernel registers109 regs / 0 spill103 regs / 0 spill
smem / occupancy43,008 B / 2 blocks/SMunchanged
suite—166/0/3

+9.72% is the P6 line's largest single-mechanism gain since r28 (the NB kernel landing, +2.56%) — and its source is not a faster mma or fewer instructions but moving one job out of the wrong place. One anchor warning: the r31 window recorded 1439.40 tok/s, yet r34's window baseline is 1364.2 — absolute values are not comparable across sessions (machine state / co-tenant load); this step's credibility rests entirely on the same-window interleaved 4-pair A/B (the later "r59b lesson" codified this rule). The master table's vs-llama column records "—": before r37 the campaign ran on the opt-in MMQ path with no whole-prefill vs-llama record to cite (the §0 reading convention).

The mechanism chain, replayed. Read the three steps in sequence: r32's census drew the map (staging 21%, epilogue ~0.4%, compute hot path uncuttable) → r33 falsified "reordering saves instructions" (SASS identity) → r34 changed the question ("must this 21% be paid inside the kernel?" — answer: no). +9.72% did not fall from the sky: it is the composite of the map (r32), elimination (r33), and the designed relocation (r34) — the campaign methodology's most complete specimen of "negative results paving the way for positive ones".

A correction to the mechanism attribution: r33 called the residual "composition" — r34 proves the word is not monolithic. The staging component inside "composition" is actually a layout-transform locality problem and can be moved away; the A-frag LDSM consumption and fp rescale remain intrinsic (r32's conclusion stands as written). The residual was split smaller again.

Where the line went next: r35 (sda scale predecode into the prepass, −0.46%, REVERTED — instructions hiding in the IMMA's shadow are not worth cutting), r36 (A-frag wavefront economics, falsified in its H1), r37 (post-parity attribution: the q6_K GEMM is 51.2% of the wall — with MMQ having made q4_K fast, the ball passed to the q6_K line) — Era C closes here, and Era D's q6_K/FA/prepass lines take the baton. This doc's plane layout [ntb][nchunk][2048/256] remains the foundation of the entire A-side stack to this day.

Follow-ups (each with its own doc): r49 discovered q/k/v and gate/up share one A, cutting prepass launches from per-GEMM to a deduplicated 193 → 110 (118.4 → 83.9 ms); r51/r52 folded quantization directly into the rms/swiglu producers (the prepass eventually 10.1 ms, with mode 2 not even writing the f32 output); r54/60 pushed the whole gate set to default-on.

6. Lessons

  1. Paying a layout transform once vs paying it restage-count times is a multiplication-level difference: transform cost = per-instance cost × the grid.y column count; wider-od projections amplify harder (gate/up is 148×, not 28×). r6's spec principle "transform out of the hot loop" — this is its complete application on the A side.
  2. Split copy and transform into separate ledgers: byte copies can be absorbed by L1/L2 and cost zero ALU; index math is pure ALU and occupies registers — once the transform leaves the kernel, the remaining copy degenerates into a bulk copy and the register pressure vanishes with it (109→103).
  3. The consumption layout is the contract; the supply location is a degree of freedom: the smem layout ldmatrix demands cannot change by one byte (r22's swizzle stands), but "who arranges it, at what frequency" is renegotiable — correctness is backstopped by the byte-exact validator.
  4. ncu being unavailable does not block a landing: the register audit, the prepass timing, and the wall clock are three independent pieces of evidence sufficient to corroborate the mechanism — a profiler is one source of evidence, not a precondition for landing.
  5. Leave a SASS-identical control behind for the replaced thing: the old kernel surviving intact = the A/B integrity gate gets its anchor for free; "new kernel + old kernel frozen" is the cheapest experimental design.

← 36-r33-hybrid-inner-loop · Index · 38-r35-scale-predecode →

38 · r35 — sda scale predecode: a total SASS win, a wall-clock tie (REVERTED)

Result: 7B q4_K BT-kernel prefill 1493.2 → 1486.3 tok/s (−0.46%, inside the noise band; bar +1.5%). At the SASS level the lever fully cashed out — SHF 64→0, HADD2 64→10, net −130 instructions — yet the wall clock did not move: the decode instructions had been executing in the IMMA's shadow all along; cutting them frees no critical resource. Commit: 6112db3 (docs record only; the code change was reverted, cmp-verified identical to HEAD, no trace in the current tree). Date: 2026-09-05 (the work ran late on 09-04 into the morning of 09-05).

1. Background — where things stood

r34 (2026-09-04) had just landed the quantize-transpose prepass, P6's largest single-mechanism gain: +9.72% (1364.2 → 1496.8 tok/s @ the 3354-tok window). Its mechanism moved the A-side layout transform wholesale out of the kernel — quantize_q8_0_pad40_t pre-produces the qa8 / packed-d|ssum planes transposed into the mma consumption layout, and mmq_raw_nb_bt_kernel's A staging degenerated from "per-element index math + swizzle" to a bulk LDG→STS. r32's census had said staging was 21% of kernel instructions; after r34 that path is "zero-math".

So where is the residual? r35's first step was a new four-region instruction census of the post-r34 BT kernel: prolog+stage0 387 / in-loop staging 286 / compute-kd 1514 / epilogue 117. compute-kd is 66% — the only heavyweight region. And inside compute-kd, the largest "nameable, cuttable" block is the sda d|ssum decode: roughly 64 SHF sign-extensions + 64 HADD2s (f16→f32) + 64 I2FPs (i32→f32) per kt — all spent restoring, on the spot, the d (f16, low 16 bits) and ssum (i16, high 16 bits) packed into one u32 into the f32/i32 operands the mma rescale needs.

Hence r35's hypothesis: the decode is pure ALU, and the prepass runs once per token while the kernel runs it per (block, kt) — hoist it into the prepass and the compute loop's each level nets ~192 fewer ALU instructions; the wall clock should move. The intuition is not baseless: r22 had proven that "per-ldmatrix address ALU eats the freed wavefronts", showing ALU had genuinely blocked progress in this class of kernel. Without this step the census stops at an unverified named candidate; with it, win or lose, the boundary of the "composition residual" gets drawn tighter (the other candidate, A-frag LDSM, is left for r36).

2. Principle — the GPU mechanism

First dissect the "decode" chain to SASS grain. In the A-side scale plane, each (token, chunk) is one u32:

bit  0..15 : d     (f16 bit pattern, the quantization scale)
bit 16..31 : ssum  (i16, the sum of the chunk's 32 q8 values, used by the rank-1 term)

The rescale needs d as f32 and ssum as i32. Recovering them from the packed u32 is three SASS instructions: SHF.R.S32.HI (arithmetic-shift the high half-word into a sign-extended i32), HADD2.F32 (f16→f32 conversion, taking one of HADD2's dual half-word lanes), and I2FP (i32→f32). All of them issue on the FP/INT pipe.

But the compute loop's master is the tensor pipe: per kt a warp issues 64 IMMA (mma.m16n8k32.s8) — the throughput mainstay confirmed repeatedly since r12. The key is that on GB10's SMs the FP/INT pipe and the tensor pipe are parallel issue resources — once the 64 IMMA fill the tensor pipe, the compiler scheduler stuffs the data-independent ALU (no dependency on the mmas) into the idle issue slots between them. These instructions occupy slots that would otherwise idle and lengthen no critical path. r35, after the fact, named this phenomenon "the IMMA shadow": instructions in the shadow are free.

The predecode scheme's benefit and costs are both easy to compute:

  • Benefit: 64 SHF + 64 HADD2 + 64 I2FP ≈ 192 fewer ALU per kt in the compute loop.
  • Cost 1: the scale plane goes from 4 B/token/chunk to 8 B (an f32 d plane + an i32 ssum plane), smem 43,008 → 45,056 B, and the read side's LDS instruction count rises (one packed u32 plane becomes two planes).
  • Cost 2: addressing two planes occupies more registers than one packed plane.

If the ALU pipe is the bottleneck, benefit > cost; if the IMMA pipe is the bottleneck and the ALU is in the shadow, benefit = 0 while the costs are paid in full — exactly the two worlds this experiment had to distinguish.

3. Implementation

3.1 Design choices (why this shape)

  • Change the prepass, not the B side. d/ssum are A-side (activation-quantization) properties, and the prepass quantize_q8_0_pad40_t already computes them — merely packed as f16|i16. The hoisting direction is natural: once per token in the prepass vs once per (block, kt) in the kernel.
  • The layout skeleton stays. Keep r31's q-major region-split addressing (already proven bank-conflict-free); only swap "one packed u32 plane" for "[d f32][ssum i32] two planes", growing each chunk's smem scale region from 256 B to 512 B.
  • Budget first, act second. The r34 kernel's smem = qa8 (8×64×32 = 16,384) + sda_q (8×64×4 = 2,048) + qb_raw (128×128 = 16,384) + sds (8×128×8 = 8,192) = 43,008 B; the scale plane's +2,048 B brings 45,056 B, still inside the 2-blocks/SM dynamic smem budget. On registers, r34 is 103 regs / 0 spill and the planar addressing was expected to add ~10 — as long as it stays under 128, 2 blocks/SM survives (256 thr × 128 regs × 2 blocks = 65,536, exactly the per-SM register-file ceiling).

3.2 Key code

BEFORE (an era-tree excerpt; after r35's revert the current tree matches this) — the decode segment in the compute loop that r35 wanted to delete (mmq_raw_nb_bt_kernel, executed per kt):

const uint32_t* sda_blk = sda_q + (size_t)kd * MMQ_NBI
                          + (size_t)(lane >> 2) * 4;
const uint4 s0 = *(const uint4*)(sda_blk);
const uint4 s1 = *(const uint4*)(sda_blk + 32);
#pragma unroll
for (int g = 0; g < 4; g++) {
    float da_q[2];
    int sa_q[2];
    const unsigned w0 = g == 0 ? s0.x : (g == 1 ? s0.z : (g == 2 ? s1.x : s1.z));
    const unsigned w1 = g == 0 ? s0.y : (g == 1 ? s0.w : (g == 2 ? s1.y : s1.w));
    da_q[0] = h2f((unsigned short)(w0 & 0xFFFF));  // f16 d → f32   (HADD2.F32)
    sa_q[0] = (int)(short)(w0 >> 16);              // i16 ssum → i32 (SHF.R.S32.HI)
    da_q[1] = h2f((unsigned short)(w1 & 0xFFFF));
    sa_q[1] = (int)(short)(w1 >> 16);              //                (I2FP at the dma multiply)
    const float dma[2] = { da_q[0] * (float)sa_q[0],
                           da_q[1] * (float)sa_q[1] };
    /* … rank-1 + main-term rescale, multiplied by dsv/dmv and accumulated into sum[] … */
}

r35's AFTER shape (reverted; narrated from the record): on the prepass side, the packing segment above

// Packed d|ssum (r31 Q-major region split of the old sda_q).
const int g = r >> 4, t15 = r & 15, q = t15 & 7, half = t15 >> 3;
const int rg = g >> 1, gsel = g & 1;
size_t sbase = ((size_t)tb * nchunk + b) * MMQ_A_SDASZ
               + (rg * 32 + q * 4 + gsel * 2 + half) * 4;
__half dh = __float2half(d);
uint16_t dbits = *reinterpret_cast<uint16_t*>(&dh);
*reinterpret_cast<uint32_t*>(ysda + sbase) =
    (uint32_t)dbits | ((uint32_t)(uint16_t)ssum << 16);

became a direct write of two planes (*(float*) for d, *(int*) for ssum — 8 B/token/chunk, with f32/i32 exact representations of the quantized values), and the kernel side read straight into da_q[]/sa_q[] — the h2f/sign-extension trio vanishing from the source.

3.3 Pitfalls

  • The smem budget was being squeezed a third time. r31 had already compressed sda_q from 4,096 to 2,048 B, and r35 added 2,048 B back — an "smem for ALU" trade needs a budget table; 45,056 B is not far from the 2-blocks/SM ceiling, and any further inflation would push occupancy down (the r38-era KDR=4 lesson is the same mechanism in reverse).
  • Deleted instructions grow back elsewhere. SASS showed LDS.128 32→48: with one packed u32 plane becoming two, each (kt, lane) issues more scale reads. "Net effect" must be read in both directions — this round netted −130, but structurally the LDS increase is half of the later tie explanation.
  • Registers rose instead of falling (103 → 113). Planar addressing added arithmetic; ptxas spent 10 more registers — still under the 128 red line, but the "deleting code = saving registers" intuition does not hold.

4. Verification

  • Parity dump 1/0 + greedy-32 byte-identical: an f32 d plane replacing "f16 pack → in-kernel h2f" is theoretically bit-exact (the f32 value is __half2float's exact expansion), but the unchanged rescale multiply order had to be proven — defends against numeric-path drift.
  • SASS census (cuobjdump) before/after: proves the change actually landed in the SASS (SHF 64→0, HADD2 64→10, LDS.128 32→48, net −130) — defends against the r30-style "cancelled out by the compiler's CSE; wrote an air lever"; r30's lesson is precisely to read the SASS before concluding, and this time the SASS really did change.
  • 5-round interleaved A/B, median taken: defends against co-tenant drift reading noise as signal (the campaign's standard practice before r59b).
  • cmp-verified = HEAD: confirms the post-revert working tree is byte-identical to the pre-change state.

5. Results (with the veto mechanism)

Layerbefore → afterVerdict
SASS SHF64 → 0lever cashed out
SASS HADD264 → 10lever cashed out
SASS LDS.12832 → 48cost cashed out
SASS net instructions2304 → 2174 (−130)lever cashed out
ptxas113 regs / 0 spill, smem 45,056 B, 2 blocks/SMbudget held
Parity / greedy-321/0 / byte-identicalnumerics clean
Wall (5-round median)1493.2 → 1486.3 tok/s (−0.46%)inside the noise band, NEUTRAL

Veto mechanism: 128 int/fp ALU instructions deleted, the wall clock moved 0.0%, and 16 more LDS.128 were paid. The only self-consistent explanation: these decode instructions were scheduled in the idle FP/INT issue slots between the 64-per-kt IMMA (the IMMA shadow) all along — they never competed with the tensor pipe for any resource; delete what is free and of course the wall does not move, and the LDS increase was absorbed by same-band noise. The compute loop is therefore not ALU-bound, and the residual is the inherent combination of A-frag LDSM consumption + fp rescale — closing the loop with r32/r33's conclusions; the "decode instruction class" is closed as a wall-clock lever from here on.

When a retry is worthwhile: (a) if some tiling rework makes the loop no longer tensor-bound (e.g. enlarging A-frag reuse so IMMA pressure drops and the ALU surfaces), predecode becomes a candidate again; (b) the B side's analogous decode (d/dmin per super-block) later took a different road — r56/r59 pre-multiplied d·sc into a registration-time f32 plane (W_dsc), and that one succeeded because it paired with the cp.async pipeline and the B-side decode genuinely competed with the mma — the same word "predecode", but completely different criteria (in-shadow or not).

6. Lessons

  1. Instructions hiding in the tensor-core's shadow are free — cut only work that competes with the mma pipe; cutting shadow instructions merely vacates idle issue slots.
  2. Instruction count is not the wall clock: −130 SASS = 0.0% wall; ask "which pipe do these instructions run on, and is that pipe saturated" before deciding to act.
  3. "Delete instructions" levers tend to grow bytes back elsewhere (LDS.128 32→48) — a census must always read the net effect; a single class's decrease may be a transfer, not an elimination.
  4. A total SASS win + a wall-clock tie is a high-value measurement: it narrows the composition residual to the LDSM/IMMA + fp-rescale body itself, directly framing r36's wavefront-economics question.

← 37-r34-quantize-transpose-prepass · Index · 39-r36-a-frag-wavefront →

39 · r36 — A-frag wavefront economics: H1 falsified (MEAS-ONLY, no code change)

Result: ncu injected into the bt kernel successfully for the first time: minfer 6.156 vs llama 3.507 shared wavefronts/IMMA (1.76×), with the LDSM share exactly 4.00× — but minfer simultaneously runs 1.85× the wavefronts/s (39.0 vs 21.1 G/s) and a level IMMA rate (6.33 vs 6.02 G-IMMA/s). If MIO were truly scarce, both kernels would be capped by the same wf/s ceiling — the causal chain breaks, H1 (the LDSM→plain-LDS swap) is falsified; H2 (issue-slot deficit) does not exist either. No code change; the bt kernel is at per-IMMA parity with llama. Commit: f44fc44 (docs-only; HEAD untouched). Date: 2026-09-05.

1. Background — where things stood

r35 had struck the entire "decode ALU" instruction class off the wall-clock lever list: −130 SASS instructions bought 0.0% wall clock, the mechanism being that those instructions hide in the IMMA's shadow. With that, only one named candidate in the r34 kernel's census had never been tested under its own metric: the A-frag (activation fragment) LDSM supply.

The candidate's provenance: r32/r33/r35 repeatedly described the residual as "the inherent combination of A-frag LDSM consumption + fp rescale". And llama.cpp's MMQ kernel offers a glaring contrast — its B-frags (weights) go through load_generic (plain LDS), the source comment saying outright "faster than load_ldmatrix", while minfer's A-frags all go through ldmatrix.sync.aligned.m8n8.x4. Hence this round's H1: swapping minfer's A-frag supply from LDSM to plain-LDS (load_generic-style) should reduce shared-memory wavefronts and thereby speed up this IMMA-bound loop — with the implicit premise that MIO (the shared-memory wavefront channel) is a scarce resource. Standing beside it, H2 was the era's other bottleneck theory: the loop is issue-slot-deficit (insufficient per-cycle issue rate) bound.

This round also had a debt to repay: r34/r35's ncu censuses were both missing due to platform injection failures (r34 landed on the regs/prepass/wall triple of indirect evidence). For the wavefront argument to stand, the kernel's real counters had to be obtained first.

2. Principle — the GPU mechanism

What a wavefront is. The L1TEX data pipe services shared-memory accesses in wavefronts: one warp-level instruction splits into several 128 B service cycles. A conflict-free 32-lane LDS is about 1 wavefront; one ldmatrix.m8n8.x4 moves 4 8×8 b16 tiles = 4×128 B = 512 B = 4.0 wavefronts (the record's self-consistency check: 20,873,216 wf / 5,218,304 inst = 4.0, matching bit for bit).

Metric definitions. The canonical ratio for wavefront pressure is

wavefronts/IMMA = Σ smsp__sass_l1tex_data_pipe_lsu_wavefronts_mem_shared_op_{ldsm,ld,st}
                  ÷ smsp__inst_executed_pipe_tensor_subpipe_imma

The denominator is IMMA (not thread instructions) because this loop's "output" is the tensor-core multiply-accumulate; wavefronts are the input serving it.

The scarcity test (this round's methodological core). "Consumes more of a resource per unit of output" and "is limited by that resource" are two different propositions. The MIO pipe's service rate is a fixed per-SM value — if it were scarce, any kernel would be capped by the same wavefronts/s ceiling, and whoever hits the line first stops. So the criterion lives on the throughput side:

wavefronts/s = (wavefronts/IMMA) × (IMMA/s)

If minfer's wf/IMMA is 1.76× while the IMMA rate is level, then wf/s must be 1.76× — running 1.85× above llama's operating point without slowing down means the ceiling is still far away and MIO has plenty of headroom. This "throughput accounting" (work/op → work/s → compare against the ceiling) is the entire mechanism by which this round falsifies H1.

Why "changing the access style" cannot save wavefronts. H1's intuition comes from llama's plain-LDS being "faster". But an LDSM.x4 moves 512 B = 4 wf, and a conflict-free plain-LDS over the same 512 B is still 4 wf — same bytes → same wavefronts, unless the geometry changes. H1's original argument compared LDSM.x4 (4 wf) against "a single tile's plain-LDS" (1 wf) — apples to oranges; llama's own plain-LDS averages 1.60 wf/inst too (B-frags via load_generic, equally not single-wavefront).

llama's real advantage: the A-frag reuse rate. Per 32-k chunk, llama loads 8 A-frags that serve 64 mmas (0.125 LDSM/IMMA); minfer loads 4 A-frags, each serving only 2 mmas (0.500 LDSM/IMMA) — the 4× reuse gap comes from the warp division of labor: llama's warp iterates od inside the kernel (the same A-frag spans 8 od-tile columns), while minfer's warp owns a single 16-od narrow strip (j0w = warp * 16), so each A-frag naturally spans fewer mmas. Reuse is a tiling property, not an access-style property — the root cause no LDS swap can fix.

3. Implementation

3.1 Design choices (why measure instead of write)

  • Unlock injection first, numbers later. r34/r35's ncu failure was fixed in r36's first step: running ncu under the sudo -n env LD_LIBRARY_PATH=... prefix made injection succeed; without the prefix the driver returns ERR_NVGPUCTRPERM / "Unknown Error on device 0" — exactly the wall r34/r35 hit. Every later P6 attribution round reuses this prefix.
  • Try H1 under its own metric. The earlier rounds' lessons (r30: SASS first; r32/r33: the compiler has often already done what you intended) all point at the same principle — a named candidate must first prove a gap on the counter where it could win, then prove the gap is the bottleneck; only passing both gates earns code. H1 passed the first gate (wf/IMMA 1.76×) and died on the spot at the second (throughput cap).
  • Contrast configuration: minfer = qwen2.5-7b q4_k_m, nt=3325 prefill (the bt path); llama = mul_mat_q, llama-bench -p 512, launch 1 each. The nt values were not paired — the hazard did not surface this round; r37's matched-nt re-test exposed it (see §5's correction note).

3.2 Key code

The accused A-frag supply (era-tree mmq_raw_nb_bt_kernel; r36 changed no code, and the current tree has evolved past r59 — this is the round's form):

const int j0w = warp * 16;          // the warp owns a single 16-od narrow strip ← the structural root of low reuse
…
const unsigned l12m = (unsigned)(lane & 12) * 32;
const unsigned grc = (unsigned)(((lane & 3) << 1) + ((lane >> 4) & 1)
                         ^ ((lane >> 2) & 3)) << 4;
unsigned G[4];                       // r22: precomputed 8 A-frag byte offsets (XOR swizzle)
#pragma unroll
for (int g = 0; g < 4; g++)
    G[g] = (unsigned)g * 512 + l12m + ((g & 1) ? (grc ^ 64u) : grc);
…
for (int kt = 0; kt < nktile; ++kt) {
    …
    int a[4][4], b[2][2];
    #pragma unroll
    for (int g = 0; g < 4; g++) {            // 4 ldmatrix.x4 per chunk
        const uint8_t* p = qat + G[g];
        unsigned r0_, r1_, r2_, r3_;
        asm volatile(
            "ldmatrix.sync.aligned.m8n8.x4.shared.b16 "
            "{%0,%1,%2,%3}, [%4];\n"
            : "=r"(r0_), "=r"(r1_), "=r"(r2_), "=r"(r3_)
            : "r"((unsigned)__cvta_generic_to_shared(p)));
        a[g][0] = (int)r0_; a[g][1] = (int)r1_;
        a[g][2] = (int)r2_; a[g][3] = (int)r3_;
    }
    /* 8 mmas per chunk: mmq_mma_k32(clow[g][nh], a[g], b[nh]), g=0..3, nh=0..1
       → 4 LDSM / 8 IMMA = 0.500 LDSM/IMMA; each a[g] serves exactly nh's 2 mmas */

Arithmetic re-check: 4 ldmatrix.x4/chunk × 4 wf = 16 wf over 8 mmas → 2.000 LDSM-wf/IMMA, matching the ncu reading bit for bit; on llama's side 0.125 LDSM-inst/IMMA × 4 = 0.500 wf/IMMA. The entire 4× gap is derivable from the warp division of labor — H1's "access style" narrative has no footing in the code.

3.3 Pitfalls

  • ERR_NVGPUCTRPERM: performance counters need driver-level permission; plain injection fails outright with "Unknown Error on device 0". The sudo -n env LD_LIBRARY_PATH=... prefix is the machine-reproducible fix (permissions + the injection library path, fixed together).
  • H1's original argument was apples-to-oranges: LDSM.x4 (4 tiles, 4 wf) versus a single-tile plain-LDS (1 wf) — payloads differ 4×, so the comparison is meaningless; a fair comparison must lock the bytes.
  • The unpaired-nt measurement planted a landmine: minfer@3325 vs llama@512 mixed two nt scales. This round's per-IMMA conclusion (1.05×) stands, but r37 revealed it holds only at prefill-scale nt — at short nt the tile-prologue amortization degrades bt's wall/IMMA to 1.43×. Cross-kernel comparisons must always lock nt.

4. Verification

  • Metric self-consistency check: minfer LDSM inst 0.500/IMMA × 4 wf = 2.000 LDSM-wf/IMMA, matching the counter reading bit for bit — rules out a counter piped to the wrong pipe.
  • launch 1 each: a single launch, no re-entry — rules out CUDA Graph / prewarm polluting the counts.
  • Throughput and ratio cross-validation: 6.156 wf/IMMA × 6.33 G-IMMA/s = 38.97 ≈ 39.0 G wf/s — two independent sources (the ratio and the throughput) close.
  • issue_active cross-check of H2: 0.457 vs 0.365 (minfer higher) — the issue-slot deficit does not exist either; both bottleneck theories cleared in one round.

5. Results

Wavefronts per IMMA (minfer bt @nt=3325 vs llama mul_mat_q @nt=512):

per-IMMAminfer btllamaratio
LDSM wf2.0000.5004.00×
LDS wf3.5002.1631.62×
ST wf0.6560.8440.78×
total wf6.1563.5071.76×
LDSM inst0.5000.1254.00×
LDS inst0.7501.3490.56×

The throughput side (the three-line ledger that falsifies H1):

minfer btllamaratio
wavefronts/s39.0 G/s21.1 G/s1.85×
IMMA/s6.33 G/s6.02 G/s1.05× (level)
issue_active/cyc/sched0.4570.365minfer higher

The ruling: minfer sustains its per-IMMA tensor rate while running 1.85× the wavefronts/s — MIO is ~1.85× away from its cap and is not a scarce resource; the extra wavefronts hide in the tensor shadow (the same mechanism as r35's ALU shadow — a second consecutive round of "delete/swap supply → no wall-clock response"). H1 falsified; H2 at its endpoint (no deficit). The bt mma kernel is at per-IMMA parity with llama, and the remaining prefill gap lies outside addressable mma-structure levers (per-tile prologue/wave amortization, the quantize prepass, fixup). No code change; HEAD untouched.

The r37 correction note (a boundary condition for later readers): the two sides of the table above ran different nt (3325 vs 512). r37's matched-nt re-test showed bt's wall/IMMA at llama's 1.43× when nt≈511 — per-IMMA parity holds only at prefill-scale nt; short nt is diluted by the prologue. Always cite this round's "per-IMMA parity" conclusion with that nt clause attached.

6. Lessons

  1. The throughput-accounting method: a high work/op ≠ being limited by the resource behind that op; compute work/s first, then compare against the resource's fixed-rate ceiling — "uses a lot" and "is stuck on it" differ by exactly one capping test.
  2. H1-class access-style swaps die of byte equality: LDSM.x4 and a conflict-free plain-LDS over the same payload are both 512 B = 4 wavefronts; an access-style swap has wavefront effects only when the geometry changes.
  3. A-frag reuse is a tiling property: the 0.125 vs 0.500 LDSM/IMMA gap comes from whether the warp iterates od inside the kernel (llama) or each guards a 16-od narrow strip (minfer) — moving it requires re-cutting the warp division of labor, not swapping a load instruction.
  4. After two consecutive same-mechanism falsifications (r35's ALU shadow, r36's wavefront shadow), the "find a faster supply inside the bt kernel" line can close as a whole — the wall-clock gap must live outside the kernel or in other kernels, which is exactly where r37's whole-wall attribution starts.

← 38-r35-scale-predecode · Index · 40-r37-post-parity-attribution →

40 · r37 — Post-parity whole-prefill attribution: the wall clock re-decomposed (MEAS-ONLY, no code change)

Result: 3325-tok prefill, the all-four-gates BT path: GPU busy 2139.6 ms / wall 2190 ms = 1521 tok/s, vs-llama 2.15× (the first whole-wall number on the 3325-eq anchor). Decomposition: q6_K GEMM 1094.7 ms = 51.2% (368.9 µs/GMAC vs llama 57.8 → 6.38×), q4_K bt 600.0 ms (28.0%, 1.15× = parity); the q6_K ffn_down class alone is 1063.5 ms (48.6%) — larger than the entire q4_K bt GEMM (600 ms). Priority queue: ① put q6_K on a raw-byte kernel (ceiling ~2720 tok/s) ② q4_K short-nt amortization ③ FA structure. No code change. Commit: ea234f1 (docs-only). Date: 2026-09-05.

1. Background — where things stood

r34–r36 had pushed the q4_K BT-kernel line to its end: r34's quantize-transpose prepass +9.72% (1364.2 → 1496.8 tok/s), while r35 and r36 falsified in succession the hypothesis that "cuttable supply remains inside the kernel" — the ALU hides in the IMMA's shadow (r35), the wavefronts hide in the same shadow (r36), and the bt kernel reached per-IMMA parity with llama's mul_mat_q (6.33 vs 6.02 G-IMMA/s). The campaign was on the opt-in MMQ path (MINFER_MMQ=1) at this point, and none of the prior MMQ rows had ever recorded a whole-wall vs-llama number on the 3325-eq anchor (the default f16 path had been stuck at 1.43× since P5).

But the 2.15× whole-wall gap remained — it has to live somewhere. r35/r36's falsifications delivered an exclusionary conclusion: the gap is not in the bt kernel's inner loop. That is precisely the license to turn to whole-wall attribution: since the single-kernel microbenchmark had declared itself "not guilty", the only honest next move was to take the whole wall apart launch by launch and see which kernel class the time actually lands in. It also answers a deeper anxiety: every +2%, +7%, +9% since r12 had been kernel-level A/B, with never a family portrait of the whole prefill wall to answer "who should the next +X% be spent on".

Where this stalls without the step: the candidate list is empty (r32's finite-lever sweep, r33's hybrid port, r35's decode, r36's LDSM all cleared out); continuing to hunt levers inside the q4_K bt kernel is doubling down on a falsified direction; and for the campaign to reach 1.0×, it must first know the composition of the remaining 2.15×.

2. Principle — attribution methodology

This round has no new kernel; its "principle" is three attribution-method decisions, each learned from a real lesson of the previous rounds.

Decision one: whole-wall nsys finds #1; matched-nt pairing sets the per-class ratios. A single-kernel ncu answers "who stalls inside this kernel" but cannot answer "is this kernel worth optimizing". Whole-wall attribution needs two levels of measurement: full-graph nsys buckets GPU busy time by launch (which kernel class ate how many milliseconds), then each kernel class gets its wall/GMAC ratio measured separately at a paired nt. Neither level works alone: without the bucketing you do not know whom to pair; without the pairing you have only absolute milliseconds and cannot tell "model structure" (q6_K simply has more FLOPs) from "implementation gap".

Decision two: wall/GMAC normalization. The only fair cross-kernel, cross-implementation unit is "wall clock per GMAC". minfer and llama run the same GGUF weights over the same shapes → the two sides' GMACs are strictly equal, so the wall ratio is the efficiency ratio. q4_K bt is compared against mul_mat_q<12>, q6_K mmq_nt<7,2> against mul_mat_q<14>, each yielding a dimensionless multiplier, and the classes add up into a GEMM total (84.0 vs 35.2 µs/GMAC = 2.38×).

Decision three: lock nt first, ratios second. r36's lesson was cashed the very next round: it had measured minfer@3325 against llama@512 and read per-IMMA 1.05× — this round's matched-nt re-test (both sides nt≈512, minfer's 511-tok prompt vs llama-bench -p 512 -n 0 -r 1 -t 8) shows bt's wall/IMMA is 1.43× (minfer 379.9 µs / 4.23 G-IMMA/s vs llama 265.7 µs / 6.04 G-IMMA/s). The gap comes from tile prologue / wave amortization: at prefill-scale nt it is diluted across thousands of inner tiles into invisibility; at short nt it stands exposed. Per-IMMA parity is cited with an nt clause from now on.

The wall-relevance precondition check. Kernel time and wall time are separated by the launch gap: this round's GPU busy 2139.6 ms vs wall 2190 ms — a gap of 50.4 ms (2.3%) — busy≈wall, so the kernel-level decomposition is valid as an approximation of the whole wall (this "verify before attributing" move is exactly the vaccine against r46's later FA trap of "the kernel shrank but the wall did not respond").

The GMAC arithmetic cross-check. q6_K ffn_down alone: 1063.5 ms @ 5.5 TFLOPS; q4_K gate/up: 63.6 TFLOPS (the whole bt-kernel class 60.0 TFLOPs) — the same ffn line differs 11.5× per-MAC between the two weight types. That is the quantitative meaning of "structural deficit": it is not that llama's q6_K is inherently dearer (its 57.8 µs/GMAC is only 1.84× its own q4_K's 31.4, consistent with K-quant unpack cost) — it is that minfer's q6_K runs an old road that none of r12–r34 ever modernized.

Why q6_K never benefited from the modernization — the structural inventory. Review every upgrade r12–r34 landed on the q4_K bt kernel and the generic q6_K has none of them: no ldmatrix B-frag supply (qb is a plain array, read word by word), no raw-nibble smem layout (B values are expanded to int at staging), no r18-style bulk copy (every tile redoes the nibble assembly), no r34 prepass transpose (the A side still goes through mmq_stage_a's per-chunk reshuffle) — and it carries an extra layer of q6_K-inherent complexity: 16 16-element sub-blocks (q4_K has 8 32-element sub-blocks) force KSPLIT=2: one mma split into m16n8k16 × 2 with two independent int accumulators (clow/chigh), and smem keeps both sds/sds1 scale planes plus an sdm min plane. It is not "a slow version of the q4_K kernel" — it is a generation-earlier architecture carrying 51.2% of the wall; the attribution round's value is turning this into an indictment with numbers.

3. Implementation

3.1 Design choices (why these two measurements and dgxspark state)

  • The "all four gates" BT path: the attribution subject must be the campaign's best configuration at the time (MMQ on + BT routing on), otherwise the buckets describe a path nobody runs.
  • co-tenant idle @0% verified: the drift lessons of the r12–r25 era (absolute values not comparable across windows) are more lethal in an attribution round — the bucketed milliseconds are to be treated as numbers "comparable against llama in the same window", so the machine must be clean. This round explicitly records co-tenant 0%.
  • The bucketing's practical granularity: the nsys trace aggregates launches by kernel symbol name (not by op semantics), so the "q6_K GEMM" bucket holds all mmq_nt<7,2> launches for both the attn_v and ffn_down weight classes; ffn_down's 13 launches × ~80 ms were split out of that bucket by shape/od. Symbol-name aggregation's blind spot is same-name-different-type — fortunately mmq_nt's TYPE is a template parameter, so the symbol name carries <7,2> and the bucket boundaries are naturally clean.
  • The single-class ncu spot check: a matched q GEMM single launch (grid 8×28, 1,605,632 IMMA bit-identical on both sides — the direct consequence of same weights, same shape) grounds the class-level 6.38× in single-kernel evidence.

3.2 Key code: why q6_K sits on the old road

The routing side (era tree): the bt entry launch_mmq_raw_nb_bt_nt's guards mean only q4_K benefits; every other type falls back to generic in the dispatch's default arm:

extern "C" int launch_mmq_raw_nb_bt_nt(int type_id, …, int kd) {
    (void)type_id;
    if (kd != 8) return 0;              // bt exists only in the KD=8 raw-nibble shape
    if (qa8g == 0 || sdag == 0) return 0;   // no prepass planes → clean fallback
    …
    mmq_raw_nb_bt_kernel<8><<<grid, 256, smem, stream>>>(w, qa8g, sdag, …);
    …
}
// dispatch (q6_K = type 7, falls into default):
default: MMQ_LAUNCH((mmq_nt_kernel<7, 2, false>)); break;   // KSPLIT=2

The generic q6_K re-derives the B values from the raw weights for every tile it consumes — mmq_stage_b<7>'s per-tile redo (against bt's "bulk-copy the raw bytes + unpack in registers at mma time", there is neither a raw-byte bulk path nor a pre-expanded plane here):

int s = (2 * c + half) % 16;                       // q6_K: 16 16-element sub-blocks
int chunk = s >> 3, g = (s >> 1) & 3, is = s & 1;
const uint8_t* ql = blk + chunk * 64 + (g & 1) * 32 + is * 16;  // the low 4-bit plane
const uint8_t* qh = blk + 128 + chunk * 32 + is * 16;           // the high 2-bit plane
#pragma unroll
for (int w = 4 * half; w < 4 * half + 4; w++) {
    …
    uint32_t nib = (g < 2) ? (QL & 0x0F0F0F0Fu) : ((QL >> 4) & 0x0F0F0F0Fu);
    uint32_t hi  = ((QH >> (2 * g)) & 0x03030303u) << 4;        // high-bit assembly
    qb[r * MMQ_WS + w] = __vsubss4((int)(nib | hi), 0x20202020); // −32 → signed 6-bit
}
if (half == 0) {                                   // each row also carries two 8-bit scales
    ds[r]  = d * (float)(int8_t)blk[192 + (2 * c) % 16];
    ds1[r] = d * (float)(int8_t)blk[192 + (2 * c + 1) % 16];
}

Add the kernel header's double-buffered smem layout (the seven planes qa/qb/ssa/sda/sds/sds1/sdm) and KSPLIT=2 (two mma.m16n8k16 + independent accumulators, because q6_K has a scale group every 16 elements): this is the shape that none of r12–r34's modernizations (ldmatrix, raw-nibble, BT, prepass) ever touched. The attribution round's whole meaning is to turn that "never modernized" into an indictment with milliseconds attached.

3.3 Pitfalls

  • r36's mixed-nt measurement was corrected this round: per-IMMA 1.05× (mixed nt) → wall/IMMA 1.43× (matched nt≈511). The correction itself went into the record — not papering over a previous round's measurement defect is why this record can be trusted.
  • busy ≠ wall must be verified first: the 50.4 ms launch-gap deficit (2.3%) is on record, and only then was the kernel-level decomposition allowed to approximate the whole wall; r46 later added a dedicated counter-example check for "kernel faster, wall unmoved".
  • The absolute-millisecond window changes with this round: 3325-tok (3325-eq) becomes the MMQ era's anchor; the earlier @2K/@3354 numbers are no longer directly comparable (machine state + anchor both changed).

4. Verification

  • co-tenant idle @0%: defends against machine drift reading the bucketed numbers as implementation gaps.
  • matched-nt pairing (minfer's 511-tok vs llama-bench -p 512 -n 0 -r 1 -t 8, launch 1 each): defends against nt amortization polluting the class-level ratios — r36's lesson promoted to a gate.
  • identical-weights → identical-GMAC: validates the precondition of the wall/GMAC normalization (grid 8×28, 1,605,632 IMMA identical on both sides).
  • Two-level cross-check: the full-graph bucketing (51.2%) and the matched-nt ratio (6.38×) point at the same culprit, and the inequality that q6_K ffn_down alone (1063.5 ms) exceeds the entire q4_K bt GEMM (600.0 ms) orders the priorities — the conclusion does not rest on a single measurement.

5. Results

Whole-wall decomposition (full-graph nsys, GPU busy 2139.6 ms, wall 2190 ms = 1521 tok/s):

Kernel classTimeShare of busyNotes
GEMM q6_K (attn_v + ffn_down, generic mmq_nt<7,2>)1094.7 ms51.2%368.9 µs/GMAC
GEMM q4_K (bt)600.0 ms28.0%60.0 TFLOPs
FA prefill attention124.7 ms5.8%5.7×
quantize prepass86.8 ms4.1%1.25× llama (partly a per-shared-A 2× redundancy, fixed only in r49)
swiglu~3.9%
rest~7%

Within that, q6_K ffn_down alone, 13 launches × ~80 ms = 1063.5 ms (48.6% of the whole wall), @5.5 TFLOPS — larger than the entire q4_K bt GEMM.

matched-nt wall/GMAC (this round's headline table):

GEMM classminferllamaratio
q4_K (bt vs mul_mat_q<12>)36.2 µs/GMAC31.4 µs/GMAC1.15× (parity-grade)
q6_K (mmq_nt<7> vs mul_mat_q<14>)368.9 µs/GMAC57.8 µs/GMAC6.38×
GEMM total84.0 µs/GMAC35.2 µs/GMAC2.38×

The priority queue (ordered by recoverable milliseconds): ① put q6_K on a raw byte-width kernel — at the q4_K bt rate of 63.6 TFLOPs, about −970 ms → ceiling ~2720 tok/s; ② q4_K short-nt amortization (1.43× @511, 1.15× at prefill nt); ③ FA structure (124.7 ms, 5.7×).

The ceiling number's arithmetic chain is worth walking in full, because it is the priority's price tag: lowering q6_K's two GEMMs (1094.7 ms) to the q4_K bt per-GMAC rate (368.9 → ~60 µs/GMAC, a 6.15× speedup) saves about 1094.7 − 1094.7/6.15 ≈ −917 ms; rounding up with the same bucket's dsc/stage residue gives ~−970 ms; wall 2190 − 970 = 1220 ms → 3325 tok ÷ 1.220 s = ~2725 ≈ 2720 tok/s. All three steps use this round's measured milliseconds, with no extrapolated parameters — r47's later measurement (q6_K 196.4 ms, wall 1274 ms) landed in the same interval, showing the pricing was conservative.

Priority ① delivered, plus the hidden tax (looking back from r38–r41): the raw-byte port chartered here was delivered by Era D's four-hit combo, round by round:

RoundLeverWhole wallq6_K class-level metric
r38q6_K BT-style raw-byte mma kernel (KSPLIT=2, KDR=4)+2.87%368.9 → 221.8 µs/GMAC (1.66×)
r39KDR=2 double buffering (A+B pipelined)+13.3%attn_v kernel −19.7%
r40__launch_bounds__(256,3) third resident block+13.0%kernel −23%
r41B-expand widened to uint4 groups+30.7%kernel 1.70 → 0.654 ms (−61.5%)

Together they cut the q6_K GEMM from 1094.7 ms to 196.4 ms (r47's re-measurement) — but once q6_K goes down the BT path it must pass through the same quantize-transpose prepass, so the prepass grew +31.6 ms (the hidden tax); the net q6_K wall-clock gain = +866.7 ms. Lesson: a landing's benefit accounting must subtract every hidden tax it creates, otherwise the next attribution round "loses" some of the already-delivered milliseconds.

Why wall decompositions expire (looking back from r47): r37's table, re-measured at r47, went "stale exactly as predicted" — q6_K 51.2% → 15.8%, wall 2190 → 1274 ms (1521 → 2610 tok/s), vs-llama 2.15× → 1.27×, and FA topped the table as the #1 structural residual (125.8 ms, 5.72×). A decomposition table is a snapshot of "the current kernel mix": every landed lever re-ranks it. So the correct way to attribute is to re-measure the whole wall every time one line converges, and any citation of an old decomposition must carry its version number.

Campaign state at r37: 1521 tok/s, 2.15× vs-llama.

6. Lessons

  1. Re-attribute the whole wall after each line converges, then pick the lever — after r35/r36's consecutive falsifications, the only way out was to ask "where in the whole wall is the gap", and the answer (51.2% in another kernel) is forever invisible from a single-kernel viewpoint.
  2. Optimizing one weight-type line exposes the next: the MMQ modernization made q4_K reach parity, which cast the never-modernized q6_K as "slower than f16" — a lever's value is relative, decided by the wall's composition.
  3. Lock the variables before normalizing: mixed-nt turned per-IMMA 1.05× into wall 1.43×; matched controls are the ticket into cross-kernel comparison (r36's measurement defect was promoted to a gate this round).
  4. A landing's net gain = the milliseconds won − the hidden taxes (prepass +31.6 ms), and decomposition tables have a shelf life — together these are the complete answer to "why wall decompositions expire".

← 39-r36-a-frag-wavefront · Index · 41-r38-q6k-bt-rawbyte-mma →

41 · r38 — q6_K BT-style raw-byte mma kernel (LANDED)

Result: 7B whole-prefill 1518.4 → 1561.9 tok/s (+2.87%, 3/3 pairs); matched-nt q6_K GEMM 368.9 → 221.8 µs/GMAC (1.66×). Commit: 75aabb9. Date: 2026-09-05.

1. Background — where things stood

On 2026-09-05 the P6 MMQ campaign reached r38 — the starting point of Era D (q6_K / FA / prepass / promotion, r38–r60).

Before this, r28–r34 had already brought the q4_K MMQ line to convergence: r28 traded a raw-nibble B side plus lower smem for 2 blocks/SM, r29's kd-loop unroll cashed in the instruction-count reduction once those 2 blocks/SM were in place, r31 removed the bank conflict on scale reads, and r34 moved the A-side layout transform out of the kernel entirely (the quantize-transpose prepass, +9.72%). r37 then ran a post-parity whole-prefill attribution (master table row 51): whole-prefill 1521 tok/s, 2.15× vs llama.cpp — but the same attribution table also showed that the q6_K GEMM alone cost 1094.7 ms, 51.2% of the entire prefill wall clock, at 6.38×/GMAC unit throughput (i.e. each GMAC takes 6.38× as long as on the q4_K MMQ line, which sat at ~57.8 µs/GMAC at the time, inferred from 368.9/6.38).

In other words: the MMQ redesign made q4_K fast and left q6_K on a path slower than f16 (row 51's one-line conclusion: "MMQ made q4_K fast and left q6_K on a slower-than-f16 path — the next lever is a different kernel"). In the 7B q4_k_m model the q6_K tensors are attn_v and ffn_down (the r38 commit message verbatim: "q6_K GEMMs (attn_v + ffn_down, the r37 51.2% residual at 368.9 uS/GMAC)"); every prefill passes through them, and a 51.2% share of the wall clock makes them the next biggest lever.

It is worth a quick look at how the q4_K line converged — because r38's kernel shape is its template: r28's raw-nibble B-side kernel (1375.2 → 1410.4, +2.56%, squeezing smem down to 45,056 B to buy 2 blocks/SM); r29's kd-loop unroll (+2.80%, "the instruction cut r25 judged wall-inert became real at 2 blocks/SM"); r31's q-major sda repack (+1.07%, LDS.64 32→0); r34's quantize-transpose prepass (1364.2 → 1496.8, +9.72%, moving the entire A-side layout transform out of the kernel). The common skeleton of all four steps: raw-byte weight streaming, low smem for high residency, index arithmetic kept out of the loop, A side preprocessed by a prepass. r38's job was to prove this skeleton ports wholesale onto q6_K, whose layout is completely different.

Why "a different kernel" instead of tuning mmq_nt<7,2> further? The 6.38× per-GMAC gap is structural: the generic kernel runs every type through the same raw-byte-stream processing on the B side, and q6_K's cross-plane (ql+qh) bitfield reassembly cannot be pushed into that framework without per-element in-loop ALU; the ceiling of parameter tuning is far below moving the reassembly into staging once and for all. r37's attribution table prescribed exactly that: "different kernel".

Why was q6_K so slow? At that point it ran the pre-r28 generic int8-GEMM kernel mmq_nt_kernel<7,2,0>. That kernel copes reasonably with types like q4_K — contiguous nibbles, sparse scales — but q6_K's packing is hostile to a hot loop: a weight value has to be reassembled across two planes, ql (low 4 bits) and qh (high 2 bits), and a scale arrives every 16 elements — all of the staging phase's per-element bitfield work sits inside the hot loop. r37's conclusion was "the next lever is a different kernel": give q6_K a dedicated mma kernel modeled on the r28/r34 shape that worked (raw-byte streaming + an expanded B + the BT A-side shell) and drive the reassembly arithmetic out of the hot loop.

That is r38. First fix a layout misreading, then build the kernel, and finally take an occupancy beating.

2. Principle — the GPU mechanism

2.1 The real q6_K layout (this doc's core correction)

Start from the definition in src/block.rs on the current tree (GGUF/ggml convention: 210 B per block, 256 elements):

#![allow(unused)]
fn main() {
// src/block.rs — Q6_K — 6-bit super-block quantization, 256 elements
// 16 blocks of 16 elements each, effectively 6.5625 bits per weight
#[repr(C)]
pub struct BlockQ6_K {
    pub ql: [u8; 128],    // quants, lower 4 bits (QK_K/2)
    pub qh: [u8; 64],     // quants, upper 2 bits (QK_K/4)
    pub scales: [i8; 16], // scales, quantized with 8 bits (QK_K/16)
    pub d: Fp16,          // super-block scale
}
// size assertion: 2 + 16 + 128 + 64 = 210 B (Q6KB = 210)
}

Three key facts:

  • The sub-block granularity is 16, not 32. scales[16] holds 16 int8 scales, one per 16 elements; d is the f16 common factor of the whole 256-element super-block. The dequantized value is d · sc[sub] · (q − 32), with q the 6-bit code 0..63 (so the expanded value is centered on −32..31).
  • An element's low 4 bits live in ql, its high 2 bits in qh, and the two planes interleave under different bitfield rules (see the expand_q6_elem closed form in §3.2).
  • The kernel's loop unit is still the 32-element chunk (8 per super-block — the task language's QI6_K=8 count: QK_K/32 = 8; llama.cpp's own QI6_K macro normalizes the other way — 32 iterations × 8 elements each, the same 8×32 grid). The pre-r38 working model, "q6_K = 8 32-element sub-blocks" (isomorphic to q4_K's 8×32), is wrong: one 32-element chunk spans two 16-element sub-blocks with different scales, sc[2c] and sc[2c+1].

This correction directly determines the mma structure: mma.m16n8k32 consumes 32 k at a time, but its integer accumulator has no "scale every 16 k" breakpoint; a single scale would multiply one sub-block's contribution by the wrong factor. Hence KSPLIT=2 (a new concept in this doc: split one 32-k chunk into two mma.m16n8k16, k=0..15 and k=16..31, each rescaled by its own 16-element sub-block's scale).

Walk one concrete element through expand_q6_elem (§3.2) to pin down the cross-plane interleaving — take elem = 70 (super-block 0, the 4th 16-element sub-block, the 6th element of the 2nd 32-chunk):

QuantityExpressionValueMeaning
m70 & 316offset within the 32-chunk
it70 >> 70first 128-element half
n70 & 12770offset within the half
ql positionit·64 + (n&63) = 6, shift (70>>6)·4 = 4high nibble of ql[6]low 4 bits
qh positionit·32 + m = 6, shift ((70>>5)&3)·2 = 4bits 4..5 of qh[6]high 2 bits
sub-block scaleelem/16 = 4sc[4]direct evidence of 16-element sub-block granularity

Note that the ql index follows n (offset within the half) while the qh index follows m (offset within the chunk) — two interleavings with different strides are exactly the arithmetic reason q6_K cannot be treated as 8×32.

2.2 Rescale arithmetic and a rounding trap

q4_K's rescale has two terms (d and dmin both multiply the accumulator); q6_K has no dmin, so a single-term rescale sum += da·dsc suffices, where da is the A-side per-token-block quantization scale (the packed d|ssum word from the r34 prepass) and dsc = d · sc[16-sub-block] is the B-side per-16-sub-block scale.

One easy numerical trap: fusing the two 16-sub-block integer accumulators before multiplying by the scale (the fused form) is not the same rounding sequence as two independent +=. Measured: the fused form deviates 1.2e-3 from the CPU reference, just past the 1e-3 parity gate; two independent += pass. Integer accumulators clow/chigh kept separate, f32 accumulation kept separate — that is the form that passes the gate.

2.3 Expanded B and the smem budget

New concept, expanded-B (the expanded B plane): during staging, each super-block's ql+qh bitfield reassembly is computed ahead of time and written into smem as a one-byte-per-element plane of centered int8 (−32..31), so in the hot loop a B fragment is a plain int8 read — ql/qh reassembly, −32 centering, and bitfield shifts all leave the hot loop (the q6_K version of the "index arithmetic stays out of the loop" lesson verified over and over in r21/r22/r31). The cost is smem: the raw nibble stream is 4 bits/element, the expanded int8 plane is 8 bits/element.

New concept, KDR (k-decode rate): how many 32-element chunks each kt iteration stages. KDR=8 = stage an entire 256-element super-block in one go; KDR=4 = half of one. Both the row width of the B expanded plane and every smem plane scale linearly with KDR.

The kernel tile is MMQ_NBI=64 tokens × MMQ_NBJ=128 od rows, 256 threads (8 warps, 16 od rows per warp). The launcher's smem arithmetic (r38 commit text verbatim):

smem = KDR * MMQ_NBI * 32   // qa8      (A's swizzled q8 plane)
     + KDR * MMQ_NBI * 4    // sda_q    (A, one packed d|ssum word per token)
     + MMQ_NBJ * KDR * 32   // qb_exp   (B expanded centered-int8 plane)
     + KDR * MMQ_NBJ * 8    // sds      (per chunk×row float2 dsc0/dsc1)
  • KDR=4: 8192 + 1024 + 16384 + 4096 = 29,696 B → 2 blocks/SM;
  • KDR=8: every term doubles = 59,392 B → 1 block/SM.

This is the occupancy cliff. Spread the warps out: 256 threads = 8 warps per block. GB10's per-SM warp ceiling is 48 (derivable from r40's measurements: 3 blocks × 8 = 24 theoretical warps correspond to 18.12 achieved warps/SM = 37.74%, and 18.12/0.3774 = 48). So:

  • KDR=8 (59,392 B) → 1 block/SM = 8 warps/SM = 16.7% theoretical occupancy;
  • KDR=4 (29,696 B) → 2 blocks/SM = 16 warps/SM = 33.3%.

This is exactly the q4_K line's pre-r28 disease (before r28 the wide kernel ran 1 block/SM; r28 bought 2 blocks with 45,056 B). The depth-vs-occupancy trade-off (r5's lesson) re-appears on every new kernel: KDR=8, staging a whole super-block per iteration, looks like it saves iteration overhead but actually pushes the whole kernel back to latency-bound — r38 paid to reconfirm this lesson with a same-day A/B.

2.4 The mma shape in summary

Per kd (one 32-chunk) per warp:

  • A side: 4 groups of ldmatrix.sync.aligned.m8n8.x4 (g=0..3, covering the 4 16-token sub-tiles of the 64 tokens); the A plane is the q8 pre-transposed by the r34 prepass — identical to the r34 BT shell, weight-type agnostic;
  • B side: 2 n half-tiles (nh=0/1, 8 rows each) × 2 k half-tiles (b[nh][0] = k 0..15 → 16-sub-block 2c, b[nh][1] = k 16..31 → 16-sub-block 2c+1);
  • mma: 4 m sub-tiles × 2 n half-tiles × 2 k half-tiles = 16 mma.m16n8k16, integer accumulators clow[4][2][4] / chigh[4][2][4];
  • epilogue: two independent +=, acc += da·dsc0·clow and acc += da·dsc1·chigh.

3. Implementation

3.1 Design choices (why this shape and not another)

  1. Fix the layout model first, then draw the mma structure. In task order, kernel design started directly after r37's attribution named q6_K; but the true first step was checking block_q6_K's sub-block granularity against the CPU reference — the existence of sc[16] alone kills the "8×32, single-scale" m16n8k32 plan and forces KSPLIT=2. Discovering this after building on the wrong model would mean reworking the entire fragment layout.
  2. Expand B rather than feed ldmatrix raw bits. The q4_K NB kernel (r28) does raw-nibble smem plus small in-loop expansion; q6_K's bitfield reassembly (two planes, two bit widths) is far more expensive than q4_K's nibble split — in the hot loop that is 5+ ALU ops per element. Expanded to centered int8, the B-side hot loop is a straight read; the cost (smem doubling) is absorbed by choosing KDR=4.
  3. Reuse the r34 BT shell verbatim on the A side. A is the activation side, weight-type agnostic: the prepass pre-transposes qa8/sda, and the kernel does bulk uint4 LDG→STS. This concentrates the new kernel's delta on the B side, minimizing the risk surface.
  4. KDR=4, not 8. See the arithmetic in §2.3: 59,392 B at 1 block/SM is a measured regression (1097.8 tok/s); 29,696 B at 2 blocks/SM is what sustains the occupancy r28 bought.
  5. Entry gate and clean fallback. The new kernel runs only when every row's id is a multiple of 256 ((id/32) % 8 == 0, i.e. whole-super-block aligned); if any launcher cap/parameter check fails it returns 0 and the caller falls back cleanly to the generic mmq_nt<7,2> — the GPU-safety rule "capability-gate failure takes an explicit fallback, never a silent downgrade" realized as the fallback path in the type-dispatch layer. Both block_stride forms are supported: the raw 210 B weight stream and 7e②'s 224 B padded repack (parity verified on both sides).

3.2 Key code

The element-expansion closed form (r38 commit 75aabb9, added to src/cuda_kernels.cu; the current tree still carries the same function body). elem is the element number 0..255 within a super-block; ql/qh point at in-block offsets 0 and 128:

__device__ __forceinline__ int expand_q6_elem(const uint8_t* ql, const uint8_t* qh, int elem) {
    int m  = elem & 31;                   // offset within the 32-chunk
    int it = elem >> 7;                   // 0/1: first/second 128-element half
    int n  = elem & 127;                  // offset within the half
    int ql_idx   = it * 64 + (n & 63);
    int ql_shift = (n >> 6) * 4;          // 0 or 4 (low/high nibble)
    int qh_idx   = it * 32 + m;
    int qh_shift = ((n >> 5) & 3) * 2;    // 0,2,4,6 (2-bit fields)
    int v = ((ql[ql_idx] >> ql_shift) & 0x0F)
          | (((qh[qh_idx] >> qh_shift) & 0x03) << 4);
    return v - 32;                        // centered to -32..31
}

Segment by segment: the low 4 bits come from the low/high nibble of ql[it*64 + (n&63)] (bit 6 of n decides which), the high 2 bits from a 2-bit field of qh[it*32 + m] (bits 5..4 of n pick which 2-bit field) — note that qh's index follows m (offset within the chunk) while ql's follows n (offset within the half): that is exactly where the "16-element sub-blocks, ql/qh interleaved at different strides" structure comes from. The −32 centering puts the expanded values directly in the mma-friendly symmetric range.

The B-expansion + dsc staging macro (same commit; called once per kt iteration). The expansion part:

/* ---- B: expand KDR*32-chunk super-block (half at KDR=4) ---- */
const int sb    = ((kt) * KDR) >> 3;          // super-block number this kt covers
const int cbase = ((kt) * KDR) & 7;           // chunk offset within that sb
for (int x = threadIdx.x; x < MMQ_NBJ * (KDR * 32); x += blockDim.x) {
    const int jj  = x / (KDR * 32), bec = x % (KDR * 32);
    const int j   = j0 + jj;                  // od row
    const int elem = cbase * 32 + bec;        // element number within the super-block
    int v = 0;
    if (j < od && sb < nsb) {
        const uint8_t* blk = W + (size_t)j * ((size_t)nsb * bstride)
            + (size_t)sb * bstride;
        v = expand_q6_elem(blk, blk + 128, elem);
    }
    qb_exp[(size_t)jj * (KDR * 32) + bec] = (uint8_t)v;
}

The dsc (= d·sc) pair — a new concept in this doc, dsc: one float2(dsc0, dsc1) per (chunk, row); they are the rescale factors of the two 16-sub-blocks of that 32-chunk:

const uint8_t* blk = W + (size_t)j * ((size_t)nsb * bstride)
    + (size_t)(c >> 3) * bstride;
const float d  = h2f(*(const uint16_t*)(blk + 208));   // f16 super-block factor (offset 208)
const int s0   = 2 * (c & 7);                          // 16-sub-block pair index
dsc0 = d * (float)(int8_t)blk[192 + s0];               // sc[16] starts at offset 192
dsc1 = d * (float)(int8_t)blk[192 + s0 + 1];

The two offsets 192/208 are the layout table from §2.1: scales occupies 192..207 and d occupies 208..209. At KDR=4 one kt touches only half of a super-block (cbase 0..3 or 4..7), hence the "half at KDR=4" comment.

The A-side bulk copy (the r34 BT shell; introduced in r38 and kept unchanged by r39 — the excerpt below is the untouched A portion of the r39 diff, word-identical to its r38 introduction): the A plane is the qa8/sda pre-transposed by the prepass, and staging is pure uint4 copying with zero arithmetic:

/* ---- A: bulk LDG->STS of the pre-transposed qa8/sda (no math) ---- */
const size_t qbase = ((size_t)blockIdx.x * nchunk + (size_t)(kt) * KDR) * MMQ_A_QASZ;
for (int off = threadIdx.x; off < (KDR * MMQ_NBI * 32) / 16; off += blockDim.x)
    ((uint4*)(qa8))[off] = ((const uint4*)(qa8g + qbase))[off];
const size_t sbase = ((size_t)blockIdx.x * nchunk + (size_t)(kt) * KDR) * MMQ_A_SDASZ;
for (int off = threadIdx.x; off < (KDR * MMQ_NBI * 4) / 16; off += blockDim.x)
    ((uint4*)(sda_q))[off] = ((const uint4*)(sdag + sbase))[off];

The per-token packed d|ssum (sda_q, one uint32) moves together with the A plane — the hot loop's rescale needs only that word plus the B-side dsc, and never touches raw activations again. The A side is of one piece with r34's q4_K BT kernel, which is exactly where "the new kernel's delta lives on the B side" lands.

The hot loop's B-fragment read (the entire payoff of expansion — no bitfield ops, no ldmatrix, two int reads):

// B-frag: straight int8 read from the expanded plane. Each mma.k16
// uses b[nh][0] (k=0..15, sub 2c) or b[nh][1] (k=16..31, sub 2c+1).
#pragma unroll
for (int nh = 0; nh < 2; nh++) {
    const int jj = j0w + nh * 8 + (lane >> 2);
    const uint8_t* qs = qb_exp + (size_t)jj * (KDR * 32) + (size_t)kd * 32;
    b[nh][0] = *(const int*)(qs + (lane & 3) * 4);
    b[nh][1] = *(const int*)(qs + 16 + (lane & 3) * 4);
}

The A side takes its fragments with 4 groups of ldmatrix.sync.aligned.m8n8.x4 from the swizzled qa8 plane (byte-identical in origin to the r34 BT kernel), after which 16 mma.m16n8k16 fill the two integer accumulator sets clow/chigh.

The launcher's smem arithmetic and fallback (r38 commit, launch_mmq_raw_nb_bt_q6k_nt):

constexpr int KDR = 4;
const int smem = KDR * MMQ_NBI * 32   // qa8
               + KDR * MMQ_NBI * 4    // sda_q (one uint32 per token)
               + MMQ_NBJ * KDR * 32   // qb_exp (centered int8, half super-block)
               + KDR * MMQ_NBJ * 8;   // sds (float2 dsc pair = 8B)
dim3 grid((nt + MMQ_NBI - 1) / MMQ_NBI, (od + MMQ_NBJ - 1) / MMQ_NBJ);
cudaFuncSetAttribute(&mmq_raw_nb_bt_q6k_kernel<KDR>,
                     cudaFuncAttributeMaxDynamicSharedMemorySize, smem);
...
mmq_raw_nb_bt_q6k_kernel<KDR><<<grid, 256, smem, stream>>>(...);
// any cudaFuncSetAttribute/launch failure: return 0 -> the caller falls back cleanly to mmq_nt<7,2>

At commit time the dispatch gate was type_id == 7 && MINFER_MMQ_Q6K_NB=1 && (id/32)%8 == 0 (opt-in). The same spot on the current tree (src/cuda.rs, around lines 3170-3238) shows the gate's evolved form: r60 flipped MINFER_MMQ_Q6K_NB/MINFER_MMQ_RAW to default-on (opt out with "0"), and r53/r56 added w_exp/w_dsc pre-expanded-plane pointers to the launcher (null on miss, falling back to this doc's in-kernel expansion path) — this doc only concerns r38's null/no-plane form. The gate skeleton follows (current-tree code, comments mark the r38 difference):

#![allow(unused)]
fn main() {
if type_id == 7                              // Q6_K
    && Self::mmq_gate_on("MINFER_MMQ_Q6K_NB")   // r38: opt-in, requires =1
    && Self::mmq_gate_on("MINFER_MMQ_RAW")      // r38: opt-in, requires =1
    && (id / 32) % 8 == 0                       // whole-super-block aligned
{
    let nchunk = (id / 32) as i32;
    ... // launcher returning 0 falls back cleanly to the generic mmq_nt<7,2>
}
}

3.3 Pitfalls

  1. The wrong layout model is the costliest trap. The "q6_K = 8 32-element sub-blocks" model explains a lot of phenomena convincingly (8 chunks, 210 B, 256 elements all check out) — only scales[16]'s 16 does not. An m16n8k32 single-scale plan designed on the wrong model blows up parity at ~1e0 magnitude; checking the layout element by element against the CPU reference is what exposed the 16×16 truth.
  2. Fused rescale rounding overshoots the gate. If the two 16-sub-block contributions are fused into one expression before multiplying by the scale, the deviation from the CPU reference is 1.2e-3, just past the 1e-3 gate; splitting into two independent += is clean. Lesson: the association order of integer accumulators and f32 accumulation is part of the parity contract.
  3. The KDR=8 smem cliff. Whole-super-block staging (59,392 B) is not a "more depth, fewer iterations" micro-optimization — it drops 2 blocks/SM straight back to 1 block/SM, and whole-prefill collapses from the ~1518–1533 band to 1097.8 tok/s. The ghost of the r7–r8 era ("wide KD=8 first measured 2124 = phantom (silent smem-cap failure)") returned from the other direction: this time there was no silent failure — the launch legally succeeded and performance collapsed. A legal cap ≠ a cap worth using.
  4. Unit-test the expansion mapping before integration. The element→super-block mapping (the index arithmetic of expand_q6_elem) ran as a standalone verifier with 0 mismatches before integration (a verbatim application of r28's "validate layout maps standalone before integration" lesson).

4. Verification

  • Layout verifier: the element→super-block mapping of expand_q6_elem ran standalone over all elements, 0 mismatches (defends against the kind of structural error in §3.3#1 sneaking into the kernel).
  • ptxas resource audit: KDR=4 compiles to 85 regs / 0 spill (defends against accidental register pressure breaking the 2 blocks/SM occupancy premise).
  • Parity gate: logits deviation vs the baseline path below the 1e-3 scale, run once for each of the raw 210 B and padded 224 B block_stride forms (defends against numerical regressions in bitfield reassembly/rescaling, and against one of the two weight byte streams going untested — the two layouts differ in offset arithmetic: bstride only changes the row pitch and the in-block 192/208 offsets are shared, but read-side overrun behavior differs, so both must pass).
  • greedy-32 byte-identical: the greedy 32-token output stream matches the pre-change binary byte for byte (defends against "parity numbers pass but argmax flips on a knife edge").
  • Interleaved A/B measurement: same window, same binary, 3/3 pairs positive (defends against fake deltas manufactured by machine-state drift).

5. Results

Metricbefore → afterNote
whole-prefill (7B, same-window A/B median)1518.4 → 1561.9 tok/s (+2.87%, 3/3)master row 52
matched-nt q6_K GEMM368.9 → 221.8 µs/GMAC (1.66×)below the ≥2× project bar
KDR=8 variantregressed to 1097.8 tok/s59,392 B → 1 block/SM, vetoed
ptxas85 regs / 0 spill (KDR=4)occupancy premise holds
vs llama.cpp (3325-eq anchor)2.13× (r37 was 2.15×)the relative value falls as the engine itself gets faster

Three readings:

  1. 1.66× < 2× landed anyway, because the ≥2× bar set at project time meant "same league as the q4_K MMQ line", and the old q6_K path (368.9 µs/GMAC, 6.38×/GMAC) was rotten enough that even 1.66× cut the unit cost by nearly 40%. A strictly-positive change has no reason to stay outside the gate. The master table records this as "LANDED, < 2x".
  2. The occupancy lever is decisive: the same kernel, KDR=8 → 1097.8 (1 block/SM), KDR=4 → +2.87% (2 blocks/SM). The difference between shallow and deep staging is not tuning; it is the difference between 16.7% and 33.3% theoretical occupancy.
  3. The attn_v kernel is still latency-bound (compute only 16.7%) — the hook this doc leaves behind: the kernel got faster but is still waiting; r39's pipelining and r41's load width both start from this number.

q6_K line postscript (this doc is the first step of a four-step arc; all numbers from master table rows 52-55):

Stepq6_K kernelwhole-prefill
r37 baseline368.9 µs/GMAC (6.38×/GMAC)1521 tok/s
r38 (this doc)221.8 µs/GMAC (1.66×)1561.9 (+2.87%)
r39 KDR=2 double bufferattn_v −19.7%1777.5 (+13.3%)
r40 third resident blockkernel −23%2015.6 (+13.0%)
r41 B-expand uint4 widen0.654 ms (−61.5%)2605.2 (+30.7%)

This doc's 1.66× looks like the smallest win, but it established the kernel skeleton the next three steps share — without that skeleton, pipelining, residency, and load width have nowhere to land.

6. Lessons

  1. Check the layout against the CPU reference before designing the mma structure — a single field, scales[16], vetoed the entire 8×32 plan; a layout error reworked at the fragment-layout layer costs everything.
  2. Depth vs occupancy goes up for auction again on every new kernel (r5's lesson, Nth rerun): for a parameter like KDR — "how many k to stage per iteration" — compute smem → blocks/SM first and decide from that, never from intuition.
  3. Land strictly-positive changes even below the project bar: the bar is a guess made at project time; the measured cost of the alternative is the fact. "LANDED, < 2x" is a perfectly legitimate state.
  4. The verifier's place is before integration: a 0-mismatch standalone mapping verifier institutionalizes r28's lesson — a layout bug costs one minute in the verifier, an afternoon in the parity matrix.

← 40 · Index · 42 · r39 q6_K KDR=2 double-buffer →

42 · r39 — q6_K KDR=2 double-buffer (LANDED)

Result: 7B whole-prefill 1568.7 → 1777.5 tok/s (+13.3%); attn_v q6_K kernel −19.7% (2,549,248 → 2,046,848 ns). Commit: f2b9e54. Date: 2026-09-05.

1. Background — where things stood

r38 (the previous doc) had just given q6_K a BT-style raw-byte mma kernel: matched-nt unit cost 368.9 → 221.8 µs/GMAC, whole-prefill +2.87%. But r38's verification data left one conspicuous hook: the attn_v kernel was still latency-bound, with compute at only 16.7% — most of the SM's issue slots were spinning idle; the kernel was waiting on data, not computing.

Waiting on what? r38's main loop was single-buffered:

for (kt...) {
    if (kt > 0) RAW_STAGE_Q6K_BT(kt);   // stage the kt panel this iteration needs
    __syncthreads();                     // whole block waits until staging is complete
    ... mma compute(kt) ...              // only then compute
    __syncthreads();                     // loop-tail fence: protects the single buffer from next round's overwrite
}

Staging (the A-side bulk uint4 copy + the B-side ql/qh bitfield expansion + the dsc reads) is squeezed between two __syncthreads(): every warp's LDG latency and expansion ALU are fully exposed — no compute overlaps them. All warps stage together, wait together for the slowest one, and only then start computing together. Compute at 16.7% is the direct reading of this serial structure: within one kt period, the staging segment is far longer than the compute segment.

The fix was validated once back in r20 (that was the A-side split-phase staging lesson: "the gap carrier is long_scoreboard in the LDG→STS chains"); this time the same idea moves to the q6_K kernel's entire staging layer, pipelining the B-side ALU expansion along with it: double buffering — two copies of every per-kt plane, with kt+1's expansion written into the other buffer while kt computes. The code comments spell out that this is the pipeline scheme the generic kernel mmq_nt<7,2> already had, now ported to the q6_K-specific kernel.

The cost of the serial structure can be computed from r38's shape (in its KDR=4 form, per block per kt iteration; derived from the kernel structure, not a measured breakdown): the staging segment must move the A-side 9,216 B (qa8 8,192 + sda_q 1,024) as uint4s, expand the B-side 128 rows × 128 elements = 16,384 elements (each one ql LDG.U8 + one qh LDG.U8 + bitfield ALU), and read 512 dsc pairs; the compute segment is 8 warps × 4 kd × 16 mma.m16n8k16 = 512 mmas plus the epilogue. The two segments are of the same magnitude, and the serial structure makes them add into the critical path — compute at 16.7% means staging holds an overwhelming share of that path.

Only one question remains: where does the smem come from.

2. Principle — the GPU mechanism

2.1 The fence economics of double buffering

Single-buffered (r38), the critical path per kt:

[all warps] stage(kt)  →  barrier  →  [all warps] compute(kt)  →  barrier
     ~LDG latency + expansion ALU        0 overlap              mma

Double-buffered (r39):

prologue: stage(0 → buf0); barrier
loop kt:  stage(kt+1 → buf^1)   ── concurrent with the line below ──▶  compute(kt on buf)
          (LDG latency hidden under the mma issue stream)
          barrier   ← the loop-tail fence does double duty (see §2.3)

Drawn as a timeline (each cell is one critical-path segment; the two rows are concurrent activities on the same SM):

r38 single-buffer:  ─[stage 0]─🚧─[compute 0]─🚧─[stage 1]─🚧─[compute 1]─🚧─
r39 double-buffer:  ─[stage 0]─🚧─[compute 0]─🚧─[compute 1]─🚧─[compute 2]─🚧─
                          [stage 1 ]─[stage 2 ]─[stage 3 ]      ← riding on the compute segments

In r38's timeline, stage and compute alternate in exclusive possession; in r39 the compute segments stretch out (covering the stages), and total time ≈ stage(prologue) + Σ compute, not Σ(stage + compute).

The essence of the gain is pure overlap: staging's workload is not one byte smaller, compute's workload is not one byte smaller; what changed is only that the two no longer wait on each other. This is r20's conclusion replayed at the staging layer — latency is not eliminated, it is hidden inside issue slots that were idle anyway. The compute 16.7% r38 left behind shows the idle slots were plentiful, so the room for overlap to cash in was large.

A counterexample worth contrasting: the cp.async double buffering tried on the q4_K raw kernel in the r13 era (784786d) measured ~0 and was vetoed — at that time the kernel was L2-throughput bound, and re-timing the same bytes bought nothing. The same mechanism (double-buffered staging) is a dead lever on a throughput-bound kernel and +13% on a latency-bound kernel: the lever is determined by the kernel's bottleneck regime, not by the mechanism itself. The compute 16.7% r38 left behind was the ticket that said "this is a latency regime" before r39 even started.

2.2 The smem budget: price both planes

Double buffering is not free: smem doubles. The form r38 landed was KDR=4 single-buffer = 29,696 B (qa8 8192 + sda_q 1024 + qb_exp 16384 + sds 4096). Keeping KDR=4 and double-buffering directly gives every term ×2 = 59,392 B = 1 block/SM — exactly the cliff r38 had just measured-vetoed with KDR=8 (1097.8 tok/s). The smem budget must price the A and B planes together; looking only at the B expanded plane (qb_exp, the largest term at 16384 B) creates the illusion that "there is still headroom".

The solution is to halve KDR to buy buffers: KDR=2 × double buffer. Per-plane accounting (MMQ_NBI=64, MMQ_NBJ=128; qa8 = KDR·NBI·32, sda_q = KDR·NBI·4, qb_exp = NBJ·KDR·32, sds = KDR·NBJ·8):

Planer38: KDR=4 single bufferVetoed: KDR=4 double bufferr39: KDR=2 double buffer
qa88,19216,3848,192
sda_q1,0242,0481,024
qb_exp16,38432,76816,384
sds4,0968,1924,096
Total29,69659,39229,696

29,696 B — exactly the same footprint as r38's KDR=4 single buffer, 2 blocks/SM preserved, but each iteration now gains real compute/staging overlap. The cost of KDR dropping 4→2: each kt advances only 64 k (two 32-chunks), kt iterations double and mmas per iteration halve — iteration overhead rises slightly and A/B plane reuse drops slightly. Trading the overlap gain against this dilution, the account is positive (§5 measures +13.3%).

2.3 How to place the fences

Double buffering saves one of the two per-iteration fences (plus one in the prologue), but both duties of the remaining fence must hold:

  1. Separate compute(kt−1, buf^1) from stage(kt+1, buf^1) — staging writes exactly the buffer the previous iteration's compute read, so the write must not start until every warp has finished reading;
  2. Separate stage(kt+1, buf^1) from compute(kt+1, buf^1) — compute reads exactly the buffer this iteration's stage wrote, so the read must not start until every warp has finished writing.

One __syncthreads() at the loop tail carries both: it is simultaneously the rendezvous for "the previous round's compute is fully done" (allowing the next round's stage to reuse the buffer) and the rendezvous for "this round's stage is fully done" (allowing the next round's compute to read). Drop either duty and you have a data race — which also explains why the "double-buffer B only" variant in §3.1 never even left the gate on correctness grounds. Fence accounting compared:

FormFences per iterationStaging vs compute
r38 single buffer2 (after stage + after compute)serial: all warps stage together, wait together, compute together
r39 double buffer1 + 1 in the prologueoverlapped: stage(kt+1) rides on compute(kt)

Note the gain is not "one fence fewer" per se (that is only tens of cycles per iteration); it is that the interval between fences changes from mutually exclusive to concurrent.

2.4 The vetoed variant: double-buffer B only (KDR=4)

The intuitive plan is to keep KDR=4 and give only the largest plane (qb_exp) a second copy, halving the doubling cost. It fails on correctness alone: with the A side still single-buffered, stage(kt+1)'s A-side bulk copy would overwrite A while compute(kt) is still reading it — A and B staging happen atomically inside the same macro (A first, then B), so pipelining requires doubling everything, and doubling everything at KDR=4 is the 59,392 B cliff. The variant dies on §2.2's "price both planes together" before any performance test.

3. Implementation

3.1 Design choices (why this shape and not another)

  1. Port an existing scheme rather than invent a new one: the generic kernel mmq_nt_kernel<7,2,0> has long been a double-buffered pipeline; the q6_K-specific kernel copies the same buf ^= 1 skeleton, reducing correctness risk.
  2. KDR 4→2 rather than hard-doubling smem: the arithmetic in §2.2 — the same 29,696 B buys overlap, not depth; r38's KDR=8 regression (1097.8) was a data point from 20 minutes earlier, leaving zero room for illusions about 59,392 B.
  3. Pipeline A and B together: the A side is a bulk uint4 LDG→STS (the r34 prepass transport), the B side is per-byte bitfield expansion + dsc reads; both live in the same RAW_STAGE macro and both move into the second buffer.
  4. Hold the register budget at 0 spill: double buffering introduces per-plane stride constants and a buffer selector; ptxas recorded 85 → 87 regs, still 0 spill, and the 2 blocks/SM occupancy premise is untouched.

3.2 Key code

The excerpts below all come from r39 commit f2b9e54's diff to src/cuda_kernels.cu (+77/−29).

smem planes ×2 and per-buffer stride:

// r39: DOUBLE-BUFFERED staging — two copies of every per-kt plane so kt+1's
// global->smem expansion (the ql+qh recomb) overlaps kt's compute, hiding the
// B-staging latency that left r38 latency-bound. Layout per buffer b below.
uint8_t* qa8    = mmq_q6k_sh;                          // [2][KDR*NBI*32]
uint8_t* sda_q  = qa8    + 2 * KDR * MMQ_NBI * 32;     // [2][KDR*NBI*4]
uint8_t* qb_exp = sda_q  + 2 * KDR * MMQ_NBI * 4;      // [2][NBJ*KDR*32]
float2*  sds    = reinterpret_cast<float2*>(qb_exp + 2 * MMQ_NBJ * KDR * 32);
const int qa8_stride    = KDR * MMQ_NBI * 32;
const int sdaq_stride   = KDR * MMQ_NBI * 4;
const int qbexp_stride  = MMQ_NBJ * KDR * 32;
const int sds_stride    = KDR * MMQ_NBJ;

The base addresses of all four planes reserve two copies, and the macro selects the buffer via b * stride — the staging macro's body itself is unchanged word for word (A's uint4 copy, B's expand_q6_elem expansion, the dsc read all as-is); only the destination pointers become the + (size_t)(b) * stride offset versions:

#define RAW_STAGE_Q6K_BT(kt, b)                                                \
    do {                                                                       \
        uint8_t* qa8b   = qa8    + (size_t)(b) * qa8_stride;                   \
        uint8_t* sdaqb  = sda_q  + (size_t)(b) * sdaq_stride;                  \
        uint8_t* qbexpb = qb_exp + (size_t)(b) * qbexp_stride;                 \
        float2*  sdsb   = sds    + (size_t)(b) * sds_stride;                   \
        /* A: bulk LDG->STS of the pre-transposed qa8/sda (no math)      */    \
        ... ((uint4*)(qa8b))[off]   = ((const uint4*)(qa8g + qbase))[off];     \
        ... ((uint4*)(sdaqb))[off]  = ((const uint4*)(sdag + sbase))[off];     \
        /* B: expand KDR*32-chunk super-block ...                              \
        ... qbexpb[...] = (uint8_t)v;  /* expand_q6_elem, same as r38 */       \

Main loop: prologue + interleave. r38's for { if(kt>0) stage(kt); barrier; compute } is rewritten as:

RAW_STAGE_Q6K_BT(0, 0);
__syncthreads();

int buf = 0;
for (int kt = 0; kt < nktile; ++kt, buf ^= 1) {
    // Overlap kt+1's global->smem expansion with kt's compute: stage into the
    // OTHER buffer (buf^1) while reading buffer buf (the mmq_nt<7,2> pipeline).
    if (kt + 1 < nktile) RAW_STAGE_Q6K_BT(kt + 1, buf ^ 1);

    const uint8_t*    qa8c   = qa8    + (size_t)buf * qa8_stride;
    const uint32_t*   sdaqc  = ... sda_q  + (size_t)buf * sdaq_stride;
    const uint8_t*    qbexpc = qb_exp + (size_t)buf * qbexp_stride;
    const float2*     sdsc   = sds    + (size_t)buf * sds_stride;

    for (int kd = 0; kd < KDR; kd++) {          // KDR=2: two 32-chunks per kt
        ...
        const uint8_t* qat = qa8c + (size_t)kd * MMQ_NBI * 32;
        ... ldmatrix A-frag / int8 B-frag / 16× mma.m16n8k16 / two += rescales ...
    }
    __syncthreads();   // loop-tail fence: the two duties of §2.3 (a plain fence at r39 time)
}

Key points: every compute-side pointer becomes the per-buffer *c version (in the diff, the three single-line hunks at 5451/5470/5483 are just reference replacements like sds→sdsc in the epilogue); the loop-tail __syncthreads() stays — it is now the only in-loop fence. The current tree (src/cuda_kernels.cu lines 6905-6929) still carries this skeleton; r53/r56 merely layered cp.async group waits (gemm_cp_wait1) on top, which this doc does not expand on.

Change-surface accounting: src/cuda_kernels.cu +77/−29 lines and src/cuda.rs 2 lines — the genuinely "new" parts are only three: the smem plane table, the macro signature gaining the (kt, b) pair, and the main loop's interleave structure; the staging macro's body (uint4 copy, expand_q6_elem, dsc read) and the compute body are untouched word for word. Minimal diff surface = minimal verification surface: the first explanation for parity/greedy going all-green is that the arithmetic path changed zero bits.

Launcher: KDR 4→2 + smem expression ×2:

constexpr int KDR = 2;
const int smem = 2 * KDR * MMQ_NBI * 32   // qa8  (double-buffered)
               + 2 * KDR * MMQ_NBI * 4    // sda_q (double-buffered)
               + 2 * MMQ_NBJ * KDR * 32   // qb_exp (double-buffered)
               + 2 * KDR * MMQ_NBJ * 8;   // sds  (double-buffered)

The dispatch gate, the (id/32)%8==0 whole-super-block alignment check, and the return-0 clean fallback on failure are all unchanged; the cuda.rs diff is only 2 lines (1+/1−, comment-only).

3.3 Pitfalls

  1. The half-pipeline of "double-buffer only the big plane": KDR=4 + B-only double buffering looks like it saves 12 KB, but in reality compute(kt) reads A while stage(kt+1) overwrites A — a data race, vetoed on correctness alone (§2.4). The minimal complete unit of pipelining is "all inputs of one kt iteration", not "the largest plane".
  2. The smem quote illusion: looking only at the B expanded plane (16 KB) suggests ample headroom; adding back A (8 KB), sda_q (1 KB), and sds (4 KB) reveals KDR=4 double buffer = 59,392 B. Budgets must be priced per "whole-iteration input".
  3. Watch small register creep: the stride constants ×4 plus the buffer selector took ptxas from 85 → 87 regs. 0 spill survived, but r40 would prove how thin this margin was (87 → 80, squeezing hard for 3 blocks/SM, is exactly where a 4 B spill got squeezed out).
  4. Two mechanical traps when restructuring staging: the prologue stage, moved out of the loop, must be immediately followed by __syncthreads() (prevents a first-iteration race); typed-pointer strides scale by element, not byte (sds is float2*, so +1 = 8 B, hence sds_stride = KDR * MMQ_NBJ without the ×8) — both are silent killers of the "structure right, offsets wrong" kind; recount the fence accounting point by point and re-derive each stride.

4. Verification

  • parity 1/0: logits deviation vs the baseline path within the 1e-3 gate — double buffering changes only buffer assignment and fence placement, no arithmetic order, so in theory it should be near-bitwise; the 1/0 record confirms no read/write race pollution from a misplaced fence (defends against losing either of §2.3's two duties).
  • greedy byte-identical: the greedy output stream matches the pre-change one (defends against argmax knife-edge flips).
  • ptxas: 87 regs / 0 spill (defends against register pressure breaking 2 blocks/SM).
  • suite 166/0/3: full regression; notation passed/failed/skipped (defends against other quantization paths being collateral damage — this diff touched the RAW_STAGE macro; q4_K's same-named macro is independent, but the suite is the final "didn't break anyone else" evidence).
  • A/B interleaved 3/3: same-window, same-binary pairing (defends against machine-drift fake deltas).

Why parity should be near-bitwise: double buffering changes no arithmetic — the expansion formula, the mma sequence, the two += of the rescale, the accumulation order are all as they were; what changes is only which of the two buffers data lands in and where the fences sit. As long as the fence accounting is right (§2.3), the output is bit-identical to single buffering. So this round's parity gate is not really defending against arithmetic regressions but against read/write races from misplaced fences — errors of that kind feature sporadic dirty data, and the parity 1/0 and greedy byte-identical gates catch it together.

  • nsys kernel timing: attn_v kernel 2,549,248 → 2,046,848 ns (defends against attribution errors like "the wall clock moved but really some other kernel got slower").

5. Results

Metricbefore → afterNote
whole-prefill (7B, same-window A/B median)1568.7 → 1777.5 tok/s (+13.3%)master row 53
attn_v q6_K kernel (nsys)2,549,248 → 2,046,848 ns (−19.7%)direct evidence the exposed latency was overlapped
kernel compute share16.7% (r38) → 21.5%idle issue slots reduced, but still far from saturated
ptxas85 → 87 regs, 0 spillthe cost of the double-buffer pointers
smem29,696 B (unchanged)KDR 4→2 × double buffer = same footprint
vs llama.cpp (3325-eq anchor)2.13× → 1.87×the engine got faster, so the relative multiple falls

Three readings:

  1. +13.3% ≫ r38's +2.87%: r38 bought occupancy (2 blocks/SM); r39 bought overlap. Occupancy cannot save exposed latency while the structure is fence-serial — more warps, but within each iteration all warps wait on staging together. Overlap hides the staging segment directly inside the compute segment: a structural deletion of time.
  2. The ratio between the attn_v kernel's −19.7% and the wall clock's +13.3% is also self-consistent: attn_v is one of q6_K's two big tensors, and a kernel-level −19.7% amortized over whole-prefill with a discount lands at the +13.3% order of magnitude.
  3. compute 21.5% is still a hook: once the overlap cashed in, the next bottleneck shows itself — latency's source shifts from "serial staging" to "per-byte B-expansion loads" (32 per-byte LDG.U8), which r41's uint4 widening will take up; and occupancy still sits at 2 blocks/SM, where r40's third resident block will also take a share. Both lines (r39→r41 latency, r39→r40 occupancy) start from this step's profile data.

Baseline drift note: r38 landed reporting 1561.9; this doc's in-window baseline is 1568.7 — absolute values from different session windows of the same code are not comparable (§0 table-reading convention), so deltas are always taken from same-window A/B pairs (+13.3%), never subtracted across windows.

q6_K line postscript: after this doc the occupancy and load-width lines run in parallel — r40 third resident block (+13.0%), r41 B-expand uint4 widen (+30.7%, kernel 1.70 → 0.654 ms); after r41 the q6_K line sat at only 1.27× vs llama, finally closing in the r45–r53 cp.async bundle. This doc's smem 29,696 B and fence skeleton lived on past r53 (the current tree's lines 6905-6929 are still the RAW_STAGE(0,0) + buf ^= 1 skeleton).

6. Lessons

  1. Once occupancy is bought, the next lever is pipelining, not a deeper tile: +2.87% (occupancy) was followed directly by +13.3% (overlap) — same kernel, same tile, same smem.
  2. The gain is "pure overlap": workloads unchanged, arithmetic unchanged; only the arrangement of buffers and fences changes. This is also why its parity risk is extremely low (no numeric order changes).
  3. Price the smem budget per "whole-iteration input", A and B planes together; a half-pipeline (doubling only the big plane) is disqualified on correctness alone.
  4. Porting an existing scheme beats inventing one: the mmq_nt<7,2> pipeline skeleton was reused as-is, concentrating correctness risk on a single fence accounting.

← 41 · r38 q6_K BT-style raw-byte mma · Index · 43 · r40 third resident block →

43 · r40 — __launch_bounds__(256,3) third resident block (LANDED)

Result: 7B whole-prefill 1784.0 → 2015.6 tok/s (+13.0%, 3/3 interleaved pairs + one independent pair); attn_v q6_K kernel 2.05 → 1.58 ms (−23%). Commit: 65ecef7. Date: 2026-09-05.

1. Background — where things stood

r39's double buffering freed the q6_K kernel from "serial staging" (+13.3%), but the kernel was still latency-bound: compute share 21.5%, No-Eligible (the share of cycles with no eligible warp at the issue slots) 74.5%, active warps/sched 3.87. The SM was waiting on latency, and with only 2 resident blocks on the SM = 16 schedulable warps, the supply of latency cover was capped by occupancy.

Of occupancy's three constraints, the q6_K kernel's state after r39 landed was:

  • smem: 29,696 B/block × 3 = 89,088 B < 102,400 B (GB10's per-SM opt-in ceiling) — smem had allowed a 3rd block all along;
  • registers: r39 recorded 87 regs/0 spill. 87 × 768 threads (3 blocks × 256) = 66,816 > 65,536 (the per-SM register file) — registers pinned residency at 2;
  • block size/other: 256 threads × 3 = 768 ≤ the per-SM thread ceiling; no obstacle.

Conclusion: a pure register-budget problem. Squeeze per-thread registers from 87 to within 80 and the 3rd block fits. The audit of the three constraints:

Constraint2 blocks/SM (status quo)3 blocks/SM (target)Verdict
smem29,696 × 2 = 59,392 B29,696 × 3 = 89,088 < 102,400allowed long ago
Registers87 × 512 = 44,544 ≤ 65,53687 × 768 = 66,816 > 65,536the sole veto
Threads512 ≤ per-SM ceiling768 ≤ per-SM ceilingallowed

The occupancy ladder is one this campaign has climbed repeatedly: r28 used smem (45,056 B) to buy q4_K 2 blocks/SM (+2.56%); r39 cashed in q6_K's latency hiding at 2 blocks/SM; r40 is registers' turn — the same scale, weighed a third time.

The question "can a 3rd block be reached by trimming regs across the 85.33 line" had in fact been on record since r38 landed (the tail of r38's commit message carried the todo "whether a 3rd block is reachable by trimming regs < 85"); r39's record promoted it to the next lever (r40's commit message verbatim: "Adds the r39-named lever"); after r39 every condition was ripe, and r40 cashed it in with one hint line. Read as three points in time, this is the documentation system working normally: phenomenon (85 regs blocks residency) → on record (r38 todo) → promoted to lever (r39) → cashed in (r40) — skipped steps usually happen when the phenomenon never got written down.

The only obstacle was psychological: squeezing to 80 regs means spilling (registers overflowing to local memory), and "0 spill" had been the admission convention for new kernels throughout the campaign (r28 123 regs/0 spill, r38 85 regs/0 spill, r39 87 regs/0 spill — every doc's ptxas audit treated 0 spill as a green metric). r40's real work was not writing that line of code; it was proving experimentally that 4 B of spill is immaterial, demoting the convention from "rule" to "heuristic".

The contrast in cost structure is also worth recording: the implementation is a one-time single line; the forensics is ten variant builds plus a round of ncu. For this class of "one-line lever" in the campaign, the doc is usually a hundred times the size of the code — because all the transferable knowledge is in the evidence chain, not in that line of code.

2. Principle — the GPU mechanism

2.1 The register arithmetic: where 80 comes from

The second parameter of __launch_bounds__(maxThreadsPerBlock, minBlocksPerMultiprocessor) is a contract handed to ptxas: guarantee at least 3 blocks of 256 threads can be resident simultaneously. ptxas back-derives the per-thread register ceiling from it:

per-SM register file        = 65,536 32-bit registers
3 blocks × 256 threads      = 768 threads
65,536 / 768                = 85.33 → rounded down to the 8-registers-per-thread allocation granularity → ceiling 80

87 regs × 768 = 66,816 exceeds the register file, so ptxas's natural allocation only allows 2 blocks; capped at 80, the 1 extra live value (ptxas actually needs 81) has to spill. 4 B spill = 1 32-bit value going to local memory (thread-local, backed by L1/L2), adding one store/load pair per pass through the staging path.

Both directions must be checked for the occupancy verdict to stand: at 2 blocks, 87 × 512 = 44,544 is far within 65,536 — 87 regs is fine in itself; the problem only appears in the 768-thread multiplication. The allocation granularity of 8 comes from an implementation detail: ptxas allocates registers in whole warp (32-thread) segments in units of 8/thread; the division result 85.33 can never be granted, and the nearest grantable step is 80.

smem side-check: 29,696 × 3 = 89,088 ≤ 102,400; smem does not block. So the launch occupancy limit moves from (regs=2, smem=3) to (regs=3, smem=3) — the first time both limiters read 3 together.

The raw evidence form of ptxas -Xptxas -v: compile with -Xptxas -v and ptxas prints one resource line per entry. The corresponding lines from the two builds (output format verbatim, numbers from the record):

ptxas info : Compiling entry function '...mmq_raw_nb_bt_q6k_kernel<2>'
ptxas info : Used 87 registers, 0 bytes spill stores, 0 bytes spill loads      // r39 build
ptxas info : Used 80 registers, 4 bytes spill stores, 4 bytes spill loads      // r40 build (after the hint)

"4 bytes spill stores/loads" is exactly that 1 32-bit live value's store/load pair, 4 B each — matching §2.1's "one extra store/load pair" word for word. That line is the entire compile-time consequence of r40's entire code change.

2.2 Why 4 B of spill is immaterial here

The default judgment "spill is a loss" comes from compute-bound kernel intuition: spill's local accesses insert into the hot loop and steal issue slots. But this kernel's measured state is No-Eligible 74.5% (ncu issue statistics: the share of issue-eligible cycles in which no warp is eligible, i.e. the idle rate of an SM that "wants to issue but has nothing to issue") — for three quarters of the cycles the SM wants to issue and has no warp to issue; that is insufficient latency cover, not insufficient issue bandwidth. The two sides of the scale:

  • Gain side: resident warps 16 → 24 (theoretical +50%). The total schedulable latency tolerance grows linearly with resident warps, and r39 had just left a large amount of LDG/expansion latency sitting in the pipeline, which 8 more warps are exactly positioned to absorb.
  • Cost side: 1 spill slot = one pair of 4 B local accesses per iteration on the staging path, L1-hit territory, and thoroughly buried under the parallelism of 24 warps.

The scale's verdict rests on measurement, not reasoning: +13.0% wall, kernel −23%. The "0 spill gate" is empirically falsified here — it defends against "a spill avalanche caused by runaway register pressure", not "any spill is guilty".

To make the gain-side mechanism explicit: a latency-bound kernel eats occupancy because memory requests in flight = resident warps × pending loads per warp. r39's double buffering made each warp pend more loads (the expanded-B LDG stream), but only 16 warps on the SM were available to cushion that latency; the 3rd block raises the latency-cushioning warps to 24, which is what moves No-Eligible from 74.5% to 70.1%. The cost-side spill goes to local memory: on first spill ptxas allocates a fixed per-thread local slot at launch, after which it is just an STL/LDL pair for that 1 live value — 4 B, L1-hit territory, and located on the staging path (already the latency-dominated segment, where one more access pair is fully absorbed by the parallelism of 24 warps).

Reading the issue-slot economics: No-Eligible 70% means that of every 10 cycles, only about 3 per SM have a warp eligible to issue. The two roads to more throughput are (a) make each issue do more (ILP/wider loads — r41's route) or (b) reduce the no-warp-eligible cycles (more resident warps — this doc's route). The two roads have different gain ceilings: this doc lowered No-Eligible by only 4.4 percentage points yet took +13.0% — because the added issues concentrate in the stall segments of the critical path; that disproportion is a typical reading for a latency-dominated kernel.

2.3 The relay with r39/r41

r39 raised compute share to 21.5% but No-Eligible stayed 74.5%; r40 lowers No-Eligible to 70.1% and raises compute to 29.06% — once the occupancy line's gain cashed in, the residual bottleneck is the L1TEX scoreboard (measured at the r40 point: 10.5 cy/warp, ~70% of warp time; r41 attributes it precisely to 32 per-byte LDG.E.U8 and presses it further to 3.6 cy). The occupancy and load-width lines each own a segment and do not substitute for each other.

2.4 Why stop at 3: the 4th block's ceiling

r40 also ruled out the next step while it was there. 4 blocks/SM requires: smem 29,696 × 4 = 118,784 B > 102,400 B — smem vetoes outright; even with enough smem, the register ceiling would drop to 65,536/1024 = 64/thread (64 at the granularity of 8), a further 16 down from 80, and the spill surface would roll from 1 slot into a sheet. So 3 is this kernel's natural residency endpoint on GB10 — r40 got there in one step, with no "squeeze one more block" tail.

3. Implementation

3.1 Design choices (why this shape and not another)

  1. Hint, not hand-trimming: let ptxas decide at the 80-cap which value to give up (it has whole-kernel live-range data); the human sets the contract and does not meddle in the allocation. Measured: ptxas's choice was 1 4 B spill.
  2. Sweep hand-trimmed variants before accepting spill: to confirm 4 B was not the lazy option, a Task-2 10-variant register-trimming sweep was run (listed verbatim in the commit message: G-recomputed, smem-base fold, de-unroll, assume, div→shift, per-g A-frag, epilogue recompute, fused mma+scale, etc.) — all landed at 80 regs/4 B spill or worse. Conclusion: this ptxas needs 81 live registers, and the 4 B spill is a hard floor, not a matter of effort.
  3. Don't touch KDR/double buffering/dispatch gate: r40 is fully orthogonal to r38/r39's mechanisms — a one-line change, output-neutral.

3.2 Key code

The entire code change is one line in src/cuda_kernels.cu (the complete code diff of 65ecef7):

 template <int KDR>
-__global__ void __launch_bounds__(256) mmq_raw_nb_bt_q6k_kernel(
+__global__ void __launch_bounds__(256, 3) mmq_raw_nb_bt_q6k_kernel(
     const uint8_t* __restrict__ W, const uint8_t* __restrict__ qa8g,
     const uint8_t* __restrict__ sdag, float* __restrict__ C,
     int nt, int od, int id, int nchunk, int bstride

__launch_bounds__'s first parameter (256) locks the thread count; the second (3) is the minimum-resident-blocks contract. The current tree (cuda_kernels.cu:6701) keeps this line as-is, and r53 additionally gave the template an EXP parameter (<KDR, bool>); the hint never changed again.

ptxas -Xptxas -v forensics is this doc's "implementation" centerpiece. Compile with -Xptxas -v, read each kernel's register/spill report line; the key numbers of the two builds:

Buildregs/threadspillResidency ceiling (regs side)
r39 (no hint)8702 blocks (87×768 > 65,536)
r40 (hint (256,3))801 slot = 4 B3 blocks (80×768 = 61,440 ≤ 65,536)

The same-table audit of the 10 hand-trimmed variants all came out ≥4 B spill — the evidence chain for the 4 B floor is complete (numbers from the r40 commit message and master row 54). The variants named in the commit message and each one's register-saving hypothesis:

VariantHypothesisResult
G-recomputedrecompute the 4 ldmatrix offsets each time, saving G[4]80/4B
smem-base foldfold the smem base-address arithmetic80/4B
de-unrollun-roll the kd #pragma unroll to shorten live ranges80/4B (or worse)
assumealignment/aliasing assumptions to help ptxas narrow80/4B
div→shiftreplace division with shifts to save temporaries80/4B
per-g A-fragshrink the A-fragment register footprint80/4B
epilogue recomputerecompute epilogue addresses80/4B
fused mma+scalefuse mma and the rescale path80/4B

(The sweep was 10 variants in total; the table lists the 8 named in the commit message.) Conclusion: at the 80-cap, this ptxas needs 81 live registers no matter how things are arranged — the 4 B spill is a hard floor, not a matter of effort. accept-the-spill went from compromise to evidence-backed decision.

The hint's survival through later evolution is also on record: after the current tree's r53 added the EXP boolean to the template, instantiations remained mmq_raw_nb_bt_q6k_kernel<KDR, true/false> and the prewarm lines are mmq_raw_nb_bt_q6k_kernel<2, true> / <2, false> (cuda_kernels.cu lines 7140-7141); __launch_bounds__(256, 3) was never reverted. A forward rule worth setting down: launch_bounds is a per-instantiation contract — every time the template gains an instantiation (<KDR, bool> is one → two), the -Xptxas -v audit must be rerun for that instance; the hint only constrains ptxas's allocation target and does not guarantee a new instance can still reach 80 regs at 0/small spill. When r53 landed, this was exactly one of the items needing reconfirmation beyond "r41's uint4 widen held 80/4B".

The ncu Occupancy section's residency evidence (two launch-limit lines from r40's verification record; §5 has the achieved values):

Block Limit Registers                3      ← 80 regs × 768 threads = 61,440 ≤ 65,536
Block Limit Shared Mem               3      ← 29,696 × 3 = 89,088 ≤ 102,400
Block Limit Warps                    6      ← 48-warp ceiling / 8 warps per block

The minimum of the three Block Limits sets the residency ceiling: before, min(2, 3, 6) = 2; after, min(3, 3, 6) = 3 — registers went from "single veto" to a joint decider tied with smem, which is precisely the hint's semantic goal.

3.3 Pitfalls

  1. The inertia of the 0-spill convention: three consecutive steps (r28/r38/r39) recorded 0 spill as green in their audit tables, nearly mistaking a heuristic for an admission rule. Half of r40's contribution is the number; the other half is writing down this convention's scope of validity.
  2. The register granularity trap: 65,536/768 = 85.33 — plan for 85 and ptxas still cannot grant it; the allocation granularity is 8, so the effective ceiling is 80. When planning occupancy, round down to the granularity, not to the division result.
  3. smem allows ≠ can be resident: when r39 landed, the smem side already fit 3 blocks (89,088 < 102,400), but the launch occupancy limit's regs=2 cast the single veto. Check occupancy bottlenecks item by item; never infer from "smem unchanged".

4. Verification

  • parity 1/0: logits deviation within the gate (the campaign's parity gate: run the same prompt against the baseline binary and compare max |Δlogits| ≤ 1e-3) — a register hint touches no arithmetic path (defends against numerical regressions introduced by collateral changes like "touched code while squeezing registers").
  • greedy-32 byte-identical: the commit message states "register hint is output-neutral" — the same arithmetic with only the resource allocation changed (defends against argmax knife-edge flips).
  • ptxas resource audit: 80 regs/1×4 B spill, and the launch occupancy limit reads regs=3, smem=3 — direct evidence the hint took effect (defends against "hint written but ptxas ignored it").
  • ncu occupancy recheck: 3 blocks/SM confirmed — achieved 18.12 warps/SM = 37.74% (against the 48-warp ceiling; theoretical 24 = 50%), active warps/sched 3.87 → 4.48 (defends against "3 blocks in theory but never 3 in practice"). The gap between achieved 18.12 and theoretical 24 is the normal loss of scheduling tails and divergence (wave tails, uneven block progress), not evidence the 3rd block failed to be resident — the launch limit and active warps counters corroborate each other: the ceiling really reached 3, and the average activation count rose with it from the 16 magnitude to 18+ (under the old shape the average cannot pass the theoretical ceiling of 16).
  • A/B interleaved 3/3 + one independent pair: same-window, same-binary pairs, all positive (defends against machine-drift fake deltas).
  • suite 166/0/3: full regression (defends against collateral damage to others).

5. Results

Metricbefore → afterNote
whole-prefill (7B, same-window A/B median)1784.0 → 2015.6 tok/s (+13.0%, 3/3 + one independent pair)master row 54
attn_v q6_K kernel2.05 → 1.58 ms (−23%)the half-more of residency cashed in
compute share21.52% → 29.06%issue slots keep backfilling
No-Eligible74.5% → 70.1%latency cover improved but still the main bottleneck
active warps/sched3.87 → 4.48ncu scheduler view
occupancy (achieved)18.12 warps/SM = 37.74% (3 blocks confirmed)theoretical 24 warps = 50%
matched-nt q6_K~221.8 → ~171 µs/GMAC (est.)still ~3.0× vs llama's 57.8
vs llama.cpp (3325-eq anchor)1.87× → 1.65×the engine got faster, so the relative multiple falls

Three readings:

  1. One line for 13%: the change is one constant in a template parameter position; the gain comes from §2.2's scale — +50% theoretical resident warps against 1 4 B spill slot. Occupancy-class levers can have a very high price/performance ratio on latency-dominated kernels.
  2. Three consecutive steps compounding: r38 (new kernel, +2.87%) → r39 (overlap, +13.3%) → r40 (residency, +13.0%), the q6_K line pressed from 368.9 µs/GMAC down to ~171 and whole-prefill pushed from the 1518 band past 2000. Each step's "Results" section hands the next step its state (the 16.7% → 21.5% compute curve, No-Eligible 74.5%). The kernel −23% vs wall +13.0% ratio is also self-consistent: in r37's attribution the q6_K GEMM was about half the wall clock, so cutting the wall-critical segment 23% amortizes to +11–13% over the whole wall — the two layers of numbers cross-check each other.
  3. The residual is named: No-Eligible still 70.1%, L1TEX scoreboard 10.5 cy/warp ≈ 70% of warp time — the next lever is load width (r41's uint4 B-expand cuts that number to 3.6 cy and whole-prefill gains another +30.7%), not residency (3 blocks reached; no cheap smem/register headroom remains).
  4. Unit-cost trajectory: matched-nt q6_K ~171 µs/GMAC, still ~3.0× vs llama's 57.8 (r40 commit message verbatim), but the internal decomposition has changed — the commit message also records "the residual is KSPLIT=2's intrinsic 2x mma.k16 plus memory-latency stall, not residency": the residency line is exhausted, and half the remaining gap is the doubled mma structural cost q6_K's 16-sub-block layout imposes, half is load latency. r41 attacks the latter.

Baseline drift note: r39 landed reporting 1777.5; this doc's in-window baseline is 1784.0 — absolute values from different session windows of the same code are not comparable (§0 table-reading convention); deltas are always taken from same-window A/B (+13.0%, 3/3 plus one independent pair), never subtracted across windows.

q6_K line postscript: this doc is the third step of the four-step arc (r38 skeleton → r39 overlap → r40 residency → r41 load width); after r41 the q6_K line had 1.27× left vs llama, and the line closed in the r45–r53 cp.async bundle. This doc's __launch_bounds__(256, 3) is the only kernel parameter in the four steps never touched again — once the residency contract is set right, all later evolution (template parameters, cp.async, pre-expanded planes) happens inside its resource frame.

6. Lessons

  1. 0 spill is a heuristic, not a law: measure spill's actual cost first, then decide how much register pressure to pay to avoid it; on a latency-bound kernel, a 4 B spill slot is no match for +50% resident warps.
  2. Check occupancy's three constraints item by item: smem, registers (capped after rounding up to the 8 granularity), and thread count each veto independently; "smem fits" does not imply "the block gets in".
  3. Let ptxas be the allocator and the human the auditor: hint sets the contract, -Xptxas -v produces the evidence, the variant sweep sets the floor — the correct division of labor for register trimming.
  4. Before buying residency, confirm the kernel really is latency-dominated: the No-Eligible/active-warps counters are the grounds for buying occupancy; on a kernel whose compute is already saturated, buying residency only dilutes each warp's smem/L1 quota.

← 42 · r39 q6_K KDR=2 double-buffer · Index · 44 · r41 q6_K B-expand uint4 widen →

44 · r41 — q6_K B-expand widened to uint4 group loads (LANDED)

Result: whole-prefill 1979.9 → 2605.2 tok/s (+30.7%, 7B q4_k_m @pp3314-eq, same-window A/B median); q6_K attn_v GEMM kernel 1.70 → 0.654 ms (−61.5%); the long_scoreboard (L1TEX) share of warp time 85.5% → 33.6%. The q6_K line's biggest single-step lever to that point. Commit: b891e1b (code, src/cuda_kernels.cu +58/−10) + aa82e8f (record). Date: 2026-09-05.

1. Background — where things stood

r37's whole-prefill attribution swapped the target for the whole campaign: the MMQ redesign had brought q4_K-class weights to near llama.cpp, but q6_K was still on a path slower than f16 — the q6_K GEMM alone ate 1094.7 ms, 51.2% of the entire prefill wall clock, at 6.38×/GMAC efficiency. That is: the next doubling of the default path lay not in q4_K's tile shape but in "building q6_K a real kernel".

So P6 landed three steps in a row, all on mmq_raw_nb_bt_q6k_kernel:

  • r38 (75aabb9): the BT-style raw-byte kernel. q6_K is not 8 32-element sub-blocks but 16 16-element sub-blocks, and one 32-k chunk spans two sub-blocks with different scales, so it uses mma.m16n8k16 (KSPLIT=2) + an independent dsc rescale per half; the B tile is expanded into centered int8 (−32..31) at staging, moving the recomb (nibble + 2-bit field assembly + −32) out of the hot loop. +2.87%.
  • r39 (f2b9e54): KDR=2 double buffering — two copies of each A/B staging plane, kt+1's global→smem expansion overlapping kt's compute. +13.3%.
  • r40 (65ecef7): __launch_bounds__(256, 3) forcing a third resident block. The 0-spill dogma was falsified (80 regs + 4 B spill beats 87 regs at 2 blocks), +13.0%, reaching 2015.6 tok/s.

At r41's start the occupancy lever was spent (3 blocks/SM, 80 regs), yet the kernel was still latency-bound: compute at only 27%, No-Eligible as high as 74.5%. First on r40's candidate-lever list was the B-expand's fetch style — and ncu's Warp State data pointed the finger exactly there: 85.5% of warp-stall cycles were long_scoreboard (13.7 of the 16.0 cy/inst CPIStall), the characteristic signature of L1TEX memory latency.

Where this step's absence would stall: three resident blocks had already pushed "use other warps to cover this warp's wait" to its limit, but the bulk of long_scoreboard is a within-warp serial dependency chain — a load issued by this warp is consumed by this warp's next instruction, and the latency has nowhere to hide. More scheduling tuning, more tile tuning, would all skirt the real bottleneck.

The CPIStall accounting. Of ncu's Warp State average 16.0 cy per inst of issue spacing, 13.7 cy are booked to long_scoreboard — i.e. 85.5% of each warp's issue time is spent "waiting for an L1TEX round trip". After r40 the measurement was 18.12 warps/SM (37.74% occupancy): resident warps were not scarce, but 74.5% No-Eligible says they lacked issuable instructions — every recomb hangs on its own load's result. Occupancy can only hide cross-warp latency; a chain of 32-level byte loads feeding one ALU stays exposed level by level inside the warp.

2. Principle — the GPU mechanism

What B-expand is. A q6_K 256-element super-block is packed into 210 B (224 B after padded registration). The in-block byte map:

offset   0 …… 127        128 …… 191       192 …… 207      208-209   210-223(padded)
content  ql: low nibbles  qh: per-element  sc[0..15]:      d: f16    14 B padding
         (256 4-bit       high 2-bit       the 16 sub-     scale     (for alignment)
         nibbles, 2 elems fields (64 B,    blocks' i8
         per byte)        4 elems per byte) scales)

An element's value = the 4-bit nibble in ql (low) | the corresponding 2-bit field in qh << 4, then −32 centered overall. The trouble is byte sharing: one ql byte serves two elements (nibble shift 0/4), one qh byte serves four elements (2-bit field shift 0/2/4/6) — the same byte is sliced at 2-bit steps across elements. The BT kernel's smem B plane stores the expanded centered int8 (1 B per element, range −32..31) and mma eats int8 directly; so the whole recomb happens at staging:

v = ((ql[i] >> shift) & 0x0F) | (((qh[j] >> shift2) & 0x03) << 4)   // then −32

The bottleneck's arithmetic. At KDR=2 one kt covers 64 elements per row; the B tile is MMQ_NBJ=128 rows × 64 elements, and the kernel dispatches in 16-element groups: ng = 128 × (2×32)/16 = 1024 groups, spread over 256 threads — exactly 4 groups per thread per kt. Before widening, each group issues 32 LDG.E.U8 (16 ql + 16 qh), and the SASS shows staging sits at the top of the kt loop with load results consumed immediately by recomb→STS — the full L1TEX round-trip latency of every byte load is exposed on the critical path. That is the 13.7 cy/inst long_scoreboard: not a bandwidth shortage, but strings of serially-awaited short loads.

Why uint4 cures it. The padded 224 B row pitch = 14×16, so any block base blk = W + j·(nsb·224) + sb·224 is 16-aligned; the offsets of the group's two runs can be enumerated straight from the code — the ql run blk + it0·64 + gg·16 (it0∈{0,1}, gg∈{0..3}) lands on {0,16,32,48,64,80,96,112}, the qh run blk + 128 + it0·32 + (gg&1)·16 lands on {128,144,160,176} — all multiples of 16, with the qh run strictly inside the qh region (128..191). When alignment holds, 16 contiguous bytes are one LDG.E.128 (uint4):

Loads per groupPer thread per ktWhole block per kt
Before widening32× LDG.E.U84 groups × 32 = 12832,768
After widening2× LDG.E.1284 groups × 2 = 82,048

A roughly 16× cut in load instructions. Scoreboard events fall roughly proportionally with instruction count, while the recomb shift/mask ALU now operates on 32-bit wide words in registers (4 bytes per word, four words per uint4; _Pragma("unroll") lets ptxas expand the byte selection into a static BFI/SHF sequence) — ALU volume is nearly unchanged.

Register neutrality is the key to composability. The widening introduces only 4 uint4 (16 registers) of transient occupancy, and the ptxas landing point stays 80 regs / 4 B spill — not one register of the 3 blocks/SM budget that r40 bought with +50% resident warps is touched. A pure load-width change can stack with the occupancy lever only if it demands no repayment of the register budget; that is also why it is cheaper than "deeper software pipeline"-class schemes (the latter already paid smem doubling once in r39).

Which structure the change lands in. r39's double buffering keeps two copies of each of a kt's four staging planes; the smem budget can be computed from the layout at the kernel head (KDR=2, MMQ_NBI=64, MMQ_NBJ=128):

per buffer copy:
  qa8    [KDR·NBI·32]      = 2·64·32  = 4096 B   (A: q8 plane, swizzled)
  sda_q  [KDR·NBI·4]       = 2·64·4   =  512 B   (A: packed d|ssum)
  qb_exp [NBJ·KDR·32]      = 128·2·32 = 8192 B   (B: expanded centered int8 ← r41's target)
  sds    [KDR·NBJ]·float2  = 2·128·8  = 2048 B   (the dsc plane)
  subtotal                          14,848 B × 2 copies = 29,696 B

These 29,696 B are exactly the source of r39's record "KDR=2 hits the same 29,696 B" — r41's widening adds not one byte of smem; the qb_exp plane keeps its size, only the instructions filling it get fewer. In-block division of labor: 256 threads = 8 warps, each warp owns 16 consecutive od rows (j0w = warp*16), and B-expand's ng dispatch is split linearly across threads, so the same group's ql/qh uint4s always land on the same thread — the recomb therefore completes entirely in registers, with no cross-thread exchange.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Widen inside the kernel rather than going back to load-time pre-expansion. r18 once moved B pre-expansion to load time and was vetoed for its +5.8 GB cost; r41 does not repeat that — widening changes only fetch width; the plane's byte size is untouched. (The load-time pre-expansion direction was later re-realized on q6_K by r53 in the form of the W_exp plane — a different trade: +1.52 GB to make staging a pure copy.)
  • (bstride & 15) == 0 as the runtime gate. Real models always register with the padded 224 B layout (16-aligned, gate open); the raw 210-B layout used by tests (210 = 13×16+2, unaligned) keeps the scalar path. One gate guarantees two things at once: the alignment safety of uint4 accesses and bit equivalence (the two paths agree element by element).
  • A closed form per group, not per-element expand_q6_elem calls. Map (cbase, gg) to it0/qsh/qh_shift, and derive all 16 outputs from the 2 uint4 wide words via register shifts. The closed form was verified against the per-element version before landing: 512,000 elements, 0 mismatches.

3.2 Key code

Where the change sits in the main loop. The r41-era kt main loop is r39's double-buffer shape: at the top of each iteration, stage the next tile first (filling buf^1), then run this tile's mma compute on buf —

// Loop skeleton of the r41 era (the gemm_cp_wait* lines still on the tree are
// r45/r53/r56's cp.async additions; at r41 time this spot had only a plain __syncthreads)
RAW_STAGE_Q6K_BT(0, 0);
__syncthreads();
for (int kt = 0; kt < nktile; ++kt, buf ^= 1) {
    if (kt + 1 < nktile) RAW_STAGE_Q6K_BT(kt + 1, buf ^ 1);  // top: stage the next tile
    __syncthreads();
    /* … KDR rounds of ldmatrix + mma.m16n8k16 + dsc rescale on buf … */
}

The SASS confirms the staging macro sits at the top of the iteration, ahead of compute issue — exactly r41's gain shape: the string of short loads queues once at the loop head, the mma compute then starts, whereas before widening the same stretch was 128 byte loads queued one by one and consumed one by one by recomb.

Before — the scalar per-byte path (the shape all q6_K took before r41; today it survives on the tree as the raw 210-B fallback). expand_q6_elem addresses each element independently, one byte at a time:

// src/cuda_kernels.cu — scalar expansion (2 byte reads per element)
__device__ __forceinline__ int expand_q6_elem(const uint8_t* ql, const uint8_t* qh, int elem) {
    int m  = elem & 31;
    int it = elem >> 7;
    int n  = elem & 127;
    int ql_idx   = it * 64 + (n & 63);
    int ql_shift = (n >> 6) * 4;          // 0 or 4 (low/high nibble)
    int qh_idx   = it * 32 + m;
    int qh_shift = ((n >> 5) & 3) * 2;    // 0,2,4,6 (2-bit fields)
    int v = ((ql[ql_idx] >> ql_shift) & 0x0F)
          | (((qh[qh_idx] >> qh_shift) & 0x03) << 4);
    return v - 32;
}

The per-element staging branch that calls it (2 byte LDGs per element — the very source of the 85.5% long_scoreboard):

// before: 16 elements per group = 32 LDG.E.U8, results consumed immediately by recomb
for (int x = threadIdx.x; x < MMQ_NBJ * (KDR * 32); x += blockDim.x) {
    const int jj = x / (KDR * 32), bec = x % (KDR * 32);
    const int j = j0 + jj;
    const int elem = cbase * 32 + bec;   /* element index within the super-block */
    int v = 0;
    if (j < od && sb < nsb) {
        const uint8_t* blk = W + (size_t)j * ((size_t)nsb * bstride)
            + (size_t)sb * bstride;
        v = expand_q6_elem(blk, blk + 128, elem);
    }
    qbexpb[(size_t)jj * (KDR * 32) + bec] = (uint8_t)v;
}

After — r41's uint4 group expansion (the tree's (bstride & 15) == 0 branch, i.e. the live path for real models):

// r41: 16-elem group expand via uint4 ql+qh global loads.
// padded 224B stride is 16-aligned → one uint4 each for the ql and qh runs, replacing 32 byte LDGs
const int it0 = (cbase >> 2) & 1;
const int qsh = ((cbase >> 1) & 1) * 4;
const int ng = MMQ_NBJ * (KDR * 32) / 16;
for (int g = threadIdx.x; g < ng; g += blockDim.x) {
    const int jj = g / 4, gg = g & 3;
    const int j = j0 + jj;
    uint8_t* out = qbexpb + (size_t)jj * (KDR * 32) + gg * 16;
    if (j < od && sb < nsb) {
        const uint8_t* blk = W + (size_t)j * ((size_t)nsb * bstride)
            + (size_t)sb * bstride;
        const uint4 qlv = *(const uint4*)(blk + it0*64 + gg*16);      // ql: one LDG.E.128
        const uint4 qhv = *(const uint4*)(blk + 128 + it0*32          // qh: one LDG.E.128
                          + (gg & 1) * 16);
        const int qhs = qsh + ((gg >> 1) & 1) * 2;
        const uint32_t qx = qlv.x, qy = qlv.y, qz = qlv.z, qw = qlv.w;
        const uint32_t hx = qhv.x, hy = qhv.y, hz = qhv.z, hw = qhv.w;
        _Pragma("unroll")
        for (int e = 0; e < 16; e++) {                                 // recomb entirely in registers
            const int sidx = e >> 2;                                   // which word holds byte e
            const int sh = (e & 3) * 8;
            const uint32_t qsel = (sidx == 0) ? qx : (sidx == 1) ? qy
                              : (sidx == 2) ? qz : qw;
            const uint32_t hsel = (sidx == 0) ? hx : (sidx == 1) ? hy
                              : (sidx == 2) ? hz : hw;
            const uint8_t qb_ = (uint8_t)((qsel >> sh) & 0xFF);
            const uint8_t hb_ = (uint8_t)((hsel >> sh) & 0xFF);
            out[e] = (uint8_t)((((qb_ >> qsh) & 0xF)                   // low nibble
                | (((hb_ >> qhs) & 3) << 4)) - 32);                    // high 2 bits, −32 centered
        }
    } else {
        _Pragma("unroll")
        for (int e = 0; e < 16; e++) out[e] = 0;                       // zero-fill out-of-range rows
    }
}

Side-by-side: the load side goes from "2 byte LDGs per element" to "2 128-bit LDGs per group"; the addressing side absorbs the scalar version's it/ql_idx/qh_idx divisions into the group coordinates gg = g & 3, jj = g / 4 closed forms; the consumption side goes from "one shift chain per byte" to "expansion shifts over four bytes per 32-bit word". Total dispatch volume is unchanged (still MMQ_NBJ × KDR×32 elements); what changes is the width and count of each memory access.

Where bstride comes from — the launch side decides by registration layout, and the same kernel serves both layouts:

#![allow(unused)]
fn main() {
// src/cuda.rs — prefill MMQ launch side: block_stride chosen by Q6_K registration layout
let block_stride: i32 = if ttype == TensorType::Q6_K && padded_q6k {
    224                                    // 7e② padded repack → the r41 uint4 gate opens
} else {
    210                                    // raw GGUF → the scalar fallback arm
};
}

Inside the kernel, (bstride & 15) == 0 says at a glance which arm to take: real models (padded) get the uint4 path, and the 210 B raw test layout falls back to scalar automatically — one SASS serves both layouts, no second kernel needed.

3.3 Pitfalls

  • The SASS placement trap (foreshadowing r43). Double-buffered staging sits at the top of the kt loop, and after widening all uint4s still issue ahead of compute. The first recomb ALU (LOP3/SHF) still waits out the whole load round trip — what r41 cut was the instruction count, not the length of the latency chain. r43's PC-sampling then quantified this residual: the recomb consumer side still held 45% of the remaining stall.
  • uint4 alignment is not a soft constraint. *(const uint4*)(blk + ...) requires 16 B alignment in hardware; a misaligned address faults outright. Alignment is guaranteed by the padded stride arithmetic (224 ≡ 0 mod 16, all in-group offsets multiples of 16) — but only for the padded layout. That is why the gate must exist, not as a "testing convenience".
  • The nibble/2-bit bookkeeping is easy to get wrong. The ql nibble choice (low/high 4 bits) and the qh 2-bit field (shifts 0/2/4/6) vary with the (cbase, gg) combination; one wrong >> in the closed form is a silent bit flip. The closed form was first compared element-by-element against the scalar version on the host for 512,000 elements, and only entered the kernel at 0 mismatches.
  • Keep and test both arms of the gate. Widening holds only for the padded layout, but the raw 210-B layout is the one parity testing uses — if the raw path were made to error out for convenience, the parity gate could only ever run on padded. Coexisting arms let the parity gate verify once each under the two fetch shapes — raw (expand_q6_elem per-element expansion) and padded (uint4 closed-form expansion) — against the same reference output: layout and fetch shape get tested as two separated variables.

4. Verification

  • Gate-1 closed-form verifier: the uint4 expansion's output vs expand_q6_elem element by element, 512,000 elements 0 mismatches — defends against algebra errors in the closed form (the bookkeeping trap above).
  • Parity dump (raw + padded) 1/0: kernel output bit-identical to the CPU reference — defends against layout/alignment assumptions drifting on the real registration path.
  • greedy-32 byte-identical: the end-to-end decoded token stream unchanged — defends against the cumulative class of errors where "the numbers look right but the decode trajectory drifted".
  • suite 166/0/3: full graph-path regression — defends against the change spilling into branches beyond q6_K.
  • ncu Warp State before/after: long_scoreboard 85.5% → 33.6% (13.7 → 3.6 cy), L1/TEX throughput 14.8% → 33.6%, compute 27.0% → 37.8% — the mechanism's evidence chain: the widening's gain really did land on the stall that was attributed.

5. Results

Metricbeforeafter
whole-prefill (7B q4_k_m, same-window A/B median)1979.92605.2 (+30.7%)
q6_K attn_v GEMM kernel1.70 ms0.654 ms (−61.5%)
long_scoreboard85.5% (13.7 cy/inst)33.6% (3.6 cy)
L1/TEX throughput14.8%33.6%
compute (pipe utilization)27.0%37.8%
registers / spill80 / 4 B80 / 4 B (unchanged)

vs-llama advanced from 1.65× after r40 to 1.27×. This is the q6_K line's biggest single-step lever (against r39's +13.3% and r40's +13.0%), register-neutral, and with zero memory cost. The remaining 33.6% at this point pointed mainly at the dsc d·sc reads (unmerged narrow loads) plus KSPLIT=2's intrinsic overhead — this attribution directly begat r42 (widening the dsc reads) and r43 (PC-sampling precise attribution).

This step's place in the campaign: at r38's start q6_K's per-GMAC cost was 368.9 µs/GMAC, and the three steps r39/r40/r41 narrowed it down to this step's −61.5% kernel time; the q6_K GEMM was no longer a "slower than f16" burden. But whose the 33.6% residual stall was — the answer of the moment (narrow dsc reads) got only the consumer side of the 26 percentage points right — r42 used a zero-gain experiment to veto the width theory, and r43's PC-sampling finally broke it into the actionable list recomb 45% / A-staging 28% / dsc consumers 26%. r41's "register-neutral" property also verified, on the r40/r41 combination, that levers can stack; that lesson was cited repeatedly in the later cp.async series (r45/r53).

(Comparison note: the 2015.6 recorded in the r40 chapter is that session window's absolute value; r41's A/B re-anchored at 1979.9 within its own session window. Cross-window absolute values are not comparable — read only same-window deltas.)

6. Lessons

  1. The q6_K line's biggest single-step lever was load width, not scheduling: byte-granularity global loads in a staging loop are a textbook L1TEX scoreboard factory — 32 serially consumed LDG.E.U8 cannot be hidden even at three resident blocks.
  2. Register-neutral levers compose with occupancy levers: check the ptxas landing point before starting; a widening that repays no register debt is free, and anything that does repay (a deeper pipeline, more buffers) must be priced against it.
  3. Widening cuts instruction count, not the latency chain: the first consumer ALU still eats the full load round trip — the residual stall needs PC-sampling to find the consuming instruction before it can be broken down further (r43).

← 43 · r40 third resident block · Index · 45 · r42 stage-wide dsc scale read →

45 · r42 — q6_K stage-wide dsc scale reads (REVERTED)

Result: whole-prefill 2607.5 → 2602.6 tok/s (−0.19%, within the noise band); L1TEX throughput 33.64 → 26.45% (the dsc data volume really did fall), but the long_scoreboard share did not move at 33.6% (3.6 → 3.8 cy) → the premise was falsified, reverted. r41's "the residual is mainly dsc reads" attribution does not hold for the dsc path. Commit: a1421e6 (record-only commit — the code change was reverted; the tree is at the r41 shape, cmp-match with HEAD). Date: 2026-09-05.

Code provenance note: r42's code change is gone from the tree, and the record commit a1421e6 contains docs only (verified with --stat: only docs/CUDA_OPTIMIZATION.md +52 and docs/LLAMA-CPP-MMQ-ANALYSIS.md +27). Per STYLE hard rule 0, this doc takes its evidence from narration + current-tree code (the post-revert shape, which is also the pre-change shape); the trialed change's shape is described from the record's text.

1. Background — where things stood

r41's uint4 group loads cut the B-expand fetch width 16-fold, long_scoreboard fell from 85.5% to 33.6%, and whole-prefill gained +30.7% in one step. The winner handily left behind the next map: r41's record explicitly spelled out its own residual's composition — "the remaining stall (33.6% L1TEX) is now mainly the dsc d·sc byte/16-bit reads (uncoalesced) plus the KSPLIT=2 intrinsic".

To r42 this attribution looked almost like a success that could be copied directly: widening B-expand's narrow loads won, so why wouldn't widening dsc's narrow loads win? The dsc reads' shape is three narrow LDGs per (row, chunk): the f16 d sits at in-block offset 208 (1×u16), the two sub-block scales sc[2c%16], sc[(2c+1)%16] start at offset 192 (2×u8); a KDR=2 kt window is 6 narrow loads per row. Fewer than B-expand's 32, but the same "narrow, scattered, immediately consumed" shape — and the scale reads are mixed into the staging loop, each one a potential seed of a scoreboard event.

Where this step's absence would stall: the 33.6% long_scoreboard is the next visible ceiling; if it really is driven by dsc's narrow loads, one more widening of the same kind should knock it down. r42's value is not in the gain but in using one clean experiment to veto this shortcut, forcing r43's instrumentation upgrade.

How reasonable the hypothesis was at the time. From the layout, the dsc reads are even more "scattered" than B-expand: B-expand's ql/qh are contiguous runs within a block (r41's uint4 merges at least within a group), while dsc's three reads sit at three separate in-block offsets (192+s0, 192+s0+1, 208); across rows, adjacent od rows j and j+1 have block bases nsb·224 B apart — the two rows' dsc reads never land in the same sector. So the intuition "these narrow reads are manufacturing L1TEX pressure" is entirely sound — after r42 we know it holds in the throughput sense, not the critical-path latency sense. The intuition was not wrong; what was wrong was treating throughput pressure as the stall source.

2. Principle — the GPU mechanism

The trialed change (as described by the record; not on the tree). Widen the dsc reads per stage window: within a KDR=2 window, the two chunks' 4 scale bytes merge into one u32 (4 contiguous B from blk+192+s0; s0 is even and the block base is 16-aligned, so 4 B alignment holds) and d merges into one u32 — each (row, window) drops from 6 narrow loads to 2 wide loads (~3× instruction cut on this path); the multiplication order is unchanged and dsc = d·sc is bit-identical; the same (bstride & 15) == 0 alignment gate.

The hypothesis it tests, and why the hypothesis could be wrong. r41's win mechanism was twofold: it cut the number of loads, and those loads were serially consumed (recomb→STS follows right behind), so the L1TEX round-trip latency could not be hidden. The dsc reads satisfy the first half (many, narrow loads) but not the equivalent form of the second half — if the dsc loads' results are not eaten immediately by a tightly dependent instruction, then each load's latency is already covered by other warps or by subsequent independent instructions, and swapping 6 narrow loads for 2 wide ones just moves fewer bytes and issues a few instructions that were never blocking.

The divide between the two metrics (this doc's core concept).

  • L1TEX throughput %: a data-plane metric — the byte volume the L1TEX pipe moves per unit time as a fraction of peak. Many narrow loads, scattered bytes → a high number.
  • long_scoreboard (CPIStall): a latency-plane metric — cycles a warp spins because some instruction's input dependency (an earlier load's result) is unmet. It cares only about unhidden round trips on the critical path.

r41 happened to move both at once (what it cut was a serially consumed chain), which made people assume they always move together. r42's experiment design happens to pull them apart: bytes down, throughput down, stall unmoved — proving that on the dsc path only the data plane moved. Where was the latency hiding? The record's refined answer was later nailed down by r43: at the consumer end (I2F.S8 waiting on the d·sc result), which widening the load end can never reach.

The arithmetic of bytes vs transactions (why "traffic falls" ≠ "latency falls"). The data the dsc path must move is fixed: per (row, chunk) = 1×u16 d + 2×u8 sc ≈ 4 B; per block per kt (128 rows × 2 chunks) = 1,024 B — not one byte more or less before or after widening. What changes is the transaction structure: in the narrow shape these are 768 LDG.E.U8/U16 (most landing on different bytes of the same 32 B sector → low sector utilization, many instructions); in the wide shape 256 u32s → transactions ÷3, sectors fall with them, and pipe utilization (throughput %) drops from 33.64% to 26.45%. But the L1TEX round-trip latency behind each load is not one cycle shorter — if those latencies were never consumed by the critical path anyway, the saved transaction overhead shows up only as a −1.8% kernel duration, not in the stall. That is the complete mechanism of "traffic down, stall unchanged".

3. Implementation

3.1 Design choices (why this shape and not another)

  • Inherit r41's stage-level granularity: one dispatch per thread per kt window, dsc's two wide loads folded into the same staging macro, no new synchronization, no new buffers — the variables converge to "only the load width changed", which is what makes the experiment clean.
  • Why widen only to (row, window) granularity and not merge across rows: dsc's mergeable bytes exist only within one row's block (the sc region from 192 + the d at 208, ≤16 B apart); scales across od rows are nsb·224 B apart, and "merging" them is not widening but gather — swapping one contiguous-segment load for several scattered-address accesses only increases transactions. So (row, window) is the only legal widening granularity under this layout; r42's shape is not design conservatism, it is geometry.
  • Bit equivalence first: the multiplication order of scales and d keeps the per-chunk original order (d·sc0 first, then d·sc1); the widening changes fetching, not arithmetic, and dsc and the final output stay bit-identical — the verification gates (parity/greedy) need no exemptions.
  • The same alignment gate: reuse (bstride & 15) == 0; the raw 210-B test path keeps narrow reads.

3.2 Key code

The trialed change is not on the tree (reverted; see the note at the top). What the tree keeps is the narrow-read shape of the dsc reads — the object r42 worked on, and the post-revert status quo (after r53 it survives in the RAW_STAGE_Q6K_BT macro as the fallback when the W_dsc plane is absent):

// src/cuda_kernels.cu — the dsc section of RAW_STAGE_Q6K_BT (narrow-read fallback arm)
// per (row, chunk): 1×u16 d (offset 208) + 2×u8 sc (from offset 192) = 3 narrow LDGs
for (int x = threadIdx.x; x < MMQ_NBJ * KDR; x += blockDim.x) {
    const int r = x % MMQ_NBJ, kd = x / MMQ_NBJ;
    const int j = j0 + r, c = (kt) * KDR + kd;
    float dsc0 = 0.0f, dsc1 = 0.0f;
    if (j < od && c < nchunk) {
        const uint8_t* blk = W + (size_t)j * ((size_t)nsb * bstride)
            + (size_t)(c >> 3) * bstride;
        const float d = h2f(*(const uint16_t*)(blk + 208));       // f16 d: narrow read 1
        const int s0 = 2 * (c & 7);
        dsc0 = d * (float)(int8_t)blk[192 + s0];                  // sc even: narrow read 2 (→ I2F.S8)
        dsc1 = d * (float)(int8_t)blk[192 + s0 + 1];              // sc odd: narrow read 3 (→ I2F.S8)
    }
    sdsb[(size_t)kd * MMQ_NBJ + r] = make_float2(dsc0, dsc1);     // consumption: I2F right behind the load
}

Of the three narrow reads, the two int8_t→float conversions are exactly r43's PC-sampling-named I2F.S8 consumers — the stall is booked to those two conversions, not to the loads. r42 worked the load end, which is prescribing at the wrong address.

What the real consumer-end fix looks like — r56's W_dsc f32 plane (precomputed d·sc at registration; the live shape on the tree):

// src/cuda_kernels.cu — the dsc section's W_dsc arm (r56): I2F and the scale reads leave the hot loop together
if (W_dsc != nullptr) {
    const int nc2 = MMQ_NBJ / 2; /* 16-B groups per kd (2 float2s) */
    for (int g = threadIdx.x; g < KDR * nc2; g += blockDim.x) {
        const int kdd = g / nc2, m = g % nc2;
        const int j = j0 + 2 * m;
        const bool full = (j + 1 < od);
        gemm_cp16(                                                     // cp.async: also made asynchronous along the way
            (__half*)(void*)(sdsb + (size_t)kdd * MMQ_NBJ + 2 * m),
            (const __half*)(const void*)(W_dsc
                + ((size_t)(c0d + kdd) * (size_t)od + (size_t)j) * 8),
            full);
    }
}

The contrast makes the lever-class difference visible: r42 swapped the fetch shape (6 narrow → 2 wide, I2F still there); r56 swapped the work's location (d·sc computed at registration, leaving one contiguous 16 B copy in the kernel, I2F gone entirely). The latter landed at +2.35% — which also proves in reverse that r42's veto was not "this path is hopeless" but "this lever class is hopeless".

3.3 Pitfalls

  • No pit at the correctness gate: parity 1/0 (raw+padded) and greedy-32 byte-identical passed first try — the arithmetic never moved, only the fetch shape did. This step's "pitfalls" are all in diagnosis.
  • The "the metric improved" trap: ncu showed L1TEX throughput 33.64 → 26.45%, Memory 30.50 → 26.77%, kernel duration −1.8% — three numbers all "getting better". But watching only throughput or duration would misread −1.8% as "right direction, insufficient force", leading to doubling down on this path. When the constraining metrics (stall share 33.6%, wall −0.19%) don't move, all the improvement metrics moved for nothing.
  • No-Eligible 58.0 → 58.6% is also a signal: the issue-port wait share did not fall, further corroborating that the bottleneck is not issue-side volume but waiting on the dependency chain.

4. Verification

  • Parity dump (raw + padded) 1/0: kernel output bit-identical to the CPU reference — defends against fetch/reassembly errors introduced by the widened loads (the byte-order trap when reassembling scale bytes into a u32).
  • greedy-32 byte-identical: the end-to-end decoded token stream unchanged — defends against the cumulative class of errors where the numbers "look right" but the decode trajectory drifts.
  • ncu L1TEX throughput / Memory throughput before/after: confirms the mechanism really moved (bytes/transactions really did fall) — rules out the mundane explanation "the change never took effect", guaranteeing what is vetoed is the hypothesis, not the implementation.
  • ncu CPIStall long_scoreboard before/after: 3.6 → 3.8 cy, share 33.6% → 33.6% — the key veto evidence: the target symptom did not move.
  • No-Eligible before/after: 58.0% → 58.6% — the issue-port wait share did not fall, corroborating the bottleneck is not on the issue side.
  • whole-prefill A/B: 2607.5 → 2602.6 (−0.19%), same-window interleaved median — far below the +1.5% landing bar, the wall-level final word.

5. Results

Metricbefore (r41 tree)after (r42 trial)
whole-prefill2607.52602.6 (−0.19%, noise; bar +1.5%)
kernel duration—−1.8%
L1TEX throughput33.64%26.45%
Memory throughput30.50%26.77%
long_scoreboard3.6 cy / 33.6%3.8 cy / 33.6% (unmoved)
No-Eligible58.0%58.6%

Veto mechanism: the change drove its direct mechanism (dsc byte volume, fetch count) fully into place, yet both the attributed target symptom (33.6% long_scoreboard) and the final gain (wall) stayed unmoved — the hypothesis "the width of dsc's narrow loads drives the residual stall" was cleanly falsified. Keeping a zero-gain width branch would only add a SASS-comparison burden to every later step, so the whole segment was reverted and the tree keeps the r41 shape.

Metric behavior of r41 vs r42 side by side (same kernel, same metrics, two back-to-back rounds):

Metricr41 (uint4 B-expand, +30.7%)r42 (wide dsc reads, −0.19%)
L1TEX throughput14.8% → 33.6% (↑: the pipe really got busy)33.64% → 26.45% (↓: transactions really got fewer)
long_scoreboard85.5% → 33.6% (↓ moved with it)33.6% → 33.6% (did not move at all)
compute27.0% → 37.8%— (no improvement)
whole-prefill+30.7%−0.19%

r41's widening hit a serially consumed chain, so the two metrics had to move together; r42's widening hit transaction overhead that was not on the critical path, so only the throughput plane moved. Read side by side, the decoupling of "throughput improved" from "wall improved" is plain — this comparison was the most persuasive attribution evidence before r43.

Under what future conditions a retry is worthwhile: only when the next profile shows long_scoreboard booked to dsc's load instructions (not their consumer ALUs) does widening become a candidate again. r43's PC-sampling then delivered the final verdict: the dsc-side stall is booked to the I2F.S8 consumer (26% of the residual) — a latency source, not a width source. The fix that finally landed therefore changed lever class: r56's W_dsc f32 plane precomputes d·sc at registration, moving the consumer end out of the hot loop together with the scale reads.

6. Lessons

  1. Cutting bytes is not cutting latency: the stall is booked to the consuming instruction; profiles must chase down the dependency's downstream op, not stop at the load.
  2. Throughput-class and stall-class metrics can move in opposite directions: r41 making them move together is the special case (a serially consumed chain), not the rule; every time, list the constraining metric separately and check whether it moved.
  3. A zero-gain but mechanism-complete experiment is an instrument: r42's −0.19% bought the conclusion "dsc width is innocent" and directly named the next tool — warp-stall source-level sampling (r43).

← 44 · r41 q6_K B-expand uint4 widen · Index · 46 · r43 PC-sampling attribution →

46 · r43 — PC-sampling attribution + pre-expand-B parity FAIL (MEAS-ONLY + REVERTED)

Result: the 33.6% long_scoreboard was named down to the consuming instructions: B-expand recomb (LOP3/SHF) 45% + A-side staging (STS.128) 28% + dsc consumers (I2F.S8) 26%. The trial "pre-expand B at registration (the W_exp plane)" was byte-perfect (readback 0/17,920) but parity FAILED (diff 448 @ index 554) and could not be isolated within budget → fully reverted; the paradox was solved by r44's one-line root cause (dense stride mismatch). Commit: b7fa305 (record-only commit — the code never landed; the tree is at the r41 shape, cmp-match with HEAD). Date: 2026-09-05.

Code provenance note: r43's trial code was reverted and the record commit contains docs only (verified with --stat: two docs, +78/+42). Per STYLE hard rule 0, this doc's attribution and failure narrative come from the docs/CUDA_OPTIMIZATION.md P6 r43 chapter + master table row 57 (sample counts quoted from its mirror LLAMA-CPP-MMQ-ANALYSIS.md §11.23); the code excerpts testify with the current tree's final correct form of the mechanism (r53's expand_q6k_dense + the two-gate test) — these are precisely the surviving versions of the two gates r43 built back then.

1. Background — where things stood

r42 turned "widen the dsc narrow reads" into a clean zero-gain experiment: byte counts and throughput both moved, yet the 33.6% long_scoreboard did not budge one notch. r42's parting conclusion was unambiguous: stop guessing; bring in instrumentation that resolves individual instructions — "Warp-Stall-Sampling source attribution on mmq_raw_nb_bt_q6k_kernel<2> to identify the actual instruction behind the 33.6%".

Where the instrument generation gap was: the ncu Warp State used in r41/r42 only answers "which class of wait is the warp stuck in" (long_scoreboard = an L1TEX dependency), not "stuck on which instruction". The byte counters of the r13/r21 era (sectors, queues) are a data-plane view. 33.6% had now survived two rounds un-dismantled, which means the mental model of it ("which load causes it") was wrong in direction — a tool giving a PC (instruction-address)-level distribution was needed.

Meanwhile, the campaign-level situation: after r41, whole-prefill 2605.2 (vs-llama 1.27×), with the q6_K kernel still the largest single node on the default path. Even after r41's 16× cut in load count, 33.6% of stall remained; if that number were "irreducible", the q6_K line would be over — so r43's second task was to decompose the 33.6% into an actionable lever list, and even a failure had to end knowing why it failed.

2. Principle — the GPU mechanism

The sampling mechanism. ncu's source-counter sampling (--set full --section SourceCounters, viewed in the --page source source-level view) periodically hardware-samples each warp during kernel execution and books the warp's state at that moment (issue-stalled long scoreboard, waiting on a barrier, etc.) to the instruction at that warp's current PC, aggregating per-instruction counters such as pcsamp_warps_issue_stalled_long_scoreboard.

This is the third tier of the campaign's instrumentation lineage. The previous two tiers each had blind spots:

InstrumentRounds usedQuestion answeredBlind spot
Byte/sector counters (r13 forensics, r21 sectors −28.6%)r13, r21data plane: who moves how many bytessectors saved, wall unmoved (r21's "stall conservation" lesson)
Warp State summary (since r20)r20, r41, r42the class shares of stall (long_scoreboard 85.5% → 33.6%)class-level, not instruction-level — r42 misfired because of it
PC-sampling source-level sampling (this step)r43which instruction's PC carries the stallsampling is statistical; it needs sufficient sample volume

Once r42's zero-gain experiment sealed off the "add more width" road, the third tier was the natural next step: to keep dismantling the 33.6%, we had to know which specific instructions carry it.

Why it books to the "consumer" and not the "loader". long_scoreboard's semantics: this warp wants to issue some instruction, but one of its input operands depends on a load that has not yet returned (the L1TEX round trip is incomplete). The PC where the warp stalls is the instruction waiting for the input — i.e. the consumer. That is:

  • The Warp State summary says "warps are waiting on L1TEX";
  • PC-sampling says further "it is this I2F / this LOP3 / this STS that is waiting".

For a load → ALU consume chain, the load side can change width at will (r42 proved it useless); as long as the consumer-side tight dependency remains, the stall is booked to the consuming instruction. That is the mechanical explanation of r42's phenomenon, and the source of this doc's first payoff.

What this instrument can and cannot tell you. It can name the 33.6% share down to instructions (this doc's deliverable), but it does not explain "why this instruction waits this long" — where the latency comes from (L1 round trip, bank conflicts, dependency-chain depth) still needs mechanism hypotheses backed by SASS. r43's usage is therefore two-stage: first sample-and-name (45/28/26), then explain each mechanically and derive the lever class. One more limitation: sampling is statistical, and with insufficient sample volume the shares of low-frequency instructions are untrustworthy — this round's 51,605 samples with all three big heads in the thousands make the attribution base solid, but the fine items "below 4th place" should not be quoted.

The sample-volume account. Of this round's 51,605 samples, 16,647 landed on long_scoreboard = 32.3%, consistent with the 33.6% share reported by the Warp State summary — two independent instruments interlock, and the attribution's foundation is stable.

3. Implementation

3.1 The attribution result: three consuming instructions split the 33.6%

Source-level sampling of mmq_raw_nb_bt_q6k_kernel<2> (post-r41 shape) split the long_scoreboard share across three consuming instructions:

ShareConsuming instructionWaiting onMechanism reading
45%LOP3.LUT 0xff + SHF (the B-expand recomb's first ALU group)the uint4 ql/qh load round trips after r41's wideningr41 cut the count, but double buffering put staging at the top of the kt loop with recomb right behind — the within-warp tight dependency chain remains; r39's double buffering only hides the cross-warp part
28%STS.128 (the A-side qa8/sda bulk LDG→STS copy)the A-side staging's global load round tripspure copy latency; also falsifies r42's side hypothesis — the A reads are contiguous merged uint4 accesses, there is no "strided A"
26%I2F.S8 (dsc's (float)(int8_t)sc conversion)the d/scale byte load round tripsstamps r42's verdict: dsc is a latency source, width innocent — widening loads can never reach a stall booked to the consumer

The three sum to ≈ 33.6% (45+28+26 = 99% of the named share). The mechanism reading of each:

  • 45% recomb (LOP3/SHF): after r41 the recomb chain is "uint4 load → first byte extract+mask (byte-wise and/or of the LOP3.LUT 0xff kind) → shift-assemble → −32 → STS". PC-sampling books the stall to the first ALU — it waits out the entire load round trip. So this 45% is the other face of r41's legacy: loads went from 32 to 2, but the tight "staging→consume" dependency structure survived intact, and the first consumer's wait is the chain's entire exposure. This rules out scheduling micro-tuning like "split the recomb differently/denser" — either eliminate the consumer end (pre-expansion), or make the loads asynchronous (cp.async).
  • 28% A-staging (STS.128): in the A-side qa8/sda bulk LDG→STS, the STS is the consumer of the load results, so the stall books to it. Its value is falsifying r42's side hypothesis — r42 suspected the A reads were strided accesses across rows and blocks, but the A reads are inherently contiguous uint4s (the prepass pre-transposed pad40 plane); the 28% is pure copy latency, not "strided A".
  • 26% dsc consumer (I2F.S8): the narrow read→conversion of (float)(int8_t)blk[192+s0]. It mechanizes r42's verdict: a stall booked to the conversion means the load end is irrelevant at any width — only moving the d·sc computation out of the kernel (r56's W_dsc plane) or moving the bytes into an asynchronous pipe gives this 26% room to move.

The list's execution value is in its ordering: the biggest head is the recomb's consumer end — taking the recomb out of the hot loop entirely (not changing its input's width) is the correct lever. That is §3.2's trial.

3.2 The trial: pre-expand B at registration (the W_exp plane)

The idea is isomorphic to what r38 did — moving the recomb out of the mma inner loop — but goes further: out of the entire kernel. At load time, expand the padded raw-W once into a centered-int8 dense plane W_exp (od × id, row pitch = id, super-block pitch = 256), so the kernel's B staging degenerates from "read ql/qh + recomb" into a pure bulk copy (same shape as the A side), and the 45% LOP3/SHF samples should vanish wholesale.

Why registration time is the correct place to do the work. The recomb is a pure per-weight function: the same weight's every byte is consumed repeatedly across one prefill (every output tile re-reads B), but computed only once at registration — the amortization is infinite. The cost is an od × id-byte dense plane (the q6_K tensors of q4_K_m 7B total ~1.5 GB in magnitude; r53's landing measured +1.52 GB device). At r43 this was a clear trade: 1.5 GB for a 45% sample share — worth trying, provided the correctness gates pass first.

The registration-time expander surviving on the tree (the corrected post-r44 form; the trial's "intended semantics" matches it):

#![allow(unused)]
fn main() {
// src/cuda.rs — expand_q6k_dense: padded raw-W → dense centered-int8 plane
// input row pitch nbe*224 (padded); output row pitch id, super-block pitch 256 (DENSE)
pub fn expand_q6k_dense(padded: &[u8], od: usize, id: usize) -> Vec<u8> {
    const Q6KB: usize = 210;
    const Q6KPB: usize = 224;
    let nbe = id / 256;
    let row_len = nbe * Q6KPB;
    let mut out = vec![0u8; od * id];
    for j in 0..od {
        let prow = &padded[j * row_len..(j + 1) * row_len];
        let orow = &mut out[j * id..(j + 1) * id];
        for sb in 0..nbe {
            let blk = &prow[sb * Q6KPB..sb * Q6KPB + Q6KB];
            let (ql, qh) = blk.split_at(128);
            let obase = &mut orow[sb * 256..sb * 256 + 256];   // dense sb pitch = 256
            for it in 0..2usize {
                for r in 0..64usize {
                    let qlb = ql[it * 64 + r];
                    let qhb = qh[it * 32 + (r & 31)];
                    let s0 = (r >> 5) * 2;
                    let e0 = it * 128 + r;
                    obase[e0] = ((qlb & 0xF) | (((qhb >> s0) & 3) << 4)).wrapping_sub(32);
                    obase[e0 + 64] =                            // e+64 shares the same ql byte
                        (((qlb >> 4) & 0xF) | (((qhb >> (s0 + 4)) & 3) << 4)).wrapping_sub(32);
                }
            }
        }
    }
    out
}
}

Note the juxtaposition of the two strides: input side sb * Q6KPB (padded 224), output side sb * 256 (dense). This is exactly the pit r44 would expose — see §3.3.

And the r43 trial's kernel arm consumed the plane as (per the record) "the raw W pointer expression + reading the W_exp buffer" — i.e. indexing a densely laid-out plane with the padded raw-W address algebra W + j·(nsb·bstride) + sb·bstride. The correct kernel arm as written after r53 (for contrast):

// src/cuda_kernels.cu — EXP arm (r53 landed shape): dense indexing W_exp + j*id + sb*256 + cbase*32
const int nc = (KDR * 32) / 16;   /* 16B groups per row */
const int ncopy = MMQ_NBJ * nc;
for (int g = threadIdx.x; g < ncopy; g += blockDim.x) {
    const int jj = g / nc, cc = g % nc;
    const int j = j0 + jj;
    const bool full = (j < od) && (sb < nsb);
    const uint8_t* src = W_exp + (size_t)j * id          // dense row pitch = id (not nsb*bstride!)
        + (size_t)sb * 256                               // dense sb pitch = 256 (not bstride!)
        + (size_t)(cbase * 32 + cc * 16);
    gemm_cp16((__half*)(void*)(qbexpb + (size_t)jj * (KDR * 32) + cc * 16),
              (const __half*)(const void*)src, full);    // pure copy: the recomb was completed at registration
}

3.3 Pitfalls: the byte-correct but parity-red paradox

The two gates' readings came out inverted:

  • Content gate green: the W_exp device readback vs the host mirror was 0/17,920 mismatches — the plane itself is byte-correct.
  • Output gate red: kernel parity failed, max diff 448 @ index 554, and this signature was bit-identical across three kernel variants (uint4 copy of W_exp, per-byte copy of W_exp, and a control arm expanding raw W in place via expand_q6_elem which passed, while any variant consuming W_exp failed).

The control arm passing narrowed the suspicion to the extreme: branch, bounds, staging buffers, and the staging→mma path are all shared and correct; the only thing that follows the failing variants is the W_exp buffer itself — a "byte-correct, consumption-fails" combination. r43 could not isolate this kernel-side interaction within budget (the hypothesis of the moment was some in-kernel aliasing) and reverted by discipline (cmp-verify = HEAD r41).

In hindsight (r44's post-mortem language), the pass/fail split of the three variants was already pointing at the answer: "that expression on W" was exactly all-right, "the same expression on W_exp" exactly all-wrong — the only degree of freedom in the difference is the data source's layout, and the failure did not vary with the copy method (uint4/per-byte), placing the problem in the address → data mapping, not the copy's execution. The missing step at the time was writing the two layouts' strides side by side and doing one line of arithmetic; r44 did exactly that line (r43's three variants "split exactly as that predicts"). The lesson is not "should have tried more within budget" but: for any change where two layouts coexist, step one is writing the two stride sets' difference into the verification checklist.

r44's one-line root cause demolishes the paradox: the address expression and the plane layout mismatch. r43's kernel indexed the densely laid-out W_exp (row pitch id, sb pitch 256) with the padded raw-W strides (row pitch nsb·bstride, sb pitch bstride = 224). Take id=256 (nsb=1) and do the concrete arithmetic:

r43's address:      W_exp + j*(nsb*bstride) + sb*bstride = W_exp + j*224 + sb*224
the correct address: W_exp + j*id          + sb*256      = W_exp + j*256 + sb*256

at j=1, sb=0: it reads [224 .. 224+224) — in the dense plane that is
  row 0's [224..256)  (32 B, the last 32 elements of row 0)
+ row 1's [0..224)    (224 B, the first 224 elements of row 1)

That is, except for j=0, every row's window slides wholesale into "the previous row's tail + this row's head" — element-level misalignment, yet every byte read is a legal centered int8 (some other element's value in the −32..31 range). So the output is a "systematically wrong, but not absurd" diff of 448 (@ index 554), not obviously-fake garbage. The control arm expanding raw W passed precisely because raw W is the padded layout — the expression happens to be right for it. The content gate only verifies CONTENT; OFFSET is decided by the address expression; the two gates each guard half, and missing either leaks.

4. Verification

  • PC-sampling sample-volume self-consistency: 16,647/51,605 = 32.3% vs the Warp State summary's 33.6% — the attribution base interlocks with the existing instrumentation (defends against sampling bias / the new instrument reading the wrong object).
  • Content gate (readback): W_exp device readback vs host mirror 0/17,920 — verifies the plane's content (defends against expansion-algebra errors; its surviving version is the host+device two-segment assertion cuda_q6k_exp_dense_byte_exact on the tree, see below).
  • Output gate (parity dump): diff 448 @ 554 — verifies offsets and end-to-end semantics (defends against "content right, addresses wrong").
  • Control-arm variants: expanding raw W in-branch passed, consuming W_exp failed — the bisection instrument (compressing suspicion from the whole path to a single buffer).

The live form of this two-gate setup on the tree (written in the r53 era; the semantics are exactly r43's two gates):

#![allow(unused)]
fn main() {
// src/graph/cuda_backend.rs — cuda_q6k_exp_dense_byte_exact (excerpt)
// host expander vs independent scalar mirror (content gate, host half)
let host = crate::cuda::CudaState::expand_q6k_dense(&padded, od, id);
let hmis = host.iter().zip(want.iter()).filter(|(a, b)| a != b).count();
assert_eq!(hmis, 0, "expand_q6k_dense vs mirror ({od}x{id})");
// device upload + pinned readback vs the same mirror (content gate, device half)
let p = state.get_weight_ptr(&exp_name).expect("W_exp registered");
let mut got = vec![0u8; od * id];
state.copy_from_device_pinned(p, &mut got);
let dmis = got.iter().zip(want.iter()).filter(|(a, b)| a != b).count();
assert_eq!(dmis, 0, "device W_exp vs mirror ({od}x{id})");
}

(r43's lesson was then written into r53's landing preconditions: the content gate must exist simultaneously with the kernel-parity gate — the test above covers only the content gate; offset correctness is guarded separately by the kernel arm's dense indexing + the parity dump.)

5. Results

  • Attribution (this step's deliverable): the 33.6% long_scoreboard = B-expand recomb 45% + A-staging STS.128 28% + dsc I2F.S8 consumers 26%; the residual is not irreducible — it is within-warp staging latency exposure (double buffering hides across warps, not across a warp's internal dependency chain).
  • Trial (this step's veto): pre-expand-B's W_exp was byte-correct (0/17,920) but parity FAILED (diff 448 @ 554, same signature in all three variants) → not isolable within budget → REVERTED, the tree cmp-matches HEAD; whole-prefill stays around 2605.2 (vs-llama 1.27×), no wall change.
  • Aftermath (forward reference): r44 dissolved the paradox with a one-line root cause (dense/padded stride mismatch); the corrected version went parity-green but wall-neutral (−0.42%) — removing the recomb merely transferred the latency's carrier; the final landing was r53's basket (r44's de-work + r45's de-wait + cp.async) at +5.03%. r43's attribution list was exactly that basket's design input.

The q6_K convergence judgment after r43 (in the record's own terms): the residual is not irreducible — it is within-warp staging latency exposure, with two classes of physical lever: cp.async-ify the staging of raw A (and pre-expanded B) (llama.cpp's structure), or a register-constrained split-phase. This judgment set the direction of all three following rounds:

RoundLeverKernel effectWall effectStatus
r41B-expand uint4 widen1.70 → 0.654 ms+30.7%LANDED
r42dsc read widening−1.8%−0.19%REVERTED
r43PC-sampling attribution + pre-expand-B—— (parity FAIL, reverted)MEAS + REVERTED
r44W_exp stride fix (recomb vanishes)−10.9% cycles−0.42%REVERTED
r45A-side cp.async (de-wait)−10.2%−0.34%REVERTED
r53r44+r45 merged basket + cp.async Bffn_down −20.5%+5.03% (3024.7 → 3176.9)LANDED

Viewed alone, r43 is one attribution plus one failure; viewed in the campaign line, it turned the 33.6% from "a number" into "three executable levers", and r44/r45's "each wall-neutral alone, over the wall merged" is the direct product of testing r43's list item by item.

6. Lessons

  1. Attribute stalls to consuming instructions, not loading instructions: long_scoreboard books to the instruction waiting for input; changing load width (r42) cannot reach a stall booked to the consumer — only eliminating the consumer-side dependency (moving it out of the hot loop) or switching to an asynchronous path works.
  2. Byte-correct ≠ offset-correct: readback only verifies CONTENT; data landing on the wrong address expression is more dangerous than "obviously wrong data" — the output stays in a legal value range and the failure signature is stable, steering the investigation toward wrong hypotheses like in-kernel aliasing.
  3. The control variant is the cheapest bisection instrument: a "same branch, swapped data source" control (raw W passes / W_exp fails) compresses the suspicion to a single buffer in one step, faster than any static review.
  4. MEAS-ONLY still counts as landing: this round's instruction-level list directly fed r44/r45/r53's designs — one failure with clear attribution beats one success with a vague mechanism.

← 45 · r42 stage-wide dsc scale read · Index · 47 · r44 W_exp stride mismatch root cause →

47 · r44 — W_exp stride mismatch root cause: fix goes parity all-green but wall-neutral (REVERTED)

Result: r43's pre-expand-B parity mystery was dissolved by a one-line indexing fix — W_exp (the dense plane, row stride = id, super-block stride = 256) was being addressed with the padded raw-W address expression (row stride = nsb·bstride, sb stride = bstride = 224). After the fix, parity cuda_prefill_mmq went 1/0 all green and ncu kernel elapsed cycles went 1,390,776 → 1,239,847 (−10.9%, the recomb ALU and its ~45% samples vanished entirely) — but whole-prefill went 2595.9 → 2584.9 tok/s (−0.42%, noise), below the +1.5% bar → REVERTED. Conclusion: after r41 the q6_K GEMM is no longer the prefill wall's bottleneck; a kernel-level win on the converged line cannot reach the wall. The same physical latency merely changed the instruction set it hides behind (load→ALU-recomb became load→STS copy, the same L1TEX exposure); the long_scoreboard share 34.9% → 57.1% is a denominator effect. Next lever: cp.async (r45). Commit: 6d02017. Date: 2026-09-05.

Provenance note: the commits cited here (r43 b7fa305, r44 6d02017) are both docs-only commits — the experimental code was reverted after measurement, before commit (cmp-match HEAD r41), so the diffs contain no code. The code excerpts in this doc all come from the current tree: expand_q6k_dense (the r53-landed host-side expander) and the EXP=true/EXP=false staging branches of mmq_raw_nb_bt_q6k_kernel — these are precisely the final landed form of r44's "one-line fix", with the address expressions word-identical (the current tree's expand_q6k_dense doc-comment even signs itself "P6 r44 / MMQ-analysis §11.24").

1. Background — where things stood

1.1 A line climbing fast

On r44's day the P6 q6_K line was in the steepest climb of the whole campaign. The trajectory of the previous four rounds: r38 BT-style raw-byte mma kernel (1518.4 → 1561.9, +2.87%, while also discovering that "q6_K = 8 sub-blocks × 32" is wrong — it is actually 16 sub-blocks × 16, with KSPLIT=2 each carrying an independent dsc); r39 KDR=2 double buffering (→ 1777.5, +13.3%, overlapping the B-expansion's staging latency with compute); r40 __launch_bounds__(256,3) third resident block (→ 2015.6, +13.0%, while falsifying the "0-spill gate"); r41 B-expand uint4 widening (→ 2605.2, +30.7%, kernel 1.70 → 0.654 ms, merging 32 per-byte LDG.E.U8 into 2 uint4s). Whole-prefill rose from 1518 to ~2600 tok/s, and vs-llama closed from 2.13× to 1.27×.

But r41 also left a precise residual reading: L1TEX scoreboard 85.5% → 33.6%, with the remaining stall's composition then unknown. r42 tried "widen the dsc scale reads" — bytes fell, the stall did not move at all (−0.19% wall) — the first hint that "the intuition about the stall's source might be wrong". r43 used PC-sampling (pcsamp_warps_issue_stalled_long_scoreboard + --page source) to attribute the stall by consuming instruction and produced that decisive table: B-expansion recomb (LOP3/SHF) 45% + A-side staging STS.128 28% + dsc's I2F.S8 consumption point 26%. All three sources point at one physical process: B's global→smem round trip exposing L1TEX latency.

There is a reasoning gap that only became visible afterwards, worth writing down: what the attribution table gives is sample shares, not wall-clock seconds. The step from "45% of samples are on the recomb" to "deleting the recomb wins back 45% of the stall" implicitly assumes that slice is a serial critical path — an assumption that held when q6_K was the wall and does not hold after r41. r44's wall-neutrality is the first empirical failure of that implicit assumption; r47 would write "wall decompositions have a shelf life" as a formal lesson.

1.2 The seemingly direct road: move the recomb out of the hot loop

The attribution table lays out a superficially direct road: 45% of the stall samples hang on the recomb ALU, and the recomb is a pure function — each output element depends on a single byte pair from (ql, qh), unrelated to tokens or any runtime state. So why not move it out of the kernel? Pre-expand once at weight registration into the centered-int8 plane W_exp, and the kernel's B staging degenerates into a pure bulk copy: all ql/qh loads and nibble-unpack ALU deleted from the hot loop, and by the attribution logic the 45% of samples should largely disappear.

The idea is not a new invention — r18 tried "load-time B pre-expansion" (bulk-copy staging, KD=8 +0.9% noise, KD=4 −19%, plus a +5.8 GB VRAM cost) and was reverted. r43's version differed in two ways: the timing moved from "load" to "registration" (effective with the NB-BT gate, serving only the q6_K type), and r43 had the attribution table's endorsement — a clear 45% target. Then came that famous parity FAIL:

  • The W_exp plane itself passed the device-side readback 0/17,920 bytes, all correct;
  • cuda_prefill_mmq parity FAILED, diff 448 @ index 554;
  • The two failing pre-expansion variants (uint4 version, per-byte version) had bit-identical diffs, while the control arm expanding raw W in-kernel passed.

Not isolable within budget, r43 ended at cmp-match HEAD, leaving a verdict r44 would take over: "byte-correct data at the wrong address expression is worse than obviously-wrong data — the readback validates CONTENT, not OFFSETS (r44 resolves it)".

1.3 r44's dual task

So r44's task had two layers. The first is diagnosis: solve the parity mystery, or "pre-expand B" cannot even get one clean measurement. The second is adjudication: how much wall clock is the corrected pre-expansion actually worth — the direct cash-in of r43's attribution table, and the touchstone for how much oil is left in the whole q6_K line. The two layers' answers landed the same day, and the second layer's answer (wall-neutral) shaped the campaign far more than the first (a one-line fix): together with the next day's r45 it declared the q6_K line converged and pushed the campaign's entire budget toward FA and prepass (the direction of r46/r47).

A footnote on the time axis: r44's work ultimately survived in two forms — as knowledge (the root cause + the "kernel win ≠ wall win" verdict + the pointer that cp.async is the real lever), fed into r45's design that same day; and as code (the dense indexing + the host expander), landing verbatim in r53's bundle two days later. A REVERTED doc is not a wasted step — provided the "veto mechanism + retry conditions" are written completely enough that future rounds can cite them precisely.

2. Principle — the GPU mechanism

2.1 What W_exp is (the definition at first appearance)

The W_exp plane: the q6_K weight tensor pre-expanded, according to its mathematical shape, into a dense centered-int8 byte plane — od × id bytes, one weight element per byte, the value range already shifted by −32 to center (q6_K's nibble + 2-bit combined encoding lands in ±127 after subtracting 32). It eliminates two things: encoding (an element's information is split across 4 bits of ql and 2 bits of qh) and padding (the raw layout pads each super-block's 210 content bytes to 224). The concept debuted in r18 (load-time pre-expansion, reverted for its VRAM cost); the r43/r44 version builds on demand at registration and serves only the NB-BT q6_K path. It must output centered int8 (−32 centered) because r38's KSPLIT=2 mma structure takes centered int8 as the B-operand contract directly (the single-term scaling of sum += da·dsc is built on that value range) — W_exp is not "another storage format"; it freezes the transform every kernel loop body was doing into data, with the feeding end's mma contract unchanged by a single character.

2.2 Precise definitions of the two layouts

q6_K's block_q6_K = ql[128] + qh[64] + sc[16] + d[2] = 210 bytes of content, describing 256 elements (16 16-element sub-blocks, each sub-block a d·sc scale pair). The two planes fork from here:

  • raw W (padded): each super-block's 210 content bytes padded to bstride = 224 (16-byte alignment, the premise of r41's uint4 widening). row stride = nsb · bstride (nsb = id/256), sb stride = bstride = 224. This is the data's original form in VRAM and cannot change (prefill MMQ's block_stride 224 depends on it; the dequant/embed fallbacks read it).
  • W_exp (dense): exactly 256 bytes per super-block (one element per byte, no padding). row stride = id (= nsb · 256), sb stride = 256.

The density difference is 256/224 ≈ 1.14×: the dense plane pays 14% more bytes for the unconditional address identity "element i is at offset i". The two layouts side by side (one super-block, a two-row sketch with nsb = 2):

raw W (padded, bstride=224)          W_exp (dense, sb stride=256)
row 0: [sb0: 210B content|14B pad][sb1: 210B|14B pad]   row 0: [sb0's 256 elements][sb1's 256 elements]
        ^0          ^210    ^224          ^434  ^448            ^0                ^256
row 1: base = 2*224 = 448            row 1: base = 2*256 = 512
misread row1@448 = dense row 1's elements 192..255 spliced with row 2's elements 0..191

(In the diagram ^ marks byte offsets. On the raw side each sb is 210 bytes of content + 14 bytes of padding; on the dense side each sb is exactly 256 bytes, one element per byte. The misalignment accumulates from the second row on.)

2.3 The arithmetic of the mismatch: a byte-level worked example

The wrong expression W_exp + j·(nsb·bstride) + sb·bstride pages through the dense pointer at padded strides. With the smallest id = 256 (nsb = 1), the row stride is wrongly 224:

  • Row 0 has no drift (offset 0): the 256 bytes read are exactly dense row 0 — but the in-row sb drift is 0 only because nsb=1;
  • Row 2: the correct base 2·256 = 512, the wrong base 2·224 = 448 = dense row 1's byte 192. So the kernel's "row 2" = dense row 1's elements 192..255 (64 bytes) spliced with dense row 2's elements 0..191 (192 bytes) — the bytes inside the window are real weight bytes; the window's assembly is wrong;
  • For id = 3584 (7B attn_v, nsb = 14): wrong row stride 3136, drifting 14×32 = 448 B per row; for id = 5120 (ffn_down, nsb = 20): 640 B per row. The larger the row number, the further the window read sits from the real data.

This explains the diff's shape: not wholesale garbage, but output errors that are "contiguous within each 256-element window, misaligned between windows" (448 @ index 554, a magnitude consistent with f32 accumulation differences). The drift amplifies with shape:

ShapensbWrong row stride (nsb·224)Correct row stride (id)Drift per row
id = 2561224 B256 B32 B
id = 3584 (7B attn_v)143,136 B3,584 B448 B
id = 5120 (7B ffn_down)204,480 B5,120 B640 B

The larger the row number, the further the window sits from the real data — no row except row 0 is correct. It also explains why the three variants sharing the same expression had bit-identical diffs — if the bug lived in some variant-private mechanism (staging order, a race, a barrier, register allocation), variants with different mechanisms could not produce the same diff. The fingerprint points at their only shared part: the address expression.

2.4 How much VRAM the plane costs — the lever's cost side

W_exp's size is od × id bytes, one plane per q6_K tensor on the NB-BT path. The 7B inventory: 14 attn_v + 14 ffn_down + output.weight, totaling 1,521,237,632 B ≈ 1.52 GB (r53's precise measurement; r43's "+15 MB" estimate at project time erred by counting one ffn_down's increment as the whole cost — which also explains why that path has exactly 27 launches). 1.52 GB is not small: it is the slimmed-down edition of the same tax as r18's "+5.8 GB reverted", and the reason r54 later built the MINFER_MMQ_Q6K_EXP=0 opt-out (−5.04% for 1.52 GB back). This doc records only the cost side: for any "pre-transformed plane" lever, the VRAM account belongs in the project proposal next to the latency account.

2.5 Why readback cannot catch it

r43's gate was a device-side readback comparison of the W_exp plane (0/17,920). Readback verifies content: the bytes the host wrote match the bytes the device reads. It does not verify offsets: which expression the kernel uses to read the plane is invisible to readback. A fully correct producer + a consumer paging by the wrong catalogue = "byte-correct data at wrong offsets". This combination is more dangerous than "obviously wrong data": it grants strong confidence (the plane is correct) and pushes the investigation toward the wrong direction of kernel-side races/aliasing. The lesson in full: the end-to-end correctness gate for a pre-transformed plane must include consumer-side parity; readback only rules out "written wrong", not "read wrong". r53 later paired expand_q6k_dense with an independent scalar-mirror test (host mirror AND device plane readback, both 0 mismatch) — institutionalizing this lesson.

2.6 Why "fixed correctly yet unprofitable" — latency changes form, it does not vanish

The pre-expansion's paper gain is deleting the recomb ALU. But B's physical latency never disappeared — it is the "global load → smem landing" L1TEX round trip, independent of what instruction consumes it. r41's shape: load → recomb ALU consumes → STS write-back, with the ALU consumption point absorbing the stall-wait; after pre-expansion: load → pure STS copy, the same stall-wait hanging on a thinner instruction stream. The ncu readings are this mechanism's complete signature: elapsed cycles −10.9% (the ALU and its samples really did vanish) while Warp-Cycles/Issued-Inst 11.70 → 17.84 (the wait amortized per instruction grew) and the long_scoreboard share 34.9% → 57.1% (denominator effect: total cycles shrank, the stall's absolute volume stayed roughly constant, so the share naturally rose). The same physical latency hides under a different instruction mix — this is the companion piece to r42's "cutting bytes does not cut latency". What can actually hide this latency inside compute gaps is the asynchronous copy (cp.async), which is r45's subject; §11.24's closing judgment is blunt: "The lever that actually hides it is cp.async (llama.cpp structure)" — llama.cpp's MMQ reference implementation uses cp.async precisely to move staging latency off the critical path, part of its instruction-stream advantage.

3. Implementation

3.1 Diagnosis and design choices: a one-line address fix, not a staging rewrite

The diagnosis is a four-step evidence chain, each step eliminating a class of hypotheses:

  1. Readback all-correct (0/17,920) → rules out "the host expander wrote it wrong": the plane's content is correct.
  2. The failing variants (uint4 pre-expansion, per-byte pre-expansion) have bit-identical diffs → rules out all variant-private mechanisms (staging order, races, barriers, register allocation): two variants with different mechanisms cannot produce the same diff unless the error is in what they share.
  3. The control arm expanding raw W in-kernel passes → the only shared code that is "right for raw W, wrong for W_exp" is the address expression: the data layout differs, the expression is the same.
  4. Conclusion: the expression pages through the dense plane at padded strides. The test is ready-made — swap the expression to the dense strides and parity should turn green; if it does not, a second cause remains and we return to step 1.

Parity turned green on the first try; the single-cause verdict holds.

The fix's shape is therefore pure address arithmetic: swap the base-address expression pointing at W_exp in B staging from the padded form to the dense form, touching no staging mechanism.

Two reasons to choose the minimal diff. First, the diagnosis had already locked suspicion onto the expression; a larger change would pollute this single-cause verdict (if parity stayed red after the fix, a second cause would exist). Second, the fix must preserve every budget constraint r38–r41 had banked: __launch_bounds__(256, 3)'s third resident block requires ≤80 registers — after the address terms change from j·(nsb·bstride) + sb·bstride to j·id + sb·256, the multiplication structure is unchanged (id and 256 are both compile-time-observable), and ptxas landed at 80 regs / 28 B spill; the 3-block budget held.

The dense indexing also has a friendly property later cashed in by r53: this NB-BT path's launch gate requires id % 256 == 0, so every 16-element group start of W_exp + j·id + sb·256 + cbase·32 is naturally 16B-aligned — when switching to cp.async 16B copies, the alignment premise is already in place.

One alternatives-check on "why build the dense plane at all": could we skip W_exp and let the kernel read raw W in the padded layout doing "nothing"? No — that is exactly r41's status quo; the recomb must stay in the hot loop and the 45% of samples have nowhere to go. Could the kernel read raw W but skip the recomb? No — in the raw layout the element-to-byte mapping is not an identity to begin with; the recomb is that mapping. So "pure copy" has as its sole precondition an element-ordered plane, and the dense W_exp is its minimal implementation; the stride mismatch was an address bug in landing that plane, not a design flaw.

3.2 Key code

The wrong expression (it is still in the tree today — but it serves raw W, for which it is correct). The current tree's src/cuda_kernels.cu EXP=false branch (r41's in-kernel expansion, the raw 210-B padded layout):

// Inside the RAW_STAGE_Q6K_BT macro, EXP=false && (bstride & 15) == 0 branch:
//   blk points at one super-block of raw W — row stride is nsb*bstride,
//   sb stride is bstride=224. This is correct for PADDED raw W.
const uint8_t* blk = W + (size_t)j * ((size_t)nsb * bstride)
                       + (size_t)sb * bstride;          // ← r43 used it verbatim on W_exp
const uint4 qlv = *(const uint4*)(blk + it0*64 + gg*16);
const uint4 qhv = *(const uint4*)(blk + 128 + it0*32 + (gg & 1) * 16);

r43's pre-expand-B variant swapped W for W_exp, deleted the recomb, and kept this stride line — the kernel then paged through the dense plane at 224/row, 224/sb. That the expand_q6_elem control arm passed is precisely because its data source W is the padded layout: expression and data matched.

The "before" that W_exp replaces — the in-kernel recomb (current tree expand_q6_elem, introduced in r38 and still the EXP=false path's core after r41's widening):

__device__ __forceinline__ int expand_q6_elem(const uint8_t* ql, const uint8_t* qh, int elem) {
    int m  = elem & 31;
    int it = elem >> 7;
    int n  = elem & 127;
    int ql_idx   = it * 64 + (n & 63);
    int ql_shift = (n >> 6) * 4;          // 0 or 4 (low/high nibble)
    int qh_idx   = it * 32 + m;
    int qh_shift = ((n >> 5) & 3) * 2;    // 0,2,4,6 (2-bit field)
    int v = ((ql[ql_idx] >> ql_shift) & 0x0F)
          | (((qh[qh_idx] >> qh_shift) & 0x03) << 4);
    return v - 32;
}

What pre-expansion does is move this bit arithmetic (plus its ql/qh loads) into expand_q6k_dense, run once at registration; the kernel side's B staging goes from "load→recomb→STS" to "pure copy". The three code blocks above read together are the complete before/after: expand_q6_elem (the work deleted) → the wrong/correct src expression (where the mismatch lives) → the host expander (the work's new home).

The correct dense indexing (current tree EXP=true branch, the landed form of r44's fix):

// EXP=true: W_exp is the dense centered-int8 plane produced by expand_q6k_dense
// (od x id, row stride = id, super-block stride = 256).
const int sb     = (kt * KDR) >> 3;            // which super-block this kt window lands in
const int cbase  = (kt * KDR) & 7;             // 32-element chunk offset within the sb
const int nc     = (KDR * 32) / 16;            // 16B chunks per row
for (int g = threadIdx.x; g < MMQ_NBJ * nc; g += blockDim.x) {
    const int jj = g / nc, cc = g % nc;
    const int j = j0 + jj;
    const bool full = (j < od) && (sb < nsb);
    const uint8_t* src = W_exp + (size_t)j * id     // ← row stride = id (dense)
        + (size_t)sb * 256                          // ← sb stride = 256 (dense)
        + (size_t)(cbase * 32 + cc * 16);           // ← offset within the sb (same in both layouts)
    // 16B alignment: id is a multiple of 256 ⇒ j*id and sb*256 both preserve it
}

Read in contrast: the two segments' only structural difference is nsb·bstride → id and bstride → 256. The entirety of r44's "root cause" is these two substitutions — after them, the in-sb offset cbase·32 + cc·16 needs no change, because it describes the element-space part common to both layouts.

The host-side expander (current tree src/cuda.rs, the r53-landed version; the expander logic in r44's experiment was isomorphic) — its doc-comment also documents the dense output layout, exactly the contract the consumer must match:

#![allow(unused)]
fn main() {
/// host mirror of the device `expand_q6_elem` (P6 r44 / MMQ-analysis
/// §11.24). Output: `od * id` bytes, `out[j * id + sb * 256 + e]` =
/// super-block element e of row j — the exact tile the kernel's staging
/// used to recomb. Requires `id % 256 == 0` (the NB-BT launch gate).
/// Two output elements per (ql, qh) byte pair: e = it*128+r and
/// e = it*128+r+64 share ql[it*64+r] (nibble shifts 0/4) and
/// qh[it*32 + (r&31)] (2-bit-field shifts 2*(r>>5) / 2*((r>>5)+2)).
pub fn expand_q6k_dense(padded: &[u8], od: usize, id: usize) -> Vec<u8> {
    const Q6KB: usize = 210;    // block_q6_K content bytes
    const Q6KPB: usize = 224;   // padded block stride (raw W side)
    let nbe = id / 256;
    let row_len = nbe * Q6KPB;                      // raw-side row stride = nsb*224
    let mut out = vec![0u8; od * id];               // dense side: od*id, row stride = id
    for j in 0..od {
        let prow = &padded[j * row_len..(j + 1) * row_len];      // 224-stride read
        let orow = &mut out[j * id..(j + 1) * id];               // id-stride write
        for sb in 0..nbe {
            let blk = &prow[sb * Q6KPB..sb * Q6KPB + Q6KB];
            let (ql, qh) = blk.split_at(128);
            let obase = &mut orow[sb * 256..sb * 256 + 256];     // 256-stride write
            for it in 0..2usize {
                for r in 0..64usize {
                    let qlb = ql[it * 64 + r];
                    let qhb = qh[it * 32 + (r & 31)];
                    let s0 = (r >> 5) * 2;
                    let e0 = it * 128 + r;
                    obase[e0] = ((qlb & 0xF) | (((qhb >> s0) & 3) << 4)).wrapping_sub(32);
                    obase[e0 + 64] =
                        (((qlb >> 4) & 0xF) | (((qhb >> (s0 + 4)) & 3) << 4)).wrapping_sub(32);
                }
            }
        }
    }
    out
}
}

Note that the host expander internally touches both layouts at once (the 224 read of row_len, the 256 write of orow). Transform code where two stride sets coexist is exactly the soil in which mismatches get written: the consumer copies the wrong side, with no compile-time or runtime signal.

The assembly relationship of the three segments: expand_q6k_dense (host, once at registration) produces W_exp → the kernel's EXP=true branch does a pure copy with dense indexing (r44's fix site; from r53 swapped to gemm_cp16 with src-size zero-fill of the tail rows beyond od) → expand_q6_elem keeps serving raw W only when EXP=false (W_exp absent/fallback). One dispatch chain with two data sources and two address contracts — which is also why the launch path needs an explicit label distinguishing them (r54's three-state label), making "which contract is in force" observable at runtime.

3.3 Pitfalls

  1. The readback gate is a content gate, not an offset gate. The perfect 0/17,920 readback granted strong confidence that "the plane is correct", which pushed the investigation toward the wrong direction of kernel-side races/aliasing. A pre-transformed plane's correctness gate must include consumer-side parity.
  2. Bit-identical cross-variant diffs are the fingerprint of shared address arithmetic. The failing variants producing the same diff (448 @ 554) = ruling out all variant-private mechanisms (staging order, barriers, register allocation), contracting the suspicion to the shared address expression; stacked with the control arm ("in-kernel expansion of raw W passes"), the only surviving explanation is expression-layout mismatch. This reasoning pattern is more transferable than the fix itself.
  3. The denominator effect on share-class metrics. After cycles −10.9%, the long_scoreboard share went 34.9% → 57.1%, which looks "worse" at a glance. Reading ncu shares requires the absolute quantities alongside: Warp-Cycles/Issued-Inst 11.70 → 17.84 (per-instruction stall-wait lengthening) coexisting with falling elapsed cycles is the complete signature of "latency changing form"; reading only the share yields the opposite conclusion.
  4. The 80 regs / 28 B spill budget check cannot be skipped. After address-arithmetic changes, register allocation is globally re-shuffled — r40's +13% depends on 3 blocks/SM, so any address-expression change requires re-reading the ptxas output to confirm the budget held.
  5. Transform code where two layouts coexist is a hotbed of stride confusion. The host expander writes both row_len (224) and orow (256); the consumer copying the wrong side gets no signal. The defense is writing the layout contract as a doc-comment (that is exactly what the current tree's function comment does) and making the consumer's index expressions match the comment item by item.

4. Verification

GateReadingWhat it defends against
parity cuda_prefill_mmq (GPU vs CPU reference)before: diff 448 @ index 554 (failing variants bit-identical); after: 1/0expression-layout mismatch — the direct evidence for this doc's root-cause verdict: one stride substitution eliminating all differences establishes the single cause
ptxas registers/spill80 regs / 28 B spillfixing the address while losing the third resident block (r40's +13% depends on 3 blocks/SM)
ncu base-vs-fix elapsed cycles1,390,776 → 1,239,847 (−10.9%)rules out "the mechanism never took effect" — the recomb ALU and its samples really did vanish
ncu Warp-Cycles/Issued-Inst + long_scoreboard share11.70 → 17.84; 34.9% → 57.1%the self-consistency check of the mechanism explanation: latency changing form + the denominator effect
whole-prefill interleaved 3× (A/B within one binary)2595.9 → 2584.9 (−0.42%)machine drift faking a trend; −0.42% is within the noise band, the verdict "wall-neutral"

The gates' ordering is itself a discipline: mechanism gates first (parity/ptxas/SASS), then performance gates (ncu/wall clock). Reversed — seeing wall-neutrality first and reverting — the root-cause fix would be discarded as "a useless change", and r53's bundle would lose its first component; celebrating on seeing kernel −10.9% first — the wall-neutral truth would be buried under mechanism excitement. Both directions have been erred before; this ordering table is the vaccine.

5. Results

Layerbefore (r43's failure state / r41 baseline)after (r44's corrected state)Verdict
paritydiff 448 @ index 554 (same diff in failing variants)cuda_prefill_mmq 1/0green
ptxas80 regs / 4 B spill (r41 state)80 regs / 28 B spill3-block budget kept
kernel elapsed cycles1,390,7761,239,847 (−10.9%)mechanism took effect
Warp-Cycles/Issued-Inst11.7017.84latency changed form
long_scoreboard share34.9%57.1% (absolute volume roughly unchanged)denominator effect
whole-prefill (interleaved 3×)2595.9 tok/s2584.9 (−0.42%)wall-neutral

Measurement protocol per §0 convention: all numbers come from same-binary interleaved A/B medians within one session window — this window's 2595.9/2584.9 are local readings of the "post-r41 ~2600 tier"; cross-session absolute values are not comparable (machine state drifts); the kernel-level comparison comes from ncu sampling of matched-nt at the same shape. vs-llama stays at 1.27× — the wall does not move, so the ratio does not move.

Veto mechanism (why reverted): r41's +30.7% had already pulled the q6_K GEMM off the wall's bottleneck seat (at r37 it was still 51.2% of wall). After that, anything improved inside the q6_K kernel — mechanism confirmed, parity all green, kernel −10.9% — no longer touches the critical path. The recomb the pre-expansion deleted merely converted "B's global→smem latency" from ALU-consumption form into STS-copy form, with an equivalent L1TEX exposure. Per the campaign rules (+1.5% bar, REVERTED keeps cmp-match HEAD), r44 reverted; the root cause and the "wall-neutral" verdict entered the record.

Worth emphasizing what the revert preserved: what was reverted is the diff, not the knowledge. Three things entered the record and were cited directly by later rounds — (1) the one-line root cause (the dense indexing formula), quoted verbatim in r53's bundle comments; (2) the mechanism explanation "the same physical latency changes instruction form to hide", which became standard practice for reading ncu share-class metrics; (3) the convergence evidence chain that "the q6_K kernel is no longer the wall" (r44 + r45, two independent mechanism lines), without which r47's fresh wall decomposition would not have known to measure FA and prepass instead of continuing to grind the GEMM.

Under what future conditions a retry is worthwhile: when the q6_K GEMM becomes part of the wall again, or when pre-expansion can be packaged with a "hide the wait" mechanism (cp.async). The condition cashed in precisely at r53: the EXP=true branch (dense indexing + cp.async pure copy + src-size zero-fill of tail rows) merged with r45's group-count pipeline into one basket — B staging became a pure copy with "no recomb ALU, no register round trips, no ql/qh reads", the latency handed to the asynchronous units. Results: ffn_down kernel 16.06 → 12.76 ms (−20.5%, r44's −10.9% and r45's −10.2% approximately adding), attn_v −15.9%, whole-prefill 3024.7 → 3176.9 (+5.03%). r44's work deletion and r45's wait deletion stacked because they are mechanically orthogonal (one deletes ALU, one hides latency), and the bundle's timing let the deletions touch the wall again. r54 then gave this 1.52 GB W_exp plane the MINFER_MMQ_Q6K_EXP=0 three-state opt-out (default on / off-plane taking the byte-identical r41 fallback / a fallback label distinguishing intentional from accidental). r44 therefore belongs to this directory's most important category: code correct, mechanism confirmed, no wall-clock value solo — not dead, WAITING.

The cash-in ledger of the r44/r45 "WAITING" pair (two same-day vetoes, both flipped positive two days later inside one basket):

MechanismSolo reading (r44/r45, both REVERTED)Bundle readingCashed in
Pre-expanded B (deletes the recomb ALU, dense indexing)kernel −10.9%, wall −0.42%ffn_down kernel −20.5% (approximately adding with r45)r53
cp.async staging (hides the global→smem wait)kernel −10.2%, wall −0.34%whole-prefill +5.03% (3024.7 → 3176.9)r53 (B side) + r56 (A side, another +2.35%)

6. Lessons

  1. When two layouts coexist in a data transform, the consumer's address expression must be checked against the data source's actual layout — the bstride/224 vs 256 confusion has no compile-time or runtime signal; only consumer-side parity can catch it.
  2. Readback verifies content, not offsets; the complete gate for a pre-transformed plane = producer readback + consumer parity, neither dispensable.
  3. A bit-identical failure across variants is the fingerprint of shared address arithmetic — contract to the shared part before acting; the control experiment (in-kernel expansion of raw W passing) is the key contracting step.
  4. On the converged line, kernel-level wins cannot reach the wall: mechanism confirmed ≠ wall-clock value; a "correct but wall-neutral" change should be recorded with its mechanism and annotated with retry conditions, not discarded (r53 cashed in r44, r56 cashed in r45 — "not dead, WAITING" is a discipline this campaign has verified repeatedly).

← 46-r43-pc-sampling-attribution · Index · 48-r45-cpasync-q6k-a-staging →

48 · r45 — cp.async for the q6_K A-side staging: mechanism confirmed, wall-neutral (REVERTED)

Result: the A-side bulk LDG→STS replaced with explicit PTX cp.async.cg.shared.global (16 B, group-count pipeline): kernel Duration 659,680 → 592,448 ns (−10.2%), long_scoreboard absolute volume 4.38 → 3.57 cy (−18.4%), Compute(SM) 37.4 → 42.1%, registers 80 / LOCAL 0 (r41's 4 B spill vanished along the way) — but whole-prefill 2604.5 → 2595.6 tok/s (−0.34%, noise, below the +1.5% bar) → REVERTED. The q6_K line declares convergence: after r41 this GEMM is no longer the prefill bottleneck, and a faster kernel cannot reach the wall. SASS forensics trap: cp.async in SASS is LDGSTS.E.BYPASS.128, not the literal CP.ASYNC — the gate must grep LDGSTS. Commit: 9825ffd. Date: 2026-09-05.

Provenance note: r45's code change was reverted after measurement (cmp-match HEAD r41), and 9825ffd is a docs-only commit whose diff contains no code. The excerpts here come from the current tree's src/cuda_kernels.cu: gemm_cp16 and the A-side cp.async staging are the final landed form of r45's mechanism (r53 packaged the B side, r56 installed the same mechanism back on the A side, with comments explicitly signing "r45's mechanism"), and the group-count pipeline main loop is unchanged word for word.

1. Background — where things stood

r44 had just dissolved the W_exp stride mismatch while also delivering a costlier conclusion: pre-expanding B (deleting the recomb ALU) went parity all-green and kernel −10.9%, but whole-prefill −0.42% — after r41 the q6_K GEMM is no longer the wall. MMQ-analysis §11.24's closing wrote that pre-expansion is a "correctness fix for the W_exp addressing but a dead end for wall perf", and that the lever that can actually hide B's global→smem staging latency inside compute is cp.async (llama.cpp's structure). r45's implementation cashes in that lever. Where does its absence stall? Of the attribution table's two biggest items (recomb 45% + A-staging 28%), r44 addressed only the former; the A side's 28% is inherent to the synchronous LDG→STS chain — as long as staging is "load into registers, then write smem", this latency must queue in the warp's instruction stream every tile. And whether the q6_K line still has a second half (the stacked gains of bundle form) depends entirely on whether both mechanisms can be proven effective — r45 is the other half of that proof.

1.1 Why the A side first, and the timeline

Three reasons to start with the A side. First, the A side is the pure bulk-copy candidate: after r34 the BT route moved activation quantization into the prepass, so the qa8/sda the kernel reads are pre-transposed q8 planes and staging is a straight per-16 B transport with no per-element transform — cp.async is a natural drop-in. The B side was still stuck on the recomb (ql+qh unpack); its pure-copy form needed W_exp's fix (r44's dense indexing) as a premise. Second, in r43's PC-sampling attribution the A-side staging STS.128 held 28% of stall samples, the second-biggest source, worth a shot. Third, this doubles as a cheap convergence test: if hiding the attribution's second-biggest latency speeds the kernel ~10% and the wall still does not move, then "the q6_K line has converged" is not a conjecture but an empirical fact, and the budget should pivot wholesale.

Background numbers: after r41 whole-prefill ~2600 tok/s (vs-llama 1.27×), the q6_K kernel matched-nt ~0.59 ms, and the campaign bar is +1.5% (whole-prefill). r44 had already set an iron rule: kernel-level wins do not automatically convert to wall clock on the converged line. r45 would either become the counterexample or nail the rule down.

The timeline is itself a signal: r44 and r45 were two consecutive rounds of the same evening (docs commits at 21:38 and 21:59); half an hour after r44's §11.24 closing wrote "cp.async is the real lever", r45's implementation began — attribution, root cause, and mechanism all completed in one day, the fastest hypothesis→verification loop of the entire campaign.

2. Principle — the GPU mechanism

2.1 The synchronous chain vs the asynchronous copy

The synchronous staging instruction sequence is LDG.128 (global → register) followed by STS.128 (register → smem). Between the two instructions sits the full L1TEX/DRAM round-trip latency: the warp issues the LDG and then stall-waits on the STS's data dependency (long_scoreboard), the register serving only as the latency's display stand. r44 already proved: swapping the latency's consumer from the recomb ALU to the STS copy leaves the exposure equivalent.

A small account: on the synchronous chain, every 16 B moved costs two instructions (LDG+STS) plus one stall-wait; cp.async is one instruction, zero stall-waits. Halving the instruction count is only the secondary gain (r25 proved instruction-count cuts are wall-inert at 1 block/SM); the main gain is the stall-wait leaving the warp's critical path — not one byte moved changed, what changed is what the warp waits for. This explains why r45's kernel −10.2% is worth more than r25-class pure instruction cuts (~0%): it moves the wait structure, not the instruction total.

cp.async.cg.shared.global [dst], [src], 16 is a different physical path: Ampere's asynchronous copy unit (SASS-level LDGSTS) writes the 16 bytes from global directly into smem, bypassing the register file. The issuing warp does not wait for the data to land; completion is accounted per commit-group — cp.async.commit_group bundles all previously issued copies into one group, and cp.async.wait_group N waits until at most N groups are outstanding. The interleave of "transport" and "compute" thus changes from an accident of instruction scheduling into an explicit contract: issue copies → do other compute → wait on groups → barrier → consume.

Three qualifier/semantics details deserve their own listing (r53/r56 both reused this primitive set):

  • .cg (cache-global): the copy goes through L2, bypassing L1. Staging data is consumed once per byte (once it enters the mma it is done), so caching in L1 is pure waste; .ca (cache-all) is for repeatedly-read data.
  • The src-size 4th operand: in cp.async.cg [dst],[src],16,sz, sz ∈ [0,16] controls the bytes actually moved, with the remainder zero-filled. r45's A-side planes are always full (the prepass zero-fills the padded rows), so it passes full=true's 16; r53's B-side W_exp tail rows (beyond od) rely entirely on sz=0 zero-fill — the same instruction serves both "pure copy" and "copy with zero-fill" semantics.
  • Groups are ordered: commit_group completes in issue order (in-order group completion), which is the premise for reasoning precisely that "at wait_group 1 exactly the previous group has landed and the newest is still in flight".

One alternatives-check on "why group counting": among the synchronization granularities available inside a kernel, __syncthreads is a block-wide full stop (back to synchronous semantics), spin-on-flag needs threadfence + atomic polling (one extra global round trip per tile, with an expensive correctness argument), and CUDA events cannot be used inside a kernel — commit/wait_group is the only in-kernel asynchronous wait primitive that is both cheap and precisely reasonable about. Group counting's "coarseness" (you cannot wait for one specific group, only for "N groups remaining") is exactly digested by double buffering's "alternate consumption" structure: the only wait shape ever needed is "the previous group landed, the newest is in flight".

2.2 The arithmetic of the group-count pipeline

The kernel already has r39's double buffering (KDR=2, two staging buffers per kt) and r40's 3 blocks/SM. r45 turns the wait into group counting:

  • Each main-loop round first issues the copies for kt+1 into buffer buf^1 and commits;
  • wait_group 1 — allows the newest group (kt+1's) to still be in flight, requiring only that the previous group (kt's) has landed, exactly covering the buffer this round consumes;
  • The last tile has no "next group" to lean on, so wait_group 0 drains all outstanding groups;
  • One __syncthreads() after the wait: cp.async completion visibility is per-thread, and smem writes must pass a block-level barrier to be visible to other threads in the block.

Using the first few rounds to unroll the pipeline (group numbers in commit order; the pre-loop already staged tile0→buf0 and committed group 1):

prologue:  stage(tile0 → buf0) … commit group1 … barrier
kt=0,buf0: stage(tile1 → buf1)=group2 → wait1 ⇒ group1 landed (group2 in flight) → barrier → compute buf0
kt=1,buf1: stage(tile2 → buf0)=group3 → wait1 ⇒ group2 landed (group3 in flight) → barrier → compute buf1
kt=2,buf0: stage(tile3 → buf1)=group4 → wait1 ⇒ group3 landed (group4 in flight) → barrier → compute buf0
last:      no next tile               → wait0 ⇒ all landed              → barrier → compute the last buffer

There is exactly one invariant: when wait1 returns, the copy group of the buffer this round consumes has necessarily landed, and the groups still in flight write only the other buffer. In-order group completion guarantees this correspondence without any timing assumptions — precisely why group counting is cheaper than "drain every round" and safer than "hand-rolled events".

The latency is thus pushed into the compute window: kt's copies are already in flight while kt−1's mma chain executes, and the warp's pre-consumption stall-wait shrinks to the group's tail residual. r20's split-phase lesson (long_scoreboard's carrier is the LDG→STS chain) finally has a hardware-level solution in this structure — r20 relied on splitting load and consume into two phases to let ptxas overlap them; cp.async simply removes "transport" from the warp's instruction stream.

2.3 The alignment premise and the A planes' luck

cp.async's 16 B form requires both source and destination 16B-aligned. The A side's two planes satisfy this naturally, and the geometry reads straight out of the staging loop: the qa8 segment moves KDR·NBI·32 bytes per (tile, kt) window (32 bytes of q8 per chunk per row), the sda segment moves KDR·NBI·4 bytes (4 bytes of scale pair per chunk per row), and both have whole numbers of 16 B chunks as the per-thread copy granularity — no straddling partial chunks, so full=true always holds (the prepass-produced planes are laid out on a 16 B grid with padded rows zero-filled). The B side could not do this yet (the recomb is per-element ALU with no contiguous-16 B copy semantics), which is the physical basis for "A first, B later"; B's pure-copy form waited for r53 to complete it with expand_q6k_dense's dense plane (that path also needs cp.async's src-size zero-fill semantics for tail rows beyond od — see §2.1's qualifier list).

2.4 The relationship with r44: two orthogonal mechanisms

Put r44/r45 on one mechanism map and they address two different facets of the same physical latency: r44 deletes extra work on the compute side (the recomb ALU), r45 deletes exposure time on the wait side (the LDG→STS stall-wait). The mechanisms neither depend on nor cancel each other — the embryo of r53's later basket thesis: "mechanisms that overlap in traffic but not in mechanism compose". "Traffic overlap" (both reduce the same L1TEX round trips) with "mechanism non-overlap" (one deletes ALU, one hides waits) can stack because deleting work reduces the number of instructions to issue while hiding waits reduces the cycles instructions stall — they act on the instruction stream's numerator and denominator respectively, and tightening both does not cancel out. The counterexample is applying the same mechanism twice (r42's further dsc widening had no gain): for a wait that no longer exists, a second "hide" has no object. At the time no one could prove wall-clock composability (each was wall-neutral solo), but r45's design already consciously preserved composability with r44: swapping the A side to cp.async does not touch the B side's structure, so when the B side later becomes a pure copy (awaiting W_exp) the two naturally fit together.

3. Implementation

3.1 Design choices: A-side solo + group-count waits, budget untouched

The change is deliberately narrow: only the qa8/sda two segments of the A side inside the RAW_STAGE_Q6K_BT macro and the main loop's wait structure; the B-side recomb (r41 form), the KDR=2 double buffering, and the 3-block residency budget are all untouched. The wait uses group counting (wait1 in-loop, wait0 in the last round) rather than wait_group 0 every round — the latter regresses to synchronous semantics and the pipeline would be for nothing. r39-era buffer WAR barrier (consumers must finish reading before overwrite) is kept: cp.async only changes "which execution unit writes smem" and grants no exemption from write-after-read hazards.

Since the code was ultimately reverted, the diff's coverage is recorded here as narration (the mechanism skeleton is visible verbatim in the current tree's r53/r56 landed versions, see §3.2): (1) four new device functions gemm_cp16/gemm_cp_commit/gemm_cp_wait1/gemm_cp_wait0 (inside the Ampere guard); (2) the qa8/sda segments of the RAW_STAGE_Q6K_BT macro changed from reinterpret_cast<const uint4*> loads + STS.128 to per-16 B-chunk gemm_cp16 (B side untouched); (3) the main loop changed the double-buffer switch point from "stage, then synchronously wait" to "stage, then group-count wait". The rejected alternatives: wait_group 0 every round (back to synchronous); cp.async on the B side too (impossible — the recomb is per-element ALU with no copyable contiguous-16 B semantics); the __pipeline_memcpy_async intrinsic (a pitfall, see §3.3#1).

3.2 Key code

The explicit PTX copy primitives (current tree src/cuda_kernels.cu 4483–4491, introduced in r45 and in use since):

__device__ __forceinline__ void gemm_cp16(__half* smem_dst, const __half* gsrc, bool full) {
    unsigned d = (unsigned)__cvta_generic_to_shared(smem_dst);
    int sz = full ? 16 : 0; // src-size 0 => zero-fill the 16B chunk
    asm volatile("cp.async.cg.shared.global [%0], [%1], 16, %2;\n" ::"r"(d),
                 "l"(gsrc), "r"(sz));
}
__device__ __forceinline__ void gemm_cp_commit() { asm volatile("cp.async.commit_group;\n"); }
__device__ __forceinline__ void gemm_cp_wait1() { asm volatile("cp.async.wait_group 1;\n"); }
__device__ __forceinline__ void gemm_cp_wait0() { asm volatile("cp.async.wait_group 0;\n"); }

Three details: __cvta_generic_to_shared converts a generic pointer into a 32-bit shared-window address (required by PTX's shared state space); the 4th operand is the src-size qualifier, and at sz=0 the hardware zero-fills the whole 16 B chunk (r53 uses this for tail rows beyond od); the .cg qualifier goes through L2 bypassing L1 (staging data is consumed once, so L1 caching is meaningless).

The A-side staging loop (current tree, inside the RAW_STAGE macro; the comment signs r45):

/* ---- A: r56 cp.async bulk staging of the pre-transposed qa8/sda --*/
/* (r45's mechanism on top of r53: the sync LDG->STS exposed its      */
/*  global latency at the top of every staging phase; cp.async hands  */
/*  it to the async unit and the group wait below hides it under the  */
/*  previous tile's compute. Bytes identical - the plane is always    */
/*  full: the prepass zero-fills the padded rows.)                    */
{
    const size_t qbase = ((size_t)blockIdx.x * nchunk + (size_t)(kt) * KDR) * MMQ_A_QASZ;
    for (int off = threadIdx.x; off < (KDR * MMQ_NBI * 32) / 16; off += blockDim.x)
        gemm_cp16((__half*)(void*)(qa8b + (size_t)off * 16),
                  (const __half*)(const void*)(qa8g + qbase + (size_t)off * 16),
                  true);
    const size_t sbase = ((size_t)blockIdx.x * nchunk + (size_t)(kt) * KDR) * MMQ_A_SDASZ;
    for (int off = threadIdx.x; off < (KDR * MMQ_NBI * 4) / 16; off += blockDim.x)
        gemm_cp16((__half*)(void*)(sdaqb + (size_t)off * 16),
                  (const __half*)(const void*)(sdag + sbase + (size_t)off * 16),
                  true);
}

The group-count pipeline main loop (current tree 6905–6924):

RAW_STAGE_Q6K_BT(0, 0);          // tile 0 synchronously preloaded (group already committed)
__syncthreads();

int buf = 0;
for (int kt = 0; kt < nktile; ++kt, buf ^= 1) {
    // stage kt+1 into the OTHER buffer while reading buffer buf
    if (kt + 1 < nktile) {
        RAW_STAGE_Q6K_BT(kt + 1, buf ^ 1);
        // two groups pending (kt's and kt+1's); wait until only kt+1's
        // remains — group(kt) has landed, buf^1's copies stay in flight
        gemm_cp_wait1();
    } else {
        gemm_cp_wait0();  // last tile: drain every outstanding group
    }
    __syncthreads();  // cross-thread visibility of the kt buffer's async copies
    // ... kt's mma compute consumes buffer buf; buffer WAR is guaranteed by the existing barrier ...
}

The only difference from r45's shape at the time: the B side was still r41's recomb (the EXP=false branch), dsc was still a scalar read, and the commit sat at the end of RAW_STAGE — the mechanism skeleton (group counting + the visibility barrier + WAR preservation) matches the above word for word.

The assembly relationship of the three excerpts: gemm_cp16/commit/wait are the primitive layer (16 B semantics + group accounting); the RAW_STAGE macro's A segment is the issue layer of "whole numbers of chunks per thread" — it returns as soon as it issues, never waiting for data; the main loop is the scheduling layer — using group counts to express "wait for whose copies before consuming whom". All three layers are independently testable (SASS counts, byte equivalence inside the macro, the loop invariant), and r45's verification gates were built along exactly these three layers.

3.3 Pitfalls

  1. __pipeline_memcpy_async silently falls back. The first version used the CUDA intrinsic: the compiler could not prove a generic uint8_t* (byte pointer) is 16B-aligned, so it did not emit cp.async and silently fell back to LDG+STS — it compiled, was semantically correct, and performed identically, the only tell being 0 copy instructions in the SASS. A textbook case of "intrinsic ≠ hardware mechanism": a mechanism must be accepted via SASS, not via compiling. Only the explicit PTX (copy size 16 as a compile-time constant) produced LDGSTS.
  2. The SASS mnemonic trap: cp.async is LDGSTS.E.BYPASS.128. Grep for CP.ASYNC and you conclude wrongly that nothing was issued. The gate is always cuobjdump -sass | grep LDGSTS (24 hits in the <2> instantiation). This entered the 77-verification-methodology forensics checklist.
  3. An unexpected dividend in the register account. Once cp.async removed the register way-station, REG 80 / LOCAL 0 — the 4 B spill that r40-era's "10 hand-trim variants could not squeeze out" was gone. Register pressure determined by the staging structure is harder than any register-level micro-tuning.
  4. The visibility barrier can be neither omitted nor duplicated. The __syncthreads() after each round's wait is a cross-thread visibility requirement (cp.async completion is per-thread semantics); the barrier the buffer WAR depends on is a different one — the two barriers have different duties, and merging or deleting either produces intermittent dirty data.
  5. The N in wait_group N is "the number of groups allowed to remain unfinished", not "wait for the first N groups". wait_group 1 = wait until at most 1 group is outstanding (i.e. everything before the second-to-last has landed); reading the semantics backwards turns the pipeline half-synchronous or leaves dangling reads. The correctness argument for group-count code must land on §2.2's trace table, not on intuition.

4. Verification

GateReadingWhat it defends against
SASS gate: cuobjdump -sass grep LDGSTS24 hits in the <2> instantiation"the intrinsic silently fell back" — confirms the physical mechanism is really running rather than the compiler downgrading; this doc's most important gate
parity cuda_prefill_mmq (GPU vs CPU reference)1/0data misalignment introduced by swapping the staging mechanism (the copied bytes are bit-equivalent to LDG+STS: the planes are always full, no partial chunks)
greedy-32 byte comparisonbyte-identicalany drift on the numeric path propagating into sampling
suite166/0/3the whole-engine regression net
ptxas registers/spillREG 80 / LOCAL 0the 3-block residency budget (r40's +13% depends on it), plus checking cp.async's expected register dividend
ncu base-vs-r45 (interleaved, same window)Duration −10.2%, longsb absolute −18.4%, Compute(SM) 37.42→42.13%three-line self-consistency at the mechanism layer: faster, fewer stall-waits, higher compute share — the exclusion of "just measurement noise"
whole-prefill interleaved 3×2604.5 → 2595.6 (−0.34%)machine drift faking a trend; the verdict "wall-neutral, bar not met"

Together, ncu's mechanism-layer readings are the complete signature of cp.async working: Elapsed Cycles 1,408,834 → 1,261,636 (−10.4%) and Warp-Cycles/Issued-Inst 11.55 → 10.60 improving in the same direction — the contrast with r44's "cycles fall but per-instruction stall rises" (latency changing form): cp.async really did hide the wait inside compute, not just reshuffle the instruction mix.

5. Results

Layerbefore (r41/r44 baseline)after (r45)Verdict
SASS0 cp.async (intrinsic fallback)LDGSTS ×24mechanism confirmed
kernel Duration659,680 ns592,448 ns (−10.2%)mechanism took effect
long_scoreboard (absolute)4.38 cy3.57 cy (−18.4%)the wait really was hidden
Compute(SM)37.42%42.13%compute share rose
registers/spill80 / 4 B80 / 0 Bunexpected dividend
whole-prefill (interleaved 3×)2604.5 tok/s2595.6 (−0.34%)wall-neutral, bar not met

Measurement protocol as in r44: same session window, the same pair of binaries (base / r45) interleaved 3 rounds of whole-prefill taking the median; ncu is matched-nt sampling at the same shape. This window's ~2600 tok/s absolute value is meaningful only within this window. vs-llama stays at 1.27×.

Veto mechanism (why reverted): the +1.5% bar was not met, and −0.34% has no directionality even. Merged with r44 into a complete convergence proof: r44 deleted work (the recomb ALU, −10.9% cycles) and the wall did not move; r45 hid waits (cp.async, −10.2% duration, −18% longsb) and the wall did not move either — the two biggest stall sources (45% + 28%) were each handled by a correct mechanism, and whole-prefill did not move an inch. Only one conclusion remains: after r41 the q6_K GEMM is not the wall, and the remaining ~3.0× vs-llama gap of the matched-nt kernel (against 57.8 µs/GMAC) no longer drives whole-prefill. The q6_K line declares CONVERGED; further kernel-level tuning (cp.async B, split-phase, pre-expansion) is expected wall-neutral in solo form. Reverted by discipline (cmp-match HEAD).

The immediate consequence of the convergence declaration is worth recording: it is not a negative "we're done" but an authorization for budget migration. r46 (23:00 that day) turned to audit FA, and r47 redid the converged-domain wall decomposition the next day (q6_K 1094.7 → 196.4 ms, falling from 51.2% to 15.8%; FA rose to the #1 structural residual) — both rounds' project proposals cite r44/r45's dual-mechanism convergence evidence directly. Had these two rounds' wall-neutrality been vaguely recorded as "the attempts didn't work", the later budget allocation would have had no basis; precisely because the veto mechanism spelled out "deleting work didn't work + hiding waits didn't work ⇒ the slice is not the wall", the conclusion could be safely extrapolated.

Under what future conditions a retry is worthwhile: when the A-side wait becomes the critical path again. r56 cashed this in precisely: after r53's package (W_exp pure copy + cp.async B staging, +5.03%) landed, the wall's composition moved — "A-side staging STS + dsc consumption" became the q6_K line's remaining bulk, and r56 installed r45's mechanism back on the A side verbatim. r56's readings:

Layerafter r53 → r56Verdict
ffn_down kernel (ncu)12.76 → 12.01 ms (−5.9%)the A-side cp.async took effect
attn_v kernel−4.2%same
whole-prefill3138.6 → 3212.5 (+2.35%, distributions separated)r45's mechanism touched the wall for the first time
SASSboth instantiations (<2,true> and <2,false>) contain LDGSTSthe mechanism fully in place

r45's own words were "Not dead, WAITING", and r53/r56 proved the second half. Together with doc 47, this doc forms the campaign's most important pair of counter-intuitive samples: a mechanism's value is not a property of the mechanism; it is a property of the wall.

6. Lessons

  1. A mechanism can be correct, confirmed, and simultaneously worthless — wall-clock value depends on what remains on the critical path, not on how elegant the mechanism is; a vetoed mechanism must enter the record carrying its "when to retry" conditions.
  2. An intrinsic is not a mechanism: the lesson of __pipeline_memcpy_async silently falling back to LDG+STS is that any "I used hardware feature X" claim must carry SASS-level evidence.
  3. Grep SASS with the real mnemonics: cp.async is called LDGSTS.E.BYPASS.128 in SASS — get the search keyword wrong and every conclusion is wrong.
  4. Pin a convergence verdict with two independent mechanism lines: deleting work does not move the wall + hiding waits does not move the wall is far stronger than a single piece of evidence; only then is the post-convergence budget migration (q6_K → FA/prepass) justified.

← 47-r44-wexp-stride-mismatch · Index · 49-r46-fap1-fa-audit →

49 · r46 (FAP1) — FA audit + occupancy/bank-conflict levers: kernel −11% but wall-neutral (REVERTED)

Result: the audit of fa_prefill_f16kv overturned the latent assumption "FA is a scalar kernel" — it has long been wmma m16n16k16 + online softmax; the real problem is occupancy starvation (69.38 KB smem/block → 1 block/SM, 16.64% occ, ncu estimated speedup headroom 68.7%) plus the bank-conflicted S/P smem round trip (row stride ≡ 0 mod 32 banks, MIO scoreboard 36%). Levers: FA_TKV 64→32 (~43.8 KB → 2 blocks/SM) + S/P row padding (+8 f32) + incidentally fixing a hidden launcher smem over-request (3*FA_TQ is correct only when FA_TQ==FA_TKV). Result: FA kernel 5.16 → 4.58 ms (−11%), 2 blocks/SM, SM busy 40% — whole-prefill only +0.27% (< the +1.5% bar; cutting 16 ms from the 124.7 ms FA slice is noise-level) → REVERTED. Named its successor FAP2: register-resident softmax, deleting the S/P round trip entirely (r48 landed +5.6%). Commit: a186f51. Date: 2026-09-05.

Provenance note: r46's code change was reverted after measurement, and a186f51 is a docs-only commit. The r46-era FA kernel excerpts here (FA_TKV=64, the Sf/Pf smem round trip, the old launcher formula) come from the historical tree before r48 landed (a bounded line range of git show d38744d~1:src/cuda_kernels.cu) — the last time this code existed in the repo; the current-tree excerpts (the FA_TKV=32 define, the +8 padding comments) are the parts that survived r46/r48, comments unchanged verbatim.

1. Background — where things stood

1.1 A stale map

On the afternoon of 2026-09-05, the q6_K line had swallowed two wall-neutral results in a row: r44 (pre-expanding B deletes the recomb, kernel −10.9%, wall −0.42%) and r45 (cp.async hides the A-side wait, kernel −10.2%, wall −0.34%). Two independent mechanism lines proved the same thing: after r41 the q6_K GEMM is no longer the prefill wall's bottleneck. That is good news (a line converged) and bad news (where does the next shot go?). The r37-era wall decomposition (q6_K GEMM at 51.2%) was stale — continuing to allocate budget by the old map would spend the whole campaign on irrelevant slices.

FA (prefill flash-attention, fa_prefill_f16kv) was the biggest question mark hanging in the archive. Its timeline: 8n (cb66fca) landed the tiled FA taking attention from 176 ms/layer to 8.5 ms/layer; P5·0 (86ca78c) put P·V on the tensor cores (10.06 → 4.24 ms/layer); P5·3 (fc07c04) fixed the K/V staging's 8-way ldmatrix conflict. FA had not been touched since, and the r23-era reading was "FA's 2.5×/layer gap is structural (llama keeps 128-wide KV tiles)" — a judgment frozen in the f16-path era, and after the MMQ redesign (r28–r41) turned the whole wall's composition over, no one had re-measured FA's relative position.

r23's judgment also carried an unverified assumption worth auditing today: the comparison object back when FA was called "structurally slow" was llama.cpp's 128-wide KV tile geometry; but "to what extent our FA already uses tensor cores" had itself not been fully audited since P5·0. If the audit found large scalar stretches in the kernel, the FAP was a 20×-class rewrite opportunity; if it found an already all-tensor-core structure, the opportunity was elsewhere.

One easily overlooked piece of context on the timeline: by 2026-09-05 three rounds in a row had been reverted (r43 parity FAIL, r44 wall-neutral, r45 wall-neutral), all on q6_K. r46 was the day's first move leaving q6_K — executing exactly the budget migration r44/r45's convergence declaration authorized. From this moment the campaign's narrative switched from "the depth of one GEMM line" to "the sweep of the whole wall"; r47's whole-wall decomposition and the FA/prepass line of r48–r52 are all downstream of this switch.

1.2 FAP1's project shape: audit first

So FAP1's first step was not writing code but auditing: what is this kernel actually now? The audit also had to answer a second question: what share of the wall does the FA slice hold now, and is the lever's ceiling enough to clear the bar? The two answers together decide whether FAP1 "acts" or "books and moves on". In hindsight, this project order saved the session — of the audit's three findings, the first (the structure is innocent) closed an imagined rewrite, the third (the slice's ceiling) vetoed the day's lever, and the second (the mechanism list of occupancy + conflicts) was inherited verbatim by FAP2 two months later. Each finding paid for itself, but only together did they constitute a complete decision.

2. Principle — the GPU mechanism

2.0 The audit method: three evidence lines

The audit ran along three parallel lines, cross-checking each other:

  1. Structure line (read the code): read fa_prefill_f16kv section by section, annotating each compute stage's execution unit (wmma / ALU / smem round trip) — answering "what is it".
  2. Resource line (do the accounting): derive each block's byte count from the kernel's smem layout declarations, divide by the device's smem/warp budget to derive blocks/SM and the occupancy ceiling — answering "is it starved".
  3. Behavior line (ncu): occupancy counters, estimated speedup, stall composition (MIO scoreboard etc.) — answering "which class of resource is the bottleneck on".

The three lines converge in §2.1–2.3: structure innocent, resources starved, behavior consistent.

2.1 Audit finding one: structure innocent — wmma + online softmax are already there

The structure checklist confirmed by reading the kernel section by section (historical tree, FA_TKV=64 era):

  • QK^T: wmma::mma_sync (m16n16k16, f16 A/B, f32 accumulate), each warp holding several 16×16 accumulator fragments (fc[0]/fc[1] in the excerpt below cover one 32-column group), looping over hd in steps of 16 — the Q row blocks × K column blocks run entirely on the tensor cores.
  • online softmax: per-KV-tile m/l/alpha state (block-shared arrays msh/lsh/alpha, one entry per row) + exponential rescaling — the standard flash-attention shape, not a "one-shot whole-softmax".
  • P·V: P enters wmma as an f16 A-operand from Pf via load_matrix_sync, V is the B-operand — the second tensor-core stage.

The "FA is a scalar kernel" assumption is falsified — a rewrite-class opportunity does not exist. That conclusion alone paid for the audit: it shut down the "FAP = a big rewrite" fantasy and pointed the campaign at small, precise levers. The audit also supplied a reference coordinate: llama.cpp's fattn keeps 128-wide KV tiles (the source of r23's old judgment), while we run 64 — narrower tiles mean worse amortization of per-tile fixed overhead, a second structural disadvantage beyond §2.2's occupancy starvation, but one that cannot be fixed by "widen the tile" (the r24/P5·neg lesson: wider tiles are worse at 1 block/SM); it can only be discussed after occupancy is solved.

2.2 Audit finding two: occupancy starvation — the 69.38 KB arithmetic

The kernel's smem ledger (FA_TQ=64, FA_TKV=64, hd=128, staging row stride = hd+8 = 136 halves):

PlaneSizeBytes
Qs (q tile, f16)64×136×217,408 B
Ks (K tile, f16)64×136×217,408 B
Vs (V tile, f16)64×136×217,408 B
Sf/Pf (aliased reuse: S f32 → P f16)64×64×416,384 B
msh/lsh/alpha (online softmax state)3×64×4768 B
Total69,376 B ≈ 69.38 KB

69.38 KB/block → each SM fits only 1 block (2 blocks need 138.8 KB, over the budget ceiling); 256 threads = 8 warps, which against 48 warp slots is 16.64% occupancy. ncu's estimated-speedup reading: 68.7% — the SM spends most of its time without enough resident warps to fill the latency.

Why 16.7% is especially lethal for this kernel: FA's main loop is the serial chain "stage KV tile → QK^T → online softmax → P·V", and every segment of the chain has dependency stall-waits (staging waits on global, softmax waits on smem, mma waits on operands). Occupancy's job is to let other warps' instructions fill those stall holes. 1 block/SM means the holes can only be filled by the same block's 7 sibling warps, which share the same serial chain's rhythm (aligned to the same tile boundaries); 2 blocks/SM introduces a second block in a different phase, with a very different hole-filling probability. FA's KV tile loop runs dozens of iterations (nt/FA_TKV rounds), and each round's stall residual × dozens of rounds is the main source of ncu's 68.7% estimated headroom. For a kernel that passes dozens of KV tiles per layer, double buffering is the only source of overlap — and double buffering can only hide so much depth; occupancy is its missing amplifier.

2.3 Audit finding three: the S/P round trip's bank conflict — 256 B ≡ 0 mod 128 B

S/P makes a full round trip through shared memory: QK^T's wmma accumulator store_matrix_sync into Sf (f32) → softmax reads Sf, writes Pf (f16, aliasing the same smem) → P·V's load_matrix_sync reads back from Pf. The problem is the row stride: Sf's row stride = FA_TKV f32 = 64×4 = 256 B; Pf's row stride = FA_PSTR = FA_TKV×2 halves = 256 B. 32 banks × 4 B = 128 B per bank cycle, and 256 B ≡ 0 (mod 128 B) — every row starts on the same bank group.

Unroll one concrete access: P·V's load_matrix_sync(pa, &Pf[...], FA_PSTR) takes 8 rows × 16 halves at once, and every ldmatrix row lands on a start address ≡ 0 (mod 256 B) — i.e. all 8 rows hit the same bank group and the access serializes into 8 beats. The softmax's row rotation (one warp handling rr = warp, warp+8, warp+16... in turn) also re-calibrates to the same bank group on every row change. This is the same disease P5·3 fixed on K/V staging ("256 B rows ≡ 0 mod 32 banks = 8-way ldmatrix conflicts"; the fix then was +8 half row padding, making the stride 272 B ≡ 16 mod 128 B) — this time growing on S/P. Corroborating evidence: MIO scoreboard holds 36% of stalls — the smem round trip's queue pressure; a tax of "8 extra beats per round trip per tile × dozens of tiles per layer" lands squarely on the MIO pipe.

2.4 The lever arithmetic: FA_TKV 64→32 + row padding

  • Occupancy: after halving FA_TKV, Ks/Vs each fall from 17,408 B to 8,704 B and Sf/Pf from 16,384 B to 8,192 B: total Qs 17,408 + Ks 8,704 + Vs 8,704 + Sf/Pf 8,192 + state 768 = 43,776 B ≈ 43.8 KB → 2 blocks per SM (87.6 KB) → 16 warps / 48 slots = 33.3% occupancy, doubled.
  • Bank conflicts: S/P rows each padded +8 f32, the stride becoming 288 B ≡ 32 (mod 128 B) — cross-row accesses land on staggered bank groups.
  • Cost forecast: halving FA_TKV doubles the KV tile count per layer (nt/FA_TKV rounds), and every tile's softmax state updates (the three shared-memory arrays m/l/alpha read/written), masking, and tile-boundary synchronization all double — this cost was outweighed at the time by "occupancy doubled" (net −11%), and r50 (FA_TKV 32→16) later proved it to be this lever class's intrinsic tax rate: the occupancy gain is sublinear (each added block yields fewer new hole-filling opportunities) while the per-tile cost is linear (tile count × fixed cost), so the two curves must cross. Parked here for now; r50 measures the intersection.

The two smem ledgers, before and after the lever, side by side (hd=128, sstr=136 halves):

PlaneFA_TKV=64 (before)FA_TKV=32 (after)
Qs64×136×2 = 17,408 B17,408 B (unchanged)
Ks64×136×2 = 17,408 B32×136×2 = 8,704 B
Vs64×136×2 = 17,408 B32×136×2 = 8,704 B
Sf/Pf64×64×4 = 16,384 B64×32×4 = 8,192 B
msh/lsh/alpha768 B768 B
Total69,376 B → 1 block/SM43,776 B → 2 blocks/SM

Note Qs unchanged and Ks/Vs/Sf halved linearly — this is the occupancy lever's "subtraction" essence: it adds nothing, only presses the request under the threshold, unlocking the SM's second set of resources (the second block's warp slots).

  • Why the padding is +8 f32: the Sf rows (f32) and Pf rows (f16, aliasing the same smem) each gain 8 elements of their own width — the Sf row stride becomes (FA_TKV+8)×4 = 288 B ≡ 32 (mod 128 B), and Pf's likewise counted in halves. 288 and 128's greatest common divisor is 32, so cross-row accesses stagger by 8 banks — no longer the 0-offset full conflict, nor the exactly-halved (128 B) 2-way. +8 is the same dose P5·3 validated on staging (272 B ≡ 16 mod 128 B, the half version), not an arbitrary odd number.

2.5 The hidden killer: the launcher's smem over-request

The launcher's requested dynamic smem sets the occupancy ceiling. The old formula:

smem = 3 * FA_TQ * (hd + 8) * 2  +  FA_TQ * FA_TKV * 4  +  3 * FA_TQ * 4

The first term assumes all three planes Q/K/V have FA_TQ rows — true only when FA_TQ == FA_TKV (then 3×64 = 64+2×64, exactly equal). With FA_TKV changed to 32, the first term's real requirement is (FA_TQ + 2*FA_TKV)*(hd+8)*2 = 34,816 B, yet the formula still requests 52,224 B (17.4 KB too many, effectively counting Vs as a full FA_TQ-sized plane); adding Sf/Pf and the state, the total request is 61,184 B ≈ 61.2 KB → 2×61.2 = 122.4 KB still over the ceiling → still 1 block/SM. That is: without fixing this line, all of FA_TKV=32's occupancy gain silently zeroes out, while functionality, parity, and even kernel time can all look "normal" — occupancy goes from 1 block to 1 block, with no error whatsoever. This is the doc's most valuable pit (see §3.3).

The word "hidden" deserves unpacking too: this bug was correct during the years FA_TQ == FA_TKV (the two expressions are identical), so it survived three rounds of changes (8n → P5·0 → P5·3) unharmed. A latent bug's signature is "the equivalence holds under the current parameters" — the first asymmetric parameter (r46 was this campaign's first experiment with asymmetric FA tiles) detonates it. Writing 3 * FA_TQ or (FA_TQ + 2*FA_TKV) made no observable difference while the equality held; the difference only shows when the first symmetry-breaking change arrives.

2.6 The wall-clock ceiling: why this lever was doomed to fall below the bar

The FA slice was about 124.7 ms of the wall at the time (r47's subsequent precise attribution: 125.8 ms = 10.2% of 1239 ms GPU busy). The lever's mechanism ceiling is kernel −11% ≈ −16 ms, against a ~1.27 s whole-prefill wall a ~1.3% nominal ceiling — even fully cashed in and serially added to the wall, it is already under the +1.5% bar. The measured +0.27% (within the noise band) says even this much barely surfaced — the A/B noise band is of ±2% magnitude, and +0.27% is indistinguishable from 0.

Re-doing the arithmetic of "why even the nominal ceiling is this small": the lever's 11% magnitude looks decent, but it acts on a slice holding only 10% of the wall, so the product is 1.1–1.3%. Lever value = slice share × lever magnitude; once multiplied, any small factor can kill the product — this arithmetic later became the standard back-of-envelope before project approval (r55's swiglu roofline bound and r47's FAP2 estimate 2× → −4.9% use the same formula). The structure of the conclusion matters: the mechanism holds (kernel −11%), the lever is too weak (slice × magnitude = ceiling < bar) — what is vetoed is "this lever × this timing", not "the FA slice" (a distinction r47 immediately corrects, see §5).

The contrast of "same mechanism, different wall" is also worth recording: before 8n, attention (the LOCAL-memory-accumulation version) was 76% of whole-prefill, and a kernel improvement of the same magnitude in that era was a +8%-class wall event; by r46's day it was worth +0.27%. A mechanism's value is not conserved — every time the wall's composition changes, every historical lever's value must be re-priced. That is the arithmetic essence of "wall decompositions have a shelf life" (r47).

3. Implementation

3.1 Design choices: an audited small lever, not a rewrite

The audit's conclusions directly shaped the implementation: the structure (wmma + online softmax) stays untouched; only three quantities move — FA_TKV 64→32 (one #define, everything else symbolic), the S/P row stride +8 f32 (two stride constants), and the launcher's smem formula (one line). Each is a minimal diff that can be reverted independently.

"Everything else symbolic" deserves emphasis: every tile-related quantity in the kernel (columns per tile, softmax's column ownership, mask boundaries, the grid's token tiling) must be derived from FA_TQ/FA_TKV rather than hardcoded — r46's diff is small because the FA code has maintained this discipline since 8n; conversely, any hardcoded constant silently breaks when tiles change (r50 runs into this from another direction: the tail tile's O write-out reuses smem as a 64×128 f32 buffer, and that 32 KB is an implicit constant no symbol covers, so the launcher must take the max to be safe). There is one more reason not to rewrite: P5·0/P5·3's records already prove this FA skeleton's tile geometry is a measured local optimum; tearing it down has no audit evidence behind it.

3.2 Key code

First draw the kernel's per-tile execution flow, marking where each of r46's three levers lands (the +8 padding on K/V staging is pre-existing from P5·3; the S/P padding and TKV are new in r46):

per block (one 64-token q tile × one head):
  stage Q → Qs [FA_TQ rows]                     ← +8 padded row stride (P5·3)
  for kt in 0..nt step FA_TKV:                  ← r46: step 64→32 (lever 1)
    stage K/V → Ks/Vs [FA_TKV rows]             ← +8 padded row stride (P5·3)
    QK^T (wmma) → fc fragments
    store fc → Sf [row stride = FA_TKV f32]     ← r46: +8 f32 padding (lever 2)
    softmax: read Sf / update m,l,alpha / write Pf   ← the S/P smem round trip (deleted in r48)
    P·V (wmma): load Pf [row stride = FA_PSTR]  ← likewise protected by the padding
    rescale O
  writeout O
launcher: smem formula                          ← r46: 3*FA_TQ → real row counts (lever 3)

The r46-era smem layout and S/P round trip (historical tree d38744d~1, pre-r48) — the definitions and plane layout:

#define FA_TQ 64
#define FA_TKV 64                                    // ← r46's lever: change to 32
#define FA_PSTR (FA_TKV * 2) // probs row stride in halves (256B): probs row r
                             // aliases only Sf row r's first half, already read
                             // by the same thread — no cross-thread race
...
__half* Ks = Qs + FA_TQ * sstr;
__half* Vs = Ks + FA_TKV * sstr;
float*  Sf = reinterpret_cast<float*>(Vs + FA_TKV * sstr);
__half* Pf = reinterpret_cast<__half*>(Sf);        // alias: probs after softmax
float*  msh  = reinterpret_cast<float*>(Sf + FA_TQ * FA_TKV);
float*  lsh  = msh + FA_TQ;
float*  alpha = lsh + FA_TQ;

S's landing and the softmax round trip (store_matrix_sync out to Sf → read Sf for max/sum → write Pf; row stride FA_TKV f32 = 256 B, i.e. §2.3's conflict source). Also read the online softmax's state structure along the way: msh/lsh/alpha are block-shared arrays, one entry per row — m/l updates, the cross-tile alpha rescaling, and the final O scaling all pass through them, once per KV tile; the S/P smem round trip stacks on every step of this state machine. In the whole "warp-per-row" loop, one row's 64 columns are split 2 per lane across 32 lanes, with max/sum reduced via a __shfl_xor tree — that part is clean design; the problem is only the smem transport before and after it:

wmma::store_matrix_sync(&Sf[swm * 16 * FA_TKV + swk * 32], fc[0], FA_TKV, wmma::mem_row_major);
wmma::store_matrix_sync(&Sf[swm * 16 * FA_TKV + swk * 32 + 16], fc[1], FA_TKV, wmma::mem_row_major);
...
for (int rr = warp; rr < FA_TQ; rr += 8) {         // one WARP per row
    int c0 = lane * 2, c1 = c0 + 1;
    float s0 = v0 ? Sf[rr * FA_TKV + c0] : -INFINITY;   // ← round-trip stop 2: read Sf
    ...
    Pf[rr * FA_PSTR + c0] = __float2half(p0);           // ← round-trip stop 3: write Pf
    Pf[rr * FA_PSTR + c1] = __float2half(p1);
}
// round-trip stop 4 (P·V): ldmatrix takes 8×16 from Pf at row stride FA_PSTR — 256B ≡ 0,
// 8 rows in the same bank group = 8-way conflict:
wmma::load_matrix_sync(pa, &Pf[wm * 16 * FA_PSTR + kk0], FA_PSTR);

The old launcher formula and its hidden over-request (same historical tree):

// old: the first term counts K/V as FA_TQ rows too — correct only when FA_TQ==FA_TKV
size_t smem = (size_t)3 * FA_TQ * (hd + 8) * 2
            + (size_t)FA_TQ * FA_TKV * 4 + 3 * FA_TQ * 4;
// r46's fix (at the time): the staging term counts the planes' real rows —
//   (FA_TQ + 2*FA_TKV)*(hd + 8)*2 + FA_TQ*FA_TKV*4 + 3*FA_TQ*4
//   → at TKV=32 the request goes 61,184 → 43,776 B, unlocking 2 blocks/SM.

The r46 legacy surviving in the current tree (fa_prefill_f16kv and the launcher, the FAP2 version) — FA_TKV 32 is exactly the value r46 introduced; the launcher's smem formula has been changed to count the planes' real rows, and the current tree's comment signs this fix directly:

#define FA_TQ 64
#define FA_TKV 32                          // ← r46's 64→32 survives to this day
...
// Padded smem row stride: hd=128 halves = 256B ≡ 0 mod 32 banks makes
// every wmma ldmatrix row land on the same bank group (8-way conflict
// per load). +8 halves (272B) shifts each row by 4 banks.
const int sstr = hd + 8;
int launch_fa_prefill_f16kv(...) {
    // Qs + Ks + Vs only (S/P no longer go through shared memory). sstr = hd+8
    // padding; Ks/Vs are FA_TKV rows (the r46 launcher's 3*FA_TQ bug is gone).
    size_t smem = ((size_t)FA_TQ + 2 * FA_TKV) * (hd + 8) * 2;
    ...
    dim3 grid((nt + FA_TQ - 1) / FA_TQ, nh, 1);
    fa_prefill_f16kv<<<grid, 128, smem, stream>>>(...);   // 256→128 threads is r48's change
}

The S/P round trip itself was deleted entirely in r48 (S/P no longer enter smem, which is why the current launcher has no Sf/Pf term at all), but the "row stride +8" padding idea lives on as the staging padding (sstr = hd+8) — the same 272 B rule that P5·3 introduced, r46 reused on S/P, and r48 kept on the K/V/Q staging.

3.3 Pitfalls

  1. The launcher smem over-request is the lever's silent killer. 3*FA_TQ happens to equal (FA_TQ+2*FA_TKV) when FA_TQ==FA_TKV, harmless for years; the first asymmetric tile change makes it over-request by 17.4 KB, pressing 2 blocks/SM back to 1 — no error, no functional anomaly, kernel time nearly unchanged; the only observation point is the occupancy counters. For any "change tile size to buy occupancy" experiment, step one is finding the line where the launcher formula's equality with the real plane geometry fails.
  2. The 256 B ≡ 0 mod 128 B bank rule recurs. P5·3 fixed the staging rows, and r46 hit the same disease on the S/P rows; only r48's outright deletion of the S/P round trip cured it. Any smem layout whose "row stride is a multiple of the warp-visible width" must first pass this congruence check.
  3. The fixed-cost tax of halving tiles exists from day one. FA_TKV halved → tile count doubled → per-tile softmax state updates/barrier counts doubled. In r46 the occupancy doubling covered it (net −11%); in r50 (32→16) it overtook (wall-neutral) — two segments of the same tax-rate curve, and r46+r50 together closed the entire "shrink the FA tile" axis.
  4. The audit precedes the lever. Acting on the assumption and "rewriting FA as tensor core" outright would have burned the whole session on a kernel that was already tensor-core. Half an hour of code reading saved a directional error.
  5. The post-lever utilization reading must point to the next step. After the lever landed, SM busy was only 40% — even at 2 blocks/SM, the SM still had long idle stretches (staging latency and the softmax serial chain remained). At the time this reading was not a "not good enough" setback but FAP2's signpost: the remaining problem was not in scheduling geometry (occupancy doubled, bar unmet) but in the data path (the S/P round trip itself) — delete it, and busy and the wall move together (r48's FAP2 version is exactly what improved SM busy and the wall simultaneously).

4. Verification

GateReadingWhat it defends against
Structure audit's three readingswmma/online-softmax in place; smem ledger 69.38 KB; ncu occ 16.64% + est speedup 68.7%"investing in the wrong kernel" — the audit is itself the evidence chain, and every later number checks back against §2's arithmetic
Measured occupancy2 blocks/SM after FA_TKV=32 + the launcher fix§3.3#1's silent over-request — occupancy must be measured, never inferred from "the code changed"
Kernel timing5.16 → 4.58 ms (−11%), SM busy 40%rules out "the mechanism never took effect"
whole-prefill interleaved A/B+0.27% (< the +1.5% bar)machine drift faking a trend; same-window same-binary interleave, the verdict "lever too weak"

Note: r46's record lists no parity/greedy-specific gate (the lever is a pure scheduling-geometry change; the numeric path is untouched); r50 of the same month added this lesson after the fact — any FA_TKV change alters the online softmax's accumulation order (the length of each tile's m/l/alpha rescaling chain changes), so a strict greedy byte-identity gate is inherently unsatisfiable for this class of change (the FA campaign's new gate note, hit again in r57). r46 happened just before that gate note was discovered; after r50, the verification baseline for FA-class changes became "the parity gate + tolerance-style comparison", with the byte-identity gate reserved for changes that do not touch accumulation order.

5. Results

Layerbefore (r45's reverted state)after (r46)Verdict
Structure audit"FA might be scalar" (unverified assumption)wmma m16n16k16 + online softmax in placeno rewrite-class opportunity
smem/occupancy69.38 KB → 1 block/SM (16.64% occ)43.8 KB → 2 blocks/SM (33.3% occ)lever took effect
S/P conflictrow stride 256 B ≡ 0 (MIO 36%)+8 f32 padding staggered the banksmechanism took effect
FA kernel5.16 ms4.58 ms (−11%)mechanism took effect
whole-prefillbaseline+0.27% (nominal ceiling ~1.3%)< the +1.5% bar

Measurement protocol as in r44/r45: same-window same-binary interleaved A/B; kernel-level is matched-nt nsys/ncu sampling. The A/B noise band is of ±2% magnitude, and +0.27% is statistically indistinguishable from 0 — this is both the evidence for "below the bar" and a reminder that any <1% wall reading alone cannot ground a conclusion; it must be triangulated with mechanism-layer readings (kernel timing, occupancy, SM busy). vs-llama stays at 1.27×.

Veto mechanism (why reverted): +0.27% is far below the +1.5% bar. The arithmetic had already sealed it: FA slice ~124.7 ms × lever magnitude 11% ≈ −16 ms, a ~1.3% nominal ceiling against the ~1.27 s wall — even perfectly cashed in, it cannot clear the bar, and the measured value merely confirmed the ceiling was not even reached. r46's written conclusion at the time was "FA is not the wall-critical path in the converged-GEMM domain"; reverted by discipline (code not kept, mechanism filed).

The post-hoc correction and retry conditions (this doc's most important turn): the immediately following r47 redid the converged-domain wall decomposition with fresh nsys, overturning the "FA is not wall-relevant" half of that sentence — FA = 125.8 ms = 10.2% of wall = 5.72× vs-llama, the whole engine's #1 structural residual. r46 and r47 do not contradict; they veto/endorse different propositions: r46 proves "this lever (the −11% occupancy/conflict fix) cannot reach the bar on this slice"; r47 proves "this slice deserves a bigger lever". FAP2 was thus chartered: register-resident softmax (delete the S/P smem round trip entirely rather than pad it), estimated 2× → −63 ms = −4.9% wall — r48 cashed it: FA kernel 5.16 → 2.12 ms (2.43×), whole-prefill 2603.5 → 2749.9 (+5.6%), FA's wall share 10.2% → ~4.7%. r46's mechanism findings (occupancy starvation + the S/P round trip) were inherited verbatim by r48; only the solution upgraded from "optimize the round trip" to "eliminate the round trip".

Two more rounds on the same lever axis closed the "shrink the tile" axis for good:

  • r50 (FA_TKV 32→16): 3 blocks/SM achieved, but the wall −0.5%/−0.01% — the occupancy gain was canceled by the doubled per-tile synchronization/softmax overhead (the fixed-cost tax forecast in §2.4 overtaking), and greedy-32 byte-identity was lost (inherent to FA_TKV changes). Merged with r46 into the conclusion: shrinking the FA tile is a dead lever (both points, r46 + r50, negative).
  • r57 (FA KV staging double buffering): FA_TQ 48 + double-buffered K/V; both attempts lost greedy-32 identity at token 19 — r50's "the FA tile is a size" notice striking again.

With that, the FAP lever map settled: beyond the occupancy axis (r46/r50/r57, dead), the only big lever left was eliminating the S/P round trip itself (r48, +5.6%) — exactly the value of r46's audit-produced mechanism list: it marked "where the blood is" correctly; the day's bandage was merely too small.

One more note on the campaign-wide impact: after r48 landed FA held ~4.7% of the wall, and r47's decomposition simultaneously named the A-quantize prepass's shared-A dedup (r49, +2.32%) and producer folding (r51/r52, +1.89%/+5.45%) — FAP1's same-day "book it and move on" decision ultimately cashed in all the mechanism information it preserved, in the form of "FAP2 + the prepass line". r46's round's direct output list:

#OutputFormDownstream cash-in
1The structural conclusion "FA is already wmma + online softmax"audit recordspared a directional rewrite
2The mechanism list of occupancy starvation + the S/P round tripaudit recordr48's FAP2 accepted it wholesale
3The launcher smem formula fix (staging term counts real rows)code (survived)permanent; signed in the current tree's comments
4FA_TKV = 32code (survived)permanent; r48/r50's baseline
5FAP2 chartered (2× → −4.9% wall estimate)recordr48 cashed +5.6%

Judged as "REVERTED", this is the directory's densest-output doc: of five outputs, four survived across rounds; the code revert only discarded "the occupancy lever itself" — the one thing already proven insufficient at the time.

6. Lessons

  1. Positive mechanism / negative wall clock is a real outcome class: record the mechanism and re-rank the slice, rather than discarding both — "kernel −11% but wall +0.27%" hides the next step's map.
  2. A wall-neutral measurement vetoes "this lever", not "this slice": judging a slice's value requires a fresh whole-wall decomposition (r47), never extrapolation from the old ledger or a single lever's result.
  3. Check the launcher before any occupancy-class lever: the smem request formula's equality with the real plane geometry holds only for certain tile combinations; change the tile without checking the launcher and the gain silently zeroes out with no error at all.
  4. Audit before coding: half an hour of structural audit closed an imagined 20× rewrite and pointed the budget at the FAP2 that actually worked.

← 48-r45-cpasync-q6k-a-staging · Index · 50-r47-converged-wall-decomposition →

50 · r47 — converged-era whole-wall re-decomposition (MEAS-ONLY)

Result: r37's attribution table expired exactly as forecast. Re-measured on the current best gate set after the whole r38–r45 q6_K package landed: q6_K GEMM 1094.7 → 196.4 ms (51.2% → 15.8% of the wall), the whole wall 2190 → 1274 ms (1521 → 2610 tok/s, same-window interleaved band 2585–2623), vs-llama 2.15× → 1.27×. q4_K at 33.2 µs/GMAC (1.06×) and q6_K at 65.0 µs/GMAC (1.13×) both reached parity — the GEMMs are no longer the wall; FA at 125.8 ms = 5.72× became the #1 structural residual (10.2% of GPU busy), falsifying r46's "smaller/overlapped" assumption; the quantize prepass grew +31.6 ms (the hidden tax of the q6_K BT port). Recommended next lever: FAP2 register-resident softmax (estimated at 2× → −63 ms = −4.9% of the wall, clearing the +1.5% bar), followed by the A-quantize shared-A dedup. Commit: 11e3640 (docs-only record commit, no code change). Date: 2026-09-05.

1. Background — where things stood

The r38–r41 q6_K campaign was a string of textbook landings: the BT-style raw-byte kernel (+2.87%), KDR=2 double buffering (+13.3%), the third resident block (+13.0%), the B-expand uint4 widening (+30.7%) — four steps pulled q6_K from 368.9 µs/GMAC down to the 65 µs magnitude and pushed whole-prefill from 1521 all the way to 2605 tok/s. But the next three levers hit the wall in succession: r42 (dsc widening −0.19%), r44 (the W_exp stride fix went parity green but the wall −0.42%), r45 (cp.async A-side staging, wall −0.34%) — all REVERTED. r45's conclusion was blunt: the q6_K line has converged — the kernel is no longer the binding constraint, and tuning the q6_K kernel further cannot move the wall.

r46 (FAP1) supplied another instance of the same lesson. The audit found the FA kernel (fa_prefill_f16kv) already wmma + online softmax, with the bottleneck being occupancy (69.38 KB smem → 1 block/SM) and the S/P smem round trip's bank conflict; after the fix the kernel went 5.16 → 4.58 ms (−11%), fully mechanism-positive, yet the wall moved only +0.27% — below the +1.5% bar, REVERTED. "Mechanism-positive / wall-negative" became a real outcome class: whether a kernel-level 2× improvement opportunity is worth pursuing depends on its share of the wall and on whether it is really on the critical path.

This is exactly where r37's rule applies: "attribute the WHOLE wall after every convergence". r37's attribution table drove the entire r38–r41 priority queue (q6_K raw kernel first), and that table was drawn in a world of 2139.6 ms GPU busy: q6_K 51.2%, FA 5.7×, prepass 86.8 ms. Eight landing rounds later, both uses of that table had expired — the denominator (the whole wall) shrank 40%, and every slice's share drifted. Without redrawing the table, the next lever choice would rest on stale data.

Two candidate directions sat on the table, neither backed by current wall-level evidence: one is the FAP2 r46 named (register-resident softmax, deleting the S/P smem round trip outright), the other the quantize prepass redundancy r37 had already flagged ("1.25× llama, partly 2× per-shared-A redundancy"). r47 is a pure measurement round: without changing a line of code, slice the whole wall open under the current best gate set and rank the two candidates with the freshest numbers.

"Current best gate set" has concrete content at this moment: the P5-era TM=128 wide-tile GEMM, r20's split-phase A staging, r22's qa8 XOR swizzle, the r28–r32 NB kernel family, r34's quantize-transpose prepass, and the r38–r41 q6_K BT package (KSPLIT=2 raw-byte kernel, KDR=2 double buffer, 3 resident blocks, uint4 B-expand), all on by default; the REVERTEDs of r42/r44/r45 and r46's REVERTED mean the FA and q6_K kernels sit at their respective convergence points. In other words, this table portrays "the engine after every proven lever has been pulled" — precisely why it deserves the word "converged" — and precisely because of that, its differences from r37's table can be attributed directly to the four rounds r38–r41, with no in-flight changes mixed in.

One more methodological detail in the r46→r47 handoff worth recording: r46's grounds for vetoing the FA lever were "small slice share + mechanism in doubt", but that judgment used r37-era shares. The question r47 must answer is therefore sharp — after the GEMM collapsed from 51.2% to parity, is FA's 125.8 ms still "small"? The answer is no (10.2% busy, 5.72× ratio), which is exactly the value of re-measuring: a share judgment's shelf life does not outlast one convergence.

2. Principle — the GPU mechanism

The attribution's dimensions. One decomposition = run nsys over the complete prefill forward and group GPU busy time by kernel role: each weight type's GEMM (q4_K / q6_K / f16…), attention (FA), the quantize prepass, and the elementwise/copy long tail. Normalization within a role uses GMAC: a matmul class's slice time ÷ that class's total GMAC count gives µs/GMAC. The point of GMAC normalization is stripping shape — q6_K's attn_v and ffn_down differ by an order of magnitude in size, but µs/GMAC is directly comparable and directly matches llama.cpp's reference implementation's same-class numbers; the ratio of the two is that line's parity multiple (the llama reference is provided by the bench build ca3d5a3e1).

Worked example (r37's q6_K row): for the 3325-token prefill, the q6_K matmuls (attn_v + ffn_down, 28 layers) have a fixed total GMAC count, and 1094.7 ms divided into it = 368.9 µs/GMAC; llama at the same shape is 57.8 µs/GMAC, a ratio of 6.38×. r47 re-measures: 196.4 ms ÷ the same GMAC total = 65.0 µs/GMAC, a ratio of 1.13×. Note the numerator shrank 5.6× between the rounds while the denominator (the GMAC total) varies only with token count and both rounds share the anchor — so the ratio's improvement comes entirely from the numerator. That is what normalization buys: separating "the wall got smaller" from "the kernel got better". Likewise q4_K: 598.1 ms ÷ its GMAC total = 33.2 µs/GMAC, 1.06× vs llama.

The denominator effect: shares squeeze each other by nature. r44 supplies a miniature sample of this principle: inside the q6_K kernel, after fixing the recomb the kernel duration fell −10.9%, but long_scoreboard's share rose from 33.6% to 57.1% — the numerator fell, the denominator fell with it, and the remaining stall's share rose instead. The wall-level decomposition behaves the same: after the q6_K slice collapsed, FA's and prepass's shares rise passively even if their absolute values are unchanged (124.7/2190 = 5.7% → 125.8/1239.2 = 10.2%, FA's absolute duration differing by only 1.1 ms). When reading a decomposition table you must read absolute durations and shares together, or you will misread "the denominator effect" as "the residual worsening".

Why decompositions expire. The wall is a denominator; a landing changes the numerators and the denominator. At r37, q6_K held 51.2%; four q6_K rounds pressed it to 196.4 ms, but the whole wall also shrank from 2190 to 1274 ms — if nothing outside q6_K had changed, its share would have fallen from 51.2% to 196.4/1274 = 15.4%; the actual 15.8% says the other slices barely moved. Ranking is relative: once r37's top lever was done, the second tier (FA 5.7×, prepass redundancy) automatically moved to the front of the window, but their absolute shares must be re-measured — r46 had already proven FA's 124.7 ms slice "not worth it" in the then-GEMM-dominated world; now that the GEMM has collapsed to parity, the same 125.8 ms slice is the biggest non-parity residual.

Why 5.72× is called "structural". This multiple is not r47's discovery — r23 named it structural back in the f16-path decomposition era ("FA's 2.5×/layer gap is structural: llama keeps 128-wide KV tiles"): llama's attention tile geometry holds the KV panel at a completely different width, so the per-layer attention gap on the two sides is decided by tiling choice, not instruction efficiency. r37 measured 5.7×, r47 measured 5.72× — across all the GEMM landings of r38–r45 this number did not move an inch, which is exactly what validates the word "structural": nothing on the GEMM side reaches it. It also predicts two things: per-instruction micro-tuning will not converge it (r46's padding route was tried and failed), and only a geometry-class rewrite (r48's whole-row warp tile) has a chance.

Hidden taxes must be booked. A landing's net contribution to the wall = its direct gain − the indirect cost it introduces. r38–r41's direct gain sits on the q6_K slice: 1094.7 − 196.4 = 898.3 ms. But r47 finds the quantize prepass rose from 86.8 to 118.4 ms (+31.6 ms): the q6_K BT port also connected attn_v's and ffn_down's A side to r34's quantize_q8_0_pad40_t prepass, and the quantization work that used to run on the generic mmq_nt<7,2> path was "moved" into the prepass. Net wall gain = 898.3 − 31.6 = 866.7 ms — the direction unchanged, but the per-lever accounting is only honest with the tax included.

Converting shares into budget. Once you have the shares, estimating a candidate lever's ceiling is one line of arithmetic. FAP2's claim is that deleting the S/P smem round trip makes the FA kernel ~2×: 125.8 ms → ~63 ms, saving 63 ms; 63 / 1274 = 4.9% of the wall, more than three times the +1.5% bar. The prepass dedup's claim is that q/k/v and gate/up each share one A (see doc 52 for detail), estimated at the ~40–60 ms order from r37's "~2× redundancy" reading — both candidates clear the bar, and the ranking depends on which mechanism is more certain.

Candidate leverMechanism claimBudget arithmeticvs the bar
FAP2 register-resident softmaxdelete the S/P smem round trip → FA kernel ~2×125.8/2 ≈ 63 ms = 63/1274 ≈ 4.9% of wallthreefold headroom, #1
A-quantize shared-A dedupsame A not re-quantized → prepass slims downredundancy ~43% of launches × 0.6 ms each ≈ 40–60 msclears the bar, #2 (simpler mechanism, lower risk)
q4_K short-nt dilutionmatched-nt 1.17× readingactually 1.06× at prefill nt — nothing to harvestshelved

The most informative line of the budget arithmetic is the third: the biggest slice, at a 48.3% share, has a budget of zero — because its ratio is 1.06×. The decomposition table's whole value is putting "share" and "ratio" side by side; missing either one misranks the queue.

3. Implementation

3.1 Measurement design: zero code changes, same window, same tools

r47 is a deliberate "no-code" round: the repo tree sits at the post-r45/r46-convergence state (r46's change already reverted), and the only artifact is the docs commit 11e3640. The measurement protocol follows r37's template:

  • Load: the 3325-token prefill (the same anchor as r37), which keeps µs/GMAC directly comparable with r37;
  • Tools: full-graph nsys for per-kernel durations and the GPU busy total; matched-nt ncu for each weight type's per-IMMA/µs-GMAC ratio;
  • Gate set: every landed lever (the q6_K BT package, TM=128, the MMQ paths, etc.) on by default, with no A/B variable — this table measures "what the engine looks like now", not "the increment of some change";
  • Same-window interleave: whole-prefill rates sampled interleaved within one session window, reporting the median and the band (2585–2623), honoring the master table's reading convention — cross-session absolute values are not comparable.

The evidence actually collected (item-for-item isomorphic to r37, keeping the two tables mutually readable):

  • the nsys full-graph per-kernel duration table → bucketed and summed by (weight-type GEMM / FA / prepass / elementwise), giving each slice's absolute milliseconds;
  • ncu matched-nt (the bench shape at nt≈511) → each weight type's per-IMMA / µs-GMAC ratio, normalized against the llama reference;
  • whole-prefill interleaved A/B (same binary, same window) → the 2610 tok/s median and the 2585–2623 band;
  • no A/B variable, no env flips — this round's only "control" is r37's old table itself.

3.2 r37 vs r47: the two tables read side by side

Slicer37 (2026-09-05, morning)r47 (same day, after r38–r45)Change
GPU busy2139.6 ms1239.2 ms−900.4 ms
Whole wall2190 ms (1521 tok/s)1274 ms (2610 tok/s)−916 ms
vs-llama2.15×1.27×−0.88×
q6_K GEMM1094.7 ms = 51.2% (368.9 µs/GMAC, 6.38×)196.4 ms = 15.8% (65.0 µs/GMAC, 1.13×)−898.3 ms, parity reached
q4_K GEMM (BT)~600 ms (60.0 TFLOPs, matched-nt 1.15×)598.1 ms = 48.3% (33.2 µs/GMAC, 1.06×)flat, now parity-class
FA attention124.7 ms (5.7×)125.8 ms (5.72×) = 10.2% busyflat, promoted to #1 structural residual
quantize prepass86.8 ms (1.25× llama)118.4 ms = 9.6% busy (+31.6 ms)the hidden tax of the q6_K BT port
Rest (elementwise/rope/store_kv/lm_head/copy long tail)~230 ms~200 msthe four named slices total 1038.7 ms; the remainder is the long tail

Three points from reading them together:

  1. q6_K's win is real, but must be netted: 898.3 ms of direct gain minus the +31.6 ms prepass tax is a net +866.7 ms, the overwhelming majority of the GPU busy reduction (900.4 ms) — the q6_K line really was the only big mover of these eight rounds.
  2. The biggest slice is already parity-class: q4_K's 598.1 ms is the largest single slice at 48.3%, but 1.06× means llama could not be much faster either — a big share ≠ a lever. Residual ranking must read share and ratio together: FA is the largest product of the two (10.2% × 5.72×).
  3. The prepass went from background noise to the third-biggest slice: 9.6% busy, with a clear mechanistic redundancy (the same A quantized repeatedly) — r49's foreshadowing.

The decomposition's self-consistency can be written directly as an equation (STYLE's "if a diagram doesn't fit, write the equation" usage):

GPU busy 1239.2 ms
  = q4_K GEMM   598.1  (48.3%  — 33.2 µs/GMAC = 1.06×, parity class, no lever)
  + q6_K GEMM   196.4  (15.8%  — 65.0 µs/GMAC = 1.13×, converged this round)
  + FA attn     125.8  (10.2%  — 5.72×,        #1 structural residual)
  + prepass     118.4  ( 9.6%  — ~43% launch redundancy, #2)
  + long tail   ~200.5 (16.1%  — rms/swiglu/rope/store_kv/lm_head/copies)
wall 1274 ms = busy 1239.2 + launch gap/D2H ~35 ms   →  3325 tok ÷ 1.274 s ≈ 2610 tok/s

The q4_K row deserves a pause: 1.06× does not mean "we are only 6% behind"; it means "the llama reference is simply this fast on this class of matmul" — jointly determined by the physical limits of IMMA throughput, smem bandwidth, and DRAM traffic. At this point q4_K's 598.1 ms stops being an "optimization target" and becomes "terrain": any proposal to shave time off it must first explain how it would beat the reference by more than 6%. This is also why after r47 the campaign never touched the q4_K kernel body again (r58's failed transplant was a structure port, not kernel tuning).

The four named slices total 1038.7 ms, 83.8% of the 1239.2 ms GPU busy; the remaining ~200 ms is the elementwise and copy long tail (rms/swiglu/rope/store_kv, embedding, lm_head, and the GPU-side work ahead of the D2H readback). Later rounds proved this long tail worth slicing too: r51/r52 measured its fused-producers slice at 151.9 ms and cut it to 86.5 ms with producer fusion (−5.45% of the wall), and r55 first used a roofline to prove swiglu already sits at 89% of the bandwidth roof (242 GB/s) — the long tail is not miscellaneous; it is an unranked pool of candidate levers; r47 leaving it as a watch item was the right restraint.

3.3 Pitfalls: two attribution traps

Pit one: hidden taxes not booked. Looking only at the q6_K slice (−898.3 ms) and not the prepass (+31.6 ms) overstates the q6_K campaign's books by 3.5%. The prepass's growth mechanism is subtle: it is not that some commit "got slower", but that r38's BT port changed which work takes the prepass road — attn_v/ffn_down's A quantization migrated from the generic kernel path to the shared prepass. This class of "work relocation" is exposed only by the difference between two decompositions.

Pit two: taking the matched-nt dilution reading at face value. matched-nt ncu measures q4_K per-IMMA at 1.17× on the nt≈511 bench shape (even 1.43× in the r37 era) — far worse than the 1.06× at prefill nt. r47 explicitly rules this a short-nt-specific prologue dilution effect (prologue/staging overhead amortized over a small nt), not a real optimizable gap, and shelves it. The lesson is isomorphic to r46's: kernel-level ratios must be read at the right shape, or you schedule phantom levers.

Pit three: reading cross-session absolute values as increments. 1521 (r37's window) and 2610 (r47's window) are not two numbers from the same machine in the same state — co-tenant load and clock policy both move absolute anchors (the master table reading convention was later reinforced painfully in r59b: "anchor every A/B baseline behaviorally in the same window"). r47's numbers are all same-window interleaved medians; cross-round comparison runs only on ratios and shares, never on absolute tok/s differences. This discipline cost nothing this round and was worth a +26.2% → +11.1% correction in r59.

4. Verification

This is a measurement round; what is "verified" is the decomposition's own credibility:

  • Same-window interleaved median: whole-prefill reported with its same-window band of 2585–2623 tok/s, avoiding single-sample machine-state noise (the master table reading convention + a forerunner of r59b's lesson);
  • The busy-vs-wall reconciliation: GPU busy 1239.2 ms vs wall 1274 ms, the ~35 ms difference being launch gap and D2H readback — a reasonable magnitude, showing the slice summation missed no big item;
  • Slice-sum self-consistency: the four named slices 1038.7 ms + long tail ~200 ms ≈ busy 1239.2 ms, with each slice's duration from nsys's per-kernel sums cross-checked against ncu's per-kernel duration × launch count (r37 protocol's two independent sources interlocking);
  • Anchor regression: vs-llama uses the same-window llama-bench 3325-eq anchor (the ~3400 tok/s family), 2610/3400 ≈ 1.27×, subtractable in the same coordinate system as r37's 2.15×;
  • r48's retrospective check: r47 predicted FAP2 2× → −63 ms; r48 measured FA kernel 5.16 → 2.12 ms (2.43×) and whole-prefill 2603.5 → 2749.9 (+5.6%, ~68 ms of wall saved) — the predicted −63 ms and the measured saving agree within noise. The decomposition's ranking was cashed in by the subsequent landing.
  • r49's retrospective check: three mutually corroborating levels of independent evidence for the prepass redundancy (r37's 1.25×-llama annotation → r47's 9.6% busy → r49's 193→110 census and +2.32%) — r47's "mechanistic redundancy" judgment lands precisely in the launch census.

5. Results

The measured wall (3325-token prefill, current best gate set, no code changes):

  • GPU busy 1239.2 ms, whole wall 1274 ms, 2610 tok/s (interleaved band 2585–2623), vs-llama 1.27×;
  • q6_K GEMM 1094.7 → 196.4 ms (51.2% → 15.8%), 65.0 µs/GMAC = 1.13× (parity-class);
  • q4_K GEMM 598.1 ms = 48.3% busy, 33.2 µs/GMAC = 1.06× (parity-class);
  • FA 125.8 ms = 10.2% busy, 5.72× — the #1 structural residual; r46's "FA smaller/overlapped" assumption does not hold (the slice surfaced intact after the GEMM's collapse);
  • quantize prepass 118.4 ms = 9.6% busy (+31.6 ms since r37, the hidden tax of the BT port).

The produced priority queue (r47's direct deliverable):

  1. FAP2 register-resident softmax: estimated 2× → −63 ms = −4.9% of the wall, threefold headroom over the bar — r48 subsequently cashed it at +5.6% (2603.5 → 2749.9);
  2. A-quantize shared-A dedup: q/k/v share the same normed, gate/up share the same normed2, and the prepass re-runs per matmul — r49 cashed it at +2.32% (2734.1 → 2797.5);
  3. q4_K matched-nt dilution: a short-nt-only effect, shelved (later restarted in another form in r58's q4_K BT structure spec).

The prediction-vs-cash-in comparison (this table is the measurement round's "correctness verification" — a wrong ranking would have wasted every later round):

r47 predictionCashing roundPredictedMeasured
FAP2 2× → −4.9% wallr48−63 ms+5.6% (~68 ms of wall saved) ✓
prepass shared-A dedupr49the ~40–60 ms order+2.32% (−34.5 ms prepass, ~28 ms of wall cashed) ✓
q4_K dilution shelvedr58 (restarted)—the q4_K BT port −12.6% REVERTED — the "shelved" judgment held in r58's window too

Campaign coordinates: 1521 (r37) → 2610 (r47) → 2750 (r48) → 2798 (r49) → 2856/3011/3177 (the prepass-and-q6_K bundle line of r51/r52/r53) → 3590.8 (r59b's final figure); vs-llama 2.15× → 1.27× → 1.21× → 1.18× → 1.05× → 1.080×. r47 sits at the inflection of this curve: it confirmed both GEMM lines at parity and switched the campaign from "fix the slowest kernel" to "harvest the remaining structure by wall share" — every landed gain of the following six rounds (r48/r49/r51/r52/r53/r56) finds its slice in this decomposition table.

Equally worth recording is what r47 deliberately did not do: it did not subdivide the ~200 ms long tail (which elementwise kernels hold what, how far swiglu sits from the bandwidth roof), merely flagging it as a watch item. Subdividing the tail was r51/r52's (the 151.9 ms pre-producer-fusion attribution) and r55's (the swiglu 242 GB/s = 89% roofline capping audit) work. One decomposition needs to answer only "what is next"; slicing every slice to the leaves is the next round's job — the decomposition's granularity follows the decision, not completeness.

6. Lessons

  1. Wall decompositions have a shelf life: after every convergence (a line reaching parity, or a kernel leaving the critical path) the attribution must be re-run, or the priority queue rests on stale shares — r37's table served r38–r41 precisely, r47's table served r48–r49 precisely, and no table spans two convergence periods.
  2. Net the landings: direct gain minus hidden tax (q6_K's −898.3 ms paired with prepass's +31.6 ms) is the campaign-level accounting; "work relocation" taxes surface only in the differential between before/after decompositions.
  3. A big share ≠ a lever: rank residuals by share and parity ratio together; the biggest already-parity slice (q4_K at 48.3%) has no room to move; mechanism-positive/wall-negative (r46) and mechanism-positive/share-not-yet (FA at r37's time) are both real outcome classes.
  4. Read kernel-level ratios at the right shape: the matched-nt dilution (1.17× @nt≈511 vs 1.06× @prefill nt) is a prologue-amortization shape artifact; not taking it seriously saves an entire wrong optimization line.

← 49 · Index · 51 →

51 · r48 (FAP2) — register-resident softmax: deleting the S/P smem round trip outright (LANDED)

Result: the S/P shared-memory round trip in fa_prefill_f16kv is deleted outright — softmax runs directly on the QK^T wmma accumulator fragments, and P is built in registers, in place, as the matrix_a operand of P·V. FA kernel 5.16 → 2.12 ms (2.43×, meeting the ≤2.6 ms target), whole-prefill 2603.5 → 2749.9 tok/s (+5.6%), FA share of the wall 10.2% → ~4.7%, vs-llama 1.27× → 1.21×. ncu: MIO/shared-scoreboard stalls vanish (top stall becomes global K/V loads), shared wavefronts 44% → 21%, 124 regs / 0 spill. Gates: parity ×3 green, greedy-32 byte-identical, suite 166/0/3. Commit: d38744d (code, src/cuda_kernels.cu only) + 7e2ee62 (record). Date: 2026-09-06.

1. Background — where things stood

r46 (FAP1) is the direct predecessor of this step. The audit confirmed that the FA prefill kernel (fa_prefill_f16kv, introduced in the 8n era, reworked through the P5.0/P5.3 rounds) already was wmma m16n16k16 + online softmax, not a scalar implementation; the real defects were two: 69.38 KB of smem → 1 block/SM (16.64% occupancy), and a bank-conflicted S/P smem round trip (S written to smem, softmax read out, P written back, P·V read back again, with row strides landing in the same bank group). FAP1's lever was conservative: FA_TKV 64→32 + S/P row padding + a launcher over-allocation fix — kernel 5.16 → 4.58 ms (−11%), the mechanism was entirely positive, but the wall gained only +0.27%, below the +1.5% bar, REVERTED. FAP2 survived as a name: "register-resident softmax, eliminate the S/P round trip for good."

Then r47's converged-regime wall decomposition changed FA's situation: once the q4_K/q6_K GEMMs both reached parity (1.06×/1.13×), FA at 125.8 ms = 5.72× and 10.2% of GPU busy time surfaced as the #1 structural residual — r46's "FA slice is smaller / gets overlapped" hypothesis was falsified. r47 used that same audit to run the budget: if deleting the S/P round trip buys a ~2× kernel improvement, that is −63 ms = −4.9% of the wall, clearing the bar three times over. FAP2 moved up from "mechanically worth doing" to "wall-level top priority."

In other words, the task facing r48 was very concrete: no longer another occupancy patch on the old structure, but tearing down the structural assumption that "S must pass through shared memory" itself. In the old structure softmax cooperated across warps on a per-row basis, so the data had to land in smem; to make it register-resident, the geometric relationship between warps and data had to change first.

2. Principle — the GPU mechanism

Why S/P went through smem — a geometry problem, not an implementation problem. wmma accumulator fragments live only in a single warp's registers. The old geometry was 256 threads (8 warps), warp mapping wm = warp>>1 (four 16-row q blocks) × wn = warp&1 (the left/right halves of the FA_TKV=64 columns): one q block's S was computed by two warps cooperating, while online softmax needs the max/sum over the whole row — the data had to be handed across warps, and smem was the only channel. The per-KV-tile round-trip bill: S store (64×64 f32 = 16 KB) + softmax read (16 KB) + P store-back (64×64 f16 = 8 KB) + ldmatrix read for P·V (8 warps × 4 k-steps × 512 B = 16 KB) ≈ 56 KB/tile of smem traffic, plus small round trips for the three state arrays msh/lsh/alpha, and the store-side row stride ≡ 0 mod 32 banks (8-way conflict). r46 tried to treat the conflict with padding and measured only −11% kernel — the round trip itself was still there.

The new geometry: one warp owns a full row block. r48 changes the tile to 4 warps (128 threads) × FA_TQ=64 rows × FA_TKV=32 columns: each warp exclusively owns one 16-row q block × all 32 KV columns. The QK^T wmma accumulator fc[2] (two 16×16 fragments) is now warp-private — softmax can stay in registers, and so can P.

4-lane row groups and the butterfly. Documented layout of the m16n16k16 f32 accumulator: lane L holds rows L>>2 and (L>>2)+8, columns 2·(L&3), 2·(L&3)+1 (and the +8 offsets). So all 32 columns of one row land in exactly 4 lanes (l = 0..3). The in-row max/sum reduction needs only a two-step butterfly with __shfl_xor offsets 1 and 2 — the old implementation was a full-warp 32-lane five-step reduction (offsets 16..1). Each lane maintains the m/l state for two rows (m0/m1/l0/l1), fully register-resident; the three smem arrays msh/lsh/alpha disappear entirely.

The f32 accumulator ↔ f16 matrix_a lane-map identity. The A operand of P·V is an f16 matrix_a fragment (m16n16k16 row_major), and its element→x[i] mapping is exactly the same as the f32 accumulator fragment's (the PTX ISA specifies the same per-lane layout for both fragment kinds). This identity (flagged unit-validated in the commit message) means P needs no sts/lds/shuffle at all: write __float2half(p) straight into pa[cc].x[i] and the same register becomes the A operand of the next wmma in place. The S→P "round trip" degenerates into an in-register type conversion.

The register ledger: why 124 regs fit. Per-warp resident fragment state (approximate accounting): O accumulator acc[8] = 8 × 8 = 64 f32; QK^T accumulator fc[2] = 16 f32; P operand pa[2] = 16 f16 (8 32-bit registers); softmax linearized arrays sm/sm1_/gcol = 8+8+8 = 24; m/l state and alpha 8 more; ~120 live registers in total — matching the measured 124 regs / 0 spill. The same ledger explains two design boundaries: why each warp can only take FA_TKV=32 (fc/pa/sm all grow linearly with FA_TKV; 64 would need +60 regs and would certainly spill), and why doubling the O accumulator from the old acc[4] to acc[8] did not blow the budget — the old implementation gave each warp only the wn half-width (64 output columns), the new one gives each warp all 128 output columns of its full row block; what the width costs is paid for by deleting the S/P smem round trip, and the net ledger is favorable.

The major-ness trap. The B operand of QK^T is K^T (k = head dim, n = KV position), and K is physically stored row-major as [kv][hd] — which for wmma is exactly a col_major B (B[k][n] = ptr[n·ldm + k]); the B operand of P·V is V itself (k = KV position, n = head dim), where the same physical layout must be a row_major B. The two consumers read the same tile along two different axes, so their major-ness flags are necessarily opposite. During development both were written as row_major: QK^T actually computed Q·K (untransposed) and the parity max err blew up from the 1e-4 class to 0.54, recovering only when K was set back to col_major.

Barrier and occupancy reconciliation. The per-KV-tile __syncthreads count drops from 4 to 3 (the one after QK^T, which existed to serve softmax's cross-warp reads of Sf, now has no readers). smem 69.38 → 34.82 KB = (FA_TQ + 2·FA_TKV)·(hd+8)·2 = 128·136·2 = 34,816 B, blocks go 1 → 2/SM. Note that warp-level occupancy did not change: before, 1 block × 8 warps; now, 2 blocks × 4 warps — 8 warps/SM either way. The gain comes from two places: the disappearance of the MIO/shared-scoreboard round trip (the 36% item in ncu goes straight to zero), and the phase interleaving of two independent blocks (while one block waits at a barrier the other can issue; in the single-block era a barrier idled the whole SM). r46 got −11% from occupancy alone (2 blocks × 8 warps); r48 got 2.43× by "deleting work" — a clean contrast.

The mechanism fingerprint in ncu. shared wavefronts fell from 44% to 21% — the remaining 21% is the smem write/read of Q/K/V staging itself (the part that was always supposed to be there); the vanished half corroborates the zeroing of the MIO/shared-scoreboard stalls: the S/P round-trip traffic evaporated at the counter level, it was not hidden in some other bucket. The top stall becoming L1TEX scoreboard (global K/V loads) says the bottleneck has ceded to where data genuinely moves — a kernel whose top stall changes from "traffic it built itself" to "inputs it must move" is the healthy signature of a structural change.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Full-row warp tile (4 warps) instead of keeping 8 warps: a cross-warp distributed-fragment softmax (lane reductions spanning warps, P reassembled via shuffles) is theoretically possible, but it would trade the 4-lane butterfly for a mixed smem/shuffle protocol between warps, and the complexity returns to the starting point. 4 warps × 16 rows = 64 rows = FA_TQ — the geometry closes.
  • FA_TKV=32 is part of the new geometry, not just an occupancy knob: each warp's accumulator is acc[8] (covering all output columns of hd=128; the old implementation gave each warp only the half-width acc[4]), and a 124-reg / 0-spill budget exactly fits; any larger FA_TKV doubles the fragment count of fc[]/pa[] and blows the registers; any smaller was the direction r50 later falsified.
  • 128 threads remain enough for staging: per-tile K/V staging = 2 × FA_TKV × hd × 2 B = 16 KB, spread over 128 threads that is 128 B/thread = 8 uint4 cp.async per thread — issue width is not lacking; and once softmax is register-resident, the extra 4 warps have no work to do on S and only widen the barrier-arrival wait surface.
  • Staging stays cp.async + padded stride (sstr = hd+8), and the launcher over-allocation r46 fixed is rebuilt with the new formula. The pre-sm80 build target (sm_75) keeps the synchronous staging fallback path, and its loop stride is parameterized on nthreads the same way (see the sync branch of fa_stage_kv_async in the current tree) — the same class of trap must be plugged on both paths.
  • The tail tile's O write-back reuses smem: the Qs/Ks/Vs region is idle after the KV loop, and 64×128 f32 = 32 KB < the 34.8 KB budget (this reuse later became r50's launcher lesson: two smem users must take the max).

A term-by-term comparison of old vs new geometry (this table is the diff's "semantic summary"):

DimensionOld (pre-r48)New (r48)
Threads / warps256 / 8128 / 4
Warp mappingwm=warp>>1 × wn=warp&1 (half column block)wm=warp (full row block × all columns)
Where S livesaccumulator → Sf smem (f32, 64×64)accumulator fragment fc[] (registers throughout)
Softmax shapeone row per warp, 32-lane 5-step reduction4-lane row groups, 2-step butterfly, 2 rows per lane
Where P livesPf smem (f16, FA_PSTR) → ldmatrix read backmatrix_a fragment pa[] (built in place)
m/l/alphamsh/lsh/alpha smem arraysregisters m0/m1/l0/l1
O accumulatoracc[4] (half width, 64 columns per warp)acc[8] (full width, 128 columns per warp)
Barriers per tile43
smem layoutQs+Ks+Vs+Sf(+Pf alias)+msh/lsh/alphaQs+Ks+Vs (34,816 B)

3.2 Key code

Old structure (d38744d^, 8 warps splitting the columns in half): each warp computes only a 16×32 half of the column block, and S must land in Sf via store_matrix_sync:

const int warp = tid >> 5; // 0..7
const int wm = warp >> 1;  // q 16-block: 4
const int wn = warp & 1;   // 64-dim chunk: 2
...
    wmma::mma_sync(fc[0], fa, fb[0], fc[0]);
    wmma::mma_sync(fc[1], fa, fb[1], fc[1]);
    }
    // S lands in smem: row stride FA_TKV=64 f32 = 256B ≡ 0 mod 32 banks (the conflict source)
    wmma::store_matrix_sync(&Sf[swm * 16 * FA_TKV + swk * 32],      fc[0], FA_TKV, wmma::mem_row_major);
    wmma::store_matrix_sync(&Sf[swm * 16 * FA_TKV + swk * 32 + 16], fc[1], FA_TKV, wmma::mem_row_major);
}
__syncthreads();   // softmax is a cross-warp consumer; all of S must be in place

The old softmax was "one row per warp" (8 warps sweeping the 64 rows in turn), each lane holding 2 columns with a full-warp five-step reduction; P was written back to Pf and m/l/alpha written back to smem:

for (int rr = warp; rr < FA_TQ; rr += 8) {   // one warp per row, 8 rows/warp
    int c0 = lane * 2, c1 = c0 + 1;
    float s0 = v0 ? Sf[rr * FA_TKV + c0] : -INFINITY;   // read S back
    float m_new = fmaxf(s0, s1);
    for (int off = 16; off > 0; off >>= 1)              // full-warp 5-step reduction
        m_new = fmaxf(m_new, __shfl_xor_sync(0xffffffffu, m_new, off));
    float m_old = msh[rr];                              // state in smem too
    a = (m_old == -INFINITY) ? 0.0f : __expf(m_old - m_new);
    Pf[rr * FA_PSTR + c0] = __float2half(p0);           // P written back to smem
    ...
    if (lane == 0) { lsh[rr] = lsh[rr] * a + sum; msh[rr] = m_new; }
    if (lane == 0) alpha[rr] = a;                       // alpha broadcast back to smem
}
__syncthreads();   // P·V's ldmatrix readers live in other warps

After that, P·V also had to load_matrix_sync Pf back from smem — one trip there, one trip back; S and P were each moved twice.

New structure (current tree, src/cuda_kernels.cu; the kernel is unchanged since d38744d): the warp mapping and all state move into registers — each lane derives the two rows and column group it holds from the fragment layout:

// FAP2: 4 warps (128 threads), warp wm owns a full 16-query-row block x
// all FA_TKV KV columns. S lives in wmma accumulators and is softmaxed in
// place; P is converted f32->f16 into the A fragment of P@V — the S/P
// shared round trip and its bank conflicts are gone entirely.
const int warp = tid >> 5;     // 0..3
const int wm  = warp;          // 16-query-row block
const int lane = tid & 31;
const int l  = lane & 3;       // 4-lane row group
const int r0 = lane >> 2;      // fragment local row 0 (0..7)
const int r1 = r0 + 8;         // fragment local row 1
const int c0 = 2 * l;          // cols (2l, 2l+1, 2l+8, 2l+9) per fragment
...
wmma::fragment<wmma::accumulator, 16, 16, 16, float> acc[8];  // full hd=128 output
float m0 = -INFINITY, m1 = -INFINITY, l0 = 0.0f, l1 = 0.0f;   // state in registers

QK^T: K stays a col_major B (= the K^T semantics), and S stays directly in the accumulator fragment fc[]:

wmma::fragment<wmma::accumulator, 16, 16, 16, float> fc[FA_TKV / 16];
for (int d = 0; d < hd; d += 16) {
    wmma::fragment<wmma::matrix_a, 16, 16, 16, __half, wmma::row_major> fa;
    wmma::fragment<wmma::matrix_b, 16, 16, 16, __half, wmma::col_major> fb[FA_TKV / 16];
    for (int cc = 0; cc < FA_TKV / 16; cc++)
        wmma::load_matrix_sync(fb[cc], &Ks[cc * 16 * sstr + d], sstr);
    wmma::load_matrix_sync(fa, &Qs[wm * 16 * sstr + d], sstr);
    for (int cc = 0; cc < FA_TKV / 16; cc++)
        wmma::mma_sync(fc[cc], fa, fb[cc], fc[cc]);   // S stays resident in fc
}
__syncthreads();   // the old structure had its 2nd barrier here (cross-warp Sf reads) — no longer needed

Online softmax on the fragments: first linearize fc[].x[i] into (fragment, quad) arrays and record each element's global column number (masking has to be re-evaluated per element); the in-row reduction uses only the 4-lane butterfly:

float sm[FA_TKV / 16 * 4], sm1_[FA_TKV / 16 * 4];
int gcol[FA_TKV / 16 * 4];
for (int cc = 0; cc < FA_TKV / 16; cc++) {          // x[0,1,4,5]→row r0
    sm[quad+0] = fc[cc].x[0]; ... sm1_[quad+3] = fc[cc].x[7];
    gcol[quad+0] = kt + cc * 16 + c0;  ...           // global KV column numbers
}
for (int q = 0; q < FA_TKV / 16 * 4; q++) {          // causal + range masking
    bool v0 = (gcol[q] <= qpos0) && (gcol[q] < kv_end);
    if (v0) mnew0 = fmaxf(mnew0, sm[q]); ...
}
for (int off = 1; off <= 2; off <<= 1) {             // 4-lane row group, 2 steps
    mnew0 = fmaxf(mnew0, __shfl_xor_sync(0xffffffffu, mnew0, off));
    mnew1 = fmaxf(mnew1, __shfl_xor_sync(0xffffffffu, mnew1, off));
}
const int fresh0 = (m0 == -INFINITY);
float a0 = fresh0 ? 0.0f : __expf(m0 - mnew0);
if (mnew0 == -INFINITY) a0 = 1.0f;                   // fully masked tile: state untouched
...
for (int q = ...) { p0[q] = valid ? __expf(sm[q] - mnew0) : 0.0f; sum0 += p0[q]; }
for (int off = 1; off <= 2; off <<= 1) sum0 += __shfl_xor_sync(..., sum0, off);
if (mnew0 != -INFINITY) m0 = mnew0;
l0 = l0 * a0 + sum0;                                 // m/l update, no smem

Then the pivotal move of the whole piece: P is built in place as the A operand of P·V. The f32 accumulator and the f16 row_major matrix_a have identical per-lane layouts, so pa.x[i] = __float2half(p) writes element by element; the O accumulator is rescaled by per-row alpha (same layout: x[0,1,4,5]→row r0, x[2,3,6,7]→row r1); V for P·V is a row_major B:

// rescale O by per-row alpha (m16n16 f32 accumulator layout)
for (int ob = 0; ob < 8; ob++) {
    acc[ob].x[0] *= aa0; acc[ob].x[1] *= aa0;   // row r0
    acc[ob].x[2] *= aa1; acc[ob].x[3] *= aa1;   // row r1
    acc[ob].x[4] *= aa0; acc[ob].x[5] *= aa0;
    acc[ob].x[6] *= aa1; acc[ob].x[7] *= aa1;
}
// Build the P@V A-operand IN PLACE: matrix_a m16n16k16 row_major and the
// f32 accumulator use the SAME (row,col) layout → pa.x[i] = f(p_i), element by element.
wmma::fragment<wmma::matrix_a, 16, 16, 16, __half, wmma::row_major> pa[FA_TKV / 16];
for (int cc = 0; cc < FA_TKV / 16; cc++) {
    pa[cc].x[0] = __float2half(p0[quad + 0]);
    pa[cc].x[1] = __float2half(p0[quad + 1]);
    pa[cc].x[2] = __float2half(p1[quad + 0]);   // x[2,3] is row r1
    ...
}
// acc = acc*alpha + P·V; V (B) is row_major — K in QK^T is col_major,
// the two operands' major-ness is opposite (both row_major ⇒ both matmuls wrong).
for (int kk0 = 0; kk0 < FA_TKV; kk0 += 16)
    for (int ob = 0; ob < 8; ob++) {
        wmma::fragment<wmma::matrix_b, 16, 16, 16, __half, wmma::row_major> vb;
        wmma::load_matrix_sync(vb, &Vs[kk0 * sstr + ob * 16], sstr);
        wmma::mma_sync(acc[ob], pa[kk0 / 16], vb, acc[ob]);
    }
__syncthreads();   // 3 barriers per tile (was 4)

Launcher: the smem formula shrinks from "Qs+Ks+Vs+Sf/Pf" to pure staging, and the r46 fix's explicit opt-in and failure fallback are preserved (NOT silent):

// Qs + Ks + Vs only (S/P no longer go through shared memory). sstr = hd+8
// padding; Ks/Vs are FA_TKV rows (the r46 launcher's 3*FA_TQ bug is gone).
size_t smem = ((size_t)FA_TQ + 2 * FA_TKV) * (hd + 8) * 2;   // 128*136*2 = 34,816 B
...cudaFuncSetAttribute(..., MaxDynamicSharedMemorySize, (int)smem);  // failure → fall back to the old kernel
dim3 grid((nt + FA_TQ - 1) / FA_TQ, nh, 1);
fa_prefill_f16kv<<<grid, 128, smem, stream>>>(...);          // 256 → 128 threads

3.3 Pitfalls

  1. The double major-ness mistake (detailed in §2): the B operands of QK^T and P·V share the same physical layout but are semantic transposes of each other, so the flags must be one col_major and one row_major. The symptom is extremely misleading: the kernel "runs", and parity max err goes from the 1e-4 class to 0.54 — not a precision regression but wrong math. After the fix, both lane maps were validated standalone before integration.
  2. Halving the threads vs hardcoded strides: the staging loop's stride was originally written for 256 threads; changing the launch to 128 without parameterizing the stride (fa_stage_kv_async(..., tid, 128), Q loads i += 128) means cp.async moves only half of each K/V tile and the other half is uninitialized smem — silently wrong data. r46's launcher over-allocation was the same class of "parameters didn't follow the geometry" trap.
  3. Fragment masks must be recomputed per element: in the old row loop a lane's column numbers were contiguous c0/c1; in the new layout one lane's 32 columns are scattered as (2l, 2l+1, 2l+8, 2l+9) per fragment, so the causal mask and the kv_end range mask must be evaluated element by element on each gcol[q], and a fully masked tile keeps the −INF state (alpha=1, P=0, l untouched) — the semantics were aligned branch by branch with the old implementation, and greedy-32 byte-identical confirms it.

4. Verification

GateNumbersWhat it defends against
cuda_fa_prefill_attention_parity (src/graph/cuda_backend.rs)1/0 greenpoint-by-point comparison of a small GQA graph (nh=4, nk=2, hd=128, nt=100, f16 KV) against the reference — pins both lane maps (the accumulator rescale mapping + P's matrix_a identity) and the major-ness combination, catching "it runs but the math is wrong" (the major-ness bug's 0.54 is exactly what it caught)
cuda_prefill7/0 greenwhole-graph prefill numeric regression, catching the kernel swap changing downstream nodes (rms/rope/FFN chains) in a real layer sequence
cuda_prefill_mmq_parity1/0 greennumerics when the MMQ GEMM path coexists with FA in the same graph, catching breakage from attention-side changes coupling into the GEMM side
greedy-32 byte-identicalbyte-for-bytecatches ULP-level reordering drifting into the sampler; r48's rewrite preserved it (r50/r57 later proved FA tile-size changes do not always preserve it — see doc 53)
suite166/0/3full regression (CPU/Metal/graph paths), catching CUDA-layer changes leaking into shared code
ncu reconciliationMIO stall zeroed, shared wavefronts 44→21%, 124 regs/0 spillcatches "faster but mechanism unexplained" — confirms the gain really comes from deleting the S/P round trip, and that 5.16 → 2.12 ms clears the preset ≤2.6 ms target line

5. Results

LevelbeforeafterΔ
FA kernel (3325-tok prefill, ncu)5.16 ms2.12 ms2.43×
whole-prefill (same-window interleaved ×3, median)2603.5 tok/s2749.9 tok/s+5.6%
FA share of the wall10.2%~4.7%halved
vs-llama (3325-eq anchor)1.27×1.21×−0.06×
smem / block69.38 KB (1 block/SM)34.82 KB (2 blocks/SM)−50%
barriers per tile / smem S/P traffic4 / ~56 KB3 / 0round trip deleted

r47's prediction was 2× → −63 ms (−4.9% of the wall); measured, FA saved ~68 ms and the wall gained +5.6% — the prediction landed and slightly exceeded. whole-prefill's +5.6% is far past the +1.5% bar, and this is the campaign's first case of a structural ">2× residual" being removed wholesale.

The same S/P defect, two fixes — the gap between the final numbers is exactly the gap between "working around" and "deleting":

Routesmemblocks/SMS/P round tripkernelwall
r46 (FAP1): FA_TKV 64→32 + row padding69.38 → 43.8 KB1 → 2kept (conflict reduced)5.16 → 4.58 ms (−11%)+0.27% (REVERTED)
r48 (FAP2): full-row warp tile + register softmax69.38 → 34.82 KB1 → 2deleted5.16 → 2.12 ms (2.43×)+5.6% (LANDED)

Follow-on coordinates: r49 trimmed the prepass further (+2.32%); the FA line was not touched again until r50/r57 — both times it hit the greedy-identity wall on tile size, and r48's geometry remains FA's stable chassis (r50's launcher lesson — the 32 KB of tail-tile O write-back smem reuse must be counted in the launcher's max — is the sequel to the last item in this doc's §3.1).

6. Lessons

  1. Deleting a round trip beats optimizing around it: r46's padding on the S/P round trip bought kernel −11%; r48 deleting the round trip outright is 2.43×. Rounds where the mechanism is positive but the wall is negative (r46) are usually patches on a structure that should have been deleted.
  2. Register residency is a geometric property: whether softmax can stay in registers depends on whether a warp exclusively owns complete data rows — change the warp×tile geometry first, then talk about "keeping data out of memory." Keeping the data layout and only swapping load/store instructions never yields this magnitude.
  3. Judge wmma operand major-ness per matmul: the same physical memory is a col_major B in QK^T and a row_major B in P·V — the two consumption directions are mutual transposes. Validate lane maps standalone before integrating; "the kernel runs" is not evidence, parity numbers are.
  4. Halving threads requires auditing every loop unrolled by thread count: constants that "follow the launch geometry" like staging strides, once hardcoded, turn a thread reduction into silently dropping half the data.

← 50 · r47 converged-regime wall decomposition · Index · 52 →

52 · r49 — A-quantize prepass shared-A dedup: consecutive-window memoization (LANDED)

Result: the q/k/v attention GEMMs consume the same normed activation and gate/up consume the same normed2, but the MMQ path's A-quantize prepass previously re-ran once per matmul (193 launches on a 3325-tok prefill). CudaState gains a consecutive-window cache keyed on (src device pointer, nt, id): on a hit the quantize_q8_0_pad40_t launch is skipped entirely. prepass 193 → 110 launches (118.4 → 83.9 ms, 9.6% → 7.4% of GPU busy), GEMM launches constant at 193 (only the redundant prepass is deleted, not the GEMMs); whole-prefill 2734.1 → 2797.5 tok/s (+2.32%), vs-llama 1.21× → 1.18×. Hits are byte-identical by construction (quantize is a pure function of (x, nt, id)); parity ×3 green, greedy-32 byte-identical, suite 166/0/3. Commit: 87a75a3 (src/cuda.rs + src/graph/cuda_backend.rs; graph.rs untouched — a hard rule). Date: 2026-09-06.

1. Background — where things stood

After r48 cashed in the first item of r47's priority queue (FAP2) at +5.6%, the second item in the queue was the prepass. This thread had been hanging in the campaign for three rounds:

  • r34 introduced the quantize-transpose prepass itself (quantize_q8_0_pad40_t, see doc 37): moving the A-side layout transform out of the kernel, +9.72% — pure profit at the time;
  • r37's attribution annotated it: "Quantize prepass 86.8 ms (1.25× llama, partly 2× per-shared-A redundancy)" — part of the 25% over llama was exactly repeated quantization;
  • r47's re-decomposition pushed it to the front: the prepass had grown to 118.4 ms = 9.6% of GPU busy (a +31.6 ms "hidden tax": the q6_K BT port wired the A sides of attn_v/ffn_down onto the same prepass), the third-largest slice of the wall, with a clear redundancy mechanism.

The redundancy's source is a graph-topology fact, not a kernel defect. Qwen2's 7 prefill matmuls per layer consume 4 distinct A inputs:

matmulA sourceshared?
q, k, vnormed (attn rms output)3 GEMMs share 1 copy
gate, upnormed2 (ffn rms output)2 GEMMs share 1 copy
attn_oattention outputexclusive
downswiglu outputexclusive

The builder unrolls by dataflow, so normed is referenced once by each of the three matmul nodes; the MMQ path unconditionally re-ran the prepass before every matmul — of the 7 quantizations per layer, 3 were recomputations (the last two of q/k/v + the last one of gate/up). Across 28 layers that is ~84 redundant launches, all of them the smallest, fastest kernels, and their launch overhead + repeated reads of x genuinely occupied 9.6% of GPU busy.

What one prepass launch actually does (r34's legacy, see doc 37): read the f32 activations x (nt × id floats), compute the amax per 32-element block, rintf-scale into q8_0, zero-pad to pad40, and produce two planes — the native form q8 (nt × (id/32) × 40 B, read block by block by the fallback kernels) and the transposed form qa8 [ntb][nchunk][2048] + sda [ntb][nchunk][256] (ntb = ceil(nt/64), nchunk = id/32, used by the BT/NB kernels for bulk staging). The cost is proportional to nt × id bytes plus a launch's fixed overhead; on the 3325-token, d=3584 shape one such launch amortizes to ~0.6 ms (118.4 ms / 193). The price of doing it three times is not compute — it is pure repeated data movement and launch queuing.

One easily underestimated property of this post-r48 residual: it modifies no math — the quantized bytes are identical, and the only issue is that "the same bytes were computed three times." All of the risk in such a change lives in cache freshness: when may it hit, when must it be invalidated.

2. Principle — the GPU mechanism

Purity is the linchpin. quantize_q8_0_pad40_t(x, qa8, sda, id, nt, ...)'s output is a pure function of (x, nt, id) — per-block amax, rintf scaling, pad40 zero fill, no cross-call state whatsoever. So "a second request with the same (src pointer, nt, id)" and "re-running it" are indistinguishable at the byte level, and a hit is equivalent. This collapses the correctness question into one: does an identical key guarantee identical input?

The semantic gap of the pointer key. The key is (src device pointer, nt, id) rather than a buffer id or node id because the state layer only sees raw pointers. But the graph allocator (liveness allocator) reuses pool buffer ids between nodes: normed's buffer may be reallocated as some other node's output after q is done with it, and later a fresh normed may land on the same device address. Same pointer ≠ same data — the key alone cannot distinguish "a second consumption of the same A" from "new data landed in an old address."

Conservative invalidation rules close the gap. r49's approach is to do no write tracking at all, and instead tighten the window with two rules:

  1. Valid only across "consecutive MatMul nodes": any non-MatMul node clears the cache when it executes. The graph executes in build order (topological order), so matmuls sharing an A are naturally adjacent in the builder's output; within the window no buffer can be rewritten (no other node runs). A late buffer-id reuse necessarily happens at some non-MatMul node — and that node has already cleared the cache.
  2. Clear at split boundaries / per execution: cleared at synchronize. The next graph execution hands the same pool buffer ids to different data, and the (src, nt, id) key reproduces verbatim — residue across executions is the most dangerous kind of false hit.

Neither exception chases "one more hit"; both only chase "impossible to get wrong." The cost is real: non-adjacent same-A consumers (which do not exist in this campaign's graphs) are collateral-damaged into misses — the loss runs in the conservative direction, costing performance, not correctness.

Three rejected alternatives in the design space, recorded here so nobody "optimizes" back into them:

  • A cache persisting across executions (keeping the planes by (ptr, nt, id) into the next execution): steps directly on the split-boundary problem — the same pool buffer id holds different data next execution, the key reproduces verbatim, and the false hit is a silent error. Infeasible unless content fingerprints are introduced (hash verification means reading the planes back, which costs more than recomputing);
  • Write-tracking invalidation (hooking every allocator write, invalidating a cached pointer when written): requires the allocator to expose write events and welds cache semantics into a module that has nothing to do with it (alloc.rs is shared by all backends) — high complexity, hard to prove, all for "a slightly larger window";
  • Adding buffer id to the key: the state layer cannot get the id (it sees raw device pointers), and threading the id down means changing the matmul dispatch signature — which violates the "graph.rs / dispatch layer untouched" hard rule.

Two lines of backend code bought all the correctness the first two alternatives were chasing — that is the engineering meaning of the "scheduling-window property" judgment.

Why this is a "scheduling-window" problem. The redundancy is not a property of the kernel (quantize_q8_0_pad40_t is beyond reproach) but a property of the node stream — the same input got scheduled three times inside a window. So the cache must take the form of "a window memo attached to the node stream," not "a constant cache attached to weights/model": it starts from zero on every execution, grows with the node stream, and is invalidated when the window closes.

Why the scratch must live outside the pool. The cache's output planes qa8/sda live in CudaState's dedicated scratch (buf_qa8_t/buf_sda_t transposed, buf_q8_prefill native) — outside the graph allocator's pool. Inside the pool, the allocator could hand this "cache" to some node as its output, and the next layer writing it would clobber the cache contents; out-of-pool scratch guarantees no node output ever aliases the cached planes.

3. Implementation

3.1 Design choices (why this shape and not another)

  • The cache lives in CudaState (backend layer), graph.rs untouched — this round's hard rule. Window semantics are the execution layer's scheduling knowledge; the graph-construction layer (shared by all backends) should not know whether some backend memoizes. graph.rs untouched also means zero risk to the CPU/Metal paths.
  • Key (src ptr, nt, id), with a physical-pointer check after a key hit: get_or_grow reallocs on a larger miss, so a cached plane's old address may be stale — the hit condition requires not just key equality but that the cached record's qa8/sda pointers match the actual pointers after growth. nt and id must be in the key: the same A pointer quantizes differently under different graph shapes (different nt) or different GEMM dimensions — the pure function's argument is the triple, not the bare pointer.
  • Both quantization forms go into the cache: transposed pad40_t (qa8 [ntb][nchunk][2048] + sda [ntb][nchunk][256], for the BT/NB kernels) and native pad40 (nt × id/32 × 40 B, for the fallback kernels), distinguished by the transposed flag — the two forms' plane contents differ, and the cache must not mix them.
  • No new env gate: the dedup rides the MINFER_MMQ path, behavior is byte-level identical to not caching, and there is no semantic switch worth A/B-ing.
  • OOM front-loading: the native plane's buffer is get_or_grow-ed once before the GEMMs, so an allocation failure errors at the matmul entry instead of landing on the q4_K fallback kernel's final launch as a null dereference.

3.2 Key code

The cache itself (src/cuda.rs, current tree; the dead_write field is r52's skip-write guard, not present in r49's original). Four invariants written into the struct's doc comment first — the code is just their implementation:

  1. Valid only across consecutive prefill-MMQ MatMul nodes (any other node clears it);
  2. Cleared at split boundaries / per execution (no leakage across executions);
  3. The cached planes live in dedicated out-of-pool scratch (never alias node outputs);
  4. The quantize output is a pure function of (src, nt, id) (a hit ≡ a recompute, byte-identical).
#![allow(unused)]
fn main() {
/// r49: consecutive-window memoization of the MMQ A-quantize prepass.
/// Correctness relies on two rules (both conservative, no write-tracking):
///   * It is valid ONLY across CONSECUTIVE prefill-MMQ MatMul nodes. Any other
///     node kind clears it (see `CudaBackend::execute_node_inner`), so a late
///     buffer-id reuse by the liveness allocator can never alias the cached A.
///   * It is cleared at split boundaries (`CudaBackend::synchronize`) so a
///     cache from a previous graph execution never leaks stale data into a
///     later one (the same pool buffer id holds different data each step).
/// The buffers are the state-level buf_qa8_t/buf_sda_t (transposed) and
/// buf_q8_prefill (native) scratch — dedicated allocations OUTSIDE the
/// graph allocator pool ... The quantize output is a pure function of
/// (src, nt, id), so a hit is byte-identical to a recompute.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
struct MmqCache {
    active: bool,                     // whether an entry is recorded
    key: (usize, usize, usize),       // (src device pointer, nt, id)
    transposed: bool,                 // true: qa8_t/sda_t planes; false: native q8
    dead_write: bool,                 // r52: mode-2 skip-write entry (later round)
    qa8: usize, sda: usize, q8: usize, // physical addresses of the cached planes (validated on hit)
}
}

The complete hit/miss path (mmq_quantize_transposed, hit validation + miss recompute + record; r52's mid-section dead-write guard omitted):

#![allow(unused)]
fn main() {
let need_qa8 = (ntb as usize) * (id as usize / 32) * 2048;
let need_sda = (ntb as usize) * (id as usize / 32) * 256;
let key = (x as usize, nt as usize, id as usize);
let mut cache = self.mmq_cache.lock().unwrap();
if cache.active && cache.key == key && cache.transposed {
    // get_or_grow may have reallocated on a larger miss: validate the
    // physical pointers so a grown buffer is never reused stale.
    let qa8 = Self::get_or_grow(&self.buf_qa8_t, need_qa8) as usize;
    let sda = Self::get_or_grow(&self.buf_sda_t, need_sda) as usize;
    if qa8 == cache.qa8 && sda == cache.sda {
        return (qa8, sda);            // HIT: zero launches, the pointers ARE the last quantize result
    }
}
... (r52's dead-write rejection path)
let qa8 = Self::get_or_grow(&self.buf_qa8_t, need_qa8);
let sda = Self::get_or_grow(&self.buf_sda_t, need_sda);
launch_quantize_q8_0_pad40_t(
    x,
    qa8 as *mut u8,
    sda as *mut u8,
    id,
    nt,
    nchunk,
    ntb,
    stream,
);                                    // MISS: recompute verbatim
cache.active = true;
cache.key = key;
cache.transposed = true;
cache.qa8 = qa8 as usize;
cache.sda = sda as usize;
cache.q8 = 0;                         // r52 also clears the dead_write flag here
}

The symmetric implementation for the native (non-transposed) form — the pad40 plane (buf_q8_prefill) used by the fallback NB/wide/narrow kernels, same window rules, with the transposed=false gate keeping the two forms from colliding:

#![allow(unused)]
fn main() {
let need = (nt as usize) * (id as usize / 32) * 40;   // native pad40 plane size
let key = (x as usize, id as usize, nt as usize);
let mut cache = self.mmq_cache.lock().unwrap();
if cache.active && cache.key == key && !cache.transposed {
    let q8 = Self::get_or_grow(&self.buf_q8_prefill, need) as usize;
    if q8 == cache.q8 {
        return q8;                    // HIT
    }
}
... (r52's dead-write rejection path)
let q8 = Self::get_or_grow(&self.buf_q8_prefill, need);
launch_quantize_q8_0_pad40(x, q8 as *mut u8, id, nt, stream);   // MISS
cache.active = true;
cache.key = key;
cache.transposed = false;
cache.q8 = q8 as usize;
}

The consumer side is untouched — the matmul dispatch just feeds the returned (qa8, sda) pointers to the GEMM launcher (q6_K BT path, src/cuda.rs):

#![allow(unused)]
fn main() {
// r49: A-quantize prepass via the consecutive-window cache —
// a same-A (q/k/v, gate/up) matmul reuses qa8g/sdag without a
// fresh quantize launch.
let (qa8g, sdag) = self.mmq_quantize_transposed(
    x as *const f32, id as i32, nt as i32, nchunk, ntb, stream,
);
}

The two anchors of window invalidation (all of 87a75a3's changes to src/graph/cuda_backend.rs, 13 lines total):

#![allow(unused)]
fn main() {
// execute_node_inner: any non-MatMul node clears it (conservative — no write tracking)
if !matches!(&node.op, Op::MatMul { .. }) {
    self.state.clear_mmq_cache();
}
}
#![allow(unused)]
fn main() {
// synchronize: the cache never leaks across graph executions (the same pool
// buffer id holds different data on the next execution)
self.state.clear_mmq_cache();
}

(Later rounds added Op::FusedFFN to the preserve set — its input plane was just recorded by the fused rms epilogue and its internal gu matmul is the first consumer, so the window semantics match MatMul→MatMul; that was D3-5's business — r49's rules are exactly the two above.)

OOM front-loading (the matmul entry first guarantees the native plane is allocatable):

#![allow(unused)]
fn main() {
// r49: the native pad40 buffer is sized upfront so an OOM surfaces here
// (before any GEMM launch) instead of as a null deref in the q4_K
// fallback's final `launch_mmq_raw_nt`.
if Self::get_or_grow(&self.buf_q8_prefill, nt * (id / 32) * 40).is_null() {
    return Err("cuda: prefill MMQ q8 scratch OOM".to_string());
}
}

3.3 Pitfalls

  1. Parity cannot reach the HIT path. The parity harness clears the cache between cases, so all three parity tests exercise only the miss path — the most correctness-critical part of caching (hit equivalence) is exactly what the unit gates cannot reach. r49's solution was to hand the HIT path to greedy-32 byte-identical: in a 32-step greedy generation every layer's q/k/v and gate/up go through the hit path, and any false hit shows up in the byte stream. Lesson: pair a "cannot-be-wrong" construction with an end-to-end identity gate that can actually reach it.
  2. Hit validation must check the physical pointers, not just the key. get_or_grow's realloc semantics mean that with an identical key the cached plane may have been moved/grown — a key hit + matching physical pointers is a real hit. Skipping this layer reads a stale address in a "small prefill first, then a large prefill" session.
  3. Launch counts and milliseconds are two different ledgers. The 83 deleted launches were all narrow width-3584 planes (the duplicate copies of normed/normed2); the 110 kept ones include 28 wide width-18944 planes (swiglu output) — launches dropped to 57%, duration only to 71% (118.4 → 83.9 ms). Estimating the time win from the launch share would overestimate by nearly half (predicted −50.9 ms vs actual −34.5 ms): the deleted launches happened to be the smallest batch.
  4. Do not force-fit the theoretical account to single digits. 4 distinct A copies per layer × 28 layers = 112 expected surviving launches; 110 measured — the 2-launch gap comes from graph-level tails (the lm_head input after the final norm, etc.) and a few layers' shape differences. Argue the mechanism with ratios (0.571 predicted vs 0.570 measured; time-weighted 0.734 vs 0.709 measured), not by reconciling launch by launch — failing to match single digits is not evidence of a bug, it is the limit of model granularity.

4. Verification

GateNumbersWhat it defends against
launch census (nsys)prepass 193 → 110, GEMM constant at 193the mechanism gate — proves only the redundant prepass was deleted and the GEMM side is untouched (if the GEMM count had changed too, real work was deleted by mistake)
cuda_prefill_mmq_parity1/0 greenMMQ numeric regression (exercises the miss path — the parity framework clears the cache between cases), catching a broken recompute path
cuda_prefill7/0 greenwhole-graph prefill numerics, catching any downstream drift introduced by the cache (also miss-only)
cuda_fa_prefill_attention_parity1/0 greenFA and MMQ coexisting in one graph, catching this round's changes spilling onto r48's freshly landed attention
greedy-32 byte-identicalbyte-for-bytethe HIT path's end-to-end identity gate — every layer's q/k/v, gate/up hits many times; any false hit or staleness hole shows up in the byte stream (the path parity cannot reach, see pitfall 1 in 3.3)
suite166/0/3full regression, including the CPU path — structural corroboration that graph.rs was untouched
A/B interleaved ×3 median2734.1 → 2797.5, distributions fully separated+2.32% clears the +1.5% bar, catching machine-state noise being read as a gain

5. Results

MetricbeforeafterΔ
prepass launches (3325-tok prefill)193110−83 (−43%)
prepass duration118.4 ms (9.6% busy)83.9 ms (7.4%)−34.5 ms
GEMM launches193193unchanged
whole-prefill (same-window interleaved ×3 median)2734.1 tok/s2797.5 tok/s+2.32%
vs-llama (3325-eq anchor)1.21×1.18×−0.03×

The mechanism's two-ratio cross-check (7 matmuls / 4 distinct A copies per layer):

  • launch ratio: theoretical 4/7 = 0.571 → 193 × 0.571 ≈ 110.3, measured 110;
  • duration ratio: linearly weighted by plane width, the surviving share = (3×3584 + 18944)/(6×3584 + 18944) = 0.734 → 118.4 × 0.734 ≈ 86.9 ms, measured 83.9 ms (−3%, made up by the omitted fixed launch overhead).

Both independent ratios lock onto the measured values — what was deleted really is the single thing "3 duplicate A copies per layer," with no other effect mixed in. The wall clock realized ~28 ms (the wall change corresponding to +2.32%), slightly less than the 34.5 ms of kernel-time savings — the typical discount when launch-level savings cash into wall time (the same pattern recurs in r51: prepass 83.0 → 10.1 ms bought wall +1.89%).

Follow-on coordinates: the prepass line ran to its end in r51/r52 — producer fusion cut launches from 110 to 28 (prepass 83.0 → 10.1 ms), and skip-write mode then waived the f32 output writes entirely (fused producers 151.9 → 86.5 ms); r49's MmqCache and its window rules became the foundation of both rounds verbatim (the fused producers record their quantized planes straight into the same cache). This cache later grew a second job as well:

Later roundIncrement on MmqCacheNumbers
r51 (doc 54)fused rms/swiglu records directly into the cache (record_mmq_cache_transposed), leaving the prepass only woprepass 110 → 28 launches, 83.0 → 10.1 ms; +1.89%
r52 (doc 55)mode-2 skip-write: the f32 output write is waived, and the dead_write guard turns window violations into loud errorsfused producers 151.9 → 86.5 ms; +5.45%
D3-5 1athe decode-side isomorph (record_mmq_cache_native, MmqCache consult skipping the standalone quantize)standalone quantize 4448 → 964 launches (−78%)

A small change that began as "don't quantize the same A three times" eventually grew into the backbone of the A-quantize supply line on both the prefill and decode sides — the conservative window semantics (key + two invalidation rules) did not change by a single word from r49 through the D series; that is the compound interest of "invariants first."

6. Lessons

  1. Redundant work on shared inputs is a scheduling-window property: memoize on the node stream — the key (pointer+shape) buys the fast path, and conservative window invalidation (any foreign node clears + execution-boundary clears) buys correctness; no write tracking, and prefer a missed hit over a wrong one.
  2. Swap the gate when the unit harness cannot reach the path: cache-hit equivalence cannot enter the parity framework that clears caches, so end-to-end greedy byte-for-byte identity is the only verification that reaches it — assign gates by "which gate can exercise which path," not by piling gates on.
  3. The backend layer's scheduling knowledge stays out of the graph-construction layer: zero changes to graph.rs made the dedup structurally risk-free for the CPU/Metal paths; window semantics are the executor's private property.
  4. Launch count and duration are two different profit ledgers: what gets deleted tends to be the smallest launches (shared narrow planes), and what stays tends to be the wide planes — weight estimates by bytes, not by launch counts.

← 51 · r48 FAP2 register-resident softmax · Index · 53 →

53 · r50 — FA_TKV 32→16 occupancy experiment (REVERTED)

Result: fa_prefill_f16kv's KV tile shrinks from 32 columns to 16, smem drops from 34.8 KB to ~33 KB, resident blocks 2 → 3 blocks/SM — parity green, suite 166/0/3, but the whole-prefill wall clock is neutral-to-negative (two rounds of interleaved A/B: 3-pair median −0.5%, 5-pair median −0.01%); more fatally, greedy-32 is no longer byte-identical (an argmax boundary flipped by ULP-level float regrouping). The occupancy gain is cancelled by the doubled per-tile sync/softmax overhead — tile shrinking on FA is a dead lever, and it inherently breaks byte identity. Commit: 9128468 (docs-only — the code change was reverted after measurement and never landed as a code commit). Date: 2026-09-06.

1. Background — where things stood

r48 (FAP2, register-resident softmax) was the heaviest blow the FA line has taken so far: S/P no longer pass through shared memory, softmax runs directly on the QK^T wmma accumulator fragments, and P becomes the A operand of P·V in registers, in place. The FA kernel dropped from 5.16 ms to 2.12 ms (2.43×), whole prefill 2603.5 → 2749.9 tok/s, and FA's share of the whole wall shrank from 10.2% to ~4.7%. r49 (A-quantize prepass shared-A dedup) took another +2.32%, pushing the baseline to 2797.5 tok/s and vs-llama to 1.18×.

r48's record left the FA line one explicit "next lever": after FAP2 the kernel's top stall became global K/V loads, and r48's ncu device report showed the resource limiting resident blocks is smem (34.82 KB per block → 2 blocks/SM; 124 registers only limits to 4 blocks). Cutting FA_TKV from 32 to 16 halves the K/V staging smem demand linearly, so resident blocks could theoretically reach 3 — less per-byte latency exposure, the classic occupancy logic.

This thread had appeared twice before, and both times it ended as "the kernel moved, the wall didn't":

RoundChangeKernelWallOutcome
r23 (f16 era)FA_TKV raisedocc 16.7 → 32.68%, kernel −6.7%−0.3%REVERTED
r46 (FAP1)FA_TKV 64→32 + S/P row padding5.16 → 4.58 ms (−11%)+0.27%REVERTED
r50 (this doc)FA_TKV 32→16— (not timed separately)−0.5% / −0.01%REVERTED

Three rounds of same-direction evidence were already enough to constitute a mechanism judgment, but r50 was still worth doing, for two reasons: first, it is a one-#define experiment (r48 symbolized every geometric quantity, so the marginal cost was near zero); second, r46's veto reason was "FA is not yet the wall's critical path," whereas after r48 FA is down to 4.7% — "does occupancy finally pay off at a smaller share" had never been directly answered. The answer is no, and the way it was answered is worth more than the numbers: this round hit two boundaries at once — occupancy is ineffective on FA (the mechanism boundary), and the strict greedy byte-identity gate is unsatisfiable by any FA_TKV change (the verification boundary, which r57 would hit again).

2. Principle — the GPU mechanism

The kernel's shape after FAP2. fa_prefill_f16kv is a 4-warp (128-thread) full-row warp-tile: each warp owns 16 query rows × all FA_TKV KV columns. Each KV tile's lifecycle is:

  1. fa_stage_kv_async: cp.async moves the K/V tile into smem (row stride sstr = hd+8, zero-filled out of range);
  2. cp.async.wait_group 0 + __syncthreads();
  3. QK^T: 8 wmma steps over the hd dimension, producing fc[FA_TKV/16] 16×16 f32 accumulator fragments;
  4. second __syncthreads();
  5. fragment-resident online softmax: per tile, a max reduction (2 __shfl_xor), __expf, a sum reduction (2 __shfl_xor), the running m/l update, and 8 O accumulator fragments × alpha rescale;
  6. P converted to f16 in place as the A operand, P·V runs wmma, and a third __syncthreads() closes the tile.

Why halving the tile does not speed it up. The key is splitting the cost per KV column into two parts:

  • wmma arithmetic: the number of QK^T and P·V mmas per column is independent of FA_TKV (the total is fixed);
  • per-tile fixed costs: 3 __syncthreads, cp.async commit/wait, the softmax shfl reduction chains, the O-fragment rescale — these are billed per tile.

FA_TKV 32→16 doubles the tile count, so the second class of cost doubles exactly. And the occupancy gain? Occupancy solves "latency hiding": more resident warps let DRAM latency be filled by other warps' issue. But after r48 the kernel's K/V traffic is not that large (7B @3325 tok: per layer KV = 2 × 3325 × 128 × 2 B × 4 kv-head ≈ 6.8 MB, 28 layers ≈ 191 MB), and 2 blocks/SM already hides the latency; a 3rd block only sends more warps to contend for the same L1TEX/DRAM bandwidth and slices each block's L2 working set smaller. Gain ≈ 0, cost = doubled fixed costs → net effect inside the noise band. This is exactly r46's lesson repeating verbatim: occupancy gain cancelled by doubled per-tile sync/softmax overhead.

The occupancy account (derived from r48's ncu device report). GB10 (CC 12.1) has 100 KB of smem available per SM:

FA_TKVdynamic smem per blocksmem limitregister limit (124 regs)actual residency
32 (current tree)(64+64)×136×2 = 34.8 KB⌊100/34.8⌋ = 242
16 (experiment)max(formula, 32 KB) ≈ 33.0 KB⌊100/33.0⌋ = 343

The record notes: this 3 blocks/SM was derived from r48's device configuration — ncu/nsys could not be re-run in that window — but it is self-consistent (34.8→2 and 33.0→3 straddle the integer-division boundary exactly), and the wall-clock result proves that even with 3 blocks truly in hand it would not have been worth it.

Why byte identity is necessarily lost. Online softmax's math is independent of tile partitioning at infinite precision, but f32 is not infinite precision. The running state carries across tiles:

l0 = l0 * a0 + sum0;      // once per tile: both l and O are recomputed under the current tile's grouping
acc[ob].x[j] *= aa0;      // O accumulators rescaled by alpha, 8 fragments × 8 lanes per tile

Doubling the tile count regroups both the f32 multiply-add chain of l and the rescale chain of acc: the summation order of Σexp and the associativity boundaries of l·a + s all change → each number drifts by ~1 ULP. The vast majority of tokens have top-2 logit gaps far larger than a ULP and argmax is unaffected; but greedy sampling is pure argmax, so if any single token happens to stand on a knife edge where the top-2 gap is ~1e-7, the generation stream forks there and never looks back. The parity gate cannot see it: parity fixtures compare numeric error within tolerance (measured max err ~1e-4, the numbers themselves correct) and never do a whole-stream byte comparison. The only gate that can see it is the greedy byte-identity gate.

3. Implementation

Archival status (flagged up front): r50's code change was a one-line experiment in the working tree that disappeared with the revert once measurement was done — there is no code commit; what survives is only the docs commit 9128468 (the record + all numbers). This doc's "before" excerpts therefore come from the current tree (i.e., the post-revert FA_TKV=32 state, the code still running today); the "after" side has only the formulas and a one-line diff narration from the record — no code to excerpt.

3.1 Design choices (why this shape and not another)

The experiment was deliberately kept minimal: not one character changed except #define FA_TKV 32 → 16. r48's rewrite had already symbolized every geometric quantity in the kernel — the QK^T fragment array fc[FA_TKV/16], P·V's k-loop bound kk0 < FA_TKV, cp.async staging's count c < FA_TKV*hd/8, the softmax fragment arrays [FA_TKV/16*4], Vs = Ks + FA_TKV*sstr, grid/launch — so a one-line define is a complete re-parameterization. This is the implicit asset r48 left behind: symbolized geometry makes tile size an enumerable one-dimensional experiment.

The change immediately flushed out one non-symbolized constant that had slipped through: the tail tile's O write-back reuses smem as a 64×128 f32 stage buffer — FA_TQ * hd * 4 = 32768 B = 32 KB, independent of FA_TKV. At FA_TKV=32 the staging formula yields 34.8 KB ≥ 32 KB and the constant is masked; at 16 the formula yields 25.5 KB < 32 KB and the tail block overruns directly. So the launcher must take the max of the two:

smem = max((FA_TQ + 2*FA_TKV)*(hd+8)*2,   // Qs + Ks + Vs staging
           FA_TQ * hd * 4)                 // tail tile's O write-back stage (f32)
     = 32 KB   (at FA_TKV=16)

3.2 Key code

The symbolized geometry itself (current tree src/cuda_kernels.cu, r48's legacy — the r50 experiment only worked as a one-line change because of it):

// S and P live entirely in wmma accumulator fragments (FAP2 register-resident
// softmax — NO Sf/Pf shared round trip, no m/l/alpha shared arrays). Each warp
// owns a full 16-query-row block x all FA_TKV KV columns, so the online softmax
// (per-row max/sum on the fragments) and the P·V contraction (build the f16
// A-operand from the scaled fragments in place) are both warp-local.
#define FA_TQ 64
#define FA_TKV 32            // r50 experiment: 32 → 16 (one-line change, later reverted)

The tile loop's symbolized consumption of FA_TKV (current tree lines 4045–4058) — tile count = kv_end / FA_TKV, so halving the tile doubles the iteration count:

for (int kt = 0; kt < kv_end; kt += FA_TKV) {
    // stage K/V tile (padded stride, zero-filled beyond kv_end)
    fa_stage_kv_async(k, v, Ks, Vs, kt, kv_end, hk, hd, stride_kv, sstr, tid, 128);
    ...
    __syncthreads();                      // ← per-tile fixed cost #1
    // S = Q · K^T via wmma (reduction over hd), this warp's full row block.
    // fc[0] = kv cols [0,16), fc[1] = [16, FA_TKV) of this tile.
    wmma::fragment<wmma::accumulator, 16, 16, 16, float> fc[FA_TKV / 16];
    ...
    __syncthreads();                      // ← per-tile fixed cost #2

The online softmax's per-tile reduction chains (current tree lines 4099–4123) — 2 max-shfl + 2 sum-shfl run once per tile, the bulk of the "billed per tile" fixed costs:

#pragma unroll
for (int off = 1; off <= 2; off <<= 1) {
    mnew0 = fmaxf(mnew0, __shfl_xor_sync(0xffffffffu, mnew0, off));  // max tree
    mnew1 = fmaxf(mnew1, __shfl_xor_sync(0xffffffffu, mnew1, off));
}
...
#pragma unroll
for (int off = 1; off <= 2; off <<= 1) {
    sum0 += __shfl_xor_sync(0xffffffffu, sum0, off);                 // sum tree
    sum1 += __shfl_xor_sync(0xffffffffu, sum1, off);
}
if (mnew0 != -INFINITY) m0 = mnew0;
l0 = l0 * a0 + sum0; l1 = l1 * a1 + sum1;      // ← the once-per-tile f32 regrouping

The O-fragment rescale (current tree lines 4132–4138) — this is the statement that breaks byte identity, forming the cross-tile f32 state chain together with l0 = l0*a0 + sum0:

#pragma unroll
for (int ob = 0; ob < 8; ob++) {
    acc[ob].x[0] *= aa0; acc[ob].x[1] *= aa0;  // x[0,1,4,5] → row r0
    acc[ob].x[2] *= aa1; acc[ob].x[3] *= aa1;  // x[2,3,6,7] → row r1
    acc[ob].x[4] *= aa0; acc[ob].x[5] *= aa0;
    acc[ob].x[6] *= aa1; acc[ob].x[7] *= aa1;
}

The tail tile's O write-back reusing smem as the 32 KB stage (current tree lines 4186–4189) — the non-symbolized constant r50 flushed out:

} else {
    // Qs/Ks/Vs regions are free after the KV loop: contiguous staging for
    // the 64x128 f32 O tile (32 KB < the 34.8 KB smem budget).
    float* stage = reinterpret_cast<float*>(smem);

The launcher (current tree lines 4216–4218) — after the revert only the first term is needed; during the r50 experiment it was max(first term, FA_TQ*hd*4):

// Qs + Ks + Vs only (S/P no longer go through shared memory). sstr = hd+8
// padding; Ks/Vs are FA_TKV rows (the r46 launcher's 3*FA_TQ bug is gone).
size_t smem = ((size_t)FA_TQ + 2 * FA_TKV) * (hd + 8) * 2;

3.3 Pitfalls

  1. The masked non-symbolized constant. The 32 KB O write-back stage is independent of FA_TKV but was masked by the staging formula's headroom (34.8 > 32). The correct procedure for a symbolized change is to first grep every smem allocation/reuse point for "the largest requirement" rather than trusting the derived formula alone — this time the launcher came within a hair of under-allocating and the tail block overran.
  2. "Parity green" ≠ "safe to land". The parity fixtures were all green (max err 1e-4), the suite 166/0/3 all green, and only greedy-32 blew the fork open in the token stream. Tolerance gates and byte gates test two different fault classes — the former tests "are the numbers right," the latter tests "is the path exactly on the baseline's rails."
  3. The occupancy numbers could not be re-verified on the spot. ncu/nsys could not run in that window, and 3 blocks/SM was derived from r48's device report. The derivation is self-consistent and the wall-clock conclusion is unambiguous, but the record honestly labels the evidence grade — a derived value never poses as a measured one.

4. Verification

  • parity suite (numeric tolerance ≤1e-4, multiple shapes): proves the math is correct — guards against "the change broke the numbers."
  • suite 166/0/3 (166 passed / 0 failed / 3 ignored): the regression surface — guards against "touching FA broke something else."
  • greedy-32 byte identity (-n 32 --greedy --seed 42, 2K prompt; whole-stream diff of dump /tmp/g32_r50_base.txt vs /tmp/g32_r50_new.txt): guards against "ULP regrouping quietly changing the sampling path" — this doc's protagonist gate, and the only one that caught the problem.
  • interleaved A/B measurement (same window, same binary): the 3-pair series base 2834.1 / 2832.9 / 2805.7 (med 2832.9) against the new med, −0.5%; a further 5-pair series at −0.01% — guards against machine drift being read as a gain.
  • (occupancy: derived from r48's ncu device report, see §2 — the evidence grade is noted in the record.)

5. Results

  • Correctness: parity green; suite 166/0/3; greedy-32 NOT byte-identical — the fork point is one argmax knife-edge token, after which the generated text diverges from "…95% DRAM-bound." to "…95% smem-bound." and each remains coherent. This is benign ULP-level float-regrouping noise (the numbers themselves are right), but against the strict byte gate it is a fail.
  • Wall clock (7B @3325 tok, same-window interleaved A/B medians): 3-pair series base med 2832.9 → −0.5%; 5-pair series −0.01%. Both rounds sit on the negative side of the ±2% noise band — the occupancy gain (2→3 blocks/SM) never cashed into anything visible on the wall.
  • Veto mechanism: after r48 FA is only ~4.7% of the whole wall, and its top stall (global K/V loads) is already hidden by 2 blocks/SM; halving FA_TKV doubles the per-tile fixed costs (3×__syncthreads, cp.async commit/wait, the softmax shfl chains, the O-fragment rescale), exactly cancelling the theoretical gain. Three rounds of same-direction evidence (r23 + r46 + r50): tile shrinking / occupancy on FA is a dead lever.
  • Under what future conditions a retry is worthwhile: only when FA becomes the wall's critical path again AND the campaign accepts downgrading FA-class changes from "greedy byte-identical" to a tolerance-grade gate package (hard argmax gate + penalty-free greedy + control groups — this package was later formalized in the decode campaign's D3a). Under the strict byte-gate policy, no FA_TKV/FA_TQ change can satisfy that gate — this is not an implementation defect but an inherent property of online softmax's regrouping math (r57's FA_TQ experiment hit the same wall again).

6. Lessons

  1. Occupancy is not a universal lever: first ask "is the top stall still un-hidden," then look at resident block count — when latency is already hidden, adding blocks only adds contention, and the per-tile-billed fixed costs double exactly as tiles shrink.
  2. Changing tile size = changing float accumulation order: any change that regroups f32 reductions/rescales under new boundaries produces ULP drift, and the greedy-32 byte gate will catch the one knife-edge token parity can never see.
  3. Symbolized geometry is an investment in "one-dimensional experiments": r48 wrote every geometric quantity as a function of FA_TKV, which is the only reason r50 could run a complete experiment with one define — and it incidentally exposed the single non-symbolized constant (the 32 KB O write-back stage).
  4. Archive dead levers too: three same-direction negatives (r23+r46+r50) mean the "FA tile-size" direction never needs to be tried again; a negative result's value is closing the whole question, not just this one diff.

← 52-r49-a-quantize-shared-dedup · Index · 54-r51-producer-fused-a-quantize →

54 · r51 — producer-fused A-quantize mode 1 (LANDED)

Result: the MMQ A-quantize prepass is folded into its producer kernels — rms_norm_quant_f32_t / swiglu_quant_f32_t write the pad40_t transposed quantized plane in the same pass that produces the f32 output (byte-identical to the standalone prepass). prepass launches 110 → 28 (only wo remains), prepass time 83.0 → 10.1 ms; the fused swiglu runs 4.09 ms per call vs 5.46 ms for the original pair (−25%); whole prefill 2803.4 → 2856.4 (+1.89%), vs-llama 1.18× → 1.16×. Commit: cf1ed4b (code, +421 lines) + bf0c986 (record). Date: 2026-09-06.

1. Background — where things stood

With r48 (FAP2 register-resident softmax, +5.6%) and r49 (A-quantize prepass shared-A dedup, +2.32%) landed, the baseline stood at 2797.5 tok/s, vs-llama 1.18×. r47's converged-regime wall decomposition had split the remaining wall into three pieces: q6_K GEMM (15.8% at the time, already driven to 1.13× and near convergence by r38–r41), FA (r48 had cut it to ~4.7%), and a hidden tax nobody had looked at squarely — the quantize prepass.

r49's dedup had already cut the prepass from 193 launches to 110 (118.4 → 83.9 ms), but it only removed redundancy — quantizing once when the same A is shared by the q/k/v GEMMs. Each of the remaining 110 launches was still a standalone kernel that "reads the f32 activations once, writes the transposed q8 plane once," occupying ~7.4% of the converged wall. And every one of those 110 reads data its producer just finished writing: in the qwen2 prefill graph, every rms_norm / swiglu output exclusively feeds the next GEMM group's A input — input-norm → q/k/v, ffn-norm → gate/up, swiglu → down, output-norm → lm_head. The data is freshly written out of L2, then read back by another kernel.

This is the other half of the layout-transformation locality lesson r34 (the quantize-transpose prepass, +9.72%) already taught: back then r34 moved the layout transform out of the GEMM kernels into a standalone prepass, winning on "transform once, let the GEMMs reuse it"; but the price was that the transformed data had to be consumed again. r49 deleted the repeated transforms; r51 deletes the distance between the transform and the production — why not have the producer write the plane while it's at it?

The appeal of this direction is its gain structure: it touches no GEMM, no FA, no hot loop's instruction stream — it just sews two back-to-back kernels into one. Mathematically it is pure reordering — quantize is a pure function of the f32 values, the f32 output is bit-unchanged, the plane is bit-unchanged. All the risk concentrates in the scheduling window: the plane must be ready before the consumer arrives, must be keyed correctly, and its tail fill must match the standalone prepass exactly.

2. Principle — the GPU mechanism

2.1 What the pad40_t plane is and how the GEMM eats it

The A side of MMQ (int8 tensor-core GEMM) does not consume row-major q8_0 but the transposed + swizzled layout r34 established: the token dimension is grouped in 64s (MMQ_A_BLK = 64, aligned with MMQ's NBI) and padded ("pad40": each token's q8_0 block is 40 B, 64-token blocks aligned). The standalone prepass quantize_q8_0_pad40_t handles one (token, chunk) per thread — a chunk is 32 consecutive elements — and emits two planes:

  • yqs [ntb][nchunk][2048]: the swizzled qs bytes for each (64-token block, chunk). The write offset (((t4*2 + (u>>2)) ^ xswz) << 4) + (u&3)*4 (t4 = r&3, xswz = (r>>2)&7) is the XOR swizzle finalized in r27 — when the BT kernel reads A in warp slices, this arrangement guarantees the same warp's 16 B blocks land in different bank groups, conflict-free;
  • ysda [ntb][nchunk][256]: the packed d|ssum scale — the f16 d (= amax/127) and the int16 ssum (the integer sum of the quantized values, used for the min-term correction) packed into 4 B, laid out per r31's q-major region-split.

The BT kernel's A staging is therefore pure bulk copy (cp.async after r45): the qa8/sda planes are moved into smem directly by (block, k-tile), and the A side does zero index arithmetic in the hot loop. That is the plane's whole reason to exist — quantize, transpose, and swizzle all leave the hot loop.

2.2 The byte-count account

7B @3325 tok. rms-class (d=3584): nchunk = 3584/32 = 112, ntb = ⌈3325/64⌉ = 52 → plane = 52×112×2048 = 11.9 MB qs + 1.49 MB sda; f32 activations 3325×3584×4 = 47.7 MB. swiglu-class (nf=18944): nchunk = 592 → plane = 63.1 MB qs + 7.9 MB sda; f32 output 3325×18944×4 = 251.9 MB.

Producerproduction body (DRAM-grade)what the standalone prepass pays extraafter fusing
rms d=3584read x 47.7 + write y 47.7 MBanother 47.7 MB read (L2 likely cold by then) + 13.4 MB plane writethe re-read hits L1/L2, near-free on the DRAM side
swiglu nf=18944read gate+up 503.7 + write dst 251.9 MBanother 251.9 MB read + 71 MB plane writesame

The key mechanism: the fused kernel's phase 2 re-reads the rows this block just wrote — swiglu is one block per token row, rms is one block per 8 rows (one warp per row) — so those bytes are 100% still in L1/L2 (a swiglu dst row is 19.5 KB; one block's working set is far smaller than L1/L2 capacity). The standalone prepass instead arrives hundreds of microseconds later, after kernels like the GEMMs and FA have flushed L2, and most likely falls to DRAM. What is saved is not the byte count, it is the bytes' temperature — r34's lesson taken to its end: consume the data where it is hottest.

2.3 Why it is bit-identical

The fusion changes no floating-point expression: phase 1's rms / silu·mul elementwise formulas are exactly those of the standalone kernels, and the reductions' lane mappings and order are identical → the f32 output is bit-identical; phase 2's quantize body is a verbatim copy of quantize_q8_0_pad40_t's code (the same amax grouping — serial amax over each 32-element chunk, the same rintf/clamp, the same swizzled writes) → the plane is bit-identical. Quantization is a pure function independent per (token, chunk), so identical input bits + identical code necessarily produce identical output bits — this demotes "did the fusion change the numbers" from a verification problem to a code-review problem. (Contrast r50's lesson: there the grouping of f32 reductions changed, which necessarily drifts ULPs; here no grouping changed.)

2.4 Where the gain concentrates, and where it doesn't

The phase 2 the fusion adds to each producer is pure added work (L1/L2 reads + plane writes), so the gain = the deleted standalone prepass − the addition. swiglu's production body is large (503.7 MB read + 251.9 MB write) and the standalone down-prepass would have re-read another 251.9 MB → fused 4.09 ms vs 5.46 ms for the pair (−25%). rms's production body is small and itself bandwidth-bound (read x + write y ≈ 2×47.7 MB); at d=3584 it is 0.77 ms vs 0.77 ms — a wash. This unevenness is itself data: the gain exists exactly where the saved thing is a DRAM re-read.

3. Implementation

3.1 Design choices (why this shape and not another)

Why rms's phase 2 re-reads y instead of staying register-resident. rms's phase 1 is one warp per row with lanes consuming d/4 float4s strided; phase 2's quantize body needs one (token, chunk) thread to take an amax over 32 consecutive elements — the two lane→element mappings don't line up. Register residency would require handing values across threads (smem relay + regrouping the amax), i.e. rewriting the quantize body: losing the "verbatim copy" structural guarantee of bit-identity, in exchange for saving one L1-hit re-read. The design choice is to keep the verbatim copy: one __syncthreads(), then phase 2 re-reads y from L1/L2.

Why one __syncthreads is enough. Phase 2 reads only the rows this block wrote: rms's block writes 8 rows and phase 2 quantizes the same 8; swiglu's block writes 1 row and quantizes the same 1. The plane addresses are partitioned by (token-block, chunk) with zero overlap between blocks — intra-block synchronization suffices; there is no cross-block dependency and no grid-level sync needed.

Why swiglu rejects the register-resident variant (by design). If swiglu organized phase 1 by the quantize body's thread mapping (thread-per-chunk), each thread would write 32 strided f32s of dst — non-coalesced f32 stores = 8× sector amplification, catastrophic on a ~254 MB f32 output stream. So the shape is inverted: phase 1 uses block-per-token-row coalesced float4 silu·mul (same elementwise formula as swiglu_f32), and after __syncthreads() phase 2's threads re-read this block's row (just written, hot in L1/L2) through the verbatim quantize body. The record also left a hook for the next step: a smem-staged dst or register-resident quantize in phase 2 "could squeeze out another ~0.5–1 ms per launch" — that is r52's mode 2.

Why the grid covers the 64-padded token count. Every byte of the transposed plane must be written: the GEMM reads all ntb×64 rows (out-of-range token rows never write C back, but if the plane's tail rows are not zero-filled, scratch reuse would carry the previous graph's stale bytes into amax/ssum — the result would still be blocked by C's write-back guard, but the plane would no longer be deterministic). The standalone prepass's grid naturally covers the padded total; the fused kernel must replicate that: rms's grid = ntb × (64/8) blocks of 256 threads, and phase 2 runs the quantize body on padded rows too (t < n false → emit zeros); swiglu's grid = ntb*64 blocks, with padded rows running only the zero fill.

Why it hangs off r49's MmqCache instead of a new channel. r49's cache key is (f32 output device pointer, nt, id), and the consumer prefill_mmq already consults it. The fused kernel keys the plane under the same key (key = f32 output pointer), so the consumer changes by zero lines. The scratch sizes are exactly those of mmq_quantize_transposed, guaranteeing get_or_grow returns the same pointer at hit validation — otherwise a grown buffer would make a "hit" read a misaligned plane.

The gate design: MINFER_MMQ_A_FUSE=1 AND the full MMQ gate set (RAW/RAW_NB/A_TRANSPOSE/Q6K_NB) + rows ≥ 16 + dim % 256 == 0. The first four prevent "writing a fused plane for a GEMM that would re-quantize natively anyway" (correct but pure waste); rows ≥ 16 keeps decode (nt==1), the capture window, FusedFFN, and short prefills on the unfused pair; dim % 256 == 0 satisfies the transposed GEMM's nchunk % 8 requirement. OOM returns Err and the caller falls back to the unfused pair.

3.2 Key code

The fused rms kernel (current tree src/cuda_kernels.cu lines 840–880; phase 1 shares rms_norm_f32's lane mapping, phase 2 is a verbatim quantize body):

__global__ void rms_norm_quant_f32_t(
    const float* __restrict__ x, const float* __restrict__ w,
    float* __restrict__ y,
    uint8_t* __restrict__ yqs,   // [ntb][nchunk][2048] swizzled qs plane
    uint8_t* __restrict__ ysda,  // [ntb][nchunk][256] packed d|ssum
    int d, float eps, int n, int nchunk, int ntb
) {
    // Phase 1: rms_norm — one warp per row, lane mapping and accumulation
    // order identical to rms_norm_f32 (bit-identical output).
    const int row = blockIdx.x * RMSQ_RPB + (threadIdx.x >> 5);
    ...
        ss = warp_reduce_sum(ss);
        float scale = rsqrtf(ss / (float)d + eps);
        ...
            y4[i].x = xv.x * scale * wv.x;      // ← f32 output: bit-identical to rms_norm_f32
            ...
    __syncthreads();
    // Phase 2: quantize this block's rows into the pad40_t plane — the
    // quantize_q8_0_pad40_t body verbatim (one thread per (token, chunk),
    // strided over the block's 8 rows). The grid covers the 64-padded token
    // count, so the padded tail rows are zero-filled exactly like the
    // standalone prepass (deterministic plane regardless of scratch reuse).

Phase 2's quantize body (current tree lines 887–916) — verbatim-identical to the standalone prepass, only the data source switched from x to the just-written y:

        float dsc = 0.0f; int ssum = 0; uint32_t packed[8];
        ...
        if (t < n) {
            const float* src = y + (size_t)t * d + b * 32;   // ← the row just written, hot in L1/L2
            float4 sv[8];
            #pragma unroll
            for (int v = 0; v < 8; v++)
                sv[v] = *reinterpret_cast<const float4*>(src + 4 * v);
            float am = 0.0f;
            ... amax grouping ...
            dsc = am / 127.0f;
            ...
                int q = int(rintf(e[j] * di));               // ← verbatim-identical to the standalone body
                q = max(-128, min(127, q));
                ...
        }                                                     // t >= n → emit zeros (tail fill)

The fused swiglu kernel (current tree lines 943–968) — block-per-row coalesced phase 1 + re-read phase 2:

__global__ void swiglu_quant_f32_t(
    const float* __restrict__ gate, const float* __restrict__ up,
    float* __restrict__ dst,
    uint8_t* __restrict__ yqs, uint8_t* __restrict__ ysda,
    int dim, int nt, int nchunk, int ntb
) {
    const int t = blockIdx.x;  // one token row per block (grid = ntb*64)
    if (t < nt) {
        ...
        for (int i = threadIdx.x; i < n4; i += blockDim.x) {   // coalesced float4
            float4 gv = g4[i]; float4 uv = u4[i]; float4 ov;
            ov.x = (gv.x / (1.0f + expf(-gv.x))) * uv.x;       // same formula as swiglu_f32
            ...
            o4[i] = ov;
        }
    }
    __syncthreads();
    for (int b = threadIdx.x; b < nchunk; b += blockDim.x) {   // phase 2 re-reads this row

The launcher's grid geometry (current tree lines 3740–3758) — where the padded coverage is settled:

void launch_rms_norm_quant_f32_t(...) {
    int grid = ntb * (MMQ_A_BLK / RMSQ_RPB);        // covers all padded rows
    rms_norm_quant_f32_t<<<grid, RMSQ_RPB * WARP, 0, stream>>>(...);
}
void launch_swiglu_quant_f32_t(...) {
    swiglu_quant_f32_t<<<ntb * MMQ_A_BLK, 256, 0, stream>>>(...);   // one block per token
}

The host side (current tree src/cuda.rs lines 2978–3011) — same scratch, same cache key, prefill_mmq untouched:

#![allow(unused)]
fn main() {
pub fn rms_norm_quant(&self, x, w, y, d, n, eps) -> Result<(), String> {
    let nchunk = d / 32;
    let ntb = n.div_ceil(64);
    let qa8 = Self::get_or_grow(&self.buf_qa8_t, ntb * nchunk * 2048);  // same as the standalone prepass
    let sda = Self::get_or_grow(&self.buf_sda_t, ntb * nchunk * 256);
    if qa8.is_null() || sda.is_null() {
        return Err("cuda: rms_norm_quant plane OOM".to_string());       // OOM → fall back
    }
    unsafe {
        launch_rms_norm_quant_f32_t(x as *const f32, w as *const f32,
                                    y as *mut f32, qa8 as *mut u8,
                                    sda as *mut u8, d as i32, eps,
                                    n as i32, nchunk as i32, ntb as i32, stream);
    }
    // register the plane into the r49 MmqCache, key = f32 output pointer;
    // dead_write=false (mode 1 did write the f32)
    self.record_mmq_cache_transposed(y as usize, n, d, qa8 as usize, sda as usize, false);
    Ok(())
}
}

The consumer side's hit validation (current tree src/cuda.rs lines 2784–2796) — "the registered plane really got eaten" closes the loop through this half; the physical pointers must match:

#![allow(unused)]
fn main() {
let key = (x as usize, nt as usize, id as usize);
let mut cache = self.mmq_cache.lock().unwrap();
if cache.active && cache.key == key && cache.transposed {
    // get_or_grow may have reallocated on a larger miss: validate the
    // physical pointers so a grown buffer is never reused stale.
    let qa8 = Self::get_or_grow(&self.buf_qa8_t, need_qa8) as usize;
    let sda = Self::get_or_grow(&self.buf_sda_t, need_sda) as usize;
    if qa8 == cache.qa8 && sda == cache.sda {
        return (qa8, sda);          // ← hit: the standalone quantize launch is skipped
    }
}
}

The "before" for contrast: the standalone prepass path that got fused away (current tree src/cuda.rs lines 2813–2832) — in mode 0/fallback the consumer still goes through here, launching a standalone quantize_q8_0_pad40_t and filling the cache itself; r51's fusion only makes the "miss launch" step stop happening after particular producers:

#![allow(unused)]
fn main() {
// (before) r49 cache miss → standalone prepass launch + record
let qa8 = Self::get_or_grow(&self.buf_qa8_t, need_qa8);
let sda = Self::get_or_grow(&self.buf_sda_t, need_sda);
launch_quantize_q8_0_pad40_t(x, qa8 as *mut u8, sda as *mut u8,
                             id, nt, nchunk, ntb, stream);
cache.active = true;
cache.key = key;
cache.transposed = true;
cache.dead_write = false;       // mode 1: the f32 src was written, safe to re-quantize
cache.qa8 = qa8 as usize;
cache.sda = sda as usize;
}

The caller-side gate (current tree src/graph/cuda_backend.rs lines 536–538, SwiGLU arm; the RmsNorm arm is isomorphic):

#![allow(unused)]
fn main() {
let dim = node.out_shape[0];
let rows = if dim > 0 { n / dim } else { 0 };
if dim > 0 && n == dim * rows && rows >= 16 && dim % 256 == 0 {
    match self.state.mmq_a_fuse_mode() { 1 => { /* swiglu_quant, fall back to the unfused pair on failure */ } ... }
}

3.3 Pitfalls

  1. The temptation and price of register residency. The "cleverest" shape (phase 1 keeps the silu·mul results in registers and quantizes them directly, saving the re-read) was rejected at design time: the thread-per-chunk mapping makes the f32 dst stores non-coalesced (8× sector amplification, unacceptable on a ~254 MB stream), and it breaks the verbatim-copy bit-identity guarantee. Lesson: pick a fusion shape by "which side's store pattern is coalesced," not by "which one saves a read."
  2. Tail fill is part of correctness, not an optimization. If the padded rows don't get zeros, the plane is non-deterministic under scratch reuse; although C's write-back guard blocks the wrong values, the "plane bit-identical to the standalone prepass" property breaks and the verification distorts with it. The fused kernels replicate the standalone behavior exactly with grid coverage of the padded total + the t < n guard — the verifier specifically proves this with poisoned buffers (see §4).
  3. The uneven gain is mechanism, not defect. fused rms at d=3584 is a 0.77 vs 0.77 wash — rms's production body is small and itself bandwidth-bound, so the fusion only trades in one L1/L2 re-read. Expecting "everywhere +25%" would misjudge the landing as a failure; the real headliner is swiglu→down.
  4. wo was deliberately left out of v1. The producer of the attention output (wo's A input) is the FA kernel — its output layout is not isomorphic to token rows, so v1 doesn't touch it. The remaining 28 launches correspond exactly to 1 wo per layer, and that arithmetic is itself the proof that "fusion coverage is complete."

4. Verification

  • standalone verifier (nvcc-compiled alone, valid_r51.cu): fused kernels vs (standalone producer + standalone prepass) byte-equal on 11 shapes — including nt=3354/129/64/63/1, d=896 nt=300, d=640 nt=100, swiglu nf=18944, covering both the 7B and 14B geometries. Division of labor among the shape classes: nt=1/63/64/129 specifically attack the 64-token block boundaries (single block, one-short-of-full, exactly full, spanning+tail), nt=3354 is the real whole-prefill shape, and d=896/640 test nchunk divisibility at non-primary sizes. The plane buffers are pre-poisoned (filled with garbage before running) — any unwritten plane byte (especially the padded tail) surfaces immediately. Guards against: the fusion changing bits, tail fill leaks.
  • parity ×3 (tolerance 1e-4, end-to-end logits): guards against end-to-end math breakage.
  • greedy-32 byte identity: under mode 1 both the f32 outputs and the planes should be bit-identical to the old path — accumulation order untouched, so the gate should be green, and it is (contrast r50's lesson: that gate catches "regrouping"; there is none here).
  • suite 166/0/3: the regression surface.
  • A/B interleaved measurement (same window, same machine, distributions fully separated): guards against machine drift being read as a gain.

5. Results

  • prepass: launches 110 → 28 (only wo remains — its producer is the FA kernel, deliberately untouched in v1); time 83.0 → 10.1 ms.

  • kernel level (whole-prefill totals): fused rms 54 × 0.77 ms (vs 0.77 wash), fused swiglu 27 × 4.09 ms (vs 5.46 ms for the original swiglu+quantize pair, −25%); net kernel time −31 ms.

  • wall clock (7B @3325 tok): 2803.4 → 2856.4 (+1.89%), A/B distributions fully separated; vs-llama 1.18× → 1.16×. Gain composition: the swiglu→down fusion (−1.37 ms per launch) plus launch-count-halving-scale scheduling savings.

  • the wall after landing (the signpost for the next round):

    Residualtimeshare of wall
    q6_K GEMM197.8 ms18%
    fused swiglu110.5 ms10%
    FA52.9 ms4.8%

    fused swiglu is promoted to a visible residual, and its f32 write + L1/L2 re-read (~251 MB per layer) is exactly r52's mode-2 target.

  • suite 166/0/3; greedy byte-identical; parity ×3.

6. Lessons

  1. Layout-transformation locality taken to its end: r34 moved the transform out of the GEMM and won once; r51 moves the transform back into the producer and wins again — two faces of the same coin: consume the data at its hottest moment and place.
  2. A verbatim copy is a fusion's bit-identity proof: making phase 2 a verbatim copy of the fused-away kernel's code turns "did the numbers change" from a verification problem into a review problem; the verifier's job narrows to coverage (poisoned buffers + tail shapes).
  3. Fusion gains concentrate where the production body is large: reconcile launch by launch (swiglu −25%, rms a wash) and "where mode 2 is worth doing" becomes a reading, not a guess.
  4. AND the gate with the full set: writing a fused plane for a GEMM that would re-quantize natively is correct waste — the fusion gate must be jointly true with the consumer path's gate set.

← 53-r50-fa-tkv-16 · Index · 55-r52-skip-write-mode2 →

55 · r52 — skip-write mode 2: skipping the f32 intermediate write-out (LANDED)

Result: MINFER_MMQ_A_FUSE=2 — the fused producers become register-resident, computing the pad40_t plane and never writing the f32 output at all (the *_nw kernels). fused rms 41.5 → 22.4 ms, fused swiglu 110.5 → 64.1 ms (−42%), fused producers total 151.9 → 86.5 ms, prefill window −6.1%; whole prefill 2855.7 → 3011.3 (+5.45%), vs-llama 1.16× → 1.09×. Before landing, the greedy gate caught an OOB transcription error that parity could not see. Commit: 910d967 (code, +534/−55) + fb659f7 (record). Date: 2026-09-06.

1. Background — where things stood

r51 folded quantization into the producers, leaving the prepass at 28 launches and 10.1 ms, whole prefill +1.89%. But r51's post-landing wall decomposition lit up the next target brightly: fused swiglu at 110.5 ms, 10% of the wall — it reads gate and up at 251.9 MB each (7B's ffn intermediate dim is 18944), writes the f32 dst of 251.9 MB, and then phase 2 re-reads those 251.9 MB from L1/L2 to quantize. r51's record left an explicit hook: a register-resident quantize in phase 2 "could squeeze out another ~0.5–1 ms per launch."

One step further is the logical endgame: if quantization happens in registers, does anyone still read the f32 output? In the prefill graph, the only consumers of the rms/swiglu outputs are the immediately following MatMul groups, and after r51 those consumers read the plane (via r49's MmqCache) — the f32 buffer itself is no longer touched by anyone. The f32 write + L1/L2 re-read in mode 1 is a complete dead code path: written out, read back, and then — quantized once — never touched again. Mode 2 deletes that path wholesale: the producer computes the f32 values directly from its inputs (expressions verbatim-unchanged), completes quantization in registers, and writes only the plane.

The correctness conditions here are an order of magnitude harsher than mode 1's: "nobody reads the f32 output" is not a kernel property, it is a graph-topology property. Any missed reader (debug dump, trace, the legacy GEMM path, a non-adjacent consumer) would read a buffer that was never written — garbage. So the bulk of r52's work is not in the kernels (the kernels are actually simpler) but in the window safety proof and the loud-failure mechanism for violations.

2. Principle — the GPU mechanism

2.1 The traffic account that gets saved

Per-layer swiglu (7B, nf=18944, 3325 tokens):

PathDRAM/L1L2 traffic
mode 1 (r51)read gate 251.9 + read up 251.9 + write dst 251.9 + phase 2 re-read 251.9 (hot in L1/L2) + write plane ~71 MB
mode 2 (r52)read gate 251.9 + read up 251.9 + write plane ~71 MB

That saves ~503.8 MB per layer of f32 write+read round trip (28 layers ≈ 14.1 GB/prefill of SM-level traffic; the re-read half mostly never hit DRAM anyway, but L1/L2 bandwidth and the store/load instructions are real money). rms is analogous: mode 1 pays a 47.7 y-write + 47.7 re-read, mode 2 is left with only the 47.7 x-read + plane write. Measured: swiglu 110.5 → 64.1 ms (−42%), rms 41.5 → 22.4 ms.

Note that what is saved is traffic, not memory: the f32 output buffer is still allocated by the graph allocator (it is still the MmqCache key and still the resolution target of the node's output buffer) — it is just never written.

2.2 Why the register-resident lane remapping is byte-safe

In mode 2 the quantization is no longer one thread serially walking 32 values; instead 8 lanes each compute 4 values, and an __shfl_xor tree regroups the amax and ssum. Regrouping usually means float reordering — r50 just taught us that is the source of ULPs — but here the two quantities being reduced happen to both be exact:

  • amax = max(|v0|…|v3|) then the shfl tree: fmaxf is exact under any associative/commutative order (no NaN in this value domain);
  • ssum = Σq (int) then the shfl tree: integer addition is exact.

And the f32 values feeding the quantization themselves: v = xv * scale * wv (rms) and silu(g)*u (swiglu) are verbatim elementwise copies, and scale's reduction lane mapping is the same as mode 1's → identical bits. The element-level rintf/clamp are unchanged. So the plane is byte-identical to mode 1 / the standalone prepass — r50's lesson applied in reverse: regrouping only bites on inexact operations (f32 sum chains); max and int-sum can be reordered freely. This is also the theoretical basis for greedy-32 being byte-identical across A_FUSE=1/2.

2.3 The skip-write window conditions (why not writing is dared)

The f32 output buffer becomes memory promised unread. The promise is justified node by node, each condition paired with a concrete "what happens otherwise":

#Window conditionGuarantee mechanismIf violated
1the only consumers of the rms/swiglu outputs are the immediately following consecutive plain MatMul groupsgraph build order = execution order (builder topology); MmqCache's consecutive-window rule clears the cache at any non-MatMul nodea later node reads the f32 → reads an unwritten buffer
2the residual adds the pre-norm buffer (x), not yqwen2 topology: attn_out + x_in, ffn_out + x_mid — the add's inputs are the norm's input sideif the deleted y were referenced by a residual → garbage residual
3the graph-tail G3 segment runs only n_out=1 rowsthat segment falls outside the fusion gate (rows ≥ 16)—
4FusedQkv/FusedFFN are decode-onlygated on nt==1; neither op appears in prefill graphs—
5RmsNorm/SwiGLU are not in-place opsthey do not appear in §5's aliasing rule, so the output buffer is not any other node's input"skip-write" under aliasing would starve other readers

Runtime degradation conditions stack on top of the static proof: any debug/trace reader present (MINFER_GRAPH_DUMP layer-0 node dump / MINFER_DUMP_DIR / MINFER_TRACE / viz live capture) or the legacy prefill GEMM enabled (its kernels read the f32 A directly under MINFER_NO_PREFILL_GEMM=1) → mode 2 degrades to mode 1 (still taking r51's gains, just not skipping the write).

3. Implementation

3.1 Design choices (why this shape and not another)

A three-tier degradation ladder: mode 2 → (OOM / window readers / degradation conditions) → mode 1 → (OOM) → the unfused pair. Every tier is correct, and every tier either writes or doesn't write the f32; the caller falls back tier by tier. This determines that mode 2's failure mode is always "a performance loss," never "a wrong result."

A loud backstop, not a quiet assumption. The window proof is an argument, and arguments go stale — the moment the graph structure changes (say, some op starts reading the rms output), the proof silently expires. So r52 added a dead_write flag to MmqCache: mode 2 records its plane with dead_write=true; any path that would re-quantize that buffer (the transposed or native prepass) refuses to execute when it hits a dead_write entry, returning (0,0)/0, and the caller errors out the whole A path — a loud error replaces reading garbage.

The key is unchanged. The plane is still keyed on the (unwritten) f32 output pointer, exactly as in r51 — the consumer side prefill_mmq changes by zero lines, and the mode-1/mode-2 difference is sealed inside mmq_a_fuse_mode()'s return value and the two new kernels.

The swiglu-nw mapping choice is covered in §3.3 pitfall (b): coalesced rounds won over naive thread-per-chunk — this one was measured, not designed.

3.2 Key code

The mode-2 rms kernel (current tree src/cuda_kernels.cu lines 1037–1131). Phase 1 shares mode 1's lane mapping and reduction order (scale bit-identical), with no y store; phase 2 has each warp sweep its own row, with 8-lane-group exchanges of amax/ssum; padded tail rows take a dedicated zero-fill arm:

__global__ void rms_norm_quant_nw_f32_t(
    const float* __restrict__ x, const float* __restrict__ w,
    uint8_t* __restrict__ yqs, uint8_t* __restrict__ ysda,
    int d, float eps, int n, int nchunk, int ntb
) {
    ...
    // Phase 1: identical to rms_norm_quant_f32_t — same lane mapping and
    // accumulation order over x, so `scale` is bit-identical. No y store.
    ...
        ss = warp_reduce_sum(ss);
        scale = rsqrtf(ss / (float)d + eps);
    ...
    if (row >= n) {
        // Padded-tail row: zero-fill this row's plane slots exactly like the
        // standalone prepass (deterministic plane regardless of scratch
        // reuse). grid = ntb*(64/RPB) covers every padded row.
        for (int b = lane; b < nchunk; b += WARP) { ... = 0; ... }
        return;
    }
    // Phase 2: quantize THIS warp's row (warp-uniform row => the shfl_xor
    // reductions below never see divergence). Chunk c = 4k + lane/8 covers
    // float4s 32k+lane; d % 256 == 0 makes d4 % 32 == 0 (exact loop).
    for (int k = 0; k < d4 / 32; k++) {
        float4 xv = x4[k * 32 + lane];
        float4 wv = w4[k * 32 + lane];
        // y expression verbatim from rms_norm_quant_f32_t phase 1 (bit-identical
        // f32 values — the mode-1 kernel quantizes these after a memory
        // round-trip, which is exact for f32).
        float v0 = xv.x * scale * wv.x;  ...
        float am = fmaxf(fmaxf(fabsf(v0), fabsf(v1)), fmaxf(fabsf(v2), fabsf(v3)));
        // 8-lane group reduce (lanes [g8*8, g8*8+8) own one 32-float chunk).
        am = fmaxf(am, __shfl_xor_sync(0xffffffffu, am, 4));   // max: exact, safe to reorder
        am = fmaxf(am, __shfl_xor_sync(0xffffffffu, am, 2));
        am = fmaxf(am, __shfl_xor_sync(0xffffffffu, am, 1));
        ...
        int ssum = (q0 + q1) + (q2 + q3);
        ssum += __shfl_xor_sync(0xffffffffu, ssum, 4);          // int-sum: exact
        ...
        const int b = k * 4 + (lane >> 3);   // chunk index (the pre-landing OOB was on this line)

The mode-2 swiglu kernel (current tree lines 1144–1210) — the coalesced-rounds scheme: lane l reads gate/up's float4 f = round stride + threadIdx.x (one 512 B access per warp, i.e. r51's phase-1 pattern), silu·mul completes in registers, and one 8-lane group is one chunk; dim % 256 == 0 guarantees every 8-lane group maps onto whole chunks:

    // Mode-2 swiglu: no dst write, single phase. Per token row, the block sweeps
    // the row's float4s in COALESCED rounds (lane l loads float4 rd*256+l of
    // gate/up — 512 B per warp access, the r51 phase-1 pattern), computes
    // silu*mul in registers, and quantizes via the rms-nw lane mapping ...
    // (A naive per-thread chunk mapping — one thread quantizing a whole 128-B
    // chunk — makes the gate/up loads lane-strided and measured SLOWER than
    // the mode-1 write+re-read.)
    for (int f = threadIdx.x; f < ...; f += blockDim.x) {
        const bool active = t < nt && f < n4;
        float v0 = 0.0f, ...;
        if (active) {
            float4 gv = g4[f]; float4 uv = u4[f];
            v0 = (gv.x / (1.0f + expf(-gv.x))) * uv.x;   // silu*mul verbatim
            ...
        }
        ... 8-lane shfl amax/ssum, same mapping as rms-nw ...
        if (f < n4) {
            // t >= nt rows land here with all-zero v/am/ssum — exactly the
            // standalone prepass's deterministic padded-tail zero-fill.
            ... write qs/swizzled + write d|ssum when u==0 ...

Host side: the mode-2 wrapper records with dead_write = true (current tree src/cuda.rs lines 3094/3129):

#![allow(unused)]
fn main() {
// the f32 output y serves only as the MmqCache key (the consumer matmul's
// A pointer); the buffer is promised dead
self.record_mmq_cache_transposed(y as usize, n, d, qa8 as usize, sda as usize, true);
}

The backstop: a re-quantize path that hits a dead_write entry refuses (current tree src/cuda.rs lines 2797–2811; the native path at lines 2855–2865 is isomorphic):

#![allow(unused)]
fn main() {
// r52: a mode-2 (skip-write) fused producer left the f32 src UNWRITTEN
// and promised its consumers would hit this cache. Reaching the launch
// path with a dead-write entry for the SAME buffer means the window
// guarantee broke ... re-quantizing would read the dead buffer's garbage.
// Refuse: the callers treat (0, 0) as a failed path and prefill_mmq
// errors out loudly instead of silently producing wrong results.
if cache.active && cache.dead_write && cache.key.0 == x as usize {
    eprintln!("minfer/cuda: MMQ A-quantize refused: mode-2 dead-write A ...");
    return (0, 0);
}
}

The scheduling end of the degradation ladder (current tree src/graph/cuda_backend.rs lines 539–566):

#![allow(unused)]
fn main() {
match self.state.mmq_a_fuse_mode() {
    2 => {
        if self.state.swiglu_quant_nw(...).is_ok() { return Ok(()); }  // mode 2
        if self.state.swiglu_quant(...).is_ok() { return Ok(()); }     // → mode 1
    }                                                                   // → unfused pair
    1 => { /* swiglu_quant, fall back to the unfused pair on failure */ }
}

Mode 2's enabling conditions (current tree src/cuda.rs mmq_a_fuse_mode; the reader-detection branch — the function in the current tree also carries r60's default-on modification; at r52 the semantics were that an explicit MINFER_MMQ_A_FUSE=2 enables it):

#![allow(unused)]
fn main() {
2 if mode2_possible
    && std::env::var_os("MINFER_GRAPH_DUMP").is_none()
    && std::env::var_os("MINFER_DUMP_DIR").is_none()
    && !crate::trace::enabled()
    && !crate::live::enabled() =>
{
    2
}
// Mode 2 requested but a window-reader/fallback condition is
// active: keep the r51 fused semantics (write the f32 output).
2 => 1,
}

3.3 Pitfalls

  1. One transcription error, two wrong faces. Before landing, rms-nw wrote the chunk index as k*8 + lane/8 — the correct form is k*4 + (lane>>3): each round k covers 32 float4s = 4 chunks (8 lanes/chunk × 4 float4s/lane), lane>>3 ∈ {0,1,2,3}, hence b = k*4 + lane>>3. Written as k*8 + lane/8, adjacent k rounds' b values overlap and the last round runs off the plane → IllegalAddress; and CUDA's async error only surfaces at the next API call, where it masqueraded as a string of fake "OOM"s. What caught it was greedy-32 (fork → investigate immediately), with parity all green — the parity fixtures never run the rms-nw path at all. Two lessons: when transcribing an index, first count how many chunks each round covers; and error masquerading lies — the first response to an error is to check the most recent kernel write boundary, not to trust the error's name.
  2. Register-resident ≠ faster; it depends on the store pattern. The first swiglu-nw used the naive thread-per-chunk mapping (one thread quantizes a whole 128 B chunk): it saved the re-read, but the gate/up loads became lane-strided — the warp access shattered into non-coalesced segments. Measured, the fused swiglu went 110.5 → 114.0 ms (slower than the re-read), and whole prefill only +1.39%. Only after switching to coalesced rounds did it reach 64.1 ms. r51's reason for rejecting register residency (non-coalesced f32 store = 8× sector amplification) re-established itself in a different spot under the nw shape: warp-level 512 B coalesced access is the floor — whoever breaks it dies.
  3. When the verifier breaks in an environment, swap the gate honestly. r51's poisoned-plane standalone verifier (valid_r51.cu) SIGBUSes in this environment and cannot run. r52's byte-identity is instead covered by greedy-32 being byte-identical across A_FUSE=1/2 (the two modes must produce exactly the same generation stream — this gate is sensitive to any kernel-level bit difference). The record states the downgrade explicitly, leaving no illusion of "verified."
  4. A promise needs a penalty clause. The window proof covers every reader that exists today, but proofs don't stop the future; the dead_write refusal mechanism turns "the proof expired" from silently reading garbage into an error at startup — the maintainability of skip-write-class optimizations is entirely staked on this layer.

4. Verification

  • greedy-32 byte-identical across A_FUSE=1/2: this doc's workhorse gate — both a proxy proof of kernel-level bit identity (§3.3 pitfall 3) and the catcher of the OOB transcription error. Guards against: the nw mapping changing bits, OOB, window violations.
  • parity ×3: end-to-end numeric correctness — but it has zero coverage of rms-nw (the fixtures never go through that path), honestly noted in the record; this doc is the best object lesson in "which gate sees what."
  • suite 166/0/3: the regression surface.
  • A/B interleaved measurement (same window, same binary): +5.45% with distributions fully separated.
  • window safety proof (the 5 conditions of §2.3, checked node by node): an argument, not a test — its staleness is backstopped by the dead_write mechanism.

5. Results

  • Kernel level (whole-prefill totals): fused rms 41.5 → 22.4 ms; fused swiglu 110.5 → 64.1 ms (−42%); fused producers total 151.9 → 86.5 ms; prefill window −6.1%.

  • Wall clock (7B @3325 tok): 2855.7 → 3011.3 (+5.45%), distributions fully separated; vs-llama 1.16× → 1.09×.

  • Correctness: parity ×3; greedy-32 byte-identical under both A_FUSE=1 and =2; suite 166/0/3.

  • The mode ladder (the quick reference for all subsequent A-FUSE behavior):

    MINFER_MMQ_A_FUSEbehaviorf32 write?
    2 (window-safe)_nw kernels: register-resident quantizeno (dead_write=true)
    2 (readers/degradation conditions present)r51's _t kernelsyes
    1r51's _t kernelsyes
    0/otherunfused pair (producer + standalone quantize)yes
  • The signpost for the next round: with the fused producers slimmed down, the wall's headliner returns to the q6_K GEMM — r45 (cp.async A staging) and r44 (W_exp), two mechanisms each measured "wall-neutral" alone, were waiting to be retried as a package, which is r53's basket thesis.

6. Lessons

  1. Skip-write safety = the completeness of the consumer enumeration: prove the window node by node, degrade automatically when readers are present, and backstop proof expiry with a loud refusal mechanism — all three, or none.
  2. The greedy byte-identity gate sees corruption parity cannot: the fixtures' path coverage defines parity's blind spots; the bit-identity gate is the only end-to-end gate sensitive to "any kernel-level bit difference."
  3. Async CUDA errors change faces: IllegalAddress can masquerade as OOM — the first response to an error should be checking the most recent kernel write boundary, not trusting the error's name.
  4. Coalesced access beats register cleverness: any remapping that costs a warp's 512 B coalesced shape should be costed at worst-case sector amplification before a line is written — r51's design veto and r52's measured retreat are the same physics.

← 54-r51-producer-fused-a-quantize · Index · 56-r53-q6k-wexp-cpasync-bundle →

56 · r53 — q6_K bundle: W_exp pre-expansion plane + cp.async B staging (LANDED)

Result: r44 (pre-expanded dense W_exp — deletes WORK) and r45 (cp.async staging — deletes WAIT), two mechanisms each wall-neutral on their own, are bundled into the q6_K BT kernel: B staging becomes a pure cp.async bulk copy out of the W_exp plane. The ffn_down kernel goes 16.06 → 12.76 ms (−20.5%) (≈ r44's −10.9% + r45's −10.2% stacking near-additively), attn_v −15.9%; whole prefill 3024.7 → 3176.9 (+5.03%) = 1.05× vs-llama — the q6_K line closes here. The price is +1.52 GB of device memory. The doc also lends its name to a rule: a fallback-correct optimization nearly landed silently as "zero fast-path launches." Commit: 83fee77 (code, +265/−20) + 4907d9f (record). Date: 2026-09-06.

1. Background — where things stood

r44 and r45 are two special negative results on the q6_K line: both parity green, both kernel-faster, both wall-neutral — and both reverted.

  • r44 (the dense-index W_exp fix): the pre-expanded-B idea first failed parity in r43 (the dense plane was indexed by the padded raw-W stride); r44 found that one line's root cause, parity went green and the kernel −10.9% — but the wall was −0.42%. The mechanism reading at the time: deleting the recomb ALU merely changed the latency's shape, and the synchronous LDG→STS memory latency stood exposed at the top of the staging phase as before.
  • r45 (cp.async A-side staging): handing the A staging's latency to the async units, kernel −10.2% and long_scoreboard −18% — but the wall −0.34%. The mechanism reading: after r41 the q6_K GEMM was no longer the wall's headliner, and a kernel that isn't the wall can never reach the wall no matter how fast it gets.

Each fell below the +1.5% bar, so under the single-shot rule both should revert. But they share one property that makes "revert" different from "rejection": they act on the same traffic path (staging) with non-overlapping mechanisms — r44 deletes WORK (the 16-shift/or/subtract-per-element recombination arithmetic), r45 deletes WAIT (memory-latency exposure). Alone, each one's gain was absorbed by the remaining cost dominated by the other: delete only WORK and WAIT takes over; delete only WAIT (and on the A side at that) and the kernel isn't the wall. This class of "overlapping in traffic, orthogonal in mechanism" combination is the basket the whole campaign had been collecting — r18's EB/SB pre-expansion machinery, the W_exp plane fixed up by r43/r44, r45's pipeline scheme — all inventory on the shelf.

r51/r52 moved the fulcrum: fused producers slimmed from 151.9 ms to 86.5 ms, swiglu fell from 10% of the wall to ~6.4%, and the wall's headliner turned back to the q6_K GEMM. The wall came back around, so the inventory shipped — the basket thesis's first full redemption: put r44's mechanism on the B side (r45 had only done the A side back then), then use r45's group-count pipeline + visibility barrier to hide the copy latency, so the same staging code sheds both WORK and WAIT.

2. Principle — the GPU mechanism

2.1 q6_K layout background (what this doc needs to stand alone)

block_q6_K is ql[128] + qh[64] + sc[16] + d[2] = 16 16-element sub-blocks (not 8×32) — one 32-k chunk spans two sub-blocks with different scales, hence the kernel's KSPLIT=2: two mma.m16n8k16 passes, each with its own dsc, with the plain rescale sum += da·dsc (no dmin term, verified against the CPU reference in r38). The B side's "centered int8" expansion: each (ql nibble, qh 2-bit field) pair combines into an 8-bit value and subtracts 32, giving a centered int8 in −32..31 — the r41 kernel did this in the hot loop every tile (uint4-widened ql/qh reads + recombination ALU).

2.2 The three costs of r41's B staging and what the bundle covers

Form① ql/qh reads (bytes)② recomb ALU③ staging latency
r41 (EXP=false, kept in the current tree)uint4-widenedin the hot loopsynchronous, exposed at the top of every tile
r53 (EXP=true)none (reads the dense plane)none (computed once at registration)cp.async, hidden inside compute

r41's uint4 widening had already cut ① by ~16×; r53 kills ①② with W_exp and ③ with cp.async. ffn_down −20.5% ≈ r44's −10.9% (deleting ②) + r45's −10.2% (deleting ③), near-additive — the two act on the same section of the same kernel and don't crowd each other out. Single-shot, each bought only half the kernel improvement, and r45 landed in the window when the wall had turned away — wall-neutrality is a joint verdict of "timing + magnitude," not a mechanism's death sentence. A −20.5% kernel improvement × q6_K's ~18–20% share of the wall ≈ +4–5% wall clock, clearing the bar exactly.

2.3 The W_exp plane and the memory account

At registration expand_q6k_dense expands the padded q6_K into dense centered-int8: id bytes per row (1 B/element), super-block stride 256; the scales (dsc) are still read separately by the kernel per the r38 scheme. The q6_K tensor inventory of 7B Q4_K_M:

Tensor classcountper-tensor od×idsubtotal
attn_v141.8 MB25.2 MB
ffn_down1467.9 MB950.6 MB
output.weight1544.6 MB544.6 MB
total1,521,237,632 B ≈ 1.52 GB

(The task brief's "+15 MB" estimate took one ffn_down's increment as the whole cost; this order-of-magnitude error incidentally explains where the q6_K launch count comes from — the number of W_exp planes = the number of q6_K GEMMs ≈ 27.) That is the bundle's purchase price: 1.52 GB for +5.03%.

2.4 How the latency is hidden: r45's pipeline scheme ported to the B side

r39's double buffering provided the structure for "stage the next tile while computing this tile"; r45 added the async-side discipline: all staging goes through cp.async and commits once per tile, and the main loop uses wait_group 1 (waiting only for the previous group to land, letting this group's copies stay in flight) + one __syncthreads (cross-thread visibility). r53 puts the B copies into the same commit group, with A/B sharing one wait discipline — the staging latency leaves the critical path entirely.

Explicit-PTX cp.async. The copies use gemm_cp16 (inline cp.async.cg.shared.global [dst], [src], 16, sz), where the src-size qualifier does the out-of-range zero fill: with full=false, sz=0 writes that 16 B chunk as all zeros — rows beyond od (tiles past the weight's rows) are deterministically zero. With EXP=false the template branch disappears at compile time and the SASS is byte-identical to r41's — the fallback is not a runtime if, it is a different compiled artifact.

3. Implementation

3.1 Design choices (why this shape and not another)

Template <KDR, bool EXP> instead of a runtime branch. The fast path and the r41 fallback share one kernel source, and if (EXP) resolves at compile time: EXP=false's SASS must be exactly r41's (the fallback's byte identity is guaranteed by construction, not retroactively by tests); EXP=true's branch costs zero. This also reduces "map miss → fall back" to a single pointer lookup — no second codebase to maintain.

The dense index respects r44's root cause. W_exp + j*id + sb*256 + cbase*32 + cc*16 — the row stride is id (dense), not the padded 224; r43's parity FAIL came precisely from using the padded stride there. The gate's id % 256 == 0 simultaneously guarantees the NB kernel's launch geometry and the dense index's 16 B alignment (cp.async 16 B chunks don't cross misaligned boundaries).

The map key = the padded weight's device pointer (the one prefill_mmq holds), value = the W_exp plane pointer. Plane names carry a geometry encoding ({name}__exp{od}x{id}), so a same-name different-shape re-registration can never silently reuse the old plane. Registration failure (alloc/upload) → the map stays empty → the kernel takes the EXP=false fallback + a loud eprintln once per process.

3.2 Key code

The host expander (current tree src/cuda.rs lines 1845–1872; the two output elements share the low/high nibbles of one (ql,qh) byte pair):

#![allow(unused)]
fn main() {
pub fn expand_q6k_dense(padded: &[u8], od: usize, id: usize) -> Vec<u8> {
    const Q6KB: usize = 210;      // raw super-block: ql[128]+qh[64]+sc[16]+d[2]
    const Q6KPB: usize = 224;     // padded stride
    let nbe = id / 256;
    let row_len = nbe * Q6KPB;
    let mut out = vec![0u8; od * id];          // dense: row stride = id
    for j in 0..od {
        let prow = &padded[j * row_len..(j + 1) * row_len];
        let orow = &mut out[j * id..(j + 1) * id];
        for sb in 0..nbe {
            let blk = &prow[sb * Q6KPB..sb * Q6KPB + Q6KB];
            let (ql, qh) = blk.split_at(128);
            let obase = &mut orow[sb * 256..sb * 256 + 256];
            for it in 0..2usize {               // 256 elements = 2 × 128
                for r in 0..64usize {
                    let qlb = ql[it * 64 + r];
                    let qhb = qh[it * 32 + (r & 31)];
                    let s0 = (r >> 5) * 2;      // the 2-bit field's phase
                    let e0 = it * 128 + r;
                    obase[e0] = ((qlb & 0xF) | (((qhb >> s0) & 3) << 4)).wrapping_sub(32);
                    obase[e0 + 64] =
                        (((qlb >> 4) & 0xF) | (((qhb >> (s0 + 4)) & 3) << 4)).wrapping_sub(32);
                }
            }
        }
    }
    out
}
}

Registration and the map (r53's original form, verified identical to the current tree; src/cuda.rs lines 1614–1644):

#![allow(unused)]
fn main() {
pub fn register_weight_q6k_exp(&self, name: &str, padded: &[u8], od: usize, id: usize) {
    // geometry-encoded sibling name: a same-name different-shape
    // re-registration can never collide with (and silently reuse) a stale
    // plane of the same byte size but a different od/id layout.
    let exp_name = format!("{name}__exp{od}x{id}");
    let exp = Self::expand_q6k_dense(padded, od, id);
    self.register_weight(&exp_name, &exp);
    // the MAP is keyed by the PADDED weight's device pointer (what
    // prefill_mmq holds); the value is the W_exp plane's pointer
    if let Some(wp) = self.get_weight_ptr(name) {
        if let Some(ep) = self.get_weight_ptr(&exp_name) {
            if !wp.is_null() && !ep.is_null() {
                self.q6k_exp.lock().unwrap().insert(wp as usize, CudaPtr(ep));
                return;
            }
        }
    }
    ... // once per process: falls back to the r41 in-kernel expand
}
}

The kernel-side EXP branch (current tree src/cuda_kernels.cu lines 6764–6789) — staging is a pure copy, and the dense index respects r44's root cause to the line:

if (EXP) {
    /* r53 bundle: the ql+qh recomb + -32 centering ran ONCE at
     * registration (expand_q6k_dense -> dense centered-int8 plane
     * W_exp: od x id, row stride = id, super-block stride = 256),
     * so the staging is a pure cp.async bulk copy (explicit PTX)
     * from W_exp — no recomb ALU, no register round-trip, no
     * ql/qh reads ... Rows beyond od zero-fill via
     * the cp.async src-size qualifier (gemm_cp16 full=0). */
    const int nc = (KDR * 32) / 16;   /* 16B chunks per row */
    const int ncopy = MMQ_NBJ * nc;
    for (int g = threadIdx.x; g < ncopy; g += blockDim.x) {
        const int jj = g / nc, cc = g % nc;
        const int j = j0 + jj;
        const bool full = (j < od) && (sb < nsb);
        const uint8_t* src = W_exp + (size_t)j * id
            + (size_t)sb * 256
            + (size_t)(cbase * 32 + cc * 16);        // ← r44's root-cause line: stride id
        gemm_cp16((__half*)(void*)(qbexpb + (size_t)jj * (KDR * 32) + cc * 16),
                  (const __half*)(const void*)src, full);
    }

The copy primitives themselves (current tree lines 4483–4491) — three __forceinline__ wrappers are the whole PTX surface:

__device__ __forceinline__ void gemm_cp16(__half* smem_dst, const __half* gsrc, bool full) {
    unsigned d = (unsigned)__cvta_generic_to_shared(smem_dst);
    int sz = full ? 16 : 0; // src-size 0 => zero-fill the 16B chunk
    asm volatile("cp.async.cg.shared.global [%0], [%1], 16, %2;\n" ::"r"(d),
                 "l"(gsrc), "r"(sz));
}
__device__ __forceinline__ void gemm_cp_commit() { asm volatile("cp.async.commit_group;\n"); }
__device__ __forceinline__ void gemm_cp_wait1() { asm volatile("cp.async.wait_group 1;\n"); }
__device__ __forceinline__ void gemm_cp_wait0() { asm volatile("cp.async.wait_group 0;\n"); }

r45's wait discipline (current tree lines 6905–6924) — the B copies join the same commit group, and wait_group 1 lets this group's copies fly while waiting only for the previous one:

    RAW_STAGE_Q6K_BT(0, 0);
    __syncthreads();

    int buf = 0;
    for (int kt = 0; kt < nktile; ++kt, buf ^= 1) {
        // Overlap kt+1's global->smem staging with kt's compute: stage into the
        // OTHER buffer (buf^1) while reading buffer buf (the mmq_nt<7,2> pipeline).
        if (kt + 1 < nktile) {
            RAW_STAGE_Q6K_BT(kt + 1, buf ^ 1);
            // r53 (EXP) / r56: two groups are pending (kt's and kt+1's); wait
            // until only kt+1's remains — group(kt), the cp.async copies ...
            // has landed, while buf^1's copies stay in flight under kt's
            // compute (in-order group completion).
            gemm_cp_wait1();
        } else {
            gemm_cp_wait0();  // last tile: drain every outstanding group
        }
        __syncthreads();  // r53/r56: cross-thread visibility of the kt
                          // buffer's async copies before compute

(For contrast: the EXP=false fallback's elementwise expansion expand_q6_elem still sits at current tree lines 6687–6698 — ((ql[ql_idx] >> ql_shift) & 0xF) | (((qh[qh_idx] >> qh_shift) & 0x03) << 4) then minus 32 — exactly the arithmetic W_exp computes once at registration.)

The byte-exactness test (current tree src/graph/cuda_backend.rs lines 3994–4062, cuda_q6k_exp_dense_byte_exact) — a three-way reconciliation of independent scalar mirror + host expansion + device read-back:

#![allow(unused)]
fn main() {
// independent scalar mirror, straight from the device formula
let mut want = vec![0u8; od * id];
for j in 0..od { for sb in 0..nbe { ... for e in 0..256usize {
    ...
    want[j * id + sb * 256 + e] = (v as i32 - 32) as u8;
} } }
// host-side production expander vs the mirror
let host = crate::cuda::CudaState::expand_q6k_dense(&padded, od, id);
let hmis = host.iter().zip(want.iter()).filter(|(a, b)| a != b).count();
assert_eq!(hmis, 0, "expand_q6k_dense vs mirror ({od}x{id})");
// device upload path: build + read back + compare
state.register_weight_q6k_padded(&name, &raw, od, id);
state.register_weight_q6k_exp(&name, &padded, od, id);
...
state.copy_from_device_pinned(p, &mut got);
let dmis = got.iter().zip(want.iter()).filter(|(a, b)| a != b).count();
assert_eq!(dmis, 0, "device W_exp vs mirror ({od}x{id})");
}

3.3 Pitfalls

  1. The liveness near-miss — this doc's namesake incident. The first integrated build: parity green, greedy green, suite green — but flip on MINFER_MMQ_RAW_NB_DEBUG and the fast-path launch count was zero. The root cause was one line: the map insert keyed on the exp buffer's own pointer, while the kernel looks up by the padded weight pointer → always miss → always the EXP=false fallback. The fallback is the byte-identical r41 path, so no correctness gate could see it; the only symptom was "+5.03% didn't show up." The fix = key on the padded pointer, and the debug label immediately counted 27 W_exp-cp.async launches.
  2. The SASS trap of explicit PTX. Once cp.async is written as inline PTX, a wrong constraint/type can make ptxas skip emitting LDGSTS and silently degrade. The SASS check confirmed LDGSTS ×12 — the evidence chain must reach the SASS level; a correct PTX source does not imply a correct SASS.
  3. An order-of-magnitude memory estimate error. The task brief estimated "+15 MB"; measured 1.52 GB — one ffn_down's increment (exp 67.9 MB − padded ~59 MB) was taken as the whole-model cost. Memory-for-speed decisions must be summed over the tensor inventory, not extrapolated from a single instance.
  4. The register budget re-check. The new branch lands in a __launch_bounds__(256, 3) kernel: <2,true> compiled to 80 regs / 24 B stack — the 3-block budget held, and r53 did not buy its gain with r40's occupancy.

4. Verification

  • W_exp byte-exactness cargo test (cuda_q6k_exp_dense_byte_exact, shapes (64,256)/(40,512)/(24,768) covering multiple super-blocks and multiple od): independent scalar mirror (transcribed straight from the device formula) vs the host expander, then vs the device-plane read-back (pinned copy), 0 mismatches. Guards against: a recurrence of r43/44's stride-mismatch class — this plane's only historical failure mode.
  • SASS check: LDGSTS ×12. Guards against: explicit-PTX cp.async degradation.
  • liveness label (MINFER_MMQ_RAW_NB_DEBUG): per-launch-path counts (the current tree's label is r54's refined three-state version: exp=off = deliberately off, fallback! = a plane was expected but missed — "reverted" and "off" must not look alike). Guards against: a fast path that silently never engages — the exclusive gate for this doc's near-miss.
  • parity ×3: end-to-end numerics. greedy byte identity (a 178-character stream): the modes and the fallback share the accumulation order, so the gate should be green — and it is.
  • A/B interleaved measurement: 3024.7 → 3176.9, distributions separated (base max < new min).
  • suite 167/0/3: +1 new gate test (the W_exp byte-exactness test joined the regular suite).

5. Results

  • Kernel level: ffn_down 16.06 → 12.76 ms (−20.5%) — r44's −10.9% and r45's −10.2% stacking near-additively; attn_v −15.9%.
  • Wall clock (7B @3325 tok): 3024.7 → 3176.9 (+5.03%); vs-llama 1.09× → 1.05× — the q6_K line closes here (at r37 it was still the worst kernel at 6.38×/GMAC; after r38–r41's +2.9/+13.3/+13.0/+30.7%, this doc completes the last stretch).
  • The price: +1.52 GB of device memory (W_exp = the sum of od×id over 14 attn_v + 14 ffn_down + output.weight); <2,true> at 80 regs / 24 B stack, 3-block occupancy preserved.
  • Correctness: W_exp byte-exact (0 mismatches); parity ×3; greedy identical; suite 167/0/3.
  • Left for the next round: the 1.52 GB purchase price needs an exit ticket (r54's MINFER_MMQ_Q6K_EXP=0 — measured −5.04% to buy back 1.52 GB, with the launch-path label distinguishing "deliberately off" from "accidental fallback"); the A side's dsc consumption and A staging wait become the new headliner (r56's W_dsc f32 plane + A cp.async, where r45's mechanism finally lands on the A side).

6. Lessons

  1. The basket theorem: optimizations that overlap in traffic and are orthogonal in mechanism stack near-additively — mechanisms that are wall-neutral alone are inventory, not scrap; stockpile them and ship the bundle when the wall turns back.
  2. A fallback-correct optimization must have a "the fast path is actually alive" check: parity and greedy are fully blind to "a fast path that silently never engages" — label the launch paths and count fast-path launches; that is this class's exclusive gate.
  3. Do the full account before trading memory for speed: a plane's cost = the sum of od×id over all target tensors; single-instance extrapolation is off by two orders of magnitude.
  4. Template the two forms: the fast path and the fallback compile from one source (EXP=false SASS = r41), byte identity is guaranteed by construction, and the fallback's cost reduces to one lookup.

← 55-r52-skip-write-mode2 · Index · 57-r54-q6k-exp-optout →

57 · r54 — MINFER_MMQ_Q6K_EXP: an exit valve for the 1.52 GB W_exp plane (LANDED)

Result: the default path is byte-unchanged (3181.0 tok/s, same noise band as r53's landed 3176.9); MINFER_MMQ_Q6K_EXP=0 trades −5.04% of whole-prefill speed (3020.7 tok/s) for 1.52 GB of device memory back (measured per PID: 7636 → 6182 MiB, Δ1454 MiB ≈ the plane's census value of 1,521,237,632 B). Both modes are parity/greedy all-green, and the liveness census of the 27 q6_K launches holds 27×/0 in both directions. Commit: 3252e96 (code) + b860b7e (record). Date: 2026-09-06.

1. Background — where things stood

r53 had just closed out the q6_K line: bundling r44's "pre-expanded dense B plane W_exp" (removing the recombination work) and r45's cp.async staging (removing the wait) into the NB-BT q6_K kernel took whole-prefill 3024.7 → 3176.9 tok/s (+5.03%), with the vs-llama multiple tightening from 1.09× to 1.05×. The price was written on the same line: +1.52 GB of device memory.

What is this 1.52 GB? W_exp is a "dense centered-int8 plane expanded per padded q6_K tensor at od × id bytes" — the q6_K ql/qh nibble stream expanded once at registration into a dense 1-byte-per-element int8, so the kernel's B staging degenerates into a pure cp.async copy. On 7B q4_k_m the tensors that go through q6_K are 14 attn_v + 14 ffn_down + output.weight, totaling 1,521,237,632 B ≈ 1.52 GB (r53's record also incidentally corrected the task's "+15 MB" underestimate — that estimate had taken one ffn_down's increment as the whole price). GB10's 127 GB unified memory doesn't feel it, but minfer is an engine shipped to users: on an 8–16 GB card, carrying an extra 1.5 GB for a default-on +5% turns "can it run at all" into "am I willing to buy another card."

And the shape at the time made this tax impossible to decline: the plane registration hung under the MINFER_MMQ_Q6K_NB gate, and that gate also controlled the kernel itself. To turn the plane off you had to turn the kernel off too, back to the f16 path — a retreat counted in tens of percent, not "5% less." r53's failure semantics covered only one case: on allocation failure the map stays empty, the kernel falls back, and it eprintlns loudly. A user "wanting a smaller memory footprint" was not among them — the allocation would succeed, and the memory would be eaten.

So r54's task was to fit this memory-for-speed switch with an exit valve. It sounds like a one-line env check; the real constraints were three:

  1. No second numeric path may be introduced. With the plane opted out, the kernel must still be there, and its output must be bit-identical to the default path. Fortunately r53 had already templated the kernel as <KDR, EXP>: the EXP=false instantiation compiles the cp.async B branch out entirely, keeping the r41-era "in-kernel uint4 ql+qh expansion" — a path that was itself a parity-green, in-service implementation, compiled into the binary since r53. The opt-out merely means not building the plane and letting the dispatch map-miss fall into this existing instantiation. Both instantiations live in the same cubin; the env only changes whether the host side registers the plane.
  2. The plane gate must be decoupled from the kernel gate. MINFER_MMQ_Q6K_NB controls the NB-BT kernel itself; a user may want that kernel (its in-kernel expansion is not slow either) but not the plane. So EXP is an independent switch, ANDed with NB — the user picks any of four combinations: kernel+plane (default), kernel without plane (saving 1.52 GB), all-f16, and the meaningless "kernel off, plane on" (the plane then isn't built either — the registration gate incidentally guarantees no reader-less plane is left behind).
  3. A deliberate fallback must be distinguishable from an accidental one. r53's namesake lesson: parity/greedy cannot see a "fast path that silently never runs" — on r53's landing day, a wrong map key sent all 27 launches down the slow path while parity and greedy stayed all green; the liveness label was what counted it out. r54 widens the B-path label from two states to three: a user-initiated off (exp=off) and a registration-bug map-miss (fallback!) must look different in the log.

Once this step was done, the default path had not moved by a single bit — its whole value is that it gave r56 (the W_dsc plane) and r60 (the promotion flipping defaults on) a reusable "default-on + 0 opt-out" template. r60's promotion definition copies this shape verbatim: all six MMQ gates become "default-on + 0 opt-out."

Two execution-level notes: first, this is a user-approved independent switch (the record's own words: "Independent switch (user-approved)") — it is not on any bundled thesis's chain; it is a pure product decision. Second, it was deliberately done before r55's convergence verdict: making "the plane is optional" an accomplished fact first is what makes the convergence statement's 1.05× a 1.05× with an escape hatch.

2. Principle — the GPU mechanism

2.1 What the 1.52 GB buys while on

The NB-BT q6_K kernel must stage the B (weights) into shared memory in KDR=2 batches for every output tile. The two paths differ entirely in the staging phase:

  • EXP=true (r53 default): B staging is a pure cp.async 16 B-chunk copy from the W_exp plane. Dense index W_exp + j*id + sb*256 + cbase*32 + cc*16 (r44's stride root-cause fix), 16 B alignment guaranteed by the id % 256 == 0 gate; rows beyond od are zero-filled by gemm_cp16's src-size 0. No ql+qh recombination ALU, no register round trip, no per-nibble ql/qh reads; the copy latency goes to the async units, hidden under the previous tile's compute by the group wait. r44 quantified "removing the work" at −10.9% of kernel time and r45 quantified "removing the wait" at −10.2% — r53 proved the two are near-additive (ffn_down kernel −20.5%: 16.06 → 12.76 ms; attn_v −15.9%).
  • EXP=false (the r41 path): the staging phase itself reads the raw ql/qh nibble stream and recombines the centered int8 on the mma loop's critical path with shift/mask ops (the per-16-element-group version of v = ((ql[..]>>qsh)&0xF) | (((qh[..]>>qh_shift)&3)<<4) - 32). This path has been in service since r41 — r41's uint4-ization merged 16 per-byte LDGs into one 16 B read, cutting the L1TEX scoreboard exposure (85.5% → 33.6%) — and r54 wrote zero new lines for it.

So the performance price of turning the plane off = giving back exactly what r44+r45+r53 bought. The whole-prefill arithmetic: 3020.7 / 3181.0 = 0.9496, i.e. −5.04% — almost symmetric with r53's +5.03%. The symmetry is no coincidence: r53's gain was precisely "the kernel stops doing these two things"; r54 puts them back in the kernel and the gain flows away down the same path.

InstantiationB staging formoriginrole in r54
<KDR, EXP=true>pure cp.async copy from the W_exp planer53default path (3181.0 tok/s)
<KDR, EXP=false>in-kernel uint4 ql+qh expansionr41the EXP=0 path + all map-miss fallbacks (3020.7 tok/s)

Both instantiations are in the same cubin (template compilation artifacts); the env changes only the host-side "register the plane or not."

2.2 Why the byte counts reconcile

W_exp per tensor = od × id bytes (1 B centered int8 per element). The census value 1,521,237,632 B = 1450.9 MiB; the measured Δ is 1454 MiB, a gap within page granularity and allocator overhead — "≈ the census" holds. Conversely this also verifies that with EXP=0 all of the plane memory comes back: the registration early-return happens before any sibling allocation, so device memory sits at its pre-r53 level rather than "built but unused."

2.3 Why byte-identity is a structural guarantee, not a verification result

The q6_K expansion is a pure integer transform: ql's 4-bit nibble and qh's 2-bit high bits compose a 6-bit value, minus 32 to center — no rounding, no floating point, no ordering issues. Both EXP instantiations feed the same mma element-identical values; the only difference is whether the expansion happens on the host (at registration, expand_q6k_dense) or in-kernel (r41's uint4 branch). So as long as the dispatch really lands in <KDR, EXP=false>, bit-identity is constructed, not something a tolerance gate has to rescue. r54's verification budget therefore went almost entirely into "the dispatch really lands there" (the liveness census); parity is a routine confirmation.

2.4 The liveness label mechanism

Under MINFER_MMQ_RAW_NB_DEBUG=1, every q6_K launch eprintlns one line: "the path currently taken." Its necessity comes from r53's field test: parity/greedy cannot distinguish "the fast path is fast" from "the fast path never ran but the slow path isn't slow enough to give itself away." The label bins every launch into a countable bucket, and a census of 27×/0 (27 fast, 0 accidental) is the evidence that "the mechanism is actually running."

2.5 New concepts, first appearance

  • W_exp plane: a registration-time pre-expanded dense int8 copy of the q6_K weights, od × id B per tensor, introduced in r53, opt-out-able via MINFER_MMQ_Q6K_EXP=0 from r54 on.
  • in-kernel expand (the r41 path): the pre-existing implementation where, absent a plane, the kernel reads ql/qh as uint4 in the staging phase and recombines them with shifts.
  • liveness label: the path-annotating eprintln under MINFER_MMQ_RAW_NB_DEBUG=1; used to count fast-path launches and guard against silent fallbacks.
  • per-PID memory measurement: on GB10, nvidia-smi's aggregate usage field reads [N/A]; per-process usage requires --query-compute-apps.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Default-on, not default-off: the campaign's default discipline is "default = the verified fastest path" — performance responsibility sits on the default, and the tradeoff is handed to an explicit opt-out. The reverse (default-off, opt-in for speed) would hide 5% behind an environment variable nobody reads. r60's promotion "default-on + 0 opt-out" inherits exactly this ordering.
  • Default-on, "0" to opt out: the gate is written std::env::var("MINFER_MMQ_Q6K_EXP").as_deref() != Ok("0"). Rust's env::var returns Err for an unset variable, and after as_deref() the miss is still not Ok("0"), so the inequality holds — unset is equivalent to "1", i.e. r53's behavior. This "default is on" phrasing deliberately avoids boilerplate like unwrap_or("1"), makes an explicit "0" the only opt-out channel, and later became r60's promotion-standard pattern.
  • Early-return at the registration point, not a dispatch-time fork: EXP=0 makes the registration function return before the sibling is even built. "Build the plane but don't use it at dispatch" would cost the full 1.52 GB and leave the switch nominal; "check the env at dispatch and pick the instantiation" would leak the env from registration semantics into execution semantics, one extra lookup per launch. The early return means the env is read exactly once, at load.
  • AND semantics: EXP is only meaningful while the NB kernel is alive, so plane registration hangs under mmq_gate_on("MINFER_MMQ_Q6K_NB") && EXP != "0" && id % 256 == 0; Q6K_NB=0 turns off kernel and plane together, so "the plane exists but its consumer is gone" orphan memory cannot occur.
  • Failure semantics inherited from r53: when plane alloc/upload fails, the map stays empty, the kernel falls back to in-kernel expand, and one loud eprintln fires — r54 doesn't touch that path, it only makes sure the path now has a name (see the labels in 3.2).

3.2 Key code

The registration gate (current tree src/cuda.rs, the EXP check introduced in r54; the od % 2 == 0 dsc lines are a r56 addition, see doc 59):

#![allow(unused)]
fn main() {
// r54: MINFER_MMQ_Q6K_EXP decouples the plane from the kernel gate —
// unset/"1" keeps the r53 default (build it), explicit "0" skips the
// build entirely (registration early-returns; device memory stays at
// the pre-r53 level) so dispatch map-misses into the EXP=false r41
// in-kernel expand. ANDed with Q6K_NB: EXP only matters when the NB
// kernel is live.
if Self::mmq_gate_on("MINFER_MMQ_Q6K_NB")
    && std::env::var("MINFER_MMQ_Q6K_EXP").as_deref() != Ok("0")
    && id % 256 == 0
{
    self.register_weight_q6k_exp(name, &padded, od, id);
    if od % 2 == 0 {                       // r56: the W_dsc plane rides the same gate
        self.register_weight_q6k_dsc(name, &padded, od, id);
    }
}
}

register_weight_q6k_exp's map semantics (same file, current tree) — the key is the padded weight's device pointer (r53's liveness incident was exactly a mistake here), the value is the plane pointer:

#![allow(unused)]
fn main() {
pub fn register_weight_q6k_exp(&self, name: &str, padded: &[u8], od: usize, id: usize) {
    // geometry-encoded sibling name: a same-name different-shape
    // re-registration can never collide with (and silently reuse) a stale
    // plane of the same byte size but a different od/id layout.
    let exp_name = format!("{name}__exp{od}x{id}");
    let exp = Self::expand_q6k_dense(padded, od, id);
    self.register_weight(&exp_name, &exp);
    // the MAP is keyed by the PADDED weight's device pointer (what
    // prefill_mmq holds); the value is the W_exp plane's pointer
    if let Some(wp) = self.get_weight_ptr(name) {
        if let Some(ep) = self.get_weight_ptr(&exp_name) { /* insert(wp, ep) */ }
    }
}
}

The dispatch-side lookup (inside the current tree's prefill_mmq) — with EXP=0 the map has no entry at all, unwrap_or(null) leaves w_exp a null pointer, and the launcher selects the <KDR, EXP=false> instantiation accordingly:

#![allow(unused)]
fn main() {
// r53: pre-expanded B plane lookup by the padded weight's
// device pointer (null on miss -> launcher selects the r41
// in-kernel-expand instantiation).
let w_exp = self
    .q6k_exp
    .lock()
    .unwrap()
    .get(&(wptr as usize))
    .map(|cp| cp.0)
    .unwrap_or(std::ptr::null_mut());
// r56 (Session E item 2b): the precomputed dsc f32-pair plane
// (null on miss -> the r41 scalar dsc path in-kernel).
let w_dsc = self
    .q6k_dsc
    .lock()
    .unwrap()
    .get(&(wptr as usize))
    .map(|cp| cp.0)
    .unwrap_or(std::ptr::null_mut());
}

The kernel's two instantiations (current tree src/cuda_kernels.cu; the W_dsc parameter in the signature is likewise a r56 addition; the EXP=false branch is r41's uint4 expansion preserved verbatim):

template <int KDR, bool EXP>
__global__ void __launch_bounds__(256, 3) mmq_raw_nb_bt_q6k_kernel(
    const uint8_t* __restrict__ W, const uint8_t* __restrict__ W_exp,
    const uint8_t* __restrict__ W_dsc, ...
) {
    ...
    if (EXP) {
        /* r53 bundle: the ql+qh recomb + -32 centering ran ONCE at
         * registration ... so the staging is a pure cp.async bulk copy
         * (explicit PTX) from W_exp — no recomb ALU, no register
         * round-trip, no ql/qh reads ... Dense index:
         * W_exp + j*id + sb*256 + cbase*32 (16B-aligned: id is a
         * multiple of 256 on this path). Rows beyond od zero-fill via
         * the cp.async src-size qualifier (gemm_cp16 full=0). */
        ...
    } else if ((bstride & 15) == 0) {
        /* r41: 16-elem group expand via uint4 ql+qh global loads. ...
         * Element-for-element identical to expand_q6_elem. */
        ...
    }

The B path's three-state label (current tree src/cuda.rs, under MINFER_MMQ_RAW_NB_DEBUG=1):

#![allow(unused)]
fn main() {
// r54: name WHY the r41 in-kernel expand is running —
// "exp=off" is the intentional MINFER_MMQ_Q6K_EXP=0
// switch; "fallback!" means a W_exp build was expected
// (padded weight, EXP gate on) but the map missed
// (alloc/upload failure or a registration bug). Raw
// 210-B weights never get a plane -> kept unqualified.
let b = if !w_exp.is_null() {
    "W_exp-cp.async"
} else if !padded_q6k {
    "in-kernel-expand"
} else if std::env::var("MINFER_MMQ_Q6K_EXP").as_deref() == Ok("0") {
    "in-kernel-expand(exp=off)"
} else {
    "in-kernel-expand(fallback!)"
};
}

The four states' semantics: W_exp-cp.async = the fast path; in-kernel-expand (unqualified) = raw 210 B weights that never had a plane to build — not a fallback; exp=off = the user switched it off deliberately; fallback! = should have been built but wasn't — a bug. The label also reports r39's KDR=2 and A-transpose together, so one census reads the whole path.

3.3 Pitfalls

  • The label must distinguish "never eligible" from "eligible but not built": raw 210 B q6_K weights can never get a plane; if that class were also labeled fallback!, the liveness census would drown in legitimate in-kernel paths and the real bug would be invisible. The three-state design's core is separating exactly these two.
  • GB10's aggregate memory reading is [N/A]: on a unified-memory architecture nvidia-smi's total fields report nothing; only nvidia-smi --query-compute-apps=pid,used_memory --format=csv reads per-process usage — this trap nearly demoted "measured memory" into "citing r53's census value." The measured 7636/6182 MiB sits 3 MiB from the census, confirming that the per-PID reading is the engine's true footprint.
  • Co-tenant suite noise: this window's suite once showed 165/2 (a 46 GB sglang co-tenant was running); both cases re-ran individually with --exact and passed — judged a co-tenant flake, unrelated to this change, and finally recorded as 167/0/3.
  • r53's map-key lesson is an inherited constraint: if the map were keyed on the plane's own pointer, dispatch's lookup by padded pointer would miss forever — r54's registration code and three-state label are both designed around "this class of error must be visible on the spot."

4. Verification

  • parity ×3, run once per mode: guards against math errors; per §2.3 this should be a construction guarantee, and parity here is a routine confirmation.
  • greedy byte-identical, three-way comparison: exp=1 vs exp=0, plus both against r53's recorded greedy stream. Guards against "the path switch changing float accumulation" — theoretically impossible, confirmed in measurement; it also covers the r41 branch's in-service health on the current tree in the exp=0 mode.
  • liveness census 27×/0, both directions: with exp=1, all 27 q6_K launches are W_exp-cp.async with 0 accidental fallbacks; with exp=0, all exp=off with 0 fallback!. Guards against r53's "parity/greedy all green but the fast path never ran" (back then a wrong map key, caught by exactly this discipline).
  • memory measured per PID: 7636 vs 6182 MiB. Guards against "assumed byte counts" — r53's "+15 MB" underestimate is what an assumption costs; the measurement matching the census's 1450.9 MiB confirms the opted-out memory truly came back.
  • suite 167/0/3: guards against cross-feature regressions (including r53's new gate test); the 165/2 co-tenant false positives were excluded per the "isolated --exact re-run" procedure.

5. Results

Modewhole prefillvs r53 defaultdevice memory (per-PID)B-path label
default / =1 (exp=on)3181.0 tok/sunchanged (r53 was 3176.9, same noise band)7636 MiB27× W_exp-cp.async, 0 fallbacks
MINFER_MMQ_Q6K_EXP=03020.7 tok/s−5.04%6182 MiB (−1454 MiB ≈ 1.52 GB)27× exp=off, 0 fallback!

vs-llama (the 3324.4 anchor, campaign accounting) is recorded from r53's 1.05× to 1.04×. Each direction gets what it wants: the default is the verified fastest path; EXP=0 is a legitimate downshift for memory-constrained cards, and the downshifted path is also a parity/greedy-verified in-service implementation, not a second-class citizen.

Put −1454 MiB in a user's terms: it is about 19% of an 8 GB card's memory budget and 9.5% of a 16 GB card's — for the former, that is the magnitude of "can I also load one more LoRA / one more KV copy," not a negligible rounding error. This is also why the exit valve had to be in place before r56/r60 kept adding planes to this chain: every plane added raises the valve's value by a notch.

Two follow-ons that show the switch's shape was right:

  • r56 rides it directly: the W_dsc plane's registration hangs under the same Q6K_EXP != "0" gate, so EXP=0 opts out of all plane memory in one move (another 363 MB). Had r54 not raised this gate first, stripping the plane after r56 would have meant re-verifying two stacked features.
  • r60 canonized the pattern: the promotion's definition is exactly "default-on + 0 opt-out" (the r54 pattern) — all six MMQ gates flipped on it, with MINFER_MMQ=0 kept as the legacy-f16 escape hatch; r54 supplied the first precedent.

6. Lessons

  1. A memory-for-speed knob needs the trio: a byte-identical fallback compiled into the binary, a liveness label that distinguishes "deliberate fallback" from "accidental fallback," and a measured (not derived) memory delta.
  2. Fit the exit valve before the next feature rides the gate: r56's dsc plane and r60's promotion both inherit this gate directly; stripping the plane afterwards means re-verifying two stacked features.
  3. On GB10, verify memory per PID: the aggregate reading is [N/A]; nvidia-smi --query-compute-apps is the instrument that works.
  4. The exit valve's fallback path must be a first-class citizen: EXP=false is not a degraded implementation but the in-service path since r41 preserved verbatim — which lets the verification budget go entirely into "the path really gets taken."

← 56 · Index · 58 →

58 · r55 — swiglu roofline audit + one-shot prefill CUDA-Graph: both closed on the record (CLOSED)

Result: both low-risk levers are closed before any code was written by measurement — fused-swiglu already runs at 242 GB/s ≈ 89% of the 273 GB/s spec roofline, so even a perfect kernel's ceiling gain is only +0.74% (< the +1.5% bar); the reclaimable part of one-shot prefill capture (the recurring gaps) is ≤ 0.1% (the rest is one-time host stalls and capture-ILLEGAL mid-window cudaMallocs). The campaign verdict: CONVERGED (3181 vs 3324.4 = 1.05×). Commit: 83c3c67 (record-only commit — zero tree changes, baseline binary cmp-verified). Date: 2026-09-06.

1. Background — where things stood

r53 landed the q6_K B-side bundle (+5.03%), r54 fitted the exit valve onto the 1.52 GB plane, and whole prefill stood at 3181 tok/s, vs-llama 1.05×. With the q6_K line, the FA line, and the A-quantize line all closed out, Session D — as the "basket final round" — took inventory of the remaining low-risk levers and picked two targets that looked most within reach:

  1. Speeding up fused-swiglu again. r51/r52 had already fused silu+mul into the producer (mode 1 writes the f32 + q8 plane; mode 2 skips even the f32, producing q8 in registers). This kernel is an already-optimized object, not a beginner's job: prepass launches went 193 → 110 (r49's shared-A dedup) → 28 (r51's fusion), prepass time 118.4 → 83.9 → 10.1 ms; mode-2 then took fused swiglu from 110.5 → 64.1 ms (−42%). The proposal argued there was still something to mine — after all, what remains is "pure bandwidth" work, and everyone wants to squeeze it once more.
  2. CUDA-Graph capture for the one-shot prefill. R3-B's 3-run protocol auto-captures repeated same-nt prefills (the server/multi-turn paths benefit); but a CLI one-shot prefill never reaches a 3rd run and has always run bare. nsys shows ~7.84 ms of idle inside the window — "capturing those gaps away" sounds like free money.

The campaign had an iron rule in force at this point: the +1.5% bar. It is not arbitrary: the A/B interleaved measurement window noise is ±2% (co-tenant drift, machine state), and anything below it can neither be proven nor reproduced — every lever landed since r37 sits above it, and every REVERTED "right direction but too small" lever (r44/r45 alone at −0.34%/−0.42%) sits below it. r55 changed the working method: derive the ceiling first, then decide whether to write code. This round's output is not code but two veto records with numbers in them — "documented skip" is this campaign's formal disposition category: the numbers, the veto mechanism, and the retry conditions go into the record, the lever is terminated, and it will not come back next session as an "obviously doable" low-hanging fruit.

Baseline sanity first: in a window with a co-tenant running, the measurement read 3144.4–3151.4 (against 3181 on a quiet machine, recorded in footnote 2), and the binary cmp-matched HEAD — confirming the measurement object had not drifted, and all readings were taken on the same object.

2. Principle — the GPU mechanism

2.1 The roofline-bound-before-coding method

A bandwidth-bound kernel's time lower bound is:

t_min = bytes_min / BW_ceiling

GB10's trap is that ncu has no dram__* counters — the "measure the DRAM bytes" road does not exist, so bytes_min must be analytically derived from the access pattern, with ncu's sector counts used only to prove the access pattern (width, coalescing), never to count bytes. BW_ceiling takes the spec value of 273 GB/s. That choice biases the bound in the right, conservative direction: real achievable bandwidth ≤ spec, so the true headroom can only be smaller than computed. The inference chain then closes: if even a spec-perfect kernel cannot clear the bar, no implementation can — that is the "bound kills the lever" logic, and it is far cheaper than "implement, then measure": the whole chain is one ncu run plus some arithmetic.

2.2 The swiglu traffic audit: why the proposal's estimate was 2× low

swiglu_quant_nw_f32_t (r52's mode-2 kernel) does, per element: read gate's f32, read up's f32, compute silu·mul, quantize into the q8 plane. Its access shape is hard-coded in the kernel — one token row per block, one float4 per lane:

const float4* g4 = reinterpret_cast<const float4*>(gate + (size_t)t * dim);
const float4* u4 = reinterpret_cast<const float4*>(up + (size_t)t * dim);
float4 gv = g4[f];
float4 uv = u4[f];
v0 = (gv.x / (1.0f + expf(-gv.x))) * uv.x;   // silu(g)*u, per float4

ncu's verdict has two parts:

  • The read side is f32, full stop. Measured read traffic is 502.2 MB = exactly 2 × nt × dim × 4 B — that equation is itself the proof: if the reads were f16 or q8, the bytes would be half or a quarter. The proposal's traffic estimate counted a narrower width and came in a full 2× low; implement with that estimate and you only discover the bar was never reachable after finishing.
  • Vectorization is already maxed. sectors/request = 16, and arithmetic fits it exactly: one warp request = 32 lanes × 16 B (float4) = 512 B = 16 × 32 B sectors. Every byte is already inside a used transfer; there is no "switch to vectorized reads" headroom to mine.

The write side's minimum bytes derive too: mode 2 writes no f32 output (r52's core), only the q8 plane and the sda plane — 36 B per 32-element block (32 int8s + the f16 dsc) plus each block's sda share, totaling ≈ 70.9 MB. So the 573.1 MB composition is checkable: read 502.2 (the proposal estimated half of that) + write ≈ 70.9.

Total minimum DRAM traffic 573.1 MB, measured duration 2.367 ms → 242 GB/s = 89% of roofline. 94.6% occupancy, with a stall profile of a latency-bound pure stream — this is what "an already-maximized stream" looks like: no coalescible accesses, no raisable occupancy; the remaining 11% is the latency gap inherent near DRAM, not something the kernel's shape can claw back. So the ceiling gain:

573.1 MB / 273 GB/s = 2.099 ms (ideal)
2.367 − 2.099 = 0.268 ms per launch
× all swiglu launches ≈ 7.2 ms ≈ +0.74% whole-prefill  <  +1.5% bar

Even a perfect kernel falls 0.76 percentage points short — the skip is computed, not felt.

The method's applicability boundary is also written down: this bound is only tight for bandwidth-bound kernels. If ncu shows low occupancy, or stalls stuck on ALU or synchronization (rather than long_scoreboard-class memory stalls), the kernel is latency/compute-bound, the byte lower bound no longer represents achievable time, and you must switch to occupancy/dependency-chain analysis. swiglu's 94.6% occupancy + pure-stream stall profile falls exactly inside the bound's domain — part of "read the profile first, then pick the analysis tool."

2.3 The one-shot capture account: three ingredients inside the 7.84 ms idle

First the capture mechanism: stream capture records the kernels/copies issued in sequence on a stream inside the window as a graph, and one cudaGraphLaunch replays all of it — at replay the host no longer walks the driver per launch. It removes the per-launch host overhead and gaps, and pays off more the more times the same graph is replayed. Decode lives on it precisely because the same decode graph replays hundreds of times; R3-B's 3-run protocol was likewise designed for "the same nt prefill appearing repeatedly" (capture on the 3rd occurrence, cost amortized by subsequent replays). A one-shot prefill occurs exactly once, so nothing gets amortized.

What it cannot remove, and cannot digest, nsys splits the 7.84 ms idle into three classes:

Ingredientmagnitudecan capture reclaim it?
two one-time ~3 ms host stalls (minfer's own host code, not inside CUDA APIs)~6 msno — they live in host logic outside the capture window; replay is irrelevant to them, and they happen only once
mid-window cudaMalloc0.78 msno — and it is capture-ILLEGAL: synchronous allocation and other potentially-implicitly-synchronizing APIs are forbidden inside a stream-capture window, so the graph cannot even be built; legalizing it means changing the allocator (the pre-grow/pre-warm route)
recurring inter-launch gaps~0.1%yes — the only thing replay can remove

The reclaimable part is ≤ 0.1%, an order of magnitude below the bar. The capture mechanism itself is not wrong — what's wrong is aiming it at a scenario whose replay count is 1.

The comparison can be made concrete: the decode campaign later quantified the pool replay can save — the graph-gap pool is on the order of 2 µs/launch (D4-4's PDL probe went after exactly it, ultimately abandoned over the co-residency tax). Reasoning backwards from that number to the one-shot prefill: even if all recurring gaps were reclaimable, their absolute size lives in that 2 µs/launch pool — consistent with nsys's ~0.1%.

3. Implementation

3.1 Design choices (why this shape and not another)

  • The bar registers first; the numbers speak after: the +1.5% is fixed in writing before measurement; once the two levers' bounds come out, the conclusions follow automatically, with no sunk cost of "build it and see if it's good." This is isomorphic to another standing campaign discipline: the "pre-registered bars" of many later rounds (e.g. D4-3's ≤~32 µs bar) are the same practice.
  • Attribution freshness first: r47 had just overturned r37's attribution table (q6_K fell from 51.2% to 15.4% — decompositions go stale as levers land). r55's audit ran on the converged shape (the tree after r53/r54), so neither the object nor the numbers were stale; this is itself a transferable practice — any roofline/attribution audit should first ask "whose engine version is this decomposition of."
  • Skips get recorded too: the measured numbers, the veto mechanism, and the conditions worth a retry all go into the record. A documented skip's value is terminating a lever — this doc's §5 retry conditions are written for exactly that.
  • Zero tree changes: this round touches no source file, and the baseline binary is cmp-verified against HEAD — every measurement is taken on a clean object; 83c3c67 contains only the record.

3.2 Key code

The audited mode-2 swiglu kernel (src/cuda_kernels.cu) — note it is already "fused all the way": reads float4, writes packed q8, no f32 round trip (the f32 write is r51's mode 1; r52's mode 2 saved that too):

__global__ void swiglu_quant_nw_f32_t(
    const float* __restrict__ gate,
    const float* __restrict__ up,
    uint8_t* __restrict__ yqs,   // [ntb][nchunk][2048]
    uint8_t* __restrict__ ysda,  // [ntb][nchunk][256]
    int dim, int nt, int nchunk, int ntb
) {
    const int t = blockIdx.x;  // one token row per block (grid = ntb*64)
    ...
        if (active) {
            const float4* g4 = reinterpret_cast<const float4*>(gate + (size_t)t * dim);
            const float4* u4 = reinterpret_cast<const float4*>(up + (size_t)t * dim);
            float4 gv = g4[f];
            float4 uv = u4[f];
            v0 = (gv.x / (1.0f + expf(-gv.x))) * uv.x;   // silu(g)*u, per float4
            ...
        }
        float am = fmaxf(fmaxf(fabsf(v0), fabsf(v1)), fmaxf(fabsf(v2), fabsf(v3)));
        // 8-lane group reduce ... dsc = am/127 ... rintf quantization into the packed word

The question the roofline audit asks: how else could a "perfect kernel" change it? Item by item: the read bytes don't change (the f32 inputs are what they are; the width is not this kernel's choice); the write bytes don't change (the q8 plane is the downstream MMQ's fixed input format); the float4 accesses already push sectors/request to 16; occupancy is already 94.6%. No bytes left to change means no time left to mine.

The capture-side window primitive (src/cuda.rs) — the audit confirms it is per-stream, and any host-side sync/allocation inside the window is illegal:

#![allow(unused)]
fn main() {
pub fn graph_begin_capture(&self) -> bool {
    let stream = self.stream();
    let err = unsafe { cudaStreamBeginCapture(stream, 1) };
    ...
}
}

That 0.78 ms cudaMalloc in the nsys trace lands while this window is open — the problem is not a slow kernel, it is that the window cannot legally open. To eat that bite, the allocation must first become pre-grow (r59's rider later took that road), not by forcing capture on top.

"Prefill never enters a capture window" is an explicit invariant in the current tree (src/cuda.rs, at the f16 GEMM dispatch):

#![allow(unused)]
fn main() {
// Prefill never enters a CUDA Graph capture window (8g①
// decode-only gate), so the on-demand scratch grow is safe.
if nt > 1 && id <= 8192 {
    let q8 = Self::get_or_grow(
        &self.buf_q8_prefill,
        nt * (id / 32) * Q8B,
    );
    ...
}

The 3-run protocol and its supporting conventions also live in prefill_mmq's doc comment (same file, current tree):

#![allow(unused)]
fn main() {
/// R1: int8 MMQ prefill GEMM — quantize activations to q8_0 (pad40
/// blocks; ...) ... The q8 scratch follows the same
/// grow-on-demand lifecycle as the f16 path's buf_f16_x: the 3-run
/// capture protocol sizes it before the capture window opens.
pub fn prefill_mmq(
}

That comment also explains why an on-demand scratch grow is allowed on the prefill path — capture-illegal allocations only matter inside a capture window, and prefill never enters one; the 3-run protocol sizes the buffers before the window opens. What r55's "one-shot prefill capture" proposal was really asking: should this 8g① invariant be overturned for the CLI one-shot scenario? With the reclaimable part ≤ 0.1%, the answer is no — and overturning it would first require fixing every grow point inside the window (each one a potential capture-illegal allocation).

3.3 Pitfalls

  • The proposal's traffic estimate was 2× low: it estimated the read side at a narrow width. One ncu run, and the identity 502.2 = 2 × nt × dim × 4 sentenced it outright. Had it been implemented first and measured later, a full round of implementation + debugging would have been wasted on an unwinnable position — and implementers tend to find reasons to keep investing after a "so close to the bar" result.
  • GB10 has no dram__* counters: "measured bytes" does not exist; only analytic derivation + ncu sector counts as corroboration. This is a platform limitation — a bound must state its derivation chain explicitly, otherwise the bound itself cannot be re-checked.
  • Capture legality only becomes visible by walking nsys ms by ms: the mid-window cudaMalloc is perfectly legal in code review (the allocator belongs there) and illegal only in the capture context; "capture-illegal" is a runtime fact no type system will catch.
  • Baseline sanity in a co-tenant window: 3144.4–3151.4 vs 3181 on a quiet machine — confirm the object hasn't drifted before the numbers mean anything (footnote 2's accounting).

4. Verification

  • ncu sectors/request = 16: proves the accesses are maximally vectorized — guards against the illusion that "vectorization headroom remains."
  • byte-identity cross-check: 502.2 MB ≡ 2 × nt × dim × 4 B — guards against misjudging the read width (exactly where the proposal's estimate went wrong). Same on the write side: 573.1 − 502.2 = 70.9 MB matches the q8/sda plane arithmetic — guards against a missing term in the total traffic.
  • baseline cmp check: the binary matches HEAD — guards against measuring the wrong object.
  • nsys ingredient classification: every idle ms tagged one-time / recurring / capture-illegal — guards against the linear extrapolation that "all 7.84 ms could be captured away."
  • pre-registered bar: the +1.5% fixed before measurement — guards against moving the goalposts after the fact.

5. Results

Both levers are closed; the numbers:

Levermeasuredceiling gainverdict
fused-swiglu rewrite573.1 MB / 2.367 ms = 242 GB/s = 89% of roofline (94.6% occ, a latency-limited pure stream)≤ +0.74% (0.268 ms × all launches)skip: even a perfect kernel is < the +1.5% bar
one-shot prefill captureidle 7.84 ms = one-time host stalls ~6 ms + cudaMalloc 0.78 ms (capture-illegal) + recurring gaps ~0.1%≤ +0.1%skip: the reclaimable part is an order of magnitude below the bar

Residual leads on the record (left for later sessions): rms_nw at 153 GB/s = 56% of roofline (ideal +0.98%); the host stalls deserve a root-cause pass; tail pre-grow +0.1%. Note that rms_nw, this "last lead," has an ideal gain of +0.98% that is itself below the +1.5% bar — done alone it also fails the gate, and it is only worth being a member of some future bundle. That is the full meaning of the convergence verdict: not just "no lever ≥ bar," but "even the remaining leads' ceilings summed cannot constitute an independent next step."

The campaign convergence statement (this round's core output): whole-prefill tightened from r37's 2.15× to 3181 vs 3324.4 = 1.05× (r53/r54). The wall decomposition and each component's closing state:

Wall componentsharestate
q4_K GEMM63.2%closed — unless the q8_1 GEMM-prologue fusion, a step-function, is attempted
q6_K GEMM15.4%closed
fused producers8.8%swiglu closed this round; rms is the last lead (56% of roofline)
FA5.3%the 2.43× is taken
host~1%stall root-cause is a recorded pass

No identified lever is ≥ +1.5%; the next tier is the step-function q8_1 prologue fusion (llama.cpp's route of folding activation quantization into the GEMM main loop). Campaign verdict: CONVERGED (under the current gate set — r56/r59 later reopened it with the bundling mechanism, see docs 59 and 62).

Retry conditions (otherwise these two skips would be re-proposed forever): swiglu — only if the q8_1 prologue fusion changes its traffic pattern (the read side stops being independent f32 streams) or the bar is lowered below +0.74%; capture — only if the allocator is capture-legalized (pre-grow) AND the host stall is root-caused into the launch path, and even then the ceiling remains the recurring gaps ≈ 0.1%.

6. Lessons

  1. Roofline-bound-before-coding: derive the byte lower bound first, then decide whether to write code — if a perfect kernel cannot clear the bar, no implementation can. Both levers cost one ncu run, not one implementation cycle.
  2. A proposal's traffic estimate must be audited by ncu first: the swiglu proposal underestimated 2× (narrow-width read side); the cost of an estimate error surviving an entire implementation cycle far exceeds one audit.
  3. Capture's gain accounting is the recurring part: one-time host stalls and capture-illegal operations are not in the reclaimable set; for a replay-count-of-1 scenario, capture's ceiling is the inter-launch gaps.
  4. A documented skip is a deliverable: numbers + veto mechanism + retry conditions in the record — only then is a lever truly closed.

← 57 · Index · 59 →

59 · r56 — q6_K A-side bundle: A cp.async + W_dsc f32 plane (LANDED)

Result: whole prefill 3138.6 → 3212.5 tok/s (+2.35%, vs-llama 1.035×); ffn_down kernel 12.76 → 12.01 ms (−5.9%), attn_v −4.2%; both the <2,true> and <2,false> instantiations reach 80 regs / 0 stack with LDGSTS in the SASS; the cost is the +363.2 MB W_dsc plane. parity ×3, the greedy 453-character stream byte-identical, liveness 27×/0. Commit: 4cf7c74 (code) + 29084de (record). Date: 2026-09-06.

1. Background — where things stood

r53 closed out the q6_K B side as a bundle (W_exp plane + cp.async, +5.03%), and r54 fitted the exit valve. Session E's basket list had two items left: item 1 was FA KV staging double buffering (next doc, REVERTED), and item 2 is this doc — q6_K's two remaining residuals.

Both residuals had been named in r43's PC-sampling attribution. At that time the q6_K kernel's stall profile was three pieces: B-expand recombination 45% + A-side staging STS 28% + dsc I2F consumption 26%. r44/r45 killed the 45% piece (each wall-neutral alone), and r53 bundled them into the wall. So after r53, the record carried §11.32's verdict verbatim: "A-side staging STS (~28%) + the dsc I2F consumer (~26%) — a W_dsc f32 plane would be the symmetric next bundle member, and an A-side cp.async redo could now compose".

Two terms need expanding:

  • The A side: each kt, the NB-BT kernel must move the activation side's qa8 (the pad40 transposed q8 plane) and sda (the d/ssum scale plane) into shared memory. In the r53 shape, B was already cp.async while A was still synchronous LDG→STS — the global-read latency at the top of every staging phase stood exposed (r45 quantified this mechanism's kernel-level payoff back then: kernel −10.2%, long_scoreboard stall −18%, but the wall only −0.34% — the q6_K GEMM was no longer the wall at the time).
  • The dsc consumer: every (chunk,row) reads the int8 scale at blk[192+s0] and the f16 d at blk[208] from the raw weight block, converts via I2F, then multiplies — 26% of r43's stall mass sat on the I2F. r42 tried a stage-level scale read and got NEUTRAL: cutting dsc bytes cannot cut dsc latency; the stall is on the consumer side.

r45's history is this doc's foreshadowing: A-side cp.async was tried once alone, was REVERTED, on the grounds that "the q6_K GEMM is no longer the bottleneck — a faster kernel not on the wall can never reach the wall." The bundle thesis predicted: once B becomes a pure copy and the wall moves onto the A-side wait and dsc consumption, the same mechanism turns positive again. r56 is that prediction's verification. The three mechanisms' "provenance" and "what they remove" line up into a table:

Mechanismwhat it removesquantitative evidence
r44 W_exp dense planethe B recombination's WORKkernel −10.9% (wall −0.42% at the time)
r45 A-side cp.asyncthe staging's WAITkernel −10.2%, longsb −18% (wall −0.34% at the time)
r53 B bundleboth of the above togetherwhole prefill +5.03%
r56 A cp.async + W_dscthe remaining A WAIT + dsc I2F WORKwhole prefill +2.35% (this doc)

2. Principle — the GPU mechanism

2.1 The pipeline structure after the bundle

After r56, each kt's staging becomes one cp.async stream: all four planes — A's qa8, A's sda, B's W_exp, and dsc — are issued with explicit-PTX gemm_cp16, and gemm_cp_commit() moves to the end of the RAW_STAGE macro — exactly one commit group per kt binds the four planes together. The waits in the main loop are unconditional:

  • middle tiles: gemm_cp_wait1() — two groups in flight (kt's and kt+1's); wait until only kt+1's remains, kt's copies (in-order group completion) have landed, and kt+1's copies keep flying under kt's compute;
  • the last tile: gemm_cp_wait0() — drain every group.

The unconditional wait is a correctness requirement: group completion is in issue order, and any data-conditional skipping of the wait would fork the group ordering between the two EXP states.

2.2 The W_dsc plane: turning the 26% I2F stall into one 16 B copy

At registration, precompute each (chunk c, row j) dsc pair: out[(c*od + j)*8 .. +8] = float2(d·sc[2(c&7)], d·sc[2(c&7)+1]). The chunk-major layout is the key — the kernel's per-kt staging window is exactly "the contiguous MMQ_NBJ rows of chunk c0+kd," and under chunk-major those bytes are contiguous, so one gemm_cp16 per pair (2 float2 = 16 B) moves them. The raw path's bill: per (row, chunk), 2 non-adjacent byte reads (int8 scale) + 1 f16 read + 3 I2F + 2 f32 multiplies; the plane path: 1 vectorized 16 B copy, zero ALU.

Plane size: nchunk × od × 8 B = (id/32) · od · 8 = od · id / 4. The q6_K tensors of 7B q4_k_m total 363.2 MB — a quarter of r53's 1.52 GB, consistent with the arithmetic "one f32 pair per 4 B bytes."

Why chunk-major is the layout's right answer: the kernel's per-kt staging window is "rows j0 .. j0+MMQ_NBJ of chunk c0+kd." If the plane were row-major (plane[j*nchunk + c]), adjacent rows of the same chunk would sit nchunk × 8 bytes apart in memory — one copy per row, impossible to coalesce into 16 B. Chunk-major (plane[(c*od + j)*8]) puts a chunk's rows together: row j's 8 bytes are immediately followed by row j+1's 8 bytes — one pair per row, two pairs per chunk — exactly composing the 16 B cp.async transfer unit. The layout follows the consumption window; this is the same design law repeatedly verified since r31 (the q-major sda repack).

2.3 Why bit-identity is constructed

The f32 stored in the plane must be bit-identical to the f32 the r41 scalar path computes on the fly, otherwise parity stops being a formality and becomes a bet. The construction guarantee comes from three points: f16→f32 is exact (half::f16::to_f32, i.e. __half2float); i8→f32 is exact; exactly one f32 multiply, with FMA contraction forbidden on both sides — d * sc0 is a single multiply, and any contraction or reordering moves the last bit. The kernel side changed only "where this product comes from," not "how it is computed."

2.4 The even-od gate

The kernel issues 16 B chunks by row pairs and zero-fills whole pairs: a pair is either fully inside od or fully outside it (src-size 0 zero-fill). A concrete boundary: at od = 4N the last pair is rows (4N−2, 4N−1) and full = (j+1 < od) holds for both; at od = 4N+1 the last pair is (4N, 4N+1), where row 4N+1 does not exist — if it were still issued, row 4N's scale would be written "as a pair" with wrong content or overwritten by the zero fill, and the scale lost would be exactly that real odd row. Hence the registration gate adds od % 2 == 0, and odd tensors map-miss back to the r41 scalar path. The zero-fill granularity must equal the data structure's granularity (here = the row pair) — another form of the r44-class stride accident, blocked at registration time.

3. Implementation

3.1 Design choices (why this shape and not another)

  • One round bundles two mechanisms, not two rounds: r53 already proved that changes "overlapping in traffic, non-overlapping in mechanism" can add (r44's "remove the work" + r45's "remove the wait"); A-cp.async (remove the wait) and W_dsc (remove the work) are exactly that thesis's next pair.
  • All the scaffolding is reused from r53: the geometry-encoded sibling name ({name}__dsc{od}x{id}, guarding against a same-name different-shape stale plane being silently reused), the map keyed by the padded weight's device pointer, and on allocation failure an empty map + a loud eprintln once per process.
  • Riding r54's gate: W_dsc registration hangs under Q6K_NB && Q6K_EXP != "0" && id%256==0 (+ od%2==0) — EXP=0 opts out of all plane memory in one move (1.52 GB + 363 MB), so r54's "exit valve" semantics are not quietly broken by the new plane.
  • A null plane = the r41 scalar path, byte-identical: raw weights, allocation failure, and odd od all land here.

3.2 Key code

The cp.async primitives and group ops (src/cuda_kernels.cu) — the src-size qualifier is the zero-fill mechanism:

__device__ __forceinline__ void gemm_cp16(__half* smem_dst, const __half* gsrc, bool full) {
    unsigned d = (unsigned)__cvta_generic_to_shared(smem_dst);
    int sz = full ? 16 : 0; // src-size 0 => zero-fill the 16B chunk
    asm volatile("cp.async.cg.shared.global [%0], [%1], 16, %2;\n" ::"r"(d),
                 "l"(gsrc), "r"(sz));
}
__device__ __forceinline__ void gemm_cp_commit() { asm volatile("cp.async.commit_group;\n"); }
__device__ __forceinline__ void gemm_cp_wait1() { asm volatile("cp.async.wait_group 1;\n"); }
__device__ __forceinline__ void gemm_cp_wait0() { asm volatile("cp.async.wait_group 0;\n"); }

The A side's cp.async conversion (inside the RAW_STAGE_Q6K_BT macro; r53's synchronous LDG→STS replaced; bytes unchanged — the prepass already fills the padded rows):

/* ---- A: r56 cp.async bulk staging of the pre-transposed qa8/sda --*/
/* (r45's mechanism on top of r53: the sync LDG->STS exposed its     */
/* global latency at the top of every staging phase; cp.async hands  */
/* it to the async unit and the group wait below hides it under the  */
/* previous tile's compute. Bytes identical - the plane is always    */
/* full: the prepass zero-fills the padded rows.)                    */
{
    const size_t qbase = ((size_t)blockIdx.x * nchunk + (size_t)(kt) * KDR) * MMQ_A_QASZ;
    for (int off = threadIdx.x; off < (KDR * MMQ_NBI * 32) / 16; off += blockDim.x)
        gemm_cp16((__half*)(void*)(qa8b + (size_t)off * 16),
                  (const __half*)(const void*)(qa8g + qbase + (size_t)off * 16), true);
    const size_t sbase = ((size_t)blockIdx.x * nchunk + (size_t)(kt) * KDR) * MMQ_A_SDASZ;
    for (int off = threadIdx.x; off < (KDR * MMQ_NBI * 4) / 16; off += blockDim.x)
        gemm_cp16((__half*)(void*)(sdaqb + (size_t)off * 16),
                  (const __half*)(const void*)(sdag + sbase + (size_t)off * 16), true);
}

The dsc pairs' 16 B streamed copy (same macro; the scalar else branch keeps r41 verbatim):

/* ---- B: dsc pair (d*sc[2c%16], d*sc[(2c+1)%16]) per (chunk,row) ----*/
/* r56: W_dsc f32 plane (registration-time precompute; chunk-major    */
/* layout plane[c*od + j] = float2(d*sc0, d*sc1)) turns the scalar    */
/* blk[192+..]/blk[208] loads + I2F (a leading r43 residual stall     */
/* post-r53) into a contiguous 16-B cp.async stream inside the same   */
/* per-kt commit group. Null plane = the r41 scalar path.             */
{
    const int c0d = (kt) * KDR;
    if (W_dsc != nullptr) {
        const int nc2 = MMQ_NBJ / 2; /* 16-B chunks (2 float2) per kd */
        for (int g = threadIdx.x; g < KDR * nc2; g += blockDim.x) {
            const int kdd = g / nc2, m = g % nc2;
            const int j = j0 + 2 * m;
            /* od even (registration gate) => a pair is either fully */
            /* valid or fully beyond od (src-size zero-fill).        */
            const bool full = (j + 1 < od);
            gemm_cp16((__half*)(void*)(sdsb + (size_t)kdd * MMQ_NBJ + 2 * m),
                      (const __half*)(const void*)(W_dsc
                          + ((size_t)(c0d + kdd) * (size_t)od + (size_t)j) * 8),
                      full);
        }
    } else { /* r41 scalar: d = h2f(blk[208]); dsc = d * (i8)blk[192+s0] ... */ }
}
...
gemm_cp_commit();   /* r56: commit moved to the END of RAW_STAGE — one group per kt covers A+B+dsc */

The main loop's wait structure (middle tiles wait1 / last tile wait0):

RAW_STAGE_Q6K_BT(0, 0);
__syncthreads();
int buf = 0;
for (int kt = 0; kt < nktile; ++kt, buf ^= 1) {
    if (kt + 1 < nktile) {
        RAW_STAGE_Q6K_BT(kt + 1, buf ^ 1);
        /* two groups pending; wait until only kt+1's remains — group(kt)
         * (r56: A + dsc too, not just B) has landed, buf^1's copies stay
         * in flight under kt's compute (in-order group completion). */
        gemm_cp_wait1();
    } else {
        gemm_cp_wait0();  // last tile: drain every outstanding group
    }
    __syncthreads();  /* r53/r56: cross-thread visibility of the kt buffer */

The host-side expand_q6k_dsc (src/cuda.rs) — the comment's "bit-identical by construction" three elements are §2.3:

#![allow(unused)]
fn main() {
/// Output: `nchunk * od * 8` bytes, `out[(c*od + j)*8..+8]` =
/// float2(d*sc[2(c&7)], d*sc[2(c&7)+1]) — chunk-major so the kernel's
/// per-kt staging ... is a pure 16-B cp.async stream. Bit-identical to
/// the in-kernel r41 scalar computation: exact f16->f32 (half::f16, =
/// __half2float), exact i8->f32, one IEEE f32 multiply, no FMA
/// contraction on either side.
pub fn expand_q6k_dsc(padded: &[u8], od: usize, id: usize) -> Vec<u8> {
    ...
    let d = half::f16::from_bits(d_bits).to_f32();
    for cc in 0..8usize {
        let s0 = 2 * cc;
        let sc0 = blk[192 + s0] as i8 as f32;
        let sc1 = blk[192 + s0 + 1] as i8 as f32;
        let idx = ((sb * 8 + cc) * od + j) * 8;
        out[idx..idx + 4].copy_from_slice(&(d * sc0).to_bits().to_le_bytes());
        out[idx + 4..idx + 8].copy_from_slice(&(d * sc1).to_bits().to_le_bytes());
    }
}

The liveness label's A-side extension (current tree src/cuda.rs) — r56 also adds the A/dsc paths to the label r53 erected ("a fallback-correct fast path needs a visible label, parity cannot see it"); the census in §4 reads this composite output:

#![allow(unused)]
fn main() {
// r56: name the A/dsc staging paths too (liveness
// check per the r53 lesson — a fallback-correct fast
// path needs a visible label, parity cannot see it).
let a = if !w_dsc.is_null() {
    "A=cp.async DSC=f32-plane"
} else {
    "A=cp.async DSC=scalar"
};
eprintln!(
    "minfer/cuda: mmq raw NB-BT q6_K kernel active \
     (r56 {}, r54 B={}, r39 KDR=2 double-buffer, \
     A-transpose)",
    a, b
);
}

3.3 Pitfalls

  • sudo -n strips environment variables (this doc's most retellable pit): running the profile elevated via plain sudo -n ncu ... has sudo strip the user environment by default — gate variables like MINFER_MMQ_Q6K_NB all vanish, and ncu silently profiles the legacy path: the numbers "look normal" but measure a kernel that is not even running. Symptom chain: the filter was set to the NB-BT q6_K kernel name → the profiled data's shape looked unfamiliar → open the ncu report's "Available Kernels" list and match the filter term: zero hits — the NB-BT q6_K kernel is not in the list at all, meaning the gate was off. Fix: sudo -E, inline the env in the command (sudo MINFER_MMQ_Q6K_NB=1 ncu ...), or just run as the same user.

  • r53's 24 B stack disappears: after r56 both instantiations land at 80 regs / 0 stack — planarizing the dsc incidentally unloads the scalar path's register pressure inside <2,false>; for the first time the two states' occupancy budgets match exactly.

  • The "false regression" of suite 167/1: cuda_conversation_multiturn_reuse failed — git stash bisection verified it fails the same way on a clean HEAD: pre-existing, environment-sensitive, unrelated to this change. Not chased.

  • A co-tenant outlier in the A/B distribution: with that point excluded, the distributions fully separate; without excluding it, the median gets dragged flat — the campaign's "distribution separation, not mean comparison" accounting saves the day again.

4. Verification

  • cuda_q6k_dsc_dense_byte_exact: independent scalar mirror vs the host expander + device-plane read-back, 3 shapes, 0 mismatches — guards against plane byte misalignment (the r44-class stride error).
  • parity ×3: guards against numeric path changes (per §2.3 this should be a construction guarantee; a routine confirmation).
  • greedy 453-character stream byte-identical: guards against accumulated drift the parity fixtures cannot cover.
  • liveness census 27× A=cp.async DSC=f32-plane, B=W_exp-cp.async, 0 fallbacks: guards against "parity/greedy all green but the fast path never ran" (r53's original lesson; r56 adds the A/dsc paths to the label too).
  • regs/stack and SASS check: both instantiations 80 regs / 0 stack, LDGSTS present: guards against the explicit-PTX trap — if LDGSTS is missing from the SASS, the compiler degraded cp.async into a synchronous copy.
  • ncu kernel level: ffn_down 12.76 → 12.01 ms (−5.9%), attn_v −4.2% — guards against attributing wall noise to kernel improvement.

5. Results

Metricbefore → afterΔ
whole prefill (distributions separated, one co-tenant outlier excluded)3138.6 → 3212.5 tok/s+2.35%
vs-llama (the 3324.4 anchor)1.05× → 1.035×—
ffn_down kernel12.76 → 12.01 ms−5.9%
attn_v kernel—−4.2%
device memory—+363.2 MB (W_dsc = od·id/4)
registers/stackr53's <2,true> 80/24B → both states 80 / 0—
suite167/1 (the only failure = cuda_conversation_multiturn_reuse, fails the same way on a clean HEAD, bisect-verified)—

The bundle thesis's second score: r45's mechanism landed exactly as predicted once the wall moved over — with B a pure copy, the A-side wait and dsc consumption are the wall, and dismantling both together moves the wall. Worth recording this score's "timing recipe": r45 (9825ffd, row 45) → (the wall moves) → r56 (4cf7c74), with the ten rounds r46–r55 in between; the mechanism was not forgotten in the interim, it was just waiting for the attribution to refresh. This differs from "abandoning a failed direction" — what should be abandoned is the hypothesis ("the A-side wait is now worth dismantling"), and what should be kept is the verified mechanism (cp.async staging itself).

This score also draws the boundary: basket logic is not a universal transplant template — the immediately following r57 (FA KV double buffering) and r58 (porting cp.async-db2 to q4_K) were both REVERTED (docs 60 and 61); a mechanism's value depends on what it replaces.

6. Lessons

  1. After the wall moves, yesterday's wall-neutral mechanism becomes today's positive gain — a reverted lever is worth retrying after the attribution refreshes, provided it is re-attributed (the r45 → r56 arc followed "look at where the wall is" throughout).
  2. A registration-time plane must be "bit-identical by construction": exact conversions, exactly one multiply, FMA contraction forbidden on both sides — bit-level consistency backstopped by tests will not survive compiler version changes.
  3. sudo stripping the env is profiling's silent poison: running ncu elevated strips the gate variables along with everything else, making you profile the legacy path while believing you're testing the new kernel; a zero-hit ncu "Available Kernels" list is its fingerprint.
  4. A paired-staging validity gate must align pair-wise (od even): the zero-fill granularity must equal the data structure's granularity; a half-valid boundary row silently loses data.

← 58 · Index · 60 →

60 · r57 — FA KV staging double buffering (REVERTED)

Result: both attempts fail the greedy-32 byte-identity gate — attempt 1 forked at token 19 (the fix uncovered a real bug in the attempt code: the tail-tile's stage→global O copy was wrapped in a warp guard); attempt 2, bug fixed, still forked, with the residual being r50-class inherent rounding drift: FA_TQ is also a tile size, so the "r50's caution doesn't apply" premise was falsified. REVERTED per the two-attempt stop rule; the FA kernel keeps the FA_TQ=64 / FA_TKV=32 / per-tile synchronous drain shape. Session E's basket net gain is left to item 2 alone (r56, +2.35%). Commit: c3268cc (record-only commit). Date: 2026-09-06. Code provenance note: the code attempts never entered the repository (a revert destroys them), so git show cannot reference the changes themselves; this doc's code excerpts are from the current tree (= the post-revert state), and the attempted shape is reconstructed by narration.

1. Background — where things stood

After r48 (FAP2) moved the FA prefill kernel's softmax into registers, the FA kernel went 5.16 → 2.12 ms (2.43×) and whole prefill +5.6%; r55's convergence statement had FA at 5.3% of the wall, the largest single residual outside the GEMMs. Session E's basket lined up five items: item 1 = FA KV staging double buffering (this doc), item 2 = the q6_K A-side bundle (r56, LANDED), items 3/4/5 = rms_nw roofline, host-stall root-cause, and tail pre-grow (all unreached — budget exhausted). FA went first because the q6_K line had just used cp.async + double buffering (KDR=2) to hide the "staging wait" inside compute (the r39/r53/r56 trio), and the same trick looked like it could transplant directly onto FA's KV movement.

The FA kernel's KV supply shape at the time (current tree, i.e. the in-service post-revert form): each KV tile iteration issues cp.async at its top but immediately follows with a wait_group 0 synchronous drain — DRAM latency is exposed once per KV tile iteration. This is exactly the target shape of the "remove the wait" (WAIT) class of lever as r45/r56 defined it:

  • in service: stage(kt) → commit → wait 0 → syncthreads → compute. The serial k-loop pays the full latency once per tile. The tile count also lives here: using the pp3314 anchor as an example, each (q-block, head)'s KV chain passes ~104 tiles (3314/32), and each tile's serial head = issue + DRAM latency + barrier — that is the pool "hide the latency" wants to eat.
  • the attempt: a prologue fetches tile 0 first; inside the loop the issue changes to kt+FA_TKV into buf^1 with wait_group 1 — kt+1's copies fly under kt's QK^T/softmax/P·V.

But there is a prior record that must be routed around: r50 (the FA_TKV 32→16 occupancy experiment) — it pushed to 3 blocks/SM, but the occupancy gain was eaten by the doubled per-tile sync/softmax overhead (−0.5%/−0.01%), and it lost greedy byte identity, concluding "FA_TKV reduction is a dead lever that also breaks byte-identity." Earlier still, r46 (FAP1) took FA_TKV 64→32 + S/P row padding to kernel −11% but only +0.27% whole-prefill (FA was not the wall-critical path then), REVERTED — the FA line was only truly closed by r48 (FAP2) register softmax. r57's premise was written explicitly: FA_TKV stays 32 (KV tile boundaries unmoved → each row's KV-column accumulation grouping and online-softmax rescale points unmoved), and only FA_TQ shrinks 64→48 to free shared memory for a second K/V buffer — "r50's caution is only about the KV tile and does not apply to TQ." This doc is the record of that premise being falsified.

The smem arithmetic (the design's starting point, checked item by item): in service, (FA_TQ + 2×FA_TKV) × (hd+8) × 2 = (64+64)×136×2 = 34,816 B; the attempt's composition is Q at 48 rows 48×136×2 = 13,056 B + two copies each of K/V 4 × 32×136×2 = 34,816 B, totaling 47,872 B, targeting 2 blocks/SM. FA_TQ=48 also carries a by-product: 128 threads = 4 warps and 48 rows = 3 sixteen-row blocks, so warp 3 no longer owns Q rows — it is assigned to staging full-time.

2. Principle — the GPU mechanism

2.1 What double buffering buys in this class of kernel

The q6_K NB-BT kernel's KDR=2 double buffering (r39) is a campaign-verified template: each kt has two staging planes, kt+1's global→smem copy overlaps kt's mma compute, and wait_group 1 waits only for "the previous group" to land (cp.async groups complete in issue order — a semantic guarantee). Its record on q6_K: 1568.7 → 1777.5 tok/s (+13.3%), attn_v kernel −19.7% — on the premise that staging latency was genuinely on q6_K's critical path then. The gain mechanism lifts the global-read latency (hundreds of cycles) off the serial chain and hands it to the async units. FA's KV supply is structurally isomorphic: the KV tile stream feeds the online-softmax main loop, and the main loop does three compute segments per tile (QK^T, softmax, P·V) — enough to shelter one copy. r57's bet was "isomorphic ⇒ same gain."

2.2 The in-service form's byte and latency account

The current-tree kernel's (src/cuda_kernels.cu) tile constants and smem layout:

#define FA_TQ 64
#define FA_TKV 32
// Shared layout (dynamic, ~35 KB — opt-in via cudaFuncSetAttribute):
//   Qs [64*hd] f16   q tile (scale folded in, f16 for the tensor-core QK^T)
//   Ks [FA_TKV*hd] f16   K tile      Vs [FA_TKV*hd] f16  V tile

The staging helper is 16 B cp.async (introduced in P5.3, fc07c04), with out-of-range rows zero-filled via src-size 0:

__device__ __forceinline__ void fa_stage_kv_async(
    const __half* __restrict__ k, const __half* __restrict__ v,
    __half* Ks, __half* Vs, int kt, int kv_end, ...) {
#if __CUDA_ARCH__ >= 800
    for (int c = tid; c < FA_TKV * hd / 8; c += nthreads) {
        int r = (c * 8) / hd, d = (c * 8) % hd;
        int p = kt + r;
        bool full = p < kv_end;
        unsigned kd = (unsigned)__cvta_generic_to_shared(Ks + r * sstr + d);
        ...
        int sz = full ? 16 : 0;
        asm volatile("cp.async.cg.shared.global [%0], [%1], 16, %2;\n" ::"r"(kd),
                     "l"(k + (size_t)p * stride_kv + hk * hd + d), "r"(sz));
        asm volatile("cp.async.cg.shared.global [%0], [%1], 16, %2;\n" ::"r"(vd),
                     "l"(v + (size_t)p * stride_kv + hk * hd + d), "r"(sz));
    }

But the main loop's call pattern is single-buffered + fully synchronous drain — cp.async merely changes the copy's issuer, and not a single beat of latency is hidden:

for (int kt = 0; kt < kv_end; kt += FA_TKV) {
    // stage K/V tile (padded stride, zero-filled beyond kv_end)
    fa_stage_kv_async(k, v, Ks, Vs, kt, kv_end, hk, hd, stride_kv, sstr, tid, 128);
#if __CUDA_ARCH__ >= 800
    asm volatile("cp.async.commit_group;\n");
    asm volatile("cp.async.wait_group 0;\n");   // ← drains immediately: zero overlap
#endif
    __syncthreads();
    // S = Q·K^T via wmma ... in service: the full DRAM latency exposed once at the top of every tile

The launcher-side smem account (current tree src/cuda_kernels.cu) — the 34,816 B is computed here, and the r57 attempt changed exactly this one expression:

int launch_fa_prefill_f16kv(...) {
    // Qs + Ks + Vs only (S/P no longer go through shared memory). sstr = hd+8
    // padding; Ks/Vs are FA_TKV rows (the r46 launcher's 3*FA_TQ bug is gone).
    size_t smem = ((size_t)FA_TQ + 2 * FA_TKV) * (hd + 8) * 2;
    //   = (64 + 64) × 136 × 2 = 34,816 B (in service)
    //   attempt = (48 + 4×32) × 136 × 2 = 47,872 B (Q shrunk rows + K/V double buffering)

(A note in passing: the comment block above the helper says "Overlapped with the previous tile's … via double buffering," which does not match this wait-0 call site — it is a stale historical comment; in an audit, trust the call site.)

2.3 Why byte-identity "should" hold (and where it actually broke)

The per-row float accumulation chain depends on two things: the KV tile boundaries (which fix the online-softmax m/l rescale points) and each row's KV-column-to-lane grouping (which fixes the summation's float associativity order). The former deserves a sentence of expansion: at the end of every KV tile, online-softmax performs m_new = max(m_old, tile_max) and an O rescale by exp(m_old − m_new) — every time a KV tile boundary moves, the rescale points move and the float chain changes shape; that is why r50 lost identity the moment it touched TKV. The premise assumed both of these are determined by FA_TKV alone, and that TQ only changes "which warp owns which 16 rows" — pure scheduling, outside the float chain. The measured outcome (attempt 2 still forks) overturns it: any change to tile geometry perturbs the summation grouping somewhere in the float chain, TQ included; the specific bit path was not chased to the bottom — the two-attempt stop rule triggered first. From r50 (TKV) to r57 (TQ), two independent samples support the conclusion: for this kernel, the tile constants are part of the numeric contract.

The conclusion also has a campaign-level corroboration: D3-4's tolerance calibration measured the end-to-end magnitude of any accumulation-order change at max|Δlogits| 0.38/0.39 (14B/7B) — that is, if a tile change really moved the float chain, the output would drift systematically at this magnitude and the greedy stream would fork sooner or later; there is no mild version where "the tile changed but only the last bit wobbles." r57's fork shape (the greedy-32 stream going wholly off after token 19) matches this magnitude exactly.

2.4 The attempted shape (reconstructed by narration)

A four-part pipelined structure, each item paired with the in-service form's change point:

  1. prologue: before entering the KV loop, stage tile 0 into buf 0 (commit + wait_group 0 + syncthreads) — the first data is in place, and the loop body no longer has the "wait, then compute" serial head.
  2. in-loop issue: each iteration issues the K/V copies of kt + FA_TKV into buf^1 (double-buffer alternation), commited as a group.
  3. the wait: wait_group 1 — kt's group has landed (groups complete in issue order) and kt+1's copies stay in flight.
  4. warp division of labor: FA_TQ=48 → 3 sixteen-row blocks; warps 0–2 own Q rows and do QK^T/softmax/P·V, while warp 3 does staging full-time (issue and zero-fill) — otherwise it idles through the whole loop.

Total smem 47,872 B (§1's arithmetic), targeting 2 blocks/SM. The numeric expectation: FA_TKV=32 unmoved → per-row accumulation order unmoved → byte-identical — an expectation that died in §2.3.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Touch TQ only, not TKV: quarantining r50's prior record outside "the KV tile boundary," the presumed float-chain entry point — this isolation assumption is exactly what got falsified, but it was the only design that could preserve byte-identity at the time.
  • The two-attempt stop rule applies in advance: a standing campaign discipline — at most two attempts per lever; if the second still hits a red gate, revert. It guards against the "so close" sunk-cost spiral; r57 is the rule's textbook execution (attempt 1 fixed a real bug, attempt 2 falsified the premise, stop).
  • Gates first, performance last: greedy-32 byte identity is the first gate, before any performance measurement — until the correctness gate is green, performance numbers are meaningless (this doc ends with no wall numbers precisely because the gate never went green). This was also the campaign's consistent gate order before the D series.

3.2 Key code

See §2.2 — the two excerpts of the in-service form (= the post-revert status quo) are "what the revert returned to." The attempted change's shape is narrated in §2.4; its diff has no commit to cite (see the code provenance note at the top).

3.3 Pitfalls

  • A real bug (in the attempt code, not in the in-service kernel): after attempt 1 forked, a section-by-section hunt found the tail-tile's stage→global O copy written inside a warp guard — tail rows outside the guard would never be written out, and greedy-32 forked at token 19. This bug explains all of attempt 1's forking, and the fix itself was correct; it was not an incumbent bug (the in-service kernel never forked).
  • The premise falsified: with the real bug fixed, attempt 2 still forked — the residual is §2.3's inherent rounding drift, independent of implementation quality. This kind of drift "cannot be fixed"; only the numeric contract can be changed.
  • Rebuild md5 is unusable as revert verification: the rebuilt binary's md5 after the revert differs from pre-attempt — nvcc builds are not bit-deterministic; whether a revert is clean is judged by behavior (the greedy stream + perf sanity), not by md5.

4. Verification

  • greedy-32 byte identity: the first gate and the fatal one — attempt 1 used it to catch the warp-guard bug, attempt 2 used it to falsify TQ-independence. Guards against: any float-chain perturbation that "shouldn't change in theory."
  • perf sanity (post-revert): 3222.4 ≈ the landed median 3212.5, and the greedy stream byte-identical to the record — guards against "an unclean revert."
  • the two-attempt stop: stop at the second red gate; no third TQ/buffering combination is chased — guards against sunk-cost-driven endless debugging.

5. Results

Itemresult
attempt 1greedy-32 forked at token 19 → traced to a tail-tile O copy warp-guard bug in the attempt code (a real bug, fixed)
attempt 2still forked after the bug fix → r50-class inherent rounding drift; FA_TQ is also a tile size
dispositionREVERTED per the two-attempt stop rule; no wall numbers (the correctness gate never went green, so performance measurement would be meaningless)
revert verificationrebuild md5 differs (expected — builds are not bit-deterministic); the greedy stream matches the record; perf sanity 3222.4 ≈ 3212.5
basket settlementSession E's net gain rests on item 2 alone (r56, +2.35%); items 3/4/5 (rms_nw roofline, host stalls, tail pre-grow) unreached — budget exhausted

The veto mechanism: this lever's value proposition was "a pure-overlap gain with zero numeric risk," and that proposition rests on "TQ stays out of the float chain" — falsified outright by attempt 2. With byte-identity as a hard gate, the mechanism is unconditionally unlandable.

Retry conditions: double buffering becomes discussable again only when some FA change deliberately gives up byte-identity — at that point it passes the tolerance gates as part of that change, not as an independent "free" lever. The tolerance gate's concrete shape already has a calibrated precedent in the campaign (D3-4 h4w): an end-to-end max|Δlogits| 0.30–0.39-class tolerance (the inherent magnitude of any accumulation-order change), a hard argmax gate, greedy flips attributed one by one (sampler knife-edge vs kernel drift), and a temp-0.8 control group. In other words, "accepting non-identical output" is not a relaxation of verification but an exchange for a more expensive yet executable verification.

6. Lessons

  1. r50's caution generalized into a law: any FA tile-size change (TKV or TQ) breaks byte-identity — tile geometry is the float summation grouping itself; every tile constant is part of the numeric contract.
  2. A "free knob" is only a hypothesis until it has been shown to stay out of the accumulation chain: before building the mechanism, spend ten minutes verifying the invariance premise with one dump gate; r57 ran the order backwards and paid two full implementation rounds.
  3. Failed attempts also pay rent: the tail-tile O copy's warp-guard bug would have remained a landmine without this hunt — the red gate did not run in vain. The negative conclusion "FA_TQ is also a tile size" itself entered the campaign record, and it blocks all future "touching only TQ should be fine" variant proposals.
  4. Judge a revert by behavior, not md5: nvcc builds are not bit-deterministic; the greedy stream + perf sanity are the evidence of "a clean revert."

← 59 · Index · 61 →

61 · r58 — q4_K BT spec + cp.async-db2 transplant (REVERTED)

Result: 7B pp3314-eq same-window paired A/B: transplanted 2819.3 vs baseline 3227.6 tok/s = −12.6% → reverted. The transplant itself was functionally clean (greedy-32 byte-identical, parity ×3 all green) — it lost on the performance model, not on correctness. The campaign's first case of "mechanism transplant succeeded, value formula failed". Commit: no repo change (code reverted without leaving a commit; the recorded commit 093ae41 is docs only). Date: 2026-09-06 (Session F).

1. Background — where things stood

By the time r57 was reverted, the q6_K line had already "closed its door": the r53 W_exp + cp.async B staging bundle had pushed 7B prefill to 3176.9 tok/s, the r56 A-side bundle (A cp.async + W_dsc plane) pushed it further to 3212.5, the r54 opt-out gate confirmed the plane was worth its memory, and the r55 roofline audit showed swiglu already at 89% of the bandwidth roof with nothing to harvest from a prefill CUDA-Graph either — the campaign was formally declared CONVERGED at r55, after which r56 still dug +2.35% out of the "already closed" pile. r57 (FA KV double buffer) broke byte identity and was reverted.

At this point vs-llama stood at 1.035× (3212.5 vs the 3324.42 clean-machine anchor). The last region never structurally examined was the q4_K bt kernel — the raw-nibble NB kernel shaped back in the r28 era, then widened in tile. r47's converged-regime breakdown had given it a respectable number: the q4_K GEMM sat only 1.06× behind llama, which at first glance read as "already at parity". But 1.06× spread over the whole prefill is still several hundred milliseconds, and r56 had just proven the q6_K staging mechanism (cp.async + pre-expanded plane) transplants as-is onto another family.

Session F's plan therefore had two steps: Phase 1a was pure measurement — take the q4_K bt kernel apart and attribute where the 1.06× actually lived; Phase 2 would rank the levers by that attribution. The conclusion of the first step spawned this transplant experiment: since the r39+r53+r56 staging pipeline had scored three consecutive hits of +13.3/+5.03/+2.35% on q6_K, porting it to q4_K bt looked like a "free" checklist item.

That "looked like" is exactly where this document's lesson lives.

2. Principle — the GPU mechanism: the pipeline value formula

2.1 The Phase-1a structural attribution

The baseline was re-measured first: 3219.6 median (consistent with the 3212.5 recorded at the r56 landing — the window was healthy). Then a kernel-busy breakdown of the landed configuration:

  • Whole-prefill kernel busy = 984.2 ms, of which q4_K bt = 622.4 ms (63.2%), 166 launches in four classes:
ClassThroughputRelative state
gate/up (ffn_gu)7.51 G-IMMA/sbest
q/o (attn_q, attn_o)7.46 G-IMMA/sbest
ffn_down5.94 G-IMMA/s21% below its own steady state
k/v (attn_k, attn_v)5.05 G-IMMA/s27.8% ceil-wave loss
  • matched-nt ncu against llama: we run 7.5 vs llama's ~6–6.5 G-IMMA/s (on the major classes) — the mma loop itself is not behind; it is ahead. The 1.06× gap does not live in the IMMA.

Conclusion: the q4_K residual lives in three places — (1) per-kt staging exposure in the ffn_down class (each k-tile's staging phase tops out with a stretch of unmasked global latency); (2) ceil-wave tail quantization (27.8% in the k/v class, ≈ 26.1 ms ≈ 2.6% of the whole prefill); (3) launch gaps.

2.2 The transplant's value formula

The top delta r58 chose was to move the q6_K three-piece set over: KDR=2 double buffering + cp.async A/sds staging + sb-parity cp.async B window (measured 46,080 B smem, 122 regs, LDGSTS appearing in the SASS). The instinct that it "should earn" came from the q6_K track record. But putting the two transplants side by side in the same formula exposes where the instinct went wrong:

Transplant value ≈ exposed latency hidden − pipeline cost
Exposed latency ≈ f(how expensive the staging it replaces is)  ← the two families differ wildly here
Pipeline cost ≈ higher barrier density + lookahead depth × issue slots + smem/register pressure

The q6_K transplant (won): before r41, B staging was the ql+qh recombination (recomb ALU) + per-byte LDG + the dsc I2F conversion — the staging phase itself was long and expensive. KDR=2 double buffering let that expensive work for kt+1 overlap kt's compute, hiding a whole stretch of genuine ALU/latency cost in compute's shadow; r53/r56 then replaced the remaining pure copies with cp.async, hiding the LDG scoreboard latency. Only an expensive thing-being-replaced leaves something to hide.

The q4_K bt side (lost): look at the staging macros of mmq_raw_nb_bt_kernel in the current tree (after r34's quantize-transpose prepass, the A side is already a pure uint4 bulk copy; the B-side qs is also a pure uint4 copy):

// src/cuda_kernels.cu — RAW_STAGE_NB_BT of mmq_raw_nb_bt_kernel (current tree,
// the shape after r59; at r58 there was only this "pure copy" staging, without
// the DSC branch below)
/* ---- A: bulk LDG->STS of the pre-transposed qa8/sda (no math) ----*/
{
    const size_t qbase = ((size_t)blockIdx.x * nchunk + (size_t)(kt) * KDR) * MMQ_A_QASZ;
    for (int off = threadIdx.x; off < (KDR * MMQ_NBI * 32) / 16; off += blockDim.x)
        ((uint4*)(qa8))[off] = ((const uint4*)(qa8g + qbase))[off];      // 16B bulk copy
    const size_t sbase = ((size_t)blockIdx.x * nchunk + (size_t)(kt) * KDR) * MMQ_A_SDASZ;
    for (int off = threadIdx.x; off < (KDR * MMQ_NBI * 4) / 16; off += blockDim.x)
        ((uint4*)(sda_q))[off] = ((const uint4*)(sdag + sbase))[off];    // 16B bulk copy
}
/* ---- B: bulk raw qs super-block copy (r18-style, no staging ALU) */
{
    const int sb = ((kt) * KDR) >> 3;
    for (int off = threadIdx.x; off < MMQ_NBJ * 8; off += blockDim.x) {
        const int jj = off >> 3, c8 = off & 7;
        const int j = j0 + jj;
        uint4 v = make_uint4(0, 0, 0, 0);
        if (j < od && sb < nsb)
            v = *(const uint4*)(W + (size_t)j * ((size_t)nsb * 144)
                + (size_t)sb * 144 + 16 + (size_t)c8 * 16);
        *(uint4*)(qb_raw + (size_t)jj * 128 + (size_t)c8 * 16) = v;
    }
}

No recomb, no I2F, no per-byte LDG — the A-side layout transform moved into the prepass at r34, and the B side was confirmed a pure copy in the r18 attempt. The staging phase itself was already near its floor. Injecting KDR=2 double buffering at this point buys:

  1. 4× the barrier density: at KDR=8 you synchronize once per 8 chunks; KDR=2 makes it once per 2 chunks — sync cost ×4;
  2. 1-deep lookahead: double buffering can only run one tile ahead, and on GB10 the global→smem round-trip latency is on the order of ~600–900 cyc; the copy time of one 64-token × KDR=2 tile is far shorter than that latency, so the lookahead cannot fill the latency hole at all;
  3. No register/ALU savings at all: q6_K's double buffering incidentally moved the recomb's register round-trip out of the compute phase; q4_K never had that round-trip to move.

The first term of the formula ≈ 0, the second is strictly positive — net value negative. −12.6% is the measured size of that negative value (much larger than the expected "small negative", because the 4× barrier density also disturbed the issue schedule that had been tracking well).

This is the mirror image of the r45 lesson: r45 said "a non-bottleneck kernel getting faster does not make the wall faster" (horizontal: swap in another kernel); r58 says "a mechanism's cost depends on the granularity of what it replaces" (vertical: the same mechanism hosted on different costs). Together they form the complete transplant criterion: quantify the host's staging-cost structure first, then decide whether to transplant.

3. Implementation

3.1 Design choices (why this shape and not another)

The transplant copied the q6_K landing version wholesale (r39's KDR=2 + r53's cp.async B window + r56's cp.async A), adapting only the q4_K layout:

  • q6_K's B-side cp.async source was the W_exp dense plane; q4_K has no W_exp, so the B side remained the 144 B super-block raw qs copy, but switched to cp.async (sb-parity rotating window);
  • the q4_K dsc (scale) side kept the per-(chunk, row) get_scale_min_k4 decode (the W_dsc plane was then still item 1 of the r58 Phase-2 spec, not yet implemented);
  • smem budget: two copies of A (qa8+sda) + two B windows = 46,080 B, 122 regs — 0 spill, resident block count unchanged.

"Build the minimal transplant first, then measure" was itself the right call (all gates green, one clean measurement); the mistake was skipping §2.2's arithmetic.

3.2 Key code

What the transplanted pipeline looks like (q6_K side, current tree): the transplanted source vanished with the revert, but its host mechanism lives on in the q6_K kernel — the staging macros of current-tree mmq_raw_nb_bt_q6k_kernel are exactly the set moved to q4_K (the comments show the three mechanisms layered):

// src/cuda_kernels.cu — mmq_raw_nb_bt_q6k_kernel (the r39+r53+r56 combined shape)
// r39: DOUBLE-BUFFERED staging — two copies of every per-kt plane so kt+1's
// global->smem expansion (the ql+qh recomb) overlaps kt's compute, hiding the
// B-staging latency that left r38 latency-bound.
uint8_t*  qa8    = mmq_q6k_sh;                       // [2][KDR*NBI*32]
uint8_t*  sda_q  = qa8 + 2 * KDR * MMQ_NBI * 32;     // [2][KDR*NBI*4]
uint8_t*  qb_exp = sda_q + 2 * KDR * MMQ_NBI * 4;    // [2][NBJ*KDR*32]
float2*   sds    = ...;                              // [2][KDR*NBJ]
...
/* ---- A: r56 cp.async bulk staging of the pre-transposed qa8/sda --*/
/* (r45's mechanism on top of r53: the sync LDG->STS exposed its global
 * latency at the top of every staging phase; cp.async hands it to the
 * async unit and the group wait below hides it under the previous
 * tile's compute. ...) */
for (int off = threadIdx.x; off < (KDR * MMQ_NBI * 32) / 16; off += blockDim.x)
    gemm_cp16((__half*)(void*)(qa8b + (size_t)off * 16),
              (const __half*)(const void*)(qa8g + qbase + (size_t)off * 16),
              true);
/* ---- B: KDR*32-chunk super-block window ... r53 bundle: the ql+qh recomb
 * + -32 centering ran ONCE at registration (expand_q6k_dense -> W_exp),
 * so the staging is a pure cp.async bulk copy from W_exp — no recomb ALU,
 * no register round-trip, no ql/qh reads ... */

Note that the premise under which this code works on q6_K is precisely that q6_K's staging was once expensive (recomb + I2F), so "hiding it" had real value; once r53/r56 eliminated the expensive part too, the q6_K side's own cp.async gains had already converged into the bundle's +5.03/+2.35%. Moving the same mechanism onto q4_K, whose staging was already cheap, makes the gain term vanish outright.

3.3 Pitfalls

Three bugs the gates stopped: the transplant itself was well built — all three bugs were caught before landing:

  1. Word/byte pointer confusion (staging side): the sda plane's stepping used a uint32_t* as if it were a byte pointer, jumping 4 B per step — the sda plane was trampled at 4× stride. The parity dump exposed it.
  2. The same confusion (compute side): the consumer side committed it again — the r44-class stride mismatch of "bytes correct, offsets misaligned". This time the greedy-32 byte stream called it first (output diverges from some token onward).
  3. Wrong B-window rotation period: the B window is a whole super-block (144 B, covering 8 chunks), while at KDR=2 a super-block completes only every 4 kt — the window must rotate on super-block parity (every 4 kt), not on kt. Rotating on kt means buffer 1 is never written and the compute side reads the previous round's stale data. Most insidious: the nt=13 dump looked clean — pure stale-node luck (the reused smem happened to still hold the correct data). Only after fixing the rotation period did everything truly go green.

Pitfall 3 is the generic disease of double-buffer changes: the rotation period must follow "how many kt one copy of data covers", not kt itself. When KDR changes, every kt-periodic implicit assumption must be re-audited.

4. Verification

  • greedy-32 byte stream: defends against "wrong math but small deviation" — it was the first to catch both pointer-confusion bugs; after the fixes, IDENTICAL.
  • parity dump ×3 (against the CPU reference): defends against "systematically wrong values that greedy happens not to trip on" — it exposed bug 1; after the fix, ×3 all green.
  • nt=13 small-shape dump: specifically defends against tail/small-shape boundary errors — with the stale-node analysis it flushed out bug 3.
  • Same-window paired A/B: the baseline was re-measured first at 3219.6 (consistent with the landed record), ensuring the −12.6% reading was clean (no r59b-style baseline contamination — re-measured on the spot).

After all functional gates went green the performance was still −12.6%, which is what made the revert legitimate: not "done wrong", but "even done right it should not be done".

5. Results

  • Transplant vs baseline: 2819.3 vs 3227.6 tok/s = −12.6% (same-window paired; baseline window median 3219.6).
  • Functional state: greedy-32 IDENTICAL, parity ×3 green, 0 spill, LDGSTS present in the SASS — the mechanism transplant itself fully succeeded.
  • Veto mechanism: the second term of the value formula (4× barrier density + 1-deep lookahead that cannot cover a 600–900 cyc latency + zero ALU savings) is strictly negative, and the first term ≈ 0 (q4_K staging was already a pure copy). After the revert the baseline returned to 3219.6.
  • Retry conditions: double buffering/cp.async re-enters the candidate list only if q4_K bt's staging becomes expensive again (e.g. new staging-phase ALU is introduced, or the A-side layout transform is forced back into the kernel) — r59's W_dsc plane went the opposite direction (eliminating the remaining staging ALU rather than hiding it).

The Phase-2 spec was produced ranked by attribution (handed to r59): (1) q4_K W_dsc plane (r56's scaffolding moved onto the other 63% of busy); (2) wave re-tiling (+0.3–0.8%, byte-identical); (3) fused ffn_gu concat; (4) riders (host-side small items like prewarm/pre-grow).

6. Lessons

  1. The cost of a mechanism depends on the granularity of the thing it replaces — such a mechanism is not free; it is AMORTIZATION-BOUND. The mirror of r45: before transplanting, first compute how much "exposed latency to hide" remains on the host.
  2. A double buffer's value = whether lookahead depth × tile copy time can cover the memory round-trip latency; a 1-deep lookahead is worth zero against a ~600–900 cyc latency, while the barrier density is a real 4×.
  3. The rotation period of a staging window/buffer follows the data's coverage span (super-block), not kt; changing KDR requires re-auditing every kt-periodic assumption.
  4. "The dump looks clean" is not the same as "no stale data was read" — stale-node luck can let a wrong configuration pass a small-shape dump; attack rotation/reuse bugs with shapes spanning multiple rotations.

← 60 · Index · 62 →

62 · r59 — q4_K W_dsc plane + riders (LANDED, Δ corrected by r59b)

Result: clean basis +11.1% (3232.0 → 3590.8 tok/s, finalized by r59b); the +26.2% recorded on the spot in a co-tenant window was voided because the baseline binary was contaminated (see doc 63). Mechanism level: q4_K bt kernel busy 762.59 → 526.82 ms = −30.9%, regs 124 → 105, at the cost of +1456 MB device memory (the W_dsc plane). Commit: 36a481f (code, + 15c04ba docs; the feb37de recorded in the original text is the unreachable pre-amend copy). Date: 2026-09-06 (Session F, same day as r58).

1. Background — where things stood

r58's teardown attribution of the q4_K bt kernel left behind an evidence-ranked Phase-2 spec, first place going to the q4_K W_dsc plane — "move the scaffolding r56 validated on q6_K onto the other 63% of busy". r58 also proved the converse: pipelining (double buffering/cp.async) is negative on q4_K because its staging is already a pure copy; but the same attribution noted that one genuine ALU residual remains in the staging phase — the rank-1 scale decode for every (chunk, od-row) pair. Eliminating residual ALU (rather than hiding it) is fully compatible with r58's lesson.

On the measurement window: this session ran on a machine carrying a 46 GB sglang co-tenant (r59b later proved an idle co-tenant produces no tax at all — the on-the-spot attribution was wrong). This doc explains the mechanism first; numbers use the clean basis corrected by r59b, with on-the-spot recorded values labeled as such.

2. Principle — the GPU mechanism: turning branchy decoding into a 16 B stream

2.1 What the eliminated cost looks like

The q4_K bt kernel's per-k-tile staging must compute, for the B side, the two coefficients of the rank-1 rescale (d·sc, −dmin·m). The original path executed, for every (chunk, od-row) pair:

get_scale_min_k4(c & 7, blk + 4, &sc, &m)   // branchy packed 6-bit scale decode
d    = h2f(*(uint16_t*)blk)                 // f16 → f32 ×2
dmin = h2f(*(uint16_t*)(blk + 2))
dv = d * (float)sc;  mv = -(dmin * (float)m)  // two int→f32 conversion multiplies

Scale arithmetic: one staging pass per block tile handles MMQ_NBJ × KDR = 128 × 8 = 1024 (chunk, row) pairs, each walking the branchy get_scale_min_k4 (two different bit-extraction paths, cc < 4 vs cc ≥ 4) — differing chunk distributions across a warp cause divergence; times 2 h2f + 2 multiplies. This is pure staging-phase ALU — exactly the single exception to r58's "staging is already a pure copy" conclusion.

2.2 Planarization: one-time pre-decode at registration

The W_dsc plane moves this decode to load time: for each q4_K tensor, build a chunk-major f32-pair plane:

plane bytes = nchunk × od × 8 B = (id/32) × od × 8 = od·id/4 B per tensor
out[(c·od + j)·8 .. +8] = float2(d·sc[c&7], −dmin·m[c&7])

chunk-major (c outermost) is the key layout choice: the tile the kernel reads at kt is exactly the contiguous rectangle rows j0..j0+128 × chunks c0..c0+7 — stored row-major that is 128 scattered 8 B reads; stored chunk-major it is one regular 16 B gemm_cp16 per row pair (two float2 are exactly 16 B aligned), flowing straight into the cp.async channel introduced by r53.

Memory account: the q4_K weights of 7B q4_k_m are actually 5.8 GB (the r58 spec's ~1.07 GB estimate used the wrong byte mass of 4.29 GB), od·id/4 ≈ 5.8 GB / 4 → measured +1456 MB. This cost belongs, alongside r56's q6_K W_dsc (+363 MB) and r53's W_exp (+1.52 GB), to the "plane for ALU" family; r54's opt-out mode proved such trades are acceptable on a 10-GB-class device — but an exit door must be provided.

2.3 Why it is bit-identical

The mma-side rescale demands the plane's coefficients match the in-kernel computation exactly; three guarantees:

  1. exact f16→f32 (half::f16::to_f32 ≡ device __half2float);
  2. exact u8→f32;
  3. exactly one IEEE f32 multiply, one exact negation, and FMA contraction forbidden on both sides — the moment the host compiler fuses d*s with a later add into an FMA, bit-identity breaks.

The u8 6-bit scale (q4_K's packed sc/m) differs from q6_K's i8 dsc: q6_K's dsc is an f16 pair, while q4_K's scale is 6-bit unsigned integers + f16 d/dmin, so the plane stores the multiplied f32 products rather than the raw scales — which is also why it needs only 8 B per pair.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Kernel templated <KDR, DSC>: DSC=true takes the plane stream, DSC=false keeps the original scalar decode — two instances of the same kernel, degrading per launch when the plane is missing, not an either/or at compile time.
  • Geometry-encoded sibling names {name}__q4dsc{od}x{id}: the W_exp pattern — the same logical weight name does not collide across geometries, and the map keyed by the original weight's device pointer hits in O(1).
  • Registration gate mirrors the dispatch gate: the plane is consumed only by the NB-BT kernel, so the registration condition = the kernel's dispatch condition (RAW_NB + A_TRANSPOSE) + geometry gates (id % 256 == 0, same as the kernel's launch gate; od % 2 == 0 guarantees a row pair's cp.async is either fully valid or fully out-of-bounds zero-padded).
  • Failure must be loud: an alloc/upload failure leaves the map empty → the kernel takes DSC=false with a once-per-process eprintln; liveness labels distinguish DSC=f32-plane / DSC=in-kernel(dsc=off) / DSC=in-kernel(fallback!) (r53's lesson: a fast path that degrades correctly must carry a visible label — parity cannot see the dispatch path).

3.2 Key code

Registration side (expand at load + upload + build the map; current-tree src/cuda.rs):

#![allow(unused)]
fn main() {
// src/cuda.rs — expand_q4k_dsc (r59): one-time pre-decode at registration
pub fn expand_q4k_dsc(raw: &[u8], od: usize, id: usize) -> Vec<u8> {
    const Q4KB: usize = 144;
    let nsb = id / 256;
    let nchunk = id / 32;
    let row_len = nsb * Q4KB;
    let mut out = vec![0u8; nchunk * od * 8];      // od·id/4 B, chunk-major
    for j in 0..od {
        let prow = &raw[j * row_len..(j + 1) * row_len];
        for sb in 0..nsb {
            let blk = &prow[sb * Q4KB..sb * Q4KB + Q4KB];
            let d = half::f16::from_bits(u16::from_le_bytes([blk[0], blk[1]])).to_f32();
            let dmin = half::f16::from_bits(u16::from_le_bytes([blk[2], blk[3]])).to_f32();
            let sc = &blk[4..16]; // 12 packed 6-bit scales+mins
            for cc in 0..8usize {
                // host mirror of the device get_scale_min_k4 (cuda_kernels.cu)
                let (s, m) = if cc < 4 {
                    (sc[cc] & 63, sc[cc + 4] & 63)
                } else {
                    ((sc[cc + 4] & 0xF) | ((sc[cc - 4] >> 6) << 4),
                     (sc[cc + 4] >> 4) | ((sc[cc] >> 6) << 4))
                };
                let idx = ((sb * 8 + cc) * od + j) * 8;
                out[idx..idx + 4].copy_from_slice(&(d * (s as f32)).to_bits().to_le_bytes());
                out[idx + 4..idx + 8]
                    .copy_from_slice(&(-(dmin * (m as f32))).to_bits().to_le_bytes());
            }
        }
    }
    out
}
}

Dispatch side (each matmul fetches the plane pointer from the map; a miss yields null → the DSC=false instance):

#![allow(unused)]
fn main() {
// src/cuda.rs — q4_K RAW dispatch arm (r59; the default-on shape after r60)
if type_id == 5 && Self::mmq_gate_on("MINFER_MMQ_RAW") && (id / 32) % 8 == 0 {
    ...
    if nb && at && kd == 8 {
        let (qa8g, sdag) = self.mmq_quantize_transposed(...);   // r34/r49 prepass
        // r59: the q4_K W_dsc f32-pair plane (null on miss ->
        // the DSC=false in-kernel scalar decode instantiation).
        let w_dsc = self.q4k_dsc.lock().unwrap()
            .get(&(wptr as usize)).map(|cp| cp.0)
            .unwrap_or(std::ptr::null_mut());
        nb_ok = qa8g != 0 && sdag != 0
            && launch_mmq_raw_nb_bt_nt(type_id, wptr as *const u8,
                w_dsc as *const u8, qa8g as *const u8, sdag as *const u8,
                out as *mut f32, nt as i32, od as i32, id as i32,
                nchunk, stream, kd) == 1;
}

The kernel-side DSC=true branch was already quoted in doc 61: one gemm_cp16 per (kdd, mm) streams the 16 B at W_dsc + ((c0d+kdd)·od + j)·8 (two float2 — the adjacent two rows of a row pair's four coefficients) into smem, and gemm_cp_commit() commits them in the same group as A/B — no branch or h2f remains in the staging phase.

Two implementation details worth calling out:

  • The 16 B payload = one pair of od rows. The loop variable mm of nc2 = MMQ_NBJ/2 corresponds to the row pair (j0+2·mm, j0+2·mm+1) — one 16 B cp.async carries, for the same chunk, the float2(d·sc, −dmin·m) of two adjacent rows. This is why the registration gate demands od % 2 == 0: a row pair is either entirely valid or entirely out-of-bounds (cp.async's src-size qualifier zero-fills the whole pair); there is no half-valid third state.
  • The compute side cannot tell the two instances apart. DSC=true and false write the same slots of the same smem array sds (sds[kd·NBJ + r], one float2 per row), and the mma-side rescale code is unchanged word for word — bit-identity therefore holds structurally, not because the two sides happen to compute equal values.

3.3 riders: moving one-time costs out of the measurement window

The same commit carried in three host-side riders left over from r57 (prewarm_prefill()):

  1. Kernel module preload — minfer_prewarm_kernels() sweeps cudaFuncGetAttributes over the MMQ/FA/fused launch set, forcing the fatbin to load outside the measurement window (r58 CUPTI: ~3 ms host stall on each side of the first mode-2 swiglu / first bt matmul);
  2. Pinned readback pre-grow — a 4 MB cudaHostAlloc allocated up front (otherwise it is the "0.78 ms tail malloc" at the first logits readback);
  3. MmqCache scratch pre-grow — buf_q8_prefill/buf_qa8_t/buf_sda_t sized for a nominal 4096-token prefill at the largest registered nchunk, so the first prefill's get_or_grow hits directly (otherwise a surprise ~150 MB cudaMalloc lands mid-window). MMQ gating: with MMQ off, the plane is dead weight.

4. Verification

  • Byte-exactness 0 mismatch: expand_q4k_dsc (the host mirror, including the handwritten get_scale_min_k4 mirror) compared byte-for-byte against the device side (the cuda_q4k_dsc_dense_byte_exact test, a matrix of od×id shapes) — defends against "the plane decoded wrongly but the deviation falls inside scale tolerance".
  • parity ×3 + the MINFER_MMQ_Q4K_DSC=0 contrast: flip the env var on the same binary; both the DSC=true and DSC=false paths must pass parity — defends against the one-sided trap of "plane path wrong, contrast path green".
  • greedy byte-identity, both paths: likewise, each path compared against the baseline stream.
  • liveness 166×/0: all 166 launches hit the plane, zero fallbacks — defends against "registered the fast path but silently degraded at dispatch" (the institutionalized r53 lesson).
  • suite 169/0/3: defends against cross-shape regressions.

5. Results

On-the-spot record (co-tenant window, basis later corrected by r59b): interleaved 5× ×2 series, baseline 2836.3/2843.2 → new 3574.7/3588.8 = +26.1/+26.2%; at the time the baseline reading below 3219.6 was attributed to a −12% "co-tenant tax". r59b proved the baseline binary itself was the r58 delta build (a −12.5% defect); the true clean delta is +11.1% (3232.0 → 3590.8), and the co-tenant-tax story is voided — doc 63 has the correction process and the protocol rule.

Mechanism-level evidence (unaffected by baseline contamination; all same-binary before/after):

  • ncu matched-nt: ffn_down-q4_K −34.6% (4.30 → 6.58 G-IMMA/s), gate/up −35%, regs 124 → 105 (scalar-decode register pressure gone);
  • nsys census: q4_K bt busy 762.59 → 526.82 ms = −30.9%;
  • but ffn_down is only −5.2% — r58's premise was half right: the winners were gate/up (−37%) and q/o (−18%), the classes with a large staging-phase decode ALU/I2F share; ffn_down's deficit is L2 reuse-shaped (each launch re-reads the 61.6 MB qa8 plane), not decode-shaped — the plane cannot help;
  • item 3 (od re-tile) skipped with evidence: it needs ≤85 regs, DSC=true is 105, estimated ~+0.2%, below the co-tenant noise floor.

Memory: +1456 MB (measured; r58's 1.07 GB estimate erred on q4_K byte mass). End-of-doc state: default 7B pp3314 ~3581 tok/s = 1.080× vs llama-bench 3323.29.

6. Lessons

  1. The same symptom (a kernel class being slow) can come from opposite root causes — per-kt decode ALU and per-launch A-plane DRAM traffic look identical in an attribution table but demand opposite fixes; classify "decode-shaped or reuse-shaped" first, then pick the lever.
  2. The baseline binary is a measuring instrument: when it is not built from the code you think it is, every delta it produces is fiction (r59b formalized this into a protocol rule).
  3. Memory estimates must use real byte mass: q4_K is actually 5.8 GB, not 4.29 GB — a 36% difference in plane cost; cost accounting for plane-type changes belongs on the registration code, not on the spec table.
  4. Eliminating ALU beats hiding ALU: r58 proved hiding (pipelining) is negative on a pure-copy host; this doc proves eliminating (planarization) is real at the same site — first ask "can this cost be made not to exist", then ask "can it be hidden".

← 61 · Index · 63 →

63 · r59b — clean re-measurement + baseline-contamination correction (measurement round)

Result: finalized numbers 3590.8 vs 3232.0 = +11.1% (replacing r59's on-the-spot +26.2%; the co-tenant-tax attribution is voided); headline figure ~3581 tok/s = 1.080× llama-bench 3323.29 @pp3314 — the campaign uses this pair from here on. No code changes. Commit: 074ca94 (docs). Date: 2026-09-06 (Session F wrap-up, same day).

1. Background — where things stood

After r59 landed the W_dsc plane it measured a +26.1/+26.2% interleaved series in the co-tenant window, and explained "the baseline reads only 2836/2843, below the landed record of 3219.6" as a −12% co-tenant tax. Both numbers were written into the master table at the time (row 73).

But +26.2% is an order of magnitude above every same-family lever the campaign had landed (q6_K's W_dsc plane +2.35%, the W_exp bundle +5.03%), and the direct ncu evidence for r59's mechanism change (staging-phase decode ALU disappearing) is ffn_down −34.6%, gate/up −35% — kernel-level truths, but extrapolating them across the whole prefill cannot support +26%. And the "co-tenant tax" explanation was never independently verified: no anchor, no tax. Session F paused for a dedicated measurement audit — this round wrote not one line of engine code, yet rewrote row 73's Δ column and left the campaign its most important protocol rule.

Two questions: (1) is this window trustworthy (does the co-tenant actually tax anything)? (2) which code was r59's baseline binary actually built from?

2. Principle — the GPU mechanism: the instrument theory of A/B measurement

2.1 The delta is a quotient of two instruments

The same-window paired A/B reading is new_base / base_base. The campaign's conventions guarantee "same window" (removing machine drift), but one implicit premise was never checked: that each binary really is the code it claims to be. The baseline binary is a measuring instrument — when it is not built from the code you think it is, the numerator and denominator still exist, but the value is fiction. And contamination is multiplicative: a baseline deficient by −12.5% inflates every delta by ~1.125×, and the bigger the change, the bigger the absolute inflation.

2.2 Behavioral anchoring: the instrument must be checked against a known answer

The only way to verify an instrument is to have it measure a known answer. Two anchors:

  • Recorded anchor: some historical binary/config has published medians (e.g. the r58-era clean record 3219.6, llama-bench's 3324.42);
  • Rebuild anchor: git worktree a clean rebuild of the baseline commit — bypassing /tmp-snapshot provenance entirely, generating the instrument straight from source.

A mismatch on either condemns the baseline — this is not new measurement, it is putting test weights on the instrument.

2.3 Why an idle co-tenant should theoretically collect no tax

The "co-tenant tax" intuition comes from resource contention (SM time, memory bandwidth, L2 capacity). A 0%-util process that merely sits resident in memory participates in none of these: its pages lie in DRAM, it occupies no SMs, no bandwidth, and its L2 lines get evicted normally. The only theoretical tax is a negligible physical term. So "an idle resident collecting a 12% tax" fails mechanically — the correct suspect is the instrument, not the neighbor. r59b turned this expectation into a measured conclusion (§4).

2.4 The three failure modes of measurement (the campaign's own history)

r59b's rules were not written in a vacuum — all three failure modes have priors in the campaign:

  1. Window drift: the same machine drifts −9% to +38% between sessions (the r12–r25 era recorded in footnote 2). Countermeasure: same-window interleaved pairing — the two readings of each A/B round must be produced interleaved within one time window; absolute values across sessions are not comparable.
  2. Co-tenant load: a live co-tenant is a real tax — r55's baseline sanity read 3144.4–3151.4 under a live co-tenant vs 3181 on a quiet machine. An idle resident is not (proven in §3.1). Countermeasure: distinguish "a neighbor holding bandwidth" from "a neighbor lying in memory".
  3. Binary drift: this doc's protagonist — you think you are measuring A vs B, but you are actually measuring B vs C. The first two modes are covered by the interleaved-pairing protocol; this one had no defense before r59.

What the three modes share: all contaminate the delta without producing any anomalous signal — readings are self-consistent, variance is normal, the mechanism evidence (ncu/nsys) is all real. The only defense is proactive anchoring.

3. Implementation: three audit rounds

3.1 Round 1: window validation (the co-tenant's innocence proof)

Two things on the machine carrying the 46 GB sglang co-tenant:

  1. Utilization sampling: 15 samples, all 0% util — the co-tenant is an idle resident;
  2. Two independent anchors:
    • llama-bench re-run: 3323.29 ± 3.08, 0.03% from the clean-machine-era 3324.42;
    • the known binary behind the 3219.6 record re-measured: 3217.0 (−0.08%).

Both anchors calibrated → an idle resident produces no measurable tax. The window is clean, the "co-tenant tax" hypothesis loses its footing, and suspicion turns to the instrument.

3.2 Round 2: binary-drift test (the smoking gun)

r59's session baseline was a /tmp snapshot (/tmp/minfer_pre_r59), which per the session record should have been "pre-r58 baseline code". Three "same baseline code" binaries compared in the same window:

BinaryReading (tok/s)Delta vs the healthy anchor 3217
/tmp/minfer_pre_r58 (r58 session's baseline snapshot)3217.0— (healthy)
/tmp/minfer_pre_r59 (r59 session's baseline snapshot)2824.7−12.2%
Fresh worktree rebuild of the same commit3232.0+0.5% (healthy)

minfer_pre_r58 and minfer_pre_r59 claim to be the same code yet differ by −12.2% in behavior; the fresh rebuild is healthy. The conclusion is unique: the r59 session mistook the r58-delta binary (the "new" A/B build carrying the −12.6% transplant) for its baseline snapshot. The fingerprint matches: −12.2% ≈ the −12.6% measured in r58's A/B — the snapshot contained that build.

3.3 Round 3: finalized measurement (all rebuilt anchors)

Both baseline and HEAD were cleanly rebuilt from source, interleaved in the same window:

  • fresh HEAD rebuild: 3590.8 median (all 10 runs within 3553.8–3591.5);
  • fresh baseline rebuild: 3232.0;
  • delta = +11.1% (corroborating the +11.5% on the "r58-era clean record" basis — self-consistent when the denominator is the 3219.6 record value);
  • combined median (10 HEAD runs aggregated) 3580.7 → the headline ~3581 tok/s;
  • same-window vs-llama: 3590.8 / 3323.29 = 1.080× (minfer ahead);
  • the memory two-mode census reproduced incidentally: +1454 MB (57189 vs 55735 MiB, minus the resident co-tenant set), matching r59's +1456 MB — the mechanism-change side evidence closes.

3.4 The multiplicative anatomy of the contamination

Decomposing r59's on-the-spot readings against the finalized numbers makes the multiplicative contamination obvious:

Quantityr59 on-the-spot (contaminated)r59b finalized (clean)
New build reading3574.7/3588.83590.8 (fresh HEAD rebuild)
Baseline reading2836.3/2843.23232.0 (fresh worktree rebuild)
delta+26.1/+26.2%+11.1%

The numerator is nearly identical in both windows (3588.8 co-tenant vs 3590.8 clean — an idle neighbor is again harmless); all the inflation comes from the denominator: contaminated 2843.2 vs healthy 3232.0 = −12.0%, interlocking with the binary-drift test's −12.2% and the r58 transplant's A/B delta −12.6%. A −12% denominator defect amplifies a true +11.1% into +26% — a delta is the product of the numerator's truth and the denominator's fiction.

4. Verification (this audit's own gates)

  • Fingerprint match: the contaminated −12.2% nearly coincides with the r58 transplant's A/B delta −12.6% — a verdict requires an independent source explaining the wrong value, not just another mystery.
  • Two anchors cross-confirming: the llama-bench anchor (external program) and the record anchor (the campaign's own historical binary) matched independently at 0.03%/0.08% — the window conclusion transfers only when both instruments are healthy.
  • Rebuild reproduction: the fresh-worktree baseline 3232.0 falls in the healthy band (3217–3232), ruling out "the code itself regressed".
  • Orthogonal-quantity cross-check: the +1454 MB memory delta matches r59's +1456 MB — the correction overturns only the baseline, not the mechanism evidence (ncu/nsys kernel numbers are same-binary before/after differences, unaffected by contamination).

5. Results

  • row 73 corrected: +26.2% → +11.1%; the "co-tenant tax −12%" attribution is voided (footnote 1 permanently marked). The master table's Perf column now only admits same-window anchored values. The r59 section's original text (including the reading series) is preserved with a CORRECTION note added in place — history is not erased; the correction is overlaid as a bound annotation so later readers see the full shape of the error.
  • campaign headline finalized: 7B pp3314 ~3581 tok/s (combined median 3580.7), 1.080× vs llama-bench 3323.29 @pp3314 — the "verified 1.080× path" crowned at r60 refers to this pair of numbers.
  • Protocol output (row 74's one-line lesson): every A/B baseline must first be behaviorally anchored in the same window — re-measure a binary with a known record, or git worktree-rebuild the baseline commit — before the delta is trustworthy.
  • Corollary: idle co-tenancy is equivalent to clean; never infer a "co-tenant tax" without an anchor. Most of the earlier r12–r25-era "box drift −9% to +38%" confusion would have been avoided by this one rule.

6. Lessons

  1. The baseline binary is a measuring instrument, not scenery — calibrate before every A/B (re-measure a known-record binary, or rebuild in a worktree); three minutes once, saves an entire round of wrong attribution.
  2. Multiplicative contamination: with the baseline deficient by −12.5%, a true +11.1% reads as +26% — the bigger a delta is hyped, the more suspect the denominator first.
  3. An idle neighbor collects no tax: 0% util × 15 samples + two-anchor verification is permanent; when a window misbehaves, check the instrument before the neighbor.
  4. /tmp snapshots have no provenance: a snapshot's filename carries no code identity; any cross-session binary must be behaviorally anchored or rebuilt outright before use.

← 62 · Index · 64 →

64 · r60 — the coronation: flipping the verified gate set to default-on (PROMOTION, LANDED)

Result: the default path = the verified 1.080× path — 7B pp3314 default ≈3578–3599 tok/s (headline ~3581); MINFER_MMQ=0 falls back to legacy f16 (measured ~2226 in this window; the clean-class documented value is ~2353). Plane memory +3.27 GB (default 9484 MiB vs planes-off 6217 MiB). decode (nt==1) untouched. Along the way, a bisect caught and fixed a pre-existing mode-2 flaw on mixed-quant models (multiturn_reuse gate 168/1/3). Commit: 57edcf6 (+ 7029ee4 docs). Date: 2026-09-06 (Session F finale).

1. Background — where things stood

By r59b, every gate of P6's r34–r59 gate set had been individually verified in opt-in state: r34's quantize-transpose prepass, r28/r29's NB kernel family, r38–r41's q6_K line, r48/r49's FA and prepass dedup, r51/r52's fused producers (mode 1/2), r53/r56's plane bundle, r59's q4_K W_dsc. r59b delivered the final numbers: clean 3590.8 tok/s, 1.080× vs llama-bench 3323.29 @pp3314. Exactly one move of the campaign remained — flip the default path.

Why the default was still f16 is written in this campaign's history: at R1 (2026-08-31), the then-untuned MMQ kernels managed only ~2.5–3 GMAC/s/matmul under GPU contention while the f16 w16-cache path ran ~8–11, 7B @2K ~155 vs ~630–880 tok/s — at that time defaulting to f16 was the only right call. Twenty-six rounds later the ranking inverted: the MMQ stack 3580.7 vs f16 ~2353 (clean-class). Flipping the default is therefore not "flipping a flag" but a full measurement task: six gates move from opt-in to opt-out at once, every former A/B escape hatch must survive under the new semantics, and every existing mechanism that interacts with the default path (CUDA Graph capture, the decode path, mixed-quant models, the memory budget) must be re-exercised.

This is the campaign's only step that "changes no kernel, only decisions" yet demands the most verification.

2. Principle — the GPU mechanism: promotion is a measurement problem

2.1 The r54-pattern gate flip

The six promoted gates uniformly flip to opt-out semantics (unset or any non-"0" value = ON, i.e. the verified-best path; explicit "0" = pre-r60 behavior):

GateWhat it controlsWhat "0" reverts to
MINFER_MMQprefill uses int8 MMQ GEMMlegacy f16 w16-cache path
MINFER_MMQ_RAWraw-byte staging variant (q4_K whole super-block)per-type legacy kernels
MINFER_MMQ_RAW_NBNB (raw-nibble) kernel familyqb8 pre-expanded form
MINFER_MMQ_A_TRANSPOSEr34 quantize-transpose prepassin-kernel A layout transform
MINFER_MMQ_Q6K_NBq6_K NB pathq6_K f16 fallback
MINFER_MMQ_A_FUSEunset = mode 2 (skip-write fused producer); "1"/"2" keep the r51/r52 semantics; "0"/unknown = offindependent quantize prepass

MINFER_MMQ_Q6K_EXP / MINFER_MMQ_Q4K_DSC, already default-on since r54/r59, keep their semantics. All reads are single-sourced in CudaState::mmq_gate_on(name) — unset maps to true (fail-open to the verified path is deliberate: the release config = the verified config; no stray value should pull the user off it).

2.2 The four classes of "things that must not move"

The risk of flipping the default is not the new path (verified by 26 rounds of A/B) but whether the default-path change touches mechanisms that never interacted with these gates:

  1. decode neutrality: decode (nt==1) never ran MMQ; after promotion no MINFER_MMQ* read may appear on the nt==1 path — otherwise decode behavior drifts with env vars.
  2. reuse-identity neutrality: graph topology is decided by GraphParams/CParams, and reuse compares params — these structures must stay env-free. The gates are allowed to act in exactly two places: dispatch-time kernel choice (nt>=16 matmul dispatch) and registration-time plane building (the loader); CUDA Graph capture bakes each process-constant choice into replay, which is inherently safe.
  3. per-gate liveness: under the default environment every gate must actually be alive (r53's lesson: a fallback-correct fast path must carry a visible label — parity cannot see the dispatch path).
  4. mixed-quant degradation + two-way memory: not every model is all-q4_K/q6_K — mode-2's skip-write producer is unsound under a weight mix that NB-BT cannot consume (the protagonist of §3.3); memory must be reported in both directions.

2.3 The gate × proof matrix

The six gates share one proof set but each has its own failure modes — verification is booked pairwise as "gate × evidence":

ProofMINFER_MMQRAWRAW_NBA_TRANSPOSEQ6K_NBA_FUSE(mode 2)
parity ×3 default env✓ (entry arm)✓✓✓ (prepass layout)✓✓ (skip-write semantics)
greedy-32 vs snapshot+gated✓✓✓✓✓✓
dispatch-label liveness✓✓✓✓✓✓ (166× hits)
opt-out "0" aliveness✓ (f16 path)✓✓✓✓✓ ("0"/unknown=off)
decode/Reuse neutrality✓ (shared by the family)————✓ (rows/n ≥ 16 guard)
Memory both ways✓ (+3.27 GB / 20.5 GB)—————

The "snapshot+gated" contrast means: the greedy stream produced by the r59b-finalized binary plus the full opt-in gate env-var set (the configuration those 26 A/B rounds verified), compared byte-for-byte against the new default's stream with zero env vars — the two must be the same numeric path for promotion to be merely a relocation of the switch.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Single-sourced semantics: all six gate reads go through mmq_gate_on; gate semantics henceforth change in one place.
  • mode-2 becomes default with a degradation ladder: in mmq_a_fuse_mode() unset maps to 2; any window reader (GRAPH_DUMP/DUMP_DIR/trace/viz) or fallback condition (NO_PREFILL_GEMM, non-NB-BT weight mix) present → degrade to mode 1 (fused but writes f32) — keeping r51's gain, giving up only the skip.
  • The weight-mix flag lives at registration: nb_bt_only is an AtomicBool initialized true, cleared by the loader when it registers a non-qualifying weight. Registration precedes the first forward, so zero per-node cost.
  • The compute-capability gate stays: mmq_active() still requires cc >= 800 (mma.m16n8k32 s8 needs sm_80+; sm_75 only has k16) — promotion changes the environment semantics, not the hardware applicability. A side effect points the right way too: the loader uses mmq_active() to decide whether to skip the f16 cache warm pass (the MMQ path reads raw bytes; the w16 copy is dead weight) — with the default on, non-MMQ machines automatically fall back to the f16 path and rebuild their own cache, and both paths' memory accounts hold.

3.2 Key code

The gate semantics proper (src/cuda.rs):

#![allow(unused)]
fn main() {
// r60: the promoted MMQ gate semantics — unset / any non-"0" value =
// ON (the verified 1.080x path), explicit "0" = opt-out to the pre-r60
// disabled/f16 behavior (the r54 `MINFER_MMQ_Q6K_EXP` pattern).
// Single-sourced: every promoted MINFER_MMQ_* dispatch and
// plane-registration read goes through this.
pub fn mmq_gate_on(name: &str) -> bool {
    std::env::var(name).map_or(true, |v| v != "0")
}

// r60: the loaders call this when they register a quantized weight that
// is NOT NB-BT-consumable (not q4_K/q6_K), or a 2-D F32 matmul weight:
// mode-2 skip-write fused producers become unsound for such mixes ...
pub fn clear_mmq_nb_bt_only(&self) {
    self.nb_bt_only.store(false, std::sync::atomic::Ordering::Relaxed);
}
}

Mode selection (default 2 + triple degradation ladder):

#![allow(unused)]
fn main() {
pub fn mmq_a_fuse_mode(&self) -> u8 {
    if !(self.mmq_active()
        && Self::mmq_gate_on("MINFER_MMQ_RAW")
        && Self::mmq_gate_on("MINFER_MMQ_RAW_NB")
        && Self::mmq_gate_on("MINFER_MMQ_A_TRANSPOSE")
        && Self::mmq_gate_on("MINFER_MMQ_Q6K_NB"))
    { return 0; }
    // r60 promotion: unset = mode 2 (the verified-best skip-write fused
    // producers); "1"/"2" keep the r51/r52 override semantics; "0" (and
    // any other unrecognized value, as before r60) = off.
    let requested = match std::env::var("MINFER_MMQ_A_FUSE").as_deref() {
        Err(std::env::VarError::NotPresent) => 2,
        Ok("1") => 1,
        Ok("2") => 2,
        _ => 0,
    };
    // r60: mode 2 additionally requires the NB-BT-only weight mix —
    // a mixed-quant model degrades to mode 1 regardless of how mode 2
    // was requested (default or explicit "2").
    let mode2_possible = requested == 2
        && !Self::no_prefill_gemm()
        && self.nb_bt_only.load(std::sync::atomic::Ordering::Relaxed);
    match requested {
        1 => 1,
        2 if mode2_possible
            && std::env::var_os("MINFER_GRAPH_DUMP").is_none()
            && std::env::var_os("MINFER_DUMP_DIR").is_none()
            && !crate::trace::enabled()
            && !crate::live::enabled() =>
        { 2 }
        2 => 1,   // window reader/fallback condition active: keep the r51 fused semantics
        ...
    }
}
}

The guards at the MMQ dispatch entry (after promotion they guard exactly the default path):

#![allow(unused)]
fn main() {
// src/cuda.rs — matmul dispatch header (excerpt)
// MINFER_MMQ-gated (r60 PROMOTION: default ON — the promoted 1.080x path;
// `MINFER_MMQ=0` = the f16 wmma path). ... id % 32 == 0 covers the block
// math of every type (q6_K runs as k32 chunks with dual 16-sub rescale).
if nt >= 16
    && id % 32 == 0
    && !Self::no_prefill_gemm()
    && matches!(ttype, TensorType::Q4_0 | TensorType::Q4_1 | ... )
}

The loader's registration arms (the two clearing points for mixed-quant + 2-D F32, src/models/qwen2/loader.rs):

#![allow(unused)]
fn main() {
cuda.register_weight(&ti.name, tensor.data());
// r60: a non-NB-BT-consumable quantized weight (not q4_K/q6_K) makes
// mode-2 skip-write fused producers unsound — see CudaState::clear_mmq_nb_bt_only.
if !matches!(ttype, TensorType::Q4_K | TensorType::Q6_K) {
    cuda.clear_mmq_nb_bt_only();
}
... // r59 W_dsc registration gate (RAW_NB/A_TRANSPOSE default-on after r60)
} else if ttype == TensorType::F32 {
    cuda.register_weight(&ti.name, tensor.data());
    // r60: a 2-D F32 weight is an f32 MATMUL weight (norms/biases are
    // 1-D) — its GEMM reads the f32 A directly, so a mode-2 skip-write
    // producer upstream would feed it a dead buffer.
    if tensor.shape.len() == 2 {
        cuda.clear_mmq_nb_bt_only();
    }
}
}

3.3 Pitfalls: the pre-existing mode-2 flaw the bisect caught

After the flip, suite gate #7 (multiturn_reuse, 0.5b q4_0 fixture) ran 168/1/3 for the first time — one failure. The troubleshooting itself is this doc's methodology sample:

  1. A triple stash bisect: clean HEAD + default env (= simulating how promotion runs) PASS; clean HEAD + gated env (MINFER_MMQ=1 ..., i.e. the verified opt-in configuration) also FAIL; the r60 build also FAIL.
  2. Conclusion: the failure is a pre-existing flaw of the verified configuration itself, exposed by promotion (0.5b's q4_0 weights are not an NB-BT-consumable quant type — the mode-2 fused producer skipped the f32 write, while a generic mmq_nt consumer later in the graph legitimately demanded a re-quantization of that A → read a dead buffer). Promotion introduced no new error; it merely let a long-buried landmine step into the default path.
  3. Fix, not revert (r60's principle: fix the wart the default now exposes, loudly): mode-2 producers are conditioned on the registration- time nb_bt_only flag; mixed-quant models degrade mode 2 → mode 1 (writes both the plane and f32, correct everywhere).
  4. A subtler exposure fixed along the way: 2-D F32 weights. Their GEMM reads the f32 A directly and bypasses MmqCache, so r52's dead-write backstop (a loud refusal on cache miss) cannot catch it at all — after the producer skip-writes, the f32 GEMM would silently read a dead buffer. The refusal path never had this exposure; this is also why the clearing point sits at registration rather than consumption.

4. Verification

  • decode neutrality: grep proves zero MINFER_MMQ* reads on the nt==1 path (plus a rows/n >= 16 guard before every mmq_a_fuse_mode call); decode -n 16 --greedy byte-identical; tg128 45.2 = 45.2 flat — defends against "prefill gates leaking into decode".
  • reuse-identity neutrality: GraphParams/CParams env-free, supports_op/supports_fused env-free (the graph side only mentions those env names in comments, zero reads) — defends against "gates affect graph topology, breaking reuse / drifting CUDA Graph replay".
  • parity ×3 default env 9/9 — defends against "some cross term of the default combination never individually verified".
  • greedy-32: default vs snapshot+gated env byte-identical (453 B stream) — proves the default path and the verified opt-in path are the same numeric path; promotion only moved the switch.
  • opt-out aliveness: MINFER_MMQ=0 → zero mmq occurrences in dispatch labels; f16 spot ~2226 tok/s (this window; clean-class documented ~2353) — defends against "a rusted-shut escape hatch".
  • 0.5b q4_k_m smoke: post-fix byte-identical — after the fix it runs the mode-1 producer (the plane is still usable: the plane is a pure function of A), the mixed-quant degradation path really works.
  • suite 169/0/3: one transient SIGSEGV did not reproduce (overcommitted-pool-hazard class, recorded).
  • memory both ways: default-on 9484 MiB; planes-off 6217 MiB (+3.27 GB plane cost); MINFER_MMQ=0 ~20.5 GB — the escape hatch is ~11 GB heavier than the default (the f16 cache is dead weight).

5. Results

  • Prefill A/B (baseline contrast before the fix): base 3592.1 vs new default 3578.0 (−0.39%, overlapping intervals) — no gate leaked; post-fix re-measure +0.44% (the other direction, noise; mode 2 confirmed active on 7B).
  • Finalized default: 7B pp3314 ≈3578–3599 (headline ~3581) = 1.080× vs llama-bench 3323.29 — r59b's verified values become the release default unchanged.
  • Documentation rule (in force from here): default = the verified 1.080× path; MINFER_MMQ=0 = legacy f16 path. From R1's ~441 tok/s first parity-clean MMQ measurement to the finalized ~3581, the campaign arc is 8.1×.
  • A glance at the campaign arc (same-anchor series, master-table Perf column): R1 opt-in 441 → r34 prepass 1496.8 → r39 KDR=2 1777.5 → r40 third resident block 2015.6 → r41 uint4 B-expand 2605.2 → r48 register softmax 2749.9 → r52 skip-write 3011.3 → r53 W_exp bundle 3176.9 → r56 pre-q4_K 3212.5 → r59 W_dsc → r59b finalized 3590.8 / headline ~3581 = the r60 default. Each hop's mechanism is in its numbered step doc.
  • Status: LANDED. Six gates default-on, every gate with a "0" exit, decode and reuse identity proven unmoved, mixed-quant auto-degrades, memory reported both ways.

6. Lessons

  1. Promotion is a measurement problem, not a flag flip: decode neutrality, reuse neutrality, per-gate liveness, mixed-quant degradation, two-way memory — only when all five are measured does the default deserve to flip.
  2. After flipping a default, bisect config vs code on any new failure (clean HEAD default / clean HEAD gated / new build, three-way): all three failing = a pre-existing flaw of the verified configuration exposed — fix it loudly, do not revert the promotion.
  3. skip-write-type optimizations are sound only relative to a consumer set: when betting that "this output will never be read", condition the bet on a registration-time-decidable predicate (nb_bt_only), not on hopes; mixed configs need automatic degradation, not crashes or silent corruption.
  4. Escape hatches must report memory too: here the opt-out path (f16 cache ~20.5 GB) is 11 GB heavier than the default (~9.5 GB) — the "conservative old path" may be the aggressive one on memory.

← 63 · Index · 65 →

65 · D1 decode attribution: split-attention staging depth is the only wall that grows with KV (measurement round, CLOSED)

Result: of 7B q4_k_m decode @1641 KV's 0.91 ms/step wall-clock increment vs tg128, 100% comes from the single kernel gqa_attn_split_partial (1.98 → 34.1 µs/launch); ncu proves it is memory-LATENCY-bound (76.5% long_scoreboard), not byte bandwidth and not insufficient parallelism — 32 splits/head already supply ample parallelism. The ATTN_SPLITS sweep measured a dead end (and is not bitwise-safe); staging-depth-class changes were probe-verified bit-identical — this "free-knob list" directly authorized D2. Commit: no repo change (measurement only; artifacts in /tmp/d1/, ephemeral, key numbers inlined here and in docs/CUDA_OPTIMIZATION.md §2D). Date: 2026-09-07.

1. Background — where things stood

On 2026-09-07, r60 had just crowned the whole MMQ prefill line: 7B pp3314 ~3581 tok/s = 1.080× vs llama.cpp; the prefill campaign had converged. One phase of the CUDA campaign remained unconquered: decode. The decode state at the time (master-table D1 row, 7B q4_k_m / DGX Spark GB10):

  • tg128 (KV~1, i.e. the average KV depth across a 128-token generation): 49.3 tok/s vs llama 49.41 — already at parity;
  • @1641 KV (generation continuing after a 1641-token prompt): 47.2 tok/s vs the llama 49.41-class — −4.5% (0.956×).

That combination itself carries attribution information: tg128 flat while @1641 slow means the gap carrier is a computation that grows with KV depth — in decode's kernel list, the only such things are the KV-cache readers (attention) and KV-dependent elementwise ops. But "it is attention" is not "where attention is slow": it could be byte bandwidth, parallelism, or dependency-chain latency — three diagnoses mapping to three entirely different levers (more bandwidth / more parallelism / restructure staging). D1 set out to separate the three without changing one line of repo code.

Third, gate economics: decode is where bitwise gates are most sensitive (attention's float summation order reacts to any structural change); writing code before knowing "which changes are free" most likely lands in tolerance-gate territory, dragging the session into a parity swamp. Classify the freedoms first, then pick a free one — D1's methodological bet, later cashed by D2.

2. Principle — the GPU mechanism

2.1 The kernel under attribution: the shape of decode split-attention

Op::Attn's nt==1 path dispatches to gqa_attn_split_partial<KV> (current tree src/cuda_kernels.cu, the 1-warp body in attn_split_1w_body):

#define ATTN_SPLITS 32

// each block = 1 warp = (1 split, 1 q head);
// lane l owns 4 consecutive dims (hd=128 → 32 lanes × 4 = 128)
int chunk = (nkv + SPLITS - 1) / SPLITS;   // @1641 → 52 rows/split
int lo = sp * chunk;
int hi = min(nkv, lo + chunk);

float4 q4 = /* this lane's 4 q dims */;
float mx = -INFINITY, S = 0.0f;
float4 oc = make_float4(0.0f, 0.0f, 0.0f, 0.0f);   // R4-rewrite product: oc lives in registers

for (int base = lo; base < hi; base += 4) {        // 4 rows per batch
    int nr = min(4, hi - base);                    // warp-uniform
    /* ...per row: 4-dim dot → 5-step shfl butterfly reduction → expf ×2
       → online-softmax update of (mx, S, oc) ... */
}

The grid is (ATTN_SPLITS=32, n_head=28) = 896 single-warp blocks (7B: 28 q heads, 4 KV heads, GQA 7:1); each block serially online-softmax-scans its split's 52 rows (@1641), 4 rows per batch. This shape is the product of the R4 dimension-parallel rewrite (commit 70f57db) — it replaced the earlier llama-style LOCAL float4 oc[32] accumulator (~80 MB local-memory traffic per layer) with one register float4 per lane, lifting @2K decode from 39.2 to 43.2–45.1 tok/s at the time.

2.2 Three candidate bottlenecks, three counter criteria

Decode @1641 attention reads K+V once each per layer:

per-layer KV bytes = 2 (K+V) × 1641 rows × 512 kv-dim (4 KV heads × hd 128) × 2 B (f16)
                   ≈ 3.36 MB; × 28 layers = 94 MB/step

Each candidate bottleneck has its own counter criteria:

CandidateCriteriaIf it holds, the lever is
Byte bandwidth (byte roofline)DRAM/L2 byte traffic at the roofline, high SM occupancycompress KV (f16 KV already done), reduce re-reads
Insufficient parallelismwaves << 1, SMs mostly idleincrease splits/block count
Memory latency chain (latency roofline)bytes × time far below the roofline, but stalls concentrated in long_scoreboardrework load scheduling (staging) — no byte change, no parallelism change

The criteria's vehicle is ncu's warp-stall sampling and occupancy counters (the r43-established rule: attribute a stall to the consumer instruction waiting on it). The result of this step was the third of the three, with a precise quantitative decomposition:

in-situ 34.1 µs/launch ≈ 12 µs byte time (@273 GB/s) + ~22 µs exposed latency

That is: two-thirds of the time is neither moving bytes nor computing — it is waiting for loads to return.

2.3 Why a latency chain: a 4-row window buys only 4 rows of load-level parallelism

The mechanism's root is the loop structure (the shape at D1 time): K rows are staged at batch start, but each V row's load issues inside the dependency chain — the old code comment claimed "V's address is known, the compiler will hoist the inline V load above the softmax chain", an assertion falsified by measurement (the falsification process is doc 66, which is exactly where D2's entire gain comes from). In reality each V load queues behind the shfl/expf chain, issue point = consume point, so:

52 rows serial × ~1 exposed memory latency per row ≈ 22 µs of waiting
4 rows per batch → load-level parallelism capped at 4 rows

Reference frame: llama.cpp's decode attention uses the same flash-decoding skeleton, but consumes a 256-row × 128-dim window with an 8-warp block — its latency hiding comes from per-block load depth (dozens of loads in flight per block), not our per-row chained structure. That previewed two levers of different magnitude: the small lever = hoisting loads inside the existing 1-warp structure (D2, bitwise-free); the big lever = changing the block/work mapping (D3a/tolerance-class, doc 68).

2.4 Why not parallelism: the ATTN_SPLITS sweep

Intuitively "32 splits not enough → go to 64/128" is the cheapest parallelism lever. Measured dead end, and instructively so:

  1. partial time flat: split-kernel time unchanged across 32/64/128 — the latency chain is per-row; adding blocks does not shorten it;
  2. combine cost 2–3×: doubling the split count doubles the combine reduction's branch count;
  3. not bitwise-safe: any split-count change reorders float summation (outputs differ with ndiff ≈ 3.6e-3, max|Δ| ~3e-9 — the r50/r57 class, not byte-identical).

Item 3 is D1's most important taxonomic output: it cuts decode attention's change space into two classes —

  • staging-depth class (same split ranges, same row order, same per-row ops; only load scheduling moves): probe-verified bit-identical (ndiff=0) → free knobs;
  • split/block structure class (split count, window shape, warp count): necessarily reorders summation → must go through tolerance gates (D3a's calibration package, doc 68).

D2 was about exhausting the first class's freedoms; the second class did not reopen until D3a, in the form of a "calibrated tolerance package".

3. Implementation (the measurement method)

No code changes in this step; §3 records how the three measurement instruments were built and their respective pitfalls.

3.1 nsys per-kernel census: taking the decode step apart

The attribution's core is an nsys per-kernel census: on 7B, take two KV anchors (tg128 at KV~1 and @1641 post-prompt decode), sample a stretch of steady decode steps at each, align by kernel name, and compare per launch. Design points:

  • same-window anchoring: the two anchors' measurements were interleaved within one machine-state window (r59b's lesson — absolute values across windows are not comparable, only in-window differences count);
  • the alignment unit is the launch, not the op: in a decode step the same kernel fires once per layer; align, build a "per-kernel per-step total time" table, then difference the two anchors;
  • additivity check: the difference table must add up to the wall-clock difference. Measured gqa_attn_split_partial 1.98 → 34.1 µs/launch (KV 1.64 → 1641), × 28 layers ≈ +0.90 ms/step, and the two anchors' wall-clock difference is 0.91 ms/step — a single kernel explains 100% of the increment; every other kernel (all MMVQ, rms, quantize, combine) is flat. The attribution closes, leaving no room for "something else hides elsewhere".

3.2 The NCUE probe methodology: nix binaries cannot be ncu-sampled directly → verbatim standalone probe

D1's second instrument is ncu counters (warp-stall sampling, occupancy, registers), and this GB10 has an unavoidable toolchain reality:

  • a nix-built release binary cannot be directly attached and sampled by ncu. One side is device-counter permission (ERR_NVGPUCTRPERM, recorded since the R1 chapter): ncu must go through the sudo -n env LD_LIBRARY_PATH=... protocol (methodology doc no. 77 §2.5; plain sudo strips env vars and ncu silently profiles the legacy path — the r56 lesson); the other: the engine runs on CUDA-graph capture/replay, and ncu's serialized replay disturbs the wall clock and makes it hard to anchor "which logical node is the Nth launch in the replay".

The solution is a verbatim standalone probe: extract the gqa_attn_split_partial kernel body unchanged into a standalone .cu, compile with nvcc directly (-O3 -arch=sm_121a), reproduce the launch in the probe with minfer's real flags/grid/arguments, and attach ncu to the probe process. Three disciplines:

  1. verbatim: the probe kernel is line-identical to the repo kernel — only then do probe counter conclusions qualify for extrapolation to in-situ (the later D3a/D3-6 probes all followed this);
  2. same flags: grid, ATTN_SPLITS, hd, scale, partial layout all take the engine's real values;
  3. nsys owns wall clock, ncu owns structure: ncu serializes replay, so per-kernel times are distorted; use it only for structural readings (occupancy/sectors/stalls); all time conclusions come from in-situ nsys (doc no. 77 §2.5's division of labor).

The probe offered two switchable memory modes, a choice that later proved decisive:

  • hot-L2: small KV run repeatedly, KV resident in L2 — short latency, benefits compressed;
  • cold-DRAM: 28 layers of weights/KV rotating, forced to start from DRAM — real decode's cache state.

During D1 the probe incidentally measured a staging variant: hot-L2 mode showed only −11%, a signal that looked "worth doing but not dazzling"; re-measured in D2 under cold-DRAM, −42%. Mode choice underestimated the gain 4× — decode's KV reads are cold-DRAM-shaped at real scale.

3.3 Pitfalls

  • ncu's stall-attribution direction: PC-sampling books a stall under the consumer instruction waiting on it, not the producing load (the r20/r43 rule). Reading D1's 76.5% long_scoreboard, the attribution target is the chain's consumer (expf/oc update), but mechanically you must reason back to "the load that was not issued early".
  • Do not over-read wave counts: 0.78 waves is evidence that "the SM array is not full", but the kernel is 896 single-warp blocks; not fitting in one wave does not mean parallelism is the bottleneck — the stall distribution proves the bottleneck is each warp's own chain.
  • The probe-vs-in-situ gap must have a name: the gap between probe (hot-L2) −11% and the later in-situ nsys −43% decomposes into "cache temperature + neighboring kernels' L2 interference" — do not use probe absolute times as wall-clock predictions (this calculation became D2's three-way evidence chain: probe −42% / nsys −43% / wall clock +2.0%, doc 66).

4. Verification (measurement validity)

A measurement round's "gates" are not bitwise/greedy but the credibility of the measurement itself. D1 used four:

  • additivity gate: the per-kernel difference sum (0.90 ms/step) matches the wall-clock difference (0.91 ms/step) — defends against missing or double-counted attribution;
  • same-window gate: all anchors measured interleaved within one machine-state window — defends against reading co-tenant drift as a KV effect (r59b-class);
  • verbatim gate: the probe kernel is line-identical to the repo kernel with the same flags — defends against invalid extrapolation of probe conclusions;
  • bitwise classification gate: the classification of "which changes are free" was itself probe-verified — staging-depth class ndiff=0 (verified on nkv ∈ {1, 29, 52, 512, 1641}, later extended by D2 to 45/45), split-count class ndiff ≈ 3.6e-3 — defends against treating a non-free change as free and starting work.

5. Results

Status: MEASURED (measurement only, no repo changes). The deliverables are three tables:

① Wall-clock attribution (7B q4_k_m, GB10, same-window interleaved):

Anchorminferllama.cppDelta
tg12849.3 tok/s49.41parity
@1641 KV47.2 tok/s49.41-class−4.5% (0.956×)

100% of the 0.91 ms/step increment comes from gqa_attn_split_partial (1.98 → 34.1 µs/launch, the only kernel that grows with KV; combine at 96.0 µs/step etc. are all flat).

② Kernel diagnosis (ncu, verbatim standalone probe, minfer flags): 76.5% long_scoreboard stall; all other pipes ≤ 12%; 40 regs; 0.78 waves. Byte decomposition: 34.1 µs ≈ 12 µs bytes + ~22 µs exposed latency → memory-LATENCY-bound. Consistency check: 12 µs @273 GB/s ≈ 3.3 MB = exactly one layer's 1641 × 512-dim K+V bytes; moving the full 94 MB of KV per step takes only ~0.34 ms, while attention spends ~0.95 ms per step — not a byte bottleneck, and not parallelism (32 splits already give 896 blocks).

③ Free-knob classification (probe-verified):

Change classbitwiseVerdict
staging-depth (load scheduling, same ranges/order/ops)bit-identical (ndiff=0)free → D2 acts
ATTN_SPLITS / split-countndiff ≈ 3.6e-3, max|Δ|~3e-9dead end + needs tolerance → sweep measured flat, closed
block/work remapping (llama-vec-style multi-warp)necessarily reordersleft for D3a (tolerance gate, doc 68)

Residual micro-lever registry (targets for later sessions, numbers archived): f32_bits_to_i32 (positions repeatedly converted, ~0.5%/step, later collected by D3-7 2c); the short-KV combine's empty-split reads (~90 µs/step, later vetoed by the D3b-2 analysis as bitwise-unreachable, see doc 67).

6. Lessons

  1. "The only kernel that grows with input scale" is attribution's first cut: two-anchor per-kernel differencing + the additivity check cuts a multi-kernel mystery into a single kernel in one stroke — make that cut before mechanisms.
  2. The three diagnoses — latency chain, bytes, parallelism — must be separated by counters, not guessed: their levers are mutually exclusive, and a wrong guess costs a whole session (D1 separated them in one stroke with 76.5% long_scoreboard + the byte decomposition).
  3. Under nix/sandbox, ncu's entrance is the verbatim standalone probe: the engine binary (especially on CUDA-graph replay) is not a legitimate sampling target; the probe must be line-identical with the same flags, and ncu yields only structural readings while nsys yields time.
  4. Calibrate the "bitwise-free class" before writing code: the classification that staging-depth is free while split-count is not let D2 land in a day and made D3a's tolerance package necessary — a measurement round's most valuable output is often not numbers but the freedom map.

← 64 · Index · 66 →

66 · D2: explicit K+V register staging (LANDED, +2.0% @1641) and the cp.async negative result

Result: 7B decode @1641 KV 47.2 → 48.2 tok/s (+2.0%), gqa_attn_split_partial kernel 34.1 → 19.4 µs/launch (−43%); probe (cold-DRAM) −42% / nsys −43% / wall clock +2.0% three-way consistent; bitwise-identical (greedy-32/256 byte-identical streams, suite 169/0/3). tg128 flat (vs llama: 0.956× → 0.975×, gap −4.5% → −2.4%). Attached: the 5 variants measured to death in the D1/D2 windows (NR=8 register window, cp.async smem pipelines ×3, pair lookahead) — all bitwise-safe, all slower than the landed form; veto mechanism inside. Commit: 0730c15. Date: 2026-09-07.

1. Background — where things stood

D1 (doc 65) had attributed 100% of 7B decode @1641's 0.91 ms/step wall-clock gap to gqa_attn_split_partial, with the diagnosis: 34.1 µs/launch ≈ 12 µs bytes + ~22 µs of exposed memory latency — each V row's load issues inside the online-softmax dependency chain, and the 4-row batch window provides only 4 rows of load-level parallelism. D1 had also calibrated the freedoms: staging-depth-class changes (same split ranges, same row order, same per-row ops; only load scheduling moves) are probe-verified bit-identical; split-count / block-structure classes necessarily reorder float summation and need tolerance gates.

D2's task was thus defined very narrowly: without touching any arithmetic, move the K and V row loads out of the dependency chain. This was the only unused item on D1's "free knob" list. D1's probe had given a conservative signal (−11% for a staging variant in hot-L2 mode); re-measured under cold-DRAM it was −42% outright — decode's KV reads are cold-DRAM-shaped at real scale, so this lever is 4× the hot-L2 signal.

One obstacle had to be cleared first, and it was dramatic in its own right: an old comment lying in the kernel claimed "V's addresses are known, the compiler will hoist the inline V loads above the softmax chain". That comment is wrong — it is both the entire source of this optimization's gain and a good lesson: trusting a comment is no substitute for dumping the SASS once.

2. Principle — the GPU mechanism

2.1 The old form's latency chain

The pre-D2 execution structure for each 4-row batch:

batch start:  K rows ×4 loads (staged, outside the chain)
loop:         for each row j:
                dot (4 mults) → 5-step shfl butterfly → 2 × expf  ← serial dependency chain
                → V row load   ← inline, issue point = consume point, queued behind the chain
                → oc accumulate update ← consumes V

The key is that the shfl/expf chain is a loop-carried dependency (mx/S updated row by row); the compiler's load reordering does not dare (or did not) cross it to hoist the V load: the V load's issue is deferred to near its consume point, and the instruction before the consume point is expf. So every row pays one load latency (hundreds of ns on GB10's cold-DRAM path), 52 rows × one each ≈ 20+ µs of pure waiting — exactly matching the ~22 µs exposed latency D1 decomposed.

"Why doesn't the compiler hoist it itself" is worth recording: the inline load sits inside an if (live) branch, and with a large loop body and a tight register budget, ptxas's scheduling window is insufficient to lift it to batch start; this is not a compiler bug but an ownership-of-gain problem — only the code's author knows "this batch's V can all issue before the chain", and that knowledge must be written explicitly into the source.

2.2 The new form: loads back-to-back

D2 stages all of each batch's K and V into registers before the first softmax step:

batch start:  K rows ×4 loads + V rows ×4 loads  ← 16 LDGs back-to-back, all outside the chain
loop:         for each row j: dot → shfl → expf → oc update (consuming the in-register v4[j])

The batch-start back-to-back loads went from 4 (K only) to 8 (K+V) — the V loads no longer hang row by row on their consume points but issue early outside the chain alongside K, and the load latency overlaps the serial chain — overlapping not just this batch but the previous batch's residual chain. The SASS is the most direct evidence:

16 LDG.E.CONSTANT clustered at 0x7a0–0x9f0; the chain's first SHFL.BFLY/MUFU.EX2 at 0xde0.

2.3 The arithmetic: why the wall clock gains only +2.0%

The kernel −43% but the wall clock only +2.0% — that ratio is dictated by the decode step's anatomy:

kernel savings: 28 layers × (34.1 − 19.4) µs ≈ 0.41 ms/step
step time @1641: ≈21 ms (47.2 tok/s)
0.41 / 21 ≈ +2.0%  ✓

attention @1641 is ~4.5% of 7B's decode step (28 × 34.1 µs ≈ 0.95 ms / step ~21 ms; the same anatomy measured on 14B in D3-1 shows an even lower short-KV share: 0.75%) — D2 cut 43% of it, landing on the wall as +2.0% (0.41 ms / 21 ms). This is also the conversion rate that recurs throughout the later decode campaign: kernel-level big wins ⇒ small wall-clock wins; only moving multiple kernels together (the D3-5/D3-7/D3-8/D4-4 bundles) stacks up to the +5% class.

Register cost (ptxas -Xptxas -v): 40 → 52–58 regs (f16) / 72 (f32), STACK/LOCAL 0, occupancy unchanged (capped at 24 blocks/SM) — 52–58 regs did not push anything out of the SM; the gain is net.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Registers, not shared memory: each K/V row is only 16 B per lane per row (f16 KV 8 B/lane/matrix); smem staging would add an LDS round-trip and a barrier; registers are the zero-cost home for a load. The negative-results table (§5) confirms the smem routes are all slower.
  • Window = the existing 4-row batch (NR=4): the batch size is already a warp-uniform control-flow boundary; widening to 8 rows (NR=8) measured slower — in-flight load count hits a wall (§5).
  • Scheduling only, no arithmetic: same batch rows, same row order, same per-row op sequence — float summation order bit-for-bit unchanged; that is the entire content of bitwise-identical by construction, and the boundary D1's taxonomy authorized.

3.2 Key code

before (the diff of commit 0730c15, the "-" side, i.e. the D1-era shape):

for (int base = lo; base < hi; base += 4) {
    int nr = min(4, hi - base); // warp-uniform
    // Stage K for the whole batch first; the V addresses are already
    // known, so the compiler hoists those loads above the softmax chain.
    float4 k4[4];                                   // ← the old comment's claim, which is wrong
    #pragma unroll
    for (int j = 0; j < 4; j++) {
        k4[j] = (live && j < nr)
            ? kv_ld4<KV>(k + (size_t)(base + j) * stride_kv + hk * hd + d0)
            : make_float4(0.0f, 0.0f, 0.0f, 0.0f);
    }
    #pragma unroll
    for (int j = 0; j < 4; j++) {
        ...
        S = S * corr + e;
        mx = nmx;
        if (live) {
            // ← inline V load: issue point = consume point, queued behind the shfl/expf chain
            float4 v4 = kv_ld4<KV>(v + (size_t)(base + j) * stride_kv + hk * hd + d0);
            oc.x = oc.x * corr + e * v4.x;
            ...
        }
    }
}

after (the diff "+" side; the current tree carries it in attn_split_1w_body, src/cuda_kernels.cu:2830):

for (int base = lo; base < hi; base += 4) {
    int nr = min(4, hi - base); // warp-uniform
    // D2: stage BOTH K and V for the whole 4-row window before the first
    // softmax step. All 8 row loads then issue back-to-back and their
    // latency overlaps the serial chain; the old form relied on the
    // compiler hoisting the inline V loads, which it does not do across
    // the shfl/softmax dependency chain (D1 probe: −11% hot-L2, D2 probe:
    // −42% cold-DRAM vs inline V; bitwise-identical — same rows, same
    // order, same per-row ops, only the load scheduling changes).
    float4 k4[4], v4[4];
    #pragma unroll
    for (int j = 0; j < 4; j++) {           // before the chain: 8 row loads back-to-back
        k4[j] = (live && j < nr)
            ? kv_ld4<KV>(k + (size_t)(base + j) * stride_kv + hk * hd + d0)
            : make_float4(0.0f, 0.0f, 0.0f, 0.0f);
        v4[j] = (live && j < nr)
            ? kv_ld4<KV>(v + (size_t)(base + j) * stride_kv + hk * hd + d0)
            : make_float4(0.0f, 0.0f, 0.0f, 0.0f);
    }
    #pragma unroll
    for (int j = 0; j < 4; j++) {
        if (j >= nr) break; // warp-uniform: all lanes exit together
        float d = q4.x * k4[j].x + q4.y * k4[j].y
                + q4.z * k4[j].z + q4.w * k4[j].w;
        #pragma unroll
        for (int off = 16; off > 0; off >>= 1)
            d += __shfl_xor_sync(0xFFFFFFFF, d, off);   // the serial chain, unchanged
        float s = d * scale;
        float nmx = fmaxf(mx, s);
        float corr = expf(mx - nmx);
        float e = expf(s - nmx);
        S = S * corr + e;
        mx = nmx;
        if (live) {
            float4 vv = v4[j];                  // ← consume the register; no load issued
            oc.x = oc.x * corr + e * vv.x;
            oc.y = oc.y * corr + e * vv.y;
            oc.z = oc.z * corr + e * vv.z;
            oc.w = oc.w * corr + e * vv.w;
        }
    }
}

Side by side: the loop body (shfl/expf/oc update) is untouched character for character; the only structural change is the v4 array filled at batch start and v4[j] replacing the inline load in the row loop. The kv_ld4<KV> template (f32 reads float4 directly / f16 converts two half2 into a float4, src/cuda_kernels.cu:2784) is also untouched.

Archival note: this code was later lifted verbatim into the device function attn_split_1w_body at D3-4 L1 (doc 69), sharing one source with the 4-warp hybrid kernel — the comment states "math is byte-identical to the pre-refactor kernel". The semantics quoted in this doc are those of the current tree's attn_split_1w_body.

3.3 Pitfalls

  • The comment lied: the old comment asserted the compiler would hoist the inline V load. Measurement falsified it — in the SASS of the inline form the V loads crowd next to their consume points, not at batch start. Lesson: before optimizing, cuobjdump -sass; do not trust comments (doc no. 77's SASS-first rule).
  • NR=8's in-flight ceiling: widening the window from 4 to 8 rows to hide one more latency layer measured 22.7 vs 19.9 µs — slower; per-lane in-flight load count (and register pressure) has a hard cap, and deepening the window is not free (§5).
  • The probe's wait-group bug (r59b-class): the first cp.async probe reported several pipeline variants DIVERGENT and the rest "coincidentally OK". The root cause was in the probe, not the concept: cp.async.wait_group <STAGES-1> only guarantees the oldest group's completion when exactly STAGES groups are in flight; when the loop tail has fewer than STAGES iterations it must be wait_group 0. Any cp.async ring over a runtime trip count needs this branchy wait. Unfixed, the negative-results table would have been archived under the wrong reason — "concept error" instead of "granularity error".

4. Verification

All gates green (what each defends):

  • probe memcmp gate: the gqa_attn_split_partial partial buffers compared byte-for-byte, ndiff=0, 45/45 @ nkv ∈ {1, 29, 52, 512, 1641}, across all candidate variants — defends against a staging change quietly touching arithmetic;
  • greedy stream gate: greedy-32 and greedy-256 token streams vs the pre-change binary byte-identical — defends against end-to-end cumulative drift and sampler interactions;
  • parity suite: the trio (cuda_prefill_mmq_parity, cuda_prefill_capture_bit_parity_pp16_pp300, cuda_fa_prefill_attention_parity) + cuda_attn_split_decode_parity + q4k/q6k decode MMVQ parity — defends against collateral damage outside attention (prefill/capture/MMVQ);
  • suite: 169/0/3 — defends against regressions elsewhere in the repo;
  • three-way consistency: probe −42% / in-situ nsys −43% / wall clock +2.0% — the kernel-level numbers of two independent instruments agree, and the wall clock lands on the prediction via §2.2's anatomy — defends against pseudo-gains that "look good on only one instrument".

5. Results

EvidencebaselineD2Δ
probe, cold-DRAM 28-layer rotation, nkv=164134.6 µs19.9 µs−42%
nsys in-situ per-launch µs @164134.119.4−43%
decode -n 128 @1641 (3× interleaved median)47.2 tok/s48.2 tok/s+2.0%
decode tg128 (KV~1)49.449.4flat
combine kernel96.0 µs/step95.7 µs/stepflat
ptxas40 regs (f16)52–58 regs (f16), 72 (f32), STACK/LOCAL 0occupancy unchanged (24 blocks/SM cap-bound)

vs llama.cpp: tg@1641 47.2/49.41 = 0.956× → 48.2/49.41 = 0.975× (gap −4.5% → −2.4%); tg128 stays at parity. 7B decode thus enters llama's ±2.5% band, and later sessions (the D3 series) kept grinding from this base.

5.1 Negative results (do not retry blindly)

All bitwise-safe (after fixing §3.3's probe wait-group bug), all measured slower than the landed form's 19.9 µs in the same cold-DRAM mode:

Variantcold-DRAM µsvs 19.9
register staging NR=8 (window doubled in depth)22.7slower
cp.async smem K-only pipeline NR=4 S=2 / S=323.4 / 22.4slower
cp.async smem K-only pipeline NR=8 S=230.9far slower
cp.async smem K+V pipeline NR=4 S=221.0slower
pair-unrolled register lookahead (95 regs)22.1slower

Veto mechanism (under what conditions a retry is worthwhile):

  • LDGSTS granularity: f16 KV is only 8 B per lane per row — cp.async's LDGSTS.E.BYPASS.128 vocabulary collapses at this granularity, and the smem route forces one LDS round-trip; the direct LDG-into-register route pays neither tax. cp.async remains the right tool — but only where the staging tile is ≥16 B/lane (the MMQ kernels; r45/r53/r56 all rely on it); decode attention is not in that class unless a KV-layout change (e.g. dpl, doc 76) widens the row granularity.
  • Window depth: NR=8 hits the hard cap of in-flight loads/register budget, not "not enough hiding". Unless the SM microarchitecture generation changes (more in-flight loads per warp), this direction is not worth retrying.

5.2 Residuals and what came after

The landed kernel's latency decomposition: 19.4 µs ≈ 12 µs bytes + ~7 µs for the serial chain itself — the chain became the floor. Going further requires a block/work-mapping change (a llama-vec-style multi-warp 256-row cooperative rewrite), unreachable under byte-identity and requiring tolerance gates (D3a did it and was reverted for rpw pathology, doc 68; D3-4 ultimately recovered it at long KV via hybrid dispatch, doc 69). The two micro-levers f32_bits_to_i32 (~0.5%/step) and the combine's empty-split reads (~90 µs/step) were left for later (D3-7 2c / D3b-2).

6. Lessons

  1. SASS first, comments second: a "the compiler will do it for me" comment let the inline V load pay 22 µs/launch of latency for nothing; the first step of any load-scheduling optimization is always dumping the SASS to see the issue points.
  2. Calibrate the bitwise-free class with a probe first, then maximize it: D1's classification (staging-depth free) + D2's execution (8 loads back-to-back) form the standard rhythm of "measurement authorizes → narrow change → three-way verified"; a −43% kernel required no parity price at all.
  3. Archive negative results with their mechanisms: the 5 variants died of LDGSTS granularity (8 B/lane) and the in-flight hard cap, not "cp.async is bad" — write the mechanism down, and the next layout change reveals which premise moved and whether a retry pays.
  4. Instrument bugs forge concept errors: one wait_group tail-semantics bug nearly wrote "cp.async doesn't work" into the archive; instruments must prove themselves first (r59b's baseline contamination and this doc's probe divergence are the same species).

← 65 · Index · 67 →

67 · D3: 14B decode attribution (D3-1) + the D3b bitwise MMVQ triple (1b LANDED; 1a/1c REVERTED)

Result: D3-1 measurement round — 14B decode's wall-clock gap (43.84 vs llama 41.14 ms/step) is not attention (at short KV attention is only 0.75% of the step): about half is three MMVQ stragglers (attn_v-q6K 134.8 GB/s on padded-f32, ffn_down-q6K 198.9, the output head 200.1, vs the 220–225 GB/s class, totaling ≈ +1.2 ms/step), the other half is the elementwise/launch chain (97 rms + 265 quantize + ~700 sub-2 µs launches); the whole step is 93.4% weight streaming, launch gaps already better than llama. D3b (bar: ≥ +0.4% at the target site, bitwise-gated): 1b q6_k_q8_mmvq_v2_pf pipelining LANDED (14B tg128 +0.44%, @3254 +0.24%; 7B +3.0%/+2.7% — later voided by D4-2, see §5); 1a both routes REVERTED (MMVQ routing is not bitwise; the NSG 2→1 remap is bitwise-green but the kernel +9.6%); 1c 160-thread blocks REVERTED (the gain mechanism does not exist on this SM). Commit: f1825b5 (1b); 1a/1c were reverted after measurement, no repo trace (patches in /tmp/d3/, ephemeral). Date: 2026-09-07.

1. Background — where things stood

D2 (doc 66) ground 7B decode @1641 to 48.2 tok/s (0.975× vs llama) with tg128 at parity — 7B was inside llama's ±2.5% circle. The campaign's next small model was Qwen2.5-14B q4_k_m: 48 layers, hidden 5120, GQA 40:8, decode tg128 ~22.8 tok/s vs llama 24.31 (0.94×), worse at long KV. 14B's weight volume is ~2× 7B's, and decode streams the whole weight set every step, so per-kernel effective GB/s becomes decode's first-class metric for the first time — gaps that in the 7B era were explained by launch counts and attention shape become, at 14B, "which matmuls fail to reach the stream rate they should".

Meanwhile D3a (the 4-warp fattn-vec rewrite, doc 68) was opening another front: structural attention rewrites must go through tolerance gates. D3's session design therefore split into two tracks — D3-1 attribution + D3b doing only bitwise levers (this doc), with attention structure left to D3a/D3-4 (docs 68/69). Bar pre-registered: a lever survives only if measured ≥ +0.4% at its target site (the kernel/anchor it claims to save), otherwise revert.

2. Principle — the GPU mechanism

2.1 The anatomy of the 14B decode step

The quantitative anatomy from D3-1's nsys census (method in §3.1):

  • per step 43.84 ms (llama.cpp 41.14, gap 2.70 ms / 6.6%);
  • 93.4% of kernel time is weight streaming (the MMVQ/MMQ family passing all weights through DRAM) — decode is essentially a "per-token pass over the full weight set"; the launch-gaps item is already better than llama (the engine's CUDA-graph replay wins on per-launch overhead — not a gap source);
  • at short KV attention is only 0.75% of the step — the wall-clock gap has nothing to do with it;
  • the gap's composition: ~half = three MMVQ stragglers running at low stream rates (§2.2), totaling ≈ +1.2 ms/step; ~the other half = the elementwise/launch chain (97 rms + 265 quantize + ~700 sub-2 µs launches — this chain was later collected in three batches by D3-5/D3-7/D3-8, docs 70/72/73).

An important contrast item: attention is the only kernel in the 14B step that grows with KV (at nkv=3254, +3.33 ms/step, 67% of the DRAM byte floor, long_scoreboard-bound) — it matters, but its lever lives at long KV and needs tolerance gates, outside this doc's bitwise track.

2.2 The three stragglers' arithmetic

A decode matmul's effective stream rate = weight bytes / kernel time. The 14B census computed this table per kernel at the tg128 and @3254 anchors; three fell below the "220–225 GB/s class":

KernelShape (od × id)Dispatch pathGB/sShortfall
attn_v-q6K1024 × 5120 (5.24M elem, 11 layers)padded-f32 (the 24M gate excludes it)134.8−40%
ffn_down-q6K5120 × 13824 (id 13824 → npair 432)MMVQ v2198.9−10%
output head lm_head-q6K152064 × 5120 (npair 160)MMVQ v2200.1−10%

attn_v's byte arithmetic: per-layer weights = 1024 rows × (5120/256 = 20 super-blocks) × 224 B (padded stride) = 4.59 MB; 134.8 GB/s ⇒ ~34 µs/launch, matching nsys's 36.4 µs. The three together ≈ +1.2 ms/step — exactly half the gap.

2.3 MMVQ vs padded-f32: why the 24M gate keeps attn_v out

The 8e era measured decode MMVQ's shape crossover on device (archived in the src/cuda.rs:2589 comment): below od*id < ~24M padded-f32 wins (od 512 → MMVQ 4.5× slower, od 896 → 3.0×, 2048×4864 → 1.66×); at large shapes MMVQ wins (7B ffn_down 1.5×, lm_head 1.4×). Mechanism:

  • small od ⇒ only 1–2 units per thread (a unit = half of 32 elements, q6_K's dp4a dot unit): q5/q6's nibble byte reads are inherently scattered (unit granularity 16–64 B); with too few units those scattered loads' latency is fully exposed; the padded-f32 kernel is a block dequant + FMA loop with coalesced reads and is actually faster at small shapes;
  • large od ⇒ many units per thread: dp4a's arithmetic-density advantage
    • uint4 widened reads (the R2 rework) outweigh the scattered-load disadvantage.

14B attn_v is od 1024 × id 5120 = 5.24M < 24M ⇒ excluded from MMVQ by the gate, landing on the padded-f32 kernel — yet it is exactly one of the kernels that most deserves MMVQ at that shape (D3-7 2b later lowered the gate to 4M and rescued it, doc 72; this doc's 1a handles "what could be done in the code of that moment").

2.4 v2's second-serial-unit problem (pf's mechanism)

q6_k_q8_mmvq_v2 is a u-loop form: each thread for (u = tid; u < npair; u += 256) accumulates unit by unit. With npair ≤ 256 each thread has one unit — no problem; at npair > 256 (14B ffn_down: id 13824 → npair 432) threads 0..175 get a second unit — and that second unit's weight + q8 loads issue only after the first unit's accumulation, serially hanging off the critical path. The 198.9 vs 220-class gap is mostly this exposed load latency.

The fix is of D2's lineage: move only load scheduling, not arithmetic — issue both units' loads before any accumulation (bitwise-free, staging-depth class).

3. Implementation

3.1 The D3-1 method: nsys census + per-kernel GB/s

The same rig as D1 (doc 65) with two upgrades: the anchors become 14B's tg128 + @3254; the census table annotates every matmul kernel with effective GB/s (weight bytes / nsys time), turning "who is below class level" into a visible ranking rather than a feeling. The full report lives in /tmp/d3/D3_FINDINGS.md (ephemeral); the key numbers are inlined in the master table's D3b rows and this doc.

3.2 D3b-1b: q6_k_q8_mmvq_v2_pf, LANDED (f1825b5)

before — v2's u-loop (src/cuda_kernels.cu:1536, current tree, untouched by 1b):

float acc = 0.0f;
for (int u = threadIdx.x; u < npair; u += 256) {   // ← when npair>256 the second
    const int kbx = u >> 3, pair = u & 7;          //    pass's loads issue only
    const uint8_t* blk = weights + (size_t)row * row_stride + (size_t)kbx * blk_stride;
    ...                                             //    after the first pass accumulates
    acc += d8 * sc0 * d * (float)dot0 + d8 * sc1 * d * (float)dot1;
}

after — the pf form splits the unit into load/acc halves, both loads back-to-back:

__global__ void __launch_bounds__(256) q6_k_q8_mmvq_v2_pf(
    const uint8_t* __restrict__ weights, const uint8_t* __restrict__ acts8,
    float* __restrict__ output, int od, int id, int nt, int blk_stride
) {
    const int row = blockIdx.x;
    const int t = blockIdx.y;
    const int nbe = id >> 8;
    const int row_stride = nbe * blk_stride;
    const int npair = id >> 5;
    const uint8_t* x8row = acts8 + (size_t)t * (id >> 5) * Q8PB;
    const uint8_t* wrow = weights + (size_t)row * row_stride;

    float acc = 0.0f;
    const int u0 = threadIdx.x;
    const int u1 = u0 + 256;                 // the same thread's second unit
    if (u0 < npair) {
        Q6kUnitRegs r0, r1;
        q6k_unit_load(u0, wrow, x8row, blk_stride, &r0);      // load #1
        const bool two = u1 < npair;       // npair > blockDim (guaranteed by the dispatch gate)
        if (two) q6k_unit_load(u1, wrow, x8row, blk_stride, &r1); // load #2 immediately after
        q6k_unit_acc(&r0, acc);                                // only then accumulate
        if (two) q6k_unit_acc(&r1, acc);
    }
    mmvq_block_reduce(acc, output, od, t);
}

q6k_unit_load/q6k_unit_acc (src/cuda_kernels.cu:1590/1611) mechanically split the v2 loop body: the load half gathers uint4 ql/qh + scales + the q8 row pointers into the register group Q6kUnitRegs; the acc half is the original dp4a tree. The bitwise argument (written twice, in the file comment and the commit): same thread→unit mapping (u = tid, tid+256), same per-unit dp4a tree, per-unit accumulation statements character-identical (same FMA contraction shape: acc += d8 * sc0 * d * dot0 + d8 * sc1 * d * dot1), ascending u order, same block reduction — the only thing moved is load scheduling, i.e. D2's staging-depth precedent reused on MMVQ.

Dispatch gate: id > 8192 (14B ffn_down's npair 432 was the only target shape at the time) goes to pf, everything else to the v2 loop. Note: the gate had no upper bound — the consequence of that omission is §5's D4-2 CORRECTION; the current tree (src/cuda.rs:4293) has id > 8192 && id <= 16384.

3.3 D3b-1a: rescuing attn_v from padded-f32, both routes dead

Goal: rescue attn_v's 134.8 GB/s (1024 × 5120) from the padded-f32 kernel.

(a) MMVQ routing (lower the 24M gate) — not bitwise, out immediately. MMVQ quantizes activations to q8 and then dp4a; the padded kernel consumes f32 directly. Different accumulation semantics ⇒ a zero-byte gate is impossible; any routing change must go through D3a's tolerance session. Within the bitwise track this road is closed (D3-7 2b later took the tolerance track and landed, doc 72).

(b) NSG 2→1 row→warp remap — bitwise-green, measured a loss. The padded kernel (q6_k_q8_mmvq_padded family, NR0=2 rows/warp, NSG=2 warps/block):

const int NR0 = 2;                  // 2 rows per warp
const int NSG = 2;                  // 2 warps per block (block = 4 rows)
int r0 = (blockIdx.x * NSG + warp_id) * NR0;

The remap: kernel+launcher change NSG 2→1 together, rows still 2/warp, warp count doubled (as recorded); each warp still streams the full activation row for its 2 rows (f32 id 5120 = 20 KB), so y re-read traffic grows 1:1 with the warp count. Result: all bitwise gates green, but the nsys same-window kernel went 36.4 → 39.9 µs (+9.6%), wall tg128 −0.74% / @3254 −0.33% → reverted.

Veto mechanism: at this shape the padded kernel is not warp-starved (134.8 GB/s is an L2/latency composition problem, not a parallelism deficit); within byte-identity the padded kernel's only free knob is rows-per-warp, and 2 is already the sweet spot.

3.4 D3b-1c: dynamic output-head block size — the mechanism does not exist

lm_head npair = 160 (id 5120) means 96 threads of a 256-thread block idle. The change: launch 160-thread blocks (5 warps, rounded to warp count), with mmvq_block_reduce bounded by the actual warp count (idle threads contribute only exact +0.0 terms ⇒ bitwise-safe; the reducer src/cuda_kernels.cu:1269 has a fixed 8-slot warp_sums[8], extra slots zero for a narrow block). Bitwise all green, but 14B wall +0.04% / +0.09% — noise level.

Veto mechanism (an arithmetic veto, not measurement uncertainty): GB10's SM thread limit is 1536. 6 × 256-thread blocks already fill the 1536 thread slots (960 live); 9 × 160 = 1440 live — the gain mechanism of "eliminating idle threads" does not exist on this SM: the live-thread increment a narrow block buys (1440 vs 960) is bounded neither by thread slots nor by occupancy (at this shape the blocks are far below the SM's block cap). 7B @1641's +0.26% (SEP) was real but sub-bar (npair 112 → 144 idle, a smaller magnitude).

3.5 Pitfalls

  • dump-gate pool-slot aliasing: informational node dumps like node{3,5,8}_prefill.f32 read recycled pool slots, and node→slot identity drifts with binary layout — in 1a's contrast they appeared as "contents byte-equal but slots swapped". From then on the bitwise gate checks only logits (both phases)/kv/decode-node files; same-size prefill node slot differences are recorded as aliasing, not value differences. (This artifact was cited repeatedly in D3-5/D3-8 as the "documented class".)
  • Dispatch gate with only a lower bound: 1b's id > 8192 had no upper bound, letting 7B ffn_down (npair 592) silently drop units — every bitwise gate compared v2_pf-to-v2_pf binaries and can never see work the dispatch drops (details and cost in §5). This is the campaign's most expensive pitfall, dug out and fixed by D4-2 (doc 74).

4. Verification

1b's (the only landed item) gate chain, and what each gate defends:

  • dump gate: 114/114 MINFER_GRAPH_DUMP files vs the pre-change binary byte-identical (logits prefill+decode, kv0, all 48 layers' KV, node dumps) — defends against a kernel scheduling change moving any value;
  • greedy gate: the -n 256 token stream byte-identical — defends against end-to-end cumulative drift;
  • suite: 169/0/3 (+ the FA trio, split-decode parity, replay-bit-parity by name) — defends against collateral damage to attention/prefill/capture;
  • the by-construction argument: same unit mapping, same dp4a tree, character-identical accumulation statements, ascending u, same reduction — defends against the fluke of "gates green but mechanism wrong";
  • window anchoring: 14B tg128 22.80–23.11 (D3-1 window 22.81), @3254 21.01–21.11 (21.25); 7B tg128 48.03 (guard 49.4, co-tenant window; 49.47 after 1b clears the guard), @1641 46.66 (guard 48.2); pp512 guard unmoved (prefill unchanged, spot-check 2055 tok/s) — defends against co-tenant drift disguised as gain (the r59b rule).

1a(b) and 1c passed the same bitwise gates (all green) and were then vetoed on the wall clock — the bitwise gates prove "not broken", the wall clock proves "not better"; neither suffices alone.

5. Results

LeverStatusKernelWall (interleaved 3× median)Verdict
1b v2_pf (npair>256 dual-unit loads hoisted)LANDED f1825b5—14B tg128 22.80→22.90 (+0.44%), @3254 21.01→21.06 (+0.24%, SEP); 7B tg128 48.03→49.47 (+3.0%, min-new > max-base), @1641 46.66→47.94 (+2.7%, SEP)landed; 7B numbers voided, see CORRECTION below
1a(a) MMVQ routing (lower the 24M gate)REVERTED——not bitwise (q8 vs f32 accumulation semantics), needs a tolerance session → landed separately as D3-7 2b
1a(b) NSG 2→1 remapREVERTED36.4→39.9 µs (+9.6%)tg128 −0.74%, @3254 −0.33%y re-reads grow 1:1 with warps; the padded kernel is not warp-starved
1c 160-thread blocksREVERTED—14B +0.04%/+0.09%; 7B @1641 +0.26% (sub-bar)1536 threads/SM: the gain mechanism does not exist
D3b-2 short-KV combine skipvetoed by analysis (not implemented)——see below

(SEP = strict separation in same-window interleaved pairing: min-new > max-base, paired intervals non-overlapping — a stronger verdict than medians, defined by the r59b window protocol, see doc 77.)

D3b-2 (vetoed by analysis, the door closed before code): the single-split path is bitwise-equal to 32-split partial+combine only at nkv=1; a real decode step's combine merges ~nkv/chunk live partials with exp(mx−gmx) rescaling, reordering the float summation relative to the serial online-softmax chain (D1 measured split-count reordering: ndiff ≈ 3.6e-3, max|Δ| ~3e-9). And the split grid is frozen by CUDA-graph replay capture (the kernel reads positions at runtime; branching at capture time would lock the captured step's shape for the whole session). The bitwise-reachable residual — combine early-exit on empty splits (contributing exact +0.0) — is worth ≤ ~15 µs/step, below every bar; the real ~166 µs/step prize needs tolerance gates.

D4-2 CORRECTION (2026-09-09, must be read together with this doc): 1b's two 7B legs are void. v2_pf's dispatch gate id > 8192 had no upper bound; 7B ffn_down (id 18944 → npair 592) entered and each thread's u1 = tid+256 covers only up to 511 — units 512..591 (80 of 592 = 13.5% of the down-q6K work) were never computed. 7B's +3.0%/+2.7% was mostly "13.5% less work computed", not pipelining gain; and since all bitwise gates compare v2_pf-to-v2_pf binaries, structurally the dropped units are invisible. The fix (b31084c, doc 74) added the upper bound id ≤ 16384 (npair ≤ 512) to the gate, with higher shapes taking the v2 loop form; after re-anchoring 7B returned to its honest baseline. The pipelining gain itself is the 14B-scale +0.2–0.4% class — the two 14B legs (+0.44%/+0.24%) are unaffected and stand.

Window anchors (this session's interleaved pre-binary medians): 14B tg128 22.80–23.11, @3254 21.01–21.11; 7B tg128 48.03 (guard 49.4), @1641 46.66 (guard 48.2); pp512 guard unmoved.

6. Lessons

  1. decode's attribution unit is GB/s, not tok/s: in a step that is 93.4% weight streaming, "which matmuls sit below the 220–225 GB/s class" is an executable question list; the three-straggler list directly generated the 1a/1b/1c levers and the later D3-7 2b landing.
  2. Admit the bitwise track's boundary in advance: routing-class changes (q8 vs f32) are out of bounds by nature; remap-class changes (y re-reads 1:1) are in bounds by nature but the mechanism may not exist — compute the mechanism first (warp count, y bytes, SM thread slots), then write code; 1c could have been vetoed by the 1536-thread arithmetic before any code, but was instead vetoed only after writing and measuring.
  3. Dispatch gates need bounds on both sides; bitwise gates cannot see dropped work: 1b's unbounded-upper gate let 7B decode run for two days missing 13.5% of its down-proj work with every gate green — tests can only prove "two binaries did the same thing", not "dispatch made the kernel do all the work" (D4-2's core lesson and this doc's most expensive line).
  4. Analytical vetoes are also output: D3b-2 used D1's reordering evidence
    • the replay-freeze mechanism to close a seemingly ~166 µs/step prize to ≤15 µs, saving an entire tolerance session — run the bitwise-feasibility math before writing code.

← 66 · Index · 68 →

68 · D3a — the 4-warp fattn-vec-style split-attention rewrite (REVERTED) + tolerance-gate calibration

Result: the rewrite itself was vetoed — 7B @1641 split kernel 21.1 → 34.7 µs (+64%, the rows-per-warp pathology), 14B @3254 73.4 → 68.4 µs (−6.9%) but the wall clock only −0.95% (inside the ±2% noise band) — both bars failed, and the whole thing was reverted per the r44 precedent. The session's durable output is the tolerance-gate calibration: a kernel-level 1e-7 numerical difference drifts end-to-end logits by O(0.4) on 28–48-layer models; the "end-to-end max|Δlogits| ≤ 1e-3" gate proposed at D3-1 is unsatisfiable for any rewrite that changes accumulation order (retired); the usable gate set = kernel-vs-CPU ≤1e-4 (real outlier-magnitude data) + argmax HARD gate (top-2 margin > 0.1) + greedy −n 256 ×5 seeds + temp-0.8 sampling contrast + suite + interleaved A/B. Commit: no repo change (the experiment patch is kept at /tmp/d3/d3a_kernel_patch.diff, full measurement record /tmp/d3/D3A_FINDINGS.md; the docs-only commit is a05af20). Date: 2026-09-07.

1. Background — where things stood

The D-series teardown campaign reached its third stage. D1 attributed 100% of 7B @1641 decode's KV-scaling wall to the single split-attention kernel gqa_attn_split_partial (34.1 µs/launch, 76.5% long_scoreboard) and proved the ATTN_SPLITS sweep a dead end (changing the split count reorders float summation, no longer bitwise). D2 landed the first win with "explicit K+V register staging": all 8 loads of the 4-row window issued early, kernel 34.1 → 19.4 µs, wall clock +2.0% (a bitwise-class change, cheap gates).

Then D3-1 finished dismantling the 14B @3254 decode wall, where the attention term is bigger: split kernel 73.4 µs/layer, still 50% above the 48.9 µs byte floor. D3b's bitwise MMVQ levers (1a/1b/1c/2) cleared the GEMM side's low fruit, but the attention residual was the next big target. While dismantling, D3-1 had also drafted a tolerance gate for future non-bitwise levers: "end-to-end max|Δlogits| ≤ 1e-3". Nobody had validated that draft — this doc is its first combat test, and it lost (see §5.2).

Lever 2's direction was written in the campaign plan long before: change decode split attention from the 1-warp serial structure to llama.cpp fattn-vec's 4-warp structure (source-verified @ca3d5a3e1). The chain of assumptions at the time: the incumbent kernel is "serially dense" — every lane works on every row, 32 threads carrying all 128 dims; llama's 4-warp structure hands rows to 4 warps in 32-row windows, each window doing only one online-softmax rescale, with shorter latency chains and more warps per SM (D3a measured __launch_bounds__(128,8) → 32 warps/SM vs the incumbent's 24). It looked like winning on both ends. D3a's task was to port this structure onto minfer's f16-KV decode path and measure what it is worth.

Where things stall without this step: the 14B @3254 attention residual (73.4 − 48.9 ≈ 24.5 µs/layer × 48 layers ≈ 1.2 ms/step) was the largest known single item, and D2's "issue-point hoisting" trick was already spent; the remaining levers are all structural.

2. Principle — the GPU mechanism: the rows-per-warp pathology

2.1 The geometry of the two generations

The incumbent 1-warp kernel (D2-staged): grid dim3(ATTN_SPLITS=32, n_head), each block 32 threads (1 warp) covering one split's chunk = ceil(nkv/32) rows × hd dims. Each lane permanently owns 4 consecutive dims (d0 = lane_id * 4) and, for every row, does one "4-dim dot + full-warp butterfly reduce + online-softmax update". However many rows, every lane is busy — that is "serially dense": no idle slots, at the price of a per-row reduce dependency chain (D2 relieved the load side with early issue, but the softmax chain itself remains row-by-row).

D3a's 4-warp structure (fattn-vec style): 128 threads (4 warps) per block; K/V stream straight from global (never into smem); Q lives in registers (16 dims per thread = 4 × float4); the warp's four 8-lane subgroups each claim a row (butterfly sums within a subgroup); each warp handles a 32-row window and does exactly one online-softmax rescale per window; probabilities stage into a 128-float smem row, and a final 4-warp LSE-merge writes one partial.

2.2 The pathology: rpw decides slot utilization

The key parameter is rows-per-warp:

chunk = ceil(nkv / 32)          # rows assigned to each split
rpw   = ceil(chunk / 4)         # split evenly over 4 warps → rows per warp

The 32-row window's lane-slot mapping is fixed: 8 passes × 4 subgroups, subgroup g claiming window rows 8g..8g+7. That is, a window has exactly 32 lane slots, and when the actual live row count is rpw the slot utilization = rpw/32:

ShapenkvchunkrpwSlot utilizationConsequence
14B @325432541022681%no pathology, 4-warp wins
7B @16411641521341% (59% idle)subgroups 2–3 idle as whole groups
7B tg128128418%3 of 4 warps exit outright; the live warp uses 8/32 lanes

And the per-block fixed cost is amortized over rpw rows: a 4-warp block's Q preload is 128 threads × 64 B = 8 KB (the incumbent 1-warp only 32 × 16 B = 512 B), plus the epilogue's syncthreads/staging and partial write-out. At rpw=13 these costs spread over 13 rows; at rpw=1 over 1 row — total collapse. Conversely the incumbent kernel has no such problem: 0.58–0.83 waves fully resident, every lane busy on every row. The fattn-vec structure has net gain only when rpw = ceil(ceil(nkv/32)/4) ≳ 16 — this doc's first conclusion, which later became D3-4 L1's dispatch threshold (nkv ≥ 1921) directly.

2.3 Why a kernel win may not be a wall-clock win

In one 14B @3254 decode step attention is only 7.4%. So a kernel −6.9% folds into roughly +0.5% on the wall clock — below the +1.5% session bar and inside the ±2% A/B noise band. "A win on a kernel that is not the wall" does not reach the wall (r45's old conclusion) — the veto's second leg.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Grid untouched: dim3(ATTN_SPLITS, n_head), independent of nkv → CUDA-graph capture/replay unaffected. This is D1's iron rule: the grid must not depend on positions.
  • Occupancy first: __launch_bounds__(128, 8) compiles to REG 64 / STACK 0 / SMEM 8960 B = 32 warps/SM (incumbent 24). smem is only the probs row (128 floats) + the epilogue's vkq_s (4×512 floats) + mx_sh/s_sh.
  • Q all in registers: each thread holds its own 16-dim slice (4 × float4, redundantly replicated across the 4 subgroups — llama's same trick); K/V stream straight from global, saving smem K/V staging bandwidth.
  • Dispatch keyed only on hd==128: f16-KV + hd==128 (the decode shape of Qwen2.5/Qwen3) takes the new kernel; hd≠128 keeps the incumbent (parity fixtures undisturbed).
  • LSE-merge epilogue: the 4 warps each hold their own (mx, S, oc) state, merged through smem into one partial — the partial layout stays exactly the incumbent kernel's, so the combine kernel needs no change.

3.2 Key code

D3a's patch lives on in the current tree inside the "hybrid kernel body" (when D3-4 L1 landed, it installed the D3a body verbatim as gqa_attn_split_partial_hybrid), so the excerpts below come from the current tree src/cuda_kernels.cu; D3a's original form of the day (gqa_attn_split_partial_h4w as the sole hd==128 path) is in /tmp/d3/d3a_kernel_patch.diff.

before — the incumbent 1-warp body (D2-staged, row-level softmax chain):

// All of the 4-row window's K+V issued early (D2), then row-by-row online-softmax updates
#pragma unroll
for (int j = 0; j < 4; j++) {
    if (j >= nr) break; // warp-uniform: all lanes exit together
    // full-row dot: this lane's 4 dims + full-warp butterfly, once per row
    float d = q4.x*k4[j].x + q4.y*k4[j].y + q4.z*k4[j].z + q4.w*k4[j].w;
    #pragma unroll
    for (int off = 16; off > 0; off >>= 1)
        d += __shfl_xor_sync(0xFFFFFFFF, d, off);
    float s = d * scale;
    float nmx = fmaxf(mx, s);
    float corr = expf(mx - nmx);
    float e = expf(s - nmx);
    S = S * corr + e;
    mx = nmx;
    if (live) {
        float4 vv = v4[j];
        oc.x = oc.x * corr + e * vv.x;
        oc.y = oc.y * corr + e * vv.y;   // one rescale (corr) per row
    }
}

after — the h4w body: warp stripes + 32-row windows (excerpted from the current tree's attn_split_h4w_body):

// Balanced contiguous stripes: warp w owns rows [lo + w*rpw, +rpw).
const int rpw = (chunk + 3) >> 2;
const int wlo = lo + w * rpw;
const int wend = min(hi, wlo + rpw);

__shared__ float probs[H4W_NTHREADS]; // per-warp 32-float prob stage
float* pw = probs + w * 32;

for (int b = wlo; b < wend; b += 32) {
    const int wl = min(32, wend - b); // rows in this window (warp-uniform)
    ...
    #pragma unroll
    for (int p = 0; p < 8; p++) {
        if (p >= np) break;
        const int row = b + g * 8 + p;      // subgroup g claims 8 rows
        float d = 0.0f;
        if (row < wend) {
            const __half* krow = k + row * stride_kv + hk * hd + 16 * t;
            const uint4 ka = *reinterpret_cast<const uint4*>(krow);
            const uint4 kb = *reinterpret_cast<const uint4*>(krow + 8);
            d = h4w_dot8(ka, qc0, qc1) + h4w_dot8(kb, qc2, qc3);
        }
        float s = h4w_subgroup_sum8(d) * scale;  // 8-lane butterfly
        if (row >= wend) s = -INFINITY;
        mx_new = fmaxf(mx_new, s);
        if (t == p) kq = s;                 // lane (g,t) keeps row (g,p)'s score
    }
    // window max across subgroups (llama: offsets nthreads_KQ..WARP_SIZE)
    for (int off = 8; off < 32; off <<= 1)
        mx_new = fmaxf(mx_new, __shfl_xor_sync(0xFFFFFFFFu, mx_new, off));
    const float wsc = expf(mx - mx_new);    // one rescale per window
    mx = mx_new;
    kq = expf(kq - mx);
    S = S * wsc + kq;
    ...
}

LSE-merge epilogue (4 warp states → one partial):

__shared__ __align__(16) float vkq_s[4 * 512];
__shared__ float mx_sh[4];
__shared__ float s_sh[4];
float Sw = warp_reduce_sum(S);
if (lane == 0) { mx_sh[w] = mx; s_sh[w] = Sw; }
__syncthreads();
const float gmax = fmaxf(fmaxf(mx_sh[0], mx_sh[1]), fmaxf(mx_sh[2], mx_sh[3]));
const float wsc = expf(mx - gmax); // idle warp: exp(-1e38 - gmax) == 0
// ...the four float4 accumulators each multiply wsc, then write to vkq_s by (w, g) slot...
// finally sum across the 4 warps per dim (stride-128 walk, bank-conflict-free), write dst

3.3 Pitfalls

The shfl_sync deadlock (this doc's single biggest lesson, now in Appendix-B). The first version put the subgroup reduction inside the row-validity guard:

if (row < wend) {                       // wrong form: lanes diverge
    d = ...;
    s = h4w_subgroup_sum8(d) * scale;   // __shfl_xor_sync(0xFFFFFFFF, ...)
}

__shfl_xor_sync(0xFFFFFFFF, v, off, 8)'s mask names all 32 lanes, but the 8-lane subgroup's butterfly only needs its own subgroup present — when different subgroups disagree on row < wend (inevitable when rpw does not divide the window; e.g. rpw=13's second window has only 5 rows), some lanes never reach the shuffle point, the names in the mask can never be redeemed, and the warp spins at the hardware level. Symptom: the standalone probe hung at nkv=3; in cargo test it presented as a GPU-spin hang. The fix (the comment at current-tree lines 3050–3055 is its tombstone): compute contributions conditionally (invalid rows default 0.0f), run the reduction unconditionally for all lanes, then mask invalid rows to -INFINITY afterwards:

float s = h4w_subgroup_sum8(d) * scale;  // unconditional: all lanes present
if (row >= wend) s = -INFINITY;          // mask afterwards

Finite initial values. An idle warp / empty stripe's mx must not be -INFINITY (exp(-INF - (-INF)) produces NaN); the h4w body uses -1e38f as the base: exp(-1e38 - gmax) == 0 holds exactly, so an empty stripe's warp contributes exactly zero.

V-side loads and FMAs must be masked together. Rows at the window tail have probability exactly 0, but the KV bytes behind them were never written — "probability 0 so the product is 0" does not hold; what gets read is a NaN from garbage bytes. When the load predicate turns off, the FMA must turn off too (current-tree lines 3074–3078 comment).

4. Verification

This kernel was numerically all-green — every gate passed, which is why it later became the tolerance-gate calibration vehicle. What each gate defends:

  • kernel-level parity probe (defends against math errors): h4w vs the CPU reference swept over all nkv 3..4096 (including 14B's exact in-situ shapes and every chunk boundary) ≤ 1.3e-7; same inputs vs the incumbent kernel ≤ 8.9e-8; on real outlier-magnitude data (residual |q|~50, V outliers ±127) new-vs-CPU 6.5e-5 vs old-vs-CPU 3.8e-5 — the same error class.
  • in-situ per-layer TRACE (defends against layout/indexing errors): NO_CUDA_GRAPH + MINFER_TRACE capturing per-node attn outputs (14B): prefill attention bitwise; decode per-layer deltas L0 2.4e-7, L1–L11 bitwise 0.0 (the f16-KV store is a noise gate — sub-ULP reordering noise is quantized away at each layer's K/V write-back), L12+ 1e-3..5e-1 (noise crosses the f16 rounding boundary and is amplified through outlier-dim cancellation).
  • argmax HARD gate (defends against sampling flips): every dump step's argmax byte-identical, with top-2 margins 4.95 / 10.97, far from the 0.1 threshold.
  • greedy × seeds + sampling contrast (defends against distribution drift): greedy −n 256 × 5 seeds × both models, 0/10 diverged by a single token; the temp 0.8 seed-7 sampling contrast fully identical.
  • suite (defends against the regression surface): 169/0/3 + the FA trio
    • split-decode parity (the parity test drives the new kernel with the hd=128/n_ctx 4200 shape — reverted along with the code).
  • interleaved A/B (defends against window drift): same-window 3× interleaved medians.

Two calibration findings about the dump gate (later written into D3b's aliasing note):

  1. decode node dumps (node{2,3,5,8,11}_decode) read aliased pool slots (node11 = kv_load's dump is only 20 KB, not the 16.7 MB KV region) — node-level diff noise comes from slot aliasing, not values.
  2. The KV-region diff between -n 1 and -n 2 dumps is a pre-existing per-step row rewrite: reproducible pre-vs-pre on the incumbent binary too (layer-10 K's rows 17–20 change between decode steps 1→2). Rule: gate on logits + the final step's KV, and use the same -n on both sides.

5. Results

5.1 The veto numbers and mechanism

nsys (NO_CUDA_GRAPH, bench −n 8 averaged):

Shapeincumbenth4wΔ
14B @3254 split kernel73.4 µs68.4 µs−6.9% (= 71.5% of the 48.9 µs byte floor, was 67%)
7B @1641 split kernel21.1 µs34.7 µs+64%

Wall clock (same-window interleaved 3× medians, tok/s):

ShapeprepostΔbarVerdict
14B @325421.0420.84−0.95%≥ +1.5%fail (a −6.9% kernel is worth ~+0.5% wall: attention is 7.4% of the step)
7B @164147.5345.91−3.4%≥ 47.9fail (tight clusters, same direction as the kernel's +64%)
7B tg128——noise-level—the rpw=1 pathology crushes the theoretical gain

Veto mechanism (why reverted, when a retry is worthwhile): the rpw pathology is structural — the 32-row window's lane-slot mapping is frozen (8 passes × 4 subgroups); whenever rpw < ~16, idle slots and the diluted per-block fixed costs eat everything. This is not tunable; the geometry does not fit the shape. The revert followed the r44 precedent (a parity-green, wall-clock sub-bar kernel change is not kept). Retry conditions are on record: multi-warp decode attention must first solve the rpw pathology — either a subgroup-dense row mapping (letting live rows fill all 32 slots) or a runtime switch back to the 1-warp body by rpw. The latter is the form D3-4 L1 adopted (doc 69).

5.2 The durable output: tolerance-gate calibration

The kernel was numerically all-green (probe 1.3e-7), but the "end-to-end max|Δlogits| ≤ 1e-3" gate proposed at D3-1 failed:

end-to-end logits max|Δ| = 0.376 (14B, 48 layers) vs 0.389 (7B, 28 layers)

The two depths are the same magnitude ⇒ this is a depth-independent error class, not a bug. Mechanism: a kernel-level 1e-7 accumulation-order difference is quantized and amplified layer by layer through f16 rounding boundaries (L12+ already at 5e-1), reaching O(0.4) at the logits. Conclusion: on 28–48-layer models, "end-to-end logits ≤ 1e-3" is unsatisfiable for any rewrite that changes accumulation order (the r50/r57 lesson generalized to decode).

The usable tolerance-gate package (all demonstrated green on this session's experimental kernel):

  1. kernel-level parity vs CPU ≤ ~1e-4, and it must be validated on real outlier-magnitude data;
  2. the argmax HARD gate: every dump step's argmax byte-identical with top-2 margin > 0.1 (measured margins 4.95 / 10.97);
  3. greedy −n 256 × 5 seeds × both models with 0/10 single-token divergence + the temp 0.8 sampling contrast;
  4. suite + FA trio + split-decode parity (new shapes must cover the new dispatch path);
  5. the same-window interleaved A/B bar.

This gate set was adopted directly by D3-4 L1's h4w arm (doc 69), D3-6, and D3-7 2b.

6. Lessons

  1. Shared shape parameters decide a structural rewrite's life or death: the 4-warp fattn-vec structure's efficiency is a staircase function of rpw = ceil(ceil(nkv/32)/4), winning only at ≥16 — before porting any "rows-to-warps" kernel, compute the target shape's rows-per-warp; do not write code first.
  2. __shfl_sync's mask is a whole-warp promise: with 0xFFFFFFFF in the mask, all 32 lanes must reach the point; any lane-divergent guard must move out of the reduction, into "compute contributions conditionally, reduce unconditionally, mask afterwards".
  3. A kernel-level win is not a wall-clock win: a −6.9% kernel that is 7.4% of a step folds to ~+0.5%; before crossing the bar, compute "what fraction of the whole step is this kernel".
  4. An end-to-end logits tolerance is a property of the error class, not of bugs: any accumulation-order change necessarily amplifies to O(0.4) across dozens of f16 store layers — tolerance gates must be pinned to discrete invariants like argmax/greedy/sampling, not to absolute thresholds on continuous quantities.

← 67 · Index · 69 →

69 · D3-4 — L1 hybrid rpw dispatch (LANDED) + L2 window-prefetch pipelining (REVERTED) + long-prompt dump calibration

Result: L1 landed (22336b2) — the dual-kernel self-gating rpw dispatch: 14B @3254 split kernel 72.1 → 62.11 µs (−13.9%) plus a 1.5 µs dud launch, wall clock 21.20 → 21.33 (+0.61%, SEP), all three 7B guards held, and the 1-warp arm is bitwise. L2 reverted — window-level K/V prefetch inside the h4w body: 62.1 → 66.46 µs (+7%); occupancy halved (32→16 warps/SM) + wave growth (3.33→4.44) outweighed the shortened chain; at 79% of the 48.9 µs byte floor the kernel is bytes+tail-bound, not chain-bound. Also archived as calibrations: two pre-existing behaviors — long-prompt dump non-determinism and the CLI n_ctx headroom. Commit: 22336b2 (neither L2 nor the in-kernel fallback form was committed). Date: 2026-09-07.

1. Background — where things stood

D3a (doc 68) ported llama's fattn-vec-style 4-warp split attention; the kernel was numerically all-green but hit the rows-per-warp pathology: rpw = ceil(ceil(nkv/32)/4) is 26 at 14B @3254 (winning 6.9%), 13 at 7B @1641 (+64%), and 1 at tg128 (collapse). After the full revert, one clear rescue path remained, already written into D3a's record: "the 4-warp path serves only dense chunks; at rpw < ~16 fall back to the D2-staged 1-warp body within the same kernel — the branch is nkv-uniform and replay-safe."

Meanwhile the window's other bottlenecks did not sit still. Baseline anchors (same-window interleaved 3× medians, pre side): 14B tg128 23.81→22.96-class (drifting), @3254 21.90→21.20; 7B tg128 50.90→50.28, @1641 49.32→48.78. The sglang co-tenant was resident throughout and the window drifts ±2% — so every judgment must be a same-window interleaved A/B; cross-session absolutes are not comparable (r59b's old rule).

D3-1's teardown also left a "distance account": at 14B @3254 the attention residual (73.4 − 48.9 µs byte floor) is the largest single item, and D3-4's brief made L2 (window-level K/V prefetch) the main lever: 62.1 µs (after L1 landed) is 79% of the floor, the K phase issues 8 serial load→reduce per 32-row window, the V phase 7 load→FMA — if D2's issue-point early-issue trick moved to window granularity, it should in theory harvest a large part of the remaining 21%.

So D3-4 did three things in one session: L1 landed D3a's legacy in dispatch form; L2 tested the brief's prefetch hypothesis; and two metrological traps (long-prompt dump, CLI n_ctx) were calibrated and archived along the way.

2. Principle — the GPU mechanism: the dual-kernel self-gate and the economics of a dud launch

2.1 Why "launch both kernels" is replay-safe

The dispatch condition is rpw = ceil(ceil(nkv/32)/4) ≥ 16 (i.e. nkv ≥ 1921). Both roads:

  • grid shape is nkv-independent: both kernels use the static dim3(ATTN_SPLITS=32, n_head) grid (h4w has H4W_NTHREADS=128 threads, the incumbent 32) — the captured launch configuration never changes, so replay is legal.
  • liveness is recomputed per execution on device: every kernel re-reads positions[0] for nkv at every execution, computes rpw itself, and then exactly one kernel works while the other early-exits wholesale. The branch condition is uniform across the whole grid (positions[0] is a launch-level scalar) — no inter-warp divergence, per-nkv output deterministic.
  • the cost: one extra dud launch per layer, measured ~1.3–1.5 µs. In exchange, each geometry serves its own rpw range without polluting the other.

2.2 Why the first cut (in-kernel fallback) was vetoed: block geometry is destiny

The first cut put both bodies in one 128-thread kernel: rpw ≥ 16 takes the h4w body, otherwise the 1-warp body (only the first 32 threads work). Bitwise all-green (7B @1845-token dump identical per file), but 7B @1641 measured 35.4 vs 19.8 µs (+78%, nsys). The mechanism is not math but geometry:

GB10 caps every SM at 1536 threads.
A 128-thread block → at most 1536/128 = 12 blocks per SM,
of which only 1 warp per block works on the 1-warp arm → 12 working warps/SM.
A 32-thread block → 1536/32 = 48 block slots; the incumbent form runs 24–32 working warps/SM.

The 1-warp body is serially dense; its throughput model is "working warps per SM" — a 128-thread block cuts its resident warp count by more than half. The body did not change; the configuration killed it. This is direct evidence for "geometry, not math": any shared-kernel form that "stuffs the 1-warp body into a bigger block" need not be tried again.

2.3 L2's hypothesis and counter-hypothesis: when prefetch is not free

L2's hypothesis chain: 62.1 µs = 79% × the 48.9 µs floor ⇒ 21% is harvestable chain overhead; D2 proved issue-point early issue is bitwise and free (when registers are plentiful). The counter-hypothesis (known only afterwards): this kernel already presses against ~131 KB/SM of in-flight loads (≈ 30× the latency-BW product) — its in-flight bytes were already oversaturated; the bottleneck is the byte count itself (the composition of the L2 5× re-reads) + the wave tail, not each chain's depth. To fit the prefetch ring's buffers into registers, __launch_bounds__'s minBlocks must drop from 8 to 4 (register budget 64→128), and occupancy goes 32→16 warps/SM, waves 3.33→4.44 — the occupancy tax for shortening the chain is 1:2, and a shortened chain cannot save a bytes-bound kernel.

3. Implementation

3.1 Design choices (why this shape and not another)

  • Dual kernels, shared device function: the incumbent 1-warp body was extracted into attn_split_1w_body (math and indexing byte-identical, only the index setup moved to the caller), and the incumbent kernel and the hybrid dispatch share that one source — the 1-warp arm's bitwise property is guaranteed structurally by "the same body".
  • The h4w body reuses D3a verbatim: the probe already verified ≤1.3e-7 vs CPU and the tolerance class is calibrated; no reinvention.
  • The gate lives on device, not host: the host cannot know each decode step's nkv in advance (that would need a sync); the kernel reads positions[0] and gates itself — zero host-side logic, zero sync.
  • Dual launch only when hd==128; other head dims (including the hd=8 parity fixture) keep the single incumbent launch (rpw_gate=0), leaving the parity surface undisturbed.

3.2 Key code

The incumbent kernel's rpw_gate early exit (current tree src/cuda_kernels.cu):

template <typename KV>
__global__ void gqa_attn_split_partial(..., int rpw_gate) {
    // D3-4 L1 dual-kernel dispatch: when rpw_gate > 0 and the 4-warp kernel
    // owns this nkv (rpw >= rpw_gate), exit before touching anything — the
    // hybrid kernel writes the partial. The branch is nkv-uniform across the
    // whole grid (positions[0] is launch-wide), so this stays replay-safe,
    // and for every nkv the incumbent path takes, the arithmetic in the body
    // is unchanged (bitwise; dump-memcmp gated).
    if (rpw_gate > 0) {
        const int nkv0 = positions[0] + 1;
        const int chunk0 = (nkv0 + ATTN_SPLITS - 1) / ATTN_SPLITS;
        if (((chunk0 + 3) >> 2) >= rpw_gate) return;   // h4w owns this nkv
    }
    attn_split_1w_body<KV>(q, k, v, partial, positions[0] + 1, ...);
}

The hybrid kernel's dual gate:

#define H4W_NTHREADS 128 // 4 warps per block
#define H4W_MIN_RPW 16   // 4-warp body only when rows/warp amortize the window

__global__ void __launch_bounds__(H4W_NTHREADS, 8)
gqa_attn_split_partial_hybrid(const float* q, const __half* k, const __half* v,
                              float* partial, const int* positions, ...)
{
    const int nkv = positions[0] + 1;
    const int chunk = (nkv + ATTN_SPLITS - 1) / ATTN_SPLITS;
    if (((chunk + 3) >> 2) < H4W_MIN_RPW) return;      // rpw < 16 → exit
    attn_split_h4w_body(q, k, v, partial, nkv, blockIdx.x, blockIdx.y, ...);
}

The launcher's dual launch (launch_gqa_attn_split_f16kv):

if (hd == 128) {
    gqa_attn_split_partial<__half><<<dim3(ATTN_SPLITS, n_head), 32, 0, stream>>>(
        q, (const __half*)k, (const __half*)v, partial, positions,
        n_head, n_head_kv, hd, scale, pstr, H4W_MIN_RPW);      // 1-warp arm
    gqa_attn_split_partial_hybrid<<<dim3(ATTN_SPLITS, n_head), H4W_NTHREADS, 0, stream>>>(
        q, (const __half*)k, (const __half*)v, partial, positions,
        n_head, n_head_kv, hd, scale, pstr);                    // h4w arm
} else {
    gqa_attn_split_partial<__half><<<dim3(ATTN_SPLITS, n_head), 32, 0, stream>>>(
        q, ..., /*rpw_gate=*/0);                                // single launch
}
gqa_attn_split_combine<<<dim3(1, n_head), hd, 0, stream>>>(partial, o, ...);

The two kernels write the same partial region, but per nkv exactly one is alive — combine is unaware the dispatch exists. The dud arm's whole-grid early exit is a launch-level scalar branch under both 32/128-thread blocks, no divergence cost, only the ~1.3–1.5 µs launch itself.

3.3 Pitfalls

  • The in-kernel fallback's +78% (§2.2): bitwise-green ≠ landable. The 1-warp body is sensitive to block size; any shared-block form must compute working-warps/SM before speaking.
  • The attn_split_1w_body extraction must preserve the math byte for byte: the extraction left "each split's row-range computation" in the caller and the body takes only (nkv, sp, h, ...) — same rows, same order, same row-level ops, only the index setup moved. The dump-memcmp gate confirmed the 1-warp arm's output is bit-identical to pre (7B @1845: 71/71 files identical, see §4).
  • A co-tenant outlier on the pre side: 7B @1641's pre series contained one 44.40 co-tenant outlier rep — the median held, and the record explicitly annotates the outlier's attribution rather than silently dropping it.

4. Verification

  • The 1-warp arm's bitwise dump gate (defends against "the shared extraction changed the math"): 7B @1845-token prompt, logits prefill+decode, all KV, decode nodes — 71/71 files byte-identical; the 3 node{3,5,8}_prefill diffs are D3b's already-calibrated slot-aliasing trio, reproducible pre-vs-pre.
  • The h4w arm's tolerance class (D3a's gate set, doc 68 §5.2): 7B @2800 (rpw=25, h4w territory) max|Δlogits| 0.309 (the calibrated 0.39 class); argmax identical every dump step, margin 0.716 (HARD gate > 0.1); upper KV layers show the f16 noise pattern (kv0–7 bitwise, kv8–27 drifting on the decode side).
  • greedy × seeds + sampling contrast (defends against distribution drift): greedy −n 256 × 5 seeds × both models, h4w-regime prompts: exactly one divergence each at the regime entry point (1/256 = 0.4% < 2%), coherent continuation afterwards (no repeated degradation); the temp 0.8 seed-7 contrast identical.
  • suite (defends against the regression surface): 169/0/3, with the parity test extended by one hd=128/n_ctx-4200 shape whose pos0 sweep crosses exactly the rpw 15/16 dispatch boundary (nkv 1920/1921) — parity coverage on both sides of the dispatch boundary.
  • guards (defend against regressions at other shapes): 7B ≥ 49.0 / ≥ 47.9 and 14B tg128 ≥ 22.7 all held.
  • nsys kernel level (defends against "wall-clock noise hiding the kernel truth"): NO_CUDA_GRAPH, bench −n 8 averaged over the last 384.

5. Results

5.1 L1 (LANDED, 22336b2)

Kernel level (nsys, NO_CUDA_GRAPH):

ShapeprepostΔ
14B @3254 split72.1 µs62.11 µs−13.9% (plus a 1.5 µs dud launch)
7B @1641 split—20.7 µs + 1.3 µs dudthe incumbent path's min equals pre → the body undisturbed

(In the D3a era the same shape was 73.4 → 68.4 — the hybrid's h4w arm is faster than D3a's full form, because small-rpw shapes no longer pollute the same bin.)

Wall clock (same-window interleaved 3× medians, tok/s):

ShapeprepostΔVerdict
14B @325421.2021.33+0.61%SEP: min-new 21.25 > max-base 21.23
14B tg12822.9622.94−0.09%guard ≥ 22.7 held
7B tg12850.2850.20−0.16%guard ≥ 49.0 held
7B @164148.7848.68−0.20%guard ≥ 47.9 held (pre contained a 44.40 co-tenant outlier)

The distance account (post-L1, 14B @3254): minfer 21.33 t/s = 46.88 ms/step vs llama 24.32 = 41.12 ms → still 5.76 ms short (−12.3%). Known lever list: attention residual (62.1−48.9)×48 = 0.63 ms; D3-1's three MMVQ stragglers (attn_v-q6K 0.28 + ffn_down-q6K 0.49 + output-head 0.44) = 1.21 ms; D3c elementwise fusion ≈ 1.0 ms (projected +2.1% wall, not implemented this session). Together ≈ 2.84 ms = 49% of the gap → landing all of them reaches only ~22.6 t/s (0.93×); the remaining ~2.9 ms is the matmul account (D3-1's wall-effective 194.9 vs llama 207.6 GB/s) + launch-structure slack. 7B: tg128 1.016× (ahead), @1641 0.985× — the 7B decode campaign is de facto closed.

5.2 L2 (REVERTED, uncommitted)

patch_l2.py: K-phase software pipeline (+2 uint4 per thread) + a 4-deep V-bulk ring (+8 uint4), __launch_bounds__ minBlocks 8→4 traded for register budget (64→128 cap).

Metricpre (h4w)postΔ
14B @3254 h4w kernel62.1 µs66.46 µs+7%
occupancy32 warps/SM16 warps/SMhalved
waves (1280 blocks)3.334.44+33%

Veto mechanism: at the saturation point of 131 KB/SM in-flight loads (≈30× the latency-BW product), shortening a single chain's depth produces no time — the time lives in the byte count (the L2 5× re-read composition) and the wave tail. D2's "issue-point moves are free" holds only when registers are plentiful; at a 64-reg budget any pipeline funding is a 1:2 occupancy trade. The bitwise gates were never reached (the revert came before them). Retry conditions: prefetch has room for discussion only after the L2 re-read byte count itself is cut (a bytes-side lever, e.g. GQA q-head batching — which D3-6 later showed is also not the residual) or a register-free staging is found.

5.3 Metrology calibrations (RECORDED, no code)

  • long-prompt dump non-determinism: at prompts ≥ 2.8K tokens the MINFER_GRAPH_DUMP PREFILL-phase files (all kv*_prefill, logits_prefill, prefill nodes) are non-deterministic even pre-vs-pre (wholesale, garbage-magnitude — an aliasing problem in the dump read path); decode-phase dumps stay deterministic. Rule: for long prompts the dump gate must be anchored pre-vs-pre on the exact shape, and gate only decode-phase files.
  • CLI n_ctx headroom: prompts above the default n_ctx 4096 leave zero generation headroom (the position N exceeds n_ctx N panic; the record cites graph.rs:367, the current-tree line has drifted to 418); bench unaffected. Until n_ctx sizing is fixed, use shapes with prompt + n ≤ 4096 for long-prompt greedy gates.

6. Lessons

  1. One dud launch per layer (~1.4 µs) is fair rent for "geometry self-gating": a device-side nkv-uniform branch + static grid preserves replay, far cheaper than a host-side synchronized decision.
  2. A bitwise-green kernel can still be +78%: the shared-kernel fallback changed the block geometry (working warps/SM), not the math — a body is sensitive to "how big a block it lives in".
  3. Occupancy and chain depth trade 1:2 (under a 64-reg budget): adding a pipeline ring to a kernel whose in-flight bytes are already saturated buys a chain shortening that cannot beat halved warp residency.
  4. A dump gate's determinism must be calibrated per phase and per shape: at long prompts even pre-vs-pre prefill-phase dumps are non-deterministic — gate only on the phase proven deterministic.

← 68 · Index · 70 →

70 · D3-5 — decode-alignment plan Stage 1: fused-producer decode A-quantize (LANDED) + negative analysis of the output-head/ffn_down geometry levers (1b/1c)

Result: 1a landed (3230b2b) — rms_norm_quant_pad40 and swiglu_quant_pad40 write a pad40 q8 plane alongside the f32 output; decode MMVQ consults via decode_quantize_native (MmqCache, keyed by the producer's f32 output pointer) and on a hit skips the standalone quantize: 14B @3254 standalone quantize launches 4448 → 964 (−78%), wall clock 14B tg128 +1.30%, @3254 +1.50%, 7B tg128 +1.66%, @1641 +1.43% (all SEP; the q8 bytes are bitwise by construction). 1b/1c analysis-negative — neither the output-head od-split/512-thread nor the ffn_down 512-thread single-unit geometry variant has a measurable mechanism; per the "measure the mechanism before writing code" rule, not built. Commit: 3230b2b (1b/1c no repo change). Date: 2026-09-08.

1. Background — where things stood

After D3-4 landed the hybrid rpw dispatch (doc 69), 14B @3254's distance account read 46.88 vs 41.12 ms/step (−12.3%), and the largest class in the known lever list was neither GEMM nor attention but launch structure: D3-1's census counted ~700 sub-2 µs launches per decode step, of which ~265 are standalone quantize_q8_0_pad40 — one hangs in front of every decode MMVQ matmul, ~1.7 µs each, ~0.46 ms/step total. More embarrassing, these quantizes repeat work: one row of attn_norm output is re-quantized 3 times per layer before the q/k/v matmuls. r49 had already prescribed for the same disease on the prefill side (the A-quantize prepass's shared-A window memoization), and r51/r52 advanced it to "the producer writes the q8 plane directly" (mode 1/2) — prefill's quantize launches fell from 193 to 28. The decode side was still the old world.

This doc is Stage 1 (1a) of the decode-alignment plan, plus the paper settlement of two geometry levers listed in the brief (1b/1c). Anchors (same-window interleaved 3× medians, pre side): 14B tg128 22.97 (interleaved pre 23.05), @3254 21.06 (one 15.87 co-tenant outlier rep; interleaved pre 21.36); 7B tg128 49.90 (interleaved 49.86), @1641 48.71 (interleaved 48.47). The sglang co-tenant was resident throughout (idle co-residency is equivalent to a clean machine; per-rep outliers land on both sides and the median carries).

Where things stall without this step: launch structure is a length-independent fixed tax — tg128 and @3254 fire the same number of quantize launches per step. It presses on every shape's wall clock and aligns directly with CUDA-graph's "one launch ≈ 2 µs pool" account: 78% fewer quantize launches ≈ reclaiming nearly 350 launch slots per step.

2. Principle — the GPU mechanism: the three-layer account of moving quantization back to the producer

2.1 The launch account

One decode step fires ~265 quantize launches × ~1.7 µs ≈ 0.46 ms/step. These launches waste twice: (a) the launch itself plus the f32 row's global round-trip (writing the q8 plane means first reading the f32 row back); (b) the same row is repeatedly quantized by adjacent matmuls (attn_norm output 3×/layer). Both follow from the layout choice of "hanging quantization on the consumer side".

2.2 Why not inline-per-block (the A-side amplification)

The most obvious form is to have each MMVQ row-block quantize its own A rows. Dead on paper: every row-block would re-read the whole A row (f32), multiplying A-side L2 traffic by od × 14 KB — one ffn_gu layer is on the order of +18.6 GB/step. This is exactly r49's shared-A lesson: the quantization input is L2-hot to the producer, while the consumer is a weight-streaming machine; making the consumer go back and re-read A swaps the cheapest read for the most expensive one. So the only correct landing spot is the producer.

2.3 Constructive bitwise and the MmqCache window

The q8 plane's bit-for-bit invariance is argued constructively, requiring no belief in numerical coincidence:

  • max is exact for any associativity — amax's fmaxf chain is the same number however associated;
  • rintf/clamp are per-element — every byte of the quantized payload depends only on its own f32 value.

So as long as the epilogue reuses the standalone kernel's per-32-block body (quantize_pad40_block verbatim), the q8 bytes are bitwise. After the producer writes the f32 output it barriers, and the epilogue re-reads the row it just wrote — L1-hot, cheaper than the standalone reading f32 from global.

Cache correctness follows r49's MmqCache window rules: the plane is a pure function of (src, nt, id); any non-(MatMul|FusedFFN) node clears entries; synchronize clears at execution boundaries. The hit path must also re-verify the physical pointer (one get_or_grow between record and consult may relocate the plane — an entry whose q8 moved falls back to the standalone launch, re-quantizing from the live f32 src). FusedFFN joins the cache-clear preserve set: its input is ffn_norm's rms output and its internal gu matmul is that src's first consumer — the same "contiguous consumer window" as MatMul→MatMul. attn_o keeps the standalone quantize: its producer is the attention node, untouched this session (regression-risk/benefit ratio unfavorable).

3. Implementation

3.1 Design choices (why this shape and not another)

  • A shared epilogue body: the quantize_pad40_block device function is simultaneously the standalone kernel's and both fused producers' body — "the epilogue IS the standalone body" is the code-level form of the bitwise claim.
  • The native (pad40, non-transposed) form: decode MMVQ consumes the pad40 native plane (only prefill's MMQ wants the transposed pad40_t), so the producer records a native entry (transposed=false).
  • The opt-out ring: MINFER_NO_DECODE_A_FUSE=1 skips both consult and record, restoring a path bit-identical to pre-D3-5 — the A/B gate.
  • Caller-side gates: the rms arm n == 1 && d % 32 == 0; the swiglu arm n % 32 == 0 (whole 32-blocks; supported models' fused-FFN intermediate is always a multiple of 32).

3.2 Key code

The shared per-32-block body (current tree src/cuda_kernels.cu, the standalone and fused epilogues share one body):

// 40B layout: 2B f16 d, 2B pad, 32B int8 payload (offset 4),
// 4B i32 sum of the quantized values (offset 36 — the pad40 slack).
__device__ __forceinline__ void quantize_pad40_block(
    const float* __restrict__ src, uint8_t* __restrict__ dst
) {
    float4 sv[8];
    #pragma unroll
    for (int v = 0; v < 8; v++)
        sv[v] = *reinterpret_cast<const float4*>(src + 4 * v);
    float am = 0.0f;
    #pragma unroll
    for (int v = 0; v < 8; v++)
        am = fmaxf(am, fmaxf(fmaxf(fabsf(sv[v].x), fabsf(sv[v].y)),
                             fmaxf(fabsf(sv[v].z), fabsf(sv[v].w))));
    float d = am / 127.0f;
    float di = (d != 0.0f) ? 1.0f / d : 0.0f;
    *reinterpret_cast<__half*>(dst) = __float2half(d);
    ...
    for (int v = 0; v < 8; v++) {          // rintf + clamp per element → bitwise
        const float* e = &sv[v].x;
        uint32_t p = 0;
        for (int j = 0; j < 4; j++) {
            int q = int(rintf(e[j] * di));
            q = max(-128, min(127, q));
            p |= (uint32_t)(uint8_t)(int8_t)q << (8 * j);
            s += q;
        }
        packed[v] = p;
    }
    ...
}

The fused rms epilogue (the rms body verbatim first, then a whole-block barrier + re-reading the row it just wrote — L1-hot):

    // epilogue: whole block arrives (all threads share `row`), then each
    // thread quantizes blocks tid, tid+blockDim.x, ... of its own row.
    __syncthreads();
    int nb = d / 32;
    const float* src = y + (size_t)row * d;
    uint8_t* dst = q8 + (size_t)row * nb * Q8PB;
    for (int b = tid; b < nb; b += blockDim.x)
        quantize_pad40_block(src + (size_t)b * 32, dst + (size_t)b * Q8PB);

The fused swiglu (before: the 6-line swiglu_f32_off body; after: the body preserved verbatim + barrier + 8 threads quantizing the block's 8 output blocks):

// Block bx wrote output elements [bx*256, bx*256+256) = quant blocks
// bx*8 .. bx*8+7, so 8 threads per block re-read them (L1-hot) and quantize;
// across the grid this is the same thread-count as the standalone kernel
// (one thread per 32-block). REQUIRES the 256-thread launch geometry.
__global__ void swiglu_quant_pad40(float* __restrict__ buf,
                                   uint8_t* __restrict__ q8, int n, int off) {
    int tid = blockIdx.x * blockDim.x + threadIdx.x;
    if (tid < n) {                                  // the body verbatim (guarded, no early return)
        float g = buf[tid];
        buf[tid] = (g / (1.0f + expf(-g))) * buf[off + tid];
    }
    __syncthreads();
    int b = blockIdx.x * 8 + (int)threadIdx.x;
    if (threadIdx.x < 8 && b < (n >> 5))
        quantize_pad40_block(buf + (size_t)b * 32, q8 + (size_t)b * Q8PB);
}

The consumer-side consult (decode_quantize_native):

#![allow(unused)]
fn main() {
fn decode_quantize_native(&self, x: *const f32, id: usize, nt: usize) -> *mut u8 {
    let need = nt * (id / 32) * 40;
    let key = (x as usize, nt, id);
    if !Self::no_decode_a_fuse() {
        let cache = self.mmq_cache.lock().unwrap();
        if cache.active && !cache.transposed && !cache.dead_write
            && cache.key == key {
            let q8 = Self::get_or_grow(&self.buf_q8_decode, need) as usize;
            if q8 == cache.q8 {                    // re-verify the physical pointer
                return q8 as *mut u8;              // hit: skip the standalone launch
            }
        }
    }
    let q8 = Self::get_or_grow(&self.buf_q8_decode, need) as *mut u8;
    unsafe { launch_quantize_q8_0_pad40(x, q8, id as i32, nt as i32, stream); }
    if !Self::no_decode_a_fuse() {
        self.record_mmq_cache_native(key.0, nt, id, q8 as usize);
    }
    q8
}
}

All three decode MMVQ entries (q4_K/q6_K/q5_K classes) begin with this one consult; on miss the behavior is exactly pre-D3-5 (launch + record). The producer-side hook sits in the backend's rms/swiglu arms:

#![allow(unused)]
fn main() {
// D3-5 1a: decode (n==1) producers fuse the pad40 q8 epilogue — bit-identical
// f32 y and q8 bytes, one launch fewer per producer, and the following decode
// matmul group skips its standalone quantize.
if n == 1 && d % 32 == 0 && !CudaState::no_decode_a_fuse() {
    self.state.rms_norm_quant_on_gpu(x, wptr, out, d, n, eps);
    return Ok(());
}
}

3.3 Pitfalls

  • Barrier completeness under "no early return": the swiglu body of swiglu_quant_pad40 was originally the if (tid >= n) return; early-return form — the fused version must become guarded (every thread reaches the barrier), otherwise threads in tail blocks return early and the __syncthreads() deadlocks (an inherent clause of the GPU safety rules).
  • Launch-geometry coupling: swiglu_quant_pad40's epilogue depends on the 256-thread-block invariant "block bx wrote [bx256, bx256+256)", stated as REQUIRES in the kernel comment — geometry and body must not evolve separately.
  • A new member of the cache window: without FusedFFN in the preserve set, ffn_norm's rms plane would be cleared when the FusedFFN node executes and the gu matmul's consult would miss forever (functionally correct, optimization inert) — it must be treated like MatMul.

4. Verification

  • bitwise probe (defends against "the epilogue is not the standalone body"): new test cuda_decode_a_quant_fuse_bitwise: fused vs standalone q8 buffers memcmp-equal at 14B hidden size and the ffn_down shape; f32 producer outputs bit-identical; MMVQ matmul outputs bit-identical via the cache-hit path.
  • suite (defends against the regression surface): 170/0/3.
  • dump gate (defends against end-to-end numerical drift): 14B short prompt −n 4: logits_{prefill,decode} + all kv* byte-identical; the 3 same-size node-dump diffs (node{2,3,11}_decode) are the calibrated pool-slot-aliasing class — reproducible both pre-vs-pre and post-vs-post (the same binary's diff set covers all of pre-vs-post's differences, i.e. no numerical delta).
  • greedy byte-for-byte (defends against sampling flips): the −n 256 token stream byte-identical vs pre on both models (14B @2799-token prompt, 7B @1847).
  • Operational note: prompt_3k3 (5451 tokens) panics at the n_ctx 4096 default on both binaries (the headroom bug D3-4 calibrated) — until n_ctx sizing is fixed, long-prompt greedy gates must use ≤2.8K-token prompts.
  • nsys (defends against "the launch account is miscounted"): 14B @3254 (NO_CUDA_GRAPH, bench −n 8).

5. Results

5.1 1a (LANDED, 3230b2b)

nsys launch account (14B @3254):

MetricprepostΔ
standalone quantize_q8_0_pad40 launches4448964−78% (the remaining ~50/step = the attn_o class)
total kernels2940225918−3484
sub-2µs launches1517111723−3448 (= the quantize delta)
swiglu kernel2.06 µs2.31 µsepilogue +0.25 µs, counted into the wall clock
fused rms—~+0.2 µssame

Wall clock (same-window interleaved 3× medians, all SEP):

ShapeprepostΔ
14B tg12823.0523.35+1.30%
14B @325421.3621.68+1.50%
7B tg12849.8650.69+1.66%
7B @164148.4749.16+1.43%

Wins at all lengths, both models — the launch/round-trip tax is length-independent, matching the prediction. Against the pre-D3-5 anchors: 14B tg128 +1.65% (session bar ≥ +1.5% cleared), @3254 +2.95%. Against llama.cpp: 14B 0.961× (tg128) / 0.891× (@3254); 7B 1.026× (tg128, ahead) / 0.995× (@1641, parity) — the 7B decode campaign closes. 1a itself reclaimed ~0.56–0.69 ms/step (the quantize launches' −78% plus their launch gaps).

The distance account (post-1a, 14B @3254): minfer 21.68 t/s = 46.12 ms/step vs llama 24.32 = 41.12 ms → still 5.00 ms short (−10.8%). List: attention residual 0.63 ms (Stage 2's GQA q-head batching); the attn_v-q6K straggler 0.28 ms; the output head 0.44 ms (1b judged mechanism-free); the rms/elementwise chain 1.0 ms (97 rms launches — the fused epilogue made rms slightly heavier; the next elementwise lever is rms launch consolidation). Together ≈ 2.35 ms = 47% of the remaining gap.

5.2 1b: output-head od-split / 512-thread rework (ANALYSIS-NEGATIVE, not implemented)

The brief's occupancy account, computed first: lm_head q6_K's npair=160 → 96 of 256 threads idle (the D3b-1c knob, already measured neutral); 6×256-thread blocks/SM = 288 resident rows, while D3b-1c's 9×160 = 432-row form is also neutral — rows-in-flight is not the limiter; the head sustains 47.6 blocks/µs while ffn_gu demonstrates 76/µs — block dispatch rate is not the limiter either. The two-row 512-thread form could be bitwise (each 256-thread half-block reduces its own row with the same 8-warp tree), but it moves none of the quantities already measured neutral. Per the brief's own rule (pick the variant with a MEASURED mechanism), not built. The head's 200.1 vs the 220–225 GB/s class remains unexplained by any geometry knob — the same class as D3b-1a's conclusion on the padded kernel (a latency/L2 composition, not a parallelism deficit).

5.3 1c: ffn_down-q6K 512-thread single-unit (ANALYSIS-NEGATIVE, not implemented)

Not bitwise against the landed q6_k_q8_mmvq_v2_pf: in the 256-thread form thread t adds fma(u_t) + fma(u_{t+256}) into the same float accumulator before the block reduction; in the 512-thread form those two units live in different threads and their sum moves into the 16-warp cross-warp reduction order — a different float sum (the r50/r57 class). A bitwise-emulating form (an smem pair-exchange so thread t still adds u_t+u_{t+256} first) needs an extra barrier and buys zero resident-parallelism gain (5120 rows = 18 waves in both forms); and the exposed path 1c meant to treat is already fixed by v2_pf's dual-unit early issue.

5.4 Follow-ups (the list this session left)

(1) attn_o keeps the standalone quantize (48/step, ~0.08 ms) — fusing it into the attention combine epilogue would touch D3-4's hybrid kernel pair; low value, high regression risk, skipped. (2) Stage 2: GQA q-head batching (0.63 ms) is the largest remaining single item. (3) rms launch consolidation (97 rms → fewer, wider; the 1.15 ms class) is the remaining elementwise mass. (4) The long-prompt n_ctx headroom fix (pre-existing; still blocking long-prompt greedy gates).

6. Lessons

  1. The correct landing spot for quantization launches is the producer: the quantization input is L1/L2-hot to the producer and a global out-of-cache read to the consumer; inline-per-block makes the most expensive reader re-read A (the od×14 KB amplification), producer-fusion lets the cheapest reader write it in passing.
  2. Constructive bitwise precedes numerical verification: max is exact for any associativity and rintf/clamp are per-element — write the epilogue as the standalone body's verbatim reuse and bitwise goes from "measured" to "proven"; the probe is just a recheck.
  3. A fused kernel's body fidelity includes control flow: the original body's if (...) return; early return deadlocks under the new barrier semantics and must become a guard; "body verbatim" must preserve control flow too.
  4. Analysis-negative output is also numbers: 1b's three candidate mechanisms (idle threads, resident rows, dispatch rate) each got a measured-neutral value attached, closing a branch more cheaply than building another kernel.

← 69 · Index · 71 →

71 · D3-6 — GQA q-head batched attention: all gates green, still reverted — the 5× L2 re-read was not the residual (REVERTED)

Result: ncu lts__t_sectors 2,121,671 → 472,583 (multiple of the analytic 1-pass minimum of 458K 4.63× → 1.03×) — the 5× K/V re-read was eliminated exactly as designed; yet live kernel time was flat: h4w 63.76 → batched 64.79 µs/layer (+1.6%, the target had been ≤54 µs). Conclusion (mechanism pinned): the 5× L2 re-read was already fully latency-hidden in the live regime; the attention residual (62.1 − 48.9 = 13.2 µs/layer ≈ 0.63 ms/step) is exposed latency, not traffic — every bytes-side lever is hereby ruled out. All correctness gates went green before the revert: parity ≤1e-4 (gqa=5/7, all shapes), argmax HARD gate (margins 2.187 / 0.557), rp=1.0 greedy byte-identical on both models. Commit: no code commit (the GQA-batched kernel existed only as /tmp/d3/patch_2a.py and was lost with /tmp; what survived is docs commit 92b0712 plus the GQA geometry coverage left in the test table). Date: 2026-09-08.

1. Background — where things stood

After D3-5 landed the fused-producer decode A quantization (doc 70), the distance ledger of the 14B decode campaign read: @3254 long context 21.68 t/s = 46.12 ms/step, llama.cpp 24.32 t/s = 41.12 ms — still 5.00 ms short (−10.8%). D3-5's follow-up list laid out the known levers and priced each one: attention residual ≈ 0.63 ms, attn_v-q6K padded-f32 tail ≈ 0.28 ms, output head ≈ 0.44 ms, rms/elementwise chain ≈ 1.0 ms — about 2.35 ms combined, 47% of the remaining gap. The attention residual was the largest single lever, and the only one that came with "a clear mechanism hypothesis": when D3-4 landed the hybrid rpw dual-kernel dispatch, the judgment it left behind was — the part of attention time (62.1 µs/layer) above the DRAM floor (48.9 µs/layer) comes from the GQA 5× re-read of K/V, the so-called "L2 composition".

The provenance of that hypothesis is worth recording, because it decides what this step means. D1's attribution session initially concluded the decode split-attention was latency-bound (staging depth was the only kernel that grew with KV); D3-4, when landing the 4-warp h4w body, re-attributed the residual to L2 re-read composition — but that was an inference, not an experiment: at the time there was no control that "cut the re-read". In other words, at the opening of D3-6 both attributions were still alive, and they made opposite predictions for the same experiment. Which one was right decided whether Stage 3's attention direction would be "cut traffic" or "cut latency".

The re-read geometry is straightforward: 14B is 40 q heads : 8 kv heads (gqa=5). The h4w split-attention grid is (ATTN_SPLITS=32, n_head=40) — each block serves one q head. The K/V rows of the same kv head are read once each by 5 mutually independent blocks (different SMs, no shared L1), so L2 sees 5× the sectors. 7B is more extreme: 28:4, gqa=7. D3-4's own words recorded it as "L2 composition" — the residual is composed of re-read traffic.

The hypothesis deserved an immediate lottery ticket, because its payout path is remarkably clean on paper: put all gqa q heads of the same kv head into one block, loop the windows reading K/V once, with the five warps each computing their own q — traffic returns to 1× and the math does not move a single bit. The D3-6 session (base /tmp/d3/minfer_pre_d36, sha1 92485d3c, HEAD 6085928 window; sglang co-tenant present throughout, same-window interleaved 3× median method, exactly the D3-5 protocol; guards: 7B tg128 ≥ 49.0 / @1641 ≥ 47.9, 14B tg128 ≥ 22.7) built it, all gates went green, and then the combined ncu/nsys evidence vetoed it. What was vetoed was not code quality but the mechanism itself — and that is precisely this step's most valuable output: it re-judged Stage 3's attention direction from "cut traffic" to "cut exposed latency".

One session-boundary note: item 2b from the D3-4/D3-1 follow-up (attn_v-q6K → MMVQ routing) did not get touched this time (no time); it passed intact to the next session — it is doc 72's 2b.

2. Principle — the GPU mechanism

The byte ledger of the GQA re-read. The K+V that 14B per-layer decode attention must read at nkv 3254 is: nkv × n_head_kv × hd × 2 B × 2 (K+V) = 3254 × 8 × 128 × 2 × 2 ≈ 13.3 MB. At GB10's ~273 GB/s effective bandwidth, one full read has a DRAM floor ≈ 48.9 µs — matching the measured floor. The h4w grid is (ATTN_SPLITS, n_head), and every q head's block must read its kv head's whole stripe once, so the request count on the L2 side is 5× (the DRAM side stays near 1×, because after the first read the same rows are resident in L2 and the other 4 passes hit L2). ncu's later measurement matched this analysis: the analytic minimum of one 1-pass ≈ 458K sectors (13.3 MB ÷ 32 B/sector ≈ 416K, the remainder is Q and partial writes), and h4w measured 2,121,671 = 4.63× (the part short of 5× is the first pass, which goes to DRAM — it does not travel the duplicated L2-sector counting path).

The cache hierarchy decides at which level the "5× re-read" happens. L1 is private per SM: the 5 q-head blocks almost certainly land on 5 different SMs, their loads of the same rows cannot see each other, and all of it floods into L2. The batched shape puts the 5 warps into the same block (same SM, shared L1): five warps read the same K row in the same window step — the first warp's load brings the row into L1, and the other four hit it directly. That is "L1 catches the sibling warps' same-address loads", and it is the microscopic mechanism by which sectors can drop to 1.03×.

The batched shape's promise and its price. Promise: grid (ATTN_SPLITS, n_kv_heads), 32*gqa threads per block (gqa=5 → 160, gqa=7 → 224), K/V traffic ÷5, paper time should fall from 63.76 µs toward the traffic-dominated limit (projection ≤54 µs). Price: blocks 40 → 8, threads per block 128 → 160/224. Totals reconciliation: h4w's grid (ATTN_SPLITS=32, 40 heads) = 1280 blocks × 4 warps = 5120 warps; batched (32, 8 kv heads) = 256 blocks × 5 warps = 1280 warps — total warps ÷4, blocks ÷5. What is resident on an SM is no longer "short 4-warp blocks, many small scheduling units" but "long 5-warp blocks, coarse-grained scheduling": the 5 warps in one block must enter and exit the same window loop together, and any warp's long stripe drags the whole block; conversely, the 5 copies of QK^T/V work for the same kv head share one KV read, and intra-block L1 reuse is the entire source of the gain. The partial buffer is shape-invariant: ATTN_SPLITS × n_head × pstr × 4 B = 32 × 40 × 136 × 4 ≈ 696 KB (pstr = (4+hd+3)&!3 = 136), and both shapes write the same [sp][head] layout.

The critical counterweight: the latency roofline. Decode attention is a GEMV-class kernel: each warp's time is a serial dependency chain (fetch K row → 8-dim partial dot → subgroup reduction → window max/rescale → fetch V row → FMA accumulate), 32 rows per window, windows back to back, with only one rescale's worth of parallel slack between windows. If that chain's depth decides kernel time and L2 bandwidth has slack anyway, then "reading 5× more" just makes L2 work extra in the shadow of other work — it never enters the critical path. That is exactly D1's original attribution (latency-bound); D3-4's "L2 composition" was an over-correction of it. D3-6's experiment design was therefore naturally a hypothesis trial: the traffic hypothesis predicts 63.76 → ~45–54 µs, the latency hypothesis predicts flat.

A contrast that must be thought through first: the serialized profiler sees the opposite regime. ncu serializes replay by default: cold L2, no concurrent memory-overlap — traffic dominates only in that regime. So ncu showing −28% while nsys live shows +1.6% is not a contradiction: the two measure two different machines. Live time is what the wall clock wants; ncu's value is proving that the batched shape "really did cut the traffic", which pins the veto reason at "traffic is not on the critical path" rather than "the batched shape was built wrong".

3. Implementation

3.1 Design choices (why this shape and not another)

  • warp = q head, window loop shared. The batched shape did not rewrite the window math: the per-warp 32-row window pass inside attn_split_h4w_body was extracted verbatim into a shared device function h4w_warp_windows, parameterized per warp (q head, stripe range). K/V base addresses are resolved once by hk, so the 5 warps share each window. Verbatim reuse of the window math is the precondition for the gate package not exploding — the parity package only has to prove "the mapping changed, the math did not", and that is the backbone of the bitwise-capable argument.
  • per-warp epilogue, 8/16 butterfly. h4w has 4 warps co-writing one q head's partial (cross-warp LSE merge in the block epilogue, see excerpt C below). In the batched shape each warp's (mx, S) is rescaled online to its final state within the window loop — warp-local, no cross-warp merge needed; the only thing left to fold is each warp's 4 row-class oc copies (oc0..oc3, one 16-dim slice per 8-lane subgroup), summed with an intra-warp 8/16 butterfly. That is one fewer sync step than h4w's epilogue.
  • Static grid, dispatch slots untouched. D3-4's self-gating is kept: grid/block contain no nkv (replay-safe), and the kernel self-gates internally on rpw >= H4W_MIN_RPW. The batched body and the 4-warp hybrid occupy the same rpw≥16 dispatch slot (active for gqa 2..8); at gqa=1 the batched shape degenerates to "one block, one warp" — zero benefit — and gqa>8 pushes the block into the other occupancy class at 288+ threads — both stay on the 4-warp hybrid; the incumbent 1-warp body for rpw<16 stays bitwise.
  • Shapes not chosen: smem-tile cooperative staging (put KV into shared memory and hand it out to warps) was ruled out at the time — D2's cp.async series has a record of negative results at exactly this granularity. The batched shape's selling point is precisely "move only the mapping, not the math, not the staging": if even that does not gain, the bytes side is hopeless.

3.2 Key code

⚠️ Code survival note: the GQA-batched kernel never landed as a commit — it lived as a patch in /tmp/d3/patch_2a.py, was reverted after measurement, and /tmp has since been wiped. Excerpts (A)(B)(C)(D)(F)(G) below all come from the current tree and show the h4w structure that the batched shape reused verbatim / inherited (window loop = the original text that was extracted into h4w_warp_windows; block epilogue = the contrast to the step the batched shape eliminated; self-gating = the dispatch slot verbatim); the batched body itself is narrated as a reconstruction from (E)'s test-table comments and the record. (H) is the gate that survived the revert — GQA geometry now permanently covers the live h4w kernel.

Excerpt A · the h4w body's thread mapping and stripe partition (current tree, attn_split_h4w_body) — the batched shape keeps the lane/w/t/g semantics and changes only h = blockIdx.y (q head) to h = hk*gqa + w, and the block thread count to 32*gqa:

// src/cuda_kernels.cu (current tree; the h4w body landed in D3a, promoted in D3-4)
__device__ __forceinline__ void attn_split_h4w_body(
    const float* __restrict__ q, const __half* __restrict__ k,
    const __half* __restrict__ v, float* __restrict__ partial,
    int nkv, int sp, int h, int nh, int nk, int hd, float scale, int pstr
) {
    const int gqa = nh / nk;
    const int hk = h / gqa;                    // ← the batched shape moves this from "computed per warp"
    const size_t stride_kv = (size_t)nk * hd;  //   up to the block boundary: warp w IS q head hk*gqa+w
    const int lane = threadIdx.x & 31;
    const int w = threadIdx.x >> 5;  // warp id
    const int t = lane & 7;          // 16-dim slice (hd=128 -> 8 slices)
    const int g = lane >> 3;         // row slot within a 4-row pass
    // Same device-side range split as the 1-warp kernel (identical [sp] rows,
    // so the combine sees the same split partitioning).
    const int chunk = (nkv + ATTN_SPLITS - 1) / ATTN_SPLITS;
    const int lo = sp * chunk;
    const int hi = min(nkv, lo + chunk);
    // Balanced contiguous stripes: warp w owns rows [lo + w*rpw, +rpw).
    const int rpw = (chunk + 3) >> 2;
    const int wlo = lo + w * rpw;
    const int wend = min(hi, wlo + rpw);

Excerpt B · the window-loop core (current tree, the part the batched shape reuses verbatim) — each warp advances its stripe in 32-row windows: 8-lane subgroups produce 4 rows of QK^T, in-window butterfly max, one rescale, probs staged through smem to feed the V accumulation. Note the K row address carries hk * hd — the batched shape is precisely 5 warps concurrently issuing loads to this same address, which is what L1 absorbs:

    // (window loop; probs is a per-warp 32-float stage: __shared__ float probs[H4W_NTHREADS]; float* pw = probs + w * 32;)
    for (int b = wlo; b < wend; b += 32) {
        const int wl = min(32, wend - b); // rows in this window (warp-uniform)
        float kq = -INFINITY;             // this lane's row score
        float mx_new = mx;
        const int np = min(8, wl);
        #pragma unroll
        for (int p = 0; p < 8; p++) {
            if (p >= np) break;
            const int row = b + g * 8 + p;
            float d = 0.0f;
            if (row < wend) {
                const __half* krow = k + row * stride_kv + hk * hd + 16 * t;
                const uint4 ka = *reinterpret_cast<const uint4*>(krow);   // 8 halves/uint4
                const uint4 kb = *reinterpret_cast<const uint4*>(krow + 8);
                d = h4w_dot8(ka, qc0, qc1) + h4w_dot8(kb, qc2, qc3);
            }
            // The subgroup reduce runs for ALL lanes: a shfl_sync with the full
            // mask deadlocks when subgroups diverge on row validity (found by
            // the standalone probe at nkv=3), so the guard selects the score
            // AFTER the reduction.
            float s = h4w_subgroup_sum8(d) * scale;
            if (row >= wend) s = -INFINITY;
            mx_new = fmaxf(mx_new, s);
            if (t == p) kq = s; // lane (g,t) keeps subgroup g's row (g,p)
        }
        #pragma unroll
        for (int off = 8; off < 32; off <<= 1)   // cross-subgroup window max
            mx_new = fmaxf(mx_new, __shfl_xor_sync(0xFFFFFFFFu, mx_new, off));
        const float wsc = expf(mx - mx_new);     // one rescale per 32-row window
        mx = mx_new;
        kq = expf(kq - mx);
        S = S * wsc + kq;
        /* oc0..oc3 scale the same way *= wsc; __syncwarp(); pw[lane] = kq; __syncwarp();
           V accumulation: subgroup g takes row b+4p+g's prob × V row, oc copies accumulate */
    }

Excerpt C · h4w's block epilogue (current tree) — the contrast to the step the batched shape eliminates. h4w's 4 warps co-write one q head, so it must merge LSE across warps: 4 copies of (mx, S) go into smem, gmax is computed, each warp rescales its own oc, then a stride-128 walk sums and writes the partial. In the batched shape this entire passage disappears — each warp's own (mx, S) is already final, and the partial slots are written per q head:

    // ── block epilogue: LSE-merge the 4 warp states, write ONE partial ──
    // (w,g) copy of dim d lands at vkq_s[w*512 + g*128 + d]: lane (g,t) stores
    // dims [16t,+16) at offset w*512 + g*128 + 16t, so the final per-dim sum
    // is a stride-128 walk — bank-conflict-free across threads.
    __shared__ __align__(16) float vkq_s[4 * 512];
    __shared__ float mx_sh[4];
    __shared__ float s_sh[4];
    float Sw = warp_reduce_sum(S);
    if (lane == 0) { mx_sh[w] = mx; s_sh[w] = Sw; }
    __syncthreads();
    const float gmax = fmaxf(fmaxf(mx_sh[0], mx_sh[1]), fmaxf(mx_sh[2], mx_sh[3]));
    const float wsc = expf(mx - gmax); // idle warp: exp(-1e38 - gmax) == 0
    /* oc0..oc3 scale the same way *= wsc; the four float4 copies are written into vkq_s; __syncthreads();
       if (tid < hd): stride-128 walk accumulating the four copies; dst[0]/dst[1] are written by tid==0,
       merging the LSE pair with s_sh[w]*expf(mx_sh[w]-gmax) */

Excerpt D · the self-gating hybrid kernel (current tree, the dispatch slot the batched shape inherits) — grid/block contain no nkv at all (positions[0] is read device-side), so every replay re-gates itself; the batched body occupies this same rpw >= H4W_MIN_RPW slot:

// src/cuda_kernels.cu (current tree)
#define H4W_NTHREADS 128 // 4 warps per block
#define H4W_MIN_RPW 16   // 4-warp body only when rows/warp amortize the window

__global__ void __launch_bounds__(H4W_NTHREADS, 8)
gqa_attn_split_partial_hybrid(
    const float* __restrict__ q, const __half* __restrict__ k,
    const __half* __restrict__ v, float* __restrict__ partial,
    const int* positions, int nh, int nk, int hd, float scale, int pstr
) {
    const int nkv = positions[0] + 1;
    const int chunk = (nkv + ATTN_SPLITS - 1) / ATTN_SPLITS;
    if (((chunk + 3) >> 2) < H4W_MIN_RPW) return;   // rpw < 16 -> incumbent body owns
    attn_split_h4w_body(q, k, v, partial, nkv, blockIdx.x, blockIdx.y,
                        nh, nk, hd, scale, pstr);
}

Excerpt E · the batched body itself (reconstructed narration) — the patch's shape, reconstructed from the record and the test comments: __global__ void gqa_attn_split_partial_gqa(...), grid = (ATTN_SPLITS, n_kv_heads), block = 32*gqa threads; warp w serves q head hk*gqa + w (hk = blockIdx.y); each warp calls the extracted h4w_warp_windows (excerpt B's loop, with the K/V base hk*hd shared by the five warps); the epilogue is per-warp — warp-local (mx, S) writes its partial slot directly, and the 4 row-class oc copies are folded in-warp with the 8/16 butterfly. The rpw self-gating condition matches excerpt D; enabled for gqa 2..8, falling back to the 4-warp hybrid at gqa=1 and >8.

Excerpt F · Rust-side dispatch (current tree src/cuda.rs) — the fixed-grid partial layout is a contract shared by both kernels (idle splits write mx=-INF/S=0, which the combine weights to zero):

#![allow(unused)]
fn main() {
// src/cuda.rs (current tree, excerpted comments)
/// ... in cuda_kernels.cu (fixed grid — the graph-replay capture depends on
/// it; idle splits write an mx=-INF/S=0 partial the combine weights to zero).
pub fn gqa_attn_split(&self, q, k, v, o, positions, nh, nk, hd, scale, f16_kv) {
    let pstr = ((4 + hd + 3) & !3) as i32;
    const ATTN_SPLITS: usize = 32; // mirrors #define ATTN_SPLITS in cuda_kernels.cu
    let need = ATTN_SPLITS * nh * (pstr as usize) * 4;
    let partial = Self::get_or_grow(&self.buf_attn_partial, need);
    // f16_kv ? launch_gqa_attn_split_f16kv(..) : launch_gqa_attn_split_f32kv(..)
    //   -- both launchers enter the hybrid/batched body when rpw>=16
}
}

Excerpt G · the dual-kernel launcher (current tree, the dispatch structure the batched shape was inserted into) — the shape of D3-4's self-gating dispatch at the launcher layer: two static launches + combine; the batched shape back then took the hybrid's place in the rpw>=16 slot (gqa 2..8) and inherited the "dud launch" mechanism unchanged:

// src/cuda_kernels.cu (current tree, f16kv launcher; excerpted comments)
// D3-4 L1: hd == 128 (Qwen2.5/Qwen3 decode shapes) dual-kernel
// self-gating dispatch on rpw = ceil(ceil(nkv/ATTN_SPLITS)/4):
// rpw >= H4W_MIN_RPW (nkv >= 1921) -> the 4-warp fattn-vec-style kernel;
// rpw < 16 -> the incumbent 32-thread D2-staged kernel. Both launches are
// static (grid/block nkv-independent), so CUDA-graph capture/replay is
// unaffected; each kernel re-reads positions[0] on every replay and
// exactly one is live for the current nkv (the rpw branch is
// nkv-uniform). The dud launch costs ~1-2 us/layer but keeps the
// small-rpw shapes on the incumbent geometry — running the 1-warp body
// inside 128-thread blocks caps the SM at 12 working warps (1536/128)
// and measured +78% kernel at 7B @1641 (35.4 vs 19.8 us, nsys).
if (hd == 128) {
    gqa_attn_split_partial<__half><<<dim3(ATTN_SPLITS, n_head), 32, 0, stream>>>(/*...*/);
    gqa_attn_split_partial_hybrid<<<dim3(ATTN_SPLITS, n_head), H4W_NTHREADS, 0, stream>>>(/*...*/);
}
gqa_attn_split_combine<<<dim3(1, n_head), hd, 0, stream>>>(partial, o, n_head, hd, pstr);

Excerpt H · the gate that survived the revert (current tree src/graph/cuda_backend.rs) — the two GQA geometries added to the parity test table for 2a were kept intact. Note that the comments still speak in pre-revert language ("drives the GQA-batched kernel"); after the revert these shapes actually cover the live h4w hybrid in the nkv≥1921 range — in the master table's words, "these shapes now permanently cover the LIVE h4w kernel too":

#![allow(unused)]
fn main() {
// src/graph/cuda_backend.rs (current tree, split-decode parity test-table entries)
// D3-6 2a: the 14B GQA geometry (40:8, gqa=5) drives the
// GQA-batched kernel (grid (ATTN_SPLITS, 8), 160 threads) on the
// same pos0 sweep — the 1920/1921 boundary picks between the
// bitwise 1-warp incumbent (nkv 1920) and the batched body
// (nkv 1921), and 4094/4095 cover full-window chunk tails.
(40usize, 8usize, 128usize, 4200usize,
 [2usize, 32, 63, 64, 127, 128, 1023, 1919, 1920, 4094, 4095]),
// D3-6 2a: the 7B GQA geometry (28:4, gqa=7 → 224-thread blocks).
// nkv 2808 (pos0 2807) reproduces the in-situ decode-step shape at
// the divergence point seen in the 7B greedy gate.
(28usize, 4usize, 128usize, 4200usize,
 [2usize, 32, 63, 64, 127, 128, 1919, 1920, 2807, 4094, 4095]),
}

(The pos0 table is nkv=pos0+1: a full sweep of 3..4096; 1920/1921 is exactly the rpw 15/16 dispatch boundary; 2808 is the in-situ decode-step shape where the 7B greedy gate hit its knife-edge.)

3.3 Pitfalls

  • The full-mask shfl_sync deadlock lesson directly constrained the extraction. That comment in excerpt B ("a shfl_sync with the full mask deadlocks when subgroups diverge on row validity — found by the standalone probe at nkv=3") was bought with a deadlock back in the D3a era: the subgroup reduction must have ALL lanes participate, and row validity is selected with -INFINITY after the reduction. The window-loop extraction into h4w_warp_windows must not break this — in the batched shape different warps can have different stripe lengths, so the divergence surface is larger than in the original shape.
  • The occupancy structure moved, and it moved a lot. h4w is __launch_bounds__(128, 8) (4 warps/block, targeting 8 blocks/SM); the batched shape is 160 (gqa=5) / 224 (gqa=7) threads with blocks 40→8. The organization of resident warps on the SM is completely different — this is exactly the experimental face of the "wave-structure bound" hypothesis, and also one natural explanation of the +1.6% reading: the L2 wait saved by going 1-pass was eaten back by coarser-grained block scheduling.
  • The epilogue simplification was free — and bought nothing. h4w's block epilogue must merge LSE across warps (excerpt C's mx_sh/s_sh + stride-128 walk); the batched shape's per-warp LSE is already final, so the epilogue is only the in-warp 8/16 butterfly — one fewer __syncthreads round, with no live-time gain, again pointing to the bottleneck not being in the epilogue.
  • The batched shape inherited the dud-launch tax. The dead launch in the dual-kernel dispatch costs ~1–2 µs/layer (excerpt G's comment, verbatim) — the price D3-4 paid to keep small-rpw shapes on the incumbent geometry. The batched shape entered through the same slot mechanism and did not touch that ledger; it is the dud split ~0.070 ms/step item in the D3-7 census (48 layers × ~1.5 µs).
  • The partial-layout contract must not break. The partial slot layout ([sp][head] × pstr) and the idle-split semantics (mx=-INF/S=0) are the combine kernel's input contract; the batched shape changed the meaning of the head dimension (the 40 q-head slots stay as they are), and the combine needs zero changes — this is the structural precondition for "moving only the mapping" to hold.

4. Verification

All gates ran before the revert decision — this revert was a mechanism veto, not a correctness veto:

  • Kernel-level parity (defends against kernel math errors): the split-decode parity test was extended with the 14B (40:8, gqa=5) and 7B (28:4, gqa=7) geometries, a full nkv 3..4096 sweep (including the in-situ 2808 divergence-point shape and the rpw 15/16 boundary), plus the D3a outlier calibration (the f16-representable residual class with |q|~57 and V ±144) — all ≤1e-4 green against the CPU reference. The outlier calibration test itself (current tree) — pushing the activation dynamic range to the f16-representable ceiling of |q|~57 / |v|~144, the batched body is still ≤1e-4 vs the CPU reference:
#![allow(unused)]
fn main() {
// src/graph/cuda_backend.rs (current tree, test excerpt)
// D3-6 2a: kernel-level gate of the calibrated tolerance package on
// realistic outlier-scale data (docs/CUDA_OPTIMIZATION.md §2D D3a:
// residual |q|~50, V outliers ±127 — the h4w body measured 6.5e-5
// vs CPU on this class, the incumbent 3.8e-5). The GQA-batched body
// shares the h4w window loop verbatim, so the same ≤1e-4 bound
// applies; 14B geometry (40:8, gqa=5) inside the batched regime
// (nkv 1921 / 4096, f16 KV).
let qs: Vec<f32> = (0..nh * hd)
    .map(|i| (((i * 37) % 19) as f32 / 5.0 - 1.9) * 30.0)  // |q| ~57
    .collect();
let vs: Vec<f32> = (0..nkt)
    .map(|i| (((i * 57) % 11) as f32 / 3.0 - 1.8) * 80.0)  // |v| ~144
    .collect();
assert_close(&format!("attn_split_gqa_batched_outlier(nkv={nkv})"),
             &agot, &aref, 1e-4);
}

These shapes therefore entered the gate set permanently: whether or not the batched shape lives, the live h4w kernel is from now on covered by real GQA geometry.

  • Dump gate (defends against end-to-end numeric drift): both models, a 2.8K-token prompt (fully inside the batched regime), −n 4: logits_prefill/kv0_prefill/early-layer KV bitwise (14B kv0–15, 7B kv0–4); logits_decode max|Δ| 0.265 (14B) / 0.259 (7B) — the calibrated 0.39-class; argmax HARD gate green, margins 2.187 / 0.557 (top-2 is far away, so a flip can only be a mechanistic error). Upper-layer KV decode-side drift = the documented f16-boundary class; the node{2,3,5,8,11}_decode and node{3,5,8}_prefill pool-slot aliasing diffs reproduce pre-vs-pre (a same-binary self-diff covers all of the pre-vs-post diff — no numeric delta).
  • Greedy + the sampler-vs-kernel attribution gate newly established by this step (defends against misreading a sampler knife-edge as kernel drift) — this step's durable methodological contribution:
    • Under --repeat-penalty 1.0 (bare argmax), the −n 256 streams are byte-identical on both models — across 512 steps the kernel never flipped a single bare argmax. This is a strong gate on kernel numerics: an argmax stream with no sampling machinery in the way is maximally sensitive to kernel numeric noise.
    • Under the default repeat penalty, each run had exactly one knife-edge flip (1/256 = 0.4%, far below the 2% line): 14B at step 63, where the bare top-2 probgap is only 0.0167 (logit gap 0.037) and the penalty swapped the order; 7B at step 8, where the post-penalty winner is bare rank 6+ (outside the traced top-5). After the flip the text stays coherent with no repetition degeneracy; the flip's next step is visible in MINFER_TRACE as a changed embed node input, while steps 0..k−1 align one by one with the baseline's top-5 down to the drift class — the cascade structure is self-consistent.
    • The temp-0.8 seed-7 sampled control is byte-identical on both models (the incumbent regime is symmetric on both sides, so sampling reordering cannot happen).
  • Suite 170/0/3 (including the FA trio and the extended split-decode parity).

5. Results (REVERTED + veto mechanism)

The window anchors are the same as D3-5 (14B tg128 23.35 / @3254 21.68; 7B tg128 50.69 / @1641 49.16; guards 7B ≥49.0 / ≥47.9, 14B tg128 ≥22.7). The batched shape does not move the wall clock, so all of these are the post-revert status quo. The combined two-instrument measurement (nsys: NO_CUDA_GRAPH, bench −n 8, mean of the last 384 steps; ncu: 2-capture serialized replay of the same launches):

kernellts__t_sectors.sumvs 1× analytic (458K)nsys (live)ncu (serialized)
h4w hybrid (5× re-read)2,121,6714.63×63.76 µs150.8 µs
GQA-batched (1× read)472,5831.03×64.79 µs108.7 µs
  • The traffic hypothesis was falsified: the batched shape cut L2 sectors to 1.03× of the analytic minimum (L1 caught the sibling warps' same-address loads), yet live time went +1.6% (the projection had been ≤54 µs). The 5× re-read was already completely hidden by the latency roofline in the live regime — L2 has bandwidth slack, and the kernel is limited by the latency/dependency chain + wave structure. D1's original attribution (latency-bound) stands; D3-4's "L2 composition" note is corrected.
  • Why +1.6% was enough to veto: this campaign's bar for a "win" is +1.5% (the landing bar, calibrated at D3-5); this lever's projected gain was −15% to −25% territory (63.76 → ≤54 µs). A measured reverse +1.6% means the mechanism direction is wrong as a whole — not "the gain was eaten by noise" but "the gain does not exist". When a prediction of this magnitude fails, the correct action is to revert and revise the attribution, not to tune and retry.
  • ncu's −28% is another regime: serialized replay, cold cache, no concurrent overlap — the traffic-dominated world. It proves the batched shape "really did cut traffic", and it proves that this fact is worthless in the live regime.
  • The re-judgment for Stage 3: the attention residual of 13.2 µs/layer ≈ 0.63 ms/step cannot be reached by any bytes-side traffic lever. The next attention lever must cut exposed latency: cross-window prefetch (while preserving occupancy — D3-4 L2's register-pipeline form lost 1:2 on the occupancy cost), or smem-tile cooperative staging (untried, but D2's cp.async negative results make this granularity historically disfavored). Retry conditions: a shape that can shorten the dependency-chain depth without reducing resident warps — otherwise the "elimination" of the re-read never converts into time.
  • The distance ledger at session close (wall clock untouched): −5.00 ms (−10.8%) unchanged; the known list becomes rms merge ≈1.0 ms + attn_v routing ≈0.28 ms (handed to D3-7) + output head 0.44 ms (mechanism missing) + attention latency lever ≤0.63 ms (mechanism in doubt) — 1.7–2.35 ms combined = 34–47% of the gap, the rest being matmul aggregation + launch-structure slack, forming Stage 3's decision input.

6. Lessons

  1. Counter evidence cannot overrule live time. Traffic ÷4.5 with time flat has only one self-consistent explanation: traffic is not on the critical path; a serialized profiler (ncu) measures the opposite regime of cold cache and no overlap, and its −28% is exactly what the "traffic-dominated world" looks like — only reading the two numbers together pins the mechanism.
  2. gate-green ≠ land. All correctness gates green only proves the kernel is correct; a performance-mechanism error still gets reverted. The revert decision can (and should) happen after the gates are all green — gates are admission conditions, not landing conditions.
  3. (durable gate methodology) greedy divergence ≠ kernel drift. The argmax under the default repeat penalty is a knife-edge among up to 64 penalized candidates, and a flip cascades on the next step (a different token enters embed). The clean kernel-numerics gate is byte-identical rp=1.0 greedy streams; attribute each flip in the default-penalty stream with a per-step logits_top trace to get the bare top-2 probgap — only flips with a large probgap are worth suspecting the kernel.
  4. Keep the test shapes of negative results. The gqa=5/7 geometries + the 2808 shape + the rpw boundary were kept intact after the revert, going from "the batched shape's gates" to "permanent coverage of the live h4w" — the next person touching attention will not have to rediscover these shapes.

← 70 · Index · 72 →

72 · D3-7 — attn_v-q6K MMVQ routing (2b) + rms wide-block / positions memo (2c): the Stage-2 closing ledger (LANDED ×2)

Result: 2b — the Q6_K decode dispatch gate od*id >= 24M lowered to >= 4M, taking 14B attn_v (od 1024 × id 5120 = 5.24M, 11/48 layers) off the padded-f32 kernel: 33.16 → 24.32 µs (−26.7%, ~177 GB/s), ×11 layers ≈ 97 µs/step; attn_v and attn_o share an MmqCache hit, so the standalone quantize count is unchanged. 2c — rms wide-block 32→128 threads (9.43 → 5.66 µs, −40%, −0.35 ms/step) + a positions_i32 per-execution-window memo (239.6 → 1.2 launches/step, −0.28 ms/step), both bitwise. Combined wall clock (SEP): 14B @3254 21.57 → 21.95 (+1.76%), tg128 +1.80%; 7B +1.04%/+1.07%; distance to parity −10.8% → −9.7%. Commit: 1088bc8 (2b, src/cuda.rs only, +11−1) + b3b6077 (2c, cuda_kernels.cu + graph/cuda_backend.rs). Date: 2026-09-08.

1. Background — where things stood

The D3-6 session left this closing window two legacies. The first is negative: the attention residual was judged to be exposed latency and every bytes-side lever ruled dead (doc 71) — the largest single item on the Stage-2 list (0.63 ms) moved from "winning-able" to "bookkeeping". The second is deferred: 2b (attn_v-q6K → MMVQ routing) went untouched last session for lack of time, but its mechanism estimate had long been sitting in D3-1's census — attn_v runs the padded-f32 kernel at only 134.8 GB/s, while its q6_K MMVQ siblings (ffn_up/down, attn_q/k/o) hold the 200–225 GB/s class; D3-1 estimated +0.28 ms/step.

So D3-7's (base /tmp/d3/minfer_pre_d37, sha1 b29f2ae6, HEAD 92b0712) positioning was very clear: settle the two "small and certain" items on the list, clearing a clean ledger for Stage 3's big move (the decode-GEMM program). Window anchors (pre_d37, 3× interleaved medians): 14B tg128 23.36 / @3254 21.65; 7B tg128 49.91 / @1641 48.31 (7B shows co-tenant drift relative to the D3-5 window, so everything is same-window interleaved A/B; sglang co-tenant present).

2c's working basis came directly from the decode launch census taken after D3-5 (14B @3254, decode-class launches):

kernelµs eachlaunches/stepms/step
rms_norm_quant_pad409.4394.60.892
f32_bits_to_i321.15239.60.276
add_bias——0.181
store_kv——0.179
attention combine——0.171
rope——0.139
add——0.139
swiglu——0.110
standalone quantize——0.084
dud split launch——0.070

Together ≈ a 2.2 ms/step ocean of sub-6µs kernels. The census also exposed one thing: the 14B decode qkv chain is not fused (per layer rope ×2 + store ×2 + bias ×3; FusedQKV is inactive on CUDA) — recorded as a front-row Stage-3 item ≈ 0.45 ms/step, which is doc 73's entry point. This session settles only the top two rows of the table (0.892 + 0.276); the rest is left to their respective larger levers.

2. Principle — the GPU mechanism

2b — why the crossover gate wrongly killed attn_v. The Q6_K/Q5_K decode dispatch gate from the 8e era is od*id >= 24M, based on the on-device micro-bench of the time (padded f32 vs mmvq): small shapes lose under MMVQ — od 512 → 4.5× slower, od 896 → 3.0×, 2048×4864 → 1.66× — because when each thread gets only 1–2 16-element units to amortize, the uncoalesced cost of the q5/q6 byte loads has nowhere to go; large shapes win under dp4a — 7B ffn_down 3584×18944 → 1.5× faster, lm_head 152064×3584 → 1.4×. But that crossover line was fitted through a handful of od points. Decode MMVQ's structure is one 256-thread block per row (the row's work split across npair 16-element units), so the longer the id, the more work per block and the flatter the amortization of the byte-load cost. attn_v is a medium-od, long-id shape (od 1024, id 5120 → npair 160 units/row → 1024 blocks × 256 threads) — exactly the class between the fitted points that the gate constant mis-killed. The measured −26.7% (~177 GB/s) confirms: the direction is right, but it still falls short of the 220-225 sibling class — consistent with the 8e data's small-shape warning at od≈1024, just not losing to padded-f32.

Why the standalone quantize count does not grow. D3-5 1a's MmqCache (producers record their q8 plane; later matmuls consult the cache, keyed by the producer's f32 output pointer) holds automatically under 2b's routing: attn_v's producer is the attention node (there is no fused quantize to speak of), but attn_v and attn_o consume the same attention output buffer and the same id — attn_v's decode_quantize_native arrives first (record), attn_o arrives later (hit). So 2b merely "moves one quantize that had to happen anyway earlier and shares it"; the standalone quantize launch count stays at 964 (trace-verified). The net gain is pure kernel-time difference: 33.16 → 24.32 µs × 11 layers.

2c(i) — why the 32-thread rms is latency-bound. rms_norm_quant_pad40 in its original form is one 32-thread warp per row. hidden 5120 → d4 = 1280 float4s, i.e. 40 serial float4 loads per lane, and no other warp on the SM to fill the load latency — pure latency-bound. Widening to 128 threads: 10 loads per lane + 4 concurrent warps, giving the load pipeline material to fill the holes. Is bitwise preserved? Yes, and the argument is structural: the reduction is still carried by lanes 0..31, the i += WARP stride keeps the original form's element→lane mapping and per-lane serial accumulation chain, and the warp_reduce_sum tree is untouched; scale is broadcast through smem to the whole block; the y write and the quantize epilogue are per-element / per-32-block independent, so the wider thread mapping cannot move any output bit. #pragma unroll 8 only deepens the load pipeline. The remaining 5.66 µs is launch (~1.3 µs) + the epilogue's q8 write + the w-row stream — this item is already near the structural floor for a kernel of this class.

2c(ii) — where the 240 redundant conversions come from. Each graph layer has rope ×2, KvcacheStore ×2, Attn ×1 — 5 node classes consuming the same positions I32 input (f32::from_bits bit patterns, filled by fill_input_i32), and the CUDA backend calls positions_i32(id) at every consuming node to convert it to native int32 on device (the rope/store/attention kernels read const int* positions) — 48 layers × 5 = 240 of the same conversion per step, 1.15 µs each, pure redundancy. The memo's semantic window is clean: within one graph execution, the same input buffer's bit patterns do not change (the host fills them before execution), so "one conversion per buffer per execution window" is mathematically identical. The key (buf id, pool_gen) handles free-list reuse: when the same id is returned and taken out again, pool_gen has changed and the cache invalidates naturally.

3. Implementation

3.1 Design choices (why this shape and not another)

  • 2b changes the gate constant; no special-case routing. The precondition was the GGUF census (census_q6k.py) first: sweep the q6_K weight shapes of all supported models and confirm that within the (4M, 24M) interval only 14B attn_v exists (7B attn_v is od 512 × id 3584 = 1.8M and stays on padded-f32, so the 7B stream is untouched byte for byte). With the census as backstop, 24_000_000 → 4_000_000 is the minimal change surface — no new dispatch branch, no new kernel, no new env var (the MINFER_NO_KQ_MMVQ=1 semantics are kept as-is).
  • 2c: census first, then act — and hit only the top two items. In the sub-6µs ocean, add_bias/rope/store_kv are on the list too, but they fall in the blast radius of qkv-chain fusion (D3-8) — this session settles only rms (0.892 ms) and bits_to_i32 (0.276 ms), avoiding two mechanisms muddying one measurement window.
  • Both sub-levers must be bitwise. 2b already introduced the first decode numeric freedom this campaign has tolerated (f32→q8 activation rounding, tolerance-gated per the D3a package); if 2c also changed numbers, the gate attribution would tangle. A bitwise 2c makes "measuring 2b in isolation" possible — wrap 2b on both sides with MINFER_NO_KQ_MMVQ=1, and every remaining diff belongs to 2c, whose diff must be zero.

3.2 Key code

Excerpt A · the Q6_K dispatch gate (current tree src/cuda.rs, the landing spot of 1088bc8) — the before/after is one gate-constant line: od * id >= 24_000_000 → >= 4_000_000 (that commit's entire change is this plus comments, +11−1):

#![allow(unused)]
fn main() {
// src/cuda.rs — TensorType::Q6_K decode arm (current tree)
// Shape gate: see the Q5_K arm comment (measured od*id
// crossover ~24M elements; below it the padded f32 kernel's
// coalesced loop wins, above it MMVQ's dp4a wins).
// D3-7 2b: gate lowered 24M -> 4M for the attn_v class —
// the 14B attn_v (od 1024 x id 5120 = 5.24M, 11 layers)
// sat on the padded-f32 kernel at 134.8 GB/s (D3-1 census)
// while its q6_K MMVQ siblings sustain the 200-225 GB/s
// class; attn_v/attn_o share the attention-output buffer
// and id, so the D3-5 MmqCache dedupes their standalone
// quantize to one launch. GGUF census: no other q6_K shape
// falls in (4M, 24M) (7B attn_v 1.8M stays padded-f32).
// Tolerance-gated (f32->q8 activation rounding): D3a
// package; MINFER_NO_KQ_MMVQ=1 keeps the padded kernel.
if nt == 1 && id % 32 == 0 && od * id >= 4_000_000 && !Self::no_kq_mmvq() {
    self.q6_k_decode_mmvq(wptr, x, out, od, id, nt, padded_q6k);
    Ok(())
} else if padded_q6k {
    launch!(launch_q6_k_f32_matmul_padded)
} else {
    launch!(launch_q6_k_f32_matmul)
}
}

Once dispatched into q6_k_decode_mmvq, the shapes split naturally: id 5120 → npair 160 → the plain v2 kernel, 1024 blocks × 256 threads — no new kernel form was added for attn_v.

Excerpt A2 · the MmqCache consult entry (current tree src/cuda.rs, head of q6_k_decode_mmvq) — the "attn_v quantize moved earlier and shared" that 2b depends on happens at the entry of every decode MMVQ; in the current tree this spot also carries the post-D4-4 dpl fast path (later than this step, see doc 76), so the excerpt takes only the D3-5-era consult semantics:

#![allow(unused)]
fn main() {
// src/cuda.rs (current tree)
pub fn q6_k_decode_mmvq(&self, wptr, x, out, od, id, nt, blk_stride_padded) {
    // D3-5 1a: consult the MmqCache first — the fused decode producers
    // (rms_norm_quant_on_gpu / swiglu_quant_off_on_gpu) recorded their
    // pad40 plane, so the matmul group following the producer skips the
    // standalone quantize launch entirely (MINFER_NO_DECODE_A_FUSE=1
    // restores the unconditional standalone launch).
    let q8 = self.decode_quantize_native(x as *const f32, id, nt);
    /* ... the D4-4 dpl fast path and padded/v2 dispatch (later than this step) ... */
}
}

attn_v's decode_quantize_native records here (its producer is the attention node, so the cache misses → quantize on the spot and record), and attn_o's identical call then hits — 2b's "quantize count unchanged" is this one line of code serving two consumers.

Excerpt B · the Q5_K arm — the code evidence for 2c's gate trap (current tree src/cuda.rs): no_kq_mmvq() does not control Q6_K only — the Q5_K decode arm reads it too (8e-era behavior). This is the mechanical root of "a one-sided env control fabricates drift":

#![allow(unused)]
fn main() {
// src/cuda.rs — TensorType::Q5_K decode arm (current tree; landed in 8e)
// Shape gate measured on-device (dbg micro-bench, padded f32
// vs mmvq): od*id < ~24M elements loses (od 512 → 4.5x
// slower, 896 → 3.0x, 2048x4864 → 1.66x) ...
// MINFER_NO_KQ_MMVQ=1 forces f32.
if nt == 1 && od * id >= 24_000_000 && !Self::no_kq_mmvq() {
    self.q5_k_decode_mmvq(wptr, x, out, od, id, nt);
    Ok(())
} else {
    launch!(launch_q5_k_f32_matmul)
}
}

Excerpt C · rms wide-block (current tree src/cuda_kernels.cu, the kernel side of b3b6077) — the reduction is lane-for-lane equivalent; only the block is widened from 32 to 128 (launcher: rms_norm_quant_pad40<<<n, 128, 0, stream>>>):

// src/cuda_kernels.cu (current tree)
__global__ void __launch_bounds__(128) rms_norm_quant_pad40(
    const float* __restrict__ x, const float* __restrict__ w,
    float* __restrict__ y, uint8_t* __restrict__ q8,
    int d, float eps, int n
) {
    int row = blockIdx.x;
    if (row >= n) return;
    int tid = threadIdx.x;
    int d4 = d / 4;
    // D3-7 2c: wide-block geometry (launch picks 32 or 128 threads). The
    // reduction is bitwise-preserved: lanes 0..31 keep the exact 32-thread
    // form's element->lane mapping, serial per-lane accumulation order and
    // warp_reduce_sum tree; the unroll only deepens load pipelining. scale
    // reaches the whole block through shared memory. The write and quantize
    // loops are per-element / per-32-block independent, so their wider
    // thread mapping cannot change any output bit.
    __shared__ float s_scale;
    float scale;
    if (tid < WARP) {
        const float4* x4 = reinterpret_cast<const float4*>(x + row * d);
        float ss = 0.0f;
        #pragma unroll 8
        for (int i = tid; i < d4; i += WARP) {   // serial chain in the original form's order
            float4 v = x4[i];
            ss += v.x * v.x + v.y * v.y + v.z * v.z + v.w * v.w;
        }
        ss = warp_reduce_sum(ss);
        scale = rsqrtf(ss / (float)d + eps);
        if (tid == 0) s_scale = scale;
    }
    __syncthreads();
    scale = s_scale;                              // broadcast to the 96 non-reducing threads
    /* write y: for (i = tid; i < d4; i += blockDim.x) — 128-way independent float4 writes
       quantize epilogue: for (b = tid; b < nb; b += blockDim.x)
                        quantize_pad40_block(src + b*32, dst + b*Q8PB); */
}

The Rust-side entry rms_norm_quant_on_gpu (src/cuda.rs) calls record_mmq_cache_native(y, n, d, q8) right after the launch — D3-5's MmqCache record point, and the foundation of 2b's "attn_v quantize moved earlier and shared" chain:

#![allow(unused)]
fn main() {
// src/cuda.rs (current tree)
pub fn rms_norm_quant_on_gpu(&self, x, w, y, d, n, eps) {
    let q8 = Self::get_or_grow(&self.buf_q8_decode, n * (d / 32) * 40);
    unsafe { launch_rms_norm_quant_pad40(x, w, y, q8, d as i32, eps, n as i32, stream); }
    self.record_mmq_cache_native(y as usize, n, d, q8 as usize);
}
}

Excerpt D · the positions_i32 memo (current tree src/graph/cuda_backend.rs) — key (id, pool_gen), one conversion per execution window:

#![allow(unused)]
fn main() {
// src/graph/cuda_backend.rs (current tree)
fn positions_i32(&mut self, id: usize) -> Result<*mut std::ffi::c_void, String> {
    // D3-7 2c: one conversion per execution window per input buffer.
    // Capture mode: the first consumer's launch is recorded at capture
    // time and replay re-executes it every step (memo hits are never
    // recorded). Non-capture mode: synchronize() clears the memo at the
    // execution boundary, so each step re-converts exactly once.
    if self.pos_memo == Some((id, self.pool_gen)) {
        return Ok(self.pos_scratch);
    }
    let src = self.ptr_of(id)?;
    let bytes = self.pool[id].bytes;
    /* scratch growth path: after a fresh cuda_malloc, pool_gen += 1,
       which invalidates and re-captures any captured graph exec that may
       still embed the old scratch pointer */
    self.state.bits_to_i32(src, self.pos_scratch, bytes / 4);
    self.pos_memo = Some((id, self.pool_gen));
    Ok(self.pos_scratch)
}
}

Capture safety is a key design clause, with semantics for both modes: during CUDA Graph capture only the first consuming node's conversion launch gets recorded into the graph — memo hits never emit, and replay re-executes that one conversion each step; in non-capture mode synchronize() clears the memo at the execution boundary (pos_memo = None, right next to r49's MmqCache clearing point), guaranteeing the boundary of "one execution window". Sharing the clearing point also means the MmqCache and the pos memo always invalidate in lockstep — the half-state "cache alive, conversion not run" cannot occur.

Excerpt D2 · the kernel the memo eliminated, in full (current tree src/cuda_kernels.cu) — f32_bits_to_i32 is 8 lines in total; the 0.276 ms/step buys its emission and scheduling × 240:

// src/cuda_kernels.cu (current tree)
__global__ void f32_bits_to_i32(
    const float* __restrict__ src,
    int* __restrict__ dst,
    int n
) {
    int tid = blockIdx.x * blockDim.x + threadIdx.x;
    if (tid >= n) return;
    dst[tid] = __float_as_int(src[tid]);
}

A kernel of this class spends its wall clock almost entirely on launch tax and the memory round trip — which is why the "once per execution window" memo converts 239.6 → 1.2 directly into time.

3.3 Pitfalls

  • The one-sided MINFER_NO_KQ_MMVQ=1 control is a fake-drift factory. The first 2c-only control set this env on the post side only, and the dump gate immediately showed a 0.22-logit "drift" (prefill+decode appearing together, depth-independent — not a propagating per-layer numeric difference but the whole path swapped). Mechanism: the env also reverts the Q5_K decode arm (excerpt B, pre-existing 8e behavior), so one side runs Q5_K through MMVQ and the other through f32 activations — the fake drift has nothing to do with 2b. The rule was therefore finalized: control envs must be set on both sides. From then on the 2c dump gate was double-sidedly wrapped throughout.
  • Co-tenant pollution cost a window. 2b's tg128 extended measurement window was polluted on the post side by co-tenant activity; the whole window was discarded rather than forced into shape, and the landing evidence rests on the clean @3254 8-pair window (21.605 → 21.72, 7/7 pairs all positive, sign-test p≈0.008). The strict SEP differed by 0.05% (excluding one co-tenant rep: min-new-excl-outlier 21.66 vs max-base 21.67); per the D3-5 precedent it was carried on the median basis.
  • MmqCache depends on consumption order. attn_v's producer is the attention node (not a fusable producer like rms/swiglu), so the dedup holds only if attn_v's quantize executes before attn_o — a same-buffer + same-id ordered consumption pair. The graph build order guarantees this order; if attn_o arrived first, the record/hit roles would swap with the conclusion unchanged; if the two ids differed, 2b would push the standalone quantize count above 964 — the trace-verified 964-unchanged is part of the gate.
  • The census counts are decode-class averages. The 94.6/239.6 in the table are non-integers — they are the means over a 16-step trace (diluted by boundary effects like graph rebuilds), not per-step constants. Read the census mean-to-mean; do not check it against single-step integers.

4. Verification

  • 2b (the D3a tolerance package, defending against f32→q8 activation rounding drift):
    • dump gate: logits_decode max|Δ| 0.254 (inside the calibrated 0.39-class), logits_prefill byte-identical; argmax HARD gate green (margin 1.915); kv0 bitwise, kv1+ showing decode-side f16-boundary drift — appearing earlier than D3-6's sub-ULP reorder class, as expected: input-quantization noise is ~4e-3 relative vs the reorder's 1e-7, so of course the former hits the f16 boundary first. The per-layer reading: layer-0's KV is written by prefill activations (this step does not touch prefill numerics) so it is bitwise; from attn_v rounding onward the decode-step hidden states carry a ~4e-3-class relative difference that amplifies with depth — the kv1+ drift gradient is exactly the shape of this propagation chain. The 0.39-class headroom from the D3a calibration package (h4w measured 6.5e-5 and incumbent 3.8e-5 on the outlier class with residual |q|~50 and V ±127) exists precisely for this class of input-quantization noise.
    • the greedy gate uses D3-6's newly established attribution method: rp=1.0 (penalty waived) greedy −n 256 byte-identical on both models (the clean kernel-numerics gate); the default-penalty stream = exactly one knife-edge event per 256 steps (5/5 seeds, coherent after the flip — in the mid-run repetition of "near the" → "as the", the penalized argmax was sitting among the repetition candidates anyway).
    • temp-0.8 control: 7B byte-identical, 14B reordered — top-p reordering under a 0.22-logit drift is expected and not a gate failure.
  • 2c (bitwise end-to-end, defending against "bitwise by construction" becoming a slogan): dump gate with MINFER_NO_KQ_MMVQ=1 set on both sides (after isolating 2b, 2c's diff must be zero): 109/114 files byte-identical, the 5 diffs = the documented pool-slot aliasing class; logits in both phases + all KV byte-identical; 7B greedy 5/5 seeds + temp-0.8 control byte-identical.
  • Suite 170/0/3 (the combined 2b+2c window; the full log was in /tmp/d3/suite_2b2c_full.log at the time and has since vanished with /tmp).

5. Results

2b (nsys, 14B @3254, same window): attn_v kernel 33.16 → 24.32 µs (−26.7%, ~177 GB/s weight stream — short of the 220-225 sibling class, consistent with the 8e small-shape crossover data's warning at od≈1024); ×11 layers ≈ 97 µs/step; standalone quantize launch count unchanged (964). Wall clock: @3254 21.605 → 21.72 (+0.42%, 8-pair median, 7/7 clean pairs positive, sign-test p≈0.008); tg128 clean window +0.26% (below the +0.3% session landing bar; the extended window was discarded as co-tenant-polluted) — landed on the @3254 evidence + the measured mechanism.

2c (nsys, reconciled against the same-window census): rms_norm_quant_pad40 9.43 → 5.66 µs (−40%; 94.6–96 launches/step); f32_bits_to_i32 239.6 → 1.2 launches/step; wall-clock effect ≈ −0.62 ms/step (rms −0.348 + bits −0.275).

Cumulative (2b+2c vs pre_d37, interleaved 3× medians, SEP strict):

configpre_d37postΔ
14B tg12823.3123.73+1.80%
14B @325421.5721.95+1.76%
7B tg12849.9550.47+1.04%
7B @164148.4748.99+1.07%

All guards hold (14B tg128 ≥ 22.7; 7B ≥ 49.0 / ≥ 47.9). Against the window anchors: 14B tg128 +1.59%, @3254 +1.39%. 2c's standalone contribution ≈ +1.4% (cumulative minus 2b's kernel projection of +0.21–0.3%), matching the census projection — where the census aimed and how much it hit, both ends reconcile.

Distance to parity (post-D3-7, 14B @3254): 21.95 t/s = 45.56 ms/step vs llama 41.12 ms → −4.44 ms (−9.7%) (after D3-6 it was −5.00 ms / −10.8%). The Stage-3 front-row list (ordered by size):

  1. Matmul aggregation (~2.9 ms, the largest block): D3-1's wall-effective 194.9 GB/s vs llama's implied ~207.6 — the decode-GEMM program (q8_1-prologue fusion / llama-class MMVQ+GEMM rework) is the only lever class that reaches it.
  2. Launch-structure residual (~0.7–1.0 ms): led by the unfused qkv chain ≈ 0.45 ms/step (rope ×2 + store ×2 + bias ×3 per layer — a CUDA attn_bias_rope_store is needed), followed by the add_bias/add/store_kv/swiglu small-kernel tail and the dud split launch (~0.07 ms).
  3. Exposed-latency attention (≤0.63 ms, mechanism in doubt): the bytes side is ruled dead (D3-6); smem-tile cooperative staging is the untried remainder, but D2's cp.async negative results make this granularity historically disfavored.
  4. Mechanism-missing tail: output head 200.1 GB/s (0.44 ms, no geometry knob at D3-5 1b), ffn_down-q6K 198.9 GB/s (0.49 ms).

7B decode stays closed out (tg128 1.021× vs llama, @1641 0.992× — ahead/even). The rest of the sub-6µs ocean (add_bias/store_kv/rope/add/swiglu ≈ 0.75 ms/step) was untouched by this step — it belongs to the qkv-chain fusion (doc 73) and larger levers; this census table is kept as Stage 3's comparison baseline.

6. Lessons

  1. census-first, census-close. 2c's targets (0.892 + 0.276 ms) and its acceptance (239.6 → 1.2, 9.43 → 5.66) come from the same launch×µs table — the data decides where to strike, and the same table adjudicates whether the strike landed, with no second basis of measurement introduced.
  2. Read the side-effect surface of a control env in the code before designing the control. MINFER_NO_KQ_MMVQ=1 is nominally a Q6_K gate but actually also flips the Q5_K arm — a one-sided setting manufactured a 0.22-logit fake drift. Rule: switch-type controls are set on both sides, and prefer isolating variables with the double-sided diff's "must be zero".
  3. A crossover gate constant is not a universal constant. 24M was fitted through a handful of od points and wrongly killed the "medium-od, long-id" attn_v. Before changing a gate: run the shape census first (a full-model sweep found only this one shape in (4M,24M)), then let measurement adjudicate — analytic boundary + full sweep + small landing step; missing any one of the three invites a crash.
  4. Settle steps that add tolerance freedom separately from bitwise steps. 2b (tolerance-gated) and 2c (bitwise) landed in the same window, but the gates were read under strict isolation: with NO_KQ_MMVQ wrapping 2b on both sides, 2c must show zero diff — clean attribution is a low-cost gate to maintain.

← 71 · Index · 73 →

73 · D3-8 — G4 FusedQKV ported to CUDA: both layer classes covered, 14B short-KV breaks through parity (LANDED)

Result: one attn_bias_rope_store_f32 launch replaces the per-layer 7-launch chain of add_bias×3 + rope×2 + store_kv×2 (math verbatim, bitwise); both layer classes landed — the concat class (Op::FusedQKV, 24/48 14B layers) + the mixed-quant class (new Op::QkvBiasRopeStore, CUDA-only, 24/48 14B layers) — decode coverage 48/48; launches −310/decode-step (−22.5%, 22098 → 17130); 14B tg128 23.81 → 24.55 (+3.11%, isolation +3.15%), breaking through parity (24.55 vs llama 24.31, 1.010×), @3254 +1.63% (isolation +2.28%); distance to parity −9.7% → −7.9%. Commit: 3857633 (class 1 concat, 5 files +652) + a448a4a (class 2 mixed-quant, 8 files +212). Date: 2026-09-08.

1. Background — where things stood

While settling rms and bits_to_i32, D3-7's launch census incidentally exposed a structural fact: the 14B decode qkv chain on CUDA is completely unfused — per layer rope ×2, KvcacheStore ×2, add_bias ×3, i.e. 7 sub-2µs small kernels × 48 layers ≈ 0.45 ms/step. And this is not territory where a new design must take risks: the Metal backend has had Op::FusedQKV since the graph era (G4) — a concat matmul (loader-registered blk.{i}.attn_qkv rows = wq|wk|wv) + one attn_bias_rope_store fused pass, with decode logits bit-identical to the unfused path (verified on both 0.5B and 7B). The CUDA backend simply never advertised Op::FusedQKV.

So D3-8's (base /tmp/d3/minfer_pre_d38, sha1 ccb7df8f, HEAD 81e5d6e) positioning is Stage-3 Tier A: port a fusion whose semantics were already verified on another backend, settling the largest item of D3-7's "launch-structure residual (~0.7–1.0 ms)" (the qkv chain, 0.45 ms). The window carried one complication that had to be handled: the machine is ~1.8% faster than the D3-7 window (different co-tenant state; anchors 14B tg128 23.81 / @3254 22.04, 7B tg128 50.45 / @1641 48.91) — cross-window comparison would charge that drift to the fusion, so everything ran with the dual basis of same-window interleaved A/B + same-binary isolation A/B (post vs post + MINFER_NO_FUSE_QKV=1), with isolation as the primary criterion.

A second complication discovered within the session is why this step has two commits: of 14B's 48 layers only 24 layers have wq|wk|wv in the same quantization type (what the concat class can absorb); the other 24 layers are mixed-quant (e.g. Q6_K attn_v mixed into Q4_K q/k — D3-7 2b's routing put attn_v onto MMVQ and thereby froze the mixed-quant layer set). The concat matmul requires a single-ttype dispatch, so mixed-quant layers need a second road: three separate matmuls (no bias) + one epilogue. Doing class 1 alone covers 50%, which amounts to leaving half of the 0.45 ms on the table.

2. Principle — the GPU mechanism

What the 7-launch chain's tax is. The per-layer q/k/v-side work of decode (nt==1) is itself negligible: bias+rope for 5120 q elements, bias+rope/store for 1024 k/v elements each — each kernel executes in 1–2 µs, while each launch's fixed overhead is ~1.3–1.7 µs (the sub-2µs ocean of the D3-1/D3-7 census). 7 kernels × 48 layers = 336 launches/step, of which execution is less than half — the textbook shape of a launch-bound chain: merging the launch count converts almost directly into wall-clock savings. The nsys ledger is in §5: bias −144, rope −96, store −96, fused +48 — 7 → 1 per layer.

The per-layer ledger split by class (nt==1): unfused = 3 matmuls + 3 bias + 2 rope + 2 store = 10 launches; class 1 = 1 concat matmul + 1 epilogue = 2 (−8/layer, of which the matmul merge contributes −2); class 2 = 3 matmuls + 1 epilogue = 4 (−6/layer, the three matmuls kept as-is). 48 layers, class 1/2 half each: −8×24 − 6×24 = −336, plus the concat layers' standalone quantize +24 — the same magnitude as the measured −310/step (reconciled item by item against the dispatch table in §5). The reason class 2 does not concat is also in this ledger: repacking mixed-quant weights into one concat plane needs a loader-level repack and extra memory, while the epilogue form touches the weights zero — the gain of 6 fewer launches per layer does not require moving 1 byte of weights.

The fusion's bitwise legitimacy. The epilogue is not a "reimplementation"; it is a verbatim transplant: rope uses rope_f32's neox pairing (j, j+half), the same freq/theta expression, the same cosf/sinf — elementwise bitwise; bias is the same two-operand one-add (add_bias_f32's two operands); the store address (dst[pos * nkt + j]) and conversion are verbatim the same as store_kv_f32/store_kv_f16 (__float2half is RN, the same rounding as the unfused f16 store's scalar tail). Mathematically this is putting 7 verbatim-identical function bodies into one grid index — so the gates can be bitwise rather than tolerance.

Why the concat matmul is bitwise. Class 1 merges 3 matmuls into 1 (14B od 7168 × id 5120, 7B od 4608 × id 3584, Q4_K); the premise that makes this bitwise lives in decode MMVQ's structure: the kernel serves one row per 256-thread block and dispatch looks only at (ttype, id, nt) — zero coupling between rows. Concat's row i and the original independent matmul's row i are the same weight bytes × the same activation vector through the same reduction; there is no cross-row reorganization, so concat ⊂ row-wise identity. This argument was explicitly proven by a probe (§4), not left on paper.

pointer-form: one kernel, two uses. The q/k/v the kernel receives are section bases: the concat class passes the three sections of the concat matmul output (q=base, k=base+nqt, v=base+2·nkt), and the mixed-quant class passes three separate matmul outputs. The data-layout difference is pushed into the caller's pointer arithmetic, and the kernel semantics are singular — this is the key design for covering 100% of layers, and the core reason this is a "port" rather than a "rewrite".

capture/replay safety. pos = positions[0] is read device-side (nt==1, no host scalar crosses the launch) and replay re-reads the device buffer every step; grid/block are pure shape functions (nqt/2 + nkt/2 + nkt threads, 256 threads per block — 14B is 2560+512+1024 = 4096 threads → 16 blocks, 7B is 2560 threads → 10 blocks) with no nkv/n_past in them — CUDA Graph captures once and replays per frame. The positions i32 conversion is supplied by D3-7's just-landed memo (exactly one conversion per step, stable pointer).

3. Implementation

3.1 Design choices (why this shape and not another)

  • One epilogue kernel, two pointer forms (rather than writing a separate kernel for mixed-quant): there is only one math surface, so the gate needs proving only once; mixed-quant's difference lives entirely in "where the q/k/v pointers point", which belongs to the caller.
  • Class 2 uses a new Op instead of being crammed into Op::FusedQKV: FusedQKV's semantics carry the concat weight name (FusedQkvMeta.qkv_weight), which mixed-quant does not have — its inputs are three separate matmul outputs. The new Op::QkvBiasRopeStore + QkvBiasRopeStoreMeta (= FusedQkvMeta minus the concat weight) keeps dispatch, the scheduler's kv_pair(layer) resolution, and the CPU backend's loud-Err set each clean.
  • q's in-place alias rule unchanged: the builder wires attention to the epilogue node (q's matmul buffer has exactly one consumer), the allocator keeps the §5 alias (the Silu/RoPE family extended + ensure_kv), and the backend adds one more layer of insurance — "D2D copy if the alias does not hold" (the RoPE arm's pattern).
  • Class 2 is CUDA-only: the emission condition is cfg(feature="cuda") + CudaState present; Metal keeps the unfused chain for mixed-quant layers (supports_fused returns false for the new op) — bitwise-neutral vs pre-D3-8. A porting task does not reopen the Metal battlefield.
  • The gate enters the reuse identity: nt==1 && gpu && fuse_qkv (class 2 additionally requires CUDA present) is part of the graph params — MINFER_NO_FUSE_QKV=1 forces a rebuild, making A/B reliable.

3.2 Key code

Excerpt A · the fused kernel (current tree src/cuda_kernels.cu, landed in 3857633) — the grid is one linear index: q's rope pairs (nqt/2) + k's rope pairs (nkt/2) + v elements (nkt), one work unit per thread:

// src/cuda_kernels.cu (current tree; header comment excerpt: D3-8 CUDA port of Metal's
// kernel_attn_bias_rope_store ... The rope math is VERBATIM rope_f32 (NEOX
// pairing (j, j+half), same freq/theta expression and cosf/sinf —
// bitwise-identical per-element results) ... positions[0] is read
// device-side (nt==1; no host scalar crosses the launch — CUDA Graph
// capture/replay safe). Thread mapping (Metal's): one thread per
// (head, d < hd/2) rope pair for q and k, one thread per v element →
// grid = nqt/2 + nkt/2 + nkt.
__global__ void attn_bias_rope_store_f32(
    float* __restrict__ q, float* __restrict__ k, float* __restrict__ v,
    const float* __restrict__ bias_q, const float* __restrict__ bias_k,
    const float* __restrict__ bias_v,
    float* __restrict__ kv_k, float* __restrict__ kv_v,
    int nqt, int nkt, int hd,
    float freq_base, float freq_scale,
    const int* positions, int kv_is_f16
) {
    const int half_dim = hd / 2;
    const int qpairs = nqt / 2;
    const int kpairs = nkt / 2;
    const int total = qpairs + kpairs + nkt;
    const int u = blockIdx.x * blockDim.x + threadIdx.x;
    if (u >= total) return;
    const int pos = positions[0];              // device-side: capture-safe

    if (u < qpairs) {
        // q section: bias + rope in place (attention reads q at offset 0)
        const int head = u / half_dim;
        const int d    = u % half_dim;
        const int j  = head * hd + d;          // ← the caller already picked the section pointer: concat or separate buffers
        const int j2 = j + half_dim;
        float x0 = q[j]  + bias_q[j];
        float x1 = q[j2] + bias_q[j2];
        float freq = freq_scale / powf(freq_base, (2.0f * d) / hd);
        float theta = pos * freq;
        float cs = cosf(theta), sn = sinf(theta);
        q[j]  = x0 * cs - x1 * sn;             // verbatim rope_f32 neox pairing
        q[j2] = x0 * sn + x1 * cs;
    } else if (u < qpairs + kpairs) {
        // k section: bias + rope in place + store into the K region
        /* after the same-form rope:
           k[j] = r0; k[j2] = r1;
           ((__half*)kv_k)[(size_t)pos * nkt + j]  = __float2half(r0);  // RN
           ... f32 branch: kv_k[(size_t)pos * nkt + j] = r0; */
    } else {
        // v section: bias + store into the V region
        const int j = u - qpairs - kpairs;
        const float val = v[j] + bias_v[j];
        v[j] = val;
        /* ((__half*)kv_v)[(size_t)pos * nkt + j] = __float2half(val); */
    }
}

Excerpt B · the launcher and the Rust entry (current tree) — the launcher comment states outright that this is Metal's dispatch_1d shape; the Rust-side CudaState::attn_bias_rope_store only forwards pointers:

// src/cuda_kernels.cu (current tree)
// D3-8: fused decode QKV epilogue launcher — 256-thread blocks over the
// flat (nqt/2 + nkt/2 + nkt) thread mapping (Metal's dispatch_1d shape).
void launch_attn_bias_rope_store(
    float* q, float* k, float* v,
    const void* bias_q, const void* bias_k, const void* bias_v,
    void* kv_k, void* kv_v,
    int nqt, int nkt, int hd, float freq_base, float freq_scale,
    const int* positions, int kv_is_f16, cudaStream_t stream
) {
    const int total = nqt / 2 + nkt / 2 + nkt;
    const int block = 256;
    const int grid = (total + block - 1) / block;
    attn_bias_rope_store_f32<<<grid, block, 0, stream>>>(q, k, v, ...);
}
#![allow(unused)]
fn main() {
// src/cuda.rs (current tree, excerpted comments)
/// D3-8: fused decode QKV epilogue (G4 CUDA port of Metal's
/// `attn_bias_rope_store`). `q`/`k`/`v` are POINTER-FORM section bases:
/// the concat class passes sections of the concat matmul output [q|k|v]
/// (nt==1), the mixed-quant class passes the three separate matmul
/// outputs. Biases added per section, q/k roped in place (math verbatim
/// `rope_f32`), k/v stored into the persistent regions at the same
/// addresses as `store_kv_f32`/`store_kv_f16` (f32 or f16 per `kv_is_f16`).
pub fn attn_bias_rope_store(&self, q, k, v, bias_q, bias_k, bias_v,
                            kv_k, kv_v, nqt, nkt, hd, freq_base,
                            freq_scale, positions, kv_is_f16) {
    unsafe { launch_attn_bias_rope_store(/* pointer forwarding */); }
}
}

Excerpt C · the builder's two-class dispatch (current tree src/models/qwen2/graph.rs) — the fuse_qkv gate + the class 1/class 2/else three arms:

#![allow(unused)]
fn main() {
// src/models/qwen2/graph.rs (current tree)
let fuse_qkv = nt == 1
    && params.cparams.gpu
    && params.cparams.fuse_qkv
    && l.bq.is_some() && l.bk.is_some() && l.bv.is_some();
// class 2 is CUDA-only: on macOS (feature off) the mixed-quant
// layers keep the unfused chain, bitwise-neutral vs pre-D3-8.
#[cfg(feature = "cuda")]
let qkv_epilogue_ok = crate::cuda::CudaState::get().is_some();
#[cfg(not(feature = "cuda"))]
let qkv_epilogue_ok = false;
let (q, kv) = if fuse_qkv && Self::qkv_concat_available(&l.wq, &l.wk, &l.wv) {
    // class 1: one concat matmul (blk.{il}.attn_qkv) + the fused epilogue;
    // q sits at concat offset 0 (nt==1, no stride issue), K/V already
    // written by the fused store
    let qkv = b.fused_qkv(normed, inp_pos, il, FusedQkvMeta { /* .. */ });
    let kv = b.kvcache_load(il, nkt, n_ctx, nk);
    (qkv, kv)
} else if fuse_qkv && qkv_epilogue_ok {
    // D3-8 class 2 (mixed quant types): separate matmuls without
    // bias, then one epilogue pass (bias×3 + rope×2 + store×2 → 1).
    // Attention is wired to the epilogue node so q's matmul buffer
    // has exactly one consumer (in-place alias rule, §5).
    let q = b.matmul(normed, l.wq.as_ref().unwrap(), None);  // no bias
    let k = b.matmul(normed, l.wk.as_ref().unwrap(), None);
    let v = b.matmul(normed, l.wv.as_ref().unwrap(), None);
    let q = b.qkv_bias_rope_store(q, k, v, inp_pos, il,
        QkvBiasRopeStoreMeta { bias_q: .., bias_k: .., bias_v: ..,
            nqt: nh * hd, nkt, hd, freq_base: hp.rope_freq_base,
            freq_scale: hp.rope_freq_scale, rope_style: hp.rope_style,
            kv_elems: nkt * n_ctx });
    let kv = b.kvcache_load(il, nkt, n_ctx, nk);
    (q, kv)
} else { /* unfused chain: 3 matmuls (with bias) + 3 bias + 2 rope + 2 store */ };
}

(The blk.{i}.attn_qkv concat weight is registered on the backend registry by the loader — register_weight("blk.{i}.attn_qkv", ...); the 28 lines 3857633 added to the loader are exactly this registration chain.)

Excerpt D · the backend dispatch arm (current tree src/graph/cuda_backend.rs) — guards + alias insurance + a single launch:

#![allow(unused)]
fn main() {
// src/graph/cuda_backend.rs (current tree; arm header comment excerpt: ... the q matmul buffer
// has exactly one consumer); fall back to a D2D copy if it ever doesn't
// (RoPE arm pattern). Replaces the 7-launch small-kernel tail
// (add_bias×3 + rope×2 + store×2) → 1 launch.
Op::QkvBiasRopeStore { layer } => {
    let meta = /* NodeMeta::QkvBiasRopeStore(m); loud-Err when missing */;
    let nt = node.out_shape[1];
    if nt != 1 {
        return Err(format!("cuda: {}: QkvBiasRopeStore is decode (nt==1) only, got nt={nt}", node.name));
    }
    if !matches!(meta.rope_style, RopeStyle::NonInterleaved) {
        return Err(format!("cuda: {}: qkv epilogue rope style {:#?} not supported \
                            (rope_f32 is neox/non-interleaved only)", node.name, meta.rope_style));
    }
    if meta.hd == 0 || meta.hd % 2 != 0 { /* loud-Err: head dim must be even */ }
    let (k_id, v_id) = kv_pair.ok_or_else(|| format!("KV regions for layer {layer} not allocated"))?;
    // q: in-place (out aliases the q input; copy when it doesn't)
    if in_bufs[0] != out_buf {
        self.copy_d2d(in_bufs[0], out_buf)?;
    }
    /* bias pointer resolution (unregistered → loud-Err); pos = self.positions_i32(in_bufs[3])?
       — supplied by the D3-7 memo, one conversion per step */
    self.state.attn_bias_rope_store(
        self.ptr_of(out_buf)?, self.ptr_of(in_bufs[1])?, self.ptr_of(in_bufs[2])?,
        bq, bk, bv, self.ptr_of(k_id)?, self.ptr_of(v_id)?,
        meta.nqt, meta.nkt, meta.hd, meta.freq_base, meta.freq_scale,
        pos, self.kv_f16,
    );
}

Excerpt E · the allocator's alias extension (current tree src/graph/alloc.rs) — the epilogue joins the Silu/RoPE family, and ensure_kv runs for that layer:

#![allow(unused)]
fn main() {
// src/graph/alloc.rs (current tree)
Op::Silu | Op::RoPE { .. } | Op::QkvBiasRopeStore { .. } => {
    // D3-8: the mixed-quant QKV epilogue also needs the layer's
    // persistent KV regions (it stores k/v like FusedQKV).
    if let Op::QkvBiasRopeStore { layer } = &node.op {
        let kv_elems = match &node.meta {
            NodeMeta::QkvBiasRopeStore(m) => m.kv_elems,
            _ => node.n_elements(),
        };
        self.ensure_kv(*layer, backend, kv_elems);
    }
    // In-place elementwise transforms: alias the input buffer ...
    if last_use[id] > i { /* the output aliases src[0]'s buffer */ }
}

Excerpt F · the probe test's two pointer forms (current tree src/cuda.rs, cuda_fused_qkv_epilogue_bitwise) — the front line of the bitwise gate: on the same concat-matmul output, run the fused form (form 1: section bases = base + offsets) and the separate-buffer form (form 2), and compare byte-for-byte against the unfused 7-launch chain over the q/k/v sections and the KV rows; all combinations of the 14B/7B geometries × f32/f16 KV:

#![allow(unused)]
fn main() {
// src/cuda.rs (current tree, test excerpt)
/// D3-8 probe A: fused `attn_bias_rope_store` vs the unfused chain
/// (add_bias×3 + rope×2 + store_kv×2) on the SAME concat-matmul output —
/// bitwise on the q/k/v sections AND both KV regions, f32 + f16 KV,
/// 14B + 7B shapes, BOTH pointer forms (concat sections AND three
/// separate buffers — the mixed-quant class-2 wiring).
// ---- FUSED form 1: concat buffer, section pointers ----
let base = d_fused as *mut u8;
let (q1, k1, v1) = (
    d_fused,
    unsafe { base.add(nqt * 4) } as *mut std::ffi::c_void,         // k section
    unsafe { base.add((nqt + nkt) * 4) } as *mut std::ffi::c_void, // v section
);
st.attn_bias_rope_store(q1, k1, v1, dbq, dbk, dbv, dk_f, dv_f, ...);
// ---- FUSED form 2: three separate buffers (class-2 shape) ----
let d_q2 = dev_alloc(nqt * 4);
let d_k2 = dev_alloc(nkt * 4);
let d_v2 = dev_alloc(nkt * 4);
/* same kernel, same biases, compared byte-for-byte against the unfused chain */
}

Reading the two commits' change surfaces together: 3857633 (class 1) = cuda.rs +389 (wrapper+probe), cuda_kernels.cu +113 (kernel+launcher), cuda_backend.rs +103 (the FusedQKV arm), qwen2/graph.rs +21 (wiring), loader.rs +28 (concat weight registration); a448a4a (class 2) completes the graph infrastructure: ops.rs (Op+Meta), builder.rs, alloc.rs, scheduler.rs (kv_pair resolution), cpu_backend.rs (the loud-Err set), json.rs (dump-graph/trace naming), cuda_backend.rs, qwen2/graph.rs. The kernel and probe were in place with the class 1 commit, and class 2 reuses the same kernel — that is the landing order of "one math surface, two entry points".

3.3 Pitfalls

  • The launch ledger is not the ideal −384. Beyond the paper ledger of 7 → 1 per layer (336 − 48), two more items must be counted: the 24 concat layers' 3 matmuls → 1 (−48), and the concat matmul's shared-A standalone quantize +24 — D3-5's MmqCache record window only covers the pre-existing "producer → following matmul group" consumption pairs, and the concat created a new consumption combination the memo window did not catch. The measured total is −310/decode-step (16-step trace, −4968 total).
  • The gate script nearly mistook tok/s for divergence. The only pre/post difference in the greedy battery's raw output was the perf banner's tok/s number — every run's "first divergence at byte ~400" was 1789.6 vs 1788.7 tok/s. The rule was finalized: strip timing lines before comparing.
  • The node{N}_* dumps are an instrument limitation, not a gate. These informational dumps read a whole recycled pool slot, and their content at dump time depends on binary layout: node3 diffs within the same binary; node5/8 diff pre-vs-post even under MINFER_NO_FUSE_QKV=1 (where graph and loader behavior are identical). The real structural gates are the prefill DOT graph byte-identical (topology unchanged) + logits/KV fully byte-identical.
  • The f16 store's rounding mode must be verbatim. __float2half is RN — exactly the same rounding as the unfused f16 store's scalar tail, which is why f16 KV's bitwise holds. Any refactor that "conveniently rewrites the conversion" will crash on this gate.
  • The rope style guard must be loud. The kernel is a verbatim transplant of the neox (non-interleaved) pairing; silently running another rope style is far more dangerous than a loud-Err — the arm returns an explicit Err. Guards of the same kind: nt==1 (the epilogue is decode-only) and hd even (the rope pairing requires it).
  • The mixed-quant layer set is not static. D3-7 2b routed 14B attn_v into MMVQ, freezing the "11 attn_v layers are Q6_K" layer set; this step's qkv_concat_available decides layer by layer at build time rather than hard-coding counts — the gate is the per-layer reconciliation of the census's (D3-7) 24/48 against the measured trace, not a hard-coded expectation.

4. Verification

  • (i) Probe tests (defend the kernel math and dispatch equivalence): double green — the fused epilogue vs the unfused 7-launch chain bitwise on the q/k/v sections and the KV rows (both pointer forms, f32+f16 KV, 14B+7B shapes); concat matmul vs the three separate matmuls bitwise (14B od 7168 × id 5120, 7B od 4608 × id 3584, Q4_K) — the argument "decode MMVQ dispatches row by row ⇒ concat ⊂ row-wise identity" explicitly proven.
  • (ii) Dump gate (defends against end-to-end numeric drift): MINFER_GRAPH_DUMP, 14B prompt_short + 7B prompt_1k7, −n 4: logits_{prefill,decode} + all kv{layer}_*.f32 byte-identical pre-vs-post — 98/98 @14B, 58/58 @7B (the informational node dumps see 3.3; not used as gates).
  • (iii) Greedy battery (defends against sampler/per-token behavior changes): −n 256, 5/5 seeds × both models byte-identical; rp=1.0 identical; the MINFER_NO_FUSE_QKV=1 control identical; the temp-0.8 sampled control identical — the only pre/post diff in the raw output is the perf banner's tok/s (see 3.3).
  • (iv) Suite 172/0/3 (D3-7's 170 + 2 new probes); prefill undisturbed (pp3254 1833 t/s pre-vs-post — the fusion is gated at build time on nt==1).
  • Coverage audit: the decode-layer dispatch of MINFER_GRAPH_TRACE=1 reconciles layer by layer with D3-7's GGUF census — 48/48 @14B (24 concat + 24 epilogue), 28/28 @7B (14 + 14).

5. Results

Mechanism (nsys, 14B @3254, NO_CUDA_GRAPH, the same 16-decode-step trace): total kernel launches 22098 → 17130 (−22.5%); per decode step ≈ −310:

dispatch itemΔ/stepnote
add_bias−1443 × 48 layers → folded into the epilogue
rope−962 × 48 layers → folded into the epilogue
store_kv−962 × 48 layers → folded into the epilogue
attn_bias_rope_store+481 fused pass per layer
q4_K MMVQ−4824 concat layers: 3 matmuls → 1
standalone quantize+24concat shared-A, a new consumption pair the D3-5 memo window does not cover

Net wall clock ≈ −0.9 ms/step, exceeding D3-7's census projection of 0.45 ms/step for the qkv chain — the difference comes from launch gaps (336 → 48 emission points, each ~1.3–1.7 µs of fixed overhead no longer paid one by one) and class 1's matmul merge (−48); neither of these shows up in the census's "kernel time" basis — only the launch count reconciles.

Wall clock (interleaved 3-pair medians, every pair clean-separated; isolation = post vs post+NO_FUSE_QKV, same binary):

configpre_d38postΔisolation
14B tg12823.8124.55+3.11%+3.15% (24.53 vs 23.78)
14B @325422.0422.40+1.63%+2.28% (22.46 vs 21.96)
7B tg12850.4550.98+1.05%+1.03% (50.92 vs 50.40)
7B @164148.9149.51+1.23%+1.00% (49.44 vs 48.95)

Guards all hold with margin (14B tg128 ≥ 22.7, @3254 ≥ 21.5; 7B ≥ 49.0 / ≥ 47.9). @3254's +2.28% (isolation basis) exceeds the +0.8% all-fusion landing bar — class 2's epilogue (the "tail lever") landed in the same commit as class 1, saving a separate tail measurement window.

Why isolation is the primary criterion. This window's machine is ~1.8% faster than the D3-7 window (co-tenant state); the pre_d38 anchor and D3-7's post anchor have no comparability whatsoever. Same-binary post vs post+MINFER_NO_FUSE_QKV=1 differs by only the fusion switch itself under the same machine condition, so the drift term is eliminated entirely — which is why both Δ columns in the table must pass the isolation test to count (all four pass: +3.15/+2.28/+1.03/+1.00%).

Distance to parity (post-D3-8, 14B @3254): 22.40 t/s = 44.64 ms/step vs llama 41.12 ms → −3.52 ms (−7.9%) (post-D3-7 was −4.44 ms / −9.7%). 14B tg128 breaks through parity this window: 24.55 vs llama 24.31 (1.010×); 7B stays ahead (tg128 1.032×, @1641 1.002×). The remaining 14B @3254 list: matmul aggregation ~2.9 ms (the decode-GEMM program), exposed-latency attention ≤0.63 ms (mechanism in doubt), output head 0.44 ms, ffn_down-q6K 0.49 ms, launch-structure residual (the add/add/swiglu tail, dud split ~0.07 ms) — the qkv-chain item closes here.

6. Lessons

  1. Porting an already-verified fusion is an order of magnitude cheaper than writing a new one. Metal G4's semantics (verbatim math + bitwise gates) came across as-is, so the risk of the 652-line insertion concentrated on "wiring" rather than "math" — the probe proved concat ⊂ row-wise identity first, and only then did the graph infrastructure (Op/Meta/alloc/scheduler) dare to spread out in one go.
  2. pointer-form decouples "where the data is" from the kernel semantics. Section bases let the concat and separate-buffer shapes share one kernel — mixed-quant models do not have to give up per-layer reality for the fusion, and coverage went 50% → 100%.
  3. The gate script must defend against itself. The only "divergence" in the raw output was the perf banner's tok/s — strip the timing lines before comparing, or every gate run chases ghosts.
  4. Informational dumps do not enter gates. Diagnostic output reading recycled pool slots is inherently binary-layout dependent; gates only trust semantically stable artifacts like logits/KV/DOT, and diffs of diagnostic output are calibrated with "same-binary self-diff" before being interpreted.

← 72 · Index · 74 →

74 · D4-2 — B0 latent correctness fix + all bitwise occupancy/prefetch axes closed (LANDED)

Result: the q6_K MMVQ v2_pf dispatch had only a lower bound (id > 8192) and no upper bound, while the kernel processes only two units per thread (npair ≤ 512) — 7B ffn_down (id 18944, npair 592) had silently dropped 80/592 = 13.5% of the down-projection work on every decode step since D3b-1b; first-step logits deviated by as much as 4.79 yet argmax survived, and every prior gate compared v2_pf against v2_pf, so nobody saw it. The fix (b31084c) gives the pipelined route an upper bound id ≤ 16384; taller rows take the v2 loop form (identical per-unit arithmetic and ascending accumulation order); the 7B anchors were re-anchored 50.67 → 49.30 / 49.84 → 48.49 (−2.7%, the honest price of correctness), 14B bitwise-unchanged. The same session closed Lever A arithmetically (llama L2 prefetch: the mechanism is inert on our decode shapes, pre-build veto) and Lever B (B1 __launch_bounds__ register squeeze, B2 160-thread right-sizing, both killed with mechanism), with zero perf regressions. Commit: b31084c (B0) + docs commit. Date: 2026-09-09.

1. Background — where things stood

The map D4-1's census (/tmp/d4/D4_DESIGN.md) drew for decode was: minfer's decode GEMV aggregation is already ahead of llama by 0.85 ms — the q4_K-class kernels run at 84% of DRAM peak, and the q6_K class still carries a small occupancy deficit; the real bulk is the attention structure (read at the time as ~2.53 ms/step) plus odds and ends of tail. D4-2's budget was about 3 hours, with the discipline carried over from r44/r46: every lever gets its mechanism arithmetic written before deciding whether to build, and every rejection must leave reproducible numbers behind.

Spreading the census numbers out: one 14B @3254 step is roughly 3.32 ms attention + ~3.6 ms decode GEMV (of which the three q6_K shapes are the bulk) + the remaining tail, against llama's same-window full step of ~5.3 ms — so "decode GEMV is already ahead" is the standing premise, and D4-2's three lines each targeted one thing: attention (left for D4-3; this doc only hands over the decomposition-update brief in §5), q6_K occupancy (Lever B), and a mechanism checklist of "what else have we not copied" (Lever A — the only decode mechanism in llama's kernels that we lack is L2 prefetch).

Of the session's three target classes, B0 was originally just routine audit: check whether each decode kernel's dispatch gate actually matches the shape set it covers. The motivation for this audit came from the previous phase's experience — the decode dispatch layer had branches bolted on repeatedly through the D3 series (D3-7's attn_v MMVQ routing, D3b-1b's tall-row pipelined branch, FusedQKV's two layer classes), and every added branch is one more round of "shape set × kernel coverage" alignment risk. When D3b-1b landed q6_k_q8_mmvq_v2_pf, its gate read id > 8192: the intent was "tall rows (npair > 256, i.e. id > 8192) take the pipelined form", but the pipelined kernel's coverage ceiling is 512 units, and nothing forced the gate's upper bound to align with it.

The design doc /tmp/d4/D4_DESIGN.md organized the session into four levers: B0 (the correctness audit's product), A (porting the L2 weight prefetch from llama's GB10 kernels — it was the only mechanism on the llama side we lacked at the D4-1 reconciliation), B (the bitwise occupancy axis of q6_K decode: register budget, pipeline depth, block geometry), C (PDL, recorded as follow-up). Every lever had pass/veto criteria pre-registered, and the A/B axes only admitted bitwise changes — any loosening of the numeric path was excluded from this session (that belongs to "tolerance-gated redesign", see §5's 0.42 ms true-deficit conclusion).

The consequence on 7B is concrete: Qwen2.5-7B has 10 q6_K layers of ffn_down (id 18944 → npair 592). With 592 > 512, units 512..591 are touched by no thread — on every decode step, on every such row, 13.5% of the down-projection dot product simply does not exist. And 14B's corresponding shape is id 13824 → npair 432 ≤ 512, which was always correct. So of the "7B +3.0%/+2.7%" gain D3b-1b recorded at the time, the vast majority was actually the skipped work, not pipelining (§1's correction note already voided D3b-1b's 7B half-row; 14B's +0.44%/+0.24% is the true magnitude of pipelining itself).

The dropped part has one more layer of stealth: 592 = 512 + 80, and 80 is exactly the number of units a 256-thread round can cover (80 ≤ 256) — that is, the loss happens on the third logical iteration, while the kernel's first two rounds (0..255, 256..511) are perfectly normal; any validation that only spot-checks the front half of a row sees flawless data. The bug's shape gives it natural immunity to "sampling validation"; only a full-coverage gate (complete logits against a reference) can expose it.

Why could this bug lie dormant for an entire D3b-1b→D4-2 cycle? Two reasons compound. First, argmax survives: dropping 13.5% of the down-projection is a fixed systematic gap (the same 80 units dropped every step); the logits are pulled off wholesale, but top-1 never flips on the tested prompts — the max|Δlogit| 4.79 damage is visible only in the dump, and at the generated-text level it "looks normal". Second, every gate compared v2_pf against v2_pf: D3b-1b's bitwise gates (114/114 dump memcmp, greedy byte-for-byte) compared the pre-fix and post-fix v2_pf binaries — the bug existed on both sides of the comparison, so the memcmp was always equal. v1 (the fully-covering loop kernel), as the semantic reference, was never pulled into 7B's decode gates.

That is the origin of this doc's headline-level lesson: in a campaign "optimizing correctness", the first job is confirming the engine computes the right thing.

2. Principle — the GPU mechanism

The v2 unit geometry and where the 512 bound comes from. Inside q6_K's 256-element super-block there are 8 is-pairs (each pair = two 16-element sub-blocks sharing the same set of ql/qh bytes and each carrying its own 2-bit scale). The v2 kernel's mapping is one thread per pair (32 elements): npair = id/32, and a 256-thread block covers 256 pairs per round. The v2 loop form iterates for (u = threadIdx.x; u < npair; u += 256) and covers any npair; v2_pf (pipelined) unrolls this loop to exactly two iterations — u0 = tid, u1 = tid + 256 — with both units' weight and activation loads issued before accumulation, moving the second unit's weight-load latency off the critical path (this is the source of D3b-1b's +0.2–0.4%). The price of the two-way unroll is registers: v2_pf is 48 regs/thread vs v2's 40 — the register budget buys exactly this dual-unit pipeline. So v2_pf covers npair ≤ 512, and this ceiling is decided by the kernel's structure, not a dispatch parameter.

7B ffn_down's id 18944 → npair 592. u1 = tid + 256 reaches at most 255 + 256 = 511; units 512..591 fall outside every thread's two test and are loaded and accumulated by no one, ever. 80/592 = 13.5% — which matches the −2.7% the 7B anchors dropped after re-anchoring (down-q6K is about 20% of 7B's per-step weight stream; 13.5% × 20% ≈ 2.7% — the arithmetic is self-consistent).

Lever A: why llama's L2 weight prefetch is inert on our side. llama's mmvq kernel issues an L2 prefetch for "the next round's weight blocks" inside its K loop; the prefetch distance is 2 loop iterations = 2·bpi 256-element blocks, where blocks_per_iter = vdr·nwarps·32/qi: from vecdotq.cuh, VDR_Q4_K_Q8_1_MMVQ=2 and QI4_K=16 (and VDR_Q6_K_Q8_1_MMVQ=1, QI6_K=8) both give 8 threads per 256-element block; on the GB10 decode path ncols_dst=1 and nwarps=4, so bpi = 4·nwarps = 16 and the prefetch distance is 32 blocks. The prefetch actually fires only when bpr > 2·bpi = 32 (when a row has ≤ 32 blocks, the "32 blocks ahead" address is already outside the row and the path does nothing): in 14B only ffn_down (bpr = 13824/256 = 54) qualifies; all id-5120 shapes (gu/qkv/q/o, lm_head, attn_v, bpr = 20) never prefetch.

Our decode MMVQ form, meanwhile, is one thread per unit, 256 threads per round: npair is 80 (gu/qkv/qo), 216 (down-q4K), 160 (lm_head/attn_v) by shape, and v2_pf's 432 unrolls into u0/u1 — every decode shape's K loop is exactly 1 iteration. What llama's prefetch needs is "a distance 2 rounds out"; we have no "2 rounds out" to prefetch into. Then look at the mechanism's own payoff: where llama's prefetch actually engages (ffn_down-q4K) it runs 224.1 GB/s, below our prefetch-less 228.6 GB/s. Both families sit at 75–84% of GB10's DRAM peak — in the bandwidth-saturated regime, L2 prefetch produces no new bandwidth; it only hides latency inside the overlap window, and the 8-threads-per-block MMVQ form already hides latency through multi-block parallelism. Conclusion: closed pre-build, no build, no commit (the D4-1 §3 precedent).

A clarification of "8 threads per 256-element block", because it is the root of the two families' shape difference: llama's MMVQ has 8 threads jointly consume one 256-element super-block (qi 32-element chunks × vdr register pairs per thread, e.g. Q4_K's QI=16, VDR=2 → 2×16 nibbles per thread), 256 threads eat 32 blocks per round, and the loop iterates as many rounds as the row has blocks — ffn_down's bpr=54 gives 54/32 ≈ 1.7 rounds, so a loop body exists and prefetch has something to attach to. Our v2 form, in reverse, presses a whole pair (32 elements) into one thread's register set (4×uint4 ql/qh + q8 slots), and 256 threads eat 256 pairs per round = an entire row with id ≤ 8192 — the loop body is flattened to iteration count 1 on most shapes. These are two legitimate form choices: llama uses shallow threads × deep loops (prefetch has a target), we use wide threads × zero loops (no prefetch needed), and D4-1's measurement already ruled the latter not slower.

Why the bandwidth-saturated regime kills both the A and B axes at once. The number 75–84% of DRAM peak is the denominator of the whole session: decode GEMV is a pure streaming workload (one 32-element dot product per weight byte, arithmetic intensity 0.03 FLOP/B); piling more resident warps into the SM merely issues more concurrent load requests — while bandwidth has slack this lifts throughput (that is how the q4_K class reached 228.6 GB/s), but with only 16–25% headroom left, an occupancy increment's real effect is a longer request queue, latency hidden, throughput unmoved. That is the unified explanation of why B1 (+1 block probed at −0.75%) and B2 (+3 blocks probed at +1.53%) both fail to pay, and it is the same coin as Lever A's conclusion "prefetch produces no bandwidth in the saturated regime". Conversely, what is a real lever in this regime is a layout that still sits below 93% effective byte rate (q6_K's 224B stride streams out 14 dead bytes per block) — it cuts dead bytes out of the request stream directly, needing no extra concurrency; D4-4 picked up that thread (the dpl in doc 76).

Lever B: the occupancy axis's mechanism ledger. Fresh ncu (2025.3.1, sudo recipe, 14B @3254) confirmed v2_pf holds 48 regs/thread → Block-Limit-Registers 5 blocks/SM, below Block-Limit-Warps 6 (sibling control kernels at 40/39 regs both reach 6 blocks). Three candidate directions: (B1) use __launch_bounds__(256,6) to force the compiler down to 40 registers, freeing room for a 6th block; (B1c) the reverse — hand npair-432 to the v2 loop (1 unit/thread), trading 6 blocks × 1-unit-MLP against 5 blocks × 2-unit-MLP; (B2) right-size 256 threads to 160 (zero idle warps for the npair-160 shapes, 9 resident blocks × 5 active warps vs 6 × 5). Each one's mechanism expectation is "occupancy +1 block → more warps to hide latency", but each must pass both the bitwise and the wall-clock gates — occupancy itself is not the goal.

B1c's matchup deserves expansion, because it is the only "free" control of the three (both kernels already exist, both already bitwise-gated). 5 blocks × 2-unit-MLP: each resident warp holds two units' weights and activations already loaded into registers, the accumulation chain being the two serial segments "dot(u0) → dot(u1)", with the second segment's load latency already hidden by the pipeline; 48 registers → 5 blocks per SM. 6 blocks × 1-unit-MLP: each warp looks at one unit per round, and load latency must be hidden by cross-block warp switching; 40 registers → 6 blocks per SM. In theory the latter has 20% more resident warps (30 → 36); measured, v2_pf wins/ties — inside a 256-thread × 4-warp decode block, the per-block reduction (mmvq_block_reduce's shfl chain) and tail-wave effects eat the gain of "2 more blocks", while the 2-unit pipeline hides latency more efficiently inside a block. That is direct evidence that "register-capped one block below the warp cap" is not a free lunch: that one block's price is 48 registers of pipeline funding, and neither cutting it (B1) nor bypassing it (B1c's motive) pays it off.

3. Implementation

3.1 Design choices (why this shape and not another)

B0's fix shape: an upper bound + rerouted landing, not a generalized kernel. Three options: (a) add a loop to v2_pf so any npair is covered — this cancels the unroll and makes short rows also pay the 48-register price; (b) split by npair at dispatch — npair ≤ 512 goes to v2_pf, everything else to the v2 loop — zero kernel changes, since the v2 loop is already the full-coverage form for any npair; (c) write a third "generalized pipelined" kernel. (b) was chosen: the v2 loop's per-unit arithmetic is identical to v2_pf's and the accumulation order is the same ascending u — exactly the bitwise invariant D3b-1b established — so the rerouted taller rows need no new correctness proof. The bound is written as id ≤ 16384 (= 512 units × 32), strictly equivalent to npair ≤ 512.

One design detail worth recording: why the gate's deciding quantity is id and not npair. At the dispatch point id is a call argument (ready-made on the host side), while npair = id/32 is a derived quantity inside the kernel — recomputing npair in the host gate and comparing is semantically equivalent to writing id ≤ 16384 directly, but the latter presses the conversion into a constant, with the conversion noted in a comment. What really matters is not the notation but the alignment obligation: between the kernel's coverage ceiling (512 units) and the gate's ceiling (16384) there must be a comment-level conversion chain, so that anyone changing one end in the future sees the other end in the same screen of code. B0's root cause was precisely that this chain was missing in D3b-1b — the kernel comment said "npair > blockDim here (dispatch-gated)", pushing the obligation onto the gate, and the gate wrote only half of it.

The choice of verification anchor: the -n 1 first-step dump. After the fix, a comparison point is needed that simultaneously decides "7B is fixed" and "14B was not touched". Full-run dump cross-binary comparison was rejected (reason in 3.3); the first step (prompt-only context, bit-identical inputs on both sides) is the only clean point; the 7B reference is the fully-covering v1 kernel — for v2_pf it is a semantic reference, not a bitwise one (see 3.3's rounding class), so the criterion is "the delta falls in the v1-vs-v2 rounding class and argmax matches", not "the delta is zero".

The pre-build discipline for Levers A/B. A was vetoed with prefetch-distance arithmetic before any code was written; B2 should likewise have been vetoed before writing code — its prior art (D3b-1c, D3-5 1b) was sitting in the §0 master table; that process lesson is recorded in 3.3.

3.2 Key code

Excerpt A · the before/after of the B0 fix (git show b31084c -- src/cuda.rs, the diff hunk in full) — the change is one condition line plus a comment, and the comment pins the bug's complete mechanism at the dispatch point:

--- a/src/cuda.rs
+++ b/src/cuda.rs
@@ -4170,7 +4170,16 @@ impl CudaState {
         unsafe {
             if Self::mmvq_v2(id) && blk_stride_padded {
                 // v2's uint4 ql/qh loads need the padded 224B stride
-                if id > 8192 {
+                // D4-2 B0 correctness fix: v2_pf processes exactly TWO units
+                // per thread (u = tid, tid+256 → npair ≤ 512), but the old
+                // gate (id > 8192) had no upper bound — npair=592 shapes
+                // (7B ffn_down id 18944, 10 q6_K layers) silently dropped
+                // units 512..591, corrupting 7B decode since D3b-1b. Guard
+                // the upper bound; taller rows take the v2 loop form, which
+                // walks any npair with identical per-unit arithmetic and the
+                // same ascending-u accumulation order (bitwise for every
+                // npair ≤ 512 shape, which keep the pipelined kernel).
+                if id > 8192 && id <= 16384 {

The current tree has since stacked B1c's MINFER_Q6K_PF A/B switch at the same spot (pipelined stays the default, "0" forces the v2 loop, consistent in semantics with r60's opt-out family):

#![allow(unused)]
fn main() {
// src/cuda.rs (current tree, the v2 dispatch section of q6_k_decode_mmvq)
if id > 8192
    && id <= 16384
    && !std::env::var("MINFER_Q6K_PF").map_or(false, |v| v == "0")
{
    // D3b-1b: tall rows (npair > 256, e.g. ffn_down id 13824)
    // run the pipelined variant (bitwise-identical, loads for
    // both serial units issue up front).
    launch_q6_k_q8_mmvq_v2_pf(wptr as *const u8, /* … */ 224, stream);
} else {
    launch_q6_k_q8_mmvq_v2(wptr as *const u8, /* … */ 224, stream);
}
}

Excerpt B · the v2_pf kernel in full (src/cuda_kernels.cu 1635–1661) — the bug's physical evidence sits in the comment on line 1655: the two test covers "is u1 out of range", but for npair > 512 shapes the u0/u1 two slots cannot hold a third unit to begin with; the kernel does not defend itself and relies entirely on the dispatch gate:

// src/cuda_kernels.cu (current tree; b31084c did not touch this kernel)
__global__ void __launch_bounds__(256) q6_k_q8_mmvq_v2_pf(
    const uint8_t* __restrict__ weights,
    const uint8_t* __restrict__ acts8,
    float* __restrict__ output,
    int od, int id, int nt, int blk_stride
) {
    const int row = blockIdx.x;
    const int t = blockIdx.y;
    const int nbe = id >> 8;
    const int row_stride = nbe * blk_stride;
    const int npair = id >> 5;                     // id/32: one is-pair per thread
    const uint8_t* x8row = acts8 + (size_t)t * (id >> 5) * Q8PB;
    const uint8_t* wrow = weights + (size_t)row * row_stride;

    float acc = 0.0f;
    const int u0 = threadIdx.x;                    // 0..255
    const int u1 = u0 + 256;                       // 256..511 ← where the upper bound comes from
    if (u0 < npair) {
        Q6kUnitRegs r0, r1;
        q6k_unit_load(u0, wrow, x8row, blk_stride, &r0);
        const bool two = u1 < npair; // npair > blockDim here (dispatch-gated)
        if (two) q6k_unit_load(u1, wrow, x8row, blk_stride, &r1);
        q6k_unit_acc(&r0, acc);                    // both units' loads issue first
        if (two) q6k_unit_acc(&r1, acc);           //  accumulation after = the pipeline itself
    }
    mmvq_block_reduce(acc, output, od, t);
}

For 7B ffn_down: npair = 592, u0 < npair is true for all 256 threads, and two = (u1 < 592) is true for all of 0..255 as well — the two slots load units 0..511 and accumulate/reduce as usual, units 512..591 vanish silently, with no out-of-bounds access, no NaN, no signal of any kind to make the old gate suspicious.

Excerpt C · the new landing for taller rows: the v2 loop form (src/cuda_kernels.cu 1536–1573 excerpt) — the full-coverage form for any npair, ascending-u accumulation; post-fix 7B ffn_down runs exactly here:

// src/cuda_kernels.cu (current tree, q6_k_q8_mmvq_v2 inner loop excerpt)
float acc = 0.0f;
for (int u = threadIdx.x; u < npair; u += 256) {   // full coverage of any npair
    const int kbx = u >> 3, pair = u & 7;
    const int s0 = 2 * pair, s1 = 2 * pair + 1;
    const uint8_t* blk = weights + (size_t)row * row_stride + (size_t)kbx * blk_stride;
    const float d = h2f(*reinterpret_cast<const uint16_t*>(blk + 208));
    const float sc0 = (float)(int8_t)blk[192 + s0];
    const float sc1 = (float)(int8_t)blk[192 + s1];
    const int chunk = pair >> 2, g = pair & 3;
    // padded 224B stride ⇒ every ql/qh piece is 16B aligned
    const uint4 qla = *reinterpret_cast<const uint4*>(blk + chunk * 64 + (g & 1) * 32);
    const uint4 qlb = *reinterpret_cast<const uint4*>(blk + chunk * 64 + (g & 1) * 32 + 16);
    const uint4 qha = *reinterpret_cast<const uint4*>(blk + 128 + chunk * 32);
    const uint4 qhb = *reinterpret_cast<const uint4*>(blk + 128 + chunk * 32 + 16);
    /* ... nibble unpack + 2-bit high-bit assembly, dp4a per v element ... */
    acc += d8 * sc0 * d * (float)dot0 + d8 * sc1 * d * (float)dot1;
}
mmvq_block_reduce(acc, output, od, t);

Its per-unit statements are verbatim the same lineage as v2_pf's (the comment on q6k_unit_acc's accumulation statement reads, in the original: "textually identical to the v2 accumulation statement") — that is the grounds on which the reroute needs no new correctness proof.

Excerpt D · the pipeline's funding structure (src/cuda_kernels.cu 1578–1601 excerpt) — where the 48 registers come from: both units' complete weight slots (4×uint4 ql/qh ×2 copies each + the scale/d scalars ×2 copies) are all held in registers before accumulation begins; this funding is exactly what B1 wanted to cut:

// src/cuda_kernels.cu (current tree; introduced in D3b-1b)
// Same mapping and arithmetic as q6_k_q8_mmvq_v2; both units' loads issue
// before either accumulates so the second unit's weight latency leaves the
// critical path. Bitwise-identical (see the file comment).
struct Q6kUnitRegs {
    uint4 qla, qlb, qha, qhb;      // one unit's ql/qh: 4×16B = 16 regs
    const uint32_t* xw;            // pointer to the q8 activation's 4B slot
    uint32_t shift;
    int g;
    float d, sc0, sc1, d8;
};                                 // ≈ 21 regs/unit × 2 copies + reduction state ≈ 48

__device__ __forceinline__ void q6k_unit_load(
    int u, const uint8_t* __restrict__ wrow, const uint8_t* __restrict__ x8row,
    int blk_stride, Q6kUnitRegs* r
) {
    const int kbx = u >> 3, pair = u & 7;
    const uint8_t* blk = wrow + (size_t)kbx * blk_stride;
    r->d = h2f(*reinterpret_cast<const uint16_t*>(blk + 208));
    r->sc0 = (float)(int8_t)blk[192 + 2 * pair];
    /* ... the uint4 loads of sc1 / ql / qh, all issued before accumulation ... */

3.3 Pitfalls

  • Cross-binary full-run dump comparison failed twice, invalid both times. After generated tokens diverge, the two runs feed different contexts into subsequent steps — from then on every logits/KV difference is cascade noise, measurable at 1e15–1e18-magnitude "deltas", unrelated to compute corruption; additionally the persistent KV region's unwritten tail slots hold pool garbage whose layout is binary-dependent. The only clean comparable point is the first decode step (-n 1: prompt-only context, identical inputs on both sides). This rule later became the D series' standard instrument (collected in the doc 77 methodology).
  • v1 is not a bitwise reference for v2. In the fix validation, 7B post-fix vs the v1 reference showed max|Δlogit| = 0.254, alarming at first glance — but 14B's own v2-vs-v1 is also 0.231: this is the v1-vs-v2 rounding class that has existed since the R2 era (argmax preserved), a pre-existing difference, not an incomplete fix. The adjudication gate must distinguish "fixed down to the reference's rounding class" from "fixed to the reference".
  • B1's spill shape. __launch_bounds__(256,6) squeezed v2_pf from 48 to 40 regs while generating STACK 40 (10 spill words) — the forced-occupancy ledger is not "8 fewer registers" but "more spill traffic in the hot loop". Seeing REG 40 + STACK 40 at the SASS level should already be the stop signal.
  • Prior-art search should happen before writing code. B2 (160-thread right-sizing) is verbatim the same lineage as D3b-1c's and D3-5 1b's conclusions ("9 blocks×160 live threads ≈ 6×256 allocated — the idle-thread gain does not exist"); those rows had been sitting in the §0 master table for a long time, and re-measuring's only value was upgrading "neutral" to "measured worse, with per-kernel control". Process correction: grep the master table's mechanism keywords before writing a kernel.

4. Verification

  • 14B pre-vs-post full dump gate: MINFER_GRAPH_DUMP 107/107 gate files byte-identical (the 7 node{N} diffs = the documented pool-slot instrument class), plus greedy per-token byte-identical — defends against the fix-only tree producing any behavioral drift on the npair ≤ 512 paths (this fix does not touch those dispatches).
  • Why the pool-slot instrument class can be exempted: those 7 node{N} files have pool slot numbers in their names, and the allocator's slot assignment order may differ between binaries (the same-named buffer lands in a different slot → different filename, same content). The adjudication protocol is "the other 107 files byte-identical + the node{N} diffs reproduce pre-vs-pre" — i.e. self-compare the unchanged binary once, and the same 7 node{N} diffs appear, proving it is instrument noise rather than behavior. This protocol was reused verbatim in D4-4 (doc 76).
  • 7B first-step logits gate (-n 1): post-fix vs the fully-covering v1 reference max|Δ| 0.254 with identical argmax (= 14B's own v2-vs-v1 rounding class of 0.231); pre-bug vs the v1 reference max|Δ| 4.72 — the same gate pins both the bug's magnitude and the fix's completeness.
  • 7B greedy stream alignment: post-fix, 7B greedy follows the v1 reference stream token by token — defends against the intermediate state of "the dump is right, generation still drifts".
  • suite 173/0/3: full-model regression, defending against the dispatch change rippling into other quantization paths (the gate change lives in the dispatch layer and should only affect the q6_K tall-row branch, but running the full suite is baseline discipline).
  • 7B's "semantic reference" gate: the fix introduced a new dispatch landing (the v2 loop taking over npair-592), a landing that had never run on 7B decode before — so beyond the dump gate, 7B greedy's token-by-token alignment against the v1 reference stream covers the "new landing × real generation loop" combination.
  • B1/B2's control gates: B2 first passed the bitwise 98/98 gate files, then used nsys for per-kernel isolated measurement, with the v2_pf/q4_K rows as the ±0.6% control group — defends against reading window drift as a kernel effect.

5. Results

The fix's own wall clock (pre = the b31084c binary, post = the final tree; interleaved 3× medians) — a fix-only tree should be identical, and it measured so:

configpre (B0)post (final)Δnote
14B tg12824.1324.01≈ 0 (noise; pair-3 straddle)fix-only tree behaviorally identical, as designed
14B @325422.1222.42≈ 0 (noise)guards hold (≥ 24.0 / ≥ 21.9)
7B tg12849.2149.29+0.2%re-anchor below
7B @164148.4148.50+0.2%guards hold

7B re-anchor (the buggy pre-D4-2 binary vs the post-fix B0 binary, 3× interleaved medians): tg128 50.67 → 49.30 (−2.7%), @1641 49.84 → 48.49 (−2.7%) — both shapes exactly −2.7%, the honest price of "starting to count the 13.5% of down-q6K work" (down-q6K ≈ 20% of 7B's per-step stream). The old 7B guards (49.0/47.9) were built on a work-dropping kernel and are voided: the new guards are tg128 ≥ 48.3 / @1641 ≥ 47.5. 14B cross-check: the pre/post medians interleave within window noise (the binary path is bitwise-identical).

Both shapes, same-day window, same fix, exactly equal drops (−2.7%/−2.7%) — that is not coincidence but the signature of arithmetic self-consistency: the fix's only behavioral change is letting ffn_down-q6K's 10 layers count the 13.5% of missing dot products back in, a workload that is constant per decode step, so the decode wall clock — insensitive to KV length — withdraws by the same proportion. Had the two shapes dropped by clearly different amounts, that would instead signal something else mixed into the measurement. The guard re-anchoring rule is recorded alongside: correct-over-fast — the guard floor is re-set to follow the "correct engine", never keeping a work-dropping path to preserve an old number; the old guard's provenance (which measurement, which binary set it) goes into the docs so the next re-anchor can trace it.

Same-window vs-llama (llama-bench ca3d5a3e1, 2026-09-09 window, aligned to the -n 128 anchors):

configminfer (post-D4-2)llamaratio
14B tg128 (KV≈0)24.0124.14 ± 0.020.995× (parity)
14B @325422.4224.14 (tg128 @ -p 3254)0.928×
14B pp3254~1816–18591634 ± 1121.11–1.14×
7B tg128 (KV≈0)49.3047.65 ± 0.091.035×
7B @164148.4947.69 ± 0.011.017×

Note: D4-1's 0.921× headline compared minfer tg64-exclusive against llama tg8, and still inside an sglang co-tenanted window — aligned to the tg128 anchors, 14B short-KV is parity, and 14B long-KV's 0.928× is the real gap (the attention structure item, handled by D4-3). 7B's ratios are on the correct engine (the pre-fix 7B numbers were flattered by the dropped work).

The operational meaning of this correction deserves to be spelled out: 0.921× (tg64-exclusive/tg8) and 0.995× (tg128/tg128) differ by 7 points with no code change at all — only the anchor alignment and the co-tenanted window changed. Decode A/B comparisons must (1) use the same -n anchor on both sides, (2) use the same window or explicitly state the window drift magnitude, (3) compare median to median. After D4-2 these three became the standard for decode comparisons (collected in the doc 77 methodology).

The three levers' closing numbers:

  • Lever A (llama L2 prefetch): closed pre-build. Mechanism ledger: the prefetch distance 2·bpi = 32 blocks only engages for rows with bpr > 32 (in 14B only ffn_down, bpr 54); our kernel runs exactly 1 K iteration per decode shape, with no "2 rounds out" to prefetch into; where llama's own prefetch engages it runs 224.1 GB/s < our prefetch-less 228.6 GB/s; both families sit at 75–84% of GB10's DRAM peak. No build, no commit.
  • B1 (__launch_bounds__(256,6)): SASS REG 48→40 + STACK 40 (10 spill words); probe tg128 +0.25%, @3254 −0.75% (22.54→22.37 medians) → KILLED — the spill cost exceeds the 6th block's gain.
  • B1c (rerouting npair-432 to the v2 loop, MINFER_Q6K_PF=0): 6 blocks × 1-unit-MLP against 5 blocks × 2-unit-MLP — v2_pf wins/ties (tg128 24.08 vs 24.02/24.08; @3254 median 22.28 vs 22.16, one noisy pair each) → v2_pf kept by default; the env is kept as a documented opt-out.
  • B2 (160-thread right-sizing, = the D3b-1c re-measurement): bitwise 98/98, but nsys per-kernel: lm_head 3243.7 → 3292.7 µs (+1.51%), attn_v 24.41 → 25.15 µs (+3.04%), v2_pf/q4_K rows ±0.6% (control group) → KILLED.

The B line's stopping rule (pre-registered in the brief): q6_K's occupancy gap cannot be closed within this kernel family under the bitwise constraint. The limiter record stands as originally judged: v2_pf holds 48 regs for the dual-unit pipeline, pinned by registers one block below the warp cap; cutting registers (B1) pays spill, cutting the pipeline (the B1c control) loses intra-block latency hiding, changing block geometry (B2) measures worse — all three roads were walked to the end, and the geometry already sits at the measured optimum. The remaining 0.42 ms of true deficit can only be reached by tolerance-gated redesign (numeric slack traded for layout freedom) or the attention structure lever, neither of which belongs to this session.

@3254 gap decomposition update (the brief handed to D4-3): (1) q6_K's "205 GB/s" census rate counts 210-byte blocks while the padded stride actually streams 224 bytes — the true DRAM rates are ≈ 218.9 (ffn_down) / 221.6 (lm_head) / 189.5 (attn_v), against the q4_K class's 228.6, making the true q6_K-vs-class deficit ≈ 0.42 ms/step, not 1.1–1.2; (2) the bitwise occupancy axis is fully closed (above), and realizing this 0.42 ms requires a tolerance-gated redesign; (3) the ~2.53 ms attention structure item is untouched and remains D4-3's prize (later corrected by D4-3 to ~1.6 ms — the anchor had stood on a llama-bench artifact); (4) 7B decode numbers before b31084c are incomparable with anything after it.

Lever C (PDL): recorded as follow-up, not built this session — the interaction risk with CUDA-graph capture needs its own session (llama attaches cudaGridDependencySynchronize + programmatic stream serialization to every decode kernel; the target is the ~0.3 ms launch-gap residual the graphs left behind). This item was formally closed in D4-4 L2 (see doc 76).

The ledger at session close: one correctness fix landed (7B is trustworthy from here on), four levers closed with measured mechanism (A pre-build, B1/B2 killed, B1c control archived), zero perf regressions, and the corrected @3254 gap decomposition handed to D4-3 — plus one process correction (master-table prior-art search moved up front) and one instrument rule (-n 1 first-step dump) entering the methodology library. This is exactly the expected output shape of a "Tier B" session: not every line gains performance, but every remaining millisecond has an owner.

6. Lessons

  1. A dispatch gate is a correctness surface, not a performance surface. Any dispatch of the form "the kernel only covers ≤ X" must have a dual upper bound; the coverage assertion belongs in the kernel comment with the host gate mechanically aligned to it (npair ≤ 512 ⇔ id ≤ 16384), or the next new shape is the next 7B.
  2. The -n 1 first-step dump is the only clean cross-binary comparable point. Dump differences after generation diverges = cascade noise + unwritten KV tail-slot garbage; a 1e15–1e18 "delta" is not corruption.
  3. After a correctness fix, anchors must be re-set. A guard set on a wrong engine is a negative asset (the old 7B guards flattered dropped work as speed); the −2.7% re-anchor is correctness's price and belongs in the record, not hidden.
  4. Grep the master table's mechanism keywords before writing kernel code. B2's prior art sat in the table for two sessions; pre-build arithmetic (Lever A's approach) is an order of magnitude cheaper than post-build measurement.

← 73 · Index · 75 →

75 · D4-3 — attention structure rewrite attempt 2 NO-GO + D4-1's llama target was a llama-bench artifact (CLOSED)

Result: the split-attention kernel vec_attn aligned to llama's fattn-vec geometry (five axes pb × threads × rows × minb × STG, a 690-row sweep, 0 correctness skips) best result 41.07 µs @14B/@3254 (bar ≤~32, 1.64× off the current dispatch kernel), 16.22 µs @7B/@1641 (bar ≤~10.35, 1.53×) — folded into in-situ wall clock that is +1.6–1.8% < the +2% integration bar, NO-GO; the attention line closes with a "measurement correction". The session's real output is the artifact identification: ncu proves llama-bench's own decode fattn-vec (grid (1,2,40)) loads a constant 5,427,200 B at any context ≈ one 128-row KV iteration per block = of 3255 KV rows only 256 covered (7.9%), while llama-cli's decode (grid (1,7,40)) loads 53.2/142.7 MB scaling with context (corroborated by a mid-context recall A/B). Honest llama decode attention ≈ 2.0–2.1 TB/s ≈ 1.7 ms/step — minfer's 3.32 ms is a ~1.9× gap, not 4.9×; 14B @3254's true gap shrinks from 13.4% to ~10% (attention ~3.5% of it). Commit: ce01048 (docs-only, no repo code change). Date: 2026-09-09.

1. Background — where things stood

The brief handed over at D4-2's close said: in one 14B @3254 step the attention structure accounts for ~2.53 ms, the largest single piece of the remaining gap. That number's provenance is D4-1's trace: llama's decode attention at 14B/@3254 recorded 6.94 µs split + 7.20 µs combine = 14.02 µs/layer per layer, 48 layers ≈ 0.67 ms/step; minfer's split+combine measured 3.32 ms/step. The reading at the time was "llama uses some structure we lack", and the byte ledger was spread out: 14B @3254's KV reads ≈ 48 layers × 66.8 MB per step (K+V f16, 3255 rows × 40 heads × 128 dims × 2 B × 2 tables), read inside 0.333 ms → ~9.6 TB/s effective L2 read rate — an order of magnitude above our attention kernels' 1.0–1.4 TB/s. The prize was estimated at 2.3–2.5 ms/step.

That reading already had a contradiction buried in it: 9.6 TB/s exceeds GB10's L2 fabric capability. But D4-1's conclusion kept it when written into the brief, because "llama did it" was the trace's direct output. D4-3's task was therefore twofold: (1) build a llama-geometry kernel per the brief and probe it against a pre-registered GO bar; (2) if it cannot be built, explain where llama's 9.6 TB/s came from — "our kernel structure is inadequate" and "the target itself was wrong" are two completely different follow-up roads.

The session ran probe-first with a hard go/no-go: write the probe first, register the bar first, run the sweep first; any integration only after a GO. The GO bar was pre-registered as: a kernel with llama's structure (pb splits × windows × subgroup lanes, Q resident in registers, K/V read straight from global, probabilities broadcast with shfl — no smem touched in the hot loop) must simultaneously reach ≥2× the current dispatch kernel on both 14B/@3254 (kernel-total ≤ ~32 µs) and 7B/@1641 (≤ ~10.35 µs); the integration bar was separately set at +2% wall clock (2% of a ~40 ms 14B step ≈ 0.8 ms). Both bars were written down before any numbers ran — that is why this doc can close cleanly on a "NO-GO".

2. Principle — the GPU mechanism

The current dispatch kernel's shape (the probe's incumbent baseline). minfer's decode attention is D3-4's hybrid rpw dual-kernel split: gqa_attn_split_partial_hybrid (4-warp form) or gqa_attn_split_partial (1-warp form) slices KV by ATTN_SPLITS, each block claims a segment and advances an online softmax row by row (each lane claims 4 consecutive dims, the full-row dot reduced with warp shfl), partials written to global; gqa_attn_split_combine merges. 14B/@3254 measures 63.8 + 3.6 = 67.4 µs (nsys per-kernel), 7B/@1641 is 20.7 + 4.1 = 24.8 µs. The probe copied this kernel family's device bodies verbatim into /tmp/d4/probe_attn2.cu as the baseline (attn_split_1w_body / gqa_attn_split_partial_hybrid / gqa_attn_split_combine), guaranteeing the sweep's control is the same math.

llama fattn-vec's geometry vs ours. llama's flash_attn_ext_vec (in ncu: flash_attn_ext_vec<128,1,F16,F16,false>) is organized as pb (parallel blocks) × KV windows: each block claims a window of KV (R rows), within the window K/V are read directly from global, Q is resident in registers, inter-row probabilities are broadcast with warp shfl, and window boundaries write partial results back for the combine to merge. Its block count = pb × n_head_kv (grid shaped (1, 7, 40): y dim 7 splits, z dim 40 KV heads), while we were fixed at 2 splits at the time. For a low-occupancy kernel at ≤2 blocks/SM, performance is decided by the SASS's load batching: in the natural load→dot→shfl interleaved loop, the next LDG waits for the current softmax dependency chain to finish before issuing; staging the whole window's K/V upfront (a run of back-to-back LDGs) lets the load latencies overlap each other — this difference measured 1.0 vs 1.7 TB/s in this probe (see §5).

Why 9.6 TB/s is a reductio, not a target. GB10's L2 bandwidth is on the order of ~6-7 TB/s and DRAM ~1.9-2.3 TB/s peak (both families' decode kernels sit in the measured 75–84% of DRAM peak = the 1.4–1.9 TB/s-class interval). If llama truly read 66.8 MB × 48 layers = 3.2 GB from L2 inside 0.333 ms, it would need 9.6 TB/s — beyond the fabric, physically impossible. Conversely, if it read only the bytes the trace's split kernel touched (5.43 MB × 48 = 260 MB ÷ 0.333 ms = 0.78 TB/s), it lands exactly in the latency-bound class all our decode kernels occupy. D4-3's core hypothesis test sits right here: the premise of the 9.6 TB/s ledger (that the traced kernels cover all the KV) must be directly verified or refuted by ncu's byte counts.

The artifact's mechanism hypothesis (the part left undecided). llama-bench's host side picks flash_attn_ext_vec's pb as 2 (grid (1,2,40)), and its split/window loop runs only about one 128-row iteration at all three contexts KV≈1024/2474/3255 — i.e. the bench path's kernel actually covers only the first 256 KV rows. Why ca3d5a3e1's occupancy loop picks pb=2 in that window while llama-cli's identical loop picks pb=7 — the record explicitly states this was not pinned down ("the observed facts above are unambiguous and reproducible; the binary is b10665-ca3d5a3e1 per its own banner"). This doc only pins down the reproducible byte facts and leaves the host-side selection logic for upstream to check.

3. Implementation

3.1 Design choices (why this shape and not another)

Probe structure: verbatim incumbent copies + a family of parameterized kernels. /tmp/d4/probe_attn2.cu (nvcc -O3 -arch=sm_121a) holds three kinds of things: (1) the incumbent bodies copied verbatim from src/cuda_kernels.cu (the baseline's credibility comes from "zero rewriting"); (2) the llama-geometry template kernel vec_attn<T,R,MINB,STG> (T = pb, R = window rows, MINB = minimum blocks per window, STG = the load-scheduling axis); (3) a double-precision CPU reference. The five-axis sweep: pb ∈ {2,4,8,16,32}, threads ∈ {32,64,128}, R ∈ {32,64,128}, minb ∈ {1,4,8}, STG ∈ {0,1,2} (0 = chunked staging every 8 rows; 1 = whole-window K+V staging; 2 = STG1 + next-window K prefetch).

The timing protocol aligned to decode's real shape: batch-of-32, min-of-8, with two rotating L2-hot KV copies (simulating decode's KV-resident-in-L2 state). The correctness gate 0.05 abs + adversarial outliers (q injection ±57 every 911 rows, v injection ±138 every 1543 rows) — the outlier injections specifically guard against the fake green of "softmax masking large errors on random data". All 690 sweep rows passed the gate, 0 correctness skips.

Code survival note for the probe: /tmp/d4/probe_attn2.cu and the sweep raw data /tmp/d4/d43_sweep5.csv are session artifacts, not in the repo; the probe's GO/NO-GO verdict and all key numbers were finalized in ce01048's docs commit (the same treatment as the r10/r11 precedent: a measurement session's first disk write is docs). The 3.2 excerpts below fall into two classes by content — the incumbent excerpts come from the current tree (the probe is verbatim identical to it), and the vec_attn form is reconstructed from the record.

3.2 Key code

Excerpt A · the incumbent's hot loop (src/cuda_kernels.cu 2803–2872 excerpt, the probe baseline = verbatim same lineage) — note the D2 comment: staging the 4-row window's K+V before computing was a load-scheduling correction bought at −11%/−42% by the D1/D2 probes; vec_attn's STG axis pushes exactly this idea to whole-window granularity:

// src/cuda_kernels.cu (current tree, attn_split_1w_body hot loop excerpt)
for (int base = lo; base < hi; base += 4) {
    int nr = min(4, hi - base); // warp-uniform
    // D2: stage BOTH K and V for the whole 4-row window before the first
    // softmax step. All 8 row loads then issue back-to-back and their
    // latency overlaps the serial chain; the old form relied on the
    // compiler hoisting the inline V loads, which it does not do across
    // the shfl/softmax dependency chain (D1 probe: −11% hot-L2, D2 probe:
    // −42% cold-DRAM vs inline V; bitwise-identical — same rows, same
    // order, same per-row ops, only the load scheduling changes).
    float4 k4[4], v4[4];
    #pragma unroll
    for (int j = 0; j < 4; j++) {
        k4[j] = (live && j < nr)
            ? kv_ld4<KV>(k + (size_t)(base + j) * stride_kv + hk * hd + d0)
            : make_float4(0.0f, 0.0f, 0.0f, 0.0f);
        v4[j] = (live && j < nr)
            ? kv_ld4<KV>(v + (size_t)(base + j) * stride_kv + hk * hd + d0)
            : make_float4(0.0f, 0.0f, 0.0f, 0.0f);
    }
    #pragma unroll
    for (int j = 0; j < 4; j++) {
        if (j >= nr) break; // warp-uniform: all lanes exit together
        float d = q4.x * k4[j].x + q4.y * k4[j].y + q4.z * k4[j].z + q4.w * k4[j].w;
        #pragma unroll
        for (int off = 16; off > 0; off >>= 1)
            d += __shfl_xor_sync(0xFFFFFFFF, d, off);   // full-row dot reduction
        float s = d * scale;
        float nmx = fmaxf(mx, s);                        // online softmax step
        float corr = expf(mx - nmx);
        float e = expf(s - nmx);
        S = S * corr + e;
        mx = nmx;
        if (live) {
            float4 vv = v4[j];
            oc.x = oc.x * corr + e * vv.x;  /* ... oc.y/z/w same form ... */
        }
    }
}

Excerpt B · the hybrid dual-kernel's self-gating dispatch (src/cuda_kernels.cu 2894–2908) — the sweep's third baseline (h4w) and the "current dispatch kernel" 67.4/24.8 µs both measure through this path; the gate's warp-uniform condition is the key to replay safety:

// src/cuda_kernels.cu (current tree, gqa_attn_split_partial dispatch arm excerpt)
// D3-4 L1 dual-kernel dispatch: when rpw_gate > 0 and the 4-warp kernel
// owns this nkv (rpw >= rpw_gate), exit before touching anything — the
// hybrid kernel writes the partial. The branch is nkv-uniform across the
// whole grid (positions[0] is launch-wide), so this stays replay-safe,
// and for every nkv the incumbent path takes, the arithmetic in the body
// is unchanged (bitwise; dump-memcmp gated).
if (rpw_gate > 0) {
    const int nkv0 = positions[0] + 1;
    const int chunk0 = (nkv0 + ATTN_SPLITS - 1) / ATTN_SPLITS;
    if (((chunk0 + 3) >> 2) >= rpw_gate) return;
}
attn_split_1w_body<KV>(q, k, v, partial, positions[0] + 1,
                       blockIdx.x, blockIdx.y, nh, nk, hd, scale, pstr,
                       threadIdx.x);

Excerpt C · vec_attn's reconstructed form (reconstructed from the probe record, not surviving code) — the skeleton of the llama geometry: a block claims (split, window), Q resident in registers, online softmax row by row after STG=1's whole-window staging; the difference from excerpt A is only the row-grouping granularity and the staging depth, the math being the same online softmax:

// /tmp/d4/probe_attn2.cu (session artifact; form reconstructed from the record)
template <int T_PB, int R_WIN, int MINB, int STG>
__global__ void vec_attn(const float* q, const f16* k, const f16* v,
                         float* partial, int nkv, /* … */) {
    // per block: 1 KV head × 1 window (R rows); Q lives in registers, K/V
    // read straight from global.
    // STG=0: chunked staging every CS=8 rows; STG=1: the whole window's
    // K+V staged in one go; STG=2: STG1 + next-window K prefetch.
    float4 kreg[R_WIN], vreg[R_WIN];
    if (STG >= 1) {
        #pragma unroll
        for (int r = 0; r < R_WIN; r++) {          // whole-window batched LDG —
            kreg[r] = ld_kv(k, base + r);           //  issued back-to-back, latencies overlap
            vreg[r] = ld_kv(v, base + r);
        }
    }
    for (int r = 0; r < R_WIN; r++) {
        float d = warp_dot(q, STG >= 1 ? kreg[r] : ld_kv(k, base + r));
        // shfl reduction → online softmax → P accumulates into oc (same form as excerpt A)
    }
    // write the partial at the window tail; the combine merges (same combine as the incumbent)
}

3.3 Pitfalls

  • Both probe bugs were caught by the correctness gate, not "explained away" by the numbers: (1) the subgroup row-map missed the warp stripe offset — every 32nd row's dot was misaligned; (2) one refactor dropped the chunked V-pass's initial staging — the STG=0 variant was slow and wrong overall. Both manifested as correctness FAILs (the outlier-injection design happened to amplify them), and the swEEP produced no "sick numbers". That is the value of the probe-first process: the correctness gate hangs inside the probe, so a bad kernel never survives to reach the main tree.
  • ncu byte counts must be read via the LDG instruction composition, not the total alone: the bench-path kernel issues 44 LDGs per warp ≈ K 16 + V 16 + Q 4 + mask 8 — exactly one 128-row KV iteration's worth. The reason "byte-identical at KV 1024/2474/3255" holds is that the three captures' LDG counts and address patterns are completely identical; if the total byte count does not change while the context changes 3×, that is far stronger artifact evidence than "slow".
  • Artifact identification takes three-way agreement to be final: ncu byte counts (5.43 MB constant vs 53.2/142.7 MB scaling with ctx) + behavioral A/B (mid-context passphrase recall) + reductio (9.6 TB/s beyond the fabric). Any single piece still leaves room for reinterpretation.

4. Verification

  • Probe correctness gate (0.05 abs + adversarial outliers + double CPU ref): defends against the sweep numbers coming from a mis-computing kernel; 690 rows, 0 skips.
  • Baseline same-lineage check: the incumbent bodies were copied verbatim from the current tree — defends against a fake GO/NO-GO caused by "the probe baseline is faster/slower than the real kernel".
  • The artifact identification's ncu evidence: the bench decode flash_attn_ext_vec<128,1,F16,F16,false> shows SASS global-load bytes 5,427,200 B bit-identical across the KV≈1024/2474/3255 captures, LDG/warp = 44; llama-cli decode (grid (1,7,40)) loads 53.2 MB at KV≈2477 (-c 2610) (≈ the 50.7 MB full-coverage arithmetic) and 142.7 MB at KV≈4869 (-c 0) — scaling with context.
  • Behavioral A/B (recall gate): the mid-context passphrase (KESTREL-5150, ~row 1200/2470) is answered correctly by llama-cli under both -c 0 and the bench-like -c 2610 — proving the CLI path's attention really reads the full context; llama-bench's tg output never performs any recall check, so it cannot notice on its own that the bench path drops 92% of the KV.
  • Reductio cross-check: taking the trace reading as full coverage, 48 layers × 66.8 MB ÷ 0.333 ms = 9.6 TB/s L2 — beyond the fabric, impossible; taking ncu's measured bytes, 5.43 MB × 48 ÷ 0.333 ms = 0.78 TB/s — the same class as the other latency-bound kernels. Both directions rule out "9.6 TB/s is real".

5. Results

Sweep results (all passing the correctness gate; split+combine totals, µs):

shapebest vec geometrysplitcombinetotalprobe baseline (1w/h4w)in-situ dispatch todayGO bar
14B @3254pb=4 T=128 R=64 STG=138.892.1741.0749.74 / 55.2667.4 (63.8+3.6)≤~32 → FAIL (1.64×)
7B @1641pb=16 T=32 R=32 STG=112.174.0416.2220.45 / 32.7824.8 (20.7+4.1)≤~10.35 → FAIL (1.53×)
7B @512pb=4 T=128 R=32 STG=18.042.1210.1611.61 / 24.59~11.8—

The geometry plateau is flat (pb 2–8 within 1.5 µs on 14B), and the STG axis enters noise on the plateau — once loading is batched per window, more scheduling depth buys nothing at this block-count level. vec_attn is indeed ~17–21% faster than the incumbent 1w baseline, but still 1.5–1.6× short of the bar.

The GO/NO-GO arithmetic (veto mechanism): the probe→in-situ conversion factor was calibrated on the h4w baseline (probe 63.8 / in-situ 51.2 = 1.25): 14B 41.07 ÷ 1.25 ≈ 48.5 + 3.6 combine ≈ 52 µs → 0.72–0.82 ms saved per step ≈ +1.6–1.8% wall clock < the +2% integration bar; 7B ≈ +0.7%. Both bars fail → no integration, the line closes. The honest remaining headroom (14B ~1.6 ms/step) requires ~2.1 TB/s effective rate — llama's own measured rate — corresponding to pb=7 (280-block)-class geometry plus a staged-load recipe beyond the sweep plateau; that is a new session with a new bar (≤25–32 µs @14B), not a continuation of D4-3.

THE ARTIFACT — the correction of D4-1's target (this doc's archival centerpiece):

capturegridSASS global-load bytesLDG/warpKV rows covered
llama-bench decode, KV≈1024 / 2474 / 3255(1,2,40)5,427,200 — bit-identical across all three44 ≈ one 128-row iteration (K 16 + V 16 + Q 4 + mask 8)2 splits × 128 = 256/3255 (7.9%)
llama-cli decode, KV≈2477 (-c 2610)(1,7,40)53.2 MB (≈ the 50.7 MB full-coverage arithmetic)—full
llama-cli decode, KV≈4869 (-c 0)(1,7,40)142.7 MB, scaling with ctx—full

The bench path kernel's load traffic is context-independent (KV 1k or 3.3k are both 68 KB per block), i.e. its 2-split grid runs only ~one 128-row iteration per block and covers only the first 256 rows; the CLI path picks 7 splits and full coverage. Recalibrated: llama's honest full-context decode fattn-vec @14B runs at ≈2.0–2.1 TB/s effective (53.2 MB/60 µs-ncu ≈ 25 µs live @2477; 142.7/171.5 ≈ 70 µs live @4869), i.e. ~33–36 µs per layer including combine at @3254 ≈ 1.7 ms/step, not 0.69. minfer's 3.32 ms/step is a ~1.9× gap, not 4.9×; 14B @3254's honest wall-clock gap to llama shrinks from 13.4% to ~10%, of which attention is ~3.5%.

Verdict: the D-series attention line closes with a "measurement correction". D4-1's census itself was not wrong (those two kernels really did run 48×6.94 + 48×7.20 µs — the ledger balances); the error was reading those two kernels as "full-context attention". The follow-up road is rewritten accordingly: (1) 14B decode's truly recoverable space against llama is ~10% of wall clock, of which attention is only ~3.5% and needs ~2.1 TB/s to reach llama's level; (2) future attention sessions start from (1,7,40)-class geometry + explicit load staging, with the bar set at ≤25–32 µs @14B; (3) no minfer-side code was changed by this session — ce01048 is docs-only, and the probe numbers and the artifact verdict are the entire output.

6. Lessons

  1. Never quote llama-bench's long-context tg rates as attention targets. The bench path's attention kernel can skip most of the KV and still produce a tg number (bench never validates recall); to check an attention target, use a llama-cli recall A/B or ncu byte counts.
  2. The reductio of "the target is impossible" should be done as early as possible. 9.6 TB/s beyond the fabric was a contradiction computable at the brief stage; doing the dimensional check before building the kernel would have saved half of this session's kernel sweep.
  3. A pre-registered bar is the NO-GO's talisman. 41.07 µs was the sweep's best number; without a bar, "39% faster" would have tempted an integration doomed to fall below the integration line. Bar written first, numbers run after — only then can a NO-GO close cleanly.
  4. Artifact identification needs a three-way evidence chain: byte counts (ncu) + behavioral validation (recall A/B) + a physical ceiling (reductio) — any single piece of evidence can be "reinterpreted"; only three-way agreement is worth rewriting a conclusion over.

← 74 · Index · 76 →

76 · D4-4 — the endgame kernel session: dpl dense split-plane q6_K decode MMVQ lands (bitwise); PDL and fused-FFN closed with mechanism (LANDED)

Result: L1 (LANDED, bitwise) — the padded 256-element/224B MMVQ row layout streams 14 dead bytes per super-block (recorded as 215/256 = 84% useful); dpl rearranges each row into [ql: nbe×128][qh: nbe×64][sc: nbe×16][d: nbe×2] = nbe·210 B of content with row stride (nbe·210+15)&~15 — same values + same accumulation order → bitwise. Probe: ffn_down 176.5 → 212.9 GB/s content (−17.1%), lm_head 208.3 → 250.0 (−16.7%). Wall clock (3× interleaved A/B medians): 14B tg128 23.28 → 24.57 (+5.53%) / @3254 22.00 → 22.94 (+4.27%); 7B tg128 47.55 → 51.20 (+7.68%) / @1641 46.44 → 50.18 (+8.05%) — the bar was 14B @3254 ≥ +0.4%, over-delivered 10×. L2 (PDL) probe green but in-situ NO-GO (same-binary env-flip: 14B tg128 −2.6%/−1.8%) → reverted; L3 (fused gate+up+SwiGLU+q8 Form B) probe +28.2% (wave quantization + per-row latency exposure) → NO-GO. Endgame vs-llama: 14B 1.018× tg128 / 0.950× @3254; 7B 1.074× / 1.052×. Commit: ffce151 (L1 code: src/cuda.rs +177, src/cuda_kernels.cu +101) + docs commit. Date: 2026-09-09.

1. Background — where things stood

This is the closing round of the decode campaign (D1→D4-4, 12 sessions). After D4-2 fixed 7B and closed the entire bitwise occupancy axis, the @3254 gap decomposition had converged to three rows: attention ~1.6 ms (D4-3's verdict: llama's true rate is ~2.1 TB/s and reaching it needs a separate session), the true q6_K-vs-q4_K-class deficit ~0.42 ms/step (true DRAM rates ffn_down 218.9 / lm_head 221.6 / attn_v 189.5 vs the q4_K class's 228.6; the bitwise occupancy axis proved unable to squeeze it out), and a ~0.3 ms launch-gap residual (CUDA graphs already banked most of it). D4-4's three levers each target one row: L1 hits the 0.42 ms (a layout lever), L2 hits the 0.3 ms residual (PDL), L3 hits FusedFFN's swiglu round trip (~106 µs/step execution + 2.2 µs launch).

L1's starting point is the precise thread D4-2 left behind: the padded layout itself is the deficit. q6_K's GGUF block is 210 B of content (ql 128 + qh 64 + sc 16 + d 2), and the historical layout pads it to a 224 B stride (a ggml_pad-alignment artifact serving prefill MMQ's uint4 alignment) — for every super-block it streams, the decode kernel pays DRAM traffic for 14 information-free bytes. Read in reverse, D4-2's conclusion that "the padded kernels are already optimal for their layout" says: change the layout, not the kernel math. The 84% figure means zeroing the dead bytes multiplies the q6_K class's theoretical stream-rate ceiling directly by 224/210 ≈ 1.067 — and the bulk of the 0.42 ms deficit sits exactly in that magnitude.

Session discipline carries over from D4-2: the baseline binary stored at /tmp/d4/minfer_pre_d44; all A/B are same-window interleaved 3 pairs taking the median (sglang co-tenant resident, ±1–2% window drift, pair medians decide); every lever probes for numbers first, passes the correctness gates, and only then discusses integration; L2/L3 both have explicit bars and veto paths.

The three levers' pre-registered bars are worth juxtaposing, because their width differences are themselves information: L1's integration bar was set at 14B @3254 ≥ +0.4% (the probe's −17% kernel-level signal folds into wall clock at roughly +0.5~1%, so the bar only demands "still positive after folding losses"); L2's bar is +0.3% (PDL's gain ceiling is ~0.3 ms/step, near-threshold to begin with); L3 has no wall-clock bar — a probe-level comparison decides life or death directly. Bar width is inversely proportional to confidence in the mechanism: for a layout lever with a "clear mechanism, trustworthy probe", the bar is loose; for a scheduling lever where "the probe may not represent in-situ", the deciding gate sits directly on the final shape's isolated A/B. This gradient proved fully right at the close: L1 over-delivered 10×, and L2/L3 were cleanly vetoed at their respective deciding gates.

2. Principle — the GPU mechanism

The padded layout's true cost on DRAM. Decode MMVQ is pure streaming: each thread reads one unit's ql/qh/sc/d + the matching q8 activation. Under the padded 224B stride, a super-block's 210 B of content spans 7 32B DRAM sectors (210 = 6.5625 sectors; the 7th sector holds only 18 B of content + 14 B of pad) — those 14 B are paid on every block, because the next block starts at the 224B boundary. The record summarized the effective byte rate as 215/256 = 84% useful (a different basis than the pure-content 210/224 = 93.75%: counting the d/scale sector fragmentation into the traffic, the effective fraction is lower). On either basis the conclusion is the same: 6–16% of the decode kernel's request stream is dead traffic, and this dead traffic is not a "necessary cost" of the bandwidth-saturated regime — it depends only on how the bytes are arranged.

The dpl split-plane's layout ledger. dpl (dense split-plane) lays each row's four components out contiguously: [ql: nbe×128][qh: nbe×64][sc: nbe×16][d: nbe×2], with nbe·210 B of content per row and row stride (nbe·210+15)&~15 (padded only to 16B alignment, not to a block boundary). 14B ffn_down (nbe 54): 54·210 = 11340 → stride 11344, only 4 B extra per row, 0.07 B amortized per block — the dead-byte rate falls from 6.25% to 0.035%. The key constraint is uint4 alignment: within the ql/qh sections every 16B is aligned (both 128 and 64 are multiples of 16), and row bases are 16B aligned (guaranteed by the stride), so every uint4 load in the kernel is untouched. Same values, same order → bitwise.

The two layouts side by side (nbe super-blocks, i the block index):

componentpadded (status quo)dpl (new plane)
ql (128 B per block)offset 0..128 within block isection base +0, offset i·128
qh (64 B per block)offset 128..192 within block isection base +nbe·128, offset i·64
sc (16 B per block)offset 192..208 within block isection base +nbe·192, offset i·16
d (2 B per block)offset 208..210 within block isection base +nbe·208, offset i·2
dead bytes between blocks14 B/block (210 → 224 stride)~0.07 B/block (diluted by the row-tail pad)
row stridenbe·224(nbe·210+15)&~15

On the kernel side the conversion happens only at the section bases (compile-time constants + an nbe multiplication); the intra-block offset formulas are unchanged byte for byte — that is where the "same mapping" invariant is implemented.

Why bitwise survives. Three layers of invariant: (1) every unit's operand values are unchanged — the rearrangement is a pure byte move, and the ql/qh/sc/d bytes' mapping to threads keeps its one-to-one correspondence through the section-base conversion; (2) the unit→thread mapping is unchanged (u = tid, tid+256 and the loop form's u += 256 are both untouched); (3) the accumulation order is unchanged (ascending u, the same q6k_unit_acc statement). The sufficient condition for floating-point bitwise is "same operands, same order"; all three layers hold, so both the probe and the unit tests do memcmp-level comparison rather than tolerance comparison.

L2's mechanism: PDL (Programmatic Dependent Launch). cudaLaunchAttributeProgrammaticStreamSerialization (PSS) lets the next kernel launch early, before the previous kernel has finished draining, with cudaGridDependencySynchronize() at its entry waiting for the dependency data to be ready — what is saved is the inter-kernel launch gap (the pool of ~2 µs/launch in the graph, ≈0.3 ms/step). The price is co-residency: the next kernel's waiting blocks occupy SM slots early and contend with the previous kernel's finishing wave. For compute-tail kernels (14B's attention h4w, lm_head — they are not pure-bandwidth streams) this tax is real money; the decode chain is a serial chain of 13 kernels and every interface pays this toll once.

PDL's relationship to D4-2's recorded "known risk" needs spelling out: D4-2's Lever C deferred PDL because of "the interaction risk with CUDA-graph capture". L2's probe attacked that risk first — attaching the PSS launch attribute to each kernel node during capture; capture+instantiate passed in one go, 200 replays were stable, and the PDL-graph output was bitwise against plain-eager execution (driver 580.173.02). That is, the risk D4-2 worried about was falsified, and what actually killed L2 was a different mechanism nobody anticipated then: the co-residency tax. The probe's control line of "24-kernel compute-bound chain +2.8%" was added as a completeness check and ended up supplying the whole lever's verdict — PDL's gain ceiling is the launch gap (~2 µs per interface), its cost is co-residency interference (potentially > 2 µs per interface for a compute-tail), and the ratio depends on the nature of the kernels on the chain, not on any PDL parameter.

L3's mechanism: wave quantization + latency exposure. Form B (fully fused) has a grid of one thread block per 32-value output q8 block → nf/32 = 432 blocks (the 27648×5120 shape), which at 6 blocks/SM occupancy is 1.5 waves — the second wave fills only half the machine, a structural waste of 10–17%. Meanwhile each block performs 64 serial row dot products, each paying a reduce-barrier chain at its end; without cross-row register staging, load latency is exposed row by row; with staging, register pressure doubles, occupancy drops to 5 blocks/SM, and the tail wave gets worse. Both horns point to the same conclusion: a fusion geometry that preserves bitwise values must lose on a 1.5-wave grid.

Spreading L3's arithmetic out, because it is the template for every "fuse XX" proposal. The incumbent pair (the gu-concat matmul + swiglu_quant_pad40) both have row-shaped grids: od rows × 256 threads, 27648 rows = 27648 blocks, far more than one wave, with negligible wave-quantization loss; the two kernels each independently saturate bandwidth. Form B bundles 32 rows into one block (64 256-thread row passes inside the block), dropping the block count to 432 — these 432 blocks must hide latency against each other, and their only overlap mechanism is occupancy; 1.5 waves means the machine idles 1/3 of its SM slots during the second wave. What the fusion saves (one f32 write-read round trip of ~32 KB/row + one launch of 2.2 µs) is far less than the 1.5-wave idling (~17% × the ~400 µs order). Numerically: +28.2% ≈ wave quantization ~+15% + per-row latency exposure ~+13%, matching the two mechanisms' independent estimates.

3. Implementation

3.1 Design choices (why this shape and not another)

L1: a sibling plane, not a replacement. The padded plane keeps serving three things: prefill MMQ (the NB-BT kernels at block_stride 224), the derivation source of the W_exp/W_dsc planes, and the dequant-f16/embed-gather fallback. dpl builds one extra {name}__dpl copy inside register_weight_q6k_padded, hung off a q6k_dpl map keyed by the padded weight's device pointer; decode dispatch consults the plane before the padded path — on a hit it runs the dpl kernel, on a miss it falls back to padded — and the MINFER_Q6K_DPL=0 opt-out semantics match the r60 family (default = the verified path, "0" = opt out and reclaim the memory). The shape gate id % 256 == 0 (exact nbe, the same precondition as the W_exp plane): every real decode shape (5120/13824/18944) satisfies it.

The memory ledger: the dpl plane duplicates the q6_K weights' entire content bytes (nbe·210 B per row + the row-tail pad) — 14B's q6_K weight set ~2.0 GB, 7B's ~0.9 GB. This is a direct trade of memory for DRAM traffic: GB10 has 128 GB of unified memory and 14B Q4_K_M's full weights are on the order of ~8 GB, so a +2.0 GB resident plane fits the budget; and the opt-out switch guarantees low-memory scenarios can keep only the padded plane (performance back to the D4-2 state, bitwise unchanged). Construction happens once at weight registration (host-side repack → register_weight onto device), so decode pays zero extra; the build cost is just a few memcpys at load.

The control eliminated at the probe stage: gs (group-split, rearranged q4_K-style into 16B groups) measured −7.9%/−10.7% in the probe (short of dpl's −17% class) → strictly dominated by dpl, not integrated.

L2: probe first, integrate later; the deciding gate is a same-binary env-flip. PDL's known risk (interaction with CUDA graph capture) was verified to "not hold" in a standalone probe before daring to touch the tree; after integration passed all the bitwise gates, the deciding gate used the same binary flipping MINFER_PDL for the A/B — stripping "code differences" out of the measurement entirely, leaving only the mechanism's effect.

L3: Form B measured first, Form A vetoed by arithmetic. Form B (fully fused, quantize included) is the probe of the gain ceiling: if even it cannot make money, Form A with its standalone quantize has even less of a chance; Form A is capped by arithmetic (~2–5 µs/step < bar) without burning probe time.

3.2 Key code

Excerpt A · building the dpl plane (src/cuda.rs 1545–1576, inside register_weight_q6k_padded) — the repack is four copy_from_slices: cut each 210 B block from the padded raw GGUF bytes and place them into the ql/qh/sc/d sections:

#![allow(unused)]
fn main() {
// src/cuda.rs (current tree = the block introduced by ffce151)
if Self::mmq_gate_on("MINFER_Q6K_DPL") && id % 256 == 0 {
    let nbe = id / 256;
    let dpl_row = (nbe * 210 + 15) & !15usize;   // 210B content per block, rows padded to 16B
    let raw_row = nbe * 210;
    let mut dpl = vec![0u8; od * dpl_row];
    for r in 0..od {
        let src = &data[r * raw_row..(r + 1) * raw_row];
        let dst = &mut dpl[r * dpl_row..(r + 1) * dpl_row];
        let (ql, rest) = dst.split_at_mut(nbe * 128);
        let (qh, rest) = rest.split_at_mut(nbe * 64);
        let (sc, dd) = rest.split_at_mut(nbe * 16);
        for ib in 0..nbe {
            let blk = &src[ib * 210..(ib + 1) * 210];   // raw 210B block
            ql[ib * 128..(ib + 1) * 128].copy_from_slice(&blk[..128]);
            qh[ib * 64..(ib + 1) * 64].copy_from_slice(&blk[128..192]);
            sc[ib * 16..(ib + 1) * 16].copy_from_slice(&blk[192..208]);
            dd[ib * 2..(ib + 1) * 2].copy_from_slice(&blk[208..210]);
        }
    }
    let dpl_name = format!("{name}__dpl");
    self.register_weight(&dpl_name, &dpl);
    // …the q6k_dpl map: keyed by the padded weight's device pointer, hanging the dpl pointer…
}
}

Excerpt B · the dpl side's unit load (src/cuda_kernels.cu 1673–1697) — field-for-field matched to the padded q6k_unit_load, the only difference being that the address formula changes from "intra-block offsets" to "section base + intra-block offset"; all uint4 loads' alignment is guaranteed by the row stride:

// src/cuda_kernels.cu (current tree = introduced by ffce151)
__device__ __forceinline__ void q6k_unit_load_dpl(
    int u, const uint8_t* __restrict__ wrow, const uint8_t* __restrict__ x8row,
    int nbe, Q6kUnitRegs* r
) {
    const int kbx = u >> 3, pair = u & 7;
    const uint8_t* ql_row = wrow;                      // the four sections each contiguous:
    const uint8_t* qh_row = wrow + (size_t)nbe * 128;  //  same values, same order,
    const uint8_t* sc_row = wrow + (size_t)nbe * 192;  //  only a different arrangement
    const uint8_t* d_row  = wrow + (size_t)nbe * 208;
    r->d = h2f(*reinterpret_cast<const uint16_t*>(d_row + (size_t)kbx * 2));
    r->sc0 = (float)(int8_t)sc_row[(size_t)kbx * 16 + 2 * pair];
    r->sc1 = (float)(int8_t)sc_row[(size_t)kbx * 16 + 2 * pair + 1];
    const int chunk = pair >> 2, g = pair & 3;
    r->shift = 2 * g;
    r->g = g;
    const uint8_t* qlp = ql_row + (size_t)kbx * 128 + chunk * 64 + (g & 1) * 32;
    r->qla = *reinterpret_cast<const uint4*>(qlp);       // row stride 16B aligned
    r->qlb = *reinterpret_cast<const uint4*>(qlp + 16);  //  ⇒ uint4 legal
    const uint8_t* qhp = qh_row + (size_t)kbx * 64 + chunk * 32;
    r->qha = *reinterpret_cast<const uint4*>(qhp);
    r->qhb = *reinterpret_cast<const uint4*>(qhp + 16);
    const uint8_t* x8 = x8row + (size_t)u * Q8PB;
    r->d8 = h2f(*reinterpret_cast<const uint16_t*>(x8));
    r->xw = reinterpret_cast<const uint32_t*>(x8 + 4);
}

Excerpt C · the dpl dispatch fast path (src/cuda.rs 4234–4267 excerpt, top of q6_k_decode_mmvq) — the pf-vs-loop shape gate is verbatim the same as the padded side's (including D4-2's id ≤ 16384 upper bound and the MINFER_Q6K_PF semantics):

#![allow(unused)]
fn main() {
// src/cuda.rs (current tree = introduced by ffce151)
// D4-4 L1: dense split-plane fast path (bitwise — see the
// registration comment). Falls through to the padded kernels when
// the plane is absent (MINFER_Q6K_DPL=0 / id not a multiple of
// 256 / map miss).
if blk_stride_padded && Self::mmq_gate_on("MINFER_Q6K_DPL") {
    let dwp = self.q6k_dpl.lock().unwrap()
        .get(&(wptr as usize)).map(|cp| cp.0);
    if let Some(dwp) = dwp {
        unsafe {
            let nbe = (id >> 8) as i32;
            if id > 8192 && id <= 16384
                && !std::env::var("MINFER_Q6K_PF").map_or(false, |v| v == "0")
            {
                launch_q6_k_q8_mmvq_v2_pf_dpl(/* … */, nbe, stream);
            } else {
                launch_q6_k_q8_mmvq_v2_dpl(/* … */, nbe, stream);
            }
        }
    }
}
// …a miss (or the gate off) falls through naturally and continues to the padded dispatch…
}

The kernels themselves (q6_k_q8_mmvq_v2_pf_dpl / q6_k_q8_mmvq_v2_dpl) differ from the padded versions only in the row_stride formula and the q6k_unit_load_dpl call, with both the dual-unit/loop forms complete — the npair-432 class takes pf (with D4-2's upper bound), the rest take the loop.

Excerpt D · the bitwise unit test's A/B skeleton (src/cuda.rs from 5371, cuda_q6k_dpl_bitwise excerpt) — the same raw bytes registered twice (once with dpl on, once with MINFER_Q6K_DPL=0), and both kernel forms (pf shape od 512/id 8960, loop shape od 4096/id 1024) must produce bit-exact output:

#![allow(unused)]
fn main() {
// src/cuda.rs (current tree, unit test excerpt)
std::env::remove_var("MINFER_Q6K_DPL");
st.register_weight_q6k_padded("d44_dpl_a", &raw, od, id);   // build the dpl plane
std::env::set_var("MINFER_Q6K_DPL", "0");
st.register_weight_q6k_padded("d44_dpl_b", &raw, od, id);   // pure padded
std::env::remove_var("MINFER_Q6K_DPL");
/* ...run decode MMVQ on both sides, comparing outputs byte for byte... */
assert!(a == b, "dpl-vs-padded decode outputs must be bit-identical (od {od} id {id})");
}

3.3 Pitfalls

  • The dpl row base is not naturally 16B aligned. The first probe version used row stride = nbe·210: at nbe 54 the row bases alternate onto 2B alignment and the uint4 loads fault on misalignment outright. The fix is the (nbe·210+15)&~15 stride — at most 15 B of pad per row (0.07 B amortized per block) in exchange for every uint4 being legal. The general form of this trap: intra-section alignment must be backstopped by the row stride; when changing a layout, the alignment responsibility moves from "inside the block" to "the row tail" and must be settled explicitly.
  • The NaN trap in bitwise checking. When the probe filled d with random f16 bytes, the NaNs in the two outputs each compared NaN != NaN and the memcmp-class comparison read DIFF with max|Δ| = 0 — a fake difference. Fix: synthesize d as 0x3C00 (1.0) rather than random bit patterns. Any new quantization path with a bitwise gate should first screen random data for "which bit patterns become NaN".
  • L2's warning signal appeared in the probe, not after integration: the 24-kernel pure compute-bound chain got +2.8% slower under PDL (the bandwidth chain −0.1%, neutral) — the co-residency tax had already shown itself in the probe. The integration's env-flip merely confirmed it as −2.6%/−1.8% on the real decode chain. Lesson: do not expect a negative probe signal to turn positive after integration; it is the same mechanism.
  • The suite's window flakes are archived as "isolated rerun green": one full-suite run failed cuda_graph_recaptures_on_pool_gen_change + cuda_q4_0_prefill_q8_0_gemm_parity — caused by the co-tenanted window; both passed in isolation and on rerun, archived as window flakes rather than defects. The adjudication basis is "the failure mode has no mechanistic connection to the change under test".

3.4 L2's integration shape (reverted, form archived)

PSS is not a one-line switch — the integration points sit on both sides of each decode-chain kernel's launch and entry. Launch side: at graph capture, cudaLaunchAttributeProgrammaticStreamSerialization is attached to each of the 13 kernel nodes (attribute value 1, effective only on sm_90+, which GB10 satisfies); kernel side: the first line of every PSS kernel's entry calls pdl_sync (a thin wrapper over cudaGridDependencySynchronize(), a no-op on non-sm_90 compile targets). The 13 kernels cover the whole decode chain: rms (2 sites) / swiglu / activation quantize / add / the matmuls / rope-store / attention split + combine. The MINFER_PDL env gates the whole attribute set — this is precisely the precondition for §4's "same-binary env-flip": flip the switch and the same binary runs the same graph once with and once without PDL.

The revert itself was also a checkpoint: the MINFER_PDL=1 path lives on in the historical commits and the tree keeps only L1 (the PSS attribute code removed wholesale, no dead switch left) — consistent with the D series' convention of "reverted = clean tree + complete record".

3.5 L3's Form B kernel shape (probe, never entered the tree)

Form B's shape, archived (from the probe record): grid = nf/32 blocks; each block first computes the 32 gate rows' dot products with the verbatim 256-thread row-unit mapping (q4_K bitwise dots, in registers), then the 32 up rows; silu(g)·u completes in registers; finally the in-block call to the verbatim quantize_pad40_block produces the q8 — the numeric path is identical to the split pair, which is the basis for verifying that it "loses only on scheduling, not on math". The probe measured it same-shape (27648×5120 q4_K) and same-protocol against the incumbent pair (gu-concat matmul + swiglu_quant_pad40): 412.1 vs 321.4 µs. Form A (fusing gu+swiglu → f32, keeping the standalone quantize) was not probed: the round trip it saves is a strict subset of Form B's (one fewer f32 write-read), and with Form B losing 28%, Form A's arithmetic ceiling of ~2–5 µs/step cannot even reach the bar.

4. Verification

The gate chain for L1/L2/L3 each (one sentence per gate on what it defends):

gatecoverswhat it defends
probe bitwise (/tmp/d4/probe_l1_dpl.cu, memcmp vs padded)L1, both shapes ffn_down + lm_headthe layout change touching the numeric path
unit test cuda_q6k_dpl_bitwise (in the suite)L1, both kernel forms (od 512/id 8960 pf-form, od 4096/id 1024 loop-form)dispatch branches outside the probe shapes drifting
14B -n 1 first-step dump (107 identical + 7 node{N})L1integration touching numerics or dispatch behavior (node{N} is D4-2's documented pool-slot instrument class, reproduces pre-vs-pre)
7B -n 1 first-step dump (72 + 2)L1same, 7B side
greedy rp=1.0 byte-for-byte (both models)L1the intermediate state of "dump right, generation drifting"
suite 174/0/3 (with the new dpl test)L1full-model regression; one window flake archived as isolated-rerun green
PDL probe (capture/replay/200 replays/PDL-graph vs eager bitwise)L2the CUDA-graph interaction risk (outcome: did not materialize)
same-binary env-flip isolated A/B (MINFER_PDL, 3 pairs)L2's deciding gatecode differences leaking into the measurement; only the mechanism's effect remains (readings −2.6%/−1.8% → reverted)
Form B probe same-shape against the incumbentL3a mathematically correct fusion losing money on scheduling (+28.2% → NO-GO)

5. Results

L1 wall clock (3× interleaved A/B medians, pre = /tmp/d4/minfer_pre_d44, post = the final tree):

ConfigprepostΔ
14B tg12823.2824.57+5.53%
14B @325422.0022.94+4.27%
7B tg12847.5551.20+7.68%
7B @164146.4450.18+8.05%

The bar was 14B @3254 ≥ +0.4%; measured +4.27%, over-delivering 10× — the bulk of the 0.42 ms true deficit was cashed by the layout lever in one move (a 14B @3254 step is ~40 ms, and +0.94 ms lands right in the magnitude of the 0.42 ms q6_K deficit plus knock-on gains). Probe level: ffn_down (od 5120, id 13824) 176.5 → 212.9 GB/s content (kernel mean time −17.1%), lm_head (od 152064, id 5120) 208.3 → 250.0 (−16.7%). After L1, 14B's q6_K true DRAM class: ffn_down ~213, lm_head ~250 GB/s — the q4_K decode class is 228.6, and the residual gap is the explanation space of that recorded 84%-useful ratio applied in reverse to lm_head (lm_head's dpl stream rate is already above the q4_K class). Memory cost: +2.0 GB (14B), +0.9 GB (7B); MINFER_Q6K_DPL=0 can always back out.

Reconciling the wall clock with the probe's magnitudes: what dpl cuts is the request-stream dead bytes of q6_K's three shapes; in one 14B @3254 step these three shapes' kernel time totals on the order of 4 ms, so −1617% ≈ −0.65 ms, plus the L2 sector-hit improvement and tail-wave tweaks from removing the dead-byte sector effect, landing in the measured +0.94 ms wall clock is self-consistent; 7B's ratios are higher (+7.68/+8.05%) because 7B's q6_K ffn_down takes a larger share of the per-step weight stream (10 layers × id 18944, where 14B is 48 layers × id 13824 with ffn_down only part of that).

L2 (PDL) closing numbers: the probe was all green — on driver 580.173.02 the PSS-attributed graph capture+instantiate worked, 200 replays were stable, and the PDL-graph was bitwise against plain-eager (the graphs interaction risk D4-2 worried about did not materialize); but in the same probe the 24-kernel compute-bound chain ran +2.8% slower (the bandwidth chain −0.1%, neutral) as the first warning. The integration (all 13 decode-chain kernels carrying PSS + the entry pdl_sync, gated by MINFER_PDL) passed all the bitwise gates and finally lost at the deciding gate: 14B tg128 −2.6%/−1.8%, 7B ≈ 0, @3254 within noise → reverted. Mechanism: PSS lets the next kernel's waiting blocks co-reside on the SM with the current kernel's finishing wave — 14B's compute-tail kernels (attention h4w, lm_head) lose more slots to the contention than they save from the launch gap; the target pool (~2 µs/launch ≈ 0.3 ms/step) had already been mostly banked by D3-4's CUDA graph capture. Retry conditions: worth another look when the decode chain becomes purely bandwidth-dominated (e.g. attention is no longer a compute-tail) or the launch gap grows (leaving graph capture).

L3 (fused gate+up+SwiGLU+q8, Form B) closing numbers: probe (/tmp/d4/probe_l3_fuse.cu, 27648×5120 q4_K) +28.2% (412.1 vs 321.4 µs; 193.2 vs 247.8 GB/s content). Decomposition: grid = nf/32 = 432 blocks = 1.5 waves at 6 blocks/SM (wave quantization, structural ~+10–17%) + 64 serial row dot products per block each paying the reduce-barrier chain (cross-row register staging would hide it, but staging doubles registers → 5 blocks/SM and a worse tail wave). Form A (fusing only gu+swiglu → f32, keeping the standalone quantize) has an arithmetic ceiling of ~2–5 µs/step, below the bar, not probed. The line's closing ledger: the swiglu round trip is a real cost (~106 µs/step execution + 2.2 µs launch), but every fusion geometry that preserves bitwise values loses more to wave quantization than it saves — unless tolerance gating is accepted (changing the numeric path), this line does not reopen.

Endgame vs-llama (q4_k_m, llama-bench ca3d5a3e1):

ModelConfigminfer (D4-4)llamaratio
14B (48L)tg12824.5724.141.018×
14B (48L)@325422.9424.140.950×
7Btg12851.2047.651.074×
7B@164150.1847.691.052×

The decode program's endgame state: 7B at or above parity on every measured shape; 14B short-KV above parity and 5% off at @3254 — the remaining honest gap is attention (~1.9×, D4-3's measurement correction; the next lever is a tolerance-gated attention rewrite, a separate session per D4-3's bar). The whole D series' (D1→D4-4) net movement: 14B tg128 22.81 → 24.57, @3.3K 20.9 → 22.94; 7B tg128 → 51.20 — ahead of or at parity with llama on every measured shape.

Each lever's archived state at the close, so the next session can pick up directly:

leverstatuscarrier on the treereopen conditions
L1 dpl q6_KLANDED (bitwise)ffce151, MINFER_Q6K_DPL ("0" opt-out)none — the bulk of the 0.42 ms deficit is already realized
L2 PDL decode chainprobe green / in-situ revertednone (form in 3.4)the chain becomes purely bandwidth-dominated, or the launch gap grows (leaving graph capture)
L3 fused gu+SwiGLU+q8Form B probe NO-GOnone (form in 3.5)revisit when tolerance gating is accepted (numeric path loosened)
attention rewriteline closed (measurement correction)—new session: (1,7,40) geometry + explicit staging, bar ≤25–32 µs @14B

6. Lessons

  1. Layout is a first-class decode-GEMV lever, ranked ahead of kernel math. Same kernel, same 48 regs, same occupancy; swapping only the 224B stride for 210B+16B alignment moved all four wall clocks +4.3~8.1% — first count how many bytes in the request stream are dead, then talk microarchitecture.
  2. Bitwise can be designed for; it is not luck. Same values (a pure byte move) + same mapping (unit→thread untouched) + same order (ascending accumulation) — with all three invariants in place, memcmp is the gate; only when a layer loosens is tolerance needed.
  3. Do not expect a negative probe signal to turn positive after integration. PDL's +2.8% compute-chain warning and the post-integration −2.6%/−1.8% are the same mechanism (the co-residency tax); seeing the reversal in the probe should have been the stop — the integration cost (launch attributes + gating for 13 kernels) could have been saved.
  4. The same-binary env-flip is the strongest A/B for isolating a mechanism's effect. No code difference, no compile difference — what is measured is the mechanism itself; any optimization with an env switch should use this move as its deciding gate (L1's MINFER_Q6K_DPL=0 can serve as the same-style recheck at any time).

← 75 · Index · 77 →

80 · D5-0 — speculative-decoding cost model: measured baseline, acceptance, and the d=2 gate (MEAS-ONLY)

Result: draft-simple speculative decoding on GB10 is conditionally viable in the d=2 regime only, and the go/no-go now rests on one measurable number. Measured: 7B q4_k_m target 54.3 tok/s, 0.5B q4_0 draft 342.2 tok/s (CUDA) / 73.3 tok/s (CPU); real greedy acceptance of the 0.5B-on-7B pair (via llama.cpp speculative-simple) p ≈ 0.68–0.70 across prose and code prompts and both draft quants. Break-even requires per-position acceptance p* = 0.73 (d=2) / 0.81 (d=4) / 0.90 (d=8) — measured p is below the line at every d unless the target's nt=d+1 verify batch achieves the BT-MMQ amortization. At d=2 the requirement is ≥ 2.5× amortization at nt=3 (the D4-1 anchor is 2.7× at nt=4); projected outcome at the anchor is 1.04–1.05×, ceiling ~1.2×. The CPU-draft cross-device variant is dead by measurement (73.3 vs 54.3 tok/s = 1.35×, no break-even at any p or d). Commit: docs only (measurement only; artifacts in /tmp/d50/, ephemeral, key numbers inlined here). Date: 2026-09-10.

1. Background — where things stood

The D4-4 session closed the decode-kernel campaign with the note that "nt=4 → BT-MMQ 2.7× cheaper/step = the spec-decode foundation" (step doc 08's 0.14-wave collapse at M=1 is the same fact read from the other side). The D5 plan (SPECULATIVE-DECODING-PLAN.md) made D5-0 the go/no-go gate: measure the cost model before building any engine plumbing. Three questions had to be answered with numbers, not priors:

  1. what does the draft model actually cost per token on dgxspark (same-GPU and cross-device)?
  2. what acceptance rate does a real 0.5B-on-7B same-family pair sustain?
  3. what does the target's verify batch have to cost for the round to win?

2. Principle — the GPU mechanism

One speculative round with draft length d and per-position conditional acceptance p:

  • draft phase: d+1 serial nt==1 decode steps on the draft model (latency-bound; on the same GPU they are pure added latency);
  • process re-eval: the draft must ingest the accepted target tokens to sync its KV (llama's measured ≈ 2(1+d) draft token-evaluations factor);
  • verification: ONE target forward at nt = d+1 — every row needs logits (n_out = d+1), and because the batch rides the decode graph's BT-MMQ path, its per-token cost is NOT C_T(1): this is exactly the regime where D4-1 measured 2.7× per-token amortization at nt=4;
  • output: 1 + E[a] tokens, E[a] = Σ_{i=1..d} p^i for constant p.

Round cost ≈ C_T(d+1) + 2(d+1)·C_D(1); break-even against the baseline 1/C_T(1) solves for the required p given measured C_T, C_D — or for the required verify amortization given measured p. Both directions are used below. The subtlety the gate turns on: where between M=1 (0.14-wave collapse) and M=4 (2.7×) does M=3 land? Linear interpolation of the win says ~2.28× (below the 2.5× requirement); tile-regime step behavior could give the full 2.7× (M=3 and M=4 may hit the identical tile configuration). Only a measured C_T(3) settles it.

3. Implementation

No minfer code changed. The measurement harness:

  1. minfer cost battery (3 interleaved rounds, -p 0 -n 128 -r 1, medians, GPU idle at 0% util, backend engagement confirmed by the CPU-vs-CUDA spread below):

    ./target/release/minfer bench -p 0 -n 128 -r 1 -o json <gguf>          # CUDA
    MINFER_DISABLE_CUDA=1 ./target/release/minfer bench ... <gguf>          # CPU
    
  2. acceptance measurement — reused llama.cpp's binary instead of building the minfer loop first (the acceptance rate is an engine-independent property of the model pair at greedy: draft argmax vs target argmax along the target's own trajectory):

    llama-speculative-simple -m <7B q4_k_m> -md <0.5B draft> \
      -ngl 99 -ngld 99 --spec-type draft-simple --spec-draft-n-max <d> \
      --temp 0 -n 256 -p <prompt>
    
  3. cost model (/tmp/d50/cost_model.py, ephemeral): break-even solver in both directions + projection grid; key formulas inlined above and in §5.

4. Verification

  • Backend engagement: 0.5B q4_0 reads 342.2 tok/s (CUDA) vs 73.3 (CPU, 20-thread Grace) — the 4.7× spread confirms the CUDA path was measured.
  • Baseline cross-check: llama-bench tg128 on the same 7B = 49.57 ± 0.09 tok/s; minfer's 54.3 = 1.096×, inside the campaign's post-r60 band (1.074×/1.018× were the D4-4 numbers) — the baseline is trustworthy.
  • Acceptance stability: per-position conditional p back-solved from the aggregate rates is consistent across draft quant (q4_0 42.6% vs q4_k_m 48.9% aggregate at d=4 → p ≈ 0.685 vs 0.68) and workload (prose 42.6% vs code 44.1% aggregate) — no cherry-picked regime.
  • Model: the greedy acceptance measured on llama.cpp transfers to minfer because minfer's 7B greedy output already matches llama's (campaign gate) and the draft is the same family; residual engine differences affect the verify cost (modeled separately), not the acceptance.

5. Results

Measured medians (3 interleaved rounds, tg128, GB10, quiet):

configtok/s
7B q4_k_m CUDA (target)54.3 (C_T(1) = 18.42 ms)
0.5B q4_0 CUDA342.2
0.5B q4_0 CPU73.3
0.5B q4_k_m CUDA365.0
0.5B q5_k_m CUDA321.8

Measured acceptance (llama.cpp, greedy, 0.5B q4_0 draft unless noted):

dpromptaggregate acceptconditional p
2prose140/238 = 58.8%≈ 0.69
4prose164/385 = 42.6%≈ 0.685
8prose173/692 = 25.0%≈ 0.65
4prose (q4_k_m draft)172/352 = 48.9%≈ 0.68
4code83/188 = 44.1%≈ 0.69

Break-even (minfer measured costs, D4-1 anchor 2.7×/token at nt=4, saturating; draft cost = 2(d+1) token-evals):

drequired p*measured pverdict at anchor
20.730.68–0.69marginal lose → needs verify amortization ≥ 2.5× at nt=3
40.810.685lose; required amortization ≥ 4.5× at nt=5 — implausible
80.900.65lose decisively

Projected d=2 wall-clock vs the 54.3 tok/s baseline, as a function of the (measured-later) nt=3 verify amortization:

amortization @nt=3p=0.68p=0.69p=0.75 (optimistic)
2.0×0.87×0.88×0.94×
2.7× (anchor)1.04×1.05×1.12×
3.5×1.18×1.20×1.28×

Verdict: conditional go, d=2 only. The single gate variable is minfer's measured C_T(3) — the D5-1 phase is re-ordered to build the n_out = d+1 verify-batch primitive FIRST and measure it before any of the loop plumbing. Stop rule: if measured nt=3 amortization < 2.5×, the campaign stops after the primitive with the negative documented. Honest ceiling even on success: ~1.05–1.2× decode tok/s — the risk the plan pre-registered.

6. Lessons

  • Measure the gate variable with someone else's binary: the acceptance rate — the number the whole gate hangs on — came from llama.cpp's speculative-simple in minutes, with zero minfer code. Building the loop first to "find out" would have been days for the same number.
  • The cross-device fallback died by measurement, not argument: 73.3 vs 54.3 tok/s (1.35×) kills the CPU-draft idea at any acceptance — an assumption ("draft several times faster") that survived until it met a number.
  • Interpolation is not measurement: the linear interpolation of the BT-MMQ win (nt=3 ≈ 2.28×) sits BELOW the 2.5× requirement while the tile-regime step hypothesis puts nt=3 at the full 2.7×. The two hypotheses disagree about the entire fate of the campaign, and only the primitive measurement arbitrates. This is why D5-1's first deliverable is a number, not a feature.
  • Acceptance is workload-stable within a model family (p ≈ 0.65–0.69 across prose/code/draft-quant here) — but it decays hard with depth (aggregate 58.8% → 25.0% from d=2 to d=8), which is why deeper drafts lose twice: more draft steps AND a worse mix.

← 79 · Index

Speculative Decoding (D5-R) — Plan

Status: D5-R CLOSED (2026-09-12, docs 83–88): all stages landed or closed by measurement. 14B d=2 = 1.42×/1.68× (95%/88% of llama's absolute speed, same-window); d=8 retired (0.63× e2e); graph-capture prize verified already banked (R3-B); open leads localized (doc 89): row marginal = q4_K 2.4 + q6_K 1.0 + norm/quant 0.7 ms/row; fix menu re-priced by measurement (doc 90): R-rows-per-block CLOSED (~0 at d=2 — chain gives act rows L2-hot; only small-M tensor-core mma remains as a lever of size), chain hygiene de-priced (~0.1 ms/row); mma audit (doc 91) + K-split (doc 92): BT GEMM is tensor-core; K-split shipped and auto-ksplit enabled on the default path by decision (doc 92 §3b: default C_T(9) 86.4→73.0; multi-MMVQ keeps nt 2..8, residual is a tile-geometry redesign; d=8 needs code-p≈0.755, measured 0.731); doc 93/94: greedy identity fixed (spec output now byte-identical to sequential: batched verify attention made bitwise position-invariant + the spec penalty window capped to repeat_last_n) and the d=8 door crossed with the q4_k_m draft (code 76.5% ≥ 0.755: d=8 = 46.6 tok/s vs d=2 43.6, 96% of llama; prose stays d=2 — deep acceptance collapses); doc 95: adaptive draft depth shipped (--spec-draft-adaptive) — per-round d from beta-smoothed per-depth acceptance + min-of-4 cost windows, 10% switch hysteresis, optimism for unseen depths; identity boundary pinned: verify nt ≤ 8 bitwise (multi-MMVQ), nt=9 (BT-MMQ lm_head) tolerance-class → adaptive capped at d=7; all four gates pass and adaptive beats the best static on both code cells (46.5/48.4 vs 43.9/46.6 tok/s — the true optimum dmean ≈ 3.5 was never in the static sweep, which a CLI bug had silently pinned to d=8); doc 96 Phase 0: the nt=9 verify path profiled — BT-MMQ = 84% of GPU time, smem scoreboard stalls 40–59% inside the BT kernels (stop-gate ≥30% → tile/staging work may proceed, re-priced separately; the bitwise-preserving lever is cp.async double-buffered staging); doc 97: conversation/server integration LANDED — --cnv --spec-draft and serve --spec-draft are live with byte-identical identity gates (all frontends now share one spec implementation; a pre-existing plain-server cross-request KV-contamination bug was found and fixed in passing). doc 98: the 96-Phase-1 cp.async staging lever measured NULL → reverted. doc 99: the fast-verify P0/P1 pass corrected the stall attribution (L1TEX latency at a register-file-capped 16 warps/SM — not smem, not staging, not scheduling), landed one bitwise micro-fix (B-plane XOR swizzle), measured three further scheduling levers NULL, and re-priced the tolerance-class knob down to single-digit percent → not built. doc 100: the last priced lever (aligned qs plane) measured timing-NULL under interleaved A/B, and exposed a ±7–12% clock-ramp band that invalidates sequential benchmark comparisons on dgxspark. The performance line stays closed — now with the mechanism mapped and the measurement method corrected (doc 101: time-budget warmups; the headline table re-measured tight — adaptive 35.9 prose / 44.8 code tok/s, 1.42×/1.77× sequential). The draft-scale sweep (doc 102) closes the last speculative-quality axis: bigger drafts lose (acceptance is bounded by the target's predictability), so the 0.5B Q4_K_M draft stays; cross-family drafts are identity-safe (4/4 byte-identical vs sequential) and the sweep flushed a latent qwen3-loader namespaced-registration bug. Multi-MMVQ verify + adaptive d (doc 95) is the final production configuration; doc 103 extends the same verify machinery to q8_0/q4_0-weighted targets (the q8_0 verify batch had been running the f32-activation kernel — 7B Q8_0 spec e2e jumps 26.8 → 68.0 tok/s, +154%); doc 104's p32 planes lift it again to 79.0 tok/s = 2.47× sequential. Speculative decoding reopened by decision after the doc 81 §4.3 errata invalidated the original closure's external pillar and doc 82 restored the batching invariant. The previous plan (closed 2026-09-10, "no loop plumbing will be built") is retired; its errors are recorded in the Appendix.

The reference study is LLAMA-CPP-SPECULATIVE-ANALYSIS.md (llama.cpp draft-simple, source-verified: speculator framework §3, draft model setup §4, drafting loop §5, verification §6, KV rollback §7, cost model §9, minfer port notes §11). Records: docs 80, 81, 82.

1. Why reopened — the corrected evidence

Every number below is measured (docs 80–82 + the doc 81 §4.3 corrected batteries):

  • Doc 82 fixed the dispatch hole: verify amortization at nt=3 went 0.52× → 2.14× (14B q4_k_m, C_T(3)=56.6 ms vs C_T(1)=39.0); the batched verify is now a weights-once round, not a per-row restream.
  • llama.cpp's own speculative decoding demonstrably wins on GB10: 1.53–1.60× (7B) / 1.86–2.43× (14B) with the generic Qwen2.5-0.5B q4_0 draft (the original "1.00× external anchor" was a measurement artifact — see Appendix).
  • The engine-independent terms are known: per-position greedy acceptance p ≈ 0.74 (14B pair) / 0.69 (7B), draft cost C_D ≈ 2.9 ms/token — identical for both engines, as is C_T(1) (minfer 39.0 vs llama 41.8 ms on the 14B).
  • Predicted minfer economics (measured components, doc 81 §4.3 addendum): d=2 → 1.34× on the 14B from plumbing alone today; d=8 → 1.09× until the verify marginal shrinks. The entire gap to llama is the batched-verify marginal (minfer 3.9/7.8 vs llama ~1.0–1.5 ms per extra token).
  • The original go/no-go bar was mis-derived (doc 82 §5): it took the per-token marginal as ε. With the corrected cost model, break-even acceptance at d=2 with an efficient verify is p* ≈ 0.18–0.35 — the measured p clears it comfortably.

2. Scope

In scope (D5-R): draft-simple, greedy first

One draft GGUF (Qwen2.5-0.5B q4_0) drafting greedily for a target (Qwen2.5-14B q4_k_m primary — the 7B is marginal at d=2 — verified in one nt = d+1 batched forward through the existing forward_graph_cached primitive). d=2 is the gate configuration; d=8 the end state.

Deferred (not in D5-R)

draft-mtp / draft-eagle3 / draft-dflash (a trained draft is the obvious follow-up lever — llama.cpp natively supports all three spec types — but each needs target-side feature seams or new architectures); n-gram stretch; temp>0 target-authoritative verification; server multi-slot.

3. Mechanism — one round

With draft length d and accept count a (greedy: keep while the target's argmax equals the drafted token):

draft phase     draft model: d sequential nt==1 forwards (second GraphCache)
verification    target: ONE nt=d+1 forward, n_out=d+1 (rows: next_token +
                the d drafted candidates, positions contiguous)
accept/cut      greedy argmax per verify row; accept while equal to the
                drafted token; always emit ≥ 1 token (the deepest accepted
                row's logits = the bonus)
KV              NO rollback kernels: the target wrote d+1 KV rows in place;
                rejected slots are simply overwritten next round
                ("positions are data" — the graph design makes rollback a
                position-bookkeeping operation)
commit          append accepted tokens; next round seeds from the bonus

Greedy equivalence is the primary correctness gate, re-scoped by measurement (doc 83 §3.4): every emitted token is the target sampler's own decision on its verify row, but the verify graph (nt=d+1, Prefill kernels) differs numerically from the nt=1 decode graph by ~0.01–0.05 logits, so sub-margin tokens may flap — both streams are valid greedy chains of the same model (AGENTS rule-9 class). Exact identity is gated on (a) the accept-rule unit tests, (b) the d=0 fallback, (c) near-tie attribution — not on bit-equality across kernel assignments.

4. Design — stage ① minimal loop

Deliberately smaller than the retired plan's machinery:

  • src/spec.rs — one concrete SpecEngine (no Speculator trait chain; the chain seat opens in a later stage if a second speculator type lands): holds the draft model (second load_model + its own GraphCache), runs d draft forwards + 1 verify forward per round, returns 1..=d+1 tokens plus the last row's logits.
  • Vocab compatibility gate at load (kept from the D5 analysis §4.2): same vocab type, BOS/EOS, size delta ≤ 128, token-text equality.
  • CLI: --spec-draft <model> + --spec-draft-n <d> (default 2), round stats (rounds / drafted / accepted / tokens-per-round) to stderr. The spec = off path stays byte-identical to today.
  • Sampling: greedy argmax over the returned n_out rows (stage ① only; temp>0 chains deferred). Logits readback is nt×n_vocab×4 B ≈ 1.8 MB/round at d=2 — acceptable for the loop; on-GPU argmax is a stage ④ candidate.

5. Stages & gates

StageWorkGate
① d=2 loop (greedy) — LANDED 2026-09-12, doc 83src/spec.rs + CLI wiringG1 (re-scoped by measurement): (a) accept-rule unit tests with synthetic logits; (b) d=0 fallback == non-spec path to one exact-tie flap (buffer-placement numerics, any draft quant); (c) self-draft divergences attributed to near-ties (first-flap margin 0.043). Batched-verify (nt=d+1 Prefill graph) vs nt=1 decode kernels differ ~0.01–0.05 logits — same rule-9 class; exact identity returns only with nt-invariant accumulation (stage ④ candidate). G2: 179 tests green; off-path untouched. G3: per-round stats on stderr
② end-to-end battery — LANDED 2026-09-12, doc 84same-window dual-engine protocol (3 reps × prose/code × 4 cells)1.33×/1.59× ≥ 1.2× PASS; llama same-window 1.64×/2.08×; gap fully priced: verify row marginal 8.8 vs 2.5 ms/row → 1.62× recoverable
③ verify-marginal attributionncu on the nt=3 and nt=9 rounds: nt=9 GEMM M-pad waste (doc 82 multi-MMVQ caps at nt≤8), dp4a utilization, attention query-tiling (KV read once per nt rows vs per row), logits/samplingone session; a cost ledger with per-item ms
④ kernel attackper ③'s ledger: multi-MMVQ extended to nt=9–16 (16-lane accumulators) and/or small-M GEMM tiles; graph capture for the fixed verify shapes (kills the +1.2 ms eager round overhead)marginal 7.8 → ≤2.5 ms/tok (14B), then → ~1.5
⑤ d=8 + tuningre-measure the d=8 economics; adaptive d (truncate at acceptance collapse); stretch: ngramd=8 ≥ 1.5× end-to-end (14B)

Each stage records per cuda_optimization_steps/STYLE.md (English, six sections, real numbers) — docs 83+.

6. Risks

RiskMitigation
The verify marginal may not reach llama's ~1.5 ms/token (their small-M MMQ tuning is multi-campaign depth, MMQ-analysis §7–12)The ladder is staged: d=2 pays from plumbing alone (1.34×); ③ prices each candidate before any kernel work
Acceptance p is prompt-dependent (code 3.1×, prose 1.7× in the battery)Report per-prompt, never averages alone; stats per round
n_out = d+1 lm_head pathAlready validated end-to-end by the specverify instrument (docs 81–82); n_out=1 decode path untouched when spec=off
Draft model doubles footprint0.5 GB on top of 9 GB — fine on GB10; opt-in flag only
Metal parityCUDA first (campaign home turf); Metal is a follow-up, not a gate

7. Acceptance gates (campaign rules apply unchanged)

Greedy token-identity vs the non-spec path (the spec path is an optimization, never a behavior change, at temp=0); full suite green; interleaved same-window A/B medians; acceptance-rate and tokens/round instrumentation on every round; every stage documented per STYLE.md.

Appendix — what the D5 closure got wrong (2026-09-10 → 09-12)

Kept brief; details in docs 80/81 (§4.3) / 82 (§5):

  1. The go/no-go bar was mis-derived. The 2.5×-amortization gate treated the per-token marginal (attention + q8 quantize + dp4a compute) as ε; the correct model is C_T(nt) ≈ weights + nt·token-work, so break-even acceptance at d=2 is far lower than the p* = 0.73 the old model printed.
  2. The external anchor was a measurement artifact. The llama-cli batteries passed -md without --spec-type draft-simple (default none) — the draft was silently never loaded, so "llama.cpp also lands at 1.00×" was base-vs-base. Corrected: 1.53–2.43×.
  3. The instrument's summary field was wrong (amortization computed as C_T(1)/C_T(nt), missing the nt factor; fixed 2026-09-12; doc tables were computed from raw medians and unaffected).

Net: the closure rested on a real 0.52× measurement but a miscalibrated bar and a void anchor. Doc 82 fixed the underlying dispatch hole (0.52× → 2.14×), the corrected batteries showed the strategy wins on GB10, and the campaign reopens as D5-R.

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0017 — Speculative decoding refuses the features its identity contract cannot carry

GGUF tooling — convert, quantize, split (F6, #49)

minfer reads GGUF (a reader in src/gguf.rs); F6 adds the writer and the three subcommands that use it. This document is the writer's contract, the subcommands' exact shapes, the supported/refused sets, and the reference used to verify each claim. It is the design and the implementation record in one file.

  • Reader: src/gguf.rs — GgufContext::init_from_data, load_gguf_model, split_file_info, resolve_splits.
  • Writer: src/gguf_write.rs — GgufWriter, TensorSpec, encode_kv, write_single, write_split, split_assignment.
  • Encoders: src/quantize.rs — QuantTarget, quantize_row, decode_to_f32.
  • Converter: src/convert.rs — safetensors + config.json + tokenizer.json parsing, Conversion, QuantizePlan.
  • Subcommands: src/tooling.rs — run_convert, run_quantize, run_split.

1. The writer contract

The parser is the contract: gguf_write emits exactly the bytes GgufContext::init_from_data reads back, and the unit tests in src/gguf_write.rs assert that (write → parse → compare). The layout is llama.cpp's gguf_write_to_file layout.

1.1 Header

magic      "GGUF"              (4 bytes)
version    u32 = 3
n_tensors  i64                 (little-endian)
n_kv       i64

1.2 Metadata (KV) encoding — every GgufType the reader supports

A key is a u64 length followed by its UTF-8 bytes. Then:

GgufTypetagpayload
Uint8, Int8, Bool0, 1, 71 byte
Uint16, Int162, 32 bytes LE
Uint32, Int32, Float324, 5, 64 bytes LE
Uint64, Int64, Float6410, 11, 128 bytes LE
String8u64 length + UTF-8 bytes
Array9element u32 tag, u64 count, then count element payloads

An array of strings is count length-prefixed strings; an array of a numeric type is the raw little-endian values (the reader's GgufKv.data already holds that form, so the writer writes it verbatim). This covers all 13 GgufType variants; there is no metadata type the writer cannot emit.

general.alignment is read back by the parser (u32, powers of two). The writer does not invent it: the caller passes the alignment (32 for a fresh conversion, the source's own ctx.alignment for a rewrite/split), and a source file's general.alignment key is copied through by the rewrite. A single-file rewrite with a default-aligned source is therefore metadata-equivalent to it, key for key (asserted by f6_rewriting_a_gguf_is_bitwise_and_metadata_equivalent).

1.3 Tensor index

For each tensor, in write order:

name      string (must be < 64 bytes)
n_dims    u32   (trailing 1s dropped, at least 1)
ne[0..n_dims]  i64 each
type      u32   (GgmlType)
offset    u64   (relative to the start of the data section)

n_dims is ggml_n_dims(ne): the highest index whose dimension is not 1, so [896, 1, 1, 1] is written with n_dims = 1. The parser reconstructs the same ne (trailing 1s) and derives nb from ne and the type, so the writer never writes strides.

1.4 Alignment

After the tensor index the writer pads to alignment (the parser seeks from its current position to ggml_pad(tell, alignment)), and pads after every tensor's payload to the same alignment. The parser's data-section size is

size = Σ ggml_pad(nbytes(t_i), alignment)          (in tensor order)

and it requires info[i].offset == Σ_{j<i} ggml_pad(nbytes(t_j), alignment) — so the writer computes exactly those offsets and lays the data down in the same order. nbytes is (Π ne) / blck_size × type_size, the same arithmetic as GgufTensorInfo::nbytes.

1.5 Per-tensor byte layout for each emitted type

The writer copies or encodes the bytes that the graph's kernels already consume; it does not add row padding. A tensor of type T with row length ne[0]:

Typebytes per blockelements/blockrow byteslayout
F3241ne[0] × 4LE f32
F1621ne[0] × 2LE half bits
BF1621ne[0] × 2LE bf16 bits (f32's top 16 bits)
Q4_01832ne[0]/32 × 18d: f16, qs[16] (element j low nibble, j+16 high)
Q4_12032ne[0]/32 × 20d: f16, m: f16, qs[16]
Q5_02232ne[0]/32 × 22d: f16, qh: u32 (5th bits, element j bit j), qs[16]
Q5_12432ne[0]/32 × 24d: f16, m: f16, qh: u32, qs[16]
Q8_03432ne[0]/32 × 34d: f16, qs[32]: i8
Q4_K144256ne[0]/256 × 144d: f16, dmin: f16, scales[12] (8×6-bit scale + 8×6-bit min, get_scale_min_k4), qs[128] (element j low nibble, j+32 high, 64 elements per 32-byte group)
Q5_K176256ne[0]/256 × 176as Q4_K, plus qh[32]: the 5th bit of element n+j is bit m1 of qh[j] and of n+j+32 is m2, m1/m2 shifting left by 2 per 64-element group
Q6_K210256ne[0]/256 × 210ql[128] (16 sub-blocks of 16, L[j+l]&0xF / L[j+l+64]&0xF low, L[j+l+32]/L[j+l+96] high), qh[64] (the 2 high bits of each of the four 32-element quarters, shifted 0/2/4/6), scales[16]: i8, d: f16
Q2_K/Q3_K/Q8_K and every I-quant———copied verbatim (rewrite/split only; no encoder)

The writer validates that ne[0] % blck_size == 0, that every dimension is ≥ 1, that names are unique and < 64 bytes, and that the provider hands it exactly nbytes — a short or long payload is a loud error, never a file whose later tensors are shifted.

1.6 The multi-part convention

The reader's convention, which the writer must satisfy:

  • Count in metadata: split.count (total parts).
  • Index in metadata: split.no — 0-based. (llama.cpp's LLM_KV_SPLIT_NO spelling; not split.index.)
  • split.tensors.count (total tensors across parts) is written too; llama.cpp emits it, the reader ignores it.
  • File names: {prefix}-NNNNN-of-MMMMM.gguf, 1-based, five digits; resolve_splits builds all M names from the entry point's name.
  • The entry point is part 0 — the reader opens it, and builds the remaining M - 1 file names from it.

write_split:

  1. Assigns tensors to parts greedily by padded data size (split_assignment), never splitting a tensor. A tensor larger than --max-size is refused with its own size, because the request cannot be honoured.
  2. Gives every part the full metadata plus its own split.no, the shared split.count, and split.tensors.count — this is what makes each part parse standalone, which is what the reader does before merging.
  3. Preserves the global tensor order: part 0's tensors first, then part 1's, … The merged index (each part's info concatenated in part order) therefore equals the single-file index exactly.
  4. When the assignment is a single part, writes a plain {stem}.gguf with no split.* keys at all — one part is not a split, and the loader then sees an ordinary single file.

The reader's failure modes are #[test]-covered on a synthetic 2-part split: a missing part, a part whose split.no does not match its position, and a filename count that disagrees with split.count all fail the load.


2. The subcommands

minfer convert  <hf-model-dir> <out.gguf> [--outtype f16|bf16|f32] [--split-max-size BYTES]
minfer quantize <in.gguf> <out.gguf> --type <target> [--split-max-size BYTES]
minfer split    <in.gguf> <out-dir> --max-size BYTES [--stem NAME]

All three are dispatched before the global option parser (like bench / specverify), so their flags do not collide with the inference options, and none of them initializes a GPU backend.

N accepts plain bytes or a binary suffix (1K, 512M, 2G). For convert/quantize with --split-max-size, the second positional is the single-file output path and the parts are written beside it using its stem; for split, the second positional is the output directory. Exit code is 0 on success and 1 on any refusal, with the reason on stderr.

2.1 convert — HuggingFace → GGUF

Inputs in <hf-model-dir>: config.json, tokenizer.json, tokenizer_config.json, model.safetensors (or a model.safetensors.index.json shard map), optional generation_config.json.

Tensor mapping (HF → GGUF):

HuggingFaceGGUF
model.embed_tokens.weighttoken_embd.weight
model.norm.weightoutput_norm.weight
lm_head.weightoutput.weight (absent for a tied model)
model.layers.N.input_layernorm.weightblk.N.attn_norm.weight
model.layers.N.self_attn.q_proj.{weight,bias}blk.N.attn_q.{weight,bias}
…k_proj / …v_proj / …o_projblk.N.attn_k / attn_v / attn_output
model.layers.N.post_attention_layernorm.weightblk.N.ffn_norm.weight
model.layers.N.mlp.{gate,up,down}_proj.weightblk.N.ffn_{gate,up,down}.weight
…self_attn.rotary_emb.inv_freqdropped (llama.cpp drops it too; RoPE is recomputed)

Any other tensor name is a refusal. Shapes are reversed: HF [out, in] becomes GGUF ne = [in, out, 1, 1], and the data stays row-major [out][in].

Metadata written (marked strict = read by Tokenizer::load / hparams_from_gguf; a file missing one is refused at load):

  • general.architecture = "qwen2", general.type = "model", general.name
  • strict qwen2.{block_count,context_length,embedding_length,feed_forward_length}
  • strict qwen2.attention.{head_count,head_count_kv,layer_norm_rms_epsilon}
  • strict qwen2.rope.freq_base
  • general.file_type (1 = f16, 32 = bf16, 0 = f32; llama.cpp's llama_ftype numbers), general.quantization_version = 2
  • general.sampling.{top_k,top_p,temp,penalty_repeat} from generation_config.json
  • strict tokenizer.ggml.model = "gpt2", strict tokenizer.ggml.pre = "qwen2"
  • strict tokenizer.ggml.tokens (all vocab_size ids), tokenizer.ggml.token_type, tokenizer.ggml.merges, tokenizer.ggml.{eos,bos,padding}_token_id, tokenizer.ggml.add_bos_token
  • strict tokenizer.chat_template (required; absent is a refusal)

tokenizer.ggml.token_type follows llama.cpp's get_vocab_base: type 1 (NORMAL) for the base vocabulary, type 3 (CONTROL) for an added token that is flagged special or shaped <|…|>, type 4 (USER_DEFINED) for the other added tokens, and type 5 (UNUSED) for ids in [|vocab|, vocab_size) with the [PAD{id}] placeholder. tokenizer.ggml.scores is not written (llama.cpp does not either; the strict loader defaults missing scores to 0).

--outtype f16 writes 2-D tensors as f16 and 1-D tensors (norms, biases) as f32 — llama.cpp's "except 1d tensors" rule, and what the engine's f32 norm/bias path consumes. --outtype f32 writes everything as f32. --outtype bf16 (#142) is the same shape with a bf16 2-D payload: f32 → bf16 is round-to-nearest-even (ggml_compute_fp32_to_bf16, including the quiet-NaN rule), and 1-D stays f32 — conversion/base.py's n_dims <= 1 rule, which byte parity against llama-quantize --pure … BF16 confirms (§4.1.1).

2.2 quantize — re-encode a single-file GGUF

Targets with an implemented and byte-verified encoder: q4_0, q4_1, q5_0, q5_1, q8_0, q4_K, q5_K, q6_K, plus the f16 / f32 element casts.

  • For a quant target only 2-D float tensors are quantized. 1-D tensors (norms, biases) and any tensor whose row length is not a multiple of the target block size keep their source type, and the CLI prints the list — llama.cpp's rule, not a silent choice.

  • The K-quant targets write one uniform type — they are not llama.cpp's _M mixtures. q4_K/q5_K are the CLI aliases for LLAMA_FTYPE_MOSTLY_Q4_K_M / _Q5_K_M in llama-quantize (and their general.file_type numbers here are llama.cpp's 15/17/18), but minfer quantize --type q4_K puts every encodable 2-D tensor at q4_K. That is what llama-quantize --pure writes, and it is the reference the byte-parity gate uses; the mixture planner (per-layer Q6_K bumps, the OUTPUT/tied-embedding branch, use_more_bits) is not implemented — #203. A file from llama-quantize … q4_K without --pure is therefore not what this command produces.

  • A K-quant row must be a multiple of 256. QK_K is 256, and a 2-D tensor whose ne[0] is not a multiple of it cannot be encoded at the requested type; minfer demotes it exactly the way llama.cpp's tensor_type_fallback does:

    requesteddemoted to
    q4_Kq5_0
    q5_Kq5_1
    q6_Kq8_0

    and to f16 when even a 32-element block does not divide the row. The CLI prints the demoted list (quantize: N tensor(s) demoted from q4_K (row length not a multiple of 256; llama.cpp's tensor_type_fallback): …). This is not cosmetic: the 0.5B's hidden size is 896 = 3.5 × 256, so 145 of its 290 tensors take the demotion and only the 24 ffn_down tensors (ne[0] = 4864 = 19 × 256) reach the q4_K/q5_K/q6_K encoder. §4.2 gates a second source (hidden 1024) where every 2-D tensor does. For the legacy targets (block size 32) the plan is unchanged: a 2-D row that is not a multiple of 32 keeps its source type, exactly as before #140.

  • The f16 cast follows the same "except 1d tensors" rule: 2-D tensors become f16, 1-D tensors (norms, biases) keep their source type — f32 in every file minfer convert --outtype f16 or llama-quantize … F16 writes, and the only type the engine's norm/bias path reads (an f16 norm is a file the engine cannot run). The CLI prints the preserved list exactly as for a quant target. minfer quantize --type f16 now produces a runnable file this way (#169).

  • The f32 cast is the one target that converts every tensor, 1-D included, to f32.

  • On a tied model (no output.weight), a sub-8-bit legacy target quantizes the shared token_embd.weight at q8_0 — llama.cpp's tied-embedding policy. The CLI prints this too. (A K-quant target skips the mixture, so its tied embedding goes to the requested type and then to that type's fallback.)

  • A multi-part source is refused (quantize one file, then split).

  • K-quant/I-quant sources are refused: there is no dequantizer, so re-encoding them would produce wrong weights.

2.3 split — single file → parts

Verbatim tensor copy (no re-encoding) under {out-dir}, using the source's own alignment and metadata. A multi-part input, an already-split-looking filename, and a model with no tensors are refused.

--max-size N is required — there is no silent default part size. --stem NAME names the output: the parts are written as {stem}-NNNNN-of-MMMMM.gguf, or as {stem}.gguf when everything fits in a single part (which then carries no split.* keys at all). Without --stem the stem is the input file's own name minus .gguf, or model when it has none.


3. Supported / refused sets and the exact refusal texts

minfer convert: unsupported architecture (model_type = "llama",
  architectures = ["LlamaForCausalLM"]); minfer's converter supports
  'Qwen2ForCausalLM' (model_type "qwen2") only

minfer convert: HuggingFace tensor 'model.layers.0.mlp.gate_up_proj.weight' has
  no GGUF mapping in minfer's Qwen2 converter; supported tensors are
  model.embed_tokens, model.norm, lm_head, and model.layers.N.{input_layernorm,
  post_attention_layernorm,self_attn.{q,k,v,o}_proj,mlp.{gate,up,down}_proj} —
  minfer refuses to drop a weight it does not recognise

minfer convert: tensor 'x' has safetensors dtype 'I8'; minfer converts bf16,
  f16 and f32 only

minfer convert: unknown --outtype "q4_0"; supported: f16, bf16, f32

minfer convert: <dir>/tokenizer_config.json has no 'chat_template'; minfer's
  strict loader renders the model's own template, so a converted GGUF without
  one is not accepted

minfer quantize: --type is required (no silent default: a wrong target would
  write wrong weights); supported: q4_0, q4_1, q5_0, q5_1, q8_0, q4_K, q5_K,
  q6_K, f16, f32

minfer quantize: target "q2_K" is a known GGUF type but minfer has no weight
  encoder for it (minfer can only read it, so writing it would emit wrong
  weights); supported encoder targets: q4_0, q4_1, q5_0, q5_1, q8_0, q4_K,
  q5_K, q6_K, f16, f32

minfer quantize: unknown quant target "banana"; supported: q4_0, q4_1, q5_0,
  q5_1, q8_0, q4_K, q5_K, q6_K, f16, f32

minfer quantize: tensor 'blk.0.ffn_down.weight' has type q4_K, which minfer
  cannot decode (no dequantizer for it); re-quantizing a K-quant/I-quant source
  is unsupported — start from an f16 or f32 GGUF

minfer quantize: <path> is a multi-part split (N parts); quantize a
  single-file GGUF, then split the result

GGUF split: tensor 'output.weight' is 144643072 bytes, larger than the
  67108864-byte part cap — a tensor cannot be split across parts; raise
  --split-max-size (at least 144643072 bytes) or use a smaller-capable format

minfer split: <path> is already a multi-part split (N parts); point at a
  single-file GGUF

minfer split: --max-size is required (no silent default part size)

size mismatch for <path>: expected 1048576 bytes, got 1572864 bytes (the server
  may have ignored the resume Range, or the download was truncated); removed the
  partial file

What each conversion step is, exactly

StepExactness
f16 → f16 (copy), f32 → f32 (copy)bit-exact
f16 → f32, bf16 → f32bit-exact (exponent and mantissa are preserved; bf16 is f32's top 16 bits)
bf16 → f16exact in the mantissa (bf16's 8 mantissa bits fit f16's 10) but not in the exponent range: a bf16 value outside f16's range becomes ±inf, and one below f16's smallest normal (2^-14) is rounded onto f16's subnormal grid (below 2^-24 it flushes to zero). No saturation is applied, so the loss is visible rather than a silently clamped weight.
f32 → f16not exact — round-to-nearest-even, and the same exponent-range rule
f32 → bf16 (#142)not exact — round-to-nearest-even on f32's top 16 bits (ggml_compute_fp32_to_bf16, NaN forced quiet). bf16 keeps f32's exponent range, so there is no overflow, only mantissa rounding
bf16 → bf16 (#142)bit-exact: the identity, since bf16 is f32's top half. Routed through the RNE encoder rather than copied, so a NaN payload is quieted exactly as the reference does.
f16 → bf16 (#142)not exact — f16's 10-bit mantissa is rounded to bf16's 7 stored bits; no overflow, because bf16's exponent range covers f16's
rewrite / split (any type)bit-exact: the payload is copied
f16/f32 → q4_0/q4_1/q5_0/q5_1/q8_0lossy by construction (that is the point); the encoder is byte-identical to llama.cpp's reference
q4_0…q8_0 → f32the exact stored values (a dequantize, not a re-quantize)

4. Verification: references and tolerances

4.1 The reference the HF conversion was checked against

Two independent references, both on the same Qwen/Qwen2.5-0.5B-Instruct checkpoint (bf16 safetensors downloaded to /tmp/f6-work/hf-src):

  1. llama.cpp's converter (convert_hf_to_gguf.py --outtype f16 with torch 2.14.0 / transformers 5.17.0) → ref-f16.gguf. All 290 tensor payloads are byte-identical to minfer's output (sha256 per tensor, compared by name). Metadata is equivalent for every value the loader reads; the only differences are key order and that llama.cpp also writes the cosmetic general.size_label.
  2. llama.cpp itself, via llama-cli/the formula below, for the logits. Running the two different files through minfer gives bitwise-identical logits and an identical greedy continuation, which is the strong claim: the files carry the same weights and minfer computes the same thing from them.

Because the source is bf16 and the output f16, "bit-exact" here is the bf16 → f16 row above: exact in the mantissa. It is not a claim that no information was lost. The mantissa half is true — bf16 has 7 stored mantissa bits (8 with the implicit one) and f16 has 10, so no bf16 mantissa is rounded by f16 — but the exponent range is not: f16's smallest normal is 2^-14 and its smallest subnormal is 2^-24, so a bf16 value below 2^-14 lands on f16's subnormal grid (and is rounded there), and one below 2^-24 flushes to zero. On this checkpoint that is 123 024 values in the 169 2-D tensors (measured 2026-09-27, f142_bf16_output_runs_within_the_stated_bound); every one of them is an f16 subnormal. This was found while checking #142's premise that "every bf16 value is exactly representable in f16" — it is not, and that is why the bf16-vs-f16 file comparison is a stated bound, not a bitwise claim (§4.3).

4.1.1 The bf16 writer (#142)

--outtype bf16 is accepted and writes 2-D BF16 / 1-D F32. The reference is a pure cast by llama.cpp, but from the f32 conversion rather than the f16 one:

minfer convert ~/.cache/minfer/f6-src/hf/Qwen2.5-0.5B-Instruct \
  ~/.cache/minfer/f6-src/qwen2.5-0.5b-instruct-f32.gguf --outtype f32
llama-quantize --pure ~/.cache/minfer/f6-src/qwen2.5-0.5b-instruct-f32.gguf \
  ~/.cache/minfer/f6-src/ref/qwen2.5-0.5b-bf16-from-f32.gguf BF16
MINFER_F142_LLAMACPP_BF16=~/.cache/minfer/f6-src/ref/qwen2.5-0.5b-bf16-from-f32.gguf \
  cargo test --release --bin minfer f142_bf16_conversion_is_byte_identical_to_the_reference -- --ignored

minfer convert --outtype f32 is exact for a bf16 source, so the cast is the f32→bf16 RNE projection of the checkpoint's own values. Result (measured 2026-09-27, aarch64): 290/290 tensor payloads byte-identical — 169 BF16 2-D tensors, 121 F32 1-D tensors. The 1-D rule agrees because conversion/base.py sets data_qtype = F32 for n_dims <= 1 on every file type, and llama-quantize's tensor_allows_quantization returns the source type for a 1-D tensor.

Why not the f16 source. #142's Source B (llama-quantize --pure <f16>.gguf … BF16) was tried first and is not byte-identical here: casting the f16 file cannot recover the 123 024 subnormal values that the f16 conversion already rounded, so 126 575 bytes across all 169 2-D tensors differ (measured 2026-09-27, per-tensor byte diff). That is the reference's input, not minfer's writer — which is exactly what the f32-source reference isolates. The writer has no torch/transformers converter to check against on dgxspark (convert_hf_to_gguf.py --outtype bf16 needs torch); the f32-source cast is the strongest available reference, and the missing direct converter reference is #209.

4.2 The reference the encoders were checked against

llama-quantize ref-f16.gguf ref-<type>.gguf <type> on the same f16 source. For each of q4_0, q4_1, q5_0, q5_1, q8_0, all 290/290 tensors are byte-identical — against the recorded dgxspark build: gcc 13.3.0, -ffp-contract=fast, llama.cpp revision unrecorded (tests/fixtures/f6-fixtures.json records it per content, §4.2.2) (measured 2026-09-27, dgxspark (aarch64, GB10 sm_121), MINFER_F6_F16_GGUF=… cargo test --release --bin minfer f6_quantize_encoder_is_byte_identical_to_llamacpp -- --ignored; the env_path convention and the per-type command are below). The claim is conditional on all three of those — the compiler, the effective -ffp-contract, and the llama.cpp revision — because a llama-quantize binary is a (compiler, flags, revision) triple and the reference is that binary's output. That qualifier is not decoration, and this section records it because the number is otherwise a statement about one build of llama-quantize: measured 2026-10-07, the same llama.cpp source built with the one flag changed does not match, and a differently-built reference made the same gate red on a platform this project ships on (#334). The K-quant half of the claim is compiler-sensitive as well as flag-sensitive, which §4.2.1 records (#349).

Which build, and how it was identified. The dgxspark producer is a GCC 13.3.0 -O3 -DNDEBUG build with no explicit -ffp-contract (so GCC's default fast). It is identified by reproduction rather than by a recorded version string — nothing recorded the binary's own source revision, which is part of what #205 still owes. Measured 2026-10-07 on dgxspark, one variable at a time, from ~/git/reading/llama.cpp HEAD 050dde50c and the cached f16 source:

llama-quantize buildref/qwen2.5-0.5b-q4_0.gguf sha256
the cached build/bin/llama-quantize (GCC 13.3.0, -O3 -DNDEBUG)04634958ae0289b8c557c4d50491d225b23728326fd6d5bbfabedd299a89b3b9
fresh build of HEAD 050dde50c, same flags04634958… — identical, so the flag is the only variable
that same build + -ffp-contract=offea94611ef461e5a65c1af0b6c7189d731735a31e310b54b48ad9b94fa2521930

The Mac's reference (macbook (macOS 27.0.1, Apple M4 Pro), 2026-10-07; the binary built 2026-09-01 from 458681e1d, CMAKE_C_FLAGS_RELEASE=-O3 -DNDEBUG, no -ffp-contract=fast) is sha256 51c2b000… and disagrees with minfer in 168 of 290 q4_0 tensors — every difference a single data nibble, zero scale bytes (blk.0.attn_k.weight: exactly 12 of 64512 bytes).

Why. llama.cpp computes x*id + c in quantize_row_q4_0_ref, and whether the compiler contracts that across statements is a rounding decision: under -ffp-contract=fast the expression is one FMA, otherwise fmul + fadd, and the two pick a different quant at an exact rounding boundary. The Mac's quantize_row_q4_0_ref disassembles to fmul.4s + fadd.4s with no fmla. GCC contracts by default and Apple clang does not without the flag, so both sides are right about their own build; minfer's encoder uses f32::mul_add unconditionally, i.e. it reproduces the contracting one. This is the mirror of the experiment this section used to record alone — minfer without mul_add against the dgxspark reference differs in 12 of 64512 bytes on the same tensor, the same count — which is what makes "the reference's build" the whole content of the claim. Q8_0 uses a single multiply and matched without it.

What the gate does about it. f6_quantize_encoder_is_byte_identical_to_llamacpp prints the reference's path, size and the recorded build it established it matched (build -ffp-contract=fast, recorded reference gcc 13.3.0 … on a pass). On a mismatch it does not report an encoder defect first: it re-encodes the whole source with the uncontracted arithmetic — quantize::FmaContract::Off through quantize_row_with, which is exact for the legacy quants and a model of the K-quants' per-expression fusion (§4.2.1) — and then it asks the manifest which recorded content the file is, because re-encoding alone cannot tell a different compiler from a corrupted file: both match neither model. The three byte comparisons — Fast against the reference, the uncontracted model against it, and the file's digest against tests/fixtures/f6-fixtures.json (§4.2.2) — plus the entry the digest matched give five named outcomes, and the classifier that decides them is pure and unit-tested (the_f6_parity_verdict_names_the_recorded_build):

outcomewhenwhat the gate does
reproducesthe Fast model matchespass, and the pass line names the recorded compiler and flag
flag mismatchonly the uncontracted model matchesfails with the build named: "the reference is NOT the -ffp-contract=fast build this encoder reproduces … rebuild llama-quantize with -ffp-contract=fast" (#334)
not the recorded contentthe digest matches no recorded content for the paththe resolver of §4.2.2 has already refused the file by name and digest; a path outside the cache is not a fixture at all
a recorded foreign buildthe digest is recorded, but the entry is not the path's authoritative_reference, and neither model matchesskips loudly (a passing run with a [f6 parity] SKIP line) naming the build the file is, the authoritative build the claim was measured against, the reason (§4.2.1's compiler sensitivity), and this box's own cc --version
the authoritative build, not reproducedthe digest is the authoritative_reference and neither model matchesfails: the encoder no longer reproduces the reference the claim is asserted against — a defect, not a compiler difference

The #[ignore] reason names the same prerequisite. The identity is a whole-file, per-tensor byte comparison in every outcome. The fixture it resolves is additionally verified against the manifest of §4.2.2 first, so which file is being compared against is checked too — a stale or replaced reference in the cache refuses the run instead of silently becoming the reference.

The pure encoder gate (quantize::tests) additionally pins the block layout against hand-checked reference numbers: the scale, the j / j+16 nibble packing, the 5th-bit plane, type_size/blck_size, and the zero-block case.

4.2.1 The K-quants (#140)

The K-quant reference functions (quantize_row_q4_K_ref / q5_K_ref / q6_K_ref) are search quantizers: make_qkx2_quants scans 21 (q4_K) or 16 (q5_K) candidate scale/min pairs and make_qx_quants re-derives the least-squares scale for 19 candidate iscale values. Every candidate is evaluated in an a*b + c shape, so the FMA contraction is a rounding decision, and one ULP picks a different quant. Matching llama.cpp took four distinct findings, each read off the disassembly of the production object (objdump -d build/ggml/src/CMakeFiles/ggml-base.dir/ggml-quants.c.o), then confirmed by byte parity:

  1. sum_x2 += x*x is an FMA. The q4_K/q5_K per-element weight is av_x + |x| with av_x = sqrt(sum(x²)/32), and the sum is contracted (fmadd s0, s1, s1, s0). The weight feeds the search's error metric.
  2. nearest_int(a*b) folds the magic constant into the product's FMA. The reference's round-to-nearest-even trick becomes fmadd a, b, #12582912.0 (fmov w0, #0x4b400000) and only then masks the mantissa — the product is not rounded before the add. nearest_int(a * b) in Rust rounds twice and picks a different integer at a boundary; the port needs nearest_int_mul.
  3. a*b - c*d contracts per expression, not per shape. In make_qkx2_quants the discriminant D = sum_w*sum_l2 - sum_l*sum_l fuses its left product (fmul + fnmsub), while this_scale = sum_w*sum_xl - sum_x*sum_l fuses its right one (fmul + fmsub). Both are a different ULP from the plain expression; the port writes each explicitly.
  4. A scalar loop and its vectorized twin can disagree. In make_qx_quants the initial accumulation loop is scalar and uses fmadd, while the 19-candidate search loop is 4-wide vectorized and computes plain fmul products with an in-order fadd reduction — no FMA at all. Matching only the scalar form left 137 of 424 random 16-element groups differing from the reference; matching the vectorized form as well made all 424 equal.

Reference procedure (uniform encoders). llama-quantize's q4_K/q5_K are CLI aliases for the Q4_K_M/Q5_K_M mixtures; the uniform encoder this project implements is what --pure selects. The gate's reference must therefore be:

llama-quantize --pure <f16>.gguf <ref>-q4_K.gguf q4_K    # likewise q5_K, q6_K
llama-quantize          <f16>.gguf <ref>-q4_0.gguf q4_0   # legacy: no mixture to disable

Sources and results (measured 2026-09-27, aarch64). The f16 sources live in the persistent cache ~/.cache/minfer/f6-src/ (never /tmp):

sourcehidden2-D tensors that reach the K encoder
qwen2.5-0.5b-instruct-f16.gguf (from Qwen/Qwen2.5-0.5B-Instruct, bf16 safetensors → minfer convert --outtype f16)89624 of 144 — 145 rows are not a multiple of 256 and take tensor_type_fallback, 121 tensors are 1-D copies
qwen3-0.6b-f16.gguf (minfer quantize <Qwen3-0.6B-Q8_0.gguf> … --type f16)1024197 of 197 — no demotion, every 2-D tensor

MINFER_F6_QUANT_TYPE=q4_K|q5_K|q6_K with MINFER_F6_F16_GGUF / MINFER_F6_LLAMACPP_QUANT set to the pair above: 290/290 tensors byte-identical on the 0.5B (encoded-as q5_0: 145, f32: 121, q4_K: 24) and 310/310 on the 0.6B (encoded-as f32: 113, q4_K: 197), for each of the three types — against the same recorded dgxspark build as §4.2: gcc 13.3.0, -ffp-contract=fast. The gate prints that breakdown, so "byte-identical" always carries how many tensors the new encoder actually saw; since #349 it names the recorded compiler too. Legacy regression on the same 0.5B source: q4_0/q4_1/q5_0/q5_1/q8_0 all still 290/290 against the reference of §4.2 (the K-quant rows above are against that same build, and their per-expression fusion is not reproducible by the Off variant — that variant is the gate's provenance probe, not a second encoder).

The K-quant claim is compiler-sensitive, and the record says which compiler. Measured 2026-10-07 on macbook (macOS 27.0.1, Apple M4 Pro): after rebuilding llama-quantize from c479922ac with -DCMAKE_C_FLAGS_RELEASE="-O3 -DNDEBUG -ffp-contract=fast" (the flag confirmed by disassembly — fmul + fadd → fmadd/fmla, Apple clang 21.0.0), the five legacy quads pass 290/290 and every K-quant fails: tensor blk.0.ffn_down.weight payload differs (2285 of 2451456 bytes) … build unmatched for q4_K, 97 of 2996224 for q5_K, 269 of 3575040 for q6_K — the reference matches neither minfer's FmaContract::Fast nor its uncontracted Off model. git log 050dde50c..c479922ac -- ggml/src/ggml-quants.c touches only the q3_K/i-quant encoders, so the residual is Apple-clang-vs-GCC codegen of the search quantizers above, not -ffp-contract.

The decision this forces, recorded. The byte-parity claim is asserted against the content tests/fixtures/f6-fixtures.json marks authoritative_reference (§4.2.2) — for every ref/… path exactly one, the dgxspark gcc 13.3.0 -ffp-contract=fast build. A reference whose digest is recorded but is not that entry is a different compiler's build of the same source: the gate names the file's build, the authoritative one and the reason and skips loudly, because a permanently red gate on a supported platform is the failure mode the project has already filed once (docs/ARCHITECTURE-EXECUTION-PLAN.md §14 row 6). A mismatch against the authoritative entry, or against a file no entry records, still fails: the skip is never available to the claim's own reference, and a corrupted or unrecorded file cannot be excused as a compiler difference.

Reproduce the sources and references:

mkdir -p ~/.cache/minfer/f6-src/ref ~/.cache/minfer/f6-src/hf/Qwen2.5-0.5B-Instruct
# 1. the HF checkpoint (5 files, ~988 MB)
for f in config.json tokenizer.json tokenizer_config.json generation_config.json model.safetensors; do
  curl -sL -o ~/.cache/minfer/f6-src/hf/Qwen2.5-0.5B-Instruct/$f \
    https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct/resolve/main/$f
done
# 2. the two common f16 sources
minfer convert ~/.cache/minfer/f6-src/hf/Qwen2.5-0.5B-Instruct \
  ~/.cache/minfer/f6-src/qwen2.5-0.5b-instruct-f16.gguf --outtype f16
minfer quantize ~/.cache/minfer/models/hf/Qwen/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf \
  ~/.cache/minfer/f6-src/qwen3-0.6b-f16.gguf --type f16
# 3. the llama-quantize references
for t in q4_0 q4_1 q5_0 q5_1 q8_0; do
  llama-quantize ~/.cache/minfer/f6-src/qwen2.5-0.5b-instruct-f16.gguf \
    ~/.cache/minfer/f6-src/ref/qwen2.5-0.5b-$t.gguf $t
done
for m in qwen2.5-0.5b qwen3-0.6b; do for t in q4_K q5_K q6_K; do
  llama-quantize --pure ~/.cache/minfer/f6-src/$([ $m = qwen2.5-0.5b ] && echo qwen2.5-0.5b-instruct || echo qwen3-0.6b)-f16.gguf \
    ~/.cache/minfer/f6-src/ref/$m-$t.gguf $t
done; done

4.2.2 The fixture manifest (#205)

The f16 sources and the llama-quantize references are inputs the gates do not produce, so until #205 nothing noticed when the cache held something else: a stale or replaced reference was compared against silently and the gate stayed green. tests/fixtures/f6-fixtures.json is the record — one entry per artifact content identity, carrying path, bytes, sha256 (or a sha256_prefix where the 2026-10-07 table truncated it), the exact producer command, the producer's identity (minfer_commit for a minfer-produced file; llamacpp_commit + compiler + ffp_contract + cflags for a llama-produced one), date and an absolute box label (gate contract rule 5). Two entries for one path are a recorded divergence, and each of them carries a divergence_notes entry explaining it: the f16 source and its bf16 cast differ by minfer producer version (dgxspark, 2026-09-27, commit unrecorded — a Mac regeneration at master ab34a72 produced other bytes, and bf16-from-f32, whose input is identical on both boxes, is identical too), while the ref/qwen2.5-0.5b-* references differ by the reference build (§4.2: the two boxes' llama-quantizes — dgxspark gcc 13.3.0 and the Mac's Apple clang 21.0.0).

Exactly one of a ref/… path's entries carries authoritative_reference: true (#349): the content the byte-parity claim is asserted against, which for every path today is the dgxspark gcc 13.3.0 -ffp-contract=fast build. The mark is what lets the parity gate's verdict say which recorded build it is looking at, and therefore whether a mismatch against both encoder models may be excused as a different compiler (§4.2.1) or must be reported as an encoder defect. --check enforces it as S6: a path with a llama-quantize record has exactly one marked entry, the mark is on a llama-quantize entry built with -ffp-contract=fast, and it appears nowhere else.

python3 scripts/check_f6_fixtures.py --check     # CI: the manifest's shape + the tree cross-check
python3 scripts/check_f6_fixtures.py --verify    # + every cached file's bytes and sha256
python3 scripts/check_f6_fixtures.py --selftest  # the checker's own cases, including a tampered copy
python3 scripts/check_f6_fixtures.py --file /tmp/copy.gguf   # one file

--verify names the file, every recorded digest and the actual one when they disagree, and reports a file no entry names; --check additionally requires that every ~/.cache/minfer/f6-src/… fixture the source tree spells is an entry, that every .gguf a recorded producer command names is too, and (S7) that the command runs the program its producer_kind names and writes the entry's own path. The other half runs inside the gates: the shared fixture resolver verifies a path it hands back (env_path → src/tooling/tests/f6_fixtures.rs), so the documented cargo test … --ignored invocation refuses a tampered cache by name and digest rather than comparing against it. MINFER_F6_CACHE relocates the cache root for an experiment — which is how the refusal is demonstrated without touching ~/.cache/minfer/f6-src/ — and a deliberate reference belongs outside the cache root, where a path the manifest does not name is not a fixture at all.

The record itself is relocatable too, with MINFER_F6_MANIFEST (#354): the manifest-side twin of MINFER_F6_CACHE, honoured by both halves — the gate's reader (src/tooling/tests/f6_fixtures.rs) and the checker (--manifest still wins over it) — so a test can point the gate at a /tmp record instead of editing a tracked file:

MINFER_F6_MANIFEST=/tmp/f6-fixtures.json cargo test --release --bin minfer -- \
    the_manifest_override_drives_the_whole_resolver
MINFER_F6_MANIFEST=/tmp/f6-fixtures.json python3 scripts/check_f6_fixtures.py --check

Both halves honour it loudly: a value that is set but empty, or a record that is missing, unreadable or malformed, refuses by name and never falls back to tests/fixtures/f6-fixtures.json — the empty case is a usage error (exit 2) in the checker. The checker's --regenerate (§4.2.3) refuses the variable outright, because re-recording into a manifest chosen by the environment is the accident the override must not create; --manifest <path> is that mode's auditable spelling. The override is what makes the parity gate's fifth verdict testable: with a fabricated record that names a reference's content as a non-authoritative build, ParityVerdict::RecordedForeignBuild — the loud [f6 parity] SKIP — is exercised end-to-end from a unit test and from a real gate run, instead of by editing the tracked manifest under a cp backup (#349).

What it does not cover. The manifest checks content, not truth: nothing else recorded those bytes, so a wrong digest in the record is accepted. Every entry carries a full sha256 since PR #350 re-captured the eleven Mac references (the 2026-10-07 table had kept 8 hex characters), so --strict-digests passes with no WEAK line; the reader still accepts a sha256_prefix entry weakly, for a digest that has to be re-captured again. The hf/ checkpoint's revision and f164/minfer-f16.gguf's producer are recorded (they were the other two gaps that PR closed), but a llama-quantize reference is still re-recorded by hand: --regenerate (§4.2.3) refuses one on purpose.

4.2.3 Regenerating from the manifest (#345)

--verify checks the cached bytes against the record; it does not produce them. --regenerate closes that gap: it re-runs the recorded producer command for the selected entries and re-records the content identity from what the run wrote.

python3 scripts/check_f6_fixtures.py --regenerate --only qwen2.5-0.5b-instruct-f32.gguf
python3 scripts/check_f6_fixtures.py --regenerate --box 'dgxspark (aarch64, GB10 sm_121)'
python3 scripts/check_f6_fixtures.py --regenerate --dry-run     # classify, run nothing (CI)

--only takes a manifest path or PATH@BOX (the form that selects one of a recorded divergence's contents); --box selects every entry recorded against that box label. The mode is a verification gate first:

  • It runs the producer the record names, or refuses. The command's program is resolved before anything runs — minfer on PATH and then ./target/release/minfer, curl for an hf-download entry, and the entry's own llamacpp_binary for a llama-quantize one. There is no fallback between producers: a missing llama-quantize is a refusal naming the path it looked for, never a reason to run minfer instead.
  • Regenerating a llama-quantize reference is out of scope. The byte-parity claim of §4.2 is about one compiler's build, so the mode refuses such an entry — naming the llamacpp_binary it records and whether that binary exists here — and says to re-run it on the box that records it. Even a box that has the build takes the refusal: a half-implemented compiler-identity check would be worse than not running it.
  • The run must reproduce the record. The content the producer wrote is hashed and compared with the entry's full sha256. A match re-records bytes and the date of the reproducing run (sha256 is equal by construction) and prints every change; a second run the same day writes nothing, so the mode is idempotent. A digest that matches no recorded content for the path is a finding (exit 1), not an update — the record's producer no longer reproduces the record — and the message names both digests, the bytes and the commit that ran.
  • The producer identity is checked before the run. A minfer entry whose minfer_commit differs from what runs here (MINFER_F6_COMMIT, else the tree's HEAD) is refused (exit 3) naming both commits: that run is not the producer the entry names, so its bytes cannot be recorded against it. An unrecorded identity may run — there is no claim to contradict — but a differing digest is then a finding like any other, never a "producer version" excuse.
  • It writes only sha256/bytes/date. authoritative_reference, the producer command and every identity field are left exactly as they were, so the mode can neither invent provenance nor move the content the byte-parity claim is asserted against. --check (S1–S7) re-runs on the result before it is kept.
  • A rejected run cannot destroy a fixture. The existing cache file is renamed to <path>.regen-before before the producer starts (a rename, not a second copy); a finding, a refusal after a run, or a producer failure restores it and keeps the rejected content at <path>.regen-rejected; a verified success removes the pre-run copy, which is byte-identical to the new file.
  • --dry-run classifies every selected entry — RUN with the ~-expanded command, or REFUSED with the reason — and writes nothing. CI runs it, so a manifest whose command has the wrong program, the wrong output path or no command at all fails there (S7).

Exit codes: 0 clean, 1 a finding or a producer failure, 3 a refused target, 2 a selection that matches nothing. A refusal is the honest per-box answer to "this producer cannot run here", not a silent skip; --strict-runnable turns it into a failure for a box that is expected to hold every producer.

Measured (dgxspark (aarch64, GB10 sm_121), 2026-10-07). --regenerate --only qwen2.5-0.5b-instruct-f32.gguf re-ran minfer convert … --outtype f32 and the file came back byte-identical — 6894f9ea3eb79e29…, 1 982 078 784 B, the recorded digest and size — so the only manifest change was the date (2026-09-27 → 2026-10-07); running it again the same day wrote nothing at all. The eleven ref/… entries refuse by name: the Mac's build-fpc binary is not on this box, and the dgxspark one is refused as out of scope even though it is.

4.3 Tolerances

ComparisonToleranceMeasured
rewrite vs source logitsbitwise (assert_eq!)equal
minfer-converted vs llama.cpp-converted logits (same engine)bitwiseequal
split vs unsplit logitsbitwiseequal
f16 → q8_0 logits, same contextmax |Δ| ≤ 1.0 and greedy text identicalmax |Δ| = 0.481, mean 0.082, max |logit| = 18.43 (2.6% of the largest logit); greedy continuation identical
f16 → q4_K logits, same context, minfer quantize output (0.5B)max |Δ| ≤ 0.30 × max |logit| and the first greedy token identicalmax |Δ| = 2.72, mean 0.44, max |logit| = 18.43 (14.8%); greedy [12095, 13, 1084, 374] identical to the f16 source
f16 → q5_K logits, same context (0.5B; the file is mostly Q5_1 after the fallback)samemax |Δ| = 4.10, mean 0.83 (22.3%); greedy [12095, 13, 12095, 374] — 3 of 4 tokens; the f16 source's 1084 flips to 12095 at token 2. The file is byte-identical to llama-quantize --pure … q5_K's, so the flip is the quantisation, not minfer
f16 → q6_K logits, same context (0.5B)samemax |Δ| = 0.996, mean 0.144 (5.4%); greedy identical
f16 → q4_K/q5_K/q6_K logits (Qwen3-0.6B, hidden 1024 — every 2-D tensor K-encoded)samemax |Δ| = 3.71 / 2.39 / 1.19, mean 0.69 / 0.43 / 0.21, max |logit| = 19.99 (18.6% / 11.9% / 5.9%); greedy [12095, 13, 576, 6722] identical for all three
an f16 file's CUDA logits vs the same file's CPU logits (#141, 34-token prompt, ctx 512, Qwen2.5-0.5B-Instruct f16)max |Δ| ≤ 0.01 and max relative ≤ 1e-3, greedy continuation identicalmax |Δ| = 7.34e-5, mean 1.26e-5, max |logit| = 18.43 (4.0e-6 relative); greedy [12095, 13, 1084, 374] on both
that f16 file under llama.cpp (same prompt, --temp 0)—Paris., the same greedy continuation minfer produces on CPU and CUDA
a bf16 file's CPU logits vs the f16 file's, same context (0.5B, bf16 source)max |Δ| ≤ 1e-4 and max relative ≤ 1e-5 and greedy continuation identical (the bitwise expectation of #142's text is refuted by measurement, §4.1)max |Δ| = 2.29e-5, mean 3.19e-6, max |logit| = 18.43 (1.24e-6 relative); 147 357 of 151 936 logits differ in the last bits, and the whole difference is attributed to the 123 024 f16-subnormal weight values the f16 file rounds (the bf16 file carries the checkpoint's exact values); greedy [12095, 13, 1084, 374] on both

The K-quant run gate (f6_k_quant_output_runs_within_the_stated_bound) states its bound as max |Δlogit| ≤ 0.30 × max |logit| and the first greedy token identical before measuring, prints all six measurements, and is run twice (0.5B and 0.6B). The 0.30 bound was set after the first, too-optimistic pass of 2.0 absolute (taken from the q8_0 row) came back at 2.72 on the 0.5B; the measured worst is 22.3%, so the bound has headroom of under 1.4×, not an order of magnitude. The full greedy continuation is reported, not asserted, for the reason the q5_K row gives.


5. The download size gate

download::http_download used to ignore the expected size it was handed: curl -C - could exit 0 with a file of the wrong length, which the "already cached" check (a size comparison) would then never see, because it was never populated. F6 adds download::check_downloaded_size(path, expected):

  • expected == None → the length is reported but not judged (the remote size could not be determined; guessing would reject valid files);
  • expected == Some(n) and the file is n bytes → accepted;
  • otherwise → an error naming both sizes, and the partial file is removed so the next attempt cannot resume onto it.

The gate is unit-tested at the pure level and end-to-end against a local TcpListener HTTP server, with two modes: a correct 206 Partial Content resume (accepted, bytes equal), and a server that answers with a correct-looking 206 header for the requested range but ships the whole object (curl appends, exits 0, the file is larger than expected → rejected and removed). No external network is used.


6. Known gaps (follow-ups)

  • No K-quant mixture planner. The three encoders landed in #140 and minfer quantize --type q4_K writes a uniform file (llama-quantize --pure). llama.cpp's Q4_K_M/Q5_K_M per-tensor policy — the OUTPUT/tied-embedding branch, the use_more_bits Q6_K bumps for attn_v and ffn_down, the fused-QKV rule — is not implemented, so llama-quantize … q4_K (no --pure) is not reproduced. Tracked as #203.
  • The legacy targets keep tensor_type_fallback's first step only by accident. For a 2-D tensor whose row length is not a multiple of the target's block size, minfer keeps the source type; llama.cpp demotes it (q4_0 → F16) and, for the K targets, to a smaller-block type (implemented here in §2.2). For the f16 sources every gate uses the two agree, because the source type already is F16. An f32 source with such a row would differ. Filed as #204.
  • The byte-parity chain is a recipe, not a script. The f16 sources and the llama-quantize references live in ~/.cache/minfer/f6-src/ and are regenerated by the recipe in §4.2.1; nothing checks that they are the ones the record describes, and a stale cache entry would silently be compared against. #205 closed that: the record is tests/fixtures/f6-fixtures.json (§4.2.2), the checker is scripts/check_f6_fixtures.py, and the gates verify the fixture they resolve. #345 then made the recipe runnable: --regenerate (§4.2.3) re-runs a recorded producer and re-records the content it produces, idempotently, and refuses a target it cannot verify. What remains by hand: a llama-quantize reference (deliberately out of scope, §4.2.3) and any digest whose run is at a minfer commit the entry does not name — which is also why the cross-box f16 divergence below is explained (a producer-version difference) but not settled (one side's minfer commit is unrecorded).
  • qwen2.5-0.5b-instruct-f16.gguf is not byte-identical across boxes. dgxspark's copy is aef12ad44a60d2dd… (2026-09-27) and the Mac's is a26884ee1286c1d3… (2026-10-07, master ab34a72), while their f32 conversions (6894f9ea3eb79e29…) and the bf16-from-f32 reference (688109f4c9a4a8ca…) are identical — so the cause is a producer version difference, not a platform one, and it is only readable at all if the record names the minfer commit that produced each fixture. That is [#205]'s field, not this one's; the two boxes' full tables are on #333 and #205.
  • f16 weights run on CPU, CUDA and Metal for both architectures (Qwen2/Qwen2.5 and Qwen3). Op::MatMul/Op::GetRows dispatch f16 on the CPU (one weight row at a time, never an f32 copy of the weights) and the dot is vectorized — AVX2 F16C / aarch64 baseline NEON FCVTL, an f64 scalar oracle, MINFER_NO_NEON=1 forcing scalar, and the multi-token prefill decoding each row once and threading the row loop through the shared CPU pool: measured on the 0.5B at a 34-token prefill 3.2 → 217 tok/s (10.71s → 0.16s; 25.0 tok/s with the vectorized dot alone). CUDA registers the raw 2 B/element weights and converts in-register (f16_f32_matmul_vec / _scalar, embed_rows_f16), so the f16 file keeps its memory advantage (0.5B: 948 MiB of device weights vs ~1.9 GiB dequantized); an f16 prefill does not enter the int8 MMQ GEMM (which streams quantized bytes) and runs the f32-activation kernel instead. Metal gained the same pair in #164 (kernel_f16_f32_matmul + kernel_get_rows_f16, src/metal/kernels/f16.metal; both loaders' Metal arm registers F32/F16/BF16), so an f16 GGUF no longer falls to the CPU there — and its prefill runs the same f32-activation kernel, not a simdgroup GEMM. Completed by #141 (qwen2) and #167 (qwen3's loader and its graph type gate, plus the one shared registration rule in models::weight_reg); the F6 half was #49. The file contract is 2-D f16 and 1-D f32, and minfer quantize --type f16 now honours it (§2.2, #169); the CUDA norm path also refuses a non-f32 norm weight (a registered length other than d*4 bytes) instead of reading d*4 bytes out of a d*2 buffer (#169).
  • bf16 weights run on CPU, CUDA and Metal. #142 added --outtype bf16 and the CPU weight path (Op::MatMul decodes one bf16 row at a time via vec_ops::mat_mul_bf16, Op::GetRows decodes bf16 embedding rows, 1-D stays f32); #208 then registered the type and added the kernels on both devices (bf16_f32_matmul_vec/_scalar + embed_rows_bf16 on CUDA, kernel_bf16_f32_matmul + kernel_get_rows_bf16 on Metal), each device half with its own exactness gate and real-model gate. bf16 is not an MMQ format, so a bf16 prefill runs the f32-activation kernel on both, and neither device registers the attn_qkv/ffn_gu concat copies (cuda::concat_rows has no 2 B/element arm), so bf16 runs unfused.
  • The bf16 converter reference is a cast, not the converter. §4.1.1 checks minfer's bf16 output against llama-quantize --pure <f32>.gguf … BF16, because convert_hf_to_gguf.py --outtype bf16 needs torch (absent here). The cast is a valid per-tensor byte reference but shares minfer's RNE rule by construction; the direct converter check is #209.
  • general.size_label is not written (cosmetic; llama.cpp derives it from the parameter count). Every other key llama.cpp writes for this architecture is written, with the same value.
  • Tied-embedding policy follows llama.cpp for the legacy quants only. For an untied model, llama.cpp promotes output.weight to Q6_K under a sub-8-bit target; minfer keeps it at the requested target (Q6_K has no encoder yet). Documented, not silently different-by-accident.
  • minfer convert holds one tensor in memory at a time (it streams from the safetensors file into the writer), so its footprint is the largest single tensor, not the model. quantize and split stream through the mmap. The output file itself is a full copy: a 0.5B f16 conversion is ~948 MiB.

7. Using a converted model

# 1. HF -> f16 GGUF
minfer convert /path/to/Qwen2.5-0.5B-Instruct /tmp/model-f16.gguf --outtype f16
minfer /tmp/model-f16.gguf "The capital of France is" -n 8 --greedy

# 2. quantize it (and the engine runs the quantized weight kernels)
minfer quantize /tmp/model-f16.gguf /tmp/model-q4_0.gguf --type q4_0
minfer /tmp/model-q4_0.gguf "The capital of France is" -n 8 --greedy

# 3. split a single file for transport
minfer split /tmp/model-q4_0.gguf /tmp/parts --max-size 200M
minfer /tmp/parts/model-q4_0-00001-of-00003.gguf "The capital of France is"

# round-trip (rewrite with the writer), bit-exact
minfer split /tmp/model-q4_0.gguf /tmp/rewrite --max-size 4G   # one part -> a plain file

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0020 — A quantized file is byte-identical to llama-quantize, or it is wrong
  • ADR-0021 — bf16 is a round-to-nearest-even cast, and 1-D tensors stay f32
  • ADR-0024 — A machine-read record is not book content
  • ADR-0007 — No ML frameworks: every operator is hand-written

minfer Debug Dump Mechanism

Feature: --features debug_dump Controlled by: MINFER_DUMP_DIR environment variable

Overview

The debug_dump feature writes raw f32 binary files of hidden states and logits at specific points during inference. Combined with the scripts/compare_layers.py tool (which compares against llama.cpp reference dumps), this enables precise layer-by-layer numerical debugging.

Performance: When not enabled (cargo build --release), all dump code is eliminated at compile time — zero instructions, zero branches. When enabled, overhead is one OnceLock env var read + file write per dump point.


Dump Points

│  forward pass
│
├─ ⓪ minfer_dump_prompt.txt ─────────────────────────────── text dump, NOT a hidden state
│     file: minfer_dump_prompt.txt
│     dump: rendered chat template prompt (for validation against llama prompt)
│
├─ ① embed_out ─────────────────────────────────────────── token_embd (Q5_0 dequant) → hidden
│     file: minfer_dump_embed_out.f32
│     shape: [nt * ne]   (nt = num tokens, ne = hidden dim)
│     dump: embedding lookup output, before any transformer layers
│
└─  for each layer N (0..23):
      │
      ├─ ⑥ layer0_bn ────────────────────────────── RMSNorm(hidden, attn_norm) → bn
      │     file: minfer_dump_layer0_bn.f32  (only layer 0)
      │     shape: [nt * ne]
      │     dump: RMSNorm output, verification target for verify_rmsnorm.py
      │
      ├─ WQ matmul: bn × WQ → bq  [Q5_0 dequant]
      │     dump ⑧: minfer_dump_layer0_bq.f32 (first 32 values)
      │     verification target for verify_matmul.py
      │
      ├─ add_bias(bq) / add_bias(bk) / add_bias(bv)
      ├─ RoPE(bq, bk)
      │     dump ⑫: minfer_dump_layer0_bq_rope.f32 (bq after RoPE)
      │
      ├─ store KV cache
      ├─ GQA attention (f32)
      │     dump ⑬: minfer_dump_layer0_ba.f32 (attention output)
      │
      ├─ WO matmul: ba × WO → bn                            [Q5_0 dequant]
      ├─ residual: hidden += bn
      │
      ├─ ② layer{N}_attn_out ──────────────────────────── attention output
      │     file: minfer_dump_layer{N}_attn_out.f32
      │     shape: [nt * ne]
      │     dump: hidden state after attention branch, before FFN
      │
      ├─ RMSNorm(hidden, ffn_norm) → ffn_in
      ├─ gate/up matmul: ffn_in × Wg/Wu → bg, bf           [Q5_0 dequant × 2]
      │     dump ⑨: minfer_dump_layer0_bg.f32 (first 32 values)
      │
      ├─ SwiGLU: silu(bg) * bf → bg
      │     dump ⑭: minfer_dump_layer0_swiglu.f32 (bg after SwiGLU)
      │
      ├─ down matmul: bg × Wd → bn                          [Q5_0 / Q4_K / Q6_K dequant]
      │     dump ⑩: minfer_dump_layer0_fd.f32 (first 32 values)
      │
      ├─ residual: hidden += bn
      │
      └─ ③ layer{N}_out ───────────────────────────────── FFN output
            file: minfer_dump_layer{N}_out.f32
            shape: [nt * ne]
            dump: hidden state after full layer (attention + FFN + residuals)

   after all layers:

   ├─ RMSNorm(hidden, output_norm) → bn
   ├─ LM head: bn × output.weight → logits                 [Q8_0 matmul]
   ├─ add output_bias
   │
   ├─ ⑤ last_norm ────────────────────────────────────── post-RMSNorm hidden
   │     file: minfer_dump_last_norm.f32
   │
   └─ ④ logits ────────────────────────────────────────── final logits
         file: minfer_dump_logits.f32
         shape: [nt * n_vocab]
         dump: raw logits before sampling


   ─── Utility dumps (not layer-specific) ───

   ⑦ minfer_dump_q8_quant_verify.txt ──────────────────── Q8_0 quantize
         fired once on first quantize_row_q8_0_buf call
         dump: amax, d, x[0], x[1], x[16], q[0], q[1], q[16]
         verification target for verify_q8_quant.py

minfer Side

Feature flag

# Cargo.toml
[features]
debug_dump = []

Core module: src/dump.rs

#![allow(unused)]
fn main() {
// For float arrays (hidden states, logits)
pub fn maybe_dump(name: &str, data: &[f32])

// For text (prompts)
pub fn maybe_dump_text(name: &str, text: &str)
}
  • Both controlled by MINFER_DUMP_DIR env var
  • If not set → no-op (returns immediately)
  • If set → maybe_dump writes raw f32 bytes to {MINFER_DUMP_DIR}/{name}.f32
  • If set → maybe_dump_text writes UTF-8 string to {MINFER_DUMP_DIR}/{name}.txt
  • Uses OnceLock to cache the env var check (read once, reused)

Build & run

# Normal build — zero overhead, dump code eliminated at compile time
cargo build --release

# Debug build — dump enabled
cargo build --release --features debug_dump

# Run with CPU path + dump
MINFER_DISABLE_MPS=1 MINFER_DUMP_DIR=/tmp \
  cargo run --release --features debug_dump -- <model> "Hello"

Python Side

scripts/dump_llama_ref.py

Generates llama.cpp reference hidden states for comparison.

# Bare text (no chat template wrapping)
uv run python -m scripts.dump_llama_ref \
  --model <path-to-gguf> \
  --prompt "Hello" \
  --output ./llama_ref

# With chat template — reads tokenizer.chat_template from GGUF,
# renders with Jinja2, producing the same prompt minfer would use
uv run python -m scripts.dump_llama_ref \
  --model <path-to-gguf> \
  --prompt "Hello" --chat \
  --output ./llama_ref

Output (per layer):

FileShapeContent
layer{N}_hidden_states.npy[hidden_dim]Last-token hidden state after N layers
logits_prefill.npy[vocab_size]Final logits
token_ids.npy[seq_len]Input token IDs
prompt.txt—Rendered prompt text (for validation against minfer's minfer_dump_prompt.txt)

How it works: Uses the "truncated model" technique — creates a fake GGUF with block_count=N (metadata only, zero-copy weights), runs llama.eval(), and extracts the embedding output via llama_get_embeddings(). Since a decoder-only transformer is feed-forward, the first N layers produce identical outputs to a full model.

When --chat is specified, the script reads tokenizer.chat_template from GGUF metadata, renders it with Jinja2 using messages=[{"role":"user", "content": prompt}] and add_generation_prompt=True, producing a prompt identical to minfer's chat template rendering.

scripts/compare_layers.py

Compares minfer dumps against llama.cpp reference.

uv run python -m scripts.compare_layers \
  --llama-dir ./llama_ref \
  --minfer-dir /tmp \
  --hidden-dim 896

Layer mapping: llama layer N (N=1..24) = minfer layer N-1 (0..23). Both dump the hidden state AFTER the corresponding layer's computation.

Prompt validation: Before layer comparison, the script reads prompt.txt (from llama reference) and minfer_dump_prompt.txt (from minfer dump). If both exist and differ, the script immediately aborts with PROMPT MISMATCH — comparison aborted and sys.exit(1).

Output: Per-layer:

ColumnMeaning
minfer RMSRMS of minfer's hidden state
llama RMSRMS of llama.cpp's hidden state
ratioRMS ratio (should be ~1.000)
cosCosine similarity (≥0.999 = match)

Plus logits comparison with top token and cosine.


Full Workflow

# Step 1: Generate llama.cpp reference (one-time, ~20 min for 24-layer model)
#         --chat ensures the SAME prompt as minfer (rendered from tokenizer.chat_template)
uv run python -m scripts.dump_llama_ref \
  --model ~/.cache/minfer/models/hf/Qwen/Qwen2.5-0.5B-Instruct-GGUF/qwen2.5-0.5b-instruct-q4_k_m.gguf \
  --prompt "Hello" --chat --output ./llama_ref

# Step 2: Run minfer with debug dump (CPU path)
#         MINFER_DUMP_DIR enables all dump points including prompt
MINFER_DISABLE_MPS=1 MINFER_DUMP_DIR=/tmp \
  cargo run --release --features debug_dump \
  -- <model> "Hello"

# Step 3: Compare
#         Script validates prompt files match before comparing layers
uv run python -m scripts.compare_layers \
  --llama-dir ./llama_ref --minfer-dir /tmp --hidden-dim 896

Diagnostic Logic

Given compare_layers.py output showing first divergence at some layer, use the sequence below to isolate the bug:

Divergence patternLlama refMinfer dumpDiagnosis
Prompt mismatchprompt.txtminfer_dump_prompt.txtTemplate rendering or tokenizer differs between llama and minfer
embed_out already divergedlayer1_hiddenembed_outQ5_0 embedding dequant is wrong
layerN_out diverged, layerN_attn_out oklayer{N}_hidden vs layer{N-1}_hiddenattn_out vs layer_outFFN matmul is wrong (gate/up/down Q5_0 dequant)
layerN_attn_out already divergedlayer{N-1}_hiddenlayer{N-1}_out vs layer{N}_attn_outAttention matmul is wrong (Q/K/V/WO Q5_0 dequant)
24 layers all match, logits divergelayer24_hiddenlayer23_outoutput.weight (Q8_0) matmul is wrong
All cosine ≥ 0.999, logits match——Bug is elsewhere: sampler, or model architecture hparams mismatch

Layer numbering

llama.cpp dumpminfer dumpComputation
layer1_hidden_states.npyminfer_dump_layer0_out.f32Hidden after layer 0 (attention + FFN)
layerN_hidden_states.npyminfer_dump_layer{N-1}_out.f32Hidden after layer N-1
logits_prefill.npyminfer_dump_logits.f32Final logits

Path Verification Status (2026-07-28)

All CPU inference paths verified correct through automated cross-validation against gguf.quants.dequantize() (validated against llama.cpp C reference in gguf-py/tests/test_quants.py).

Verification Results

#PathScriptMethodResult
1Q5_0 embedding dequantverify_q5_embed.pyvs minfer_dump_embed_out.f32✅ exact match (8 values identical)
2RMSNormverify_rmsnorm.py --comparevs minfer_dump_layer0_bn.f32✅ cosine = 1.0000000000
3Q8_0 quantizationverify_q8_quant.pyvs minfer_dump_q8_quant_verify.txt✅ q[0]=q[1]=q[16] identical
4Q4_K scalar dot producttest_q4k_dot_simpleunit test✅ passing
5Q8_0 scalar dot producttest_q8k_dot_simpleunit test✅ passing
6Q6_K scalar dot productreference_dot_q6kunit test✅ passing
7Row stride (all tensors)dump_tensors.py vs matmul wsmanual✅ correct
⑧WQ matmul outputverify_matmul.py vs bq dumpdot product✅ cos = 1.0000000000
⑨FFN gate matmul outputverify_matmul.py vs bg dumpdot product✅ cos = 1.0000000000
⑩FFN down matmul outputverify_matmul.py vs fd dumpSwiGLU + dot✅ cos = 0.9999848730
⑫RoPE rotationverify_rope.py vs bq_rope dumpfreq + sin/cos✅ cos = 0.9999999942
⑬GQA attention (nkv=1)verify_attention.py vs ba dumpV lookup✅ cos = 0.9999839613
⑭SwiGLU activationmanual vs bg dumpsilu×up✅ plausible

Despite all 14 paths being verified correct, the Q5_K_M model still produces garbled output. Model file confirmed working with llama-cli. Root cause is a subtle integration issue not captured by individual verification.

Diagnostic Status (2026-07-29)

ItemStatusMethod
Q5_0 dequant formula✅gguf.quants.dequantize() (C-validated)
Weight tensor layout✅Raw bytes match Python
RMSNorm✅cosine = 1.0 against reference
Q8_0 quantization✅amax/d/q values match
Unit tests (Q4_K/Q8_0/Q6_K)✅all passing
Row stride✅formula matches GGUF physical layout
Matmul outputs (WQ/gate/down)✅cosine = 1.0
RoPE✅cosine = 0.9999999942
GQA attention✅cosine = 0.9999839613
SwiGLU✅values plausible
Residual connections✅fd contribution matches exactly
Model metadata comparison✅Q4_0 vs Q5_K_M: identical architecture
llama-cli on Q5_K_M✅produces correct output
Per-layer verification⚪inconsistent due to f32/f64 precision diff
Cross-model per-layer (Q4 vs Q5)⚪weights differ, cannot compare
KV cache integration⬜not verified
Generation loop interaction⬜not verified

Verification Method

gguf.quants.dequantize() ── GGUF file ──→ Python reference values
        │                                       │
        │ (validated against C)                  │ compare
        │                                       │
minfer's own computation  ── forward pass ──→ minfer dump (.f32 / .txt)

The Python side uses gguf.quants.dequantize() from gguf-py, which is the same implementation validated in llama.cpp's test_quants.py (quantize + dequant must be bit-exact against the C reference). This eliminates the "Python formula might be wrong" concern — the verification chain traces back to llama.cpp's C implementation.

Row Stride Verification (2026-07-28)

The hypothesis that matmul ws (computed as (id/blck_size)*type_size) diverges from the GGUF physical row stride was tested and disproven:

TensorTypeShapeGGUF row strideMatmul wsMatch
blk.0.ffn_downQ6_K[4864,896]3,9903,990✅
blk.11.ffn_downQ4_K[4864,896]2,7362,736✅
blk.0.ffn_gateQ5_0[896,4864]616616✅
blk.0.attn_qQ5_0[896,896]616616✅

The formulas are inherently consistent: both derive from (ne[0]/blck_size)*type_size.

Matmul Output Verification (2026-07-28)

All three matmul types verified correct using full forward-computation in Python, including bias and SwiGLU:

TensorTypeShapeCosineMethod
WQQ5_0[896,896]1.0000000000RMSNorm + bias
FFN gateQ5_0[896,4864]1.0000000000post-attn RMSNorm
FFN downQ6_K[4864,896]0.9999848730SwiGLU + dot product

Despite all 10 paths being verified correct, the Q5_K_M model still produces garbled output. Remaining unverified: RoPE, GQA attention, SwiGLU, KV cache, residual connections. Root cause remains unidentified.

Verification Scripts

The verification scripts in scripts/verify_*.py provide standalone Python reference implementations for independent validation of minfer's computation. All dequantization uses gguf.quants.dequantize() — the same Python implementation validated against llama.cpp's C reference in gguf-py/tests/test_quants.py (quantize + dequant must be bit-exact).

ScriptVerifiesUsage
verify_embed.pyToken embedding dequant (auto-detect type)uv run python -m scripts.verify_embed --model <gguf> --token-id <id>
verify_rmsnorm.pyRMSNorm output (bn)uv run python -m scripts.verify_rmsnorm --model <gguf> --layer 0 --token-id <id> --compare <dump>
dump_tensors.pyTensor layout (offset/n_bytes/shape/type)uv run python -m scripts.dump_tensors --model <gguf>

These read the same GGUF file as minfer, perform the same computation using the validated gguf.quants.dequantize(), and print key values for comparison with minfer dumps (under --features debug_dump).

Two additional verification paths are covered by Rust unit tests:

TestFileVerifies
test_q4k_dot_simplequants.rsQ4_K × Q8_0 scalar dot product
test_q8k_dot_simplequants.rsQ8_0 × Q8_0 scalar dot product

Output Files (minfer)

FileDump pointShapeBytes (Qwen2.5-0.5B, prompt="Hello"≈30 tokens)
minfer_dump_prompt.txt⓪ prompttext~200 B
minfer_dump_embed_out.f32① embedding30 × 896 = 26,880~105 KB
minfer_dump_layer0_bn.f32⑥ RMSNorm30 × 896 = 26,880~105 KB
minfer_dump_layer0_attn_out.f32② attention26,880~105 KB
minfer_dump_layer0_out.f32③ FFN26,880~105 KB
...
minfer_dump_layer23_attn_out.f32② attention26,880~105 KB
minfer_dump_layer23_out.f32③ FFN26,880~105 KB
minfer_dump_logits.f32④ logits30 × 151,936 = 4,558,080~17.4 MB
minfer_dump_q8_quant_verify.txt⑦ Q8_0 quantize1 line text~100 B

Total: 24 × 2 × 105 KB + 17.4 MB ≈ 22 MB for a 24-layer model.

The gate contract

Scope: what a gate is in this repository, and the five rules this campaign learned the hard way. A gate is any test whose verdict is a claim about the engine — "this path ran", "this value is right", "this is not slower" — as opposed to a test that pins a pure function's output. This document is the one home for those rules: AGENTS.md carries only the point-of-use facts (the command, the wrapper's default, the current dated counts) and points here.

The rules are stated once, here. The per-ticket records that produced them live in ARCHITECTURE-EXECUTION-PLAN.md, CUDA-BACKEND-DESIGN.md and BUILD.md, and are linked as precedent rather than restated. Reading a record is how you check a rule against the instance it came from; this page is what you apply while writing the next gate.

1. Assert the value, not a relation between two code paths

A gate that compares mode A against mode B is blind to any fault the two modes share. Ask what the assertion would do if the implementation were wrong in the same way on both sides: if the answer is "pass", the gate needs a value arm.

Precedent. cuda_map_window_matches_the_span_over_the_same_rows swept f32/f16/q8_0 and compared each mode against another; dropping the Q8_0 block base in kv4 left it green, because both sides of the comparison read the same wrong base. A single-row arm that compares the window against the dequantized cell was added and catches it (mutation-checked, #87, recorded in ARCHITECTURE-EXECUTION-PLAN.md §"All three acceptance gates, mutation-checked"). Rules 1 and 2 are a pair: the control arm below is what makes a relative assertion meaningful, and this rule is what makes an absolute one necessary.

Apply it. Every device value gate should include one arm whose expected value is computed independently of the path under test (a dequantized reference, a scalar oracle, a pinned literal). A mode-vs-mode comparison is the fast arm, never the only arm.

The isolation instance of this rule is #173: a gate that compared two snapshots of a process-global timing table was replaced by exact values read from a per-scheduler sink, because a concurrent execution could move the relation (record: ARCHITECTURE-EXECUTION-PLAN.md §"#173"). The device twin is #185: the F5 "host stalls removed" gate read the process-wide cuda::stream_sync_count() and now reads the CudaBackend's own counter, because a concurrent device test's syncs landed inside the delta (4160 vs 728 in the run that exposed it). Rule 1 is about a shared code path; a shared destination is the same hazard, and a value read from an owned table is immune to it.

The orphaned-entry-point instance (#218, #223). A gate must exercise the production entry point, never a test-only helper that mirrors it. #218 found that the #145 gate cuda_prefill_smem_optin_covers_every_launchable_instantiation called gemm_prefill_smem_init — an eager startup sweep whose CudaState::try_new call site #188 had quietly deleted. The helper still worked, so the gate stayed green while production performed no such opt-in at all; the dead-code pass then annotated the orphan #[cfg_attr(not(test), allow(dead_code))] instead of asking why a production-looking function had no production caller. A mirrored helper is the same hazard as rule 1's shared code path: the gate and the production path can be wrong together, and the mirror is what makes it look certified. Ask of every gate: is the function it calls reachable from a production entry point? If the answer is a test-only sweep, the gate's claim is about the sweep, not the engine — either drive the real entry point (the #218 gates drive gemm_smem_optin through a real forward and through the production launcher) or rename the gate so its claim states the mirror it tests. The campaign corollary: a dead-code diagnostic that allows a test-reachable production-looking item is a question deferred, not a warning silenced.

That question has two legitimate answers. Deleting the item is one. #223 took the other and restored the production call — CudaState::try_new drives the eager pre-warm again, through the same per-instantiation cache the launcher reads, so the attribute is set before any CudaBackend (the only holder of a capture window) can exist. Annotating the item without answering the question is neither.

The unrun-configuration instance (#332). A gate is only as wide as the code its configuration actually compiles. scripts/check_dead_code_oracle.py judges liveness by stripping every allow(dead_code) and reading rustc; on Linux it can say nothing about a #[cfg(target_os = "macos")] module, and a run that reported "0 additions" there would be read as coverage of code it never compiled. The checker therefore carries the platform as an explicit configuration (--config macos), refuses it by name on a non-Mac host, and records the state in the manifest: macos = "unjudged" while no measurement exists — a verdict absent, never a pass — and, since it was run on a Mac 2026-10-09 (macbook (macOS 27.0.1, Apple M4 Pro), 42 items, 0 additions after seeding the set), the judged [[macos]] section instead. The list of configurations is part of a gate's claim: adding one is how a blind spot closes, and the marker is how a not-yet-run one stays visible until someone runs it.

Corollary — the dead-code annotation rules

AGENTS.md Core Convention 5 states the one-line rule; this is the full contract. Moved here from AGENTS.md (which keeps the point-of-use line and links back).

The non-test build is warning-free, and gated. src/main.rs carries #![cfg_attr(not(test), deny(warnings))], so any new warning fails cargo build --release on every platform — and the bin half of cargo test --release, which all three build jobs exercise. PR #216 cleared the 40-diagnostic dead-code backlog; this keeps it cleared. The not(test) scope is deliberate: the test build's own warnings (unused locals and imports inside test code) are a separate, tracked cleanup, not part of this gate. When code is only reachable in some configuration, silence it at the item with #[cfg_attr(<cfg>, allow(dead_code))] and say why — never widen the crate-level gate. The <cfg> must name the configuration in which the item is unused — not(feature = "cuda") for a device-only reader, not(test) for a test-only one — never a blanket not(test) when the real gate is a feature (#243). A genuinely test-only item names the consuming test (or the feature cfg) in that note; an item that looks like production code and has no production caller is not silenced but asked about — why is there no caller? — because its own comment can claim a caller that a later cleanup deleted while the annotation kept the build green (the #218 orphan; answered in #223). A deliberately retained deferred item — a variant, dtype or field the data model can express but no architecture constructs yet — says in its note what would construct or read it (a CLI spelling, a kernel, an architecture), so "kept for the vocabulary" is a statement about the design and not a restatement of the lint (#244). Both halves of that convention are mechanical since #254: scripts/check_dead_code_annotations.py (CI check-docs) rejects a bare allow and an annotation without a reason, and the stripped oracle scripts/check_dead_code_oracle.py (last step of test-linux-cpu / build-linux-cuda; the third configuration, --config macos, is the plain cargo check --release run on a Mac, refused on any other host — #332) fails when a new allow hides an item rustc reports once every annotation is stripped — compared against scripts/dead-code-baseline.toml — beside this checker, per ADR-0024 — whose addition is a decision made in the same PR. Those two jobs are skipped when the change classifier finds no code change (#457; the decision is ADR-0025), and that cannot hide a new annotation: a new allow(dead_code) is a source change, so the classification runs them — and so is a baseline-only edit, because the manifest lives under scripts/, which the classifier treats as rust.

2. A control arm must differ in the property under test

A negative control that is rejected by an earlier check never exercises the check the gate is about. The control must be constructed so that the only thing that can refuse it is the property under test.

Precedent. The q4_K W_dsc plane gate (#165, #167) originally used a q8_0 payload as its negative. A q8_0 payload is longer than a q4_K payload, so the payload-length check rejected it before the type check ran — the gate could not see a type-gate bypass. The fix was an equal-ratio q4_0 arm: q4_0 has q4_K's exact bytes-per-element ratio (od * (id / 32) * 18 in the unit test), so only the type gate can refuse it. See models::weight_reg::tests::q4k_weights_register_the_dsc_plane_only_under_every_gate.

Apply it. For each condition in a conjunctive rule, ask which single arm can fail only because of that condition. If the control is refused earlier (a length check, an alignment check, an unset flag), it is not testing the condition.

The degenerate-input instance. An arm can also fail to differ when the fixture lets the difference cancel, even though the code under test is wrong. #186's new decode arm compared the packed int K dot against the dequantized-f32 reference — but it stored a single KV cell, so the softmax had one key and the K score cancelled out of the output entirely; mutating the __dp4a block base (elem >> 5 → elem >> 4) still passed. The fix is to make the difference observable through the arm's own data — here, two cells read through an explicit [0, 2) span, so the score reaches the softmax weights — and then the same mutation is red (max |Δ| = 0.35126442). This is the same family as #145's self-clearing assertion: before trusting a control, ask what the fixture does to the quantity the control is supposed to move, and assert that the quantity can move at all.

3. Every gate needs mutation evidence

Break the thing the gate guards — with the implementation, not the test — watch the gate go red, revert. A gate that has never failed is a gate that has not been shown to test anything.

Precedent. Five separate gates in this campaign passed for the wrong reason and were found only by deliberately breaking the implementation: relation-blind (rule 1), a self-clearing assertion, process-global registry contamination, a wrong-message assertion, and a wrong-name assertion. The records are in CUDA-BACKEND-DESIGN.md (the #147, #151, #165 hardening sections each end with a "Mutations" paragraph) and ARCHITECTURE-EXECUTION-PLAN.md.

The seam makes it cheap. src/testfail.rs is the one failure-injection switch (issue #171). One environment variable at the command line replaces the bespoke mock the earlier tickets each had to write:

sitechokepointwhat the mutation looks like
forward_batchserver::chat::guarded_forward_batcha panic through the real guard → the 500 path (#151)
alloc_in_poolgraph::alloc::GraphAllocator::alloc_in_poolthe allocation refuses before the pool is touched
execute_nodegraph::scheduler::BackendScheduler::executethe backend dispatch refuses
register_weightmodels::weight_reg::register_cuda_weightthe CUDA weight registrar panics
metal_cross_copygraph::metal_backend::MetalBackend::cross_enqueuethe staging copy's MTLSharedEvent signal is suppressed, so phase B's bounded wait takes its real timeout branch (#137)
launch:* / attr:*src/cuda/kernels/*.cu (minfer_launch_ok / minfer_smem_optin)the real CUDA call is driven into failure (#147)

The matching rule is exact-token and comma-separated (MINFER_TEST_CALL_FAIL=forward_batch,launch:gemm_f16_f16; all for every Rust site). The seam is off whenever the variable is unset, which is CI, every normal run and the compute-sanitizer run; the property is pinned by testfail::tests::the_seam_is_off_by_default. It is test-only: a chokepoint is an ordinary call that returns Ok/does nothing in production.

The observation half. A gate that must prove a path executed cannot read the path's own answer — a dispatch function naming its branch is self-certifying. testfail::note_checked(site) / checked(site) are bumped by the chokepoint itself (the shape #141's vectorized f16 dot needed: F16_SIMD_PATH_CALLS, asserted by vec_ops::tests::f16_dot_uses_the_vectorized_path). A gate asserts the counter advanced instead of trusting the dispatch's report.

The work-bound twin (#160). Rule 4's progress assertion has its own mutation lever: MINFER_TEST_TICK drives BatchEngine::tick into one of two faults. =wedge returns from the step without advancing work_units, so the per-step progress assertion fires on the step that wedged the engine; =spin advances the counter but never completes a run, so only step_budget catches it. Both are read once per process and unset in every production, CI and default run, exactly like MINFER_TEST_CALL_FAIL.

Honest scope — presence is checkable, truth is not. A script can require that a mutation transcript exists in a PR body or a record; it cannot check that the mutation was real, that the gate was the one that failed, or that the revert was byte-identical. Rule 3 is a discipline, not an enforced property. The durable part of #171 is making the experiment one line, because the cost is what made the discipline get skipped. The machine-checkable half of the sibling rule 5 (a suite-count consistency check) stays separate, on #94.

4. Bound a gate's runtime by work, not by seconds

A verdict that depends on how fast the machine was at that moment is a verdict about the machine. Absolute deadlines and single wall-clock ratios both failed on a loaded box in this campaign.

Precedent. BUILD.md §Tests and ARCHITECTURE-EXECUTION-PLAN.md §"#154" / §"#158" hold the records:

  • #154: server_batch_matches_serial_and_is_faster measured two whole workloads once, sequentially, and asserted t_serial > t_batched. Under a parallel harness the first phase absorbed the start-up wave: 21.20s batched vs 9.95s serial = 0.47x in parallel against 1.50x serially. It now interleaves matched rounds and asserts the median of the per-round serial/batched ratios > 1.0, so a verdict is robust to up to half the rounds being disturbed.
  • #158: published_metrics_move_as_requests_are_served bounded a run by absolute wall-clock deadlines and asserted the engine had gone idle, so 16 extra CPU spinners made it panic (28 passed / 1 failed). It now bounds work: BatchEngine::work_units must advance on every step that leaves the engine busy, plus step_budget(prompt, max_tokens). The same spinner run is green, and failure detection went from 120s to 0.18s.

Apply it. Prefer a counted invariant (steps, rows, tokens, work units, entries walked) over a clock. When a timing relation is unavoidable, use interleaved matched rounds and a median (never two sequential sums), fix the round count in advance, and print every per-round value so a loaded result is auditable. A process-hang watchdog may keep a generous timeout, but it must not be a gate's only failure signal.

#160 finished the sweep #158's audit opened. Every while engine.busy() stepper in the #[ignore]d server gates now runs through one shared per-step WorkBound — a progress assertion plus step_budget — so a wedge in BatchEngine::tick fails the gating step in seconds instead of hanging the suite. #196 closed the case #160 could not: the two gates that drive the production serve_loop were outside any in-test bound, because a wedge kept the loop busy and the test thread never returned. serve_loop itself now carries the same invariant as a production guard — STALL_STEP_LIMIT consecutive steps that left the engine busy without advancing BatchEngine::work_units end the loop, answer every live and queued request once with a 500 server_error, and move the minfer_worker_stalled_total counter — so both gates are wedge-proof with no test-side deadline, and a wedged server no longer spins at 100% CPU with clients left hanging. The wall-clock bounds that remain are named and say at the call site why each is only a backstop: the serve_loop feeder's poll terminator (FEEDER_POLL_BACKSTOP, now the last resort for a worker whose published metrics never settle rather than the only way out of a wedge), and the two run_cli child-process ceilings (1800s for the real-model sessions, 60s for the no-model registry cases) are cross-process hang guards, env-overridable through MINFER_CLI_WATCHDOG_SECS. The mutation seam is MINFER_TEST_TICK (see §3); the dated records and transcripts are #160 and #196 in ARCHITECTURE-EXECUTION-PLAN.md §test-infrastructure.

5. A device number carries its date, device and command

A count or a timing without provenance cannot be audited and cannot be corrected when it drifts.

Precedent. The AGENTS.md suite counts drifted twice by hand-editing; #94 records three instances of the class. The fix is not more diligence, it is making every number self-describing.

Apply it. Write device and suite numbers as <counts>, <device>, <command>, <date> — for example 36 passed / 0 failed, dgxspark (aarch64, GB10 sm_121), FEATURES=cuda scripts/real_model_gates.sh, 2026-09-25. A number whose command you cannot name is not evidence; delete it rather than re-state it.

Name the box absolutely. A record is read on machines other than the one that produced it — an agent on an x64 CUDA box reads AGENTS.md too — so box is the machine's own name, never a relative one: no this box, my machine or local. The labels in use today are dgxspark (aarch64, GB10 sm_121) (the maintainer's DGX Spark — one physical machine, so its CPU and CUDA rows share the label) and x86_64 (CI runner) (GitHub's runner). A newly used machine gets its own label — its hostname, or another equally stable id, plus the arch/device that matters for the counts — and its own [[counts]] rows; a suffix or a re-used label would make two machines' numbers indistinguishable. The label is part of the machine-checked identity: it is the box field in scripts/test-baselines.toml and the --box argument of scripts/check_baselines.py --check-live, so a rename must move both the manifest and the prose in the same commit. A box whose name has no manifest row fails --check-live loudly, which is deliberate: a fabricated box must not pass vacuously.

The PR body's Mac verification section is this rule on the one path CI cannot run (#335). build-macos compiles the crate and — since #303 — the test target, but it has no Metal device, so on the macOS/Metal path the Mac-local run is the evidence. The template requires the section (.github/PULL_REQUEST_TEMPLATE.md, enforced by the check-pr-body job and scripts/check_pr_body.py), and its content convention is the one stated above: <box label> / <date> / <command> with the machine's own absolute label, or N/A — <reason> when the change touches no macOS-only file and no shared layer's macOS arm. It is the place where what was verified on which box, and what was not becomes part of the record — the record a Linux-green PR otherwise lacks. Why it earns a section rather than a sentence in someone's memory: the predicate is not confined to src/metal* — git grep -F 'cfg(target_os = "macos")' -- src/ finds 111 sites in 36 files today (the issue that introduced the section measured 35 / 108 on 2026-10-07; #299 and #329 moved the row), and they include shared layers (graph/alloc.rs, graph/kvcache.rs, graph/kvformat.rs, graph/scheduler.rs, graph/backend.rs, graph/registry.rs, graph/fusion.rs, models/weight_reg.rs, every models/*/graph*.rs) — while the macOS-only assertions no Linux job executes are the 53 unit tests the current recorded rows differ by (macOS 541 − aarch64 488, docs/status.toml 2026-10-07) and 11 integration tests (macOS 21 − aarch64 10, the four #![cfg(target_os = "macos")] binaries in tests/), plus the op matrix's Metal column and every performance number. The checker enforces that the section is present and filled; whether it is true is the reviewer's job (#175).

Prose anchors: name the section, not the line range

This is a documentation convention, not a sixth gate rule — it belongs here because check_doc_line_anchors.py is the gate it supports, and because the reason it exists is the gate's own boundary.

A path:NNN anchor in the docs is checked for three questions, each of which fails the run: does the file and the line exist; does a backticked symbol sitting next to the anchor still live within ±25 lines of it (#266); and is that symbol inside the cited range (rule E, #339). The ±25-line window is loose on purpose (a symbol 20 lines away is a neighbour, not a move), which is why the narrower in-range claim is checked separately: a symbol that is in the file but not in the cited range fails, reported as range-miss in the summary and by --list with [range-miss: …]. A call-site citation is not a miss: if the range names the symbol where it is called, the identifier is inside the range and the anchor passes. Neither test can see a range that still resolves but no longer holds the text the citing sentence describes, and that is not a bug to fix with a further rule: measured for #327, 173 anchors carry an adjacent backticked non-symbol token, and 99 of the 144 live ones do not contain that token inside their own range — 69 % of the population, and the tokens are mostly code expressions (ins[1][t].to_bits() as usize, nt >= 9) or fragments of a neighbouring quoted sentence, never the claim. Even the narrowest useful spelling ("≥2 pure-lowercase words") leaves one hit — an anchor citing build.rs for `ar rcs` that no longer held it — and that one was a real drift, which is the point: the detectable subset is a coincidence, not a rule. Rule E is the symbol half of that question, where the adjacent token is mechanical; the non-symbol half stays a convention.

So the convention is:

  • A prose anchor points at a section heading. Cite the heading by name and number (docs/COMPUTE-GRAPH-DESIGN.md §7.3 "In-place execution and the aliasing rule") next to the range, so a reader who lands on the wrong text has a stable second locator. A heading rename is a visible edit; a 200-line insert is not.
  • Prefer a symbol anchor where one exists. check_doc_line_anchors.py rule C is the one content-aware test that is mechanical: put the backticked symbol next to the anchor and the checker follows it through a move. A bare range is the fallback, not the default.
  • A symbol anchor's range must hold that symbol (rule E). Write the symbol and cite the lines that contain it — a definition citation starts at the definition, and a NNN-MMM range must not start hundreds of lines early. A trailing () is part of the token, so `metal_available()` is checked like `metal_available`. If the sentence is about a call, cite the call site; the identifier is there too. When neither is possible, cite a section heading (the bullet above) rather than a range that holds something else.
  • A range that quotes prose is the weakest form. Do not backtick the quoted words to "make them checkable": the adjacency rule was written for identifiers, and applying it to prose judges the neighbouring sentence (measured above).

Rule E shipped as a note from #339 to #371 because of what a failure would have cost then and because it does not catch the whole class. Measured on 7991218, 79 non-frozen anchors carried an adjacent backticked symbol whose identifier was not inside their own range (84 with the trailing () stripped; 46 cited a NNN-MMM range and 38 a single line, and 81 of the 84 hold the identifier in the target file at all), and they spanned 18 docs — 21 in 13-decode-loop-graph-reuse.md, 12 in cuda_tutorial/02-minimal-cuda.md, 11 in 14-metal-backend.md, 7 in ARCHITECTURE-ROADMAP.md. 26 of the 84 lived in walkthrough docs that state the revision their lines were verified against (lines verified at commit e7fa0da), where re-pointing the anchor would falsify the record — the same reason the FROZEN set exists. #371 re-pointed all 81 — the revision-pinned docs included, whose claim was re-stated at that PR's base rather than frozen — and promoted the rule: a range miss fails the plain check-docs run, and --strict-symbols promotes only the two heuristic classes (symbol-far / symbol-foreign). And two of the three anchors #339 was filed for are invisible to the rule: :256 cited metal_backend.rs:345-1102 for execute_node, a range that does contain the definition at 843 (containment cannot see a range that starts 527 lines early), and :586's synchronize anchor carries no adjacent symbol at all (`self.submit_pending()` is not an identifier), so no adjacency rule can reach it. Both were re-pointed by hand in the #339 PR, as was the () (:526 cites metal_backend.rs:2098-2100 now). What remains of the class is its other half — a range that carries no symbol at all, so no adjacency rule can see whether it still holds the text its sentence describes; that bare-range sweep is #336 and #356.

The remaining bare ranges are a tracked sweep, not a silent one: 878 of the 955 resolved anchors are bare ranges across 33 docs (docs/cuda_tutorial/** alone carries 276), and converting them is its own ticket.

The drift rule: re-point by the map, in the same PR

Rule C follows a symbol through a move, but a bare range has nothing to follow: a commit that inserts a line in an anchored file leaves every anchor below it resolving and wrong. That is the gap scripts/check_anchor_drift.py (#344) closes, and it runs in check-docs on pull_request events. It diffs HEAD against the PR's merge base — git diff -U0 origin/<base_ref>...HEAD, which is why that job's fetch-depth: 0 is load-bearing twice — reduces each file to its hunks, and compares the base revision's anchors with the head tree's. Anchors are paired on the doc line with its numbers normalised (so a re-pointed line still pairs), which is PR #343's own proof (":NNN" -> ":N" multisets identical) applied per line:

  • the pair does not carry the mapped numbers → stale, printed as doc:line → target:old (now new), exit 1;
  • a cited endpoint was deleted by the range, so the map has no image for it → ambiguous: reported with the reason, never guessed, and promoted to a failure by --strict;
  • the range never touched the target, or the doc already carries the mapped numbers → silent.

The rule for a PR author is therefore: moving a line in an anchored file obliges you to re-point every anchor into it, by the mapped delta, in the same PR — and the cheapest way to need that less often is rule C above, a symbol anchor the map cannot strand. The gate exists because the rule was missed twice in one round: #329's dispatch change moved 7 anchors off by one (docs 07 and 14), and #299's weight accounting moved 69 across eight docs by 1–3. Both were caught by a hand-written old→new line map, which is exactly what the gate replaces — and which it beats: re-run over #299's range it named 73, the 69 the re-point commit fixed plus 4 it missed. Those four are cited here by the revision that carried them, not by a live anchor: the PR that added this section re-pointed all of them, so the four doc lines that held the stale numbers now hold the corrected ones — line 98 and line 452 of docs/ARCHITECTURE-ROADMAP.md, line 56 of docs/SOURCE-LAYOUT-PLAN.md and line 705 of docs/inference_e2e_walkthrough/03-model-dispatch-weights.md, where the cited graph/alloc.rs range moved from 134-137 to 135-138. Read that line at 506b26c to see the defect; read it today to see the fix.

Two boundaries are deliberate. A base anchor whose doc line was rewritten beyond its numbers is not compared — the author touched that line, and pairing it would be a guess; --list prints it as not compared, so "why did this pass?" has an answer. The FROZEN set above is the same set in both modes: a record frozen against a past revision is not drift. The other boundary — a bare :NNN continuation — is no longer a blind spot: it is rule D below.

A bare :NNN continuation attaches to the anchor before it

A citation often names a file once and then a second range beside it:

`docs/ARCHITECTURE-ROADMAP.md:NNN` … `graph/alloc.rs:NNN`, `:MMM`

The :MMM is a bare continuation — a backticked span whose whole content is :NNN or :NNN-MMM. ANCHOR requires a path.ext, so before #355 neither check_doc_line_anchors.py nor check_anchor_drift.py could see it: a re-point that fixed the visible graph/alloc.rs:NNN left :MMM behind, silently. That is not hypothetical — it is how #347's seven rows survived 387fe91 and da7a35a, and check_anchor_drift.py reported green over each. (The numbers are placeholders: this page states the rule, it does not cite a revision, so it carries no live anchor for the gates to keep re-pointing.)

The convention, stated once here and implemented as rule D:

  • A bare continuation attaches to the nearest preceding path.ext:NNN match on the same doc line. "Nearest preceding" is exact: the window is one doc line, the direction is backwards only, and the nearest earlier anchor wins — so a line carrying two path anchors and two continuations attaches each number to the anchor immediately before it, not to the first one on the line. The continuation is then resolved and range-checked against that file, exactly like a written path:NNN, and the drift gate maps it through that file's hunks in the same way — a stranded continuation is doc:line (bare continuation) → path:old (now new), exit 1.
  • A continuation that follows no anchor on its line stays silent. So does one that comes before every anchor on its line, and one whose path anchor is itself external, ambiguous or frozen. The rule is syntactic, so the silent class is not a guess about intent: nothing on the line names the file, and attaching to the nearest anchor in either direction is wrong where it is not empty. Measured on f754ee3, the two standalone :1289-1321 spans in docs/ARCHITECTURE-ROADMAP.md (its line 204 and its line 619) each continue a file named on the previous line, while the anchor that follows on their own line names metal_backend.rs — so a forward window would fail both, naming the wrong file, and the backward one has nothing to attach to.
  • The window is the line, not the paragraph or the table cell. Measured on f754ee3 with the issue's criterion, the docs carry 58 bare continuations on anchor-carrying lines: 38 resolve, 11 were out of range (the visible path anchor had been re-pointed by the #261/#263 split and the continuation left at its pre-split absolute number — fixed in the #355 PR), 2 are in the frozen docs/ARCHITECTURE-EXECUTION-PLAN.md, and 7 are silent by the two rules above. The FROZEN exemption covers bare continuations exactly as it covers path anchors. #367 swept this class: the seven silent spans were re-written in the explicit form — five live citations, the two in the frozen docs/ARCHITECTURE-EXECUTION-PLAN.md left under that exemption — so the live silent count is 0. The checker's own counts moved with it: unattached 93 → 89 and checked 1045 → 1051, because a written path is an anchor the gates can see. What stays silent is not a citation: this page's own §"Prose anchors" illustrates the rule with the two doc-line numbers of the #339 rows, and the frozen docs/ARCHITECTURE-EXECUTION-PLAN.md keeps its two under the exemption above. Writing a path is the whole fix; none of them would become a citation by a wider window.
  • Prefer the explicit form when the two citations are in different files. The rule attaches the number to the nearest anchor, not to the file the sentence meant: 03-kernels-elementwise.md's (`src/cuda/methods.rs:NNN`, launch at `:MMM`) meant the launcher that the #262 split moved to src/cuda/methods/prefill_f16.rs, so the fix names that file. The explicit form is what #367 chose for the residual: widening the window to the paragraph or the table cell was rejected, because a forward window is exactly wrong in the case that motivated the rule — the two standalone 1289-1321 spans in docs/ARCHITECTURE-ROADMAP.md continue a file named on the previous line, while the anchor that follows on their own line names metal_backend.rs. #336 is the neighbouring sweep: bare ranges that carry no symbol at all, so no rule can see whether they still hold the text their sentence describes.

The convention for the content of a range is unchanged from §"Prose anchors" above: a heading or a symbol is still the stable locator, and a bare range is still the fallback.

Writing the next gate — checklist

  1. What value does it assert, and how is that value computed independently of the path under test? (rule 1)
  2. Is the function the gate calls reachable from a production entry point — or is the claim about a test-only mirror? (rule 1, the #218 instance)
  3. Which single arm can fail only because of the property under test, and is it refused by an earlier check? (rule 2)
  4. What is the mutation — which implementation line do you break, and does the transcript show this gate going red? Use MINFER_TEST_CALL_FAIL=<site>; if the counter is what certifies the run, testfail::note_checked it too. (rule 3)
  5. Is the runtime bounded by work? If not, are the rounds interleaved and the assertion on a median that prints its inputs? (rule 4)
  6. Does every number in the PR body carry its date, device and command? (rule 5 — on a macOS/Metal change that is the Mac verification line)

How the shape is enforced

The facts rules 3 and 5 ask a PR to state in prose have a mechanical floor, and so does the Mac record of §5. The PR body is rendered from .github/PULL_REQUEST_TEMPLATE.md, and scripts/check_pr_body.py refuses a body whose required headings are missing or whose three checked sections — Bar named before measuring, Mutation evidence and Mac verification — are empty or left at their template placeholder (a stated N/A — <reason> is filled). The check-pr-body job in ci.yml runs it on pull_request events only, after the checker's own --selftest cases. It reads the body through env:, so it needs no token and works on a fork PR (#175, #335). This is the shape, not the rules: the rules stay stated once, above.

The doc gates (check-docs)

Moved here from AGENTS.md, which keeps the two-line pointer.

Doc status is machine-checked. Two ledgers, each beside the checker that reads it (ADR-0023): scripts/status.toml is the source of truth for the plan's phase counters / next: sentence / baseline commits, and scripts/test-baselines.toml for the suite counts in TEST-BASELINES.md. scripts/check_status.py --check and scripts/check_baselines.py --check (CI check-docs) fail when the prose disagrees, naming the file, line and both values — edit the ledger, not a counter. The suite ledger's one live-checkable row is compared against the real cargo test log by scripts/check_baselines.py --check-live in test-linux-cpu — a job the change classifier keeps running whenever that ledger moves, while a docs-only change skips it (ADR-0025). The same job runs scripts/build_book.sh (pinned mdBook + a sha384-checked Mermaid download) and scripts/check_docs_links.py, which fails on a relative link whose target does not exist, and scripts/check_doc_line_anchors.py (#266), which fails on a path:NNN anchor whose file or line is gone — and, when the anchor names a backticked symbol, on a symbol that has left the file it points at; a bare :NNN continuation attaches to the nearest preceding path:NNN on its own doc line and is resolved and range-checked against that file the same way, while one that follows no anchor on its line stays silent (#355); the historical records whose anchors cite a pre-split revision are frozen inside it, one reason apiece, and an exemption that no longer covers a failure fails too. Resolving is not pointing, so the same job runs scripts/check_anchor_drift.py (#344) on pull_request: it diffs HEAD against the PR's merge base (origin/<base_ref>...HEAD, the second reason the checkout is fetch-depth: 0), maps every anchored file's old→new lines, and fails an anchor the range moved without re-pointing — naming it doc:line → target:old (now new), and for a bare continuation doc:line (bare continuation) → target:old (now new). A cited line the range deleted has no image in the map, so it is reported for a human instead of guessed, and the same FROZEN set exempts the same records. It named 73 anchors over #299's range, the 69 its re-point commit fixed plus 4 it missed. It also runs scripts/check_f6_fixtures.py (#205), which checks the shape of the F6 fixture manifest tests/fixtures/f6-fixtures.json, cross-checks every ~/.cache/minfer/f6-src/… fixture the source tree names against it (and every .gguf path a recorded producer command names), pins each command's shape and classifies it with --regenerate --dry-run (#345 re-runs one and re-records what it produced), and whose --selftest proves a tampered copy is refused by name and digest; the byte half runs where the cache exists, in the F6 gates themselves. Open doc debt: #62 (docs/USAGE.md after the elastic partition). Test health was #82, closed. A finding from a gate run belongs in the records too — a red baseline makes every later gate run ambiguous.

Honest scope

This document is a contract, not a linter. Nothing in CI parses this page: rules 1, 2 and 4 are review discipline and rule 3's truth is unverifiable by construction. What the repository does enforce is the concrete part — the seam is off by default under a CI-covered test, check_docs_links.py keeps this page reachable, the per-ticket records keep the instances auditable, and the check-pr-body job requires the checked facts — the two gate sections and the Mac record — to be present in the PR body, never to be true (§"How the shape is enforced"). Treat the rules as the questions a reviewer must be able to answer from the PR, not as property tests.

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0010 — The identity gate: bitwise by default, a named tolerance class otherwise
  • ADR-0023 — The machine ledgers live beside their checkers
  • ADR-0024 — A machine-read record is not book content
  • ADR-0025 — A docs-only change runs the docs gate, not the compilers

Test baselines — the recorded suite measurements

Why this page exists. AGENTS.md is the always-loaded agent index and has to stay small, but these records are long and every device run makes them longer. They live here; AGENTS.md links to them. How to run each command is in AGENTS.md ("Build & Run").

Machine-checked. scripts/test-baselines.toml is the source of truth for every number on this page, and scripts/check_baselines.py --check (CI job check-docs) fails when this prose disagrees with it — naming the file, the line and both values. Edit that file, not a counter. Each record carries its date, device and command. Only the x86_64 (CI runner) row is compared live, by the test-linux-cpu job, against the cargo log it just collected (scripts/check_baselines.py --check-live); every other row is a recorded measurement, refreshed only by a real run, because CI has no GPU and no aarch64 runner.

The box label is absolute, never relative. dgxspark (aarch64, GB10 sm_121) is the maintainer's DGX Spark — one physical machine, which is why both its CPU and CUDA rows carry the same label — and x86_64 (CI runner) is GitHub's runner. Never write this box or local in a record: the agent reading it may be on another machine. The convention itself is stated once in gate contract rule 5. A record is written the moment it is measured, in the same commit as the run it describes.

Records

  • CPU unit, box dgxspark (aarch64, GB10 sm_121), cargo test --release, 2026-10-07: 490 passed / 0 failed / 40 ignored unit + 10 / 0 / 6 integration. (The +4 over the 2026-10-06 measurement are #205's three fixture-manifest tests — the manifest reader, the outside-the-cache arm and the tampered-cache refusal — plus #349's parity-verdict classifier; the +1 on top is #306's format-aware physical-shift gate; the next +1 is #362's platform-neutral every_device_gathers_the_attn_map, which moves the CPU rows together with the macOS ones; #354's two platform-neutral manifest-override tests — the resolver arm and the broken-override refusal — then the latest two, 488 → 490, measured on dgxspark 2026-10-07. #310 adds one pure test of the packed routing matrix, packed_route_covers_the_matrix, but it does not move this row: it lives under src/metal/, which #[cfg(target_os = "macos")] mod metal excludes from every non-macOS build. The aarch64 row is unchanged by #310, and a re-measurement on dgxspark 2026-10-08 confirmed 490 / 0 / 40.)
  • CPU unit, box x86_64 (CI runner), cargo test --release, 2026-10-08: 491 passed / 0 failed / 41 ignored unit + 10 / 0 / 6 integration. (This is the first direct measurement of the x86_64 suite: #56 adds three cfg(all(test, target_arch = "x86_64")) K-quant bitwise gates — avx2_q8k_dots_match_scalar_bitwise, avx512_q8k_dots_match_scalar_bitwise, dispatch_q8k_dots_match_scalar_bitwise — plus one #[ignore]d timing harness kquant_simd_dot_speedup. The aarch64 row (490) still carries two quants::neon_correctness tests that x86_64 lacks, so the old "x86_64 = aarch64 − 2" passed-count relation becomes x86_64 = aarch64 − 2 (neon) + 3 (x86_64 K-quant gates) = aarch64 + 1 (491); the ignored count is now machine-dependent: x86_64 41 = aarch64 40 + the x86_64-only harness. The live check_baselines.py --check-live comparison in test-linux-cpu is the authority for this row; #310 does not move it (its routing-matrix test lives under the macOS-gated src/metal/).)
  • CPU real-model set, box dgxspark (aarch64, GB10 sm_121), PARALLEL=0 scripts/real_model_gates.sh and the default parallel form, 2026-10-06: 40 / 0 each (the device-gated members of the set skip on a CPU build, so it counts 40 of the CUDA set's 46).
  • CUDA unit, box dgxspark (aarch64, GB10 sm_121), scripts/cuda_test.sh, 2026-10-07: 576 / 0 / 46. (The row read 526 before #162; #153 adds two device unit gates plus one pure kvformat gate and moves the two cuda::kv_dtype_tests to the tag mapping they now assert — 536/37 → 539/38; #144 adds the packed fused-epilogue and packed FA-prefill device gates — 539/38 → 541/38; #185 adds the device-path guard's test, the explicit-auto-budget test and the per-backend stream-sync counter gate — 541/38 → 544/38; #188 adds the capture probe and the concurrent two-engine gate — 544/38 → 545/39; #189 adds one pure test of the S4 A/B's sign-test statistic off the recorded loaded distribution — 545/39 → 546/39; #196 adds the two wedge-proof serve_loop stall gates — 546/39 → 548/39; #140 adds the K-quant encoder tests (7 passed / 1 ignored) — 548/39 → 555/40; #142 adds the bf16 writer + CPU-path tests (7 passed / 2 ignored) — 555/40 → 562/42; #202's three launch sites add no test, so the row is unchanged by it; #218 adds the two fresh-process prefill-GEMM smem opt-in gates (cuda::issue218_tests) and the captured-graph invariant arm — 562/42 → 565/42; #223 adds the fresh-process eager pre-warm gate (cuda::issue223_tests), which asserts the runtime guarantee before any launch — 565/42 → 566/42. #240/#241 deletes the dead legacy layer_gpu/device_entry wrapper layer and, with it, the one feature-independent guard test (the guard's only caller was dead), so 566/42 → 565/42 and the CPU rows lose that same test (481 → 480 on aarch64, 479 → 478 in CI). #138 adds one feature-independent allocator drain gate plus the device multi-input deferred-wait gate — 565/42 → 567/42, and the CPU rows 480 → 481 on aarch64 and 478 → 479 in CI. #208's CUDA half adds two non-ignored device gates (cuda_backend::tests::weights::cuda_bf16_matmul_matches_the_exact_shift_reference and ..._cuda_bf16_embed_gather_matches_the_exact_shift_reference) and one ignored device gate (f208_bf16_weights_run_on_the_cuda_device) — 567/42 → 570/45, the CPU rows 481/36 → 482/39 on aarch64 and 479/38 → 480/39 in CI, and each CUDA real-model set 42 → 45. #208's Metal half adds the two macOS-only Metal kernel-exactness gates and one ignored real-model gate (f208_bf16_weights_run_on_the_metal_device, a no-op off macOS, so it merely raises the ignored count) — 570/45 → 570/46, the CPU rows 482/39 → 482/40 on aarch64 and 480/39 → 480/40 in CI, and each real-model set one higher; #205's three fixture-manifest gates and #349's parity-verdict classifier then add four feature-independent tests (570/46 → 574/46) and #306's format-aware physical-shift gate the fifth — 574/46 → 575/46; #362's platform-neutral every_device_gathers_the_attn_map the sixth — 575/46 → 576/46, re-measured on dgxspark 2026-10-07.) The row is a recorded measurement and carries a #207 projection: docs/status.toml names the CPU row on the same box that it shares a test binary with, and scripts/check_status.py --check prints — never fails on — the projected count when that row has moved since this one was measured.)
  • CUDA real-model set (0.5B config), box dgxspark (aarch64, GB10 sm_121), FEATURES=cuda scripts/real_model_gates.sh, 2026-10-06: 46 / 0.
  • CUDA real-model set (Qwen3-0.6B config), box dgxspark (aarch64, GB10 sm_121), FEATURES=cuda scripts/real_model_gates.sh, 2026-10-06: 46 / 0.
  • compute-sanitizer --tool memcheck over the CUDA unit suite, box dgxspark (aarch64, GB10 sm_121), scripts/cuda_test.sh under the sanitizer, 2026-10-01: 0 API errors (565 / 0 / 42).
  • macOS unit, box macbook (macOS 27.0.1, Apple M4 Pro), cargo test --release --no-fail-fast, 2026-10-08: 553 passed / 0 failed / 45 ignored unit + 21 / 0 / 6 integration. (The +3 over the previous 550 are #310's new gates — the pure packed_route_covers_the_matrix (the packed routing matrix; it lives under the macOS-gated src/metal/, so this row is the only one it moves) plus the macOS-only metal_packed_decode_stage_matches_the_native_read (the mechanism-A/B differential) and metal_packed_window_stages_at_the_absolute_row (the absolute-row staging arm the zero-lo_min gates cannot see); [#310] is now enabled (READS_PACKED_KV = true, so MINFER_CACHE_TYPE=q8_0 loads and runs on Metal — mechanism A reads packed cells natively in the decode flash family, mechanism B stages the window to an f32 scratch buffer and runs the unchanged f32 prefill/window families); the earlier +5 are [#310]'s packed-Q8_0 KV gates in graph::metal_backend::tests::packed_kv — the byte-identical store, the causal decode/prefill, the one-range attn_span and the kv_map window, and the C5 FLAG_PACKED round trip — plus the +1 ignored real-model gate, so the ignored row moved 44 → 45; the +1 under the 542 baseline is #369's fast-map gate window_map_flash_matches_the_cpu_reference; the +1 under that is #359's fast-windowed-kernel gate window_flash_matches_the_cpu_reference; the +3 under that are #362's two Metal kv_map window gates plus the pure every_device_gathers_the_attn_map; the earlier +4 are #205's three fixture-manifest tests plus #349's parity-verdict classifier, and the +1 on top is #306's format-aware physical-shift gate — the +3 added from a Mac, the rest from dgxspark and platform-neutral. #354's two platform-neutral manifest-override tests are derived here (543 → 545), not re-measured on a Mac. The +2 over the #54 baseline are #329's dispatch-refusal gate and #299's weights-charged E4 gate; the +1 ignored is #315's macOS-only #[ignore]d windowed-prefill measurement harness, so the passed count it changes is not this one. The macOS suite was red for the whole Metal round; this is the baseline a later Mac run diffs against — docs/METAL-BACKEND-DESIGN.md §7.4 — and it is green because the round fixed its three production bugs. Not run in CI: build-macos compiles the test target since #303 but has no Metal device.)
  • macOS real-model set, box macbook (macOS 27.0.1, Apple M4 Pro), PARALLEL=0 scripts/real_model_gates.sh (serial), 2026-10-08: 45 / 0 on both cached models — the set is 45 because #310's macOS-only #[ignore]d metal_q8_0_kv_answers_like_f32_on_a_real_model gate joined the --ignored set (44 → 45), after #315's gate had taken it 42 → 43 (since #310 enabled Metal's packed read, the gate drives the production capability directly, no test seam), and the previously failing server::batch::tests::kv_sharing::a_store_inside_a_shared_prefix_takes_a_private_row now passes: Metal gathers the set-valued kv_map window since #362, so its shared-prefix slot reads the donor's rows in place instead of copying them (recorded in docs/METAL-BACKEND-DESIGN.md §7.4). No residual.

Metal Runtime Parameter Alignment Audit (minfer ↔ llama.cpp)

Principle: any parameter that llama.cpp decides at compile time (#define) or via configuration, minfer must NOT hardcode — it must be runtime-selected (device auto-detection) or configurable (env-var). Compile-time constants are only allowed when no runtime mechanism makes sense, and must note which llama configuration they correspond to.

Created 2026-08-06. Reference llama.cpp commit 88b47a755.

Audit Table

#Parameterminfer (hardcoded location)llama.cpp mechanism (location)StatusAction
AGEMM tile (prefill)constexpr NR0=64, NR1=32 in the 5 _mm_f32 kernels (metal.metal:768/892/1006/1120/1234); dispatch grid (nt/32, od/64) (metal.rs:290)Compile #define N_MM_BLOCK_X/Y, SZ_SIMDGROUP → NRA=64, NRB=128 (ggml-metal-impl.h:8-15); #ifdef GGML_METAL_HAS_TENSOR selects the mpp::tensor_ops::matmul2d kernel vs the legacy simdgroup kernel; runtime has_tensor (MTLGPUFamilyMetal4_GGML) picks nr0/nr1/smem 64/128/4096 vs 64/32/6144 (ggml-metal-device.cpp:748-759)✅ MOOT on M4 Pro (2026-08-06)CLOSED: llama DISABLES the tensor GEMM for pre-M5 devices (ggml-metal-device.m:713-725 — "M4 no significant difference", M2 "5% slower"; enabled only on M5/M6/A19/A20 or GGML_METAL_TENSOR_ENABLE=1). minfer's 64×32 simdgroup is what llama uses on the M4 Pro. The prefill gap is NOT the GEMM — it is attention (see "Prefill Gap" in METAL_OPTIMIZATIONS.md)
Bmul_vec tile (decode)Per-kernel const short NR0/NSG: Q5_0=4/2, Q8_0=2/4, Q6_K=2/2, Q4_K=2/2 (metal.metal:183…1707)Per-quant N_R0_*/N_SG_* header constants (ggml-metal-impl.h:24-74), picked via get_pipeline_mul_mv (ggml-metal-device.cpp:766+)✅ values match llama exactly (verified 2026-08-06)Audit whether the values should be centralized / device-dependent; no change needed now
CAttention chunk sizesplit_chunks = clamp(nkv/16, 1, 32) (metal.rs:1468-1470), overridable via MINFER_ATTN_CHUNKS; Bc=32 tile (metal.metal:2587)OP_FLASH_ATTN_EXT_NCPSG/NQPSG compile constants; nwg=32, nsg runtime-computed (ggml-metal-ops.cpp:2726-2975)⚠️ default hardcoded but env-overridableConfirm the env mechanism suffices; document the design difference (minfer adaptive chunks vs llama fixed C)
DKV cache typeMINFER_CACHE_TYPE env (default f32) (metal.rs:78-80)Compile-time (f16 default, baked into the context)✅ minfer is more flexibleNo change
EDispatch threadgroup/gridHardcoded TG (32,4)/(64,1)/(128,1) in dispatch calls (metal.rs:290-627)Computed per-op from pipeline nr0/nr1/nsg (ggml-metal-ops.cpp)⚠️ matmuls verified aligned (2026-08-06 #6); others unauditedAudit each for shape/device dependence
FElementwise/small kernelsfloat4, fixed thread countssimilar✅Low priority

Key Parameter Mapping (A — the GEMM tile)

llama compile:  ggml-metal-impl.h  SZ_SIMDGROUP=16, N_MM_NK=2, N_MM_BLOCK_X=4,
                N_MM_BLOCK_Y=2, N_MM_SIMD_GROUP_X=2, N_MM_SIMD_GROUP_Y=2
                → NRA = 16*2*2 = 64, NRB = 16*4*2 = 128   (// TODO: become function constants)
llama compile:  #ifdef GGML_METAL_HAS_TENSOR → <metal_tensor> + <MetalPerformancePrimitives/...>
                → kernel_mul_mm uses mpp::tensor_ops::matmul2d (128×64)
llama runtime:  has_tensor = [dev supportsFamily:MTLGPUFamilyMetal4_GGML]
                → nr0/nr1/smem: 64/128/4096 (M4)   or   64/32/6144 (fallback)
minfer today:   constexpr NR0=64, NR1=32 (the fallback path, hardcoded)

TODOs

  1. A: GEMM tile — CLOSED 2026-08-06: llama disables the mpp/tensor GEMM on M4 Pro (pre-M5); minfer's 64×32 simdgroup IS llama's M4 path. No port needed.
  2. B: centralize/audit mul_vec NR0/NSG (confirm vs llama N_R0_*/N_SG_*).
  3. C: confirm attention chunk default mechanism (env override suffices).
  4. E: audit dispatch grid/threadgroup params for shape/device dependence.
  5. Audit method: full grep of constexpr/dispatch constants in metal.metal/metal.rs vs llama ggml-metal-impl.h / -device.cpp / -ops.cpp.
  6. Prefill attention — FIXED 2026-08-11: the 2026-08-06 "low value" verdict was based on an OLD llama baseline (the gap was measured ~1.2-1.3x back then; with llama-Metal rebuilt it is 3.8x at pp430). The classic kernel_gqa_attn_f32 (sequential KV-loop, grid (nt,nk), barrier per 32-row tile) was measured at ~100ms/430tok (48% of prefill, ~25x llama's attention). A 3-pass parallel attention (kernel_attn_scores/kernel_softmax_attn/kernel_attn_output, one 256-thread TG per (t,h) row, all barrier-free) cut it to ~30ms: pp430 212→144ms (~32%), pp30 44→40ms. GQA via per-head hk=h/gqa (the broadcast-GEMM idea abandoned — a 2D GEMM can't produce the per-head 3D scores tensor). See METAL_OPTIMIZATIONS.md §3.4. (llama flash half8x8 with simdgroup matrices would close the residual ~2.3x but is still a ~600-line port for ~30ms of a one-time prefill — lower priority now that the attention kernel itself is fixed.) MINFER_SKIP_ATTN applies during prefill for profiling; MINFER_NO_MATMUL_ATTN=1 restores the classic kernel for A/B.
  7. Plan reconciliation (2026-08-12) — the 2026-08-06 "accept the architecture floor" verdict is REVOKED; the goal is to match llama.cpp performance. Tracking + action path: METAL_OPTIMIZATIONS.md §0 (single progress table) + §4 (the only action path: Xcode GUI per-kernel trace → flash-attention port for the decode non-matmul 4× gap → prefill GEMM execution efficiency toward ~7 TFLOPs/s). Open audit items B/C/E (#2-5 above) stay open but are secondary to §4; grid-shape probe (prefill_gemm_throughput_profile, 3.5-5.4 TFLOPs/s by nt) is the first prefill step. 2D-simdgroup GEMM + bf16 staging explicitly dropped (see METAL_OPTIMIZATIONS.md §4.3).

CUDA Tutorial for minfer contributors — from zero to optimization

This tutorial takes a minfer contributor who has never written CUDA to the point where they can read every kernel in src/cuda/kernels/*.cu, modify the CUDA backend safely, and follow the optimization records in docs/cuda_optimization_steps/ — including reproducing a historical step's measurement. It teaches by reading minfer's real code: each kernel chapter walks an actual __global__ function from this repository, line by line, with the CPU counterpart from the inference walkthrough as the mirror.

Who this is for. You can read Rust (minfer is pure Rust), you have read walkthrough chapters 09– 13 — especially 10 · CPU matmul kernels and 11 · attention & KV — and you have never written a CUDA kernel. Every CUDA concept is defined where it first appears.

The machine. Everything in this tutorial was written and verified on the target platform itself: an NVIDIA GB10 (DGX Spark, aarch64) with CUDA 13.0 (/usr/local/cuda/bin/nvcc). Toy programs were compiled and run for real; their outputs are quoted as observed. To build minfer with the CUDA backend: cargo build --release --features cuda (details and pitfalls: docs/BUILD.md).

What you will be able to do afterwards

  • Predict what a kernel launch does — threads, blocks, memory traffic — before running it.
  • Read any kernel in src/cuda/kernels/*.cu and any dispatch site in src/graph/cuda_backend.rs.
  • Explain why minfer's decode path is GEMV + CUDA Graph replay while prefill is int8 MMQ GEMM.
  • Profile with ncu/nsys, read the numbers, and prove an optimization instead of trusting it (the campaign's gate chain).

The ladder — read in order

ChapterPartWhat it adds
01 · What kind of machine is a GPU1 — mental modelHost/device, SIMT (thread/warp/block/grid), memory hierarchy, why LLM inference fits — plus your first compiled-and-run kernel (Toy #1)
02 · The minimal CUDA you actually need2 — language surfaceKernel syntax & indexing, error checking & sync semantics, device memory & minfer's Rust/FFI wrapper, streams, the build system (Toys #2–#3)
03 · Reading minfer's kernels I3a — first real kernelsElementwise (add_f32), quantized-weight dequant (dequant_q4_0_f16), embedding gather (embed_rows_q4_0)
04 · Reading minfer's kernels II3b — the matmul ladderScalar GEMV → vectorized GEMV → tiled GEMM (gemm_f16_nt_kernel_t) → int8 MMQ pointer
05 · Reading minfer's kernels III3c — attention + hostFlash-attention-style prefill (fa_prefill_f16kv), fused decode tail (attn_bias_rope_store_f32), KV in device memory, the Rust Backend layer, CUDA Graph capture/replay
06 · Optimization methods4 — from reading to changingProfiling with ncu/nsys, the technique catalog (each anchored to the step that used it), the verification gate chain, three exercises
07 · Where to go next5 — the mapReading order for the reference docs, nvcc/ncu cheat sheet, pitfall list, toy index
flowchart LR
    A["01 mental model<br/>(Toy #1)"] --> B["02 minimal CUDA<br/>(Toys #2–#3)"]
    B --> C["03 kernels I<br/>elementwise · dequant · embed"]
    C --> D["04 kernels II<br/>GEMV → tiled GEMM → MMQ"]
    D --> E["05 kernels III<br/>attention · host side · graphs"]
    E --> F["06 optimization<br/>profile · change · verify"]
    F --> G["07 the map<br/>reference docs · appendix"]
DocumentRoleThis tutorial's policy
docs/CUDA-TECH-PRIMER.mdtechnique reference — every term and technique, explained at depthone-line mention + link; never re-explained here
docs/CUDA-BACKEND-DESIGN.mdbackend design & implementation phaseslinked for design history; this tutorial teaches the code as it stands
docs/CUDA_OPTIMIZATION.md + docs/cuda_optimization_steps/the optimization campaign: live status + 70+ step recordschapter 06 anchors each technique to the step that used it
walkthrough 14 · Metal / 15 · CUDAbackend chapters of the inference serieschapter 15 is the "what" of the backend — this series is the "how CUDA works" underneath it
docs/GPU_SAFETY.mdhard safety rules for GPU codecited where the rules come from; read it before touching backends
docs/GLOSSARY.mdcampaign glossaryfallback for any term this series defines too briefly

Conventions

  • Prose in English; code, env vars and paths as-is. Writing contract: STYLE.md (same voice rules as the inference walkthrough).
  • Line numbers were verified against the tree at the time each chapter was written; the function/kernel name is the stable address, the line number a convenience. Toy outputs are real runs on the GB10 with CUDA 13.0.
  • Toys are standalone .cu files embedded in the chapters; compile commands are given verbatim and none of them are part of minfer's build.

Start reading: 01 · What kind of machine is a GPU.

01 · What kind of machine is a GPU

Part: Part 1 — the mental model (chapters 01–02: concepts first, then the minimal CUDA toolkit). Prereq: docs/inference_e2e_walkthrough/09–13 — you know minfer's CPU pipeline: prefill (doc 09), the quantized matmul kernels and the thread pool (doc 10), attention and vec ops (doc 11), the decode loop (doc 13). Zero CUDA knowledge is assumed. Code: none — concepts + one toy. minfer's kernel source reading starts in ch. 03; the platform summary below defers to docs/CUDA-TECH-PRIMER.md.

1. Background — where this sits

Through walkthrough docs 09–13 you followed one full inference run on the CPU. Prefill pushed the whole prompt through the compute graph in a single forward pass; decode then looped one token at a time; every matmul was quantized weight rows multiplied against Q8_0 (8-bit, block-scaled) activations; and a small persistent thread pool split each matmul's output rows across the CPU's cores. Doc 15 then showed that the same graph can execute on a CUDA backend — but it treated the GPU as a black box that receives tensors and returns logits. This tutorial ladder opens that box.

First, why the box exists at all. Doc 10's decode arithmetic is the whole motivation: a 7B Q4_K_M model keeps ~4.4 GB of quantized weights, and a decode step must stream all of them from RAM for every single token, because with one token row there is nothing to amortize the weight reads over. On a CPU-class memory system moving ~60–100 GB/s, that alone caps decode at roughly 15–25 tokens/s — no amount of clever SIMD changes the ceiling, because the bottleneck is bytes moved, not math done (doc 10 §2). A GPU's answer is not smarter arithmetic. Its answer is orders of magnitude more parallel work in flight, so the memory system is kept saturated instead of idling between cache misses — the same argument doc 14 makes for Metal on Apple silicon: thousands of GPU lanes keep the memory controller saturated in a way a score of CPU cores cannot. minfer grew a CUDA backend because on the target machine that parallelism is what lets the 4.4 GB weight stream run at 75–84% of the memory system's peak rate instead of trickling through a handful of cores.

But a GPU is a genuinely different kind of machine, with its own execution model, its own memory hierarchy, and its own ways to lose performance. So the ladder is taught in steps:

ChapterWhat it adds
01 (this one)The mental model: host vs device, SIMT execution, the memory hierarchy, and one runnable toy (vector add).
02The minimal CUDA you actually need: kernel syntax in full, the runtime API surface, and the error-handling discipline.
03–05Reading minfer's real kernels in src/cuda/kernels/*.cu, easiest first: elementwise → decode GEMV/matvec → prefill GEMM and attention.
06Optimization: the techniques the CUDA campaign actually measured, each linked to the step record that used it.

Three runnable toys carry the hands-on thread (vector add here; an elementwise/SiLU toy and a streams toy later; they are indexed in 07-where-next.md's appendix). After this chapter you should be able to do four things: define every term in the header of any CUDA kernel launch; predict what a kernel launch does before you run it; do the bytes-vs-bandwidth arithmetic that decides whether a kernel is memory-bound or compute-bound; and say why minfer's decode path is shaped the way it is. You will also have compiled and run a CUDA program on this GB10.

2. Principle — the concepts

2.1 Host and device — two machines, two address spaces

A CUDA program runs on two computers at once, and CUDA has words for both. The host is the CPU side — for minfer, the Rust process. The device is the GPU side — the .cu code compiled by nvcc. They are separate machines with separate address spaces: a pointer on the host means nothing on the device, and a pointer the device returns means nothing to the host. Every byte that crosses the boundary crosses it explicitly, through a copy call (cudaMemcpy) or through memory that is deliberately shared. This is the first mental shift from CPU programming, where a &[f32] is just a &[f32] no matter which function reads it.

The consequence that shapes everything else in minfer: the data transfer is the cost. Copying bytes between host and device is no faster than any other memory traffic — often slower — so a design that ships weights to the GPU per forward pass pays a transfer tax on every token. minfer's design (TECH-PRIMER §1, §6.1) is the direct answer: upload every weight tensor once at load time, leave it resident on the device for the process lifetime, and never copy it again. After that, decode moves essentially zero bytes across the host↔device boundary — positions and token ids ride over as a few dozen bytes per step, written into device buffers. The Phase-7 step record measured what the opposite extreme costs: a transfer-dominated prefill shape ran at 30.7 tok/s where the weight-resident design later reached 1204 tok/s on the same model (TECH-PRIMER §6.2, step docs 01–02).

The machine this runs on — and that this tutorial's toys run on — is an NVIDIA GB10 (DGX Spark). In a few lines, from docs/CUDA-TECH-PRIMER.md §1 (read it for the depth): it is a Grace-Blackwell "superchip" where a 20-core Arm CPU — the same machine that runs minfer's NEON+SDOT CPU backend — and a Blackwell-class GPU share one coherent memory pool (~128 GB, LPDDR5x). The GPU reports compute capability 12.1 (sm_121 in nvcc's naming; docs/GLOSSARY.md), the driver here is 580.173.02. One warning the record has already paid for: "unified memory" does not mean the program can treat CPU and GPU pointers as one — minfer deliberately uses plain cudaMalloc device pools and weight-resident uploads, never cudaMallocManaged. Unified physics, still two address spaces.

For the concrete machine, these are the numbers I queried on the GB10 with the 11-line probe in §5 (you should run it too; they are the launch-planning constants for every later chapter):

Property (queried via cudaGetDeviceProperties)GB10 value
Name / compute capabilityNVIDIA GB10 / 12.1 (sm_121)
SMs (streaming multiprocessors — §2.2)48
Warp size32 threads
Max resident threads per SM / per block1536 / 1024
Shared memory per block48 KiB
L2 cache24 MiB
Reported global memory121.6 GiB

2.2 SIMT — thread, warp, block, grid, SM

Now the execution model. A kernel is an ordinary-looking function (marked __global__ in the code) with one twist: it does not run once — it runs once per thread, on many thousands of threads at once. You write "the work of one element", and the hardware clones it across the whole dataset. A thread is one of those clones: it has its own program counter, its own registers, and a built-in idea of which element it is responsible for. The launch call — kernel<<<grid, block>>>(args) — does not call the function; it schedules the function to be executed by a specific number of threads, and returns immediately.

The launch parameters name two levels of that thread army. A block is a group of threads (up to 1024) that is scheduled onto exactly one SM — a streaming multiprocessor, the GPU's physical compute unit, with its own registers, schedulers, and cache. A block never splits across SMs and never migrates; threads inside a block can cooperate (share a scratchpad, wait on barriers — chapters 03–04). The grid is the collection of all blocks one launch creates. So the full picture of vec_add<<<GRID, BLOCK>>> is: GRID blocks × BLOCK threads, every block housed by some SM, all executing the same few lines of code on different data.

Inside the SM, threads do not actually proceed one at a time. They are organized into warps of exactly 32 threads, and the warp is the unit the hardware actually steps: all 32 threads of a warp fetch and execute the same instruction at the same time, each on its own data. That design is called SIMT — single instruction, multiple threads — and it is why GPUs are cheap per thread: one instruction decoder and scheduler serves 32 lanes. It is also why if statements have a hidden price: if some threads of a warp take the if branch and others the else, the warp executes both branches serially, masking off half its lanes each time (divergent branch — a term chapter 03 will make concrete).

A classroom analogy, carried as far as it holds: the kernel is a worksheet with one exercise ("add a[i] and b[i], write c[i]"). A thread is one student doing one exercise. A warp is a row of 32 desks that must move together — the teacher (the SM) reads each instruction once and the whole row executes it in unison. A block is a classroom: up to 1024 students who share a blackboard (shared memory) and can coordinate among themselves; the whole class is assigned to one physical room (the SM) for its entire stay. The grid is the school day: every classroom launched at once. Where the analogy lies: classrooms are virtual. One SM hosts several blocks simultaneously (up to its resource limits), and when there are more blocks than fit, they run in batches — see occupancy below.

Two quantities you will meet in every minfer performance discussion fall out of this picture. Occupancy is how much of an SM's hosting capacity — thread slots, registers, shared memory — a kernel actually uses; an SM at full occupancy has enough resident warps to switch to a different warp the moment the current one stalls (e.g. waiting on a memory load), which is how GPUs hide latency. Waves is the grid-level version: the GPU runs blocks in batches of "(blocks resident per SM) × (SM count)"; total blocks divided by that is the wave count, and a kernel whose grid is 1.5 waves wastes half of the second, nearly-empty wave (wave quantization — a recurring villain in the campaign record: a fused-kernel probe at 1.5 waves measured +28.2% slower, TECH-PRIMER §4). The GB10 has 48 SMs accepting up to 1536 threads each = 73,728 resident threads; with minfer's usual 256-thread blocks that is 6 blocks per SM, 288 blocks resident at once. The toy in §3 launches 262,144 blocks — about 910 waves. minfer's decode GEMMs at batch size 1 launch fewer blocks than one wave holds — 0.14 waves in the record — which is exactly why they amortize so well when batched (TECH-PRIMER §4).

Contrast with the CPU pool you already know (walkthrough doc 10). minfer's CPU backend parallelizes a matmul with a hand-built persistent pool (get_pool, src/kernel/pool.rs:258): workers are OS threads spawned once because spawning measured ~170 µs (src/kernel/pool.rs:7, the file's measurement comment) — against a per-token budget of a few milliseconds; they spin on an atomic generation counter to wake in microseconds; the work unit is a chunk of output rows (pool::chunk, src/kernel/pool.rs:224); and correctness rests on "each row belongs to exactly one worker", so the result is bit-identical at any thread count. Around 20 workers exist on dgxspark — one per core, because on the CPU, parallelism is expensive and scarce.

The GPU inverts nearly every line of that design:

CPU pool (doc 10)GPU grid (this chapter)
Unit of parallelismOS threadhardware thread (in warps)
How many~cores (20 on GB10)thousands resident; millions launched
Cost of one morereal (stack, wake latency)~zero until SM slots fill
Work unita chunk of output rowsone output element (or a small tile, ch. 04)
Who schedulesyou (atomics, chunking, spin)the hardware (block scheduler)
Your jobsplit rows fairly, wake workersgive every thread a unique index; arrange memory (§2.3); launch enough blocks
Parallelism cost floorthread spawn ~170 µskernel launch ~2–7 µs (TECH-PRIMER §8)

So when "one core becomes 48 SMs × 1536 thread slots", what changes conceptually is not "more threads of the same kind". It is that the unit of scheduling moves out of your hands: you stop managing workers and start describing an army of independent index-holders, then spend your effort on the two things the hardware will not do for you — mapping indices to memory locations efficiently (coalescing, §2.3) and launching enough blocks to keep all SMs fed (occupancy and waves).

2.3 The memory hierarchy — where the bytes actually are

The second mental shift is about memory. The GPU's memory system is a stack of levels, each smaller and faster than the one below, and a kernel's performance is largely which level its data comes from. From the thread's point of view:

  • Registers — private per thread, where your local variables live. Fastest by far. (Per-thread, ~0 cycles — TECH-PRIMER §5.1.)
  • Shared memory — a small scratchpad (48 KiB per block here) you manage explicitly: a block stages data from the big pool into it so all its threads can reuse the bytes at low cost. You opt in with __shared__ declarations; chapter 04 exercises it.
  • L1 / L2 caches — hardware-managed, shared across blocks on an SM (L1) and across the chip (L2; 24 MiB here).
  • Device memory — the big pool (121.6 GiB reported), called global memory in CUDA because it is visible to every thread of every block via plain pointers. This is where cudaMalloc puts tensors, where weights and KV cache live.

The magnitude table, using the latency classes TECH-PRIMER §5.1 records for this platform plus the two bandwidth figures this chapter can defend:

LevelScopeWho manages itTypical latencyBandwidth character
Registersone threadthe compiler~0 cycleseffectively free per operand
Shared memoryone blockyou (__shared__)~30 cycles (tens)very high, but tiny capacity
L1 / L2per-SM / chiphardware~200 / ~400 cycleshigh; 24 MiB total here
Device (global) memoryall threadscudaMalloc~600+ cycles (hundreds)the ceiling that matters: ~273 GB/s spec (docs/GLOSSARY.md: GB10's unified LPDDR5x, shared CPU+GPU); ~225–229 GB/s measured by this chapter's toy (§3)

Memory bandwidth is the term for that last number: how many bytes per second the memory system can stream, regardless of how much compute sits around it. It deserves care with vocabulary, because GPUs differ wildly here. High-end datacenter GPUs stream from HBM — High Bandwidth Memory, DRAM stacks mounted beside the die — at multiple TB/s. The GB10 instead shares one LPDDR5x pool between CPU and GPU, spec'd at ~273 GB/s (GLOSSARY). That is the roofline every later chapter measures against: minfer's decode kernels sustain 75–84% of it (step doc 74), the campaign's best-pure-stream kernel hit 89% (r55, step doc 58), and this chapter's trivial toy will land at ~83%. Bandwidth is the decode wall, so every GPU decision in minfer is ultimately a decision about bytes.

One paragraph preview of the most important byte-saving idea. The 32 threads of a warp execute the same load instruction; if their 32 addresses are adjacent (say 32 consecutive floats), the hardware folds them into the minimum number of wide memory transactions — 128-byte sectors. That is coalescing. If the addresses are scattered instead, the same instruction can generate many separate transactions and burn multiples of the bandwidth for the same data. TECH-PRIMER §5.2 records the real stakes: repacking one weight layout so its bytes sat contiguously (the dpl repack) cut content traffic by 17.1% on one kernel — same math, same values, purely a memory-map change. Coalescing gets its full treatment with exercises in chapters 03–04; for now, remember the rule of thumb: adjacent threads should read adjacent memory.

2.4 Why LLM inference loves (and hates) GPUs

Put the two previous sections together and you get the performance model that organizes everything else. A kernel's time is bounded by the larger of two costs: the bytes it must move through the memory system (bytes ÷ bandwidth) or the arithmetic it must do (operations ÷ peak throughput). The ratio of operations to bytes — arithmetic intensity — decides which bound you sit against: low intensity means memory-bound (the memory is the bottleneck), high intensity means compute-bound. LLM inference has one of each phase, which is why minfer's two paths look so different.

Prefill (walkthrough doc 09) runs the whole prompt at once: nt token rows go through every matmul together, so each matmul is a true GEMM (General Matrix-Multiply — matrix × matrix). Every weight byte is reused nt times inside the arithmetic, so the more tokens you batch, the more compute each byte buys: intensity grows with nt, and large-shape prefill is throughput-bound — it wants maximum math, which is what tensor cores (the SM's matrix-multiply hardware) provide. minfer's prefill path is an int8 MMQ GEMM — matrix-matrix quantized: weights stay quantized, activations are quantized on the fly to 8-bit so the integer multiply hardware multiplies the throughput (walkthrough doc 09 for the pipeline, TECH-PRIMER §6.2 for the kernel families) — plus fa_prefill_f16kv, a FlashAttention-style tiled attention kernel. The record's headline of what "the GPU is good at this" means: swapping the transfer-dominated prefill for weight-resident tensor-core GEMMs took 7B prefill from 30.7 to 1204 tok/s (39×, TECH-PRIMER §6.2).

Decode (walkthrough doc 13) is the opposite regime. One token per forward: ~250 matmuls per token (doc 10 §3), and every one of them degenerates to a GEMV (matrix-×-vector: one activation row against the whole weight matrix). Each weight byte now travels from device memory to be used in exactly one multiply-add — the measured class is ~0.03 FLOP per byte (step doc 74) — so decode is always memory-bound: its time is essentially "weight bytes ÷ bandwidth", full stop. That is why minfer's decode path is a GEMV-style matvec family (q*_q8_mmvq, dp4a integer dots per weight row) whose only goal is to stream bytes at the highest rate the memory system allows (TECH-PRIMER §6.2), and why the campaign's decode attribution speaks in GB/s, not tok/s (step doc 67). It is also why decode needs a second trick: the step is a chain of ~13 small kernels, each costing ~2–7 µs of CPU-side launch overhead — negligible for one kernel, but the chain re-runs for every token, and hundreds of launches per step (the prefill chain runs ~380) make launch overhead a first-class cost. CUDA Graphs (TECH-PRIMER §8) record the whole launch sequence once (capture) and re-issue it with one call (replay), collapsing that overhead; MINFER_NO_CUDA_GRAPH=1 reverts to per-kernel launches and is the standard A/B control in the step records.

And note what both paths never do after load: talk to the host. Weights resident (§2.1), activations and KV living in device pools — the transfer cost of inference was paid once, at model load. The rest of this tutorial lives inside the device.

3. Toy #1 — your first kernel: vector add

Everything above, in 58 lines. c[i] = a[i] + b[i] for 67 million elements: the "hello world" of CUDA, and — not coincidentally — the same shape as minfer's elementwise kernels (add_f32, add_bias_f32: TECH-PRIMER §6.4). The N is chosen large on purpose: 3 × 256 MiB = 768 MiB of traffic, far beyond the 24 MiB L2, so the timing measures real device-memory streaming and not cache or launch effects.

// toy1_vec_add.cu — Toy #1 (CUDA tutorial ch. 01): c[i] = a[i] + b[i], on the GPU.
// Verified with CUDA 13.0 (/usr/local/cuda/bin/nvcc) on NVIDIA GB10, -arch=sm_121.
#include <cstdio>
#include <cstdlib>
#include <cmath>

// A "kernel": one function body, executed by MANY GPU threads at once.
__global__ void vec_add(const float* a, const float* b, float* c, int n) {
    int i = blockIdx.x * blockDim.x + threadIdx.x;  // this thread's global index
    if (i < n) c[i] = a[i] + b[i];                  // guard: the last block may overrun n
}

int main() {
    const int    N     = 1 << 26;                    // 67,108,864 elements (256 MiB per array)
    const size_t bytes = N * sizeof(float);          // 3 arrays => ~768 MiB of DRAM traffic below
    const int    BLOCK = 256;                        // threads per block (minfer's usual choice)
    const int    GRID  = (N + BLOCK - 1) / BLOCK;    // ceil-div: enough blocks to cover all N

    // host = CPU memory; device = GPU memory. Two separate address spaces.
    float *ha = (float*)malloc(bytes), *hb = (float*)malloc(bytes), *hc = (float*)malloc(bytes);
    for (int i = 0; i < N; i++) { ha[i] = i * 0.001f; hb[i] = 1.0f; }

    float *da, *db, *dc;                             // device-side pointers
    cudaMalloc(&da, bytes);                          // allocate GPU memory for a
    cudaMalloc(&db, bytes);
    cudaMalloc(&dc, bytes);
    cudaMemcpy(da, ha, bytes, cudaMemcpyHostToDevice);  // copy a: CPU -> GPU
    cudaMemcpy(db, hb, bytes, cudaMemcpyHostToDevice);  // copy b: CPU -> GPU

    // Two events on the default stream bracket the kernel; they read GPU time.
    cudaEvent_t t0, t1;
    cudaEventCreate(&t0);
    cudaEventCreate(&t1);
    cudaEventRecord(t0);
    vec_add<<<GRID, BLOCK>>>(da, db, dc, N);         // the launch: GRID blocks x BLOCK threads
    cudaError_t err = cudaGetLastError();            // launch errors surface HERE, asynchronously
    if (err != cudaSuccess) { printf("launch failed: %s\n", cudaGetErrorString(err)); return 1; }
    cudaEventRecord(t1);
    cudaDeviceSynchronize();                         // block the CPU until the GPU is truly done
    float ms = 0.0f;
    cudaEventElapsedTime(&ms, t0, t1);               // measured time between the two events

    cudaMemcpy(hc, dc, bytes, cudaMemcpyDeviceToHost);  // read the result back: GPU -> CPU

    for (int i = 0; i < N; i++) {                    // verify against a plain CPU loop
        if (fabsf(hc[i] - (ha[i] + hb[i])) > 1e-5f) {
            printf("MISMATCH at %d: %f vs %f\n", i, hc[i], ha[i] + hb[i]);
            return 1;
        }
    }
    double gbs = 3.0 * bytes / (ms * 1e-3) / 1e9;    // read a + read b + write c
    printf("OK  n=%d grid=%d block=%d  kernel=%.3f ms  ~%.0f GB/s effective\n",
           N, GRID, BLOCK, ms, gbs);

    cudaFree(da); cudaFree(db); cudaFree(dc);        // free GPU memory
    free(ha); free(hb); free(hc);
    return 0;
}

Read it in five groups.

The kernel (lines 8–11). __global__ (pronounced "global") marks a function compiled for the device, launchable from the host — that is CUDA's word for "kernel". The body is the work of one thread: read a[i], read b[i], write c[i]. There is no loop over N — the loop is replaced by the launch. The interesting line is the index formula i = blockIdx.x * blockDim.x + threadIdx.x. Every thread carries three built-in variables identifying it: threadIdx.x (my position within my block), blockIdx.x (my block's position in the grid), and blockDim.x (threads per block, the same everywhere). The product-plus-sum turns the two-level (block, thread) address into one global element index. Note the types: these are 32-bit integers, so i overflows past 2³¹−1 — a real bug class for large tensors that chapter 02 returns to. The if (i < n) guard exists because GRID is rounded up: with N = 100 and BLOCK = 256, one block of 256 threads is launched and threads 100–255 must do nothing. Omitting the guard writes past the buffer — on the GPU that silently corrupts memory or faults later, which is why the guard is never optional even when the current N divides evenly.

The mapping, drawn. A miniature of the same launch with n = 8 and block = 4 (grid = ⌈8/4⌉ = 2) — every element of c is claimed by exactly one thread, the GPU analogue of doc 10's "each row belongs to exactly one worker":

launch vec_add<<<grid=2, block=4>>>(a, b, c, n=8)

        block 0                           block 1
   (blockIdx.x = 0)                  (blockIdx.x = 1)
   threadIdx:  0   1   2   3         0   1   2   3
   global i:   0   1   2   3         4   5   6   7
               |   |   |   |         |   |   |   |
   c:        c[0] c[1] c[2] c[3]   c[4] c[5] c[6] c[7]

   i = blockIdx.x * blockDim.x + threadIdx.x        (blockDim.x = 4)
   block 1, thread 2  ->  i = 1*4 + 2 = 6  ->  c[6] = a[6] + b[6]

Host and device memory (lines 19–28). malloc gives host pointers ha, hb, hc; cudaMalloc gives device pointers da, db, dc — same C syntax, different world, and the two pointer kinds are not interchangeable (§2.1). cudaMemcpy(dst, src, bytes, kind) is the explicit bridge; the kind enum (cudaMemcpyHostToDevice / cudaMemcpyDeviceToHost) says which way. These two copies are the toy's real "load time": 512 MiB cross the boundary before any math happens. In minfer that cost is paid once per weight at load and never again; here we pay it per run because the program is short.

Timing and launch (lines 30–41). A stream is an ordered queue of GPU work — launches and copies land in it and execute in order; the code so far used the implicit default stream. A cudaEvent is a marker you place in a stream; when the GPU reaches it, it timestamps it. cudaEventRecord(t0) … launch … cudaEventRecord(t1) therefore measures the kernel's time on the GPU's own clock — unlike a stopwatch around the launch call, which would measure the (asynchronous!) enqueue and not the work. The launch line vec_add<<<GRID, BLOCK>>>(...) is the only non-C++ syntax in the file: it says "run this kernel with GRID blocks of BLOCK threads". And note the check right after it: cudaGetLastError(). Kernel launches are asynchronous — the line returns immediately, and an invalid launch (too many threads per block, no device, bad config) does not raise an error at the call site the way a Rust Result would; the error is parked until you ask. Not asking is how "silently wrong" GPU programs are born; minfer's whole backend checks errors at every step (docs/GPU_SAFETY.md culture). cudaDeviceSynchronize() then blocks the host until all queued device work is done — mandatory before reading results back, because the cudaMemcpy readback is stream-ordered but the host must not race ahead of the GPU. (Chapter 02 replaces the blunt device-wide sync with the precise stream-ordered tools.)

Verify, report, free (lines 43–57). The result is read back and compared against a plain CPU loop — the same "compare each path against its own reference" discipline minfer's step records use (TECH-PRIMER §9), in miniature: the GPU's answer for this kernel is bitwise identical to the CPU's (same f32 adds, same order — no reduction is involved), so a tolerance is generosity, not necessity, here. The printout does the byte accounting for §4: vector add reads a and b and writes c, so it moves 3 × N × 4 bytes of device traffic; dividing by the measured seconds gives an effective bandwidth. Finally every cudaMalloc is paired with a cudaFree — in minfer this pairing is wrapped in RAII device-buffer types (src/cuda.rs) so it cannot be forgotten.

Compile, run, and the actual observed output (CUDA 13.0, V13.0.88, /usr/local/cuda/bin/nvcc; NVIDIA GB10, driver 580.173.02):

$ /usr/local/cuda/bin/nvcc -O2 -arch=sm_121 toy1_vec_add.cu -o toy1_vec_add && ./toy1_vec_add
OK  n=67108864 grid=262144 block=256  kernel=3.587 ms  ~224 GB/s effective

Five runs measured 3.514–3.587 ms (~224–229 GB/s effective). -arch=sm_121 tells nvcc to emit machine code for this GPU's compute capability 12.1 (TECH-PRIMER §3 explains the SASS/PTX machinery minfer's build.rs automates). Look at the numbers and connect them to §2.2: 262,144 blocks × 256 threads = 67,108,864 thread instances, ~910 waves over 48 SMs, each wave alive for ~4 µs — and the whole 768 MiB streaming job is over in 3.5 ms.

4. Performance intuition — bytes over bandwidth

Now do to this toy what the campaign does to every kernel: bound it from below, measure it, and explain the gap.

Bytes. Vector add reads two float arrays and writes one: 3 × N × 4 = 3 × 67,108,864 × 4 = 805,306,368 bytes ≈ 805 MB (768 MiB). No kernel that must touch every element can beat the time it takes the memory system to move those bytes — that is the roofline idea of §2.4, in its simplest case.

Lower bound. At the documented GB10 roofline of ~273 GB/s (GLOSSARY; calibrated by the campaign's r55 audit, step doc 09), the floor is 805.3 MB ÷ 273 GB/s ≈ 2.95 ms.

Measured. The toy measured 3.51–3.59 ms, i.e. an effective 225–229 GB/s = 82–84% of the roofline. The remaining ~15% is the ordinary tax of real streaming: DRAM latency not fully hidden, L2/write-back effects, clock behavior — not a bug in our kernel. Two reference points say this is exactly the right neighborhood: minfer's decode matvec kernels sustain 75–84% of this same peak (step doc 74), and the campaign's best pure stream — fused swiglu, audited at 89% (r55, step doc 58) — only ~6 points better. A one-evening toy lands in the class that the campaign needed months to tune toward, because vector add has no quantization to unpack, no gather, no reuse — it is already the pure byte stream that decode is fighting to become. (Doc 15's number for minfer's decode weight streaming, the "200–225 GB/s class", is the same magnitude you just measured.)

Arithmetic intensity, seen rather than computed. Vector add does 2 FLOPs per 12 bytes moved ≈ 0.17 FLOP/byte. Decode GEMV sits at ~0.03 (step doc 74). Both are deep on the memory-bound side — and you can see it in the numbers without knowing the GPU's peak FLOPs: the kernel's time tracked the byte count (805 MB → 3.5 ms at ~83% of bandwidth), not any arithmetic budget. Whenever a kernel's time is predictable from its bytes, it is memory-bound; that one test is most of chapter 04's toolkit.

What breaks the simple model — the small end. Re-run the same toy smaller and the model collapses honestly: at N = 2²⁰ (12 MiB of traffic), the same binary measured 0.127–0.155 ms — bytes shrank 64× but time only ~25×, so the effective bandwidth collapsed to ~80–100 GB/s; at N = 2¹² it measured 0.057–0.105 ms on an idle GPU, where fixed path costs and clocks ramping from idle dominate and the "bandwidth" printed is meaningless (~1 GB/s). Below some size, time stops tracking bytes: launch latency, memory latency (not bandwidth — too few blocks in flight to hide it), and measurement hygiene take over. This is not a corner case — it is minfer's decode regime (grids at 0.14 waves; the D1 attribution's verdict of memory-latency-bound with 76.5% of stalls on long scoreboard waits, TECH-PRIMER §6.3), and it is why the campaign insists on same-window interleaved A/B medians (TECH-PRIMER §1): an idle integrated GPU measures nothing the way a busy one does.

Checklist of ways this kernel could be slow or wrong (each will get its chapter): non-adjacent per-thread addresses → no coalescing, multiple transactions per warp (§2.3); a grid smaller than one wave → SMs idle; the int i index overflowing past 2³¹ elements; reading c back before the synchronize → stale bytes; and skipping the cudaGetLastError check → an invalid launch fails silently and the CPU verification "passes" against uninitialized memory — the failure mode chapter 02's error-handling discipline exists to prevent.

5. Try it

Save the §3 block as toy1_vec_add.cu in any scratch directory (the filename the toy index in 07-where-next.md uses), then:

$ export PATH=/usr/local/cuda/bin:$PATH     # nvcc is not on every shell's PATH
$ nvcc -O2 -arch=sm_121 toy1_vec_add.cu -o toy1_vec_add && ./toy1_vec_add
OK  n=67108864 grid=262144 block=256  kernel=3.587 ms  ~224 GB/s effective

(Without the export, use the full path: /usr/local/cuda/bin/nvcc …. Verified with CUDA 13.0, V13.0.88, on the GB10, driver 580.173.02.)

Experiments, each one a one-line edit away:

  • N = 1 << 20 — watch effective bandwidth collapse (~80–100 GB/s): the fixed-cost/latency regime of §4, decode's home turf.
  • BLOCK = 32 (one warp per block) vs BLOCK = 1024 — the time should barely move for this trivial kernel; note for later that minfer chose 256 (__launch_bounds__(256), TECH-PRIMER §4.1) for reasons that matter in real kernels (registers, occupancy tuning), not for vector add.
  • N = 1000003 (odd) with the if (i < n) guard deleted — ~1,500 threads now write ~768 bytes past the end of hc: heap corruption that may not fault, and may not even trip the CPU check. Never trust "it divided evenly last time"; the guard is free insurance.
  • See the machine, not just the timing: nsys profile ./toy1_vec_add opens the Nsight Systems timeline; you should see the two H2D copies, the kernel (~3.5 ms), and the D2H copy as separate stream-ordered items — §2.1 and the event story, drawn.

And the device-properties probe used in §2.1 — the 11 lines that print this chapter's platform table (48 SMs, warp 32, 1536 threads/SM, 48 KiB shared, 24 MiB L2):

// devprobe.cu — query the device constants every later chapter assumes.
#include <cstdio>
int main() {
    cudaDeviceProp p;
    cudaGetDeviceProperties(&p, 0);
    printf("name=%s cc=%d.%d SMs=%d warpSize=%d maxThreadsPerSM=%d maxThreadsPerBlock=%d\n",
           p.name, p.major, p.minor, p.multiProcessorCount, p.warpSize,
           p.maxThreadsPerMultiProcessor, p.maxThreadsPerBlock);
    printf("sharedMemPerBlock=%zu KiB L2=%d MiB globalMem=%.1f GiB\n",
           p.sharedMemPerBlock / 1024, p.l2CacheSize >> 20, p.totalGlobalMem / 1073741824.0);
    return 0;
}
$ nvcc -O2 -arch=sm_121 devprobe.cu -o devprobe && ./devprobe
name=NVIDIA GB10 cc=12.1 SMs=48 warpSize=32 maxThreadsPerSM=1536 maxThreadsPerBlock=1024
sharedMemPerBlock=48 KiB L2=24 MiB globalMem=121.6 GiB

6. Cross-references

  • 02 · The minimal CUDA you actually need (next): kernel syntax in full, the runtime API surface, and the error-handling discipline this chapter only gestured at with cudaGetLastError.
  • docs/CUDA-TECH-PRIMER.md §1 — the platform in depth (GB10, unified memory, weight-resident design); §4 — the programming-model vocabulary used throughout (grid/block/warp, occupancy, waves, __launch_bounds__); §5 — the memory hierarchy and coalescing at reference depth.
  • docs/inference_e2e_walkthrough/10 — the CPU counterpart of this chapter's execution model: the persistent thread pool, row-chunked work, and the byte arithmetic that makes decode bandwidth-bound on the CPU too.
  • docs/inference_e2e_walkthrough/13 — the decode loop whose kernel chain (GEMV + CUDA Graph replay) chapters 03–05 read line by line.
  • docs/GLOSSARY.md — the campaign glossary, for any term this chapter defines only in passing (sm_121, LPDDR5x, roofline).

← Index · 02 · The minimal CUDA you actually need →

02 · The minimal CUDA you actually need

Part: Part 2 — the minimal language/API surface. Prereq: 01 · What kind of machine is a GPU — you know what a thread, warp (a group of 32 threads that execute in lockstep), block and grid are, and what a kernel launch does conceptually. Code: src/cuda.rs, src/cuda/kernels/*.cu, build.rs — every file:line was verified against the current tree at writing time (the function name is the stable address, the line number a convenience).

1. Background — where this sits

Chapter 01 gave you the execution model in the abstract: a kernel is a function that runs once per thread, threads come in blocks, blocks come in a grid, and the GPU schedules blocks onto its SMs (Streaming Multiprocessors — the GPU's processor units, each running many warps concurrently). This chapter turns that model into the concrete API surface minfer actually uses.

The surface is small. CUDA (NVIDIA's C/C++ extension plus its runtime library) exposes hundreds of API calls, but minfer's device layer — src/cuda.rs (1,014 lines) plus the src/cuda/methods/*.rs families #262 split out — is built on roughly twenty of them, one language feature (__global__ functions) and one launch syntax (<<<...>>>). The five things you cannot read minfer's CUDA code without: kernel syntax and indexing (§2.1); error checking and sync semantics (§2.2); device memory and the Rust FFI layer (§2.3); streams and asynchronous execution (§2.4); and the build pipeline (§2.5).

Scope note: the deep technique material — coalescing (when consecutive threads access consecutive addresses so the hardware merges their loads into one wide transaction), occupancy, tiling, the MMQ (Matrix-Multiply Quantization) prefill path — is deliberately deferred to docs/CUDA-TECH-PRIMER.md and Part 4; this chapter teaches just enough of the language and runtime for those to read like prose.

Throughout we anchor on one real model: Qwen2.5-0.5B, hidden size 896, FFN (Feed-Forward Network) intermediate width 4864. The walkthrough prints these shapes from the real graph (docs/inference_e2e_walkthrough/05-graph-builder-ir.md:184-188 (§2 "The main event: one forward pass, node by node") — ffn_gate [896 -> 4864], ffn_down [4864 -> 896]) and from the GGUF (docs/inference_e2e_walkthrough/02-gguf-load.md:232 (§2 "Quantized sizing: 32 values in 18 bytes") — token_embd.weight [896, 151936]). Both numbers are small enough to do the arithmetic in your head, and they are the exact shapes the elementwise kernels in src/cuda/kernels/*.cu process on every decode step.

2. Principle — the concepts

2.1 Kernel syntax and thread indexing

The three qualifiers. CUDA adds three function qualifiers to C++; they answer one question — who can call this function, from where?

  • __global__ — a kernel: a function callable from the host (host = the CPU and its memory, as opposed to the device = the GPU and its memory) that executes on the device, once per thread. Always returns void.
  • __device__ — callable only from device code. An ordinary inline function on the GPU: it creates no threads, it runs inside the calling thread.
  • __host__ — an ordinary CPU function (the default). Written explicitly only for functions that compile in both worlds — __host__ __device__.

src/cuda/kernels/*.cu uses exactly this split: __global__ for the 89 kernels (the count of the pre-#263 single TU; 102 today) (the count from grep -c '__global__' src/cuda_kernels.cu at that revision), plain __device__ helpers for shared math, and ordinary C++ for the host-side launcher functions below.

The launch syntax and the index formula. A kernel launch is a function call preceded by a <<<...>>> configuration clause, and inside the kernel the single most important line in CUDA programming tells each thread which element it owns:

silu<<<grid, block, shared_mem_bytes, stream>>>(dx, dy, n);
//    └──┬──┘ └─┬─┘ └─────┬─────┘ └──┬──┘
//     blocks   threads   optional   optional
//     (dim3)   per block shared mem  stream

int i = blockIdx.x * blockDim.x + threadIdx.x;
//       └────┬───┘   └───┬───┘   └────┬────┘
//       which block  threads per   my lane
//       (grid coord) block        (inside the block)

grid and block are dim3 values — three-dimensional sizes (x, y, z); a plain integer fills x and y/z default to 1, so block = 256 means "256 threads in a 1-D line". With blockDim.x = 256, block 0 owns i ∈ [0, 256), block 1 owns [256, 512), and so on. Consecutive threads get consecutive i — the property that makes memory access coalesced, and why 1-D elementwise kernels are the fastest thing a GPU does. The last two launch arguments default to 0 and the default stream (§2.4).

Toy #2 — a SiLU kernel, with a CPU reference. SiLU (Sigmoid Linear Unit, the activation in every Qwen FFN: silu(x) = x / (1 + exp(-x))) is minfer's most-launched elementwise op, so it is the right first kernel. (Toy #1, vector add, lives in chapter 01; the toy index is appendix C of 07.) The error checks below are not decoration — §2.2 explains the two different places an error can surface:

// toy2_silu.cu — a SiLU elementwise kernel + CPU reference comparison.
// Toy #2 of the minfer CUDA tutorial (chapter 02). Verified with CUDA 13.0
// on GB10 (aarch64, sm_121): nvcc -arch=sm_121 -O2 toy2_silu.cu -o toy2 && ./toy2
#include <cstdio>
#include <cmath>
#include <cstdlib>
#include <cuda_runtime.h>

#define CK(x) do { cudaError_t e_ = (x); if (e_ != cudaSuccess) { \
    printf("CUDA error %s at %s:%d\n", cudaGetErrorString(e_), __FILE__, __LINE__); \
    return 1; } } while (0)

__global__ void silu(const float* x, float* y, int n) {
    int i = blockIdx.x * blockDim.x + threadIdx.x;  // global 1-D index
    if (i >= n) return;                    // guard: grid is ceil-div, so the
    y[i] = x[i] / (1.0f + expf(-x[i]));   // last block is partly idle
}

int main(int argc, char** argv) {
    int n = (argc > 1) ? atoi(argv[1]) : 4864;   // Qwen2.5-0.5B FFN width
    float *hx = new float[n], *hy = new float[n], *ref = new float[n];
    for (int i = 0; i < n; i++) hx[i] = -8.0f + 16.0f * (float)i / (n - 1);

    float *dx = nullptr, *dy = nullptr;
    CK(cudaMalloc(&dx, n * sizeof(float)));
    CK(cudaMalloc(&dy, n * sizeof(float)));
    CK(cudaMemcpy(dx, hx, n * sizeof(float), cudaMemcpyHostToDevice));

    int block = 256, grid = (n + block - 1) / block;   // ceil-div
    printf("n=%d block=%d grid=%d -> %d threads launched, %d idle\n",
           n, block, grid, grid * block, grid * block - n);

    silu<<<grid, block>>>(dx, dy, n);
    CK(cudaGetLastError());          // launch errors surface HERE (async!)
    CK(cudaDeviceSynchronize());     // run errors surface HERE
    CK(cudaMemcpy(hy, dy, n * sizeof(float), cudaMemcpyDeviceToHost));

    for (int i = 0; i < n; i++) ref[i] = hx[i] / (1.0f + expf(-hx[i]));
    float maxerr = 0.0f; int worst = 0;
    for (int i = 0; i < n; i++) {
        float e = fabsf(hy[i] - ref[i]);
        if (e > maxerr) { maxerr = e; worst = i; }
    }
    printf("x=%.6f gpu=%.6f cpu=%.6f | maxerr=%.3e %s\n",
           hx[worst], hy[worst], ref[worst], maxerr, maxerr < 1e-6f ? "PASS" : "FAIL");
    CK(cudaFree(dx)); CK(cudaFree(dy));
    delete[] hx; delete[] hy; delete[] ref;
    return 0;
}

Compile and run on the GB10 (CUDA 13.0 at /usr/local/cuda/bin/nvcc, aarch64; -arch=sm_121 targets its compute capability 12.1 — §2.5). The toy takes an optional size, so both grid regimes are observable:

$ /usr/local/cuda/bin/nvcc -arch=sm_121 -O2 toy2_silu.cu -o toy2 && ./toy2
n=4864 block=256 grid=19 -> 4864 threads launched, 0 idle
x=2.679828 gpu=2.507852 cpu=2.507852 | maxerr=4.768e-07 PASS
$ ./toy2 896
n=896 block=256 grid=4 -> 1024 threads launched, 128 idle
x=2.815642 gpu=2.656601 cpu=2.656602 | maxerr=2.384e-07 PASS

Three lessons from the real output. First, grid sizing is ceil-div, and the guard is what makes it safe: (n + block - 1) / block is the standard integer ceiling division (every launcher in minfer uses it verbatim), and it over-issues threads whenever n is not a multiple of 256 — for the hidden size 896, 896/256 = 3.5, so 4 blocks × 256 = 1024 threads start and 128 must do nothing, which is exactly what if (i >= n) return; is for. Without the guard those threads read and write past the buffer end, and out-of-bounds writes on a GPU corrupt other allocations' bytes silently or surface as an async error later (§2.2). Second, the CPU reference is the whole quality model: maxerr ≈ 5e-07 is one rounding difference on expf — the GPU's expf and glibc's expf are different implementations of the same function — and minfer's test gates use the same shape, comparing against the CPU reference and reserving bit-identical for comparisons against the same code path (§3). Third, note the ratio: the kernel is 5 lines; launch, copy back, reference loop and checks dwarf it.

The 2-D formula — minfer's real add_bias_f32. A bias add works on token-major activations [rows][d]: row t is a token, column i a hidden-dimension index, and every row adds the same b[i]. minfer's kernel add_bias_f32 (src/cuda/kernels/ops_elementwise.cu:151)) maps the 2-D shape directly onto the grid:

// `add_bias_f32` (src/cuda/kernels/ops_elementwise.cu:151)
__global__ void add_bias_f32(
    float* __restrict__ y,
    const float* __restrict__ b,
    int d
) {
    int t = blockIdx.x, i = threadIdx.x + blockIdx.y * blockDim.x;
    if (i >= d) return;
    y[t * d + i] += b[i];
}

Line by line: t = blockIdx.x is the token — the launcher makes the grid's x dimension the row count, one block per token. i = threadIdx.x + blockIdx.y * blockDim.x is the column within the row — the 1-D formula with the block index taken from the y dimension. The lesson: the x/y/z grid dimensions are free coordinates; map whichever data axis is longest onto x (the one consecutive threads walk — the coalescing-friendly one) and use y/z for the slower axes. if (i >= d) return; is the toy's guard, now on the column axis; d needs no guard because the grid's x dimension is the row count. __restrict__ is a promise that this pointer is the only way the function accesses that memory — an optimization hint that frees the compiler from proving non-aliasing (two pointers referring to the same memory, which would forbid reordering loads and stores). The bias vector itself is 896 × 4 B = 3.5 KB, re-read by every token's block — small enough to stay hot in L2 cache across blocks.

The launcher builds that 2-D grid launch_add_bias_f32 (src/cuda/kernels/ops_elementwise.cu:353): 64 threads in x per block, dim3 grid(n, (d + 63) / 64, 1) — the same ceil-div over the column axis, 896 columns → 14 blocks in y. For one decode token (n = 1) the grid is 1 × 14 blocks of 64 threads — 896 threads for 896 elements, one thread per element. Note the kernel takes no row count: the grid encodes it, which makes the launcher responsible for the launch geometry matching the kernel's expectation. That is a contract, and breaking it is exactly the bug minfer documents on the Rust wrapper add_bias_f32 (src/cuda/methods/elementwise.rs:104-107): "rows is the ROW COUNT (token count) — the kernel grid maps one block row per token, so passing the total element count writes out of bounds." Passing rows * d would launch rows*d block-rows — a grid that reads and writes far past the buffer. On CUDA this does not segfault the host; it either corrupts unrelated device memory or surfaces at the next synchronization point (§2.2). The backend repeats the warning at its only call site (src/graph/cuda_backend.rs:1510-1514, the row-count comment at the add_bias_f32 call, which passes nt, the token count) — leave this kind of comment on every launch geometry you write.

minfer's real elementwise family, side by side with the toy. The kernel inventory (grep -n '__global__' src/cuda/kernels/*.cu) lists 102 kernels today (89 at the pre-#263 revision); the elementwise ones are the plainest and are all structured like your toy (global 1-D index, guard, one memory access per thread):

Kernel (pre-#263 src/cuda_kernels.cu)LinesBodyUsed for
add_f322403-2412z[i] = x[i] + y[i]residual add (Op::Add)
mul_f322416-2425z[i] = x[i] * y[i]elementwise multiply (Op::Mul)
silu_f322429-2434y[i] = v/(1+expf(-v)) in-placeFFN activation (Op::Silu)
swiglu_f322471-2481dst = silu(gate) * upunfused SwiGLU (Op::SwiGLU)
swiglu_f32_off2441-2446in-place, up at offset offfused-FFN path (decode)
add_bias_f322391-2399y[t*d+i] += b[i], 2-Dattention/FFN output bias
f32_bits_to_i322489-2497bit-reinterpret f32 → int32graph I32-input convention

add_f32 (src/cuda/kernels/ops_elementwise.cu:163) is your toy's skeleton exactly — index formula, guard, z[tid] = x[tid] + y[tid] — with two idiom changes: the parameter is named tid (thread id) and every pointer is const ... __restrict__. silu_f32 (src/cuda/kernels/ops_elementwise.cu:189) is the toy's kernel verbatim except it is in-place — one buffer, read and written through the same pointer — legal because each element is touched by exactly one thread. In-place-ness is a graph-level decision in minfer (the alias rule: a node may alias its input only when it is the sole consumer and runs on the same backend — §3 shows the D2D (device-to-device) copy the backend stages when the allocator did not alias it). Both launch through their launchers' ceil-div grid of 256-thread blocks ( launch_add_f32 (src/cuda/kernels/ops_elementwise.cu:363), launch_silu_f32 :385-392), and every launcher in the file ends with , stream): the entry-point family is uniform so the Rust side can target any stream uniformly. Why that matters is §2.4.

2.2 Error checking and synchronization semantics

The toy's last checks before trusting its output were two different calls — CK(cudaGetLastError()) right after the launch, then CK(cudaDeviceSynchronize()) before the D2H copy — and the difference is the most confusing thing about CUDA error handling. A kernel launch is asynchronous: silu<<<...>>> does not run your kernel — it enqueues it into a queue (a stream, §2.4) and returns immediately, typically in a few microseconds. The GPU pulls work off the queue when it gets to it; the host keeps running straight past the launch. So:

  • Errors detectable at enqueue time — invalid launch configuration (0 threads per block), a kernel not compiled into the binary, too many threads per block for this GPU — come back from the launch itself (and from cudaGetLastError() right after it). Cheap, synchronous checks.
  • Errors that happen while the kernel executes — out-of-bounds write, misaligned address, illegal instruction — can only be discovered by the GPU at run time, after the host has moved on; they sit in a per-thread error slot until a later CUDA call picks them up — conventionally a synchronization point: cudaDeviceSynchronize() (wait until every stream has finished all its work) or cudaStreamSynchronize(s) (wait until stream s is drained).

That is the whole "a kernel's error surfaces later" phenomenon: the queue means the host-side call that would report the error may run milliseconds before the GPU even starts the faulty kernel. If you never synchronize, you may never see the error — the process can exit with silently corrupted results. One wrinkle: an execution error is sticky. Once the context (the per-process GPU state that all your allocations and streams live in) records a fault, every subsequent CUDA call returns the same error until the context is re-created — so a misdiagnosed "error at line 4000" is usually the first sync after the real culprit. Bisect by syncing after each launch suspicion, which is what minfer's debug sync (below) is for.

How minfer wraps this. minfer checks each launch's own error and drains at sync points; the wrapper is CudaState::sync (src/cuda/methods/events.rs:143-157):

#![allow(unused)]
fn main() {
// `CudaState::sync` (src/cuda/methods/events.rs:143-157; #145 comment elided)
pub fn sync(&self) {
    let err = unsafe { cudaGetLastError() };
    if err != 0 {
        LATCHED_API_ERRORS.fetch_add(1, Ordering::Relaxed);
        eprintln!("{}", latched_api_error_message(err));
    }
    let err = unsafe { cudaStreamSynchronize(self.stream()) };
    if err != 0 {
        eprintln!("CUDA stream sync error: {} ({err})", cuda_error_name(err));
    }
}
}

cudaGetLastError() picks up whatever an earlier call latched — not evidence about the kernel (the #145 lesson); LATCHED_API_ERRORS counts it and latched_api_error_message names the observer and the real error. cudaStreamSynchronize(self.stream()) blocks until the bound stream has drained: (a) the host waits so results are trustworthy, (b) any execution error lands in the return value, checked and printed. Note what it does not do: no panic, no abort, no Result — CudaState::sync is a drain point, and the safety contract lives one level up, in the backend, where invariant violations become Err. There also used to be a per-node debug variant, debug_sync(il, label) (gated behind MINFER_CUDA_DEBUG=1), printing the layer index and label with any launch or sync error — the bisection tool. It belonged to the legacy layer_gpu surface and was deleted with it in #240/#241; the graph path's CudaState::sync() is the drain point, and MINFER_OP_TIMING=1 the per-op timing tool.

Why no sync per launch? The graph path launches ~100+ kernels per decode step; a cudaGetLastError() after each is cheap and does happen, but a cudaStreamSynchronize after each would serialize the pipeline — so the drain happens once per split, and docs/GPU_SAFETY.md rule 4 (docs/GPU_SAFETY.md:232, the CUDA section's rule 4) keeps that latch honest: "A latched error is never blamed on the kernel that just ran." The first call reports whatever an earlier call latched — not evidence about the kernel — and the error is counted, never dropped.

The GPU safety contract. docs/GPU_SAFETY.md (CUDA section, docs/GPU_SAFETY.md:225-242 (the CUDA section) translates the project-wide hard rules to the CUDA backend; read the whole doc before touching this code. In four lines: kernel-invariant violations return Err from execute_node — never a silent CPU fallback; a guard failure aborts with the node's name (docs/GPU_SAFETY.md:229). No sync inside an active capture window — a stray cudaStreamSynchronize while CUDA records a graph corrupts the capture (the 7e② "faster but wrong" incident; §2.4 builds on this, docs/GPU_SAFETY.md:230). Device memory is not host-readable via plain memcpy on GB10 — dereferencing a device pointer from the host segfaults; all D2H (device-to-host) traffic goes through cudaMemcpy (docs/GPU_SAFETY.md:231). A latched error is reported, never blamed on the kernel that just ran; same-stream ordering is the correctness contract for async fills (docs/GPU_SAFETY.md:232, :238). Rule 1's Rust shape is on the elementwise dispatch arm: the backend checks the kernel's preconditions before launching and returns a formatted Err naming the node (src/graph/cuda_backend.rs:998, the Op::Add arm) — Op::Add rejects input-size mismatches with Err(format!("cuda: {}: add input size mismatch", node.name)) before calling add_f32). That is the boundary between the two error worlds: host-visible checks are ordinary Rust Results raised before launch; device-side faults (an OOB inside a kernel) are the sticky CUDA errors picked up at the next sync() — a well-designed CUDA codebase is explicit about which world each check belongs to.

2.3 Device memory management

The three calls. Device memory (the GPU's own DRAM — gigabytes of it, not addressable by normal CPU pointers) is managed by three runtime calls: cudaMalloc(void** ptr, size_t bytes) allocates on the device and returns a device pointer (valid as a kernel argument, not valid to dereference from the host — the rule just above); cudaFree(void* ptr) frees it; and cudaMemcpy(dst, src, count, kind) copies bytes, where kind makes one function serve all three directions: H2D (host→device upload), CUDA_MEMCPY_HOST_TO_DEVICE = 1; D2H (download, = 2); D2D (device-internal, = 3) — the kind tells the driver which address spaces the pointers live in, and minfer defines the constants by hand CUDA_MEMCPY_HOST_TO_DEVICE (src/cuda.rs:94-96)). A blocking host↔device cudaMemcpy is synchronous: it does not return until the bytes have moved. For decode-speed code that is a problem (waiting for the GPU), which is why the async variants (cudaMemcpyAsync on a stream, plus pinned memory) exist — but the synchronous calls are the right default for one-shot work like weight loading, and that is how minfer uses them.

Pinned vs pageable, in one paragraph. Normal host memory ("pageable") can be swapped or moved by the OS, so the DMA (Direct Memory Access — the copy engine that transfers bytes without CPU involvement) hardware cannot be handed a stable physical address; the driver first copies your pageable buffer into an internal pinned staging area (memory locked against paging, allocated with cudaHostAlloc), then DMAs from there. Fine for one-shot uploads; but for repeated async transfers you allocate your own pinned buffers and copy into them yourself — then cudaMemcpyAsync can be enqueued on a stream and return before the bytes move, overlapping the transfer with compute. The cost: pinned memory is a scarce OS resource (over-allocating it degrades the whole machine, not just your process). minfer keeps small, ring-shaped pinned pools only where asynchrony pays — the 8-slot × 2 MB staging ring for input fills (write_input_async, src/cuda/methods/copy.rs:29), a grow-on-demand D2H readback buffer for per-step logits (PinnedBuf, src/cuda.rs:415), and the capture staging pool (CaptureStaging, src/cuda.rs:435) — with a synchronous pageable fallback whenever pinned allocation fails. docs/GPU_SAFETY.md:238 (CUDA rule 6) states the rule that makes this safe: same-stream ordering — the async fill is safe because every consumer kernel is enqueued later on the same stream.

How the Rust side wraps the C API (FFI). minfer links the CUDA runtime directly — no cuda crate, no bindgen. It declares the C functions it needs in an extern "C" block (FFI, Foreign Function Interface — the mechanism by which Rust calls C-ABI functions; extern "C" selects the C calling convention so Rust and C agree on how arguments are passed) in src/cuda/ffi_runtime.rs:

#![allow(unused)]
fn main() {
// src/cuda/ffi_runtime.rs:1-127 (excerpt — incl. cudaMemcpyAsync,
// cudaStreamSynchronize, device queries, CUDA Graph APIs)
extern "C" {
    fn dlopen(filename: *const std::ffi::c_char, flag: std::ffi::c_int) -> *mut std::ffi::c_void;
    fn cudaSetDevice(device: i32) -> i32;
    fn cudaFree(ptr: *mut std::ffi::c_void) -> i32;
    fn cudaMalloc(ptr: *mut *mut std::ffi::c_void, size: usize) -> i32;
    fn cudaGetLastError() -> i32;
    // cudaMemcpy(dst, src, count, kind), cudaHostAlloc, cudaStreamCreate, ...
}
}

How to read such a block — the skill of reading an FFI layer:

  • Types are the contract. CUDA's cudaError_t cudaMalloc(void** devPtr, size_t size) becomes fn cudaMalloc(ptr: *mut *mut c_void, size: usize) -> i32. The C enum cudaError_t is an i32 in Rust; the double pointer is Rust's *mut *mut c_void — the callee writes the allocated address through it. c_void everywhere: neither side cares what lives behind the pointer.
  • Out-parameters instead of return values. C APIs that "return" data do it by writing through a pointer argument (cudaMalloc, cudaGetDeviceCount, cudaStreamCreate); the wrapper functions in the rest of the file perform exactly the translation into Rust-style Results.
  • Raw pointers are not Send/Sync. Rust assumes a raw pointer may alias anything, so types containing them are not thread-safe by default. minfer's CudaPtr newtype CudaPtr (src/cuda.rs:71-72)) is the deliberate, documented exception — unsafe impl Send/Sync is an assertion you must defend: sound here because CUDA device allocations are process-global resources usable from any thread, and every access is funneled through Mutexes.
  • Where unsafe lives. Every call into these functions is unsafe (the compiler cannot verify C's rules). The design rule to imitate: keep unsafe thin — the extern block and the immediate calls inside small, named wrapper functions, so the safe surface (CudaState's methods) never leaks raw pointers or error codes.

RAII and the ownership models. RAII (Resource Acquisition Is Initialization — the idiom where acquiring a resource in a constructor and releasing it in a destructor ties the resource's lifetime to an object's, so drops and panics cannot leak it) is how minfer keeps the host-side CUDA resources leak-proof: every pinned allocation has a Drop impl calling cudaFreeHost (PinnedPool src/cuda.rs:397, PinnedBuf src/cuda.rs:415-424, CaptureStaging src/cuda.rs:435-532). Device memory is the interesting half-exception. A weight buffer is not RAII-managed at all: weights live for the whole process, keyed by name in the registry (weights: Mutex<HashMap<String, (CudaPtr, usize)>>, CudaState::weights (src/cuda.rs:614)), and the registry's replace rule deliberately leaks the stale buffer because a live captured CUDA Graph may still reference the old pointer register_weight (src/cuda/methods/weights.rs:18)) — leaking by decision, with a bound and a comment is legitimate, because the alternative (freeing memory a captured graph still points at) is a use-after-free the GPU hits mid-replay. Scratch buffers, in contrast, are pooled: the backend frees every pool buffer in its Drop (src/graph/cuda_backend.rs:859, impl Drop for CudaBackend), and CudaState offers the raw pair cuda_malloc/cuda_free (src/cuda/methods/buffers.rs:21) for pool use.

One complete ownership path, walked. The simplest non-toy path in the file: register a weight → use it → (never) free it.

  1. Alloc. register_weight(name, data) register_weight (src/cuda/methods/weights.rs:18-97)) first checks the registry — same name and same byte size ⇒ the device copy exists, reuse it and return (:24-31). Otherwise it calls cudaMalloc for data.len() bytes (:39-48); on OOM (out of memory) it prints the byte count and tensor name and returns without registering — the loader's all-weights-registered gate then refuses to enable the GPU backend loudly (a docs/GPU_SAFETY.md invariant). Nothing silently half-works.
  2. Copy (H2D). Still inside register_weight, a stream-ordered cudaMemcpyAsync on the context stream plus a cudaStreamSynchronize uploads the GGUF (the on-disk model format minfer parses) bytes verbatim since #188 — the legacy blocking cudaMemcpyHostToDevice participated in the legacy default stream's implicit global synchronization and could invalidate another thread's capture window (the pre-#188 probe is kept as register_weight_blocking_legacy). Quantized weights are uploaded raw and dequantized on the device by the dequant_*_f16 kernels at first use (a series fact; chapter 03 reads those kernels). On copy failure the freshly allocated buffer is freed before returning (:49-87) — the manual version of RAII: the error path releases what the success path acquired.
  3. Register. The pointer is wrapped and inserted (:93-96): self.weights.lock().unwrap().insert(name.to_string(), (CudaPtr(ptr), data.len())). From here the only way to reach the buffer is get_weight_ptr (src/cuda/methods/weights.rs:565) — the registry is the single owner, and dispatchers resolve by name at execution time.
  4. Use. A decode step later, dispatch resolves the name to a device pointer and passes it to a launcher. The Rust side of add_f32 (src/cuda/methods/elementwise.rs) is a 17-line translation unit from safe Rust to the C launcher: fetch self.stream(), then one unsafe call launch_add_f32(x as *const f32, y as *const f32, z as *mut f32, n as i32, stream). The extern declaration sits at launch_add_f32 (src/cuda/methods/elementwise.rs:25) — these launcher symbols are provided by libcuda_kernels.a, the archive build.rs produces from src/cuda/kernels/*.cu (§2.5). The usize → i32 narrowing and the c_void → *const f32 casts are the FFI layer's whole job — the kernel wants const float* and int, and this is where the host types are made to match. Note what is absent: no error check — launches are enqueued asynchronously and checked at the next sync() (§2.2).
  5. Drop. For weights: never — process-lifetime by design (the registry owns them until exit; the OS reclaims the context). For pool scratch: CudaBackend::drop frees every pool buffer with cudaFree (a context-wide call); the pre-#188 stream lock is gone, and the Drop now first binds this backend's own stream with self.bind() (src/graph/cuda_backend.rs:866) for the teardown's host-side transfers. For the grow-on-demand scratch slots there is a middle pattern, get_or_grow (src/cuda/methods.rs): if the slot's allocation is too small, free the old buffer then allocate the new one, only under the slot's own Mutex so two graph executions cannot grow the same slot concurrently.

That is the complete path: alloc → upload → registry (single owner) → name-resolve at dispatch → lifetime-by-design. Every other part of the device layer — the KV regions (src/cuda/methods/kvstore.rs), the f16 cache (prefill_f16.rs), the MMQ scratch planes (mmq_quant.rs/prefill_mmq.rs) — is a variation on this skeleton.

2.4 Streams and asynchronous execution

What a stream is. A cudaStream_t is a queue of GPU work — kernel launches, memory copies, event records — that the GPU executes in enqueue order. Every asynchronous operation you have met takes a stream argument: kernel<<<grid, block, 0, stream>>>, cudaMemcpyAsync(..., stream). The stream is what makes CUDA asynchronous at all: launching into a stream returns immediately, and the GPU drains the queue at its own pace. The semantics you build everything on: within a stream, total order — item N starts only after item N-1 finishes; that is the correctness backbone, and minfer's "same-stream ordering" rule (docs/GPU_SAFETY.md:238, CUDA rule 6) is just this sentence applied (an async H2D fill is safe because the kernel that reads the buffer is enqueued later on the same stream). Across streams, no order and potential parallelism — work in stream A and stream B may run concurrently (if resources allow). And the default stream (stream 0, used whenever you omit the argument) is special: the legacy default stream synchronizes with all other (blocking) streams — a serialization belt, not just "stream number zero".

CUDA events. A cudaEvent_t is a marker you record into a stream; it completes when every item before it in that stream has finished. Synchronize on one event to wait for a part of a pipeline; cudaEventElapsedTime between two events measures GPU time between the markers (on the device's own clock, immune to host scheduling noise) — how every number in Part 4 was measured.

Toy #3 — serialization vs overlap, measured. Do two kernels actually run in parallel on two non-default streams, and does the default stream really serialize them? The toy launches the same busy-spin kernel twice — first both into the default stream, then one each into two streams — timing each phase with events (Toy #3 of the series; verified with CUDA 13.0 on GB10, sm_121):

// toy3_streams.cu — Toy #3 of the minfer CUDA tutorial (chapter 02).
// Two identical "busy" kernels: launched back-to-back on the default stream
// they serialize; launched on two non-default streams they overlap.
// Timed with cudaEvents. Verified with CUDA 13.0 on GB10 (sm_121):
//   nvcc -arch=sm_121 -O2 toy3_streams.cu -o toy3 && ./toy3
#include <cstdio>
#include <cuda_runtime.h>

#define CK(x) do { cudaError_t e_ = (x); if (e_ != cudaSuccess) { \
    printf("CUDA error %s at %s:%d\n", cudaGetErrorString(e_), __FILE__, __LINE__); \
    return 1; } } while (0)

// Busy kernel: one small block (fits beside other work), spins ~`cycles`
// clocks, then writes one float so the work is not optimized away.
__global__ void busy(float* out, long long cycles) {
    long long start = clock64();
    while (clock64() - start < cycles) { }
    if (threadIdx.x == 0) out[blockIdx.x] = (float)(clock64() - start);
}

int main() {
    float* d = nullptr;
    CK(cudaMalloc(&d, 64 * sizeof(float)));
    cudaEvent_t beg, end;
    CK(cudaEventCreate(&beg));
    CK(cudaEventCreate(&end));
    cudaStream_t s1, s2;
    CK(cudaStreamCreate(&s1));
    CK(cudaStreamCreate(&s2));
    const long long cycles = 200000000LL;   // tune so one kernel ~ tens of ms
    float ms1 = 0.0f, ms2 = 0.0f;

    // Phase 1: both kernels on the DEFAULT stream (0) -> serialized.
    CK(cudaEventRecord(beg));
    busy<<<1, 32>>>(d, cycles);
    busy<<<1, 32>>>(d, cycles);
    CK(cudaEventRecord(end));
    CK(cudaEventSynchronize(end));
    CK(cudaEventElapsedTime(&ms1, beg, end));

    // Phase 2: one kernel per non-default stream -> overlap.
    CK(cudaEventRecord(beg));
    busy<<<1, 32, 0, s1>>>(d, cycles);
    busy<<<1, 32, 0, s2>>>(d, cycles);
    CK(cudaEventRecord(end));
    CK(cudaEventSynchronize(end));
    CK(cudaEventElapsedTime(&ms2, beg, end));

    printf("default stream : %.2f ms (kernel2 waits for kernel1)\n", ms1);
    printf("two streams    : %.2f ms (kernels run concurrently)\n", ms2);
    printf("overlap speedup: %.2fx\n", ms1 / ms2);
    CK(cudaStreamDestroy(s1)); CK(cudaStreamDestroy(s2));
    CK(cudaEventDestroy(beg)); CK(cudaEventDestroy(end));
    CK(cudaFree(d));
    return 0;
}

Observed on the GB10:

$ /usr/local/cuda/bin/nvcc -arch=sm_121 -O2 toy3_streams.cu -o toy3 && ./toy3
default stream : 159.26 ms (kernel2 waits for kernel1)
two streams    : 79.58 ms (kernels run concurrently)
overlap speedup: 2.00x

Read the arithmetic, not just the speedup: one busy kernel is ~80 ms, so the serialized pair costs ~160 ms and the overlapped pair ~80 ms — exactly 2.00×, because the two one-block kernels fit on the GPU simultaneously. The same experiment with big grids can show less overlap (two kernels that each fill every SM cannot physically run side by side — streams give the GPU permission to overlap, not extra hardware). And notice the timing pattern: cudaEventRecord(end) is enqueued, the real wait is cudaEventSynchronize(end) — the event version of the §2.2 sync story.

minfer's actual stream usage. The surprise: for all that machinery, each engine's backend runs one non-blocking stream, created when the backend is built (with_layout, src/graph/cuda_backend.rs:204); stream() (src/cuda/methods/stream.rs:18) answers that bound stream, falling back to the context stream (try_new, src/cuda/methods/init.rs:55) only for an unbound caller. Why one stream, when streams exist for overlap? First, the workload is a dependency chain — a decode step is a strict sequence (norm → matmul → rope → attention → … → lm_head) with nothing to overlap within it. Second, the real per-step overhead was launches, not gaps between kernels — ~100+ launches per decode step, each a few µs of host time — so minfer's answer was not streams but CUDA Graph capture/replay: record the whole step's launches once, then replay the graph as a single launch. Capture is per-stream (only work enqueued on the capturing stream is recorded) and opens in thread-local mode, so each backend captures into its own stream; stream_guard (src/graph/cuda_backend.rs:556) is now a None-returning shim — #188 took the stream-work serialization away, and a caller must not rely on mutual exclusion there. The state machine lives in graph_replay_step (src/graph/cuda_backend.rs:459): executions 1–2 of a split run as plain launches (warmup — llama.cpp's protocol), the 3rd opens the capture window, synchronize() closes it (instantiate + launch once + cache, close_capture_or_sync (src/graph/cuda_backend.rs:573)), and every later execution is a single graph_launch_exec (src/graph/cuda_backend.rs:496). The backend's CudaBackend::synchronize (src/graph/cuda_backend.rs:2307) is the split-boundary drain point from §2.2: bind its own stream, clear the per-execution memos, then close_capture_or_sync. Replay is gated by MINFER_NO_CUDA_GRAPH=1 (falls back to per-kernel launches — §5 makes the launch stream visible in nsys). The mental model: streams are the substrate (one stream, strict order, checked at drain points), and CUDA Graphs are the optimization on top (same stream semantics, N launches folded into 1). Chapter 06 covers the graph machinery; here you only need the stream vocabulary to read it.

2.5 Build and toolchain

Compiling toys by hand. nvcc is NVIDIA's compiler driver — it runs the host C++ compiler on the host parts and its own cicc/ptxas on the device parts, then packs both into one object file. On dgxspark it is not on the default PATH; call it with the full path (this is CUDA 13.0):

$ /usr/local/cuda/bin/nvcc --version
nvcc: NVIDIA (R) Cuda compiler driver
Copyright (c) 2005-2025 NVIDIA Corporation
Built on Wed_Aug_20_01:57:39_PM_PDT_2025
Cuda compilation tools, release 13.0, V13.0.88

The one flag you must think about is the GPU architecture: -arch=sm_NN generates SASS (the GPU's native machine code — "Streamed ASSembler") for compute capability NN, where the digits come from the device itself. On the GB10 (aarch64): nvidia-smi --query-gpu=compute_cap reports 12.1, driver 580.173.02 — so the flag used for every toy in this tutorial is -arch=sm_121, which CUDA 13.0 accepts (verify with nvcc --list-gpu-arch). -arch=native (ask the installed driver) also works on dgxspark and compiles both toys identically — either is fine; the toys pin sm_121 so the command is reproducible on paper. A binary built for a newer arch than your GPU will not load (no SASS match); one built for an older arch only runs if the binary embedded PTX (the intermediate assembly the driver can JIT — Just-In-Time compile — for a newer GPU): exactly the compat layer build.rs sets up below.

How the repo does it. cargo build --release --features cuda never calls nvcc from your shell; build.rs (the Cargo build script that runs before compilation) does the whole pipeline:

  1. Opt-in, and required once opted in. The CUDA section returns early unless CARGO_FEATURE_CUDA is set (build.rs:362, the opt-in early return) — plain builds never touch nvcc. But once the feature is requested, CUDA is required: src/cuda.rs declares the launch_* symbols that only the kernels archive provides, so every failure below panic!s with an actionable message (build.rs:368, the panic path).

  2. Find nvcc and the toolkit root. find_nvcc() (build.rs:578-601) probes CUDA_HOME/CUDA_PATH, then which nvcc, resolving to an absolute path either way; find_cuda_home() (build.rs:603-622) derives the root for -I{home}/include.

  3. Pin the host compiler (ccbin). nvcc uses the first cc/g++ on PATH as its host compiler and hard-fails when that GCC is newer than the toolkit supports (CUDA 13 rejects GCC 15 — the error surfaces confusingly inside <cmath>). detect_host_compiler() (build.rs:680-719) probes nvcc's default first, then g++-15 … g++-11, g++, clang++; the winner is passed as -ccbin (build.rs:385, :476), and MINFER_CUDA_CCBIN overrides the probe.

  4. Probe the architectures. detect_archs() (build.rs:732-763) compiles a one-line dummy kernel for every candidate from sm_70 to sm_121 detect_archs (build.rs:732)) and keeps the ones this nvcc accepts — candidates newer than the toolkit simply fail their probe and are skipped, so one list works on every CUDA version. (The floor is sm_70, not Pascal: the prefill GEMM uses WMMA tensor-core intrinsics that require Volta+, build.rs:725, the sm_70 floor rationale).

  5. Compile once, embed many targets. The single nvcc invocation (build.rs:476, the one nvcc invocation) carries -O3 -fPIC plus one -gencode arch=compute_NN,code=sm_NN per detected arch — SASS for every GPU class — plus two kinds of PTX (build.rs:469-474, the PTX pushes): a backward compute_70/72 (so an older card like a V100 can JIT forward) and a forward compute_{highest} (so a GPU newer than the newest SASS can JIT). The result is the portable fat binary: every probed arch as native SASS, plus PTX from Volta up that any future GPU can JIT.

  6. Archive and link into the Rust binary. The object is packed into libcuda_kernels.a with ar rcs (build.rs:486-495), and two cargo directives hand it to the linker (build.rs:497-498): cargo:rustc-link-search=native={out_dir} + cargo:rustc-link-lib=static=cuda_kernels. The launch_* symbols the extern "C" block of §2.3 declares are resolved against this archive — the whole seam between the two languages.

  7. cudart: static or shared. The CUDA runtime is linked according to the cuda_static feature (build.rs:415, the cudart choice): cuda_static links libcudart_static.a (plus dl/pthread) so the binary needs only the NVIDIA driver at run time; the default links libcudart.so and bakes an rpath — only when the toolkit dir is not a system dir, to avoid shadowing a nix-provided glibc (build.rs:541, the rpath guard). The lib directory is probed rather than assumed (build.rs:633, find_cuda_lib_dir), because distro packages install cudart into the multiarch dir. What is never linked is the driver library libcuda.so.1 — it is dlopen'd lazily at run time (preload_driver, src/cuda/methods/init.rs:34), the same lazy-loading trick as the dlopen declaration at the top of the extern block.

For the knobs this section skipped (cross-compiling, MINFER_CUDA_CCBIN in full, distro-vs-toolkit layout quirks), docs/BUILD.md is the reference — 73 lines, worth reading once before your first --features cuda build.

3. In minfer's code — one op, end to end

Now read one real operation through all layers at once: Op::Silu on a decode step of Qwen2.5-0.5B (n = 4864, one token). This is the path every CUDA node takes; the layered picture first (the one diagram of this chapter):

 scheduler.rs  : execute_node(node, backend)          (pure Rust, safe)
     │
     ▼
 `Op::Silu` (cuda_backend.rs:1026)   Op::Silu arm               (guards → Err or launch)
     │   in_bufs[0].id != out_buf.id?  → copy_d2d (D2D stage)
     ▼
 `silu_f32` (src/cuda/methods/elementwise.rs:145)         CudaState::silu_f32        (thin unsafe wrapper)
     │   self.stream() = this backend's bound cudaStream_t
     ▼
 `launch_silu_f32` (src/cuda/methods/elementwise.rs:39)               extern "C" launch_silu_f32 (FFI declaration)
     ▼
 `launch_silu_f32` (ops_elementwise.cu:385) launch_silu_f32            (grid = ceil-div, <<<>>>)
     ▼
 `silu_f32` (ops_elementwise.cu:189) __global__ silu_f32         (index → guard → math)
     ▼
 [ GPU: 19 blocks × 256 threads, enqueued on the stream, drains async ]
     …
 `CudaState::sync` (src/cuda/methods/events.rs:143)         CudaState::sync()          (latched errors + drain,
                                                       at the split boundary)

The scheduler calls execute_node on whichever backend was assigned at build time (assignment is decided before execution — the never-silently-fallback rule from docs/GPU_SAFETY.md). The CUDA backend's Op::Silu arm (src/graph/cuda_backend.rs:1026, the arm in execute_node_inner) is seven lines:

#![allow(unused)]
fn main() {
// `Op::Silu` (src/graph/cuda_backend.rs:1026-1032)
// In-place op (alias rule, graph rules §5): stage via D2D copy when
// the allocator did not alias the input, then run on the output.
Op::Silu => {
    if in_bufs[0].id != out_buf.id {
        self.copy_d2d(in_bufs[0], out_buf)?;
    }
    self.state.silu_f32(self.ptr_of_ref(out_buf)?, out_buf.len);
    Ok(())
}
}

Why the D2D copy? silu_f32 is an in-place kernel — it reads and writes the same buffer. The compute graph permits a node to alias its input only under the alias rule (sole consumer + same backend); when the liveness allocator did not alias the two buffers (they are distinct pool slots), the backend stages a device-to-device copy first so the in-place kernel can never write a buffer some other node still needs. copy_d2d copy_d2d (src/graph/cuda_backend.rs:1795) resolves both pool slots to device pointers, refuses on a byte-size mismatch (Err — the §2.2 contract), and enqueues cudaMemcpyDeviceToDevice on the backend's bound stream. Two GPU operations (copy + kernel) for the price of one node, both asynchronous, both ordered by the stream.

Then the descent from §2.3 step 4: CudaState::silu_f32 (src/cuda/methods/elementwise.rs) fetches the bound stream and calls the extern launcher (declared at launch_silu_f32 (src/cuda/methods/elementwise.rs:39)); the C launcher computes grid = (4864 + 255) / 256 = 19 and enqueues silu_f32<<<19, 256, 0, stream>>>; the kernel gives each of the 4,864 threads exactly one element — 19 blocks × 256 threads, zero idle (the 4864 = 19 × 256 exact fit from Toy #2). The launch returns in a few microseconds while the kernel may not even have started. Finally the CPU counterpart — this tutorial's pattern is kernel → CPU → why the GPU version looks the way it does (chapter 03 does this line by line for the whole elementwise family). The same op on CPU is vec_silu_f32 vec_silu_f32 (src/vec_ops/silu.rs:8-23)): signature pub fn vec_silu_f32(n: usize, y: &mut [f32], x: &[f32]), an x86_64 AVX2+FMA arm detected at run time vec_silu_f32_avx2 (src/vec_ops/silu.rs:27-53), 8 lanes per step), and the scalar fallback y[i] = x[i] / (1.0 + (-x[i]).exp()) — the same formula as the kernel, one explicit loop index instead of 4,864 materialized threads.

Compare the two and the design falls out:

  • Same formula, same one-element-per-iteration shape — the GPU kernel is the scalar loop turned inside out: the CPU iterates i, the GPU materializes i as a thread. Everything the CUDA version adds (index formula, guard, launcher, stream) is machinery for that inversion. The CPU side branches per machine (AVX2/FMA detected at run time), the GPU side per arch at build time (the §2.5 gencode list); both keep a scalar fallback, and neither hides a failed dispatch (the GPU arm's fallback is an Err, per the safety contract) — the backend's test gates compare against this exact CPU function (src/graph/cuda_backend/tests/elementwise.rs:159-171, the host reference through the same vec_ops) computes vec_*_f32 and asserts the kernel output matches: the toy's CPU-reference pattern, scaled to the whole op set).
  • In-place is a graph decision, not a kernel decision. The CPU signature takes y and x separately (out-of-place); the GPU kernel is in-place — on the device an extra buffer costs a full allocation and an extra stream of DRAM traffic, while on the CPU the write is nearly free. Same math, different economics — the source of most "why does the GPU version look different" moments you will have in Part 4.

4. Performance intuition

Numbers, not adjectives. All shapes are Qwen2.5-0.5B decode (one token, nt = 1) — the case this chapter's kernels actually run.

What an elementwise kernel costs. silu_f32 at n = 4864 reads 4864 × 4 B = 19.5 KB and writes 19.5 KB — ~38.9 KB of DRAM traffic per layer per step, ~0.9 MB across the 24 transformer layers (Qwen2.5-0.5B is 24 layers — docs/inference_e2e_walkthrough/05-graph-builder-ir.md:161 (§2 "The main event: one forward pass, node by node")). Compare the attention output projection (attn_q, an [896, 896] Q4_0 weight ≈ 451 KB per layer — walkthrough doc 14, docs/inference_e2e_walkthrough/14-metal-backend.md:71 (§2 "Why f32 activations on the GPU when the CPU quan")) and elementwise ops are noise in the byte budget. Their cost is therefore not bandwidth but latency: launch overhead (microseconds per launch, host-side) plus the kernel's start-to-finish time. This is why the elementwise family is where fusion lives: decode never launches silu_f32 standalone if it can help it — the fused-FFN path runs swiglu_f32_off (src/cuda/kernels/ops_elementwise.cu:201, silu+mul in one pass over the concatenated gate|up buffer) and the prefill path fuses the q8 quantization epilogue into the same kernel (swiglu_quant_pad40, src/cuda/kernels/ops_elementwise.cu:215 — its comment block is worth reading for the "no early return — every thread reaches the barrier" discipline).

Thread-count arithmetic. 4,864 threads is a tiny grid. The GB10's SM count is queried at runtime get_attr (src/cuda/methods/init.rs:112) reads CUDA_DEV_ATTR_MULTIPROC_COUNT); with a few dozen SMs each holding up to ~2048 resident threads, one 4,864-thread kernel cannot come close to filling the machine — most blocks run, finish, and leave SMs idle. That is fine for a 2-µs, latency-bound kernel, and it is the quantitative reason decode kernels in minfer are judged on bytes moved per second rather than occupancy (how full the SMs are): at this grid size there is nothing to fill with. Contrast prefill (nt = 512): the same op is 512 × 4864 ≈ 2.5M elements → 9,728 blocks of 256 threads — now the grid spans every SM many times over and bandwidth becomes the limit. The grid-size formula does not change; the regime does.

Launch overhead is the decode tax, and CUDA Graphs are the subtraction. The D4-4 record in docs/CUDA_OPTIMIZATION.md:153 ("§0 Master history table — the complete optimizat") measures what graph replay recovers at "~2 µs/launch" of graph gap; multiply by the ~100+ launches of a decode step (the same record's census) and you get hundreds of microseconds — a real fraction of a small-model decode step. That, not kernel bandwidth, is why §2.4's capture/replay machinery exists, and why MINFER_NO_CUDA_GRAPH=1 is the first knob to flip when you want to see the launch stream (§5).

What would make an elementwise kernel slow — a checklist that applies to any kernel you write. Uncoalesced access: z[tid] = x[tid] + y[tid] with consecutive tid gives each warp (32 threads) one contiguous 128-B transaction per array; a strided index (tid * 16) puts each thread's 4 B in a different transaction — up to 32× more memory transactions for the same bytes (Part 3 dissects this; for now: keep tid consecutive). A guard that kills the warp shape: if (i >= n) return; only costs when it divides within a warp (at most the last 31 threads of the grid), while a guard that diverges every warp (if (x[i] < 0) return;) forces the remaining lanes to idle through the skipped code. An idle grid: 1 block for a 4,864-element op would run 4,864 iterations inside one block — serial latency, no SM parallelism; the ceil-div grid is the standard answer. Unnecessary sync: every cudaDeviceSynchronize() between launches drains the whole device — minfer's per-split sync() (§2.2) is the entire synchronization budget of a step; add none.

Toy #3's 2.00× came from two kernels that fit side by side; when two kernels each saturate DRAM bandwidth (as two big prefill GEMMs would), two streams buy zero. minfer's one stream per engine is therefore the correct design for a dependency-chained workload, with CUDA Graphs attacking the actual overhead (launches) — keep both tools in mind and let the measurement pick.

5. Try it / Observe

The toys. Both are complete listings in this chapter (§2.1 and §2.4) — save them and run:

$ export PATH=/usr/local/cuda/bin:$PATH        # nvcc is not on PATH by default
$ nvcc -arch=sm_121 -O2 toy2_silu.cu -o toy2 && ./toy2 && ./toy2 896
$ nvcc -arch=sm_121 -O2 toy3_streams.cu -o toy3 && ./toy3

(expected outputs are pasted under each toy; -arch=native works too on this GB10 / CUDA 13.0 machine). Two experiments beyond the pasted runs:

  • In toy2, delete the if (i >= n) return; guard and run ./toy2 896 — nothing visibly breaks (the extra 128 threads write past the end of the 3.5 KB allocation into unallocated context memory), which is exactly why the §2.2 sticky-error + sync discipline exists. Run compute-sanitizer --tool memcheck ./toy2 896 to see the OOB (out-of-bounds) access reported.
  • In toy3, replace both <<<1, 32>>> launches with <<<64, 256>>> and rerun — the kernels now saturate the GPU and the "overlap" speedup shrinks toward 1.0× (stream parallelism bounded by resources).

Watch minfer's real kernels.

$ cargo build --release --features cuda
$ MINFER_NO_CUDA_GRAPH=1 ./target/release/minfer <model.gguf> "hello"
#   decode without graph replay — every kernel is a separate launch; compare
#   tokens/s against the default (replay) run to feel §4's launch tax

See the stream timeline. One nsys command gives you the §2.4 picture — kernels queued back-to-back on one stream, and (with graphs on) the replay as a single launch entry; ncu then drills into one kernel (all profiling tools live under /usr/local/cuda/bin):

$ nsys profile -o /tmp/dec --force-overwrite=true \
    ./target/release/minfer <model.gguf> "hello" -n 16
$ nsys stats --report cuda_gpu_kern_sum /tmp/dec.nsys-rep     # which kernel, how long
$ nsys stats --report cuda_gpu_trace /tmp/dec.nsys-rep | head -40   # launch timeline
$ ncu --kernel-name regex:silu --launch-count 3 --set basic \
    ./target/release/minfer <model.gguf> "hello" -n 8

6. Cross-references

← 01 · What kind of machine is a GPU · Index · 03 · Reading minfer's kernels I →

03 · Reading minfer's kernels I — elementwise, dequant, embedding

Part: Part 3a — first real kernels. Prereq: chapters 01–02 (you can write an elementwise kernel, you know the index formulas and the device memory API). Code: src/cuda/kernels/*.cu, src/block.rs, src/graph/cuda_backend.rs — all file:line citations verified against the tree at writing time (the function name is the stable address, the line number a convenience).

1. Background — where this sits

Chapters 01 and 02 gave you the mental model and the language surface: a kernel is a function every launched thread executes, threads are grouped into blocks and blocks into a grid, and the first line of nearly every kernel is the same index formula that turns those coordinates into one flat position. Chapter 02 also showed you minfer's Rust-side memory wrapper, so you know that device buffers are opaque *mut c_void handles owned by the backend.

This chapter starts the ladder the tutorial is really about: reading minfer's actual kernels. We take the three easiest kernel families in the file — the elementwise add pair, the Q4_0 weight dequantizer, and the embedding gather — and read each one the way a contributor should: source excerpt, line by line, then the CPU counterpart from the inference walkthrough, then the performance arithmetic. These are not warm-ups chosen for cuteness. Every single forward pass of the graph executes them: the residual stream is a chain of add_f32 calls, every quantized weight that feeds the f16 prefill GEMM was produced by a dequant_q*_f16 kernel, and the very first real op of any transformer forward — turning token ids into vectors — is embed_rows_q4_0 or one of its type siblings.

You will also meet the dispatch chain, the pattern every later chapter reuses:

flowchart LR
    A["Rust: Op match arm<br/>graph/cuda_backend.rs"] --> B["Rust device layer<br/>CudaState method (cuda/methods/*.rs)"]
    B --> C["C launcher: grid sizing<br/>launch_* (cuda/kernels/*.cu)"]
    C --> D["__global__ kernel<br/>one thread's worth of math"]

Keep it in mind while reading: a graph node named Add does not "run add_f32" directly. It runs a Rust match arm, which calls a CudaState method, which calls a C launcher that computes the grid, which finally launches the kernel. Four hops, each with a distinct job. When a later chapter says "the Op::MatMul arm routes to MMQ", you will know exactly which hop does the routing.

2. Principle — the concepts

2.1 The kernel inventory — the map of src/cuda/kernels/

The forensics protocol from STYLE.md starts with enumeration. Run:

grep -n '__global__' src/cuda/kernels/*.cu

Today that prints 99 __global__ void definitions — plus six prose mentions — across 17 files (10,147 lines). Nobody memorizes 99 entries; you navigate by family. Here is the map (file and line):

FamilyExamples (file, first hit line)What it doesWhere taught
Elementwise / epilogueadd_f32 (ops_elementwise.cu:163), add_bias_f32 (:151), silu_f32 (:189), rope_f32 (:263), store_kv_f16 (kv_store.cu:31)one pass over a buffer, per-element maththis chapter
Dequant to f16dequant_q8_0_f16 (gemm_wmma.cu:23) … dequant_q6_k_f16 (:174)quantized weight bytes → dense f16this chapter
Embedding gatherembed_rows_q8_0 (ops_misc.cu:160) … embed_rows_q4_1 (:671)token ids → dequantized weight rowsthis chapter
Small helpersgather_rows_f32 (ops_misc.cu:145), f32_bits_to_i32 (ops_elementwise.cu:249), convert_f32_f16_kernel (gemm_wmma.cu:203)glue: format conversion, device-side decodesthis chapter (quick-read)
Decode matvec (MMVQ)q4_k_q8_mmvq_multi (mmvq_multi.cu:25), q4_0_q8_mmvq (:446)one output row per block, dp4a integer dotsch 04
Prefill GEMMgemm_f16_nt_kernel_t (gemm_wmma.cu:367), mmq_nt_kernel (mmq_int8.cu:236)tiled tensor-core GEMM, f16 or int8ch 04
Attentiongqa_attn_f32_f16kv (attention_decode.cu:18), fa_prefill_kv (attention_prefill.cu:117)GQA attention, flash-style prefillch 05
Fused decode tailattn_bias_rope_store_f32 (kv_store.cu:118)bias×3 + rope×2 + store×2 in one launchch 05
Activation quantizequantize_q8_0_pad40_t (mmvq_aquant.cu:85), rms_norm_quant_f32_t (:169)f32 activations → q8 blocks for MMQch 04

Two structural facts the table hides, and both matter for reading:

  1. Kernels and launchers live in the same translation unit, but are different APIs. The __global__ functions are device code. The void launch_* functions (plain C++, called through FFI from Rust) own the grid arithmetic. When you want to know "how many threads does this launch", find the launcher, not the kernel.
  2. Type-generic dispatch is done with switch (type_id), not C++ templates. launch_embed_rows (ops_misc.cu:699) and launch_dequant_f16 (gemm_wmma.cu:777) each take an integer type id and select one of eight concrete kernels. The Rust side owns the id tables — and, as §3.2 will show, the two tables do not use the same numbering. That is a real trap when you first read the code.

2.2 The three thread-to-data mapping patterns

Nearly every easy kernel in these files answers one question first: what does one thread process? The elementwise/dequant/embed families use three answers:

  • One thread per element — add_f32, mul_f32, silu_f32, f32_bits_to_i32. Index = flat position. Simplest possible; used when the per-element work is trivial.
  • One thread per 32-element quant block — the dequant_q*_f16 family and the 32-element embed_rows_* variants. Q4_0 stores 32 values in 18 packed bytes; unpacking them one element per thread would mean every thread re-reading the scale byte and masking its own nibble out of a shared byte. Giving one thread the whole block lets it read the scale once and write a contiguous run of outputs.
  • One thread per N elements (vectorized) — convert_f32_f16_kernel processes 8 elements per thread so it can load a float4 pair and store __half2 pairs (§3.4). Same math, fewer memory transactions.

And the grid comes from one formula, the ceil-div you met in chapter 02:

grid = (total + block - 1) / block;   // ceil(total / block)
if (grid > 2147483647LL) grid = 2147483647LL;   // x-dimension cap

You will find that pair (or its dim3 variant) in every launcher. Two details worth internalizing now. First, the launcher computes total in a long long — products like od * (id / 32) can overflow int for large matrices, and the 64-bit accumulation happens before the cast to the grid dimension. Second, the 2147483647LL clamp is a defensive guard, not a live case in any verified model — the CUDA x-dimension limit is 2³¹−1 and minfer simply refuses to exceed it rather than launch something invalid.

2.3 Where the Rust side hands over

src/cuda/methods/*.rs declares the launchers, one extern "C" block per family (the FFI surface, e.g. launch_add_f32 (src/cuda/methods/elementwise.rs:25), launch_dequant_f16 (src/cuda/methods/prefill_f16.rs:12)) and wraps each in a small safe method on CudaState. The graph backend never sees kernel names; it sees graph ops. The three call sites this chapter follows:

  • Op::Add → CudaState::add_f32 — src/graph/cuda_backend.rs:1003
  • the MatMul bias epilogue → CudaState::add_bias_f32 — :1514
  • Op::GetRows → embed_rows_on_gpu / gather_rows_f32_on_gpu — :972 / :986

One safety rule from docs/GPU_SAFETY.md colors all of these arms: when a kernel's input violates its invariants (wrong sizes, unsupported layout), the arm returns Err from execute_node. There is no silent fallback to another backend mid-run — backend assignment was decided at build time.

3. In minfer's code

3.1 add_f32 and add_bias_f32 — the residual stream

The transformer's residual stream ("the highway where each layer's contribution is an update, not a replacement" — walkthrough doc 11 §2.0) is executed by these two kernels. In the 0.5B graph the residual adds are among the most frequent nodes (doc 05 walks the topology; doc 11 explains the pre-norm layout), which makes them the right first read: they are everywhere, and they are the "hello world" shape of this codebase.

The kernel — add_f32 (src/cuda/kernels/ops_elementwise.cu:163):

__global__ void add_f32(
    const float* __restrict__ x,
    const float* __restrict__ y,
    float* __restrict__ z,
    int n
) {
    int tid = blockIdx.x * blockDim.x + threadIdx.x;
    if (tid >= n) return;
    z[tid] = x[tid] + y[tid];
}

Line by line:

  • blockIdx.x * blockDim.x + threadIdx.x — the chapter-02 index formula. Blocks are 1-D, threads are 1-D, so every thread gets a unique flat index over the whole grid.
  • if (tid >= n) return; — the guard. The grid is ceil-div sized, so the last block usually launches more threads than there are elements; the surplus threads exit here instead of writing past the buffer. (The mismatched-shape failure mode this prevents is exactly Toy #2's exercise.)
  • z[tid] = x[tid] + y[tid]; — the whole "computation": one fused multiply-add-free, cache-line-friendly load-load-store. The __restrict__ qualifiers promise the compiler the three buffers do not alias, which lets nvcc keep the loads before the store.

The launcher — launch_add_f32 (src/cuda/kernels/ops_elementwise.cu:363):

void launch_add_f32(
    const float* x, const float* y, float* z, int n, cudaStream_t stream
) {
    int block_sz = 256;
    dim3 block(block_sz, 1, 1);
    dim3 grid((n + block_sz - 1) / block_sz, 1, 1);
    add_f32<<<grid, minfer_launch_block("launch:add_f32", block), 0, stream>>>(x, y, z, n);
}

The ceil-div is over the total element count n — not over tokens, not over rows. 256 threads per block is the sources' default block size (a multiple of 32, the warp width, so no partially-filled warp). The stream argument (cudaStream_t — a queue of device work; every kernel in one stream runs in order) is threaded through from Rust so the backend can order its own work without global synchronization. #162 brackets each launch with minfer_launch_prelude/minfer_launch_ok, which read its own error (elided above).

The Rust call chain: the backend's Op::Add arm checks that both inputs have the same element count as the output and calls CudaState::add_f32 (src/cuda/methods/elementwise.rs), which is a thin FFI shim. The dispatch site is Op::Add (src/graph/cuda_backend.rs:998):

#![allow(unused)]
fn main() {
Op::Add => {
    let n = out_buf.len;
    if in_bufs[0].len != n || in_bufs[1].len != n {
        return Err(format!("cuda: {}: add input size mismatch", node.name));
    }
    self.state.add_f32(
        self.ptr_of_ref(in_bufs[0])?,
        self.ptr_of_ref(in_bufs[1])?,
        self.ptr_of_ref(out_buf)?,
        n,
    );
    Ok(())
}
}

Note what the arm does not do: it does not know or care that this Add is a residual. Graph topology (built in GraphBuilder) decided the wiring; the backend only sees buffers and counts. And the shape mismatch is an Err — the GPU_SAFETY rule from §2.3, applied to a five-line kernel.

add_bias_f32 looks similar but solves a different indexing problem — broadcasting a one-dimensional bias across token rows. Kernel, add_bias_f32 (src/cuda/kernels/ops_elementwise.cu:151):

__global__ void add_bias_f32(
    float* __restrict__ y,
    const float* __restrict__ b,
    int d
) {
    int t = blockIdx.x, i = threadIdx.x + blockIdx.y * blockDim.x;
    if (i >= d) return;
    y[t * d + i] += b[i];
}
  • blockIdx.x is the token row t — one block per row of the [nt][d] activation buffer, not one thread per element of the whole buffer. This is the second mapping pattern of §2.2 in a 2-D disguise: the grid encodes the row, the thread index encodes the column.
  • threadIdx.x + blockIdx.y * blockDim.x — the column index. d can exceed one block's width (a 0.5B layer has d = 896; a 7B one has 3,584 or 18,944), so the launcher folds the remainder into grid.y: grid(n, (d + 63) / 64) with 64 threads per block (launch_add_bias_f32, ops_elementwise.cu:353). For 896: grid.y = 14, so a 30-token prefill launches 30 × 14 = 420 blocks of 64.
  • y[t * d + i] += b[i]; — in-place accumulate. The bias vector b[i] is read by every row, so it stays hot in L2 across the grid.

The bias call site sits inside the MatMul arm (src/graph/cuda_backend.rs:1510-1514, the add_bias_f32 epilogue), and the comment there records the one bug-prone detail of this kernel's contract:

#![allow(unused)]
fn main() {
// add_bias_f32's last argument is the ROW COUNT (nt), not
// the total element count — the kernel grid maps one block
// row per token (a wrong count writes out of bounds).
self.state.add_bias_f32(self.ptr_of_ref(out_buf)?, bptr, od, nt);
}

Pass the element count instead of the row count and grid.x becomes nt * d rows — the kernel indexes y[t * d + i] far past the buffer. The Rust wrapper's docstring add_bias_f32 (src/cuda/methods/elementwise.rs:104-107) repeats the warning. This is the chapter's first lesson in grid-shape contracts: a kernel is not just its body, it is the geometry its launcher assumes.

CPU counterpart. Doc 11 §2.0–§2.1 walks the same residual adds on the CPU (Op::Add nodes over the residual stream, vec_ops/ helpers), and doc 06 covers how the fusion pass minimizes how often they materialize. The math is identical; the only difference is who schedules the loop — the CPU does a vec_add_f32 over one core's SIMD lanes, the GPU spreads it across 26,880 threads (§4 does that arithmetic).

3.2 dequant_q4_0_f16 — unpacking 4-bit weights on device

The data layout first

Everything in this section depends on 18 bytes. BlockQ4_0 (src/block.rs:65):

#![allow(unused)]
fn main() {
// Q4_0 — 4-bit quantization, 32 elements per block (line 184-189)
// Each value is stored as a 4-bit nibble (signed, offset by 8)
// Scale is fp16
#[derive(Clone, Copy)]
#[repr(C)]
pub struct BlockQ4_0 {
    pub d: Fp16,      // delta (scale)
    pub qs: [u8; 16], // nibbles / quants (32 × 4-bit = 16 bytes)
}
}

with pub const Q4B: usize = 18 at Q4B (block.rs:30) and a compile-time assert!(core::mem::size_of::<BlockQ4_0>() == 2 + 16) at :203. So the byte layout is exactly:

one BlockQ4_0 = 18 bytes = 32 dequantized f32/f16 values
┌──────────┬─────────────────────────────────────────────┐
│ d (2 B)  │ qs[0] … qs[15] (16 B = 32 nibbles)          │
│ f16 scale│ byte j holds: LOW nibble = elem j,          │
│          │              HIGH nibble = elem j+16        │
└──────────┴─────────────────────────────────────────────┘
value = d * (nibble - 8)        // the +8 stored offset

Two conventions to burn in, because every quant kernel in these files assumes them:

  • Nibble order: element j comes from the low 4 bits of byte j; element j + 16 from the high 4 bits — q4_0_q8_0_matmul (matmul_f32act.cu:14-64) decodes exactly that, as do the embed and dequant kernels.
  • The +8 offset: minfer (like llama.cpp) stores round(v/d) + 8, so the unsigned nibble 0..15 maps back by subtracting 8 — that is the - 8.0f you will see in every Q4_0 body. Q4_K weights instead carry a per-sub-block min (walkthrough doc 10 §2.4 covers the CPU side of that pairing).

The kernel

dequant_q4_0_f16 (src/cuda/kernels/gemm_wmma.cu:38) (family header + type-id table at :10-21):

__global__ void dequant_q4_0_f16(
    const uint8_t* __restrict__ w, __half* __restrict__ out, int od, int id
) {
    int nb = id / 32;
    long long g = (long long)blockIdx.x * blockDim.x + threadIdx.x;
    if (g >= (long long)od * nb) return;
    int row = (int)(g / nb);
    const uint8_t* blk = w + g * 18;
    float d = h2f(*reinterpret_cast<const uint16_t*>(blk));
    const uint8_t* q = blk + 2;
    __half* o = out + (long long)row * id + (int)(g % nb) * 32;
    // minfer Q4_0 stores round(v/d) + 8 (same -8 offset as the matmuls).
    #pragma unroll
    for (int i = 0; i < 16; i++) {
        o[i]      = __float2half(d * (float(q[i] & 0x0F) - 8.0f));
        o[i + 16] = __float2half(d * (float(q[i] >> 4) - 8.0f));
    }
}

What one thread processes: one 32-element block — precisely one BlockQ4_0 of one weight row. Not one element, not one row. nb = id / 32 is the number of blocks per row (input dim id split into 32-value groups); the grid has od * nb threads. Reading it line by line:

  • g in long long — the global thread index, computed in 64 bits because od * nb for a 7B-class tensor is in the millions and a 32-bit intermediate is an overflow waiting for a bigger model.
  • row = g / nb and the g % nb below split the flat index back into (row, block-within-row). Same 2-D-from-1-D trick as add_bias_f32, but derived instead of launched.
  • blk = w + g * 18 — the source pointer. Because the weight stream is row-major and each row is exactly nb blocks, the flat block index g is also the byte offset in units of 18. This only works because the row length id is a multiple of 32 — the backend gates it before dispatching: the MatMul arm rejects id % 32 != 0 (src/graph/cuda_backend.rs:1481, the quant-block gate) and the f16 warm path requires id % 256 == 0 (warm_w16, src/cuda/methods/prefill_f16.rs:89).
  • h2f(...) — the shared helper h2f (src/cuda/kernels/common.cuh:97): reinterpret the 2 scale bytes as __half and convert to f32. The scale is stored f16, read once per block.
  • The unrolled loop — each byte yields two f16 outputs: & 0x0F takes the low nibble (elements 0–15), >> 4 the high nibble (elements 16–31), subtract 8, multiply by the scale, convert to half with __float2half (round-to-nearest). #pragma unroll tells nvcc to expand the 16 iterations; the loop bounds are compile-time constants, so the branch-free expansion is free code size for eliminated loop overhead.
  • o = out + row * id + (g % nb) * 32 — the destination pointer: row start in the dense [od][id] f16 matrix, plus 32 elements per block. The writes of one thread are fully contiguous.

The launcher — launch_dequant_f16 (src/cuda/kernels/gemm_wmma.cu:777):

void launch_dequant_f16(
    int type_id, const uint8_t* w, __half* out,
    int od, int id, int block_stride, cudaStream_t stream
) {
    int block = 256;
    long long total;
    switch (type_id) {
        case 7: total = (long long)od * (id / 16); break;            // q6_K
        default: total = (long long)od * (id / 32); break;           // all others
    }
    long long grid = (total + block - 1) / block;
    if (grid > 2147483647LL) grid = 2147483647LL;
    switch (type_id) {
        case 0: dequant_q8_0_f16<<<…>>>(w, out, od, id); break;
        case 1: dequant_q4_0_f16<<<…>>>(w, out, od, id); break;
        ...

Grid = ceil-div over od * nb blocks (Q6_K splits into 16-element sub-blocks, hence its special case). One launch geometry serves eight quant types; only the per-thread decode differs. #162 wraps each arm in minfer_launch_prelude/minfer_launch_block/minfer_launch_ok (elided above).

Who calls it, and when — the honest picture

The series shorthand is "quantized weights are dequantized to f16 on device at load" (STYLE.md, series facts). Reading the code gives you the precise version, which is worth knowing because it explains when you will see these kernels in a profile:

  • The persistent f16 weight cache (Phase 8p) is warmed at load time by the model loaders: enable_w16_cache (src/models/qwen2/loader.rs:685) + warm_w16 (:690) per weight. warm_w16 (src/cuda/methods/prefill_f16.rs:89) maps TensorType → the same type ids and calls w16_get, which launches the dequant). But only for models whose matmul weights total ≥ W16_ENABLE_BYTES (src/cuda.rs:597) = 2 GiB and when the int8 MMQ prefill path is not active (mmq_active (qwen2/loader.rs:683)). A 0.5B Q4_0 model is far below that bar (its largest tensor, tok_embd, is 76.6 MB in Q4_0 — §4.1) and therefore runs no dequant at load today.
  • Otherwise launch_dequant_f16 runs per call into a scratch buffer, from the f16 prefill GEMM path prefill_gemm_f16 (src/cuda/methods.rs:224), launch at w16_get (src/cuda/methods/prefill_f16.rs:119).
  • Why a cache at all: the two-pass prefill GEMM used to dequantize W on every call — "288 ms per 7B @2K forward" — because weights are immutable after registration, the dequant result is cached per weight pointer (w16_cache (src/cuda.rs:615-623)).

And why dequantize at all, when the CPU side made a point of never dequantizing at load (walkthrough doc 10 §2.1's bandwidth argument)? The answer is the weight-residency rule plus the hardware target: weights are uploaded once and stay on the GPU for the process lifetime (docs/CUDA-TECH-PRIMER.md §6.1), so an f16 copy costs device memory capacity but no per-token PCIe/host traffic; and the tensor-core GEMM wants dense f16 tiles — wmma fragments cannot read nibble-packed blocks. On the CPU, dequantizing at load would quadruple RAM and the bytes streamed per token; on the GPU it buys access to hardware the packed format cannot feed. (One line, as promised: TECH-PRIMER §6.1 + docs/CUDA-BACKEND-DESIGN.md hold the full story.)

An honest footnote from the same dispatch comment (src/cuda/methods/dispatch.rs:173-183): the default prefill today is the int8 MMQ GEMM, which streams raw quantized bytes and never touches dequant_*_f16 — MINFER_MMQ=0 escapes to the f16 wmma path that does. Both paths coexist; §5 shows how to run each.

CPU counterpart. Doc 10 §2.1 is the CPU-side dequant argument and §2.4 the K-quant pairing (Q8_K activations with precomputed bsums); the CPU dequantize helpers live in quants/ and the scalar reference formula (nibble − 8) · d appears in doc 10's parity notes. The GPU kernel is the same formula — that is the point of the parity tests.

3.3 embed_rows_q4_0 — the gather

The first real op of every forward: turn token ids into embedding vectors. The family comment (src/cuda/kernels/ops_misc.cu:139-143) states the job:

// Embedding = gather + dequantize weight rows on device (removes the CPU
// round trips around the prefill's embed and G3 tail get_rows). ids are
// I32-as-f32 bit patterns (exact for |v| < 2^24), read via __float2int_rn.
// The generic f32 gather (get_rows: out[t*n+i] = x[ids[t]*n+i]) shares the
// f32 kernel.

The kernel — embed_rows_q4_0 (src/cuda/kernels/ops_misc.cu:180):

__global__ void embed_rows_q4_0(
    const uint8_t* __restrict__ w,
    const float* __restrict__ ids,
    float* __restrict__ out,
    int n_embd, int nt
) {
    const int BS = 18; // f16 d + 16 nibble bytes
    int nb = n_embd / 32;
    int tid = blockIdx.x * blockDim.x + threadIdx.x;
    if (tid >= nt * nb) return;
    int t = tid / nb, b = tid % nb;
    int id = __float_as_int(ids[t]); // I32-as-f32 bit pattern (graph rule §4)
    const uint8_t* blk = w + ((long long)id * nb + b) * BS;
    float d = h2f(*reinterpret_cast<const uint16_t*>(blk));
    const uint8_t* q = blk + 2;
    float* o = out + (long long)t * n_embd + b * 32;
    // element j = LOW nibble of byte j; element j+16 = HIGH nibble.
    // minfer Q4_0 stores round(v/d) + 8 (same -8 offset as the matmuls and
    // the CPU embed path).
    #pragma unroll
    for (int i = 0; i < 16; i++) {
        o[i]      = d * (float(q[i] & 0x0F) - 8.0f);
        o[i + 16] = d * (float(q[i] >> 4) - 8.0f);
    }
}

What one thread processes: one 32-element block of one token's embedding row — t = tid / nb picks the token, b = tid % nb the block within that token's row. The dequant body is byte-for-byte the Q4_0 recipe from §3.2; what is new is the gather around it:

  • int id = __float_as_int(ids[t]); — the graph convention that integer inputs (token ids, positions) ride through f32 buffers as bit patterns (f32::from_bits(v) at fill time — fill_input_i32 (src/graph/alloc.rs:1903)). __float_as_int is a bit reinterpretation, not a numeric conversion: it hands back the exact i32 that was stored. A reading note in the forensics spirit: the family comment above still says "read via __float2int_rn", but the code below it bit-casts — the comment drifted, the bit-cast is the correct choice, because __float2int_rn rounds a float value and would mangle any bit pattern that is not a small float. When comment and code disagree, the code plus the parity tests win — and you just learned why this convention exists: it lets the allocator treat all inputs as bytes while keeping kernels CUDA-Graph-replayable with no host sync.
  • blk = w + ((long long)id * nb + b) * BS — the gather itself. id (the token) selects the row, b the block, BS = 18 the block stride. One long multiply guards the id * nb product (a 151,936-row table times 28 blocks overflows nothing here, but the guard is the habit).

How this maps to the weight layout convention. Repo AGENTS rule 4: weight metadata is [in, out], memory is row-major [out][in]; activations are token-major [nt][d]. For token_embd the metadata shape is [n_embd, n_vocab] = [896, 151936] on 0.5B, so memory holds 151,936 rows of 896 contiguous values — row id is exactly token id's embedding vector. The gather is then the degenerate matmul: a one-hot times the weight matrix picks one row (doc 05's graph excerpt shows the node: embed(3) GetRows — h = one token_embd row per id, [896, nt]). The output is token-major f32 [nt][896] — the layout every later kernel assumes.

Dispatch. The backend arm Op::GetRows (src/graph/cuda_backend.rs:962) matches Op::GetRows on metadata: with NodeMeta::Embed it calls embed_rows_on_gpu (:972), with plain metadata it calls the generic gather_rows_f32_on_gpu (:986) — the same GetRows node serves the embedding and the tail-row selects before lm_head (doc 09 §2.4). The Rust wrapper embed_rows_on_gpu (src/cuda/methods/gpu_act.rs:176) maps TensorType to (type_id, block_stride) — Q4_0 => (1, 18) at src/cuda/methods/gpu_act.rs:189 — and F32 embeddings skip the quant kernels entirely by calling the f32 gather (src/cuda/methods/gpu_act.rs:196). The C launcher (launch_embed_rows, ops_misc.cu:699) computes the grid per type (eight embed_rows_* kernels behind one switch) with the same one-thread-per-32-block geometry.

Note what this kernel does not do: it does not consult any f16 cache. Embedding rows are read once per token per step — a gather of nt rows out of 151,936 — so materializing the whole table as f16 would only grow the footprint. Matmul weights read every row on every call, which is why they get the cache (§3.2) and the embedding does not. Same quant format, opposite caching decisions, and the read pattern is why.

CPU counterpart. Doc 05 introduces the GetRows/embedding node and its CPU execution (b.embedding(...) → cpu_backend's get_rows), doc 09 §3.2 shows the tail-row selects in the prefill path, and doc 11 §2.1 explains why the graph has GetRows nodes at the lm_head at all. The CPU path pays a host↔device round trip for a GPU-resident model — the family comment's "removes the CPU round trips" is the reason the kernel exists.

3.4 Quick-read table — the small helpers

Three glue kernels you will meet constantly; one row each, with the one line that carries the idea.

Kernel (file:line)PurposeThe one interesting line
convert_f32_f16_kernel (gemm_wmma.cu:203)f32 activations → f16, feeding the wmma prefill GEMM:208 — base = (…blockIdx.x * blockDim.x + threadIdx.x) * 8: the thread index is multiplied by 8; each thread float4-loads 8 f32 and stores 4 __half2 (:213), 8× fewer transactions for the same traffic (the P1 comment at :206). Launcher launch_convert_f16 :833 sizes the grid over n/8.
f32_bits_to_i32 (ops_elementwise.cu:249)positions/token ids arrive as I32-as-f32 bit patterns; rope/store/attention kernels want raw int*:256 — dst[tid] = __float_as_int(src[tid]); the whole kernel is that line: one device-side bit reinterpretation pass, so the per-layer path never syncs with the host (comment :243). Rust entry bits_to_i32 src/cuda/methods/kvstore.rs:146, called from positions_i32 (cuda_backend.rs:1813, launch :1843).
gather_rows_f32 (ops_misc.cu:145)the quant-free GetRows: out[t*n+i] = x[ids[t]*n+i]:157 — out[idx] = src[(long long)id * n + i]; the classic gather: one flat index decoded into (t, i), the id looked up per thread. Same (1, 18)-style dispatch you saw in §3.3 is what routes F32 embeddings and the tail-row selects here.

All three are one-thread-per-element kernels with the usual ceil-div launcher; if §3.1 made sense, these read themselves.

4. Performance intuition

4.1 Bytes per element — before and after dequant

The fixed exchange rate of this chapter, from the block layouts (src/block.rs:30-38, walkthrough doc 02):

RepresentationBytes / element0.5B tok_embd (151,936 × 896)
Q4_0 packed18/32 = 0.5625 B76,575,744 B ≈ 76.6 MB
f16 (dequant target)2 B272,269,312 B ≈ 272.3 MB
f32 (CPU reference)4 B544,538,624 B ≈ 544.5 MB

What that does to bandwidth, both directions:

  • Every kernel that reads the f16 copy pays 3.56× the weight bytes that the packed MMQ kernels pay (2 vs 0.5625 B/elem). That is the standing cost of the f16 prefill path, and the reason the int8 MMQ GEMM — which streams raw nibbles — is the default (dispatch comment, src/cuda/methods/dispatch.rs:173-183).
  • The one-time dequant itself moves ≈ 349 MB (read 76.6 + write 272.3) per weight tensor of that size, which is why the campaign cached the result: doing it per call cost a measured 288 ms per 7B @2K forward before Phase 8p w16_get (src/cuda/methods/prefill_f16.rs:119).
  • Versus f32, the f16 copy still halves weight traffic — the same 2× argument that made the KV cache f16 (Phase 8b).

4.2 Threads launched — three real launches

Dims verified in-repo: Qwen2.5-0.5B has n_embd = 896, 24 layers (docs/QWEN2-SUPPORT.md:79 (§4 "Verified models")), and n_vocab = 151936 (walkthrough doc 09 §2.4: the 0.5B lm_head costs 30 × 896 × 151936).

Embed gather, 30-token prefill (embed_rows_q4_0): nb = 896 / 32 = 28 blocks per row, so total = nt × nb = 30 × 28 = 840 threads → grid = ceil(840 / 256) = 4 blocks → 1,024 threads launched, 840 doing work, 184 exit at the guard. Eight hundred threads to embed a whole prompt. The kernel is not bandwidth-limited here — it is launch-bound: at the campaign's "2 µs/graph-gap scale" per launch (docs/CUDA-TECH-PRIMER.md §6.4), a launch of this size costs more than its memory traffic (840 × (18 B read + 128 B written) ≈ 123 KB). This is why fused epilogues, not gather micro-optimizations, dominate decode-step latency — the story chapters 05–06 continue.

Full-table dequant, 0.5B tok_embd (dequant_q4_0_f16): total = od × nb = 151,936 × 28 = 4,254,208 threads → grid = ceil(4,254,208 / 256) = 16,618 blocks. Now the launch is fully saturated and the kernel is a pure streaming pass: 18 B read + 64 B written per thread. Watch the write amplification: each thread writes 32 f16 (64 B) from 18 source bytes — the ratio is 3.56, exactly the storage expansion of §4.1 seen per thread.

Residual add, 30-token prefill (add_f32 over [30][896]): n = 26,880 → grid = ceil(26,880 / 256) = 105 blocks. Bytes moved: read 2 × 4 B + write 4 B = 12 B per element for 2 flops — an arithmetic intensity of ~0.17 FLOP/byte. No amount of compute throughput rescues that ratio; the kernel is memory-bound by construction, and the only levers are moving fewer bytes (fusion — do not write z and re-read it) and coalescing (already free here: thread tid reads elements tid, fully contiguous).

4.3 What would make them slow

  • A wrong grid contract. Pass element count instead of row count to add_bias_f32 and you launch nt × d block-rows — out-of-bounds writes, not a slowdown but a crash or silent corruption (the call-site comment, src/graph/cuda_backend.rs:1510-1512).
  • Element-per-thread dequant. One thread per element would re-read the scale byte and re-enter the nibble 32× more often per output; the block-per-thread shape exists to amortize the scale read and emit contiguous 64 B runs.
  • Uncoalesced nibble loads. The 18-byte block reads are not full 128 B transactions, but consecutive threads read consecutive 18-byte blocks, so the hardware coalescer still streams them densely; an interleaved or transposed block order would destroy that.
  • Divergence. All three kernels guard-and-exit once, at a boundary aligned across whole warps' worth of threads (the tail block only). There is no data-dependent branching inside the loop, so the SIMT execution stays lockstep — the divergence term from chapter 01 never comes into play here.

5. Try it / Observe

Three commands, all from the repo root (build details and the ccbin/arch pitfalls: docs/BUILD.md:

# 1. build with the CUDA backend (needs nvcc on PATH or /usr/local/cuda/bin)
cargo build --release --features cuda

# 2. run one prompt through the graph on the GPU
./target/release/minfer <model.gguf> "hi"

# 3. same run, but force the prefill down the f16 dequant path of §3.2
#    (default prefill is int8 MMQ; MINFER_MMQ=0 escapes to the f16 wmma GEMM)
MINFER_MMQ=0 ./target/release/minfer <model.gguf> "hi"

To see the kernels, record a trace and open the visualizer (viz/README.md is the one-line reference: the viz page replays a recorded graph with per-node real tensor statistics — min/max/mean + a value heatmap — and the token/logit distributions per decode step):

MINFER_TRACE=/tmp/t.json ./target/release/minfer <model.gguf> "hi"   # record
cd viz && python3 -m http.server 8080     # open http://localhost:8080, load /tmp/t.json

What to look for in the trace: the GetRows node that is the first real op (embed §3.3), the residual Add nodes between every sub-block (§3.1), and — under MINFER_MMQ=0 — the prefill MatMuls that route through the dequantize-then-GEMM pair instead of MMQ. MINFER_NO_CUDA_GRAPH=1 reverts CUDA Graph replay to per-kernel launches if you want launch-level visibility.

6. Cross-references

← 02 · The minimal CUDA you actually need · Index · 04 · Reading minfer's kernels II →

04 · Reading minfer's kernels II — from GEMV to tiled GEMM

Part: Part 3b — the matmul ladder. Prereq: chapter 03 (the dispatch chain, quant block layout, index formulas — this chapter uses all three). Code: src/cuda/kernels/*.cu, src/graph/cuda_backend.rs — all file:line citations verified against the tree at writing time (the function name is the stable address, the line number a convenience).

1. Background — where this sits

Chapter 03 taught you the reading method on the easiest kernels in the file. This chapter reads the family that earns the GPU its keep — the matmuls — the heart of the tutorial for a simple reason: in a transformer forward pass, the matmuls are where nearly all the floating-point work and nearly all the weight traffic happen. Every other kernel exists to feed them or clean up after them.

The chapter is a ladder with five rungs, in the order the code itself climbed during the optimization campaign:

  1. a scalar GEMV (General Matrix-Vector multiply — one output = one dot product of a weight row with the activation vector);
  2. a vectorized GEMV, same math, 16-byte loads;
  3. the quantized decode path — MMVQ (matrix-vector quantized), integer dots over packed 4-bit weights, which is what a real decode step runs;
  4. the tiled f16 GEMM (General Matrix-Multiply) for prefill — tiling in three layers, __syncthreads() between them, wmma tensor-core fragments inside;
  5. a pointer to the int8 MMQ (matrix-matrix quantized) prefill GEMM — the current default — whose anatomy lives in the reference docs.

One word before we start, because it organizes everything: nt is the number of token rows the matmul processes — 1 during decode, hundreds during prefill. Almost every dispatch decision you are about to see is a function of nt, and §4 turns that into arithmetic: at nt == 1 the matmul degenerates into a GEMV that is memory-bound no matter how you code it; at large nt each weight byte is reused so many times that the same silicon becomes compute-bound. The ladder exists because no single kernel wins at both ends.

2. Principle — the concepts

2.1 The matmul corner of the kernel inventory

The forensics protocol starts with enumeration. Chapter 03's count is 99 __global__ void definitions across src/cuda/kernels/*.cu — 17 translation units, 10,147 lines today (the pre-#263 single TU was 10,215; re-run the grep to confirm before citing):

grep -n '__global__' src/cuda/kernels/*.cu

Of those 99, 47 are matmul-family kernels (names matching matmul/mmvq/gemm/mmq). You navigate them by regime:

FamilyExamples (file, first hit line)What it doesWhere taught
f32-activation matvecf32_f32_matmul_vec (ops_misc.cu:314), f32_f32_matmul_scalar (:370), q4_0_f32_matmul (matmul_f32act.cu:103), q6_k_f32_matmul_padded (ops_misc.cu:16)dot products against f32 activations, per weight typethis chapter (§3.1–3.2)
Decode MMVQ (int8 dots)q4_k_q8_mmvq (mmvq_skipwrite.cu:201), q6_k_q8_mmvq (:268), q4_0_q8_mmvq (mmvq_multi.cu:446), q8_0_p32_q8_mmvq (:655) (+ _v2/_multi/_pf variants)one weight row per 256-thread block, __dp4a over quantized activationsthis chapter (§3.3–3.4)
Prefill GEMM (f16)gemm_f16_nt_kernel_t (gemm_wmma.cu:367), gemm_qb_nt_kernel (gemm_fused_dequant.cu:218)tiled tensor-core GEMM over dequantized f16 weightsthis chapter (§3.5)
Prefill MMQ (int8)mmq_nt_kernel (mmq_int8.cu:236), mmq_raw_nb_kernel (mmq_nb.cu:9), mmq_raw_nb_bt_kernel (:289), mmq_raw_nb_bt_q6k_kernel (mmq_bt_q6k.cu:42), mmq_ksplit_reduce_kernel (mmq_nb.cu:589)tiled int8 tensor-core GEMM, raw weight bytesthis chapter (§3.6 — pointer only)
Activation quantizequantize_q8_0_pad40 (mmvq_aquant.cu:27), quantize_q8_0_pad40_t (:85), quantize_q8_0 (ops_elementwise.cu:12)f32 activations → 40-byte q8 blocks for the quantized pathsthis chapter (§3.3–3.6)

Three things the table does not show:

  1. The nt regimes share one dispatch function. Rust-side, one place decides GEMV-vs-GEMM: CudaState::matmul_f32_ptr_layout (src/cuda/methods/dispatch.rs:162). Every MatMul-shaped node goes through it — the plain Op::MatMul (src/graph/cuda_backend.rs:1454) arm and the decode-fused Op::FusedQKV concat matmul (matmul_f32_ptr_layout, cuda_backend.rs:1399).
  2. The v2/multi/pf suffixes are variants, not new algorithms. q4_k_q8_mmvq_v2 (mmvq_skipwrite.cu:376) is the same dp4a structure reorganized for 16-byte weight loads; _multi (mmvq_multi.cu:25) adds an in-block token loop for nt 2–8; _pf (mmvq_q6k.cu:65) is q6_K's pipelined form for tall rows. Read one, you have read the family's skeleton.
  3. Two quantize kernels serve two layouts. quantize_q8_0_pad40 (mmvq_aquant.cu:27) writes token-major q8 blocks (the GEMV/MMVQ layout), while quantize_q8_0_pad40_t (mmvq_aquant.cu:85) writes the transposed, swizzled layout the MMQ GEMM stages (§3.6). Same math — max, scale, round — different destination addresses.

2.2 GEMV: the shape decode asks for

During decode, nt == 1: the activation side of the matmul is a single vector y[id], and the op is out[r] = Σ_i W[r][i] · y[i] for r = 0 .. od-1, with the weight matrix in minfer's memory convention — metadata [in, out], memory row-major [out][in] (AGENTS rule 4) — so output r reads exactly weight row r, which is id contiguous values. The two layout facts from chapter 03 are load-bearing here: activations are token-major [nt][id], so at nt == 1 the vector y is one contiguous run that every thread reads (cached, DRAM pays once); and weight rows are contiguous, so a thread or block assigned one output row streams memory linearly — TECH-PRIMER §5.2's coalescing story applies for free.

The thread-mapping question ("what does one thread do?") has two natural answers for a GEMV, and minfer ships both:

  • one thread per output element — the thread owns out[r] and loops the dot product serially. Simple, and the weight reads are perfectly coalesced across threads (consecutive r = consecutive rows), but each thread does id dependent multiply-adds with no help from its neighbors.
  • one block per output row — 256 threads split the row's id elements, each computing a partial sum, then a two-stage reduction (warp shuffle, then shared memory across warps) produces the output. Shorter serial chains — and the real prize: it is the shape the quantized decode kernels need, because one 256-thread block can also own the decoding of a whole row's packed blocks (§3.4).

2.3 Why GEMM wants tiles: the reuse arithmetic

At prefill, nt is the prompt length (say 512). The GEMM is

C[nt][od] = A[nt][id] · B[od][id]ᵀ

and the naive loop order ("for each output element, dot one row with one column") re-reads the same weight bytes once per token: nt × the whole matrix. The fix is tiling — process the output in small rectangles so every tile of B loaded once serves all the tokens in the tile of A. Count it for the 0.5B down-projection [od=896, id=4864] at nt = 512 (dims: docs/inference_e2e_walkthrough/05-graph-builder-ir.md:162 ("§2.5 The main event: one forward pass, node by node")):

  • without reuse: each weight byte is read nt times → 2.45 MB (Q4_0) × 512 ≈ 1.25 GB of traffic for one layer, one forward.
  • with a 64-token tile: each weight byte is read nt / 64 = 8 times (once per token-tile) → ≈ 19.6 MB.

That ratio — traffic divided by ceil(nt / tile) — is the entire economic argument for the prefill GEMM's complexity, and why fa_prefill_kv (src/cuda/kernels/attention_prefill.cu:117) is shaped the same way: 64 query tokens share one K/V stream.

2.4 Tiling in three layers

"Tile" appears at three scales in a high-performance GEMM; name them now, because chapter 06's technique catalog assumes you can:

  1. Block tile — which rectangle of C one block owns. Set by the grid: blockIdx.x picks the token-tile, blockIdx.y the output-tile.
  2. Shared-memory staging — the block cooperatively copies its A-tile and B-tile from global memory into shared memory (the on-chip scratchpad from chapter 01), because every thread will re-read those tiles many times. __syncthreads() — the barrier making every thread wait until all copies are done — separates "staging" from "compute".
  3. Register tile — each thread accumulates a small sub-rectangle of the block tile in registers (or in wmma fragments — the tensor-core register layout, defined in chapter 05 §3.1). Registers are the only place an accumulator can live without paying memory traffic.
C[nt][od]                       one BLOCK TILE = 64 tokens × TM outputs
 ┌────────────┬────────────┐    ┌──────────────────────────┐
 │ blk (0,0)  │ blk (0,1)  │    │  shared memory:          │
 ├────────────┼────────────┤    │   As = A-tile  64×KS f16 │  ← staged once,
 │ blk (1,0)  │ blk (1,1)  │    │   Bs = B-tile  TM×KS f16 │     reused by all
 └────────────┴────────────┘    │  registers:              │  64·TM threads
   blockIdx.x  blockIdx.y       │   fc[j][oc] fragments    │  ← accumulated
                                └──────────────────────────┘     per thread

The k-dimension (id) is not tiled in the grid — it is walked in chunks of KS inside the block, double-buffered (stage the next chunk while the current one computes). That loop nest is what §3.5 reads.

2.5 The quantized families: MMVQ vs MMQ

Both quantized families follow the same algebra — walkthrough doc 10 §2.4 if the pairing is not reflexive yet — but they sit at opposite ends of the nt axis:

  • MMVQ (decode, nt == 1): quantize the single activation row to int8 once (a 40-byte-per-block "pad40" layout), then each weight block's packed nibbles form an integer dot with the int8 activation values via __dp4a — the SIMT instruction computing a 4-way 8-bit integer dot in one op; scales fold in at the end. One block per weight row, streaming the row once. Nothing is reusable at nt == 1, so the kernel's only job is converting bandwidth into outputs efficiently.
  • MMQ (prefill, nt ≥ 9): the same int8 idea arranged as a tiled GEMM on the int8 tensor cores — mma.m16n8k32.s8 multiplies 16×32 by 32×8 int8 tiles in hardware. Weights are staged raw (never dequantized); activations quantized once per call in a prepass.

The reason decode does not want the tensor-core path: mma.m16n8k32 computes a 16-row output tile, so at nt == 1 15 of those 16 rows are padding — the instruction's throughput is wasted, and the binding resource is weight bytes anyway (walkthrough doc 15 §2.3). Prefill does not want MMVQ because of the §2.3 table: one block per row re-reads every weight byte nt times.

3. In minfer's code

3.1 f32_f32_matmul_scalar — the scalar GEMV

The plainest matmul in the file — the shape every later kernel improves on. The kernel, f32_f32_matmul_scalar (src/cuda/kernels/ops_misc.cu:370):

// General-case fallback: one thread per (token, output) pair, scalar dot.
__global__ void f32_f32_matmul_scalar(
    const float* __restrict__ weights,
    const float* __restrict__ acts,
    float* __restrict__ output,
    int od, int id, int nt
) {
    long long idx = (long long)blockIdx.x * blockDim.x + threadIdx.x;
    if (idx >= (long long)nt * od) return;
    int t = (int)(idx / od), r = (int)(idx % od);
    const float* wr = weights + (size_t)r * id;
    const float* y = acts + (size_t)t * id;
    float acc = 0.0f;
    for (int i = 0; i < id; i++) acc += wr[i] * y[i];
    output[idx] = acc;
}

Line by line:

  • idx in long long, and the ceil-div guard — the flat thread index over nt * od outputs, widened defensively (chapter 03's overflow habit); if (idx >= nt * od) return; is the usual grid-tail guard.
  • t = idx / od, r = idx % od — decode the flat index back into (token, output). One thread owns one output element: the first of §2.2's two mappings.
  • wr = weights + r * id — the weight row. Row-major [out][in] means output r's inputs are weights[r*id .. r*id+id], contiguous. Two adjacent threads (r, r+1) therefore read two adjacent rows — coalesced at the warp level.
  • y = acts + t * id — the token's activation row. At nt == 1 every thread reads the same y bytes; they hit L1/L2 and cost DRAM once.
  • the loop, and the store — a serial scalar dot: id dependent multiply-adds, one accumulator, no unrolling, no vector loads. This is the baseline the rest of the ladder exists to beat. The store address is idx: od outputs per token are contiguous, so flat index = address.

The launcher launch_f32_f32_matmul (src/cuda/kernels/ops_misc.cu:799) is where the ladder's first fork lives — id % 8 == 0 (the vector kernel's float4 alignment requirement) routes to f32_f32_matmul_vec with grid = ceil(od/8), 64-thread blocks; otherwise the scalar kernel launches with the familiar grid = ceil(nt*od/256) ceil-div. This fork is the whole F32 dispatch — no quant gates, no MMVQ — and the scalar kernel is its general-case fallback. For the quantized types, each has a sibling with the same GEMV shape (q4_0_f32_matmul (matmul_f32act.cu:103) reads packed nibbles but keeps f32 activations; q6_k_f32_matmul_padded (ops_misc.cu:16) is the padded-stride variant), so the shape of this kernel is the GEMV shape for the whole f32-activation family.

Bytes moved — the number that decides everything. Qwen2.5-0.5B has n_embd = 896 and FFN width 4864 (docs/QWEN2-SUPPORT.md:79 ("§4 Verified models"); docs/inference_e2e_walkthrough/05-graph-builder-ir.md:162 ("§2.5 The main event: one forward pass, node by node"). Two layers, f32 weights, one decode token:

  • attention wo [896 out, 896 in]: weights = 896·896·4 B = 3.21 MB; activations 3.6 KB; outputs 3.6 KB. FLOPs = 2·896·896 ≈ 1.61 MFLOP.
  • ffn_down [896 out, 4864 in]: weights = 896·4864·4 B = 17.4 MB; FLOPs = 2·896·4864 ≈ 8.72 MFLOP.

At the GB10's documented ~273 GB/s (docs/GLOSSARY.md:127 ("L3 — Performance model")), ffn_down cannot finish faster than 17.4 MB ÷ 273 GB/s ≈ 63.7 µs — and the math it must do (8.7 MFLOP) is ~136 GFLOP/s at that pace, nothing for a GPU. That asymmetry is the whole story: a GEMV's arithmetic intensity — FLOPs per byte moved — is fixed by the shape, not by how cleverly you code it, and at 2 FLOP per 4-byte element it is 0.5 FLOP/byte. You cannot optimize a scalar GEMV into being compute-bound; you can only (a) move fewer bytes — quantize the weights (§3.4) — or (b) reuse bytes — add tokens (§3.5).

CPU counterpart. Walkthrough doc 10 (10-cpu-matmul-kernels.md) is the CPU mirror of this whole ladder: the dequant-at-load bandwidth argument (§2.1), why the activations go int8 too (§2.2), and the AVX2/NEON implementations (§3). The GPU kernel above is the same scalar reference executed once per thread instead of once per core — and doc 10's §2.1 is the same ledger §4 keeps with device numbers: weights dominate, so their bytes are the budget.

3.2 f32_f32_matmul_vec — the vectorized GEMV

Same math, three mechanical changes. The kernel header states the mapping (src/cuda/kernels/ops_misc.cu:310-313, the f32_f32_matmul_vec header comment):

// ─── F32 × F32 matmul (7e④) ───────────────────────────────────
// Same unit lane mapping as the q4_K kernel: lanes own (row, 256-elem
// chunk) pairs, float4 loads on both operands. Requires id % 8 == 0 for
// the aligned float4 loads; the scalar kernel covers the general case.

First change — the constants and the mapping (ops_misc.cu:320-336, the f32_f32_matmul_vec constants):

    const int NR0 = 4;
    const int NSG = 2;
    const int CHK = 256;

    int warp_id = threadIdx.x / WARP;
    int lane_id = threadIdx.x % WARP;
    int r0 = (blockIdx.x * NSG + warp_id) * NR0;
    if (r0 >= od) return;

    int nch = (id + CHK - 1) / CHK;
    // Step 82: the token dimension lives in this in-block loop, not in the
    // launch grid (grid.y used to be nt = one full weight re-stream per
    // token). The weight bytes for the block's rows are re-read across
    // tokens from L1, so DRAM sees one weight stream per block; nt==1
    // keeps the exact single-token op order (bitwise).
    for (int t = 0; t < nt; ++t) {
        const float* y = acts + (size_t)t * id;

One 64-thread block (2 warps — the launcher's block(64)) owns NSG × NR0 = 8 output rows: warp w takes rows r0 .. r0+3, and each lane (one of the warp's 32 threads) will own a (row, 256-element chunk) pair. The grid is ceil(od / 8) blocks — output-parallel only, no grid.y = nt.

Second change — the inner loop is a unit-lane loop with float4 loads (ops_misc.cu:342-358, the f32_f32_matmul_vec float4 unit loop):

        for (int u = lane_id; u < nch * NR0; u += WARP) {
            int ic = u % nch, rr = u / nch;
            const float* wr = weights + (size_t)(r0 + rr) * id + ic * CHK;
            const float* yc = y + ic * CHK;
            int len = min(CHK, id - ic * CHK);
            float p = 0.0f;
            // the unit's lane streams the WHOLE chunk (8 floats per pass)
            for (int i = 0; i < len; i += 8) {
                float4 a0 = *reinterpret_cast<const float4*>(wr + i);
                float4 a1 = *reinterpret_cast<const float4*>(wr + i + 4);
                float4 b0 = *reinterpret_cast<const float4*>(yc + i);
                float4 b1 = *reinterpret_cast<const float4*>(yc + i + 4);
                p += a0.x * b0.x + a0.y * b0.y + a0.z * b0.z + a0.w * b0.w
                   + a1.x * b1.x + a1.y * b1.y + a1.z * b1.z + a1.w * b1.w;
            }
            acc[rr] += p;
        }
  • u = lane_id; u += WARP — the round-robin unit loop from §2.2's second mapping, one level down: a warp owns a row, its lanes split the row's 256-element chunks. For id = 4864: nch = 19 chunks, 4 rows per warp → 76 units over 32 lanes ≈ 2.4 passes.
  • float4 — a 16-byte vector type; reinterpret_cast<const float4*> loads four consecutive f32 in one 16-byte transaction instead of four 4-byte ones — hence the launcher's id % 8 == 0 gate: each pass consumes two float4 pairs (8 elements), so misaligned id would fault.
  • acc[rr] += p — each lane keeps one partial per row (float acc[NR0], zeroed at ops_misc.cu:340); chunk partials accumulate per lane, not globally.

Third change — the reduction and the token loop (ops_misc.cu:335-365, the f32_f32_matmul_vec token loop and warp reduction): per row, warp_reduce_sum — the butterfly shuffle from chapter 02 — folds the 32 lane-partials into lane 0, which stores output[t * od + r0 + rr]. The for t loop wraps everything, so one launch serves every token; the comment calls out the bitwise constraint (the nt == 1 order is exactly preserved) and the reason it exists — the pre-Step-82 shape (grid.y = nt) re-streamed the whole weight matrix once per token (campaign record: docs/CUDA_OPTIMIZATION.md §0 row 82).

Why 16-byte transactions matter (and why it is still slow). DRAM talks in bursts, not bytes: a 4-byte load still costs a 32-byte sector, and a warp's scattered 4-byte loads can cost up to 8× the useful traffic. The float4 load aligns each lane's request to 16 bytes, so a warp's 32 lanes touch 32 × 16 = 512 contiguous bytes — fully dense, minimum transactions (TECH-PRIMER §5.2's coalescing rule, at the widest general-purpose width). But notice what did not change: the weight bytes per output, and therefore the 0.5 FLOP/byte intensity. Vectorization makes the memory system run at its best; it cannot change the shape's budget — on Qwen2.5-0.5B in f32, ffn_down decode is still a ≥ 63.7 µs layer. The byte count only drops when the weights themselves shrink — rung 3.

CPU counterpart. Doc 10 §3 does the same promotion in Rust: scalar loop → #[target_feature(enable = "avx2")] chunks with _mm256_fmadd_ps — float4 ↔ 256-bit SIMD vectors, the warp reduction ↔ a horizontal add. Same gain class, same ceiling.

3.3 The dispatch — one function decides the whole ladder

Here is the section to bookmark: every MatMul-shaped node reaches the same match arm, and the arm hands everything to one function. The Op::MatMul arm (src/graph/cuda_backend.rs:1454), abridged to the calls that matter:

#![allow(unused)]
fn main() {
Op::MatMul { transpose_b } => {
    if *transpose_b { return Err(...); }           // weights arrive row-major [out][in]
    ...
    // quant kernels address whole 32-element blocks … (id % 32 == 0 gate)
    if meta.weight_ttype != crate::tensor::TensorType::F32 && id % 32 != 0 {
        return Err(...);                           // GPU_SAFETY: Err, not fallback
    }
    ...
    self.state.matmul_f32_ptr_layout(
        wptr, meta.weight_ttype, self.ptr_of_ref(in_bufs[0])?,
        self.ptr_of_ref(out_buf)?, od, id, nt,
        self.state.is_weight_padded(&meta.weight_name),
    )?;
    if let Some(bname) = &meta.bias_name { ... }   // add_bias_f32, ch-03's epilogue
    Ok(())
}
}

The name parses as: matmul with f32 activations (_f32), raw pointers (_ptr — no tensor objects cross the FFI), and a layout flag (_layout — whether a Q6_K weight was registered with the padded 224-byte stride). The decision tree inside matmul_f32_ptr_layout (src/cuda/methods/dispatch.rs:162) has three tiers:

Tier 1 — prefill GEMM (src/cuda/methods/dispatch.rs:191-218, the prefill-GEMM gate): nt >= 9 (or nt >= 2 when the off-by-default MINFER_SMALL_M_GEMM=1 experiment routes nt 2–8 here, doc 91) and id % 32 == 0 and a supported quant type → a tiled GEMM — prefill_mmq (int8, default) or prefill_gemm_f16 (the f16 escape, §3.5) depending on mmq_active() (src/cuda/methods/policy.rs:39: the resolved device tier's MMQ flag — a table flag, or cc >= 800 — and MINFER_MMQ not 0).

Tier 2 — decode/small-batch per-type kernels: everything else falls through to a match ttype with per-type shape gates. The Q4_0 arm (src/cuda/methods/dispatch.rs:273-284, the Q4_0 arm's decode gates):

#![allow(unused)]
fn main() {
} else if nt == 1 && id >= 2048 && id % 32 == 0 && !Self::no_q40_mmvq() {
    // doc 103: decode MMVQ … claims shapes that previously ran the f32 kernel
    self.q4_0_decode_mmvq(wptr, x, out, od, id, nt);
    Ok(())
} else if nt >= 2 && nt <= 8 && id > 8192 && id % 32 == 0 && !Self::no_q40_mmvq() {
    self.q4_0_decode_mmvq_multi(wptr, x, out, od, id, nt);
    Ok(())
} else {
    launch!(launch_q4_0_f32_matmul)
}
}

and the K-quant arms have the same skeleton with their own measured gates (src/cuda/methods/dispatch.rs:323-403, the K-quant arms' measured gates): Q4_K decode MMVQ at id >= 2048, Q5_K at od*id >= 24_000_000, Q6_K at od*id >= 4_000_000 — each a documented crossover where the dp4a structure starts beating the f32-activation kernel, each with an env opt-out for A/B. Tier 3 — the F32 fallback: launch_f32_f32_matmul → vec or scalar by the id % 8 fork of §3.1.

The one-line answer to "which kernel does decode's Matmul dispatch to?":

  • decode (nt == 1): Op::MatMul (cuda_backend.rs:1454) → matmul_f32_ptr_layout (src/cuda/methods/dispatch.rs:162) → per-type MMVQ — for Q4_0 with id ≥ 2048: q4_0_decode_mmvq (src/cuda/methods/mmvq.rs:224) → launch_q4_0_q8_mmvq (mmvq_multi.cu:603) → q4_0_q8_mmvq (mmvq_multi.cu:446), after decode_quantize_native (src/cuda/methods.rs:180) has produced (or memoized, the MmqCache) the pad40 q8 activation plane via quantize_q8_0_pad40.
  • prefill (nt ≥ 9): the same arm → mmq_active() → prefill_mmq (src/cuda/methods/prefill_mmq.rs:133) → for Q4_K: transposed-A prepass quantize_q8_0_pad40_t (mmvq_aquant.cu:85) then launch_mmq_raw_nb_bt_nt (mmq_nb.cu:601) → mmq_raw_nb_bt_kernel (mmq_nb.cu:289); Q6_K has its own BT kernel mmq_raw_nb_bt_q6k_kernel (mmq_bt_q6k.cu:42); fallback shapes land on mmq_nt_kernel (mmq_int8.cu:236).

The walkthrough's "MMVQ decode", verified honestly. The e2e walkthrough's master table (doc 15 §2.2) says decode runs "per-type MMVQ + f32-activation kernels" — exactly what the code says, but the shape gates decide which matmuls take MMVQ on your model. On Qwen2.5-0.5B (id = 896 for every attention projection, gate and up; id = 4864 only for ffn_down): ffn_down (id 4864 ≥ 2048) → MMVQ (q4_0_q8_mmvq on a Q4_0 model); everything else — QKV projections, wo, gate, up, lm_head (id 896) — misses the gate and runs the f32-activation kernel q4_0_f32_matmul (matmul_f32act.cu:103). On Qwen2.5-7B (id = 3584 everywhere), every decode matmul clears the gate and the whole step is MMVQ — plus the v2 variants, since mmvq_v2(id) (src/cuda/methods/mmvq.rs:725, the mmvq_v2 gate) additionally requires id % 256 == 0 (3584 = 256·14 ✓). The gate is not an oversight: the arms' comments record the measured crossovers (small-id MMVQ loses — the uncoalesced nibble loads dominate when rows are short; the Q5_K arm prices it at src/cuda/methods/dispatch.rs:342-357). The reading habit this tutorial keeps hammering: the master table gives the structure; the gates give your model's truth.

One more dispatch consumer: the decode fused path. Chapter 05 documented Op::FusedQKV (cuda_backend.rs:1365) as concat-matmul then a fused bias+rope+store (attn_bias_rope_store_f32, kv_store.cu:118); the concat matmul inside it is the same matmul_f32_ptr_layout (cuda_backend.rs:1399) call, so the fused node and the plain Op::MatMul node make identical kernel choices at identical shapes — the fusion is in the epilogue, not the matvec.

3.4 q4_0_q8_mmvq — the real decode path

Now the kernel your 0.5B Q4_0 ffn_down actually runs at nt == 1. First its ingredients, then the code.

The activation plane. decode_quantize_native (src/cuda/methods.rs:180) quantizes the one f32 activation row into the pad40 layout — 40 bytes per 32-element block: 2-byte f16 scale, 2 bytes of padding, 32 int8 values at offset 4, and a 4-byte int32 sum at offset 36. The layout comment (mmvq_aquant.cu:23-26) says the sum feeds the MMQ min-term correction and is "invisible" to MMVQ. The writer kernel is quantize_q8_0_pad40 (mmvq_aquant.cu:27), one thread per block, tree-reduced amax — chapter 03's quantize family, already read.

The kernel — q4_0_q8_mmvq (src/cuda/kernels/mmvq_multi.cu:446):

__global__ void __launch_bounds__(256) q4_0_q8_mmvq(
    const uint8_t* __restrict__ weights,
    const uint8_t* __restrict__ acts8,
    float* __restrict__ output,
    int od, int id, int nt
) {
    const int row = blockIdx.x;
    const int t = blockIdx.y;
    const int nb = id >> 5; // dispatch gate: id % 32 == 0
    const int row_stride = nb * Q4B;
    const uint8_t* x8row = acts8 + (size_t)t * nb * Q8PB;

    float acc = 0.0f;
    for (int u = threadIdx.x; u < nb; u += 256) {
        const uint8_t* blk = weights + (size_t)row * row_stride + (size_t)u * Q4B;
        const float d4 = h2f(*reinterpret_cast<const uint16_t*>(blk));
        const uint8_t* x8b = x8row + (size_t)u * Q8PB;
        const float d8 = h2f(*reinterpret_cast<const uint16_t*>(x8b));
        const uint32_t* xw = reinterpret_cast<const uint32_t*>(x8b + 4);
        int dot = 0, sx = 0;
        #pragma unroll
        for (int v = 0; v < 4; v++) {
            // 18-B stride: payload is 2B-aligned only — two u16 halves/word
            const uint32_t w =
                (uint32_t)*reinterpret_cast<const uint16_t*>(blk + 2 + 4 * v) |
                ((uint32_t)*reinterpret_cast<const uint16_t*>(blk + 2 + 4 * v + 2) << 16);
            const uint32_t lo = w & 0x0F0F0F0F;           // elements 4v..4v+3
            const uint32_t hi = (w >> 4) & 0x0F0F0F0F;    // elements 16+4v..
            dot = __dp4a((int)lo, (int)xw[v], dot);
            dot = __dp4a((int)hi, (int)xw[v + 4], dot);
            sx  = __dp4a(0x01010101, (int)xw[v], sx);
            sx  = __dp4a(0x01010101, (int)xw[v + 4], sx);
        }
        // q4_0 value = (nibble - 8) * d  →  Σ = d * (dot - 8 * sx)
        acc += d8 * d4 * (float)(dot - 8 * sx);
    }
    mmvq_block_reduce(acc, output, od, t);
}

What one thread processes: a slice of one weight row — 32-element blocks, round-robin. The launcher launch_q4_0_q8_mmvq (mmvq_multi.cu:603) is grid(od, nt) × 256 threads, so block row owns output element out[t][row] and its 256 threads split the row's nb = id/32 quant blocks (u = threadIdx.x; u += 256). For ffn_down 0.5B: nb = 152, so each thread handles exactly one block and 104 threads idle — a tail you accept because the structure is per-row on purpose. Line by line:

  • row_stride = nb * Q4B — Q4B is 18 (common.cuh:42), chapter 03's Q4_0 block. One thread's block pointer is row·2736 + u·18 — a stride-18 walk, the row streamed linearly.
  • d4, d8 — the two f16 scales: the weight block's and the activation block's (40-byte stride, Q8PB (common.cuh:133)).
  • the 2-byte-loads-as-u32 trick — Q4_0's 18-byte stride guarantees only 2-byte alignment (family comment at mmvq_multi.cu:440-441), so a direct uint32_t load would be a misaligned-access fault on some devices; the kernel assembles each 32-bit word from two uint16_t loads. "It compiles" is not "it is defined".
  • lo / hi masks — the chapter-03 nibble convention as SIMD: one u32 word holds four packed nibbles; & 0x0F0F0F0F extracts elements 4v..4v+3 (low nibbles), >> 4 the elements 16+4v.. (high nibbles). Two __dp4as per word compute 4-way int8 dots against the matching int8 activation words — eight multiply-adds in four instructions, in integer silicon.
  • sx — the sum of the activation int8 values (dp4a of 0x01010101 × x accumulates x's bytes). Why: the stored nibble is round(v/d) + 8, so the true dot is Σ (nibble−8)·x = dot − 8·sx — the −8 correction for free, the same algebra as the CPU's dot_q4_0_q8_0 (walkthrough doc 10 §2.4).
  • acc += d8 * d4 * (dot - 8*sx) — both scales fold once per block, in f32. The integer accumulation is exact; the only rounding in the row is the two quantizations, which already happened.
  • mmvq_block_reduce (src/cuda/kernels/common.cuh:193) — the two-stage reduction of §2.2's second mapping: 5 warp shuffles, warp_sums[8] in shared memory, __syncthreads(), thread 0 adds and stores output[t*od + row].

The K-quant sibling. q4_k_q8_mmvq (mmvq_skipwrite.cu:201) has the same skeleton — grid(od, nt), 256 threads, round-robin units, dp4a, shared reduce — with the Q4_K super-block decode inside the unit loop (mmvq_skipwrite.cu:215-240): get_scale_min_k4 unpacking per-sub-block scale/min nibbles (mmvq_skipwrite.cu:221), one dp4a pair per 4 bytes of nibbles (mmvq_skipwrite.cu:235-236), and the two-term correction d8 · (s8·d·dot − m8·dm·sx) (mmvq_skipwrite.cu:239) because Q4_K stores a per-sub-block min. The v2 variant (mmvq_skipwrite.cu:376) reorganizes the same math for 16-byte uint4 weight loads (the R2 "weight-streaming" rework, docs/cuda_optimization_steps/09-r2-mmvq-weight-streaming.md), and the _multi variants (mmvq_multi.cu:25) wrap the token loop in the unit loop — the nt 2–8 regime. The differences are load widths and loop nesting, never dot algebra.

Why int8 dots at all — the intensity arithmetic. Rung 3's payoff on §3.1's budget: same ffn_down layer, now Q4_0. Weight bytes: 896 rows × 152 blocks × 18 B = 2.45 MB (was 17.4 MB — 7.1× less traffic); intensity: 2 FLOP per 0.5625-byte element = 3.56 FLOP/weight-byte (was 0.5; docs/GLOSSARY.md:126 ("L3 — Performance model") rounds this to "~1 MAC per weight-byte → bandwidth-bound"). The floor drops from 63.7 µs to 2.45 MB ÷ 273 GB/s ≈ 9.0 µs — still memory-bound (you never escape that at nt == 1), but 7× lower. And the compute to fill that bandwidth — 8.7 MFLOP in 9 µs — is trivial, which is precisely why integer dp4a silicon is enough.

CPU counterpart. Doc 10 §2.4 states the pairing rule and §3.2 implements it: dot_q4_0_q8_0 consumes Q8_0 activation blocks with the identical (nibble−8) correction and scale folding. The implementations agree because the parity gates (chapter 06) compare them to the f64 reference — note the one intentional divergence: the CPU quantizes activations to Q8_0's 34-byte blocks; the GPU here uses the 40-byte pad40 layout.

3.5 gemm_f16_nt_kernel_t — the prefill GEMM, tiled in three layers

The f16 prefill path is the tutorial's payoff for §2.4's theory: here is where block tiles, shared-memory staging, and register tiles actually live. It is also the escape path today (the int8 MMQ of §3.6 is the default), but it is the right one to read first — smaller, and every idea transfers.

How to get there. MINFER_MMQ=0 routes prefill to prefill_gemm_f16 (src/cuda/methods.rs:224), which obtains the weight as f16 (from the persistent per-weight f16 cache — w16_get, src/cuda/methods/prefill_f16.rs:119, dequantized once by chapter 03's dequant_q*_f16 — or by dequantizing into scratch on this call), converts the f32 activations once (launch_convert_f16), and launches the GEMM (launch_gemm_f16, gemm_wmma.cu:847).

The contract. C[nt, od] = A[nt, id] · B[od, id]ᵀ. The header comment is the design in six lines (gemm_wmma.cu:360-365, the gemm_f16_nt_kernel_t header comment):

// C[nt, od] = A[nt, id] · B[od, id]^T. 64 x TM output tiles (TM = 64
// baseline, 128 halves the B-panel re-reads through L2 and the per-k-step
// barrier count), k-step 32, double-buffered shared staging, 8 warps (each
// owns 32 nt rows x TM/4 od cols as 2 x TM/64 f32 fragment pairs). f32
// accumulation. Tails: nt/od masked at store, k-tail zero-filled (id % 8
// == 0 keeps the uint4 chunk loads aligned).

Layer 1 — block tiles. The launcher launch_gemm_f16 (gemm_wmma.cu:847) builds the grid dim3 grid((nt + 63) / 64, (od + TM_ - 1) / TM_) (gemm_wmma.cu:891), TM_ = 128 default per the MINFER_GEMM_TM selection (gemm_wmma.cu:851-858). Block (bx, by) owns output rows n0 = bx·64 (tokens) × m0 = by·TM (outputs). The comment at gemm_wmma.cu:389-391 records why the token axis is grid.x: consecutive blocks share one B panel (TM weight rows × id), so the L2 serves the weight stream across blocks — the weight matrix streams from DRAM ~once per forward instead of once per token-tile.

Layer 2 — shared-memory staging, double-buffered. The kernel's setup (gemm_wmma.cu:376-382) carves one dynamic shared-memory allocation into As (2 × 64×KS f16 — two buffers), Bs (2 × TM×KS f16), and Cs (a per-warp staging area for the store). The k-loop (gemm_wmma.cu:457-469):

    for (int k = 0; k < id; k += KS, buf ^= 1) {
#if __CUDA_ARCH__ >= 800
        if (k + KS < id) {
            gemm_stage_ab<TM, KS, TN, AF32>(A, B, As, Bs, Am, buf ^ 1, n0,
                                            m0, k + KS, nt, od, id, tid);
            gemm_cp_commit();
            // wait until the CURRENT tile landed (one group may stay in flight)
            gemm_cp_wait1();
        } else {
            gemm_cp_wait0();
        }
        __syncthreads();

Read it as a pipeline, not a loop: at k-step j, the block issues cp.async copies (the asynchronous global→shared copy path, TECH-PRIMER §5.3) for the next tile into buffer buf^1 while it still computes on buffer buf; gemm_cp_wait1 waits until only the current tile's copy group is outstanding, and __syncthreads() makes the whole block's math wait for the whole block's staging. KS = 32 by default (gemm_wmma.cu:859-866 — the KS=64 variant halves barrier count but its 56 KB shared appetite halves resident blocks on GB10, measured −38%; the comment prices it). This is the double-buffered staging of §2.4, verbatim.

Layer 3 — register tiles as wmma fragments. Each warp owns a 32-token × TM-output rectangle (warp w: wm = w >> 1 picks the od chunk, wn = w & 1 the 32-row token half, gemm_wmma.cu:387-388). The compute step (gemm_wmma.cu:534-549):

#pragma unroll
        for (int kh = 0; kh < KHC; kh++) {
            wmma::load_matrix_sync(fa[0], &As[buf * TN * KS + wn * 32 * KS + kh * 32], KS);
            wmma::load_matrix_sync(fa[1], &As[buf * TN * KS + (wn * 32 + 16) * KS + kh * 32], KS);
            wmma::load_matrix_sync(fa[2], &As[buf * TN * KS + wn * 32 * KS + kh * 32 + 16], KS);
            wmma::load_matrix_sync(fa[3], &As[buf * TN * KS + (wn * 32 + 16) * KS + kh * 32 + 16], KS);
#pragma unroll
            for (int oc = 0; oc < ODC; oc++) {
                wmma::load_matrix_sync(fb[0], &Bs[buf * TM * KS + (ob + oc * 16) * KS + kh * 32], KS);
                wmma::load_matrix_sync(fb[1], &Bs[buf * TM * KS + (ob + oc * 16) * KS + kh * 32 + 16], KS);
                wmma::mma_sync(fc[0][oc], fa[0], fb[0], fc[0][oc]);
                wmma::mma_sync(fc[1][oc], fa[1], fb[0], fc[1][oc]);
                wmma::mma_sync(fc[0][oc], fa[2], fb[1], fc[0][oc]);
                wmma::mma_sync(fc[1][oc], fa[3], fb[1], fc[1][oc]);
            }
        }

Chapter 05 defined fragments and mma_sync; here note the shape of the nest: per 32-wide k-slice, four A-fragments (two 16-token rows × two 16-wide k-halves) multiply two B-fragments each, accumulating into fc[j][oc] — registers for the entire k-loop. The trailing comment at gemm_wmma.cu:530-533 is a fossil of a real bug ("the v1 bug: only the first 16 k's were multiplied") — both k-halves must accumulate; fragment indexing bugs do not crash, they silently halve your dot products (the parity gates catch them, chapter 06).

The store. After the k-loop, each warp spills its fragments through Cs (shared) and writes out with bounds masks gemm_f16_nt_kernel_t (gemm_wmma.cu:367): store_matrix_sync lands the 16×16 fragment in shared memory, then lanes copy the 256 values to global C[n * od + m] for in-range (n, m) — how the kernel handles the ragged tail of a 30-token prompt without a second code path. This is also why §4's table has a prefill row that looks nothing like the decode rows: with 64 tokens per tile, each staged B byte is consumed by 64 tokens' worth of fragments before the next k-tile is staged — the reuse arithmetic of §2.3 made silicon.

Bank conflicts, in one paragraph. Shared memory is not one wide port: the hardware splits it into 32 banks (lanes) of 4 bytes each, and a warp's access is fast only when its 32 addresses hit 32 distinct banks; when several lanes address the same bank — a bank conflict — the accesses serialize into as many passes as there are colliding lanes. The classic trigger is a shared row stride that is an exact multiple of the bank count (32 floats = 128 bytes): every row's column 0 lands in bank 0, so a column-wise read across rows collapses to one bank. The standard fix is padding the stride by one bank's width — exactly what chapter 05's attention kernel does with sstr = hd + 8 (attention_prefill.cu:132), whose comment is a worked example worth rereading now that you know the term. This GEMM sidesteps the issue differently: its hot shared reads are wmma:: load_matrix_sync calls, and the fragment-load hardware handles the layout. Background: TECH-PRIMER §5 (the memory-hierarchy table and coalescing rules these bank arguments extend).

CPU counterpart. Doc 10 §3 is the honest mirror: the CPU cannot afford per-thread fragments, so it tiles across cache lines and SIMD registers, and its "shared memory" is L1/L2 with hardware coherence doing what __syncthreads() does here. Layer 2 is where the architectures genuinely diverge.

3.6 The int8 MMQ prefill — a pointer, not a tour

The default prefill path (nt ≥ 9, mmq_active(), MINFER_MMQ unset) is the campaign's flagship: the int8 MMQ GEMM, promoted default-on at r60 after measuring 1.080× vs llama.cpp on 7B Q4_K_M (docs/CUDA_OPTIMIZATION.md:133, the r60 promotion row; P6). Its anatomy is a whole reference doc, and the tutorial's policy is to link, not re-explain — but you should recognize its pieces in a profile:

  • Activation prepass: quantize_q8_0_pad40_t (mmvq_aquant.cu:85) quantizes f32 activations to int8 and writes them pre-transposed and swizzled into the exact layout the GEMM stages (mmq_nb.cu:289, fed by prefill_mmq.rs:133); llama.cpp's quantize_mmq_q8_1 design — "byte-identical … only reordered" (mmvq_aquant.cu:82-84).
  • The GEMM: mmq_raw_nb_bt_kernel (mmq_nb.cu:289) — raw quantized weight bytes staged per tile, decoded in registers next to the mma.m16n8k32.s8 instruction, per-k-block scale rescale, f32 accumulation. The q4_K route enters at launch_mmq_raw_nb_bt_nt (src/cuda/methods/prefill_mmq.rs:53 / mmq_nb.cu:601); Q6_K has its own BT kernel mmq_raw_nb_bt_q6k_kernel (mmq_bt_q6k.cu:42); non-BT-consumable shapes fall back to mmq_nt_kernel (mmq_int8.cu:236).
  • Split-K: when the grid is M-starved (small nt), doc 92's auto ksplit (prefill_mmq.rs:121-131, auto_ksplit) slices the k-range across grid.z and mmq_ksplit_reduce_kernel (mmq_nb.cu:589) adds the partials — the same split-K family as decode attention (chapter 05 §2.4).
template <int KDR, bool DSC>
__global__ void __launch_bounds__(256) mmq_raw_nb_bt_kernel(
    const uint8_t* __restrict__ W, const uint8_t* __restrict__ W_dsc,
    const uint8_t* __restrict__ qa8g,
    const uint8_t* __restrict__ sdag, float* __restrict__ C,
    int nt, int od, int id, int nchunk,
    float* __restrict__ Cpart, int ksplit
) {

Eight lines on purpose: the parameters tell the story (raw weights W, a pre-decoded scale plane W_dsc, the swizzled activation planes qa8g/sdag, a partial-output buffer Cpart for split-K). The deep read — staging, swizzles, fragments, the r34→r60 lever history — lives in docs/LLAMA-CPP-MMQ-ANALYSIS.md (the llama.cpp reference kernel, instruction by instruction) and docs/CUDA_OPTIMIZATION.md (+ its per-step records) for how minfer's grew against it.

4. Performance intuition

4.1 The arithmetic-intensity table — decode vs prefill

Arithmetic intensity (AI) is FLOPs performed per byte of memory traffic (docs/GLOSSARY.md:126 ("L3 — Performance model")). For a matmul it is decided by the shape, and the deciding variable is nt — how many token rows share each weight byte. The numbers below are this chapter's own byte/FLOP arithmetic, for one Qwen2.5-0.5B ffn_down layer [od=896, id=4864] in Q4_0 (dims: docs/QWEN2-SUPPORT.md:79 ("§4 Verified models"), walkthrough 05):

PathKernelWeight bytes movedAI (FLOP / byte)What bounds it
decode, f32 weightsf32_f32_matmul_{scalar,vec}17.4 MB0.5weight bytes, by a landslide
decode, Q4_0 (MMVQ)q4_0_q8_mmvq2.45 MB (+ ~6 KB q8 activations)≈ 3.6still weight bytes — 7.1× lower wall
decode, any type (nt == 1)—one weight stream~1 MAC/weight-byte (docs/GLOSSARY.md:126 ("L3 — Performance model"))bandwidth, always
prefill, nt = 512 (MMQ)mmq_raw_nb_bt_kernel2.45 MB weights + 3.1 MB pad40 activations + 1.8 MB output ≈ 7.4 MB≈ 600math (tensor-core) throughput

The derivation of the last row: FLOPs = 2·512·896·4864 ≈ 4.46 GFLOP; MMQ bytes ≈ 512·152·40 B (activations) + 2.45 MB (weights) + 512·896·4 B (output) ≈ 7.4 MB; 4.46e9 / 7.4e6 ≈ 600. Between the first and last row the intensity swings by three orders of magnitude — and that swing, not any kernel's cleverness, is what the dispatch gate nt >= 9 (src/cuda/methods/dispatch.rs:191) reacts to (or nt >= 2 under the off-by-default MINFER_SMALL_M_GEMM=1 experiment, doc 91).

Two consequences worth internalizing:

  • Decode's wall only moves when the bytes do. At nt == 1 no code change can raise AI; the levers are fewer bytes (quantization: 7.1× here) or fewer launches around the same bytes (chapter 05's fusion + CUDA Graph). Hence decode wins like "MMVQ +74–77% at 7B shapes" (src/cuda/methods/dispatch.rs:324-328, the 8e② record) vs prefill wins like "the whole kernel replaced".
  • Prefill's wall only moves when the math does. At AI ≈ 600 the traffic is amortized; what limits the GEMM is MACs per second — why the MMQ campaign is a story of inner-loop structure, not byte counts.

4.2 The roofline, in one paragraph

The roofline model prices any kernel as time ≥ max(FLOPs / peak-FLOPs, bytes / peak-BW) — the larger term wins (docs/GLOSSARY.md:125 ("L3 — Performance model")). On GB10 the bandwidth term uses the documented ~273 GB/s unified LPDDR5x figure (docs/GLOSSARY.md:127 ("L3 — Performance model"); chapter 01's toy measured ~225–229 GB/s of it, 01-gpu-mental-model.md:227 ("§2.3 The memory hierarchy — where the bytes actually are"), and the glossary's one-line classification is this chapter's summary: GB10 decode is memory-bound, prefill compute-bound (docs/GLOSSARY.md:124 ("L3 — Performance model")). Check it against the table: decode Q4_0 at AI ≈ 3.6 tops out near 273 GB/s × 3.6 ≈ 1 TFLOP/s if every byte were useful — the compute side is far above that, so bytes bind at every quant level. Prefill at AI ≈ 600 would need ~160 TFLOP/s to stay bandwidth-limited — the other side of the crossover — so math binds instead. The budget sentence: at nt == 1 you buy bandwidth; at nt ≥ 9 you buy MACs.

4.3 What would make each rung slow

  • Scalar GEMV, badly gridded. One thread per output is fine — the coalescing is free — but launch it with a 2-D grid over nt (the pre-Step-82 shape, §3.2's comment) and every token re-streams every byte.
  • Vec GEMV, misaligned. Drop the id % 8 gate and the float4 loads fault or split into two transactions at the row tails.
  • MMVQ, below its gate. The uncoalesced 18-byte (2-byte-aligned) weight walks only pay off because dp4a turns them into 8 MACs per load; on short rows (id < 2048, or the 24M/4M-element K-quant floors) the f32 kernels' wide coalesced loads win — that is what the per-arm gates are (src/cuda/methods/dispatch.rs:273-392, the per-type arm gates).
  • Prefill GEMM, mis-tiled. A tile that underfills the machine (TM=64 at huge od, the MINFER_GEMM_TM A/B) or a k-step whose shared appetite halves occupancy (KS=64's −38%, gemm_dynamic_smem_bytes (gemm_wmma.cu:585)) trades the §2.3 reuse away.
  • Anywhere: the silent 15/16. Force the tensor-core GEMM onto nt == 1 and 15 of every 16 mma rows are padding (walkthrough 15 §2.3) — no profiler shows an error, just a tok/s number that never improves.

5. Try it / Observe

Build once (nvcc chain and arch pitfalls: docs/BUILD.md; the GB10's nvcc is not on every shell's PATH):

export PATH=/usr/local/cuda/bin:$PATH
cargo build --release --features cuda

Run the bench — -p 512 prefill tokens (MMQ territory), -n 64 decode tokens (MMVQ territory) — flags per docs/USAGE.md (bench [-p N] [-n N] [-r N] [-o md|csv|json]); any Q4_0/Q4_K_M model works, and minfer auto-downloads given an HF name (docs/USAGE.md:16 ("Usage")):

./target/release/minfer bench -p 512 -n 64 -r 3 hf:Qwen/Qwen2-0.5B-GGUF:qwen2-0.5b-q4_0.gguf

A/B the chapter's decode rung — on a Q4_0 model MINFER_NO_Q40_MMVQ=1 reverts ffn_down (the one matmul that clears the id ≥ 2048 gate, §3.3) from MMVQ to the f32-activation kernel; MINFER_MMVQ_V1=1 swaps the v2 weight-streaming variants for the 8e originals on K-quant models:

MINFER_NO_Q40_MMVQ=1 ./target/release/minfer bench -p 512 -n 64 -r 3 <model.gguf>
MINFER_MMQ=0 ./target/release/minfer bench -p 512 -n 64 -r 3 <model.gguf>   # prefill: f16 GEMM (§3.5) instead of int8 MMQ

What to look for: the first A/B moves decode tok/s (a 7× byte-budget change on one matmul per layer — small but visible; measured decode deltas for this class: docs/CUDA_OPTIMIZATION.md §0's 8e/8e② row and the 103 record in §2 Part V); the second moves prefill pp tok/s (the f16 path streams ~3.6× the weight bytes, §4.1's table). To see which kernels your model actually dispatched, record MINFER_TRACE=/tmp/t.json on a run (trace/viz flow: chapter 03 §5).

6. Cross-references

← 03 · Reading minfer's kernels I · Index · 05 · Reading minfer's kernels III

05 · Reading minfer's kernels III — attention and the host side

Part: Part 3c — attention kernels + the Rust host layer. Prereq: chapters 03–04 (the matmul ladder, tiling, dispatch — and why decode is memory-bound). Code: src/cuda/kernels/attention_prefill.cu (fa_prefill_kv), src/cuda/kernels/kv_store.cu (attn_bias_rope_store_f32), src/graph/cuda_backend.rs (the Backend trait implementation), src/graph/scheduler.rs (splits, cross-backend copies, replay trigger).

1. Background — where this sits

Chapters 03 and 04 read the matmul kernels: how a quantized weight row becomes a dot product, how tiles map onto blocks, and how the prefill GEMM (General Matrix-Multiply) differs from the decode matvec. This chapter reads the hardest device code in the repo — the two attention kernels — and then crosses the language boundary into the Rust layer that decides which kernel launches, with which pointers, in which order.

Two things make attention special compared to the matmuls you have already read:

  • It is not a fixed-shape problem. A matmul's shape comes from the model and the token count. Attention's inner loop length is positions[t] + 1 — the number of keys accumulated so far — and that is data on the GPU, not a host integer. Both kernels in this chapter read the positions array on the device to find their work. This is minfer's graph rule 1 — "KV positions are data, not structure" (AGENTS.md:113 ("Compute Graph — core rules", rule 1)) — doing real work inside a kernel.
  • It has a serial dependency the matmuls do not have. Softmax (the exponentiate-and-normalize that turns scores into weights) needs the largest score of the whole row before any output can be finalized. The online softmax restructures that dependency into a rescaling loop; §2 teaches it from zero with a two-chunk worked example before we touch the real kernel.

The host half of the chapter follows one decode step through the Rust stack: the execute_node dispatch, the device buffer pool, the fused tail that writes K/V (key/value) into the cache, and the CUDA Graph (record-once, replay-many) machinery that removes per-kernel launch overhead. Where a design decision has history — the Phase-3 host-copy bug, the all-weights gate — we cite the record instead of retelling it.

2. Principle — softmax needs the whole row; online softmax pretends it doesn't

2.1 The problem with plain softmax

Attention computes, for every query row t and every key position p ≤ t, a score, then converts each row into probabilities with softmax and uses them to average the value rows:

s[t][p] = (q[t] · k[p]) / sqrt(hd)
a[t][p] = exp(s[t][p]) / Σ_p' exp(s[t][p'])
o[t]    = Σ_p a[t][p] · v[p]

The CPU code can do this literally: materialize the whole score row, take a max, exponentiate, sum, divide (walkthrough 11 does this in scalar Rust). On the GPU that "whole row first" step is the problem: at a 2K-token context a single head row is 2048 scores, and a naive kernel would have to store the whole score matrix somewhere — huge write+read traffic in global memory, or shared memory it does not fit in. We want to process the keys in chunks with a small fixed-size state, the way streaming code reads a huge file. The obstacle: exp(s - max) needs the max over the entire row before the first exponentiate is correct.

2.2 The online softmax: running max, running sum, rescaled output

The fix (from the FlashAttention paper — a technique the CUDA-TECH-PRIMER §6.3 lists for the prefill kernel) is to carry three small state variables per row while walking the keys:

  • m — the largest score seen so far (the running max);
  • l — the sum of exp(s - m) over the keys seen so far (the running sum, always computed against the current max);
  • o — the accumulated weighted value sum Σ exp(s - m) · v, again against the current max.

For each new chunk of scores you compute the new row max m_new, then notice that everything accumulated so far was scaled by exp(· - m_old) while it now needs to be scaled by exp(· - m_new). The correction factor is a single scalar:

α = exp(m_old - m_new)
o ← o · α + Σ_new exp(s - m_new) · v
l ← l · α + Σ_new exp(s - m_new)

At the end, o / l is exactly the softmax-weighted average — the same number the one-shot formula would have produced, up to floating-point rounding. The max is never "global"; it is eventually-consistent, and every step that used a soon-to-be-outdated max gets multiplied back into line.

2.3 A two-chunk worked example (actual arithmetic)

One query row, eight keys, processed in two chunks of four. Scores (already divided by sqrt(hd)):

chunk 1: s = [1, 2, 4, 3]      chunk 2: s = [0, 5, 2, 1]

Value rows: v₁..v₄ = 1, 2, 3, 4 and v₅..v₈ = 5, 6, 7, 8. (Numbers chosen so the arithmetic stays readable; all figures rounded to 4 decimals — sums are evaluated at full precision and then rounded, so re-multiplying the printed operands may differ in the last digit. The online and one-shot totals agree exactly before rounding.)

Reference (one-shot) pass. Global max = 5. Exponentials exp(s − 5):

e⁻⁴    e⁻³    e⁻¹    e⁻²  |  e⁻⁵    e⁰    e⁻³    e⁻⁴
0.0183 0.0498 0.3679 0.1353  0.0067 1.0000 0.0498 0.0183

Denominator l = 1.6462. Numerator o = 0.0183·1 + 0.0498·2 + 0.3679·3 + 0.1353·4 + 0.0067·5 + 1·6 + 0.0498·7 + 0.0183·8 = 8.2916. Final output 8.2916 / 1.6462 = 5.037.

Online pass, chunk 1. Max so far m = 4. Running sum against 4:

l = e¹⁻⁴ + e²⁻⁴ + e⁴⁻⁴ + e³⁻⁴ = 0.0498 + 0.1353 + 1.0000 + 0.3679 = 1.5530
o = 0.0498·1 + 0.1353·2 + 1.0000·3 + 0.3679·4            = 4.7920

Online pass, chunk 2. The new chunk contains the true max, 5, so m_new = 5 and the correction factor is α = exp(m_old − m_new) = exp(4 − 5) = 0.3679. Rescale everything carried from chunk 1, then add the new chunk computed against 5:

o ← 4.7920 · 0.3679 + (0.0067·5 + 1·6 + 0.0498·7 + 0.0183·8)
  = 1.7629         + 6.5287
  = 8.2916
l ← 1.5530 · 0.3679 + (0.0067 + 1.0000 + 0.0498 + 0.0183)
  = 0.5713         + 1.0748
  = 1.6462

Both running totals land exactly on the one-shot values. Notice why the rescale is legitimate: the accumulated o was a sum of exp(s − 4)·v terms, and multiplying by exp(4 − 5) rewrites each term as exp(s − 5)·v. The max is only a shared shift that keeps the exponentials out of overflow/underflow territory; shifting it is algebra, not approximation.

Two details that matter when you read the real kernel:

  • The first chunk is special: m starts at −∞ and α would be exp(−∞ − m_new) = 0, so the code simply skips the rescale on the first tile (see the fresh0/fresh1 flags at fa_prefill_kv (attention_prefill.cu:117)).
  • Masked keys must contribute exactly nothing: causality forbids attending to future positions, so a masked score is forced to 0.0 after the max reduction, not merely given a tiny weight — otherwise l would be polluted and the normalization would be subtly wrong (this is the (gcol[q] < qlim0) guard at fa_prefill_kv (attention_prefill.cu:117).

2.4 RoPE in two sentences, and why the tail is fused

RoPE (Rotary Position Embedding) encodes a token's position by rotating each adjacent pair of the head vector — element j with element j + hd/2 — through an angle proportional to the absolute position. Because a rotation of m followed by the inverse rotation of n cancels to m − n, a query at position m dotted with a key at position n depends only on the relative distance m − n, which is exactly the invariance attention wants. (The full math and the CPU implementation: walkthrough 11 §2.3 — not repeated here.)

Fusion means merging several small kernels into one so the intermediate data never leaves the chip and the launch count drops. In the decode graph, three tiny steps follow the QKV matmuls — add the projection bias, apply RoPE to q and k, scatter k/v into the cache — each a few microseconds of work, each a kernel launch. §3.2 reads the kernel that does all of them in one pass.

3. In minfer's code

3.1 fa_prefill_kv — flash-attention-style prefill attention

The contract. One CUDA thread block (a group of threads that runs on one SM — Streaming Multiprocessor — and can share an on-chip scratchpad called shared memory) owns a 64-token tile of queries for one head, and streams the whole key/value history for that head through a 32-key tile: grid.x = ceil(nt / 64) query token tiles, grid.y = nh heads, 128 threads = 4 warps (a warp is 32 threads that execute in lockstep) launch_fa_prefill_kv (attention_prefill.cu:394, the launcher). K/V come from the persistent cache as __half (16-bit float — the f16 KV cache from chapter 04's bandwidth story); q and the output o are f32. The kernel is gated to hd == 128 models — see the dispatch note at the end of this section.

Why this shape at all. The header comment above the prefill kernel (attention_prefill.cu:10-19) is the honest cost accounting of the kernel it replaced:

// The legacy gqa_attn_f32_f16kv launches one block per (token, head): K is
// re-read per token per head (7B @2K: ~132 GB per layer) and the hd-wide
// accumulator lives in registers (float4 oc[32] = 128 regs → spills). It
// measured 176 ms per layer (76% of the whole 2K prefill). This kernel
// tiles the q dimension: one block per (64-token q tile, head), K/V tiles
// staged in shared memory, QK^T on tensor cores, online softmax with the
// O accumulator in shared memory. K traffic drops to ~0.8 GB per layer.

One block per (token, head) re-reads every K row once per query token; tiling 64 queries together amortizes each K row across 64 consumers — chapter 03's GEMM tiling reasoning, applied to attention.

The tile constants and shared layout (attention_prefill.cu:33-34, the tile constants, and :128-135, the shared layout): FA_TQ = 64 query rows, FA_TKV = 32 key columns per iteration; shared memory holds the q tile, the K tile, and the V tile, all as __half with a padded row stride:

extern __shared__ __align__(256) uint8_t smem[];
// Padded smem row stride: hd=128 halves = 256B ≡ 0 mod 32 banks makes
// every wmma ldmatrix row land on the same bank group (8-way conflict
// per load). +8 halves (272B) shifts each row by 4 banks.
const int sstr = hd + 8;
__half* Qs = reinterpret_cast<__half*>(smem);
__half* Ks = Qs + FA_TQ * sstr;
__half* Vs = Ks + FA_TKV * sstr;

The comment is a miniature lesson in bank conflicts: shared memory is banked in 32 lanes, and when every row is exactly 256 bytes wide, the same column of consecutive rows lands in the same bank, so a multi-row access serializes. The 16-byte padding shifts each row off the hot banks — no algorithm change, just an address formula.

QKᵀ on tensor cores. wmma (Warp Matrix Multiply-Accumulate, the CUDA API for tensor-core matmul) multiplies 16×16×16 matrix fragments; a fragment is the per-lane register layout of a piece of a matrix. The Q·Kᵀ product of a 16-query-row block against the 32-column K tile is accumulated in fragment registers fa_prefill_kv (attention_prefill.cu:117):

wmma::fragment<wmma::accumulator, 16, 16, 16, float> fc[FA_TKV / 16];
for (int cc = 0; cc < FA_TKV / 16; cc++) wmma::fill_fragment(fc[cc], 0.0f);
for (int d = 0; d < hd; d += 16) {
    wmma::fragment<wmma::matrix_a, 16, 16, 16, __half, wmma::row_major> fa;
    wmma::fragment<wmma::matrix_b, 16, 16, 16, __half, wmma::col_major> fb[FA_TKV / 16];
    for (int cc = 0; cc < FA_TKV / 16; cc++)
        wmma::load_matrix_sync(fb[cc], &Ks[cc * 16 * sstr + d], sstr);
    wmma::load_matrix_sync(fa, &Qs[wm * 16 * sstr + d], sstr);
    for (int cc = 0; cc < FA_TKV / 16; cc++)
        wmma::mma_sync(fc[cc], fa, fb[cc], fc[cc]);
}

Line by line: fc[cc] accumulates this warp's 16 query rows against K-tile columns [16·cc, 16·cc+16); the loop over d walks the head dimension 16 elements at a time (the K dimension of this GEMM), loading an A-fragment (16 query rows × 16 dims, row-major) and two B-fragments (16 dims × 16 key columns, col-major — Kᵀ's layout falls out of storing K row-major) per step. The scale factor was already folded into q at load fa_prefill_kv (attention_prefill.cu:117), so the scores need no second pass.

The online softmax, fragment-resident. Now the section-2 machinery, but the "row" lives in tensor-core accumulator registers. For an m16n16 f32 accumulator each lane holds 8 elements — two fragment rows (r0, r0+8) × four column groups — and four lanes (l = 0..3) share one row, so a row max is a local loop plus a 2-step butterfly shuffle (__shfl_xor_sync exchanges a register between lanes) fa_prefill_kv (attention_prefill.cu:117):

float mnew0 = -INFINITY, mnew1 = -INFINITY;
for (int q = 0; q < FA_TKV / 16 * 4; q++) {
    // valid = causal (kv <= query pos) AND within the stored KV range
    // (rows >= kv_end are zero-staged and must NOT contribute).
    bool v0 = (gcol[q] < qlim0) && (gcol[q] < kv_end) && (CAUSAL || gcol[q] >= win_lo);
    bool v1 = (gcol[q] < qlim1) && (gcol[q] < kv_end) && (CAUSAL || gcol[q] >= win_lo);
    if (v0) mnew0 = fmaxf(mnew0, sm[q]);
    if (v1) mnew1 = fmaxf(mnew1, sm1_[q]);
}
for (int off = 1; off <= 2; off <<= 1) {
    mnew0 = fmaxf(mnew0, __shfl_xor_sync(0xffffffffu, mnew0, off));
    mnew1 = fmaxf(mnew1, __shfl_xor_sync(0xffffffffu, mnew1, off));
}

Note the three validity conditions on every element: causal masking (gcol[q] < qlim0 — the per-row exclusive limit fa_prefill_kv (attention_prefill.cu:117) reads from the device-side positions/bound array), tile-range masking (gcol[q] < kv_end — the last KV tile is zero-filled beyond the history end, and zeros must not enter the max), and the window floor (CAUSAL || gcol[q] >= win_lo, E1b). Then the classic triple — fresh flags, exponentiate against the new max, running sums — fa_prefill_kv (attention_prefill.cu:117):

const int fresh0 = (m0 == -INFINITY);
float a0 = fresh0 ? 0.0f : __expf(m0 - mnew0);   // α, the rescale factor
if (mnew0 == -INFINITY) a0 = 1.0f;               // fully-masked tile: no-op
for (int q = 0; q < FA_TKV / 16 * 4; q++) {
    p0[q] = ((gcol[q] < qlim0) && (gcol[q] < kv_end) && (CAUSAL || gcol[q] >= win_lo))
                ? __expf(sm[q] - mnew0) : 0.0f;  // masked ⇒ exactly 0
    sum0 += p0[q];
}
for (int off = 1; off <= 2; off <<= 1)             // 4-lane sum butterfly
    sum0 += __shfl_xor_sync(0xffffffffu, sum0, off);
if (mnew0 != -INFINITY) m0 = mnew0;
l0 = l0 * a0 + sum0;                               // l ← l·α + new sum

The O rescale in code. The output accumulator is 8 fragments (the 64-dim-per-16-row-block V product, hd/16 = 8 blocks). Each fragment's 8 elements interleave the two fragment rows, so the per-row α is applied by multiplying the right lanes of every fragment fa_prefill_kv (attention_prefill.cu:117):

// rescale O fragments by the per-row alpha (x[0,1,4,5] -> row r0,
// x[2,3,6,7] -> row r1 — the m16n16 f32 accumulator layout).
for (int ob = 0; ob < 8; ob++) {
    acc[ob].x[0] *= aa0; acc[ob].x[1] *= aa0;
    acc[ob].x[2] *= aa1; acc[ob].x[3] *= aa1;
    acc[ob].x[4] *= aa0; acc[ob].x[5] *= aa0;
    acc[ob].x[6] *= aa1; acc[ob].x[7] *= aa1;
}

This is §2.2's o ← o·α, 64 rows at a time, entirely in registers. Then P (the probabilities — the scaled scores) is packed into an f16 A-fragment in place fa_prefill_kv (attention_prefill.cu:117), exploiting that the m16n16 f32 accumulator and the m16k16 f16 row-major A-fragment use the same element-per-lane layout), and the P·V product is accumulated fa_prefill_kv (attention_prefill.cu:117). The comment at attention_prefill.cu:322-325 records the layout facts that make the round trip free — the kind of thing you verify once with a standalone fragment-layout test, then trust.

Write-out and K/V staging. After the KV loop, one final normalize acc / l and the fragments go to global memory; the last, partially-filled query tile stages through shared memory so out-of-range rows can be skipped fa_prefill_kv (attention_prefill.cu:117); rows whose l is 0 (fully masked) stay 0. The K/V tiles themselves arrive via fa_stage_kv_async (attention_prefill.cu:43): 16-byte cp.async transfers — an asynchronous copy that lands in shared memory without passing through registers — which zero-fill rows beyond kv_end by capping the copy size (sz = full ? 16 : 0, fa_stage_kv_async (attention_prefill.cu:43)); the pre-sm80 fallback does plain synchronous vector loads. Zero-filling lets the softmax treat out-of-range keys uniformly and exclude them with the one gcol < kv_end test instead of a second control path.

Which models take this path. The host wrapper gqa_attn_kv_prefill (src/cuda/methods/attention.rs:261) gates it:

#![allow(unused)]
fn main() {
// 8n: prefill (nt >= 64) runs the FA-style tiled attention. ...
if nt >= 2 && hd == 128 && !Self::no_fa_prefill() && layout != crate::cuda::KV_LAYOUT_F32 {
    let rc = unsafe { launch_fa_prefill_kv(...) };
    if rc == 0 { return; }
}
}

gqa_attn_kv_prefill (src/cuda/methods/attention.rs:261.) Four conditions, each with a reason: nt >= 2 (doc 86 lowered the gate from nt >= 64, since the kernel masks causally from the positions array, so verify-shaped short batches are safe); hd == 128 (FA_HQ is hard-wired to hd/4 = 32); MINFER_NO_FA_PREFILL=1 no_fa_prefill (src/cuda/methods.rs:281) as the A/B escape hatch; and layout != KV_LAYOUT_F32 (the FA tile is __half, and the launcher is instantiated only for the f16 and packed layouts, so a f32 cache keeps the general kernel). If the shared-memory opt-in fails at launch launch_fa_prefill_kv (attention_prefill.cu:394) the launcher returns −1, prints one loud warning, and the wrapper falls back to the legacy per-token kernel — the one visible fallback in the attention path, and it is announced, not silent. Note for Qwen2.5-0.5B specifically: its head dim is 64 (docs/QWEN2-SUPPORT.md:79 (§4 "Verified models") ), so 0.5B prefill runs the legacy gqa_attn_f32_f16kv kernel; fa_prefill_kv serves the hd=128 classes (Qwen2.5-7B, Qwen3-4B…). The CPU counterpart — the same online softmax in scalar Rust — is walkthrough 11 §3.2's attention arms.

3.2 attn_bias_rope_store_f32 — the fused decode tail

The contract. Decode processes one token (nt == 1). After the QKV projection matmuls, three small jobs remain before attention can run: add the attention biases (if the model has them), rotate q and k by RoPE, and write k/v into the persistent KV cache. The unfused graph spent seven launches on these (add_bias ×3, rope ×2, store_kv ×2 — the count in the kernel's header comment, kv_store.cu:100-117; TECH-PRIMER §6.4 prices the whole campaign at "−310 launches/step"). This kernel is one launch that does all of it, ending with K/V in exactly the layout the next kernel reads.

Thread mapping. The launcher launch_attn_bias_rope_store (kv_store.cu:338) is a flat 1-D grid of 256-thread blocks over total = nqt/2 + nkt/2 + nkt — one thread per RoPE pair for q (nqt/2), one per RoPE pair for k (nkt/2), one per element for v (nkt); nqt = nh·hd and nkt = nk·hd are the q and k section widths of the (single) token's QKV output. Each thread branches on which section its linear id u falls in — three sections, one kernel attn_bias_rope_store_f32 (kv_store.cu:118):

__global__ void attn_bias_rope_store_f32(
    float* __restrict__ q, float* __restrict__ k, float* __restrict__ v,
    const float* __restrict__ bias_q,          // …bias_k, bias_v likewise
    float* __restrict__ kv_k,                  // persistent K region (kv_v too)
    int nqt, int nkt, int hd,
    float freq_base, float freq_scale,
    const int* positions, const int* cells,    // C6: RoPE index, store row
    int kv_is_f16
) {
    const int half_dim = hd / 2;
    const int qpairs = nqt / 2;
    const int kpairs = nkt / 2;
    const int total = qpairs + kpairs + nkt;
    const int u = blockIdx.x * blockDim.x + threadIdx.x;
    if (u >= total) return;
    const int pos = positions[0];          // nt==1: read on device
    const int row = cells[0];              // C6: the row the store lands on

positions[0]/cells[0] are read on the device — the comment at kv_store.cu:116-117 calls this out: "no host scalar crosses the launch — CUDA Graph capture/replay safe". A captured graph freezes its kernel arguments, so a host-side n_past would be baked in and wrong on every replay; device data is re-read each time.

Section 1 — q: bias + RoPE in place attn_bias_rope_store_f32 (kv_store.cu:118):

if (u < qpairs) {
    // q section: bias + rope in place (attention reads q at offset 0)
    const int head = u / half_dim;
    const int d    = u % half_dim;
    const int base = head * hd;
    const int j  = base + d;
    const int j2 = j + half_dim;
    float x0 = q[j]  + bias_q[j];
    float x1 = q[j2] + bias_q[j2];
    float freq = freq_scale / powf(freq_base, (2.0f * d) / hd);
    float theta = pos * freq;
    float cs = cosf(theta), sn = sinf(theta);
    q[j]  = x0 * cs - x1 * sn;
    q[j2] = x0 * sn + x1 * cs;
}

This is verbatim rope_f32 (ops_elementwise.cu:263) with the bias add folded into the loads — same NEOX pairing (j, j + hd/2), same frequency expression, same cosf/sinf. "Verbatim" is a hard requirement: the fused kernel had to be bit-identical to the seven-kernel chain it replaced (the header comment, kv_store.cu:100-117, lists each correspondence) — fusion is only free when the answer does not change. The A/B proof lives in the parity test cuda_kv_f16_roundtrip_attn (cuda_backend/tests/kv.rs:1011 exercises the f16 round trip end to end).

Section 2 — k: bias + RoPE + store into the cache attn_bias_rope_store_f32 (kv_store.cu:118). The first twelve lines are the q-section math with bias_k/k swapped — same pairing, same frequency, same rotation. The new part is what happens after the rotation: the rotated values are written both back to the k buffer and into the persistent K region:

    const float r0 = x0 * cs - x1 * sn;
    const float r1 = x0 * sn + x1 * cs;
    k[j]  = r0;
    k[j2] = r1;
    if (kv_is_f16) {
        ((__half*)kv_k)[(size_t)row * nkt + j]  = __float2half(r0);
        ((__half*)kv_k)[(size_t)row * nkt + j2] = __float2half(r1);
    } else {
        kv_k[(size_t)row * nkt + j]  = r0;
        kv_k[(size_t)row * nkt + j2] = r1;
    }
}

The store lines are the whole KV-cache story in miniature: kv_k[(size_t)row * nkt + j] — the cache is a flat [cell][nkt] array, and the thread computes the scatter row itself from cells[0] (C6: positions[0] is the sequence-relative RoPE angle; the row is the allocator's). The same shape exists in the standalone store_kv_f32 (kv_store.cu:11); the f16 branch converts on store with __float2half (round-to-nearest), the identical conversion the unfused store_kv_f16 path uses — again for bit-identity.

Section 3 — v: bias + store attn_bias_rope_store_f32 (kv_store.cu:118): v gets no RoPE (only q and k are rotated), so its threads add the bias and store one element each, same row * nkt + j addressing into the V region.

Why one kernel instead of three (really seven). Two independent reasons, and they compound:

  1. Launch overhead. Every kernel launch has a fixed CPU-side cost, and the decode step is a chain of hundreds of small kernels (§4 does the arithmetic); TECH-PRIMER §6.4 puts decode chains in the "launch-overhead-bound" regime ("2 µs/graph-gap scale", docs/CUDA-TECH-PRIMER.md:300-302 (§6 "Element-wise and fused epilogue kernels") ). Three launches replaced by one saves two gaps per layer per token, plus the L2 (layer-2 cache on the GPU) round-trips of writing q/k/v out and reading them back.
  2. Producer–consumer locality. The k section writes the rotated values once, directly to the cache address gqa_attn_split (the next kernel in the layer) will read. In the unfused chain k makes three trips through the memory hierarchy — written by rope_f32, read by store_kv_f16, read again by attention — for data that never needed to leave the L2. Chapter 04's "decode is memory-bound" framing is why this matters even for a few KB.

The graph-level counterpart of this kernel is Op::FusedQKV — AGENTS rule 7: "Decode fusions: Op::FusedQKV (concat matmul + bias/rope/store)…" (AGENTS.md:119 ("Compute Graph — core rules", rule 7)), with the mechanics in TECH-PRIMER §6.4 (docs/CUDA-TECH-PRIMER.md:294-298 (§6 "Element-wise and fused epilogue kernels") ). Section 3.4 shows the Rust arm that launches it.

Two ways to call the same kernel. The q/k/v parameters are pointer-form section bases, which lets one kernel serve both decode layer classes attn_bias_rope_store_f32 (kv_store.cu:118): the concat class points all three into one concatenated matmul output (q = base, k = base + nqt, v = base + 2·nkt; the Rust arm does this pointer arithmetic at (cuda_backend.rs:1426-1433, the Op::FusedQKV section bases), and the mixed-quant class (e.g. a model where attn_v is Q6_K and cannot join the concat) points them at three separate matmul outputs (Op::QkvBiasRopeStore, arm at cuda_backend.rs:1299). One device kernel, two graph topologies, zero duplicated math.

3.3 KV in device memory — where the cache actually lives

Chapter 03 followed weights; the KV cache is the other half of GPU-resident state, and it has a different owner. Walkthrough 07 §2.5 explains the allocator side; here we look at where those regions physically sit when the CUDA backend runs, and what the layout buys the kernels of §3.1–3.2.

Ownership and lifetime. Each layer owns exactly two persistent regions, created on first use on the layer's assigned backend (GraphAllocator::ensure_kv (src/graph/alloc.rs:995)):

#![allow(unused)]
fn main() {
fn ensure_kv(&mut self, layer: usize, backend: Backend, n_embd: usize, row_elems: usize,
             n_ctx: usize) -> Result<[BufRef; 2], String> {
    let packed = row_elems != n_embd;   // a Q8_0 node stamps the packed width
    let elems = row_elems * n_ctx;      // (check_width + registry capability)
    if let Some(region) = self.kv.get(layer) { … return Ok([region.k, region.v]); }
    let k = self.alloc_persistent(&format!("kv.{layer}.k"), backend, elems);
    let v = self.alloc_persistent(&format!("kv.{layer}.v"), backend, elems);
    self.kv.insert(layer, k, v, n_embd, row_elems, n_ctx, packed); Ok([k, v])
}
}

ensure_kv is fallible: a packed row_elems is checked against the format and the registry's reads_packed_kv, and an existing region must match n_ctx, backend and packing — else Err. alloc_persistent (alloc.rs:1081) routes through the same pool allocator as everything else — on CUDA a cudaMalloc held in the backend's buffer pool (§3.4) — and registers it as never freed. Because the allocator lives in GraphCache (AGENTS rule 2, AGENTS.md:114 ("Compute Graph — core rules", rule 2)), the regions survive rebuilds and hold their contents across decode steps: two device buffers per layer nobody may recycle.

Size and layout. The size comes from the graph builder: kv_elems: nkt * n_ctx (kv_elems (src/models/qwen2/graph.rs:158)), where nkt = n_head_kv · hd (the n_kv_embd dimension) and n_ctx is the capacity. So the brief question — "[n_past][kv_heads*head_dim]?" — resolves like this in the store code:

region capacity : [n_ctx][nkt]          (nkt = n_head_kv * hd)
element (p, j)  : region[p * nkt + j]   p = absolute position, j = kv dim

The store address dst[positions[t] * nkt + j] store_kv_f32 (kv_store.cu:11) now receives the allocator-resolved cell row (C6: positions ropes, cells stores; they coincide only while a run starts at cell 0). n_past never appears in the layout — it is only ever how many leading rows are valid, and that count lives in the on-device boundary array. That is precisely AGENTS rule 1, "KV positions are data, not structure" (AGENTS.md:113 ("Compute Graph — core rules", rule 1)): the graph topology is identical at position 0 and position 2000, and the kernels discover the valid range from the data — attention_prefill.cu:206-208 records the exclusive per-row limit, and fa_prefill_kv computes kv_end = bound[last_t] + 1 at fa_prefill_kv (attention_prefill.cu:117). Two payoffs we have already met: the decode graph can be allocated once and replayed (§3.5), and attention never needs a host round trip to learn where the history ends.

  decode step, one layer (all CUDA, no host copies)
  rms_norm ──concat matmul──▶ [q | k | v] (pool buf)
                                 │ attn_bias_rope_store_f32:
                                 │  bias+RoPE q,k; K,V scattered at the resolved row
                                 ▼
  ┌─────────────────────┐  ┌─────────────────────┐
  │ kv.{layer}.k region │  │ kv.{layer}.v region │  persistent (never freed)
  │ [n_ctx][nkt] f16    │  │ [n_ctx][nkt] f16    │
  └─────────┬───────────┘  └──────────┬──────────┘
            │ gqa_attn_split reads rows 0..pos
            ▼
      attention output ──▶ wo matmul ──▶ next layer

f16 KV: half the bytes, same addresses. The per-engine kv_layout tag (cuda_backend.rs:46, the KV_LAYOUT_* code) is fixed when the backend is built from the engine's resolved KvFormat (src/graph/kvformat.rs is the single authority): load_model_configured (src/models/mod.rs:391) resolves it from MINFER_CACHE_TYPE and stamps it through GraphAllocator::set_kv_format. When the tag is f16, every store converts to __half and every attention read converts back; §3.2's kernel shows both sides of that. The store_kv_f16 header comment states the trade (kv_store.cu:23-30): "halves attention read bandwidth", and the Rust store_kv_f16 doc comment (src/cuda/methods/kvstore.rs:153-155) adds the fine print — the region stays f32-sized; the f16 view uses the first half of the bytes: allocation does not shrink, the bytes written per store and read per attention call do (§4 does the arithmetic). The correctness story for the f16 round trip is test cuda_kv_f16_roundtrip_attn (cuda_backend/tests/kv.rs:1011).

Why attention can read the regions directly. A KvcacheLoad node is not a copy — its output buffer is the K region (GraphAllocator::node_buffer (src/graph/alloc.rs:1092) maps the node to pair[0]; the CUDA arm comments "out_buf IS the region — no kernel" at (cuda_backend.rs:942, the KvcacheLoad arm comment). So the whole path — matmul, fused tail, cache, attention, next layer — touches pool device memory and crosses no host boundary. The one structural guard on that layout: attention requires hd == hd_kv and nkt == n_head_kv · hd (the kernels stride KV rows by nkt), and violations return Err, not a workaround execute_node_inner (cuda_backend.rs:916); the same guard has a GPU_SAFETY audit entry, docs/GPU_SAFETY.md:105 (§3 "Audit findings (2026-08-02) — status", the H1 finding).

3.4 The Rust host side — cuda_backend.rs as a Backend

Everything device-side so far was launched by a Rust struct implementing the Backend trait (src/graph/backend.rs: capability query, buffer pool, execute_node, host read/write, synchronize). Walkthrough 15 gives the full tour; this section reads the four parts a contributor actually touches.

Struct state CudaBackend (cuda_backend.rs:21). One field per responsibility:

#![allow(unused)]
fn main() {
pub struct CudaBackend {
    state: &'static crate::cuda::CudaState,  // device handle (singleton)
    stream: *mut std::ffi::c_void,           // #188: this engine's own stream
    kv_layout: i32,                          // KV layout tag: f32/f16/q8_0 (§3.3)
    pool: Vec<CudaBuf>, free: Vec<usize>,    // live buffers + free-list indices
    pool_gen: u64,                           // bumped on every (re)allocation
    ...
    graph_execs: Vec<CapturedGraph>,         // captured CUDA Graphs (§3.5)
    capturing: Option<(u64, (usize, usize))>,// open capture window, if any
    graphs_mode: GraphMode,                  // Enabled / Disabled
    ...
}
}

CudaState (src/cuda.rs) is the process-wide singleton: device init, the weight registry, and sync(). It owns a context stream for unbound callers, but since #188 the stream a device operation uses is bound per thread by the calling backend (bind_stream): each CudaBackend creates its own cudaStreamNonBlocking stream, so two engines no longer serialize. pool_gen looks like bookkeeping but is load-bearing for §3.5: a captured graph bakes in device pointers, so any pool churn invalidates every capture, and pool_gen detects that.

execute_node: the dispatch. The trait method (execute_node (cuda_backend.rs:2174)) is a thin wrapper: it calls execute_node_inner and, if the node failed while a capture window was open, aborts the window first (abort_capture (cuda_backend.rs:900) — a doomed window must never be closed into a cached graph). The real dispatch is one big match at execute_node_inner (cuda_backend.rs:916):

#![allow(unused)]
fn main() {
match &node.op {
    // Inputs are host-filled by the allocator; KvcacheLoad is a view
    // of the persistent K region (out_buf IS the region — no kernel).
    Op::Input | Op::KvcacheLoad { .. } => Ok(()),
    ...
}

Three representative arms:

  • Op::FusedQKV (cuda_backend.rs:1365-1453) — the decode tail of §3.2: the concat matmul (matmul_f32_ptr_layout, cuda_backend.rs:1399-1408, writing [q|k|v] into the output buffer), then the fused epilogue (fused_qkv_epilogue, cuda_backend.rs:1435-1451) with pointer-form section bases computed at cuda_backend.rs:1425-1434. KV pointers come from the scheduler's kv_pair (§3.3); guards reject nt != 1, non-neox RoPE, odd hd — each Err with the offending values (cuda_backend.rs:1372-1390).
  • Op::Attn (cuda_backend.rs:1599-1783) — the shape dispatcher of §3.1/§3.2: invariants first (cuda_backend.rs:1604-1651), then nt == 1 → gqa_attn_split (split-K flash-decoding; the 8d comment at cuda_backend.rs:1689-1693 records why: the single-warp kernel left the GPU idle, 48% of the 7B decode step per nsys), 2..=16 (f16/f32) → the batched variant (cuda_backend.rs:1711-1742), nt > 16 → the prefill kernels (cuda_backend.rs:1743-1782), where the host wrapper gqa_attn_kv_prefill (src/cuda/methods/attention.rs:261) internally routes to the FA prefill kernel.
  • Op::KvcacheStore (cuda_backend.rs:1552-1597) — the unfused prefill store: verifies the output buffer is the K region (cuda_backend.rs:1555-1560), derives nt from the element count, converts the boundary array device-side, and launches the layout's store (store_kv_q8_0/store_kv_f16/store_kv_f32, cuda_backend.rs:1581-1595) once per region.

Every arm ends Ok(()) or returns Err(String); there is no third outcome. The match's fallthrough makes the policy explicit execute_node_inner (cuda_backend.rs:916):

#![allow(unused)]
fn main() {
op => Err(format!(
    "cuda: op {op:?} has no kernel (stays on the CPU backend per supports_op)"
)),
}

Buffer pool lifecycle. The pool is a Vec<CudaBuf> (raw pointer + byte length) with a free list of indices:

#![allow(unused)]
fn main() {
fn alloc_buffer(&mut self, size: usize) -> usize {
    let _sg = self.stream_guard(); // cudaMalloc syncs the device
    let bytes = size * 4;
    if let Some(pos) = self.free.iter().position(|&id| self.pool[id].bytes == bytes) {
        let id = self.free.remove(pos);
        self.pool_gen += 1;            // pointers may have changed hands
        return id;
    }
    let ptr = <crate::cuda::CudaState>::cuda_malloc(bytes); // null on OOM (logged)
    self.pool.push(CudaBuf { ptr, bytes });
    self.pool_gen += 1;
    self.pool.len() - 1
}
}

(alloc_buffer (cuda_backend.rs:2121).) Three conventions to notice. Exact-size reuse: the free list matches byte length, so a recycled buffer is always big enough. pool_gen on every path: fresh alloc or reuse, both bump it, because both can change the node→pointer mapping a captured graph depends on. No panics on OOM: cuda_malloc logs and returns null, and the null buffer fails cleanly as an Err from ptr_of at execute time — never a panic inside the allocator (ptr_of (cuda_backend.rs:630); stream_guard is a shim, not a lock). free_buffer never calls cudaFree — it recycles (free_buffer (cuda_backend.rs:2148)), which is what lets the persistent KV regions and the per-step scratch share one arena; alloc_fresh (alloc_fresh (cuda_backend.rs:2158)) bypasses the free list when a buffer's id is still referenced elsewhere (cross-backend staging, walkthrough 07 §2.7). Drop (cuda_backend.rs:859, the Drop impl) frees the pool, the positions scratch, and every captured exec. The ownership rule wrapping all of this is AGENTS rule 8: "Backends own their buffer pools; the allocator is the single owner" (AGENTS.md:120 ("Compute Graph — core rules", rule 8)) — the GraphAllocator decides which buffer a node gets and when it dies; the backend only manages device memory behind those decisions.

read_host / write_host — and the copy rule. The asymmetry is the lesson:

#![allow(unused)]
fn main() {
fn read_host(&self, _id: usize) -> Option<&[f32]> {
    // A staged D2H transfer cannot return a borrowed slice (this method
    // takes &self; the host staging buffer would escape its guard). Use
    // `copy_to_host` via alloc.rs's copy_to_cpu CUDA arm instead.
    None
}
}

(read_host (cuda_backend.rs:2268).) Reading device memory back to the host is always an explicit, syncing copy_to_host (cuda_backend.rs:642: state.sync() then a pinned-staging readback); write_host (write_host (cuda_backend.rs:2275)) is the input-fill path — a pinned-staged async H2D copy, safe because same-stream ordering means later kernels see the data. The rule behind the asymmetry — never host-copy a GPU-pending buffer — is AGENTS rule 5 (AGENTS.md:117 ("Compute Graph — core rules", rule 5)), written in the blood of Phase 3. In three sentences: a per-node host readback inside a split whose command buffer was still open read stale (not-yet-written) data, which surfaced as an all-zero KV region and garbled output (docs/COMPUTE-GRAPH-DESIGN.md:974-976 (§7 "In-place execution and the aliasing rule") , the §7.3 "In-place execution and the aliasing rule" hard rule). The fix was not "sync more" but structural — the in-place aliasing rule plus a single sanctioned copy point at split boundaries — so the bug class has nowhere to reappear. The GPU_SAFETY audit generalizes the lesson: any change to shared mutable GPU state must be validated against a known-good reference, not just an A/B of two paths over the same corrupted state (docs/GPU_SAFETY.md:175-180 (§4a "Split-attention and float4 kernel guards", the shared-mutable-state lesson)).

synchronize and the bounded-wait rule.

#![allow(unused)]
fn main() {
fn synchronize(&mut self) {
    let _bound = self.bind();         // #188: this backend's stream
    if self.capturing.is_none() {
        let _sg = self.stream_guard();
    }
    self.state.clear_mmq_cache();     // one-execution-window scoped
    self.close_capture_or_sync(true); // closes an open capture, or plain sync
}
}

(close_capture_or_sync (cuda_backend.rs:573).) synchronize takes the blocking form and retire the non-blocking one (#138): memos expire, an open capture window closes here, and the actual wait is CudaState::sync (src/cuda/methods/events.rs:143) — cudaGetLastError checked, then cudaStreamSynchronize, and its error code checked. That is the CUDA expression of the GPU-safety rule "synchronize() is the one choke point: stream-ordered work is waited with a bounded loop and the status is checked" (TECH-PRIMER §7, docs/CUDA-TECH-PRIMER.md:317-318 (§7 "Synchronization discipline (GPU Safety, docs/GP"); the rules themselves are docs/GPU_SAFETY.md). At a split boundary the scheduler calls alloc.retire_backend`, whose CUDA arm is the deferred-wait form — §3.5 picks up.

3.5 CUDA Graph capture/replay and cross-backend splits

The problem. A decode step is a few hundred small kernel launches (§4 counts them), each paying a CPU-side cost — TECH-PRIMER §8's one-liner: "per-launch CPU overhead (~2–7 µs) is pure tax" (docs/CUDA-TECH-PRIMER.md:322-323 (§8 "CUDA Graphs — capture once, replay many (Phase 7")). CUDA Graphs (record a sequence of launches once, then submit them all with a single replay call) remove most of that tax without changing the kernels. The scheduler asks the CUDA backend, before executing a split, whether it wants to replay a capture (BackendScheduler::execute (src/graph/scheduler.rs:137); the ask itself is one line, c.graph_replay(graph.uid, split.node_range, …) at graph_replay (src/graph/scheduler.rs:295). On the backend, graph_replay_step (cuda_backend.rs:459) runs a three-run protocol: the first two executions of a (graph uid, node range) go through normal per-node launches (graph_runs counter, graph_replay_step (cuda_backend.rs:459)); on the third, the backend opens a capture window (graph_begin_capture, with no stream lock since #188 — the window opens on this backend's own stream, in cudaStreamCaptureModeThreadLocal, graph_replay_step (cuda_backend.rs:459) — from then until synchronize, every kernel the dispatch enqueues is recorded, not executed. At the boundary, close_capture_or_sync (cuda_backend.rs:573) instantiates the recorded graph, launches it once, and caches the exec; every later step replays the whole split as one graph_launch_exec call graph_replay_step (cuda_backend.rs:459). N per-node launches collapse into one.

That sounds fragile — it would be, if anything the kernels read could change between steps. Two invariants hold it up. First, positions are data (§3.3): kernel arguments (pointers, dims) are identical every step; only buffer contents change, and those are rewritten before replay — TECH-PRIMER §8's "why replay is safe in minfer's design" (docs/CUDA-TECH-PRIMER.md:338-343 (§8 "CUDA Graphs — capture once, replay many (Phase 7")). Second, pool generations: any buffer (re)allocation bumps pool_gen, and a replay whose captured pool_gen differs is destroyed and re-captured graph_replay_step (cuda_backend.rs:459).

MINFER_NO_CUDA_GRAPH=1 is the A/B revert. It forces GraphMode::Disabled at construction with_layout (cuda_backend.rs:204), which makes graph_replay_step (cuda_backend.rs:459) return false — every step runs the plain per-node launch path. It is also the recovery switch: any capture/replay failure disables graphs for the rest of the session with a loud message saying exactly that (cuda_backend.rs:499 and :545, the two "graphs disabled for this session" messages). TECH-PRIMER §8 calls it "the A/B control used by every graph-adjacent step doc" (docs/CUDA-TECH-PRIMER.md:336-337 (§8 "CUDA Graphs — capture once, replay many (Phase 7")). A related hard rule: nothing inside a capture window may sync — a debug readback corrupts the capture, the 7e② "faster but wrong" incident (docs/GPU_SAFETY.md:230 ("CUDA (Phase 7, aarch64 GB10)", rule 2, the 7e② incident)) — which is why trace/viz capture disables replay in the scheduler (BackendScheduler::execute (src/graph/scheduler.rs:137)).

The split/copy story at backend boundaries. On a mixed graph — or any graph where consecutive nodes landed on different backends — the scheduler partitions nodes into contiguous same-backend Splits (split_graph, src/graph/scheduler.rs:87) and executes each with the same boundary protocol (BackendScheduler::execute (src/graph/scheduler.rs:137)):

#![allow(unused)]
fn main() {
if let Some(pb) = prev_backend {
    if pb != split.backend {
        // 1. retire the previous backend's async work (no host block, #138)
        alloc.retire_backend(pb);
        // 1b. staged Metal/CUDA captures are valid now — read back
        flush_metal_captures(graph, alloc, &mut staged, trace_on, live_on);
        flush_cuda_captures(graph, alloc, &mut cuda_caps, trace_on, live_on);
        // 2. enqueue this split's inputs across backends
        for &inp in &split.inputs {
            alloc.copy_across(graph.uid, inp, split.backend)?;
        }
    }
}   // then execute the split's nodes (§3.4's dispatch walk)
}

Retire, then copy, then execute — the only sanctioned cross-backend copy in the system, which is how rule 5's "never host-copy a GPU-pending buffer" survives contact with multi-backend graphs (walkthrough 08 §2.4 calls this the split protocol; the copies are enqueued and each one's single wait is deferred to the consumer's first read, #138). On an all-CUDA model there is exactly one split, the boundary work vanishes, and the loop reduces to the replay check plus the dispatch walk (BackendScheduler::execute (src/graph/scheduler.rs:137); the BackendTag::Cuda arm is BackendScheduler::execute (src/graph/scheduler.rs:137), and a node whose buffer is on another backend than its split is an assignment bug and returns Err with both named (BackendScheduler::execute (src/graph/scheduler.rs:137)).

The gate: all weights registered, or Err — never silent. The kernels of §3.1–3.2 only exist for the quant types the backend implements. minfer's answer to "what if a weight has an unsupported type" is to decide at build time, all-or-nothing: CUDA participation requires a device and every weight registered with a kernel-supported type (Qwen2Graph::device (src/models/qwen2/graph.rs:776), cuda_on = … && Self::weights_on_cuda(model)). weights_on_cuda weights_on_cuda (src/models/qwen2/graph.rs:881) walks every tensor — embedding, per-layer wq/wk/wv/wo, gate/up/down, norms, biases — and on failure prints the exact loser: "CUDA GATE: weight '{}' (type {:?}) has no CUDA kernel or is not registered" — the gate is Qwen2Graph::device (src/models/qwen2/graph.rs:776). That either routes the whole model to CPU (loudly, at build time, recorded in CParams.gpu) or admits the graph as fully-GPU. What is forbidden is the third option: discovering mid-run that a kernel is missing and quietly falling back. If a weight lookup still fails inside execute_node, it is Err naming the weight (execute_node_inner (cuda_backend.rs:916)); an unhandled op is Err execute_node_inner (cuda_backend.rs:916); a kernel-invariant violation is Err with the actual values execute_node_inner (cuda_backend.rs:916). AGENTS states the contract once: "kernel-invariant violations return Err from execute_node — never a silent CPU fallback; backend assignment is decided at build time" (AGENTS.md:107 ("GPU Safety", the Err-never-fallback contract)); TECH-PRIMER §7 repeats it (docs/CUDA-TECH-PRIMER.md:312-314 (§7 "Synchronization discipline (GPU Safety, docs/GP")); the design record states the rule (docs/COMPUTE-GRAPH-DESIGN.md:1102-1104 (§9.2 "Eligibility") ); the assignment walkthrough gives the reason — a quiet fallback would mask the bug the guard exists to catch (docs/inference_e2e_walkthrough/06-assign-fusion.md:111-112`).

4. Performance intuition

Launch overhead, decoded into numbers. Count the kernels one decode step launches, directly off the dispatch table of §3.4, for Qwen2.5-0.5B (24 layers, 14 query heads / 2 KV heads, hd = 64, n_kv_embd = 128 — docs/QWEN2-SUPPORT.md:79 (§4 "Verified models") with the default decode fusions on:

per layerlaunches
RmsNorm ×2 (attn_norm, ffn_norm)2
FusedQKV — QKV matvec + attn_bias_rope_store_f322
Attn (nt = 1) — gqa_attn_split_partial + _combine2
MatMul ×2 — wo, down2
FusedFFN — gate+up concat matvec + in-place swiglu2
Add ×2 (residuals); KvcacheLoad is a view (§3.3)2
per layer12

24 layers × 12 = 288, plus embedding gather, the (memoized, §3.4) positions conversion, final norm, and lm_head ≈ 292 launches per token — counted from the dispatch table, not measured; nsys stats (§5) shows the real number for your quant and gate combination. Price it: TECH-PRIMER §8's measured band for per-launch CPU overhead is ~2–7 µs (docs/CUDA-TECH-PRIMER.md:322-323 (§8 "CUDA Graphs — capture once, replay many (Phase 7")), so the eager path spends roughly 0.6–2.0 ms per token just launching kernels — before the GPU has done any work. A captured step replays all of it with one launch call. The repo has a measured anchor for this class of win: the positions-conversion memo (§3.4) eliminated re-conversions that cost "240 launches/step … ~0.28 ms of pure launch overhead" at a 14B decode (cuda_backend.rs:68-74, the positions-memo comment) — about 1.2 µs per launch, right in TECH-PRIMER's band. §3.2's fusion is the same arithmetic at graph level — the 7-launch QKV tail becomes 1 ("−310 launches/step" across a whole model, docs/CUDA-TECH-PRIMER.md:294-298 (§6 "Element-wise and fused epilogue kernels") — and the dispatch notes price even one wasted launch at "~1-2 us/layer" (attention_decode.cu:930, the split-K dispatch note).

f16 KV bytes per token per layer. With nkt = n_head_kv · hd, each region stores nkt elements per position. Qwen2.5-0.5B: nkt = 2·64 = 128 elements → one f32 K row is 512 B, K + V together 1 KB per token per layer (the walkthrough's number: 24 KB/token across 24 layers, docs/inference_e2e_walkthrough/09-prefill-forward-path.md:253 (§2 "Sizing the context once for both phases") . With f16 KV each row is 256 B → 512 B per token per layer, 12 KB/token model-wide. Decode attention at context length p reads 2 · p such rows per layer, so the halving directly halves the attention kernel's KV traffic; at Qwen3-4B scale (n_kv_embd = 1024, 36 layers — 288 KB per position in f32, docs/inference_e2e_walkthrough/11-attention-vecops-kv.md:71 (§2 "Why the KV cache exists") that is ~144 KB per position touched, though the regions stay f32-sized in allocation (§3.3). The flip side is precision: K/V are rounded to f16 on store and every downstream kernel reads the rounded values — which is why the parity tests compare against the f16-rounded reference, not f32 (cuda_kv_f16_roundtrip_attn, cuda_backend/tests/kv.rs:1011).

Prefill attention: the tiling win in one number. The legacy per-(token, head) kernel re-read the K history once per query token per head: at 7B @2K that was ~132 GB of K traffic per layer, 176 ms, 76% of the whole 2K prefill. fa_prefill_kv amortizes each K row across a 64-query tile and stages K/V once per 32-key chunk: ~0.8 GB per layer — about 165× less traffic fa_prefill_kv (attention_prefill.cu:117). The grid at those shapes is small and regular: ceil(2048/64) = 32 query tiles × 28 heads = 896 blocks of 128 threads, each asking for ((64 + 2·32) · 136 · 2) = 34,816 B ≈ 34.8 KB of dynamic shared memory launch_fa_prefill_kv (attention_prefill.cu:394), raised via cudaFuncSetAttribute (attention_prefill.cu:417). What makes it slow, by construction: an hd ≠ 128 model silently takes the legacy path (0.5B does exactly this, §3.1); a device that refuses the shared-memory opt-in falls back with one printed warning and a "~50× slower" attention launch_fa_prefill_kv (attention_prefill.cu:394); and an unpadded shared-memory stride would re-introduce the 8-way bank conflicts the sstr = hd + 8 line exists to prevent fa_prefill_kv (attention_prefill.cu:117).

5. Try it / Observe

Build once (the nvcc chain is chapter 02's / docs/BUILD.md; the GB10's nvcc is not on every shell's PATH):

export PATH=/usr/local/cuda/bin:$PATH
cargo build --release --features cuda

# A/B the graph replay (the biggest decode lever, §3.5) — same command, env on/off:
./target/release/minfer bench -p 512 -n 128 -r 3 <model.gguf> -o md
MINFER_NO_CUDA_GRAPH=1 ./target/release/minfer bench -p 512 -n 128 -r 3 <model.gguf> -o md

# A/B the attention and fusion paths (each rebuilds the graph; AGENTS rule 7):
MINFER_NO_FA_PREFILL=1 ./target/release/minfer bench -p 512 -n 0 -r 2 <model.gguf>   # hd=128 models only
MINFER_NO_FUSE_QKV=1  ./target/release/minfer bench -p 128 -n 64 -r 3 <model.gguf>   # decode tail, §3.2

Expect the replayed runs to win on decode tok/s (the launch tax of §4); expect identical greedy output — replay is bit-parity-gated (cuda_graph_replay_bit_parity, cuda_backend/tests/capture.rs:274).

Per-node timing and values: MINFER_TRACE records every node's real output stats (decode steps included; KV regions are skipped on Metal, captured in full on CPU) for the viz page — see viz/README.md ("Real trace", viz/README.md:224; the serve-it line is at :44).

MINFER_TRACE=/tmp/t.json ./target/release/minfer <model.gguf> "Hello!" -n 5
# the trace shows the fused nodes (fused_qkv, qkv_bias_rope_store) sitting
# where seven nodes used to be; to see the launch tax itself, compare:
nsys profile -o /tmp/decode --force-overwrite true ./target/release/minfer <model.gguf> "hi" -n 32
nsys stats --report cuda_gpu_kern_sum /tmp/decode.nsys-report

Parity tests (no model file needed — synthetic weights, real GPU):

cargo test --release --features cuda cuda_fa_prefill_attention_parity -- --nocapture
cargo test --release --features cuda cuda_kv_f16_roundtrip_attn     -- --nocapture
cargo test --release --features cuda cuda_graph_replay_bit_parity   -- --nocapture

6. Cross-references

  • 06 — next: optimization + verification, where the launch/bandwidth arithmetic of §4 becomes a toolkit.
  • 04 — previous: the decode matvec and why decode is memory-bound (the premise §3.2's fusion argument stands on).
  • docs/inference_e2e_walkthrough/11-attention-vecops-kv.md — the attention / RoPE / KV math and CPU implementations this chapter deliberately does not repeat (§2.1, §2.3 especially).
  • docs/inference_e2e_walkthrough/07-allocator-liveness-kv.md — why the KV regions are persistent and who frees them (§2.5).
  • docs/inference_e2e_walkthrough/08-scheduler-execute.md — the split protocol and error contract from the scheduler's side (§2.4, §2.5).
  • docs/inference_e2e_walkthrough/15-cuda-backend.md — the end-to-end CUDA backend tour (the "what"; this chapter is the "how, line by line").
  • docs/CUDA-BACKEND-DESIGN.md — design goals + the phase-by-phase implementation record (§5) behind every "Phase 7d"-style comment quoted here.
  • docs/CUDA-TECH-PRIMER.md §7 (synchronization discipline) and §8 (CUDA Graphs) — the reference-depth versions of §3.4–3.5.
  • docs/GPU_SAFETY.md + AGENTS.md — the safety rules cited throughout (rules 1, 5, 7, 8; the Err-never-fallback contract).
  • docs/CUDA_OPTIMIZATION.md (+ docs/cuda_optimization_steps/) — the measured campaign history for every number quoted from a step doc.

← 04 · Reading minfer's kernels II · Index · 06 · Optimization methods →

06 · Optimization methods and how to verify them

Part: Part 4 — from reading to changing. Prereq: chapters 03–05 (you can now read minfer's elementwise, dequant, matmul, attention kernels and the host layer that launches them). Code: docs/cuda_optimization_steps/ (the evidence base — one document per historical optimization step), src/cuda/kernels/*.cu (the kernels each technique lives in — verified lines).

Chapters 03–05 taught you to read minfer's CUDA backend. This chapter teaches you to change it safely. That skill is two-sided: knowing the small set of techniques that actually move GPU kernels (a catalog, each anchored to a real kernel and a real campaign record), and knowing how to prove that your change helped — because on a GPU, "it compiles and prints plausible numbers" is very far from "it is faster and still correct". Both halves come from minfer's own history: 12 optimization sessions, ~60 levers tried, every landed and reverted decision recorded in docs/cuda_optimization_steps/.

1. Background — where this sits

By now you know the machinery: kernels, warps, blocks, shared memory, the CUDA Graph replay loop (chapter 05), and the matmul/attention kernel families (chapter 04). What you do not yet have is judgment: when a kernel is slow, which lever do you reach for first, and how do you know it worked?

minfer's answer to both questions is the same artifact: the campaign record in docs/cuda_optimization_steps/ (indexed by docs/CUDA_OPTIMIZATION.md, the live-status hub — its §0 master table has one row per lever with the measured delta and the one-line lesson). Every technique in this chapter's catalog names the kernel it lives in, the step document(s) that used it, and the measurement that adjudicated it. Nothing here is folklore; every entry is a link into the record.

This chapter is the get-started. Three deeper documents do the heavy lifting:

  • docs/cuda_optimization_steps/77-verification-methodology.md — the campaign's full evidence protocol (§2 and §4 of this chapter summarize it);
  • docs/CUDA_OPTIMIZATION.md — the live status of every lever (landed / reverted / measured-only) plus the env-gate reference in its Appendix A;
  • docs/CUDA-TECH-PRIMER.md — the technique reference this tutorial is forbidden to duplicate (§10 covers profiling depth, §11 the env-gate inventory).

2. Principle — profiling: how to see a kernel before you touch it

You cannot optimize what you cannot see. Two tools ship with CUDA and both are installed on dgxspark (GB10, CUDA 13.0). They answer different questions:

  • ncu (Nsight Compute, /usr/local/cuda-13.0/bin/ncu — not on every shell's PATH) profiles one kernel launch in depth: throughput percentages, occupancy, stall reasons. Use it to answer "what is this kernel's problem?"
  • nsys (Nsight Systems, /usr/local/bin/nsys — on PATH) records the timeline of the whole process: which kernels ran, in what order, how many times, and how much wall-clock time sits between them. Use it to answer "where does the step's time actually go, and how many launches am I paying for?"

What "SOL %" means. ncu's headline section is called GPU Speed Of Light Throughput — the fraction of the hardware's theoretical peak the kernel achieved. Three lines matter most. Compute (SM) Throughput is how busy the arithmetic units (the SM's execution pipes) were as a share of their peak — low means the math units are idle. Memory Throughput is how busy the narrowest memory pipe (DRAM, L2, or L1 — the max of the three) was — low means the memory system is idle. Duration is the kernel's wall time. A kernel with both numbers low (say both under 30%) is latency-bound: threads are mostly waiting (on loads, on barriers, on each other), not computing and not streaming — and the fix is rarely "more FLOPs", it is more concurrency or fewer dependencies. A kernel with Memory Throughput ~90%+ is memory-bound: the only wins left are moving fewer bytes. This one-sentence triage — both low → latency; memory high → bandwidth; compute high → math — decided most of the campaign's lever choices, and ncu prints the same advice as an OPT note when it sees the latency case.

2.1 The minimal ncu command set

Profile a real minfer build (this runs on the actual GPU — commands and numbers below are from this chapter's authoring session):

# Build first (chapter 01; details/pitfalls in docs/BUILD.md)
cargo build --release --features cuda

# Profile ONLY the q8_0 decode-MMVQ kernel, 3 launches, basic section set.
# --launch-count stops after N profiled launches; -k takes a kernel-name regex.
sudo -n env LD_LIBRARY_PATH=/usr/local/cuda-13.0/lib64 \
  /usr/local/cuda-13.0/bin/ncu --set basic --launch-count 3 \
  -k regex:q8_0_p32_q8_mmvq \
  ./target/release/minfer ~/.cache/minfer/models/hf/Qwen/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf "hi" -n 4

Why sudo -n env LD_LIBRARY_PATH=... and not a bare ncu? Two GB10 facts the campaign paid to learn (recorded in 77-verification-methodology.md §2.5): bare ncu fails with ERR_NVGPUCTRPERM — the GPU's performance counters are permission-gated, and this is expected, not a broken install; and plain sudo strips environment variables, which can silently make ncu profile a different (legacy) code path than the one you think you are measuring — so the env var must be re-injected through sudo. Both were verified live for this chapter: the bare command printed ERR_NVGPUCTRPERM; the sudo form worked. (One more local gotcha observed here: sudo ncu creates a root-owned /tmp/nvidia/, after which a plain-user nsys fails with Failed to create directory "/tmp/nvidia/nsight_systems" — run nsys with TMPDIR=/tmp/yourtmp or fix the directory ownership.)

The kernel lines of the observed output (Qwen3-0.6B Q8_0, 9-token prompt, decode q8_0_p32_q8_mmvq — the kernel at q8_0_p32_q8_mmvq (src/cuda/kernels/mmvq_multi.cu:655):

  q8_0_p32_q8_mmvq(...) (1024, 1, 1)x(256, 1, 1), Context 1, Stream 13, Device 0, CC 12.1
    Section: GPU Speed Of Light Throughput
    SM Frequency                    Ghz         2.14
    Elapsed Cycles                cycle       48,821
    Memory Throughput                 %         8.31
    Duration                         us        22.78
    L1/TEX Cache Throughput           %        11.90
    L2 Cache Throughput               %        14.93
    Compute (SM) Throughput           %         8.31
    Section: Launch Statistics
    Block Size                                     256
    Grid Size                                    1,024
    Registers Per Thread             register/thread  40

Read it like this: grid 1,024 blocks × 256 threads on a 48-SM GPU, and both SOL percentages at ~8% — this decode-step kernel (profiled serialized and cold, on a tiny model) is nowhere near either roofline; it is dominated by latency and launch-size effects. That is exactly the kind of verdict ncu gives you in one command. Do not quote ncu's Duration as a speed claim, though: ncu serializes kernels and replays them several times (the output says 9 passes), so its per-kernel times are distorted — the campaign's rule is that ncu is for structural metrics (occupancy, sectors, stall shares, register counts) and nsys is the wall-clock authority.

Useful narrower invocations, all seen in the step records:

# Count only what you need (fast, huge runs stay usable):
ncu --metrics sm__warps_active.avg.per_cycle_active,sm__throughput.avg.pct_of_peak_sustained_elapsed ...
# Export a report file for the UI instead of stdout:
ncu -o /tmp/myreport ...
# Attribute stalls to source lines (PC sampling):
ncu --set full --page source ...

2.2 The minimal nsys round trip

nsys in three sentences. Record: nsys profile -t cuda -o /tmp/trace ./target/release/minfer <model> "hi" -n 16 wraps the whole process and writes a .nsys-rep timeline of every kernel launch, memcpy, and gap. Summarize: nsys stats --report cuda_gpu_kern_sum /tmp/trace.nsys-rep prints the per-kernel table — total time, instance count, min/med/max per kernel — which is where launch-count claims come from. Read: sort by total time and the top rows are your optimization priority list.

Observed on dgxspark (same 0.6B model, nsys stats --report cuda_gpu_kern_sum, top rows quoted as printed):

 Time (%)  Total Time (ns)  Instances  Avg (ns)  Med (ns)  ...  Name
 --------  ---------------  ---------  --------  --------       ----
     51.2       10,285,344        193  53,291.9  50,592.0       void mmq_nt_kernel<...>
     29.0        5,837,152        229  25,489.7  11,648.0       q8_0_f32_matmul(...)
      7.2        1,439,456        113  12,738.5  13,696.0       q8_0_p32_q8_mmvq(...)
      2.0          410,464        223   1,840.6   1,440.0       rms_norm_f32(...)
      1.7          351,232        167   2,103.2   2,048.0       quantize_q8_0_pad40(...)
      1.5          307,360        168   1,829.5   1,824.0       store_kv_f16(...)

This is the instrument behind some of the campaign's most-cited numbers: the D3-8 fusion verdict "total launches −310 per decode step" and the D3-5 verdict "standalone quantize_q8_0_pad40: 4448 → 964 launches per trace" are nsys instance counts, not beliefs (step docs 73 and 70. When you propose a fusion, this table is how you measure what you deleted.

2.3 The evidence discipline — the 15-line version

The full protocol is 77-verification-methodology.md (read it before your first real A/B — §4 of this chapter explains why each gate exists). The tool protocol in compressed form:

  1. ncu needs sudo -n env LD_LIBRARY_PATH=... — plain sudo strips env vars and has silently profiled the wrong (legacy) path before (the r56 lesson, doc 56).
  2. GB10 has no dram__*, launch__grid_size, or shared-sector counters — bandwidth rooflines are derived from lts__t_sectors_aperture_device (L2 sector count × 32 B per sector) plus analytic byte counts (r55).
  3. ncu serializes; nsys is the wall-clock authority — use ncu for occupancy/sectors/stall structure only, never for headline durations.
  4. PC-sampling attributes a stall to the consumer instruction that is waiting, not the producer that is slow (r20/r43) — do not read causality backwards.
  5. SASS first: cuobjdump -sass ./binary | grep LDGSTS before writing a cp.async lever — verify the compiler actually emitted the instruction you are betting on (r45; ptxas -Xptxas -v gives the register/spill budget).
  6. A/B means same-binary, env-gated, interleaved — flip a MINFER_* flag, never rebuild between sides (§4 of this chapter).
  7. Counters + SASS + a reductio must agree before a mechanism claim is allowed to stand (the campaign's counter-forensics rule, doc 18).

Items 1–5 are CUDA-TECH-PRIMER.md §10's territory at reference depth; item 6 is §11's.

3. In minfer's code — the technique catalog

Ten techniques cover nearly everything the campaign did. Each entry follows the same fixed shape: the principle in three sentences, where it lives in minfer's kernels (file:line, verified on the current tree), which step record(s) used it, and the one measurement you would run. Links go to the step documents; CUDA_OPTIMIZATION.md §0 is the cross-reference index if you want the surrounding session.

3.1 Memory coalescing

A warp's 32 lanes issue one memory instruction together, and the hardware coalesces (merges) those per-lane addresses into as few 32-byte sectors as the pattern allows: consecutive lanes touching consecutive addresses is one transaction per warp, while a strided or scattered pattern splits into one transaction per lane. Uncoalesced access therefore multiplies memory transactions without multiplying useful bytes — the classic silent tax. The first question about any data-parallel kernel is thus "what byte does lane i touch, as a function of i?"

  • Where minfer uses it: the elementwise family is written warp-dense — e.g. store_kv_f16 maps one lane to four consecutive floats — the load src + t*nkt + j, store_kv_f16 (src/cuda/kernels/kv_store.cu:31, the float4 read at :42), and the MMVQ decode kernels' shape gate explicitly protects against the uncoalesced case — doc 06 records that small shapes lose because "1–2 units per thread expose the uncoalesced q5/q6 byte loads" (src/cuda/methods/dispatch.rs:350-357, the q5_k MMVQ shape-gate comment).
  • Step records: 11-p5-gemm-tiles-fa-rewrite.md (P5·1: the 1-element-per-thread elementwise kernels "left 15/16 of every transaction unused"; vectorizing to 4/8 elements per lane: 1435 → 1493 tok/s, +4%) and, as the cautionary tale, 26-r21-coalesced-block-linear-a.md (making A-staging perfectly coalesced cut sectors −28.6% yet moved the wall −2.3% — sectors were not the binding constraint; the lesson "stall mass is conserved" lives there).
  • How you'd measure it: ncu's memory tables — sectors per request (l1tex__t_sectors_pipe_lsu_mem_global_op_ld.sum ÷ instruction count; 1.0-ish for dense float loads, 4–32 for strided ones) or simply the L1/TEX Throughput % in the SOL section; A/B via the P5·1 pattern (same kernel, scalar vs vectorized body).

3.2 Shared-memory tiling (chapter 04's GEMM)

DRAM is far too slow to feed tensor cores from directly, so a GEMM block first stages the tile of A and B it needs into shared memory — the on-chip, per-block scratch — and every warp then reads its fragments from there many times. The tile size is the whole trade: bigger tiles cut how often each weight byte is re-read from DRAM/L2 but eat shared memory, and shared memory per block limits how many blocks fit per SM (occupancy, §3.4). Chapter 04 walked gemm_f16_nt_kernel_t line by line; the design arithmetic is in step doc 02.

  • Where minfer uses it: gemm_f16_nt_kernel_t (src/cuda/kernels/gemm_wmma.cu:367) — TN=64 × TM tile, KS=32 k-step, dynamic smem (extern __shared__ at src/cuda/kernels/gemm_wmma.cu:376); the MMQ GEMM family tiles the same way with raw quantized bytes (mmq_raw_nb_kernel, src/cuda/kernels/mmq_nb.cu:9, BT successor mmq_raw_nb_bt_kernel :289; q6_K variant mmq_raw_nb_bt_q6k_kernel src/cuda/kernels/mmq_bt_q6k.cu:42).
  • Step records: 02-wmma-f16-prefill-gemm-8m.md (the 64×64×32 tile turned prefill from 30.7 → 1204 tok/s, 39×, by cutting weight re-reads from nt× to ~1×; also records the two negative tile experiments: KS=64 −38% and TM=256 −3%) and 11-p5-gemm-tiles-fa-rewrite.md (TM=128, +30% — "halves B-panel L2 re-reads and barriers per FLOP").
  • How you'd measure it: whole-prefill tok/s A/B (the tile changes bytes moved per FLOP — lts__t_sectors per GMAC before/after) plus ncu occupancy (§3.4) because the smem budget is what occupancy pays with.

3.3 Vectorized loads (float4 / uint4)

A 32-bit scalar load moves 4 bytes with a full instruction and a full transaction slot; a 128-bit float4/uint4 load moves 16 bytes in one instruction. Vectorizing turns N small loads into N/4 wide ones — fewer instructions, fewer sectors, and wider (LDG.128) transactions — but only when the address is 16-byte aligned and the data layout actually puts 16 useful bytes together. When it applies it is one of the cheapest levers there is; when the layout does not cooperate it is a parity bug factory (nibble offsets, alignment).

  • Where minfer uses it: store_kv_f16 (float4 load + two __half2 stores, store_kv_f16 (src/cuda/kernels/kv_store.cu:31); the q6_K B-expand reads packed data as uint4 groups; the q8_0 p32 decode planes are designed around the uint4* row pointer (q8_0_p32_q8_mmvq, src/cuda/kernels/mmvq_multi.cu:655, the row pointer and the __ldg group loads: src/cuda/kernels/mmvq_multi.cu:665-671).
  • Step records: 11-p5-gemm-tiles-fa-rewrite.md (P5·1, +4%); 44-r41-q6k-bexpand-uint4.md (widening 32 per-byte loads to uint4 groups: q6_K GEMM kernel −61.5%, the long_scoreboard stall share 85.5% → 33.6%, whole prefill +30.7% — the single biggest q6_K lever); 79-phase8-coverage-batch.md (8q: Q5_0's 22-byte block makes its qh word not 4-byte-aligned — two u16 loads instead of one u32, which eliminated the CPU fallback); 76-d4-4-dpl-q6k-final.md and 104-q80-p32-split-plane.md (repacking planes so that uint4 loads align — layout work for vectorization).
  • How you'd measure it: ncu l1tex__t_sectors_pipe_lsu_mem_global_op_ld.sum and the long_scoreboard stall share (smsp__average_warps_issue_stalled_long_scoreboard...) before/after — doc 44's headline is exactly those two counters — plus the same-binary A/B.

3.4 Occupancy & block-size tuning

Occupancy is how many warps are resident per SM, and resident warps are the latency cover: when one warp stalls on a load, the SM schedules another. Per-SM residency is capped by three budgets — threads, registers, and shared memory — so "make the kernel more comfortable" (bigger smem, more registers) can directly cost performance by evicting a resident block, and "spend a little" (a spill, a smaller staging buffer) can buy a whole block back. The campaign's occupancy ladder is the best-documented arc in the record: r28 bought a 2nd block by shrinking smem, r39/40 bought a 3rd by spending registers.

  • Where minfer uses it: __launch_bounds__(256, 3) on the q6_K BT GEMM (mmq_raw_nb_bt_q6k_kernel, src/cuda/kernels/mmq_bt_q6k.cu:42) — the compiler limit of 80 regs/thread to fit 3 blocks/SM (80 × 768 threads = 61,440 ≤ the 65,536-register file, vs 87 regs → only 2 blocks); the q4_K NB kernel's 43,008 B smem budget → 2 blocks/SM (mmq_raw_nb_kernel, src/cuda/kernels/mmq_nb.cu:9, the :15 comment; doc 31 measures the r28 45,056 B); the MMVQ family's __launch_bounds__(256) everywhere.
  • Step records: 31-r28-nb-kernel-2blocks.md (smem 45,056 B ⇒ 2 blocks/SM, +2.56%; ncu sm__warps_active.avg.per_cycle_active 16.17 confirms residency), 43-r40-third-resident-block.md (the register-arithmetic table; +13.0% from one line; also falsified the "0 spill" dogma — 4 B of spill was immaterial), 42-r39-q6k-kdr2-double-buffer.md (+13.3% at 2 blocks/SM), and the negatives: 11-p5-gemm-tiles-fa-rewrite.md (KS=64: depth traded for occupancy, −38%), 53-r50-fa-tkv-16.md (occupancy ↑ cancelled by per-tile overhead — reverted), 67-d3-14b-attribution-bitwise-mmvq.md (block-size right-sizing measured neutral at 14B shapes).
  • How you'd measure it: ncu's Occupancy section (Theoretical vs Achieved Occupancy, Block Limit Registers/Shared Mem/Warps — which budget is the binding one) + ptxas -Xptxas -v for the register/spill numbers before you build; A/B on the whole-prefill bar (+1.5%).

3.5 Warp divergence in quantized kernels

A warp executes one instruction at a time for all 32 lanes; when lanes take different sides of a data-dependent branch, the two sides run serially and the warp pays for the union of both paths (divergence). In quantized kernels the danger looks like a per-element switch on block type or a per-nibble edge case; the structural cure is to make branch conditions warp-uniform (every lane of the warp takes the same side) — or to move the choice out of the kernel entirely. minfer does both: quantization types are dispatched as separate per-type kernels (the type is a template/dispatch-level constant, never a runtime branch inside the loop), and where a fused producer reduces across a warp it explicitly documents the uniformity invariant.

  • Where minfer uses it: the per-type kernel families (q4_k_q8_mmvq, src/cuda/kernels/mmvq_skipwrite.cu:201; q5_k at :320, q6_k at :268) mean the hot loop never asks "which type am I?"; the fused rms+quantize producer quantizes "THIS warp's row (warp-uniform row ⇒ the shfl_xor reductions below never see divergence)" — the comment is at rms_norm_quant_nw_f32_t (src/cuda/kernels/mmvq_skipwrite.cu:25, the comment :71-72); the dequant kernels (dequant_q4_0_f16, src/cuda/kernels/gemm_wmma.cu:38) have a single uniform body with a bounds check only. Divergence still shows up in the accounting: doc 43 attributes part of the gap between achieved and theoretical occupancy to wave-tail divergence, and doc 62 rejects a chunk-distribution scheme precisely because "differing chunk distributions across a warp cause divergence".
  • Step records: 79-phase8-coverage-batch.md (8e/8e② — per-type kernels and the launch table instead of in-kernel type switches), 55-r52-skip-write-mode2.md (the skip-write quantizer's reduction is arranged so lanes never diverge), 62-r59-q4k-wdsc-plane.md (divergence as a design veto).
  • How you'd measure it: ncu's Scheduler Statistics — the "Warp Cycles Per Issued Instruction" and branch-efficiency metrics (smsp__sass_average_data_branch_divergence...) — plus the SASS census approach of doc 30 when in doubt.

3.6 Kernel fusion (attn_bias_rope_store_f32, Op::FusedQKV/FusedFFN)

Every kernel launch has a fixed price — launch latency, the input/output round trip through memory, the scheduler gap between kernels. Fusing merges a chain of small kernels into one so intermediate values stay in registers or shared memory and N launches become 1. On minfer's decode path the fusion targets were chosen by counting launches with nsys: the per-layer chain add_bias×3 + rope×2 + store_kv×2 (7 launches) became one kernel, and nsys counted the difference.

  • Where minfer uses it: attn_bias_rope_store_f32 attn_bias_rope_store_f32 (src/cuda/kernels/kv_store.cu:118) — one launch replaces the 7-launch decode tail; the graph ops Op::FusedQKV (concat matmul + fused epilogue) and Op::FusedFFN (gate|up concat matmul + in-place swiglu) are declared in Op::FusedQKV (src/graph/ops.rs:158), Op::FusedFFN (:177), executed in execute_node_inner (src/graph/cuda_backend.rs:916; the Op::FusedFFN arm :1236, the Op::FusedQKV arm :1365), and gated at build time in fuse_qkv (src/models/qwen2/graph.rs:131-138). The decode A-quantize fusion (swiglu_quant_pad40, rms_norm_quant_pad40, swiglu_quant_pad40 (src/cuda/kernels/ops_elementwise.cu:215), rms_norm_quant_pad40 (:90) writes the quantized activation plane beside the f32 output so the following matmul skips a standalone quantize launch.
  • Step records: 73-d3-8-fusedqkv-port.md (D3-8: −310 launches/decode-step, +1.63% @14B — and note the fusion is bit-identical because each piece is verbatim the unfused math), 70-d3-5-fused-producer-a-quantize.md (D3-5: standalone quantize 4448 → 964 launches per trace), 72-d3-7-attnv-mmvq-rms.md (2c: f32_bits_to_i32 239.6 → 1.2 launches/step via one-execution-window memoization — fusion's sibling, caching).
  • How you'd measure it: nsys stats --report cuda_gpu_kern_sum instance counts before/after (launches deleted are the point), then the A/B gate: MINFER_NO_FUSE_QKV=1 / MINFER_NO_FUSE_FFN=1 flip the same binary to the unfused topology (they are part of the graph-reuse identity — the rebuild is forced for you; try_reuse (src/graph/cache.rs:69).

3.7 CUDA Graph launch amortization (MINFER_NO_CUDA_GRAPH=1)

A decode step is hundreds of tiny kernels; each launch costs microseconds of host + driver work. CUDA Graph records the whole kernel sequence once (capture) and replays it with a single call — the kernels are identical, only the launch overhead is amortized. minfer captures the whole decode step (Phase 7d) and repeatedly-identical prefills (R3-B, after a 3-run "will this repeat?" protocol — one-shot CLI prefills never capture, and r55 measured that capture would be a pure loss for them plus a capture-illegal mid-window malloc).

  • Where minfer uses it: CudaBackend (src/graph/cuda_backend.rs:21) reads MINFER_NO_CUDA_GRAPH (replay off → per-kernel eager launches); the capture/replay machinery and its pool-generation invalidation live in that file's CudaBackend (chapter 05 walked it).
  • Step records: 01-phase7-cuda-backend.md (7d decode capture/replay), 07-r3-small-model-overhead.md (R3-B prefill capture default-on; the 3-run protocol rationale; also A1/A2 — the non-graph overheads a timeline exposes), 58-r55-swiglu-roofline-prefill-graph.md (the measured skip for one-shot prefills).
  • How you'd measure it: the env-gate A/B. Measured for this chapter on Qwen3-0.6B Q8_0, bench -p 256 -n 32 -r 2, two interleaved pairs (author's run, 2026-09-14, quiet box): graph on tg32 224.45 / 223.29 tok/s vs graph off 183.31 / 181.00 — ≈ +23% decode, 2/2 pairs with clean separation; pp256 7510/7411 vs 7319/7305 (~+2%). The cited campaign-side number for the same gate at 14B scale is in doc 88 (88-d5-r-stage5-final-battery.md): "48.08 captured vs 50.15 eager" for the prefill-capture cell.

3.8 Async copy / double buffering (cp.async, KDR=2)

cp.async (SASS: LDGSTS) copies global→shared asynchronously: the warp issues the copy for the next tile and keeps computing the current one, with commit_group/wait_group N as the pacing barrier. Double buffering gives the pipeline two buffers per plane (fill one while computing the other), which turns the staging latency from a wall into overlap — if the kernel is actually staging-latency-bound, and if the extra smem does not cost occupancy (§3.4's trap: r39 notes doubling every plane at KDR=4 is exactly the 1-block/SM mistake).

  • Where minfer uses it: the q6_K BT GEMM stages every per-kt plane twice ("r39: DOUBLE-BUFFERED staging — two copies of every per-kt plane so kt+1's global→smem expansion overlaps kt's compute", comment at mmq_raw_nb_bt_q6k_kernel (src/cuda/kernels/mmq_bt_q6k.cu:42, the r39 comment :52-54); the A/B/dsc staging planes ride cp.async (the r53/r56 bundles); gemm_f16_nt_kernel_t double-buffers its A/B panels (doc 02 §2.4).
  • Step records (this technique has both spectacular wins and instructive nulls): 02-wmma-f16-prefill-gemm-8m.md (8m② cp.async: 31 → 35 TFLOPS), 42-r39-q6k-kdr2-double-buffer.md (+13.3%), 56-r53-q6k-wexp-cpasync-bundle.md (+5.03% — and the "nearly landed silently" liveness lesson), 59-r56-q6k-a-cpasync-wdsc.md (+2.35%, LDGSTS verified in SASS); the negatives: 17-staging-shape-family.md (cp.async-db neutral — the kernel was L2-throughput-bound, not MLP-starved), 61-r58-q4k-bt-cpasync-transplant.md (−12.6%: the mechanism's cost scales with what it replaces), 66-d2-kv-register-staging.md (three cp.async attention pipelines all slower than register staging), and 98-bt-cpasync-null.md (the final null that closed the staging line on GB10).
  • How you'd measure it: cuobjdump -sass | grep LDGSTS (is the async copy even emitted?), ncu long_scoreboard share before/after, kernel µs from nsys (not ncu), and the wall A/B with the +1.5% bar.

3.9 f16 KV storage

KV cache rows are read once per decode step per head, every step, for the whole context; storing them as f16 instead of f32 halves those bytes while attention math stays f32-accumulated. The win grows linearly with context, and the cost is none at the byte level — f16 KV is a pure traffic cut — but it is a policy decision (which tensors convert, who else reads the region), so it is gated and load-time-decided.

  • Where minfer uses it: store_kv_f16 (src/cuda/kernels/kv_store.cu:31) and the f16-KV attention mirror gqa_attn_f32_f16kv (src/cuda/kernels/attention_decode.cu:18); the policy is per engine now — KvFormat (src/graph/kvformat.rs, stamped into CUDA's kv_layout tag): auto-f16 on a GPU at n_layers × n_kv_embd ≥ 8192; MINFER_CACHE_TYPE=f32|f16|q8_0 override.
  • Step records: 79-phase8-coverage-batch.md (8b: 7B @2K decode +11%; the caveat on record — MINFER_GRAPH_DUMP reads KV as f32, so dump and f16-KV are incompatible on the debug path).
  • How you'd measure it: decode tok/s A/B across context lengths (the delta widens with KV size — doc 79's framing), or nsys attention-kernel time at two context lengths; MINFER_CACHE_TYPE=f32 is the A/B side of the same binary.

3.10 int8 MMQ prefill (weights quantized, activations quantized on the fly)

The f16 route (§3.2) dequantizes weights once and runs one f16 GEMM; the MMQ route keeps weights in their quantized bytes end-to-end and quantizes the activations to int8 (q8_0) in a prepass, so the GEMM streams ~4× fewer weight bytes and multiplies on int8 tensor cores (mma.m16n8k32). That is what made minfer's prefill reach llama.cpp parity (1.080× at r59b) — and it is the single deepest line in the campaign: the design was reverse-engineered from llama.cpp (the MMQ analysis doc), then re-derived kernel by kernel over ~30 rounds.

  • Where minfer uses it: the A-plane prepass (quantize_q8_0_pad40_t, src/cuda/kernels/mmvq_aquant.cu:85 — the pre-transposed, 64-token-blocked layout), the raw-byte NB/BT GEMM family (mmq_raw_nb_bt_kernel, src/cuda/kernels/mmq_nb.cu:289; q6_K variant mmq_raw_nb_bt_q6k_kernel, src/cuda/kernels/mmq_bt_q6k.cu:42), dispatched for nt ≥ 9 under the MINFER_MMQ gate read through CudaState::mmq_gate_on (src/cuda/methods/policy.rs:14).
  • Step records: 08-r1-int8-mmq-prefill-gemm.md (R1: parity-first strategy — "parity-clean but ~2.9 TMAC/s vs llama ~24: the 8× gap was unprofiled"), the r9→r59 redesign ladder (37-r34-quantize-transpose-prepass.md +9.72% layout-transform locality; 44-r41-q6k-bexpand-uint4.md +30.7%; 62-r59-q4k-wdsc-plane.md +11.1% W_dsc plane; 63-r59b-clean-remeasure.md the baseline-correction that produced the definitive 1.080×), 64-r60-promotion-default-on.md (the verified gate set flipped default-on). The llama.cpp side of the story — what their MMQ does, instruction for instruction — is docs/LLAMA-CPP-MMQ-ANALYSIS.md (§11 mirrors the redesign rounds one for one).
  • How you'd measure it: whole-prefill tok/s interleaved A/B with MINFER_MMQ=0 as the f16-escape side of the same binary (Appendix A of CUDA_OPTIMIZATION.md carries the full gate list and the memory cost of each plane), plus lts__t_sectors per GMAC for the byte-stream claim.
  • Persistent kernels (blocks that loop over tiles instead of one-tile-per- block): tried once — r24's scheduling ladder — measured −3.3% and reverted (29-r24-scheduling-ladder.md); no wave-quantization tail existed to remove on these shapes.
  • TMA (Tensor Memory Accelerator) and cluster launch: no step record mentions them; the campaign's staging levers were all cp.async/smem/L2, and nothing in the current kernels needs hardware-managed tensor tile movement. They are future-hardware levers, not GB10 campaign levers.
  • PDL (programmatic dependent launch) — the closest relative of the above that was fully integrated: measured −2.6%/−1.8% in-situ and reverted, with a standing warning against co-residency tricks on this workload (76-d4-4-dpl-q6k-final.md; CUDA-TECH-PRIMER.md §11's PDL note).

4. The verification discipline — why the gates exist

Everything in §3 came with a measurement, and the measurements all share one skeleton — the campaign's gate chain, written up in full in 77-verification-methodology.md. This section is the why: every gate can state in one sentence which defect class it defends — doc 77's first lesson is that a gate which cannot is ritual, not verification. The doc's three villain names are phantom gain, correctness erosion, and baseline drift.

Gate 1 — the parity trio (numeric correctness). Three independent parity checks run before any landing: cuda_prefill_mmq (1/0, 8 quant types × 8 shapes against the host reference — defends the quant kernels' numeric path), cuda_prefill (7/0 — the prefill graph end to end), and cuda_fa_prefill_attention_parity (1/0 — the attention kernel). The 1e-3 tolerance is informative, not lenient: a nibble-layout error (the classic quantized-kernel bug) deviates at ~1e0 magnitude while legitimate f32 rounding noise sits at ~1e-5, so the magnitude of a failure identifies its class. Catches: wrong unpacking, wrong scale folds, misaligned loads — the "it produces numbers" class of bug.

Gate 2 — greedy byte-for-byte identity (end-to-end correctness). -n 32 --greedy --seed 42 on a long prompt, pre-change binary vs candidate: the token stream must be byte-for-byte identical. This defends the class parity fixtures structurally cannot see — graph-level and memory-level corruption that happens to miss the fixture shapes: the doc's two real cases are r52's rms-kernel out-of-bounds write and r58's smem cross-write, both invisible in kernel tests and both fatal to a 2000-token generation. The known boundary: any change to tile size necessarily regroups float accumulation order, so for that class the gate is swapped for a calibrated tolerance package (kernel ≤1e-4 on outlier data, the argmax hard gate, the rp=1.0 greedy stream — the D3a calibration, doc 68). Catches: correctness erosion — the sampling chain amplifying a numerics change into visible divergence.

Gate 3 — interleaved A/B medians (performance truth). 3×/5× pairs, same binary (env-gated), same time window, alternating which side runs first; headline numbers require distribution separation (min-new > max-base), and the whole-prefill landing bar is +1.5% against a re-measured baseline. Alternating order is the design core: the GPU is shared, and back-to-back runs disguise window drift as a trend while paired medians cancel it. Catches: phantom gain (the fast path never actually taken — hence the liveness-label rule, r53/r54: an intentional fallback prints exp=off, an accidental one prints fallback!) and baseline drift (the r59/r59b story: a "baseline" binary that was actually a stale deficient build inflated a +26.2% claim down to its true +11.1%, doc 63).

Gate 4 — the suite. The device test suite (166 → 174+ over the campaign; 576 on 2026-10-07) guards collateral damage; a test that flakes under a co-tenant window is adjudicated with an isolated --exact rerun, never silently retried to green. Catches: regression in other kernels — the change you did not mean to make.

The rule for a new kernel variant. A variant is not "done" when it is faster; it is done when it has (1) the parity trio, (2) its greedy-identity or calibrated-tolerance verdict, (3) an interleaved same-binary A/B through an env gate with separation past the bar, and (4) a liveness/label check proving the fast path actually ran. The env-gate A/B pattern is the load-bearing habit: one binary, two sides, flipped by a MINFER_* switch — never rebuild between A and B, because a rebuild changes the baseline too. The campaign's full gate inventory (which switch reverts which lever) is CUDA-TECH-PRIMER.md §11 and CUDA_OPTIMIZATION.md Appendix A. One-line summary of doc 77 §4's hard-won coda: when both sides of your A/B share the same bug, the comparison is bit-identical — cross-binary comparisons must anchor on a -n 1 first-step dump (D4-2's blind-spot case).

5. Try it — three exercises, escalating

Each exercise states the task, what to measure, which step document recorded the original result, and the expected difficulty. None of them requires touching the repo (exercise (b) explicitly forbids it): work in /tmp copies.

Exercise (a) — vectorize a toy and A/B it with cudaEvents (easy)

Task. Take a scalar elementwise kernel from your chapters 01–02 toys (the vector add or SiLU toy) and write a second version where each thread handles 4 elements via one float4 load. Time both with cudaEvent pairs (device-side timestamps — the §3.7 A/B discipline in miniature: fresh input, interleaved reps, report both). A complete, compilable example is below; the A/B skeleton is the part to keep for your own kernels.

// /tmp/vec4_ab.cu — scalar vs float4 SAXPY, timed with cudaEvents (CUDA 13.0, GB10)
#include <cstdio>
#include <cuda_runtime.h>

__global__ void saxpy_scalar(int n, float a, const float* x, float* y) {
    int i = blockIdx.x * blockDim.x + threadIdx.x;
    if (i < n) y[i] = a * x[i] + y[i];              // 1 element per thread
}

__global__ void saxpy_float4(int n4, float a, const float4* x, float4* y) {
    int i = blockIdx.x * blockDim.x + threadIdx.x;
    if (i < n4) {                                   // 4 elements per thread
        float4 xv = x[i], yv = y[i];
        yv.x = a * xv.x + yv.x; yv.y = a * xv.y + yv.y;
        yv.z = a * xv.z + yv.z; yv.w = a * xv.w + yv.w;
        y[i] = yv;
    }
}

int main() {
    const int n = 1 << 26;                          // 64 Mi floats = 256 MB
    const size_t bytes = (size_t)n * sizeof(float);
    float *x, *y, *hx = (float*)malloc(bytes), *hy = (float*)malloc(bytes);
    for (int i = 0; i < n; i++) { hx[i] = 1.0f; hy[i] = 2.0f; }
    cudaMalloc(&x, bytes); cudaMalloc(&y, bytes);
    cudaMemcpy(x, hx, bytes, cudaMemcpyHostToDevice);
    const float a = 2.0f, blocks = 1024;

    cudaEvent_t beg, end;                           // events: device-side timestamps
    cudaEventCreate(&beg); cudaEventCreate(&end);
    for (int rep = 0; rep < 2; rep++) {             // 2 timed reps each, interleaved
        float ms_s = 0.f, ms_v = 0.f;
        cudaMemcpy(y, hy, bytes, cudaMemcpyHostToDevice);   // fresh y per run
        cudaEventRecord(beg);
        saxpy_scalar<<<(int)((n + blocks - 1) / blocks), (int)blocks>>>(n, a, x, y);
        cudaEventRecord(end); cudaEventSynchronize(end);
        cudaEventElapsedTime(&ms_s, beg, end);
        cudaMemcpy(y, hy, bytes, cudaMemcpyHostToDevice);   // fresh y per run
        cudaEventRecord(beg);
        saxpy_float4<<<(int)((n / 4 + blocks - 1) / blocks), (int)blocks>>>(
            n / 4, a, reinterpret_cast<const float4*>(x), reinterpret_cast<float4*>(y);
        cudaEventRecord(end); cudaEventSynchronize(end);
        cudaEventElapsedTime(&ms_v, beg, end);
        // both kernels move 3 x bytes (x read, y read, y write): GB/s = 3*bytes/ms
        printf("rep %d: scalar %.3f ms (%.0f GB/s)  float4 %.3f ms (%.0f GB/s)\n",
               rep, ms_s, 3.0 * bytes / ms_s / 1e6, ms_v, 3.0 * bytes / ms_v / 1e6);
    }
    float maxerr = 0.f;
    cudaMemcpy(hy, y, bytes, cudaMemcpyDeviceToHost);
    for (int i = 0; i < n; i++) maxerr = fmaxf(maxerr, fabsf(hy[i] - 4.0f));
    printf("max |y - 4| = %g\n", maxerr);
    cudaEventDestroy(beg); cudaEventDestroy(end);
    cudaFree(x); cudaFree(y); free(hx); free(hy);
    return 0;
}
/usr/local/cuda/bin/nvcc -O2 -arch=sm_121 /tmp/vec4_ab.cu -o /tmp/vec4_ab && /tmp/vec4_ab

Observed on the GB10 (author's run; both kernels stream 3×256 MB — read x, read y, write y):

rep 0: scalar 3.539 ms (228 GB/s)  float4 3.221 ms (250 GB/s)
rep 1: scalar 3.173 ms (254 GB/s)  float4 3.205 ms (251 GB/s)
max |y - 4| = 0

What to measure, and the honest trap. The expected lesson is not "float4 wins" — this access is already dense and coalesced, every byte of every sector is useful, so vectorization mostly removes instructions and the time is a wash inside noise. That is the point: §3.3 pays when the scalar version wastes sectors (strided, or 1 useful byte per 32-byte sector). To see a real win, make the scalar kernel touch memory with a stride (e.g. one element per 32) and watch both the A/B and ncu's sectors-per-request move. The original record of this exact effect is 11-p5-gemm-tiles-fa-rewrite.md (P5·1: +4% whole-prefill from the 15/16-unused-transaction fix).

Exercise (b) — change dequant thread granularity in a copy (medium)

Task. Copy dequant_q4_0_f16 (src/cuda/kernels/gemm_wmma.cu:38) into /tmp/dq_gran.cu with a synthetic weight buffer — do not modify the repo. The incumbent maps one thread to one 32-element block (g = blockIdx.x * blockDim.x + threadIdx.x indexes od*nb blocks; each thread writes 32 halves). Write two variants: (1) one thread per element pair (a thread unpacks one byte's two nibbles); (2) one thread per four blocks (64 halves per thread). Host-verify all three against a scalar CPU dequant (max|Δ| == 0), then profile each with ncu:

sudo -n env LD_LIBRARY_PATH=/usr/local/cuda-13.0/lib64 \
  /usr/local/cuda-13.0/bin/ncu --set basic -k regex:dq_ /tmp/dq_gran

What to measure. Duration is suggestive only (ncu serializes); compare Memory Throughput %, L1/TEX Throughput %, and the Occupancy section's Theoretical Occupancy and Block Limit lines across the three granularities. The incumbent sits in a real sweet spot — element-parallel dequant is write-bound and trivially coalesced — so your most likely finding is "the incumbent was already right", which is itself a result (the campaign's metadata: 8p's fused-dequant experiments live in 05-persistent-f16-cache-8p.md, and the aligned/unaligned load subtleties of quant blocks — the Q5_0 22-byte block — are doc 79's 8q item). Difficulty: medium — the CUDA is easy; the discipline (bit-exact host check before trusting any timing) is the exercise.

Exercise (c) — re-run one historical A/B on the current tree (hard)

Task. Pick one clear-win step doc — good first choices: doc 43 (__launch_bounds__(256,3), one-line code change) or doc 11's P5·2 (TM=128) — and re-run its A/B on the current tree using its env gate and protocol: same benchmark anchor, interleaved pairs, median-vs-median, distribution separation. For doc 43 the modern equivalent knob is the q6_K kernel's __launch_bounds__ line itself (do not modify the repo — build a scratch worktree copy in /tmp if you want to flip it); for P5·2 note the era shift first: MINFER_GEMM_TM (64/128/256, read at launch_gemm_f16 (src/cuda/kernels/gemm_wmma.cu:847) retiles the f16 wmma GEMM, which is only on the hot path when you run the escape side MINFER_MMQ=0 — exactly the A/B frame P5·2 was measured in. Then compare your numbers with the doc's recorded ones.

What to measure. Whole-prefill tok/s (or bench -p/-n cells for decode), 3+ interleaved pairs, min-new > max-base required before quoting a delta. The teaching payload is the meta-result: your absolute numbers will not match the doc's (the anchors moved; machine state differs — that is r59b's baseline-anchoring lesson, doc 63), but a healthy tree should reproduce the direction of every landed lever and the flat/negative direction of every reverted one. If a landed lever reads negative on your run, suspect your own protocol first (wrong gate value, no warmup, co-tenant window) — that suspicion reflex is the chapter's real deliverable. Difficulty: hard — not because any step is hard, but because this is the first time you own the whole gate chain; doc 77 is the rubric.

6. Cross-references

← 05 · Reading minfer's kernels III · Index · 07 →

07 · Where to go next — the map, the cheat sheet, the pitfalls

Part: Part 5 — leaving the ladder. Prereq: chapters 01–06. Code: none — pointers and reference tables.

1. You are here

You can now read every kernel in src/cuda/kernels/*.cu, you know how minfer's decode (GEMV + fused tail + CUDA Graph replay) and prefill (MMQ GEMM + fa_prefill_f16kv) paths are assembled, and you know the verification discipline that keeps kernel changes honest. This chapter hands you to the deep references — the tutorial deliberately stopped where they begin.

2. The reading map

OrderDocumentWhat it gives you nowRead how
1docs/CUDA-TECH-PRIMER.mdevery technique the campaign uses, at reference depth (platform §1, software stack §2, memory §5, kernel inventory §6, sync §7, graphs §8, numerics §9, profiling §10, env gates §11)full read once; it will feel like revisiting old friends — you met every section's topic in chapters 01–06
2docs/CUDA_OPTIMIZATION.mdthe campaign's live status: what landed, what was rejected, current best numbersread the current status block; keep it open while browsing steps
3docs/cuda_optimization_steps/70+ step records — each is a worked example of chapter 06's method: hypothesis → change → gate chain → verdictstart with 77 · verification methodology, then follow the era indexes (Part I foundations → Part IV Era D); skim failures — they teach the defect classes
4docs/CUDA-BACKEND-DESIGN.mdwhy the backend is shaped the way it is (phases, risks, llama.cpp reference map)one pass; you now recognize every phase's code
5docs/GLOSSARY.mdevery term/formula, classifiedreference — jump in when a term resurfaces
6walkthrough 14 · Metalthe same concepts on a different backend APIoptional; a good transfer test: map threadgroup↔block, MSL↔CUDA C yourself

3. Appendix A — command cheat sheet

# Toolchain (GB10, CUDA 13.0; nvcc is not on every shell's PATH)
export PATH=/usr/local/cuda/bin:$PATH
nvcc --version                       # 13.0
nvidia-smi                           # NVIDIA GB10, driver 580.x

# Compile a toy from any chapter (example: Toy #1)
nvcc toy1_vec_add.cu -o toy1 && ./toy1

# Build minfer with the CUDA backend (see docs/BUILD.md for details/pitfalls)
cargo build --release --features cuda          # + CUDA backend
cargo build --release --features cuda,cuda_static   # static cudart

# Run + observe
./target/release/minfer model.gguf "hello"     # graph path on GPU
MINFER_NO_CUDA_GRAPH=1 ./target/release/minfer model.gguf "hello"   # A/B: per-kernel launches
MINFER_NO_FUSE_QKV=1 MINFER_NO_FUSE_FFN=1 ./target/release/minfer model.gguf "hello"  # A/B: unfused
MINFER_TRACE=/tmp/t.json ./target/release/minfer model.gguf "hello"  # per-node trace (viz/README.md)

# Profile a kernel (see ch 06 §1 for how to read the output)
ncu --set basic ./target/release/minfer model.gguf "hi"
nsys profile ./target/release/minfer model.gguf "hi"

4. Appendix B — pitfalls the campaign already paid for

Each of these is a real defect class from minfer's history or the CUDA programming model; the walkthrough/GPU_SAFETY links go to the full stories.

  • Index out of bounds at the tail — every i = blockIdx.x*blockDim.x + threadIdx.x kernel needs the if (i < n) guard; ceil-div grids overshoot.
  • __syncthreads() inside a divergent branch — threads that skip the branch never arrive; the barrier hangs or corrupts. Barriers live on straight-line block-wide control flow only.
  • Forgetting the error check — kernel failures are asynchronous; without cudaGetLastError() right after launch (and a sync before readback) a broken kernel silently produces garbage. This is why docs/GPU_SAFETY.md mandates checked, bounded submission.
  • Host-copying a GPU-pending buffer — reading device memory the GPU is still writing gives stale/corrupt data; minfer's rule: never host-copy a GPU-pending buffer (the Phase-3 KV-corruption bug; docs/GPU_SAFETY.md).
  • Unbounded sync waits — "wait forever" turns a wedged kernel into a wedged process; minfer's synchronize waits bounded and checks status.
  • Silent CPU fallback — a guard failure must Err, never quietly run the op on CPU: benchmark numbers from a fallback are phantom gains (docs/GPU_SAFETY.md; backend assignment is a build-time decision).
  • Trusting an unisolated speedup — thermal/tenancy drift and one-off runs invent gains; the gate chain (ch 06 §3) exists because "it felt faster" is not evidence.

5. Appendix C — toy index

ToyChapterTeachesVerified with
#1 toy1_vec_add.cu01first kernel: launch, indexing, error checks, timingnvcc 13.0 / GB10
#2 SiLU elementwise02elementwise kernel + CPU reference comparenvcc 13.0 / GB10
#3 stream timing02default vs named stream, cudaEvent timingnvcc 13.0 / GB10

6. Cross-references

← 06 · Optimization methods · Index

Qwen3-4B Q4_K_M Performance: minfer vs llama.cpp

Analysis of where minfer (compute-graph path) stands against llama.cpp on the same GGUF file, and why — measured on the M4 Pro dev machine, 2026-08.

Setup

  • Hardware: Apple M4 Pro (10 P + 4 E CPU cores, 20-core GPU, ~273 GB/s memory bandwidth).
  • Model: Qwen3-4B-Instruct-Q4_K_M.gguf (2.32 GiB, 4.02 B params, 36 layers, hd=128 decoupled, GQA 8 kv-heads, kv dim 1024, qwen3.context_length = 40960). The same file drives both engines.
  • llama.cpp: build 749f688fc (10547), GGML_METAL=ON; measured with llama-bench (-p <n> -n <m> -ngl 99|0 -r 3..5) and cross-checked with llama-cli.
  • minfer: target/release, graph path; MINFER_TIMING=1 (sample/forward split) and the MINFER_OP_PROFILE=1 op profiler (committed with this analysis — per-op host-encode table on the first submit, one line per later submit).
  • Protocol: every number measured alone (no concurrent benchmark processes — concurrent CPU load visibly corrupts Metal numbers, see §6).

Headline numbers

Scenariominferllama.cppRatio
Metal decode (tg, tok/s)75.4–75.978.6–79.7~95 %
Metal prefill, steady-state (marginal, tok/s)~800–1000~900≈ parity
Metal prefill, first request, 241 tok552 ms pre-fix → ~370 ms post-fix (default n_ctx 4096)313 ms1.76× → ~1.2× wall
CPU decode (tok/s, -t 8)1.1 → ~52–5863–68~60× → ~80–90 %
CPU prefill (tok/s)~1.1 → ~75~211~190× → ~2.8×

Re-measured by #54 (2026-10-06, 6b95763, macbook (macOS 27.0.1, Apple M4 Pro)): minfer bench -p 241 -n 128 -r 3 Qwen3-4B-Q4_K_M.gguf gives Metal decode 74.11 ± 0.19 tok/s and prefill pp241 729.83 ± 0.20 tok/s (~330 ms) — decode is within the run-to-run range above and prefill is at the table's steady-state band. llama.cpp was not re-run for #54, so the ratio columns remain the 2026-08 record. Full re-measured rows: docs/METAL_OPTIMIZATIONS.md §0.1.

Post-fix notes: the first-request gap was a KV-sizing policy issue (§2), now fixed — the 241-token first request drops from 552 ms to ~370 ms (241×1.15 + ~106 ms instead of + ~289 ms). The CPU gap (§3) was NEON/threading — now ~50× faster, near llama parity.

Bottom line: on the GPU there is no meaningful gap — decode is ~95 % of llama and prefill is at parity in steady state; the first-request difference was a KV-size policy issue, now fixed. The CPU gap was closed from ~60× to ~80–90 % of llama via NEON/SDOT kernels + a persistent row-parallel pool (§3). Remaining items: CPU prefill (~2.9×) and f16 KV cache. Details below.

1. Metal decode: no gap (75.9 vs 79.7 tok/s)

Per-token decomposition of minfer's forward() (13.09 ms/token wall):

ComponentTimeShare
GPU execution (submit wait)~12.5 ms~98 %
Host encode of ~250 kernel dispatches~0.19 ms~1.5 %
Sampling0.08 ms0.6 %
Logits download + bookkeepingremainder—
  • Both engines are memory-bandwidth bound: every decode token reads the full 2.49 GB of quantized weights. 2.49 GB / 12.5 ms ≈ 199 GB/s, which is ~85–90 % of the M4 Pro's practical peak (~220–230 GB/s of the 273 GB/s spec). There is nothing left on the table at the kernel level: minfer's kernel_q4_k_f32_matmul is a faithful port of llama's kernel_mul_mv_q4_K_f32_impl (NR0=2 / NSG=2, stride-4 super-block thread layout) — same work, same dispatch.
  • The residual ~4 % is measurement structure (llama-bench excludes logits download/sampling bookkeeping differently) plus micro-differences: KV cache f32 vs llama's f16, and per-token host overhead.
  • MINFER_OP_PROFILE confirms nothing pathological on the host side: over a full prefill + 8 decode submits, MatMul encode totals 0.585 ms, every other op sub-ms; the decode GPU time is stable at 12.3–13.0 ms/submit.

2. Metal prefill: parity in steady state; the first request pays a KV-size tax

Both engines use simdgroup GEMMs for prefill matmuls: minfer's 64×32-tile kernel_q4_k_mm_f32 (f32 activations) vs llama's 64×64 kernel_mul_mm_q4_K_f32 (f16 activations). Both land at ~75–80 % of the GPU's fp32 peak — measured marginal throughput:

Promptminfer first-submit GPU (pre-fix)llama (llama-bench)
pp16 / 5 tok~283 ms (5 tok)238 tok/s (67 ms)
pp120 / 121 tok390 ms—
pp240 / 241 tok552 ms771 tok/s (313 ms)
pp480 / 481 tok851 ms821 tok/s (585 ms)
marginal (slope)~1.0–1.25 ms/tok~1.1–1.2 ms/tok

→ steady-state prefill ≈ 800–1000 tok/s (minfer) vs ~900 tok/s (llama).

The first-request tax (the 425-vs-771 wall-clock gap) — FIXED

minfer's single-shot CLI used to size the KV regions at max_seq_len = 40960 (Qwen3Graph::forward passed model.hparams.max_seq_len straight through, and GraphParams.cparams.n_ctx sizes the persistent per-layer KV regions):

36 layers × 2 regions × 40960 × 1024 × 4 B (f32) = 12.1 GB of Metal shared buffers

The first GPU command buffer then paid ~275 ms — a one-time Metal driver cost that scales with the total KV buffer bytes (measured first-submit GPU time vs n_ctx — same process, --cnv):

n_ctxfirst-submit GPU time
2048 (0.6 GB)87 ms
4096 (1.2 GB)106 ms
40960 (12.1 GB)289 ms

After the first submit the cost never recurs (decode submits are 12.3–13.0 ms each; a second, larger prefill in the same process — conv turn 2 — shows no fixed cost either). Note the cost is not memory commitment: peak RSS is identical (~2.1 GB) at n_ctx 4096 and 40960 — it is the Metal driver's first-use setup (buffer VA/registration) for the oversized regions, so it is strictly a one-time latency + address-space tax, never a resident-memory one.

llama.cpp avoids showing it because (a) its KV cache is sized by n_ctx (default 4096 → ~1 GB) and (b) it runs a warmup pass at model load that absorbs the one-time cost before the first real request.

Secondary consequence (pre-fix): the single-shot process held 12.1 GB of f32 KV buffers it never used (a 10-token prompt), while --cnv / --n-ctx correctly used the CLI value (conv turn-1 at n_ctx 4096: 106 ms; at 40960: 289 ms — the difference was entirely this sizing).

Fix (this commit): ModelDef::forward now takes n_ctx; the single-shot CLI passes --n-ctx (default 4096) and both Qwen2Graph::forward / Qwen3Graph::forward clamp it to the model's max_seq_len (llama.cpp clamps the same way). Result: default first-submit 289 ms → ~106–109 ms, the 12.1 GB address-space allocation disappears, and the 5-token prefill wall drops from 0.31 s to 0.11 s. Greedy output is byte-identical (KV content is unchanged — only buffer sizes). --n-ctx now also documents as applying to single-shot / --cnv, not just the server.

3. CPU: the real gap — ~60× decode, ~190× prefill — FIXED

Before (baseline): minfer CPU decode 1.1 tok/s (936 ms/token, 99.4 % in forward()); llama.cpp CPU decode 61–68 tok/s (best at -t 8; -t 10 → 60.8; -t 14 including the 4 E-cores collapses to 15.9 tok/s — E-cores hurt).

The gap decomposed into two compounding causes, both structural:

  1. Single-threaded (×~10): minfer's CPU backend had no threading. llama.cpp uses 8–10 P-core threads.
  2. Scalar dot products on ARM (×~5.7 per core): minfer's quantized dot kernels were AVX2 (x86) with a scalar fallback — no SIMD on aarch64.

After (this commit): CPU decode 1.1 → ~52–58 tok/s at -t 8 (~50×, ~80–90 % of llama's 63–68), prefill ~1.1 → ~75 tok/s (241-token). What was done:

  1. NEON + SDOT dot kernels (src/quants.rs): all eight quantized dot products got aarch64 fast paths using the ARMv8.2 sdot instruction (16 MACs/instr, emitted via inline asm — vdotq_s32 is unstable in std::arch). Bit-exact with the scalar kernels (int32 accumulation is exact; per-block float ops kept in identical order — verified by unit tests on random data).
  2. Q8_K activations for K-quant matmuls (src/block.rs, src/quants.rs, src/kernel.rs): llama.cpp's activation format for Q4_K/Q5_K/Q6_K weights — 256-element blocks with precomputed int16 per-subblock sums (bsums), so the dots never re-reduce the activation and apply one scale per 256 elements instead of 8. The kernels were restructured to llama's shape (SDOT accumulation with scales applied via vmlaq_n_s32 before the horizontal reduce: one vaddvq per superblock instead of 8). Also a NEON q8_K quantizer (bit-exact with the scalar one).
  3. Persistent CPU thread pool + row-parallel matmuls (src/kernel.rs): per-call thread::scope spawning costs ~170 µs (measured) — too slow for ~250 matmuls/token — so a persistent pool (atomic generation handoff + spin/yield workers, main thread participates as the last worker) dispatches each matmul's rows in ~1–3 µs. Output is bit-identical to single-threaded (each row is computed by exactly one worker with the identical code path). A generic par_for extends the same pool to other ops (the attention is parallelized over heads).
  4. NEON vector ops (src/vec_ops.rs): vec_dot_f32, vec_muladd_f32, vec_scale_f32, vec_add_f32, and an in-place softmax (fast polynomial exp) — the CPU attention/norm path was fully scalar before.
  5. CLI -t/--threads <N> (default: macOS P-core count via hw.perflevel0.logicalcpu, 10 on M4 Pro; measured best is 8 for the 4B — E-cores hurt, matching llama). MINFER_NO_NEON=1 forces the scalar kernels for A/B.

Correctness: NEON vs scalar, threaded vs single-threaded, and the whole NEON+threads pipeline vs the scalar single-thread baseline are all byte-identical on the greedy output; new unit tests cover the NEON kernels against the scalar references on random data. The CPU activation quantization change (q8_0 → q8_K for K-quants) shifts logits slightly, so CPU-vs-Metal greedy text can now flip at near-tie boundaries (both are coherent); Metal output is unchanged.

CPU prefill (241 tok): ~75 tok/s vs llama 211 — the remaining prefill gap is the serial activation quantization (NEON now) plus no token-level parallelism; the GPU path remains the fast prefill route (~800+ tok/s).

4. Secondary: KV cache f32 vs f16 (long-context decode)

Decode attention re-reads the whole KV: 36 × 4 kv-heads × 128 hd × 2 (K+V) × 4 B × C = 147 KB × C per token. At 191 GB/s:

Contextminfer (f32)llama (f16)decode slowdown vs llama
2048+1.6 ms/tok+0.8 ms/tok~6 %
4096+3.2 ms/tok+1.6 ms/tok~13 %
8192+6.3 ms/tok+3.2 ms/tok~24 %

Also halves the KV read traffic and the §2 first-submit tax. (llama.cpp's default kv_cache_type is f16; minfer stores f32.)

5. Recommendations (ordered by measured impact)

  1. Fix single-shot KV sizing DONE (commit d5b8023) — ModelDef::forward takes n_ctx; single-shot passes --n-ctx (default 4096), clamped to the model's max_seq_len. First-request latency 289 → ~106 ms, 12.1 GB address-space allocation gone; greedy output byte-identical.
  2. CPU backend: NEON dot products + threading DONE (this commit) — SDOT kernels + Q8_K activations + persistent row-parallel pool + NEON vec ops + -t/--threads: CPU decode 1.1 → ~52–58 tok/s (~50×, ~80–90 % of llama), CPU prefill ~1.1 → ~75 tok/s. Details in §3.
  3. KV cache f16/bf16: halves KV traffic and the §2 first-submit tax; worth ~6–24 % on long-context decode.
  4. CPU prefill (optional): token-level parallelism or a GEMM path to close the 75-vs-211 gap; the GPU path already handles prefill at ~800+ tok/s.

6. Methodology notes / pitfalls seen

  • Never benchmark two engines concurrently: an earlier llama-bench Metal run executed while the CPU benchmark was running measured 27.5 tok/s (vs 79.7 alone) with 17 % variance. Metal decode still needs CPU cores to encode ~250 kernel dispatches per token.
  • llama-bench's backend column lists available devices ("MTL,BLAS") even for -ngl 0 — the -ngl value is the authoritative offload control.
  • llama-cli perf lines: "Prompt: … t/s" = prefill, "Generation: … t/s" = decode.
  • minfer's first prefill in a fresh process includes the KV-size first-submit tax (§2; post-fix ~106 ms at default n_ctx 4096); --cnv (n_ctx 4096) is the clean way to measure steady-state prefill, or subtract the measured tax from the single-shot total.
  • Raw instrumentation: MINFER_OP_PROFILE=1 (op-encode table + per-submit GPU time), MINFER_TIMING=1 (sample vs forward per token).

Appendix: raw measurements

WhatValue
minfer Metal decode, n=6475.9 tok/s (forward 13.09 ms/tok, GPU 12.5)
minfer Metal decode, n=1675.4 tok/s
minfer CPU decode, n=641.1 tok/s (forward 936 ms/tok)
llama-bench Metal tg6479.73 ± 0.14 tok/s
llama-cli Metal Generation, n=4878.6 tok/s
llama-bench CPU tg32/64 (t=10)60.74 ± 6.4 / 66.71 ± 3.2 tok/s
llama-cli CPU Generation (t=8 / 10 / 14)63–69 / 60.8 / 15.9 tok/s
minfer CPU decode post-fix (-t 1 / 8 / 10)10.4 / 48–58 / 50.9 tok/s
minfer CPU prefill 241 tok post-fix (-t 8)~70–75 tok/s
llama-bench CPU pp240 (t=8)202.7 ± 1.3 tok/s
llama-bench Metal pp16/240/480238 / 771 / 821 tok/s
llama-bench CPU pp240211 tok/s
minfer prefill GPU (first submit, pre-fix)5 tok: 283 ms · 121: 390 ms · 241: 552 ms · 481: 851 ms
minfer prefill GPU vs n_ctx (10 tok, pre-fix)2048: 87 ms · 4096: 106 ms · 40960: 289 ms
minfer first-submit, post-fix (n_ctx 4096 / 2048 / 8192 / 65536→clamped)109 / 99 / 127 / 308 ms
minfer first-submit RSS (n_ctx 4096 vs 40960)2.14 vs 2.12 GB — tax is not resident memory
minfer conv turn-2 delta prefill (n_ctx 4096)16 tok in 12.3 ms (no tax)
minfer 5-token prefill wall (pre → post fix)0.31 s → 0.11 s
greedy 64-token output, pre vs post fixbyte-identical

CUDA & GPU Technology Primer — every technique minfer uses, explained

This is the reference for the CUDA and NVIDIA-GPU technologies the minfer inference engine actually uses — what each one is, how minfer uses it, where it lives in the code, and what the optimization campaign (step docs cuda_optimization_steps/01–80) learned about it. It is a teaching document first: every concept is explained before it is used. Deep-dive narratives stay in the step docs; this file is the map.

0. How to read this document — the three layers (and which layer each term lives in)

The campaign's vocabulary mixes three different technical domains. When a term confuses you, first ask which layer it belongs to:

LayerDomainTerms that live hereDecided by
AlgorithmLLM inference algorithms (the speculative-decoding literature: Leviathan/Chen 2023; llama.cpp's draft-simple)d (draft length, e.g. "d=2"), acceptance rate p, E[a] = Σp^i, break-evenprobability & decision math — hardware-independent
Performance modelcomputer-architecture performance analysis (the roofline model, arithmetic intensity)"amortization" (spreading fixed costs over more work), memory-bound vs compute-bound, bytes-per-tokenmeasured bytes/FLOPs against hardware ceilings
Micro-architecture (kernel implementation)GPU GEMM kernel engineering (tiling, CUTLASS-style vocabulary)tile shape, "tile-regime", wave quantization, occupancy, __launch_bounds__, coalescingthe kernel's tiling configuration vs the shape it is launched with

A concrete chain from the D5-0 gate (doc 80) showing all three at once:

d=2 (algorithm layer: draft 2 tokens per round)
  → the verify forward is a batched nt=3 decode step (a SHAPE)
    → which tile-regime does M=3 land in? (micro-architecture layer)
      → 2.28x (interpolation) or 2.7x (same-tile-as-M=4)? = the amortization
        (performance-model layer: fixed weight-traffic cost spread over rows)
        → ≥ 2.5x needed for break-even at measured p≈0.68
          → back up: is the algorithm parameter d=2 worth it?

No single layer could answer the gate question — the algorithm parameter's fate was decided by a micro-architectural regime measurement. That cross-layer interaction is exactly what the campaign's step docs record, and what this primer maps.

Full taxonomy: a corpus-wide audit extended these three layers to seven (adding Platform & tooling, Engine architecture, Methodology) and classified every term and formula in the campaign docs — see GLOSSARY.md.

Code surfaces referenced throughout:

FileLinesRole
src/cuda.rs~5.4kDevice layer: hand-written CUDA bindings, context/state, weight upload & registration, buffer concat, KV type management
src/cuda/kernels/*.cu + common.cuh~10kEvery CUDA kernel (102 __global__ functions) + their launch_* C-shim wrappers, in 17 translation units since #263, compiled by nvcc into libcuda_kernels.a
src/graph/cuda_backend.rs~6.1kThe Backend trait implementation: buffer pools, per-node dispatch, CUDA Graph capture/replay, synchronization
build.rs—nvcc orchestration: compiler discovery, host-compiler pinning, arch detection, SASS/PTX emission, cudart linking

1. The platform — what "GB10 / DGX Spark" means for this code

minfer's CUDA campaign ran on an NVIDIA GB10 (DGX Spark): a Grace-Blackwell superchip where a 20-core Arm CPU (the same machine runs minfer's NEON+SDOT CPU backend) and a Blackwell-class GPU share one coherent memory pool (~128 GB). Practical consequences that show up everywhere in the record:

  • The GPU reports a Blackwell-class compute capability — the build covers sm_70 … sm_121 (see §3); the campaign machine (driver 580.173.02) uses one of the newest targets in that list.
  • Unified memory does NOT mean minfer should use cudaMallocManaged — it doesn't. All device memory is plain cudaMalloc pools (§5); "resident weights" (§6.1) is the deliberate design: upload once, keep on the GPU, never page across the link.
  • CPU and GPU measurements interleave on one machine — which is why every campaign number is a same-window interleaved A/B median (co-tenant load once moved a kernel's measurement by 2.7×; step doc 15).

2. The software stack — how Rust reaches the GPU

minfer uses no CUDA Rust ecosystem crates (no bindgen, no cuda-rust, no cudarc). The stack is four layers, each thin:

src/graph/scheduler.rs          Rust: picks the backend per node, splits the graph
src/graph/cuda_backend.rs       Rust: buffer pools, dispatch, capture/replay, sync
        │  extern "C" — hand-declared symbols (cuda.rs + backend.rs)
        ▼
libcuda_kernels.a               nvcc-compiled from src/cuda/kernels/*.cu:
  launch_q6_k_q8_mmvq(...)        every kernel + a C `launch_*` shim per kernel
        │                         (the shim contains the <<<grid,block>>> syntax,
        ▼                          which is CUDA-C++ only)
libcudart.so / libcudart_static.a   the CUDA runtime API (cudaMalloc, cudaGraph…)
        ▼
libcuda.so.1                    the kernel-mode driver (preloaded via dlopen,
                                 see §2.1)

Why a C shim for launches: Rust has no <<<>>> launch syntax; rather than generate it, each kernel gets a ~20-line C wrapper (launch_<name>(...)) that takes plain pointers and ints, computes the launch configuration, and calls the kernel. Rust declares these as extern "C" blocks (cuda.rs) — the entire binding layer is hand-written and auditable (114 launch_* declarations).

Why dlopen libcuda first (cuda.rs): libcudart itself dlopens the driver library libcuda.so.1. Preloading it from the well-known locations (RTLD_NOW | RTLD_GLOBAL) guarantees cudart reuses the already-loaded driver instead of failing on a stub-dependency resolution — a robustness detail that matters on non-standard installs.

Static vs shared runtime: by default minfer links the shared libcudart.so; the cuda_static cargo feature links libcudart_static.a instead (plus pthread/dl) so the binary carries no cudart dependency — at the cost of a larger binary. build.rs wires both link modes.

3. nvcc and the build system (build.rs)

nvcc is NVIDIA's compiler driver: it compiles .cu files by splitting them into host C++ (compiled by a host compiler) and device code (compiled by NVVM/PTXAS into GPU machine code). minfer's build must orchestrate it because plain cargo build never touches nvcc — the CUDA compile only runs when the cuda feature is requested (an early footgun: any build with an nvcc on PATH used to attempt the compile; build.rs now gates strictly on the feature).

Key mechanics, all in build.rs:

  • Compiler discovery: find_nvcc() / find_cuda_home() locate the toolkit; CUDA_HOME style overrides apply.
  • Host-compiler pinning (-ccbin): nvcc inherits the first cc/g++ on PATH as its host compiler and hard-fails when that compiler is newer than the toolkit supports. build.rs only pins -ccbin when the default is actually rejected — otherwise the environment's choice is left untouched (documented in docs/BUILD.md).
  • Architecture detection (detect_archs): GPU machine code is architecture-specific. build.rs probes nvcc with a dummy .cu compiled for each candidate sm_70, 72, 75, 80, 86, 89, 90, 100, 103, 110, 120, 121 and keeps the ones the toolkit accepts.
  • SASS + backward-JIT PTX: for each detected arch it emits arch=compute_<a>,code=sm_<a> — i.e. SASS (the arch-specific machine code, fastest launch, no JIT) — plus a compute_70/72 PTX (NVIDIA's virtual-ISA assembly) so older GPUs JIT-forward to a working image even without exact SASS.
  • Output: one static library libcuda_kernels.a per build, exposing the launch_* shims; the Rust linker is pointed at it and at the cudart link path.

Terminology used in the campaign: SASS = the actual per-architecture instruction stream (read with cuobjdump -sass / nvdisasm; step doc 30's "opcode census" counted LDG/IMAD/HMMA instructions there). PTX = the virtual assembly JIT'd at load time. sm_XX / compute_XX = the architecture generations.

4. The CUDA programming model — the vocabulary every kernel section uses

Every term in this section also appears, classified by layer, in the corpus-wide GLOSSARY.md (L4/L5).

  • Host / device: the CPU side (Rust) and the GPU side (.cu). They have separate address spaces; all data crossing the boundary goes through explicit copies (cudaMemcpy) or pinned-memory staging (§5.4).
  • Kernel (__global__): a function executed by many GPU threads at once. minfer has ~65 of them (the inventory in §6).
  • Grid / block / warp / thread: a launch specifies a grid of blocks (each block scheduled onto one SM — streaming multiprocessor, the GPU's compute unit); a block contains up to 1024 threads; threads execute in 32-thread warps (the SIMD unit — every thread in a warp executes the same instruction). minfer's kernels mostly use 256-thread blocks (__launch_bounds__(256), §4.1) or 128 (rms pad40).
  • __launch_bounds__(max_threads, min_blocks): a compiler hint capping register usage so a block/thread combination is guaranteed launchable — D4-2 tested __launch_bounds__(256,6) explicitly (48→40 regs, but +40 B stack spill → measured −0.75% → reverted). Lesson: launch bounds trade registers for spill; measure, don't assume.
  • Streams: an ordered queue of GPU work. minfer creates one stream (cudaStreamCreate) per backend; all launches and copies are stream-ordered so synchronize() semantics stay simple (§7).
  • Occupancy & waves: an SM can host a limited number of concurrent blocks (bounded by registers/shared memory/threads). Occupancy = how much of that capacity is used. The GPU runs blocks in waves — total blocks ÷ (blocks resident per SM × SM count). Wave quantization is a recurring campaign villain: a grid of 1.14 waves runs as 2 waves (the second nearly empty) — e.g. D4-4's fused-FFN probe grid of 1.5 waves explained its +28.2% loss, and M=1 decode GEMMs at 0.14 waves are the reason batched decode (spec-decode's verify, doc 80) amortizes so well.
  • Occupancy ≠ performance: r24's scheduling-structure ladder showed grid rearrangements that "improved" occupancy yet measured flat — the +1.5% whole-prefill landing bar was calibrated in that round.

5. Memory — hierarchy, movement, and the rules the campaign burned in

5.1 The hierarchy

LevelScopeLatency classminfer use
Registersper-thread~0 cyclesaccumulator tiles (float acc[8]), the ≤85-regs ceiling that gated r59's wave re-tile
Shared memoryper-block, user-managed~30 cyclesstaging tiles for wmma GEMM (8m), swizzle space (r22), smem over-cap silently killed r8's first wide tile (attr-set + launch both failed — "suspiciously fast" readings must be checked against resource caps)
L1 / L2 cacheper-SM / chip~200 / ~400 cyclesweight L2 residency tested in r19 (see below)
DRAM (unified pool)device~600+ cyclesweights, KV, activations; byte bandwidth through this level is the decode wall (D1–D3 attribution)

5.2 Coalescing and vectorized access

A warp's 32 threads should touch contiguous memory so the hardware folds their accesses into the minimum number of 128-byte transactions (coalescing). minfer's kernels load weights as uint4 (16-byte vector loads, the widest LDG) along unit-aligned rows; the dpl repack (D4-4) exists precisely to make this true end-to-end — the stock Q6_K row layout wasted 14 of every 224 bytes on padding (84% useful), and the dense split-plane repack ([ql][qh][sc][d] planes, 210 B content, 16 B-aligned stride) turned padding bytes into bandwidth: −17.1% content traffic on ffn_down, +5.5% 14B tg128.

5.3 Cache-control intrinsics (and their negative results)

  • __ldg (read-only, non-coherent load path): r19 tested marking the entire weight stream __ldg — imperceptible; then a persisting L2 window (cudaAccessPolicyWindow pinning weights in L2) — catastrophic. Lesson recorded in doc 24: L2 residency management is the wrong lever when the working set is streamed once per token anyway.
  • cp.async (asynchronous shared-memory copy, Ampere+): the 8m prefill GEMM stages weight/activation tiles into shared memory with cp.async so the copy pipeline overlaps the mma pipeline; the r21 "coalesced block-linear A staging" variant (removing the XOR swizzle in favor of linear layout) measured as a stall-mass wash → reverted (the swizzle stays).

5.4 Pinned host memory (page-locked staging)

Device↔host copies through pageable memory bounce via an internal driver staging buffer and serialize. minfer instead allocates pinned host buffers (cudaHostAlloc) — one grow-on-demand staging buffer (R3-A2) capped at 128 MB per split — and copies results into it device→host directly. MINFER_NO_PINNED_READBACK=1 reverts to plain cudaMemcpy. The capture-readback path (§8) reads logits from this pinned buffer immediately after each captured node's launch.

5.5 Buffer pools (no cudaMalloc in the hot path)

cudaMalloc is expensive (driver round-trip). CudaBackend owns buffer pools: allocation hands out pool slots (bumped pool_gen tracks generations), the graph allocator's liveness analysis (AGENTS §"Compute Graph" rules) reuses slots, and input buffers are never freed mid-graph. All pools are plain device memory — no unified/managed allocations anywhere.

6. The kernel inventory — every __global__ family, what it does and why

6.1 Weight-resident model (Phase 7 / 8p)

At load, every weight tensor is uploaded once and stays on the GPU for the process lifetime (cuda.rs registers them; execution gates on all-weights-registered). The 8p round added a persistent f16 weight cache (+ fused dequant-in-GEMM) for shapes where dequant-on-the-fly beats quantized GEMM. This is the foundation of everything else: decode has zero host↔device weight traffic.

6.2 The matmul families (the campaign's main battlefield)

Kernel familyCompute primitiveUsed forStep docs
q*_f32_matmul (q4_0, q4_1, q5_0, q5_1, q4_k, q5_k, q6_k)CUDA cores, per-thread dot over dequantized weightspre-Phase-7 / fallback paths, GPU reads f32 activations—
gemm_f16_nt_kernel_t (+ gemm_qb_nt)Tensor cores via nvcuda::wmma (f16 in, f32 accumulate)prefill GEMM on the f16 weight cache (8m) — 30.7→1204 tok/s (39×)02, 05
mmq_nt_kernel, mmq_raw_nt/wide/nb/bt(_q6k)integer dp4a/imad over int8-quantized activations × quantized weightsprefill MMQ (R1 parity-first, opt-in → default), the r5–r59 line — 8.1× arc. "BT" names the raw-byte BT-style kernel variant of this family (introduced r38 as the q6_K BT-style kernel; mmq_raw_nb_bt_kernel) — the form whose batched (nt>1) executions the D series measured; "BT-MMQ 2.7× at nt=4" (doc 80's gate anchor) refers to this family running the multi-token batch08, 13–31, 33–40, 67, 76
q*_q8_mmvq, v2, v2_pf, _dpldp4a (4-way int8 dot) per weight rowdecode matrix-vector (one token): q4_K/q5_K/q6_K; v2 = D3b bitwise re-tile, pf = D4-2 padded-form dispatch, dpl = D4-4 dense split-plane06, 67, 74, 76
dequant_*_f16per-block dequant to f16feeds the f16 GEMM path05
embed_rows_*, gather_rows_f32row-gatherembedding lookup on GPU—

The two activation regimes (a core convention, AGENTS §Core Conventions): CPU quantizes activations to Q8_0 on the fly (Q8_K for K-quant weights) and runs integer math; GPU prefill MMQ also uses int8 activations, but the GPU's f32/wmma paths read f32 activations directly. CPU-vs-GPU logits therefore differ by design — each path is verified against its own reference (§9).

Quantized weight formats (GGUF v3, src/block.rs = the repr(C) ggml-common layout): super-block schemes where a group of 256 values carries scales/mins at coarser granularity (Q4_K: 8 sub-blocks of 32 + a super-scale per 256; Q6_K: 16-byte scales + 2-bit high bits packed with the 6-bit lows). The GPU kernels unpack these layouts directly from the byte stream — no pre-dequantization — which is why the dpl layout experiment (a different byte arrangement of the same values) could be bitwise-safe: same values, same accumulation order, only the memory map changed.

6.3 Attention kernels

  • fa_prefill_f16kv — the 8n FlashAttention-style tiled prefill attention: online softmax (rescaling the running max/sum instead of materializing the score matrix), the O accumulator held in registers, 256 B P-stride to avoid a score-clobber race. 20× over the naive path.
  • gqa_attn_f32 (h4w = 4-warp variant) and gqa_attn_f32_f16kv — decode attention with GQA (grouped-query: N query heads share KV heads; the q-head batching experiment D3-6 lost to the 5× L2 re-read it created).
  • gqa_attn_split_partial + gqa_attn_split_combine — flash-decoding split-K (8d, later R4's dim-parallel rewrite): split the KV sequence across blocks, each computes partial (num, den, max) triples, a combine kernel folds them. The D1 attribution found split_partial was 100% of the KV-growth wall (memory-LATENCY-bound, 76.5% long_scoreboard) — the number that authorized D2/D3.

6.4 Element-wise and fused epilogue kernels

rms_norm_* (with _quant fused output-quantize variants, _nw no-wide, _pad40), silu_f32, swiglu_* (gate×up + SiLU), rope_f32, add_bias_f32, add_f32, mul_f32, store_kv_f32/f16, quantize_q8_0 (_pad40, _t transposed), f32_bits_to_i32 (positions as bits — AGENTS rule 4), convert_f32_f16. The fused epilogues are the D3-5/D3-8 story: attn_bias_rope_store_f32 replaced a 7-launch chain (bias×3 + rope×2 + store×2) with one launch (−310 launches/step), and producer-fused A-quantize (MINFER_MMQ_A_FUSE) folds activation quantization into the producer kernels (quantize launches −78%).

Why fusion matters on this GPU: decode chains are launch-overhead-bound (2 µs/graph-gap scale); every kernel boundary costs a launch + memory round-trip through L2. The fusion ledger lives in step docs 70, 73, 76.

7. Synchronization discipline (GPU Safety, docs/GPU_SAFETY.md)

The hard rules that bound every kernel addition:

  • submit() waits bounded + checks status — never an unbounded block on the GPU; failures abort with actual values, never silently.
  • No early return past a threadgroup_barrier() (Metal phrasing; the CUDA analogue is __syncthreads()) — divergent barrier exits deadlock or corrupt.
  • Kernel-invariant violations return Err from execute_node — never a silent CPU fallback mid-graph; backend assignment is decided at build time (supports_op), and the graph records GPU participation in CParams.gpu.
  • Never host-copy a GPU-pending buffer — the Phase-3 KV-corruption bug class; copies cross-backend happen only at scheduler split boundaries.
  • synchronize() (backend.rs) is the one choke point: stream-ordered work is waited with a bounded loop and the status is checked.

8. CUDA Graphs — capture once, replay many (Phase 7d)

Decode = the same ~13-kernel chain re-launched hundreds of times; per-launch CPU overhead (~2–7 µs) is pure tax. CUDA Graphs fix this: a sequence of launches is recorded (captured) into a graph object, then replayed with one call — driver-side launch overhead collapses.

minfer's machinery (cuda_backend.rs):

  • Capture eligibility: a (graph uid, split node-range) is captured after a 3rd consecutive direct-launch warmup (llama.cpp's warmup-twice heuristic adapted); captures are valid only for the pool_gen they were captured at — any pool re-allocation invalidates them (the buffers the graph's pointers reference must be the same memory).
  • Prefill capture is default-on now (MINFER_NO_PREFILL_CAPTURE=1 opts out; replay at real-prefill scale validated it).
  • MINFER_NO_CUDA_GRAPH=1 force-disables (the A/B control used by every graph-adjacent step doc).
  • Why replay is safe in minfer's design: AGENTS rule 1 — "KV positions are data, not structure". The graph's kernel arguments (pointers, dims) are identical every step; only buffer contents (positions, tokens) change, and those are written into the same device buffers before replay. This is also exactly the property the spec-decode verify batch needs (doc 80: fixed d ⇒ one captured verify graph, positions refilled host-side per round).
  • Validation: MINFER_GRAPH_DUMP replays capture real-data dumps; the standard gate is replay-dump byte-identity vs eager launches.

9. Numerics — precision, accumulation, and the verification gates

  • f32 accumulation everywhere in tensor-core paths: wmma accumulators are f32 (mma_sync with f32 D/C operands); r15's probe of f16 accumulation was a dead end (numerics + throughput both lost).
  • Tolerance gate vs bitwise: the campaign's two-tier standard. New math gets a tolerance gate (CPU reference in double, 0.05 abs with adversarial outliers, per step doc 68's calibration) plus bitwise-identity tests whenever a rewrite claims "same math" (byte-identical dumps, greedy outputs byte-identical across seeds). The rounding classes are understood, not hand-waved: D4-2's B0 bug (7B dropped 13.5% of down-proj rows) showed up first as max|Δ| 0.254 first-step logits at a documented rounding class.
  • The dump instruments (--features debug_dump, MINFER_DUMP_DIR, MINFER_GRAPH_DUMP, MINFER_TRACE): per-node real-data dumps that make "bitwise vs eager" and "KV rows identical" checkable claims. Pool-slot recycling makes some informational dumps binary-layout-dependent — a documented instrument limitation, not a bug.

10. Profiling & forensics — how the campaign knows what it knows

  • nsys (Nsight Systems): timeline traces → launch counts per decode step (the D3-8 −310 launches/step claim) and wall decomposition.
  • ncu (Nsight Compute): per-kernel counters → the D1 verdict ("memory-LATENCY-bound, 76.5% long_scoreboard" = threads stalled waiting on DRAM loads) and the D4-3 llama-bench artifact bust (a decode kernel loading a constant 5.4 MB regardless of context = not real work).
  • SASS census (step doc 30): counting actual machine instructions to explain instruction-mix changes.
  • Counter forensics (doc 18): reconciling achieved bytes/FLOPs against hardware ceilings — a claim survives only if the counters, the SASS, and a reductio all agree.

11. The env-gate inventory (A/B controls)

Every lever is env-gated so the same binary runs both sides of an A/B (never a rebuild between A and B — the campaign's method):

GateMeaning
MINFER_DISABLE_CUDA=1force CPU (backend selection off)
MINFER_NO_CUDA_GRAPH=1disable capture/replay (eager launches)
MINFER_NO_PREFILL_CAPTURE=1disable prefill-graph capture
MINFER_NO_PINNED_READBACK=1revert to plain cudaMemcpy readback
MINFER_NO_FUSE_QKV / MINFER_NO_FUSE_FFNrevert the decode fusions
MINFER_MMQ_A_FUSEproducer-fused activation quantize
MINFER_Q6K_DPL=0opt out of the dpl split-plane layout
MINFER_Q6K_PF=0revert v2_pf padded-form dispatch
MINFER_PDL(closed NO-GO, D4-4) programmatic dependent launch
MINFER_NO_MPS=1Metal: force CPU (macOS paths)
MINFER_NO_NEON=1CPU (aarch64): force scalar
MINFER_NO_AVX2=1CPU (x86): force the whole quants AVX2 layer scalar
MINFER_NO_AVX512=1CPU (x86): drop just the AVX-512/VNNI K-quant dots to AVX2
MINFER_GRAPH_DUMP / MINFER_TRACE / MINFER_DUMP_DIRverification instruments (§9)

PDL (programmatic dependent launch — letting the next kernel launch before the previous finishes via cudaLaunchAttributeProgrammaticStreamSerialization

  • cudaGridDependencySynchronize()): the one CUDA feature the campaign integrated fully and then reverted — in-situ A/B read −2.6%/−1.8% on 14B because early-launch co-residency taxes compute-tail kernels more than the graph-gap it recovers (doc 76). A standing warning against co-residency tricks on this workload.

12. Where to go next

  • Build specifics (toolkit install, ccbin, cudart modes): docs/BUILD.md
  • GPU safety rules: docs/GPU_SAFETY.md
  • The campaign's full history, one doc per step: docs/cuda_optimization_steps/README.md
  • The compute graph these kernels serve: docs/COMPUTE-GRAPH-DESIGN.md
  • The spec-decode consumer of the nt>1 regime: docs/SPECULATIVE-DECODING-PLAN.md

Glossary — every term and formula in the CUDA campaign docs, by layer

This is the consolidated vocabulary of the CUDA campaign corpus (CUDA_OPTIMIZATION.md, CUDA-TECH-PRIMER.md, the 80 step docs in cuda_optimization_steps/, the four analysis/plan docs, and the build/safety references). Every entry is corpus-verified: it appears in at least one of those documents.

0. The layer taxonomy

The primer's §0 originally introduced three layers (Algorithm / Performance model / Micro-architecture) to explain the D5-0 gate chain. The full-corpus audit showed that roughly half of the vocabulary does not fit any of those three, so the taxonomy is extended to seven layers:

LayerDomainTypical questions it answers
L1 AlgorithmLLM inference algorithms & their mathWhat does the model compute? What is the optimal d?
L2 Numerics & data formatsquant formats, byte layouts, rounding, tolerancesHow are numbers represented and how wrong may they be?
L3 Performance modelroofline, bandwidth, amortization, break-evenHow fast should it go, and why isn't it?
L4 Micro-architecturetiling, tensor cores, SASS, shared memory, schedulingHow does the kernel actually use the machine?
L5 Platform & toolingCUDA API/runtime, nvcc/PTX/SASS toolchain, profilers, env gatesWhat do we drive the hardware with, and how do we look inside?
L6 Engine architectureminfer's compute graph, backends, allocator, fusion, safety rulesHow is the engine itself structured?
L7 Methodologymeasurement discipline, parity gates, forensics, doc conventionsHow do we know a number is real?

The D5-0 gate chain still illustrates why layers matter — the same terms now map to L1/L4/L3 explicitly:

d=2 (L1) → verify = batched nt=3 decode step (shape) → which tile-regime (L4)
  → amortization ≥ 2.5x? (L3) → break-even at p≈0.68 (L3) → go/no-go on d=2 (L1)

Source-doc tags used below

TagDocument
PRIMERCUDA-TECH-PRIMER.md
HUBCUDA_OPTIMIZATION.md (campaign hub + live status)
STEPScuda_optimization_steps/NN-*.md (numbered step doc)
MMQLLAMA-CPP-MMQ-ANALYSIS.md
SPECLLAMA-CPP-SPECULATIVE-ANALYSIS.md
D5SPECULATIVE-DECODING-PLAN.md
GRAPHCOMPUTE-GRAPH-DESIGN.md
BACKENDCUDA-BACKEND-DESIGN.md
BUILDBUILD.md
SAFETYGPU_SAFETY.md
METALMETAL_OPTIMIZATIONS.md (its §5.6 is a doc-local symbol glossary)

L1 — Algorithm: LLM inference algorithms & math

Term / FormulaOne-line meaningSource
prefillOne forward pass over the whole prompt, batched, compute-bound.PRIMER, STEPS 20+
decodeOne token per step (nt=1), memory-bound: streams all weights every token.PRIMER, STEPS
logitsThe [nt][vocab] output of the final matmul; input to sampling.GRAPH, STEPS
greedy decodingAlways take argmax of logits — deterministic, used for parity gates.HUB, STEPS
top-k / top-p / temperature / repeat-penaltyThe sampler chain in sampler.rs, defaults matching llama.cpp (0.8/0.95/1.1).GRAPH
BPE tokenizerByte-pair-encoding tokenizer parsed from GGUF metadata; special-token match matters (DeepSeek-R1 distill).GRAPH
RoPERotary position embedding applied to Q/K in-place, fused into decode QKV.PRIMER, GRAPH
RMSNormRoot-mean-square layer norm; fused with residual add in decode kernels.PRIMER, GRAPH
SiLU / SwiGLUActivation and gated FFN; SwiGLU is fused (gate+up concat + silu-mul).PRIMER, GRAPH
softmaxAttention normalization; FA prefill uses tiled online softmax, FAP2 keeps it register-resident.STEPS 46-48
GQAGrouped-query attention: fewer KV heads than Q heads; decode attention must replicate K/V.STEPS 26
KV cachePersistent per-layer K/V tensors grown token by token; the "positions as data" rule keeps topology fixed.GRAPH, PRIMER
n_pastNumber of cached tokens; must never appear in graph topology, only as data.GRAPH
draft modelSmall model (0.5B) proposing d tokens the big model verifies.D5, SPEC
draft length dTokens drafted per speculative round (d=2/4/8 measured).D5
acceptance rate pProbability the target accepts a drafted token; measured p≈0.68–0.70 (llama speculative-simple, greedy).D5, STEPS 80
E[a] = Σ_{i=1..d} p^iExpected tokens accepted per round under independence.D5, PRIMER
verify passOne batched forward (nt = d+1) checking all drafted tokens at target cost C_T(n).D5
speculative speedup ruleWin only if E[a]·C_T(1) > C_T(d+1) + d·C_D — the whole D5 plan reduces to this inequality.D5, STEPS 80
MTP (multi-token prediction)Draft source using MTP heads (DeepSeek-V3 / Qwen3-Next GGUFs); unavailable to minfer's dense models.SPEC, D5
EAGLE-3 / DFlash / DSpark / n-gram self-speculatorsThe other draft mechanisms in llama.cpp's speculative zoo, outside draft-simple's scope.SPEC
MoE + MLAArchitecture prerequisite for MTP drafts — its own future campaign; minfer targets dense.SPEC, D5

L2 — Numerics & data formats

Term / FormulaOne-line meaningSource
GGUF v3Single-file model format; multi-part files merge into one tensor index, entry = part 0.GRAPH, BACKEND
Q4_0 / Q4_1 / Q5_0 / Q5_1 / Q8_0Legacy block quants: fixed block of 32 values + fp scale(s); q8_0 also used for activations.PRIMER, BACKEND
Q4_K / Q5_K / Q6_KK-quants: super-block of 256 split into 8 (or 16 for q6_K) sub-blocks with packed multi-bit scales.PRIMER, MMQ
super-block / sub-blockq4_K: 8×32 with 6-bit scales packed 4-per-32-bit word; q6_K: 16×16 with 8-bit scales + ql/qh nibble halves.MMQ, PRIMER
d, dmin, minPer-block fp16 scale, and for K-quant the per-super-block scale of scales / offset.PRIMER
dscPer-super-block scale descriptor staging (q4_K DSC path; MINFER_MMQ_Q4K_DSC).STEPS 38-39
ql / qhLow/high nibble halves of packed 4-bit weights (q6_K: ql+qh interleaved by bit plane).MMQ
SWAR unpackSIMD-within-register nibble→int8 expansion via bit masks instead of per-byte ops (r30).STEPS 30
dequantize / dequantConverting packed blocks to arithmetic values inside the kernel; raw kernels avoid materializing f16.MMQ, STEPS
raw-byte / raw-nibble kernelsOperate directly on packed bytes ("BT") or nibbles ("NB") without a dequant f16 round-trip.MMQ, STEPS 28-38
round-trip (quant↔dequant)The correctness check pattern: quantize, dequantize, compare against fp reference.STEPS
tolerance gateNumerical acceptance: abs err ≤ 0.05 vs a CPU reference computed in f64.HUB, STEPS 78
bitwise identityStronger gate: refactor must produce byte-identical outputs (fused vs unfused is bit-identical by design).GRAPH, STEPS 78
f16 storage vs f32 accumulateWeights/activations may be stored f16 (__half) but mma/dp4a accumulate in f32 or int32.PRIMER, STEPS
int8 prefill activationsCUDA prefill quantizes activations to int8 for IMMA GEMM; decode MMVQ reads f32.PRIMER, BACKEND
Q8_0 activation quant (CPU)CPU quantizes activations on the fly; GPU reads f32 — logits differ by design, compare per-path.GRAPH
f32-accumulate mmar15 experiment: accumulate tensor-core results in fp32 registers instead of int32.STEPS 15
__expf scale pathq6_K dsc rebuild uses __expf; parity-guarded (W_exp debug, r44/r54 gates).STEPS 44, 54
W16 cacheCUDA-side cache of weights converted to f16 for some paths (MINFER_NO_W16CACHE to disable).STEPS 34+
split-k (dpl)"dpl" = split-plane B layout used by the final q6_K BT kernel (doc 76).STEPS 76
MMVQ uint4 sub-pairsVectorized 16-byte loads split per-thread sub-pairs in the MMVQ weight-streaming rework.STEPS 12+
QI8_1llama.cpp MMQ tiling constant: int8-activation tile width per 32-k chunk (= QK8_1/(4·QR8_1) = 8).MMQ

Tensor-layout symbols (from METAL §5.6)

SymbolOne-line meaningSource
n_embd / n_head / nkModel hidden size, query-head count, KV-head count; gqa = n_head/nk.METAL
hd / hd_kvAttention head dim and KV head dim (may differ under GQA).METAL
nt / nkv / nktTokens in the batch (decode nt==1), KV positions used, KV capacity.METAL
od / id / nfMatmul output/input dims (weight rows/cols) and FFN intermediate dim.METAL
positionsPer-token KV position array; nkv = positions[t] + 1.METAL
ne00..ne33ggml tensor dims: ne0x = dim0 of the x-th src, ne1x = dim1, etc.METAL
nb10..nb33ggml byte strides per dim for src1 (nb10 elem stride, nb11 row/token stride).METAL
ns10 / ns20Element counts per head/row/token (nb11/nb10, nb21/nb20) — flash-KV inner-loop stride.METAL
nwg / nsgWorkgroups and simdgroups per threadgroup (Metal launch geometry).METAL

L3 — Performance model

Term / FormulaOne-line meaningSource
memory-bound / compute-boundLimited by bytes moved vs FLOPs issued; GB10 decode is memory-bound, prefill compute-bound.PRIMER
roofline modelPerformance ceiling = min(peak FLOPs, AI × peak bandwidth); AI = FLOPs per byte.PRIMER
arithmetic intensity (AI)FLOPs per byte of traffic; decode GEMM at nt=1 is ~1 MAC/weight-byte → bandwidth-bound.PRIMER
GB/s, TB/sEffective bandwidth; GB10 unified LPDDR5x ~273 GB/s shared CPU+GPU.PRIMER, STEPS
tok/sDecode throughput in tokens per second (headline metric: 7B q4_k_m CUDA 54.3).HUB, STEPS 80
MAC / GMACMultiply-accumulate; GMAC = 10⁹ MACs; TMAC/s = 10¹² MACs per second (kernel-level throughput).STEPS 65-73
M/GMACSASS instructions issued per GMAC — the instruction-stream efficiency metric (llama 6.06 vs ours 10.14).STEPS 65-72
amortizationSpreading fixed weight traffic over more rows: nt=4 batched decode gives BT-MMQ 2.7×.PRIMER, STEPS 80
C_T(n)Cost of a target verify at batch n; C_T(1)=18.42 ms for 7B q4_k_m CUDA.STEPS 80
C_DDraft-model cost per token (0.5B: 2.92 ms CUDA → CPU-draft dead at 1.35×).STEPS 80
break-even p*Minimum acceptance rate for speculative win: 0.73 / 0.81 / 0.90 at d=2/4/8.STEPS 80
D5-0 gateCondition to proceed: measured nt=3 verify amortization ≥ 2.5×.STEPS 80
batched-decode regiment = tokens per decode step; nt=4 is the campaign's anchor amortization point.PRIMER, STEPS
KV traffic sharePer-token bytes = weights (dominant) + KV read/write + logits; quantized KV shrinks the KV share.PRIMER
3× gap attributionMethod of splitting the llama.cpp-vs-minfer wall-clock gap into per-kernel shares before optimizing.STEPS 12, 47, 65
wavefront countShared-memory work serialized per wavefront — 1.76× wavefronts/IMMA at equal IMMA rate meant inefficiency, not scarcity.STEPS 36
bytes-per-tokenDecomposition of decode memory traffic; the roofline input for every decode optimization.PRIMER

L4 — Micro-architecture (kernel implementation)

Term / FormulaOne-line meaningSource
SM (streaming multiprocessor)The GPU core unit; GB10 has 6144 CUDA cores across SMs; occupancy counts blocks/SM.PRIMER
blocks/SMResident blocks per SM; NB kernel uses 2, q6_K BT uses 3 (r40 probe).STEPS 28, 40
occupancyRatio of resident warps to maximum; raised by lowering registers/smem per block.STEPS 12, 40
register pressureToo many registers per thread kills occupancy; measured via ptxas -v spill output.STEPS 12-73
__launch_bounds__Compiler directive capping registers/threads to hit a target occupancy.PRIMER, STEPS
warp32 threads executing in lockstep; divergence inside a warp serializes paths.PRIMER
tile / tile shapeThe M×N×K block a kernel iterates over; TM=128/256 x-tile widening experiments (r13, r23).PRIMER, MMQ
tile-regimeWhich pre-tuned launch/tile configuration an (M,N,K) shape lands in — the D5-0 pivot concept.PRIMER, STEPS 80
wave quantizationPartial last wave of blocks leaves SMs idle; small GEMMs must size grids to avoid it.PRIMER, STEPS 12
mma.sync m16n8k16Tensor-core int8 matrix-multiply-accumulate instruction; the BT GEMM inner op.PRIMER, MMQ
wmmaLegacy warp-level matrix API; used for FA prefill P·V on tensor cores.STEPS 20
dp4a4-way int8 dot-product instruction; the MMVQ decode path's inner op.PRIMER, MMQ
IMMA / tensor pipeThe int8 tensor-core hardware pipe; smsp__inst_executed_pipe_tensor_subpipe_imma counts it.STEPS 36, 65
LDSM / ldmatrixLoads an 8×8 f16 fragment into registers laid out for mma; A-fragment reuse ratio 0.125 vs 0.5 was the llama edge.STEPS 36, 65
A-fragment / B-fragmentThe mma operand fragments each warp holds; reuse rate decides LDSM traffic.STEPS 36, 65
LDGSTS / cp.asyncAsync global→shared copy bypassing registers; the BT kernels stage A/B/dsc with it.PRIMER, STEPS 45
cp.async-db2cp.async with 2-stage double buffering (MINFER_MMQ_RAW_* sched gates).STEPS 45-54
LDG / STS / LDSGlobal load, shared store, shared load — the synchronous counterpart trio.STEPS, PRIMER
SASS instruction names (IMAD, I2F, F2I, FMUL, FFMA, LOP3, SHF, PRMT, LEA, SEL, CS2R, IADD3, FADD, BRA)The assembly opcodes read in SASS forensics to count real work per loop.STEPS 14-73
shared memory / dynamic smemOn-chip scratchpad; sized via cudaFuncAttributeMaxDynamicSharedMemorySize.PRIMER, BACKEND
bank conflictsSimultaneous LDS hits to the same bank serialize; r22 removed them via layout.STEPS 22
swizzleXOR-based shared-memory address permutation to avoid bank conflicts.PRIMER, STEPS 22
scoreboard stallWarp waiting on a memory dependency tracked by L1TEX scoreboard (r41 attack).STEPS 41
MIO pipeMemory-IO instruction queue; shown not scarce (r36) — A-fragment reuse was.STEPS 36
coalescingWarp-wide global accesses touching contiguous lines; r21 coalesced A staging.STEPS 21
L2 window / access policy windowPinning a buffer's residency in L2 via cudaAccessPolicyWindow (MINFER_MMQ_L2WIN).STEPS 23
KSPLITSplitting the K reduction across blocks with atomic/partial adds (q6_K KSPLIT=2).STEPS 39
KS (k-step)K elements processed per inner iteration (KS=64 GEMM).STEPS 12
KD / KDRK-depth unroll factor and K-depth register pipeline depth (q6_K KDR=4, =8 regressed).STEPS 39-40
double bufferingOverlapping stage(n+1) loads with compute(n); the NB/BT staging pattern.STEPS 9, 45
software pipeliningRestructuring the loop so load/compute phases of different iterations overlap.STEPS 24
unroll (kd-loop)Compiler/pragma loop unrolling to expose ILP; r28 "NB kd-loop unroll".STEPS 28
epilogueThe post-mma tail: scaling, output store; r32 cut its cost.STEPS 32
B pre-format / quantize-transpose prepassRepacking B (weights) offline into kernel-friendly layout (r34).STEPS 34
MMVQ_PARAMETERS_GB10llama.cpp launch-config constant table for GB10 MMVQ, adopted by minfer.STEPS 12
block reduceWarp/block-wide reduction for logits accumulation (mmvq_block_reduce).STEPS 12
PDL / programmatic dependent launchOverlapping dependent kernel launch tails: cudaGridDependencySynchronize + programmatic stream serialization attribute.STEPS 74, PRIMER
griddepcontrolThe SASS/PTX-level instruction pair behind PDL.STEPS 74
wave (n)One full pass of all resident blocks; kernel iteration wave counting for sched analysis.STEPS 36
elect.syncWarp election intrinsic seen in SASS forensics.STEPS 65
MMQ_TILE_NE_K / MMQ_TILE_Y_Kllama.cpp MMQ shared-memory tile pitches (B tile 32+4 ints; y-tile row stride 36 ints = 144 B, the +4 avoids bank conflicts).MMQ

L5 — Platform & tooling

Term / FormulaOne-line meaningSource
GB10 / DGX SparkThe target machine: Grace 20-core ARM + Blackwell GPU, unified LPDDR5x, sm_121.PRIMER, BUILD
sm_XX / compute_XXGPU arch targets; build emits SASS for sm_70…sm_121 probes + PTX compute_70/72 for backward JIT.BUILD, PRIMER
nvcc / ptxasCUDA compiler and its SASS backend; ptxas -v gives register/spill counts.BUILD, STEPS
PTXVirtual ISA JIT-compiled at load; the backward-compatibility artifact.BUILD, PRIMER
SASSReal GPU assembly; forensic disassembly via cuobjdump/nvdisasm.STEPS 14+
-ccbin pinningForcing nvcc's host compiler when the default is rejected (MINFER_CUDA_CCBIN).BUILD
detect_archsbuild.rs probing sm_70…sm_121 by compiling a dummy .cu per arch.BUILD
libcuda_kernels.aStatic kernel archive nvcc produces, linked into the Rust binary.BUILD
cuda_static featureLink cudart statically so no libcudart.so is needed at runtime.BUILD
CUDA driver vs runtime APIlibcuda low-level vs libcudart convenience layer; minfer binds runtime via hand-written externs.PRIMER, BACKEND
CUDA GraphsCaptured kernel sequence replayed with one launch (cudaStreamBeginCapture/cudaGraphLaunch/cudaGraphInstantiate); MINFER_NO_CUDA_GRAPH to disable.BACKEND, STEPS 19
graph capture (prefill)Pre-capturing the prefill segment; default-ON since R3-B (MINFER_NO_PREFILL_CAPTURE=1 opts out).STEPS 57+
pinned memoryPage-locked host memory (cudaHostAlloc) for fast H2D/D2H; readback path has a kill switch.BACKEND, STEPS
cudaMemcpyAsync / streamsAsync copies on streams (cudaStreamCreate); decode uses graph launch, prefill streams.BACKEND
cudaMallocManaged / unified memoryMemory visible to both CPU and GPU (used once; avoided on GB10 due to bandwidth sharing).PRIMER
cudaFuncSetAttributeRuntime call to raise per-kernel dynamic smem limits.BACKEND
error guards (cudaGetLastError)Every launch checks the error; cudaErrorMisalignedAddress was a real campaign bug.STEPS 13, SAFETY
ncu (Nsight Compute)Kernel profiler: sm__warps_active, lts__t_sectors, smsp__inst_executed metric families.STEPS 12+
nsys (Nsight Systems)Timeline profiler for wall decomposition and graph-launch analysis.STEPS 57+
locked clocksnvidia-smi -lgc fixes GPU clocks so medians are comparable across runs.STEPS 77
MINFER_* env gates~40 kill-switch env vars (MINFER_MMQ_RAW, MINFER_MMQ_A_TRANSPOSE, MINFER_PDL, MINFER_FUSED_B, …) toggling one experiment at a time.HUB, STEPS
cuobjdump / nvdisasmTools producing the SASS listings used in forensics.STEPS 25
MINFER_TRACE / MINFER_GRAPH_DUMPPer-node real-data trace and graph dumps for viz tooling.GRAPH

L6 — Engine architecture

Term / FormulaOne-line meaningSource
ComputeGraph / CNodeThe declarative graph and its nodes; inference = build → assign → fuse → allocate → execute.GRAPH
Op enum / NodeMetaTyped operations and per-node metadata in graph/ops.rs.GRAPH
GraphBuilderDeterministic graph construction; identical GraphParams ⇒ identical topology.GRAPH
GraphParams / params-only reuseReuse check compares parameters only (GraphCache::try_reuse), never data.GRAPH
CParams.gpuParticipation flag recording whether the run used the GPU backend.GRAPH, BACKEND
Backend traitsupports_op/supports_fused, buffer pool, execute_node, host IO, synchronize — the CUDA backend is the worked example.GRAPH, BACKEND
GraphAllocator / livenessSingle buffer owner; allocates by liveness in build order; persistent KV regions survive rebuilds.GRAPH
kv_pair / persistent KV regionsEach layer owns two allocator regions (K/V) that outlive a decode step.GRAPH
scheduler (assign → split → execute)Assigns backends, splits the graph at backend boundaries, executes splits serially.GRAPH
split boundaryCross-backend sync/copy point; one Metal command buffer per split (CUDA: one stream).GRAPH
FusionPassBuild-time fusion of QKV (bias+rope+store) and FFN (swiglu); fused vs unfused bit-identical; gated by env.GRAPH
positions-as-dataRule 1: topology never depends on n_past — the precondition for decode reuse and CUDA graphs.GRAPH
fill_input_i32Integer inputs stored via f32::from_bits so the f32-typed input buffer carries token ids.GRAPH
in-place aliasing ruleSilu/RoPE alias their input (sole consumer + same backend); never host-copy a GPU-pending buffer.GRAPH, SAFETY
ModelDef traitPer-architecture forward/build_graph/forward_graph; models live in models/<name>/.GRAPH
weight layout conventionMetadata [in, out], memory row-major [out][in], activations token-major [nt][d].GRAPH
guard failure = abortKernel-invariant violations return Err from execute_node — never silent CPU fallback; guards print actual values.SAFETY
submit() bounded waitGPU submission waits bounded and checks status — never blocks forever.SAFETY
no early return past barrierMetal/CUDA rule: no exit path may skip a threadgroup_barrier/__syncthreads.SAFETY
prefill captureBackend feature storing the captured prefill graph for replay.BACKEND
IR (intermediate representation)The graph as an op-level IR; fusion makes fused ops first-class IR citizens.GRAPH
NodeId / DTypeNode handle and tensor data-type enum carried by every CNode.GRAPH
GetRowsRow-selection op: embedding lookup, and the n_out tail-row optimization (G3).GRAPH
BatchMatMulBatched matmul op (shared activation quantization, Q4_0); composable with fusion.GRAPH
FusedOp / supports_fusedFusion capability tag (currently only SwiGLU) checked by the fusion pass against each backend.GRAPH
n_out tail-row optimizationAfter the final wo, run FFN/norm/lm_head only on the tail n_out rows (llama inp_out_ids style); GraphParams.n_out joins the reuse decision.GRAPH
DOT / JSON exportgraph/dot.rs and graph/json.rs render the graph for viz.GRAPH

L7 — Methodology

Term / FormulaOne-line meaningSource
interleaved same-window A/BAlternating minfer/llama.cpp runs inside one time window so thermal/clock drift cancels.STEPS 77, HUB
median of NTake the median of repeated runs; means are corrupted by outliers on shared hardware.STEPS 77
pre-registered barThe success threshold is written into the step doc before measuring.STEPS 77
parity gateGreedy output must match llama.cpp (or CPU f64 ref within 0.05) before any perf number counts.HUB, STEPS 77-78
correctness batchDocs 78/79: batch verification sweeps over all kernels/quant types.STEPS 78-79
MEAS-ONLYStep status: measurement without landing code (e.g. doc 80).HUB
LANDED / REVERTEDStep outcome statuses in the hub tables.HUB
r-numbers (r1…r76)Experiment numbering across the MMQ/GEMM campaign; each step doc records one.HUB, MMQ
Era A/B/C/DCampaign phases: baseline (A), MMVQ (B), MMQ GEMM (C), decode/GEMM (D).HUB
Direction-A/B, Session A–FNamed experiment tracks within a phase (e.g. raw-nibble vs BT; quantize-fusion sessions).STEPS 28-56
SASS forensicsExplaining a perf delta by diffing disassembly (r25 opcode diff) instead of guessing.STEPS 25, 65
wall decompositionSplitting end-to-end time per kernel/phase before optimizing anything.STEPS 12, 47, 65
counter-guided iterationNext experiment chosen by the ncu counter that bounds the kernel (occupancy → wavefronts → IMMA rate).STEPS 36-73
one-variable-at-a-timeEach experiment toggles exactly one gate/env var; everything else frozen.STEPS 77
llama.cpp as reference$HOME/git/reading/llama.cpp is the ground truth for both parity and technique adoption.MMQ, SPEC
doc-per-step conventionEvery step writes one numbered record with a fixed six-section structure (STYLE.md).STEPS STYLE
measurement artifactsRaw bench JSONs kept under /tmp per step and reported, never committed.STEPS 77-80
G1 / G2 / G3Graph-refactor Phase-9 sub-experiment labels (attention dispatch, rms_norm_256, n_out tail-row).GRAPH

llama.cpp Metal (MPS) End-to-End Inference Path — Reference Baseline

This document records, line-by-line, llama.cpp's end-to-end GPU inference implementation on Apple Silicon (Metal / MPS backend), as the reference starting point for comparing minfer against llama.cpp. Each row gives the execution order, purpose, and source location (file:line), plus a minfer equivalent comparison column.

Scope: the full chain from "model weight loading" to "next-token logits readback" (including batch preparation, graph build, scheduler, Metal execution, KV cache), with the Metal backend internals recorded at the finest granularity. The CPU-side sampler is not expanded (non-Metal path).

0. Version & reading conventions

  • llama.cpp baseline: master @ 8a832e4bf (2026-08-20). This revision uses the per-arch graph-build API (src/models/ + llm_graph_context); the Metal backend spans 6 files under ggml/src/ggml-metal/.
  • minfer baseline: master @ ad040ff (2026-08-20).
  • Paths are relative to each repo root (llama.cpp / minfer).
  • Comparison-column value convention:
    • Match = minfer has an equivalent implementation;
    • N/A = minfer has no such step (architectural difference);
    • Similar = functionally equivalent but structurally different (noted).
FileRole
ggml/src/ggml-metal/ggml-metal.cppMetal backend interface (buffer types, set/get tensor, graph_compute entry)
ggml/src/ggml-metal/ggml-metal-device.mLow-level MTLDevice/MTLCommandQueue, encoder, buffer allocation, kernel library loading
ggml/src/ggml-metal/ggml-metal-device.cppPipeline (kernel instance) lookup/compilation, op support table
ggml/src/ggml-metal/ggml-metal-context.mMulti-command-buffer scheduling, graph compute, tensor set/get
ggml/src/ggml-metal/ggml-metal-ops.cppPer-op encoding (encoder setup + kernel dispatch + fusion)
ggml/src/ggml-metal/ggml-metal.metalMetal kernel source
ggml/src/ggml-metal/ggml-metal-impl.hQuantized block structs, dequant functions, threadgroup constants, function-constant offsets
ggml/src/ggml-backend.cppBackend scheduler (split / alloc / compute)
src/llama-graph.cppGraph-build helpers (build_*, llm_graph_context)
src/llama-context.cppdecode / process_ubatch / graph_compute / logits readback
src/llama-kv-cache.cppKV cache (allocation, slot lookup, in-graph write/read)
src/llama-model.cpp / src/llama-model-loader.cppModel loading & weight registration
src/models/qwen2.cppQwen2 architecture graph build

1. High-level overview (12 phases)

#PhasePurposellama.cpp locationminfer equivalent
P0Backend & scheduler initCreate Metal backend, MTLCommandQueue, kernel library, scheduler; pre-reserve compute buffersggml-metal.cpp:689 ggml-metal-context.m:84 llama-context.cpp:581 ggml-backend.cpp:1792src/metal/ops.rs (MpsState::try_new)
P1Model loading / weight registrationAllocate Metal buffers per tensor and upload quantized weightsllama-model.cpp:1401 llama-model-loader.cpp:1426 ggml-metal-device.m:1631src/models/qwen2/loader.rs + src/metal/ops.rs (register_weight)
P2Batch preparation & microbatchingSplit the API batch into micro-batches, reserve host output buffersllama-context.cpp:1635 llama-batch.cpp:25src/main.rs (single batch, no split) "N/A"
P3Compute graph build (Qwen2)Build the ggml compute graph (DFS topological order)src/models/qwen2.cpp:53 llama-graph.cpp ggml.c:7188src/models/qwen2/graph.rs:438 (declarative graph build)
P4Scheduler split & allocationAssign nodes to backends, split into runs, gallocr allocationggml-backend.cpp:1936→:1055"N/A" (single MPS backend, static buffers)
P5Scheduler computePer-split: copy inputs, call backend graph_computeggml-backend.cpp:1594 ggml-metal.cpp:535forward.rs:88-134 (single CB, all layers)
P6Metal graph computeMulti-command-buffer encode (main thread + n_cb workers)ggml-metal-context.m:438 :663src/metal/ops.rs (submit, single CB)
P7Per-op encoding & concurrency/barrierFilter empty nodes, concurrency check, insert memoryBarrierggml-metal-ops.cpp:175 device.m:513src/metal/ (barrier), :327 (dispatch_2d)
P8Per-op kernel dispatchPer op type: set pipeline + args + threadgroupsggml-metal-ops.cpp:265-497src/metal/ (quant_matmul_f32_on_gpu_buf) et al.
P9GPU kernel executionMetal shader computeggml-metal.metalsrc/metal/kernels/
P10KV cacheIn-graph write (set_rows) & read (flash/matmul direct)llama-kv-cache.cpp:1301 llama-graph.cpp:2800src/metal/ops.rs (store_kv) + src/cache.rs
P11logits/embd readbackGPU→host copyggml-metal-context.m:351 llama-context.cpp:1854src/metal/runtime.rs (output_norm_gpu then download_logits, now n_out rows / 608 KB)

2. Detailed step table (end-to-end execution order)

P0 Backend & scheduler init (once per process)

#StepPurposellama.cpp locationminfer equivalent
0.1Register Metal backendConstruct backend per device, call ggml_metal_initggml-metal.cpp:689—
0.2Create struct ggml_metal contextMTLDevice + shared MTLCommandQueue, load kernel library, create concurrent dispatch queue, fusion/concurrency flagsggml-metal-context.m:84-175src/metal/ops.rs (MpsState singleton) — 2026-08-21: kernel library is now a build-time precompiled .metallib embedded in the binary (build.rs, llama's -O3 flags, newLibraryWithData), with a runtime newLibraryWithSource fallback when the toolchain is absent
0.3Init low-level deviceMTLCreateSystemDefaultDevice + newCommandQueue, probe capabilities (simdgroup_mm / unified_memory / bfloat / tensor)ggml-metal-device.m:714-760src/metal/ (new_device capability probe)
0.4tensor-API gatehas_tensor defaults to OFF for pre-M5/M6/A19/A20 (disabled on M4)ggml-metal-device.m:753-760"N/A" (llama itself doesn't use the tensor API on M4)
0.5Create schedulerggml_backend_sched_new (backend array + gallocr + events)ggml-backend.cpp:1792"N/A"
0.6Pre-reserve worst-case graphsReserve pp (prefill) and tg (decode) graphsllama-context.cpp:630-657"N/A" (static buffers)

P1 Model loading / weight registration (once per process)

#StepPurposellama.cpp locationminfer equivalent
1.1Load architecture tensorsQwen2's load_arch_tensors creates all weight tensorssrc/models/qwen2.cpp:19-47src/models/qwen2/loader.rs
1.2Per-layer device splitLayers 0..i_gpu_start-1 stay on CPU; the rest split to GPU devices by free memoryllama-model.cpp:1314-1323"N/A" (all-on-GPU or MINFER_DISABLE_MPS all-CPU)
1.3Allocate Metal buffersggml_backend_alloc_ctx_tensors_from_buft → ggml_metal_buffer_init; mmap path ggml_metal_buffer_mapllama-model.cpp:1637 ggml-metal-device.m:1631,1701src/metal/ (register_part: ONE page-aligned newBufferWithBytesNoCopy per mmap'd part) — 2026-08-21: weights are (buffer, byte-offset) into the part buffer, llama's exact design
1.4Buffer storage modeshared = newBufferWithBytesNoCopy (mmap/weights); private = newBufferWithLengthggml-metal-device.m:1668,1673src/metal/ (mmap parts StorageModeShared NoCopy; scratch/KV buffers StorageModeShared copies)
1.5Mark weight buffersGGML_BACKEND_BUFFER_USAGE_WEIGHTSllama-model.cpp:1657—
1.6Upload weight datammap direct reference; non-mmap uses blit set_tensor_asyncllama-model-loader.cpp:1548,1558 ggml-metal-context.m:3072026-08-21: zero-copy — weights are Borrowed slices of the mmap'd GGUF (Tensor.data: Cow<'static,[u8]>) wrapped by newBufferWithBytesNoCopy at the part level; no memcpy anywhere

P2 Batch preparation & microbatching (per decode)

#StepPurposellama.cpp locationminfer equivalent
2.1Init batch allocatorllama_batch_allocr::initsrc/llama-batch.cpp:25—
2.2Split micro-batchesmemory->init_batch, retry on failure (cache optimization)llama-context.cpp:1828-1856 llama-kv-cache.cpp:698"N/A" (minfer single batch; -n 0 = whole-segment prefill)
2.3Reserve host outputoutput_reserve fixed-size logits/embd buffersllama-context.cpp:2032forward.rs:141 (new logits vec each call)
2.4Microbatch loopCall process_ubatch per ubatchllama-context.cpp:1879-1900—

P3 Compute graph build (Qwen2)

#StepPurposellama.cpp locationminfer equivalent
3.1Entryllama_model::build_graph → build_arch_graphllama-model.cpp:2457src/models/qwen2/graph.rs:438 (forward)
3.2Input embdbuild_inp_embd: token ids + ggml_get_rows(tok_embd, inp_tokens)llama-graph.cpp:2284src/metal/ (embed_tokens_gpu, get_rows — 2026-08-21: all minfer quants on GPU + dispatched into the MAIN command buffer, llama-graph-style single submit, #38/#39)
3.3Position inputbuild_inp_posllama-graph.cpp:2373src/metal/ops.rs (upload_positions)
3.4KV graph inputsbuild_attn_inp_kv (k_idxs/v_idxs, mask, rotation tensors)llama-graph.cpp:2729src/metal/ (store_kv uses pos_buf)
3.5Per layer: attn_normbuild_norm (RMSNorm + Mul + optional Add)llama-graph.cpp:1556src/metal/ops.rs (rms_norm) + :648 (add)
3.6Per layer: QKVbuild_qkv (3× build_lora_mm = ggml_mul_mat for wq/wk/wv)llama-graph.cpp:1592src/metal/ (3× quant_matmul)
3.7Per layer: RoPEggml_rope_ext (once each for Q, K)src/models/qwen2.cpp:86-96src/metal/ops.rs (rope_f32 ×2)
3.8Per layer: KV writemctx_cur->cpy_k/cpy_v → ggml_set_rowsllama-graph.cpp:2800-2801 llama-kv-cache.cpp:1301,1336src/metal/ops.rs (store_kv)
3.9Per layer: attentionbuild_attn → build_attn_mha: flash path ggml_flash_attn_ext; non-flash mul_mat(k,q)+soft_max_ext+mul_mat(v,kq)llama-graph.cpp:2517,2557src/metal/ops.rs (attn_flash_prefill) / :741 (gqa_attn_f32)
3.10Per layer: wo + residualbuild_attn inner build_lora_mm(wo) + ggml_addllama-graph.cpp:2677 src/models/qwen2.cpp:110src/metal/ (wo matmul) + :648 (add)
3.11Per layer: ffn_normbuild_normsrc/models/qwen2.cpp:114rms_norm
3.12Per layer: FFNbuild_ffn: SILU-gated mul(gate,up) + mul_mat(down)llama-graph.cpp:1669src/metal/ops.rs (swiglu) + :392 (down matmul) — 2026-08-21: last layer runs on n_out rows only (llama's get_rows reduction, #34)
3.13Per layer: residualggml_addsrc/models/qwen2.cpp:127add_f32 — 2026-08-21: last layer's both residuals on the tail n_out rows (add_f32_off, #34)
3.14Output normbuild_norm (result_norm)src/models/qwen2.cpp:137src/metal/runtime.rs (rms_norm inside output_norm_gpu) — 2026-08-21: also output-rows-only (n_out), matching llama
3.15lm_headbuild_lora_mm(model.output) + optional biassrc/models/qwen2.cpp:145-150src/metal/runtime.rs (output GEMM) — 2026-08-21: now output-rows-only (n_out), matching llama
3.16Node orderingggml_build_forward_expand → ggml_build_forward_impl → ggml_visit_parents_graph (DFS, parents before children)ggml.c:7188,7120"N/A" (minfer encodes imperatively in layer order)

P4 Scheduler split & allocation

#StepPurposellama.cpp locationminfer equivalent
4.1Split graphggml_backend_sched_split_graph: 5-pass node→backend assignment (weight's backend decides MUL_MAT), build splits, insert cross-backend tensor_copyggml-backend.cpp:1055-1443"N/A" (all Metal)
4.2Allocate memoryggml_gallocr_alloc_graph (retry via reserve_n on failure)ggml-backend.cpp:1562-1585"N/A" (static buffers)

P5 Scheduler compute (per split)

#StepPurposellama.cpp locationminfer equivalent
5.1Copy split inputsCopy cross-backend srcs to the split's backend (INPUT flag → sync copy; MoE copies only used experts; else async)ggml-backend.cpp:1555-1671"N/A"
5.2Call backend computeggml_backend_graph_compute_async → Metal's ggml_backend_metal_graph_computeggml-backend.cpp:1678 ggml-metal.cpp:535forward.rs:134 (cb.submit)
5.3Event recordMTLEvent signal for multi-copy scenariosggml-backend.cpp:1717-1721"N/A"

P6 Metal graph compute (multi-command-buffer scheme)

#StepPurposellama.cpp locationminfer equivalent
6.1Split workn_main = MAX(64, 0.1*n_nodes); first n_nodes_0 nodes encoded by main thread, rest split evenly by n_cbggml-metal-context.m:445-466—
6.2Main-thread encodeCreate cmd_bufs[n_cb], enqueue, encode_async(n_cb)ggml-metal-context.m:510-523—
6.3Worker encodedispatch_apply(n_cb, d_queue, encode_async) encodes remaining CBs concurrentlyggml-metal-context.m:530-550—
6.4encode_async blockPer CB: compute node range, ggml_metal_op_init → loop ggml_metal_op_encode → ggml_metal_op_free → commitggml-metal-context.m:676-721—
6.5Async returngraph_compute returns immediately (only capture mode waits + checks status)ggml-metal-context.m:557-611src/metal/ops.rs (submit blocks + 10 s cap)
6.6Synchronizeggml_metal_synchronize: wait + check all CB statuses, set has_error on failureggml-metal-context.m:239-295submit() MTLCommandBufferStatus check

n_cb value: ggml_backend_metal_set_n_cb(backend, 1) (ggml-metal.cpp:612,707), ggml_metal_set_n_cb caps at GGML_METAL_MAX_COMMAND_BUFFERS (context.m:665). That is 2 CBs (1 main + 1 worker). minfer uses a single CB for all layers (forward.rs:76-134).

P7 Per-op encoding & concurrency/barrier model

#StepPurposellama.cpp locationminfer equivalent
7.1Create encoderggml_metal_encoder_init: MTLDispatchTypeConcurrent (when use_concurrency) or serialggml-metal-ops.cpp:42 ggml-metal-device.m:464src/metal/ops.rs (cmd_buffer: new_compute_command_encoder, serial)
7.2Filter empty nodesSkip empty / no-op nodesggml-metal-ops.cpp:55-62"N/A" (minfer has no empty-node concept)
7.3Op support checkggml_metal_device_supports_op big switchggml-metal-ops.cpp:201 ggml-metal-device.m:1086quant-type checks in layer_gpu
7.4Concurrency checkIf the current node's read/write ranges conflict with existing mem_ranges, insert memoryBarrierWithScope:MTLBarrierScopeBuffers and clear ranges; else record ranges and run concurrentlyggml-metal-ops.cpp:159-173,220-225 ggml-metal-device.m:513src/metal/ (barrier: after every dispatch) + :333
7.5Op dispatch switchDispatch to ggml_metal_op_* by node->op; returns fusion count n_fuseggml-metal-ops.cpp:265-497encode in fixed sequence inside layer_gpu

Key difference: llama uses mem_ranges for dependency-aware barriers (non-conflicting adjacent ops run concurrently in the same encoder, MTLDispatchTypeConcurrent); minfer inserts an unconditional barrier after every dispatch (dispatch_2d → barrier(), src/metal/).

P8 Per-op kernel dispatch (forward-path ops)

ggml_opllama.cpp encoderSelected kernel (variants)llama.cpp locationminfer equivalent
MUL_MATggml_metal_op_mul_mat (3-way selection, §3.2)kernel_mul_mm_* / kernel_mul_mv_ext_* / kernel_mul_mv_*ops.cpp:2299-2541src/metal/ (quant_matmul_f32_on_gpu_buf) + :352 (gemm_dispatch)
FLASH_ATTN_EXTggml_metal_op_flash_attn_ext_kv_f16 / _pad / _blk / main kernel / _vec / _vec_reduceops.cpp:2990-3492src/metal/ops.rs (attn_flash_prefill)
RMS_NORMggml_metal_op_norm (fuses Mul+Add)kernel_rms_norm_fuse_implops.cpp:3887-4006 ggml-src/metal/kernels/qkv_fused.metalsrc/metal/ops.rs (rms_norm) + separate :648 (add)
ROPEggml_metal_op_ropekernel_rope_norm/neox/multi/visionops.cpp:4025-4126 ggml-src/metal/kernels/fa_prefill.metalsrc/metal/ops.rs (rope_f32)
ADD/SUB/MUL/DIVggml_metal_op_bin (ADD fusion ×8)kernel_add / kernel_mul (n_fuse specialization)ops.cpp:3578src/metal/ops.rs (add_f32 single op)
GET_ROWSggml_metal_op_get_rowskernel_get_rows_q/_fops.cpp:1165 ggml-src/metal/kernels/src/metal/ops.rs (embed_tokens_gpu → 2026-08-21: all minfer-supported quants — Q4_0/Q4_1/Q5_0/Q5_1/Q8_0 (32-elem) + Q4_K/Q6_K/Q5_K (256-elem), matching llama's template coverage)
SET_ROWSggml_metal_op_set_rowskernel_set_rows_*ops.cpp:1210 ggml-src/metal/kernels/src/metal/ops.rs (store_kv dedicated kernel)
CPY/DUP/CONTggml_metal_op_cpykernel_cpy_t_t/_f32_q/_q_f32ops.cpp:2078 ggml-src/metal/kernels/"N/A" (f16 KV converted directly by store_kv)

P9 GPU kernel execution (key kernels)

kernelPurposellama.cpp locationminfer equivalent
kernel_mul_mm<...>simdgroup/tensor matmul (64×32 tile, §3.2 variants)ggml-src/metal/kernels/ (template + instantiations)src/metal/kernels/mul_mm.metal (kernel_q4_0_mm_f32) and 7 more mm kernels
kernel_mul_mv_*mat-vec (decode, per quant type)ggml-src/metal/kernels/fa_decode.metal (q4_0), :8498 (q4_K), etc.src/metal/kernels/ *_f32_matmul kernels
kernel_mul_mv_ext_*small-batch (ne11∈[2,8]) mat-mvggml-src/metal/kernels/fa_prefill.metal"N/A"
kernel_flash_attn_ext_kv_f16quantized KV → f16 dequant pre-pass (Q4_0/1, Q5_0/1, Q8_0)ggml-src/metal/kernels/"N/A" (minfer KV stores f32/f16 raw, MINFER_CACHE_TYPE=f16; a packed Q8_0 cache is also stored, and mechanism B is the analogous per-window dequant stage — not llama's kv_f16 variant)
kernel_flash_attn_ext_padpad pre-pass for partial KV blocksggml-src/metal/kernels/src/metal/kernels/: kernel_kv_tail_pad (equivalent)
kernel_flash_attn_ext_blkmask pre-pass (nqptg/ncpsg blocks)ggml-src/metal/kernels/inline causal mask (kernel_flash_attn_blk_f32)
kernel_flash_attn_ext / _implflash attention main kernel (half8x8)ggml-src/metal/kernels/,7184src/metal/kernels/: kernel_flash_attn_blk_f32
kernel_flash_attn_ext_vec / _vec_reducedecode small-batch flash (half4x4, ne01<20)ggml-src/metal/kernels/,7980src/metal/kernels/: kernel_flash_attn_ext_f32
kernel_rms_norm_fuse_implRMSNorm + Mul + Add fusionggml-src/metal/kernels/qkv_fused.metalrms_norm_256 + separate add
kernel_soft_max*non-flash path softmaxggml-src/metal/kernels/mul_f32act_kquant.metal,2117"N/A" (inlined in flash; or a dedicated softmax kernel)
kernel_rope_*RoPEggml-src/metal/kernels/fa_prefill.metalsrc/metal/kernels/: kernel_rope_f32
kernel_get_rows_*embedding lookupggml-src/metal/kernels/,10092src/metal/kernels/: kernel_get_rows_q4_0/q4_1/q5_0/q5_1/q8_0/q4_k/q6_k/q5_k (templates kernel_get_rows_q32/_q256)

P10 KV cache

#StepPurposellama.cpp locationminfer equivalent
10.1KV tensor creationggml_new_tensor_3d(ctx, type_k/v, n_embd_k_gqa, kv_size, n_stream), default F16llama-kv-cache.cpp:231-232src/cache.rs (GPU KV f16 auto for 7B class / f32 for small, MINFER_CACHE_TYPE overrides)
10.2Per-layer device allocationper-layer backend buft, allocate + clearllama-kv-cache.cpp:299,307src/metal/ (KV buffer allocation)
10.3Slot lookupfind_slot: ring-buffer cell range + k/v idx tensorsllama-kv-cache.cpp:894src/metal/: store_kv writes by pos_buf
10.4In-graph writecpy_k/cpy_v → ggml_set_rows (K always cache-row indexed; V per FA/non-FA layout)llama-kv-cache.cpp:1301-1389src/metal/ops.rs (store_kv, nkt/nt strides)
10.5In-graph readflash reads cache tensor directly; non-flash mul_mat(k,q)/mul_mat(v,kq)llama-graph.cpp:2807-2808,2491,2535attn_flash_prefill / gqa_attn_f32 read KV buffers directly

Layout: llama KV = f16 [nkv][nk*hd], token stride nk*hd*elem (after llama-graph.cpp permute, flash receives nb11=nk*hd*elem). minfer uses the same layout but f32 by default, f16 optional (MINFER_CACHE_TYPE=f16).

P11 logits/embd readback

#StepPurposellama.cpp locationminfer equivalent
11.1Locate backendggml_backend_sched_get_tensor_backend(t_logits)llama-context.cpp:1948—
11.2Async readbackggml_backend_tensor_get_async → ggml_metal_get_tensor_async: newBufferWithBytesNoCopy wraps host memory + blit encoder GPU→host, queued into cmd_bufs_extllama-context.cpp:1854 ggml-metal-context.m:351-391src/metal/ (copy_from_gpu: Shared buffer direct memcpy, no blit) — 2026-08-21: now n_out×nv (608 KB for single output; was 301 MB)
11.3Synchronizebefore the next decode, ggml_backend_sched_synchronize waits for the blitggml-backend.cppsubmit() blocks + download_logits

3. Supplementary mapping tables

3.1 Per-op dispatch switch (ggml-metal-ops.cpp:265-497)

Complete forward-path mapping (non-forward ops omitted):

ggml_ophandlerfusion
CONCATggml_metal_op_concat—
ADD/SUB/MUL/DIVggml_metal_op_binADD ×N (up to 8 consecutive ADDs → 1 dispatch); Snake/GEGLU specialization
ADD_IDggml_metal_op_add_id—
SOFT_MAXggml_metal_op_soft_max—
MUL_MATggml_metal_op_mul_mat—
MUL_MAT_IDggml_metal_op_mul_mat_id (MoE)—
GET_ROWS / SET_ROWSop_get_rows / op_set_rows—
NORM / RMS_NORMggml_metal_op_normRMSNorm + Mul(weight) + Add(bias) in 1 kernel
ROPE / ROPE_BACKggml_metal_op_rope—
FLASH_ATTN_EXTggml_metal_op_flash_attn_extQK^T + softmax + PV single kernel (+aux pad/blk/kv_f16/vec_reduce)
DUP / CPY / CONTggml_metal_op_cpy—
SILU_BACK / GLUop_silu_back / op_glu (training/gating)—

3.2 MUL_MAT 3-way kernel selection (ggml-metal-ops.cpp:2336-2538)

BranchTrigger conditionkernel / pipelinethreadgroupsllama.cpp location
① mat-mv extsrc1=f32, ne00%128==0, src0 type in supported set, ne11∈[2,8] (K-quants need ne11∈[4,8])kernel_mul_mv_ext_* (nsg=2, nxpsg per ne00: 16/8/4)(ne01/r0ptg, ne11/r1ptg, ne12·ne13), 32×nsgops.cpp:2340-2439 device.cpp:706
② simdgroup MMnon-transposed, has_simdgroup_mm, ne00>=64, ne11>8 — **minfer #40: GEMM now dispatches for `nt≥2 && (od≥2048nt≥9)` (was nt≥16), closing the nt∈[9,15] gap (7B pp12 16.6→124 t/s ≈ llama)**kernel_mul_mm_<t0>_<t1> (function-constants bc_inp/bc_out/ne12/ne13/r2/r3)
③ mat-vecotherwisekernel_mul_mv_* (per quant type)nsg/nr0 per typeops.cpp:2491-2538 device.cpp:801+

Function-constant offsets (ggml-metal-impl.h:99-100): FC_MUL_MV=600, FC_MUL_MM=700; MM uses 700-705 (bc_inp/bc_out/ne12/ne13/r2/r3).

On M4, llama disables the tensor API (§0.4) and actually uses branch ②'s legacy simdgroup_matrix path — which is level-for-level equivalent to minfer's kernel_q4_k_mm_f32 (src/metal/kernels/fa_prefill.metal, 64×32 tile, 32×4 threads, 8192 B smem) (see minfer docs/METAL_OPTIMIZATIONS.md §3.6).

3.3 FLASH_ATTN_EXT variant selection (ggml-metal-ops.cpp:2990-3492)

TestVariantTrigger condition
use_vec_vec + _vec_reducene01 < 20 && ne00 % 32 == 0 (decode small batch, half4x4)
use_kv_f16first kernel_flash_attn_ext_kv_f16 dequantizes KV→f16KV type ∈ {Q4_0,Q4_1,Q5_0,Q5_1,Q8_0} (new in #27390)
has_kvpadfirst kernel_flash_attn_ext_padne11 % ncpsg != 0 (KV not a multiple of ncpsg)
has_maskfirst kernel_flash_attn_ext_blkmask present (block pre-pass)
main kernelkernel_flash_attn_ext (half8x8, nqptg=8/ncpsg=64, nsg=ne00>=512?8:4)prefill (non-vec path)

Middle-buffer layout (ops.cpp:3055-3065): after dst, in order pad → blk → tmp → kv_f16 (sizes computed by ggml_metal_op_flash_attn_ext_extra_*).

3.4 Fusion rule summary

Fusionllama.cppminfer
RMSNorm + Mul(weight) + Add(bias)kernel_rms_norm_fuse_impl (ops.cpp:3929-3974)not fused: rms_norm + separate add_f32
Consecutive ADD ×Nop_bin (ops.cpp:3195+)single add_f32
flash attention (QK^T+softmax+PV)single kernel + aux passesattn_flash_prefill (src/metal/ops.rs)
KV write + RoPERoPE writes into the KV path (graph k/v expanded together)store_kv dedicated kernel
GLU/SiLUop_glu / op_snake_fusedswiglu_f32 (single kernel)
mul_mat + bias / residualnot fused (bias/residual is a separate kernel_add)same (separate add_f32 after wo)

3.5 Multi-CB and encoder/barrier model comparison

Itemllama.cppminfer
CB countn_cb=1 → 2 CBs (main thread 64 nodes + 1 worker)1 CB (all 28 layers + output)
Encode parallelismdispatch_apply multi-thread concurrent encodesingle-threaded sequential encode
encoder dispatch typeMTLDispatchTypeConcurrent (on by default, GGML_METAL_CONCURRENCY_DISABLE to off)serial (new_compute_command_encoder)
barrierdependency-aware (mem_ranges conflict only inserts memoryBarrierWithScope)unconditional memoryBarrierWithScope after every dispatch
commit/waitmain thread returns async, synchronize waits explicitly; 10 s timeout guardsubmit() blocks on completed handler (10 s timeout + status check)

4. Key data structures

StructDefined atPurpose / key fields
struct ggml_metal (ggml_metal_t)ggml-metal-context.m:26device, library, d_queue, n_cb, cmd_bufs[] (cmd_bufs[n_cb+1]), encode_async block, cmd_bufs_ext, cmd_buf_last, has_error
struct ggml_metal_deviceggml-metal-device.m:521mtl_device, mtl_queue (globally shared), rsets (residency sets), library, props, addr_virt
struct ggml_metal_encoderggml-metal-device.m:460wraps MTLComputeCommandEncoder
struct ggml_metal_libraryggml-metal-device.m:97MTLLibrary + cached MTLComputePipelineState map + lock; newLibraryWithSource (device.m:234)
struct ggml_metal_bufferggml-metal-device.m (buffer_init:1631)buffers[] ({id<MTLBuffer>, offs}), is_shared, rset
struct ggml_metal_opggml-metal-ops.cpp:28per-CB encode state: enc, mem_ranges, filtered idxs[], fusion flags
struct ggml_backend_schedggml-backend.cpp (sched_new:1792)backend array, splits[], galloc, node/leaf_backend_ids[], events[b][c], graph_copy
struct ggml_backend_sched_splitggml-backend.cpp:1055+{backend_id, i_start, i_end, n_inputs, inputs[], graph}
llm_graph_contextllama-graph.hctx0, gf, hparams/cparams, sched, res
llm_graph_resultllama-graph.ht_inp_tokens, t_logits, t_embd, inputs[], compute ctx + ggml_cgraph
ggml_metal_pipeline_with_paramsggml-metal-ops.h{pipeline, nr0, nr1, nsg, smem} — everything a single kernel dispatch needs

5. minfer comparison notes (for later comparison work)

  1. Architecture difference: llama.cpp = declarative ggml graph (topological nodes → scheduler → backend); minfer = imperative (forward.rs encodes layer-by-layer directly into a single MPS command buffer). No scheduler/allocator layer; src/cache.rs holds KV directly. 1b. Output-rows reduction (2026-08-21): llama shrinks the graph to n_outputs rows after the last attention (get_rows(cur, inp_out_ids) + get_rows(inpSA, inp_out_ids), qwen2.cpp:106-108) → the last layer's FFN + both residuals + final norm + lm_head all run on 1 row. minfer mirrors the FULL reduction: final norm + lm_head via n_out (#32), and the last layer's FFN + residuals via layer_gpu(n_out, is_last) (#34) — minfer's total graph work (≈6.26 TFLOP) now exactly equals llama's.
  2. Level-for-level equivalence proven (minfer docs/METAL_OPTIMIZATIONS.md §3.4/§3.6): the prefill GEMM kernels (kernel_mul_mm vs kernel_q*_mm_f32) match at source/IR/smem/dispatch/runtime-compile level; this table's P8/P9 rows are the comparison anchors.
  3. Fusion gap: llama's RMSNorm+Mul+Add, ADD×N, single-kernel flash fusion vs minfer's mostly-separate dispatches (§3.4) — the source of the per-layer dispatch-count difference in decode/prefill.
  4. KV format: llama defaults f16 + optional quantized KV (since #27390, a kv_f16 dequant pass); minfer auto-selects f16 for the 7B class / f32 for small models (#37, MINFER_CACHE_TYPE=f16/f32 overrides), and since C4 also has a packed Q8_0 cache — CUDA since C4 S2b, Metal since #310, MINFER_CACHE_TYPE=q8_0.
  5. Multi-CB: llama 2-CB concurrent encode; minfer single CB all layers. Measured (minfer §3.6): llama's 2-CB split is slower in a pure-GEMM replay — not a speed source.
  6. Barrier: llama dependency-aware; minfer barriers after every dispatch. Measured free (§3.6) — not a gap source.

llama.cpp Compute Graph Design Analysis

This document analyzes the compute graph architecture design of llama.cpp and its role in end-to-end inference. Based on llama.cpp source code (2026-08 version).


1. Core Data Structures

1.1 ggml_cgraph — The Compute Graph Itself

Defined in ggml/src/ggml-impl.h:329:

struct ggml_cgraph {
    int size;             // maximum number of nodes/leafs/grads/grad_accs
    int n_nodes;          // number of operator nodes currently in use
    int n_leafs;          // number of constant leaf nodes (weights, etc.)
    ggml_tensor ** nodes; // mutable tensors (operator nodes), topologically ordered
    ggml_tensor ** leafs; // constant tensors (immutable data)
    ggml_tensor ** grads; // gradients (for training, nullptr during inference)
    ggml_tensor ** grad_accs;
    int32_t * use_counts;
    ggml_hash_set visited_hash_set;
    enum ggml_cgraph_eval_order order; // LEFT_TO_RIGHT or RIGHT_TO_LEFT
    uint64_t uid;  // graph identifier, used for reuse detection (0 means not set)
};

Key design: nodes is a topologically ordered operator sequence. Each ggml_tensor node's src[] array points to predecessor nodes, forming a DAG. During execution, forward computation proceeds in nodes[0..n_nodes-1] order. uid is a graph identifier for recognizing identical graph topologies across calls — e.g. the CUDA backend keys its CUDA Graph cache on uid (0 means unset/ignored).

1.2 llm_graph_context — Graph Builder Base Class

Defined in src/llama-graph.h:950, this is the base class for all model graph construction:

struct llm_graph_context {
    const llm_arch arch;
    const llama_hparams & hparams;
    const llama_cparams & cparams;
    const llama_ubatch  & ubatch;
    // ... model dimension parameters (n_embd, n_layer, n_head, n_rot, ...)

    ggml_backend_sched_t sched;
    ggml_backend_t backend_cpu;
    const llama_memory_context_i * mctx; // KV cache memory context

    ggml_context * ctx0; // ggml memory pool (nodes are allocated here)
    ggml_cgraph  * gf;  // the graph to be filled

    llm_graph_result * res; // output result container

    // common builder methods
    ggml_tensor * build_inp_embd(...);
    ggml_tensor * build_norm(...);
    ggml_tensor * build_lora_mm(...);
    ggml_tensor * build_qkv(...);
    // ... etc.
};

Each concrete model (e.g., llama_model_llama, llama_model_qwen2) inherits from this class and implements the specific forward logic in its graph::graph() constructor.

1.3 llm_graph_result — Graph Execution Result Container

Defined in src/llama-graph.h:859:

class llm_graph_result {
    ggml_tensor * t_inp_tokens;  // input token ids
    ggml_tensor * t_logits;      // output logits
    ggml_tensor * t_embd;        // hidden state (for embedding extraction)
    ggml_tensor * t_embd_pooled; // pooled embedding
    ggml_tensor * t_h_nextn;     // hidden state for MTP/NextN

    std::vector<ggml_tensor *> t_layer_inp; // per-layer input (for speculative decoding)

    std::vector<llm_graph_input_ptr> inputs;  // input tensor set
    std::vector<llm_graph_fused_node> fused_nodes; // fused nodes

    ggml_context_ptr ctx_compute;
    ggml_cgraph * gf;
    int64_t max_nodes;
};

Graph reuse detection: llm_graph_result::can_reuse(params) (declared at src/llama-graph.h:888; the comparison logic is in llm_graph_params::allow_reuse(), src/llama-graph.h:738) compares whether the new params are topologically equivalent to the previous step's params: arch, gtype, cvec, loras, cparams flag bits (embeddings/causal_attn/nextn_layer_offset), ubatch structure (n_tokens/n_seq_tokens/n_seqs/equal_seqs, and the seq id set when equal_seqs is split), n_outputs, sampler set (including its output tensor bindings). Key invariant: the graph topology is a deterministic function of these params —— if the params are equivalent, the topology is necessarily identical. Therefore, on reuse, the graph is not rebuilt and ggml_backend_sched_alloc_graph() is not called, only the input tensor data is updated (res->set_inputs(&ubatch)). Note that n_past does not participate in the comparison: the KV cache position is data (the idx/position input tensors filled at each step), not graph structure.


2. End-to-End Inference Flow

2.1 Call Chain

llama_decode(batch)
  └─ llama_context::decode(batch)
       ├─ balloc->init(batch)          // initialize batch allocator
       ├─ while (has ubatch):
       │    └─ process_ubatch(ubatch, gtype, mctx)
       │         ├─ 1. mctx->apply()                    // apply KV cache memory context
       │         ├─ 2. check graph reuse
       │         │     if (res->can_reuse(gparams)):
       │         │         reuse previous graph directly
       │         │     else:
       │         │         model.build_graph(gparams)     // build ggml_cgraph
       │         │           └─ dispatch → llama_model_xxx::build_arch_graph()
       │         │                 └─ construct graph object
       │         │         ggml_backend_sched_alloc_graph(sched, gf)
       │         ├─ 3. res->set_inputs(&ubatch)          // copy input data to GPU tensors
       │         └─ 4. graph_compute(gf)                 // execute computation
       │              └─ ggml_backend_sched_graph_compute_async(sched, gf)
       ├─ extract logits / embeddings
       └─ sample next token

2.2 Graph Construction Process

Using a standard Transformer decoder as an example (e.g., src/models/llama.cpp):

// 1. input embedding
inpL = build_inp_embd(model.tok_embd);  // token ids → embedding

// 2. per-layer processing
for (int il = 0; il < n_layer; ++il) {
    inpSA = inpL;  // residual connection save point

    // Pre-norm
    cur = build_norm(inpL, model.layers[il].attn_norm, nullptr, LLM_NORM_RMS, il);

    // Q/K/V projection (may include LoRA)
    auto [Qcur, Kcur, Vcur] = build_qkv(layer, cur, ...);

    // RoPE positional encoding
    Qcur = ggml_rope_ext(ctx0, Qcur, inp_pos, nullptr, n_rot, ...);
    Kcur = ggml_rope_ext(ctx0, Kcur, inp_pos, nullptr, n_rot, ...);

    // KV Cache write + Attention (KV write and wo projection are both done inside build_attn)
    cur = build_attn(inp_attn, layer.wo, layer.wo_b, layer.wo_s,
                     Qcur, Kcur, Vcur, nullptr, nullptr, nullptr, kq_scale, il);

    // residual connection
    inpL = ggml_add(ctx0, inpSA, cur);

    // FFN (SwiGLU: silu(gate(x)) * up(x))
    cur = build_norm(inpL, layer.ffn_norm, nullptr, LLM_NORM_RMS, il);
    gate = build_lora_mm(layer.ffn_gate, cur);
    up   = build_lora_mm(layer.ffn_up,   cur);
    cur  = ggml_mul(ctx0, ggml_silu(ctx0, gate), up);
    cur  = build_lora_mm(layer.ffn_down, cur);
    inpL = ggml_add(ctx0, inpL, cur);  // residual
}

// 3. output layer
cur = build_norm(inpL, model.output_norm, nullptr, LLM_NORM_RMS, -1);
cur = build_lora_mm(model.output, cur);  // lm_head
ggml_build_forward_expand(gf, cur);       // register into graph
res->t_logits = cur;

Each build_xxx() call creates a ggml tensor node on ctx0, automatically establishing src[] dependency relationships. Finally, ggml_build_forward_expand() recursively adds the result node and all its predecessors into gf->nodes[].

2.3 Backend Scheduler Execution

ggml_backend_sched_graph_compute_async(sched, gf)
  └─ ggml_backend_sched_split_graph(sched, gf)  // if needed
       ├─ Pass 1: assign backend for each node (prefer GPU)
       ├─ Pass 2: expand GPU coverage up/down to reduce cross-backend copies
       └─ Pass 3: split by backend assignment into splits (contiguous subgraphs)
  └─ for each split:
       ├─ sync previous split (ensure previous split has completed)
       ├─ copy inputs to split's backend (insert copy nodes for cross-backend transfers)
       ├─ ggml_backend_graph_compute(backend, subgraph)
       └─ copy outputs to next split's backend

3. Backend Scheduler Multi-Backend Dispatch

3.1 split_graph() — Graph Partitioning Algorithm

Defined in ggml/src/ggml-backend.cpp:1055.

Core idea: Split the entire ggml_cgraph into several contiguous subgraphs (splits) by backend assignment, with each split executed on the same backend.

The actual algorithm is 5 passes (not a simple "three-pass scan"):

  1. Pass 1 — Initial assignment: iterate all leaves and nodes; for tensors without an explicit backend, assign by backend_id_from_cur() (i.e. the backend of the buffer holding the data, typically the GPU where weights live), without overriding user-specified assignments.

  2. Pass 2 — Assignment expansion: a total of 4 sub-passes (in order: expand GPU down → expand GPU up → expand rest down → expand rest up). The first two only expand non-CPU GPU backends (cleared and skipped when cur_backend_id == n_backends - 1, i.e. CPU); the last two expand the remaining unassigned nodes to the current backend (including CPU). Result: CPU is used only when weights are on CPU, or there is a CPU-only op between GPUs; unsupported ops are left empty for later handling.

  3. Pass 3 — Upgrade + fallback: for already-assigned nodes, if a backend with the "same buffer type and higher priority" supports the op and all srcs are compatible, upgrade (e.g. when BLAS/CPU share the host buffer type, a CPU node can be upgraded to BLAS); for still-unassigned nodes, choose the backend that "supports the most already-assigned inputs".

  4. Pass 4 — src/view completion: a view node shares its backend with its view_src; remaining unassigned srcs inherit the dst's backend; if still empty, choose the first supporting backend (GGML_ASSERT guarantees one exists —— therefore a CPU fallback must exist).

  5. Pass 5 — Splitting: iterate nodes[]; apart from "adjacent nodes' backend changes", the following cases also start a new split: the current node's weight src (GGML_BACKEND_BUFFER_USAGE_WEIGHTS) is on a different and incompatible backend (in which case the previous split's GPU memory can be reused); a split's input/output tensor count exceeds the limit. This produces the splits[] sequence, and records the input/output tensors each split needs to copy across backends.

3.2 graph_compute() — Per-Split Execution

for (int i = 0; i < n_splits; i++) {
    split = &splits[i];

    // 1. Ensure previous split has completed (buffer may be reused)
    if (prev_backend != split->backend) {
        ggml_backend_synchronize(prev_backend);
    }

    // 2. Copy input tensors to current split's backend
    for (input : split->inputs) {
        ggml_backend_tensor_copy(input, split_backend);
    }

    // 3. Execute current split's subgraph
    ggml_backend_graph_compute(split_backend, split->cgraph);

    // 4. Copy output tensors to next split's backend
    //    (via events for async operation, does not block current backend)
}

Implementation details (ggml_backend_sched_compute_splits(), ggml-backend.cpp:1594):

  • Synchronization with the previous split is done via ggml_backend_event (event_synchronize/event_wait), degrading to ggml_backend_synchronize only when no event exists.
  • User input tensors (GGML_TENSOR_FLAG_INPUT) must be copied immediately and synchronously, to prevent the user from modifying data before the copy completes.
  • MoE weight optimization: when the split's first node is MUL_MAT_ID (MoE expert matmul) and the input weights are in a host buffer, it reads the expert id tensor and copies only the contiguous expert blocks used this time to the GPU (including trailing padding to prevent NaN), significantly reducing cross-backend copy volume.

3.3 Multi-Backend Scenario Example

Graph:  [Embedding(CPU)] → [MatMul(GPU)] → [RMSNorm(CPU)] → [MatMul(GPU)] → [LMHead(CPU)]

Split 1: Embedding         (CPU)
Split 2: MatMul            (GPU)
Split 3: RMSNorm           (CPU)  — GPU does not support this op
Split 4: MatMul            (GPU)
Split 5: LMHead            (CPU)

Tensor copy nodes are automatically inserted between splits. Async transfer via ggml_backend_event does not block GPU computation.


4. Graph Reuse Optimization

4.1 Reuse Conditions

The reuse decision is implemented by llm_graph_params::allow_reuse() (src/llama-graph.h:738), called by llm_graph_result::can_reuse() (src/llama-graph.h:888). Comparison items:

  • arch — model architecture
  • gtype — graph type (decode/prefill/MTP draft, etc.)
  • cvec / loras / cross — adapter pointers
  • cparams.embeddings / cparams.causal_attn / cparams.nextn_layer_offset etc.
  • ubatch structure: n_tokens, n_seq_tokens, n_seqs, n_seqs_unq, equal_seqs, token/embd input shapes; when equal_seqs is split, it also compares each seq's seq_id one by one
  • n_outputs — the number of output tokens
  • samplers — sampler set (and the output[i]/seq_id[i][0] bindings when a sampler is present)

n_past (the KV cache position) does not participate in the comparison —— it only affects the input data (the values of the idx/position tensors), not the graph topology.

4.2 Reuse Benefits

  • Skip graph reconstruction: Avoid build_graph() rebuilding the DAG
  • Skip memory allocation: ggml_backend_sched_alloc_graph() is an expensive operation
  • Retain buffers: GPU buffers are not released, reused directly

For the decode phase (n_tokens=1 each time), the graph topology is nearly unchanged, resulting in very high reuse rates.


5. Key Graph Construction Patterns

5.1 build_inp_embd() — Input Embedding

token_ids [n_tokens] → ggml_get_rows(embeddings) → inpL [n_embd, n_tokens]

When the input is embeddings rather than token ids (e.g., multimodal), ubatch.embd is used directly.

5.2 build_norm() — Layer Normalization

Supports both RMSNorm and LayerNorm, distinguished by the LLM_NORM_RMS / LLM_NORM enum.

5.3 build_qkv() — Q/K/V Projection

Combines the Q, K, V projections into one or more matrix multiplications (supports both fused wqkv and separate wq/wk/wv paths), returning a struct (defined at src/llama-graph.h:937):

struct llm_graph_qkv {
    ggml_tensor * q; // [n_embd_head, n_head,    n_tokens]
    ggml_tensor * k; // [n_embd_head, n_head_kv, n_tokens]
    ggml_tensor * v; // [n_embd_head, n_head_kv, n_tokens]
};

5.4 build_lora_mm() — Matrix Multiplication with LoRA

Actual signature (src/llama-graph.h:1006, weights first): build_lora_mm(w, cur, w_s = nullptr):

ggml_tensor * build_lora_mm(ggml_tensor * w, ggml_tensor * cur,
                            ggml_tensor * w_s = nullptr) {
    res = ggml_mul_mat(w, cur);        // w @ cur
    if (w_s) res = ggml_mul(res, w_s); // per-tensor scale
    for (lora : *loras) {
        ab_cur = lora.b @ (lora.a @ cur);  // B @ (A @ cur)
        res = ggml_add(res, ggml_scale(ab_cur, scale));
    }
}

LoRA adapters are statically unrolled during graph construction, adding no runtime branching. Note: whether LoRA is unrolled affects the loras param and thus the allow_reuse() reuse decision (switching a LoRA adapter changes the graph topology and triggers a rebuild).

5.5 build_attn() — Attention

Supports multiple attention modes:

  • MHA (Multi-Head Attention)
  • GQA (Grouped-Query Attention)
  • MLA (Multi-head Latent Attention, DeepSeek)
  • SWA (Sliding Window Attention)
  • Flash Attention (when cparams.flash_attn is enabled)

5.6 KV Cache Interaction

inp_attn = build_attn_inp_kv()      // create KV cache input descriptor (register idx/mask etc. input tensors)
...
cur = build_attn(inp_attn, wo, ..., Qcur, Kcur, Vcur, kq_scale, il)
        ├─ ggml_build_forward_expand(Qcur/Kcur/Vcur)   // expand first, prevent reordering
        ├─ mctx_cur->cpy_k(ctx0, Kcur, k_idxs, il)     // ★ write KV cache (done inside build_attn)
        ├─ k = mctx_cur->get_k(ctx0, il)               // read history K (a view of the cache tensor)
        ├─ v = ggml_view_4d(ctx0, k, ...)              // V is a tail view of the K tensor (when KV is stored together)
        └─ build_attn_mha(q, k, v, kq_mask, ...)       // attention

Note: the inp_attn->set_input_kv() interface does not exist. The KV write is a graph node done inside build_attn() via mctx_cur->cpy_k() (the write indices come from llm_graph_input_attn_kv::self_k_idxs, filled each step by set_input(ubatch)); the K/V read is a view of the KV cache tensor (get_k()), so the graph topology is independent of n_past —— the position exists only in the values of the input tensors, which is exactly why graph reuse works at every decode step.

KV cache memory management is abstracted by the llama_memory_context_i interface, supporting multiple implementations (standard cache, ISWA, DSA, MTP, etc.).


6. Comparison with minfer

Dimensionllama.cppminfer
Execution modelDeclarative DAG, build graph first then executeImperative forward, compute while building
Memory managementBackend scheduler auto-allocates + reusesManual tensor lifecycle management
Multi-backendAuto split + async copyManual GPU dispatch (src/metal/)
Operator fusionDone at graph-construction time by the model code / backend kernel layer (LLM_FUSED_OP_FLASH_ATTN etc.; the CUDA backend uses ggml_can_fuse to fuse op sequences at the kernel layer; CUDA Graph capture is keyed on uid), and the scheduler has no general fusion passAlready has manual fusion (GPU: swiglu/attn_bias_rope_store/fused qkv+gu kernel; CPU: batched matmul)
Graph reusecan_reuse() skips reconstructionRecompute every step
Complexity~3800 lines graph framework + ~100-200 lines per modelNo independent graph layer, direct implementation in forward.rs
FlexibilityNew models only need to inherit llm_graph_contextNew models require writing complete forward
Debugging capabilityggml_graph_dump_dot() exports DOT graphNo graph structure, difficult to visualize globally

7. Design Summary

llama.cpp's compute graph architecture is one of its core competitive advantages:

  1. Graph-execution separation: Building ggml_cgraph is a pure CPU operation, fast and side-effect-free; execution is uniformly dispatched by the backend scheduler
  2. Multi-backend transparency: Model code is unaware of specific hardware; the scheduler automatically handles GPU offload and data transfer
  3. Graph reuse: For fixed batch size decode scenarios, skipping reconstruction significantly reduces latency
  4. Extensibility: Each model architecture only needs to inherit llm_graph_context and implement graph construction, reusing all infrastructure
  5. Debuggability: DOT format export, callback mechanism, and tensor naming facilitate tracing

The tradeoff is higher code complexity, but this architecture makes llama.cpp a true inference runtime (rather than a simple matrix multiplication library), capable of efficiently supporting hardware configurations from single CPU to multi-GPU.


Reference Files

FileContent
ggml/src/ggml-impl.h:329ggml_cgraph struct definition
ggml/src/ggml.cGraph operation implementations (ggml_new_graph, ggml_build_forward_expand, ggml_graph_dump_dot)
ggml/src/ggml-backend.cpp:1055ggml_backend_sched_split_graph() graph partitioning (5 passes)
ggml/src/ggml-backend.cpp:1594ggml_backend_sched_compute_splits() per-split execution (including MoE expert partial copy)
ggml/src/ggml-backend.cpp:1961ggml_backend_sched_graph_compute_async() entry
ggml/include/ggml-backend.hBackend scheduler API documentation
src/llama-graph.h:738llm_graph_params::allow_reuse() reuse decision
src/llama-graph.h:859llm_graph_result result container
src/llama-graph.h:888llm_graph_result::can_reuse()
src/llama-graph.h:950llm_graph_context graph builder base class
src/llama-graph.cppset_input()/build_attn() (including KV write cpy_k)/build_lora_mm() implementation
src/llama-context.cpp:1325process_ubatch() end-to-end flow
src/llama-context.cpp:2475graph_compute() execution entry point
src/models/llama.cppLLaMA model graph construction example
src/models/qwen2.cpp:53Qwen2 model graph constructor (llama_model_qwen2::graph::graph)
src/models/models.hAll model architecture declarations (nested struct graph : public llm_graph_context)

llama.cpp MMQ — quantized-weight int8 tensor-core matmul: structure, numerics, SASS census, and minfer contrast

This document consolidates every established fact about llama.cpp's MMQ ("matrix-matrix quantized") GEMM from minfer's CUDA optimization campaign (P6 r6–r25, docs/CUDA_OPTIMIZATION.md), cross-checked line-by-line against the actual source. It is the analysis side of the "why is llama.cpp MMQ fast" question; the empirical records it consolidates are the r20 stall table and the r25 SASS opcode-class census, both reproduced verbatim below.

Sources. llama.cpp at ca3d5a3e1 (matches the mul_mat_q<12,128,0> that was profiled): ggml/src/ggml-cuda/{mmq.cuh, mmq-vec-dot.cuh, mmq-load-tiles.cuh, mma.cuh, mmq-config-ampere.cuh, mmq.cu, quantize.cu}, plus ggml/src/ggml-common.h and ggml/src/ggml-quants.c. minfer at the r25 HEAD: src/cuda_kernels.cu (today src/cuda/kernels/mmq_raw.cu; mmq_raw_wide_nt_kernel<KDR> + mmq_stage_b). Where the campaign document and the source disagree, the discrepancy is listed in §Corrections rather than silently repeated. (This write-up was re-verified against the source line-by-line by the post-r25 audit; the audit's corrections are the resolved items in §Corrections.)

Reading convention: x = src0 = the weight matrix (only quantized tensors reach MMQ; it is kept RAW in smem); y = src1 = the activation tokens (quantized to q8_1 on device); dst is the fp32 output. In mul_mat_q_process_tile the weight tile is tile_x and the activation tile is tile_y (mmq.cuh:889-891).


1. Scope & dispatch

MMQ is the quantized-weight path for ggml_mul_mat on CUDA: it runs the integer tensor-core MMA (mma.m16n8k32.s32.s8.s8.s32) directly over the raw quantized weight bytes and a q8_1-quantized copy of the activations, instead of dequantizing weights to f16 and running an f16 wmma GEMM. There is no f32/f16 weight materialization anywhere in the hot loop (unlike minfer's default 8p f16 wmma path).

Which file implements it. ggml-cuda/mmq.cuh (kernel + launch), mmq-vec-dot.cuh (per-warp dot/accumulate), mmq-load-tiles.cuh (weight staging), mma.cuh (tile/mma wrapper), mmq-config-*.cuh (per-arch config tables), and mmq.cu (host dispatch). The mul_mat_q kernel family is the entry point (mmq.cuh:952-1237).

When it dispatches. ggml_cuda_should_use_mmq (mmq.cu:259-386) decides. For Q4_K the type is in the supported switch (mmq.cu:277). The decisive rule on NVIDIA is Turing+: if (turing_mma_available(cc)) return true; (mmq.cu:312-314) — i.e. on GB10 (Blackwell, sm_121; turing_mma_available = NVIDIA && highest-compiled-arch ≥ Turing, common.cuh:348-350) MMQ is chosen unconditionally for every supported quantized type, for any batch size. (Unconditionally for the MMQ-vs-cuBLAS question — the ggml_cuda_mul_mat chain upstream asks MMVQ first for ne11 ≤ 8, so small batches never reach this predicate; see §12 for the full chain and the MMVQ kernel structure.) So the "prefill threshold" framing is an AMD-only idea: the ne11 < MMQ_DP4A_MAX_BATCH_SIZE (=64) gate (mmq.cu:327, constant at mmq.cuh:8) sits inside the NVIDIA-only branch (if (GGML_CUDA_CC_IS_NVIDIA(cc)), mmq.cu:326-328) and applies only when turing_mma_available is false; the AMD/RDNA gates are separate (mmq.cu:330-385, incl. the ne11 <= 128/256 RDNA branch at mmq.cu:337-345). On GB10 the only extra requirement is ≥48 KiB per-block smem (mmq.cu:303-310).

Ubatch shape. The "nt-512" the campaign profiled is not a dispatch threshold; it is the ubatch (ubatch size) that llama.cpp splits the prefill into. -p 2600 becomes 512-token ubatches (llama-bench: -p 2600 → 512-token ubs), so each mul_mat_q sees src1→ne1 = 512 tokens → 4 J=128 token tiles (ntx = 512/128 = 4). The campaign profiled mul_mat_q<12,128,0> at exactly this shape.

The q4_K instantiation. The profiled kernel is mul_mat_q<GGML_TYPE_Q4_K, 128, false> (= <12,128,0>: type 12 = Q4_K, J=128, fallback=0). Its config, from ggml_cuda_mmq_get_config_ampere (mmq-config-ampere.cuh:172): CASE(GGML_TYPE_Q4_K, 256, 1, 128, 128, GGML_CUDA_MMQ_SRAM_LAYOUT_Q8_1, MMQ_ITER_K, true, false) → nthreads=256, occupancy=1, I=128 (od-rows), J=128 (tokens), sram_layout=Q8_1, K_vram=MMQ_ITER_K=256, stream_k=true, fallback=false. Template meaning per ggml_cuda_mmq_config (mmq.cuh:165-178): I = SRAM tile width in src0→ne1 / dst→ne0 (the od-dim), J = SRAM tile width in src1→ne1 / dst→ne1 (the token dim), sram_layout = the weight-tile byte layout, K_vram = logical K per inner loop (= MMQ_ITER_K = 256, mmq.cuh:9). fallback toggles out-of-bounds guards in the od direction (selected at mmq.cuh:1560-1566 by nrows_x % 128 == 0).

MoE batch edge cases. ggml_cuda_mul_mat_q has a second (MoE, ids != nullptr) arm (mmq.cu:179-256) that only differs in how the activation rows are gathered and scattered: it builds an ids_src1/ids_dst inverse map + expert_bounds (mmq.cu:187-189) via ggml_cuda_launch_mm_ids_helper (mmq.cu:200-201), and for the gate/up broadcast case (dedup_bcast = ne11 == 1 && n_expert_used > 1, mmq.cu:193) it quantizes each token once and scatters to its compact rows through the ids_src1 map (quantize_scatter_mmq_q8_1_cuda, mmq.cu:233-234). Both arms pad the activation's inner dim to a row-multiple: ne10_padded = GGML_PAD(ne10, MATRIX_ROW_PADDING) (mmq.cu:120), which sizes the src1_q8_1 buffer (mmq.cu:136-138, 205-207); the s12/s13 strides and the ntx/nty grid are computed from ne10_padded, not ne10.

2. Block / warp / tile geometry

ParamValueSource
threads/block256 (8 warps)config nthreads=256; mmq-config-ampere.cuh:172
warp size32mmq.cuh:969
block od tile (I)128mmq-config-ampere.cuh:172
block token tile (J)128mmq-config-ampere.cuh:172
rows_per_warp32 (J=128 → J>=48 && J%16==0 → 32)mmq.cuh:180-186
ntx (x-minitiles/warp)rows_per_warp / tile_C::I = 32/16 = 2mmq-vec-dot.cuh:377
accumulator regssum[J*I/(nwarps*warp_size)] = 128·128/256 = sum[64]mmq.cuh:903
mma per vec_dot (32-k chunk)j0:8 × k01:4 (step QI8_1=8) × ntx:2 = 64 mma.m16n8k32 (16 per k01 sub-iter)mmq-vec-dot.cuh:414-440

Warp-to-tile mapping (mmq-vec-dot.cuh:389-390): the 8 warps split into 4 od-groups by i0 = (threadIdx.y/ntx)*rows_per_warp = (ty/2)*32, and within each od-group the two warps de-interleave the token stream via y += (ty%ntx)*(tile_C::J*MMQ_TILE_Y_K) (mmq-vec-dot.cuh:379). Each od-group covers rows_per_warp=32 od-rows (ntx=2 × tile_C::I=16 minitiles) and the full J=128 token tile, but each warp covers only 64 token columns: a warp does J/(ntx·tile_C::J) = 128/(2·8) = 8 j0-steps of tile_C::J=8 token columns, and the pair jointly covers all 128. Each j0-step × k01-step × n issues one tile_C::ne = 16·8/32 = 4-register C fragment per minitile (mma.cuh:227) — so the 64 mma instances per 32-k chunk (per vec_dot call) fill the 64 sum registers, each accumulated 4× (once per k01 sub-iteration).

tile_C / fragment mapping. The mma wrapper is mma(D, A, B) with tile<16,8,int> D (C), A, and tile<8,8,int> B (mmq-vec-dot.cuh:370-372), which expands to mma.sync.aligned.m16n8k32.row.col.s32.s8.s8.s32 (mma.cuh:946). Because each int holds four int8 (K=32 packed into K/4=8 words), tile<16,8,int>::ne = 16·8/32 = 4 registers per thread (mma.cuh:227). The CAMPAIGN'S "I·J/32 = 4 regs per m16n8k32 C" reading is confirmed (tile<I,J,T,DATA_LAYOUT_I_MAJOR>: ne = I*J/32 on the NVIDIA Turing+ branch, mma.cuh:226-227); the I·J/64 = 2 figure belongs to the AMD MFMA branch (#if defined(AMD_MFMA_AVAILABLE), mma.cuh:107-108) and was corrected in r11.

Lane-level C-fragment map (NVIDIA tile<16,8,int>, DATA_LAYOUT_I_MAJOR). For the C (sum) fragment each lane holds ne = 4 values mapped by get_i(l) = (l/2)*8 + threadIdx.x/4 and get_j(l) = (threadIdx.x%4)*2 + (l%2) (mma.cuh:245,262 — note the audit's pointer to mmq-vec-dot.cuh is off by file; the helpers live in mma.cuh). So a lane owns a 2×2 block of the 16×8 C tile at rows {threadIdx.x/4, threadIdx.x/4+8} and columns {(threadIdx.x%4)*2, (threadIdx.x%4)*2+1}: l=0,1 sit on row tid/4 (cols (tid%4)*2 and +1), l=2,3 on row tid/4+8. Because tile_C::J=8, these per-lane columns are the token columns the mma writes, and the get_j map is what ties the C fragment's columns to the j0-stepped token tile.

3. Activation pipeline

The activation (src1, f32) is quantized to block_q8_1_mmq once per GEMM launch, before the kernel, not per block. In ggml_cuda_mul_mat_q (mmq.cu:85-256) a vmem pool buffer src1_q8_1 is allocated (mmq.cu:138) and filled by a separate quantize kernel — quantize_mmq_q8_1 (quantize.cu:458) launched via quantize_mmq_q8_1_cuda (quantize.cu:575) inside ggml_cuda_mul_mat_q (mmq.cu:156-157 non-MoE, :236-237 MoE; quantize_scatter_mmq_q8_1_cuda for the MoE dedup_bcast arm, :233-234) — then tile_y is bulk-loaded from it (mmq.cuh:909-939). So llama does not fuse the activation quantize into the MMQ kernel; it runs a per-GEMM q8_1 quantize pass as a separate kernel (once per launch, not per block in the hot loop). The minfer asymmetry stands — llama quantizes the activations once per GEMM call while minfer's default path pays its own r23 "convert f32→f16" 73 ms per-launch tax — but the cost is of the same order: minfer's r23 default-path accounting should credit llama with a quantize-pass cost similar to its own convert tax. (The q8_1 quantize is cheaper than minfer's f32→f16 convert; the point is that "llama pays zero" is not correct.)

block_q8_1_mmq (mmq.cuh:27-46) is a 128-element block (QK8_1_MMQ = 4·QK8_1 = 128): a leading 16-byte union of scales (d4[4], ds4[4], or d2s6[8]) plus int8_t qs[128]; sizeof == 144 B (mmq.cuh:56-57). The layout comment (mmq.cuh:28-36) states the y data is grouped into 128-value blocks, transposed, and each block padded with 16 bytes, the pad reused to store the block scale and partial sum — this is the "d/ssum in pad bytes" claim. For Q4_K/Q5_K the DS4 layout is used (mmq.cuh:82-84): half2 ds4[4] carries one 16-bit scale + one 16-bit partial sum per 32 values (d0,s0,d1,s1,…).

Inside the kernel the activation tile tile_y has row stride MMQ_TILE_Y_K = 36 ints = 144 B (mmq.cuh:119; MMQ_TILE_NE_K + MMQ_TILE_NE_K/QI8_1 = 32 + 4, since QI8_1 = QK8_1/(4·QR8_1) = 32/4 = 8, ggml-common.h:124,258). The row is [scale || qs]: the q8_1 scale union — half2 ds4[4] in the DS4 layout used for Q4_K/Q5_K, i.e. 4 ints — is read at (half2*)y and the 32-int qs plane at y+4 (mmq-vec-dot.cuh:383-384). So the smem row stride equals the global block_q8_1_mmq size (144 B, mmq.cuh:56-57); the two were only "different" earlier because the scale word was miscounted as 1 int instead of the DS4 layout's 4 ints.

4. Weight staging & smem layout

load_tiles_q4_K (mmq-load-tiles.cuh:703-812) stages the raw q4_K weight into tile_x in the MMA data layout. Per 256-k super-block per row the weight is 144 B raw (block_q4_K = d[2] + dmin[2] + scales[12] + qs[128]). In smem the weight stays raw nibbles, one per byte (0..15), not expanded to signed int8:

x_qs[i*sram_stride + 16*(txi/8) + txi%8 + 0] = (qs0 >> 0) & 0x0F0F0F0F;   // mmq-load-tiles.cuh:736
x_qs[i*sram_stride + 16*(txi/8) + txi%8 + 8] = (qs0 >> 4) & 0x0F0F0F0F;   // mmq-load-tiles.cuh:737

The 0x0F0F0F0F mask isolates each 4-bit nibble into its own byte; no __vsubss4 centering (no dmin subtraction at this point) — the dmin is instead folded into the scale: `x_dm[i*sram_stride

  • 4*ksc + l] = (bxi->dm * make_half2(1.0f,-1.0f)) * make_half2(sc8[l], m8[l])(mmq-load-tiles.cuh:772-777), so the half2 holds(d·sc, −dmin·m)` per 32-value sub-block.

The smem row stride is sram_stride = ggml_cuda_mmq_get_sram_stride(GGML_CUDA_MMQ_SRAM_LAYOUT_Q8_1) = 2·MMQ_TILE_NE_K + 2·MMQ_TILE_NE_K/QI8_1 + 4 = 64 + 8 + 4 = **76 ints (304 B)** (mmq.cuh:137; 2·MMQ_TILE_NE_K/QI8_1 = (2·32)/8 = 8, not 2 — the earlier "70 ints (280 B)" undercounted the second term). K%8 == 4 is statically enforced (mmq.cuh:153-159). The "+4" is the 16-byte pad that makes consecutive rows rotate bank phases; the nibble plane occupies the leading 2·MMQ_TILE_NE_K = 64 ints and the half2 scale plane follows at offset 64 (mmq-vec-dot.cuh:381-382).

Barrier structure. Per MMQ_ITER_K = 256 k-iteration, mul_mat_q_process_tile (mmq.cuh:907-940) does load_tiles(weight) + stage y-half-1 → barrier → vec_dot(k00=0) → barrier → stage y-half-2 → barrier → vec_dot(k00=32) → barrier = 4 barriers per 256 k = 2 per 128 k, identical to minfer at MMQ_KD=8 (256-k) — this is what the r20 "barrier density at parity" finding verified (structurally established from the source, r20).

Synchronous staging, no cp.async. The weight and activation tiles are staged with plain global→(register)→smem stores (mmq.cuh:907-939), ordered by __syncthreads. There is no cp.async / TMA pipeline in the classic (non-Blackwell-fp4) MMQ path; the latency hiding comes solely from having enough resident warps, not from prefetch depth.

Q6_K specialization — different scale layout from q4_K. Q6_K uses GGML_CUDA_MMQ_SRAM_LAYOUT_Q6_K with its own load (ggml_cuda_mmq_load_tiles_q6_K, mmq-load-tiles.cuh:938) and vec_dot (ggml_cuda_mmq_vec_dot_q6_K_q8_1_mma, mmq-vec-dot.cuh:1018), and no raw-nibble dmin fold. Its smem row stride is 2·MMQ_TILE_NE_K + MMQ_TILE_NE_K/QI6_K + MMQ_TILE_NE_K/8 + 7 = 64 + 1 + 4 + 7 = 76 ints (mmq.cuh:143; QI6_K = QK_K/(4·QR6_K) = 256/8 = 32, ggml-common.h:139). The q6_K nibbles are fully centered at load: x_qs[...] = __vsubss4(ql | qh, 0x20202020) (mmq-load-tiles.cuh:982-983) subtracts 32 from each byte, producing signed int8 in [−32,31] — no dmin term later (unlike q4_K, which keeps raw 0..15 nibbles and removes dmin at accumulate). The scale layout is split rather than a half2: one float d per row (x_df[i*(MMQ_TILE_NE_K/QI6_K) + i/QI6_K] = bxi->d, mmq-load-tiles.cuh:1003) plus an int8 scale per 16-value sub-block (x_sc, mmq-load-tiles.cuh:1021; unpacked per byte at mmq-vec-dot.cuh:1115-1121). The rescale is therefore two-stage: tmp = (C0·scA0 + C1·scA1)·dB accumulated per j0-step, then sum += tmp·dA (mmq-vec-dot.cuh:1161,1170) — i.e. the signed int8 scale is folded with the float d at the end, a materially different flow from q4_K's sum += dmA·dsB·C.

5. Launch schedule

The host launch is launch_mul_mat_q (mmq.cuh:1393-1473). It reads nsm (mmq.cuh:1397) and decides between an xy-tiled grid and the stream-k grid:

const int ntiles_dst = ntx*nty*ntzw;                            // mmq.cuh:1439
const int tiles_nwaves = (ntiles_dst + nsm - 1)/nsm;            // mmq.cuh:1440
const int tiles_efficiency_percent = 100*ntiles_dst/(nsm*tiles_nwaves);  // mmq.cuh:1441
block_nums_stream_k = (NVIDIA && efficiency >= 90) ? ntiles_dst : nsm;    // mmq.cuh:1442

ntx/nty derivation. The grid dims are host-derived from the tile sizes, not hardcoded: nty = (nrows_x + I - 1)/I and ntx = (ncols_max + J - 1)/J (mmq.cuh:1410-1411); the kernel recomputes nty = (nrows_x + I - 1)/I from the passed nrows_x (mmq.cuh:974) and receives ntx as a fast-divisor param (ntx_fd, mmq.cuh:1421). This parameterization is what drives the stream-k ntiles_dst = ntx·nty·ntzw (mmq.cuh:1439) and the fixup ntiles_dst % blocks != 0 decision below — so ntx/nty are dictated by ncols_max/nrows_x (the ubatch shape), not by a dispatch threshold.

For the nt-512 q-proj (ntx=4, nty=28): ntiles_dst = 112, tiles_nwaves = ceil(112/48) = 3, efficiency = 100·112/144 = 77.8% < 90% → block_nums_stream_k = nsm = 48 → grid (48,1,1) persistent blocks, plus the fixup kernel because 112 % 48 ≠ 0 (fixup_needed mmq.cuh:1446).

Why the fixup exists. Stream-k splits the K-range (blocks_per_ne00 super-block count) across the 48 blocks (mmq.cuh:1066-1074), so a block may work on the tail of one output tile and the head of another. Two blocks can non-deterministically contribute to the same output tile, and because the accumulation is fp, the order of partial-sum adds is not reproducible. mul_mat_q_stream_k_fixup (mmq.cuh:1239-1375) runs as a second launch (grid (48,4,1), mmq.cuh:1454), reads the tmp_last_tile partials written by the fixup=true last iterator (mmq.cuh:1232-1236) and combines them into dst. This is the numeric cost of stream-k: the campaign measured the fixup at +34 μs (r20).

Write-back epilogue. The accumulator is already fp32 end-to-end: the C fragment is an int32 mma result, but it is converted at the rescale sum += dmA.x·dsB.x·C.x + dmA.y·dsB.y (mmq-vec-dot.cuh:434-437) and sum[] is float (mmq.cuh:903). So write_back is a plain per-value global store — dst[ids_dst[j]*stride + i] = sum[(j0/tile_C::J + n)*tile_C::ne + l] (ggml_cuda_mmq_write_back_mma, mmq.cuh:473-525, esp. :519; no NVFP4/y_scale rescale on the Q4_K path). For stream-k the fixup=true arm writes contiguous per-block partials instead: write_back(sum, ids_dst, tmp_fixup + blockIdx.x*(J*I), y_scale, I, I, J) (mmq.cuh:943), and the separate mul_mat_q_stream_k_fixup second launch runs only when ntiles_dst % block_nums_stream_k.x != 0 (fixup_needed, mmq.cuh:1446; allocated tmp_fixup :1450-1451, fixup launch :1464-1469).

Occupancy math. nbytes_shared = nbs_ids + nbs_x + GGML_PAD(nbs_y, nthreads·sizeof(int)) (mmq.cuh:1386-1391): nbs_ids = J·4 = 512, nbs_x = I·sram_stride·4 = 128·76·4 = 38,912, nbs_y = J·144 = 18,432. At the profiled I=J=128 shape GGML_PAD(18,432, 1024) = 18,432 (18,432 = 18×1024, so no pad step — the pad only applies for J not a multiple of 64, e.g. J=80) → 57,856 B ≈ 56.5 KB per block. With ~99 KB shared/SM on GB10 that is 1 block/SM, and it is additionally pinned to 1 by __launch_bounds__(nthreads, 1) (mmq.cuh:953). 8 warps over 4 schedulers → 2.00 active warps/sched — matching the r20 capture. The occupancy lever that raw-nibble smem buys is real for smaller J tiles (e.g. J=64 → nbs_y = 9,216, total 48,384 B ≈ 47.3 KB → 2 blocks/SM); at the profiled I=J=128 shape llama is 1 block/SM exactly like minfer.

6. Compute loop & numerics

The q4_K compute loop is the universal ggml_cuda_mmq_vec_dot_q8_1_q8_1_mma (dispatched at mmq.cuh:770), with the weight dequantize done once in load_tiles_q4_K. Core (mmq-vec-dot.cuh:369-442):

  • Loop granularity (NVIDIA path). vec_dot is called once per 32-k chunk (per k00), and each call issues 64 mma: j0 = {0,16,…,112} (8 steps, stride ntx·tile_C::J = 16) × k01 = {0,8,16,24} (4 steps, step QI8_1=8) × n = {0,1} (2) (mmq-vec-dot.cuh:414-440). A-frags = tile_A A[ntx][MMQ_TILE_NE_K/QI8_1] = A[2][4] = 8 A-frags (mmq-vec-dot.cuh:386, loaded once before the j0 loop), B-frags = 32 tile_B (one per j0×k01, :420, reused across the 2 n-iterations). Because the sum index does not depend on k01, each sum slot is accumulated 4× per vec_dot (once per k01 sub-iteration). (The campaign's "16 mma per 32-k chunk" is the per-k01 count, 8 j0 × 2 n.)
  • A-fragments tile_A[ntx] loaded by load_ldmatrix(A[n], x_qs + (i0+n·tile_A::I)·sram_stride + k0, sram_stride) (mmq-vec-dot.cuh:397) — ldmatrix.m8n8.x4 over the raw-nibble rows.
  • B-fragments tile_B loaded by load_generic(B, y_qs + j0·MMQ_TILE_Y_K + k01, MMQ_TILE_Y_K) (mmq-vec-dot.cuh:420) — plain LDS (the source comment: "faster than load_ldmatrix").
  • mma(C, A[n][k01/QI8_1], B) → mma.m16n8k32.row.col.s32.s8.s8.s32, int32 accumulate (mma.cuh:946; int-only C/D — the f32-accumulate spelling is rejected by ptxas, r15).
  • Per-chunk rescale. For each C value (mmq-vec-dot.cuh:434-437): sum[i] += dmA.x·dsB.x·C.x + dmA.y·dsB.y, where dmA = __half22float2(x_dm[...]) and dsB = __half22float2(y_dm[...]) (mmq-vec-dot.cuh:426,408). dmA.x = weight scale d·sc, dmA.y = −dmin·m, dsB.x = activation scale d, dsB.y = activation partial sum ssum. So the dmin correction is a rank-1 fold in (token, od-col) applied at accumulate time, exactly the term the campaign's r15 dma = da·(float)sa fold mirrors.

get_scale_min_k4 semantics (ggml-quants.c:880-887): decodes the packed 12-byte scales array of a q4_K/q5_K super-block into per-32-value (d_scale, dmin) pairs. For j < 4: d = scales[j] & 63, m = scales[j+4] & 63; else d = (scales[j+4]&0xF) | ((scales[j-4]>>6)<<4), m = (scales[j+4]>>4) | ((scales[j]>>6)<<4) — i.e. six 6-bit and six 4-bit-composed codes. The kernel does not call it directly; the mma path's load_tiles_q4_K applies the equivalent unpack via unpack_scales_q45_K (mmq-load-tiles.cuh:766-767). The semantic is: each 256-value super-block has 8 sub-blocks, each with its own (scale, dmin); a value v ∈ [0,15] dequantizes to d·s·v − dmin·m.

Where "dequant-at-use" happens. The weight nibbles are raw in smem (0..15); they are not signed-centered and not expanded at staging. The mma consumes them as the int8 operand (so the accumulator holds Σ nibble·act, with values in the unsigned range), and the centering (dmin) offset is removed by the dmA.y·dsB.y term in the fp rescale. So llama's "dequant" is split: nibble isolation at load (0x0F mask), dmin removal at accumulate. In contrast minfer byte-expands the nibbles to per-k int8 during staging (qb8) and folds dmin into a float2 (d, dmin·m) scale.

Rescale precision. At the CUDA level the rescale is fp32 (the float2 from __half22float2 multiplied and accumulated in fp32, mmq-vec-dot.cuh:434-437). See §Corrections for the tension between this source-level reading and the r25 census attribution.

7. Measured profile (r20 + r25, matched nt-512 q-proj)

Both captures are on the same matched layer-0 q-proj GEMM (nt=512, id=od=3584). mul_mat_q<12,128,0> runs grid (48,1,1) + fixup (grid 48,4,1). minfer mmq_raw_wide_nt_kernel<..> grid (4,28) = 112 short blocks, 2.33 waves.

r20 stall table (per-issue-active warp ratio; llama = launch 0 of 8, identical at launch 6):

metricminfer <4>minfer <8>llama.cpp <12,128>
duration q-proj (μs)632.4609.6263.6
issue /cyc/sched0.16–0.260.200.42
eligible /cyc0.22–0.280.280.64
warps active /cyc2.002.002.00
warp inst / tile (k)499488356
IMMA mma-inst1,605,6321,605,6321,605,632
long_scoreboard6.225.761.15
wait0.630.570.58
barrier0.260.170.19
short_scoreboard0.340.530.12
not_selected0.410.400.50
math_pipe_throttle0.360.340.47
mio_throttle0.110.280.25
lg_throttle0.330.600.09
dispatch_stall0.290.280.31
LDS bank conflicts3,211,2643,211,26442

r20 readings that survive: issue efficiency (0.42 vs 0.26) at equal occupancy (2.00 warps/sched) is the carrier; the stall gap is long_scoreboard (6.22 vs 1.15, 97% of the named excess — minfer warps spend ~86% of resident cycles on global-load latency vs llama ~24%); the r13 stall table was double-distorted (pre-r14 kernel at nt-2630 vs llama nt-512); IMMA is EXACTLY at parity (1,605,632 = mma.m16n8k32 count for the same GEMM). Launch shape (stream-k vs 3-wave), occupancy, tensor work and memory bytes are ruled out; the fixup costs +34 μs.

r25 SASS opcode-class census (per 128×128 output tile, fully reduced over K=id; pred_on / 32 → warp, validated to ~1.4% of smsp__inst_executed.sum):

SASS class (ncu opcode metric)ours/tiletheirs/tiledelta/tileratioverdict
integer ALU (IADD3/IMAD/LEA/SHF/SEL/ISETP/LOP3)113,55244,087+69,4652.58×← 77% of surplus
FP32 FMUL (rescale)72,59257,344+15,2481.27×dequant-rescale
conversion (I2FP/F2I)72,60058,254+14,3461.25×int-mma→fp32
misc (NOP/CS2R)13,3365,851+7,4852.28×loop/init
control-flow (BRA/isync)3,5841,160+2,4243.09×loop control
uniform datapath (UR)1,80855+1,75333×uniform regs
FP32 FFMA (rescale/accum)114,688114,688+01.00×identical
bit (LOP3/PRMT/SHF)8456−4480.02×(theirs more)
fp16 HADD2/HFMA path15,23236,400−21,1680.42×(theirs more)
memory (LDG/STS/LDS/LDSM)31,81636,836−5,0200.86×(theirs more)

Totals reconcile: ours smsp__inst_executed.sum = 49,900,928 → 445,544 warp-inst/tile; theirs 39,844,864 → 355,758; surplus +89,786/tile (25.2%). Dedicated warp-family memory metrics (per tile): global_ld ours 6,944 / theirs 6,384; LDSM ours 8,064 / theirs 1,792 (+6,272, 4.5×); shared_ld ours 8,960 / theirs 19,346 (2.2× more); shared_st ours 5,376 / theirs 7,730.

r20/r23 duration. The matched nt-512 q-proj GEMM is 263.6 μs (llama) + ~34 μs fixup; the per-tile wall is 113 μs (llama) vs 211 μs (minfer). The r23 cross-model figure for the full prefill GEMM class is llama MMQ 39.5 μs/GMAC (incl. fixup) vs minfer's default f16 path 59.5 μs/GMAC + 73 ms convert tax. LDS bank conflicts: llama ~42 vs minfer 3.2M (r20) / 16.86M op_ld + 6.47M op_st (r22, before the XOR swizzle zeroed op_ld).

8. Why it is fast — the design's logic

The campaign's cross-cutting conclusion (r13, r15, r20, r22, r25) is that llama's MMQ wins on two coupled axes, and the census (r25) proves the winning mechanism is not the mma or the memory path (both at parity/over-parity) but the support instruction stream and the occupancy that the smem budget buys:

  1. Raw-nibble weight plane halves the smem budget. mmq-load-tiles.cuh keeps the weight as 4-bit nibbles (one per byte, 0x0F mask) in the 70-int stride; the expanded-byte form (minfer's qb8 int8 per-k) is ~2× the bytes. At a smaller J this is the difference between 1 and 2 blocks/SM, i.e. between ~2 and ~4 warps/scheduler latency coverage. The tradeoff it accepts: the mma consumes raw 0..15 nibbles (accumulator carries the unsigned dot) and the dmin centering is deferred to a rank-1 fp rescale term.
  2. q8_1 pre-quantization of activations, once per GEMM — the Q8_1 SRAM layout carries the per-32 scale and the partial sum in the 16-B pad word, so the activation operand needs no in-loop dequant; dsB.y (partial sum) is available for free.
  3. fp16 scale path keeps the rescale op count low (per r25 attribution): the dequant-rescale runs through 16-bit half2 scale values rather than a full fp32 I2F → FP32 FMUL chain, so the per-chunk sum += dmA·dsB·C + dmA.y·dsB.y is few instructions. minfer's fp32 rescale pays more FMUL + I2FP (72,592 + 72,600 vs 57,344 + 58,254).
  4. Tight index math / no redundant work per MAC. The 32-od-row × 64-token warp shape (the two warps of an od-group jointly cover the full 128 tokens; see §2) halves the A-fragment loads per MAC vs a 16-row warp (r13: "their warp covers 2× the od-rows"), and the B-fragment is loaded once per (warp, j0-step) via plain LDS. This is the warp shape minfer's r17 remap tested — measured NEUTRAL for minfer at 1 block/SM (see §10 Direction C).
  5. The mma work itself is at hard parity — IMMA and FFMA are byte-identical between the engines (r20: 1,605,632; r25: FFMA 114,688/tile both, IMMA 2×MAC both). Compute is never the gap.

The tradeoffs it accepts: (a) the nibble is expanded/dequantized at use inside the accumulate (a per-32-k rescale term rather than a fully pre-centering staging); (b) the stream-k schedule requires a separate fixup pass (+34 μs) to reorder fp partial sums; (c) the tile is coupled to the ubatch shape — J=128 is chosen to minimize ntx for the 512-token ubatch, and mul_mat_q_switch_J picks the largest J that fits ntiles_x (mmq.cuh:1484-1500).

9. Contrast with minfer's MMQ

The campaign's kernel is mmq_raw_wide_nt_kernel<KDR> (mmq_raw.cu:222): block 128 od × 128 tokens, warp = 16 od-rows × full 128-token tile, 8 A-frags + 2 B-frags = 16 chains, sum[64]. Its staging layout (mmq_raw.cu:222):

operandminfer wide (r22/r25 HEAD)llama MMQ
activationqa8 32-B chunks, XOR-swizzled granules, d/ssum in sda_qQ8_1 SRAM, scale+partial in the 16-B pad
weightqb8 expanded per-k int8 (nibbles 0..15, one per byte), slot-major 48-B strideraw nibbles (0x0F mask), 70-int stride
scalesds float2 (d, dmin·m) — fp32 rescalehalf2 (d·sc, −dmin·m) — fp16-op code path
stagingsingle-buffer synchronous, split-phase (r20)synchronous, 4 barriers/256-k

The design differences and their measured consequences:

  • Expanded B (qb8) + fp32 rescale. minfer's smem budget at KD=8 is 98,304 B → 1 block/SM, vs llama's 56.5 KB (1 block/SM at I=J=128, but with the headroom to reach 2 at smaller J). The r25 census attributes minfer's +89,786 surplus to integer ALU +69,465 (77%) (address/predicate math for the 8-A-frag LDSM + staging bounds) and fp32 dequant-rescale +29,594 (FMUL+I2FP).
  • Measured efficiency. minfer issue 0.25–0.26 vs llama 0.42 (r13/r20/r25; r15 noted a sub-1:1 instruction→duration tracking at stall-bound SM% ≈ 31); llama eligible 0.64 vs minfer 0.28. minfer's long_scoreboard 6.22 (pre-r20) → 2.92 (post-split-phase), but the freed stalls re-saturated on lg_throttle 0.33→2.38 (r20). The remaining occupancy-bound residual ≈ 1.4× (minfer 211 μs/tile vs llama 113 μs/tile + 34 μs fixup).

The quadruple-confirmed paradigm verdict (r20, r21–r22, r24, r25). Four independent lever families all land on the same wall: (r20) the gap carrier is long-scoreboard/latency, not barriers or bytes; (r21/r22) the staging is instruction/issue-bound — coalesced staging and swizzle change sector/LDSM-conflict counts but move wall >0, and any "saving" that adds ALU loses; (r24) tile-order swizzle and persistent blocks are regressive — the scheduling structure family is closed; (r25) a −38% integer-ALU cut (kd #pragma unroll), which halves the instruction surplus, moves the wall by <0.5% — the kernel is issue/occupancy-bound, not instruction-count-bound. The binding resource is "more resident warps per SM", i.e. 2 blocks/SM, i.e. the smem budget; llama's 86% (r23) memory-throughput GEMM illustrates the other half — the raw-byte B-stream (4.5 bit/weight) at high SOL is what beats an f16 16-bit stream by ~1.5×/MAC.

10. Implications for a redesign

The campaign's lever map (each family measured-closed in r13–r25) leaves exactly three untried, and all three are on the occupancy and instruction-width axes, not the memory/scheduling axes:

  • Direction A — raw-nibble B in smem to halve the smem budget → 2 blocks/SM (the occupancy lever no measured family touched). Move the weight nibble expansion out of the staging byte-expansion (qb8) and keep the raw 4-bit plane (0x0F mask) so the B tile costs ~half the bytes. This is the one lever that directly attacks 1 block/SM → 2.00 warps/sched (the r25 "occupancy/latency-hiding" conclusion) without adding instructions. It requires the mma to consume unsigned nibbles with a rank-1 dmin fold (arithmetic identical to llama's; parity-safe, no add-reordering).
  • Direction B — fp16 half2 rescale (numerics gated). Shift the dequant-rescale from fp32 (FMUL+I2FP, minfer's +29.6k/tile) to the half2 path, cutting the largest remaining fp surplus. Gated on numerics because the parity gate is 1e-3 and minfer's fp32 rescale is what the campaign kept for exactness; the r15 "f32-accumulate mma does not exist" result also caps how far this can go.
  • Direction C — revisit the 2×-od-row warp shape after A. The warp is currently 16 od-rows; llama's 32-od-row shape (2 minitiles) halves the A-fragment load work per MAC. This was measured NEUTRAL in r17 at 1 block/SM (instructions −5.8%, wall ~0), so it is only meaningful once A (2 blocks/SM) has changed the latency regime.

Risks / gates to respect with any of these: (1) parity — the q4_K dmin fold and the 0x0F mask must reproduce get_scale_min_k4 semantics bit-identically under the 1e-3 tolerance; (2) token identity — any fp add-reordering (e.g. a stream-k k-split) breaks the greedy token-identity gate, which is the harder gate than the 1e-3 numeric one (r24 rung-3 decision); (3) smem-cap guard — the launcher must re-derive smem and refuse KD=8 if it exceeds the device cap (the silent-attr-failure regression of the r7-era wide tile); (4) the op_st conflict mass (6.47M, r22) remains the store-side residual and would need re-tiling if a future shape change touches the B staging stores.


Corrections & source-level nuances

The following are places where the campaign document's claims differ from, or are not fully reconciled with, the source at ca3d5a3e1. The empirical core (r20 stall table, r25 census) is unchanged; these are precision notes.

The numbered items below fold in the post-r25 source audit. Items marked (resolved) supersede the corresponding claim; the r13/r25 historical notes are kept but flagged where the audit corrects them.

  1. "65-int padded stride" ≠ source. CUDA_OPTIMIZATION.md r13 describes llama's weight plane as "raw-nibble (65-int padded stride, bank-rotating)". The source's sram_stride for the Q4_K (Q8_1) layout is 76 ints (304 B) — 2·32 + 2·32/8 + 4 (mmq.cuh:137), K%8==4 enforced (mmq.cuh:153-159). The "16 B pad" half matches (the +4 ints), but the "65-int" figure was not reproduced; treat the verified stride as 76. (resolved — the audit confirmed the stride but carried the earlier "70 ints" slip; the 2·32/QI8_1 term is 8, so the stride is 76.)
  2. "fp16 dequant-rescale" vs source-level fp32. CUDA_OPTIMIZATION.md r25 attributes llama's throughput advantage to "llama dequantizes to fp16 (HADD2/HFMA path)". At the CUDA level the q8_1×q8_1_mma rescale is fp32 — sum += dmA.x·dsB.x·C.x + dmA.y·dsB.y with dmA/dsB as float2 produced by __half22float2 (mmq-vec-dot.cuh:426,434-437) — the half2 is the storage of the scale, and the arithmetic is fp32. The census's high llama fp16 opcode count (36,400) is therefore empirical but its origin is not fully resolved from the source (the half2 scale construction in load_tiles_q4_K, mmq-load-tiles.cuh:772-777, is the most likely contributor); the campaign's "fp16 rescale" wording is an interpretation, not a source-verified arithmetic description. The direction of the binary advantage for llama (fewer fp32 FMUL/I2FP, more fp16 ALU) is real and is what the census records. (still interpreted, not source-verified.)
  3. MMQ_TILE_Y_K = 36 ints = 144 B (resolved). The doc and CUDA_OPTIMIZATION.md quoted 33 ints / 132 B (MMQ_TILE_NE_K + MMQ_TILE_NE_K/QI8_1 = 32+1). The Q8_1 DS4 scale word is 4 ints (half2 ds4[4]), not 1, so MMQ_TILE_Y_K = 32 + 32/QI8_1 = 32 + 4 = 36 ints = 144 B (mmq.cuh:119; QI8_1 = QK8_1/(4·QR8_1) = 32/4 = 8, ggml-common.h:124,258). The smem row stride therefore equals the global block_q8_1_mmq size (144 B, mmq.cuh:56-57); the earlier "global 144 B ≠ smem-row 132 B" distinction (old item 3) is superseded — the two were only different because the scale word was miscounted as 1 int.
  4. "2 blocks/SM → 4 warps/scheduler" as an explanation of the profiled result. The profiled mul_mat_q<12,128,0> runs at 1 block/SM (56.5 KB > 99/2, plus __launch_bounds__(…, 1)), exactly like minfer. The 2-blocks/SM benefit is the design's occupancy lever that materializes at smaller J tiles; it should not be read as the measured occupancy of the nt-512 q-proj capture. (Not verified at any J in this campaign.)
  5. "x_qs" naming. CUDA_OPTIMIZATION.md and this doc use x for the weight and y/tile_y for the q8_1 activation (per mul_mat ordering src0=weight, src1=activation). In block_q8_1_mmq the int8 data field is qs; in-kernel it is read at y+4 (mmq-vec-dot.cuh:383). Any reference to "x_qs" as the activation data in the campaign should be read as the weight's nibble plane (x_qs inside load_tiles_*), not the activation; the activation qs is y_qs.
  6. Compute-loop granularity (resolved). Per vec_dot call (one 32-k k00 chunk) there are 64 mma: j0 8 steps × k01 4 steps (step QI8_1=8) × n 2 (mmq-vec-dot.cuh:414-440). The campaign's "16 mma per 32-k chunk" is the per-k01 count (8 j0 × 2 n). A-frags = A[2][4] = 8, B-frags = 32, and each sum slot is accumulated 4× per vec_dot.
  7. "llama pays zero convert tax" ≠ source (resolved). The activations are quantized by a separate kernel — quantize_mmq_q8_1 (quantize.cu:458), launched by quantize_mmq_q8_1_cuda (quantize.cu:575) — inside ggml_cuda_mul_mat_q (mmq.cu:156-157 non-MoE, :236-237 MoE). It is not fused in-kernel. minfer's r23 default-path accounting should credit llama with a quantize-pass cost of similar order to minfer's convert tax (the q8_1 quantize is cheaper than f32→f16, but "llama pays zero" is wrong).
  8. nbs_y pad (resolved). nbytes_shared = nbs_ids + nbs_x + GGML_PAD(nbs_y, nthreads·sizeof(int)) (mmq.cuh:1386-1391). At J=128, GGML_PAD(18,432, 1024) = 18,432 (18,432 = 18×1024), i.e. no pad step for the profiled shape. The total at this shape is 512 + 128·76·4 + 18,432 = 57,856 B (≈56.5 KB), not the earlier 54,784 B (53.5 KB) — the audit kept the total unchanged while fixing the pad, but the nbs_x = I·sram_stride·4 term also rises to 38,912 because sram_stride is 76. The pad only applies for J not a multiple of 64.
  9. tile<16,8,int>::ne cite (resolved). ne = I·J/32 = 4 is on the NVIDIA Turing+ branch at mma.cuh:227; mma.cuh:108 is the AMD MFMA I·J/64 branch.
  10. Warp token coverage (resolved). Each warp covers 64 token columns, not the full 128: y += (ty%ntx)*(tile_C::J*MMQ_TILE_Y_K) (mmq-vec-dot.cuh:379) de-interleaves the od-group pair, whose two warps jointly cover 128. This is the warp shape minfer's r17 remap tested (measured NEUTRAL for minfer).
  11. Dispatch edge (resolved). The ne11 < MMQ_DP4A_MAX_BATCH_SIZE gate (mmq.cu:327) sits inside the NVIDIA-only branch (GGML_CUDA_CC_IS_NVIDIA, mmq.cu:326-328); the AMD/RDNA gates are separate (mmq.cu:330-385).
  12. Audit citation notes. The GAP A lane-map helpers get_i/get_j for tile<16,8,int> live in mma.cuh:245,262 (the audit's mmq-vec-dot.cuh:245,262 pointer is off by file — the helpers are in mma.cuh, which mmq-vec-dot.cuh includes). The GAP E host ntx is computed at mmq.cuh:1410-1411 (not :136-158), and the kernel recomputes nty at mmq.cuh:974.

11. minfer redesign design — direction A (raw-nibble B smem, 2 blocks/SM)

This is the Phase-1 design (document only, src/ untouched) for minfer's next CUDA MMQ GEMM kernel: the raw-nibble-smem variant targeting 2 blocks/SM. It is the one occupancy lever that no measured family in the r13–r25 campaign touched, per §9–§10. Everything below is derived from the r-numbered facts in docs/CUDA_OPTIMIZATION.md and the llama.cpp source at ca3d5a3e1; a Phase-2 implementer can build the kernel from this section alone without re-deriving the geometry.

Headline hypothesis (the bet). The r25 census verdict is exact: IMMA and FFMA are at hard parity (FFMA 114,688/tile both; IMMA 2×MAC both), the surplus is 100% support instructions (integer/address ALU +69,465/tile = 77% of the +89,786 surplus), and a −38% integer-ALU cut moved wall by <0.5% (r25). The kernel is therefore issue/occupancy-bound, not instruction-count-bound; the binding resource is more resident warps per SM. Direction A buys exactly that — 2 blocks/SM → ~4 active warps/scheduler (from 2.00) — by shrinking the weight B smem to the raw-packed 4-bit form and accepting a small in-loop B-expansion instruction cost. The bet: the occupancy gain hides latency better than the added instructions hurt. Per §10 this is the one untried lever that directly attacks 1 block/SM → 2.00 warps/sched without adding instructions to the A/balance path.

The one hard constraint. 2 blocks/SM on GB10 (48 SMs; shared/SM ≈ 99 KB usable, opt-in per-block ≈ 99 KB, §2) requires per-block dynamic smem ≤ ~49.5 KB. The current wide kernel at KD=8 is 98,304 B (mmq_raw.cu:222) → hard-pinned 1 block/SM (2.00 warps/sched, r20). The B weight tile is the dominant term and the only one that shrinks by switching representation.

11.1 Tile geometry

Block = 256 threads (8 warps), one block per (token-tile × od-tile). Smem byte formulas (KDR = chunks/super-block = 256/32 = 8 at KD=8; region map mmq_raw.cu:222):

regionper blockelementnotes
QA8 (activation)KDR·T·32chunk q8 planesr22 XOR swizzle, r20 split-phase staging
SDA (act d/ssum)KDR·T·8uint2 (d f16 | ssum i16)token pair (t,t+8) per LDS.64
QB EXP (weight)8·O·481 byte/nibble, 48B slotexpanded-qb8 (current, mmq_raw.cu:222)
QB RAW (weight)O·128 qs plane (or O·144 full super-block)2 nibbles/byteraw-packed GGUF qs
SDS (weight scale)KDR·O·8float2 (d·sc, −dmin·m)r15 rank-1 rescale

Candidate geometry table (bytes; KD=8, KDR=8; T=tokens, O=od-rows):

geomQA8SDAQB_expSDSexp totalQB_rawraw totalblocks/SM (exp / raw)
64×12816,3844,09649,1528,19277,82416,38445,0561 / 2
128×6432,7688,19224,5764,09669,6328,19253,2481 / 1
64×6416,3844,09624,5764,09649,1528,19232,7682 / 2–3

(exp = expanded-qb8 with 48B slot; raw = 2-nibbles/byte qs plane. blocks/SM = per-block ≤ 49.5 KB ⇒ 2. QB_raw holds the qs plane only; the header scales are staged into SDS separately.)

Justification (A-reuse vs B-reuse, r23 + r24). The measured re-read amplification at 128×128 is A ≈ 327 MB vs B ≈ 152 MB ⇒ A dominates 2.1:1 (CUDA_OPTIMIZATION.md:314-319). A-re-reads scale with the od-tile count (od/O); B-re-reads with the token-tile count (nt/T). The dominant stream is A (activation), so the geometry must keep O large and take any shrink on T:

  • 64×128 — od/O unchanged ⇒ A re-reads stay at base (327 MB); nt/T doubles ⇒ B re-reads → 304 MB. Total 631 MB. B-reuse is sacrificed (the smaller, reusable stream — r24 "B is the reusable operand in dispatch windows"), A-reuse (the 2:1 dominant stream) is preserved.
  • 64×64 (the only expanded-qb8 2-block shape) — od/O=2× ⇒ A re-reads → 654 MB, B → 304 MB, total 958 MB. Doubles the dominant stream: strictly worse than 64×128.
  • 128×64 — 53,248 B raw → 1 block (fails the goal).

Chosen primary geometry: 64 tokens × 128 od, KDR=8 (KD=8), raw-packed QB → 45,056 B → 2 blocks/SM. It is the only shape that both (a) reaches 2 blocks/SM and (b) keeps the od tile at 128 so the dominant A-re-read stream is untouched.

11.2 Warp shape & register budget

Keep the per-warp structure (each warp owns a private 16-od-row slice and reads the full T-token tile, mmq_raw.cu:222) now over T=64: 8 warps × 16 od-rows = 128 od. mma.m16n8k32 maps m = token, n = od (epilogue C[i·od + j], mmq_raw.cu:222). Using the audit C-lane map get_i(l) = (l/2)*8 + tid/4, get_j(l) = (tid%4)*2 + (l%2) (mma.cuh:245,262) and tile<16,8,int>::ne = I·J/32 = 4 (mma.cuh:226-227):

quantitycurrent 128×128new 64×128
token-groups per warp (m-steps, T/16)84
od-groups per warp (n-steps, 16/8)22
mma per 32-k chunk168
sum[] size (groups × 4 C regs)sum[64]sum[32]
A-frag ldmatrix per chunk84
B-frag per chunk1 ldmatrix.x41 raw-expand
chains per thread168

Register estimate. Dropping sum[64]→sum[32] (−32), halving the A-frag array (a[8][4]→a[4][4], −16) and the temp C array (clow[8][2][4]→clow[4][2][4], −32) is partially offset by the raw-B in-loop expansion temps. Estimated ~110–130 regs (vs current 141–149, r22) — well under the 255-spill cliff (the r22 Lever-2 255-reg + 112B-spill is the failure mode to avoid). The halved A-frag count also lowers the per-chunk shared-load count, partly offsetting the B-expansion ALU (§11.3).

11.3 Staging plan & the B-representation decision

The single unmeasured decision in Direction A is how the B weight enters smem:

(Option 1) expanded-qb8 + ldmatrix (current). Stage the raw weight, nibble-isolate at staging (0x0F0F0F0F, 1 byte/nibble), 48B slot-major, load B-frags with ONE ldmatrix.x4 (mmq_raw.cu:222, 4685-4698). B smem = 8·O·48 = 49,152 B @ O=128 ⇒ 64×128 totals 77,824 B ⇒ 1 block/SM. It does not reach 2 blocks at any geometry keeping O=128. Rejected.

(Option 2) raw-nibbles + in-loop expansion (chosen). Stage the raw GGUF qs plane, 2 nibbles/byte (mask at USE, not at stage — the inverse payload of llama's x_qs[...]=(qs0>>0)&0x0F0F0F0F, mmq-load-tiles.cuh:736-737), O×128 B (@ O=128: 16,384 B), as a bulk copy (no staging ALU — mirroring r18's bulk-copy staging into a raw region). B-frags are then assembled in the compute loop: LDS the packed bytes + PRMT/SHF/LOP3 to spread each 2-nibble byte into two int8. B smem = 16,384 B ⇒ 64×128 totals 45,056 B ⇒ 2 blocks/SM. The header scales are staged separately into SDS (float2, KDR·O·8), keeping the per-chunk rescale byte-identical to the current kernel.

The tradeoff, quantified from the r25 census. Option 2 raises the in-loop instruction count because the memory-side fact is already inverted: the census had ours LDSM 8,064/tile vs theirs 1,792 and ours shared_ld 8,960 vs theirs 19,346 — minfer already uses the leaner (ldmatrix) B path. Option 2 moves B back to plain-LDS + unpack: per (warp, chunk) the B-fragment costs ~1 LDS (was 1 ldmatrix, ~saved 0) + ~30–60 ALU to unpack 8 rows × 32 nibbles → int8. Over the per-tile stream that is ~+5–10% warp-inst on the integer-ALU class (already the dominant surplus at 113,552/tile). The counterweight: the A-side drops 8 → 4 ldmatrix per chunk (−4 LDSM/chunk), so the net instruction move is ~+3–6% total — an order of magnitude smaller than the r25-introspection baseline. Since r25 showed a −38% integer cut moves wall <0.5%, a +few-% instruction change is expected to be ~wall-inert provided the occupancy lever fires — which is exactly the bet under test.

Kept from minfer / adopted from llama. Keep split-phase A staging (r20) and the qa8 XOR swizzle (r22); keep the rank-1 two-term rescale (r15/r16) and the f16-scale-with-fp32-rescale (per §Corrections item 2 the fp16-vs-fp32 attribution is unresolved at the source level, and the fp32 rescale is what the campaign kept for exactness). Do NOT copy llama's sram layout (sram_stride=76 ints, mmq.cuh:137) — it is 1-byte-per-nibble and is not the halving. What is adopted from llama is only the concept of a raw nibble B plane, but packed 2/byte (the block_q4_K.qs[128] plane, ggml-common.h), which is what actually halves the smem. The dispatch guard (id/32)%8 == 0 (src/cuda/methods/init.rs:111) applies unchanged (same pad40 q8 quantize, src/cuda/methods/init.rs:119).

11.4 Numerics

Raw-nibble semantics are exactly the r13-era two-term rescale (the parity-safe form, r15/r16-verified):

  • The mma consumes the unsigned 0..15 nibble as the int8 B operand (upper nibble zero ⇒ positive int8), so the int accumulator holds C_int = Σ_k nib(k)·act(k), nib ∈ [0,15].
  • At accumulate (per chunk, fp32): sum += da·dsv·C_int + dma·dmv with da = act d, dsv = d·sc, dma = da·ssum, dmv = −dmin·m (mmq_raw.cu:222). This is the exact d·s·nib − dmin·m dequant form. There is no (nib − m) centering anywhere — that fold is proven wrong for q4_K because the dmin offset is per-sub-block-scaled (−dmin·m, not a fixed subtraction), which is the "82.896 diff mode" lesson. mma is .s32.s8.s8.s32 (mmq_raw.cu:222); the f32-accumulate spelling does not exist (r15).
  • fp32 write-back epilogue (adopt llama's): sum[] is already fp32 at the rescale, so the epilogue is a plain per-value C[i·od + j] = sum[...] global store (mmq_raw.cu:222) — no fp16 anywhere in the mma→store path (llama's Q4_K write_back is likewise a plain fp32 store, mmq.cuh:519).

11.5 Risk table

#riskconsequenceguard / signal
1Nibble-layout mistake at in-loop unpack (wrong nibble=k, sign-extend the high 4 bits, double dmin)the r13-era 82.896 max-diff parity modecuda_prefill_mmq parity arm (§11.6) must run BEFORE the first perf run; a garbage-magnitude diff (like r14's uint4-tiling corrupting qb8) = layout bug; ~1e-6 diff = legit fp rounding.
2Token identity — any fp add reorderinggreedy token identity divergesthe design does NOT reorder (same per-chunk mma + same two-term fp fold order); still gate on greedy-32 (r24 rung-3 convention).
3Smem-cap overrun (the r7-era silent-attr-failure regression, phantom 2124)launcher quietly fallback/corruptslauncher re-derives smem and return 0 (→ narrow fallback, src/cuda/methods/init.rs:140-152) if over cap; cudaFuncSetAttribute result checked (mmq_raw.cu:222, 4795-4798). KD=8 @ 45,056 B safe; KD=16 or O=256 would not be.
4Register spill at KD=8 (in-loop B-expand temps + sum[32])ptxas → 255 regs + local spill (the r22 Lever-2 failure)-Xptxas -v gate: expect ~110–130 regs, 0 spill; REG > 160 → risk.
5Occupancy gained but wall flat (ncu ~4 warps/sched, duration unchanged)falsifies the occupancy hypothesisthis is the designed kill criterion (§11.8), not a bug — it closes the line.
6A-side re-staging for the smaller Tmore A per od-tileA re-reads unchanged (od/O held at 128); only B re-reads grow (the designed sacrifice).

11.6 Phase-2 gate plan (in order)

  1. Build + correctness. cargo build --features cuda; new env gate MINFER_MMQ_RAW_NB=1 selects the raw-nibble variant as a parallel kernel — it never replaces mmq_raw_wide_nt_kernel in this phase; the existing wide kernel remains the default raw path.
  2. Parity. MINFER_MMQ=1 MINFER_MMQ_RAW=1 MINFER_MMQ_RAW_WIDE=1 MINFER_MMQ_RAW_KD=4 (KD defaults to 8; the KD=8 default path is likewise gated). Gate: ≤1e-3 max-diff parity vs the host reference.
  3. Greedy token identity. greedy-32 output identical to the default path (r24 rung-3).
  4. Perf (interleaved 3× medians, relative bar). ≥ +1.5% over the re-measured baseline (r24 convention; current baseline KD=8 ~1385, KD=4 ~1366 tok/s). The absolute ≥1350 bar is superseded by the relative bar because the baseline already clears 1350.
  5. ncu occupancy + stall re-check. sm__warps_active/sched should read ~4 at 2 blocks/SM; long_scoreboard should fall from ~2.92 (post-r20) toward llama's ~1.15; duration/GMAC (r13 target ≤107.7 μs/GMAC; llama parity ≈41 μs/GMAC — but the gate is the relative wall bar).
  6. Suite. 166/0/3.

11.7 Kill criteria (variant abandoned)

  • Parity unresolvable after 3 attempts (a nibble/dmin layout bug surviving 3 fixes), or
  • Occupancy achieved (ncu reads ~4 warps/sched) but wall < baseline — which falsifies the occupancy hypothesis and closes Direction A (it would show that, like r23's FA TKV=32 2-blocks/SM, the 2× k-loop fixed costs / added in-loop ALU consume the latency-hiding gain), or
  • Register spill at KD=8 that cannot be recovered without dropping to a smaller O.

11.8 The per-round records — consolidated into the step documents

The round-by-round records that followed here (r28 → r59, formerly §11.8–§11.37) were maintained round for round in BOTH this document and the per-step documents, doubling the surface that had to be kept honest; on 2026-09-10 they were consolidated. The authoritative copies are:

  • CUDA_OPTIMIZATION.md §0 — one row per round (commit, measured delta, status, one-line lesson), rows r28–r59;
  • cuda_optimization_steps/31–64 — one standalone chapter per round (31 = r28 NB kernel … 64 = r60 promotion), indexed by the master table's Part III/IV tables.

Anchors for older cross-references: former §11.8 = r28 Phase-2 outcome, §11.26 = r47, §11.34 = r55, §11.36 = r58, §11.37 = r59. The design body above (§11.1–§11.7) keeps the Direction-A design: tile geometry, warp shape, staging plan, numerics, risk table, gate plan, and kill criteria.

12. The other side of the dispatch — MMVQ and the small-M chain

Recorded after step doc 81 (2026-09-10). Doc 81 measured minfer's batched verify step at 0.52× per-token amortization — nt=2–15 falls into a legacy kernel whose grid is grid(od/4, nt), i.e. one full weight re-stream per token at ~125 GB/s — and closed the D5 campaign. The external reference (doc 81 §4.1) showed llama.cpp landing at 1.00× on the same pair. The mechanism behind both numbers is the dispatch chain upstream of MMQ, which §1 does not cover: on small batch sizes llama.cpp never reaches the should_use_mmq question, because MMVQ answers first — and the MMVQ kernel's structure makes weight traffic independent of M.

12.1 The full quantized mul_mat dispatch chain — no hole

ggml_cuda_mul_mat (ggml-cuda.cu:1836–1872) dispatches quantized weights in this order:

  1. MMVF — should_use_mmvf (ggml-cuda.cu:1840): f16-activation vector kernel over thin matrices / small batches (GGML types without quant vec-dot support);
  2. MMF — should_use_mmf (ggml-cuda.cu:1860): f32/f16 GEMV for small batches on f16/f32 weights;
  3. MMVQ — should_use_mmvq (ggml-cuda.cu:1864, impl mmvq.cu:318): the quantized vector kernel, gate MMVQ_MAX_BATCH_SIZE = 8 (mmvq.cuh:3), with per-arch tuned threshold tables (mmvq.cu:320–361):
    • Ada / RTX 4090: Q2_K ≤ 4, Q3_K ≤ 6, everything else ≤ 8;
    • Blackwell / RTX 5090: Q2_K–Q4_K ≤ 5, Q5_K ≤ 6, Q6_K ≤ 7;
    • DGX Spark GB10 (mmvq.cu:348–355, "tuned on DGX Spark GB10"): Q2_K ≤ 6, everything else ≤ 8 — the reference tunes the 2–8 window per chip; it never leaves it unhandled;
  4. MMQ — should_use_mmq (ggml-cuda.cu:1868, impl mmq.cu:259): Turing+ mma → true for any ne11 (§1);
  5. cuBLAS — only for types MMQ does not support.

So every M has a dedicated kernel: 1–8 → MMVQ, 9–∞ → MMQ. minfer's hole (nt=2–15 → legacy *_f32_matmul, from the nt >= 16 tiled-GEMM gate at src/cuda/methods/weights.rs:227 and the nt == 1 MMVQ gates, e.g. src/cuda/methods/weights.rs:305) is exactly the seam between two tuned regimes that this chain does not have.

12.2 The MMVQ kernel — M lives in registers, not in the grid

mul_mat_vec_q (mmvq.cu:585). The launch grid is (od_rows / rows_per_cuda_block, channel, sample) — ne11 is not a grid dimension (rows_per_cuda_block comes from the per-arch MMVQ parameter table, calc_rows_per_block, mmvq.cu:563, 602). Each block owns a few weight rows, streams each weight block ONCE per K-iteration, and dots it against ALL M tokens:

float tmp[ncols_dst][rows_per_cuda_block] = {{0.0f}};   // mmvq.cu:693
for (int kbx = ...; kbx < blocks_per_row_x; kbx += blocks_per_iter) {  // K loop
    ...                                                  // weight block loads
    for (int j = 0; j < ncols_dst; ++j) {                // mmvq.cu:724 — ALL M tokens
        for (int i = 0; i < rows_per_cuda_block; ++i)
            tmp[j][i] += vec_dot_q_cuda(vx, &y[j*stride_col_y + kby], ...);

The kernel is template-instantiated per c_ncols_dst = 1…8; the M token activations are register/L1-resident, so DRAM weight traffic equals the nt=1 case by construction. GB10 even has a dedicated L2-prefetch branch inside this K loop (__CUDA_ARCH__ == GGML_CUDA_CC_DGX_SPARK, mmvq.cu:704–723, mmvq_prefetch_l2 at distance 2 K-iterations).

12.3 MMQ at small M — the M-tile floors at 8

For ne11 above the MMVQ gate, MMQ's M-tile J is selected by ggml_cuda_mmq_get_J_max (mmq.cuh:366–374): the largest per-arch configured tile ≤ ne11, stepping down in multiples of 8 (J ≤ 512, config tables mmq-config-*.cuh). M=9–15 runs as a partial J=8 tile sweep — weights still stream once; an under-filled tile costs compute, not bandwidth. (Tile I/J semantics: §2.)

12.4 The invariant, and the contrast with minfer

Design invariant: M never appears in the launch grid as a dimension that multiplies weight traffic — it is either a register loop (MMVQ, ≤8) or a tiled dimension with partial-tile masking (MMQ, ≥9). minfer's legacy kernel violates it — launch_q4_k_f32_matmul (src/cuda/kernels/matmul_f32act.cu:672): dim3 grid((od + NR0*NSG - 1)/(NR0*NSG), nt, 1), grid.y = nt, one full weight pass per token, on a 64-thread f32-dequantizing kernel at ~125 GB/s (vs the MMVQ path's ~238 GB/s). That is the measured 34.9 ms/token linear regime of doc 81 §4 (4.36 GiB / 125 GB/s ≈ 34.9 ms), and the 14× gap at nt=8 (277.5 ms vs ≈ 18–20 ms weights-once):

ntllama.cpp pathweight passesminfer pathweight passes
1MMVQ1MMVQ (nt == 1 gate)1
2–8MMVQ (tokens in registers)1legacy f32 kernelnt
9–15MMQ (partial J=8 tiles)1legacy f32 kernelnt
≥ 16MMQ1tiled MMQ GEMM1

Fix shape for minfer (a future-feature fix, not a D5 revival — doc 81 §5): lift the MMVQ nt == 1 gate to nt <= 8 (the register-loop structure extends directly; the activations of 8 tokens are 8·id/32 q8 blocks ≈ negligible), or let the tiled GEMM accept nt < 16 with partial-tile masking. Either restores the weights-once invariant; §11's Direction-A raw-nibble kernel would also close it if it ever lands.

llama.cpp Small-Draft-Model Speculative Decoding — Source Analysis

Source basis: llama.cpp master @ commit 050dde50c (checked out at ~/git/reading/llama.cpp). Scope: COMMON_SPECULATIVE_TYPE_DRAFT_SIMPLE ("draft-simple") — the classic standalone small draft model speculative decoding — plus the orchestration framework it runs in. The framework today also hosts EAGLE-3, MTP, DFlash/DSpark and four n-gram self-speculators; this doc covers what they share (the lifecycle, verification and rollback machinery) and what is specific to the small-draft-model path.

Upstream user-facing documentation: docs/speculative.md in the llama.cpp repo. This document is the implementation-level companion: every claim below is anchored to a file:line.


1. TL;DR

  • The draft model is any small GGUF with a compatible vocab; it runs in its own llama_context with the same n_ctx as the target and is decoded once per verification round.
  • Drafting is greedy argmax over the draft logits, gated by an optional confidence floor (p_min on the top-k-renormalized probability) and capped by n_max (default 3).
  • Verification never injects a draft token: the target model evaluates [id_last, draft...] in one batch and the target's own sampler chain decides token by token; a draft token is kept only when the target samples the same token. Output distribution is therefore exactly the target sampler's, and every round produces ≥ 1 token for free (the replacement/final sample).
  • Rollback of rejected drafts is done by plain KV seq_rm when the memory supports partial removal (unified KV cache); otherwise the code snapshots state (common_prompt_checkpoint) and replays the accepted prefix as a "guaranteed draft" next round.
  • The draft context stays synchronized with the accepted history by replaying every target batch through it without logits (process()), so the per-round draft cost is one seed decode
    • up to n_max sequential draft decodes + one replay batch.

2. File map

ConcernLocation
Speculator framework + all implementationscommon/speculative.cpp (2980 ln) / common/speculative.h
Params structs (common_params_speculative*)common/common.h:171–395
Acceptance sampling (sample_and_accept_n)common/sampling.cpp:678–715
Canonical single-slot loopexamples/speculative-simple/speculative-simple.cpp (377 ln)
Multi-slot server integrationtools/server/server-context.cpp (drafting ~2965, accept ~3881)
Context capability probecommon/common.cpp:1583 (common_context_can_seq_rm)
CLI/env optionscommon/arg.cpp (~4150–4330), docs/speculative.md
Checkpoint (state snapshot)common/common.h:1165 (common_prompt_checkpoint)

3. The pluggable speculator framework

3.1 Types and priority chain

common_speculative_type (common/common.h:171): none, draft-simple, draft-eagle3, draft-mtp, draft-dflash, draft-dspark, ngram-simple, ngram-map-k, ngram-map-k4v, ngram-mod, ngram-cache (11 total, static_asserted at speculative.cpp:2615).

--spec-type accepts a comma-separated list. common_speculative_init (speculative.cpp:2602) instantiates one impl per enabled type in a fixed priority order: all n-gram impls first (they are free lookups — if they produce a draft, the draft model never runs), then the draft-model impls (draft-simple first among them, speculative.cpp:2619–2629).

Chaining works through the per-sequence drafting flag (common_speculative_draft_params.drafting, speculative.h:58): common_speculative_draft (speculative.cpp:2790) walks the impl list; the first impl that fills *dp.result clears the flag, so later impls skip that sequence; empty results fall through to the next impl.

Note: draft-simple is not auto-selected. Passing -md with a plain small model keeps types = {none} → common_speculative_init returns nullptr → speculation silently off. You must pass --spec-type draft-simple (the GGUF-metadata auto-detection at speculative.cpp:2284 only recognizes MTP / DFlash / DSpark drafts; see §4.3).

3.2 Impl interface

Base class common_speculative_impl (speculative.cpp:138):

struct common_speculative_impl {
    const common_speculative_type type;
    uint32_t n_seq;
    int32_t  n_max;                    // effective max draft length of this impl
    // built-in accounting: n_call_begin/draft/accept, n_gen_drafts, n_acc_drafts,
    // n_gen_tokens, n_acc_tokens, n_acc_tokens_per_pos, t_begin/draft/accept_us
    virtual void begin(seq_id, const llama_tokens & prompt) = 0; // new-generation refresh
    virtual bool process(const llama_batch & batch) = 0;         // observe a target batch
    virtual void draft(common_speculative_draft_params_vec & dp) = 0;
    virtual void accept(seq_id, uint16_t n_accepted, bool is_other) = 0;
    virtual bool get_state(...) const; virtual void set_state(...);   // optional serialize
};

Public lifecycle (all null-safe no-ops when speculation is disabled):

StepAPIdraft-simple behavior
new generationbegin(seq, prompt)no-op
every target decodeprocess(batch)replay batch on draft ctx, logits = nullptr
per rounddraft()fill *result for flagged seqs (see §5.3)
after verificationaccept(seq, n_accepted)no-op (rollback is the caller's job)

The per-seq draft request is a plain struct the caller fills (common_speculative_draft_params, speculative.h:53):

bool          drafting;   // ask this impl for a draft this round
int32_t       n_max;      // per-round cap (context/predict budget), -1 = impl default
llama_pos     n_past;     // position of id_last (seed)
llama_token   id_last;    // seed token
const llama_tokens * prompt;  // current sequence (input for n-gram impls)
llama_tokens * result;         // output; caller owns the vector

The outer common_speculative object (speculative.cpp:2177) is just dparams[n_seq] + impls[] + impl_last[seq] (impl_last routes the later accept() call to whichever impl actually produced the last draft, speculative.cpp:2875).

4. Draft model & context setup

4.1 Model/context creation

common_speculative_init_from_params → common_speculative_init_result (speculative.cpp:2515–2560): loads the draft GGUF from params.speculative.draft.mparams.path and creates a dedicated context. Context params (speculative.cpp:2532–2538):

  • cparams.n_ctx = llama_n_ctx(ctx_tgt) — the draft must hold the same sequence (prompt + all accepted tokens), not just n_max extra tokens;
  • cparams.n_rs_seq = 0 — the draft context never uses RS-bounded rollback;
  • cparams.ctx_other = ctx_tgt — cross-context linkage.

Draft-side knobs are derived by common_base_params_to_speculative (speculative.cpp:2460): model path / -ngld / -devd / tensor-buft overrides / -td threads from the common_params_speculative_draft block, KV dtypes cache_type_k/v default F16 (-ctkd/-ctvd), and n_outputs_max = n_parallel (the draft never needs logits rows from process(); draft() requests logits per row it decodes).

4.2 Vocab compatibility (hard gate)

common_speculative_are_compatible (speculative.cpp:67–130), checked in the impl constructor and fatal on mismatch:

  1. same vocab type (SPM/BPE/WPM/UGM);
  2. same BOS add-flag and id; same EOS add-flag and id;
  3. |n_vocab_tgt − n_vocab_dft| ≤ 128 (SPEC_VOCAB_MAX_SIZE_DIFFERENCE, line 30);
  4. token text byte-equality for every id from 5 (SPEC_VOCAB_CHECK_START_TOKEN_ID, line 31) through min(n_vocab) — catches renames even when sizes match.

The draft-simple constructor also asserts n_seq == llama_n_seq_max(ctx_dft) (speculative.cpp:247) — the draft context must have one KV slot per target sequence.

4.3 Type auto-detection from GGUF metadata

common_speculative_types_from_gguf (speculative.cpp:2284–2319) reads only metadata:

  • general.architecture == "dflash" → draft-dflash, or draft-dspark if a markov_w1.weight tensor (the Markov head) exists;
  • otherwise, presence of blk.{n_layer-1}.nextn.eh_proj.weight → draft-mtp;
  • anything else → {} (types stay none — see §3.1 note).

5. draft-simple internals (speculative.cpp:179–389)

5.1 Constructor: the draft sampler

One common_sampler per sequence (speculative.cpp:226–236):

common_params_sampling params;
params.no_perf = false;
params.top_k   = 10;
params.samplers = { COMMON_SAMPLER_TYPE_TOP_K };   // explicit chain: only top-k
smpl.reset(common_sampler_init(llama_get_model(ctx_dft), params));

Although the explicit chain is only TOP_K(10), common_sampler_init always appends a dist sampler at the end (common/sampling.cpp, "default: sample from distribution"), so common_sampler_sample computes softmax probabilities over the top-10 candidates. The draft loop then ignores the sampled token and takes the argmax instead:

common_sampler_sample(smpl, ctx_dft, i_batch, true);        // populates candidates
const auto * cur_p = common_sampler_get_candidates(smpl, true);  // sorted by p desc
const llama_token id = cur_p->data[0].id;                   // ← greedy argmax
if (cur_p->data[0].p < params.p_min) { /* stop this seq */ }

(speculative.cpp:322–342) Two consequences worth pinning down:

  • Drafting is deterministic given the logits (argmax), independent of RNG seed. The dist sampler runs (consuming RNG, normalizing p) but its pick is discarded.
  • data[0].p is the argmax's probability renormalized within the top-10, so p_min (default 0.0, --draft-p-min) means "the argmax must hold at least p_min of the top-10 mass"; with the default it never truncates a draft early.

The own llama_batch is sized to the draft context's n_batch (speculative.cpp:207).

5.2 process(): free KV synchronization

bool process(const llama_batch & batch) override {
    llama_batch batch_dft = batch;
    batch_dft.logits = nullptr;          // no output rows needed
    return llama_decode(ctx_dft, batch_dft) == 0;
}

(speculative.cpp:262–277) Every target decode batch — prompt prefill (speculative-simple.cpp:135) and each verification batch (speculative-simple.cpp:237) — is replayed verbatim (same tokens, positions, seq ids) on the draft context. This keeps the draft KV cache identical to the target's accepted history without any bookkeeping; the draft never re-prefills.

5.3 draft(): the greedy drafting loop

Request shape (caller side, speculative-simple.cpp:188–196): {drafting=true, n_max, n_past, id_last, &prompt_tgt, &draft}. Implementation (speculative.cpp:279–384), with all drafting sequences batched together:

for each seq with dp.drafting:
    common_sampler_reset(smpl[seq])
    batch += { id_last @ dp.n_past, logits = true }        // seed row
decode(batch)                                              // 1 batched decode

i = 0
while n_drafting > 0:
    clear(batch); i_batch = 0
    for each still-drafting seq:
        common_sampler_sample(smpl[seq], ctx_dft, i_batch++, true)
        cur_p = candidates (sorted)
        id = cur_p->data[0].id                             // greedy
        if cur_p->data[0].p < params.p_min: drop seq       // confidence floor
        common_sampler_accept(smpl[seq], id, true)
        result.push_back(id)
        if result.size() >= min(params.n_max, dp.n_max): drop seq
        batch += { id @ dp.n_past + i + 1, logits = true } // next input row
    if batch empty: break
    decode(batch); ++i                                     // evaluate new tokens

for each seq: if result.size() < params.n_min: result.clear()   // too short → no spec

Early-stop conditions: p_min floor, global --draft-draft-n-max/--draft-max (params.draft.n_max, default 3), the per-round caller cap dp.n_max (context / n_predict budget, §6.2), and decode failure. n_min (--draft-min, default 0) discards drafts shorter than it entirely — the round then runs as plain decoding. accept() and begin() are no-ops for this impl (speculative.cpp:386–388, 258–260).

5.4 vestigial knob

p_split ("speculative decoding split probability", default 0.1) is still parsed (arg.cpp:4209) and documented, but no code on this master reads it — the old "draft only with probability p_split" heuristic is gone; drafting happens every round. backend_sampling (default on, --spec-draft-backend-sampling) lets the draft's sampling run on the backend (common_sampler_sample short-circuits on llama_get_sampled_token_ith, sampling.cpp:610).

6. Verification: the target-side loop

6.1 One round of examples/speculative-simple

Invariants entering each round (speculative-simple.cpp:141–152):

  • prompt_tgt holds the committed tokens at positions [0, n_past); prompt_tgt.size() == n_past;
  • id_last is the token at position n_past, not yet evaluated by either model (it was only sampled from the previous round's last logits row, or is the prompt's last token);
  • both KV caches hold exactly [0, n_past).
while true:
  if draft.empty():                                    // no replay pending
      ckpt.update_pos(n_tokens, tgt_pos_min, tgt_pos_max)
      ckpt.update_dft(ctx_dft)                         // only if ctx can't seq_rm partial
      n_draft_max = min(n_ctx - n_past - 2,            // leave room for id_last + shift
                        n_predict - n_predict - 1)     // generation budget
      draft_params = {drafting, n_draft_max, n_past, id_last, &prompt_tgt, &draft}
      common_speculative_draft(spec)
      if !draft.empty() && ctx can't seq_rm partial: ckpt.update_tgt(ctx_tgt)
      // roll the draft ctx back to [0, n_past): undo the draft() decodes
      ckpt.load_dft(ctx_dft) or seq_rm(ctx_dft, ckpt.pos_max + 1, -1)

  batch_tgt = { id_last @ n_past++ } + draft[i] @ n_past + i   // logits on every row
  llama_decode(ctx_tgt, batch_tgt)
  common_speculative_process(spec, batch_tgt)          // replay on draft ctx (no logits)

  ids = common_sampler_sample_and_accept_n(smpl, ctx_tgt, draft)   // §6.3
  if partial acceptance && ctx can't seq_rm partial:
      draft = move(ids); ckpt.load_tgt/dft; restore sampler clone; continue   // replay
  common_speculative_accept(spec, seq, ids.size() - 1)

  n_past += ids.size() - 1                             // committed accepted tokens
  for id in ids: prompt_tgt.push_back(id_last); id_last = id; print; EOG check
  draft.clear()
  seq_rm(ctx_tgt, n_past, -1); seq_rm(ctx_dft, n_past, -1)     // drop rejected tail

(speculative-simple.cpp:162–342)

Position/n_past bookkeeping: the seed sits at position P (n_past++ post-increment), drafts at P+1+i. After full acceptance of all k drafts, n_past = P+k+1 and seq_rm(n_past, -1) removes nothing; on partial acceptance of a < k drafts, it removes the rejected KV rows at P+a+1 … P+k plus the replacement token's row at P+a+1, which is re-evaluated as the next round's seed. Nothing is wasted: the replacement token becomes the next id_last.

6.2 Draft budget

n_draft_max = n_ctx − n_past − 2 (speculative-simple.cpp:181, server equivalent server_slot::get_n_draft_max, server-context.cpp:483): the −2 reserves the seed row and one slot for context shift. It is additionally clamped by the remaining n_predict budget and by dp.n_max inside draft(). The target must be able to emit logits for 1 + n_draft rows per sequence — common_speculative_get_output_limits(n_batch, n_parallel, n_draft) (speculative.cpp:2589) computes {total = min(n_batch, n_parallel·(1+n_draft)), per_seq = min(n_batch, 1+n_draft)} and feeds n_outputs_max / n_outputs_max_per_seq (include/llama.h:364–366).

6.3 Acceptance algorithm

common_sampler_sample_and_accept_n (common/sampling.cpp:678–706):

for (i = 0; i < draft.size(); i++) {
    id = common_sampler_sample(gsmpl, ctx, idxs[i], grammar_first);  // target chain
    common_sampler_accept(gsmpl, id, true);                          // target state advances
    result.push_back(id);
    if (draft[i] != id) break;                                       // mismatch → replace
}
if (i == draft.size()) { id = sample(row idxs[i]); accept; push; }   // bonus token

Semantics:

  • Row i of the target verification batch holds the logits after consuming the token at position P+i; it predicts the token at P+i+1 — the same slot draft[i] proposes. The comparison is therefore apples-to-apples.
  • The pushed token is always the target's own sample. Draft tokens are only ever checked, never injected. The output distribution equals the target sampler chain exactly — for any temperature/grammar/penalty configuration, no distribution correction is needed.
  • The first mismatch replaces the rejected draft token with the target's sample and stops, so result.size() − 1 = accepted draft count and result.size() ≥ 1 always: every round emits at least the seed token even when the draft is fully rejected (GGML_ASSERT at speculative-simple.cpp:262).
  • With a greedy target chain this is exactly classical speculative rejection sampling. With a sampled chain it is more conservative than the theoretical p_d/p_t accept test (a draft token is kept only when the target sampler draws the identical token), trading a little acceptance rate for exactness and simplicity — a deliberate llama.cpp design choice; the doc comment at speculative-simple.cpp:251–257 spells out the "the sampler would have to sample that same token" contract.

Grammar interaction: grammar_first=true is passed for the draft side only; on the target side the user's grammar applies through the normal chain (grammar-based resampling inside common_sampler_sample, sampling.cpp:646–675).

7. KV rollback: seq_rm vs checkpoints vs replay

Whether partial acceptance needs special handling depends on the memory module, probed at runtime by common_context_can_seq_rm (common/common.cpp:1583–1627): decode 2 tokens, then attempt seq_rm(mem, 0, 1, -1):

Probe resultMemoryRollback strategy
PART (removal ok)unified KV cacheplain seq_rm of the rejected tail — no snapshots
RS (n_rs_seq > 0)recurrent w/ snapshotsseq_rm bounded by n_rs_seq (common.h:992); checkpoint only when a rollback exceeds it
FULL (removal fails)e.g. pure recurrentcheckpoint save/restore + replay

common_prompt_checkpoint (common/common.h:1165) holds {n_tokens, pos_min, pos_max, data_tgt, data_dft, data_spec} and save/loads both contexts' per-seq state with LLAMA_STATE_SEQ_FLAGS_PARTIAL_ONLY (include/llama.h:912). data_spec additionally stashes impl-internal state (used by EAGLE-3's deferred boundary, not by draft-simple).

The replay trick (speculative-simple.cpp:267–290, server: slot.spec_is_replay, server-context.cpp:3922): on partial acceptance with a non-removable memory, the context is restored to the pre-round checkpoint and draft ← ids — the already-target-approved tokens are re-submitted as the next draft. The next verification round accepts them wholesale (the target sampler is restored from a clone, smpl_save, so the draws repeat), effectively batch-committing the accepted prefix through the normal machinery.

The draft context is always rolled back to [0, n_past) before the next draft() call (speculative-simple.cpp:206–213), because draft() re-seeds by decoding id_last and the process() replay of the verification batch would otherwise double-insert those rows. The server flags the resulting re-evaluation as a known optimization target (TAG_SPEC_AVOID_DRAFT_REEVAL, server-context.cpp:3041).

8. Server integration (multi-slot)

tools/server/server-context.cpp runs one common_speculative for all slots with n_seq = n_parallel; per-slot state lives on the slot (spec_draft, spec_i_batch, spec_ckpt, spec_prompt, spec_is_replay, spec_synth_rng).

  • Drafting (update_slots, server-context.cpp:2965–3032): every SLOT_STATE_GENERATING slot that can batch together gets its draft params filled (id_last = slot.sampled, the last accepted token); all drafts are produced by a single common_speculative_draft call inside queue_tasks.yield_to_queue(...) — the draft() loop's per-seq batching (§5.3) turns this into one decode per step across slots. n_draft_max per slot from get_n_draft_max() (server-context.cpp:483).
  • Checkpoints are taken per slot after drafting when the target/draft context requires it (server-context.cpp:3034–3080), including the RS-bounded rule (draft.size() > llama_n_rs_seq(ctx) → checkpoint, lines 3055–3060).
  • Acceptance (server-context.cpp:3881–3951): same sample_and_accept_n with the slot's sampler (spec_i_batch records which rows of the chunked target batch belong to this slot's verification); rollback identical to §7; spec_is_replay marks replay rounds and adjusts statistics so replayed tokens are not double counted (server-context.cpp:3956–3958).
  • Accepted tokens are committed to the slot prompt, streamed (process_token), and both memories are truncated from pos_next() (server-context.cpp:3976–3982).
  • Speculative state travels with server-side state checkpoints: common_speculative_get_state/set_state stash data_spec with the slot's checkpoint (server-context.cpp:2349–2350, 3355–3356).

9. Cost model — when does it win?

Per round with draft length d and acceptance a (a ≤ d), measured in model token-evaluations:

WorkTargetDraft
draft phase—seed + d tokens, d+1 sequential decode steps (latency-bound)
verification1 + d rows in one batched decode1 + d tokens in one replay batch (throughput-bound, cheap)
rollbackO(rejected) cell removal or snapshot restoresame

Net: each round costs the target one batched decode of 1+d rows (≈ prefill-like efficiency) and yields 1+a tokens. The draft model pays ≈ 2(1+d) token-evals (the factor 2 is the re-evaluation described in §7 — flagged TAG_SPEC_AVOID_DRAFT_REEVAL). So the draft model must be several times faster per token than the target for draft-simple to break even; with GPU-offloaded target and CPU draft (or vice versa) the sequential draft decodes can hide behind other work. Acceptance statistics (common_speculative_print_stats, speculative.cpp:2936–2980) report mean accepted length 1 + n_acc_tokens/n_call_accept and per-position acceptance rates — the practical knob is n_max: beyond the position where the per-position acceptance rate collapses, drafted tokens only cost decode steps (also benchmarked cheaply via synthetic acceptance, §10).

Postscript (2026-09-10, after step doc 81). The "one batched decode ≈ prefill-like efficiency" assumption in the table above is dispatch-dependent, not a law. Measured on GB10 with the same 7B-target/0.5B-draft pair: llama.cpp's own speculative-simple nets 1.00× end-to-end (the draft cost eats the whole gain at real acceptance p≈0.69 — the efficiency ceiling holds, the win does not), while minfer's batched path is not prefill-like at nt=2–15 (legacy kernel, one weight re-stream per token → per-token amortization 0.52×) and would have lost 2× before drafting even started. The mechanism on both sides — llama.cpp's MMVQ≤8/MMQ≥9 chain vs minfer's grid(od/4, nt) hole — is recorded in LLAMA-CPP-MMQ-ANALYSIS.md §12; the measurements are step doc 81. Consequence: minfer's D5 campaign is closed by measurement; a future multi-token feature would first need the small-M dispatch fix.

10. Observability & benchmarking hooks

  • Stats: impl-level counters printed by common_speculative_print_stats (speculative.cpp:2936–2980); speculative-simple additionally prints n_draft / n_predict / n_drafted / n_accept / accept% (speculative-simple.cpp:354–358).
  • Synthetic acceptance (benchmarking only, output is invalid): --spec-synth-rates P0,P1,... (unconditional per-position acceptance probabilities, must be finite, in [0,1], monotonically non-increasing — validated in common_speculative_synth_rates_resolve, speculative.cpp:2379–2411) or --spec-synth-len L (binary-searches a constant conditional p with p + p² + … + p^k = L − 1). The server then replaces real verification with server_sample_and_accept_synth (server-context.cpp:57–100), which draws u ~ U(0,1) per position against synth_probs[i], never accepts drafted EOG tokens, and keeps grammar/reasoning state consistent by not advancing it on synthetic tokens.

11. Notes for minfer

What a minfer port of draft-simple would need, mapped to the current graph architecture:

  1. Two graphs, two contexts. Target and draft are separate models → separate ComputeGraphs / GraphCaches (each with its own GraphParams identity). The draft context's n_ctx must match the target's (the KV regions live in the allocator, so both allocators must reserve the same prompt span; n_rs_seq = 0 analog: no snapshot path).
  2. Vocab gate (§4.2) is cheap metadata work over the GGUF tokenizer tables — same checks (vocab type, BOS/EOS, ≤128 size delta, token text equality from id 5).
  3. process() = replay batch without logits. In minfer terms: run the draft graph with the same positions/tokens inputs but mark the logits node dead (no host copy). The minfer rule "KV positions are data, not structure" is exactly what makes the replay possible: the same prefill-shaped graph serves both target and draft.
  4. draft() = decode loop with nt==1 graphs on the draft model, greedy = argmax over the logits node; p_min needs the top-k renormalized probability (minfer's sampler.rs top-k path can provide it).
  5. Rollback: minfer's per-layer persistent KV regions make seq_rm-style truncation a per-layer region rewrite (positions ≥ n_past dropped) — no snapshot machinery needed for dense models, matching the PART fast path.
  6. Verification is pure sampler work (sample_and_accept_n semantics: sample each row with the user chain, keep while equal, always emit ≥ 1 token) and slots naturally onto minfer's sampler.rs.

12. Source index (quick reference)

SymbolLocation
common_speculative_type enumcommon/common.h:171–183
common_params_speculative{,_draft,_ngram_*}common/common.h:324–395
common_speculative_draft_paramscommon/speculative.h:53–72
vocab compatibilitycommon/speculative.cpp:67–130
common_speculative_impl basecommon/speculative.cpp:138–177
impl_draft_simple (ctor/process/draft/accept)common/speculative.cpp:179–389
GGUF type auto-detectioncommon/speculative.cpp:2284–2319
common_speculative_n_maxcommon/speculative.cpp:2329–2377
synthetic rates resolvecommon/speculative.cpp:2379–2504
common_base_params_to_speculativecommon/speculative.cpp:2460–2503
common_speculative_init_result (draft ctx)common/speculative.cpp:2506–2587
common_speculative_init (dispatch/priority)common/speculative.cpp:2602–2745
common_speculative_draft (chain)common/speculative.cpp:2790–2873
common_speculative_acceptcommon/speculative.cpp:2875–2909
common_speculative_print_statscommon/speculative.cpp:2936–2980
output limitscommon/speculative.cpp:2589–2598
common_sampler_sample_and_accept_ncommon/sampling.cpp:678–715
common_sampler_sample (backend-sampling shortcut, chain apply)common/sampling.cpp:594–676
dist sampler (softmax + selected)src/llama-sampler.cpp:1150+
canonical loopexamples/speculative-simple/speculative-simple.cpp:162–342
common_context_can_seq_rm probecommon/common.cpp:1583–1627
common_context_seq_rm_typecommon/common.h:988–994
common_prompt_checkpointcommon/common.h:1165+
server drafting / checkpointstools/server/server-context.cpp:2965–3080
server accept / replaytools/server/server-context.cpp:3881–4004
server_sample_and_accept_synthtools/server/server-context.cpp:57–100
slot draft budgettools/server/server-context.cpp:483–500
CLI options (--spec-*)common/arg.cpp:4150–4330, upstream docs/speculative.md