Backends — the three execution engines
minfer runs one declarative compute graph (build → assign → fuse → allocate →
execute, docs/COMPUTE-GRAPH-DESIGN.md) on three interchangeable backends:
| CPU | Metal | CUDA | |
|---|---|---|---|
| Executor | src/graph/cpu_backend.rs | src/graph/metal_backend.rs | src/graph/cuda_backend.rs |
| Device layer | — (std threads) | src/metal/ + src/metal/kernels/ shaders | src/cuda.rs + src/cuda/kernels/*.cu |
| Platform | any | macOS (Apple GPU) | NVIDIA, opt-in --features cuda |
| Assign priority | last (always answers) | first on macOS | second, when built in |
| Activations | quantized to Q8_0 (Q8_K for K-quant weights) | read as f32 | f32; int8 MMQ for prefill |
| KV cache | f32 regions, or packed Q8_0 (MINFER_CACHE_TYPE=q8_0 — C4, 3.76× smaller) | f32, f16 or packed Q8_0 (C4 S2b, #310) | f32, f16 or packed Q8_0 (C4 S2b) |
| Async model | synchronous | one MpsCommandBuffer per split | stream + CUDA Graph capture/replay |
| Deep dives | walkthrough 10, 11, CPU optimizations | walkthrough 14, Metal optimizations | walkthrough 15, backend plan, campaign |
The src/<backend>/kernels/ paths and the CUDA src/cuda/methods/*.rs families are
the source layout plan (#261):
the CUDA half has landed (src/cuda/kernels/*.cu, src/cuda/methods/, src/quants/,
src/vec_ops/, src/kernel/), while the Metal split (src/metal/kernels/*.metal) and #53's
DeviceMemory answer are the Mac round. The Metal shaders are still the single src/metal/kernels/.
This page is the overview: what the backend contract is, how nodes land on a backend, and where the three differ. The linked pages carry the per-backend detail.
1. The contract: the Backend trait
Every backend implements one trait (src/graph/backend.rs:21-98). The
scheduler knows nothing about Metal or CUDA specifics — it only talks to this
surface:
| Method | Meaning |
|---|---|
supports_op(op, dtype) | capability query, asked per node at build time |
supports_fused(fused) | gates the fusion pass — fused IR nodes are only produced where a kernel exists (fusion.rs:72) |
alloc_buffer / free_buffer | the backend's own buffer pool, sized in f32 elements |
alloc_fresh | same, but bypasses the recycle free list — split-boundary staging needs ids whose physical contents are still referenced later in the same execute (backend.rs:36-45) |
execute_node(node, in_bufs, out_buf, kv_pair) | run one node; kv_pair carries the layer's persistent (K, V) region ids for KV ops; the output may alias an input (in-place ops) |
read_host / write_host | host access to a pool buffer — direct slices on CPU, staged transfers on GPU |
synchronize | wait for async work: CPU no-op; Metal submits the pending command buffer; CUDA closes a capture window if one is open |
graph_replay (CUDA only) | replay a previously captured CUDA Graph for a split (backend.rs:94-97, feature-gated; default no-op) |
Two supporting traits/rules ride along:
KvProvider::kv_pair(layer)— each layer owns two persistent KV regions (K and V) that live in the backend's pool and survive graph rebuilds (backend.rs:12-19; allocator detail in walkthrough 07).- The allocator is the single owner. Backends own pools, but buffers are only created through the allocator's liveness pass; input buffers are never freed, and in-place aliasing is decided at allocation time, not by kernels.
2. How a node gets its backend
Assignment happens once, at build time — never mid-run:
GraphAllocator::supports(op, dtype)walks the registered backends highest priority first: Metal → CUDA → CPU (src/graph/alloc.rs:140-153). The first backend whosesupports_opanswerstruewins that node. The CPU backend supports the full op set, so it always terminates the walk.- GPU participation is gated on weights: a GPU backend only claims ops
once all weight tensors are registered on it; the gate fails → the run
aborts with the actual values, never a silent CPU fallback
(
docs/GPU_SAFETY.md). - Whether any GPU took part is recorded in
CParams.gpu, which is part of the graph-reuse identity — a run that switches between GPU and CPU gets a differentGraphParamsand therefore a rebuilt graph (walkthrough 13).
Because assignment is per node, one forward pass can mix backends. The
scheduler cuts the node list into splits — maximal runs of the same
backend — and at every split boundary it synchronizes the previous backend and
copies split inputs across (copy_across, a host round trip through shared
memory; src/graph/scheduler.rs:176). Metal encodes one
MpsCommandBuffer per split and submits it at synchronize. Mechanically
this is
walkthrough 08 §3.
The error contract at execution time mirrors the build-time gate:
kernel-invariant violations return Err from execute_node — never a
silent fallback to another backend (e.g. a KV-store position ≥ n_ctx, an
attention head geometry mismatch, a device-limit shortfall). docs/GPU_SAFETY.md
is the binding rules page for the GPU backends (bounded submit waits, no early
return past a threadgroup_barrier, device limits queried at runtime).
3. What differs between the three
CPU — deterministic, zero-setup, bit-identical
- Quantized weights straight from the GGUF mmap; activations quantized to Q8_0 (32-value blocks) or Q8_K (306-byte super-blocks) per matmul — the int8×int8 dot kernels are AVX2 / NEON+SDOT with scalar fallbacks (walkthrough 10).
- Thread parallelism is ownership-based (one matmul row / one attention
head per worker), so results are bit-identical for any
--threadsvalue — the property the greedy-token verification gates rely on. - KV regions are f32; scores are computed in f32.
Metal — zero-copy on unified memory
- Weights register with
newBufferWithBytesNoCopy: the GGUF bytes are the Metal buffer, no copy, on Apple Silicon's unified memory (walkthrough 14). - Activations stay f32; a per-op dispatch matrix picks handwritten shaders (3 matmul tiers, 5 attention variants, rms_norm 2 widths) and decode fusions.
- Optional f16 KV regions halve attention bandwidth (
MINFER_CACHE_TYPE=f16).
CPU KV cache types — MINFER_CACHE_TYPE
f32 (default) · f16 (resolves to f32 here: this path has no f16 KV kernel) ·
q8_0 (C4: packed Q8_0 cells, ceil(n_kv_embd/32 * 34) bytes per cell instead of
4 * n_kv_embd, so the regions are 3.76× smaller; the store quantizes and the attention
reads the packed blocks directly — the K score is a Q8_0 × Q8_0 dot against the quantized
query, V accumulates out of the cell, and S1's dequantize-into-a-scratch pass is gone).
An unknown value is refused on every device, and a backend without a packed-read kernel refuses
q8_0 loudly rather than run f32 — the answer is the registry's reads_packed_kv, which is **true
for the CPU (C4 S1/S2a), CUDA (C4 S2b) and Metal (#310 enabled Metal's
packed read, whose attention window rides #44; the three C4 items left on
#87 — the packed fused epilogue, the packed FA prefill
and the dp4a packed dot — landed (#144 items 1+3, #186 item 2; the CUDA residual is #212). A physical context shift
(kv_rm/kv_shift) works on a packed region: the survivors move verbatim and are
re-rope/re-quantized one row at a time. MINFER_NO_FUSED_Q8_KV=1 restores the S1 read path
(the A/B of standing rule 3; measured 1.16× at ctx 512 and 1.31× at ctx 2048 in the fused
read's favour). The tolerance class against f32 is named, never bitwise: see the C4 record
in docs/ARCHITECTURE-EXECUTION-PLAN.md §5.
CUDA — opt-in, campaign-tuned
- Built only with
--features cuda(plain builds never touch nvcc);--features cuda,cuda_staticlinks cudart statically for deployment (docs/BUILD.md). - Prefill runs int8 MMQ tensor-core GEMMs; decode runs weight-streaming MMVQ kernels; attention is split-KV with a combine pass — the whole arc is the CUDA optimization campaign (r5–r60, D1–D4-4).
- Captures decode-shaped splits as CUDA Graphs and replays them
(
graph_replay;MINFER_NO_CUDA_GRAPH=1to disable) — the only backend with a capture/replay protocol on the trait.
Fusion capability differs per backend
The fusion pass consults supports_fused before producing fused IR nodes, so
the same model builds a different graph per backend: Metal accepts
SwiGLU (metal_backend.rs:705-709); CUDA accepts SwiGLU
(cuda_backend.rs:1303-1305) and carries its own fused decode kernels from
the campaign (see the CUDA docs for the inventory). The decode fusions
(Op::FusedQKV, Op::FusedFFN) are env-revertable
(MINFER_NO_FUSE_QKV=1 / MINFER_NO_FUSE_FFN=1) and are part of the reuse
identity. Fused vs unfused is bit-identical; when comparing, the unfused path
must still run the FusionPass.
4. Support matrix and forcing a backend
- Quant formats per backend (including the Metal prefill GEMM dispatch
window and CUDA MMQ notes):
docs/SUPPORT-MATRIX.md. Short version: Q4_0/Q4_1/Q5_0/Q5_1/Q8_0/Q4_K/Q5_K/Q6_K everywhere; Q2_K/Q3_K/I-quants nowhere. MINFER_DISABLE_MPS=1— force the CPU backend on macOS.MINFER_NO_NEON=1(aarch64) — drop the CPU NEON layer to scalar (A/B lever).MINFER_NO_AVX2=1(x86) — drop the whole CPU quants AVX2 layer to scalar (A/B lever);MINFER_NO_AVX512=1drops just the AVX-512/VNNI K-quant dots to AVX2.MINFER_NO_CUDA_GRAPH=1,MINFER_CACHE_TYPE=f32|f16|q8_0(q8_0= packed, on CPU + CUDA + Metal since #310, C4),MINFER_NO_FUSED_Q8_KV=1(keep S1's dequantizing read of a packed cache, for the A/B),MINFER_NO_FUSE_QKV/MINFER_NO_FUSE_FFN— per-backend behavior levers.- Which backend to expect: the startup banner and
MINFER_TRACE/MINFER_GRAPH_TRACEshow per-node assignments (walkthrough 08 §4).
Numerics across backends are not identical by design — CPU quantizes activations, GPUs read f32 (and CUDA prefill quantizes differently still), so CPU-vs-GPU logits differ; every path is verified against its own reference (greedy output equality with llama.cpp where noted in the support matrix).
5. Reading order
- walkthrough 08 — the scheduler — splits, copies, execution.
- Backend episodes of the walkthrough: 10 / 11 (CPU), 14 (Metal), 15 (CUDA).
docs/GPU_SAFETY.md— the hard rules before touching GPU code.- Per-backend history:
docs/CPU_OPTIMIZATIONS.md,docs/METAL_OPTIMIZATIONS.md,docs/CUDA-BACKEND-DESIGN.md+docs/CUDA_OPTIMIZATION.md.