Backends — the three execution engines

minfer runs one declarative compute graph (build → assign → fuse → allocate → execute, docs/COMPUTE-GRAPH-DESIGN.md) on three interchangeable backends:

CPUMetalCUDA
Executorsrc/graph/cpu_backend.rssrc/graph/metal_backend.rssrc/graph/cuda_backend.rs
Device layer— (std threads)src/metal/ + src/metal/kernels/ shaderssrc/cuda.rs + src/cuda/kernels/*.cu
PlatformanymacOS (Apple GPU)NVIDIA, opt-in --features cuda
Assign prioritylast (always answers)first on macOSsecond, when built in
Activationsquantized to Q8_0 (Q8_K for K-quant weights)read as f32f32; int8 MMQ for prefill
KV cachef32 regions, or packed Q8_0 (MINFER_CACHE_TYPE=q8_0 — C4, 3.76× smaller)f32, f16 or packed Q8_0 (C4 S2b, #310)f32, f16 or packed Q8_0 (C4 S2b)
Async modelsynchronousone MpsCommandBuffer per splitstream + CUDA Graph capture/replay
Deep diveswalkthrough 10, 11, CPU optimizationswalkthrough 14, Metal optimizationswalkthrough 15, backend plan, campaign

The src/<backend>/kernels/ paths and the CUDA src/cuda/methods/*.rs families are the source layout plan (#261): the CUDA half has landed (src/cuda/kernels/*.cu, src/cuda/methods/, src/quants/, src/vec_ops/, src/kernel/), while the Metal split (src/metal/kernels/*.metal) and #53's DeviceMemory answer are the Mac round. The Metal shaders are still the single src/metal/kernels/.

This page is the overview: what the backend contract is, how nodes land on a backend, and where the three differ. The linked pages carry the per-backend detail.

1. The contract: the Backend trait

Every backend implements one trait (src/graph/backend.rs:21-98). The scheduler knows nothing about Metal or CUDA specifics — it only talks to this surface:

MethodMeaning
supports_op(op, dtype)capability query, asked per node at build time
supports_fused(fused)gates the fusion pass — fused IR nodes are only produced where a kernel exists (fusion.rs:72)
alloc_buffer / free_bufferthe backend's own buffer pool, sized in f32 elements
alloc_freshsame, but bypasses the recycle free list — split-boundary staging needs ids whose physical contents are still referenced later in the same execute (backend.rs:36-45)
execute_node(node, in_bufs, out_buf, kv_pair)run one node; kv_pair carries the layer's persistent (K, V) region ids for KV ops; the output may alias an input (in-place ops)
read_host / write_hosthost access to a pool buffer — direct slices on CPU, staged transfers on GPU
synchronizewait for async work: CPU no-op; Metal submits the pending command buffer; CUDA closes a capture window if one is open
graph_replay (CUDA only)replay a previously captured CUDA Graph for a split (backend.rs:94-97, feature-gated; default no-op)

Two supporting traits/rules ride along:

  • KvProvider::kv_pair(layer) — each layer owns two persistent KV regions (K and V) that live in the backend's pool and survive graph rebuilds (backend.rs:12-19; allocator detail in walkthrough 07).
  • The allocator is the single owner. Backends own pools, but buffers are only created through the allocator's liveness pass; input buffers are never freed, and in-place aliasing is decided at allocation time, not by kernels.

2. How a node gets its backend

Assignment happens once, at build time — never mid-run:

  1. GraphAllocator::supports(op, dtype) walks the registered backends highest priority first: Metal → CUDA → CPU (src/graph/alloc.rs:140-153). The first backend whose supports_op answers true wins that node. The CPU backend supports the full op set, so it always terminates the walk.
  2. GPU participation is gated on weights: a GPU backend only claims ops once all weight tensors are registered on it; the gate fails → the run aborts with the actual values, never a silent CPU fallback (docs/GPU_SAFETY.md).
  3. Whether any GPU took part is recorded in CParams.gpu, which is part of the graph-reuse identity — a run that switches between GPU and CPU gets a different GraphParams and therefore a rebuilt graph (walkthrough 13).

Because assignment is per node, one forward pass can mix backends. The scheduler cuts the node list into splits — maximal runs of the same backend — and at every split boundary it synchronizes the previous backend and copies split inputs across (copy_across, a host round trip through shared memory; src/graph/scheduler.rs:176). Metal encodes one MpsCommandBuffer per split and submits it at synchronize. Mechanically this is walkthrough 08 §3.

The error contract at execution time mirrors the build-time gate: kernel-invariant violations return Err from execute_node — never a silent fallback to another backend (e.g. a KV-store position ≥ n_ctx, an attention head geometry mismatch, a device-limit shortfall). docs/GPU_SAFETY.md is the binding rules page for the GPU backends (bounded submit waits, no early return past a threadgroup_barrier, device limits queried at runtime).

3. What differs between the three

CPU — deterministic, zero-setup, bit-identical

  • Quantized weights straight from the GGUF mmap; activations quantized to Q8_0 (32-value blocks) or Q8_K (306-byte super-blocks) per matmul — the int8×int8 dot kernels are AVX2 / NEON+SDOT with scalar fallbacks (walkthrough 10).
  • Thread parallelism is ownership-based (one matmul row / one attention head per worker), so results are bit-identical for any --threads value — the property the greedy-token verification gates rely on.
  • KV regions are f32; scores are computed in f32.

Metal — zero-copy on unified memory

  • Weights register with newBufferWithBytesNoCopy: the GGUF bytes are the Metal buffer, no copy, on Apple Silicon's unified memory (walkthrough 14).
  • Activations stay f32; a per-op dispatch matrix picks handwritten shaders (3 matmul tiers, 5 attention variants, rms_norm 2 widths) and decode fusions.
  • Optional f16 KV regions halve attention bandwidth (MINFER_CACHE_TYPE=f16).

CPU KV cache types — MINFER_CACHE_TYPE

f32 (default) · f16 (resolves to f32 here: this path has no f16 KV kernel) · q8_0 (C4: packed Q8_0 cells, ceil(n_kv_embd/32 * 34) bytes per cell instead of 4 * n_kv_embd, so the regions are 3.76× smaller; the store quantizes and the attention reads the packed blocks directly — the K score is a Q8_0 × Q8_0 dot against the quantized query, V accumulates out of the cell, and S1's dequantize-into-a-scratch pass is gone). An unknown value is refused on every device, and a backend without a packed-read kernel refuses q8_0 loudly rather than run f32 — the answer is the registry's reads_packed_kv, which is **true for the CPU (C4 S1/S2a), CUDA (C4 S2b) and Metal (#310 enabled Metal's packed read, whose attention window rides #44; the three C4 items left on #87 — the packed fused epilogue, the packed FA prefill and the dp4a packed dot — landed (#144 items 1+3, #186 item 2; the CUDA residual is #212). A physical context shift (kv_rm/kv_shift) works on a packed region: the survivors move verbatim and are re-rope/re-quantized one row at a time. MINFER_NO_FUSED_Q8_KV=1 restores the S1 read path (the A/B of standing rule 3; measured 1.16× at ctx 512 and 1.31× at ctx 2048 in the fused read's favour). The tolerance class against f32 is named, never bitwise: see the C4 record in docs/ARCHITECTURE-EXECUTION-PLAN.md §5.

CUDA — opt-in, campaign-tuned

  • Built only with --features cuda (plain builds never touch nvcc); --features cuda,cuda_static links cudart statically for deployment (docs/BUILD.md).
  • Prefill runs int8 MMQ tensor-core GEMMs; decode runs weight-streaming MMVQ kernels; attention is split-KV with a combine pass — the whole arc is the CUDA optimization campaign (r5–r60, D1–D4-4).
  • Captures decode-shaped splits as CUDA Graphs and replays them (graph_replay; MINFER_NO_CUDA_GRAPH=1 to disable) — the only backend with a capture/replay protocol on the trait.

Fusion capability differs per backend

The fusion pass consults supports_fused before producing fused IR nodes, so the same model builds a different graph per backend: Metal accepts SwiGLU (metal_backend.rs:705-709); CUDA accepts SwiGLU (cuda_backend.rs:1303-1305) and carries its own fused decode kernels from the campaign (see the CUDA docs for the inventory). The decode fusions (Op::FusedQKV, Op::FusedFFN) are env-revertable (MINFER_NO_FUSE_QKV=1 / MINFER_NO_FUSE_FFN=1) and are part of the reuse identity. Fused vs unfused is bit-identical; when comparing, the unfused path must still run the FusionPass.

4. Support matrix and forcing a backend

  • Quant formats per backend (including the Metal prefill GEMM dispatch window and CUDA MMQ notes): docs/SUPPORT-MATRIX.md. Short version: Q4_0/Q4_1/Q5_0/Q5_1/Q8_0/Q4_K/Q5_K/Q6_K everywhere; Q2_K/Q3_K/I-quants nowhere.
  • MINFER_DISABLE_MPS=1 — force the CPU backend on macOS.
  • MINFER_NO_NEON=1 (aarch64) — drop the CPU NEON layer to scalar (A/B lever).
  • MINFER_NO_AVX2=1 (x86) — drop the whole CPU quants AVX2 layer to scalar (A/B lever); MINFER_NO_AVX512=1 drops just the AVX-512/VNNI K-quant dots to AVX2.
  • MINFER_NO_CUDA_GRAPH=1, MINFER_CACHE_TYPE=f32|f16|q8_0 (q8_0 = packed, on CPU + CUDA + Metal since #310, C4), MINFER_NO_FUSED_Q8_KV=1 (keep S1's dequantizing read of a packed cache, for the A/B), MINFER_NO_FUSE_QKV / MINFER_NO_FUSE_FFN — per-backend behavior levers.
  • Which backend to expect: the startup banner and MINFER_TRACE / MINFER_GRAPH_TRACE show per-node assignments (walkthrough 08 §4).

Numerics across backends are not identical by design — CPU quantizes activations, GPUs read f32 (and CUDA prefill quantizes differently still), so CPU-vs-GPU logits differ; every path is verified against its own reference (greedy output equality with llama.cpp where noted in the support matrix).

5. Reading order

  1. walkthrough 08 — the scheduler — splits, copies, execution.
  2. Backend episodes of the walkthrough: 10 / 11 (CPU), 14 (Metal), 15 (CUDA).
  3. docs/GPU_SAFETY.md — the hard rules before touching GPU code.
  4. Per-backend history: docs/CPU_OPTIMIZATIONS.md, docs/METAL_OPTIMIZATIONS.md, docs/CUDA-BACKEND-DESIGN.md + docs/CUDA_OPTIMIZATION.md.