minfer Architecture
A pure-Rust LLM inference engine written from scratch, inspired by llama.cpp, with zero ML framework dependencies. Inference runs through a declarative compute graph (builder → scheduler → per-backend kernels), modeled on llama.cpp's
ggml_cgraph+ backend scheduler. This document describes the overall design: module responsibilities, the compute graph pipeline, the CPU / Metal backend layering, quantization, adding a new model architecture, and adding a new backend. The pre-graph imperative forward is preserved at the end as an appendix (Appendix A).
1. Design Principles
- No ML framework — attention, RMSNorm, RoPE, SiLU, softmax are all
handwritten. Only 5 external crates:
rand,regex,half,serde/serde_json,minijinja. - Declarative compute graph, not an imperative loop — the forward pass is
built as a pure
ComputeGraph(no side effects at build time), then assigned to backends, fused, allocated, and executed by the scheduler. This mirrors llama.cpp (ggml_cgraph+ggml_backend_sched) and enables graph reuse, per-op backend assignment, DOT export, and a clean path to new backends. Design and implementation record:docs/COMPUTE-GRAPH-DESIGN.md. - Bytes-in / bytes-out tensors — weight tensors are raw
&[u8]; SIMD dot-product kernels (AVX2 / Metal shaders) operate on byte slices matching the exact GGML quantized block layout (repr(C)inblock.rs). - Activations stay f32 — CPU matmuls quantize activations to Q8_0 on-the-fly; the Metal backend reads f32 activations directly for all weight types (matching llama.cpp's Metal backend).
- Backend assignment is a build-time decision, never a silent mid-execution
fallback —
supports_opdecides which ops run where; cross-backend transfers happen at split boundaries; kernel-invariant violations abort (gpu_abort/Err), they never silently fall back to CPU.
2. Module Map
| Module | Responsibility |
|---|---|
main.rs | CLI, GGUF load, chat template, prefill → autoregressive generation loop, timing |
graph/ | Compute graph core — IR, builder, scheduler, backends, reuse cache (see §4 table) |
gguf.rs | GGUF v3 parser (metadata KV + tensor table + data blob), multi-part (split) support, ggml_pad alignment |
block.rs | 20+ quantized block types as repr(C) structs + fp16 conversions, matching ggml-common.h |
quants.rs + src/quants/*.rs | AVX2+FMA / NEON+SDOT dot-product kernels (Q4_0/Q4_1/Q5_0/Q5_1/Q8_0/K-quants × Q8_0/Q8_K) + f32→Q8_0/Q8_K quantization, scalar fallback. The module file is the decider; the parts are src/quants/{dot_q4_0,dot_q4_1,dot_q5,dot_q8_0,kquant,quantize_q8_0,quantize_q8_k,avx2,neon}.rs (#264, the source layout plan) |
kernel.rs + src/kernel/*.rs | Quantized matmul dispatch (Q4_0/Q4_1/Q5_0/Q5_1/Q4_K/Q5_K/Q6_K/Q8_0) over activations, CPU scalar fallback, the shared worker pool, and the shared embed_tokens row getter |
vec_ops.rs + src/vec_ops/*.rs | SIMD vector ops: RMSNorm, RoPE (Qwen2/Llama styles), softmax, SiLU, add/scale/mul, plus the f16/bf16 weight-row dots |
tensor.rs | 4D Tensor (type/shape/strides/Vec<u8> data), ggml-compatible strides & byte sizing |
sampler.rs | Repeat/frequency/presence penalties → top-k → top-p → temperature, seeded StdRng |
tokenizer.rs | Self-contained BPE tokenizer, loaded from GGUF metadata (no tiktoken) |
template.rs | GGUF chat_template rendering via minijinja + a Python-str-method hook (F7/#50), so Qwen3's think-block template renders; a template that cannot be rendered is a loud refusal, never a silent ChatML fallback (docs/CHAT-TEMPLATE-AND-TOKENIZER-DESIGN.md) |
models/ | Architecture implementations. mod.rs has the ModelDef trait + factory dispatch |
models/qwen2/ | Qwen2/Qwen2.5: mod.rs (model struct + trait impl), graph.rs (Qwen2Graph::build/forward), loader.rs (GGUF weights + hparams) |
models/qwen3/ | Qwen3 (dense): same triple — decoupled head dim + per-head Q/K norm (qk_norm); its ChatML+<think> template renders since F7 |
conversation.rs | Multi-turn session state (append-only KV) behind --cnv |
server/ | OpenAI-compatible HTTP server (serve): axum + tokio, multi-slot, /v1/chat/completions streaming; viz.rs serves the viz page |
src/metal/ + src/metal/kernels/*.metal | Apple MPS (Metal) backend: L1 runtime (MpsState/MetalDevice), L2 command-buffer encoding, L3 shaders (the legacy whole-layer layer_gpu is retained for tests). The split into src/metal/{runtime,encode,ops,policy}.rs + src/metal/kernels/ is the source layout plan (#265) |
cuda.rs + src/cuda/kernels/*.cu | NVIDIA CUDA device layer + launch layer + kernels (feature-gated --features cuda); executed through the graph via graph/cuda_backend.rs. The split is the source layout plan (#262, #263) |
device_tier.rs | cc-keyed device tier table + selector (measured GB10 row, llama.cpp-adopted consumer rows, GENERIC fallback); resolved once at init, feeds the MMQ gate, smem feasibility and plane-VRAM budget checks (docs 105–106) |
download/ | Hugging Face Hub + Ollama download, cached-name resolution, resume support |
dump.rs | Per-layer hidden-state debug dump (gated by --features debug_dump) |
bench.rs / live.rs / trace.rs / spec_verify.rs | llama-bench-style bench, live viz inference host, per-node trace export (MINFER_TRACE), D5 verify-step gate bench |
The backend layers — runtime, launch, kernels, executors
Each device backend is organised along four layers, and only the last one is polymorphic:
| layer | CUDA | Metal | CPU |
|---|---|---|---|
| L1 device/runtime | src/cuda/methods/{init,weights,stream,buffers,copy,events,capture}.rs | src/metal/runtime.rs | — (std threads) |
| L2 launch/dispatch | src/cuda/methods/<family>.rs + the extern "C" declarations each family owns (under src/cuda/methods.rs) | src/metal/encode.rs + ops.rs | src/kernel/*.rs |
| L3 kernel sources | src/cuda/kernels/*.cu + common.cuh | src/metal/kernels/*.metal | src/quants/*.rs, src/vec_ops/*.rs |
| L4 graph executor | src/graph/cuda_backend.rs | src/graph/metal_backend.rs | src/graph/cpu_backend.rs |
(The CUDA paths in that table have landed (#262,
#263); the Metal ones are still the target of Step 4
(#265, Mac-local). docs/SOURCE-LAYOUT-PLAN.md §8 has
the file-by-file tree and each step updates this table as it lands.)
Two rules follow, and they are why the crate keeps one flat interface instead of a directory per layer:
- Device is the first axis, the layer is the second.
Backend+registry.rs(src/graph/backend.rs) are the one device seam — llama.cpp keepsggml-backend*.cpp/hflat beside its per-device directories for the same reason. L2 cannot leave L1: the CUDA launchers are inherent methods ofCudaState, the Metal ones methods ofMpsCommandBuffer; splitting them out would be a type refactor, not a file move. - A
commonneeds a second real implementation. A shared abstraction is added only when at least two backends implement it and at least two callers use it with the same semantics. The one candidate today isallocplan::DeviceMemory(CUDA answers it; the Metal half is #53). CPU'squants.rs/vec_ops.rsare deliberately not a device-private layer: they are the crate's numeric kernel library, consumed bygraph/kvformat.rsandgraph/cuda_backend.rs.
src/graph/ — the compute graph core
| File | Role |
|---|---|
mod.rs | ComputeGraph (topo-validated node list + inputs/outputs), CNode, DType, BufRef; re-exports the Backend handle from registry.rs |
registry.rs | The backend registry (F4): the fixed id space (Backend::CPU/METAL/CUDA), the name surface and its three startup refusals, each entry's priority + capability record + pool/host-I/O hooks, and the --backend / MINFER_BACKENDS filter. Contract: docs/BACKEND-REGISTRY-DESIGN.md |
ops.rs | Op enum (full payload PartialEq), NodeMeta, AttnMode, FusedOp |
builder.rs | GraphBuilder — declarative construction (embedding/rms_norm/matmul/rope/attn/kvcache/…) |
alloc.rs | Per-backend liveness allocator + persistent per-layer KV regions + KvProvider |
backend.rs | Backend trait + KvProvider |
cpu_backend.rs | CPU execution (wraps kernel.rs + vec_ops.rs) |
metal_backend.rs | Metal execution (per-op MPS kernels; cfg(target_os = "macos")) |
cuda_backend.rs | CUDA execution (feature-gated): int8 MMQ prefill + MMVQ decode (default-on), split-KV attention, CUDA Graph capture/replay |
scheduler.rs | assign → split → execute (+ cross-backend copies at split boundaries) |
fusion.rs | Pattern-matching fusion (SwiGLU), gated by backend supports_fused |
cache.rs | GraphCache — params-only deterministic graph reuse |
params.rs | GraphParams/CParams/GraphType — the reuse identity |
dot.rs | Graphviz DOT export (--dump-graph) |
json.rs | Graph JSON export for the viz page (--dump-graph-json, MINFER_TRACE) |
3. Inference Pipeline
The top-level flow lives in main.rs. The whole engine is a single-pass
prefill followed by an autoregressive decode loop; both call
ModelDef::forward, which routes through the compute graph. Every forward is
build → assign → fuse → alloc → execute (the subgraph below), but the
built graph is cached per GraphParams and reused — only the first call with
new params pays the build/assign/alloc cost.
flowchart TD
A["CLI args: model, prompt, flags"] --> B{"resolve model<br/>download::resolve"}
B -->|local path| C["load GGUF v3<br/>single or split parts (mmap)"]
B -->|"hf:… / ollama:…"| D["auto-download → path"]
B -->|cached name| C
C --> E["parse metadata KV + tensor table"]
E --> F["init GPU backend<br/>MPS / CUDA"]
F --> G["load model<br/>dispatch on general.architecture"]
G --> MODE{"mode?"}
MODE -->|"--cnv"| CV["conversation REPL<br/>(run_conversation)"]
MODE -->|"serve"| SV["OpenAI-compatible HTTP server<br/>multi-slot, streaming"]
MODE -->|"viz"| VZ["viz server (page + live SSE)"]
MODE -->|single shot| H["load BPE tokenizer from GGUF"]
H --> J{"no-template?"}
J -->|no| K["render chat template<br/>GGUF chat_template via minijinja<br/>(unrenderable → loud refusal;<br/>no template at all → ChatML)"]
J -->|yes| L["raw prompt"]
K --> M["tokenize prompt<br/>n_ctx = max(--n-ctx, prompt len)<br/>(sizes the KV regions once)"]
L --> M
M --> N["PREFILL<br/>graph forward, all prompt tokens at once"]
N --> O["last-token logits (n_out = 1)"]
O --> P{"DECODE loop<br/>while generated < n_predict"}
P --> Q["sample next token<br/>penalties → top-k → top-p → temp"]
Q --> R{"stop token or<br/>stop string match?"}
R -->|yes| S["done"]
R -->|no| T["append token, decode+print"]
T --> U["graph forward, single token<br/>KV persists in the allocator"]
U --> P
subgraph GRAPH["every forward: build → assign → fuse → alloc → execute"]
G1["GraphBuilder<br/>build_graph (pure IR)"] --> G2["assign backends<br/>priority Metal → CUDA → CPU"]
G2 --> G3["fuse<br/>SwiGLU + decode fusions (gated)"]
G3 --> G4["alloc<br/>liveness + persistent KV"]
G4 --> G5["execute<br/>per split, cross-backend copies"]
end
N -.->|"GraphCache: params-only reuse"| GRAPH
U -.->|"GraphCache: params-only reuse"| GRAPH
Generation parameters (defaults match llama.cpp): temp=0.8, top_k=40,
top_p=0.95, repeat_penalty=1.1 (last 64 tokens), frequency_penalty=0.0,
presence_penalty=0.0, seed=42, n_ctx=4096, n_predict=512. Sampling
applies the three penalties in one pass, then top-k → top-p → temperature.
Timing is dual-caliber: Prefill: = prompt tokens / prefill wall time;
Generated: = generated tokens / decode wall time (pure decode, matches
llama-bench "Generation" caliber); Total: = blended.
4. Compute Graph Architecture
Inference = build a ComputeGraph (pure, side-effect free) → assign backends
→ fuse → allocate → execute. The graph is built once per distinct
GraphParams and reused (decode steps reuse the same graph; the model's
forward() routes through Qwen2Graph::forward).
4.1 The IR
ComputeGraph= topologically orderednodes(builder appends sources before consumers),inputs,outputs,uid(for CUDA-Graph-style reuse).CNode=op+srcdependencies +out_shape/out_dtype+backend(None until assigned) +meta(weight names, rope/attn params).Opcarries full payloads (RmsNorm{eps},MatMul{transpose_b},KvcacheStore{layer}…) so graphs are structurally comparable.
4.2 Builder (GraphBuilder)
Per-architecture code calls builder methods (mirroring llama.cpp's
llm_graph_context): embedding, rms_norm, qk_norm (Qwen3 per-head Q/K
norm), matmul, get_rows, rope, silu, add, mul, swiglu,
softmax, attn, kvcache_store/kvcache_load, plus the decode-fusion
constructors fused_qkv/qkv_bias_rope_store/fused_ffn. Building is pure —
no computation happens at build time.
4.3 Scheduler pipeline
assign_backends → fuse → alloc_graph → execute
- assign_backends — capability-driven: each node gets the highest-priority
backend whose
supports_opreturns true (priority Metal → CUDA → CPU). Weight registration decides GPU feasibility. - fuse — pattern matching (
Mul(Silu(X),Y) → SwiGLU) gated per backend bysupports_fused(no double-fusion with hand-written kernels). The plan's second rule (RoPE(Add(X,B)) → FusedBiasRope) was removed: no backend ever claimed the capability, and the bias+rope work ships as the builder's decode nodes (FusedQKV/QkvBiasRopeStore, whoseattn_bias_rope_storekernel subsumes the bias+rope part).BatchMatMulstays deferred (single-output IR limitation,COMPUTE-GRAPH-DESIGN.md§5.4). - alloc_graph — per-backend liveness allocator: buffers shared between nodes whose live ranges don't overlap; persistent per-layer KV regions survive rebuilds; in-place ops alias their input buffer (see §4.5).
- execute — per split (contiguous same-backend runs): sync the previous
backend, copy split inputs across backends (
copy_across, a host round trip via shared memory), run the nodes, then a final sync. Metal batches oneMpsCommandBufferper split, submitted atsynchronize().
4.4 Reuse (GraphCache)
Params-only deterministic reuse (llama.cpp allow_reuse invariant):
GraphParams = n_tokens / n_out (tail rows) / gtype /
cparams (n_ctx, flash_attn, gpu, fuse_qkv, fuse_ffn,
explicit_span) / weights_version deterministically determines the topology —
equal params ⇒
identical graph.
n_past is deliberately absent (it is execution data). CParams.gpu records
backend participation so a backend toggle forces a rebuild. GraphCache owns
the allocator (so the KV regions persist across rebuilds, e.g. the
prefill→decode transition) and try_reuse compares params only; debug builds
assert structural consistency (Op: PartialEq).
4.5 Core invariants (must not be violated)
- KV positions are data, not structure.
KvcacheStore/Loadcarry only the layer index; write positions come from thepositionsinput node. Topology never depends onn_past. - Each layer owns TWO persistent KV regions (K and V), resolved by
kv_pair(layer)(KvProvider). The store node's output buffer is the K region; backends write the V sibling viakv_pair. - GGUF weight layout: metadata
[in, out](ne[0] fastest), memory[out][in]row-major → matmulod = shape[1],id = shape[0]. Activations: shape metadata[d, nt, 1, 1], memory token-major[nt][d]. I32 inputs are stored asf32::from_bitsbit patterns (fill_input_i32). - In-place ops (
Silu,RoPE) alias their input buffer — the allocator maps the output to the input'sBufRef(only when the input's sole consumer is this op AND it is on the same backend). Never host-copy a GPU-pending buffer: a hostcopy_inof a producer that is encoded but not submitted reads stale data (the Phase-3 KV-corruption bug). Cross-backend in-place inputs get a fresh buffer (the producer completed before the split boundary, so the copy is safe there). - Execution follows build order (a valid topological order by construction) — guarantees a KV store executes before the attention that reads it. Nodes with no allocated buffer (dead, e.g. fusion orphans) are skipped.
- CPU vs GPU activation paths differ numerically: CPU matmuls are Q8_0×Q8_0 (activation-quantized), the GPU backends read f32 activations for all weight types (CUDA prefill additionally offers the default-on int8 MMQ path). Compare each path against its own reference, not against the other.
4.6 Per-layer computation (Qwen2) — as built by graph.rs
flowchart LR
A["token_embd lookup (GetRows)"] --> B["hidden"]
B --> C["RMSNorm attn_norm"]
C --> D["WQ / WK / WV matmuls + bias"]
D --> E["RoPE on Q and K (in-place alias)"]
E --> F["kvcache_store: K/V → layer regions"]
F --> G["GQA attention<br/>Q·K^T → softmax → ·V"]
G --> H["WO matmul + bias"]
H --> I["+ residual → hidden"]
I --> J["RMSNorm ffn_norm"]
J --> K["FFN gate + up matmuls"]
K --> L["SiLU(gate) × up (fused SwiGLU)"]
L --> M["FFN down matmul"]
M --> N["+ residual → hidden"]
N --> O["next layer / output_norm"]
GQA: each query head h maps to KV head hk = h / gqa. The KV head dimension
is independent (n_kv_embd read from the K weight's ne[1]), so Qwen2.5-0.5B
(n_embd=896, n_head=14, hd=64, n_kv_embd=128) strides correctly.
4.7 Prefill vs decode
| Phase | nt (tokens) | Notes |
|---|---|---|
| Prefill | > 1 | graph type Prefill; KV store writes all positions, attention reads the full written prefix |
| Decode | 1 | graph type Decode; same topology as prefill modulo nt → the graph is rebuilt once (KV persists), then reused for every subsequent token |
The graph path keeps K/V on the executing backend (attention and KV are on the same backend by construction), so there is no per-token KV drain.
5. Backend Layering
5.1 The Backend trait (graph/backend.rs)
Backend (the thing the IR stores on a node) is not this trait: since F4 it is a Copy
handle — an id into src/graph/registry.rs — and the trait below is implemented by the pools
(CPU, Metal, CUDA) and held by the registry entry. The name, id, priority and capability record
live on the handle/entry, not on the trait (Backend::name() is registry.rs:105):
#![allow(unused)] fn main() { pub trait Backend: Send + Sync { fn kv_pair(&self, layer: usize) -> Option<(usize, usize)>; // persistent per-layer K/V regions fn supports_op(&self, op: &Op, dtype: DType) -> bool; fn supports_fused(&self, fused: &FusedOp) -> bool; fn supports_attn_span(&self) -> bool { false } // E1: explicit attention windows (CPU/CUDA/Metal) fn weights_bytes(&self) -> usize { 0 } // E4: bytes the memory budget is charged for fn alloc_buffer(&mut self, size: usize) -> usize; // backend's own pool fn free_buffer(&mut self, id: usize); fn pool_len(&self) -> usize; fn alloc_fresh(&mut self, size: usize) -> usize; // bypasses the recycle free list (split-boundary staging) fn execute_node(&mut self, node: &CNode, in_bufs: &[usize], out_buf: usize, kv_pair: Option<(usize, usize)>) -> Result<(), String>; fn copy_cells(&mut self, dst: BufRef, src: BufRef, dst_row: usize, src_row: usize, rows: usize, elems_per_cell: usize) -> Result<(), String>; // C3 compaction / C2 shift fn read_host(&self, id: usize) -> Option<&[f32]>; fn write_host(&mut self, id: usize, data: &[f32]) -> Result<(), String>; fn write_host_window(&mut self, id: usize, offset: usize, data: &[f32]) -> Result<(), String>; fn synchronize(&mut self); fn retire(&mut self) { self.synchronize(); } // release device objects for a dropped pool // CUDA only: try to replay a captured graph for (uid, range); capture is // gated to decode-shaped graphs. Default impl returns false. #[cfg(feature = "cuda")] fn graph_replay(&mut self, uid: u64, range: (usize, usize), nt_hint: Option<usize>) -> bool; } }
- CPU (
cpu_backend.rs):Vec<f32>pool; executes viakernel.rs+vec_ops.rs; F32 weights use a plain f32 matmul, quantized weights usecpu_quant_matmul_f32(Q8_0-activation path). - Metal (
metal_backend.rs, macOS): shared-memoryMTLBufferpool; per-op dispatch toMpsState's kernels (rms_norm, quant_matmul_f32_on_gpu_buf, rope, silu, add, mul, swiglu, embed_tokens_gpu, store_kv, gqa_attn_f32); one command buffer per split. Weights resolve by name from MpsState's registry (weight_buf(name) -> (buffer, offset)). - CUDA (
cuda_backend.rs, feature-gated--features cuda): wraps thecuda.rsdevice layer — per-op dispatch with int8 MMQ prefill + MMVQ decode (default-on;MINFER_MMQ=0reverts; the MMQ gate resolves from thedevice_tier.rsdevice-tier table), split-KV attention, and CUDA Graph capture/replay keyed on the graphuid(graph_replay, decode-shaped only). Implementation record:docs/CUDA-BACKEND-DESIGN.md; per-step optimization history indocs/CUDA_OPTIMIZATION.md.
The allocator owns every backend pool (single source of truth); the scheduler orchestrates assignment, cross-backend copies, and sync.
5.2 Selection rules
Assignment asks each node in the registry's priority order — Metal 300, CUDA 200, CPU 100
(src/graph/registry.rs:72-75) — for the first backend that has a pool, is not fenced off, and
reports the op supported; the CPU is tried last and is always allowed, so the answer is total
(GraphAllocator::supports_for, src/graph/alloc.rs:403). The fence is --backend <name> /
MINFER_BACKENDS=<csv>, and an unknown name, a name this build does not contain and a name this
machine cannot use are three distinct startup refusals. The two orders, the name surface and
those refusals are docs/BACKEND-REGISTRY-DESIGN.md's contract; the user-facing flags are
docs/USAGE.md.
- Metal: all graph weights must be GPU-registered (
Qwen2Graph::weights_on_gpumirrors the old per-layer check).MINFER_DISABLE_MPS=1forces CPU. - CUDA: feature-gated (
--features cuda), and every graph weight must be GPU-registered with a device present. A node whose block the E5 offload plan left on the CPU never reaches the device (supports_foranswersCPUfor it, so a partially offloaded model cannot run a block whose weights were never registered there).MINFER_DISABLE_CUDA=1is presence-checked and forces the CPU. - CPU: always available; AVX2 dispatch via
is_x86_feature_detected!("avx2")on x86, NEON+SDOT on aarch64, scalar fallback elsewhere (MINFER_NO_NEON=1forces scalar; the K-quant dots also have an AVX-512/VNNI path since #56, gated byMINFER_NO_AVX512=1andMINFER_NO_AVX2=1).
5.3 GPU safety
All Metal submits wait bounded (10 s) and check status; no early return past a
threadgroup_barrier; device limits queried at runtime, never hardcoded; guard
failures gpu_abort with actual values. In the graph architecture,
kernel-invariant violations return Err from execute_node and must NOT be
treated as a silent CPU fallback — backend assignment at build time decides
where ops run; only genuine support limitations (e.g. Raw weights) select the
CPU backend. See docs/GPU_SAFETY.md.
6. Quantization & Tensor Layout
- Weight layout (
block.rs):repr(C)blocks matchingggml-common.h. Q4_0/Q4_1/Q5_0/Q5_1/Q8_0 = 32-value blocks; Q4_K/Q5_K/Q6_K = 256-value super-blocks. Supported: Q4_0, Q4_1, Q8_0, Q4_K, Q6_K, Q5_0, Q5_1, Q5_K (CPU + Metal), F32/F16 norms & biases. - Activation quantization: CPU matmuls quantize f32 activations to Q8_0
on-the-fly (
quants.rs); Metal reads f32 directly. - GGUF parsing (
gguf.rs):ggml_pad(x, n) = (x + n - 1) & !(n - 1)alignment; tensor strides computed from type block size exactly like ggml; split multi-part models are merged into one tensor index.
flowchart LR
A["GGUF file"] --> B["metadata KV<br/>hparams + tokenizer + template"]
A --> C["tensor table<br/>name / type / shape / offset"]
C --> D["quantized data blob"]
D --> E["Tensor: type, shape, strides, Vec<u8>"]
B --> F["HParams"]
B --> G["Tokenizer"]
B --> H["Chat template"]
7. KV Cache
The graph path owns the KV cache: two persistent regions per layer (K and
V), sized n_kv_embd × n_ctx, allocated in the backend pool the layer runs
on. kv_pair(layer) resolves them; the KV store node writes K/V at the
positions carried by the positions input; attention reads the written prefix.
The regions live inside the GraphCache's allocator and survive graph
rebuilds (the prefill→decode transition). MINFER_CACHE_TYPE=f16 selects an
f16 GPU cache where the kernels support it. The pre-graph cache.rs KVCache
type — and the vestigial &mut KVCache argument of ModelDef::forward — was
deleted in #252; these regions
are the only KV store.
8. Adding a New Architecture
- Create
src/models/<name>/withmod.rs,graph.rs,loader.rs. - Add a
matchbranch insrc/models/mod.rs::load_model()for the newgeneral.architecturevalue. - In
loader.rs: defineHParams(includingn_kv_embd) andLayerWeights, parse them from GGUF metadata (bothqwen2.*andllama.*prefixes are accepted). - In
graph.rs: implementbuild_graph(&self, params: &GraphParams) -> ComputeGraphdeterministically in params (the reuse invariant), usingGraphBuilder— mirror llama.cpp'sllm_graph_contextbuilder methods. - In
mod.rs: implementModelDef(forward,build_graph,forward_graph,forward_graph_cached,as_any,special_tokens,n_layer/n_head_kv/n_embd_head/n_kv_embd/n_vocab,rope_style).models/qwen3/is the worked example of a second architecture (decoupled head dim + per-head Q/K norm). - If needed, add a chat template format in
template.rs.
Architectures that share Qwen2's tensor naming convention (LLaMA, Mistral,
Phi) are the easiest ports. The RopeStyle enum defines both Qwen2
(non-interleaved) and Llama (interleaved) pairings, but only the
non-interleaved form is wired up: both loaders hard-code it
(models/qwen2/loader.rs:155, models/qwen3/loader.rs:173) and the CUDA
backend refuses the interleaved style (graph/cuda_backend.rs:1522). A family
that needs interleaved RoPE requires a loader change plus a CUDA kernel — see
docs/MODEL-SUPPORT-ROADMAP.md, "Cost model: what a port actually costs".
9. Model Download
download/mod.rs resolves hf:<repo>[:<quant>] and ollama:<model>[:<tag>]
URIs, downloads via curl with size-checked resume, quant-matches single or
split files case-insensitively, and stores them under
~/.cache/minfer/models. Cached filenames can be used directly as the model
argument.
10. Related Documentation
| Topic | Location |
|---|---|
| Compute graph design + rewrite plan + implementation record (per-phase commits) | docs/COMPUTE-GRAPH-DESIGN.md |
| llama.cpp compute-graph design analysis (ggml_cgraph / scheduler / reuse) | docs/LLAMA-COMPUTE-GRAPH.md |
| Metal backend optimizations / gap analysis (primary tracking) | docs/METAL_OPTIMIZATIONS.md |
| GPU safety conventions + audit | docs/GPU_SAFETY.md |
| CPU backend optimizations | docs/CPU_OPTIMIZATIONS.md |
| CUDA optimization history + per-step records | docs/CUDA_OPTIMIZATION.md (+ docs/cuda_optimization_steps/) |
| Debug dump format | docs/debug-dump.md |
| Historical bugs / debugging notes | docs/BUG-6-KV-CACHE-INDEXING.md, docs/QWEN2.5-*, docs/DEBUGGING-* |
Appendix A — Legacy Imperative Architecture (removed in Phase 6)
Historical reference. The imperative per-layer forward (
models/qwen2/forward.rs) was the engine's core until the compute graph replaced it (Phase 6); it is preserved here verbatim in structure so old notes, benchmarks, and kernel analyses remain interpretable. Do not treat this as the current design.
A.1 Design stance (then)
The engine ran a direct per-layer forward loop instead of a compute graph: "simpler, easier to trace, and the whole layer can be fused onto the GPU." The GPU fallback was per-layer and safe: a layer that could not run on the GPU (e.g. unsupported weight type) submitted partial GPU work, downloaded the hidden state, and continued on the CPU.
A.2 The old forward pass
ModelDef::forward(tokens, positions, kv) was implemented in
models/qwen2/forward.rs as an imperative loop:
flowchart LR
A["token_embd lookup"] --> B["hidden"]
B --> C["RMSNorm attn_norm"]
C --> D["WQ / WK / WV matmuls + bias"]
D --> E["RoPE on Q and K"]
E --> F["store K/V into KV cache"]
F --> G["GQA attention<br/>Q·K^T → softmax → ·V"]
G --> H["WO matmul + bias"]
H --> I["+ residual → hidden"]
I --> J["RMSNorm ffn_norm"]
J --> K["FFN gate + up matmuls"]
K --> L["SiLU(gate) × up"]
L --> M["FFN down matmul"]
M --> N["+ residual → hidden"]
N --> O["next layer / output_norm"]
Decode-time optimizations: fused QKV (nt==1) via a concatenated
blk.{il}.attn_qkv weight, fused bias+rope+store (attn_bias_rope_store),
fused SwiGLU kernel, and the last layer computed only the tail n_out rows
(an inp_out_ids-style partial-row optimization). The n_out tail-row
optimization was subsequently carried into the graph path as the G3 work:
the builder inserts GetRows(wo, tail_ids) (and the same for the residual)
before the last layer's FFN, so the last FFN, the final norm and lm_head all
run on n_out rows only (models/qwen2/graph.rs:64, :228-231;
GraphParams.n_out is part of the reuse identity).
A.3 Old backend layering & fallback
flowchart TD
A["forward nt tokens"] --> B{"embedding on GPU?"}
B -->|yes| C["GPU embed lookup → buf_hidden"]
B -->|no| D["CPU embed → upload hidden"]
D --> E["upload positions"]
C --> E
E --> F{"per-layer: layer_gpu ok?"}
F -->|"yes, all layers"| G["output_norm_gpu<br/>on GPU"]
F -->|"no at layer i"| H["submit partial GPU work<br/>download hidden, sync KV to CPU"]
H --> I["CPU loop from layer i"]
G -->|"output on GPU"| J["download logits → return"]
G -->|"output fell back"| I
I --> K["output_norm + output matmul on CPU"]
K --> L["return logits"]
Selection rules (then):
- Metal: layer 0 must have all 7 weight matrices + norms registered on the
GPU. Within a layer all 7 matrices must be all Q4 group (Q4_0/Q4_1) or
all QK group (Q4_K/Q5_0/Q6_K); Q5_1/Q5_K use the f32 path and are exempt.
MINFER_DISABLE_MPS=1forces CPU. - CUDA (
--features cuda): requires every layer's 7 matrices to be all Q4_0/Q4_1 or all Q4_K/Q6_K. Decode replays a captured CUDA Graph. - CPU: always available; AVX2 dispatch via
is_x86_feature_detected!("avx2"), scalar fallback elsewhere (plus the AVX-512/VNNI K-quant dots, #56).
The GPU path skipped the per-token CPU→GPU KV drain (no sync_kv_to_cpu)
because GPU-layer failure is deterministic by weight type — the sync only
happened in the fallback branch.
A.4 Old KV cache
cache.rs provided an architecture-agnostic per-layer cache: k/v were
pre-allocated Vec<f32> of max_size × dim, size tracked the current
sequence length. store_multi wrote K/V for many positions at once (prefill);
decode wrote one. The GPU maintained its own buffers and sync_kv_to_cpu
copied them back only on the CPU-fallback path. MINFER_CACHE_TYPE=f16
selected an f16 GPU cache (opt-in).
A.5 Old "Adding a New Architecture"
- Create
src/models/<name>/withmod.rs,forward.rs,loader.rs. - Add a
matchbranch insrc/models/mod.rs::load_model(). - Define
HParams(includingn_kv_embd) andLayerWeightsinloader.rs. - Implement the per-layer forward pass in
forward.rsusingkernel::,vec_ops::,cache::. - Implement
ModelDefinmod.rs. - Add a chat template format in
template.rsif needed.
Decisions governing this document
This page is the current contract; the decisions behind it are frozen in the ADR corpus:
- ADR-0001 — Inference runs through one declarative compute graph
- ADR-0012 — Device is the first axis, the layer the second — and no premature
common - ADR-0002 — Topology is a function of
GraphParamsalone, sopositionscannot be structure - ADR-0007 — No ML frameworks: every operator is hand-written
- ADR-0008 — GPU safety: bounded waits, no early return past a barrier, runtime device limits
- ADR-0009 — A failure is an error, never a silent fallback
- ADR-0013 — The CPU quantizes activations to Q8_0; a device reads f32