minfer Architecture

A pure-Rust LLM inference engine written from scratch, inspired by llama.cpp, with zero ML framework dependencies. Inference runs through a declarative compute graph (builder → scheduler → per-backend kernels), modeled on llama.cpp's ggml_cgraph + backend scheduler. This document describes the overall design: module responsibilities, the compute graph pipeline, the CPU / Metal backend layering, quantization, adding a new model architecture, and adding a new backend. The pre-graph imperative forward is preserved at the end as an appendix (Appendix A).


1. Design Principles

  1. No ML framework — attention, RMSNorm, RoPE, SiLU, softmax are all handwritten. Only 5 external crates: rand, regex, half, serde/serde_json, minijinja.
  2. Declarative compute graph, not an imperative loop — the forward pass is built as a pure ComputeGraph (no side effects at build time), then assigned to backends, fused, allocated, and executed by the scheduler. This mirrors llama.cpp (ggml_cgraph + ggml_backend_sched) and enables graph reuse, per-op backend assignment, DOT export, and a clean path to new backends. Design and implementation record: docs/COMPUTE-GRAPH-DESIGN.md.
  3. Bytes-in / bytes-out tensors — weight tensors are raw &[u8]; SIMD dot-product kernels (AVX2 / Metal shaders) operate on byte slices matching the exact GGML quantized block layout (repr(C) in block.rs).
  4. Activations stay f32 — CPU matmuls quantize activations to Q8_0 on-the-fly; the Metal backend reads f32 activations directly for all weight types (matching llama.cpp's Metal backend).
  5. Backend assignment is a build-time decision, never a silent mid-execution fallback — supports_op decides which ops run where; cross-backend transfers happen at split boundaries; kernel-invariant violations abort (gpu_abort / Err), they never silently fall back to CPU.

2. Module Map

ModuleResponsibility
main.rsCLI, GGUF load, chat template, prefill → autoregressive generation loop, timing
graph/Compute graph core — IR, builder, scheduler, backends, reuse cache (see §4 table)
gguf.rsGGUF v3 parser (metadata KV + tensor table + data blob), multi-part (split) support, ggml_pad alignment
block.rs20+ quantized block types as repr(C) structs + fp16 conversions, matching ggml-common.h
quants.rs + src/quants/*.rsAVX2+FMA / NEON+SDOT dot-product kernels (Q4_0/Q4_1/Q5_0/Q5_1/Q8_0/K-quants × Q8_0/Q8_K) + f32→Q8_0/Q8_K quantization, scalar fallback. The module file is the decider; the parts are src/quants/{dot_q4_0,dot_q4_1,dot_q5,dot_q8_0,kquant,quantize_q8_0,quantize_q8_k,avx2,neon}.rs (#264, the source layout plan)
kernel.rs + src/kernel/*.rsQuantized matmul dispatch (Q4_0/Q4_1/Q5_0/Q5_1/Q4_K/Q5_K/Q6_K/Q8_0) over activations, CPU scalar fallback, the shared worker pool, and the shared embed_tokens row getter
vec_ops.rs + src/vec_ops/*.rsSIMD vector ops: RMSNorm, RoPE (Qwen2/Llama styles), softmax, SiLU, add/scale/mul, plus the f16/bf16 weight-row dots
tensor.rs4D Tensor (type/shape/strides/Vec<u8> data), ggml-compatible strides & byte sizing
sampler.rsRepeat/frequency/presence penalties → top-k → top-p → temperature, seeded StdRng
tokenizer.rsSelf-contained BPE tokenizer, loaded from GGUF metadata (no tiktoken)
template.rsGGUF chat_template rendering via minijinja + a Python-str-method hook (F7/#50), so Qwen3's think-block template renders; a template that cannot be rendered is a loud refusal, never a silent ChatML fallback (docs/CHAT-TEMPLATE-AND-TOKENIZER-DESIGN.md)
models/Architecture implementations. mod.rs has the ModelDef trait + factory dispatch
models/qwen2/Qwen2/Qwen2.5: mod.rs (model struct + trait impl), graph.rs (Qwen2Graph::build/forward), loader.rs (GGUF weights + hparams)
models/qwen3/Qwen3 (dense): same triple — decoupled head dim + per-head Q/K norm (qk_norm); its ChatML+<think> template renders since F7
conversation.rsMulti-turn session state (append-only KV) behind --cnv
server/OpenAI-compatible HTTP server (serve): axum + tokio, multi-slot, /v1/chat/completions streaming; viz.rs serves the viz page
src/metal/ + src/metal/kernels/*.metalApple MPS (Metal) backend: L1 runtime (MpsState/MetalDevice), L2 command-buffer encoding, L3 shaders (the legacy whole-layer layer_gpu is retained for tests). The split into src/metal/{runtime,encode,ops,policy}.rs + src/metal/kernels/ is the source layout plan (#265)
cuda.rs + src/cuda/kernels/*.cuNVIDIA CUDA device layer + launch layer + kernels (feature-gated --features cuda); executed through the graph via graph/cuda_backend.rs. The split is the source layout plan (#262, #263)
device_tier.rscc-keyed device tier table + selector (measured GB10 row, llama.cpp-adopted consumer rows, GENERIC fallback); resolved once at init, feeds the MMQ gate, smem feasibility and plane-VRAM budget checks (docs 105–106)
download/Hugging Face Hub + Ollama download, cached-name resolution, resume support
dump.rsPer-layer hidden-state debug dump (gated by --features debug_dump)
bench.rs / live.rs / trace.rs / spec_verify.rsllama-bench-style bench, live viz inference host, per-node trace export (MINFER_TRACE), D5 verify-step gate bench

The backend layers — runtime, launch, kernels, executors

Each device backend is organised along four layers, and only the last one is polymorphic:

layerCUDAMetalCPU
L1 device/runtimesrc/cuda/methods/{init,weights,stream,buffers,copy,events,capture}.rssrc/metal/runtime.rs— (std threads)
L2 launch/dispatchsrc/cuda/methods/<family>.rs + the extern "C" declarations each family owns (under src/cuda/methods.rs)src/metal/encode.rs + ops.rssrc/kernel/*.rs
L3 kernel sourcessrc/cuda/kernels/*.cu + common.cuhsrc/metal/kernels/*.metalsrc/quants/*.rs, src/vec_ops/*.rs
L4 graph executorsrc/graph/cuda_backend.rssrc/graph/metal_backend.rssrc/graph/cpu_backend.rs

(The CUDA paths in that table have landed (#262, #263); the Metal ones are still the target of Step 4 (#265, Mac-local). docs/SOURCE-LAYOUT-PLAN.md §8 has the file-by-file tree and each step updates this table as it lands.)

Two rules follow, and they are why the crate keeps one flat interface instead of a directory per layer:

  1. Device is the first axis, the layer is the second. Backend + registry.rs (src/graph/backend.rs) are the one device seam — llama.cpp keeps ggml-backend*.cpp/h flat beside its per-device directories for the same reason. L2 cannot leave L1: the CUDA launchers are inherent methods of CudaState, the Metal ones methods of MpsCommandBuffer; splitting them out would be a type refactor, not a file move.
  2. A common needs a second real implementation. A shared abstraction is added only when at least two backends implement it and at least two callers use it with the same semantics. The one candidate today is allocplan::DeviceMemory (CUDA answers it; the Metal half is #53). CPU's quants.rs/vec_ops.rs are deliberately not a device-private layer: they are the crate's numeric kernel library, consumed by graph/kvformat.rs and graph/cuda_backend.rs.

src/graph/ — the compute graph core

FileRole
mod.rsComputeGraph (topo-validated node list + inputs/outputs), CNode, DType, BufRef; re-exports the Backend handle from registry.rs
registry.rsThe backend registry (F4): the fixed id space (Backend::CPU/METAL/CUDA), the name surface and its three startup refusals, each entry's priority + capability record + pool/host-I/O hooks, and the --backend / MINFER_BACKENDS filter. Contract: docs/BACKEND-REGISTRY-DESIGN.md
ops.rsOp enum (full payload PartialEq), NodeMeta, AttnMode, FusedOp
builder.rsGraphBuilder — declarative construction (embedding/rms_norm/matmul/rope/attn/kvcache/…)
alloc.rsPer-backend liveness allocator + persistent per-layer KV regions + KvProvider
backend.rsBackend trait + KvProvider
cpu_backend.rsCPU execution (wraps kernel.rs + vec_ops.rs)
metal_backend.rsMetal execution (per-op MPS kernels; cfg(target_os = "macos"))
cuda_backend.rsCUDA execution (feature-gated): int8 MMQ prefill + MMVQ decode (default-on), split-KV attention, CUDA Graph capture/replay
scheduler.rsassign → split → execute (+ cross-backend copies at split boundaries)
fusion.rsPattern-matching fusion (SwiGLU), gated by backend supports_fused
cache.rsGraphCache — params-only deterministic graph reuse
params.rsGraphParams/CParams/GraphType — the reuse identity
dot.rsGraphviz DOT export (--dump-graph)
json.rsGraph JSON export for the viz page (--dump-graph-json, MINFER_TRACE)

3. Inference Pipeline

The top-level flow lives in main.rs. The whole engine is a single-pass prefill followed by an autoregressive decode loop; both call ModelDef::forward, which routes through the compute graph. Every forward is build → assign → fuse → alloc → execute (the subgraph below), but the built graph is cached per GraphParams and reused — only the first call with new params pays the build/assign/alloc cost.

flowchart TD
    A["CLI args: model, prompt, flags"] --> B{"resolve model<br/>download::resolve"}
    B -->|local path| C["load GGUF v3<br/>single or split parts (mmap)"]
    B -->|"hf:… / ollama:…"| D["auto-download → path"]
    B -->|cached name| C
    C --> E["parse metadata KV + tensor table"]
    E --> F["init GPU backend<br/>MPS / CUDA"]
    F --> G["load model<br/>dispatch on general.architecture"]
    G --> MODE{"mode?"}
    MODE -->|"--cnv"| CV["conversation REPL<br/>(run_conversation)"]
    MODE -->|"serve"| SV["OpenAI-compatible HTTP server<br/>multi-slot, streaming"]
    MODE -->|"viz"| VZ["viz server (page + live SSE)"]
    MODE -->|single shot| H["load BPE tokenizer from GGUF"]
    H --> J{"no-template?"}
    J -->|no| K["render chat template<br/>GGUF chat_template via minijinja<br/>(unrenderable → loud refusal;<br/>no template at all → ChatML)"]
    J -->|yes| L["raw prompt"]
    K --> M["tokenize prompt<br/>n_ctx = max(--n-ctx, prompt len)<br/>(sizes the KV regions once)"]
    L --> M
    M --> N["PREFILL<br/>graph forward, all prompt tokens at once"]
    N --> O["last-token logits (n_out = 1)"]
    O --> P{"DECODE loop<br/>while generated < n_predict"}
    P --> Q["sample next token<br/>penalties → top-k → top-p → temp"]
    Q --> R{"stop token or<br/>stop string match?"}
    R -->|yes| S["done"]
    R -->|no| T["append token, decode+print"]
    T --> U["graph forward, single token<br/>KV persists in the allocator"]
    U --> P

    subgraph GRAPH["every forward: build → assign → fuse → alloc → execute"]
        G1["GraphBuilder<br/>build_graph (pure IR)"] --> G2["assign backends<br/>priority Metal → CUDA → CPU"]
        G2 --> G3["fuse<br/>SwiGLU + decode fusions (gated)"]
        G3 --> G4["alloc<br/>liveness + persistent KV"]
        G4 --> G5["execute<br/>per split, cross-backend copies"]
    end
    N -.->|"GraphCache: params-only reuse"| GRAPH
    U -.->|"GraphCache: params-only reuse"| GRAPH

Generation parameters (defaults match llama.cpp): temp=0.8, top_k=40, top_p=0.95, repeat_penalty=1.1 (last 64 tokens), frequency_penalty=0.0, presence_penalty=0.0, seed=42, n_ctx=4096, n_predict=512. Sampling applies the three penalties in one pass, then top-k → top-p → temperature.

Timing is dual-caliber: Prefill: = prompt tokens / prefill wall time; Generated: = generated tokens / decode wall time (pure decode, matches llama-bench "Generation" caliber); Total: = blended.


4. Compute Graph Architecture

Inference = build a ComputeGraph (pure, side-effect free) → assign backends → fuse → allocate → execute. The graph is built once per distinct GraphParams and reused (decode steps reuse the same graph; the model's forward() routes through Qwen2Graph::forward).

4.1 The IR

  • ComputeGraph = topologically ordered nodes (builder appends sources before consumers), inputs, outputs, uid (for CUDA-Graph-style reuse).
  • CNode = op + src dependencies + out_shape/out_dtype + backend (None until assigned) + meta (weight names, rope/attn params).
  • Op carries full payloads (RmsNorm{eps}, MatMul{transpose_b}, KvcacheStore{layer} …) so graphs are structurally comparable.

4.2 Builder (GraphBuilder)

Per-architecture code calls builder methods (mirroring llama.cpp's llm_graph_context): embedding, rms_norm, qk_norm (Qwen3 per-head Q/K norm), matmul, get_rows, rope, silu, add, mul, swiglu, softmax, attn, kvcache_store/kvcache_load, plus the decode-fusion constructors fused_qkv/qkv_bias_rope_store/fused_ffn. Building is pure — no computation happens at build time.

4.3 Scheduler pipeline

assign_backends → fuse → alloc_graph → execute
  1. assign_backends — capability-driven: each node gets the highest-priority backend whose supports_op returns true (priority Metal → CUDA → CPU). Weight registration decides GPU feasibility.
  2. fuse — pattern matching (Mul(Silu(X),Y) → SwiGLU) gated per backend by supports_fused (no double-fusion with hand-written kernels). The plan's second rule (RoPE(Add(X,B)) → FusedBiasRope) was removed: no backend ever claimed the capability, and the bias+rope work ships as the builder's decode nodes (FusedQKV/QkvBiasRopeStore, whose attn_bias_rope_store kernel subsumes the bias+rope part). BatchMatMul stays deferred (single-output IR limitation, COMPUTE-GRAPH-DESIGN.md §5.4).
  3. alloc_graph — per-backend liveness allocator: buffers shared between nodes whose live ranges don't overlap; persistent per-layer KV regions survive rebuilds; in-place ops alias their input buffer (see §4.5).
  4. execute — per split (contiguous same-backend runs): sync the previous backend, copy split inputs across backends (copy_across, a host round trip via shared memory), run the nodes, then a final sync. Metal batches one MpsCommandBuffer per split, submitted at synchronize().

4.4 Reuse (GraphCache)

Params-only deterministic reuse (llama.cpp allow_reuse invariant): GraphParams = n_tokens / n_out (tail rows) / gtype / cparams (n_ctx, flash_attn, gpu, fuse_qkv, fuse_ffn, explicit_span) / weights_version deterministically determines the topology — equal params ⇒ identical graph. n_past is deliberately absent (it is execution data). CParams.gpu records backend participation so a backend toggle forces a rebuild. GraphCache owns the allocator (so the KV regions persist across rebuilds, e.g. the prefill→decode transition) and try_reuse compares params only; debug builds assert structural consistency (Op: PartialEq).

4.5 Core invariants (must not be violated)

  1. KV positions are data, not structure. KvcacheStore/Load carry only the layer index; write positions come from the positions input node. Topology never depends on n_past.
  2. Each layer owns TWO persistent KV regions (K and V), resolved by kv_pair(layer) (KvProvider). The store node's output buffer is the K region; backends write the V sibling via kv_pair.
  3. GGUF weight layout: metadata [in, out] (ne[0] fastest), memory [out][in] row-major → matmul od = shape[1], id = shape[0]. Activations: shape metadata [d, nt, 1, 1], memory token-major [nt][d]. I32 inputs are stored as f32::from_bits bit patterns (fill_input_i32).
  4. In-place ops (Silu, RoPE) alias their input buffer — the allocator maps the output to the input's BufRef (only when the input's sole consumer is this op AND it is on the same backend). Never host-copy a GPU-pending buffer: a host copy_in of a producer that is encoded but not submitted reads stale data (the Phase-3 KV-corruption bug). Cross-backend in-place inputs get a fresh buffer (the producer completed before the split boundary, so the copy is safe there).
  5. Execution follows build order (a valid topological order by construction) — guarantees a KV store executes before the attention that reads it. Nodes with no allocated buffer (dead, e.g. fusion orphans) are skipped.
  6. CPU vs GPU activation paths differ numerically: CPU matmuls are Q8_0×Q8_0 (activation-quantized), the GPU backends read f32 activations for all weight types (CUDA prefill additionally offers the default-on int8 MMQ path). Compare each path against its own reference, not against the other.

4.6 Per-layer computation (Qwen2) — as built by graph.rs

flowchart LR
    A["token_embd lookup (GetRows)"] --> B["hidden"]
    B --> C["RMSNorm attn_norm"]
    C --> D["WQ / WK / WV matmuls + bias"]
    D --> E["RoPE on Q and K (in-place alias)"]
    E --> F["kvcache_store: K/V → layer regions"]
    F --> G["GQA attention<br/>Q·K^T → softmax → ·V"]
    G --> H["WO matmul + bias"]
    H --> I["+ residual → hidden"]
    I --> J["RMSNorm ffn_norm"]
    J --> K["FFN gate + up matmuls"]
    K --> L["SiLU(gate) × up (fused SwiGLU)"]
    L --> M["FFN down matmul"]
    M --> N["+ residual → hidden"]
    N --> O["next layer / output_norm"]

GQA: each query head h maps to KV head hk = h / gqa. The KV head dimension is independent (n_kv_embd read from the K weight's ne[1]), so Qwen2.5-0.5B (n_embd=896, n_head=14, hd=64, n_kv_embd=128) strides correctly.

4.7 Prefill vs decode

Phasent (tokens)Notes
Prefill> 1graph type Prefill; KV store writes all positions, attention reads the full written prefix
Decode1graph type Decode; same topology as prefill modulo nt → the graph is rebuilt once (KV persists), then reused for every subsequent token

The graph path keeps K/V on the executing backend (attention and KV are on the same backend by construction), so there is no per-token KV drain.


5. Backend Layering

5.1 The Backend trait (graph/backend.rs)

Backend (the thing the IR stores on a node) is not this trait: since F4 it is a Copy handle — an id into src/graph/registry.rs — and the trait below is implemented by the pools (CPU, Metal, CUDA) and held by the registry entry. The name, id, priority and capability record live on the handle/entry, not on the trait (Backend::name() is registry.rs:105):

#![allow(unused)]
fn main() {
pub trait Backend: Send + Sync {
    fn kv_pair(&self, layer: usize) -> Option<(usize, usize)>;   // persistent per-layer K/V regions
    fn supports_op(&self, op: &Op, dtype: DType) -> bool;
    fn supports_fused(&self, fused: &FusedOp) -> bool;
    fn supports_attn_span(&self) -> bool { false }       // E1: explicit attention windows (CPU/CUDA/Metal)
    fn weights_bytes(&self) -> usize { 0 }               // E4: bytes the memory budget is charged for
    fn alloc_buffer(&mut self, size: usize) -> usize;    // backend's own pool
    fn free_buffer(&mut self, id: usize);
    fn pool_len(&self) -> usize;
    fn alloc_fresh(&mut self, size: usize) -> usize;     // bypasses the recycle free list (split-boundary staging)
    fn execute_node(&mut self, node: &CNode, in_bufs: &[usize],
                    out_buf: usize, kv_pair: Option<(usize, usize)>) -> Result<(), String>;
    fn copy_cells(&mut self, dst: BufRef, src: BufRef, dst_row: usize, src_row: usize,
                  rows: usize, elems_per_cell: usize) -> Result<(), String>;   // C3 compaction / C2 shift
    fn read_host(&self, id: usize) -> Option<&[f32]>;
    fn write_host(&mut self, id: usize, data: &[f32]) -> Result<(), String>;
    fn write_host_window(&mut self, id: usize, offset: usize, data: &[f32]) -> Result<(), String>;
    fn synchronize(&mut self);
    fn retire(&mut self) { self.synchronize(); }         // release device objects for a dropped pool
    // CUDA only: try to replay a captured graph for (uid, range); capture is
    // gated to decode-shaped graphs. Default impl returns false.
    #[cfg(feature = "cuda")]
    fn graph_replay(&mut self, uid: u64, range: (usize, usize), nt_hint: Option<usize>) -> bool;
}
}
  • CPU (cpu_backend.rs): Vec<f32> pool; executes via kernel.rs + vec_ops.rs; F32 weights use a plain f32 matmul, quantized weights use cpu_quant_matmul_f32 (Q8_0-activation path).
  • Metal (metal_backend.rs, macOS): shared-memory MTLBuffer pool; per-op dispatch to MpsState's kernels (rms_norm, quant_matmul_f32_on_gpu_buf, rope, silu, add, mul, swiglu, embed_tokens_gpu, store_kv, gqa_attn_f32); one command buffer per split. Weights resolve by name from MpsState's registry (weight_buf(name) -> (buffer, offset)).
  • CUDA (cuda_backend.rs, feature-gated --features cuda): wraps the cuda.rs device layer — per-op dispatch with int8 MMQ prefill + MMVQ decode (default-on; MINFER_MMQ=0 reverts; the MMQ gate resolves from the device_tier.rs device-tier table), split-KV attention, and CUDA Graph capture/replay keyed on the graph uid (graph_replay, decode-shaped only). Implementation record: docs/CUDA-BACKEND-DESIGN.md; per-step optimization history in docs/CUDA_OPTIMIZATION.md.

The allocator owns every backend pool (single source of truth); the scheduler orchestrates assignment, cross-backend copies, and sync.

5.2 Selection rules

Assignment asks each node in the registry's priority order — Metal 300, CUDA 200, CPU 100 (src/graph/registry.rs:72-75) — for the first backend that has a pool, is not fenced off, and reports the op supported; the CPU is tried last and is always allowed, so the answer is total (GraphAllocator::supports_for, src/graph/alloc.rs:403). The fence is --backend <name> / MINFER_BACKENDS=<csv>, and an unknown name, a name this build does not contain and a name this machine cannot use are three distinct startup refusals. The two orders, the name surface and those refusals are docs/BACKEND-REGISTRY-DESIGN.md's contract; the user-facing flags are docs/USAGE.md.

  • Metal: all graph weights must be GPU-registered (Qwen2Graph::weights_on_gpu mirrors the old per-layer check). MINFER_DISABLE_MPS=1 forces CPU.
  • CUDA: feature-gated (--features cuda), and every graph weight must be GPU-registered with a device present. A node whose block the E5 offload plan left on the CPU never reaches the device (supports_for answers CPU for it, so a partially offloaded model cannot run a block whose weights were never registered there). MINFER_DISABLE_CUDA=1 is presence-checked and forces the CPU.
  • CPU: always available; AVX2 dispatch via is_x86_feature_detected!("avx2") on x86, NEON+SDOT on aarch64, scalar fallback elsewhere (MINFER_NO_NEON=1 forces scalar; the K-quant dots also have an AVX-512/VNNI path since #56, gated by MINFER_NO_AVX512=1 and MINFER_NO_AVX2=1).

5.3 GPU safety

All Metal submits wait bounded (10 s) and check status; no early return past a threadgroup_barrier; device limits queried at runtime, never hardcoded; guard failures gpu_abort with actual values. In the graph architecture, kernel-invariant violations return Err from execute_node and must NOT be treated as a silent CPU fallback — backend assignment at build time decides where ops run; only genuine support limitations (e.g. Raw weights) select the CPU backend. See docs/GPU_SAFETY.md.


6. Quantization & Tensor Layout

  • Weight layout (block.rs): repr(C) blocks matching ggml-common.h. Q4_0/Q4_1/Q5_0/Q5_1/Q8_0 = 32-value blocks; Q4_K/Q5_K/Q6_K = 256-value super-blocks. Supported: Q4_0, Q4_1, Q8_0, Q4_K, Q6_K, Q5_0, Q5_1, Q5_K (CPU + Metal), F32/F16 norms & biases.
  • Activation quantization: CPU matmuls quantize f32 activations to Q8_0 on-the-fly (quants.rs); Metal reads f32 directly.
  • GGUF parsing (gguf.rs): ggml_pad(x, n) = (x + n - 1) & !(n - 1) alignment; tensor strides computed from type block size exactly like ggml; split multi-part models are merged into one tensor index.
flowchart LR
    A["GGUF file"] --> B["metadata KV<br/>hparams + tokenizer + template"]
    A --> C["tensor table<br/>name / type / shape / offset"]
    C --> D["quantized data blob"]
    D --> E["Tensor: type, shape, strides, Vec&lt;u8&gt;"]
    B --> F["HParams"]
    B --> G["Tokenizer"]
    B --> H["Chat template"]

7. KV Cache

The graph path owns the KV cache: two persistent regions per layer (K and V), sized n_kv_embd × n_ctx, allocated in the backend pool the layer runs on. kv_pair(layer) resolves them; the KV store node writes K/V at the positions carried by the positions input; attention reads the written prefix. The regions live inside the GraphCache's allocator and survive graph rebuilds (the prefill→decode transition). MINFER_CACHE_TYPE=f16 selects an f16 GPU cache where the kernels support it. The pre-graph cache.rs KVCache type — and the vestigial &mut KVCache argument of ModelDef::forward — was deleted in #252; these regions are the only KV store.


8. Adding a New Architecture

  1. Create src/models/<name>/ with mod.rs, graph.rs, loader.rs.
  2. Add a match branch in src/models/mod.rs::load_model() for the new general.architecture value.
  3. In loader.rs: define HParams (including n_kv_embd) and LayerWeights, parse them from GGUF metadata (both qwen2.* and llama.* prefixes are accepted).
  4. In graph.rs: implement build_graph(&self, params: &GraphParams) -> ComputeGraph deterministically in params (the reuse invariant), using GraphBuilder — mirror llama.cpp's llm_graph_context builder methods.
  5. In mod.rs: implement ModelDef (forward, build_graph, forward_graph, forward_graph_cached, as_any, special_tokens, n_layer/n_head_kv/n_embd_head/n_kv_embd/n_vocab, rope_style). models/qwen3/ is the worked example of a second architecture (decoupled head dim + per-head Q/K norm).
  6. If needed, add a chat template format in template.rs.

Architectures that share Qwen2's tensor naming convention (LLaMA, Mistral, Phi) are the easiest ports. The RopeStyle enum defines both Qwen2 (non-interleaved) and Llama (interleaved) pairings, but only the non-interleaved form is wired up: both loaders hard-code it (models/qwen2/loader.rs:155, models/qwen3/loader.rs:173) and the CUDA backend refuses the interleaved style (graph/cuda_backend.rs:1522). A family that needs interleaved RoPE requires a loader change plus a CUDA kernel — see docs/MODEL-SUPPORT-ROADMAP.md, "Cost model: what a port actually costs".


9. Model Download

download/mod.rs resolves hf:<repo>[:<quant>] and ollama:<model>[:<tag>] URIs, downloads via curl with size-checked resume, quant-matches single or split files case-insensitively, and stores them under ~/.cache/minfer/models. Cached filenames can be used directly as the model argument.


TopicLocation
Compute graph design + rewrite plan + implementation record (per-phase commits)docs/COMPUTE-GRAPH-DESIGN.md
llama.cpp compute-graph design analysis (ggml_cgraph / scheduler / reuse)docs/LLAMA-COMPUTE-GRAPH.md
Metal backend optimizations / gap analysis (primary tracking)docs/METAL_OPTIMIZATIONS.md
GPU safety conventions + auditdocs/GPU_SAFETY.md
CPU backend optimizationsdocs/CPU_OPTIMIZATIONS.md
CUDA optimization history + per-step recordsdocs/CUDA_OPTIMIZATION.md (+ docs/cuda_optimization_steps/)
Debug dump formatdocs/debug-dump.md
Historical bugs / debugging notesdocs/BUG-6-KV-CACHE-INDEXING.md, docs/QWEN2.5-*, docs/DEBUGGING-*

Appendix A — Legacy Imperative Architecture (removed in Phase 6)

Historical reference. The imperative per-layer forward (models/qwen2/forward.rs) was the engine's core until the compute graph replaced it (Phase 6); it is preserved here verbatim in structure so old notes, benchmarks, and kernel analyses remain interpretable. Do not treat this as the current design.

A.1 Design stance (then)

The engine ran a direct per-layer forward loop instead of a compute graph: "simpler, easier to trace, and the whole layer can be fused onto the GPU." The GPU fallback was per-layer and safe: a layer that could not run on the GPU (e.g. unsupported weight type) submitted partial GPU work, downloaded the hidden state, and continued on the CPU.

A.2 The old forward pass

ModelDef::forward(tokens, positions, kv) was implemented in models/qwen2/forward.rs as an imperative loop:

flowchart LR
    A["token_embd lookup"] --> B["hidden"]
    B --> C["RMSNorm attn_norm"]
    C --> D["WQ / WK / WV matmuls + bias"]
    D --> E["RoPE on Q and K"]
    E --> F["store K/V into KV cache"]
    F --> G["GQA attention<br/>Q·K^T → softmax → ·V"]
    G --> H["WO matmul + bias"]
    H --> I["+ residual → hidden"]
    I --> J["RMSNorm ffn_norm"]
    J --> K["FFN gate + up matmuls"]
    K --> L["SiLU(gate) × up"]
    L --> M["FFN down matmul"]
    M --> N["+ residual → hidden"]
    N --> O["next layer / output_norm"]

Decode-time optimizations: fused QKV (nt==1) via a concatenated blk.{il}.attn_qkv weight, fused bias+rope+store (attn_bias_rope_store), fused SwiGLU kernel, and the last layer computed only the tail n_out rows (an inp_out_ids-style partial-row optimization). The n_out tail-row optimization was subsequently carried into the graph path as the G3 work: the builder inserts GetRows(wo, tail_ids) (and the same for the residual) before the last layer's FFN, so the last FFN, the final norm and lm_head all run on n_out rows only (models/qwen2/graph.rs:64, :228-231; GraphParams.n_out is part of the reuse identity).

A.3 Old backend layering & fallback

flowchart TD
    A["forward nt tokens"] --> B{"embedding on GPU?"}
    B -->|yes| C["GPU embed lookup → buf_hidden"]
    B -->|no| D["CPU embed → upload hidden"]
    D --> E["upload positions"]
    C --> E
    E --> F{"per-layer: layer_gpu ok?"}
    F -->|"yes, all layers"| G["output_norm_gpu<br/>on GPU"]
    F -->|"no at layer i"| H["submit partial GPU work<br/>download hidden, sync KV to CPU"]
    H --> I["CPU loop from layer i"]
    G -->|"output on GPU"| J["download logits → return"]
    G -->|"output fell back"| I
    I --> K["output_norm + output matmul on CPU"]
    K --> L["return logits"]

Selection rules (then):

  • Metal: layer 0 must have all 7 weight matrices + norms registered on the GPU. Within a layer all 7 matrices must be all Q4 group (Q4_0/Q4_1) or all QK group (Q4_K/Q5_0/Q6_K); Q5_1/Q5_K use the f32 path and are exempt. MINFER_DISABLE_MPS=1 forces CPU.
  • CUDA (--features cuda): requires every layer's 7 matrices to be all Q4_0/Q4_1 or all Q4_K/Q6_K. Decode replays a captured CUDA Graph.
  • CPU: always available; AVX2 dispatch via is_x86_feature_detected!("avx2"), scalar fallback elsewhere (plus the AVX-512/VNNI K-quant dots, #56).

The GPU path skipped the per-token CPU→GPU KV drain (no sync_kv_to_cpu) because GPU-layer failure is deterministic by weight type — the sync only happened in the fallback branch.

A.4 Old KV cache

cache.rs provided an architecture-agnostic per-layer cache: k/v were pre-allocated Vec<f32> of max_size × dim, size tracked the current sequence length. store_multi wrote K/V for many positions at once (prefill); decode wrote one. The GPU maintained its own buffers and sync_kv_to_cpu copied them back only on the CPU-fallback path. MINFER_CACHE_TYPE=f16 selected an f16 GPU cache (opt-in).

A.5 Old "Adding a New Architecture"

  1. Create src/models/<name>/ with mod.rs, forward.rs, loader.rs.
  2. Add a match branch in src/models/mod.rs::load_model().
  3. Define HParams (including n_kv_embd) and LayerWeights in loader.rs.
  4. Implement the per-layer forward pass in forward.rs using kernel::, vec_ops::, cache::.
  5. Implement ModelDef in mod.rs.
  6. Add a chat template format in template.rs if needed.

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0001 — Inference runs through one declarative compute graph
  • ADR-0012 — Device is the first axis, the layer the second — and no premature common
  • ADR-0002 — Topology is a function of GraphParams alone, so positions cannot be structure
  • ADR-0007 — No ML frameworks: every operator is hand-written
  • ADR-0008 — GPU safety: bounded waits, no early return past a barrier, runtime device limits
  • ADR-0009 — A failure is an error, never a silent fallback
  • ADR-0013 — The CPU quantizes activations to Q8_0; a device reads f32