Usage
Examples use the built binary (cargo build --release first — see the
Build section of the project README); cargo run --release -- … works identically.
./target/release/minfer <model> [prompt] [OPTIONS]
<model> can be a local path, a download URI, or a cached model name:
| Format | Example |
|---|---|
| Local file | ~/models/qwen2.gguf, ./model.gguf, /abs/model.gguf |
| Hugging Face | hf:Qwen/Qwen2-0.5B-GGUF:qwen2-0.5b-q4_0.gguf (auto-download) |
| Ollama | ollama:qwen2.5:0.5b (pull) |
| Cached model name | qwen2.5-0.5b-instruct-q4_0 (resolved from ~/.cache/minfer/models, see list) |
If prompt is omitted, reads from stdin. Run minfer --help for the full
option list; the subcommands are:
| Command | Purpose |
|---|---|
<model> [prompt] [OPTIONS] | single-shot generation |
serve [--port N] [--n-ctx N] [--n-slots N] <model> | OpenAI-compatible HTTP server |
info <model> | print GGUF metadata + key tensors |
download hf <repo> [quant] / download ollama <model>[:tag] | fetch models |
list | list locally cached models |
viz [--port N] <model> | self-contained viz server (default port 8081) |
bench [-p N] [-n N] [-r N] [-o md|csv|json] <model> | perf test: pp<P> prefill / tg<T> decode, mean ± stddev over reps |
specverify [-p N] [-r N] [-o json|md] <model> | D5-1a gate bench: batched verify cost C_T(nt) + per-token amortization at deep KV |
Sampling and runtime options:
--temp <T>— sampling temperature (default 0.8;--greedy= greedy decoding)--top-k <K>/--top-p <P>— top-K / nucleus sampling (defaults 40 / 0.95)--repeat-penalty <N>— repeat penalty (default 1.1; 1.0 = off), plus--frequency-penalty/--presence-penalty- F3 sampler set (#48) — every default below
leaves the pre-F3 chain unchanged:
--min-p <P>— drop tokens whose probability is belowP * max(0 = off; 1.0 = argmax only)--typical <P>— locally typical sampling (1.0 = off; 0 keeps the single most typical token)--xtc-probability <P>/--xtc-threshold <T>— XTC: with probabilityP, exclude the top choices whose probability is at leastT(T<= 0.5; 0 = off)--dry-multiplier <N>— DRY (Don't Repeat Yourself) penalty strength (0 = off), with--dry-base <N>(default 1.75),--dry-allowed-length <N>(default 2),--dry-penalty-last-n <N>(default 64) and--dry-sequence-breakers <L>— restart sequences as token ids, e.g.--dry-sequence-breakers 198;13,2(;between sequences,,between ids)--mirostat <0|1|2>— mirostat off / v1 / v2, with--mirostat-tau <N>(target surprise in bits, default 5.0),--mirostat-eta <N>(learning rate, default 0.1) and--mirostat-m <N>(v1 estimator window, default 100). In mirostat mode the temperature is ignored (mirostat'smutruncation subsumes it);--temp 0still means greedy. Mirostat cannot be combined with--spec-draft.--logit-bias <L>— add to raw logits:ID:BIASpairs separated by,, repeatable, e.g.--logit-bias 15043:-2.0,198:1.5. A token id outside the vocabulary, or a bias outside[-100, 100], is refused at startup. A nonsensical value for any of these is refused at startup (exit 1), never silently ignored.
- F2 constrained decoding (#47) — mutually
exclusive; an unsupported construct is refused at startup (exit 1) with the offending token:
--grammar <FILE>/--grammar-str <GBNF>— constrain sampling to a GBNF grammar--json-schema <FILE>/--json-schema-str <JSON>— constrain sampling to a JSON Schema (compiled to GBNF internally) The grammar is compiled once against the loaded vocabulary; the mask is applied inside the sampler, after the penalties/DRY and before the greedy shortcut, so--greedyrespects it.--spec-draftcannot be combined with a grammar (a verify round samples several rows from one automaton state). The accepted GBNF and JSON-Schema subsets, and every construct that is refused, are catalogued in GRAMMAR-DESIGN.md; the honest narrowing to know about is that object properties are accepted in declaration order andoneOfis compiled asanyOf. Example:
minfer model.gguf "Give me a person" --json-schema-str \ '{"type":"object","properties":{"name":{"type":"string"},"age":{"type":"integer","minimum":0, "maximum":150}},"required":["name","age"],"additionalProperties":false}' # -> {"age": 25, "name": "John Doe"} --stop <STR>— stop generation at this string (repeatable)-n, --n-predict <N>— max tokens to generate (default 512)--seed <N>— RNG seed for sampling- Speculative decoding (D5-R, ADR-0017) —
a draft model proposes, the target verifies the proposals in one batched forward:
--spec-draft <model>— the draft model (any supported GGUF)--spec-draft-n <N>— drafted tokens per round (default 2, so the verify batch isN + 1rows)--spec-draft-adaptive— pickdper round from the per-depth acceptance and cost EWMA, capped at 8 unless--spec-draft-nsets a smaller cap--spec-draft-nand--spec-draft-adaptiveare no-ops without--spec-draft. Speculative decoding refuses mirostat and a grammar (both named above), because a verify round samples several rows from one shared RNG, or from one automaton state.
--n-ctx <N>— sizes the KV cache (clamped to the model's max context)-t, --threads <N>— CPU worker threads--gpu <N>— CUDA device index; unset auto-selects the highest compute capability. Ignored on CPU and Metal, where there is nothing to index.--backend <name>— F4: fence the run to a set of backends. Repeatable, and both the comma-separated (--backend cpu,cuda) and the=-joined (--backend=cpu) forms work; the namescpu,metalandcudaare matched case-insensitively.cpustays allowed whatever is named — it is the fallback — so--backend cpuis how a run is forced onto the CPU. Unset = every backend this build has;MINFER_BACKENDSis the environment form of the same surface, and the flag wins when both are set. It is extracted before subcommand dispatch, soserve,viz,benchandspecverifyhonour it too. Three distinct refusals, each naming what it rejected: an unknown name (unknown backend 'nope'; known backends are: cpu, metal, cuda), a name this build does not contain (backend 'cuda' is known but not compiled into this build: the CUDA backend is compiled only with --features cuda), and a name this machine cannot use (backend 'cuda' is compiled in but not available on this machine: <why>).--gpu-layers <N|auto>— E5: run the firstNtransformer blocks on the device and the rest on the CPU (0= CPU only; unset/MINFER_GPU_LAYERS= every block the device can hold, the pre-E5 behaviour;auto= as many as the budget allows). The placement is printed at load (offload: 4 of 24 blocks on cuda, 20 on cpu; embed/output on cpu (32.0 MiB of device weights; --gpu-layers 4)), and the tensors outside the blocks —token_embd, the final norm andlm_head— stay on the CPU unless every block is offloaded.MINFER_GPU_MEM <MiB>— the weight budgetautofits into; unset = three quarters of the device's free bytes (the same default the activation gate uses).autoalso holds back a quarter of that budget for the KV arenas and the activation pool, and reports what it decided (offload: … auto: 5 of 24 blocks fit — weights budget 64 MiB, 16 MiB reserved for KV/activations; MINFER_GPU_MEM=64 MiB). Without a device that reports free memory (Metal today)autoneedsMINFER_GPU_MEM.
Three flags describe the run instead of changing how it decodes:
--meta— dump the GGUF's metadata KV pairs and the key tensors (token_embd, the norms,lm_head) — the same two dumpsminfer info <model>prints. Unlikeinfoit does not stop there: the model still loads and generation proceeds, so it is the flag form for "show me the file, then answer".--dump-graph <PATH>— export the compute graph the run just built (build → assign → fusion, with the real backend assignment and the same fusion environment toggles as the live path) as Graphviz DOT. It runs one prefill, writes the file, printsGraph DOT exported to <PATH> (<n> nodes)and exits before decoding.--dump-graph-json <PATH>— the same graph as the JSON theviz/web visualizer reads. It printsGraph JSON exported to <PATH> (<n> nodes, <kind>), where<kind>names the phase the graph was built for, and also exits before decoding.
Chat templates and tokenizer
Chat rendering uses the model's own tokenizer.chat_template from the GGUF. The
published Qwen2.5/Qwen3 templates use Python string methods
(content.split('</think>'), .lstrip('\n')) that minijinja does not provide
natively; minfer supplies them through minijinja's unknown-method hook with
CPython semantics, so those templates render as published (this is what makes
Qwen3's think-block handling and tool-call formatting reach the model).
Reference renderings — multi-turn, a system message, a generation prompt, a
<think>-reasoning turn — are committed under tests/fixtures/chat/ with their
provenance.
A template the engine cannot compile or render is an error, never a generic prompt:
Error: chat template error — unsupported template construct: unsupported Python str
method `splitlines` (template line 41); minfer refuses to fall back to a generic
ChatML prompt. Supported Python str methods: capitalize, count, endswith, find,
join, lower, lstrip, replace, rfind, rsplit, rstrip, split, startswith, strip,
title, upper.
The CLI exits before inference, serve/viz refuse to start (checked before the
worker thread is spawned), and a per-request refusal is an HTTP 400. The generic
ChatML renderer survives only for a GGUF that has no tokenizer.chat_template at
all, and the startup path prints a one-line notice when that happens.
--no-template still bypasses template rendering entirely.
The tokenizer is byte-level BPE and is equally strict about what it does not
implement. tokenizer.ggml.pre selects the pre-tokenization rule — qwen2
(alias deepseek-r1-qwen) or qwen35 — and an unknown or missing value, a
tokenizer.ggml.model that is not gpt2, an empty merge table, or a vocabulary
missing any of the 256 byte tokens refuses the load with the offending value
named. Token ids match the reference (transformers AutoTokenizer, or llama.cpp
llama-tokenize on the same GGUF) byte for byte over the corpus committed in
tests/fixtures/tokenizer/.
Multi-turn conversation
--cnv (docs/CLI-CONVERSATION-PLAN.md): append-only KV + incremental template
rendering — each turn only prefills the new message delta, the whole
conversation accumulates in the KV cache:
./target/release/minfer --cnv qwen2.5-0.5b-instruct-q4_0 # interactive REPL
./target/release/minfer --cnv -st qwen2.5-0.5b-instruct-q4_0 "hi" # single turn
In-conversation commands: /exit /quit, /clear, /regen (regenerate the
last reply), /help; EOF (Ctrl+D) exits.
Conversation options:
-st, --single-turn— run one turn, then exit--system <STR>— system prompt-mli, --multiline-input— submit input on an empty line--color on|off|auto— color output (default auto = tty)--session <FILE>(with--cnv) — save/load the conversation.FILEis the history as JSON;FILE.kvis a KV session companion written next to it (C5) that carries the rows those messages were rendered into, plus the host state they belong to. On start a matching companion is resumed and the history is not re-prefilled — the run printsresumed N message(s) and M KV row(s) … — 0 tokens prefilled; anything that does not match this run (another--n-ctx, another model'sn_kv_embd, anotherMINFER_CACHE_TYPE, a history the user edited, an older file version) prints the reason and falls back to re-rendering the JSON, which is always correct if slower. On overflow the oldest turns are dropped automatically and generation continues: the dropped turn's KV rows are removed in place and the tail is re-based (Phase C / C2), so only the new turn's delta is prefilled;MINFER_NO_CONTEXT_SHIFT=1forces the older, exact "drop the turns and re-prefill the rest" behaviour (an engine that cannot move rows — e.g. Metal, where it is Phase G — falls back to that path on its own and says so on stderr)
Qwen3-style <think>…</think> reasoning blocks are gray-highlighted
(single-shot mode too, when stdout is a terminal or MINFER_COLOR=1).
OpenAI-compatible HTTP server
Continuous batching (the worker composes one decode batch across the active slots
instead of one forward per slot — Phase E / E2) is on by default when the
model's forwards run on CUDA or Metal, and off on CPU (E6; Metal joined in
#44 part (b), 2026-10-06). The default was
decided by measurement, not assumption: on this project's reference CPU batching measured
slower than serving requests one at a time (0.49x on 7B Q4_K_M, 0.88x on 0.5B Q4_0,
--n-slots 4), while on the GB10 it measured 1.97x faster (7B Q4_K_M, four identical
prompts, equal work, --n-slots 4, default settings).
Both readings are dated, and both are historical. The CPU pair is the E2 record's step 3
(2026-09-17); the GPU 1.97x is E6's re-fetch with the default and no environment variable
(2026-09-19, matching E2's 1.9x of 2026-09-18 within noise). All three predate the
cross-request prefix sharing described below (C8a/C8b, 2026-09-21/22), so the CPU figure
prices a per-request prefill that sharing has since made avoidable, and the reason the E2
record gave for it — concurrency forfeits the cross-request prefix reuse each slot otherwise
keeps — no longer holds. They are history, not a current claim; the boxes, the workloads and
the tables are in the E2 and E6 records of
ARCHITECTURE-EXECUTION-PLAN.md §5.
--n-slots does not cap a request. The arena is divided among the slots as an initial,
elastic share only: the server starts from n_ctx / n_slots per slot, but a request's real
bound is the whole n_ctx. When a request needs more than its share, admission reclaims
cells from idle slots (releasing their cached prefixes) and the allocator moves whatever runs
are in the way, so a busy neighbour cannot block it (C7/C7b, the same plan §5). The serial
path (MINFER_BATCH=0) is the exception: its per-slot graph region really is n_ctx / n_slots cells.
--slots-file <PATH>(batched engine only) — resume the server's slot contexts fromPATHat startup and rewrite it after every completed request (C5 S2). A request whose prompt matches a restored slot's tokens prefills only its own delta; the file also carries the KV rows, so a restart no longer re-prefills the conversations that had finished. The startup line prints the size it will write per request (about 12 MiB for a 0.5B/512-row arena), because that cost is a decision. A snapshot from another--n-slotsor--n-ctx(or another model / KV element type) is refused loudly and the server starts empty.MINFER_BATCH=1forces batching on (this is how to batch on CPU, for experiments or for a machine where your own measurement says it wins).MINFER_BATCH=0forces it off.- Any other value warns and uses the device default.
- Metal joined the default only once it could take the node the batched path needs: until
#44 part (b) (2026-10-06)
supports_attn_span()was false there, so batching would have failed loudly instead of serving and Metal stayed opt-in. The explicit switch is no longer needed. - The server prints its choice at startup:
[server] batching: on (device cuda; MINFER_BATCH=1 forces it on, =0 forces it off).
Two fuse-related switches are easy to confuse (D3):
MINFER_NO_FUSE_QKV=1/MINFER_NO_FUSE_FFN=1disable the corresponding decode fusion (the decoder builds the plain matmul/rope/store — or gate+up+ silu+mul — path instead).MINFER_FFN_COMPOSITION=1keeps the FFN fusion but builds it as the proven composition (concat matmul + gate/up windows + in-place SwiGLU) instead of the hand-written fused node. It is the reference the A/B is run against, and it is ignored with a warning on a backend without offset views (Metal until G5).
The KV cache type is its own switch (C4): MINFER_CACHE_TYPE=f32|f16|q8_0, strict — an
unknown value fails the load on every device, f16 resolves to f32 on the CPU (no f16
KV kernel there) and q8_0 is refused where the attention kernel has no packed read — it is
supported on the CPU, CUDA and Metal (Metal since
#310). A packed q8_0 cache is 3.76×
smaller and, since C4 S2, is read by a fused Q8_0 × Q8_0 K dot with V accumulated out of
the cell; MINFER_NO_FUSED_Q8_KV=1 restores the older dequantize-into-a-scratch read for
the A/B. Details, numbers and the named tolerance class: docs/BACKENDS.md,
docs/ARCHITECTURE-EXECUTION-PLAN.md §5 (C4).
Two environment switches around the GPU are easy to get wrong:
MINFER_DISABLE_CUDAis checked for presence, not value: setting it to0disables CUDA (and therefore also turns the batching default off, since the model then runs on CPU). To force the CPU path deliberately useMINFER_DISABLE_CUDA=1; to use the GPU, leave it unset.- On CUDA, batched prefills stay per request by construction (
fa_prefilltiles one query tile against one KV window — E1b), so--n-slotsconcurrency still pays one prefill per request there. The batched-decode win is unaffected. - Prefix reuse across slots is a share, not a copy, wherever the attention kernel gathers a
kv_map(CPU, CUDA, and Metal since #362): admission points the arriving request at the donor slot's rows and no byte is copied. A device that cannot gather, and any device at all underMINFER_NO_KV_SHARE=1(the A/B switch), falls back to C8a's row copy — either way the arriving request prefills only its own suffix.
Slot saturation
A request that arrives while every engine slot is busy is rejected loudly, not queued (the decision #121 pinned down). The worker answers
- non-streaming:
503 Service Unavailablewith{"error":{"code":503,"message":"no idle slot","type":"unavailable_error"}}; - streaming: the event stream has already started, so the status line is
200and the signal is adata:error frame carrying the same error object (followed by[DONE]) — not an empty stream a client could mistake for a generation that produced nothing.
minfer_jobs_dropped_total moves by exactly one per rejected request. The alternative —
holding the request until a slot frees, which minfer_queue_depth would then measure and
which the pre-E2 plan assumed (see OPENAI-CHAT-API-PLAN.md §"Slot Lifecycle") — is a
feature request, #150; the serial path
(MINFER_BATCH=0) still queues, because its single worker pulls one job at a time from
the same channel.
Structured output (F2)
/v1/chat/completions accepts OpenAI's response_format and a llama.cpp-style grammar
extension:
| Field | Effect |
|---|---|
"response_format": {"type": "text"} (or absent) | no constraint |
"response_format": {"type": "json_object"} | any single JSON value |
"response_format": {"type": "json_schema", "json_schema": {"name": "person", "schema": {…}}} | the compiled schema (name and strict are accepted and ignored) |
"grammar": "root ::= …" | a GBNF grammar inline |
grammar together with a non-text response_format is a 400 (one grammar per request, never a
precedence rule), and so is any unsupported GBNF/schema construct — the schema is compiled on the
handler side, before the request takes a slot, so the error is an HTTP 400 with the offending
construct named, not a truncated generation:
curl -s localhost:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
"messages": [{"role":"user","content":"Give me a person"}],
"temperature": 0, "max_tokens": 40,
"response_format": {"type":"json_schema","json_schema":{"name":"person","schema":
{"type":"object","properties":{"name":{"type":"string"},"age":{"type":"integer","minimum":0,
"maximum":150}},"required":["name","age"],"additionalProperties":false}}}}'
# {"choices":[{"message":{"content":"{\n \"age\": 25,\n \"name\": \"Qwen\"\n}"}}]}
A grammar is per request: the compiled automaton is shared by Arc and each request (serial
or batch slot) carries its own position, so slots cannot perturb each other. A response body never
ends mid-character: if a generation stops with a partial UTF-8 sequence pending, those bytes are
dropped (with a printed note) rather than decoded to U+FFFD.
Metrics and observability (F8)
GET /metrics returns a Prometheus text
snapshot (content type text/plain; version=0.0.4; charset=utf-8). It is served
from its own router state, so a scrape needs neither the tokenizer nor the job
channel and keeps answering while the model is busy; rendering only reads atomics,
so scraping cannot perturb generation.
Family names and units (every family is minfer_-prefixed; *_bytes is bytes,
the per-op family is seconds, everything else is a count):
- Request lifecycle:
minfer_requests_total,minfer_requests_completed_total,minfer_requests_rejected_total(refused before queueing: the server is draining, or the worker is gone),minfer_requests_in_flight(accepted and not yet finished — the drain surface),minfer_jobs_dropped_total(the worker could not place the job),minfer_worker_stalled_total(the worker's counted no-progress bound tripped: every live and queued request was answered500 the worker stalledand the worker stopped — #196). - Depth:
minfer_queue_depth(accepted - admitted: the channel backlog plus the worker's pending deque — the one number neither thread can see alone),minfer_worker_pending_jobs,minfer_requests_running(occupying an engine slot now). - Throughput:
minfer_prompt_tokens_total,minfer_completion_tokens_total, andminfer_completion_tokens_per_second— generated tokens/s over a trailing 16-second window (a lifetime average would keep reporting a startup burst on an idle server). Tokens are counted where the response is produced, so the batched and serial paths agree; a client that disconnects before its answer is complete is not counted, because those tokens were never delivered. - Drain:
minfer_draining(0/1),minfer_drain_abandoned_requests(still in flight when the deadline expired; 0 is a clean drain). - Allocator (E4
MemoryReport, for the backend the server runs on):minfer_memory_{weights,pool,live,peak_live,budget,headroom}_bytes,minfer_memory_idle_slots,minfer_memory_reserved_classes. On an unbounded backend (CPU/Metal with no budget) thebudget/headroomfamilies are omitted, not reported as 0. - KV arena:
minfer_kv_{layers,rows,region_bytes},minfer_kv_packed(1 for a packed Q8_0 cache),minfer_kv_{reserved,owned,shared,free}_cells,minfer_kv_{free_runs,sequences}, and the C3/C8b countersminfer_kv_{defrags,cells_moved,cows,cow_cells}_total. - Per-op timing (present only when the flag below is set):
minfer_op_seconds_total{op="matmul"}andminfer_op_calls_total{op=…}. The interval is the scheduler's per-node dispatch, so it includes the backend's own prologue and excludes split-level syncs and cross-backend copies; only ops that actually ran appear.
Occupancy is a live reading: the worker republishes the allocator's numbers after every step. On the batched path that is the one shared arena; the serial path has one arena per slot, so it reports the arena of the slot that served the last request.
Flags:
MINFER_OP_TIMING— presence-checked (any value), off by default. Turns on the per-op timing above. Off, the scheduler never reads the clock and the timing family is absent from a scrape; on, the numbers reported change and the numbers computed do not (a greedy run is identical with and without it).MINFER_DRAIN_MS— how long a graceful shutdown may take, in milliseconds (default30000). A value that is not a whole number of milliseconds is reported and the default is used;0means "stop now". See below.
MINFER_OP_TIMING=1 ./target/release/minfer serve --n-ctx 4096 --n-slots 1 qwen2.5-0.5b-instruct-q4_0
curl -s http://127.0.0.1:8080/metrics
Graceful shutdown (F8)
On SIGINT or SIGTERM the server stops accepting new work and lets the
requests already accepted finish, up to MINFER_DRAIN_MS:
- the listener is closed (a brand-new connection gets a connection error), and
a request that arrives on an already-accepted connection gets
503— it is counted inminfer_requests_rejected_total; - in-flight responses are allowed to complete;
- at the deadline the server logs how many requests were still in flight,
records that count in
minfer_drain_abandoned_requests, and exits. It never waits on the worker indefinitely — an SSE client that never disconnects cannot keep the process alive.
With no signal the server runs forever exactly as before, and the --slots-file
snapshot (written after every completed request) is unaffected.
./target/release/minfer serve --n-ctx 4096 --n-slots 1 qwen2.5-0.5b-instruct-q4_0
# POST /v1/chat/completions (stream + non-stream)
# GET /v1/models, GET /health
# GET /metrics (Prometheus text; see "Metrics and observability" above)
Performance testing (bench)
./target/release/minfer bench -r 3 <model> # pp512 + tg128, markdown table
./target/release/minfer bench -p 3314 -n 128 <model> # campaign-shape pp/tg
pp<P> ingests P prompt tokens (prefill-only, generate nothing); tg<T>
prefills the context then decodes T tokens. -p 0 / -n 0 skip a test, -r
sets the measured reps (1 untimed warmup each), -o csv|json emits the same
fields machine-readable, --n-ctx only ever grows the auto KV sizing
(P+T+16, clamped to the model's context length).
Verify-step gate bench (specverify)
./target/release/minfer specverify -p 512 -r 40 -o json <model>
Measures the batched verify-step cost C_T(nt) (nt = 1, 3, 5 by default) at
a fixed deep KV depth and reports the per-token amortization
nt·C_T(1)/C_T(nt) — the D5 speculative-decoding gate instrument (step doc
81). -p sets the depth, -r the timed reps (3 untimed warmups each),
MINFER_SPECVERIFY_NTS=1,3,16 overrides the phase list,
MINFER_SPECVERIFY_NOUT=1 forces prefill-style n_out=1. Exit code is 0
whenever the measurement completes; the PASS/FAIL verdict is in the JSON.
GGUF tooling — convert, quantize, split (F6)
# HuggingFace Qwen2 checkpoint -> GGUF (f16, or f32 for a lossless archive)
./target/release/minfer convert /path/to/Qwen2.5-0.5B-Instruct out.gguf --outtype f16
# quantize an existing single-file GGUF (q4_0, q4_1, q5_0, q5_1, q8_0, f16, f32)
./target/release/minfer quantize out.gguf out-q4_0.gguf --type q4_0
# split a single file into transport-sized parts (any type; tensor bytes copied verbatim)
./target/release/minfer split out-q4_0.gguf /tmp/parts --max-size 200M
convert reads config.json, tokenizer.json, tokenizer_config.json and
model.safetensors (single file or an index.json shard map) and writes the
GGUF v3 metadata the engine's strict loader requires: tokenizer.ggml.model = gpt2, tokenizer.ggml.pre = qwen2, the full 256-token byte vocabulary, the
merge table, the special-token ids and tokenizer.chat_template. Supported
architectures: Qwen2ForCausalLM only. 1-D tensors stay f32 under --outtype f16 (llama.cpp's rule); --outtype f32 is exact for bf16/f16 sources.
quantize re-encodes 2-D float weights; 1-D norms/biases keep their source
type, and on a tied model a sub-8-bit target quantizes the shared
token_embd.weight at q8_0 (both are printed). K-quants and I-quants are
refused by name — minfer can read them but has no encoder that has been
verified against llama.cpp.
split writes {stem}-NNNNN-of-MMMMM.gguf parts with
split.no/split.count/split.tensors.count; the loader reads part 0 and
merges every part into one tensor index. --split-max-size/--max-size accept
bytes or a K/M/G suffix, and a tensor is never split across parts.
Full contract, supported/refused sets and verification references:
GGUF-TOOLING.md.
Examples
# Local model
./target/release/minfer ~/models/qwen2-0.5b-q4_0.gguf "What is the capital of France?"
# Cached model by name (no full path needed)
./target/release/minfer qwen2.5-0.5b-instruct-q4_0 "Hello"
# Auto-download from Hugging Face + run (quant auto-detected, splits included)
./target/release/minfer hf:Qwen/Qwen2.5-0.5B-Instruct-GGUF:qwen2.5-0.5b-instruct-q4_0.gguf "Hello"
# Inspect GGUF metadata + key tensors
./target/release/minfer info qwen2.5-0.5b-instruct-q4_0
# List available GGUF files in a HF repo (without downloading)
./target/release/minfer download hf Qwen/Qwen2.5-0.5B-Instruct-GGUF
# Pull from Ollama and create a symlink
./target/release/minfer download ollama qwen2.5:0.5b
# List locally cached models
./target/release/minfer list
Sampler pipeline and invariants (moved from AGENTS.md)
sampler.rs (#48): one SamplerConfig drives one pipeline — logit bias → penalties (repeat / frequency / presence, last 64 tokens) → DRY → grammar mask (F2) → greedy shortcut (temp == 0) → top-k → typical → top-p → min-p → XTC → temperature or mirostat v1/v2, seeded StdRng. Every F3 knob defaults to a no-op, so the default path is bit-identical to the pre-#48 chain (pinned by test_default_pipeline_matches_the_pinned_pre_f3_sequence and, through the grammar-aware entry point, by test_default_pipeline_matches_the_pinned_pre_f2_sequence). SamplerConfig::validate refuses nonsensical values at CLI startup / HTTP 400 (never clamps silently), and logit-bias token ids are checked against the vocabulary. Mirostat's mu is caller-owned state (MirostatState: one per run / session / request / batch slot); speculative decoding refuses mirostat (--spec-draft), because a verify round samples several rows from one shared RNG. DRY sequence breakers are token-id sequences (--dry-sequence-breakers 198;13,2); llama.cpp's string form needs a tokenizer port (follow-up). CLI: --temp --greedy --top-k --top-p --repeat-penalty --frequency-penalty --presence-penalty --min-p --typical --xtc-probability --xtc-threshold --dry-multiplier --dry-base --dry-allowed-length --dry-penalty-last-n --dry-sequence-breakers --mirostat --mirostat-tau --mirostat-eta --mirostat-m --logit-bias --grammar --grammar-str --json-schema --json-schema-str -n --seed -t.
Constrained decoding (F2, #47). src/grammar.rs compiles a GBNF grammar (or a JSON Schema, through a generated GBNF) into one pushdown automaton — a flat program per rule, a set of {rule, pc} call stacks — and turns it into a per-state token bitset. The mask is applied inside sample_with_config_grammar, after DRY and before the greedy shortcut: every stage before it only shifts logits and every stage after it only removes candidates, so a forbidden token can never be chosen, and the mask consumes no RNG (mirostat/DRY are unperturbed). The compiled Arc<Grammar> is per request; the mutable GrammarState is per run, exactly like MirostatState. Token advancement is byte-level correct (a piece may be one byte of a multi-byte character); a token whose pending bytes can never complete to an accepted codepoint is rejected, EOG is legal only at an accepting state with no pending bytes, and no allowed token is a loud stop (SampleError::NoAllowedToken), never an arbitrary token. Unsupported GBNF or schema constructs are startup/400 refusals — never a silent guess; the accepted subset and every refusal are catalogued in docs/GRAMMAR-DESIGN.md. Server: response_format (json_object / json_schema) plus a grammar extension field; the two together are a 400. --spec-draft + a grammar is refused (a verify round samples several rows from one automaton state).
Decisions governing this document
This page is the current contract for the CLI surface; the decisions behind the flags and behaviours it documents are frozen in the ADR corpus:
- ADR-0006 — The KV storage format is a per-engine gate, not a process-wide global
- ADR-0014 — A KV session is a versioned, checksummed file — never a memory dump
- ADR-0015 — The offload
autofit takes a prefix, not a knapsack - ADR-0017 — Speculative decoding refuses the features its identity contract cannot carry
- ADR-0018 — The grammar mask is one stage inside the single sampler pipeline
- ADR-0019 — A chat template that cannot be rendered refuses the load
- ADR-0021 — bf16 is a round-to-nearest-even cast, and 1-D tensors stay f32