04 · Tokenizer + chat template
Stage: model dispatch + weights (03) → tokenizer + template → graph build (05). The weights are registered; nothing has run yet. This stage turns the string you typed into the integer token list that every later document operates on — the raw material of the whole rest of the series. Code:
src/tokenizer.rs(Tokenizer::load:353,encode:619,decode_bytes:661),src/template.rs(render_template:603,render_messages:494),src/main.rs(template block :1404-1439,get_chat_template:2449, decode loop :1749-1768) — lines verified at commit15fa45c.
1. Background — where this stage sits
Doc 03 left the engine with the GGUF memory-mapped and parsed, the
architecture dispatched (Qwen2 or Qwen3), and every weight tensor registered
in the compute-graph allocator — possibly on the GPU. But not one byte of your
prompt has been touched: you typed "What is 2+2?", and the model, so far,
has no idea it exists.
This stage closes that gap, and it has two halves. First, the chat template: a string, stored in the GGUF metadata, that describes how a conversation is spelled out in the exact marker format the model was trained on — your raw prompt is wrapped in that format before anything else happens. Second, the tokenizer: the code that turns that wrapped text into a list of integers.
Those integers are called token ids. A token is the model's unit of text —
a short chunk such as "What", " is", or a single character — and each
distinct chunk the model knows has a number. The model cannot read characters at
all. Its very first layer is a lookup table (the embedding matrix) that maps the
integer 3838 to a vector of floats; integer 3837 maps to a different vector.
Feed it raw characters and there is simply no table entry — nothing downstream
can run. That is why this stage gates the entire series: the token list is the
input to graph build (05), the prefill forward (09), the sampler (12), and the
decode loop (13).
The template half matters just as much, and it is the more surprising one. A
chat model was not trained to continue arbitrary text; it was trained to answer
when it sees a very specific arrangement of marker strings like
<|im_start|>user. Get that arrangement wrong and a perfectly good model
produces garbage — it will happily continue your sentence instead of
answering it. The markers are not decoration; they are the protocol.
2. Principle — how it works and why
2.1 The stage in one picture
prompt: "What is 2+2?"
│
▼
get_chat_template() reads GGUF metadata key "tokenizer.chat_template"
│ (missing, or --no-template → use the raw prompt)
▼
template::render_template() minijinja (+ a Python-`str`-method hook) renders
│ with add_generation_prompt=true; a template it cannot
│ render is a LOUD error naming the construct (F7/#50)
▼
"<|im_start|>system\nYou are a helpful assistant.<|im_end|>\n
<|im_start|>user\nWhat is 2+2?<|im_end|>\n
<|im_start|>assistant\n"
│
▼
Tokenizer::encode()
├─ 1. special-token scan whole strings like <|im_end|> become one id
├─ 2. pre-tokenize the `tokenizer.ggml.pre` rule (qwen2 / qwen35)
│ splits into word / number / punctuation pieces
├─ 3. byte-encode map every raw byte to a printable unicode char
└─ 4. greedy BPE merges merge the adjacent pair with the lowest rank
▼
Vec<u32> [151644, 3838, 374, 220, 17, 10, 17, 30, 151645, 151648, 198]
│
└──► doc 05: graph build consumes these ids (positions 0,1,2,…)
2.2 Why does the model need a template at all?
Start with what a language model fundamentally does: given a sequence of
tokens, predict a probability for every token in the vocabulary of what comes
next. A base model (trained only on raw documents) uses this to continue
text: prompt it with "The capital of France is" and it predicts " Paris" —
or equally " known", because continuing documents is its whole job.
A chat model is a base model that went through a second training phase (usually called instruction tuning or alignment). The training data in that phase was conversations, serialized in a fixed format:
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
What is 2+2?<|im_end|>
<|im_start|>assistant
Those <|im_start|> / <|im_end|> strings are special tokens — vocabulary
entries that were reserved during training and shown to the model millions of
times as turn boundaries. The model's chat behavior lives entirely in that
format: during alignment training, every example was markers, a user turn,
<|im_start|>assistant, then an answer — so the model learned the conditional
distribution "text that follows <|im_start|>assistant": answers, not
continuations.
Now the punchline about the last line. The rendered prompt ends with
<|im_start|>assistant\n — an empty assistant turn opener. This is what the
flag add_generation_prompt controls, and it is not optional. With the
opener, the model's next-token distribution is "the first token of an
assistant answer", and it says "2+2 equals 4...". Without it, the prompt
ends inside the user turn, and the model keeps writing the user turn — more
question text, or a stray <|im_end|>. It will not answer.
So the template is not a display nicety; it selects which distribution the
model samples from. minfer always passes add_generation_prompt=true on the
CLI path (src/main.rs), because a one-shot prompt is by definition a
"generate the assistant's next turn" request.
One more piece: bos and eos. bos (beginning of sequence) is a token
some model families expect at the very start of every input; eos (end of
sequence) is the token the model was trained to emit when it is done talking.
The template context exposes bos_token as a variable (src/template.rs);
whether a bos marker appears is the template string's choice, not the
engine's. For eos, minfer does not rely on the template — the GGUF metadata
carries tokenizer.ggml.eos_token_id and <|im_end|>'s id directly (§2.5),
and the decode loop treats them as stop signals.
2.3 The tokenizer: byte-level BPE, end to end
BPE (Byte Pair Encoding) is the algorithm that decided which chunks of text become tokens. It was run once, months before you ever run inference, on a huge training corpus:
- Start with every single byte as its own token (256 of them).
- Count which pair of adjacent tokens occurs most often in the corpus; merge that pair into a new token; repeat. Each merge gets a merge rank — the order number in which it was learned. Rank 0 was learned first, i.e. it was the most frequent pair in the whole corpus.
- Stop when the vocabulary reaches its target size — for Qwen models,
151,936 entries (
docs/QWEN3-SUPPORT-PLAN.md:37).
The training-time result ships inside the model file: the GGUF metadata carries
the token strings (tokenizer.ggml.tokens), the merge list in rank order
(tokenizer.ggml.merges), and the special-token types. Encoding is simply
replaying those merges on your text.
A first example — byte-exact, because it comes straight from minfer's test
suite, whose expected ids were cross-checked against llama.cpp
(src/tokenizer.rs):
text: <|User|>What is 2+2?<|Assistant|><think>\n
ids: 151644 3838 374 220 17 10 17 30 151645 151648 198
| id | stored token text | what a human sees | how it was produced |
|---|---|---|---|
| 151644 | <|User|> | (role marker) | special token, matched whole, before BPE |
| 3838 | What | What | pre-token piece whose characters merge into the stored entry |
| 374 | Ġis | ␣is | regex piece " is"; already a single vocab entry |
| 220 | Ġ | ␣ | regex \s+ piece — the lone space before a digit |
| 17 | 2 | 2 | regex \p{N} piece — one digit only |
| 10 | + | + | regex punctuation piece |
| 151645 | <|Assistant|> | (role marker) | special token |
| 151648 | <think> | (reasoning marker) | special token |
| 198 | Ċ | newline | regex \s*[\r\n]+ piece |
Two oddities in that table are the pre-tokenizer at work. The rule named by
tokenizer.ggml.pre (qwen2 for Qwen2.5 and Qwen3; F7/#50) splits text into
word pieces, single digits, and punctuation runs before any merging happens —
merges can never cross a piece boundary. Digits are matched one at a time
(\p{N} matches exactly one in the qwen2 rule), which is why 2+2 costs
four tokens and why models are famously weak at long arithmetic: every digit is
a separate concept. And a space before a digit attaches to nothing (the word
rule only glues a leading space to letters), so it becomes a bare Ġ —
token 220.
Now the greedy merge loop itself. minfer splits the piece into characters and
repeatedly merges the adjacent pair with the lowest rank. A toy illustration (invented ranks): if the piece
is [m][i][n][f][e][r] and ("i","n") has rank 88 while every other adjacent
pair ranks higher, [i][n] fuses first; the scan then repeats on the shorter
list until no adjacent pair is in the merge table, and each surviving piece is
looked up in the vocab. The real ranks live in the GGUF; the real loop is
excerpted in §3.2.3.
Why lowest rank first, and why does that give good tokenizations? Because rank
order is frequency order from training. The first merges ever learned were
the most common byte pairs; later merges built on earlier ones. Replaying
lowest-rank-first reconstructs the same segmentation the vocabulary was built
for, so your text is cut exactly the way the model saw text cut during
training. The practical payoff is compression: common words were merged
thousands of merges ago and exist as single tokens, so "the" costs one
position instead of three. That matters downstream because every cost in
this engine scales with token count — prefill matmuls, KV cache size (each
token reserves a K and V row per layer), and decode latency per generated
token.
Why byte-level? Because the alphabet is bytes, not characters. Before
merging, every raw byte 0–255 is mapped to a printable unicode character
(build_byte_to_unicode, src/tokenizer.rs; printable ASCII and most
Latin-1 map to themselves, the rest get chars from code point 256 upward —
space becomes Ġ, newline Ċ), and every vocab entry is stored in that
mapped form. The consequence: any byte string round-trips — Chinese, emoji,
binary junk — and decode (§2.4) can always invert the mapping exactly. A
character-level tokenizer cannot make that promise: a character the vocabulary
never saw has no representation at all.
2.4 Decode: ids back to bytes, and why bytes and not a String
Generation runs the same table backwards. decode_bytes concatenates
id_to_token[id] for each id — producing the mapped-form text, e.g. Ġis —
then maps every character back to its raw byte via the reverse table
(unicode_to_byte). Out come raw bytes, exactly the bytes that were encoded.
Why insist on bytes rather than a Rust String? Because a multi-byte UTF-8
character can be split across two tokens. Consider the Chinese character 中
(U+4E2D), whose UTF-8 encoding is the three bytes E4 B8 AD; minfer's test
(src/tokenizer.rs) contains a token whose mapped text is ä¸ — the
mapped forms of exactly those three bytes. If the model emits the first two
bytes of the character in one token and the third in the next, a per-token
String::from_utf8_lossy conversion would stamp a � (U+FFFD replacement
character) into your output stream permanently — the bytes were already
thrown away. decode_bytes never attempts the conversion: it emits raw
bytes, so E4 B8 + AD reassembles perfectly wherever they land. The tests
pin this: decode_bytes_keeps_multibyte_bytes (src/tokenizer.rs).
2.5 Special tokens: ids with a job
A special token is a vocabulary entry that is not a piece of human text but
a control signal: <|im_start|>, <|im_end|>, <think>, <|User|>, and so
on. In the GGUF they are flagged by tokenizer.ggml.token_type — the values
the code checks are 3 (control) and 4 (user-defined) (src/tokenizer.rs).
They get special treatment at both ends of the pipeline:
- Encode: a special token must survive as one id — the BPE machinery would
otherwise shred
<|im_end|>into ordinary character pieces. minfer scans for special-token strings before running BPE on each segment (src/tokenizer.rs), matching the earliest position first and the longest string at a given position. This is not cosmetic: DeepSeek-R1-style markers<|User|>use fullwidth unicode bars that the pre-tokenizer rule would happily split apart; the regression test at :533 keeps them intact. - Decode/generate: the ids of eos and
<|im_end|>are handed to the generation loop as stop sentinels — when the sampler produces one, the engine stops instead of appending it. They are also fed into the sampler's penalty window (src/main.rs, doc 12). The ids come fromModelDef::special_tokens()(doc 03), sourced from GGUF metadata:tokenizer.ggml.eos_token_id, plus a lookup of<|im_end|>that falls back to the eos id (src/models/qwen2/loader.rs:130-131). One vocabulary, two directions, and a set of reserved ids that act as the protocol's punctuation.
3. Implementation
3.1 Data in / data out
Input data — GGUF metadata (parsed in doc 02; the tokenizer reads it via
GgufContext, src/tokenizer.rs):
| GGUF key | Type | Lands in |
|---|---|---|
tokenizer.ggml.tokens | string array (~151,936 entries for Qwen) | id_to_token: Vec<String>, inverted into vocab: HashMap<String,u32> |
tokenizer.ggml.scores | f32 array | id_to_score (loaded for llama.cpp parity, unused) |
tokenizer.ggml.token_type | i32 array (1 normal, 3 control, 4 user-defined) | id_to_type, drives the special-token table |
tokenizer.ggml.merges | string array "A B" per merge, in rank order | merges: HashMap<(String,String), usize> — pair → rank |
tokenizer.ggml.pre | string (qwen2, qwen35, …) | pre: PreTokenizer — which rule split() applies (F7/#50). An unknown or missing value refuses the whole load |
tokenizer.ggml.bos_token_id / eos_token_id | u32 | bos_token / eos_token |
tokenizer.chat_template | one long string | passed to minijinja verbatim |
Note what is not here: no tokenizer model file, no external vocabulary. The
vocab ships inside the GGUF because the model file already had to describe its
own output layer (output.weight is [n_embd, 151936] — the vocabulary size
is baked into the weight shape), so the conversion tool writes the matching
token table alongside it.
The flow is: &str prompt + metadata → rendered String → Vec<u32> →
ctx = max(--n-ctx, ids.len()) (src/main.rs), which sizes the persistent
KV regions once for the whole run → forward(&ids, positions 0..n) (doc 05+).
During generation the direction reverses: one sampled id per step →
decode_bytes → raw bytes → stdout/SSE. The template's token cost is real
memory: the rendered wrapper becomes part of the prompt, and the prompt length
feeds ctx — a template that bloats the prompt bloats the KV allocation.
3.2 Key code
3.2.1 Template selection and rendering (CLI path)
The whole template stage in main.rs is deliberately small — read the template
out of metadata, render, encode:
#![allow(unused)] fn main() { // src/main.rs — the whole CLI template stage (F7/#50) // === Chat template (need tokenizer for bos_token text) === let processed = if no_template { prompt.clone() } else if let Some(tmpl) = get_chat_template(&gguf_model.parts[0].data) { let bos_text = tokenizer .id_to_token .get(tokenizer.bos_token as usize) .map(|s| s.as_str()) .unwrap_or(""); // An unrenderable template refuses the run here, before inference — // `validate` renders a canary conversation through it and a failure // names the construct and the template line. if let Err(e) = template::validate(&tmpl) { eprintln!("Error: {}", e.message()); std::process::exit(1); } match template::render_template(&tmpl, &prompt, true, bos_text) { Ok(p) => p, Err(e) => { eprintln!("Error: {}", e.message()); std::process::exit(1); } } } else { eprintln!( "Notice: this GGUF has no tokenizer.chat_template; using the generic ChatML renderer" ); prompt.clone() }; let input_ids = tokenizer.encode(&processed); }
Three branches, in priority order: --no-template bypasses everything and
tokenizes the raw prompt (useful for base models and for comparing token
counts); otherwise get_chat_template pulls tokenizer.chat_template from
the GGUF metadata bytes (a tiny re-parse of metadata only —
src/main.rs); with no template key at all, the raw prompt is used
as-is. The literal true argument to render_template is
add_generation_prompt — §2.2 explained why it must always be on for a
one-shot prompt. If the tokenizer produces zero ids, the run aborts: an empty
token list would leave the graph builder with no tokens to embed.
The renderer wraps minijinja, a small Jinja-compatible template engine, plus a
Python str-method hook (F7/#50). That hook is what lets the published
Qwen2.5/Qwen3 templates run at all: they use Python string syntax
(message.content.split('</think>').lstrip('\n')), and minijinja strings expose
no methods, so the engine registers
Environment::set_unknown_method_callback and implements the methods with
CPython semantics (split, lstrip/rstrip/strip with a character set,
replace, startswith/endswith, join, find, …):
#![allow(unused)] fn main() { // src/template.rs — the environment every render goes through fn environment() -> Environment<'static> { let mut env = Environment::new(); env.set_unknown_method_callback(unknown_method); env.add_function("raise_exception", |msg: String| -> Result<Value, Error> { Err(Error::new(ErrorKind::InvalidOperation, msg)) }); env } }
The context exposes exactly what real chat templates expect: messages,
add_generation_prompt, bos_token, and tools as none (transformers'
default), so {% if tools %} branches take the no-tools path.
A template that cannot be compiled (add_template fails) or cannot be
rendered (render_messages returns Err(TemplateError)) is a refusal,
not a fallback. The refusal names the construct and the template line:
chat template error — unsupported template construct: unsupported Python str
method `splitlines` (template line 41); minfer refuses to fall back to a generic
ChatML prompt. Supported Python str methods: capitalize, count, endswith, find,
join, lower, lstrip, replace, rfind, rsplit, rstrip, split, startswith, strip,
title, upper
The CLI prints that and exits; serve/viz run template::validate once at
startup and refuse to start; a per-request failure is an HTTP 400; a
conversation turn reports it. The hand-written ChatML renderer
(fallback_chatml_messages, src/template.rs) survives for exactly one case —
a GGUF with no tokenizer.chat_template key at all, where there is no model
format to lose (that is the branch that prints the Notice: line above).
ChatML is the marker convention Qwen models are trained on and the de-facto
lingua franca of chat templates, which is why it is a reasonable default for a
GGUF that carries no template. It is deliberately not used as a fallback when a
template exists: the F7 design record
(docs/CHAT-TEMPLATE-AND-TOKENIZER-DESIGN.md) shows what that silently cost —
Qwen3's think-block extraction and tool-call formatting, fed back verbatim.
3.2.2 Loading the tokenizer from GGUF metadata
Tokenizer::load (fallible since F7/#50: it returns Result<Self, String> and
refuses a tokenizer it cannot reproduce byte for byte) walks the metadata
key-value list once per data kind. The
token strings become the id_to_token vector and an inverted vocab map
(:114-118); special-token ids and types fill special_tokens
(:135-143) plus bos_token / eos_token / im_end (:145-149). BPE's
data structure is built at :120-133: for each tokenizer.ggml.merges
entry — a string like "Ġ t", two space-separated halves — the code splits on
the first space and inserts merges[(first, second)] = i, where i is the
array index. That index is the merge rank, because converters write merges
in the order they were learned. Splitting on the first space is enough
because each half is one byte-encoded string with no literal spaces in it
(spaces were mapped to Ġ precisely so they could never appear inside a
half).
Special tokens need one more data structure, and its comment explains the invariant:
#![allow(unused)] fn main() { // src/tokenizer.rs, 172-179 (the <|im_start|>/<|im_end|>/eos // fallback inserts between, described in the text below) // Merge GGUF special tokens (type 3/4) with hardcoded fallbacks, then // group by first char with longest-first ordering inside each group // (an earliest-position, longest-match scan needs both). let mut special_by_first: HashMap<char, Vec<(String, u32)>> = HashMap::new(); for (pat, id) in merged { let first = pat.chars().next().unwrap_or('\0'); special_by_first.entry(first).or_default().push((pat, id)); } for group in special_by_first.values_mut() { group.sort_by(|a, b| b.0.len().cmp(&a.0.len())); } }
The skipped middle starts from special_tokens.clone() and
defensively inserts <|im_start|>, <|im_end|>, and the eos token only when
the GGUF did not already provide them (contains_key guards): some converted
models mark their specials as ordinary type-1 tokens, so minfer hardcodes the
ChatML markers as fallbacks, and real metadata always wins. The resulting
special_by_first index buckets patterns by their first character; the encode
scan (next) will jump straight to the bucket for the character it is looking
at instead of testing every pattern against every position.
3.2.3 Encode: specials first, then regex, then merges
The top-level encode is a loop over "segments": text up to the next special
token goes through BPE, the special token becomes a single id, repeat
(src/tokenizer.rs, doc comment at :273-279):
#![allow(unused)] fn main() { // src/tokenizer.rs (fn head at :280-283, final `result` at :309-310) loop { // Find the earliest position where any special token starts. let mut earliest: Option<(usize, u32, usize)> = None; // (byte_pos, id, byte_len) 'scan: for (ci, ch) in remaining.char_indices() { if let Some(group) = self.special_by_first.get(&ch) { let rest = &remaining[ci..]; for (pat, id) in group { if rest.starts_with(pat.as_str()) { earliest = Some((ci, *id, pat.len())); break 'scan; // group is longest-first; earliest char wins } } } } if let Some((pos, id, len)) = earliest { // Encode text before the special token if pos > 0 { result.extend(self.encode_bpe(&remaining[..pos])); } result.push(id); remaining = &remaining[pos + len..]; } else { // No more special tokens, encode the rest result.extend(self.encode_bpe(remaining)); break; } } }
The double ordering matters: the outer scan takes the first character that
starts any special token ("earliest position wins"); within one position,
the bucket is sorted longest-first, so the first starts_with hit is the
longest match ("<think▁begin|> beats <think>"). The dedicated tests
special_token_earliest_position_wins and
longest_special_token_wins_at_same_position pin both rules.
Inside a segment, encode_bpe (src/tokenizer.rs) applies the pre-tokenization
rule named by tokenizer.ggml.pre (F7/#50) over the text — for qwen2, which
Qwen2.5 and Qwen3 both select:
(?:'[sS]|'[tT]|'[rR][eE]|'[vV][eE]|'[mM]|'[lL][lL]|'[dD])|[^\r\n\p{L}\p{N}]?\p{L}+|\p{N}| ?[^\s\p{L}\p{N}]+[\r\n]*|\s*[\r\n]+|\s+(?!\S)|\s+
The alternation, read left to right: contractions ('s, 't, 're…), an
optional leading space glued to a run of letters, a single digit, an optional
leading space glued to punctuation, line breaks, and any other whitespace. It is
implemented as a hand-written scan rather than a regex pattern because the
Rust regex crate has no lookahead and the (?!\S) on the whitespace
alternative is load-bearing for duplicated spaces; PreTokenizer::split returns
byte slices that tile the input exactly, pinned byte for byte against CPython
regex on the model's own tokenizer.json pattern by the committed split
fixtures. Each piece becomes one piece here: byte_encode (the per-byte
character mapping from §2.3) maps its bytes to printable chars, and the piece
goes to bpe_encode. These piece boundaries are sacred — merges never cross
them — which is why " is" and "2" never fuse into one token no matter what
the merge table says.
Then the merge loop itself:
#![allow(unused)] fn main() { // src/tokenizer.rs (whole-piece shortcut at :218-221, lookup tail at :251-254) loop { // Find the best merge (lowest rank) let mut best_rank: Option<usize> = None; let mut best_idx: Option<usize> = None; for i in 0..word.len().saturating_sub(1) { let pair = (word[i].clone(), word[i + 1].clone()); if let Some(&rank) = self.merges.get(&pair) { if best_rank.is_none() || rank < best_rank.unwrap() { best_rank = Some(rank); best_idx = Some(i); } } } if best_idx.is_none() { break; } // Merge at best_idx let idx = best_idx.unwrap(); let merged = format!("{}{}", word[idx], word[idx + 1]); word.splice(idx..=idx + 1, std::iter::once(merged)); } }
Read it as: loop {scan every adjacent pair, keep the lowest-rank one, splice
it}, until no pair is in the merge table. There is deliberately no
"whole piece is already a vocab entry" shortcut (F7/#50 removed one): a
vocabulary entry is not necessarily reachable through merges — Qwen3.5 holds a
Devanagari cluster as one entry with no rank producing it, and every reference
splits it — so the loop always starts from single characters, exactly like HF
tokenizers and llama.cpp. After it, each surviving piece is looked up in the
vocab, and a piece that is not there falls back one token per byte
(byte_fallback), which Tokenizer::load guarantees is possible by refusing a
vocabulary that lacks any of the 256 byte tokens. The old tail mapped such a
piece to id 0 (unwrap_or(0)) — see §3.4 for why that was a bug. Complexity is
O(pieces²) per word with tiny constants; tokenization runs once per prompt, so
it is not a hot path.
3.2.4 Decode: ids → bytes → streamed text
#![allow(unused)] fn main() { // src/tokenizer.rs pub fn decode_bytes(&self, ids: &[u32]) -> Vec<u8> { let mut encoded = String::new(); for &id in ids { if (id as usize) < self.id_to_token.len() { let token = &self.id_to_token[id as usize]; encoded.push_str(token); } } // Reverse byte-level encoding let mut result = Vec::new(); for c in encoded.chars() { if let Some(&b) = self.unicode_to_byte.get(&c) { result.push(b); } else { // Fallback: encode the char as UTF-8 let mut buf = [0u8; 4]; let s = c.encode_utf8(&mut buf); result.extend_from_slice(s.as_bytes()); } } result } }
Two quiet robustness choices: out-of-range ids are skipped — a corrupted
sample cannot panic the stream (tested at :466-471) — and characters not
in the reverse byte map (vocab entries holding genuine unicode text rather
than byte-mapped forms) pass through as their UTF-8 bytes (:336-341).
The streaming holdback, used by the server and conversation paths, is
complete_utf8_prefix_len (src/tokenizer.rs). Its doc comment says
it "mirrors llama.cpp's format_incomplete_utf8 holdback", and the mechanism
is just the UTF-8 length grammar walked once: the lead byte of a sequence
determines its length (1 for ASCII, 2–4 for multi-byte, judged by the
0xC0/0xE0/0xF0 masks on the top bits), so the function scans forward
until the next character would run past the end of the buffer, and returns the
offset where the complete prefix ends — everything from there onward waits for
the next token. The callers wire it into per-step emission: server/chat.rs:162
appends newly decoded bytes to a full accumulator and server/chat.rs:176
flushes up to emitted + complete_utf8_prefix_len(&full[emitted..]);
conversation.rs:548/564 does the same for the REPL. (The CLI decode loop
instead writes raw bytes straight to stdout and lets the terminal assemble the
glyph; complete_utf8_prefix_len_holds_incomplete_trailing at :474 pins the
grammar.)
And the consumer end — how the ids the sampler produces meet this decoder in
the CLI decode loop (doc 13 walks the loop; here, only the tokenizer-relevant
lines). The sampled id first runs the stop-sentinel check, then is recorded
for the penalty window (generated.push / prev_tokens.push,
:903-907):
#![allow(unused)] fn main() { // src/main.rs, 909-918 if is_stop_token(sampled.token_id, &special) { break; } // Stop-string detection on the FULL byte stream before emitting. full.extend_from_slice(&tokenizer.decode_bytes(&[sampled.token_id])); if let Some(cut) = sampler::match_stop_suffix(&full, &stop_refs) { full.truncate(cut); if cut > emitted { hi.feed(&full[emitted..]); emitted = full.len(); } break; } }
special came from model.special_tokens() (src/main.rs), and
is_stop_token (src/main.rs) is the two-line sentinel check
id == special.eos || Some(id) == special.im_end — the SpecialTokens struct
itself is two fields (eos, im_end: Option<u32>, src/models/mod.rs).
Note the byte-accumulation pattern: decoded bytes go into full before
emission (the rest of the loop, :919-922, flushes newly completed bytes), so
a stop string (--stop "Let me think") that straddles two tokens is caught in
the accumulated stream — the same byte-first philosophy as §2.4. The tokens
handed to prev_tokens also feed the sampler's repeat-penalty window (doc
12), so prompt and generated token ids influence penalties from the very
first step (src/main.rs).
3.3 Design choices (why this shape and not another)
Why a chat template at all — why not just tokenize the prompt?
Because tokenization is the wrong layer for chat behavior. The model was
aligned on a marker protocol; the only way to reach its "answer mode" is to
reproduce that protocol byte-for-byte at the input (§2.2). This gets chat
behavior out of the same weights by changing only the prompt text, where
separate per-mode models or hidden role channels would multiply the model. The
cost: template rendering is a compatibility surface (§3.4's minijinja
gotcha), and template correctness is invisible until the model answers wrong
— hence the debug dump (§4) and the --no-template bypass.
Why self-contained BPE instead of a tokenizer crate?
The obvious alternative is a dependency — tokenizers (HuggingFace) or
tiktoken — bringing a large dependency tree and its own version skew.
minfer's constraint is zero ML framework deps (ARCHITECTURE.md §1: five crates
total, minijinja the newest). The decisive fact is that the tokenizer's
data already ships in the GGUF — tokens, scores, types, merges,
special-token flags — so a crate would mostly re-read the same tables and add
a second source of truth; the algorithm itself is ~80 lines (§3.2.3). The
honest trade-off is exactness risk: BPE implementations differ in
pre-tokenization details, and a divergence silently changes every id. minfer
buys that risk down with llama.cpp-parity tests —
special_tokens_match_as_single_ids_before_bpe hardcodes ids copied from
llama.cpp's tokenizer as the expected output (src/tokenizer.rs), so
any divergence fails CI rather than shipping as subtly different model
behavior.
Why greedy lowest-rank merging?
Because the merge table is a frequency-ordered construction history, and
lowest-rank-first replay is the inverse operation (§2.3): it reproduces the
segmentation the model was trained on, with no search — one linear scan per
merge round. The alternatives are worse on both axes: longest-first or
highest-rank-first produce segmentations the model never saw (the pieces
would still be valid tokens, but the embedding each maps to was trained on
different contexts — quality quietly degrades), and optimal segmentation
search (minimize token count, e.g. Viterbi) costs orders of magnitude more.
The compression payoff is concrete: single-token " is" costs one position
instead of three — one fewer row through every attention head and one fewer
KV row per layer, multiplied by every layer.
Why add_generation_prompt=true (and hard-coded on the CLI path)?
The rendered prompt must end with the empty assistant-turn opener
(<|im_start|>assistant\n), or the model's next-token distribution is
"continue whatever turn is open" — usually a continuation of the user's own
text (§2.2). It is hard-coded true on the CLI path (src/main.rs)
because a one-shot CLI prompt is definitionally a "start the assistant's turn"
request. The multi-turn paths pass it explicitly too — and only on the final
render: format_single (src/template.rs), the incremental renderer
behind --cnv, renders the recorded past with false and only the new state
with true, then diffs the two strings so the KV cache is appended with just
the delta — a mid-history render must not append an opener, or the KV would
contain an assistant header that never led to an answer.
Why byte-level (and why decode in bytes)?
Byte-level BPE gives total input coverage (§2.3) and byte decode gives
lossless streaming (§2.4). A string-oriented decoder would corrupt output
precisely in the cases that matter — emoji, CJK — and permanently:
from_utf8_lossy cannot be undone. The design keeps lossy conversion strictly
at the presentation edge (the doc comment on decode,
src/tokenizer.rs, says streaming paths must use decode_bytes),
never in the data path.
Why match special tokens before BPE, with earliest-then-longest rules?
Special tokens are protocol punctuation; letting BPE see them destroys their
meaning (and with R1-style fullwidth markers, the regex pieces can never
recombine into the special string — the test comment at
src/tokenizer.rs says exactly this). Earliest-position-wins matches
how a human reads: the leftmost marker is the next structural event.
Longest-at-position-wins disambiguates prefixes (<think vs
<think▁begin|>); any other priority would be arbitrary.
3.4 Pitfalls & invariants
The minijinja 2.21 gotcha — fixed in F7, and it is why the hook exists.
minfer uses minijinja = "2" with default-features = false. In minijinja 2.21
template strings are Rust strings and expose no str methods, so Qwen3's
shipped chat_template (Python method syntax:
message.content.split('</think>'), .lstrip('\n')) used to fail at render
time (unknown method: string has no method named split) and silently fall back
to ChatML — losing think-block extraction, tool-call formatting and
enable_thinking (docs/QWEN3-SUPPORT-PLAN.md §5 gotcha #9 keeps the
historical record). F7 (#50)
installs minijinja's unknown-method callback and implements those methods with
CPython semantics, so the model's own template runs. The new failure mode is
the opposite of the old one: if a template needs a construct the hook does not
implement you get Error: chat template error — unsupported template construct: … and no inference — never a generic prompt that silently changes model
behaviour.
Specials must never reach BPE. The whole-string scan happens before
encode_bpe, and the regression test exists because the R1 template broke
otherwise. Invariant: a new special-token source must join
merged/special_by_first before encode runs.
Unknown pieces used to map to id 0 — now they cannot. bpe_encode's tail
maps an unknown piece to one token per byte, and Tokenizer::load refuses a
vocabulary that lacks any of the 256 byte tokens (F7/#50). Before that, a
vocab/merges inconsistency silently yielded the vocabulary's first entry and one
word of output was consistently garbage. The empty-encode guard in main.rs
still catches the louder "nothing encoded at all" failure.
Byte-decode invariant: no lossy conversion in the streaming path. The
lossy decode (src/tokenizer.rs) exists for tests only — its doc
comment says so explicitly. Streaming paths must pair decode_bytes with
complete_utf8_prefix_len, or multi-byte characters split across tokens become
permanent U+FFFD in the transcript.
Template and conversation modes are coupled. --cnv refuses
--no-template (src/main.rs) because the conversation session's
append-only KV scheme requires template rendering to compute what the next
turn appends.
Template output feeds KV sizing. ctx = max(n_ctx, prompt_len)
(src/main.rs): the rendered prompt's token count participates in sizing
the persistent KV regions (doc 07). A runaway template (e.g. one that
duplicates history) does not just slow prefill — it changes the allocation.
4. Observe & verify
- The two printed counts. Every CLI run prints
Vocabulary: 151936 tokens(src/main.rs— the number for Qwen-family models) and thenPrompt: {} tokens(printed right after tokenization). Run the same prompt with and without--no-template: the difference is exactly the boilerplate the template added (system turn, role markers, the assistant opener). - See the rendered prompt. Build with
--features debug_dumpand setMINFER_DUMP_DIR:crate::dump::maybe_dump_text("minfer_dump_prompt", …)(src/main.rs) writes the post-template, pre-tokenization string — the literal<|im_start|>…<|im_end|>…<|im_start|>assistanttext of §2.2. Format reference:docs/debug-dump.md. - Unit tests are the fastest oracle — and the llama.cpp cross-check.
cargo test tokenizer::covers the byte round-trip (decode_bytes_reverses_byte_encoding, the CJKdecode_bytes_keeps_multibyte_bytes), the holdback grammar (complete_utf8_prefix_len_holds_incomplete_trailing), and all three special-token rules — including the id list copied from llama.cpp, and the real-modeltoken_ids_match_the_referencegate (5 cached models × 52 corpus entries) that fails on any id shift.cargo test template::covers rendering byte-for-byte against the committed transformers references, the loud refusal, and the incrementalformat_singlediff semantics (format_single_diffs_only_new_user_messagerenders the real Qwen ChatML template shape,src/template.rs). - Per-token text in traces and the loud refusal. With
MINFER_TRACE=<path>, the decode loop attaches each sampled token's decoded text to the trace (crate::trace::set_token,src/main.rs) — handy for spotting byte-fallback pieces from §3.4. A template failure printsError: chat template error — …on stderr and stops before inference: the model's own template is not renderable and the engine refused to substitute a different prompt. The only line that means ChatML replaced a template isNotice: this GGUF has no tokenizer.chat_template; using the generic ChatML renderer.
5. Cross-references
- 02 — GGUF load: where
tokenizer.ggml.*andtokenizer.chat_templatecome from — this stage is a metadata consumer. - 03 — Model dispatch and weights: provides
ModelDef::special_tokens()(eos,im_end) and the weights this stage's ids will drive. - 05 — Graph build: the IR and the builder: consumes
the token list;
GraphParamsand the KV sizing thatctx = max(n_ctx, prompt_len)feeds. - 09 — Prefill forward path: the ids become the
embedding lookup rows with positions
0..len. - 12 — Sampler: stop sentinels and the penalty window that receives prompt + generated ids; stop-string byte matching from §3.2.4.
- 13 — Decode loop + graph reuse: the loop
whose per-token
decode_bytes+ holdback streaming this doc set up; also the multi-turn path whereformat_singlerenders only the appended turn. docs/QWEN3-SUPPORT-PLAN.md§5 #9: the full minijinja 2.21 record — the exact template lines that failed and the fallback consequences, kept as the history of the bug F7 fixed.docs/CHAT-TEMPLATE-AND-TOKENIZER-DESIGN.md: the F7 contract — accepted and refused template constructs, the loud refusal, the pre-tokenizer rules, and the reference behind every gate.docs/CHAT-TEMPLATE-AND-TOKENIZER-DESIGN.md: the F7 contract — accepted and refused template constructs, the loud refusal, the pre-tokenizer rules, and the reference behind every gate.docs/OPENAI-CHAT-API-PLAN.mdanddocs/CLI-CONVERSATION-PLAN.md: the server-side template handling (render_messages,tools) and the incremental-render design (format_single) behind multi-turn sessions.docs/USAGE.md: every flag this stage reads (--no-template,--stop,--cnv).docs/GLOSSARY.md: backstop definitions (BPE, merge rank, ChatML, bos/eos).ARCHITECTURE.md§3: the flowchart rows this doc expands (template render → tokenize → prefill).
← 03 — Model dispatch and weights · Index · 05 — Graph build: the IR and the builder →