Glossary — every term and formula in the CUDA campaign docs, by layer

This is the consolidated vocabulary of the CUDA campaign corpus (CUDA_OPTIMIZATION.md, CUDA-TECH-PRIMER.md, the 80 step docs in cuda_optimization_steps/, the four analysis/plan docs, and the build/safety references). Every entry is corpus-verified: it appears in at least one of those documents.

0. The layer taxonomy

The primer's §0 originally introduced three layers (Algorithm / Performance model / Micro-architecture) to explain the D5-0 gate chain. The full-corpus audit showed that roughly half of the vocabulary does not fit any of those three, so the taxonomy is extended to seven layers:

LayerDomainTypical questions it answers
L1 AlgorithmLLM inference algorithms & their mathWhat does the model compute? What is the optimal d?
L2 Numerics & data formatsquant formats, byte layouts, rounding, tolerancesHow are numbers represented and how wrong may they be?
L3 Performance modelroofline, bandwidth, amortization, break-evenHow fast should it go, and why isn't it?
L4 Micro-architecturetiling, tensor cores, SASS, shared memory, schedulingHow does the kernel actually use the machine?
L5 Platform & toolingCUDA API/runtime, nvcc/PTX/SASS toolchain, profilers, env gatesWhat do we drive the hardware with, and how do we look inside?
L6 Engine architectureminfer's compute graph, backends, allocator, fusion, safety rulesHow is the engine itself structured?
L7 Methodologymeasurement discipline, parity gates, forensics, doc conventionsHow do we know a number is real?

The D5-0 gate chain still illustrates why layers matter — the same terms now map to L1/L4/L3 explicitly:

d=2 (L1) → verify = batched nt=3 decode step (shape) → which tile-regime (L4)
  → amortization ≥ 2.5x? (L3) → break-even at p≈0.68 (L3) → go/no-go on d=2 (L1)

Source-doc tags used below

TagDocument
PRIMERCUDA-TECH-PRIMER.md
HUBCUDA_OPTIMIZATION.md (campaign hub + live status)
STEPScuda_optimization_steps/NN-*.md (numbered step doc)
MMQLLAMA-CPP-MMQ-ANALYSIS.md
SPECLLAMA-CPP-SPECULATIVE-ANALYSIS.md
D5SPECULATIVE-DECODING-PLAN.md
GRAPHCOMPUTE-GRAPH-DESIGN.md
BACKENDCUDA-BACKEND-DESIGN.md
BUILDBUILD.md
SAFETYGPU_SAFETY.md
METALMETAL_OPTIMIZATIONS.md (its §5.6 is a doc-local symbol glossary)

L1 — Algorithm: LLM inference algorithms & math

Term / FormulaOne-line meaningSource
prefillOne forward pass over the whole prompt, batched, compute-bound.PRIMER, STEPS 20+
decodeOne token per step (nt=1), memory-bound: streams all weights every token.PRIMER, STEPS
logitsThe [nt][vocab] output of the final matmul; input to sampling.GRAPH, STEPS
greedy decodingAlways take argmax of logits — deterministic, used for parity gates.HUB, STEPS
top-k / top-p / temperature / repeat-penaltyThe sampler chain in sampler.rs, defaults matching llama.cpp (0.8/0.95/1.1).GRAPH
BPE tokenizerByte-pair-encoding tokenizer parsed from GGUF metadata; special-token match matters (DeepSeek-R1 distill).GRAPH
RoPERotary position embedding applied to Q/K in-place, fused into decode QKV.PRIMER, GRAPH
RMSNormRoot-mean-square layer norm; fused with residual add in decode kernels.PRIMER, GRAPH
SiLU / SwiGLUActivation and gated FFN; SwiGLU is fused (gate+up concat + silu-mul).PRIMER, GRAPH
softmaxAttention normalization; FA prefill uses tiled online softmax, FAP2 keeps it register-resident.STEPS 46-48
GQAGrouped-query attention: fewer KV heads than Q heads; decode attention must replicate K/V.STEPS 26
KV cachePersistent per-layer K/V tensors grown token by token; the "positions as data" rule keeps topology fixed.GRAPH, PRIMER
n_pastNumber of cached tokens; must never appear in graph topology, only as data.GRAPH
draft modelSmall model (0.5B) proposing d tokens the big model verifies.D5, SPEC
draft length dTokens drafted per speculative round (d=2/4/8 measured).D5
acceptance rate pProbability the target accepts a drafted token; measured p≈0.68–0.70 (llama speculative-simple, greedy).D5, STEPS 80
E[a] = Σ_{i=1..d} p^iExpected tokens accepted per round under independence.D5, PRIMER
verify passOne batched forward (nt = d+1) checking all drafted tokens at target cost C_T(n).D5
speculative speedup ruleWin only if E[a]·C_T(1) > C_T(d+1) + d·C_D — the whole D5 plan reduces to this inequality.D5, STEPS 80
MTP (multi-token prediction)Draft source using MTP heads (DeepSeek-V3 / Qwen3-Next GGUFs); unavailable to minfer's dense models.SPEC, D5
EAGLE-3 / DFlash / DSpark / n-gram self-speculatorsThe other draft mechanisms in llama.cpp's speculative zoo, outside draft-simple's scope.SPEC
MoE + MLAArchitecture prerequisite for MTP drafts — its own future campaign; minfer targets dense.SPEC, D5

L2 — Numerics & data formats

Term / FormulaOne-line meaningSource
GGUF v3Single-file model format; multi-part files merge into one tensor index, entry = part 0.GRAPH, BACKEND
Q4_0 / Q4_1 / Q5_0 / Q5_1 / Q8_0Legacy block quants: fixed block of 32 values + fp scale(s); q8_0 also used for activations.PRIMER, BACKEND
Q4_K / Q5_K / Q6_KK-quants: super-block of 256 split into 8 (or 16 for q6_K) sub-blocks with packed multi-bit scales.PRIMER, MMQ
super-block / sub-blockq4_K: 8×32 with 6-bit scales packed 4-per-32-bit word; q6_K: 16×16 with 8-bit scales + ql/qh nibble halves.MMQ, PRIMER
d, dmin, minPer-block fp16 scale, and for K-quant the per-super-block scale of scales / offset.PRIMER
dscPer-super-block scale descriptor staging (q4_K DSC path; MINFER_MMQ_Q4K_DSC).STEPS 38-39
ql / qhLow/high nibble halves of packed 4-bit weights (q6_K: ql+qh interleaved by bit plane).MMQ
SWAR unpackSIMD-within-register nibble→int8 expansion via bit masks instead of per-byte ops (r30).STEPS 30
dequantize / dequantConverting packed blocks to arithmetic values inside the kernel; raw kernels avoid materializing f16.MMQ, STEPS
raw-byte / raw-nibble kernelsOperate directly on packed bytes ("BT") or nibbles ("NB") without a dequant f16 round-trip.MMQ, STEPS 28-38
round-trip (quant↔dequant)The correctness check pattern: quantize, dequantize, compare against fp reference.STEPS
tolerance gateNumerical acceptance: abs err ≤ 0.05 vs a CPU reference computed in f64.HUB, STEPS 78
bitwise identityStronger gate: refactor must produce byte-identical outputs (fused vs unfused is bit-identical by design).GRAPH, STEPS 78
f16 storage vs f32 accumulateWeights/activations may be stored f16 (__half) but mma/dp4a accumulate in f32 or int32.PRIMER, STEPS
int8 prefill activationsCUDA prefill quantizes activations to int8 for IMMA GEMM; decode MMVQ reads f32.PRIMER, BACKEND
Q8_0 activation quant (CPU)CPU quantizes activations on the fly; GPU reads f32 — logits differ by design, compare per-path.GRAPH
f32-accumulate mmar15 experiment: accumulate tensor-core results in fp32 registers instead of int32.STEPS 15
__expf scale pathq6_K dsc rebuild uses __expf; parity-guarded (W_exp debug, r44/r54 gates).STEPS 44, 54
W16 cacheCUDA-side cache of weights converted to f16 for some paths (MINFER_NO_W16CACHE to disable).STEPS 34+
split-k (dpl)"dpl" = split-plane B layout used by the final q6_K BT kernel (doc 76).STEPS 76
MMVQ uint4 sub-pairsVectorized 16-byte loads split per-thread sub-pairs in the MMVQ weight-streaming rework.STEPS 12+
QI8_1llama.cpp MMQ tiling constant: int8-activation tile width per 32-k chunk (= QK8_1/(4·QR8_1) = 8).MMQ

Tensor-layout symbols (from METAL §5.6)

SymbolOne-line meaningSource
n_embd / n_head / nkModel hidden size, query-head count, KV-head count; gqa = n_head/nk.METAL
hd / hd_kvAttention head dim and KV head dim (may differ under GQA).METAL
nt / nkv / nktTokens in the batch (decode nt==1), KV positions used, KV capacity.METAL
od / id / nfMatmul output/input dims (weight rows/cols) and FFN intermediate dim.METAL
positionsPer-token KV position array; nkv = positions[t] + 1.METAL
ne00..ne33ggml tensor dims: ne0x = dim0 of the x-th src, ne1x = dim1, etc.METAL
nb10..nb33ggml byte strides per dim for src1 (nb10 elem stride, nb11 row/token stride).METAL
ns10 / ns20Element counts per head/row/token (nb11/nb10, nb21/nb20) — flash-KV inner-loop stride.METAL
nwg / nsgWorkgroups and simdgroups per threadgroup (Metal launch geometry).METAL

L3 — Performance model

Term / FormulaOne-line meaningSource
memory-bound / compute-boundLimited by bytes moved vs FLOPs issued; GB10 decode is memory-bound, prefill compute-bound.PRIMER
roofline modelPerformance ceiling = min(peak FLOPs, AI × peak bandwidth); AI = FLOPs per byte.PRIMER
arithmetic intensity (AI)FLOPs per byte of traffic; decode GEMM at nt=1 is ~1 MAC/weight-byte → bandwidth-bound.PRIMER
GB/s, TB/sEffective bandwidth; GB10 unified LPDDR5x ~273 GB/s shared CPU+GPU.PRIMER, STEPS
tok/sDecode throughput in tokens per second (headline metric: 7B q4_k_m CUDA 54.3).HUB, STEPS 80
MAC / GMACMultiply-accumulate; GMAC = 10⁹ MACs; TMAC/s = 10¹² MACs per second (kernel-level throughput).STEPS 65-73
M/GMACSASS instructions issued per GMAC — the instruction-stream efficiency metric (llama 6.06 vs ours 10.14).STEPS 65-72
amortizationSpreading fixed weight traffic over more rows: nt=4 batched decode gives BT-MMQ 2.7×.PRIMER, STEPS 80
C_T(n)Cost of a target verify at batch n; C_T(1)=18.42 ms for 7B q4_k_m CUDA.STEPS 80
C_DDraft-model cost per token (0.5B: 2.92 ms CUDA → CPU-draft dead at 1.35×).STEPS 80
break-even p*Minimum acceptance rate for speculative win: 0.73 / 0.81 / 0.90 at d=2/4/8.STEPS 80
D5-0 gateCondition to proceed: measured nt=3 verify amortization ≥ 2.5×.STEPS 80
batched-decode regiment = tokens per decode step; nt=4 is the campaign's anchor amortization point.PRIMER, STEPS
KV traffic sharePer-token bytes = weights (dominant) + KV read/write + logits; quantized KV shrinks the KV share.PRIMER
3× gap attributionMethod of splitting the llama.cpp-vs-minfer wall-clock gap into per-kernel shares before optimizing.STEPS 12, 47, 65
wavefront countShared-memory work serialized per wavefront — 1.76× wavefronts/IMMA at equal IMMA rate meant inefficiency, not scarcity.STEPS 36
bytes-per-tokenDecomposition of decode memory traffic; the roofline input for every decode optimization.PRIMER

L4 — Micro-architecture (kernel implementation)

Term / FormulaOne-line meaningSource
SM (streaming multiprocessor)The GPU core unit; GB10 has 6144 CUDA cores across SMs; occupancy counts blocks/SM.PRIMER
blocks/SMResident blocks per SM; NB kernel uses 2, q6_K BT uses 3 (r40 probe).STEPS 28, 40
occupancyRatio of resident warps to maximum; raised by lowering registers/smem per block.STEPS 12, 40
register pressureToo many registers per thread kills occupancy; measured via ptxas -v spill output.STEPS 12-73
__launch_bounds__Compiler directive capping registers/threads to hit a target occupancy.PRIMER, STEPS
warp32 threads executing in lockstep; divergence inside a warp serializes paths.PRIMER
tile / tile shapeThe M×N×K block a kernel iterates over; TM=128/256 x-tile widening experiments (r13, r23).PRIMER, MMQ
tile-regimeWhich pre-tuned launch/tile configuration an (M,N,K) shape lands in — the D5-0 pivot concept.PRIMER, STEPS 80
wave quantizationPartial last wave of blocks leaves SMs idle; small GEMMs must size grids to avoid it.PRIMER, STEPS 12
mma.sync m16n8k16Tensor-core int8 matrix-multiply-accumulate instruction; the BT GEMM inner op.PRIMER, MMQ
wmmaLegacy warp-level matrix API; used for FA prefill P·V on tensor cores.STEPS 20
dp4a4-way int8 dot-product instruction; the MMVQ decode path's inner op.PRIMER, MMQ
IMMA / tensor pipeThe int8 tensor-core hardware pipe; smsp__inst_executed_pipe_tensor_subpipe_imma counts it.STEPS 36, 65
LDSM / ldmatrixLoads an 8×8 f16 fragment into registers laid out for mma; A-fragment reuse ratio 0.125 vs 0.5 was the llama edge.STEPS 36, 65
A-fragment / B-fragmentThe mma operand fragments each warp holds; reuse rate decides LDSM traffic.STEPS 36, 65
LDGSTS / cp.asyncAsync global→shared copy bypassing registers; the BT kernels stage A/B/dsc with it.PRIMER, STEPS 45
cp.async-db2cp.async with 2-stage double buffering (MINFER_MMQ_RAW_* sched gates).STEPS 45-54
LDG / STS / LDSGlobal load, shared store, shared load — the synchronous counterpart trio.STEPS, PRIMER
SASS instruction names (IMAD, I2F, F2I, FMUL, FFMA, LOP3, SHF, PRMT, LEA, SEL, CS2R, IADD3, FADD, BRA)The assembly opcodes read in SASS forensics to count real work per loop.STEPS 14-73
shared memory / dynamic smemOn-chip scratchpad; sized via cudaFuncAttributeMaxDynamicSharedMemorySize.PRIMER, BACKEND
bank conflictsSimultaneous LDS hits to the same bank serialize; r22 removed them via layout.STEPS 22
swizzleXOR-based shared-memory address permutation to avoid bank conflicts.PRIMER, STEPS 22
scoreboard stallWarp waiting on a memory dependency tracked by L1TEX scoreboard (r41 attack).STEPS 41
MIO pipeMemory-IO instruction queue; shown not scarce (r36) — A-fragment reuse was.STEPS 36
coalescingWarp-wide global accesses touching contiguous lines; r21 coalesced A staging.STEPS 21
L2 window / access policy windowPinning a buffer's residency in L2 via cudaAccessPolicyWindow (MINFER_MMQ_L2WIN).STEPS 23
KSPLITSplitting the K reduction across blocks with atomic/partial adds (q6_K KSPLIT=2).STEPS 39
KS (k-step)K elements processed per inner iteration (KS=64 GEMM).STEPS 12
KD / KDRK-depth unroll factor and K-depth register pipeline depth (q6_K KDR=4, =8 regressed).STEPS 39-40
double bufferingOverlapping stage(n+1) loads with compute(n); the NB/BT staging pattern.STEPS 9, 45
software pipeliningRestructuring the loop so load/compute phases of different iterations overlap.STEPS 24
unroll (kd-loop)Compiler/pragma loop unrolling to expose ILP; r28 "NB kd-loop unroll".STEPS 28
epilogueThe post-mma tail: scaling, output store; r32 cut its cost.STEPS 32
B pre-format / quantize-transpose prepassRepacking B (weights) offline into kernel-friendly layout (r34).STEPS 34
MMVQ_PARAMETERS_GB10llama.cpp launch-config constant table for GB10 MMVQ, adopted by minfer.STEPS 12
block reduceWarp/block-wide reduction for logits accumulation (mmvq_block_reduce).STEPS 12
PDL / programmatic dependent launchOverlapping dependent kernel launch tails: cudaGridDependencySynchronize + programmatic stream serialization attribute.STEPS 74, PRIMER
griddepcontrolThe SASS/PTX-level instruction pair behind PDL.STEPS 74
wave (n)One full pass of all resident blocks; kernel iteration wave counting for sched analysis.STEPS 36
elect.syncWarp election intrinsic seen in SASS forensics.STEPS 65
MMQ_TILE_NE_K / MMQ_TILE_Y_Kllama.cpp MMQ shared-memory tile pitches (B tile 32+4 ints; y-tile row stride 36 ints = 144 B, the +4 avoids bank conflicts).MMQ

L5 — Platform & tooling

Term / FormulaOne-line meaningSource
GB10 / DGX SparkThe target machine: Grace 20-core ARM + Blackwell GPU, unified LPDDR5x, sm_121.PRIMER, BUILD
sm_XX / compute_XXGPU arch targets; build emits SASS for sm_70…sm_121 probes + PTX compute_70/72 for backward JIT.BUILD, PRIMER
nvcc / ptxasCUDA compiler and its SASS backend; ptxas -v gives register/spill counts.BUILD, STEPS
PTXVirtual ISA JIT-compiled at load; the backward-compatibility artifact.BUILD, PRIMER
SASSReal GPU assembly; forensic disassembly via cuobjdump/nvdisasm.STEPS 14+
-ccbin pinningForcing nvcc's host compiler when the default is rejected (MINFER_CUDA_CCBIN).BUILD
detect_archsbuild.rs probing sm_70…sm_121 by compiling a dummy .cu per arch.BUILD
libcuda_kernels.aStatic kernel archive nvcc produces, linked into the Rust binary.BUILD
cuda_static featureLink cudart statically so no libcudart.so is needed at runtime.BUILD
CUDA driver vs runtime APIlibcuda low-level vs libcudart convenience layer; minfer binds runtime via hand-written externs.PRIMER, BACKEND
CUDA GraphsCaptured kernel sequence replayed with one launch (cudaStreamBeginCapture/cudaGraphLaunch/cudaGraphInstantiate); MINFER_NO_CUDA_GRAPH to disable.BACKEND, STEPS 19
graph capture (prefill)Pre-capturing the prefill segment; default-ON since R3-B (MINFER_NO_PREFILL_CAPTURE=1 opts out).STEPS 57+
pinned memoryPage-locked host memory (cudaHostAlloc) for fast H2D/D2H; readback path has a kill switch.BACKEND, STEPS
cudaMemcpyAsync / streamsAsync copies on streams (cudaStreamCreate); decode uses graph launch, prefill streams.BACKEND
cudaMallocManaged / unified memoryMemory visible to both CPU and GPU (used once; avoided on GB10 due to bandwidth sharing).PRIMER
cudaFuncSetAttributeRuntime call to raise per-kernel dynamic smem limits.BACKEND
error guards (cudaGetLastError)Every launch checks the error; cudaErrorMisalignedAddress was a real campaign bug.STEPS 13, SAFETY
ncu (Nsight Compute)Kernel profiler: sm__warps_active, lts__t_sectors, smsp__inst_executed metric families.STEPS 12+
nsys (Nsight Systems)Timeline profiler for wall decomposition and graph-launch analysis.STEPS 57+
locked clocksnvidia-smi -lgc fixes GPU clocks so medians are comparable across runs.STEPS 77
MINFER_* env gates~40 kill-switch env vars (MINFER_MMQ_RAW, MINFER_MMQ_A_TRANSPOSE, MINFER_PDL, MINFER_FUSED_B, …) toggling one experiment at a time.HUB, STEPS
cuobjdump / nvdisasmTools producing the SASS listings used in forensics.STEPS 25
MINFER_TRACE / MINFER_GRAPH_DUMPPer-node real-data trace and graph dumps for viz tooling.GRAPH

L6 — Engine architecture

Term / FormulaOne-line meaningSource
ComputeGraph / CNodeThe declarative graph and its nodes; inference = build → assign → fuse → allocate → execute.GRAPH
Op enum / NodeMetaTyped operations and per-node metadata in graph/ops.rs.GRAPH
GraphBuilderDeterministic graph construction; identical GraphParams ⇒ identical topology.GRAPH
GraphParams / params-only reuseReuse check compares parameters only (GraphCache::try_reuse), never data.GRAPH
CParams.gpuParticipation flag recording whether the run used the GPU backend.GRAPH, BACKEND
Backend traitsupports_op/supports_fused, buffer pool, execute_node, host IO, synchronize — the CUDA backend is the worked example.GRAPH, BACKEND
GraphAllocator / livenessSingle buffer owner; allocates by liveness in build order; persistent KV regions survive rebuilds.GRAPH
kv_pair / persistent KV regionsEach layer owns two allocator regions (K/V) that outlive a decode step.GRAPH
scheduler (assign → split → execute)Assigns backends, splits the graph at backend boundaries, executes splits serially.GRAPH
split boundaryCross-backend sync/copy point; one Metal command buffer per split (CUDA: one stream).GRAPH
FusionPassBuild-time fusion of QKV (bias+rope+store) and FFN (swiglu); fused vs unfused bit-identical; gated by env.GRAPH
positions-as-dataRule 1: topology never depends on n_past — the precondition for decode reuse and CUDA graphs.GRAPH
fill_input_i32Integer inputs stored via f32::from_bits so the f32-typed input buffer carries token ids.GRAPH
in-place aliasing ruleSilu/RoPE alias their input (sole consumer + same backend); never host-copy a GPU-pending buffer.GRAPH, SAFETY
ModelDef traitPer-architecture forward/build_graph/forward_graph; models live in models/<name>/.GRAPH
weight layout conventionMetadata [in, out], memory row-major [out][in], activations token-major [nt][d].GRAPH
guard failure = abortKernel-invariant violations return Err from execute_node — never silent CPU fallback; guards print actual values.SAFETY
submit() bounded waitGPU submission waits bounded and checks status — never blocks forever.SAFETY
no early return past barrierMetal/CUDA rule: no exit path may skip a threadgroup_barrier/__syncthreads.SAFETY
prefill captureBackend feature storing the captured prefill graph for replay.BACKEND
IR (intermediate representation)The graph as an op-level IR; fusion makes fused ops first-class IR citizens.GRAPH
NodeId / DTypeNode handle and tensor data-type enum carried by every CNode.GRAPH
GetRowsRow-selection op: embedding lookup, and the n_out tail-row optimization (G3).GRAPH
BatchMatMulBatched matmul op (shared activation quantization, Q4_0); composable with fusion.GRAPH
FusedOp / supports_fusedFusion capability tag (currently only SwiGLU) checked by the fusion pass against each backend.GRAPH
n_out tail-row optimizationAfter the final wo, run FFN/norm/lm_head only on the tail n_out rows (llama inp_out_ids style); GraphParams.n_out joins the reuse decision.GRAPH
DOT / JSON exportgraph/dot.rs and graph/json.rs render the graph for viz.GRAPH

L7 — Methodology

Term / FormulaOne-line meaningSource
interleaved same-window A/BAlternating minfer/llama.cpp runs inside one time window so thermal/clock drift cancels.STEPS 77, HUB
median of NTake the median of repeated runs; means are corrupted by outliers on shared hardware.STEPS 77
pre-registered barThe success threshold is written into the step doc before measuring.STEPS 77
parity gateGreedy output must match llama.cpp (or CPU f64 ref within 0.05) before any perf number counts.HUB, STEPS 77-78
correctness batchDocs 78/79: batch verification sweeps over all kernels/quant types.STEPS 78-79
MEAS-ONLYStep status: measurement without landing code (e.g. doc 80).HUB
LANDED / REVERTEDStep outcome statuses in the hub tables.HUB
r-numbers (r1…r76)Experiment numbering across the MMQ/GEMM campaign; each step doc records one.HUB, MMQ
Era A/B/C/DCampaign phases: baseline (A), MMVQ (B), MMQ GEMM (C), decode/GEMM (D).HUB
Direction-A/B, Session A–FNamed experiment tracks within a phase (e.g. raw-nibble vs BT; quantize-fusion sessions).STEPS 28-56
SASS forensicsExplaining a perf delta by diffing disassembly (r25 opcode diff) instead of guessing.STEPS 25, 65
wall decompositionSplitting end-to-end time per kernel/phase before optimizing anything.STEPS 12, 47, 65
counter-guided iterationNext experiment chosen by the ncu counter that bounds the kernel (occupancy → wavefronts → IMMA rate).STEPS 36-73
one-variable-at-a-timeEach experiment toggles exactly one gate/env var; everything else frozen.STEPS 77
llama.cpp as reference$HOME/git/reading/llama.cpp is the ground truth for both parity and technique adoption.MMQ, SPEC
doc-per-step conventionEvery step writes one numbered record with a fixed six-section structure (STYLE.md).STEPS STYLE
measurement artifactsRaw bench JSONs kept under /tmp per step and reported, never committed.STEPS 77-80
G1 / G2 / G3Graph-refactor Phase-9 sub-experiment labels (attention dispatch, rms_norm_256, n_out tail-row).GRAPH