| 83 | d5-r-stage1-spec-loop | the greedy d=2 loop (--spec-draft): second GraphCache, lazy accept loop (unit-tested), namespaced weight registries + nb_bt_only global-mix fix — the three single-model assumptions a second model breaks; 14B d=2 = 1.34×/1.58× | 🟢 |
| 84 | d5-r-stage2-dual-engine-battery | same-window dual-engine protocol (3 reps × prose/code × 4 cells): minfer 1.33×/1.59× vs llama 1.64×/2.08× — the whole gap = verify row marginal (8.8 vs 2.5 ms/row) | 🟢 |
| 85 | d5-r-stage3-verify-marginal-ledger | nsys per-kernel ledger: nt=3 marginal 17.6 ms = attention nt 2–63 hole 9.0 (legacy per-(token,head) kernel vs the 0.8 ms split path) + matmul 8.1 + elt 1.9 + idle 0.7; nt=9 = dispatch cliff onto padded GEMM | 📏 |
| 86 | d5-r-stage4a-attention-verify-shapes | one gate: fa_prefill nt≥64 → nt≥2 — C_T(3) 56.9→48.8, C_T(9) 101.3→86.4; e2e 1.42×/1.68× (code ≥ llama's same-window 1.64×); ledger projection validated ~5% | 🟢 |
| 87 | d5-r-stage4b-multi-mmvq-nt16-closed | multi-MMVQ nt 9–16: groups-of-8 = parity (weights re-streamed per group), acc[16] = register spill (111 ms) → doc-82 GEMM boundary stands; d=8 retired (0.71× prose projected, 0.63× measured in doc 88) | 🔴 |
| 88 | d5-r-stage5-final-battery | final battery: minfer d=2 1.42×/1.68× (35.7/42.5 tok/s) = 95%/88% of llama's absolute speed; capture prize verified already banked (R3-B); D5-R closes | 🟢 |
| 89 | d5-r-row-marginal-localization | the absolute-gap leader localized without ncu: cold-L2 real-kernel bench + chain nsys + ablation — q4_K 2.4 / q6_K 1.0 / norm-quant 0.7 ms/row; ~half the matmul term = block-per-row activation re-read; fix menu priced (R-rows-per-block, small-M mma, chain hygiene) | 📏 |
| 90 | d5-r-rrows-per-block-closed | menu item 1 implemented → measured → reverted: ~0 at d=2 (chain keeps act rows L2-hot; block-parallel latency hiding dominates utilization); nt≥6 flatten recorded; small-M mma re-confirmed as the only lever of size | 🔴 |
| 91 | mma-path-block-starvation | BT GEMM already mma.m16n8k32; small-M floor root-caused to block starvation (ntb=1 → 40 blocks); conditional double-buffer shipped (bitwise-safe, prefill guarded); K-split designed as the fix that revives d=8 | 🔬 |
| 92 | ksplit-shipped-flip-resolved-no | K-split (grid.z + deterministic reduce) shipped for both BT kernels behind the gate; C_T(9) 86.4→72.9 but the ≤55 flip condition failed — multi-MMVQ stays production; residual = per-tile staging serialization; nt 9..64 auto-ksplit → §3b: enabled on the default path by user decision (default C_T(9) 73.0; d=8 still acceptance-bound) | ✅ |
| 93 | draft-quant-and-greedy-identity | draft-quant swap is a mixed knob (±3 pts acceptance, opposite signs per cell); greedy identity test FAILS — spec ≠ sequential, flips traced to batched verify attention/softmax; nt-invariance campaign proposed with the identity test as acceptance criterion | 🔬 |
| 94 | greedy-identity-and-d8-crossing | greedy identity achieved (spec = sequential byte-for-byte; attention nt-invariance + penalty-window cap); d=8 crosses on code with q4_k_m (46.6 tok/s, 96% of llama); prose stays d=2 | ✅ |
| 95 | adaptive-draft-depth | adaptive per-round draft depth (beta acceptance + min-window costs + 10% hysteresis + optimism for unseen depths); identity boundary pinned — verify nt ≤ 8 bitwise (multi-MMVQ), nt=9 BT lm_head tolerance-class → adaptive capped at d=7; all four gates pass; adaptive BEATS the best static on both code cells (46.5/48.4 vs 43.9/46.6 tok/s) — the buggy static sweep had never measured d=3..7 | ✅ |
| 96 | nt9-profile-phase0 | nt=9 verify profiled (stop-gate measured): BT-MMQ = 84% of GPU time, attention ~1%; ncu: both BT kernels at ~20% of both roofs, smem scoreboard stalls = 40–59% of warp cycles → doc 92's staging serialization confirmed dominant; gate verdict PROCEED; cp.async double-buffer staged as the bitwise-preserving Phase-1 lever (narrow EV: prefill lever / identity-relaxed d=8, not spec throughput) — Phase 1 re-priced separately, below 战役 97 | ✅ P0 |
| 97 | spec-conversation-server | speculative decoding in --cnv and serve (Engine trait hooks + SpecAwareEngine + sibling spec loops mirroring the plain decode token-for-token); server position/termination contract (seed carry, mid-batch stop/EOG ends the turn); pre-existing plain-server bug fixed: cross-request slot GraphCache reuse leaked stale KV (identical requests hid it) → per-request cache reset | ✅ |
| 98 | bt-cpasync-null | 96 Phase 1 measured: cp.async double-buffered staging brought to the q4_K BT kernel (full r56 treatment, dbuf extended to ksplit) → NULL on GB10 (kernel µs / C_T / pp512 all baseline-within-noise; dbuf on/off identical) — the BT stall is compute-side (ldmatrix→mma chains), not staging; remaining levers are tolerance-class → patch reverted per doc-90 discipline, D5-R closed at its identity-safe ceiling | ⚫ |
| 99 | fastverify-p0-p1 | fast-verify P0/P1: doc 98's tolerance-class claim corrected (int mma exact → wider bitwise-safe set); fragment prefetch / non-volatile mma / cp.async all measured NULL; B-plane XOR swizzle landed (shared wavefront excess 40%→0, −4.3% kernel instance, bitwise 4/4); pc-sampling re-attributes the stall to L1TEX latency × 16-warp occupancy ceiling → knob not built, EV re-priced down | ✅ |
| 100 | qs-plane-drift | the last priced lever (aligned qs plane) measured timing-NULL under interleaved A/B; sequential "−12.4%" was clock-ramp drift (±7–12% band, 208 MHz idle → 3 GHz) — sequential before/after runs on dgxspark are invalid instruments; repack family closed, D5-R fully closed | ⚫ |
| 101 | steady-state-method | doc 100's rule made executable: time-budget warmups in specverify/bench; headline table re-measured tight (pp512 2083 ± 7.2; adaptive 35.9 prose / 44.8 code; d8 collapse reproduced) — doc-95 absolutes confirmed as drift artifacts, structure exact | ✅ |
| 102 | draft-scale-sweep | bigger drafts lose (acceptance bounded by the target, not draft capacity) — default draft stays 0.5B Q4_K_M; flushed + fixed a latent qwen3-loader namespaced-registration bug (Q6_K padded weights under raw name → both models to CPU); post-EOS token-text gate relaxed to warning, cross-family identity 4/4 | ✅ fix + null |
| 103 | q40-q80-mmvq | q4_0/q8_0 decode joins the MMVQ family (8e structure, NEW CODE ONLY — every landed K-quant kernel/arm untouched, size-floored gates + MINFER_NO_Q40_MMVQ/MINFER_NO_Q80_MMVQ fallbacks): 7B q4_0 tg128 48.1→56.8 (+18.1%), 7B q8_0 26.7→28.8 (+7.9%), 7B q8_0 spec e2e 26.8→68.0 tok/s (+154%); same-file llama.cpp comparison — q4_0 decode 1.045× faster than llama.cpp, q8_0 closed 0.87×→0.94×; pp512 unchanged ✓, suite 187/0/3, q8_0 identity 4/4 | ✅ |
| 104 | q80-p32-split-plane | nsys located doc 103's remaining q8_0 gap inside the kernel (97% of decode GPU time; 16 scattered 2B weight loads per block at a 34B lane stride = ~2x q4_0's L1TEX wavefront cost per byte) → p32 split planes (payload 32B/block 16B-aligned for uint4 x2 + dense 2B d plane, traffic unchanged, raw registration untouched, byte-equal outputs): 7B Q8_0 tg128 28.5→31.9-32.1 (+12.3% over f32) = parity with llama.cpp (0.87x→0.94x→1.00x), spec e2e 69.1→79.0 tok/s = 2.47x sequential; q4_0-p32 and 36B-pad variants measured and rejected | ✅ |
| 105 | device-tier-tables | T-series T1: cc-keyed device tier table + selector (device_tier.rs, pure data, offline-tested) — GB10 Measured, consumer rows Adopted from llama.cpp (Blackwell K-quant caps 5/6/7, Orin K-quants→1, Turing mmq=false ruling #4), GENERIC fallback; MINFER_DEVICE_TIER override; mmq gate = resolved tier flag; batch caps tabled but unwired (R8). Encoding fix: runtime cc is 1201, table keys llama-encoded via conversion | ✅ |
| 106 | query-formula-gates | T-series T2: auto-ksplit SM-count parameterization (target max(256,2*SM), GB10-invariant), BT smem feasibility gate (single-source C formula + cuda_mmq_smem_bytes(); R2 fixed — externs now query the selected device, not device 0), plane VRAM budget gate on all optional planes (p32/W_exp/W_dsc) for 8 GB unified-memory devices. Tile candidates + cap activation + T3 deferred per plan | ✅ |
| 107 | c4-packed-q8-kv-cuda | C4 #144: the packed Q8_0 cache's two missing tuned routes — the packed fused decode epilogue (attn_bias_rope_store_q8_0, one thread per (head, 32-element K block) and per V block, the store's own quantizer) and the packed FA prefill (fa_prefill_kv<CAUSAL,MAP,LAYOUT> dequantizes each block into the same f16 tile; general kernel stays the fallback). Qwen3-0.6B pp2048 564.5 → 8231.1 tok/s (14.58x; packed/f16 14.7x → 1.038x); 0.5B tg128 q8_0 161.5 → 170.1 (1.052x, the pre-registered 1.15x bar missed — the ticket's 1.18x was the f16-weight arm's cut). Two new device gates, three mutations, 124 launch sites | 🟢 |
| 108 | c4-dp4a-packed-q8-kv-cuda | #186: the dp4a packed Q8_0 K dot — the decode split-K body accumulates int (__dp4a) against a per-(head, block)-quantized query and scales by d_q*d_k once per block; K is never converted to float (V still is). Bar named first (load-attributable share ≥ 10%), measured 20.3% on the 0.5B's hd-64 decode where both layouts share one 1-warp geometry; landed at 0.5B tg128 171.95 → 193.30 (1.124x, packed/f16 1.396x → 1.242x) and Qwen3-0.6B tg128 123.98 → 136.66 (parity with its f16 136.73); pp2048 flat. Real-model class re-measured (tail 2.479504, argmax 0.5527, greedy 9/9); prefill/verify left alone by measurement | 🟢 |
| 109 | c4-packed-q8-kv-l1-request | #202: the packed Q8_0 KV cell's L1 request count, no layout change — the four s8 K/V quant loads per 4 elements become two u16 loads (34k + 2 + 4m is always 2-byte aligned even when only odd blocks are 4-byte aligned). The counter moves exactly as the ticket's mechanism predicted — packed/f16 L1 load-sector ratio 1.712x → 0.9845x (792 904 → 455 840), instructions −2.17% — but the pre-registered +2% tg128 bar was NOT cleared (+0.27%: 193.13 → 193.65, medians of 5 interleaved rounds) and the kernel is only 1.4-1.7% faster (nsys). A partial refutation: the 1.23x packed/f16 decode residual is not L1-request-bound (not L2: 0.56x, not instructions: 1.10x, not sectors: 0.98x). Landed counter-only, byte-identical, no layout/session/CPU change; the latency hypothesis is filed for the next attribution | 🟢 counter-only |