Test baselines — the recorded suite measurements
Why this page exists. AGENTS.md is the always-loaded agent index and has to stay
small, but these records are long and every device run makes them longer. They live here;
AGENTS.md links to them. How to run each command is in AGENTS.md ("Build & Run").
Machine-checked. scripts/test-baselines.toml
is the source of truth for every number on this page, and scripts/check_baselines.py --check
(CI job check-docs) fails when this prose disagrees with it — naming the file, the line and both
values. Edit that file, not a counter. Each record carries its date, device and command. Only the
x86_64 (CI runner) row is compared live, by the test-linux-cpu job, against the cargo log it
just collected (scripts/check_baselines.py --check-live); every other row
is a recorded measurement, refreshed only by a real run, because CI has no GPU and no
aarch64 runner.
The box label is absolute, never relative. dgxspark (aarch64, GB10 sm_121) is the
maintainer's DGX Spark — one physical machine, which is why both its CPU and CUDA rows carry
the same label — and x86_64 (CI runner) is GitHub's runner. Never write this box or
local in a record: the agent reading it may be on another machine. The convention itself is
stated once in gate contract rule 5. A record is written the moment it
is measured, in the same commit as the run it describes.
Records
- CPU unit, box
dgxspark (aarch64, GB10 sm_121),cargo test --release, 2026-10-07: 490 passed / 0 failed / 40 ignored unit + 10 / 0 / 6 integration. (The +4 over the 2026-10-06 measurement are #205's three fixture-manifest tests — the manifest reader, the outside-the-cache arm and the tampered-cache refusal — plus #349's parity-verdict classifier; the +1 on top is #306's format-aware physical-shift gate; the next +1 is #362's platform-neutralevery_device_gathers_the_attn_map, which moves the CPU rows together with the macOS ones; #354's two platform-neutral manifest-override tests — the resolver arm and the broken-override refusal — then the latest two, 488 → 490, measured ondgxspark2026-10-07. #310 adds one pure test of the packed routing matrix,packed_route_covers_the_matrix, but it does not move this row: it lives undersrc/metal/, which#[cfg(target_os = "macos")] mod metalexcludes from every non-macOS build. The aarch64 row is unchanged by #310, and a re-measurement ondgxspark2026-10-08 confirmed 490 / 0 / 40.) - CPU unit, box
x86_64 (CI runner),cargo test --release, 2026-10-08: 491 passed / 0 failed / 41 ignored unit + 10 / 0 / 6 integration. (This is the first direct measurement of the x86_64 suite: #56 adds threecfg(all(test, target_arch = "x86_64"))K-quant bitwise gates —avx2_q8k_dots_match_scalar_bitwise,avx512_q8k_dots_match_scalar_bitwise,dispatch_q8k_dots_match_scalar_bitwise— plus one#[ignore]d timing harnesskquant_simd_dot_speedup. The aarch64 row (490) still carries twoquants::neon_correctnesstests that x86_64 lacks, so the old "x86_64 = aarch64 − 2" passed-count relation becomes x86_64 = aarch64 − 2 (neon) + 3 (x86_64 K-quant gates) = aarch64 + 1 (491); the ignored count is now machine-dependent: x86_64 41 = aarch64 40 + the x86_64-only harness. The livecheck_baselines.py --check-livecomparison intest-linux-cpuis the authority for this row; #310 does not move it (its routing-matrix test lives under the macOS-gatedsrc/metal/).) - CPU real-model set, box
dgxspark (aarch64, GB10 sm_121),PARALLEL=0 scripts/real_model_gates.shand the default parallel form, 2026-10-06: 40 / 0 each (the device-gated members of the set skip on a CPU build, so it counts 40 of the CUDA set's 46). - CUDA unit, box
dgxspark (aarch64, GB10 sm_121),scripts/cuda_test.sh, 2026-10-07: 576 / 0 / 46. (The row read 526 before #162; #153 adds two device unit gates plus one purekvformatgate and moves the twocuda::kv_dtype_teststo the tag mapping they now assert — 536/37 → 539/38; #144 adds the packed fused-epilogue and packed FA-prefill device gates — 539/38 → 541/38; #185 adds the device-path guard's test, the explicit-auto-budget test and the per-backend stream-sync counter gate — 541/38 → 544/38; #188 adds the capture probe and the concurrent two-engine gate — 544/38 → 545/39; #189 adds one pure test of the S4 A/B's sign-test statistic off the recorded loaded distribution — 545/39 → 546/39; #196 adds the two wedge-proofserve_loopstall gates — 546/39 → 548/39; #140 adds the K-quant encoder tests (7 passed / 1 ignored) — 548/39 → 555/40; #142 adds the bf16 writer + CPU-path tests (7 passed / 2 ignored) — 555/40 → 562/42; #202's three launch sites add no test, so the row is unchanged by it; #218 adds the two fresh-process prefill-GEMM smem opt-in gates (cuda::issue218_tests) and the captured-graph invariant arm — 562/42 → 565/42; #223 adds the fresh-process eager pre-warm gate (cuda::issue223_tests), which asserts the runtime guarantee before any launch — 565/42 → 566/42. #240/#241 deletes the dead legacylayer_gpu/device_entrywrapper layer and, with it, the one feature-independent guard test (the guard's only caller was dead), so 566/42 → 565/42 and the CPU rows lose that same test (481 → 480 on aarch64, 479 → 478 in CI). #138 adds one feature-independent allocator drain gate plus the device multi-input deferred-wait gate — 565/42 → 567/42, and the CPU rows 480 → 481 on aarch64 and 478 → 479 in CI. #208's CUDA half adds two non-ignored device gates (cuda_backend::tests::weights::cuda_bf16_matmul_matches_the_exact_shift_referenceand..._cuda_bf16_embed_gather_matches_the_exact_shift_reference) and one ignored device gate (f208_bf16_weights_run_on_the_cuda_device) — 567/42 → 570/45, the CPU rows 481/36 → 482/39 on aarch64 and 479/38 → 480/39 in CI, and each CUDA real-model set 42 → 45. #208's Metal half adds the two macOS-only Metal kernel-exactness gates and one ignored real-model gate (f208_bf16_weights_run_on_the_metal_device, a no-op off macOS, so it merely raises the ignored count) — 570/45 → 570/46, the CPU rows 482/39 → 482/40 on aarch64 and 480/39 → 480/40 in CI, and each real-model set one higher; #205's three fixture-manifest gates and #349's parity-verdict classifier then add four feature-independent tests (570/46 → 574/46) and #306's format-aware physical-shift gate the fifth — 574/46 → 575/46; #362's platform-neutralevery_device_gathers_the_attn_mapthe sixth — 575/46 → 576/46, re-measured ondgxspark2026-10-07.) The row is a recorded measurement and carries a #207 projection:docs/status.tomlnames the CPU row on the same box that it shares a test binary with, andscripts/check_status.py --checkprints — never fails on — the projected count when that row has moved since this one was measured.) - CUDA real-model set (0.5B config), box
dgxspark (aarch64, GB10 sm_121),FEATURES=cuda scripts/real_model_gates.sh, 2026-10-06: 46 / 0. - CUDA real-model set (Qwen3-0.6B config), box
dgxspark (aarch64, GB10 sm_121),FEATURES=cuda scripts/real_model_gates.sh, 2026-10-06: 46 / 0. compute-sanitizer --tool memcheckover the CUDA unit suite, boxdgxspark (aarch64, GB10 sm_121),scripts/cuda_test.shunder the sanitizer, 2026-10-01: 0 API errors (565 / 0 / 42).- macOS unit, box
macbook (macOS 27.0.1, Apple M4 Pro),cargo test --release --no-fail-fast, 2026-10-08: 553 passed / 0 failed / 45 ignored unit + 21 / 0 / 6 integration. (The +3 over the previous 550 are #310's new gates — the purepacked_route_covers_the_matrix(the packed routing matrix; it lives under the macOS-gatedsrc/metal/, so this row is the only one it moves) plus the macOS-onlymetal_packed_decode_stage_matches_the_native_read(the mechanism-A/B differential) andmetal_packed_window_stages_at_the_absolute_row(the absolute-row staging arm the zero-lo_mingates cannot see); [#310] is now enabled (READS_PACKED_KV = true, soMINFER_CACHE_TYPE=q8_0loads and runs on Metal — mechanism A reads packed cells natively in the decode flash family, mechanism B stages the window to an f32 scratch buffer and runs the unchanged f32 prefill/window families); the earlier +5 are [#310]'s packed-Q8_0 KV gates ingraph::metal_backend::tests::packed_kv— the byte-identical store, the causal decode/prefill, the one-rangeattn_spanand thekv_mapwindow, and the C5FLAG_PACKEDround trip — plus the +1 ignored real-model gate, so the ignored row moved 44 → 45; the +1 under the 542 baseline is #369's fast-map gatewindow_map_flash_matches_the_cpu_reference; the +1 under that is #359's fast-windowed-kernel gatewindow_flash_matches_the_cpu_reference; the +3 under that are #362's two Metalkv_mapwindow gates plus the pureevery_device_gathers_the_attn_map; the earlier +4 are #205's three fixture-manifest tests plus #349's parity-verdict classifier, and the +1 on top is #306's format-aware physical-shift gate — the +3 added from a Mac, the rest fromdgxsparkand platform-neutral. #354's two platform-neutral manifest-override tests are derived here (543 → 545), not re-measured on a Mac. The +2 over the #54 baseline are #329's dispatch-refusal gate and #299's weights-charged E4 gate; the +1 ignored is #315's macOS-only#[ignore]d windowed-prefill measurement harness, so the passed count it changes is not this one. The macOS suite was red for the whole Metal round; this is the baseline a later Mac run diffs against —docs/METAL-BACKEND-DESIGN.md§7.4 — and it is green because the round fixed its three production bugs. Not run in CI:build-macoscompiles the test target since #303 but has no Metal device.) - macOS real-model set, box
macbook (macOS 27.0.1, Apple M4 Pro),PARALLEL=0 scripts/real_model_gates.sh(serial), 2026-10-08: 45 / 0 on both cached models — the set is 45 because #310's macOS-only#[ignore]dmetal_q8_0_kv_answers_like_f32_on_a_real_modelgate joined the--ignoredset (44 → 45), after #315's gate had taken it 42 → 43 (since #310 enabled Metal's packed read, the gate drives the production capability directly, no test seam), and the previously failingserver::batch::tests::kv_sharing::a_store_inside_a_shared_prefix_takes_a_private_rownow passes: Metal gathers the set-valuedkv_mapwindow since #362, so its shared-prefix slot reads the donor's rows in place instead of copying them (recorded indocs/METAL-BACKEND-DESIGN.md§7.4). No residual.