Test baselines — the recorded suite measurements

Why this page exists. AGENTS.md is the always-loaded agent index and has to stay small, but these records are long and every device run makes them longer. They live here; AGENTS.md links to them. How to run each command is in AGENTS.md ("Build & Run").

Machine-checked. scripts/test-baselines.toml is the source of truth for every number on this page, and scripts/check_baselines.py --check (CI job check-docs) fails when this prose disagrees with it — naming the file, the line and both values. Edit that file, not a counter. Each record carries its date, device and command. Only the x86_64 (CI runner) row is compared live, by the test-linux-cpu job, against the cargo log it just collected (scripts/check_baselines.py --check-live); every other row is a recorded measurement, refreshed only by a real run, because CI has no GPU and no aarch64 runner.

The box label is absolute, never relative. dgxspark (aarch64, GB10 sm_121) is the maintainer's DGX Spark — one physical machine, which is why both its CPU and CUDA rows carry the same label — and x86_64 (CI runner) is GitHub's runner. Never write this box or local in a record: the agent reading it may be on another machine. The convention itself is stated once in gate contract rule 5. A record is written the moment it is measured, in the same commit as the run it describes.

Records

  • CPU unit, box dgxspark (aarch64, GB10 sm_121), cargo test --release, 2026-10-07: 490 passed / 0 failed / 40 ignored unit + 10 / 0 / 6 integration. (The +4 over the 2026-10-06 measurement are #205's three fixture-manifest tests — the manifest reader, the outside-the-cache arm and the tampered-cache refusal — plus #349's parity-verdict classifier; the +1 on top is #306's format-aware physical-shift gate; the next +1 is #362's platform-neutral every_device_gathers_the_attn_map, which moves the CPU rows together with the macOS ones; #354's two platform-neutral manifest-override tests — the resolver arm and the broken-override refusal — then the latest two, 488 → 490, measured on dgxspark 2026-10-07. #310 adds one pure test of the packed routing matrix, packed_route_covers_the_matrix, but it does not move this row: it lives under src/metal/, which #[cfg(target_os = "macos")] mod metal excludes from every non-macOS build. The aarch64 row is unchanged by #310, and a re-measurement on dgxspark 2026-10-08 confirmed 490 / 0 / 40.)
  • CPU unit, box x86_64 (CI runner), cargo test --release, 2026-10-08: 491 passed / 0 failed / 41 ignored unit + 10 / 0 / 6 integration. (This is the first direct measurement of the x86_64 suite: #56 adds three cfg(all(test, target_arch = "x86_64")) K-quant bitwise gates — avx2_q8k_dots_match_scalar_bitwise, avx512_q8k_dots_match_scalar_bitwise, dispatch_q8k_dots_match_scalar_bitwise — plus one #[ignore]d timing harness kquant_simd_dot_speedup. The aarch64 row (490) still carries two quants::neon_correctness tests that x86_64 lacks, so the old "x86_64 = aarch64 − 2" passed-count relation becomes x86_64 = aarch64 − 2 (neon) + 3 (x86_64 K-quant gates) = aarch64 + 1 (491); the ignored count is now machine-dependent: x86_64 41 = aarch64 40 + the x86_64-only harness. The live check_baselines.py --check-live comparison in test-linux-cpu is the authority for this row; #310 does not move it (its routing-matrix test lives under the macOS-gated src/metal/).)
  • CPU real-model set, box dgxspark (aarch64, GB10 sm_121), PARALLEL=0 scripts/real_model_gates.sh and the default parallel form, 2026-10-06: 40 / 0 each (the device-gated members of the set skip on a CPU build, so it counts 40 of the CUDA set's 46).
  • CUDA unit, box dgxspark (aarch64, GB10 sm_121), scripts/cuda_test.sh, 2026-10-07: 576 / 0 / 46. (The row read 526 before #162; #153 adds two device unit gates plus one pure kvformat gate and moves the two cuda::kv_dtype_tests to the tag mapping they now assert — 536/37 → 539/38; #144 adds the packed fused-epilogue and packed FA-prefill device gates — 539/38 → 541/38; #185 adds the device-path guard's test, the explicit-auto-budget test and the per-backend stream-sync counter gate — 541/38 → 544/38; #188 adds the capture probe and the concurrent two-engine gate — 544/38 → 545/39; #189 adds one pure test of the S4 A/B's sign-test statistic off the recorded loaded distribution — 545/39 → 546/39; #196 adds the two wedge-proof serve_loop stall gates — 546/39 → 548/39; #140 adds the K-quant encoder tests (7 passed / 1 ignored) — 548/39 → 555/40; #142 adds the bf16 writer + CPU-path tests (7 passed / 2 ignored) — 555/40 → 562/42; #202's three launch sites add no test, so the row is unchanged by it; #218 adds the two fresh-process prefill-GEMM smem opt-in gates (cuda::issue218_tests) and the captured-graph invariant arm — 562/42 → 565/42; #223 adds the fresh-process eager pre-warm gate (cuda::issue223_tests), which asserts the runtime guarantee before any launch — 565/42 → 566/42. #240/#241 deletes the dead legacy layer_gpu/device_entry wrapper layer and, with it, the one feature-independent guard test (the guard's only caller was dead), so 566/42 → 565/42 and the CPU rows lose that same test (481 → 480 on aarch64, 479 → 478 in CI). #138 adds one feature-independent allocator drain gate plus the device multi-input deferred-wait gate — 565/42 → 567/42, and the CPU rows 480 → 481 on aarch64 and 478 → 479 in CI. #208's CUDA half adds two non-ignored device gates (cuda_backend::tests::weights::cuda_bf16_matmul_matches_the_exact_shift_reference and ..._cuda_bf16_embed_gather_matches_the_exact_shift_reference) and one ignored device gate (f208_bf16_weights_run_on_the_cuda_device) — 567/42 → 570/45, the CPU rows 481/36 → 482/39 on aarch64 and 479/38 → 480/39 in CI, and each CUDA real-model set 42 → 45. #208's Metal half adds the two macOS-only Metal kernel-exactness gates and one ignored real-model gate (f208_bf16_weights_run_on_the_metal_device, a no-op off macOS, so it merely raises the ignored count) — 570/45 → 570/46, the CPU rows 482/39 → 482/40 on aarch64 and 480/39 → 480/40 in CI, and each real-model set one higher; #205's three fixture-manifest gates and #349's parity-verdict classifier then add four feature-independent tests (570/46 → 574/46) and #306's format-aware physical-shift gate the fifth — 574/46 → 575/46; #362's platform-neutral every_device_gathers_the_attn_map the sixth — 575/46 → 576/46, re-measured on dgxspark 2026-10-07.) The row is a recorded measurement and carries a #207 projection: docs/status.toml names the CPU row on the same box that it shares a test binary with, and scripts/check_status.py --check prints — never fails on — the projected count when that row has moved since this one was measured.)
  • CUDA real-model set (0.5B config), box dgxspark (aarch64, GB10 sm_121), FEATURES=cuda scripts/real_model_gates.sh, 2026-10-06: 46 / 0.
  • CUDA real-model set (Qwen3-0.6B config), box dgxspark (aarch64, GB10 sm_121), FEATURES=cuda scripts/real_model_gates.sh, 2026-10-06: 46 / 0.
  • compute-sanitizer --tool memcheck over the CUDA unit suite, box dgxspark (aarch64, GB10 sm_121), scripts/cuda_test.sh under the sanitizer, 2026-10-01: 0 API errors (565 / 0 / 42).
  • macOS unit, box macbook (macOS 27.0.1, Apple M4 Pro), cargo test --release --no-fail-fast, 2026-10-08: 553 passed / 0 failed / 45 ignored unit + 21 / 0 / 6 integration. (The +3 over the previous 550 are #310's new gates — the pure packed_route_covers_the_matrix (the packed routing matrix; it lives under the macOS-gated src/metal/, so this row is the only one it moves) plus the macOS-only metal_packed_decode_stage_matches_the_native_read (the mechanism-A/B differential) and metal_packed_window_stages_at_the_absolute_row (the absolute-row staging arm the zero-lo_min gates cannot see); [#310] is now enabled (READS_PACKED_KV = true, so MINFER_CACHE_TYPE=q8_0 loads and runs on Metal — mechanism A reads packed cells natively in the decode flash family, mechanism B stages the window to an f32 scratch buffer and runs the unchanged f32 prefill/window families); the earlier +5 are [#310]'s packed-Q8_0 KV gates in graph::metal_backend::tests::packed_kv — the byte-identical store, the causal decode/prefill, the one-range attn_span and the kv_map window, and the C5 FLAG_PACKED round trip — plus the +1 ignored real-model gate, so the ignored row moved 44 → 45; the +1 under the 542 baseline is #369's fast-map gate window_map_flash_matches_the_cpu_reference; the +1 under that is #359's fast-windowed-kernel gate window_flash_matches_the_cpu_reference; the +3 under that are #362's two Metal kv_map window gates plus the pure every_device_gathers_the_attn_map; the earlier +4 are #205's three fixture-manifest tests plus #349's parity-verdict classifier, and the +1 on top is #306's format-aware physical-shift gate — the +3 added from a Mac, the rest from dgxspark and platform-neutral. #354's two platform-neutral manifest-override tests are derived here (543 → 545), not re-measured on a Mac. The +2 over the #54 baseline are #329's dispatch-refusal gate and #299's weights-charged E4 gate; the +1 ignored is #315's macOS-only #[ignore]d windowed-prefill measurement harness, so the passed count it changes is not this one. The macOS suite was red for the whole Metal round; this is the baseline a later Mac run diffs against — docs/METAL-BACKEND-DESIGN.md §7.4 — and it is green because the round fixed its three production bugs. Not run in CI: build-macos compiles the test target since #303 but has no Metal device.)
  • macOS real-model set, box macbook (macOS 27.0.1, Apple M4 Pro), PARALLEL=0 scripts/real_model_gates.sh (serial), 2026-10-08: 45 / 0 on both cached models — the set is 45 because #310's macOS-only #[ignore]d metal_q8_0_kv_answers_like_f32_on_a_real_model gate joined the --ignored set (44 → 45), after #315's gate had taken it 42 → 43 (since #310 enabled Metal's packed read, the gate drives the production capability directly, no test seam), and the previously failing server::batch::tests::kv_sharing::a_store_inside_a_shared_prefix_takes_a_private_row now passes: Metal gathers the set-valued kv_map window since #362, so its shared-prefix slot reads the donor's rows in place instead of copying them (recorded in docs/METAL-BACKEND-DESIGN.md §7.4). No residual.