F4 — Backend registry (design record)

Ticket: #57 ("[F4] Backend registry instead of the compile-time enum"). This document is the design-first artifact: it is committed before the implementation and states the registry's shape, the name surface, the ordering rule and the exact refusals. Gaps and the backlog item live in docs/ARCHITECTURE-ROADMAP.md; the ticket record lives in docs/ARCHITECTURE-EXECUTION-PLAN.md.

1. Why the enum was the problem

Backend was enum Backend { CPU, Metal, Cuda } in src/graph/mod.rs, and every consumer matched on it. GraphAllocator alone dispatched twelve operations through match backend { … }; BackendScheduler::execute matched to pick the executing pool and again for the MINFER_TRACE capture path; and the fusion wiring built a Vec<&dyn Backend> by hand and found a node's backend with a position(|b| b.name() == "cuda") lookup that existed only because the match could not express "whichever device is present". Adding a backend therefore meant editing the allocator, the scheduler, the fusion wiring, the exporters, the KV-session tag table and the op matrix; and none of it told a user anything — there was no way to ask for a backend by name or to learn that a requested one did not exist.

That argument is a decision, and it is frozen in ADR-0001 (one device seam, build-time assignment) and ADR-0011 (the id space and the name surface). This page keeps the shape that replaced it — the handle, the registered set, the ordering rule and the refusals.

2. The handle: Backend(u16)

Backend stays a cheap, Copy, Hash, Eq, Ord value; it stops being an enum. It is an opaque index into a fixed id space:

#![allow(unused)]
fn main() {
pub struct Backend(u16);

impl Backend {
    pub const CPU:   Backend = Backend(0);
    pub const METAL: Backend = Backend(1);
    pub const CUDA:  Backend = Backend(2);
}
}

The id space is compile-time and configuration-independent: cpu = 0, metal = 1, cuda = 2 on every build, whether or not the backend is compiled in. That is deliberate — the id is already on disk and in exports:

  • the KV-session backend tag (graph/kvsession.rs, tag_of / backend_of_tag) is a u32 in a versioned file, so the numbering is a file-format contract;
  • graph/json.rs / graph/dot.rs name and index the backend in exported graph documents;
  • Ord is derived from the id, and the pre-F4 derived Ord on the enum was declaration order — CPU < Metal < Cuda.

Backend::index(), Backend::from_index() and Backend::name() are the only ways to leave and re-enter the id space; kvsession's tag functions and the fusion index are now one-line calls to them instead of three-arm matches.

Debug is implemented by hand so diagnostics print exactly what the enum printed (CPU, Metal, Cuda) — no log line or error message changes shape.

2.1 Two orders, stated once

The registry has two orders and they are not the same; conflating them is the failure this document exists to prevent.

orderwhat it fixesauthority
identity (Backend::index())the on-disk KV-session tag, the exported graph's backend index, Ord, the fusion pass's backend listBackend::CPU/METAL/CUDA
priority (assignment preference)which backend supports_for offers firstBackendEntry::priority, descending

Before F4 the priority order was implicit in the statement order of GraphAllocator::supports_for: Metal, then CUDA, then CPU. It is preserved exactly, and it is now a number on each entry — Metal 300, CUDA 200, CPU 100 — so the ordering gate can pin it and a future accidental reordering fails a test instead of silently changing where a graph's nodes land.

Determinism rule. The assignment pass fixes the graph's topology, which is part of GraphParams' reuse identity (standing rule 3). Nothing in the registry may depend on HashMap iteration order:

  • the registry table is a fixed-size array indexed by the handle id — not a HashMap;
  • Registry::by_priority() sorts by (Reverse(priority), index), so two entries can never tie and the result is a total order derived from the pinned numbers;
  • BackendFilter is a [bool; N_BACKENDS] indexed by id, never a set that is iterated;
  • the fusion-pass backend list is built by walking the identity order, so the index it hands the pass is Backend::index() on every run.

3. The registry table

One BackendEntry per backend, built once at startup (registry(), a OnceLock, is forced by main before anything else runs):

#![allow(unused)]
fn main() {
pub struct BackendCaps {
    pub reads_packed_kv: bool,          // the #87 seam, §8 — the only field
}

pub struct BackendEntry {
    pub handle: Backend,
    pub name: &'static str,             // "cpu" | "metal" | "cuda"
    pub priority: u16,
    pub caps: BackendCaps,
    pub pool:     fn(&GraphAllocator) -> Option<&dyn BackendTrait>,
    pub pool_mut: fn(&mut GraphAllocator) -> Option<&mut dyn BackendTrait>,
    pub host_read: fn(&GraphAllocator, usize) -> Option<Vec<f32>>,
    // F5 (#58): the split boundary's two phases (§11).
    pub copy_cross:  fn(&mut GraphAllocator, u64, NodeId, Backend) -> Result<bool, String>,
    pub await_cross: fn(&mut GraphAllocator, u64, NodeId, Backend) -> Result<(), String>,
    pub kv_format: fn(&GraphAllocator) -> KvFormat,
    pub enable:    fn(&mut GraphAllocator) -> bool,
    pub unavailable: fn() -> Option<&'static str>,
}
}

Each backend module owns its entry and its hooks (cpu_backend::entry(), metal_backend::entry(), cuda_backend::entry()), and Registry::build() calls their register() — that call site is the only place a backend is introduced.

The capability authority is the module-level item, stated once (#244, 2026-10-01). Each backend defines its op×dtype matrix and its fusion matrix as module-level free functions (cpu_backend::supports_op, cuda_backend::supports_fused, …) and its attn_span answer as a module-level constant (cpu_backend::SUPPORTS_ATTN_SPAN, …). The impl Backend for X methods are one-line forwards to them, and the assignment pass reads the trait method (graph::backend_takes → dyn Backend::supports_op). The registry carries no copy of those three answers. The supports_op / supports_fused / supports_attn_span fields that used to sit here were written by every entry() and read only by registry::tests — a mirror with no production reader, which is why #244 deleted them (option (b) of that ticket's escalation). The gate that used to compare the field against the trait now compares the module-level function against the trait (registry::tests::registry_caps_match_the_backend_trait), i.e. the authority the field merely mirrored; putting the two on one line is no longer possible even in principle.

BackendCaps therefore carries exactly one field, reads_packed_kv, and that one is read in production because the question is asked about a format, not an engine: GraphAllocator::ensure_kv's packed-region refusal and KvFormat::supports call registry::reads_packed_kv(backend), where no &self is available (§8). It is the one capability the registry itself carries.

pool / pool_mut are why "sync" and "copy" are not a match any more: the allocator's dispatch helpers ask the entry for the pool and then call the Backend trait method (synchronize, copy_cells, write_host, …) on it. The two backends that have a native host-read path different from the trait's borrowed read_host (CUDA's stream-ordered copy_to_host) say so in host_read, which is why copy_to_cpu and the CUDA KV debug read all collapse to one call without a match.

enable is the lazy "a session names this backend, so bring its pool up" hook; unavailable is the runtime probe (device present, not disabled) used by the availability refusal in §5.

copy_cross / await_cross are F5's addition (§11): the split boundary's cross-backend staging copy, split into enqueue and wait so a backend with device memory can make the transfer asynchronous and name exactly where the consumer waits. They follow the same rule as the rest of the entry — the allocator never asks "is this the CPU?", it asks the entry.

4. What no longer matches on Backend

sitebeforeafter
alloc.rs pool ops (12)match backend { Backend::CPU => …, Metal => …, Cuda => … }(entry.pool_mut)(self) → trait method
alloc.rs supports_forhardcoded Metal-then-CUDA-then-CPU if let chainregistry().by_priority() + entry.caps
alloc.rs kv_load enablethree-arm match with per-cfg fallbacksentry.enable + unavailable
alloc.rs kv_element_formatmatch with per-cfg fallbacksentry.kv_format (takes the allocator: the answer is the engine's resolved format, per #99/#153 — the CPU field and the CUDA backend's kv_layout — not a process global)
alloc.rs copy_cells_in_poolmatch + per-cfg stringsentry.pool_mut + registry-aware refusal
scheduler.rs executematch split.backend { … }alloc.pool_mut(split.backend)
scheduler.rs read_host_buffermatchentry.host_read
json.rs / qwen2,3/graph.rs fusion wiringhand-built Vec + index matchalloc.fusion_backends() + Backend::index()
json.rs / dot.rs namingmatch on the enumBackend::name() / entry
kvsession.rs tagsmatch 0/1/2Backend::index() / from_index()
op_matrix.rs (test matrix)match tag { … }equality on the handle + registry caps
ensure_kv packed check, KvFormat::supportsbackend != Backend::CPU / matches!(device, Device::Cpu)entry.caps.reads_packed_kv

Everything the enum used to decide is now either a registry field or a trait call behind the entry's pool hook.

5. The name surface

Accepted names: cpu, metal, cuda. Comparison is case-insensitive and whitespace is trimmed ( Metal is metal); the canonical spelling is lower-case and is what every message uses.

Two spellings request a set of backends:

  • --backend <name> on the CLI (repeatable, and comma-separated values are accepted: --backend cpu,metal). It is extracted from argv before any subcommand dispatch, so serve, viz, run and bench all honour it.
  • MINFER_BACKENDS=<csv> in the environment, for callers that cannot pass a flag. The flag wins when both are present.

The request is a fence: it removes backends from participation. It never adds one, and it is not an offload policy (--gpu-layers / MINFER_GPU_LAYERS stay the authority for how many blocks the device holds). cpu is always allowed, whether or not it is named: it is the universal fallback and a graph must always be assignable, so --backend cuda means "the device if it can take the node, the CPU otherwise", and --backend cpu is the useful spelling — force the CPU exactly as MINFER_DISABLE_MPS=1 does, but for every device.

Unset (the default) means "every backend this build has and this machine can use" — the pre-F4 behaviour, byte for byte. The pre-existing MINFER_DISABLE_MPS and MINFER_DISABLE_CUDA fences keep working unchanged, both presence-checked: they are read by the device layer (MpsState, CudaState) and a fenced backend is simply never available, whatever the name surface says.

Where the fence is enforced:

  1. Device participation — Qwen2Graph::device / Qwen3Graph::device (the single authority CParams.gpu and the server's batching default read) return Cpu for a fenced device, so the builder never emits a device-only fused node (FusedQKV, QkvBiasNorm, FusedFFN, QkvBiasRopeStore) that the CPU cannot execute.
  2. Assignment — GraphAllocator::supports_for skips a fenced backend, so a node is never placed on a pool whose weights were never registered there.

Both read one process-wide filter installed at startup (registry::active_filter()), which is what makes them unable to disagree.

6. Refusals — three classes, three messages, always loud

A name is resolved in two stages, because the compile-time answer and the runtime answer are different questions and a backend that is compiled out must not be silently treated as absent.

Stage 1 — names, before anything else in main. Purely a function of the name and the compile-time registry (no device is touched, so it is covered by CI). Two failures:

Error: unknown backend 'gpu2'; known backends are: cpu, metal, cuda
Error: backend 'cuda' is known but not compiled into this build: the CUDA
       backend is compiled only with --features cuda

The first is "no such backend name" — always an error, on every build. The second is "the name exists, this binary does not contain it", and it names which of the two situations it is. A metal name on Linux and a cuda name on a default build both land here, with the reason spelled out (Metal is compiled only on macOS (target_os = "macos")).

Stage 2 — availability, once the device layer is up. Still startup (before the model is loaded), but now the device layer can be asked. A compiled-in backend that this machine cannot use is a third, different message:

Error: backend 'cuda' is compiled in but not available on this machine: no CUDA
       device is available, or CUDA is disabled (MINFER_DISABLE_CUDA)

A cuda name on a --features cuda build with MINFER_DISABLE_CUDA=1, or with no device, lands here — not in the "not compiled" bucket, and never in a silent fallback to the CPU.

Only a named backend is checked at stage 2. The filter therefore carries two facts per backend — allowed (may this run use it?) and requested (did the request name it?) — because the default request means "whatever this build can use". Preserving that distinction is what keeps the pre-existing MINFER_DISABLE_CUDA / MINFER_DISABLE_MPS flags meaning "run on the CPU" rather than turning them into "refuse to start": with no --backend / MINFER_BACKENDS, nothing was named, so stage 2 has nothing to check and the run proceeds on the CPU exactly as it did before F4. Naming the fenced backend (--backend cuda with MINFER_DISABLE_CUDA=1) is a refusal, because the user asked for it.

Stage 2 runs before the model path is resolved on every path that will run a model (run/serve/viz in main, and bench/specverify in their own run), so an unrelated failure — a missing file, a bad GGUF — cannot preempt it.

Both stages exit non-zero and print nothing else about backends. The reverse direction is also loud: no code path may drop a named backend and keep going.

7. Feature gates

configurationmetal entrycuda entrymetal namecuda name
default Linuxabsentabsentnot compilednot compiled
--features cudaabsentpresentnot compiledresolves
macOSpresentabsentresolvesnot compiled
macOS + --features cudapresentpresentresolvesresolves

The registry's registered set is the compile-time one: cuda_backend::register is #[cfg(feature = "cuda")], metal_backend::register is #[cfg(target_os = "macos")]. The names are not gated: Backend::name() and the known-name list are unconditional, so an unregistered name still resolves to a handle and still produces the accurate "not compiled into this build" refusal instead of "unknown backend". The ordering gate pins the registered set and the priority order per configuration (all four rows above), so a change to either is visible in CI on every configuration.

8. The seam for a later per-format capability query (#87)

#87 wants the registry to answer "can this backend read a packed q8_0 KV region?" instead of a hardcoded CPU-only test. The registry carries exactly that field — BackendCaps::reads_packed_kv — and it is used by this ticket's own code, not reserved for the future: GraphAllocator::ensure_kv's packed-region refusal and KvFormat::supports both read it, replacing backend != Backend::CPU and matches!(device, Device::Cpu) respectively. That is why it is a field and not dead abstraction: there is one authority for the answer, and #87 is the work that flips CUDA's and Metal's value to true (and adds their kernels). No other per-format query is added here.

9. What the registry is not

  • Not a plugin/dlopen system: the set of backends is fixed at compile time. The registry makes the set data instead of control flow; it does not make it extensible at runtime.
  • Not a device-selection policy: --gpu, --gpu-layers and MINFER_GPU_LAYERS keep their meanings.
  • Not a weight registry: GraphAllocator::register_weight (CPU) and CudaState::register_weight (CUDA) are unchanged; the "all weights registered" gate is per architecture and stays where it is.
  • Not a place to move Device: models::Device stays the coarse "the device participates" fact that the server's batching default reads, and gains a Device::backend() mapping so the two id spaces have one bridge.

10. Acceptance and the gates

  • Behaviour preservation. The identity and priority orders, the supports_op / supports_fused / supports_attn_span answers, the sync/copy arms and the memory accounting are unchanged; the existing suites are the measurement, and the order/name gates pin the parts a suite would not notice.
  • Gates
    1. registry::tests::names_resolve_and_unknown_names_are_refused — the known names resolve to the pinned handles, an unknown name and a compiled-out name are distinct loud errors, and diagnostics keep the pre-F4 spelling (pure; CI covers it).
    2. registry::tests::the_registered_set_and_priority_order_are_pinned — the registered set, the priority order and the priority numbers, per configuration, plus Ord and a fresh allocator's deterministic answer.
    3. registry::tests::the_name_surface_fences_devices_and_keeps_cpu — the fence surface (comma/repeat spellings, cpu always admitted, the flag winning over the environment, both refusals).
    4. registry::tests::the_packed_kv_capability_is_the_registrys_answer and registry::tests::registry_caps_match_the_backend_trait — the #87 seam is the field both C4 gates read; and the capability answer is one authority: each Backend trait method forwards to its backend module's own free function / constant, and the assignment pass reads the trait. #244 deleted the three mirrored BackendCaps fields, so the registry holds no second copy to disagree with (§3).
    5. alloc::tests::a_fresh_allocator_inherits_the_runs_backend_filter — the fence reaches the assignment pass through the same active filter.
    6. tests/backend_registry_cli.rs — the process level: an unknown name exits non-zero naming the accepted set, a compiled-out name gives the other message, a known name passes the gate (the failure moves on to the model), bench honours the flag, --help documents it, and an unnamed device disabled by the pre-existing flags is still "run on the CPU".
    7. alloc::tests::cuda_the_backend_fence_moves_assignment_off_a_usable_device — #[ignore]d (needs a device): the fence moves assignment off an enabled, usable device without tearing it down.
  • Mutation evidence. (a) making an unknown name fall back to the default backend fails gate 1; (b) swapping two priorities fails gate 2; (b2) perturbing one priority value without reordering also fails gate 2, which is why the expected numbers are literals. All three are reverted; the observed failure output is recorded in the ticket record.

11. F5 — the async staging copy and its synchronization points (#58)

Ticket: #58 ("[F5] Async cross-backend copies and events"). The implementation record (measurements, mutation evidence, honest scope) lives in docs/ARCHITECTURE-EXECUTION-PLAN.md §F5; the scheduler-side contract is also summarized in docs/COMPUTE-GRAPH-DESIGN.md §3.4. This section is the registry contract those records refer to.

11.1 What the hot path is, exactly

BackendScheduler::execute partitions a graph into contiguous same-backend splits (assign_backends → split_graph). When a value produced by one split is consumed by a split on another backend, the boundary must move it; the scheduler does that through GraphAllocator::copy_across once per entry of Split::inputs. Those copies are the hot path this ticket is about — they run on every forward, on the critical path of every decode step of a partially offloaded model.

Everything else that reads device memory back to the host is legitimately host-visible and out of scope, and is enumerated here so "zero blocking copies" is not read as "zero device→host copies anywhere":

sitewhy it blocks on purpose
GraphAllocator::copy_to_cpu on the logits / output paththe run's answer; the caller is about to read it
GraphAllocator::copy_kv_to_cpu (KV session save, --session)a file write; no overlap to exploit
Copy / debug dumps, MINFER_GRAPH_DUMP, doc dumpsdiagnostics
MINFER_TRACE / viz capture (CudaBackend::copy_to_host fallback, CaptureStaging)already batched asynchronously; the per-node fallback is for tensors above the staging ceiling
weight / tokenizer loading, fill_input, write_hosthost→device or host-only; no device leg to wait on

Before F5 a single CUDA→host boundary input cost two host stalls: a full cudaStreamSynchronize plus a blocking cudaMemcpy D2H inside CudaBackend::copy_to_host. The counters below are how that became a number (graph::copystats, per allocator, read by the gates).

11.2 The two phases

copy_across (phase A) resolves/allocates the destination staging buffer, marks the entry pending, increments copies, and calls the source backend's copy_cross. A second request for the same (graph, node, destination) while that entry is still pending is the same transfer — one staging buffer, one unchanged source node — so it is a no-op and is not counted again (#138). Under the F5 boundary that state was unreachable (the first copy was always awaited before the second was enqueued); the deferral below is what makes it reachable, and re-issuing it would duplicate the transfer and leave the first copy's pinned slab and event behind, because the allocator clears one pending key per entry.

await_cross (phase B) calls the source backend's await_cross, clears the pending flag and increments waits. One phase-B call per phase-A call is the contract, and GraphAllocator::cross_input — the checked accessor — refuses a still-pending entry with a loud Err naming the missing wait, so dropping a wait can never be bought with a silent read of in-flight data. The scheduler reads a consumer input through GraphAllocator::cross_input_ready, which issues a pending entry's wait at that first use and then calls cross_input; whatever nothing downstream reads is drained once, after the last split (GraphAllocator::drain_cross_pending). The wait is still issued exactly once per copy, at the latest safe point rather than at the boundary (#138).

Phase A returning Ok(true) means "this backend issued the transfer"; Ok(false) means "declined — use the synchronous host round trip". Declining is not a silent CPU fallback: the allocator does perform the pair, just synchronously, and the boundary counters record it as a blocking copy.

Mechanism choice on Metal (#137). Three primitives were candidates. A command-buffer completion handler can only notify the host — it is not waitable, so it cannot express phase B at all. A plain MTLEvent is device-scoped and exposes no host wait. MTLSharedEvent is the one primitive that both MTLCommandBuffer::encodeSignalEvent / encodeWaitForEvent accept and that exposes a bounded host wait (waitUntilSignaledValue:timeoutMS:); it is therefore what the port uses, and it is also the reserved device-side mechanism. The two waits are enumerated in §11.3. The blit is encoded into the producer split's own command buffer rather than a dedicated one; §11.5 records why that is a correctness requirement, not tidiness.

backendphase A (copy_cross)phase B (await_cross)
cpusynchronous host round trip — the CPU has no device memory, so there is no transfer to make asynchronous. The device leg of CPU→device is the destination pool's own stream-ordered write_host (pinned + cudaMemcpyAsync, 7e⑥), which never blocked the host either.documented no-op — nothing was enqueued that needs waiting for; the destination device orders its own fill on its stream. Still counted, so the one-wait-per-copy contract is backend-independent.
cudadevice→host: cudaMemcpyAsync D2H into a pinned slab + cudaEventRecord, both stream-ordered after the producing kernels. Any other destination declines: a device→device staging copy (unreachable — copy_across early-returns when the source and destination backends match), CUDA→Metal on a macOS+CUDA build.device→host: cudaEventSynchronize — the host block, and the only one the async path takes for that copy — then the bytes are published into the staging buffer. A device consumer uses cudaStreamWaitEvent (no host block).
metaldevice→host: a MTLBlitCommandEncoder copy of the source StorageModeShared MTLBuffer into a fresh shared staging buffer, plus MTLCommandBuffer::encodeSignalEvent on a fresh MTLSharedEvent. Both are encoded into the producer split's own command buffer (one command buffer per split — see §11.5). Any other destination declines: a Metal→Metal staging copy is unreachable (copy_across early-returns when the source and destination backends match) and CUDA does not run on Apple Silicon, so no device→device pair exists on macOS.device→host: MTLSharedEvent::waitUntilSignaledValue:timeoutMS: with a 10 s bound — the host block, and the only one the async path takes for that copy — then the bytes are published into the staging buffer. A timeout is a loud Err naming the value waited for, the observed signaledValue and the command buffer's real status() / error(); never an unbounded block (MetalBackend::cross_take). A device consumer would use MTLCommandBuffer::encodeWaitForEvent (no host block); it is reserved, because the device→device pair is unreachable on macOS.

11.3 The synchronization points, enumerated

Every execution, in order:

  1. The boundary retire (BackendScheduler::execute step 1) — GraphAllocator::retire_backend(previous). Not a copy: it retires the previous backend's kernels (and, for CUDA, closes an open graph-capture window and clears the MMQ memoization). Since #138 it calls Backend::retire, whose default body is synchronize — a backend whose boundary work is a submission (Metal) still blocks here — while CUDA overrides it: its close is already stream-ordered with the copies of step 2 (both run on the backend's own stream), so the cudaStreamSynchronize that synchronize adds orders nothing new and is not taken. sync_backend (the blocking form) remains the after-the-last-split flush.
  2. Phase A, per staged input — copy_across. A CUDA→host copy enqueues cudaMemcpyAsync + cudaEventRecord; a CUDA→device copy would use cudaMemcpyAsync D2D; the CPU does its host memcpy; Metal encodes a MTLBlitCommandEncoder copy plus encodeSignalEvent into the producer split's own command buffer (it runs before that split's retire, so the copy is ordered with the kernels that wrote its source — §11.5). No host wait here, and all of the boundary's inputs are enqueued before any of them is waited on.
  3. Phase B, per staged input — await_cross, issued at the consumer's first read of that staging buffer (or, for an entry nothing reads, by the end-of-execution drain):
    • device→host, CUDA: cudaEventSynchronize on the event recorded in step 2. Invariant that makes it necessary: the D2H destination is host memory, and the host is about to read it; without the wait the consumer reads bytes the DMA may not have written yet. Deferring it to the read is what lets the other copies of the same boundary stay in flight.
    • device→host, Metal: MTLSharedEvent::waitUntilSignaledValue:timeoutMS: with a 10 s bound, then a host read of the shared staging buffer. Invariant: the event is signaled at the end of the producer split's command buffer, and the host's StorageModeShared read is only ordered after that signal; the wait is deferred to the consumer's read exactly as CUDA's is. A timeout is a loud Err naming the real status. The bound and the loudness are the GPU-safety contract (docs/GPU_SAFETY.md §2.4).
    • host→device: no host wait; the fill is ordered on the consuming backend's own stream ahead of the kernels that read it. Invariant: stream order — the copy and the first consumer share one stream, so no host synchronization is needed and none is taken. This is the device-consumer direction the ticket's first bullet names, and it is already host-free by construction.
    • device→device (no backend implements it today): cudaStreamWaitEvent on the consuming stream, or MTLCommandBuffer::encodeWaitForEvent with the signalling MTLSharedEvent, would be the mechanism; the host never blocks. It stays reserved: copy_across early-returns when source and destination match, and no other device destination is reachable on a macOS build (CUDA does not run on Apple Silicon), so there is no pair to exercise. See the census note in the execution plan.
    • The contract invariant, independent of backend: one phase-B wait per phase-A copy, enforced by cross_input's refusal.
  4. The next boundary's retire, or the final flush after the last split — sync_backend, which for CUDA is also where an open capture window closes.

The gated evidence for 2–3 is copystats::CrossCopyStats (per allocator: copies, waits, deferred_waits, blocking_host_copies, async_host_copies, event_syncs, stream_waits), CudaBackend::blocking_readback_count() (a device-level count of blocking cudaMemcpy D2H calls), CudaBackend::stream_sync_count() (host stalls; per backend since #185, the process-wide counter deleted in #242), CudaBackend::cross_inflight_peak() (how many staging copies were enqueued but not yet waited on at once — the device-side overlap metric of #138) and, on Metal, MetalBackend::sync_readback_count() (host reads of a StorageModeShared pool buffer — there is no blocking-copy API to count, so this is the device-level analogue; the async path never moves it). MINFER_SYNC_COPIES=1 restores the pre-F5 synchronous path as the bitwise reference; copystats::set_sync_for_test is its programmatic form.

11.4 Overlap: what F5 delivers, and what the deferred wait adds

F5 delivered the async substrate + documented waits: no blocking copy on the boundary path, one explicit wait per staged input, and one measured reduction in host stalls (the per-copy stream synchronizations inside copy_to_host are gone).

#138 makes the wait late. The boundary only enqueues; each staged entry's single wait is issued at the consumer's first read of it, and the end-of-execution drain covers an entry nothing reads. Two things follow, both measured on the 0.5B mixed-offload gate:

  • the boundary's own cudaStreamSynchronize is gone (it retired the producer before copies that are already stream-ordered behind it — one full stream sync per CUDA→CPU boundary, 21 → 0 over the gate's 7 forwards), and
  • the boundary's copies stay in flight while the consumer works, so a boundary with several staged inputs holds several transfers at once (2 where the F5 enqueue-then-wait order cannot exceed 1).

Resolved by #300: there is no cross-split overlap to gain on the reachable macOS topology, and §11.6 records the measured negative result — the boundary blit stays on the producer's command buffer, which is both correct (it reads the source in the submission that wrote it) and not slower than the explicit-dependency alternative. The split loop remains strictly sequential, and a host-side consumer must wait by definition. What the deferral buys is that the wait happens where the data is needed, not where it was produced.

11.5 The Metal port (#137): why the blit shares the split's command buffer

MetalBackend keeps one MpsCommandBuffer per split (compute-graph rule 8); the producer submits and bounded-waits it in Backend::retire. The first cut of this port gave each staging copy its own command buffer, committed from copy_cross after the retire. That is a second submission overlapping the next split's, and on the measurement box it changed the kernels' results: the same mixed 0.5B offload graph run twice with the copy mode toggled differed by max |Δlogit| ≈ 1.4, while each mode compared with itself was bitwise stable. The staged bytes were identical (the source was read at enqueue and compared with the staging buffer at the wait), so the divergence was the extra in-flight command buffer perturbing Metal kernel execution — a pre-existing fragility of the backend, not the copy.

Encoding the blit into the producer split's own command buffer fixes it and restores the one-command-buffer invariant: copy_across for a Metal→CPU boundary now runs before BackendScheduler::execute calls retire_backend, so MetalBackend::cross_enqueue appends an MTLBlitCommandEncoder pass (and encodeSignalEvent) to the still-open split buffer, and the split's own submission carries the copy behind the kernels that wrote its source. A source with no open split buffer gets a standalone buffer that cross_enqueue submits itself; the synchronous reference (MINFER_SYNC_COPIES=1) keeps the post-retire order, because its host read must see the producer's finished output.

Measured on macbook (macOS 27.0.1, Apple M4 Pro) — the F5 S3 record in docs/ARCHITECTURE-EXECUTION-PLAN.md.

11.6 The #300 result: the serialization is a missing dependency, and there is no overlap to gain

Ticket: #300 ("true cross-split overlap after #137"). It asked for either a measured overlap or a recorded negative result beside the §11.5 note. The result is a negative one, with the mechanism measured on macbook (macOS 27.0.1, Apple M4 Pro) at ad707c7 (2026-10-09), recorded in 9df405d.

The #137 divergence is a missing-dependency race, not a kernel perturbation. §11.5 inferred from the then-red baseline that "the extra in-flight command buffer perturbed Metal kernel execution". Isolated, the mechanism is simpler and deterministic. The pathological first cut is a standalone boundary command buffer submitted from copy_cross before retire submits the producer's split buffer, with no dependency on that buffer. On the shared MpsState command queue the standalone buffer is committed first, so the blit reads the producer's StorageModeShared source window before the producer's kernels wrote it, and the consumer's wait on an already-signaled event returns the previous forward's bytes. On models::qwen2::graph::tests::offload_copy::async_cross_copies_never_block_and_stay_bitwise_identical_on_metal that cut reads max |Δlogit| = 26.718678 at step 0, the same value on two runs (the §11.5 "staged bytes identical" check was stale-to-stale, which is why it looked consistent).

Option A — a separate boundary buffer with an explicit dependency — is correct. Committing the boundary blit on its own MTLCommandQueue, with encodeWaitForEvent on an event the producer signals at its split end before the blit and its own event for the consumer, restores bitwise identity: the real-model gate reads max |Δlogit| = 0 over the prefill + 6-decode loop (twice) with the counters unchanged — copies=35 waits=35 deferred_waits=35 blocking_host_copies=0 async_host_copies=21 event_syncs=21 sync_readbacks=0 async against copies=35 waits=35 blocking_host_copies=21 sync_readbacks=21 sync — and the cheap 3-node gate passes (copies=2 waits=2 deferred=2 blocking=0 async_host=1 event_syncs=1 sync_readbacks=0). So the perturbation is avoidable: it is a dependency the shared queue did not supply, not a fragility of Metal kernel execution.

But Option A buys no measurable cross-split overlap, so the production design stays. Cross-split overlap needs a second split whose work can run beside the copy. On the reachable macOS topology there is none: the E5 mixed graph (4 of 24 blocks on the device) is a single CPU → Metal → CPU sequence, so there is exactly one device→host boundary per forward; that boundary's consumer is the host CPU split, whose first node directly reads the dominant staged tensor (the hidden state), so the copy is on the consumer's critical path by definition; and the boundary's other two staged tensors are read within the CPU split's first eight nodes (add at node 46, cells at 52, attn_span at 54 in the decode graph), so a separate buffer could only overlap a blit of two tiny index/window tensors with a handful of host-side node setups. A general non-blocking producer retire cannot be made safe without reordering the next device submission after the boundary blit, which reintroduces exactly the serialization it would remove. A device→device pair (Metal→Metal) would be the one topology with real overlap, and copy_across early-returns on it while no CUDA device exists on macOS — it stays reserved (§11.3). The boundary blit therefore keeps riding the producer's command buffer; it is bitwise, and on this topology it is not slower than Option A.

Reproduction. Both cuts are MINFER_300_*-gated temporary patches to MetalBackend::cross_enqueue (a standalone cmd_buffer() + submit() before retire for the pathological cut; a second queue + encodeWaitForEvent / encodeSignalEvent for Option A) — neither is in the tree. The acceptance command is the same in every arm:

cargo test --release --bin minfer async_cross_copies_never_block_and_stay_bitwise_identical_on_metal -- --ignored --test-threads=1 --nocapture

Decisions governing this document

This page is the current contract; the decisions behind it are frozen in the ADR corpus:

  • ADR-0001 — Inference runs through one declarative compute graph
  • ADR-0011 — Backend ids are a file-format contract: appended, never renumbered
  • ADR-0006 — The KV storage format is a per-engine gate, not a process-wide global
  • ADR-0008 — GPU safety: bounded waits, no early return past a barrier, runtime device limits
  • ADR-0009 — A failure is an error, never a silent fallback