F4 — Backend registry (design record)
Ticket: #57 ("[F4] Backend
registry instead of the compile-time enum"). This document is the
design-first artifact: it is committed before the implementation and states
the registry's shape, the name surface, the ordering rule and the exact
refusals. Gaps and the backlog item live in
docs/ARCHITECTURE-ROADMAP.md; the ticket record lives in
docs/ARCHITECTURE-EXECUTION-PLAN.md.
1. Why the enum was the problem
Backend was enum Backend { CPU, Metal, Cuda } in src/graph/mod.rs, and every consumer
matched on it. GraphAllocator alone dispatched twelve operations through
match backend { … }; BackendScheduler::execute matched to pick the executing pool and again
for the MINFER_TRACE capture path; and the fusion wiring built a Vec<&dyn Backend> by hand and
found a node's backend with a position(|b| b.name() == "cuda") lookup that existed only because
the match could not express "whichever device is present". Adding a backend therefore meant
editing the allocator, the scheduler, the fusion wiring, the exporters, the KV-session tag table
and the op matrix; and none of it told a user anything — there was no way to ask for a backend by
name or to learn that a requested one did not exist.
That argument is a decision, and it is frozen in ADR-0001 (one device seam, build-time assignment) and ADR-0011 (the id space and the name surface). This page keeps the shape that replaced it — the handle, the registered set, the ordering rule and the refusals.
2. The handle: Backend(u16)
Backend stays a cheap, Copy, Hash, Eq, Ord value; it stops being an
enum. It is an opaque index into a fixed id space:
#![allow(unused)] fn main() { pub struct Backend(u16); impl Backend { pub const CPU: Backend = Backend(0); pub const METAL: Backend = Backend(1); pub const CUDA: Backend = Backend(2); } }
The id space is compile-time and configuration-independent: cpu = 0,
metal = 1, cuda = 2 on every build, whether or not the backend is compiled
in. That is deliberate — the id is already on disk and in exports:
- the KV-session backend tag (
graph/kvsession.rs,tag_of/backend_of_tag) is au32in a versioned file, so the numbering is a file-format contract; graph/json.rs/graph/dot.rsname and index the backend in exported graph documents;Ordis derived from the id, and the pre-F4 derivedOrdon the enum was declaration order —CPU < Metal < Cuda.
Backend::index(), Backend::from_index() and Backend::name() are the only
ways to leave and re-enter the id space; kvsession's tag functions and the
fusion index are now one-line calls to them instead of three-arm matches.
Debug is implemented by hand so diagnostics print exactly what the enum
printed (CPU, Metal, Cuda) — no log line or error message changes shape.
2.1 Two orders, stated once
The registry has two orders and they are not the same; conflating them is the failure this document exists to prevent.
| order | what it fixes | authority |
|---|---|---|
identity (Backend::index()) | the on-disk KV-session tag, the exported graph's backend index, Ord, the fusion pass's backend list | Backend::CPU/METAL/CUDA |
| priority (assignment preference) | which backend supports_for offers first | BackendEntry::priority, descending |
Before F4 the priority order was implicit in the statement order of
GraphAllocator::supports_for: Metal, then CUDA, then CPU. It is preserved
exactly, and it is now a number on each entry — Metal 300, CUDA 200, CPU 100 —
so the ordering gate can pin it and a future accidental reordering fails a test
instead of silently changing where a graph's nodes land.
Determinism rule. The assignment pass fixes the graph's topology, which is
part of GraphParams' reuse identity (standing rule 3). Nothing in the registry
may depend on HashMap iteration order:
- the registry table is a fixed-size array indexed by the handle id — not a
HashMap; Registry::by_priority()sorts by(Reverse(priority), index), so two entries can never tie and the result is a total order derived from the pinned numbers;BackendFilteris a[bool; N_BACKENDS]indexed by id, never a set that is iterated;- the fusion-pass backend list is built by walking the identity order, so the
index it hands the pass is
Backend::index()on every run.
3. The registry table
One BackendEntry per backend, built once at startup (registry(), a
OnceLock, is forced by main before anything else runs):
#![allow(unused)] fn main() { pub struct BackendCaps { pub reads_packed_kv: bool, // the #87 seam, §8 — the only field } pub struct BackendEntry { pub handle: Backend, pub name: &'static str, // "cpu" | "metal" | "cuda" pub priority: u16, pub caps: BackendCaps, pub pool: fn(&GraphAllocator) -> Option<&dyn BackendTrait>, pub pool_mut: fn(&mut GraphAllocator) -> Option<&mut dyn BackendTrait>, pub host_read: fn(&GraphAllocator, usize) -> Option<Vec<f32>>, // F5 (#58): the split boundary's two phases (§11). pub copy_cross: fn(&mut GraphAllocator, u64, NodeId, Backend) -> Result<bool, String>, pub await_cross: fn(&mut GraphAllocator, u64, NodeId, Backend) -> Result<(), String>, pub kv_format: fn(&GraphAllocator) -> KvFormat, pub enable: fn(&mut GraphAllocator) -> bool, pub unavailable: fn() -> Option<&'static str>, } }
Each backend module owns its entry and its hooks (cpu_backend::entry(),
metal_backend::entry(), cuda_backend::entry()), and Registry::build()
calls their register() — that call site is the only place a backend is
introduced.
The capability authority is the module-level item, stated once (#244,
2026-10-01). Each backend defines its op×dtype matrix and its fusion matrix as
module-level free functions (cpu_backend::supports_op, cuda_backend::supports_fused,
…) and its attn_span answer as a module-level constant
(cpu_backend::SUPPORTS_ATTN_SPAN, …). The impl Backend for X methods are
one-line forwards to them, and the assignment pass reads the trait method
(graph::backend_takes → dyn Backend::supports_op). The registry carries no
copy of those three answers. The supports_op / supports_fused /
supports_attn_span fields that used to sit here were written by every
entry() and read only by registry::tests — a mirror with no production
reader, which is why #244 deleted them (option (b) of that ticket's
escalation). The gate that used to compare the field against the trait now
compares the module-level function against the trait
(registry::tests::registry_caps_match_the_backend_trait), i.e. the authority
the field merely mirrored; putting the two on one line is no longer possible
even in principle.
BackendCaps therefore carries exactly one field, reads_packed_kv, and that
one is read in production because the question is asked about a format,
not an engine: GraphAllocator::ensure_kv's packed-region refusal and
KvFormat::supports call registry::reads_packed_kv(backend), where no &self
is available (§8). It is the one capability the registry itself carries.
pool / pool_mut are why "sync" and "copy" are not a match any more: the
allocator's dispatch helpers ask the entry for the pool and then call the
Backend trait method (synchronize, copy_cells, write_host, …) on it.
The two backends that have a native host-read path different from the trait's
borrowed read_host (CUDA's stream-ordered copy_to_host) say so in
host_read, which is why copy_to_cpu and the CUDA KV debug read all
collapse to one call without a match.
enable is the lazy "a session names this backend, so bring its pool up"
hook; unavailable is the runtime probe (device present, not disabled) used by
the availability refusal in §5.
copy_cross / await_cross are F5's addition (§11): the split boundary's
cross-backend staging copy, split into enqueue and wait so a backend with
device memory can make the transfer asynchronous and name exactly where the
consumer waits. They follow the same rule as the rest of the entry — the
allocator never asks "is this the CPU?", it asks the entry.
4. What no longer matches on Backend
| site | before | after |
|---|---|---|
alloc.rs pool ops (12) | match backend { Backend::CPU => …, Metal => …, Cuda => … } | (entry.pool_mut)(self) → trait method |
alloc.rs supports_for | hardcoded Metal-then-CUDA-then-CPU if let chain | registry().by_priority() + entry.caps |
alloc.rs kv_load enable | three-arm match with per-cfg fallbacks | entry.enable + unavailable |
alloc.rs kv_element_format | match with per-cfg fallbacks | entry.kv_format (takes the allocator: the answer is the engine's resolved format, per #99/#153 — the CPU field and the CUDA backend's kv_layout — not a process global) |
alloc.rs copy_cells_in_pool | match + per-cfg strings | entry.pool_mut + registry-aware refusal |
scheduler.rs execute | match split.backend { … } | alloc.pool_mut(split.backend) |
scheduler.rs read_host_buffer | match | entry.host_read |
json.rs / qwen2,3/graph.rs fusion wiring | hand-built Vec + index match | alloc.fusion_backends() + Backend::index() |
json.rs / dot.rs naming | match on the enum | Backend::name() / entry |
kvsession.rs tags | match 0/1/2 | Backend::index() / from_index() |
op_matrix.rs (test matrix) | match tag { … } | equality on the handle + registry caps |
ensure_kv packed check, KvFormat::supports | backend != Backend::CPU / matches!(device, Device::Cpu) | entry.caps.reads_packed_kv |
Everything the enum used to decide is now either a registry field or a trait call behind the entry's pool hook.
5. The name surface
Accepted names: cpu, metal, cuda. Comparison is case-insensitive and
whitespace is trimmed ( Metal is metal); the canonical spelling is
lower-case and is what every message uses.
Two spellings request a set of backends:
--backend <name>on the CLI (repeatable, and comma-separated values are accepted:--backend cpu,metal). It is extracted fromargvbefore any subcommand dispatch, soserve,viz,runandbenchall honour it.MINFER_BACKENDS=<csv>in the environment, for callers that cannot pass a flag. The flag wins when both are present.
The request is a fence: it removes backends from participation. It never
adds one, and it is not an offload policy (--gpu-layers / MINFER_GPU_LAYERS
stay the authority for how many blocks the device holds). cpu is always
allowed, whether or not it is named: it is the universal fallback and a graph
must always be assignable, so --backend cuda means "the device if it can take
the node, the CPU otherwise", and --backend cpu is the useful spelling —
force the CPU exactly as MINFER_DISABLE_MPS=1 does, but for every device.
Unset (the default) means "every backend this build has and this machine can
use" — the pre-F4 behaviour, byte for byte. The pre-existing
MINFER_DISABLE_MPS and MINFER_DISABLE_CUDA fences keep working unchanged,
both presence-checked: they are read by the device layer (MpsState,
CudaState) and a fenced backend is simply never available, whatever the name
surface says.
Where the fence is enforced:
- Device participation —
Qwen2Graph::device/Qwen3Graph::device(the single authorityCParams.gpuand the server's batching default read) returnCpufor a fenced device, so the builder never emits a device-only fused node (FusedQKV,QkvBiasNorm,FusedFFN,QkvBiasRopeStore) that the CPU cannot execute. - Assignment —
GraphAllocator::supports_forskips a fenced backend, so a node is never placed on a pool whose weights were never registered there.
Both read one process-wide filter installed at startup
(registry::active_filter()), which is what makes them unable to disagree.
6. Refusals — three classes, three messages, always loud
A name is resolved in two stages, because the compile-time answer and the runtime answer are different questions and a backend that is compiled out must not be silently treated as absent.
Stage 1 — names, before anything else in main. Purely a function of the
name and the compile-time registry (no device is touched, so it is covered by
CI). Two failures:
Error: unknown backend 'gpu2'; known backends are: cpu, metal, cuda
Error: backend 'cuda' is known but not compiled into this build: the CUDA
backend is compiled only with --features cuda
The first is "no such backend name" — always an error, on every build. The
second is "the name exists, this binary does not contain it", and it names which
of the two situations it is. A metal name on Linux and a cuda name on a
default build both land here, with the reason spelled out (Metal is compiled only on macOS (target_os = "macos")).
Stage 2 — availability, once the device layer is up. Still startup (before the model is loaded), but now the device layer can be asked. A compiled-in backend that this machine cannot use is a third, different message:
Error: backend 'cuda' is compiled in but not available on this machine: no CUDA
device is available, or CUDA is disabled (MINFER_DISABLE_CUDA)
A cuda name on a --features cuda build with MINFER_DISABLE_CUDA=1, or with
no device, lands here — not in the "not compiled" bucket, and never in a silent
fallback to the CPU.
Only a named backend is checked at stage 2. The filter therefore carries
two facts per backend — allowed (may this run use it?) and requested (did the
request name it?) — because the default request means "whatever this build can
use". Preserving that distinction is what keeps the pre-existing
MINFER_DISABLE_CUDA / MINFER_DISABLE_MPS flags meaning "run on the CPU"
rather than turning them into "refuse to start": with no --backend /
MINFER_BACKENDS, nothing was named, so stage 2 has nothing to check and the
run proceeds on the CPU exactly as it did before F4. Naming the fenced backend
(--backend cuda with MINFER_DISABLE_CUDA=1) is a refusal, because the user
asked for it.
Stage 2 runs before the model path is resolved on every path that will run a
model (run/serve/viz in main, and bench/specverify in their own
run), so an unrelated failure — a missing file, a bad GGUF — cannot preempt it.
Both stages exit non-zero and print nothing else about backends. The reverse direction is also loud: no code path may drop a named backend and keep going.
7. Feature gates
| configuration | metal entry | cuda entry | metal name | cuda name |
|---|---|---|---|---|
| default Linux | absent | absent | not compiled | not compiled |
--features cuda | absent | present | not compiled | resolves |
| macOS | present | absent | resolves | not compiled |
macOS + --features cuda | present | present | resolves | resolves |
The registry's registered set is the compile-time one: cuda_backend::register
is #[cfg(feature = "cuda")], metal_backend::register is
#[cfg(target_os = "macos")]. The names are not gated: Backend::name() and
the known-name list are unconditional, so an unregistered name still resolves to
a handle and still produces the accurate "not compiled into this build" refusal
instead of "unknown backend". The ordering gate pins the registered set and the
priority order per configuration (all four rows above), so a change to either is
visible in CI on every configuration.
8. The seam for a later per-format capability query (#87)
#87 wants the registry to answer
"can this backend read a packed q8_0 KV region?" instead of a hardcoded
CPU-only test. The registry carries exactly that field —
BackendCaps::reads_packed_kv — and it is used by this ticket's own code,
not reserved for the future: GraphAllocator::ensure_kv's packed-region
refusal and KvFormat::supports both read it, replacing
backend != Backend::CPU and matches!(device, Device::Cpu) respectively. That
is why it is a field and not dead abstraction: there is one authority for the
answer, and #87 is the work that flips CUDA's and Metal's value to true (and
adds their kernels). No other per-format query is added here.
9. What the registry is not
- Not a plugin/
dlopensystem: the set of backends is fixed at compile time. The registry makes the set data instead of control flow; it does not make it extensible at runtime. - Not a device-selection policy:
--gpu,--gpu-layersandMINFER_GPU_LAYERSkeep their meanings. - Not a weight registry:
GraphAllocator::register_weight(CPU) andCudaState::register_weight(CUDA) are unchanged; the "all weights registered" gate is per architecture and stays where it is. - Not a place to move
Device:models::Devicestays the coarse "the device participates" fact that the server's batching default reads, and gains aDevice::backend()mapping so the two id spaces have one bridge.
10. Acceptance and the gates
- Behaviour preservation. The identity and priority orders, the
supports_op/supports_fused/supports_attn_spananswers, the sync/copy arms and the memory accounting are unchanged; the existing suites are the measurement, and the order/name gates pin the parts a suite would not notice. - Gates
registry::tests::names_resolve_and_unknown_names_are_refused— the known names resolve to the pinned handles, an unknown name and a compiled-out name are distinct loud errors, and diagnostics keep the pre-F4 spelling (pure; CI covers it).registry::tests::the_registered_set_and_priority_order_are_pinned— the registered set, the priority order and the priority numbers, per configuration, plusOrdand a fresh allocator's deterministic answer.registry::tests::the_name_surface_fences_devices_and_keeps_cpu— the fence surface (comma/repeat spellings,cpualways admitted, the flag winning over the environment, both refusals).registry::tests::the_packed_kv_capability_is_the_registrys_answerandregistry::tests::registry_caps_match_the_backend_trait— the #87 seam is the field both C4 gates read; and the capability answer is one authority: eachBackendtrait method forwards to its backend module's own free function / constant, and the assignment pass reads the trait. #244 deleted the three mirroredBackendCapsfields, so the registry holds no second copy to disagree with (§3).alloc::tests::a_fresh_allocator_inherits_the_runs_backend_filter— the fence reaches the assignment pass through the same active filter.tests/backend_registry_cli.rs— the process level: an unknown name exits non-zero naming the accepted set, a compiled-out name gives the other message, a known name passes the gate (the failure moves on to the model),benchhonours the flag,--helpdocuments it, and an unnamed device disabled by the pre-existing flags is still "run on the CPU".alloc::tests::cuda_the_backend_fence_moves_assignment_off_a_usable_device—#[ignore]d (needs a device): the fence moves assignment off an enabled, usable device without tearing it down.
- Mutation evidence. (a) making an unknown name fall back to the default backend fails gate 1; (b) swapping two priorities fails gate 2; (b2) perturbing one priority value without reordering also fails gate 2, which is why the expected numbers are literals. All three are reverted; the observed failure output is recorded in the ticket record.
11. F5 — the async staging copy and its synchronization points (#58)
Ticket: #58 ("[F5] Async
cross-backend copies and events"). The implementation record (measurements,
mutation evidence, honest scope) lives in
docs/ARCHITECTURE-EXECUTION-PLAN.md §F5; the scheduler-side contract is also
summarized in docs/COMPUTE-GRAPH-DESIGN.md §3.4. This section is the registry
contract those records refer to.
11.1 What the hot path is, exactly
BackendScheduler::execute partitions a graph into contiguous same-backend
splits (assign_backends → split_graph). When a value produced by one split is
consumed by a split on another backend, the boundary must move it; the
scheduler does that through GraphAllocator::copy_across once per entry of
Split::inputs. Those copies are the hot path this ticket is about — they
run on every forward, on the critical path of every decode step of a partially
offloaded model.
Everything else that reads device memory back to the host is legitimately host-visible and out of scope, and is enumerated here so "zero blocking copies" is not read as "zero device→host copies anywhere":
| site | why it blocks on purpose |
|---|---|
GraphAllocator::copy_to_cpu on the logits / output path | the run's answer; the caller is about to read it |
GraphAllocator::copy_kv_to_cpu (KV session save, --session) | a file write; no overlap to exploit |
Copy / debug dumps, MINFER_GRAPH_DUMP, doc dumps | diagnostics |
MINFER_TRACE / viz capture (CudaBackend::copy_to_host fallback, CaptureStaging) | already batched asynchronously; the per-node fallback is for tensors above the staging ceiling |
weight / tokenizer loading, fill_input, write_host | host→device or host-only; no device leg to wait on |
Before F5 a single CUDA→host boundary input cost two host stalls: a full
cudaStreamSynchronize plus a blocking cudaMemcpy D2H inside
CudaBackend::copy_to_host. The counters below are how that became a number
(graph::copystats, per allocator, read by the gates).
11.2 The two phases
copy_across (phase A) resolves/allocates the destination staging buffer, marks
the entry pending, increments copies, and calls the source backend's
copy_cross. A second request for the same (graph, node, destination) while
that entry is still pending is the same transfer — one staging buffer, one
unchanged source node — so it is a no-op and is not counted again (#138). Under
the F5 boundary that state was unreachable (the first copy was always awaited
before the second was enqueued); the deferral below is what makes it reachable,
and re-issuing it would duplicate the transfer and leave the first copy's pinned
slab and event behind, because the allocator clears one pending key per entry.
await_cross (phase B) calls the source backend's await_cross, clears the
pending flag and increments waits. One phase-B call per phase-A call is the
contract, and GraphAllocator::cross_input — the checked accessor — refuses a
still-pending entry with a loud Err naming the missing wait, so dropping a wait
can never be bought with a silent read of in-flight data. The scheduler reads a
consumer input through GraphAllocator::cross_input_ready, which issues a
pending entry's wait at that first use and then calls cross_input; whatever
nothing downstream reads is drained once, after the last split
(GraphAllocator::drain_cross_pending). The wait is still issued exactly once
per copy, at the latest safe point rather than at the boundary (#138).
Phase A returning Ok(true) means "this backend issued the transfer";
Ok(false) means "declined — use the synchronous host round trip". Declining
is not a silent CPU fallback: the allocator does perform the pair, just
synchronously, and the boundary counters record it as a blocking copy.
Mechanism choice on Metal (#137). Three primitives were candidates. A
command-buffer completion handler can only notify the host — it is not
waitable, so it cannot express phase B at all. A plain MTLEvent is
device-scoped and exposes no host wait. MTLSharedEvent is the one primitive
that both MTLCommandBuffer::encodeSignalEvent / encodeWaitForEvent accept
and that exposes a bounded host wait
(waitUntilSignaledValue:timeoutMS:); it is therefore what the port uses, and it
is also the reserved device-side mechanism. The two waits are enumerated in
§11.3. The blit is encoded into the producer split's own command buffer rather
than a dedicated one; §11.5 records why that is a correctness requirement, not
tidiness.
| backend | phase A (copy_cross) | phase B (await_cross) |
|---|---|---|
| cpu | synchronous host round trip — the CPU has no device memory, so there is no transfer to make asynchronous. The device leg of CPU→device is the destination pool's own stream-ordered write_host (pinned + cudaMemcpyAsync, 7e⑥), which never blocked the host either. | documented no-op — nothing was enqueued that needs waiting for; the destination device orders its own fill on its stream. Still counted, so the one-wait-per-copy contract is backend-independent. |
| cuda | device→host: cudaMemcpyAsync D2H into a pinned slab + cudaEventRecord, both stream-ordered after the producing kernels. Any other destination declines: a device→device staging copy (unreachable — copy_across early-returns when the source and destination backends match), CUDA→Metal on a macOS+CUDA build. | device→host: cudaEventSynchronize — the host block, and the only one the async path takes for that copy — then the bytes are published into the staging buffer. A device consumer uses cudaStreamWaitEvent (no host block). |
| metal | device→host: a MTLBlitCommandEncoder copy of the source StorageModeShared MTLBuffer into a fresh shared staging buffer, plus MTLCommandBuffer::encodeSignalEvent on a fresh MTLSharedEvent. Both are encoded into the producer split's own command buffer (one command buffer per split — see §11.5). Any other destination declines: a Metal→Metal staging copy is unreachable (copy_across early-returns when the source and destination backends match) and CUDA does not run on Apple Silicon, so no device→device pair exists on macOS. | device→host: MTLSharedEvent::waitUntilSignaledValue:timeoutMS: with a 10 s bound — the host block, and the only one the async path takes for that copy — then the bytes are published into the staging buffer. A timeout is a loud Err naming the value waited for, the observed signaledValue and the command buffer's real status() / error(); never an unbounded block (MetalBackend::cross_take). A device consumer would use MTLCommandBuffer::encodeWaitForEvent (no host block); it is reserved, because the device→device pair is unreachable on macOS. |
11.3 The synchronization points, enumerated
Every execution, in order:
- The boundary retire (
BackendScheduler::executestep 1) —GraphAllocator::retire_backend(previous). Not a copy: it retires the previous backend's kernels (and, for CUDA, closes an open graph-capture window and clears the MMQ memoization). Since #138 it callsBackend::retire, whose default body issynchronize— a backend whose boundary work is a submission (Metal) still blocks here — while CUDA overrides it: its close is already stream-ordered with the copies of step 2 (both run on the backend's own stream), so thecudaStreamSynchronizethatsynchronizeadds orders nothing new and is not taken.sync_backend(the blocking form) remains the after-the-last-split flush. - Phase A, per staged input —
copy_across. A CUDA→host copy enqueuescudaMemcpyAsync+cudaEventRecord; a CUDA→device copy would usecudaMemcpyAsyncD2D; the CPU does its host memcpy; Metal encodes aMTLBlitCommandEncodercopy plusencodeSignalEventinto the producer split's own command buffer (it runs before that split'sretire, so the copy is ordered with the kernels that wrote its source — §11.5). No host wait here, and all of the boundary's inputs are enqueued before any of them is waited on. - Phase B, per staged input —
await_cross, issued at the consumer's first read of that staging buffer (or, for an entry nothing reads, by the end-of-execution drain):- device→host, CUDA:
cudaEventSynchronizeon the event recorded in step 2. Invariant that makes it necessary: the D2H destination is host memory, and the host is about to read it; without the wait the consumer reads bytes the DMA may not have written yet. Deferring it to the read is what lets the other copies of the same boundary stay in flight. - device→host, Metal:
MTLSharedEvent::waitUntilSignaledValue:timeoutMS:with a 10 s bound, then a host read of the shared staging buffer. Invariant: the event is signaled at the end of the producer split's command buffer, and the host'sStorageModeSharedread is only ordered after that signal; the wait is deferred to the consumer's read exactly as CUDA's is. A timeout is a loudErrnaming the real status. The bound and the loudness are the GPU-safety contract (docs/GPU_SAFETY.md§2.4). - host→device: no host wait; the fill is ordered on the consuming backend's own stream ahead of the kernels that read it. Invariant: stream order — the copy and the first consumer share one stream, so no host synchronization is needed and none is taken. This is the device-consumer direction the ticket's first bullet names, and it is already host-free by construction.
- device→device (no backend implements it today):
cudaStreamWaitEventon the consuming stream, orMTLCommandBuffer::encodeWaitForEventwith the signallingMTLSharedEvent, would be the mechanism; the host never blocks. It stays reserved:copy_acrossearly-returns when source and destination match, and no other device destination is reachable on a macOS build (CUDA does not run on Apple Silicon), so there is no pair to exercise. See the census note in the execution plan. - The contract invariant, independent of backend: one phase-B wait per
phase-A copy, enforced by
cross_input's refusal.
- device→host, CUDA:
- The next boundary's retire, or the final flush after the last split —
sync_backend, which for CUDA is also where an open capture window closes.
The gated evidence for 2–3 is copystats::CrossCopyStats (per allocator:
copies, waits, deferred_waits, blocking_host_copies, async_host_copies,
event_syncs, stream_waits), CudaBackend::blocking_readback_count() (a
device-level count of blocking cudaMemcpy D2H calls),
CudaBackend::stream_sync_count() (host stalls; per backend since #185, the
process-wide counter deleted in #242), CudaBackend::cross_inflight_peak()
(how many staging copies were enqueued but not yet waited on at once — the
device-side overlap metric of #138) and, on Metal,
MetalBackend::sync_readback_count() (host reads of a StorageModeShared pool
buffer — there is no blocking-copy API to count, so this is the device-level
analogue; the async path never moves it).
MINFER_SYNC_COPIES=1 restores the pre-F5 synchronous path as the bitwise
reference; copystats::set_sync_for_test is its programmatic form.
11.4 Overlap: what F5 delivers, and what the deferred wait adds
F5 delivered the async substrate + documented waits: no blocking copy on the
boundary path, one explicit wait per staged input, and one measured reduction in
host stalls (the per-copy stream synchronizations inside copy_to_host are
gone).
#138 makes the wait late. The boundary only enqueues; each staged entry's single wait is issued at the consumer's first read of it, and the end-of-execution drain covers an entry nothing reads. Two things follow, both measured on the 0.5B mixed-offload gate:
- the boundary's own
cudaStreamSynchronizeis gone (it retired the producer before copies that are already stream-ordered behind it — one full stream sync per CUDA→CPU boundary, 21 → 0 over the gate's 7 forwards), and - the boundary's copies stay in flight while the consumer works, so a boundary with several staged inputs holds several transfers at once (2 where the F5 enqueue-then-wait order cannot exceed 1).
Resolved by #300: there is no cross-split overlap to gain on the reachable macOS topology, and §11.6 records the measured negative result — the boundary blit stays on the producer's command buffer, which is both correct (it reads the source in the submission that wrote it) and not slower than the explicit-dependency alternative. The split loop remains strictly sequential, and a host-side consumer must wait by definition. What the deferral buys is that the wait happens where the data is needed, not where it was produced.
11.5 The Metal port (#137): why the blit shares the split's command buffer
MetalBackend keeps one MpsCommandBuffer per split (compute-graph rule 8);
the producer submits and bounded-waits it in Backend::retire. The first cut of
this port gave each staging copy its own command buffer, committed from
copy_cross after the retire. That is a second submission overlapping the next
split's, and on the measurement box it changed the kernels' results: the same
mixed 0.5B offload graph run twice with the copy mode toggled differed by
max |Δlogit| ≈ 1.4, while each mode compared with itself was bitwise stable. The
staged bytes were identical (the source was read at enqueue and compared with
the staging buffer at the wait), so the divergence was the extra in-flight
command buffer perturbing Metal kernel execution — a pre-existing fragility of the
backend, not the copy.
Encoding the blit into the producer split's own command buffer fixes it and
restores the one-command-buffer invariant: copy_across for a Metal→CPU boundary
now runs before BackendScheduler::execute calls retire_backend, so
MetalBackend::cross_enqueue appends an MTLBlitCommandEncoder pass (and
encodeSignalEvent) to the still-open split buffer, and the split's own
submission carries the copy behind the kernels that wrote its source. A source
with no open split buffer gets a standalone buffer that cross_enqueue submits
itself; the synchronous reference (MINFER_SYNC_COPIES=1) keeps the post-retire
order, because its host read must see the producer's finished output.
Measured on macbook (macOS 27.0.1, Apple M4 Pro) — the F5 S3 record in
docs/ARCHITECTURE-EXECUTION-PLAN.md.
11.6 The #300 result: the serialization is a missing dependency, and there is no overlap to gain
Ticket: #300 ("true cross-split overlap after #137"). It asked for either a
measured overlap or a recorded negative result beside the §11.5 note. The result
is a negative one, with the mechanism measured on
macbook (macOS 27.0.1, Apple M4 Pro) at ad707c7 (2026-10-09), recorded in
9df405d.
The #137 divergence is a missing-dependency race, not a kernel perturbation.
§11.5 inferred from the then-red baseline that "the extra in-flight command buffer
perturbed Metal kernel execution". Isolated, the mechanism is simpler and
deterministic. The pathological first cut is a standalone boundary command buffer
submitted from copy_cross before retire submits the producer's split
buffer, with no dependency on that buffer. On the shared MpsState command
queue the standalone buffer is committed first, so the blit reads the producer's
StorageModeShared source window before the producer's kernels wrote it, and
the consumer's wait on an already-signaled event returns the previous forward's
bytes. On models::qwen2::graph::tests::offload_copy::async_cross_copies_never_block_and_stay_bitwise_identical_on_metal
that cut reads max |Δlogit| = 26.718678 at step 0, the same value on two runs
(the §11.5 "staged bytes identical" check was stale-to-stale, which is why it
looked consistent).
Option A — a separate boundary buffer with an explicit dependency — is
correct. Committing the boundary blit on its own MTLCommandQueue, with
encodeWaitForEvent on an event the producer signals at its split end before the
blit and its own event for the consumer, restores bitwise identity: the real-model
gate reads max |Δlogit| = 0 over the prefill + 6-decode loop (twice) with the
counters unchanged — copies=35 waits=35 deferred_waits=35 blocking_host_copies=0 async_host_copies=21 event_syncs=21 sync_readbacks=0 async against
copies=35 waits=35 blocking_host_copies=21 sync_readbacks=21 sync — and the cheap
3-node gate passes (copies=2 waits=2 deferred=2 blocking=0 async_host=1 event_syncs=1 sync_readbacks=0). So the perturbation is avoidable: it is a
dependency the shared queue did not supply, not a fragility of Metal kernel
execution.
But Option A buys no measurable cross-split overlap, so the production design
stays. Cross-split overlap needs a second split whose work can run beside the
copy. On the reachable macOS topology there is none: the E5 mixed graph (4 of 24
blocks on the device) is a single CPU → Metal → CPU sequence, so there is exactly
one device→host boundary per forward; that boundary's consumer is the host
CPU split, whose first node directly reads the dominant staged tensor (the hidden
state), so the copy is on the consumer's critical path by definition; and the
boundary's other two staged tensors are read within the CPU split's first eight
nodes (add at node 46, cells at 52, attn_span at 54 in the decode graph), so
a separate buffer could only overlap a blit of two tiny index/window tensors with
a handful of host-side node setups. A general non-blocking producer retire
cannot be made safe without reordering the next device submission after the
boundary blit, which reintroduces exactly the serialization it would remove. A
device→device pair (Metal→Metal) would be the one topology with real overlap, and
copy_across early-returns on it while no CUDA device exists on macOS — it stays
reserved (§11.3). The boundary blit therefore keeps riding the producer's command
buffer; it is bitwise, and on this topology it is not slower than Option A.
Reproduction. Both cuts are MINFER_300_*-gated temporary patches to
MetalBackend::cross_enqueue (a standalone cmd_buffer() + submit() before
retire for the pathological cut; a second queue + encodeWaitForEvent /
encodeSignalEvent for Option A) — neither is in the tree. The acceptance command
is the same in every arm:
cargo test --release --bin minfer async_cross_copies_never_block_and_stay_bitwise_identical_on_metal -- --ignored --test-threads=1 --nocapture
Decisions governing this document
This page is the current contract; the decisions behind it are frozen in the ADR corpus:
- ADR-0001 — Inference runs through one declarative compute graph
- ADR-0011 — Backend ids are a file-format contract: appended, never renumbered
- ADR-0006 — The KV storage format is a per-engine gate, not a process-wide global
- ADR-0008 — GPU safety: bounded waits, no early return past a barrier, runtime device limits
- ADR-0009 — A failure is an error, never a silent fallback