The FW3.4 dense-Arwic pair triggered the +/-20% stop rule (+33.5% CPU
p50, 14x frame allocation). This slice removes the three measured
costs without changing GPU command order (the referee suites assert
identical recorded call sequences):
- WalkFrameDriver: Collect (ONE walk per frame - no GPU work; leaf
calls and flush points become a recorded event list; the driver
absorbed the renderer collection pass and exposes the visited sets)
+ Replay (prepare the whole stream once, then replay events,
interleaving DrawOrderedRange with leaf calls in the exact recorded
order). RunFrame = Collect+Replay for existing callers.
- WbDrawDispatcher: SubmitOrderedStream split into PrepareOrderedStream
(all sections + commands + merge runs uploaded once per frame) and
DrawOrderedRange (bind-once latch; per-run pipeline + DrawIdOffset +
DrawIndirectRangeRhi). Load-bearing correctness catch from the
implementation round: merge runs take FORCED BREAKS at the recorded
event marks - whole-stream merging must not fuse two segments that
retail separates with a leaf GPU call (shell, punch); the straddle
assert stays as a dead-code safety net.
- WalkProductionWorldData: WalkFrameStaticRecords carries an
ArraySegment into a per-frame grow-only arena; the per-cell
fresh-array copies (the 1.9 MB/frame alloc p50) are gone - zero
steady-state allocation after warmup.
Suites (lead-verified): full Release build 0 warnings; hermetic
6,758/0; Walk lane 209/1; InstalledDat Walk conformance 40/1
untouched. Next: the dense-Arwic re-measure against the same-session
baseline.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The walk-order submission layer over the existing RHI (plan section FW2):
- OrderedDrawStream: append-only walk-ordered draw commands
(GroupKey + transform + per-instance data + WalkDrawStage + cell
provenance), struct-of-arrays with one lockstep Reset (#193 shape).
The PortalPunch stage exists but has no FW2 submission path - the
submitter throws on it; punch emission lands with FW3 wiring.
- WbDrawDispatcher.OrderedStream partial: per-instance-first emission
(the deferred-alpha shape - command i owns instance i, walk order
survives into the indirect array), each SSBO section written once,
then one DrawIndirectRangeRhi call per maximal merge run. Runs are
built by pure-CPU BuildOrderedMergeRuns and may never span a stage,
pipeline-bucket, or cull boundary; ValidateMergeRun re-checks every
emitted run and throws (the campaign fail-loud rule). Nothing is
sorted, reordered, or dropped: N commands in, N indirect commands
out, covered exactly once.
- WorldDepthContract: retail world depth verified verbatim from the
decomp - Render::zfuncVal @0x00820e1c = 0x2, SetDepthBufferMode
@0x005a2d10 writes the enum directly as D3DRS_ZFUNC so the value IS
D3DCMP_LESS, applied by the surface-state applier @0x0059c80a with
Z-write toggled by blend; the LESSEQUAL sites are GameSky::Draw-local.
Seven world pipeline sites now cite the named constant (no value
changes).
- Plan updated: FW1 status block + gate amendment (the ten pose-stamped
retail traces supersede re-expressing the old-builder replay
fixtures; those retire with the old builder at FW4 and their
scenario classes re-verify at the FW3/FW4 connected gates).
Known FW2 scope notes recorded in the code: the building-detail
overlay replay is production wiring (FW3); the _drawCullModes scratch
may not interleave with a mid-flight RetailAlphaQueue scope (FW3
sequencing constraint). The pixel A/B equivalence proof rides FW3's
cutover toggle where a walk-driven scene first exists.
Suites: full Release build 0 warnings; Walk lane 154/1 skip;
hermetic 6,714/0 (+27 new).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>