fix(streaming): #339 — a packed EnvCell geom id no longer misroutes into the 32-bit prepare arm
Some checks are pending
Headless portability / portable-headless (ubuntu-latest) (push) Waiting to run
Headless portability / portable-headless (windows-latest) (push) Waiting to run
Headless portability / linux-graphical (push) Waiting to run
Headless portability / linux-vulkan (push) Waiting to run

The reveal hang (three live occurrences: portal-space, login twice) was
an unhandled OverflowException on the render frame's readiness
evaluation: EnsureRenderDataReady found a packed 64-bit EnvCell geometry
id OWNED but DESCRIPTOR-LESS — the release path removes the descriptor
while render data parks on the LRU, IncrementRefCount restores ownership
on a revisit, and the scheduler's PrepareEnvCellGeomMeshDataAsync has
not yet re-registered — and fell through to the Setup/GfxObj arm, whose
checked((uint)id) cast threw. After that the reveal was never evaluated
again.

The fix corrects the TYPE DISPATCH rather than suppressing anything:
packed ids (bit 33, GetEnvCellGeomId) answer "not yet" in the
acquire-to-prepare window — the true answer, since the scheduler
re-registers on the same landblock build — and PrepareMeshDataAsync's
blind cast becomes a typed, loud invariant failure naming the id kind
and the issue, so a future caller repeating the confusion gets a
diagnosis instead of three live hangs.

Validation: the crash was DETERMINISTIC at login cell 0xA8B4002F (two
consecutive hard failures); with the fix the same login revealed
cleanly and a full session — 16,585 entities, five portal generations,
the user-passed Session-B dungeon gate — ran with zero overflows and
zero guard fires. Clean-room suite: 11,253 passed / 6 skipped /
0 failed.

Found because the Session-B gate launch finally captured the stack the
earlier #339 hangs never printed. #343 (the wounded loop's Reset-in-
render-loop shutdown) and #344 (the mid-teleport world-frame race the
same evening surfaced) are filed separately and unfixed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Erik 2026-08-07 08:37:43 +02:00
parent 42de5f18ff
commit 205379c6d6
2 changed files with 51 additions and 1 deletions

View file

@ -356,7 +356,21 @@ after which the reveal is never evaluated again. Demote-then-revisit
ordering explains the intermittency and why the same destination passed
twice earlier: the window only exists after an evict/re-acquire cycle.
**Status: mechanism established; ROOT-CAUSE FIX NOT YET DESIGNED.** The fix
**Status: FIXED 2026-08-07 evening, LIVE-VALIDATED the same night.** The fix
is the type-dispatch correction, not a suppression: `EnsureRenderDataReady`
now answers "not yet" for a packed id in the acquire→prepare window (the
scheduler's `PrepareEnvCellGeomMeshDataAsync` re-registers the descriptor on
the same landblock build, so not-ready is the true answer, not a dodge), and
`PrepareMeshDataAsync` converts the blind checked cast into a typed, loud
invariant failure naming the id kind and this issue. Validation: the crash
was DETERMINISTIC at login cell 0xA8B4002F (two consecutive hard failures);
with the fix, the same login revealed cleanly and a full play session
(16,585 entities, five portal generations, the Session-B dungeon gate) ran
with zero overflows and zero guard fires. The original design note below is
retained; the "band-aid" candidates it names remain rejected — this fix
corrects the dispatch, it does not swallow the error.
**Original status: mechanism established; ROOT-CAUSE FIX NOT YET DESIGNED.** The fix
is NOT "catch the exception" and NOT "skip 64-bit ids" (band-aids both):
the acquire/re-request ordering must make owned-implies-descriptor an
invariant again, or EnsureRenderDataReady must legitimately re-request the