The verdict splits, and not where anyone expected. Measured on one machine, one
build, one day, physical console, uncapped, 4x MSAA on both arms, validation off:
Stationary LIGHT scene (6,675 entities, 780-870 FPS)
CPU p50 GL 1.127 ms -> VK 1.294 ms MISS +14.8%
CPU p99 GL 1.407 ms -> VK 1.531 ms MISS +8.8%
GPU p50 GL 0.651 ms -> VK 0.160 ms PASS -75.4%
Alloc/frm GL 77,664 B -> VK 11,440 B PASS -85.3%
Working GL 943.7 MiB -> VK 877.1 MiB PASS -7.0%
Stationary DENSE scene (21,024 entities, identical on both arms)
CPU p50 GL 5.934 ms -> VK 5.775 ms PASS -2.7%
CPU p99 GL 8.867 ms -> VK 7.354 ms PASS -17.1%
GPU p50 GL 1.673 ms -> VK 0.909 ms PASS -45.7%
Alloc/frm GL 82,016 B -> VK 15,752 B PASS -80.8%
Process CPU GL 1.246 cores -> VK 1.016 PASS -18.5% (Windows, not ours)
Canonical nine-stop route, identical world at all nine stops
Frames GL 30,378 -> VK 38,683 PASS +27.3%
CPU p50 GL 11.718 ms -> VK 9.166 ms PASS -21.8%
GPU p99 GL 2.325 ms -> VK 1.193 ms PASS -48.7%
Vulkan loses two rows in exactly one configuration: a stationary field at a frame
rate no player will ever see. The reason is measured rather than argued. A
temporary probe on BOTH arms, now stripped, attributes 0.148 ms/frame to required
Vulkan WSI and synchronisation calls - vkQueuePresentKHR 0.070, vkQueueSubmit2
0.027, the timeline wait 0.026, vkAcquireNextImageKHR 0.025 - against roughly
0.014 ms for GL's whole SwapBuffers. That cost is FIXED per frame, so it is 12%
of a 1.13 ms frame, 2.5% of a 5.9 ms one and under 1% of a dense-town frame,
while the GPU and allocation savings scale with the work. The sign of the CPU
comparison flips as soon as the frame contains a town.
The campaign's named cost centre is closed rather than carried a fourth time.
Bindings 4, 6, 7 and 8 costing a descriptor write per draw - forward-carried
since V6i-3 as the thing to fix if CPU were short - measures 0.031 ms for ALL
~216 draws of the frame, about 140 ns each and 2.4% of it. No Vulkan code was
changed to chase the miss: every lever the V8 row named was already taken
(coherent rings, one submit per frame), irrelevant to p50 (pipeline pre-warm),
measured and small (descriptors), or would have traded real memory for nothing
(a fourth swapchain image, when acquire is call cost and not waiting).
The methodological finding is worth reading before the numbers. The R6 soak is
NOT the vehicle the founding numbers came from - the G5 production profile states
its own conditions and they exclude the probe, the artifact owner and the
screenshot oracle - and it is biased AGAINST Vulkan, because
VulkanGraphicsContext arms retainBackbufferCapture exactly when
ACDREAM_AUTOMATION_ARTIFACT_DIR is set, making every Vulkan frame copy the whole
swapchain image while GL reads on demand. On one binary in one hour the soak
reports CPU p50 7.3 ms and 2,531 KiB/frame where the ordinary profile reports
1.13 ms and 77 KiB. The route table above is therefore conservative in Vulkan's
favour: it wins on the vehicle that charges it extra.
Gates: Release build green; App tests 4,152 / 3 skipped, the pre-slice baseline,
no #250-family failure; strict GL offline pixel gate against 13c8733d at 1.95e-05
(11 px of 563,200), inside the 9-31 band, so GL did not move; one connected
Vulkan run with VK_LAYER_KHRONOS_validation proven inserted by the loader at zero
errors and zero warnings; and BOTH R6 soaks green - Vulkan 506.6 s and GL 506.8 s,
zero failures, graceful exits - which discharges the soak half of V7's
outstanding list. RenderDoc is not installed on this machine, so that capture
carries to V10 with a cause rather than as an omission.
The recommendation: proceed to V10 and amend the acceptance table rather than
waive it, naming the scene and pacing the floor is judged at. Two natural
candidates are already in the evidence and Vulkan passes both outright. The
opposite reading - that the light-scene rows disqualify the cutover - is
available and has been given the same measurement space. That call is the user's
and this slice does not make it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>