acdream/tests/AcDream.App.Tests/Rendering/Gpu/RecordingGpuDevice.cs
Erik eced67d038 feat(render): Campaign V slice V6l commit 2 - the portal mask draws on Vulkan
Contract amendment 2 of three, and V4g's remaining half behind it. Plan section
5.5.16 defect 2: PortalDepthMaskRenderer's two-pass punch (#117) is built on
glStencilFunc/glStencilOp/glStencilMask, GpuPipelineDescription carried no
stencil state at all, and nothing else can express it - so the renderer stayed
raw GL, invisible to the Vulkan arm, and V4g's "stencil/depth-mask pipelines"
row could not be written.

The amendment splits the way core Vulkan 1.3 splits. The ENABLE and the
attachment intent are baked: GpuPipelineDescription.StencilTest, false by
default so no pipeline in the tree changed. The per-draw compare, three outcome
ops, reference and both masks are a GpuStencilState that the pipeline carries as
a DEFAULT and IGpuPassEncoder.SetStencil overrides - exactly the split cull
mode, front face and depth write already have, and exactly what
VK_DYNAMIC_STATE_STENCIL_OP/_COMPARE_MASK/_WRITE_MASK/_REFERENCE make dynamic.
The four stencil dynamic states are declared ONLY by a pipeline that tests
stencil: declaring a dynamic state obliges every draw with the pipeline to have
set it, so adding them unconditionally would make every existing pipeline depend
on a call none of them make. GpuStencilOp carries three values because the punch
uses three - Replace marks, Equal gates, Zero self-cleans - and a fourth would
be a facility with no consumer.

The arm. Three pipelines, not one, because depth COMPARE is not dynamic in the
contract and the punch's two passes differ in it: mark tests LEQUAL and writes
no depth, punch tests ALWAYS and writes, seal is ALWAYS + write with no stencil.
All three write no colour, which is what retail's "COLOR-INVISIBLE triangle fan"
means. The fan is expanded to a triangle LIST on the CPU - the contract has no
fan topology and Vulkan's is not portable - which is exact: triangle i is
(v0, v[i+1], v[i+2]), the same triangles in the same order.

portal_depth.{vert,frag} is a new committed shader pair, and this is the ONE
renderer in the campaign whose two arms do not share a source. Its clip planes
have to travel in the TerrainClip uniform block at binding 2, which is already
precisely this shape and already read by terrain_modern.vert and sky.vert - but
on GL that binding is held globally by ClipFrame for terrain, so a portal draw
that rebound it would leave every later terrain draw in the frame reading the
wrong region. The GL arm therefore keeps its inline program.
PortalDepthShaderParityTests is the tripwire: retail's far-Z constant
(0.99999988, from DrawPortalPolyInternal 0x0059bc90), #129's capped mark-bias
expression and the eight-half-plane loop are asserted to appear in both. Both
are deleted at V11. 9/10 shader pairs now compile to SPIR-V.

Two GL-side gaps closed while the state was being extended, both of section 7.1
rule 1's class rather than new work. GlAmbientCapabilityState now saves and
restores the stencil test, function, ops and both masks - the portal punch draws
mid-frame among renderers that are still raw GL and assume the test is off - and
the COLOUR MASK, which had no consumer until a colour-invisible pipeline existed
and whose absence would have blacked out every raw-GL renderer after such a
pass.

PortalTunnelPresentation was re-read and confirmed as V6k left it: it clears
depth and draws into the active viewport, binds no framebuffer of its own, and
needs no port for section 5.4's sake. It remains unported on the Vulkan arm -
the composition uses NullLocalPlayerTeleportPresentation there - which is an
absence on the V7 list, not a defect.

Gates. Release build green. App tests 4,129/3 skips; complete Release suite
9,192/5 (one solution-wide run reported a single App failure that did not
reproduce in two subsequent runs, solution-wide or alone - the documented
rerun-singly flake class). Strict GL offline pixel gate against 08ffe141:
2.31e-05, 13 differing pixels of 563,200, inside the documented 9-31 band. GL
connected -Runs 3: 3/3 RENDERED on the desktop witness and 3/3 on the client
capture. One offline Vulkan run with VK_LAYER_KHRONOS_validation proven inserted
by the loader: zero validation errors, zero warnings, a captured world frame.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 17:36:45 +02:00

547 lines
19 KiB
C#

using System.Numerics;
using AcDream.App.Rendering;
using AcDream.App.Rendering.Gpu;
namespace AcDream.App.Tests.Rendering.Gpu;
/// <summary>
/// One recorded RHI call. Renderer tests assert against the ordered sequence
/// instead of against a live driver, which is what keeps the App suite runnable
/// on a machine with no GPU while renderers migrate onto <see cref="IGpuDevice"/>
/// during Campaign V.
/// </summary>
internal abstract record GpuRecordedCall;
internal sealed record GpuRecordedFrameBegin(long Serial, int SlotIndex) : GpuRecordedCall;
internal sealed record GpuRecordedFrameEnd(long Serial) : GpuRecordedCall;
internal sealed record GpuRecordedRingAllocation(GpuRingUsage Usage, int ByteCount, uint OffsetBytes) : GpuRecordedCall;
internal sealed record GpuRecordedPassBegin(string Name, int SampleCount) : GpuRecordedCall;
internal sealed record GpuRecordedPassEnd(string Name) : GpuRecordedCall;
internal sealed record GpuRecordedPipelineBind(string PipelineName) : GpuRecordedCall;
internal sealed record GpuRecordedStorageBind(uint Binding, string BufferName, uint OffsetBytes, uint SizeBytes)
: GpuRecordedCall;
internal sealed record GpuRecordedUniformBind(uint Binding, string BufferName, uint OffsetBytes, uint SizeBytes)
: GpuRecordedCall;
internal sealed record GpuRecordedVertexBind(uint Binding, string BufferName, uint OffsetBytes) : GpuRecordedCall;
internal sealed record GpuRecordedStencil(GpuStencilState Stencil) : GpuRecordedCall;
internal sealed record GpuRecordedIndexBind(string BufferName, uint OffsetBytes, GpuIndexType IndexType)
: GpuRecordedCall;
internal sealed record GpuRecordedPushConstants(GpuPushConstants Constants) : GpuRecordedCall;
internal sealed record GpuRecordedViewport(int X, int Y, int Width, int Height) : GpuRecordedCall;
internal sealed record GpuRecordedScissor(int X, int Y, int Width, int Height) : GpuRecordedCall;
internal sealed record GpuRecordedCullMode(GpuCullMode CullMode) : GpuRecordedCall;
internal sealed record GpuRecordedFrontFace(GpuFrontFace FrontFace) : GpuRecordedCall;
internal sealed record GpuRecordedDepthWrite(bool Enabled) : GpuRecordedCall;
internal sealed record GpuRecordedDrawIndexed(
uint IndexCount,
uint InstanceCount,
uint FirstIndex,
int VertexOffset,
uint FirstInstance) : GpuRecordedCall;
internal sealed record GpuRecordedDraw(
uint VertexCount,
uint InstanceCount,
uint FirstVertex,
uint FirstInstance) : GpuRecordedCall;
internal sealed record GpuRecordedMultiDrawIndirect(
string BufferName,
uint OffsetBytes,
uint DrawCount,
uint StrideBytes) : GpuRecordedCall;
internal sealed record GpuRecordedTextureRegistration(string TextureName, GpuSamplerDescription Sampler, uint Slot)
: GpuRecordedCall;
internal sealed record GpuRecordedTextureRelease(uint Slot) : GpuRecordedCall;
/// <summary>
/// In-memory <see cref="IGpuDevice"/> that owns no driver objects. Ring
/// allocations are backed by a real byte array, so a test can drive a renderer
/// and then read back exactly what it wrote — the same bytes a driver would have
/// seen. Everything else is recorded into <see cref="Calls"/> in submission order.
/// </summary>
internal sealed class RecordingGpuDevice : IGpuDevice
{
private const int DefaultRingCapacityBytes = 8 * 1024 * 1024;
private readonly List<GpuRecordedCall> _calls = [];
private readonly List<Action> _queuedActions = [];
private readonly Dictionary<GpuSamplerDescription, RecordingGpuSampler> _samplers = [];
private readonly byte[] _ring;
private readonly Stack<uint> _freeTextureSlots = new();
private uint _nextTextureSlot;
private uint _ringCursor;
private long _serial;
private RecordingGpuFrame? _openFrame;
private bool _disposed;
public RecordingGpuDevice(int ringCapacityBytes = DefaultRingCapacityBytes)
{
ArgumentOutOfRangeException.ThrowIfLessThan(ringCapacityBytes, 1);
_ring = new byte[ringCapacityBytes];
RingBuffer = new RecordingGpuBuffer(new GpuBufferDescription(
"test-ring",
ringCapacityBytes,
GpuBufferUsage.Storage | GpuBufferUsage.Uniform | GpuBufferUsage.Indirect,
GpuMemoryResidency.HostWritable));
RecordingGpuTexture placeholder = new("default-white", GpuTextureKind.Texture2D, GpuTextureFormat.Rgba8Unorm, 1, 1, 1, 1);
DefaultTextureSlot = RegisterTexture(placeholder, CreateSampler(GpuSamplerDescription.UiNearest));
}
/// <summary>Every recorded call, in submission order.</summary>
public IReadOnlyList<GpuRecordedCall> Calls => _calls;
/// <summary>Backing store for ring allocations, so tests can read what a renderer wrote.</summary>
public ReadOnlySpan<byte> RingBytes => _ring;
/// <summary>Number of ring bytes handed out during the currently open (or most recent) frame.</summary>
public uint RingBytesAllocated => _ringCursor;
public int OpenFrameCount { get; private set; }
public int LiveTextureSlotCount => (int)_nextTextureSlot - _freeTextureSlots.Count;
public GpuBackendKind Backend => GpuBackendKind.Recording;
public GpuCapabilityRecord Capabilities { get; init; } = new()
{
Backend = GpuBackendKind.Recording,
DeviceName = "recording",
DriverInfo = "in-memory test double",
ApiVersion = "n/a",
MaxTextureTableSlots = GpuBindingModel.TextureTableCapacity,
MaxStorageBufferBindings = GpuBindingModel.StorageBindingCount,
MaxPushConstantBytes = GpuBindingModel.MaxPushConstantBytes,
MinStorageBufferOffsetAlignment = 256,
MinUniformBufferOffsetAlignment = 256,
MaxClipDistances = GpuBindingModel.ClipPlanesPerSlot,
MaxSampleCount = 8,
SupportsMultiDrawIndirect = true,
SupportsDrawParameters = true,
SupportsTextureCompressionBc = true,
SupportsTimestampQueries = true,
SupportsPersistentlyMappedRings = true,
};
public IGpuResourceRetirementQueue Retirement => ImmediateGpuResourceRetirementQueue.Instance;
public IGpuTimerPool Timers { get; } = new RecordingGpuTimerPool();
public GpuTextureSlot DefaultTextureSlot { get; }
public void Clear() => _calls.Clear();
public IGpuBuffer CreateBuffer(in GpuBufferDescription description) =>
new RecordingGpuBuffer(description);
/// <summary>
/// Campaign V slice V6i-2: every image this device made, in creation order.
/// A caller that creates its own textures internally — the world texture
/// arrays and the terrain atlas do — has no other way to assert what landed
/// on them.
/// </summary>
public IReadOnlyList<RecordingGpuTexture> CreatedTextures => _createdTextures;
private readonly List<RecordingGpuTexture> _createdTextures = [];
public IGpuTexture CreateTexture(in GpuTextureDescription description)
{
RecordingGpuTexture texture = new(
description.Name,
description.Kind,
description.Format,
description.Width,
description.Height,
description.LayerCount,
description.MipLevelCount);
_createdTextures.Add(texture);
return texture;
}
public IGpuSampler CreateSampler(in GpuSamplerDescription description)
{
if (_samplers.TryGetValue(description, out RecordingGpuSampler? existing))
return existing;
RecordingGpuSampler created = new(description);
_samplers.Add(description, created);
return created;
}
public IGpuPipeline CreatePipeline(GpuPipelineDescription description)
{
ArgumentNullException.ThrowIfNull(description);
return new RecordingGpuPipeline(description);
}
public IGpuRenderTarget CreateRenderTarget(in GpuRenderTargetDescription description) =>
new RecordingGpuRenderTarget(description);
public GpuTextureSlot RegisterTexture(IGpuTexture texture, IGpuSampler sampler)
{
ArgumentNullException.ThrowIfNull(texture);
ArgumentNullException.ThrowIfNull(sampler);
uint slot = _freeTextureSlots.Count > 0 ? _freeTextureSlots.Pop() : _nextTextureSlot++;
_calls.Add(new GpuRecordedTextureRegistration(texture.Name, sampler.Description, slot));
return new GpuTextureSlot(slot);
}
public void ReleaseTextureSlot(GpuTextureSlot slot)
{
if (!slot.IsAssigned)
throw new ArgumentException("Cannot release an unassigned texture slot.", nameof(slot));
_freeTextureSlots.Push(slot.Index);
_calls.Add(new GpuRecordedTextureRelease(slot.Index));
}
public IGpuFrame BeginFrame()
{
ObjectDisposedException.ThrowIf(_disposed, this);
if (_openFrame is not null)
throw new InvalidOperationException("The previous frame must end before another begins.");
_ringCursor = 0;
long serial = ++_serial;
int slotIndex = (int)((serial - 1) % 2);
_calls.Add(new GpuRecordedFrameBegin(serial, slotIndex));
OpenFrameCount++;
_openFrame = new RecordingGpuFrame(this, serial, slotIndex);
return _openFrame;
}
public void QueueDeviceAction(Action action)
{
ArgumentNullException.ThrowIfNull(action);
_queuedActions.Add(action);
}
public void ProcessDeviceActions()
{
Action[] pending = [.. _queuedActions];
_queuedActions.Clear();
foreach (Action action in pending)
action();
}
public byte[] CaptureBackbuffer(int width, int height)
{
ArgumentOutOfRangeException.ThrowIfNegativeOrZero(width);
ArgumentOutOfRangeException.ThrowIfNegativeOrZero(height);
return new byte[checked(width * height * 4)];
}
public void WaitIdle() => ProcessDeviceActions();
public void Dispose() => _disposed = true;
internal void Record(GpuRecordedCall call) => _calls.Add(call);
internal GpuRingAllocation Allocate(int byteCount, GpuRingUsage usage)
{
ArgumentOutOfRangeException.ThrowIfNegative(byteCount);
uint alignment = usage switch
{
GpuRingUsage.Storage => Capabilities.MinStorageBufferOffsetAlignment,
GpuRingUsage.Uniform => Capabilities.MinUniformBufferOffsetAlignment,
_ => 4u,
};
uint aligned = AlignUp(_ringCursor, alignment);
if (aligned + (uint)byteCount > (uint)_ring.Length)
{
throw new InvalidOperationException(
$"Ring allocation of {byteCount} bytes for {usage} exceeds the {_ring.Length}-byte test ring.");
}
_ringCursor = aligned + (uint)byteCount;
_calls.Add(new GpuRecordedRingAllocation(usage, byteCount, aligned));
return new GpuRingAllocation(RingBuffer, aligned, _ring.AsSpan((int)aligned, byteCount));
}
internal IGpuBuffer RingBuffer { get; }
internal void CloseFrame(RecordingGpuFrame frame)
{
if (!ReferenceEquals(_openFrame, frame))
return;
_calls.Add(new GpuRecordedFrameEnd(frame.Serial));
OpenFrameCount--;
_openFrame = null;
}
private static uint AlignUp(uint value, uint alignment) =>
alignment <= 1 ? value : (value + alignment - 1) / alignment * alignment;
}
internal sealed class RecordingGpuFrame(RecordingGpuDevice device, long serial, int slotIndex) : IGpuFrame
{
private bool _ended;
public int SlotIndex { get; } = slotIndex;
public long Serial { get; } = serial;
public GpuRingAllocation AllocateRing(int byteCount, GpuRingUsage usage) => device.Allocate(byteCount, usage);
public IGpuPassEncoder BeginPass(GpuPassDescription description)
{
ArgumentNullException.ThrowIfNull(description);
device.Record(new GpuRecordedPassBegin(description.Name, description.SampleCount));
return new RecordingGpuPassEncoder(device, description);
}
public void End()
{
if (_ended)
return;
_ended = true;
device.CloseFrame(this);
}
public void Dispose() => End();
}
internal sealed class RecordingGpuPassEncoder(RecordingGpuDevice device, GpuPassDescription pass) : IGpuPassEncoder
{
private bool _closed;
public GpuPassDescription Pass { get; } = pass;
public void BindPipeline(IGpuPipeline pipeline)
{
ArgumentNullException.ThrowIfNull(pipeline);
device.Record(new GpuRecordedPipelineBind(pipeline.Description.Name));
}
public void BindStorageBuffer(uint binding, IGpuBuffer buffer, uint offsetBytes, uint sizeBytes)
{
ArgumentNullException.ThrowIfNull(buffer);
device.Record(new GpuRecordedStorageBind(binding, buffer.Name, offsetBytes, sizeBytes));
}
public void BindUniformBuffer(uint binding, IGpuBuffer buffer, uint offsetBytes, uint sizeBytes)
{
ArgumentNullException.ThrowIfNull(buffer);
device.Record(new GpuRecordedUniformBind(binding, buffer.Name, offsetBytes, sizeBytes));
}
public void BindVertexBuffer(uint binding, IGpuBuffer buffer, uint offsetBytes)
{
ArgumentNullException.ThrowIfNull(buffer);
device.Record(new GpuRecordedVertexBind(binding, buffer.Name, offsetBytes));
}
public void BindIndexBuffer(IGpuBuffer buffer, uint offsetBytes, GpuIndexType indexType)
{
ArgumentNullException.ThrowIfNull(buffer);
device.Record(new GpuRecordedIndexBind(buffer.Name, offsetBytes, indexType));
}
public void SetPushConstants(in GpuPushConstants constants) =>
device.Record(new GpuRecordedPushConstants(constants));
public void SetViewport(int x, int y, int width, int height) =>
device.Record(new GpuRecordedViewport(x, y, width, height));
public void SetScissor(int x, int y, int width, int height) =>
device.Record(new GpuRecordedScissor(x, y, width, height));
public void SetCullMode(GpuCullMode cullMode) => device.Record(new GpuRecordedCullMode(cullMode));
public void SetFrontFace(GpuFrontFace frontFace) => device.Record(new GpuRecordedFrontFace(frontFace));
public void SetStencil(in GpuStencilState stencil) => device.Record(new GpuRecordedStencil(stencil));
public void SetDepthWrite(bool enabled) => device.Record(new GpuRecordedDepthWrite(enabled));
public void DrawIndexed(uint indexCount, uint instanceCount, uint firstIndex, int vertexOffset, uint firstInstance) =>
device.Record(new GpuRecordedDrawIndexed(indexCount, instanceCount, firstIndex, vertexOffset, firstInstance));
public void Draw(uint vertexCount, uint instanceCount, uint firstVertex, uint firstInstance) =>
device.Record(new GpuRecordedDraw(vertexCount, instanceCount, firstVertex, firstInstance));
public void MultiDrawIndexedIndirect(IGpuBuffer commands, uint offsetBytes, uint drawCount, uint strideBytes)
{
ArgumentNullException.ThrowIfNull(commands);
device.Record(new GpuRecordedMultiDrawIndirect(commands.Name, offsetBytes, drawCount, strideBytes));
}
public IDisposable BeginTimerScope(string scopeName) => NullDisposable.Instance;
public void Dispose()
{
if (_closed)
return;
_closed = true;
device.Record(new GpuRecordedPassEnd(Pass.Name));
}
private sealed class NullDisposable : IDisposable
{
public static NullDisposable Instance { get; } = new();
public void Dispose()
{
}
}
}
internal sealed class RecordingGpuBuffer(GpuBufferDescription description) : IGpuBuffer
{
private readonly byte[] _storage = new byte[description.SizeBytes];
public string Name { get; } = description.Name;
public long SizeBytes { get; } = description.SizeBytes;
public GpuBufferUsage Usage { get; } = description.Usage;
public GpuMemoryResidency Residency { get; } = description.Residency;
public bool IsDisposed { get; private set; }
public void Upload(long offsetBytes, ReadOnlySpan<byte> data) =>
data.CopyTo(_storage.AsSpan((int)offsetBytes, data.Length));
public void CopyTo(IGpuBuffer destination, long sourceOffsetBytes, long destinationOffsetBytes, long byteCount)
{
ArgumentNullException.ThrowIfNull(destination);
if (destination is not RecordingGpuBuffer target)
throw new ArgumentException("Recording buffers can only copy to recording buffers.", nameof(destination));
_storage.AsSpan((int)sourceOffsetBytes, (int)byteCount)
.CopyTo(target._storage.AsSpan((int)destinationOffsetBytes, (int)byteCount));
}
public void Read(long offsetBytes, Span<byte> destination) =>
_storage.AsSpan((int)offsetBytes, destination.Length).CopyTo(destination);
public void Dispose() => IsDisposed = true;
}
internal sealed class RecordingGpuTexture(
string name,
GpuTextureKind kind,
GpuTextureFormat format,
int width,
int height,
int layerCount,
int mipLevelCount) : IGpuTexture
{
private readonly List<(int MipLevel, int Layer, int ByteCount)> _uploads = [];
public string Name { get; } = name;
public GpuTextureKind Kind { get; } = kind;
public GpuTextureFormat Format { get; } = format;
public int Width { get; } = width;
public int Height { get; } = height;
public int LayerCount { get; } = layerCount;
public int MipLevelCount { get; } = mipLevelCount;
public bool MipChainGenerated { get; private set; }
public bool IsDisposed { get; private set; }
public IReadOnlyList<(int MipLevel, int Layer, int ByteCount)> Uploads => _uploads;
public void Upload(int mipLevel, int layer, ReadOnlySpan<byte> data) =>
_uploads.Add((mipLevel, layer, data.Length));
public void GenerateMipChain() => MipChainGenerated = true;
public void Dispose() => IsDisposed = true;
}
internal sealed class RecordingGpuSampler(GpuSamplerDescription description) : IGpuSampler
{
public GpuSamplerDescription Description { get; } = description;
public bool IsDisposed { get; private set; }
public void Dispose() => IsDisposed = true;
}
internal sealed class RecordingGpuPipeline(GpuPipelineDescription description) : IGpuPipeline
{
public GpuPipelineDescription Description { get; } = description;
public bool IsDisposed { get; private set; }
public void Dispose() => IsDisposed = true;
}
internal sealed class RecordingGpuRenderTarget : IGpuRenderTarget
{
public RecordingGpuRenderTarget(GpuRenderTargetDescription description)
{
Description = description;
ColorTexture = new RecordingGpuTexture(
$"{description.Name}-color",
GpuTextureKind.Texture2D,
description.ColorFormat,
description.Width,
description.Height,
layerCount: 1,
mipLevelCount: 1);
}
public GpuRenderTargetDescription Description { get; }
public IGpuTexture ColorTexture { get; }
public bool IsDisposed { get; private set; }
public void Dispose() => IsDisposed = true;
}
internal sealed class RecordingGpuTimerPool : IGpuTimerPool
{
public bool IsSupported => false;
public bool TryResolve(string scopeName, out double milliseconds)
{
milliseconds = 0d;
return false;
}
}
/// <summary>Convenience helpers so renderer tests read as assertions, not as list surgery.</summary>
internal static class RecordingGpuDeviceAssertions
{
public static IEnumerable<T> OfKind<T>(this RecordingGpuDevice device) where T : GpuRecordedCall =>
device.Calls.OfType<T>();
public static Vector4 ClearColorOf(this GpuPassDescription pass) => pass.Color.ClearColor;
}