acdream/tests/AcDream.App.Tests/Rendering/Gpu/RecordingGpuDevice.cs
Erik c8d0f70bbe feat(render): Campaign V slice V6i-2 commit 2 — world texture creation crosses to IGpuTexture
Plan §5.5.11 recorded what V4t deliberately left behind: it moved the table
ENTRY of every world texture to the device and kept CREATION with the caches,
because "creating world textures through IGpuTexture is real remaining work and
it belongs with the Vulkan world arm, which is the first thing that cannot use a
GL handle at all." §5.5.12 item 1 handed it forward and named the missing piece
exactly — "an ITextureArray implementation over IGpuTexture, not a codec",
because V6b's BlockCompressionCodec and BlockCompressionMipChain already supply
the BC chains. This is that work.

IWorldTextureArray is the seam, and the slot is what crosses it. Before this
commit ObjectMeshManager read BindlessWrapHandle/BindlessClampHandle off the
concrete GL array and interned them into the device table itself. A 64-bit
ARB_bindless_texture handle has no Vulkan spelling, so the array now answers the
question the caller was really asking — ResolveSlot(wrapping) — and each arm gets
there its own way: ManagedGLTextureArray makes the same idempotent interning call
one level down, and RhiWorldTextureArray returns a pair it registered at
construction. ReleaseTextureSlots replaces the snapshot dictionary the manager
kept for the same reason, and still runs only once physical retirement completes.

Which implementation exists is decided ONCE, by the IWorldTextureArrayFactory
composition builds — plan §3.1's no-runtime-fork rule. Everything above the seam
(capacity policy, slot allocation, ref counting, layer retirement, empty-atlas
eviction, and the whole of ObjectMeshManager's atlas policy) is written once and
branches on nothing.

Three things the RHI array does differently, each because the backends genuinely
differ rather than by choice: BC mip chains are CPU-built through
BlockCompressionMipChain, since Vulkan cannot blit into a compressed image, while
RGBA8 uses the device's blit; filtering lives in an immutable sampler rather than
a texture parameter, so both address modes are registered up front exactly as the
GL array holds two resident handles; and RGB8/A8/Rgba32f are refused at creation
with the reason named. A8 is the interesting refusal — the GL array serves it by
swizzling R into A, and a Vulkan swizzle lives in the image VIEW, which the pinned
GpuTextureDescription does not describe. A silent substitution would render wrong
and look like a shader bug.

TerrainAtlas gains the second construction path V6i drafted and reverted. The
decode is factored out and shared, so both arms read the same DATs, in the same
order, with the same resize-to-max policy; only the upload forks.
ICompositeTextureArrayBackend gains its RHI arm, which is four small methods
because that seam was already a seam.

The Vulkan arm is EXERCISED, not merely present. That is the whole reason the
V6i draft was reverted rather than landed — "built then reverted because nothing
exercised it" — and it is the same failure §5.5.12 measured twice in the
descriptor layouts. So the composition host now builds the real terrain atlas
through IGpuDevice.CreateTexture on the arm with no GL context, and creates and
releases one shared array of each format family plus one composite array at
startup. Creation only; nothing draws them. Releasing them in the same statement
covers one thing a retained bundle would not — that both slot pairs come back and
the images route through the retirement queue.

Gates: Release build; App tests 4,104 / 3 skips; strict GL offline pixel gate vs
0ca802cd 3.20e-05 (18 px of 563,200, inside the documented 9–31 px control band);
GL connected tools/run-repeat-connected-gate.ps1 -Runs 3 at 3/3 RENDERED on the
desktop witness AND 3/3 on the client capture; one Vulkan composition-host run
with VK_LAYER_KHRONOS_validation proven inserted by the loader at zero errors,
zero warnings, no [shutdown] diagnostic, and a captured frame. That run built
terrain-atlas 512x512x33 with 10 mip levels, terrain-alpha-atlas 512x512x8, RGBA8
64x64x32 (slots 3/4, 174,720 mip bytes blitted), BC1 64x64x32 (slots 5/6, 696 mip
bytes encoded) and composite 32x32x8 (slot 7).

One whole-suite run failed Issue181WallPressEquilibriumTests once; it passed
alone and did not recur in five further runs. Seven test classes mutate the same
process-global CameraDiagnostics switches with no xUnit collection isolation, and
this diff touches no camera, visibility or physics code. A separate run of the
UNCHANGED parent tree failed a different zero-allocation test, which is `#250`'s
documented class. Both are filed rather than attributed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 13:57:43 +02:00

544 lines
19 KiB
C#

using System.Numerics;
using AcDream.App.Rendering;
using AcDream.App.Rendering.Gpu;
namespace AcDream.App.Tests.Rendering.Gpu;
/// <summary>
/// One recorded RHI call. Renderer tests assert against the ordered sequence
/// instead of against a live driver, which is what keeps the App suite runnable
/// on a machine with no GPU while renderers migrate onto <see cref="IGpuDevice"/>
/// during Campaign V.
/// </summary>
internal abstract record GpuRecordedCall;
internal sealed record GpuRecordedFrameBegin(long Serial, int SlotIndex) : GpuRecordedCall;
internal sealed record GpuRecordedFrameEnd(long Serial) : GpuRecordedCall;
internal sealed record GpuRecordedRingAllocation(GpuRingUsage Usage, int ByteCount, uint OffsetBytes) : GpuRecordedCall;
internal sealed record GpuRecordedPassBegin(string Name, int SampleCount) : GpuRecordedCall;
internal sealed record GpuRecordedPassEnd(string Name) : GpuRecordedCall;
internal sealed record GpuRecordedPipelineBind(string PipelineName) : GpuRecordedCall;
internal sealed record GpuRecordedStorageBind(uint Binding, string BufferName, uint OffsetBytes, uint SizeBytes)
: GpuRecordedCall;
internal sealed record GpuRecordedUniformBind(uint Binding, string BufferName, uint OffsetBytes, uint SizeBytes)
: GpuRecordedCall;
internal sealed record GpuRecordedVertexBind(string BufferName, uint OffsetBytes) : GpuRecordedCall;
internal sealed record GpuRecordedIndexBind(string BufferName, uint OffsetBytes, GpuIndexType IndexType)
: GpuRecordedCall;
internal sealed record GpuRecordedPushConstants(GpuPushConstants Constants) : GpuRecordedCall;
internal sealed record GpuRecordedViewport(int X, int Y, int Width, int Height) : GpuRecordedCall;
internal sealed record GpuRecordedScissor(int X, int Y, int Width, int Height) : GpuRecordedCall;
internal sealed record GpuRecordedCullMode(GpuCullMode CullMode) : GpuRecordedCall;
internal sealed record GpuRecordedFrontFace(GpuFrontFace FrontFace) : GpuRecordedCall;
internal sealed record GpuRecordedDepthWrite(bool Enabled) : GpuRecordedCall;
internal sealed record GpuRecordedDrawIndexed(
uint IndexCount,
uint InstanceCount,
uint FirstIndex,
int VertexOffset,
uint FirstInstance) : GpuRecordedCall;
internal sealed record GpuRecordedDraw(
uint VertexCount,
uint InstanceCount,
uint FirstVertex,
uint FirstInstance) : GpuRecordedCall;
internal sealed record GpuRecordedMultiDrawIndirect(
string BufferName,
uint OffsetBytes,
uint DrawCount,
uint StrideBytes) : GpuRecordedCall;
internal sealed record GpuRecordedTextureRegistration(string TextureName, GpuSamplerDescription Sampler, uint Slot)
: GpuRecordedCall;
internal sealed record GpuRecordedTextureRelease(uint Slot) : GpuRecordedCall;
/// <summary>
/// In-memory <see cref="IGpuDevice"/> that owns no driver objects. Ring
/// allocations are backed by a real byte array, so a test can drive a renderer
/// and then read back exactly what it wrote — the same bytes a driver would have
/// seen. Everything else is recorded into <see cref="Calls"/> in submission order.
/// </summary>
internal sealed class RecordingGpuDevice : IGpuDevice
{
private const int DefaultRingCapacityBytes = 8 * 1024 * 1024;
private readonly List<GpuRecordedCall> _calls = [];
private readonly List<Action> _queuedActions = [];
private readonly Dictionary<GpuSamplerDescription, RecordingGpuSampler> _samplers = [];
private readonly byte[] _ring;
private readonly Stack<uint> _freeTextureSlots = new();
private uint _nextTextureSlot;
private uint _ringCursor;
private long _serial;
private RecordingGpuFrame? _openFrame;
private bool _disposed;
public RecordingGpuDevice(int ringCapacityBytes = DefaultRingCapacityBytes)
{
ArgumentOutOfRangeException.ThrowIfLessThan(ringCapacityBytes, 1);
_ring = new byte[ringCapacityBytes];
RingBuffer = new RecordingGpuBuffer(new GpuBufferDescription(
"test-ring",
ringCapacityBytes,
GpuBufferUsage.Storage | GpuBufferUsage.Uniform | GpuBufferUsage.Indirect,
GpuMemoryResidency.HostWritable));
RecordingGpuTexture placeholder = new("default-white", GpuTextureKind.Texture2D, GpuTextureFormat.Rgba8Unorm, 1, 1, 1, 1);
DefaultTextureSlot = RegisterTexture(placeholder, CreateSampler(GpuSamplerDescription.UiNearest));
}
/// <summary>Every recorded call, in submission order.</summary>
public IReadOnlyList<GpuRecordedCall> Calls => _calls;
/// <summary>Backing store for ring allocations, so tests can read what a renderer wrote.</summary>
public ReadOnlySpan<byte> RingBytes => _ring;
/// <summary>Number of ring bytes handed out during the currently open (or most recent) frame.</summary>
public uint RingBytesAllocated => _ringCursor;
public int OpenFrameCount { get; private set; }
public int LiveTextureSlotCount => (int)_nextTextureSlot - _freeTextureSlots.Count;
public GpuBackendKind Backend => GpuBackendKind.Recording;
public GpuCapabilityRecord Capabilities { get; init; } = new()
{
Backend = GpuBackendKind.Recording,
DeviceName = "recording",
DriverInfo = "in-memory test double",
ApiVersion = "n/a",
MaxTextureTableSlots = GpuBindingModel.TextureTableCapacity,
MaxStorageBufferBindings = GpuBindingModel.StorageBindingCount,
MaxPushConstantBytes = GpuBindingModel.MaxPushConstantBytes,
MinStorageBufferOffsetAlignment = 256,
MinUniformBufferOffsetAlignment = 256,
MaxClipDistances = GpuBindingModel.ClipPlanesPerSlot,
MaxSampleCount = 8,
SupportsMultiDrawIndirect = true,
SupportsDrawParameters = true,
SupportsTextureCompressionBc = true,
SupportsTimestampQueries = true,
SupportsPersistentlyMappedRings = true,
};
public IGpuResourceRetirementQueue Retirement => ImmediateGpuResourceRetirementQueue.Instance;
public IGpuTimerPool Timers { get; } = new RecordingGpuTimerPool();
public GpuTextureSlot DefaultTextureSlot { get; }
public void Clear() => _calls.Clear();
public IGpuBuffer CreateBuffer(in GpuBufferDescription description) =>
new RecordingGpuBuffer(description);
/// <summary>
/// Campaign V slice V6i-2: every image this device made, in creation order.
/// A caller that creates its own textures internally — the world texture
/// arrays and the terrain atlas do — has no other way to assert what landed
/// on them.
/// </summary>
public IReadOnlyList<RecordingGpuTexture> CreatedTextures => _createdTextures;
private readonly List<RecordingGpuTexture> _createdTextures = [];
public IGpuTexture CreateTexture(in GpuTextureDescription description)
{
RecordingGpuTexture texture = new(
description.Name,
description.Kind,
description.Format,
description.Width,
description.Height,
description.LayerCount,
description.MipLevelCount);
_createdTextures.Add(texture);
return texture;
}
public IGpuSampler CreateSampler(in GpuSamplerDescription description)
{
if (_samplers.TryGetValue(description, out RecordingGpuSampler? existing))
return existing;
RecordingGpuSampler created = new(description);
_samplers.Add(description, created);
return created;
}
public IGpuPipeline CreatePipeline(GpuPipelineDescription description)
{
ArgumentNullException.ThrowIfNull(description);
return new RecordingGpuPipeline(description);
}
public IGpuRenderTarget CreateRenderTarget(in GpuRenderTargetDescription description) =>
new RecordingGpuRenderTarget(description);
public GpuTextureSlot RegisterTexture(IGpuTexture texture, IGpuSampler sampler)
{
ArgumentNullException.ThrowIfNull(texture);
ArgumentNullException.ThrowIfNull(sampler);
uint slot = _freeTextureSlots.Count > 0 ? _freeTextureSlots.Pop() : _nextTextureSlot++;
_calls.Add(new GpuRecordedTextureRegistration(texture.Name, sampler.Description, slot));
return new GpuTextureSlot(slot);
}
public void ReleaseTextureSlot(GpuTextureSlot slot)
{
if (!slot.IsAssigned)
throw new ArgumentException("Cannot release an unassigned texture slot.", nameof(slot));
_freeTextureSlots.Push(slot.Index);
_calls.Add(new GpuRecordedTextureRelease(slot.Index));
}
public IGpuFrame BeginFrame()
{
ObjectDisposedException.ThrowIf(_disposed, this);
if (_openFrame is not null)
throw new InvalidOperationException("The previous frame must end before another begins.");
_ringCursor = 0;
long serial = ++_serial;
int slotIndex = (int)((serial - 1) % 2);
_calls.Add(new GpuRecordedFrameBegin(serial, slotIndex));
OpenFrameCount++;
_openFrame = new RecordingGpuFrame(this, serial, slotIndex);
return _openFrame;
}
public void QueueDeviceAction(Action action)
{
ArgumentNullException.ThrowIfNull(action);
_queuedActions.Add(action);
}
public void ProcessDeviceActions()
{
Action[] pending = [.. _queuedActions];
_queuedActions.Clear();
foreach (Action action in pending)
action();
}
public byte[] CaptureBackbuffer(int width, int height)
{
ArgumentOutOfRangeException.ThrowIfNegativeOrZero(width);
ArgumentOutOfRangeException.ThrowIfNegativeOrZero(height);
return new byte[checked(width * height * 4)];
}
public void WaitIdle() => ProcessDeviceActions();
public void Dispose() => _disposed = true;
internal void Record(GpuRecordedCall call) => _calls.Add(call);
internal GpuRingAllocation Allocate(int byteCount, GpuRingUsage usage)
{
ArgumentOutOfRangeException.ThrowIfNegative(byteCount);
uint alignment = usage switch
{
GpuRingUsage.Storage => Capabilities.MinStorageBufferOffsetAlignment,
GpuRingUsage.Uniform => Capabilities.MinUniformBufferOffsetAlignment,
_ => 4u,
};
uint aligned = AlignUp(_ringCursor, alignment);
if (aligned + (uint)byteCount > (uint)_ring.Length)
{
throw new InvalidOperationException(
$"Ring allocation of {byteCount} bytes for {usage} exceeds the {_ring.Length}-byte test ring.");
}
_ringCursor = aligned + (uint)byteCount;
_calls.Add(new GpuRecordedRingAllocation(usage, byteCount, aligned));
return new GpuRingAllocation(RingBuffer, aligned, _ring.AsSpan((int)aligned, byteCount));
}
internal IGpuBuffer RingBuffer { get; }
internal void CloseFrame(RecordingGpuFrame frame)
{
if (!ReferenceEquals(_openFrame, frame))
return;
_calls.Add(new GpuRecordedFrameEnd(frame.Serial));
OpenFrameCount--;
_openFrame = null;
}
private static uint AlignUp(uint value, uint alignment) =>
alignment <= 1 ? value : (value + alignment - 1) / alignment * alignment;
}
internal sealed class RecordingGpuFrame(RecordingGpuDevice device, long serial, int slotIndex) : IGpuFrame
{
private bool _ended;
public int SlotIndex { get; } = slotIndex;
public long Serial { get; } = serial;
public GpuRingAllocation AllocateRing(int byteCount, GpuRingUsage usage) => device.Allocate(byteCount, usage);
public IGpuPassEncoder BeginPass(GpuPassDescription description)
{
ArgumentNullException.ThrowIfNull(description);
device.Record(new GpuRecordedPassBegin(description.Name, description.SampleCount));
return new RecordingGpuPassEncoder(device, description);
}
public void End()
{
if (_ended)
return;
_ended = true;
device.CloseFrame(this);
}
public void Dispose() => End();
}
internal sealed class RecordingGpuPassEncoder(RecordingGpuDevice device, GpuPassDescription pass) : IGpuPassEncoder
{
private bool _closed;
public GpuPassDescription Pass { get; } = pass;
public void BindPipeline(IGpuPipeline pipeline)
{
ArgumentNullException.ThrowIfNull(pipeline);
device.Record(new GpuRecordedPipelineBind(pipeline.Description.Name));
}
public void BindStorageBuffer(uint binding, IGpuBuffer buffer, uint offsetBytes, uint sizeBytes)
{
ArgumentNullException.ThrowIfNull(buffer);
device.Record(new GpuRecordedStorageBind(binding, buffer.Name, offsetBytes, sizeBytes));
}
public void BindUniformBuffer(uint binding, IGpuBuffer buffer, uint offsetBytes, uint sizeBytes)
{
ArgumentNullException.ThrowIfNull(buffer);
device.Record(new GpuRecordedUniformBind(binding, buffer.Name, offsetBytes, sizeBytes));
}
public void BindVertexBuffer(IGpuBuffer buffer, uint offsetBytes)
{
ArgumentNullException.ThrowIfNull(buffer);
device.Record(new GpuRecordedVertexBind(buffer.Name, offsetBytes));
}
public void BindIndexBuffer(IGpuBuffer buffer, uint offsetBytes, GpuIndexType indexType)
{
ArgumentNullException.ThrowIfNull(buffer);
device.Record(new GpuRecordedIndexBind(buffer.Name, offsetBytes, indexType));
}
public void SetPushConstants(in GpuPushConstants constants) =>
device.Record(new GpuRecordedPushConstants(constants));
public void SetViewport(int x, int y, int width, int height) =>
device.Record(new GpuRecordedViewport(x, y, width, height));
public void SetScissor(int x, int y, int width, int height) =>
device.Record(new GpuRecordedScissor(x, y, width, height));
public void SetCullMode(GpuCullMode cullMode) => device.Record(new GpuRecordedCullMode(cullMode));
public void SetFrontFace(GpuFrontFace frontFace) => device.Record(new GpuRecordedFrontFace(frontFace));
public void SetDepthWrite(bool enabled) => device.Record(new GpuRecordedDepthWrite(enabled));
public void DrawIndexed(uint indexCount, uint instanceCount, uint firstIndex, int vertexOffset, uint firstInstance) =>
device.Record(new GpuRecordedDrawIndexed(indexCount, instanceCount, firstIndex, vertexOffset, firstInstance));
public void Draw(uint vertexCount, uint instanceCount, uint firstVertex, uint firstInstance) =>
device.Record(new GpuRecordedDraw(vertexCount, instanceCount, firstVertex, firstInstance));
public void MultiDrawIndexedIndirect(IGpuBuffer commands, uint offsetBytes, uint drawCount, uint strideBytes)
{
ArgumentNullException.ThrowIfNull(commands);
device.Record(new GpuRecordedMultiDrawIndirect(commands.Name, offsetBytes, drawCount, strideBytes));
}
public IDisposable BeginTimerScope(string scopeName) => NullDisposable.Instance;
public void Dispose()
{
if (_closed)
return;
_closed = true;
device.Record(new GpuRecordedPassEnd(Pass.Name));
}
private sealed class NullDisposable : IDisposable
{
public static NullDisposable Instance { get; } = new();
public void Dispose()
{
}
}
}
internal sealed class RecordingGpuBuffer(GpuBufferDescription description) : IGpuBuffer
{
private readonly byte[] _storage = new byte[description.SizeBytes];
public string Name { get; } = description.Name;
public long SizeBytes { get; } = description.SizeBytes;
public GpuBufferUsage Usage { get; } = description.Usage;
public GpuMemoryResidency Residency { get; } = description.Residency;
public bool IsDisposed { get; private set; }
public void Upload(long offsetBytes, ReadOnlySpan<byte> data) =>
data.CopyTo(_storage.AsSpan((int)offsetBytes, data.Length));
public void CopyTo(IGpuBuffer destination, long sourceOffsetBytes, long destinationOffsetBytes, long byteCount)
{
ArgumentNullException.ThrowIfNull(destination);
if (destination is not RecordingGpuBuffer target)
throw new ArgumentException("Recording buffers can only copy to recording buffers.", nameof(destination));
_storage.AsSpan((int)sourceOffsetBytes, (int)byteCount)
.CopyTo(target._storage.AsSpan((int)destinationOffsetBytes, (int)byteCount));
}
public void Read(long offsetBytes, Span<byte> destination) =>
_storage.AsSpan((int)offsetBytes, destination.Length).CopyTo(destination);
public void Dispose() => IsDisposed = true;
}
internal sealed class RecordingGpuTexture(
string name,
GpuTextureKind kind,
GpuTextureFormat format,
int width,
int height,
int layerCount,
int mipLevelCount) : IGpuTexture
{
private readonly List<(int MipLevel, int Layer, int ByteCount)> _uploads = [];
public string Name { get; } = name;
public GpuTextureKind Kind { get; } = kind;
public GpuTextureFormat Format { get; } = format;
public int Width { get; } = width;
public int Height { get; } = height;
public int LayerCount { get; } = layerCount;
public int MipLevelCount { get; } = mipLevelCount;
public bool MipChainGenerated { get; private set; }
public bool IsDisposed { get; private set; }
public IReadOnlyList<(int MipLevel, int Layer, int ByteCount)> Uploads => _uploads;
public void Upload(int mipLevel, int layer, ReadOnlySpan<byte> data) =>
_uploads.Add((mipLevel, layer, data.Length));
public void GenerateMipChain() => MipChainGenerated = true;
public void Dispose() => IsDisposed = true;
}
internal sealed class RecordingGpuSampler(GpuSamplerDescription description) : IGpuSampler
{
public GpuSamplerDescription Description { get; } = description;
public bool IsDisposed { get; private set; }
public void Dispose() => IsDisposed = true;
}
internal sealed class RecordingGpuPipeline(GpuPipelineDescription description) : IGpuPipeline
{
public GpuPipelineDescription Description { get; } = description;
public bool IsDisposed { get; private set; }
public void Dispose() => IsDisposed = true;
}
internal sealed class RecordingGpuRenderTarget : IGpuRenderTarget
{
public RecordingGpuRenderTarget(GpuRenderTargetDescription description)
{
Description = description;
ColorTexture = new RecordingGpuTexture(
$"{description.Name}-color",
GpuTextureKind.Texture2D,
description.ColorFormat,
description.Width,
description.Height,
layerCount: 1,
mipLevelCount: 1);
}
public GpuRenderTargetDescription Description { get; }
public IGpuTexture ColorTexture { get; }
public bool IsDisposed { get; private set; }
public void Dispose() => IsDisposed = true;
}
internal sealed class RecordingGpuTimerPool : IGpuTimerPool
{
public bool IsSupported => false;
public bool TryResolve(string scopeName, out double milliseconds)
{
milliseconds = 0d;
return false;
}
}
/// <summary>Convenience helpers so renderer tests read as assertions, not as list surgery.</summary>
internal static class RecordingGpuDeviceAssertions
{
public static IEnumerable<T> OfKind<T>(this RecordingGpuDevice device) where T : GpuRecordedCall =>
device.Calls.OfType<T>();
public static Vector4 ClearColorOf(this GpuPassDescription pass) => pass.Color.ClearColor;
}