GSplat Depth Sorting — Phased Implementation Plan (Option 3a)

Status: Spike + Phase 0 MERGED TO MAIN (PR #511, squash e1e079e6); shader-symmetry + perf-validation follow-ups MERGED (PR #523, squash fd243382); Phase 1 IMPLEMENTED (branch gsplat-depth-sorting-phase1): texture-backed splat storage + always-on aSortedIndex, per-node gsplat materials, syncGSplatMaterialWithGeometry commit rebind, texture-lifetime-pinned-to-geometry via dispose listener. Implementation deltas vs the Rev 4 text: (a) maxTextureSize needed net-new plumbing (RendererCapabilities.maxTextureSize → module-scoped element-texture-layout.ts (née splat-texture-layout), the gpu-byte-budget live-authority pattern — the pool ctor is too far from the renderer); (b) texture bytes are counted by stashing the texture on geometry.userData (visible to estimateGeometryBytes) rather than a PooledBuffer field + per-site dispose edits; (c) pick vElementId now derives from aSortedIndex (storage slot) instead of gl_InstanceID — identical under identity ordering, correct-by-construction once Phase 2 permutes; (d) TSL textureSize() is typed uint (WGSL convention) while GLSL returns int — the width read is wrapped in int(...), without which the generated GLSL fails to compile on the forceWebGL backend (risk-register class #1, found by the parity suite’s non-vacuousness guard). Phase 1 MERGED as PR #535 (+ follow-up coverage PRs #537/#538). Phase 2 MERGED (PR #540, squash 2471327f): sort_splats_by_depth WASM kernel + frounded TS twin (exact-permutation parity), persistent Comlink SortWorker (workers/sort-worker.ts), main-thread rendering/depth-sort-coordinator.ts (generation guard, one-in-flight-per-node, committedData demotion guard), commit wiring + LayersPanel mode-switch hook, non-vacuous E2E gate on the new front-to-back test_gsplats_normal_overlap_reversed fixture (fail-first verified). Implementation deltas vs the Rev 4 text: (a) the throughput floor moved from the estimated 100M splats/s to a measured 50M/s, report-only unless LUXAR_PERF_QUIET_HOST=1 (details in §5’s exit-criteria note); (b) the camera/render/reprocess callbacks are injected via configureDepthSort at app init (SceneLoader deliberately owns no camera state, so commit-time model-view comes from the coordinator, not the loader); (c) an ordering resolve additionally requires the mesh’s committedData stamp to survive — LOD demotion returns the geometry to the evictable pool, and the cleared stamp is exactly that signal; (d) the sort-time model-view is derived from FRESHLY updated matrices (mesh.updateWorldMatrix + a local inverse(camera.matrixWorld)) — the renderer-maintained camera.matrixWorldInverse / mesh.matrixWorld caches are stale when a commit fires before the next frame (first commit of a load, idle-paused loop); (e) a queued re-sort is drained even when the in-flight sort RPC fails, and the TS twin’s key clamp uses a < comparison to reproduce Rust’s f32::min(NaN, 65535) == 65535 semantics exactly (post-merge double-check hardening). Phase 3 IMPLEMENTED (PR #553, branch depth-sort-phase3): the 'depth-sort-scheduler' per-frame callback + evaluateDepthSortPerFrame in the coordinator, config/sections/depth-sort/ (enabled/angleThresholdDeg/translationFraction), ?depthSort=0 escape hatch, and the ‘Depth Sort’ monitor line. Implementation deltas vs the §6 text: (a) the tracked pose is the model-view matrix’s z-row (unit axis direction + normalized offset m14/|axis|) rather than a viewDir+position pair — the kernel sorts by view-space z = axis·p + offset, so that row fully determines the permutation; translation ORTHOGONAL to the view axis provably cannot change the ordering and never dispatches (the translation trigger compares offsets along the axis against translationFraction × boundingSphere.radius, catching the behind-camera-set changes that view-axis motion causes); (b) the monitor integration is a third persistent update-profiler root (‘Depth Sort’, beginDepthSortPass(), mirroring the ‘LOD Refinement’ pass isolation) rather than an update-session stage — camera-triggered sorts run outside any update session; each pass records the dispatch→applied round-trip latency plus splat count and ordering-upload bytes; (c) ?depthSort=0 disables the ENTIRE subsystem (worker never spawns, commits keep identity ordering, the mode-switch reprocess hook goes inert) — not just the scheduler — because a determinism escape hatch that still sorted at commit time would not pin output; (d) hysteresis is dispatch-updates-reference (each dispatch re-records the pose), in-flight nodes are skipped rather than queued (the queued-re-sort slot stays reserved for commits, preserving Phase 2’s bounded-failure-retry invariant), and dispatch is additionally gated on effective mesh visibility (own flag AND every ancestor — an LOD level can be a hidden partition GROUP whose member meshes stay visible=true), a surviving committedData stamp, and the live blending mode staying order-dependent; (e) configureDepthSort takes a live getCamera GETTER, not a camera reference — the ortho-mode toggle REPLACES the scene manager’s camera object, and a captured init-time reference would freeze sorting at the abandoned perspective pose (latent in Phase 2’s commit sorts, consequential once the per-frame scheduler exists; the lod-group-registry getter precedent). Post-Phase-3 hardening campaign MERGED (2026-07, PRs #568/#569/#571/#572/#575/#577 + review follow-up): null-camera first-dispatch recovery + empty-commit pose hygiene; capacity-clamp consistency (worker count = written texel count — the OOB-ordering fix); worker-center release on LOD demotion; ONE GLOBAL cross-node renderOrder domain (collect-then-assign: wrapper groups by mean member view-z, exact BSP ranks within a wrapper — replaces the per-wrapper 0..N−1 ranks that were incomparable across wrappers/leaves); same-count commits preserve the previous permutation (no 1-frame identity flash while scrubbing); normal-mode picking writes real projected depth (front-most wins) instead of brightness-as-depth. Follow-ups (2026-07): the cross-node renderOrder machinery (BSP part lookup/FKN traversal/rank memo/global assignment) was extracted to rendering/depth-sort-coordinator/render-order.ts behind a three-function per-frame API (PR #591); the renderOrder pass is deliberately NOT gated on SortWorker availability — cross-node mesh order survives a never-constructed worker (CSP-blocked script), losing only the worker-dependent within-mesh order (PR #596). Worker-startup resilience IMPLEMENTED (branch sortworker-init-resilience): the worker is now spawned at APP INIT (warmUpDepthSortWorker, gated on depthSortEnabled) instead of by the first order-dependent commit, and its init deadline became configurable. Rationale: the first commit lands while the main thread and data-worker pool are saturated decoding, and the worker’s Comlink reply must be dispatched on that same thread — measured at 3M points, the worker logged its own readiness (WASM up, answered) while the main thread hit the 30 s deadline, terminated it, and cached the rejection for the session, leaving every order-dependent node (all four geometry types, normal ∪ volumetric) in storage order until reload. Deltas — (a)-(d) are issue #1694’s fix, which this branch builds on rather than contributes (the transient-vs-permanent classification, the bounded starved retry with its backoff and self-wake, the late re-registration sweep, and the init-epoch guard); (e) is this branch’s own (app-init warm-up above, the shared init guard with its messageerror arm, the configurable deadline, and the monitor’s UNAVAILABLE note): (a) the failure taxonomy splits — constructor throw / onerror / onmessageerror / an initialize() REJECTION stay permanent and latch, only the deadline is transient; (b) a transient miss still terminates the worker and keeps the rejected initPromise cached — so a commit can never spend an attempt — and only maybeRetryStarvedWorkerInit (driven by evaluateDepthSortPerFrame past its loader-idle gate, on a growing backoff with a self-wake for the idle-paused render loop) drops that rejection and spawns a fresh worker, bounded at 3 attempts in total, after which it latches; (c) a successful late init re-registers every tracked order-dependent node via reregisterAfterLateWorkerInit (the fresh worker holds none of their centers); (d) every module write past an await in the init closure is epoch-guarded, because the deadline timer lives in that closure and a dispose cannot cancel it — an orphaned one would otherwise classify a miss against the next app’s healthy worker; and the resulting verdict is observable via getDepthSortWorkerStatus() (idle/ready/starved/failed + the deadline-miss count); (e) the hand-rolled init race was replaced by the data pool’s shared worker-pool/lifecycle/init-with-guard.ts (gaining the onmessageerror arm), the deadline moved to config.depthSort.workerInitTimeoutMs, and giving up is now REPORTED — isDepthSortAvailable() drives a depth sort UNAVAILABLE note in the monitor footer, per issue #705’s never-silently-unsorted rule. Risk register §10, changelogs §11. Scope: Viewer-only. No .gsplats.zarr / .luxar.zarr format change, no Python-side change. Goal: Correct order-dependent transparency (normal blending mode) for Gaussian splats at 10M+ splats, via texture-backed splat storage + a per-instance ordering attribute + asynchronous worker depth sorting. Non-goals (follow-ups, not this plan): WebGPU compute-shader sorting (“Option 3b”), weighted-blended OIT, global cross-node sorting. (Points AND Lines sorting symmetry have since LANDED — see §8; all three instanced geometry types share the storage + sorting machinery. Mesh has since joined the SORTING half only: same registration, worker and kernel, but its apply permutes geometry.index rather than aSortedIndex, so every statement about instanced-quad STORAGE in this document remains three-geometry — see MESH_NODE_SPEC.md §6.3.)

Related reading: src/rendering/README.md (the retired interleaved-attributes.ts module — whose Float16 narrowing contract this plan superseded — was deleted once the lines migration left it with no consumers).


1. Problem statement (facts on the ground)

  • All three geometry types draw as instanced quads: THREE.Mesh + InstancedBufferGeometry, 4-vertex/6-index base quad, per-instance data in one InstancedInterleavedBuffer (rendering/gsplat-geometry.ts, stride 13 floats = 52 B/splat: aCenter(3) aCholesky01/23/45(2+2+2) aAmplitude(1) aColor(3)).

  • WebGL2/WebGPU instancing draws instances strictly in instance-buffer order. There is no per-instance index buffer. Reordering today means rewriting 52 B × N.

  • No sorting exists anywhere in the render path. renderer.sortObjects (default) only sorts whole Object3Ds. Within a mesh, splat order = ladder/commit order.

  • normal mode (rendering/blending-state.ts ~:189-201) is SrcAlpha/OneMinusSrcAlpha alpha-over — non-commutative, therefore wrong in arbitrary order. additive/luminous/max are commutative and unaffected. 2026-07 update: the sorted domain is now the needsDepthSort predicate = normal ∪ volumetric — the emission–absorption mode (VOLUMETRIC_BLENDING_SPEC.md) is the second non-commutative mode and rides this exact infrastructure (same SortWorker, same renderOrder pass, same mode-switch transitions; a normal↔volumetric switch is a sorted→sorted no-op).

  • GSplat normal-mode alpha is FIXED (Phase 0, merged): under LUXAR_NORMAL_PREMULT the fragment emits real premultiplied coverage alpha paired with getGSplatNormalBlendingState() (materials/gsplat/shader-glsl.ts ~:408-425, blending-state.ts ~:238); all other modes keep the alpha = 1.0 contract. What remains wrong — and what Phases 1-3 fix — is that alpha-over is still applied in storage order, not depth order.

  • Every commit already does a full CPU repack + full GPU re-upload of the interleaved buffer (gpu-buffer-pool/gsplats-adapter.ts::updateGeometry → writePooledAttribute per attribute), plus O(N) bounding-box and max-Cholesky-row-norm loops. Progressive ladder refinement re-runs this for the entire prefix on every level (data/gsplats/gsplats-progressive-loader.ts memoized concat → commit-gsplats-geometry.ts).

  • The projection worker (workers/data-worker/projection/gsplats.ts) is a stateless Comlink RPC over a round-robin pool; inputs are deliberately structured-cloned (transfer would detach SliceCache arrays — data-processor-gsplats.ts warning comment), outputs are transferred back. Plain-3D data never touches this worker: the gate is useWebWorkers && splatCount > 1000 && ndim > 3 (data-processor-gsplats.ts ~:221), and the ndim==3 fast path is an in-process .slice() copy (projection/gsplats.ts:148-166). Any sorting design must therefore NOT live inside the projection call — it would run on the main thread for exactly the datasets that need sorting most.

  • Vertex texture fetch is already in production: the colormap LUT is sampled in the gsplat vertex shader (shader-glsl.ts ~:311-316, #ifdef USE_COLORMAP).

  • Pick nodes share the visual mesh’s geometry and pick materials are already per-node (node-factory/create-gsplats-node.ts:68-77). Pick shaders read the same five instanced attributes (picking/gsplat/shaders.ts ~:39-43, pick.tsl.ts ~:89-94). Post-commit material sync has an established home: rendering/material-sync-helpers.ts (points radiusScale precedent, pushes to render + pick materials via userData.pickNode).

  • THREE r184 blend/attribute facts verified against node_modules/three:

    • Integer instanced attributes work out of the box: a Uint32Array attribute binds via vertexAttribIPointer because its GL type is UNSIGNED_INT (three.module.js:1950), matching a GLSL in uint; the WebGPU path derives uint32 vertex format from the array constructor.

    • material.premultipliedAlpha = true is a trap on the TSL path: NodeMaterial.setup() injects an automatic output-RGB×alpha transform when the flag is set (three.webgpu.js:21817-21822). A shader that already premultiplies would be premultiplied twice on TSL but not GLSL — a guaranteed parity break. Do not use the flag.

    • The known WebGPU→WebGL2-bridge gl.getError() issue is specifically separate alpha-channel blend equation/factor state (documented in materials/gsplat/material-tsl.ts:190-212, :300-320); CustomBlending itself is fine — max mode ships CustomBlending + MaxEquation through the TSL path today.

Overlapping-node authoring rule

Depth sorting is exact only within one node. Cross-node renderOrder uses BSP partition order where available and a mean-view-z approximation otherwise, so two spatially overlapping normal/volumetric nodes can exchange relative order as the camera moves. For nodes with the same geometry type, blend mode, and effective opacity, merge them into one node and pass partition={"max_elements": N}: disjoint BSP cells can then be ordered exactly back-to-front while preserving placement and appearance. This authoring path is available to Points, Lines, Mesh, and Gaussian Splats. Put the displayed dimensions first to keep exact BSP ordering. If the geometry types, modes, or opacities differ, merging cannot preserve the authored material; use additive only for an emissive medium, or separate the bounds.

That escape hatch must not be used against depth-writing geometry. additive is the only mode that ignores the depth buffer, so it paints through opaque geometry and through a full-opacity depth-writing normal Lines or Mesh layer. Use luminous for the same additive light-summing with depth testing. Finalization reports both hazards when co-visible world-space node bounds overlap. To avoid noisy warnings from coarse AABBs, the two-sorted-node warning requires one box to contain the other; the additive/depth-writing warning still fires on any positive spatial overlap. Explicitly authored or inherited additive is treated as intentional X-ray rendering and suppresses the latter warning. Neither diagnostic rewrites authored modes or defaults.

Bandwidth argument for the architecture: at 10M splats, re-sorting by buffer rewrite costs 520 MB/sort. With splat data in a texture and only a Uint32 ordering attribute per instance, a re-sort uploads 40 MB — 13× less — through the attribute addUpdateRange machinery that already exists.


2. Target architecture

                       ┌────────────────────────────────────────────┐
                       │ Splat texture (per pooled geometry)        │
   commit (slice/LOD)  │ RGBA32F, width 4096, 4 texels/splat        │
  ────────────────────▶│ [cx cy cz amp][L00 L10 L11 L20]            │
   full/partial upload │ [L21 L22 cr cg][cb — — —]  (— = reserved)  │
                       └────────────────────────────────────────────┘
                                          ▲ texelFetch ×4
┌──────────────┐   ordering Uint32Array   │
│ SortWorker   │  ────────────────────▶ aSortedIndex (per-instance, │
│ (dedicated,  │   (transferable)        Uint32, DynamicDrawUsage)  │
│  stateful)   │                          │
│ centers cache│                          ▼
│ u16 counting │                    vertex shader:
│ sort (WASM + │                    idx = int(aSortedIndex)
│  TS twin)    │                    splat = texelFetch(uSplatTex, addr(idx))
└──────────────┘                    (rest of the pipeline unchanged)
      ▲
      │ view matrix (64 B) per sort; centers transferred once per commit
      │
sort triggers: commit (Phase 2) + per-frame camera-delta scheduler (Phase 3)

Two orders coexist by design:

  • Storage order (texture rows) stays the canonical commit order — the additive-ladder concat order. Never permuted.

  • Draw order (aSortedIndex) is the view-dependent permutation. Identity for order-independent modes; back-to-front for normal.

Single owner of ordering: the SortWorker computes every non-identity ordering. The projection pipeline (worker-pool or in-process) is never involved — see the ndim gate fact in §1. This keeps one code path for 3D and nD data, avoids main-thread sorts, and avoids threading sort state through the stateless round-robin pool.

This is the SparkJS decomposition adapted to Luxar’s per-node (not global-accumulator) scene graph, nD projection pipeline, and dual GLSL/TSL backend.

Pinned design decisions

Decision

Choice

Rationale

Splat storage

One THREE.DataTexture RGBA32F per pooled geometry, NearestFilter, width 4096, 4 texels/splat

texelFetch is core WebGL2; float16 vertex-format pitfalls (see revert dd7c4478 — the since-deleted interleaved-attributes.ts header documented the contract) don’t exist for textures; 4-texel padding (64 B vs 52 B) buys addressing simplicity + room for a future per-splat scalar; a later RGBA16F narrowing (32 B/splat) is a data-conversion-only change

Addressing

1024 splats/row; `addr(i, t) = ivec2(((i & LUXAR_SPLAT_ROW_MASK) << 2)

t, i >> LUXAR_SPLAT_ROW_SHIFT)` with the constants emitted as compile-time defines / TSL factory params from the actual texture width

Max splats/node

maxTextureSize × splatsPerRow (4096-wide: 16.7M at height 16384, 8.4M at 8192)

Height, not width, is the scaling axis. Nodes beyond the cap should use tiles/adaptive recipes (250K/part idiom); warn + clamp; DataArrayTexture is the documented escape hatch if low-maxTextureSize devices matter in the field

Ordering

aSortedIndex (Uint32Array InstancedBufferAttribute, DynamicDrawUsage), always present, identity by default

One shader path, no define flips/recompiles on mode switch; keeps THREE’s _maxInstanceCount derivation working (an instanced attribute must exist); 4 B/splat is negligible. Uint32→in uint binding is native in r184 (§1); the parity harness must cover it on both backends

Attribute path

Deleted, not kept alongside

No dual shader maintenance (“no backwards compatibility burden”, “complete before perfect”); vertex buffers drop to 2 (quad + ordering) — well under WebGPU compat-mode’s maxVertexBuffers=8

normal-mode blend state (gsplats)

CustomBlending + AddEquation + One / OneMinusSrcAlpha, symmetric alpha channel (no blendEquationAlpha/blendSrcAlpha/blendDstAlpha overrides), premultipliedAlpha flag never set, transparent: true, depthTest: true, depthWrite: false

One/OneMinusSrcAlpha because the shader premultiplies. Symmetric alpha state avoids the documented WebGPU-bridge gl.getError() issue; the premultipliedAlpha flag would double-premultiply on TSL (§1). depthWrite: false unconditionally — a coverage-alpha fragment (alpha as low as 1e-4 passes the shader’s discards) must never write depth or splat footprints punch occlusion halos; sorted transparency never depth-writes

Material sharing

GSplat visual materials become per-node (skip the LRU in getGSplatMaterial, still register() for camera broadcasts); create with _layerMaterialCloned: true

uSplatTex is a per-node uniform; the clone-on-divergence dance (colormap path, LayersPanel first-touch clone) collapses into “always own your material”. Pick materials are already per-node

Texture ownership

The pool entry owns the texture: PooledBuffer gains splatTexture, created WITH the geometry at allocation and disposed WITH it (geometry.dispose() + splatTexture.dispose() in the evictors/dispose()); commit rebinds it via a new syncGSplatMaterialWithGeometry in material-sync-helpers.ts (render + pick material, same shape as the points radiusScale sync)

Texture lifetime = geometry lifetime under the post-#511 contract: growth is release + reacquire, never an in-place reallocation — an in-place texture swap on a rendered geometry/material is the same strand class the buffer fix killed (Info.memoryMap pins replaced resources; see the grow-leak probe). A grown node simply gets a fresh geometry+texture pair; no carry-forward is needed because every commit rewrites the full count. Unlike InterleavedBuffer, THREE.Texture IS an EventDispatcher with backend dispose listeners, so pool-side splatTexture.dispose() frees GPU memory deterministically

Sort keys

Camera-space depth, min/max-normalized to uint16 (65536-bucket counting sort, back-to-front), WASM kernel + mandatory TS twin

O(N) and memory-bound: expect ~100-300M keys/s in a worker → ~1-10 ms at 1M, ~30-100 ms at 10M. NOT SparkJS’s raw-f16-bit-pattern keys: Luxar scenes span physical units from nm to km, so raw camera-space depths can overflow f16 (>65504 → every splat keys to Inf) or underflow to 0 — either collapses the ordering entirely. Per-sort min/max normalization is scale-invariant (antimatter15-splat-style), same O(N), and 65536 bins across a node’s depth extent is ample for blending order. Guard zmax==zmin → identity

Sort execution

One dedicated persistent SortWorker (module-scoped state, own Comlink endpoint), NOT the round-robin data-worker pool, and NOT inside projection

Sorting needs per-node retained centers; round-robin would scatter state, and projection doesn’t even run in a worker for ndim==3 (§1). SharedArrayBuffer rejected: COOP/COEP would ripple into luxar serve, the export bundle’s serve.py, and the Go launchers

Sort gating

Only nodes whose effective composed blending mode is order-dependent (normal) get registration/sorts; everything else keeps identity ordering

Zero cost for the default (additive) path. (opaque at fractional opacity is also order-dependent but degenerate — explicitly out of scope, documented in the mode docs)

2.1 Sorting execution model

CPU-in-worker, not GPU. On the production WebGL2 backend a GPU sort means fragment-shader bitonic ping-pong: O(n log²n) fullscreen passes (~300 passes at 10M) with no way to write a vertex-consumable buffer from a shader — strictly worse than an O(n) CPU counting sort in a worker. SparkJS’s own hybrid (GPU distance compute → RGBA8 readback → CPU radix) exists because its splats are GPU-generated (procedural/skinned) and live only in textures; Luxar’s projection pipeline hands us CPU-side centers for free, so computing keys in the worker skips the readback path entirely. The GPU answer is Option 3b (WebGPU compute radix), which replaces the worker, not the interface. Revisit only if projection itself ever moves GPU-side.

Full snapshot sorts, not incremental/progressive maintenance. Counting sort is already O(n) with memory-bound constants — frame-coherence tricks (nearly-sorted adaptive sorts, “resort only what moved”) cannot beat O(n) and add worst-case cliffs; this is why every serious 3DGS implementation uses radix/counting. The “progressive” axis in this design is staleness, not partial computation: render with the last complete order while the next one computes (bounded by the ~3° trigger, so the stale order is never more than ~one threshold wrong). Time-distribution happens across nodes, not within a sort: the per-frame scheduler dispatches every past-threshold node in the same tick and the single SortWorker serializes them (a one-per-frame coverage×staleness priority queue was designed but never implemented — measured 2026-07: MatrixCity frame tails are load-driven, not dispatch-burst-driven, so the budget was dropped at the evidence gate); tiles/adaptive partitions cap per-sort N at ~250K (~1-3 ms each), making per-node sorts naturally bite-sized. Coarse-first bucket refinement (approximate order now, exact later) is rejected: complexity for artifacts that async full sorts already avoid.

A permutation must swap atomically — that’s where the double buffering lives. A partially-applied ordering is not a reordering, it is a corrupt permutation (splats drawn twice/zero times → flicker), so partial visual application is never acceptable. Three tiers:

  1. Main-thread apply needs no explicit double buffer: JS single-threadedness means aSortedIndex.set(ordering) + needsUpdate completes between frames; the draw only ever sees a complete order. Driver-side buffer renaming under DynamicDrawUsage is the implicit GPU double buffer.

  2. Worker boundary: ping-pong a recycled pair of transferables — main thread returns the previous ordering buffer with the next sort request, worker fills and transfers back. Zero steady-state allocation (SparkJS’s persistent-buffer pattern across postMessage).

  3. Escape hatch for 10M-scale upload stalls (40 MB ≈ 5-8 ms on the GL thread): an A/B pair of ordering attributes with a uniform selector — stream the new order into the inactive attribute across frames via partial addUpdateRange, then flip the uniform (atomic swap, amortized upload, +4 B/splat). Ship without it; add only if the Phase-3 perf gate shows frame drops attributable to ordering uploads. RESOLVED (2026-07, perf lever L8): the perf campaign measured the stall far above the estimate — sort-adjacent frame p99 of 119–563 ms at 10M splats (Mac M4 Max, synthetic orbit; the full-buffer memcpy dominates, not just the upload) — and it shipped as chunked apply instead of the A/B pair (no shader change, no +4 B/splat): orderings > 1M indices stream one 1M-index (4 MB) slice per rendered frame through per-slice addUpdateRange, driven by the Phase-3 per-frame scheduler; while a stream runs the node is dispatch-gated (no new sorts) and a newer ordering is held (latest wins) until the stream completes — restarting mid-stream under a continuous orbit never converged (element-storage.ts chunked-apply note has the full mechanics + the documented transient old/new-mix trade the A/B design would have avoided — judged acceptable: bounded ≤ ceil(n/1M) frames of duplicate/omit shimmer, same visual class as stale-order frames). Classic WebGL only; the WebGPU backends ignore attribute ranges and keep the single-shot path. SUPERSEDED (2026-07-30) — tier 3 SHIPPED. The chunked apply’s “bounded transient mix” premise did not survive measurement: under a CONTINUOUS orbit a new sort arrives about as fast as a stream drains, so the mix is the steady state, not a transient. Measured on main with a browser probe that decodes the element texture and validates the drawn index buffer every sampled frame: 27–33% of frames on the 1.9M-splat volumetric visible_human_head (27,613 elements double-drawn) and 70–80% of frames on the 8M-point global_rivers_earth terrain (up to 1,022,162 double-drawn = 12.8% of the node), duplicates always beginning exactly at the chunk boundary; an A/B that raised the threshold above the node size gave a valid permutation in every frame. So the A/B pair shipped as specced: aSortedIndex/aSortedIndexB + a runtime uSortedIndexSlot selector, slices streaming into the INACTIVE buffer and the slot flipping on completion. BOTH buffers are allocated at attach, so the +4 B/element is paid by every node, sorting or not. The pair was FIRST aliased onto one buffer and split on a node’s first sort (free for commutative-mode nodes) — that was a native-WebGPU black-screen bug, caught only by an explicit ?renderer=webgpu A/B against main and worth recording as a hazard class: three keys a render pipeline’s vertex-buffer LAYOUT by BufferAttribute identity (WebGPUAttributeUtils.createShaderVertexBuffers), but RenderObject.getGeometryCacheKey hashes only attribute NAMES/itemSize/normalized, and RenderObjects.get answers needsGeometryUpdate with a bare setGeometry() that refreshes the attribute list while leaving Pipelines’ cached pipeline alone. So splitting the alias after first render bound three vertex buffers into a two-buffer layout: every later attribute shifted down a slot and aQuadCorner read the ordering buffer’s u32s as vec2<f32>, collapsing every quad — a black scene, no validation error, no console warning. Classic WebGL binds attributes by program location and is structurally immune; so is tsl-shader-parity, which runs WebGPURenderer({ forceWebGL: true }) — a TSL graph through the WebGL2 bridge, which never builds a WebGPU vertex layout at all. Rule: for any geometry a WebGPU pipeline has drawn, the attribute SET and each attribute’s object identity are immutable. Guarded by a unit invariant test (splat-texture-storage.test.ts, “keeps the ordering attribute OBJECTS fixed for the geometry lifetime”), since native-WebGPU E2E is not runnable in CI. Slicing stays classic-WebGL-gated (WebGPU writes one slice then flips); double-buffering is unconditional. The dispatch apply-gate is GONE — with the drawn buffer always whole there is nothing to protect. Cost, measured on the 10M orbit bench (M4 Max, 3 runs/arm): sort-adjacent frame p99 ~77 ms → ~92 ms with 12 → 24 sorts; median and p95 unchanged; still far under the 119–563 ms the chunking exists to prevent. A throttle on “an ordering is already queued” was built and rejected at the evidence gate — it cut sorts to ~15 but moved p99 only to ~89 ms (inside noise) while costing 41% more staleness (fitted sort-axis lag on the 8M fast orbit: mean 44° → 63°). Waiting for a WHOLE ordering instead of showing a partly-applied one does cost freshness: same 8M fast orbit, lag mean 36.5° before vs 44.3° after — the deliberate price of never drawing a corrupt permutation.

Stability as a feature: the counting sort is stable, so equal-key splats (same 1/65536 depth bin) retain storage order — which is the additive-ladder energy order. Ties resolve toward “most important splat wins the blend order”, for free.


3. Phase 0 — Correct alpha for normal mode (prerequisite)

Without this, sorted or not, gsplat normal mode cannot reveal the framebuffer behind a splat. Independent, small, immediately shippable.

Changes

  1. Fragment shaders (materials/gsplat/shader-glsl.ts + shader-tsl.ts): under a new LUXAR_NORMAL_PREMULT define (GLSL) / factory flag (TSL), emit premultiplied coverage alpha:

    float a = clamp(intensity * uOpacity, 0.0, 1.0);   // coverage, clamped for OneMinusSrcAlpha sanity
    fragColor = vec4(gammaColor * intensity * uOpacity, a);  // RGB stays unclamped HDR
    

    Default (no define) output is byte-identical to today — additive/luminous/max keep the exact vec4(rgb, 1.0) contract. (As shipped, additive blending is unified onto the shared blending-state helpers on both backends — the GLSL alpha-MaxEquation overflow guard this section originally cited was deleted during the same campaign, with the raw-scene-hdr capture path sanitizing alpha at readback instead; see §10.) Note the deliberate HDR asymmetry: RGB carries unclamped intensity × uOpacity, alpha carries the clamped coverage — an emitter-with-occlusion model, correct for 1-α destination attenuation.

  2. Blend state: add getGSplatNormalBlendingState() to rendering/blending-state.ts returning the pinned state from §2 (CustomBlending, One/OneMinusSrcAlpha, symmetric alpha, transparent true, depthWrite false). No opacity parameter — unlike the generic normal entry, depthWrite is unconditionally false here, so the state is opacity-independent. Do NOT widen getCompleteBlendingState — the gsplat GLSL wrapper is already the documented divergence point, and the generic normal entry (with its opacity >= 0.99 depthWrite flip) stays correct for lines’ straight-alpha output (points now force depthWrite off as well — see getPointBlendingState, #1002). Consume it from applyBlendingMode in both material-glsl.ts (which hand-manages custom factors today) and material-tsl.ts — whose comment claiming “TSL needs no gsplat-specific override” becomes stale and must be updated, since with real alpha output NormalBlending’s SrcAlpha factor would now double-multiply. LayersPanel’s applyBlendingStateToMaterial fallback is unaffected (it already prefers the material’s own applyBlendingMode).

  3. Depth trade-off, documented: with depthWrite: false, a gsplat normal layer no longer occludes additive layers behind it (the generic normal mode’s opacity >= 0.99 rationale). Accepted: coverage-alpha splats writing depth is strictly worse (halo artifacts from near-transparent fragments that pass the 1e-4 discards).

  4. Docs: update the BlendingMode docs in material-manager/factories.ts:35-45 and materials/gsplat/README.md (“alpha-always-1 contract” becomes “alpha = 1.0 except normal”).

Tests / exit criteria

  • TSL/GLSL parity spec covers the new define/flag combination (guards the double-premult trap by construction).

  • Unit: getGSplatNormalBlendingState state shape; mode-switch round-trip normal↔additive restores the exact prior state.

  • Visual E2E fixture: two overlapping splats, normal mode, opacity 0.5 — background must show through (fails today). Add to tests/fixtures/generate_test_data.py (single source of truth; unit-test globalSetup auto-generates).

  • Existing additive/max visual baselines unchanged (regression gate). Run the WebGPU diagnostic path (?renderer=webgpu and &webgpuForceWebgl) on the fixture — the bridge’s alpha-state sensitivity is exactly what this phase touches.

Risk: low. Purely additive define + one new blend-state helper. normal blending will still be order-dependent after this phase — visibly better, occasionally wrong on overlap, fully fixed by Phases 1–3.


4. Phase 1 — Texture-backed splat storage + always-on ordering attribute

The structural phase. After it, rendering is visually identical (identity ordering), but splat data lives in a texture and draw order is decoupled from storage order.

Changes, by layer

  1. Pool adapter (rendering/gpu-buffer-pool/gsplats-adapter.ts):

    • createGSplatsGeometry(capacity) builds: quad + index (unchanged), one aSortedIndex InstancedBufferAttribute (Uint32Array, DynamicDrawUsage), and a pooled DataTexture (RGBA32F, width × ceil(capacity/splatsPerRow), NearestFilter, generateMipmaps=false, flipY=false), width = min(4096, caps.maxTextureSize) with the addressing defines derived from it.

    • PooledBuffer (in pool-stats.ts) gains splatTexture: THREE.DataTexture. Growth follows the post-#511 contract: release + reacquire — growGSplatsGeometry/rebuildInterleavedBuffer no longer exist, and Phase 1 must NOT reintroduce an in-place texture reallocation (same strand class the buffer fix killed). An undersized acquire releases the geometry+texture pair to the pool intact and falls through to best-fit/fresh allocation; no carry-forward (every commit rewrites the full count). The evictors and pool.dispose() dispose splatTexture alongside the geometry. The existing _lastAcquireRebuilt → invalidateRenderObjectFor(mesh) contract covers the geometry swap — and since the flagged-issues pass, invalidateRenderObjectFor also eagerly re-points userData.pickNode at the new geometry, which Phase 1 inherits for free; the texture swap is covered by the material rebind (below).

    • updateGeometry(...) replaces the six writePooledAttribute calls with one fused pass writing the staged arrays (centers3D, the 6-stride choleskyFactors, amplitudes, colors) into the texture’s backing Float32Array in texel layout, then texture.needsUpdate = true; fills aSortedIndex[0..count) = identity (addUpdateRange); keeps instanceCount, the _maxInstanceCount delete, and the bbox + max-row-norm handling exactly as today (since perf lever L1, the projection’s fused-scan bounds metadata supplies both when present; the loops remain as the fallback). Fewer passes over the data than the current per-attribute writes — this phase should be neutral-to-faster on commit CPU.

    • Non-pool fallback (gsplat-geometry.ts::createInstancedGSplatsMesh / updateInstancedGSplatsMesh) gets the same treatment with a mesh-owned texture. The fallback’s rebuilt branch now swaps in a FRESH geometry and disposes the old one (post-#511); the texture rides the same swap — build fresh, dispose old, never rebind in place. (The 2026-07-12 broken-fallback finding was fixed in PR #503.)

  2. Byte accounting: gpu-buffer-pool/geometry-bytes.ts::estimateGeometryBytes counts capacity × 64 B texture + capacity × 4 B ordering — 68 B/splat vs 52 B today, ≈ +31% VRAM (state this in the shared gpu-byte-budget.ts docs; the future RGBA16F narrowing brings it to 36 B, −31%). Byte-budget eviction and dispose() free the texture with the geometry.

  3. Shaders — all four gsplat stacks in one PR (visual GLSL/TSL, picking GLSL/TSL):

    • Delete the five in/attribute() per-instance declarations; add uniform highp sampler2D uSplatTex; + in uint aSortedIndex; (TSL: attribute('aSortedIndex', 'uint'), textureLoad).

    • Prologue: int idx = int(aSortedIndex); then four texelFetch(uSplatTex, addr(idx, t), 0) reads reconstructing center/cholesky/amplitude/color into the exact local names the existing math uses — zero changes downstream of the prologue (covariance projection, Jacobian, eigen-expansion, the unified perspectiveNearFade, unconditional coverage fade, screenSize un-flip, and fragment kernel all untouched). TSL house rule applies: the vertex is Fn-traced — the prologue is .toVar() STATEMENTS inside the existing Fn(() => {...}), never free expression trees (the uninitialized-eigen bug class).

    • unpackCholesky3D() takes the fetched values as parameters instead of reading attributes; addressing constants come from the width-derived defines / TSL params.

  4. Materials (materials/gsplat/material-glsl.ts + material-tsl.ts): add uSplatTex uniform + updateSplatTexture(tex) (modelled on updateColormapTexture). material-manager.ts::getGSplatMaterial bypasses the LRU (per-node materials), still registers for updateCameraParams broadcast; create-gsplats-node.ts sets _layerMaterialCloned: true and drops the colormap clone dance. Verify per-node material disposal on node teardown — the existing clone-lifecycle path (material-manager/lifecycle.ts subscribeToDispose) should already cover it; confirm with a leak assertion in the pool tests.

  5. Commit path: add syncGSplatMaterialWithGeometry(mesh) to rendering/material-sync-helpers.ts (points-radiusScale precedent): rebind uSplatTex from the acquired pool entry on the visual material and, via userData.pickNode, the pick material. Called from commit-gsplats-geometry.ts after acquire. The noop stamp-only path is untouched.

  6. Cleanup: buildGSplatAttributeSpecs/bindInterleavedAttributes in gsplat-geometry.ts are deleted or absorbed (packCholeskyForShader was later retired outright by perf lever L1 — the 6-stride choleskyFactors3D now flows from projection to the texel writer directly); GSPLATS_ATTRIBUTE_SPECS in the adapter is replaced by the texel-layout writer. Points/Lines interleaved paths are untouched (symmetry is restored when/if they migrate — see §8).

Tests / exit criteria

  • Unit: texel address math incl. non-4096 widths; fused writer round-trip (write staged arrays → read back texel layout); growth = release + reacquire returns a FRESH geometry+texture pair with the old pair pooled intact (NO carry-forward — the forbidden mechanism); byte accounting; material-disposal leak check.

  • Parity: tsl-shader-parity.spec.ts green on the texture-fetch shaders and the uint attribute on both backends.

  • E2E: full gsplat suite (geometry-types, blending-modes, real-dataset-loading, picking specs) green with unchanged visual baselines — Phase 1 must be pixel-identical (f32 attribute reads vs f32 texelFetch are bit-exact).

  • Perf gate: Update Buffers profiler stage not regressed on the COMMITTED timelapse-nav harness (scripts/perf/timelapse-nav-bench.mjs — before/after, both renderers; wall-clock is noise-dominated, use frame percentiles + info.memory); VRAM within the +31% envelope on demo scenes.

  • Texture-lifecycle leak gate: scripts/perf/grow-leak-probe.mjs on native WebGPU (--chrome; see scripts/perf/README.md for the three adapter traps) — after forced grows + LRU evict, info.memory (textures AND attributes) returns to the clean floor. This is the probe that proved the buffer-strand fix; Phase 1 must pass it for textures.

  • Codegen snapshots (LUXAR_UPDATE_SNAPSHOTS=1, FRESH server — the spec false-passes against a stale :5173) regenerate for all gsplat surfaces; the full parity-variant inventory stays green: base/gamma-one/colormap, gsplat-offcenter (fragcoord convention), gsplat-tiny-sigma-fade/-reject (coverage fade), behind-camera trio, gsplat-normal-premult.

  • Context-loss E2E (webgl-errors.spec.ts): textures rebuilt on restore (recommit path).

Risks: texture-unit pressure (uSplatTex + uColormapTex in the vertex stage = 2 of ≥16 guaranteed — fine); maxTextureSize=4096-class devices cap at 4.2M splats/node (warn + clamp; tiles idiom is 250K/part); WebGPU textureLoad parity (covered by the harness + ?renderer=webgpu&webgpuForceWebgl diagnostic).


5. Phase 2 — SortWorker + depth-ordered commits

Introduces the sorting engine and wires it at commit time. After this phase, normal mode is correct within a frame or two of any load/slice-change and only goes stale while orbiting (fixed in Phase 3). Ongoing and camera-driven sorting lives exclusively in the SortWorker, never in the projection call (see §1: plain-3D projection doesn’t run in a worker at all). Implementation delta (Phase 2, 2026-08): the first ordering after an eligible instanced commit may also be computed on the main thread within the shared syncSortMaxElements frame budget. Indexed Mesh is excluded because its ordering is applied through geometry.index; async registration and all subsequent sorting are unchanged.

Changes

  1. WASM kernel + TS twin (the one new compute primitive):

    • wasm/rust/src/depth_sort.rs: sort_splats_by_depth(centers3: &[f32], model_view: &[f32; 16], ordering: &mut [u32], count) -> u32 — pass 1: camera-space z per splat into an f32 scratch + running min/max; pass 2: normalize to uint16 keys (floor((z - zmin) / (zmax - zmin) * 65535), zmax==zmin → identity) + 65536-bucket histogram; prefix-sum reversed for back-to-front; pass 3: stable scatter. Behind-camera splats key to the far bucket (the vertex shader already degenerates them). Raw f16-bit-pattern keys (SparkJS) are deliberately NOT used — Luxar’s physical-unit range (nm..km) overflows/underflows f16 and collapses the ordering; normalized integer keys are scale-invariant. Export from lib.rs; no validate_ndim involvement (input is always projected 3D centers).

    • wasm/typescript/depth-sort.ts: 1:1 reference implementation (repo parity contract), plus the parity test pairing.

  2. SortWorker (workers/sort-worker.ts + workers/sort-worker/): dedicated persistent Comlink worker (module-scoped state, same idiom as the data-worker’s WasmCtx; initialize(wasmPath?) forwards the embedder’s WASM path exactly like data-worker.ts:53). API:

    • registerNode(nodeId, generation, centers3 /* transferred */, count) — called after commit for order-dependent nodes. generation is a per-node monotonic counter incremented on every non-noop commit (stored on the mesh userData next to committedData) — NOT loadedViewVersion: ladder-refinement appends share a view version while count grows, and an ordering must only ever apply to the exact commit it was computed for (a shorter stale permutation applied to a grown buffer is corrupt, not just stale). Noop stamp-only commits leave the generation unchanged, so an in-flight sort stays valid across them. The staged processed.centers3D is transferred, not cloned: the projection output is always a fresh buffer (worker path transfers ownership in; the ndim==3 fast path is a .slice() copy — projection/gsplats.ts:148-150), the commit’s texture-write/bbox loops are the last main-thread readers, and the memoized-concat noop identity keys on sourceData (pre-projection), never on processed.*. Zero-copy registration.

    • sort(nodeId, generation, modelView) → { generation, ordering /* transferred */ }.

    • releaseNode(nodeId) — wired to node disposal/pool release; dispose() terminates the worker on app teardown.

  3. Commit wiring (data-processor-gsplats.ts / commit-gsplats-geometry.ts): when the node’s blending mode is order-dependent — read from the material’s live userData.blendingMode, the value that actually drives the blend state regardless of who set it (creation attrs or a later applyBlendingMode from the LayersPanel compose chain) — the commit (a) writes identity ordering as in Phase 1, (b) registers/refreshes the node with the SortWorker, and (c) requests one sort with the current camera.matrixWorldInverse × mesh.matrixWorld. On resolve: generation-check (drop stale), aSortedIndex.set(ordering) + addUpdateRange(0, count) + needsUpdate + request a frame. At most one in-flight sort per node from day one.

  4. Mode-switch hook: switching a gsplat layer TO normal cannot simply “register+sort” — the SortWorker has no centers for a node that was order-independent at its last commit (registration is gated, and the staged arrays were transferred/discarded). Instead the switch invalidates the node’s noop stamp (userData.committedData = undefined) and requests a view update: the memoized-concat noop path would otherwise skip re-projection entirely, so clearing the stamp forces a standard reprocess→commit, which registers + sorts like any commit (SliceCache still holds the source nD data — this is one O(N) re-projection per mode switch, a rare user action). Switching AWAY just stops future sorts; a sorted order is harmless under commutative modes — no identity reset needed.

Tests / exit criteria

  • Rust unit tests (make test-wasm) + TS twin parity: ordering is a permutation, back-to-front monotone, stable, behind-camera splats last; unit-extremes cases (nm-scale ~1e-6 and km-scale ~1e6 coordinate magnitudes, plus zmax==zmin) produce valid orderings — the exact failure mode raw-f16 keys would have; benchmark-wasm entry with an absolute throughput floor. Implementation delta (Phase 2, measured 2026-07): the original ≥100M splats/s target was a pre-implementation estimate; the release WASM kernel measures ~74M splats/s at 1M worst-case spatially-incoherent splats on an M-class dev machine (~96-114M/s native Rust — the gap is WASM execution overhead + wasm-bindgen boundary copies; real Morton-coherent gsplat data sorts faster). Two faster-looking variants (2×8-bit LSD radix; branchless multi-lane min/max + interleaved partial histograms) were measured and REJECTED — both slower than the simple three-pass counting sort (rationale in depth_sort.rs). The floor (perf-budget.test.ts, opt-in pnpm test:perf) is ≥50M splats/s, report-only unless LUXAR_PERF_QUIET_HOST=1 explicitly enables quiet-host enforcement — below the worst-case measurement, far above any real regression (debug build, comparison sort, alloc-per-element). Budget impact: a worst-case 10M-splat sort is ~135-190 ms in the background SortWorker — above §2’s 30-100 ms target but off the frame path; acceptable staleness (Phase 3 sorts are throttled and generation-guarded). Follow-up headroom if ever needed: WASM simd128, or a stateful register/sort WASM API keeping centers wasm-resident across camera-orbit re-sorts (eliminates the 12MB/sort copy-in).

  • Unit (mock worker): generation-guard drops stale results; single-in-flight; release-on-dispose; transfer-detach assertions (registered buffers unusable on main thread afterwards — intentional).

  • E2E: the Phase-0 overlap fixture renders correctly in normal mode after load settle, from the load-time camera; additive baselines still unchanged (identity path).

Risk: low-moderate — new worker lifecycle, but the generation guard and single-in-flight rule are small state machines with unit coverage.


6. Phase 3 — Camera-triggered re-sort

The live feature: sorted order tracks the camera. Pure scheduling on top of Phase 2’s engine.

Changes

  1. Scheduler: new per-frame callback 'depth-sort-scheduler' registered in core/app/init/pipeline.ts beside 'lod-group-selector' (which already owns the per-frame camera scratch pattern — same no-per-frame-allocation invariant). Per registered order-dependent node: track view direction + position at last completed sort; when dot(viewDir, lastViewDir) < cos(θ) or translation exceeds a bounds-relative fraction, dispatch a sort (Phase 2’s single-in-flight + generation guard already handle overlap/staleness). Frames between dispatch and resolve render the previous order — bounded staleness, no popping-to-wrong, standard 3DGS behavior. Skip dispatch while the node’s loader has a pending view update (same signal the refinement loop consults).

  2. Config: new config/sections/depth-sort/ (data/types/validate/README, following the adaptive-dpr section shape): enabled (default true), angleThresholdDeg (default ≈3°), translationFraction (default 0.05). URL escape hatch ?depthSort=0 in config/url-params.ts (pattern: ?gpuBudgetMB=) — pins identity ordering for deterministic E2E/visual runs.

  3. Monitor/profiling: sort latency + bytes as an update-profiler stage and a data-loading-monitor line, per the monitor design language (semantic color only, tabular numerals).

Tests / exit criteria

  • Unit (mock worker): threshold hysteresis; no dispatch during pending loads; ?depthSort=0 inert.

  • E2E: orbit around the overlap fixture in normal mode with settle-based assertions — final settled frame matches the sorted baseline from every tested angle; ?depthSort=0 reproduces Phase-2 behavior (determinism guard).

  • Perf: at 1M sorted splats, orbiting holds target FPS with sort round-trip < 50 ms and zero main-thread stalls > 4 ms attributable to sorting (ordering upload is the only main-thread GPU cost: 4 MB @ 1M).

Risks: worker memory (retained centers = 12 B/splat; 120 MB at 10M — acceptable against the 640 MB texture, releaseNode bounds it, and f16 center retention is a documented follow-up halving); visible re-sort “settling” during fast orbits (tune θ; it replaces being always wrong today).


7. Phase 4 (optional, perf) — Partial texture uploads

Not required for correctness — a measured-win phase banking the storage refactor’s dividend. Key implementation finding (2026-07): three r184 already ships texture-side partial uploads — Texture.updateRanges / addUpdateRange(startFloat, countFloat), honored by the classic WebGLRenderer (webgl/WebGLTextures.js::updateTexture: it takes the whole-image texSubImage2D path ONLY when updateRanges is empty, else uploads per-range). Since luxar’s default backend is classic WebGL, no copyTextureToTexture staging path is needed for the win; the WebGPU backends ignore the ranges and re-upload the whole image (correct, just not yet partial). Three sub-stages:

Stage 1 — slack elimination (LANDED 2026-07)

Every non-noop commit re-uploaded the entire capacity-sized RGBA32F texture, including the pool’s 1.5× growth headroom and best-fit slack rows. writeSplatTexels now registers per-row updateRanges over [0, count) via the dirty-range helper (today registerElementTexelDirtyRange in element-storage.ts), so only the live rows upload. Row-split because the WebGL path uploads each range with height = 1; splat×4-texel alignment on a width-multiple-of-4 texture guarantees no row straddle; ranges self-collapse into one contiguous span per call (the WebGPU backends never clear them). Above FULL_UPLOAD_ROW_FRACTION (0.75) of rows dirty, it falls back to the single full-image upload (per-row call overhead outweighs the saving). Measured (gsplats_4d_neuromast_2ch, classic WebGL): initial-load upload 3.31 MB → 2.20 MB (33%, the fresh-alloc headroom); per-commit slack-scrub case 5.81 MB → 2.36 MB (59%). Pixel-identical (texel content unchanged; shaders only read [0, count)).

Stage 2 — append-only writes for ladder streaming (LANDED 2026-07, all three geometries)

Progressive ladder refinement re-uploaded the whole committed prefix per level. An append-only commit writes only the new splat span: writeSplatTexels(…, { fromSplat }) skips the prefix write loop and registers just [fromSplat, count) (the registerSplatTexelDirtyRange general span already supported this from Stage 1).

Design finding (revises the earlier plan): no projection-kernel contract change is needed. The nD→3D projection is per-splat-independent and order-preserving, and the loader appends LOD levels in order, so a k-level input concat is a byte-identical prefix of the (k+1)-level concat. Under an unchanged view state the (k+1)-level projection’s first prevCount visible outputs are therefore byte-identical to the previous commit’s entire output — so outputPrefixCount === prevCount (the committed visibleSplatCount) is deducible, not measured from a returned visibility mask. The trust chain is: (1) a forward-chained prefix-lineage WeakMap (types/prefix-lineage.ts), stamped in the loader’s concatenateMemoized and checked in commit by identity against getCommittedData(mesh) — a match proves same generation (a view change bumps the loader reset generation → no parent → append rejected → full rewrite), genuine extension, and that the GPU still holds that parent’s projection; (2) the append gate also requires !attributesRebuilt && geometry === prevGeometry (pool reused this node’s buffers in place — which also bounds count ≤ capacity), gpuPrefixIntact, splatCount > prevCount, and committedTruncate === readTruncate(mesh) (the one hole: uTruncate is a material uniform outside viewStatesEqual). The splat texture remains suffix-only, but aSortedIndex resets to full identity: preserving an old depth-sorted prefix and appending a storage-order suffix makes alpha-over draw two independently ordered populations until the commit-triggered sort lands. The bbox recompute stays full (CPU-only, no GPU upload). A context-restore full-dirty hook (NodeFactory.rebuildAfterContextRestore) marks every splat texture + aSortedIndex full-dirty and clears gpuPrefixIntact, because a WebGL context loss zeroes the GPU buffers while the CPU mirror survives — without the flag, a following append would downgrade the pending full upload to a stale-prefix partial.

The GSplat identity reset is a perceptual fallback, not a settled correctness fix for #1987: the reported persistent gastrulation ghost still lacks a visual repro. The change treats the hybrid prefix/suffix ordering in dense translucent splats as more likely to read as a detached second population, so the interim draw now gives up the useful sorted prefix and renders the whole grown node in storage order until the fresh permutation lands. The cost is one SortWorker round-trip plus, on the classic WebGL backend, ceil(count / SORTED_INDEX_CHUNK_ELEMENTS) drawn apply frames per append — about 1–2 frames for the roughly 200k-splat gastrulation node, but several frames for multi-million-splat nodes. Holding instanceCount at prevCount was rejected: the coordinator would have to raise it when an ordering lands, but no ordering is guaranteed under commutative blending, ?depthSort=0, or an unavailable SortWorker, so the node could stall indefinitely. Points and Lines retain the hybrid because they have no corresponding observed artifact and preserving their valid sorted prefix minimizes interim disorder; their tests explicitly pin that choice.

Three-geometry symmetric (LANDED 2026-07, Points/Lines): the same append is writePointTexels(…, { fromPoint }) for Points (since the points texture migration, §8 / PR #630) and writeLineTexels(…, { fromSegment }) for Lines (segment units; originally landed as writeInterleavedAttribute(…, { fromInstance }) and carried over unchanged by the lines texture migration, §8 / PR-C), both routed through the pool facade’s uniform fromInstance option (their ranged uploads already eliminated slack, so this was their remaining O(k·N)→O(N) win). Order-preservation verified: the Points zero-radius drop is a per-level order-preserving forward scan run before concat, and Lines clipping is an order-preserving drop + in-place endpoint clip (never splits 1→2, never reorders) — so prefix identity holds for both under the same lineage trust chain (the shared helper now lives in types/prefix-lineage.ts). Their gates mirror the gsplat one minus preserveOrdering/committedTruncate (no permutation, no external projection input), plus one extra conjunct: optional-field presence must match the committed parent — the Points concat is all-or-nothing per field (concatOptionalField), so a Float32↔absent flip re-fills the column with a constant that need not match the committed prefix; same for Lines colors/sharpness whose presence flip re-fills the prefix through the interpolation kernel. A prerequisite fix rode along: concatenateLinesData now fills missing-part sharpness with 0.5 (the projection’s null-sharpness default), not 0.0 — a mixed-sharpness ladder previously popped razor-sharp on the sharpness-less parts even on full rewrites. The context-restore hook covers all three types through one shared branch (element textures + aSortedIndex marked full-dirty + gpuPrefixIntact cleared — lines’ historical interleaved-buffer branch died with their texture migration). The shared lineage WeakMap is retention-capped: setPrefixParent deletes the superseded parent’s own entry (chain depth 1) and every commit consumes-and-clears the entry after its gate check, so intermediate concat results of a ladder are never pinned via lineage (a WeakMap holds values strongly while the key lives — an uncapped chain retained ~(n−1)/2 × the final CPU arrays on an n-level ladder).

Stage 3 — WebGPU partial-upload parity ([POST] — MEASURED-REJECT 2026-07-25, revisit criteria below)

Built and measured per the measure-first doctrine; rejected at the ≥10% in-app bar (archived as a closed PR with the implementation + all numbers). What the campaign established:

  • Raw mechanism win is real and large. three r184’s WebGPU backend ignores texture updateRanges (whole-image queue.writeTexture per version bump; only attribute ranges and DataArrayTexture.layerUpdates are consumed). Probed on Dawn/Metal (M4 Max headless Chrome): writeTexture ≈ 2.2 GB/s, so a 5M-splat element texture (305 MB) costs ~129 ms per full re-upload vs ~7 ms for a ladder-append span (18×) and ~56 ms for a 40%-rows scrub span; per-ROW ranged calls beat even a single contiguous block. The working implementation (an instance-level wrapper on WebGPUBackend.updateTexture consuming the existing per-row ranges via ranged writeTexture + consume-and-clear, knee = ∞) lives in the closed PR.

  • In-app, no available workload clears the bar. Timelapse-nav (neuromast, ~8 MB textures): 0% — uploads hide inside the frame. Synthetic 5M full-dirty recommits: −1.6% (100% dirty moves the same bytes either way). Append-shaped ladder loads (visible-human 1.9M single ladder — 4096×2800 capacity-padded texture, 183 MB per full re-upload; MatrixCity 13.6M ≈ 250k-splat tiles ⇒ ≤16 MB per texture ≈ one frame per full re-upload): flat on every metric over the full load window (fixed-45 s-window A/B reruns, 2026-07-26; an earlier “+11–20% tail improvement” reading came from a ~7 s partial-window instrument artifact — the v1 bench’s settle detector silently broke — and is superseded; raw JSONs archived on the closed PR’s branch under delme/webgpu-writetexture-probe/bench-results/). Loads are decode-dominated.

  • Revisit when any of these hold: (a) WebGPU becomes a default/primary backend; (b) single-texture ladders at ≥5M splats ship as real datasets (each ladder chunk then costs ~16 dropped frames on main vs ~1 with the wrapper for small append chunks — a doubling ladder’s ~50% tail chunk is still ~8 frames); (c) an upstream three release consumes texture updateRanges (file the PR with the probe numbers — writeTexture already supports sub-rectangles with no bytesPerRow alignment constraint, and common/Textures.js:357 invokes texture.onUpdate on every backend, so consume-and-clear semantics compose cleanly with luxar’s pendingFullUpload guard). The archive branch carries REVIVAL-NOTES.md (audit 2026-07-26): read it before reviving — leading items are the knee=∞ extrapolation past 40% dirty and the (now-pinned) RGBA32F format assumption.

Correctness on WebGPU remains automatic (full re-upload); the classic-WebGL Stage 1/2 wins are unaffected.

Gate Stages 2/3 on the Update Buffers + Load Arrays profiler stages over the timelapse-nav benchmark (see the decode-bottlenecks findings) — ship only with numbers. Exit criterion: refinement passes upload O(new splats), not O(total).


8. Explicit deferrals

  • Points/Lines symmetry (LANDED 2026-07 — completed by the volumetric phases 3–4 arc): Points storage LANDED — the lift happened exactly as designed: the storage layer generalized into element-texture-layout.ts (parameterized ElementTextureLayout; gsplat = 4 texels, point = 3: center+radius / rgb+sharpness / scalar+alpha-reserved) + element-storage.ts (attach / dirty-range / writeSortedIndex*), and points now render from uPointTex via aSortedIndex with per-node materials (52 B/point vs 36 interleaved; the alpha slot pre-positions volumetric Phase 3’s RGBA colors). Points sort integration LANDED (the arc’s PR-B): the coordinator went geometry-neutral (noteGSplatsCommit/noteGSplatsBlendingModeSwitch → noteDepthSortCommit/noteDepthSortBlendingModeSwitch) and the points commit registers at both pool and non-pool paths with a LAZY centers provider (fresh Float32 copy of data.positions, paid only on the sorted path — the committed/lineage array’s buffer must never itself be transferred; the copy also widens Float16). Points normal mode is now correctly back-to-front sorted, joins the camera-motion re-sort scheduler, the cross-node global renderOrder scale, and the preserveOrdering same-count-recommit prior; the LayersPanel mode-switch hook, per-mesh disposal release, and lazy-LOD demotion release all cover points. Order-dependence is judged directly on needsDepthSort(mode) for every geometry type: the interim effectiveGeometryMode downgrade helper (which kept points/lines volumetric additive-unsorted until the volumetric math shipped) became identity once volumetric Phase 4 landed and was DELETED from rendering/blending-state.ts — points volumetric has sorted since volumetric phase 3, lines since phase 4. Lines storage + sort integration LANDED (the arc’s PR-C, completing the symmetry): 6 texels/segment (LINE_TEXTURE_LAYOUT, width floor(4096/6)·6 = 4092; layout in line-geometry.ts — startPos+startWidth / endPos+endWidth / startColor+startSharpness / endColor+endSharpness / segLen+per-endpoint cap-suppression scalars / scalars+reserved alphas, texel5.zw pre-positioning volumetric Phase 4’s per-endpoint opacity), rendered from uLineTex via aSortedIndex with per-node materials. The migration retired three lines-only mechanisms in one stroke: the interleaved-buffer packing, the colormap ATTRIBUTE-SET toggle (the fixed layout always carries the scalar slots, so hasScalars no longer bucketing the pool nor rebuilding geometry — presence rides the userData.hasScalars stamp like points), and the material-manager LRU (the last cached material kind; all visual materials are now per node). The commit registers LAZY segment MIDPOINTS ((start+end)/2, fresh Float32 — the standard approximation; artifacts only when long segments interleave, subdivision if ever needed) and shares preserveOrdering, the append fast path (writeLineTexels({fromSegment})), the LayersPanel mode-switch hook, and disposal/demotion releases. Lines volumetric now sorts too (volumetric Phase 4 shipped 2026-07-24): lines render the real emission–absorption math and the same lazy midpoint provider feeds the sort through needsDepthSort(mode) — zero coordinator change. Pick vElementId reads aSortedIndex (storage slot); the context-restore hook covers lines through the same element-texture full-dirty branch as points/gsplats.

  • Option 3b (WebGPU compute sort): writes the same aSortedIndex/ordering interface from a compute pass (three r184 renderer.computeAsync + storage() TSL nodes), eliminating the worker round-trip on the WebGPU backend. Requires a parity-policy carve-out (a compute sort has no GLSL twin — functional-equivalence tests on ordering output instead).

  • RGBA16F splat texture (32 B/splat): pure data-conversion change post-Phase 1; textures avoid every vertex-format pitfall that sank the dd7c4478 interleaved narrowing.

  • Weighted-blended OIT: a sort-free translucent mode candidate; orthogonal to and composable with this plan (it would reuse Phase 0’s real-alpha output).

  • opaque mode at fractional opacity: order-dependent today, stays so — degenerate configuration, documented rather than fixed.


9. Sequencing, sizing, and ship gates

Phase

Ships alone?

Size

Hard gate before merge

0 — normal-mode alpha

Yes (visible improvement)

S (2 shader pairs, 1 blend-state helper)

Parity (incl. WebGPU bridge run) + overlap fixture + unchanged additive baselines

1 — texture storage

Yes (invisible; perf-neutral)

L (adapter, 4 shader stacks, materials, pool, budget, sync helper)

Pixel-identical E2E; commit-perf non-regression

2 — SortWorker + sort-at-commit

Yes (correct at rest)

M/L (WASM+TS kernel, SortWorker, commit wiring)

Kernel parity + benchmark floor + correct-after-settle E2E

3 — live re-sort

Yes (the feature)

M (scheduler, config, monitor)

Orbit E2E + FPS/latency budget + ?depthSort=0 determinism

4 — partial uploads

Optional (Stages 1–2 landed, all three geometries)

S (Stage 1) / M (Stage 2 gsplats) / S (Stage 2 Points/Lines)

Measured upload win; Stage 2 GSplats intentionally changes the interim append ordering as documented above

One PR per phase, in order; each leaves main shippable. Phase 1 is the only high-blast-radius change and is deliberately behavior-preserving so its review is a pure refactor review. Full E2E suite before each merge per repo policy; make test-wasm + benchmark-wasm for Phase 2; README updates (rendering/README.md architecture notes, materials/gsplat/README.md contract, workers/ README) ride each phase’s PR.


10. Risk register (ranked; what is actually hard about this)

  1. Three r184 half-supported vertex paths — the twice-burned pattern (highest risk). This repo has already been bitten twice by r184 vertex-pipeline gaps: GLSLNodeBuilder hardcoding gl_PointSize = 1.0 (forced the Points→instanced-quad migration) and gpuType: HalfFloatType silently ignored on interleaved attributes (the dd7c4478 Float16 revert that shipped a regression). Phase 1 bets on two more under-exercised paths: a uint instanced attribute and vertex-stage textureLoad, each through both the raw-GLSL ShaderMaterial path and the TSL node-builder (including WebGPURenderer’s WebGL2-fallback GLSL emission). Raw WebGL2 support is verified (§1); what is NOT verified is what three’s node builders emit. Mitigation — mandatory spike before Phase 1 is scoped: a 1-day throwaway probe rendering one quad-instanced mesh with a uint attribute + vertex texelFetch on all three surfaces (WebGL, WebGPU-native, webgpuForceWebgl). Fallback if uint attributes are broken in the node builder: a float attribute carries exact integers to 2^24 = 16.7M — exactly our per-node cap, so it works but eliminates headroom.

    SPIKE VERDICT (2026-07-13, three 0.184.0 — RISK RETIRED, all three surfaces PASS). Probe: 16 instanced quads, Uint32Array aSortedIndex carrying the permutation (7i+3) mod 16, colors vertex-fetched from an RGBA32F DataTexture (NearestFilter); control row fetches by gl_InstanceID/instanceIndex. Verified by pixel classification against a 16-color palette.

    Surface

    Backend reported

    uint attr + vertex fetch

    WebGLRenderer + GLSL3 ShaderMaterial (in uint + texelFetch)

    webgl2

    PASS

    WebGPURenderer + TSL (attribute('…','uint') + textureLoad)

    webgpu-native (real adapter, headless --enable-unsafe-webgpu)

    PASS

    WebGPURenderer({forceWebGL:true}) — TSL→GLSL emission

    webgl2-fallback

    PASS

    Operational notes for Phase 1 tests: (a) WebGPURenderer.readRenderTargetPixelsAsync(rt,x,y,w,h) RETURNS the buffer — no out-param, unlike the WebGL variant (the existing dispatch in post-processing/hdr/pixel-utils.ts handles this); (b) drawImage from a WebGPU canvas races presentation (observed all-black readbacks) — sample via render-target readback, never via 2D-canvas copy; (c) readback row order differs per backend (framebufferYDown) — disambiguate with a known asymmetry. Spike code: delme/spike-uint-vtf/ (gitignored, throwaway).

    DEBT-REMEDIATION UPDATE (2026-07-13). The vacuous-parity issue this spike’s methodology flagged is FIXED: harness meshes now use production assembly, every sprite parity test carries a non-vacuousness guard + a footprint-invariant per-covered-pixel metric, and two REAL production TSL bugs the vacuousness had hidden were found and fixed (branch-scoped materialization of shared eigen values in shader-tsl.ts/pick.tsl.ts — quads degenerate for near-diagonal Σ2D; both vertex stages now trace inside Fn() with explicit .toVar() statements — the TSL house rule going forward, which #1683 later had to EXTEND to the fragment stage: a value shared between colorNode and depthNode must be assigned in an unconditional prologue both entry points call first, because the order three builds the two entry points in is not part of its API). Post-fix, GLSL↔TSL parity is PIXEL-EXACT (per-covered diff 0.00) on every variant. The same remediation also SUPERSEDED part of §3’s Phase-0 wording: the GLSL additive alpha-MaxEquation guard referenced there was subsequently deleted (additive blending unified onto the shared helpers; the raw-scene-hdr capture path now sanitizes alpha at readback for all geometry types). RESOLVED (same day, /double-check pass): the production “bowtie quads” under ?renderer=webgpu were the SECOND half of the fragcoord story — both TSL fragments anchored the Gaussian at screenCoordinate (TOP-LEFT origin on both backends by three’s contract; the fallback emits size.y - gl_FragCoord.y) while vCenterScreen is bottom-left, mirroring the kernel about the horizontal midline. Centered mirror-symmetric parity fixtures are structurally blind to a y-flip, which is how pixel-exact parity coexisted with broken production. Fixed by reconstructing the bottom-left fragcoord in both factories — the un-flip term is three’s screenSize (the bound render target’s size, the exact term the builder’s own flip used; the app-stamped uResolution differs whenever the target isn’t drawing-buffer-sized, e.g. under SSAA), i.e. vec2(x, screenSize.y - y); a permanently off-center gsplat-offcenter parity variant now guards the whole convention class. Production WebGPU gsplat rendering verified pixel-equivalent to WebGL on both fixtures.

  2. Product risk: normal mode may not look right on HDR scientific data (second, and the cheapest to retire). Alpha-over compositing happens in linear HDR before tone mapping; coverage alpha is clamp(intensity·opacity), so dim splats (amplitude ≪ 1, common in microscopy) occlude almost nothing (near-additive look) while saturated ones occlude hard — and the aesthetic verdict on real data is unknown. The nightmare is discovering after Phases 1-3 that sorted normal isn’t what anyone wants. Mitigation: Phase 0 is deliberately first and standalone — evaluate on real datasets (neuromast, h2afva) at the Phase-0 checkpoint before any Phase 1-3 investment. Order-dependence artifacts will still exist at that checkpoint; judge color/occlusion character, not popping.

    FIRST REAL-DATA VERDICT (2026-07-29, PR #916): POSITIVE on a heavy-background light-sheet volume. The Tribolium embryo demo (demo_gsplats_3d_tribolium_embryo.py, 965×1871×991 Zeiss LightSheet Z.1 volume, amplitudes scaled 0.03 — squarely the “amplitude ≪ 1” microscopy regime this risk worries about) now ships blending_mode="normal" precisely because it looks better than the accumulating modes. Under volumetric the volume’s heavy diffuse background integrates along every ray and saturates into a solid slab with the embryo buried inside it; under normal the background stops accumulating and the cell-surface nuclei resolve individually. Judged as this risk asks — color/occlusion character at a fixed camera, not popping.

    Two caveats on how far this generalizes. (a) It is one dataset at one viewpoint, judged by eye against a live A/B of the same scene; it says nothing about the async ordering state machine (risk 4) or about scenes whose subject is the accumulated volume. (b) The dominant mechanism was not the coverage-alpha behaviour this risk reasons about — it was the peak-vs-ray-integral projection split (usesPeakProjection), which removes the background accumulation outright regardless of how little dim splats occlude one another. So the prediction above (“dim splats occlude almost nothing ⇒ near-additive look”) is right about the alpha-over term yet still under-predicts how different normal looks from additive/volumetric on such data, because the projection change dominates. Because nothing sums any more, the scene needed +1.97 stops of exposure to sit at a normal level — worth expecting for any future normal-mode conversion.

    FOLLOW-UP VERDICT (2026-08-22, PR #1880): REVERSED after tuning the coupled volumetric controls. The same Tribolium scene now ships blending_mode="volumetric" with strong absorption (κ=3.13) and opacity held down to 0.06. That pairing preserves the emission-absorption depth cue without letting the diffuse background saturate into the slab described above. The original fixed-opacity comparison remains useful history, but it was not a verdict on volumetric after jointly tuning absorption and emission. The scene still uses +1.97 stops: under the new settings it lifts the deliberately low-opacity ray integral rather than compensating for peak projection, so exposure alone is not diagnostic of the blending mode.

    BASIS NOTE (same PR). Both verdicts above argue about how a blending mode handles this volume’s “heavy diffuse background” — and the fit itself then changed underneath them. The floor now subtracts the specimen’s own ~675-count background rather than only the ~205-count detector offset, which removes ~84% of the scene’s emitted mass while leaving nuclei brightness roughly unchanged. The background both entries reason about is therefore largely no longer in the data, so a future A/B on this scene starts from a different basis than either verdict, and neither should be read as a general statement about accumulating modes on heavy-background data.

  3. Phase 1 blast radius under a pixel-identical bar. Four shader stacks rewritten simultaneously, the material system migrated to per-node (touching registration/camera-broadcast/dispose lifecycle — the same neighborhood as the known LRU-dispose engine-review finding), and the pool adapter rebuilt — with “unchanged visual baselines” as the merge gate. The gate is the right call (it makes review tractable) but it concentrates the campaign’s regression risk in one PR. Mitigation: land it right after a green-baseline refresh; lean on ?dpr=1 pinning; budget review time like a rendering-engine change, not a refactor.

  4. The async ordering state machine. Individually trivial rules (single-in-flight, generation match, length == instanceCount, identity on reuse, release on dispose) — but their product across commits, ladder appends, mode switches, pool reuse, dataset switches, LOD-child swaps, and context loss is where corrupt-permutation flicker hides. Mitigation: enforce the apply-invariant (generation equal AND ordering.length == current count, else drop silently) in exactly one function; property-test the state machine with a mock worker (random interleavings).

  5. Vertex-texture-fetch and RGBA32F-upload performance variance. 16 texelFetches per splat (4 verts × 4 texels) vs one interleaved cache-line read today; VTF is fast on desktop/Apple GPUs (SparkJS ships this at scale) but historically weaker on some mobile tiles, and full RGBA32F texSubImage2D uploads have driver-dependent pitch/conversion costs. The Phase-1 perf gate catches it; escape hatches are pre-designed (RGBA16F halves both footprint and fetch bandwidth).

  6. Logistics that bite late: the SortWorker is a new Vite ?worker chunk that must ride into luxar export offline bundles and the native launchers (add an export smoke test); WASM-path forwarding into a second worker; parity-harness extension to uint-attribute/texture-fetch shaders; E2E stale-server pitfall vs “pixel-identical” claims.


11. Adversarial re-check changelogs (all findings applied above)

Rev 4 (2026-07-14, pre-Phase-1 re-base; Rev 4.1 doc-truth pass same day). Phase 1 was written against pool machinery that PR #511’s own debt-remediation later deleted; re-based onto the shipped contracts:

  • Growth is release + reacquire (growGSplatsGeometry/rebuildInterleavedBuffer are gone): the splat texture is created and disposed WITH its pool geometry; in-place texture reallocation on a rendered geometry is forbidden — it is the exact Info.memoryMap strand class the buffer fix killed (empirically: +1 pinned generation per grow, permanent, on both native backends’ pre-fix code). No carry-forward; commits rewrite the full count.

  • Non-pool fallback follows the fresh-swap+dispose pattern (broken-fallback finding long since fixed in #503).

  • The texelFetch prologue must be Fn-traced .toVar() statements (TSL house rule from the uninitialized-eigen bug); downstream now includes the unified perspectiveNearFade, unconditional coverage fade, and screenSize un-flip — all untouched by the prologue swap.

  • invalidateRenderObjectFor now also re-points userData.pickNode at the new geometry — Phase 1’s pick side inherits the geometry swap for free; only the pick MATERIAL’s uSplatTex rebind rides the new sync helper.

  • Exit criteria hardened with the committed tooling: scripts/perf/timelapse-nav-bench.mjs (perf gate) and scripts/perf/grow-leak-probe.mjs on native WebGPU (texture-lifecycle leak gate), plus the post-campaign parity-variant inventory and snapshot-regen discipline (fresh-server rule).

  • Status: Phase 0 merged to main (squash e1e079e6); B9a-c shader symmetry + native-WebGPU validation in PR #523.

Rev 3

  1. Sort keys changed from raw f16 bit patterns to min/max-normalized uint16 (65536 buckets). SparkJS’s f16 keys assume scene-scale depths; Luxar scenes carry physical units from nm to km, so raw camera-space depths can overflow f16 (all splats → Inf bucket) or underflow to 0 (all splats → one bucket) — either silently destroys the ordering. Normalized integer keys are scale-invariant at identical cost; unit-extremes tests added.

  2. Mode-switch data-flow gap closed: switching a layer TO normal post-commit found the SortWorker without centers (registration is mode-gated) — and a naive re-request would hit the memoized noop path and skip re-projection. Fix: invalidate committedData + request a view update (one O(N) reprocess per switch).

  3. generation pinned: per-node monotonic non-noop-commit counter — loadedViewVersion is insufficient because ladder appends share a view version while the count grows, and a short stale permutation applied to a grown buffer is corrupt, not merely stale. Noop commits keep in-flight sorts valid.

  4. Blend-state completeness: transparent: true added to the pinned gsplat-normal state (transparent-pass classification / WebGPU pipeline key); vestigial opacity parameter dropped from getGSplatNormalBlendingState() (depthWrite is unconditionally false, so the state is opacity-independent).

  5. Sort gating source corrected: gate on the material’s live userData.blendingMode (what actually renders) rather than “the attrs-composer chain” — composition only runs via the LayersPanel; at initial load the material’s mode comes from creation attrs.

Rev 2

  1. Structural: sorting moved out of the projection call entirely — plain-3D data (ndim > 3 worker gate) never reaches the projection worker, so the original “sort in the projection worker” design would have sorted 10M splats on the main thread for the most common case. The SortWorker (originally Phase 3) is now the single owner of ordering and lands in Phase 2; ProjectionViewState is no longer touched.

  2. TSL double-premultiplication trap: material.premultipliedAlpha injects an automatic RGB×alpha output node on NodeMaterial (three.webgpu.js:21817) — banned; blend factors set explicitly instead.

  3. WebGPU-bridge constraint scoped correctly: the gl.getError() issue is separate alpha-channel equation/factor state, not CustomBlending per se (max mode ships CustomBlending on TSL today) — the premult-normal state uses CustomBlending with a symmetric alpha channel.

  4. Depth-state gap closed: gsplat premult-normal pins depthWrite: false unconditionally (generic normal’s opacity >= 0.99 flip would let near-transparent splat fragments write depth → occlusion halos); trade-off documented.

  5. Blend-state API shape fixed: new getGSplatNormalBlendingState() helper instead of widening the shared getCompleteBlendingState (which has no geometry-type axis); the stale “TSL needs no gsplat override” comment is called out for update.

  6. Arithmetic: VRAM delta corrected from +23% to +31% (68 B vs 52 B — the ordering attribute counts too).

  7. Addressing consistency: bit-math constants were hardcoded for width 4096 while the width was capability-dependent — now compile-time defines/TSL params derived from actual width; “16384²” phrasing corrected (height is the scaling axis).

  8. Uint32 attribute binding made explicit: verified r184 auto-selects vertexAttribIPointer for Uint32Array (three.module.js:1950) — noted as a relied-upon fact with parity coverage.

  9. Spurious sortKeyCache texel label removed (was never specified).

  10. Sort-throughput numbers reconciled: ~30-100 ms at 10M and a ≥100M splats/s benchmark floor (previous table said 10-30 ms while the gate implied 200 ms).

  11. Texture rebind rides material-sync-helpers.ts (established post-commit render+pick sync seam) instead of ad-hoc commit code.

  12. Transfer-safety proven, not assumed: ndim==3 fast path .slice()s (projection/gsplats.ts:148-150), so registering centers with the SortWorker by transfer can never detach a SliceCache buffer; added transfer-detach assertions to Phase 2 tests, plus SortWorker WASM-path init, app-teardown dispose(), and a per-node material-disposal leak check as explicit items.