# GSplat Depth Sorting — Phased Implementation Plan (Option 3a) > **Status**: Spike + Phase 0 MERGED TO MAIN (PR #511, squash `e1e079e6`); shader-symmetry + perf-validation follow-ups MERGED (PR #523, squash `fd243382`); **Phase 1 IMPLEMENTED** (branch `gsplat-depth-sorting-phase1`): texture-backed splat storage + always-on `aSortedIndex`, per-node gsplat materials, `syncGSplatMaterialWithGeometry` commit rebind, texture-lifetime-pinned-to-geometry via dispose listener. Implementation deltas vs the Rev 4 text: (a) `maxTextureSize` needed net-new plumbing (`RendererCapabilities.maxTextureSize` → module-scoped `element-texture-layout.ts` (née splat-texture-layout), the `gpu-byte-budget` live-authority pattern — the pool ctor is too far from the renderer); (b) texture bytes are counted by stashing the texture on `geometry.userData` (visible to `estimateGeometryBytes`) rather than a `PooledBuffer` field + per-site dispose edits; (c) pick `vElementId` now derives from `aSortedIndex` (storage slot) instead of `gl_InstanceID` — identical under identity ordering, correct-by-construction once Phase 2 permutes; (d) TSL `textureSize()` is typed `uint` (WGSL convention) while GLSL returns `int` — the width read is wrapped in `int(...)`, without which the generated GLSL fails to compile on the forceWebGL backend (risk-register class #1, found by the parity suite's non-vacuousness guard). Phase 1 MERGED as PR #535 (+ follow-up coverage PRs #537/#538). **Phase 2 MERGED** (PR #540, squash `2471327f`): `sort_splats_by_depth` WASM kernel + frounded TS twin (exact-permutation parity), persistent Comlink SortWorker (`workers/sort-worker.ts`), main-thread `rendering/depth-sort-coordinator.ts` (generation guard, one-in-flight-per-node, committedData demotion guard), commit wiring + LayersPanel mode-switch hook, non-vacuous E2E gate on the new front-to-back `test_gsplats_normal_overlap_reversed` fixture (fail-first verified). Implementation deltas vs the Rev 4 text: (a) the throughput floor moved from the estimated 100M splats/s to a measured 50M/s, report-only unless `LUXAR_PERF_QUIET_HOST=1` (details in §5's exit-criteria note); (b) the camera/render/reprocess callbacks are injected via `configureDepthSort` at app init (SceneLoader deliberately owns no camera state, so commit-time model-view comes from the coordinator, not the loader); (c) an ordering resolve additionally requires the mesh's `committedData` stamp to survive — LOD demotion returns the geometry to the evictable pool, and the cleared stamp is exactly that signal; (d) the sort-time model-view is derived from FRESHLY updated matrices (`mesh.updateWorldMatrix` + a local `inverse(camera.matrixWorld)`) — the renderer-maintained `camera.matrixWorldInverse` / `mesh.matrixWorld` caches are stale when a commit fires before the next frame (first commit of a load, idle-paused loop); (e) a queued re-sort is drained even when the in-flight sort RPC fails, and the TS twin's key clamp uses a `<` comparison to reproduce Rust's `f32::min(NaN, 65535) == 65535` semantics exactly (post-merge double-check hardening). **Phase 3 IMPLEMENTED** (PR #553, branch `depth-sort-phase3`): the `'depth-sort-scheduler'` per-frame callback + `evaluateDepthSortPerFrame` in the coordinator, `config/sections/depth-sort/` (`enabled`/`angleThresholdDeg`/`translationFraction`), `?depthSort=0` escape hatch, and the 'Depth Sort' monitor line. Implementation deltas vs the §6 text: (a) the tracked pose is the model-view matrix's z-row (unit axis direction + normalized offset `m14/|axis|`) rather than a viewDir+position pair — the kernel sorts by view-space z = axis·p + offset, so that row fully determines the permutation; translation ORTHOGONAL to the view axis provably cannot change the ordering and never dispatches (the translation trigger compares offsets along the axis against `translationFraction × boundingSphere.radius`, catching the behind-camera-set changes that view-axis motion causes); (b) the monitor integration is a third persistent update-profiler root ('Depth Sort', `beginDepthSortPass()`, mirroring the 'LOD Refinement' pass isolation) rather than an update-session *stage* — camera-triggered sorts run outside any update session; each pass records the dispatch→applied round-trip latency plus splat count and ordering-upload bytes; (c) `?depthSort=0` disables the ENTIRE subsystem (worker never spawns, commits keep identity ordering, the mode-switch reprocess hook goes inert) — not just the scheduler — because a determinism escape hatch that still sorted at commit time would not pin output; (d) hysteresis is dispatch-updates-reference (each dispatch re-records the pose), in-flight nodes are skipped rather than queued (the queued-re-sort slot stays reserved for commits, preserving Phase 2's bounded-failure-retry invariant), and dispatch is additionally gated on effective mesh visibility (own flag AND every ancestor — an LOD level can be a hidden partition GROUP whose member meshes stay visible=true), a surviving `committedData` stamp, and the live blending mode staying order-dependent; (e) `configureDepthSort` takes a live `getCamera` GETTER, not a camera reference — the ortho-mode toggle REPLACES the scene manager's camera object, and a captured init-time reference would freeze sorting at the abandoned perspective pose (latent in Phase 2's commit sorts, consequential once the per-frame scheduler exists; the `lod-group-registry` getter precedent). **Post-Phase-3 hardening campaign MERGED** (2026-07, PRs #568/#569/#571/#572/#575/#577 + review follow-up): null-camera first-dispatch recovery + empty-commit pose hygiene; capacity-clamp consistency (worker count = written texel count — the OOB-ordering fix); worker-center release on LOD demotion; ONE GLOBAL cross-node renderOrder domain (collect-then-assign: wrapper groups by mean member view-z, exact BSP ranks within a wrapper — replaces the per-wrapper 0..N−1 ranks that were incomparable across wrappers/leaves); same-count commits preserve the previous permutation (no 1-frame identity flash while scrubbing); normal-mode picking writes real projected depth (front-most wins) instead of brightness-as-depth. **Follow-ups (2026-07)**: the cross-node renderOrder machinery (BSP part lookup/FKN traversal/rank memo/global assignment) was extracted to `rendering/depth-sort-coordinator/render-order.ts` behind a three-function per-frame API (PR #591); the renderOrder pass is deliberately NOT gated on SortWorker availability — cross-node mesh order survives a never-constructed worker (CSP-blocked script), losing only the worker-dependent within-mesh order (PR #596). **Worker-startup resilience IMPLEMENTED** (branch `sortworker-init-resilience`): the worker is now spawned at APP INIT (`warmUpDepthSortWorker`, gated on `depthSortEnabled`) instead of by the first order-dependent commit, and its init deadline became configurable. Rationale: the first commit lands while the main thread and data-worker pool are saturated decoding, and the worker's Comlink reply must be dispatched on that same thread — measured at 3M points, the worker logged its own readiness (WASM up, answered) while the main thread hit the 30 s deadline, terminated it, and cached the rejection for the session, leaving every order-dependent node (all four geometry types, `normal` ∪ `volumetric`) in storage order until reload. Deltas — (a)-(d) are issue **#1694**'s fix, which this branch builds on rather than contributes (the transient-vs-permanent classification, the bounded starved retry with its backoff and self-wake, the late re-registration sweep, and the init-epoch guard); (e) is this branch's own (app-init warm-up above, the shared init guard with its `messageerror` arm, the configurable deadline, and the monitor's UNAVAILABLE note): (a) the failure taxonomy splits — constructor throw / `onerror` / `onmessageerror` / an `initialize()` REJECTION stay permanent and latch, only the deadline is transient; (b) a transient miss still terminates the worker and keeps the rejected `initPromise` cached — so a commit can never spend an attempt — and only `maybeRetryStarvedWorkerInit` (driven by `evaluateDepthSortPerFrame` past its loader-idle gate, on a growing backoff with a self-wake for the idle-paused render loop) drops that rejection and spawns a fresh worker, bounded at 3 attempts in total, after which it latches; (c) a successful late init re-registers every tracked order-dependent node via `reregisterAfterLateWorkerInit` (the fresh worker holds none of their centers); (d) every module write past an `await` in the init closure is epoch-guarded, because the deadline timer lives in that closure and a dispose cannot cancel it — an orphaned one would otherwise classify a miss against the next app's healthy worker; and the resulting verdict is observable via `getDepthSortWorkerStatus()` (`idle`/`ready`/`starved`/`failed` + the deadline-miss count); (e) the hand-rolled init race was replaced by the data pool's shared `worker-pool/lifecycle/init-with-guard.ts` (gaining the `onmessageerror` arm), the deadline moved to `config.depthSort.workerInitTimeoutMs`, and giving up is now REPORTED — `isDepthSortAvailable()` drives a `depth sort UNAVAILABLE` note in the monitor footer, per issue #705's never-silently-unsorted rule. Risk register §10, changelogs §11. > **Scope**: Viewer-only. No `.gsplats.zarr` / `.luxar.zarr` format change, no Python-side change. > **Goal**: Correct order-dependent transparency (`normal` blending mode) for Gaussian splats at 10M+ splats, via texture-backed splat storage + a per-instance ordering attribute + asynchronous worker depth sorting. > **Non-goals** (follow-ups, not this plan): WebGPU compute-shader sorting ("Option 3b"), weighted-blended OIT, global cross-node sorting. (Points AND Lines sorting symmetry have since LANDED — see §8; all three instanced geometry types share the storage + sorting machinery. **Mesh** has since joined the SORTING half only: same registration, worker and kernel, but its apply permutes `geometry.index` rather than `aSortedIndex`, so every statement about instanced-quad STORAGE in this document remains three-geometry — see `MESH_NODE_SPEC.md` §6.3.) Related reading: `src/rendering/README.md` (the retired `interleaved-attributes.ts` module — whose Float16 narrowing contract this plan superseded — was deleted once the lines migration left it with no consumers). --- ## 1. Problem statement (facts on the ground) - All three geometry types draw as **instanced quads**: `THREE.Mesh` + `InstancedBufferGeometry`, 4-vertex/6-index base quad, per-instance data in one `InstancedInterleavedBuffer` (`rendering/gsplat-geometry.ts`, stride 13 floats = **52 B/splat**: `aCenter`(3) `aCholesky01/23/45`(2+2+2) `aAmplitude`(1) `aColor`(3)). - WebGL2/WebGPU instancing draws instances **strictly in instance-buffer order**. There is no per-instance index buffer. Reordering today means rewriting 52 B × N. - **No sorting exists anywhere** in the render path. `renderer.sortObjects` (default) only sorts whole `Object3D`s. Within a mesh, splat order = ladder/commit order. - `normal` mode (`rendering/blending-state.ts` ~:189-201) is SrcAlpha/OneMinusSrcAlpha alpha-over — **non-commutative**, therefore wrong in arbitrary order. `additive`/`luminous`/`max` are commutative and unaffected. **2026-07 update**: the sorted domain is now the `needsDepthSort` predicate = `normal` ∪ `volumetric` — the emission–absorption mode (VOLUMETRIC_BLENDING_SPEC.md) is the second non-commutative mode and rides this exact infrastructure (same SortWorker, same renderOrder pass, same mode-switch transitions; a `normal`↔`volumetric` switch is a sorted→sorted no-op). - **GSplat `normal`-mode alpha is FIXED (Phase 0, merged)**: under `LUXAR_NORMAL_PREMULT` the fragment emits real premultiplied coverage alpha paired with `getGSplatNormalBlendingState()` (`materials/gsplat/shader-glsl.ts` ~:408-425, `blending-state.ts` ~:238); all other modes keep the `alpha = 1.0` contract. What remains wrong — and what Phases 1-3 fix — is that alpha-over is still applied in **storage order**, not depth order. - Every commit already does a **full CPU repack + full GPU re-upload** of the interleaved buffer (`gpu-buffer-pool/gsplats-adapter.ts::updateGeometry` → `writePooledAttribute` per attribute), plus O(N) bounding-box and max-Cholesky-row-norm loops. Progressive ladder refinement re-runs this for the entire prefix on every level (`data/gsplats/gsplats-progressive-loader.ts` memoized concat → `commit-gsplats-geometry.ts`). - The projection worker (`workers/data-worker/projection/gsplats.ts`) is a **stateless** Comlink RPC over a round-robin pool; inputs are deliberately structured-cloned (transfer would detach SliceCache arrays — `data-processor-gsplats.ts` warning comment), outputs are transferred back. **Plain-3D data never touches this worker**: the gate is `useWebWorkers && splatCount > 1000 && ndim > 3` (`data-processor-gsplats.ts` ~:221), and the ndim==3 fast path is an in-process `.slice()` copy (`projection/gsplats.ts:148-166`). Any sorting design must therefore NOT live inside the projection call — it would run on the main thread for exactly the datasets that need sorting most. - Vertex texture fetch is **already in production**: the colormap LUT is sampled in the gsplat _vertex_ shader (`shader-glsl.ts` ~:311-316, `#ifdef USE_COLORMAP`). - Pick nodes **share the visual mesh's geometry** and pick materials are **already per-node** (`node-factory/create-gsplats-node.ts:68-77`). Pick shaders read the same five instanced attributes (`picking/gsplat/shaders.ts` ~:39-43, `pick.tsl.ts` ~:89-94). Post-commit material sync has an established home: `rendering/material-sync-helpers.ts` (points radiusScale precedent, pushes to render + pick materials via `userData.pickNode`). - THREE r184 blend/attribute facts verified against `node_modules/three`: - Integer instanced attributes work out of the box: a `Uint32Array` attribute binds via `vertexAttribIPointer` because its GL type is `UNSIGNED_INT` (`three.module.js:1950`), matching a GLSL `in uint`; the WebGPU path derives `uint32` vertex format from the array constructor. - `material.premultipliedAlpha = true` is a **trap on the TSL path**: `NodeMaterial.setup()` injects an automatic output-RGB×alpha transform when the flag is set (`three.webgpu.js:21817-21822`). A shader that already premultiplies would be premultiplied twice on TSL but not GLSL — a guaranteed parity break. Do not use the flag. - The known WebGPU→WebGL2-bridge `gl.getError()` issue is specifically **separate alpha-channel blend equation/factor state** (documented in `materials/gsplat/material-tsl.ts:190-212, :300-320`); `CustomBlending` itself is fine — `max` mode ships `CustomBlending + MaxEquation` through the TSL path today. ### Overlapping-node authoring rule Depth sorting is exact only within one node. Cross-node `renderOrder` uses BSP partition order where available and a mean-view-z approximation otherwise, so two spatially overlapping `normal`/`volumetric` nodes can exchange relative order as the camera moves. For nodes with the same geometry type, blend mode, and effective opacity, merge them into one node and pass `partition={"max_elements": N}`: disjoint BSP cells can then be ordered exactly back-to-front while preserving placement and appearance. This authoring path is available to Points, Lines, Mesh, and Gaussian Splats. Put the displayed dimensions first to keep exact BSP ordering. If the geometry types, modes, or opacities differ, merging cannot preserve the authored material; use `additive` only for an emissive medium, or separate the bounds. That escape hatch must not be used against depth-writing geometry. `additive` is the only mode that ignores the depth buffer, so it paints through `opaque` geometry and through a full-opacity depth-writing `normal` Lines or Mesh layer. Use `luminous` for the same additive light-summing with depth testing. Finalization reports both hazards when co-visible world-space node bounds overlap. To avoid noisy warnings from coarse AABBs, the two-sorted-node warning requires one box to contain the other; the additive/depth-writing warning still fires on any positive spatial overlap. Explicitly authored or inherited `additive` is treated as intentional X-ray rendering and suppresses the latter warning. Neither diagnostic rewrites authored modes or defaults. **Bandwidth argument for the architecture**: at 10M splats, re-sorting by buffer rewrite costs 520 MB/sort. With splat data in a texture and only a `Uint32` ordering attribute per instance, a re-sort uploads **40 MB** — 13× less — through the attribute `addUpdateRange` machinery that already exists. --- ## 2. Target architecture ``` ┌────────────────────────────────────────────┐ │ Splat texture (per pooled geometry) │ commit (slice/LOD) │ RGBA32F, width 4096, 4 texels/splat │ ────────────────────▶│ [cx cy cz amp][L00 L10 L11 L20] │ full/partial upload │ [L21 L22 cr cg][cb — — —] (— = reserved) │ └────────────────────────────────────────────┘ ▲ texelFetch ×4 ┌──────────────┐ ordering Uint32Array │ │ SortWorker │ ────────────────────▶ aSortedIndex (per-instance, │ │ (dedicated, │ (transferable) Uint32, DynamicDrawUsage) │ │ stateful) │ │ │ centers cache│ ▼ │ u16 counting │ vertex shader: │ sort (WASM + │ idx = int(aSortedIndex) │ TS twin) │ splat = texelFetch(uSplatTex, addr(idx)) └──────────────┘ (rest of the pipeline unchanged) ▲ │ view matrix (64 B) per sort; centers transferred once per commit │ sort triggers: commit (Phase 2) + per-frame camera-delta scheduler (Phase 3) ``` Two orders coexist by design: - **Storage order** (texture rows) stays the canonical commit order — the additive-ladder concat order. Never permuted. - **Draw order** (`aSortedIndex`) is the view-dependent permutation. Identity for order-independent modes; back-to-front for `normal`. **Single owner of ordering**: the `SortWorker` computes every non-identity ordering. The projection pipeline (worker-pool or in-process) is never involved — see the ndim gate fact in §1. This keeps one code path for 3D and nD data, avoids main-thread sorts, and avoids threading sort state through the stateless round-robin pool. This is the SparkJS decomposition adapted to Luxar's per-node (not global-accumulator) scene graph, nD projection pipeline, and dual GLSL/TSL backend. ### Pinned design decisions | Decision | Choice | Rationale | | ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Splat storage | One `THREE.DataTexture` RGBA32F per pooled geometry, `NearestFilter`, width 4096, 4 texels/splat | texelFetch is core WebGL2; float16 vertex-format pitfalls (see revert `dd7c4478` — the since-deleted `interleaved-attributes.ts` header documented the contract) don't exist for textures; 4-texel padding (64 B vs 52 B) buys addressing simplicity + room for a future per-splat scalar; a later RGBA16F narrowing (32 B/splat) is a data-conversion-only change | | Addressing | 1024 splats/row; `addr(i, t) = ivec2(((i & LUXAR_SPLAT_ROW_MASK) << 2) | t, i >> LUXAR_SPLAT_ROW_SHIFT)` with the constants emitted as **compile-time defines / TSL factory params from the actual texture width** | Power-of-two bit ops, no row-straddle (width % 4 == 0). Defines (not literals) so a `maxTextureSize < 4096` device (WebGL2 floor is 2048) degrades to a narrower texture instead of corrupt addressing | | Max splats/node | `maxTextureSize × splatsPerRow` (4096-wide: 16.7M at height 16384, 8.4M at 8192) | Height, not width, is the scaling axis. Nodes beyond the cap should use `tiles`/`adaptive` recipes (250K/part idiom); warn + clamp; `DataArrayTexture` is the documented escape hatch if low-`maxTextureSize` devices matter in the field | | Ordering | `aSortedIndex` (`Uint32Array` `InstancedBufferAttribute`, `DynamicDrawUsage`), **always present, identity by default** | One shader path, no define flips/recompiles on mode switch; keeps THREE's `_maxInstanceCount` derivation working (an instanced attribute must exist); 4 B/splat is negligible. Uint32→`in uint` binding is native in r184 (§1); the parity harness must cover it on both backends | | Attribute path | **Deleted, not kept alongside** | No dual shader maintenance ("no backwards compatibility burden", "complete before perfect"); vertex buffers drop to 2 (quad + ordering) — well under WebGPU compat-mode's `maxVertexBuffers=8` | | `normal`-mode blend state (gsplats) | `CustomBlending + AddEquation + One / OneMinusSrcAlpha`, **symmetric alpha channel** (no `blendEquationAlpha`/`blendSrcAlpha`/`blendDstAlpha` overrides), `premultipliedAlpha` flag **never set**, `transparent: true`, `depthTest: true`, `depthWrite: false` | One/OneMinusSrcAlpha because the shader premultiplies. Symmetric alpha state avoids the documented WebGPU-bridge `gl.getError()` issue; the `premultipliedAlpha` flag would double-premultiply on TSL (§1). `depthWrite: false` unconditionally — a coverage-alpha fragment (alpha as low as 1e-4 passes the shader's discards) must never write depth or splat footprints punch occlusion halos; sorted transparency never depth-writes | | Material sharing | GSplat visual materials become **per-node** (skip the LRU in `getGSplatMaterial`, still `register()` for camera broadcasts); create with `_layerMaterialCloned: true` | `uSplatTex` is a per-node uniform; the clone-on-divergence dance (colormap path, LayersPanel first-touch clone) collapses into "always own your material". Pick materials are already per-node | | Texture ownership | The **pool entry** owns the texture: `PooledBuffer` gains `splatTexture`, created WITH the geometry at allocation and disposed WITH it (`geometry.dispose()` + `splatTexture.dispose()` in the evictors/`dispose()`); commit rebinds it via a new `syncGSplatMaterialWithGeometry` in `material-sync-helpers.ts` (render + pick material, same shape as the points radiusScale sync) | Texture lifetime = geometry lifetime under the post-#511 contract: **growth is release + reacquire, never an in-place reallocation** — an in-place texture swap on a rendered geometry/material is the same strand class the buffer fix killed (`Info.memoryMap` pins replaced resources; see the grow-leak probe). A grown node simply gets a fresh geometry+texture pair; no carry-forward is needed because every commit rewrites the full count. Unlike `InterleavedBuffer`, `THREE.Texture` IS an EventDispatcher with backend dispose listeners, so pool-side `splatTexture.dispose()` frees GPU memory deterministically | | Sort keys | Camera-space depth, **min/max-normalized to uint16** (65536-bucket counting sort, back-to-front), WASM kernel + mandatory TS twin | O(N) and memory-bound: expect ~100-300M keys/s in a worker → **~1-10 ms at 1M, ~30-100 ms at 10M**. NOT SparkJS's raw-f16-bit-pattern keys: Luxar scenes span physical units from nm to km, so raw camera-space depths can overflow f16 (>65504 → every splat keys to Inf) or underflow to 0 — either collapses the ordering entirely. Per-sort min/max normalization is scale-invariant (antimatter15-splat-style), same O(N), and 65536 bins across a node's depth extent is ample for blending order. Guard zmax==zmin → identity | | Sort execution | One dedicated persistent `SortWorker` (module-scoped state, own Comlink endpoint), NOT the round-robin data-worker pool, and NOT inside projection | Sorting needs per-node retained centers; round-robin would scatter state, and projection doesn't even run in a worker for ndim==3 (§1). `SharedArrayBuffer` rejected: COOP/COEP would ripple into `luxar serve`, the export bundle's `serve.py`, and the Go launchers | | Sort gating | Only nodes whose **effective composed** blending mode is order-dependent (`normal`) get registration/sorts; everything else keeps identity ordering | Zero cost for the default (`additive`) path. (`opaque` at fractional opacity is also order-dependent but degenerate — explicitly out of scope, documented in the mode docs) | ### 2.1 Sorting execution model **CPU-in-worker, not GPU.** On the production WebGL2 backend a GPU sort means fragment-shader bitonic ping-pong: O(n log²n) fullscreen passes (~300 passes at 10M) with no way to write a vertex-consumable buffer from a shader — strictly worse than an O(n) CPU counting sort in a worker. SparkJS's own hybrid (GPU distance compute → RGBA8 readback → CPU radix) exists because its splats are GPU-_generated_ (procedural/skinned) and live only in textures; Luxar's projection pipeline hands us CPU-side centers for free, so computing keys in the worker skips the readback path entirely. The GPU answer is Option 3b (WebGPU compute radix), which replaces the worker, not the interface. Revisit only if projection itself ever moves GPU-side. **Full snapshot sorts, not incremental/progressive maintenance.** Counting sort is already O(n) with memory-bound constants — frame-coherence tricks (nearly-sorted adaptive sorts, "resort only what moved") cannot beat O(n) and add worst-case cliffs; this is why every serious 3DGS implementation uses radix/counting. The "progressive" axis in this design is **staleness, not partial computation**: render with the last complete order while the next one computes (bounded by the ~3° trigger, so the stale order is never more than ~one threshold wrong). Time-distribution happens **across nodes**, not within a sort: the per-frame scheduler dispatches every past-threshold node in the same tick and the single SortWorker serializes them (a one-per-frame coverage×staleness priority queue was designed but never implemented — measured 2026-07: MatrixCity frame tails are load-driven, not dispatch-burst-driven, so the budget was dropped at the evidence gate); `tiles`/`adaptive` partitions cap per-sort N at ~250K (~1-3 ms each), making per-node sorts naturally bite-sized. Coarse-first bucket refinement (approximate order now, exact later) is rejected: complexity for artifacts that async full sorts already avoid. **A permutation must swap atomically — that's where the double buffering lives.** A partially-applied ordering is not a reordering, it is a _corrupt permutation_ (splats drawn twice/zero times → flicker), so partial visual application is never acceptable. Three tiers: 1. **Main-thread apply** needs no explicit double buffer: JS single-threadedness means `aSortedIndex.set(ordering)` + `needsUpdate` completes between frames; the draw only ever sees a complete order. Driver-side buffer renaming under `DynamicDrawUsage` is the implicit GPU double buffer. 2. **Worker boundary**: ping-pong a recycled pair of transferables — main thread returns the previous ordering buffer with the next sort request, worker fills and transfers back. Zero steady-state allocation (SparkJS's persistent-buffer pattern across `postMessage`). 3. **Escape hatch for 10M-scale upload stalls** (40 MB ≈ 5-8 ms on the GL thread): an A/B pair of ordering attributes with a uniform selector — stream the new order into the inactive attribute across frames via partial `addUpdateRange`, then flip the uniform (atomic swap, amortized upload, +4 B/splat). Ship without it; add only if the Phase-3 perf gate shows frame drops attributable to ordering uploads. **RESOLVED (2026-07, perf lever L8):** the perf campaign measured the stall far above the estimate — sort-adjacent frame p99 of 119–563 ms at 10M splats (Mac M4 Max, synthetic orbit; the full-buffer memcpy dominates, not just the upload) — and it shipped as **chunked apply** instead of the A/B pair (no shader change, no +4 B/splat): orderings > 1M indices stream one 1M-index (4 MB) slice per rendered frame through per-slice `addUpdateRange`, driven by the Phase-3 per-frame scheduler; while a stream runs the node is dispatch-gated (no new sorts) and a newer ordering is held (latest wins) until the stream completes — restarting mid-stream under a continuous orbit never converged (`element-storage.ts` chunked-apply note has the full mechanics + the documented transient old/new-mix trade the A/B design would have avoided — judged acceptable: bounded ≤ ceil(n/1M) frames of duplicate/omit shimmer, same visual class as stale-order frames). Classic WebGL only; the WebGPU backends ignore attribute ranges and keep the single-shot path. **SUPERSEDED (2026-07-30) — tier 3 SHIPPED.** The chunked apply's "bounded transient mix" premise did not survive measurement: under a CONTINUOUS orbit a new sort arrives about as fast as a stream drains, so the mix is the steady state, not a transient. Measured on main with a browser probe that decodes the element texture and validates the drawn index buffer every sampled frame: **27–33% of frames** on the 1.9M-splat volumetric `visible_human_head` (27,613 elements double-drawn) and **70–80% of frames** on the 8M-point `global_rivers_earth` terrain (up to 1,022,162 double-drawn = 12.8% of the node), duplicates always beginning exactly at the chunk boundary; an A/B that raised the threshold above the node size gave a valid permutation in every frame. So the A/B pair shipped as specced: `aSortedIndex`/`aSortedIndexB` + a runtime `uSortedIndexSlot` selector, slices streaming into the INACTIVE buffer and the slot flipping on completion. BOTH buffers are allocated at attach, so the +4 B/element is paid by every node, sorting or not. The pair was FIRST aliased onto one buffer and split on a node's first sort (free for commutative-mode nodes) — **that was a native-WebGPU black-screen bug**, caught only by an explicit `?renderer=webgpu` A/B against main and worth recording as a hazard class: three keys a render pipeline's vertex-buffer LAYOUT by `BufferAttribute` identity (`WebGPUAttributeUtils.createShaderVertexBuffers`), but `RenderObject.getGeometryCacheKey` hashes only attribute NAMES/itemSize/normalized, and `RenderObjects.get` answers `needsGeometryUpdate` with a bare `setGeometry()` that refreshes the attribute list while leaving `Pipelines`' cached pipeline alone. So splitting the alias after first render bound three vertex buffers into a two-buffer layout: every later attribute shifted down a slot and `aQuadCorner` read the ordering buffer's `u32`s as `vec2`, collapsing every quad — a black scene, no validation error, no console warning. Classic WebGL binds attributes by program location and is structurally immune; so is `tsl-shader-parity`, which runs `WebGPURenderer({ forceWebGL: true })` — a TSL graph through the WebGL2 bridge, which never builds a WebGPU vertex layout at all. **Rule: for any geometry a WebGPU pipeline has drawn, the attribute SET and each attribute's object identity are immutable.** Guarded by a unit invariant test (`splat-texture-storage.test.ts`, "keeps the ordering attribute OBJECTS fixed for the geometry lifetime"), since native-WebGPU E2E is not runnable in CI. Slicing stays classic-WebGL-gated (WebGPU writes one slice then flips); double-buffering is unconditional. The dispatch apply-gate is GONE — with the drawn buffer always whole there is nothing to protect. Cost, measured on the 10M orbit bench (M4 Max, 3 runs/arm): sort-adjacent frame p99 ~77 ms → ~92 ms with 12 → 24 sorts; median and p95 unchanged; still far under the 119–563 ms the chunking exists to prevent. A throttle on "an ordering is already queued" was built and rejected at the evidence gate — it cut sorts to ~15 but moved p99 only to ~89 ms (inside noise) while costing 41% more staleness (fitted sort-axis lag on the 8M fast orbit: mean 44° → 63°). Waiting for a WHOLE ordering instead of showing a partly-applied one does cost freshness: same 8M fast orbit, lag mean 36.5° before vs 44.3° after — the deliberate price of never drawing a corrupt permutation. **Stability as a feature**: the counting sort is stable, so equal-key splats (same 1/65536 depth bin) retain storage order — which is the additive-ladder energy order. Ties resolve toward "most important splat wins the blend order", for free. --- ## 3. Phase 0 — Correct alpha for `normal` mode (prerequisite) Without this, sorted or not, gsplat `normal` mode cannot reveal the framebuffer behind a splat. Independent, small, immediately shippable. **Changes** 1. **Fragment shaders** (`materials/gsplat/shader-glsl.ts` + `shader-tsl.ts`): under a new `LUXAR_NORMAL_PREMULT` define (GLSL) / factory flag (TSL), emit premultiplied coverage alpha: ```glsl float a = clamp(intensity * uOpacity, 0.0, 1.0); // coverage, clamped for OneMinusSrcAlpha sanity fragColor = vec4(gammaColor * intensity * uOpacity, a); // RGB stays unclamped HDR ``` Default (no define) output is byte-identical to today — additive/luminous/max keep the exact `vec4(rgb, 1.0)` contract. (As shipped, additive blending is unified onto the shared blending-state helpers on both backends — the GLSL alpha-MaxEquation overflow guard this section originally cited was deleted during the same campaign, with the `raw-scene-hdr` capture path sanitizing alpha at readback instead; see §10.) Note the deliberate HDR asymmetry: RGB carries unclamped `intensity × uOpacity`, alpha carries the clamped coverage — an emitter-with-occlusion model, correct for `1-α` destination attenuation. 2. **Blend state**: add `getGSplatNormalBlendingState()` to `rendering/blending-state.ts` returning the pinned state from §2 (CustomBlending, One/OneMinusSrcAlpha, symmetric alpha, transparent **true**, depthWrite **false**). No `opacity` parameter — unlike the generic `normal` entry, depthWrite is unconditionally false here, so the state is opacity-independent. Do NOT widen `getCompleteBlendingState` — the gsplat GLSL wrapper is already the documented divergence point, and the generic `normal` entry (with its `opacity >= 0.99` depthWrite flip) stays correct for lines' straight-alpha output (points now force depthWrite off as well — see `getPointBlendingState`, #1002). Consume it from `applyBlendingMode` in **both** `material-glsl.ts` (which hand-manages custom factors today) and `material-tsl.ts` — whose comment claiming "TSL needs no gsplat-specific override" becomes stale and must be updated, since with real alpha output `NormalBlending`'s SrcAlpha factor would now double-multiply. LayersPanel's `applyBlendingStateToMaterial` fallback is unaffected (it already prefers the material's own `applyBlendingMode`). 3. **Depth trade-off, documented**: with `depthWrite: false`, a gsplat `normal` layer no longer occludes additive layers behind it (the generic normal mode's `opacity >= 0.99` rationale). Accepted: coverage-alpha splats writing depth is strictly worse (halo artifacts from near-transparent fragments that pass the `1e-4` discards). 4. **Docs**: update the `BlendingMode` docs in `material-manager/factories.ts:35-45` and `materials/gsplat/README.md` ("alpha-always-1 contract" becomes "alpha = 1.0 except `normal`"). **Tests / exit criteria** - TSL/GLSL parity spec covers the new define/flag combination (guards the double-premult trap by construction). - Unit: `getGSplatNormalBlendingState` state shape; mode-switch round-trip normal↔additive restores the exact prior state. - Visual E2E fixture: two overlapping splats, `normal` mode, opacity 0.5 — background must show through (fails today). Add to `tests/fixtures/generate_test_data.py` (single source of truth; unit-test globalSetup auto-generates). - Existing additive/max visual baselines unchanged (regression gate). Run the WebGPU diagnostic path (`?renderer=webgpu` and `&webgpuForceWebgl`) on the fixture — the bridge's alpha-state sensitivity is exactly what this phase touches. **Risk**: low. Purely additive define + one new blend-state helper. `normal` blending will still be order-_dependent_ after this phase — visibly better, occasionally wrong on overlap, fully fixed by Phases 1–3. --- ## 4. Phase 1 — Texture-backed splat storage + always-on ordering attribute The structural phase. After it, rendering is visually identical (identity ordering), but splat data lives in a texture and draw order is decoupled from storage order. **Changes, by layer** 1. **Pool adapter** (`rendering/gpu-buffer-pool/gsplats-adapter.ts`): - `createGSplatsGeometry(capacity)` builds: quad + index (unchanged), one `aSortedIndex` `InstancedBufferAttribute` (Uint32Array, `DynamicDrawUsage`), and a pooled `DataTexture` (RGBA32F, `width × ceil(capacity/splatsPerRow)`, `NearestFilter`, `generateMipmaps=false`, `flipY=false`), width = `min(4096, caps.maxTextureSize)` with the addressing defines derived from it. - `PooledBuffer` (in `pool-stats.ts`) gains `splatTexture: THREE.DataTexture`. **Growth follows the post-#511 contract: release + reacquire** — `growGSplatsGeometry`/`rebuildInterleavedBuffer` no longer exist, and Phase 1 must NOT reintroduce an in-place texture reallocation (same strand class the buffer fix killed). An undersized acquire releases the geometry+texture pair to the pool intact and falls through to best-fit/fresh allocation; no carry-forward (every commit rewrites the full count). The evictors and `pool.dispose()` dispose `splatTexture` alongside the geometry. The existing `_lastAcquireRebuilt` → `invalidateRenderObjectFor(mesh)` contract covers the geometry swap — and since the flagged-issues pass, `invalidateRenderObjectFor` also eagerly re-points `userData.pickNode` at the new geometry, which Phase 1 inherits for free; the texture swap is covered by the material rebind (below). - `updateGeometry(...)` replaces the six `writePooledAttribute` calls with **one fused pass** writing the staged arrays (`centers3D`, the 6-stride `choleskyFactors`, `amplitudes`, `colors`) into the texture's backing `Float32Array` in texel layout, then `texture.needsUpdate = true`; fills `aSortedIndex[0..count) = identity` (`addUpdateRange`); keeps `instanceCount`, the `_maxInstanceCount` delete, and the bbox + max-row-norm handling exactly as today (since perf lever L1, the projection's fused-scan `bounds` metadata supplies both when present; the loops remain as the fallback). Fewer passes over the data than the current per-attribute writes — this phase should be **neutral-to-faster** on commit CPU. - Non-pool fallback (`gsplat-geometry.ts::createInstancedGSplatsMesh` / `updateInstancedGSplatsMesh`) gets the same treatment with a mesh-owned texture. The fallback's rebuilt branch now swaps in a FRESH geometry and disposes the old one (post-#511); the texture rides the same swap — build fresh, dispose old, never rebind in place. (The 2026-07-12 broken-fallback finding was fixed in PR #503.) 2. **Byte accounting**: `gpu-buffer-pool/geometry-bytes.ts::estimateGeometryBytes` counts `capacity × 64 B` texture + `capacity × 4 B` ordering — **68 B/splat vs 52 B today, ≈ +31% VRAM** (state this in the shared `gpu-byte-budget.ts` docs; the future RGBA16F narrowing brings it to 36 B, −31%). Byte-budget eviction and `dispose()` free the texture with the geometry. 3. **Shaders — all four gsplat stacks in one PR** (visual GLSL/TSL, picking GLSL/TSL): - Delete the five `in`/`attribute()` per-instance declarations; add `uniform highp sampler2D uSplatTex;` + `in uint aSortedIndex;` (TSL: `attribute('aSortedIndex', 'uint')`, `textureLoad`). - Prologue: `int idx = int(aSortedIndex);` then four `texelFetch(uSplatTex, addr(idx, t), 0)` reads reconstructing `center/cholesky/amplitude/color` into the exact local names the existing math uses — **zero changes downstream** of the prologue (covariance projection, Jacobian, eigen-expansion, the unified `perspectiveNearFade`, unconditional coverage fade, `screenSize` un-flip, and fragment kernel all untouched). TSL house rule applies: the vertex is Fn-traced — the prologue is `.toVar()` STATEMENTS inside the existing `Fn(() => {...})`, never free expression trees (the uninitialized-eigen bug class). - `unpackCholesky3D()` takes the fetched values as parameters instead of reading attributes; addressing constants come from the width-derived defines / TSL params. 4. **Materials** (`materials/gsplat/material-glsl.ts` + `material-tsl.ts`): add `uSplatTex` uniform + `updateSplatTexture(tex)` (modelled on `updateColormapTexture`). `material-manager.ts::getGSplatMaterial` bypasses the LRU (per-node materials), still registers for `updateCameraParams` broadcast; `create-gsplats-node.ts` sets `_layerMaterialCloned: true` and drops the colormap clone dance. **Verify per-node material disposal on node teardown** — the existing clone-lifecycle path (`material-manager/lifecycle.ts` subscribeToDispose) should already cover it; confirm with a leak assertion in the pool tests. 5. **Commit path**: add `syncGSplatMaterialWithGeometry(mesh)` to `rendering/material-sync-helpers.ts` (points-radiusScale precedent): rebind `uSplatTex` from the acquired pool entry on the visual material and, via `userData.pickNode`, the pick material. Called from `commit-gsplats-geometry.ts` after acquire. The noop stamp-only path is untouched. 6. **Cleanup**: `buildGSplatAttributeSpecs`/`bindInterleavedAttributes` in `gsplat-geometry.ts` are deleted or absorbed (`packCholeskyForShader` was later retired outright by perf lever L1 — the 6-stride `choleskyFactors3D` now flows from projection to the texel writer directly); `GSPLATS_ATTRIBUTE_SPECS` in the adapter is replaced by the texel-layout writer. Points/Lines interleaved paths are untouched (symmetry is restored when/if they migrate — see §8). **Tests / exit criteria** - Unit: texel address math incl. non-4096 widths; fused writer round-trip (write staged arrays → read back texel layout); growth = release + reacquire returns a FRESH geometry+texture pair with the old pair pooled intact (NO carry-forward — the forbidden mechanism); byte accounting; material-disposal leak check. - Parity: `tsl-shader-parity.spec.ts` green on the texture-fetch shaders **and** the uint attribute on both backends. - E2E: full gsplat suite (`geometry-types`, `blending-modes`, `real-dataset-loading`, picking specs) green with **unchanged visual baselines** — Phase 1 must be pixel-identical (f32 attribute reads vs f32 texelFetch are bit-exact). - Perf gate: `Update Buffers` profiler stage not regressed on the COMMITTED timelapse-nav harness (`scripts/perf/timelapse-nav-bench.mjs` — before/after, both renderers; wall-clock is noise-dominated, use frame percentiles + `info.memory`); VRAM within the +31% envelope on demo scenes. - **Texture-lifecycle leak gate**: `scripts/perf/grow-leak-probe.mjs` on native WebGPU (`--chrome`; see scripts/perf/README.md for the three adapter traps) — after forced grows + LRU evict, `info.memory` (textures AND attributes) returns to the clean floor. This is the probe that proved the buffer-strand fix; Phase 1 must pass it for textures. - Codegen snapshots (`LUXAR_UPDATE_SNAPSHOTS=1`, FRESH server — the spec false-passes against a stale :5173) regenerate for all gsplat surfaces; the full parity-variant inventory stays green: base/gamma-one/colormap, `gsplat-offcenter` (fragcoord convention), `gsplat-tiny-sigma-fade`/`-reject` (coverage fade), behind-camera trio, `gsplat-normal-premult`. - Context-loss E2E (`webgl-errors.spec.ts`): textures rebuilt on restore (recommit path). **Risks**: texture-unit pressure (uSplatTex + uColormapTex in the vertex stage = 2 of ≥16 guaranteed — fine); `maxTextureSize=4096`-class devices cap at 4.2M splats/node (warn + clamp; tiles idiom is 250K/part); WebGPU `textureLoad` parity (covered by the harness + `?renderer=webgpu&webgpuForceWebgl` diagnostic). --- ## 5. Phase 2 — SortWorker + depth-ordered commits Introduces the sorting engine and wires it at commit time. After this phase, `normal` mode is correct within a frame or two of any load/slice-change and only goes stale while orbiting (fixed in Phase 3). Ongoing and camera-driven sorting lives **exclusively** in the SortWorker, never in the projection call (see §1: plain-3D projection doesn't run in a worker at all). **Implementation delta (Phase 2, 2026-08)**: the first ordering after an eligible instanced commit may also be computed on the main thread within the shared `syncSortMaxElements` frame budget. Indexed Mesh is excluded because its ordering is applied through `geometry.index`; async registration and all subsequent sorting are unchanged. **Changes** 1. **WASM kernel + TS twin** (the one new compute primitive): - `wasm/rust/src/depth_sort.rs`: `sort_splats_by_depth(centers3: &[f32], model_view: &[f32; 16], ordering: &mut [u32], count) -> u32` — pass 1: camera-space `z` per splat into an f32 scratch + running min/max; pass 2: normalize to uint16 keys (`floor((z - zmin) / (zmax - zmin) * 65535)`, zmax==zmin → identity) + 65536-bucket histogram; prefix-sum reversed for back-to-front; pass 3: stable scatter. Behind-camera splats key to the far bucket (the vertex shader already degenerates them). Raw f16-bit-pattern keys (SparkJS) are deliberately NOT used — Luxar's physical-unit range (nm..km) overflows/underflows f16 and collapses the ordering; normalized integer keys are scale-invariant. Export from `lib.rs`; no `validate_ndim` involvement (input is always projected 3D centers). - `wasm/typescript/depth-sort.ts`: 1:1 reference implementation (repo parity contract), plus the parity test pairing. 2. **`SortWorker`** (`workers/sort-worker.ts` + `workers/sort-worker/`): dedicated persistent Comlink worker (module-scoped state, same idiom as the data-worker's `WasmCtx`; `initialize(wasmPath?)` forwards the embedder's WASM path exactly like `data-worker.ts:53`). API: - `registerNode(nodeId, generation, centers3 /* transferred */, count)` — called after commit for order-dependent nodes. **`generation` is a per-node monotonic counter incremented on every non-noop commit** (stored on the mesh userData next to `committedData`) — NOT `loadedViewVersion`: ladder-refinement appends share a view version while `count` grows, and an ordering must only ever apply to the exact commit it was computed for (a shorter stale permutation applied to a grown buffer is corrupt, not just stale). Noop stamp-only commits leave the generation unchanged, so an in-flight sort stays valid across them. The staged `processed.centers3D` is **transferred, not cloned**: the projection output is always a fresh buffer (worker path transfers ownership in; the ndim==3 fast path is a `.slice()` copy — `projection/gsplats.ts:148-150`), the commit's texture-write/bbox loops are the last main-thread readers, and the memoized-concat noop identity keys on `sourceData` (pre-projection), never on `processed.*`. Zero-copy registration. - `sort(nodeId, generation, modelView) → { generation, ordering /* transferred */ }`. - `releaseNode(nodeId)` — wired to node disposal/pool release; `dispose()` terminates the worker on app teardown. 3. **Commit wiring** (`data-processor-gsplats.ts` / `commit-gsplats-geometry.ts`): when the node's blending mode is order-dependent — read from the material's **live** `userData.blendingMode`, the value that actually drives the blend state regardless of who set it (creation attrs or a later `applyBlendingMode` from the LayersPanel compose chain) — the commit (a) writes identity ordering as in Phase 1, (b) registers/refreshes the node with the SortWorker, and (c) requests one sort with the current `camera.matrixWorldInverse × mesh.matrixWorld`. On resolve: **generation-check (drop stale)**, `aSortedIndex.set(ordering)` + `addUpdateRange(0, count)` + `needsUpdate` + request a frame. At most one in-flight sort per node from day one. 4. **Mode-switch hook**: switching a gsplat layer TO `normal` cannot simply "register+sort" — the SortWorker has no centers for a node that was order-independent at its last commit (registration is gated, and the staged arrays were transferred/discarded). Instead the switch **invalidates the node's noop stamp (`userData.committedData = undefined`) and requests a view update**: the memoized-concat noop path would otherwise skip re-projection entirely, so clearing the stamp forces a standard reprocess→commit, which registers + sorts like any commit (SliceCache still holds the source nD data — this is one O(N) re-projection per mode switch, a rare user action). Switching AWAY just stops future sorts; a sorted order is harmless under commutative modes — no identity reset needed. **Tests / exit criteria** - Rust unit tests (`make test-wasm`) + TS twin parity: ordering is a permutation, back-to-front monotone, stable, behind-camera splats last; **unit-extremes cases** (nm-scale ~1e-6 and km-scale ~1e6 coordinate magnitudes, plus zmax==zmin) produce valid orderings — the exact failure mode raw-f16 keys would have; `benchmark-wasm` entry with an absolute throughput floor. **Implementation delta (Phase 2, measured 2026-07)**: the original ≥100M splats/s target was a pre-implementation estimate; the release WASM kernel measures ~74M splats/s at 1M **worst-case spatially-incoherent** splats on an M-class dev machine (~96-114M/s native Rust — the gap is WASM execution overhead + wasm-bindgen boundary copies; real Morton-coherent gsplat data sorts faster). Two faster-looking variants (2×8-bit LSD radix; branchless multi-lane min/max + interleaved partial histograms) were measured and REJECTED — both slower than the simple three-pass counting sort (rationale in `depth_sort.rs`). The floor (`perf-budget.test.ts`, opt-in `pnpm test:perf`) is **≥50M splats/s**, report-only unless `LUXAR_PERF_QUIET_HOST=1` explicitly enables quiet-host enforcement — below the worst-case measurement, far above any real regression (debug build, comparison sort, alloc-per-element). Budget impact: a worst-case 10M-splat sort is ~135-190 ms in the background SortWorker — above §2's 30-100 ms target but off the frame path; acceptable staleness (Phase 3 sorts are throttled and generation-guarded). Follow-up headroom if ever needed: WASM `simd128`, or a stateful register/sort WASM API keeping centers wasm-resident across camera-orbit re-sorts (eliminates the 12MB/sort copy-in). - Unit (mock worker): generation-guard drops stale results; single-in-flight; release-on-dispose; transfer-detach assertions (registered buffers unusable on main thread afterwards — intentional). - E2E: the Phase-0 overlap fixture renders **correctly** in `normal` mode after load settle, from the load-time camera; additive baselines still unchanged (identity path). **Risk**: low-moderate — new worker lifecycle, but the generation guard and single-in-flight rule are small state machines with unit coverage. --- ## 6. Phase 3 — Camera-triggered re-sort The live feature: sorted order tracks the camera. Pure scheduling on top of Phase 2's engine. **Changes** 1. **Scheduler**: new per-frame callback `'depth-sort-scheduler'` registered in `core/app/init/pipeline.ts` beside `'lod-group-selector'` (which already owns the per-frame camera scratch pattern — same no-per-frame-allocation invariant). Per registered order-dependent node: track view direction + position at last completed sort; when `dot(viewDir, lastViewDir) < cos(θ)` or translation exceeds a bounds-relative fraction, dispatch a sort (Phase 2's single-in-flight + generation guard already handle overlap/staleness). Frames between dispatch and resolve render the previous order — bounded staleness, no popping-to-wrong, standard 3DGS behavior. Skip dispatch while the node's loader has a pending view update (same signal the refinement loop consults). 2. **Config**: new `config/sections/depth-sort/` (`data/types/validate/README`, following the `adaptive-dpr` section shape): `enabled` (default true), `angleThresholdDeg` (default ≈3°), `translationFraction` (default 0.05). URL escape hatch `?depthSort=0` in `config/url-params.ts` (pattern: `?gpuBudgetMB=`) — pins identity ordering for deterministic E2E/visual runs. 3. **Monitor/profiling**: sort latency + bytes as an update-profiler stage and a data-loading-monitor line, per the monitor design language (semantic color only, tabular numerals). **Tests / exit criteria** - Unit (mock worker): threshold hysteresis; no dispatch during pending loads; `?depthSort=0` inert. - E2E: orbit around the overlap fixture in `normal` mode with settle-based assertions — final settled frame matches the sorted baseline from every tested angle; `?depthSort=0` reproduces Phase-2 behavior (determinism guard). - Perf: at 1M sorted splats, orbiting holds target FPS with sort round-trip < 50 ms and zero main-thread stalls > 4 ms attributable to sorting (ordering upload is the only main-thread GPU cost: 4 MB @ 1M). **Risks**: worker memory (retained centers = 12 B/splat; 120 MB at 10M — acceptable against the 640 MB texture, `releaseNode` bounds it, and f16 center retention is a documented follow-up halving); visible re-sort "settling" during fast orbits (tune θ; it replaces being _always wrong_ today). --- ## 7. Phase 4 (optional, perf) — Partial texture uploads Not required for correctness — a measured-win phase banking the storage refactor's dividend. **Key implementation finding (2026-07):** three r184 already ships texture-side partial uploads — `Texture.updateRanges` / `addUpdateRange(startFloat, countFloat)`, honored by the classic `WebGLRenderer` (`webgl/WebGLTextures.js::updateTexture`: it takes the whole-image `texSubImage2D` path ONLY when `updateRanges` is empty, else uploads per-range). Since luxar's **default backend is classic WebGL**, no `copyTextureToTexture` staging path is needed for the win; the WebGPU backends ignore the ranges and re-upload the whole image (correct, just not yet partial). Three sub-stages: ### Stage 1 — slack elimination (LANDED 2026-07) Every non-noop commit re-uploaded the **entire capacity-sized** RGBA32F texture, including the pool's 1.5× growth headroom and best-fit slack rows. `writeSplatTexels` now registers per-row `updateRanges` over `[0, count)` via the dirty-range helper (today `registerElementTexelDirtyRange` in `element-storage.ts`), so only the live rows upload. Row-split because the WebGL path uploads each range with `height = 1`; splat×4-texel alignment on a width-multiple-of-4 texture guarantees no row straddle; ranges self-collapse into one contiguous span per call (the WebGPU backends never clear them). Above `FULL_UPLOAD_ROW_FRACTION` (0.75) of rows dirty, it falls back to the single full-image upload (per-row call overhead outweighs the saving). **Measured** (`gsplats_4d_neuromast_2ch`, classic WebGL): initial-load upload 3.31 MB → 2.20 MB (33%, the fresh-alloc headroom); per-commit slack-scrub case 5.81 MB → 2.36 MB (59%). Pixel-identical (texel content unchanged; shaders only read `[0, count)`). ### Stage 2 — append-only writes for ladder streaming (LANDED 2026-07, all three geometries) Progressive ladder refinement re-uploaded the whole committed prefix per level. An append-only commit writes only the **new** splat span: `writeSplatTexels(…, { fromSplat })` skips the prefix write loop and registers just `[fromSplat, count)` (the `registerSplatTexelDirtyRange` general span already supported this from Stage 1). **Design finding (revises the earlier plan): no projection-kernel contract change is needed.** The nD→3D projection is per-splat-independent and order-preserving, and the loader appends LOD levels in order, so a k-level input concat is a byte-identical *prefix* of the (k+1)-level concat. Under an unchanged view state the (k+1)-level projection's first `prevCount` visible outputs are therefore byte-identical to the previous commit's entire output — so `outputPrefixCount === prevCount` (the committed `visibleSplatCount`) is **deducible**, not measured from a returned visibility mask. The trust chain is: (1) a forward-chained prefix-lineage `WeakMap` (`types/prefix-lineage.ts`), stamped in the loader's `concatenateMemoized` and checked in commit by identity against `getCommittedData(mesh)` — a match proves same generation (a view change bumps the loader reset generation → no parent → append rejected → full rewrite), genuine extension, and that the GPU still holds that parent's projection; (2) the append gate also requires `!attributesRebuilt && geometry === prevGeometry` (pool reused this node's buffers in place — which also bounds `count ≤ capacity`), `gpuPrefixIntact`, `splatCount > prevCount`, and `committedTruncate === readTruncate(mesh)` (the one hole: `uTruncate` is a material uniform outside `viewStatesEqual`). The splat texture remains suffix-only, but `aSortedIndex` resets to full identity: preserving an old depth-sorted prefix and appending a storage-order suffix makes alpha-over draw two independently ordered populations until the commit-triggered sort lands. The bbox recompute stays full (CPU-only, no GPU upload). A **context-restore full-dirty hook** (`NodeFactory.rebuildAfterContextRestore`) marks every splat texture + `aSortedIndex` full-dirty and clears `gpuPrefixIntact`, because a WebGL context loss zeroes the GPU buffers while the CPU mirror survives — without the flag, a following append would downgrade the pending full upload to a stale-prefix partial. The GSplat identity reset is a **perceptual fallback, not a settled correctness fix for #1987**: the reported persistent gastrulation ghost still lacks a visual repro. The change treats the hybrid prefix/suffix ordering in dense translucent splats as more likely to read as a detached second population, so the interim draw now gives up the useful sorted prefix and renders the whole grown node in storage order until the fresh permutation lands. The cost is one SortWorker round-trip plus, on the classic WebGL backend, `ceil(count / SORTED_INDEX_CHUNK_ELEMENTS)` drawn apply frames per append — about 1–2 frames for the roughly 200k-splat gastrulation node, but several frames for multi-million-splat nodes. Holding `instanceCount` at `prevCount` was rejected: the coordinator would have to raise it when an ordering lands, but no ordering is guaranteed under commutative blending, `?depthSort=0`, or an unavailable SortWorker, so the node could stall indefinitely. Points and Lines retain the hybrid because they have no corresponding observed artifact and preserving their valid sorted prefix minimizes interim disorder; their tests explicitly pin that choice. **Three-geometry symmetric (LANDED 2026-07, Points/Lines):** the same append is `writePointTexels(…, { fromPoint })` for Points (since the points texture migration, §8 / PR #630) and `writeLineTexels(…, { fromSegment })` for Lines (segment units; originally landed as `writeInterleavedAttribute(…, { fromInstance })` and carried over unchanged by the lines texture migration, §8 / PR-C), both routed through the pool facade's uniform `fromInstance` option (their ranged uploads already eliminated slack, so this was their remaining O(k·N)→O(N) win). Order-preservation verified: the Points zero-radius drop is a per-level order-preserving forward scan run *before* concat, and Lines clipping is an order-preserving drop + in-place endpoint clip (never splits 1→2, never reorders) — so prefix identity holds for both under the same lineage trust chain (the shared helper now lives in `types/prefix-lineage.ts`). Their gates mirror the gsplat one minus `preserveOrdering`/`committedTruncate` (no permutation, no external projection input), plus one extra conjunct: **optional-field presence must match the committed parent** — the Points concat is all-or-nothing per field (`concatOptionalField`), so a Float32↔absent flip re-fills the column with a constant that need not match the committed prefix; same for Lines colors/sharpness whose presence flip re-fills the prefix through the interpolation kernel. A prerequisite fix rode along: `concatenateLinesData` now fills missing-part **sharpness with 0.5** (the projection's null-sharpness default), not 0.0 — a mixed-sharpness ladder previously popped razor-sharp on the sharpness-less parts even on full rewrites. The context-restore hook covers all three types through one shared branch (element textures + `aSortedIndex` marked full-dirty + `gpuPrefixIntact` cleared — lines' historical interleaved-buffer branch died with their texture migration). The shared lineage WeakMap is retention-capped: `setPrefixParent` deletes the superseded parent's own entry (chain depth 1) and every commit consumes-and-clears the entry after its gate check, so intermediate concat results of a ladder are never pinned via lineage (a WeakMap holds values strongly while the key lives — an uncapped chain retained ~(n−1)/2 × the final CPU arrays on an n-level ladder). ### Stage 3 — WebGPU partial-upload parity ([POST] — MEASURED-REJECT 2026-07-25, revisit criteria below) Built and measured per the measure-first doctrine; **rejected at the ≥10% in-app bar** (archived as a closed PR with the implementation + all numbers). What the campaign established: - **Raw mechanism win is real and large.** three r184's WebGPU backend ignores texture `updateRanges` (whole-image `queue.writeTexture` per `version` bump; only attribute ranges and `DataArrayTexture.layerUpdates` are consumed). Probed on Dawn/Metal (M4 Max headless Chrome): `writeTexture` ≈ 2.2 GB/s, so a 5M-splat element texture (305 MB) costs ~129 ms per full re-upload vs ~7 ms for a ladder-append span (18×) and ~56 ms for a 40%-rows scrub span; per-ROW ranged calls beat even a single contiguous block. The working implementation (an instance-level wrapper on `WebGPUBackend.updateTexture` consuming the existing per-row ranges via ranged `writeTexture` + consume-and-clear, knee = ∞) lives in the closed PR. - **In-app, no available workload clears the bar.** Timelapse-nav (neuromast, ~8 MB textures): 0% — uploads hide inside the frame. Synthetic 5M full-dirty recommits: −1.6% (100% dirty moves the same bytes either way). Append-shaped ladder loads (visible-human 1.9M single ladder — 4096×2800 capacity-padded texture, 183 MB per full re-upload; MatrixCity 13.6M ≈ 250k-splat tiles ⇒ ≤16 MB per texture ≈ one frame per full re-upload): **flat on every metric** over the full load window (fixed-45 s-window A/B reruns, 2026-07-26; an earlier "+11–20% tail improvement" reading came from a ~7 s partial-window instrument artifact — the v1 bench's settle detector silently broke — and is superseded; raw JSONs archived on the closed PR's branch under `delme/webgpu-writetexture-probe/bench-results/`). Loads are decode-dominated. - **Revisit when any of these hold:** (a) WebGPU becomes a default/primary backend; (b) single-texture ladders at ≥5M splats ship as real datasets (each ladder chunk then costs ~16 dropped frames on main vs ~1 with the wrapper for small append chunks — a doubling ladder's ~50% tail chunk is still ~8 frames); (c) an upstream three release consumes texture `updateRanges` (file the PR with the probe numbers — `writeTexture` already supports sub-rectangles with no `bytesPerRow` alignment constraint, and `common/Textures.js:357` invokes `texture.onUpdate` on every backend, so consume-and-clear semantics compose cleanly with luxar's `pendingFullUpload` guard). The archive branch carries `REVIVAL-NOTES.md` (audit 2026-07-26): read it before reviving — leading items are the knee=∞ extrapolation past 40% dirty and the (now-pinned) RGBA32F format assumption. Correctness on WebGPU remains automatic (full re-upload); the classic-WebGL Stage 1/2 wins are unaffected. Gate Stages 2/3 on the `Update Buffers` + `Load Arrays` profiler stages over the timelapse-nav benchmark (see the decode-bottlenecks findings) — ship only with numbers. Exit criterion: refinement passes upload O(new splats), not O(total). --- ## 8. Explicit deferrals - **Points/Lines symmetry** (LANDED 2026-07 — completed by the volumetric phases 3–4 arc): **Points storage LANDED** — the lift happened exactly as designed: the storage layer generalized into `element-texture-layout.ts` (parameterized `ElementTextureLayout`; gsplat = 4 texels, point = 3: center+radius / rgb+sharpness / scalar+alpha-reserved) + `element-storage.ts` (attach / dirty-range / `writeSortedIndex*`), and points now render from `uPointTex` via `aSortedIndex` with per-node materials (52 B/point vs 36 interleaved; the alpha slot pre-positions volumetric Phase 3's RGBA colors). **Points sort integration LANDED** (the arc's PR-B): the coordinator went geometry-neutral (`noteGSplatsCommit`/`noteGSplatsBlendingModeSwitch` → `noteDepthSortCommit`/`noteDepthSortBlendingModeSwitch`) and the points commit registers at both pool and non-pool paths with a LAZY centers provider (fresh Float32 copy of `data.positions`, paid only on the sorted path — the committed/lineage array's buffer must never itself be transferred; the copy also widens Float16). Points `normal` mode is now correctly back-to-front sorted, joins the camera-motion re-sort scheduler, the cross-node global `renderOrder` scale, and the `preserveOrdering` same-count-recommit prior; the LayersPanel mode-switch hook, per-mesh disposal release, and lazy-LOD demotion release all cover points. Order-dependence is judged directly on `needsDepthSort(mode)` for every geometry type: the interim `effectiveGeometryMode` downgrade helper (which kept points/lines `volumetric` additive-unsorted until the volumetric math shipped) became identity once volumetric Phase 4 landed and was DELETED from `rendering/blending-state.ts` — points `volumetric` has sorted since volumetric phase 3, lines since phase 4. **Lines storage + sort integration LANDED** (the arc's PR-C, completing the symmetry): 6 texels/segment (`LINE_TEXTURE_LAYOUT`, width `floor(4096/6)·6 = 4092`; layout in `line-geometry.ts` — startPos+startWidth / endPos+endWidth / startColor+startSharpness / endColor+endSharpness / segLen+per-endpoint cap-suppression scalars / scalars+reserved alphas, texel5.zw pre-positioning volumetric Phase 4's per-endpoint opacity), rendered from `uLineTex` via `aSortedIndex` with per-node materials. The migration retired three lines-only mechanisms in one stroke: the interleaved-buffer packing, the colormap ATTRIBUTE-SET toggle (the fixed layout always carries the scalar slots, so `hasScalars` no longer bucketing the pool nor rebuilding geometry — presence rides the `userData.hasScalars` stamp like points), and the material-manager LRU (the last cached material kind; all visual materials are now per node). The commit registers LAZY segment MIDPOINTS (`(start+end)/2`, fresh Float32 — the standard approximation; artifacts only when long segments interleave, subdivision if ever needed) and shares `preserveOrdering`, the append fast path (`writeLineTexels({fromSegment})`), the LayersPanel mode-switch hook, and disposal/demotion releases. Lines `volumetric` now sorts too (volumetric Phase 4 shipped 2026-07-24): lines render the real emission–absorption math and the same lazy midpoint provider feeds the sort through `needsDepthSort(mode)` — zero coordinator change. Pick `vElementId` reads `aSortedIndex` (storage slot); the context-restore hook covers lines through the same element-texture full-dirty branch as points/gsplats. - **Option 3b (WebGPU compute sort)**: writes the same `aSortedIndex`/ordering interface from a compute pass (three r184 `renderer.computeAsync` + `storage()` TSL nodes), eliminating the worker round-trip on the WebGPU backend. Requires a parity-policy carve-out (a compute sort has no GLSL twin — functional-equivalence tests on ordering output instead). - **RGBA16F splat texture** (32 B/splat): pure data-conversion change post-Phase 1; textures avoid every vertex-format pitfall that sank the `dd7c4478` interleaved narrowing. - **Weighted-blended OIT**: a sort-free `translucent` mode candidate; orthogonal to and composable with this plan (it would reuse Phase 0's real-alpha output). - **`opaque` mode at fractional opacity**: order-dependent today, stays so — degenerate configuration, documented rather than fixed. --- ## 9. Sequencing, sizing, and ship gates | Phase | Ships alone? | Size | Hard gate before merge | | ------------------------------- | ----------------------------- | ------------------------------------------------------------------ | --------------------------------------------------------------------------------- | | 0 — normal-mode alpha | Yes (visible improvement) | S (2 shader pairs, 1 blend-state helper) | Parity (incl. WebGPU bridge run) + overlap fixture + unchanged additive baselines | | 1 — texture storage | Yes (invisible; perf-neutral) | L (adapter, 4 shader stacks, materials, pool, budget, sync helper) | Pixel-identical E2E; commit-perf non-regression | | 2 — SortWorker + sort-at-commit | Yes (correct at rest) | M/L (WASM+TS kernel, SortWorker, commit wiring) | Kernel parity + benchmark floor + correct-after-settle E2E | | 3 — live re-sort | Yes (the feature) | M (scheduler, config, monitor) | Orbit E2E + FPS/latency budget + `?depthSort=0` determinism | | 4 — partial uploads | Optional (Stages 1–2 landed, all three geometries) | S (Stage 1) / M (Stage 2 gsplats) / S (Stage 2 Points/Lines) | Measured upload win; Stage 2 GSplats intentionally changes the interim append ordering as documented above | One PR per phase, in order; each leaves `main` shippable. Phase 1 is the only high-blast-radius change and is deliberately behavior-preserving so its review is a pure refactor review. Full E2E suite before each merge per repo policy; `make test-wasm` + `benchmark-wasm` for Phase 2; README updates (`rendering/README.md` architecture notes, `materials/gsplat/README.md` contract, `workers/` README) ride each phase's PR. --- ## 10. Risk register (ranked; what is actually hard about this) 1. **Three r184 half-supported vertex paths — the twice-burned pattern (highest risk).** This repo has already been bitten twice by r184 vertex-pipeline gaps: `GLSLNodeBuilder` hardcoding `gl_PointSize = 1.0` (forced the Points→instanced-quad migration) and `gpuType: HalfFloatType` silently ignored on interleaved attributes (the `dd7c4478` Float16 revert that shipped a regression). Phase 1 bets on two more under-exercised paths: a **uint instanced attribute** and **vertex-stage `textureLoad`**, each through _both_ the raw-GLSL ShaderMaterial path and the TSL node-builder (including WebGPURenderer's WebGL2-fallback GLSL emission). Raw WebGL2 support is verified (§1); what is NOT verified is what three's node builders emit. **Mitigation — mandatory spike before Phase 1 is scoped**: a 1-day throwaway probe rendering one quad-instanced mesh with a uint attribute + vertex texelFetch on all three surfaces (WebGL, WebGPU-native, `webgpuForceWebgl`). Fallback if uint attributes are broken in the node builder: a float attribute carries exact integers to 2^24 = 16.7M — exactly our per-node cap, so it works but eliminates headroom. > **SPIKE VERDICT (2026-07-13, three 0.184.0 — RISK RETIRED, all three surfaces PASS).** > Probe: 16 instanced quads, `Uint32Array` `aSortedIndex` carrying the permutation `(7i+3) mod 16`, colors vertex-fetched from an RGBA32F `DataTexture` (`NearestFilter`); control row fetches by `gl_InstanceID`/`instanceIndex`. Verified by pixel classification against a 16-color palette. > > | Surface | Backend reported | uint attr + vertex fetch | > | ------------------------------------------------------------------- | --------------------------------------------------------------- | ------------------------ | > | `WebGLRenderer` + GLSL3 `ShaderMaterial` (`in uint` + `texelFetch`) | webgl2 | **PASS** | > | `WebGPURenderer` + TSL (`attribute('…','uint')` + `textureLoad`) | webgpu-native (real adapter, headless `--enable-unsafe-webgpu`) | **PASS** | > | `WebGPURenderer({forceWebGL:true})` — TSL→GLSL emission | webgl2-fallback | **PASS** | > > Operational notes for Phase 1 tests: (a) `WebGPURenderer.readRenderTargetPixelsAsync(rt,x,y,w,h)` RETURNS the buffer — no out-param, unlike the WebGL variant (the existing dispatch in `post-processing/hdr/pixel-utils.ts` handles this); (b) `drawImage` from a WebGPU canvas races presentation (observed all-black readbacks) — sample via render-target readback, never via 2D-canvas copy; (c) readback row order differs per backend (`framebufferYDown`) — disambiguate with a known asymmetry. Spike code: `delme/spike-uint-vtf/` (gitignored, throwaway). > > **DEBT-REMEDIATION UPDATE (2026-07-13).** The vacuous-parity issue this spike's methodology flagged is FIXED: harness meshes now use production assembly, every sprite parity test carries a non-vacuousness guard + a footprint-invariant per-covered-pixel metric, and two REAL production TSL bugs the vacuousness had hidden were found and fixed (branch-scoped materialization of shared eigen values in `shader-tsl.ts`/`pick.tsl.ts` — quads degenerate for near-diagonal Σ2D; both vertex stages now trace inside `Fn()` with explicit `.toVar()` statements — the TSL house rule going forward, which #1683 later had to EXTEND to the fragment stage: a value shared between `colorNode` and `depthNode` must be assigned in an unconditional prologue both entry points call first, because the order three builds the two entry points in is not part of its API). Post-fix, GLSL↔TSL parity is PIXEL-EXACT (per-covered diff 0.00) on every variant. The same remediation also SUPERSEDED part of §3's Phase-0 wording: the GLSL additive alpha-MaxEquation guard referenced there was subsequently deleted (additive blending unified onto the shared helpers; the `raw-scene-hdr` capture path now sanitizes alpha at readback for all geometry types). **RESOLVED (same day, /double-check pass): the production "bowtie quads" under `?renderer=webgpu` were the SECOND half of the fragcoord story** — both TSL fragments anchored the Gaussian at `screenCoordinate` (TOP-LEFT origin on both backends by three's contract; the fallback emits `size.y - gl_FragCoord.y`) while `vCenterScreen` is bottom-left, mirroring the kernel about the horizontal midline. Centered mirror-symmetric parity fixtures are structurally blind to a y-flip, which is how pixel-exact parity coexisted with broken production. Fixed by reconstructing the bottom-left fragcoord in both factories — the un-flip term is three's `screenSize` (the bound render target's size, the exact term the builder's own flip used; the app-stamped `uResolution` differs whenever the target isn't drawing-buffer-sized, e.g. under SSAA), i.e. `vec2(x, screenSize.y - y)`; a permanently off-center `gsplat-offcenter` parity variant now guards the whole convention class. Production WebGPU gsplat rendering verified pixel-equivalent to WebGL on both fixtures. 2. **Product risk: `normal` mode may not look right on HDR scientific data (second, and the cheapest to retire).** Alpha-over compositing happens in linear HDR _before_ tone mapping; coverage alpha is `clamp(intensity·opacity)`, so dim splats (amplitude ≪ 1, common in microscopy) occlude almost nothing (near-additive look) while saturated ones occlude hard — and the aesthetic verdict on real data is unknown. The nightmare is discovering after Phases 1-3 that sorted `normal` isn't what anyone wants. **Mitigation: Phase 0 is deliberately first and standalone — evaluate on real datasets (neuromast, h2afva) at the Phase-0 checkpoint before any Phase 1-3 investment.** Order-_dependence_ artifacts will still exist at that checkpoint; judge color/occlusion character, not popping. > > **FIRST REAL-DATA VERDICT (2026-07-29, PR #916): POSITIVE on a heavy-background light-sheet volume.** The Tribolium embryo demo (`demo_gsplats_3d_tribolium_embryo.py`, 965×1871×991 Zeiss LightSheet Z.1 volume, amplitudes scaled 0.03 — squarely the "amplitude ≪ 1" microscopy regime this risk worries about) now ships `blending_mode="normal"` precisely *because* it looks better than the accumulating modes. Under `volumetric` the volume's heavy diffuse background integrates along every ray and saturates into a solid slab with the embryo buried inside it; under `normal` the background stops accumulating and the cell-surface nuclei resolve individually. Judged as this risk asks — color/occlusion character at a fixed camera, not popping. > > Two caveats on how far this generalizes. (a) It is **one dataset at one viewpoint, judged by eye** against a live A/B of the same scene; it says nothing about the async ordering state machine (risk 4) or about scenes whose subject *is* the accumulated volume. (b) The dominant mechanism was **not** the coverage-alpha behaviour this risk reasons about — it was the peak-vs-ray-integral projection split (`usesPeakProjection`), which removes the background accumulation outright regardless of how little dim splats occlude one another. So the prediction above ("dim splats occlude almost nothing ⇒ near-additive look") is right about the alpha-over term yet still under-predicts how different `normal` looks from `additive`/`volumetric` on such data, because the projection change dominates. Because nothing sums any more, the scene needed +1.97 stops of exposure to sit at a normal level — worth expecting for any future `normal`-mode conversion. > > **FOLLOW-UP VERDICT (2026-08-22, PR #1880): REVERSED after tuning the coupled volumetric controls.** The same Tribolium scene now ships `blending_mode="volumetric"` with strong absorption (κ=3.13) and opacity held down to 0.06. That pairing preserves the emission-absorption depth cue without letting the diffuse background saturate into the slab described above. The original fixed-opacity comparison remains useful history, but it was not a verdict on volumetric after jointly tuning absorption and emission. The scene still uses +1.97 stops: under the new settings it lifts the deliberately low-opacity ray integral rather than compensating for peak projection, so exposure alone is not diagnostic of the blending mode. > > **BASIS NOTE (same PR).** Both verdicts above argue about how a blending mode handles this volume's "heavy diffuse background" — and the fit itself then changed underneath them. The floor now subtracts the specimen's own ~675-count background rather than only the ~205-count detector offset, which removes ~84% of the scene's emitted mass while leaving nuclei brightness roughly unchanged. The background both entries reason about is therefore largely no longer in the data, so a future A/B on this scene starts from a different basis than either verdict, and neither should be read as a general statement about accumulating modes on heavy-background data. 3. **Phase 1 blast radius under a pixel-identical bar.** Four shader stacks rewritten simultaneously, the material system migrated to per-node (touching registration/camera-broadcast/dispose lifecycle — the same neighborhood as the known LRU-dispose engine-review finding), and the pool adapter rebuilt — with "unchanged visual baselines" as the merge gate. The gate is the right call (it makes review tractable) but it concentrates the campaign's regression risk in one PR. Mitigation: land it right after a green-baseline refresh; lean on `?dpr=1` pinning; budget review time like a rendering-engine change, not a refactor. 4. **The async ordering state machine.** Individually trivial rules (single-in-flight, generation match, length == instanceCount, identity on reuse, release on dispose) — but their product across commits, ladder appends, mode switches, pool reuse, dataset switches, LOD-child swaps, and context loss is where corrupt-permutation flicker hides. Mitigation: enforce the apply-invariant (**generation equal AND ordering.length == current count, else drop silently**) in exactly one function; property-test the state machine with a mock worker (random interleavings). 5. **Vertex-texture-fetch and RGBA32F-upload performance variance.** 16 texelFetches per splat (4 verts × 4 texels) vs one interleaved cache-line read today; VTF is fast on desktop/Apple GPUs (SparkJS ships this at scale) but historically weaker on some mobile tiles, and full RGBA32F `texSubImage2D` uploads have driver-dependent pitch/conversion costs. The Phase-1 perf gate catches it; escape hatches are pre-designed (RGBA16F halves both footprint and fetch bandwidth). 6. **Logistics that bite late**: the SortWorker is a new Vite `?worker` chunk that must ride into `luxar export` offline bundles and the native launchers (add an export smoke test); WASM-path forwarding into a _second_ worker; parity-harness extension to uint-attribute/texture-fetch shaders; E2E stale-server pitfall vs "pixel-identical" claims. --- ## 11. Adversarial re-check changelogs (all findings applied above) **Rev 4 (2026-07-14, pre-Phase-1 re-base; Rev 4.1 doc-truth pass same day).** Phase 1 was written against pool machinery that PR #511's own debt-remediation later deleted; re-based onto the shipped contracts: - **Growth is release + reacquire** (`growGSplatsGeometry`/`rebuildInterleavedBuffer` are gone): the splat texture is created and disposed WITH its pool geometry; in-place texture reallocation on a rendered geometry is forbidden — it is the exact `Info.memoryMap` strand class the buffer fix killed (empirically: +1 pinned generation per grow, permanent, on both native backends' pre-fix code). No carry-forward; commits rewrite the full count. - Non-pool fallback follows the fresh-swap+dispose pattern (broken-fallback finding long since fixed in #503). - The texelFetch prologue must be Fn-traced `.toVar()` statements (TSL house rule from the uninitialized-eigen bug); downstream now includes the unified `perspectiveNearFade`, unconditional coverage fade, and `screenSize` un-flip — all untouched by the prologue swap. - `invalidateRenderObjectFor` now also re-points `userData.pickNode` at the new geometry — Phase 1's pick side inherits the geometry swap for free; only the pick MATERIAL's `uSplatTex` rebind rides the new sync helper. - Exit criteria hardened with the committed tooling: `scripts/perf/timelapse-nav-bench.mjs` (perf gate) and `scripts/perf/grow-leak-probe.mjs` on native WebGPU (texture-lifecycle leak gate), plus the post-campaign parity-variant inventory and snapshot-regen discipline (fresh-server rule). - Status: Phase 0 merged to main (squash `e1e079e6`); B9a-c shader symmetry + native-WebGPU validation in PR #523. ### Rev 3 13. **Sort keys changed from raw f16 bit patterns to min/max-normalized uint16** (65536 buckets). SparkJS's f16 keys assume scene-scale depths; Luxar scenes carry physical units from nm to km, so raw camera-space depths can overflow f16 (all splats → Inf bucket) or underflow to 0 (all splats → one bucket) — either silently destroys the ordering. Normalized integer keys are scale-invariant at identical cost; unit-extremes tests added. 14. **Mode-switch data-flow gap closed**: switching a layer TO `normal` post-commit found the SortWorker without centers (registration is mode-gated) — and a naive re-request would hit the memoized noop path and skip re-projection. Fix: invalidate `committedData` + request a view update (one O(N) reprocess per switch). 15. **`generation` pinned**: per-node monotonic non-noop-commit counter — `loadedViewVersion` is insufficient because ladder appends share a view version while the count grows, and a short stale permutation applied to a grown buffer is corrupt, not merely stale. Noop commits keep in-flight sorts valid. 16. **Blend-state completeness**: `transparent: true` added to the pinned gsplat-normal state (transparent-pass classification / WebGPU pipeline key); vestigial `opacity` parameter dropped from `getGSplatNormalBlendingState()` (depthWrite is unconditionally false, so the state is opacity-independent). 17. **Sort gating source corrected**: gate on the material's live `userData.blendingMode` (what actually renders) rather than "the attrs-composer chain" — composition only runs via the LayersPanel; at initial load the material's mode comes from creation attrs. ### Rev 2 1. **Structural**: sorting moved out of the projection call entirely — plain-3D data (`ndim > 3` worker gate) never reaches the projection worker, so the original "sort in the projection worker" design would have sorted 10M splats on the main thread for the most common case. The SortWorker (originally Phase 3) is now the single owner of ordering and lands in Phase 2; `ProjectionViewState` is no longer touched. 2. **TSL double-premultiplication trap**: `material.premultipliedAlpha` injects an automatic RGB×alpha output node on NodeMaterial (`three.webgpu.js:21817`) — banned; blend factors set explicitly instead. 3. **WebGPU-bridge constraint scoped correctly**: the `gl.getError()` issue is separate _alpha-channel_ equation/factor state, not CustomBlending per se (`max` mode ships CustomBlending on TSL today) — the premult-normal state uses CustomBlending with a symmetric alpha channel. 4. **Depth-state gap closed**: gsplat premult-normal pins `depthWrite: false` unconditionally (generic normal's `opacity >= 0.99` flip would let near-transparent splat fragments write depth → occlusion halos); trade-off documented. 5. **Blend-state API shape fixed**: new `getGSplatNormalBlendingState()` helper instead of widening the shared `getCompleteBlendingState` (which has no geometry-type axis); the stale "TSL needs no gsplat override" comment is called out for update. 6. **Arithmetic**: VRAM delta corrected from +23% to **+31%** (68 B vs 52 B — the ordering attribute counts too). 7. **Addressing consistency**: bit-math constants were hardcoded for width 4096 while the width was capability-dependent — now compile-time defines/TSL params derived from actual width; "16384²" phrasing corrected (height is the scaling axis). 8. **Uint32 attribute binding made explicit**: verified r184 auto-selects `vertexAttribIPointer` for Uint32Array (`three.module.js:1950`) — noted as a relied-upon fact with parity coverage. 9. **Spurious `sortKeyCache` texel label** removed (was never specified). 10. **Sort-throughput numbers reconciled**: ~30-100 ms at 10M and a ≥100M splats/s benchmark floor (previous table said 10-30 ms while the gate implied 200 ms). 11. **Texture rebind rides `material-sync-helpers.ts`** (established post-commit render+pick sync seam) instead of ad-hoc commit code. 12. **Transfer-safety proven, not assumed**: ndim==3 fast path `.slice()`s (`projection/gsplats.ts:148-150`), so registering centers with the SortWorker by transfer can never detach a SliceCache buffer; added transfer-detach assertions to Phase 2 tests, plus SortWorker WASM-path init, app-teardown `dispose()`, and a per-node material-disposal leak check as explicit items.