# Luxar Zarr Format Specification > For a version-policy and migration summary that distinguishes this scene format from the gsplats format, see [Formats & Migration](./FORMAT_AND_MIGRATION.md). ## Version: 0.2 **New in v0.2 — the self-identifying root header.** The root group now carries `format_version` + `format_type: "luxar_zarr"` (mirroring the standalone gsplats header, so a reader can tell the two store kinds apart from the root attrs alone) and `luxar_software_version` (the `luxar.__version__` that wrote the scene — provenance only, **excluded from `content_hash`**). The 0.1 key `luxar_version` is no longer written; every reader still accepts it as the legacy spelling, so published 0.1 stores load unchanged. Version checking follows one rule in Python and the viewer — see [Formats & Migration](./FORMAT_AND_MIGRATION.md#versioning-policy). **Features in v0.1:** - Chunk-based spatial index for efficient nD point queries - Points are reordered using Morton/Hilbert space-filling curves for spatial locality - Compound ordering: discrete dimensions (time/channel) + spatial dimensions - HDR color support (float32 values in/out; stored quantized per-channel true-log) - Transform system with matrix transposition for THREE.js compatibility This document specifies the Zarr-based storage format used by Luxar for high-performance 3D and nD scientific visualization. ## Overview The Luxar Zarr format is a hierarchical data structure designed for efficient storage and streaming of large-scale scientific data with support for arbitrary dimensionality, transformations, and rendering attributes. ## End-to-End Data Flow This diagram shows how data flows from Python creation through storage to WebGL rendering: ``` ┌────────────────────────────────────────────────────────────────────────────┐ │ PYTHON LAYER (luxar packages) │ ├────────────────────────────────────────────────────────────────────────────┤ │ │ │ luxar.core │ │ ┌──────────────────┐ │ │ │ Scene, Points │ Define scene graph with transforms │ │ │ Dimensions │ positions: float32[N, D] │ │ └────────┬─────────┘ colors: float32[N, 3] │ │ │ radii: float32[N] │ │ ↓ │ │ luxar.validation │ │ ┌──────────────────┐ │ │ │ Type checking │ Validate shapes, ranges, semantic types │ │ │ Shape validation │ Ensure data consistency │ │ └────────┬─────────┘ │ │ │ │ │ ↓ │ │ luxar.encoding │ │ ┌──────────────────┐ │ │ │ Semantic typing │ COORDINATE → uint16 fixed-point (AUTO/MEM), f32 (PREC) │ │ │ Quantization │ COLOR → uint8 (SDR), geolog u16 (HDR) │ │ │ Broadcasting │ Uniform values → single scalar │ │ └────────┬─────────┘ │ │ │ │ │ ↓ │ │ luxar.io │ │ ┌──────────────────┐ │ │ │ Spatial ordering │ Morton/Hilbert space-filling curves │ │ │ Compound sort │ Discrete dims (time) → spatial (x,y,z) │ │ │ Chunking │ Split into 16KB-256KB chunks (target 64KB) │ │ │ AABB calculation │ Per-chunk bounding boxes │ │ └────────┬─────────┘ │ │ │ │ └───────────┼─────────────────────────────────────────────────────────────────┘ │ ↓ Write ┌───────────────────────────────────────────────────────────────────────────┐ │ STORAGE LAYER (Zarr) │ ├───────────────────────────────────────────────────────────────────────────┤ │ │ │ scene.luxar.zarr/ │ │ ├── zarr.json Scene attributes and consolidated metadata │ │ └── node_name/ │ │ ├── positions/ Blosc(zstd-9) compressed uint16 chunks │ │ ├── colors/ Blosc compressed uint8/float32 │ │ ├── radii/ Compressed or broadcast scalar │ │ └── chunk_bounds/ AABB per chunk for spatial queries │ │ [xmin,xmax, ymin,ymax, zmin,zmax, ...] │ │ │ │ Compression: 2-10× (blosc/zstd) + quantization: 2-4× = 4-40× total │ │ │ └───────────┬───────────────────────────────────────────────────────────────┘ │ ↓ HTTP/zarrita ┌───────────────────────────────────────────────────────────────────────────┐ │ TYPESCRIPT LAYER (luxar-viewer) │ ├───────────────────────────────────────────────────────────────────────────┤ │ │ │ cache/ │ │ ┌──────────────────┐ │ │ │ S-cache + L0 │ Decoded slices + decompressed chunks (RAM) │ │ │ L1: Memory LRU │ ~100MB, ~1μs access │ │ │ L2: OPFS │ ~2GB, ~1ms access │ │ │ HTTP fetch │ Unlimited, ~100ms access │ │ │ Prefetcher │ Adjacent chunks (±1 in each dimension) │ │ └────────┬─────────┘ │ │ │ │ │ ↓ │ │ data/ │ │ ┌──────────────────┐ │ │ │ Scene loader │ Parse scene attributes, build THREE.js scene graph │ │ │ Spatial index │ Query chunk_bounds for AABB intersection │ │ │ Array decoder │ Dequantize uint16 → float32 │ │ │ nD slicer │ Hypersphere visibility (effective radius) │ │ └────────┬─────────┘ │ │ │ │ │ ↓ │ │ rendering/ │ │ ┌──────────────────┐ │ │ │ Material manager │ Create shaders with world-space sizing │ │ │ Post-processing │ Bloom, tone mapping, detector noise │ │ └────────┬─────────┘ │ │ │ │ └───────────┼─────────────────────────────────────────────────────────────────┘ │ ↓ BufferGeometry ┌───────────────────────────────────────────────────────────────────────────┐ │ WEBGL LAYER │ ├───────────────────────────────────────────────────────────────────────────┤ │ │ │ Vertex Shader │ │ ┌──────────────────┐ │ │ │ Transform points │ Apply 4×4 matrices (scene, view, projection) │ │ │ Size calculation │ Angular diameter → pixel size │ │ │ Color pass │ Pass attributes to fragment shader │ │ └────────┬─────────┘ │ │ │ │ │ ↓ │ │ Fragment Shader │ │ ┌──────────────────┐ │ │ │ Gaussian kernel │ Smooth point splatting │ │ │ HDR rendering │ Float16 framebuffer for >1.0 colors │ │ │ Alpha blending │ Additive/premultiplied modes │ │ └────────┬─────────┘ │ │ │ │ │ ↓ │ │ Display: 60 FPS interactive visualization of 100K-10M elements │ │ │ └───────────────────────────────────────────────────────────────────────────┘ **Key Performance Characteristics:** - Python write: ~1-5M points/sec (spatial ordering overhead) - Compression ratio: 4-40× (quantization + blosc) - Network bandwidth: 50-500KB/sec for smooth navigation - Cache hit rate: 80-95% with prefetching - GPU rendering: 100K-10M elements at 60 FPS ``` ## Format Structure The tree below shows the default Zarr format-3 container layout: ``` scene.luxar.zarr/ ├── zarr.json # Scene attributes and consolidated metadata ├── / # Scene nodes (groups or points) │ ├── zarr.json # Node attributes (includes spatial index metadata) │ ├── positions/ # Point positions (required for points, spatially sorted) │ ├── colors/ # Point colors (optional, same order as positions) │ ├── radii/ # Point radii (optional, same order as positions) │ ├── sharpnesses/ # Point sharpness (optional, same order as positions) │ ├── chunk_bounds/ # Chunk bounding boxes for spatial queries (optional) │ ├── label_offsets/ # Per-element label byte offsets, CSR-style (optional) │ ├── label_bytes/ # Concatenated UTF-8 label strings (optional) │ ├── key_offsets/ # Per-element key byte offsets, CSR-style (optional) │ ├── key_bytes/ # Concatenated UTF-8 key strings (optional) │ ├── image_label_offsets/ # Per-element image byte offsets, CSR-style (optional) │ ├── image_label_bytes/ # Concatenated encoded image blobs (optional) │ └── / # Nested child nodes (recursive structure) └── overlays/ # Screen-space overlays (optional) └── / # Individual overlay ├── zarr.json # Overlay attributes (type, position, style, visible_range, hover) └── image. # Raw image file (image overlays only; exact name matches payload; bytes folded into content_hash at compile time) ``` See [Compatibility Notes](#compatibility-notes) for the format-2 container layout; the attributes and node structure are the same in both formats. ### Compression & the `luxar_delta_v1` filter All arrays use Blosc zstd level 9 with a width-aware shuffle policy (byte shuffle for multi-byte integer codes, no shuffle for uint8/floats — see `luxar.encoding.compression`). Additionally, any quantized uint8/uint16 code array (positions/vertices/centers, Cholesky halves, radii/widths/amplitudes, sharpness, colors) MAY carry the Luxar-owned `luxar_delta_v1` transform. In a format-3 store it appears as `{"name": "luxar_delta_v1", "configuration": {"cols": C, "bits": 8|16}}` in the array's `zarr.json` codec chain: columnar per-chunk delta+zigzag residuals, applied probe-gated at encode time (only where it measurably shrinks the store — 12-16% whole-store, lossless). It is a pure storage transform below the `encoding` attrs; zarr/zarrita undo it during whole-chunk reconstruction, so decode and random access are unchanged. Readers need the codec available: an installed Luxar package auto-registers both format spellings through the `numcodecs.codecs` and `zarr.codecs` entry points (`import luxar.encoding` also registers them); the viewer registers `numcodecs.luxar_delta_v1` for format 2 and the bare `luxar_delta_v1` for format 3. Readers without Luxar installed fail loudly (unknown codec), never silently. Full wire-format spec: `docs/specs/GSPLATS_ZARR_FORMAT.md` § "The `luxar_delta_v1` delta filter". ## Scene-Level Attributes The root group's attributes contain scene-wide configuration: | Root attr | Written by | Meaning | |---|---|---| | `format_version` | 0.2+ | Scene format version (`"0.2"`). Checked by every reader against the contract's `supported` set: supported → silent, same-major newer minor → warn and load, anything else → refused. | | `format_type` | 0.2+ | `"luxar_zarr"` — identifies a compiled scene (a detached `.gsplats.zarr` root says `"gsplats_zarr"`). | | `type` | all | `"scene"`. | | `luxar_software_version` | 0.2+ | The `luxar.__version__` that wrote the store. Provenance only — **excluded from `content_hash`** by both hashers, so two releases compiling the same scene agree on the digest and a `luxar optimize` restamp never churns viewer caches. | | `luxar_version` | 0.1 only | The legacy version key. Read as a fallback when `format_version` is absent; never written by a current compiler. | | `content_hash` | all | Post-order xxhash64 build identity for the store: array bytes, storage identity, attrs (minus this key and `luxar_software_version`) and child digests, computed by `io/_compiler/finalize/hashing.py`. Every finite float nested in hashed **group** attrs is persisted at 12 significant decimal digits before hashing, absorbing platform drift below roughly 1e-12 relative; larger differences and exact array encoding attrs must be stabilized by their producers. The viewer validates its cache against it. | | `scene_dimensions` | all | The nD dimension table (below). | No `timestamp` is written on a scene: the digest must be reproducible across two compiles of the same script. Geometric-log scalar and per-channel companded rails are stored at float32 precision, covariance certificates at four significant digits, and fitted coordinate-grid steps at twelve significant digits. Those producer rules, together with the group-attribute grid above, suppress the observed libm/reduction drift while keeping exact array decoder metadata intact; larger differences remain producer defects and can still change the digest. ```javascript { "format_version": "0.2", "format_type": "luxar_zarr", "type": "scene", "luxar_software_version": "2026.9.16", // provenance; NOT part of content_hash "units": "um", // Physical units (nm, um, mm, cm, m, meter, metre, km, inch, foot, px, au, s) "scene_dimensions": { // Scene-level dimension specification (REQUIRED for nD data) "dimensions": [ { "name": "x", "unit": "um", "range": [-100.0, 100.0], // Optional bounds "step": 1.0, // Optional navigation step size "display": true, // Whether dimension is displayed (max 3) "discrete": false, // Whether dimension has discrete values "cyclic": false, // Whether dimension wraps around "scale": 1.0, // Physical scale factor "spatial": true, // Whether points extend through this dimension "description": "X axis" // Optional description }, // ... more dimensions ] } } ``` Additional JSON-serializable scene metadata may be authored through `scene.attrs`, for example `title`, `description`, or `sample`. Mutations made while the `LuxarZarrCompiler` is open write through to the root `.zattrs`; mutating the live mapping after finalization only updates its in-memory cache and emits a warning. The structured root attrs Luxar owns — `scene_dimensions`, `viewer_config`, `citation` — each have a dedicated validating API (`Scene.dimensions`, `Scene.viewer_config`, `create_scene(citation=...)`). Author them there, not through `scene.attrs`: the mapping accepts free-form keys at the root, so writing one of those by hand skips its validation and leaves the `Scene` object's own copy stale. ### `citation` (optional root attr) `citation` credits whoever produced the data the scene shows. It is written to the root attributes rather than to a companion file or a web page so the attribution travels with the store — through `luxar export`, through a copy of the `.luxar.zarr` handed to someone without the repo, and into any viewer that opens it. ```javascript { "citation": { "short": "OpenCell (Cho et al. 2022); embeddings by cytoself (Kobayashi et al. 2022)", // REQUIRED: single line, what a UI renders "ref": "Cho / Kobayashi et al. 2022", // optional compact caption reference, max 40 chars "doi": "10.1126/science.abi6983", // optional, bare DOI (no https://doi.org/ prefix) "license": "CC BY-SA 4.0", // optional "url": "https://example.org/dataset" // optional } } ``` Only `short` is required: a citation that cannot be displayed is not a citation. `ref` is the optional compact form used where the full byline does not fit, notably the bundled demos' bottom-right captions; it is refused above 40 characters rather than truncated. A store carrying `ref` is intentionally a newer format-contract payload; feeding those attrs back through an older Luxar authoring path fails citation validation because that version does not recognise the key. Every field must be a single line of printable text — every line break and control character is refused, including a bare `\r` and a bidi override, since a credit that renders differently from the string that was stored is a spoofing risk rather than a cosmetic one. `doi`, when present, must match the full `10./` shape, so the near-misses people paste (a URL, a `doi:` prefix, a registrant whose suffix was lost) are refused rather than stored. `url`, when present, must be `http://` or `https://`: it is the one field a UI turns into a link, and a citation travels inside data that is copied and published onward, so an executable or inline-payload scheme is not storable here. The key is **absent** — not `null`, not an empty object — when the scene owes no credit, which is the normal case for procedurally generated data. Readers should therefore treat a missing `citation` as "unknown or not applicable" and never as a claim that the data has no author. Written by passing `citation=` to `LuxarZarrCompiler.create_scene(...)`. The payload is validated at write time (`luxar.core.citation.validate_citation`), so a malformed citation raises instead of being baked into every copy of the data. ### `input_digests` (optional root attr) `input_digests` maps each resolved input artifact's basename to the lowercase sha256 digest of the bytes the demo resolver accepted. The map is canonical by basename and records every manifest artifact resolved in the process before the scene was stamped; basenames are therefore expected to be unique within that set. The attribute participates in the root `content_hash`, so changing the resolved input generation changes the scene's content identity even when its geometry is otherwise unchanged. This differs from the root `environment/` sidecar group, which is derived metadata and is explicitly excluded from `content_hash`. ### `incomplete` (optional root attr) `incomplete` (boolean) is written to the root `.zattrs` only when the writer aborted before `finalize()` completed — an exception or Ctrl-C propagated out of the `LuxarZarrCompiler` context, or `finalize()` itself failed midway. A successful `finalize()` never leaves this marker. `LuxarScene.load` refuses to load a store carrying `incomplete: true`, since it may be missing nodes or consolidated metadata. ## Node Types ### 1. Group Nodes Group nodes organize the scene hierarchy and can contain child nodes. **Attributes (.zattrs):** ```javascript { "type": "group", "transform": [1,0,0,0, 0,1,0,0, 0,0,1,0, 0,0,0,1], // 4x4 matrix as 16-element array "nd_transform": { // Optional, per-dimension transforms "Time": {"scale": 0.001, "offset": 50.0}, // for non-displayed dimensions "Channel": {"permutation": [2, 1, 0]} }, "opacity": 1.0, // 0.0-1.0, inherited by children "absorption": 1.0, // volumetric mode's kappa (>= 0, default 1.0; multiplicative) "gamma": 1.0, // 0.1-10.0, per-node gamma correction "intensity": 1.0, // 0.0-100.0, per-node linear color gain — but on a // colormapped node this is the scalar display window, // not a gain (see Scalar Colormap Attributes) "offset": 0.0, // -10.0-10.0, per-node additive brightness shift (black level) "blending_mode": "additive", // normal, additive, max, opaque, luminous, volumetric — written // only when explicitly set; unset ⇒ inherited from the // nearest ancestor that sets it (viewer default: additive) "layer_order": 20, // optional: authored cross-layer draw order // (higher = nearer the camera). Never stamped — // absence means "use the inferred ordering". "colormap": "viridis", // Optional palette; nearest-setter-wins for rendering. GSplats // inherit it directly; Points / Lines / Mesh scalar leaves must // still author their own colormap today (see Rendering Attribute // Composition) "layer": false, // Optional: if true, node appears in the viewer's Layers panel "visible": true, // Optional: initial visibility when the scene loads (default true) "child_index": 0 // Insertion order among siblings (stamped on add). The viewer // sorts siblings by this so the scene graph / layers panel // follow napari-style add order, not zarr's alphabetical // consolidated-metadata enumeration. Absent → enumeration order. } ``` #### Volumetric blending and `absorption` Set `blending_mode` to `"volumetric"` for emission–absorption compositing: each element emits light while attenuating elements behind it. The viewer supports this mode for Points, Lines, and GSplats and depth-sorts the order-dependent geometry back-to-front. `absorption` is the node-level coefficient κ. It is non-negative, defaults to `1.0`, and composes multiplicatively through the scene hierarchy. It is inert in the other blending modes. At κ = 0, volumetric rendering reaches the additive limit; increasing κ produces stronger self-occlusion. When colors carry an RGBA alpha channel, the alpha is converted to optical-depth weight in volumetric mode instead of being treated only as a linear contribution scale. For the rendering equations, blend-state contract, and per-geometry details, see the [Volumetric Blending specification](../specs/VOLUMETRIC_BLENDING_SPEC.md). ### 2. Group Kinds — Specialized `group` Nodes A `Group` may carry an optional `kind` attribute that turns it into a specialized container with viewer-aware semantics. Specialized groups also carry a `display_type` attribute (one of `"points"`, `"lines"`, `"gsplats"`) — the layers panel uses this for the user-facing type label, so a layer reads as one logical entity of `display_type` rather than as a "group". #### `kind: "lod"` — Level-of-Detail group Picks **one of N alternative children** at runtime based on the current view. Each child carries a `coverage_fraction` threshold; the group's `selector` attr names the UNITS. Under `selector: "screen-area"` (what every auto-derived ladder stamps) a threshold is a literal **screen-area fraction**: the viewer projects the LOD group's bbox to a screen-space rect, takes that rect's area as a fraction of the viewport area, and renders the **finest** child whose threshold is satisfied (with 10% asymmetric hysteresis on the downgrade direction to suppress flicker). The derived whole-object ladder is `[0, …, 1/8, 1/4, 1/2]` — full detail while the object occupies at least half the screen, one level coarser per halving of occupied area — identically on any monitor/viewport size (the metric is built from NDC fractions; the rect is clipped to the viewport, off-screen reads 0, and sub-pixel-thin content ramps to its clipped linear span — see the normative metric definition in `docs/specs/GSPLATS_ZARR_FORMAT.md`). Under the legacy `selector: "coverage"` (older stores, and explicitly authored `coverage_fractions=[...]` lists) the thresholds are diagonal-metric units in `[0, 4]`: the viewer compares them against the projected bbox diagonal over `FILL_FACTOR=0.5 ×` the fitted screen axis (`min(width, height)` — the extent the camera framing actually fits). Under a PERSPECTIVE camera, a group whose bounds reach the camera's near plane has no meaningful projection (the homogeneous divide degenerates), so both selectors saturate to the **finest** child; an ORTHOGRAPHIC projection never degenerates, so the ordinary clipped metric applies directly (a camera inside a large group still reads full coverage naturally — see the normative rules in `docs/specs/GSPLATS_ZARR_FORMAT.md`). `kind="lod"` is **geometry-agnostic**: children can be points, lines, gsplats, or themselves specialized groups (e.g. a Partition group inside an LOD group). The finest child's resolved `display_type` becomes the LOD group's `display_type`. **Standalone `.gsplats.zarr` root**: a `kind=lod` group is also a valid root of a standalone `.gsplats.zarr` file — the file root IS the node (current standalone format version: **v3.4**, see `docs/specs/GSPLATS_ZARR_FORMAT.md`). The viewer opens such a file directly (`?src=.gsplats.zarr`) and frames on its `position_bounds`. On-disk, children are `child_/` in **coarsest→finest** order; the writer always stamps `default_level: 0` (the coarsest child) — a progressive-load hint (render cheap first, then refine), deliberately decoupled from the data-model default (the finest level the `.centers` accessor returns). **Attributes (.zattrs):** ```javascript { "type": "group", "kind": "lod", "display_type": "gsplats", // Resolved at write time from the finest // child; the layers panel uses this as // the user-facing layer type. "selector": "screen-area", // Units of the children's coverage_fraction: // "screen-area" (derived ladders) or // "coverage" (legacy diagonal metric). "default_level": 0, // 0-based initial active level (coarsest→finest). // Seeds the "Active level" dropdown in the // Layers panel; does not lock the runtime // selector by itself. "transform": [1,0,0,0, 0,1,0,0, 0,0,1,0, 0,0,0,1], "nd_transform": { ... }, // Optional, same shape as on plain Group nodes "opacity": 1.0, // Compositing — inherited by children "absorption": 1.0, // volumetric mode's kappa (>= 0, default 1.0; multiplicative) "gamma": 1.0, "intensity": 1.0, "offset": 0.0, "blending_mode": "additive", // Only when explicitly set (unset ⇒ inherited) "layer": false, // Optional: expose in the Layers panel with // an "Active level" dropdown + "N LODs" badge "visible": true } ``` **Children**: - Subgroup naming is **not** enforced; Python's convenience API writes `child_0`, `child_1`, … in **coarsest→finest** order, and the loader treats insertion order as authoritative. - Each child's `.zattrs` MUST carry a `"coverage_fraction"` in the group's selector units (`[0, 1]` for `screen-area`; `[0, 4]` for legacy `coverage`). Values must be strictly monotonic increasing in coarsest→finest order, and the coarsest is always `coverage_fraction: 0.0` (always applicable). A **whole-object** ladder (the auto-derivation: SCREEN-OCCUPANCY HALVING, independent of element counts) anchors its finest at `coverage_fraction: 0.5` — full detail while the object occupies at least half the screen — with each coarser level halving the threshold (`…, 1/8, 1/4, 1/2`). `luxar restamp-lod --anchor` may explicitly re-anchor a stored whole-object ladder at any fraction in `(0, 1]`; partition-bound ladders remain at `1.0`. The `1.0` fills-screen ceiling (the tile alone occupying the whole screen) holds a level until the node is larger still. That fills-screen anchor is the right one whenever a **spatial partition** is part of the switch, because a tile's projected rect is intrinsically a fraction of the whole object's; a whole-object-anchored ladder there would put every tile on its finest level while the object is merely full-frame. Every producer emits it automatically once it can see the binding: the `adaptive` / `overview` gsplat recipes; the two gsplat writers' topology-aware fallback (a `kind=partition` crossed on the way down); scene-side adders that detect a `kind=partition` ancestor; and `add_points(partition=..., substitutive_lod=...)`, which verifies a multi-part fine branch directly. In that overview topology the group's bbox is the whole object, so the anchor is the recipe contract (coarse at opening frame, fine on zoom), not tile geometry. An **explicitly authored** `coverage_fractions=[...]` list always wins over all of them. The rule assumes >= 2 parts. Every producer that can see the final sibling count excludes a **one-part** partition and falls back to the whole-object `0.5` anchor — `--recipe adaptive`, both gsplat writers, and the Points overview composition do, which matters because a dataset below `--max-elements` yields exactly that shape. The partition-ancestor scene-adder route cannot check it (part 0's ladder is derived before part 1 exists); the compiler's finalize pass warns when it sees the result, without rewriting it. - A child MAY carry `"lod_bounds": {"min": [...], "max": [...]}` with the same nD axis order and shape as `position_bounds`. This is a producer-chosen robust extent (for example percentile bounds that exclude a sparse tail), used **only** to size the node for either LOD selector metric. Frustum gating, eviction, framing, clipping, and scene ranges continue to use the complete `position_bounds`, so excluded outliers remain visible and resident when they should. The robust bound must stay within `position_bounds`; a bound that is not contained is rejected and the child falls back to `position_bounds`. Missing or malformed `lod_bounds` fall back to that child's `position_bounds`; producers should therefore stamp every child in a ladder when they want one consistent robust extent. Decimation, culling, and filtering must recompute or remove the derived bound. Producer-side authoring policy is tracked in #1655. - Children themselves are standard nodes — they retain their own `type` (`gsplats` / `points` / `lines` / `group`, possibly with their own `kind` attr) and full attr set. **Builder API (Python):** ```python # Manual (hand-authored thresholds default to the legacy # selector="coverage" diagonal units; pass selector="screen-area" to author # literal screen-area fractions): lod = scene.add_lod_group("multires") lod.add_gsplats_from_data("child_0", coarse_data, coverage_fraction=0.0) lod.add_gsplats_from_data("child_1", medium_data, coverage_fraction=0.5) lod.add_gsplats_from_data("child_2", fine_data, coverage_fraction=1.0) # Convenience (auto-derives coverage_fractions by screen-occupancy halving — # selector="screen-area": coarsest 0.0, halving up to the finest's # half-screen-area anchor 0.5): scene.add_gsplats_from_data( "multires", flat_data, lod_group=dict(compression_factor=4, levels=2), additive_lod=dict(n_lods=4), ) ``` #### `kind: "partition"` — Spatial-decomposition group Decomposes a single large leaf node (10M+ elements) into N smaller children for per-child frustum culling and per-child LOD. The user adds one node; the writer partitions it via recursive BSP at compile time (balanced **median** split by default; `midpoint` and `sah` rules are also available). The viewer renders all children that intersect the camera frustum simultaneously. A frustum-only per-frame selector also gates fetch and eviction for off-screen parts; unlike an LOD selector, it never substitutes one child for another. Children are **homogeneous**: every child's resolved `display_type` must match the wrapper's (you cannot decompose a single logical layer into mixed-type parts). **Standalone `.gsplats.zarr` root**: a `kind=partition` group is also a valid root of a standalone `.gsplats.zarr` file. The viewer opens it directly (`?src=.gsplats.zarr`) and frames on `position_bounds`. The `luxar gsplat partition` CLI writes this shape via spatial BSP (`--parts` / `--max-elements` / `--rule median|midpoint|sah`). **Attributes (.zattrs):** ```javascript { "type": "group", "kind": "partition", "display_type": "points", // All children resolve to this type. "max_elements": 1000000, // Per-part cap that drove the BSP recursion. "bsp_tree": { // Optional recursive tree; axis is a position-column "axis": 0, // index mapped through displayDims by the viewer; // any unmapped split axis rejects the whole tree. "split": 0.0, "left": { "part": 0 }, "right": { "part": 1 } }, "position_bounds": { // Union of children's bboxes — lets "min": [-10, -10, -10], // picking / framing / scene-bounds-cache "max": [10, 10, 10] // treat the layer as one logical entity. }, "transform": [1,0,0,0, 0,1,0,0, 0,0,1,0, 0,0,0,1], "nd_transform": { ... }, // Optional, same shape as on plain Group nodes "opacity": 1.0, "absorption": 1.0, // volumetric mode's kappa (>= 0, default 1.0; multiplicative) "gamma": 1.0, "intensity": 1.0, "offset": 0.0, "blending_mode": "additive", // Only when explicitly set (unset ⇒ inherited) "layer": false, // Optional: expose in the Layers panel with // a "N parts" badge "visible": true } ``` For points and Gaussian splats, each split plane exactly separates the child bounds. Lines and mesh are partitioned atomically by polyline and face centroid, respectively, so their vertices may cross a split plane; their stored tree is a stable approximate order rather than an exact painter's-order separation. **Children**: - Subgroup naming is **not** enforced; the convenience kwarg writes `part_0`, `part_1`, … in BSP recursion order. Names may have gaps when an empty region is omitted; the contiguous `child_index` attr is the identity used by `bsp_tree` leaves and the viewer. - Each child is a standard `points` / `lines` / `gsplats` node (or itself a kind=lod / kind=partition group). All must resolve to the same `display_type`. **Builder API (Python):** ```python # Convenience kwarg on the leaf adders — partition applies at compile time # and the user never sees the wrapper unless they inspect the zarr: scene.add_points("pts", positions, partition=True) # default cap scene.add_points("pts", positions, partition=dict(max_elements=500_000)) scene.add_gsplats("splats", centers, amplitudes, cholesky, partition=dict(max_elements=2_000_000, rule="median")) # Manual (explicit tree construction): part = scene.add_partition_group("manual", display_type="points", max_elements=500_000, layer=True) part.add_points("part_0", subset0) part.add_points("part_1", subset1) ``` `max_elements` and `rule` are the only accepted keys in a `partition` dict; unknown keys raise an error before any node is written. #### Multi-additive LOD (progressive loading) — Points / Lines / GSplats All three leaf types (Points, Lines, GSplats) support a uniform **multi-additive-LOD** layout for progressive loading. A leaf node with `n_additive_sublods > 1` carries `additive_/` subgroups under its path; each subgroup is a fully-formed leaf node of the same type, carrying a subset of the parent's data plus its own spatial index. The viewer's progressive loader concatenates loaded levels for render and refines toward the full data over `requestAnimationFrame()` frames once initial paint commits. Mesh uses the same subgroup layout but is documented separately (see the *Mesh Node* section): its levels are a partition of the FACES, each re-indexing its own vertex table, so a level is not a subset of the parent's arrays and carries no spatial index — and the ladder is a *reveal*, restricted to orderings whose every prefix is one connected patch. **Parent attributes (.zattrs):** ```javascript { "type": "points", // or "lines" / "gsplats" "n_points": 10000, // (or n_segments / n_splats — total across levels) "n_additive_sublods": 4, "position_bounds": { "min": [...], "max": [...] }, /* compositing attrs ride here */ } ``` **Subgroup naming:** `additive_/` for `i = 0..n-1`, in **coarsest → finest** order. The convenience writer (`additive_lod=` kwarg on `add_points` / `add_lines` / `add_gsplats_from_data`) emits the subgroups in this convention. **Ladder provenance:** the ordering method that built the ladder is not a top-level attr — it rides in the quality-stamp dicts, as `lod_method` inside the parent's `level_stats` and inside each subgroup's `lod_stats` (alongside `lod_level` / `lod_n_lods` / `lod_breakpoints_kind`). Readers treat absence as "unstamped". **Per-type unit:** - **Points** — per-element. Each subgroup contains a subset of `positions` + per-element attrs (`colors` / `radii` / `sharpnesses` / `scalars`). - **GSplats** — per-element. Each subgroup contains a subset of `centers` / `amplitudes` / `cholesky_factors_diag` (+ `cholesky_factors_offdiag`) / `colors`. (Since format v3.1 the in-memory packed `cholesky_factors` is stored on disk split into `_diag` + `_offdiag` so each can be encoded independently; `_offdiag` is absent for 1D splats. Legacy v3.0 files store a single packed `cholesky_factors`, read via a presence-detect fallback. Format **v3.2** renamed the `kind=lod` selector attrs to the coverage semantics described above — `selector: "coverage"` + per-child `coverage_fraction`; **v3.3** added the optional `luxar_delta_v1` filter on quantized code arrays. The current format is **v3.4**, which adds the `selector: "screen-area"` mode (literal screen-area-fraction thresholds, stamped by every derived ladder); see `docs/specs/GSPLATS_ZARR_FORMAT.md`, the authoritative gsplats format spec.) - **Lines** — per-polyline. Each subgroup contains WHOLE polylines (vertices + their segments). Segment indices are local to the subgroup so topology stays valid during partial loads; the viewer offset-adjusts on concatenation. Supports all four `line_type` variants (`segments` / `polyline` / `loop` / `indexed`). **Labels (Points / Lines):** per-element string labels are NOT stored per subgroup. No single level's array is what a pick index addresses, since the viewer's loader concatenates the levels it has loaded into one buffer — so the label CSR (`label_offsets` + `label_bytes`) lives on the **parent** node, which consequently carries `has_labels: true`, and the `additive_` subgroups carry no label arrays and no `has_labels`. Index `k` of the parent CSR is the `k`-th element of the concatenation of `additive_` in order (coarsest → finest), each level in its own stored (spatially reordered) order — the same on-disk index space a flat labelled leaf's CSR uses, just spanning the levels. Labels are all-or-nothing across a ladder: a partially-labelled ladder is rejected at write time. The viewer commits levels coarsest-first, so a fully-loaded ladder maps straight through — index `k` is committed slot `k`. The **committed buffer** is not in general a prefix of this union, though: the per-level loader compacts out elements culled by the current nD slice and fetches only the chunk ranges a query intersects, so slots shift. For **Points** that shift is corrected: the viewer composes each level's own visible-slot → on-disk-index map (the map issue #1421 / PR #1425 introduced for **flat** nodes) into this union index space, offsetting level `i` by the preceding levels' on-disk counts, so a labelled Points ladder resolves exactly under an nD slice too (issue #1439). Where the levels' metadata is inconsistent the viewer publishes no map at all and falls back to the raw committed slot, rather than composing an index it cannot trust. For **Lines** no map is composed across the levels, so a laddered lines node still resolves at the raw committed slot. Two further notes. Under the documented `partition=`-outer + `additive_lod=`-inner composition the CSR lands on each `part_` ladder parent, which is exactly where the viewer looks it up (the hit leaf scene node — the outermost `kind=partition` wrapper is the reported path only), so that composition resolves too, through the same per-geometry path as an unpartitioned ladder. And for Lines the CSR is per-vertex, matching the flat Lines writer, while the viewer's Lines pick id is a per-segment storage slot. Issue #1424 bridged that granularity for **flat** Lines nodes — a labelled flat lines node resolves the picked segment's slot back to that segment's start vertex row in the stored ordering — but it does so through the same visible-slot → on-disk-index map a **lines** ladder does not publish, so across a lines ladder the hover only lands on the right string when every element carries the same one (there is no broadcast label form — `labels` is always one entry per element). #1439 carried that map over the levels for Points only; a laddered Lines node is still open. **Builder API (Python):** ```python # Same kwarg surface across all three leaf types. scene.add_points("pts", positions, additive_lod=True) # 4-level default scene.add_lines("ln", verts, widths, line_type="segments", # explicit dict additive_lod=dict(n_lods=3, method="salience")) scene.add_gsplats_from_data("splats", data, # gsplats kwarg additive_lod=dict(n_lods=4)) ``` Composition with `partition=` is "partition-outer, additive-LOD-inner": each spatial part gets its own LOD ladder. **Points and Lines only.** On gsplats the combination is refused: the writer every laddered gsplats leaf goes through (`add_gsplats_multi_lod_impl`) has no partition step at all, so `add_gsplats_from_data(..., additive_lod=..., partition=...)` raises rather than silently dropping one of the two, and so does `add_gsplats_from_file` on a store whose leaves already carry ladders (`gsplat lod --recipe tiles|overview|adaptive`, `batch-fit merge --recipe stream`). Collapse those with `gsplat flatten` if you need to re-partition them — or pass `additive_lod=False`, which flattens each level to a single sub-LOD in memory and partitions normally, on either door. Partitioned gsplats ladders are still *readable*: the paragraph above describes what the pipeline itself produces, which the viewer resolves the same way it resolves a Points one. ### 3. Points Nodes Points nodes contain the actual point data. **Attributes (.zattrs):** ```javascript { "type": "points", "transform": [1,0,0,0, 0,1,0,0, 0,0,1,0, 0,0,0,1], "nd_transform": { // Optional, per-dimension transforms "Time": {"scale": 0.001, "offset": 50.0} // for non-displayed dimensions }, "opacity": 1.0, "absorption": 1.0, // volumetric mode's kappa (>= 0, default 1.0; multiplicative) "gamma": 1.0, "intensity": 1.0, "offset": 0.0, "blending_mode": "additive", // or "normal", "max", "opaque", "luminous", "volumetric" — // written only when explicitly set (unset ⇒ inherited) "layer": false, // Optional: if true, node appears in the viewer's Layers panel "visible": true, // Optional: initial visibility when the scene loads (default true) "n_points": 10000, "max_radius": 2.5, "extend_to_all": ["Time", "Channel"] // Optional: extend visibility to all values of these dimensions } ``` The on-disk `extend_to_all` value is **always a resolved list of dimension names**, never the `"all"` sentinel the Python `add_*` API accepts: the sentinel is expanded to the scene's non-displayed dimension names before anything is written (including onto each `additive_/` sub-LOD of a laddered node). The viewer relies on that invariant and reads the attr as a `string[]`. **Data Arrays:** The dtypes below describe the default `EncodingMode.AUTO`. Every array self-describes its on-disk encoding via an `encoding` attr in its `.zattrs` (see `luxar.encoding` and *Array Encodings* below); readers dispatch on `encoding.name` and decode to float32 (or the array's original integer dtype, e.g. uint8 colors stay uint8). `PRECISION` stores raw float32 everywhere; `MEMORY` quantizes more aggressively for the wide-range geolog family (8-bit where AUTO uses 16-bit — HDR colors, wide-range positive scalars); coordinates stay uint16 and bounded scalars pick 8-vs-16 bits from their dynamic range identically in both modes. Uniform arrays are stored as a single `broadcasted` value and byte-identical duplicates as an `array_ref`, regardless of mode. Chunking is **byte-based**, not a fixed element count: the first-dimension chunk length is derived from the 64 KB target (`TARGET_CHUNK_BYTES` ÷ bytes-per-row for the array's *input* dtype — computed before encoding, so float32 rows even when the stored code is uint8/uint16). When spatial ordering is enabled (the default), each per-element array's first-axis chunk is sized to its own dtype byte budget rounded down to a **multiple** of the spatial index's `chunk_size` atom (never below one atom) — so a chunk-index range always falls inside a whole zarr chunk, and a large scene issues far fewer requests because most arrays pack several index chunks per zarr chunk. This applies to Points, Lines and GSplats arrays alike, and to the *standalone* `.gsplats.zarr` tree writer too — it shares the same writer, which is what keeps a scene leaf and a standalone leaf byte-identical (see `tests/test_scene_leaf_parity.py`). GSPLATS_ZARR_FORMAT.md §8 specifies the same per-array sizing. #### positions/ (Required) - **Shape:** `(N, D)` where N = number of points, D = dimensionality - **Dtype/Encoding:** `uint16` per-axis fixed-point (`linear_perchannel_u16`) under AUTO/MEMORY — each axis quantized over its own `[min, max]` to 65536 levels, decoded back to float32 on read (visually lossless, ~2× smaller). `float32` under PRECISION, or when a per-axis extent ≥ 2¹⁶ forces the float32 fallback (uint16 could no longer resolve a unit step). - **Chunks:** `(chunk_rows, D)` — byte-based / spatial-index-aligned (see above) - **Compression:** Blosc with zstd, level 9 (width-aware shuffle policy) - **Description:** Point positions in D-dimensional space (never broadcast) #### colors/ (Optional) - **Shape:** `(N, 3)` for RGB or `(N, 4)` for RGBA - **Dtype:** `uint16` (`geolog_perchannel_u16`, HDR default) / `uint8` (SDR `rgb_uint8`, or HDR under MEMORY) / `float32` (PRECISION). HDR colors are quantized per channel on a true-log grid (uniform relative precision, code 0 reserved for exact zeros) and decoded back to float32. - **Chunks:** `(chunk_rows, 3|4)` — byte-based / spatial-index-aligned - **Compression:** Blosc with zstd, level 9 (width-aware shuffle policy) - **Description:** HDR RGB colors in normalized range - **SDR Range:** 0.0-1.0 (standard dynamic range) - **HDR Range:** Values > 1.0 represent HDR brightness - **Typical HDR:** 0.0-10.0 (extreme brightness) - **Note:** Values are NOT in 0-255 range; use 0.0-1.0 for normal colors - **Alpha (optional 4th channel):** per-point opacity α ∈ [0, 1] (never HDR; the SDR/HDR autodetect scans RGB only). Every blending mode scales a point's contribution by α; `volumetric` maps it into optical depth w = −ln(1−α) — see VOLUMETRIC_BLENDING_SPEC.md §5.4.1. No format-version bump: readers key off the array shape, and codecs are channel-agnostic. - **Default:** White (1.0, 1.0, 1.0) if not provided #### radii/ (Optional) - **Shape:** `(N,)` - **Dtype/Encoding:** POSITIVE_SCALAR — under AUTO, quantized to `bounded_scalar_uint8`/`bounded_scalar_uint16` (rescale-first, anchored at the array's own `[min, max]`) or, for wide dynamic range (> 65536:1), `geolog_scalar_uint16` (geometric-log grid, code 0 reserved for exact zeros). `float32` under PRECISION; `broadcasted` when uniform. - **Deduplication:** `false` — chunk bounds depend on this array's own positive-scalar quantization grid. - **Chunks:** `(chunk_rows,)` — byte-based / spatial-index-aligned - **Compression:** Blosc with zstd, level 9 (width-aware shuffle policy) - **Description:** Point radii in scene units - **Default:** 0.5 if not provided (`DEFAULT_POINT_RADIUS` in `typing_utils/constants.py`, mirrored in the viewer's `config/constants.ts`). The same value is what the viewer draws a radii-less node with and what the spatial index expands a no-radii chunk's bounds by. - **Validation:** All values must be positive - **Shader contract:** radii do NOT scale with the node's `transform` — a node-level scale repositions point centers but leaves the rendered disc size unchanged (size attributes are applied after the model transform). This is deliberate and shared with Lines `widths/`; gsplats differ (their covariances transform with the node). Bake the desired world size into the radii themselves when scaling a node. #### sharpnesses/ (Optional) - **Shape:** `(N,)` - **Dtype/Encoding:** BOUNDED_SCALAR with fixed bounds `(0.0, 1.0)` — under AUTO, quantized to `bounded_scalar_uint8`/`uint16`; `float32` under PRECISION; `broadcasted` when uniform. - **Chunks:** `(chunk_rows,)` — byte-based / spatial-index-aligned - **Compression:** Blosc with zstd, level 9 (width-aware shuffle policy) - **Description:** Point edge sharpness — a normalised `[0, 1]` knob. The viewer maps it to the super-Gaussian falloff exponent `β = 2^(6s − 2)`: `s = 0.5 → β = 2` (a true Gaussian), higher `s` → harder/crisper edge (β up to 16), lower `s` → peakier cusp (β down to 0.25). - **Default:** 0.5 (→ β = 2, Gaussian) if not provided - **Validation:** All values must be in `[0, 1]` ### 4. Lines Nodes Lines nodes contain polyline/segment data. All four user-facing line types (`segments`, `polyline`, `loop`, `indexed`) are converted to a unified **indexed representation** at write time — a `segments` array of vertex-index pairs — so the on-disk layout is identical for every type; the user's original choice is recorded in `original_line_type`. Joint continuity is defined by shared **indices**, not equal coordinates. Two segment endpoints stored as separate vertex rows remain independent even when their coordinates match, so connected thick curves should use `polyline` or `indexed` authoring with every joint referenced through one shared vertex row. **Attributes (.zattrs):** ```javascript { "type": "lines", "n_vertices": 10000, "n_segments": 9999, "ndim": 3, "original_line_type": "polyline", // "segments" | "polyline" | "loop" | "indexed" "has_colors": true, "has_sharpness": false, "max_width": 1.5, "position_bounds": {"min": [...], "max": [...]}, "ordering": "hilbert", // or "morton" / "none" "vertex_ordering": { ... }, // Vertex spatial-index metadata (D-space) "segment_ordering": { ... }, // Segment spatial-index metadata (2×D-space) "join": "miter", // Optional, LINES ONLY: join style at // degree-2 polyline joints — // "miter" (default) or "none". // Unset ⇒ inherited, then the viewer // default. Compositing: set it on the // layer, not on internal children. /* transform, nd_transform, opacity, absorption, gamma, intensity, offset, blending_mode, layer, visible — same as Points */ } ``` `join` selects what the vertex stage does where two segments of a polyline meet. Without join geometry the turn leaves an uncovered circular sector on the outside of the bend and a double-covered lens inside — dark ticks along the convex edge of a thick curve, bright ticks along the concave one. `"miter"` rotates each quad's end edge onto the shared miter edge so the two TILE: coverage becomes a partition, so there is nothing to sum and every blending mode is correct by construction. It is gated in-shader by a rendered-width threshold (a sub-pixel wedge is invisible, and thin lines are the million-segment ones) and by a miter limit, so `"none"` is rarely worth authoring. A session-wide `?lineJoin=none|miter` overrides whatever the file says. Unrecognised values are rejected at write time rather than silently meaning "no joins". Lines use **dual spatial indexing**: vertices are curve-ordered in D-space (like Points) and segments are independently curve-ordered in (2×D)-space (concatenating both endpoints), each with its own chunk-bounds array (`vertex_chunk_bounds` / `segment_chunk_bounds`, both `(num_chunks, D, 2)` float32 — segment *bounds* are deliberately D-space even though the segment *ordering* sorts in 2×D, so both support view-frustum intersection tests directly). **Data Arrays** (same AUTO/PRECISION/MEMORY conventions as Points; all per-vertex arrays are reordered by the vertex sort): #### vertices/ (Required) - **Shape:** `(N, D)` — vertex positions - **Dtype/Encoding:** COORDINATE, same as Points `positions/`: `linear_perchannel_u16` under AUTO/MEMORY (float32 under PRECISION or the ≥ 2¹⁶-extent fallback). Never deduplicated to an `array_ref` and never LUT-encoded — the lines spatial-index loader reads it as raw chunked zarr with no structural-encoding dispatch (grid-snapped vertices would otherwise store as LUT indices). #### segments/ (Required, auto-generated) - **Shape:** `(M, 2)` — vertex-index pairs, indices local to this node - **Dtype/Encoding:** INDEX — stored as the smallest unsigned integer dtype that fits the max index (`uint8`/`uint16`/`uint32`), read raw (never LUT-encoded or deduplicated). #### widths/ (Required) - **Shape:** `(N,)` — per-vertex line widths (scene units, must be positive) - **Dtype/Encoding:** POSITIVE_SCALAR, same rules as Points `radii/` (`bounded_scalar_uint8/16` or `geolog_scalar_uint16` under AUTO; float32 under PRECISION; `broadcasted` when a scalar width is given). - **Shader contract:** like Points `radii/`, widths do NOT scale with the node's `transform` — a node-level scale repositions vertices but leaves the rendered line width unchanged. #### colors/ (Optional) - **Shape:** `(N, 3)` — per-vertex RGB, same COLOR encoding rules as Points (SDR → `rgb_uint8`; HDR → `geolog_perchannel_u16` under AUTO). #### sharpnesses/ (Optional) - **Shape:** `(N,)` — per-vertex edge sharpness, same BOUNDED_SCALAR `[0, 1]` rules and semantics as Points `sharpnesses/`. #### scalars/ (Optional) - **Shape:** `(N,)` — **per-vertex** colormap scalars (matching `widths`, not per-segment); declared via `has_scalars` / `scalar_data_range` / `colormap` attrs (see *Scalar Colormap Attributes* below). Per-vertex labels (`label_offsets`/`label_bytes`), keys (`key_offsets`/`key_bytes`), and image labels (`image_label_offsets`/`image_label_bytes`) are supported with the same CSR-style layout as Points (see *Per-Element Labels* and *Per-Element Keys*). Because the string channels are per-vertex while the viewer picks whole *segments*, hover and selection on a lines node report the picked segment's **start** vertex. Two consequences follow from that convention: on a segment that the current slice clips only partially the reported start vertex may lie entirely outside the visible slab (what is drawn starts at the clipped position, not at the stored vertex), and **any vertex that is never a segment's start is unreachable by hovering** — its label, key, or image label can never be read. Which vertices those are depends on `original_line_type` (the segment pairs are built by `luxar.io._ordering.lines.convert_to_indexed`): - **`segments`** — the pairs are consecutive disjoint vertices `(0,1)`, `(2,3)`, …, so **every odd-numbered vertex** (in authored order) is only ever an end: half of the label array is unreachable. Author the label you want shown on the even-numbered vertex of each pair. - **`polyline`** — the pairs are `(0,1)`, `(1,2)`, …, so only the final vertex is unreachable. - **`loop`** — the last vertex connects back to the first, so every vertex starts a segment and all labels are reachable. - **`indexed`** — whichever subset the supplied `indices` never place first, plus any vertex no segment references at all. ### 5. Mesh Nodes Mesh nodes contain triangle-surface data — isosurfaces, segmentation boundaries, organ and cortical meshes. They are the only node type that describes a *connected, opaque surface* rather than a set of soft per-element primitives. ✅ **Writable and renderable.** Mesh nodes are written, read and reported by `luxar info`, and the viewer loads, shades and picks them too — the whole vertical ships today (see `docs/specs/MESH_NODE_SPEC.md` §11). The format contract still distinguishes the two: `geometry_types` (the writable leaf vocabulary) and `loader_types` (the viewer-drawable subset) — and both now include `mesh`. Two structural differences from the other three types: - **No per-element size.** A triangle's extent comes from its own vertices, so there is no `radii` / `widths` / `cholesky_factors` analogue — and a mesh contributes **zero extent padding** to `position_bounds`. - **Topology is load-bearing.** `faces` is an index array, so unlike a coordinate or colour array it cannot tolerate lossy or aliasing encoding: it is written with dedup and LUT encoding **disabled** (see below). **Attributes (.zattrs):** ```javascript { "type": "mesh", "n_vertices": 10000, "n_faces": 19996, "ndim": 3, "has_normals": true, "normal_dims": [0, 1, 2], // REQUIRED iff has_normals — see below "has_colors": true, "has_scalars": false, "has_uvs": true, "has_texture": true, "texture_encoding": "raw", // "raw" | "png" | "webp" | "jpeg" | "ktx2" "texture_width": 2048, "texture_height": 1024, "texture_channels": 3, "texture_color_space": "srgb", // "srgb" | "linear" "shading": "smooth", // "smooth" | "flat" | "none" "double_sided": true, "position_bounds": {"min": [...], "max": [...]}, "ordering": "none", // always "none" in v1 (no spatial index) // ... plus the standard render attrs (opacity, gamma, intensity, offset, // absorption, blending_mode, colormap, layer, transform, nd_transform, // extend_to_all) and mesh-only appearance attrs (ambient, shade_exponent, // specular, shininess, alpha_cutoff, texture_filter, texture_wrap), // plus slab_tolerance, plus the opt-in material family: "material": "physical", // "luxar" (default when absent) | "physical" "roughness": 0.4, // physical knobs, each in [0, 1], written "metalness": 1.0, // only when authored; absent = three's "clearcoat": 1.0, // own default "clearcoat_roughness": 0.1, "iridescence": 0.0, "sheen": 0.0, "sheen_color": "#ffffff", // "#rrggbb" "transmission": 1.0, // the glass family: [0, 1] "ior": 1.5, // [1, 2.333] "thickness": 0.4, // >= 0, scene units "attenuation_color": "#f6d148", // "#rrggbb" "attenuation_distance": 0.3, // > 0, scene units; absent = none "dispersion": 0.5, // >= 0 "refract_data": true // Phase 3: draw after, and refract, the emissive data } ``` The seven house-shader appearance attrs control shading and texture sampling; `slab_tolerance` controls nD membership loading; `material` selects the material family and unlocks the fourteen physically based knobs (`docs/guides/specs/MESH_PHYSICAL_MATERIALS_SPEC.md`). All twenty-three mesh-only authored attrs are rejected on points, lines, Gaussian splats, and groups. A physical knob without `material: "physical"` is refused at authoring; a physical mesh refuses `ambient` / `shade_exponent` / `specular` / `shininess`, `blending_mode`, `colormap`, a texture and `shading: "none"`; and `thickness`, `attenuation_color`, `attenuation_distance`, `dispersion` and `refract_data` are refused without a `transmission` above zero — none of these pairings has a meaning. `refract_data` (Phase 3) makes a glass draw after, and refract, the points, lines and splats behind it, while data in front of it stays crisp on top (the viewer partitions each data fragment by depth against the glass). #### vertices/ (Required) - **Shape:** `(V, D)` — nD vertex positions, exactly like `Lines.vertices`. - **Encoding:** `COORDINATE`, written with `deduplicate=false` and `allow_lut=false`. The loader reads it as raw chunked zarr and does not resolve `array_ref`, so dedup would silently drop geometry for a byte-identical sibling, and LUT encoding of grid-snapped coordinates would decode as garbage. #### faces/ (Required) - **Shape:** `(F, 3)` — triangle vertex indices. - **On-disk dtype:** ⚠️ **any unsigned integer width — a reader must not assume `uint32`.** `uint32` is the writer's canonical logical dtype, but the `INDEX` encoder narrows integer arrays losslessly by observed value range, so a 4-vertex mesh stores `uint8` and a 60k-vertex one `uint16`. The narrowing is reversible (the original dtype is recorded in the array's `encoding` attr) and verified exact in every encoding mode, but a consumer that hardcodes `uint32` will reject valid stores. Accept any integer dtype and widen on read. - **Encoding:** `INDEX`, `deduplicate=false`, `allow_lut=false` — same reasoning as `Lines.segments`. Note `allow_lut=false` matters here beyond the raw-read argument: grid-structured index values are exactly what LUT encoding targets, so without it a regular mesh would be a prime candidate for it. - **Winding:** counter-clockwise as seen with the mesh's authored spatial triple in ascending index order. For a 3D mesh that triple is `[0,1,2]`; for an nD mesh it is `sorted(normal_dims)` when normals are present. No winding can be counter-clockwise under *every* 3D projection of an nD mesh, so the viewer restores front-facing winding only when the displayed set matches that frame. - Writers must keep every index in `[0, n_vertices)`. An out-of-range index is not a rendering artefact: it reads past the vertex buffer, which aborts the whole WASM module in the Rust culling kernel. #### normals/ (Optional) - **Shape:** `(V, 3)` — **always** 3-component, even for an nD mesh. - **Encoding:** `COORDINATE` (per-axis `uint16` over each component's own `[-1, 1]` range — a free 2× over float32 — and it correctly blocks broadcasting, since a normal is always per-vertex). - **Paired with a required `normal_dims` attr.** Normals are a *display-space* quantity, meaningful only for the three displayed dimensions, so the store must record which three they describe. ⚠️ Do **not** store normals against an implicit "first three dimensions": for a `(t, x, y, z)` mesh those are `(t, x, y)` and such a normal is meaningless. The viewer uses stored normals only when `shading == "smooth"` **and** `normal_dims` equals the active `displayDims`, and otherwise computes flat face normals from the projected triangle — so a wrong-but-well-formed triple degrades shading rather than corrupting it, while an ill-formed one is rejected at write time. - Zero-length normals are **warned about, not rejected**: degenerate triangles legitimately produce them and the renderer epsilon-guards its `normalize`. #### colors/ (Optional) - **Shape:** `(V, 3)` or `(V, 4)` — per-vertex RGB or RGBA, same encoding rules as Points. The optional 4th component is per-vertex **opacity**. #### scalars/ (Optional) - **Shape:** `(V,)` — per-vertex colormap scalars; declared via `has_scalars` / `scalar_data_range` / `colormap` (see *Scalar Colormap Attributes* below). Per-vertex labels (`label_offsets`/`label_bytes`), keys (`key_offsets`/`key_bytes`), and image labels (`image_label_offsets`/`image_label_bytes`) use the same CSR-style layout as Points (see *Per-Element Labels* and *Per-Element Keys*). **Not written for a mesh node:** no spatial index (`ordering` is always `"none"`). #### nD slicing: whole-triangle cull A mesh slices differently from the other three geometry types, and the difference is visible. Points, Lines and GSplats are collections of independent elements, so slicing keeps or drops each element on its own — and Lines goes further, *clipping* a segment that straddles the slice and interpolating its attributes at the cut. A triangle cannot be handled that way cheaply: cutting one against an nD slab yields a polygon that has to be re-triangulated, with new vertices and interpolated attributes, every frame the slice moves. Luxar does not do that. The rule is: > A triangle is drawn **iff all three of its vertices** fall inside the slice slab. Two consequences follow, and both are worth knowing before you author a mesh with hidden dimensions. **A cut surface has a ragged edge.** Because whole triangles are kept or dropped, the boundary follows triangle edges rather than the slice plane. On a well-tessellated surface sliced with a slab comparable to its edge length this reads as a slightly jagged edge. On a *coarse* mesh with a thin slab it can drop whole regions — if no triangle has all three vertices inside, nothing is drawn. **On a continuous hidden dimension you get a slab, not a section.** There is no interpolation, so there is no such thing as an exact cross-section: what you see is "the surface near this slice", of finite thickness. The viewer says so once per node, by name, in an `info` log line. `slab_tolerance` is the control over that thickness: ```python scene.add_mesh( "surface", vertices, faces, normals=normals, normal_dims=[0, 1, 2], slab_tolerance=2.5, # vertices within ±2.5 cells (a 5-cell slab); default 1.0 ) ``` It is a half-width measured in cells of the hidden dimension's own `step`, must be strictly positive, and defaults to one cell. Raising it thickens the slab (more surface shown, more of it away from the slice); lowering it thins the slab toward the degenerate case above. It applies **only** to *continuous* hidden dimensions — a *discrete* one (time, channel, or any axis with `categories`) uses a half-cell membership rule instead and ignores the attr. Discrete hidden dimensions are the dominant real case for a mesh, and they have none of the problems in this section: a timepoint either matches or it does not. Mesh is the only geometry type whose slab is tunable, and the reason is that it has nothing to measure. The other three derive their tolerance from a per-element extent — a point's `radii`, a line's `widths`, a splat's truncated `sigma` — that a mesh vertex simply does not have, so the thickness is chosen rather than read off the data. If your mesh's hidden dimensions are continuous *and* spatial, and you need a true planar section, a mesh node is the wrong representation today — fit the volume as Gaussian splats instead, which slice exactly. Exact nD triangle clipping is a deliberate non-goal for now; see `docs/specs/MESH_NODE_SPEC.md` §5 and §9. **Additive sub-LOD subgroups (`additive_/`) ARE written, but only for a reveal.** A prefix of an index buffer is a holed surface, not a coarse one, so the ladder is not a level of detail for a mesh and `add_mesh(additive_lod=…)` accepts only `method="radial"` — a concentric-shell *reveal*, whose every prefix is a contiguous partial surface. Three consequences are visible on disk: each level re-indexes its own gathered vertex table (so the parent's `n_vertices` exceeds the source count by the boundary duplication, exactly as `kind=partition` parts do); no level carries `energy_fraction_cum` and the parent carries no `reference_energy`, because a reveal prefix is a partial object at full brightness rather than a dim version of the whole; and `has_labels`, `has_keys`, and `has_image_labels` are all **unset**, because one source vertex maps to a slot in every level that touches it, so a union annotation CSR spanning levels has no well-defined index space. Annotated meshes take `substitutive_lod=` or `partition=`, both of which keep their per-element annotations. A mesh **may** be a child of a `kind=partition` group; `add_mesh(partition=…)` writes exactly that, with each part carrying its own gathered-and-renumbered vertex table (vertices on a cut are duplicated between neighbouring parts). A mesh may equally be a child of a `kind=lod` group — `add_mesh(substitutive_lod=…)` writes that shape, with each coarse level a decimated copy of the surface. Each child stamps `level_stats.geometric_error`, the measured source-to-level collapse bound normalized by the source bounding-box diagonal; this is separate from mixture `quality` and carries no energy fields. The two cannot be combined in one call. ### 6. Sound Nodes Sound nodes hold an audio clip — an ambient bed, a narration bound to a story step, or a spatial source that gets louder as the camera approaches. They are the one node type that is **heard rather than drawn** (design: `docs/guides/specs/SOUND_SPEC.md`). `sound` is in the format contract's `node_types` but **not** in `geometry_types` or `loader_types`: it carries no elements, no blending mode, no LOD and no picking, so none of the geometry vocabularies list it. The viewer dispatches it to its audio engine by `type` alone. **Layout:** ``` /sounds/hum_hsp70/ # a group under any group, like other nodes zarr.json # type: "sound", attrs below (.zattrs at format 2) positions # (K, ndim) plain float32 (no quantizing encoder) — ABSENT for a clip live everywhere audio.mp3 # the clip, a plain store key (or audio.m4a for AAC) ``` **Attributes:** ```javascript { "type": "sound", "spatial": false, // true → K PositionalAudio voices at `positions` "trigger": "once", // "continuous" | "once" | "on_depart" | "on_arrive" "delay_ms": 800.0, // after the trigger fires "gain": 1.0, // per-node linear gain "bus": "voice", // "ambient" | "voice" | "effects" "loop": false, // derived: trigger == "continuous" "fade_in_ms": 0.0, "fade_out_ms": 0.0, "license": "CC0", // REQUIRED provenance, all three "attribution": "Freesound user X", "source_url": "https://…", "format": "mp3", // sniffed from the bytes: "mp3" | "aac" "audio_file": "audio.mp3", // the plain key holding the clip "duration_ms": 31240.0, // optional, stamped when a tag reader is available "channels": 2, // optional, same source; checked (== 4) for an ambisonic clip "attach_to": "Story 3: hsp70", // optional: follow that node's bounding-box centre "ambisonic": "foa", // optional: a 4-channel AmbiX FIELD (AAC only, never spatial) "has_positions": true, "n_positions": 1, "ndim": 4, "extend_to_all": ["time"], // as for points "ordering": "none", // never a spatial index "position_bounds": {...}, // spatial only; NOT folded into the scene bounds // spatial only — PannerNode knobs, absent = viewer default from the scene scale "distance_model": "inverse", "ref_distance": 2.0, "max_distance": 30.0, "rolloff": 1.0, "cone_inner_deg": 90, "cone_outer_deg": 180, "cone_outer_gain": 0.2, "orientation": [0, 0, -1], "layer": true, "visible": true, "transform": [...], "nd_transform": {...} } ``` **Audibility is the slab rule.** A sound node's `positions` rows are tested against the hidden-dimension slice exactly like a points node's: a row whose hidden coordinates fall inside the slab is live, others are silent, and `extend_to_all` makes a row live everywhere along a dimension. A node with no `positions` is live everywhere. The Python `hidden={"story": 3}` sugar writes one row at `story=3` (other columns 0) and `extend_to_all` over every other hidden dimension. **Triggers.** `continuous` and `once` follow the slab's edges. `on_depart` / `on_arrive` follow the waypoint driver's events (`viewer_config.waypoints`): the node fires when a story flight leaves / lands on the waypoint whose `when` clause its row satisfies (a node without rows belongs to every waypoint). An `on_arrive` clip still fades out when its story is left; an `on_depart` clip plays out. **`attach_to`.** The NAME of another node: the source follows that node's bounding-box centre in the viewer ("the cluster hums" without authoring coordinates). Spatial by default; combine with `hidden=` to bind it to a value. Mutually exclusive with `positions`. **`ambisonic: "foa"`.** The clip is a first-order ambisonic FIELD — four AmbiX channels (ACN order `W, Y, Z, X`, SN3D) — that the viewer decodes to stereo and rotates against the camera so the field stays fixed to the world. AAC only (MP3 holds two channels); never spatial and never positioned (`hidden=` still decides when it is live). The writer refuses the node when a tag reader reports a channel count other than 4. **Formats.** MP3 and AAC (`.m4a` / ADTS) are accepted and sniffed from the payload, never from the filename. Ogg/Opus is refused because Safari cannot decode it; WAV and FLAC are refused as the wrong size class for a hosted store. **Not written for a sound node:** no appearance attrs (`opacity`, `colormap`, `blending_mode`, … are refused by the writer), no spatial index, and no contribution to the root `position_bounds` — a source at the far corner of a dataset must not push the opening framing out. **`content_hash` covers the clip.** The bytes of the file named by `audio_file` are folded into the node's digest through the same payload step that covers an overlay's `image_file` (`io/_compiler/finalize/hashing.py::PAYLOAD_FILE_ATTRS`), and `luxar optimize` carries the file into a re-chunked store for the same reason. **Viewer-side defaults** live in `viewer_config.audio` (`AudioConfig`: `enabled`, `master_gain`, `panning_model` `"equalpower"` | `"HRTF"`, per-bus gains, and `duck_db` — the ambient attenuation while anything on the voice bus plays). ## Scalar Colormap Attributes For Points, Lines and Mesh, an optional per-element `scalars` zarr array enables colormap-driven shading. The presence and configuration are declared through three attrs on the data node (next to `type`, `opacity`, etc.): ```json { "type": "points", "has_scalars": true, "scalar_data_range": [0.0, 1.0], "colormap": "viridis" } ``` - **`has_scalars: bool`** — gates whether the loader opens the `scalars` zarr array. Set by the writer when a scalar array exists; ignored otherwise. - **`scalar_data_range: [min, max]`** — input range used to normalize scalars to `[0, 1]` before the LUT lookup. Required when `has_scalars` is true; defaults to `[0, 1]` if omitted. **`intensity`/`offset` on a colormapped node are the display window, not a gain.** When a colormap is active, an authored `intensity`/`offset` defines the scalar display *window* (the value→LUT mapping) exactly as the Layers panel's range control does — the post-LUT color gain stays at identity, so the value is never applied twice. A *non-identity* leaf-authored `intensity`/`offset` pair therefore *replaces* `scalar_data_range` as the window (`window = [-offset/intensity, (1 - offset)/intensity]`). The decision is by value, matching the Layers panel: an explicitly authored identity pair (`intensity: 1.0`, `offset: 0.0`) behaves exactly like an unauthored one and keeps the `scalar_data_range` window. An ancestor-only gain is instead folded onto the declared `scalar_data_range`. This keeps the load-time render identical to the post-interaction (Layers-panel) render, and mirrors how gsplat nodes treat their `amplitude_data_range`. A direct-color node with no colormap still treats `intensity`/`offset` as an ordinary post-shading gain. - **`colormap: 'viridis' | 'plasma' | ... | 'custom'`** — selects a built-in LUT (15+ available) or `'custom'` to enable a user-supplied LUT sibling array. **Custom LUT**: when `colormap = 'custom'`, the scene loader looks for a sibling array named `colormap_lut` (alongside the data node, not nested inside `scalars`). Shape: `[256, 3]` (RGB) or `[256, 4]` (RGBA), dtype `uint8`. The loader passes the raw bytes through to `getColormapTexture` which builds a 256×1 DataTexture; invalid lengths fall back to viridis with a warning. Custom LUT textures are cached per-app with a bounded LRU (16 entries) and disposed on scene unload (B.2 of the viewer-code-review-rerun hardening pass). **Scalar dtype**: the `scalars` zarr array may be `float32`, `float16`, or `uint8`. Uint8 scalars are kept in their native dtype through the loader and accumulator and widened to Float32 at the GPU upload boundary (A.2 of the same hardening pass). **Per-vertex Lines**: the Lines `scalars` array is per-vertex (matching `widths`), not per-segment. Projection interpolates between endpoints at clipped boundaries so the LUT lookup at a slice edge uses the correct value. ## Environment (baked scene lighting) A `material: "physical"` mesh is lit by the viewer's scene environment (`docs/guides/specs/MESH_PHYSICAL_MATERIALS_SPEC.md` §3.3). Two root-level things describe it: **`viewer_config.environment`** (optional; Python `EnvironmentConfig`): ```json "environment": { "source": "scene", // "room" (default) | "scene" | "hdri" "probe": "auto", // "auto" | "node:" | [x, y, z] "resolution": 128, // cube face size for a "scene" capture, 16-1024 "intensity": 1.0, // scene.environmentIntensity, >= 0 "url": "env/studio.hdr" // "hdri" only; store-relative or absolute HTTP(S) } ``` `"scene"` makes the viewer capture the environment from the scene itself (an exact `CubeCamera` render from the probe), so metals and glass reflect the data they sit in. House-shaded meshes, points, lines and splats never read the environment, so the block changes nothing about them. An HDRI URL is limited to 2048 characters; absolute URLs must use HTTP(S) and must not contain embedded credentials. Protocol-relative and active-content schemes such as `data:` or `javascript:` are refused. **`environment/` group** (optional; written by `luxar env bake` / `luxar env attach`, never by the compiler): the six captured cube faces, prefiltered by the viewer at load in milliseconds so a published scene pays no live capture. ``` environment/ # a SIDECAR: no `type`, no `kind` attr ├── zarr.json # attrs, below └── faces-3f9a1c02/ # (6, H, W, 4) uint16 — IEEE half-float bits, RGBA ``` ```json { "format": "cube-faces-half", "faces": "faces-3f9a1c02", // the LIVE array; a re-bake is a NEW name "sample_format": "half-float-bits", "shape": [6, 128, 128, 4], "face_order": ["px", "nx", "py", "ny", "pz", "nz"], // three's CubeTexture order "coordinate_system": "webgl", // the backend that captured it "probe": {"spec": "auto", "position": [0.0, 0.0, 0.0]}, "resolution": 128, "scene_content_hash": "…", // the root content_hash it was baked against "appearance": {"viewer_config": {}}, "baked_at": "2026-09-06T00:00:00Z", "viewer_version": "…", "content_hash": "…" // the GROUP's own digest (tooling only) } ``` Three rules make it safe. The group carries neither `type` nor `kind`, so the viewer's node discovery skips it as a metadata sidecar; `LuxarScene.nodes` and `luxar info` consult `RESERVED_ROOT_GROUPS`, while `luxar optimize` and scene hashing skip `ENVIRONMENT_GROUP` directly. The compiler refuses a user node named `environment`. The group is **excluded from the scene `content_hash`**, so attaching a map never changes the root digest: the `scene_content_hash` guard is exact (the viewer ignores a map whose digest is not the root's, saying so in the console), a visitor's warm cache survives a bake, and attaching the same map twice writes nothing. And the faces array is named by its own digest, so a re-bake is a new path a caching viewer cannot serve stale. `uint16` rather than `float16` because the viewer's zarr reader needs a `Float16Array` for `/` directly. A case-insensitive static host may therefore return the real metadata document for a dangling case-shifted name even though the compiler hashes it as absent; the result is a broken image response, not a valid overlay. Producers should avoid all payload names that collide case-insensitively with zarr metadata documents. Note: the overlay `blend_mode` is a screen-space-overlay compositing concept (how the 2D overlay image blends over the rendered frame) — distinct from the scene-node attribute `blending_mode` that controls 3D geometry blending. **Video overlay** (`overlay_video`): ```json { "type": "overlay_video", "position": [0.06, 0.5], "anchor": "center-left", "video_file": "video.webm", "poster_file": "poster.png", "size": [0.26, null], "loop": true, "autoplay": true, "muted": true, "playback_rate": 1.0, "blend_mode": "normal", "z_index": 3 } ``` The clip is stored verbatim beside the overlay as `video.webm` or `video.mp4` (the compiler sniffs the container — an EBML header or an `ftyp` box — and refuses anything else); an optional `poster_file` follows the image-overlay payload rules. A `null` height in `size` keeps the clip's own aspect ratio. `autoplay` requires `muted` (browsers block un-muted autoplay), and the viewer plays a clip only while its `visible_range` matches, pausing it otherwise. Transparency travels as a **stacked alpha matte**, `"alpha_matte": "stacked"`: the frame is the colour on top and the alpha channel as a grey matte of the same size below — one ordinary opaque clip twice as tall (ffmpeg `split[c][a];[a]alphaextract[a];[c][a]vstack`) — and the viewer recombines the halves in a shader onto a canvas, so the clip is transparent in every browser, Safari and WKWebView included (the exported native app). A VP9 WebM with an alpha plane (`yuva420p`) is still accepted as a plain clip, but only Chrome and Firefox render its alpha; Safari decodes it and drops the alpha, showing the clip on a black square. Without WebGL the viewer shows a stacked clip as is (colour over matte). **HTML overlay** (`overlay_html`): ```json { "type": "overlay_html", "position": [0.01, 0.5], "html": "

Some formatted text

", "width": 0.3, "interactive": true, "z_index": 2 } ``` Note: the viewer sanitizes `html` against both a tag allowlist and an attribute allowlist at render time, so hand-authored values are still constrained. A tag outside the allowlist is unwrapped — it disappears while its children are kept (the one exception is `