I/O Package

The io package handles reading and writing Luxar scenes to Zarr format with spatial indexing.

Input/Output operations for Luxar data.

Overview

The I/O package provides two main classes and a filterable capacity warning:

  • LuxarZarrCompiler: Progressive writer for creating Zarr datasets

  • LuxarScene: Reader for loading and querying Zarr datasets

  • ElementCapacityWarning: An authored node may exceed the viewer’s conservative per-node element-texture capacity

Both classes work together to enable efficient streaming of massive n-dimensional scenes with points, lines, Gaussian splats, and other primitives.

exception luxar.io.ElementCapacityWarning[source]

A node may exceed the viewer’s per-node element-texture capacity.

LuxarZarrCompiler

The compiler writes data progressively to Zarr format without caching in memory, enabling creation of TB-scale datasets on GB-scale machines.

class luxar.io.LuxarZarrCompiler(store_path: Optional[Union[str, Path]]=None, compressor: CompressorLike = <width-aware default compressor>, version: str = '0.2', enable_spatial_index: bool = True, encoding_mode: EncodingMode = EncodingMode.AUTO, ordering_method: SpatialOrderingMethod = 'hilbert', float16_allowed: bool = False, auto_partition_max_elements: int | None = None, gsplat_amplitude_bits: Optional[Literal['auto', 8, 16]]=None)[source]

Bases: ZarrWriterProtocol

Progressive Zarr compiler with context manager support.

This compiler writes data immediately to Zarr without keeping it in memory, enabling processing of datasets larger than available RAM.

Parameters:
  • store_path – Path where the Zarr store will be created. A .zip path normalizes the inner store name and publishes a single-file archive after finalization.

  • compressor – Compression configuration for datasets

  • version – Luxar format version

  • enable_spatial_index – Whether to build spatial indices for points

  • encoding_mode – Encoding mode for array storage

  • gsplat_amplitude_bits – Optional AUTO amplitude tier for GSplats only

Examples

Basic usage with context manager: >>> dims = Dimensions.default_3d() >>> with LuxarZarrCompiler(‘output.luxar.zarr’) as compiler: … scene = compiler.create_scene(dimensions=dims) … positions = np.random.randn(10000, 3).astype(np.float32) … scene.add_points(‘points’, positions)

Compile directly to a single-file archive: >>> with LuxarZarrCompiler(‘output.luxar.zarr.zip’) as compiler: … scene = compiler.create_scene(dimensions=Dimensions.default_3d()) … scene.add_points(‘points’, positions)

With HDR colors and custom dimensions: >>> dims = Dimensions([ … Dimension(‘x’, unit=’um’), … Dimension(‘y’, unit=’um’), … Dimension(‘z’, unit=’um’) … ]) >>> with LuxarZarrCompiler(‘scene.luxar.zarr’) as compiler: … scene = compiler.create_scene(dimensions=dims) … # HDR colors with values > 1.0 … colors = np.random.rand(1000, 3).astype(np.float32) * 5.0 … scene.add_points(‘bright_points’, positions, colors=colors)

Large datasets (partitioned into multiple nodes): >>> dims = Dimensions.default_3d() >>> with LuxarZarrCompiler(‘huge.luxar.zarr’, ordering_method=”hilbert”) as compiler: … scene = compiler.create_scene(dimensions=dims) … # Process chunks one at a time, each becomes a separate node … for i in range(100): … chunk_positions, chunk_colors = load_chunk(i) # 10M points … scene.add_points(f’chunk_{i}’, chunk_positions, colors=chunk_colors) … # Each node: sorted, has chunk_bounds, memory freed after write

__init__(store_path: Optional[Union[str, Path]]=None, compressor: CompressorLike = <width-aware default compressor>, version: str = '0.2', enable_spatial_index: bool = True, encoding_mode: EncodingMode = EncodingMode.AUTO, ordering_method: SpatialOrderingMethod = 'hilbert', float16_allowed: bool = False, auto_partition_max_elements: int | None = None, gsplat_amplitude_bits: Optional[Literal['auto', 8, 16]]=None) → None[source]

Initialize the Zarr compiler.

Parameters:
  • store_path – Path for the Zarr store, or None for temporary. A .zip path normalizes the inner store name and publishes a single-file archive after finalization.

  • compressor – Compressor for datasets

  • version – Luxar format version

  • enable_spatial_index – Whether to use spatial ordering for points/gsplats (default: True)

  • encoding_mode – Encoding mode for array storage (AUTO/PRECISION/MEMORY)

  • ordering_method – Spatial ordering method (“morton” or “hilbert”, default: “hilbert”)

  • float16_allowed – Allow float16 encoding in MEMORY mode (default: False for TypeScript compatibility)

  • auto_partition_max_elements – If set, add_points and add_gsplats automatically apply partition=dict(max_elements=N) when the input element count exceeds N. User-explicit partition= at the call site always wins. Default None (opt-in, no auto-partition). Useful for large datasets where you want a per-part frustum-cull benefit without explicit per-call boilerplate.

  • gsplat_amplitude_bits – Optional AUTO amplitude tier for GSplats only. 8 selects uint8 geometric-log amplitudes when the adaptive encoder would otherwise use more than 8 bits; 16 retains the historical default. "auto" resolves each GSplat node from its recorded fitting/source_dtype. Points radii and Lines widths are unaffected.

Note

Physical units should be specified per-dimension using the Dimensions system when calling create_scene(dimensions=…). This provides more flexibility for multi-dimensional data.

__enter__() → LuxarZarrCompiler[source]

Enter context manager.

__exit__(exc_type: Any, exc_val: Any, exc_tb: Any) → None[source]

Exit context manager, finalizing only on a clean exit.

If an exception is propagating out of the with block (a write error, or Ctrl-C / KeyboardInterrupt) and the store was NOT already finalized, the store is left UNfinalized and a root incomplete marker is stamped so a directory artifact is detectable; archive staging is discarded instead and the compiler becomes unusable. Finalizing here would seal a partial store as a valid, hash-stamped scene (the root type='scene' attr is written up front), silently corrupting downstream consumers. The marker is stamped only when the store was not already finalized — if the body finalized explicitly (e.g. Scene.to_zarr()) and an unrelated exception then raised, the complete store is left untouched. The marker write is best-effort so it can never mask the original exception. The temporary directory (if any) is cleaned up on every path.

create_scene(dimensions: Dimensions, viewer_config: 'ViewerConfig' | None = None, citation: Mapping[str, str] | None = None) → Scene[source]

Create a scene with this compiler as writer.

Parameters:
  • dimensions – Dimension specification for the scene (REQUIRED). Scene dimensions are the single source of truth for the coordinate system and must always be specified.

  • viewer_config – Optional viewer configuration hints. Stored in the zarr file and read by the viewer at load time as scene-specific defaults.

  • citation – Optional credit for whoever produced the underlying dataset – {"short", "ref"?, "doi"?, "license"?, "url"?}. Stored in the root attributes so it travels with the data. None means there is no external dataset to credit.

Returns:

Scene object configured with this compiler as writer

Raises:

ValueError – If dimensions is None or invalid

write_group(path: str, *, _transform_normalized: bool = False, **attrs: Any) → None[source]

Create a group in the Zarr store.

Parameters:
  • path – Path for the group within the store

  • _transform_normalized – Internal control parameter (part of the writer protocol, never persisted). The Node/Scene API pre-normalizes transform/nd_transform (its attrs cache must hold the column-major form for the transform getter) and passes True so an already column-major matrix is not transposed a second time. Raw compiler-API callers leave the default False and go through the normalization gate below (issue #678).

  • **attrs – Attributes to attach to the group

delete_group_attr(path: str, key: str) → None[source]

Remove an attribute from a group in the Zarr store.

Parameters:
  • path – Path for the group within the store

  • key – Attribute key to remove

node_exists(path: str) → bool[source]

Return whether a node path exists for an internal rollback guard.

delete_node(path: str) → None[source]

Delete a subtree during rollback, without editing the scene graph.

snapshot_rollback_state() → Tuple[dict[str, List[float]] | None, frozenset[Tuple[str, str]], bool, Tuple[dict[tuple, tuple[str, str]], dict[str, str]], Tuple[bool, bool] | None][source]

Capture mutable authoring state changed by geometry writes.

The write-only metadata cache is intentionally excluded: it has no readers and cannot affect later output after its subtree is deleted.

restore_rollback_state(state: Tuple[dict[str, List[float]] | None, frozenset[Tuple[str, str]], bool, Tuple[dict[tuple, tuple[str, str]], dict[str, str]], Tuple[bool, bool] | None]) → None[source]

Restore compiler state captured before a rolled-back write.

transaction(path: str) → Iterator[None][source]

Roll back writes below path while preserving the original error.

write_points(path: str, positions: NDArray[float32] | NDArray[float16], colors: NDArray[float32] | NDArray[uint8] | NDArray[uint16] | tuple | list | None = None, radii: NDArray[float32] | NDArray[float16] | NDArray[uint8] | float | None = None, sharpness: NDArray[float32] | NDArray[float16] | NDArray[uint8] | float | None = None, scalars: NDArray[float32] | NDArray[float16] | NDArray[uint8] | float | None = None, labels: Sequence[str] | None = None, image_labels: Any | None = None, keys: Sequence[str] | None = None, **attrs: Any) → Dict[str, Any][source]

Write points data progressively to Zarr.

Data is written immediately to disk without being kept in memory. If spatial ordering is enabled, points are reordered using Morton/Hilbert curves.

Scalar convenience: radii, sharpness, scalars, and colors accept scalars: - radii=0.5 → all points get radius 0.5 (broadcasted) - colors=[1.0, 0, 0] → all points red (broadcasted) - sharpness=0.5 → all points standard Gaussian (broadcasted; normalized [0, 1] knob) - scalars=0.5 → all points get scalar 0.5 (broadcasted)

Parameters:
  • path – Path for the points within the store

  • positions – Point positions of shape (N, D)

  • colors – Colors - array (N, 3), RGB tuple/list, or None

  • radii – Radii - array (N,), scalar float, or None

  • sharpness – Sharpness - array (N,), scalar float, or None

  • scalars – Scalars for colormap lookup - array (N,), scalar float, or None

  • labels – Optional list of strings, one per point. Stored as CSR-encoded label_offsets + label_bytes arrays for hover tooltips.

  • image_labels – Optional per-element images for hover thumbnails. Accepts List[bytes], List[PIL.Image], List[ndarray], List[Path], or Dict[int, Any] for sparse assignment.

  • keys – Optional list of machine-readable strings, one per point. Stored as CSR-encoded key_offsets + key_bytes arrays for link / copy templates to substitute as {hover_key}. Independent of labels.

  • **attrs – Additional attributes

Returns:

Metadata dictionary about the written data

write_lines(path: str, vertices: NDArray[float32] | NDArray[float16], widths: NDArray[float32] | NDArray[float16] | NDArray[uint8] | float, colors: NDArray[float32] | NDArray[uint8] | NDArray[uint16] | tuple | list | None = None, sharpness: NDArray[float32] | NDArray[float16] | NDArray[uint8] | float | None = None, scalars: NDArray[float32] | NDArray[float16] | NDArray[uint8] | float | None = None, indices: NDArray[uint32] | None = None, line_type: str = 'polyline', labels: Sequence[str] | None = None, image_labels: Any | None = None, keys: Sequence[str] | None = None, **attrs: Any) → dict[str, Any][source]

Write lines data to Zarr with dual spatial indexing.

Lines support dual spatial indexing for efficient lazy loading: - Vertices ordered in D-space (like Points) - Segments ordered in (2×D)-space (concatenating both endpoints)

All line types are internally converted to indexed representation.

Scalar convenience: widths, colors, and sharpness accept scalars: - widths=0.1 → all vertices get width 0.1 - colors=[1.0, 0, 0] → all vertices red - sharpness=0.5 → uniform sharpness (normalized [0, 1] knob; 0.5 = Gaussian)

Parameters:
  • path – Path for the lines within the store

  • vertices – Vertex positions of shape (N, D)

  • widths – Line widths - array (N,) or scalar float

  • colors – Colors - array (N, 3), RGB tuple/list, or None

  • sharpness – Sharpness - array (N,), scalar float, or None

  • scalars – Scalars for colormap lookup - array (N,), scalar float, or None

  • indices – Vertex-index pairs for indexed lines, as flat (2E,) elements or an (E, 2) pair array. Joint continuity requires connected edges to reference the same vertex row.

  • line_type – Type of line connectivity. Use polyline for one chain or indexed for multiple chains / graph topology.

  • labels – Optional list of strings, one per vertex. Stored as CSR-encoded label_offsets + label_bytes arrays for hover tooltips.

  • image_labels – Optional per-element images for hover thumbnails.

  • keys – Optional list of machine-readable strings, one per vertex. Stored as CSR-encoded key_offsets + key_bytes arrays for link / copy templates to substitute as {hover_key}. Independent of labels.

  • **attrs – Additional attributes

Returns:

Metadata dictionary about the written lines

write_mesh(path: str, vertices: NDArray[float32] | NDArray[float16], faces: NDArray[uint32], normals: NDArray[float32] | None = None, normal_dims: Sequence[int] | None = None, colors: NDArray[float32] | NDArray[uint8] | NDArray[uint16] | tuple | list | None = None, scalars: NDArray[float32] | NDArray[float16] | NDArray[uint8] | float | None = None, uvs: NDArray[float32] | None = None, texture: NDArray[Any] | None = None, texture_encoding: str = 'raw', texture_width: int | None = None, texture_height: int | None = None, texture_channels: int | None = None, texture_color_space: str = 'srgb', texture_ktx2_mode: str = 'uastc', texture_ktx2_quality: int | None = None, texture_ktx2_rdo_l: float | None = None, texture_ktx2_zcmp: int | None = None, shading: str | None = None, double_sided: bool = True, labels: Sequence[str] | None = None, image_labels: Any | None = None, keys: Sequence[str] | None = None, **attrs: Any) → dict[str, Any][source]

Write a triangle mesh to Zarr.

Mesh is the surface geometry type: vertices in nD plus a faces triangle-index array. Unlike Points / Lines / GSplats it carries no per-element size (a triangle’s extent comes from its own vertices), and it has no spatial index — the loader is whole-node. This call writes one leaf; LOD is layered on by the caller, not by this method — see write_mesh_multi_lod() (additive reveal levels) below.

Scalar convenience: colors accepts a broadcast RGB(A) tuple/list, and scalars a single value, exactly as the sibling writers do.

Parameters:
  • path – Path for the mesh within the store.

  • vertices – Vertex positions of shape (V, D).

  • faces – Triangle vertex indices, (F, 3) or flat (3F,). Wound counter-clockwise as seen with the mesh’s authored spatial triple in ascending index order.

  • normals – Optional per-vertex normals, shape (V, 3). Requires normal_dims.

  • normal_dims – The three dimension indices the normals describe — required with normals and rejected without them. Normals are a display-space quantity, so the store must say which three dimensions they belong to; an implicit “first three” is wrong for any mesh whose leading dimension is not spatial.

  • colors – Colors — array (V, 3|4), RGB(A) tuple/list, or None. A 4th component is per-vertex opacity.

  • scalars – Scalars for colormap lookup — array (V,), scalar, or None.

  • shading – "smooth", "flat", or unlit "none". Defaults to "smooth" when normals are supplied, else "flat". An explicit value is stored as given: "flat" renders a faceted surface even with normals present, "smooth" without normals falls back to derived flat normals at render time, and "none" computes no lighting normal.

  • double_sided – Whether back faces render. True by default.

  • labels – Optional per-vertex strings for hover tooltips (CSR-encoded).

  • image_labels – Optional per-vertex images for hover thumbnails.

  • keys – Optional list of machine-readable strings, one per vertex. Stored as CSR-encoded key_offsets + key_bytes arrays for link / copy templates to substitute as {hover_key}. Independent of labels.

  • uvs – Optional (V, 2) per-vertex texture coordinates. Required with texture and refused without it. Values outside [0, 1] are legal and tile under texture_wrap="repeat".

  • texture – Optional base-colour image. (H, W, C) array under texture_encoding="raw" or "ktx2", else a 1-D uint8 array of encoded bytes. Mutually exclusive with colors and colormap — a mesh has one base-colour source.

  • texture_encoding – raw | png | webp | jpeg | ktx2. KTX2 accepts uint8 RGB/RGBA input and requires the Khronos toktx executable, version 4.1.0 or newer. HDR requires raw.

  • texture_ktx2_mode – uastc (default) or etc1s.

  • texture_ktx2_quality – Codec quality; defaults to 2 for UASTC and 128 for ETC1S.

  • texture_ktx2_rdo_l – UASTC RDO lambda in [0.001, 10.0]; defaults to 0.25. Lower values preserve more quality and produce larger files; 0 disables RDO while retaining zstd compression.

  • texture_ktx2_zcmp – UASTC zstd level in [1, 22]; defaults to 9.

  • texture_width – Declared width. Required for encoded payloads, where it cannot be read without decoding; read off the array for raw or ktx2, and refused if it disagrees.

  • texture_height – Declared height. Same contract as texture_width.

  • texture_channels – Declared channels — 1, 3 or 4. Same contract.

  • texture_color_space – srgb (default) or linear. An ordinary PNG/JPEG is sRGB-encoded; declaring it wrong gives a subtly over-dark or washed-out surface rather than an obvious failure. HDR raw textures must be declared linear.

  • **attrs – Additional attributes.

Returns:

Metadata dictionary about the written mesh.

write_sound(path: str, payload: bytes, fmt: str, positions: NDArray[float32] | NDArray[float16] | None, *, sound_attrs: dict[str, Any], **attrs: Any) → dict[str, Any][source]

Write a sound node to Zarr.

The clip lands as a plain store key (audio.mp3 / audio.m4a) next to the node’s zarr documents and the optional positions array is written like a mesh’s vertices (whole-node, no spatial index). The bytes are folded into content_hash at finalize through PAYLOAD_FILE_ATTRS and never touch the scene bounds. See luxar.io._compiler.geometry_writers.sound.

write_points_multi_lod(path: str, levels: List[Dict[str, Any]], *, extend_to_all: List[str] | None = None, **attrs: Any) → Dict[str, Any][source]

Write multi-additive-LOD Points: parent node + additive_<i>/ subgroups.

Each additive_<i>/ subgroup is a fully-formed Points node (written via write_points()) carrying that level’s data arrays + its own spatial index. The parent group carries the global position_bounds, n_additive_sublods, the compositing attrs (opacity / gamma / colormap / blending_mode / transform / nd_transform / layer / visible) — these inherit down to subgroups via the viewer’s scene-graph composition at render time.

String channels: per-element labels and keys are each written as one CSR pair on the parent (stamping has_labels / has_keys); the additive_<i> subgroups carry neither. Each parent CSR’s index space is the committed union — the concatenation of the levels in additive_0 … additive_{n-1} order, each level in its own stored (spatially reordered) order — because the viewer’s progressive loader concatenates loaded levels into one buffer. Each channel is independently all-or-nothing across the ladder.

Parameters:
  • path – Path for the points node within the store.

  • levels – List of per-level dicts with keys positions / colors / radii / sharpness / scalars / labels / keys. positions is required; others may be None. Each string channel must be present on every level or on none.

  • extend_to_all – Forwarded to each per-level write. Already-resolved dimension NAMES — the caller expands the "all" sentinel, which must never reach disk (it is stamped verbatim here, onto the parent group and every sub-LOD).

  • **attrs – Additional parent-node attrs (compositing, colormap, etc.).

Returns:

Aggregate metadata dict with type / n_points / n_additive_sublods / position_bounds / levels.

write_lines_multi_lod(path: str, levels: List[Dict[str, Any]], *, extend_to_all: List[str] | None = None, **attrs: Any) → Dict[str, Any][source]

Write multi-additive-LOD Lines.

Mirrors write_points_multi_lod(). Each level dict carries vertices + widths + colors / sharpness / scalars / labels / keys + segments (local index pairs into that level’s vertices) + n_polylines. Each subgroup is written via write_lines() with line_type='indexed' and the local segment indices.

String channels (per-VERTEX for Lines) are written as one CSR pair per present labels / keys channel on the parent, which stamps has_labels / has_keys. Each pair describes the committed union: the concatenation of the levels in additive_0 … additive_{n-1} order, each level in its own stored (spatially reordered) order. The additive_<i> subgroups carry neither. Each channel is independently all-or-nothing across the ladder.

write_mesh_multi_lod(path: str, levels: List[Dict[str, Any]], *, extend_to_all: List[str] | None = None, **attrs: Any) → Dict[str, Any][source]

Write multi-additive-LOD Mesh: parent node + additive_<i>/ subgroups.

Mirrors write_points_multi_lod() in shape — a parent carrying the totals, the global position_bounds, n_additive_sublods and the compositing attrs, with one fully-formed leaf per level underneath. Each level dict carries what write_mesh() needs: vertices + faces (already re-indexed into that level’s OWN vertex table by luxar.mesh.split.split_mesh_by_faces()) plus optional normals / normal_dims / colors / scalars / shading / double_sided / _scalar_data_range, and lod_stats.

A mesh ladder is a REVEAL: the levels are a partition of the faces, coarsest (innermost shell) first, and the viewer forms level i by concatenating levels 0..i. Level counts are therefore FACE counts, and the parent’s n_vertices / n_faces are the sums over levels — the vertex sum exceeds the source vertex count by the shell-boundary duplication the re-indexing costs (see luxar.mesh.split.duplication_factor()).

No labels, and deliberately no union label CSR. The three sibling ladders write ONE CSR pair on the parent spanning the levels, because their elements are independent rows and the concatenation of the levels is a well-defined index space. A mesh level RE-INDEXES its own vertices, so a single source vertex on a shell boundary maps to a slot in several levels (and a vertex no level’s faces reference maps to none): the union index space is ill-defined, and any CSR written over it would pair labels with the wrong vertices. So the parent carries no has_labels / has_keys and a level carrying either is refused outright rather than silently dropped.

Parameters:
  • path – Path for the mesh node within the store.

  • levels – One dict per additive level, coarsest first. vertices and faces are required; every other key is optional.

  • extend_to_all – Forwarded to each per-level write. Already-resolved dimension NAMES — the caller expands the "all" sentinel, which must never reach disk (it is stamped verbatim here, onto the parent group and every sub-LOD).

  • **attrs – Additional parent-node attrs (compositing, colormap, …).

Returns:

Aggregate metadata dict with type / n_vertices / n_faces / n_additive_sublods / position_bounds / levels.

Raises:

ValueError – If levels is empty, an attr is invalid, or any level carries labels or keys.

write_gsplats(path: str, centers: NDArray[float32] | NDArray[float16], amplitudes: NDArray[float32] | NDArray[float16] | NDArray[uint8] | float, cholesky_factors: NDArray[float32], colors: NDArray[float32] | NDArray[uint8] | NDArray[uint16] | tuple | list | None = None, label_ids: ndarray[Any, Any] | Sequence[int] | None = None, label_vocabulary: dict[int, str] | None = None, labels: Sequence[str] | None = None, image_labels: Any | None = None, keys: Sequence[str] | None = None, _source_dtype: str | None = None, **attrs: Any) → dict[str, Any][source]

Write Gaussian splats data to Zarr (single-LOD, flat layout).

Scalar convenience: amplitudes and colors accept scalars: - amplitudes=1.0 → all splats get amplitude 1.0 - colors=[1.0, 0, 0] → all splats red

Parameters:
  • path – Path for the gsplats within the store

  • centers – Splat centers of shape (N, D)

  • amplitudes – Amplitudes - array (N,) or scalar float

  • cholesky_factors – Packed Cholesky factors of shape (N, k)

  • colors – Colors - array (N, 3), RGB tuple/list, or None

  • label_ids – Optional non-negative integer class id per splat.

  • label_vocabulary – Explicit mapping from stored class ids to names.

  • labels – Optional list of strings, one per splat. Stored as CSR-encoded label_offsets + label_bytes arrays for hover tooltips.

  • image_labels – Optional per-element images for hover thumbnails.

  • keys – Optional list of machine-readable strings, one per splat. Stored as CSR-encoded key_offsets + key_bytes arrays for link / copy templates to substitute as {hover_key}. Independent of labels.

  • **attrs – Additional attributes

Returns:

Metadata dictionary about the written gsplats

write_gsplat_leaf_subtree(path: str, leaf: Any, _source_dtype: str | None = None, **attrs: Any) → dict[str, Any][source]

Write a GSplatLeaf (single set or additive ladder) into the scene.

This is the scene-side seam onto         write_gsplat_leaf() — the single authoring path also used by the standalone .gsplats.zarr writer. There is no parallel additive-ladder writer: the additive additive_<i>/ subgroups, their attrs, spatial ordering, chunking and position_bounds come from the exact same code a standalone file uses, so a scene additive ladder is byte-identical to a standalone one by construction.

Used for GSplatData embeds, whose arrays are always full per-splat (no uniform-Cholesky / scalar-amplitude / labels / keys — those leaf-only scene features stay on write_gsplats()).

Parameters:
  • path – Path for the gsplats node within the store.

  • leaf – A GSplatLeaf (1 sub-LOD → flat leaf; >1 → additive ladder).

  • **attrs – Node attributes (opacity, blending_mode, colormap, extend_to_all, truncation_radius, transform, …).

Returns:

Aggregate metadata dict (incl. position_bounds).

create_resizable_dataset(path: str, dtype: dtype, shape: Tuple[int, ...], maxshape: Tuple[int | None, ...] | None = None, chunks: bool | int | Tuple[int, ...] | None = True) → Array[source]

Create a resizable dataset for streaming writes.

Parameters:
  • path – Path for the dataset

  • dtype – Data type

  • maxshape – ACCEPTED AND IGNORED. An h5py-compatibility argument that zarr has never enforced — a zarr array has no pre-declared ceiling, .resize() simply works — so this has always been inert. zarr 2 accepted and dropped it; zarr 3 removed the kwarg outright, so it is no longer forwarded at all. Kept on the signature because callers pass it and removing it would be a gratuitous break, but it constrains nothing.

  • shape – Initial shape

  • chunks – Chunking configuration

Returns:

Zarr dataset handle

finalize() → None[source]

Finalize the Zarr store with metadata consolidation.

Runs under suppress_payload_member_warning(): the finalize walkers (bounds, hashing, LOD backfill, amplitude windows) enumerate every group through zarr, which warns once per plain payload key it meets (an overlay’s image.png, a sound node’s audio.mp3). Those keys are part of the format, so the advice is noise here — and under -W error it used to turn every such compile into “Could not finalize Zarr store”.

property store_path: str

Get the live directory store, or the finalized archive path.

property final_store_path: str

Get the directory store or archive path produced by finalization.

Key Features

  • Progressive Writing: No memory caching - stream directly to disk

  • Spatial Indexing: Automatic Morton/Hilbert ordering for fast queries

  • Chunk Bounds: Precomputed bounding boxes for O(chunks) spatial queries

  • Context Manager: Automatic finalization with with statement

  • Hierarchy Support: Create nested groups with full path notation

Example:

from luxar.io import LuxarZarrCompiler
from luxar.core.dimensions import Dimensions
import numpy as np

dims = Dimensions.default_3d()
with LuxarZarrCompiler('scene.luxar.zarr') as compiler:
    scene = compiler.create_scene(dimensions=dims)

    # Add root-level points
    compiler.write_points('cloud', positions, colors)

    # Create hierarchy with full paths
    compiler.write_points('GroupA/SubPoints', positions2, colors2)

LuxarScene

The reader provides efficient access to Zarr datasets with lazy loading and spatial queries.

class luxar.io.LuxarScene(root: Group, path: Path)[source]

Bases: object

Read-only access to a Luxar zarr scene.

Provides Python API to read and validate Luxar scene files. Enables: - Round-trip testing (write → read → verify) - Python-based scene analysis and inspection - Format validation and debugging - Data extraction for processing

Example

>>> scene = LuxarScene.load('scene.zarr')
>>> aprint(f"Version: {scene.version}")
>>> aprint(f"Nodes: {[n['name'] for n in scene.nodes]}")
>>> points = scene.get_points('my_cloud')
>>> positions = points['positions']  # Decoded array
__init__(root: Group, path: Path)[source]

Initialize from an open zarr group.

Parameters:
  • root – Open zarr group (root of scene)

  • path – Path to the zarr store

classmethod load(path: str | Path) → LuxarScene[source]

Load a Luxar scene from a zarr store.

Parameters:

path – Path to the zarr store

Returns:

LuxarScene instance

Raises:
  • FileNotFoundError – If path doesn’t exist

  • ValueError – If not a valid Luxar scene, or if the store is marked incomplete (the writer exited with an error before finalizing, so nodes/metadata may be missing).

property path: Path

Path to the zarr store.

property version: str

Scene format version (format_version, else the 0.1 legacy key).

property root_attrs: Dict[str, Any]

All root attributes.

property dimensions: Dimensions | None

Scene dimensions if defined.

property viewer_config: ViewerConfig | None

Scene viewer configuration if defined.

The read-only mirror of luxar.core.scene.Scene.viewer_config, and parsed the same way — the attr is a plain dict on disk, so a consumer that means to CARRY the config into a new scene (luxar mesh lod rewriting a ladder, say) needs the object, not the dict. Losing it is not cosmetic: a scene that dropped tone_mapping silently falls back to the viewer’s ACES default, which shifts the hues of a custom colormap LUT.

Returns:

The parsed ViewerConfig, or None when the scene declares none.

property nodes: List[Dict[str, Any]]

List all nodes with metadata.

Returns list of dicts with node info: - name: Node name (path from root) - type: ‘points’, ‘gsplats’, ‘lines’, or ‘group’ - Additional metadata depending on type

list_points() → List[str][source]

Names of all points nodes.

list_gsplats() → List[str][source]

Names of all gsplats nodes.

list_lines() → List[str][source]

Names of all lines nodes.

list_meshes() → List[str][source]

Names of all mesh nodes.

list_sounds() → List[str][source]

Names of all sound nodes.

list_groups() → List[str][source]

Names of all group nodes.

has_node(name: str) → bool[source]

Check if a node exists.

get_node_type(name: str) → str[source]

Get the type of a node.

get_node_metadata(name: str) → Dict[str, Any][source]

Get metadata for a node (no data arrays).

Note: transforms are returned as raw column-major lists. Use get_group() for automatic transform conversion.

get_colormap_lut(name: str) → ndarray | None[source]

Get a node’s custom colormap LUT, or None if it has none.

A node whose colormap attr is the sentinel "custom" carries the actual colors in a colormap_lut dataset — the writer resolves any colormap that is not one of the builtin names (an ndarray LUT, but also a plain matplotlib/colorcet name) to that pair, so the viewer needs neither library at display time.

Returned raw, not through _decode_array: the LUT is written as a plain unencoded (256, 3) uint8 dataset, so there is no encoding header for the decoder to read.

Parameters:

name – Name of the node

Returns:

The (256, 3) uint8 LUT, or None when the node has no colormap_lut dataset (a builtin or absent colormap).

Raises:

KeyError – If node doesn’t exist

get_group(name: str) → Dict[str, Any][source]

Get group node metadata with parsed transform.

Parameters:

name – Name of the group node

Returns:

Dictionary with group metadata. Transform (if present) is converted from column-major list to a 4x4 NumPy matrix.

Raises:
get_points(name: str, *, flatten: bool = False) → PointsData[source]

Load points node data with automatic decoding.

Parameters:
  • name – Name of the points node

  • flatten – When true, read the finest substitutive level, concatenate every partition part, and concatenate every disjoint additive increment. This reconstructs the complete finest-level point cloud from an authored structure. Substitutive children are ordered coarsest to finest, and additive children are disjoint increments. chunk_bounds is not returned because leaf chunk metadata does not align with concatenated rows. This is a whole-node, non-streaming read; use get_point_array() when only one field is needed. The default keeps the legacy flat-leaf behavior.

Returns:

positions, colors, radii, sharpness, chunk_bounds, metadata. Supports dict-style access for backward compatibility (e.g. data["positions"]).

Return type:

PointsData with fields

Raises:
  • KeyError – If node doesn’t exist

  • ValueError – If node is not a points node, or its structured leaves are empty or carry inconsistent optional arrays.

get_point_array(name: str, array_name: str, *, flatten: bool = False) → ndarray | None[source]

Decode one points array, optionally across the finest structure.

array_name uses the PointsData field spelling: positions, colors, radii, sharpness, or chunk_bounds. Flattened reads reject chunk_bounds because leaf chunk metadata does not align with concatenated rows. Reading one field is useful when another layer shares a large array_ref and decoding the complete node would materialize an unnecessary second copy.

get_gsplats(name: str) → GSplatsData[source]

Load gsplats node data with automatic decoding.

Parameters:

name – Name of the gsplats node

Returns:

centers, amplitudes, cholesky_factors, colors, chunk_bounds, metadata. Supports dict-style access for backward compatibility.

Return type:

GSplatsData with fields

Raises:
get_lines(name: str) → LinesData[source]

Load lines node data with automatic decoding.

Parameters:

name – Name of the lines node

Returns:

vertices, widths, colors, sharpness, segments, metadata. Supports dict-style access for backward compatibility.

Return type:

LinesData with fields

Raises:
get_mesh(name: str) → MeshData[source]

Load mesh node data with automatic decoding.

Parameters:

name – Name of the mesh node

Returns:

vertices, faces, normals, normal_dims, colors, scalars, metadata. Supports dict-style access for backward compatibility.

Return type:

MeshData with fields

Raises:
  • KeyError – If node doesn’t exist

  • ValueError – If node is not a mesh node, or is missing a required array

Key Features

  • Lazy Loading: Only load metadata initially, fetch data on demand

  • Spatial Queries: Efficient chunk-based queries using precomputed bounds

  • Node Discovery: List and inspect all nodes in the scene graph

  • Type-Safe Access: Type-specific getters (get_points, get_lines, get_gsplats)

Example:

from luxar.io import LuxarScene

# Load scene (metadata only)
scene = LuxarScene.load('scene.luxar.zarr')

# List available nodes
print(scene.list_points())
print(scene.list_lines())

# Get node data (lazy loading)
points_node = scene.get_points('cloud')
positions = points_node.positions  # NumPy array attribute

Spatial Ordering

Spatial ordering functions for Morton/Hilbert curves and chunk bounds calculation.

Spatial ordering algorithms for Points, Lines, and GSplats.

Public facade over the private io/_ordering/ subpackage:

  • curves/ — Morton & Hilbert space-filling curve encoders

  • grid — float → integer grid quantization

  • compound — the shared lexsort-barrier + curve primitive

  • points / gsplats / lines — per-geometry sorts and chunk-bounds builders (parallel structure per the three-geometry symmetry)

Everything is re-exported here so the historical public paths (luxar.io.ordering.<name>) keep resolving unchanged.

luxar.io.ordering.morton_encode_nd(coords: ndarray, bits_per_dim: int = 16) → ndarray[source]

Encode nD integer coordinates to Morton codes via bit interleaving.

Uses a Numba JIT-compiled kernel when available, falling back to vectorized NumPy.

Parameters:
  • coords – Integer coordinates, shape (N, d)

  • bits_per_dim – Bits to use per dimension (default 16)

Returns:

Morton codes, shape (N,), dtype uint64

luxar.io.ordering.morton_encode_128bit(coords: ndarray, bits_per_dim: int) → tuple[ndarray, ndarray][source]

Encode nD integer coordinates to 128-bit Morton codes as (high, low) pairs.

For high-dimensional data (> 6 dims), 64-bit Morton codes have insufficient precision. This function produces 128-bit codes as paired uint64 values.

Parameters:
  • coords – Integer coordinates, shape (N, d)

  • bits_per_dim – Bits to use per dimension

Returns:

Upper 64 bits of Morton codes, shape (N,), dtype uint64 low: Lower 64 bits of Morton codes, shape (N,), dtype uint64

Return type:

high

luxar.io.ordering.hilbert_encode_nd(coords: ndarray, bits_per_dim: int = 16) → ndarray[source]

Encode nD integer coordinates to Hilbert curve indices.

Uses a Numba JIT-compiled kernel to avoid Python-loop overhead. Falls back to the hilbertcurve library if Numba is unavailable.

Parameters:
  • coords – Integer coordinates, shape (N, d)

  • bits_per_dim – Bits to use per dimension (default 16)

Returns:

Hilbert indices, shape (N,), dtype uint64

luxar.io.ordering.normalize_coords_to_grid(coords: ndarray, min_coords: ndarray, max_coords: ndarray, resolution: int) → ndarray[source]

Normalize floating coordinates to an integer grid.

luxar.io.ordering.compute_auto_resolution(coords: ndarray, max_resolution: int = 65536) → int[source]

Compute appropriate resolution based on data spread.

Parameters:
  • coords – Float coordinates, shape (N, d)

  • max_resolution – Maximum resolution (default 65536)

Returns:

Resolution as power of 2, capped at max_resolution

luxar.io.ordering.detect_barrier_dims(centers: ndarray, max_cardinality: int = 1024) → list[int][source]

Heuristically identify categorical/barrier axes in a GSplat center array.

A standalone .gsplats.zarr carries no per-dimension descriptors, so when neither an explicit barrier_dims nor persisted coarsen_dims provenance is available this conservatively infers which axes behave like a categorical stack (time, channel): an axis qualifies iff its values are integers AND take few distinct values (<= max_cardinality and << N). Continuous spatial float coordinates never qualify.

Conservative in the safe direction. A false NEGATIVE (missing a barrier) only causes over-fetch — no worse than pure spatial ordering. A false POSITIVE (flagging a spatial axis) gives it tight (epsilon-padded) chunk bounds with no σ expansion, so a spatially-extended splat can fall outside its chunk bounds and be dropped from a query — a correctness bug. So both guards err toward NOT flagging: the integer test is strict (rtol=0, absolute tolerance only — a large-magnitude continuous coordinate is never “close enough” to an integer), and the n_unique * 4 <= n guard rejects a fine integer spatial grid (many distinct values relative to N) that is not a true category.

Subordinate by design: callers apply explicit barrier_dims and coarsen_dims complements first (the scene compiler passes scene Dimension.discrete dims; the batch merge passes the stacked-time axis), using this only as the last resort for provenance-less standalone files.

Parameters:
  • centers – Splat centers, shape (N, d).

  • max_cardinality – Max distinct values for an axis to count as categorical.

Returns:

Sorted list of barrier column indices (possibly empty).

luxar.io.ordering.sort_points_compound(positions: ndarray, dimensions: list[Dimension], method: Literal['morton', 'hilbert'] = 'hilbert') → tuple[ndarray, dict][source]

Sort Points using compound ordering (discrete dims → spatial curve).

This implements the compound ordering strategy from luxar.io spec: - Primary sort: Discrete dimensions (lexicographic) - Secondary sort: Morton/Hilbert code of spatial dimensions

Parameters:
  • positions – Point positions, shape (N, d), all dimensions

  • dimensions – Dimension objects defining discrete/spatial properties

  • method – Spatial curve method (“morton” or “hilbert”), default “hilbert”

Returns:

Indices to reorder points metadata: Dict with ordering metadata

Return type:

sort_indices

luxar.io.ordering.sort_splats_spatial(centers: ndarray, method: Literal['morton', 'hilbert'] = 'hilbert', resolution: int | None = None, slice_dims: Sequence[int] | None = None) → tuple[ndarray, dict][source]

Sort GSplats using barrier-aware compound spatial ordering.

When slice_dims names one or more categorical/barrier axes (e.g. time or channel), splats are grouped by those axes first (lexicographic) and a Morton/Hilbert curve orders spatially within each barrier value — so a chunk never straddles two timepoints. This mirrors sort_points_compound() and is what keeps per-timepoint reads local (see the io ordering README).

With slice_dims empty/None this is pure spatial ordering over all center columns — the historical behavior, unchanged for 3D data.

Parameters:
  • centers – Splat centers, shape (N, d), float32

  • method – Spatial curve method (“morton” or “hilbert”)

  • resolution – Ignored (kept for signature stability; the grid resolution is derived per-axis from the bit budget, as it always has been).

  • slice_dims – Barrier/categorical column indices to order by first. An index outside [0, d) raises ValueError (see _normalise_slice_dims()) — the SAME check compute_chunk_bounds_gsplats() applies, so the failure lands at the first door. apply_gsplat_spatial_ordering hands the same list to both, and a negative index is a genuine sort-side bug of its own (it lands in the barrier set AND in ordering_dims) — see _normalise_slice_dims() for the full argument.

Returns:

Indices to reorder splats metadata: Dict with ordering metadata (incl. slice_dims / ordering_dims)

Return type:

sort_indices

Raises:

ValueError – If a slice_dims entry is outside [0, d).

luxar.io.ordering.compute_chunk_bounds_points(positions: ndarray, radii: ndarray | float | None, chunk_size: int, slice_dims: list[int] | None = None, *, coord_slack: NDArray[float64] | None = None, scalar_slack: float | None = None) → ndarray[source]

Compute chunk bounding boxes for Points (includes radius extent).

The pad IS the footprint: [min - r, max + r] is exactly the set of query positions for which some point in the chunk can still be visible, so the bound is right in both directions rather than merely wide enough. The interval is computed in float64 and narrowed to the float32 store with OUTWARD rounding (see _store_outward_f32()), so the “never tighter” half holds against the padded radii at every coordinate magnitude — not only where a 0.5 pad happens to survive a round-to-nearest store.

SCALAR QUANTISATION SLACK (scalar_slack): radii is itself a POSITIVE_SCALAR and may decode larger than the authored value. This single per-array pad is added to the radius on SPATIAL dimensions only; barrier dimensions still receive no footprint expansion. The compiler supplies it from positive_scalar_round_trip_slack(); a direct caller that omits it gets authored-radius bounds.

QUANTISATION SLACK (coord_slack): these bounds are computed from the positions AS HANDED IN, but under the default AUTO encoding the positions themselves are stored as per-axis uint16 fixed point (the COORDINATE path in luxar.encoding._encoders.perchannel), so a DECODED position can land up to half a quantum (extent/131070) outside a bound derived from the authored one — 7.6e-3 at an extent of 1000, well above the float32 ULP the outward store closes, and enough for the reader to skip the chunk entirely. coord_slack is that displacement, per axis, and is added outward on EVERY dimension: on top of the radius on a spatial axis and on top of _BARRIER_BOUND_EPS on a barrier axis (the epsilon is float-boundary safety, the slack is a displacement — they add). The compiler supplies it from the encoder’s own predicate (coordinate_round_trip_slack()), which returns zero for an axis the encoder stores exactly (a gridded time/channel axis, a constant axis, a float32 fallback — and the whole array when the write will really store a LUT, which the caller declares with that predicate’s allow_lut). A direct caller that omits it gets bounds for the authored coordinates only (issue #1655).

radii=None does not mean “no extent” — a points node that stores no radii array is drawn with the renderer’s default radius, so the bounds are expanded by DEFAULT_POINT_RADIUS (luxar.typing_utils.constants), exactly as if a scalar radius of that value had been passed.

Reader-side note: for a radii-less node the viewer runs NO per-point effective-radius cull (data/points/effective-radius-calculator.ts and data/points/projection.ts both gate on radii), and its hidden-dim query tolerance falls back to defaultMaxRadius = 0.1 because no max_radius attr is stamped without radii. An honest (wider) bound is therefore paid in over-DRAW, not merely over-fetch: nothing culls the extra points, so every point of every chunk the window touches is rendered. And the cost is relative to the AXIS, not to the pad — the selection window on a hidden continuous axis grows from chunk_extent + 2*(0.01 + 0.1) to chunk_extent + 2*(0.5 + 0.1), a few percent on an axis spanning ~100 units but EVERY chunk on an axis spanning ~1 unit. It is still the right trade against the old under-fetch, which dropped points with nothing to notice. Shrinking that tolerance to a float-safety epsilon now that the bounds are honest is issue #1655 items 2 and 3; nothing on the reader side changes here.

CRITICAL: Radius expansion is only applied to SPATIAL dimensions, not discrete dimensions. Discrete dimensions (slice_dims) represent categorical values like time steps, channels, or orbital indices. A point at orbital=0 should NOT extend into orbital=3 space - they are separate categories. Those axes get only the tiny float-boundary epsilon _BARRIER_BOUND_EPS, on every path.

Parameters:
  • positions – Point positions (already sorted), shape (N, d)

  • radii – Point radii (already sorted), shape (N,), a broadcast scalar, or None (⇒ the renderer’s DEFAULT_POINT_RADIUS)

  • chunk_size – Number of points per chunk

  • slice_dims – Indices of discrete (non-spatial) dimensions where radius expansion should NOT be applied. Default: None (apply to all dims). An index outside [0, d) raises ValueError (see _normalise_slice_dims()).

  • coord_slack – Per-axis outward pad, shape (d,), covering how far the STORE can move a coordinate from the value passed here (see the QUANTISATION SLACK note above). Default None ⇒ zero on every axis. Validated by _normalise_coord_slack().

  • scalar_slack – Outward pad covering how far the stored radius can exceed the authored radius. Applied on spatial dimensions only; overflow produces conservative infinite spatial bounds.

Returns:

Bounding boxes, shape (num_chunks, d, 2)

Return type:

chunk_bounds

luxar.io.ordering.compute_chunk_bounds_gsplats(centers: ndarray, cholesky_factors: ndarray, chunk_size: int, coverage_sigma: float = 2.75, slice_dims: Sequence[int] | None = None, *, coord_slack: ndarray | None = None) → ndarray[source]

Compute chunk bounding boxes for GSplats (includes ellipsoidal extent).

CRITICAL: the ellipsoidal (coverage_sigma·σ) extent is applied only to SPATIAL axes. slice_dims name categorical/barrier axes (time, channel): a splat at time=0 must not extend into time=1’s bounds, so those axes get only a tight float-boundary epsilon (_BARRIER_BOUND_EPS). This mirrors compute_chunk_bounds_points() and keeps a chunk’s barrier-axis footprint from straddling categories, which is what makes single-timepoint queries fetch only their own chunks.

The extent is accumulated in float64 and narrowed to the float32 store with OUTWARD rounding (see _store_outward_f32_array()), so a stored bound is never tighter than the footprint at any coordinate magnitude — not only where a small σ happens to survive float32 arithmetic and a round-to-nearest store.

coord_slack extends that guarantee to the DECODED centers. Under AUTO, a non-gridded axis is stored as per-axis uint16 fixed point and can move by half a quantum (extent/131070). The compiler asks the encoder for that displacement using the centers’ resolved mode and adds it on every axis: on top of coverage_sigma·σ spatially, and on top of _BARRIER_BOUND_EPS on a barrier axis. Exact paths answer zero/None (a gridded time/channel axis, LUT encoding, or float32 escalation), so their bounds remain unchanged. A direct caller that omits coord_slack gets bounds for the authored centers only.

KNOWN GAP — THE DECODED CHOLESKY EXTENT IS NOT COVERED. cholesky_factors is quantised in its own right (the diagonal uses log_perchannel_u8 under AUTO), and these bounds are built from the values as handed in. A decoded diagonal can therefore produce a LARGER σ than the one the coverage_sigma·σ pad was sized for, letting the rendered footprint escape the stored bound. This is the extent half of the footprint: coord_slack closes the center-coordinate half and does nothing for this one, which needs the same encoder-reported round-trip treatment as decoded point radii and line widths.

Parameters:
  • centers – Splat centers (already sorted), shape (N, d)

  • cholesky_factors – Packed Cholesky factors (already sorted), shape (N, k), or a single shared row (1, k) reused for every chunk when all splats have a uniform (identical) Cholesky factorization

  • chunk_size – Number of splats per chunk

  • coverage_sigma – Coverage radius in standard deviations. This is the gsplat truncation_radius under its spatial-ordering name — the same quantity the LOD path spells truncation_sigmas. Two compiler sites bind it to the dataset’s own radius: io/_compiler/gsplat_tree.py::_write_single_splat_set (by keyword) and the scene path geometry_writers/gsplats.py → gsplat_assembly.py::apply_gsplat_spatial_ordering, which passes truncation_radius positionally. The default here only applies to a direct call.

  • slice_dims – Barrier/categorical dimension indices (no σ expansion). Default None → expand all axes (historical behavior). An index outside [0, d) raises ValueError (see _normalise_slice_dims()).

  • coord_slack – Per-axis outward pad covering the center encoder’s round-trip displacement. Default None → zero on every axis.

Returns:

Bounding boxes, shape (num_chunks, d, 2)

Return type:

chunk_bounds

luxar.io.ordering.convert_to_indexed(n_vertices: int, line_type: str, indices: ndarray | None) → ndarray[source]

Convert any line type to indexed segment pairs.

All line types are internally converted to the unified indexed representation for efficient spatial ordering and storage.

Parameters:
  • n_vertices – Number of vertices in the lines

  • line_type – One of “segments”, “polyline”, “loop”, “indexed”

  • indices – For “indexed” type, the user-provided index array

Returns:

(S, 2) uint32 array of vertex index pairs

Return type:

segments

Raises:

ValueError – If line_type is invalid or indices missing for “indexed” type

luxar.io.ordering.sort_segments_compound(segment_coords_2d: ndarray, dimensions: list[Dimension], method: Literal['morton', 'hilbert'] = 'hilbert') → tuple[ndarray, dict][source]

Sort segments using compound ordering in (2×D)-space.

Segments are represented as (2×D)-dimensional points by concatenating both endpoint coordinates. This captures the full geometric nature of segments (position, orientation, length) for spatial coherence.

The compound ordering strategy is: - Primary sort: Discrete dimensions from both endpoints (lexicographic) - Secondary sort: Morton/Hilbert code of spatial dimensions from both endpoints

Parameters:
  • segment_coords_2d – Segment coordinates in (2×D)-space, shape (S, 2*d) First d columns are start point, second d columns are end point.

  • dimensions – Original Dimension objects (may have more dims than data)

  • method – Spatial curve method (“morton” or “hilbert”), default “hilbert”

Returns:

Indices to reorder segments metadata: Dict with ordering metadata in (2×D)-space

Return type:

sort_indices

luxar.io.ordering.order_lines_spatial(vertices: ndarray, segments: ndarray, dimensions: list[Dimension], method: Literal['morton', 'hilbert'] = 'hilbert') → tuple[ndarray, ndarray, ndarray, ndarray, dict][source]

Apply dual spatial ordering to lines using Morton or Hilbert curves.

Lines have dual ordering: 1. Vertices are ordered in D-space (like Points) 2. Segments are ordered in (2×D)-space (concatenating both endpoints)

This is the main entry point for Lines spatial ordering.

Parameters:
  • vertices – (V, D) float32 vertex positions

  • segments – (S, 2) uint32 index pairs (from convert_to_indexed)

  • dimensions – List of Dimension objects (may have more dims than data)

  • method – Ordering method (“morton” or “hilbert”)

Returns:

(V, D) float32 reordered vertex positions sorted_segments: (S, 2) uint32 reordered and remapped segment indices vertex_sort_indices: Indices to recover original vertex order segment_sort_indices: Indices to recover original segment order metadata: Dict with “vertex_ordering” and “segment_ordering” sub-dicts

Return type:

sorted_vertices

luxar.io.ordering.compute_vertex_chunk_bounds(vertices: ndarray, chunk_size: int, slice_dims: list[int] | None = None, *, coord_slack: NDArray[float64] | None = None) → ndarray[source]

Compute chunk bounding boxes for vertices (no radius/width expansion).

Spatial axes carry no pad (a vertex is a point), so the interval is exact; the barrier epsilon, however, is a small ABSOLUTE quantity. Both are computed in float64 and narrowed to the float32 store with OUTWARD rounding (see _store_outward_f32_array()), so a stored bound is never tighter than the footprint at any coordinate magnitude.

QUANTISATION SLACK (coord_slack): that holds for the vertices AS HANDED IN. Under the default AUTO encoding the vertices are stored as per-axis uint16 fixed point, so a DECODED vertex can sit up to half a quantum (extent/131070) outside a bound derived from the authored one on a non-gridded axis. coord_slack is that displacement, per axis, added outward on EVERY dimension — including on top of _BARRIER_BOUND_EPS on a barrier axis. The compiler supplies it from the encoder’s own predicate; a direct caller that omits it gets bounds for the authored vertices only. See compute_chunk_bounds_points().

Parameters:
  • vertices – Vertex positions (already sorted), shape (V, D)

  • chunk_size – Number of vertices per chunk

  • slice_dims – Indices of discrete (non-spatial) dimensions (padded by a float-boundary epsilon only — the reader’s per-dimension tolerance owns the query reach). An index outside [0, D) raises ValueError (see _normalise_slice_dims()).

  • coord_slack – Per-axis outward pad, shape (D,), covering how far the store can move a vertex from the value passed here. Default None ⇒ zero on every axis. Validated by _normalise_coord_slack().

Returns:

Bounding boxes, shape (num_chunks, D, 2)

Return type:

chunk_bounds

luxar.io.ordering.compute_segment_chunk_bounds(vertices: ndarray, segments: ndarray, widths: ndarray, chunk_size: int, slice_dims: list[int] | None = None, *, coord_slack: NDArray[float64] | None = None, scalar_slack: float | None = None) → ndarray[source]

Compute chunk bounding boxes for segments (includes line width).

Segment bounds are in D-dimensional space (not 2×D) for view frustum intersection tests. Each segment’s bounds include the line width extent.

The width interval is accumulated in float64 and narrowed to the float32 store with OUTWARD rounding (see _store_outward_f32_array()), so a stored bound is never tighter than the footprint of the AUTHORED widths at any coordinate magnitude — not only where a small width happens to survive float32 arithmetic and a round-to-nearest store.

SCALAR QUANTISATION SLACK (scalar_slack): widths is itself a POSITIVE_SCALAR and may decode larger than the authored value. This single per-array pad is added to the width on SPATIAL dimensions only; barrier dimensions still receive no footprint expansion. The compiler supplies it from positive_scalar_round_trip_slack(); a direct caller that omits it gets authored-width bounds.

QUANTISATION SLACK (coord_slack): that holds for the vertices AS HANDED IN. Under the default AUTO encoding the vertices are stored as per-axis uint16 fixed point, so a DECODED endpoint can sit up to half a quantum (extent/131070) outside a bound derived from the authored one on a non-gridded axis. coord_slack is that displacement, per axis, added outward on EVERY dimension — on top of the width on a spatial axis and on top of _BARRIER_BOUND_EPS on a barrier axis. The compiler supplies it from the encoder’s own predicate (the SAME vector it gives compute_vertex_chunk_bounds(): both builders are in D-space over the same vertices array); a direct caller that omits it gets bounds for the authored vertices only. See compute_chunk_bounds_points().

IMPORTANT: widths must be a full (V,) array. Broadcast widths should be expanded with np.full(V, width_value) before calling this function.

Parameters:
  • vertices – Vertex positions (already sorted), shape (V, D)

  • segments – Segment index pairs (already sorted), shape (S, 2)

  • widths – Vertex widths (already sorted), shape (V,)

  • chunk_size – Number of segments per chunk

  • slice_dims – Indices of discrete (non-spatial) dimensions (padded by a float-boundary epsilon only — the reader’s per-dimension tolerance owns the query reach). An index outside [0, D) raises ValueError (see _normalise_slice_dims()).

  • coord_slack – Per-axis outward pad, shape (D,), covering how far the store can move a vertex from the value passed here. Default None ⇒ zero on every axis. Validated by _normalise_coord_slack().

  • scalar_slack – Outward pad covering how far the stored width can exceed the authored width. Applied on spatial dimensions only; overflow produces conservative infinite spatial bounds.

Returns:

Bounding boxes, shape (num_chunks, D, 2)

Return type:

chunk_bounds

Morton and Hilbert Curves

Space-filling curves map n-dimensional coordinates to 1-dimensional ordering while preserving spatial locality. This enables:

  • Better Compression: Similar values group together (2-10x improvement)

  • Faster Queries: Scan fewer chunks for spatial queries (10-100x speedup)

  • Cache Efficiency: Adjacent points in space are adjacent in memory

See Architecture and Core Concepts for detailed explanation of spatial indexing strategy.

Writer Protocol

class luxar.io.writer.ZarrWriterProtocol(*args, **kwargs)[source]

Protocol for progressive Zarr writing.

This interface enables different implementations of Zarr writers while maintaining a consistent API for the scene graph nodes. Writers implementing this protocol handle immediate data persistence without keeping data in memory.

store: zarr.Group
write_group(path: str, *, _transform_normalized: bool = False, **attrs: Any) → None[source]

Create a group structure in the Zarr store.

Parameters:
  • path – Path within the Zarr store for the group

  • _transform_normalized – Internal control parameter (NOT a group attribute — never persisted). Set to True by the Node/Scene API to declare that transform / nd_transform in attrs are already in the on-disk normalized form (column-major flat list / validated dict), so the writer must not normalize them a second time (the conversion is not idempotent). All other callers leave the default False and get full normalization + validation.

  • **attrs – Attributes to attach to the group

write_points(path: str, positions: NDArray[float32] | NDArray[float16], colors: NDArray[float32] | NDArray[uint8] | NDArray[uint16] | tuple | list | None = None, radii: NDArray[float32] | NDArray[float16] | NDArray[uint8] | float | None = None, sharpness: NDArray[float32] | NDArray[float16] | NDArray[uint8] | float | None = None, scalars: NDArray[float32] | NDArray[float16] | NDArray[uint8] | float | None = None, labels: Sequence[str] | None = None, image_labels: Any | None = None, keys: Sequence[str] | None = None, **attrs: Any) → Dict[str, Any][source]

Write points data immediately to Zarr.

Data is written directly to disk without being kept in memory. Only metadata about the written data is returned.

The actual dtypes used for storage depend on EncodingMode: - AUTO mode: Analyzes data and selects optimal encoding - PRECISION mode: Uses float32 for maximum precision - MEMORY mode: Aggressively quantizes (float16/uint8)

Scalar Convenience (v1.4.0): Optional attributes accept scalars: - radii=0.5 instead of np.full(N, 0.5) - colors=(1.0, 0, 0) instead of np.full((N, 3), [1,0,0]) - sharpness=0.5 instead of np.full(N, 0.5)

Parameters:
  • path – Path within the Zarr store for this points

  • positions – Point positions array of shape (N, D) - float32 or float16 (NOT scalar - positions must be full arrays)

  • colors – Optional - array of shape (N, 3), tuple/list (R,G,B), or None

  • radii – Optional - array of shape (N,), scalar float, or None

  • sharpness – Optional - array of shape (N,), scalar float, or None

  • scalars – Optional - array of shape (N,), scalar float, or None. Used for colormap lookup when a colormap is applied.

  • labels – Optional list of strings, one per point, for hover tooltips

  • image_labels – Optional per-element images for hover thumbnails

  • keys – Optional machine-readable strings, one per point, for link and copy templates.

  • **attrs – Additional attributes for the points

Returns:

  • n_points: Number of points written

  • ndim: Dimensionality of the points

  • path: Path where data was written

  • has_colors: Whether colors were written

  • has_radii: Whether radii were written

  • has_sharpness: Whether sharpness was written

Return type:

Dictionary containing only metadata about the written data

write_lines(path: str, vertices: NDArray[float32] | NDArray[float16], widths: NDArray[float32] | NDArray[float16] | NDArray[uint8] | float, colors: NDArray[float32] | NDArray[uint8] | NDArray[uint16] | tuple | list | None = None, sharpness: NDArray[float32] | NDArray[float16] | NDArray[uint8] | float | None = None, scalars: NDArray[float32] | NDArray[float16] | NDArray[uint8] | float | None = None, indices: NDArray[uint32] | None = None, line_type: str = 'polyline', labels: Sequence[str] | None = None, image_labels: Any | None = None, keys: Sequence[str] | None = None, **attrs: Any) → Dict[str, Any][source]

Write lines data immediately to Zarr.

Scalar Convenience (v1.4.0): Uniform attributes accept scalars: - widths=0.1 instead of np.full(N, 0.1) - colors=(1.0, 0, 0) instead of np.full((N, 3), [1,0,0]) - sharpness=0.5 instead of np.full(N, 0.5)

Parameters:
  • path – Path within the Zarr store for this lines node

  • vertices – Vertex positions array of shape (N, D)

  • widths – Line widths - array of shape (N,) or scalar float

  • colors – Optional - array of shape (N, 3), tuple/list (R,G,B), or None

  • sharpness – Optional - array of shape (N,), scalar float, or None

  • scalars – Optional - array of shape (N,), scalar float, or None. Used for colormap lookup when a colormap is applied.

  • indices – Optional vertex indices for indexed line type

  • line_type – Type of line connectivity

  • labels – Optional list of strings, one per vertex, for hover tooltips

  • image_labels – Optional per-element images for hover thumbnails

  • keys – Optional machine-readable strings, one per vertex, for link and copy templates.

  • **attrs – Additional attributes for the lines

Returns:

Dictionary containing metadata about the written lines

write_mesh(path: str, vertices: NDArray[float32] | NDArray[float16], faces: NDArray[uint32], normals: NDArray[float32] | None = None, normal_dims: Sequence[int] | None = None, colors: NDArray[float32] | NDArray[uint8] | NDArray[uint16] | tuple | list | None = None, scalars: NDArray[float32] | NDArray[float16] | NDArray[uint8] | float | None = None, uvs: NDArray[float32] | None = None, texture: NDArray[Any] | None = None, texture_encoding: str = 'raw', texture_width: int | None = None, texture_height: int | None = None, texture_channels: int | None = None, texture_color_space: str = 'srgb', texture_ktx2_mode: str = 'uastc', texture_ktx2_quality: int | None = None, texture_ktx2_rdo_l: float | None = None, texture_ktx2_zcmp: int | None = None, shading: str | None = None, double_sided: bool = True, labels: Sequence[str] | None = None, image_labels: Any | None = None, keys: Sequence[str] | None = None, **attrs: Any) → Dict[str, Any][source]

Write a triangle mesh immediately to Zarr.

The surface geometry type: nD vertices plus a faces triangle-index array. It carries no per-element size — a triangle’s extent comes from its own vertices — and has no spatial index: the loader is whole-node. This call writes one leaf; LOD is layered on by the caller, not by this method — see write_mesh_multi_lod (additive reveal levels) below.

Parameters:
  • path – Path within the Zarr store for this mesh node

  • vertices – Vertex positions of shape (V, D)

  • faces – Triangle vertex indices, (F, 3) or flat (3F,)

  • normals – Optional per-vertex normals of shape (V, 3); requires normal_dims

  • normal_dims – The three dimension indices normals describes — required with normals, rejected without

  • colors – Optional - array of shape (V, 3|4), tuple/list (R,G,B[,A]), or None

  • scalars – Optional - array of shape (V,), scalar float, or None. Used for colormap lookup when a colormap is applied.

  • uvs – Optional (V, 2) per-vertex texture coordinates. Required with texture and refused without it. Values outside [0, 1] are legal and tile under texture_wrap="repeat".

  • texture – Optional base-colour image. (H, W, C) array under texture_encoding="raw" or "ktx2", else a 1-D uint8 array of encoded bytes. Mutually exclusive with colors and colormap — a mesh has one base-colour source.

  • texture_encoding – raw | png | webp | jpeg | ktx2. KTX2 accepts uint8 RGB/RGBA input and requires the Khronos toktx executable, version 4.1.0 or newer. HDR requires raw.

  • texture_ktx2_mode – uastc (default) or etc1s.

  • texture_ktx2_quality – Codec quality; defaults to 2 for UASTC and 128 for ETC1S.

  • texture_ktx2_rdo_l – UASTC RDO lambda in [0.001, 10.0]; defaults to 0.25. Lower values preserve more quality and produce larger files; 0 disables RDO while retaining zstd compression.

  • texture_ktx2_zcmp – UASTC zstd level in [1, 22]; defaults to 9.

  • texture_width – Declared width. Required for encoded payloads, where it cannot be read without decoding; read off the array for raw or ktx2, and refused if it disagrees.

  • texture_height – Declared height. Same contract as texture_width.

  • texture_channels – Declared channels — 1, 3 or 4. Same contract.

  • texture_color_space – srgb (default) or linear. An ordinary PNG/JPEG is sRGB-encoded; declaring it wrong gives a subtly over-dark or washed-out surface rather than an obvious failure. HDR raw textures must be declared linear.

  • shading – "smooth" / "flat" / unlit "none"; defaults by normal presence

  • double_sided – Whether back faces render (default True)

  • labels – Optional list of strings, one per vertex, for hover tooltips

  • image_labels – Optional per-element images for hover thumbnails

  • keys – Optional machine-readable strings, one per vertex, for link and copy templates.

  • **attrs – Additional attributes for the mesh

Returns:

Dictionary containing metadata about the written mesh

write_gsplats(path: str, centers: NDArray[float32] | NDArray[float16], amplitudes: NDArray[float32] | NDArray[float16] | NDArray[uint8] | float, cholesky_factors: NDArray[float32], colors: NDArray[float32] | NDArray[uint8] | NDArray[uint16] | tuple | list | None = None, label_ids: ndarray[Any, Any] | Sequence[int] | None = None, label_vocabulary: dict[int, str] | None = None, labels: Sequence[str] | None = None, image_labels: Any | None = None, keys: Sequence[str] | None = None, _source_dtype: str | None = None, **attrs: Any) → Dict[str, Any][source]

Write Gaussian splats data immediately to Zarr.

Scalar Convenience (v1.4.0): Uniform attributes accept scalars: - amplitudes=1.0 instead of np.full(N, 1.0) - colors=(1.0, 0, 0) instead of np.full((N, 3), [1,0,0])

Parameters:
  • path – Path within the Zarr store for this gsplats node

  • centers – Splat center positions array of shape (N, D)

  • amplitudes – Amplitude values - array of shape (N,) or scalar float

  • cholesky_factors – Packed Cholesky factors array of shape (N, k)

  • colors – Optional - array of shape (N, 3), tuple/list (R,G,B), or None

  • label_ids – Optional non-negative integer class id per splat

  • label_vocabulary – Explicit mapping from stored class ids to names

  • labels – Optional list of strings, one per splat, for hover tooltips

  • image_labels – Optional per-element images for hover thumbnails

  • keys – Optional machine-readable strings, one per splat, for link and copy templates.

  • **attrs – Additional attributes for the gsplats

Returns:

Dictionary containing metadata about the written gsplats

write_points_multi_lod(path: str, levels: list, *, extend_to_all: List[str] | None = None, **attrs: Any) → dict[source]

Write multi-additive-LOD Points: parent node + additive_<i>/ subgroups.

Each level is a dict with positions plus optional colors / radii / sharpness / scalars / labels / keys. See luxar.core.group.lod.points.make_additive_lod_points() for the level-construction helper that produces the input.

labels and keys are independently all-or-nothing across the ladder and are NOT written per level. Each present channel gets one CSR pair on the parent, spanning the levels in stored order; the subgroups carry neither channel.

write_lines_multi_lod(path: str, levels: list, *, extend_to_all: List[str] | None = None, **attrs: Any) → dict[source]

Write multi-additive-LOD Lines (polyline-level granularity).

Each level is a dict with vertices + widths + segments plus optional colors / sharpness / scalars / labels / keys and n_polylines (summed into the returned metadata only — the parent group does not stamp it). See luxar.core.group.lod.lines.make_additive_lod_lines() for the helper that produces the input.

labels and keys are independently all-or-nothing across the ladder and are NOT written per level. Each present channel gets one per-vertex CSR pair on the parent, spanning the levels in stored order; the subgroups carry neither channel.

write_sound(path: str, payload: bytes, fmt: str, positions: NDArray[float32] | NDArray[float16] | None, *, sound_attrs: dict[str, Any], **attrs: Any) → dict[str, Any][source]

Write a sound node: an opaque MP3/AAC clip + optional positions.

fmt is "mp3" or "aac" (already sniffed by the adder); sound_attrs carries the validated playback / spatial / licence knobs and **attrs the compositing pass-throughs. Returns the node metadata.

write_mesh_multi_lod(path: str, levels: list, *, extend_to_all: List[str] | None = None, **attrs: Any) → dict[source]

Write multi-additive-LOD Mesh (a reveal ladder, face granularity).

Each level is a dict with vertices + faces — the faces already re-indexed into that level’s own gathered vertex table — plus optional normals / normal_dims / colors / scalars / shading / double_sided / _scalar_data_range and lod_stats. See luxar.core.group.lod.mesh.make_additive_lod_mesh() for the helper that produces the face groups and luxar.mesh.split.split_mesh_by_faces() for the re-indexing.

labels is REFUSED, not carried — the one place this diverges from its three siblings. They write ONE union CSR on the parent spanning the levels; a mesh level re-indexes its own vertices, so a single source vertex maps to a slot in several levels and the union index space is ill-defined. Labels survive on the substitutive_lod= and partition= paths instead.

create_resizable_dataset(path: NodePath, dtype: np.dtype, shape: Tuple[int, ...], maxshape: MaxShape = None, chunks: ChunkSpec = True) → zarr.Array[source]

Create a resizable dataset.

Part of writer protocol for potential future extensions.

Parameters:
  • path – Path for the dataset within the Zarr store

  • dtype – Data type for the dataset (numpy dtype)

  • shape – Initial shape of the dataset

  • maxshape – Accepted and ignored — see LuxarZarrCompiler.create_resizable_dataset(). zarr has never enforced a maximum shape; the argument is h5py heritage.

  • chunks – Chunk configuration for the dataset

Returns:

Zarr array handle that supports resizing and slicing

delete_group_attr(path: str, key: str) → None[source]

Remove an attribute from a group in the Zarr store.

Parameters:
  • path – Path within the Zarr store for the group

  • key – Attribute key to remove

node_exists(path: str) → bool[source]

Return whether a node path exists for an internal rollback guard.

delete_node(path: str) → None[source]

Delete a subtree during rollback, without editing the scene graph.

snapshot_rollback_state() → Tuple[dict[str, List[float]] | None, frozenset[Tuple[str, str]], bool, Tuple[dict[tuple, tuple[str, str]], dict[str, str]], Tuple[bool, bool] | None][source]

Capture compiler state that deleted geometry writes may have changed.

restore_rollback_state(state: Tuple[dict[str, List[float]] | None, frozenset[Tuple[str, str]], bool, Tuple[dict[tuple, tuple[str, str]], dict[str, str]], Tuple[bool, bool] | None]) → None[source]

Restore compiler state captured before a rolled-back write.

transaction(path: str) → ContextManager[None][source]

Roll back a newly-created subtree and compiler state on failure.

finalize() → None[source]

Finalize the Zarr store.

Performs any necessary cleanup, metadata consolidation, or optimization steps before closing the store.

property store_path: str

Get the current path to the underlying Zarr store.

Writers may use a staging directory while active and return a different published path after finalize().

Returns:

Current staging or finalized store path.

property final_store_path: str

Get the path where the finalized Zarr store will be published.

__init__(*args, **kwargs)

ZarrWriterProtocol is a Protocol (abstract interface), not a concrete class. It defines the writer contract that any Zarr writer implementation must satisfy. This allows for alternative implementations (e.g., remote writers, streaming writers) while maintaining compatibility with the compiler.

The contract is enforced, not merely declared: LuxarZarrCompiler is checked against it by mypy with no type: ignore[override] escapes, and test_writer_protocol_agreement.py additionally pins the parameter names, order and defaults of every method it declares. Order matters to a caller because these are positional-or-keyword parameters, and a parameter present in one signature but not the other renumbers every argument after it.

Volume Loading

Load an nD image / volume from .zarr, .zarr.zip, .tiff, .npy, or .npz — with channel/timepoint selection and axis-order handling — for Gaussian splat fitting and calibration.

Multi-format volume loading (domain layer).

Loads a spatial volume from .npy / .npz / .zarr / .zarr.zip / .tiff / imageio-supported files, with OME-Zarr-aware positional slicing and an explicit --axes override. This is reusable domain logic (no CLI/Typer coupling): missing optional readers raise ImportError with an install hint, which the CLI converts to a clean exit. Previously lived in luxar.cli.gsplat_config; moved here so domain code (e.g. the denoise pipeline) no longer imports upward into the CLI layer.

luxar.io.volume.load_volume(path: Path, channel: int | None = None, timepoint: int | None = None, array_key: str | None = None, axes: str | None = None, info: dict | None = None, region: Tuple[slice, ...] | None = None) → ndarray[source]

Load a volume from various file formats.

Supported formats:

.npy — NumPy binary (numpy, base dep) .npz — NumPy compressed (numpy, base dep) .zarr — Zarr array/group, including OME-ZARR 5D (zarr, base dep) .tiff / .tif — TIFF image (tifffile, optional: pip install luxar[io]) other — Fallback via imageio (optional: pip install luxar[io])

Parameters:
  • path – Path to the volume file

  • channel – Channel index for 4D/5D+ OME-ZARR data. With axes, this is a flat row-major index across every channel-like axis. If None, defaults to 0 when slicing is needed; for 4D arrays without axes, None returns the array as-is.

  • timepoint – Timepoint index for 5D+ OME-ZARR data. If None, defaults to 0 when slicing is needed; for 4D arrays, None returns the array as-is.

  • array_key – Array key within .npz or .zarr files

  • axes – Explicit per-dimension axis labels (e.g. "z,c,y,x") overriding the positional TCZYX/CZYX/ZYX heuristic — for data whose axis order differs. The single time axis is sliced by timepoint; the flat channel index is decoded across all channel-like axes. Those axes are dropped and spatial axes are kept in the given order.

  • info – Optional dict, populated with source_dtype — the element type of the array AS STORED, captured before the float32 cast below. This is the only place it is knowable: the returned array is always float32, so a consumer that wants to quote a size (e.g. the denominator of a compression ratio) would otherwise describe the working copy and overstate it by the cast’s inflation factor — exactly 2x for the 16-bit acquisitions most microscopy produces.

  • region – Optional spatial slices applied before materializing a zarr array. Non-zarr formats are sliced after loading.

Returns:

Volume as float32 numpy array (>=2D)

Raises:
  • ImportError – If an optional reader (tifffile/imageio) is required but not installed. Callers that want a clean CLI exit should catch this.

  • ValueError – On an invalid key or a sub-2D volume.

OME-Zarr Discovery

Discover the shape, axis labels, and voxel size of an OME-Zarr multiscale dataset without loading the pixel data.

OME-Zarr shape discovery (domain layer).

Discovers the shape and axis structure of an OME-Zarr / NGFF dataset (T/C/Z/Y/X layout, voxel size, unit, resolution levels), with fallbacks to a custom axes attribute and a shape-based heuristic. Reusable domain logic with no CLI/Typer coupling. Previously lived in luxar.cli.gsplat_config.

luxar.io.ome_zarr.discover_ome_zarr_shape(path: Path, axes_override: List[str] | None = None, array_key: str | None = None) → OMEZarrInfo[source]

Discover the shape and axis structure of an OME-Zarr dataset.

Which array is read is ONE rule shared with the two volume-loading entry points (_select_zarr_array()), so a re-fit that re-opens a store cannot land on a different (e.g. downsampled) array than the command that read its shape.

Parses that array’s NGFF multiscales metadata in BOTH OME-Zarr layouts — top-level (0.4) and nested under an ome key (0.5); see resolve_ngff_attrs(). The block is read off the GROUP THAT OWNS the selected array when that group’s block is evidence about it (_owner_ngff_attrs()), off the root otherwise. A block whose axes count disagrees with the SELECTED array’s ndim is not metadata about that array (a 5D image beside its 3D labels/…) and is skipped. Falls back to a custom axes attribute — root-first, owner second, see _usable_custom_axes() — then to a shape-based heuristic (5D→TCZYX, 4D→CZYX, 3D→ZYX) for non-NGFF zarr stores. That last fallback GUESSES the T/C roles and recovers no voxel size, so for an ambiguous (≥4D) store it says so on the console — stating whether nothing was declared or something was declared but unusable — and points at axes_override.

Accepts both plain .zarr directories and .zarr.zip archives — zarr’s ZipStore handles the latter transparently.

Parameters:
  • path – Path to the .zarr store or .zarr.zip archive.

  • axes_override – Explicit axis labels (e.g. ["time","channel","z","y","x"]). Overrides axis classification and names while retaining positional NGFF units/scales from the selected array when available.

  • array_key –

    Key path to a specific array within the zarr store (e.g. "h2afva/fused"). When provided, skips auto-selection and navigates directly to this array (which may itself be a group — then its own full-resolution level is taken).

    The datasets[] entry for the voxel size is matched against the array that was SELECTED, whether or not a key was passed — so a coarser pyramid level reports its own spacing, not level 0’s, and that holds for the auto path too (which can perfectly well land on a coarser level, e.g. when level 0’s declared path does not resolve). The match is exact, but on the NORMALISED spellings, so a block writing its own levels explicitly relative ("./0" for the array at "0") still names them; see _selected_dataset().

Returns:

OMEZarrInfo with discovered metadata.

Raises:

ValueError – If the zarr store has no arrays, array_key is not found, or the store is unreadable.

luxar.io.ome_zarr.resolve_ngff_attrs(attrs: Mapping[str, Any]) → Dict[str, Any][source]

The mapping that actually carries multiscales, for either OME-Zarr layout.

OME-Zarr 0.4 puts multiscales at the top level of a node’s attributes. OME-Zarr 0.5 nests the whole NGFF block one level down under an ome key. Reading only the 0.4 spelling on a 0.5 store fails SILENTLY — no metadata is found, so axis roles get guessed from the shape and no physical voxel size is recovered — which is why every reader goes through here instead of spelling attrs["multiscales"] itself.

The layout can NOT be inferred from the store’s zarr format version: a zarr v3 store written by a 0.4-era tool has the v3 chunk layout with 0.4 (top-level) attributes. Both spellings are therefore always tried.

The predicate is deliberately multiscales-specific rather than “any NGFF key”: a store may carry a top-level 0.4 multiscales and an ome block holding only rendering metadata (omero), and selecting that block on the strength of omero alone would throw away the pyramid the caller came for — the exact silent mis-read this function exists to prevent, from the other side.

Presence of the key is not enough either, for the same reason: a store carrying a good top-level 0.4 pyramid alongside {"ome": {"multiscales": []}} would have the pyramid discarded and be reported as declaring an empty multiscales — false of the store. The nested block therefore wins only when its multiscales is a NON-EMPTY list, or when the top level declares no multiscales at all (in which case the nested one, empty or malformed as it may be, is the only thing the store said, and handing it back is what lets the caller report “declared but unusable” instead of “nothing declared”).

Defensive by design — ome may be absent, not a mapping, or a mapping with no multiscales. In each of those cases the top-level attributes are returned, so a caller’s own “no multiscales here” handling runs as it did before rather than this raising.

Parameters:

attrs – A zarr node’s user attributes (e.g. dict(group.attrs)).

Returns:

Either the nested ome block (0.5) or attrs itself (0.4), as a plain dict.

luxar.io.ome_zarr.ngff_scale_transform(transforms: Any) → List[float] | None[source]

The scale vector of a NGFF coordinateTransformations list, if any.

The list is SEARCHED for the type == "scale" entry rather than indexed at [0]: the spec allows a translation (or any other transform) to come first, and transforms[0]["scale"] then raises on a perfectly valid store.

Never raises. Anything malformed — not a list, no scale entry, a scale that is not a sequence, or a component that is not a number (null, a non-numeric string) — yields None, i.e. “no spacing declared”, which is how every other malformed-metadata path in this module degrades. Numeric STRINGS still convert, since that is how some writers spell a float.

class luxar.io.ome_zarr.OMEZarrInfo(axes: ~typing.List[str], shape: ~typing.Tuple[int, ...], n_timepoints: int, n_channels: int, channel_axes: ~typing.List[str], channel_shape: ~typing.Tuple[int, ...], spatial_shape: ~typing.Tuple[int, ...], spatial_axes: ~typing.List[str], time_axis: int | None = None, channel_indices: ~typing.Tuple[int, ...] = (), spatial_indices: ~typing.Tuple[int, ...] = (), axis_units: ~typing.List[str | None] = <factory>, axis_scales: ~typing.Tuple[float, ...] | None = None, voxel_size: ~typing.Tuple[float, ...] | None = None, unit: str | None = None, resolution_levels: int = 1, path: ~pathlib.Path | None = None)[source]

Bases: object

Metadata about an OME-Zarr dataset’s structure.

axes: List[str]

Axis labels, e.g. ["t", "c", "z", "y", "x"].

shape: Tuple[int, ...]

Full array shape at highest resolution.

n_timepoints: int

Size of the T dimension (1 if absent).

n_channels: int

Number of flat channel tasks (product of channel-like axes, or 1).

channel_axes: List[str]

Axis labels folded into the flat channel task index.

channel_shape: Tuple[int, ...]

Shape of axes folded into the flat channel task index.

spatial_shape: Tuple[int, ...]

ZYX (or YX) portion of the shape.

spatial_axes: List[str]

Spatial axis labels, e.g. ["z", "y", "x"].

time_axis: int | None = None

Index into shape of the time axis, or None when there is none.

Part of the (time_axis, channel_indices, spatial_indices) trio below: the decomposition discovery ACTUALLY used to derive n_timepoints / n_channels / spatial_shape.

channel_indices: Tuple[int, ...] = ()

Indices into shape of the axes folded into the flat channel index.

In the order they are folded, so decode_flat_channel_index(c, channel_shape) maps position-for-position onto them.

spatial_indices: Tuple[int, ...] = ()

Indices into shape of the spatial axes, in spatial_shape order.

Publishing all three indices closes a standing hazard: a consumer that needs to know which axis played which role had to RE-DERIVE it from axes with a second vocabulary, and the vocabularies disagree in both directions. NGFF metadata is classified by the type field, so a channel axis named stain is a channel here but not to any name-driven rule, while an axis typed view falls through to SPATIAL here but is channel-like to classify_axis_labels(). Two vocabularies deciding the same question is how a consumer silently plans against a layout discovery never reported — read these fields instead of re-classifying the labels.

axis_units: List[str | None]

Physical unit for each source axis, position-for-position with axes.

axis_scales: Tuple[float, ...] | None = None

Composed NGFF scale for every source axis, or None when unavailable.

voxel_size: Tuple[float, ...] | None = None

Physical spacing from coordinateTransformations (spatial axes only).

unit: str | None = None

Physical unit string (e.g. "micrometer").

resolution_levels: int = 1

Number of multiscale levels.

path: Path | None = None

Path to the zarr store.

__post_init__() → None[source]

Keep positional axis metadata aligned with the discovered axes.

__init__(axes: ~typing.List[str], shape: ~typing.Tuple[int, ...], n_timepoints: int, n_channels: int, channel_axes: ~typing.List[str], channel_shape: ~typing.Tuple[int, ...], spatial_shape: ~typing.Tuple[int, ...], spatial_axes: ~typing.List[str], time_axis: int | None = None, channel_indices: ~typing.Tuple[int, ...] = (), spatial_indices: ~typing.Tuple[int, ...] = (), axis_units: ~typing.List[str | None] = <factory>, axis_scales: ~typing.Tuple[float, ...] | None = None, voxel_size: ~typing.Tuple[float, ...] | None = None, unit: str | None = None, resolution_levels: int = 1, path: ~pathlib.Path | None = None) → None

Chunk-Layout Optimization

Re-chunk an existing store for streaming in one structure-preserving pass — values, codecs, the spatial-index grid, the zarr format version and every attribute except the cache guard (the root’s content_hash, restamped, and the chunk_layout summary beside it) all survive; only zarr chunk shapes change. Backs the luxar optimize command and the luxar info --stats chunk diagnostic.

Re-chunk an existing zarr store in one structure-preserving pass.

luxar optimize exists because a store that is already on disk is usually badly chunked for STREAMING, and regenerating it is not an option: the source volume may be gone, the fit may have cost GPU-days, and the copy on Zenodo has a DOI. The demo corpus measured 606,349 files for 2.94 GB — an average file of 5.1 KB, 97.0% of them under 16 KB — against a TARGET_CHUNK_BYTES of 64 KB. Cold-loading one demo from object storage took 245 s / 9,390 requests as written and 51 s / 2,348 requests after this pass. It was never bandwidth-bound (38.6 MB over 245 s); it was round trips.

Fixing the WRITER (which #1713 did) does not subsume this, for three reasons. It cannot reach data that already exists. It sizes its byte budget from the array’s INPUT dtype — float32 positions, before the encoder quantizes them to uint16 — so its chunks land 2-4x under target, whereas a post-hoc pass reads the STORED dtype and sees the real itemsize. And authoring and serving want different targets, which is a flag rather than a rebuild.

What is preserved, exactly

Everything except the zarr chunk grid: array VALUES bit-for-bit, dtype, compressor, filters, serializer, fill_value, memory order, the on-disk zarr FORMAT (a v2 store stays v2 — format conversion is gsplat migrate-format’s job, not this one), and every group and array attribute EXCEPT the two the pass is contractually required to move: the root’s content_hash, which is restamped, and the chunk_layout summary written beside it. Those two are the cache-invalidation guard described below, and dropping the restamp ships a re-chunked store under the source’s hash. The plain non-zarr files a group’s attrs name — an overlay image, which no array or group API reaches — are copied across byte-for-byte with everything else, or the pass refuses rather than dropping them (_copy_payload_files() states which names it will not write and why). Codecs are reused from the SOURCE array rather than re-derived, because an omitted compressor is not “no compressor” (zarr’s "auto" is Blosc/lz4 at format 2 and zstd at format 3) and some Luxar arrays are deliberately RAW.

The spatial-index contract

chunk_size and chunk_bounds ARE the viewer’s partition grid, and this pass must not move it. Every emitted chunk is therefore a whole multiple of the node’s chunk_size atom — rounded DOWN from the byte budget, never below one atom — so a partition’s row range still falls inside a single zarr chunk and no row-range read straddles a boundary. The bounds arrays themselves are never re-chunked.

The cache-invalidation hazard

The viewer’s MultiLevelCachingStore validates its persistent cache by comparing content_hash against the remote, and that cache holds ENCODED CHUNKS KEYED BY CHUNK INDEX. A re-chunk that leaves the hash where it was is therefore the worst thing this pass could do: chunk key 0/0 covers a different row range while a warm client believes itself up to date, so it serves bytes that no longer mean what their keys say. Silent wrong data, on the one path with no error to raise.

Both hashers now fold layout in themselves. compute_content_hashes() hashes each array’s STORAGE IDENTITY — name, shape, dtype, chunks, shards, codec ids and the array’s own attrs — before its values, and the .gsplats.zarr stamp (save_gsplats._stamp_content_hash) folds the same identity terms over metadata alone; both key each child group by its NAME. Only the value walk also folds the bytes of a group’s plain payload files — the gsplat stamp has no payload handling at all. So recomputing over the re-chunked output lands on a different digest by construction. What makes that reach the viewer is the RESTAMP: nothing else rewrites the stored content_hash, and the stored one is what gets compared. Suppress it and the output ships under the source’s hash however far the grid moved.

The chunk_layout summary attr the pass writes on the root is folded in too, since attrs are hashed. It used to be the ONLY thing moving the digest — back then, suppressing it was measured to leave the output hash byte-identical to the source’s — and against a layout-aware hasher it is now belt-and-braces there. It is kept because it documents what the pass did, and because it is still the whole guard in the hash-less case below.

The attr is written for a NON-Luxar store too, and that is not tidiness either. A store with a kind marker but no content_hash gets no restamp, so the viewer’s MultiLevelCachingStore validation queue falls back to a SHA-256 of the raw root document bytes. At format 3 the root zarr.json carries the chunk grid, so that token moves whatever we write; at format 2 the probe reads .zattrs, which — without chunk_layout — would be byte-identical to the source’s, and the warm cache would again believe itself current. “Suppress the attr on non-Luxar stores” is therefore a plausible-sounding simplification that silently breaks exactly the hash-less format-2 case.

luxar.io.optimize.optimize_store(source_path: str | Path, dest_path: str | Path, *, target_bytes: int = 65536, profile: str | None = None, overwrite: bool = False, verify: bool = False, generic: bool = False) → OptimizePlan[source]

Copy source_path to dest_path, re-chunked, values untouched.

Raises rather than writing in place: the pass reads the source while writing the destination, so a destination that IS the source — or contains it, or lives inside it — is refused (_check_destination_path()). overwrite only governs replacing a DIFFERENT existing output, and only one that is a zarr store or an empty directory; it is opt-in because “publish under a new URL prefix” is the real fix for a warm client cache and cannot be enforced from here.

All-or-nothing. The output is built in a hidden sibling directory and renamed onto dest_path only after the copy — and verify, when asked for — has succeeded, so a failure leaves neither a half-written store at the user’s path nor a damaged previous one. A previous store is moved ASIDE rather than deleted first (_replace()), so a failed swap restores it instead of costing both copies.

luxar.io.optimize.plan_optimization(root: Group, *, target_bytes: int = 65536, profile: str | None = None) → OptimizePlan[source]

Plan the re-chunk of root without writing anything.

Every leaf is planned independently — a kind=lod level, a kind=partition part and an additive_<i> rung each carry their own atom — so nesting needs no special case beyond the recursive walk.

luxar.io.optimize.summarize_chunk_layout(root: Group) → ChunkLayoutSummary[source]

Measure a store’s streaming shape: chunk sizes and request count.

luxar.io.optimize.summarize_plan(plan: OptimizePlan) → ChunkLayoutSummary[source]

The same diagnostic, off a plan that has already been walked.

luxar info --stats wants both the summary and a real plan (the “try luxar optimize” hint is gated on one), and the corpus this tool exists for holds 606,349 files — so the two share one walk rather than opening every array twice.

Counts OBJECTS, not nominal grid cells: a shard is one file however many chunks it packs, and a (0, D) placeholder is none at all.

luxar.io.optimize.resolve_target_bytes(*, target_bytes: int | None = None, target_kb: int | None = None, profile: str | None = None) → int[source]

Turn the three mutually exclusive size flags into one byte budget.

luxar.io.optimize.CHUNK_PROFILES: dict[str, int] = {'archive': 1048576, 'hosting': 262144, 'local': 65536}

Byte targets selected by --profile.

local is the authoring default: a 64 KB chunk is the balance point between request count and over-fetch when a read costs microseconds. hosting is the top of the documented 16-256 KB band — over object storage a round trip costs ~100 ms and over-fetching a few hundred KB is free by comparison. archive deliberately leaves that band: an archived store is not being streamed, and its only real cost is the file COUNT (upload time, inode pressure, per-object storage minimums).

The larger targets trade PARTIAL-QUERY bytes for FULL-LOAD requests, so “object storage, therefore hosting” is not unconditional. A Points or GSplats node is not loaded whole — the viewer turns visible index chunks into element ranges of chunk_size atoms, and one atom-hit costs one zarr chunk whatever its size. Measured on a real store (atom 2340, uint16 (N, 3)): local 4 atoms / 54.8 KB per partial hit, hosting 18 atoms / 246.8 KB (4.5x), archive 74 atoms / 1014.6 KB (18x). Size up when the access pattern is “load the node whole”; stay on local when the viewer will be slicing into a large one.

On an ANIMATED node (a non-displayed axis the user plays) the useful measure is frames-per-chunk: group the ordered chunk_bounds atoms exactly as the planned zarr chunks will group them, then measure each group’s inclusive hidden-axis span. The writer already orders slice-major, so a 1 MB archive chunk of a 250-frame Lines node holds ~6 frames. The viewer now prefetches the nearest next zarr chunk boundary across a node’s arrays while those preceding frames play (#2686); before that, every boundary produced a measured 0.4-0.9 s stall on a 7 Mbps link. A LADDERED played node is the opposite case: a coarse rung that fits one chunk stays cache-resident and serves every timepoint (#2377), which is why the 4D splat demos ship at 1 MB. The per-array planner therefore stays byte based, while the whole-store plan warns when an actively played, un-laddered node would exceed two frames per chunk across multiple chunks.

class luxar.io.optimize.OptimizePlan(target_bytes: int, profile: str | None, arrays: list[ArrayPlan] = <factory>, playback_warnings: list[PlaybackChunkWarning] = <factory>)[source]

Bases: object

The whole-store plan — what optimize_store() will do, in advance.

target_bytes: int
profile: str | None
arrays: list[ArrayPlan]
playback_warnings: list[PlaybackChunkWarning]
property n_rechunked: int

How many arrays get a new chunk grid.

property source_n_chunks: int

Chunk files the input store holds — the requests a full load costs now.

property target_n_chunks: int

Chunk files the output will hold — the requests it will cost instead.

__init__(target_bytes: int, profile: str | None, arrays: list[ArrayPlan] = <factory>, playback_warnings: list[PlaybackChunkWarning] = <factory>) → None
class luxar.io.optimize.ArrayPlan(path: str, shape: tuple[int, ...], dtype: str, itemsize: int, source_chunks: tuple[int, ...], target_chunks: tuple[int, ...], source_n_chunks: int, target_n_chunks: int, atom: int | None, source_file_grid: tuple[int, ...], skip_reason: str = '')[source]

Bases: object

What will happen to one array, and why.

path: str
shape: tuple[int, ...]
dtype: str
itemsize: int

dtype.itemsize, carried rather than re-derived from dtype. str(np.dtype(...)) is NOT a round trip for every dtype — a structured dtype stringifies to "[('a', '<i4'), ('b', '<f8')]" and numpy’s variable-width StringDType to "StringDType()", neither of which np.dtype() accepts — so re-parsing it to reach the itemsize raised TypeError on exactly the third-party stores (AnnData/cellxgene, an OME-Zarr label table) --generic exists for: luxar info --stats exited 1 on a store its own --format json path handled, and optimize --dry-run died with a traceback while the real copy of the same store succeeded.

source_chunks: tuple[int, ...]
target_chunks: tuple[int, ...]
source_n_chunks: int
target_n_chunks: int
atom: int | None
source_file_grid: tuple[int, ...]

The grid whose cells are FILES — the shard grid when the array is sharded, its chunk grid otherwise. Recorded rather than re-derived so summarize_plan() can measure a store off the plan’s single walk.

skip_reason: str = ''

Empty when the array is being re-chunked; otherwise the reason it is not.

property rechunked: bool

True when this array gets a new chunk grid.

property target_chunk_bytes: int

Nominal payload of the planned chunk, in bytes.

property source_file_bytes: int

Nominal payload of one SOURCE object (a shard when sharded).

__init__(path: str, shape: tuple[int, ...], dtype: str, itemsize: int, source_chunks: tuple[int, ...], target_chunks: tuple[int, ...], source_n_chunks: int, target_n_chunks: int, atom: int | None, source_file_grid: tuple[int, ...], skip_reason: str = '') → None
class luxar.io.optimize.ChunkLayoutSummary(n_arrays: int, n_chunks: int, mean_chunk_bytes: float, n_arrays_under_floor: int)[source]

Bases: object

The streaming-shape diagnostic luxar info --stats reports.

Computed off the same walk the optimizer plans from, so “is this store worth optimizing?” is answerable without hosting it first.

n_arrays: int

Arrays that produce at least one object. A (0, D) array_ref placeholder writes no chunk and costs no request, so counting it would put arrays nobody fetches in the denominator of the floor share below — measured on a 6-node scene with deduplicated positions, that reported “23/24 arrays (96%)” where the honest answer is 13/14 (93%).

n_chunks: int
mean_chunk_bytes: float

Chunk-count-weighted mean of the NOMINAL chunk payload, in bytes.

n_arrays_under_floor: int

Arrays whose nominal chunk payload is below MIN_CHUNK_BYTES. Single-chunk arrays count: a store of a hundred tiny arrays is request-heavy for exactly the reason a store of tiny chunks is. Arrays that fetch nothing do not.

property share_under_floor: float

Fraction of arrays chunked below the 16 KB floor, in [0, 1].

__init__(n_arrays: int, n_chunks: int, mean_chunk_bytes: float, n_arrays_under_floor: int) → None