Demo Site Runbook
How the public Luxar demo gallery and viewer are hosted, and how to publish a new wave of demos to them without breaking anything.
This is an operational document. It records the architecture, the publish
sequence, and — most importantly — the failure modes that produce a plausible
wrong answer rather than an error. No credentials appear here; they live on
the operator’s machine.
The publishing harness named below (gen_landing.py, snapshot_hashes.py,
compare_snapshot.py, and the upload script) also lives outside this
repository on that machine.
1. Architecture
Three hostnames on the luxarviewer.dev zone, each serving a different thing:
Hostname |
Serves |
Backed by |
|---|---|---|
|
The viewer alone, at the root |
Cloudflare Pages project |
|
The gallery page, its media, and its viewer |
Cloudflare Pages project |
|
The |
Cloudflare R2 bucket |
Neither Pages project has a functions/ directory or an R2 binding. Both
are pure static assets. This is the single most important property of the
setup, and it is deliberate: Pages Functions share the Workers free limit of
100,000 requests/day, a limit that fails rather than bills. Serving the
corpus through a Function put one invocation on every chunk fetch, which does
not survive real traffic.
Instead:
Scene data is fetched cross-origin from
data.luxarviewer.dev, an R2 custom domain that bypasses Workers entirely. The gallery emits absolute?src=https://data.luxarviewer.dev/...URLs for this reason. (Requires CORS — §4.3.)Demo-site gallery media (
demos.luxarviewer.dev/media/*, ~180 files, ~485 MiB) are deployed files, served by Pages’ CDN for free and unmetered.Root-README media (
data.luxarviewer.dev/media/<sha256-prefix>.<ext>) are direct R2 objects. Their content-addressed keys are immutable and are recorded inscripts/gallery/media-manifest.json.Source archives (
data.luxarviewer.dev/inputs/<slug>/) may be the only copy of a demo’s pinned inputs. They are not disposable scene-build outputs.Stable demo deep links (
demos.luxarviewer.dev/d/<demo-key>) redirect to that scene under the current dated data prefix. The root README’s gallery titles depend on these routes, so every publish wave must preserve and update them alongside the gallery.
The historical marker header x-luxar-fn: r2 is how to confirm no Worker is in
the path. It should now appear on nothing:
curl -sI "https://demos.luxarviewer.dev/media/earthquakes.webp" | grep -i x-luxar-fn
curl -sI "https://data.luxarviewer.dev/data/<prefix>/<store>.luxar.zarr/<existing-object>" | grep -i x-luxar-fn
# no output from either = correct
Beware when checking this: responses cached before the Function was removed still carry the header. Add a cache-busting query string, or you will conclude the Worker is still deployed when it is not (§3.2 again, applied to your own verification).
A verify-everything sweep of the whole chain lives in the operator’s harness; its essentials are reproduced in §5.
1.1 The viewer takes an absolute src
https://luxarviewer.dev/?src=<absolute-url> opens any store the browser can
reach, which is the point of hosting the viewer at the apex — it is a general
tool, not a demo appendage. That requires CORS on whatever origin holds the
data (§4.3). The gallery uses the same absolute form because its data lives on
the separate R2 hostname; a relative /data/... URL has no server on the Pages
origin.
2. Publishing a wave
A “wave” is: some demos changed on dev, rebuild them, and update the site.
build the changed demos
-> capture gallery media (stills + orbit videos)
-> luxar optimize --profile archive # 1 MB chunk target
-> hash-compare against the LAST LOCAL BUILD
-> audit the complete local scene inventory
-> upload only what changed, to a NEW dated prefix
-> rebuild the page against that prefix, with an ABSOLUTE data host
-> deploy (page + media as static assets)
-> audit the live site
-> purge the superseded prefix
The published corpus deliberately uses archive rather than the general
object-storage hosting profile. On a representative live store it reduced
the chunk count from 5,004 to 224 (22×), accepting larger partial reads in
exchange for far fewer stored objects. Re-measure browser traffic and request
cost before changing that tradeoff.
Read any playback warning after the optimize step. For an un-laddered played
store, re-run at hosting and re-measure requests and per-step bytes before
publishing; §3.21 records the current exception and measurement procedure.
make generate-gallery-datasets automatically runs both built-scene auditors
against only the stores it generated in that invocation, then reports on the
complete local inventory. Both audits gate a newly generated store without
letting an unrelated stale local store block it. Before uploading a wave, run
the complete inventory commands directly as well. A store that fails either
generated-store gate is not re-gated on a later idempotent run because it is
then an already-present neighbour: fix the demo and regenerate with --force
(or delete the store) before continuing.
hatch run check-demo-ladders --require-scenes
hatch run check-scene-credits --require-scenes
Keep this immediately before upload, after the final local build and
optimization. --require-scenes is load-bearing: an empty or wrongly located
inventory must fail rather than produce a green “inspected nothing” result.
Both direct full-inventory commands are the backstop and must exit zero before
upload.
Before rebuilding earthquakes, ocean_currents_earth,
global_rivers_earth, or biodiversity_planetary_scale, put KTX-Software’s
toktx on PATH. Otherwise the basemap silently downgrades to WebP or JPEG;
grep the build log for authoring the Earth basemap as before publishing.
The page generator takes the data prefix as an argument, so pointing a wave at a new prefix is a parameter change, not an edit:
python gen_landing.py gallery.json deploy/index.html \
https://data.luxarviewer.dev/data/<new-prefix>
Passing a site-relative /data/<prefix> there is the mistake to avoid: it
still renders a working page, but every chunk fetch becomes same-origin and
needs something on the gallery origin to serve it. The audit asserts the URLs
are absolute for exactly this reason.
2.1 Always publish to a new dated prefix
Prefixes are dated (data/2026-08-26a/). Never overwrite a live prefix in
place. luxar optimize assigns a fresh content_hash by design, and a
warm viewer cache validating on an unchanged hash would serve stale chunks. A
new prefix sidesteps the whole class of problem: new URL, no stale cache, and
the old prefix stays intact as a rollback until the audit passes.
Unchanged stores are server-side copied within R2 rather than re-uploaded — same bytes, no egress, and it keeps the wave cheap.
2.2 Hash-compare correctly, or not at all
To decide which stores actually changed, compare this local build against the
previous local build. Do not compare a local build against the published
store: optimize gives every output a new content_hash, so that pairing
reports “changed” for everything and is meaningless. This has already cost one
near-miss 1.7 GB needless republish.
Snapshot hashes before rebuilding, then diff:
python snapshot_hashes.py > before.json # walk datasets/demos, record content_hash
# ... rebuild ...
python compare_snapshot.py before.json # report only genuine changes
2.3 “Unchanged” and “never built” are different answers
A store that failed to rebuild (missing optional dependency, no GPU, no
credentials) still has its old hash on disk and will report IDENTICAL. That is
indistinguishable from a successful no-op rebuild unless you also check
freshness. This caught cell_tracking_challenge reporting IDENTICAL at 64 hours
old.
Check freshness on FILES, not on the store directory. A directory’s mtime
updates only when its direct children change, so stat on
<store>.luxar.zarr can report a fresh timestamp over a tree nothing rewrote.
Not hypothetical: desi_galaxies passed a directory-mtime check while all
9,064 of its files were four days old.
Walk the tree and count how many files predate the run:
mtimes = [p.stat().st_mtime for p in store.rglob("*") if p.is_file()]
stale = sum(1 for m in mtimes if m < run_started)
# a real rebuild leaves ~0 stale; a skipped one leaves ~all
2.4 A demo may reuse a cached scene and still exit 0
Some demos short-circuit when their output already exists. desi_galaxies
prints Using cached scene: …, emits its own warning that the cached scene is
stale, then finishes with Dataset generated at … and exit status 0. A wave
driver reading exit codes learns nothing.
The staleness it warned about was real and had shipped: both LOD ladders held a finest-level node of 9,751,955 points against its own 4,000,000-point demo ceiling, which “can silently lose their tail on a 4096-class GPU”. The published tile carried that for six days.
That 4,000,000 is not a universal cap — it is a safety margin local to that
demo (SCENE_MAX_POINTS_PER_NODE, and its comment says so). The conservative
4096-class per-node caps are per geometry type, in typing_utils/constants.py:
geometry |
cap |
|---|---|
Lines |
2,793,472 segments |
Points |
5,591,040 points |
GSplats |
4,194,304 splats |
DESI’s 9,751,955 breaches the real 5,591,040 Points cap too, so that warning was not a false positive. Comparing a points node against 4,000,000 over-flags only the band from 4,000,000 to 5,591,040.
Two remedies, both printed by the demo itself:
luxar demo run <key> -- --recompute # recompute from source
rm -rf datasets/demos/<key>.luxar.zarr # or drop it and let the demo unpack
# the current record-fetched asset
Grep build logs for Using cached scene after any sweep. One gallery tile in
the 86-entry manifest took that path — few enough to miss, and the one that had
a real defect behind it.
2.5 Rebuild against the commit you think you are on
A sweep is only valid for the code it ran against, and dev moves under a long
one. A full pass over the 86 gallery tiles takes roughly ninety minutes here,
during which several PRs can land. Record the HEAD the sweep started from, and
re-check it at the end; if a store-affecting commit landed mid-sweep, the
results are stale for every demo it touches.
Discarding a partial sweep and restarting at current dev is usually cheaper
than finishing one you know is stale and then reasoning about which subset to
redo — that subset calculation is where scoping errors happen.
3. Hazards that fail silently
Every item here produced a plausible wrong answer in production rather than an error. They are grouped by what lies to you.
3.1 Stale media (historical — and how the fix works now)
Scene URLs are dated; media URLs are not. When media was served from R2
through a Pages Function, re-uploading a still to the same key left every edge
serving the old bytes, and the only remedy was a CACHE_EPOCH constant in
pages/functions/_r2.js that had to be bumped by hand on every wave that
re-captured media. Forgetting it produced a correct page with last week’s
thumbnails — and it did, at least once.
Moving media to Pages static assets removed that failure mode rather than mitigating it: a deployment replaces the asset and Pages invalidates its own CDN, so there is no epoch to remember. Re-capturing media now just means re-running the deploy.
Do not reintroduce a hand-maintained cache-buster. If media ever moves back behind a long-TTL rule on another host, it needs hashed filenames instead — see §4.2.
3.2 A cache policy change does not reach cached objects
Changing R2 CORS, or any response-header policy, affects only objects fetched after the change. Anything already in the edge cache keeps serving the old response until its TTL expires. With a 30-day TTL that is a month of breakage that looks fine on every fresh test you run.
Test the cached path, not a cache-busted one:
curl -s -o /dev/null "$URL" # warm it
curl -sD - -o /dev/null -H "Origin: https://luxarviewer.dev" "$URL" \
| grep -iE 'cf-cache-status|access-control-allow-origin'
# want: cf-cache-status: HIT *and* the CORS header present
If cached objects lack the header, purge (Cloudflare → Caching → Configuration → Purge Everything). There is no narrower fix.
3.3 curl -I misreports cache status
HEAD requests report cf-cache-status: DYNAMIC even where a GET cleanly
shows MISS → HIT → HIT. Diagnosing cache behaviour with -I will tell you
your cache rule failed when it is working. Use -o /dev/null -D - with a real
GET.
3.4 Zarr v2 vs v3 — never name a metadata document
The corpus holds both on-disk formats, permanently. Any tool that hardcodes
.zattrs, .zarray, .zgroup, .zmetadata, or zarr.json will silently
mis-handle half the stores. In zarr v3 there is one zarr.json per node with
attributes nested under attributes, and the consolidated index lives at
zarr.json → consolidated_metadata → metadata.
Two rules, learned the hard way:
Read through a bi-format helper, never a literal document name. A link audit that hardcoded
zarr.jsonscored a v2 store’s 9-byte “Not found” body as “no links” and passed.Absence must raise, not return empty. Two separate probes both reported “no
part_Nfound → correct” while walking zero nodes. A checker that cannot distinguish “nothing there” from “could not look” is worse than no checker.
3.5 Scale estimates: measure the affected, not the eligible
Counting things that could be affected instead of measuring what is has produced order-of-magnitude errors twice — a tile estimate wrong by >10×, and a model predicting 37,152 wasted requests on a scene that emits 5. Static models over the corpus are upper bounds. Ground-truth the head of the distribution in a real browser before reporting a number.
3.6 R2 billing: 404s count
R2 bills per operation. Cloudflare’s pricing page exempts exactly one error class — HTTP 401 — and says nothing about 404, which is billed as a Class B read. In practice the edge caches 404s too, so with a cache rule in place R2 sees one per URL per PoP per TTL and the cost collapses. The real cost of a 404 storm is first-paint latency, not the bill.
3.7 ffmpeg: every encoder option must precede the output filename
With ffmpeg 6.1.1, put -passlogfile after the pass-2 output filename:
ffmpeg -y -i in.webm -c:v libvpx-vp9 -b:v 480k -pass 2 -an -row-mt 1 out.webm -passlogfile P
warns that the option is trailing:
Trailing option(s) found in the command: may be ignored.
Options placed after an output filename apply to the next output, so ffmpeg
ignores the custom prefix and looks for the default ffmpeg2pass-0.log instead
of the P-0.log that a correctly ordered pass 1 wrote. Pass 2 exits 251:
Error opening file ffmpeg2pass-0.log.
[vost#0:0/libvpx-vp9] Error reading log file 'ffmpeg2pass-0.log' for pass-2 encoding
Error opening output file out.webm.
The reverse mismatch is equally broken: a trailing pass-1 option writes its
statistics to ffmpeg2pass-0.log, then a correctly ordered pass 2 looks for
P-0.log. If both passes trail the option, they both use the default filename
and can appear to work despite the broken ordering.
What makes this worth a numbered hazard rather than a footnote is the false
explanation waiting next to it. WebM/Matroska output can carry no stream
timestamps — including the gallery masters that ffmpeg assembles from explicit
per-angle screenshots — so ffprobe reports duration_ts=N/A and
nb_frames=N/A. “The input is undecodable” is therefore plausible and wrong.
Confirm decodability before blaming the input:
ffmpeg -v error -stats -i in.webm -f null - # reports frame=120 -- it decodes fine
Correct ordering:
ffmpeg -y -i in.webm -c:v libvpx-vp9 -b:v 480k -pass 1 -passlogfile P -an -f null /dev/null
ffmpeg -y -i in.webm -c:v libvpx-vp9 -b:v 480k -pass 2 -passlogfile P -an -row-mt 1 out.webm
The gallery harness normally encodes VP9 by quality (-crf 24 -b:v 0); a
target-bitrate two-pass re-encode is a last resort after the recapture controls
in §4.2.
3.8 Judge a re-encode on frames, never on byte count
Hitting a size target says nothing about whether the tile still depicts its
subject. hilbert_curve_3d re-encoded from 10 MB to 295 KB hit a 300 KB target
exactly and lost the fine wire detail that is the subject — the cube’s outline
survived, so every automated check passed. Extract matched frames from source
and output and look at them.
Measured knee for that scene: 295 KB visibly degraded, 587 KB resolved throughout, 1174 KB indistinguishable from source. Dense point clouds and high-motion synthetic scenes need roughly double what a smooth microscopy volume does.
For the normal size-control workflow, use the sanctioned WEBM_CRF or
WEBP_QUALITY knobs in §4.2, recapture, and compare representative frames.
3.9 Quality figures from before 2026-08-23 use the wrong basis
#1914 (2026-08-23) fixed five call sites that scored a fit against the raw
volume, when a fit reconstructs V - image_min. A figure stamped before that
date is invalid in a direction that depends on the volume’s pedestal and cannot
be corrected by arithmetic. Archives rebuilt afterward can carry correct
figures.
Consequences for this site:
Check the archive root
timestampand the sidecarquality_notedate before using a figure in a page, note, or comparison. Re-measure pre-fix or unstamped figures with current code against one materialised reference volume.A figure is only comparable to another measured on the same side of 2026-08-23. Comparing a pinned figure to a fresh one measures the scoring change, not the data.
The public page is unaffected: no rendered field carries a quality figure (
gen_landing.pynever readsnote). The exposure is in the manifestnotefields, which are internal.
3.10 One path, two generations: the closed Git-LFS migration hazard
Before #2354, a Zenodo entry could pin an in-repo Git-LFS generation with
sha256 / bytes and a different record generation with
hosted_sha256 / hosted_bytes. This explained an apparent contradiction for
cmu1_ch0.gsplats.zarr.zip: the former archive was a flat laddered leaf, while
the record archive was a four-part kind=partition tree.
#2354 removed the Zenodo payloads from Git LFS and collapsed each entry onto the record contract. The active cmu1 pin is now:
cmu1_ch0 sha256 7d605c3c... bytes 45,697,890 <- current record pin
There is no longer a repo generation selected by ordinary resolution. A stale cache from before the teardown fails the current pin and is downloaded again. The optional hosted fields remain supported only for legacy manifests.
The historical consequence worth retaining
The migration exposed a general rule: a tile’s structure is a property of the artifact generation, not just the recipe. The same path can name unrelated archive topology after a re-upload, and byte size alone does not identify which generation was measured. Hash the artifact before attaching structural claims to it, and inspect the scene-build call as well as the archive: re-adding raw arrays discards archive topology, while grafting a partition or nested tree preserves it.
The former dual contract covered 23 files across fourteen datasets, and every
pair differed. That inventory and its repo/record size ratios described the
staged migration, not the current manifest, so it was retired with the in-repo
payloads. The structural examples remain useful: cells3d changed size but
stayed flat, cryoem_virus gained a five-rung ladder, and cmu1 changed from a
flat leaf to a four-part partition. Only a digest-confirmed, kind-based read
settles topology.
Derive topology from each group’s declared kind (children of kind=lod are
substitutive levels, children of kind=partition are parts) rather than from
node-name patterns. And note that no kind attr anywhere means flat, not
unreadable.
3.11 Flattening only cuts requests when archive topology reaches the scene
Five authoring paths have different outcomes from the same archive change:
authoring call |
effect of flattening the archive |
|---|---|
load into |
archive topology is discarded; bytes and splat content can still change |
load into |
additive rungs and substitutive levels are lowered into the scene, so scene node count changes with the artifact |
|
a matrix-shaped leaf/LOD is re-authored through |
|
grafted scene node count changes with the artifact |
|
structure is authored locally by the recipe |
Applied to the current demo code:
h2afva_timelapsenow supplies a matrix-shaped ladder, soadd_gsplats_from_filere-authors it throughGSplatData;h2afva_stackstill supplies a nested partition and is grafted. Only the stack inherits its archive tree node-for-node.codex_pancreasfits locally throughsave_with_lod(recipe="adaptive")and grafts its own output, so its groups are authored rather than inherited. Its structure changes through recipe settings such asmax_elements, not by flattening a separately supplied artifact.cmu1_pathologygrafts the record’s partition, so its archive structure reaches the scene node-for-node (3.10). Flattening that archive is a real load win, but measure the resulting topology before claiming it.cryoem_virus,milkyway_dust,dapi,multichannel, andopencell_map4passGSplatDatathrough, so their digest-confirmed hosted ladders and levels become scene nodes rather than being flattened by the demo.
So split the claim per demo before promising a load win. “Fewer nodes” and “fewer bytes” are different wins; only paths that preserve archive topology get the first, and some demos are already correct.
3.12 Before flattening, compute the RESIDENT count — the total will mislead you
Section 7 gives the counting rule (only the resident slice counts, and its two traps). This is the flatten-specific consequence, because flattening is where the wrong number is most tempting: it collapses a tree into one node, and that node’s total is what the compiler prints.
Worked case. The pinned h2afva_51tp generation is the result of the 3.17
rebuild (flatten → lod --recipe stream → optimize): one 4D leaf with a
twelve-step progressive ladder and 121,163,285 splats:
node total, 51 timepoints 121,163,285 <- what ElementCapacityWarning prints
resident slice, worst case 2,629,840 <- what the GPU commits
cap (4096-class GPU) 4,194,304 -> 1.59x UNDER, fine
Read as a total that is 28.89x over cap and looks like a blocker. It is not: only
one of the 51 timepoints is ever resident. _lod_policy.py says so directly —
“the compiler also warns on the node total rather than the resident slice, so
that warning is expected for a sliced nD node that satisfies the runtime limit.”
An ElementCapacityWarning on a sliced nD node is not a finding.
The exception is a STATIC object, which has no hidden axis to reduce the
committed set. The current cmu1 record generation (sha256; ch0
7d605c3c...) flattens its three channels to 8,823,953 / 9,924,486 / 10,830,790
splats. All are 2D — nothing to slice on — and every channel exceeds the cap, so
the overflow shows as a Hilbert-contiguous clean-edged hole that reads as
missing data. There, parts stop being optional and become load-bearing.
Do not try to confirm that by opening the tile: it renders whole on a
developer Mac, because a Metal maxTextureSize=16384 path caps at 16,777,216
rather than 4,194,304. A clean render on your own hardware carries no information
about the floor — see 3.16.
So the check before flattening is arithmetic, not a run: a kind=lod group
contributes only its finest child, a kind=partition sums its parts,
additive_N rungs are deltas that re-partition the level’s own content, so never
add them on top of it — and equally, a ladder never reduces the leaf below its
own total, so it is not a remedy for an over-cap node (3.16) — then divide by the
hidden-axis extent if there is one. Only if the resident figure exceeds the
cap does the demo need partition=dict(max_elements=…) landing in the same
change as the flatten.
3.13 A ladder’s counts are in different UNITS per geometry, and the wrong one writes zero rungs
On add_lines, an explicit counts list is in polylines while
"stream:<c>" is in vertices. Measured on 4,000 polylines x 27 vertices:
counts=[39062, 78124, 108000] -> clamped to the polyline count -> 0 RUNGS
"stream:39062" -> 3 rungs (39,069 / 39,069 / 29,862 vertices)
The clamp is silent: no error, no warning, and a store with zero rungs still
loads and still renders. check_demo_ladders.py is the only thing that catches
it, and only above 200,000 elements — so a mid-sized demo can ship a ladder that
does not exist.
3.14 One viewport does not validate a substitutive ladder
A whole-object substitutive ladder anchors its finest level at 0.5 screen occupancy, and adding levels cannot move that anchor. So the viewport aspect ratio — not the ladder — can decide whether a scene is bounded.
Measured on cosmicflows_laniakea_full at the authored pose: two substitutive
levels settled at ~892k segments on 16:9, while the wide outer basins still
selected ~2.1M and ~3.2M fine segments at 4:3 and 1:1.
A ladder that looks bounded on a 16:9 capture can therefore be unbounded on a
square window. Validate at several aspect ratios, or prefer an additive ladder
where the geometry allows one (check Section 7.4 first for indexed lines),
whose prefix is bounded by construction rather than by framing.
3.15 A streaming ladder’s first rung is first paint — size it in bytes, not chunks
desi’s 2,000-element first rung (SCENE_FIRST_CHUNK,
demo_desi_galaxies.py:207) is sized so its eager coarsest substitutive
level lands in one zarr chunk. An additive-only leaf has no coarse level, so
its first rung is first paint, and 2,000 elements is far below a sensible
download budget.
Measured group counts before and after converting six embedding demos from substitutive to additive:
demo |
today |
at |
at |
|---|---|---|---|
|
13 |
12 |
7 |
|
13 |
12 |
7 |
|
17 |
15 |
11 |
|
18 |
14 |
9 |
|
17 |
17 |
13 |
|
18 |
18 |
13 |
39,062 is a 200 ms budget at 25 Mbps and 16 B/element. At 2,000 every extra rung is another node, so the ladder costs requests without buying a faster first paint — the same accounting trap as counting nodes instead of bytes, one level down.
These stream:39062 measurements predate the sliced-node share floor. A node
with hidden dimensions now starts at max(39062, ceil(n / 8)), so the six rows
above rebuild with 5/5/7/6/9/9 groups respectively.
3.16 State the window before you read the number, and check it is longer than the phenomenon
The most expensive error in this campaign was not a wrong mechanism. It was a phenomenon that did not exist, and two mechanisms invented in later work to explain it.
The false observation
gsplats_2d_cmu1_pathology’s live tile appeared to commit 10,295,708 of its
store’s 20,591,415 elements — exactly 50.0%, per channel — with isLoading
stuck true, no console error, and half the data apparently unreachable. It was
called unpublishable, held from a wave, filed as an open problem, and carried
into later work as evidence for an unbounded-loop hazard.
What is actually true
Progressive: 4/4 LODs loaded (7093383 splats) — complete @ 63,442ms
Progressive: 4/4 LODs loaded (6601413 splats) — complete @ 64,183ms
Progressive: 4/4 LODs loaded (6896619 splats) — complete @ 67,864ms
totalElements 20,591,415 isLoading false
progress lines 12 total, 0.18/s
Every channel reaches 4/4. The three sum to 20,591,415 — the whole store. It completes in about 68 seconds. Nothing is truncated, nothing stalls, and at 0.18 lines/s nothing spins.
Why the number looked real
The gap between the last level-1 line (41,750 ms) and the first completion (63,442 ms) is 21.7 seconds of silence. The probe declared the total “settled” after three identical samples at 5 s intervals — 15 seconds.
15 < 21.7. The measurement stopped inside a quiet period mid-load and reported
a snapshot as a terminal state. isLoading: true was not a symptom; it was the
literal truth.
Why it survived so long
The artifact was stable and reproducible. Two independent probe runs agreed to the unit, because both shared the 15 s settle rule. Reproducibility measured the rule, not the system.
It was quantitatively beautiful. 50.0% in every channel, across three different leaf totals, matching a cap prediction exactly. That is what bought it credibility — and the exactness came from all three channels having 4 rungs, so any stop after level 1 yields 50% everywhere. Three “independent” confirmations were one constraint counted three times.
The search terms encoded the hypothesis. The first grep set was
clamp|truncat|exceed|capacity|MAX_SPLATS. The answer was inProgressive: 4/4 … — complete, which that set structurally could not match. A filter built from a theory can only ever confirm it.
The rules
Fix the observation window before looking at the result, and justify it against the expected duration of the thing being measured. Twelve rungs at the 10–26 s per level actually observed is minutes; a 15 s settle rule and even a 180 s cap are both inside that range.
“Stopped changing” is not “finished.” Prefer a signal that matches the phase being measured:
— completefor a full ladder, orisLoading: falseonly for the first committed view, rather than inferring either from a stationary number.A stable artifact is not a real effect. If two runs agree, check they do not share a stopping rule, a cache, or a filter.
Grep for what the code says, not for what you suspect. Loaders, policies and compilers usually log their own reason; find that string first.
What genuinely survived
The load is slow for an ordinary reason: the store is 20,591,415 splats, roughly 172 MB, delivered in ~68 s — about 20 Mbps. That is a large download, not a defect. A request-count explanation was considered and measured false: the published store has 508 chunks with a nominal ~1024 KB uncompressed chunk shape (the
archiveprofile’s 1 MB target), against 50,316 chunks at 1.0–2.7 KB in the upstream.gsplats.zarrarchives.optimize --profile archivedoes re-chunk grafted subtrees, so the archives’ fragmentation never reaches a published tile. It does still hit whoever downloads those archives directly — a demo build pays 38 MB in 16,852 pieces — which is an authoring-side fix worth making upstream.The cap risk is real, and a clean render does not test it. Each channel is 6.6M–7.1M splats with no hidden axis. It renders whole on a 16384-class developer GPU — the probe launches
--use-angle=metal, wheremaxTextureSizeis 16384, giving a gsplats cap of4096 x 16384 / 4 = 16,777,216. 6.9M is comfortably under that.constants.pyis explicit: “maxTextureSize is a GPU property (16384 on modern desktop, 4096 on the conservative floor), so the only bound an AUTHOR can rely on is the 4096-class one.” So a clean render on developer hardware is the expected observation and carries no information about the floor — on a 4096-class GPU the same node clamps and loses a Hilbert-contiguous wedge (#1957 erased the North Atlantic by clamping 2.3% of a Lines node). Never validate a cap question on one GPU class.The authoring guard sees the accumulated quantity.
warn_if_over_element_capruns once at the parent with the ladder total; rung checks are deliberately suppressed as redundant, and the graft path performs the same aggregate check. The 6.6M–7.1M CMU-1 channels therefore warn against the 4,194,304-splat conservative floor.An additive ladder still never bounds the committed set — a prefix converges to 100% of the leaf. Only a partition, or a hidden axis to slice on, reduces what is resident.
3.17 Dropping substitutive LODs without re-chunking leaves the store far slower than it needs to be
Removing substitutive levels is the right call for most single-object scenes, but it is half an operation. The recipe is three steps, in this order:
flatten -> lod --recipe stream -> optimize --profile archive
The ordering cost is independently measured on the Drosophila 500-timepoint
archive: requests per timepoint step fell from 173 to 2 after
optimize --profile archive
(demo_gsplats_4d_drosophila_embryogenesis.py:163-165).
The reason is not inherited source chunking: flatten and lod rewrite every
array at the 64 KB authoring target, discarding even an existing 1 MB layout.
Running optimize before either command is therefore undone, and a timepoint
slice again spans many small chunks. Additive-only and re-chunking are a
package, and optimize must run last so every rewritten array gets the 1 MB
layout.
The h2afva_51tp rebuild exposed two further traps:
Don’t stop at
flatten. A bare flat leaf loses the ladder entirely; the target is a leaf plus rungs (lod --recipe stream), which is what_lod_policy.py’sstream 4 nodesrow describes.A partition can be worth nothing on a time-stacked node. The writer already lexsorts by the time barrier, so per-timepoint chunk locality exists without any partition — adding one buys no request reduction there. (It still earns its place on a static node over the element cap, 3.12.)
Result now pinned for h2afva_51tp: 1,873,559,527 → 1,115,714,088 bytes
(−40.5%), 176 substitutive levels → 0, and the former partitioned tree → one
4D leaf with a twelve-step progressive ladder. The chunk count fell from
125,751 to 2,316, with all 51 timepoints intact at uniform spacing and none
blended.
Superseded at the SCENE level on 2026-09-10. The pinned archive stays one
laddered leaf, but demo_gsplats_4d_h2afva_timelapse.py now re-authors it at
build time into a kind=partition of one part per TIMEPOINT (51 parts, each with
eight equal-count rungs: 277,131 splats in rung 0 for the 2,217,045-splat
reference frame, with every increment under the 900 K cap), cached beside the
download. Measured cold with the former capped-stream recipe against the single
leaf and against 44 spatial parts at identical chunking: the single leaf’s
global ladder re-streamed from its bottom on every slice (3-27% of the frame
resident while a step loads); spatial parts paid ~30 MB per step; time parts
fetch exactly one part per step and never starve a slice. The “a partition buys
nothing on a time-stacked node” rule above is about SPATIAL parts; time parts
are a different structure.
3.18 A probe must emit the evidence that its own window was valid
3.16 says to check the observation window is longer than the phenomenon. That is advice, and advice does not run. Make it a reported field instead, so a meaningless run announces itself.
Worked example from an ad-hoc browser-console probe; there is no checked-in script to rerun. The depth-sort flashing bug lives only on count-changed commits; equal-count re-commits take a path that always worked. Three runs against the same live bundle reported:
host A firstCommitWaitMs 501 commits 99 countChanged 99 distinctSteps 89 unsorted 0 -> FIXED
host B firstCommitWaitMs (none — fixed 8 s warm-up) commits 73 countChanged 0 distinctSteps 1 unsorted 0 -> INCONCLUSIVE
host B firstCommitWaitMs 1752 commits 85 countChanged 0 distinctSteps 1 unsorted 0 -> INCONCLUSIVE
Host B’s runs looked like passes on the headline numbers and were worth nothing: zero denominators, with every observed commit an equal-count progressive re-commit on a path that never had the bug.
The 1752 ms diagnostic did not refute a cold-cache explanation. In the viewer,
isLoading === false reports first-commit latency, not full-ladder completion, and
the debug interface is installed only after the initial loadDataset call returns.
Progressive refinement may continue afterwards; the 85 re-commits show that it did.
The number therefore bounds time to the debug interface and first committed view,
not time to a settled ladder, so a load-bound observation window remained plausible.
Open the viewer with ?debug, keep the 180 s first-commit timeout separate from the
30 s observation window, and publish both the wait and its exit reason:
const firstCommitTimeoutMs = 180000;
const observationWindowMs = 30000;
const tWait = Date.now();
let firstCommitObserved = false;
while (Date.now() - tWait < firstCommitTimeoutMs) {
const st = window.__luxarDebug?.getState?.();
if (st && st.isLoading === false) {
firstCommitObserved = true;
break;
}
await new Promise(r => setTimeout(r, 500));
}
window.__firstCommitWaitMs = Date.now() - tWait;
window.__firstCommitObserved = firstCommitObserved;
window.__observationWindowMs = observationWindowMs;
if (!firstCommitObserved) {
throw new Error('first-commit timeout; probe inconclusive');
}
// Run the measurement for exactly observationWindowMs from here.
Rules:
Gate on a signal that matches the phase being measured, not a fixed sleep.
isLoading: falseis suitable for first commit; full-ladder work needs the loader’sProgressive: n/n LODs loaded — completesignal. A hardcoded warm-up is only an untested claim about the system’s timescale.Report the wait and the exit reason.
firstCommitObserved: falsemeans the 180 s cap expired and the run is inconclusive; the elapsed value alone cannot distinguish timeout from success. This timeout precedes the separate 30 s observation window, so do not compare one duration to the other.Print the denominator next to every ratio.
0 unsortedand0 of 0render identically in a summary line and mean opposite things — one is a pass, the other is no measurement. A zero denominator is never a pass; emit an explicitINCONCLUSIVEverdict instead of a green one.Assert that the driver ran, not just that the output looks clean. The effect here lives on time-axis steps, so the probe must report how many distinct axis positions it observed. One host saw 89 distinct timepoints with counts ranging 10,229–30,296; another saw 85 commits at a single position, and only the step-count field distinguishes “the fix works” from “nothing was exercised”. Where a probe depends on the system animating itself, have it detect that and drive the axis directly when it is not.
Note both failures here were the same mistake by different hands: a fixed 8 s warm-up written by the session that had already documented that a stopping rule is a claim about timescale, and a 15 s settle rule (3.16) written by the session that had just relayed that lesson. Knowing the rule is not the control; emitting the diagnostic is.
3.19 Node reduction cuts requests only when arrays fit in one chunk
Section 2’s rule — hosted cost is requests, not bytes — is right, but “one
request per node” is an upper bound, not a measurement. After
optimize --profile archive (1 MB target) a big array spans many chunks while a
small one spans exactly one, so cutting node count only cuts fetches in one of two
regimes. The following historical counts come from consolidated metadata at the
retired published prefix 2026-08-27b. The live prefix is per store (8.1; resolve
it from the /d/<demo-key> route): data/2026-09-12 backs most rows here, while
biodiversity_planetary_scale is at data/2026-09-19. arrays excludes
zero-shaped array_ref placeholders, chunks is the root
chunk_layout.chunks_after value, and the eager columns apply the viewer’s
default_level deferral rule in
packages/luxar-viewer/src/data/scene-loader/nodes/load-lod-group-node.ts. They
count every array under that default-level subtree, including all additive rungs;
the first-rung columns keep only additive_0 inside each ladder:
store |
groups |
arrays |
chunks |
eager arrays |
eager chunks |
first-rung arrays |
first-rung chunks |
whole chunks:arrays |
|---|---|---|---|---|---|---|---|---|
|
2767 |
10320 |
10320 |
3440 |
3440 |
860 |
860 |
1.00 |
|
87 |
320 |
389 |
80 |
80 |
10 |
10 |
1.22 |
|
29 |
108 |
187 |
108 |
187 |
48 |
75 |
1.73 |
|
15 |
79 |
224 |
79 |
224 |
79 |
224 |
2.84 |
|
18 |
60 |
508 |
60 |
508 |
15 |
127 |
8.47 |
Use the eager columns for the converged default-level load and the whole-store
ratio for the cost to reach full detail. Use the first-rung columns for a
first-paint claim (3.15). laniakea is the only store in this table with no
additive ladder, so it alone has identical eager and first-rung totals.
Two regimes, and the ratio tells you which one you are in:
Node-bound (ratio ≈ 1). Every eager array is a single chunk, so removing an eagerly loaded array removes approximately one request from the converged default-level load. In the retired
2026-08-27bgeneration,codex_pancreaswas exactly 1.00 both store-wide and for its eager subset; that load converged at 3,440 arrays/chunks, not all 10,320, while its first committed rung needed 860. The 63-request stacked-leaf versus 689-request partition measurement inpackages/luxar/src/luxar/demos/_lod_policy.pyis not a reusable sublinear node-to-request law: it compares a byte-bound leaf with a node-bound partition.Byte-bound (ratio >> 1). Arrays span many chunks, so request count tracks total bytes and is nearly indifferent to node count.
cmu1’s converged load fetches 508 chunks from 60 arrays, while its first rung fetches 127 from 15; both have ratio 8.47, so halving its node count would barely move either.
The ratio is a property of a pipeline STAGE, not of a store. optimize is
what moves these stores toward the node-bound regime. The same historical
published roots recorded both pipeline stages in chunk_layout, so this
comparison was reproducible from their zarr.json files while the retired
2026-08-27b prefix was live. The live prefix is per store (8.1; resolve it
from the /d/<demo-key> route): data/2026-09-12 backs most rows here, while
biodiversity_planetary_scale is at data/2026-09-19. The table is retained
as a historical measurement, not re-measured:
store |
arrays |
chunks before |
before ratio |
chunks after |
after ratio |
|---|---|---|---|---|---|
|
10320 |
19888 |
1.93 |
10320 |
1.00 |
|
320 |
8600 |
26.88 |
389 |
1.22 |
|
108 |
2081 |
19.27 |
187 |
1.73 |
|
79 |
5004 |
63.34 |
224 |
2.84 |
|
60 |
7580 |
126.33 |
508 |
8.47 |
biodiversity_planetary_scale therefore reads 19.27 before optimize against
1.73 after it — same generation and structure, but a factor of 11 fewer chunks
per physical array because optimize re-chunks to a 1 MB target. This is consistent
with 3.17’s measurement in the other direction (173 requests as-built, 2 after
optimize).
So a node reduction’s payoff is contingent on the publish step: optimization can move a small-array store into the node-bound regime, but it does not guarantee that outcome. Halving eager arrays is a direct total-load request win on a node-bound published artefact and close to meaningless on a byte-bound as-built one. The win belongs to the combination, not to the authoring change alone.
Practical consequence: before claiming a node reduction buys a faster load, check the chunks:arrays ratio of the artefact you will actually serve — post-optimize, and of the generation you are publishing, not whichever one happens to be live. A store with many small nodes gains directly; a store with few large ones gains almost nothing and its lever is total bytes instead (3.17).
Corollary for the other direction: adding nodes is only expensive in the node-bound regime. A per-part additive ladder that multiplies groups is cheap on a byte-bound store and costly on a node-bound one — so the same structural change has opposite cost depending on chunk layout.
Measured: re-tiling codex_pancreas at a 4x coarser partition
Measured 2026-08-28 in commit 99c1b0820 on the since-deleted
runbook-followup-codex-and-review-loop branch; hosted topology re-checked
2026-09-24. The rebuilt row’s eager and first-rung pairs were derived from that
live re-check. The then-live 2026-08-27b store was rebuilt at a 4x coarser
partition (172 parts → 43) and taken through optimize --profile archive:
generation |
parts |
geometry groups |
all groups |
arrays |
chunks |
eager arrays |
eager chunks |
first-rung arrays |
first-rung chunks |
chunks:arrays |
|---|---|---|---|---|---|---|---|---|---|---|
published |
172 |
2764 |
2767 |
10320 |
10320 |
3440 |
3440 |
860 |
860 |
1.00 |
rebuilt |
43 |
700 |
703 |
2580 |
2828 |
860 |
860 |
215 |
215 |
1.10 |
The three extra groups in the all-group count are the overlay wrapper and its
two text overlays; stating both counts keeps the structural comparison
like-for-like. The old prefix has since been retired. The current 2026-09-12
root metadata independently confirms that the hosted store is the 43-part
rebuild: 700 geometry groups (703 including overlays), arrays_total=2580, and
chunks_after=2828.
The full-detail chunk total fell by 3.65x, short of the 4.00x reduction in
arrays because the coarser partition made each array larger. That shortfall is
whole-store only: both the converged eager load and first committed rung fell by
the full 4.00x, to 860 and 215 arrays/chunks respectively, and the eager
subset remains exactly 1.00 chunks:arrays. The whole-store ratio moved from
1.00 → 1.10 toward byte-bound without leaving the node-bound regime. Only
re-running optimize gives the real request count.
Worked example: is a per-part ladder worth its nodes?
This was an ad hoc local full-data experiment on 2026-08-28, based on
demo_nuclear_pore_complex.py; neither the probe nor its output was checked in.
It is not the published 2026-08-27b preview, which has 41,288 points in one
unpartitioned node. The local experiment used 9,874,128 elements and a 32-part
BSP, with both variants taken through optimize --profile archive:
version groups arrays chunks first commit
un-laddered 36 160 224 9,874,128 elements
4 rungs/part 161 640 672 ~1,250,000 elements
Post-optimize chunks:arrays is 1.40 un-laddered and 1.05 laddered, so both variants are node-bound and the 128 rung groups cost real fetches: +448 requests to reach full detail, 3x the un-laddered total.
But that is the wrong total to compare. First paint needs only the first rung of each part — roughly a quarter of the arrays, ~128-160 chunks — against all 224 for the un-laddered store, which must also commit 9.87M elements in one go. So the ladder is cheaper at first paint on both axes and more expensive only in the total to reach full detail, which arrives progressively and off the critical path.
For a gallery tile that is the right trade. State it that way round: a ladder does not reduce total requests, it moves them after first paint.
3.20 A pin audit expires on the next upload, and local resolution can mask a bad hosted pin
Two separate hazards that showed up together.
Local resolution does not prove the hosted pin
Before #2354, section 3.10’s resolution order was verified local cache, then in-repo git-LFS payload, then Zenodo. An ordinary local run could therefore prove only the packaged generation, not the record pin.
Use the shipped cold-fetch harness before any future removal:
make check-cold-fetch hides the in-repo payload directory, starts from an
empty cache and verifies the record against its manifest pin. For teardown, run
hatch run python scripts/verify_cold_fetch.py --require-verified N with the
expected target count after the final upload, so a dormant or missed variant
cannot turn an incomplete audit green. A digest mismatch aborts with a checksum
error rather than accepting a silently wrong generation.
An audit expires on the next write, not on a timer
During the migration, gsplats_cmu1_pathology pinned record sizes 84,492,218 /
93,189,978 / 100,336,780 against a draft holding 45,697,890 / 50,699,910 /
53,698,264, and
git log --all -S'45697890' -- packages/luxar/src/luxar/demos/data_manifest.json
returned nothing at the time. The pins later entered history in c5cb10207
and 77d98994e, so the same query is no longer empty.
Read that result carefully — but do not over-read it. It does not mean nobody had pinned the draft: the correct values existed at that moment as an uncommitted edit in another worktree.
Both halves were true and it is worth keeping them apart:
Every committed manifest was stale, the one that ships included. Only committed pins ship, so the publish-time failure was entirely real for as long as the right values lived nowhere durable.
The uncommitted edit explains why the risk existed, not why it was absent. Work in a worktree protects nobody — the same “unpublished work is invisible” shape as a store built but never uploaded.
What the query genuinely cannot do is tell those apart: git log --all -S<value>
returning empty means “nothing has committed this value”, which reads
identically for “wrong everywhere” and “right but not yet durable”. Reach for a
second signal before naming the defect. Run make check-zenodo-live first: its
audit prints PROBABLY THE WRONG MANIFEST, NOT BROKEN RECORDS when at least 66%
of the pins disagree, before listing the full-path git log check.
The audit that had cleared the manifest read 35 match / 0 differ / 7 not in any draft; those 7 were uploaded afterwards. And the session doing the uploading went green → red inside one sitting: upload four archives, audit green, upload three more, audit red with “description says 278.0 MB, actually 150.1 MB”.
So the rule is not “re-run before publish” — it is:
Treat any audit older than the last write as void. A green is a statement about the instant it ran.
Write pins after the upload they describe, and commit them immediately; a pin in a worktree protects nobody.
Re-check the record’s prose too. Descriptions carry sizes and part counts and go stale with the same write.
The partial re-pin left one archive on its deliberate previous generation
PR #2333 resolved this for five of the eight restructured archives —
cmu1_ch0/1/2, cryoem_virus, and the hosted-only h2afva_51tp (whose
sha256/bytes moved in place, since with no in-repo payload a second contract
would be ambiguous). Each moved file keeps its outgoing digest in
superseded_sha256, so one previous cache generation reads as stale rather than
corrupt.
The paired migration in #2334 moved vh_head and ct_atlas together with their
positionally indexed sidecars, bringing seven of the eight onto the restructured
generation. Their fits now also carry the colours or labels natively, so a writer
reorder cannot separate the payload from the splats it describes.
Only milkyway_dust remains on its previous pin on purpose. Its coarse levels
are what get selected when the galaxy is orbited at range; a flat streaming ladder
would regress the zoomed-out view (3.14).
That publish-order constraint was discharged before #2354. The
milkyway_dust file on record 21912280 was rolled back to the manifest’s
10,647,985-byte pin, and records 21912280 and 22118695 were published
together on 2026-09-02. Cold fetches therefore receive the generation the
manifest verifies. The general rule remains: never publish a record whose bytes
disagree with the committed pin, because a checksum ValueError is an integrity
failure rather than an ordinary DatasetUnavailable fallback. The live gallery
tile is indifferent because it serves an already-derived scene and never
consults the pin (3.10).
The 2026-09-18 v2.0.0 wave published records 22804363, 22807312, and
22804386: cc-by gained the corrected neuromast pair, the h2afva 51-timepoint
archive moved to a four-step ladder, and the Drosophila timelapse moved to the
unculled 128-million-splat build. The manifest repin retires the temporary
neuromast R2 override and makes existing Drosophila caches fetch the new 1.1 GB
archive.
3.21 Animated, un-laddered nodes need chunk-boundary-aware prefetch
The 2026-09-02 wave re-chunked every store with optimize --profile archive.
On collision_animated (a 4.66 M-vertex Lines node over 250 frames, no ladder)
that made a 1 MB vertex chunk hold ~6 frames; every ~6 frames the playhead
crossed into the next chunk of each of the four arrays and waited for a
0.25-0.85 MB fetch (0.4-0.9 s on a 7 Mbps link, cold cache, 40 steps = 65
requests / 21 MB). The old hidden-axis lookahead was one step deep, so it could
not hide a boundary that far ahead. The viewer now uses each array’s zarr chunk
shape plus the index atom bounds to prefetch the next boundary in playback
direction (#2686). Two things were NOT the cause, and the first diagnosis
wrongly blamed one of them: the ordering is already slice-major (compound
ordering, vertex_ordering.slice_dims == [3], every atom inside one frame), and
the old prefetcher did run — only one step ahead.
For this node the mean added lead is about 2.9 frames with archive, 1.2
frames with hosting, and 1.0 frame with local; the profile choice still
governs how much time the prefetch has to hide each request.
Rules that fall out, alongside 3.17 and the #2377 residency finding:
Compute frames-per-chunk before choosing a profile for an animated node: group consecutive
chunk_boundsatoms exactly as each planned zarr chunk will group them, then measure the inclusive time span of each group.Laddered + played (splat timelapses): 1 MB, so the coarse rung is one resident chunk. Un-laddered + played (animated lines/points):
hosting(~1.5 frames per chunk) is the middle of the request/byte trade;localis 3,500 chunks against the Functions request cap. Not animated: smaller is a pure win.optimizekeeps each array’s byte-based chunk planner, then adds a warn-only per-node check for the playback/cache policy it can verify. The check is gated onviewer_config.animation(playing: true), not on “has a hidden axis” — a keypress-navigated axis is the refine case. It warns when an un-laddered spatial-index node would exceed two frames per chunk across multiple chunks; a whole array in ONE chunk is benign because there is no boundary to cross. Read the atom’s hidden-axis span fromchunk_bounds— a store with one NODE per hidden coordinate (hilbert_curve_3d) misreads when the scene range is divided by chunk count. Wave 2026-09-11:collision_animatedplays at 1.8 frames/chunk onhosting;cloudalso starts with playback enabled but has not yet been measured on its published layout. Ocean’s tentacles would be 28 frames/chunk onarchiveif that axis were ever played.For a TIME-PARTED store (h2afva) the profile rationale is request count, not rung residency: each part is one frame loaded whole, so there is no cross-timepoint rung to keep warm. Measured on the published
archivelayout (2026-09-11,chunk_layoutread first): 127,192 chunks at the writer default -> 3,793 onarchive(33.5x fewer); a coarsest-rung traverse of all 51 parts is 255 requests (5 chunks per part), the finest rung 1,020. The 30 B/splat one-array-per-rung estimate (~1,400) was 5.5x pessimistic — a rung carries five arrays andarchivepacks them far better than a per-array byte calculation predicts — so treat that estimate as a safe go/no-go upper bound, never as the expected cost. (A first “measurement” of 1,326 on this store was the writer default, notarchive.)Read
chunk_layoutBEFORE measuring chunk cost, and check it on EVERY store in a wave, not a sample: its absence is the only reliable sign that a scene was never optimized, and stores built on another machine by another operator are exactly the ones that slip through. A chunk-cost number without the layout it was measured on is not a number.Resolve demo ids versus store names through §8.1 before using any operator list. Gallery capture selectors accept demo ids and only warn on unmatched store names, so a mis-keyed pass can exit 0 with placeholder tiles and still satisfy a tile-count guard.
3.22 Guard the artefact you ship, not only the inputs you fed it
A publish deployed a gallery page carrying 9 tiles instead of 85, and every guard passed. They were all reasonable guards — and all of them checked inputs:
store-count parity on R2 85 == 85 PASS
chromatrace leakage 0 PASS
media staged / oversize 180 / 0 PASS
relative /data/ URLs 0 PASS
tiles on the rendered page (not checked)
Cause: build_gallery_data.py skips any demo whose store is absent from the stage
directory —
if not hosted.is_dir():
continue
— and the stage held only the nine stores this wave re-chunked. Earlier waves re-chunked everything, so the stage was incidentally complete and the gap never showed. The data was correct throughout: all 85 stores were at the new prefix and parity confirmed it. Only the page was wrong, and the page is what a visitor sees.
Two guards, and the pairing is the point:
Stage completeness, before the page is built. Count store directories (or symlinks —
is_dir()follows them) in the stage and require it to equal what the prefix serves. This one says why.Tile count, after the render. Count
href="https://luxarviewer.dev/?src="occurrences and require one per store; also accept legacyhref="/viewer/index.html?src="links while/viewer/remains served. Separately require every tile to reference a prefix that is still live — either the wave’s new prefix or an earlier one still serving — since a stale-prefix tile passes a bare count. This one says that, and is the last line before deploy.
Verify a new guard against the broken artefact, not only the fixed one. Both were run against the page that actually deployed: 9 vs 85, abort. A guard only tested on a good input is an assumption.
The general rule: a green pipeline over correct inputs does not imply a correct output. Ask what a visitor receives and assert that directly. Every other check here is a proxy for it.
Recovery, for the record: layered symlinks into the stage (full base wave, then each later wave’s overrides, newest winning so per-store byte counts stay accurate), rebuild, re-render, redeploy — one changed file. Roughly four minutes degraded.
4. Cloudflare configuration
4.1 Cache rule on the data subdomain
Rules → Cache Rules, filter hostname equals data.luxarviewer.dev, action
“Eligible for cache” + “Ignore cache-control header and use this TTL”.
TTL is a tradeoff, not a tuning knob: prefixes are dated and therefore immutable, so a long TTL is safe for data. But the same rule governs how long a mistake persists — a CORS or header change is invisible to already-cached objects for the full TTL (§3.2). One day is a reasonable compromise; anything longer wants a purge in the change procedure.
4.2 Media cache safety
Demo-site /media/* URLs are not dated. They are Pages static assets, which
is safe because a deploy invalidates them (§3.1). Root-README media also live
under /media/*, but on data.luxarviewer.dev; those objects use SHA-256-derived
filenames, so the long-TTL data-host cache rule is safe by construction. Any
other media moved behind that rule needs a short TTL or hashed filenames
first, otherwise a re-captured still is unfixable without a full purge.
Two Pages limits bound the Pages set: 25 MiB per file and 20,000 files.
The count is comfortable (180), but the largest clip sits at 24.87 MiB —
0.13 MiB under the cap. A single oversized file fails the whole deployment, so
the gallery harness checks every PNG, WebP and WebM immediately after it is
written. It warns at 20 MiB, fails at the 25 MiB boundary, and prints the total
plus the five largest files at the end of the run. Treat a warning as a prompt
to choose a deliberate encoding adjustment with WEBM_CRF or WEBP_QUALITY
in generate-gallery.spec.ts, then recapture and inspect the affected demos;
do not silently trade quality for size with an automatic re-encode loop. Judge
the result on matched frames rather than bytes alone (§3.8).
4.3 CORS on the R2 bucket
Required for both viewers now, since the gallery also fetches data
cross-origin. AllowedOrigins must list every hostname that hosts a viewer:
https://luxarviewer.dev apex viewer
https://demos.luxarviewer.dev gallery viewer
https://luxar-demos.pages.dev Pages fallback URL
Per-deployment preview URLs (<hash>.luxar-demos.pages.dev) cannot be
enumerated, so data will not load in a Pages preview deploy. Verify against
the canonical hostname.
Directory .luxar.zarr stores, including the entire published corpus, fetch
metadata and chunks with simple GETs. The identity watchdog also uses
If-None-Match so an unchanged root metadata document can return a bodyless
304. They require Access-Control-Allow-Origin, if-none-match in
AllowedHeaders, and ETag in Access-Control-Expose-Headers. Without the
allowed request header, browsers reject the conditional probe at CORS
preflight; the viewer falls back to unconditional probes, which remain correct
but re-download the root metadata document.
Zipped .zarr.zip stores use byte-range requests. Their host must honour
Range, include range in AllowedHeaders, and expose Content-Range,
Content-Length, Accept-Ranges, and ETag. The range headers let the viewer
validate partial responses; ETag preserves archive identity across
cross-origin cache validation.
Verify both paths. Curl does not enforce CORS, so inspect the response headers explicitly:
curl -sI -H "Origin: https://luxarviewer.dev" "$DIRECTORY_OBJECT_URL"
# want: access-control-allow-origin, access-control-expose-headers including "etag"
curl -sI -X OPTIONS -H "Origin: https://luxarviewer.dev" \
-H "Access-Control-Request-Method: GET" \
-H "Access-Control-Request-Headers: if-none-match" "$DIRECTORY_OBJECT_URL"
# want: 204, access-control-allow-headers including "if-none-match"
curl -sI -X OPTIONS -H "Origin: https://luxarviewer.dev" \
-H "Access-Control-Request-Method: GET" \
-H "Access-Control-Request-Headers: range" "$ZIP_URL"
# want: 204, access-control-allow-headers including "range"
curl -sD - -o /dev/null -H "Origin: https://luxarviewer.dev" \
-H "Range: bytes=0-0" "$ZIP_URL"
# want: 206 and access-control-expose-headers listing all four headers above
4.4 One hostname, one Pages project
A domain attached to two Pages projects is ambiguous. After splitting the viewer out of the gallery, confirm each project claims only its own:
curl -s "https://api.cloudflare.com/client/v4/accounts/$CF_ACCOUNT_ID/pages/projects/<project>/domains" \
-H "Authorization: Bearer $CF_API_TOKEN"
5. Verifying a wave
Run all of these; each catches a class the others cannot see.
Live-site audit — every tile, scene URL, still and video resolves; no placeholders; bylines match
DEMO_META; bucket store count equals tile count. Target: 0 problems.Browser probe — load a sample of scenes in the real viewer, assert non-zero elements and no console errors. A store can be perfectly published and still render nothing.
Cache and CORS — §3.2 and §4.3, against cached objects.
Pre-flight before purging the old prefix — deleting from R2 is irreversible, so first prove nothing references it: grep the gallery page, the apex page, and the deployed JS bundles for the old prefix string, and confirm every store in the old prefix also exists in the new one.
5.1 Do not publish these
chromatrace_choir_umap and chromatrace_choir_umap_sequence are unpublished
research and must never reach the bucket. The external publish script carries
both an exclusion list and a hard abort guard that fails the run if their media
appears in the upload set. Keep both — the list alone has no teeth.
6. Framing tiles
Gallery framing lives in scripts/gallery/manifest.json; see
scripts/gallery/README.md for the field list. Two things worth knowing here:
fillTargetdefaults to 0.95, which is wrong for round or dense subjects. A globe or a capsid at 0.95 is cropped past its own silhouette and reads as a texture wall. 0.65 is the value that keeps recurring for these.Judge a tile on the render, not on the metric. The
coveragescore uses a percentile bounding box that ignores exactly the frame-edge outliers that make a tile unreadable — it rated a broken framing 93% and the correct one 63%.border-litis the better signal, but it too is inflated by legitimately bright limbs (a luminous cloud shell gains rim luminance seen edge-on). Look at the image.Worst-orbit-pose border-lit does not answer whether the still is framed correctly. An elongated subject projects wider as it rotates. Tribolium’s correct poster frame still measures 48.4% because the embryo reaches the edge at
rock +15°; shrinking it to satisfy that warning would underfill the tile.
7. Reading a scene’s LOD and partition structure
Structure decisions get made from these numbers, so getting the accounting right matters more than it looks. Every rule below is here because assuming the obvious reading produced a wrong answer.
For a remote zip, the central directory is enough to inspect its topology without downloading the full payload. Fetch it with zip-aware tooling or a progressively larger trailing byte range: the directory size scales with the entry count and can be several MiB, so a fixed tail length is not a reliable bound.
7.1 Substitutive levels are alternatives; additive rungs are deltas
A kind=lod group’s child_0/1/2 are decimated copies of one another, so a
scene’s content is the max over levels, never the sum. Within one level, the
additive_N rungs are deltas and do sum — verified on desi_galaxies, whose
child_0 rungs run 2000, 2000, 4000, 8000, 16000 … and total exactly 152,262,
matching the level.
Summing across levels instead reported that store as “11,123,187 elements” when it holds 9,751,955 with 1.37M of ladder redundancy above it.
A logical node’s size is likewise the sum over its parts, not the largest
single array. desi_galaxies’ finest level reads 900,000 if you take the
biggest positions array and 9,751,955 if you total its parts.
And the rule that moves the most numbers: only the RESIDENT slice counts. A
node stacked on a hidden axis is measured per hidden coordinate, not by its
total — demos/_lod_policy.py states this. human_multiome_peak_umap totals
6,248,730 across six attribute views but is 1,041,455 resident; measuring the
total over-flags it against any cap.
Two traps inside that:
Measuring one part under-reports by the part count — a first pass read
nuclear_pore_complexat 164,633 when it is 4,937,064 resident per state, a 30x error, because the path measured was a singlepart_N.Summing each part’s largest slice over-reports, because different parts peak on different hidden coordinates. Group globally by hidden coordinate first, then take the max.
Never infer element counts from physical array shapes. An array_ref encoding
stores a deduplicated array with a zero first dimension; use the node’s
n_points / n_splats, or encoding.original_shape when inspecting that array.
7.2 shape=[0] means array_ref
A byte-identical duplicate of another array in the same store is encoded as an
empty (0,) / (0, D) array with encoding.name="array_ref"; target names
the source array and original_shape records the logical shape. Readers resolve
the target. A zero-shaped array with that encoding is normal; one without it is
wrong. See Array Encodings.
7.3 The BSP tree is bsp_tree, on the kind=partition wrapper
New stores use zarr format 3: each node’s attributes are nested under
attributes in its zarr.json; there is no .zattrs. A format-2-only scanner
therefore finds nothing and reports, wrongly, that partitions carry no BSP
metadata. Use Luxar’s bi-format metadata readers when inspecting stores.
The serialized form is a nested dict with left / right and an axis per
internal node. The algebra on it lives in core/group/partition.py:
prune_serialized_bsp_tree, map_serialized_bsp_tree,
reconstruct_serialized_bsp_tree, serialized_bsp_tree_separates,
serialized_bsp_tree_straddles_centers,
serialized_bsp_tree_axis_overlap_floors, and persist_pruned_bsp_tree.
Historical measurement from the retired published prefix 2026-08-27b. The
live prefix is per store (8.1; resolve it from the /d/<demo-key> route):
data/2026-09-12 backs ocean_currents_earth, while
biodiversity_planetary_scale is at data/2026-09-19:
store |
node |
depth |
leaves |
axes |
|---|---|---|---|---|
|
|
4 |
16 |
0,1,2 |
|
|
2 |
3 |
0,1 |
|
|
1 |
2 |
2 |
A depth-1 single-axis entry like By taxon & period is a planar cut rather than
a spatial tree, and is worth checking: the serialized axis is a centre-column
index the viewer maps through displayDims, so a split on a non-displayed
dimension culls nothing.
7.4 Additive LOD refuses topology-changing indexed edges
indexed is the only line type that groups connected vertices into whole
component units, via lod/lines.py::_indexed_connected_components; for chain
inputs, those units are whole polylines. A multi-level ladder can only write
each component back as a chain in ascending vertex order, so that rebuild —
not the authored edge list — is what the rungs would carry. It is also exactly
what the writer now checks the authored edges against, before that node is
written.
A plain add_lines(..., line_type="indexed", additive_lod=...) that resolves
to more than one level raises ValueError unless the rebuild matches: every
connected component’s authored undirected edge multiset — multiplicity
included — must equal its consecutive-vertex pairs. Edge direction and row
order do not matter; a duplicated edge does. A request that resolves to a
single level is unaffected and writes the authored edges flat.
Converting to segments lifts the restriction, but identify_polylines
returns n // 2 arrays of shape (2,), so the laddering unit becomes a single
segment. A prefix is then scattered segments, i.e. fragmented polylines rather
than whole ones.
Additive LOD composed under substitutive_lod= uses the same
indexed_components_are_chains contract. Only a finest indexed level whose
topology fails that check loses its ladder; the coarse levels are lifted gsplat
clouds and keep their ladder where one applies. An explicit additive request
emits a UserWarning; the default composed ladder is skipped with an
informational message.
partition= plus indexed Lines and an explicit additive ladder is preflighted
against the whole node before the first part is written. That eager check is
intentionally stricter than waiting for each part’s resolved ladder, but avoids
leaving a partially written partition when the topology cannot be preserved.
The refusal shipped in 172b686e6 (#2321):
lod/lines.py::_validate_indexed_ladder_edges runs at both multi-level return
paths of make_additive_lod_lines and raises rather than warns, so an explicit
request that cannot be honoured fails loudly instead of writing rewritten
edges.
7.5 Compare like with like
Published and local copies of the same store can differ structurally, not
just in freshness. desi_galaxies has no partition when published and a
depth-2 BSP locally; nuclear_pore_complex is 41,288 elements published and
9,874,128 across both local states after a demo change. Label the source of every
number, and never put both in one table.
8. Current state, and what a fresh operator needs
Updated 2026-09-25. Read this before touching the site; several items below change what the earlier sections tell you to do.
8.1 Where the site stands right now
Gallery data prefixes, re-resolved from the stable
/d/routes on 2026-09-25:data/2026-09-12backs 64 tiles anddata/2026-09-16backs 16 tiles. The later waves are one tile atdata/2026-09-18(gsplats_4d_drosophila_embryogenesis), two atdata/2026-09-19(biodiversity_planetary_scale,gsplats_3d_opencell_map4), two atdata/2026-09-20(gsplats_4d_h2afva_timelapse,gsplats_4d_neuromast_2ch), and three atdata/2026-09-22(gsplats_3d_blastocyst_dapi_nuclei,gsplats_3d_blastocyst_multichannel,protein_landscape).There is no rollback prefix.
2026-09-01was purged after verification. Recovery is a rebuild from the record archives (§8.3) plus a redeploy, not a repoint. Do not plan around a fallback that does not exist — confirm withrclone lsf r2:luxar-demos/data --dirs-onlyrather than assuming.Gallery, derived from
scripts/gallery/manifest.jsonon 2026-09-25: 88 tiles and 86 videos. Every tile retains a still; two are deliberately still-only (§8.2).95 stable
/d/routes (§8.1.1).Tiles and
/d/routes open the STANDALONE viewer,https://luxarviewer.dev/?src=<data URL>, in a new tab — not the gallery’s own copy. The gallery used to ship a second viewer build at/viewer/, so every release had to be deployed to two Pages projects and they drifted (the standalone once sat 3 days stale).demos.luxarviewer.dev/viewer/is STILL SERVED for links already in the wild; retire it by replacing the deployed copy with/viewer/* https://luxarviewer.dev/:splat 302, and only then drop the viewer overlay from the deploy tree and collapse the stamp/asset checks to one host. Verified before switching: the data host’s CORS allow-list already includesluxarviewer.dev, and a tile URL carries only?src=— all per-demo framing lives in each store’s bakedviewer_config.
8.1.1 Stable per-demo routes — /d/<demo-key>
https://demos.luxarviewer.dev/d/<demo-key> 302s to the viewer at that
store’s current prefix. Anything durable — the README, a paper, an issue, a
chat message — should link to that, never to a dated ?src= URL, which breaks
on the next deploy and breaks silently (the origin answers a miss with
200 text/html, so the reader gets a blank viewer and nothing goes red).
Generated by scripts/gallery/gen_redirects.py (in this repo, with tests),
which writes a static Pages _redirects. Regenerate it on every deploy or
the routes rot exactly like the links they replace:
rclone lsf r2:luxar-demos/data \
--recursive --max-depth 2 --dirs-only > /tmp/live_stores.txt
python scripts/gallery/gen_redirects.py \
--prefix <newest-prefix> --live-stores /tmp/live_stores.txt \
-o <deploy-tree>/_redirects --check-contract
The recursive listing supplies <prefix>/<store> entries across every live
prefix. --max-depth 2 stops at the store directory instead of walking every
group, array, and chunk directory inside each Zarr store. --prefix is the
fallback only for bare store names in a hand-authored list. After an incremental
publish, a store may appear under both its old and new dated prefixes; the
lexicographically newest dated prefix wins, matching the deployment convention,
and the older copy remains the rollback. Prefixes must be one path segment, so
running the command from one level above data/ fails rather than generating
targets containing a doubled data/ segment.
Before publishing, verify that the generated # prefixes: header lists the live
prefixes actually used by routes and spot-check every newly published demo’s
encoded target:
grep '^# prefixes:' <deploy-tree>/_redirects
grep '^/d/<demo-key> ' <deploy-tree>/_redirects
The header intentionally uses # prefixes: (plural); update any deploy-side
verification that still matches the former singular key. Do not hand-stitch
fragments or replace this with a single-prefix union invocation.
Two things it handles that a reimplementation gets wrong:
A demo key is not always its store name — 7 of 88 manifest entries differ:
cosmicflows_laniakea→cosmicflows_laniakea_full;nd_transforms→nd_transforms_bench;ppi_flow_field→ppi_flow_field_full;particle_collision_animated→collision_animated;zebrahub_velocity_streamlines→zebrahub_velocity_streamlines_standard;gsplats_interop_observatory→gsplats_interop_observatory_rubin; andgsplats_interop_spz_scaniverse→gsplats_interop_spz_hornedlizard. The two interop demos also emitgsplats_interop_observatory_gemini-southandgsplats_interop_spz_racoonfamily, respectively; those secondary outputs have no manifest entry and therefore no stable route, but still belong in operator upload/rebuild lists. The retirednetwork_performancedemo likewise servedperformance_test. Gallery capture/generation selectors use demo ids; served R2 directories and operator upload/rebuild lists use store names. This generator takes the store from the manifest entry’sdatasetfield, never itsid, and emits routes for both spellings.--check-contractfails the build naming any README-linked key without a route. The root README links 29 tile titles at these routes, so a key rename is a breaking change. If a rename is genuinely needed, add an alias route for the old key so existing links keep resolving.
There is deliberately no /d/* catch-all: an unknown key should not quietly
land on the gallery, because that hides a typo behind a page that looks fine.
Also run, report-only, at deploy time:
python scripts/gallery/audit_readme_demo_count.py --page <deploy-tree>/index.html
The README states the live demo count in two places. scripts/sync_demo_counts.py
owns those claims (and the matching docs/index.rst line), projecting the length
of scripts/gallery/manifest.json onto them under the CI gate
hatch run check-demo-counts; the deploy-time audit reads the same claim table
and checks it against the tiles actually in the built page, which CI cannot see.
8.2 noOrbitVideo — two tiles are deliberately static
Since #2477, gsplats_lod_embryo_line and gsplats_2d_codex_pancreas carry
"noOrbitVideo": true in scripts/gallery/manifest.json (verified
2026-09-25). Regenerating the gallery therefore keeps both tiles static. Their
subjects are planar arrangements viewed face-on, so any rock swings them toward
edge-on: embryo_line’s mean luminance swings 13x over ±20°, twice per loop,
which reads as violent flashing. Measured, at ±6° it is still 8.2x — no
amplitude fixes it. Same convention as exotic_surfaces (“near face-on only,
NO orbit video”).
The .webp IS the animated loop — build_gallery_data.py uses it as the
tile’s still, and has_media is video or still. So dropping only the
.webm leaves the pulsing in place; the flag encodes a static single-frame
webp, skips the webm, and removes any stale staged webm. Their published
.webm objects were deleted from both the deploy tree and R2. If you re-enable
orbit video for either demo, regenerate and publish both media variants rather
than reviving one old object by hand.
8.3 Build from the record archives — and verify provenance
The site is normally built from the Zenodo record generation, held at
~/luxar-zenodo-archives/ (52 files, MANIFEST.json, SHA256SUMS). Datasets
with a manifest base_url instead use their pinned mirror archives under
inputs/<slug>/; preserve those objects because they may have no Zenodo or Git
LFS copy. Verify the local record archive before use — it takes seconds and the
stale-generation incident below began with an unverified local archive:
cd ~/luxar-zenodo-archives && shasum -c SHA256SUMS # expect: 52 OK, 0 failed
Verified on 2026-09-25, the manifest declares one current checksum contract per
artifact and contains no hosted_sha256 fields: #2454 fixed the stale-cache
acceptance and repair paths, and #2431 removed the former dual pins. The resolver
still supports that legacy field, so keep the dormant guard below. Do not assume
a verified cache entry is the wrong generation. Instead, verify the generation
the scene actually records. Every manifest-backed demo stamps a canonical
scene-root input_digests map from the exact bytes accepted by the resolver.
When stamped before finalization, as the demos do, that map participates in
content_hash (#2475):
SCENE=/path/to/scene.luxar.zarr
hatch run python - "$SCENE" <<'PY'
import json
import sys
from pathlib import Path
from luxar._zarr_compat import read_node_attrs
attrs = read_node_attrs(Path(sys.argv[1]))
if attrs is None:
sys.exit(f"no readable scene at {sys.argv[1]}")
print(json.dumps(attrs.get("input_digests", {}), indent=2, sort_keys=True))
PY
Compare each basename and digest with the record’s SHA256SUMS, or with the
manifest sha256 for a dataset using a pinned mirror. This is the primary
post-build generation check. An empty map means either that the demo recorded no
manifest inputs or that the stamp is absent, including a scene built before
#2475. For a manifest-backed demo that should have archive inputs, treat an empty
map as failed provenance and rebuild it. As a log-side cross-check, also confirm
that input files were verified, no superseded generation was actually used, and
the legacy dual-contract warning never appeared:
grep -c "SHA256 verified" <build-log> # want: > 0
grep -c "SUPERSEDED" <build-log> # want: 0 (an earlier generation was used)
grep -c "record hosts a newer build" <build-log> # want: 0 (dormant legacy guard)
Checking element counts is not enough on its own. The stale-generation incident produced large pass-through discrepancies, but a transformed demo can produce a larger discrepancy legitimately:
demo |
archive/reference |
compiled/served |
interpretation |
|---|---|---|---|
|
29,579,229 |
20,591,415 |
stale generation |
|
99,921 |
40,253 |
stale generation |
|
4,173,532 |
1,283,624 |
expected compilation |
Compare scene-from-old against scene-from-new, or use input_digests as above;
do not infer the input generation from the size of one compiled scene.
8.4 Environment traps on this machine
Each of these cost hours and none announces itself:
A missing
grimpwheel brokehatch runbox-wide. The source build needed rustc 1.94 while the box had 1.92, killing everymake test-*/check-*and gallery capture. The workaround was to call the built venv interpreter withPYTHONPATH=packages/luxar/src; noteluxar.cliis a package exposing a Typerapp, so-m luxar.cliwill not run — drive the app object. Upstream fixed this on 2026-09-11 by backfilling the standard CPython 3.14 macOS arm64 wheels for 3.16/3.17.urllib’s default User-Agent is blocked at the edge. It presents as a uniform failure across every URL, which reads like broken routes rather than a blocked client. Use a browser UA. Never trust a negative HTTP result without showing the tool.Never run two Chromium capture/probe jobs concurrently. It wedged both, cost ~2 h, produced a 1-hour Playwright timeout on
page.evaluate, and left a manifest edit unrestored. Any script that edits a tracked file should restore it in a trap, not at the end.
8.5 What is in this repo, and what is not
In the repo (versioned, tested, reviewable):
scripts/gallery/gen_redirects.py, audit_readme_demo_count.py,
check_tile_staleness.py, verify_media.py, manifest.json,
media-manifest.json, and their tests.
Not in the repo — the operator harness at
~/workspace/luxar-hosting-harness/, which holds the deploy shell scripts
(build_page.sh, publish_refresh.sh, deploy_gallery_static.sh),
gen_landing.py, build_gallery_data.py, and the verification sweeps.
That directory is not a git repository — no history, no remote, no backup.
Treat it as a single point of failure and copy anything you change.
pipeline/site-2026-09-02/ there holds the tooling, evidence and a README for
the 2026-09-02 publish.
Bulk data: ~/workspace/luxar-site-staged-2026-09-02/stores/ holds the 86
staged stores (7.8 GB). Read its README first — it is the local rebuild, and
22 stores differ from what the site serves because those were server-side copied
rather than uploaded. For a true mirror, rclone copy from the live prefix.
8.6 Minimum competent update
To change what the site serves:
Verify the archives (§8.3), then verify
~/.cache/luxaragainst them.Rebuild the affected demos; confirm the logs show hosted-digest reads and no stale warnings.
Re-chunk each store with its published profile, read per store — 86 are
archive(1 MB), 4 arelocal(64 KB). Do not apply one profile to all.Publish to a new dated prefix; server-side copy the unchanged stores from the current one rather than re-uploading them.
Rebuild the page, regenerate
_redirects(§8.1.1), deploy.Verify: store-count parity, every served
content_hashmatches what you published, the live page carries a tile per store all on the new prefix, and a render probe loads every scene with zero failed nodes.Only then purge the old prefix — and re-check the live page references it nowhere first.
Never publish chromatrace_choir_umap or chromatrace_choir_umap_sequence
(§5.1), and never publish a Zenodo record — that is Loic’s manual step.