Sound Spec — ambient and spatial audio as scene-graph nodes
Status: Design agreed with the project owner on 2026-09-06; all four phases
are implemented in draft PR #2566 (feat/sound-layer, stacked on PR #2536, which
it needs for the story waypoints its triggers use; it stays a draft until #2536
merges). Implementation decisions and deviations are marked inline below.
1. Motivation
A scene can already be seen from a camera pose and read through overlays.
The kiosk (see REMOTE_CONTROL_SPEC.md) wants it to be heard too: an ambient
bed that changes with the story, a voice that narrates a cluster on arrival, and
sounds that come from places in the data and get louder as the camera approaches.
Sound is a layer on top of rendering and independent of it — like overlays — but
spatial sources have a position in the scene’s reference frame, which makes them
nodes, not overlays.
Decisions taken (owner, 2026-09-06):
Question |
Decision |
|---|---|
Playback at the event |
Room speakers, stereo. Equal-power panning by default; HRTF authorable. |
First uses |
Narration per story on arrival; a continuous ambient bed; spatial cluster sounds. |
Assets |
Stored in the zarr as opaque files, like overlay images. |
Autoplay rule |
Both: Chrome kiosk flag when available, a one-time “tap to enable sound” gate as fallback. |
Data model |
A first-class |
Narration |
Text-to-speech at scene build time: auto-detect macOS |
Ambient / cluster clips |
Recorded CC0 clips (Freesound and the like), licence recorded per node. |
Default |
Authored per scene; viewer default ON with a mute control shown only when the scene has sound nodes. |
Extras |
Recording panel captures audio; ambisonic (first-order) beds rotating with the camera — decoded by a dependency-free native graph, NOT Omnitone (Phase 4 decision: Omnitone decodes to binaural only, wrong on room speakers; fetches HRIRs from a CDN at runtime, impossible on an offline kiosk; unmaintained). |
Branch |
New branch and PR stacked on #2536. |
2. Framework
Web Audio API through three’s own wrappers. No new dependency.
THREE.AudioListener attaches to the camera and follows it every frame;
THREE.PositionalAudio is an Object3D wrapping a PannerNode (position,
orientation, distance model, cone, HRTF or equal-power panning);
THREE.Audio is the non-spatial sibling. Three’s wrappers already handle the
AudioContext singleton, buffer sources, gain, loop and play/stop.
Not chosen: Resonance Audio (archived); Howler / Tone.js (nothing for playback we lack; Tone matters only for synthesis, which is out of scope); Omnitone (kept for Phase 4 ambisonics only). Doppler no longer exists in Web Audio; nothing lost.
3. Data model
3.2 On disk
/sounds/hum_hsp70/ # a group under any group, like other nodes
zarr.json # type: "sound", attrs below
positions # (K, ndim) PLAIN float32 (not the quantizing
# encoder: a one-row array collapses to code 0
# under per-channel uint16), absent for non-spatial
audio.mp3 # opaque
attrs: type, spatial, trigger, delay_ms, gain, bus, loop (derived), fade_in_ms,
fade_out_ms, distance_model, ref_distance, max_distance, rolloff, cone_*,
orientation, attach_to, format, duration_ms, license, attribution,
source_url, layer (bool), extend_to_all, transform, nd_transform,
ambisonic ("foa" for a first-order AmbiX ACN/SN3D field; absent otherwise),
channels (stamped when a tag reader is present)
format is the CODEC (mp3 / aac), never the spatial layout: an ambisonic
field is format: "aac" + ambisonic: "foa", and the writer refuses a field
that is not four-channel. World front is −Z; the viewer decodes with nine gains
over the dipoles plus a virtual-cardioid pair at ±60° for the room’s stereo.
content_hash covers the audio bytes (the same gap #1720 closed for overlay
PNGs must not reopen).
3.3 viewer_config.audio
ViewerConfig(audio=AudioConfig(
enabled=True, # viewer default when the scene has sound nodes
master_gain=0.8,
panning_model="equalpower", # or "HRTF" for headphones
buses={"ambient": 0.6, "voice": 1.0, "effects": 0.8},
duck_db=-9.0, # ambient attenuation while voice plays
))
4. Viewer architecture
4.1 src/audio/
AudioEngine— owns theAudioContext(via three’s listener), the master gain, the three buses and the ducker; handles the autoplay gate (§4.4); exposessetMasterGain,setMuted,setPanningModel,play(name),stop(name), and emitssound-started/sound-ended.SoundNode— one per scene node: aTHREE.PositionalAudio(spatial) orTHREE.Audio(non-spatial) parented under the node’s transform group, its buffer decoded once from the opaque file, fades implemented on the node gain. Reacts to the slab result the same way a points node does: audible ⇄ silent with fades,continuousrestarts on the rising edge,oncefires on it.Listener — a
THREE.AudioListenerchild of the camera. Three updates its position and orientation from the camera’s world matrix each frame; nothing per-frame is added on the Luxar side.Loader — a
SoundWholeNodeLoader(whole node, no chunking): readspositions, runs the existing per-vertex slab kernel for audibility, fetches the clip through the opaque-file reader.
4.2 Layers panel and rail
Sound nodes are layers: eye = mute that node, plus a per-node gain slider and the licence tooltip. A speaker rail item (mute all, master gain) appears only when the loaded scene has at least one sound node; its muted state persists like other rendering settings.
4.3 Triggers and the waypoint driver
WaypointDriver gains two events on the embedder bus,
waypoint-departed {index} and waypoint-arrived {index, completed} (arrival
fires when the flyTo promise resolves, or immediately after a snap). Sound
nodes with on_depart / on_arrive subscribe and match the waypoint’s when
against their own row. A flight cancelled by the visitor still resolves
(completed: false), so narration still starts — but from wherever the camera
stopped, which is the right behaviour. A flight SUPERSEDED by a newer waypoint
never fires waypoint-arrived, so two narrations cannot overlap when a visitor
steps quickly. An on_arrive clip fades out on the slab’s falling edge (moving
on cuts the previous story’s narration); an on_depart clip plays out.
4.4 Autoplay
Browsers refuse to start an AudioContext without a user gesture on the page.
Kiosk: launch Chrome with
--autoplay-policy=no-user-gesture-required(documented next to the kiosk block inREMOTE_CONTROL_SPEC.md§4.3). The context starts on load.Fallback: if the context is not
runningafter load, the engine shows a minimal “Tap to enable sound” gate (an overlay, dismissed by the first pointer or key event anywhere), resumes the context, and only then startscontinuousnodes — and re-runs the rising edges from a silent baseline, so an openingoncenarration is not lost to the tap (Phase 1 clarification). The remote API reportsaudio.stateso a controller can tell the display needs its tap.The overlay is decided from
context.state, NOT from the outcome ofresume(): a context held by an autoplay policy may leave that promise unsettled, so a.then()/.catch()is not a reliable place to decide anything.The gate is not one-shot. It also opens from
onstatechangereachingrunning, from a one-shot pointer/key listener that retries the resume, and fromenableSound(). Without those the context stays suspended for the whole session unless the listener happens to toggle the mute.Blocked is not muted.
isBlocked()— sound wanted, context not running — is reported separately fromisMuted(), and the rail’s Sound button renders it as a third state whose click resumes the context rather than toggling a mute the listener never set.At most one deferred waypoint trigger per trigger kind. A trigger that cannot start while the gate is shut is held, and superseded by the next waypoint event of the same kind. An arrival therefore does not discard the departure trigger from the same transition. Replacing the prior trigger prevents a listener who walks N stops before sound starts from hearing all N clips at once when it does.
4.5 Remote API (extends REMOTE_CONTROL_SPEC.md)
setAudio({ masterGain?, muted?, buses?, panningModel? }), playSound(name),
stopSound(name), getViewerState().audio = { state, muted, masterGain, buses, playing: [names] }, events sound-started, sound-ended, and the two
waypoint events above.
5. The stories demo
Narration:
luxar.demos._narration.synthesise(text, voice, cache_dir): macOSsay→afconvertto.m4awhen available, or OpenAI TTS whenLUXAR_NARRATION_ENGINE=openaiexplicitly selects it andOPENAI_API_KEYauthenticates the request; otherwise a warning and no narration node. Clips are cached by hash of (text, voice, engine) under the demo cache, so a rebuild with unchanged text costs nothing. OpenAI clips are stamped CC0; macOSsayclips are stamped for personal, non-commercial use only and must not be published. Each story’s narration is its panel text (title, facts, open question) read in order,trigger="on_arrive",delay_ms=600, busvoice.Ambient bed: one CC0 clip,
trigger="continuous",fade 1500 ms, busambient(per-story beds or filters later). Phase 1 ships “Calm Ambient 1 (Synthwave 4k)” by The Cynic Project (cynicmusic.com), CC0, from OpenGameArt (https://opengameart.org/content/calm-ambient-1-synthwave-4k), fetched checksum-pinned throughcached_downloadinto theesm3_protein_storiescache: soft evolving pads, no percussion, ~2.6 min loop. It replaced a Freesound clip the owner found “too industrial and harsh” — the bed must be warm and unobtrusive under the narration, and a clip is judged by listening to the WHOLE loop, since many start gently and turn harsh.Cluster sounds (Phase 2): one spatial source per story at the cluster centre, distance falloff tuned from the story’s waypoint camera, live only at that story.
6. Phases
Phase |
Scope |
|---|---|
1 |
|
2 |
|
3 |
Recording panel “Include Audio”: the master-gain tap into real-time WebM (Opus mime first; the offline path stays silent) |
4 |
Ambisonic beds ( |
All four phases shipped together in draft PR #2566 (2026-09-06).
Out of scope: sonification (procedural data-driven sound), live TTS over the remote API (Phase D of the remote-control spec owns it), multichannel outputs.
7. Risks and open questions
Format: MP3 everywhere; AAC on Safari/Chrome/Firefox is fine too; Opus is not decodable on Safari. The writer refuses Ogg.
Store size: a 30 s narration clip is ~0.5 MB as MP3; ten stories plus beds ≈ 10–15 MB, acceptable. Hosted demos may prefer URL sources later.
Clicks: every start/stop goes through a short gain ramp (≥ 30 ms);
continuousnodes fade on the slab edge.Clock:
on_arrivetiming uses the flight promise, not the audio clock; a cancelled flight still resolves. Documented as intended.Two audibility rules? No: a sound node’s audibility is the slab rule the points use;
hidden=is sugar for one row. Overlays keepvisible_range(screen space, no position). The waypointwhenclause is the third use of the same vocabulary.Should the rail mute also silence the autoplay gate’s need? Yes: a muted scene never shows the gate.