Fully decode a per-channel quantized array (log_perchannel_* /
signed_log_perchannel_* / linear_perchannel_*) into a Float32 output.
This makes per-channel a first-class self-decoded encoding — like
quantized / lut / broadcasted — so the decode layer owns dequantization
and consumers receive decoded float32 (no consumer-side dequant). Mirrors the
Python ArrayDecoder._decode_*_perchannel: raw integer levels are dequantized
with the array's own per-column col_lo/col_hi scales.
elementsPerItem is the column count C (the per-channel dimension, e.g. ndim
for positions/centers, d for the Cholesky diagonal): the flattened output is
[item*C + col], so col = globalIndex % C.
Ranges are fetched CONCURRENTLY (like loadDirect / loadQuantized):
destination offsets are precomputed, each range writes its own disjoint
output span, and the global flattened index offsets[i] + j keeps the
column phase exact regardless of resolution order. Above the worker
threshold, decode runs in the worker pool on the WASM per-channel kernels
(decode_*_perchannel_*) — the log/signed-log Cholesky pair costs an
expm1 per element, which is real main-thread work at millions of splats.
Below the threshold (or on worker failure) the main-thread
makePerChannelDequant closure decodes bit-identically — both paths do the
same f64 math on the same f64 scales.
Fully decode a per-channel quantized array (
log_perchannel_*/signed_log_perchannel_*/linear_perchannel_*) into a Float32output.This makes per-channel a first-class self-decoded encoding — like
quantized/lut/broadcasted— so the decode layer owns dequantization and consumers receive decoded float32 (no consumer-side dequant). Mirrors the PythonArrayDecoder._decode_*_perchannel: raw integer levels are dequantized with the array's own per-columncol_lo/col_hiscales.elementsPerItemis the column count C (the per-channel dimension, e.g. ndim for positions/centers, d for the Cholesky diagonal): the flattened output is[item*C + col], socol = globalIndex % C.Ranges are fetched CONCURRENTLY (like
loadDirect/loadQuantized): destination offsets are precomputed, each range writes its own disjoint output span, and the global flattened indexoffsets[i] + jkeeps the column phase exact regardless of resolution order. Above the worker threshold, decode runs in the worker pool on the WASM per-channel kernels (decode_*_perchannel_*) — the log/signed-log Cholesky pair costs anexpm1per element, which is real main-thread work at millions of splats. Below the threshold (or on worker failure) the main-threadmakePerChannelDequantclosure decodes bit-identically — both paths do the same f64 math on the same f64 scales.