CloudDelta

View Source

2D and 3D point-cloud compression in Elixir. A cloud is treated as a set of points: Morton-order, optional quantization, delta encode, zlib.

Honest status

v0.1 advertised 7.99:1 lossless compression and “0.01 bits/delta”. Those numbers were not real:

  • The published metric summed unique Huffman code lengths and counted one 8-bit permutation array. It never measured the bytes compress/1 wrote, never stored a Huffman tree, and never decoded residuals.
  • Independent sorting of X and Y breaks pairing. Restoring the original points requires two permutations (2 n log2 n bits). That cost alone makes the old scheme expand on unique float32 data.
  • The shipped encoder stored every float again inside a Huffman tree and used 8-bit indices (broken for n > 255). Actual binaries were ~4× larger than raw float32, and round-trips were not lossless.

This tree replaces that pipeline. Ratios below are raw_f32_bytes / encoded_bytes of a real binary, checked by decompressing it.

What it does now

ModeReconstructionWhen it wins
mode: :quantized (default, 16-bit)Bounded error, about one quantumStructured / clustered / gridded clouds
mode: :losslessBit-exact float32Near-incompressible unique floats; aims to match or beat zlib on the packed points

Unique random float32s are already near entropy. No lossless codec will turn those into 8:1. Competence here means: never lie about size, actually round-trip, and beat naive zlib on clouds with spatial structure.

Synthetic 2D, n=10,000:

Patternzlib rawlossless16-bit12-bit
random1.131.432.403.99
clustered1.151.602.835.25
grid1.301.655.0516.16

Real 3D (4 Sep 2026). Ratio vs packed f32. Quantized LAZ (laspy/lazrs) and G-PCC (tmc3 v23-rc2, octree, geom-only, no angular mode) use the same integer grid as CloudDelta. Draco 1.5.7 uses its own -qp on the original floats.

Lossless (bit-exact float32; LAZ/G-PCC do not apply):

DatasetnzlibzstdCloudDelta
bun000 range scan40,2561.772.362.86
bunny zipper35,9471.101.091.31
armadillo172,9741.741.601.77
KITTI Velodyne 000000125,6351.561.351.81
Autzen ALS trim110,0001.971.932.66

12-bit, encoded bytes. “vs X” is CloudDelta / X (below 1 means we are smaller):

DatasetCDDracoLAZG-PCCvs Dracovs LAZvs G-PCC
bun00074,04686,47637,35953,8690.86×1.98×1.37×
zipper92,94387,486115,60575,4371.06×0.80×1.23×
armadillo353,563330,227692,264238,1601.07×0.51×1.48×
KITTI205,67999,678155,907153,7022.06×1.32×1.34×
Autzen259,816149,395190,676200,6191.74×1.36×1.30×

Morton+delta beats quantize-then-zstd on every cloud. It does not beat G-PCC on any of these sets, and it loses badly to Draco on LiDAR at 8-bit (KITTI 44,778 B vs Draco 8,437 B). LAZ wins the structured range scan and loses on dense reconstructions. The method is in the same conversation as the standards — a different, simpler pipeline — not a replacement for them.

Encode time, 12-bit, median of 3 wall-clock runs (Apple M-series, 4 Sep 2026). CloudDelta time is compress/2 (quantize + Morton + zlib). Draco and G-PCC times are the encoder process after the input file is written. LAZ time is laspy write/compress only (no Python startup).

DatasetnCloudDeltaDracoLAZG-PCCCD / DracoCD / G-PCC
bun00040,25640 ms11 ms3.6 ms42 ms3.8×0.96×
zipper35,94736 ms11 ms4.2 ms44 ms3.4×0.83×
armadillo172,974258 ms46 ms9.1 ms158 ms5.6×1.6×
KITTI125,635193 ms30 ms8.3 ms84 ms6.4×2.3×
Autzen110,000177 ms26 ms6.4 ms100 ms6.8×1.8×

Elixir vs C++ is most of the Draco gap. Against the G-PCC reference encoder the gap is small on the Stanford scans and about 2× on LiDAR. LAZ is in another speed class. Lossless CloudDelta is 2–5× slower than zlib and ~10–20× slower than zstd; that is the cost of the spatial pass.

Decode time, 12-bit, same method (uncompress_points/1, draco_decoder, laspy read, tmc3 --mode=1):

DatasetCloudDeltaDracoLAZG-PCCCD / DracoCD / G-PCC
bun00015 ms5.5 ms4.5 ms51 ms2.8×0.30×
zipper14 ms4.8 ms5.7 ms50 ms2.8×0.28×
armadillo84 ms14 ms9.7 ms201 ms6.2×0.42×
KITTI64 ms10 ms7.3 ms131 ms6.2×0.49×
Autzen61 ms8.8 ms7.0 ms127 ms7.0×0.48×

Decode is CloudDelta’s better number: about 3× faster than its own encode, and 2–3.5× faster than G-PCC decode on every set. Still 3–7× behind Draco, and LAZ remains the speed class of its own. Encode pays the Morton sort; decode is zlib inflate plus a prefix-sum walk.

Usage

{x, y} = CloudDelta.Benchmark.generate_dataset(10_000, :clustered)

compressed = CloudDelta.compress({x, y}, mode: :quantized, bits: 16)
{x2, y2} = CloudDelta.uncompress(compressed)

lossless = CloudDelta.compress({x, y}, mode: :lossless, preserve_order: true)
true = CloudDelta.check_compression({x, y})

CloudDelta.stats({x, y}, mode: :quantized, bits: 16)

Options:

  • :mode — :quantized (default) or :lossless
  • :bits — quantization bits per axis, 4..24 (default 16)
  • :preserve_order — restore input order (default false)

Benchmark

CloudDelta.Benchmark.run_benchmark_suite()

# Real clouds + RD curve + zlib/zstd/Draco/LAZ/G-PCC
CloudDelta.Compare.run()

Put Stanford bunny/ and Armadillo.ply under priv/datasets/ (see 3D Scanning Repository). Optional LiDAR: a KITTI Velodyne .bin and/or a LAS/LAZ under priv/datasets/lidar/. LAZ needs laspy + lazrs (see .tools/.venv). G-PCC needs tmc3 from MPEG TMC13 (TMC3 or .tools/tmc13/build/tmc3/tmc3).

Installation

def deps do
  [
    {:cloud_delta, "~> 0.2.0"}
  ]
end

License

MIT