Stream ETL benchmark: SimdJson vs Jason

View Source

Source revision: 392ffd196312a5e40a0e74a0fa0ae1ef6e1bc361
Jason version: 1.4.5

Both workflows parse the same JSON and calculate the same reduction. The SimdJson workflow uses SimdJson.stream/2; the Jason workflow uses Jason.decode!/1 followed by equivalent lookup and reduction work. The flat million-row fixture has three fields per row (id, value, and unselected ignored); both workflows project id and value from every row. This measures narrow-row streaming ETL, not how select/2 scales as more fields are requested from one wide row.

Acceptance

ResultMillion-row batchSimdJson process peakSimdJson/Jason process-peak ratioMaximum allowed ratioMaximum allowed SimdJson peak
PASS1,000123.93 MiB0.4×0.6×128.0 MiB

Side-by-side measurements

All values are medians (p50).

FixtureRowsBatchWorkflowTotal latency (ms)Time to first row (ms)Rows/sWorker process peak (MiB)Whole-VM RSS peak (MiB)
small100128SimdJson.stream/2 + reduce2.1962.186455370.01 MiB205.57 MiB
small100128Jason.decode!/1 + lookup/reduce0.0540.04918518520.0 MiB205.88 MiB
small1001000SimdJson.stream/2 + reduce2.2352.226447430.01 MiB205.88 MiB
small1001000Jason.decode!/1 + lookup/reduce0.040.03625000000.0 MiB206.04 MiB
medium10000128SimdJson.stream/2 + reduce193.7632.61516092.17 MiB194.88 MiB
medium10000128Jason.decode!/1 + lookup/reduce6.4796.08715434485.53 MiB150.24 MiB
medium100001000SimdJson.stream/2 + reduce25.7932.8813877023.0 MiB156.72 MiB
medium100001000Jason.decode!/1 + lookup/reduce6.3825.93515669078.0 MiB152.68 MiB
million1000000128SimdJson.stream/2 + reduce19122.38838.22252295123.71 MiB526.98 MiB
million1000000128Jason.decode!/1 + lookup/reduce1523.3831473.292656434371.53 MiB1136.87 MiB
million10000001000SimdJson.stream/2 + reduce2890.09737.549346009123.93 MiB675.71 MiB
million10000001000Jason.decode!/1 + lookup/reduce1545.1661494.392647180309.81 MiB1403.47 MiB

SimdJson relative to Jason

Values below 1.0× mean SimdJson used less time or memory than Jason; values above 1.0× mean it used more.

FixtureBatchTotal latencyTime to first rowWorker process peakWhole-VM RSS peak
small12840.667×44.612×n/a0.998×
small100055.875×61.833×n/a0.999×
medium12829.906×0.429×0.392×1.297×
medium10004.042×0.485×0.375×1.026×
million12812.553×0.026×0.333×0.464×
million10001.87×0.025×0.4×0.481×

Measurement definitions

  • Worker process peak is the highest BEAM memory observed for the isolated benchmark worker. Compare it only with the other workflow's worker-process value.
  • Whole-VM RSS peak is the highest resident-set size observed for the entire Erlang VM, including BEAM heaps, native allocations, resident mapped pages, loaded code, shared libraries, and allocator retention. Compare it only with the other workflow's RSS value.
  • Absolute RSS is contextual rather than library-exclusive because both workflows run in the same long-lived VM. File-backed RSS qualifications therefore measure increase from a pre-operation baseline.
  • This benchmark exercises the binary-based stream/2 API. The separate stream_file/2 qualification covers bounded file-backed parser memory.
  • Total latency and throughput are informational. The release acceptance threshold is based on the million-row worker-process peak.