Evaluation result with aggregate score, per-example rows, and call errors.
Rows keep the original example, prediction, normalized metric score, pass/fail state, feedback, metric metadata, and any program error. Optimizers use the same structure that you can inspect in tests and notebooks.
save_as_json/2 and save_as_csv/2 port DSPy's save_as_json/save_as_csv
evaluator options (dspy/evaluate/evaluate.py). Deviation from upstream: the
file surface lives here on the result — the rows are already plain data — not
as evaluator constructor options. Row semantics match upstream's
_prepare_results_output: example fields merged with prediction fields (a
key present on both sides becomes example_<key> and pred_<key>), plus a
score column (upstream names this column after the metric function; Imp
metrics are anonymous functions, so the column is always score). An
Imp.History value serializes as %{"messages" => [...]}, matching
upstream's History-in-example serialization.
Summary
Functions
Writes the per-example rows to path as CSV.
Writes the per-example rows to path as a JSON array of flat objects.
Types
Functions
Writes the per-example rows to path as CSV.
Mirrors DSPy Evaluate(save_as_csv=...). The header is the union of row
keys (sorted, score last; upstream takes the first row's keys and fails
on ragged rows — the union keeps every column and is deterministic).
Non-scalar cells are JSON-encoded, mirroring upstream's stringified dicts.
Saving an empty result raises: an empty file with no header would be a
silent failure.
Writes the per-example rows to path as a JSON array of flat objects.
Mirrors DSPy Evaluate(save_as_json=...): one object per devset row with
the example fields, the prediction fields, and the score.