Imp.Experiment (Imp v0.5.0)

Copy Markdown View Source

One public, fail-closed optimize/select/test lifecycle.

check/5 evaluates the baseline and optimized programs on selection data, chooses the higher score (retaining baseline on a tie), and only then evaluates the selected program on untouched test data. The returned artifact contains both parameter snapshots and can be applied to a freshly reconstructed trusted program with Imp.Optimizer.Artifact.apply/4. Set compare_baseline_on_test: true to evaluate the baseline on the same ordered test rows after the selected artifact has been built and applied.

Evaluation failures remain ordered failure_score diagnostic rows while the configured error budget has not been exhausted. max_errors: 0 (the Experiment default) stops on the first ordinary failure, a positive finite budget stops when that many failures have occurred, and :infinity retains every ordinary failure. Operational-safety failures always escape immediately.

For a noisy model, set evaluation_options: [repetitions: n, aggregation: :mean] to repeat each outer selection and test evaluation over the same ordered rows. Use repetitions: [selection: n, test: m] when the selection decision and final test estimate require different repeat counts. The default is one pass. Repetitions change only the family-independent Experiment admission and reporting boundary; an optimizer's internal candidate evaluations remain under that optimizer's own documented policy.

A result describes this program, metric, data split, and run configuration. Repeat the lifecycle across representative tasks and conditions before generalizing an optimizer's effectiveness.

Summary

Functions

Runs the canonical train → selection → selected-only test lifecycle.

Functions

check(program, optimizer, data, metric, opts \\ [])

@spec check(struct(), struct(), Imp.Experiment.Data.t(), function(), keyword()) ::
  {:ok, Imp.Experiment.Result.t()} | {:error, map()}

Runs the canonical train → selection → selected-only test lifecycle.