One public, fail-closed optimize/select/test lifecycle.
check/5 evaluates the baseline and optimized programs on selection data,
chooses the higher score (retaining baseline on a tie), and only then evaluates
the selected program on untouched test data. The returned artifact contains
both parameter snapshots and can be applied to a freshly reconstructed trusted
program with Imp.Optimizer.Artifact.apply/4. Set
compare_baseline_on_test: true to evaluate the baseline on the same ordered
test rows after the selected artifact has been built and applied.
Evaluation failures remain ordered failure_score diagnostic rows while the
configured error budget has not been exhausted. max_errors: 0 (the
Experiment default) stops on the first ordinary failure, a positive finite
budget stops when that many failures have occurred, and :infinity retains
every ordinary failure. Operational-safety failures always escape
immediately.
For a noisy model, set evaluation_options: [repetitions: n, aggregation: :mean] to repeat each outer selection and test evaluation over
the same ordered rows. Use repetitions: [selection: n, test: m] when the
selection decision and final test estimate require different repeat counts.
The default is one pass. Repetitions change only the family-independent
Experiment admission and reporting boundary; an optimizer's internal
candidate evaluations remain under that optimizer's own documented policy.
A result describes this program, metric, data split, and run configuration. Repeat the lifecycle across representative tasks and conditions before generalizing an optimizer's effectiveness.
Summary
Functions
Runs the canonical train → selection → selected-only test lifecycle.
Functions
@spec check(struct(), struct(), Imp.Experiment.Data.t(), function(), keyword()) :: {:ok, Imp.Experiment.Result.t()} | {:error, map()}
Runs the canonical train → selection → selected-only test lifecycle.