Imp.Optimizer.GEPA (Imp v0.5.0)

Copy Markdown View Source

Program-level GEPA optimizer for Imp programs.

The optimizer exposes every predictor through Imp.ProgramParameters, evaluates named candidate maps through a trace-rich adapter, applies strict minibatch improvement before validation, and maintains the source-shaped per-instance Pareto archive in its internal optimization engine. It rewrites predictor instructions; it does not change tool descriptions.

A real proposal source is mandatory for optimization. Set :reflection_lm or :reflection_strategy; GEPA never fabricates an instruction from the current prompt or diagnostic records. Baseline-only generations: 0 runs do not require a proposal source.

The default is DSPy's GEPA

Without an :execution_profile, GEPA runs :gepa_v0_1_4_merge, the semantics of dspy.GEPA with its defaults, or :gepa_v0_1_4 when use_merge: false is given, as in dspy.GEPA(use_merge=False). These profiles seal the pinned single-proposal search: CPython's persisted MT19937 stream is shared by Pareto selection and Fisher-Yates minibatch sampling, perfect minibatches are skipped at 1.0, evaluation caching is disabled, and one failed batched reflection is retried once as the corresponding single task. Candidate proposal remains serial, while :num_threads may bound the same concurrent row evaluation used by the source artifact. Like pinned GEPA, max_metric_calls is checked between iterations: an iteration that legally starts is allowed to finish. Separate internal metric/reflection envelopes bound that legal overshoot; they are not alternate stopping rules. The budget is :max_metric_calls, or :max_full_evaluations converted as DSPy converts max_full_evals, to that many evaluations of the training and validation sets together; giving both raises. With neither, :generations derives the budget. The reflection call limit is derived from the budget too: :max_reflection_calls raises under the default, and under an explicit pinned profile it may only restate the derived limit. The merge profile uses the source-authenticated common-ancestor merge path of DSPy's ordinary GEPA treatment.

Under these profiles the reflection record mode, parent and component selection, minibatch sampling, proposal selection, frontier, evaluation and acceptance policies, ComBee and proposal concurrency are fixed at DSPy's values, and giving one of them another value raises, as do :feedback_fn and :reflection_strategy. Some of these (module_selector: :all or a custom selector, candidate_selection_strategy: :current_best, a custom proposer) are DSPy options that the pinned profiles do not support yet. The options DSPy passes through to GEPA (:merge_val_overlap_floor, :max_merge_invocations, :merge_acceptance_policy, :stopper, :max_reflection_cost, :callbacks, :component_feedback) are accepted.

execution_profile: :beam_native is Imp's own search, where every option is available. It also caches evaluations, merges only with use_merge: true, never skips a perfect minibatch, draws from a BEAM RNG, takes :generations as its iteration count, and counts :max_full_evaluations beside :max_metric_calls.

What the reflection model reads

Each reflection call asks for a new instruction for one predictor, from records of that predictor's calls on the minibatch. Under the DSPy profiles (reflection_record_mode: :gepa_v0_1_4) a record is DSPy's: Inputs, Generated Outputs and Feedback for one call of the predictor per example, with values in their JSON spelling. When the program calls the predictor more than once, as an agent loop calls its step predictor once per turn, that call is drawn at random, keyed on :seed. A row whose program failed gives no record, and when a component chosen for reflection has no records the iteration ends without a reflection call.

An Imp.History input is shown as Context, one line per turn. For an agent (Imp.react/3, Imp.Predict.ReActV2) that is the whole finished run, taken from the prediction's :history metadata whichever step was drawn: every turn's thought, tool calls and tool results, and the outputs, which the history holds unless finish_on or the extractor ended the turn. The tools the predictor offers its model are shown as tools, with their names, descriptions and arguments. Feedback is the metric's feedback, or :component_feedback's, or "This trajectory got a score of 0.0." (with the row's score) when there is none.

reflection_record_mode: :beam_native, available under execution_profile: :beam_native and its default there, gives each record the example's inputs, the program's outputs, the feedback, the score and the whole trace, and reflects even when there are no records.

Other options

:candidate_selection_strategy controls the parent program sampled for the next reflection. It defaults to pinned GEPA's :pareto policy; under :beam_native it also accepts :current_best, the other released built-ins, or a validated custom Imp.Optimizer.GEPA.CandidateSelector module/struct.

The :callbacks option accepts callback modules or {module, context} tuples implementing any subset of the documented GEPA callback contract. Hooks are synchronous and observational; failures are isolated from optimization.

:component_feedback maps predictor names to strict arity-one callbacks. These callbacks shape reflective minibatches and are part of optimization; invalid names, invalid output, and callback failures stop the run.

:proposal_concurrency (:beam_native) enables first-party GEPA speculative parallel proposals. Contexts are sampled sequentially from one archive and RNG snapshot, expensive proposal phases run concurrently, and all effects are applied by proposal slot. This is separate from ComBee aggregation: it does not combine worker proposals or use map-shuffle-reduce voting.

:module_selector picks which named components each reflective mutation updates: :round_robin (default, one component per mutation in declared program order) and, under :beam_native, :all (every component per mutation), an arity-five function, or a selector module/struct implementing the Imp.Optimizer.GEPA.ModuleSelector contract. Custom selectors receive the engine state, captured trajectories, minibatch scores, candidate index, and candidate map, and must return a non-empty list of the candidate's component names; anything else raises.

:combee (:beam_native) accepts true or its documented keyword options. ComBee duplicates and deterministically shuffles reflection records, reduces floor(sqrt(n)) balanced groups concurrently, and performs one ordered final reduction. :proposal_timeout bounds reflection work and inherits :timeout when omitted; a finite nested ComBee timeout is an additional upper bound.

Summary

Functions

Compiles a program and returns a safe, checksummed parameter artifact.

Compiles a program and returns the optimizer report independently of program metadata support.

Functions

compile_with_artifact(optimizer, program, trainset, devset, opts \\ [])

Compiles a program and returns a safe, checksummed parameter artifact.

This is the durable path for consumer-defined multi-predictor modules. The artifact contains only named predictor signatures, demonstrations, and configuration plus the GEPA report. It never serializes the consumer module, LMs, adapters, callbacks, or other executable runtime state. Reconstruct the trusted program in the deploying application and apply the artifact with Imp.Optimizer.Artifact.apply/4.

compile_with_report(optimizer, program, trainset, devset, opts \\ [])

Compiles a program and returns the optimizer report independently of program metadata support.

new(metric, opts \\ [])