Imp.Optimizer.GRPO (Imp v0.5.0)

Copy Markdown View Source

Provider-neutral, iterative module-level mmGRPO compilation.

Groups align calls by {predictor, relative invocation} across rollouts and propagate the program-level reward to each aligned completion. A trainer owns only the reinforcement session lifecycle and model artifact. Optional durable session checkpoints reconcile accepted dispatches after caller crashes and resume completed optimizer steps. Provider callbacks execute under explicit deadlines in isolated unlinked tasks, so callback implementations must not rely on the caller's process dictionary or mailbox.

Pinned DSPy 3.2.1 rejects distinct student LMs even though its unfinished job routing is keyed by LM. Imp makes that lifecycle explicit: programs with distinct student LMs use one independent trainer session and artifact per LM. Predictor groups are routed only to the session that owns those predictors, and only those predictors are rebound to that session's current or final artifact. Student groups run in stable predictor order; a later group therefore observes earlier completed groups in the program. This deliberately avoids pretending that one trained artifact represents every predictor. A durable multi-student checkpoint records completed artifacts plus the active child-session checkpoint, so resume never retrains a completed student or guesses an unknown mutating outcome.

Durable jobs require Imp.Optimizer.GRPO.Callback values for reward and validation logic. They bind a trusted module/function to a consumer-owned versioned id and JSON-safe configuration digest. Bare functions remain available only when checkpoint_path is absent; Imp refuses them before trainer activity rather than persisting compiler-local function identity as a false restart contract.

timeout bounds each per-example rollout and validation evaluation (default 5000ms) and is a BEAM-native execution option: pass a larger value or :infinity when rollouts are slow — agentic or environment-backed programs routinely run for minutes and would otherwise be killed at the 5s default.

checkpoint_selection: :best_validation predeclares validation-only selection across trained checkpoints. Validation is greedy for the bundled TRL LM, scores are finite scalars, higher is better, and the earliest checkpoint wins ties. The final trainer state remains retained for trajectory integrity even when an earlier content-verified artifact is deployed. The base model and untouched test data are never candidates in this selector.

Summary

Functions

new(reward_fn, opts \\ [])