Provider-neutral, iterative module-level mmGRPO compilation.
Groups align calls by {predictor, relative invocation} across rollouts and
propagate the program-level reward to each aligned completion. A trainer owns
only the reinforcement session lifecycle and model artifact. Optional durable
session checkpoints reconcile accepted dispatches after caller crashes and
resume completed optimizer steps. Provider callbacks execute under explicit
deadlines in isolated unlinked tasks, so callback implementations must not
rely on the caller's process dictionary or mailbox.
Pinned DSPy 3.2.1 rejects distinct student LMs even though its unfinished job routing is keyed by LM. Imp makes that lifecycle explicit: programs with distinct student LMs use one independent trainer session and artifact per LM. Predictor groups are routed only to the session that owns those predictors, and only those predictors are rebound to that session's current or final artifact. Student groups run in stable predictor order; a later group therefore observes earlier completed groups in the program. This deliberately avoids pretending that one trained artifact represents every predictor. A durable multi-student checkpoint records completed artifacts plus the active child-session checkpoint, so resume never retrains a completed student or guesses an unknown mutating outcome.
Durable jobs require Imp.Optimizer.GRPO.Callback values for reward and
validation logic. They bind a trusted module/function to a consumer-owned
versioned id and JSON-safe configuration digest. Bare functions remain
available only when checkpoint_path is absent; Imp refuses them before
trainer activity rather than persisting compiler-local function identity as
a false restart contract.
timeout bounds each per-example rollout and validation evaluation (default
5000ms) and is a BEAM-native execution option: pass a larger value or
:infinity when rollouts are slow — agentic or environment-backed programs
routinely run for minutes and would otherwise be killed at the 5s default.
checkpoint_selection: :best_validation predeclares validation-only
selection across trained checkpoints. Validation is greedy for the bundled
TRL LM, scores are finite scalars, higher is better, and the earliest
checkpoint wins ties. The final trainer state remains retained for trajectory
integrity even when an earlier content-verified artifact is deployed. The
base model and untouched test data are never candidates in this selector.