Local, process-supervised TRL reinforcement backend.
The backend is intentionally narrow: it speaks the versioned
Imp.Clients.TRLProtocol to the bundled worker and supports only GRPO.
Constructing the backend does not install Python packages or download a
model. The configured Python environment and model tree must already exist.
The bundled default contract pins Qwen2.5-0.5B, TRL 1.6.0, and one durable MPS LoRA update; the bundled two-step contract exercises durable continuation. Both accept arbitrary Imp-rendered prompt groups and finite external rewards. Ordered groups are source-bound and batched through official TRL. Later steps restore the prior adapter, optimizer, scheduler, Trainer state, and explicit MPS RNG before continuing. Experiment-specific assertions such as a required tensor change belong in an explicit contract; they are not imposed on ordinary training, where uniform-reward groups may truthfully produce no-op steps. Mismatched rollout or step budgets fail before the worker loads the model.
Public GRPO.train_kwargs may override the narrow data-only training keys
learning_rate, beta, loss_type, and scale_rewards. They are validated
before worker startup, sealed into the session identity, and written to a
session-owned runtime contract. Arbitrary Python kwargs are never forwarded.