Local causal-LM view of the model owned by a running TRLTrainer session.
It exists so Imp's ordinary program rollouts and the subsequent TRL update use the same pinned policy process. It cannot start a worker or download a model on its own. Training LMs sample by default; artifact deployments are constructed in greedy mode. A capable single-field enum adapter may bind an exact finite choice set during greedy deployment, which the causal policy scores and selects using its own logits rather than post-hoc output repair. Sampled training rejects that constraint because a choice-normalized policy requires a different objective than TRL's ordinary token-policy GRPO.
Summary
Types
@type t() :: %Imp.Clients.TRLLM{ artifact_sha256: String.t() | nil, generation_mode: :sample | :greedy, model: String.t(), response_field: atom(), rollout_source: :model_generated | :controlled_external, timeout: pos_integer(), worker_key: term() }