Imp.Clients.TRLLM (Imp v0.5.0)

Copy Markdown View Source

Local causal-LM view of the model owned by a running TRLTrainer session.

It exists so Imp's ordinary program rollouts and the subsequent TRL update use the same pinned policy process. It cannot start a worker or download a model on its own. Training LMs sample by default; artifact deployments are constructed in greedy mode. A capable single-field enum adapter may bind an exact finite choice set during greedy deployment, which the causal policy scores and selects using its own logits rather than post-hoc output repair. Sampled training rejects that constraint because a choice-normalized policy requires a different objective than TRL's ordinary token-policy GRPO.

Summary

Types

t()

@type t() :: %Imp.Clients.TRLLM{
  artifact_sha256: String.t() | nil,
  generation_mode: :sample | :greedy,
  model: String.t(),
  response_field: atom(),
  rollout_source: :model_generated | :controlled_external,
  timeout: pos_integer(),
  worker_key: term()
}