Imp.Predict.RLM (Imp v0.5.0)

Copy Markdown View Source

Recursive Language Model module.

RLM is not retrieval-augmented generation. It is an inference-time strategy for large or awkward contexts: inputs are exposed as variables in a persistent constrained-Elixir environment, and a controller LM iteratively writes code until that code submits structured output.

This implementation uses a BEAM-safe sandbox for production control. The primary controller response is %{reasoning: "...", code: "..."}. Safe code supports persistent assignment, bounded comprehensions and transformations, llm_query/1, llm_query_batched/1, recurse/2, load/1, registered tools, print/1, and submit/1.

Controller iterations, recursion depth, and optional wall time are bounded separately. A shared max_llm_calls ledger covers only one-shot sub-LM work from llm_query*, depth-limit rlm_query* fallbacks, and sub-LM calls made inside recursive children. Root and child controller turns, extraction, and compaction generations are not charged. Generated source is parsed but never evaluated by Code.eval_*; only an explicit AST allowlist executes, with atom-safe parsing and an interpreter step budget.

Tool execution is policy-gated. Denied, crashing, or policy-crashing registered-tool effects return {:error, {:rlm_tool_error, reason}} to the interpreter, which records the redacted failure and permits controller repair.

Summary

Functions

Runs the RLM loop.

Closes the optional persistent RLM environment.

Creates an RLM controller loop.

Creates a lazy value handle that an RLM controller can load explicitly.

Functions

call(rlm, inputs)

Runs the RLM loop.

The controller LM returns reasoning and constrained Elixir code. A successful submit/1 call returns a Imp.Prediction with :rlm_trace metadata.

close(rlm)

Closes the optional persistent RLM environment.

new(signature, opts \\ [])

Creates an RLM controller loop.

Options:

  • :lm - controller LM.
  • :sub_lm - LM used for llm_query calls; defaults to :lm.
  • :tools - list of Imp.Tool values available to tool calls.
  • :tool_policy - which tools the model may call; see Imp.ToolPolicy.
  • :max_iterations - maximum controller turns, as DSPy's dspy.RLM names it.
  • :max_llm_calls - shared limit for one-shot sub-LM calls. This includes llm_query*, depth-limit rlm_query* fallbacks, and sub-LM work inside recursive children; it excludes controller, extraction, and compaction calls.
  • :max_time_ms - optional deadline for the complete RLM call. When omitted, RLM effects have no configured deadline.
  • :max_recursion_depth - how many levels of child RLMs rlm_query* and recurse/2 may start below the root; defaults to 1, one level. At the limit rlm_query* falls back to a one-shot sub-LM query and recurse/2 fails.
  • :max_interpreter_steps - AST execution steps per controller turn.
  • :max_interpreter_value_bytes - maximum serialized size of an interpreter value.
  • :max_interpreter_effects - external effects allowed per controller turn.
  • :max_preview_chars - characters of each variable's printed value the controller sees each turn.
  • :max_observation_chars - truncation limit for string observations.
  • :compaction / :compaction_threshold_pct - summarize root history at a model-context fraction.
  • :compaction_context_tokens - context limit paired with the explicit chars/4 token-estimation fallback; Imp.LM currently exposes no standard tokenizer/context metadata.
  • :persistent - retain the constrained namespace across calls; release it with close/1.

sandbox_serializable(name, loader, opts \\ [])

Creates a lazy value handle that an RLM controller can load explicitly.