Imp 03: Evaluate and optimize

Copy Markdown View Source
imp_checkout? = fn path ->
  is_binary(path) and File.regular?(Path.join(path, "mix.exs")) and
    File.regular?(Path.join(path, "lib/imp.ex"))
end

explicit_repo = System.get_env("IMP_PATH")

if explicit_repo && not imp_checkout?.(Path.expand(explicit_repo)) do
  raise "IMP_PATH does not point to an Imp source checkout or unpacked package"
end

repo =
  [explicit_repo, Path.expand("..", __DIR__), File.cwd!()]
  |> Enum.reject(&is_nil/1)
  |> Enum.map(&Path.expand/1)
  |> Enum.find(imp_checkout?)

if repo do
  # Prefer the notebook's own Imp checkout even when Livebook was launched
  # from an unrelated Mix project. Unpacked package archives may omit
  # mix.lock, so pin it only when the source checkout actually provides it.
  install_opts =
    if File.regular?(Path.join(repo, "mix.lock")),
      do: [lockfile: Path.join(repo, "mix.lock")],
      else: []

  Mix.install([{:imp, path: repo}], install_opts)
else
  # Standalone notebook: install the released package from Hex.
  Mix.install([{:imp, "~> 0.5"}])
end

# Live cells run when LIVE_PROVIDER=1 and a key is set: OPENAI_API_KEY, or
# OPENROUTER_API_KEY for the same model through OpenRouter. In your own code,
# any ReqLLM model string works.
live_lm = fn opts ->
  model = System.get_env("OPENAI_MODEL", "gpt-5.4-mini")

  cond do
    System.get_env("LIVE_PROVIDER") != "1" ->
      {:skip, "Set LIVE_PROVIDER=1 and OPENAI_API_KEY to run this cell."}

    key = System.get_env("OPENAI_API_KEY") ->
      {:ok, Imp.req_llm("openai:" <> model, Keyword.put(opts, :api_key, key))}

    key = System.get_env("OPENROUTER_API_KEY") ->
      {:ok, Imp.req_llm("openrouter:openai/" <> model, Keyword.put(opts, :api_key, key))}

    true ->
      {:skip, "Set OPENAI_API_KEY to run this cell."}
  end
end

Build train and dev sets

Livebook 02 showed what a program sends. This notebook gives a program a scorecard: optimizers are only useful after examples and a metric say what better means.

Optimizers need two kinds of examples. The train set is material they can use to build candidates. The dev set is the held-out scorecard used to choose between those candidates.

trainset = [
  Imp.example(question: "2+2?", answer: "4") |> Imp.with_inputs(:question),
  Imp.example(question: "3+3?", answer: "6") |> Imp.with_inputs(:question)
]

devset = [
  Imp.example(question: "2 plus 2?", answer: "4") |> Imp.with_inputs(:question)
]

Define a metric

The metric is the telos of an optimization run. Imp can search demos, instructions, or artifacts, but it can only improve what the metric can see.

metric = Imp.exact_match(:answer)

Evaluate a baseline

Start with a baseline before optimizing. A baseline tells you whether the metric, dev set, adapter, and LM are wired together before search adds motion.

lm =
  Imp.LM.Static.new(
    handler: fn messages, _opts ->
      prompt = Enum.map_join(messages, "\n", & &1.content)
      if prompt =~ "answer: 4" or prompt =~ "Always answer 4",
        do: %{answer: "4"},
        else: %{answer: "unknown"}
    end
  )

program = Imp.predict("question -> answer", lm: lm)
Imp.evaluate(program, devset, metric)

Labeled few-shot

The simplest improvement is to attach known-good examples as demos. This is not magic training; it is ordinary data being rendered by the adapter.

compiled =
  Imp.Optimizer.LabeledFewShot.new(k: 1)
  |> then(&Imp.optimize!(program, &1, trainset))

Imp.evaluate(compiled, devset, metric)

Random search tries several demo subsets and keeps the candidate that scores best. Its value is not sophistication; its value is that it creates an inspectable optimization report.

compiled =
  metric
  |> Imp.Optimizer.BootstrapFewShotWithRandomSearch.new(num_candidate_programs: 3, max_bootstrapped_demos: 1)
  |> then(&Imp.optimize!(program, &1, trainset, devset))

{
  Imp.evaluate(compiled, devset, metric),
  Imp.Optimizer.Report.fetch(compiled)
}

Instruction search changes the program instructions and keeps the candidate that scores best against the dev set.

Use this when the examples are fine but the task wording is the bottleneck.

compiled =
  Imp.Optimizer.InstructionSearch.compile(
    program,
    metric,
    trainset,
    devset,
    ["Answer unknown.", "Always answer 4."]
  )

Imp.Optimizer.Report.fetch(compiled)

Optimize anything

Use artifact optimization when the thing you want to improve is not an Imp program yet: a config file, policy text, rubric, template, or other named artifact.

This is the broader Imp philosophy in miniature: define an artifact, define how to score it, then let the system propose and evaluate changes.

result =
  Imp.Optimize.Anything.run(
    "mode=slow",
    fn candidate ->
      cond do
        candidate =~ "mode=fast" and candidate =~ "timeout=5" -> 1.0
        candidate =~ "mode=fast" -> 0.5
        true -> 0.0
      end
    end,
    config: [
        engine: [max_candidate_proposals: 2, parallel: false],
        reflection: [
          custom_candidate_proposer: fn candidate, component, _records, _iteration ->
            current = Map.fetch!(candidate, component)

            if current =~ "mode=fast",
              do: current <> "\ntimeout=5",
              else: current <> "\nmode=fast"
          end
        ]
    ]
  )

IO.inspect(result, label: "Optimize Anything result")

Evaluate a live provider

This cell spends a few provider calls only when LIVE_PROVIDER=1 and provider credentials are present. It proves that the same evaluate -> optimize -> report path works with a real LM, while keeping the dataset intentionally tiny.

case live_lm.(max_tokens: 100) do
  {:ok, lm} ->
    live_program =
      Imp.predict(
        Imp.signature(
          "question -> answer: string",
          "Return JSON only. The answer field must be exactly the requested numeral."
        ),
        lm: lm,
        adapter: Imp.Adapter.JSON,
        config: [json_retries: 1]
      )

    live_trainset = [
      Imp.example(question: "Return the answer exactly 4.", answer: "4")
      |> Imp.with_inputs(:question)
    ]

    live_devset = [
      Imp.example(question: "Return the answer exactly 4.", answer: "4")
      |> Imp.with_inputs(:question)
    ]

    live_metric = fn _example, prediction ->
      prediction
      |> Imp.get(:answer, "")
      |> to_string()
      |> String.trim()
      |> Kernel.==("4")
    end

    baseline = Imp.evaluate(live_program, live_devset, live_metric)

    unless baseline.score == 1.0 do
      raise "live provider returned an invalid evaluation: #{inspect(baseline)}"
    end

    compiled =
      Imp.optimize!(
        live_program,
        Imp.Optimizer.BootstrapFewShotWithRandomSearch.new(live_metric, num_candidate_programs: 2, max_bootstrapped_demos: 1),
        live_trainset,
        live_devset
      )

    %{
      baseline_score: baseline.score,
      compiled_score: Imp.evaluate(compiled, live_devset, live_metric).score,
      optimizer_report: Imp.Optimizer.Report.fetch(compiled)
    }

  skip ->
    skip
end

GEPA-style reflection

GEPA-style optimization keeps per-example diagnostics, reflects on misses, and uses Pareto pressure so an improvement for one case does not erase performance on another.

The key idea is not "make a bigger prompt." The key idea is to preserve useful diagnostics from failures and use them as material for the next candidate.

result =
  Imp.Optimize.Anything.run(
    "Base",
    fn candidate, requirement ->
      if String.contains?(candidate, requirement), do: 1.0, else: {0.0, %{feedback: requirement}}
    end,
    dataset: ["Paris", "concise"],
    valset: ["Paris"],
    config: [
        engine: [max_candidate_proposals: 1, parallel: false],
        reflection: [
          custom_candidate_proposer: fn _candidate, _component, _records, _iteration ->
            "Paris\nconcise"
          end
        ]
    ]
  )

IO.inspect(result, label: "GEPA-style reflection result")

Next: open livebooks/04_tools_agents_mcp_rlm.livemd when the program needs explicit tools, external catalogs, or bounded recursive exploration.