We have a router, a metric, and a baseline. An optimizer takes a program, training examples, and a metric, and returns a new program that does better on them. We don't edit the prompt; we give the optimizer data and let it make the changes.
Showing the model solved tickets
The simplest optimizer, LabeledFewShot, attaches labeled training examples to
the program as demonstrations. The model sees solved tickets before it sees
ours:
improved = Imp.optimize!(router, Imp.Optimizer.LabeledFewShot.new(k: 8), trainset)
Imp.evaluate(improved, testset, metric, num_threads: 8).score
#=> 0.75The router went from 0.25 to 0.75 on tickets it never saw. In six more runs, each in a fresh VM, it went from 0.2–0.45 to 0.75–0.85. Optimizing made no model calls; the only cost is a longer prompt.
improved is a new value. router is unchanged, so we can compare the two, or
keep both.
What changed
The improvement is data we can read. The optimizer sampled eight training tickets, with a fixed seed, so it picks the same eight every time:
for demo <- improved.demos, do: {Imp.get(demo, :ticket), Imp.get(demo, :team)}{"Our VAT number is wrong on the latest invoice.", "atlas"}
{"Card payments have been timing out at checkout since 9am.", "harbor"}
{"Two-factor codes are not being accepted.", "beacon"}
{"I cancelled in April but was still charged in May.", "atlas"}
{"How do I switch from monthly to annual billing?", "atlas"}
{"Does the API support filtering by created date?", "quill"}
{"Receipt emails have not been sent for any order since the deploy.", "harbor"}
{"Webhooks stopped being delivered around midnight UTC.", "harbor"}and Imp renders them as earlier turns of the conversation, before our ticket. The first pair:
{:ok, prediction} = Imp.call(improved, %{ticket: "We were charged twice this month."})
[_system, user, assistant | _rest] = prediction.metadata.trace.messages
IO.puts(user.content <> "\n\n" <> assistant.content)[[ ## ticket ## ]]
Our VAT number is wrong on the latest invoice.
{
"team": "atlas"
}Seven more pairs follow, and then our ticket.
Nothing about the model changed, and no prompt was written by hand. The
examples show what our squad names mean, and the model generalizes from them.
Before we ship improved, we can read its demos like any other change to
our code, and they are saved with the program.
Searching for better instructions
Demonstrations are one lever; the instruction is another. Imp.Optimizer.GEPA
runs the program on training tickets, shows a stronger model where it failed,
and has that model write a better instruction. It keeps the candidates that
score best on the development set. A metric can return feedback as well as a
score, to say why a prediction failed, and GEPA passes it along:
strong_lm = Imp.req_llm("openai:gpt-5.4", api_key: System.fetch_env!("OPENAI_API_KEY"))
feedback_metric = fn example, prediction ->
expected = Imp.get(example, :team)
got = Imp.get(prediction, :team)
if got == expected,
do: %{score: 1.0, feedback: "Correct: #{expected}."},
else: %{score: 0.0, feedback: "Wrong: this ticket belongs to #{expected}, not #{got}."}
end
optimizer =
Imp.Optimizer.GEPA.new(feedback_metric, reflection_lm: strong_lm, max_metric_calls: 150, num_threads: 8)
searched = Imp.optimize!(router, optimizer, trainset, devset)
Imp.evaluate(searched, testset, metric, num_threads: 8).score
#=> 0.95GEPA takes the development set as a fourth argument, so the test set stays
unseen. max_metric_calls is the budget: this run took about 70 seconds and
cost about eight cents. The instruction it wrote is part of the program:
IO.puts(searched.signature.instructions)You are given a support ticket as input in this format:
- `ticket`: the text of the support request
Your task is to classify the ticket to the correct owning squad and output only the squad name.
Valid squads and routing rules:
- `beacon`: account access and security issues, including login/authentication problems, MFA/two-factor issues, password resets, account lockouts, and other security-related requests
- `quill`: product feature requests, how-to/product usage questions, data export requests, and issues involving Slack integrations or alerting/alerts integrations
- `harbor`: general product/platform issues not covered by another squad, including dashboard/site performance problems, API errors/outages, email delivery/communications issues, and order or commerce-related communications such as missing receipt emails
- `atlas`: billing and invoicing issues only, including duplicate charges, being charged twice, invoice disputes, and invoice/payment issues
Important classification guidance:
- Infer the underlying issue/topic from the whole ticket, not just keywords
- Prefer the most specific matching squad
- Do not default to `atlas`; use it only for clear billing/invoicing matters
- Tickets about dashboard slowness or intermittent 502/API failures belong to `harbor`
- Tickets asking how to export data to CSV belong to `quill`
- If a ticket does not match `beacon`, `quill`, or `atlas` specifically, route it to `harbor`
Output requirements:
- Return only the squad name
- Do not include explanations
- Do not include labels, markdown, punctuation, or any extra text beyond the exact team valueGEPA read the tickets it got wrong and wrote down what each squad owns, the thing our labels mean and the model was never told. It also made choices worth a second look before we ship: atlas takes billing "only", and anything that matches no squad goes to harbor. We can review those because the optimizer's output is text.
Three runs, each in a fresh VM and routed through OpenRouter to the same models, scored 0.95, 0.9 and 0.95, took 69 to 78 seconds, and cost eight to fourteen cents each. The eight demonstrations scored 0.75 to 0.85 and cost nothing to compile. Here the instruction is the better lever: the model lacked a description of the squads, and GEPA wrote one from its failures. Demonstrations are cheaper and a good first step; the two combine, and Choosing an optimizer compares the rest.
Livebook 03 runs these optimizers in a notebook, offline or with a key.
Next: Saving and loading →