Metrics for evaluation, selection, and optimization.
In Imp, a metric is just an Elixir function that scores a program output
against an example. Evaluation and optimizers accept a few convenient return
shapes: booleans, numbers, maps with :score and :feedback, predictions
that carry score fields, or %Imp.Metrics.Result{} values. Imp normalizes
those shapes before it computes averages, chooses best-of-N candidates, or
feeds optimizer reports.
Use the small built-ins for deterministic local tasks and write ordinary functions when the task needs domain judgment.
A metric takes (example, prediction) or (example, prediction, trace).
The trace is nil when a program is evaluated (Imp.evaluate/4, and the
validation scoring optimizers do through it), and the program's trace when an
optimizer bootstraps demos from it, as DSPy's trace=None switch has it: a
metric can score continuously for evaluation and pass or fail a demo while
compiling.
Example
iex> metric = Imp.Metrics.exact_match(:answer)
iex> example = Imp.example(question: "Capital?", answer: "Paris")
iex> prediction = Imp.prediction(answer: "paris")
iex> metric.(example, prediction)
true
iex> result = Imp.Metrics.normalize_result(%{score: 0.75, feedback: "partial"})
iex> {result.score, result.passed?, result.feedback}
{0.75, true, "partial"}
Summary
Functions
Builds a metric that passes when an answer appears in a predicted context field.
Classifies a normalized answer as "yes_no", "numeric", "short_span", or "long_span".
Returns a structured single-row classification result.
Computes accuracy and macro/micro/weighted F1 for classification rows.
DPR normalization (dspy/dsp/utils/dpr.py DPR_normalize): NFD, tokenize
with the DPR SimpleTokenizer pattern, lowercase each token.
Returns exact-match truth after normalize_text/1.
Builds an evaluator metric that compares one prediction field to an example field.
Returns a structured extractive-QA metric result.
Computes token F1 after normalize_text/1.
Returns feedback attached to a normalized metric return value.
DPR has_answer: whether any tokenized answer occurs as a contiguous
token subsequence of the DPR-normalized text.
Computes HotPotQA-style F1 mirroring DSPy's hotpot_f1_score/HotPotF1.
Converts a metric return value into %Imp.Metrics.Result{}.
Normalizes answer text exactly like DSPy's dspy.evaluate.metrics.normalize_text.
Returns whether a metric return value passes after normalization.
Returns whether any passage contains any answer, mirroring DSPy's
_passage_match (dspy/evaluate/metrics.py) over DPR has_answer
(dspy/dsp/utils/dpr.py).
Computes recall of expected evidence ids in retrieved documents.
Returns the normalized numeric score for a metric return value.
Describes how a predicted span relates to the expected answer span.
Functions
Builds a metric that passes when an answer appears in a predicted context field.
Mirrors DSPy's answer_passage_match: each passage is checked separately
(a string context counts as a single passage) and matching uses DPR
has_answer token-sequence semantics via passage_match/2, so an answer
can never match inside an unrelated word and never across a passage seam.
Classifies a normalized answer as "yes_no", "numeric", "short_span", or "long_span".
Returns a structured single-row classification result.
The comparison uses normalize_text/1 so label capitalization and light
punctuation differences do not matter. Aggregate classification metrics can
be computed with classification_report/2.
Computes accuracy and macro/micro/weighted F1 for classification rows.
Rows can be {gold, predicted} tuples or maps with :gold/:predicted
(or string-keyed equivalents).
DPR normalization (dspy/dsp/utils/dpr.py DPR_normalize): NFD, tokenize
with the DPR SimpleTokenizer pattern, lowercase each token.
Returns exact-match truth after normalize_text/1.
When given multiple acceptable answers, any match passes.
Builds an evaluator metric that compares one prediction field to an example field.
This is the normal first metric for classification, short-span QA, and beginner optimizer examples.
Mirrors DSPy's answer_exact_match (dspy/evaluate/metrics.py): when the
example field holds a LIST of acceptable answers, the prediction passes if
it matches ANY member after normalize_text/1.
iex> metric = Imp.Metrics.exact_match(:answer)
iex> example = Imp.example(question: "What is 1+1?", answer: ["2", "two"])
iex> metric.(example, Imp.prediction(answer: "Two"))
true
Returns a structured extractive-QA metric result.
The result uses exact match as the pass/fail score and includes F1, answer type, and span-relation metadata for dashboards and parity reports.
Computes token F1 after normalize_text/1.
Duplicate token overlap is counted, matching extractive QA metric behavior. When given multiple acceptable answers, the best F1 is returned.
Returns feedback attached to a normalized metric return value.
DPR has_answer: whether any tokenized answer occurs as a contiguous
token subsequence of the DPR-normalized text.
Faithful to DSPy's port of Facebook DPR, including the edge where an empty tokenized answer matches any text.
Computes HotPotQA-style F1 mirroring DSPy's hotpot_f1_score/HotPotF1.
Identical to f1/2 except that when either normalized side is one of the
special HotPotQA labels "yes", "no", or "noanswer" and the sides
differ, the score is 0.0 — the same gating the official HotPotQA evaluation
script (hotpot_evaluate_v1.py) applies. When given multiple acceptable
answers, the best score is returned.
Converts a metric return value into %Imp.Metrics.Result{}.
Accepted values include %Result{}, %Imp.Prediction{}, maps, booleans,
and numbers. Unknown values become a failed result with diagnostic feedback.
Normalizes answer text exactly like DSPy's dspy.evaluate.metrics.normalize_text.
The SQuAD-style pipeline, in DSPy's order: Unicode NFD normalization,
lowercasing, deletion (not substitution) of Python's string.punctuation
characters (ASCII only — Unicode punctuation is kept), word-boundary English
article removal, and whitespace collapse. Differential parity with real
DSPy 3.2.1 is pinned in test/metrics_dspy_parity_test.exs.
Returns whether a metric return value passes after normalization.
Returns whether any passage contains any answer, mirroring DSPy's
_passage_match (dspy/evaluate/metrics.py) over DPR has_answer
(dspy/dsp/utils/dpr.py).
Answers and passages both go through normalize_text/1 and DPR
tokenization; an answer matches only as a contiguous token sequence within
a single passage.
Computes recall of expected evidence ids in retrieved documents.
prediction may be a %Imp.Prediction{} with RAG retrieval metadata or a
list of retrieved document maps. Expected ids can be strings or atoms.
Returns the normalized numeric score for a metric return value.
Describes how a predicted span relates to the expected answer span.