Imp.Metrics (Imp v0.5.0)

Copy Markdown View Source

Metrics for evaluation, selection, and optimization.

In Imp, a metric is just an Elixir function that scores a program output against an example. Evaluation and optimizers accept a few convenient return shapes: booleans, numbers, maps with :score and :feedback, predictions that carry score fields, or %Imp.Metrics.Result{} values. Imp normalizes those shapes before it computes averages, chooses best-of-N candidates, or feeds optimizer reports.

Use the small built-ins for deterministic local tasks and write ordinary functions when the task needs domain judgment.

A metric takes (example, prediction) or (example, prediction, trace). The trace is nil when a program is evaluated (Imp.evaluate/4, and the validation scoring optimizers do through it), and the program's trace when an optimizer bootstraps demos from it, as DSPy's trace=None switch has it: a metric can score continuously for evaluation and pass or fail a demo while compiling.

Example

iex> metric = Imp.Metrics.exact_match(:answer)
iex> example = Imp.example(question: "Capital?", answer: "Paris")
iex> prediction = Imp.prediction(answer: "paris")
iex> metric.(example, prediction)
true

iex> result = Imp.Metrics.normalize_result(%{score: 0.75, feedback: "partial"})
iex> {result.score, result.passed?, result.feedback}
{0.75, true, "partial"}

Summary

Functions

Builds a metric that passes when an answer appears in a predicted context field.

Classifies a normalized answer as "yes_no", "numeric", "short_span", or "long_span".

Returns a structured single-row classification result.

Computes accuracy and macro/micro/weighted F1 for classification rows.

DPR normalization (dspy/dsp/utils/dpr.py DPR_normalize): NFD, tokenize with the DPR SimpleTokenizer pattern, lowercase each token.

Returns exact-match truth after normalize_text/1.

Builds an evaluator metric that compares one prediction field to an example field.

Returns a structured extractive-QA metric result.

Computes token F1 after normalize_text/1.

Returns feedback attached to a normalized metric return value.

DPR has_answer: whether any tokenized answer occurs as a contiguous token subsequence of the DPR-normalized text.

Computes HotPotQA-style F1 mirroring DSPy's hotpot_f1_score/HotPotF1.

Converts a metric return value into %Imp.Metrics.Result{}.

Normalizes answer text exactly like DSPy's dspy.evaluate.metrics.normalize_text.

Returns whether a metric return value passes after normalization.

Returns whether any passage contains any answer, mirroring DSPy's _passage_match (dspy/evaluate/metrics.py) over DPR has_answer (dspy/dsp/utils/dpr.py).

Computes recall of expected evidence ids in retrieved documents.

Returns the normalized numeric score for a metric return value.

Describes how a predicted span relates to the expected answer span.

Functions

answer_passage_match(answer_field \\ :answer, context_field \\ :context)

Builds a metric that passes when an answer appears in a predicted context field.

Mirrors DSPy's answer_passage_match: each passage is checked separately (a string context counts as a single passage) and matching uses DPR has_answer token-sequence semantics via passage_match/2, so an answer can never match inside an unrelated word and never across a passage seam.

answer_type(answer)

Classifies a normalized answer as "yes_no", "numeric", "short_span", or "long_span".

classification(prediction, label, opts \\ [])

Returns a structured single-row classification result.

The comparison uses normalize_text/1 so label capitalization and light punctuation differences do not matter. Aggregate classification metrics can be computed with classification_report/2.

classification_report(rows, opts \\ [])

Computes accuracy and macro/micro/weighted F1 for classification rows.

Rows can be {gold, predicted} tuples or maps with :gold/:predicted (or string-keyed equivalents).

dpr_normalize(text)

DPR normalization (dspy/dsp/utils/dpr.py DPR_normalize): NFD, tokenize with the DPR SimpleTokenizer pattern, lowercase each token.

em(prediction, answers)

Returns exact-match truth after normalize_text/1.

When given multiple acceptable answers, any match passes.

exact_match(field \\ :answer)

Builds an evaluator metric that compares one prediction field to an example field.

This is the normal first metric for classification, short-span QA, and beginner optimizer examples.

Mirrors DSPy's answer_exact_match (dspy/evaluate/metrics.py): when the example field holds a LIST of acceptable answers, the prediction passes if it matches ANY member after normalize_text/1.

iex> metric = Imp.Metrics.exact_match(:answer)
iex> example = Imp.example(question: "What is 1+1?", answer: ["2", "two"])
iex> metric.(example, Imp.prediction(answer: "Two"))
true

extractive_qa(prediction, answer, opts \\ [])

Returns a structured extractive-QA metric result.

The result uses exact match as the pass/fail score and includes F1, answer type, and span-relation metadata for dashboards and parity reports.

f1(prediction, answers)

Computes token F1 after normalize_text/1.

Duplicate token overlap is counted, matching extractive QA metric behavior. When given multiple acceptable answers, the best F1 is returned.

feedback(value)

Returns feedback attached to a normalized metric return value.

has_answer(tokenized_answers, text)

DPR has_answer: whether any tokenized answer occurs as a contiguous token subsequence of the DPR-normalized text.

Faithful to DSPy's port of Facebook DPR, including the edge where an empty tokenized answer matches any text.

hotpot_f1(prediction, answers)

Computes HotPotQA-style F1 mirroring DSPy's hotpot_f1_score/HotPotF1.

Identical to f1/2 except that when either normalized side is one of the special HotPotQA labels "yes", "no", or "noanswer" and the sides differ, the score is 0.0 — the same gating the official HotPotQA evaluation script (hotpot_evaluate_v1.py) applies. When given multiple acceptable answers, the best score is returned.

normalize_result(result)

Converts a metric return value into %Imp.Metrics.Result{}.

Accepted values include %Result{}, %Imp.Prediction{}, maps, booleans, and numbers. Unknown values become a failed result with diagnostic feedback.

normalize_text(value)

Normalizes answer text exactly like DSPy's dspy.evaluate.metrics.normalize_text.

The SQuAD-style pipeline, in DSPy's order: Unicode NFD normalization, lowercasing, deletion (not substitution) of Python's string.punctuation characters (ASCII only — Unicode punctuation is kept), word-boundary English article removal, and whitespace collapse. Differential parity with real DSPy 3.2.1 is pinned in test/metrics_dspy_parity_test.exs.

pass?(value)

Returns whether a metric return value passes after normalization.

passage_match(passages, answers)

Returns whether any passage contains any answer, mirroring DSPy's _passage_match (dspy/evaluate/metrics.py) over DPR has_answer (dspy/dsp/utils/dpr.py).

Answers and passages both go through normalize_text/1 and DPR tokenization; an answer matches only as a contiguous token sequence within a single passage.

retrieval_recall(prediction, expected_ids, opts \\ [])

Computes recall of expected evidence ids in retrieved documents.

prediction may be a %Imp.Prediction{} with RAG retrieval metadata or a list of retrieved document maps. Expected ids can be strings or atoms.

score(value)

Returns the normalized numeric score for a metric return value.

span_relation(prediction, answer)

Describes how a predicted span relates to the expected answer span.