Confidence-aware evaluation for enum-constrained JSON classification.
The implementation follows GEPA commit
65df4325e3fb4781cf2ab17dd144d6ce2f7b98fe for joint token logprob,
scoring formulas, feedback buckets, and alternative formatting. Raw
confidence is diagnostic metadata, not a maximized objective. The exposed
:confidence_quality objective is the configured correctness-aware scoring
strategy's value in [0, 1]; incorrect predictions always receive 0.0.
Missing logprobs fail closed by default. fallback: :accuracy must be
selected explicitly to continue with accuracy only. In that mode the
:confidence_quality objective is omitted and unavailability is recorded in
metric metadata. This avoids treating either missing logprobs or confidence
in an incorrect prediction as optimization quality.
Summary
Functions
Evaluates a prediction and returns a normalized metric result map.
Builds source-faithful reflective feedback from correctness and raw confidence.
Types
Functions
@spec evaluate(Imp.Example.t(), Imp.Prediction.t() | nil, keyword()) :: map()
Evaluates a prediction and returns a normalized metric result map.
Builds source-faithful reflective feedback from correctness and raw confidence.