Held-out reliability analysis for raw constrained-label confidence.
Raw token confidence is treated as an uncalibrated score. Histogram fitting estimates probability of correctness from a distinct calibration split. Evaluation IDs and source IDs must be unique within each split; evaluation, source, and optional group identities must be disjoint across splits.
A fitted histogram exposes class balance, bin occupancy, and its complete mapping. All-correct, all-wrong, and single-bin fits are retained for diagnostics but marked non-authoritative and cannot calibrate held-out data.
Summary
Functions
Whether a fit has mixed outcomes and at least two supported bins.
Maps one raw score to held-out empirical probability of correctness.
Compares raw and calibrated Brier score on the same held-out records.
Evaluates an authoritative histogram on identity-disjoint held-out sources.
Fits fixed-width histogram calibration on uniquely identified sources.
Returns a deterministic, JSON-safe description of a fitted histogram.
Computes Brier score, ECE, reliability buckets, abstention, and prompt drift.
Types
Functions
Whether a fit has mixed outcomes and at least two supported bins.
Maps one raw score to held-out empirical probability of correctness.
Compares raw and calibrated Brier score on the same held-out records.
Evaluates an authoritative histogram on identity-disjoint held-out sources.
Fits fixed-width histogram calibration on uniquely identified sources.
Returns a deterministic, JSON-safe description of a fitted histogram.
Computes Brier score, ECE, reliability buckets, abstention, and prompt drift.