An evaluation run containing metrics and per-case results.
Stores the configuration used, aggregate metrics, and detailed results for each test case to enable drill-down into failures.