ReqLLM.Evaluation (ReqLLM v1.26.0)

View Source

Evaluates one text or JSON state against named questions.

This is a model call, not a test run. Providers handle their own request and answer formats. An evaluation model does not need to support chat generation.

Summary

Functions

Evaluates one state against a map of named questions.

Lists model specs that evaluate/4 can call with an installed adapter.

Functions

evaluate(model_spec, state, questions, opts \\ [])

@spec evaluate(ReqLLM.model_input(), String.t() | map() | list(), map(), keyword()) ::
  {:ok, ReqLLM.Response.t()} | {:error, term()}

Evaluates one state against a map of named questions.

questions = %{
  department: %{
    type: :choice,
    instructions: "Which team should handle this?",
    criteria: %{billing: "Billing and refunds", support: "Other requests"}
  },
  urgent: %{type: :boolean, instructions: "Is this urgent?"}
}

{:ok, result} = ReqLLM.evaluate("typesafe:jev-latest", "Please refund me", questions)
result.object["department"]["choice"]
result.object["urgent"]["probability"]

Choice and score answers may include probabilities and confidence. Providers may support more question types. OpenRouter evaluations accept routing preferences such as provider_options: [openrouter_provider: %{zdr: true}]. The response stores named answers in object with string keys. provider_meta.raw_response keeps the original provider data.

evaluate!(model_spec, state, questions, opts \\ [])

@spec evaluate!(ReqLLM.model_input(), String.t() | map() | list(), map(), keyword()) ::
  ReqLLM.Response.t() | no_return()

Same as evaluate/4, but raises on error.

models()

@spec models() :: [String.t()]

Lists model specs that evaluate/4 can call with an installed adapter.

Catalog evaluation metadata can also describe models for providers that ReqLLM does not yet support. Those models are excluded.