LlamaCppEx.Decision (LlamaCppEx v0.8.55)

Copy Markdown View Source

Typed decisions with a decision model: llama.cpp's /v1/systemone API, in-process.

A decision model answers typed questions about a state in one forward pass per prompt; no token is generated. Each answer is a probability distribution, so it comes with a confidence instead of free text. This is llama-server's TypeSafe-compatible System One API, ported into the NIF: same request shape, same prompts, same answers.

Models

The model file says what kind of decision model it is ("<arch>.decision.type"); model_type/1 reads it. Supported types:

  • :openjev, :lev, :nimble - read the logits of one label token per option (:nimble lists every question of the request in each prompt, :lev asks each choice twice with the options reversed).
  • :kev - scores each option by a dot product of hidden states.
  • :laya - a ModernBERT encoder; one score per [MASK] marker.
  • :clef - reads every question in one prompt and decides them jointly.

Decision GGUFs are published in ggml-org's "Decision models" collection on Hugging Face, for example ggml-org/Laya-GGUF, ggml-org/Clef-Flash-GGUF and ggml-org/OpenJev-GGUF.

Questions

questions maps a question id to a question. Each question has:

  • :type - :choice, :score or :noul (strings work too).
  • :instructions - the question. A string, or any JSON-encodable term, which the model is given as JSON text.
  • :criteria - depends on :type:
    • :choice - the options, mapping each option key to its description (nil for none). The number of options is capped by the model: 52 for openjev, 255 for laya and clef.
    • :score - a list of 2 to 10 level descriptions, lowest first.
    • :noul - optional; %{"true" => ..., "false" => ...} descriptions.

The questions of a request are answered independently, except with clef, which decides them jointly.

Order matters to the model: the options of a choice are shown in the order given, and with clef and nimble so are the questions. A map is taken in Elixir's map order (sorted keys for up to 32 entries); pass a keyword list or a list of {key, value} tuples to fix the order yourself. Ids and option keys come back exactly as given, atoms included.

State

A string, or any JSON-encodable term (a map, a list), which the model is given as JSON text. A list of chat messages, or a map with a "messages" list, works the same way.

Answers

  • :choice - %{type: :choice, choice: key, probabilities: %{key => p}, confidence: c}
  • :score - %{type: :score, score: s, legend: %{0 => desc, ...}, probabilities: %{0 => p, ...}, confidence: c}, where score is the probability-weighted level index and can fall between two levels.
  • :noul - %{type: :noul, noul: p}, the probability that the answer is true.

confidence runs from 0 (every option equally likely) to 1. Probabilities are scaled with the temperatures stored in the model file; they are not guaranteed to be calibrated for your data.

Example

{:ok, model} = LlamaCppEx.load_model("Laya-Q8_0.gguf", n_gpu_layers: -1)
{:ok, decision} = LlamaCppEx.Decision.new(model)

{:ok, %{answers: answers}} =
  LlamaCppEx.Decision.decide(
    decision,
    "Customer message: I was charged twice for my order last week.",
    route: [
      type: :choice,
      instructions: "Which team should handle this?",
      criteria: [billing: "payments and refunds", shipping: nil, technical: nil]
    ],
    angry: [type: :noul, instructions: "Is the customer angry?"],
    urgency: [
      type: :score,
      instructions: "How urgent is this?",
      criteria: ["can wait", "this week", "today", "right now"]
    ]
  )

answers.route.choice
#=> :billing (0.992 with Laya-Q8_0)

Limits

  • Text only. Upstream also takes images for openjev and clef through libmtmd, which this build does not link; a request with images is rejected.
  • laya and clef evaluate the whole prompt in one batch, so it has to fit in :n_batch (2048 tokens by default, see new/2).
  • A %Decision{} drives one context and is not safe to use from two processes at once, like LlamaCppEx.Context. Give each process its own, or serialize calls through one process.

Summary

Functions

Answers questions about state.

The decision type of a model.

Creates a decision engine on a new context for model.

Types

answer()

@type answer() ::
  %{
    type: :choice,
    choice: id(),
    probabilities: %{optional(id()) => float()},
    confidence: float()
  }
  | %{
      type: :score,
      score: float(),
      legend: %{optional(non_neg_integer()) => term()},
      probabilities: %{optional(non_neg_integer()) => float()},
      confidence: float()
    }
  | %{type: :noul, noul: float()}

decision_type()

@type decision_type() :: :openjev | :lev | :kev | :nimble | :laya | :clef

id()

@type id() :: atom() | String.t()

question()

@type question() ::
  %{optional(atom() | String.t()) => term()} | [{atom() | String.t(), term()}]

questions()

@type questions() :: %{optional(id()) => question()} | [{id(), question()}]

result()

@type result() :: %{
  answers: %{optional(id()) => answer()},
  usage: %{input_tokens: non_neg_integer(), output_tokens: 0}
}

t()

@type t() :: %LlamaCppEx.Decision{
  context: LlamaCppEx.Context.t(),
  ref: reference(),
  type: decision_type()
}

Functions

decide(decision, state, questions)

@spec decide(t(), term(), questions()) :: {:ok, result()} | {:error, String.t()}

Answers questions about state.

See the module doc for the shape of state, questions and the answers. Returns {:error, reason} for a request the model rejects (a missing :instructions, an unknown :type, too many options, a prompt that does not fit the context) with the same messages as llama-server.

Raises ArgumentError if two question ids, or two option keys of one question, turn into the same string.

model_type(model)

@spec model_type(LlamaCppEx.Model.t()) :: decision_type() | :unknown | nil

The decision type of a model.

Returns nil for a model that is not a decision model and :unknown for a decision type this build does not support.

new(model, opts \\ [])

@spec new(
  LlamaCppEx.Model.t(),
  keyword()
) :: {:ok, t()} | {:error, String.t()}

Creates a decision engine on a new context for model.

The context is shaped for the model's decision type: laya, kev and clef read the embeddings output, so their context has embeddings on and no pooling; the types that share a prompt prefix across questions (openjev, lev, kev, nimble) get a second sequence to evaluate that prefix once per request.

Options

  • :n_ctx - Context size. Defaults to 4096. A prompt holds the state, one question and its options (all questions for clef and nimble), so a long state needs a larger context.
  • :n_batch - Max tokens per batch. For laya, kev and clef it is also the micro-batch size and caps the prompt that can be evaluated at once; defaults to min(n_ctx, 2048) for them and to LlamaCppEx.Context's default otherwise.

Any LlamaCppEx.Context.tuning_option_keys/0 option (:n_threads, :flash_attn, ...) is passed through to LlamaCppEx.Context.create/2.