Typed decisions with a decision model: llama.cpp's /v1/systemone API,
in-process.
A decision model answers typed questions about a state in one forward pass
per prompt; no token is generated. Each answer is a probability distribution,
so it comes with a confidence instead of free text. This is llama-server's
TypeSafe-compatible System One API,
ported into the NIF: same request shape, same prompts, same answers.
Models
The model file says what kind of decision model it is
("<arch>.decision.type"); model_type/1 reads it. Supported types:
:openjev,:lev,:nimble- read the logits of one label token per option (:nimblelists every question of the request in each prompt,:levasks each choice twice with the options reversed).:kev- scores each option by a dot product of hidden states.:laya- a ModernBERT encoder; one score per[MASK]marker.:clef- reads every question in one prompt and decides them jointly.
Decision GGUFs are published in ggml-org's "Decision models" collection on
Hugging Face, for example ggml-org/Laya-GGUF, ggml-org/Clef-Flash-GGUF and
ggml-org/OpenJev-GGUF.
Questions
questions maps a question id to a question. Each question has:
:type-:choice,:scoreor:noul(strings work too).:instructions- the question. A string, or any JSON-encodable term, which the model is given as JSON text.:criteria- depends on:type::choice- the options, mapping each option key to its description (nilfor none). The number of options is capped by the model: 52 for openjev, 255 for laya and clef.:score- a list of 2 to 10 level descriptions, lowest first.:noul- optional;%{"true" => ..., "false" => ...}descriptions.
The questions of a request are answered independently, except with clef, which decides them jointly.
Order matters to the model: the options of a choice are shown in the order
given, and with clef and nimble so are the questions. A map is taken in
Elixir's map order (sorted keys for up to 32 entries); pass a keyword list or
a list of {key, value} tuples to fix the order yourself. Ids and option keys
come back exactly as given, atoms included.
State
A string, or any JSON-encodable term (a map, a list), which the model is given
as JSON text. A list of chat messages, or a map with a "messages" list,
works the same way.
Answers
:choice-%{type: :choice, choice: key, probabilities: %{key => p}, confidence: c}:score-%{type: :score, score: s, legend: %{0 => desc, ...}, probabilities: %{0 => p, ...}, confidence: c}, wherescoreis the probability-weighted level index and can fall between two levels.:noul-%{type: :noul, noul: p}, the probability that the answer is true.
confidence runs from 0 (every option equally likely) to 1. Probabilities
are scaled with the temperatures stored in the model file; they are not
guaranteed to be calibrated for your data.
Example
{:ok, model} = LlamaCppEx.load_model("Laya-Q8_0.gguf", n_gpu_layers: -1)
{:ok, decision} = LlamaCppEx.Decision.new(model)
{:ok, %{answers: answers}} =
LlamaCppEx.Decision.decide(
decision,
"Customer message: I was charged twice for my order last week.",
route: [
type: :choice,
instructions: "Which team should handle this?",
criteria: [billing: "payments and refunds", shipping: nil, technical: nil]
],
angry: [type: :noul, instructions: "Is the customer angry?"],
urgency: [
type: :score,
instructions: "How urgent is this?",
criteria: ["can wait", "this week", "today", "right now"]
]
)
answers.route.choice
#=> :billing (0.992 with Laya-Q8_0)Limits
- Text only. Upstream also takes images for openjev and clef through libmtmd, which this build does not link; a request with images is rejected.
- laya and clef evaluate the whole prompt in one batch, so it has to fit in
:n_batch(2048 tokens by default, seenew/2). - A
%Decision{}drives one context and is not safe to use from two processes at once, likeLlamaCppEx.Context. Give each process its own, or serialize calls through one process.
Summary
Functions
Answers questions about state.
The decision type of a model.
Creates a decision engine on a new context for model.
Types
@type answer() :: %{ type: :choice, choice: id(), probabilities: %{optional(id()) => float()}, confidence: float() } | %{ type: :score, score: float(), legend: %{optional(non_neg_integer()) => term()}, probabilities: %{optional(non_neg_integer()) => float()}, confidence: float() } | %{type: :noul, noul: float()}
@type decision_type() :: :openjev | :lev | :kev | :nimble | :laya | :clef
@type result() :: %{ answers: %{optional(id()) => answer()}, usage: %{input_tokens: non_neg_integer(), output_tokens: 0} }
@type t() :: %LlamaCppEx.Decision{ context: LlamaCppEx.Context.t(), ref: reference(), type: decision_type() }
Functions
Answers questions about state.
See the module doc for the shape of state, questions and the answers.
Returns {:error, reason} for a request the model rejects (a missing
:instructions, an unknown :type, too many options, a prompt that does not
fit the context) with the same messages as llama-server.
Raises ArgumentError if two question ids, or two option keys of one
question, turn into the same string.
@spec model_type(LlamaCppEx.Model.t()) :: decision_type() | :unknown | nil
The decision type of a model.
Returns nil for a model that is not a decision model and :unknown for a
decision type this build does not support.
@spec new( LlamaCppEx.Model.t(), keyword() ) :: {:ok, t()} | {:error, String.t()}
Creates a decision engine on a new context for model.
The context is shaped for the model's decision type: laya, kev and clef read the embeddings output, so their context has embeddings on and no pooling; the types that share a prompt prefix across questions (openjev, lev, kev, nimble) get a second sequence to evaluate that prefix once per request.
Options
:n_ctx- Context size. Defaults to4096. A prompt holds the state, one question and its options (all questions for clef and nimble), so a long state needs a larger context.:n_batch- Max tokens per batch. For laya, kev and clef it is also the micro-batch size and caps the prompt that can be evaluated at once; defaults tomin(n_ctx, 2048)for them and toLlamaCppEx.Context's default otherwise.
Any LlamaCppEx.Context.tuning_option_keys/0 option (:n_threads,
:flash_attn, ...) is passed through to LlamaCppEx.Context.create/2.