Modules
Elixir bindings for llama.cpp.
Library-level supervision tree.
Chat template formatting using llama.cpp's Jinja template engine.
OpenAI-compatible chat completion response struct.
OpenAI-compatible streaming chat completion chunk struct.
Inference context with KV cache.
Typed decisions with a decision model: llama.cpp's /v1/systemone API,
in-process.
Generate embeddings from text using an embedding model.
Converts JSON Schema to GBNF grammar for constrained generation.
Download GGUF models from HuggingFace Hub.
Multi-Token Prediction (MTP) speculative decoding.
Model loading and introspection.
Holds multiple models resident and routes requests to them by id.
Behaviour for the model I/O the manager performs on the write path.
Advisory, placement-aware memory budgeting for LlamaCppEx.ModelManager.
A single resident-model record held in the LlamaCppEx.ModelManager ETS table.
Default LlamaCppEx.ModelManager.Backend implementation.
Opt-in supervisor for the multi-model manager.
Single owner for option policy shared across the public entry points.
Remote ggml devices over the llama.cpp RPC backend.
The worker side of the llama.cpp RPC backend: serves this node's devices to a remote client.
Token sampling configuration.
Converts Ecto schema modules to JSON Schema maps for structured output.
GenServer for continuous batched multi-sequence inference.
Behavior for batch building strategies.
Balanced batching strategy.
Shared batch-assembly helpers used by the batching strategies.
Decode-maximal batching strategy.
Prefill-priority batching strategy.
Parser for <think>...</think> blocks in thinking model output.
Text tokenization and detokenization.