LLMProxy applies authentication, model access, limits, accounting, and tracing at the provider execution boundary. In-process, ReqLLM, SafeRPC, and HTTP callers therefore receive the same controls.
API keys
Create an application key through the storage facade:
{:ok, api_key, raw_key} =
LLMProxy.Storage.create_key("production-worker", %{
allowed_models: ["fast", "coding"],
trace_requests: true
})raw_key is returned once. LLMProxy stores its hash and looks up the record for later requests. Deliver the raw key through your normal secret-management channel.
HTTP clients authenticate with either header:
Authorization: Bearer sk-proxy-...
x-api-key: sk-proxy-...In-process callers pass the raw key, an API-key schema/map, or %LLMProxy.Actor{}.
The configured master key represents operator access. It bypasses ordinary model and quota checks and should be reserved for bootstrap, administration, and trusted service operations.
Model access
Set allowed_models on an API key to constrain it to public catalog names:
%{
allowed_models: ["fast", "coding"]
}Clients never need direct access to upstream provider names or credentials. Catalog aliases are the policy boundary.
Budget limits
Composable limits operate over stored usage windows:
%{
budget_limits: [
LLMProxy.Limit.cost(:day, 10.00),
LLMProxy.Limit.requests(:minute, 60),
LLMProxy.Limit.input_tokens(:four_hours, 100_000),
LLMProxy.Limit.output_tokens(:week, 500_000)
]
}Supported metrics:
:cost_usd:requests:input_tokens:output_tokens:cache_read_tokens:cache_write_tokens
Supported windows:
:minute:hour:four_hours:day:week:month
Stored maps may use equivalent strings such as "cost_usd", "4h", "24h", "7d", and "30d".
Existing fixed quota fields remain supported for four-hour and weekly token/message limits. New integrations should prefer budget_limits when they need several metrics or windows.
Usage and cost
Successful requests record:
- input and output tokens;
- cache read and write tokens;
- estimated USD cost from LLMDB pricing;
- provider and upstream model;
- request duration and time to first token when available;
- public model, tags, metadata, and trace ID;
- the API key and optional message record responsible for the call.
LLMProxy.Response exposes usage and routing information to in-process callers:
response.usage
response.provider_name
response.model
response.trace_id
response.cache_hitHTTP wire responses retain their protocol shape. The trace ID is returned through x-request-id and x-llm-proxy-trace-id headers.
Traces and messages
Set trace_requests: true on an API key to persist request and response bodies for that key. This is useful for debugging and feedback workflows, but it may capture sensitive prompts and model output.
Before enabling body tracing:
- define retention and deletion policy;
- restrict access to trace storage and admin surfaces;
- avoid tracing keys that carry secrets or regulated data unless storage controls permit it;
- understand that metadata and tags may also contain user-defined values.
Message logs and traces share request identifiers so operators can move from aggregate usage to an individual call.
Feedback
Submit feedback by request_id or trace_id through:
POST /feedback
POST /v1/feedbackThe feedback record links to the stored trace and API key. Use it for operator review or downstream evaluation systems; LLMProxy does not bundle a hosted evaluation product.
Telemetry events
LLMProxy emits :telemetry events for routing attempts and circuit-breaker transitions.
Routing event prefixes include:
[:llm_proxy, :routing, :attempt, :start]
[:llm_proxy, :routing, :attempt, :stop]
[:llm_proxy, :routing, :attempt, :exception]
[:llm_proxy, :routing, :stream_attempt, :start]Circuit event prefixes include:
[:llm_proxy, :circuit, :open]
[:llm_proxy, :circuit, :half_open]
[:llm_proxy, :circuit, :closed]
[:llm_proxy, :circuit, :skip]Metadata identifies provider, upstream model, and timeout. Stop and exception events add result-specific measurements or metadata.
OpenTelemetry
Provider execution creates spans named by operation, such as:
llm_proxy.provider.call
llm_proxy.provider.streamSpans include provider and model attributes plus the LLMProxy trace ID. HTTP, Ecto, and Req integrations are instrumented through their OpenTelemetry packages.
Standalone releases enable OTLP export when OTEL_EXPORTER_OTLP_ENDPOINT is set. Without it, trace export is disabled.
Operational admin
The optional Incant integration presents API keys, provider tokens, traces, messages, and an operations dashboard without making Incant a runtime requirement for the gateway.
See Admin Integration for topology and access boundaries.