ReqLLM.Providers.LMStudio (ReqLLM v1.26.0)

View Source

Local inference through LM Studio's OpenAI-compatible API.

Supports chat, streaming, tool calling, image inputs, structured output, and embeddings when the selected model supports them. Model identifiers come from LM Studio's /v1/models endpoint; catalog membership is not required.

model = ReqLLM.model!(%{provider: :lmstudio, id: "my-local-model"})
ReqLLM.generate_text(model, "Hello!", reasoning_effort: :none)

The default endpoint is http://127.0.0.1:1234/v1. Override it with base_url on a request or model, or config :req_llm, :lmstudio, base_url: "...".

Authentication is optional. For servers requiring a token, use api_key, config :req_llm, :lmstudio_api_key, or LMSTUDIO_API_KEY, in that order.

Provider options include ttl (JIT-loaded model idle time in seconds), repeat_penalty, and response_format. Context size and model loading are configured in LM Studio. See the LM Studio provider guide for details.

Summary

Functions

Attaches the standard ReqLLM request and response pipeline with optional auth.

Builds a Finch SSE request with the same model, endpoint, and auth as chat.

Extends the OpenAI-compatible body with LM Studio's chat inference controls.

Default implementation of decode_response/1.

Default implementation of decode_stream_event/2.

Returns the provider name used in diagnostics.

Default implementation of encode_body/1.

Default implementation of extract_usage/2.

Prepares chat, structured-output, or embedding requests.

Normalizes reasoning effort while retaining canonical atoms for validation.

Functions

attach(request, model_input, user_opts)

Attaches the standard ReqLLM request and response pipeline with optional auth.

A configured LM Studio key becomes a bearer token; an absent key adds no auth. Provider model identifiers are preserved on the wire.

attach_stream(model, context, opts, finch_name)

Builds a Finch SSE request with the same model, endpoint, and auth as chat.

Uses the provider body encoder to preserve tools, images, structured-output schemas, and translated options. Construction failures return an API error.

base_url()

build_body(request)

Extends the OpenAI-compatible body with LM Studio's chat inference controls.

TTL, repetition penalty, and reasoning effort are emitted for chat and object requests. Embeddings retain the default embedding body format.

decode_response(request_response)

Default implementation of decode_response/1.

Handles success/error responses with standard ReqLLM.Response creation.

decode_stream_event(event, model)

Default implementation of decode_stream_event/2.

Decodes SSE events using OpenAI-compatible format.

default_base_url()

default_env_key()

Callback implementation for ReqLLM.Provider.default_env_key/0.

display_name()

Returns the provider name used in diagnostics.

encode_body(request)

Default implementation of encode_body/1.

Encodes request body using OpenAI-compatible format for chat and embedding operations.

extract_usage(body, model)

Default implementation of extract_usage/2.

Extracts usage data from standard usage field in response body.

prepare_request(operation, model_spec, input, opts)

Prepares chat, structured-output, or embedding requests.

Explicit model maps are normalized without requiring catalog membership. Structured output uses the compiled schema as a native JSON response format. Unsupported operations return a structured parameter error.

provider_extended_generation_schema()

provider_id()

provider_schema()

supported_provider_options()

translate_options(operation, model, opts)

Normalizes reasoning effort while retaining canonical atoms for validation.

The default effort is omitted and :max is clamped to :xhigh with a warning. Unsupported values are omitted with a warning. Reprocessing is idempotent; conversion to the wire string occurs only when building the request body.