ChatModel adapter using the req_llm library as the HTTP/LLM backend.
Provides access to any provider supported by req_llm (Anthropic, OpenAI, Google Gemini, Groq, Ollama, AWS Bedrock, etc.) through the unified LangChain framework.
Model Specification
The model field takes a req_llm-format specifier string: "provider:model_id".
Usage
alias LangChain.ChatModels.ChatReqLLM
alias LangChain.Chains.LLMChain
alias LangChain.Message
# Anthropic via req_llm
llm = ChatReqLLM.new!(%{model: "anthropic:claude-haiku-4-5"})
# OpenAI
llm = ChatReqLLM.new!(%{model: "openai:gpt-4o"})
# Ollama local model
llm = ChatReqLLM.new!(%{model: "ollama:llama3", base_url: "http://localhost:11434"})
# Groq with streaming
llm = ChatReqLLM.new!(%{model: "groq:llama-3.3-70b-versatile", stream: true})
{:ok, chain} =
%{llm: llm}
|> LLMChain.new!()
|> LLMChain.add_message(Message.new_user!("Hello!"))
|> LLMChain.run()Tool Use
Tools are translated to req_llm format automatically. The callback field in the
req_llm Tool struct is set to a stub — tool execution remains the LLMChain's
responsibility, as with all other ChatModel adapters.
Provider Options
Provider-specific options (e.g. thinking, tool_choice, seed) can be passed
via provider_opts:
ChatReqLLM.new!(%{
model: "anthropic:claude-haiku-4-5",
provider_opts: %{"thinking" => %{"type" => "enabled", "budget_tokens" => 2000}}
})Narration (Assistant Phase)
The OpenAI Responses API labels each assistant message item with a
phase: "commentary" for the model narrating what it is about to do,
often just before a tool call, and "final_answer" for its reply. Each
label becomes a marker on the text
LangChain.Message.ContentPart it belongs to, "narration" or
"answer" (see LangChain.Message.ContentPart.utterance/1). Text with no
label is left unmarked, which is every other provider.
A message whose text is entirely narration and has no tool calls is not an
answer (LangChain.Message.narration?/1), so
LangChain.Chains.LLMChain calls the model again to finish its turn.
The marker goes back with assistant history, because the model degrades when its own commentary is replayed unlabelled. Keep it wherever the conversation is stored.
How much of this a given req_llm reports depends on its version. When it
labels each content part, two items stay two parts and a streaming consumer
knows a part is narration from its first token. When it labels only the
message, the parts are rebuilt from the per-item text it reports alongside,
and while streaming the marker arrives with the terminal chunk, after the
text.
Stream Completion
A streamed turn is finished when a chunk marks the response terminal, which req_llm derives from a provider's terminal event. A stream whose body simply ends carries no such chunk, and the accumulated text would then never become a message.
The finish reason req_llm records alongside the stream says which case this
is. When it reports a finished response (:stop, :tool_calls, :length,
:content_filter), only the marker went missing and the turn is closed here
with that status. When it reports anything else, including the :incomplete
of a body that ended without a recognized terminal event, the response is
left unfinished so LangChain.Chains.LLMChain reports it as an
"incomplete_stream" error and retries rather than admitting half a
response into the conversation. Both paths log a warning.
Connection Retry Behavior
The retry_count option controls how many times a request is retried when
a pooled HTTP connection turns out to be stale (server closed it between
requests). This is a transport-level issue where retrying with a fresh
connection is the correct response.
Only closed-connection errors are retried. Timeouts, rate limits (429), overloaded (529), authentication errors, and invalid requests all return immediately -- they are not problems that a simple retry will fix.
retry_count | Total HTTP requests |
|---|---|
0 | 1 (no retries) |
1 | 2 (1 initial + 1 retry) |
2 (default) | 3 (1 initial + 2 retries) |
Req's built-in HTTP retry is disabled to prevent the two retry layers from compounding. See GitHub issue #503.
When running LLM calls from a background job queue (e.g., Oban) that has its
own retry logic, set retry_count: 0 so there are no hidden retries:
ChatReqLLM.new!(%{model: "...", retry_count: 0})
Summary
Functions
Call the LLM via req_llm with a prompt or list of messages.
Convert a LangChain ContentPart to a ReqLLM.Message.ContentPart.
Convert a ReqLLM.Response to a LangChain.Message.
Convert a single LangChain.Function to a ReqLLM.Tool with a stub callback.
Convert a list of LangChain Function structs to ReqLLM.Tool structs.
Convert a single LangChain Message to a list of ReqLLM.Message structs.
Convert a list of LangChain messages to a ReqLLM.Context.
Create a ChatReqLLM configuration.
Create a ChatReqLLM configuration, raising on error if invalid.
Determine if an error should be retried via a fallback LLM.
Translate a req_llm finish_reason atom to a LangChain.Message status atom.
Translate a single ReqLLM.StreamChunk to a list of LangChain.MessageDelta structs.
Translate a req_llm usage map to a LangChain.TokenUsage struct.
Types
Functions
Call the LLM via req_llm with a prompt or list of messages.
@spec content_part_to_req_llm(LangChain.Message.ContentPart.t()) :: ReqLLM.Message.ContentPart.t() | nil
Convert a LangChain ContentPart to a ReqLLM.Message.ContentPart.
Returns nil for unsupported types (they are filtered out of the content list).
@spec do_process_response(t(), ReqLLM.Response.t()) :: LangChain.Message.t() | {:error, LangChain.LangChainError.t()}
Convert a ReqLLM.Response to a LangChain.Message.
@spec function_to_req_llm_tool(LangChain.Function.t()) :: ReqLLM.Tool.t()
Convert a single LangChain.Function to a ReqLLM.Tool with a stub callback.
The stub callback is never invoked in normal LangChain operation — the tool definition is only used for schema generation (telling the LLM what tools exist).
The :strict flag is passed through so providers that support it (e.g. OpenAI
structured outputs) can enforce the parameter schema. Both the :parameters
(list of LangChain.FunctionParam) and :parameters_schema (raw JSONSchema
map) forms are supported, mirroring LangChain.ChatModels.ChatOpenAI.
@spec functions_to_req_llm_tools([LangChain.Function.t()] | nil) :: [ReqLLM.Tool.t()]
Convert a list of LangChain Function structs to ReqLLM.Tool structs.
Each tool gets a stub callback — tool execution remains the LLMChain's responsibility.
@spec message_to_req_llm_messages(LangChain.Message.t()) :: [ReqLLM.Message.t()]
Convert a single LangChain Message to a list of ReqLLM.Message structs.
Most roles map 1-to-1. The :tool role expands to one message per ToolResult.
@spec messages_to_req_llm_context([LangChain.Message.t()]) :: ReqLLM.Context.t()
Convert a list of LangChain messages to a ReqLLM.Context.
Tool messages are expanded: a single LangChain :tool message (which may carry
multiple ToolResult structs) becomes one ReqLLM.Message per result, matching
the one-result-per-message convention expected by OpenAI-compatible providers.
@spec new(attrs :: map()) :: {:ok, t()} | {:error, Ecto.Changeset.t()}
Create a ChatReqLLM configuration.
Create a ChatReqLLM configuration, raising on error if invalid.
@spec retry_on_fallback?(LangChain.LangChainError.t()) :: boolean()
Determine if an error should be retried via a fallback LLM.
Translate a req_llm finish_reason atom to a LangChain.Message status atom.
@spec translate_stream_chunk(ReqLLM.StreamChunk.t()) :: [LangChain.MessageDelta.t()]
Translate a single ReqLLM.StreamChunk to a list of LangChain.MessageDelta structs.
Returns an empty list for chunks that produce no LangChain deltas (e.g. empty content, non-terminal metadata).
@spec translate_usage(map() | nil) :: LangChain.TokenUsage.t() | nil
Translate a req_llm usage map to a LangChain.TokenUsage struct.
req_llm hands every provider's usage over as a complete snapshot of the
message so far, and a provider may report more than one over a stream.
Anthropic reports two, one opening the message and one closing it, which
TokenUsage.add/2 combines by keeping the larger count per class.
An unreported class arrives normalized to zero rather than omitted, so the
derived :total_tokens a snapshot carries describes only that snapshot.
Read a combined total from TokenUsage.total/1.