ReqLLM.Providers.OpenAI.ResponsesAPI (ReqLLM v1.26.0)

View Source

OpenAI Responses API driver for Responses endpoint models.

Implements the ReqLLM.Providers.OpenAI.API behaviour for OpenAI's Responses endpoint, which provides extended reasoning capabilities for advanced models.

Endpoint

/v1/responses

Supported Models

Models with "api": "responses" metadata:

  • o-series: o1, o3, o4, o1-preview, o1-mini
  • GPT-4.1 series: gpt-4.1, gpt-4.1-mini
  • GPT-5 series: gpt-5, gpt-5-preview

Capabilities

  • Reasoning: Extended thinking with explicit reasoning token tracking when supported
  • Streaming: SSE-based streaming with reasoning deltas and usage events
  • Tools: Function calling with responses-specific format
  • Reasoning effort: Control computation intensity (minimal, low, medium, high)
  • Enhanced usage: Separate tracking of reasoning vs output tokens

Encoding Specifics

  • Input messages use input_text content type instead of text
  • Token limits use max_output_tokens instead of max_tokens
  • Tool choice format: {type: "function", name: "tool_name"}
  • Reasoning effort: {effort: "high"} format
  • Code Interpreter: tool maps such as %{"type" => "code_interpreter", "container" => ...} are passed through unchanged to OpenAI. Both object containers (%{"type" => "auto", ...}) and string container IDs are supported. This is Responses-API-only; Chat Completions does not support the code_interpreter tool type.

Code Interpreter

Raw code_interpreter_* output items from the response are collected in response.provider_meta["code_interpreter"]["items"] and are excluded from normal text and function tool-call extraction.

Decoding

Non-streaming Responses

Aggregates multiple output segment types:

  • output_text segments → text content
  • reasoning segments (summary + content) → thinking content
  • function_call segments → tool_call parts

Streaming Events

  • response.output_text.delta → text chunks
  • response.reasoning.delta / response.reasoning_text.delta → thinking chunks
  • response.reasoning_summary_text.delta → thinking chunks whose metadata carries item_id, output_index and summary_index
  • response.reasoning_summary_part.added / .done → meta chunks with a reasoning_summary_part map (status, item_id, output_index, summary_index, text) marking summary part boundaries
  • response.output_item.done with a compaction item → :content_part chunk holding a :provider_block that later requests replay verbatim
  • response.usage → usage metrics with reasoning_tokens
  • response.completed → terminal event with finish_reason
  • response.incomplete → terminal event for truncated responses
  • response.failed → terminal event with finish_reason: :error and the failure message
  • error → terminal event with finish_reason: :error and the error message

Usage Normalization

Extracts reasoning tokens from usage.output_tokens_details.reasoning_tokens and provides:

  • :reasoning_tokens - Primary field (recommended)
  • :reasoning - Backward-compatibility alias (deprecated)

Summary

Functions

build_body(request)

build_compact_body(context, model_name, opts, request \\ nil)

@spec build_compact_body(
  ReqLLM.Context.t(),
  String.t(),
  map() | keyword(),
  map() | nil
) :: map()

Builds the request body for POST /responses/compact.

Uses previous_response_id from provider_options when present; otherwise encodes the context into an input array, replaying reasoning and compaction items so the service can compact the full prior state.

compact_path()

@spec compact_path() :: String.t()

Path of the Responses API compaction endpoint.

decode_stream_event(event, model, state)

init_stream_state()