ReqLLM. Images. OpenAICompatible
(ReqLLM v1.26.0)
View Source
Shared codec for providers that speak the OpenAI Images wire format.
Currently used by ReqLLM.Providers.OpenAI (via
ReqLLM.Providers.OpenAI.ImagesAPI) and ReqLLM.Providers.Azure. It owns the
encoding rules — option translation, body/multipart construction, and response
decoding — so no provider has to restate them.
Option pipeline
Options flow through two steps in a fixed order, and each step has exactly one job:
1. validate_options/1 reject what no translation can express
2. translate_options/2 resolve, drop, and warn on validated inputvalidate_options/1 runs before ReqLLM.Provider.Options.process/4, so
translate_options/2 — which process/4 invokes through the provider's
ReqLLM.Provider.translate_options/3 callback — only ever sees input it can
express. That ordering is why translate_options/2 may return {opts, warnings} and never an error: everything unrepresentable was already
rejected, and everything else is a lossy-but-valid transformation reported as
a warning through :on_unsupported.
translate_options/2 also scopes the model-specific fields: the gpt-image-only
:background, :moderation, :output_compression, and :input_fidelity, and
the DALL-E-only response_format: :url, are dropped with a warning when the
target model or endpoint has no field for them.
Wire format
Generation is a JSON POST to path(:generation); editing is a multipart POST
to path(:edit), signalled by a non-nil :source_image.
Streaming generation is the same POST with "stream": true and an optional
"partial_images" count, answered as SSE. Each data: line is a JSON object:
image_generation.partial_image events carry a preview frame
(b64_json, partial_image_index, and the echoed size, quality,
background, output_format), and the terminal image_generation.completed
event carries the final image plus usage. Preview frames are opaque even
when background is transparent. decode_stream_event/2 turns those
events into ReqLLM.StreamChunks.
Summary
Functions
Builds the JSON body map for the generations endpoint.
Decodes an Images API response into a canonical ReqLLM.Response.
Decodes one streaming generations SSE event into ReqLLM.StreamChunks.
Builds the Req :form_multipart keyword list for the edits endpoint.
Whether a model id (or LLMDB.Model) is in the gpt-image family, the only
Images API models that serve streaming generations.
Normalizes image generation input into a {:ok, context, prompt} tuple.
Returns true when the options describe an image edit rather than a generation.
MIME type for an image output_format given as an atom or wire string.
Endpoint path for an image :generation or :edit.
Image option keys providers should register on the Req request.
JSON body for a streaming generations request.
Headers for a streaming generations request: the provider's auth headers,
JSON content type, SSE accept, and any custom headers from :req_http_options.
Translates generic image options into what the Images API accepts.
Rejects image options the Images API cannot express under any translation.
Rejects image options the streaming generations endpoint cannot serve.
Functions
Builds the JSON body map for the generations endpoint.
Accepts a map or keyword list with :model, :prompt, and the optional
image generation options (:n, :size, :quality, :style, :user,
:output_format, :response_format, :background, :moderation,
:output_compression).
Expects options that have already been through translate_options/2, which
resolves :aspect_ratio into :size and drops options the Images API has no
field for.
:model must be the catalog model id rather than a provider-side alias: it
decides whether response_format is a legal field for the target model.
Callers that send a different identifier on the wire (e.g. an Azure
deployment name) should replace "model" in the returned map afterwards.
@spec decode_response({Req.Request.t(), Req.Response.t()}) :: {Req.Request.t(), Req.Response.t() | Exception.t()}
Decodes an Images API response into a canonical ReqLLM.Response.
Non-2xx statuses are returned as a ReqLLM.Error.API.Response for the caller
to surface; providers with their own error extraction should route those
through it before reaching here.
@spec decode_stream_event(map(), LLMDB.Model.t()) :: [ReqLLM.StreamChunk.t()]
Decodes one streaming generations SSE event into ReqLLM.StreamChunks.
image_generation.partial_image becomes a :content_part chunk carrying the
preview frame. The part's metadata says partial?: true with its
partial_image_index; the chunk's metadata says stream_only?: true so the
frame reaches live consumers but not the assembled response.
image_generation.completed becomes the final :content_part (part metadata
partial?: false) followed by a terminal :meta chunk with usage and
provider metadata, which ends the stream. Error events, and image events whose
b64_json is missing or not valid base64, become a terminal error meta so the
stream fails fast instead of waiting for the receive timeout; anything else
decodes to nothing.
Builds the Req :form_multipart keyword list for the edits endpoint.
Required keys in opts: :model, :prompt, :source_image. Optional keys
(:mask, :n, :size, :quality, :output_format, :user, :background,
:input_fidelity, :output_compression, and the *_media_type companions)
are added only when present.
@spec gpt_image_model?(LLMDB.Model.t() | String.t() | nil) :: boolean()
Whether a model id (or LLMDB.Model) is in the gpt-image family, the only
Images API models that serve streaming generations.
Matches by catalog family metadata or by the gpt-image id prefix, on the
provider model id when set. Used both to gate ReqLLM.Images.stream_image/3
and to route provider decode_stream_event/3 callbacks to
decode_stream_event/2.
@spec image_context(term(), keyword()) :: {:ok, ReqLLM.Context.t(), String.t()} | {:error, term()}
Normalizes image generation input into a {:ok, context, prompt} tuple.
Uses an existing :context option when present, otherwise normalizes the
prompt/messages input. The prompt is the text content of the last user
message; an empty prompt is an error.
Returns true when the options describe an image edit rather than a generation.
An edit is signalled by a non-nil :source_image. An explicitly nil
:source_image is treated as a generation, since a multipart edit request
cannot be built without image bytes.
MIME type for an image output_format given as an atom or wire string.
Defaults to "image/png", the format every Images API model emits unless told
otherwise.
@spec path(:generation | :edit) :: String.t()
Endpoint path for an image :generation or :edit.
Providers that mount the Images API under a different prefix (Azure's deployment-scoped routes) build on top of these suffixes.
@spec request_option_keys() :: [atom()]
Image option keys providers should register on the Req request.
Covers every option that can reach the wire, plus :prompt. Plumbing options
(:provider_options, :receive_timeout, …) are deliberately excluded —
request-building layers register and merge those themselves.
JSON body for a streaming generations request.
Same as build_generation_body/1 with the prompt, model id, and "stream": true
applied on top of the processed options.
Headers for a streaming generations request: the provider's auth headers,
JSON content type, SSE accept, and any custom headers from :req_http_options.
Translates generic image options into what the Images API accepts.
Returns {opts, warnings} in the shape the ReqLLM.Provider.translate_options/3
callback expects, so providers sharing this codec route through it and the
transformations surface through :on_unsupported.
Drops options the API has no field for (:seed, :negative_prompt, and
:style outside DALL-E 3), maps the DALL-E quality names (:standard/:hd)
onto the gpt-image ones, and resolves :aspect_ratio into the nearest :size
the model offers — the only place that resolution happens.
The gpt-image-only fields are scoped the same way: :background,
:moderation, :output_compression, and :input_fidelity are dropped for
DALL-E; :output_compression is dropped for :png output; :moderation is
dropped for edits (the edits endpoint has no such field); :input_fidelity is
dropped for generations and for gpt-image-1-mini (gpt-image-2 accepts and
ignores it); response_format: :url is dropped for gpt-image, which only ever
returns image bytes.
Assumes validate_options/1 has already run: a malformed :aspect_ratio is
left untouched rather than raising, since it should never get this far.
model_id must be the catalog model id, since the accepted fields and sizes
differ between the gpt-image and DALL-E families.
@spec validate_options(keyword()) :: :ok | {:error, Exception.t()}
Rejects image options the Images API cannot express under any translation.
Run this before ReqLLM.Provider.Options.process/4, so that
translate_options/2 receives only representable input. Rejects a malformed
:aspect_ratio, a :mask without a :source_image, and a :transparent
background with :jpeg output (JPEG has no alpha channel); a well-formed
:aspect_ratio is left for translate_options/2 to resolve.
@spec validate_stream_options(keyword()) :: :ok | {:error, Exception.t()}
Rejects image options the streaming generations endpoint cannot serve.
Streaming is generation-only (:source_image selects the multipart edits
endpoint, which has no SSE form) and single-image (:n must be 1 when set).
ReqLLM.Images.stream_image/3 runs this before starting the stream, and the
OpenAI and Azure attach_stream/4 image paths run it again so a direct
ReqLLM.Streaming.start_stream/4 caller gets the same guarantees.