Access AI models through Microsoft Azure's enterprise cloud platform. Supports OpenAI models (GPT-4, GPT-4o, o1, o3 series) and Anthropic Claude models with full tool calling and streaming support.

Configuration

Azure uses API key authentication with deployment-based routing.

Environment Variables

AZURE_OPENAI_API_KEY=your-api-key

Provider Options

ReqLLM.generate_text(
  "azure:gpt-4o",
  "Hello",
  base_url: "https://my-resource.openai.azure.com/openai",
  deployment: "my-gpt4-deployment"
)

Model Specs

For the full model-spec workflow, see Model Specs.

Azure is a strong fit for the full explicit model specification path because base_url is part of the model metadata. Use exact Azure model IDs from LLM Catalog when possible, and use ReqLLM.model!/1 when you need to pin base_url or work ahead of the registry. deployment remains a request option.

Key Differences from Direct Provider APIs

  1. Custom endpoints: Each Azure resource has a unique base URL (https://{resource}.openai.azure.com/openai)

  2. Deployment-based routing: Models are accessed via deployments, not model names.

    • OpenAI: /deployments/{deployment}/chat/completions?api-version={version}
    • Anthropic: /v1/messages (model specified in body, like native Anthropic API)
  3. API key authentication: Uses api-key header for all model families

  4. No model field in body: The deployment ID in the URL determines the model

Provider Options

Passed via :provider_options keyword or as top-level options:

base_url (Required)

  • Type: String
  • Purpose: Azure resource endpoint
  • Format: https://{resource-name}.openai.azure.com/openai
  • Example: base_url: "https://my-company.openai.azure.com/openai"
  • Note: Must be customized for your Azure resource

deployment

  • Type: String
  • Default: Uses model.id (e.g., gpt-4o)
  • Purpose: Azure deployment name that determines which model is used
  • Example: deployment: "my-gpt4-deployment"
  • Note: The deployment name is configured when you deploy a model in Azure

api_version

  • Type: String
  • Default: "2025-04-01-preview"
  • Purpose: Azure API version
  • Example: provider_options: [api_version: "2024-10-01-preview"]
  • Note: Check Azure documentation for supported versions

api_key

  • Type: String
  • Purpose: Azure API key
  • Fallback: AZURE_OPENAI_API_KEY env var
  • Example: api_key: "your-api-key"

Examples

Basic Usage (OpenAI)

{:ok, response} = ReqLLM.generate_text(
  "azure:gpt-4o",
  "What is Elixir?",
  base_url: "https://my-resource.openai.azure.com/openai",
  deployment: "my-gpt4-deployment"
)

Basic Usage (Anthropic Claude)

{:ok, response} = ReqLLM.generate_text(
  "azure:claude-3-sonnet",
  "What is Elixir?",
  base_url: "https://my-resource.openai.azure.com/openai",
  deployment: "my-claude-deployment"
)

Streaming

{:ok, response} = ReqLLM.stream_text(
  "azure:gpt-4o",
  "Tell me a story",
  base_url: "https://my-resource.openai.azure.com/openai",
  deployment: "my-gpt4-deployment"
)

ReqLLM.StreamResponse.tokens(response)
|> Stream.each(&IO.write/1)
|> Stream.run()

Tool Calling

tools = [
  ReqLLM.tool(
    name: "get_weather",
    description: "Get weather for a location",
    parameter_schema: [location: [type: :string, required: true]],
    callback: &MyApp.Weather.fetch/1
  )
]

{:ok, response} = ReqLLM.generate_text(
  "azure:gpt-4o",
  "What's the weather in Paris?",
  base_url: "https://my-resource.openai.azure.com/openai",
  deployment: "my-gpt4-deployment",
  tools: tools
)

Embeddings

{:ok, embedding} = ReqLLM.generate_embedding(
  "azure:text-embedding-3-small",
  "Hello world",
  base_url: "https://my-resource.openai.azure.com/openai",
  deployment: "my-embedding-deployment"
)

Image Generation

Available for gpt-image models (gpt-image-1, gpt-image-1.5, gpt-image-2) on Azure OpenAI resources. Foundry endpoints (.services.ai.azure.com) are not supported for images.

{:ok, response} = ReqLLM.generate_image(
  "azure:gpt-image-1",
  "A watercolor painting of a lighthouse",
  base_url: "https://my-resource.openai.azure.com/openai",
  deployment: "my-image-deployment",
  size: "1024x1024"
)

[image] = ReqLLM.Response.images(response)
File.write!("lighthouse.png", image.data)

Image editing (multipart) with a source image and optional mask:

{:ok, response} = ReqLLM.generate_image(
  "azure:gpt-image-1",
  "Make the sky stormy",
  base_url: "https://my-resource.openai.azure.com/openai",
  deployment: "my-image-deployment",
  source_image: File.read!("lighthouse.png")
)

Endpoint routing follows the base_url you supply:

base_url shapeRequest pathDeployment sent as
https://<resource>.openai.azure.com/openai/deployments/<deployment>/images/{generations,edits}?api-version=…URL path segment
https://<resource>.openai.azure.com/openai/v1/images/{generations,edits}model field in the body/form
https://<resource>.services.ai.azure.com—not supported (returns an error)

Notes:

  • All three gpt-image models work on either endpoint format. A DeploymentNotFound (HTTP 404) means the deployment name does not exist on the resource, not that the model is unavailable — deployment names are set at deployment-creation time and frequently differ from the model id, so pass deployment: explicitly.
  • Cost is billed per token for these models. response.usage carries the input_tokens/output_tokens reported by the Images API alongside image_usage.
  • aspect_ratio resolves to the nearest size the model offers ("16:9" → 1536x1024); an explicit size wins. seed and negative_prompt are rejected up front, since the Images API has no such fields.
  • Azure supports :png and :jpeg output. ReqLLM rejects output_format: :webp before it sends the request.
  • The GPT Image options work as on OpenAI: background: :transparent (PNG output), output_compression with output_format: :jpeg, input_fidelity on edits, and quality: :auto | :low | :medium | :high (plus :xhigh | :max where the deployed model offers them). moderation is forwarded unchanged; Azure does not document it, so a rejection surfaces as an API error.

  • Non-image models are rejected before the request is sent — azure:gpt-4o and azure:claude-* both return ReqLLM.Error.Invalid.Parameter rather than failing at the API.
  • Image responses expose provider metadata under response.provider_meta["azure"].

See the Image Generation guide for the full option list.

Structured Output

schema = [
  name: [type: :string, required: true],
  age: [type: :pos_integer, required: true]
]

{:ok, person} = ReqLLM.generate_object(
  "azure:gpt-4o",
  "Generate a fictional person",
  schema,
  base_url: "https://my-resource.openai.azure.com/openai",
  deployment: "my-gpt4-deployment"
)

Extended Thinking (Claude)

{:ok, response} = ReqLLM.generate_text(
  "azure:claude-3-sonnet",
  "Solve this complex problem step by step",
  base_url: "https://my-resource.openai.azure.com/openai",
  deployment: "my-claude-deployment",
  reasoning_effort: :medium
)

Responses API (reasoning summaries, reasoning context, compaction)

GPT-5 family and codex models are routed to the Responses API: /responses on the v1 GA base URL, or /responses?api-version=... on the legacy base URL. The request body is built by the same encoder as the OpenAI provider, so the Responses-only options work under the azure: namespace too:

{:ok, response} =
  ReqLLM.generate_text(
    "azure:gpt-5.4",
    "Solve this carefully",
    base_url: "https://my-resource.openai.azure.com/openai",
    deployment: "gpt-5.4",
    reasoning_effort: :medium,
    provider_options: [
      reasoning_summary: :auto,
      reasoning_context: :current_turn,
      store: false
    ]
  )
  • reasoning_summary (:auto | :concise | :detailed): summaries are decoded into message.reasoning_details; streaming emits :thinking chunks with summary_index metadata and reasoning_summary_part meta chunks at part boundaries.

  • reasoning_context (:current_turn | :all_turns): which earlier reasoning items the model renders (GPT-5.4 and later; GPT-5.6 defaults to :all_turns).

  • store, previous_response_id, prompt_cache_key: stateless replay and response chaining.
  • context_management: server-side compaction, e.g. [%{type: "compaction", compact_threshold: 200_000}].

Manual compaction uses ReqLLM.compact_context/3, which posts to /responses/compact on the v1 GA base URL (https://<resource>.openai.azure.com/openai/v1). The legacy api-version surface answers /responses/compact?api-version=... with a server error at the time of writing, so use the v1 GA base URL for compaction:

{:ok, compacted} =
  ReqLLM.compact_context("azure:gpt-5.4", first.context,
    base_url: "https://my-resource.openai.azure.com/openai",
    deployment: "gpt-5.4"
  )

next = ReqLLM.Context.append(compacted.context, ReqLLM.Context.user("Add a booking form."))

The compacted context preserves the complete returned API window, including retained messages and tool items, in its original order. Do not remove items from this window before the next request.

Compaction items come back as :provider_block content parts and are replayed automatically on later requests. See the OpenAI guide for the full flow.

Supported Models

OpenAI GPT-4 Family

  • azure:gpt-4o - Latest multimodal model
  • azure:gpt-4o-mini - Smaller, faster variant
  • azure:gpt-4 - Original GPT-4
  • azure:gpt-4-turbo - Faster GPT-4 variant

OpenAI Reasoning Models

  • azure:o1 - Advanced reasoning model
  • azure:o1-mini - Smaller reasoning model
  • azure:o3 - Latest reasoning model
  • azure:o3-mini - Smaller o3 variant

Note: Reasoning models use max_completion_tokens instead of max_tokens. ReqLLM handles this translation automatically.

OpenAI Embedding Models

  • azure:text-embedding-3-small - Small, efficient embeddings
  • azure:text-embedding-3-large - Higher quality embeddings
  • azure:text-embedding-ada-002 - Legacy embedding model

OpenAI Image Models

  • azure:gpt-image-1 - GPT Image generation and editing
  • azure:gpt-image-1.5 - Improved GPT Image model
  • azure:gpt-image-2 - Latest GPT Image model

Anthropic Claude Models

  • azure:claude-3-opus - Most capable Claude model
  • azure:claude-3-sonnet - Balanced performance
  • azure:claude-3-haiku - Fast, efficient model
  • azure:claude-3-5-sonnet - Latest Claude 3.5 Sonnet

Note: Claude models support extended thinking via reasoning_effort option.

Wire Format Notes

OpenAI Models

  • Endpoint: /deployments/{deployment}/chat/completions
  • API: OpenAI Chat Completions format (model field omitted)
  • Responses API models: /responses and /responses/compact (deployment sent as the body model on v1 GA and Foundry URLs)

Anthropic Models

  • Endpoint: /v1/messages (model specified in request body)
  • API: Anthropic Messages format
  • Headers: Includes anthropic-version: 2023-06-01 and x-api-key

Common

  • Authentication: api-key header for all model families
  • Streaming: Standard Server-Sent Events (SSE)
  • API Version: Required query parameter on all requests

All differences handled automatically by ReqLLM.

Error Handling

Common error scenarios:

  • Missing API key: Set AZURE_OPENAI_API_KEY or pass api_key option
  • Invalid deployment: Ensure the deployment name matches your Azure resource
  • Placeholder base_url: Must provide your actual resource URL
  • Unsupported API version: Check Azure documentation for supported versions

Resources