openai.AzureClient

Azure OpenAI Service chat completions.

Reference version

Signature

class openai.AzureClient

Azure OpenAI Service chat completions.

Azure differs from OpenAI in three ways, all of them config:

  • the deployment is part of the URL — https://{resource_name}.openai.azure.com/openai/deployments/{deployment_id}/chat/completions — with the API version as a query parameter;
  • the credential rides in an api-key header, not Authorization: Bearer;
  • a token ceiling is effectively mandatory, so 4096 is injected when the caller set neither max_tokens nor max_completion_tokens (including through request_body). Passing request_body = {"max_tokens": null} removes it again — the merge runs after the default is applied.

Supply either base_url, which names the deployment in its path, OR deployment_id (with resource_name or AZURE_OPENAI_ENDPOINT). There is no "azure/<model>" shorthand: Azure serves only the deployments created on a resource, under whatever names they were given, so a model name does not identify one.

Source:<builtin>/openai/azure.bamlbytes 1044–11722

Fields

model

string

resource_name

string | null

deployment_id

string | null

api_version

string | null

base_url

ai.Credential | null

api_key

ai.Credential | null

request_body

baml.json.json | null

headers

map<string, string> | null

query_params

map<string, string> | null

temperature

float | null

max_tokens

int | null

max_completion_tokens

int | null

top_p

float | null

stop

string[] | null

seed

int | null

capture_wire

bool

timeout

baml.time.Duration | ai.LlmTimeoutOptions | null

Static methods

function

new

(
model: string,
resource_name: string | null = …,
deployment_id: string | null = …,
api_version: string | null = …,
base_url: ai.Credential | null = …,
api_key: ai.Credential | null = …,
request_body: baml.json.json | null = …,
headers: map<string, string> | null = …,
query_params: map<string, string> | null = …,
temperature: float | null = …,
max_tokens: int | null = …,
max_completion_tokens: int | null = …,
top_p: float | null = …,
stop: string[] | null = …,
seed: int | null = …,
capture_wire: bool = …,
timeout: baml.time.Duration | ai.LlmTimeoutOptions | null = …
) -> openai.AzureClient throws never

Creates a client. Construction reads no environment variable, so declaring one is always safe; credentials and endpoints resolve when a request is built.

Instance methods

function

compat

(
self
) -> openai.internal.ChatCompat throws baml.errors.Io | baml.errors.ParseError | ai.errors.InvalidRequest

The provider record: everything the shared chat core needs to know about this endpoint.

function

params

(
self,
preview: bool = …
) -> openai.internal.ChatParams throws baml.errors.Io | baml.errors.ParseError

The per-request parameters handed to the shared chat core.

Parameters

  • preview: omit the credential, for rendering a request without sending it.
function

resolved_api_key

(self) -> string | null throws baml.errors.Io | baml.errors.ParseError

The request-time credential: the explicit value, else AZURE_OPENAI_API_KEY. Azure sends it in an api-key header rather than as a bearer token.

function

resolved_base_url

(
self
) -> string throws baml.errors.Io | baml.errors.ParseError | ai.errors.InvalidRequest

The deployment base this client posts to, resolved in the order Azure users actually configure things:

  1. an explicit base_url (a literal or an env.NAME ref), used verbatim — it already names the deployment;
  2. deployment_id with resource_name, which builds https://{resource}.openai.azure.com/openai/deployments/{deployment};
  3. deployment_id with AZURE_OPENAI_ENDPOINT — the variable the official Azure OpenAI SDKs read — treated as the resource ORIGIN and given the same /openai/deployments/{deployment} path as rung 2.

The engine reports the mutually-exclusive shapes at compile time with per-key spans; a BAML constructor can only report at request time, so the error names every source that was checked.

Implementations

ai.Client for openai.AzureClient

Instance methods

function

default_cache_args

(self) -> baml.json.json throws never

What a ${cache()} sends as prompt_cache_breakpoint on the part before it, judged from model by the same rules as openai.ChatClient: {"mode": "explicit"}, or null for GPT-5.5 and earlier, which reject the field. Azure routes on the deployment, not model, so this is only as good as model's name for what the deployment serves.

Provisioned (PTU-M) deployments are not supported: Azure rejects prompt_cache_breakpoint on them, and nothing here can tell one from a Standard deployment, so a ${cache()} on a GPT-5.6 PTU-M deployment is a 400.

function

id

(self) -> string throws never
function

invoke

(
self,
input: ai.ModelTurnInput
) -> ai.ModelTurn throws baml.errors.Timeout | baml.errors.UnknownError | reflect.errors.CompilationError | ai.errors.Failure
function

render

(
self,
input: ai.ModelTurnInput
) -> baml.http.Request throws baml.errors.Timeout | baml.errors.UnknownError | reflect.errors.CompilationError | ai.errors.Failure

Source:<builtin>/openai/azure.bamlbytes 9570–11209

ai.stream.StreamingClient for openai.AzureClient

Instance methods

function

invoke_stream

(
self,
input: ai.ModelTurnInput
) -> ai.stream.TurnStream throws baml.errors.Timeout | baml.errors.UnknownError | reflect.errors.CompilationError | ai.errors.Failure

Source:<builtin>/openai/azure.bamlbytes 11215–11720