ai.LlmTimeoutOptions

The four deadlines of one LLM request attempt. Each retry or fallback attempt starts its own clocks.

Reference version

Signature

class ai.LlmTimeoutOptions

The four deadlines of one LLM request attempt. Each retry or fallback attempt starts its own clocks.

A zero or negative Duration means no limit, on any field.

let client = openai.ChatClient.new(
    model = "gpt-4o-mini",
    timeout = ai.LlmTimeoutOptions {
        timeout: baml.time.Duration.from_seconds(120),
        connect_timeout: baml.time.Duration.from_seconds(10),
        first_token_timeout: baml.time.Duration.from_seconds(30),
        stream_token_timeout: baml.time.Duration.from_seconds(10),
    },
);

Source:<builtin>/ai/timeouts.bamlbytes 804–3149

Fields

timeout

baml.time.Duration | null

End-to-end total for one attempt, connecting included. Forwarded to the underlying baml.http.send / baml.http.send_sse. null is 5 minutes.

connect_timeout

baml.time.Duration | null

Limit on establishing the connection. null is 10 seconds.

first_token_timeout

baml.time.Duration | null

Streaming only: limit on receiving the first non-empty content delta (text, reasoning or tool input). Starts when the request attempt is dispatched, so it includes connecting, sending the request and waiting for response headers. Keep-alives, response metadata and empty deltas do not end it. null is not enforced.

stream_token_timeout

baml.time.Duration | null

Streaming only: limit on the time between non-empty content deltas. Starts once the first non-empty content delta has arrived, and only another one resets it. null is not enforced.

Static methods

function

_resolve

(
timeout: baml.time.Duration | ai.LlmTimeoutOptions | null
) -> ai.LlmTimeoutOptions throws never

(internal) A client's timeout argument as options: a bare Duration is the end-to-end total, and null leaves every deadline at its default.

Instance methods

function

_apply

(self, request: baml.http.Request) -> baml.http.Request throws never

(internal) Puts the two transport deadlines on request, which baml.http then enforces, and returns it. The two token deadlines are enforced by ai.stream.TurnStream.