OpenAI API timeout errors: connection timeout vs. read timeout (and how to fix each)

An OpenAI API timeout isn't a single condition — there are two distinct kinds with different causes and different fixes. Getting them confused leads to either over-aggressive retries (and potential duplicate charges) or under-configured clients that timeout on legitimate long-running jobs. This page covers both, including how to configure timeouts in the Python SDK and when it's safe to retry.

The 30-second answer

Verified against the current SDK — 5 September 2026

Everything in this section was run, not assumed. Environment: openai==3.8.0 (installed fresh with pip install openai), httpx2==2.12.0, Python 3.11.15, Linux. Timeouts were reproduced against a local HTTP server that stalls (read timeout) and a non-routable address (connect timeout), so no OpenAI credits were spent and the numbers are deterministic.

1. import httpx no longer works in a fresh install. The 3.x SDK depends on the httpx2 package, not httpx. Copying the widely-circulated OpenAI(timeout=Timeout(...)) snippet into a new virtualenv fails before you ever reach a timeout:

$ pip install openai
$ python -c "import httpx"
ModuleNotFoundError: No module named 'httpx'
$ python -c "import openai; print(openai.__version__, openai.Timeout)"
3.8.0 <class 'openai.Timeout'>

Use the SDK's own re-export — from openai import Timeout — which works on both the 1.x/2.x (httpx) and 3.x (httpx2) lines. Every code sample on this page has been updated to it.

2. The defaults, read from openai._constants in 3.8.0:

DEFAULT_TIMEOUT      Timeout(connect=5.0, read=600, write=600, pool=600)
DEFAULT_MAX_RETRIES  2
INITIAL_RETRY_DELAY  0.5s    MAX_RETRY_DELAY  8.0s

So the 10-minute read timeout described below is still current.

3. Both connect and read timeouts raise openai.APITimeoutError — with the same message. In 3.8.0 a connect timeout did not surface as APIConnectionError; both paths produced APITimeoutError: Request timed out. Do not branch your retry logic on the exception class alone if you need to tell them apart — measure elapsed time against your configured connect value, or set max_retries=0 and inspect the underlying httpx2 exception.

4. The duplicate-request risk on read timeout is real, and it is the SDK doing it, not you. With the default max_retries=2, a completion that hits a 1-second read timeout was re-sent automatically; the mock server logged three arrivals with x-stainless-retry-count 0, 1, 2, and the client gave up after 4.7 s total. With max_retries=0 the server saw exactly one request and the client returned after 1.0 s:

READ TIMEOUT read=1s, default max_retries (2)
  -> APITimeoutError: Request timed out.
  -> server received 3 request(s), x-stainless-retry-count seen: ['0', '1', '2'], client gave up after 4.7s

READ TIMEOUT read=1s, max_retries=0
  -> APITimeoutError: Request timed out.
  -> server received 1 request(s), x-stainless-retry-count seen: ['0'], client gave up after 1.0s

If the model was still generating when the first request timed out, each automatic retry is a fresh, billable generation. For non-idempotent or expensive calls, construct the client with max_retries=0 and retry deliberately in your own code.

5. Per-request overrides accept either a float or a Timeout. Both timeout=0.5 and timeout=Timeout(0.5, connect=1.0) on chat.completions.create() were honoured and raised APITimeoutError as expected.

What was not re-tested: behaviour against OpenAI's live endpoint under load, and reasoning-model generation times. Those sections below are unchanged and should be read as guidance, not measurement.

The two types of timeout — and why the distinction matters

The OpenAI Python SDK uses an httpx-family HTTP transport (httpx in the 1.x/2.x SDKs; the httpx2 package in the 3.x SDKs — see the verification section above). Either way the transport distinguishes four timeout phases; in practice two of them matter most for API usage:

Timeout typeWhat it meansPython exceptionSafe to retry?
Connect timeoutTCP connection to OpenAI's servers not established within the limithttpx.ConnectTimeout → openai.APIConnectionError (1.x/2.x); openai.APITimeoutError observed in 3.8.0Yes — request never reached OpenAI
Read timeoutConnected, but no data received from the server within the limithttpx.ReadTimeout → openai.APITimeoutErrorDepends (see below)
Write timeoutSending the request body took too longhttpx.WriteTimeoutYes — request not accepted
Pool timeoutWaiting for an available connection from the poolhttpx.PoolTimeoutYes — request not sent

The OpenAI SDK historically wrapped these into two exception classes: openai.APIConnectionError (couldn't connect) and openai.APITimeoutError (connected but timed out waiting for the response). In the 3.8.0 test above, a connect timeout also raised APITimeoutError, so treat the class split as version-dependent. In practice, catching both separately and applying different retry logic gives you the safest behavior.

Default timeout values in the OpenAI Python SDK

The default timeout configuration (unchanged from the v1.x SDK through 3.8.0, verified above) is:

# OpenAI SDK default — from openai._constants (v1.x through 3.8.0):
# Timeout(timeout=600.0, connect=5.0)
# That means:
#   connect: 5 seconds
#   read:    600 seconds (10 minutes)
#   write:   600 seconds
#   pool:    600 seconds

The 10-minute read timeout is intentionally generous to accommodate long completions. You'll hit it most often with reasoning models on hard tasks, or when OpenAI is under load and generation slows down. For interactive applications where you'd rather fail fast and show an error, you'll want to shorten it.

How to configure custom timeouts

Pass an httpx.Timeout object to the OpenAI client constructor. This applies to all requests made with that client instance.

from openai import OpenAI
from openai import Timeout

# Tight timeouts for an interactive app: fail fast if slow
client_fast = OpenAI(
    timeout=Timeout(30.0, connect=5.0)
    # connect=5s, read/write/pool=30s
)

# Relaxed timeouts for a batch job with o1 or long completions
client_batch = OpenAI(
    timeout=Timeout(600.0, connect=10.0)
    # connect=10s, read/write/pool=600s (SDK default, made explicit)
)

# Full four-value control
client_custom = OpenAI(
    timeout=Timeout(
        connect=5.0,
        read=120.0,
        write=10.0,
        pool=5.0
    )
)

You can also override the timeout on a per-request basis by passing it to the method call directly:

response = client.chat.completions.create(
    model="o1",
    messages=[{"role": "user", "content": "Solve this complex problem..."}],
    timeout=Timeout(300.0, connect=5.0)  # per-request override
)

Common causes of slow or timing-out responses

If you're hitting timeouts even with generous settings, the issue is usually one of these:

Timeouts with streaming responses

Streaming changes the timeout calculus. In streaming mode, the read timeout applies to the interval between received chunks, not the total response time. This means:

For robust streaming, implement a client-side watchdog that tracks time since the last received chunk:

import time
from openai import OpenAI

client = OpenAI()
CHUNK_STALL_TIMEOUT = 30  # seconds between chunks before we give up

last_chunk_time = time.monotonic()

with client.chat.completions.stream(
    model="gpt-5.4",
    messages=[{"role": "user", "content": "Write a long essay..."}],
) as stream:
    for chunk in stream:
        now = time.monotonic()
        if now - last_chunk_time > CHUNK_STALL_TIMEOUT:
            raise TimeoutError("Stream stalled — no chunk received in 30s")
        last_chunk_time = now
        if chunk.choices[0].delta.content:
            print(chunk.choices[0].delta.content, end="", flush=True)

When is it safe to retry?

This is the most important practical question, because retrying at the wrong time can result in two completed requests being billed.

FAQ

Is an OpenAI API timeout the same as a 503 Service Unavailable? No. A timeout is a client-side condition — your HTTP client gave up waiting. A 503 is a server response meaning OpenAI accepted the connection but is currently unable to serve the request. Both can have similar causes (server overload), but they're handled differently. A 503 gives you an HTTP response to parse; a timeout gives you a Python exception with no HTTP status.

Does OpenAI support a request-level timeout separate from the SDK-level one? Yes — you can pass timeout= directly to any method call (e.g., client.chat.completions.create(..., timeout=60)), which overrides the client-level default for that single request. The value can be a float (seconds) or a full httpx.Timeout object for granular control.

Why am I getting timeouts only on o1/o3 but not on GPT-4o? Reasoning models run an internal chain-of-thought pass before producing output — this is where almost all of the latency lives. With high reasoning effort on a complex problem, o1 and o3 can take 2–5 minutes or more. This is expected. Set your timeout to at least 300–600 seconds for these models, and use streaming so your application stays responsive while waiting for output.

Verified on 5 September 2026 against openai 3.8.0 / httpx2 2.12.0 / Python 3.11.15 with a local mock server (connect, read, retry and per-request-override behaviour reproduced; outputs shown above). Sections describing live-endpoint behaviour under load were not re-measured and are marked as guidance.