OpenAI API timeout errors: connection timeout vs. read timeout (and how to fix each)
An OpenAI API timeout isn't a single condition — there are two distinct kinds with different causes and different fixes. Getting them confused leads to either over-aggressive retries (and potential duplicate charges) or under-configured clients that timeout on legitimate long-running jobs. This page covers both, including how to configure timeouts in the Python SDK and when it's safe to retry.
The 30-second answer
- Connection timeout: your client couldn't reach OpenAI's servers at all. Usually a network issue. Safe to retry immediately.
- Read timeout: connected fine, but the response took longer than your timeout allows. The server may still be generating — retrying mid-stream risks a duplicate completion.
- The OpenAI SDK's default read timeout is 600 seconds (10 minutes) — generous for most tasks, but reasoning models (o1/o3) with high effort can exceed it.
- Check status.openai.com first — if OpenAI is having a degraded service event, no amount of timeout tuning will help.
Verified against the current SDK — 5 September 2026
Everything in this section was run, not assumed. Environment: openai==3.8.0 (installed fresh with pip install openai), httpx2==2.12.0, Python 3.11.15, Linux. Timeouts were reproduced against a local HTTP server that stalls (read timeout) and a non-routable address (connect timeout), so no OpenAI credits were spent and the numbers are deterministic.
1. import httpx no longer works in a fresh install. The 3.x SDK depends on the httpx2 package, not httpx. Copying the widely-circulated OpenAI(timeout=Timeout(...)) snippet into a new virtualenv fails before you ever reach a timeout:
$ pip install openai $ python -c "import httpx" ModuleNotFoundError: No module named 'httpx' $ python -c "import openai; print(openai.__version__, openai.Timeout)" 3.8.0 <class 'openai.Timeout'>
Use the SDK's own re-export — from openai import Timeout — which works on both the 1.x/2.x (httpx) and 3.x (httpx2) lines. Every code sample on this page has been updated to it.
2. The defaults, read from openai._constants in 3.8.0:
DEFAULT_TIMEOUT Timeout(connect=5.0, read=600, write=600, pool=600) DEFAULT_MAX_RETRIES 2 INITIAL_RETRY_DELAY 0.5s MAX_RETRY_DELAY 8.0s
So the 10-minute read timeout described below is still current.
3. Both connect and read timeouts raise openai.APITimeoutError — with the same message. In 3.8.0 a connect timeout did not surface as APIConnectionError; both paths produced APITimeoutError: Request timed out. Do not branch your retry logic on the exception class alone if you need to tell them apart — measure elapsed time against your configured connect value, or set max_retries=0 and inspect the underlying httpx2 exception.
4. The duplicate-request risk on read timeout is real, and it is the SDK doing it, not you. With the default max_retries=2, a completion that hits a 1-second read timeout was re-sent automatically; the mock server logged three arrivals with x-stainless-retry-count 0, 1, 2, and the client gave up after 4.7 s total. With max_retries=0 the server saw exactly one request and the client returned after 1.0 s:
READ TIMEOUT read=1s, default max_retries (2) -> APITimeoutError: Request timed out. -> server received 3 request(s), x-stainless-retry-count seen: ['0', '1', '2'], client gave up after 4.7s READ TIMEOUT read=1s, max_retries=0 -> APITimeoutError: Request timed out. -> server received 1 request(s), x-stainless-retry-count seen: ['0'], client gave up after 1.0s
If the model was still generating when the first request timed out, each automatic retry is a fresh, billable generation. For non-idempotent or expensive calls, construct the client with max_retries=0 and retry deliberately in your own code.
5. Per-request overrides accept either a float or a Timeout. Both timeout=0.5 and timeout=Timeout(0.5, connect=1.0) on chat.completions.create() were honoured and raised APITimeoutError as expected.
What was not re-tested: behaviour against OpenAI's live endpoint under load, and reasoning-model generation times. Those sections below are unchanged and should be read as guidance, not measurement.
The two types of timeout — and why the distinction matters
The OpenAI Python SDK uses an httpx-family HTTP transport (httpx in the 1.x/2.x SDKs; the httpx2 package in the 3.x SDKs — see the verification section above). Either way the transport distinguishes four timeout phases; in practice two of them matter most for API usage:
| Timeout type | What it means | Python exception | Safe to retry? |
|---|---|---|---|
| Connect timeout | TCP connection to OpenAI's servers not established within the limit | httpx.ConnectTimeout → openai.APIConnectionError (1.x/2.x); openai.APITimeoutError observed in 3.8.0 | Yes — request never reached OpenAI |
| Read timeout | Connected, but no data received from the server within the limit | httpx.ReadTimeout → openai.APITimeoutError | Depends (see below) |
| Write timeout | Sending the request body took too long | httpx.WriteTimeout | Yes — request not accepted |
| Pool timeout | Waiting for an available connection from the pool | httpx.PoolTimeout | Yes — request not sent |
The OpenAI SDK historically wrapped these into two exception classes: openai.APIConnectionError (couldn't connect) and openai.APITimeoutError (connected but timed out waiting for the response). In the 3.8.0 test above, a connect timeout also raised APITimeoutError, so treat the class split as version-dependent. In practice, catching both separately and applying different retry logic gives you the safest behavior.
Default timeout values in the OpenAI Python SDK
The default timeout configuration (unchanged from the v1.x SDK through 3.8.0, verified above) is:
# OpenAI SDK default — from openai._constants (v1.x through 3.8.0):
# Timeout(timeout=600.0, connect=5.0)
# That means:
# connect: 5 seconds
# read: 600 seconds (10 minutes)
# write: 600 seconds
# pool: 600 seconds
The 10-minute read timeout is intentionally generous to accommodate long completions. You'll hit it most often with reasoning models on hard tasks, or when OpenAI is under load and generation slows down. For interactive applications where you'd rather fail fast and show an error, you'll want to shorten it.
How to configure custom timeouts
Pass an httpx.Timeout object to the OpenAI client constructor. This applies to all requests made with that client instance.
from openai import OpenAI
from openai import Timeout
# Tight timeouts for an interactive app: fail fast if slow
client_fast = OpenAI(
timeout=Timeout(30.0, connect=5.0)
# connect=5s, read/write/pool=30s
)
# Relaxed timeouts for a batch job with o1 or long completions
client_batch = OpenAI(
timeout=Timeout(600.0, connect=10.0)
# connect=10s, read/write/pool=600s (SDK default, made explicit)
)
# Full four-value control
client_custom = OpenAI(
timeout=Timeout(
connect=5.0,
read=120.0,
write=10.0,
pool=5.0
)
)
You can also override the timeout on a per-request basis by passing it to the method call directly:
response = client.chat.completions.create(
model="o1",
messages=[{"role": "user", "content": "Solve this complex problem..."}],
timeout=Timeout(300.0, connect=5.0) # per-request override
)
Common causes of slow or timing-out responses
If you're hitting timeouts even with generous settings, the issue is usually one of these:
- Large
max_tokenson a slow model. Every additional output token takes time. Settingmax_tokens=4096when you typically get 200-token responses doesn't cause slowness — but if you actually generate 4000+ tokens, that takes time. Setmax_tokensto a realistic ceiling. - Reasoning models with high effort.
o1ando3-mini(ando3in preview) run an internal "thinking" pass before producing output. With"reasoning_effort": "high", a complex task can take two to five minutes. This is expected behavior — don't shorten your timeout for these models. Use streaming so you get first-token acknowledgment early. - Network path instability. OpenAI uses Cloudflare and regional infrastructure — if your server's route to OpenAI's endpoints is flapping, you'll see intermittent read stalls that look like timeouts. Try from a different network or region to diagnose.
- OpenAI service degradation. OpenAI publishes real-time status at status.openai.com. Degraded API response time is the most common incident type and it produces exactly the symptoms of a timeout — slow or no response. Check this before changing any code.
Timeouts with streaming responses
Streaming changes the timeout calculus. In streaming mode, the read timeout applies to the interval between received chunks, not the total response time. This means:
- You usually get the first token quickly (within a few seconds), which resets the read-timeout clock.
- If the stream stalls mid-generation — no new chunks for longer than your read timeout — you'll get a
ReadTimeouteven if the total elapsed time is well under your limit. - A stream stall typically indicates server-side slowdown under load, not a bug in your code.
For robust streaming, implement a client-side watchdog that tracks time since the last received chunk:
import time
from openai import OpenAI
client = OpenAI()
CHUNK_STALL_TIMEOUT = 30 # seconds between chunks before we give up
last_chunk_time = time.monotonic()
with client.chat.completions.stream(
model="gpt-5.4",
messages=[{"role": "user", "content": "Write a long essay..."}],
) as stream:
for chunk in stream:
now = time.monotonic()
if now - last_chunk_time > CHUNK_STALL_TIMEOUT:
raise TimeoutError("Stream stalled — no chunk received in 30s")
last_chunk_time = now
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
When is it safe to retry?
This is the most important practical question, because retrying at the wrong time can result in two completed requests being billed.
- Connection timeout (
APIConnectionErroron 1.x/2.x;APITimeoutErrorwithin yourconnectbudget on 3.x): always safe to retry. The request never reached OpenAI's servers, so there's nothing to duplicate. - Read timeout before the first token: almost always safe to retry. OpenAI either didn't receive the request or hasn't started processing it. Use exponential backoff.
- Read timeout mid-stream: not safe to retry blindly. The server received your request and was generating output. A retry will create a new completion. If idempotency matters (e.g., you're writing to a database), check whether you received any partial output before retrying, and decide whether to retry the full request or continue from where you left off.
- Note: OpenAI does not use HTTP 529. That status code is specific to Anthropic's API. OpenAI uses 503 for service unavailability (with a
Retry-Afterheader) and 429 for rate/quota limits. A true timeout is a client-side condition — no HTTP status code involved — which is why distinguishing it from server-error codes matters.
FAQ
Is an OpenAI API timeout the same as a 503 Service Unavailable? No. A timeout is a client-side condition — your HTTP client gave up waiting. A 503 is a server response meaning OpenAI accepted the connection but is currently unable to serve the request. Both can have similar causes (server overload), but they're handled differently. A 503 gives you an HTTP response to parse; a timeout gives you a Python exception with no HTTP status.
Does OpenAI support a request-level timeout separate from the SDK-level one? Yes — you can pass timeout= directly to any method call (e.g., client.chat.completions.create(..., timeout=60)), which overrides the client-level default for that single request. The value can be a float (seconds) or a full httpx.Timeout object for granular control.
Why am I getting timeouts only on o1/o3 but not on GPT-4o? Reasoning models run an internal chain-of-thought pass before producing output — this is where almost all of the latency lives. With high reasoning effort on a complex problem, o1 and o3 can take 2–5 minutes or more. This is expected. Set your timeout to at least 300–600 seconds for these models, and use streaming so your application stays responsive while waiting for output.
Verified on 5 September 2026 against openai 3.8.0 / httpx2 2.12.0 / Python 3.11.15 with a local mock server (connect, read, retry and per-request-override behaviour reproduced; outputs shown above). Sections describing live-endpoint behaviour under load were not re-measured and are marked as guidance.