429 backoff with jitter in Python: OpenAI's advice, Sume's headers
OpenAI recommends exponential backoff with random jitter on 429. A Python status poller that honors Sume's retry-after first and falls back to backoff.

Wait the retry-after seconds when the 429 carries them, and use exponential backoff with random jitter when it does not. OpenAI's rate-limit guide describes exponential backoff as waiting briefly after a failure and increasing the delay after each retry, and says random jitter stops clients retrying at the same moment. Sume's docs say to use retry-after when present.
OpenAI facts are from its Rate limits guide (listed under Sources) and Sume facts from Errors and rate limits, read 2026-09-30.
What does each guide tell a client to do?
OpenAI says a 429 may include a Retry-After header giving seconds to wait. Sume says to back off on 429, use retry-after when present, and not retry unsafe submit requests without an Idempotency-Key.
| Rule | OpenAI | Sume |
|---|---|---|
| Wait signal | Retry-After header, may be present | retry-after when present |
| Fallback | Exponential backoff | Back off; exponential for polling |
| Avoid herding | Random jitter | SDK polls are jittered |
| Paid submit retry | Not covered here | Only with an Idempotency-Key |
How do I poll a Sume job with that policy?
Reads are cheap on Sume, but a 429 on a status read means the read failed, not the job. The job keeps running and billing, so retry the read instead of resubmitting.
import asyncio
import os
import random
import httpx
async def get_status(client: httpx.AsyncClient, job_id: str, attempts: int = 5):
for attempt in range(attempts):
r = await client.get(f"https://api.sume.com/v1/jobs/{job_id}/status")
if r.status_code != 429:
r.raise_for_status()
return r.json()
wait = float(r.headers.get("retry-after", 2**attempt))
await asyncio.sleep(wait + random.uniform(0, 1))
raise RuntimeError("still rate limited")
async def main():
headers = {"x-api-key": os.environ["SUME_API_KEY"]}
async with httpx.AsyncClient(headers=headers) as client:
print(await get_status(client, "job_123"))
asyncio.run(main())Does the Sume SDK already do this?
Yes for TypeScript: createSumeClient retries 408, 429, 5xx and transport failures twice by default with exponential backoff and jitter, honoring retry-after. A POST is retried only when it carries an Idempotency-Key. The run helpers also tolerate six consecutive transient read failures before giving up.
What are the limits of this answer?
Keep the loop bounded and log the x-sume-request-id of the last failure. A timeout in your poller does not cancel the job; store the job id and read it again later.
Sources
Related posts
More in Developers
- OpenAI 429 slow_down vs 503 server_is_overloaded, and Sume
OpenAI now splits 429 slow_down (traffic ramping too fast) from 503 server_is_overloaded. Here is how each maps to Sume's 429 scope and 503 retry rules.
- OpenAI Agents SDK imageGenerationTool action edit and Sume
imageGenerationTool() takes action: 'generate' | 'edit' | 'auto' in openai-agents-js 0.18.0. To edit with Sume instead, wrap its image route as a function tool.
- OpenAI Agents SDK MCP require_approval for Sume write tools
Use require_approval with a tool_names list to gate Sume write tools by name, or connect with OAuth mcp:read so those tools are never visible.
- OpenAI Agents SDK MCP tool_input_guardrails for Sume spend
tool_input_guardrails on an Agents SDK MCP server can reject a call before it runs. For Sume paid tools, check idempotency_key and max_spend_usd.
Written by Sume