429 backoff with jitter in Python: OpenAI's advice, Sume's headers

OpenAI recommends exponential backoff with random jitter on 429. A Python status poller that honors Sume's retry-after first and falls back to backoff.

4 min readSume
All posts

Wait the retry-after seconds when the 429 carries them, and use exponential backoff with random jitter when it does not. OpenAI's rate-limit guide describes exponential backoff as waiting briefly after a failure and increasing the delay after each retry, and says random jitter stops clients retrying at the same moment. Sume's docs say to use retry-after when present.

OpenAI facts are from its Rate limits guide (listed under Sources) and Sume facts from Errors and rate limits, read 2026-09-30.

What does each guide tell a client to do?

OpenAI says a 429 may include a Retry-After header giving seconds to wait. Sume says to back off on 429, use retry-after when present, and not retry unsafe submit requests without an Idempotency-Key.

Retry guidance (OpenAI guide and Errors and rate limits, read 2026-09-30)
RuleOpenAISume
Wait signalRetry-After header, may be presentretry-after when present
FallbackExponential backoffBack off; exponential for polling
Avoid herdingRandom jitterSDK polls are jittered
Paid submit retryNot covered hereOnly with an Idempotency-Key

How do I poll a Sume job with that policy?

Reads are cheap on Sume, but a 429 on a status read means the read failed, not the job. The job keeps running and billing, so retry the read instead of resubmitting.

import asyncio
import os
import random

import httpx


async def get_status(client: httpx.AsyncClient, job_id: str, attempts: int = 5):
    for attempt in range(attempts):
        r = await client.get(f"https://api.sume.com/v1/jobs/{job_id}/status")
        if r.status_code != 429:
            r.raise_for_status()
            return r.json()
        wait = float(r.headers.get("retry-after", 2**attempt))
        await asyncio.sleep(wait + random.uniform(0, 1))
    raise RuntimeError("still rate limited")


async def main():
    headers = {"x-api-key": os.environ["SUME_API_KEY"]}
    async with httpx.AsyncClient(headers=headers) as client:
        print(await get_status(client, "job_123"))


asyncio.run(main())

Does the Sume SDK already do this?

Yes for TypeScript: createSumeClient retries 408, 429, 5xx and transport failures twice by default with exponential backoff and jitter, honoring retry-after. A POST is retried only when it carries an Idempotency-Key. The run helpers also tolerate six consecutive transient read failures before giving up.

What are the limits of this answer?

Keep the loop bounded and log the x-sume-request-id of the last failure. A timeout in your poller does not cancel the job; store the job id and read it again later.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume