OpenAI 429 slow_down vs 503 server_is_overloaded, and Sume

OpenAI now splits 429 slow_down (traffic ramping too fast) from 503 server_is_overloaded. Here is how each maps to Sume's 429 scope and 503 retry rules.

4 min readSume
All posts

OpenAI's changelog, read 2026-09-30, says traffic that increases too quickly can return a 429 with the code slow_down, while temporary model overload returns a 503 with server_is_overloaded. On Sume the two ideas live in different places: a 429 rate_limited names the budget in error.details.scope, and a provider_capacity_exceeded response means Sume's provider dispatch queue is full.

These are Sume's own error contracts. They do not translate OpenAI's codes, and Sume does not return slow_down or server_is_overloaded on this page's evidence.

What does OpenAI say to do?

The changelog says both responses may include Retry-After. When present, wait at least that long before retrying; when missing, use exponential backoff.

How does each case map to Sume?

Sume's rate limit docs give reads and writes separate budgets, so "a tight status-poll loop cannot 429 your own submits". A 429 says which budget it spent in error.details.scope, either read or write, and the response carries retry-after. For capacity, the errors page lists provider_capacity_exceeded.

OpenAI changelog codes against Sume error rules, read 2026-09-30.
SituationOpenAISume
Requests too fast429 slow_down429, error.details.scope is read or write; wait retry-after
Provider capacity503 server_is_overloadedprovider_capacity_exceeded; retry later with the same idempotency key
Retry headerRetry-After may be presentretry-after is sent on 429

How should my retry code differ?

For a 429, read ratelimit-remaining and back off on retry-after instead of counting requests yourself. For provider_capacity_exceeded, retry later but keep the same Idempotency-Key, so a job that was in fact created is not created twice; see idempotency keys for AI video APIs.

async function withRetry(call, tries = 4) {
  for (let i = 0; i < tries; i++) {
    const res = await call();
    if (res.status !== 429 && res.status !== 503) return res;
    const wait = Number(res.headers.get("retry-after")) || 2 ** i;
    await new Promise((r) => setTimeout(r, wait * 1000));
  }
  throw new Error("still rate limited or at capacity");
}

Is every Sume 429 a request-rate problem?

No. queue_full also arrives as a 429 and means Sume cannot accept another paid job for the workspace until one finishes or is canceled; see queue full vs concurrency full. Request rate and generation concurrency are separate limits, and raising one does not raise the other.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume