OpenAI Agents Python MCP backoff ceiling vs Sume retry-after

Set the MCP retry backoff ceiling low enough that a Sume 429 with retry-after is honored first, and never retry a paid create without its idempotency key.

4 min readSume
All posts

The openai-agents-python v0.20.0 notes add a configurable retry backoff ceiling for MCP and derive streamable HTTP retry backoff from the backoffs already taken. Against Sume, the ceiling is a cap on your own waiting, not a replacement for the server's hint: when a 429 carries retry-after, wait at least that long, and treat queue_full as a different case from ordinary rate limiting.

The release page says only that these changes exist; it does not list option names in the part read on 2026-10-01, so check the SDK docs for the exact parameter. Sume behavior is from Errors and credits and Jobs and results.

What does Sume say to do on a 429?

The docs are short: "Back off when you receive 429. Use retry-after when present. Do not retry unsafe submit requests without an Idempotency-Key." Responses can also carry ratelimit-limit, ratelimit-remaining and ratelimit-reset. See rate limit headers for what each one means.

Is queue_full the same as a rate limit?

No. Both return 429, but they have different codes. rate_limited means too many requests in the current window. queue_full means the workspace's generation concurrency plus queue capacity is full, so Sume cannot accept another paid generation job until a queued or processing job finishes or is canceled. Being at concurrency by itself is not an error: valid jobs are accepted as queued while queue capacity remains. A growing exponential backoff does not free a full queue, so branch on the code field rather than the status alone. The longer explanation is in queue_full vs concurrency full.

How should a ceiling interact with the server hint?

Retry decisions against Sume responses, from the docs read 2026-10-01.
ResponseWhat the docs sayClient behavior
429 rate_limited with retry-afterUse retry-after when presentWait at least that long, even if your ceiling is lower
429 queue_fullNeeds an existing job to finish or be canceledStop submitting new paid jobs; read job status first
provider_capacity_exceededRetry later with the same idempotency keyBackoff is fine; reuse the key
Submit retryReuse the same Idempotency-KeyThe retry returns the original job instead of billing twice

What about waits on a slow job?

Backoff is for failed requests. A job that is simply still running should be waited on with jobs_wait: on remote MCP timeout_seconds defaults to 50 and is capped at 55, and wait_slice_expired means call jobs_wait again with the same ids, never resubmit the create. On the REST side, follow next_poll_after_seconds when present and back off otherwise. See jobs_wait for long video jobs.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume