429 after a traffic spike: ramp up generation jobs in waves
Anthropic warns a sharp usage increase can trigger 429 acceleration limits. For Sume jobs, size each wave from generation_limits and the in-flight budget.

If a 429 shows up right after a jump in traffic, ramp up in waves instead of firing everything at once. Anthropic's page says a sharp increase in usage can trigger acceleration limits and advises ramping gradually. For Sume generation jobs, size each wave from the generation_limits snapshot on submit responses and refresh it before the next wave.
Anthropic facts are from its Rate limits page (listed under Sources) and Sume facts from Generation admission, read 2026-09-30.
What does Anthropic say about sudden increases?
Its note says you might get 429 errors because of acceleration limits if your organization has a sharp increase in usage, and to avoid them you should ramp up traffic gradually and maintain consistent usage patterns. Sume's docs describe no equivalent acceleration limit; its 429s are rate_limited and queue_full.
How do I size a wave of Sume jobs?
The docs give a budget for new in-flight work: max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue_capacity_remaining. wave_size_hint is only a submission hint and is never a concurrency limit. At zero headroom, wait and refresh before submitting more.
| Snapshot field | Value |
|---|---|
| concurrency_limit | 100 |
| queued_jobs_limit | 500 |
| queue_capacity_remaining (no jobs running) | 600 |
| wave_size_hint | 450 |
| New in-flight budget with 30 processing, 10 queued | 60 |
What if I get queue_full anyway?
queue_full means every accepted slot is used. Stop adding work, poll existing jobs until one is terminal, cancel queued jobs you no longer need, and retry with the same idempotency key once capacity opens. Concurrency being full alone is not an error: valid jobs are accepted as queued.
Does a slower ramp change cost?
No. Submit pacing decides when jobs start, not what they cost; the estimate is reserved at submit and captured on success. Do not resubmit a paid request because a local worker timed out; reuse the Idempotency-Key for retries of the same intent.
Sources
Related posts
More in Developers
- Activepieces 10-minute flow timeout and long Sume video jobs
Activepieces Cloud caps a flow run at 10 active minutes. Submit the Sume job async, pause the flow, and resume on the result instead of polling in a loop.
- Activepieces subflow retried: keep Sume jobs from billing twice
A retried Activepieces subflow can re-run your Sume submit. Derive the Idempotency-Key from a stable input, not the attempt, so a retry returns the same job.
- Activepieces sync webhook 408 at 30 seconds: use Sume async
An Activepieces /sync webhook returns 408 when no response arrives in 30 seconds. Call Sume in async mode, which answers 202 with a job id at once.
- Can an AI agent debug a failed webhook on Sume over MCP?
Over MCP an agent can read job status, events and delivery status to see why a webhook did not arrive. Redeliver is a REST call that needs jobs:write.
Written by Sume