HeyGen 429 Retry-After vs Sume rate_limited and queue_full
HeyGen sends 429 with a Retry-After header in seconds. Sume sends 429 rate_limited with retry-after when present, and a separate 429 queue_full.

On HeyGen, a 429 carries a Retry-After header with the seconds to wait. On Sume, read retry-after when it is present, and tell two 429 codes apart: rate_limited (too many requests) and queue_full (no room for another paid generation job). Only the first is fixed by waiting out a request window.
HeyGen facts are from its Usage Limits page; Sume facts from Errors and rate limits and Generation admission, read 2026-10-01.
What does HeyGen return when a limit is exceeded?
HeyGen says all endpoints enforce rate limits, and that exceeding one returns 429 Too Many Requests with a Retry-After header giving the number of seconds to wait. Exceeding its concurrent-job limit also returns 429 with Retry-After.
Which headers does Sume send?
Public API responses can include ratelimit-limit, ratelimit-remaining, ratelimit-reset and retry-after. The docs say to back off on 429, use retry-after when present, and not to retry unsafe submit requests without an Idempotency-Key.
Why are there two kinds of 429 on Sume?
Generation admission separates the controls. Submit rate limits cover request volume and fail with 429 rate_limited; the fix is backoff plus an idempotency key. queue_full means workspace concurrency plus queue capacity is full, so Sume cannot accept another paid job until an existing queued or processing job finishes or is canceled. Full concurrency alone is not an error: valid jobs are accepted as queued while queue capacity remains.
| Code | Meaning | What to do |
|---|---|---|
rate_limited | Too many requests in the current window | Back off, use retry-after, retry with the same idempotency key |
queue_full | Concurrency plus queue capacity is full | Wait for or cancel an existing job; do not just loop |
How do I write one handler for both?
Check the code before sleeping. This sketch waits only for rate_limited and reuses the same Idempotency-Key.
async function submit(body, key) {
for (let attempt = 0; attempt < 4; attempt++) {
const res = await fetch("https://api.sume.com/v1/avatar-1.0/talking-video", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.SUME_API_KEY,
"Content-Type": "application/json",
"Idempotency-Key": key,
},
body: JSON.stringify(body),
});
if (res.status !== 429) return res;
const { error } = await res.json();
if (error.code !== "rate_limited") return { queueFull: true, error };
const wait = Number(res.headers.get("retry-after")) || 2 ** attempt;
await new Promise((r) => setTimeout(r, wait * 1000));
}
throw new Error("still rate limited");
}Sources
Related posts
More in Developers
- HeyGen asset upload max size: 32 MB, and what Sume does instead
HeyGen's POST /v3/assets caps files at 32 MB, URLs included. Sume has no asset step: requests take public HTTPS URLs, and oversized bodies return 413.
- HeyGen audio file size limit (MP3/WAV 32 MB) vs Sume 10 MB
HeyGen audio_asset_id takes MP3 or WAV up to 32 MB. Sume's Fabric audio_url must be on the Sume media host and at most 10 MB. What each rule means.
- HeyGen audio to video max length: 10 or 30 min vs Sume 300 s
HeyGen's pages say 30 minutes per audio-to-video request and 10 minutes for avatar audio input. Sume's Fabric route caps duration_seconds at 300.
- HeyGen avatar script limit: 5,000 characters vs Sume's 4-60 s
HeyGen caps avatar script text at 5,000 characters. Sume has no character cap in its docs: it accepts scripts estimated at 4-60 seconds inclusive.
Written by Sume