Luma API 429 requests per minute: sliding window vs Sume
Luma counts requests in a sliding 60-second window and returns 429 if RPM or concurrent jobs fails. Sume returns 429 rate_limited: back off, reuse the key.

Luma's Agents API refuses a generation request with HTTP 429 when either its requests-per-minute check or its concurrent-jobs check fails; RPM is measured over a sliding 60-second window. Sume also answers 429 when request volume passes an abuse-protection limit, with code rate_limited: back off, use retry-after when present, and retry with the same idempotency key.
How does Luma's sliding window work?
Per Luma's rate-limit guide, read 2026-10-01, each request is timestamped and the API counts requests in the last 60 seconds. There is no fixed reset boundary, so you do not get a full refill at a set minute: requests age out one by one. Successful POST /v1/generations responses (201) carry X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset, and a 429 for RPM adds Retry-After.
What does Sume return on a rate limit?
The generation admission table lists 429 rate_limited as "API request volume exceeded an abuse-protection limit", with the client behavior "back off using retry-after when present". Submit rate limits are one of four separate controls; read, status and list endpoints can also be limited and should be treated as polling backpressure, not generation concurrency.
How do the two compare?
| Topic | Luma Agents API | Sume |
|---|---|---|
| Window | Sliding 60 seconds | Not described as a window in the admission docs |
| Second limit on the same 429 | Concurrent jobs | Separate queue_full code |
| Wait hint | Retry-After on RPM 429 | retry-after when present |
| Retry safety | Not covered in the page read | Idempotency-Key; a replay returns the original job |
| Support handle | X-Request-Id | x-sume-request-id on every response |
How should I retry a Sume submit?
Send an Idempotency-Key so a retry cannot create a second paid job, wait for retry-after seconds, then resubmit with the same key and body. Reusing a key for a different payload returns 409 idempotency_conflict.
async function submit(body, key) {
for (let attempt = 0; attempt < 5; attempt++) {
const res = await fetch("https://api.sume.com/v1/videos", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.SUME_API_KEY,
"Content-Type": "application/json",
"Idempotency-Key": key,
},
body: JSON.stringify(body),
});
if (res.status !== 429) return res;
const wait = Number(res.headers.get("retry-after") ?? 2 ** attempt);
await new Promise((r) => setTimeout(r, wait * 1000));
}
throw new Error("still rate limited");
}What do I quote when I ask for help?
The errors page says every response carries x-sume-request-id. Quote it with the error code, and do not send API keys or raw media URLs. For wait-time detail see how long to wait on a 429.
Sources
Related posts
More in Developers
- Luma API concurrent jobs limit vs Sume plan concurrency
Luma caps active generations per API client and answers 429 when full. Sume ties concurrency to your plan and queues extra jobs until queue_full.
- Luma API video URL expires after 1 hour: what to do on Sume
Luma Agents API video URLs are presigned and expire after 1 hour. On Sume, a completed job is fetched with your API key from the /content endpoint.
- Luma API X-Request-Id vs Sume x-sume-request-id
Luma's X-Request-Id echoes your header or is generated. Sume sends x-sume-request-id on every response; quote it, plus error.code, when you contact support.
- Luma Ray 3.2 extend with a generation id vs Sume per-clip jobs
Luma extends video from one prior generation id as start or end frame. Sume docs describe frame_images with first_frame or last_frame, a new job per clip.
Written by Sume