BFL 24 concurrent requests, 6 for flux-kontext-max, vs Sume
BFL caps concurrent requests at 24 (6 for flux-kontext-max). Sume treats concurrency as a dispatch limit: extra jobs queue until you hit 429 queue_full.

BFL's integration guidelines list a maximum of 24 concurrent requests for most endpoints and 6 for flux-kontext-max, with exponential backoff on 429. Sume's docs describe a different shape: concurrency is a dispatch limit, so a submit past the processing limit is still accepted as queued until queue capacity runs out.
BFL figures are from its guidelines page and Sume behavior is from Generation admission, both read 2026-10-01.
What does BFL say the limits are?
The rate-limiting note gives three rules: at most 24 concurrent requests for most endpoints, at most 6 for flux-kontext-max, and exponential backoff for 429 responses. The same page also suggests a queue system for high-volume applications, which means the waiting happens in your code.
How does Sume treat concurrency?
The docs say: "Concurrency is a dispatch limit, not a submit limit." If the workspace is at its generation concurrency limit, Sume can still accept jobs as queued while queue capacity remains, and workers move them to processing later.
The cap that rejects work is queue capacity: new paid generation submissions fail with 429 queue_full. Request volume is separate and fails with 429 rate_limited.
Which control applies where?
| Question | BFL guidelines | Sume docs |
|---|---|---|
| Concurrent cap | 24 for most endpoints, 6 for flux-kontext-max | Workspace concurrency_limit on processing jobs |
| Submit past the cap | Not described beyond backoff on 429 | Accepted as queued while capacity remains |
| Hard stop | 429, back off exponentially | 429 queue_full |
| Request volume | Same 429 guidance | 429 rate_limited, back off |
How should I size a batch on Sume?
Read the generation_limits returned by submit responses and GET /v1/balance. The docs size new in-flight work as concurrency_limit minus active and queued jobs, capped by queue_capacity_remaining. wave_size_hint is "Not a concurrency limit, override, or processing width", so do not use it as one.
Retry a queue_full with the same idempotency key after jobs finish. For general retry handling see 429 and retry-after.
Sources
Related posts
More in Developers
- BFL 402 and 429 retry rules, and the same split on Sume
BFL raises on 402 and backs off on 429. Sume splits the same way: 402 insufficient_credits is a stop, 429 queue_full or rate_limited means wait and retry.
- BFL polling_url on api.bfl.ai vs Sume status_url
BFL says to always poll the polling_url it returns. Sume's job envelope carries status_url and result_url for the same reason: follow them, do not build URLs.
- BFL webhooks vs polling_url, and Sume webhook mode with polling
BFL says webhook users need no polling_url change. On Sume, webhook mode still returns status_url, so verify the signed callback and keep polling as a backup.
- Boost dull video colors by API: the vibrance filter intensity
Sume's video filter allowlists vibrance, with intensity from -2 to 2 and a default of 0. A small positive value lifts muted color; a negative one mutes it.
Written by Sume