BFL 24 concurrent requests, 6 for flux-kontext-max, vs Sume

BFL caps concurrent requests at 24 (6 for flux-kontext-max). Sume treats concurrency as a dispatch limit: extra jobs queue until you hit 429 queue_full.

4 min readSume
All posts

BFL's integration guidelines list a maximum of 24 concurrent requests for most endpoints and 6 for flux-kontext-max, with exponential backoff on 429. Sume's docs describe a different shape: concurrency is a dispatch limit, so a submit past the processing limit is still accepted as queued until queue capacity runs out.

BFL figures are from its guidelines page and Sume behavior is from Generation admission, both read 2026-10-01.

What does BFL say the limits are?

The rate-limiting note gives three rules: at most 24 concurrent requests for most endpoints, at most 6 for flux-kontext-max, and exponential backoff for 429 responses. The same page also suggests a queue system for high-volume applications, which means the waiting happens in your code.

How does Sume treat concurrency?

The docs say: "Concurrency is a dispatch limit, not a submit limit." If the workspace is at its generation concurrency limit, Sume can still accept jobs as queued while queue capacity remains, and workers move them to processing later.

The cap that rejects work is queue capacity: new paid generation submissions fail with 429 queue_full. Request volume is separate and fails with 429 rate_limited.

Which control applies where?

Limits as documented, read 2026-10-01.
QuestionBFL guidelinesSume docs
Concurrent cap24 for most endpoints, 6 for flux-kontext-maxWorkspace concurrency_limit on processing jobs
Submit past the capNot described beyond backoff on 429Accepted as queued while capacity remains
Hard stop429, back off exponentially429 queue_full
Request volumeSame 429 guidance429 rate_limited, back off

How should I size a batch on Sume?

Read the generation_limits returned by submit responses and GET /v1/balance. The docs size new in-flight work as concurrency_limit minus active and queued jobs, capped by queue_capacity_remaining. wave_size_hint is "Not a concurrency limit, override, or processing width", so do not use it as one.

Retry a queue_full with the same idempotency key after jobs finish. For general retry handling see 429 and retry-after.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume