ElevenLabs per-key concurrency caps vs Sume's plan limit
ElevenLabs lets enterprise service account keys carry TTS, music and dubbing concurrency limits. Sume has one plan-based limit and no per-key setting.

The ElevenLabs changelog added optional tts_concurrency_limit, music_concurrency_limit and dubbing_concurrency_limit integers for enterprise service account API keys, so a cap can be set per key and per product. Sume has no per-key setting: generation concurrency is plan-only and applies to the workspace, and the effective value is returned as generation_limits.concurrency_limit.
What did ElevenLabs add?
The changelog lists the fields on both the create and the update service account API key endpoints.
| Endpoint | New optional field |
|---|---|
| POST /v1/service-accounts/{service_account_user_id}/api-keys | tts_concurrency_limit, music_concurrency_limit, dubbing_concurrency_limit |
| PATCH /v1/service-accounts/{service_account_user_id}/api-keys/{api_key_id} | The same three fields |
How is concurrency set on Sume?
The generation admission page says generation concurrency is plan-only: prepaid top-ups do not raise the processing concurrency limit, and admin overrides may raise the effective concurrency_limit. The dashboard Concurrency tab is the source of truth, exposed as generation_limits.concurrency_limit.
Does a higher request rate raise concurrency?
No. The authentication docs say how many generations run at once is governed separately by your plan's concurrency limit, and that raising your request rate does not raise it. Jobs above the limit wait in the queue; see queue full vs concurrency full.
What if I want to keep one service from using all the slots?
The sources read here describe no per-key or per-product split on Sume. You would limit it in your own submit code, for example with a client-side semaphore per service. Compare the vendor TTS case in ElevenLabs TTS concurrency vs Sume queued jobs.
Sources
Related posts
More in Developers
- 429 enforced_spend_limit_reached: why retrying never works
A 429 with no retry-after can be a monthly spend cap, not a rate limit. What Anthropic's page says, and how Sume separates 402, 429 rate_limited and queue_full.
- WAV 16-bit PCM vs 32-bit float: what audio detach returns
Sume audio detach writes wav as 16-bit PCM (pcm_s16le) by default. If your camera records 32-bit float, here is what that means and which options you get.
- Face swap API: quality has no default, unlike avatar videos
Avatar videos default to plus when quality is omitted. Sume's Beta face swap has no default: quality is required and takes standard, plus or max.
- Firecrawl cache: maxAge 2 days vs Sume's fresh: true
Firecrawl scrape caches for 2 days by default (maxAge 172800000 ms). Sume's scrape has a boolean fresh, default false; fresh: true bypasses cached content.
Written by Sume