Google Chirp 3 allows 200 requests a minute; Sume limits by plan

Google lists per-minute request quotas by voice type: Chirp 3 at 200, Studio 500, Standard 1,000. Sume TTS is limited by workspace concurrency on your plan.

4 min readSume
All posts

Google Cloud Text-to-Speech rates requests per minute, per project, by voice type: Standard 1,000, Neural2 1,000, Studio 500, Chirp 3 at 200, Long Audio Synthesis 100 and Voice Cloning 30. Sume has no per-minute number for TTS. It limits how many paid jobs process at once in a workspace, by plan: 1, 4, 8 or 20.

Google's quotas are from its quotas page, read 2026-10-01; Sume's from Generation admission.

What are Sume's plan limits?

The docs list processing concurrency and queue capacity by plan, and say to prefer the effective generation_limits.concurrency_limit field over the static table, since overrides can change it.

Sume default generation limits by plan, from Generation admission, read 2026-10-01.
PlanProcessing at onceQueue capacity
Free15
Pro420
Startup840
Scale20100

Can I compare 200 a minute with 4 at once?

Not directly. A rate is requests over time; concurrency is jobs running together, so throughput depends on how long a job takes, which neither source states for speech. If a job took 5 seconds, 4 at once would be about 48 a minute, but that 5 seconds is an assumption. Measure your own.

What happens when I send more than my plan runs?

Extra valid jobs queue and run as slots open. A full queue gives 429 queue_full; submit volume above an abuse limit gives 429 rate_limited, which you retry with backoff and an idempotency key. Prepaid top-ups do not raise concurrency.

Where do I read my own limit?

In the dashboard Concurrency tab, and on each submit response as generation_limits. Size a batch from that field, not from the table above.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume