Gemini rate limits: per project, not per API key. And Sume?

Google says Gemini rate limits apply per project, not per API key, so a second key adds nothing. What Sume's docs say about plan limits and how to read them.

4 min readSume
All posts

Google's Gemini rate-limit page says limits are applied per project, not per API key, so creating more keys in one project does not raise them. Sume's docs say every API key gets a request budget per minute, set by the subscription plan of the workspace the key belongs to, with separate read and write budgets.

Google facts are from its rate limits page (listed under Sources) and Sume facts from Authentication, read 2026-09-30.

What does each page say about the scope of a limit?

Google measures limits across requests per minute, input tokens per minute and requests per day, per project. Sume states a per-minute request budget for each key and ties its size to the plan of the workspace.

Limit scope side by side (Google page and Authentication, read 2026-09-30)
Gemini APISume API
Scope namedPer project, not per API keyEach key, by workspace plan
DimensionsRPM, input TPM, RPDRequests per minute; reads and writes separate
Live value on SumeNot covered hereratelimit-limit header on every response

Does a second Sume key double my budget?

The docs do not say that keys share or add budgets, so do not plan around either. What they do say: keys are workspace-scoped, the plan sets the number, and ratelimit-limit on the response is the authority for the deployment you are talking to. Read it from your own traffic instead of multiplying keys.

Which limit is not about requests at all?

Generation concurrency. Sume governs how many generations run at once separately, through the plan's concurrency limit on the generation_limits object, and a higher request rate does not raise it. Keep API keys on your server and give each service its own key so one can be revoked alone from the API Keys dashboard.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume