Gemini Batch API rate limits vs Sume's 100-run bulk queue
Gemini Batch API allows 100 concurrent batch requests and a 2GB input file. Sume bulk runs queue up to 100 Format runs with concurrency 1 to 16.

Gemini's Batch API limits are 100 concurrent batch requests, a 2GB input file, a 20GB file storage limit and a per-model cap on enqueued tokens. Sume's bulk-run queue has a different unit: up to 100 Format runs per queue, with a concurrency window from 1 to 16.
Gemini numbers are from its rate-limits page; Sume numbers from Bulk runs and the API reference, read 2026-10-01.
What are the Gemini Batch limits?
Batch requests have their own limits, separate from non-batch calls: concurrent batch requests 100, input file size 2GB, file storage 20GB. Enqueued tokens are capped per model across all your active batch jobs, and the figure depends on your usage tier.
What are Sume's bulk-run limits?
POST /v1/formats/:format_id/bulk-runs queues up to 100 runs with a concurrency window of 1 to 16 and returns a 202 queue receipt. It needs formats:write. Each item has the same body as a single run.
| Limit | Gemini Batch API | Sume bulk runs |
|---|---|---|
| Items in flight | 100 concurrent batch requests | concurrency 1 to 16 |
| Items per submission | Bounded by 2GB input file | 1 to 100 items |
| Size budget | Enqueued tokens per model | Workspace generation concurrency applies to children |
| Completion signal | Not covered here | No queue webhook; poll status_url |
What should I know before sizing a batch?
Three documented behaviors matter. The queue has no webhook; communication.webhook_url is per item, so poll status_url for queue-level progress. completed means every item is terminal, not that all succeeded, so branch on counts.failed. And mint a fresh Idempotency-Key per batch, because replaying a spent key returns 202 with the old queue.
How do I decide between them?
They solve different jobs: Gemini batches are model calls billed by tokens, while a Sume queue renders saved Formats. If you need to submit more than 100 items, split them into several queues, keeping in mind that workspace concurrency still caps the children. Practical steps are in bulk Format runs and Gemini usage tiers vs plan concurrency.
Sources
Related posts
More in Developers
- Gemini image_size 1k rejected: use 1K, and Sume resolution tiers
Gemini rejects a lowercase image_size such as 1k; use an uppercase K. Sume's images API takes a resolution tier, and only values the model lists.
- Omni 1.1 Flash start and end frame API on Sume
Omni 1.1 Flash can render between two keyframes, including looping clips. On Sume, Omni takes image_url plus end_image_url; frame_images covers other models.
- Gemini Omni previous_interaction_id vs a Sume job per request
Google chains Omni 1.1 Flash extensions with previous_interaction_id. Sume has no such field: each request is its own job, and an edit takes a video_url.
- Text to speech mulaw 8000 Hz: Gemini 3.8 and Sume TTS
Sume TTS can return pcm_mulaw or pcm_alaw at 8000 Hz in wav or raw containers. Gemini 3.8 TTS does it with audio/mulaw and audio/alaw mime types.
Written by Sume