Gemini API speaker diarization: 8 speakers; Sume STT has no labels
Gemini 3.5 Transcribe diarizes up to 8 speakers, 3+ experimental, within 30 minutes per request. Sume STT returns word timings but no speaker labels.

The Gemini API supports speaker diarization on gemini-3.5-transcribe for up to 8 speakers, with attribution for 3 or more speakers marked experimental. Sume STT does not expose diarization at all: its result has text, word timings and optional sentence segments, so speaker labels would have to come from somewhere else.
Gemini facts are from its model page, read 2026-09-30. Sume facts are from the OpenAPI document and MCP tool description for sume/stt-1.0.
What are the Gemini diarization limits?
The model page lists diarization as supported on file processing, not on the live streaming endpoint. File processing is limited to 30 minutes per request when diarization or word-level timestamps are enabled, against 1 hour otherwise. The page also notes that diarization is incompatible with custom vocabulary.
| Item | Gemini 3.5 Transcribe | Sume STT 1.0 |
|---|---|---|
| Speaker labels | Up to 8 speakers | None returned |
| 3+ speakers | Experimental | Not applicable |
| Request option | Supported | Body is closed; diarize is fixed server-side and rejected if sent |
| Output units | Word annotations | words[] with start, end; optional sentences |
What does a Sume STT result contain?
Each completed job has text, language fields when available, and words[] entries of { word, start, end } in seconds. Sending segmentation: { "mode": "sentence" } adds sentence segments derived from the word timings. There is no speaker key in either list.
How can I get speaker turns anyway?
If you control the recording, capture each speaker on a separate file, transcribe each one as its own Sume job, and merge by word start. That gives you the speaker from the file name, not from a model. For a mixed recording, you need a labelling step of your own after transcription. The interview workflow shows the transcription half.
When does the Gemini route fit?
When your requirement is automatic labels on a mixed recording of eight or fewer speakers and the file is under 30 minutes, Gemini documents that feature and Sume does not. When you need timings, sentences and a job API, Sume covers that.
Sources
Related posts
More in Models
- Gemini Omni video editing: the EEA, UK and under-10-second rules
Google says editing or extending uploaded videos with Omni is unavailable in the EEA, Switzerland and the UK, and uploads must be under 10 seconds.
- Gemini Omni extend video: Google's 40 s cap vs Sume's modes
Google's Gemini Omni can extend a video by 3-10 s up to 40 s. Sume's Omni row has no extend mode; here is what it offers instead.
- Gemini Omni reference video length: 3 seconds per clip on Sume
On Sume, each Gemini Omni Flash reference video can be at most 3 seconds, with up to 3 clips. Trim longer footage first, then address clips as VIDEO_REF tags.
- Gemini TTS API: which engine Sume's TTS routes to
Gemini 3.8 TTS is not in Sume's TTS Router, which lists Cartesia Sonic ids only. What Sume's TTS takes for voice, engine and length, with a checklist.
Written by Sume