Gemini API speaker diarization: 8 speakers; Sume STT has no labels

Gemini 3.5 Transcribe diarizes up to 8 speakers, 3+ experimental, within 30 minutes per request. Sume STT returns word timings but no speaker labels.

4 min readSume
All posts

The Gemini API supports speaker diarization on gemini-3.5-transcribe for up to 8 speakers, with attribution for 3 or more speakers marked experimental. Sume STT does not expose diarization at all: its result has text, word timings and optional sentence segments, so speaker labels would have to come from somewhere else.

Gemini facts are from its model page, read 2026-09-30. Sume facts are from the OpenAPI document and MCP tool description for sume/stt-1.0.

What are the Gemini diarization limits?

The model page lists diarization as supported on file processing, not on the live streaming endpoint. File processing is limited to 30 minutes per request when diarization or word-level timestamps are enabled, against 1 hour otherwise. The page also notes that diarization is incompatible with custom vocabulary.

Diarization support, read 2026-09-30.
ItemGemini 3.5 TranscribeSume STT 1.0
Speaker labelsUp to 8 speakersNone returned
3+ speakersExperimentalNot applicable
Request optionSupportedBody is closed; diarize is fixed server-side and rejected if sent
Output unitsWord annotationswords[] with start, end; optional sentences

What does a Sume STT result contain?

Each completed job has text, language fields when available, and words[] entries of { word, start, end } in seconds. Sending segmentation: { "mode": "sentence" } adds sentence segments derived from the word timings. There is no speaker key in either list.

How can I get speaker turns anyway?

If you control the recording, capture each speaker on a separate file, transcribe each one as its own Sume job, and merge by word start. That gives you the speaker from the file name, not from a model. For a mixed recording, you need a labelling step of your own after transcription. The interview workflow shows the transcription half.

When does the Gemini route fit?

When your requirement is automatic labels on a mixed recording of eight or fewer speakers and the file is under 30 minutes, Gemini documents that feature and Sume does not. When you need timings, sentences and a job API, Sume covers that.

Sources

Related posts

More in Models

All Models posts

Written by Sume