Gemini TTS two-speaker limit: three-voice dialogue with Sume concat

Gemini TTS configures two speakers per request. For three voices on Sume, make one TTS job per voice and join up to 20 parts with Timeline audio concat.

4 min readSume
All posts

Gemini's speech guide configures two speakers in speech_config.speakers, so a three-voice scene needs more than one request there. On Sume the pattern is the same by design: each TTS job uses one voice, and Timeline audio concat joins up to 20 ordered parts into one file.

Vendor facts are from Google's speech generation guide (read 2026-10-01); Sume facts are from the OpenAPI file and Timeline audio docs.

How many speakers does Gemini take per request?

The guide says multi-speaker dialogue is configured with two speakers in speech_config.speakers. Single-speaker generation uses an array; multi-speaker uses an object with a speakers list.

How does Sume handle several voices?

A TTS job takes its voice from the top-level avatar_id / avatar_handle, or voice.id. For a dialogue, submit one job per line or per speaker turn, each with its own voice, then join them.

Concat limits from the docs, read 2026-10-01.
ItemValue
Parts per concat1 to 20, ordered
Join typeSample-domain, no re-TTS, no silence at seams
Output formatswav (default) or mp3
Produced audioAt most 1800 s

Which output format keeps the seams clean?

Concat joins in the sample domain. The docs say mp3 re-adds priming padding at every edge, and to keep wav when the file will be joined again or drives lip-sync. Request wav for the individual lines too.

What if I have more than 20 lines?

Concat in groups, then concat the group files; every URL must already be this workspace's media.sume.com audio. The multiple voices guide covers the one-voice-per-job pattern in more detail.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume