Gemini TTS two-speaker limit: three-voice dialogue with Sume concat
Gemini TTS configures two speakers per request. For three voices on Sume, make one TTS job per voice and join up to 20 parts with Timeline audio concat.

Gemini's speech guide configures two speakers in speech_config.speakers, so a three-voice scene needs more than one request there. On Sume the pattern is the same by design: each TTS job uses one voice, and Timeline audio concat joins up to 20 ordered parts into one file.
Vendor facts are from Google's speech generation guide (read 2026-10-01); Sume facts are from the OpenAPI file and Timeline audio docs.
How many speakers does Gemini take per request?
The guide says multi-speaker dialogue is configured with two speakers in speech_config.speakers. Single-speaker generation uses an array; multi-speaker uses an object with a speakers list.
How does Sume handle several voices?
A TTS job takes its voice from the top-level avatar_id / avatar_handle, or voice.id. For a dialogue, submit one job per line or per speaker turn, each with its own voice, then join them.
| Item | Value |
|---|---|
| Parts per concat | 1 to 20, ordered |
| Join type | Sample-domain, no re-TTS, no silence at seams |
| Output formats | wav (default) or mp3 |
| Produced audio | At most 1800 s |
Which output format keeps the seams clean?
Concat joins in the sample domain. The docs say mp3 re-adds priming padding at every edge, and to keep wav when the file will be joined again or drives lip-sync. Request wav for the individual lines too.
What if I have more than 20 lines?
Concat in groups, then concat the group files; every URL must already be this workspace's media.sume.com audio. The multiple voices guide covers the one-voice-per-job pattern in more detail.
Sources
Related posts
More in Use cases
- Change a garment's colour in a photo with a mask_url edit
Recolour clothing with a mask: send the photo as an input reference, a public HTTPS mask_url, and a colour prompt to ChatGPT Image 2.5 on POST /v1/images.
- HeyGen avatar new outfit with reference images vs Sume
HeyGen prompt avatars take avatar_id plus up to three reference_images for a new outfit. Sume's photo input takes one image_url per avatar.
- Add a hook title to the first seconds of a video with one cue
Send one authored cue with start 0 and end 3 to POST /v1/video-captions and Sume burns that hook text into the clip, with no speech-to-text step.
- Ken Burns effect API: Timeline stills are static holds
Sume Timeline holds a still image static; motion is accepted and ignored with a motion_ignored warning. zoompan is on the video filter allowlist.
Written by Sume