Gemini voice replication: consent audio and a 10-30 s sample

Gemini API voice replication needs a 10-30 second source_audio and a consent_audio from the same speaker. Sume's TTS call takes a ready voice id instead.

4 min readSume
All posts

The Gemini API docs say every voice replication request needs two real human recordings from the same adult speaker: a 10 to 30 second source_audio clip and a consent_audio clip reciting a mandatory consent statement. Sume's TTS 1.0 request takes neither; it takes a voice id that already exists.

What does Gemini require to replicate a voice?

Per the page read 2026-10-01, 24 kHz mono 16-bit WAV is recommended for both clips. The consent recording must be the same speaker reciting the consent statement in one of the supported languages.

Gemini voice replication inputs, from the Gemini API docs, read 2026-10-01.
InputRequirement
source_audio10-30 seconds of clean, natural speech from the speaker
consent_audioSame speaker reciting the consent statement
SpeakerThe same adult speaker for both clips

What does a Sume TTS request take instead?

A transcript plus a voice selector: an avatar reference, or voice.id. The id must be a TTS voice UUID or a Voices library id, "not a voice name from another TTS ecosystem"; other shapes get 400 public_reason=invalid_voice_id before a job is queued or credits are reserved. See the API reference.

Does the language still matter?

Yes. Set language for every non-English transcript; omitted, it defaults to English at the provider. Audio longer than 1200 seconds fails with tts_duration_exceeded and no credit capture.

Where do I read about cloning on Sume?

See voice cloning API. For Gemini speech on Sume, see Gemini TTS: which engine Sume routes. Get written consent from any speaker before you clone their voice anywhere.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume