Gemini TTS API: which engine Sume's TTS routes to
Gemini 3.8 TTS is not in Sume's TTS Router, which lists Cartesia Sonic ids only. What Sume's TTS takes for voice, engine and length, with a checklist.

Sume does not route to Gemini TTS. The Sume TTS Router catalog is Sonic only: model is an enum of sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview, and sume/tts-1.0 has no engine picker at all. If your plan depends on a Gemini model, call Google's API; if it depends on a voice and a finished audio file, read the checklist below.
Gemini facts are from Google's launch post for Gemini 3.8 Flash TTS and Flash-Lite TTS. Sume facts are from the API reference OpenAPI description, read 2026-09-30.
What does Google say Gemini 3.8 TTS offers?
Google's post lists new voices created from natural-language prompts, 2,000+ production-ready voices, voice replication from a 30-second sample, two-speaker scene staging, and more than 100 languages and dialects. It says every clip is watermarked with SynthID, and that voice replication requires a verbal consent recording. Availability is the Gemini API, Google AI Studio, Gemini Notebook and Google Vids; the article does not give pricing.
What does Sume's TTS take instead?
A Sume TTS request takes a transcript and a voice selector: a top-level avatar_id or avatar_handle, or voice.id. The reference says the discoverable route is the avatar list, GET /v1/avatar-1.0/avatars, using an avatar whose voice.status is ready. There is no field for describing a voice in words.
Sume's TTS is an asynchronous job: you submit, then poll or receive a webhook. Synthesized audio longer than 1200 seconds fails with tts_duration_exceeded, and no credit is captured.
| Question | Gemini 3.8 TTS (Google) | Sume TTS |
|---|---|---|
| Voice from a text prompt | Yes, per the post | No field for it |
| Pick a voice | 2,000+ ready voices | avatar_id, avatar_handle or voice.id |
| Pick an engine | Gemini 3.8 Flash or Flash-Lite | TTS Router model: Sonic ids only |
| Length cap | Not stated in the post | 1200 s per job |
| Delivery | Gemini API and Google apps | Async job, poll or webhook |
How do I pick a Sonic model on Sume?
Send the Sonic id in model to POST /v1/tts-router/generate. The list of ids comes from GET /v1/tts-router/models; an unknown id fails with 400 model_not_found.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: gemini-checklist-001" \
-d '{
"model": "sonic-3.6",
"transcript": "Your order ships tomorrow.",
"avatar_handle": "acme",
"language": "en"
}'What should I do if I need Gemini voices?
Use Google's API for the Gemini audio, then bring the file into your pipeline. Sume's model ids and what each Sonic version changes are in Cartesia Sonic 3.6 API model ids, and the general request shape is in the text-to-speech API guide.
Sources
Related posts
More in Models
- Generative Extend any aspect ratio: Sume video ratios
Premiere's Generative Extend now takes any aspect ratio. Sume's gemini-omni-flash-1.1 makes 16:9 or 9:16 clips; for other shapes, conform with Timeline fit.
- Genjutsu Object Swap API: Sume serves Motion Transfer only
Higgsfield pairs Motion Transfer with Object Swap in Genjutsu. Sume's Genjutsu row serves Motion Transfer only; Object Swap is rejected by name.
- GPT Image 2.5 cached input price: does Sume pass it on?
OpenAI lists cached input at $1.25 text and $2 image per million tokens for Sunburst. Sume's docs give estimates and reserved cost, not a cache discount.
- GPT Image 2 supported aspect ratios on Sume: 5:4, 9:8, 4:5
On Sume, gpt-image-2 accepts 5:4, 9:8 and 4:5 on top of the common ratios, plus custom pixels with a 3:1 limit. The exact list and the 9:8 size to send.
Written by Sume