gemini-3.1-flash-tts-preview replacement: 3.8 Flash-Lite TTS
Google names gemini-3.8-flash-lite-tts as the replacement for the 3.1 preview TTS model. What that changes for you, and what stays fixed on Sume TTS.

Google's changelog says gemini-3.8-flash-lite-tts is built to replace gemini-3.1-flash-tts-preview for high-throughput production and real-time voice agent cascades. If you call Sume TTS instead, nothing in your request names a Gemini model: TTS 1.0 has no engine picker, so a Gemini model swap does not change your Sume calls.
Google's facts are from its changelog (Sep 22, 2026) and speech-generation guide; Sume's are from the API reference. Read 2026-10-01.
What does the Gemini change involve?
The Sep 22 changelog entry announces gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts as generally available, plus a Voices endpoint at /v1beta/voices. The guide says to pass the verbatim transcript in input, attach turn-level styling with a speech_metadata annotation, and that unary requests return WAV (audio/wav) with a RIFF header by default. Check the guide itself before you change a production call.
| Topic | Gemini 3.8 TTS (guide) | Sume TTS 1.0 (API reference) |
|---|---|---|
| Model choice | gemini-3.8-flash-tts or gemini-3.8-flash-lite-tts | No engine picker on 1.0 |
| Explicit model | In the request | POST /v1/tts-router/generate with a required catalog model id |
| Credentials | Google key | Sume API key only |
| Default audio | WAV for unary requests | mp3, 44100 Hz, 128000 bit rate |
How do I pick an explicit model on Sume?
Use the TTS Router. Its model field is documented as a required TTS Router catalog model id (pass-through). List what is available with GET /v1/tts-router/models, and read the catalog rather than assuming a Gemini id is present.
What stays the same on Sume?
Voices: a voice.id must be a TTS voice UUID or a Voices library id (voi_ plus 32 hex), or you send an avatar reference. Authentication stays a Sume API key; no provider credentials are sent.
What should I do next?
If you call Gemini directly, follow Google's own migration notes. If you call Sume, re-test a sample after any catalog change. Related: which Gemini TTS engine Sume routes.
Sources
Related posts
More in Developers
- Gemini TTS API polling: the Sume audio job pattern
Gemini 3.8 Flash TTS went GA on Sep 22. On Sume, speech is a job: submit async, store the job id, poll the status URL, and never resubmit while it runs.
- Gemini TTS reads stage directions aloud: Sume emotion field
Gemini 3.8 TTS speaks its text field verbatim, so style goes in speech_metadata. On Sume, keep the transcript clean and put delivery in generation_config.
- Gemini API file size limit: 2GB per file, 20GB storage
Gemini's rate-limits page lists a 2GB input file limit and 20GB file storage under Batch API limits. Sume takes public HTTPS URLs instead of uploaded files.
- Gemini Batch API rate limits vs Sume's 100-run bulk queue
Gemini Batch API allows 100 concurrent batch requests and a 2GB input file. Sume bulk runs queue up to 100 Format runs with concurrency 1 to 16.
Written by Sume