Speech to text custom vocabulary: Sume STT takes a language hint only

Gemini 3.5 Transcribe biases up to 1,000 custom terms. Sume STT has no vocabulary field, only a language_code hint, so fix names after transcription.

4 min readSume
All posts

Custom vocabulary for speech to text means passing a list of terms so the model spells names and jargon correctly. Gemini 3.5 Transcribe supports up to 1,000 such terms; Sume STT (sume/stt-1.0) has no vocabulary field. Its request accepts audio_url, language_code, duration_seconds, segmentation, metadata and the job-mode fields (mode, webhook_url, wait_timeout_seconds), and rejects anything else.

Vendor facts are from the Gemini model page, read 2026-09-30; Sume facts from the OpenAPI document.

How does Gemini describe custom vocabulary?

The model page lists custom vocabulary biasing at up to 1,000 terms, adding that customers typically see best results with up to 100. It is incompatible with diarization and with word-level timestamps, so you choose between biasing and those features on a single request. OpenAI's speech-to-text guide is in the same area; its quickstart model, gpt-transcribe, is the one the guide recommends for recorded speech.

What does Sume STT accept?

The request schema is closed (additionalProperties: false), so a vocabulary or keywords key would fail validation. The one language control is language_code, described as an optional BCP-47 or provider language hint; omit it for auto-detect. Word timings are always returned.

Term-biasing options, read 2026-09-30.
OptionGemini 3.5 TranscribeSume STT 1.0
Custom term listUp to 1,000 termsNot accepted
Language hintAuto-detects 85+ languageslanguage_code or auto-detect
Combined with word timingsIncompatibleTimings always returned

How do I get product names right on Sume?

Set language_code when you know the language, then correct terms after transcription with a replacement list you maintain, matched against text or against words[] so the timings stay valid. See language hint versus auto-detect.

Is there a vocabulary control on the text-to-speech side?

Yes, in the other direction: the TTS request has an optional pronunciation_dict_id, described as an optional pronunciation dictionary id. That changes how words are spoken, not how they are recognised. See text-to-speech pronunciation.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume