Deepgram Aura-2 languages: seven, and how Sume TTS sets a language

Deepgram lists English, Spanish, German, French, Dutch, Italian and Japanese for Aura. Sume TTS takes a language field per request, matched to the voice.

4 min readSume
All posts

Deepgram's TTS models page lists seven Aura languages: English, Spanish, German, French, Dutch, Italian and Japanese. Sume TTS does not publish a fixed language list in its request schema. It takes a language field (2 to 16 characters, BCP-47 or ISO-639) that must match the voice you picked, and warns before a charge if it does not.

Deepgram's list is from its models page, read 2026-10-01; Sume's from the OpenAPI schema behind the API reference.

How many voices does each Aura language have?

The page gives a voice count per language.

Aura voices per language, from Deepgram's models page, read 2026-10-01.
LanguageVoices
English (en)37
Spanish (es)18
Dutch (nl)9
Italian (it)9
German (de)7
Japanese (ja)5
French (fr)2

How does Sume decide which language is spoken?

From the language field. The schema says to set it for every non-English transcript: if omitted it defaults to English at the provider, though Sume infers Korean or Japanese from a transcript that is only Hangul or kana. It also says never to translate a non-English request into English.

What happens if the voice and language disagree?

A known mismatch returns 409 tts_voice_language_mismatch before any job or charge. Retry with confirm_language_mismatch: true only after you have checked the voice, transcript and language. Regional tags compare by primary language, so en-GB against an English voice does not trip it.

Which one should I choose for Japanese or Dutch?

Both appear workable on paper, but a language list is not an audio test. Deepgram names Dutch and Japanese; on Sume, pass the language and listen to the result. For the Sonic side of that question, see Sonic 3.6 languages.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume