ElevenLabs Eleven v4 API: what it is and how to reach it

Eleven v4 and v4 Turbo launched September 28, 2026 on ElevenLabs' own API. What ElevenLabs says they do, and which TTS ids Sume's router lists.

4 min readSume
All posts

Eleven v4 is ElevenLabs' new text-to-speech model, announced September 28, 2026, together with a low-latency variant, Eleven v4 Turbo. ElevenLabs says both are available in ElevenAgents, ElevenCreative and through ElevenAPI, its own API. Sume's TTS Router does not list an Eleven model: its model field accepts Sonic ids only.

The Eleven facts are from ElevenLabs' launch post and models page, read 2026-09-29. The Sume facts are from the TTS Router in the Sume API reference.

What does ElevenLabs say Eleven v4 does?

These are ElevenLabs' own descriptions, not measurements we ran.

From ElevenLabs' launch post and models page, read 2026-09-29.
TopicWhat ElevenLabs says
ModelsEleven v4 and its low-latency variant, Eleven v4 Turbo
LanguagesBoth support more than 90 languages
Delivery controlDescribe delivery in natural language or use inline tags such as [laughs] or [said angrily in French accent]
PronunciationSupport for International Phonetic Alphabet (IPA) phonemes is improved
Voice cloningInstant Voice Clones can capture a voice from 10 seconds of audio; Professional Voice Clones are supported
WhereElevenAgents, ElevenCreative and ElevenAPI

Which text-to-speech models does Sume list?

POST /v1/tts-router/generate requires a model that is one of five ids: sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview (the last is described in the catalog as a provider beta channel). An id outside that list fails with 400 model_not_found. Sume's TTS 1.0 route has no engine picker and rejects model with a 400.

So a request for Eleven v4 through Sume does not match the schema. For the Sonic side of the same question, see Text to speech API and Multilingual text to speech API.

How do delivery tags compare with Sume's controls?

ElevenLabs describes inline tags as a way to steer delivery inside the text. Sume's TTS schema documents no inline tags or SSML. It has an optional generation_config with speed, volume and an emotion string, and a pronunciation_dict_id field for custom pronunciations. Text to speech pronunciation covers the dictionary.

Does ElevenLabs say where each model runs?

ElevenLabs' models page says Eleven v4 is available via its Text to Dialogue API and Eleven v4 Turbo via the Text to Dialogue websocket. It describes v4 for character voiceovers, emotional dialogue, audiobooks and multilingual projects, and Turbo for support agents, AI assistants and interactive characters. Follow ElevenLabs' documentation for model ids, request shapes and pricing; none of those are Sume facts, and we did not test them.

What should I check before I choose a model?

  • ElevenLabs' rankings and listener-preference figures are its own claims from its own tests. Listen to your own script in each model you can reach.
  • Latency figures in ElevenLabs' posts are its medians for Eleven v4 Turbo; measure with your own text and region.
  • If your pipeline is Sume Timeline or Avatar, check the voice you need is one Sume's TTS routes serve before you plan around another vendor's model.

Sources

Related posts

More in Models

All Models posts

Written by Sume