HeyGen ElevenLabs v3 model_id per request vs Sume TTS
HeyGen's speech endpoint takes settings.model_id such as eleven_v3 per request. Sume TTS 1.0 hides model ids; its separate TTS Router lists models explicitly.

HeyGen's POST /v3/voices/speech now accepts settings.model_id with engine: "elevenlabs", so one request can pick eleven_v4, eleven_v3 or another listed model. Sume's TTS 1.0 does the opposite: its tool contract forbids provider model ids and any model picker. A separate TTS Router lists explicit models, and its v1 is Sonic only.
HeyGen facts are from its September 2026 changelog; Sume facts are from the OpenAPI description of the TTS Router and the TTS tool description in the Sume codebase, read 2026-09-30. See also the Public API reference.
What does HeyGen's model_id do?
The changelog lists eleven_v4, eleven_v3, eleven_multilingual_v2, eleven_flash_v2, eleven_flash_v2_5 and eleven_turbo_v2_5. The setting overrides the voice's saved model preference for that request only; omit settings to keep the preference. The older elevenlabs_v3 engine is deprecated.
How does Sume choose a speech model?
For TTS 1.0 (POST /v1/tts-1.0/generate, public model id sume/tts-1.0), you pick a voice, not a model. The tool description says there is no provider model id, credentials or model/model_id picker, and a request over 1200 seconds fails with tts_duration_exceeded. The TTS Router is a catalog of explicit pass-through models, invoked with the model in the body.
| Route | Model choice | Notes |
|---|---|---|
HeyGen /v3/voices/speech | settings.model_id per request | ElevenLabs engine only |
| Sume TTS 1.0 | None; choose a voice | Separate from the router |
| Sume TTS Router | Model in the request body | v1 is Sonic only |
Can I get an ElevenLabs model on Sume?
Not from what the docs quoted here show: the router's v1 catalog is Sonic only, and TTS 1.0 exposes no model choice. If a pipeline depends on a named ElevenLabs model, keep that call outside Sume or adjust the voice and script instead.
What should I change when porting a speech call?
Drop model_id, send the transcript with a voice id, and read results by job. For the Sonic route, see ElevenLabs API alternative: text to speech with Sonic.
Sources
Related posts
More in Models
- HeyGen photo avatar hand gestures vs Sume Motion Control
HeyGen's motion_prompt still needs an animation reference for Avatar V photo avatars. On Sume, body motion is a separate Motion Control route, not a prompt.
- HeyGen text to video API: 5-15 s, 768p vs Sume durations
heygen-video-1 makes 5-15 second clips at 480p or 768p from a 5,000-character prompt. Sume lists durations and resolutions per model in its catalog.
- HeyGen reference to video: 9 images, 3 videos, 3 audio
HeyGen's heygen-video-1 reference mode takes 9 images, 3 videos and 3 audio files, 12 in total. Here is how Sume's reference limits compare.
- Hy-Image-3.5-Preview API limits vs Sume's image limits
Tencent's Hy-Image-3.5-Preview takes 256 to 8192 px edges, up to 4096x4096 and 20 references. Sume lists no Hy Image model; its own limits are below.
Written by Sume