Hume Octave 2 description field vs Sume TTS emotion controls
Hume lists a voice description field; acting instructions are coming soon. Sume TTS 1.0 has generation_config: emotion text, speed 0.6 to 1.5, volume 0.5 to 2.

Hume's overview says Octave 2 (preview) supports a description field for voice customization and lists "Acting instructions" as coming soon. Sume TTS 1.0 already has a small set of controls under generation_config: an emotion string of up to 64 characters, speed from 0.6 to 1.5 and volume from 0.5 to 2.
Hume's page is the TTS overview, read 2026-10-01; Sume's controls are from the OpenAPI schema behind the API reference.
How do the controls differ?
One is free text about the voice; the other is a short guide plus two numbers.
| Control | Hume Octave 2 (preview) | Sume TTS 1.0 |
|---|---|---|
| Describe the voice | description field | Not in the request; pick a voice id or avatar |
| Direct the performance | Acting instructions, coming soon | emotion, up to 64 characters |
| Pace | Not listed in this overview | speed 0.6 to 1.5 |
| Loudness | Not listed in this overview | volume 0.5 to 2 |
What does the Sume emotion field do?
The schema calls it an "optional emotion guide for generation". It is one short phrase, not a script of stage directions, and the docs do not list an allowed vocabulary. Try a few phrases on a sample and keep the ones you like.
What does a request with all three look like?
Controls sit inside generation_config.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"transcript": "We did it. We actually did it.",
"avatar_handle": "@narrator",
"generation_config": { "emotion": "relieved and happy", "speed": 0.95, "volume": 1.2 }
}'Should I wait for Hume's acting instructions?
Only if the performance control is what you need and you can use Hume. The page gives no date for the feature, so do not plan around it. For what Sume ships today, see TTS with emotion, speed and volume.
Sources
Related posts
More in Developers
- Hume Octave takes 5,000 characters per utterance in MP3, WAV or PCM
Hume's TTS overview lists 5,000 characters per utterance and MP3, WAV or PCM output. Sume TTS 1.0 takes 20,000 characters and mp3, wav or raw containers.
- Ideogram 4 describe image to JSON prompt vs Sume reference input
Ideogram's describe endpoint returns a structured json_prompt with optional bounding boxes. The Sume docs list no such route; pass the image as a reference.
- Ideogram ad localizer API: exact_copy vs a Sume edit prompt
Ideogram's ad-localizer takes one language per call and an exact_copy mapping. On Sume you send the ad as a reference and spell out the copy in the prompt.
- Ideogram async generation_id polling vs Sume's 202 job envelope
Ideogram returns images directly unless async or webhook_url is set, then you poll /v2/generations. Sume blocks up to 30s, then returns a 202 job.
Written by Sume