Qwen TTS inline tags vs Sume's emotion and speed controls

Qwen-Audio-3.0-TTS reads tags like [gasp] and [angry] in text. Sume TTS uses generation_config with an emotion guide and speed from 0.6 to 1.5.

4 min readSume
All posts

Qwen-Audio-3.0-TTS accepts natural language such as "say this angrily" or inline tags like [gasp], [giggles] and [angry] in the target text. Sume TTS does not document an inline tag syntax. It exposes a generation_config object with volume, speed and an optional emotion guide.

Qwen facts are from the Tongyi Lab post, read 2026-10-01. Sume facts are from the OpenAPI spec for POST /v1/tts-1.0/generate.

What does generation_config accept?

The spec calls it optional volume, speed and emotion controls. Emotion is a free string of 1 to 64 characters, described as an optional emotion guide for generation. It is not a numeric scale.

Sume TTS generation_config fields from the OpenAPI spec, read 2026-10-01.
FieldType and range
volumeNumber, 0.5 to 2.0 multiplier
speedNumber, 0.6 to 1.5 multiplier
emotionString, 1 to 64 characters, optional guide

Can you put [gasp] in a Sume transcript?

The docs do not say tags like [gasp] are interpreted, so do not rely on them. Test a short sample first. The documented way to steer delivery is emotion and speed.

How do you direct a read in practice?

Keep the transcript plain, put the intent in emotion, and adjust speed in small steps within 0.6 to 1.5. Generate two or three takes, and compare them by ear.

const body = {
  transcript: "I did not expect that.",
  avatar_id: process.env.SUME_AVATAR_ID,
  generation_config: { emotion: "surprised", speed: 1.1 },
};
console.log(JSON.stringify(body));

Where does the audio go next?

The result comes through the job flow in Jobs and results. For ads and short videos, see AI voiceover for ads.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume