TTS API for video narration: script to voice track to MP4

Narrate a video with Sume's TTS API: send up to 20,000 characters, get a voice file, then lay it under clips with Timeline 1.0.

4 min readSume
All posts

For video narration, call Sume's TTS API with your script to get a voice file, then use that file as the audio spine of a Timeline 1.0 render that places your clips over it. POST /v1/tts-1.0/generate takes up to 20,000 characters and returns a Sume-hosted file; POST /v1/timeline-1.0/render returns one MP4.

TTS facts are from the Sume API reference and Timeline facts from Timeline 1.0, read 2026-09-29. TTS is $0.0475 per 1,000 characters plus a 5.5% agent fee by default.

How do I make the narration file?

Pick a voice: the avatar_handle of an avatar whose voice is ready, or a voice.id you hold. Ask for WAV if the file will be joined or re-cut; the reference says to pass wav/pcm_s16le/44100 for that use.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: narration-001" \
  -d '{
    "transcript": "Every morning starts with the same small ritual. This is ours.",
    "avatar_handle": "narrator",
    "language": "en",
    "output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 },
    "timestamps": { "words": true }
  }'

How do I put the voice under the clips?

Give Timeline 1.0 the narration as audio.url, its length as audio.duration_seconds, and the clips as video[] slots with on-spine start times. Every URL must be a media.sume.com file from your workspace; import outside files first. Word timings from the TTS job tell you where each sentence starts, so you can time each clip to its line.

From Timeline 1.0 and the TTS reference, read 2026-09-29.
NeedField
Narration as the spineaudio.url and audio.duration_seconds (1-1800)
Clip placementvideo[].source_url, start, duration
Music under the voicesoundtrack with gain_db, loop, duck_db
OutputOne MP4, 1080x1920 by default

What are the limits?

  • TTS audio over 1,200 seconds fails with tts_duration_exceeded; split a long script.
  • Set language for any non-English script.
  • Timeline 1.0 is billed per output minute at the rate on API pricing.
  • The TTS route is not streaming; poll or use a webhook.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume