TTS API for video narration: script to voice track to MP4
Narrate a video with Sume's TTS API: send up to 20,000 characters, get a voice file, then lay it under clips with Timeline 1.0.

For video narration, call Sume's TTS API with your script to get a voice file, then use that file as the audio spine of a Timeline 1.0 render that places your clips over it. POST /v1/tts-1.0/generate takes up to 20,000 characters and returns a Sume-hosted file; POST /v1/timeline-1.0/render returns one MP4.
TTS facts are from the Sume API reference and Timeline facts from Timeline 1.0, read 2026-09-29. TTS is $0.0475 per 1,000 characters plus a 5.5% agent fee by default.
How do I make the narration file?
Pick a voice: the avatar_handle of an avatar whose voice is ready, or a voice.id you hold. Ask for WAV if the file will be joined or re-cut; the reference says to pass wav/pcm_s16le/44100 for that use.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: narration-001" \
-d '{
"transcript": "Every morning starts with the same small ritual. This is ours.",
"avatar_handle": "narrator",
"language": "en",
"output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 },
"timestamps": { "words": true }
}'How do I put the voice under the clips?
Give Timeline 1.0 the narration as audio.url, its length as audio.duration_seconds, and the clips as video[] slots with on-spine start times. Every URL must be a media.sume.com file from your workspace; import outside files first. Word timings from the TTS job tell you where each sentence starts, so you can time each clip to its line.
| Need | Field |
|---|---|
| Narration as the spine | audio.url and audio.duration_seconds (1-1800) |
| Clip placement | video[].source_url, start, duration |
| Music under the voice | soundtrack with gain_db, loop, duck_db |
| Output | One MP4, 1080x1920 by default |
What are the limits?
- TTS audio over 1,200 seconds fails with
tts_duration_exceeded; split a long script. - Set
languagefor any non-English script. - Timeline 1.0 is billed per output minute at the rate on API pricing.
- The TTS route is not streaming; poll or use a webhook.
Sources
Related posts
More in Use cases
- Remix a viral TikTok video with AI: what Sume does and does not do
Trending search returns TikTok metadata and watch URLs, not video files. To remix, ingest a clip you own, read its shots and text, then generate new footage.
- Virtual twilight: turn a daytime exterior into a dusk photo
Virtual twilight turns a daytime house photo into a dusk shot with a sunset sky and lit windows. How to make one with AI, and what to check after.
- Word document to video AI: turn a .docx into a narrated video
To turn a Word document into a video, extract its text and images first. Sume's agent takes text in input and images as attachments; a .docx isn't one.
- Safety training videos for employees, made with AI
Make safety training videos for employees with an AI presenter: one hazard per clip, your own site photos, captions, and versions in your crew's languages.
Written by Sume