Replace video audio with an AI voice (and the lip-sync catch)

Sume has no one-call audio swap: detach the audio, transcribe it, make a new voice with TTS, and lay it on a Timeline. Lips will not re-sync to the new voice.

4 min readSume
All posts

You can replace a video's audio with an AI voice on Sume by chaining four existing routes: detach the original audio, transcribe it, generate the new voice with text-to-speech, and put that file on a Timeline over the original picture. It is a pipeline you run, not a single "convert" button, and the mouth in the picture will not re-sync to the new voice.

Descript's changelog of 2026-09-17 describes a "Convert to AI Speech" action that reads the same script back in a voice you pick. Sume's equivalent is built from the parts documented in Audio detach and Timeline 1.0.

Which steps does the swap take?

Each step is its own job, so keep the output of one as the input of the next.

Audio-swap pipeline on Sume, from the docs read 2026-09-30.
StepRouteNotes
1. Detach audioPOST /v1/audio-detachTakes one workspace media.sume.com video; the video is untouched
2. TranscriptPOST /v1/stt-1.0/transcribeSpeech-to-text, $0.01 per audio minute
3. New voicePOST /v1/tts-1.0/generateText-to-speech, $0.0475 per 1,000 characters
4. CombineTimeline 1.0audio.url is one Sume-hosted spine; render $0.10 per output minute

Do I need to detach the audio first?

You need the words, and detaching gives a wav that speech-to-text accepts. The detach docs say the default output is the format that timeline_create audio.url and speech-to-text want. If you already have the script, skip steps 1 and 2 and send it straight to text-to-speech. See detach for speech-to-text for the format options.

Where does the new voice go?

Into the Timeline program. audio.url takes one Sume-hosted spine, and audio.duration_seconds (required, 1 to 1800 seconds) sets the output length. Put your original clip in video[].source_url. The hosted MCP server lists tts_create and stt_create for the same flow, as described in MCP tools and gates.

Will the lips match the new voice?

No. Sume's model overview says video models do not lip-sync to generated TTS or to a later voice-over. If the person's mouth is visible and talking, a replaced voice will drift from it, so this suits footage where lips are off camera, screen recordings, or B-roll. For a speaking face, the docs route it through an avatar workflow instead; read lip sync to a voiceover for the limits.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume