Replace video audio with an AI voice (and the lip-sync catch)
Sume has no one-call audio swap: detach the audio, transcribe it, make a new voice with TTS, and lay it on a Timeline. Lips will not re-sync to the new voice.

You can replace a video's audio with an AI voice on Sume by chaining four existing routes: detach the original audio, transcribe it, generate the new voice with text-to-speech, and put that file on a Timeline over the original picture. It is a pipeline you run, not a single "convert" button, and the mouth in the picture will not re-sync to the new voice.
Descript's changelog of 2026-09-17 describes a "Convert to AI Speech" action that reads the same script back in a voice you pick. Sume's equivalent is built from the parts documented in Audio detach and Timeline 1.0.
Which steps does the swap take?
Each step is its own job, so keep the output of one as the input of the next.
| Step | Route | Notes |
|---|---|---|
| 1. Detach audio | POST /v1/audio-detach | Takes one workspace media.sume.com video; the video is untouched |
| 2. Transcript | POST /v1/stt-1.0/transcribe | Speech-to-text, $0.01 per audio minute |
| 3. New voice | POST /v1/tts-1.0/generate | Text-to-speech, $0.0475 per 1,000 characters |
| 4. Combine | Timeline 1.0 | audio.url is one Sume-hosted spine; render $0.10 per output minute |
Do I need to detach the audio first?
You need the words, and detaching gives a wav that speech-to-text accepts. The detach docs say the default output is the format that timeline_create audio.url and speech-to-text want. If you already have the script, skip steps 1 and 2 and send it straight to text-to-speech. See detach for speech-to-text for the format options.
Where does the new voice go?
Into the Timeline program. audio.url takes one Sume-hosted spine, and audio.duration_seconds (required, 1 to 1800 seconds) sets the output length. Put your original clip in video[].source_url. The hosted MCP server lists tts_create and stt_create for the same flow, as described in MCP tools and gates.
Will the lips match the new voice?
No. Sume's model overview says video models do not lip-sync to generated TTS or to a later voice-over. If the person's mouth is visible and talking, a replaced voice will drift from it, so this suits footage where lips are off camera, screen recordings, or B-roll. For a speaking face, the docs route it through an avatar workflow instead; read lip sync to a voiceover for the limits.
Sources
Related posts
More in Use cases
- AI ad resizer: one image, several aspect ratios, one API
Sume has no resizer button. Send the same source image to POST /v1/images once per aspect_ratio (1:1, 4:5, 9:16, 16:9) and recompose each ad slot.
- Seedance 2.5 draft mode: a 480p-then-1080p loop on Sume
Higgsfield added Seedance 2.5 Draft Mode. On Sume there is no draft model: render at resolution 480p, pick a take, then re-request at 1080p.
- Seedream 5.0 pro bbox and point tags vs region edits on Sume
Seedream 5.0 pro edits a region from <bbox> and <point> tags on a 0-999 grid. Sume lists other Seedream ids; use mask_url on ChatGPT Image 2.5 for regions.
- Shopify's recommended 2048 x 2048 product image with gpt-image-2.5
Shopify.dev recommends 2048 x 2048 px for product images. That is a valid custom image_size on Sume's gpt-image-2.5: 2048 is a multiple of 16 and 4,194,304 px.
Written by Sume