Why AI video models do not lip sync to a voiceover, and what to use

Sume docs: video models do not lip-sync to generated speech or a voice-over. For a talking face use Avatar video or the lip-sync endpoint (still plus audio).

4 min readSume
All posts

A text-to-video or image-to-video model does not lip sync to a voiceover you add afterward, and Sume's models docs say so directly: video models do not lip-sync to generated TTS or to a later voice-over, so a talking face is never a video-model clip with narration laid underneath. For speech, use a path built for it: Avatar video from a script, or the lip-sync endpoint that takes a still and an audio clip.

This is from the Models overview, read 2026-09-29.

Which path fits which job?

Speaking-face options, from the Sume docs cited below, read 2026-09-29.
NeedEndpointNotes
A presenter reads a scriptPOST /v1/avatar-1.0/talking-videoOne of script or video_inputs; 4 to 60 seconds
A photo speaks audio you already havePOST /v1/minimax/h3-max/lip-syncA still plus audio, 5 to 14.8 seconds, list price × 1.25
Wordless motion, B-roll, product shotsPOST /v1/videosNo lip sync; generate_audio for native sound
Your own clip, new facePOST /v1/models/sume/avatar-face-swap/v1.0/runsBeta; source about 4 to 15 seconds with usable audio

What about video models that generate audio?

Several catalog models can generate audio with generate_audio, and some accept audio references. That is native sound made with the video. It is not the same as syncing lips to a specific recording you supply, and the docs do not promise lip sync from it. Check each model's generate_audio and supported_input_references at GET /v1/videos/models.

How do I put speech under B-roll then?

Cut B-roll for the parts where no one is on camera and let the avatar or lip-sync clip carry the on-camera lines. Join them with a Timeline render, using the speech as the audio spine. See talking head vs B-roll.

Where do I read the details?

Lip sync from a photo and audio covers the still-plus-audio call. Avatar, lip sync or motion control helps you pick one.

How do I check what a video model accepts?

GET /v1/videos/models lists each model with these fields.

Model discovery fields, from the Sume Video generation docs, read 2026-09-29.
FieldTells you
generate_audioWhether it can generate a soundtrack
supported_input_referencesWhich of image_url, video_url, audio_url it takes
supported_durationsLengths in whole seconds
supported_resolutions, supported_aspect_ratiosFrame options

Sources

Related posts

More in Models

All Models posts

Written by Sume