Why AI video models do not lip sync to a voiceover, and what to use
Sume docs: video models do not lip-sync to generated speech or a voice-over. For a talking face use Avatar video or the lip-sync endpoint (still plus audio).

A text-to-video or image-to-video model does not lip sync to a voiceover you add afterward, and Sume's models docs say so directly: video models do not lip-sync to generated TTS or to a later voice-over, so a talking face is never a video-model clip with narration laid underneath. For speech, use a path built for it: Avatar video from a script, or the lip-sync endpoint that takes a still and an audio clip.
This is from the Models overview, read 2026-09-29.
Which path fits which job?
| Need | Endpoint | Notes |
|---|---|---|
| A presenter reads a script | POST /v1/avatar-1.0/talking-video | One of script or video_inputs; 4 to 60 seconds |
| A photo speaks audio you already have | POST /v1/minimax/h3-max/lip-sync | A still plus audio, 5 to 14.8 seconds, list price × 1.25 |
| Wordless motion, B-roll, product shots | POST /v1/videos | No lip sync; generate_audio for native sound |
| Your own clip, new face | POST /v1/models/sume/avatar-face-swap/v1.0/runs | Beta; source about 4 to 15 seconds with usable audio |
What about video models that generate audio?
Several catalog models can generate audio with generate_audio, and some accept audio references. That is native sound made with the video. It is not the same as syncing lips to a specific recording you supply, and the docs do not promise lip sync from it. Check each model's generate_audio and supported_input_references at GET /v1/videos/models.
How do I put speech under B-roll then?
Cut B-roll for the parts where no one is on camera and let the avatar or lip-sync clip carry the on-camera lines. Join them with a Timeline render, using the speech as the audio spine. See talking head vs B-roll.
Where do I read the details?
Lip sync from a photo and audio covers the still-plus-audio call. Avatar, lip sync or motion control helps you pick one.
How do I check what a video model accepts?
GET /v1/videos/models lists each model with these fields.
| Field | Tells you |
|---|---|
generate_audio | Whether it can generate a soundtrack |
supported_input_references | Which of image_url, video_url, audio_url it takes |
supported_durations | Lengths in whole seconds |
supported_resolutions, supported_aspect_ratios | Frame options |
Sources
Related posts
More in Models
- Batch edit photos with AI: one edit across many photos
To batch edit photos with AI, send the same edit prompt once per photo, with that photo as the reference. How to script it, pace it, and what it costs.
- AI outfit change video: change clothes in a clip with AI
Change someone's outfit in an existing video with an AI video-to-video edit: name the garment, what it becomes, and what must stay the same.
- Claude Sonnet 5.5 or Opus 5.5 for an agent that calls a video tool?
Sonnet 5.5 and Opus 5.5 differ in price and release date on Anthropic's pages. Which facts matter when the tool is Sume's video generation, and what stays same?
- How to colorize a black and white video with AI
Colorize a black and white video with an AI video-to-video edit: send the clip and a prompt naming the colors. The model invents them.
Written by Sume