Shortest video an AI dub or face swap accepts: 5 s vs 4 s
Synthesia's dubbing page says a video must be at least 5 seconds. Sume's Beta face swap plans for about 4-15 seconds and avatar videos take 4-60. Side by side.

Synthesia's AI dubbing page says "Video must be at least 5s long". Sume does not offer a dubbing product in the docs read for this post; its nearest video-in tool, Beta face swap, plans for source videos of about 4-15 seconds with usable audio, and avatar videos take a 4-60 second script.
What does Synthesia's dubbing page list?
Besides the 5-second minimum, the page lists "140+ languages and dialects", accepts "MP4, MOV, and WebM files", and says the "first minute of video is completely free". It adds that "Subtitles are auto-generated for every dubbed video". These are Synthesia's statements about its own product.
What length does each Sume video tool take?
The face swap docs say Beta worker validation targets source videos "currently planned for about 4-15 seconds" with usable audio, and the request takes only avatar_handle, video_url and quality. The avatar video docs accept an estimated 4-60 seconds.
| Tool | Shortest | Longest stated |
|---|---|---|
| Synthesia AI dubbing | 5 s | Not stated on the page read |
| Sume face swap (Beta) | About 4 s | About 15 s |
| Sume avatar video | 4 s | 60 s |
Is face swap the same as dubbing?
No. Face swap puts a Sume avatar's face onto your source video; it does not translate speech. The docs list transcripts and prompts among unsupported fields on that endpoint. If you need translated audio, see lip sync vs dubbing for how the pieces differ.
What if my clip is shorter than the minimum?
For face swap, the docs give a planned range of about 4-15 seconds and do not say how a shorter clip is handled. Test with a real sample before batching.
Sources
Related posts
More in Sume Avatar 1.0
- Does a face-swapped video need a YouTube AI disclosure?
YouTube asks for disclosure when content makes a real person appear to say or do something they didn't. What that means for a Sume Beta face-swap output.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar video previews: approve the first frame before rendering
Create an avatar video preview to get first-frame stills, regenerate them if needed, then call generate-video on the preview id to render the final video.
- How to create a reusable AI avatar with the Sume Avatar 1.0 API
Send POST /v1/avatar-1.0/generate with an avatar_handle and a prompt, profile, or image input. Poll the job, then reuse the handle for avatar videos.
Written by Sume