Prime Video AI lip sync vs Sume: human voice on a talking image
Amazon uses AI and VFX to lip-sync shows to human-dubbed audio. Sume cannot re-sync existing footage, but Fabric can animate a still to a human recording.

Sume cannot re-lip-sync existing footage to a new dub the way Prime Video does. What Sume can do is animate a still image or avatar to an audio file, including a human recording, with veed/fabric-1.0. Read 2026-10-01 from Amazon's page and the Sume Models docs.
What did Amazon announce?
Amazon says it uses AI and VFX to lip-sync titles to human-dubbed audio. It names Maxton Hall seasons 1 and 2 in English globally, with season 3 on Dec 9, and says more titles are planned.
What is the closest Sume workflow?
POST /v1/veed/fabric-1.0 takes an audio_url on the Sume host (max 10 MB), duration_seconds 1 to 300, resolution 480p or 720p, and exactly one of image_url or avatar_id/avatar_handle. Record a human voice, import it, and Fabric makes the face speak it.
What does it not do?
It does not take a video of a person and change their mouth. The input is a single image or avatar, not footage. Sume's video models do not lip-sync to a TTS file either, per the Models docs.
When is it the right fit?
Spokesperson clips, explainers and localised intros where one still face talks. For a film or series with many cuts and angles, use the vendor tooling built for that.
What are the input limits for Fabric?
Limits from the Sume Models docs, as quoted above.
| Field | Limit |
|---|---|
| audio_url | Sume host, max 10 MB |
| duration_seconds | 1 to 300 |
| resolution | 480p or 720p |
| Visual input | Exactly one of image_url or avatar_id/avatar_handle |
Sources
Related posts
More in Sume Avatar 1.0
- Synthesia Roleplay Sessions vs Sume scripted scenario clips
Synthesia launched Roleplay and Survey Sessions on Oct 1. Sume has no live roleplay, but can render the scenario prompt as a short avatar clip.
- Synthesia Survey Sessions vs Sume: not interactive, but scripted
Synthesia's Survey Sessions avatar asks questions live and reports themes. Sume renders avatar question clips from a script, with no live conversation.
- VEED Fabric Emotions audio tags vs Sume Fabric audio_url
VEED's Fabric Emotions adds inline tags like [excited] in the script. On Sume, Fabric takes a finished audio_url; emotion lives in the TTS step.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
Written by Sume