Prime Video AI lip sync vs Sume: human voice on a talking image

Amazon uses AI and VFX to lip-sync shows to human-dubbed audio. Sume cannot re-sync existing footage, but Fabric can animate a still to a human recording.

4 min readSume
All posts

Sume cannot re-lip-sync existing footage to a new dub the way Prime Video does. What Sume can do is animate a still image or avatar to an audio file, including a human recording, with veed/fabric-1.0. Read 2026-10-01 from Amazon's page and the Sume Models docs.

What did Amazon announce?

Amazon says it uses AI and VFX to lip-sync titles to human-dubbed audio. It names Maxton Hall seasons 1 and 2 in English globally, with season 3 on Dec 9, and says more titles are planned.

What is the closest Sume workflow?

POST /v1/veed/fabric-1.0 takes an audio_url on the Sume host (max 10 MB), duration_seconds 1 to 300, resolution 480p or 720p, and exactly one of image_url or avatar_id/avatar_handle. Record a human voice, import it, and Fabric makes the face speak it.

What does it not do?

It does not take a video of a person and change their mouth. The input is a single image or avatar, not footage. Sume's video models do not lip-sync to a TTS file either, per the Models docs.

When is it the right fit?

Spokesperson clips, explainers and localised intros where one still face talks. For a film or series with many cuts and angles, use the vendor tooling built for that.

What are the input limits for Fabric?

Limits from the Sume Models docs, as quoted above.

Fabric 1.0 request limits, read 2026-10-01.
FieldLimit
audio_urlSume host, max 10 MB
duration_seconds1 to 300
resolution480p or 720p
Visual inputExactly one of image_url or avatar_id/avatar_handle

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume