Synthesia animated scene preview vs Sume first-frame stills
Synthesia previews one scene with a fully animated avatar. Sume previews are first-frame stills with no motion, approved before the full render.
Synthesia's update lets you preview one scene with a fully animated avatar from the editor. Sume's preview does not animate: POST /v1/avatar-video-previews creates a first-frame still, so you can approve composition before starting the full talking-video render. Motion and lip movement are judged only on the final render.
Synthesia's wording is from its updates page; Sume's from Avatar video previews, read 2026-10-01.
What does the Synthesia preview do?
The updates page describes previewing an avatar's animations for a single scene from the editor, so you can review and iterate on avatar performance without generating or re-generating the whole video.
What does a Sume preview show?
The docs say previews generate the first-frame still stage without starting the full talking-video render. For multi-scene video_inputs you get one still per scene. Captions are stored on create and applied only at generate-video; preview stills are never caption-burned.
| Question | Synthesia scene preview | Sume preview |
|---|---|---|
| Shows avatar animation | Yes, per its update | No, a still |
| Unit | One scene | One still per scene |
| Captions in preview | Not stated | Never burned into stills |
How do I go from preview to final video?
Four routes cover the flow: create POST /v1/avatar-video-previews, read GET /v1/avatar-video-previews/:id, redo stills with POST .../regenerate, then POST .../generate-video. The create body matches Avatar Video fields, with exactly one of script or video_inputs; quality defaults to plus. See first-frame previews for avatar video.
When is a still enough?
Use it to check framing, product placement and scene setup. If what you need to review is gesture or speech performance, a still cannot show it; that needs the rendered clip. For a direct render without the preview stage, use Generate avatar video.
Sources
Related posts
More in Use cases
- Synthesia avatar sharing vs Sume workspace-scoped avatars
Synthesia shares an avatar's base person with teammates or a workspace. Sume API keys are workspace-scoped, and avatars are referenced by handle.
- TikTok Ad Network asset specs: banner 640x100 and where Sume stops
TikTok Ad Network wants video at 1280x720, 720x1280 or 720x720 and banners at 600x500, 640x200 or 640x100. What Sume renders natively and what needs a crop.
- TikTok carousel ad specs: 2-35 images and three sizes
A standard TikTok carousel takes 2 to 35 JPG or PNG images at 1200x628, 640x640 or 720x1280. Here is how to map those onto Sume image settings.
- TikTok carousel ad music: mp3 required, silent on Ad Network
TikTok carousel ads need an .mp3 of 2 s or more, up to 10M. Library music is silent on Ad Network placements, so upload your own file there.
Written by Sume