Synthesia Interactive Avatar API vs Sume rendered avatar clips

Synthesia headlines a live Interactive Avatar API. Sume avatar video is script-driven and rendered as a job: submit, poll, then fetch the clip.

4 min readSume
All posts

Synthesia's updates page headlines "Build live avatars into your product with Interactive Avatar API". Sume's avatar API is not that kind of product: POST /v1/avatar-1.0/talking-video is script-driven and rendered, so you submit a script, poll a job, and receive a finished video.

The Synthesia item is a headline only on the page I read, so this post makes no claims about its latency or protocol. Sume details are from Generate avatar video, read 2026-10-01.

What does the Synthesia headline say?

The updates page lists "Build live avatars into your product with Interactive Avatar API" next to a note that Roleplay Sessions now has a plan for every team. Read the API's own documentation before relying on it; the updates page gives no technical detail.

What does a Sume avatar request return?

Avatar videos turn a ready avatar into a talking video from exactly one of script or video_inputs. Scripts and multi-scene plans are accepted when the estimated duration is 4 to 60 seconds inclusive; longer ones should be shortened or split into jobs. You poll job status, events and result. Completed results can include public media.sume.com video artifacts plus public-safe fields such as preview_image_url and scene_previews.

A bounded synchronous wait exists, but its maximum is 30 seconds, and if the job is still queued the response says so; treat the job as the source of truth.

What about turning a still and audio into a clip?

The models page lists VEED Fabric 1.0 at POST /v1/veed/fabric-1.0 for talking still plus audio clips. That is also a rendered job, not a live session.

Live versus rendered, as documented and read 2026-10-01.
NeedFits a rendered Sume job?
Pre-approved scripted explainerYes
Per-scene review before final renderYes, with first-frame previews
Open-ended conversation with a viewerNo; a job returns a finished clip

How should I decide?

If the viewer's input arrives while the avatar is on screen, you need a live product. If the words are known before the video exists, render clips. For the same split on another vendor, see live avatar vs video avatar API.

Sources

Related posts

More in Models

All Models posts

Written by Sume