HeyGen cinematic avatar API: 3 looks per shot vs Sume's one

HeyGen cinematic_avatar takes a prompt and 1-3 avatar looks with no script. Sume Avatar Video resolves one avatar per final video from a script.

4 min readSume
All posts

HeyGen's cinematic avatar API is POST /v3/videos with type: "cinematic_avatar": a prompt of 1 to 10,000 characters plus an array of 1 to 3 avatar look ids, with no script or voice. Sume's Avatar Video works differently: it is script-driven and the docs say current execution supports one resolved avatar per final video.

What does the HeyGen cinematic avatar request take?

Per the HeyGen page read 2026-10-01, prompt is the creative brief for the shot, avatar_id is an array of 1 to 3 look ids so more than one avatar can feature in the same shot, and optional references steer style or motion. Looks and references share a budget of at most 3 videos and 9 images. The example sets aspect_ratio, resolution and duration.

What does Sume Avatar Video take instead?

Requests to POST /v1/avatar-1.0/talking-video provide exactly one of script or video_inputs, plus a ready avatar by avatar_handle. Scene direction goes in scene: { "type": "prompt", "prompt": "..." }. The docs state: "Current execution supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene." See Generate avatar video.

How do the two compare?

Shot-level inputs as documented, read 2026-10-01.
InputHeyGen cinematic_avatarSume Avatar Video
Driverprompt (1-10,000 chars)script or video_inputs
Avatars per request1 to 3 looksOne resolved avatar per final video
VoiceNone neededSpoken scenes use type: "text"
Scene directionIn the promptscene prompt or photo

How do I get two characters in one Sume video?

Not in one avatar video. Render each speaker as its own job and join the results; two-host podcast with one avatar per video shows the split.

Sources

Related posts

More in Models

All Models posts

Written by Sume