Blog

Sume Avatar 1.0

Sume Avatar 1.0 by API: reusable avatars, talking video from a script, multi-scene and two-speaker videos, previews, B-roll, and Face Swap (Beta).

Start with: Introducing Sume Avatar 1.0

Talking head video API: avatar vs lip sync vs motion control

Pick a Sume talking-head route by input: a script and a ready avatar, your own audio and a still, or a driving video. Limits and per-second prices.

How to make an AI avatar video longer than 60 seconds

One Sume avatar video job covers an estimated 4-60 seconds. Split the script, render each part, and join the parts with Timeline 1.0 for a longer video.

Add B-roll to a talking-head avatar video with the Sume API

Keep an Avatar 1.0 clip's voice as one Timeline spine, cut away to B-roll slots, and cut back with source_in equal to start so the lips stay on the words.

AI avatar conversation video API: two speakers in one dialogue

A Sume avatar video job resolves one avatar, so a two-person dialogue is one job per turn, cut together in speaking order with Timeline 1.0.

AI avatar video with your product and background via API

Add product_image and a scene prompt or photo to a Sume Avatar 1.0 talking video, then check the first frame in a preview before the full render.

Stock AI avatars API: find a ready-made avatar for a talking video

Search Sume's avatar catalog with POST /v1/avatar-catalog/search, pick a ready public avatar, and pass its handle to a talking video without creating one.

How to create a reusable AI avatar with the Sume Avatar 1.0 API

Send POST /v1/avatar-1.0/generate with an avatar_handle and a prompt, profile, or image input. Poll the job, then reuse the handle for avatar videos.

Talking avatar video API: generate an avatar video from a script

POST /v1/avatar-1.0/talking-video turns a ready avatar and a 4-60 second script into a talking video. Options, captions, polling, and per-second rates.

Multi-scene avatar video API: build one video from ordered scenes

Send ordered video_inputs instead of one script to compose spoken and silent scenes into one 4-60 second avatar video. Scene fields, rules, and limits.

Avatar video previews: approve the first frame before rendering

Create an avatar video preview to get first-frame stills, regenerate them if needed, then call generate-video on the preview id to render the final video.

Avatar Face Swap API (Beta): apply an avatar face to a video

Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.

Introducing Sume Avatar 1.0

Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.