Vidu API: S2-Avatar live voice vs Sume avatar video jobs

Vidu S2-Avatar is a real-time voice model. Sume has no live session: you submit a script and a scene photo to an avatar job and fetch the finished video.

4 min readSume
All posts

If you searched "Vidu API" after the S2-Avatar launch: Vidu describes a real-time interactive model, and Sume does not ship a live session. Sume's avatar endpoint renders a script into a finished talking video, which you fetch when the job ends.

Vidu facts are from its platform update log; Sume facts are from Generate avatar video and Webhooks, all read 2026-09-30.

What does Vidu say S2-Avatar does?

The September 15, 2026 entry lists "Vidu S2-Avatar: Real-Time Interactive Model". It supports real-time voice interaction and complex motion control, and reference images can be used for product interaction, outfit changes and background replacement. The page says nothing more about request fields, so check Vidu's own docs before you plan around it.

What does Sume do instead?

Sume's docs say avatar videos "turn a ready avatar into a script-driven talking video" through POST /v1/avatar-1.0/talking-video. You send an avatar_handle plus exactly one of script or video_inputs. The result is a file, not a stream, so there is no way to speak to it mid-render.

For a background, set scene to { "type": "photo", "image_url": "https://..." } for a photo reference, or { "type": "prompt", "prompt": "..." } for scene direction. The image must be a fetchable public HTTPS URL.

How do the two compare on the points Vidu lists?

Vidu S2-Avatar as listed on its update page versus Sume avatar video docs, read 2026-09-30.
NeedVidu S2-Avatar (vendor page)Sume avatar video (docs)
Live voice interactionListedNot offered; script in, video out
Background changeReference imagesscene photo or prompt
Product in frameReference imagesOptional product_image
DeliveryReal-time sessionAsync job, status poll or webhook

How do I get the finished video?

Poll the job or use a webhook. Sume sends terminal job events only: job.completed, job.failed and job.canceled, with no progress or partial deliveries. Scripts must land at 4-60 seconds of estimated duration, so split longer ones into several jobs.

When should I pick a live model instead?

If a viewer must talk to the avatar and hear an answer immediately, you need a live session product, and Sume's rendered jobs will not cover that. If you can script the lines ahead of time, a rendered job works; see real-time AI avatar vs video avatar API for the longer comparison.

Sources

Related posts

More in Models

All Models posts

Written by Sume