Synthesia Roleplay Sessions vs Sume scripted scenario clips

Synthesia launched Roleplay and Survey Sessions on Oct 1. Sume has no live roleplay, but can render the scenario prompt as a short avatar clip.

4 min readSume
All posts

Sume has no live roleplay avatar. It can render the customer or manager lines of a training scenario as avatar clips of 4 to 60 seconds. Synthesia's Sessions page, read 2026-10-01, lists Roleplay and Survey Sessions as live.

What can I render?

Each scenario line becomes one POST /v1/avatar-1.0/talking-video job with a script. One avatar speaks per final video, and silence beats are available with voice.type silence. Aspect ratios include 16:9 and 9:16.

Scenario clips on Sume, from the Avatar videos docs, read 2026-10-01.
ItemSume avatar clip
EndpointPOST /v1/avatar-1.0/talking-video
Clip length4 to 60 seconds
Speakers per final videoOne avatar
Silence beatsvoice.type silence
Aspect ratiosInclude 16:9 and 9:16

How does a learner respond?

Your app plays a clip, records the learner and sends the audio to Sume speech-to-text. The branching and feedback are your code. Sume does not score the answer.

Which should I pick?

Pick Synthesia Sessions for a live back-and-forth. Pick Sume when you want pre-rendered scenarios you can store, version and ship in your own LMS.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume