Synthesia dub speaker attribution vs Sume speech-to-text and audio
Synthesia lets Enterprise users reassign and rename speakers in a dub. Sume gives you audio detach, speech-to-text and audio joins; speaker mapping is yours.

Synthesia now lets Enterprise users reassign, create and rename speakers in a dubbed video with one click. Sume has no speaker editor in the pages read here: it gives you audio detach, a speech-to-text route and audio joins, and the mapping of who said what stays in your own code.
Synthesia facts are from its updates page (entry dated 9/3/2026); Sume facts from Audio detach and Timeline audio, read 2026-10-01.
What did Synthesia change?
The entry says you can reassign, create and rename speakers in a dubbed video with one click, available on Enterprise plans. It is an in-product edit.
Which Sume pieces touch a dub?
| Step | Surface |
|---|---|
| Pull the audio from one video | Audio detach: one workspace media.sume.com video in, a new audio artifact out |
| Transcribe | POST /v1/stt-1.0/transcribe |
| Join audio for one render | Timeline 1.0 audio.parts[] |
Why detach first?
Detach returns a sample-exact wav by default, which the docs name as what Timeline audio and speech-to-text want. The source video is untouched. Poll GET /v1/jobs/:id/status and read GET /v1/jobs/:id/result; there is no per-resource GET for a detach.
Where do speakers get assigned?
The Sume pages cited here do not describe a speaker-attribution editor, so treat it as your own step: keep a table of time ranges and speaker names next to the transcript, and rebuild the audio from those ranges when a mapping changes. For a worked pipeline see build an AI dubbing pipeline.
What should I check before relying on this?
Read the speech-to-text docs for the exact response fields; this post only cites the route. If you need a one-click speaker editor inside the dub, that is what Synthesia describes.
Sources
Related posts
More in Use cases
- Synthesia Avatar Builder credits: 14 per option, and Sume jobs
Synthesia Avatar Builder charges 14 credits per generated option. Sume creates an avatar with one job per request via POST /v1/avatar-1.0/generate.
- Synthesia brand kit fonts vs Sume caption fonts: Hangul only
Synthesia Motion Graphics now use brand kit fonts. Sume's caption font field takes a Hangul face only, so Latin brand fonts cannot be set there.
- Synthesia bulk download of 20 videos vs Sume result URLs
Synthesia can bulk download up to 20 videos as .mp4 files. On Sume you list avatar-video resources and fetch each job's result for its media.sume.com URL.
- Synthesia CSV bulk personalization vs Sume bulk runs (1-100)
Synthesia maps CSV or XLSX columns to template variables. Sume bulk runs take a JSON array of 1 to 100 Format runs with a concurrency window of 1 to 16.
Written by Sume