Sync dubbing in 92 languages vs building a dub on Sume
Sync's built-in dubbing now lists 92 languages with speaker detection. Sume has no dubbing endpoint: chain transcription with a language hint, translation, TTS.

Sync's August 31, 2026 changelog says its built-in dubbing now supports 92 languages, detects speakers automatically and returns lossless audio for the final dub. Sume has no single dubbing endpoint: you assemble transcription with an optional language hint, your own translation step, and text-to-speech.
Sync facts are from its changelog; Sume facts from the OpenAPI document and the Models and Avatar videos docs, read 2026-10-01.
What changed in Sync's dubbing?
The changelog entry says dubs can be created across the App, API, Adobe Premiere Pro and DaVinci Resolve. An earlier entry describes translating a video and lip-syncing it in one step, with translation, voice cloning and lipsync handled by Sync, listed there at 29 languages. The 92 figure is the newer one.
What do I assemble on Sume?
| Step | Sume control from the docs |
|---|---|
| Transcribe | language_code: optional BCP-47 or provider hint such as en or ko; omit for auto-detect |
| Translate | Your own step; Sume docs list no translation endpoint in this flow |
| Speak | TTS with pronunciation_dict_id and generation_config for speed, volume, emotion |
| Picture | A talking-face shot is Fabric with a still plus TTS audio |
What does Sume not do here?
The Models docs say video models do not lip-sync to generated TTS or to a later voice-over. So a dub laid over an existing clip will not have re-synced lips; if you need a face that speaks the new language, generate the speaking shot from a still plus the new audio. Speaker detection is not described in the Sume pages I read, so split speakers yourself before transcribing. The speech transcription rate card lists $0.01 per audio minute.
What about Korean captions?
Avatar video captions have a rule: a Korean script with a Latin-only style such as slam, punch or tiktok-green is rejected with 400 caption_hangul_text_latin_style, so pick a Hangul style for Korean speech. For the full chain see build an AI dubbing pipeline.
Sources
Related posts
More in Use cases
- Synthesia dub speaker attribution vs Sume speech-to-text and audio
Synthesia lets Enterprise users reassign and rename speakers in a dub. Sume gives you audio detach, speech-to-text and audio joins; speaker mapping is yours.
- Synthesia Avatar Builder credits: 14 per option, and Sume jobs
Synthesia Avatar Builder charges 14 credits per generated option. Sume creates an avatar with one job per request via POST /v1/avatar-1.0/generate.
- Synthesia brand kit fonts vs Sume caption fonts: Hangul only
Synthesia Motion Graphics now use brand kit fonts. Sume's caption font field takes a Hangul face only, so Latin brand fonts cannot be set there.
- Synthesia bulk download of 20 videos vs Sume result URLs
Synthesia can bulk download up to 20 videos as .mp4 files. On Sume you list avatar-video resources and fetch each job's result for its media.sume.com URL.
Written by Sume