Synthesia dubbing file size limit: 5 GB or 2.5 hours

Synthesia dubbing accepts uploads up to 5 GB or 2.5 hours, 4K, in .mp4, .webm or .mov. Sume's limits sit at the audio stage: a 10 MB audio_url and 300 s.

4 min readSume
All posts

Synthesia's dubbing page says an uploaded file can be up to 5 GB or 2.5 hours, at up to 4K (3840x2160), in .mp4, .webm or .mov. A YouTube link has no duration limit. Sume does not publish a whole-video dubbing upload limit; the limits that matter in a Sume pipeline are per stage, such as a 10 MB audio_url and a 300 second duration_seconds on the Fabric endpoint.

Synthesia numbers are from its dubbing page, read 2026-10-01. Sume numbers are from the OpenAPI document and Avatar videos.

What are the Synthesia dubbing limits?

The upload row also lists frame rates of 23.98 to 60 fps and audio sample rates of 8 to 96 kHz. The Dubbing page can dub up to 10 videos in bulk. Under advanced options, a duration setting of Adaptive (the default) changes playback speed to fit the translation, while Original keeps the video speed and adjusts only the voiceover.

Input limits as written in each source, read 2026-10-01.
WhereLimit
Synthesia uploadUp to 5 GB or 2.5 hours, up to 4K
Synthesia YouTube linkNo duration limit
Synthesia bulk dubbingUp to 10 videos
Sume Fabric audio_urlSume-hosted public HTTPS URL, max 10 MB
Sume Fabric duration_seconds1 to 300
Sume Avatar VideoEstimated duration 4-60 seconds

Why are the numbers so different?

They bound different things. Synthesia's numbers cap the source video you hand over. Sume's Fabric endpoint takes a still image plus an audio file, and the OpenAPI schema says non-Sume hosts are rejected for audio_url, so audio must already live on the Sume media host. Compare against a TTS segment, not a feature film.

Can Sume put new speech on an existing video?

The models docs say video models do not lip-sync to generated TTS or to a later voice-over, and that every on-camera speaking shot is Fabric with an accepted still plus TTS. So a Sume talking face is generated from a still and audio, not re-voiced from a source video. For a speech-only pipeline, see build an AI dubbing pipeline with STT, translate and TTS.

What should I do with a long source?

Cut the audio into segments that each stay under 10 MB and 300 seconds, upload each to the Sume media host, and make one request per segment. If you only need a Synthesia-sized file dubbed as is, that is a question for Synthesia's page, not Sume's.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume