AI dubbing rate per minute: a step-by-step cost breakdown
An AI dub's rate per minute is the sum of its steps: transcribe, translate, speak, render. Sume's per-step prices and a worked 1- and 10-minute dub.

An AI dubbing rate per minute is the sum of the steps a dub takes: transcribe the original speech, translate it, speak the translation, and put the new voice under the video. On Sume those steps are metered separately, and with a 900-character translated script per minute a 10-minute dub works out to at most $0.15275 per minute before the 5.5% agent fee, translation not included.
Rates are read from the code behind API pricing. Limits come from the TTS schema in the Sume API reference (the OpenAPI document behind the API reference docs), Video inspect and Timeline 1.0, read on 2026-09-29. Translate a video's voiceover by API has the requests for each step.
What does one minute of AI dubbing cost?
The script length is an assumption: 900 characters of translated text per minute of video. Count your own translated script, because speech is billed per character, spaces and punctuation included. A 1-minute dub comes to at most $0.15275. Probing the clip is free; only its transcript is billed.
| Step | Rate | 1-minute video | 10-minute video |
|---|---|---|---|
Transcribe (video inspect, transcribe: true) | $0.01 per audio minute | $0.01 | $0.1 |
| Translate | Your own step | Not a Sume charge | Not a Sume charge |
| Speak the translation | $0.0475 per 1,000 characters | $0.04275 | $0.4275 |
| Render video over the new voice | Up to $0.10 per output minute | Up to $0.1 | Up to $1 |
| Total | Up to $0.15275 | Up to $1.5275 |
Why is the render an "up to" price?
Timeline 1.0 reserves its rate per started output minute. In the current pricing code the render then captures its own compute cost, never above that reservation, so the row shows the ceiling. Failed jobs release or refund their reservation where applicable.
What limits change the math for long videos?
- The video must already be a file in your workspace on
media.sume.com, such as an earlier Sume job's output: video inspect and Timeline read nothing else. - The transcript reserves by
duration_seconds, at most 600 seconds; omit it and 1 minute is reserved. A clip with no audio track fails withinspect_source_has_no_audio. - One TTS request takes up to 20,000 characters and fails with
tts_duration_exceeded, with no credit captured, past 1,200 seconds of audio. - Set TTS
languagefor every non-English transcript; omitted, it defaults to English. - A Timeline render outputs 1 to 1,800 seconds.
What does this price not include?
- Translation. You translate the lines yourself or with another tool, and that cost sits outside Sume's rates.
- The original music and effects. In the current code a render plays only its audio spine and an optional
soundtrack, so each clip's own audio is dropped. Supply a Sume-hosted bed assoundtrackif you have one. - Lip sync. Sume's model docs say video models do not lip-sync to a later voice-over, so a speaker on camera won't match the new voice; Lip sync vs dubbing explains the difference.
Sources
Related posts
More in Pricing
- 4K AI image: generate at 4K or upscale a 1K image? The cost
Nano Banana can render at a 4K tier on Sume. Image Upscale 1.0 enlarges an existing image. What each costs per image and what each needs as input.
- Why an AI image costs list × 1.25 on Sume: worked examples
Sume's image catalog prices a model at its list price times 1.25, rounded up to a cent. Worked examples for five models and where the number shows.
- Am I charged when an AI image generation fails or is canceled?
On Sume's image API a completed image is billed in full and a failed or canceled one is not. How the reserve, capture and refund work, and what a 402 means.
- AI image generator for commercial use: what the terms allow
Commercial use of AI images is set by each tool's terms and plan, not the model. On Sume, paid plans include it and Sume claims no ownership of outputs.
Written by Sume