Multilingual video captions: language is a hint, not a font
On Sume video captions, `language` only tells speech-to-text what to expect. Look and font come from `style`, `design` and `font`, never from the language.

To caption a Spanish, German, French, Portuguese or Italian video on Sume, send the clip to POST /v1/video-captions and treat language as a transcription hint only. The docs say it plainly: language never selects the style or the font. You choose the look separately with style, design and font.
Descript's changelog announced gradient caption fills and filler-word cleanup in five more languages (Spanish, German, French, Portuguese, Italian). Those are Descript editor features; this page covers what the Sume captions route does with the same languages. Facts below are from the Video captions docs, read 2026-10-01.
What does the language field actually do?
The docs describe it as a speech-to-text hint (ko, en, and so on). Leave it out and the language is detected automatically. It does not restyle anything: a style you name renders as named, whatever language the speech is in.
The only place wording changes the look is an omitted style. Then the wording decides: slam for Latin text, black-outline for Korean. Spanish, German, French, Portuguese and Italian are all written in the Latin alphabet, so an omitted style gives you slam.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: caption-es-001" \
-d '{
"video_url": "https://example.com/clip-es.mp4",
"language": "es",
"style": "punch"
}'Which field controls which part of the caption?
| Field | What it controls |
|---|---|
language | Speech-to-text hint; omit for auto-detect |
style | The look and the motion; slam, punch, tiktok-green, korean-ad and the Hangul identities |
design | Per-request overrides of a style's colours, typography, placement, phrasing and motion |
font | A Hangul face for the chosen style; Hangul styles only |
script_text | Script alignment: the words you want burned in |
How do I change the colours or gradient look?
Use design. A style is a set of design tokens and design overrides them for one request, so one key changes one thing. The documented colour fields are base, active, stroke, accent, accent_deep and card, given as hex, rgb()/rgba() or transparent. The docs do not describe a gradient fill, so do not plan around one. Note that design is not supported on punch or tiktok-green; see caption design overrides.
What does a multilingual batch cost?
The caption job price is on the rate card and in GET /v1/catalog; the docs state $0.20 for videos up to 60 seconds under the current fixed estimate. Confirm it there before a large batch.
Should I pass language for every video?
If you know it, pass it; if the batch mixes languages, omit it and let detection run. Either way, pick style once for the whole batch so every language shares one look. If the caption text drifts from what was said, pass script_text for alignment (see Video captions).
Sources
Related posts
More in Use cases
- ElevenLabs agent hold audio: 180 s, 40 MB; make it with Sume
ElevenLabs agent hold audio must be MP3 or WAV, up to 40 MB and 180 seconds. Steer a Sume track's length in the prompt, trim with Timeline audio, export mp3.
- FLUX Virtual Try-On v2: 4 MP inputs and a Sume reference edit
BFL's vto-v2 keeps inputs up to 4 MP as-is. Sume has no try-on endpoint in these docs; a garment swap is a reference edit with public HTTPS images.
- Seamless looping food video with AI: same first and last frame
To loop a food clip, send one image as image_url and end_image_url to Gemini Omni Flash 1.1, which Google says suits seamless loops. Clips run 3 to 10 seconds.
- Gemini TTS two-speaker limit: three-voice dialogue with Sume concat
Gemini TTS configures two speakers per request. For three voices on Sume, make one TTS job per voice and join up to 20 parts with Timeline audio concat.
Written by Sume