Kling 4.0 on-screen text vs Sume burned-in captions

Kling 4.0 renders on-screen text inside the video. Sume burns captions onto a finished clip; language is a speech-to-text hint, and cues add authored text.

4 min readSume
All posts

Kling 4.0's page says it generates on-screen text in 9 languages, plus emoji, that stays readable as the shot moves. That text is drawn by the video model. Sume's route is different: POST /v1/video-captions burns captions onto an existing public video URL, and its language field is a speech-to-text hint, not a translation setting.

Sume's behavior is from Video captions, read 2026-10-01.

Where does the text come from in each case?

With Kling, the model renders the text as part of the picture while it generates. With Sume's captions, the text is added after the video exists, from speech in the clip, from your script_text, or from authored overlay cues. The docs say to prefer this when you already have a finished clip.

Two ways to get text on screen, read 2026-10-01.
QuestionKling 4.0 (vendor page)Sume video captions
When is text addedDuring generationAfter, on a finished clip
Languages9, per the pagelanguage hints speech-to-text (ko, en, and so on); omit to auto-detect
Exact wordingPrompt-drivenscript_text aligns to speech; cues or segments burn authored text

Does the language field translate my captions?

No. The docs describe language as a speech-to-text hint and say it never selects the style or the font. They describe no translation step. If you want captions in another language, supply the translated copy yourself as cues, each with text, start and end.

How do I burn exact text on a silent clip?

Speech-to-captions needs audible speech; a silent clip fails as caption_no_speech. Pass cues or segments to burn authored overlay copy without speech recognition, as in captions for silent clips. script_text, words, cues and segments are mutually exclusive.

What does a caption job cost and need?

The docs say each accepted standalone job reserves and captures $0.20 for videos up to 60 seconds under the current fixed estimate, and tell you to confirm live pricing in GET /v1/catalog. video_url must be a fetchable public HTTPS video. Korean copy needs a Hangul style; Latin styles return 400. See burn captions onto video.

Sources

Related posts

More in Models

All Models posts

Written by Sume