Captions vs subtitles: what's the difference?

Captions write down what is heard, sound cues included, for viewers who can't hear. Subtitles translate the dialogue for viewers who don't speak it.

4 min readSume
All posts

Captions and subtitles differ in who they are for. Captions write down what is heard, in the video's own language, for viewers who can't hear the audio, so they include sound cues such as [music] or [door slams] as well as the dialogue. Subtitles translate the dialogue for viewers who can hear it but don't understand the language, so they carry the dialogue only.

Both are text on screen, and the same tools make both. The Sume facts below come from the Video captions docs and the Sume API reference, read on 2026-09-29; anything called current behavior is read from Sume's code. Whether captions are closed (a track viewers turn on) or open (burned into the picture) is a separate question, covered in hard subtitles vs soft subtitles.

What is the difference between closed captions and subtitles?

The same split, with delivery added. Closed captions are captions a viewer can switch on or off. Subtitles can also be switched on or off, but their text is the dialogue in another language. So the questions to ask are:

  • Who is reading? Viewers who can't hear the sound need captions, including non-speech sounds and, where it matters, who is talking.
  • Which language? Captions stay in the spoken language. Subtitles are in the viewer's language.
  • Can it be turned off? Closed captions and soft subtitles can. Open captions and hard subtitles are part of the picture and always show.
  • Is the sound off? Viewers who can hear but watch muted read the same captions, so captions of the spoken words serve them too.

How are captions generated automatically?

Automatic captions come from speech-to-text: a model transcribes the speech and times each word, and the words are drawn over the video. Sume's caption job (POST /v1/video-captions) works this way. Send a public HTTPS video_url, and it transcribes the speech and burns the words into a new video. language is only a speech-to-text hint; omit it for automatic detection.

Omit any wording field and the job burns the speech-to-text wording: a transcript of the speech. Sound descriptions and speaker names are yours to add, as the next section shows. To fix the wording but keep the speech timing, send your script as script_text: Sume keeps the speech-to-text timings and aligns the burned wording to your script.

How do I add sound cues or subtitles in another language?

Write the lines yourself and send them as cues: each has text, start, and end in seconds, and Sume burns exactly that copy at those times with no speech-to-text. A cue can be a sound line such as [music], so captions can carry what speech-to-text leaves out. For translated subtitles, cues also carry text in the viewer's language; Korean subtitles for an English video and bilingual subtitles walk through that.

From Video captions and the Sume API reference, read 2026-09-29.
You wantSend to /v1/video-captionsWhat gets burned
Captions of the speechvideo_url only (optional language hint)The speech-to-text wording
Captions with your exact wordingvideo_url + script_textYour script, on speech-to-text timings
Captions with sound cuesvideo_url + cues (text, start, end)Exactly your lines, at your times
Subtitles in another languagevideo_url + cues in that languageExactly your lines, at your times

What are the limits, and what does it cost?

  • The result is burned in: open captions or hard subtitles that viewers can't turn off. The job returns a captioned video_url; raw transcripts are not part of its public contract.
  • Sume takes no .srt upload. Pass phrase-level text as cues instead; burn an SRT file into a video shows the conversion.
  • A job takes up to 200 cues, each timed within 60 seconds. In current code it refuses a source over 60 seconds, and a source with no audio stream even when you send cues.
  • Korean text on the Latin styles slam, punch, or tiktok-green is refused with 400 (caption_hangul_text_latin_style); pick a Hangul style.
  • Each accepted job reserves and captures $0.20 for videos up to 60 seconds, plus a 5.5% agent fee by default.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume