Final Cut Pro convert closed captions: burn cues in Sume
Final Cut Pro 12.3 converts closed captions to subtitles. To burn existing text onto a video elsewhere, send cues or segments to Sume; SRT is unsupported.

Final Cut Pro 12.3 added a way to convert existing closed captions to subtitles. The Sume docs describe no closed-caption converter, but if you have the text and timings, POST /v1/video-captions burns them onto a video as cues or segments. SRT uploads are unsupported, so you map each line to text, start and end.
Apple's line is from its release notes; Sume's from Video captions, read 2026-10-01.
What did Apple add in 12.3?
Under workflow enhancements the notes list: easily convert existing closed captions to subtitles, quickly select all subtitles with a single command, and adjust position, rotation, scale and alignment for multiple selected subtitles from the inspector. All of that happens inside the editor.
How do I burn existing caption text with Sume?
Pass cues (or segments), each with text, start and end in seconds. The docs call this an authored overlay that skips speech-to-text. script_text, words, cues and segments are mutually exclusive, so send only one.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: cc-burn-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/clean.mp4",
"cues": [
{ "text": "[door closes]", "start": 1.2, "end": 2.4 },
{ "text": "We start Monday.", "start": 2.4, "end": 4.8 }
]
}'Why not upload the caption file?
The docs state that SRT uploads and provider task ids are unsupported. The docs describe no reading of closed-caption tracks from your video, so send the text in the request.
| Input | Accepted? |
|---|---|
cues / segments | Yes, skip speech-to-text |
words (word-level) | Yes, skip speech-to-text |
| SRT upload | No |
| Provider task id | No |
Can I change the look afterwards?
Yes. To re-burn the same video under a different style, pass source_caption_id instead of video_url. The docs say Sume reuses that caption's source video and the word timings it already has, and billing is unchanged because a restyle is still a render. Standalone jobs reserve $0.20 USD for videos up to 60 seconds under the current fixed estimate.
What about silent clips?
A silent clip fails as caption_no_speech unless you pass cues, which is the usual route for existing text; see silent clips and overlay cues and captions vs subtitles.
Sources
Related posts
More in Use cases
- Final Cut Pro Edit Detection: sample stills to see the cuts
Final Cut Pro 12.3 Edit Detection splits a rendered video at shot changes. Sume does not detect cuts; video-frames samples stills at an fps so you can see them.
- Final Cut Pro Generate Captions: US English only, Korean fix
Apple's Generate Captions in Final Cut Pro 12.3 is U.S. English only. For Korean speech, Sume video captions takes a language hint and Hangul caption styles.
- Firefly Composite API for product photos vs Sume reference edit
Firefly Composite Operations blend a product photo into a generated scene. On Sume, send the photo as an input reference, with an optional mask_url.
- Fix color banding in AI video: the deband filter API
Visible steps in skies and gradients can be softened with the deband filter in Sume's video filter, with noise as a fallback. It returns a new MP4. Check free.
Written by Sume