Threads podcast transcript post: burn the quote onto the clip
For a podcast quote clip, burn captions from your exact transcript text with Sume video captions script_text, or authored cues. SRT uploads are not accepted.

Threads' podcast toolkit includes transcript posts for sharing highlights. If you also want the quote burned onto a video clip, Sume POST /v1/video-captions can do it from the exact wording you supply: script_text aligns the burned-in text to your script, and cues burn authored text with no speech-to-text.
Threads details are from Meta's announcement; caption behavior from Video captions, read 2026-10-01.
What does Threads offer for transcripts?
Meta's post says you can share standout moments and featured guests with transcript, video, and guest card posts. It does not describe how a transcript post is built, so this page covers only the video side: a clip whose on-screen words match your transcript.
How do I keep the wording exact?
Pass script_text. The docs say Sume keeps the speech-to-text word timings as the timing source of truth and aligns burned-in wording to your script. That needs audible speech in the clip, and alignment can fail with typed errors such as script_alignment_mismatch; see the mismatch fix.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: quote-clip-001" \
-d '{
"video_url": "https://example.com/quote-clip.mp4",
"script_text": "Your exact quote, as written in the transcript."
}'Which input fits which case?
| Input | Behavior |
|---|---|
| None | Speech-to-text transcribes the clip |
script_text | Aligns burned-in wording to your script |
cues or segments | Authored overlay with text, start, end; skips speech-to-text |
words | Also a fixed-copy input |
Can I upload an SRT file?
No. The constraints say SRT uploads and provider task ids are unsupported; pass phrase-level text as cues or segments. script_text, words, cues and segments are mutually exclusive. A silent clip fails as caption_no_speech, so use cues for it, as in silent clips with overlay cues.
What must the video URL look like?
A fetchable public HTTPS video URL. Localhost, private-network, non-HTTPS and signed or private URLs are rejected, so a clip you cut earlier must be reachable that way. The input is an existing finished clip; captions are burned in, not delivered as a sidecar.
Sources
Related posts
More in Use cases
- TikTok Ad Network asset specs: banner 640x100 and where Sume stops
TikTok Ad Network wants video at 1280x720, 720x1280 or 720x720 and banners at 600x500, 640x200 or 640x100. What Sume renders natively and what needs a crop.
- TikTok carousel ad specs: 2-35 images and three sizes
A standard TikTok carousel takes 2 to 35 JPG or PNG images at 1200x628, 640x640 or 720x1280. Here is how to map those onto Sume image settings.
- TikTok carousel ad music: mp3 required, silent on Ad Network
TikTok carousel ads need an .mp3 of 2 s or more, up to 10M. Library music is silent on Ad Network placements, so upload your own file there.
- TikTok Global App Bundle image ad sizes: 9:16, 16:9, 1:1
TikTok's Global App Bundle image ads need 9:16 at 720x1280 or more, 16:9 at 1280x720 or more, or 1:1 at 640x640 or more, as JPG or PNG within 100 MB.
Written by Sume