HeyGen translation SRT captions vs Sume burned-in captions
HeyGen now generates caption sidecars for every translation and lipsync job. Sume burns captions into the video and does not accept SRT as input.

HeyGen's August 2026 change makes caption sidecars always available for completed video translation and lipsync jobs, so you download a separate caption file. Sume works the other way: captions are burned into the video, and SRT is not accepted as input.
HeyGen facts are from its API changelog; Sume facts from Video captions and Generate avatar video, read 2026-09-30.
What changed in HeyGen translation captions?
The changelog says caption-generation flags on video translation and lipsync requests are deprecated and ignored. Captions are always generated, and clients choose whether to display or download them. enable_caption is the flag named as deprecated. Pronunciation edits affect generated speech only; the changelog says captions and subtitle files keep the script's original spelling.
How do Sume captions differ?
Sume's standalone route, POST /v1/video-captions, takes a public HTTPS video_url and returns a captioned video as a job. It also accepts style, font, language, script_text, and authored cues or segments. SRT uploads and provider task ids are unsupported; pass phrase-level text as cues instead.
| Question | HeyGen (changelog) | Sume (docs) |
|---|---|---|
| Default output | Caption sidecars for every completed translation or lipsync job | Captioned video; captions burned in |
| Opt-in flag | enable_caption deprecated and ignored | Inline captions on avatar video, or the standalone route |
| SRT as input | Not stated in the changelog entry | Unsupported |
| Failure behavior | Not stated | Inline caption failure soft-fails; the job can still succeed |
What happens to captions inside an avatar job?
Inline captions on an avatar video do not create a separate billed video-caption job. Caption stage failures soft-fail: the avatar job can still succeed with a clean primary video_url and captions.status=failed. See avatar video captions fail but video succeeds.
Which should I pick for a translated video?
If a platform needs a separate text track, a sidecar file is the shape that fits, and HeyGen now produces one by default. If you want the words in the picture on every player, burned-in is the shape Sume produces. Sume gives you no SRT file to hand to a platform, so the choice is about delivery format, not quality. The difference is explained further in captions vs subtitles.
Sources
Related posts
More in Integrations
- Instagram Audio API for Reels ads: prepare the video side
Meta's Instagram Audio API can find royalty-free replacements for copyrighted Reels music. Sume does not call it, but can drop the audio from your video.
- Add Sume as an MCP server in Junie CLI (mcp.json)
Add Sume to Junie CLI's mcp.json as a remote server with url and headers, or leave headers out and use Authorize for OAuth. Both paths, step by step.
- Kling models in Runway MCP vs Sume: what maps to what
Runway MCP added four Kling entries on Sep 18, 2026. Sume lists a Kling v3 Pro row and a Kling motion control MCP tool; it has no Kling O3 row.
- Make webhook "Queue is full" 400: what Sume does and how to recover
When a Make webhook answers 400 Queue is full or 429, Sume retries up to 10 times at a fixed gap. If all fail, redeliver the job's event once the queue drains.
Written by Sume