caption_no_speech: caption a silent clip with cues
caption_no_speech means the clip has no audible speech. Sume's fix is next_action use_overlay_captions: send cues with text, start and end and skip ASR.

caption_no_speech is the failure Sume gives when a caption job has no audible speech to transcribe. It carries next_action: use_overlay_captions: resend the job with cues (or segments), each with text, start and end, and Sume burns that copy without running speech-to-text.
Auto captions of the kind Descript redesigned on 2026-09-17 assume someone is talking. The rule here is in Video captions.
What does the cues body look like?
Times are in seconds. This body burns two overlay cards on a silent clip.
{
"video_url": "https://media.sume.com/artifacts/example/silent.mp4",
"cues": [
{ "text": "Spring collection", "start": 0.0, "end": 2.0 },
{ "text": "Out now", "start": 2.0, "end": 4.0 }
]
}Which inputs can I combine?
Only one wording source per request.
| Input | Runs speech-to-text? | Use it for |
|---|---|---|
| none | Yes | Burn what is spoken |
script_text | Yes, for timings | Align your wording to the speech |
words | No | Word-level timing you already have |
cues / segments | No | Phrase-level overlay cards, silent clips |
Can I upload an SRT instead?
No. The docs say SRT uploads and provider task ids are unsupported; pass phrase-level text as cues or segments. script_text, words, cues and segments are mutually exclusive, so a request with two of them is invalid.
When does this matter for ads?
A clip with music and no speech still needs words on screen for viewers who watch with the sound off. See captions on Facebook feed video ads for that context, then author the cues yourself.
Sources
Related posts
More in Developers
- catalog_list shows a capability that tools_list does not
On Sume's hosted MCP, catalog_list lists public API capabilities and can name ones with no MCP tool. tools_list is the session's real tool contract.
- Reframe a video around the speaker by API: what a crop can do
Descript's Center Active Speaker (Beta) follows whoever talks. Sume's video filter crop is one fixed rectangle per clip, so cut per speaker and crop each part.
- chatgpt-image-latest shuts down Dec 1: pin a Sume model id
OpenAI retires the chatgpt-image-latest alias on Dec 1, 2026. On Sume, pin openai/gpt-image-2.5 or send sume/auto, and read ids from GET /v1/images/models.
- Claude Code AGENTS.md without CLAUDE.md: Sume skills install
Claude Code reads AGENTS.md when no CLAUDE.md exists. sume skills install writes to .agents/skills or .claude/skills, so check which one you got.
Written by Sume