Add a hook title to the first seconds of a video with one cue
Send one authored cue with start 0 and end 3 to POST /v1/video-captions and Sume burns that hook text into the clip, with no speech-to-text step.

Sume has no dedicated hook-title field. You get the same result by sending one authored cue to POST /v1/video-captions with start 0 and end 3: the text is burned over those seconds and speech-to-text is skipped.
Submagic's create-project docs show a hookTitle option with text, template, top and size. Sume's route is different: it is the general caption route, and you place the card yourself. Facts below are from Video captions, read 2026-10-01.
How do I send a hook as a cue?
Pass cues (or segments) with text, start and end in seconds. The docs say this burns exactly that copy at those times without ASR, and script_text, words, cues and segments are mutually exclusive, so a hook cue replaces spoken-word captions on that job. video_url must be a public HTTPS video URL.
const res = await fetch("https://api.sume.com/v1/video-captions", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.SUME_API_KEY,
"Content-Type": "application/json",
"Idempotency-Key": "hook-title-001",
},
body: JSON.stringify({
video_url: "https://media.sume.com/artifacts/example/clean.mp4",
cues: [{ text: "Stop scrolling: watch this", start: 0, end: 3 }],
}),
});
console.log(res.status, await res.json());Can I move the hook to the top of the frame?
The design.placement group has anchor_ratio, the line centre as a fraction of frame height, and landscape_anchor_ratio for wide frames. A small value sits near the top. design is not supported on the punch and tiktok-green styles. The docs do not show a cue-only example with design, so render one short clip and check the position before a batch.
What does it cost and what can go wrong?
| Topic | What the docs say |
|---|---|
| Fixed price | $0.20 per accepted job, videos up to 60 seconds |
| Speech needed | No for cues; a silent clip fails as caption_no_speech only on the speech path |
| Korean copy | Use a Hangul style; Latin styles return caption_hangul_text_latin_style |
| Source | Public HTTPS URL; private and signed URLs are rejected |
How is this different from word captions?
Word captions follow speech. A hook is one authored card with fixed timing. If the clip also needs spoken captions, run them as a separate pass; see burn captions onto a video. Silent source clips use the same cue path, covered in silent clips and overlay cues.
Sources
Related posts
More in Use cases
- Ken Burns effect API: Timeline stills are static holds
Sume Timeline holds a still image static; motion is accepted and ignored with a motion_ignored warning. zoompan is on the video filter allowlist.
- AI kids story video generator: scenes of 2 to 30 seconds
Make a children's story video as one scene per request: wan-3.0 takes 2 to 30 seconds per scene, then join up to 200 slots in a Timeline 1.0 render.
- AI real estate walkthrough from photos: chain first/last-frame clips
Kling 4.0 takes up to 10 keyframes for a walkthrough. Sume takes a first and last frame per clip, so chain one clip per room pair and join them on a timeline.
- LinkedIn ad image size and AI-generated images in Campaign Manager
LinkedIn single image ads accept JPG, PNG or GIF up to 5 MB, and Campaign Manager has its own AI image tool. Which format to set on a Sume upload.
Written by Sume