Add a hook title to the first seconds of a video with one cue

Send one authored cue with start 0 and end 3 to POST /v1/video-captions and Sume burns that hook text into the clip, with no speech-to-text step.

4 min readSume
All posts

Sume has no dedicated hook-title field. You get the same result by sending one authored cue to POST /v1/video-captions with start 0 and end 3: the text is burned over those seconds and speech-to-text is skipped.

Submagic's create-project docs show a hookTitle option with text, template, top and size. Sume's route is different: it is the general caption route, and you place the card yourself. Facts below are from Video captions, read 2026-10-01.

How do I send a hook as a cue?

Pass cues (or segments) with text, start and end in seconds. The docs say this burns exactly that copy at those times without ASR, and script_text, words, cues and segments are mutually exclusive, so a hook cue replaces spoken-word captions on that job. video_url must be a public HTTPS video URL.

const res = await fetch("https://api.sume.com/v1/video-captions", {
  method: "POST",
  headers: {
    Authorization: "Bearer " + process.env.SUME_API_KEY,
    "Content-Type": "application/json",
    "Idempotency-Key": "hook-title-001",
  },
  body: JSON.stringify({
    video_url: "https://media.sume.com/artifacts/example/clean.mp4",
    cues: [{ text: "Stop scrolling: watch this", start: 0, end: 3 }],
  }),
});
console.log(res.status, await res.json());

Can I move the hook to the top of the frame?

The design.placement group has anchor_ratio, the line centre as a fraction of frame height, and landscape_anchor_ratio for wide frames. A small value sits near the top. design is not supported on the punch and tiktok-green styles. The docs do not show a cue-only example with design, so render one short clip and check the position before a batch.

What does it cost and what can go wrong?

Hook-cue behavior from the Sume docs, read 2026-10-01.
TopicWhat the docs say
Fixed price$0.20 per accepted job, videos up to 60 seconds
Speech neededNo for cues; a silent clip fails as caption_no_speech only on the speech path
Korean copyUse a Hangul style; Latin styles return caption_hangul_text_latin_style
SourcePublic HTTPS URL; private and signed URLs are rejected

How is this different from word captions?

Word captions follow speech. A hook is one authored card with fixed timing. If the clip also needs spoken captions, run them as a separate pass; see burn captions onto a video. Silent source clips use the same cue path, covered in silent clips and overlay cues.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume