Captions misspell your brand name: fix it with authored cues

Sume video captions have no custom dictionary field. Pass cues with your own spelling and speech-to-text is skipped, so the brand name prints as written.

4 min readSume
All posts

The Sume video captions endpoint has no dictionary or custom-vocabulary field, so you cannot register a brand name in advance. To control the spelling, send cues (or segments) with your own text, start and end: the docs say authored overlay cues skip speech-to-text, so nothing is guessed.

Submagic's create-project request shows a dictionary array of terms. The Sume request fields below are from the video captions docs, read 2026-10-01.

Which caption fields control the words?

Required is video_url. The optional text-related fields differ in who decides the words.

Where the caption words come from, per the Sume video captions docs read 2026-10-01.
InputWords come from
Nothing but video_urlSpeech-to-text on the audio
script_textYour script, aligned to the speech
cues / segmentsYour text and timings; no speech-to-text

When does script_text help?

script_text still needs audible speech, since it aligns your script to what is said. If the script and the audio disagree, see caption script alignment mismatch. It is the lighter fix when the speaker follows the script word for word and only the spelling is the problem.

How do I send authored cues?

Each cue needs text, start and end. You supply the timings, so take them from your own transcript or from the script you recorded against. The language field is only a speech-to-text hint (ko, en), and it does not change how authored text is drawn.

{
  "video_url": "https://example.com/launch.mp4",
  "cues": [
    { "text": "Meet Acme-Flow", "start": 0.0, "end": 1.6 },
    { "text": "built for small teams", "start": 1.6, "end": 3.2 }
  ]
}

Does it cost anything extra?

Each accepted standalone caption job reserves and captures $0.20 USD for videos up to 60 seconds under the docs' current fixed estimate, whichever input you use. If you want a free look at the audio first, a video inspect takes an optional language_code when transcribe is on; probe and stills are unbilled.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume