Search footage for a spoken phrase with an API: words with timing

DaVinci Resolve 21 lists IntelliSearch. To find a spoken phrase in a clip by API, transcribe with video inspect and search words[] in your own code.

4 min readSume
All posts

To find where a phrase is spoken in a clip, run POST /v1/video-inspect with transcribe: true, then search the returned words[] in your own code and use the match time to pull a still. The docs list a transcript and stills, and no search endpoint over footage.

Blackmagic's DaVinci Resolve 21 page lists "Search Content with AI IntelliSearch" among its AI tools. This post covers finding spoken words in one clip with the public routes, from the Video inspect docs, read 2026-09-30.

What does the transcript give me to search?

With transcribe: true the resource carries transcript with text, words[], optional sentence segments[] (set segmentation.mode: "sentence") and an audio_url. Set language_code (for example en or ko) as a hint, or omit it to auto-detect. The rate is $0.01 per audio minute; probe and stills stay unbilled.

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: find-phrase-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "frames": false,
    "transcribe": true,
    "language_code": "en"
  }'

How do I jump from a match to a picture?

The docs name words[] but this page does not list its fields, so read a real response to see how each word's time is named. Take the matched word's time and call POST /v1/video-frames with at: [t] to get a durable still at that instant, at source size. Every at value must satisfy 0 <= t < duration. For a skim, video inspect offers seek: "fast", which snaps each still to the keyframe at or before its instant; keep precise when the timestamp must match.

What are the limits?

The docs cap video inspect at a source of 1800 s, and a silent clip fails with inspect_source_has_no_audio. Check probe.has_audio first with frames: false. This route is about spoken words; the docs show no search over what is on screen.

What each Sume route returns for a search workflow (https://docs.sume.com/models/video-inspect, read 2026-09-30)
NeedRouteReturns
Spoken words with timingPOST /v1/video-inspectwords[], optional segments[]
Still at the matchPOST /v1/video-framesframes[{t,url,width,height}]
Is there audio at allvideo-inspect with frames falseprobe.has_audio

Sources

Related posts

More in Developers

All Developers posts

Written by Sume