Extract on-screen text from a reference video with an API

Reference ingest reads on-screen text at source resolution, returns lines with boxes, spans and confidence, and flags low-confidence lines instead of guessing.

4 min readSume
All posts

To pull the on-screen text out of a reference video, call POST /v1/reference-ingest. The manifest's text_tracks[] holds each line it read, with the text, a normalised box, a time span, a confidence and a role_hint. Nothing is corrected: a line below the confidence threshold is marked needs_verification and comes with a crop at the clip's native resolution.

Availability matters. Per the Reference ingest page read 2026-09-29, the endpoint is on in development and opt-in in production, so confirm it is listed for your workspace first.

Which languages does it read?

The OCR pass is set up for Korean and Latin text. ocr.languages defaults to ["ko", "en"]. It reads deduplicated frame states at source resolution and merges matches across frames into lines, so a caption that stays on screen for three seconds is one line, not ninety.

How do I tune it?

OCR options, from Reference ingest, read 2026-09-29.
FieldMeaningDefault
ocr.languagesLanguages to read["ko", "en"]
ocr.fpsText-state sampling rate, 0.5 to 2. OCR runs on deduplicated states, at most five frames1
ocr.min_confidence_attach_cropLines under this value are needs_verification and carry a crop0.85
curl -X POST https://api.sume.com/v1/reference-ingest \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: ref-ocr-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/reference.mp4",
    "ocr": { "languages": ["en"], "fps": 2, "min_confidence_attach_crop": 0.9 }
  }'

What should I do with a low-confidence line?

Read the crop, then, if you still need a pixel-level look, extract a frame at a manifest time with the video frames endpoint. The docs say to do this at most once per uncertain[] entry. Do not paste an unverified OCR line into a brief or an ad; a wrong price or claim copied from a reference is a real error, and the manifest tells you exactly which lines to check.

What can go wrong?

  • source_too_long_for_reference_ingest: the clip is over 300 seconds.
  • source_no_video_stream: the file has no video.
  • ffmpeg_fields_rejected: you sent a filter or codec key. Sume compiles every pass itself.
  • 400 for semantic: true, which is refused with reference_ingest_semantic_unavailable until that pass ships.

How do I get the crops back?

delivery.inline_strip and delivery.inline_crops control what the Sume Agent host attaches as inline images: the strip is on by default, and crops are none, low_confidence_only (default) or all. Over plain HTTP you read the manifest and its crops from the response. The hosted MCP reference_ingest tool returns the manifest as text.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume