What a Sume video analysis returns for each scene

Each scene in a stored video analysis has start and end seconds, summary, visual, shot type, on-screen text, confidence and keyframes. Field-by-field list.

4 min readSume
All posts

A Sume video analysis returns a typed scenes[] array. Each scene has contiguous start_seconds and end_seconds, a summary, visual, optional shot_type and camera_motion, on_screen_text, audio, a keyframe_url, keyframes and a confidence from 0 to 1. This is the legacy surface; stored rows stay readable.

What is in each scene?

The video analyses docs describe a video-understanding model segmenting the MP4 into typed scenes, then ffmpeg extracting at least one JPEG still per second of each scene. The current analysis_version is 1.1.

Scene fields, read 2026-09-30 from the Sume docs.
FieldWhat it holds
start_seconds / end_seconds / duration_secondsContiguous timing for the scene
summary, visualText description of the scene
shot_type, camera_motionOptional
on_screen_textText shown in the scene
audio{ speech, has_speech, has_music } only if include_transcript was true, otherwise null
keyframe_urlRepresentative JPEG, or null with a keyframe_mirror_failed warning
keyframesArray of { t, url }, one still per whole second; a failed second is url: null
confidence0 to 1

How do I get the spoken words per scene?

Set include_transcript to true when you create the analysis; each scene's audio then carries the spoken speech. Without it, audio is null. There is no separate one-second text or context track.

How accurate is it?

The docs give one caveat: longer videos are less accurate, and short clips give the most reliable results. Each scene also carries a confidence value. It is not virality prediction and it does not generate video.

Should I build on it?

The page says not to start new work on this surface. For probe facts, stills and an optional transcript, use video inspect, which does not return typed scenes.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume