What a Sume video analysis returns for each scene
Each scene in a stored video analysis has start and end seconds, summary, visual, shot type, on-screen text, confidence and keyframes. Field-by-field list.

A Sume video analysis returns a typed scenes[] array. Each scene has contiguous start_seconds and end_seconds, a summary, visual, optional shot_type and camera_motion, on_screen_text, audio, a keyframe_url, keyframes and a confidence from 0 to 1. This is the legacy surface; stored rows stay readable.
What is in each scene?
The video analyses docs describe a video-understanding model segmenting the MP4 into typed scenes, then ffmpeg extracting at least one JPEG still per second of each scene. The current analysis_version is 1.1.
| Field | What it holds |
|---|---|
| start_seconds / end_seconds / duration_seconds | Contiguous timing for the scene |
| summary, visual | Text description of the scene |
| shot_type, camera_motion | Optional |
| on_screen_text | Text shown in the scene |
| audio | { speech, has_speech, has_music } only if include_transcript was true, otherwise null |
| keyframe_url | Representative JPEG, or null with a keyframe_mirror_failed warning |
| keyframes | Array of { t, url }, one still per whole second; a failed second is url: null |
| confidence | 0 to 1 |
How do I get the spoken words per scene?
Set include_transcript to true when you create the analysis; each scene's audio then carries the spoken speech. Without it, audio is null. There is no separate one-second text or context track.
How accurate is it?
The docs give one caveat: longer videos are less accurate, and short clips give the most reliable results. Each scene also carries a confidence value. It is not virality prediction and it does not generate video.
Should I build on it?
The page says not to start new work on this surface. For probe facts, stills and an optional transcript, use video inspect, which does not return typed scenes.
Sources
Related posts
More in Developers
- Video analysis API with a YouTube link: 422 unsupported_source
Sume's video analyses reject YouTube URLs with 422 unsupported_source. Import the video you own to media.sume.com first; rules for the other bad inputs.
- Video API duration: whole seconds only, no 7.5 s clips
Sume's video duration is an integer and each model lists whole-second supported_durations. Ranges per model and how to pick one before you submit a job.
- Editing upload limits: Descript's GB tiers vs Sume's seconds caps
Descript lists per-file size caps by plan (1, 10, 20, 50 GB). Sume's media routes cap by clip duration instead: 300 s or 1800 s in, up to 900 s out.
- MCP tools for editing video: the hosted Sume list, trim to timeline
Descript and Runway both ship MCP servers. The hosted Sume MCP exposes video_trim, video_filter, audio_detach, video_frames_create and the timeline tools.
Written by Sume