Reframe a video around the speaker by API: what a crop can do
Descript's Center Active Speaker (Beta) follows whoever talks. Sume's video filter crop is one fixed rectangle per clip, so cut per speaker and crop each part.

The Sume docs list no speaker tracking. The crop op in POST /v1/video-filter cuts one fixed rectangle, given as fractions of the frame, from the whole clip. To follow two speakers you trim the clip per speaker turn, crop each piece with its own rectangle, and join the pieces in a Timeline.
Descript's help page describes Center Active Speaker as a way to "automatically reframe your video around whoever's speaking" and marks it Beta. This post covers the manual route, from the Video filter docs, read 2026-09-30.
How does a crop rectangle work?
crop takes x and y in [0, 1] and width and height in [0.05, 1], with x + width <= 1 and y + height <= 1. The compiler even-rounds for yuv420p. A rectangle outside the frame fails with video_filter_crop_out_of_bounds. Check first with the free POST /v1/video-filter/check.
curl -X POST https://api.sume.com/v1/video-filter \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: crop-left-speaker-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/two-shot.mp4",
"ops": [{ "op": "crop", "x": 0, "y": 0, "width": 0.5, "height": 1 }]
}'How do I switch between speakers?
Decide the turns in your own code (for example from a transcript). Then, for each turn, cut the range with POST /v1/video-trim, crop that piece with the rectangle that frames the speaker, and drop the resulting MP4s into video[] of POST /v1/timeline-1.0/render in order. Each step is its own job, so a two-speaker interview of ten turns is roughly ten trims, ten crops and one render.
What does this cost and where does it stop?
The docs list $0.02 per trim job, $0.02 per filter encode (the check is free) and $0.10 per output minute for the render. The crop is static, so a speaker who moves inside a turn drifts out of frame. The docs say source clips for the filter are at most 300 s.
| Step | Route | Public rate |
|---|---|---|
| Cut one speaker turn | POST /v1/video-trim | $0.02 per job |
| Crop to the speaker | POST /v1/video-filter | $0.02 per encode |
| Preflight the crop | POST /v1/video-filter/check | Free |
| Join the pieces | POST /v1/timeline-1.0/render | $0.10 per ceil(output minute) |
Sources
Related posts
More in Developers
- chatgpt-image-latest shuts down Dec 1: pin a Sume model id
OpenAI retires the chatgpt-image-latest alias on Dec 1, 2026. On Sume, pin openai/gpt-image-2.5 or send sume/auto, and read ids from GET /v1/images/models.
- Claude Code AGENTS.md without CLAUDE.md: Sume skills install
Claude Code reads AGENTS.md when no CLAUDE.md exists. sume skills install writes to .agents/skills or .claude/skills, so check which one you got.
- Claude Code auto mode with Sume's MCP: paid-call gates still apply
In Claude Code auto mode a classifier reviews each action. On Sume's MCP, paid calls still need idempotency_key, and dry_run and max_spend_usd stay available.
- Claude Code --bare and MCP: name the Sume server with a key
Claude Code --bare connects only MCP servers named on the command line. To use Sume there, name the server with an API-key header; no browser sign-in needed.
Written by Sume