Pull product stills from a video with the video-frames API
Video frames returns jpeg or png stills from a Sume-hosted clip at set seconds. Google lists 500 x 500 minimum, so extract from the clip before captions.

POST /v1/video-frames returns up to 24 durable jpeg or png stills from a Sume-hosted clip at the seconds you list, and it is unbilled. Google Merchant Center lists at least 500 x 500 pixels and no promotional elements or content covering the product, so extract stills from the clean clip, not the captioned cut.
How do I ask for stills?
Send video_url (a media.sume.com clip) and exactly one of at[] or fps. at[] takes 1 to 24 seconds, each at least 0 and below the clip length; format is jpeg (default) or png; max_edge clamps the long edge from 16 to 2160. Submit always returns 202; read the result at GET /v1/video-frames/{id}.
curl -X POST https://api.sume.com/v1/video-frames \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stills-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/product.mp4",
"at": [0.5, 2, 4],
"format": "png"
}'What comes back?
When resource_status is ready, frames holds t, url, width and height for each still, hosted as durable artf_ images. One instant that fails comes back with a null url and does not fail the job.
How does this line up with Google's page?
The rules below are Google's; the frames columns are from the video frames docs.
| Google line | Video frames lever |
|---|---|
| At least 500 x 500 pixels | Omit max_edge to keep the source size, then read width and height |
| No promotional elements or covering content | Extract from the clip before captions or overlays |
| JPEG or PNG accepted | format is jpeg or png |
What should I check?
Pick times where the product is clear of hands and other objects. Google's page in 2026-09-30 recommends 1500x1500 pixels or above, so check the still's width and height in the result against it.
Sources
Related posts
More in Use cases
- Use a frame from your footage as a reference for AI video
Pull a still at a time you pick with video frames (unbilled), then pass it as image_url or reference_image_urls to a Sume video model. Steps and limits.
- Remove background from video with AI: what Sume covers
Sume has no video background-removal endpoint: RMBG 1.0 takes a still image. What you can do instead is a prompt edit with video_to_video, or a crop.
- Replace video audio with an AI voice (and the lip-sync catch)
Sume has no one-call audio swap: detach the audio, transcribe it, make a new voice with TTS, and lay it on a Timeline. Lips will not re-sync to the new voice.
- AI ad resizer: one image, several aspect ratios, one API
Sume has no resizer button. Send the same source image to POST /v1/images once per aspect_ratio (1:1, 4:5, 9:16, 16:9) and recompose each ad slot.
Written by Sume