Veo 3.1 first and last frame: how the API works
Veo 3.1 takes a first image and a last image and generates the motion between them. Here is Google's request shape and where the same input works on Sume.

Veo 3.1 can start from a first image and end on a last image: you pass the first as image and the last as lastFrame, and the model generates the motion between them. Google marks this as available for Veo 3.1 models only, and lastFrame must be used together with image. Sume's video catalog does not currently list a Veo model. On Sume, the same idea is frame_images with first_frame and last_frame entries, on the models below.
Google's parameters are from its Veo page; Sume's are from the Video generation docs, read 2026-09-29.
What does Google's Veo 3.1 last-frame request need?
Google's parameter table lists image (an initial image to animate) and lastFrame (the final image for an interpolation video to transition, only with image). Aspect ratio is 16:9 (default) or 9:16, duration is 4, 6 or 8 seconds, and resolution defaults to 720p. Google calls this interpolation; personGeneration for it is allow_adult only.
How do I send first and last frames on Sume?
Use frame_images on POST /v1/videos. Each entry needs a frame_type of first_frame or last_frame. A last frame without a first frame is refused. If you also send input_references, frame_images wins and the request is image-to-video.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: frames-001" \
-d '{
"model": "seedance-2.5",
"prompt": "The door opens and morning light fills the room.",
"frame_images": [
{"type":"image_url","image_url":{"url":"https://example.com/closed.png"},"frame_type":"first_frame"},
{"type":"image_url","image_url":{"url":"https://example.com/open.png"},"frame_type":"last_frame"}
],
"duration": 6
}'Which Sume models accept a last frame?
From the catalog code, the models that list both frame types are below. grok-imagine-video-1.5 takes a first frame only, and it needs one.
| Model | Frame types | Clip length |
|---|---|---|
seedance-2.5 | first, last | 4 to 30 s |
seedance-2, seedance-2-fast, seedance-2-mini | first, last | 4 to 15 s |
kling-3 | first, last | 4 to 15 s |
wan-3.0 | first, last | 2 to 30 s |
minimax-h3, minimax-h3-max | first, last | 5 to 15 s |
gemini-omni-flash-1.1 | first, last | 3 to 10 s |
grok-imagine-video-1.5 | first only | 4 to 15 s |
What should the two frames look like?
Keep both images the same size and framing, and use public HTTPS URLs. The model invents everything between them, so if the two frames differ a lot, ask for one simple move in the prompt. Check the first frames of the result before you use it; if the ending drifts, shorten the duration.
Sources
Related posts
More in Models
- Veo 3.1 Lite, Fast or standard: how to choose
Veo 3.1 has three preview models: standard, Fast and Lite. Lite drops 4K and extension; Fast targets speed. What each does, plus Sume's tiered Seedance.
- Veo 3.1 reference images: three max, and the 8-second rule
Veo 3.1 takes up to three reference images, and a request with them must run 8 seconds. Sume does not list Veo; its reference-capable models take more.
- Veo 3.1 vs Kling 3.0: how to choose by inputs and limits
Google's Veo 3.1 limits next to the kling-3 entry in Sume's catalog: clip length, resolution, references and audio. Sume lists kling-3 and does not list Veo.
- Choosing an AI video API as a developer: a technical checklist
Choose an AI video API on job model, not demos: async submit and poll, idempotent retries, webhooks, live capability discovery, a billing model you can read
Written by Sume