Veo 3.1 reference images: three max, and the 8-second rule
Veo 3.1 takes up to three reference images, and a request with them must run 8 seconds. Sume does not list Veo; its reference-capable models take more.

Veo 3.1 accepts up to three reference images, sent with a referenceType of "asset" to keep a subject's appearance, and Google says a request with reference images must use an 8-second duration. Sume does not list Veo in its video catalog, so there is no Veo request to send there; the models it does list take more references, but with different rules per model.
Veo facts are from Google's Veo page; Sume facts are from Video generation and Video Router. Both read 2026-09-29.
What does Google say about Veo 3.1 references?
The page describes reference images as "up to three images to be used as style and content references". Its referenceType value "asset" is the one for preserving how a subject looks. Durations are 4, 6 or 8 seconds, and 8 is required with reference images, as it is for 1080p and 4K. Portrait and landscape are both listed: 9:16 and 16:9.
How many references do the models on Sume take?
Sume sends references in input_references on POST /v1/videos, and a model accepts only the types its supported_input_references lists. Caps differ by model:
| Model | Images | Video clips | Audio |
|---|---|---|---|
| Veo 3.1 (Google page; not on Sume) | Up to 3 | Not listed | Not listed |
wan-3.0 | Up to 10 | Up to 5 | Up to 5 |
gemini-omni-flash-1.1 | Up to 10 | Up to 3, each 3 s or less | No |
seedance-2.5 | Yes | Yes | Yes |
Does a reference force a clip length on Sume?
No model on Sume ties references to one duration the way Google's page does for Veo. Each model keeps its own range: gemini-omni-flash-1.1 runs 3 to 10 seconds and wan-3.0 2 to 30, whether or not you attach references. One rule from the docs does apply everywhere: if a request carries both frame_images and input_references, frame_images wins and the call is image-to-video.
How do I send references on Sume?
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: refs-001" \
-d '{
"model": "wan-3.0",
"prompt": "The character from the reference walks through a market",
"resolution": "720p",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/hero.png"}}
]
}'Which one fits a three-image brief?
If your brief has three or fewer references and an 8-second clip is fine, Veo's limits are not a constraint, but you would call Google's API for it. If you need more than three images, or video and audio references in the same request, check supported_input_references in GET /v1/videos/models and pick from what is listed.
Sources
Related posts
More in Models
- Veo 3.1 vs Kling 3.0: how to choose by inputs and limits
Google's Veo 3.1 limits next to the kling-3 entry in Sume's catalog: clip length, resolution, references and audio. Sume lists kling-3 and does not list Veo.
- Choosing an AI video API as a developer: a technical checklist
Choose an AI video API on job model, not demos: async submit and poll, idempotent retries, webhooks, live capability discovery, a billing model you can read
- AI video with audio: which models always add sound, which are optional
On Sume, minimax-h3, minimax-h3-max and gemini-omni-flash-1.1 always add native audio; kling-3, Seedance and Wan make it optional; grok-imagine-video-1.5 none.
- sume/auto or a pinned video model: how to decide
Use sume/auto when you accept 3 to 10 second clips at 16:9 or 9:16 and do not care which family runs. Pin a model id for a specific input, length or audio.
Written by Sume