Veo 3.1 API: how to call it, and what Sume lists instead

The Veo 3.1 API is Google's, called through the Gemini API with preview model codes. Sume does not list Veo; here are the catalog models with the same inputs.

5 min readSume
All posts

The Veo 3.1 API is Google's: you call it through the Gemini API with a model code such as veo-3.1-generate-preview, poll a long-running operation, and download the result. It makes 8-second-class clips (4, 6 or 8 seconds) at 720p, 1080p or 4K with audio always on. Sume's video catalog does not currently list a Veo model. If you need Veo itself, use Google's API; if you need the same kinds of inputs from Sume, the table below maps them.

Google facts come from its Veo page, read 2026-09-29. Sume facts come from the Video generation docs and the catalog code behind GET /v1/videos/models, read the same day.

How do I call the Veo 3.1 API?

Google's REST example sends a POST to models/veo-3.1-generate-preview:predictLongRunning with your key in x-goog-api-key, then polls the returned operation until done is true and downloads the video from the result. The three Veo 3.1 codes on that page are veo-3.1-generate-preview, veo-3.1-fast-generate-preview and veo-3.1-lite-generate-preview, and the page marks all three Preview.

curl -s "https://generativelanguage.googleapis.com/v1beta/models/veo-3.1-generate-preview:predictLongRunning" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -X POST \
  -d '{"instances":[{"prompt":"A red convertible on a coastal road at sunset, engine roaring."}]}'

Does Sume have Veo 3.1?

Sume's video catalog does not currently list a Veo model. The listed ids include seedance-2.5, seedance-2-mini, seedance-2, seedance-2-fast, kling-3, wan-3.0, grok-imagine-video-1.5, minimax-h3, minimax-h3-max and gemini-omni-flash-1.1. Ask GET /v1/videos/models for the live list before you build on it, because the catalog changes.

The /v1/videos route follows the OpenRouter video shape, so a client written for it changes only the base URL and key. Model ids on Sume are bare catalog ids with no provider prefix.

Which Sume models take the inputs Veo 3.1 takes?

Veo 3.1 takes text, a first image, a last image, up to three reference images, and audio it generates itself. These catalog models accept the matching inputs on Sume.

Veo 3.1 inputs from Google's Veo page against Sume catalog fields, read 2026-09-29.
InputVeo 3.1 (Google)On Sume
First and last frameYes, image plus lastFrameframe_images on seedance-2.5, seedance-2, kling-3, wan-3.0, minimax-h3, gemini-omni-flash-1.1
Reference imagesUp to 3input_references on seedance-2.5, wan-3.0 (up to 10), minimax-h3 (up to 9), gemini-omni-flash-1.1 (up to 10)
Native audioAlways onAlways on for minimax-h3 and gemini-omni-flash-1.1; generate_audio on kling-3 and Seedance
4KYes, 8 s onlygemini-omni-flash-1.1 lists 4K, 3 to 10 s
Longest clip8 s15 s on most models, 30 s on seedance-2.5 and wan-3.0

What do I lose by leaving Veo?

Veo-only features stay Veo-only: Google's page describes extending a Veo-generated video by 7 seconds, up to 20 times, and that works only on videos Veo made. Sume does not list an extend call; chaining clips is the workaround. Match the model to the shot, not the brand, and test one short clip first.

Sources

Related posts

More in Models

All Models posts

Written by Sume