Omni 1.1 Flash start and end frame API on Sume
Omni 1.1 Flash can render between two keyframes, including looping clips. On Sume, Omni takes image_url plus end_image_url; frame_images covers other models.

On Sume, give gemini-omni-flash-1.1 a start frame as image_url and an end frame as end_image_url through the Video Router. Google says Omni 1.1 generates continuous video between two keyframes, which suits camera orbits, zoom transitions and seamless loops; to loop, use the same image for both.
Google's claim is from its Omni 1.1 Flash developer post. Sume's request shapes are from Video Router and Video generation, read 2026-10-01.
What does Google say about start and end frames?
The post says you can specify the starting and ending frames of a shot to get smooth transitions and camera movements, and names complex orbits, zoom transitions and seamless looping clips as use cases. Its examples ask for "one continuous shot, no jump cuts", and one ends on a frame that matches the scene from the beginning.
Which Sume fields carry the two frames?
Two surfaces exist, and they differ. In the Video Router, the image_to_video capability takes image_url plus an optional end_image_url. In the generic Videos API, frame_images entries each carry a frame_type of first_frame or last_frame, and each model advertises what it accepts in supported_frame_images. For Omni, follow the Video Router row and check the catalog before assuming a field.
If you send both frame_images and input_references to the Videos API, frame_images takes precedence and the request is image-to-video.
What limits apply to the clip itself?
| Setting | Documented value |
|---|---|
| Duration | 3–10 seconds |
| Resolution | 360p, 720p, 1080p or 4K |
| Aspect ratio | 16:9 or 9:16 |
| Audio | Native synced audio; the Video Router says generate_audio: false is rejected |
| Frame fields | image_url + end_image_url (Video Router) |
How do I make a looping clip?
Send the same image as both the start and end frame, describe a motion that returns to its starting pose in the prompt, and keep the duration inside 3–10 seconds. Whether the loop looks seamless depends on the prompt and image; Sume's docs make no promise about it. For the neighboring interpolation topic see Gemini Omni first and last frame interpolation.
Sources
Related posts
More in Developers
- Text to speech mulaw 8000 Hz: Gemini 3.8 and Sume TTS
Sume TTS can return pcm_mulaw or pcm_alaw at 8000 Hz in wav or raw containers. Gemini 3.8 TTS does it with audio/mulaw and audio/alaw mime types.
- Gemini TTS voice design voice_ id vs Sume voice ids
Gemini voice design returns a persistent voice_ id from a text prompt. Sume TTS accepts only a voice UUID or voi_ library id, and rejects other shapes with 400.
- GPT Image 2 input_fidelity: omit it; Sume returns 400
OpenAI says to omit input_fidelity for gpt-image-2 because inputs run at high fidelity. Sume lists no such field and rejects unlisted parameters with 400.
- GPT Image moderation_blocked vs Sume content_policy_rejected
OpenAI returns moderation_blocked with moderation_details. On Sume, policy refusals are grouped under content_policy_rejected. What to read and when to retry.
Written by Sume