Seamless looping food video with AI: same first and last frame
To loop a food clip, send one image as image_url and end_image_url to Gemini Omni Flash 1.1, which Google says suits seamless loops. Clips run 3 to 10 seconds.

Use one image as both the first and the last frame. Google's Gemini Omni 1.1 Flash post says specifying the starting and ending frames makes it suitable for seamless looping clips, and Sume's Video Router takes a start frame as image_url and an optional end_image_url for gemini-omni-flash-1.1, with 3 to 10 second clips.
Vendor wording is from the Google post, Sume's from the Video Router docs, both read 2026-10-01.
How do I set up the loop?
The Video Router docs say the image_to_video capability takes image_url plus an optional end_image_url. For a loop, send your food still as both, and write a prompt where the motion returns to rest: steam rising and fading, sauce settling, a slow turntable that ends where it began.
// POST /v1/video-router/generate
{
"model": "gemini-omni-flash-1.1",
"prompt": "Steam drifts up from the bowl and fades; the camera holds still.",
"duration": 6,
"image_url": "https://media.sume.com/your-dish.png",
"end_image_url": "https://media.sume.com/your-dish.png"
}What does the catalog confirm for this model?
| Property | Value |
|---|---|
| Duration | 3 to 10 seconds |
| Resolutions | 360p, 720p, 1080p, 4K |
| Aspect ratios | 16:9 or 9:16 |
| Frame fields | image_url plus optional end_image_url |
| Audio | Native synced audio |
Is the loop point guaranteed to be invisible?
No. Matching end frames makes the jump between the last and first frame small, but motion, lighting and steam can still disagree at the seam, so play the clip on repeat before publishing. Google also lists 360p previews for iteration; check Sume's catalog entry for which resolutions you can request. If the loop is too short, increase video length by looping explains repeating it in a player or timeline.
Sources
Related posts
More in Use cases
- Gemini TTS two-speaker limit: three-voice dialogue with Sume concat
Gemini TTS configures two speakers per request. For three voices on Sume, make one TTS job per voice and join up to 20 parts with Timeline audio concat.
- Change a garment's colour in a photo with a mask_url edit
Recolour clothing with a mask: send the photo as an input reference, a public HTTPS mask_url, and a colour prompt to ChatGPT Image 2.5 on POST /v1/images.
- HeyGen avatar new outfit with reference images vs Sume
HeyGen prompt avatars take avatar_id plus up to three reference_images for a new outfit. Sume's photo input takes one image_url per avatar.
- Add a hook title to the first seconds of a video with one cue
Send one authored cue with start 0 and end 3 to POST /v1/video-captions and Sume burns that hook text into the clip, with no speech-to-text step.
Written by Sume