Grok Imagine video editing: 8.7 s on xAI, Omni edit on Sume
xAI caps Grok Imagine video edits at 8.7 seconds. Sume's Grok row has no video-to-video; edits go through gemini-omni-flash-1.1 with a video_url.

Grok Imagine video editing is in xAI's docs with an 8.7 second cap on the edited video. The Sume Grok row has video_to_video: false; to edit a clip on Sume, send prompt plus video_url to gemini-omni-flash-1.1.
xAI's cap is from its video docs, read 2026-09-30. Sume's edit rules are from the Video Router docs and catalog.
What does xAI say about editing?
Its feature list includes video editing, and it states that edited video is capped at 8.7 seconds.
What is Sume's edit route?
The Video Router lists video_to_video (edit) for gemini-omni-flash-1.1: send video_url, with a prompt that describes the edit. resolution is optional and defaults to 720p. aspect_ratio is rejected and duration is not sent.
| Where | Edit mode |
|---|---|
| xAI, Grok | Video editing, 8.7 s cap |
| Sume, Grok | None (video_to_video: false) |
Sume, gemini-omni-flash-1.1 | video_url + prompt, resolution optional (default 720p) |
What are the input rules?
video_url is the edit source, not a reference: it cannot be combined with image_url, end_image_url or reference_*_urls. Prompting tips are in Gemini Omni video edit prompts.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: video-router-edit-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "Replace the bottle with an apple. Keep everything else the same.",
"video_url": "https://example.com/clip.mp4",
"resolution": "720p",
"mode": "async"
}'What should I do?
Do not port a Grok edit call to the Sume Grok row. Rewrite it as an Omni edit, and check the Omni duration limits for your source clip.
Sources
Related posts
More in Models
- Grok Imagine video sound: xAI audio vs Sume's silent row
xAI's Grok Imagine video makes audio unless generate_audio is False. Sume's grok-imagine-video-1.5 row lists audio: false and rejects generate_audio.
- HeyGen ElevenLabs v3 model_id per request vs Sume TTS
HeyGen's speech endpoint takes settings.model_id such as eleven_v3 per request. Sume TTS 1.0 hides model ids; its separate TTS Router lists models explicitly.
- HeyGen photo avatar hand gestures vs Sume Motion Control
HeyGen's motion_prompt still needs an animation reference for Avatar V photo avatars. On Sume, body motion is a separate Motion Control route, not a prompt.
- HeyGen text to video API: 5-15 s, 768p vs Sume durations
heygen-video-1 makes 5-15 second clips at 480p or 768p from a 5,000-character prompt. Sume lists durations and resolutions per model in its catalog.
Written by Sume