AI video editing with a text prompt: MiniMax H3 and Sume's edit path
MiniMax H3 is described as editing existing video from instructions. What the vendor says, what Sume's H3 ids accept, and the id with an edit field.

MiniMax says H3 does generalized reference and editing across text, images, video and audio, so you can change an existing clip by describing the change. On Sume the two H3 ids take a reference video with a prompt, but the docs give no dedicated edit mode for them; the id the docs name for a video_url edit is gemini-omni-flash-1.1, through the Video Router.
Vendor claims are from MiniMax's announcement, the Hailuo AI page and fal's prompting guide; Sume facts from the Video generation and Video Router docs and request validation code, read 2026-09-29.
What edits does MiniMax describe?
The Hailuo AI page lists replacing, adding or removing people and objects, and changing backgrounds, lighting, effects, dialogue and voice, with untouched content staying stable. fal's guide describes "precise localized edits to video you already have" through H3's reference-to-video endpoint, and shows prompts such as "Replace the cat in the video with a dog". Its advice is to state each change and pair it with what stays stable.
Can I edit a video with minimax-h3 on Sume?
You can send a video as a reference: up to 3 videos, each 2–15 seconds and 15 seconds combined, alongside images and audio, 12 files at most. You then write the edit as the prompt. Sume does not document that as an edit mode, does not promise the rest of the shot is preserved, and the reference does not set the output length: duration does. Try it on a short clip and compare before relying on it.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: h3-edit-001" \
-d '{
"model": "minimax-h3",
"prompt": "Use Video 1 as the source shot. Replace the cat with a dog; keep the room, camera move and lighting unchanged.",
"input_references": [
{ "type": "video_url", "video_url": { "url": "https://example.com/cat-shot.mp4" } }
],
"resolution": "768p",
"duration": 5
}'Which Sume model has a dedicated edit field?
gemini-omni-flash-1.1. Send video_url and a prompt to POST /v1/video-router/generate; the prompt describes the edit, and aspect_ratio is rejected. Other ids reject video_url. Edit a video with a prompt walks through that request.
| `minimax-h3` with a video reference | `gemini-omni-flash-1.1` with `video_url` | |
|---|---|---|
| Endpoint | POST /v1/videos | POST /v1/video-router/generate |
| Source clip goes in | input_references as video_url | video_url |
| Documented as an edit mode | No | Yes |
| Sound | Native stereo, always on | Native synced audio, always on |
| Source length | 2–15 s per reference video | No aspect_ratio; duration is not sent |
Does MiniMax rank H3 for editing?
The Hailuo AI page says blind human testing puts H3 second in text-to-video with audio and first in video editing with audio. The page attributes that to Artificial Analysis Video Arena scores dated 2026-07-31. It is the vendor's claim, so this post does not rely on it.
Sources
Related posts
More in Models
- Nano Banana 2 Lite API: what Google shipped and what Sume lists
Nano Banana 2 Lite is Google's gemini-3.1-flash-lite-image model. What Google says it does, and which Nano Banana models Sume's image API lists today.
- Nano Banana Pro vs Nano Banana 2: which id to send on Sume
Nano Banana Pro and Nano Banana 2 share tiers and 10 references on Sume; they differ in extra aspect ratios and in which tiers change the price.
- Open-source video model vs API: which to use for Wan-class video
Weights you run yourself, or a hosted video API? What the choice changes for cost, setup and limits, with Wan 3.0's per-second API price as the worked example.
- Remove shadow from photo with AI: edit it or cut it out
To remove a shadow from a photo with AI, send the photo to an image-edit model, name the one shadow to remove, and list what stays. Prompts, cost, limits.
Written by Sume