MiniMax H3 first_frame and reference roles can't mix: Sume rule
MiniMax H3 treats first/last frames and reference roles as mutually exclusive. On Sume, frame_images takes precedence over input_references. Which to send.

MiniMax's H3 create-task docs say image-to-video and reference-to-video are mutually exclusive: if any reference role appears in content, first_frame and last_frame must not, and the two cannot be mixed. Sume's video docs describe a different rule: when both frame_images and input_references are sent, frame_images takes precedence and the request is treated as image-to-video.
Sources: the MiniMax create-task page and Sume's Video generation docs, read 2026-10-01.
What does the MiniMax rule say?
Reference-to-video takes any combination of reference images (role=reference_image), reference videos (role=reference_video), and reference audio (role=reference_audio). Image-to-video uses the first and last frame roles. Put either family in content and the other must be absent, "and vice versa".
What does Sume do with both fields?
Sume has two request fields. frame_images carries first or last frame images (each entry needs a frame_type of first_frame or last_frame) and means image-to-video. input_references carries style or content references the model uses as guidance rather than exact frames, and means reference-to-video. Send both and frame_images takes precedence. The docs describe the outcome as image-to-video; they do not describe an error for the mix.
| Request | MiniMax H3 docs | Sume docs |
|---|---|---|
| Frames only | Image-to-video | frame_images: image-to-video |
| References only | Reference-to-video | input_references: reference-to-video |
| Both | Not allowed | frame_images takes precedence; treated as image-to-video |
What should I send for each goal?
Pick the mode first and send only its field. If you need a known opening or closing image, send frame_images. If you want the look or content of references to guide a fresh clip, send input_references alone. Do not rely on precedence to combine them: the reference entries would not shape the result as reference-to-video.
Only models whose supported_input_references lists a type accept that type. Audio and video references are honored by MiniMax H3 and MiniMax H3 Max among others, so check the catalog entry before sending them. More on the limits in MiniMax H3 reference-to-video limits.
{
"model": "minimax-h3",
"prompt": "A character walking through a forest",
"frame_images": [
{
"type": "image_url",
"image_url": { "url": "https://example.com/first-frame.png" },
"frame_type": "first_frame"
}
],
"resolution": "768p"
}Sources
Related posts
- MiniMax H3 reference limits: 9 images, 3 videos, 3 audio, 12 total
- First and last frame to video with MiniMax H3: request and rules
- MiniMax H3 Max reference images: 2 free, then per image
- MiniMax H3 API: H3 and H3 Max video at native 768p with stereo audio
- Video to video motion transfer with AI: MiniMax H3 reference video
More in Developers
- MiniMax H3 reference input limits: 30 MB, 50 MB, 64 MB request
MiniMax H3 caps images at 30 MB, reference video at 50 MB and the request at 64 MB, and says to use public URLs. A pre-flight list, and how Sume takes URLs.
- Modal 150-second web timeout and 303 redirect: Sume polling
Modal web endpoints return a 303 redirect after 150 seconds. Sume returns 202 with status_url and result_url, so a client polls and follows no redirects.
- Capacity fallback vs Sume queued jobs at full concurrency
Runway's Model Router can fall back to another model at its concurrency limit. Sume instead accepts the job as queued on the same model until a slot opens.
- Nano Banana batch API: Gemini's 24 h batch vs Sume async jobs
Gemini's Batch API trades up to 24 hours of turnaround for higher rate limits. Sume has no batch tier for images: send async or webhook jobs per request.
Written by Sume