What is MiniMax H3 Max? The post-trained variant, explained
MiniMax H3 Max is a variant post-trained by fal.ai on MiniMax H3 for faster generation. Its resolutions, lengths and modes, and how Sume lists it.

MiniMax H3 Max is a video model post-trained by fal.ai on MiniMax H3 and optimized for high-speed generation, according to MiniMax's own API docs. It is a variant of H3, not a newer generation: same text, image-frame and reference modes, but with 480p and 768p as its mainstream outputs, and 5–15 second clips.
Vendor facts are from MiniMax's API docs and fal's model page, and Sume facts from its Video generation and Video Router docs, all read 2026-09-29.
How does H3 Max differ from H3?
| Property | MiniMax H3 | MiniMax H3 Max |
|---|---|---|
| Origin | MiniMax's open multimodal video model | Post-trained by fal.ai on H3, tuned for speed |
| MiniMax API docs resolutions | 768P and 2K | 480P and 768P |
| MiniMax API docs duration | 4–15 seconds, integers | 5–15 seconds, integers |
| Sume id | minimax-h3: 480p and 768p native, 2K and 4K priced upscales | minimax-h3-max: 480p, 768p and 1080p |
| Modes on Sume | Text, first/last frame, reference | Text, first/last frame, reference |
Why does H3 Max list 1080p on Sume and fal but not in MiniMax's docs?
fal's model page lists 480p, 768p and 1080p for H3 Max. Sume's docs describe its 1080p as a latent refinement from native 768p. MiniMax's own docs page lists only 480P and 768P for it. When resolutions differ by source, the provider you call decides; read supported_resolutions from GET /v1/videos/models on Sume.
Does H3 Max keep the native stereo sound?
On Sume, yes: the docs say minimax-h3-max is text-to-video, first/last-frame and reference-to-video with native stereo audio, and that it reports image, video and audio supported_input_references. Sume rejects generate_audio: false for both H3 ids.
Should I start with H3 or H3 Max?
The Sume docs describe H3 Max as the 768p variant tuned for speed. If you need 2K, only minimax-h3 offers it in code, as an upscale. If you want 1080p, use minimax-h3-max. Both are billed as list × 1.25; the per-second rates differ by id.
Sources
Related posts
More in Models
- MiniMax H3 Max lip sync API: a still plus audio, 5 to 14.8 seconds
Sume runs MiniMax H3 Max lip sync at POST /v1/minimax/h3-max/lip-sync: send a still and Sume-hosted audio of 5 to 14.8 seconds. Body, resolutions and price.
- AI video with native stereo sound: MiniMax H3 audio through the API
MiniMax H3 generates native stereo sound in the same pass as the video. On Sume audio is always on, generate_audio false is refused, and audio steers.
- MiniMax H3 open weights on Hugging Face: what is in the release
MiniMax H3 has a model card on Hugging Face: two checkpoints, a 33B-parameter Transformer, a community license and a 4-GPU serving example. What to check first.
- AI video editing with a text prompt: MiniMax H3 and Sume's edit path
MiniMax H3 is described as editing existing video from instructions. What the vendor says, what Sume's H3 ids accept, and the id with an edit field.
Written by Sume