MiniMax H3 in ComfyUI or through a hosted API: what differs
ComfyUI documents three MiniMax H3 workflows with local model files; a hosted API returns a job. The limits that change between the two, from the docs.

ComfyUI runs MiniMax H3 from local model files in text-to-video, image-to-video and reference-to-video workflows, while a hosted API such as Sume's POST /v1/videos takes the same three modes as one JSON request and returns a job. Pick ComfyUI when you need graph-level control, and the API when you need to call it from code without GPUs.
ComfyUI facts are from its H3 tutorial and MiniMax's model card; Sume facts are from the Video generation docs, all read 2026-09-29.
What does ComfyUI need for H3?
Per its tutorial: a diffusion model variant (fl2va for text and image-to-video, ref2va for reference-to-video), a text encoder, a video VAE and an audio VAE. Optional turbo LoRAs cut sampling to 8 steps for text and image-to-video and 4 steps for reference-to-video, against a 20-step default. The tutorial also documents a MiniMaxH3AddGuide node that anchors images or audio at any frame, and latent noise masks for regenerating part of a clip.
How do the limits compare?
| Topic | ComfyUI and model card | Sume `minimax-h3` |
|---|---|---|
| Native canvas | 768 px short edge, sizes rounded to a multiple of 32 | Native 480p and 768p; 2K and 4K are priced upscales |
| Length | Duration snaps to a 17-frame-per-block grid at 24 fps | 5–15 seconds, whole seconds |
| References | Up to 9 images, 3 videos, 3 audio clips | Image, video and audio types; counts are not in the docs |
| Where it runs | Your hardware | Sume's provider, billed per output second |
Can I mask and regenerate part of a clip through the API?
Sume does not list that. Its H3 ids take a prompt, frames or references; the latent-mask editing in ComfyUI is a graph feature. If you need to change one region of a clip in Sume, the docs list gemini-omni-flash-1.1's video_url edit mode through the Video Router, a different model.
Which one should I start with?
If you have not run H3 before, one hosted request shows what the model does with your prompt before you set up drivers, encoders and VAEs. If you already have the GPUs and want turbo LoRAs or guide anchoring, start in ComfyUI. Check the license on the model card either way.
Sources
Related posts
More in Developers
- First and last frame to video with MiniMax H3: request and rules
MiniMax H3 fills the motion between an opening and a closing image. On Sume send frame_images with first_frame and last_frame; ratio, length and price rules.
- Multi-shot AI video prompts: MiniMax H3 shot labels and timecodes
MiniMax H3 models multiple shots natively. Write [Shot 1] labels or timecoded blocks in the prompt; syntax from the vendor guides and a Sume request.
- MiniMax H3 prompt guide: reference jobs, negative direction, length
The MiniMax H3 prompting rules that change output: give each reference a job, write timecodes, say what to avoid, and stay under 7,000 characters.
- MiniMax H3 API in Python: submit, poll and download a video job
A runnable Python example for MiniMax H3 on Sume: POST /v1/videos, poll the job until completed, then download the clip from unsigned_urls. Uses httpx.
Written by Sume