2K AI video generator API: MiniMax H3 at 2K on Sume
MiniMax H3 makes video up to 2K. On Sume you ask for resolution 2K on minimax-h3, an upscale of native 768p. Request, price for 10 seconds, and limits.

To generate 2K video by API on Sume, send POST /v1/videos with model: "minimax-h3" and resolution: "2K". MiniMax's announcement describes H3 as generating up to 15 seconds at 2K; Sume's docs describe its 2K and 4K as priced upscales of the native 768p output.
Vendor facts are from MiniMax's announcement and model card, and Sume facts from the Video generation docs and pricing code, read 2026-09-29.
How do I ask for 2K?
Set resolution to 2K on minimax-h3. It does not appear in the model's supported_resolutions, so do not depend on that list to find it. minimax-h3-max does not accept 2K or 4K in current code.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: h3-2k-001" \
-d '{
"model": "minimax-h3",
"prompt": "A slow dolly along a rain-lit market street at night, neon signs reflecting on wet stone",
"resolution": "2K",
"aspect_ratio": "16:9",
"duration": 8
}'What does a 2K clip cost?
Per output second by resolution, reserved on submit at the provider list × 1.25, plus the 5.5% agent fee by default.
| Resolution | 10 seconds |
|---|---|
| 768p (native) | $0.75 |
| 2K (upscale) | $1.63 |
| 4K (upscale) | $2.00 |
Is Sume's 2K the same pipeline MiniMax describes?
MiniMax's model card describes a 2K step, H3-Regenerate-2K, which the card lists as an upscaling module that reuses context from the 768p base. Sume's docs say only that 2K and 4K are upscales of native 768p that are priced if requested; they do not say which provider step runs, so this post does not claim one.
What are the length and ratio limits?
5–15 seconds in whole seconds, and aspect ratios 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16. For anything that is not 2K, remember that 720p is refused on H3: use 768p.
Sources
Related posts
More in Models
- What is MiniMax H3 Max? The post-trained variant, explained
MiniMax H3 Max is a variant post-trained by fal.ai on MiniMax H3 for faster generation. Its resolutions, lengths and modes, and how Sume lists it.
- MiniMax H3 Max lip sync API: a still plus audio, 5 to 14.8 seconds
Sume runs MiniMax H3 Max lip sync at POST /v1/minimax/h3-max/lip-sync: send a still and Sume-hosted audio of 5 to 14.8 seconds. Body, resolutions and price.
- AI video with native stereo sound: MiniMax H3 audio through the API
MiniMax H3 generates native stereo sound in the same pass as the video. On Sume audio is always on, generate_audio false is refused, and audio steers.
- MiniMax H3 open weights on Hugging Face: what is in the release
MiniMax H3 has a model card on Hugging Face: two checkpoints, a 33B-parameter Transformer, a community license and a 4-GPU serving example. What to check first.
Written by Sume