MiniMax H3 Max reference audio needs an image or video on Sume
Krea says MiniMax H3 Max audio references cannot be sent alone; Sume's reference_audio_urls has the same rule. Limits for images and videos compared.

You cannot send reference audio on its own to MiniMax H3 Max reference-to-video, and Sume has the same rule: reference_audio_urls requires at least one reference image or video. Krea's September 1 changelog adds one limit Sume's video page does not name: up to 7 reference videos, 15 seconds combined, reported with reference_video_seconds.
Krea facts are from its changelog (preview feature); Sume facts are from the video input fields and video generation docs, read 2026-10-01.
What does Krea say about H3 Max references?
MiniMax H3 Max supports reference-to-video in preview. You pass image, video, and audio references together to keep characters, motion, and voice consistent. Krea states the video limit as up to 7 clips, 15 seconds combined, and says audio references cannot be sent alone.
How do the reference limits compare?
The counts differ, so do not copy a payload from one API to the other.
| Reference | Krea (H3 Max preview) | Sume video request |
|---|---|---|
| Images | Not stated in the snapshot | reference_image_urls: 1 to 9 URLs |
| Videos | Up to 7, 15 seconds combined | reference_video_urls: 1 to 3 URLs |
| Audio | Cannot be sent alone | reference_audio_urls: 1 to 3 URLs; needs at least one image or video |
Which Sume models take audio and video references?
The video generation docs say only models whose supported_input_references lists a type accept it. Audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3, and MiniMax H3 Max. Check GET /v1/models for the model you intend to use before you submit.
For a related limit on another vendor, see Kling 4.0: five reference videos.
What do I do if I only have a voice track?
Add at least one reference image or video of the subject, then pass the audio next to it. A lone reference_audio_urls array fails the rule above.
Sources
Related posts
More in Models
- LTX-2.5 Cinemagraph LoRA vs image-to-video on Sume
The LTX-2.5 Cinemagraph LoRA is an image-to-video adapter for selective motion. On Sume, start from first_frame and describe the motion in the prompt.
- Luma layers API: 10 RGBA layers vs Sume's flat image result
Luma type layering splits one image into up to 10 ordered RGBA PNG layers. Sume returns flat images at data[].url; n counts images, not layers.
- Ray 3.2 API aspect ratios and resolutions, and Sume's lists
Luma lists six Ray 3.2 aspect ratios (9:16 to 21:9) at 540p, 720p and 1080p. Sume reports ratios per model in supported_aspect_ratios and rejects size.
- Luma Scenes Ray 3.2 or Seedance 2: pin a model or use auto
Luma says Ray 3.2 and Seedance 2 are not interchangeable. On Sume, pin a catalog model for a known clip length, or send sume/auto and let Sume pick.
Written by Sume