MiniMax H3 voice reference: match a voice with an audio clip

MiniMax H3 accepts up to 3 audio clips as references. How to write the prompt, the limits Sume enforces, and the rule that audio cannot be the only reference.

4 min readSume
All posts

To make a MiniMax H3 character speak in a given voice, send the recording as an audio reference and name it in the prompt, for example "The character says: ... Match the voice in Audio 1." That wording is fal's guide for H3 reference-to-video; on Sume the audio goes in input_references next to an image or a video, because audio cannot be the only reference.

Vendor facts are from fal's prompting guide and MiniMax's model card; limits are from the Sume Video generation docs and validation code, read 2026-09-29.

What does the request look like?

One image for the face, one audio clip for the voice, and a prompt that assigns each a job.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: h3-voice-001" \
  -d '{
    "model": "minimax-h3",
    "prompt": "Use Image 1 for the presenter. She looks into camera and says: Welcome back. Match the voice in Audio 1.",
    "input_references": [
      { "type": "image_url", "image_url": { "url": "https://example.com/presenter.png" } },
      { "type": "audio_url", "audio_url": { "url": "https://example.com/voice-sample.mp3" } }
    ],
    "resolution": "768p",
    "aspect_ratio": "9:16",
    "duration": 8
  }'

What are the audio limits?

Reference limits on minimax-h3 and minimax-h3-max, from the model card and Sume code, read 2026-09-29.
LimitValue
Audio clipsUp to 3
LengthEach 2–15 seconds, 15 seconds combined
Total referencesAt most 12 across images, videos and audio
Audio aloneNot allowed; pair with an image or a video
URLsPublic HTTPS

Is this voice cloning?

fal's guide calls it "voice cloning from a recording". Sume's docs list audio references as honored by both H3 ids and say nothing about consent or rights. Use only voices you have the right to reproduce.

What if I already have the finished speech?

Then a lip-sync product fits better: Sume's MiniMax H3 Max lip-sync endpoint takes a still and your audio, and the output length follows the audio.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume