MiniMax H3 voice reference: match a voice with an audio clip
MiniMax H3 accepts up to 3 audio clips as references. How to write the prompt, the limits Sume enforces, and the rule that audio cannot be the only reference.

To make a MiniMax H3 character speak in a given voice, send the recording as an audio reference and name it in the prompt, for example "The character says: ... Match the voice in Audio 1." That wording is fal's guide for H3 reference-to-video; on Sume the audio goes in input_references next to an image or a video, because audio cannot be the only reference.
Vendor facts are from fal's prompting guide and MiniMax's model card; limits are from the Sume Video generation docs and validation code, read 2026-09-29.
What does the request look like?
One image for the face, one audio clip for the voice, and a prompt that assigns each a job.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: h3-voice-001" \
-d '{
"model": "minimax-h3",
"prompt": "Use Image 1 for the presenter. She looks into camera and says: Welcome back. Match the voice in Audio 1.",
"input_references": [
{ "type": "image_url", "image_url": { "url": "https://example.com/presenter.png" } },
{ "type": "audio_url", "audio_url": { "url": "https://example.com/voice-sample.mp3" } }
],
"resolution": "768p",
"aspect_ratio": "9:16",
"duration": 8
}'What are the audio limits?
| Limit | Value |
|---|---|
| Audio clips | Up to 3 |
| Length | Each 2–15 seconds, 15 seconds combined |
| Total references | At most 12 across images, videos and audio |
| Audio alone | Not allowed; pair with an image or a video |
| URLs | Public HTTPS |
Is this voice cloning?
fal's guide calls it "voice cloning from a recording". Sume's docs list audio references as honored by both H3 ids and say nothing about consent or rights. Use only voices you have the right to reproduce.
What if I already have the finished speech?
Then a lip-sync product fits better: Sume's MiniMax H3 Max lip-sync endpoint takes a still and your audio, and the output length follows the audio.
Sources
Related posts
More in Developers
- Promise.allSettled vs Promise.all for a batch of API jobs
Promise.allSettled waits for every promise and reports each outcome; Promise.all rejects on the first failure. For paid API jobs, use allSettled.
- Python API rate limiting: stay under a per-minute limit
Pace Python API calls with an asyncio limiter set under the API's per-minute budget, keep polling on its own budget, and back off on 429 retry-after.
- Python requests default timeout: there isn't one
Python Requests has no default timeout: without timeout= a call can hang indefinitely. Set (connect, read) on every call, and keep it short for job APIs.
- Real time speech to text API: what a file-based API can do
Sume's speech to text API isn't real time: it transcribes recordings of up to 10 minutes at a URL. Chunked recordings give near-live transcripts.
Written by Sume