Multi-shot AI video prompts: MiniMax H3 shot labels and timecodes
MiniMax H3 models multiple shots natively. Write [Shot 1] labels or timecoded blocks in the prompt; syntax from the vendor guides and a Sume request.

To get several shots in one MiniMax H3 clip, write them into the prompt: MiniMax's prompt guide labels shots [Shot 1], [Shot 2] and gives later shots a timestamp such as "At 00:03.500, the camera cuts to...", while fal's guide uses timecoded blocks like [0–2 seconds]. The model does multi-shot natively, and Sume has no separate shots field: the storyboard is prompt text.
Syntax is from MiniMax's prompt guide and fal's guide, and MiniMax's announcement for the native multi-shot claim; request fields from the Sume Video generation docs, read 2026-09-29.
What does a multi-shot prompt look like?
Keep to the length you buy: duration is 5–15 seconds on Sume, so timecodes past your duration have nothing to land on. The shot order below adds up to 10 seconds.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: h3-shots-001" \
-d '{
"model": "minimax-h3",
"prompt": "[0-3 seconds] Wide shot of a lighthouse at dusk, waves breaking. [3-6 seconds] The camera cuts to a close-up of the keeper lighting the lamp. [6-10 seconds] Pull back outside as the beam sweeps across the water. Sound: wind, gulls, a low cello.",
"resolution": "768p",
"aspect_ratio": "16:9",
"duration": 10
}'Which syntax should I use?
| Source | Convention |
|---|---|
| MiniMax prompt guide | [Shot 1], [Shot 2] labels; later shots timestamped; transitions worded as "the camera cuts to", "the shot transitions to" or "the shot changes to" |
| fal prompting guide | Timecoded blocks such as [0–2 seconds] ... [10–15 seconds]; recommends storyboarding inside the prompt, with up to 7,000 characters |
What keeps a multi-shot clip coherent?
MiniMax's guide says to establish the style at the start of Shot 1, keep the character consistent and describe observable changes. fal's advice is to describe transitions as physical events rather than named effects. If you have character images, send them as references and name them in the prompt so each shot reuses the same look.
Does Sume validate the shots?
No. Sume checks the request envelope, such as duration, resolution and aspect ratio, and sends the prompt to the provider. Whether shot 2 lands at 3 seconds is the model's behavior, so check the output.
Sources
Related posts
More in Developers
- MiniMax H3 API in Python: submit, poll and download a video job
A runnable Python example for MiniMax H3 on Sume: POST /v1/videos, poll the job until completed, then download the clip from unsigned_urls. Uses httpx.
- MiniMax H3 reference limits: 9 images, 3 videos, 3 audio, 12 total
The MiniMax H3 reference-to-video limits and the errors Sume returns when you cross them. Counts, clip lengths, the audio-only rule and image price.
- Video to video motion transfer with AI: MiniMax H3 reference video
MiniMax H3 lists V2V motion transfer. Send a motion video and character images as references on Sume, write each reference's job, and mind the limits.
- MiniMax H3 voice reference: match a voice with an audio clip
MiniMax H3 accepts up to 3 audio clips as references. How to write the prompt, the limits Sume enforces, and the rule that audio cannot be the only reference.
Written by Sume