Wan 3.0 omni reference video: images, clips and audio in one prompt
Omni reference means one request carries mixed references. Wan 3.0 takes up to 10 images, 5 videos and 5 audio tracks; on Sume, send them as input_references.

Omni reference video is video generation where one request carries references of several kinds at once, such as character images, a motion clip and an audio track. Wan 3.0 does this: it accepts up to 10 images, up to 5 video clips and up to 5 audio tracks, and on Sume you send them in input_references to wan-3.0.
The vendor limits are from fal's Wan 3 page; Alibaba's API reference says references are addressed in the prompt as Image 1, Video 1 and so on, in list order. Sume's fields are from Video generation. All read 2026-09-29.
What can each reference do?
The docs say references are visual guidance rather than exact frames. What each type is used for is up to your prompt, so name it.
| Type | `type` value | Cap | Typical prompt role |
|---|---|---|---|
| Image | image_url | 10 | Who or what appears, or the style |
| Video | video_url | 5, 15 s in total | A motion or camera pattern to follow |
| Audio | audio_url | 5, 15 s in total | Sound or a voice to build on |
How do I send a mixed request?
Put each reference in input_references with its type, and say in the prompt what each is for. Do not also send frame_images: when both are present the request becomes image-to-video and the references are dropped.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: wan-omni-001" \
-d '{
"model": "wan-3.0",
"prompt": "The woman in Image 1 walks the way the person in Video 1 does, along a beach at sunset",
"duration": 8,
"resolution": "720p",
"input_references": [
{ "type": "image_url", "image_url": { "url": "https://example.com/woman.png" } },
{ "type": "video_url", "video_url": { "url": "https://example.com/walk.mp4" } }
]
}'Does the Image 1 wording work on Sume?
Alibaba documents it for its own API. Sume's docs do not describe indexed names for wan-3.0, so treat it as a prompt convention to test, not a guarantee, and describe each reference in plain words as well.
Which other models take video and audio references?
Per the Video generation docs, the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max honor audio and video references; Gemini Omni Flash 1.1 takes video but not audio. Read supported_input_references from GET /v1/videos/models before you build on one.
Sources
Related posts
More in Models
- Wan 3.0 thinking mode: why document and web page input needs it
fal says Wan 3.0 needs thinking mode to read a document or web page; Alibaba says prompt_extend must be true. What each means, and why Sume's wan-3.0 skips it.
- Wan 3.0 vertical video: 9:16 for Reels, Shorts and TikTok
Wan 3.0 supports 9:16. Send aspect_ratio 9:16 to wan-3.0 for a vertical clip of 2 to 30 seconds at 480p, 720p or 1080p. Accepted ratios and cost.
- Wan 3.0 vs MiniMax H3: clip length, ratios and reference caps
Wan 3.0 runs 2 to 30 s at up to 1080p; MiniMax H3 runs 5 to 15 s at 480p or 768p. Where they differ on Sume: ratios, reference caps and audio inputs.
- Wan 3.0 vs Seedance 2.0: which model to call when
Wan 3.0 and Seedance 2.0 are both on Sume's video API. Wan runs 2 to 30 s and bills per second; Seedance 2.0 runs 4 to 15 s with a 21:9 frame. Pick by the job.
Written by Sume