Suno v6 audio and video input vs Sume: text and one image
Suno v6 says you can create with text, audio, images and video. Sume's music request takes a text prompt and one optional image; audio and video are not inputs.

Of the four input types Suno v6 names, Sume's music models accept two: a text prompt and one optional image. Audio and video are not inputs to POST /v1/music-1.0/generate or the Music Router. The request has no field for them.
Suno's list is from its v6 announcement. Sume's facts are from Music 1.0 and Music Router, read 2026-09-30.
What inputs does Suno v6 list?
The announcement says you can "Create with text, audio, images and video", and also lists building a mashup from multiple sources and sampling or isolating audio to make a new beat. It does not give limits for those inputs.
Which of those can I send to Sume?
Two of the four. The docs say a request is a text prompt plus optional image conditioning, and the image must be public HTTPS. A music request carries nothing else that could hold a sound or a clip.
| Input type | Suno v6 (announcement) | Sume music request |
|---|---|---|
| Text | Listed | Yes, prompt |
| Images | Listed | Yes, one optional image_url |
| Audio | Listed | No |
| Video | Listed | No |
Can I use a video's sound some other way?
Sume can pull the audio track out of a hosted video with audio detach, and join or split Sume-hosted audio with Timeline audio. Those produce audio files; they do not feed a music request, which has no audio field.
What should I do for a picture-led track?
Describe the mood, tempo and instruments in prompt and pass the still as image_url. For the image-input side across engines, see Lyria 3.5 reference images vs Sume's image_url.
Sources
Related posts
More in Models
- Suno v6-wild: is there a less predictable setting on Sume?
Suno v6-wild is a paid model that is less predictable. Sume's music request has no seed, temperature or guidance field, so variety comes from the prompt.
- Tavus Phoenix-4.5: photo rules vs a Sume photo avatar
Tavus Phoenix-4.5 starts a face from a photo in minutes and allows glasses, jewelry and hair. Sume creates an avatar from a public HTTPS image URL.
- TikTok Symphony with Seedance 2.5: 30-second AI video ads
TikTok Symphony now generates up to 30 seconds with Seedance 2.5, in select markets. The same 30-second length is available by API on Sume as seedance-2.5.
- Vidu API: S2-Avatar live voice vs Sume avatar video jobs
Vidu S2-Avatar is a real-time voice model. Sume has no live session: you submit a script and a scene photo to an avatar job and fetch the finished video.
Written by Sume