Lyria 3.5 takes up to 10 images; Sume music takes one image_url
Google's Lyria guide allows up to 10 images as input. Sume's music docs describe one optional public HTTPS image_url. What that changes.

Google's Lyria guide says you can give a prompt up to 10 images. Sume's Music 1.0 docs describe one optional image_url, a public HTTPS link. Read 2026-10-01. If you want more than one picture to set the mood, describe the rest in words.
How do I use one image?
Pass image_url with your prompt, for example a cover or a moodboard frame, and say in the text what the image means for the music.
| Service | Image input |
|---|---|
| Google Lyria guide | Up to 10 images |
| Sume Music 1.0 | One optional public HTTPS image_url |
What if I have a moodboard of ten?
Pick the most telling frame, or write the mood, palette and setting into the prompt (up to 5,000 characters). Sume's docs make no promise about how closely music follows a picture.
Does the image replace the text prompt?
No. Keep the text prompt (up to 5,000 characters) as the main instruction and use the image as extra context.
Sources
Related posts
More in Models
- Midjourney edit model image references: 4 vs Sume's ranges
Midjourney's V8.2 edit model takes up to 4 image references. On Sume the ceiling is per model: GPT Image 2.5 takes up to 16. Read the descriptor first.
- minimax/hailuo-3 on OpenRouter vs Sume minimax-h3 id
OpenRouter names the model minimax/hailuo-3. Sume uses bare catalog ids, so the matching id is minimax-h3, native 480p/768p, with 2K and 4K as priced upscales.
- Nano Banana multiple images at once: n range vs Gemini's count
Google says Gemini won't always return the exact image count you ask for in a prompt. On Sume, set n, and read each model's n range from the catalog.
- Nano Banana video to image: poster from a video via Sume stills
Google's Gemini API takes a video as context for a thumbnail or poster. Sume's images API takes image references only, so grab a still frame first.
Written by Sume