ElevenLabs Music Audio Reference vs Sume Music's image input

ElevenLabs Audio Reference guides Music v2/v2.5 with a track of about 30 seconds. Sume Music 1.0 takes a text prompt and one optional public HTTPS image URL.

4 min readSume
All posts

Sume Music 1.0 has no audio reference input. Its request takes a text prompt and one optional public HTTPS image_url, so to steer a track you describe the sound in words, or pass a still image that sets the mood. ElevenLabs' Audio Reference, by contrast, lets you upload a short track to guide a Music v2 or v2.5 generation.

ElevenLabs facts are from its Eleven Music page; Sume facts from the Music 1.0 docs, both read 2026-10-01.

What does ElevenLabs Audio Reference do?

The page says Audio Reference lets you upload a short audio track to guide the style and sound of a new Music v2 or v2.5 generation, alongside your text prompt. You upload up to approximately 30 seconds in commonly used audio formats, and every reference is screened for copyright compliance. It influences overall sound, production style, instrumentation, tempo and mood, and the result stays a newly generated composition. It is available on all paid plans, and the page describes it in the web app under Music, then Generations.

What does Sume Music 1.0 accept instead?

Inputs compared, from the two pages read 2026-10-01.
InputElevenLabs Music pageSume Music 1.0 docs
Text promptYesRequired, 1 to 5000 characters
Audio referenceAbout 30 s, v2/v2.5, paid plansNot a request field
ImageNot listed on the pageOptional public HTTPS image_url
OutputNot stated on the pageAudio artifact, typically audio/mpeg on media.sume.com
PriceSee ElevenLabs pricing$0.125 per audio

How do I steer a Sume track without a reference track?

Put the reference into the brief. The OpenAPI description suggests a per-scene music brief: emotion, genre, a BPM number, key and mode, lead instruments with texture, one named arc moment, era and production, then "Instrumental, no vocals." Exclusions go in the positive prompt, because negative_prompt is not supported. For examples of that structure see AI music prompt examples.

If your reference is visual, such as a mood still, send it as image_url. The docs' own example prompts for "the mood of the reference still". Price does not vary by prompt length or by the optional image.

What should I do if I need audio-to-audio matching?

Use ElevenLabs for that step, or write the characteristics you hear (tempo, instruments, mood) into the Sume prompt. Do not send duration or duration_seconds; Music 1.0 rejects them. Poll GET /v1/jobs/{id}/status, then read the audio artifact from result.artifacts[].

Sources

Related posts

More in Developers

All Developers posts

Written by Sume