ElevenLabs Music Audio Reference vs Sume Music's image input
ElevenLabs Audio Reference guides Music v2/v2.5 with a track of about 30 seconds. Sume Music 1.0 takes a text prompt and one optional public HTTPS image URL.

Sume Music 1.0 has no audio reference input. Its request takes a text prompt and one optional public HTTPS image_url, so to steer a track you describe the sound in words, or pass a still image that sets the mood. ElevenLabs' Audio Reference, by contrast, lets you upload a short track to guide a Music v2 or v2.5 generation.
ElevenLabs facts are from its Eleven Music page; Sume facts from the Music 1.0 docs, both read 2026-10-01.
What does ElevenLabs Audio Reference do?
The page says Audio Reference lets you upload a short audio track to guide the style and sound of a new Music v2 or v2.5 generation, alongside your text prompt. You upload up to approximately 30 seconds in commonly used audio formats, and every reference is screened for copyright compliance. It influences overall sound, production style, instrumentation, tempo and mood, and the result stays a newly generated composition. It is available on all paid plans, and the page describes it in the web app under Music, then Generations.
What does Sume Music 1.0 accept instead?
| Input | ElevenLabs Music page | Sume Music 1.0 docs |
|---|---|---|
| Text prompt | Yes | Required, 1 to 5000 characters |
| Audio reference | About 30 s, v2/v2.5, paid plans | Not a request field |
| Image | Not listed on the page | Optional public HTTPS image_url |
| Output | Not stated on the page | Audio artifact, typically audio/mpeg on media.sume.com |
| Price | See ElevenLabs pricing | $0.125 per audio |
How do I steer a Sume track without a reference track?
Put the reference into the brief. The OpenAPI description suggests a per-scene music brief: emotion, genre, a BPM number, key and mode, lead instruments with texture, one named arc moment, era and production, then "Instrumental, no vocals." Exclusions go in the positive prompt, because negative_prompt is not supported. For examples of that structure see AI music prompt examples.
If your reference is visual, such as a mood still, send it as image_url. The docs' own example prompts for "the mood of the reference still". Price does not vary by prompt length or by the optional image.
What should I do if I need audio-to-audio matching?
Use ElevenLabs for that step, or write the characteristics you hear (tempo, instruments, mood) into the Sume prompt. Do not send duration or duration_seconds; Music 1.0 rejects them. Poll GET /v1/jobs/{id}/status, then read the audio artifact from result.artifacts[].
Sources
Related posts
More in Developers
- ElevenLabs STT 3 GB / 10 hour limit vs Sume STT 600 s jobs
ElevenLabs speech to text accepts files up to 3 GB and 10 hours in standard mode. Sume STT 1.0 reserves up to 10 minutes per job, so cut long recordings.
- eleven_turbo_v2_5 deprecated: use Flash; Sume TTS Router ids
ElevenLabs lists eleven_turbo_v2_5 as deprecated and suggests eleven_flash_v2_5. Sume's TTS Router pins an explicit catalog id, so list the ids first.
- ElevenLabs Voice Design 409: wait and retry, or new key on Sume?
Voice Design now returns 409 while a preview generation is in progress; ElevenLabs says retry. On Sume, reuse an Idempotency-Key only for the same payload.
- Signing the EU transparency Code of Practice after 27 July 2026
The Commission says you can still sign the AI-content transparency code after 27 July 2026; that date only set the initial list. Which section fits an API team?
Written by Sume