AI music stems: does the API return separate tracks?

Sume's music API returns one audio file, not stems. What Lyria 3.5's documentation covers, and how to get parts you can mix separately.

4 min readSume
All posts

No: Sume's music API returns one mixed audio file per job, not stems. The result has an audio artifact in result.artifacts[] and, when present, a lyrics field, and the Music Router and Music 1.0 docs list no stem, layer or loop option. Google's Lyria page, read for this post, does not document stems either.

Sume facts are from the Music Router, Music 1.0 and Timeline audio docs; Google's from its Lyria page. Read 2026-09-29. If a stem separator is on your list, Sume does not currently list one.

What does a music job return?

From the Sume Music Router docs, read 2026-09-29.
FieldContent
result.artifacts[] where type is audioOne track, typically audio/mpeg on media.sume.com
result.lyricsModel-reported lyrics or section map, when present
Stems, layers, loop pointsNot documented

How can I get parts I can mix separately?

Generate the parts as their own tracks, from separate briefs, and mix them yourself. The prompt is the only control, so describe one role per brief and keep the tempo and key identical.

  • Brief 1: "Drums and bass only, 100 BPM, E minor. A 30-second track."
  • Brief 2: "Pad and piano only, 100 BPM, E minor. A 30-second track."
  • Repeat the tempo and key in every brief. Neither is a guaranteed setting, so measure each take.
  • Expect the parts to be independent performances: they will not be sample-locked to each other the way real stems are.
curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: stem-drums-001" \
  -d '{
    "prompt": "Drums and bass only, no melody, 100 BPM, E minor. Tight kick, round bass. A 30-second track. Instrumental, no vocals."
  }'

Can I split or edit the track I got?

In time, yes; by instrument, no. Timeline audio split cuts a Sume-hosted file into 1 to 20 time ranges and returns each as its own file, WAV by default, and concat joins up to 20 parts. Those tools work on time, not on which instrument is playing, and the docs describe no re-synthesis. The docs also list no audio input for the music request: only prompt and an optional image_url, so you cannot send a track in to be separated or remixed.

What if I only need the voice separate from the music?

Do not mix them until the end. Generate the narration with text-to-speech, the music with the Music Router, and combine them in Timeline 1.0, where a soundtrack bed can sit under the voice with gain_db, loop, fade_out_seconds and duck_db. That gives you a voice track and a music track that stay separate until the render.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume