Flow Music edits a song by section: how to do it by API

Google Flow Music edits songs by section; its Lyria API page says multi-prompt editing is unsupported. On Sume, cut two takes and join the parts you keep.

4 min readSume
All posts

Google Flow Music advertises the ability to "precisely edit your song, section by section", while the Lyria API page says iterative editing through multiple prompts is not supported. The Sume docs list no song-edit endpoint; a substitute the docs support is to generate more than one take, cut each with timeline audio, and join the sections you keep.

Sume steps are from the Music Router docs and Timeline audio docs, read 2026-09-30.

What do Google's pages say?

The two pages describe different surfaces.

Editing on Google's pages and Sume's docs, read 2026-09-30.
SurfaceWhat the page says
Flow Music app"Precisely edit your song, section by section"; uses Lyria 3 Pro; iOS now, Android coming soon
Lyria API (Gemini)"iterative editing or refining a generated clip through multiple prompts is not supported"
Sume Music RouterOne prompt per job; lyria-3-pro is a routable id; no edit endpoint in the docs

How do I keep a good intro and swap the chorus?

Generate a second take with the same brief and a changed chorus line, keep both files, and build a new file from ranges. The docs require this workspace's media.sume.com audio and do not list accepted input formats or say whether a generated track can be used directly, so try one job (or import first with POST /v1/media-imports). Each parts[] item is { url, source_in?, duration? }, up to 20 parts, joined in the sample domain with no silence at the seams. Section timings are your own listening notes: result.lyrics may carry a model-reported section map, but the Music docs say it is not an audio measurement.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: song-splice-001" \
  -d '{
    "operation": "concat",
    "parts": [
      { "url": "https://media.sume.com/artifacts/artf_demo/take1.mp3", "source_in": 0, "duration": 16 },
      { "url": "https://media.sume.com/artifacts/artf_demo/take2.mp3", "source_in": 16, "duration": 24 }
    ]
  }'

What are the limits of this approach?

The join is sample-exact with no added silence, but the music itself changes at the seam, so it suits takes that share tempo and key, which you can ask for in the prompt. Nothing here re-sings a lyric or edits a stem. It is a splice, and a concat job is $0.01.

How do I submit the job and fetch the track?

A music request takes mode async, sync, subscribe or webhook. With sync or subscribe, wait_timeout_seconds is 0 to 30; with webhook, webhook_url must be a public HTTPS callback. Send an Idempotency-Key on the submit, then poll GET /v1/jobs/{job_id}/status and read GET /v1/jobs/{job_id}/result. The audio is the entry in result.artifacts[] where type is audio, hosted on media.sume.com; raw provider URLs are not public outputs. Run the second take as its own job with its own key, then splice.

The prompt is 1 to 5000 characters. Put exclusions in the positive prompt ("Instrumental, no vocals"), because a non-empty negative_prompt is refused. Full field list: Music Router docs.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume