Flow Music edits a song by section: how to do it by API
Google Flow Music edits songs by section; its Lyria API page says multi-prompt editing is unsupported. On Sume, cut two takes and join the parts you keep.

Google Flow Music advertises the ability to "precisely edit your song, section by section", while the Lyria API page says iterative editing through multiple prompts is not supported. The Sume docs list no song-edit endpoint; a substitute the docs support is to generate more than one take, cut each with timeline audio, and join the sections you keep.
Sume steps are from the Music Router docs and Timeline audio docs, read 2026-09-30.
What do Google's pages say?
The two pages describe different surfaces.
| Surface | What the page says |
|---|---|
| Flow Music app | "Precisely edit your song, section by section"; uses Lyria 3 Pro; iOS now, Android coming soon |
| Lyria API (Gemini) | "iterative editing or refining a generated clip through multiple prompts is not supported" |
| Sume Music Router | One prompt per job; lyria-3-pro is a routable id; no edit endpoint in the docs |
How do I keep a good intro and swap the chorus?
Generate a second take with the same brief and a changed chorus line, keep both files, and build a new file from ranges. The docs require this workspace's media.sume.com audio and do not list accepted input formats or say whether a generated track can be used directly, so try one job (or import first with POST /v1/media-imports). Each parts[] item is { url, source_in?, duration? }, up to 20 parts, joined in the sample domain with no silence at the seams. Section timings are your own listening notes: result.lyrics may carry a model-reported section map, but the Music docs say it is not an audio measurement.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: song-splice-001" \
-d '{
"operation": "concat",
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/take1.mp3", "source_in": 0, "duration": 16 },
{ "url": "https://media.sume.com/artifacts/artf_demo/take2.mp3", "source_in": 16, "duration": 24 }
]
}'What are the limits of this approach?
The join is sample-exact with no added silence, but the music itself changes at the seam, so it suits takes that share tempo and key, which you can ask for in the prompt. Nothing here re-sings a lyric or edits a stem. It is a splice, and a concat job is $0.01.
How do I submit the job and fetch the track?
A music request takes mode async, sync, subscribe or webhook. With sync or subscribe, wait_timeout_seconds is 0 to 30; with webhook, webhook_url must be a public HTTPS callback. Send an Idempotency-Key on the submit, then poll GET /v1/jobs/{job_id}/status and read GET /v1/jobs/{job_id}/result. The audio is the entry in result.artifacts[] where type is audio, hosted on media.sume.com; raw provider URLs are not public outputs. Run the second take as its own job with its own key, then splice.
The prompt is 1 to 5000 characters. Put exclusions in the positive prompt ("Instrumental, no vocals"), because a non-empty negative_prompt is refused. Full field list: Music Router docs.
Sources
Related posts
More in Use cases
- FLUX Deblur: 4 MP input vs Sume's 1-4x image upscale
FLUX Deblur sharpens a blurry image with no prompt or mask, up to 4 MP. Sume has no deblur endpoint; its image upscaler takes a factor of 1 to 4.
- FLUX Erase: mask rules and dilate_pixels vs a Sume mask edit
FLUX Erase removes a masked object with no prompt and a dilate_pixels setting. Sume has no erase endpoint; masked edits use mask_url with a prompt.
- FLUX Outpainting: 4 MP canvas cap vs Sume's image size options
FLUX Outpainting extends an image onto a canvas of up to 4 megapixels with offsets. Sume documents no outpainting mode; use image_size or aspect_ratio.
- GivingTuesday 2026 video ideas: a captioned 3-scene ask
GivingTuesday is December 1, 2026. Build a short nonprofit ask as three avatar scenes (hook, impact, CTA) with burned-in captions and a music bed.
Written by Sume