Combine more than 20 audio files: nest the concat jobs
Timeline audio concat takes 1 to 20 parts per job. For 45 files, concat in batches of 20 or fewer, then concat the batch outputs. Four jobs, $0.04 flat.

One timeline audio concat takes 1 to 20 parts[], so 45 files need two passes: three concats of 20, 20 and 5 parts, then one concat of those three outputs. That is four jobs at $0.01 flat per job, $0.04 in total, and each produced file must stay within 1800 seconds.
How many jobs does N files need?
The count follows the 20-part limit in the Timeline audio docs. The final pass must also fit in 20 parts, so past 400 files you need a third pass.
| Files | Jobs | Fee |
|---|---|---|
| 20 or fewer | 1 | $0.01 |
| 45 | 3 + 1 = 4 | $0.04 |
| 400 | 20 + 1 = 21 | $0.21 |
| 401 or more | A third pass is needed | Add $0.01 per job |
Is the join still gapless?
Each concat is a sample-domain join with no silence at the seams and no re-synthesis. Joining the joined files should keep that property when the batch outputs are wav (the default), though the docs describe the property per job, not across passes. The docs warn that mp3 output re-adds priming padding at every edge, so keep output.format wav until the last step.
How do I run the second pass?
Each result is kind: timeline_audio with one audio_url and segments[] giving index, start and duration_seconds. Feed the three audio_url values into a new concat. They are media.sume.com files, which is what the concat URL rule asks for, but the docs do not spell out that a concat output is accepted as a part, so run one small test first. Each job needs its own Idempotency-Key, and every URL must be this workspace's media.sume.com audio.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: concat-pass-2" \
-d '{
"operation": "concat",
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/batch1.wav" },
{ "url": "https://media.sume.com/artifacts/artf_demo/batch2.wav" },
{ "url": "https://media.sume.com/artifacts/artf_demo/batch3.wav" }
]
}'What if the total passes 1800 seconds?
Produced audio is capped at 1800 seconds, so a longer run needs separate final files.
How does the job run?
Timeline jobs default to mode: "async". Pass mode: "sync" to wait up to 30 seconds for a 200 finished job, or you get 202 and poll GET /v1/jobs/:id/status and GET /v1/jobs/:id/result; there is no separate GET for the audio or render job. Idempotency-Key is required on the render and audio jobs, and every URL must already be this workspace's media.sume.com audio or video. Wait for each pass to finish before you feed its audio_url into the next one.
Hosted MCP has timeline_audio and timeline_create; the flow is the tool, then jobs_wait, then the result tool. Details are in the Timeline audio docs and Timeline 1.0 docs.
Sources
Related posts
More in Media tools
- Extract audio from a video URL: why example.com is refused
Audio detach only reads a video already on this workspace's media.sume.com. An outside URL fails unsupported_media_source, so import it first, then detach.
- Lyria 3.5 outputs MP3 or WAV: what you get from Sume music
Google lists MP3 by default or WAV for Lyria 3.5. Sume's music request has no format field and returns an audio file; timeline audio can make a wav.
- Start a video's audio 45 seconds into a song: audio.source_in
Timeline 1.0's audio.source_in sets the in-point into a single audio spine. Output length stays duration_seconds, and it is illegal with parts or silence mode.
- Timeline audio segments: re-base video start times after concat
After a concat, use segments[] (index, start, duration_seconds) as the on-spine start of each video slot in Timeline 1.0. Declared starts are authoritative.
Written by Sume