How to split an audio file into parts with an API
Split one audio file into parts at the times you choose with Sume's Timeline audio split: up to 20 ranges, each returned as its own wav or mp3 file.

To split an audio file into parts, send POST /v1/timeline-1.0/audio with operation: "split", the file's url, and a ranges[] list of { start, end } seconds. Each range comes back as its own durable audio file, and you can send up to 20 ranges in one job.
The fields below are from the Sume docs page Timeline audio, read 2026-09-29.
What does the split request look like?
The url must be audio already in this workspace's media.sume.com space, such as the output of an earlier Sume job. An Idempotency-Key header is required. Leave end off a range to take the rest of the file.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: timeline-audio-split-001" \
-d '{
"operation": "split",
"url": "https://media.sume.com/artifacts/artf_demo/spine.wav",
"ranges": [{ "start": 0, "end": 12.4 }, { "start": 12.4 }]
}'What do I get back?
Poll GET /v1/jobs/:id/status until the job is terminal, then read GET /v1/jobs/:id/result. The result is kind: timeline_audio with segments[], and each segment has its own audio_url. There is no GET-by-id route for this tool, so the job envelope is the only way to read it.
The default output is wav, which is sample-exact. Set output.format to mp3 for smaller files; the docs say mp3 re-adds priming padding at every edge, so keep wav if the parts will be joined again.
How many ranges, and can they overlap?
A job takes 1 to 20 ranges, and ranges may overlap. Produced audio is limited to 1800 seconds. The refusals are stable codes you can branch on.
| Rule or code | Meaning |
|---|---|
ranges[] 1 to 20 | Each { start, end? }; end omitted means the rest of the file |
audio_split_requires_url / audio_split_requires_ranges | Split is missing its url or ranges |
audio_split_takes_no_parts | Split was sent with a concat field, parts |
audio_range_end_before_start | A range end is at or before its start |
unsupported_media_source / source_not_found | Off-host or dead URL |
Should I split by time or by silence?
This tool cuts at the times you send: a range is a start and an optional end, and the docs describe no other way to choose a boundary. Find the boundaries first, for example from the word timings in a transcript, and pass them as start and end. Because ranges may overlap, you can leave a short lead-in before each boundary without cutting the previous part. Each returned segment is a separate file, so you can send the parts to other tools one by one.
What if the audio is inside a video?
Detach it first. Audio detach turns one video's track into a wav or mp3, and the docs say to detach once and then split with Timeline audio when you need many ranges from one track.
Sources
Related posts
More in Media tools
- Spotify audio ad specs: length, file, and loudness rules
Spotify audio ads run up to 30 s as MP3, WAV or OGG: 44.1 kHz stereo, 192–320 kbps, -16 LUFS, up to 50 MB, plus a 600×600-minimum companion.
- Spotify error: audio and video streams are different durations
Spotify rejects a video episode whose audio and video tracks start and end at different times. Re-export them to match, then check the length with a probe.
- Spotify video podcast specs: MP4, 1080p 16:9, H.264
Spotify for Creators recommends a 16:9 MP4 at 1080p or higher, H.264 and AAC-LC, 24 to 60 FPS, one video and one audio track. How to render and check it.
- AI sound effects generator API: an open-weight SFX model, then Sume
Sume lists no sound-effects route. Stable Audio 3 Small SFX is open-weight, so you can run it yourself, import the files to Sume and mix them in Timeline 1.0.
Written by Sume