How to combine audio files without gaps using an API

Join audio files into one gapless file with Sume's Timeline audio concat: up to 20 parts, sample-domain, with segment offsets returned for re-basing.

4 min readSume
All posts

To combine audio files with no gap, send POST /v1/timeline-1.0/audio with operation: "concat" and an ordered parts[] list of 1 to 20 audio URLs. The join is done in the sample domain, so nothing is re-synthesized and no silence is added at the seams.

The fields below are from the Sume docs page Timeline audio, read 2026-09-29.

What does the concat request look like?

Each part is { url, source_in?, duration? }. Use source_in and duration to take only a slice of a part. Every URL must already be this workspace's media.sume.com audio, and an Idempotency-Key header is required. Do not send a top-level url or ranges.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: timeline-audio-concat-001" \
  -d '{
    "operation": "concat",
    "parts": [
      { "url": "https://media.sume.com/artifacts/artf_demo/line1.wav" },
      { "url": "https://media.sume.com/artifacts/artf_demo/line2.wav", "source_in": 0.1, "duration": 1.8 }
    ]
  }'

What comes back, and how do I read it?

The default mode is async, so you poll GET /v1/jobs/:id/status and then read GET /v1/jobs/:id/result. You can pass mode: "sync" to wait up to 30 seconds for a finished job. The result is kind: timeline_audio with one audio_url, duration_seconds, and segments[]; each segment has an index, a start and its duration_seconds.

Those segment offsets are what you re-base Timeline 1.0 video[].start against if the joined file becomes the audio spine of a render, described on the Timeline 1.0 page.

Which format should I choose for the joined file?

output.format is wav by default (pcm_s16le, sample-exact) or mp3. The docs say mp3 is smaller but re-adds priming padding at every edge, and that wav is the one to keep when the file will be joined again or drives lip-sync. Produced audio is limited to 1800 seconds.

Concat rules and refusals from the Timeline audio docs, read 2026-09-29.
Rule or codeMeaning
parts[] 1 to 20Ordered; each { url, source_in?, duration? }
audio_concat_requires_partsConcat sent without parts
audio_concat_takes_no_url / audio_concat_takes_no_rangesConcat sent with a split field
audio_parts_channel_mismatchParts do not share one channel layout (found by the worker)
unsupported_media_source / source_not_foundOff-host or dead URL

What can go wrong with the parts?

Two refusals are worth knowing. A part that is not this workspace's media.sume.com audio, such as the output of an earlier Sume job, is refused as unsupported_media_source or source_not_found. And parts with different channel layouts, such as one mono and one stereo, fail audio_parts_channel_mismatch when the worker reads them, so make the parts match before you submit. If a slice of a part is enough, source_in and duration trim it inside the join, which saves a separate trim step.

When should I skip this job?

The docs say a join needed only inside one render belongs on Timeline 1.0 audio.parts[] instead, so you skip this job. Use Timeline audio when you want a reusable file: for example one narration built from several spoken lines that you will feed to another tool.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume