Gap or click when joining audio files: use wav, not mp3
Sume's Timeline audio joins up to 20 Sume-hosted parts with no silence at the seams. Its docs say to keep wav, because mp3 adds padding at every edge.

To join audio files without a gap, join them as wav. Sume's Timeline audio (POST /v1/timeline-1.0/audio, operation concat) joins Sume-hosted audio in the sample domain, with no re-synthesis and no silence at the seams. Its schema says wav (pcm_s16le, the default) stays sample-exact, while mp3 is smaller but re-introduces priming padding at every edge.
The facts are from the Timeline audio request schema in the Sume API reference and the Timeline audio docs, read 2026-09-29.
What does the join do?
parts holds 1 to 20 Sume-hosted HTTPS audio URLs, joined in order. Each part may carry source_in and duration to cut before it joins. All parts must share one channel layout, or the request fails with audio_parts_channel_mismatch. The result is a durable media.sume.com file.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: audio-join-001" \
-d '{
"operation": "concat",
"parts": [
{ "url": "https://media.sume.com/artifacts/example/line-1.wav" },
{ "url": "https://media.sume.com/artifacts/example/line-2.wav" }
],
"output": { "format": "wav" }
}'Which output format should I ask for?
output.format is optional and defaults to wav. The schema gives two values.
| `output.format` | What the schema says |
|---|---|
wav (default, pcm_s16le) | Stays sample-exact, so the result can be joined again or used to drive Avatar 1.0 image-to-video |
mp3 | Smaller, but re-introduces priming padding at every edge |
Why does mp3 add a gap?
Sume's schema states the effect and not the cause: mp3 output re-introduces priming padding at every edge, so if you hear a pause at a seam in an mp3 join, that is the documented behavior to look at first. The wav default avoids the padding, so the result can be joined again or used to drive Avatar 1.0 image-to-video without picking up encoder delay.
Pick mp3 only for the last file you hand to a listener, and keep the working files wav. If you join twice, for example a voiceover line and then a music tail, the schema says the padding returns at every edge of an mp3, so join the wav files first and ask for mp3 only on the final join. Listen to the seams of your own files after the join; the docs make no numeric promise about how long the padding is.
What are the limits?
- The join takes 1 to 20 parts per call, so a longer script needs a join of joins. Keep those intermediates wav too.
- Only Sume-hosted URLs are accepted, per the schema. A part from anywhere else does not match the request schema, so join audio that already lives on
media.sume.comin your workspace. Each part can still carry its ownsource_inandduration, so trimming happens in the same call. - If the joined audio is needed only inside one video, the docs say to use
audio.parts[]onPOST /v1/timeline-1.0/renderand skip this job. - For layering music under a voice, see Mix voice with background music.
Sources
Related posts
More in Models
- Gemini Live Avatar vs a video avatar API: live or rendered?
Gemini 3.8 Live Avatar streams a lip-synced face in a conversation. Sume's avatar video renders a finished clip from a script. Which job needs which.
- Gemini Live Avatar: 97 languages and SynthID, explained
Google says Gemini 3.8 Live Avatar covers 97 languages and adds SynthID to audio and video. What each claim covers, and what it says about other avatar tools.
- Gemini Omni 1.1 Flash 360p drafts: what a cheap preview costs
Gemini Omni Flash 1.1 can render 360p previews. On Sume a 360p second is billed at 30% of a 720p second, so drafts cost less than finals.
- Gemini Omni Flash 4K: clip length, upscaling and price
Gemini Omni Flash 1.1 renders 4K clips of 3 to 10 seconds on Sume, and Google labels its 4K output upscaled. Length versus Veo 3.1, plus cost.
Written by Sume