Duck background music under a voiceover with the Timeline API
Set soundtrack.duck_db (0 to 20) on POST /v1/timeline-1.0/render so the music dips under your voiceover spine. It needs a real spine; silence mode is refused.

To duck music under a voiceover with Sume, render with Timeline 1.0 and add a soundtrack whose duck_db is between 0 and 20. The voiceover is the render's audio spine, the soundtrack is the music bed, and duck_db sets how far the bed dips while the spine speaks. It needs a real spine: duck_db with audio.mode: "silence" is refused as duck_requires_audio_spine.
Fields come from the Timeline 1.0 docs, read 2026-09-29. The docs do not say how the dip is shaped in time, so listen to a short render before you set a value for a whole batch.
What does the request look like?
Every URL must already be a media.sume.com file in your workspace (import first with POST /v1/media-imports). Idempotency-Key is required.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: duck-001" \
-d '{
"audio": {
"url": "https://media.sume.com/artifacts/artf_demo/voice.wav",
"duration_seconds": 24
},
"soundtrack": {
"url": "https://media.sume.com/artifacts/artf_demo/bed.mp3",
"loop": true,
"duck_db": 12,
"fade_out_seconds": 2
},
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4", "start": 0, "duration": 24 }
]
}'Which soundtrack fields exist?
| Field | Meaning |
|---|---|
url | The music bed, a Sume-hosted file |
gain_db | Level of the bed |
loop | Repeat the bed to fill the render |
fade_out_seconds | At most 10 |
duck_db | 0 to 20; needs a real spine, not silence |
What does it cost?
The render is $0.10 per output minute, reserved as ceil(audio.duration_seconds / 60) minutes; a 24 second render reserves one. The docs note no provider inference is involved, only worker ffmpeg. You can call POST /v1/timeline-1.0/plan first: it is unbilled and returns billable_minutes and estimated_cost_usd_micros without creating a job.
What else can trip it?
soundtrack_fade_exceeds_output: the fade is longer than the spine.video[0].startmust be 0, or the render fails withtimeline_must_start_at_zero.- Off-host URLs are rejected as
unsupported_media_source.
How do I make the voiceover spine?
If your voiceover is several audio files, list them in audio.parts[] (at most 20, joined gaplessly in the sample domain with no re-synthesis) when they are only needed inside this render. If you need one reusable file, POST /v1/timeline-1.0/audio with operation: "concat" joins up to 20 parts for $0.01 per job and returns the offsets to line your video[].start values up against.
Sources
Related posts
More in Developers
- Render a silent video from clips with the Timeline API
Set audio.mode to silence and a duration_seconds on POST /v1/timeline-1.0/render to join clips with no audio file. Which fields are illegal there, and pricing.
- Timeline error too_many_chained_transitions: how to fix it
Timeline 1.0 refuses more than 8 adjacent fades with too_many_chained_transitions. Insert a hard cut. Also transition_too_long, transition_not_frame_aligned.
- Render a vertical 1080x1920 video from clips with an API
Timeline 1.0 defaults to a 1080x1920 MP4. Set output width, height and fps, and pick fit cover, contain, stretch or blur for clips that do not match the frame.
- Let the API pick the video model: sume/auto for vertical UGC clips
Send model sume/auto to POST /v1/videos and Sume picks the family. The response echoes sume/auto and never names the model. When to pin a model instead.
Written by Sume