A 15-second 'previously on' recap from one episode with source_in
Detach an episode's audio once, then list three moments twice in one Timeline 1.0 render: video[].source_in for picture, audio.parts source_in for sound.

A recap that opens episode two with three moments from episode one does not need three trim jobs. Timeline 1.0 can point each video[] slot at a different source_in of the same file, and audio.parts[] can slice the matching sound out of one detached audio file with its own source_in and duration. Detach once, render once, and the result is a vertical MP4 by default.
Why do I need to list each moment twice?
Because a Timeline render takes its sound from the audio spine, not from the clips. The picture comes from video[] and the sound from audio.url or audio.parts, and nothing links them. Give a moment the same in-point in both places and the same length, and picture and sound stay in step. The parts join in the sample domain, with no gap, and there can be up to 20 of them.
How do I get the audio and the in-points?
Detach the episode's audio with Audio detach; it returns a sample-exact wav by default and caps output at 900 seconds, which covers an episode of a few minutes. To pick moments, run video inspect with a frames.at list to look at stills near candidate times; stills are unbilled. Then write down the start second of each moment.
What does the recap request look like?
Three moments of 5, 4 and 6 seconds total 15 seconds, so audio.duration_seconds is 15 and the slots start at 0, 5 and 9. The audio parts' lengths must add up to at least the declared duration or the render is refused with audio_parts_shorter_than_duration. Add a short transition on the slots after the first if you want a softer join; the docs cap it at 1 second.
| Moment | source_in (s) | Slot start (s) | Duration (s) |
|---|---|---|---|
| 1 | 12 | 0 | 5 |
| 2 | 41.5 | 5 | 4 |
| 3 | 78 | 9 | 6 |
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: recap-ep1-001" \
-d '{
"audio": {
"duration_seconds": 15,
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/ep1.wav", "source_in": 12, "duration": 5 },
{ "url": "https://media.sume.com/artifacts/artf_demo/ep1.wav", "source_in": 41.5, "duration": 4 },
{ "url": "https://media.sume.com/artifacts/artf_demo/ep1.wav", "source_in": 78, "duration": 6 }
]
},
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/ep1.mp4", "source_in": 12, "start": 0, "duration": 5 },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/ep1.mp4", "source_in": 41.5, "start": 5, "duration": 4 },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/ep1.mp4", "source_in": 78, "start": 9, "duration": 6 }
]
}'What are the limits to plan around?
Audio detach is billed per job at the docs' listed rate and the render per output minute, so a 15-second recap is one detach plus one rounded-up minute. Use POST /v1/timeline-1.0/plan first; it compiles the document, returns segment_count and an estimate, and does not download media or reserve credits. It cannot predict warnings for short sources that get padded or looped.
Sources
Related posts
More in Use cases
- Prime Big Deal Days beauty video ads: avatar clips, 4-60 s
Make beauty product ads for Prime Big Deal Days with Avatar Video: scripts must estimate to 4-60 seconds, and product_image is optional. Sizes and prices.
- Bulk product videos for Prime Big Deal Days: 100 Format runs
Queue up to 100 Format runs in one bulk request with a concurrency window of 1 to 16, then poll one queue id. Setup, limits and the traps to avoid.
- ProRes 4444 or PNG sequence from AI video: Sume returns MP4
Sume's video jobs return an MP4, not ProRes 4444 or a PNG sequence. Frame stills at times you name come from video frames as images. What to use instead.
- Qwen TTS inline tags vs Sume's emotion and speed controls
Qwen-Audio-3.0-TTS reads tags like [gasp] and [angry] in text. Sume TTS uses generation_config with an emotion guide and speed from 0.6 to 1.5.
Written by Sume