A 15-second 'previously on' recap from one episode with source_in

Detach an episode's audio once, then list three moments twice in one Timeline 1.0 render: video[].source_in for picture, audio.parts source_in for sound.

4 min readSume
All posts

A recap that opens episode two with three moments from episode one does not need three trim jobs. Timeline 1.0 can point each video[] slot at a different source_in of the same file, and audio.parts[] can slice the matching sound out of one detached audio file with its own source_in and duration. Detach once, render once, and the result is a vertical MP4 by default.

Why do I need to list each moment twice?

Because a Timeline render takes its sound from the audio spine, not from the clips. The picture comes from video[] and the sound from audio.url or audio.parts, and nothing links them. Give a moment the same in-point in both places and the same length, and picture and sound stay in step. The parts join in the sample domain, with no gap, and there can be up to 20 of them.

How do I get the audio and the in-points?

Detach the episode's audio with Audio detach; it returns a sample-exact wav by default and caps output at 900 seconds, which covers an episode of a few minutes. To pick moments, run video inspect with a frames.at list to look at stills near candidate times; stills are unbilled. Then write down the start second of each moment.

What does the recap request look like?

Three moments of 5, 4 and 6 seconds total 15 seconds, so audio.duration_seconds is 15 and the slots start at 0, 5 and 9. The audio parts' lengths must add up to at least the declared duration or the render is refused with audio_parts_shorter_than_duration. Add a short transition on the slots after the first if you want a softer join; the docs cap it at 1 second.

The three recap moments used in the example request, from the Timeline 1.0 docs, read 2026-10-01.
Momentsource_in (s)Slot start (s)Duration (s)
11205
241.554
37896
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: recap-ep1-001" \
  -d '{
    "audio": {
      "duration_seconds": 15,
      "parts": [
        { "url": "https://media.sume.com/artifacts/artf_demo/ep1.wav", "source_in": 12, "duration": 5 },
        { "url": "https://media.sume.com/artifacts/artf_demo/ep1.wav", "source_in": 41.5, "duration": 4 },
        { "url": "https://media.sume.com/artifacts/artf_demo/ep1.wav", "source_in": 78, "duration": 6 }
      ]
    },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_demo/ep1.mp4", "source_in": 12, "start": 0, "duration": 5 },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/ep1.mp4", "source_in": 41.5, "start": 5, "duration": 4 },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/ep1.mp4", "source_in": 78, "start": 9, "duration": 6 }
    ]
  }'

What are the limits to plan around?

Audio detach is billed per job at the docs' listed rate and the render per output minute, so a 15-second recap is one detach plus one rounded-up minute. Use POST /v1/timeline-1.0/plan first; it compiles the document, returns segment_count and an estimate, and does not download media or reserve credits. It cannot predict warnings for short sources that get padded or looped.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume