Lecture slides beside the speaker: one Timeline compose shot per slide

Timeline compose stacks one still and one video in a frame, up to 300 seconds a shot. Make one stack per slide, then join them in Timeline 1.0.

5 min readSume
All posts

A lecture video with the slide on one side and the speaker on the other is a stack of one still and one video, and Timeline compose builds exactly that: one still plus one video on screen at the same time, returned as one MP4. A compose is one shot, capped at 300 seconds, so a lecture needs one compose per slide, joined afterwards by Timeline 1.0.

Which layout puts the slide beside the speaker?

Stack layout keys from the Timeline compose docs, read 2026-10-01.
KeyValue for a side-by-side lecture
operationstack
layout.splitvertical (a left/right split)
layout.image_regionleft or right on a vertical split
layout.ratioThe still's share of the frame, 0.1 to 0.9; the video takes the rest
outputwidth 1920, height 1080 for a landscape lecture

How do I make one shot per slide?

Export each slide as an image and upload it to your workspace. For each slide, send a compose with that slide as image.url and the same lecture video as video.url, with video.source_in at the moment the slide appears and video.duration for how long it stays up. The still is held for the whole shot and cannot lengthen it, because the output length always comes from the video layer. Each compose is a separate job.

curl -X POST https://api.sume.com/v1/timeline-1.0/compose \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: lecture-slide-03" \
  -d '{
    "operation": "stack",
    "image": { "url": "https://media.sume.com/artifacts/artf_demo/slide-03.png" },
    "video": {
      "url": "https://media.sume.com/artifacts/artf_demo/lecture.mp4",
      "source_in": 184,
      "duration": 96
    },
    "layout": { "split": "vertical", "image_region": "left", "ratio": 0.65 },
    "output": { "width": 1920, "height": 1080 }
  }'

How do I join the shots with the lecture audio?

Compose passes the video's audio through, but a Timeline 1.0 render takes its audio from the spine you declare. Detach the lecture audio once with Audio detach, which returns a sample-exact wav by default, and send it as audio.url. Then list each compose result in video[] with increasing start values so the slots together cover audio.duration_seconds. The first slot must start at 0, and a slot may trail the spine by at most 0.5 seconds.

A detach returns at most 900 seconds of audio, so a lecture longer than 15 minutes needs one detach per range. Pass those files as audio.parts (up to 20 gapless slices) instead of audio.url. The source video can be at most 1,800 seconds, and a Timeline holds up to 200 video[] slots.

What does a stacked lecture cost?

The docs list $0.02 flat per compose job and $0.01 per detach job, plus the Timeline render at $0.10 per output minute. A lecture with 20 slides is 20 compose jobs, one detach, and one render. Use POST /v1/timeline-1.0/plan to check the final document without billing.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume