Lecture slides beside the speaker: one Timeline compose shot per slide
Timeline compose stacks one still and one video in a frame, up to 300 seconds a shot. Make one stack per slide, then join them in Timeline 1.0.

A lecture video with the slide on one side and the speaker on the other is a stack of one still and one video, and Timeline compose builds exactly that: one still plus one video on screen at the same time, returned as one MP4. A compose is one shot, capped at 300 seconds, so a lecture needs one compose per slide, joined afterwards by Timeline 1.0.
Which layout puts the slide beside the speaker?
| Key | Value for a side-by-side lecture |
|---|---|
operation | stack |
layout.split | vertical (a left/right split) |
layout.image_region | left or right on a vertical split |
layout.ratio | The still's share of the frame, 0.1 to 0.9; the video takes the rest |
output | width 1920, height 1080 for a landscape lecture |
How do I make one shot per slide?
Export each slide as an image and upload it to your workspace. For each slide, send a compose with that slide as image.url and the same lecture video as video.url, with video.source_in at the moment the slide appears and video.duration for how long it stays up. The still is held for the whole shot and cannot lengthen it, because the output length always comes from the video layer. Each compose is a separate job.
curl -X POST https://api.sume.com/v1/timeline-1.0/compose \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: lecture-slide-03" \
-d '{
"operation": "stack",
"image": { "url": "https://media.sume.com/artifacts/artf_demo/slide-03.png" },
"video": {
"url": "https://media.sume.com/artifacts/artf_demo/lecture.mp4",
"source_in": 184,
"duration": 96
},
"layout": { "split": "vertical", "image_region": "left", "ratio": 0.65 },
"output": { "width": 1920, "height": 1080 }
}'How do I join the shots with the lecture audio?
Compose passes the video's audio through, but a Timeline 1.0 render takes its audio from the spine you declare. Detach the lecture audio once with Audio detach, which returns a sample-exact wav by default, and send it as audio.url. Then list each compose result in video[] with increasing start values so the slots together cover audio.duration_seconds. The first slot must start at 0, and a slot may trail the spine by at most 0.5 seconds.
A detach returns at most 900 seconds of audio, so a lecture longer than 15 minutes needs one detach per range. Pass those files as audio.parts (up to 20 gapless slices) instead of audio.url. The source video can be at most 1,800 seconds, and a Timeline holds up to 200 video[] slots.
What does a stacked lecture cost?
The docs list $0.02 flat per compose job and $0.01 per detach job, plus the Timeline render at $0.10 per output minute. A lecture with 20 slides is 20 compose jobs, one detach, and one render. Use POST /v1/timeline-1.0/plan to check the final document without billing.
Sources
Related posts
More in Use cases
- LinkedIn ad image size and AI-generated images in Campaign Manager
LinkedIn single image ads accept JPG, PNG or GIF up to 5 MB, and Campaign Manager has its own AI image tool. Which format to set on a Sume upload.
- LinkedIn carousel ad specs: 2 to 10 cards, images only
LinkedIn carousel ads take 2 to 10 image cards up to 10 MB each, 1080x1080 recommended, no video. How to plan the cards with Sume's four-images-per-job limit.
- LinkedIn in-stream video ads: up to 90 seconds, in testing
LinkedIn's video ad page says in-stream ads can run up to 90 seconds and are in testing. How to trim a clip to length and make a horizontal and vertical pair.
- LinkedIn 4:5 ad image: 720x900, mobile only, no side borders
LinkedIn's vertical single image ad is 4:5 at 720x900 and serves on mobile only; 1:1.91 vertical images get side borders. How to request 4:5 on Sume.
Written by Sume