Talking head video with graphics: layouts and overlays
Put a slide, product shot or logo on screen with a talking head: split the frame or lay the graphic over the presenter. How to do it with an API.

A talking head video with graphics shows a still, such as a slide, a product shot, a title card or a logo, on screen with the presenter. There are two layouts: split the frame so the graphic takes one part and the presenter the rest, or lay the graphic over the presenter. With Sume you render the presenter as an AI avatar clip, then combine it with one still per job in Timeline compose: stack splits the frame, overlay lays the still on top, and the presenter's speech passes through.
The facts come from Timeline compose, Timeline 1.0 and Generate avatar video, read on 2026-09-29. The endpoint itself is covered in Timeline compose and audio API; this post is about talking-head layouts.
Which talking head layouts can I build?
Compose takes operation (stack or overlay), image.url and video.url, plus an optional layout. The full key list is in Timeline compose and audio API; these are the settings behind each talking-head layout.
ratiois the still's share of the frame, 0.1–0.9; the video takes the rest.- An overlay still keeps its aspect ratio. Sizing a logo is covered in how to add a logo or watermark to a video.
| Layout | `operation` | `layout` keys |
|---|---|---|
| Graphic on top, presenter below (half-banner) | stack | The defaults: split: horizontal, image_region: top, ratio: 0.5 |
| Presenter on top, graphic below | stack | split: horizontal, image_region: bottom |
| Graphic beside the presenter | stack | split: vertical, image_region: left or right |
| Logo, title or lower band over the presenter | overlay | position top, center or bottom; size with width_ratio (default 0.9 of the width) |
How do I put a graphic on a talking head video?
First render the presenter: POST /v1/avatar-1.0/talking-video returns a public media.sume.com video when the job completes. Then send compose the still and that clip, with an Idempotency-Key. Set output to the frame you want; the default is 1080×1920. The result is one MP4 at video_url.
curl -X POST https://api.sume.com/v1/timeline-1.0/compose \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: slide-over-presenter-001" \
-d '{
"operation": "stack",
"image": { "url": "https://media.sume.com/artifacts/artf_demo/slide.png" },
"video": { "url": "https://media.sume.com/artifacts/artf_demo/talk.mp4" },
"layout": { "split": "horizontal", "image_region": "top", "ratio": 0.4 }
}'Where does the graphic have to come from?
Both URLs must already be this workspace's media.sume.com files. Off-host URLs such as https://example.com/slide.png are refused, image.url must be a still, and video.url must be a video. Sume's generated outputs are Sume-hosted under media.sume.com, so a still made on Sume qualifies. A slide or logo file that sits only on your computer or your own site doesn't. Put any text into the still itself: compose has no text field.
Can the graphic change during the video?
Not within one compose. It takes one still, holds it for the whole clip, and takes its length from the video. To show a different graphic per section:
- Cut the talking head into sections with Video trim.
- Compose each section with its own still.
- Join the sections on Timeline 1.0. Timeline builds its sound from one audio spine and, in current code, drops each clip's own sound, so give it the presenter's original voice as
audio.url, for example from Audio detach.
What can't it do, and what does it cost?
Compose puts a still and a video together, never two videos, so a screen recording or a second camera beside the presenter isn't possible here, and neither is a moving graphic. The ceiling is 300 seconds of output. A compose is $0.02 per job, and in current code it is reserved at that price and charged at the job's own compute cost, never above it. The talking head is billed per second: $0.184/s standard, $0.245/s plus, $0.55/s max (no product image). Each is plus a 5.5% agent fee by default. To put a clip on each slide of a deck instead, see AI avatars for PowerPoint presentations.
Sources
Related posts
More in Sume Avatar 1.0
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
- Avatar video previews: approve the first frame before rendering
Create an avatar video preview to get first-frame stills, regenerate them if needed, then call generate-video on the preview id to render the final video.
- How to create a reusable AI avatar with the Sume Avatar 1.0 API
Send POST /v1/avatar-1.0/generate with an avatar_handle and a prompt, profile, or image input. Poll the job, then reuse the handle for avatar videos.
Written by Sume