Customer service training videos: role-plays made with AI
Customer service training videos as short AI role-plays: which scenarios to script, how to build a two-person scene, captions, and what each clip costs.

A useful customer service training video is a short role-play: a customer says something difficult, the employee answers the way you want it handled, and a narrator names the skill being shown. You can make a library of these with AI presenters instead of actors, one scenario per video, each under a minute so staff can watch one before a shift.
The Sume facts below come from the Generate avatar video, Create new avatar, Audio detach, Timeline 1.0 and Video captions docs, read on 2026-09-29. Points marked as current behavior are read from Sume's code.
Which customer service scenarios should I script?
Pick the conversations your team gets wrong or finds hard, and write each as three to six short turns plus one line from the narrator. Examples:
- An upset customer whose order is late: acknowledge, give a fact, give a next step.
- A refund request outside your policy: say no without arguing, offer what you can do.
- A question the employee can't answer: the hand-off to someone who can.
- Retail: a return at the counter without a receipt.
- Healthcare front desk: rescheduling a missed appointment. Use invented names and details, never real patient information.
- The same scene twice, handled badly and then well, if your trainers teach by contrast.
How do I make a two-person role-play with AI?
Sume's current execution supports one resolved avatar per final video, so the customer and the employee never share a shot. Create one presenter per part once (customer, employee, and a narrator if you use one), each from a prompt, profile traits or a reference photo, and reuse the same cast across every scenario so staff recognize the roles. Render one talking clip per turn. Each clip needs a script Sume estimates at 4 to 60 seconds, so give every reply enough words to fill four seconds; a one-word answer on its own falls short.
Then cut the turns together in speaking order in one Timeline render. In current code the render drops each clip's own sound, so each turn's voice is detached first and laid on the render's audio spine. AI avatar conversation video: two speakers shows that join step by step.
When a policy changes, re-render only the turns whose words changed and run the join again; the presenters and the other turns stay as they are.
Should training videos have captions?
Caption them, so staff can watch on a shop floor or with the sound off. Send the finished scene to POST /v1/video-captions; script_text aligns the burned words to your script. In current code the caption job refuses a video longer than 60 seconds or one without an audio stream, which is one more reason to keep each scenario under a minute.
How much do customer service training videos cost?
Each step is billed per call from one prepaid balance. The presenters are a one-time cost; the clips are billed per second, so a scene's price follows its length.
| Step | Call | Price |
|---|---|---|
| Presenters: customer, employee, narrator (once) | POST /v1/avatar-1.0/generate | $0.95 per avatar, each |
| One clip per turn, default quality | POST /v1/avatar-1.0/talking-video | $0.245 per second |
| Pull each turn's voice | POST /v1/audio-detach | $0.01 per job |
| Cut the turns together | POST /v1/timeline-1.0/render | $0.10 per output minute |
| Burned-in captions | POST /v1/video-captions | $0.20 per job, for videos up to 60 seconds |
What are the limits?
- No shared shot: each clip holds one presenter.
- English only: in current code the avatar route speaks English. What languages can an AI avatar speak covers the route for other languages.
- Up to 20 turns per joined scene when the voices go in as
audio.parts[]. - The render reads only this workspace's
media.sume.comfiles, such as the clips and audio Sume returned; don't plan on uploading recordings of real calls.
Sources
Related posts
More in Use cases
- Dental marketing videos with AI: what to make, what to skip
Dental marketing videos AI can make without a patient: a meet-the-dentist clip, first-visit and FAQ answers, and office-photo tours. What to skip, and costs.
- Donor thank you video: one render, each donor's name
A donor thank you video can show each donor's name for the price of one render plus a caption job per name, or say each name aloud at a full render per donor.
- How to extend an AI video past 30 seconds
Sume has no extend parameter on video models. Chain clips: pull a last frame, use it as the next first_frame, then join with Timeline. Vendor limits too.
- Facebook Reels ad specs: 9:16, 1440x2560, H.264 and safe zones
Meta's ads guide for Facebook Reels lists 9:16, 1440x2560, MP4/MOV, 4 GB, H.264 and safe zones of 14% top and 35% bottom. What Sume can and cannot hit.
Written by Sume