Customer service training videos: role-plays made with AI

Customer service training videos as short AI role-plays: which scenarios to script, how to build a two-person scene, captions, and what each clip costs.

5 min readSume
All posts

A useful customer service training video is a short role-play: a customer says something difficult, the employee answers the way you want it handled, and a narrator names the skill being shown. You can make a library of these with AI presenters instead of actors, one scenario per video, each under a minute so staff can watch one before a shift.

The Sume facts below come from the Generate avatar video, Create new avatar, Audio detach, Timeline 1.0 and Video captions docs, read on 2026-09-29. Points marked as current behavior are read from Sume's code.

Which customer service scenarios should I script?

Pick the conversations your team gets wrong or finds hard, and write each as three to six short turns plus one line from the narrator. Examples:

  • An upset customer whose order is late: acknowledge, give a fact, give a next step.
  • A refund request outside your policy: say no without arguing, offer what you can do.
  • A question the employee can't answer: the hand-off to someone who can.
  • Retail: a return at the counter without a receipt.
  • Healthcare front desk: rescheduling a missed appointment. Use invented names and details, never real patient information.
  • The same scene twice, handled badly and then well, if your trainers teach by contrast.

How do I make a two-person role-play with AI?

Sume's current execution supports one resolved avatar per final video, so the customer and the employee never share a shot. Create one presenter per part once (customer, employee, and a narrator if you use one), each from a prompt, profile traits or a reference photo, and reuse the same cast across every scenario so staff recognize the roles. Render one talking clip per turn. Each clip needs a script Sume estimates at 4 to 60 seconds, so give every reply enough words to fill four seconds; a one-word answer on its own falls short.

Then cut the turns together in speaking order in one Timeline render. In current code the render drops each clip's own sound, so each turn's voice is detached first and laid on the render's audio spine. AI avatar conversation video: two speakers shows that join step by step.

When a policy changes, re-render only the turns whose words changed and run the join again; the presenters and the other turns stay as they are.

Should training videos have captions?

Caption them, so staff can watch on a shop floor or with the sound off. Send the finished scene to POST /v1/video-captions; script_text aligns the burned words to your script. In current code the caption job refuses a video longer than 60 seconds or one without an audio stream, which is one more reason to keep each scenario under a minute.

How much do customer service training videos cost?

Each step is billed per call from one prepaid balance. The presenters are a one-time cost; the clips are billed per second, so a scene's price follows its length.

From Generate avatar video, Create new avatar, Audio detach, Timeline 1.0, Video captions and the API pricing rate card, read 2026-09-29. Each rate is plus a 5.5% agent fee by default.
StepCallPrice
Presenters: customer, employee, narrator (once)POST /v1/avatar-1.0/generate$0.95 per avatar, each
One clip per turn, default qualityPOST /v1/avatar-1.0/talking-video$0.245 per second
Pull each turn's voicePOST /v1/audio-detach$0.01 per job
Cut the turns togetherPOST /v1/timeline-1.0/render$0.10 per output minute
Burned-in captionsPOST /v1/video-captions$0.20 per job, for videos up to 60 seconds

What are the limits?

  • No shared shot: each clip holds one presenter.
  • English only: in current code the avatar route speaks English. What languages can an AI avatar speak covers the route for other languages.
  • Up to 20 turns per joined scene when the voices go in as audio.parts[].
  • The render reads only this workspace's media.sume.com files, such as the clips and audio Sume returned; don't plan on uploading recordings of real calls.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume