AI whiteboard animation generator: blank board to sketch
AI can imitate whiteboard animation: a blank-board first frame, a finished-sketch last frame, the drawing in between. Words go on as captions.

An AI whiteboard animation generator can imitate the hand-drawn explainer look: give a video model a blank-board still as the first frame and the finished sketch as the last frame, and prompt the drawing in between, one idea per clip. Generated lettering can come out wrong, so draw the pictures with AI and burn the words on as captions, with the narration setting the pace.
Facts come from Sume's Video generation, Image API, Timeline 1.0 and Video captions docs and the Sume API reference, read on 2026-09-29. Limits marked as current behavior are read from Sume's code. The strokes and any drawing hand are generated, so they may not move like real drawing; watch each clip.
How does AI make a whiteboard drawing animation?
With two stills per idea. The first is the board before the idea is drawn; the last is the board after. A video model fills the frames between them, and the prompt describes the drawing, such as “a marker sketches a lightbulb from left to right, still camera, plain white board”.
- Make the finished sketch from the blank board: send the board still as a reference in
input_referencesonPOST /v1/imagesand ask for the line drawing on it, so both frames share one board. References must be public HTTPS, and models whoseinput_referencesdescriptor is{"min": 0, "max": 0}reject them. The Image API returns the sketch at a signed Sume URL, so copy the file you keep and host it at your own public HTTPS URL before you use it as a frame. - For the next idea, the last sketch becomes the next clip's first frame, so the board fills up from shot to shot.
- Only models whose
supported_frame_imageslistslast_frametake an end frame, and in current code alast_framewithout afirst_frameis refused.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: board-bulb-001" \
-d '{
"model": "seedance-2",
"prompt": "A black marker sketches a lightbulb on a white board, still camera",
"frame_images": [
{ "type": "image_url", "image_url": { "url": "https://example.com/board-blank.png" }, "frame_type": "first_frame" },
{ "type": "image_url", "image_url": { "url": "https://example.com/board-bulb.png" }, "frame_type": "last_frame" }
],
"resolution": "720p",
"aspect_ratio": "16:9",
"duration": 6
}'Can the AI write the words on the whiteboard?
Don't count on it. A model may misspell a label or garble a number, and fixing one word means generating the clip again. Keep text out of the drawings and put it on with a caption job: each cue burns exactly the text you send between its start and end seconds. Left without a style, Latin text gets slam, which in current code sets words in capitals and lays a light dark overlay over the whole frame, dimming the white board; set a style from the Video captions page if that matters.
In current code the caption job refuses a video over 60 seconds or one without an audio stream, so caption the finished, narrated video, in parts under a minute if it runs longer. Add captions to a long video shows the split.
How do I turn the clips into a whiteboard video?
Write one sentence per idea and voice the script with POST /v1/tts-1.0/generate. With timestamps.words: true and segmentation.mode: "sentence", the result carries sentence segments[] with start and end times, so each drawing clip can start where its sentence starts. Slideshow with AI voiceover shows that mapping.
Then one POST /v1/timeline-1.0/render puts the narration on the audio spine and the clips in video[] at those starts. When a sentence outlasts its clip, render.pad_mode: "freeze" holds the last frame, the finished sketch, instead of replaying the drawing. In current code each clip's own sound is dropped, so viewers hear the narration and an optional soundtrack.
How much does an AI whiteboard animation cost?
A clip per idea, one narration, one render and one caption job per minute of video. Sketch stills from the Image API are priced per model, listed on GET /v1/images/models.
| Step | Call | Price |
|---|---|---|
| Each drawing clip | POST /v1/videos | By model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models |
| Narration | POST /v1/tts-1.0/generate | $0.0475 per 1,000 characters |
| Join | POST /v1/timeline-1.0/render | $0.10 per output minute, reserved in whole minutes |
| Words on screen | POST /v1/video-captions | $0.20 per job, for videos up to 60 seconds |
What are the limits?
- No video model accepts a
seed, so a redrawn clip comes out different; keep the takes you approve. - Stills for generation must be at public HTTPS URLs; the render takes only this workspace's
media.sume.comfiles, such as the clips Sume returned. - The render's default frame is 1080×1920, vertical; set
output.widthandoutput.heightto match 16:9 clips. - For the presenter or B-roll style of explainer instead, see How to make an explainer video with AI.
Sources
Related posts
More in Use cases
- Amazon Fire tablet video ad specs: 15 s max, 16:9 or 16:10, audio
Amazon's templated Fire tablet video ads allow 15 seconds, 1280x720 or larger, 16:9 or 16:10, .mp4 under 500 MB with AAC audio at 48 kHz. Sume fit.
- Amazon Fire TV feature rotator video specs: 1920x1080, -24 LKFS
Amazon lists the Fire TV feature rotator video at 5–21 seconds, 16:9, 1920x1080, no borders, and loudness of -24 LKFS. What Sume can and cannot set.
- Amazon Fire TV inline video ad specs: 10 to 30 s, 16:9, 960x540 min
Amazon Ads lists Fire TV inline video at 10–30 seconds, 16:9, 960x540 minimum, 500 MB, 2,000 kbps and audio at 128 kbps or higher. Sume can make the picture.
- Article URL to video AI: turn a blog post link into a video
An article URL to video tool makes a short video from a post. On Sume send the article text in input and its images as attachments; the link alone is not read.
Written by Sume