How to make AI cat videos: one cat, many shots
To make AI cat videos, approve one still of your cat, start every shot from it, and join the shots under narration. Dancing and talking cats, costs.

To make AI cat videos, pick one approved picture of your cat, start every shot from that picture, generate each shot as its own short clip, and join the clips in order under a narrator or music. Starting from the same picture is what keeps the cat recognizable from scene to scene; a text prompt alone describes the cat afresh each time, so it can come out different in every clip.
The steps below use Sume's API; in the Agents tab you can ask for the same video in plain words, and the agent asks before it spends. Facts come from the Video generation, Models overview and Timeline 1.0 docs and the Sume API reference, read on 2026-09-29.
How do I keep the same cat in every shot?
Use a photo of your own cat or a still you generated and approved, then send it into every shot in one of two ways. Consistent character across AI video shots explains both inputs in detail:
- As the first frame:
frame_imageswithframe_type: "first_frame". The picture sits at a public HTTPS URL, the clip opens on it, and the motion is generated from there. Use this when the shot starts on the cat. - As a reference:
input_references. The docs call references visual guidance rather than exact frames, so the cat can drift. Use this when the shot starts elsewhere and the cat walks in. - Compare fur markings, eye color and collar across the finished clips yourself; nothing in the docs guarantees them, and a shot that drifts is regenerated.
How do I turn the shots into a cat story?
Write the story as a list of shots, one line each: “the cat wakes up on a windowsill”, “the cat chases a leaf across the kitchen”. Each line is one POST /v1/videos call; clip length depends on the model, up to 30 seconds on the longest, so check supported_durations on GET /v1/videos/models.
Then voice the story with text to speech (POST /v1/tts-1.0/generate; a transcript takes up to 20,000 characters) and join the clips in one Timeline 1.0 render with the narration as the audio spine; the render takes only this workspace's media.sume.com files, such as the clips and narration Sume returned. Make an AI video from a story covers the join. In current code each clip's own sound is dropped, so the narration and an optional music soundtrack are what viewers hear.
Can I make an AI cat dance?
Yes, through motion control instead of a prompt. POST /v1/kling/3.0/motion-control animates a still with the motion of a driving video you supply: send the cat still as image_url, a public HTTPS dance clip of up to 30 seconds as motion_video_url, and that clip's length as duration_seconds, which is required. The output length follows the driving video, and the driving video's own sound is kept unless you set keep_original_sound: false. The API reference doesn't say how a human dance maps onto a cat's body, so treat the first try as a test. Motion control API lists the other fields.
Can the cat talk?
The docs don't cover it. Sume's docs make talking shots from a still plus a voice through lip sync, and say video models do not lip-sync to generated speech or to a later voice-over. They describe presenters and speaking shots, not animal faces. For a cat story, use a narrator's voice over the shots instead of moving the cat's mouth.
How much does an AI cat video cost?
You pay per call. A one-minute story is a handful of clips, one narration, and one render.
| Step | Call | Price |
|---|---|---|
| Each shot | POST /v1/videos | By model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models |
| Dance shot | POST /v1/kling/3.0/motion-control | $0.1575 per output second |
| Narration | POST /v1/tts-1.0/generate | $0.0475 per 1,000 characters |
| Join | POST /v1/timeline-1.0/render | $0.10 per output minute, reserved in whole minutes |
Sources
Related posts
More in Use cases
- How to make an AI movie: shot by shot, then the cut
You make an AI movie shot by shot: clips of up to 30 seconds, generated from a shot list, checked, voiced, then cut together. Workflow, limits, costs.
- Law firm video marketing with an AI avatar of the attorney
Law firm video marketing with AI: short attorney intro and practice-area clips from an avatar of the lawyer, what not to generate, how it's made and billed.
- LinkedIn content credentials: what the CR label means
LinkedIn shows a C2PA icon on images and videos signed with Content Credentials; clicking it shows whether AI was used, the tool, and who signed it.
- LinkedIn video ad specs: 16:9, 1:1, 4:5, 9:16 and the 30-second loop
LinkedIn's video ads page lists 3 s to 30 min, MP4, 75 KB to 500 MB, 4:5 at 720x900, 9:16 at 720x1280, and says videos under 30 s loop. How to hit them on Sume.
Written by Sume