How to make AI cat videos: one cat, many shots

To make AI cat videos, approve one still of your cat, start every shot from it, and join the shots under narration. Dancing and talking cats, costs.

5 min readSume
All posts

To make AI cat videos, pick one approved picture of your cat, start every shot from that picture, generate each shot as its own short clip, and join the clips in order under a narrator or music. Starting from the same picture is what keeps the cat recognizable from scene to scene; a text prompt alone describes the cat afresh each time, so it can come out different in every clip.

The steps below use Sume's API; in the Agents tab you can ask for the same video in plain words, and the agent asks before it spends. Facts come from the Video generation, Models overview and Timeline 1.0 docs and the Sume API reference, read on 2026-09-29.

How do I keep the same cat in every shot?

Use a photo of your own cat or a still you generated and approved, then send it into every shot in one of two ways. Consistent character across AI video shots explains both inputs in detail:

  • As the first frame: frame_images with frame_type: "first_frame". The picture sits at a public HTTPS URL, the clip opens on it, and the motion is generated from there. Use this when the shot starts on the cat.
  • As a reference: input_references. The docs call references visual guidance rather than exact frames, so the cat can drift. Use this when the shot starts elsewhere and the cat walks in.
  • Compare fur markings, eye color and collar across the finished clips yourself; nothing in the docs guarantees them, and a shot that drifts is regenerated.

How do I turn the shots into a cat story?

Write the story as a list of shots, one line each: “the cat wakes up on a windowsill”, “the cat chases a leaf across the kitchen”. Each line is one POST /v1/videos call; clip length depends on the model, up to 30 seconds on the longest, so check supported_durations on GET /v1/videos/models.

Then voice the story with text to speech (POST /v1/tts-1.0/generate; a transcript takes up to 20,000 characters) and join the clips in one Timeline 1.0 render with the narration as the audio spine; the render takes only this workspace's media.sume.com files, such as the clips and narration Sume returned. Make an AI video from a story covers the join. In current code each clip's own sound is dropped, so the narration and an optional music soundtrack are what viewers hear.

Can I make an AI cat dance?

Yes, through motion control instead of a prompt. POST /v1/kling/3.0/motion-control animates a still with the motion of a driving video you supply: send the cat still as image_url, a public HTTPS dance clip of up to 30 seconds as motion_video_url, and that clip's length as duration_seconds, which is required. The output length follows the driving video, and the driving video's own sound is kept unless you set keep_original_sound: false. The API reference doesn't say how a human dance maps onto a cat's body, so treat the first try as a test. Motion control API lists the other fields.

Can the cat talk?

The docs don't cover it. Sume's docs make talking shots from a still plus a voice through lip sync, and say video models do not lip-sync to generated speech or to a later voice-over. They describe presenters and speaking shots, not animal faces. For a cat story, use a narrator's voice over the shots instead of moving the cat's mouth.

How much does an AI cat video cost?

You pay per call. A one-minute story is a handful of clips, one narration, and one render.

From Video generation, Timeline 1.0 and the Sume API reference, read 2026-09-29. Each rate is plus a 5.5% agent fee by default; see API pricing.
StepCallPrice
Each shotPOST /v1/videosBy model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models
Dance shotPOST /v1/kling/3.0/motion-control$0.1575 per output second
NarrationPOST /v1/tts-1.0/generate$0.0475 per 1,000 characters
JoinPOST /v1/timeline-1.0/render$0.10 per output minute, reserved in whole minutes

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume