How to make an AI movie: shot by shot, then the cut

You make an AI movie shot by shot: clips of up to 30 seconds, generated from a shot list, checked, voiced, then cut together. Workflow, limits, costs.

6 min readSume
All posts

You make an AI movie shot by shot, not in one generation: video models make clips measured in seconds, so a film is a script broken into many short shots that you generate, review and cut together, with voices and music added afterward. On Sume the longest clip is 30 seconds and one assembled render runs up to 30 minutes, so a feature-length film is several renders joined in your own editor.

Facts come from Sume's Video generation, Models, Timeline 1.0 and Video frames docs and the Sume API reference, read on 2026-09-29. Nothing here is a claim about how the film will look or where it can be shown.

Can AI make a full-length movie?

Not as one file from one prompt. seedance-2.5 and wan-3.0 accept up to 30 seconds per clip, and most other models stop at 15 seconds or less; AI video length limits lists them. A Timeline 1.0 render then joins 1 to 200 clips into one MP4 of up to 1800 seconds.

So a 90-minute feature needs at least 180 clips even if every clip ran the full 30 seconds, and at least 3 renders. Real films cut far more often, so plan on more shots, and render one scene or reel at a time.

What is the workflow for making an AI film?

The same steps a crew follows, with generation in place of the shoot. Can AI make a video from a story? covers the per-shot requests (approve a still, animate it as the first frame, voice it, join); at film scale the bookkeeping is the work:

  • Script, then a shot list: one line per shot with its length, framing, cast and the reel it belongs to. With hundreds of takes, this list is the only index you have.
  • Keep every take you approve. No video model accepts a seed, so you cannot regenerate the same clip later.
  • Review the dailies. POST /v1/video-frames pulls stills from a Sume-hosted clip at the times you name, unbilled, so you can compare faces and props across shots before you cut.
  • Cut one scene or reel per POST /v1/timeline-1.0/render, with up to 200 clips and 30 minutes each, then join the reels in your own editor.
  • Put all sound on each render's audio spine and optional soundtrack: in current code a clip's own audio is dropped. The spine's audio.parts[] takes at most 20 files, so a dialogue-heavy reel joins its lines into longer files first; POST /v1/timeline-1.0/audio concatenates up to 20 Sume-hosted audio files into one reusable file.

How do I keep characters consistent and make them talk?

Reuse approved pictures: settle each character in one still and build every shot's still from it; Consistent character across AI video shots compares the inputs, and none guarantees an identical look. Speaking shots are a separate lip-sync step, because video models do not lip-sync to generated speech or a later voice-over: a still plus a voice line through VEED Fabric 1.0, as Lip sync API shows. Give each character one voice and reuse it across the film.

How much does it cost to make an AI movie?

There is no flat price per film; each step is billed by what it produces. A full 30-minute render reserves $3.00. Every take you generate is billed, including the ones you don't use.

From Video generation, Timeline 1.0 and API pricing, read 2026-09-29. Each rate is plus a 5.5% agent fee by default.
StepCallPrice
Each shotPOST /v1/videosBy model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models
Dialogue voicePOST /v1/tts-1.0/generate$0.0475 per 1,000 characters
Talking shotPOST /v1/veed/fabric-1.0$0.1875 per audio second (720p)
ScorePOST /v1/music-router/generate$0.125 per audio
Each scene or reelPOST /v1/timeline-1.0/render$0.10 per output minute, reserved in whole minutes
Dailies stillsPOST /v1/video-framesUnbilled

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume