How to make an AI movie: shot by shot, then the cut
You make an AI movie shot by shot: clips of up to 30 seconds, generated from a shot list, checked, voiced, then cut together. Workflow, limits, costs.

You make an AI movie shot by shot, not in one generation: video models make clips measured in seconds, so a film is a script broken into many short shots that you generate, review and cut together, with voices and music added afterward. On Sume the longest clip is 30 seconds and one assembled render runs up to 30 minutes, so a feature-length film is several renders joined in your own editor.
Facts come from Sume's Video generation, Models, Timeline 1.0 and Video frames docs and the Sume API reference, read on 2026-09-29. Nothing here is a claim about how the film will look or where it can be shown.
Can AI make a full-length movie?
Not as one file from one prompt. seedance-2.5 and wan-3.0 accept up to 30 seconds per clip, and most other models stop at 15 seconds or less; AI video length limits lists them. A Timeline 1.0 render then joins 1 to 200 clips into one MP4 of up to 1800 seconds.
So a 90-minute feature needs at least 180 clips even if every clip ran the full 30 seconds, and at least 3 renders. Real films cut far more often, so plan on more shots, and render one scene or reel at a time.
What is the workflow for making an AI film?
The same steps a crew follows, with generation in place of the shoot. Can AI make a video from a story? covers the per-shot requests (approve a still, animate it as the first frame, voice it, join); at film scale the bookkeeping is the work:
- Script, then a shot list: one line per shot with its length, framing, cast and the reel it belongs to. With hundreds of takes, this list is the only index you have.
- Keep every take you approve. No video model accepts a
seed, so you cannot regenerate the same clip later. - Review the dailies.
POST /v1/video-framespulls stills from a Sume-hosted clip at the times you name, unbilled, so you can compare faces and props across shots before you cut. - Cut one scene or reel per
POST /v1/timeline-1.0/render, with up to 200 clips and 30 minutes each, then join the reels in your own editor. - Put all sound on each render's audio spine and optional
soundtrack: in current code a clip's own audio is dropped. The spine'saudio.parts[]takes at most 20 files, so a dialogue-heavy reel joins its lines into longer files first;POST /v1/timeline-1.0/audioconcatenates up to 20 Sume-hosted audio files into one reusable file.
How do I keep characters consistent and make them talk?
Reuse approved pictures: settle each character in one still and build every shot's still from it; Consistent character across AI video shots compares the inputs, and none guarantees an identical look. Speaking shots are a separate lip-sync step, because video models do not lip-sync to generated speech or a later voice-over: a still plus a voice line through VEED Fabric 1.0, as Lip sync API shows. Give each character one voice and reuse it across the film.
How much does it cost to make an AI movie?
There is no flat price per film; each step is billed by what it produces. A full 30-minute render reserves $3.00. Every take you generate is billed, including the ones you don't use.
| Step | Call | Price |
|---|---|---|
| Each shot | POST /v1/videos | By model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models |
| Dialogue voice | POST /v1/tts-1.0/generate | $0.0475 per 1,000 characters |
| Talking shot | POST /v1/veed/fabric-1.0 | $0.1875 per audio second (720p) |
| Score | POST /v1/music-router/generate | $0.125 per audio |
| Each scene or reel | POST /v1/timeline-1.0/render | $0.10 per output minute, reserved in whole minutes |
| Dailies stills | POST /v1/video-frames | Unbilled |
Sources
Related posts
More in Use cases
- Law firm video marketing with an AI avatar of the attorney
Law firm video marketing with AI: short attorney intro and practice-area clips from an avatar of the lawyer, what not to generate, how it's made and billed.
- LinkedIn content credentials: what the CR label means
LinkedIn shows a C2PA icon on images and videos signed with Content Credentials; clicking it shows whether AI was used, the tool, and who signed it.
- LinkedIn video ad specs: 16:9, 1:1, 4:5, 9:16 and the 30-second loop
LinkedIn's video ads page lists 3 s to 30 min, MP4, 75 KB to 500 MB, 4:5 at 720x900, 9:16 at 720x1280, and says videos under 30 s loop. How to hit them on Sume.
- Nano Banana character consistency: reference limits and a Sume request
Google documents character reference slots for its Nano Banana models. The limits, and how to keep a character across images with input_references on Sume.
Written by Sume