AI movie trailer generator: shots, narration, music, titles
An AI movie trailer is short generated shots, a narrator, music that builds and burned-in title cards, cut together. How to make one, with costs.

An AI movie trailer generator builds a trailer the way an editor cuts one: short generated shots from your script or concept, a narrator line or two, music that builds, and title cards, cut into one video. The film does not need to exist; you generate the shots the trailer shows. Plan it as a shot list, not a single prompt, because each video clip is one short shot.
Facts come from Sume's Video generation, Music 1.0, Music Router, Timeline 1.0 and Video captions docs and the Sume API reference, read on 2026-09-29. Limits marked as current behavior are read from Sume's code. In the Agents tab you can brief the whole trailer in plain words; the agent asks before it spends.
How do I make a movie trailer with AI?
- Write the beats: setup, the turn, rising stakes, the title. Give each beat one or two shots of a few seconds.
- Approve a still per shot, reusing the same character pictures so the cast stays recognizable, then animate each with
POST /v1/videos, the still as thefirst_frameinframe_images. Lengths differ by model;supported_durationsonGET /v1/videos/modelslists them. - Voice the narrator with
POST /v1/tts-1.0/generate.generation_configtakes aspeedfrom 0.6 to 1.5 and a free-textemotionguide. - Generate the score (next section).
- Cut with one
POST /v1/timeline-1.0/render: shots in order,fadeordissolvetransitions of up to 1 second, andoutput.fade_in_seconds/fade_out_secondsof up to 5 seconds for the open and close. - Burn the title, tagline and date on with caption
cues.
How do I make trailer music that builds?
Send a brief to POST /v1/music-router/generate. The Music docs suggest naming an arc with one named moment, such as “breakdown to bass and claps at 0:20, full return at 0:28”, so put your build and your hit on the seconds where the title lands. Close with “Instrumental, no vocals.” and add “no spoken word” under narration. The length is asked for in the prompt; duration is rejected. The timing is a creative direction, not a guaranteed setting, so listen before you cut to it.
In the render, the narration is the audio spine and the music is the soundtrack, which duck_db (0–20) dips while the narrator speaks; AI book trailer maker shows that mix. A trailer with no narrator sets audio.mode to "silence", and the soundtrack becomes the whole track. In current code each shot's own sound is dropped.
How do I put the title and release date on screen?
As timed caption cues on the finished cut, not drawn by the video model, since the title and date must be exact. Each cue burns exactly the text you send between its start and end; AI book trailer maker shows the cues request.
Two points matter for a film trailer. Left without a style, Latin text gets slam, which in current code sets words in capitals and lays a light dark overlay on the frame. And in current code the caption job refuses a video over 60 seconds or one without an audio stream, so keep the narration or music in the cut, and caption a longer trailer in parts, as in add captions to a long video.
How much does an AI movie trailer cost?
Each part is billed by what it produces; the shots are billed by model, and every take counts, kept or not.
| Part | Call | Price |
|---|---|---|
| Each shot | POST /v1/videos | By model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models |
| Narration | POST /v1/tts-1.0/generate | $0.0475 per 1,000 characters |
| Score | POST /v1/music-router/generate | $0.125 per audio |
| Cut | POST /v1/timeline-1.0/render | $0.10 per output minute, reserved in whole minutes |
| Title cards | POST /v1/video-captions | $0.20 per job, for videos up to 60 seconds |
What are the limits?
- In current code each shot's own sound is dropped from the cut, and video models do not lip-sync to a narrator; a character speaking on screen is a separate still-plus-voice shot, covered in How to make an AI movie.
- Stills for shots must be at public HTTPS URLs; the render takes only this workspace's
media.sume.comfiles, such as the clips Sume returned. - No video model accepts a
seed, so keep every take you approve. - The render's default frame is vertical, 1080×1920; set
output.widthandoutput.heightfor a widescreen trailer. - Build the cast from characters you made.
Sources
Related posts
More in Use cases
- AI product photos: cutout, new scene, upscale, and the cost per SKU
A three-call product photo pipeline on Sume: remove the background, generate a scene from the cutout, upscale the winner. What each call costs per SKU.
- AI product photo retouching: fix defects, keep the product
AI product photo retouching: send the real photo to an image-edit model, list the defects to fix and what must not change, then check every file.
- AI real estate photo editing: fixes, prompts, and checks
AI real estate photo editing: send each listing photo to an image-edit model with one named fix and a list of what must not change, then check it.
- AI recipe video generator: step photos, voice, quantities
An AI recipe video turns a photo of each step into a short clip, reads the method aloud, and burns quantities on as text. Steps, costs and limits.
Written by Sume