AI movie trailer generator: shots, narration, music, titles

An AI movie trailer is short generated shots, a narrator, music that builds and burned-in title cards, cut together. How to make one, with costs.

5 min readSume
All posts

An AI movie trailer generator builds a trailer the way an editor cuts one: short generated shots from your script or concept, a narrator line or two, music that builds, and title cards, cut into one video. The film does not need to exist; you generate the shots the trailer shows. Plan it as a shot list, not a single prompt, because each video clip is one short shot.

Facts come from Sume's Video generation, Music 1.0, Music Router, Timeline 1.0 and Video captions docs and the Sume API reference, read on 2026-09-29. Limits marked as current behavior are read from Sume's code. In the Agents tab you can brief the whole trailer in plain words; the agent asks before it spends.

How do I make a movie trailer with AI?

  • Write the beats: setup, the turn, rising stakes, the title. Give each beat one or two shots of a few seconds.
  • Approve a still per shot, reusing the same character pictures so the cast stays recognizable, then animate each with POST /v1/videos, the still as the first_frame in frame_images. Lengths differ by model; supported_durations on GET /v1/videos/models lists them.
  • Voice the narrator with POST /v1/tts-1.0/generate. generation_config takes a speed from 0.6 to 1.5 and a free-text emotion guide.
  • Generate the score (next section).
  • Cut with one POST /v1/timeline-1.0/render: shots in order, fade or dissolve transitions of up to 1 second, and output.fade_in_seconds / fade_out_seconds of up to 5 seconds for the open and close.
  • Burn the title, tagline and date on with caption cues.

How do I make trailer music that builds?

Send a brief to POST /v1/music-router/generate. The Music docs suggest naming an arc with one named moment, such as “breakdown to bass and claps at 0:20, full return at 0:28”, so put your build and your hit on the seconds where the title lands. Close with “Instrumental, no vocals.” and add “no spoken word” under narration. The length is asked for in the prompt; duration is rejected. The timing is a creative direction, not a guaranteed setting, so listen before you cut to it.

In the render, the narration is the audio spine and the music is the soundtrack, which duck_db (0–20) dips while the narrator speaks; AI book trailer maker shows that mix. A trailer with no narrator sets audio.mode to "silence", and the soundtrack becomes the whole track. In current code each shot's own sound is dropped.

How do I put the title and release date on screen?

As timed caption cues on the finished cut, not drawn by the video model, since the title and date must be exact. Each cue burns exactly the text you send between its start and end; AI book trailer maker shows the cues request.

Two points matter for a film trailer. Left without a style, Latin text gets slam, which in current code sets words in capitals and lays a light dark overlay on the frame. And in current code the caption job refuses a video over 60 seconds or one without an audio stream, so keep the narration or music in the cut, and caption a longer trailer in parts, as in add captions to a long video.

How much does an AI movie trailer cost?

Each part is billed by what it produces; the shots are billed by model, and every take counts, kept or not.

From Video generation, Timeline 1.0, Video captions and API pricing, read 2026-09-29. Each rate is plus a 5.5% agent fee by default.
PartCallPrice
Each shotPOST /v1/videosBy model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models
NarrationPOST /v1/tts-1.0/generate$0.0475 per 1,000 characters
ScorePOST /v1/music-router/generate$0.125 per audio
CutPOST /v1/timeline-1.0/render$0.10 per output minute, reserved in whole minutes
Title cardsPOST /v1/video-captions$0.20 per job, for videos up to 60 seconds

What are the limits?

  • In current code each shot's own sound is dropped from the cut, and video models do not lip-sync to a narrator; a character speaking on screen is a separate still-plus-voice shot, covered in How to make an AI movie.
  • Stills for shots must be at public HTTPS URLs; the render takes only this workspace's media.sume.com files, such as the clips Sume returned.
  • No video model accepts a seed, so keep every take you approve.
  • The render's default frame is vertical, 1080×1920; set output.width and output.height for a widescreen trailer.
  • Build the cast from characters you made.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume