Hotel promotional video with AI, from your own photos
A hotel promotional video can be made from the photos you already have: animate each shot, add a voiced welcome and music, render wide and vertical cuts.

A hotel promotional video is a short film of the property, usually arrival, lobby, a room, and the amenities, with a spoken welcome and music, made for the website, booking pages and social feeds. You can make one without a shoot: turn each existing property photo into a slow camera-move clip with AI, put the clips in the order a guest meets them, and lay a voice-over and a music bed under them in one render, once as a 16:9 master and once as a 9:16 cut.
The Sume steps below come from the Video generation, Timeline 1.0 and Music Router docs and the Sume API reference, read on 2026-09-29. Limits marked as current behavior are read from Sume's code.
What should a hotel promotional video show?
Follow the guest's path and give each photo one small movement:
- Arrival: the entrance or facade, a slow push-in.
- Lobby: a gentle pan across the space, lights on.
- The room: a slow move from the door toward the bed and the window.
- Amenities: the pool, the spa, the restaurant, the breakfast spread.
- Close: the view or the neighborhood, then the hotel name and how to book, spoken in the voice-over.
- Keep it true. Only the first frame of each clip is your photo; every frame after it is generated, and a model can add a view, a balcony or a room size the property doesn't have. Watch every clip and cut any that shows something guests won't find.
How do I make a hotel promo video with AI?
It takes four kinds of step, each one API call. Restaurant video advertising with AI shows the same flow with a full clip request, so only the hotel choices are here:
- Animate each photo: send it at a public HTTPS URL as the
first_frameinframe_imagesonPOST /v1/videos, with one slow camera move in the prompt. Most catalog models top out at 15 seconds per clip; a few seconds per room is enough. Sendresolutionandaspect_ratioexplicitly. - Voice the welcome with Sume text-to-speech: the hotel's name, what a guest gets, how to book.
- Make a calm music bed with the Music Router; steer its length in the prompt, since
durationis rejected. - Join everything with one Timeline 1.0 render (
POST /v1/timeline-1.0/render) per cut: the voice as the audio spine, the clips in guest order asvideo[]slots, the track assoundtrackwithduck_dbso it dips under the voice. Every URL must already be amedia.sume.comfile in your workspace, such as the clips, voice and music Sume returned. - Between rooms, a slot after the first can carry a
transitionsuch asfadeordissolve, up to 1 second. - In current code the render keeps only the spine and the soundtrack, so any sound a clip generated is dropped.
How do I get a website version and a social cut?
Render the same program twice. The default output is 1080×1920, a vertical frame for Reels, TikTok and Shorts. For the website master, set output.width and output.height (even numbers from 256 to 2160), for example 1920 and 1080. Each slot's fit (cover, contain, stretch or blur) decides how a clip made for one shape fills the other. For a lobby screen that loops, see Seamless loop AI video.
People also search for a hotel promotional video template. With this setup, the render document is the template: keep the JSON, swap in new clip and voice URLs for each property or season.
How much does a hotel promo video cost?
You pay per step from your workspace USD balance. A 30-second promo is five or six clips, one voice-over, one music bed and one render per cut:
| Step | Call | Price |
|---|---|---|
| Animate each property photo | POST /v1/videos | By model, at provider list × 1.25 per clip; see pricing_skus on GET /v1/videos/models |
| Voiced welcome | POST /v1/tts-1.0/generate | $0.0475 per 1,000 characters |
| Music bed | POST /v1/music-router/generate | $0.125 per audio |
| Join into one MP4 (once per cut) | POST /v1/timeline-1.0/render | $0.10 per output minute |
What are the limits?
- Property photos must be at public HTTPS URLs; localhost, private-network and signed URLs are rejected.
- The render takes only this workspace's
media.sume.comfiles, 1 to 200 slots, and an output of 1 to 1,800 seconds. - Timeline is assembly only: sequencing, transitions and the audio spine. For on-screen room names or rates, burn captions onto the finished cut; in current code the caption job refuses a source over 60 seconds or one without an audio stream. Real estate photo to video AI shows the same flow for a listing.
Sources
Related posts
More in Use cases
- How to make a podcast trailer: length, script, AI voice
To make a podcast trailer, script an intro, highlights, a hook and a call to follow, keep it to two minutes or less, and mix the voice over a music bed.
- How to make AI cat videos: one cat, many shots
To make AI cat videos, approve one still of your cat, start every shot from it, and join the shots under narration. Dancing and talking cats, costs.
- How to make an AI movie: shot by shot, then the cut
You make an AI movie shot by shot: clips of up to 30 seconds, generated from a shot list, checked, voiced, then cut together. Workflow, limits, costs.
- Law firm video marketing with an AI avatar of the attorney
Law firm video marketing with AI: short attorney intro and practice-area clips from an avatar of the lawyer, what not to generate, how it's made and billed.
Written by Sume