AI travel video generator: your trip photos or a prompt

An AI travel video generator either animates your own trip photos or invents a place from a prompt. How each works, what to label, and costs.

5 min readSume
All posts

An AI travel video generator does one of two jobs: it animates your own trip photos, a real place you shot, into moving clips, or it invents a destination from a text prompt. The first is travel footage; the second is mood B-roll that only looks like a place, and it should never be shown as the actual hotel, beach or view someone will get.

Below is how to make both with Sume's API and join them into one reel. Facts come from the Video generation, Timeline 1.0 and Music Router docs and the Sume API reference, read on 2026-09-29; limits marked as current behavior are read from Sume's code.

How do I turn my travel photos into a video?

  • Pick one photo per shot and host each at a public HTTPS URL.
  • Animate each: send it to POST /v1/videos as the first_frame in frame_images, with one small motion in the prompt ("waves roll in, the camera pushes slowly toward the cliff"). Set aspect_ratio to 9:16 for Reels, Shorts and TikTok, or 16:9 for YouTube, on a model that lists it in supported_aspect_ratios.
  • Write a voice-over and voice it with POST /v1/tts-1.0/generate; a transcript takes up to 20,000 characters.
  • Make a music bed with POST /v1/music-router/generate from a prompt; steer its length in the prompt, since duration is rejected.
  • Join with POST /v1/timeline-1.0/render: the voice-over as audio.url, the music as soundtrack, and one video[] slot per clip. The default output is 1080×1920, a vertical frame.

What happens to each clip's own sound?

In current code the render drops it. Sound comes only from the audio spine and the optional soundtrack, so set generate_audio to false on the clips if you are adding narration and music anyway. The soundtrack takes duck_db to lower the music under the voice.

Can AI make a travel video of a place I haven't been?

Yes, from a prompt alone, and the result is invented even when the prompt names a real town. The model draws a plausible street, beach or skyline, not a record of the one that exists. Use it for intros, transitions and mood, as in AI B-roll generator, and say it is AI-generated wherever a viewer could take it for the real place.

For a tour operator or hotel, that line matters most: a generated room or view is not the room or view the guest books. Animate your own photos of the real property instead; hotel promo video with AI covers the promo side. For an aerial move from one photo, see image to drone shot AI.

How much does an AI travel video cost?

From Video generation, Timeline 1.0, Music Router, the Sume API reference and API pricing, read 2026-09-29. Each price is plus a 5.5% agent fee by default.
StepCallPrice
Animate a photoPOST /v1/videosProvider list × 1.25, by model; $1.89 for 5 s of seedance-2 at 720p 9:16
Voice-overPOST /v1/tts-1.0/generate$0.0475 per 1,000 characters
Music bedPOST /v1/music-router/generate$0.125 per audio
JoinPOST /v1/timeline-1.0/render$0.10 per output minute, reserved in whole minutes

What are the limits?

  • Most catalog models make at most 15 seconds per clip; seedance-2.5 and wan-3.0 go to 30.
  • Photos for generation must be public HTTPS URLs. The render takes only this workspace's media.sume.com files, such as the clips, voice-over and music Sume returned.
  • A render is 1–1800 seconds long with 1–200 clip slots.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume