School promotional video with AI: show the real campus

A school promotional video should show the real campus and programs. Use AI for motion, voice and text on your own photos, not a generated campus. Steps, costs.

5 min readSume
All posts

A school promotional video should show the real school: your campus, your programs, and students and staff who agreed to appear, ending on the open-day date and the application link. AI's honest job is motion, voice and text: animate your own campus photos into short clips, add a narrator or an on-camera presenter, and cut a vertical and a wide version. A generated campus or generated students would show families a school that doesn't exist.

The Sume facts below come from the Video generation, Media inputs, Generate avatar video, Timeline 1.0 and Video captions docs and Sume's Terms of Service, read on 2026-09-29. Points marked as current behavior are read from Sume's code. The same steps fit a college, university or training academy.

What should a school promo video include?

Pick one audience per video, such as families applying next year, and answer what they ask first:

  • Who the school is for: age range, focus, and what makes a day there different, in one or two sentences.
  • What a day looks like: classrooms, labs, the library, sports fields, from your own photos.
  • Programs: one short shot and one line per program you want to feature.
  • A welcome from the head of school or an admissions officer.
  • The next step: open-day date, application deadline and link, shown on screen as exact text.
  • For one event only, such as an open day, event promo video covers the poster-based version.

Can I use AI-generated students or campus shots?

Not as if they were your school. Sume's terms say generated outputs should not be treated as a factual record of real people or events, and not to use Sume to mislead people about the authenticity of generated media. Families judge a school by its buildings and people, so both should be real.

Real students and staff need permission. The terms say you represent that you have permission to use any person's likeness or voice. For minors, get a parent's or guardian's consent under your school's own media policy before a photo goes into any tool.

How do I make a school promotional video from campus photos?

Only the first frame of each clip is your photo; what moves after it is generated, so check every clip for rooms, signs or people that aren't yours.

  • Put each campus photo at a public HTTPS URL and send it as the first_frame in frame_images on POST /v1/videos, with a prompt for one slow camera move. Most catalog models top out at 15 seconds per clip.
  • For a presenter, send a script and your avatar's avatar_handle to POST /v1/avatar-1.0/talking-video with a campus photo as the scene (type: "photo"). Sume accepts a script when it estimates the video at 4 to 60 seconds; 720p is the documented resolution. In current code the avatar speaks English only, so use text-to-speech narration with language for other languages.
  • Join the clips under the voice with a Timeline 1.0 render. It reads only your workspace's media.sume.com files, and in current code each clip's own audio is dropped, so a presenter's speech has to come in as the voice track: pull it out with POST /v1/audio-detach first.
  • Render twice for two shapes: output width and height are any even numbers from 256 to 2160, and the default is 1080×1920. video[].fit takes cover, contain, stretch or blur for clips made in the other shape.
  • Burn the dates and link as authored cues. In current code the caption job refuses videos over 60 seconds or without an audio stream.

How much does a school promotional video cost?

Each piece is billed per call from one prepaid balance; a second size adds one more render and one more caption job.

From Video generation, Generate avatar video, Timeline 1.0, Video captions and the API pricing rate card, read 2026-09-29. Each rate is plus a 5.5% agent fee by default.
PieceCallPrice
Campus photo into a moving clipPOST /v1/videosBy model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models
Presenter clip, default qualityPOST /v1/avatar-1.0/talking-video$0.245 per second
Narration instead of a presenterPOST /v1/tts-1.0/generate$0.0475 per 1,000 characters
Join clips and voice (once per size)POST /v1/timeline-1.0/render$0.10 per output minute
Dates and application link on screenPOST /v1/video-captions$0.20 per job, for videos up to 60 seconds

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume