AI video generator for church: bumpers, announcements, loops

An AI video generator can make a church's sermon-series bumpers, weekly announcement clips and pre-service loops. How each is made, costs, limits.

6 min readSume
All posts

Yes, an AI video generator can make three short videos church media teams produce every week or every series: a title bumper for a sermon series, an announcement video, and a looping background for the screens before service. Each is short enough to generate in pieces, and the words on screen (series title, dates, times) go on afterwards as typed text rather than being drawn by the video model, so they are spelled exactly as you typed them.

With Sume, a volunteer can describe any of these in the Agents tab, where the agent picks the models and asks before it spends; the same work runs over the API. Facts come from the Video generation, Generate avatar video, Timeline 1.0 and Video captions docs, read on 2026-09-29; limits marked as current behavior are read from Sume's code.

How do I make a sermon series bumper?

  • Generate the scene: POST /v1/videos with a prompt (“morning light through tall windows, dust in the air, slow push-in”), or with your series artwork as the first_frame in frame_images at a public HTTPS URL so the bumper opens on it.
  • Add music from a brief to POST /v1/music-router/generate, ending with “Instrumental, no vocals.” Length is steered in the prompt.
  • Join the clip and track in a POST /v1/timeline-1.0/render (audio.mode: "silence", the track as soundtrack at gain_db: 0), then burn the series title as timed cues with POST /v1/video-captions. In current code the caption job refuses a video over 60 seconds or one with no audio stream, so add the music first and keep bumpers under a minute. Leave style out and Latin text gets slam, which in current code sets the title in capitals.

Can AI make our weekly announcement video?

Yes, in two ways. A presenter clip comes from POST /v1/avatar-1.0/talking-video: a ready avatar (avatar_handle) reads your script, which Sume accepts when it estimates the video at 4–60 seconds, so longer weeks split into several clips. 720p is the documented resolution, and aspect_ratio takes 16:9 for the sanctuary screen or 9:16 for phones. The avatar speaks English only in current code; for other languages, narrate instead. Talking avatar video API covers the request.

The other way is narrated slides: voice the week's notices with text to speech and hold one still per notice in a Timeline render, as in slideshow with voiceover. Both have speech, so both can take captions afterwards.

How do I make a looping background for the screens?

Send the same image as both the first_frame and the last_frame, so the clip ends where it starts; only models whose supported_frame_images lists last_frame accept that. Seamless loop AI video covers the check. To fill the minutes before service, repeat the clip as consecutive video[] slots in one Timeline render: up to 200 slots and 1,800 seconds (30 minutes). A loop that long can't carry burned captions, since the caption job takes 60 seconds at most in current code.

How much does it cost?

Each piece is billed per call, from one prepaid balance. A weekly set is a few video clips, one avatar clip or narration, one or two tracks, a render per piece, and a caption job per captioned piece.

From Video generation, Generate avatar video, Music Router, Timeline 1.0 and Video captions, read 2026-09-29. Each rate is plus a 5.5% agent fee by default; see API pricing.
PieceCallPrice
Scene or loop clipPOST /v1/videosBy model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models
Presenter clip (default quality, no product)POST /v1/avatar-1.0/talking-video$0.245 per second
MusicPOST /v1/music-router/generate$0.125 per audio
Join or loopPOST /v1/timeline-1.0/render$0.10 per output minute, reserved in whole minutes
Titles and datesPOST /v1/video-captions$0.20 per job, for videos up to 60 seconds

What should a church media team check?

  • Text: never let the video model write names, times or scripture references; put them in cues and read the render before Sunday.
  • Inputs: images go to POST /v1/videos as public HTTPS URLs; the Timeline render takes only this workspace's media.sume.com files, such as the clips Sume returned. In current code each clip's own sound is dropped, so sound comes from the render's audio spine and soundtrack.
  • People: if you animate photos of members, watch every clip; only the first frame is the photo.
  • One event rather than a weekly set? Event promo video with AI covers a single promo.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume