AI video generator for church: bumpers, announcements, loops
An AI video generator can make a church's sermon-series bumpers, weekly announcement clips and pre-service loops. How each is made, costs, limits.

Yes, an AI video generator can make three short videos church media teams produce every week or every series: a title bumper for a sermon series, an announcement video, and a looping background for the screens before service. Each is short enough to generate in pieces, and the words on screen (series title, dates, times) go on afterwards as typed text rather than being drawn by the video model, so they are spelled exactly as you typed them.
With Sume, a volunteer can describe any of these in the Agents tab, where the agent picks the models and asks before it spends; the same work runs over the API. Facts come from the Video generation, Generate avatar video, Timeline 1.0 and Video captions docs, read on 2026-09-29; limits marked as current behavior are read from Sume's code.
How do I make a sermon series bumper?
- Generate the scene:
POST /v1/videoswith a prompt (“morning light through tall windows, dust in the air, slow push-in”), or with your series artwork as thefirst_frameinframe_imagesat a public HTTPS URL so the bumper opens on it. - Add music from a brief to
POST /v1/music-router/generate, ending with “Instrumental, no vocals.” Length is steered in the prompt. - Join the clip and track in a
POST /v1/timeline-1.0/render(audio.mode: "silence", the track assoundtrackatgain_db: 0), then burn the series title as timedcueswithPOST /v1/video-captions. In current code the caption job refuses a video over 60 seconds or one with no audio stream, so add the music first and keep bumpers under a minute. Leavestyleout and Latin text getsslam, which in current code sets the title in capitals.
Can AI make our weekly announcement video?
Yes, in two ways. A presenter clip comes from POST /v1/avatar-1.0/talking-video: a ready avatar (avatar_handle) reads your script, which Sume accepts when it estimates the video at 4–60 seconds, so longer weeks split into several clips. 720p is the documented resolution, and aspect_ratio takes 16:9 for the sanctuary screen or 9:16 for phones. The avatar speaks English only in current code; for other languages, narrate instead. Talking avatar video API covers the request.
The other way is narrated slides: voice the week's notices with text to speech and hold one still per notice in a Timeline render, as in slideshow with voiceover. Both have speech, so both can take captions afterwards.
How do I make a looping background for the screens?
Send the same image as both the first_frame and the last_frame, so the clip ends where it starts; only models whose supported_frame_images lists last_frame accept that. Seamless loop AI video covers the check. To fill the minutes before service, repeat the clip as consecutive video[] slots in one Timeline render: up to 200 slots and 1,800 seconds (30 minutes). A loop that long can't carry burned captions, since the caption job takes 60 seconds at most in current code.
How much does it cost?
Each piece is billed per call, from one prepaid balance. A weekly set is a few video clips, one avatar clip or narration, one or two tracks, a render per piece, and a caption job per captioned piece.
| Piece | Call | Price |
|---|---|---|
| Scene or loop clip | POST /v1/videos | By model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models |
| Presenter clip (default quality, no product) | POST /v1/avatar-1.0/talking-video | $0.245 per second |
| Music | POST /v1/music-router/generate | $0.125 per audio |
| Join or loop | POST /v1/timeline-1.0/render | $0.10 per output minute, reserved in whole minutes |
| Titles and dates | POST /v1/video-captions | $0.20 per job, for videos up to 60 seconds |
What should a church media team check?
- Text: never let the video model write names, times or scripture references; put them in cues and read the render before Sunday.
- Inputs: images go to
POST /v1/videosas public HTTPS URLs; the Timeline render takes only this workspace'smedia.sume.comfiles, such as the clips Sume returned. In current code each clip's own sound is dropped, so sound comes from the render's audio spine andsoundtrack. - People: if you animate photos of members, watch every clip; only the first frame is the photo.
- One event rather than a weekly set? Event promo video with AI covers a single promo.
Sources
Related posts
More in Use cases
- AI virtual staging video: from empty room to furnished
An AI virtual staging video stages the empty-room photo first, then animates it or reveals the furniture, with the empty photo first and the staged one last.
- AI voice generator for games: voice every line as a file
An AI voice generator for games turns each dialogue line into an audio file in a character's voice. Voices, engine formats, batch runs and cost.
- AI voiceover for ads: one brand voice across every cut
Make AI voiceover for ads by pinning one voice and voicing each 6, 15 or 30 second cut as its own text to speech request. Fields, fit and cost.
- AI voiceover for short videos: 9:16 with voice and music
Make an AI voiceover for a Reel, Short or TikTok: script to speech, clips on the beat of the voice, optional music bed, one vertical MP4.
Written by Sume