AI Christmas video generator: a greeting that moves

An AI Christmas video is a few animated clips from a photo or a winter scene, an original instrumental, and your greeting burned on as typed text.

5 min readSume
All posts

An AI Christmas video generator turns a family or team photo, or a text description of a winter scene, into a few seconds of motion; you then add music and a typed greeting to make a moving Christmas card. Make it as three pieces: short generated clips, an instrumental track made from a text brief, and the greeting burned on as text, so the words are spelled exactly as you typed them.

With Sume, you can ask for the whole card in the Agents tab, where the agent picks the models and asks before it spends, or make the calls over the API below. It is the same four steps as an AI birthday video from photos, so this page covers only what changes for Christmas. Facts come from the Video generation, Music Router, Timeline 1.0 and Video captions docs, read on 2026-09-29; limits marked as current behavior are read from Sume's code.

How do I make a Christmas video with AI?

Follow the four steps of the birthday post: animate each picture with POST /v1/videos, make a track, join the clips in one Timeline 1.0 render, and burn the greeting on with POST /v1/video-captions. What is specific to a Christmas card:

  • No photo is needed. Describe a winter scene in the prompt alone (“snow falls past a lit window, the tree lights twinkle, slow push-in”): the endpoint also generates from text. With a family or team photo, send it as the first_frame in frame_images, at a public HTTPS URL.
  • Write two greetings if you send to clients as well as family: “Merry Christmas from the Lee family” for one render, “Thank you for a great year. Happy holidays from Acme.” for another. Each is its own caption job over the same edit.
  • Leave style out and Latin text gets slam, which in current code sets the greeting in capitals.

Can AI make Christmas music for the video?

It makes a track from your description, not a recording of an existing carol. The Music docs describe prompt-driven generation: write the feel in your own words (“sleigh bells, warm strings, gentle glockenspiel, 90 BPM, a soft finish at 0:28. A 30-second track. Instrumental, no vocals.”) rather than naming a song. Length is steered in the prompt; duration is rejected, and a looped soundtrack covers a track that comes back short. Add background music to a video covers loop and fade.

How do I make a vertical and a landscape version?

Render twice. The render's default output is 1080×1920, a vertical frame for phone messages and stories; for email or a website, render the same clips again with output.width and output.height set to a landscape size. Each slot's fit decides how a clip fills the frame: cover is the default, and contain, stretch and blur are the other values. With cover, check that no face is cut off at the edges, or generate a second set of clips with aspect_ratio: "16:9" instead of "9:16". Caption each render separately, since each is its own video.

How much does an AI Christmas video cost?

A 20-second card in both shapes is three video jobs, one track, two renders and two caption jobs.

From Video generation, Music Router, Timeline 1.0 and Video captions, read 2026-09-29. Each rate is plus a 5.5% agent fee by default; see API pricing.
StepCallPrice
Animate each photo or scenePOST /v1/videosBy model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models
MusicPOST /v1/music-router/generate$0.125 per audio
Join, per shapePOST /v1/timeline-1.0/render$0.10 per output minute, reserved in whole minutes
Greeting, per shapePOST /v1/video-captions$0.20 per job, for videos up to 60 seconds

What are the limits?

  • In current code the caption job refuses a video longer than 60 seconds or one with no audio stream, even when you send cues, so caption the edit after the music is in and keep it to a minute or less.
  • In current code a Timeline render drops each clip's own sound, so the track is what people hear.
  • Faces in an animated family photo can change: only the first frame is your photo. Watch each clip and generate again if someone looks different.
  • Clip length and aspect ratios vary by model; check supported_durations and supported_aspect_ratios on GET /v1/videos/models.
  • The render takes only this workspace's media.sume.com files, such as the clips and track Sume returned; photo inputs to POST /v1/videos must be public HTTPS URLs.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume