AI Santa video: a personal message with the child's name

Make an AI Santa video in three steps: a Santa character still, the message as speech, then lip sync. One short job per child, billed per second.

5 min readSume
All posts

An AI Santa video is a short clip of a Santa character saying a child's name and a personal message. You can make one yourself in three steps: create a still image of a Santa character, turn the message into speech, and lip sync the still to that speech. The character is made once; each child then needs its own speech and lip sync jobs, so ten messages are ten runs of the last two steps.

Facts come from Sume's Models overview, the TTS and VEED Fabric 1.0 schemas in the Sume API reference, and Sume's Terms, read on 2026-09-29. Prices come from the code behind API pricing. The general mechanism, and why a video model with a voice-over laid underneath won't move the lips to your words, is in How to make a photo talk with AI; this page covers what changes for Santa.

How do I make an AI video of Santa?

Make the character and the voice once, then run two jobs per child:

  • The still: a Santa character you made or have the rights to, at a public HTTPS URL. You can generate one with POST /v1/images (Image API); its result URLs are Sume-hosted and signed, so copy the file you choose to a public HTTPS location you control before the per-child runs. Don't use a photo of a real Santa performer or branded Santa artwork.
  • The voice: TTS 1.0 speaks with a voice your workspace already has, sent as voice.id (a Voices library id, voi_ plus 32 hex characters) or as a ready avatar's avatar_handle. In the Sume app today, the Voices screen lets you clone a voice or invent one from a prompt; there is no API call for that step.
  • Per child: POST /v1/tts-1.0/generate with the message as transcript, then POST /v1/veed/fabric-1.0 with the still as image_url, the speech as audio_url, and its length as duration_seconds.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: santa-emma-001" \
  -d '{
    "transcript": "Ho ho ho! Hello Emma! The elves tell me you have been so kind to your little brother this year.",
    "voice": { "id": "voi_…" },
    "language": "en"
  }'

Can Santa speak another language?

Yes, on this path. Set language on the TTS request for every non-English message; if you leave it out, the speech defaults to English. Fabric only lip-syncs the audio it receives, so it doesn't need a language setting. If the voice and the language don't match, TTS can return a mismatch warning, and you resend with confirm_language_mismatch: true only after you've listened to it.

How much does an AI Santa video cost?

Each child costs one TTS job and one Fabric job, each plus a 5.5% agent fee by default. Fabric counts audio seconds, rounded up, so a 30-second message at 720p comes to $5.625 before the fee. The still is generated once and priced by the image model you choose.

From the TTS and Fabric schemas in the Sume API reference and the code behind API pricing, read 2026-09-29.
StepLimitPrice
Message as speech: TTS 1.0Up to 20,000 characters; spaces and punctuation count$0.0475 per 1,000 characters
Lip sync: VEED Fabric 1.01–300 seconds of audio on the Sume media host, at most 10 MB; 720p by default, or 480p$0.1875 per audio second (720p); the 480p rate is on API pricing

What should I watch out for?

  • Don't use the child's photo: the child is the audience, not the subject. Sume's terms also forbid content that exploits or endangers minors.
  • Use only a face and a voice you have the rights to. The terms say you represent that you have permission to use any person's likeness or voice you submit.
  • Watch every clip before you send it. The terms call generated outputs synthetic and make you responsible for reviewing them before publication.
  • For a class or a shop's customer list, the jobs can run in a batch, but how many run at once is set by your plan's concurrency limit, not by how fast you send requests (Authentication).
  • The audio must be on the Sume media host; a recording on your own server is refused, and a TTS result qualifies.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume