AI wedding invitation video maker: art, motion, music, text

An AI wedding invitation video is your invitation art animated into short clips, set to music, with names, date and venue burned on as typed text.

5 min readSume
All posts

Yes, an AI tool can make a wedding invitation video, but in parts: animate your invitation art or a photo of the couple into short clips, set them to a music track, then burn the names, date, venue and RSVP line on as text you typed. Don't ask the video model to write the names: every frame after the first is generated, so lettering it draws can drift, while typed captions come out spelled exactly as you typed them.

With Sume, you can describe the invitation in the Agents tab, where the agent picks the models and asks before it spends, or make the same four calls over the API. Facts below come from the Video generation, Music Router, Timeline 1.0 and Video captions docs, read on 2026-09-29; limits marked as current behavior are read from Sume's code.

How do I make a wedding invitation video with AI?

The pipeline is the same four calls as an AI birthday video from photos: animate stills into clips, generate a track, join them in one render, then burn the text on. What changes for a wedding:

  • Start from a still you control: the invitation design, a save-the-date photo, or card art from an AI invitation generator with the text area left empty. It goes to POST /v1/videos as the first_frame in frame_images, at a public HTTPS URL, with a prompt for small motion such as “petals drift down, candle flames flicker, slow push-in”.
  • Brief the music for the occasion: “Warm string quartet, 70 BPM, a gentle swell at 0:20. A 30-second track. Instrumental, no vocals.” Length is steered in the prompt; the Music Router rejects duration.
  • Several events, one set of clips: a save-the-date, a mehndi or sangeet invitation and a reception invitation can reuse the same edit; each version is its own caption job with its own text.

How do I get the names and date on screen exactly right?

Burn them in as cues. Each cue is a text with a start and an end in seconds; cues skip speech-to-text and burn exactly that copy at those times. One card per detail keeps each line readable: names, then the date, then the venue, then the RSVP line.

Caption the finished edit, after the music is in. In current code the caption job refuses a video longer than 60 seconds or one with no audio stream, even when you send cues, so keep the invitation to a minute or less. Leave style out and Latin text gets slam, which in current code sets the words in capitals. How to add text over a video covers placement.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: invite-text-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/invite-edit.mp4",
    "cues": [
      { "text": "Ana & Sam are getting married", "start": 0, "end": 5 },
      { "text": "Saturday, June 12", "start": 5, "end": 9 },
      { "text": "RSVP by May 1", "start": 20, "end": 25 }
    ]
  }'

Can the invitation be in Hindi or another script?

Not as burned captions, as far as the docs go. In the caption docs, font picks a face only for Hangul styles, the face list is Hangul, and any other font name is rejected rather than substituted. No Devanagari face is documented. For an Indian wedding invitation in Hindi, put the text into the invitation art yourself and use that still as the first frame; the opening frame keeps it as designed, but generated frames after it may not, so hold the text frames short or add motion only around them.

How much does an AI wedding invitation video cost?

Four calls, four prices. A 30-second invitation with three clips uses three video jobs, one track, one render and one caption job. Each extra event version over the same edit adds one caption job.

From Video generation, Music Router, Timeline 1.0 and Video captions, read 2026-09-29. Each rate is plus a 5.5% agent fee by default; see API pricing.
StepCallPrice
Animate the artPOST /v1/videosBy model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models
MusicPOST /v1/music-router/generate$0.125 per audio
JoinPOST /v1/timeline-1.0/render$0.10 per output minute, reserved in whole minutes
Details as textPOST /v1/video-captions$0.20 per job, for videos up to 60 seconds

What should I check before sending it?

  • Faces: if you animate a photo of the couple, watch every clip. Only the first frame is your photo, so generate again if someone looks different.
  • Frame shape: pick one aspect_ratio, such as 9:16 for a phone message, for every clip. The render's default output is 1080×1920; set output.width and output.height for landscape.
  • Length per clip: it depends on the model, and the longest accept 30 seconds; check supported_durations on GET /v1/videos/models.
  • Inputs: photo URLs must be public HTTPS. The render takes only this workspace's media.sume.com files, such as the clips and track Sume returns, not files from your computer.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume