Can AI make a lyric video? Yes, if you time the lines

AI can make a lyric video: pictures under the song, plus each lyric line burned in at the time it is sung. You supply the line timings.

5 min readSume
All posts

Yes, AI can make a lyric video. A lyric video is two layers: a picture track that runs under the song, and the words of each line burned into the picture at the moment they are sung. AI tools can generate the song and the pictures and draw the text, but the line timings are the part you supply.

On Sume that is a Timeline render with the song as its sound, then a captions job that burns your lines as timed cues. The facts below come from the Video captions, Timeline 1.0 and Music Router docs and Sume's code, read on 2026-09-29.

How do I make a lyric video with AI?

Work in this order, so the words go on last:

  • Get the song as a Sume file. A track from POST /v1/music-router/generate comes back as an audio artifact; completed jobs return media.sume.com URLs.
  • Build the picture track with POST /v1/timeline-1.0/render: the song as audio.url, its length as audio.duration_seconds, and clips or stills from your workspace as video[] slots. Can AI make a music video for my song? covers choosing a clip per section.
  • Write the lyrics as cues: one line per cue, with start and end in seconds.
  • Send the rendered video's URL and the cues to POST /v1/video-captions. Cues skip speech-to-text and burn exactly that copy at those times.
curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: lyric-video-chorus-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/chorus.mp4",
    "cues": [
      { "text": "We run until the lights go down", "start": 0.4, "end": 3.2 },
      { "text": "and sing it back again", "start": 3.4, "end": 6.0 }
    ]
  }'

Can AI time the lyrics to the song for me?

Sume documents no way to time sung lyrics automatically, so plan to time them yourself. Without cues, the captions job runs speech-to-text, which needs audible speech; the docs don't cover sung vocals over music. With cues, each line appears at the start and end you give, and a job takes up to 200 cues.

A generated song can help with the words: the Music Router result carries result.lyrics, the model-reported lyrics or section map, when present. The docs don't say it includes timings, so listen through and note each line's start and end.

If you omit style, Latin text gets slam, which in current code uppercases every word. Name a style if you want another look.

Can I use my own song?

Only in part. Timeline reads only files already in your Sume workspace on media.sume.com, and there is no public route to upload a local file, so an MP3 on your computer can't be sent straight to a Timeline render; the render's sound has to be a Sume-hosted track, such as one from the Music Router. The captions job is different: it takes any public HTTPS video URL. A short video that already carries your song, hosted at a public HTTPS address, can have lyrics burned in directly.

How long can a lyric video be?

The picture track can run up to 1,800 seconds. The captions job is the limit: in current code it refuses a source longer than 60 seconds, and a source with no audio stream. A song always has audio, so the 60 seconds is what matters.

For a full song, cut the render into pieces of 60 seconds or less, caption each piece with its own cue times starting from 0, and join the pieces in a second render. Add captions to a long video by API shows the split and rejoin. Put the song back as the audio of that final render: in current code a Timeline render takes sound only from its audio track and optional soundtrack, and each clip's own sound is dropped.

How much does an AI lyric video cost?

Each step is billed on its own, plus a 5.5% agent fee by default. The render reserves by output minute and captures its own compute, never above the reservation. A full song cut into pieces adds a caption job per piece and a second render; the pictures you generate for the slots are priced by their own models.

From Video captions, Timeline 1.0, Music Router and API pricing, read 2026-09-29.
StepCallPrice
Generate the song (optional)POST /v1/music-router/generate$0.125 per audio
Lay pictures under the songPOST /v1/timeline-1.0/renderUp to $0.10 per output minute
Burn the timed lyricsPOST /v1/video-captions$0.20 per job of up to 60 seconds

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume