Can AI make a lyric video? Yes, if you time the lines
AI can make a lyric video: pictures under the song, plus each lyric line burned in at the time it is sung. You supply the line timings.

Yes, AI can make a lyric video. A lyric video is two layers: a picture track that runs under the song, and the words of each line burned into the picture at the moment they are sung. AI tools can generate the song and the pictures and draw the text, but the line timings are the part you supply.
On Sume that is a Timeline render with the song as its sound, then a captions job that burns your lines as timed cues. The facts below come from the Video captions, Timeline 1.0 and Music Router docs and Sume's code, read on 2026-09-29.
How do I make a lyric video with AI?
Work in this order, so the words go on last:
- Get the song as a Sume file. A track from
POST /v1/music-router/generatecomes back as an audio artifact; completed jobs returnmedia.sume.comURLs. - Build the picture track with
POST /v1/timeline-1.0/render: the song asaudio.url, its length asaudio.duration_seconds, and clips or stills from your workspace asvideo[]slots. Can AI make a music video for my song? covers choosing a clip per section. - Write the lyrics as cues: one line per cue, with
startandendin seconds. - Send the rendered video's URL and the cues to
POST /v1/video-captions. Cues skip speech-to-text and burn exactly that copy at those times.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: lyric-video-chorus-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/chorus.mp4",
"cues": [
{ "text": "We run until the lights go down", "start": 0.4, "end": 3.2 },
{ "text": "and sing it back again", "start": 3.4, "end": 6.0 }
]
}'Can AI time the lyrics to the song for me?
Sume documents no way to time sung lyrics automatically, so plan to time them yourself. Without cues, the captions job runs speech-to-text, which needs audible speech; the docs don't cover sung vocals over music. With cues, each line appears at the start and end you give, and a job takes up to 200 cues.
A generated song can help with the words: the Music Router result carries result.lyrics, the model-reported lyrics or section map, when present. The docs don't say it includes timings, so listen through and note each line's start and end.
If you omit style, Latin text gets slam, which in current code uppercases every word. Name a style if you want another look.
Can I use my own song?
Only in part. Timeline reads only files already in your Sume workspace on media.sume.com, and there is no public route to upload a local file, so an MP3 on your computer can't be sent straight to a Timeline render; the render's sound has to be a Sume-hosted track, such as one from the Music Router. The captions job is different: it takes any public HTTPS video URL. A short video that already carries your song, hosted at a public HTTPS address, can have lyrics burned in directly.
How long can a lyric video be?
The picture track can run up to 1,800 seconds. The captions job is the limit: in current code it refuses a source longer than 60 seconds, and a source with no audio stream. A song always has audio, so the 60 seconds is what matters.
For a full song, cut the render into pieces of 60 seconds or less, caption each piece with its own cue times starting from 0, and join the pieces in a second render. Add captions to a long video by API shows the split and rejoin. Put the song back as the audio of that final render: in current code a Timeline render takes sound only from its audio track and optional soundtrack, and each clip's own sound is dropped.
How much does an AI lyric video cost?
Each step is billed on its own, plus a 5.5% agent fee by default. The render reserves by output minute and captures its own compute, never above the reservation. A full song cut into pieces adds a caption job per piece and a second render; the pictures you generate for the slots are priced by their own models.
| Step | Call | Price |
|---|---|---|
| Generate the song (optional) | POST /v1/music-router/generate | $0.125 per audio |
| Lay pictures under the song | POST /v1/timeline-1.0/render | Up to $0.10 per output minute |
| Burn the timed lyrics | POST /v1/video-captions | $0.20 per job of up to 60 seconds |
Sources
Related posts
More in Use cases
- AI meditation music generator: calm tracks, longer sessions
An AI meditation music generator makes calm instrumental tracks from a text brief. How to prompt one, join tracks into a session, and add a voice.
- AI meditation video generator: voice, calm loop, music bed
An AI guided meditation video is a slow narration over a calm visual loop and a soft music bed that dips under the voice. How to build one, and costs.
- AI menu generator: make the artwork, type the prices
An AI menu generator works best for the artwork: backgrounds, borders, and dish photos. Type dish names and prices yourself. Sizes, prompts, and cost.
- AI movie trailer generator: shots, narration, music, titles
An AI movie trailer is short generated shots, a narrator, music that builds and burned-in title cards, cut together. How to make one, with costs.
Written by Sume