AI recipe video generator: step photos, voice, quantities
An AI recipe video turns a photo of each step into a short clip, reads the method aloud, and burns quantities on as text. Steps, costs and limits.

An AI recipe video generator makes a step-by-step cooking video without filming each step: a photo of each step becomes a short moving clip, a synthetic voice reads the method, and the ingredients and quantities are burned on as text you typed. Start from your own step photos where accuracy matters, because a generated dish shows what the model imagines, not your recipe.
Below is how to build one with Sume's API; in the Agents tab you can ask for the same video in plain words, and the agent asks before it spends. Facts come from the Video generation, Timeline 1.0 and Video captions docs and the Sume API reference, read on 2026-09-29; limits marked as current behavior are read from Sume's code.
How do I make a recipe video with AI?
- Write the method as one sentence per step, with quantities in the text: “Whisk 2 eggs with 100 ml milk.”
- Voice it in one
POST /v1/tts-1.0/generatecall withtimestamps.words: trueandsegmentation.mode: "sentence". The result carries gapless sentencesegments[]with start and end times, so each step's sentence tells you how long its shot runs. Slideshow with voiceover shows the mapping. - Animate each step: send its photo to
POST /v1/videosas thefirst_frameinframe_images, at a public HTTPS URL, with a small, plain motion (“the whisk turns slowly, steam rises”). Make each clip at least as long as its sentence; a shorter source is padded or looped by the render, with a warning. - Join: one
POST /v1/timeline-1.0/renderwith the narration asaudio.url,audio.duration_secondsset to its length, and onevideo[]slot per step, each starting at its sentence's start. In current code each clip's own sound is dropped, so viewers hear the narration and an optionalsoundtrack. - Burn the quantities: send the render's
video_urltoPOST /v1/video-captionswith timedcues(next section).
Why start from real step photos?
Only the first frame of each clip is your photo; every frame after it is generated. A model asked to show “fold in the flour” can show the wrong tool, the wrong amount, or a different pan. Your photo pins the start of each step, and the narration and captions carry the instructions, so treat the text layer as the source of truth and watch every clip before publishing. If you need clean dish photos first, AI food photography covers restyling a real dish.
How do I show ingredients and quantities on screen?
As caption cues. Each cue is a text with a start and an end in seconds; cues skip speech-to-text and burn exactly that copy at those times, and one job takes up to 200 cues. Put one ingredient or one quantity per cue and time it to the step's sentence, so the words on screen match what the voice says. Leave style out and Latin text gets slam, which in current code sets the words in capitals, so “ml” reads “ML”.
In current code the caption job refuses a video longer than 60 seconds, so a short-form recipe fits in one job. For a longer recipe, caption it in parts under 60 seconds and rejoin them, as in add captions to a long video.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: recipe-text-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/pancakes.mp4",
"cues": [
{ "text": "2 eggs + 100 ml milk", "start": 0, "end": 4.2 },
{ "text": "120 g flour", "start": 4.2, "end": 8.6 }
]
}'How much does an AI recipe video cost?
A 45-second, six-step recipe is six video clips, one narration, one render and one caption job.
| Step | Call | Price |
|---|---|---|
| Animate each step | POST /v1/videos | By model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models |
| Narration | POST /v1/tts-1.0/generate | $0.0475 per 1,000 characters |
| Join | POST /v1/timeline-1.0/render | $0.10 per output minute, reserved in whole minutes |
| Quantities as text | POST /v1/video-captions | $0.20 per job, for videos up to 60 seconds |
What are the limits?
- Clip length depends on the model, up to 30 seconds on the longest; check
supported_durationsonGET /v1/videos/models. - Step photos must be at public HTTPS URLs. The render takes only this workspace's
media.sume.comfiles, such as the clips and narration Sume returned. - The render's default output is 1080×1920, a vertical frame; set
output.widthandoutput.heightfor a landscape recipe video. - A TTS
transcripttakes up to 20,000 characters per call.
Sources
Related posts
More in Use cases
- AI Santa video: a personal message with the child's name
Make an AI Santa video in three steps: a Santa character still, the message as speech, then lip sync. One short job per child, billed per second.
- AI stock photo generator: stock-style images made to order
An AI stock photo generator makes the image a stock search would find, at your size, several per call. How to prompt it, what it costs, and the rights.
- AI text to product image: what a prompt alone can give you
AI text to product image draws a product that fits your words, not your actual product. Use it for concepts; for listing photos, add a real photo.
- AI travel video generator: your trip photos or a prompt
An AI travel video generator either animates your own trip photos or invents a place from a prompt. How each works, what to label, and costs.
Written by Sume