AI recipe video generator: step photos, voice, quantities

An AI recipe video turns a photo of each step into a short clip, reads the method aloud, and burns quantities on as text. Steps, costs and limits.

5 min readSume
All posts

An AI recipe video generator makes a step-by-step cooking video without filming each step: a photo of each step becomes a short moving clip, a synthetic voice reads the method, and the ingredients and quantities are burned on as text you typed. Start from your own step photos where accuracy matters, because a generated dish shows what the model imagines, not your recipe.

Below is how to build one with Sume's API; in the Agents tab you can ask for the same video in plain words, and the agent asks before it spends. Facts come from the Video generation, Timeline 1.0 and Video captions docs and the Sume API reference, read on 2026-09-29; limits marked as current behavior are read from Sume's code.

How do I make a recipe video with AI?

  • Write the method as one sentence per step, with quantities in the text: “Whisk 2 eggs with 100 ml milk.”
  • Voice it in one POST /v1/tts-1.0/generate call with timestamps.words: true and segmentation.mode: "sentence". The result carries gapless sentence segments[] with start and end times, so each step's sentence tells you how long its shot runs. Slideshow with voiceover shows the mapping.
  • Animate each step: send its photo to POST /v1/videos as the first_frame in frame_images, at a public HTTPS URL, with a small, plain motion (“the whisk turns slowly, steam rises”). Make each clip at least as long as its sentence; a shorter source is padded or looped by the render, with a warning.
  • Join: one POST /v1/timeline-1.0/render with the narration as audio.url, audio.duration_seconds set to its length, and one video[] slot per step, each starting at its sentence's start. In current code each clip's own sound is dropped, so viewers hear the narration and an optional soundtrack.
  • Burn the quantities: send the render's video_url to POST /v1/video-captions with timed cues (next section).

Why start from real step photos?

Only the first frame of each clip is your photo; every frame after it is generated. A model asked to show “fold in the flour” can show the wrong tool, the wrong amount, or a different pan. Your photo pins the start of each step, and the narration and captions carry the instructions, so treat the text layer as the source of truth and watch every clip before publishing. If you need clean dish photos first, AI food photography covers restyling a real dish.

How do I show ingredients and quantities on screen?

As caption cues. Each cue is a text with a start and an end in seconds; cues skip speech-to-text and burn exactly that copy at those times, and one job takes up to 200 cues. Put one ingredient or one quantity per cue and time it to the step's sentence, so the words on screen match what the voice says. Leave style out and Latin text gets slam, which in current code sets the words in capitals, so “ml” reads “ML”.

In current code the caption job refuses a video longer than 60 seconds, so a short-form recipe fits in one job. For a longer recipe, caption it in parts under 60 seconds and rejoin them, as in add captions to a long video.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: recipe-text-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/pancakes.mp4",
    "cues": [
      { "text": "2 eggs + 100 ml milk", "start": 0, "end": 4.2 },
      { "text": "120 g flour", "start": 4.2, "end": 8.6 }
    ]
  }'

How much does an AI recipe video cost?

A 45-second, six-step recipe is six video clips, one narration, one render and one caption job.

From Video generation, Timeline 1.0, Video captions and the Sume API reference, read 2026-09-29. Each rate is plus a 5.5% agent fee by default; see API pricing.
StepCallPrice
Animate each stepPOST /v1/videosBy model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models
NarrationPOST /v1/tts-1.0/generate$0.0475 per 1,000 characters
JoinPOST /v1/timeline-1.0/render$0.10 per output minute, reserved in whole minutes
Quantities as textPOST /v1/video-captions$0.20 per job, for videos up to 60 seconds

What are the limits?

  • Clip length depends on the model, up to 30 seconds on the longest; check supported_durations on GET /v1/videos/models.
  • Step photos must be at public HTTPS URLs. The render takes only this workspace's media.sume.com files, such as the clips and narration Sume returned.
  • The render's default output is 1080×1920, a vertical frame; set output.width and output.height for a landscape recipe video.
  • A TTS transcript takes up to 20,000 characters per call.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume