AI meditation video generator: voice, calm loop, music bed

An AI guided meditation video is a slow narration over a calm visual loop and a soft music bed that dips under the voice. How to build one, and costs.

5 min readSume
All posts

An AI meditation video generator builds a guided session from three parts: a slow narration voiced from your script, a calm visual loop, and a soft instrumental bed that dips under the voice. The video runs as long as the narration, so the script sets the length; on Sume one render runs up to 30 minutes.

Facts come from Sume's Video generation, Music Router, Music 1.0 and Timeline 1.0 docs and the Sume API reference, read on 2026-09-29. In the Agents tab you can ask for the session in plain words; the agent asks before it spends. Nothing here is a health or wellbeing claim.

How do I voice a guided meditation with AI?

Send the script to POST /v1/tts-1.0/generate as transcript, with a voice: voice.id, or the avatar_id of an avatar whose voice is ready. Then these settings matter for a meditation:

  • generation_config.speed is a multiplier from 0.6 to 1.5; the low end gives the unhurried pace a session needs. Listen before you commit.
  • language names the language of the script. Set it for any non-English script; left out, it defaults to English.
  • One call takes up to 20,000 characters, and audio longer than 1,200 seconds fails with tts_duration_exceeded. For a longer script, voice it in passages and list the files in the render's audio.parts[] (up to 20 slices, joined without gaps).

How do I make the calm visual loop?

Make one short clip that ends on the frame it starts on, from a single calm still such as a lake at dusk or a candle flame, with small motion and a still camera. The render then replays it for the whole session with render.pad_mode: "loop". AI lofi video generator walks through the first-and-last-frame request for that clip, and Seamless loop AI video shows how to check the two ends for a jump.

What differs for a meditation is the length: a lofi loop plays under music alone, while here the narration decides how long the loop has to run, so render the voice first and read its length.

How do I put music under the voice?

Generate the bed with POST /v1/music-router/generate. Length is asked for in the prompt, since duration is rejected. Close the brief with “Instrumental, no vocals.” and, under narration, add “no spoken word”. AI meditation music generator has calm briefs to adapt.

Then one Timeline 1.0 render joins everything: the narration is the audio spine, the music is the soundtrack, and the loop fills one video[] slot as long as the session. In current code the clip's own sound is dropped, so viewers hear only the voice and the bed.

From Timeline 1.0 and the Sume API reference, read 2026-09-29.
FieldRangeUse
audio.duration_seconds1–1800The narration's length, which is the video's length
soundtrack.gain_db−60 to 12; default −16Bed level under the voice
soundtrack.duck_db0–20; default 0How far the bed dips while the voice speaks
soundtrack.looptrue / falseRepeats a short bed until the spine ends
soundtrack.fade_out_seconds0–10Fades the bed at the end
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: meditation-001" \
  -d '{
    "audio": {
      "url": "https://media.sume.com/artifacts/artf_demo/narration.wav",
      "duration_seconds": 1200
    },
    "soundtrack": {
      "url": "https://media.sume.com/artifacts/artf_demo/calm-bed.mp3",
      "loop": true, "duck_db": 6, "fade_out_seconds": 10
    },
    "output": { "width": 1920, "height": 1080 },
    "render": { "pad_mode": "loop" },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_demo/lake-loop.mp4", "start": 0, "duration": 1200 }
    ]
  }'

How much does an AI meditation video cost?

A 20-minute session, with an assumed script length and two music takes:

Computed from API pricing and Video generation, read 2026-09-29. Before the 5.5% agent fee charged by default.
PartRateExample
Narration, 12,000 characters$0.0475 per 1,000 characters$0.57
Music bed, 2 takes$0.125 per audio$0.25
Render, 20 minutes$0.10 per output minute$2.00
Visual loop, 1 clipBy model, at provider list × 1.25See pricing_skus on GET /v1/videos/models

What are the limits?

  • One render is at most 30 minutes; for a longer session, render parts and join them in your own editor.
  • Every URL in the render must be this workspace's media.sume.com files, such as the narration, music and clip Sume returned.
  • The default output is 1080×1920, a vertical frame; set output.width and output.height for landscape.
  • duck_db needs a real spine: it is rejected when audio.mode is "silence".
  • Music is generated per brief, not picked from a library, and has no seed or duration setting, so listen to each take.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume