AI meditation music generator: calm tracks, longer sessions

An AI meditation music generator makes calm instrumental tracks from a text brief. How to prompt one, join tracks into a session, and add a voice.

5 min readSume
All posts

An AI meditation music generator makes calm instrumental tracks from a text brief: you describe a slow tempo, soft sustained instruments and no drums or vocals, and it returns a track. Single tracks run up to a few minutes, so for a 20- or 30-minute session you generate several in the same key and tempo and join them into one file.

On Sume the Music Router generates each track at $0.125 per audio generation, plus a 5.5% agent fee by default. Prompting comes from Music 1.0, joining from Timeline audio, and the voice step from the Sume API reference and Audio detach, read on 2026-09-29. Nothing here is a health or sleep claim.

How do I write a prompt for meditation music?

Use the seven-axis brief the Music docs suggest (the general template is in AI music prompt examples), with every axis pointed at stillness. The axes are creative directions, not guaranteed settings, so listen to each take:

  • Emotion, precisely: "hushed, unhurried, warm" rather than "relaxing".
  • Tempo as a number, kept slow, and the same key on every track you plan to join.
  • Two to four instruments with texture: "soft analog pad", "low drone", "distant singing bowl".
  • An arc with one gentle moment, e.g. "a second pad enters at 1:00".
  • Exclusions in the prompt itself ("no percussion"), since a non-empty negative_prompt returns 400, and the close "Instrumental, no vocals."
  • Length in the prompt ("a 3-minute track"); duration is rejected.
curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: calm-session-a-part-1" \
  -d '{
    "prompt": "Hushed, unhurried and warm ambient drone, 60 BPM, D major. Soft analog pad, low cello drone, distant singing bowl. A second pad enters at 1:00. A 3-minute track. No percussion. Instrumental, no vocals."
  }'

How do I make a 30-minute meditation track?

Join the tracks you keep with a Timeline audio concat. The join is sample-domain, with no silence at the seams, but it has no crossfade setting: each part starts where the last one ends. That is why matching key, tempo and instruments across the parts matters.

From Timeline audio, read 2026-09-29.
RuleWhat the docs say
Parts1 to 20, in order; each { url, source_in?, duration? }
InputsThis workspace's media.sume.com audio, such as your generated tracks
Channel layoutAll parts must share one, or audio_parts_channel_mismatch
Output lengthUp to 1,800 s (30 minutes)
Output formatWAV (pcm_s16le) by default, or MP3
Price$0.01 per job

How do I add a guided meditation voice over the music?

Voice the script with text to speech at a slow pace: generation_config.speed goes down to 0.6. Then mix it in a Timeline 1.0 render, with the voice as the audio spine and a music track as the soundtrack, and pull the mix back out as audio with audio detach; How to mix voice with background music walks through those steps. The render runs for its declared audio.duration_seconds, which that post sets to the voice's length, so the guided part lasts as long as the reading. For a long session:

  • A Timeline render's audio runs 1 to 1,800 seconds.
  • Audio detach outputs at most 900 seconds per job; a longer mix needs one range request per stretch.
  • For music-only minutes before or after the voice, join the guided mix and your music tracks with one concat. Parts with different channel layouts are refused (audio_parts_channel_mismatch), so compare the detach result's channels with your tracks first.

How much does AI meditation music cost?

Each generation is a flat $0.125 per audio generation; each join or detach is one small job; the voice is billed per character and the render per started output minute. The sessions below are example assumptions.

Computed from API pricing (music, text to speech, Timeline render) and the Timeline audio and Audio detach rates, read 2026-09-29. Before the 5.5% agent fee.
ExamplePrice
One calm track, 3 takes$0.38
Music-only session: 10 takes, 8 kept and joined$1.26
Guided track: 3 takes, a 3,000-character script, a 5-minute mix$1.03

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume