On hold music for business: make the music and messages

On hold music for a business is a calm track, often with spoken messages, played to waiting callers. How to generate both, in phone-ready formats.

5 min readSume
All posts

On hold music for a business is a calm instrumental track your phone system plays to waiting callers, often with short recorded messages between passes. You can license a stock track, or generate your own music and voice the messages with text to speech, then load both into the phone system in the format it asks for.

On Sume the Music Router makes the track at $0.125 per audio generation, and text to speech voices the messages at $0.0475 per 1,000 characters, each plus a 5.5% agent fee by default. Facts come from Music 1.0, the TTS schema in the Sume API reference (see API reference) and Audio detach, read on 2026-09-29.

Why does hold music sound so bad on the phone?

A phone call carries a narrow slice of the audible range and is compressed for speech, not music. Deep bass and bright highs are lost, and dense arrangements smear together. So write the brief for the phone, not for speakers:

  • Few instruments with a clear midrange lead, such as piano, nylon guitar or vibraphone.
  • A steady, moderate tempo and no big drops in level.
  • No vocals, so callers don't mistake the music for someone speaking.

How do I make hold music with AI?

Send a brief to the Music Router. Put the length and every exclusion in the prompt: there is no duration field, and a non-empty negative_prompt returns 400. Tracks run up to a few minutes. Close with "Instrumental, no vocals." as the Music docs suggest; AI music prompt examples has more briefs.

The result is an audio artifact, typically audio/mpeg (MP3), on media.sume.com. The music request has no output-format field, so you can't ask for a telephony format. If your phone system wants one for music, convert the MP3 with the phone system's own tool or an audio editor.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hold-music-main-take-1" \
  -d '{
    "prompt": "Warm, patient and friendly light jazz, 92 BPM, F major. Solo piano melody, soft upright bass, brushed snare. Even level throughout; the melody returns at 1:00. A 2-minute track. Instrumental, no vocals."
  }'

How do I record on hold messages for my business?

Voice each message as its own text to speech request and ask for the format the phone system takes: TTS can return 8 kHz μ-law or A-law WAV directly, as well as MP3 (the default, at 44,100 Hz and 128 kbps) or other PCM. Which format to pick, and how to organize a set of prompts, is covered in IVR text to speech.

  • Keep messages and music as separate files if the phone system can play them in turn. A voice mixed over music (a Timeline render, then audio detach) can't come out at 8 kHz: detach offers 16,000, 44,100 or 48,000 Hz.
  • Use one voice for every message: the same voice.id, or the same ready avatar's avatar_id.

How much does custom on hold music cost?

Music is a flat $0.125 per audio generation, with every take billed; messages are billed per character at $0.0475 per 1,000 characters, spaces and punctuation included. The counts below are example assumptions.

On use: the Sume Terms of Service say Sume does not claim ownership of generated outputs and that paid plans include commercial use as described on the pricing page. Whether a business needs any other license to play music to callers where it operates is outside what this post can answer; this is not legal advice.

Computed from the Music generation and Text to speech rows on API pricing, read 2026-09-29. Before the 5.5% agent fee.
ExampleGenerationsMessagesPrice
One track, 3 takes, no messages30$0.375
One track, 3 takes, 4 messages34$0.413
Two tracks, 3 takes each, 10 messages610$0.8688

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume