Text to speech reading a confirmation code digit by digit
To make TTS read a code one digit at a time on Sume, write the transcript with the digits separated, pin a Sonic id, and listen before you ship it.

On Sume, the safe way to get a confirmation code read one digit at a time is to write the digits separated in the transcript (for example 4 8 2 1), submit it with a pinned Sonic model id, and listen to the audio before it goes out. Sume's TTS request has no text-normalization setting, so the transcript you send is what gets spoken.
Cartesia's Sonic 3.6 notes say the model voices confirmation codes correctly without preprocessing, and its advanced guide shows an OTP read digit by digit through a normalization setting. Sume's request schema does not list that setting, so this post covers what you can do with the fields Sume does have.
What does Cartesia say about codes in Sonic 3.6?
Two statements, both read on 2026-09-30. The Sonic 3.6 page says the model follows your transcript faithfully and voices confirmation codes and heteronyms correctly without preprocessing. The advanced-capabilities guide shows a Hindi sentence containing an OTP, 4821, read the English way as four eight two one rather than as a Hindi number, by setting normalization to en-IN on Cartesia's own API.
| Question | Cartesia docs | Sume request |
|---|---|---|
| Codes read correctly without preprocessing | Stated for Sonic 3.6 | Not a Sume claim; listen to the output |
| Digit-by-digit control | normalization setting with a locale such as en-IN | No such field in the schema |
| Language of the transcript | language or locale | language, 2 to 16 characters |
How do I write the transcript for a code?
Sume reads the transcript literally: the MCP tool guidance says the literal transcript is the default in every thread. So put the shape you want heard in the text itself. Separate the digits with spaces or commas, and keep the sentence around them short.
Whether a given model then pauses or groups the digits as you hoped is something to hear, not assume. Sume's guidance is that the receipt proves input integrity, not pronunciation, so a successful job tells you the text arrived unchanged, not how it sounded.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: otp-demo-001" \
-d '{
"model": "sonic-3.6",
"transcript": "Your confirmation code is 4 8 2 1.",
"avatar_handle": "acme",
"language": "en"
}'Which model id should I pin?
On the TTS Router, model is required and is a pass-through catalog id. The schema's enum lists sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview. Pin sonic-3.6 explicitly if the Cartesia claim above is the reason you are choosing it; an alias such as sonic-latest can point somewhere else later. The ids are covered in Cartesia Sonic 3.6 API model ids.
Can a pronunciation dictionary help?
The TTS schema has an optional pronunciation_dict_id field described only as an optional pronunciation dictionary id. Sume's docs do not say how a dictionary treats digit strings, so test it on your own codes before relying on it. For general word fixes, see Text to speech pronunciation.
What should I check before sending codes to customers?
Generate a few real-looking codes, including ones with repeated digits and leading zeros, and listen to each. Keep the code itself out of logs if it is a live secret. For call and voicemail flows that use short fixed prompts, see IVR and voicemail prompts with text to speech.
Sources
Related posts
More in Use cases
- Twitch clip captions: edit text and timing, burn them in
Twitch is rolling out editable clip captions. For a clip file you edit yourself, Sume burns your own cues with text, start and end, no speech-to-text.
- Twitch vertical clips: crop an older 16:9 clip to 9:16
Twitch Dual Format makes vertical clips going forward. For a 16:9 clip without one, crop a 9:16 slice with video-filter: width 0.316, height 1.
- Walmart main image size: 2200x2200 and Sume's multiples of 16
Walmart wants a 2200x2200 px main image under 5 MB. 2200 is not a multiple of 16, so generate 2208x2208 on Sume's gpt-image-2.5, then resize.
- WordPress 7.1 client-side media processing and Sume images
WordPress 7.1 resizes and converts uploads in the browser (default quality 0.82). Sume has no output_compression, so request the format you want.
Written by Sume