TTS reads 1999-2000 wrong: write ranges and fractions as words
Cartesia does not normalize ranges like 1999-2000 or fractions like 2/3. Sume sends your transcript literally, so write them as words before you submit.

Hyphenated date ranges such as 1999-2000 are not normalized by Cartesia, and neither are fractions such as 2/3. Sume has no normalization switch to change that, so rewrite them as spoken words ("1999 to 2000", "two thirds") in the transcript before you submit it.
Cartesia behavior is from its Text normalization guide; Sume behavior is from the TTS tool description and the OpenAPI document behind the API reference. Both were read 2026-09-30.
Which written forms does Cartesia leave alone?
The guide says the normalizer runs by default and converts written forms such as 7:00 PM into spoken forms. It marks the following as not normalized and gives the workaround for each.
| Category | Example | Write it as |
|---|---|---|
| Date ranges | 1999-2000 | 1999 to 2000 |
| Date ranges | Dec 5-Dec 12 | Dec 5 to Dec 12 |
| Money ranges | $10-15 | 10 to 15 dollars |
| Fractions | 2/3 | two thirds |
| Uncommon units | V, Pa | volts, pascals |
What does Sume do with my transcript?
Sume passes your transcript as written; the TTS tool description lists a literal transcript as the default input, and the request schema lists no normalization field, so you cannot switch the provider's normalizer from Sume. A receipt proves input integrity, not pronunciation, so a green job does not mean the date was read correctly: listen to the audio.
The Cartesia guide describes a normalization control with auto, off and code values. The Sume schema does not expose it, so treat pre-normalizing as your only lever here.
How do I pre-normalize in code?
A small replace step before the request handles the forms above. Keep it conservative and check the output by ear, because a blind regex can mangle phone numbers or part numbers.
const spoken = (text) =>
text
.replace(/\b(\d{4})-(\d{4})\b/g, "$1 to $2")
.replace(/\b1\/2\b/g, "one half")
.replace(/\b2\/3\b/g, "two thirds");
console.log(spoken("The war lasted from 1999-2000, about 2/3 of a year."));What if a word still sounds wrong?
That is a pronunciation problem, not a normalization one. Sume's request has an optional pronunciation_dict_id field, described as an optional pronunciation dictionary id; see text-to-speech pronunciation for how to use it. Rewriting the word phonetically in the transcript is the other option.
Does rewriting change the length limit?
Slightly. A transcript can be 1 to 20,000 characters, and "1999 to 2000" is longer than "1999-2000". Only very long scripts need to care; for those, see TTS 1200-second limit for audiobook chapters.
Sources
Related posts
More in Developers
- TTS voice-language mismatch warning: confirm, then retry
Sume's TTS warns when a voice's language differs from `language`. No job or charge exists yet; after the user agrees, retry with `confirm_language_mismatch`.
- Twitch 2K clips: trim a 1440p clip without setting output
Video trim's output width and height are limited to 256-2160, so a 2560-wide 1440p clip cannot be conformed. Omit output and the source size is kept.
- Vercel 413 FUNCTION_PAYLOAD_TOO_LARGE with Sume video
Vercel Functions cap request and response bodies at 4.5 MB. Do not proxy a generated video through one; hand the client the media.sume.com URL instead.
- AI SDK addToolOutput clears approval: Sume idempotency_key
AI SDK 7.0.126 clears tool approvals when addToolOutput runs. If a paid Sume call replays, a stable idempotency_key is what guards the create.
Written by Sume