Text to speech streaming API: what Sume returns instead
Sume's text to speech API doesn't stream audio chunks. It returns a finished file per job; split long scripts into sentence jobs to start playback sooner.

A text to speech streaming API sends audio in chunks while the speech is still being synthesized, for example over a WebSocket or server-sent events, so playback can start before the whole text is done. Sume's text to speech API does not stream: TTS 1.0 runs an async job and returns a finished audio file, and the Developer API has no SSE or WebSocket transport today.
The facts below come from the TTS 1.0 route in the Sume API reference (the OpenAPI document behind the API reference docs) and Jobs and results, read on 2026-09-29. If your product needs audio to start within a live conversation, you need a streaming TTS service; if it can wait for files, the pattern below shortens the wait before the first sentence plays.
What does Sume's TTS API return instead of a stream?
The API reference describes TTS 1.0 as “async job + poll/webhook (non-streaming).” Every delivery mode ends with the same finished file; none sends partial audio.
| Mode | What you get and when |
|---|---|
async (default) | A job with status_url and result_url right away; poll until terminal is true, then read the result |
sync / subscribe | The same job, after a blocking wait of at most 30 seconds; if it isn't done, keep polling and don't resubmit |
webhook | A terminal callback only (job.completed, job.failed, job.canceled); no progress or partial callbacks |
events_url | A pull snapshot of job events, not a stream |
How do I start playback sooner on a long script?
Make the unit of work smaller. Split the script into sentences, submit one TTS job per sentence at once, and play the files in script order as each one completes. The first sentence can play while later ones are still in the queue. This is a pattern, not a latency guarantee: how long each job takes is not documented.
- Give each sentence job its own
Idempotency-Key, and reuse that key if you retry the same submit. - Concurrency is a dispatch limit, not a submit limit: jobs above your workspace's concurrency limit are accepted as
queuedwhile queue capacity remains, then start in turn. - Keep the same voice and
languageon every sentence job so the pieces match. - Poll each job's
status_urlat least once, even after a replayed submit, and read its result only when it isterminal;/resultanswers409 job_not_completeduntilresult_readyis true.
// Server-side Node 18+: one TTS job per sentence, played in script order.
const auth = { Authorization: `Bearer ${process.env.SUME_API_KEY}` };
const get = async (url) => (await (await fetch(url, { headers: auth })).json()).data;
async function speak(sentence, key) {
const res = await fetch("https://api.sume.com/v1/tts-1.0/generate", {
method: "POST",
headers: { ...auth, "Content-Type": "application/json", "Idempotency-Key": key },
body: JSON.stringify({ transcript: sentence, voice: { id: process.env.SUME_VOICE_ID } }),
});
if (!res.ok) throw new Error(`submit ${res.status}`);
const { status_url, result_url } = (await res.json()).data;
let job = { next_poll_after_seconds: 1 };
do {
await new Promise((r) => setTimeout(r, 1000 * (job.next_poll_after_seconds ?? 2)));
job = await get(status_url);
} while (!job.terminal);
if (job.sume_status !== "completed") throw new Error(`TTS job ${job.sume_status}`);
return (await get(result_url)).result.audio_url;
}
// script and play() are yours: the text, and your audio player.
const sentences = script.match(/[^.!?]+[.!?]+/g) ?? [script];
const jobs = sentences.map((s, i) => speak(s.trim(), `script-42-part-${i}`));
for (const job of jobs) play(await job);Can one request return audio per sentence?
Yes, but after the job completes, not during it. Send timestamps: { "words": true } and segmentation: { "mode": "sentence" }, and the result carries gapless segments[]. With a wav or raw container, each segment also gets its own audio_url; with MP3 you get the timings without segment files. That helps with captions and editing, not with starting playback early. Text to speech API in Python shows those request fields.
What does it cost, and what are the limits?
TTS 1.0 costs $0.0475 per 1,000 characters, plus a 5.5% agent fee by default; spaces and punctuation count. Each sentence job bills for its own characters, so splitting a script doesn't reduce the characters you pay for.
- Up to 20,000 characters of
transcriptper request. - Synthesized audio longer than 1,200 seconds fails with
tts_duration_exceeded, and no credits are captured. - A client-side timeout doesn't cancel a job: it keeps running and still bills.
Sources
Related posts
More in Developers
- Duck background music under a voiceover with the Timeline API
Set soundtrack.duck_db (0 to 20) on POST /v1/timeline-1.0/render so the music dips under your voiceover spine. It needs a real spine; silence mode is refused.
- Render a silent video from clips with the Timeline API
Set audio.mode to silence and a duration_seconds on POST /v1/timeline-1.0/render to join clips with no audio file. Which fields are illegal there, and pricing.
- Timeline error too_many_chained_transitions: how to fix it
Timeline 1.0 refuses more than 8 adjacent fades with too_many_chained_transitions. Insert a hard cut. Also transition_too_long, transition_not_frame_aligned.
- Render a vertical 1080x1920 video from clips with an API
Timeline 1.0 defaults to a 1080x1920 MP4. Set output width, height and fps, and pick fit cover, contain, stretch or blur for clips that do not match the frame.
Written by Sume