Text to speech streaming API: what Sume returns instead

Sume's text to speech API doesn't stream audio chunks. It returns a finished file per job; split long scripts into sentence jobs to start playback sooner.

5 min readSume
All posts

A text to speech streaming API sends audio in chunks while the speech is still being synthesized, for example over a WebSocket or server-sent events, so playback can start before the whole text is done. Sume's text to speech API does not stream: TTS 1.0 runs an async job and returns a finished audio file, and the Developer API has no SSE or WebSocket transport today.

The facts below come from the TTS 1.0 route in the Sume API reference (the OpenAPI document behind the API reference docs) and Jobs and results, read on 2026-09-29. If your product needs audio to start within a live conversation, you need a streaming TTS service; if it can wait for files, the pattern below shortens the wait before the first sentence plays.

What does Sume's TTS API return instead of a stream?

The API reference describes TTS 1.0 as “async job + poll/webhook (non-streaming).” Every delivery mode ends with the same finished file; none sends partial audio.

From the TTS 1.0 route and job modes in the Sume API reference and Jobs and results, read 2026-09-29.
ModeWhat you get and when
async (default)A job with status_url and result_url right away; poll until terminal is true, then read the result
sync / subscribeThe same job, after a blocking wait of at most 30 seconds; if it isn't done, keep polling and don't resubmit
webhookA terminal callback only (job.completed, job.failed, job.canceled); no progress or partial callbacks
events_urlA pull snapshot of job events, not a stream

How do I start playback sooner on a long script?

Make the unit of work smaller. Split the script into sentences, submit one TTS job per sentence at once, and play the files in script order as each one completes. The first sentence can play while later ones are still in the queue. This is a pattern, not a latency guarantee: how long each job takes is not documented.

  • Give each sentence job its own Idempotency-Key, and reuse that key if you retry the same submit.
  • Concurrency is a dispatch limit, not a submit limit: jobs above your workspace's concurrency limit are accepted as queued while queue capacity remains, then start in turn.
  • Keep the same voice and language on every sentence job so the pieces match.
  • Poll each job's status_url at least once, even after a replayed submit, and read its result only when it is terminal; /result answers 409 job_not_completed until result_ready is true.
// Server-side Node 18+: one TTS job per sentence, played in script order.
const auth = { Authorization: `Bearer ${process.env.SUME_API_KEY}` };
const get = async (url) => (await (await fetch(url, { headers: auth })).json()).data;

async function speak(sentence, key) {
  const res = await fetch("https://api.sume.com/v1/tts-1.0/generate", {
    method: "POST",
    headers: { ...auth, "Content-Type": "application/json", "Idempotency-Key": key },
    body: JSON.stringify({ transcript: sentence, voice: { id: process.env.SUME_VOICE_ID } }),
  });
  if (!res.ok) throw new Error(`submit ${res.status}`);
  const { status_url, result_url } = (await res.json()).data;
  let job = { next_poll_after_seconds: 1 };
  do {
    await new Promise((r) => setTimeout(r, 1000 * (job.next_poll_after_seconds ?? 2)));
    job = await get(status_url);
  } while (!job.terminal);
  if (job.sume_status !== "completed") throw new Error(`TTS job ${job.sume_status}`);
  return (await get(result_url)).result.audio_url;
}

// script and play() are yours: the text, and your audio player.
const sentences = script.match(/[^.!?]+[.!?]+/g) ?? [script];
const jobs = sentences.map((s, i) => speak(s.trim(), `script-42-part-${i}`));
for (const job of jobs) play(await job);

Can one request return audio per sentence?

Yes, but after the job completes, not during it. Send timestamps: { "words": true } and segmentation: { "mode": "sentence" }, and the result carries gapless segments[]. With a wav or raw container, each segment also gets its own audio_url; with MP3 you get the timings without segment files. That helps with captions and editing, not with starting playback early. Text to speech API in Python shows those request fields.

What does it cost, and what are the limits?

TTS 1.0 costs $0.0475 per 1,000 characters, plus a 5.5% agent fee by default; spaces and punctuation count. Each sentence job bills for its own characters, so splitting a script doesn't reduce the characters you pay for.

  • Up to 20,000 characters of transcript per request.
  • Synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded, and no credits are captured.
  • A client-side timeout doesn't cancel a job: it keeps running and still bills.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume