MAI-Voice-2.1-Flash lists 45 ms; Sume TTS is an async job

Microsoft lists MAI-Voice-2.1-Flash at about 45 ms and MAI-Voice-2.1 at about 550 ms. Sume TTS 1.0 is non-streaming: a job with a poll URL or webhook.

4 min readSume
All posts

Microsoft lists MAI-Voice-2.1-Flash at about 45 ms latency and MAI-Voice-2.1 at about 550 ms. Those figures suit a voice agent that speaks as it thinks. Sume TTS 1.0 does not stream: the OpenAPI text says "Phase 1 is async job + poll/webhook (non-streaming)". It fits narration and video, not a live call.

Microsoft's numbers are from its MAI-Voice page, read 2026-10-01; the page does not say what the latency measures, so treat it as a vendor-stated figure. Sume's are from the API reference.

What does a Sume job do instead of streaming?

It returns a job. With mode: "async" the first response carries a status_url, result_url, events_url and cancel_url. You poll status, or pass a webhook_url and receive signed terminal events (job.completed, job.failed, job.canceled); there are no partial or progress callbacks.

Can I make it feel instant with sync mode?

Partly. mode: "sync" waits up to wait_timeout_seconds, at most 30. If the job is not done by then, the response is still successful and returns the queued or processing state, so keep polling status_url instead of resubmitting. This bounds the HTTP wait, not the job.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: line-001" \
  -d '{
    "transcript": "Your order has shipped.",
    "avatar_handle": "@narrator",
    "mode": "sync",
    "wait_timeout_seconds": 20
  }'

Which use cases fit each?

Delivery model by use case, read 2026-10-01.
Use caseBetter fit
Live call or agent turnA streaming engine like Flash
Narration for a videoSume async job
Batch of voiceover linesSume async jobs with Idempotency-Key
Per-sentence timings for captionsSume timestamps.words and sentence segments

What if I need both?

Use a streaming vendor for live turns and Sume for rendered output. Do not wait on Sume inside a live conversation; its own docs say clients needing a longer wait should submit async and poll from their side.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume