Speechmatics Linden for voice agents vs Sume batch STT jobs

Speechmatics announced Linden for voice agents. Sume STT is a bounded batch job with a webhook, not a live stream, so live agents need streaming.

4 min readSume
All posts

Sume STT 1.0 is a batch job, not a live stream, so it suits recordings and not the turn-by-turn listening a voice agent does. Use it when you have a recorded file and can wait for a result or a webhook.

The Speechmatics homepage banner (read 2026-10-01) reads "Introducing Linden, our new model purpose-built for voice agents". The snapshot gives no other Linden details, so this post does not describe it.

What does a voice agent need from speech recognition?

The banner says Linden is built for voice agents, which listen live. A job that waits for a whole file does not fit that loop.

How does a Sume STT job behave?

From the schema: mode: sync and mode: subscribe are aliases for the same bounded wait, and that 30-second ceiling bounds the HTTP wait, not the job. Terminal callbacks are delivered to webhook_url. The optional duration_seconds field, used to reserve usage, accepts at most 600 seconds (10 minutes); omitted, it reserves for 1 minute.

STT 1.0 job behavior, read 2026-10-01.
ItemValue
RoutePOST /v1/stt-1.0/transcribe
HTTP wait ceiling30 seconds; the job continues
CallbackTerminal only, to webhook_url
Usage reservation hintduration_seconds, 1 to 600

Where does batch STT fit around an agent?

After the call: transcribe the recording, read word timings, and feed captions or summaries. Not during the call. See real-time speech to text for the streaming question.

What if a job outlasts the wait?

Do not resubmit; the 30-second ceiling bounds the wait, so read the job later or wait for the webhook. The Scribe v2 Realtime comparison makes the same batch-versus-realtime point.

Sources

Related posts

More in Models

All Models posts

Written by Sume