Dictation API: AssemblyAI cleaned text vs Sume STT word timings

AssemblyAI's Dictation API returns send-ready text. Sume STT returns a transcript with words[] timings and no cleanup flag, so you do the filler removal.

4 min readSume
All posts

A dictation API is a speech-to-text call that returns text ready to send, not a raw transcript. AssemblyAI's Dictation API page says it returns text with filler gone, self-corrections resolved and names spelled right. Sume's STT 1.0 returns a transcript with words[] timings and has no cleanup option, so any filler or correction handling is a step you add after the job.

AssemblyAI facts are from its product page; Sume facts are from the API reference.

What does each service return?

AssemblyAI describes the Dictation API as the first API built for dictation: users speak and it returns finished text. Sume STT 1.0 completes with public-safe text, language fields when available, and words[] word-level timings as { word, start, end } in seconds from the audio start.

Dictation output versus Sume STT output, read 2026-10-01.
PropertyAssemblyAI Dictation API (product page)Sume STT 1.0 (OpenAPI)
OutputText ready to sendTranscript text plus words[]
Filler wordsGoneNot removed by a flag
Self-correctionsResolvedNot resolved by a flag
Word timingsNot stated on the pageAlways returned

Can I turn on cleanup in Sume STT?

No. The request schema says word timings are always returned, that you send segmentation to also get sentence segments, and that provider knobs such as diarize and tag_audio_events are fixed server-side. There is no filler or self-correction option to set.

The request takes a public HTTPS audio_url, an optional language_code (omit it to auto-detect) and an optional duration_seconds from 1 to 600.

What do I do after the transcript comes back?

Filter in your own code. The timed words[] array lets you drop tokens you decide are filler and keep the rest in order. Resolving a self-correction ("Tuesday, no, Wednesday") is a language task and needs a rule or a text model on your side; the timings alone do not do it.

If your goal is a video without fillers rather than clean text, the work happens on the clip instead: see remove filler words from video. Captions are a separate surface; Video captions covers speech-to-captions, which needs audible speech.

When is raw plus timings the better fit?

When you need to know when each word was said: subtitles, highlight-as-you-read, or cutting audio at word boundaries. Cleaned text discards exactly that detail. When you only need a message to paste into a field, a dictation-style service matches the job more directly.

Sources

Related posts

More in Models

All Models posts

Written by Sume