Edit a transcript with an AI instruction: Sume's STT flow

ElevenLabs STT accepts an edit instruction and returns edited_transcript. Sume's STT has no such field: it returns text and word timings to edit yourself.

4 min readSume
All posts

Sume's speech-to-text cannot rewrite a transcript from an instruction. POST /v1/stt-1.0/transcribe returns text and word timings, and its request body is closed, so an extra instruction field is rejected. To edit, you transcribe first and apply the edit in your own step.

ElevenLabs facts are from its changelog; Sume facts from the API reference, both read 2026-09-30.

What did ElevenLabs add?

The September 28 changelog entry says batch and realtime speech-to-text accept a natural-language edit instruction of up to 2,000 characters, with the result returned in edited_transcript.

What fields does Sume's STT give me?

The request schema sets additionalProperties to false, so only the documented fields are accepted: audio_url, language_code, duration_seconds, segmentation, metadata and the usual mode, webhook_url and wait_timeout_seconds communication fields. Completed results expose text, language fields when available, and words[] as { word, start, end }.

With segmentation: { mode: "sentence" } you also get sentence segments derived from the returned word timings; the schema says it fails closed with a typed error if the provider returns no timed words.

Instruction editing vs the Sume STT flow, read 2026-09-30.
StepElevenLabs STT (changelog)Sume STT (reference)
Edit instruction fieldUp to 2,000 charactersNone; body is closed
Edited text outputedited_transcriptNot returned
Original textNot stated in the entrytext
TimingsNot stated in the entrywords[] always, optional sentences

How do I edit a transcript on Sume?

Keep the words[] timings as the source of truth, make your text edit (by hand or with your own model), and map the changed words back to their start and end seconds. Timed words are what a cut list or subtitle file needs.

Two worked examples: edit subtitles, then burn them and a paper edit from a transcript.

Why keep the original text too?

An edit that rewrites wording no longer matches the audio. If the edited text drives captions, keep the original text and word timings so you can show what was actually said and re-time any change against it.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume