Edit a transcript with an AI instruction: Sume's STT flow
ElevenLabs STT accepts an edit instruction and returns edited_transcript. Sume's STT has no such field: it returns text and word timings to edit yourself.

Sume's speech-to-text cannot rewrite a transcript from an instruction. POST /v1/stt-1.0/transcribe returns text and word timings, and its request body is closed, so an extra instruction field is rejected. To edit, you transcribe first and apply the edit in your own step.
ElevenLabs facts are from its changelog; Sume facts from the API reference, both read 2026-09-30.
What did ElevenLabs add?
The September 28 changelog entry says batch and realtime speech-to-text accept a natural-language edit instruction of up to 2,000 characters, with the result returned in edited_transcript.
What fields does Sume's STT give me?
The request schema sets additionalProperties to false, so only the documented fields are accepted: audio_url, language_code, duration_seconds, segmentation, metadata and the usual mode, webhook_url and wait_timeout_seconds communication fields. Completed results expose text, language fields when available, and words[] as { word, start, end }.
With segmentation: { mode: "sentence" } you also get sentence segments derived from the returned word timings; the schema says it fails closed with a typed error if the provider returns no timed words.
| Step | ElevenLabs STT (changelog) | Sume STT (reference) |
|---|---|---|
| Edit instruction field | Up to 2,000 characters | None; body is closed |
| Edited text output | edited_transcript | Not returned |
| Original text | Not stated in the entry | text |
| Timings | Not stated in the entry | words[] always, optional sentences |
How do I edit a transcript on Sume?
Keep the words[] timings as the source of truth, make your text edit (by hand or with your own model), and map the changed words back to their start and end seconds. Timed words are what a cut list or subtitle file needs.
Two worked examples: edit subtitles, then burn them and a paper edit from a transcript.
Why keep the original text too?
An edit that rewrites wording no longer matches the audio. If the edited text drives captions, keep the original text and word timings so you can show what was actually said and re-time any change against it.
Sources
Related posts
- Auto-generate subtitles API: transcribe, fix the text, then burn
- Paper edit by API: build a rough cut from transcript lines
- Batch transcription API: transcribe many audio files
- Transcribe a video into sentence segments with Sume: options and cost
- Medical transcription API: what Sume STT does and doesn't
More in Developers
- ElevenLabs per-key concurrency caps vs Sume's plan limit
ElevenLabs lets enterprise service account keys carry TTS, music and dubbing concurrency limits. Sume has one plan-based limit and no per-key setting.
- 429 enforced_spend_limit_reached: why retrying never works
A 429 with no retry-after can be a monthly spend cap, not a rate limit. What Anthropic's page says, and how Sume separates 402, 429 rate_limited and queue_full.
- WAV 16-bit PCM vs 32-bit float: what audio detach returns
Sume audio detach writes wav as 16-bit PCM (pcm_s16le) by default. If your camera records 32-bit float, here is what that means and which options you get.
- Face swap API: quality has no default, unlike avatar videos
Avatar videos default to plus when quality is omitted. Sume's Beta face swap has no default: quality is required and takes standard, plus or max.
Written by Sume