ElevenLabs STT edit instructions vs Sume transcribe: what Sume returns
ElevenLabs added transcript-editing instructions up to 2,000 characters to its speech to text. Sume's transcribe endpoint returns a plain transcript.

Sume's POST /v1/stt-1.0/transcribe has no transcript-editing instruction field. ElevenLabs added one to its speech to text, up to 2,000 characters, per its changelog entry dated Sep 28, read 2026-10-01.
What does Sume accept?
audio_url, language_code, duration_seconds from 1 to 600 and segmentation set to sentence. Word timings always come back.
How do the two compare?
Fields named in this post, as read on 2026-10-01.
| Provider | Transcript-editing instruction field |
|---|---|
| ElevenLabs speech to text | Yes, up to 2,000 characters |
| Sume POST /v1/stt-1.0/transcribe | No |
How do I clean a transcript?
Edit the text in your own code or with a text model, then use it as the script for TTS or captions. Keep the word timings from the original if you need alignment.
Sources
Related posts
More in Models
- ElevenLabs dialogue continuity vs Sume TTS segments: what differs
ElevenLabs Text to Dialogue links takes with previous_text and request ids. Sume TTS has no such fields; it returns gapless sentence segments instead.
- Adobe Firefly Image 5 API: native 4 MP vs Sume resolution tiers
Firefly's Image5 model is described as native 4 MP with Instruct Edit. Sume does not list a Firefly model; it uses 512, 1K, 2K and 4K tiers.
- Fish Audio Drama 3 single-word fix vs Sume sentence segments
Fish Audio says Drama 3 preview can fix a single word. Sume has no word-level repair: regenerate a sentence and join takes with Timeline audio concat.
- FLUX 1.1 Ultra raw mode and 4MP: what Sume image size accepts
Raw mode and 4MP belong to BFL's FLUX1.1 [pro] Ultra endpoint. Sume lists FLUX.2 ids and sizes images with aspect_ratio, resolution tiers and image_size.
Written by Sume