Medical transcription API: what Sume STT does and doesn't
Sume has one speech-to-text route, sume/stt-1.0, with no medical model option. What it takes and returns, so you can decide for clinical audio.

Sume does not offer a medical transcription model. It has one speech-to-text route, public model id sume/stt-1.0, and its request has no field to choose a domain or a medical vocabulary. If you need clinical audio handled by a medical-specific model, that is a different product from what the Sume reference describes.
The medical model news is from ElevenLabs' changelog; the Sume facts are from the API reference. Both read 2026-09-30. This is a feature comparison, not compliance or legal advice.
What did ElevenLabs ship?
The changelog entry for September 11 describes Scribe v2 Medical as a new batch speech recognition model specialized for medical and clinical audio, billed at the same rate as Scribe v2.
What does Sume's speech-to-text take and return?
POST /v1/stt-1.0/transcribe takes a public HTTPS audio_url. Optional fields are language_code (omit it for auto-detect), duration_seconds and segmentation. The reference says provider model ids stay internal and that provider knobs such as diarize and tag_audio_events are fixed server-side.
A completed job returns text, language fields when available, and words[] with { word, start, end } in seconds from the start of the audio.
| Item | What the reference says |
|---|---|
| Model id | sume/stt-1.0, one route |
| Input | Public HTTPS audio_url |
duration_seconds | Optional, 1 to 600; omit to reserve 1 minute |
language_code | Optional hint; omit for auto-detect |
| Word timings | Always returned in words[] |
| Domain or medical option | None in the request |
How do I decide for clinical audio?
Run a sample of your own recordings through the route and compare the text against a reference transcript you trust, paying attention to drug names and measurements. The reference gives no accuracy figure for medical terms, and this post does not either. Also check your own handling and consent obligations for health recordings with whoever advises you on that.
What about long recordings?
The reference describes duration_seconds as a value up to 10 minutes used for usage reservation. A longer consultation would need to be split into pieces before submission; batch transcription covers submitting many files, and the 30-second sync wait is a limit on the HTTP call, not the job.
Sources
Related posts
More in Models
- Seedance 1.5 Pro retires Nov 11: which Seedance ids Sume lists
ByteDance retires Seedance 1.5 Pro on November 11, 2026. Sume's video docs list seedance-2.5 and seedance-2; use bare catalog ids and check the models endpoint.
- Seedance 2.5 reference limit: BytePlus says 50, Sume varies by model
BytePlus allows up to 50 multimodal references for one Seedance 2.5 clip. Sume documents limits per model, so read the catalog for the id you send. Examples.
- Seedream 5.0 pro layer decomposition vs a Sume Seedream result
Seedream 5.0 pro can split one image into a base plus up to 16 PNG layers. Sume lists Seedream 5.0 Lite, and results are one Sume-hosted URL per image.
- Suno v6 audio and video input vs Sume: text and one image
Suno v6 says you can create with text, audio, images and video. Sume's music request takes a text prompt and one optional image; audio and video are not inputs.
Written by Sume