Medical transcription API: what Sume STT does and doesn't

Sume has one speech-to-text route, sume/stt-1.0, with no medical model option. What it takes and returns, so you can decide for clinical audio.

4 min readSume
All posts

Sume does not offer a medical transcription model. It has one speech-to-text route, public model id sume/stt-1.0, and its request has no field to choose a domain or a medical vocabulary. If you need clinical audio handled by a medical-specific model, that is a different product from what the Sume reference describes.

The medical model news is from ElevenLabs' changelog; the Sume facts are from the API reference. Both read 2026-09-30. This is a feature comparison, not compliance or legal advice.

What did ElevenLabs ship?

The changelog entry for September 11 describes Scribe v2 Medical as a new batch speech recognition model specialized for medical and clinical audio, billed at the same rate as Scribe v2.

What does Sume's speech-to-text take and return?

POST /v1/stt-1.0/transcribe takes a public HTTPS audio_url. Optional fields are language_code (omit it for auto-detect), duration_seconds and segmentation. The reference says provider model ids stay internal and that provider knobs such as diarize and tag_audio_events are fixed server-side.

A completed job returns text, language fields when available, and words[] with { word, start, end } in seconds from the start of the audio.

Sume STT request and result, from the API reference read 2026-09-30.
ItemWhat the reference says
Model idsume/stt-1.0, one route
InputPublic HTTPS audio_url
duration_secondsOptional, 1 to 600; omit to reserve 1 minute
language_codeOptional hint; omit for auto-detect
Word timingsAlways returned in words[]
Domain or medical optionNone in the request

How do I decide for clinical audio?

Run a sample of your own recordings through the route and compare the text against a reference transcript you trust, paying attention to drug names and measurements. The reference gives no accuracy figure for medical terms, and this post does not either. Also check your own handling and consent obligations for health recordings with whoever advises you on that.

What about long recordings?

The reference describes duration_seconds as a value up to 10 minutes used for usage reservation. A longer consultation would need to be split into pieces before submission; batch transcription covers submitting many files, and the 30-second sync wait is a limit on the HTTP call, not the job.

Sources

Related posts

More in Models

All Models posts

Written by Sume