Deepgram nova-3-pharma vs Sume STT: drug-name transcripts
Deepgram added nova-3-pharma for English drug names. Sume STT has one public model, sume/stt-1.0, so check each drug name against word timings.

Sume has no pharma model to pick. Speech-to-text runs on one public model id, sume/stt-1.0, and the request accepts no model selector, so you cannot ask for drug-name vocabulary the way Deepgram's new nova-3-pharma does. What Sume returns is words[] with timings, which lets a reviewer check each medication name against the audio.
Deepgram's side is from its changelog; Sume's side is from the API reference. Both read 2026-10-01.
What did Deepgram release?
The Deepgram changelog entry dated Sep 17, 2026 describes nova-3-pharma as a new Nova-3 model built for pharmaceutical vocabulary, focused on accurate drug-name recognition, for pharmacy and healthcare voice-agent workflows. It is available in English for both batch and streaming.
What does Sume STT let me choose?
Little, by design. The request takes a public audio_url, an optional language_code, and an optional duration_seconds. The schema says provider knobs such as diarize and tag_audio_events are fixed server-side, and provider model ids stay internal.
| Question | Deepgram (changelog) | Sume STT 1.0 |
|---|---|---|
| Domain model for drug names | nova-3-pharma, English | None; one public model sume/stt-1.0 |
| Language input | English model | Optional hint, for example en or ko; omit for auto-detect |
| Word timings | Not stated in the entry | Always returned in words[] as { word, start, end } |
| Provider settings | Chosen by the caller | Fixed server-side |
How do I verify a drug name with Sume?
Transcribe, then treat every medication token as unconfirmed. Use the start and end seconds on that word to cut a short clip and have a person listen to it. Word timings are a review aid, not a clinical check, and Sume's docs make no accuracy claim for pharmaceutical terms.
If a drug name is wrong, a language hint will not fix it, because the hint selects a language, not a vocabulary. Plan a human or lookup step after transcription.
What should I do next?
If you need a vocabulary-tuned medical model today, use the vendor's. If you stay with Sume, keep the job id and word timings with each transcript so reviews are repeatable. For a fuller account of what Sume STT returns for medical audio, see what Sume STT does for medical transcription.
Sources
Related posts
More in Models
- ElevenLabs character limits by model vs Sume TTS 20,000
ElevenLabs lists 5,000 characters for v3, 10,000 for v4 and 40,000 for Flash v2.5. Sume TTS 1.0 takes up to 20,000 characters in one request.
- ElevenLabs speech to speech API: what Sume offers instead
ElevenLabs lists speech-to-speech voice changer models. Sume has no voice-to-voice endpoint: transcribe with STT, edit the text, then run TTS.
- FLUX.2 flex steps and guidance: BFL has them, Sume does not
BFL lists adjustable steps and guidance only for FLUX.2 flex ($0.06/MP). Sume lists flux.2-flex but rejects unlisted parameters with 400.
- FLUX.2 max grounding search: what it is, what Sume lists
Only FLUX.2 [max] does web-grounded generation at BFL. Sume lists flux.2-pro and flux.2-flex, so grounding is not available there.
Written by Sume