Kazakh speech to text API: Nova-3 and Sume language_code
Deepgram added Kazakh to Nova-3 on Sep 3, 2026. Sume STT 1.0 takes an optional language_code hint, or auto-detect. How to send audio and check the output.

Deepgram's changelog says Nova-3 added Kazakh on Sep 3, 2026. Sume STT 1.0 does not publish a language table, so for Kazakh audio send a public HTTPS audio_url and either omit language_code for auto-detect or pass a BCP-47 hint such as kk, then read the transcript yourself.
Deepgram facts are from its changelog; Sume's from the OpenAPI schema for POST /v1/stt-1.0/transcribe, read 2026-09-30.
What did Deepgram change?
The entry says Kazakh is a new Nova-3 language, alongside improved Nova-3 monolingual models for seven existing languages, for batch and streaming workloads. It describes Deepgram's service; it does not say anything about Sume.
What does Sume STT 1.0 accept?
| Field | What the schema says |
|---|---|
audio_url | Required. Public HTTPS audio URL to transcribe |
language_code | Optional BCP-47 / provider language hint (for example en or ko). Omit for auto-detect |
duration_seconds | Optional, 1-600; maximum 10 minutes |
| Model id | Public id sume/stt-1.0; provider model ids stay internal |
Does Sume support Kazakh?
The docs do not list supported languages, and they do not name Kazakh. Do not assume it from Deepgram's changelog, because Sume's provider model ids stay internal. Run a short Kazakh sample and compare the text with what was said.
Should I set language_code or leave it out?
Try both on the same sample. Omitting it uses auto-detect; setting it gives the provider a hint. For the same trade-off in Korean, see STT language hint. Long files should go in clips of at most 10 minutes, submitted with mode: async as in Jobs and results.
Sources
Related posts
More in Models
- FLUX video edit API: 720p downscale, and video_url edits on Sume
BFL's edit-a-video endpoint downscales sources above 720p. Sume lists no FLUX video; its edit path is video_url with gemini-omni-flash-1.1 only.
- gemini-2.5-flash-image shutdown on October 2, 2026
Google lists gemini-2.5-flash-image for shutdown on October 2, 2026. Sume lists Nano Banana 2 and Pro under google/ ids to move image requests to.
- Gemini 3.5 Transcribe API limits: 1 hour vs Sume STT's 10 minutes
Gemini 3.5 Transcribe takes up to 1 hour per request, 30 minutes with diarization or word timestamps. Sume STT takes 600 seconds; split longer audio.
- Gemini API speaker diarization: 8 speakers; Sume STT has no labels
Gemini 3.5 Transcribe diarizes up to 8 speakers, 3+ experimental, within 30 minutes per request. Sume STT returns word timings but no speaker labels.
Written by Sume