gpt-transcribe expected languages vs Sume STT language_code

OpenAI's gpt-transcribe accepts multiple expected input languages. Sume STT takes one optional language_code hint or auto-detect. What to send for mixed audio.

4 min readSume
All posts

Sume STT takes at most one language hint. language_code is an optional BCP-47 or provider hint such as en or ko, and the docs say to omit it for auto-detect. For a clip that mixes two languages, omit the hint instead of picking one, then check the result. OpenAI's gpt-transcribe, per its changelog, supports multiple expected input languages.

OpenAI facts are from its changelog; Sume facts from the API reference. Read 2026-10-01.

What does OpenAI say gpt-transcribe supports?

The changelog entry says both gpt-transcribe and gpt-live-transcribe support free-form transcription context, keyword hints and multiple expected input languages, and points to its transcription guide for supported outputs. This post does not go beyond that entry.

Language input from the OpenAI changelog and the Sume API reference, read 2026-10-01.
ItemOpenAI gpt-transcribeSume STT 1.0
Expected languagesMultipleOne optional hint
No hintNot stated in the entryAuto-detect
Context and keyword hintsSupportedNot offered on the STT request

What should I send for a bilingual clip on Sume?

Omit language_code. A single hint names only one language, and the docs offer no way to name two. Read the returned text and the language fields when available, and spot-check the switch points using words[] timings.

When should I split the audio?

If one auto-detect pass gives poor results, cut the clip at the language change using the word timings, and submit each part with its own hint. That costs extra jobs, so use it only where needed. Hints for Korean are covered in language hints: Korean or auto-detect.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume