Muse Voice Transcribe keyword biasing vs Sume STT: no vocab field
Muse Voice Transcribe has keyword biasing for names and domain terms. Sume STT has no vocabulary field: its request takes a language hint and word timings.

Meta's Muse Voice Transcribe lets you improve recognition of names and domain-specific terms with keyword biasing. Sume STT has no such field. Its request is audio_url plus optional segmentation, language_code, duration_seconds, metadata and the job-delivery fields (mode, webhook_url, wait_timeout_seconds), and provider knobs are fixed server-side.
Meta's feature is from its developer guide; Sume's from the /v1/stt-1.0/transcribe schema in the API reference, read 2026-10-01. There is a longer treatment of the closed request in custom vocabulary and Sume STT.
What does Meta's keyword biasing do?
The guide says the keywords parameter is for names, jargon and product words the model would otherwise mis-hear, and that a languageBias parameter is for when you already know the language. It adds that both "nudge the model rather than constrain it, so neither guarantees an exact spelling".
What does Sume's request let me set?
The schema lists a fixed set of fields. language_code is an optional hint, and omitting it means auto-detect. Nothing biases the recognizer toward specific words.
| Control | Muse Voice Transcribe | Sume `sume/stt-1.0` |
|---|---|---|
| Boost names or jargon | keywords (repeatable) | None |
| Language hint | languageBias | language_code, optional |
| Word timings | Turn-level timestamps in the post | words[], always returned |
| Provider knobs | Not applicable | Fixed server-side |
How do I fix mis-heard names on Sume?
Treat it as a post-processing step. The result carries word timings, so you can find the wrong token, replace it in your own text, and keep the times. For wholesale rewrites, see edit a transcript with an AI instruction.
Does a language hint help with names?
The docs call language_code a hint, not a vocabulary. Setting it can help when the language is known, but Sume does not document it as a way to improve specific terms, so test it on your own clips.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-demo-001" \
-d '{
"audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
"duration_seconds": 120,
"language_code": "en"
}'Sources
Related posts
More in Models
- Nano Banana multiple images at once: n range vs Gemini's count
Google says Gemini won't always return the exact image count you ask for in a prompt. On Sume, set n, and read each model's n range from the catalog.
- Nano Banana video to image: poster from a video via Sume stills
Google's Gemini API takes a video as context for a thumbnail or poster. Sume's images API takes image references only, so grab a still frame first.
- OpenAI organization verification for GPT Image 2.5 and Sume ids
OpenAI says you may need API Organization Verification before using GPT Image models. Sume lists the same two models as openai/gpt-image-2.5 and -sunburst.
- Pika Camera Director vs a prompted video edit on Sume
Pika Camera Director reshoots a scene from new angles. On Sume, Omni Flash 1.1 edits a clip from a video_url and a prompt; there is no camera-angle parameter.
Written by Sume