OpenAI-compatible /v1/audio/speech: what Sume's TTS route is
Sume has no `/v1/audio/speech` clone. Its TTS is `POST /v1/tts-1.0/generate` or the Router, a job-based request authenticated with a Sume key.

Sume does not expose an OpenAI-format /v1/audio/speech route, so you cannot point an OpenAI SDK at it. Its text to speech is its own request: POST /v1/tts-1.0/generate, or POST /v1/tts-router/generate when you want to name a catalog model. Both create a job, and you authenticate with your Sume API key only.
What did Inworld ship?
Inworld's TTS release notes dated September 11, 2026 (read 2026-09-30) say its realtime TTS now serves POST /v1/audio/speech in OpenAI's format, so apps built on OpenAI's text-to-speech can switch by pointing the SDK at Inworld and choosing an Inworld model and voice. The notes say the official Python and Node.js SDKs work unchanged, including streaming responses.
How does Sume's request differ?
The shapes are different enough that a drop-in swap will not work.
| Aspect | Inworld (per notes) | Sume |
|---|---|---|
| Route | POST /v1/audio/speech, OpenAI format | POST /v1/tts-1.0/generate or POST /v1/tts-router/generate |
| Model | An Inworld model | TTS 1.0 has no engine picker; the Router requires model, such as sonic-3.6 |
| Voice | An Inworld voice | avatar_id, avatar_handle or voice.id |
| Auth | OpenAI SDK pointed at Inworld | Sume API key only; no provider credentials |
| Result | Speech response | A job with a status URL and audio |
What does a Sume call look like?
Send transcript plus a voice selector. On TTS 1.0, model and model_id are rejected with 400. On the Router, model is required and comes from GET /v1/tts-router/models; an unknown id fails with 400 model_not_found.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: speech-demo-001" \
-d '{
"model": "sonic-3.6",
"transcript": "Thanks for calling.",
"avatar_handle": "acme",
"language": "en"
}'Is it a streaming response?
No. The route description calls it an async job with poll or webhook delivery, non-streaming. A sync wait is clamped to 0 to 30 seconds, and the schema says that ceiling bounds the HTTP wait, not the job. If the job is not finished, the response is still a 2xx with the current state and polling URLs; keep polling and do not submit a second paid job for the same intent. Audio longer than 1200 seconds fails with tts_duration_exceeded. For latency-sensitive use, read text to speech streaming API.
Sources
Related posts
More in Developers
- OpenAI whisper-1 shutdown Feb 2027: a Sume STT alternative
OpenAI removes whisper-1 on Feb 26, 2027. Sume's STT request names no provider model, takes a public audio URL, and caps reserved duration at 10 minutes.
- OpenAPI 3.2 generator: Sume's spec is still 3.0.3
Sume's published OpenAPI document declares 3.0.3, not 3.2. Check that your generator reads 3.0.x before pointing it at api.sume.com/reference/json.
- OpenCode MCP timeout: 5000 ms tools fetch vs Sume jobs_wait
OpenCode's remote MCP timeout is in milliseconds, default 5000, and covers fetching tools. Sume's jobs_wait holds up to 55 seconds per call, a separate limit.
- OpenRouter video provider.options on Sume: rejected, not dropped
OpenRouter lists provider passthrough configuration. Sume v1 runs one backend per model, so a non-empty provider.options returns 400 unsupported_parameter.
Written by Sume