OpenRouter audio/speech returns bytes; Sume TTS returns a job

OpenRouter's /audio/speech streams raw audio bytes. Sume's TTS Router returns a job id and a media URL, so a byte-stream client needs a poll step.

4 min readSume
All posts

OpenRouter's /api/v1/audio/speech answers with a raw audio byte stream, not JSON. Sume's explicit TTS route, POST /v1/tts-router/generate, answers with a job: you get a job id first, then read a Sume media URL from the result. Porting means replacing "write the response body to a file" with "poll, then download the URL".

OpenRouter facts are from its text-to-speech guide; Sume facts from the OpenAPI file and Jobs and results, read 2026-10-01.

What does OpenRouter return from audio/speech?

The guide describes an endpoint compatible with the OpenAI Audio Speech API: send text, receive a raw audio byte stream in your chosen format. It says the response is not JSON, so you can pipe it to a file or an audio player. Its examples send model, input, voice and response_format.

What does the Sume TTS Router return instead?

The OpenAPI description calls it "Explicit pass-through text-to-speech" and requires a model taken from GET /v1/tts-router/models. Every submit mode returns the job id in its first response, and a 2xx means the job exists, not that audio is ready. The docs tell clients to poll status_url until terminal is true, then read result_url once result_ready is true, and to use the Sume media URLs from the result.

What changes when I port a byte-stream client?

Response handling compared, read 2026-10-01.
StepOpenRouter audio/speechSume TTS Router
First responseRaw audio bytesJob envelope with the job id
Get the audioWrite the body to a filePoll, then fetch the URL in the result
Model choicemodel in the bodymodel from GET /v1/tts-router/models
Text fieldinputTranscript and voice selector per the OpenAPI schema

Should I resubmit if my wait times out?

No. In sync and subscribe modes the wait is bounded at 30 seconds; if the job is not terminal, poll rather than resubmit. See also the OpenAI-compatible speech endpoint note and async TTS jobs.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume