OpenAI speech API has six formats; Sume TTS has mp3, wav, raw
OpenAI lists MP3, Opus, AAC, FLAC, WAV and PCM for speech output. Sume TTS 1.0 has three containers: mp3, wav and raw, with set sample rates and encodings.

OpenAI's text-to-speech guide lists six response formats: MP3, Opus, AAC, FLAC, WAV and PCM, with PCM described as raw 24 kHz 16-bit signed little-endian samples without a header. Sume TTS 1.0 offers three containers: mp3, wav and raw. If your pipeline needs Opus, AAC or FLAC, plan a conversion step on your side.
OpenAI's list is from its guide, read 2026-10-01; Sume's from the OpenAPI schema behind the API reference.
Which formats map to which?
Only three line up directly.
| Format | OpenAI | Sume TTS 1.0 |
|---|---|---|
| MP3 | Yes, the default | Yes, the default; 44100 Hz, 128 kbps |
| WAV | Yes | Yes, with a PCM encoding |
| PCM | Yes, 24 kHz 16-bit | raw container, choice of encoding and rate |
| Opus, AAC, FLAC | Yes | Not offered |
What can I tune on Sume's formats?
Sample rate can be 8000, 16000, 22050, 24000, 44100 or 48000 Hz. MP3 bit rate can be 32000, 64000, 96000, 128000 or 192000. WAV and raw take pcm_f32le, pcm_s16le, pcm_mulaw or pcm_alaw. MP3 cannot carry a PCM encoding.
What do I send to match OpenAI's PCM?
Raw container, 24000 Hz, signed 16-bit little-endian.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"transcript": "Match the other API.",
"avatar_handle": "@narrator",
"output_format": { "container": "raw", "sample_rate": 24000, "encoding": "pcm_s16le" }
}'Is there an OpenAI-compatible endpoint?
See the OpenAI-compatible audio speech endpoint post for what the Sume route accepts; this post covers the output formats only.
Sources
Related posts
More in Developers
- OpenRouter 402 with credits left vs Sume 402 on reserve
OpenRouter can return 402 with a positive balance when the in-flight budget is full. Sume's 402 means the estimate could not be reserved. Telling them apart.
- OpenRouter audio/speech returns bytes; Sume TTS returns a job
OpenRouter's /audio/speech streams raw audio bytes. Sume's TTS Router returns a job id and a media URL, so a byte-stream client needs a poll step.
- OpenRouter base64 input_audio vs Sume STT 1.0 audio_url
OpenRouter transcription takes base64 audio inside the JSON body. Sume STT 1.0 takes an audio_url instead, so you host the file and send a link.
- OpenRouter Batch API vs Sume async jobs: which for video?
OpenRouter's Batch API is for text and embeddings with a 24 hour window. For video files, Sume returns a job id per request: async, sync, or webhook mode.
Written by Sume