gpt-realtime-2.1 only supports v1/realtime: Sume TTS is a job

gpt-realtime-2.1 works only on v1/realtime, not Chat Completions or Batch. Sume TTS 1.0 is an async job you poll, which suits narration, not live talk.

4 min readSume
All posts

OpenAI's model page lists v1/realtime as the only supported endpoint for gpt-realtime-2.1; Chat Completions, Batch and Fine-tuning are marked not supported. Sume's TTS 1.0 route works the other way round: POST /v1/tts-1.0/generate creates an async job that you poll or receive by webhook, and it does not stream.

Facts are from the OpenAI model page and Sume's OpenAPI file and docs, read 2026-10-01.

What does v1/realtime only mean?

You cannot send gpt-realtime-2.1 a one-shot request on the usual text endpoints or queue it as a batch. Every use is a realtime session, so your client holds a live connection for as long as the audio flows.

How does a Sume TTS request differ?

The OpenAPI description of /v1/tts-1.0/generate says phase 1 is "async job + poll/webhook (non-streaming)". With mode: async the response returns status_url, result_url, events_url and cancel_url. Completed results expose mirrored audio artifacts. Related audio routes such as Timeline audio use the same envelope: GET /v1/jobs/:id/status, then GET /v1/jobs/:id/result.

Request shape, OpenAI model page and Sume docs, read 2026-10-01.
gpt-realtime-2.1Sume TTS 1.0
Endpointv1/realtime onlyPOST /v1/tts-1.0/generate
BatchNot supportedOne job per request
DeliveryLive sessionPoll status_url or webhook
Input cap128,000 context window20000 characters of transcript
OutputStreamed audioAudio file artifact

What audio format comes back?

A file. The request has an output_format object with a container of mp3, wav or raw and a sample_rate in Hz, defaulting to mp3 at 44100. Sample rates in the schema include 8000, 16000, 22050, 24000, 44100 and 48000. See TTS for video narration for using the file in a timeline.

Which one fits which job?

A live conversation where the user interrupts needs a realtime session, and Sume's TTS 1.0 route is not one: it is non-streaming. Pre-rendered narration, voiceover for a video or a phone-line prompt fits a job: submit the transcript, get the job id, poll, and download the artifact. The transcript is capped at 20000 characters, so split long scripts into several jobs.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume