Text to dialogue API: continuity between requests on Sume

Sume's TTS has no previous_text field. Keep a line seamless with one longer transcript, or join separate takes with a gapless Timeline audio concat.

4 min readSume
All posts

Sume's text-to-speech request has no field for the text before or after a line, so continuity between separate requests is not something you pass in. Put the lines that must flow into one transcript, or render takes separately and join them with Timeline audio concat.

The ElevenLabs changelog entry for Sept 28 added previous_text and future_text (max 100 characters each) and previous_request_ids / next_request_ids (max 3) to Text-to-Dialogue. Sume's OpenAPI contract, read 2026-09-30, contains none of those names.

What can I do in one Sume request?

POST /v1/tts-1.0/generate accepts a transcript of up to 20000 characters. Optional timestamps.words with segmentation.mode: "sentence" returns word timings and gapless per-sentence segments[] (segment[i].end == segment[i+1].start); audio slices additionally need emit_audio and a wav or raw container. Only "sentence" is supported in v1. If the whole exchange fits in one transcript, one request keeps its delivery in one render.

How do I join separate takes without a gap?

Import the audio, then call POST /v1/timeline-1.0/audio with operation: "concat" and up to 20 ordered parts[]. The docs describe the join as sample-domain: no re-TTS and no silence at the seams. The call needs an Idempotency-Key, and the job is polled through /v1/jobs/:id/status and /result.

Continuity options, from the ElevenLabs changelog and Sume docs, read 2026-09-30.
NeedElevenLabs Text-to-DialogueSume
Context from neighboring textprevious_text / future_textNo such field; use one transcript
Link to earlier requestsprevious_request_ids / next_request_idsNo such field
Join finished takesNot stated in the changelog entryconcat, sample-domain, no added silence

What can go wrong when joining?

Parts must share one channel layout, or the job fails with audio_parts_channel_mismatch. Keep wav output (the default) when the file will be joined again: mp3 re-adds priming padding at every edge. Joining preserves each take as rendered, so a delivery mismatch between two takes is not smoothed; it is only made gapless. Details are in Timeline audio.

Which one should I pick?

Use one transcript when the lines fit in 20000 characters and one voice setup. Use concat when takes are produced at different times or need individual retries. For per-line voices, see text to speech with multiple voices.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume