Voice cloning API: clone once in the app, speak via the API

Sume has no API route that creates a voice clone. You clone once in the app, then send the voi_ id to the text to speech API on every call.

4 min readSume
All posts

A voice cloning API has two parts: a call that turns an audio sample into a reusable voice, and a call that speaks text in that voice. On Sume only the second part is an API. You clone a voice once in the Sume app (Assets → Voices, from an audio file you upload or a recording), copy its voi_ id, and send that id as voice.id on POST https://api.sume.com/v1/tts-1.0/generate from your backend.

The API facts come from the TTS 1.0 schema in the Sume API reference, the OpenAPI document behind the API reference docs; the app steps come from the Sume app's current code. Both were read on 2026-09-29. The step-by-step clone walkthrough, with its upload limits, is in AI voiceover in your own voice.

Is there an API endpoint that creates the clone?

No. The public API has no route that creates a voice from a sample or from a description. Both happen in the app: the Voices page says “Clone a voice or invent one from a prompt”, and its create dialog offers “Upload or record audio to clone, or describe a person and we generate the voice.” So plan the clone as a one-time manual step per voice, and automate everything after it.

From the TTS 1.0 schema in the Sume API reference and the Sume app's current code, read 2026-09-29.
StepWhere it happensWhat you get
Clone a voice from an upload or a recordingSume app, Assets → VoicesA voice in your Voices library
Invent a voice from a descriptionSume app, Assets → VoicesA voice in your Voices library
Get the voice's idThe voice's Copy ID actionA Voices library id: voi_ plus 32 hex characters
Speak text in the voicePOST /v1/tts-1.0/generate with voice.idA job whose result links the audio file

How does my backend speak in the cloned voice?

Store the voi_ id in your config or database next to the voice's name, and send it on each request. Keep the API key in a server-side environment variable: Sume's authentication docs tell browser and mobile clients to call your backend, which attaches the key.

  • transcript: the text, up to 20,000 characters; spaces and punctuation count toward usage.
  • voice.id: the voi_ id. Library ids are resolved to the stored voice before the job is queued.
  • language: set it for every non-English transcript; when omitted it defaults to English.
  • Idempotency-Key header: reuse the same key when you retry a submit, so the retry returns the original job instead of billing a second one.
  • The submit returns a job; poll status_url until it is terminal, or pass a webhook_url for the terminal callback, then read result_url.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: welcome-line-001" \
  -d '{
    "transcript": "Thanks for calling. Your order is on its way.",
    "voice": { "mode": "id", "id": "voi_…" },
    "language": "en",
    "mode": "async"
  }'

What happens if the voice id is wrong?

The request fails fast. voice.id must be a TTS voice UUID or a Voices library id (voi_ plus 32 hex characters), not a voice name from another TTS product; any other shape is rejected with 400 and invalid_voice_id before a job is queued or credits are reserved. Copy the id verbatim from the app.

If you send an avatar reference (avatar_id or avatar_handle) together with voice.id, the two must match, or the request fails with 400.

What does speech in a cloned voice cost?

Speech in a cloned voice is billed as text to speech: $0.0475 per 1,000 characters, plus a 5.5% agent fee by default, counted on the transcript's characters. The Audio section of the API rate card has three rows, music generation, speech transcription and text to speech; it lists no price for the in-app clone step itself. The output defaults to MP3 at 44,100 Hz and 128 kbps; wav and raw are the other containers.

Whose voice can I clone?

Sume's Terms of Service put this on you: “You represent that you have all rights and permissions needed for the content you submit, including permission to use any person's likeness or voice.” The acceptable-use list also says not to submit voices “that you do not have permission to use.” That is what the terms say; it is not legal advice.

Paid plans include commercial use of generated outputs “as described on the pricing page,” per the same terms; text to speech commercial use covers that question.

Sources

Related posts

More in Models

All Models posts

Written by Sume