Voice cloning API: clone once in the app, speak via the API
Sume has no API route that creates a voice clone. You clone once in the app, then send the voi_ id to the text to speech API on every call.

A voice cloning API has two parts: a call that turns an audio sample into a reusable voice, and a call that speaks text in that voice. On Sume only the second part is an API. You clone a voice once in the Sume app (Assets → Voices, from an audio file you upload or a recording), copy its voi_ id, and send that id as voice.id on POST https://api.sume.com/v1/tts-1.0/generate from your backend.
The API facts come from the TTS 1.0 schema in the Sume API reference, the OpenAPI document behind the API reference docs; the app steps come from the Sume app's current code. Both were read on 2026-09-29. The step-by-step clone walkthrough, with its upload limits, is in AI voiceover in your own voice.
Is there an API endpoint that creates the clone?
No. The public API has no route that creates a voice from a sample or from a description. Both happen in the app: the Voices page says “Clone a voice or invent one from a prompt”, and its create dialog offers “Upload or record audio to clone, or describe a person and we generate the voice.” So plan the clone as a one-time manual step per voice, and automate everything after it.
| Step | Where it happens | What you get |
|---|---|---|
| Clone a voice from an upload or a recording | Sume app, Assets → Voices | A voice in your Voices library |
| Invent a voice from a description | Sume app, Assets → Voices | A voice in your Voices library |
| Get the voice's id | The voice's Copy ID action | A Voices library id: voi_ plus 32 hex characters |
| Speak text in the voice | POST /v1/tts-1.0/generate with voice.id | A job whose result links the audio file |
How does my backend speak in the cloned voice?
Store the voi_ id in your config or database next to the voice's name, and send it on each request. Keep the API key in a server-side environment variable: Sume's authentication docs tell browser and mobile clients to call your backend, which attaches the key.
transcript: the text, up to 20,000 characters; spaces and punctuation count toward usage.voice.id: thevoi_id. Library ids are resolved to the stored voice before the job is queued.language: set it for every non-English transcript; when omitted it defaults to English.Idempotency-Keyheader: reuse the same key when you retry a submit, so the retry returns the original job instead of billing a second one.- The submit returns a job; poll
status_urluntil it is terminal, or pass awebhook_urlfor the terminal callback, then readresult_url.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: welcome-line-001" \
-d '{
"transcript": "Thanks for calling. Your order is on its way.",
"voice": { "mode": "id", "id": "voi_…" },
"language": "en",
"mode": "async"
}'What happens if the voice id is wrong?
The request fails fast. voice.id must be a TTS voice UUID or a Voices library id (voi_ plus 32 hex characters), not a voice name from another TTS product; any other shape is rejected with 400 and invalid_voice_id before a job is queued or credits are reserved. Copy the id verbatim from the app.
If you send an avatar reference (avatar_id or avatar_handle) together with voice.id, the two must match, or the request fails with 400.
What does speech in a cloned voice cost?
Speech in a cloned voice is billed as text to speech: $0.0475 per 1,000 characters, plus a 5.5% agent fee by default, counted on the transcript's characters. The Audio section of the API rate card has three rows, music generation, speech transcription and text to speech; it lists no price for the in-app clone step itself. The output defaults to MP3 at 44,100 Hz and 128 kbps; wav and raw are the other containers.
Whose voice can I clone?
Sume's Terms of Service put this on you: “You represent that you have all rights and permissions needed for the content you submit, including permission to use any person's likeness or voice.” The acceptable-use list also says not to submit voices “that you do not have permission to use.” That is what the terms say; it is not legal advice.
Paid plans include commercial use of generated outputs “as described on the pricing page,” per the same terms; text to speech commercial use covers that question.
Sources
Related posts
More in Models
- Wan 3.0 1080p: how to request full HD from the API
Wan 3.0 outputs 480p, 720p or 1080p. Send resolution 1080p to wan-3.0 on POST /v1/videos, up to 30 seconds, with audio. See the request and the price.
- Wan 3.0 2-second clips: the shortest video length and what it costs
Wan 3.0's floor is 2 seconds, shorter than the 4-second minimum on Seedance 2.x models. Send duration 2 to wan-3.0; at 720p it costs $0.25 before the agent fee.
- Wan 3.0 30-second video: one clip, no stitching
Wan 3.0 makes a single clip of up to 30 seconds with audio. Set duration to 30 on wan-3.0 in POST /v1/videos. Here is the request, the cost, and the limits.
- Wan 3.0 document to video: what the model reads and what Sume passes
Alibaba's Wan 3.0 can read an uploaded document to make a video. On Sume's video API today wan-3.0 takes prompts, frames and media references, not a file input.
Written by Sume