HeyGen voice clone API: 20+ minutes, paid slot, no Sume clone

HeyGen's professional voice clone needs 1-10 recordings totaling 20+ minutes and a paid slot. Sume's TTS tool selects voices and has no clone upload.

4 min readSume
All posts

HeyGen's professional voice clone, added in September 2026, trains from 1-10 recordings of one speaker totaling 20+ minutes, and each cloned voice occupies a purchased slot. Sume's TTS tool does not train a voice: it takes a voice selector, and its description lists no voice-cloning or voice-upload widget.

HeyGen figures below are from its API changelog; Sume statements are from the TTS tool description in the hosted MCP server code and the MCP tools and gates page, read 2026-09-30.

What does HeyGen's professional clone require?

The changelog says you train a dedicated voice adapter on the HeyGen Voice model, using 1-10 recordings of the same speaker totaling 20 minutes or more. POST /v3/models/audio/voices creates the voice (or retrains one when voice_id is supplied), and you poll the voice until its status is ACTIVE.

HeyGen professional voice clone terms as written in its changelog, read 2026-09-30.
ItemHeyGen changelog
Training audio1-10 recordings, same speaker, 20+ minutes total
SlotEach voice occupies a purchased professional voice clone slot
TrainingsFive per slot per monthly billing period
Synthesis0.6 API credits per generated minute
Over the slot limitThe voice returns voice_expired until a slot is added
Short recordingInstant Clone remains available to all API plans

What does Sume's TTS tool take instead?

The tts_create tool description asks for exactly one transcript input plus a voice selector: voice.id, or avatar_id / avatar_handle, in which case Sume resolves that avatar's ready TTS voice. It states that there is no voice-cloning or voice-upload widget, so a recording of your speaker is not an input to that tool.

Spend works per character: the tool is paid per character, requires an idempotency_key, and supports dry_run to preview cost and max_spend_usd to cap it.

Which one fits my project?

If the requirement is that the voice be a specific real person trained from their own recordings, the HeyGen route described above is built for that, and its slot and credit terms apply. If you need narration in a chosen catalog voice, or the voice already attached to a Sume avatar, selecting by id is enough and nothing needs training. The Sume docs describe no voice-cloning route; see voice cloning API for the wider picture.

What should I check before committing to a clone?

Confirm you hold consent from the speaker for the recordings you upload, and read the vendor's own terms; this post is not legal advice. Then work out the slot count you need, because HeyGen's changelog ties each voice to one slot. For avatar videos where the voice comes with the avatar, see Generate avatar video for how scripts become speech.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume