HeyGen create template via API: Sume has video_inputs, not templates
HeyGen now creates templates over the API and binds variables by element_id. Sume has no stored template object; send ordered video_inputs on each request.

HeyGen's October 2026 changelog adds endpoints to create a Studio template over the API: POST /v3/templates promotes a video to a private template, and PUT /v3/templates/{template_id}/variables binds avatars, voices, images, video and audio by element_id. Sume's docs describe no stored template object; the equivalent is sending ordered video_inputs on each Avatar Video request.
HeyGen details are from its API changelog. Sume details are from Generate avatar video, read 2026-10-01.
How does HeyGen's template flow work?
Make one video, promote it to a template, declare its variables, then generate a video per viewer with POST /v3/templates/{template_id}. GET on a template returns a composition where each part has an id and bindable_as, plus an edit_version. Exact phrases can become {{name}} placeholders through match.
What replaces a template on Sume?
The request body itself. POST /v1/avatar-1.0/talking-video takes either a script or ordered video_inputs, never both. Your code is the template: keep the scene list in your repo and substitute text before each call.
| Need | HeyGen template | Sume Avatar Video |
|---|---|---|
| Reusable layout | Stored template | Your own video_inputs array |
| Per-viewer text | {{name}} placeholder | String substitution in script or input_text |
| Scene look | Bound image or video | scene prompt or photo, or per-scene background |
| Pause beat | Not covered here | voice.type: "silence" with duration |
What are the limits of that approach?
Current execution supports one resolved avatar per final video, and the total planned duration must land in the 4 to 60 second window. Multi-scene input is meant for hooks, demos or silence beats in one composed video. A stored, editable template with version control is not something the docs describe.
How do I keep per-viewer renders safe to retry?
Send an Idempotency-Key on every submit; see the idempotency comparison.
Sources
Related posts
More in Developers
- HeyGen studio video scene limit: 50 scenes, and Sume's 4-60 s plan
HeyGen studio videos allow 50 scenes and 30 minutes per scene. Sume multi-scene video_inputs share one 4-60 second window and one resolved avatar.
- Hookdeck passes gzip bytes unchanged: verify Sume on raw bytes
Hookdeck now forwards binary and gzip bodies byte for byte. Read the raw bytes before parsing and verify the Sume sume-v1 signature over them.
- Hume Octave 2 description field vs Sume TTS emotion controls
Hume lists a voice description field; acting instructions are coming soon. Sume TTS 1.0 has generation_config: emotion text, speed 0.6 to 1.5, volume 0.5 to 2.
- Hume Octave takes 5,000 characters per utterance in MP3, WAV or PCM
Hume's TTS overview lists 5,000 characters per utterance and MP3, WAV or PCM output. Sume TTS 1.0 takes 20,000 characters and mp3, wav or raw containers.
Written by Sume