HeyGen image to video from a photo API vs Sume image_url

HeyGen type image animates a PNG or JPEG with a script and voice_id, no avatar setup. Sume does it with image_url plus audio on the Fabric route.

4 min readSume
All posts

On HeyGen, POST /v3/videos with type: "image" takes a PNG or JPEG of a person plus a script and voice_id, and the page says "no avatar setup needed". On Sume the same shortcut is the Fabric route: send a public HTTPS image_url with a Sume-hosted audio_url, and no avatar record is created.

What does HeyGen's image type need?

An image of a person, as a public URL or an uploaded asset, and a voice_id. The image object replaces avatar_id; type: "image" and type: "avatar" are mutually exclusive. The result is polled with GET /v3/videos/{video_id}. Read 2026-10-01.

What does the Sume equivalent need?

POST /v1/veed/fabric-1.0 takes exactly one visual source, either image_url or avatar_id/avatar_handle, plus audio_url and duration_seconds. The schema describes image_url as a public HTTPS still image for audio-driven image-to-video. resolution is 480p or 720p, default 720p. The models page advises using the generated, inspected posed still as image_url, and avatar_handle only when the user named that avatar.

Where do the two differ?

Photo-to-talking-video inputs, read 2026-10-01.
InputHeyGen type imageSume Fabric
Image fieldimage (url or asset id)image_url
Speechscript + voice_idSume-hosted audio_url
Avatar recordNot neededNot needed with image_url
Output sizeresolution such as 1080p480p or 720p

When should I make a reusable avatar instead?

If you will reuse one face across many videos, create an avatar from a photo with input.type: "photo" on POST /v1/avatar-1.0/generate, then call it by avatar_handle. Creating it is a separate job, per the avatar docs. For a one-off clip, image_url is enough; see lip-sync with a photo and audio.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume