HeyGen avatar new outfit with reference images vs Sume
HeyGen prompt avatars take avatar_id plus up to three reference_images for a new outfit. Sume's photo input takes one image_url per avatar.
HeyGen can put an existing avatar in a new outfit by passing the look as avatar_id and up to three reference_images (url, asset_id or base64) on a prompt avatar. Sume's photo input takes a single image_url, so a new look there is a new avatar created from one reference photo.
How does HeyGen's new-outfit call work?
With type: "prompt", an avatar_id makes the referenced look's image condition the result, so the person stays recognisable while the prompt changes outfit and setting. reference_images add wardrobe, style or setting guidance; use them alone for a new character or layered on avatar_id. A crop showing only the garment keeps the reference about wardrobe. Read 2026-10-01.
What does Sume take for a new look?
POST /v1/avatar-1.0/generate with input.type: "photo" and one image_url, plus a top-level avatar_handle. The URL must be a fetchable public HTTPS image; localhost, private-network URLs, non-HTTPS URLs and non-image responses are rejected before submission, per the avatar docs. A product image is a separate field on Avatar Video; omit product_image for a productless video, per Generate avatar video.
What are the differences?
| Item | HeyGen prompt avatar | Sume photo input |
|---|---|---|
| Existing look as base | avatar_id | Not a parameter on the photo input |
| Reference images | Up to 3 reference_images | One image_url |
| Result | Saved to the referenced look's character | A new avatar under your avatar_handle |
What should I do for several outfits on Sume?
Make one photo of each outfit and create one avatar per look, each with its own handle. One new avatar per look covers naming, and retouching a look from a photo covers fixes to an existing one.
Sources
Related posts
More in Use cases
- Add a hook title to the first seconds of a video with one cue
Send one authored cue with start 0 and end 3 to POST /v1/video-captions and Sume burns that hook text into the clip, with no speech-to-text step.
- Ken Burns effect API: Timeline stills are static holds
Sume Timeline holds a still image static; motion is accepted and ignored with a motion_ignored warning. zoompan is on the video filter allowlist.
- AI kids story video generator: scenes of 2 to 30 seconds
Make a children's story video as one scene per request: wan-3.0 takes 2 to 30 seconds per scene, then join up to 200 slots in a Timeline 1.0 render.
- AI real estate walkthrough from photos: chain first/last-frame clips
Kling 4.0 takes up to 10 keyframes for a walkthrough. Sume takes a first and last frame per clip, so chain one clip per room pair and join them on a timeline.
Written by Sume