Captions AI avatar looks vs one Sume avatar handle per look
Captions added Avatar Looks to save new looks per avatar. On Sume you create one avatar handle per look from a prompt, a profile or a reference image.
Captions' August 2026 notes say you can create and save new looks for avatars and use them across Prompt to Video and other avatar projects. Sume has no separate look object in the docs I read: to get a new look you create a new avatar, from a prompt, a profile or a reference image, and keep its handle.
Captions text is from its what's-new page and Sume's from the docs, both read 2026-10-01.
What are the three ways to make an avatar on Sume?
The Create new avatar page lists three inputs. Each request creates a job; poll it until it completes, then use the returned avatar handle or resource id to generate avatar videos.
| Input | What you provide |
|---|---|
| Prompt | Describe the avatar you want |
| Profile | Structured traits (the props input type in the API) |
| Image | A reference image (input.type: "photo") |
How do I keep several looks for one character?
Create one avatar per look and name the handles so the relationship is obvious, for example one handle for the casual look and one for the studio look. The route is POST /v1/avatar-1.0/generate with a top-level avatar_handle; a leading @ is allowed and Sume stores it without it. This is a naming convention on your side, not a documented look-grouping feature.
How do I use a look in a video?
Pass the handle as avatar_handle to POST /v1/avatar-1.0/talking-video. A product image is optional: omit product_image for a productless avatar video. See Generate avatar video. For comparison with another vendor's look packs, read HeyGen look packs vs a new Sume avatar per look.
Does Sume copy a person's likeness from a video upload?
The avatar docs describe prompt, profile and image inputs only. Check the page for current inputs before building a flow that depends on anything else.
Sources
Related posts
More in Use cases
- Multilingual video captions: language is a hint, not a font
On Sume video captions, `language` only tells speech-to-text what to expect. Look and font come from `style`, `design` and `font`, never from the language.
- Chrome Web Store 440x280 promo tile: generate at 1760x1120
The Chrome Web Store small promo tile is 440x280, below GPT Image 2.5's pixel floor. Request 1760x1120 (exact 4x, same 11:7 ratio) on Sume and downscale.
- Chrome Web Store screenshots at 1280x800: request it directly
Chrome Web Store screenshots are 1280x800 (or 640x400), JPG or PNG, square corners, full bleed. 1280x800 passes GPT Image 2.5 rules, so request it on Sume.
- Keep one color, rest gray, in a video by API with colorhold
Use colorhold in Sume's video filter to keep one color and gray out the rest of a clip. Free check first, then one encode job on sources up to 300 seconds.
Written by Sume