Creatify custom avatar lipsync_input and consent video, explained
Creatify custom avatars are built from a lipsync_input MP4 plus a consent video. On Sume, a photo input with an image_url creates a reusable avatar handle.
Creatify's Custom Avatar API builds an avatar from your own videos: a lipsync_input MP4 used for lipsync training, plus a consent video, in video/mp4 or video/quicktime. Sume's avatar API takes a different route: a reference image sent as a photo input, which produces an avatar you reuse by handle.
Creatify details are from its Custom Avatar page, read 2026-10-01. Sume details are from Avatar and Avatar videos.
What does Creatify's custom avatar request need?
Two endpoints are documented. POST /api/personas_v2/ is a multipart upload of the files; POST /api/personas/ takes publicly accessible URLs as JSON. Both list lipsync_input, creator_name, gender and video_scene as required fields. The page describes uploading your own consent and lipsync videos, and then says to poll GET /api/personas/{id}/ for the approval status.
| Endpoint | Body | Required fields |
|---|---|---|
POST /api/personas_v2/ | multipart/form-data files | lipsync_input, creator_name, gender, video_scene |
POST /api/personas/ | JSON with public file URLs | lipsync_input, creator_name, gender, video_scene |
How does Sume create an avatar from an image?
Per the Avatar docs, use the image route when you have a reference image; the API input type is photo:
curl -X POST https://api.sume.com/v1/avatar-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: avatar-image-001" \
-d '{
"avatar_handle": "reference_presenter",
"input": {
"type": "photo",
"image_url": "https://example.com/reference.png"
}
}'How do I use the avatar afterwards?
Launch requests for Avatar Video reference a ready avatar with a top-level avatar_handle. The docs also say an avatar video uses one resolved avatar per final video. The image_url must be a fetchable public HTTPS URL. For multi-host layouts, see one avatar per video.
Which approach fits which input?
If what you have is footage of a person speaking, Creatify's page is the one that documents video input. If you have a still image, Sume's photo input starts from that image. Check each vendor's current terms for consent requirements before using a real person's likeness; this post is not legal advice.
Sources
Related posts
More in Developers
- Cursor custom mode with a pinned skill: install the Sume skill
Cursor lets any skill be a custom mode pinned in the chat. Install the Sume skill with sume skills install; the mode cannot waive idempotency_key or MCP scope.
- Cursor /goal with paid Sume calls: what actually caps spend
Cursor /goal keeps an agent on one objective. The goal text is no budget: bound Sume spend with generation_spend_cap_usd, max_paid_calls and dry_run.
- D-ID translation result URL: valid 24 hours, then re-fetch
D-ID's translation result_url is valid for 24 hours, so store the file or re-fetch it. Sume media.sume.com URLs do not expire. What to save and when.
- Deepgram Aura-2 allows 2,000 characters per request; Sume 20,000
Deepgram answers 413 when Aura text passes 2,000 characters. Sume TTS 1.0 accepts 1 to 20,000 characters per request, with a separate 1,200-second audio cap.
Written by Sume