Tavus Phoenix-4.5: photo rules vs a Sume photo avatar
Tavus Phoenix-4.5 starts a face from a photo in minutes and allows glasses, jewelry and hair. Sume creates an avatar from a public HTTPS image URL.
Tavus Phoenix-4.5, announced September 2, 2026, lets you start using a face from a photo or video within minutes, and allows glasses, jewelry and hair in front of the shoulders. Sume creates an avatar from a photo too, via a fetchable public HTTPS image_url, and renders scripted talking videos rather than live conversations.
Tavus claims are from its changelog; Sume's from Avatars and Generate avatar video, read 2026-09-30.
What does the Phoenix-4.5 entry say?
The face is usable in minutes, then tunes in the background into an unwatermarked version. Movement below the neck is more natural; glasses, jewelry and hair in front of the shoulders are allowed; animated human styles (cartoon, anime, Pixar-style) are supported. It became the default for training a new face on September 9, 2026.
What does a Sume photo avatar take?
The avatar create call takes an input of type photo with an image_url. The docs require a fetchable public HTTPS image URL; localhost, private-network, non-HTTPS and non-image responses are rejected before generation is submitted. The docs do not list the Phoenix-style rules above (glasses, jewelry, cartoon styles), so test your image rather than assume.
| Topic | Tavus Phoenix-4.5 (changelog) | Sume (docs) |
|---|---|---|
| Input | Photo or video | Photo image_url, public HTTPS |
| Accessories | Glasses, jewelry, hair in front of shoulders allowed | Not specified in the docs |
| Animated styles | Cartoon, anime, Pixar-style supported | Not specified in the docs |
| Output | Face model for conversations | Rendered clip; resolution is 720p |
What do I get back?
Avatar videos turn a ready avatar into a script-driven talking video, with aspect_ratio of 1:1, 3:4, 9:16, 4:3 or 16:9. To approve a first frame before spending on a full render, create a preview first; see avatar video first-frame previews. For the conversation-versus-clip choice, read real-time avatar vs video avatar API.
How should I test a tricky photo?
Create the avatar from the photo, then run a short script through a preview and look at the still. If glasses or a hairstyle render badly, swap the photo and try again. Do the same for stylized art, since the Sume docs make no promise about it.
Sources
Related posts
More in Models
- TikTok Symphony with Seedance 2.5: 30-second AI video ads
TikTok Symphony now generates up to 30 seconds with Seedance 2.5, in select markets. The same 30-second length is available by API on Sume as seedance-2.5.
- Vidu API: S2-Avatar live voice vs Sume avatar video jobs
Vidu S2-Avatar is a real-time voice model. Sume has no live session: you submit a script and a scene photo to an avatar job and fetch the finished video.
- Vidu S2-Editing: live stream edits vs a Sume Omni clip edit
Vidu S2-Editing edits an incoming video stream in real time. Sume edits one finished clip with a prompt through Gemini Omni video_to_video, as an async job.
- Wan 3.0 on Runway, or the wan-3.0 id as an API job on Sume
Runway's changelog lists Wan 3.0 in tool mode and workflows on paid plans. On Sume, wan-3.0 is a model id you call from code and poll as a job.
Written by Sume