OmniHuman 1.5 API: native 1080p vs Sume avatar 720p

BytePlus says OmniHuman 1.5 outputs native 1080p from one image plus audio. Sume's Avatar Video currently outputs 720p; its video models list 1080p separately.

4 min readSume
All posts

BytePlus describes OmniHuman 1.5 as video from a single image and multimodal prompts, with native 1080p output. Sume's Avatar Video is not that model: its resolution is currently 720p. If 1080p matters more than the avatar workflow, Sume's general video models list 1080p in their catalog entries.

What does the BytePlus page claim?

The page lists four capabilities: video from a single image with audio, image and text prompts; native 1080p-resolution output; rhythmic, emotional and multi-person performances; and gesture control with camera movement. These are vendor claims as read 2026-09-30; this post does not test them.

What limits does Sume's Avatar Video have?

From the Avatar videos docs, read 2026-09-30:

Avatar Video limits from the Sume docs, read 2026-09-30.
FieldDocumented value
resolutionCurrently 720p
aspect_ratio1:1, 3:4, 9:16, 4:3, 16:9; default 9:16
Avatars per videoOne resolved avatar per final video
Qualitystandard, plus (default), max

So is multi-person output possible on Sume?

Not in one Avatar Video. The docs say current execution supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene. For two speakers, see AI avatar conversation video with two speakers for how Sume handles that case.

Where does Sume list 1080p?

The video generation catalog returns supported_resolutions per model, with example values such as 480p, 720p and 1080p. Read that field for the model you plan to use; it is the source of truth for resolution, as described in Video generation. The quality tier max on Avatar Video is slower turnaround, not a higher resolution.

Sources

Related posts

More in Models

All Models posts

Written by Sume