HeyGen text to video API: 5-15 s, 768p vs Sume durations

heygen-video-1 makes 5-15 second clips at 480p or 768p from a 5,000-character prompt. Sume lists durations and resolutions per model in its catalog.

4 min readSume
All posts

HeyGen's heygen-video-1 text-to-video endpoint, POST /v3/models/videos, makes 5–15 second videos at 480p or 768p from a prompt of up to 5,000 characters. Sume has no single window like that: each model in the catalog advertises its own supported_durations and supported_resolutions, so read the list for the model you call.

HeyGen facts are from its September 2026 changelog; Sume facts are from Video generation, read 2026-09-30.

What does the HeyGen model accept?

Per the changelog, mode picks text_to_video, image_to_video or reference_to_video. Image-to-video follows the first frame's proportions. Creation returns 202 with a video_id; you poll until completed, failed or cancelled.

How do Sume's windows differ by model?

The Sume docs give these examples of non-uniform limits. A catalog read at request time is the source of truth.

Duration and resolution limits quoted from the Sume video docs, read 2026-09-30.
ModelDurationsResolutions
seedance-2.54–30 s480p, 720p, 1080p
wan-3.02–30 sPer catalog entry
minimax-h35–15 sNative 480p and 768p
Most other modelsUp to 15 sPer catalog entry

Which Sume model has the same 5-15 s, 768p window?

minimax-h3 is listed at 5–15 seconds with native 480p/768p, and the docs say 768p is first-class rather than 720p. That is the closest match to the HeyGen numbers; whether the output looks alike is something to test, not assume.

How do I read a model's limits before submitting?

Call the model catalog and read supported_durations, supported_resolutions and supported_frame_images for the id. Image-to-video on Video 1.0 sends prompt plus image_url as the first frame. For per-model length background, see AI video length limits by model.

Sources

Related posts

More in Models

All Models posts

Written by Sume