HeyGen audio to video max length: 10 or 30 min vs Sume 300 s
HeyGen's pages say 30 minutes per audio-to-video request and 10 minutes for avatar audio input. Sume's Fabric route caps duration_seconds at 300.

HeyGen's pages give two figures: the Audio to Video page says one request renders up to 30 minutes of audio, while the Usage Limits page lists avatar audio input at a maximum of 10 minutes (600 seconds). On Sume, the Fabric route accepts duration_seconds up to 300, so audio longer than five minutes must be split across requests.
Numbers below are as written on the pages read 2026-10-01. If your audio sits between 10 and 30 minutes, test one request before building around either figure.
What does HeyGen say about audio length?
The Audio to Video page uses POST /v3/videos with audio_url or audio_asset_id instead of script. It says the video's length follows your audio, up to 30 minutes, and tells you to split longer recordings into segments of 30 minutes or less. The Usage Limits page, under avatar input, says "Audio input: Maximum 10 minutes (600 seconds)". It also notes a 30-minute cap per scene on output video.
What is the audio cap on Sume?
The VEED Fabric 1.0 route is POST /v1/veed/fabric-1.0. Its request needs audio_url, a measured duration_seconds, and exactly one visual source. The OpenAPI schema describes duration_seconds as the audio duration used to reserve credits at admit, with a maximum of 300 and a minimum of 1. The sibling MiniMax H3 Max Lip Sync route uses the same body with audio of 5 to 14.8 seconds, per the models page.
How do the limits compare side by side?
| Source | Field or page | Stated limit |
|---|---|---|
| HeyGen Audio to Video | audio_url / audio_asset_id | Up to 30 minutes per request |
| HeyGen Usage Limits | Avatar input, audio | 10 minutes (600 seconds) |
| Sume Fabric 1.0 | duration_seconds | 1 to 300 seconds |
| Sume MiniMax H3 Max Lip Sync | Audio length | 5 to 14.8 seconds |
What should I do with audio longer than 300 seconds?
Cut it at sentence boundaries into clips of 300 seconds or less, measure each clip's length, and send one Fabric request per clip. Then join the finished clips. Lip-sync a long video over 5 minutes walks through that split-and-join approach, and split audio first for clips over 15 seconds covers the shorter route.
Sources
Related posts
More in Developers
- HeyGen avatar script limit: 5,000 characters vs Sume's 4-60 s
HeyGen caps avatar script text at 5,000 characters. Sume has no character cap in its docs: it accepts scripts estimated at 4-60 seconds inclusive.
- HeyGen download_failed error: URL rules vs Sume's media errors
HeyGen download_failed means a media URL was not public, mismatched or corrupt. Sume returns image_not_fetchable or input_media_unreachable for similar cases.
- HeyGen image to video from a photo API vs Sume image_url
HeyGen type image animates a PNG or JPEG with a script and voice_id, no avatar setup. Sume does it with image_url plus audio on the Fabric route.
- HeyGen insufficient_credit vs Sume insufficient_credits (402)
HeyGen returns code insufficient_credit with HTTP 402. Sume returns insufficient_credits, plural, also 402. Match on the exact string and the status.
Written by Sume