Creatify Aurora 2.0 max audio length: 59.7 s, and how to split
Creatify Aurora 2.0 takes up to 59.7 s of audio; Aurora v1 takes 5 minutes. Sume Avatar Video accepts 4-60 s per job, so split longer scripts across jobs.
Creatify's API reference says aurora_v2 (Aurora 2.0) takes up to 59.7 seconds of audio, while the endpoint's general audio field says 5 minutes max (the page does not say which model that figure belongs to, so do not assume it holds for Aurora 2.0). If your clip is longer than 59.7 seconds, split it before you call Aurora 2.0.
On Sume, an Avatar Video job accepts a script whose estimated duration is 4-60 seconds inclusive; longer scripts are split into multiple jobs. Vendor facts below are from Creatify's page read 2026-10-01; Sume facts are from Avatar videos.
What are the audio limits side by side?
| Path | Limit as documented |
|---|---|
Creatify aurora_v2 | Up to 59.7 seconds of audio |
| Creatify Aurora audio field (general) | 5 minutes max, mp3 and wav |
Sume Avatar Video (/v1/avatar-1.0/talking-video) | Estimated duration 4-60 seconds inclusive |
Sume Fabric duration_seconds | Maximum 300 in the OpenAPI schema |
Why is 59.7 different from 60?
The two numbers measure different things. Creatify's 59.7 is the longest audio file Aurora 2.0 accepts. Sume's 4-60 window is checked against the duration Sume estimates for the target video, and it applies to a script or a multi-scene video_inputs plan. A 59.7 second voice-over fits both, but a 60 second one fits only Sume's window.
How do I split a longer script?
Cut at sentence or scene boundaries so each part reads as 4-60 seconds, then submit one job per part with the same avatar_handle. Sume's docs say to split longer scripts into multiple jobs; send each job with its own Idempotency-Key, since a key identifies one exact request. You can join the finished clips afterwards; see assemble long-form video with the Timeline API.
Can a request inside the window still fail?
Yes. The OpenAPI description for video_inputs says requests within the duration window can still fail with public_reason content_policy_rejected when a provider rejects generated clip content. Length is a precondition, not a guarantee. If the wallet cannot fund a run you get 402 insufficient_credits instead; see Errors and credits.
Sources
Related posts
More in Models
- Creatify Boreal talking clips vs Sume's still-plus-audio route
Creatify says Boreal's gains are smallest on single-person talking clips. Sume makes every speaking shot from an accepted still plus TTS audio via Fabric.
- D-ID V4 expressives sentiment_id vs Sume's emotion string
D-ID picks a delivery with a sentiment_id preset; Sume TTS takes a free-text emotion string plus a 0.6-1.5 speed multiplier. How each one is set.
- Deepgram nova-3-pharma vs Sume STT: drug-name transcripts
Deepgram added nova-3-pharma for English drug names. Sume STT has one public model, sume/stt-1.0, so check each drug name against word timings.
- ElevenLabs character limits by model vs Sume TTS 20,000
ElevenLabs lists 5,000 characters for v3, 10,000 for v4 and 40,000 for Flash v2.5. Sume TTS 1.0 takes up to 20,000 characters in one request.
Written by Sume