Colossyan intrinsicDurationTrackReference vs Sume scene duration
Colossyan can size a scene to the actor's speech. In Sume avatar videos you set voice.duration per scene, silence beats need it, and the total stays 4-60 s.

Colossyan lets a scene take the length of its actor's speech by pointing intrinsicDurationTrackReference at a track's referenceId. In Sume's avatar-video docs the multi-scene plan carries its own duration on each voice object, and a silence beat requires one, so you declare scene timing instead of deriving it.
Colossyan facts are from its Timing page; Sume facts from Generate avatar video, both read 2026-10-01.
How does Colossyan time a scene?
Its page lists ways to set scene duration. One is an explicit duration in milliseconds. Another is intrinsicDurationTrackReference, which Colossyan resolves by working out the referenced track's length in isolation first. Whichever you use, it says everything longer than the scene is cut out.
What does a Sume scene declare?
Each item in video_inputs has a voice. A text voice carries script or input_text plus a duration (the docs example uses "duration": 3). A silence voice, voice.type: "silence", is a non-speaking beat: duration is required and script / input_text are not allowed.
| Question | Colossyan Timing page | Sume avatar-video docs |
|---|---|---|
| Explicit length | duration in ms on the scene | voice.duration on the scene |
| Length from speech | intrinsicDurationTrackReference | Not documented; declare it |
| Silent gap | Not covered on this page | voice.type: "silence" with required duration |
| Whole video | Not stated on this page | Estimated 4-60 seconds inclusive |
What should I do about a script that runs long?
Sume accepts scripts and multi-scene plans only when the estimated duration lands in 4-60 seconds, and the docs say to shorten longer scripts or split them into multiple jobs. Because the plan is declared, add up your duration values before you submit rather than relying on the speech length.
To check the first frames of a plan before a full render, see first-frame previews for avatar videos.
Is there a way to size scenes from speech on Sume?
The avatar-video page does not document one. If you need scene lengths that follow audio, measure the audio yourself and write the result into voice.duration.
Sources
Related posts
More in Developers
- Colossyan dynamicVariables vs Sume: fill placeholders yourself
Colossyan swaps {name} placeholders from dynamicVariables. Sume's avatar endpoint has no such field: fill the script yourself, one idempotent job per recipient.
- Creatify custom avatar lipsync_input and consent video, explained
Creatify custom avatars are built from a lipsync_input MP4 plus a consent video. On Sume, a photo input with an image_url creates a reusable avatar handle.
- Cursor custom mode with a pinned skill: install the Sume skill
Cursor lets any skill be a custom mode pinned in the chat. Install the Sume skill with sume skills install; the mode cannot waive idempotency_key or MCP scope.
- Cursor /goal with paid Sume calls: what actually caps spend
Cursor /goal keeps an agent on one objective. The goal text is no budget: bound Sume spend with generation_spend_cap_usd, max_paid_calls and dry_run.
Written by Sume