LTX /v1/extend limits, and chaining clips with Sume frame_images
LTX's /v1/extend is 1080p only, adds 2-20 seconds, needs 73+ input frames. Sume has no extend call here: chain clips with frame_images, then join in Timeline.

LTX's POST /v1/extend continues an existing video, but only at 1080p, adding 2 to 20 seconds, with at least 73 input frames and a 505-frame total. Sume does not expose an extend call on its Omni video id. To get a longer video, generate a next clip from the last frame of the previous one with frame_images, then join the clips with Timeline.
LTX's limits are from its post How to make longer AI videos, read 2026-10-01. Sume's side is from Video generation and Timeline 1.0.
What are the /v1/extend limits?
| Constraint | Value in the post |
|---|---|
| Resolution | 1080p only |
| New material | duration from 2 to 20 seconds |
| Input | At least 73 frames |
| Total | Capped at 505 frames including context |
| Model | ltx-2-3-pro on the API |
| Audio | Generated automatically when the input video has audio |
Does Sume have an extend call?
Not on the Omni model. The Video Router docs say previous_interaction_id / extend is not exposed for gemini-omni-flash-1.1, because it is not on the provider's published input schema. The docs read here do not describe an extend call on other rows, so the supported route is chaining.
How do I chain clips on Sume?
Pull the last frame of clip one (see get the last frame to chain clips), then send it as a frame_images entry with frame_type: "first_frame" on the next request. The docs note that if both frame_images and input_references are provided, frame_images takes precedence and the request is treated as image-to-video. Check supported_frame_images in the catalog first, since not every model accepts a frame.
Chaining restarts generation from a still, so motion and audio are not carried over the way an extend carries them. Expect to review each seam.
How do I join the clips?
Timeline 1.0 assembles one audio spine plus ordered video slots into a long-form MP4, per its docs. Import the generated files as Sume-hosted media first, then list them in video[]. The longer walkthrough is extend an AI video past 30 seconds; the Veo equivalent is in Veo 3.1 extend vs chaining.
Sources
Related posts
More in Developers
- LTX keyframe interpolation is open-source only; Sume takes last_frame
LTX's KeyframeInterpolationPipeline has no API endpoint. On Sume, frame_images accepts first_frame and last_frame on models that report support in the catalog.
- LTX retake vs Sume: fix one section of a video
LTX's POST /v1/retake regenerates one time region of a video. Sume's video edit takes a whole clip, so trim the bad section, edit it, and rejoin it.
- LTX frame counts must be 8k+1; Sume durations are whole seconds
LTX frame counts must satisfy (F-1) % 8 == 0, so 30 frames is invalid. On Sume you send whole seconds and read the allowed set from the catalog, not frames.
- Luma API 429 requests per minute: sliding window vs Sume
Luma counts requests in a sliding 60-second window and returns 429 if RPM or concurrent jobs fails. Sume returns 429 rate_limited: back off, reuse the key.
Written by Sume