Can ChatGPT make long videos? Clip limits and how to join

Not in one clip: generated clips top out at 15 or 30 seconds. ChatGPT can join many clips over one audio track into a render of up to 30 minutes.

5 min readSume
All posts

Not in one generation: a single AI-generated clip is short, at most 30 seconds on the longest models in Sume's catalog and 15 seconds on the rest. A long video is several clips joined over one audio track, and ChatGPT can request both the clips and the join through tools, up to 30 minutes per render.

ChatGPT reaches those tools through developer mode, which gives it full MCP client support; Sora is not an option, since OpenAI discontinued the Sora web and app experiences on April 26, 2026 (OpenAI Help Center). The limits below come from Sume's Video Generation and Timeline 1.0 docs, read on 2026-09-29. Sume has no official ChatGPT connector; this is a remote MCP connection to https://mcp.sume.com/mcp, which Sume's basics page says still works but is not part of the primary path today.

How long can a video from ChatGPT be?

It depends on the tool ChatGPT calls. One generate_video clip is limited by its model: each model lists its supported_durations in whole seconds, and ChatGPT can read them with the video-router_models tool. AI video length limits by model has the per-model list. A joined render is limited by the timeline instead:

From Video Generation and Timeline 1.0, read 2026-09-29.
What ChatGPT asks forToolLongest output
One clip on seedance-2.5 or wan-3.0generate_video30 seconds
One clip on any other catalog modelgenerate_video15 seconds
Many clips joined over one audio tracktimeline_create1,800 seconds (30 minutes), 1–200 slots

How does ChatGPT join clips into a longer video?

With Sume's Timeline 1.0 render, which takes one audio spine plus ordered video slots and returns one MP4. Over MCP the sequence is timeline_create, then jobs_wait, then timeline_get. One way ChatGPT can build a long video:

  • ChatGPT writes the script and splits it into shots.
  • It voices the script with tts_create, or makes a track with music_create, to get the audio spine.
  • It generates one clip per shot with generate_video. Stills also work: they are held on screen for their slot.
  • It calls timeline_create with the spine and the clips, each with a start and duration. The first slot starts at 0.

What are the limits of a long video?

  • audio.duration_seconds sets the output length, from 1 to 1,800 seconds (30 minutes). With audio.mode: "silence" there is no spine file.
  • Every URL must already be your workspace's media.sume.com artifact or asset, such as a clip ChatGPT generated. An off-host URL is refused, so your own footage from elsewhere can't go on the timeline.
  • In current code the render's sound comes only from the audio spine and the optional soundtrack; each clip's own audio is dropped.
  • One jobs_wait call holds at most 55 seconds. On wait_slice_expired, ChatGPT should call jobs_wait again with the same ids, never resubmit the paid create. MCP tool call timeouts on long video jobs explains the loop.

How much does a long video cost?

Two parts. Each clip is billed per job from your Sume workspace balance at the provider's list price × 1.25, plus a 5.5% agent fee by default. The Timeline render is listed at $0.10 per output minute on API pricing. The reserve is ceil(audio.duration_seconds / 60) minutes, so a 10-minute render reserves 10 minutes; in current code the charge is then captured at the render's own compute cost, never above that reservation.

Ask ChatGPT to run each paid tool with dry_run=true first: it previews the cost without submitting the job. max_spend_usd caps a call when you send it, and ChatGPT asks you to confirm write actions by default.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume