ChatGPT script to video: voice it, then cut to each line

ChatGPT writes the script; tools turn it into video. Voice it word for word, make one visual per line, and cut each to its sentence on a timeline.

5 min readSume
All posts

ChatGPT can write the script, but turning that exact script into a video takes tools. In developer mode ChatGPT can voice the script word for word with a text-to-speech tool, make one clip or still per line, and cut each visual to its sentence on a timeline that renders one MP4. Without those tools, what you get from the chat is the script.

ChatGPT's side comes from OpenAI's ChatGPT Developer mode guide. The tools are on Sume's hosted MCP server, per MCP tools and gates, Timeline 1.0, and the Sume API reference, read on 2026-09-29. Sume has no official ChatGPT connector: this is a remote MCP connection, and Sume's basics page says hosted MCP still works but is not the primary path today. The general method, without ChatGPT, is in Script to video AI.

How does ChatGPT turn a script into a video?

In three tool steps after the script is final. Once developer mode is on (OpenAI lists it for Pro, Plus, Business, Enterprise, and Education accounts on the web) and Sume's app has Write turned on, ChatGPT calls:

  • tts_create with the script as transcript, timestamps.words: true, and segmentation.mode: "sentence". The finished job then carries words[] with start and end seconds and gapless sentence segments[], so each sentence has a start and an end.
  • generate_video or generate_image once per line, for a clip or a still that shows it.
  • timeline_create with the voiceover as the audio spine and one video[] slot per line, each starting where its sentence starts. Then jobs_wait and timeline_get for the MP4.

Can ChatGPT make all the per-line clips in one call?

Yes, with script_run. It runs a short JavaScript program on Sume's side for turns that need three or more calls of the same shape, such as one generate_image per scene. Each call inside keeps the same gates, and each paid create still needs its own idempotency_key. The run is bounded by timeout_seconds (5–55), max_calls, and max_paid_calls, and it returns the child jobs[] to wait on, so ChatGPT follows it with jobs_wait rather than expecting finished clips.

Will ChatGPT keep every word of my script?

Only if the words it sends are yours. tts_create speaks the transcript it receives, and ChatGPT fills that field, so freeze the script first and read the tool input before you approve. OpenAI says the full JSON of each tool call's input is available in the chat and asks you to review write actions carefully. The checks worth making at each step:

  • The voiceover, clips, and stills from the earlier steps are Sume-hosted, so they qualify. A file from your computer does not: hosted MCP cannot read your disk.
  • In current code a Timeline render takes sound only from the voiceover spine and an optional soundtrack; each clip's own audio is dropped. Script to video AI covers the length and slot limits.
From Timeline 1.0, MCP tools and gates, the Sume API reference, and OpenAI's Developer mode guide, read 2026-09-29.
Tool callCheck in the inputWhy
tts_createtranscript matches your script word for wordUp to 20,000 characters are voiced as sent
tts_createtimestamps.words: true and segmentation.mode: "sentence"Sentence segments[] are the cut points
Clips and stillsdry_run: true on the first callPreviews the cost without submitting
timeline_createvideo[0].start is 0; each later start is a sentence startDeclared starts are authoritative
timeline_createEvery URL is a media.sume.com fileTimeline takes only this workspace's Sume-hosted media

How much does a script-to-video run cost?

The voiceover bills $0.0475 per 1,000 characters; the timeline render bills $0.10 per output minute, reserved per started minute; clips and stills are priced per model. Each paid call adds a 5.5% agent fee by default. OpenAI says write actions require confirmation by default, and Sume's dry_run=true previews a paid call's cost without submitting it, so ask ChatGPT to dry-run the voiceover and the clips and show you the prices first. One jobs_wait holds at most 55 seconds; MCP tool call timeouts on long video jobs covers the wait loop.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume