ChatGPT script to video: voice it, then cut to each line
ChatGPT writes the script; tools turn it into video. Voice it word for word, make one visual per line, and cut each to its sentence on a timeline.

ChatGPT can write the script, but turning that exact script into a video takes tools. In developer mode ChatGPT can voice the script word for word with a text-to-speech tool, make one clip or still per line, and cut each visual to its sentence on a timeline that renders one MP4. Without those tools, what you get from the chat is the script.
ChatGPT's side comes from OpenAI's ChatGPT Developer mode guide. The tools are on Sume's hosted MCP server, per MCP tools and gates, Timeline 1.0, and the Sume API reference, read on 2026-09-29. Sume has no official ChatGPT connector: this is a remote MCP connection, and Sume's basics page says hosted MCP still works but is not the primary path today. The general method, without ChatGPT, is in Script to video AI.
How does ChatGPT turn a script into a video?
In three tool steps after the script is final. Once developer mode is on (OpenAI lists it for Pro, Plus, Business, Enterprise, and Education accounts on the web) and Sume's app has Write turned on, ChatGPT calls:
tts_createwith the script astranscript,timestamps.words: true, andsegmentation.mode: "sentence". The finished job then carrieswords[]with start and end seconds and gapless sentencesegments[], so each sentence has a start and an end.generate_videoorgenerate_imageonce per line, for a clip or a still that shows it.timeline_createwith the voiceover as the audio spine and onevideo[]slot per line, each starting where its sentence starts. Thenjobs_waitandtimeline_getfor the MP4.
Can ChatGPT make all the per-line clips in one call?
Yes, with script_run. It runs a short JavaScript program on Sume's side for turns that need three or more calls of the same shape, such as one generate_image per scene. Each call inside keeps the same gates, and each paid create still needs its own idempotency_key. The run is bounded by timeout_seconds (5–55), max_calls, and max_paid_calls, and it returns the child jobs[] to wait on, so ChatGPT follows it with jobs_wait rather than expecting finished clips.
Will ChatGPT keep every word of my script?
Only if the words it sends are yours. tts_create speaks the transcript it receives, and ChatGPT fills that field, so freeze the script first and read the tool input before you approve. OpenAI says the full JSON of each tool call's input is available in the chat and asks you to review write actions carefully. The checks worth making at each step:
- The voiceover, clips, and stills from the earlier steps are Sume-hosted, so they qualify. A file from your computer does not: hosted MCP cannot read your disk.
- In current code a Timeline render takes sound only from the voiceover spine and an optional soundtrack; each clip's own audio is dropped. Script to video AI covers the length and slot limits.
| Tool call | Check in the input | Why |
|---|---|---|
tts_create | transcript matches your script word for word | Up to 20,000 characters are voiced as sent |
tts_create | timestamps.words: true and segmentation.mode: "sentence" | Sentence segments[] are the cut points |
| Clips and stills | dry_run: true on the first call | Previews the cost without submitting |
timeline_create | video[0].start is 0; each later start is a sentence start | Declared starts are authoritative |
timeline_create | Every URL is a media.sume.com file | Timeline takes only this workspace's Sume-hosted media |
How much does a script-to-video run cost?
The voiceover bills $0.0475 per 1,000 characters; the timeline render bills $0.10 per output minute, reserved per started minute; clips and stills are priced per model. Each paid call adds a 5.5% agent fee by default. OpenAI says write actions require confirmation by default, and Sume's dry_run=true previews a paid call's cost without submitting it, so ask ChatGPT to dry-run the voiceover and the clips and show you the prices first. One jobs_wait holds at most 55 seconds; MCP tool call timeouts on long video jobs covers the wait loop.
Sources
Related posts
More in Agents
- Claude Opus 5.5 and paid MCP tools: preview and cap spend with dry_run
Give Claude Opus 5.5 a paid video tool without a surprise bill: dry_run previews cost, max_spend_usd caps a call, idempotency_key stops double submits.
- Batched tool calls and video jobs: one jobs_wait for up to 20 clips
Anthropic says Claude Sonnet 5.5 batches tool calls more than Sonnet 5. On Sume MCP, submit several clips in parallel, then wait on all with one jobs_wait.
- Give my agent video generation: MCP, REST or Agent Completions?
Three ways to give an agent video generation on Sume: the hosted MCP server, the /v1/videos REST API as your own tool, or Agent Completions for the whole task.
- GPT-6 Astra async tool calls: what a slow video job means for agents
OpenAI's GPT-6 guide describes async tool calling with async: true and call_id. How that maps to a video job that takes minutes, and where Sume's job ids fit.
Written by Sume