Can ChatGPT make videos with sound? Audio and voice
Yes, through a video tool whose model generates audio. A voice that speaks your script is a second step: text to speech, then lip sync.

Yes, if ChatGPT calls a video tool whose model generates audio: some video models return a clip with its own soundtrack, others return silent video. A person speaking your exact script is a different job, done in two steps: text to speech for the voice, then a lip-sync model that animates a still to that audio.
Sora is not the route: OpenAI says the Sora web and app experiences were discontinued on April 26, 2026 (OpenAI Help Center). ChatGPT reaches a video tool through developer mode, which provides full MCP client support; How to add an MCP server to ChatGPT covers setup. The Sume facts below come from Video Generation, the Models page, and MCP tools and gates, read on 2026-09-29. Hosted MCP still works but is not part of Sume's primary path today (basics).
Will the video have sound by default?
On Sume's generate_video tool, yes when the model supports it. The request field generate_audio defaults to the model's audio capability, and in current code the tool's description tells ChatGPT to prefer generate_audio: true, noting that leaving it out is also audio-on. It sets false only when you ask for a silent clip.
Not every model makes audio. Each model's catalog row carries a generate_audio flag that says whether it can generate an audio track, and ChatGPT can read it with the video-router_models tool before it submits. AI video with sound lists which models make sound.
| The sound you want | Tool ChatGPT calls | What to know |
|---|---|---|
| Sound generated with the clip | generate_video | Only on models whose catalog generate_audio flag is true |
| A voice speaking your script | tts_create, then avatar-image-to-video_create | A still plus the voice audio; video models don't lip-sync |
| A silent clip | generate_video with generate_audio: false | Only on models that allow it; in current code a model that always makes audio refuses false |
Can ChatGPT make a video where someone speaks my script?
Yes, but not with a video model alone. Sume's docs say video models do not lip-sync to generated TTS or to a later voice-over, so laying narration under a generated face will not match the lips. The documented route is:
tts_createvoices the script with a Sume voice and returns an audio file. ChatGPT text to speech covers voices and prompts.avatar-image-to-video_create(VEED Fabric 1.0,veed/fabric-1.0) turns one still plus that audio into a talking clip.- The audio must be on the Sume media host and at most 10 MB, and
duration_secondsruns from 1 to 300. - Send exactly one visual source: a public HTTPS
image_url, or a ready avatar by id or handle. - Output is
480por720p, with720pthe default.
How much does a video with sound cost?
A generated clip is billed per job from your Sume workspace balance at the provider's list price × 1.25; ask ChatGPT to call the tool with dry_run=true first, which previews the cost without submitting. A speaking clip has two meters on API pricing: text to speech at $0.0475 per 1,000 characters, and VEED Fabric 1.0 at $0.1875 per audio second (720p). Each rate carries a 5.5% agent fee on top by default.
Every paid call needs an idempotency_key, and max_spend_usd caps a call only when you send it. ChatGPT asks you to confirm write actions by default before it runs them.
What doesn't this do?
- It does not read a video or audio file from your computer: hosted MCP cannot read files from your laptop.
- It does not lip-sync a text-to-video clip to your voice-over; that is the Fabric step above.
- It does not promise how the generated sound will sound. Listen to the clip before you use it.
Sources
Related posts
More in Agents
- Can Claude Cowork make videos? Yes, through a connector
Not by itself: Cowork's models output text. Add a remote MCP video tool as a custom connector and a Cowork task can render clips and images as links.
- Can Claude create product images? From your photo, via a tool
Claude can't create images, but it can send your product photo to an image tool as a reference and describe the new scene. How it works and its limits.
- Can Claude generate high resolution images? 4K via a tool
Claude can't make images itself. Through an image tool it can ask for a 4K tier on models that list one, then upscale a still up to 4x.
- Can Claude generate multiple images at once? Yes, via a tool
Claude can't draw, but through an image tool one call can ask for several images, up to 4 per model today, and Claude can send one call per scene at once.
Written by Sume