Can ChatGPT make UGC videos? Script, voice, and lip sync
Not by itself. With a video tool connected in developer mode, ChatGPT can script, voice, and lip-sync a UGC-style clip, confirming paid calls first.

Not by itself, but it can run the whole job through a video tool. A UGC-style video, a person talking to camera about a product, needs a face, a voice, and lip sync. In developer mode ChatGPT can write the script and then call tools that make a presenter still, voice the script, and lip-sync the still to that audio, asking you to confirm paid calls by default.
ChatGPT's side comes from OpenAI's ChatGPT Developer mode guide. The tools are on Sume's hosted MCP server, per MCP tools and gates, the models overview, and the Sume API reference, read on 2026-09-29. Sume has no official ChatGPT connector: this is a remote MCP connection, and Sume's basics page says hosted MCP still works but is not the primary path today. Make AI UGC videos with Claude is the same pipeline from Claude.
Which tools does ChatGPT call to make a UGC video?
Four steps, three of them paid. In Sume's current tool descriptions, video models don't lip-sync to a voiceover, so speech goes through text to speech and then a lip-sync tool:
- Every paid call adds a 5.5% agent fee by default and bills your Sume wallet.
- A wordless UGC beat, hands and product with no speech, can skip the voice and use
generate_videofrom an approved still instead.
| Step | Tool | What it takes and bills |
|---|---|---|
| Script | None: ChatGPT writes it | Free; keep it short |
| Presenter still | generate_image | A prompt; priced per model, previewed with dry_run |
| Voice | tts_create | Up to 20,000 characters and a Sume voice; $0.0475 per 1,000 characters |
| Lip sync | avatar-image-to-video_create (VEED Fabric 1.0) | One still plus Sume-hosted audio under 10 MB, 1–300 seconds; $0.1875 per audio second (720p) |
Where does the voice come from?
From your Sume workspace, not from ChatGPT. tts_create needs a voice selector: a Sume voice id you already hold, or an avatar whose voice status is ready, which the API reference calls the public route to a TTS voice. The audio it returns is Sume-hosted, which is what the lip-sync step requires; a voice file from elsewhere is refused. ChatGPT text to speech covers voice choice.
What do I ask ChatGPT?
Turn on developer mode (OpenAI lists it for Pro, Plus, Business, Enterprise, and Education accounts on the web), add https://mcp.sume.com/mcp as an app, and turn Write on at Sume's consent page, where it is off by default. Then choose Developer mode from the Plus menu and be explicit:
Use only the Sume app. Write a 15-second UGC script for
https://example.com/products/serum.jpg: a creator at a bathroom
counter says why she switched. Show me the script first.
Then generate_image a posed presenter still, tts_create the
script with my avatar's voice, and avatar-image-to-video_create
the still with that audio. dry_run each paid call first.Will ChatGPT spend money without asking?
Not by default. OpenAI says write actions require confirmation and that tools without the readOnlyHint annotation are treated as write actions. You can remember an approve choice for a tool for the rest of a conversation, so only do that once the dry-run prices look right. On Sume's side, dry_run=true previews cost without submitting, max_spend_usd caps a call when you send it, and each paid call carries an idempotency_key.
Each tool returns a job id. ChatGPT waits with jobs_wait, which holds at most 55 seconds per call, then reads jobs_result. Finished files are Sume-hosted artifacts under media.sume.com. Watch the clip before you use it: the docs make no promise about how the product or the face will look, and nothing here says how an ad will perform.
Sources
Related posts
More in Agents
- Can ChatGPT make videos for free? Plans and per-clip costs
Not for free. Free ChatGPT accounts can't turn on developer mode, and the video tool it calls needs its own paid plan and bills each clip.
- Can ChatGPT make videos from photos? Image to video now
Not with Sora, which OpenAI discontinued. ChatGPT can send a photo's public URL to an image-to-video tool over MCP and get a clip back.
- Can ChatGPT make videos with sound? Audio and voice
Yes, through a video tool whose model generates audio. A voice that speaks your script is a second step: text to speech, then lip sync.
- Can Claude Cowork make videos? Yes, through a connector
Not by itself: Cowork's models output text. Add a remote MCP video tool as a custom connector and a Cowork task can render clips and images as links.
Written by Sume