Can ChatGPT make UGC videos? Script, voice, and lip sync

Not by itself. With a video tool connected in developer mode, ChatGPT can script, voice, and lip-sync a UGC-style clip, confirming paid calls first.

5 min readSume
All posts

Not by itself, but it can run the whole job through a video tool. A UGC-style video, a person talking to camera about a product, needs a face, a voice, and lip sync. In developer mode ChatGPT can write the script and then call tools that make a presenter still, voice the script, and lip-sync the still to that audio, asking you to confirm paid calls by default.

ChatGPT's side comes from OpenAI's ChatGPT Developer mode guide. The tools are on Sume's hosted MCP server, per MCP tools and gates, the models overview, and the Sume API reference, read on 2026-09-29. Sume has no official ChatGPT connector: this is a remote MCP connection, and Sume's basics page says hosted MCP still works but is not the primary path today. Make AI UGC videos with Claude is the same pipeline from Claude.

Which tools does ChatGPT call to make a UGC video?

Four steps, three of them paid. In Sume's current tool descriptions, video models don't lip-sync to a voiceover, so speech goes through text to speech and then a lip-sync tool:

  • Every paid call adds a 5.5% agent fee by default and bills your Sume wallet.
  • A wordless UGC beat, hands and product with no speech, can skip the voice and use generate_video from an approved still instead.
From MCP tools and gates, the models overview, the Sume API reference, and API pricing, read 2026-09-29. Tool rules are current code where noted.
StepToolWhat it takes and bills
ScriptNone: ChatGPT writes itFree; keep it short
Presenter stillgenerate_imageA prompt; priced per model, previewed with dry_run
Voicetts_createUp to 20,000 characters and a Sume voice; $0.0475 per 1,000 characters
Lip syncavatar-image-to-video_create (VEED Fabric 1.0)One still plus Sume-hosted audio under 10 MB, 1–300 seconds; $0.1875 per audio second (720p)

Where does the voice come from?

From your Sume workspace, not from ChatGPT. tts_create needs a voice selector: a Sume voice id you already hold, or an avatar whose voice status is ready, which the API reference calls the public route to a TTS voice. The audio it returns is Sume-hosted, which is what the lip-sync step requires; a voice file from elsewhere is refused. ChatGPT text to speech covers voice choice.

What do I ask ChatGPT?

Turn on developer mode (OpenAI lists it for Pro, Plus, Business, Enterprise, and Education accounts on the web), add https://mcp.sume.com/mcp as an app, and turn Write on at Sume's consent page, where it is off by default. Then choose Developer mode from the Plus menu and be explicit:

Use only the Sume app. Write a 15-second UGC script for
https://example.com/products/serum.jpg: a creator at a bathroom
counter says why she switched. Show me the script first.
Then generate_image a posed presenter still, tts_create the
script with my avatar's voice, and avatar-image-to-video_create
the still with that audio. dry_run each paid call first.

Will ChatGPT spend money without asking?

Not by default. OpenAI says write actions require confirmation and that tools without the readOnlyHint annotation are treated as write actions. You can remember an approve choice for a tool for the rest of a conversation, so only do that once the dry-run prices look right. On Sume's side, dry_run=true previews cost without submitting, max_spend_usd caps a call when you send it, and each paid call carries an idempotency_key.

Each tool returns a job id. ChatGPT waits with jobs_wait, which holds at most 55 seconds per call, then reads jobs_result. Finished files are Sume-hosted artifacts under media.sume.com. Watch the clip before you use it: the docs make no promise about how the product or the face will look, and nothing here says how an ad will perform.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume