Can ChatGPT create a video of me talking? Photo plus voice
Yes, with a connected tool: in developer mode ChatGPT can voice your script and lip sync a photo of you to it over MCP, then link the finished clip.

ChatGPT can create a video of you talking only through a connected tool. In developer mode it can call tools on a remote MCP server: a text-to-speech tool voices your script, and a lip sync tool animates a photo of you to that speech and returns a link to the clip. You supply the photo as a public link and, if you want your own voice, a voice you cloned beforehand.
ChatGPT's side comes from OpenAI's ChatGPT Developer mode guide. OpenAI says the Sora web and app experiences were discontinued on April 26, 2026, and it gave September 24, 2026 as the Sora API's discontinuation date; Can ChatGPT make videos? covers that. The tools here are Sume's tts_create and avatar-image-to-video_create (VEED Fabric 1.0), from MCP tools and gates, the Models overview and the Sume API reference, read on 2026-09-29. Sume has no official ChatGPT connector, and its basics page says hosted MCP still works but is not the primary path today.
How do I get ChatGPT to make a video of me?
OpenAI lists developer mode for Pro, Plus, Business, Enterprise, and Education accounts on the web. Add https://mcp.sume.com/mcp as a developer-mode app (How to add an MCP server to ChatGPT), and turn Write on at Sume's sign-in, since it is off by default and paid tools stay hidden without it. Then choose Developer mode from the Plus menu, select the app, and name the tools:
Use the Sume app only.
1. tts_create: transcript "Hi, I'm Sam. Welcome to week one
of the course.", voice.id "voi_…", language "en",
timestamps.words true. Show me the dry_run cost first.
2. avatar-image-to-video_create with payload:
image_url "https://example.com/me.jpg", the TTS audio_url,
and its length as duration_seconds.
3. jobs_wait, then jobs_result, and give me the clip link.Can I just attach my photo in the chat?
Not as the input. The lip sync request takes the photo as image_url, a public HTTPS still, and Sume's docs say hosted MCP cannot read files from your laptop. Put the photo at an HTTPS link you control, and keep in mind that anyone with a public link can open it. The audio must be on the Sume media host, which the TTS result already is.
Can it use my own voice?
Yes, if you clone it first in the Sume app, not in ChatGPT. The Voices screen there lets you clone a voice or invent one from a prompt, and TTS takes the result as voice.id, a Voices library id (voi_ plus 32 hex characters). Set language for a script that isn't English; left out, it defaults to English. How to clone yourself with AI covers the cloning step and why this path uses your photo directly rather than an avatar made from it.
What are the limits and costs?
Each clip is one TTS job and one lip sync job, each plus a 5.5% agent fee by default. The lip sync job counts audio seconds, rounded up. On ChatGPT's side, write actions require confirmation by default, so check each tool input before you approve it.
| Step | Tool | Limit | Price |
|---|---|---|---|
| Your script as speech | tts_create | Up to 20,000 characters; spaces and punctuation count | $0.0475 per 1,000 characters |
| Your photo, talking | avatar-image-to-video_create | One still; 1–300 seconds of Sume-hosted audio, at most 10 MB; 720p by default, or 480p | $0.1875 per audio second (720p) |
Whose face and voice can I use?
Your own, or someone's who has agreed. Sume's Terms say you represent that you have all rights and permissions for what you submit, including permission to use any person's likeness or voice, and they forbid using the service to mislead people about the origin or authenticity of generated media. The terms also call generated outputs synthetic and make you responsible for reviewing them before you publish, so watch the clip before you share it.
Sources
Related posts
More in Agents
- Can ChatGPT make videos better quality? Upscaling a clip
With an upscaling tool, yes: ChatGPT can send a clip's public URL to a video upscaler over MCP. For new clips, ask for a higher resolution.
- Can ChatGPT make long videos? Clip limits and how to join
Not in one clip: generated clips top out at 15 or 30 seconds. ChatGPT can join many clips over one audio track into a render of up to 30 minutes.
- Can ChatGPT make music? Getting a song or beat as audio
To get a song or a beat as an audio file, give ChatGPT a music tool: in developer mode it can call one over MCP and hand back a link to the track.
- Can ChatGPT make music videos? Song, clips, one render
With tools, yes: ChatGPT can have a track generated, a short clip made per section, and the clips cut over the song in one MP4 render.
Written by Sume