Can ChatGPT create a video of me talking? Photo plus voice

Yes, with a connected tool: in developer mode ChatGPT can voice your script and lip sync a photo of you to it over MCP, then link the finished clip.

5 min readSume
All posts

ChatGPT can create a video of you talking only through a connected tool. In developer mode it can call tools on a remote MCP server: a text-to-speech tool voices your script, and a lip sync tool animates a photo of you to that speech and returns a link to the clip. You supply the photo as a public link and, if you want your own voice, a voice you cloned beforehand.

ChatGPT's side comes from OpenAI's ChatGPT Developer mode guide. OpenAI says the Sora web and app experiences were discontinued on April 26, 2026, and it gave September 24, 2026 as the Sora API's discontinuation date; Can ChatGPT make videos? covers that. The tools here are Sume's tts_create and avatar-image-to-video_create (VEED Fabric 1.0), from MCP tools and gates, the Models overview and the Sume API reference, read on 2026-09-29. Sume has no official ChatGPT connector, and its basics page says hosted MCP still works but is not the primary path today.

How do I get ChatGPT to make a video of me?

OpenAI lists developer mode for Pro, Plus, Business, Enterprise, and Education accounts on the web. Add https://mcp.sume.com/mcp as a developer-mode app (How to add an MCP server to ChatGPT), and turn Write on at Sume's sign-in, since it is off by default and paid tools stay hidden without it. Then choose Developer mode from the Plus menu, select the app, and name the tools:

Use the Sume app only.
1. tts_create: transcript "Hi, I'm Sam. Welcome to week one
   of the course.", voice.id "voi_…", language "en",
   timestamps.words true. Show me the dry_run cost first.
2. avatar-image-to-video_create with payload:
   image_url "https://example.com/me.jpg", the TTS audio_url,
   and its length as duration_seconds.
3. jobs_wait, then jobs_result, and give me the clip link.

Can I just attach my photo in the chat?

Not as the input. The lip sync request takes the photo as image_url, a public HTTPS still, and Sume's docs say hosted MCP cannot read files from your laptop. Put the photo at an HTTPS link you control, and keep in mind that anyone with a public link can open it. The audio must be on the Sume media host, which the TTS result already is.

Can it use my own voice?

Yes, if you clone it first in the Sume app, not in ChatGPT. The Voices screen there lets you clone a voice or invent one from a prompt, and TTS takes the result as voice.id, a Voices library id (voi_ plus 32 hex characters). Set language for a script that isn't English; left out, it defaults to English. How to clone yourself with AI covers the cloning step and why this path uses your photo directly rather than an avatar made from it.

What are the limits and costs?

Each clip is one TTS job and one lip sync job, each plus a 5.5% agent fee by default. The lip sync job counts audio seconds, rounded up. On ChatGPT's side, write actions require confirmation by default, so check each tool input before you approve it.

From the TTS and Fabric schemas in the Sume API reference, the Models overview and the code behind API pricing, read 2026-09-29.
StepToolLimitPrice
Your script as speechtts_createUp to 20,000 characters; spaces and punctuation count$0.0475 per 1,000 characters
Your photo, talkingavatar-image-to-video_createOne still; 1–300 seconds of Sume-hosted audio, at most 10 MB; 720p by default, or 480p$0.1875 per audio second (720p)

Whose face and voice can I use?

Your own, or someone's who has agreed. Sume's Terms say you represent that you have all rights and permissions for what you submit, including permission to use any person's likeness or voice, and they forbid using the service to mislead people about the origin or authenticity of generated media. The terms also call generated outputs synthetic and make you responsible for reviewing them before you publish, so watch the clip before you share it.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume