Can ChatGPT make product videos from a product photo?

Yes, with a video tool: in developer mode ChatGPT can send your product photo to a video model as the first frame and return a link to the clip.

5 min readSume
All posts

Yes, if you give ChatGPT a video tool. In developer mode, ChatGPT can pass your product photo to a video model as the clip's first frame, write a prompt about motion and camera, and hand back a link to the finished product video. OpenAI discontinued the Sora web and app experiences on April 26, 2026, so the rendering happens in the tool you connect.

OpenAI's side comes from its ChatGPT Developer mode guide and Sora discontinuation article. The tool here is generate_video on Sume's hosted MCP server, per Video Generation and MCP tools and gates, read on 2026-09-29. Sume has no official connector for ChatGPT: this is a remote MCP connection, and Sume's basics page says hosted MCP still works but is not the primary path today. Can ChatGPT make videos? covers the setup.

How does ChatGPT turn a product photo into a video?

The video model does the rendering; ChatGPT fills in the call. The docs describe two ways to hand a model an image, and they behave differently:

  • frame_images with frame_type first_frame is image-to-video: the clip starts on your photo.
  • input_references only guide the look; Product logo warping in image-to-video compares the two.
  • The photo must be reachable over public HTTPS. In current code a reference that answers 404 fails with input_media_unreachable, so paste a direct image link, not a product page.

What do I type in ChatGPT?

Choose Developer mode from the Plus menu, select the Sume app, and name the tool and the frame. Keep the prompt about motion and camera, not about redrawing the product:

Use the Sume app's generate_video tool.
First frame: https://example.com/products/bottle.jpg
Prompt: slow push-in on the bottle on a marble counter,
soft morning light moving across the label, camera only.
9:16, 5 seconds. Run a dry_run first and show me the cost.

Will the label and logo stay exact?

The docs make no promise that they will, so check every frame before you publish. A first frame pins how the clip opens; it doesn't lock the product through the motion. The post linked above covers what to do when a label drifts.

From Video Generation, read 2026-09-29. Each model advertises its own lists; ask ChatGPT to read them with video-router_models first.
SettingWhat the docs say
Aspect ratioPer model in supported_aspect_ratios; the sample lists 21:9, 16:9, 4:3, 1:1, 3:4, 9:16
DurationPer model: seedance-2.5 and wan-3.0 go up to 30 seconds; the docs cap the other models at 15 seconds or less
Photo inputA first_frame in frame_images, or input_references
WaitTypically 30 seconds to several minutes, depending on the model and parameters

What does a product video cost, and does ChatGPT ask first?

generate_video is a paid tool billed to your Sume wallet, and the price depends on the model, resolution, and length. As a range, a 10-second 720p clip from a prompt or a photo costs $0.13 to $5.82 across Sume's catalog, plus a 5.5% agent fee by default; How much does an AI video cost? breaks that down. dry_run=true previews the cost without submitting the job, and max_spend_usd caps a call when you send it.

OpenAI says write actions require confirmation by default, so ChatGPT asks before each paid call unless you choose to remember your approval for that tool in the conversation. The call returns a job id; ChatGPT waits with jobs_wait, which holds at most 55 seconds per call, then reads the clip's media.sume.com link from jobs_result. For a whole catalog of SKUs, Sume's docs point to HTTP calls from your backend rather than hosted MCP: see Product photo to video API.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume