AI outfit change video: change clothes in a clip with AI
Change someone's outfit in an existing video with an AI video-to-video edit: name the garment, what it becomes, and what must stay the same.

To change the clothes someone wears in an existing video, use an AI video-to-video edit: send the clip with a prompt that names the garment, what it should become, and what must not change, such as “Change her white T-shirt to a black leather jacket. Keep her face, hair, pose, the background and the camera move the same.” The model re-renders the clip from that prompt, without a reshoot; check that it kept what you asked it to keep.
On Sume, that edit is gemini-omni-flash-1.1 on the Video Router. The facts below come from the Video Router, Media inputs and Timeline 1.0 docs and Sume's pricing code, read on 2026-09-29.
How do I change an outfit in a video with AI?
Send POST /v1/video-router/generate with model: "gemini-omni-flash-1.1", the clip's public HTTPS URL in video_url, and the outfit change in prompt. The request shape is the docs' own edit example, whose prompt replaces one object and ends “Keep everything else the same.” Edit a video with a prompt covers polling the job and fetching the result.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: outfit-change-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "Change her white T-shirt to a black leather jacket. Keep her face, hair, pose, the background and the camera move the same.",
"video_url": "https://example.com/walk.mp4",
"resolution": "720p",
"mode": "async"
}'What does the edit accept?
The clip must sit at a fetchable public HTTPS URL; localhost, private-network and signed or private URLs are rejected. Edit mode takes no aspect_ratio or duration, resolution defaults to 720p, and native audio is always on. Edit a video with a prompt lists every field edit mode accepts and rejects.
Can I put a specific garment from a product photo on the person?
Not in the same edit. video_url cannot be combined with an image or reference URL, so the new outfit is described in words only: color, material, cut, details such as buttons or a logo. A product photo can drive a new clip generated from photos instead of your footage; Virtual try-on for Shopify and AI fashion video generator API cover that route.
What should I check in the result?
The docs don't promise what else an edit changes, or whether it keeps the source's length, framing or sound, and they make no promise about keeping a person's likeness. Watch the whole clip before you use it:
- The edges of the garment, where it meets skin, hair and the background, from frame to frame.
- Hands and arms that cross the garment.
- The face, the background and the camera move, against the source.
- The sound, since the model's native audio is always on.
What about the outfit transition trend?
The short-video outfit change trend, where a cut or a hand wipe reveals a new look, is two clips of the same person joined on a matching pose, not an edit of one clip. With both clips as Sume files, a Timeline 1.0 render joins them with a hard cut or a short wipeleft or wiperight transition of up to 1 second. In current code the render takes sound only from its own audio track and optional soundtrack, so add the music there; each clip's own sound is dropped.
How much does an AI outfit change cost?
The edit is billed per second of output by resolution, at the provider's list price × 1.25, plus a 5.5% agent fee by default. Each new outfit is a new edit, billed again.
| Resolution | Per output second |
|---|---|
720p (default) | $0.125 |
1080p | $0.1875 |
Sources
Related posts
More in Models
- Claude Sonnet 5.5 or Opus 5.5 for an agent that calls a video tool?
Sonnet 5.5 and Opus 5.5 differ in price and release date on Anthropic's pages. Which facts matter when the tool is Sume's video generation, and what stays same?
- How to colorize a black and white video with AI
Colorize a black and white video with an AI video-to-video edit: send the clip and a prompt naming the colors. The model invents them.
- FLUX.2 flex text rendering: prompt tips and the Sume model id
Black Forest Labs positions FLUX.2 [flex] for text and fine detail. How its typography guidance reads, and how to call black-forest-labs/flux.2-flex on Sume.
- FLUX.2 pro vs Nano Banana Pro API: request specs side by side
Request parameters that differ between black-forest-labs/flux.2-pro and google/nano-banana-pro on Sume: resolution tiers, aspect ratios, references and price.
Written by Sume