FLUX Virtual Try-On v2: 4 MP inputs and a Sume reference edit
BFL's vto-v2 keeps inputs up to 4 MP as-is. Sume has no try-on endpoint in these docs; a garment swap is a reference edit with public HTTPS images.
FLUX Virtual Try-On v2 (vto-v2) uses model and garment images up to 4 megapixels as-is, up from 2 in v1, with the same request format. The Sume images docs do not list a try-on endpoint; the closest path is a reference edit with input_references on an image model such as black-forest-labs/flux.2-pro.
BFL facts are from its try-on docs and release notes; Sume facts from Image generation, read 2026-10-01.
What changed in vto-v2?
BFL's release notes call vto-v2 the recommended version, focused on keeping the person's identity intact with gains in garment fidelity. Input images up to 4 MP are used as-is and the output follows the model image's resolution. The request and response format match vto-v1, so migration is a change of endpoint path (POST /v1/flux-tools/vto-v2).
What are the input limits?
Inputs above 4 MP are not rendered at full resolution: they are downscaled to about 1 MP while keeping aspect ratio. BFL suggests keeping both images near 1 MP for the best quality and latency balance, noting larger inputs raise latency.
| Version | Input handling |
|---|---|
vto-v2 | Up to 4 MP used as-is; above that, downscaled to about 1 MP |
vto-v1 | Up to 2 MP; larger inputs downscale to about 1 MP |
What does Sume accept for a garment swap?
Sume's catalog includes black-forest-labs/flux.2-pro. References go in input_references as an array of image URLs, which must be public HTTPS; localhost, private-network and non-HTTPS URLs are rejected. Whether a model takes references depends on its catalog descriptor. Put the person photo and the garment in as references and describe the swap in the prompt. This is a general edit, not BFL's try-on tool, so its output has no stated link to vto-v2 behavior.
How do I keep the person's framing?
On edit and image-to-image calls the docs recommend aspect_ratio: "auto" to match the reference; leaving the field out is not the same. For video try-on rather than stills, see virtual try-on video.
Sources
Related posts
More in Use cases
- Seamless looping food video with AI: same first and last frame
To loop a food clip, send one image as image_url and end_image_url to Gemini Omni Flash 1.1, which Google says suits seamless loops. Clips run 3 to 10 seconds.
- Gemini TTS two-speaker limit: three-voice dialogue with Sume concat
Gemini TTS configures two speakers per request. For three voices on Sume, make one TTS job per voice and join up to 20 parts with Timeline audio concat.
- Change a garment's colour in a photo with a mask_url edit
Recolour clothing with a mask: send the photo as an input reference, a public HTTPS mask_url, and a colour prompt to ChatGPT Image 2.5 on POST /v1/images.
- HeyGen avatar new outfit with reference images vs Sume
HeyGen prompt avatars take avatar_id plus up to three reference_images for a new outfit. Sume's photo input takes one image_url per avatar.
Written by Sume