Pinterest Visual Search Ads: product photo variants from one image
Pinterest Visual Search Ads put ads in visual search results. Make angle and context variants of one product photo with Sume's image-to-image reference input.

Pinterest Visual Search Ads are ads placed in Pinterest search results and Pin closeups, and visual search starts from a picture, so one product photo is rarely enough creative. With Sume you can send one reference image to POST /v1/images and ask for several angle or context variants, up to 10 images per call with n (per-model ceilings are lower).
Pinterest facts are from its newsroom post, read 2026-10-01. Sume facts are from the Image API docs. This post does not say what Pinterest will accept or rank; check Pinterest's own ad requirements.
What did Pinterest announce?
At Pinterest Presents, Pinterest announced Visual Search Ads. The post says the format places advertisers in Pinterest Search Results and Pin closeups, pairs keyword relevance with a new ad format, and will be available in beta for eligible advertisers in all Pinterest ad markets in the coming weeks. It also says Pinterest users run more than 80 billion searches per month, the vast majority visual.
Why would one product photo need variants?
Visual search matches on what the picture shows, so a single front-on packshot covers a narrow slice of queries. Variants that change the angle, the surface or the room give you more distinct pictures to test. Pinterest's post also mentions self-serve A/B testing in Performance+, which is the place to compare them.
How do I request variants from one reference?
Put the photo in input_references as a public HTTPS image_url, describe the change in prompt, and set n. The docs say reference URLs must be public HTTPS; localhost, private-network and non-HTTPS URLs are rejected before submission.
const res = await fetch("https://api.sume.com/v1/images", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.SUME_API_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "openai/gpt-image-2",
prompt: "same product on a walnut shelf, three-quarter angle, soft window light",
n: 4,
input_references: [
{ type: "image_url", image_url: { url: "https://example.com/product.jpg" } },
],
}),
});
console.log(res.status, await res.json());What limits should I check first?
| Item | What the docs say |
|---|---|
| Images per call | Up to 10 with n; per-model ceilings are lower |
| Reference inputs | Read the input_references range for the model in the catalog |
| Text-only models | input_references of min 0, max 0 reject references |
| Reference URLs | Public HTTPS only |
| Transparent output | Use Image 1.0 with transparency: true for transparent stills today |
What should I do next?
List models with GET /v1/images/models, pick one whose input_references range is above zero, and generate a small batch. Keep the product itself consistent by reviewing every variant before it goes into an ad. For a fuller pipeline see AI product photography API.
Sources
Related posts
More in Use cases
- Podcast video clip in vertical format: 1080x1920 by default
Timeline 1.0 renders 1080x1920 MP4 by default, so a podcast clip is vertical unless you set output width and height. Cut, segment and render steps.
- AI podcast intro music generator: one prompt, fixed price
Write a podcast intro as one Music Router prompt of 1 to 5000 characters. Every generation is charged the fixed Music price, whichever model id you route to.
- Podcast audio to a 45-second teaser video with Timeline source_in
Timeline 1.0 takes an audio spine with a source_in in-point and a still as a static hold, so a podcast excerpt plus cover art renders as a video with no editor.
- App Store preview poster frame: default 5 seconds, check it
Apple's default poster frame for an app preview is 5 seconds in. Extract that exact frame from your render with Sume video_frames before you upload.
Written by Sume