AI text to product image: what a prompt alone can give you

AI text to product image draws a product that fits your words, not your actual product. Use it for concepts; for listing photos, add a real photo.

4 min readSume
All posts

AI text to product image works, but from text alone the model draws a product that fits your description, not your product: it has never seen your label, shape, or exact colors, so the result won't match what you ship. Use text-only product images for concepts, mood boards, and scenes; for photos that show the item you sell, send at least one real photo of it as a reference.

Below is how both work with Sume's Image API, read on 2026-09-29. If you are asking the same question about a chat assistant, can Claude create product images? covers that route.

When is a product image from text enough?

When no one will compare the picture with a real item. Sume's terms describe generated outputs as synthetic and say they may be inaccurate, so a text-only image should not stand in for the item a buyer will receive.

  • A product that doesn't exist yet: packaging ideas, colorways, and shapes to discuss before a sample is made.
  • Mood boards and art direction for a shoot.
  • Generic scenes and props: a marble counter, a kitchen shelf, a gym bag on a bench, with no brand shown.
  • Blog headers and slides where a generic product stands in for a category.

How do I generate a product image from a text prompt?

Send a prompt with no references to POST /v1/images. Describe the product the way a photographer's brief would: the object, its material and color, the background, the light, and the angle. Ask for several versions with n, within the model's range; the docs' catalog example for Seedream 4.5 allows 1 to 4, and square 1:1 among its ratios.

curl -X POST "https://api.sume.com/v1/images" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "bytedance-seed/seedream-4.5",
    "prompt": "Studio product photo of a matte sage-green ceramic coffee mug with a thin gold rim, no logo or text, on a light oak table, soft window light from the left, three-quarter view, plain cream wall behind",
    "aspect_ratio": "1:1",
    "n": 4
  }'

How do I make the image show my actual product?

Add real photos. On models that take references, send photos of your product in input_references and describe the scene you want around it. What a model accepts is published in its catalog descriptor:

  • An input_references range of {"min": 0, "max": 0} means the model is text-to-image only and rejects references.
  • In current code, edit-capable models take up to 10 references, and ChatGPT Image 2.5 takes up to 16.
  • Reference URLs must be public HTTPS; localhost, private-network, and non-HTTPS URLs are rejected.
  • On these edit calls, send aspect_ratio: "auto" to keep the reference's shape; leaving the field out is not the same.

Which kind of product photo should I make next?

Once a real photo is in the loop, the job has a name and its own post: a clean packshot is explained in what is a packshot, a studio shot on white in white background product photo with AI, and scenes around the product in AI lifestyle product photography.

How much does an AI product image cost?

Text-only and reference calls are billed the same way: per completed image at the model's endpoint price.

From Image API, read 2026-09-29.
RuleWhat the docs say
BillingCompleted generations billed in full; failed or cancelled ones are not billed
Price per modelThe pricing line of GET /v1/images/models/{model_id}/endpoints; cost_usd × n is what you pay
Images per callUp to 10 with n; per-model ceilings are lower
Resolution tiers512, 1K, 2K, 4K, as each model's catalog lists
Result URLsSume-hosted and signed; download the files you keep

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume