GPT Image 2.5 Flare: 5 images a minute at Tier 1 vs Sume limits
OpenAI limits Flare to 5 images per minute at Tier 1 and 250 at Tier 5. Sume applies plan concurrency, queue capacity and 429 codes instead. Both tables here.

OpenAI's Flare page lists 5 images per minute (IPM) at Tier 1, rising to 250 at Tier 5. Sume's docs describe different limits: plan concurrency, a queue, and 429 codes. They do not say that Sume mirrors OpenAI's tier table, so read the dashboard for your own numbers.
OpenAI's figures are from its model page; Sume's from the generation admission docs. Both read 2026-10-01.
What are OpenAI's Flare limits?
| Tier | Tokens per minute | Images per minute |
|---|---|---|
| Tier 1 | 100,000 | 5 |
| Tier 2 | 250,000 | 20 |
| Tier 3 | 800,000 | 50 |
| Tier 4 | 3,000,000 | 150 |
| Tier 5 | 8,000,000 | 250 |
What limits does Sume apply?
Sume's admission docs list plan-based concurrency: Free 1, Pro 4, Startup 8, Scale 20, Enterprise 20. Concurrency is a dispatch limit, not a submit limit: extra jobs wait in a queue of max(3, concurrency x 5). The docs also say to prefer the generation_limits.concurrency_limit field in your dashboard over the static table.
What errors do I handle?
Three are documented: 429 queue_full when no accepted capacity is left (wait or cancel queued jobs, retry with the same idempotency key), 429 rate_limited for request volume (back off, use retry-after when present), and 402 insufficient_credits when the estimated cost cannot be reserved.
How do I send a burst?
Use mode: "async" so submits return a job at once, then read results from the job. Count the queue size from your plan before sending a batch; one that exceeds it gets queue_full rather than being slowed.
curl -X POST https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Idempotency-Key: flare-batch-0001" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-image-2.5",
"prompt": "Studio photo of a ceramic mug on linen",
"mode": "async"
}'Sources
Related posts
More in Developers
- GPT Image 2.5 largest square: 2880x2880, not 3000x3000
GPT Image 2.5 caps pixels at 8,294,400. A square tops out at 2880x2880; 3000x3000 (9,000,000) is rejected. OpenAI calls sizes above 2560x1440 experimental.
- GPT Image 2 input_fidelity: omit it; Sume returns 400
OpenAI says to omit input_fidelity for gpt-image-2 because inputs run at high fidelity. Sume lists no such field and rejects unlisted parameters with 400.
- GPT Image moderation_blocked vs Sume content_policy_rejected
OpenAI returns moderation_blocked with moderation_details. On Sume, policy refusals are grouped under content_policy_rejected. What to read and when to retry.
- GPT Image revised_prompt: what a Sume images response returns
OpenAI returns revised_prompt on the image generation call. A Sume POST /v1/images response has data[].url and usage, with no revised prompt field.
Written by Sume