GPT Image 2.5 cached input price: does Sume pass it on?
OpenAI lists cached input at $1.25 text and $2 image per million tokens for Sunburst. Sume's docs give estimates and reserved cost, not a cache discount.

OpenAI's Sunburst page lists cached input at $1.25 per million text tokens and $2 per million image tokens, but Sume's docs describe no cache discount. Budget from the endpoint's pricing lines and the estimates in the docs, not from the cached rate.
OpenAI's figures come from its model pages; Sume's from the Image API docs. Read 2026-09-30.
What does OpenAI list?
The Sunburst page lists text input at $5 (cached $1.25), image input at $8 (cached $2), and image output at $30 per million tokens. The Flare page lists the same $5, $8 and $30 rates and says "Cached inputs receive 75% discounts".
What do Sume's docs say?
Sume states that Flare and Sunburst use the same Fal token rates: $30 per million output image tokens, $8 per million input image tokens, and $5 per million input text tokens. Output estimates use OpenAI's ChatGPT Image 2.5 size and quality calculator. At 1024x1024, xhigh output is $0.09366 and max output is $0.21072 before input tokens and Sume pricing. Input token counts are estimates, and Fal rounds the total up to $0.0001.
The docs mention no cached-input rate. They also say auto quality reserves max.
| Line | OpenAI (Sunburst) | Sume docs |
|---|---|---|
| Text input | $5 (cached $1.25) | $5 |
| Image input | $8 (cached $2) | $8 |
| Image output | $30 | $30 |
| Cache discount | Listed | Not described |
What should I budget on?
Endpoint pricing lines are the amount charged to your wallet, with Sume's margin already applied, so cost_usd x n is what you pay. Read them from GET /v1/images/models/openai/gpt-image-2.5/endpoints. Completed generations are billed in full; failed or cancelled ones are not billed.
Why not assume the cache discount?
Because Sume's page does not promise it. If a cached rate matters to your volume, treat it as unconfirmed and verify the billed usage.cost on real jobs.
How do I check this myself?
To find your real cost, run a few representative jobs and compare usage.cost on each response with the estimate you computed from the token rates. The linked docs pages and the catalog endpoint show the current values, and this post reflects them as of 2026-09-30.
Sources
Related posts
More in Models
- GPT Image 2 supported aspect ratios on Sume: 5:4, 9:8, 4:5
On Sume, gpt-image-2 accepts 5:4, 9:8 and 4:5 on top of the common ratios, plus custom pixels with a 3:1 limit. The exact list and the 9:8 size to send.
- Grok Imagine image edit limit: xAI 5 sources, Sume catalog
xAI's Aug 28 release notes raise image editing to 5 source images. On Sume, Grok is edit-capable with n of 1; read input_references per model.
- grok-imagine-image-quality retires Nov 2: what about quality?
xAI retires grok-imagine-image-quality on November 2, 2026 and serves it as grok-imagine-image-2.0 at low quality. Sume's x-ai/grok-image takes no quality.
- Grok Imagine video editing: 8.7 s on xAI, Omni edit on Sume
xAI caps Grok Imagine video edits at 8.7 seconds. Sume's Grok row has no video-to-video; edits go through gemini-omni-flash-1.1 with a video_url.
Written by Sume