Cap what one Format run can spend with generation_spend_cap_usd
Every Sume Format run has a generation spend cap. Send generation_spend_cap_usd to set it: up to 500, null runs at 500, 0 is rejected. What happens at the cap.

To cap what one Format run can spend, send generation_spend_cap_usd in the body of POST /v1/formats/{handle}/{slug}/runs. Omit it to inherit the Format's own cap; send a number up to 500 for this run's ceiling; null runs at the $500 platform maximum; 0 or anything above 500 is 400. A run can never spend past its effective cap.
Rules are from Create a run, read 2026-09-29.
What does it look like?
The receipt echoes the effective cap as usage.generation_spend_cap_usd_micros.
curl -X POST https://api.sume.com/v1/formats/acme/promo/runs \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: promo-cap-001" \
-d '{
"instruction": "One 9:16 clip, three shots",
"input": { "product_name": "Aurora Headphones" },
"generation_spend_cap_usd": 8
}'What does each value do?
| You send | The run's cap |
|---|---|
| Nothing | The Format's cap ($400 when it never set one) |
| A number up to 500 | That number; above the Format's own cap is honored, not clamped |
null | $500, the platform maximum |
0, or above 500 | 400 |
What happens at the cap?
The run ends failed and usage shows how close it got. The generic format_run_failed error is also what a run that wanted to spend past its cap lands on, so compare usage.billable_amount_usd_micros with usage.generation_spend_cap_usd_micros before you raise the brief.
What does the spend figure include?
Metered generation: video, image, avatar, voice and timeline work. It excludes the agent's own LLM turn, so it is not the run's total cost, and it is a receipt figure, not an invoice; GET /v1/usage and GET /v1/balance are the billing records. The docs suggest caps around $120 for production long-form runs and a few dollars for a single-scene retry, which are their examples, not a recommendation for you.
Do scheduled runs work the same way?
Not quite: a schedule has its own cap and a per-run override can only lower it. See the schedule cap.
How do I cap a whole bulk batch?
Each bulk item is the same body as a single run, so each item can carry its own generation_spend_cap_usd. A bulk request queues 1 to 100 items with a concurrency window of 1 to 16; poll GET /v1/format-run-queues/{id}. The queue has no webhook of its own, and completed means every item is terminal, not that every item succeeded, so check counts.failed. See bulk runs.
Sources
Related posts
More in Developers
- Gemini 3.8 Flash function calling: let it start a Sume video job
Declare a function for Gemini 3.8 Flash, run it against Sume's /v1/videos when the model returns a functionCall, and reply with a functionResponse with job id.
- Golang HTTP client retry: what net/http retries on POST
Go's http.Transport retries only network errors on reused connections, and a POST only with an Idempotency-Key header. Write the 429 and 5xx loop yourself.
- GPT-6 Astra tool calling needs the Responses API: meaning for Sume
OpenAI says GPT-6 Astra supports Chat Completions but its tool calling requires Responses. Use the Responses mcp tool for Sume, not a chat.completions loop.
- GPT Image 2.5 4K: how to request a 3840x2160 image by API
To get a 4K image from GPT Image 2.5 on Sume, send image_size 3840x2160 to POST /v1/images and be ready for a 202 job response. Request, cost, and polling.
Written by Sume