Agent loops on Sume: tool calls spend writes, polls spend reads

An MCP tool call spends the write budget once for the run it creates; a jobs_status poll spends none. Reads have their own, much larger budget.

4 min readSume
All posts

An MCP tool call spends the write budget for the run it creates, once, not for the JSON-RPC request that carried it. A jobs_status poll over MCP spends no write budget at all. Reads have a separate budget, so a tight status-poll loop cannot 429 your own submits.

Why does this matter for agents?

OpenAI's changelog entry of Sep 29 added computer use to the Agents API, where agents can complete tasks in an OpenAI-hosted browser. Agent loops of this kind call tools and re-check status many times. Sume's rate limit docs say the read multiple is sized for agents: an agent harvest can hold twenty-odd jobs open and poll each of them.

What are the budgets per plan?

Per-minute request budgets from the Sume authentication docs, read 2026-09-30.
PlanWrites per minuteReads per minute
Free1204800
Pro30012000
Startup60024000
Scale120048000

What counts as a read and what as a write?

A read is any GET or HEAD (polling status_url, events_url, result_url, listing Formats or runs) plus the two POSTs that submit nothing: /v1/generation/admission-preview and the MCP endpoint itself. Everything else, such as creating runs, cancelling and uploads, is a write. Reads get forty times the plan's write number.

How should the loop react to a 429?

Read ratelimit-remaining instead of counting requests, and back off on retry-after. A 429 names the budget in error.details.scope (read or write); see 429 scope: read or write. Request rate is separate from concurrency: how many generations run at once is governed by the plan's concurrency limit on generation_limits.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume