Notion 429/529 retry_after in the body vs Sume retry-after header
Notion now puts additional_data.retry_after in 429/529 bodies. Sume sends a retry-after header on rate_limited. One small helper can read both.

Notion's 2026-09-24 changelog says 429 and 529 responses now include additional_data.retry_after in the response body. Sume's 429 rate_limited carries a retry-after header when present. A single helper can read the wait from either place, then retry.
Notion facts are from its changelog, read 2026-10-01. Sume facts are from Errors and credits and Generation admission.
Where does each API put the wait time?
Different places, so a header-only retry layer will miss Notion's value and a body-only one will miss Sume's.
| API | Status | Where the wait is |
|---|---|---|
| Notion | 429, 529 | additional_data.retry_after in the body |
| Sume | 429 rate_limited | retry-after header, when present |
| Sume | 429 queue_full | No wait value; wait for jobs to finish or cancel queued ones |
What does one retry helper look like?
Read the header first, fall back to the body, and cap the number of attempts. This sketch assumes the value is a number of seconds, so check the unit in each API's reference before relying on it.
async function waitFor(res) {
const header = res.headers.get("retry-after");
if (header) return Number(header);
const body = await res.clone().json().catch(() => null);
const fromBody = body && body.additional_data && body.additional_data.retry_after;
return fromBody ? Number(fromBody) : 1;
}
export async function send(url, init, tries = 4) {
for (let i = 0; i < tries; i++) {
const res = await fetch(url, init);
if (res.status !== 429 && res.status !== 529) return res;
const seconds = await waitFor(res);
await new Promise((r) => setTimeout(r, seconds * 1000));
}
throw new Error("still rate limited");
}Is it safe to retry a paid Sume submit?
Yes, if the same Idempotency-Key goes on every attempt. The docs say not to retry unsafe submit requests without one, and a 409 idempotency_conflict means a key was reused for a different payload. See axios retry with an idempotency key for the pattern.
Why is queue_full different?
The docs state that queue_full is different from ordinary request rate limiting: Sume cannot accept another paid generation job for the workspace until an existing queued or processing job finishes or is canceled. Backing off on a timer alone may not clear it, so read the error.code before choosing a retry policy. Compare 429 vs 503.
Sources
Related posts
More in Developers
- Notion MCP 20 calls per 10 seconds: poll Sume jobs less
Notion MCP allows 20 search and 20 data source query calls per connection every 10 seconds. Use a Sume webhook so job polling does not compete for them.
- OpenAI Agents API durable sessions and Sume job ids
A durable OpenAI Agents API session continues across turns; a Sume render is a separate job. Keep the job id in the session and re-read it, never re-create.
- OpenAI Agents SDK 0.20 httpx2 MCP headers: Sume API key
With a custom MCP HTTP client in Agents SDK 0.20, send the Sume key as an Authorization Bearer or x-api-key header. httpx2 only matters if you built the client.
- Agents SDK mcp.Client mode auto against Sume's MCP server
With MCP Python SDK v2, the Agents SDK probes the newest protocol and falls back to initialize. Sume's hosted server still serves initialize with 2025-11-25.
Written by Sume