Python API rate limiting: stay under a per-minute limit
Pace Python API calls with an asyncio limiter set under the API's per-minute budget, keep polling on its own budget, and back off on 429 retry-after.

To rate limit API calls in Python, pace them on the client: put every request through a limiter set a little under the API's per-minute limit, so calls wait their turn instead of failing. For asyncio code, aiolimiter.AsyncLimiter(max_rate, time_period) does this in a few lines. Keep status polling on its own limiter when the API budgets reads separately, and still back off on a 429 using its retry-after header, because the server's count is the one that matters.
This post is about calling someone else's API. Library facts come from the aiolimiter documentation and HTTPX: Async Support; the example budget is Sume's, from Authentication, Errors and rate limits and Generation admission, all read on 2026-09-29.
How do I rate limit requests with aiolimiter?
Create one AsyncLimiter and wrap each call in async with limiter:. aiolimiter is a leaky bucket: AsyncLimiter(100) allows 100 acquisitions per 60 seconds, the default time_period. Its docs note that up to max_rate acquisitions pass in a burst before pacing starts, so the first minute can see more calls than max_rate. To forbid bursts, set max_rate to 1 and time_period to the gap between calls: AsyncLimiter(1, 0.6) allows one call every 0.6 seconds, 100 a minute. Create the limiter inside the running event loop; reusing one across loops is not supported.
import asyncio, os
import httpx
from aiolimiter import AsyncLimiter
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
async def submit(client, writes, item):
for attempt in range(5):
async with writes: # at most one create every 0.6 s
r = await client.post(
"https://api.sume.com/v1/videos",
headers={**AUTH, "Idempotency-Key": f"clip-{item['id']}-v1"},
json={"model": "sume/auto", "prompt": item["prompt"],
"aspect_ratio": "9:16", "duration": 5},
)
if r.status_code != 429 or r.json()["error"]["code"] == "queue_full":
return r # queue_full is a capacity answer, not a pacing one
await asyncio.sleep(float(r.headers.get("retry-after", 2 ** attempt)))
return r
async def main(items):
writes = AsyncLimiter(1, 0.6) # 100 creates per minute, no burst
async with httpx.AsyncClient() as client:
return await asyncio.gather(*(submit(client, writes, i) for i in items))
asyncio.run(main([{"id": "8823", "prompt": "A product clip on a desk"}]))What number should the limiter use?
Start from the API's documented budget and leave headroom for other processes that share the key. On Sume, every key gets a per-minute budget across all of /v1, set by the workspace's plan, and reads and writes are counted separately. A read is any GET or HEAD, such as polling a job's status; creating jobs and cancelling are writes.
| Plan | Writes per minute | Reads per minute |
|---|---|---|
| Free | 120 | 4800 |
| Pro | 300 | 12000 |
| Startup | 600 | 24000 |
| Scale | 1200 | 48000 |
| Enterprise | Contact sales | Contact sales |
What should my code do when it still gets a 429?
Wait and resend with the same idempotency key. Sume's docs say to read ratelimit-remaining rather than counting requests yourself (rate limit headers explains each one) and to back off on retry-after, which gives the seconds to wait. The 429 names the budget it came from in error.details.scope (read or write), so you know which limiter to slow down. Don't retry an unsafe submit without an Idempotency-Key; Retry-After after a 429 covers the wait itself.
A second 429 code, queue_full, is not a request-rate limit: it means the workspace can't accept another paid generation job until an existing one finishes or is canceled. Pacing harder doesn't clear it.
Does a rate limiter cap how many jobs run at once?
No. A rate limiter spaces out requests; it says nothing about how many slow jobs are running. Sume's docs make the same split: request rate is not generation capacity, and raising your request rate does not raise the plan's concurrency limit. When concurrency is full, new valid jobs are accepted as queued while queue capacity remains. To cap jobs in flight from Python, hold an asyncio Semaphore from submit to the final status, alongside the limiter.
Sources
Related posts
More in Developers
- Python requests default timeout: there isn't one
Python Requests has no default timeout: without timeout= a call can hang indefinitely. Set (connect, read) on every call, and keep it short for job APIs.
- Real time speech to text API: what a file-based API can do
Sume's speech to text API isn't real time: it transcribes recordings of up to 10 minutes at a URL. Chunked recordings give near-live transcripts.
- Extract on-screen text from a reference video with an API
Reference ingest reads on-screen text at source resolution, returns lines with boxes, spans and confidence, and flags low-confidence lines instead of guessing.
- Reference video shot breakdown API: cuts, keyframes and timings
Break a reference video into shots with an API: POST /v1/reference-ingest returns frame-exact cuts, keyframes and a labeled strip for one clip, unbilled.
Written by Sume