Replicate API rate limits: 600 creates a minute, then 429

Replicate's API allows 600 prediction creates and 3,000 other requests per minute. Low credit and no card tighten it; over the limit you get a 429.

4 min readSume
All posts

Replicate's API rate limits are 600 requests per minute for creating predictions and 3,000 requests per minute for every other endpoint. Short bursts above those numbers are allowed before throttling starts, limits get stricter as your credit runs low, and a request over the limit gets HTTP 429.

These facts come from Replicate's own Rate limits page, read on 2026-09-29. Replicate can change them; check the page before you size a job runner.

What are Replicate's API rate limits?

Replicate publishes two default limits and two conditions that lower them:

From Replicate's Rate limits page, read 2026-09-29.
CaseLimit
Creating predictions600 requests per minute
All other endpoints3,000 requests per minute
Short burstsAllowed above the default limits before throttling
Credit running outStronger rate limits (no number published)
Granted credit, no payment method on file1 request per second, at most 6 requests per minute
Higher limitsContact Replicate

What happens when I hit the Replicate rate limit?

Replicate answers with status 429 and a body like { "detail" : "Request was throttled. Your rate limit resets in ~30s." }. The page doesn't document rate-limit headers, so read the detail text or back off on your own schedule.

In your client:

  • Treat 429 as retryable: wait, then send the same request again, with growing delays and some jitter.
  • Pace creates below 600 a minute with a client-side limiter rather than retrying into the wall. See client-side rate limiting in Python.
  • Count status polls too: they are not prediction creates, so they fall under the 3,000-a-minute figure for other endpoints.
  • Tell a 429 apart from a 5xx outage before retrying: see 429 vs 503.

Why am I throttled below 600 a minute?

Two account states lower the limit. First, as you approach running out of credit, Replicate applies stronger rate limits, to stop accidental overspending and to give you time to top up before you are shut off. Replicate's own advice is to set up credit auto-reload and keep your balance above $20. Second, if you've been granted credit and have no payment method on file, you're limited to 1 request per second with a maximum of 6 requests per minute. If a batch that normally runs fine starts getting 429s, check your balance first. For what Replicate charges, see Is the Replicate API free?.

How do I get higher Replicate API limits?

The page says to contact Replicate if you want higher limits. It doesn't publish a self-serve tier ladder, so there is no number to plan around above the defaults until Replicate agrees one with you.

Is a request limit the same as how many jobs run at once?

No. A rate limit caps how many HTTP requests you send in a window; a concurrency limit caps how many generations run at the same time. Replicate's rate-limits page covers only the first.

APIs split these differently, so check both numbers when you compare hosts. Sume, for example, gives each API key separate read and write request budgets and, separately, accepts jobs over its concurrency limit as `queued`; Sume API errors and rate limits has the numbers.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume