fal.ai API rate limit: concurrency from 2 up to 40

fal.ai limits how many requests run at once, not requests per minute: 2 for a new account, rising with credit purchases to 40 self-serve.

5 min readSume
All posts

fal.ai's API rate limit is a concurrency limit: it caps how many of your requests can be in progress at the same time, not how many you send per minute. A new account can run 2 requests at once, and the limit rises automatically as you buy credits, up to 40 self-serve. Queued requests over the limit wait for a free slot instead of being rejected; direct calls get a 429 that fal's SDK retries.

These facts come from fal's own Concurrency Limits page, read on 2026-09-29. fal can change them; check the page and your dashboard's Concurrency page for your current number.

How does the fal.ai concurrency limit work?

Every fal account has one global limit on requests in the IN_PROGRESS state. The page describes how it's enforced:

  • Only requests a runner is actively processing count. Requests in IN_QUEUE don't, so you can submit as many as you want to the queue.
  • Queue calls (subscribe or submit): fal checks your concurrency before dispatching. At capacity, the request stays queued and is retried with exponential backoff, with no maximum retry count.
  • Direct calls (run): the gateway checks concurrency at request time, and the fal client retries for you.
  • Requests are never dropped for concurrency, with one exception: if you set a start_timeout and it expires before a slot opens.
  • The limit applies across all endpoints you call. fal may also set extra per-endpoint limits on some high-demand models.

How do I raise my fal.ai concurrency limit?

By buying credits. The limit is set from your credit purchase history, not from a plan you pick:

From fal's Concurrency Limits page, read 2026-09-29.
Questionfal's answer
Starting limit2 concurrent requests for every new account
What raises itTotal paid invoices from the last four weeks, applied automatically
When it changesWhen a credit purchase invoice is paid, typically within a few minutes
Self-serve ceiling40 concurrent requests
Above 40, or per-endpoint allocationsContact fal's sales team (Enterprise plans)
Where to see itThe Concurrency page in the dashboard, with peak usage and a 30-day chart

What does a fal.ai 429 error mean?

If you call fal with raw HTTP, a 429 whose type is concurrent_requests_limit means you are at your concurrency limit. The response carries an X-Fal-needs-retry: 1 header, and fal says to retry with exponential backoff (1s, 2s, 4s, 8s…).

With the SDK you mostly don't see it: fal_client.run() and fal_client.subscribe() both retry 429 responses with exponential backoff up to 10 times. The difference is where the retry lives. With subscribe(), the queue requeues the request server-side with no maximum; with run(), the SDK retries the HTTP call up to 10 times. fal recommends subscribe() for production, as in its example:

import fal_client

# Queue-based call: concurrency is handled server-side
result = fal_client.subscribe("fal-ai/flux/schnell", arguments={
    "prompt": "a sunset over mountains"
})

How do other APIs handle concurrency?

The model differs by provider, so check it before you plan a bulk run. fal ties concurrency to credit purchases. Sume, for example, ties it to the plan: prepaid top-ups do not raise it, and jobs over the limit are accepted as queued until the queue is full. Video job concurrency and queueing has the numbers, and fal vs Replicate covers how fal's calls and billing work.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume