Skip to content
Prompt Words
History

Rate limiting

Also called: throttling, token bucket, request quota, API rate limit.

Capping how many requests each client may make in a period and refusing the rest, usually with HTTP 429 Too Many Requests. A token bucket allows short bursts: every request spends a token, and tokens refill at a steady rate up to a fixed capacity.

The bucket holds 5 tokens and gets 1 back every second. Each request spends one; with none left, the answer is 429.

No requests yet.

    Tokens 5

    Say it in a prompt

    Rate-limit the public API per API key with a token bucket in Redis: capacity 20 tokens, refilled at 10 tokens per second. When the bucket is empty, return 429 Too Many Requests with a Retry-After header in seconds, and send X-RateLimit-Remaining on every response.

    Seen on

    • Stripe: Allows 100 API requests per second per account in live mode and 25 in a sandbox, answers extra requests with 429 Too Many Requests, and suggests a client-side token bucket.
    • Amazon API Gateway: Throttles with the token bucket algorithm: the rate is how fast tokens are added and the burst is the bucket's capacity; throttled clients get 429.
    • GitHub REST API: Authenticated users get 5,000 requests per hour; the x-ratelimit-remaining and x-ratelimit-reset headers show what is left and when it resets.

    You might describe it as

    • stop one client from flooding the API
    • only let each key call it 10 times a second
    • a bot is hammering our endpoint

    Not to be confused with

    • Backpressure

      Rate limiting is a fixed allowance per client, and requests over it are refused; backpressure slows the sender down based on how busy the receiver is right now.

    • Circuit breaker

      Rate limiting protects your service from callers sending too much; a circuit breaker protects you from a dependency that is failing.