Rate limiting algorithms: token bucket vs leaky bucket
Fixed and sliding windows, token and leaky buckets compared, where to enforce limits, distributed limits with Redis, and how to protect OTP, login and test submission.
On this page 13 sections
- Why rate limit
- Fixed and sliding windows
- Fixed window counter
- Sliding window log
- Sliding window counter
- Token bucket and leaky bucket
- Token bucket
- Leaky bucket
- Where to enforce limits
- Distributed limits with Redis
- Endpoints to protect first: OTP, login, submit
- Key takeaways
- Frequently asked questions
Rate limiting caps how many requests a client can make in a given time and rejects the excess, usually with HTTP 429 Too Many Requests. Four algorithms cover most needs: fixed windows (simplest, but they allow double bursts at window edges), sliding windows (smoother), the token bucket (allows short bursts while holding an average rate) and the leaky bucket (drains requests at a steady rate, and is what Nginx's limit_req uses). Choose by whether you want to allow bursts, and enforce limits in layers: at the edge, at the proxy and in the application.
Why rate limit
- Fairness. One runaway script, or a browser tab refreshing the results page every second, shouldn't slow the platform for everyone else.
- Security. Limits blunt OTP guessing, credential stuffing on login pages and scraping of your question bank.
- Cost. Every OTP is a paid SMS, and "SMS bombing" a phone number through your signup form costs you money.
- Protecting others. Your SMS gateway, payment provider and email service enforce their own limits; staying under them is your job.
People use "rate limiting" and "throttling" loosely. A useful distinction: rate limiting rejects requests over the limit, while throttling slows them down, by queueing or delaying them until they fit the allowed rate. Many tools do both, as Nginx's burst setting shows below.
Fixed and sliding windows
Fixed window counter
Count requests per key, such as a user ID, in fixed windows like "this minute", and reject once the count passes the limit. It is one counter per key and easy to build with Redis's INCR and EXPIRE. The flaw is at the boundary. With a limit of 5 a minute, a client can send 5 requests at 9:59:59 and 5 more at 10:00:00: 10 requests in two seconds, all allowed.
Sliding window log
Store a timestamp for every request, for example in a Redis sorted set, and count only those in the last 60 seconds. It is exact, but memory grows with the limit, which makes it expensive for high limits.
Sliding window counter
Keep counts for the current and previous fixed windows and estimate the sliding count by weighting the previous one. With a limit of 100 a minute, 15 seconds into the current minute, the last 60 seconds include three quarters of the previous minute. If the previous minute had 80 requests and this one has 30 so far, the estimate is 30 + 80 × 0.75 = 90, so the request is allowed. It costs two counters per key and removes most of the edge burst.
Token bucket and leaky bucket
Token bucket
Picture a bucket that holds up to B tokens and refills at r tokens a second. Each request spends one token; when the bucket is empty, requests are rejected until it refills. A client that has been quiet can burst up to B requests at once, but over time can't exceed r a second. That suits human behaviour well: a student who answers five questions quickly after a long think should be able to autosave all five at once.
Leaky bucket
Requests pour into a bucket that leaks at a constant rate; if the bucket is full, new requests are dropped. Used as a queue, it smooths traffic, so whatever arrives in bursts leaves at a steady pace, which protects a fragile downstream service. Used as a meter that rejects rather than queues, it makes much the same decisions as a token bucket.
Nginx's limit_req module uses the leaky bucket method. Its burst parameter is the size of the queue: excess requests up to that size are delayed so they are processed at the configured rate, and anything beyond is rejected. Adding nodelay serves burst requests immediately while still freeing slots only at the configured rate, which behaves much like a token bucket.
| Algorithm | Allows bursts? | State per key | Best for |
|---|---|---|---|
| Fixed window | Yes, up to double at window edges | One counter | Simple daily or hourly quotas |
| Sliding window log | No | One timestamp per request | Low limits that must be exact, such as OTP sends |
| Sliding window counter | Slightly | Two counters | General API limits |
| Token bucket | Yes, up to the bucket size | Token count and a timestamp | User-facing APIs where short bursts are normal |
| Leaky bucket as a queue | Absorbed and smoothed | A queue or counter | Protecting a downstream service that needs a steady rate |
Where to enforce limits
| Layer | What it can key on | Strengths | Limits |
|---|---|---|---|
| CDN or web application firewall | IP address, headers, path | Stops floods before they reach your servers | Coarse and approximate; AWS WAF rate-based rules, for example, count over 1, 2, 5 or 10-minute windows, with a minimum limit of 10 |
| Load balancer or reverse proxy, such as Nginx | IP address, headers, cookies, URL | Cheap, and easy to set per path | Each Nginx server keeps its own counts, so three proxies allow three times the limit |
| Application | User, phone number, device, test attempt | Knows the business rules | Needs a shared store such as Redis, and adds work to each request |
In India, IP-based limits need special care. A hostel's Wi-Fi, a coaching centre's network or a mobile network can put hundreds of students behind one public IP address, so a tight per-IP limit blocks a whole classroom. Keep IP limits generous, as a flood guard, and apply the strict limits per user or per phone number. Behind a load balancer, also make sure Nginx sees the real client IP, or every request appears to come from the balancer:
# In the http block. Trust X-Forwarded-For only from the load balancer
set_real_ip_from 10.0.0.0/16;
real_ip_header X-Forwarded-For;
limit_req_zone $binary_remote_addr zone=login_ip:10m rate=10r/m;
limit_req_status 429;
server {
location = /api/auth/login/ {
limit_req zone=login_ip burst=5 nodelay;
proxy_pass http://app;
}
}
Nginx returns 503 for rejected requests by default, hence limit_req_status 429. A 10 MB zone is plenty for this: Nginx's documentation says one megabyte holds about 8,000 client states on 64-bit systems. Where your balancer can't rate-limit on its own, see our comparison of L4 vs L7 load balancers for what each layer can do.
Distributed limits with Redis
When several app servers enforce the same limit, they need shared counters, and Redis is the usual home. Redis's own documentation for INCR describes a fixed-window rate limiter and warns about a race: if a client increments a counter but fails before setting its expiry, the key is left without one. Running both steps in a Lua script, which Redis executes without interleaving other commands, fixes that. The same approach gives an atomic token bucket:
-- KEYS[1]: bucket key. ARGV: refill rate per second, capacity, now (ms)
local rate = tonumber(ARGV[1])
local capacity = tonumber(ARGV[2])
local now = tonumber(ARGV[3])
local b = redis.call('HMGET', KEYS[1], 'tokens', 'ts')
local tokens = tonumber(b[1]) or capacity
local ts = tonumber(b[2]) or now
tokens = math.min(capacity, tokens + (now - ts) / 1000 * rate)
local allowed = 0
if tokens >= 1 then tokens = tokens - 1; allowed = 1 end
redis.call('HSET', KEYS[1], 'tokens', tokens, 'ts', now)
redis.call('PEXPIRE', KEYS[1], math.ceil(capacity / rate * 1000))
return allowed
Use one key per limited thing, such as rl:otp: followed by the phone number, pass the time from a consistent clock so app servers with drifting clocks don't disagree, and decide in advance what happens if Redis is unreachable. Failing open keeps autosave working during a Redis blip; failing closed is safer for OTP sends, where abuse costs money.
Endpoints to protect first: OTP, login, submit
| Endpoint | Key | Illustrative starting limit | Why |
|---|---|---|---|
| Send OTP | Phone number, plus IP and device | 1 per 30 seconds and 5 per hour per number | Stops SMS bombing and protects your SMS budget; show a countdown in the app |
| Verify OTP | The OTP request | 5 wrong attempts, then the code expires | A 6-digit code has only a million possibilities |
| Password login | Account, and IP address | A few failures per account before growing delays; a generous per-IP ceiling | Credential stuffing, without locking out a whole hostel |
| Password reset | Email or phone number | 3 per hour | Prevents message bombing |
| Autosave during a test | Test attempt | Token bucket, for example 2 a second with bursts of 20 | Never block genuine saves; drop exact duplicates instead |
| Submit test | Test attempt | An idempotency key, so exactly one submission counts | Retries on bad networks must succeed, not be rejected |
| Question bank and search APIs | User | Per-minute and per-day quotas | Slows scraping of paid content |
The submit row matters most on exam day. A student on a weak connection may retry a submission several times; a strict rate limit would turn a network blip into a lost attempt. Make submission idempotent instead, and keep rate limits for abuse. The guide to running an online exam covers the rest of exam-day preparation.
Finally, tell clients what happened. RFC 6585 defines 429 Too Many Requests, lets you add a Retry-After header saying how long to wait, and says 429 responses must not be cached. Well-behaved apps then wait, with some random jitter, instead of retrying at once. For how limits fit into a platform built for heavy days, see handling 100,000 concurrent users.
Key takeaways
- Fixed windows are simple but allow edge bursts; sliding windows fix most of that.
- Token buckets allow natural bursts within an average rate; leaky buckets smooth traffic to a steady flow.
- Enforce limits in layers: coarse per-IP limits at the edge, strict per-user limits in the application.
- Shared IP addresses are common in hostels, coaching centres and mobile networks, so key strict limits on users.
- Protect OTP, login and reset first; make test submission idempotent rather than rate-limited.
Frequently asked questions
What is rate limiting in API?
API rate limiting restricts how many requests a client, identified by an API key, user, token or IP address, may make in a time window, such as 100 requests a minute. Requests over the limit are rejected, usually with HTTP 429 and a Retry-After header, until the window resets or the bucket refills. It keeps one client from monopolising the API, blunts abuse and protects services behind the API.
Can load balancer do rate limiting?
Many can. Software load balancers and proxies such as Nginx, HAProxy and Envoy have built-in rate limiting by IP address, header or path. On AWS, you attach AWS WAF rate-based rules to an Application Load Balancer or CloudFront distribution. These layers are good at coarse, per-IP flood control, but limits tied to users, phone numbers or test attempts usually belong in the application, backed by a shared store such as Redis.
What is rate limiting and throttling?
Both control request rates, but they respond differently. Rate limiting rejects requests above the allowed rate, usually with HTTP 429. Throttling slows them down instead, by queueing or delaying excess requests until they fit the allowed rate, or by reducing a client's quota. Nginx's limit_req shows both: requests within the burst allowance are delayed and processed at the set rate, and requests beyond it are rejected.