Rate Limiting Strategies That Actually Work at Scale
Why Rate Limiting Is Harder Than It Looks
Rate limiting sounds trivial: pick a number, reject anything above it. In production it becomes one of the subtlest problems in API infrastructure, and most teams get at least one dimension of it wrong.
The difficulty is that a rate limit is really a negotiation between three parties who disagree: the client that wants maximum throughput, the origin service that needs protection, and the business that wants neither to churn customers nor to fall over under load.
Fixed Window Versus Sliding Window
The simplest algorithm is the fixed window: count requests in a period, reset at the boundary. Its weakness is the edge. A client can send a full quota at 12:00:59 and another full quota at 12:01:01, doubling your intended rate in two seconds.
Sliding windows count over a rolling period and smooth out that burst, at the cost of storing per-client timestamps rather than a single counter. For most APIs a sliding window paired with a token bucket gives the best balance: controlled bursts with a sustainable refill rate.
The Distributed Counter Problem
Run more than one gateway node and your counters have to live somewhere every node can see. Store them in each node's memory and a client gets your limit times the number of nodes.
A shared low-latency store such as Redis fixes correctness but adds a network hop to every check. Atomic increments via Lua scripts and pipelining keep that hop cheap, but the architecture has to account for it from day one, not as an afterthought.
Choosing the Right Key
Per-IP limiting is easy and wrong for anyone behind corporate NAT. Per-client limiting keyed on an API token is fairer and lets you set different ceilings per plan tier.
Per-operation limiting protects expensive endpoints independently. A search call that scans millions of rows deserves a tighter budget than a health check. Apply limits at the operation level, not only the client level.
Communicating Limits Honestly
The best rate limiting is transparent. Return standard headers describing the remaining budget and reset time on every response, not just on rejections, so clients can slow down before they hit the wall.
An API that throttles silently is far harder to integrate with than one that tells clients exactly where they stand. Good limits make well-behaved clients easy to build.
Tuning With Data Instead of Guesses
The hardest part is picking the numbers. Measure the real request distribution per client across a normal week and set the ceiling above the 99th percentile of legitimate use.
Then watch your rejection rate after launch. A limit that never triggers is probably too loose; one that constantly trips paying customers is too tight. Treat the threshold as something you revisit, not a constant you set once.









