Atomic Redis Lua Rate Limiting for APIs: Strategies for Developers
Redis ready API rate limiting for developers: choose sliding window, token, or leaky bucket; use atomic Redis Lua scripts and return RateLimit headers.

For most public APIs, a sliding window counter is the right default: it gives near-exact accuracy with constant memory per client. Reach for a token bucket when you need to allow controlled bursts, or a leaky bucket when a fragile downstream service demands a strict and steady drain. Choose based on memory budget, burst tolerance, and how fragile your downstream systems are, and implement the enforcement logic with atomic Redis and Lua scripts backed by RateLimit headers.
TL;DR:
- A sliding window counter balances accuracy and memory efficiency, making it suitable for general public APIs with high traffic.
- Implementing rate limiters atomically with Redis Lua scripts prevents race conditions and maintains consistency across multiple application instances.
- Use the refill rate slightly above normal P99 traffic and set bucket capacity to handle 5 to 10 seconds of burst traffic without false positives.
- Inform clients of their rate limits via standardized headers and provide clear Retry-After responses to help them manage retries effectively.
- Layer enforcement at both edge (like NGINX) and application levels, and choose algorithms based on memory constraints, burst tolerance, and downstream system fragility.
Table of Contents
- How the main rate-limiting algorithms compare
- Implementation notes for each algorithm
- Building limiters that hold up across instances
- Choosing the right algorithm for your endpoint
- Communicating limits to clients through headers
- Setting refill rates, capacity, and alerts
- Where these patterns show up in production
- Fairness, abuse prevention, and the ethics of throttling
- Why sensible defaults save you from yourself later
- Where to read more on rate limiting standards
- Sources
- FAQ
How the main rate-limiting algorithms compare
Five algorithms cover almost every production case; each sits at a different point on the trade-off between memory cost, burst tolerance, and accuracy.
- Fixed window: one counter per time slot, cheapest to run, but allows a burst near the window boundary.
- Sliding window log: stores every request timestamp, giving exact counts at the cost of memory that grows with request volume.
- Sliding window counter: blends the current and previous window counts into a weighted estimate, holding memory constant per client.
- Token bucket: tracks a token count and a refill timestamp, letting clients spend saved-up capacity in short bursts.
- Leaky bucket: processes requests at a fixed output rate regardless of how they arrive, smoothing traffic for fragile downstreams.
Fixed window suits low-stakes internal limits, sliding window log suits low-volume high-stakes endpoints like login, sliding window counter suits general public APIs, token bucket suits developer-facing APIs that need burst headroom, and leaky bucket suits queues in front of rate-sensitive backends.
Implementation notes for each algorithm
Each algorithm needs a specific state shape and carries its own runtime cost, so picking one is really picking a data structure.

Fixed window stores a single counter keyed by client ID and time bucket, incremented on each request and reset when the bucket rolls over. It is simple to reason about, but a client can send a full quota’s worth of requests at the very end of one window and another full quota at the start of the next, doubling the effective rate for a short span. That makes it acceptable for coarse, low-risk limits but risky for anything security-sensitive.
Sliding window log keeps a sorted set of request timestamps per client, trimming anything older than the window on each check. It is exact, since it counts real requests rather than an estimate, but memory grows with request volume, which makes it expensive for high-traffic clients.
Sliding window counter avoids that cost by storing just two counters, one for the current window and one for the previous window, and computing a weighted estimate based on how far into the current window you are. This is the approach Cloudflare’s edge rate limiting is built around, pairing PoP-level controls with centralized counters to support large-scale traffic while keeping memory flat.
Token bucket stores a token count and a last-refill timestamp per client. On each request, you calculate tokens earned since the last check, cap the total at bucket capacity, and deduct one token if available. Refill and consume must happen atomically or two concurrent requests can each read the same token count and both succeed when only one should.
Leaky bucket comes in two flavors: a policing variant that simply drops requests exceeding the drain rate, and a shaping variant that queues them for later processing. Reach for it when the endpoint sits in front of a database or third-party service that cannot absorb spikes.
Pro Tip: Wrap refill and consume logic in a single Lua script rather than separate GET and SET calls; a read-then-write sequence across two round trips is exactly the kind of race condition atomic scripts exist to prevent.
Building limiters that hold up across instances
A rate limiter that works in one process and fails under concurrency is worse than no limiter at all, since it gives a false sense of protection.
- Store counters in Redis so every application instance reads and writes the same state instead of drifting apart.
- Use Redis Lua EVAL scripts to combine the read, refill, and consume steps into one atomic operation, as Redis’s own rate limiting tutorial recommends, since MULTI/EXEC and optimistic locking both leave gaps under high concurrency.
- For high-volume services, aggregate counts locally per instance or per point of presence before syncing to a central store, rather than hitting Redis on every single request.
- Layer a gateway-level leaky bucket, for example in NGINX, as an outer wall against abusive traffic, and keep finer per-key limits at the application layer for legitimate but bursty clients.
- Decide fail-open versus fail-closed per endpoint before an incident forces the decision: fail open on read-heavy public endpoints so a Redis outage does not take down the whole API, and fail closed on authentication or payment endpoints where letting unlimited traffic through is the bigger risk.
Pro Tip: Log every fail-open event separately from normal traffic; a Redis outage that silently disables your rate limits is the kind of failure that only shows up in an incident review.
Choosing the right algorithm for your endpoint
Run through four axes before writing any code: how much memory you can spend per client, whether bursts are legitimate or a threat, how fragile the downstream system is, and whether exact counts or estimates are acceptable.
- Public developer APIs: sliding window counter or token bucket, since both tolerate reasonable bursts without exact-count overhead.
- Payment and authentication endpoints: sliding window log for exactness, or a fail-closed default that errs toward rejecting requests over letting suspicious traffic through.
- Downstream-limited flows: leaky bucket, so the output rate never exceeds what the fragile system behind it can handle.
- High-volume, low-risk internal traffic: fixed window, trading boundary bursts for the simplest possible implementation.
If you cannot answer what happens when Redis is unreachable, you have not finished the design, regardless of which algorithm you picked.
Communicating limits to clients through headers
Servers should tell clients where they stand before those clients start guessing. An emerging IETF draft defines RateLimit and RateLimit-Policy headers, with RateLimit-Limit and RateLimit-Reset marked as required and RateLimit-Remaining recommended but optional. Many major providers still use the older vendor-prefixed X-RateLimit-* headers, which predate the draft, so supporting both during a transition period is reasonable.
- Return RateLimit-Limit, RateLimit-Remaining, and RateLimit-Reset on every response, not just on rejections.
- Send a Retry-After header whenever you reject a request with a 429 status.
- Treat these headers as hints rather than guarantees, since load can shift between the moment a header is issued and the next request.
- On the client, use exponential backoff with full jitter rather than a fixed delay, which spreads retries out instead of creating synchronized retry storms.
A production reference reports Cloudflare’s sliding window counter runs with an extremely low error rate across very large request volumes, evidence that the estimate-based approach holds up at scale without sacrificing meaningful accuracy.
Setting refill rates, capacity, and alerts
Pull p95 and p99 requests-per-second per client from your existing metrics pipeline, then set the refill rate a bit above the sustained p99 so normal usage never trips the limiter. Size the bucket capacity to absorb roughly 5 to 10 seconds of expected p99 burst, which covers a client retrying a batch job without punishing everyone else.
- Watch your 429 rate and throttle ratio as a share of total requests, not just raw counts.
- Track request latency separately for throttled and non-throttled traffic, since a spike in the former often precedes a spike in the latter.
- Alert on Redis command failures or timeouts tied to the rate limiter, since a silent failure there defeats the whole system.
- Roll out limit changes to a small traffic percentage first, then widen once the 429 rate and latency both look stable.
Pro Tip: When you tighten a limit, announce the change and the new headers to API consumers before it ships. A surprise 429 with no context generates more support tickets than the abuse it was meant to stop.
Where these patterns show up in production
Most production stacks split enforcement across two layers rather than relying on one.
- Edge: an NGINX leaky-bucket configuration absorbs abusive traffic and obvious scraping before it ever reaches application servers.
- Application: a Redis Lua script implementing token bucket or sliding window counter enforces per-key limits, generally failing open on read-only endpoints so a cache outage does not take down the API.
- Public developer APIs: sliding window counter or token bucket, tuned to the plan tier a client is on.
- Authentication endpoints: strict, low-capacity limits, often fail-closed, since letting excess traffic through here is a security risk rather than an inconvenience.
- Webhook processors: leaky bucket or a queue-backed limiter, since retries from the sending service need to be smoothed rather than rejected outright.
Fairness, abuse prevention, and the ethics of throttling
Rate limiting is a fairness mechanism as much as a technical one: it decides which requests get served when demand exceeds capacity, and that decision affects real users and real businesses. A limit set too aggressively can lock out legitimate customers during a traffic spike, while a limit set too loosely lets a small number of abusive clients degrade service for everyone else.
Publish your limits and the reasoning behind them in your API documentation, since undocumented throttling reads as arbitrary and erodes trust with the developers building on top of your API. Apply limits consistently across similar clients rather than quietly favoring some accounts, and be explicit in your terms of service about what counts as abusive traffic, such as credential stuffing or scraping, versus normal high-volume use from a paying customer.
When you throttle a client, the response itself matters ethically as well as technically. A 429 with clear headers and a Retry-After value respects the client’s time and lets their system recover gracefully, while a silent drop or a vague error forces them to guess. For multi-tenant platforms, isolate limits per tenant so one customer’s traffic spike, whether legitimate or malicious, cannot exhaust the quota of another tenant sharing the same infrastructure. That isolation is often the difference between a minor incident and a breach of trust with paying customers who did nothing wrong.

Why sensible defaults save you from yourself later
The algorithm you pick matters less than picking one and monitoring it honestly. Sliding window counter and token bucket cover most cases with minimal ops overhead, but naive client retries without backoff or header awareness will still cause outages. Keep product, SDK, and infrastructure teams talking to each other about limits before customers find out the hard way.
— Bitblade
Where to read more on rate limiting standards
Start with the IETF RateLimit header draft for the emerging standard, then Redis’s own rate limiter tutorial for implementation code.
Sources
- Build 5 Rate Limiters with Redis: Fixed Window, Sliding Window, Token Bucket, and Leaky Bucket
- draft-ietf-httpapi-ratelimit-headers-05
- API Rate Limiting Strategies: 2026 Engineering Reference
- Counting things: a lot of different things
Recommended
Frequently asked questions
The most effective strategies combine an algorithm suited to the traffic pattern, such as a sliding window counter for general APIs or a token bucket for bursty clients, with atomic enforcement so concurrent requests cannot bypass the limit. Pairing edge-level controls with application-level per-key limits, as Cloudflare’s approach illustrates, adds a second layer of protection.
Store counters or token state in a shared store like Redis, and wrap the read, refill, and consume steps in a single atomic Lua script to avoid race conditions, a pattern detailed in Redis’s rate limiter tutorial. Return clear RateLimit headers on every response and a Retry-After header on rejections so clients know how to behave.
Start by checking memory budget, burst tolerance, and how fragile your downstream systems are, then match those constraints to an algorithm: sliding window counter or token bucket for public APIs, sliding window log or fail-closed defaults for sensitive endpoints. Tune the refill rate from your measured p99 traffic and size bucket capacity to absorb a few seconds of expected burst.
API rate limiting is the practice of capping how many requests a client can make in a given period, protecting the service from overload and keeping usage fair across clients. Servers typically communicate the limit and remaining quota through response headers, an approach formalized in the IETF RateLimit header draft.
Ready to build this yourself? Get your API key — 7-day trial, no card required — or see the Blockchain API for the full endpoint reference.