Short answer
A rate limiter caps how many requests a client can make in a period. Common algorithms are token bucket (allows bursts up to a capacity, refills at a steady rate), leaky bucket (smooths output), fixed window counters and sliding window logs or counters. In distributed systems the counters usually live in a shared store such as Redis, updated atomically.
Token bucket sketch
typescriptfunction allow(bucket: Bucket, now: number): boolean {
const elapsed = (now - bucket.updatedAt) / 1000;
bucket.tokens = Math.min(bucket.capacity, bucket.tokens + elapsed * bucket.refillPerSecond);
bucket.updatedAt = now;
if (bucket.tokens < 1) return false;
bucket.tokens -= 1;
return true;
}Design decisions
- Key by user, API key, IP or endpoint.
- Return HTTP 429 with a Retry-After header.
- Decide whether to fail open or closed if the limiter store is unavailable.
- Use atomic operations or scripts to avoid race conditions.
How to answer it in an interview
- Compare fixed window edge bursts with a sliding window.
- Mention where the limiter runs: gateway, service or both.