Server
Rate limit
Rate limit is a ceiling on how many requests one client may make in a time window. Beyond it, the server refuses or delays the extras.
How it is measured
Define the key (IP, token, user), the window (per second, per minute) and the algorithm: fixed window, sliding window or token bucket. Then count requests per key and the number rejected, usually with a 429 and a Retry-After.
Measure the effect on legitimate users as well as abusers. Check the rejected requests by key, since one shared office IP can hit a per-IP limit that a single bot would.
Worked example
A WordPress login at /wp-login.php receives 3,000 POSTs in ten minutes from 40 IPs, guessing passwords. Nginx `limit_req` at 5 requests per minute per IP, with burst 3, rejects most and returns 429 to the rest.
A legitimate staff member behind the same office NAT as two colleagues hits the limit while trying to log in, so the team adds the office address to an allow list and keeps the global limit.
How it differs
A rate limit is the rule. HTTP 429 is the response a client sees when it is applied. The limit excludes any specific reply, since some systems silently drop or delay. The 429 excludes the policy that triggered it. A WAF may add limits too, but it adds content inspection as well.
Common errors
Limiting by IP when users share one address. Not returning Retry-After. Applying one limit to all endpoints. Letting the limiter state live on one server in a cluster. Setting limits so tight that your own scripts and monitors are blocked.
In practice
Add stricter limits to login, search and any endpoint that sends email or hits a paid API, with looser limits for the rest. Allow-list your uptime monitor and log every rejection with its key.