Server
Requests per second
Also called RPS.
Requests per second is the count of requests a server handles each second. It is a rate, and by itself it says nothing about how quickly each one was answered.
How it is measured
Count completed requests in one-second buckets from access logs or metrics, and report the peak as well as the average. Say which requests count: all, only 2xx, or only dynamic ones, since CDN-served static assets can dwarf the figure.
For load tests, hold the number steady and watch latency and errors. The meaningful number is the highest rate at which p95 stays under target with an error rate near zero.
Worked example
A ticket site expects its on-sale at 10:00. Logs from the last event show 90 requests per second on average but 1,150 in the first 20 seconds. The old setup handled 300 before p95 passed 2 seconds.
A load test at 1,200 requests per second against the new setup with a queue in front shows p95 at 640 ms and no errors. The team now knows the number that matters is the peak, not the daily average.
How it differs
Requests per second counts how many requests arrive or finish in a second. Throughput is the broader idea of work done per unit of time, and RPS is its usual unit for web servers. RPS excludes bytes and job counts, throughput can use any unit. Neither excludes latency, so quote both.
Common errors
Quoting the daily average when the peak is what breaks things. Counting static assets with pages. Treating a benchmark of one fast endpoint as site capacity. Ignoring failed requests in the count. Testing from the same region and network as the server.
In practice
Find your peak RPS in the last 30 days from access logs, then load-test to at least double that. Record it with the p95 beside it.