Server

Throughput

Also called RPS.

Throughput is how much work a system completes per unit of time, such as requests per second or megabytes per second. It measures output, and says nothing about how long each item took.

How it is measured

Count completed work in a time window: successful responses per second, bytes served per second, jobs finished per minute. Count only what finished, since requests that are still waiting are not throughput yet.

Report it with latency and error rate. Little's law ties it together: in-flight work equals throughput multiplied by average latency.

Worked example

A CDN-fronted image service serves 2.4 Gbps at peak, which is 9,500 requests per second at about 32 KB each. Origin throughput is only 180 requests per second, because 98 percent are cache hits at the edge.

A nightly import writes 1.1 million rows in 40 minutes, about 460 rows per second. Batching inserts of 500 rows raises it to 5,800 rows per second, with each batch taking longer than a single insert.

How it differs

Throughput counts work done per second. p95 latency shows how long the slower requests took. Throughput excludes time per request, latency excludes volume. A system can improve one and hurt the other: batching raises throughput and delays individual items. Concurrency sits between them.

Common errors

Quoting throughput without latency. Counting errors as completed work. Benchmarking against cache hits. Confusing bandwidth with request rate. Pushing throughput up until latency collapses past saturation.

In practice

When you test or tune, record throughput, p95 and error rate in the same table. Choose a target such as 400 requests per second at p95 under 500 ms, and stop pushing when the second number breaks.

See also

p95 latency, Concurrency

Sources

Count this on a real site.

Watch my website