Server
Concurrency
Concurrency is how many requests the server has started but not yet finished at one moment. It is a count of work in flight, not a rate.
How it is measured
Read it from the number of open requests, active connections, busy workers or in-flight jobs, sampled every second. Little's law links the three numbers: concurrency equals arrival rate times average time in system.
At 50 requests per second with a 200 ms average, expect about 10 in flight. If latency doubles to 400 ms at the same arrival rate, concurrency doubles to 20 without a single extra visitor.
Worked example
A PHP-FPM pool for a membership site has 16 workers. At 09:00, 60 requests per second arrive and each takes 150 ms, so about 9 workers are busy. A slow report starts taking 1.5 seconds, and with the same arrival rate 90 would be needed, so the pool pegs at 16.
The extra requests wait in the listen queue, and the load balancer sees 504 responses. The traffic never changed. The time each request held a worker did.
How it differs
Concurrency counts work in flight now. Throughput counts work finished per second. Concurrency excludes any statement about time, and throughput excludes how many were waiting. A system can hold high concurrency with low throughput when everything is stuck, which is the pattern of a slow dependency.
Common errors
Confusing concurrent users with concurrent requests. Sizing a pool by requests per second alone. Assuming a bigger limit increases speed. Ignoring that keep-alive connections count as open but not as busy. Comparing concurrency across services with different request durations.
In practice
Find the maximum concurrency your app can sustain by watching worker busy count while you run a ramp test, then set pools and limits slightly below that number. Alert when in-flight requests sit at the cap for more than a minute.