Technical
HTTP 500
Also called Internal Server Error.
HTTP 500 is the status a server returns when it hit an unexpected condition it could not handle. It is the catch-all for a bug or crash on the server side, not for a bad request.
How it is measured
Count 5xx per minute from the access log or load balancer, and as a share of all responses. Then read the error log at the same timestamp for the stack trace, because the access log only records that it failed.
A 1% error rate across a busy day is a different problem from a 100% rate on one route. An uptime monitor flags a 500 on its check URL immediately, but only on that URL.
Worked example
After a plugin update, a WordPress checkout page returns 500 on every POST. The access log shows 212 500s in 40 minutes on `/checkout/`, none on GET. The PHP error log at the same second says `Allowed memory size of 134217728 bytes exhausted`.
Raising `memory_limit` to 256M and reloading clears it. The monitor, which only polled the homepage, stayed green the whole time.
How it differs
An HTTP 500 is a generic server failure. An HTTP 502 means a gateway or proxy got a bad answer from the server behind it. The 500 comes from the app itself, while the 502 comes from the layer in front and usually means the app is down or crashing.
Common errors
Monitoring only the homepage. Returning 500 for validation failures that should be 400. Hiding the error page and logging nothing. Ignoring 500s that only happen on POST. Showing a stack trace to visitors. Calling a spike a one-off without checking the deploy time.
In practice
Alert on the 5xx rate as well as on up and down, and probe the key transactional URLs. Match each alert to a deploy time. Keep the error log retained and readable, and show visitors a plain error page with no stack trace.