Uptime
Failover
Failover is moving traffic from a failed origin to a standby. It can be automatic, or a human can flip it.
How it is measured
Time it from the first failed probe to the first successful user request on the standby. The parts are detection delay, decision (confirmations), and propagation (DNS TTL or load balancer update). Record each one.
Also count what the switch lost: writes not yet replicated, sessions on the old node, and queued jobs. That count is your actual data loss for the event.
Worked example
A listings site fails over from its primary in Dallas to a replica in Chicago. Probes run every 30 seconds with three confirmations, so detection takes 90 seconds. The DNS TTL is 120 seconds, so most traffic arrives by about 3.5 minutes. The replica was 2.1 seconds behind, so 6 listings posted in that window are missing.
On review the team moves the switch to a load balancer health check, which cuts propagation from 120 seconds to under 10. The total comes to roughly 100 seconds.
How it differs
Failover is the action of switching. Active-passive is the arrangement of keeping a standby to switch to. You can have a passive spare and no failover plan, and you can fail over between two active nodes.
Common errors
Flapping failover on a blip. Failing back before the old primary is repaired. Promoting a replica that is lagging. Forgetting application-stored IPs or connection pools. Never practicing.
In practice
Run a controlled failover on staging, then in a quiet hour on production, and time each stage. Set failback to manual so you do not bounce between two half-broken nodes.