Server

Out of memory

Also called OOM.

Out of memory is the state where a process or machine has asked for more RAM than is available, and the kernel or runtime kills something. The kill is abrupt, with no clean shutdown.

How it is measured

Look for the evidence, which is in the kernel log (`dmesg | grep -i 'killed process'`), container exit code 137, or a runtime error such as 'JavaScript heap out of memory'. Record the time, the process, and memory use just before.

Compare against the limits: cgroup limit for a container, total RAM for a VM, and the heap flag for the runtime. A container killed at 512 MB can sit on a host with plenty of free memory.

Worked example

A WooCommerce import on a 1 GB VPS reads a 60 MB CSV into an array. PHP-FPM workers each hold 120 MB, there are 8 of them, and MariaDB takes 300 MB. At the 3rd import the kernel's OOM killer ends MariaDB, the largest process, and every page returns a database error.

Capping PHP-FPM to 4 workers, streaming the CSV and adding 1 GB swap as a safety net keeps the import running with MariaDB untouched.

How it differs

Out of memory is the kill that happens when RAM runs out. A memory leak is one cause: a process that grew without bound. Swap is the buffer that delays the kill by using disk. OOM excludes the slow decline that precedes it, and the leak excludes legitimate peaks, such as a big import that simply needs more room.

Common errors

Restarting the service and moving on without reading dmesg. Assuming the app that died is the one that used the memory. Sizing for average load. Giving the runtime a heap larger than the container. Disabling the OOM killer and freezing the host.

In practice

Add a memory graph and an alert at 85 percent, check your kernel log for past kills, and set per-service limits so the thing that dies is the least important one.

See also

Memory leak, Swap

Sources

Count this on a real site.

Watch my website