🐳 Your Container Keeps Getting OOM-Killed – While the Host Sits There With Memory to Spare
A container dies with an out-of-memory error, and `free -h` on the host shows plenty of memory available – which seems like a contradiction, until you realize Docker enforces a MEMORY LIMIT scoped to the container itself (via cgroups), entirely independent of what the host machine has free; the container is killed for exceeding ITS OWN ceiling, not the host’s.
🔎 The Problem
$ docker run -m 512m my-app # The container is capped at 512MB regardless of how much RAM the host # actually has - if the application inside genuinely needs more than # 512MB under real load (a memory-hungry request, a growing in-memory # cache), the kernel'"'"'s OOM killer terminates the container'"'"'s process # the moment it crosses that limit - "Exited (137)" - even while # `free -h` on the host itself shows tens of gigabytes still free.
✅ Fix: Diagnose the Actual Memory Need, Then Set the Limit to Match
- `docker stats` while the app is under realistic load shows the container’s actual memory usage over time – if it’s consistently approaching the limit before being killed, that’s confirmation the limit itself (not a leak) is the immediate cause, and it needs to be raised to match real usage.
- If usage keeps climbing without ever leveling off, that’s a genuine memory leak inside the application, and raising the limit only delays the same crash – profile the app’s memory usage directly (heap dumps, a memory profiler for the runtime in use) rather than treating a higher limit as the fix.
- For JVM, Node, and similar managed runtimes specifically, make sure the RUNTIME’s own memory settings (`-Xmx` for Java, `–max-old-space-size` for Node) are set BELOW the container’s memory limit – a runtime that doesn’t know about the container ceiling can allocate right up to what it thinks the whole host has, only to get killed by Docker before the runtime’s own garbage collector even reacts.
⚠️ Reading Exit Code 137
- Exit code 137 specifically means the container process received SIGKILL (128 + signal 9) – it’s the reliable signature of an OOM kill (or a manual `docker kill`), distinct from an application crash that would show a different, application-specific exit code.
A container’s memory limit isn’t a suggestion tied to the host’s free memory – it’s a hard ceiling the kernel enforces regardless, and ‘the host has plenty free’ was never the number that mattered.
