Skip to content
All articles

Containers · Kubernetes

A container is a process, not a small VM

A container is an ordinary Linux process with namespaces and cgroups around it. That one fact explains the shared kernel, the PID 1 shutdown bug and the heap sized for the whole node.

· 4 min read

Three layers around a container: namespaces restrict what it can see (its own processes, network, mounts and hostname), cgroups restrict what it can use (CPU time, memory, IO), and the kernel is not restricted, shared with the host and every container. Below, who receives SIGTERM on a deploy: with CMD ./start.sh the shell is PID 1 and many shells don't forward it, so the app gets SIGKILL after the grace period; with exec java -jar app.jar, or under tini or --init, the app is PID 1 and its shutdown hook runs. A note: older runtimes report the host's memory, so a heap sized for 64 GB inside a 512 MB container gets killed.

There is no guest kernel inside a container, and nothing is virtualised. A container is an ordinary Linux process on the host; run ps on the host and you will see it. What makes it a container is two kernel features wrapped around that process:

  • Namespaces restrict what it can see: its own process tree, network stack, mounts and hostname.
  • cgroups restrict what it can use: CPU, memory and IO.

Picturing it as a small virtual machine is where most container surprises come from. Three of them turn up in support queues again and again.

The kernel is shared

Every container on a node runs on the host's kernel. A kernel vulnerability is therefore a host vulnerability, and a container escape is a real category of incident rather than a theoretical one. Containers are a convenient isolation boundary, but not one strong enough for running untrusted code from strangers. That needs a VM, Firecracker or gVisor.

Your process is PID 1

Inside its PID namespace your application is process 1, and PID 1 plays by different rules: it gets no default signal handlers, and it must reap its zombie children. Whether a deploy shuts your service down cleanly depends on which process ends up in that slot.

Container commandPID 1 isOn SIGTERM
CMD ["java","-jar","app.jar"]The JVMThe shutdown handler runs
CMD ./start.shThe shellMany shells don't forward it, so SIGKILL arrives after the grace period
Run with tini or --initAn init processForwarded, and zombie children are reaped

This is the mechanism behind a very common bug. The grace period is set correctly, the application implements shutdown correctly, and pods are still killed hard on every deploy, because a shell sits between the signal and the code.

Think Like an Engineer

The exec form CMD ["./start.sh"] removes the wrapping shell, but the script still runs under its own interpreter, and that interpreter becomes PID 1. If the script ends with exec java -jar app.jar, the JVM replaces it and receives the signal. Without that one exec, nothing has changed.

Memory limits are visible only if you look

Older runtimes report the host's memory through the usual interfaces. A JVM or Node process then sizes its heap for a 64 GB machine while it lives in a 512 MB container, and the kernel's OOM killer ends it with a SIGKILL that no shutdown code can catch. Modern JVMs are container-aware by default. A lot of tuning advice, and a lot of running code, is not: anything that sizes a pool or a cache from "available system memory" will size it for the node.

The tell is a container that dies at a limit far below what its configuration implies.

Get one diagram a week

A short article built around one engineering diagram, from the same library as these courses.

One diagram-led article a week on AI and systems engineering. We email you once to confirm, and every newsletter has an unsubscribe link. Privacy policy