Skip to content

dind: run dockerd in its own cgroup namespace on cgroup v2 - #587

Open
u9g wants to merge 1 commit into
docker-library:masterfrom
u9g:dind-cgroupns
Open

u9g wants to merge 1 commit into
docker-library:masterfrom
u9g:dind-cgroupns

Conversation

@u9g

@u9g u9g commented Sep 25, 2026

Copy link
Copy Markdown

Fixes the Kubernetes case in moby/moby#45378.

What's wrong

On cgroup v2, Kubernetes deliberately does not give privileged containers a cgroup namespace (KEP-2254; containerd's CRI skips it for privileged: true). So in a privileged docker:dind pod, /sys/fs/cgroup is the node's root cgroup:

  • dind sets up its cgroup nesting there: it moves whatever sits in the node's root cgroup.procs into /init and writes the node's root cgroup.subtree_control.
  • dockerd creates its containers in /docker/<id>, next to kubepods.slice, so every container a job starts is outside the pod's cgroup: no pod limits apply to it, it gets the root's default weight against all of kubepods.slice, and kubectl top/the metrics server don't see it.

docker run --privileged isn't affected because Docker gives the container a private cgroup namespace on cgroup v2, which is why this only shows up on Kubernetes.

The fix

If cgroup v2 is in use and the container's cgroup isn't the root of its namespace, re-exec through unshare --cgroup --mount and re-mount /sys/fs/cgroup so it is rooted at the container's cgroup. After that dind and dockerd see exactly what they see under docker run --privileged, and dockerd's /docker/<id> lands inside the pod's cgroup. This is the same approach (and the same check) the moby/buildkit image uses since moby/buildkit#6368, and what people in the moby issue have been doing by hand. If unshare can't do it (not privileged enough), a warning is printed and startup continues as before.

unshare --cgroup needs util-linux's unshare (busybox's has no --cgroup), hence util-linux-misc; it adds about 3 MB to the image.

How it was tested

On an ubuntu-24.04 GitHub runner (cgroup v2, systemd driver), with --cgroupns=host to get what Kubernetes does:

$ docker run -d --name k8s --privileged --cgroupns=host --memory 256m -e DOCKER_TLS_CERTDIR= dind-test
$ docker exec k8s docker run -d alpine sleep 300

Before, from the host:

/proc/<sleep pid>/cgroup: 0::/docker/34922192e98d...

After:

/proc/<dockerd pid>/cgroup: 0::/system.slice/docker-0d0980098692....scope/init
/proc/<sleep pid>/cgroup:   0::/system.slice/docker-0d0980098692....scope/docker/58d02c8935ab...
$ docker exec k8s docker run --rm alpine tail /dev/zero; echo $?
137

(the outer --memory 256m now reaches the inner container; before, tail /dev/zero escaped it.)

docker run -d --privileged dind-test (private cgroup namespace) is unchanged: dockerd is at 0::/init and docker run works. With only --cap-add SYS_ADMIN, startup goes on to fail where it failed before.

Kubernetes does not give privileged containers a cgroup namespace on
cgroup v2 (KEP-2254), so in a privileged dind pod /sys/fs/cgroup is the
node's root: "dind" sets up its nesting there, moving whatever sits in
the node's root cgroup into /init, and dockerd creates its containers in
/docker/<id> next to kubepods.slice, outside the pod's cgroup and so
outside its limits and accounting (moby/moby#45378).

When the container's cgroup is not the root of its namespace, enter a
new cgroup and mount namespace with unshare and re-mount /sys/fs/cgroup
so it is rooted there, the same way moby/buildkit#6368 does for the
buildkit image. The mount options are carried over because cgroup2
applies them to the whole hierarchy. If that is not possible, warn and
carry on as before.

Signed-off-by: Jason Lernerman <jasonlernerman@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant