Conversation
Kubernetes does not give privileged containers a cgroup namespace on cgroup v2 (KEP-2254), so in a privileged dind pod /sys/fs/cgroup is the node's root: "dind" sets up its nesting there, moving whatever sits in the node's root cgroup into /init, and dockerd creates its containers in /docker/<id> next to kubepods.slice, outside the pod's cgroup and so outside its limits and accounting (moby/moby#45378). When the container's cgroup is not the root of its namespace, enter a new cgroup and mount namespace with unshare and re-mount /sys/fs/cgroup so it is rooted there, the same way moby/buildkit#6368 does for the buildkit image. The mount options are carried over because cgroup2 applies them to the whole hierarchy. If that is not possible, warn and carry on as before. Signed-off-by: Jason Lernerman <jasonlernerman@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the Kubernetes case in moby/moby#45378.
What's wrong
On cgroup v2, Kubernetes deliberately does not give privileged containers a cgroup namespace (KEP-2254; containerd's CRI skips it for
privileged: true). So in a privilegeddocker:dindpod,/sys/fs/cgroupis the node's root cgroup:dindsets up its cgroup nesting there: it moves whatever sits in the node's rootcgroup.procsinto/initand writes the node's rootcgroup.subtree_control.dockerdcreates its containers in/docker/<id>, next tokubepods.slice, so every container a job starts is outside the pod's cgroup: no pod limits apply to it, it gets the root's default weight against all ofkubepods.slice, andkubectl top/the metrics server don't see it.docker run --privilegedisn't affected because Docker gives the container a private cgroup namespace on cgroup v2, which is why this only shows up on Kubernetes.The fix
If cgroup v2 is in use and the container's cgroup isn't the root of its namespace, re-exec through
unshare --cgroup --mountand re-mount/sys/fs/cgroupso it is rooted at the container's cgroup. After thatdindanddockerdsee exactly what they see underdocker run --privileged, anddockerd's/docker/<id>lands inside the pod's cgroup. This is the same approach (and the same check) themoby/buildkitimage uses since moby/buildkit#6368, and what people in the moby issue have been doing by hand. Ifunsharecan't do it (not privileged enough), a warning is printed and startup continues as before.unshare --cgroupneeds util-linux'sunshare(busybox's has no--cgroup), henceutil-linux-misc; it adds about 3 MB to the image.How it was tested
On an ubuntu-24.04 GitHub runner (cgroup v2, systemd driver), with
--cgroupns=hostto get what Kubernetes does:Before, from the host:
After:
(the outer
--memory 256mnow reaches the inner container; before,tail /dev/zeroescaped it.)docker run -d --privileged dind-test(private cgroup namespace) is unchanged:dockerdis at0::/initanddocker runworks. With only--cap-add SYS_ADMIN, startup goes on to fail where it failed before.