diff options
| author | Christian Brauner <brauner@kernel.org> | 2026-09-17 13:49:04 +0200 |
|---|---|---|
| committer | Christian Brauner <brauner@kernel.org> | 2026-09-23 09:03:18 +0200 |
| commit | 2e2142a35d809c3cc726d3b39e7aeee98d9ab71c (patch) | |
| tree | 79326485830446d475b472ddbbb981c3912e3e0b /fs | |
| parent | 76ebb69da677d2a4c5e59bf758b6424d511f5684 (diff) | |
| download | linux-2e2142a35d809c3cc726d3b39e7aeee98d9ab71c.tar.gz linux-2e2142a35d809c3cc726d3b39e7aeee98d9ab71c.zip | |
Revert "put_mnt_ns(): leave mounts connected"
This reverts commit 0342482a4d15358fe6931606caf58968de5d1d38.
Keeping every mount of a dying namespace attached to its parent changes
who owns it. An attached but unmounted mount is owned by its parent: the
parent's final mntput() is what drops the child's reference and takes
its superblock down. Before that commit only the namespace root was in
that position and every other mount dropped its own reference in
namespace_unlock().
That ownership rule turns any reference from a child's superblock back
to one of its ancestors into a cycle. The obvious case is a loop device:
unshare -m sh -c 'mount -o loop /var/tmp/img /var/tmp/mp'
mount(8) opens the image from inside the new namespace, so the loop
device's backing file pins that namespace's copy of the root mount.
When the namespace dies the loop mount stays attached to that copy and
is owned by it. The copy can't reach zero because the backing file
holds it. The backing file is only put once the ext4 superblock
releases the block device and autoclear runs, and that needs the loop
mount to go first. Nothing breaks the cycle. The superblock, the loop
device and, since the root copy keeps all of its children.
systemd-sysext sets up its loop devices in a private mount namespace and
relies on autoclear when that namespace goes away, which is how this was
found. Anything that holds a file on an ancestor from a mounted
filesystem has the same problem. That includes ecryptfs lower paths,
erofs file-backed mounts, fuse passthrough backing files.
That's a bigger fix and not an -rc change. Revert.
Fixes: 0342482a4d15 ("put_mnt_ns(): leave mounts connected")
Reported-by: Michael Vogt <michael@amutable.com>
Link: https://gist.github.com/mvo5/63ef46482349f3b1c3957d463a0c9c6f
Link: https://patch.msgid.link/20260917-work-put_mnt_ns-revert-v1-2-34d9b8679e68@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
Diffstat (limited to 'fs')
| -rw-r--r-- | fs/namespace.c | 2 |
1 files changed, 1 insertions, 1 deletions
diff --git a/fs/namespace.c b/fs/namespace.c index 5e41021eaa63..580877e46b1a 100644 --- a/fs/namespace.c +++ b/fs/namespace.c @@ -6301,7 +6301,7 @@ void put_mnt_ns(struct mnt_namespace *ns) guard(namespace_excl)(); emptied_ns = ns; guard(mount_writer)(); - umount_tree(ns->root, UMOUNT_CONNECTED); + umount_tree(ns->root, 0); } struct vfsmount *kern_mount(struct file_system_type *type) |
