diff options
| author | Mikhail Gavrilov <mikhail.v.gavrilov@gmail.com> | 2026-09-24 03:31:16 +0500 |
|---|---|---|
| committer | Dave Hansen <dave.hansen@linux.intel.com> | 2026-09-28 08:08:16 -0700 |
| commit | e3ee38c1bc0b11a8d3e63dce7050759b61864286 (patch) | |
| tree | 1d5b1ebb2a3a5d60500ce3916e4aad1bada93f44 /arch/x86/include/uapi/asm | |
| parent | d2457a7e2727dc747a61a50d65c2f259afce8c45 (diff) | |
| download | linux-next-e3ee38c1bc0b11a8d3e63dce7050759b61864286.tar.gz linux-next-e3ee38c1bc0b11a8d3e63dce7050759b61864286.zip | |
x86/mm: Drop unnecessary PMD page copy when freeing
On a box with a discrete GPU, lockdep reports a possible deadlock as
soon as kswapd shrinks the TTM page pool. The immediate cause is an
x86 commit that added an mmap_read_lock() to kernel page protection
munging code.
The huge vmap code holds the same lock over a GFP_KERNEL allocation,
which is a no-no now that reclaim can take it. That allocation is in a
page table *free* path and ends up being for dubious purposes[1].
Basically, it tries to avoid hardware setting Accessed=1 in page table
entries that are unreachable by the hardware, a non-issue.
Remove the PMD copy. Detach the original PMD page at the PUD, flush
the mid-level caches, and free the PTE tables straight from the
detached PMD page. With no allocation left, the locking issue is gone.
Lockdep splat/analysis:
WARNING: possible circular locking dependency detected
7.3.0-rc3-f6e7b42bf05b+ #183 Tainted: G U
------------------------------------------------------
kswapd0/269 is trying to acquire lock:
((init_mm).mmap_lock){++++}-{4:4}, at: change_page_attr_set_clr+0x29a/0x4a0
but task is already holding lock:
(pool_shrink_rwsem){.+.+}-{4:4}, at: ttm_pool_shrink+0xb2/0x330 [ttm]
Chain exists of:
(init_mm).mmap_lock --> fs_reclaim --> pool_shrink_rwsem
The cycle is built from three edges:
1) pool_shrink_rwsem -> (init_mm).mmap_lock
The TTM shrinker restores the caching attribute of every page it
frees, while holding pool_shrink_rwsem:
ttm_pool_shrink()
-> ttm_pool_dispose_list()
-> ttm_pool_free_page()
-> set_pages_wb()
-> change_page_attr_set_clr() [ init_mm mmap read lock ]
2) fs_reclaim -> pool_shrink_rwsem
The same shrinker, called from reclaim.
3) (init_mm).mmap_lock -> fs_reclaim
ioremap() installing a huge PUD mapping over an existing PMD table:
ioremap_page_range()
-> vmap_range_noflush()
-> vmap_try_huge_pud() [ init_mm mmap read lock ]
-> pud_free_pmd_page()
-> __get_free_page(GFP_KERNEL) [ enters reclaim ]
[ dhansen: Lots of changelog munging/trimming and merged comments from my
version of the fix. ]
Fixes: d5d8b8662e6e ("x86/mm/pat: Acquire init_mm read lock on attribute changes to avoid UAF")
Suggested-by: Pedro Falcato <pfalcato@suse.de>
Signed-off-by: Mikhail Gavrilov <mikhail.v.gavrilov@gmail.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Pedro Falcato <pfalcato@suse.de>
Link: https://lore.kernel.org/20260916062222.27347-1-mikhail.v.gavrilov@gmail.com
Link: https://lore.kernel.org/all/e11449f0-d9ad-4d1b-ab21-2be7d71fe335@intel.com/ [1]
Link: https://patch.msgid.link/20260923223116.20090-1-mikhail.v.gavrilov@gmail.com
Cc: stable@vger.kernel.org
Diffstat (limited to 'arch/x86/include/uapi/asm')
0 files changed, 0 insertions, 0 deletions
