| Age | Commit message (Collapse) | Author |
|
https://git.kernel.org/pub/scm/linux/kernel/git/jgg/iommufd.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/nolibc/linux-nolibc.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/sysctl/sysctl.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/kees/linux.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/livepatching/livepatching.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/shuah/linux-kselftest.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup.git
|
|
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/gregkh/char-misc.git
# Conflicts:
# drivers/android/binder.c
# drivers/android/binder_alloc.c
# drivers/android/binderfs.c
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext.git
|
|
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/kvms390/linux.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/kvmarm/kvmarm.git
# Conflicts:
# arch/arm64/include/asm/ptrace.h
# arch/arm64/include/asm/sysreg.h
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/paulmck/linux-rcu.git
|
|
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git
# Conflicts:
# Documentation/scheduler/index.rst
# arch/arm64/configs/defconfig
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/pcmoore/lsm.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next.git
# Conflicts:
# arch/arm64/net/bpf_jit_comp.c
# arch/x86/net/bpf_jit_comp.c
# mm/internal.h
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next.git
# Conflicts:
# drivers/net/ethernet/realtek/r8169_main.c
# net/mac80211/ieee80211_i.h
# net/mac80211/tx.c
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/hid/hid.git
|
|
# Conflicts:
# fs/coredump.c
# fs/f2fs/f2fs.h
# fs/fuse/dax.c
# fs/xfs/libxfs/xfs_btree.c
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/kbuild/linux.git
# Conflicts:
# scripts/kallsyms.c
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/mm/linux.git
# Conflicts:
# arch/arm64/kvm/mmu.c
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/kvmarm/kvmarm.git
|
|
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf.git/
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/netdev/net.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs.git
# Conflicts:
# fs/smb/server/smb2pdu.c
# fs/smb/server/vfs.c
# fs/smb/server/vfs.h
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/mszeredi/fuse.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/leitao/linux.git
|
|
The large-chunk test always requests two base pages. On interfaces with
a large MTU, that may not exceed two maximum-sized frames.
Request a power-of-two buffer larger than twice the MTU. This makes the
existing rx_buf_len check and data-integrity traffic exercise the larger
layout.
Signed-off-by: Björn Töpel <bjorn@kernel.org>
Tested-by: Breno Leitao <leitao@debian.org>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260925104417.2325213-6-bjorn@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
ofdlocks reports five results but plans for four, so even when they
all pass it ends with
# Planned tests != run tests (4 != 5)
and exits with KSFT_FAIL, which run_kselftest.sh reports as a failure.
Link: https://lore.kernel.org/20260925205502.115327-1-danishkhateeb03@gmail.com
Fixes: 33d5b13098fb ("kselftest/filelock: report each test in oftlocks separately")
Signed-off-by: Danish Khateeb <danishkhateeb03@gmail.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Assisted-by: LLM
Cc: Shuah Khan <shuah@kernel.org>
Cc: Mark Brown <broonie@kernel.org>
Cc: Jeff Layton <jlayton@kernel.org>
|
|
The uevent_filtering test shrinks the uevent socket buffer to 4 KB
although the default socket buffer size is much higher. This leads to
this test being flaky when too many unrelated uevents are fired on the
machine. They might fill up the netlink receive buffer leading to ENOBUFS
errors when trying to receive the uevents. For example, I could trigger
test failures when running triggering a lot of udev events in the
background:
$ # run multiple of that in the background:
$ while :; do sudo udevadm trigger --action=change; done &
$ sudo ./uevent_filtering
# Starting 1 tests from 1 test cases.
# RUN global.uevent_filtering ...
add@/devices/virtual/mem/fullACTION=addDEVPATH=/devices/virtual/mem/fullSUBSYSTEM=memSYNTH_UUID=0MAJOR=1MINOR=7DEVNAME=fullDEVMODE=0666SEQNUM=304458
add@/devices/virtual/mem/fullACTION=addDEVPATH=/devices/virtual/mem/fullSUBSYSTEM=memSYNTH_UUID=0MAJOR=1MINOR=7DEVNAME=fullDEVMODE=0666SEQNUM=304471
add@/devices/virtual/mem/fullACTION=addDEVPATH=/devices/virtual/mem/fullSUBSYSTEM=memSYNTH_UUID=0MAJOR=1MINOR=7DEVNAME=fullDEVMODE=0666SEQNUM=304481
add@/devices/virtual/mem/fullACTION=addDEVPATH=/devices/virtual/mem/fullSUBSYSTEM=memSYNTH_UUID=0MAJOR=1MINOR=7DEVNAME=fullDEVMODE=0666SEQNUM=349156
No buffer space available - Failed to receive uevent
# uevent_filtering.c:463:uevent_filtering:Expected 0 (0) == ret (-1)
# uevent_filtering: Test failed
# FAIL global.uevent_filtering
not ok 1 global.uevent_filtering
The default receive buffer size (SK_RMEM_MAX) is far larger than the
requested 4 KB, so keep this to make the test less flaky.
Link: https://lore.kernel.org/20260619-get-swam-a1cd4cca@mheyne-amazon
Fixes: 9d3df886d17b ("selftests: uevent filtering")
Signed-off-by: Maximilian Heyne <mheyne@amazon.de>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Cc: Christian Brauner <christianvanbrauner@gmail.com>
Cc: David S. Miller <davem@davemloft.net>
Cc: Shuah Khan <shuah@kernel.org>
Cc: Wei Yang <richard.weiyang@gmail.com>
Cc: <stable@vger.kernel.org>
|
|
The test assumes fs.nr_open is close to the default 1048576, but some
systems set it much higher (e.g. 1073741816). This is systemd's doing:
since systemd v240 (2018), PID 1 bumps fs.nr_open and fs.file-max to their
largest possible values on boot, as file descriptors are already accounted
for by memcg [1].
In that case, dup2() to nr_open + 64 requires the kernel to allocate a
file descriptor table with ~1 billion entries, which fails with ENOMEM.
On a kernel that already carries 04a2c4b4511d1, dup2() no longer fails
with ENOMEM. The allocation is now rejected up front and the caller
gets EMFILE instead, without the WARNING, but the test still fails.
Cap the nr_open value used for the test's own arithmetic to a known
reasonable base value (1048576) and restore the true original value once
the test has completed.
Link: https://lore.kernel.org/20260814165709.513263-1-khorenko@virtuozzo.com
Link: https://github.com/systemd/systemd/commit/a8b627aaed409a15260c25988970c795bf963812 [1]
Signed-off-by: Konstantin Khorenko <khorenko@virtuozzo.com>
Signed-off-by: Eva Kurchatova <eva.kurchatova@virtuozzo.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Cc: Shuah Khan <shuah@kernel.org>
Cc: Wei Yang <richard.weiyang@gmail.com>
Cc: <stable@vger.kernel.org>
|
|
In tests with multiple concurrent waiters on edge-triggered epoll
instances where an emitter writes to multiple sockets (epoll16, epoll56,
epoll58):
When the emitter performs its first write(), ep_poll_callback() fires and
wakes up both waiters because one waiter uses epoll_wait() and the other
one uses poll(). This translates to different wait queues, ep->wq for
epoll and ep->poll_wait for poll/select, which are both awoken by the
kernel because of that single write. Next, both waiter threads invoke
epoll_wait(), but since there is only one event, only one epoll_wait()
will return non-zero because of the edge-triggered mode being used (in
level-triggered mode, the kernel would re-queue the event because of
remaining unread data).
Since the second waiter sees an empty ready list, it does not increment
ctx.count and the test fails spuriously with ctx.count == 1 instead of 2.
Emitter (CPU 0) Thread 0 (CPU 1) Thread 1 (CPU 2)
=============== ================ ================
epoll_wait(e0, -1) poll(e0, -1)
[on e0->wq] [on e0->poll_wait]
write(sfd[1])
|
+--(Kernel wakes BOTH e0->wq and e0->poll_wait via callback)--+
| |
| wakes up wakes up |
| epoll_wait() reaps e1 poll() returns 1 |
| (e1 removed via ET) (wants event) |
| e0->rdllist is EMPTY | |
| count++ (count = 1) v |
| epoll_wait(e0, 0) |
| sees EMPTY list! |
| returns 0! |
| thread exits |
v |
write(sfd[3]) |
(event arrives too late!) v
EXPECT_EQ(count, 2) <-- SPURIOUS FAILURE!
Introduce waiter_entry1ap_loop() to retry poll() if the initial
epoll_wait(..., 0) yielded no events. This ensures the thread waits for
the subsequent write rather than failing immediately. Apply this helper
in epoll16, epoll56, and for both waiter threads in epoll58.
Link: https://lore.kernel.org/20260828-selftest-epoll-fix-race-v2-1-953ab57fd60a@codasip.com
Fixes: f2728fe80cef ("selftests: add epoll selftests")
Signed-off-by: Florian Schmaus <florian.schmaus@codasip.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Cc: Heiher <r@hev.cc>
Cc: Roman Penyaev <rpenyaev@suse.de>
Cc: Shuah Khan <shuah@kernel.org>
Cc: Christian Brauner <brauner@kernel.org>
|
|
preregistered by libc
On thread creation, Musl registers the private expedited memory barrier,
see pthread_create [1]. Thus, invoking the barrier command will no longer
be rejected by the kernel with EPERM. The test checking this will fail.
Check if the memory barrier command has been registered and skip the test
in this case.
Link: https://git.musl-libc.org/cgit/musl/tree/src/thread/pthread_create.c#n260 [1]
Link: https://lore.kernel.org/20260803124900.3328789-3-christian.gellermann@codasip.com
Signed-off-by: Chris Gellermann <christian.gellermann@codasip.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Tested-by: Michael Jeanson <mjeanson@efficios.com>
Cc: Ben Segall <bsegall@google.com>
Cc: Dietmar Eggemann <dietmar.eggemann@arm.com>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Juri Lelli <juri.lelli@redhat.com>
Cc: K Prateek Nayak <kprateek.nayak@amd.com>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: Mel Gorman <mgorman@suse.de>
Cc: "Paul E . McKenney" <paulmck@kernel.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Shuah Khan <shuah@kernel.org>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Valentin Schneider <vschneid@redhat.com>
Cc: Vincent Guittot <vincent.guittot@linaro.org>
Cc: Wei Yang <richard.weiyang@gmail.com>
|
|
Patch series "selftests/membarrier: Skip an unregistered memory barrier
test on Musl".
The membarrier test "membarrier MEMBARRIER_CMD_PRIVATE_EXPEDITED not
registered failure" fails in the multithreaded test scenario when using
Musl libc as the command gets preregistered implicitly during thread
creation. Skip the test if command registration is detected.
This patch (of 2):
Add a new membarrier_get_registrations() for reusage.
Link: https://lore.kernel.org/20260803124900.3328789-1-christian.gellermann@codasip.com
Link: https://lore.kernel.org/20260803124900.3328789-2-christian.gellermann@codasip.com
Signed-off-by: Chris Gellermann <christian.gellermann@codasip.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Tested-by: Michael Jeanson <mjeanson@efficios.com>
Cc: Ben Segall <bsegall@google.com>
Cc: Dietmar Eggemann <dietmar.eggemann@arm.com>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Juri Lelli <juri.lelli@redhat.com>
Cc: K Prateek Nayak <kprateek.nayak@amd.com>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: Mel Gorman <mgorman@suse.de>
Cc: "Paul E . McKenney" <paulmck@kernel.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Shuah Khan <shuah@kernel.org>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Valentin Schneider <vschneid@redhat.com>
Cc: Vincent Guittot <vincent.guittot@linaro.org>
Cc: Wei Yang <richard.weiyang@gmail.com>
|
|
Test the existence and the valid input acceptance of the newly added
quota goal complement flag sysfs file.
Link: https://lore.kernel.org/20260929080113.41708-6-sj@kernel.org
Signed-off-by: SJ Park <sj@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Cc: Shuah Khan <shuah@kernel.org>
|
|
check_huge_shmem() was required to distinguish shmem huge pages because
/proc/self/smaps reports them using a dedicated “ShmemPmdMapped”
entry, as opposed to “FilePmdMapped” for file-backed huge pages.
Now that /proc/self/smaps is no longer used to detect huge pages and
/proc/kpageflags is used instead, it is sufficient to distinguish between
file-backed and anonymous pages since the ShmemPmdMapped is also kind of
file-backed.
Therefore, remove check_huge_shmem() and use check_huge_file() instead and
cleanup khugepaged's check_huge operation in mem_ops.
Link: https://lore.kernel.org/20260924-fix_split-v8-4-cba7359d882a@arm.com
Signed-off-by: Yeoreum Yun <yeoreum.yun@arm.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Suggested-by: David Hildenbrand (Arm) <david@kernel.org>
Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
Acked-by: Zi Yan <ziy@nvidia.com>
Acked-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Reviewed-by: Sarthak Sharma <sarthak.sharma@arm.com>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Nico Pache <nico.pache@linux.dev>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: Dev Jain <dev.jain@arm.com>
Cc: Barry Song <baohua@kernel.org>
Cc: Lance Yang <lance.yang@linux.dev>
Cc: Usama Arif <usama.arif@linux.dev>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Shuah Khan <shuah@kernel.org>
Cc: Kevin Brodsky <kevin.brodsky@arm.com>
|
|
check_large_folios() only checks for large folios without distinguishing
between anonymous and file-backed huge pages.
To add huge page type checking, integrate the huge page checks into
__check_huge():
1. If hpage_size == pmd_pagesize, check PAGE_IS_HUGE instead of using
check_large_folios(), since only the mapping type matters. This
identifies PMD-mapped huge pages.
2. Otherwise, use check_large_folios() to detect large folios. This
covers mTHP cases.
3. Check the folio flags according to the huge page type.
Link: https://lore.kernel.org/20260924-fix_split-v8-3-cba7359d882a@arm.com
Signed-off-by: Yeoreum Yun <yeoreum.yun@arm.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Suggested-by: David Hildenbrand (Arm) <david@kernel.org>
Suggested-by: Zi Yan <ziy@nvidia.com>
Reviewed-by: Sarthak Sharma <sarthak.sharma@arm.com>
Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
Tested-by: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: Lorenzo Stoakes <ljs@kernel.org>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Nico Pache <nico.pache@linux.dev>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: Dev Jain <dev.jain@arm.com>
Cc: Barry Song <baohua@kernel.org>
Cc: Lance Yang <lance.yang@linux.dev>
Cc: Usama Arif <usama.arif@linux.dev>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Shuah Khan <shuah@kernel.org>
Cc: Kevin Brodsky <kevin.brodsky@arm.com>
|
|
Since glibc commit 321e1fc73f (“malloc: Enable 2MB THP by default on
AArch64”), glibc may call madvise(MADV_HUGEPAGE) for sufficiently large
allocations made by memalign().
The underlying VMA may start at a different address from the aligned
address returned by memalign(). Furthermore, a subsequent
madvise(MADV_HUGEPAGE) call does not split the VMA because the flag is
already set.
This causes split_huge_page_test to fail because the check_pmd_huge()
helpers incorrectly require the address returned by memalign() to match
the VMA start address reported in /proc/self/smaps.
Instead of relying on /proc/self/smaps, use /proc/self/pagemap and
/proc/kpageflags to detect huge-page mappings checking PAGE_IS_HUGE and
PAGE_IS_FILE according to type of huge page.
Since shmem pages are also file-backed, simply check whether the page is
file-backed.
Link: https://lore.kernel.org/20260924-fix_split-v8-2-cba7359d882a@arm.com
Fixes: 642bc52aed9c ("selftests: vm: bring common functions to a new file")
Signed-off-by: Yeoreum Yun <yeoreum.yun@arm.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Suggested-by: David Hildenbrand (Arm) <david@kernel.org>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
Tested-by: Baolin Wang <baolin.wang@linux.alibaba.com>
Acked-by: Zi Yan <ziy@nvidia.com>
Cc: Lorenzo Stoakes <ljs@kernel.org>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Nico Pache <nico.pache@linux.dev>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: Dev Jain <dev.jain@arm.com>
Cc: Barry Song <baohua@kernel.org>
Cc: Lance Yang <lance.yang@linux.dev>
Cc: Usama Arif <usama.arif@linux.dev>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Shuah Khan <shuah@kernel.org>
Cc: Kevin Brodsky <kevin.brodsky@arm.com>
|
|
Patch series "kselftest: mm: fix some failure of split_huge_page_test", v8.
split_huge_page_test can fail for the following reasons:
1. During the test, khugepaged may collapse previously split pages again,
causing intermittent failures.
2. Since glibc commit 321e1fc73f (“malloc: Enable 2MB THP by default on AArch64”),
glibc may call madvise(MADV_HUGEPAGE) for sufficiently large allocations
made by memalign(). The underlying VMA may start at a different address
from the aligned address returned by memalign(). Moreover, a subsequent
madvise(MADV_HUGEPAGE) call does not split the VMA because it already
has the same advice.
This causes the test to fail because the check_huge_xxx() helpers
incorrectly require the address returned by memalign() to match the
VMA start address reported in /proc/self/smaps.
Address these issues by applying MADV_NOHUGEPAGE after faulting in the
huge page, preventing khugepaged from collapsing it again, and instead of
relying on /proc/self/smaps, use /proc/self/pagemap and /proc/kpageflags
to detect huge-page mappings and large folios:
1. If hpage_size == pmd_pagesize, check PAGE_IS_HUGE instead of
using check_large_folios(), since only the mapping type matters.
This identifies PMD-mapped huge pages.
2. Otherwise, use check_large_folios() to detect large folios. This
covers mTHP cases.
3. Check the folio flags according to the type of huge page.
Since check_huge_shmem() was required to distinguish shmem huge pages
because /proc/self/smaps reports them using a dedicated
“ShmemPmdMapped” entry, as opposed to “FilePmdMapped” for
file-backed huge pages.
Now that /proc/self/smaps is no longer used to detect huge pages and
/proc/kpageflags is used instead, it is sufficient to distinguish between
file-backed and anonymous pages since the ShmemPmdMapped is also kind of
file-backed.
Therefore, remove check_huge_shmem() and use check_huge_file() instead.
This patch (of 4):
There're some random failure for split_huge_page_test when khugepaged
collapses pages into pmd again which had split by the test.
Prevent the khugepaged's collapses for split page by setting the mapped
pmd-huge-page with MADV_NOHUGEPAGE before split.
Link: https://lore.kernel.org/20260924-fix_split-v8-0-cba7359d882a@arm.com
Link: https://lore.kernel.org/20260924-fix_split-v8-1-cba7359d882a@arm.com
Signed-off-by: Yeoreum Yun <yeoreum.yun@arm.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Suggested-by: Kevin Brodsky <kevin.brodsky@arm.com>
Suggested-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Reviewed-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Reviewed-by: Zi Yan <ziy@nvidia.com>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Reviewed-by: Sarthak Sharma <sarthak.sharma@arm.com>
Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Nico Pache <nico.pache@linux.dev>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: Dev Jain <dev.jain@arm.com>
Cc: Barry Song <baohua@kernel.org>
Cc: Lance Yang <lance.yang@linux.dev>
Cc: Usama Arif <usama.arif@linux.dev>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Shuah Khan <shuah@kernel.org>
|
|
While inspecting selftests/mm syscall wrappers, I noticed that mlock2_()
in mlock2.h handles the syscall return value differently from other
wrappers:
int ret = syscall(__NR_mlock2, start, len, flags);
if (ret) {
errno = ret;
return -1;
}
Commit 1ddae9d67ee1 ("selftests/mm/mlock: print error on failure")
introduced this intending to make mlock2_() behave like libc by setting
errno and returning -1. However, glibc syscall(2) already returns -1 on
failure and sets positive errno. Assigning "errno = ret;" overwrites
errno with -1.
To verify this, mlock2 was disabled in the kernel (via sys_ni_syscall) to
return -ENOSYS. Testing revealed two interrelated defects:
1. In the unmodified test, mlock2_() clobbered errno to -1. The check
"if (ret && errno == ENOSYS)" in main() was bypassed, resulting in an
immediate crash in the first test:
~ # ./mlock2-tests
TAP version 13
1..15
Bail out! mlock2(0): Unknown error -1
# Planned tests != run tests (15 != 0)
# Totals: pass:0 fail:0 xfail:0 xpass:0 skip:0 error:0
(exit code: 1 - FAIL)
2. After restoring mlock2_() to directly return syscall(), errno correctly
retained ENOSYS (38), entering the ENOSYS check in main(). However, it
then called ksft_finished():
~ # ./mlock2-tests
TAP version 13
# Totals: pass:0 fail:0 xfail:0 xpass:0 skip:0 error:0
~ # echo $?
0
Because ksft_set_plan() had not been called yet (ksft_plan == 0) and
zero tests ran (ksft_pass == 0), ksft_finished() evaluated 0 == 0 as
success and exited with KSFT_PASS (code 0) without any TAP skip header.
Fix both issues by:
1. Returning the syscall() result directly in mlock2_() so that errno is
preserved.
2. Calling ksft_exit_skip() on ENOSYS so unsupported kernels report a TAP
skip ("1..0 # SKIP ...") and exit with KSFT_SKIP (code 4).
Verification on the mlock2-disabled kernel:
~ # ./mlock2-tests
TAP version 13
1..0 # SKIP mlock2() syscall is not supported
~ # echo $?
4
Re-enabling mlock2 in the kernel confirmed all 15 tests pass cleanly:
~ # ./mlock2-tests
TAP version 13
1..15
ok 1 test_mlock_lock: Locked
...
ok 15 test_mlockall_future_droppable: droppable memory not locked
# Totals: pass:15 fail:0 xfail:0 xpass:0 skip:0 error:0
~ # echo $?
0
Link: https://lore.kernel.org/20260923-selftests-mm-mlock2-fix-v1-1-750b627854c6@dgu.ac.kr
Fixes: 1ddae9d67ee1 ("selftests/mm/mlock: print error on failure")
Fixes: 65c89684896d ("selftests/mm: mlock2-tests: conform test to TAP format output")
Signed-off-by: Park Tae-sun <ts930@dgu.ac.kr>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reviewed-by: Gregory Price <gourry@gourry.net>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Reviewed-by: Muhammad Usama Anjum <usama.anjum@arm.com>
Cc: Lorenzo Stoakes <ljs@kernel.org>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Shuah Khan <shuah@kernel.org>
Cc: Brendan Jackman <brendan.jackman@linux.dev>
|
|
There are intermittent failures in collapse_max_ptes_swap() and
collapse_max_ptes_shared() when using the khugepaged_context:
// while running ./khugepaged -s 2
# Run test: collapse_max_ptes_shared (khugepaged:anon)
# Allocate huge page... OK
# Share huge page over fork()... OK
# Trigger CoW on page 1023 of 2048... OK
# Maybe collapse with max_ptes_shared exceeded.... OK
# Trigger CoW on page 1024 of 2048... Fail
Bail out! Unexpected huge page
# Planned tests != run tests (26 != 23)
# Totals: pass:23 fail:0 xfail:0 xpass:0 skip:0 error:0
# Run test: collapse_max_ptes_swap (khugepaged:anon)
# Swapout 257 of 2048 pages... OK
# Maybe collapse with max_ptes_swap exceeded.... OK
# Swapout 256 of 2048 pages... OK
Bail out! Unexpected huge page
# Planned tests != run tests (26 != 17)
# Totals: pass:17 fail:0 xfail:0 xpass:0 skip:0 error:0
This happens because khugepaged may collapse the pages before
wait_for_scan() is called, causing a sanity check that expects uncollapsed
pages to fail.
For example, in collapse_max_ptes_swap(), after faulting the pages back in
and paging out up to max_ptes_swap pages, khugepaged may collapse them
again before c->collapse() is called.
To prevent this, mark the VMA with MADV_NOHUGEPAGE after it has been
collapsed by wait_for_scan() for anon. This prevents khugepaged from
collapsing it again before c->collapse() is called.
This failure was observed on NVIDIA Spark with 16KB page.
Link: https://lore.kernel.org/20260929-fix_khugepagd_fail-v4-2-2169c18f2576@arm.com
Signed-off-by: Yeoreum Yun <yeoreum.yun@arm.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reviewed-by: Gregory Price (Meta) <gourry@gourry.net>
Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
Tested-by: Baolin Wang <baolin.wang@linux.alibaba.com>
Acked-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Cc: Zi Yan <ziy@nvidia.com>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Nico Pache <nico.pache@linux.dev>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: Dev Jain <dev.jain@arm.com>
Cc: Barry Song <baohua@kernel.org>
Cc: Lance Yang <lance.yang@linux.dev>
Cc: Usama Arif <usama.arif@linux.dev>
Cc: Kiryl Shutsemau <kas@kernel.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: David Hildenbrand <david@kernel.org>
Cc: Shuah Khan <shuah@kernel.org>
|
|
Patch series "kselftest: mm: fix intermittent failure khugepaged test", v4.
There are intermittent failures in collapse_max_ptes_swap() and
collapse_max_ptes_shared() when using the khugepaged_context:
# Run test: collapse_max_ptes_shared (khugepaged:anon)
# Allocate huge page... OK
# Share huge page over fork()... OK
# Trigger CoW on page 1023 of 2048... OK
# Maybe collapse with max_ptes_shared exceeded.... OK
# Trigger CoW on page 1024 of 2048... Fail
Bail out! Unexpected huge page
# Planned tests != run tests (26 != 23)
# Totals: pass:23 fail:0 xfail:0 xpass:0 skip:0 error:0
# Run test: collapse_max_ptes_swap (khugepaged:anon)
# Swapout 257 of 2048 pages... OK
# Maybe collapse with max_ptes_swap exceeded.... OK
# Swapout 256 of 2048 pages... OK
Bail out! Unexpected huge page
# Planned tests != run tests (26 != 17)
# Totals: pass:17 fail:0 xfail:0 xpass:0 skip:0 error:0
This happens because khugepaged may collapse the pages before
wait_for_scan() is called, causing a sanity check that expects uncollapsed
pages to fail.
For example, in collapse_max_ptes_swap(), after faulting the pages back in
and paging out up to max_ptes_swap pages, khugepaged may collapse them
again before c->collapse() is called.
To prevent this, mark the VMA with MADV_NOHUGEPAGE after it has been
collapsed by wait_for_scan() for anon. This prevents khugepaged from
collapsing it again before c->collapse() is called.
Also, fix false-positive results when a child process fails in tests such
as collapse_fork*() or collapse_max_ptes_shared():
# -------------------------
# running ./khugepaged -s 2
# -------------------------
#
# Run test: collapse_max_ptes_shared (khugepaged:anon)
# Allocate huge page... OK
# Share huge page over fork()... OK
# Trigger CoW on page 1023 of 2048... OK
# Maybe collapse with max_ptes_shared exceeded.... OK
# Trigger CoW on page 1024 of 2048... Fail
Bail out! Unexpected huge page
# Planned tests != run tests (26 != 23)
# Totals: pass:23 fail:0 xfail:0 xpass:0 skip:0 error:0 // child failed.
# Check if parent still has huge page... OK // parent hpage success
ok 24 collapse_max_ptes_shared // considered as success
...
# Totals: pass:26 fail:0 xfail:0 xpass:0 skip:0 error:0
This failure was observed on NVIDIA Spark with 16KB page.
This patch (of 2):
Although the child process in collapse_fork*() or
collapse_max_ptes_shared() reports `KSFT_FAIL`, the result is ignored
because the test only checks whether the parent's page was collapsed into
a huge page.
As a result, the test is considered successful whenever the parent's page
is a huge page, even if the child test fails, as shown below:
#
# Run test: collapse_max_ptes_shared (khugepaged:anon)
# Allocate huge page... OK
# Share huge page over fork()... OK
# Trigger CoW on page 1023 of 2048... OK
# Maybe collapse with max_ptes_shared exceeded.... OK
# Trigger CoW on page 1024 of 2048... Fail
Bail out! Unexpected huge page
# Planned tests != run tests (26 != 23)
# Totals: pass:23 fail:0 xfail:0 xpass:0 skip:0 error:0 // child failed.
# Check if parent still has huge page... OK // parent hpage success
ok 24 collapse_max_ptes_shared // considered as success
...
# Totals: pass:26 fail:0 xfail:0 xpass:0 skip:0 error:0
To address this, propagate the child's failure and skip the subsequent
check in the parent.
Link: https://lore.kernel.org/20260929-fix_khugepagd_fail-v4-0-2169c18f2576@arm.com
Link: https://lore.kernel.org/20260929-fix_khugepagd_fail-v4-1-2169c18f2576@arm.com
Signed-off-by: Yeoreum Yun <yeoreum.yun@arm.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
Reviewed-by: Gregory Price (Meta) <gourry@gourry.net>
Reviewed-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Acked-by: Zi Yan <ziy@nvidia.com>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Nico Pache <nico.pache@linux.dev>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: Dev Jain <dev.jain@arm.com>
Cc: Barry Song <baohua@kernel.org>
Cc: Lance Yang <lance.yang@linux.dev>
Cc: Usama Arif <usama.arif@linux.dev>
Cc: Kiryl Shutsemau <kas@kernel.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Shuah Khan <shuah@kernel.org>
|
|
The harness's MADV_DONTNEED thread zaps 1 to 32 pages at a time, never a
whole PMD-aligned area, and only a zap that covers a full table frees the
table itself (CONFIG_PT_RECLAIM).
Make the thread zap a whole PMD-aligned area about one iteration in 64,
and keep the fine-grained zaps as the common case. The new case frees
page tables, racing that against a collapse walking the same table.
Link: https://lore.kernel.org/20260919002451.496763-20-kirill@shutemov.name
Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Tested-by: Muhammad Usama Anjum <usama.anjum@arm.com>
Assisted-by: LLM
Cc: Alexander Gordeev <agordeev@linux.ibm.com>
Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: Barry Song <baohua@kernel.org>
Cc: "David Hildenbrand (arm)" <david@kernel.org>
Cc: Dev Jain <dev.jain@arm.com>
Cc: Hugh Dickens <hughd@google.com>
Cc: Jason Gunthorpe <jgg@ziepe.ca>
Cc: kernel-team@meta.com
Cc: Lance Yang <lance.yang@linux.dev>
Cc: Leon Romanovsky <leon@kernel.org>
Cc: Liam Howlett <liam@infradead.org>
Cc: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Mike Rapoport (Microsoft) <rppt@kernel.org>
Cc: Nico Pache (Red Hat) <nico.pache@linux.dev>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: Shuah Khan (Samsung OSG) <shuah@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Usama Arif <usama.arif@linux.dev>
Cc: "Vlastimil Babka (SUSE)" <vbabka@kernel.org>
Cc: Zi Yan <ziy@nvidia.com>
|