summaryrefslogtreecommitdiff
path: root/tools/include
AgeCommit message (Collapse)Author
25 hoursMerge branch 'for-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/nolibc/linux-nolibc.git
25 hoursMerge branch 'bitmap-for-next' of https://github.com/norov/linux.gitMark Brown
25 hoursMerge branch 'next' of https://github.com/awilliam/linux-vfio.gitMark Brown
25 hoursMerge branch 'for-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next.git # Conflicts: # arch/arm64/net/bpf_jit_comp.c # arch/x86/net/bpf_jit_comp.c # mm/internal.h
25 hoursMerge branch 'main' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next.git # Conflicts: # net/mac80211/ieee80211_i.h # net/mac80211/tx.c
26 hoursMerge branch 'fs-next' of linux-nextMark Brown
# Conflicts: # fs/coredump.c # fs/f2fs/f2fs.h # fs/fuse/dax.c # fs/xfs/libxfs/xfs_btree.c
26 hoursMerge branch 'perf-tools-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/perf/perf-tools-next.git
30 hoursMerge https://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm.git ↵David Hildenbrand (Arm)
mm-unstable into for-next Signed-off-by: David Hildenbrand (Arm) <david@kernel.org>
31 hourscompiler.h: add ASSERT_STATIC_STORAGE()Yury Norov
Patch series "Catch automatic storage in IDA and Maple Tree definitions". A 0day report [1] from the region allocation benchmark exposed a lockdep initialization bug: a stack-local Maple Tree used MTREE_INIT(), whose embedded lock has a static initializer. On the first allocation, lockdep rejected the lock address as a non-static class key and disabled locking validation. The IDA benchmark had the same issue, masked because it ran after Maple Tree had already disabled lockdep. The fix [2] switches the test to using mt_init_flags() and ida_init(). This series adds a compile-time check to the related DEFINE_IDA() and DEFINE_MTREE() declaration macros to catch the same class of mistake earlier. Patch 1 introduces ASSERT_STATIC_STORAGE(). It declares an unused static pointer initialized with the object's address, requiring that address to be a valid static initializer. Patches 2 and 3 apply the helper to IDA and Maple Tree definitions, respectively. The helper is mirrored in the tools compiler header. The existing automatic local IDAs and Maple Trees in the userspace radix-tree tests are converted to runtime initialization. The interval-tree span test keeps its existing mt_init_flags() call and uses a plain Maple Tree declaration. For example, an automatic local definition: void example(void) { DEFINE_IDA(ida); ida_destroy(&ida); } now produces: error: initializer element is not constant note: in expansion of macro 'ASSERT_STATIC_STORAGE' note: in expansion of macro 'DEFINE_IDA' File-scope definitions and static local definitions remain valid. Automatic local objects should use ida_init(), mt_init(), or mt_init_flags(). The check is limited to declaration macros. Direct uses of IDA_INIT(), MTREE_INIT(), and MTREE_INIT_EXT() remain unchanged. The helper cannot be inserted directly into those initializer expressions because it expands to a declaration. Validated by GCC and Clang checks accepting static storage and rejecting automatic storage The userspace IDR/IDA and Maple Tree test are passed as well. This patch (of 3): Static lock initializers rely on a persistent object address when lockdep assigns a lock-class key. Using such an initializer for an automatic local object can compile successfully but disable lockdep on the first lock acquisition. Add ASSERT_STATIC_STORAGE() for declaration macros that require static storage duration. It declares an unused static pointer initialized with the object's address. An automatic local object's address is not a valid static initializer, so the compiler rejects it. Mirror the helper in tools/include/linux/compiler.h because the userspace radix-tree tests include the kernel IDA and Maple Tree headers with the tools compiler definitions. For example: void example(void) { int object; ASSERT_STATIC_STORAGE(object); } GCC reports: error: initializer element is not constant name##_storage_check = &(name) ^ note: in expansion of macro 'ASSERT_STATIC_STORAGE' ASSERT_STATIC_STORAGE(object); File-scope objects and static local objects remain valid. The helper takes an object identifier and must be used as a declaration after that object has been declared. Link: https://lore.kernel.org/all/20260911155244.1406122-1-ynorov@nvidia.com/ Link: https://lore.kernel.org/20260911221444.1523311-2-ynorov@nvidia.com Link: https://download.01.org/0day-ci/archive/20260910/202609101106.771b567e-lkp@intel.com/ [1] Link: https://lore.kernel.org/all/20260911155244.1406122-1-ynorov@nvidia.com/ [2] Signed-off-by: Yury Norov <ynorov@nvidia.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Cc: Alice Ryhl <aliceryhl@google.com> Cc: Andrew Ballance <andrewjballance@gmail.com> Cc: Christopher Li <sparse@chrisli.org> Cc: Liam R. Howlett <liam@infradead.org> Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
31 hoursmm: remove unused totalram_pages_inc() and totalram_pages_dec()Tal Zussman
totalram_pages_inc() and totalram_pages_dec() have had no callers since commit 7fbc5e26123e ("memblock: extract page freeing from free_reserved_area() into a helper") and commit 287b89773d81 ("powerpc/pseries/cmm: Use adjust_managed_page_count() insted of totalram_pages_*"), respectively. Remove them. Drop the totalram_pages_inc() stub from tools mm.h too. Link: https://lore.kernel.org/20260901-mm-remove-unused-helpers-v2-2-f6474e169c23@columbia.edu Reviewed-by: Lorenzo Stoakes (ARM) <ljs@kernel.org> Reviewed-by: Mike Rapoport (Microsoft) <rppt@kernel.org> Signed-off-by: Tal Zussman <tz2294@columbia.edu> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Acked-by: David Hildenbrand (Arm) <david@kernel.org> Cc: Liam R. Howlett <liam@infradead.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Vlastimil Babka <vbabka@kernel.org>
3 daysselftests: Add additional kernel functions to tools/include/Jason Gunthorpe
These are needed by the VFIO mlx5 selftest in the following patches, which includes some headers from mlx5 and also needs a few more MMIO-related features. - DECLARE_FLEX_ARRAY in new tools/include/linux/stddef.h (wraps existing __DECLARE_FLEX_ARRAY from uapi/linux/stddef.h) - dma_wmb/dma_rmb barriers: x86 uses compiler barrier (DMA-coherent), arm64 uses dmb oshst/oshld (outer-shareable for device visibility), generic fallback uses wmb/rmb - ioread32be/iowrite32be in tools/include/asm-generic/io.h for big-endian MMIO register access Assisted-by: Claude:claude-opus-4.6 Reviewed-by: David Matlack <dmatlack@google.com> Signed-off-by: Jason Gunthorpe <jgg@nvidia.com> Link: https://lore.kernel.org/r/5-v7-c6d30e8ce1e4+3dfa6-mlx5st_jgg@nvidia.com Signed-off-by: Alex Williamson <alex@shazbot.org>
4 daystools/nolibc: add support for pread() and pwrite()Thomas Weißschuh
These are standard function to interact with a file descriptor without touching its internal file offset. Signed-off-by: Thomas Weißschuh <linux@weissschuh.net> Reviewed-by: Willy Tarreau <w@1wt.eu> Link: https://patch.msgid.link/20260927-nolibc-pread-v1-5-e9af6e47b37a@weissschuh.net
4 daystools/nolibc: introduce __NOLIBC_BITS_PER_SYSCALL_ARGThomas Weißschuh
ILP32 architectures, like x32 and n32, support 64-bit system call arguments, although they are nominally 32-bit architectures. Add a new define which provides the size of system call arguments so generic libc code can use the best system call for those. Signed-off-by: Thomas Weißschuh <linux@weissschuh.net> Reviewed-by: Willy Tarreau <w@1wt.eu> Link: https://patch.msgid.link/20260927-nolibc-pread-v1-4-e9af6e47b37a@weissschuh.net
4 daystools/nolibc: move __NOLIBC_LLARGPART() to compiler.hThomas Weißschuh
pread() and pwrite() support will require the usage of the macro from the arch-*.h headers. sys.h can not be used from those. Move the macro to a more generic header. Signed-off-by: Thomas Weißschuh <linux@weissschuh.net> Reviewed-by: Willy Tarreau <w@1wt.eu> Link: https://patch.msgid.link/20260927-nolibc-pread-v1-3-e9af6e47b37a@weissschuh.net
4 daystools/nolibc: make 64-bit systemcall argument padding genericThomas Weißschuh
Several 32-bit architectures pass a 64-bit systemcall argument as two 32-bit values requiring alignment to an even register pair. This case is currently relevant for the ftruncate64() wrapper and handled by those architectures providing custom wrappers. With the introduction of pread() and pwrite() wrappers, the same pattern would lead to a lot of repeated code. Add a #define, so that the generic code can do the padding where necessary, without any per-architecture duplication. Signed-off-by: Thomas Weißschuh <linux@weissschuh.net> Reviewed-by: Willy Tarreau <w@1wt.eu> Link: https://patch.msgid.link/20260927-nolibc-pread-v1-2-e9af6e47b37a@weissschuh.net
5 daystools/nolibc: add 64-bit parisc supportThomas Weißschuh
Extend the existing parisc support to also handle 64-bit mode. Some peculiarities: * The nonstandard ftruncate64 system call requires a workaround. * The linker defaults to main() as entry point, which needs to be switched to _start() explicitly. * The linker defaults to dynamic linking, which needs to be disabled explicitly. * With the default -Os, gcc does not use 64-bit multiplication instructions, requiring libgcc. Use -O2 instead. * There is no qemu-user available. * CONFIG_COMPAT needs to be disabled to avoid build errors due to a missing 32-bit vDSO compiler. Signed-off-by: Thomas Weißschuh <linux@weissschuh.net> Link: https://patch.msgid.link/20260914-nolibc-parisc64-v3-1-f14b320117c1@weissschuh.net
5 daystools/nolibc: fix FD_SET(), FD_CLR() and FD_ISSET() on 64-bitDanish Khateeb
Since commit feaf75658783 ("nolibc: fix fd_set type"), an fd_set is an array of unsigned long, but the FD_* macros still build their masks as an int or an unsigned int: 1 << n in FD_SET(), 1U << n in FD_CLR() and FD_ISSET(), where n is the fd modulo 64 on 64-bit architectures. So: - FD_SET() of an fd with n == 31 sets the bits for n = 31-63, as 1 << 31 is negative and gets sign-extended. - For n >= 32 the shifts are undefined. x86-64, for instance, masks the shift count, so FD_SET(40) and FD_ISSET(40) use the bit of fd 8. - FD_CLR() zero-extends ~(1U << n), so it also clears the bits for n = 32-63, whichever fd it is given. select() on fd 40, for instance, fails with EBADF, or watches fd 8 instead if that one is open. Build the masks from 1UL. Fixes: feaf75658783 ("nolibc: fix fd_set type") Assisted-by: LLM Signed-off-by: Danish Khateeb <danishkhateeb03@gmail.com> Link: https://patch.msgid.link/20260925211600.119116-3-danishkhateeb03@gmail.com Signed-off-by: Thomas Weißschuh <linux@weissschuh.net>
5 daystools/nolibc: fix readdir_r() with 64-bit directory offsetsDanish Khateeb
readdir_r() stores the result of _sys_lseek() in an int. Directory offsets are opaque cookies which can use all 64 bits: ext4, for instance, gives 64-bit processes 63-bit hashes, and 0x7fffffffffffffff as the offset after the last entry. Truncated to an int, such an offset is negative about half the time, and readdir_r() takes it for an error. Since commit 4ada5679f18d ("tools/nolibc/dirent: avoid errno in readdir_r"), readdir_r() fails at the first entry whose offset has bit 31 set, which on ext4 is usually one of the first few, and returns the truncated offset as the error number. Before that, only -1 counted as an error, which the last entry always hits: readdir_r() then returned errno, usually 0, without filling in the entry, so the caller got the previous entry a second time and never saw the last one. 32-bit processes get 31-bit hashes from ext4 and are not affected. Keep the offset in an off_t. Fixes: 665fa8dea90d ("tools/nolibc: add support for directory access") Assisted-by: LLM Signed-off-by: Danish Khateeb <danishkhateeb03@gmail.com> Link: https://patch.msgid.link/20260925211600.119116-2-danishkhateeb03@gmail.com Signed-off-by: Thomas Weißschuh <linux@weissschuh.net>
6 daysbpf: Add file descriptor interface for program streamsKumar Kartikeya Dwivedi
The existing BPF_PROG_STREAM_READ_BY_FD command only supports polling a program stream through repeated bpf() calls. It cannot block for new data or integrate with poll-based event loops. Add BPF_PROG_STREAM_OPEN to return a read-only, close-on-exec file descriptor for a selected program stream. Reads block by default and BPF_F_STREAM_NONBLOCK, the only accepted flag, provides non-blocking behavior. poll reports readable data and reports hangup once the program has been freed. Like pipes and sockets, the descriptor is not seekable and lseek fails with ESPIPE. A stream descriptor deliberately does not retain the program. Move each stream into a separately refcounted allocation so program teardown can mark it dead and wake descriptor users while outstanding descriptors drain buffered data safely. Readers sample the dead flag before looking for data, so EOF is reported only when the stream was already dead before it was found empty; data published right before teardown is never skipped. Only programs loaded through BPF_PROG_LOAD get streams. Classic BPF filters, JIT subprograms and shim programs never write to one, and kernel-side writers already resolve a subprogram to its main program, so those programs no longer carry stream state. Readiness needs its own counter. Stream capacity is charged before allocation and before an element is published to the stream log, so using that reservation as the read and poll condition can report readable data while no element exists: a blocking reader retries instead of sleeping and a lone non-blocking reader can see POLLIN followed by EAGAIN. Publish bytes with release ordering after adding elements to the lockless log, use acquire loads before consuming them or reporting readiness, limit each read to its readable snapshot and subtract only bytes actually copied. This keeps the aggregate count correct even when concurrent publishers update it out of publication order. With several readers on one stream, readiness remains advisory, as it is for pipes. The capacity counter is kept solely for enforcing the stream size limit. Wakeups are always deferred through irq_work. Stream writers run in whatever context the program runs in: NMI context for perf_event programs, sections with interrupts disabled inside bpf_spin_lock or rqspinlock critical sections since bpf_stream_vprintk() is KF_SPINLOCK_SAFE, and tracing programs attached anywhere in the kernel, including inside the wait queue and epoll code itself. Waking waiters directly from there can deadlock, and no cheap context check covers every case: on PREEMPT_RT, spinlock_t sections do not disable interrupts, so in_nmi() or irqs_disabled() cannot tell such a program apart from a benign one. Queue an irq_work item instead, as bpf_ringbuf does. Queue it only when a publication turns an empty stream readable. Readers block and pollers wait only after finding the stream empty, and the readable count never drops below zero because each read is bounded by its snapshot, so the first publication after such an observation is the one that makes the count positive, and it is the one that queues the wakeup. Publications into a stream that already holds data raise no interrupt, so a program that prints while nobody drains its stream pays for a single irq_work until the stream is emptied again. This matches bpf_ringbuf, which notifies only once the consumer has caught up. Blocking readers and level-triggered pollers re-check the readable count before waiting, so they cannot miss data, and edge-triggered epoll consumers drain until EAGAIN before waiting again, as epoll(7) requires. Synchronize pending work before releasing the final stream reference so the callback cannot outlive the stream, but only when the work was ever queued: irq_work_sync() waits for an RCU grace period on PREEMPT_RT and on architectures without an irq_work interrupt. Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260925045536.1480933-3-memxor@gmail.com
6 daystools include: Add missing stdarg.h includeIan Rogers
Header file makes use of varargs but doesn't directly include stdarg.h. Closes: https://lore.kernel.org/linux-perf-users/20260916062233.192271F000FF@smtp.kernel.org/ Reported-by: sashiko-bot@kernel.org Signed-off-by: Ian Rogers <irogers@google.com> Acked-by: Namhyung Kim <namhyung@kernel.org> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
7 daystools/nolibc: fix verrx() and errx() message formattingDanish Khateeb
verrx() passes its va_list to warnx() instead of vwarnx(). warnx() is variadic, so it takes the va_list as its first and only argument, and the conversions in the format string print the va_list itself and then whatever is left in the following argument slots. On x86_64, for instance, errx(1, "int args: %d %d %d", 1, 2, 3); prints something like "int args: -1699009032 1 2", where the first number comes from the address of the va_list and changes from run to run, and errx(1, "%d %s", 1234, "foo"); crashes with SIGSEGV, as "%s" takes a leftover argument slot for the string. errx() goes through verrx(), so every message with arguments printed by either of them is wrong. Call vwarnx() from verrx(), as verr() calls vwarn(). Fixes: 9da0f529c089 ("tools/nolibc: add err.h") Assisted-by: LLM Signed-off-by: Danish Khateeb <danishkhateeb03@gmail.com> Link: https://patch.msgid.link/20260924120958.23698-2-danishkhateeb03@gmail.com Signed-off-by: Thomas Weißschuh <linux@weissschuh.net>
7 daysbtf: Extend UAPI to support BTF location (inline site) infoAlan Maguire
Add BTF_KIND_LOC_PARAM, BTF_KIND_LOC_PROTO and BTF_KIND_LOCSEC to help represent location information for functions. BTF_KIND_LOC_PARAM is used to represent how we retrieve data at a location; either via register(s), or register+offset, a dereference of a register+offset or a constant value. BTF_KIND_LOC_PROTO represents location information about a location with multiple BTF_KIND_LOC_PARAMs. And finally BTF_KIND_LOCSEC is a set of location sites, each of which has - a BTF_KIND_FUNC function associated with the inline site - a location prototype specifying where to find the function parameters - an address offset relative to the kernel base address This can be used to support representing - a fully-inlined function at potentially multiple inline sites with potentially different parameter availability - a partially-inlined function where some _LOC_PROTOs represent inlined sites as above and others have normal _FUNC representations Also BTF_KIND_LOCSEC struct btf_loc will have two type id references; one for the associated func, the other for the loc_proto. Accordingly increase the number of m_offs references in btf_field_desc to 2. Signed-off-by: Alan Maguire <alan.maguire@oracle.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Acked-by: Eduard Zingerman <eddyz87@gmail.com> Link: https://patch.msgid.link/20260924111428.75957-2-alan.maguire@oracle.com
7 daystools/nolibc: fix WIFSIGNALED() for status 0Danish Khateeb
WIFSIGNALED() needs the subtraction to wrap around for status 0, so that only 1 to 0xff, a terminating signal with or without the core dump flag, pass the check. But status is an int, and 0 - 1 is -1, which is less than 0xff. A child that exits with 0 is therefore reported as both exited and killed, and a caller that tests WIFSIGNALED() first sees "killed by signal 0". Make the subtraction unsigned, as musl does in the same macro. Fixes: 8c934d4822c7 ("tools/nolibc: add helpers for wait() signal exits") Assisted-by: LLM Signed-off-by: Danish Khateeb <danishkhateeb03@gmail.com> Link: https://patch.msgid.link/20260923205622.1123270-2-danishkhateeb03@gmail.com [Thomas: Drop stable@ Cc] Signed-off-by: Thomas Weißschuh <linux@weissschuh.net>
7 daystools: use __pa() in virt_to_phys()Tianyi Chen
With 32-bit pointers and a 64-bit phys_addr_t, GCC warns about the pointer-to-integer cast in virt_to_phys(). It can also sign-extend pointers whose top bit is set, turning 0x80000000 into 0xffffffff80000000. Use __pa() to convert through unsigned long before widening the result. This preserves the unsigned pointer value and matches the other address conversion helpers. Reported-by: Mike Rapoport <rppt@kernel.org> Link: https://lore.kernel.org/r/179007180279.508753.4595268983388589280.b4-review@b4 Assisted-by: LLM Signed-off-by: Tianyi Chen <hi@tychen.cc> Link: https://patch.msgid.link/20260923023126.1910311-4-hi@tychen.cc Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
9 daysbpf: Add bpf_call_rcu() kfuncPuranjay Mohan
BPF programs that manage their own objects have no way to run their own logic once an RCU grace period has elapsed. bpf_obj_drop() defers a free, but returning an index to an allocator or unpinning a resource once readers are done has no equivalent. sched_ext's BPF library works around this today by pushing freed nodes onto a list and having a userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a BPF program to reclaim them. Add: int bpf_call_rcu(struct bpf_rcu_head *rh, void *map, int (*callback)(struct bpf_map *map, void *key, void *value)); @rh is a struct bpf_rcu_head embedded in a value of @map, so the callback runs as callback(map, key, value) for the element it lives in and needs no cookie. A head can only be armed once, which bounds outstanding work by the number of elements. struct bpf_rcu_head holds the callback state inline rather than a pointer to it, as bpf_timer, bpf_wq and bpf_task_work do, because there is nothing to cancel and so nothing that has to outlive the map value. That avoids an allocation and a state machine on the arming path at the cost of 48 bytes per element. An RCU callback cannot be cancelled, so everything it touches has to stay alive until it runs: - The callback is the program's text, so arming takes a program reference as bpf_timer, bpf_wq and bpf_task_work do, dropped once the callback returns. bpf_prog_inc_not_zero() also fails the arm with -EBADF once the program is dying. - The map is held by that reference through used_maps. An inner map is not, so bpf_rcu_head is rejected in one. - The field is only accepted in BPF_MAP_TYPE_ARRAY, whose elements are never freed individually. A hash element can be deleted and recycled while a callback is queued on it. - The head is disarmed before the callback runs so it can be armed again from there, which takes a new program reference before the running callback drops its own. Arming therefore fails with -EPERM once the map is held by neither a process nor bpffs. bpf_iter hands a program a writable pointer to the live element, which would let it overwrite a queued head, so bpf_iter_attach_map() rejects maps carrying one. The callback is verified non-sleepable even when the caller is sleepable, and RCU invokes it with BH disabled. Signed-off-by: Puranjay Mohan <puranjay@kernel.org> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260922200208.3203834-2-puranjay@kernel.org
9 daystools: sync coredump.h headerChristian Brauner
Sync the headers for the selftests. Link: https://patch.msgid.link/20260821-work-coredump-filter-v1-2-91f9a73ef03e@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
9 daystools: sync coredump.h headerChristian Brauner
Sync the headers for the selftests. Link: https://patch.msgid.link/20260820-work-coredump-sparse-v2-15-ba32dd718c51@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
10 daysbitfield: get rid of __MAKE_OP machineryYury Norov
The __MAKE_OP machinery hides the fixed-width bitfield helper definitions from source searches and makes the end result highly obscured and largely uncontrolled. This follows earlier discussions about making these helpers easier to find: [1], [2]. Move the explicit helpers and their shared checks into linux/bitfield-fix-width.h in both the kernel and tools headers, and include it from bitfield.h to preserve existing users. The repeated overflow check is factored into __assert_field(), preserving the original condition and compile-time diagnostics. With GCC 15.2.0 and x86-64 defconfig plus the bitfield KUnit tests, the before/after kernel builds are binary identical. The __MAKE_OP generates the following 40 functions (including the direct ____MAKE_OP(u8,u8,,) invocation): u8_encode_bits() u8_replace_bits() u8p_replace_bits() u8_get_bits() le16_encode_bits() le16_replace_bits() [dead code] le16p_replace_bits() le16_get_bits() be16_encode_bits() be16_replace_bits() [dead code] be16p_replace_bits() [dead code] be16_get_bits() u16_encode_bits() u16_replace_bits() u16p_replace_bits() u16_get_bits() le32_encode_bits() le32_replace_bits() [dead code] le32p_replace_bits() le32_get_bits() be32_encode_bits() be32_replace_bits() [dead code] be32p_replace_bits() be32_get_bits() u32_encode_bits() u32_replace_bits() u32p_replace_bits() u32_get_bits() le64_encode_bits() le64_replace_bits() [dead code] le64p_replace_bits() [dead code] le64_get_bits() be64_encode_bits() be64_replace_bits() [dead code] be64p_replace_bits() [dead code] be64_get_bits() u64_encode_bits() u64_replace_bits() u64p_replace_bits() u64_get_bits() Functions marked with [dead code] have no in-tree callers and are removed by this change. Link: https://lore.kernel.org/all/20250214073402.0129e259@kernel.org/ [1] Link: https://lore.kernel.org/all/aeub59FBHbCy-KKP@yury/ [2] Assisted-by: OpenAI Codex Reviewed-by: Akihiko Odaki <odaki@rsg.ci.i.u-tokyo.ac.jp> Acked-by: Jakub Kicinski <kuba@kernel.org> Signed-off-by: Yury Norov <ynorov@nvidia.com>
12 daystools/nolibc/parisc: consistently use %dp to refer to register %r27Thomas Weißschuh
The parisc initialization code inconsistently uses both %dp and %r27 to refer to the data pointer register. Consistently use the %dp memnonic. Suggested-by: Helge Deller <deller@gmx.de> Link: https://lore.kernel.org/lkml/495ef25e-0bac-4f98-b0da-ae0d96dcea59@gmx.de/ Signed-off-by: Thomas Weißschuh <linux@weissschuh.net> Link: https://patch.msgid.link/20260914-nolibc-parisc-dp-register-v1-1-6388f4aa8fbe@weissschuh.net
12 daysbpf: tcp: Support parse/len/write header option hooks in bpf_tcp_opsAmery Hung
Add the TCP header option callbacks to the bpf_tcp_ops struct_ops type: parse_hdr - parse the options of an incoming skb on an established connection hdr_opt_len - reserve space in the TCP header for bpf options write_hdr_opt - write the reserved bpf options These mirror the BPF_SOCK_OPS_PARSE_HDR_OPT_CB, _HDR_OPT_LEN_CB and _WRITE_HDR_OPT_CB legacy sockops callbacks, but are exposed as struct_ops members so a program can implement them with normal function signatures and per-member helper sets. The reserved header window is shared between the legacy sockops and bpf_tcp_ops paths. tcp_{syn,synack,established}_options() first run the legacy BPF_SOCK_OPS_HDR_OPT_LEN_CB and then call hdr_opt_len, so both sources accumulate into opts->bpf_opt_len; at write time the legacy options are emitted first and bpf_tcp_ops writes after them. API design bpf_tcp_ops overloads the sock_ops header-option helpers rather than introducing a new API: bpf_reserve_hdr_opt(), bpf_store_hdr_opt() and bpf_load_hdr_opt() are exposed per-member (reserve for hdr_opt_len, store/load for write_hdr_opt, load for parse_hdr) and share the existing kernel option-walking core via _bpf_sock_ops{store,load}hdr_opt(), with the bpf_tcp_ops wrappers synthesizing a temporary bpf_sock_ops_kern from the program ctx. This keeps a port from the legacy BPF_SOCK_OPS*_HDR_OPT_CB callbacks mechanical (same helper calls) and adds no new UAPI helper/kfunc surface. An alternative considered was to drop the option helpers entirely: have hdr_opt_len reserve space purely through its return value, and introduce a dedicated TCP-header-option dynptr used for both reading and writing. That is a cleaner, more self-contained interface, but it is a larger change and does not reuse the legacy helpers, making a port from sockops less mechanical. It can be pursued as a follow-up; the helper-based interface here keeps this series focused on moving the hooks to struct_ops. The hdr_opt_len fast path in tcp_established_options() is gated by cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS). Note this is a global, per-attach-type static branch: it is enabled whenever any bpf_tcp_ops is attached, even one that does not implement hdr_opt_len or that is attached to a different cgroup. In those cases the block still runs but bpf_tcp_ops_hdr_opt_len() no-ops via the per-member check in the dispatch macro. A per-member/per-cgroup gate could be added later if the extra fast-path work proves measurable. Signed-off-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com> Link: https://patch.msgid.link/20260917200542.3689605-13-ameryhung@gmail.com
12 daysbpf: Add infrastructure to support attaching struct_ops to cgroupsMartin KaFai Lau
This patch adds necessary infrastructure to attach a struct_ops map to a cgroup. The initial need was to support migrating the legacy BPF_PROG_TYPE_SOCK_OPS to a struct_ops. Recently, there are other struct_ops use cases that need to attach struct_ops to a cgroup. For example, the recent BPF OOM and memcg discussion in LSFMMBPF 2026. The motivation is to create a consistent expectation for attaching struct_ops to cgroup instead of each subsystem creating its own infrastructure. This logic includes hierarchy expectation, ordering expectation, attachment API, and rcu gp. There is already an existing implementation for attaching multiple bpf progs to a cgroup. There are also tools built around it for querying. Attaching a struct_ops map (which is a group of bpf programs) could also adhere to a similar API and potentially reuse most of the existing implementation. A couple of ideas have been tried. One of them is to use mprog.c. In terms of the amount of changes, I eventually came to the same conclusion as in commit 120933984460 ("bpf: Implement mprog API on top of existing cgroup progs"). I then shifted the focus to reusing the current {update,compute,activate,purge}_effective_progs() which has the main logic that implements the mprog API. Since then, I tried to add a 'struct cgroup *cgroup' member to the existing 'struct bpf_struct_ops_link' and link_create will create a 'struct bpf_struct_ops_link' object to be stored in the pl->link. This turns out to have more changes on both cgroup.c and bpf_struct_ops.c than I like. This patch directly reuses the 'struct bpf_cgroup_link' which cgroup.c already understands. Add 'struct bpf_map *map' to 'struct bpf_cgroup_link'. In the future, as more subsystems are extended by struct_ops, we may consider to make 'struct bpf_map *map' as a primary citizen of a link like 'struct bpf_prog *prog' and directly add 'struct bpf_map *map' to the generic 'struct bpf_link'. The pl->link could be the traditional 'prog' link or the new 'map' link. The places that need to handle them differently have already been refactored into the new prog_list_*() added in the earlier patch. In those new prog_list_*(), this patch will check "pl->link && pl->link->map", learn that it is a 'map' link and handle it correctly. The bpf_prog_array also needs to handle that its item can store the traditional 'prog' or it can store a struct_ops map. The places that need to handle them differently have also been refactored into the new bpf_cgroup_array_*() added in the earlier patch. The two differences are: - different sentinel (dummy_bpf_prog in prog vs cfi_stub in struct_ops) - the array for struct_ops may need to go through different rcu gp. The bpf_cgroup_array_*() functions use the cgroup_bpf_attach_type (ie atype) to distinguish the array is storing prog or storing struct_ops map. This patch also implements a separate struct bpf_link_ops "cgroup_struct_ops_link_ops" to have a separate link_ops implementation that only handles the cgroup's struct_ops link. Questions: - Although this patch did not change it, it is not obvious to me how the replace_effective_progs() and purge_effective_progs() handle cases when there are existing BPF_F_PREORDER progs attached in the hlist. Misc notes: - CGROUP_TCP_SOCK_OPS is added to the 'enum cgroup_bpf_attach_type'. The actual implementation of the tcp_bpf_ops (a struct_ops) will be added in the next patch. - free_after_mult_rcu_gp is added to 'struct bpf_struct_ops' such that the bpf_prog_array can have a mix of sleepable and non-sleepable prog in a struct_ops. This can tell how the bpf_prog_array should be freed. - For a struct_ops that supports cgroup attachment, it does not need to implement its own reg/unreg function. reg/unreg to a cgroup is done by the common infrastructure added in this patch. - The cgroup's struct_ops link only supports BPF_F_ALLOW_MULTI. This is enforced internally in cgroup_bpf_struct_ops_attach. This should be consistent with the current prog's link behavior in cgroup_bpf_link_attach. In the future, we may allow each subsystem to choose differently. - A cgroup_atype member is added to 'struct bpf_struct_ops'. When a subsystem struct_ops needs to support cgroup attachment, it needs to add a value to 'enum cgroup_bpf_attach_type' and then assign it to the newly added cgroup_atype member in the bpf_struct_ops. - During LINK_CREATE in syscall, the patch uses the same BPF_STRUCT_OPS (in attr->link_create.attach_type). The bpf_struct_ops_link_create learns the map and from the map it learns the st_ops. If the st_ops->cgroup_atype is not 0, it will create a cgroup's link. - When a subsystem registers a struct_ops that supports cgroup attachment, the struct_ops infrastructure will also ask the cgroup infrastructure to remember a few things. This is done by calling cgroup_bpf_struct_ops_register(). Signed-off-by: Martin KaFai Lau <martin.lau@kernel.org> Signed-off-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260917200542.3689605-10-ameryhung@gmail.com
14 daysnetlink: specs: netdev: update doc for xsk tx-checksum and zc-max-segsJakub Kicinski
Minor touch ups to the XSK-related doc strings in the netdev spec. Tx checksum is (obviously?) Layer 4 (TCP/UDP), not Layer 3 (IP). The xdp_zc_max_segs change is more subtle, I guess, but to kernel devs saying "frags" may sound like the number of skb frags. The value includes the head in this case. The default of 1 means single buffer, XDP_USE_SG sockets are refused when it is 1. Let's use the word "buffer" instead of "frag". Link: https://patch.msgid.link/20260916021608.1702731-1-kuba@kernel.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-14perf headers: Sync perf_event.h/perf_regs.h with the kernel headersDapeng Mi
Sync the UAPI header changes of supporting SIMD/eGPRs/SSP sampling into corresponding tools UAPI headers. Additionally, support the new introduced perf_event_attr fields in the perf_event_attr__fprintf and perf_event__attr_swap() helpers, and add sanity check for the new introduced __reserved_4 field in perf_attr_check(). Co-developed-by: Kan Liang <kan.liang@linux.intel.com> Signed-off-by: Kan Liang <kan.liang@linux.intel.com> Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-09-14tools/nolibc: verify that a directory is openedThomas Weißschuh
If a non-directory is opened, ENODIR should be returned from opendir()/fdopendir() right away and not only during readdir_r(). Validate the type of opened file during open. Signed-off-by: Thomas Weißschuh <linux@weissschuh.net> Reviewed-by: Willy Tarreau <w@1wt.eu> Link: https://patch.msgid.link/20260831-nolibc-fdopendir-enotdir-v1-1-8cf0e79c4e6f@weissschuh.net
2026-09-14tools/nolibc: validate directory with O_DIRECTORY in opendir()Thomas Weißschuh
Switch to O_DIRECTORY to validate that the opened file is a directory. ALso remove the call to fdopendir() in opendir() to avoid a double check of 'fd'. This will also make it easier to add a similar check for fdopendir(). Signed-off-by: Thomas Weißschuh <linux@weissschuh.net> Reviewed-by: Willy Tarreau <w@1wt.eu> Link: https://patch.msgid.link/20260831-nolibc-fdopendir-enotdir-v1-3-8cf0e79c4e6f@weissschuh.net
2026-08-31tools/nolibc: Add sendfile()Daniel Palmer
This is useful because it allows copying data up to 2GB from one fd to another without a buffer in userspace and with a single syscall. For 32bit kernels we use the sendfile64 syscall to allow for 64bit offsets to work. Signed-off-by: Daniel Palmer <daniel@thingy.jp> Link: https://patch.msgid.link/20260828103205.384589-2-daniel@thingy.jp Signed-off-by: Thomas Weißschuh <linux@weissschuh.net>
2026-08-31tools/nolibc: add support for hexagonThomas Weißschuh
A mostly straightforward new architecture with a few quirks: * Only qemu-user is supported for testing. * Clang is required for compilation. * -fsanitize=undefined is broken. Signed-off-by: Thomas Weißschuh <linux@weissschuh.net> Acked-by: Willy Tarreau <w@1wt.eu> Reviewed-by: Brian Cain <brian.cain@oss.qualcomm.com> Link: https://patch.msgid.link/20260819-nolibc-hexagon-v1-3-6bc3be591f09@weissschuh.net
2026-08-31tools/nolibc: split the architecture list into multiple linesThomas Weißschuh
The list is getting long, split it up for easier additions. Signed-off-by: Thomas Weißschuh <linux@weissschuh.net> Acked-by: Willy Tarreau <w@1wt.eu> Reviewed-by: Brian Cain <brian.cain@oss.qualcomm.com> Link: https://patch.msgid.link/20260819-nolibc-hexagon-v1-1-6bc3be591f09@weissschuh.net
2026-08-27Merge tag 'net-7.3-rc1' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net Pull networking fixes from Jakub Kicinski: "Including fixes from Bluetooth, IPSec and Netfilter. Current release - fix to a fix: - netfilter: ipset: remove need to allocate memory on delete operations Current release - regressions: - macb: drop CONFIG_OF #if block, fix build Previous releases - always broken: - stream of fixes for SCTP continues - inet: frags: strip GSO state from fragments before reassembly - virtio-net: ensure that TCP packets don't overflow gso_segs - tcp-ao: fix use-after-free of current_key on reconnect to another peer - page_pool: remove zone/policy GFP flags when allocating XArray entries - Bluetooth: L2CAP: reject accept queue add unless BT_LISTEN - tls: device: fix out-of-bounds write in tls_append_frag() - eth: bnxt: - ring the doorbell when SW USO exits early, avoid packets stuck in Tx - gate TPH enablement behind BNXT_SUPPORTS_QUEUE_API check, avoid users of older NICs seeing non-actionable warning messages - eth: qede: fix NULL pointer dereference in TPA fragment processing" * tag 'net-7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (216 commits) inet: frags: strip GSO state from fragments before reassembly net/sched: sch_htb: limit htb_classify inner-class filter hops selftests/net: packetdrill: add tcp_urg_ptr_retransmit tcp: fix corruption of urgent data on multi-segment retransmit usb: atm: usbatm: fix invalid ci_range initialization net: fec: only stop PTP if it was initialized slip: remove slip_hangup() to fix use-after-free in slip_receive_buf() net: bridge: mcast: fix use-after-free of a master VLAN's multicast context net/sched: bound qdisc_pkt_len to prevent qdisc soft lockup net: dsa: mxl862xx: enable assisted learning on CPU port net: stmmac: restore NET_IP_ALIGN in the RX DMA offset net: stmmac: drop gso_enabled_types and rely on netdev features net: stmmac: selftests: Don't test flow control for small rx fifos net: stmmac: selftests: Account for the UC filter list for filtering tests net: stmmac: dwxgmac: Account for the primary MAC address for UC filtering net: stmmac: dwmac4: Account for the primary MAC address for UC filtering net: stmmac: dwmac1000: Account for the primary MAC address for UC filtering net: stmmac: selftests: Check multiple MMC counters selftests: net: Fix slow configurations in big_tcp_tunnels.sh selftests: net: Lower threshold with csum offload off in big_tcp_tunnels.sh ...
2026-08-24xsk: align TX metadata layout across ABIsStanislav Fomichev
Add explicit padding before launch_time so xsk_tx_metadata has the same layout on 32-bit and 64-bit systems. On several architectures (csky, i386, nios2, m65k, openrisc, sh), the old native 32-bit layout put launch_time at offset 12 and had a natural size of 20 bytes. Using sizeof(struct xsk_tx_metadata) as tx_metadata_len was already rejected because the length must be a multiple of eight, so the straightforward use of the interface was broken on those ABIs. Userspace could still register a padded length of 24 bytes, though; mixing the old and new layouts then silently reads launch_time from the wrong offset and misprograms packet launch times. This intentionally replaces that incompatible layout because the affected architectures are unlikely to have any notable users. (x86_64 and arm64 have the most users and are _not_ affected) Fixes: ca4419f15abd ("xsk: Add launch time hardware offload support to XDP Tx metadata") Reviewed-by: Simon Horman <horms@kernel.org> Signed-off-by: Stanislav Fomichev <sdf@fomichev.me> Link: https://patch.msgid.link/20260819160535.1472459-2-sdf@fomichev.me Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-23Merge tag 'mm-nonmm-stable-2026-08-22-16-57' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm Pull non-MM updates from Andrew Morton: - "ocfs2/dlm: bound peer-controlled lengths in the o2dlm" (Bryam Vargas) Validate and bound all input lengths and count fields in the o2dlm migration and recovery receive handlers to prevent memory corruption and kernel panics from malformed cluster messages - "ocfs2: validate xattr entry bounds" (Cen Zhang) Validate OCFS2 extended attribute entry name and value bounds during metadata reads to prevent out-of-range memory accesses during retrieval or listing operations. - "taskstats: fix cgroupstats invalid fd handling and add selftests" (Yiyang Chen) Return -EBADF when cgroupstats receives an invalid file descriptor to prevent caller hangs and misleading success ACKs. Add a kselftest to validate valid cgroup v1 queries and verify proper error handling across different Netlink flag combinations. - "misc lib/raid/ improvements v2" (Christoph Hellwig) Improve benchmark-based algorithm selection for the XOR and RAID6 libraries, add KUnit benchmark tests, and cleanup minor implementation details. - "ocfs2: cluster: o2hb_region_pin() fixes" (Joseph Qi) Fix sleeping-in-atomic, lock order inversion and error-path cleanup bugs in o2hb_region_pin() by releasing o2hb_live_lock across sleeping configfs_depend_item() calls and using unlocked variants from callback context. Ensure failed pin attempts properly decrement user counts and unpin partially initialized heartbeat regions to prevent memory leaks and unprotected states. - "lib/ucs2_string.c: fix out-of-bounds read in ucs2_strnlen()" (Vincent Mailhol) Fix an off-by-one which could cause an out-of-bounds read. - "ocfs2: harden heartbeat teardown races" (Cen Zhang) Fix two OCFS2 heartbeat/o2net teardown races found by KASAN. - "taskstats: tidy up the cpumask command path" *Bradley Morgan) make two small cleanups in kernel/taskstats.c. - "ocfs2: validate active orphan slots during inode read" (ZhengYuan Huang) Validate active ordinary and append-DIO orphan slots read from OCFS2 dinodes at the metadata boundary to prevent corrupted slot indices from causing out-of-bounds array accesses. - "ocfs2: bound-check both readdir re-validation scans" (Zhan Xusheng) Enforce strict boundary checks on directory entry record lengths and offset calculations during OCFS2 directory re-scans to prevent out-of-bounds memory reads and directory position corruption. * tag 'mm-nonmm-stable-2026-08-22-16-57' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm: (95 commits) mailmap: fix bouncing address for Taniya Das ocfs2: bound-check dir entries in the inline-data re-validation scan ocfs2: bound-check dir entries in the readdir re-validation scan squashfs: avoid thundering-herd cache wakeups prctl: fix PR_SET_MM_AUXV losing the forced AT_NULL terminator mailmap: update email address for Linfeng Sun lib/interval_tree: fix allocation warning messages checkpatch: add NOKPROBE_SYMBOL to the whitelist of lines that can occur immediately after functions Squashfs: check block offset is not negative signal: factor out the kernel reserved si_code check ocfs2: fix readdir position truncation on 32-bit kernels ocfs2: fix cached cluster count after suballocator reclaim ocfs2: fix circular locking dependency in ocfs2_init_acl() ocfs2: validate DIO orphan slot during inode read ocfs2: validate orphan slot during inode read selftests/prctl: fix non-anonymous VMA mapping in set-anon-vma-name test MAINTAINERS: add IRC and patchwork for LTP include/linux/list.h: mark list_add and __list_add as __always_inline tools/mm: prevent page_owner_sort from truncating input hung_task: update DETECT_HUNG_TASK_BLOCKER Kconfig help ...
2026-08-20Merge tag 'mm-stable-2026-08-18-18-39' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm Pull MM updates from Andrew Morton: - "mm: drop "sub" prefix from various places" (Dev Jain) page->folio conversion and a naming cleanup - "mm/kasan: remove redundant initialization for kasan_flag_write_only" (Igor Putko) KASAN cleanup work - "mm/filemap: reduce unnecessary xarray lookups" (Chi Zhiling) Small speedup in the pagecaache read code - "mm/percpu: Fix possible NOFS/NOIO reclaim recursion" (Kaitao Cheng) Improve the vmalloc code - mainly the avoidance of GFP_KERNEL allocations when the caller asked for GFP_NOFS or GFP_NOIO - "mm/kmemleak: avoid soft lockup when scanning task stacks" (Breno Leitao) Avoid a soft lockup watchdog trigger from the kmemleak scanning code in extreme situations - "mm/page_owner: misc cleanups" (Ye Liu) Cleanups to the page_owner code. For some reason lots of people have been working on the page_owner code this cycle. - "mm: convert to walk_page_range_vma() to eliminate find_vma()" (Kefeng Wang) Simplify and accelerate the page walking library function - "mm/migrate: preparatory cleanups for batch copy and offload" (Shivank Garg) Cleanups in the migration code - "mm/page_owner: add per-fd filter infrastructure for print_mode and NUMA filtering" (Zhen Ni) Per-fd filtering to page_owner in order to reduce the sometimes vast amount of output it can produce - "mm: Refactor bootmem gigantic hugepage allocation" (Muchun Song) Fixes and preparatory cleanups around bootmem HugeTLB handling, sparse initialization ordering, and related vmemmap setup - "mm/zsmalloc: reduce lock contention in zs_free()" (Wenchao Hao) Reduce lock contention in zs_free(), which dominates the unmap path under memory pressure on Android (LMK kills) and on x86 servers running zswap-heavy workloads. Up to 1.83x improvement in microbenchmarking. - "move alloc_tag.c file under mm/" (Suren Baghdasaryan) - "samples/damon: handle damon_{start,stop}() failures" (SJ Park) Fix improper handling of damon_start(), damon_stop(), and damon_call() failures across DAMON sample modules to prevent potential memory leaks, operation disruptions and use-after-free bugs - "mm/damon/sysfs: kobject_del() directories that users can create/remove" (SJ Park) Fix delayed sysfs directory removal under DEBUG_KOBJECT_RELEASE causeing creation failures due to duplicate directory names by adding missing kobject_del() calls before creating new directories - "mm: cleanup clear_not_present_full_ptes()" (David Hildenbrand) Clean up the core pte handling code - "selftests/damon: misc fixes for test bugs" (Kunwu Chan) Fix several bugs in the DAMON selftests - "selftests/damon: fix memcg_path staging handling" (Cheng Nie) Fix a bug in _damon_sysfs.py for damos_filter memcg_path setup, and add a test case for it in sysfs.py. - "selftests/damon: test kdamond refresh_ms" (Ruslan Valiyev) Selftest coverage for DAMON's refresh_ms sysfs feature by updating the test control module and verifying that scheme stats update automatically without manual intervention - "mm/damon: five misc fixups" (Akinobu Mita) Miscellaneous DAMON fixups. - "mm/damon/core: detect internal variation above max_nr_regions/2" (Jiayuan Chen) Fix DAMON's region splitting behavior when region counts exceed half the maximum budget by dynamically scaling down the split fraction as the limit approaches, preventing large regions from staying un-split, and add corresponding KUnit test coverage - "mm: preparatory patches for PMD level swap entries" (Usama Arif) Refactor and clean up PMD softleaf helpers, call sites, and architecture flags to lay the groundwork for a follow-up series that introduces PMD page table swap entries - "mm/damon: update, optimize, and clean up doc, tests, and code" (SJ Park) Update DAMON design and ABI documentation, expands unit and selftest coverage, optimize damon_commit_target_regions(), and clean up recently added sysfs interface code for better readability - "mm/vmpressure: reduce CPU, memory and code overhead on cgroup v2" (Usama Arif) Optimize vmpressure() by skipping unnecessary work on cgroup v2 for userspace event notifications and refactor v1-only eventfd handling into mm/memcontrol-v1.c to reduce memory overhead and code complexity - "selftests/mm: refactor pkey helpers and fix mmap error handling" (Hongfu Li) Refactor pkeys shared tracing and assertion helpers into a common file, unify protection key selftests to use consistent diagnostic logging and assertions, and enforce standardized MAP_FAILED return checks for mmap() calls across the tests - "mm/damon: optimize out nr_accesses_bp" (SJ Park) Replace the error-prone, continuously updated nr_accesses_bp field in damon_region with an on-demand moving sum function, reducing structure memory overhead and avoiding state corruption bugs - "Open HugeTLB allocation routine for more generic use" (Ackerley Tng) Decouple HugeTLB folio allocation from VMA dependencies by introducing hugetlb_alloc_folio(), enabling subsystems like guest_memfd to allocate HugeTLB folios without standard VMA reservations or pseudo-VMAs - "mm/damon: provide pseudo moving sum probe_hits" (SJ Park) Integrate DAMON's probe_hits attribute counter into the pseudo moving sum infrastructure, enabling real-time, online monitoring without waiting for full aggregation intervals - "mm: Some cleanups for page allocator APIs" (Brendan Jackman) Simplify and refactor the page allocator entry points and flags by unifying allocation paths, adding internal alloc_flags arguments, and eliminating redundant __ prefixed alloc_pages variants. - "Fix incorrect access of hugetlb pte entries" (Dev Jain) Enforce the consistent use of huge_ptep_get() instead of ptep_get() for HugeTLB entries and fixes an unaligned address issue in arm64's huge_ptep_get() implementation - "mm/damon: validate all parameters in the core" (SJ Park) Consolidate parameter validation into the DAMON core specifically within damon_start() and damon_commit_ctx() to centralize error checking, eliminate caller-side redundant checks and to improve maintenance efficiency - "tools/mm/page_owner_sort: fix filtering and cleanup issues" (Yichong Chen) Rename is_need() to filter_record() for clearer return semantics, fix per-record allocation memory leaks and bound output copies in search_pattern() to address an existing buffer issue - "memcg: bail out reclaim when memcg is dying" (Jiayuan Chen) Mitigate a system-wide stall which occurs when a cgroup is removed while one of its memory control files is doing synchronous reclaim - "mm/memory-failure: add panic option for unrecoverable pages" (Breno Leitao) Introduce an opt-in vm.panic_on_unrecoverable_memory_failure sysctl that immediately panics the kernel on unrecoverable memory errors in kernel-owned pages to preserve error context and prevent delayed, silent data corruption - "mm/damon: refactor damon_{start,stop,commit}() for simple error handling" (SJ Park) Refactor the DAMON core API functions to guarantee that all contexts are fully stopped when damon_start(), damon_stop(), or damon_commit() fail, eliminating the need for complex and error-prone caller-side cleanup code - "Keep tail page private zero at free and folio split" (Zi Yan) Add checks to ensure tail_page->private is zero when freeing compound or high-order pages and when promoting tail pages during large folio splits. By validating these fields at free and split time, it allows the removal of redundant private field clearing inside prep_compound_tail() - "mm: drop redundant lru_add_drain in anon folio reuse paths" (Barry Song) Eliminate redundant lru_add_drain() calls in wp_can_reuse_anon_folio() and do_swap_page() to reduce LRU lock contention and system overhead By validating folio refcounts against the LRU cache before draining and removing unnecessary drains in the swap path, it achieves up to a 30.5% reduction in drain calls during heavy swap workloads - "mm: clean up folio LRU and swap declarations" (Jianyue Wu) Reorganize folio LRU and swap code by relocating page-cluster state to mm/swap_state.c, renaming mm/swap.c to mm/folio.c, and moving MM-internal reclaim declarations into mm/internal.h. - "userfaultfd: working set tracking for VM guest memory" (Kiryl Shutsemau) Add userfaultfd support for tracking the working set of VM guest memory, so a VMM can identify hot pages and reclaim cold ones to tiered or remote storage - "mm: remove CONFIG_HAVE_BOOTMEM_INFO_NODE (Part 2)" (David Hildenbrand) Remove the remaining pieces of CONFIG_HAVE_BOOTMEM_INFO_NODE, performing some smaller cleanups around freeing of reserved vmemmap pages on the way. - "mm/damon: update probe hits for runtime parameter commits" (SJ Park) Ensure that DAMON's probe_hits attribute counter is properly updated when monitoring intervals are changed at runtime, matching the behavior of nr_accesses. To achieve this, it refactors and renames existing helper functions for shared use, applies the updates to probe_hits, and handles edge cases in damon_probe_hits_mvsum() to maintain measurement accuracy. - "KSM: performance optimizations for rmap_walk_ksm" (xu xin) Resolve a severe KSM reverse-mapping performance bottleneck where thousands of split VMAs sharing a single anon_vma cause extended lock contention. By adding an interval-filtering check during the rmap walk, it reduces worst-case anon_vma lock hold times from over 500ms down to under 2ms, preventing application freezes and latency spikes under memory pressure. - "mm: split a couple of headers from internal.h" (Mike Rapoport) Split declarations related to mm_init, memblock, vmalloc and sparse into new headers - "KSM: use linear_page_index in collect_procs_ksm()" (xu xin) Apply the interval tree optimization from rmap_walk_ksm() to collect_procs_ksm() to avoid iterating over non-matching VMAs during KSM memory error handling. It hoists loop-invariant address initialization and restricts the anon_vma_interval_tree_foreach walk to a targeted page offset range, reducing redundant checks and improving lookup efficiency. - "selftests/mm: avoid false failures in hugetlb and KSM tests" (Sayali Patil) Fix issues in the hugetlb and KSM MM selftest categories that can report failures when the prerequisites for the tests are not satisfied - "mm/damon: introduce data attributes only monitoring" (SJ Park) Introduce attribute-weighted region management in DAMON, allowing users to prioritize specific data attributes (such as page sizes or cgroups) over or instead of access monitoring. By assigning weights to attribute probes, DAMON can completely disable access tracking and adjust monitoring regions based on weighted probe-hit counters to optimize monitoring quality for attribute-focused workloads. - "mm/hmm: Add mmap lock-drop support for userfaultfd-backed mappings" (Stanislav Kinsburskii) Extend hmm_range_fault() to support userfaultfd-backed regions by allowing the mmap lock to be dropped during fault handling via a new hmm_range_fault_locked() helper. By accepting a locked pointer and signaling retry status when lock release occurs, it enables page fault resolution in userfaultfd regions while preserving backward compatibility for existing callers. - "mm: make VMA page offset handling more consistent" (Lorenzo Stoakes) Clean up and standardize how vma->vm_pgoff is accessed and manipulated across file-backed and anonymous mappings in the kernel It introduces dedicated helper functions such as vma_start_pgoff(), vma_end_pgoff(), vma_set_pgoff() and linear_page_delta() while renaming rmap interval tree helpers to better reflect their functionality. These changes establish a cleaner foundation for future work that will unify virtual page offset indexing for all anonymous and CoW'd folios. - "mm: handle device-private PMDs in walk callbacks" (Usama Arif) Address kernel panics and state corruption caused by MM walk callbacks reaching non-present device-private PMD swap entries created during HMM migrations It ensures that functions which acquire pmd_trans_huge_lock() properly recognize device-private PMDs instead of assuming a present THP or a standard migration entry. - "mm/rmap: Refactor try_to_unmap_one" (Dev Jain) Refactor try_to_unmap_one by modularizing Hugetlb, anonymous-lazyfree, and anonymous-swapbacked logic into dedicated functions, laying the structural groundwork for batched anonymous large folio unmapping. - "Docs/ABI/damon: sysfs ABI document fixes and additions" (Song Hu) Fix typos and fills in missing entries in the DAMON sysfs ABI document - "dax/kmem: atomic whole-device hotplug via sysfs" (Gregory Price) Introduce an atomic sysfs state attribute and supporting DAX/MM infrastructure to prevent userland races when offlining and removing entire memory regions By adding an unplugged state alongside standard online modes, it enables whole-device atomic hotplug control while preserving backward compatibility. - "mm: convert more vm_flags_t users to vma_flags_t" (Lorenzo Stoakes) Continue transitioning the kernel from the deprecated vm_flags_t type to vma_flags_t across core memory management infrastructure. It replaces legacy type usage in core functions such as do_mmap(), unmapped area allocation, mm->def_vma_flags, and VMA operations like mlock, mprotect, and mremap. - "Two small patches to clean up mm/mm_slot.h" (xu xin) Refactor mm_slot.h by introducing mm_slot_remove() to unify duplicate slot deletion sequences in khugepaged and KSM. It also adds code documentation explaining why mm_slot_lookup and mm_slot_insert must remain as preprocessor macros rather than static inline functions. - "mm/damon/core: hide core-private struct fields" (SJ Park) Clean up DAMON core structures by consistently marking internal-only fields with private: comment tags to prevent improper direct access from outer layers. It enforces encapsulation across core structures including damon_region, damon_target, and damon_ctx and updates DAMON_SYSFS to interact through approved access APIs instead of exposing raw struct members. - "mm/damon: unurgent fixes for infinite loop, NULL de-ref and races" (SJ Park) Address potential infinite loops, NULL dereferences, and race conditions identified in DAMON It fixes an infinite loop triggered by extreme user configurations, a NULL pointer dereference within unit tests and minor monitoring accuracy degradation caused by subtle runtime races. - "mm/page_alloc: fixes for free_pages_nolock() on RT/UP" (Brendan Jackman) Fix an NMI safety flaw in __free_frozen_pages() where freeing pages on non-SMP or PREEMPT_RT kernels can bypass can_spin_trylock() checks via non-PCP or isolated migration paths. It also resolves potential kernel crashes and privilege escalation risks triggered when BPF tracing runs in NMI context alongside memory hotplug or large allocation frees. - "mm/page_alloc: couple of followups for recent cleanups" (Brendan Jackman) Clean up and update page allocator nomenclature, documentation, and debug assertions. It aligns internal FPI_ flags with the public "nolock" naming convention, removes outdated internal implementation details from high-level page allocator comments, and eliminates obsolete VM_BUG_ON() assertions in allocation paths. - "mm/mseal: further cleanups" (Lorenzo Stoakes) Refactor and simplify the mseal implementation by clarifying API boundaries and removing unnecessary code complexity. It replaces generic do_mseal() usage outside the syscall with a dedicated mseal_mmap_page_zero() helper for MMAP_PAGE_ZERO, eliminates mm_struct parameters to enforce that sealing applies only to current->mm, and streamlines overall logic and comments with no functional changes intended. - "mm/vmscan: fix swappiness=max and clean up per-node proactive reclaim" (Ridong Chen) Resolve reclaim behavior bugs and clean up function parameters across memory reclaim paths It fixes swappiness=max in both standard reclaim and MGLRU so unswappable anonymous memory no longer falls back to evicting page cache, ensures reclaim_store() returns accurate error codes instead of collapsing all failures into -EAGAIN, and removes the obsolete gfp_mask parameter from __node_reclaim(). - "mm: mincore: misc cleanups" (Kefeng Wang) Clean up and simplifies the mincore code. Most importantly, it removes the historical special behavior that always reports VM_PFNMAP pages as non-resident. - "mm/huge_memory: drop dead split helper variants" (Kiryl Shutsemau) Two trivial cleanups in the folio split API - "mm/damon: fix uninitialized DAMOS field and kunit exec expectation bugs" (SJ Park) Resolve minor operational and testing bugs in DAMON identified by Sashiko. It initializes the damos->last_applied field to prevent occasional efficiency degradation and fixes invalid memory accesses in DAMON KUnit tests during test failure handling. - "cleanup for stable_page_flags()" (Jinjiang Tu) Clean up and refactor stable_page_flags() used by /proc/kpageflags without altering functionality. It uses BIT_ULL() to prevent shift-overflow warnings on 64-bit flag bits, converts folio-specific flag checks to standard folio_test_*() helpers, and removes redundant CONFIG_PAGE_IDLE_FLAG handling. - "Batch unmap of uffd-wp file folios" (Dev Jain) Extend batched folio unmapping support to file folios within userfaultfd write-protect (uffd-wp) VMAs by adding batching capabilities to pte_install_uffd_wp_if_needed(). This removes special-case restrictions on uffd-wp VMAs in try_to_unmap_one(), significantly simplifying the function's control flow and complexity. - "mm/early_ioremap: clarify and clean up early_ioremap_reset()" (Sang-Heon Jeon) Clarify and clean up the architecture-specific usage of __late_set_fixmap() and __late_clear_fixmap() after early_ioremap_reset() It adds explicit documentation regarding when early_ioremap_reset() must be called and removes redundant macro definitions and reset calls in the RISC-V and ARM64 architectures. - "mm: fix reclaim storms in defrag_mode" (Johannes Weiner) Address severe performance regressions, swap storms, and spurious OOMs caused by vm.defrag_mode=1 under high memory pressure in Meta production It updates the page allocator slowpath so non-movable allocation requests actively trigger direct reclaim and direct compaction at pageblock_order scale, allowing them to claim whole pageblocks rather than spinning unproductively. - "zram: lockmap tweaks" (Sebastian Siewior) Optimize and fix lockdep tracking for zram devices by consolidating per-entry lockmaps and isolate lock classes across multiple instances This reduces memory overhead by replacing per-entry lockdep_map instances with a single map per struct zram, and assigns a dynamic lock_class_key to each instance to prevent false deadlock reports when different zram devices are backed by distinct filesystems. * tag 'mm-stable-2026-08-18-18-39' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm: (501 commits) selftests/mm: thuge-gen: fix test_shmget() for PAGE_SIZE check selftests/mm: unpoison pages in memory-failure teardown mm/shmem: downgrade final i_blocks check in shmem_evict_inode() to pr_warn() mm/khugepaged: replace mutex_lock/mutex_unlock usage with guard macro mm/zsmalloc: fix release order of locks in zs_page_migrate() Documentation: zram: remove sections numbering ksm: stop iterating VMAs when ksm_test_exit returns true mm: fold userfaultfd_rwp() to false without CONFIG_ARCH_HAS_PTE_PROTNONE mm/migrate: report RCU-tasks quiescent states in migrate_pages_batch() zram: use a custom key for each zram object zram: move lockmap to be per-zram instead per table selftests/mm: fix gup_longterm EINVAL error message mm: page_alloc: fix non-movable reclaim storm in defrag_mode mm: page_alloc: move capture_control to the page allocator mm: compaction: support non-movable compaction for pageblock requests mm: page_alloc: __GFP_FS lockdep annotation for direct compaction hugetlb: evaluate subpool free state while locked mm/damon: remove trailing semicolons after function definitions mm/damon/ops-common: prevent migration fallback to non-target nodes mm/damon: update outdated comment about DAMOS filter handling ...
2026-08-20Merge tag 'bitmap-for-7.3' of https://github.com/norov/linuxLinus Torvalds
Pull bitmap updates from Yury Norov: "The usual set of fixes, cleanups and performance improvements together with a couple of new tests: - bitmap_find_next_zero_area_off() optimization (Sunyi) - bitmap_find_next_zero_area_off(): return size when no zero area is found (Yury) - bitmap vs IDA vs Maple Tree performance test (Yury) - get rid of cpumap_print_to_pagebuf() (Yury) - use nr_node_ids in __nodemask_pr_numnodes() (Li RongQing) - bitops: make the *_bit_le functions use unsigned long (Benjamin) - bitmap scatter & gather test fix (Christophe) - use __ASSEMBLER__ in bitmap header files (Thomas)" * tag 'bitmap-for-7.3' of https://github.com/norov/linux: (25 commits) lib: test bitmap vs IDA vs Maple Tree performance for region allocations bitmap: Return size when no zero area is found media: s5p-mfc: Treat bitmap size as allocation failure crypto: ccp: Treat bitmap size as allocation failure powerpc/msi: Treat bitmap size as allocation failure ARM: dma-mapping: Treat bitmap size as allocation failure bitmap: drop bitmap_next_set_region() nodemask: reduce bitmap width to nr_node_ids in __nodemask_pr_numnodes() bitmap: Properly initialise destination bitmap for scatter & gather test lib/bitmap-str: get rid of cpumap_print_to_pagebuf() perf: Use sysfs_emit() for cpumask show callbacks PCI/sysfs: Use sysfs_emit() for cpumask show callbacks RDMA/hfi1: Use sysfs_emit() for cpumask show helper hwtracing: hisi_ptt: Use sysfs_emit() for cpumask show fpga: dfl-fme-perf: Use sysfs_emit() for cpumask show devfreq: Use sysfs_emit() for cpumask show callbacks cpu: Use sysfs_emit() for cpumask show callback x86/events: Use sysfs_emit() for cpumask show callbacks powerpc: Use sysfs_emit() for cpumask show callbacks arm: Use sysfs_emit() for cpumask show callbacks ...
2026-08-20Merge tag 'net-next-7.3' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next Pull networking updates from Jakub Kicinski: "One of the 'small improvements all over the place' releases for us. It's hard to draw any direct comparisons because summer vacations disrupted our patch processing (and presumably - generation) quite a bit. Quick and dirty count suggests we (Paolo and I) merged a very similar number of net (632) and net-next (648) patches. This is not telling the full story either because 1/3 to 1/2 of the net-next patches also *seem* like AI-driven low priority fixes, cleanups and clarifications. We are completely overwhelmed, of course. The glimmer of hope is that we secured sufficient LLM budget and access (thank you Meta!) to run reviews with multiple frontier models on each patch. This eliminates some hallucinations. That said, in terms of review, the LLMs can only do so much. The sad truth is that our APIs (especially for rare events like PCIe errors, timeouts etc) have always been racy, and now LLMs don't let us ignore that. I expect our direction for the next release will be to tweak the reviews a little bit more, but start shifting focus to letting the LLMs take care of the busy work - managing patchwork, automating common process complaints, editing commit messages, and maybe applying patches which already got "reviewed-by" tags from people we trust... Core & protocols: - A few steps lowering rtnl_lock dependence: - per-netns netdev unregistration for select SW drivers (e.g. veth, ipvlan, tunnels) - rtnl_lock-less FIB rule changes (RTM_NEWRULE and RTM_DELRULE) - prepare software drivers and TC qdiscs for rtnl_lock-less GET - Support BIG TCP (>64kB TSO) in UDP tunnels (vxlan, geneve) - Support buffers larger than PAGE_SIZE in devmem zero-copy API - Improve MPTCP handling of extreme memory pressure handling, when out-of-order queue had to be pruned - Report the per-group user count via RTM_GETMULTICAST - Expose the route deletion reason in RTM_DELROUTE - Add a SO_RIGHTS_NOTRUNC option to UNIX sockets to enable more useful handling of LSM denials when receiving SCM_RIGHTS messages: instead of truncating the message at the first blocked fd, keep every fd slot and store the LSM errno in the blocked slot - IPv6 Segment Routing - support looking up the post-encap SID (address) in a different/specified routing table - Support PRP RedBox (interlink) creation - Support per-nexthop UDP dst port in VXLAN - Continue converting getsockopt callbacks in a number of protocols to iov_iter Ethernet: - Merge initial CXL support for AMD/Solarflare NICs (shared branch with the CXL tree) - New drivers: - ADIN1140 10BASE-T1S MACPHY - Initial skeleton of Intel iXD and ZTE Dinghai drivers - High-speed NICs: - AMD/Pensando: - support firmware flashing - Cisco (enic): - SR-IOV V2 admin channel and MBOX protocol - Huawei (hns3): - support for ethtool pfc_prevention_tout - nVidia/Mellanox: - support sharing bandwidth control across interfaces of the same device - Marvell (octeontx2-pf): - link RQ page pools to netdev for Netlink stats - Google vNIC: - XDP metadata support for DQ RDA - Microsoft vNIC: - support forcing full-page RX buffers - Other NICs: - Synopsys IP: - eic7700: support for eth1 - Microchip (lan743x): - support for RMII interface - Wangxun: - support for ethtool -G and -C for VFs - add Tx timeout and PCIe error handling - Intel (igb/igc): - RSS key get/set support - support for forcing link speed without auto-negotiation - Switches: - NXP (dpaa2): - support bonding/LAG offload - Mediatek: - mt7530: EN7528 support - initial support for MT7628 - Micrel (ksz8/9): - refactoring work to move towards library model - PTP support for KSZ8463 - nVidia/Mellanox: - support rtnl-lock-less ethtool callbacks - Realtek: - rtl8366rb: use generic RTL83xx code - support SGMII and HSGMII for RTL8367S - PHYs: - Airoha: - EcoNet EN7528 PHY support - DAPU Telecom - DAPU Telecom DAP8211R(I) Gigabit PHY support - Realtek: - support RTL8261C_CG - support RTL8261D Wireless: - nl80211: per-link statistics support for multi-link operation - mac80211: AQL/airtime-fairness support for multicast - Merge Peripheral Authentication Service (PAS) / TEE support for ath12k (shared branch with the firmware/qcom tree) - New drivers: - mm81x for Morse Micro Long-Range S1G devices - nxpwifi for NXP devices (mostly forked off from mwifiex) - Driver changes: - Broadcom (brcmfmac): - DPP support, some Cypress part update - MediaTek (mt76): - mt7928 support - mt7925 NAN support - mt7996 AP powersave improvements - Qualcomm (ath12k): - much kernel infrastructure integration work - AHB platform MultiPD support - Realtek (rt89): - LED support - RTL8922DE support - dual-BT coex for RTL8922D - Intel: - new FW version support Bluetooth: - HCI: add support for Shorter Connection Interval (SCI) feature - af_bluetooth: add minimal context analysis annotations - Driver changes: - Intel: - add Bluetooth SAR revision 2 support - add vendor_reset PCI sysfs for PLDR - Mediatek: - add USB IDs for MT7902 and MT7922 devices - Realtek: - add USB IDs for 8761CU and 8852BE devices - NXP: - add M.2 Bluetooth device support using pwrseq Misc: - DPLL support for manual/numerical oscillator control (NCO) (implement in zl3073x) - MCTP support for MCTP over USB v1.1 (DMTF DSP0283) - Power-over-Ethernet: support Realtek PSE controllers - Remove the IBM EHEA driver - Remove tulip/xircom_cb driver" * tag 'net-next-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next: (1433 commits) net/mlx5e: do not HW-GRO coalesce small frames net: openvswitch: fix nf_connlabels leak in ovs_ct_init net: add missing ref_tracker_dir_exit() to alloc_netdev_mqs() net: openvswitch: fix flow mask use-after-free on flow deletion sctp: stop processing a packet once its association is deleted dpll: zl3073x: add PTP clock support dpll: zl3073x: add channel ToD, phase step and TIE operations dpll: zl3073x: scale poll interval proportionally to timeout ptp: vmclock: prevent read-only mappings from becoming writable ipv4: reject undersized MTUs in ip_do_fragment() bonding: initialize err for empty target lists net: dsa: initial support for MT7628 embedded switch net: dsa: initial MT7628 tagging driver net: phy: mediatek: add phy driver for MT7628 built-in Fast Ethernet PHYs dt-bindings: net: dsa: add MT7628 ESW net: pse-pd: realtek-pse-mcu: add UART transport net: pse-pd: realtek-pse-mcu: add I2C transport net: pse-pd: add Realtek PSE MCU core dt-bindings: net: pse-pd: add bindings for Realtek PSE MCU vsock: use sock_error() to consume sk_err after a failed connect ...
2026-08-20Merge tag 'bpf-next-7.3' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next Pull bpf updates from Daniel Borkmann: "Major changes: - Redesign the verifier error reporting: failures now carry source and instruction annotations along with the causal event history that led to them, making program rejections far easier to debug and repair (Kumar Kartikeya Dwivedi) - Add arena argument support to kfuncs and struct_ops through the new __arena and __arena__nullable suffixes (Tejun Heo, Puranjay Mohan, Kumar Kartikeya Dwivedi, Ihor Solodrai) - Signed BPF program loader rework to accommodate both BPF and security community needs where the kernel runs the signature verification at BPF_PROG_LOAD time before the LSM admission hook (Daniel Borkmann) - Add a set of ksock kfuncs which let BPF LSM and syscall programs create, connect and send on UDP sockets in order to emit telemetry data (Mahe Tardy) - Unify helper and kfunc call argument verification and classify kfunc arguments purely from BTF into a generated bpf_func_proto which is computed once at add-call time (Amery Hung) Other features and fixes: - Enable EXECMEM_ROX_CACHE for BPF allocations on x86 (Mike Rapoport) - Add bidirectional VLAN support to bpf_fib_lookup() through the new BPF_FIB_LOOKUP_VLAN and BPF_FIB_LOOKUP_VLAN_INPUT flags (Avinash Duduskar) - Infer zext_dst from static register liveness analysis to fix 32-bit zero-extension semantics, and remove the artificial limitations on pointer types eligible for spilling (Eduard Zingerman) - Inline the numeric open-coded iterator kfuncs so that bpf_for() loops no longer pay a kfunc call on every iteration (Puranjay Mohan) - Add an arena-based bitmap data structure to libarena along with serial and parallel selftests (Emil Tsalapatis) - Teach resolve_btfids to discover kfuncs from the kernel's BTF ID sets and to emit kfunc BTF decl tags, reducing the kernel build's dependency on pahole features (Ihor Solodrai) - Add BPF_F_ADJ_ROOM_DECAP_* flags to bpf_skb_adjust_room() so that tunnel decapsulation can update the GSO and encapsulation state of the skb (Nick Hudson) - Fix the ring buffer pending_pos walk and the available-data accounting on 32-bit position wrap (Israel Téllez García) - Add memory usage accounting for arena maps and fix an mmap_lock deadlock on arena lock failure (Jiayuan Chen) - Add tracing_multi link info support to the kernel UAPI and bpftool, and refactor the stack map code to run with preemption disabled (Jiri Olsa) - Support BPF_F_EGRESS in bpf_redirect_peer() to emit the skb in the egress direction of the target's peer device (Jordan Rife) - Add a KF_SPINLOCK_SAFE kfunc flag so that providers, in particular modules, can declare kfuncs safe to call under bpf_spin_lock instead of relying on the verifier's hard-coded allowlist (Kaitao Cheng) - Introduce global percpu data for BPF programs with libbpf probing and bpftool skeleton support, and stop exposing uninitialized kernel heap memory when copying per-CPU map values (Leon Hwang) - Add s390 JIT support for load-acquire and store-release instructions (Maxim Khmelevskii) - Fix a CFI mismatch in the task work callback and an arm64 KASAN false positive after bpf_throw() (Mykyta Yatsenko) - Reject writes through untrusted BTF pointers and bound the rdonly/rdwr_buf_size kfunc arguments (Nicholas Dudar) - Invalidate RCU pointers only after the final spin unlock and account for preempt and IRQ disabled regions as overlapping RCU protection (Ning Ding) - Support mixing bpf2bpf calls and tail calls on RV64, add signed operations and 32-bit atomics to the RV32 JIT, and add timed may_goto support (Pu Lehui, Kuan-Wei Chiu, Feng Jiang) - Fix a use-after-free on mm_struct in bpf_find_vma() for foreign tasks and an mmap_lock leak in the irq_work path (Sanghyun Park) - Populate mmap-able BPF array map memory lazily which makes mmap() O(1) instead of proportional to the map size (Song Liu) - Introduce a jit_required flag and reject programs with inlined helpers when no JIT is available, where the interpreter would otherwise jump into an invalid address (Tiezhu Yang) - Fix the x86 JIT per-CPU address resolution into an extended register where the REX prefix dropped the high destination register bit (Vineet Gupta) - Reject MEM_ALLOC BTF accesses past object bounds, arena frees below the arena base, and mixed arena and ordinary atomic paths (Yiyang Chen) - Fix the trampoline handling of 128-bit arguments and of return values larger than 8 bytes (Yonghong Song) - Ensure that any fault prone load is rewritten with exception table handling, and fix the arena load-acquire and atomic fetch handling in the x86, arm64, riscv and s390 JITs (Daniel Borkmann) - Many more fixes and cleanups across the verifier, arena, trampolines, sockmap, cgroup, ring buffer, x86/arm64/riscv/s390 JITs, libbpf, bpftool, resolve_btfids and selftests" * tag 'bpf-next-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next: (373 commits) selftests/bpf: Add tests for a store on a fault prone qdisc pointer selftests/bpf: Add tests for fault prone loads out of RCU pointers selftests/bpf: Add tests for pointer type merge at a shared load selftests/bpf: Remove duplicate copies of the arena spinlock qnodes selftests/bpf: Retry stat generation in cgroup_iter_memcg selftests/bpf: Test pseudo-function policy diagnostics bpf: Distinguish function references in policy diagnostics bpf: Preserve source attribution without source text selftests/bpf: Test kfunc argument diagnostics bpf: Correct kfunc argument diagnostics bpf: Use canonical stack argument names in diagnostics bpf: Preserve R0 lineage across helper calls selftests/bpf: Exercise negative optlen in cgroup getsockopt hook bpf: Reject negative optlen in cgroup getsockopt hook selftests/bpf: tc_tunnel - validate decap GSO and encapsulation state bpf: Clear decap state on skb_adjust_room shrink path bpf: Allow new DECAP flags and add guard rails bpf: Add BPF_F_ADJ_ROOM_DECAP_* flags for tunnel decapsulation bpf: Refactor masks for ADJ_ROOM flags and encap validation bpf: Name the enum for BPF_FUNC_skb_adjust_room flags ...
2026-08-18Merge tag 'locking-core-2026-08-17' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull locking updates from Ingo Molnar: "Futexes: - Use runtime constants for futex_hash computation (K Prateek Nayak, Peter Zijlstra) - Optimise the size check get_futex_key() (Sebastian Andrzej Siewior) - Avoid private hash use-after-free on final put (Felix Hoffmann) - Tell kmemleak we're not leaking __futex_queues (Peter Zijlstra) Rust integration updates: - Implement refcounted interrupt disable and SpinLockIrq for Rust (Boqun Feng, Heiko Carstens, Joel Fernandes, Lyude Paul) - Rust sync: add helpers for mb, dma_mb and friends; add generic memory barriers and use LKMM atomics instead of Rust atomics in the revocable code (Gary Guo) - Add abstraction and integrate synchronize_rcu() (Philipp Stanner) Lock debugging: - Add qspinlock contended_release tracepoint (Dmitry Ilvokhin, Peter Zijlstra) - Enable the printing of held locks of remote running tasks and print task CPU (Ingo Molnar) - percpu-rwsem: Annotate intentional data race in readers_active_check() (Sun Shaojie) Misc fixes and updates by Boqun Feng, Peter Zijlstra, Fangrui Song, Naveen Kumar Chaudhary and Thomas Huth" * tag 'locking-core-2026-08-17' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: (44 commits) rust: sync: Introduce SpinLockIrq::lock_with() and friends rust: sync: Add SpinLockIrq rust: sync: Use super::* in spinlock.rs rust: helper: Add spin_{un,}lock_irq_{enable,disable}() helpers rust: Introduce interrupt module s390/preempt: Enable HAS_SEPARATE_PREEMPT_RESCHED_BITS arm64: sched/preempt: Enable HAS_SEPARATE_PREEMPT_RESCHED_BITS preempt: Introduce HAS_SEPARATE_PREEMPT_RESCHED_BITS sched: Avoid signed comparison of preempt_count() in __cant_migrate() sched: Remove the unused preempt_offset parameter of __cant_sleep() locking: Switch to _irq_{disable,enable}() variants in cleanup guards irq: Add KUnit test for refcounted interrupt enable/disable irq,spin_lock: Add counted interrupt disabling/enabling openrisc: Include <linux/cpumask.h> in smp.h preempt: Introduce __preempt_count_{sub,add}_return() preempt: Introduce HARDIRQ_DISABLE_BITS preempt: Track NMI nesting to separate per-CPU counter futex: Tell kmemleak we're not leaking __futex_queues x86/paravirt: Trace contended_release on unlock tracing/lock: Use TRACE_EVENT_FN() for contended_release ...
2026-08-18Merge tag 'objtool-core-2026-08-17' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull objtool updates from Ingo Molnar: - Fix various klp-build bugs reported by Joe Lawrence (Josh Poimboeuf, Joe Lawrence) - Misc fixes and cleanups (Puranjay Mohan, Thomas Huth and Ingo Molnar) * tag 'objtool-core-2026-08-17' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: objtool/klp: Fix vmlinux klp relocations for EXPORT_SYMBOL_FOR_MODULES() objtool/klp: Fix .kcfi_traps special section extraction objtool/klp: Fix vmlinux .klp.symid link error for .exitcall.exit symbols objtool/klp: Fix line numbers in Module.symvers parse errors objtool/klp: Fix relocations for EXPORT_SYMBOL_FOR_MODULES() symbols objtool/klp: Allow new references to module exports objtool/klp: Don't match local symbols against exports objtool/klp: Fix cross-module klp relocation section naming objtool/klp: Explicitly disallow patching or referencing init code/data objtool/klp: Ignore replacement offset of empty x86 alternatives objtool/klp: Fix size of empty special section entries objtool/klp: Fix vmlinux .klp.symid link error for .no_trim_symbol symbols objtool/headers: Sync tools/include/linux/objtool_types.h with include/linux/objtool_types.h objtool: Replace __ASSEMBLY__ with __ASSEMBLER__ in header files objtool/klp: Fix symbol resolution for duplicate data symbols objtool/klp: Add .klp.symid for sympos disambiguation objtool/klp: Skip hidden directories when finding objects objtool/klp: Fix false module dependencies caused by dead relocs objtool/klp: Normalize Module.symvers paths to module names objtool/klp: Fix module name normalization for paths with dots
2026-08-18Merge tag 'arm64-upstream' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux Pull arm64 updates from Will Deacon: "There's a reasonable amount of stuff here, including a bunch of updates to the perf PMU drivers and some MPAM updates to expose the memory bandwidth counters via resctrl. On the architecture side, some highlights include support for BBML3 and steps towards support for an architectural NMI solution, all wrapped up in a web of fixes for latent issues identified by Sashiko. ACPI: - Combine reads of AMU counters into a single FFH feedback counter op Confidential computing: - Fix smp_processor_id() in preemptible context when retrieving an attestation token inside a realm - Convert pKVM over to a "CC platform" - Clean-up our SWIOTLB configuration in preparation for reworking the handling of encrypted/decryped DMA buffers in the dma-mapping tree CPU errata handling: - Work around broken device memory ordering on NVIDIA Olympus cores - Fix broken 'nospectre_bhb' command-line option - Select the idle loop backend instruction on the command-line CPU features: - Replace our BBML2-noabort feature with the new architectural BBML3 feature - Disable in-kernel BTI for recent versions of Clang due to issues with livepatch that are still being investigated - Clean-up documentation describing which ID register fields are exposed to userspace Interrupts: - Preliminary work towards supporting FEAT_NMI, which cleans up our IRQ entry code and fixes some latent issues with pseudo-NMI - Support for an SDEI backend to trigger an NMI backtrace Memory management: - Treat all devices as coherent when CLIDR_EL1.LoC == 0 - Fix no-map handling of sub-page-sized regions - Second attempt at unmapping the linear aliases of the kernel data and bss sections - Fix EFI runtime calls when software-PAN is enabled Miscellaneous: - Add Mark Rutland as a reviewer! - Tidy-up our futex cmpxchg logic when using the new LSUI instructions - Drop the requirement on DYNAMIC_FTRACE_WITH_CALL_OPS when selecting HAVE_DYNAMIC_FTRACE_WITH_DIRECT_CALLS - Fix a false-positive KCSCAN splat in the delay loop - Use a portable typedef for 128-bit scalar types in our UAPI headers - Non-critical fixes for Sashiko reports all over MPAM: - Hook MPAM memory bandwidth counters into resctrl's counter assignment interface - Fix a quirk in the MPAM bandwidth counting on Nvidia T241 so that it also applies to 63 bit counters Perf: - Workarounds for hardware issues in the CMN-S3 PMU (Graviton 5) and CPU PMU (NVIDIA Olympus again!) - Add support for the DDR PMU on Marvell CN20K SoCs - Add support for Picoheart implementations of the DCW PCIe PMU - Add support for Channel/Rank/Bank filtering in the CXL PMU driver - Add support for 64-bit counters in the CSPMU device - Add support for revision 2 of the CMN S3 PMU Ptrace: - Fix a decade-old bug in our handling of seccomp and tracing on syscall entry - Fix regset handling for inactive SVE and SSVE registers Selftests - Add some tests for the decade-old bug that we just tried to fix in our syscall entry path - Fix SVE test crash on SME-only CPUs" * tag 'arm64-upstream' of git://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux: (95 commits) arm64/efi: Avoid voluntary preemption with efi_mm installed arm64: bti: Disable in-kernel BTI with recent versions of Clang arm64: entry: Avoid unnecessary local_irq_disable() on kernel exit irqchip/gic-v3: make the unmasking of pseudo-NMIs explicit when handling IRQs arm64: Disable KCSAN instrumentation in delay.o arm_mpam: Disable driver unbind to avoid UAF arm_mpam: Fix a NULL pointer dereference on unbinding after an error interrupt perf: arm_pmuv3: Zero initialize hw_id branch stack field arm64: mm: Unmap kernel data/bss entirely from the linear map iommu/arm-smmu-v3-sva: Use system_supports_bbml3() to detect CPU feature perf/arm-cmn: Support CMN S3 r2 perf/arm-cmn: Plumb in new filter types perf/arm-cmn: Refactor event filter data perf/arm-cmn: Refactor event filter programming perf/arm-cmn: Rename filter variables for clarity arm64: mm: fix accidental linear mapping of no-map reserved memory tools: Ensure tools copy of linux/filter.h exports the UAPI kselftest/arm64: Fix abi test compilation errors arch: arm64: add early_param idle=<wfi|yield|nop> arm64: entry: mask DAIF before returning from C EL1 handlers ...
2026-08-18Merge tag 'nolibc-20260814-for-7.3-1' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/nolibc/linux-nolibc Pull nolibc updates from Thomas Weißschuh: - New architecture: Alpha - New library functionality: readlink(), getcwd() - Various bugfixes and cleanups * tag 'nolibc-20260814-for-7.3-1' of git://git.kernel.org/pub/scm/linux/kernel/git/nolibc/linux-nolibc: tools/nolibc: add support for Alpha tools/nolibc/powerpc: mark ctr and xer as clobbered by system call tools/nolibc: remove dead __ARCH_WANT_SYS_OLD_SELECT selftests/nolibc: add debug information tools/nolibc: mark arg1 operand in __nolibc_syscall0() as write-only selftests/nolibc: Add test for getcwd() and readlink() tools/nolibc: unistd: Add readlink() tools/nolibc: unistd: Add getcwd()
2026-08-17Merge tag 'fscrypt-for-linus' of git://git.kernel.org/pub/scm/fs/fscrypt/linuxLinus Torvalds
Pull fscrypt updates from Eric Biggers: "The main change this cycle is a significant simplification that's been overdue for a while now: standardizing on a single file contents encryption implementation in ext4 and f2fs, instead of having two. Specifically, the original filesystem-layer file contents encryption implementation is removed, and the blk-crypto implementation is now used unconditionally. blk-crypto delegates either to inline crypto hardware or to the CPU via blk-crypto-fallback. The latter is functionally equivalent to the original filesystem-layer code. The blk-crypto implementation already existed, but previously it was used only when the filesystem was mounted with "-o inlinecrypt". Now, "-o inlinecrypt" just selects whether inline crypto hardware is used. To allow maintaining that user control over hardware use, the blk-crypto API is extended with a new flag BLK_CRYPTO_CFG_ALLOW_HW. Overall, this removes quite a bit of redundant code from ext4, f2fs, and fs/crypto/. It should make things easier for ongoing filesystem efforts such as iomap support, large folios, and btrfs encryption (btrfs had already been planning to use blk-crypto exclusively.) There are two small behavior changes of note: - Direct I/O now works on encrypted files even without "-o inlinecrypt", rather than falling back to buffered I/O. This is effectively a bugfix, though I'll continue to keep an eye out for any user that may have been depending on the buffered I/O fallback. - IV_INO_LBLK_32 policies are no longer supported in certain cases that didn't make sense and have no known uses. This has been in linux-next since July 22 with no reported issues. All encryption xfstests pass on ext4 and f2fs. As usual I've also been using it on a system with an fscrypt-encrypted home directory. Of course, the blk-crypto code paths also aren't new and were already being used on many systems via the inlinecrypt mount option. In addition to the main change described above, there are a few other cleanups such as using lock guards for mutexes, improving documentation, and removing a workaround for outdated gcc versions" * tag 'fscrypt-for-linus' of git://git.kernel.org/pub/scm/fs/fscrypt/linux: (29 commits) blk-crypto: Update docs for blk-crypto-fallback motivation blk-crypto: Remove unused function blk_crypto_config_supported() fscrypt: Update docs for data path fscrypt: Remove unused function fscrypt_finalize_bounce_page() f2fs: Update outdated comment in f2fs_write_begin() fs: Update outdated comment for SB_INLINECRYPT fscrypt: Update encryption policy version docs fscrypt: Replace some variable-size memsets with fixed-size fscrypt: Add safety checks to non-block-based en/decryption fscrypt: Merge bio.c and inline_crypt.c into block.c fscrypt: Remove unused functions and workqueue fscrypt: Remove fs-layer zeroout code fscrypt: Remove fscrypt_dio_supported() fscrypt: Replace calls to fscrypt_inode_uses_inline_crypto() fs/buffer: Remove fs-layer decryption code f2fs: Remove fs-layer file contents en/decryption code ext4: Further de-generalize the bio postprocessing code ext4: Make ext4_bio_write_folio() return void ext4: Remove fs-layer file contents en/decryption code Documentation: fscrypt: Update docs for inlinecrypt ...