summaryrefslogtreecommitdiff
path: root/kernel/bpf
AgeCommit message (Collapse)Author
37 hoursMerge branch 'for-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next.git # Conflicts: # mm/internal.h
37 hoursMerge branch 'fs-next' of linux-nextMark Brown
# Conflicts: # fs/coredump.c # fs/f2fs/f2fs.h # fs/fuse/dax.c # fs/xfs/libxfs/xfs_btree.c
3 daysbpf: arena: mark arena_map_mmap() mappings VM_MIXEDMAPLorenzo Stoakes (ARM)
The bpf_map->ops->map_mmap callback invoked by bpf_map_mmap() can be set to one of ringbuf_map_mmap_kern(), ringbuf_map_mmap_user(), array_map_mmap() or arena_map_mmap(). It is convention in mm to mark mappings whose pages the kernel manages itself with VM_MIXEDMAP, so core mm knows not to treat them as ordinary page cache or anonymous memory. The map_mmap callbacks ringbuf_map_mmap_kern() and ringbuf_map_mmap_user() use remap_vmalloc_range(), which ultimately invokes vm_insert_page() and so marks the ranges VM_MIXEDMAP, and array_map_mmap() sets VM_MIXEDMAP explicitly. However, the exception to this is arena_map_mmap(), which doesn't set the flag. This patch corrects this and updates the comment to reflect it. The pages are refcounted and vm_normal_page() finds them regardless of the flag, and VM_DONTEXPAND remains set (marking the memory as VM_SPECIAL and thus unmergeable). The one effect is that NUMA balancing now skips these VMAs, as it already does for the other bpf map mappings, which is the reason array_map_mmap() gives for setting the flag. The intent of this patch is to be able to establish the invariant that only PFN-mapped or mixed map ranges may clear the VM_MAYWRITE flag, as is done in bpf_map_mmap(). Link: https://lore.kernel.org/20261003-b4-mmap-prepare-vma-flag-sanify-v4-13-a1f052500fd7@kernel.org Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com> Reviewed-by: Suren Baghdasaryan <surenb@google.com> Cc: Liam R. Howlett <liam@infradead.org> Cc: Vlastimil Babka <vbabka@kernel.org> Cc: Jann Horn <jannh@google.com> Cc: Pedro Falcato <pfalcato@suse.de> Cc: David Hildenbrand <david@kernel.org> Cc: Mike Rapoport <rppt@kernel.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Jonathan Corbet <corbet@lwn.net> Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org> Cc: Dennis Dalessandro <dennis.dalessandro@cornelisnetworks.com> Cc: Jason Gunthorpe <jgg@ziepe.ca> Cc: Leon Romanovsky <leon@kernel.org> Cc: Paul Moore <paul@paul-moore.com> Cc: Stephen Smalley <stephen.smalley.work@gmail.com> Cc: Jaroslav Kysela <perex@perex.cz> Cc: Takashi Iwai <tiwai@suse.com> Cc: Alexei Starovoitov <ast@kernel.org> Cc: Daniel Borkmann <daniel@iogearbox.net> Cc: Andrii Nakryiko <andrii@kernel.org> Cc: Eduard Zingerman <eddyz87@gmail.com> Cc: Kumar Kartikeya Dwivedi <memxor@gmail.com> Cc: Zi Yan <ziy@nvidia.com> Cc: Baolin Wang <baolin.wang@linux.alibaba.com> Cc: Nico Pache <nico.pache@linux.dev> Cc: Ryan Roberts <ryan.roberts@arm.com> Cc: Dev Jain <dev.jain@arm.com> Cc: Barry Song <baohua@kernel.org> Cc: Lance Yang <lance.yang@linux.dev> Cc: Usama Arif <usama.arif@linux.dev> Cc: Kiryl Shutsemau <kas@kernel.org> Cc: Doug Gilbert <dgilbert@interlog.com> Cc: James Bottomley <james.bottomley@hansenpartnership.com> Cc: Martin K. Petersen <mkp@kernel.org> Cc: Simona Vetter <simona@ffwll.ch> Cc: Helge Deller <deller@gmx.de> Cc: Sebastian Reichel <sre@kernel.org> Cc: John Hubbard <jhubbard@nvidia.com> Cc: Peter Xu <peterx@redhat.com> Cc: Masami Hiramatsu <mhiramat@kernel.org> Cc: Oleg Nesterov <oleg@redhat.com> Cc: Peter Zijlstra <peterz@infradead.org> Cc: Thomas Gleixner <tglx@kernel.org> Cc: Ingo Molnar <mingo@redhat.com> Cc: Borislav Petkov <bp@alien8.de> Cc: Dave Hansen <dave.hansen@linux.intel.com> Cc: Arnaldo Carvalho de Melo <acme@kernel.org> Cc: Namhyung Kim <namhyung@kernel.org> Cc: Mark Rutland <mark.rutland@arm.com> Cc: Rik van Riel <riel@surriel.com> Cc: Harry Yoo <harry@kernel.org> Cc: Juri Lelli <juri.lelli@redhat.com> Cc: Vincent Guittot <vincent.guittot@linaro.org> Cc: Maarten Lankhorst <maarten.lankhorst@linux.intel.com> Cc: Maxime Ripard <mripard@kernel.org> Cc: Thomas Zimmermann <tzimmermann@suse.de> Cc: David Airlie <airlied@gmail.com> Cc: Will Deacon <will@kernel.org> Cc: Aneesh Kumar K.V <aneesh.kumar@kernel.org> Cc: Nicholas Piggin <npiggin@gmail.com> Cc: Arnd Bergmann <arnd@arndb.de> Cc: Muchun Song <muchun.song@linux.dev> Cc: Oscar Salvador <osalvador@suse.de> Cc: Matthew Wilcox (Oracle) <willy@infradead.org> Cc: Jan Kara <jack@suse.cz> Cc: Marc Zyngier <maz@kernel.org> Cc: Oliver Upton <oupton@kernel.org> Cc: Catalin Marinas <catalin.marinas@arm.com> Cc: Madhavan Srinivasan <maddy@linux.ibm.com> Cc: Anup Patel <anup@brainfault.org> Cc: Paul Walmsley <pjw@kernel.org> Cc: Palmer Dabbelt <palmer@dabbelt.com> Cc: Albert Ou <aou@eecs.berkeley.edu> Cc: Christian Borntraeger <borntraeger@linux.ibm.com> Cc: Janosch Frank <frankja@linux.ibm.com> Cc: Claudio Imbrenda <imbrenda@linux.ibm.com> Cc: Alexander Gordeev <agordeev@linux.ibm.com> Cc: Gerald Schaefer <gerald.schaefer@linux.ibm.com> Cc: Heiko Carstens <hca@linux.ibm.com> Cc: Vasily Gorbik <gor@linux.ibm.com> Cc: David S. Miller <davem@davemloft.net> Cc: Andreas Larsson <andreas@gaisler.com> Cc: Al Viro <viro@zeniv.linux.org.uk> Cc: Christian Brauner <brauner@kernel.org> Cc: Matthew Brost <matthew.brost@intel.com> Cc: Joshua Hahn <joshua.hahnjy@gmail.com> Cc: Rakie Kim <rakie.kim@sk.com> Cc: Byungchul Park <byungchul@sk.com> Cc: Gregory Price <gourry@gourry.net> Cc: Ying Huang <ying.huang@linux.alibaba.com> Cc: Alistair Popple <apopple@nvidia.com> Cc: Chris Li <chrisl@kernel.org> Cc: Kairui Song <kasong@tencent.com> Cc: Kemeng Shi <shikemeng@huaweicloud.com> Cc: Nhat Pham <nphamcs@gmail.com> Cc: Baoquan He <baoquan.he@linux.dev> Cc: Youngjun Park <youngjun.park@lge.com> Cc: Johannes Weiner <hannes@cmpxchg.org> Cc: Qi Zheng <qi.zheng@linux.dev> Cc: Shakeel Butt <shakeel.butt@linux.dev> Cc: Axel Rasmussen <axelrasmussen@google.com> Cc: Yuanchu Xie <yuanchu@google.com> Cc: Wei Xu <weixugc@google.com> Cc: Chengming Zhou <chengming.zhou@linux.dev> Cc: Michal Hocko <mhocko@kernel.org> Cc: Miklos Szeredi <miklos@szeredi.hu> Cc: Xu Xin <xu.xin@linux.dev>
3 daysmm: make per-VMA locks available universallyDave Hansen
Patch series "mm: Unconditional per-VMA locks and cleanups", v7. tl;dr: Make per-VMA locks available in all configs. Simplify some of the per-VMA lock users now that they can rely on them being always available. Binder and networking folks: Your code is the target of the cleanups. I'm cc'ing you now on v2 because there's emerging consensus on the mm side that the approach here is sane. I'm not quite sure how this pile would get merged, but ack/review tags would be appreciated if this looks good to you. Longer version: When working on some x86 shadow stack code, it was a real pain to avoid causing recursive locking problems with mmap_lock. One way to avoid those was to avoid mmap_lock and use per-VMA locks instead. They are great, but they are not available in all configs which makes them unusable in generic code, or if you want to completely avoid mmap_lock. Make per-VMA locks available in all configs. Right now, they are only available on select architectures when SMP and MMU are enabled. But all of the primitives that per-VMA locks are built on (RCU, maple trees, refcounts) work just fine without SMP or MMU. The only real downside is that making VMAs a wee bit bigger on !MMU and !SMP builds. The upside is much cleaner code, lower complexity and less #ifdeffery. Clean up a binder VMA locking site now that it can rely on per-VMA locks. Building on top of universally-available per-VMA locks, introduce a new helper. Since the new API does not require callers to have a fallback to mmap_lock, it's much easier to use. Callers can potentially replace this very common kernel idiom: mmap_read_lock(mm); vma = vma_lookup() // fiddle with vma mmap_read_unlock(mm); with: vma = vma_start_read_unlocked(mm, address); // fiddle with vma vma_end_read(vma); Which avoids mmap_lock entirely in the fast path. Use that new API for another binder site and one in the TCP code. This patch (of 7): The per-VMA locks have been around for several years. They've had some bugs worked out of them and have seen quite wide use. However, they are still only available when architectures explicitly enable them. Remove the conditional compilation around the per-VMA locks, making them available on all architectures and configs. The approach up to now seemed to be to add ARCH_SUPPORTS_PER_VMA_LOCK when the architecture started using per-VMA locks in the fault handler. But, contrary to the naming, the Kconfig option does not really indicate whether the architecture supports per-VMA locks or not. It is more of a marker for whether the architecture is likely to benefit from per-VMA locks. To me, the most important thing side-effect of universal availability is letting per-VMA locks be used in SMP=n configs. This lets us use per-VMA locking in all x86 code without fallbacks. Overall, this just generally makes the kernel simpler. Just look at the diffstat. It also opens the door to users that want to use the per-VMA locks in common code. Doing *that* brings additional simplifications. The downside of this is adding some fields to vm_area_struct and mm_struct. There are likely ways to optimize this, especially for things like SMP=n configs. For now, do the simplest thing: use the same implementation everywhere. == Considerations for NOMMU config == NOMMU systems do not write-lock VMAs, therefore read-locking a VMA would always succeed unless VMA is detached. Therefore for NOMMU config we make vma_mark_attached() a NOOP, which keeps VMAs always in detached state. This causes VMA read-locking to always fail and the caller falls back to locking mmap_lock. The following functions will have a different implementation in NOMMU config: - vma_mark_attached(), vma_mark_detached() are made NOOPs, keeping VMAs always in a detached state and preventing assertions and refcount underflows; - vma_start_write(), vma_start_write_killable() are made NOOPs to avoid warnings in __vma_start_write() due to VMAs being detached. These functions are not used in NOMMU code but __vma_start_write() is an exported function, therefore might be used by drivers. - vma_assert_attached() is made NOOP because it's reachable from NOMMU code via split_vma()->vma_iter_store_new()->vma_iter_store_overwrite(); - vma_assert_write_locked() is asserting vma->vm_mm is write-locked, as was done before this change; - vma_assert_locked() is asserting vma->vm_mm is locked, as was done before this change; The following functions work for both MMU and NOMMU configs: - vma_lock_init() performs the same initialization as for MMU config; - mm_lock_seqcount_init(), mm_lock_seqcount_begin(), mm_lock_seqcount_end() are called from mmap_write_{lock|unlock} and update mm_lock_seq correctly. - mmap_lock_speculate_try_begin(), mmap_lock_speculate_retry() work as is because mm_lock_seq is updated correctly; - vma_start_read(), vma_start_read_locked() will always fail because VMAs are always detached; - vma_end_read() will never be called because vma_start_read() never succeeds; - vma_is_attached() always return false because VMAs are always detached; - vma_assert_detached() will never trigger because VMAs are never attached; - vma_start_read_locked() always return false because VMAs are always detached; - lock_vma_under_rcu() will be safe as the attempted read lock will bail; Changes in the following files are not affecting NOMMU config: task_mmu.c - not compiled when CONFIG_MMU=n; pagewalk.c - not compiled when CONFIG_MMU=n; userfaultfd.c - not compiled when CONFIG_MMU=n (CONFIG_USERFAULTFD depends on CONFIG_MMU); The following changes in the BPF code are made to keep NOMMU config working like before: stack_map_lock_vma() - keeps mmap_lock in NOMMU config; bpf_iter_task_vma_new() - bails out in NOMMU config; Link: https://lore.kernel.org/20260831203056.838265-1-surenb@google.com Link: https://lore.kernel.org/20260831203056.838265-2-surenb@google.com Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com> Signed-off-by: Suren Baghdasaryan <surenb@google.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Reviewed-by: Lorenzo Stoakes (ARM) <ljs@kernel.org> Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org> Acked-by: David Hildenbrand (Arm) <david@kernel.org> Cc: Liam R. Howlett <liam@infradead.org> Cc: Shakeel Butt <shakeel.butt@linux.dev> Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org> Cc: Todd Kjos <tkjos@android.com> Cc: Christian Brauner <christian@brauner.io> Cc: Carlos Llamas <cmllamas@google.com> Cc: Alice Ryhl <aliceryhl@google.com> Cc: David S. Miller <davem@davemloft.net> Cc: David Ahern <dsahern@kernel.org> Cc: Arve Hjønnevåg <arve@android.com>
3 daysMerge branch 'vfs-7.4.netfs' into vfs.allChristian Brauner
Signed-off-by: Christian Brauner <brauner@kernel.org>
3 daysMerge branch 'vfs-7.4.idmap' into vfs.allChristian Brauner
3 daysMerge branch 'vfs-7.4.file' into vfs.allChristian Brauner
Signed-off-by: Christian Brauner <brauner@kernel.org>
3 daysnetfs: Remove folio_queueDavid Howells
Remove folio_queue as it's no longer used. Signed-off-by: David Howells <dhowells@redhat.com> Link: https://patch.msgid.link/20261001091239.3343034-11-dhowells@redhat.com Reviewed-by: Paulo Alcantara <pc@manguebit.org> cc: Matthew Wilcox <willy@infradead.org> cc: Christoph Hellwig <hch@infradead.org> cc: netfs@lists.linux.dev cc: linux-fsdevel@vger.kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
4 daysbpf: Destroy callee-local dynptrs on subprog returnXu Yunxiang
A data slice is valid only while the specific dynptr that produced it remains valid. prepare_func_exit() copies a subprogram return value into the caller and then frees the callee state without applying the normal destruction semantics to dynptrs in the callee stack. A callee can therefore derive a slice from a local dynptr, return or spill the slice into its caller, and then let the dynptr disappear with the callee frame. The slice retains the id of a dynptr that is no longer in the verifier state, so the verifier continues to accept accesses through it. Before the object relationship refactor, bpf_dynptr_data() slices from referenced dynptrs also carried the shared reference id, allowing a later release to catch some of these escaped slices. The refactor made those slices precise children of their source dynptr and exposed the missing teardown as an accepted stale access even after the shared resource is released. Add destroy_dynptrs_in_stack_slots() to apply the existing dynptr teardown to an inclusive range of stack slots. Reuse it for dynptr initialization, variable-offset stack writes, and the entire callee stack before freeing the frame. The teardown rejects losing the last dynptr for a referenced resource and invalidates the dynptr and all its descendants, including slices held in the caller. Propagate errors through prepare_func_exit(). Reuse the existing teardown API and diagnostic. Update the affected callback test expectation in this change so the commit remains test-clean. Fixes: 308c7a0ae885 ("bpf: Refactor object relationship tracking and fix dynptr UAF bug") Reported-by: Sashiko <sashiko-bot@kernel.org> Suggested-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Xu Yunxiang <xyx2021@mail.ustc.edu.cn> Reviewed-by: Amery Hung <ameryhung@gmail.com> Acked-by: Ihor Solodrai <ihor.solodrai@linux.dev> Link: https://lore.kernel.org/r/CAMB2axOmdHwSrHihSRskDywkCmEGQOX+6dZPN_A9u-HhxY3UDA@mail.gmail.com Link: https://lore.kernel.org/r/CAMB2axMmrTb+UZ84UD48MwMtXbk1s9bWtreU3AOfjzU0fPT0BA@mail.gmail.com Link: https://lore.kernel.org/bpf/20260927120422.1042107-2-xyx2021@mail.ustc.edu.cn Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
6 daysbpf: Allow a variable in DATASEC that is smaller than its typeAlexei Starovoitov
LLVM splits a static of a Rust program into pieces. Every piece is a VAR with the type of the whole static: [223] STRUCT 'BpfCell<core::option::Option<...>>' size=32 vlen=1 [236] VAR '..scx_cosmos9TASK_CTXS.0' type_id=223, linkage=static [248] DATASEC '.bss' size=0 vlen=9 type_id=236 offset=24648 size=1 (VAR '..TASK_CTXS.0') and the kernel rejects such BTF with "Invalid size". Allow it. The size of a variable in DATASEC is used in two places: - btf_find_datasec_var() looks for timers, spin locks, kptrs and other special fields. btf_find_field_one() skips a variable when its size is not the size of its type, so there are no special fields in a piece. - btf_datasec_show() prints variables by their types. It would read past the piece and past the end of the map value when the piece is the last one. Don't print the pieces. "global data test #8" and "#10" in prog_tests/btf.c expect "Invalid size". BTF of both is loaded now. The map of #10 is not created, since its value is larger than the section. Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://lore.kernel.org/bpf/20261002124714.180012-10-alexei.starovoitov@gmail.com Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
6 daysbpf: Allow arguments without names in static functions in BTFAlexei Starovoitov
Some static functions in BTF of a Rust program have arguments without names: [71] FUNC_PROTO '(anon)' ret_type_id=0 vlen=1 '(anon)' type_id=72 [81] FUNC 'unwrap_failed' type_id=71 linkage=static and the kernel rejects such BTF with "Invalid arg#1". It's 5 of 25 static functions in scx_cosmos and 6 of 24 in scx_simple. The verifier looks at types of the arguments. The names are printed only, as "(anon)" when there is none. Allow such static FUNC. Global functions are checked as before. The test "func (Some arg has no name)" in prog_tests/btf.c has a static function. Make it global to keep the check. Signed-off-by: Alexei Starovoitov <ast@kernel.org> Acked-by: Alan Maguire <alan.maguire@oracle.com> Link: https://lore.kernel.org/bpf/20261002124714.180012-8-alexei.starovoitov@gmail.com Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
6 daysbpf: Allow names of Rust types and functions in BTFAlexei Starovoitov
Names of types and functions in BTF that LLVM makes for a Rust program are not C identifiers: [69] STRUCT 'NonNull<str>' size=16 vlen=1 [147] FUNC 'write_fmt<scx_cosmos::BpfStream>' type_id=146 and the kernel rejects such BTF with "Invalid name". scx_simple and scx_cosmos schedulers written in Rust have them in STRUCT, FWD and FUNC, made of letters, digits and " #&()*,:;<>[]{}", 430 characters at most. Allow any printable character in btf_name_valid_identifier(), like it's done for DATASEC. It checks names of types, functions, members, enumerators, variables and arguments, so all of them can have such characters now. The limit of KSYM_NAME_LEN stays. The name of FUNC is a part of the name of the program in kallsyms, where a space would break the parsers. Replace what is not a character of an identifier with '_' there. Tests in prog_tests/btf.c expect "Invalid name" for names with '!' and '*', which are valid now. Put a character that is not printable there. The type name '?foo' is expected to load. Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://lore.kernel.org/bpf/20261002124714.180012-6-alexei.starovoitov@gmail.com Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
6 daysbpf: Treat load and store through a number as arena accessAlexei Starovoitov
There are no address spaces in Rust. LLVM emits plain loads and stores for arena memory, without cast_kern: r3 = 0x20 ll // R_BPF_64_64 .bss, libbpf puts it into arena r2 = *(u64 *)(r3 + 0) // R3 invalid mem access 'scalar' lock *(u64 *)(r2 + 8) += r1 Add BPF_F_ARENA_SCALAR flag of BPF_PROG_LOAD. When the program is loaded with it and has an arena treat ldx, stx, st and atomics through a number as arena access. It's as safe as access through PTR_TO_ARENA: JIT adds the base of arena to the low 32 bits of the address. The program has an arena only with CAP_BPF and CAP_PERFMON. Registers of the program are not changed, since Rust compares and stores the address after the access. bpf_do_misc_fixups() copies the low 32 bits into BPF_REG_AX and the access goes through it: r2 = *(u64 *)(r3 + 0) -> w12 = w3 r2 = *(u64 *)(r12 + 0) Constant blinding needs BPF_REG_AX for immediates and leaves insns that use BPF_REG_AX alone. So the immediate of st through a number is not blinded, the rest of the program is: *(u64 *)(r3 + 0) = 1 -> w12 = w3 *(u64 *)(r12 + 0) = 1 The same insn may see PTR_TO_ARENA on another path. It works for both. It's a flag, since a number is also what the verifier makes of a pointer that went away when the program has CAP_PERFMON: a ringbuf record after bpf_ringbuf_submit(), a pointer to the packet after bpf_skb_pull_data(). Programs in C that have an arena keep "invalid mem access 'scalar'" for such bugs. JIT tells with bpf_jit_supports_arena_scalar() that it takes BPF_REG_AX as the address of arena access. With other JITs the flag is rejected with -EOPNOTSUPP. x86 and arm64 do: - x86: the fault handler finds the register that holds the address through reg2pt_regs[]. Add BPF_REG_AX, r10 of x86, there. - arm64: nothing else is needed. The JIT takes any register as the address and the fault handler reads it by its number. Compile tested only. Any number is an address of arena, NULL and small numbers included, like it is for PTR_TO_ARENA: arena code in C dereferences NULL and relies on the fault being handled. Still rejected: - insn that sees a number on one path and a pointer that is not PTR_TO_ARENA on another. - numbers passed to helpers and kfuncs, except __arena arguments. Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://lore.kernel.org/bpf/20261002124714.180012-4-alexei.starovoitov@gmail.com Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
6 daysbpf: Allow bitwise ops, shifts and mul/div on pointers with CAP_PERFMONAlexei Starovoitov
Rust's core::fmt keeps a flag in the low bit of a pointer: r2 = *(u64 *)(r1 + 0) r3 = r2 r3 &= 1 r2 >>= 1 The verifier rejects it: r0 &= 8 R0 bitwise operator &= on pointer prohibited Negation and byte swap of a pointer already produce a number when allow_ptr_leaks is set. Do the same for bitwise ops, shifts, *=, /= and %=, 64-bit and 32-bit. The result is an unknown number. With CAP_PERFMON the program can store the pointer and load it back as a number already. += and -= are not changed: they keep the pointer, 32-bit += is rejected. Pointers that allow no arithmetic and pointers that may be NULL are still rejected. Two tests in verifier_value_illegal_alu store through the result of &= and /=. They are rejected at the store now. Update expected messages. Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://lore.kernel.org/bpf/20261002124714.180012-2-alexei.starovoitov@gmail.com Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
7 daysbpf: Roll back freplace link state when bpf_arch_text_poke() failsYuan Chen
A freplace attach claims the target prog by bumping tgt_prog->aux->freplace_link_cnt and setting tr->extension_prog. bpf_arch_text_poke() then makes the extension take effect. If the poke fails, the claims are never released: the attach unwinds through bpf_link_cleanup(), which clears link->prog, so bpf_trampoline_unlink_prog() never runs. Drop the link count under ext_mutex on the error path, and set tr->extension_prog only after the poke succeeded. The count is still bumped before the poke: it blocks prog_array updates while the entry is patched. Fixes: d6083f040d5d ("bpf: Prevent tailcall infinite loop caused by freplace") Fixes: be8704ff07d2 ("bpf: Introduce dynamic program extensions") Suggested-by: Leon Hwang <leon.hwang@linux.dev> Signed-off-by: Yuan Chen <chenyuan@kylinos.cn> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Acked-by: Leon Hwang <leon.hwang@linux.dev> Link: https://patch.msgid.link/20260928091121.219333-1-chenyuan_fl@163.com
7 daysMerge git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf 7.3-rc5Alexei Starovoitov
Cross-merge BPF and other fixes after downstream PR. Conflicts: arch/arm64/net/bpf_jit_comp.c arch/x86/net/bpf_jit_comp.c Signed-off-by: Alexei Starovoitov <ast@kernel.org>
7 daysbpf: Support per-call-site kfunc specializationEmil Tsalapatis
specialize_kfunc() currently updates the canonical kfunc descriptor in place. Different call sites therefore cannot select different specializations of the same kfunc, and the selected target depends on verification order. Keep canonical descriptors unchanged for verifier lookups. Specialize a copy for each call site, then reuse or append an immutable target descriptor and store its index in the finalized call instruction. Allow space for one canonical and one specialized target per kfunc. This is conservative because most kfuncs do not specialize. Adding specialized kfunc versions does not require resorting the array because lookups are only used for deduplication and func_model lookup. Deduplication for specialized kfuncs is handled by the specialization process itself, while all specializations of a function share the same proto and func_model. Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20261002105218.6171-7-emil@etsalapatis.com
7 daysbpf: Directly store kfunc desc index in instruction off fieldEmil Tsalapatis
Currently, the verifier sort the kfunc desc table twice: Once during kfunc collection by function ID, useful during verification, and once by immediate address, useful during JIT bytecode lowering time. The latter sort can be removed by storing in the instruction the kfunc desc array index the desc is in and using that instead. Modify the JITs to follow the same convention. This change simplifies the subsequent patch that adds per-call site function specialization. Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20261002105218.6171-6-emil@etsalapatis.com
7 daysbpf: Add sleepable arena page allocation pathEmil Tsalapatis
The bpf_arena_alloc_pages() function currently only allocates pages inside a spinlock critical section with IRQs off. This forces the use of alloc_pages_nolock() in the BPF allocator, even when the caller is a sleepable BPF function. This in turn causes allocation failures even in cases where falling into the allocator slow path and possibly sleeping would eventually succeed. This can be triggered consistently by heavy BPF arena users like scx. Allocate the arena pages before taking the critical section and pass whether the caller can sleep to bpf_alloc_pages(). This lets sleepable callers use the blocking allocator while non-sleepable callers retain the no-lock allocation behavior. Fixes: b8467290edab ("bpf: arena: make arena kfuncs any context safe") Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20261002105218.6171-4-emil@etsalapatis.com
7 daysbpf: Add sleepable argument to bpf_alloc_pages()Emil Tsalapatis
bpf_alloc_pages() currently decides whether it may use the blocking page allocator from the current execution context alone. Let callers further restrict that choice by passing whether their context is sleepable. Use the blocking allocator only when both the caller and runtime context allow sleeping. Add __GFP_RETRY_MAYFAIL so this path can reclaim without invoking the OOM killer when the allocation is charged to another memcg. Existing non-sleepable callers retain the no-lock allocation behavior. Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20261002105218.6171-3-emil@etsalapatis.com
7 daysbpf: Use an llist for page allocationsEmil Tsalapatis
bpf_map_alloc_pages() does not use its map argument. Storing allocated pages in an array also forces arena callers to allocate a separate pointer array. Expose the single-page allocator as bpf_alloc_page(), rename the bulk helper to bpf_alloc_pages(), and return bulk allocations through an llist using page->pcp_llist. Add bpf_free_pages() to safely release all pages remaining on such a list. Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20261002105218.6171-2-emil@etsalapatis.com
7 daysbpf: Fix missing migration protection in __rhtab_map_lookup_and_delete_batch()Ömer Mete Kaya
bpf_mem_cache_free_rcu() uses this_cpu_ptr() which requires migration to be disabled. All callers of rhtab_delete_elem() disable migration except __rhtab_map_lookup_and_delete_batch(), which calls it under rcu_read_lock() only. On CONFIG_PREEMPT_RCU, rcu_read_lock() does not disable preemption or migration, so the task can migrate between CPUs during the delete loop, causing this_cpu_ptr() to trigger: BUG: using smp_processor_id() in preemptible [00000000] code Fix by wrapping the delete loop in migrate_disable()/migrate_enable() in __rhtab_map_lookup_and_delete_batch(), matching the migration protection that the other callers already provide. Fixes: 818e00848227 ("bpf: Implement iteration ops for resizable hashtab") Reported-by: syzbot+fd7e415d891073b83e1f@syzkaller.appspotmail.com Signed-off-by: Ömer Mete Kaya <omermetekaya0@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Acked-by: Mykyta Yatsenko <yatsenko@meta.com> Link: https://patch.msgid.link/20260929081609.557899-1-omermetekaya0@gmail.com Closes: https://syzkaller.appspot.com/bug?extid=fd7e415d891073b83e1f
8 daysbpf: Fix objects stuck in free_by_rcu_ttraceAlexei Starovoitov
do_call_rcu_ttrace() returns early when call_rcu_ttrace_in_progress is set and leaves the objects in free_by_rcu_ttrace. __free_rcu() frees waiting_for_gp_ttrace only and clears the flag. Hence the objects that free_bulk() or __free_by_rcu() added while RCU tasks trace GP was in flight stay in free_by_rcu_ttrace until free_bulk() or alloc_bulk() is called for the same bpf_mem_cache again, which may never happen. The number of such objects is not bounded. Turn call_rcu_ttrace_in_progress into three states: 0 - idle 1 - __free_rcu() is queued 2 - __free_rcu() is queued and free_by_rcu_ttrace got more objects since do_call_rcu_ttrace() sets 2. __free_rcu() does cmpxchg(1 -> 0) and starts the next GP when it fails. It cannot clear the flag first and check free_by_rcu_ttrace later, since bpf_mem_alloc_destroy() frees bpf_mem_cache without waiting for RCU callbacks when the flag is zero. Now __free_rcu() queues itself, so the one that didn't see 'draining' may do call_rcu_tasks_trace() after rcu_barrier_tasks_trace() in free_mem_alloc(). Queue it under rcu_read_lock() and do synchronize_rcu() before the barriers. Calling rcu_barrier_tasks_trace() twice works too, but creating and destroying hash maps in a loop on many cpus slows down to one free_mem_alloc() per GP and kworkers pile up. Fixes: 8d5a8011b35d ("bpf: Batch call_rcu callbacks instead of SLAB_TYPESAFE_BY_RCU.") Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://lore.kernel.org/bpf/20260930095920.601738-3-alexei.starovoitov@gmail.com Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
8 daysbpf: Factor out __do_call_rcu_ttrace()Alexei Starovoitov
Move the part of do_call_rcu_ttrace() that runs after call_rcu_ttrace_in_progress is set into __do_call_rcu_ttrace(). The next patch will call it from __free_rcu(). No functional change. Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://lore.kernel.org/bpf/20260930095920.601738-2-alexei.starovoitov@gmail.com Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
8 daysbpf: Fix packet range of pointers sharing an idAlexei Starovoitov
Since commit 022ac0750883 ("bpf: use reg->var_off instead of reg->off for pointers"), find_good_pkt_pointers() sets the range of all packet pointers sharing an id from the umax of the compared pointer, and check_packet_access() requires umax + off + size <= range. That assumes the umax of two such pointers differ by exactly their constant distance. reg_bounds_sync() breaks it when var_off tightens one umax and not the other: r4 &= 0x38 if r4 > 50 goto exit ; umax 50, var_off (0x0; 0x38) r5 = pkt + r4 ; umax 50 r6 = r5 r6 += 8 ; umax 56, not 58 Comparing r6 with pkt_end sets the range to 56, and the valid 8-byte load at r5 is rejected (50 + 8 > 56). Comparing r5 sets it to 50, and the out-of-bounds 1-byte load at r6 - 7, i.e. r5 + 1, is accepted (56 - 7 + 1 <= 50). Don't call reg_bounds_sync() on a packet pointer that keeps its id (a constant was added or subtracted) or its range (an unknown non-negative value was subtracted), so that var_off cannot tighten its umax. Only update the 32-bit bounds from var_off: reg_bounds_sanity_check() wants them constant when the lower half of var_off is, e.g. for pkt + 8. This relies on nothing else changing the 64-bit bounds of a packet pointer, which holds today. var_off of such a pointer is no longer narrowed by its bounds. Adjust three verifier_align expectations; the low bits, which the alignment checks use, don't change. veristat on the selftests shows no verdict changes and +0.8% insns in test_cls_redirect_subprogs. Fixes: 022ac0750883 ("bpf: use reg->var_off instead of reg->off for pointers") Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://lore.kernel.org/bpf/20261001145255.855630-1-alexei.starovoitov@gmail.com Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
8 daysbpf, cgroup: Fix cgroup struct_ops query for a second attach typeShakeel Butt
Two things in __cgroup_bpf_query() work only because CGROUP_TCP_SOCK_OPS is the only struct_ops attach type. It calls cgroup_bpf_enabled(atype) with an atype that find_atype_by_struct_ops_id() works out at runtime. That macro is an asm goto and needs a constant. Today the compiler can see there is only one value; add a second type and the build breaks with "impossible constraint in 'asm'". The check is not needed anyway. When the key is off nothing is attached, so progs[atype] and effective[atype] are empty and the count loop leaves total_cnt at 0. The other attach types already use that loop with no such check. Drop the check and the skip_count label. And find_atype_by_struct_ops_id() matches on type_id alone. An attach type whose subsystem is not built keeps type_id 0, so a query for type 0 finds it and returns success with nothing instead of -ENOENT. Skip such slots. Fixes: 369d9dcd8fb8 ("bpf: Add infrastructure to support attaching struct_ops to cgroups") Signed-off-by: Shakeel Butt <shakeel.butt@linux.dev> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Reviewed-by: Amery Hung <ameryhung@gmail.com> Acked-by: Yafang Shao <laoar.shao@gmail.com> Acked-by: Tejun Heo <tj@kernel.org> Link: https://patch.msgid.link/20261001141901.3225830-1-shakeel.butt@linux.dev
9 daysbpf: Correct Program Structure diagnostic contextKumar Kartikeya Dwivedi
Program Structure reports have two attribution gaps. A missing jump table is reported at the beginning of its subprogram rather than at the gotox that needs the table, and the subprogram-layout checks run before func_info and line_info are validated and installed, so their reports cannot name the source function or line. Report the missing jump table at the gotox instruction that looks it up, rather than at the start of its subprogram. func_info and line_info validation only needs the subprogram boundaries found by add_subprogs() and the LD_ABS and tail-call properties that check_subprogs() collects while scanning instructions; it does not depend on the layout checks themselves. Split the property collection into its own nested subprogram and instruction scan, together with the program's callx marker, and run bpf_check_btf_info() before check_subprogs(). The jump-boundary and fallthrough reports then carry validated source information without changing either check. CO-RE relocations are unaffected: commit c26e97721b17 ("bpf: Apply CO-RE relocations before subprogram validation") applies them before add_subprogs(), and they stay there. Only the func_info and line_info validation moves. Fixes: a8f427835394 ("bpf: Report Program Structure CFG errors") Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Daniel Borkmann <daniel@iogearbox.net> Link: https://lore.kernel.org/bpf/cf2f420c2b21de440a7dc51b1565c0f06d4b539640ee5c03384e5d77bcfb5686@mail.kernel.org/ Link: https://lore.kernel.org/bpf/20260926133048.2962553-3-memxor@gmail.com
10 daysbpf: Verify BTF_KIND_LOC_PARAM vlen, flagsAlan Maguire
Ensure that vlen and flags combinations are consistent for BTF_KIND_LOC_PARAMs. Fixes: 33c5a3278bdb ("btf: Extend UAPI to support BTF location (inline site) info") Suggested-by: Alexei Starovoitov <ast@kernel.org> Signed-off-by: Alan Maguire <alan.maguire@oracle.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260926175113.2368566-2-alan.maguire@oracle.com
10 daysbpf: Check all subprog arguments in the common pathAmery Hung
All supported BPF-subprogram argument types now pass through check_func_arg() from one guarded branch in the legacy validation loop. Replace that loop and its outgoing-stack validation with check_func_args(). The scratch prototype is indexed by BTF parameter and the call metadata carries the BTF function prototype. The existing BTF argument iteration therefore maps parameters to ABI slots for both kfunc and BPF subprogram calls, including additional slots occupied by by-value aggregates. The common checker can return non-EFAULT errors other than -EINVAL, as the existing BTF-ID path already did. This is compatible with subprog call handling: only -EFAULT is fatal. Any other error marks BTF unreliable. Static subprogs can then fall back to inline verification, while global subprogs reject the call. Remove the caller-register plumbing that the dedicated loop required. Signed-off-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260928185334.1004200-12-ameryhung@gmail.com
10 daysbpf: Check global subprog memory arguments in the common pathAmery Hung
Route fixed-size global ARG_PTR_TO_MEM arguments through check_func_arg(). Preserve their read-write access check, nullable contract, and support for BTF-defined allocated memory. Recognize subprog calls explicitly in the common checker. Global subprog stack liveness can prove that bytes in an argument are unused by the callee. Retain the existing allowance for those poisoned stack bytes when call metadata is present, and update the nullability log expectation. The verifier also rejects packet pointers when the callee may change packet data. The concrete packet-backed register state is only available while checking the call site; the independently verified callee sees generic PTR_TO_MEM. Record this property as subprog_may_change_pkt in subprog-only call metadata and retain the packet-pointer rejection in the common fixed-memory path. Signed-off-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260928185334.1004200-11-ameryhung@gmail.com
10 daysbpf: Check global subprog BTF-ID arguments in the common pathAmery Hung
Route global ARG_PTR_TO_BTF_ID arguments through check_func_arg(). Use vmlinux BTF for their type match and retain the global-subprogram rules instead of applying kfunc-only trusted-pointer validation or runtime type resolution. The common nullability check now rejects non-nullable global BTF-ID arguments before register-type matching. Update the corresponding verifier log expectations. Signed-off-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260928185334.1004200-10-ameryhung@gmail.com
10 daysbpf: Check subprog dynptr arguments in the common pathAmery Hung
Subprog and helper/kfunc dynptr arguments ultimately use the same process_dynptr_func() validation. Route the subprog dynptr branch through check_func_arg() so register type, offset, and dynptr state are checked in the common order. check_func_arg() validates the register against the PTR_TO_STACK and CONST_PTR_TO_DYNPTR compatibility set before dispatching to process_dynptr_func(). Once the subprog path uses the common checker, process_dynptr_func() has no caller that bypasses this validation. Remove its now-redundant register-type check. The existing wrong-register-type test passed a NULL local to a callee that did not use the argument. Clang left the context pointer in R1, while GCC materialized zero. The common checker rejected the values at different stages and emitted different diagnostics. Obtain a PTR_TO_BTF_ID from bpf_get_current_task_btf() and keep its callee argument live. This makes both compilers exercise the intended register-type mismatch. Signed-off-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260928185334.1004200-9-ameryhung@gmail.com
10 daysbpf: Check subprog arena arguments in the common pathAmery Hung
ARG_PTR_TO_ARENA accepts both arena pointers and scalars. The existing BPF-subprogram contract includes a constant-zero scalar; it is an arena address value rather than a conventional pointer rejected by nullability checking. Kfunc prototype generation already records this zero acceptance with PTR_MAYBE_NULL for both arena suffixes; separate JIT metadata determines whether the kernel callee receives the arena base or NULL. Apply the same call-site representation to the temporary BPF-subprogram prototype, then extend the guarded check_func_arg() path to cover arena arguments. Signed-off-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260928185334.1004200-8-ameryhung@gmail.com
10 daysbpf: Check subprog context arguments in the common pathAmery Hung
Subprog and helper/kfunc context arguments require PTR_TO_CTX with an acceptable offset. Route subprog context arguments through check_func_arg() and update the verifier-log expectation. Signed-off-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260928185334.1004200-7-ameryhung@gmail.com
10 daysbpf: Check global subprog untrusted arguments in the common pathAmery Hung
Global subprogram arguments tagged __arg_untrusted deliberately accept any caller value. The callee is verified separately with a read-only untrusted pointer, whose accesses use protected probe-read instructions. Keep the untrusted type in the canonical bpf_subprog_info used to prepare the callee state. Represent it as ARG_IGNORE only in the temporary bpf_func_proto used for call-site checking, then extend the guarded check_func_arg() path to cover it. Signed-off-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260928185334.1004200-6-ameryhung@gmail.com
10 daysbpf: Check subprog scalar arguments in the common pathAmery Hung
BPF subprogram scalar arguments are checked separately even though the common argument checker enforces the same SCALAR_VALUE requirement. Start the migration by routing the scalar branch through check_func_arg() and update the verifier-log expectation. Use the BTF parameter index for metadata and the ABI slot index for register lookup. Route extra slots of by-value aggregates through check_arg_extra_slot() so every occupied slot follows the common path. Following patches can extend this guarded common-check branch as each remaining BPF-subprogram-specific implementation is removed. Signed-off-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260928185334.1004200-5-ameryhung@gmail.com
10 daysbpf: Build argument prototypes for subprog callsAmery Hung
The common argument checker consumes a parameter-indexed bpf_func_proto, while BPF subprogram argument metadata is cached by ABI slot in compact bpf_subprog_info entries. Embedding a full prototype in every fixed subprogram entry would waste memory. Embed a scratch prototype in the verifier environment. Fill it by walking BTF parameters and the cached slots with separate cursors, mirroring the kfunc representation for parameters that occupy multiple slots. Keep the compact cache as the source of truth. The verifier environment is already heap allocated, so this avoids a separate allocation, failure path, and cleanup. Keep the existing validation loop, but make it consume the generated prototype in preparation for moving BPF subprogram calls to the common argument checker. No functional change. Signed-off-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260928185334.1004200-4-ameryhung@gmail.com
10 daysbpf: Identify subprog calls in argument metadataAmery Hung
Prepare subprog calls to use the common check_func_args() path. That checker distinguishes call kinds through bpf_call_arg_meta, but btf_check_func_arg_match() leaves both BTF and the function ID zero even though it validates a program-BTF signature. Represent a subprog call with a non-NULL BTF and a zero function ID. Define helpers as NULL BTF plus nonzero ID, and kfuncs as non-NULL BTF plus nonzero ID. Requiring a nonzero kfunc ID also prevents unavailable special-kfunc IDs from matching subprog metadata. The corresponding subprog predicate is introduced later alongside its first use. Setting BTF would expose checks that treat every BTF-backed call as a kfunc. Restrict kfunc-only BTF parameter lookup, register admission, release diagnostics, no-cast alias, and projection handling to actual kfuncs. The BTF is not modified here, but bpf_call_arg_meta::btf and several downstream consumers use non-const pointers. Match those existing types instead of broadening this series with a const-correctness cleanup. No functional change is intended. Signed-off-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260928185334.1004200-3-ameryhung@gmail.com
10 daysbpf: Fix kfunc BTF parameter lookups after wide argumentsAmery Hung
Several kfunc argument checks index BTF parameters using the ABI slot number. Parameter and slot indexes diverge after a by-value argument wider than one eightbyte, so a later argument can select the wrong BTF parameter or index past the end of the function prototype. This affects the BTF lookups in check_func_arg_nullability() and check_func_arg_release(), and the parameter passed to btf_check_iter_arg() by process_iter_arg(). check_func_arg() already tracks the BTF parameter and ABI slot separately. Pass the parameter index to all three consumers while retaining argno for register and stack-slot diagnostics. Signed-off-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260928185334.1004200-2-ameryhung@gmail.com
11 daysbpf: Skip detached progs in trampoline images that are still in useFlorent Revest (Anthropic)
bpf_tramp_image_put() makes sure a trampoline image is not freed while a task may still be running in it, but nothing similar is done for the progs called by that image. Detach drops the last prog reference right away and the prog is freed after grace periods, on the basis that a task still in the traced function skips the fexit progs once the nop at ip_after_call is patched to a jump. That leaves out a task sleeping in a sleepable prog that runs before the detached one, which no grace period waits for: CPU 0 CPU 1 in image I, sleeping in prog S detach P from I's trampoline -> new image, bpf_tramp_image_put(I) bpf_prog_put(P), last ref grace periods, P freed back from S __bpf_prog_enter(P) call P->bpf_func If S and P are fexit progs the task is already past the patched jump, and fentry only images don't have one. On x86 this is an int3 in poisoned bpf_prog_pack memory: Oops: int3: 0000 [#1] SMP NOPTI CPU: 18 UID: 0 PID: 94573 Comm: x169 Not tainted 6.18.44 #1 PREEMPT(lazy) RIP: 0010:0xffffffffc0601d8d Call Trace: <TASK> ? bpf_trampoline_6442515411+0x1a4/0x21b bpf_lsm_bprm_committed_creds+0x5/0x10 security_bprm_committed_creds+0x5f/0x70 begin_new_exec+0x2d6/0x410 ... We hit this in production when progs attached through trampolines got detached while their hooks were busy, and it was independently found with a fuzzer and KASAN. Have the JITs emit a patchable nop in front of each prog call sequence and record it in the image. When a prog is detached, patch its nop to a jump over the call sequence. Progs that stay attached keep running for the tasks that are in the image, and ip_after_call isn't needed anymore. The task can be in any image that isn't freed yet, not only in the current one. It sleeps in image I1 that calls S, P and Q, then P is detached and the trampoline moves to image I2, then Q is detached and its call is still in I1. So the trampoline keeps a list of its images until they are freed, and detaching a prog patches its nop in all of them. Images hold a reference on the trampoline for that long. On riscv and loongarch a jump of any range takes several instructions, and a task preempted in the middle of them could resume into half of the new sequence. The nop is a single instruction there, patched to a near branch through arch_bpf_trampoline_skip(). With the extra nops, BPF_MAX_TRAMP_LINKS progs no longer fit in a page on x86 and arm64 (and already didn't on powerpc), so lower the limit there like s390 does. Fixes: e21aa341785c ("bpf: Fix fexit trampoline.") Reported-by: Sechang Lim <rhkrqnwk98@gmail.com> Suggested-by: Alexei Starovoitov <ast@kernel.org> Signed-off-by: Florent Revest (Anthropic) <florent.revest@linux.dev> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260926135605.1217928-3-florent.revest@linux.dev Closes: https://lore.kernel.org/bpf/20260815071927.147049-1-zirajs7@gmail.com/
11 daysbpf: Wait for an RCU tasks grace period before freeing trampoline progsFlorent Revest (Anthropic)
When a prog is detached from a trampoline, it is freed after an RCU grace period, or an RCU tasks trace one if it is sleepable. This covers the tasks that are running the prog, since the prog's enter helper takes the matching read lock before the prog is called. On a preemptible kernel, it doesn't cover a task that was preempted in the trampoline just before the enter helper. That task holds no lock yet, and it calls the prog after it was freed: BUG: KASAN: vmalloc-out-of-bounds in __bpf_prog_enter_recur+0x3a5/0x3f0 Read of size 8 at addr ffffc90000055040 by task candidate/110 CPU: 1 UID: 0 PID: 110 Comm: candidate Not tainted 7.3.0-rc2-00014-g15071f2a1263-dirty #2 PREEMPT(full) Call Trace: <TASK> __bpf_prog_enter_recur+0x3a5/0x3f0 bpf_trampoline_6442509193+0x37/0xf1 __x64_sys_futex+0x9/0x410 do_syscall_64+0xb0/0x530 ... Wait for an RCU tasks grace period before the existing one when freeing a prog that was linked to a trampoline. An RCU tasks grace period only ends once the tasks that were preempted have run again, and bpf_tramp_image_put() already relies on it to free the image. It doesn't wait for tasks that sleep in the trampoline, the next commit takes care of those. Only progs that were linked to a trampoline can be called this way, so bpf_trampoline_add_prog() marks them and other progs are still freed as before. Fixes: e21aa341785c ("bpf: Fix fexit trampoline.") Reported-by: Junseo Lim <zirajs7@gmail.com> Signed-off-by: Florent Revest (Anthropic) <florent.revest@linux.dev> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260926135605.1217928-2-florent.revest@linux.dev Closes: https://lore.kernel.org/bpf/aqdrwVpanH3WGurX@omen-arch/
2026-09-25bpf: Add file descriptor interface for program streamsKumar Kartikeya Dwivedi
The existing BPF_PROG_STREAM_READ_BY_FD command only supports polling a program stream through repeated bpf() calls. It cannot block for new data or integrate with poll-based event loops. Add BPF_PROG_STREAM_OPEN to return a read-only, close-on-exec file descriptor for a selected program stream. Reads block by default and BPF_F_STREAM_NONBLOCK, the only accepted flag, provides non-blocking behavior. poll reports readable data and reports hangup once the program has been freed. Like pipes and sockets, the descriptor is not seekable and lseek fails with ESPIPE. A stream descriptor deliberately does not retain the program. Move each stream into a separately refcounted allocation so program teardown can mark it dead and wake descriptor users while outstanding descriptors drain buffered data safely. Readers sample the dead flag before looking for data, so EOF is reported only when the stream was already dead before it was found empty; data published right before teardown is never skipped. Only programs loaded through BPF_PROG_LOAD get streams. Classic BPF filters, JIT subprograms and shim programs never write to one, and kernel-side writers already resolve a subprogram to its main program, so those programs no longer carry stream state. Readiness needs its own counter. Stream capacity is charged before allocation and before an element is published to the stream log, so using that reservation as the read and poll condition can report readable data while no element exists: a blocking reader retries instead of sleeping and a lone non-blocking reader can see POLLIN followed by EAGAIN. Publish bytes with release ordering after adding elements to the lockless log, use acquire loads before consuming them or reporting readiness, limit each read to its readable snapshot and subtract only bytes actually copied. This keeps the aggregate count correct even when concurrent publishers update it out of publication order. With several readers on one stream, readiness remains advisory, as it is for pipes. The capacity counter is kept solely for enforcing the stream size limit. Wakeups are always deferred through irq_work. Stream writers run in whatever context the program runs in: NMI context for perf_event programs, sections with interrupts disabled inside bpf_spin_lock or rqspinlock critical sections since bpf_stream_vprintk() is KF_SPINLOCK_SAFE, and tracing programs attached anywhere in the kernel, including inside the wait queue and epoll code itself. Waking waiters directly from there can deadlock, and no cheap context check covers every case: on PREEMPT_RT, spinlock_t sections do not disable interrupts, so in_nmi() or irqs_disabled() cannot tell such a program apart from a benign one. Queue an irq_work item instead, as bpf_ringbuf does. Queue it only when a publication turns an empty stream readable. Readers block and pollers wait only after finding the stream empty, and the readable count never drops below zero because each read is bounded by its snapshot, so the first publication after such an observation is the one that makes the count positive, and it is the one that queues the wakeup. Publications into a stream that already holds data raise no interrupt, so a program that prints while nobody drains its stream pays for a single irq_work until the stream is emptied again. This matches bpf_ringbuf, which notifies only once the consumer has caught up. Blocking readers and level-triggered pollers re-check the readable count before waiting, so they cannot miss data, and edge-triggered epoll consumers drain until EAGAIN before waiting again, as epoll(7) requires. Synchronize pending work before releasing the final stream reference so the callback cannot outlive the stream, but only when the work was ever queued: irq_work_sync() waits for an RCU grace period on PREEMPT_RT and on architectures without an irq_work interrupt. Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260925045536.1480933-3-memxor@gmail.com
2026-09-25bpf: Skip zero-length stream writesKumar Kartikeya Dwivedi
A stream write whose formatted output is empty still allocates a stream element and links it into the stream log. Capacity accounting only counts payload bytes, so such elements never count against the stream limit, and a program that keeps producing empty output allocates elements without bound until a reader drains them. Return success without allocating anything when there is nothing to write, for both bpf_stream_vprintk() and staged writes. Every element in a stream log then carries at least one byte, which later changes rely on when they derive read readiness from published byte counts. Fixes: 5ab154f1463a ("bpf: Introduce BPF standard streams") Suggested-by: Emil Tsalapatis <emil@etsalapatis.com> Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260925045536.1480933-2-memxor@gmail.com
2026-09-25bpf: Retarget indirect jump targets across prologue prependsDaniel Borkmann
A few patches do not expand in place, they prepend: bpf_convert_ctx_accesses() puts the ctx save for gen_epilogue and the instructions of gen_prologue in front of insn 0, and bpf_do_misc_fixups() prepends the may_goto counter init at the start of every subprog that uses may_goto. They copy the original instruction of 'tgt_idx' into the last slot of the patch buffer, so it now lives at tgt_idx + delta, and call adjust_jmp_off() to move direct branches from tgt_idx to tgt_idx + delta. Nothing does the same for indirect branches, so a BPF_MAP_TYPE_INSN_ARRAY slot that named tgt_idx keeps naming tgt_idx, which is now the first prepended instruction, and insn_aux_data[tgt_idx].indirect_target makes the JIT emit the landing pad there. Add a __bpf_patch_insn_data() variant and pass BPF_PATCH_MOVE_TARGET at the affected call-sites to fix up the delta. The poke descriptors of direct tail calls are shifted the same way, and the original instruction is not searched for by content in that mode, as it is the last slot by construction. Fixes: 493d9e0d6083 ("bpf, x86: add support for indirect jumps") Fixes: 07ae6c130b46 ("bpf: Add helper to detect indirect jump targets") Reported-by: Nicholas Carlini <npc@anthropic.com> Suggested-by: Nicholas Carlini <npc@anthropic.com> Signed-off-by: Daniel Borkmann <daniel@iogearbox.net> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Acked-by: Anton Protopopov <a.s.protopopov@gmail.com> Acked-by: Eduard Zingerman <eddyz87@gmail.com> Link: https://patch.msgid.link/20260925175244.1136329-11-daniel@iogearbox.net
2026-09-25bpf: Clear the insn array used flag with release semanticsDaniel Borkmann
A program that fails after the JIT ran, for example in bpf_check_tail_call(), has already stored its jitted offsets and target pointers into the insn array maps it uses through bpf_prog_update_insn_ptrs() when release_insn_arrays() hands them back by clearing the used flag with a plain atomic_set(). Nothing orders those stores before the flag, so on a weakly ordered architecture they can still be in flight when the next program takes the map: it is acquired with atomic_xchg() in bpf_insn_array_init() and reset there, and the late stores then land on top of whatever the new owner has written since, be it that reset or the pointers its own JIT stored, leaving entries that point into an image which is about to be freed. Clear the flag with atomic_set_release() instead, so that everything the releasing program stored is visible before the map can be acquired again. This pairs with the fully ordered atomic_xchg() on the acquire side. Fixes: b4ce5923e780 ("bpf, x86: add new map type: instructions array") Signed-off-by: Daniel Borkmann <daniel@iogearbox.net> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260925175244.1136329-10-daniel@iogearbox.net
2026-09-25bpf: Clear insn array jitted pointers on map reuseDaniel Borkmann
An insn array records, per entry, the xlated offset, the jitted offset and a jitted target pointer in ips[]. bpf_insn_array_init() runs when the map is bound to a program and resets only xlated_off; the jitted offset and the ips[] pointer are left as the previous owner set them. A program can be verified and JITed and only then fail to load (e.g. bpf_check_tail_call()). By that point bpf_prog_update_insn_ptrs() has already filled ips[] with pointers into that program's JIT image. Thus, the map is handed back for reuse with the stale pointers intact. If the next program reuses the map, the stale pointer survives in the now active program's jump table, pointing into the previous, freed image. Fix by resetting the jitted offset and the ips[] pointer alongside the xlated offset. Fixes: b4ce5923e780 ("bpf, x86: add new map type: instructions array") Reported-by: Sandipan Roy <saroy@redhat.com> Reported-by: James Burton <jamesburton@meta.com> Signed-off-by: Daniel Borkmann <daniel@iogearbox.net> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Acked-by: Anton Protopopov <a.s.protopopov@gmail.com> Acked-by: Eduard Zingerman <eddyz87@gmail.com> Link: https://patch.msgid.link/20260925175244.1136329-9-daniel@iogearbox.net
2026-09-25bpf: Don't remove indirect jump targets in the nop removal passDaniel Borkmann
bpf_opt_remove_nops() drops every ja +0 and may_goto +0 unconditionally. A gotox target may itself be such a transparent nop: it is a verified entry of an insn array jump table that is reachable at run time through the indirect jump. bpf_insn_array_adjust_after_remove() marks it as INSN_DELETED instead, and bpf_insn_array_ready() skips INSN_DELETED entries, so the program loads with a hole in its jump table. The x86 JIT, for example, lowers gotox to a raw indirect jump, so executing that index jumps to the zeroed ips[] slot, that is, to address 0. Fix by keeping an indirect jump target in place. Fixes: 493d9e0d6083 ("bpf, x86: add support for indirect jumps") Reported-by: Sandipan Roy <saroy@redhat.com> Reported-by: James Burton <jamesburton@meta.com> Signed-off-by: Daniel Borkmann <daniel@iogearbox.net> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Acked-by: Anton Protopopov <a.s.protopopov@gmail.com> Acked-by: Eduard Zingerman <eddyz87@gmail.com> Link: https://patch.msgid.link/20260925175244.1136329-8-daniel@iogearbox.net
2026-09-25bpf: Reject indirect jumps that leave their subprogramDaniel Borkmann
The jump table of a subprog is collected in compute_subprog_jts() from the insn_array maps of the program, and a map is attributed to the subprog that contains its first entry. check_indirect_jump() instead resolves the targets from the map the gotox register actually points to, bounded only by the index range of that register, and never relates them back to the subprog of the gotox. The two disagree, so bpf_insn_successors() reports a subset of the edges the BPF program can take and a gotox can enter a subprog the CFG never walked. Close both ends: reject a map that reaches past the subprog holding its first entry right when it is collected in compute_subprog_jts(), so that every map of a program with a gotox is confined to one subprog, and confine the resolved targets of a gotox to its subprog in check_indirect_jump(). A target inside the subprog is then always in the jump table the CFG walked, which is the invariant that actually has to hold. The out-of-range check in visit_gotox_insn() cannot fire any more once every table is confined and becomes a verifier_bug_if(). Fixes: 493d9e0d6083 ("bpf, x86: add support for indirect jumps") Reported-by: James Burton <jamesburton@meta.com> Reported-by: Nuoqi Gui <gnq25@mails.tsinghua.edu.cn> Signed-off-by: Daniel Borkmann <daniel@iogearbox.net> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Acked-by: Anton Protopopov <a.s.protopopov@gmail.com> Acked-by: Eduard Zingerman <eddyz87@gmail.com> Link: https://patch.msgid.link/20260925175244.1136329-7-daniel@iogearbox.net
2026-09-25bpf: Cache the jump table of a subprogram during CFG discoveryDaniel Borkmann
create_jt() builds the jump table of the subprogram containing a gotox by copying out and sorting every insn_array map of the program, and it does so once per gotox instruction. The cost is therefore the number of gotox instructions times the number of entries in all of the maps. A program of 4003 instructions with 2000 gotox and one 500k entry map holding two distinct targets has 4000 indirect jump edges, 0.4% of the limit, and takes 343s to be rejected. The map costs next to nothing to prepare, as an unset entry is already a valid target. At the insn limit, with a single 1M entry map, the same shape extrapolates to 41 hours. All gotox instructions of a subprogram share the same jump table, so build the table of every subprogram in a single pass over the maps and let every gotox use it in place: bpf_insn_successors() resolves a gotox through its containing subprogram, and the tables are kept until the verifier environment is torn down, so nothing is copied into the instruction aux data. The edge accounting in visit_gotox_insn() moves to the first visit of the gotox, marked by the BRANCH bit of its CFG state, and a revisit returns right away since the first one already pushed every target that still needed exploring. The only user left of insn_aux_data[].jt is then the table which visit_abnormal_return_insn() allocates for tail_call and ld_{abs,ind} insns, so that bpf_insn_successors() reports the hidden exit from their subprogram. Both are recognisable by their opcode, so derive that edge in bpf_insn_successors() from the exit_idx of the containing subprogram instead, again leaving it out when the subprogram has no exit, and drop the field along with bpf_clear_insn_aux_data(). The instruction aux data then owns no allocation, thus nothing needs to be freed when insns are removed or the verifier environment is torn down. Subprograms removed as dead code free their table in adjust_subprog_starts_after_remove(), which also clears the slots that the compaction of subprog_info vacates, as they still hold copies of the moved entries and with them their table pointers. check_cfg() is then linear in the number of map entries plus the number of indirect jump edges, so what still scales now with the program is what BPF_MAX_GOTOX_EDGES bounds: gotox map entries edges before after ---------------------------------------------- 500 250000 1000 35.34s 0.07s 1000 250000 2000 82.00s 0.07s 2000 250000 4000 148.62s 0.07s 2000 125000 4000 74.54s 0.05s 2000 500000 4000 342.71s 0.19s Fixes: 493d9e0d6083 ("bpf, x86: add support for indirect jumps") Suggested-by: Eduard Zingerman <eddyz87@gmail.com> Signed-off-by: Daniel Borkmann <daniel@iogearbox.net> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260925175244.1136329-6-daniel@iogearbox.net
2026-09-25bpf: Leave out the hidden exit edge of a subprogram without an exitDaniel Borkmann
visit_abnormal_return_insn() gives tail_call and ld_{abs,ind} insns a second successor, the exit_idx of their subprogram, so that the hidden exit from the subprogram is part of the CFG. check_subprogs() only assigns exit_idx when it walks over a BPF_EXIT, and a subprogram whose last insn jumps back into itself never has one, so exit_idx stays zero and the edge points at insn 0 of the program: 0: r0 = 0 1: r0 = *(u8 *)skb[0] 2: goto -2 For a subprogram other than main, update_insn() then reads the liveness masks at a negative relative index. Such a subprogram can still load when it leaves through a bpf_throw(), so mark exit_idx as unset in that case and leave the hidden edge out, as there is no exit for it to reach. Fixes: e40f5a6bf88a ("bpf: correct stack liveness for tail calls") Fixes: ee861486e377 ("bpf: Fix ld_{abs,ind} failure path analysis in subprogs") Signed-off-by: Daniel Borkmann <daniel@iogearbox.net> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260925175244.1136329-5-daniel@iogearbox.net