| Age | Commit message (Collapse) | Author |
|
https://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next.git
# Conflicts:
# mm/internal.h
|
|
# Conflicts:
# fs/coredump.c
# fs/f2fs/f2fs.h
# fs/fuse/dax.c
# fs/xfs/libxfs/xfs_btree.c
|
|
The bpf_map->ops->map_mmap callback invoked by bpf_map_mmap() can be set to
one of ringbuf_map_mmap_kern(), ringbuf_map_mmap_user(), array_map_mmap()
or arena_map_mmap().
It is convention in mm to mark mappings whose pages the kernel manages
itself with VM_MIXEDMAP, so core mm knows not to treat them as ordinary
page cache or anonymous memory.
The map_mmap callbacks ringbuf_map_mmap_kern() and ringbuf_map_mmap_user()
use remap_vmalloc_range(), which ultimately invokes vm_insert_page() and so
marks the ranges VM_MIXEDMAP, and array_map_mmap() sets VM_MIXEDMAP
explicitly.
However, the exception to this is arena_map_mmap(), which doesn't set the
flag.
This patch corrects this and updates the comment to reflect it.
The pages are refcounted and vm_normal_page() finds them regardless of the
flag, and VM_DONTEXPAND remains set (marking the memory as VM_SPECIAL and
thus unmergeable). The one effect is that NUMA balancing now skips these
VMAs, as it already does for the other bpf map mappings, which is the
reason array_map_mmap() gives for setting the flag.
The intent of this patch is to be able to establish the invariant that only
PFN-mapped or mixed map ranges may clear the VM_MAYWRITE flag, as is done
in bpf_map_mmap().
Link: https://lore.kernel.org/20261003-b4-mmap-prepare-vma-flag-sanify-v4-13-a1f052500fd7@kernel.org
Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Jann Horn <jannh@google.com>
Cc: Pedro Falcato <pfalcato@suse.de>
Cc: David Hildenbrand <david@kernel.org>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Jonathan Corbet <corbet@lwn.net>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Dennis Dalessandro <dennis.dalessandro@cornelisnetworks.com>
Cc: Jason Gunthorpe <jgg@ziepe.ca>
Cc: Leon Romanovsky <leon@kernel.org>
Cc: Paul Moore <paul@paul-moore.com>
Cc: Stephen Smalley <stephen.smalley.work@gmail.com>
Cc: Jaroslav Kysela <perex@perex.cz>
Cc: Takashi Iwai <tiwai@suse.com>
Cc: Alexei Starovoitov <ast@kernel.org>
Cc: Daniel Borkmann <daniel@iogearbox.net>
Cc: Andrii Nakryiko <andrii@kernel.org>
Cc: Eduard Zingerman <eddyz87@gmail.com>
Cc: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Cc: Zi Yan <ziy@nvidia.com>
Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: Nico Pache <nico.pache@linux.dev>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: Dev Jain <dev.jain@arm.com>
Cc: Barry Song <baohua@kernel.org>
Cc: Lance Yang <lance.yang@linux.dev>
Cc: Usama Arif <usama.arif@linux.dev>
Cc: Kiryl Shutsemau <kas@kernel.org>
Cc: Doug Gilbert <dgilbert@interlog.com>
Cc: James Bottomley <james.bottomley@hansenpartnership.com>
Cc: Martin K. Petersen <mkp@kernel.org>
Cc: Simona Vetter <simona@ffwll.ch>
Cc: Helge Deller <deller@gmx.de>
Cc: Sebastian Reichel <sre@kernel.org>
Cc: John Hubbard <jhubbard@nvidia.com>
Cc: Peter Xu <peterx@redhat.com>
Cc: Masami Hiramatsu <mhiramat@kernel.org>
Cc: Oleg Nesterov <oleg@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Thomas Gleixner <tglx@kernel.org>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Borislav Petkov <bp@alien8.de>
Cc: Dave Hansen <dave.hansen@linux.intel.com>
Cc: Arnaldo Carvalho de Melo <acme@kernel.org>
Cc: Namhyung Kim <namhyung@kernel.org>
Cc: Mark Rutland <mark.rutland@arm.com>
Cc: Rik van Riel <riel@surriel.com>
Cc: Harry Yoo <harry@kernel.org>
Cc: Juri Lelli <juri.lelli@redhat.com>
Cc: Vincent Guittot <vincent.guittot@linaro.org>
Cc: Maarten Lankhorst <maarten.lankhorst@linux.intel.com>
Cc: Maxime Ripard <mripard@kernel.org>
Cc: Thomas Zimmermann <tzimmermann@suse.de>
Cc: David Airlie <airlied@gmail.com>
Cc: Will Deacon <will@kernel.org>
Cc: Aneesh Kumar K.V <aneesh.kumar@kernel.org>
Cc: Nicholas Piggin <npiggin@gmail.com>
Cc: Arnd Bergmann <arnd@arndb.de>
Cc: Muchun Song <muchun.song@linux.dev>
Cc: Oscar Salvador <osalvador@suse.de>
Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
Cc: Jan Kara <jack@suse.cz>
Cc: Marc Zyngier <maz@kernel.org>
Cc: Oliver Upton <oupton@kernel.org>
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Madhavan Srinivasan <maddy@linux.ibm.com>
Cc: Anup Patel <anup@brainfault.org>
Cc: Paul Walmsley <pjw@kernel.org>
Cc: Palmer Dabbelt <palmer@dabbelt.com>
Cc: Albert Ou <aou@eecs.berkeley.edu>
Cc: Christian Borntraeger <borntraeger@linux.ibm.com>
Cc: Janosch Frank <frankja@linux.ibm.com>
Cc: Claudio Imbrenda <imbrenda@linux.ibm.com>
Cc: Alexander Gordeev <agordeev@linux.ibm.com>
Cc: Gerald Schaefer <gerald.schaefer@linux.ibm.com>
Cc: Heiko Carstens <hca@linux.ibm.com>
Cc: Vasily Gorbik <gor@linux.ibm.com>
Cc: David S. Miller <davem@davemloft.net>
Cc: Andreas Larsson <andreas@gaisler.com>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Joshua Hahn <joshua.hahnjy@gmail.com>
Cc: Rakie Kim <rakie.kim@sk.com>
Cc: Byungchul Park <byungchul@sk.com>
Cc: Gregory Price <gourry@gourry.net>
Cc: Ying Huang <ying.huang@linux.alibaba.com>
Cc: Alistair Popple <apopple@nvidia.com>
Cc: Chris Li <chrisl@kernel.org>
Cc: Kairui Song <kasong@tencent.com>
Cc: Kemeng Shi <shikemeng@huaweicloud.com>
Cc: Nhat Pham <nphamcs@gmail.com>
Cc: Baoquan He <baoquan.he@linux.dev>
Cc: Youngjun Park <youngjun.park@lge.com>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: Qi Zheng <qi.zheng@linux.dev>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Axel Rasmussen <axelrasmussen@google.com>
Cc: Yuanchu Xie <yuanchu@google.com>
Cc: Wei Xu <weixugc@google.com>
Cc: Chengming Zhou <chengming.zhou@linux.dev>
Cc: Michal Hocko <mhocko@kernel.org>
Cc: Miklos Szeredi <miklos@szeredi.hu>
Cc: Xu Xin <xu.xin@linux.dev>
|
|
Patch series "mm: Unconditional per-VMA locks and cleanups", v7.
tl;dr: Make per-VMA locks available in all configs. Simplify some of the
per-VMA lock users now that they can rely on them being always available.
Binder and networking folks: Your code is the target of the cleanups. I'm
cc'ing you now on v2 because there's emerging consensus on the mm side
that the approach here is sane. I'm not quite sure how this pile would
get merged, but ack/review tags would be appreciated if this looks good to
you.
Longer version:
When working on some x86 shadow stack code, it was a real pain to avoid
causing recursive locking problems with mmap_lock. One way to avoid those
was to avoid mmap_lock and use per-VMA locks instead. They are great, but
they are not available in all configs which makes them unusable in generic
code, or if you want to completely avoid mmap_lock.
Make per-VMA locks available in all configs. Right now, they are only
available on select architectures when SMP and MMU are enabled. But all
of the primitives that per-VMA locks are built on (RCU, maple trees,
refcounts) work just fine without SMP or MMU.
The only real downside is that making VMAs a wee bit bigger on !MMU and
!SMP builds.
The upside is much cleaner code, lower complexity and less #ifdeffery.
Clean up a binder VMA locking site now that it can rely on per-VMA locks.
Building on top of universally-available per-VMA locks, introduce a new
helper. Since the new API does not require callers to have a fallback to
mmap_lock, it's much easier to use. Callers can potentially replace this
very common kernel idiom:
mmap_read_lock(mm);
vma = vma_lookup()
// fiddle with vma
mmap_read_unlock(mm);
with:
vma = vma_start_read_unlocked(mm, address);
// fiddle with vma
vma_end_read(vma);
Which avoids mmap_lock entirely in the fast path.
Use that new API for another binder site and one in the TCP code.
This patch (of 7):
The per-VMA locks have been around for several years. They've had some
bugs worked out of them and have seen quite wide use. However, they are
still only available when architectures explicitly enable them. Remove
the conditional compilation around the per-VMA locks, making them
available on all architectures and configs.
The approach up to now seemed to be to add ARCH_SUPPORTS_PER_VMA_LOCK when
the architecture started using per-VMA locks in the fault handler. But,
contrary to the naming, the Kconfig option does not really indicate
whether the architecture supports per-VMA locks or not. It is more of a
marker for whether the architecture is likely to benefit from per-VMA
locks.
To me, the most important thing side-effect of universal availability is
letting per-VMA locks be used in SMP=n configs. This lets us use
per-VMA locking in all x86 code without fallbacks.
Overall, this just generally makes the kernel simpler. Just look at the
diffstat. It also opens the door to users that want to use the per-VMA
locks in common code. Doing *that* brings additional simplifications.
The downside of this is adding some fields to vm_area_struct and
mm_struct. There are likely ways to optimize this, especially for things
like SMP=n configs. For now, do the simplest thing: use the same
implementation everywhere.
== Considerations for NOMMU config ==
NOMMU systems do not write-lock VMAs, therefore read-locking a VMA would
always succeed unless VMA is detached. Therefore for NOMMU config we make
vma_mark_attached() a NOOP, which keeps VMAs always in detached state.
This causes VMA read-locking to always fail and the caller falls back to
locking mmap_lock.
The following functions will have a different implementation in NOMMU
config:
- vma_mark_attached(), vma_mark_detached() are made NOOPs, keeping VMAs
always in a detached state and preventing assertions and refcount
underflows;
- vma_start_write(), vma_start_write_killable() are made NOOPs to avoid
warnings in __vma_start_write() due to VMAs being detached. These
functions are not used in NOMMU code but __vma_start_write() is an
exported function, therefore might be used by drivers.
- vma_assert_attached() is made NOOP because it's reachable from NOMMU
code via split_vma()->vma_iter_store_new()->vma_iter_store_overwrite();
- vma_assert_write_locked() is asserting vma->vm_mm is write-locked, as
was done before this change;
- vma_assert_locked() is asserting vma->vm_mm is locked, as was done
before this change;
The following functions work for both MMU and NOMMU configs:
- vma_lock_init() performs the same initialization as for MMU config;
- mm_lock_seqcount_init(), mm_lock_seqcount_begin(),
mm_lock_seqcount_end() are called from mmap_write_{lock|unlock} and
update mm_lock_seq correctly.
- mmap_lock_speculate_try_begin(), mmap_lock_speculate_retry() work as
is because mm_lock_seq is updated correctly;
- vma_start_read(), vma_start_read_locked() will always fail because
VMAs are always detached;
- vma_end_read() will never be called because vma_start_read() never
succeeds;
- vma_is_attached() always return false because VMAs are always
detached;
- vma_assert_detached() will never trigger because VMAs are never
attached;
- vma_start_read_locked() always return false because VMAs are always
detached;
- lock_vma_under_rcu() will be safe as the attempted read lock will bail;
Changes in the following files are not affecting NOMMU config:
task_mmu.c - not compiled when CONFIG_MMU=n;
pagewalk.c - not compiled when CONFIG_MMU=n;
userfaultfd.c - not compiled when CONFIG_MMU=n (CONFIG_USERFAULTFD depends
on CONFIG_MMU);
The following changes in the BPF code are made to keep NOMMU config
working like before:
stack_map_lock_vma() - keeps mmap_lock in NOMMU config;
bpf_iter_task_vma_new() - bails out in NOMMU config;
Link: https://lore.kernel.org/20260831203056.838265-1-surenb@google.com
Link: https://lore.kernel.org/20260831203056.838265-2-surenb@google.com
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Signed-off-by: Suren Baghdasaryan <surenb@google.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reviewed-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Todd Kjos <tkjos@android.com>
Cc: Christian Brauner <christian@brauner.io>
Cc: Carlos Llamas <cmllamas@google.com>
Cc: Alice Ryhl <aliceryhl@google.com>
Cc: David S. Miller <davem@davemloft.net>
Cc: David Ahern <dsahern@kernel.org>
Cc: Arve Hjønnevåg <arve@android.com>
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
Remove folio_queue as it's no longer used.
Signed-off-by: David Howells <dhowells@redhat.com>
Link: https://patch.msgid.link/20261001091239.3343034-11-dhowells@redhat.com
Reviewed-by: Paulo Alcantara <pc@manguebit.org>
cc: Matthew Wilcox <willy@infradead.org>
cc: Christoph Hellwig <hch@infradead.org>
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
A data slice is valid only while the specific dynptr that produced it
remains valid. prepare_func_exit() copies a subprogram return value into
the caller and then frees the callee state without applying the normal
destruction semantics to dynptrs in the callee stack.
A callee can therefore derive a slice from a local dynptr, return or spill
the slice into its caller, and then let the dynptr disappear with the
callee frame. The slice retains the id of a dynptr that is no longer in the
verifier state, so the verifier continues to accept accesses through it.
Before the object relationship refactor, bpf_dynptr_data() slices from
referenced dynptrs also carried the shared reference id, allowing a later
release to catch some of these escaped slices. The refactor made those
slices precise children of their source dynptr and exposed the missing
teardown as an accepted stale access even after the shared resource is
released.
Add destroy_dynptrs_in_stack_slots() to apply the existing dynptr teardown
to an inclusive range of stack slots. Reuse it for dynptr initialization,
variable-offset stack writes, and the entire callee stack before freeing
the frame. The teardown rejects losing the last dynptr for a referenced
resource and invalidates the dynptr and all its descendants, including
slices held in the caller. Propagate errors through prepare_func_exit().
Reuse the existing teardown API and diagnostic. Update the affected
callback test expectation in this change so the commit remains test-clean.
Fixes: 308c7a0ae885 ("bpf: Refactor object relationship tracking and fix dynptr UAF bug")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Suggested-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Xu Yunxiang <xyx2021@mail.ustc.edu.cn>
Reviewed-by: Amery Hung <ameryhung@gmail.com>
Acked-by: Ihor Solodrai <ihor.solodrai@linux.dev>
Link: https://lore.kernel.org/r/CAMB2axOmdHwSrHihSRskDywkCmEGQOX+6dZPN_A9u-HhxY3UDA@mail.gmail.com
Link: https://lore.kernel.org/r/CAMB2axMmrTb+UZ84UD48MwMtXbk1s9bWtreU3AOfjzU0fPT0BA@mail.gmail.com
Link: https://lore.kernel.org/bpf/20260927120422.1042107-2-xyx2021@mail.ustc.edu.cn
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
|
|
LLVM splits a static of a Rust program into pieces. Every piece is a VAR
with the type of the whole static:
[223] STRUCT 'BpfCell<core::option::Option<...>>' size=32 vlen=1
[236] VAR '..scx_cosmos9TASK_CTXS.0' type_id=223, linkage=static
[248] DATASEC '.bss' size=0 vlen=9
type_id=236 offset=24648 size=1 (VAR '..TASK_CTXS.0')
and the kernel rejects such BTF with "Invalid size".
Allow it. The size of a variable in DATASEC is used in two places:
- btf_find_datasec_var() looks for timers, spin locks, kptrs and other
special fields. btf_find_field_one() skips a variable when its size is
not the size of its type, so there are no special fields in a piece.
- btf_datasec_show() prints variables by their types. It would read
past the piece and past the end of the map value when the piece is the
last one. Don't print the pieces.
"global data test #8" and "#10" in prog_tests/btf.c expect "Invalid
size". BTF of both is loaded now. The map of #10 is not created, since
its value is larger than the section.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20261002124714.180012-10-alexei.starovoitov@gmail.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
|
|
Some static functions in BTF of a Rust program have arguments without
names:
[71] FUNC_PROTO '(anon)' ret_type_id=0 vlen=1
'(anon)' type_id=72
[81] FUNC 'unwrap_failed' type_id=71 linkage=static
and the kernel rejects such BTF with "Invalid arg#1". It's 5 of 25
static functions in scx_cosmos and 6 of 24 in scx_simple.
The verifier looks at types of the arguments. The names are printed only,
as "(anon)" when there is none. Allow such static FUNC. Global functions
are checked as before.
The test "func (Some arg has no name)" in prog_tests/btf.c has a static
function. Make it global to keep the check.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Alan Maguire <alan.maguire@oracle.com>
Link: https://lore.kernel.org/bpf/20261002124714.180012-8-alexei.starovoitov@gmail.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
|
|
Names of types and functions in BTF that LLVM makes for a Rust program
are not C identifiers:
[69] STRUCT 'NonNull<str>' size=16 vlen=1
[147] FUNC 'write_fmt<scx_cosmos::BpfStream>' type_id=146
and the kernel rejects such BTF with "Invalid name". scx_simple and
scx_cosmos schedulers written in Rust have them in STRUCT, FWD and FUNC,
made of letters, digits and " #&()*,:;<>[]{}", 430 characters at most.
Allow any printable character in btf_name_valid_identifier(), like it's
done for DATASEC. It checks names of types, functions, members,
enumerators, variables and arguments, so all of them can have such
characters now. The limit of KSYM_NAME_LEN stays.
The name of FUNC is a part of the name of the program in kallsyms, where
a space would break the parsers. Replace what is not a character of
an identifier with '_' there.
Tests in prog_tests/btf.c expect "Invalid name" for names with '!' and
'*', which are valid now. Put a character that is not printable there.
The type name '?foo' is expected to load.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20261002124714.180012-6-alexei.starovoitov@gmail.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
|
|
There are no address spaces in Rust. LLVM emits plain loads and stores
for arena memory, without cast_kern:
r3 = 0x20 ll // R_BPF_64_64 .bss, libbpf puts it into arena
r2 = *(u64 *)(r3 + 0) // R3 invalid mem access 'scalar'
lock *(u64 *)(r2 + 8) += r1
Add BPF_F_ARENA_SCALAR flag of BPF_PROG_LOAD. When the program is loaded
with it and has an arena treat ldx, stx, st and atomics through a number
as arena access. It's as safe as access through PTR_TO_ARENA: JIT adds
the base of arena to the low 32 bits of the address. The program has
an arena only with CAP_BPF and CAP_PERFMON.
Registers of the program are not changed, since Rust compares and
stores the address after the access. bpf_do_misc_fixups() copies
the low 32 bits into BPF_REG_AX and the access goes through it:
r2 = *(u64 *)(r3 + 0) -> w12 = w3
r2 = *(u64 *)(r12 + 0)
Constant blinding needs BPF_REG_AX for immediates and leaves insns that
use BPF_REG_AX alone. So the immediate of st through a number is not
blinded, the rest of the program is:
*(u64 *)(r3 + 0) = 1 -> w12 = w3
*(u64 *)(r12 + 0) = 1
The same insn may see PTR_TO_ARENA on another path. It works for both.
It's a flag, since a number is also what the verifier makes of a pointer
that went away when the program has CAP_PERFMON: a ringbuf record after
bpf_ringbuf_submit(), a pointer to the packet after bpf_skb_pull_data().
Programs in C that have an arena keep "invalid mem access 'scalar'" for
such bugs.
JIT tells with bpf_jit_supports_arena_scalar() that it takes BPF_REG_AX
as the address of arena access. With other JITs the flag is rejected
with -EOPNOTSUPP. x86 and arm64 do:
- x86: the fault handler finds the register that holds the address
through reg2pt_regs[]. Add BPF_REG_AX, r10 of x86, there.
- arm64: nothing else is needed. The JIT takes any register as the
address and the fault handler reads it by its number. Compile tested
only.
Any number is an address of arena, NULL and small numbers included, like
it is for PTR_TO_ARENA: arena code in C dereferences NULL and relies on
the fault being handled.
Still rejected:
- insn that sees a number on one path and a pointer that is not
PTR_TO_ARENA on another.
- numbers passed to helpers and kfuncs, except __arena arguments.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20261002124714.180012-4-alexei.starovoitov@gmail.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
|
|
Rust's core::fmt keeps a flag in the low bit of a pointer:
r2 = *(u64 *)(r1 + 0)
r3 = r2
r3 &= 1
r2 >>= 1
The verifier rejects it:
r0 &= 8
R0 bitwise operator &= on pointer prohibited
Negation and byte swap of a pointer already produce a number when
allow_ptr_leaks is set. Do the same for bitwise ops, shifts, *=, /= and
%=, 64-bit and 32-bit. The result is an unknown number. With CAP_PERFMON
the program can store the pointer and load it back as a number already.
+= and -= are not changed: they keep the pointer, 32-bit += is rejected.
Pointers that allow no arithmetic and pointers that may be NULL are
still rejected.
Two tests in verifier_value_illegal_alu store through the result of
&= and /=. They are rejected at the store now. Update expected messages.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20261002124714.180012-2-alexei.starovoitov@gmail.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
|
|
A freplace attach claims the target prog by bumping
tgt_prog->aux->freplace_link_cnt and setting tr->extension_prog.
bpf_arch_text_poke() then makes the extension take effect. If the
poke fails, the claims are never released: the attach unwinds
through bpf_link_cleanup(), which clears link->prog, so
bpf_trampoline_unlink_prog() never runs.
Drop the link count under ext_mutex on the error path, and set
tr->extension_prog only after the poke succeeded. The count is
still bumped before the poke: it blocks prog_array updates while
the entry is patched.
Fixes: d6083f040d5d ("bpf: Prevent tailcall infinite loop caused by freplace")
Fixes: be8704ff07d2 ("bpf: Introduce dynamic program extensions")
Suggested-by: Leon Hwang <leon.hwang@linux.dev>
Signed-off-by: Yuan Chen <chenyuan@kylinos.cn>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Leon Hwang <leon.hwang@linux.dev>
Link: https://patch.msgid.link/20260928091121.219333-1-chenyuan_fl@163.com
|
|
Cross-merge BPF and other fixes after downstream PR.
Conflicts:
arch/arm64/net/bpf_jit_comp.c
arch/x86/net/bpf_jit_comp.c
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
|
|
specialize_kfunc() currently updates the canonical kfunc descriptor in
place. Different call sites therefore cannot select different
specializations of the same kfunc, and the selected target depends on
verification order.
Keep canonical descriptors unchanged for verifier lookups. Specialize a
copy for each call site, then reuse or append an immutable target
descriptor and store its index in the finalized call instruction.
Allow space for one canonical and one specialized target per kfunc. This
is conservative because most kfuncs do not specialize. Adding specialized
kfunc versions does not require resorting the array because lookups are
only used for deduplication and func_model lookup. Deduplication for
specialized kfuncs is handled by the specialization process itself,
while all specializations of a function share the same proto and
func_model.
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20261002105218.6171-7-emil@etsalapatis.com
|
|
Currently, the verifier sort the kfunc desc table twice: Once during
kfunc collection by function ID, useful during verification, and
once by immediate address, useful during JIT bytecode lowering
time. The latter sort can be removed by storing in the instruction
the kfunc desc array index the desc is in and using that instead.
Modify the JITs to follow the same convention.
This change simplifies the subsequent patch that adds per-call site
function specialization.
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20261002105218.6171-6-emil@etsalapatis.com
|
|
The bpf_arena_alloc_pages() function currently only allocates pages
inside a spinlock critical section with IRQs off. This forces the use of
alloc_pages_nolock() in the BPF allocator, even when the caller is a
sleepable BPF function. This in turn causes allocation failures even in
cases where falling into the allocator slow path and possibly sleeping
would eventually succeed. This can be triggered consistently by heavy
BPF arena users like scx.
Allocate the arena pages before taking the critical section and pass
whether the caller can sleep to bpf_alloc_pages(). This lets sleepable
callers use the blocking allocator while non-sleepable callers retain
the no-lock allocation behavior.
Fixes: b8467290edab ("bpf: arena: make arena kfuncs any context safe")
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20261002105218.6171-4-emil@etsalapatis.com
|
|
bpf_alloc_pages() currently decides whether it may use the
blocking page allocator from the current execution context
alone. Let callers further restrict that choice by passing
whether their context is sleepable.
Use the blocking allocator only when both the caller and
runtime context allow sleeping. Add __GFP_RETRY_MAYFAIL so
this path can reclaim without invoking the OOM killer when
the allocation is charged to another memcg. Existing
non-sleepable callers retain the no-lock allocation behavior.
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20261002105218.6171-3-emil@etsalapatis.com
|
|
bpf_map_alloc_pages() does not use its map argument. Storing allocated
pages in an array also forces arena callers to allocate a separate
pointer array.
Expose the single-page allocator as bpf_alloc_page(), rename the bulk
helper to bpf_alloc_pages(), and return bulk allocations through an
llist using page->pcp_llist. Add bpf_free_pages() to safely release all
pages remaining on such a list.
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20261002105218.6171-2-emil@etsalapatis.com
|
|
bpf_mem_cache_free_rcu() uses this_cpu_ptr() which requires migration
to be disabled. All callers of rhtab_delete_elem() disable migration
except __rhtab_map_lookup_and_delete_batch(), which calls it under
rcu_read_lock() only.
On CONFIG_PREEMPT_RCU, rcu_read_lock() does not disable preemption or
migration, so the task can migrate between CPUs during the delete loop,
causing this_cpu_ptr() to trigger:
BUG: using smp_processor_id() in preemptible [00000000] code
Fix by wrapping the delete loop in migrate_disable()/migrate_enable()
in __rhtab_map_lookup_and_delete_batch(), matching the migration
protection that the other callers already provide.
Fixes: 818e00848227 ("bpf: Implement iteration ops for resizable hashtab")
Reported-by: syzbot+fd7e415d891073b83e1f@syzkaller.appspotmail.com
Signed-off-by: Ömer Mete Kaya <omermetekaya0@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Mykyta Yatsenko <yatsenko@meta.com>
Link: https://patch.msgid.link/20260929081609.557899-1-omermetekaya0@gmail.com
Closes: https://syzkaller.appspot.com/bug?extid=fd7e415d891073b83e1f
|
|
do_call_rcu_ttrace() returns early when call_rcu_ttrace_in_progress is set
and leaves the objects in free_by_rcu_ttrace. __free_rcu() frees
waiting_for_gp_ttrace only and clears the flag. Hence the objects that
free_bulk() or __free_by_rcu() added while RCU tasks trace GP was in flight
stay in free_by_rcu_ttrace until free_bulk() or alloc_bulk() is called for
the same bpf_mem_cache again, which may never happen. The number of such
objects is not bounded.
Turn call_rcu_ttrace_in_progress into three states:
0 - idle
1 - __free_rcu() is queued
2 - __free_rcu() is queued and free_by_rcu_ttrace got more objects since
do_call_rcu_ttrace() sets 2. __free_rcu() does cmpxchg(1 -> 0) and starts
the next GP when it fails. It cannot clear the flag first and check
free_by_rcu_ttrace later, since bpf_mem_alloc_destroy() frees bpf_mem_cache
without waiting for RCU callbacks when the flag is zero.
Now __free_rcu() queues itself, so the one that didn't see 'draining' may
do call_rcu_tasks_trace() after rcu_barrier_tasks_trace() in
free_mem_alloc(). Queue it under rcu_read_lock() and do synchronize_rcu()
before the barriers. Calling rcu_barrier_tasks_trace() twice works too, but
creating and destroying hash maps in a loop on many cpus slows down to one
free_mem_alloc() per GP and kworkers pile up.
Fixes: 8d5a8011b35d ("bpf: Batch call_rcu callbacks instead of SLAB_TYPESAFE_BY_RCU.")
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20260930095920.601738-3-alexei.starovoitov@gmail.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
|
|
Move the part of do_call_rcu_ttrace() that runs after
call_rcu_ttrace_in_progress is set into __do_call_rcu_ttrace().
The next patch will call it from __free_rcu().
No functional change.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20260930095920.601738-2-alexei.starovoitov@gmail.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
|
|
Since commit 022ac0750883 ("bpf: use reg->var_off instead of reg->off
for pointers"), find_good_pkt_pointers() sets the range of all packet
pointers sharing an id from the umax of the compared pointer, and
check_packet_access() requires umax + off + size <= range. That assumes
the umax of two such pointers differ by exactly their constant distance.
reg_bounds_sync() breaks it when var_off tightens one umax and not the
other:
r4 &= 0x38
if r4 > 50 goto exit ; umax 50, var_off (0x0; 0x38)
r5 = pkt + r4 ; umax 50
r6 = r5
r6 += 8 ; umax 56, not 58
Comparing r6 with pkt_end sets the range to 56, and the valid 8-byte
load at r5 is rejected (50 + 8 > 56). Comparing r5 sets it to 50, and
the out-of-bounds 1-byte load at r6 - 7, i.e. r5 + 1, is accepted
(56 - 7 + 1 <= 50).
Don't call reg_bounds_sync() on a packet pointer that keeps its id (a
constant was added or subtracted) or its range (an unknown non-negative
value was subtracted), so that var_off cannot tighten its umax. Only
update the 32-bit bounds from var_off: reg_bounds_sanity_check() wants
them constant when the lower half of var_off is, e.g. for pkt + 8.
This relies on nothing else changing the 64-bit bounds of a packet
pointer, which holds today.
var_off of such a pointer is no longer narrowed by its bounds. Adjust
three verifier_align expectations; the low bits, which the alignment
checks use, don't change. veristat on the selftests shows no verdict
changes and +0.8% insns in test_cls_redirect_subprogs.
Fixes: 022ac0750883 ("bpf: use reg->var_off instead of reg->off for pointers")
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20261001145255.855630-1-alexei.starovoitov@gmail.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
|
|
Two things in __cgroup_bpf_query() work only because CGROUP_TCP_SOCK_OPS is
the only struct_ops attach type.
It calls cgroup_bpf_enabled(atype) with an atype that
find_atype_by_struct_ops_id() works out at runtime. That macro is an asm
goto and needs a constant. Today the compiler can see there is only one
value; add a second type and the build breaks with "impossible constraint
in 'asm'". The check is not needed anyway. When the key is off nothing is
attached, so progs[atype] and effective[atype] are empty and the count loop
leaves total_cnt at 0. The other attach types already use that loop with
no such check. Drop the check and the skip_count label.
And find_atype_by_struct_ops_id() matches on type_id alone. An attach type
whose subsystem is not built keeps type_id 0, so a query for type 0 finds
it and returns success with nothing instead of -ENOENT. Skip such slots.
Fixes: 369d9dcd8fb8 ("bpf: Add infrastructure to support attaching struct_ops to cgroups")
Signed-off-by: Shakeel Butt <shakeel.butt@linux.dev>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Reviewed-by: Amery Hung <ameryhung@gmail.com>
Acked-by: Yafang Shao <laoar.shao@gmail.com>
Acked-by: Tejun Heo <tj@kernel.org>
Link: https://patch.msgid.link/20261001141901.3225830-1-shakeel.butt@linux.dev
|
|
Program Structure reports have two attribution gaps. A missing jump
table is reported at the beginning of its subprogram rather than at
the gotox that needs the table, and the subprogram-layout checks run
before func_info and line_info are validated and installed, so their
reports cannot name the source function or line.
Report the missing jump table at the gotox instruction that looks it up,
rather than at the start of its subprogram.
func_info and line_info validation only needs the subprogram boundaries
found by add_subprogs() and the LD_ABS and tail-call properties that
check_subprogs() collects while scanning instructions; it does not
depend on the layout checks themselves. Split the property collection
into its own nested subprogram and instruction scan, together with the
program's callx marker, and run bpf_check_btf_info() before
check_subprogs(). The jump-boundary and fallthrough reports then carry
validated source information without changing either check.
CO-RE relocations are unaffected: commit c26e97721b17 ("bpf: Apply
CO-RE relocations before subprogram validation") applies them before
add_subprogs(), and they stay there. Only the func_info and line_info
validation moves.
Fixes: a8f427835394 ("bpf: Report Program Structure CFG errors")
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://lore.kernel.org/bpf/cf2f420c2b21de440a7dc51b1565c0f06d4b539640ee5c03384e5d77bcfb5686@mail.kernel.org/
Link: https://lore.kernel.org/bpf/20260926133048.2962553-3-memxor@gmail.com
|
|
Ensure that vlen and flags combinations are consistent for
BTF_KIND_LOC_PARAMs.
Fixes: 33c5a3278bdb ("btf: Extend UAPI to support BTF location (inline site) info")
Suggested-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Alan Maguire <alan.maguire@oracle.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260926175113.2368566-2-alan.maguire@oracle.com
|
|
All supported BPF-subprogram argument types now pass through
check_func_arg() from one guarded branch in the legacy validation loop.
Replace that loop and its outgoing-stack validation with
check_func_args().
The scratch prototype is indexed by BTF parameter and the call metadata
carries the BTF function prototype. The existing BTF argument iteration
therefore maps parameters to ABI slots for both kfunc and BPF subprogram
calls, including additional slots occupied by by-value aggregates.
The common checker can return non-EFAULT errors other than -EINVAL, as
the existing BTF-ID path already did. This is compatible with subprog
call handling: only -EFAULT is fatal. Any other error marks BTF
unreliable. Static subprogs can then fall back to inline verification,
while global subprogs reject the call.
Remove the caller-register plumbing that the dedicated loop required.
Signed-off-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260928185334.1004200-12-ameryhung@gmail.com
|
|
Route fixed-size global ARG_PTR_TO_MEM arguments through
check_func_arg(). Preserve their read-write access check, nullable
contract, and support for BTF-defined allocated memory.
Recognize subprog calls explicitly in the common checker. Global
subprog stack liveness can prove that bytes in an argument are unused
by the callee. Retain the existing allowance for those poisoned stack
bytes when call metadata is present, and update the nullability log
expectation.
The verifier also rejects packet pointers when the callee may change
packet data. The concrete packet-backed register state is only available
while checking the call site; the independently verified callee sees
generic PTR_TO_MEM. Record this property as subprog_may_change_pkt in
subprog-only call metadata and retain the packet-pointer rejection in the
common fixed-memory path.
Signed-off-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260928185334.1004200-11-ameryhung@gmail.com
|
|
Route global ARG_PTR_TO_BTF_ID arguments through check_func_arg(). Use
vmlinux BTF for their type match and retain the global-subprogram rules
instead of applying kfunc-only trusted-pointer validation or runtime type
resolution.
The common nullability check now rejects non-nullable global BTF-ID
arguments before register-type matching. Update the corresponding
verifier log expectations.
Signed-off-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260928185334.1004200-10-ameryhung@gmail.com
|
|
Subprog and helper/kfunc dynptr arguments ultimately use the same
process_dynptr_func() validation. Route the subprog dynptr branch
through check_func_arg() so register type, offset, and dynptr state
are checked in the common order.
check_func_arg() validates the register against the PTR_TO_STACK and
CONST_PTR_TO_DYNPTR compatibility set before dispatching to
process_dynptr_func(). Once the subprog path uses the common checker,
process_dynptr_func() has no caller that bypasses this validation. Remove
its now-redundant register-type check.
The existing wrong-register-type test passed a NULL local to a callee
that did not use the argument. Clang left the context pointer in R1,
while GCC materialized zero. The common checker rejected the values at
different stages and emitted different diagnostics.
Obtain a PTR_TO_BTF_ID from bpf_get_current_task_btf() and keep its
callee argument live. This makes both compilers exercise the intended
register-type mismatch.
Signed-off-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260928185334.1004200-9-ameryhung@gmail.com
|
|
ARG_PTR_TO_ARENA accepts both arena pointers and scalars. The existing
BPF-subprogram contract includes a constant-zero scalar; it is an arena
address value rather than a conventional pointer rejected by
nullability checking.
Kfunc prototype generation already records this zero acceptance with
PTR_MAYBE_NULL for both arena suffixes; separate JIT metadata determines
whether the kernel callee receives the arena base or NULL. Apply the same
call-site representation to the temporary BPF-subprogram prototype, then
extend the guarded check_func_arg() path to cover arena arguments.
Signed-off-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260928185334.1004200-8-ameryhung@gmail.com
|
|
Subprog and helper/kfunc context arguments require PTR_TO_CTX with an
acceptable offset. Route subprog context arguments through
check_func_arg() and update the verifier-log expectation.
Signed-off-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260928185334.1004200-7-ameryhung@gmail.com
|
|
Global subprogram arguments tagged __arg_untrusted deliberately
accept any caller value. The callee is verified separately with a
read-only untrusted pointer, whose accesses use protected probe-read
instructions.
Keep the untrusted type in the canonical bpf_subprog_info used to
prepare the callee state. Represent it as ARG_IGNORE only in the
temporary bpf_func_proto used for call-site checking, then extend the
guarded check_func_arg() path to cover it.
Signed-off-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260928185334.1004200-6-ameryhung@gmail.com
|
|
BPF subprogram scalar arguments are checked separately even though the
common argument checker enforces the same SCALAR_VALUE requirement.
Start the migration by routing the scalar branch through check_func_arg()
and update the verifier-log expectation.
Use the BTF parameter index for metadata and the ABI slot index for
register lookup. Route extra slots of by-value aggregates through
check_arg_extra_slot() so every occupied slot follows the common path.
Following patches can extend this guarded common-check branch as each
remaining BPF-subprogram-specific implementation is removed.
Signed-off-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260928185334.1004200-5-ameryhung@gmail.com
|
|
The common argument checker consumes a parameter-indexed bpf_func_proto,
while BPF subprogram argument metadata is cached by ABI slot in compact
bpf_subprog_info entries. Embedding a full prototype in every fixed
subprogram entry would waste memory.
Embed a scratch prototype in the verifier environment. Fill it by
walking BTF parameters and the cached slots with separate cursors,
mirroring the kfunc representation for parameters that occupy multiple
slots. Keep the compact cache as the source of truth. The verifier
environment is already heap allocated, so this avoids a separate
allocation, failure path, and cleanup.
Keep the existing validation loop, but make it consume the generated
prototype in preparation for moving BPF subprogram calls to the common
argument checker. No functional change.
Signed-off-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260928185334.1004200-4-ameryhung@gmail.com
|
|
Prepare subprog calls to use the common check_func_args() path. That
checker distinguishes call kinds through bpf_call_arg_meta, but
btf_check_func_arg_match() leaves both BTF and the function ID zero even
though it validates a program-BTF signature.
Represent a subprog call with a non-NULL BTF and a zero function ID.
Define helpers as NULL BTF plus nonzero ID, and kfuncs as non-NULL BTF
plus nonzero ID. Requiring a nonzero kfunc ID also prevents unavailable
special-kfunc IDs from matching subprog metadata. The corresponding
subprog predicate is introduced later alongside its first use.
Setting BTF would expose checks that treat every BTF-backed call as a
kfunc. Restrict kfunc-only BTF parameter lookup, register admission,
release diagnostics, no-cast alias, and projection handling to actual
kfuncs.
The BTF is not modified here, but bpf_call_arg_meta::btf and several
downstream consumers use non-const pointers. Match those existing types
instead of broadening this series with a const-correctness cleanup.
No functional change is intended.
Signed-off-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260928185334.1004200-3-ameryhung@gmail.com
|
|
Several kfunc argument checks index BTF parameters using the ABI slot
number. Parameter and slot indexes diverge after a by-value argument
wider than one eightbyte, so a later argument can select the wrong BTF
parameter or index past the end of the function prototype.
This affects the BTF lookups in check_func_arg_nullability() and
check_func_arg_release(), and the parameter passed to
btf_check_iter_arg() by process_iter_arg().
check_func_arg() already tracks the BTF parameter and ABI slot
separately. Pass the parameter index to all three consumers while
retaining argno for register and stack-slot diagnostics.
Signed-off-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260928185334.1004200-2-ameryhung@gmail.com
|
|
bpf_tramp_image_put() makes sure a trampoline image is not freed while
a task may still be running in it, but nothing similar is done for the
progs called by that image. Detach drops the last prog reference right
away and the prog is freed after grace periods, on the basis that a
task still in the traced function skips the fexit progs once the nop at
ip_after_call is patched to a jump.
That leaves out a task sleeping in a sleepable prog that runs before
the detached one, which no grace period waits for:
CPU 0 CPU 1
in image I, sleeping in prog S
detach P from I's trampoline
-> new image, bpf_tramp_image_put(I)
bpf_prog_put(P), last ref
grace periods, P freed
back from S
__bpf_prog_enter(P)
call P->bpf_func
If S and P are fexit progs the task is already past the patched jump,
and fentry only images don't have one. On x86 this is an int3 in
poisoned bpf_prog_pack memory:
Oops: int3: 0000 [#1] SMP NOPTI
CPU: 18 UID: 0 PID: 94573 Comm: x169 Not tainted 6.18.44 #1 PREEMPT(lazy)
RIP: 0010:0xffffffffc0601d8d
Call Trace:
<TASK>
? bpf_trampoline_6442515411+0x1a4/0x21b
bpf_lsm_bprm_committed_creds+0x5/0x10
security_bprm_committed_creds+0x5f/0x70
begin_new_exec+0x2d6/0x410
...
We hit this in production when progs attached through trampolines got
detached while their hooks were busy, and it was independently found
with a fuzzer and KASAN.
Have the JITs emit a patchable nop in front of each prog call sequence
and record it in the image. When a prog is detached, patch its nop to a
jump over the call sequence. Progs that stay attached keep running for
the tasks that are in the image, and ip_after_call isn't needed anymore.
The task can be in any image that isn't freed yet, not only in the
current one. It sleeps in image I1 that calls S, P and Q, then P is
detached and the trampoline moves to image I2, then Q is detached and
its call is still in I1. So the trampoline keeps a list of its images
until they are freed, and detaching a prog patches its nop in all of
them. Images hold a reference on the trampoline for that long.
On riscv and loongarch a jump of any range takes several instructions,
and a task preempted in the middle of them could resume into half of the
new sequence. The nop is a single instruction there, patched to a near
branch through arch_bpf_trampoline_skip().
With the extra nops, BPF_MAX_TRAMP_LINKS progs no longer fit in a page
on x86 and arm64 (and already didn't on powerpc), so lower the limit
there like s390 does.
Fixes: e21aa341785c ("bpf: Fix fexit trampoline.")
Reported-by: Sechang Lim <rhkrqnwk98@gmail.com>
Suggested-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Florent Revest (Anthropic) <florent.revest@linux.dev>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260926135605.1217928-3-florent.revest@linux.dev
Closes: https://lore.kernel.org/bpf/20260815071927.147049-1-zirajs7@gmail.com/
|
|
When a prog is detached from a trampoline, it is freed after an RCU
grace period, or an RCU tasks trace one if it is sleepable. This covers
the tasks that are running the prog, since the prog's enter helper takes
the matching read lock before the prog is called. On a preemptible
kernel, it doesn't cover a task that was preempted in the trampoline
just before the enter helper. That task holds no lock yet, and it calls
the prog after it was freed:
BUG: KASAN: vmalloc-out-of-bounds in __bpf_prog_enter_recur+0x3a5/0x3f0
Read of size 8 at addr ffffc90000055040 by task candidate/110
CPU: 1 UID: 0 PID: 110 Comm: candidate Not tainted 7.3.0-rc2-00014-g15071f2a1263-dirty #2 PREEMPT(full)
Call Trace:
<TASK>
__bpf_prog_enter_recur+0x3a5/0x3f0
bpf_trampoline_6442509193+0x37/0xf1
__x64_sys_futex+0x9/0x410
do_syscall_64+0xb0/0x530
...
Wait for an RCU tasks grace period before the existing one when freeing
a prog that was linked to a trampoline. An RCU tasks grace period only
ends once the tasks that were preempted have run again, and
bpf_tramp_image_put() already relies on it to free the image. It doesn't
wait for tasks that sleep in the trampoline, the next commit takes care
of those.
Only progs that were linked to a trampoline can be called this way, so
bpf_trampoline_add_prog() marks them and other progs are still freed as
before.
Fixes: e21aa341785c ("bpf: Fix fexit trampoline.")
Reported-by: Junseo Lim <zirajs7@gmail.com>
Signed-off-by: Florent Revest (Anthropic) <florent.revest@linux.dev>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260926135605.1217928-2-florent.revest@linux.dev
Closes: https://lore.kernel.org/bpf/aqdrwVpanH3WGurX@omen-arch/
|
|
The existing BPF_PROG_STREAM_READ_BY_FD command only supports polling a
program stream through repeated bpf() calls. It cannot block for new data
or integrate with poll-based event loops.
Add BPF_PROG_STREAM_OPEN to return a read-only, close-on-exec file
descriptor for a selected program stream. Reads block by default and
BPF_F_STREAM_NONBLOCK, the only accepted flag, provides non-blocking
behavior. poll reports readable data and reports hangup once the program
has been freed. Like pipes and sockets, the descriptor is not seekable and
lseek fails with ESPIPE.
A stream descriptor deliberately does not retain the program. Move each
stream into a separately refcounted allocation so program teardown can mark
it dead and wake descriptor users while outstanding descriptors drain
buffered data safely. Readers sample the dead flag before looking for data,
so EOF is reported only when the stream was already dead before it was
found empty; data published right before teardown is never skipped.
Only programs loaded through BPF_PROG_LOAD get streams. Classic BPF
filters, JIT subprograms and shim programs never write to one, and
kernel-side writers already resolve a subprogram to its main program, so
those programs no longer carry stream state.
Readiness needs its own counter. Stream capacity is charged before
allocation and before an element is published to the stream log, so using
that reservation as the read and poll condition can report readable data
while no element exists: a blocking reader retries instead of sleeping and
a lone non-blocking reader can see POLLIN followed by EAGAIN. Publish bytes
with release ordering after adding elements to the lockless log, use
acquire loads before consuming them or reporting readiness, limit each read
to its readable snapshot and subtract only bytes actually copied. This
keeps the aggregate count correct even when concurrent publishers update it
out of publication order. With several readers on one stream, readiness
remains advisory, as it is for pipes. The capacity counter is kept solely
for enforcing the stream size limit.
Wakeups are always deferred through irq_work. Stream writers run in
whatever context the program runs in: NMI context for perf_event programs,
sections with interrupts disabled inside bpf_spin_lock or rqspinlock
critical sections since bpf_stream_vprintk() is KF_SPINLOCK_SAFE, and
tracing programs attached anywhere in the kernel, including inside the wait
queue and epoll code itself. Waking waiters directly from there can
deadlock, and no cheap context check covers every case: on PREEMPT_RT,
spinlock_t sections do not disable interrupts, so in_nmi() or
irqs_disabled() cannot tell such a program apart from a benign one. Queue
an irq_work item instead, as bpf_ringbuf does.
Queue it only when a publication turns an empty stream readable. Readers
block and pollers wait only after finding the stream empty, and the
readable count never drops below zero because each read is bounded by its
snapshot, so the first publication after such an observation is the one
that makes the count positive, and it is the one that queues the wakeup.
Publications into a stream that already holds data raise no interrupt, so
a program that prints while nobody drains its stream pays for a single
irq_work until the stream is emptied again. This matches bpf_ringbuf, which
notifies only once the consumer has caught up. Blocking readers and
level-triggered pollers re-check the readable count before waiting, so they
cannot miss data, and edge-triggered epoll consumers drain until EAGAIN
before waiting again, as epoll(7) requires.
Synchronize pending work before releasing the final stream reference so
the callback cannot outlive the stream, but only when the work was ever
queued: irq_work_sync() waits for an RCU grace period on PREEMPT_RT and on
architectures without an irq_work interrupt.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260925045536.1480933-3-memxor@gmail.com
|
|
A stream write whose formatted output is empty still allocates a stream
element and links it into the stream log. Capacity accounting only counts
payload bytes, so such elements never count against the stream limit, and a
program that keeps producing empty output allocates elements without bound
until a reader drains them.
Return success without allocating anything when there is nothing to write,
for both bpf_stream_vprintk() and staged writes. Every element in a stream
log then carries at least one byte, which later changes rely on when they
derive read readiness from published byte counts.
Fixes: 5ab154f1463a ("bpf: Introduce BPF standard streams")
Suggested-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260925045536.1480933-2-memxor@gmail.com
|
|
A few patches do not expand in place, they prepend: bpf_convert_ctx_accesses()
puts the ctx save for gen_epilogue and the instructions of gen_prologue in
front of insn 0, and bpf_do_misc_fixups() prepends the may_goto counter init
at the start of every subprog that uses may_goto.
They copy the original instruction of 'tgt_idx' into the last slot of the
patch buffer, so it now lives at tgt_idx + delta, and call adjust_jmp_off()
to move direct branches from tgt_idx to tgt_idx + delta.
Nothing does the same for indirect branches, so a BPF_MAP_TYPE_INSN_ARRAY
slot that named tgt_idx keeps naming tgt_idx, which is now the first prepended
instruction, and insn_aux_data[tgt_idx].indirect_target makes the JIT emit
the landing pad there. Add a __bpf_patch_insn_data() variant and pass
BPF_PATCH_MOVE_TARGET at the affected call-sites to fix up the delta. The
poke descriptors of direct tail calls are shifted the same way, and the
original instruction is not searched for by content in that mode, as it
is the last slot by construction.
Fixes: 493d9e0d6083 ("bpf, x86: add support for indirect jumps")
Fixes: 07ae6c130b46 ("bpf: Add helper to detect indirect jump targets")
Reported-by: Nicholas Carlini <npc@anthropic.com>
Suggested-by: Nicholas Carlini <npc@anthropic.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Anton Protopopov <a.s.protopopov@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260925175244.1136329-11-daniel@iogearbox.net
|
|
A program that fails after the JIT ran, for example in bpf_check_tail_call(),
has already stored its jitted offsets and target pointers into the insn array
maps it uses through bpf_prog_update_insn_ptrs() when release_insn_arrays()
hands them back by clearing the used flag with a plain atomic_set(). Nothing
orders those stores before the flag, so on a weakly ordered architecture they
can still be in flight when the next program takes the map: it is acquired
with atomic_xchg() in bpf_insn_array_init() and reset there, and the late
stores then land on top of whatever the new owner has written since, be it
that reset or the pointers its own JIT stored, leaving entries that point
into an image which is about to be freed.
Clear the flag with atomic_set_release() instead, so that everything the
releasing program stored is visible before the map can be acquired again.
This pairs with the fully ordered atomic_xchg() on the acquire side.
Fixes: b4ce5923e780 ("bpf, x86: add new map type: instructions array")
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260925175244.1136329-10-daniel@iogearbox.net
|
|
An insn array records, per entry, the xlated offset, the jitted offset and
a jitted target pointer in ips[]. bpf_insn_array_init() runs when the map
is bound to a program and resets only xlated_off; the jitted offset and the
ips[] pointer are left as the previous owner set them. A program can be
verified and JITed and only then fail to load (e.g. bpf_check_tail_call()).
By that point bpf_prog_update_insn_ptrs() has already filled ips[] with
pointers into that program's JIT image. Thus, the map is handed back for
reuse with the stale pointers intact. If the next program reuses the map,
the stale pointer survives in the now active program's jump table,
pointing into the previous, freed image. Fix by resetting the jitted offset
and the ips[] pointer alongside the xlated offset.
Fixes: b4ce5923e780 ("bpf, x86: add new map type: instructions array")
Reported-by: Sandipan Roy <saroy@redhat.com>
Reported-by: James Burton <jamesburton@meta.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Anton Protopopov <a.s.protopopov@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260925175244.1136329-9-daniel@iogearbox.net
|
|
bpf_opt_remove_nops() drops every ja +0 and may_goto +0 unconditionally.
A gotox target may itself be such a transparent nop: it is a verified
entry of an insn array jump table that is reachable at run time through
the indirect jump. bpf_insn_array_adjust_after_remove() marks it as
INSN_DELETED instead, and bpf_insn_array_ready() skips INSN_DELETED
entries, so the program loads with a hole in its jump table. The x86
JIT, for example, lowers gotox to a raw indirect jump, so executing
that index jumps to the zeroed ips[] slot, that is, to address 0.
Fix by keeping an indirect jump target in place.
Fixes: 493d9e0d6083 ("bpf, x86: add support for indirect jumps")
Reported-by: Sandipan Roy <saroy@redhat.com>
Reported-by: James Burton <jamesburton@meta.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Anton Protopopov <a.s.protopopov@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260925175244.1136329-8-daniel@iogearbox.net
|
|
The jump table of a subprog is collected in compute_subprog_jts() from the
insn_array maps of the program, and a map is attributed to the subprog that
contains its first entry. check_indirect_jump() instead resolves the targets
from the map the gotox register actually points to, bounded only by the
index range of that register, and never relates them back to the subprog
of the gotox.
The two disagree, so bpf_insn_successors() reports a subset of the edges the
BPF program can take and a gotox can enter a subprog the CFG never walked.
Close both ends: reject a map that reaches past the subprog holding its
first entry right when it is collected in compute_subprog_jts(), so that
every map of a program with a gotox is confined to one subprog, and
confine the resolved targets of a gotox to its subprog in
check_indirect_jump(). A target inside the subprog is then always in the
jump table the CFG walked, which is the invariant that actually has to
hold. The out-of-range check in visit_gotox_insn() cannot fire any more
once every table is confined and becomes a verifier_bug_if().
Fixes: 493d9e0d6083 ("bpf, x86: add support for indirect jumps")
Reported-by: James Burton <jamesburton@meta.com>
Reported-by: Nuoqi Gui <gnq25@mails.tsinghua.edu.cn>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Anton Protopopov <a.s.protopopov@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260925175244.1136329-7-daniel@iogearbox.net
|
|
create_jt() builds the jump table of the subprogram containing a gotox by
copying out and sorting every insn_array map of the program, and it does
so once per gotox instruction. The cost is therefore the number of gotox
instructions times the number of entries in all of the maps. A program of
4003 instructions with 2000 gotox and one 500k entry map holding two
distinct targets has 4000 indirect jump edges, 0.4% of the limit, and
takes 343s to be rejected. The map costs next to nothing to prepare, as
an unset entry is already a valid target. At the insn limit, with a single
1M entry map, the same shape extrapolates to 41 hours.
All gotox instructions of a subprogram share the same jump table, so
build the table of every subprogram in a single pass over the maps and
let every gotox use it in place: bpf_insn_successors() resolves a gotox
through its containing subprogram, and the tables are kept until the
verifier environment is torn down, so nothing is copied into the
instruction aux data. The edge accounting in visit_gotox_insn() moves
to the first visit of the gotox, marked by the BRANCH bit of its CFG
state, and a revisit returns right away since the first one already
pushed every target that still needed exploring.
The only user left of insn_aux_data[].jt is then the table which
visit_abnormal_return_insn() allocates for tail_call and
ld_{abs,ind} insns, so that bpf_insn_successors() reports the hidden
exit from their subprogram. Both are recognisable by their opcode, so
derive that edge in bpf_insn_successors() from the exit_idx of the
containing subprogram instead, again leaving it out when the subprogram
has no exit, and drop the field along with bpf_clear_insn_aux_data().
The instruction aux data then owns no allocation, thus nothing needs to
be freed when insns are removed or the verifier environment is torn
down.
Subprograms removed as dead code free their table in
adjust_subprog_starts_after_remove(), which also clears the slots that
the compaction of subprog_info vacates, as they still hold copies of the
moved entries and with them their table pointers.
check_cfg() is then linear in the number of map entries plus the number
of indirect jump edges, so what still scales now with the program is what
BPF_MAX_GOTOX_EDGES bounds:
gotox map entries edges before after
----------------------------------------------
500 250000 1000 35.34s 0.07s
1000 250000 2000 82.00s 0.07s
2000 250000 4000 148.62s 0.07s
2000 125000 4000 74.54s 0.05s
2000 500000 4000 342.71s 0.19s
Fixes: 493d9e0d6083 ("bpf, x86: add support for indirect jumps")
Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260925175244.1136329-6-daniel@iogearbox.net
|
|
visit_abnormal_return_insn() gives tail_call and ld_{abs,ind} insns a
second successor, the exit_idx of their subprogram, so that the hidden
exit from the subprogram is part of the CFG. check_subprogs() only
assigns exit_idx when it walks over a BPF_EXIT, and a subprogram whose
last insn jumps back into itself never has one, so exit_idx stays zero
and the edge points at insn 0 of the program:
0: r0 = 0
1: r0 = *(u8 *)skb[0]
2: goto -2
For a subprogram other than main, update_insn() then reads the liveness
masks at a negative relative index. Such a subprogram can still load
when it leaves through a bpf_throw(), so mark exit_idx as unset in that
case and leave the hidden edge out, as there is no exit for it to reach.
Fixes: e40f5a6bf88a ("bpf: correct stack liveness for tail calls")
Fixes: ee861486e377 ("bpf: Fix ld_{abs,ind} failure path analysis in subprogs")
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260925175244.1136329-5-daniel@iogearbox.net
|