diff options
| author | Puranjay Mohan <puranjay@kernel.org> | 2026-09-22 13:00:50 -0700 |
|---|---|---|
| committer | Alexei Starovoitov <ast@kernel.org> | 2026-09-22 23:44:24 +0000 |
| commit | 1832696b22078f57069a676266caaeb4b0562ad7 (patch) | |
| tree | ae54d7aee6eb09d11965776d7fecea9f7feb20cb /tools/include | |
| parent | 63b13537e6b223699dabb60ec0009b0c976e7181 (diff) | |
| download | linux-next-1832696b22078f57069a676266caaeb4b0562ad7.tar.gz linux-next-1832696b22078f57069a676266caaeb4b0562ad7.zip | |
bpf: Add bpf_call_rcu() kfunc
BPF programs that manage their own objects have no way to run their own
logic once an RCU grace period has elapsed. bpf_obj_drop() defers a
free, but returning an index to an allocator or unpinning a resource
once readers are done has no equivalent. sched_ext's BPF library works
around this today by pushing freed nodes onto a list and having a
userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a
BPF program to reclaim them.
Add:
int bpf_call_rcu(struct bpf_rcu_head *rh, void *map,
int (*callback)(struct bpf_map *map, void *key,
void *value));
@rh is a struct bpf_rcu_head embedded in a value of @map, so the
callback runs as callback(map, key, value) for the element it lives in
and needs no cookie. A head can only be armed once, which bounds
outstanding work by the number of elements.
struct bpf_rcu_head holds the callback state inline rather than a
pointer to it, as bpf_timer, bpf_wq and bpf_task_work do, because there
is nothing to cancel and so nothing that has to outlive the map value.
That avoids an allocation and a state machine on the arming path at the
cost of 48 bytes per element.
An RCU callback cannot be cancelled, so everything it touches has to
stay alive until it runs:
- The callback is the program's text, so arming takes a program
reference as bpf_timer, bpf_wq and bpf_task_work do, dropped once
the callback returns. bpf_prog_inc_not_zero() also fails the arm
with -EBADF once the program is dying.
- The map is held by that reference through used_maps. An inner map
is not, so bpf_rcu_head is rejected in one.
- The field is only accepted in BPF_MAP_TYPE_ARRAY, whose elements
are never freed individually. A hash element can be deleted and
recycled while a callback is queued on it.
- The head is disarmed before the callback runs so it can be armed
again from there, which takes a new program reference before the
running callback drops its own. Arming therefore fails with -EPERM
once the map is held by neither a process nor bpffs.
bpf_iter hands a program a writable pointer to the live element, which
would let it overwrite a queued head, so bpf_iter_attach_map() rejects
maps carrying one.
The callback is verified non-sleepable even when the caller is
sleepable, and RCU invokes it with BH disabled.
Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260922200208.3203834-2-puranjay@kernel.org
Diffstat (limited to 'tools/include')
| -rw-r--r-- | tools/include/uapi/linux/bpf.h | 4 |
1 files changed, 4 insertions, 0 deletions
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h index 6330b7d745c5..0aaa54359aeb 100644 --- a/tools/include/uapi/linux/bpf.h +++ b/tools/include/uapi/linux/bpf.h @@ -7611,6 +7611,10 @@ struct bpf_task_work { __u64 __opaque; } __attribute__((aligned(8))); +struct bpf_rcu_head { + __u64 __opaque[6]; +} __attribute__((aligned(8))); + struct bpf_wq { __u64 __opaque[2]; } __attribute__((aligned(8))); |
