summaryrefslogtreecommitdiff
path: root/tools/include
diff options
context:
space:
mode:
authorPuranjay Mohan <puranjay@kernel.org>2026-09-22 13:00:50 -0700
committerAlexei Starovoitov <ast@kernel.org>2026-09-22 23:44:24 +0000
commit1832696b22078f57069a676266caaeb4b0562ad7 (patch)
treeae54d7aee6eb09d11965776d7fecea9f7feb20cb /tools/include
parent63b13537e6b223699dabb60ec0009b0c976e7181 (diff)
downloadlinux-next-1832696b22078f57069a676266caaeb4b0562ad7.tar.gz
linux-next-1832696b22078f57069a676266caaeb4b0562ad7.zip
bpf: Add bpf_call_rcu() kfunc
BPF programs that manage their own objects have no way to run their own logic once an RCU grace period has elapsed. bpf_obj_drop() defers a free, but returning an index to an allocator or unpinning a resource once readers are done has no equivalent. sched_ext's BPF library works around this today by pushing freed nodes onto a list and having a userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a BPF program to reclaim them. Add: int bpf_call_rcu(struct bpf_rcu_head *rh, void *map, int (*callback)(struct bpf_map *map, void *key, void *value)); @rh is a struct bpf_rcu_head embedded in a value of @map, so the callback runs as callback(map, key, value) for the element it lives in and needs no cookie. A head can only be armed once, which bounds outstanding work by the number of elements. struct bpf_rcu_head holds the callback state inline rather than a pointer to it, as bpf_timer, bpf_wq and bpf_task_work do, because there is nothing to cancel and so nothing that has to outlive the map value. That avoids an allocation and a state machine on the arming path at the cost of 48 bytes per element. An RCU callback cannot be cancelled, so everything it touches has to stay alive until it runs: - The callback is the program's text, so arming takes a program reference as bpf_timer, bpf_wq and bpf_task_work do, dropped once the callback returns. bpf_prog_inc_not_zero() also fails the arm with -EBADF once the program is dying. - The map is held by that reference through used_maps. An inner map is not, so bpf_rcu_head is rejected in one. - The field is only accepted in BPF_MAP_TYPE_ARRAY, whose elements are never freed individually. A hash element can be deleted and recycled while a callback is queued on it. - The head is disarmed before the callback runs so it can be armed again from there, which takes a new program reference before the running callback drops its own. Arming therefore fails with -EPERM once the map is held by neither a process nor bpffs. bpf_iter hands a program a writable pointer to the live element, which would let it overwrite a queued head, so bpf_iter_attach_map() rejects maps carrying one. The callback is verified non-sleepable even when the caller is sleepable, and RCU invokes it with BH disabled. Signed-off-by: Puranjay Mohan <puranjay@kernel.org> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260922200208.3203834-2-puranjay@kernel.org
Diffstat (limited to 'tools/include')
-rw-r--r--tools/include/uapi/linux/bpf.h4
1 files changed, 4 insertions, 0 deletions
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 6330b7d745c5..0aaa54359aeb 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -7611,6 +7611,10 @@ struct bpf_task_work {
__u64 __opaque;
} __attribute__((aligned(8)));
+struct bpf_rcu_head {
+ __u64 __opaque[6];
+} __attribute__((aligned(8)));
+
struct bpf_wq {
__u64 __opaque[2];
} __attribute__((aligned(8)));