diff options
| author | Kumar Kartikeya Dwivedi <memxor@gmail.com> | 2026-09-25 06:55:29 +0200 |
|---|---|---|
| committer | Alexei Starovoitov <ast@kernel.org> | 2026-09-25 21:04:35 +0000 |
| commit | e831ff570908c572d7ee5389a7e03a8637d1d8ee (patch) | |
| tree | b9fd2f8a6876e875c1f0ff187c66b3ad69f8381a /tools/include/uapi | |
| parent | f17bedcd29517739a6491c1c51cc6a593f6a0332 (diff) | |
| download | linux-next-e831ff570908c572d7ee5389a7e03a8637d1d8ee.tar.gz linux-next-e831ff570908c572d7ee5389a7e03a8637d1d8ee.zip | |
bpf: Add file descriptor interface for program streams
The existing BPF_PROG_STREAM_READ_BY_FD command only supports polling a
program stream through repeated bpf() calls. It cannot block for new data
or integrate with poll-based event loops.
Add BPF_PROG_STREAM_OPEN to return a read-only, close-on-exec file
descriptor for a selected program stream. Reads block by default and
BPF_F_STREAM_NONBLOCK, the only accepted flag, provides non-blocking
behavior. poll reports readable data and reports hangup once the program
has been freed. Like pipes and sockets, the descriptor is not seekable and
lseek fails with ESPIPE.
A stream descriptor deliberately does not retain the program. Move each
stream into a separately refcounted allocation so program teardown can mark
it dead and wake descriptor users while outstanding descriptors drain
buffered data safely. Readers sample the dead flag before looking for data,
so EOF is reported only when the stream was already dead before it was
found empty; data published right before teardown is never skipped.
Only programs loaded through BPF_PROG_LOAD get streams. Classic BPF
filters, JIT subprograms and shim programs never write to one, and
kernel-side writers already resolve a subprogram to its main program, so
those programs no longer carry stream state.
Readiness needs its own counter. Stream capacity is charged before
allocation and before an element is published to the stream log, so using
that reservation as the read and poll condition can report readable data
while no element exists: a blocking reader retries instead of sleeping and
a lone non-blocking reader can see POLLIN followed by EAGAIN. Publish bytes
with release ordering after adding elements to the lockless log, use
acquire loads before consuming them or reporting readiness, limit each read
to its readable snapshot and subtract only bytes actually copied. This
keeps the aggregate count correct even when concurrent publishers update it
out of publication order. With several readers on one stream, readiness
remains advisory, as it is for pipes. The capacity counter is kept solely
for enforcing the stream size limit.
Wakeups are always deferred through irq_work. Stream writers run in
whatever context the program runs in: NMI context for perf_event programs,
sections with interrupts disabled inside bpf_spin_lock or rqspinlock
critical sections since bpf_stream_vprintk() is KF_SPINLOCK_SAFE, and
tracing programs attached anywhere in the kernel, including inside the wait
queue and epoll code itself. Waking waiters directly from there can
deadlock, and no cheap context check covers every case: on PREEMPT_RT,
spinlock_t sections do not disable interrupts, so in_nmi() or
irqs_disabled() cannot tell such a program apart from a benign one. Queue
an irq_work item instead, as bpf_ringbuf does.
Queue it only when a publication turns an empty stream readable. Readers
block and pollers wait only after finding the stream empty, and the
readable count never drops below zero because each read is bounded by its
snapshot, so the first publication after such an observation is the one
that makes the count positive, and it is the one that queues the wakeup.
Publications into a stream that already holds data raise no interrupt, so
a program that prints while nobody drains its stream pays for a single
irq_work until the stream is emptied again. This matches bpf_ringbuf, which
notifies only once the consumer has caught up. Blocking readers and
level-triggered pollers re-check the readable count before waiting, so they
cannot miss data, and edge-triggered epoll consumers drain until EAGAIN
before waiting again, as epoll(7) requires.
Synchronize pending work before releasing the final stream reference so
the callback cannot outlive the stream, but only when the work was ever
queued: irq_work_sync() waits for an RCU grace period on PREEMPT_RT and on
architectures without an irq_work interrupt.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260925045536.1480933-3-memxor@gmail.com
Diffstat (limited to 'tools/include/uapi')
| -rw-r--r-- | tools/include/uapi/linux/bpf.h | 39 |
1 files changed, 39 insertions, 0 deletions
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h index 0aaa54359aeb..4687c3310996 100644 --- a/tools/include/uapi/linux/bpf.h +++ b/tools/include/uapi/linux/bpf.h @@ -936,6 +936,33 @@ union bpf_iter_link_info { * 0 on success or -1 if an error occurred (in which case, * *errno* is set appropriately). * + * BPF_PROG_STREAM_OPEN + * Description + * Open a file descriptor for one of the BPF streams associated + * with the program identified by *prog_fd*. The stream is selected + * by *stream_id*. + * + * The returned file descriptor supports **read**\ (2) and + * **poll**\ (2). Reads block while the stream is empty unless + * **BPF_F_STREAM_NONBLOCK** is specified in *flags*. A non-blocking + * read of an empty stream fails with **EAGAIN**. + * + * **poll**\ (2) reports **POLLIN** when data is available and + * **POLLHUP** once the program has been freed, that is, after every + * reference to it, including links and other file descriptors, has + * been dropped. Hangup may lag the final release because program + * teardown is deferred. Buffered data remains readable after + * **POLLHUP** and a read returns zero after all such data has been + * consumed. + * + * The file descriptor is read-only and has the close-on-exec flag + * set. It is not seekable and **lseek**\ (2) fails with **ESPIPE**. + * *flags* may only contain **BPF_F_STREAM_NONBLOCK**. + * + * Return + * A new file descriptor (a nonnegative integer), or -1 if an + * error occurred (in which case, *errno* is set appropriately). + * * NOTES * eBPF objects (maps and programs) can be shared between processes. * @@ -993,6 +1020,7 @@ enum bpf_cmd { BPF_TOKEN_CREATE, BPF_PROG_STREAM_READ_BY_FD, BPF_PROG_ASSOC_STRUCT_OPS, + BPF_PROG_STREAM_OPEN, __MAX_BPF_CMD, BPF_COMMON_ATTRS = 1 << 16, /* Indicate carrying syscall common attrs. */ }; @@ -1524,6 +1552,11 @@ enum { BPF_STREAM_STDERR = 2, }; +/* flags for BPF_PROG_STREAM_OPEN command */ +enum { + BPF_F_STREAM_NONBLOCK = (1U << 0), +}; + union bpf_attr { struct { /* anonymous struct used by BPF_MAP_CREATE command */ __u32 map_type; /* one of enum bpf_map_type */ @@ -1950,6 +1983,12 @@ union bpf_attr { __u32 flags; } prog_assoc_struct_ops; + struct { + __u32 prog_fd; + __u32 stream_id; + __u32 flags; + } prog_stream_open; + } __attribute__((aligned(8))); /* The description below is an attempt at providing documentation to eBPF |
