summaryrefslogtreecommitdiff
path: root/tools/include/uapi
diff options
context:
space:
mode:
authorKumar Kartikeya Dwivedi <memxor@gmail.com>2026-09-25 06:55:29 +0200
committerAlexei Starovoitov <ast@kernel.org>2026-09-25 21:04:35 +0000
commite831ff570908c572d7ee5389a7e03a8637d1d8ee (patch)
treeb9fd2f8a6876e875c1f0ff187c66b3ad69f8381a /tools/include/uapi
parentf17bedcd29517739a6491c1c51cc6a593f6a0332 (diff)
downloadlinux-next-e831ff570908c572d7ee5389a7e03a8637d1d8ee.tar.gz
linux-next-e831ff570908c572d7ee5389a7e03a8637d1d8ee.zip
bpf: Add file descriptor interface for program streams
The existing BPF_PROG_STREAM_READ_BY_FD command only supports polling a program stream through repeated bpf() calls. It cannot block for new data or integrate with poll-based event loops. Add BPF_PROG_STREAM_OPEN to return a read-only, close-on-exec file descriptor for a selected program stream. Reads block by default and BPF_F_STREAM_NONBLOCK, the only accepted flag, provides non-blocking behavior. poll reports readable data and reports hangup once the program has been freed. Like pipes and sockets, the descriptor is not seekable and lseek fails with ESPIPE. A stream descriptor deliberately does not retain the program. Move each stream into a separately refcounted allocation so program teardown can mark it dead and wake descriptor users while outstanding descriptors drain buffered data safely. Readers sample the dead flag before looking for data, so EOF is reported only when the stream was already dead before it was found empty; data published right before teardown is never skipped. Only programs loaded through BPF_PROG_LOAD get streams. Classic BPF filters, JIT subprograms and shim programs never write to one, and kernel-side writers already resolve a subprogram to its main program, so those programs no longer carry stream state. Readiness needs its own counter. Stream capacity is charged before allocation and before an element is published to the stream log, so using that reservation as the read and poll condition can report readable data while no element exists: a blocking reader retries instead of sleeping and a lone non-blocking reader can see POLLIN followed by EAGAIN. Publish bytes with release ordering after adding elements to the lockless log, use acquire loads before consuming them or reporting readiness, limit each read to its readable snapshot and subtract only bytes actually copied. This keeps the aggregate count correct even when concurrent publishers update it out of publication order. With several readers on one stream, readiness remains advisory, as it is for pipes. The capacity counter is kept solely for enforcing the stream size limit. Wakeups are always deferred through irq_work. Stream writers run in whatever context the program runs in: NMI context for perf_event programs, sections with interrupts disabled inside bpf_spin_lock or rqspinlock critical sections since bpf_stream_vprintk() is KF_SPINLOCK_SAFE, and tracing programs attached anywhere in the kernel, including inside the wait queue and epoll code itself. Waking waiters directly from there can deadlock, and no cheap context check covers every case: on PREEMPT_RT, spinlock_t sections do not disable interrupts, so in_nmi() or irqs_disabled() cannot tell such a program apart from a benign one. Queue an irq_work item instead, as bpf_ringbuf does. Queue it only when a publication turns an empty stream readable. Readers block and pollers wait only after finding the stream empty, and the readable count never drops below zero because each read is bounded by its snapshot, so the first publication after such an observation is the one that makes the count positive, and it is the one that queues the wakeup. Publications into a stream that already holds data raise no interrupt, so a program that prints while nobody drains its stream pays for a single irq_work until the stream is emptied again. This matches bpf_ringbuf, which notifies only once the consumer has caught up. Blocking readers and level-triggered pollers re-check the readable count before waiting, so they cannot miss data, and edge-triggered epoll consumers drain until EAGAIN before waiting again, as epoll(7) requires. Synchronize pending work before releasing the final stream reference so the callback cannot outlive the stream, but only when the work was ever queued: irq_work_sync() waits for an RCU grace period on PREEMPT_RT and on architectures without an irq_work interrupt. Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260925045536.1480933-3-memxor@gmail.com
Diffstat (limited to 'tools/include/uapi')
-rw-r--r--tools/include/uapi/linux/bpf.h39
1 files changed, 39 insertions, 0 deletions
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 0aaa54359aeb..4687c3310996 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -936,6 +936,33 @@ union bpf_iter_link_info {
* 0 on success or -1 if an error occurred (in which case,
* *errno* is set appropriately).
*
+ * BPF_PROG_STREAM_OPEN
+ * Description
+ * Open a file descriptor for one of the BPF streams associated
+ * with the program identified by *prog_fd*. The stream is selected
+ * by *stream_id*.
+ *
+ * The returned file descriptor supports **read**\ (2) and
+ * **poll**\ (2). Reads block while the stream is empty unless
+ * **BPF_F_STREAM_NONBLOCK** is specified in *flags*. A non-blocking
+ * read of an empty stream fails with **EAGAIN**.
+ *
+ * **poll**\ (2) reports **POLLIN** when data is available and
+ * **POLLHUP** once the program has been freed, that is, after every
+ * reference to it, including links and other file descriptors, has
+ * been dropped. Hangup may lag the final release because program
+ * teardown is deferred. Buffered data remains readable after
+ * **POLLHUP** and a read returns zero after all such data has been
+ * consumed.
+ *
+ * The file descriptor is read-only and has the close-on-exec flag
+ * set. It is not seekable and **lseek**\ (2) fails with **ESPIPE**.
+ * *flags* may only contain **BPF_F_STREAM_NONBLOCK**.
+ *
+ * Return
+ * A new file descriptor (a nonnegative integer), or -1 if an
+ * error occurred (in which case, *errno* is set appropriately).
+ *
* NOTES
* eBPF objects (maps and programs) can be shared between processes.
*
@@ -993,6 +1020,7 @@ enum bpf_cmd {
BPF_TOKEN_CREATE,
BPF_PROG_STREAM_READ_BY_FD,
BPF_PROG_ASSOC_STRUCT_OPS,
+ BPF_PROG_STREAM_OPEN,
__MAX_BPF_CMD,
BPF_COMMON_ATTRS = 1 << 16, /* Indicate carrying syscall common attrs. */
};
@@ -1524,6 +1552,11 @@ enum {
BPF_STREAM_STDERR = 2,
};
+/* flags for BPF_PROG_STREAM_OPEN command */
+enum {
+ BPF_F_STREAM_NONBLOCK = (1U << 0),
+};
+
union bpf_attr {
struct { /* anonymous struct used by BPF_MAP_CREATE command */
__u32 map_type; /* one of enum bpf_map_type */
@@ -1950,6 +1983,12 @@ union bpf_attr {
__u32 flags;
} prog_assoc_struct_ops;
+ struct {
+ __u32 prog_fd;
+ __u32 stream_id;
+ __u32 flags;
+ } prog_stream_open;
+
} __attribute__((aligned(8)));
/* The description below is an attempt at providing documentation to eBPF