From 3b55f350c68a0aceff108f47f9d31f47ebffaf7b Mon Sep 17 00:00:00 2001 From: Daniel Borkmann Date: Mon, 7 Sep 2026 14:10:24 +0200 Subject: bpf: Fix bpf_skb_change_tail wrt csum partial skbs Cilium generates ICMP "frag needed" replies from BPF when a LB DSR packet exceeds the egress MTU. The reply is built by first trimming the packet down to target size via bpf_skb_change_tail(), and then pushing the ICMP error headers in front of it. The trim is rejected for skbs which carry a checksum offload, e.g. TCP packets aggregated by GRO on ingress where tcp_gro_complete() leaves the skb as CHECKSUM_PARTIAL. __bpf_skb_min_len() raises the minimum length to the end of the L4 checksum field, so a trim to 42 bytes bails out with -EINVAL given a min_len of 52 in this case, and due to that the ICMP generator fails. This is not the case if GRO is turned off. Fix this bpf_skb_change_tail() restriction and drop the checksum offload when the new length no longer covers the checksum field. The BPF program rewrites the skb into an ICMP error and computes the checksum itself anyway. Fixes: 5293efe62df8 ("bpf: add bpf_skb_change_tail helper") Reported-by: Tom Hadlaw Reported-by: Yusuke Suzuki Signed-off-by: Daniel Borkmann Link: https://lore.kernel.org/r/20260907121025.1923656-1-daniel@iogearbox.net Signed-off-by: Alexei Starovoitov --- net/core/filter.c | 11 +++++------ 1 file changed, 5 insertions(+), 6 deletions(-) (limited to 'net') diff --git a/net/core/filter.c b/net/core/filter.c index 61940e753552..8513167a858a 100644 --- a/net/core/filter.c +++ b/net/core/filter.c @@ -3961,12 +3961,6 @@ static u32 __bpf_skb_min_len(const struct sk_buff *skb) if (offset > 0) min_len = offset; } - if (skb->ip_summed == CHECKSUM_PARTIAL) { - offset = skb_checksum_start_offset(skb) + - skb->csum_offset + sizeof(__sum16); - if (offset > 0) - min_len = offset; - } return min_len; } @@ -3983,6 +3977,11 @@ static int bpf_skb_grow_rcsum(struct sk_buff *skb, unsigned int new_len) static int bpf_skb_trim_rcsum(struct sk_buff *skb, unsigned int new_len) { + if (skb->ip_summed == CHECKSUM_PARTIAL && + new_len < skb_checksum_start_offset(skb) + skb->csum_offset + + sizeof(__sum16)) + skb->ip_summed = CHECKSUM_NONE; + return __skb_trim_rcsum(skb, new_len); } -- cgit v1.2.3 From e4a62833adff6ef0fe7c0b90393204fe3c26b5c5 Mon Sep 17 00:00:00 2001 From: Weiming Shi Date: Wed, 9 Sep 2026 12:08:08 +0800 Subject: bpf: Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL An LWT_SEG6LOCAL program can invalidate its cached SRH with bpf_lwt_seg6_adjust_srh() and then call bpf_skb_pull_data(). The latter may reallocate skb->head, leaving the per-CPU SRH pointer dangling. Post-program SRH validation then writes through that pointer. Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL programs so the verifier rejects this unsafe helper combination. Other LWT program types continue to expose the helper through lwt_out_func_proto(). Fixes: 004d4b274e2a ("ipv6: sr: Add seg6local action End.BPF") Reported-by: co+adfca3e91be95776@bugs.sh Suggested-by: Alexei Starovoitov Signed-off-by: Weiming Shi Signed-off-by: Daniel Borkmann Reviewed-by: Emil Tsalapatis Closes: https://lore.kernel.org/all/GCy0KRM2IcQGoJQTjJEU9D0maBxXzEDHuQpq@bugs.sh/ Link: https://lore.kernel.org/bpf/DL9COXZQXX4V.1FN45QO2Q77ZH@gmail.com/ Link: https://lore.kernel.org/bpf/20260909040807.3885815-2-bestswngs@gmail.com --- net/core/filter.c | 2 ++ 1 file changed, 2 insertions(+) (limited to 'net') diff --git a/net/core/filter.c b/net/core/filter.c index 8513167a858a..2a84f9d01131 100644 --- a/net/core/filter.c +++ b/net/core/filter.c @@ -9044,6 +9044,8 @@ static const struct bpf_func_proto * lwt_seg6local_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog) { switch (func_id) { + case BPF_FUNC_skb_pull_data: + return NULL; #if IS_ENABLED(CONFIG_IPV6_SEG6_BPF) case BPF_FUNC_lwt_seg6_store_bytes: return &bpf_lwt_seg6_store_bytes_proto; -- cgit v1.2.3 From 01b245ba016d44861690594e10f67e026ce8552f Mon Sep 17 00:00:00 2001 From: Jiayuan Chen Date: Thu, 10 Sep 2026 19:26:26 +0800 Subject: bpf: Fix out-of-bounds read of sk_protocol in bpf_sock_destroy() sk_protocol lives in struct sock, not in struct sock_common. A timewait or request sock handed to bpf_sock_destroy() by the tcp iterator is neither, so reading sk->sk_protocol runs past the object: ================================================================== BUG: KASAN: slab-out-of-bounds in bpf_sock_destroy+0xc7/0xe0 Read of size 2 at addr ffff8881047d11b4 by task test_progs/428 Tainted: [W]=WARN Call Trace: dump_stack_lvl+0x91/0xf0 print_report+0xd1/0x630 kasan_report+0xf3/0x130 __asan_report_load2_noabort+0x14/0x30 bpf_sock_destroy+0xc7/0xe0 bpf_prog_c3dd61f9d9cd9f37_iter_tcp6_timewait+0x9f/0xb7 bpf_iter_run_prog+0x538/0xde0 bpf_iter_tcp_seq_show+0x26b/0x4b0 bpf_seq_read+0x424/0x1210 vfs_read+0x197/0xe40 ksys_read+0x119/0x240 __x64_sys_read+0x72/0xc0 x64_sys_call+0x647/0x27e0 do_syscall_64+0xe5/0x610 entry_SYSCALL_64_after_hwframe+0x76/0x7e Only check sk_protocol on full socks. tcp_abort() already knows how to deal with TIME_WAIT and NEW_SYN_RECV socks. Also fix the comment, it never matched the code. Fixes: 4ddbcb886268 ("bpf: Add bpf_sock_destroy kfunc") Reported-by: Xiang Mei (Microsoft) Closes: https://lore.kernel.org/bpf/20260702224519.800135-1-xmei5@asu.edu/ Signed-off-by: Jiayuan Chen Reviewed-by: Kuniyuki Iwashima Link: https://lore.kernel.org/r/20260910112634.152195-1-jiayuan.chen@linux.dev Signed-off-by: Alexei Starovoitov --- net/core/filter.c | 13 +++++++++---- 1 file changed, 9 insertions(+), 4 deletions(-) (limited to 'net') diff --git a/net/core/filter.c b/net/core/filter.c index 2a84f9d01131..cae43b999162 100644 --- a/net/core/filter.c +++ b/net/core/filter.c @@ -12913,8 +12913,9 @@ __bpf_kfunc_start_defs(); * @sock: Pointer to socket to be destroyed * * Return: - * On error, may return EPROTONOSUPPORT, EINVAL. - * EPROTONOSUPPORT if protocol specific destroy handler is not supported. + * On error, may return EOPNOTSUPP, or whatever the protocol specific + * destroy handler returns. + * EOPNOTSUPP if protocol specific destroy handler is not supported. * 0 otherwise */ __bpf_kfunc int bpf_sock_destroy(struct sock_common *sock) @@ -12926,8 +12927,12 @@ __bpf_kfunc int bpf_sock_destroy(struct sock_common *sock) * Supporting protocols will need to acquire sock lock in the BPF context * prior to invoking this kfunc. */ - if (!sk->sk_prot->diag_destroy || (sk->sk_protocol != IPPROTO_TCP && - sk->sk_protocol != IPPROTO_UDP)) + if (!sk->sk_prot->diag_destroy) + return -EOPNOTSUPP; + + if (sk_fullsock(sk) && + sk->sk_protocol != IPPROTO_TCP && + sk->sk_protocol != IPPROTO_UDP) return -EOPNOTSUPP; return sk->sk_prot->diag_destroy(sk, ECONNABORTED); -- cgit v1.2.3 From eaab8cab451b9502ce224cd202550375b894a467 Mon Sep 17 00:00:00 2001 From: Jiayuan Chen Date: Thu, 10 Sep 2026 19:27:28 +0800 Subject: tcp: Skip cond_resched() in inet_csk_listen_stop() under BPF context bpf_sock_destroy() runs from the tcp iterator, under rcu_read_lock(). If the sock is a listener that still has children in its accept queue, tcp_abort() ends up in inet_csk_listen_stop() and the cond_resched() there trips the debug check: BUG: sleeping function called from invalid context at net/ipv4/inet_connection_sock.c:1523 in_atomic(): 0, irqs_disabled(): 0, non_block: 0, pid: 628, name: test_progs preempt_count: 0, expected: 0 RCU nest depth: 1, expected: 0 locks held by test_progs/628: 3, last CPU#3: #0: ffff8881158cee18 (&p->lock){+.+.}-{4:4}, at: bpf_seq_read+0x56/0x1210 #1: ffff8881106bb858 (sk_lock-AF_INET6){+.+.}-{0:0}, at: bpf_iter_tcp_seq_show+0x32b/0x4b0 #2: ffffffffb435af20 (rcu_read_lock){....}-{1:3}, at: bpf_iter_run_prog+0x46b/0xde0 CPU: 3 UID: 0 PID: 628 Comm: test_progs Tainted: G W 7.2.0+ #65 PREEMPT Tainted: [W]=WARN Call Trace: dump_stack_lvl+0xc1/0xf0 dump_stack+0x10/0x20 __might_resched+0x3d2/0x610 inet_csk_listen_stop+0x7b/0xbf0 tcp_abort+0x23b/0x3b0 bpf_sock_destroy+0xfc/0x140 bpf_prog_448133d24601754f_iter_tcp6_server+0x81/0x8a bpf_iter_run_prog+0x538/0xde0 bpf_iter_tcp_seq_show+0x26b/0x4b0 bpf_seq_read+0x424/0x1210 vfs_read+0x197/0xe40 ksys_read+0x119/0x240 __x64_sys_read+0x72/0xc0 x64_sys_call+0x647/0x27e0 do_syscall_64+0xe5/0x610 entry_SYSCALL_64_after_hwframe+0x76/0x7e RIP: 0033:0x7fad39b28aca RSP: 002b:00007ffc381c61c0 EFLAGS: 00000246 ORIG_RAX: 0000000000000000 RAX: ffffffffffffffda RBX: 00007ffc381c6a88 RCX: 00007fad39b28aca RDX: 0000000000000032 RSI: 00007ffc381c6250 RDI: 0000000000000014 RBP: 00007ffc381c61e0 R08: 0000000000000000 R09: 0000000000000000 R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000003 R13: 0000000000000000 R14: 000055f077c1bbb0 R15: 00007fad3a0f3000 The commit that added the kfunc already guards lock_sock() in tcp_abort() and udp_abort() with has_current_bpf_ctx(), but missed the listener path. Do the same for the cond_resched(). The loop runs inside the iterator's rcu_read_lock(), it must not reschedule or report a quiescent state there. Fixes: 4ddbcb886268 ("bpf: Add bpf_sock_destroy kfunc") Signed-off-by: Jiayuan Chen Link: https://lore.kernel.org/r/20260910112736.153710-1-jiayuan.chen@linux.dev Signed-off-by: Alexei Starovoitov --- net/ipv4/inet_connection_sock.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) (limited to 'net') diff --git a/net/ipv4/inet_connection_sock.c b/net/ipv4/inet_connection_sock.c index 6257459bcee2..6a30f1138454 100644 --- a/net/ipv4/inet_connection_sock.c +++ b/net/ipv4/inet_connection_sock.c @@ -1520,7 +1520,8 @@ skip_child_forget: local_bh_enable(); sock_put(child); - cond_resched(); + if (!has_current_bpf_ctx()) + cond_resched(); } if (queue->fastopenq.rskq_rst_head) { /* Free all the reqs queued in rskq_rst_head. */ -- cgit v1.2.3 From 75f8cf22463d82bb1fb0239a3d485fc8f4c8ef03 Mon Sep 17 00:00:00 2001 From: Jiayuan Chen Date: Thu, 3 Sep 2026 18:09:20 +0800 Subject: bpf: Fix out-of-bounds read of rtt_min in sock_ops A sockops prog reading skops->rtt_min never checks the sk type: on the tcp_conn_request() path sock_ops->sk is a request_sock (non-full), and the ctx rewrite casts it to a tcp_sock (full) and reads rtt_min past the end of the request_sock, returning dirty adjacent memory. SEC("sockops") int prog(struct bpf_sock_ops *skops) { switch (skops->op) { case BPF_SOCK_OPS_RWND_INIT: leak = skops->rtt_min; /* reads the request_sock OOB */ ... } } For instance one such read returned rtt_min=0xffff8881, the high half of a leaked kernel pointer. Guarding that cast is exactly what SOCK_OPS_GET_FIELD() does -- it checks is_locked_tcp_sock and returns 0 when sock_ops->sk is not a locked full socket. Every other tcp_sock field in sock_ops goes through it; rtt_min is the only one open-coded, so it skips the check. Read rtt_min through SOCK_OPS_GET_FIELD() too. rtt_min is a bit special: it is a struct minmax and we only want the current min, so pass rtt_min.s[0].v. That is equivalent to the old hand-computed offset offsetof(struct tcp_sock, rtt_min) + sizeof_field(struct minmax_sample, t) (s[0] sits at rtt_min + 0 and .v at + sizeof(.t), i.e. what minmax_get() returns), so the loaded field is unchanged and only the full-sock guard is added. The two BUILD_BUG_ON()s that protected the hand-computed offset are no longer needed. Before patch: 0: r1 = *(u64 *)(r1 +0) ; r1 = skops->sk 1: r1 = *(u32 *)(r1 +2324) ; ((tcp_sock *)sk)->rtt_min.s[0].v After patch: 0: *(u64 *)(r1 +56) = r9 1: r9 = *(u8 *)(r1 +50) ; is_locked_tcp_sock 2: if r9 == 0 goto pc+4 ; not a locked full sock -> 0 3: r9 = *(u64 *)(r1 +56) 4: r1 = *(u64 *)(r1 +0) ; r1 = skops->sk 5: r1 = *(u32 *)(r1 +2324) ; rtt_min.s[0].v 6: goto pc+2 7: r9 = *(u64 *)(r1 +56) 8: r1 = 0 Fixes: 44f0e43037d3 ("bpf: Add support for reading sk_state and more") Reported-by: VEGA Signed-off-by: Jiayuan Chen Reviewed-by: Emil Tsalapatis Link: https://lore.kernel.org/r/20260903100921.113374-1-jiayuan.chen@linux.dev Signed-off-by: Alexei Starovoitov --- net/core/filter.c | 13 +------------ 1 file changed, 1 insertion(+), 12 deletions(-) (limited to 'net') diff --git a/net/core/filter.c b/net/core/filter.c index cae43b999162..532405988fd9 100644 --- a/net/core/filter.c +++ b/net/core/filter.c @@ -11106,18 +11106,7 @@ static u32 sock_ops_convert_ctx_access(enum bpf_access_type type, break; case offsetof(struct bpf_sock_ops, rtt_min): - BUILD_BUG_ON(sizeof_field(struct tcp_sock, rtt_min) != - sizeof(struct minmax)); - BUILD_BUG_ON(sizeof(struct minmax) < - sizeof(struct minmax_sample)); - - *insn++ = BPF_LDX_MEM(BPF_FIELD_SIZEOF( - struct bpf_sock_ops_kern, sk), - si->dst_reg, si->src_reg, - offsetof(struct bpf_sock_ops_kern, sk)); - *insn++ = BPF_LDX_MEM(BPF_W, si->dst_reg, si->dst_reg, - offsetof(struct tcp_sock, rtt_min) + - sizeof_field(struct minmax_sample, t)); + SOCK_OPS_GET_FIELD(rtt_min, rtt_min.s[0].v, struct tcp_sock); break; case offsetof(struct bpf_sock_ops, bpf_sock_ops_cb_flags): -- cgit v1.2.3 From 7d70a0b02d262971201fdd1e221586bdd95c2910 Mon Sep 17 00:00:00 2001 From: Zhixing Chen Date: Thu, 3 Sep 2026 18:43:58 +0800 Subject: bpf: Use kvfree() in xdp_test_run_teardown() xdp_test_run_setup() allocates xdp->frames and xdp->skbs with kvmalloc_array(). The setup error path already releases both arrays with kvfree(), while the normal teardown path still uses kfree(). Use kvfree() in xdp_test_run_teardown() as well, so the release helper matches the allocator on both paths. Signed-off-by: Zhixing Chen Reviewed-by: Emil Tsalapatis Link: https://lore.kernel.org/r/20260903104358.29228-1-running910@gmail.com Signed-off-by: Alexei Starovoitov --- net/bpf/test_run.c | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) (limited to 'net') diff --git a/net/bpf/test_run.c b/net/bpf/test_run.c index 5d51f6cb7d15..513354e928cb 100644 --- a/net/bpf/test_run.c +++ b/net/bpf/test_run.c @@ -205,8 +205,8 @@ static void xdp_test_run_teardown(struct xdp_test_data *xdp) { xdp_unreg_mem_model(&xdp->mem); page_pool_destroy(xdp->pp); - kfree(xdp->frames); - kfree(xdp->skbs); + kvfree(xdp->frames); + kvfree(xdp->skbs); } static bool frame_was_changed(const struct xdp_page_head *head) -- cgit v1.2.3 From 490a83d6386eec1d29f470c8d7331677fb46c3b7 Mon Sep 17 00:00:00 2001 From: Geliang Tang Date: Tue, 8 Sep 2026 17:08:32 +0800 Subject: bpf, sockmap: Fix self-redirect copied_seq double-counting When a BPF stream_verdict program redirects an skb back to the same socket (self-redirect with BPF_F_INGRESS), sk_psock_verdict_apply() calls tcp_eat_skb() which advances tcp_sk->copied_seq. However, the skb is then delivered to the socket's psock ingress queue and later read by tcp_bpf_recvmsg_parser(), which also advances copied_seq via the copied_from_self accounting path. This double-counting causes copied_seq to advance by 2x the actual data length, triggering: TCP recvmsg seq # bug 2: copied BF2E806, seq BF2E7FD, \ rcvnxt BF2E806, fl 0 WARNING: net/ipv4/tcp.c:2745 at tcp_recvmsg_locked+0x72b/0x2640 Call Trace: tcp_recvmsg+0x10a/0x500 sock_recvmsg+0x168/0x1d0 __sys_recvfrom+0x19a/0x2a0 __x64_sys_recvfrom+0xe4/0x1f0 do_syscall_64+0xf7/0x530 entry_SYSCALL_64_after_hwframe+0x77/0x7f cleanup rbuf bug: copied BF2E806 seq BF2E806 rcvnxt BF2E806 WARNING: net/ipv4/tcp.c:1609 at tcp_cleanup_rbuf+0xf2/0x1c0 Call Trace: tcp_recvmsg_locked+0x8d1/0x2640 tcp_recvmsg+0x10a/0x500 sock_recvmsg+0x168/0x1d0 __sys_recvfrom+0x19a/0x2a0 __x64_sys_recvfrom+0xe4/0x1f0 do_syscall_64+0xf7/0x530 entry_SYSCALL_64_after_hwframe+0x77/0x7f Fix this by converting self-redirect verdict to __SK_PASS at the beginning of sk_psock_verdict_apply(). This bypasses the __SK_REDIRECT case entirely (which calls sk_psock_eat_skb), letting the __SK_PASS path queue the skb to the psock ingress queue. The data is then read via tcp_bpf_recvmsg_parser(), which advances copied_seq exactly once through copied_from_self. Cross-socket redirects continue through __SK_REDIRECT with sk_psock_eat_skb() unchanged. Fixes: e5c6de5fa025 ("bpf, sockmap: Incorrectly handling copied_seq") Suggested-by: Jakub Sitnicki Suggested-by: Jiayuan Chen Signed-off-by: Geliang Tang Reviewed-by: Emil Tsalapatis Reviewed-by: Jiayuan Chen Link: https://lore.kernel.org/r/1a8e797a1b26e2f695aaac22ac644c2862f63466.1788858299.git.tanggeliang@kylinos.cn Signed-off-by: Alexei Starovoitov --- net/core/skmsg.c | 4 ++++ 1 file changed, 4 insertions(+) (limited to 'net') diff --git a/net/core/skmsg.c b/net/core/skmsg.c index 2521b643fa05..df385a5a961e 100644 --- a/net/core/skmsg.c +++ b/net/core/skmsg.c @@ -1000,6 +1000,10 @@ static int sk_psock_verdict_apply(struct sk_psock *psock, struct sk_buff *skb, int err = 0; u32 len, off; + if (verdict == __SK_REDIRECT && skb_bpf_ingress(skb) && + skb_bpf_redirect_fetch(skb) == psock->sk) + verdict = __SK_PASS; + switch (verdict) { case __SK_PASS: err = -EIO; -- cgit v1.2.3 From 70504de0bb627848667207bec7ccfd647deb8814 Mon Sep 17 00:00:00 2001 From: Zhiling Zou Date: Fri, 11 Sep 2026 00:18:24 +0800 Subject: xsk: Use a 32-bit compare in xsk_map_gen_lookup xsk_map_gen_lookup() loads a u32 key and compares it with max_entries using BPF_JMP_IMM. BPF immediates are sign-extended to 64 bits, so a max_entries value of 0x80000000 or higher becomes a threshold larger than every zero-extended 32-bit key. An out-of-range index then skips the bounds check and the generated lookup reads past xsk_map[]. Compare with BPF_JMP32_IMM so the check stays in 32-bit unsigned range. Fixes: e65650f291ee ("bpf: Implement map_gen_lookup() callback for XSKMAP") Reported-by: Vega Signed-off-by: Zhiling Zou Signed-off-by: Alexei Starovoitov Reviewed-by: Emil Tsalapatis Link: https://patch.msgid.link/7d2cb8e8dfaa9eb8fdff85156987a60960787dc3.1789056660.git.zhilinz@nebusec.ai Signed-off-by: Eduard Zingerman --- net/xdp/xskmap.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) (limited to 'net') diff --git a/net/xdp/xskmap.c b/net/xdp/xskmap.c index 3bff346308d0..bf00d6463c19 100644 --- a/net/xdp/xskmap.c +++ b/net/xdp/xskmap.c @@ -124,7 +124,7 @@ static int xsk_map_gen_lookup(struct bpf_map *map, struct bpf_insn *insn_buf) struct bpf_insn *insn = insn_buf; *insn++ = BPF_LDX_MEM(BPF_W, ret, index, 0); - *insn++ = BPF_JMP_IMM(BPF_JGE, ret, map->max_entries, 5); + *insn++ = BPF_JMP32_IMM(BPF_JGE, ret, map->max_entries, 5); *insn++ = BPF_ALU64_IMM(BPF_LSH, ret, ilog2(sizeof(struct xsk_sock *))); *insn++ = BPF_ALU64_IMM(BPF_ADD, mp, offsetof(struct xsk_map, xsk_map)); *insn++ = BPF_ALU64_REG(BPF_ADD, ret, mp); -- cgit v1.2.3 From 4a4852376e3a2727ea40e61143d6d7c22bb6dfad Mon Sep 17 00:00:00 2001 From: Emil Tsalapatis Date: Tue, 22 Sep 2026 17:20:20 +0000 Subject: bpf: Fix bpf_sock context code generation Currently, the ctx access code reads the rx_queue_mapping field with either a 4-byte or 2-byte load. The rest of the bits in the register are marked known zero by the verifier. However, the emitted ctx access code places in the register on certain the special value (-1) using BPF_MOV_IMM64, which gets sign-extended to turn on all the bits in the register. By shifting this value right, the program ends up with a value at runtime above what the verifier assumes is possible. Fix this by ensuring the read value is as wide as the assumed size. Use MOV32 instructions instead of MOV64 instructions to keep the upper bits zero as assumed by the verifier. Also properly report the size of the destination variable (the bpf_sock field, 4 bytes) instead of the source (the socket field, 2 bytes). Fixes: c3c16f2ea6d2 ("bpf: Add rx_queue_mapping to bpf_sock") Reported-by: Nicholas Carlini Suggested-by: Nicholas Carlini Signed-off-by: Emil Tsalapatis Signed-off-by: Alexei Starovoitov Reviewed-by: Jiayuan Chen Link: https://patch.msgid.link/20260922172028.6269-4-emil@etsalapatis.com --- net/core/filter.c | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) (limited to 'net') diff --git a/net/core/filter.c b/net/core/filter.c index 532405988fd9..70dc621672f2 100644 --- a/net/core/filter.c +++ b/net/core/filter.c @@ -10566,11 +10566,12 @@ u32 bpf_sock_convert_ctx_access(enum bpf_access_type type, target_size)); *insn++ = BPF_JMP_IMM(BPF_JNE, si->dst_reg, NO_QUEUE_MAPPING, 1); - *insn++ = BPF_MOV64_IMM(si->dst_reg, -1); + *insn++ = BPF_MOV32_IMM(si->dst_reg, -1); #else - *insn++ = BPF_MOV64_IMM(si->dst_reg, -1); - *target_size = 2; + *insn++ = BPF_MOV32_IMM(si->dst_reg, -1); #endif + *target_size = sizeof_field(struct bpf_sock, rx_queue_mapping); + break; } -- cgit v1.2.3 From 814a81c842bd88f6bd8a4ce550d560df071a5d03 Mon Sep 17 00:00:00 2001 From: Zhao Gongyi Date: Thu, 17 Sep 2026 20:10:16 +0800 Subject: bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc sock_map_alloc() only rejects max_entries == 0 and otherwise allows any u32 value. sock_map_free() then walks the sks[] array with a signed int iterator: int i; for (i = 0; i < stab->map.max_entries; i++) struct sock **psk = &stab->sks[i]; When a SOCKMAP is created with max_entries = 0xffffffff (UINT_MAX), the allocation of 32 GiB can succeed on large-memory hosts. During free the counter reaches 0x80000000, wraps to INT_MIN, is sign-extended by movslq and turned into a ~16 GiB negative offset from stab->sks, pointing far below the allocation. The faulting access is an xchg() write in sock_map_free(). Without KASAN, the same out-of-bounds write can fault on an unmapped vmalloc page or corrupt an unrelated allocation if that vmalloc address is populated. On a KASAN kernel with CONFIG_KASAN_VMALLOC=y, the shadow check for that address hits an unmapped shadow page and oopses first: BUG: unable to handle page fault for address: fffff521b59c5a00 RIP: 0010:kasan_check_range+0x107/0x190 Call Trace: sock_map_free+0x93/0x190 map_create+0x68d/0xb30 __sys_bpf+0x21e/0x2e70 Vmcore confirmed stab->map.max_entries == 0xffffffff, stab->sks == 0xffffc911ace2d000, and the faulting address sks + (s64)INT_MIN * 8 exactly at 0xffffc90dace2d000. The same buggy path is reached on the normal close()/bpf_map_free_deferred() path whenever such a map is destroyed. sock_map_alloc() used to bound its allocation size through bpf_map_charge_init(), but the bound was dropped when rlimit-based memory accounting was removed. Reject max_entries > INT_MAX at creation time so the signed iterator in sock_map_free() never sees a value that would overflow. Triggered by syzkaller and reproduced on both a 6.6-based KASAN kernel and the upstream v7.3-rc2 kernel. Fixes: 0d2c4f964050 ("bpf: Eliminate rlimit-based memory accounting for sockmap and sockhash maps") Signed-off-by: Zhao Gongyi Signed-off-by: Alexei Starovoitov Link: https://patch.msgid.link/20260917121016.48171-1-zhaogongyi@bytedance.com --- net/core/sock_map.c | 1 + 1 file changed, 1 insertion(+) (limited to 'net') diff --git a/net/core/sock_map.c b/net/core/sock_map.c index ca49bc7f8687..38df84284328 100644 --- a/net/core/sock_map.c +++ b/net/core/sock_map.c @@ -41,6 +41,7 @@ static struct bpf_map *sock_map_alloc(union bpf_attr *attr) struct bpf_stab *stab; if (attr->max_entries == 0 || + attr->max_entries > INT_MAX || attr->key_size != 4 || (attr->value_size != sizeof(u32) && attr->value_size != sizeof(u64)) || -- cgit v1.2.3