summaryrefslogtreecommitdiff
path: root/net/ipv4
AgeCommit message (Collapse)Author
13 hoursMerge branch 'for-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next.git # Conflicts: # arch/arm64/net/bpf_jit_comp.c # arch/x86/net/bpf_jit_comp.c # mm/internal.h
13 hoursMerge branch 'main' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next.git # Conflicts: # drivers/net/ethernet/realtek/r8169_main.c # net/mac80211/ieee80211_i.h # net/mac80211/tx.c
14 hoursMerge branch 'for-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/mm/linux.git # Conflicts: # arch/arm64/kvm/mmu.c
17 hourstcp: annotate lockless access to sk->sk_errQuanye Yang
BUG: KCSAN: data-race in do_recvmmsg / mptcp_recvmsg read-write (marked) to 0xffff8880134d391c of 4 bytes by task 2619 on cpu 1: instrument_atomic_read_write include/linux/instrumented.h:113 [inline] sock_error include/net/sock.h:2565 [inline] do_recvmmsg+0x50c/0x580 net/socket.c:3049 __sys_recvmmsg net/socket.c:3144 [inline] __do_sys_recvmmsg net/socket.c:3167 [inline] __se_sys_recvmmsg net/socket.c:3160 [inline] __x64_sys_recvmmsg+0x161/0x180 net/socket.c:3160 x64_sys_call+0x19c7/0x1ca0 arch/x86/include/generated/asm/syscalls_64.h:300 do_syscall_x64 arch/x86/entry/syscall_64.c:61 [inline] do_syscall_64+0xde/0x3d0 arch/x86/entry/syscall_64.c:84 entry_SYSCALL_64_after_hwframe+0x77/0x7f read to 0xffff8880134d391c of 4 bytes by task 2620 on cpu 0: tcp_recv_should_stop include/net/tcp.h:3086 [inline] mptcp_recvmsg+0x54d/0xd50 net/mptcp/protocol.c:2466 inet_recvmsg+0x204/0x210 net/ipv4/af_inet.c:894 sock_recvmsg_nosec net/socket.c:1151 [inline] sock_recvmsg+0x11a/0x140 net/socket.c:1173 ____sys_recvmsg+0x14b/0x3c0 net/socket.c:2933 ___sys_recvmsg+0x116/0x160 net/socket.c:2975 __sys_recvmsg net/socket.c:3008 [inline] __do_sys_recvmsg net/socket.c:3014 [inline] __se_sys_recvmsg net/socket.c:3011 [inline] __x64_sys_recvmsg+0xeb/0x160 net/socket.c:3011 x64_sys_call+0x1319/0x1ca0 arch/x86/include/generated/asm/syscalls_64.h:48 do_syscall_x64 arch/x86/entry/syscall_64.c:61 [inline] do_syscall_64+0xde/0x3d0 arch/x86/entry/syscall_64.c:84 entry_SYSCALL_64_after_hwframe+0x77/0x7f value changed: 0x0000006b -> 0x00000000 Reported by Kernel Concurrency Sanitizer on: CPU: 0 UID: 0 PID: 2620 Comm: syz.2.33 Not tainted 7.2.0-g39d4f32c5d53 #76 PREEMPT(full) Hardware name: QEMU Ubuntu 26.04 PC (i440FX + PIIX, 1996), BIOS 1.17.0-debian-1.17.0-1ubuntu1 04/01/2014 do_recvmmsg() and getsockopt(SO_ERROR) call sock_error() without the socket lock. sock_error() clears sk_err with xchg(), which races with unmarked loads of the same field. KCSAN reported the unmarked peek in tcp_recv_should_stop(). Annotate that helper and the other send-side peeks with READ_ONCE(). On the no-data recv and splice paths, if (sk_err) followed by sock_error() and an unconditional break can return 0 after another thread consumes the error. Call sock_error() once and only stop when it returns a non-zero error. tcp_bpf_sendmsg() read sk_err twice; fold those unmarked loads into one READ_ONCE() and use that value as the returned errno. The field is still not consumed. MPTCP is handled in the next patch. Suggested-by: Jiayuan Chen <jiayuan.chen@linux.dev> Link: https://lore.kernel.org/netdev/8bbee583-6f21-4817-bfeb-2d60057380a3@linux.dev/ Reported-by: Matthieu Baerts (NGI0) <matttbe@kernel.org> Closes: https://github.com/multipath-tcp/mptcp_net-next/issues/632 Signed-off-by: Quanye Yang <quanyeyang@proton.me> Reviewed-by: Eric Dumazet <edumazet@kernel.org> Link: https://patch.msgid.link/20260925-mptcp-sk-err-net-v5-1-0cac04d6ea48@proton.me Signed-off-by: Paolo Abeni <pabeni@redhat.com>
26 hourstcp: remove mmap_lock fallback pathDave Hansen
Previously, the per-VMA locking could fail in the face of writers which necessitates a fallback to mmap_lock. The new vma_start_read_unlocked() will wait for writers instead of failing. Use the new helper. Wait for writers. Remove the fallback to mmap_lock. The fallback removal does not affect NOMMU case because TCP_ZEROCOPY is gated on CONFIG_MMU. This really is a nice cleanup. It removes the need to pass the lock state back and forth to find_tcp_vma(). Link: https://lore.kernel.org/20260831203056.838265-6-surenb@google.com Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com> Signed-off-by: Suren Baghdasaryan <surenb@google.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Acked-by: Lorenzo Stoakes <ljs@kernel.org> Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org> Tested-by: syzbot@syzkaller.appspotmail.com Cc: Liam R. Howlett <liam@infradead.org> Cc: Shakeel Butt <shakeel.butt@linux.dev> Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org> Cc: Arve Hjønnevåg <arve@android.com> Cc: Todd Kjos <tkjos@android.com> Cc: Christian Brauner <christian@brauner.io> Cc: Carlos Llamas <cmllamas@google.com> Cc: Alice Ryhl <aliceryhl@google.com> Cc: David S. Miller <davem@davemloft.net> Cc: David Ahern <dsahern@kernel.org> Cc: David Hildenbrand (Arm) <david@kernel.org>
30 hourstcp: preserve timestamps across receive queue collapseJason Xing
When tcp_collapse() rebuilds skbs under memory pressure, the copy process doesn't include the right tstamp and hwtstamp from the old skb. And memcpy(nskb->cb, skb->cb, ...) copies has_rxtstamp, but nskb->tstamp and hwtstamps are left at zero, so tcp_recv_timestamp() ends up emitting no cmsg at all. In net timestamping case, if such an skb happens to be the last one consumed in a recvmsg() call, the application receives no RX timestamp for that call. Fix this by copying both tstamp and hwtstamp of the last skb to the new skb, matching tcp_try_coalesce()/tcp_add_backlog(). Note that the has_rxtstamp flag can still be inherited through the cb memcpy from an skb that contributes no bytes (fully covered skb left in the ofo tree by the tcp_ooo_try_coalesce() -> coalesce_done path), so set TCP_SKB_CB(nskb)->has_rxtstamp to false which makes the new block the only place setting it. Signed-off-by: Jason Xing <kerneljasonxing@gmail.com> Reviewed-by: Eric Dumazet <edumazet@kernel.org> Link: https://patch.msgid.link/20260924152529.5689-1-kerneljasonxing@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
43 hourstcp: refresh TS.Recent for accepted old ACKsJeff Jo
A TCP packet can carry new data while acknowledging traffic in the opposite direction. With overlapping traffic in both directions, a delayed packet's acknowledgment can be older than one Linux has already accepted, even when that packet fills a gap in the received data. Linux accepts the data, but tcp_ack() takes the old_ack path and skips updating TS.Recent, the timestamp saved for outgoing acknowledgments. The reply therefore echoes an older timestamp. If the sender uses this echo to measure round-trip time after a long idle period, its estimate includes the idle time and can reduce its sending rate. Update TS.Recent in old_ack using tcp_replace_ts_recent(), before SACK processing can trigger a transmission. This reuses the existing timestamp and sequence checks, including PAWS protection against old duplicate packets. ACK validation already rejects old ACKs in SYN_RECV before this path, so no additional state check is needed. Echoing the timestamp of the packet that fills the receive gap follows RFC 7323 section 4.3. In a socket reproduction with 300 seconds idle, controlled reordering and retransmission to exercise timestamp-based RTT sampling, the sender's smoothed round-trip time was 37.5 seconds without the fix and 15.5 ms with it. Fixes: 12fb3dd9dc3c ("tcp: call tcp_replace_ts_recent() from tcp_ack()") Assisted-by: LLM sparse Signed-off-by: Jeff Jo <jeffjo@openai.com> Reviewed-by: Eric Dumazet <edumazet@google.com> Signed-off-by: David S. Miller <davem@davemloft.net>
2 daysudp: fix auto-selected port colliding with an existing SO_REUSEPORT socketJiayuan Chen
Observed with two independent servers in the same process: fd1 = socket(AF_INET, SOCK_DGRAM, 0); setsockopt(fd1, SOL_SOCKET, SO_REUSEPORT, ...); bind(fd1, port 0); /* got 40000 */ fd2 = socket(AF_INET, SOCK_DGRAM, 0); setsockopt(fd2, SOL_SOCKET, SO_REUSEPORT, ...); bind(fd2, port 0); /* got 40000 as well */ Both sockets end up on the same port and join the same reuseport group, so each of them takes part of the other's datagrams. The same port sharing happens with SO_REUSEADDR, without the reuseport group. udp_lib_lport_inuse() keeps the reuse rules when it scans for a free port: a socket with SO_REUSEADDR set, or one with SO_REUSEPORT and the same uid, is not a conflict, so its port is never marked in the bitmap and the scan can hand it out again. Those rules only make sense when the user asks for a specific port. bind(0) wants a free port and has no way to know whose port it lands on. TCP already does this: inet_csk_find_open_port() passes relax=false and reuseport_ok=false, so the scan treats every port in use as a conflict whatever options the sockets have. Commit aacd9289af8b ("tcp: bind() use stronger condition for bind_conflict") did it for SO_REUSEADDR and commit 0643ee4fd1b7 ("inet: Fix get port to handle zero port number with soreuseport set") for SO_REUSEPORT. Do the same for UDP: ignore both options during a scan, keep them for an explicit port. This changes bind(0) for SO_REUSEADDR sockets once the port range is full: it used to share a port and now fails with EADDRINUSE, like TCP. Sharing a port on purpose when the range is full is a feature, TCP has net.ipv4.ip_autobind_reuse for it, off by default. UDP can get the same knob later, this patch only fixes the scan. Fixes: ba418fa357a7 ("soreuseport: UDP/IPv4 implementation") Cc: stable+noautosel@kernel.org # bind() behaviour change Signed-off-by: Jiayuan Chen <jiayuan.chen@linux.dev> Link: https://patch.msgid.link/20260928023145.301855-1-jiayuan.chen@linux.dev Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysnet: gso: limit recursive IP-in-IP segmentationZihan Xi
IP-in-IP GSO can re-enter inet_gso_segment() or ipv6_gso_segment() for each nested IP header. encap_level tracks header bytes, not callback depth, so a deep chain can exhaust the kernel stack. Making inet_gso_segment() stackable introduced unbounded IPv4 nesting; IPIP GSO/TSO later made the path reachable. The IPv6 stackable path was introduced separately and uses the same guard. Count IPv4 and IPv6 GSO handler entries in skb_gso_cb, initialized once per top-level GSO operation and preserved across GRE/UDP context changes. Use the existing IP_TUNNEL_RECURSION_LIMIT for both handlers. The first five entries pass, and the sixth returns -EINVAL before dispatching another GSO callback. Fixes: 3347c9602955 ("ipv4: gso: make inet_gso_segment() stackable") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Closes: https://lore.kernel.org/all/cover.1790157745.git.zihanx@nebusec.ai/ Assisted-by: LLM Co-developed-by: Luxing Yin <root@tr0jan.top> Signed-off-by: Luxing Yin <root@tr0jan.top> Signed-off-by: Zihan Xi <zihanx@nebusec.ai> Reviewed-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/20260924051521.32568-2-zihanx@nebusec.ai Signed-off-by: Paolo Abeni <pabeni@redhat.com>
4 daysnetfilter: fix several typos in commentsHemanth Selam
Fix typos reported by scripts/checkpatch.pl using the misspelling list in scripts/spelling.txt. Only touches comments, no code changes. Assisted-by: Cursor:claude-opus-5 Signed-off-by: Hemanth Selam <hemanth.selam@gmail.com> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
7 daysMerge git://git.kernel.org/pub/scm/linux/kernel/git/netdev/netJakub Kicinski
Cross-merge networking fixes after downstream PR (net-7.3-rc5). Conflicts: drivers/net/mdio/mdio-realtek-rtl9300.c 89a8a1eef2d44 ("net: mdio: realtek-rtl9300: fix RTL931x C22 extended page selection") cc3cb8db1eef9 ("net: mdio: realtek-rtl9300: Add page tracking") https://lore.kernel.org/arUZOqy73bE2pp0w@sirena.org.uk net/8021q/vlan_dev.c cd5dd68267c4 ("vlan: ensure sufficient headroom in vlan_dev_hard_header()") ca6ff8dd70eb ("vlan: annotate data-races in vlan_dev_priv fields") Adjacent changes: net/ipv6/ip6_gre.c dd47bcf279f1 ("ip6_gre: Call ip6erspan_tunnel_unlink_md() in ip6erspan_changelink().") cce829e2aa1d ("ip6_gre: Protect ip6gre_net.tunnels[][] with mutex.") drivers/net/ethernet/meta/fbnic/fbnic_txrx.c b5d9e9d4d0c1 ("eth: fbnic: use the Rx queue napi pointer to find the napi vector") c0aca269ec07 ("eth: fbnic: Make Rx completion coalescing configurable") drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c d68acbf93531 ("net: stmmac: selftests: Support running selftests on DSA conduits") 85ca3292d7a3 ("net: stmmac: Remove ARP offload code") drivers/net/ethernet/wangxun/libwx/wx_hw.c 3173cba11701 ("net: libwx: fix races in Tx timestamp handling") 7042c8c193e5 ("net: libwx: rename wx_pf_flags to wx_flags") Signed-off-by: Jakub Kicinski <kuba@kernel.org>
7 daysMerge tag 'net-7.3-rc5' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net Pull networking fixes from Jakub Kicinski: "Including fixes from Bluetooth, NFC and Netfilter. Every week in this release is record-setting for number of posted patches. It doesn't seem like we're creating any regressions with all these fixes, three 'Fixes' tags here point to 7.2 commits but none are true regression fixes. We're trying to keep the count down, nonetheless. Previous releases - regressions: - net: don't require the hwtstamp NDOs when a PHY provides timestamping - ipv6: fix dst leak for uncached routes - vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE packets Previous releases - always broken: - packet: use ubuf_info completion for TX_RING packets - arp: terminate device name before lookup - ipv6: do not let ipv6_find_hdr() return an offset past the packet end - udp: remove a disconnected socket from the 4-tuple hash table - sctp: discard the rest of the packet on a stale-cookie error - eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged eswitch" [ And lots of other random network driver fixes ] * tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (189 commits) tcp: prevent collapsing skbs across boundary in rtx queue vlan: ensure sufficient headroom in vlan_dev_hard_header() net/sched: sch_teql: fix shadowed err in __teql_resolve() bridge: check llc_mac_hdr_init() return value in br_send_bpdu() llc: fix skb UAF and leaks on llc_mac_hdr_init() failure llc: reserve device headroom for allocated frames gve: DQO: reject TSO packets with an out of range MSS gve: fix TX drop when GSO MSS is too small for hw gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO net: flush skb_defer_nodes in dev_cpu_dead() net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails af_packet: fix integer overflow in prb_calc_retire_blk_tmo() tipc: Fix a data race on mon->peer_cnt in mon_timeout() net: phy: intel-xway: workaround 100BASE-TX Link-Up issue net/smc: fix UAF on lgr list traversal in smcr_port_err() net/rds: size a connection's path set by the transport it ends up with nfp: hold IPsec RX state under the XArray lock net: ena: fix MMIO read buffer leak on probe failure net: ena: fix PHC cleanup on probe failure net/sched: act_ct: fix helper UAF due to extensions realloc ...
7 daystcp: fix use-after-free of retransmit_skb_hint in tcp_send_synack()Yilin Zhang
When tcp_send_synack() replaces the cloned SYN skb at the head of the retransmit queue with a copy, it frees the original with tcp_rtx_queue_unlink_and_free() and only repairs tp->highest_sack. tp->retransmit_skb_hint keeps pointing at the freed skbuff_fclone_cache object. The dangling hint is read in tcp_verify_retransmit_hint() and used as the root of the rbtree walk in tcp_xmit_retransmit_queue(). An unprivileged TFO client (sendmsg(MSG_FASTOPEN)) can arm the hint with an attacker-supplied ICMP fragmentation-needed message, after which a simultaneous open frees the armed SYN skb: BUG: KASAN: slab-use-after-free in tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316) Read of size 4 at addr ffff88800604d928 by task swapper/1/0 Call Trace: tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316) tcp_simple_retransmit (net/ipv4/tcp_input.c:3158) tcp_v4_err (net/ipv4/tcp_ipv4.c:587) Sync the hint to the copy. Fixes: c31b70c9968f ("tcp: Add logic to check for SYN w/ data in tcp_simple_retransmit") Reported-by: Kimi Security Team <bug-report@moonshot.ai> Tested-by: Weiming Shi <shiweiming@moonshot.ai> Signed-off-by: Yilin Zhang <yilinzhang@moonshot.ai> Reviewed-by: Eric Dumazet <edumazet@google.com> Link: https://patch.msgid.link/8a9dff4063a2745653b7e88ceb745d75efa16e68.1790224474.git.yilinzhang@moonshot.ai Signed-off-by: Jakub Kicinski <kuba@kernel.org>
7 daysMerge git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf 7.3-rc4Alexei Starovoitov
Cross-merge BPF and other fixes after downstream PR. Conflicts: kernel/bpf/helpers.c tools/testing/selftests/bpf/prog_tests/cb_refs.c tools/testing/selftests/bpf/prog_tests/verifier.c Signed-off-by: Alexei Starovoitov <ast@kernel.org>
7 daysMerge tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpfLinus Torvalds
Pull bpf fixes from Alexei Starovoitov: - Fix bpf_skb_change_tail() to drop the checksum offload instead of rejecting the trim of CHECKSUM_PARTIAL skbs (Daniel Borkmann) - Add KF_PERFMON kfunc flag and require CAP_PERFMON for kfuncs that read arbitrary memory and for untrusted read-only memory reads (Daniel Borkmann) - Clear scalar delta on narrowing stack spill (Daniel Borkmann) - Set up the frame pointer for the exception callback in arm64 JIT, and zero-fill other CPUs when BPF_F_CPU update creates a per-cpu hash element (Donggeun Yoo) - Various fixes (Emil Tsalapatis): - Fix bounds check underflow for skb-backed dynptrs - Fix rx_queue_mapping context access code generation in bpf_sock - Reject packet pointer arguments to subprogs that may mutate the packet - Reject ALU instructions that see arena and non-arena operands on different code paths - Fix copied_seq double-counting on sockmap self-redirect (Geliang Tang) - Fix divide-by-zero in btf_struct_walk() on a flexible array of zero-sized elements, fix out-of-bounds read of rtt_min in sock_ops (Jiayuan Chen) - Fix bpf_sock_destroy() out-of-bounds read of sk_protocol on TIME_WAIT and request socks, and sleeping under RCU when destroying a listener with pending children (Jiayuan Chen) - Fix JEQ/JNE with immediate operand in MIPS32 JIT and missing zero extension of BSWAP 16/32 in MIPS64 JIT (Johan Almbladh) - Avoid soft lockup in htab lookup[_and_delete] batch operations on large maps (Jose Fernandez) - Various fixes (Kumar Kartikeya Dwivedi): - Verify global subprogs in each sleepability context they are called from - Make post-verification instruction rewrites killable - Preserve packet pointer displacement in regsafe() - Apply CO-RE relocations before subprogram validation, restrict CO-RE poisoning to relocatable instructions, and reject truncated ldimm64 CO-RE relocations in libbpf - Assign lock identity to callback map values - Compare stack frames in regs_exact() - Bound ownership depth through local kptrs and graph roots - Fix u32 overflow in map batch operations when the map size exceeds 4GB (Masoud Aghasi) - Fix UAF in bpf memalloc due to concurrent consumption of ttrace lists in alloc_bulk() (Pu Lehui) - Allow gotox as the terminal instruction of a program or a subprogram (Siddharth Chintamaneni) - Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL, skip unsettled links in link iterator, and reject dev-bound-only programs on other devices (Weiming Shi) - Reject non-negative stack offsets in stack_slot_obj_get_spi() (Xu Yunxiang) - Check params size before reading reserved fields in bpf_crypto_ctx_create() (Yuqi Xu) - Reject max_entries > INT_MAX in sock_map_alloc() (Zhao Gongyi) - Use a 32-bit compare in xsk_map_gen_lookup() (Zhiling Zou) - Use kvfree() in xdp_test_run_teardown() (Zhixing Chen) * tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: (58 commits) selftests/bpf: Test per-cpu initialization of a BPF_F_CPU created element bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element bpf: Fix BSWAP 32 and 16 on MIPS64 bpf: Fix immediate JMP JEQ/JNE on MIPS32 bpf: Reject dev-bound-only programs on other devices bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc selftests/bpf: Test for mixed arena/nonarena code paths bpf: Prevent variable arena/non-arena register contents selftests/bpf: Test rejection of pkt args to mutating subprogs bpf: Reject pkt arguments in mutating subprogs selftests/bpf: Add selftests for rx_queue_mapping context access bpf: Fix bpf_sock context code generation selftests/bpf: Test dynptr slices past end of skb bpf: Fix bounds check for skb-backed dynptrs selftests/bpf: Reject iterator destruction through fp+0 bpf: Reject non-negative offsets in stack_slot_obj_get_spi() bpf: Check params size before reading reserved fields selftests/bpf: Check local object ownership depth bpf: Bound ownership depth through local kptrs and graph roots selftests/bpf: Cover frame changes in bounded loops ...
8 daysnet: ipconfig: bound DHCP option constructionYuqi Xu
ic_dhcp_init_options() appends the hostname (option 12), vendor-class (option 60) and client-ID (option 61) options into the fixed 312-byte bootp_pkt.exten[] buffer. Only the client-ID branch checked the remaining space; the hostname and vendor-class writes were unbounded. A 64-byte hostname together with the maximum 252-byte dhcpclass= identifier needs 18 + (2 + 64) + (2 + 252) = 338 of the 312 available bytes even before the terminating END marker, so the vendor-class memcpy runs past the end of exten[]. With CONFIG_FORTIFY_SOURCE this is reported as a field-spanning write and, when the kernel is booted with panic_on_warn=1, aborts boot with a panic. Route the optional options through a common helper that makes sure the option, its 2-byte header and the END marker all fit and drops an option that would not. Configurations with short options keep sending exactly the same bytes as before. Fixes: 130c0f47fdf9 ("ipconfig: send host-name in DHCP requests") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Assisted-by: LLM Signed-off-by: Yuqi Xu <xuyuqiabc@gmail.com> Reviewed-by: Ren Wei <weir@nebusec.ai> Reviewed-by: Simon Horman <horms@kernel.org> Link: https://patch.msgid.link/7808dfbfa2162dfd0b19f59aff5742d6e0db2abb.1789798023.git.xuyuqiabc@gmail.com Signed-off-by: Paolo Abeni <pabeni@redhat.com>
8 daysip_gre: Reject enabling collect metadata through changelinkXuanqiang Luo
ipgre_netlink_parms() can enable collect_md on an existing GRE, GRETAP or ERSPAN device. Unlike newlink, changelink does not enforce metadata tunnel uniqueness. Converting a non-metadata device can therefore replace the metadata receive entry for another device of the same type in the same netns. Deleting either device then clears the shared entry, breaking metadata receive lookup for the surviving device. If parameter validation fails after collect_md is set, deleting the modified device can also clear an entry it never owned. Reject enabling metadata mode in both changelink callbacks before any encapsulation or tunnel parameters are modified. Allow requests that repeat the metadata attribute on an existing metadata device. Fixes: 2e15ea390e6f ("ip_gre: Add support to collect tunnel metadata.") Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn> Link: https://patch.msgid.link/20260921031859.9283-1-xuanqiang.luo@linux.dev Signed-off-by: Jakub Kicinski <kuba@kernel.org>
8 daysnet: arp: terminate device name before lookupZijie Huang
The ARP ioctl copies a user-provided struct arpreq into a stack object. Its arp_dev field may contain IFNAMSIZ bytes without a NUL terminator. Such input is passed to dev_get_by_name_rcu() or __dev_get_by_name(), where strcmp() can read past the end of the stack object when a matching alternative interface name exists. Terminate the field before the lookup to prevent the out-of-bounds read. Fixes: 36fbf1e52bd3 ("net: rtnetlink: add linkprop commands to add and delete alternative ifnames") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Signed-off-by: Zijie Huang <milkory@outlook.com> Signed-off-by: Ren Wei <weir@nebusec.ai> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Link: https://patch.msgid.link/fabf02a70787d17299e4b3153eadffaf20d154b3.1789910973.git.milkory@outlook.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
8 daysfou: reject omitted FOU_ATTR_IPPROTO on FOU_ENCAP_DIRECTHui Peng
Commit 7a9bc9e3f423 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.") added NLA_POLICY_MIN(NLA_U8, 1) to fou_nl_policy[FOU_ATTR_IPPROTO], which rejects an explicitly supplied FOU_ATTR_IPPROTO == 0 attribute with -ERANGE. However, FOU_ATTR_IPPROTO is an optional netlink attribute. When a user sends FOU_CMD_ADD with FOU_ATTR_TYPE set to FOU_ENCAP_DIRECT and omits FOU_ATTR_IPPROTO entirely, nla_policy validation succeeds and parse_nl_config() leaves cfg->protocol as 0 (from memset(cfg, 0, sizeof(*cfg))). fou_create() then creates a FOU_ENCAP_DIRECT socket with fou->protocol == 0. In fou_udp_recv(), returning -fou->protocol to udp_queue_rcv_one_skb() triggers IP protocol resubmission when fou->protocol > 0, whereas returning 0 tells the UDP tunnel layer that the skb was consumed without freeing it. When fou->protocol == 0, every packet received on the socket returns 0 from fou_udp_recv() and leaks the sk_buff. Reject FOU_ENCAP_DIRECT when !cfg->protocol in fou_create() so that creating a direct encapsulation port without FOU_ATTR_IPPROTO fails with -EINVAL while leaving FOU_CMD_DEL and FOU_CMD_GET (which share parse_nl_config()) unaffected. Tested in QEMU against Linux 7.3.0-rc3 by sending a FOU_CMD_ADD Generic Netlink request with FOU_ATTR_PORT = 5555 and FOU_ATTR_TYPE = FOU_ENCAP_DIRECT while omitting FOU_ATTR_IPPROTO. On the unfixed kernel, FOU_CMD_ADD succeeds (err = 0), FOU_CMD_GET reports fou->type = 1 and fou->protocol = 0, and sending 4000 UDP packets to 127.0.0.1:5555 leaks all 4000 sk_buffs (SUnreclaim in /proc/meminfo grows from 41456 kB to 59008 kB, +17552 kB); with this patch applied, FOU_CMD_ADD is rejected with -EINVAL (-22). Fixes: 23461551c006 ("fou: Support for foo-over-udp RX path") Fixes: 7a9bc9e3f423 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.") Cc: stable@vger.kernel.org Signed-off-by: Hui Peng <benquike@gmail.com> Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn> Link: https://patch.msgid.link/20260921045920.1613098-1-benquike@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
9 daysbpf: Drop duplicate check_app_limited in tcp_bpf_pushGeliang Tang
Commit c5c37af6ecad9 ("tcp: Convert do_tcp_sendpages() to use MSG_SPLICE_PAGES") moved tcp_rate_check_app_limited() inside do_tcp_sendpages(), turning it into a wrapper around tcp_sendmsg_locked(). Later, commit ebf2e8860eea ("tcp_bpf: Inline do_tcp_sendpages as it's now a wrapper around tcp_sendmsg") inlined the wrapper in tcp_bpf_push() with direct tcp_sendmsg_locked() calls, which perform the check on every path that queues data, but kept the outer tcp_rate_check_app_limited() that was previously needed to cover do_tcp_sendpages(). The outer call is now redundant. The site changed here, tcp_bpf_push(), holds the socket lock and invokes tcp_sendmsg_locked() on every iteration. The early-return paths in tcp_sendmsg_locked() that skip tcp_rate_check_app_limited() - the MSG_ZEROCOPY allocation failure and MSG_FASTOPEN branches - return without queueing any MSG_SPLICE_PAGES data, so there is no functional consequence from omitting the outer check. A potential benefit of this change is that it facilitates future reuse of tcp_bpf_push() for sockmap support in protocols beyond TCP, such as MPTCP. Since tcp_rate_check_app_limited() is TCP-specific while sendmsg_locked() is a generic interface in struct proto_ops, this change allows us to switch to different protocols via sk->sk_socket->ops->sendmsg_locked() without carrying protocol-specific assumptions. Signed-off-by: Geliang Tang <tanggeliang@kylinos.cn> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev> Link: https://patch.msgid.link/1f7dc605b16fc0590f7cbf5a27d57271926c01ce.1790074764.git.tanggeliang@kylinos.cn
9 daystcp: remove dead code in tcp_rcv_state_process()Eric Dumazet
Since commit 3d501dd326fb ("tcp: do not accept ACK of bytes we never sent"), tcp_ack() bounds the acceptable old ACK window by min(tp->max_window, tp->bytes_acked). When sk->sk_state == TCP_SYN_RECV, tp->bytes_acked is always 0, so any segment with before(ack, prior_snd_una) immediately returns -SKB_DROP_REASON_TCP_TOO_OLD_ACK and never reaches the old_ack label (which returns 0). Therefore, tcp_ack() can only return 0 in closing states (where old ACKs are accepted), and can never return 0 in TCP_SYN_RECV. Simplify the tcp_ack() return value check in tcp_rcv_state_process() to only check for negative return values and remove the unreachable !reason branch. Signed-off-by: Eric Dumazet <edumazet@google.com> Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn> Link: https://patch.msgid.link/20260922010627.2291980-1-edumazet@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
10 daysudp: remove a disconnected socket from the 4-tuple hash tableShardul Bankar
A UDP socket bound to a specific address and port keeps its entry in the 4-tuple hash table after it is disconnected: sk binds to 127.0.0.1:21001 sk connects to 127.0.0.2:20001 // filed in the 4-tuple table sk disconnects, connect(AF_UNSPEC) // still filed, peer now 0.0.0.0:0 __udp_disconnect() takes a socket out of that table only as a side effect of ->rehash() or ->unhash(), and it skips ->rehash() when SOCK_BINDADDR_LOCK is set and ->unhash() when SOCK_BINDPORT_LOCK is set. commit 6996a2d2d0a6 ("udp: Unhash auto-bound connected sk from 4-tuple hash table when disconnected.") fixed the same end state for a wildcard-bound socket, by a path this one does not take. The entry is counted whether or not anything hits it. hash4_cnt on the hash2 slot stays raised for as long as the socket lives, so udp_has_hash4() keeps sending every packet for that address and port through the 4-tuple lookup first. On IPv6 it can also be hit. __udp_disconnect() does not clear sk_v6_daddr, so udp_v6_rehash() files the entry under the peer the socket was connected to with a zero dport, and inet6_match() compares that same field: a datagram from the former peer with a zero source port matches, and source port zero is accepted on receive. On IPv4 the peer is cleared, so a match would need a zero source address as well, which the routing layer rejects as martian. The stale sk_v6_daddr is a separate defect, not addressed here; removing the entry closes this path either way. The entry can also be relocated. __udp_disconnect() clears sk_bound_dev_if, so a subsequent SO_BINDTODEVICE calls ->rehash(), and because the receive address is still specific udp_lib_rehash() moves the entry instead of removing it, into the bucket that (rcv_saddr, num, 0, 0) hashes to -- a pure function of the address and port, so every socket reaching this state on one address and port collects in one bucket. The bucket cannot be chosen from outside, as udp_ehashfn() is seeded with a per-boot secret. This last one became reachable only with commit 644f9108f3a5 ("udp: Make rehash4 independent in udp_lib_rehash()"), which moved the hash4 handling out of a branch a disconnected socket does not take; the stale entry itself dates from the commit in Fixes. Take the socket out of the table before __udp_disconnect() runs, while it still matches how it was filed. This also reaches the wildcard case ahead of udp_lib_rehash()'s udp_unhash4() branch, leaving that branch unreachable from udp_disconnect(); removing it belongs in net-next. udp_disconnect() and udp_abort() are the only UDP entries into __udp_disconnect(), which is shared with raw, ping and l2tp sockets that are not struct udp_sock: ping_prot.obj_size is sizeof(struct inet_sock), so udp_hashed4() on one would read past the allocation. Fixes: 78c91ae2c6de ("ipv4/udp: Add 4-tuple hash for connected socket") Assisted-by: LLM Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com> Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com> Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-2-718891af0d7a@mpiricsoftware.com Signed-off-by: Paolo Abeni <pabeni@redhat.com>
10 daysudp: relocate a connected socket in the 4-tuple hash table on re-connectShardul Bankar
A connected UDP socket that connects again to a different peer is not re-filed in the 4-tuple hash table: sk binds to 127.0.0.1:21001 sk connects to 127.0.0.2:20001 // filed under hash(sk, peer1) sk connects to 127.0.0.3:20002 // still filed under hash(sk, peer1) packet from 127.0.0.3:20002 // hash(sk, peer2) misses, so the // lookup falls back to scoring the // hash2 chain for this address // and port udp_lib_hash4() returns early when the socket is already hashed, assuming ->rehash() relocates it. ->rehash() runs from __ip{4,6}_datagram_connect() only while the receive address is unset, which a second connect never is: the first connect assigns it, whether the socket was bound to a specific address or to the wildcard. commit 644f9108f3a5 ("udp: Make rehash4 independent in udp_lib_rehash()") added that early return and named connect(AF_UNSPEC) as the way around it. That workaround does not help a socket with both SOCK_BINDADDR_LOCK and SOCK_BINDPORT_LOCK set, because __udp_disconnect() skips ->rehash() for the first and ->unhash() for the second. Delivery is correct either way. Relocate the socket when the hash it is filed under differs from the one requested, which is what commit 78c91ae2c6de ("ipv4/udp: Add 4-tuple hash for connected socket") did before the early return became unconditional. It is done here under hslot->lock, which that version did not take, to match udp_lib_rehash() and udp_lib_unhash(). hslot2 is unchanged, so hash4_cnt needs no adjustment, as in udp_lib_rehash(). A first connect is unaffected, and IPv6 shares the code. With 500 sockets on the port, a re-connected socket measured 522,553 pps without this change and 2,055,078 with it. The UDP side was noted as remaining work in [1]. Link: https://lore.kernel.org/netdev/apnHqmYZQ4yzOP4N@v4bel/ [1] Fixes: 644f9108f3a5 ("udp: Make rehash4 independent in udp_lib_rehash()") Assisted-by: LLM Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com> Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com> Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-1-718891af0d7a@mpiricsoftware.com Signed-off-by: Paolo Abeni <pabeni@redhat.com>
13 daysbpf: tcp: Support parse/len/write header option hooks in bpf_tcp_opsAmery Hung
Add the TCP header option callbacks to the bpf_tcp_ops struct_ops type: parse_hdr - parse the options of an incoming skb on an established connection hdr_opt_len - reserve space in the TCP header for bpf options write_hdr_opt - write the reserved bpf options These mirror the BPF_SOCK_OPS_PARSE_HDR_OPT_CB, _HDR_OPT_LEN_CB and _WRITE_HDR_OPT_CB legacy sockops callbacks, but are exposed as struct_ops members so a program can implement them with normal function signatures and per-member helper sets. The reserved header window is shared between the legacy sockops and bpf_tcp_ops paths. tcp_{syn,synack,established}_options() first run the legacy BPF_SOCK_OPS_HDR_OPT_LEN_CB and then call hdr_opt_len, so both sources accumulate into opts->bpf_opt_len; at write time the legacy options are emitted first and bpf_tcp_ops writes after them. API design bpf_tcp_ops overloads the sock_ops header-option helpers rather than introducing a new API: bpf_reserve_hdr_opt(), bpf_store_hdr_opt() and bpf_load_hdr_opt() are exposed per-member (reserve for hdr_opt_len, store/load for write_hdr_opt, load for parse_hdr) and share the existing kernel option-walking core via _bpf_sock_ops{store,load}hdr_opt(), with the bpf_tcp_ops wrappers synthesizing a temporary bpf_sock_ops_kern from the program ctx. This keeps a port from the legacy BPF_SOCK_OPS*_HDR_OPT_CB callbacks mechanical (same helper calls) and adds no new UAPI helper/kfunc surface. An alternative considered was to drop the option helpers entirely: have hdr_opt_len reserve space purely through its return value, and introduce a dedicated TCP-header-option dynptr used for both reading and writing. That is a cleaner, more self-contained interface, but it is a larger change and does not reuse the legacy helpers, making a port from sockops less mechanical. It can be pursued as a follow-up; the helper-based interface here keeps this series focused on moving the hooks to struct_ops. The hdr_opt_len fast path in tcp_established_options() is gated by cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS). Note this is a global, per-attach-type static branch: it is enabled whenever any bpf_tcp_ops is attached, even one that does not implement hdr_opt_len or that is attached to a different cgroup. In those cases the block still runs but bpf_tcp_ops_hdr_opt_len() no-ops via the per-member check in the dispatch macro. A per-member/per-cgroup gate could be added later if the extra fast-path work proves measurable. Signed-off-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com> Link: https://patch.msgid.link/20260917200542.3689605-13-ameryhung@gmail.com
13 daysbpf: tcp: Support selected sock_ops callbacks as struct_opsAmery Hung
In LSFMMBPF 2025, I have talked about moving the BPF_PROG_TYPE_SOCK_OPS to a struct_ops interface [1]. The BPF_SOCK_OPS_*_CB enum interface has grown over time as new TCP callback points were added. A BPF_PROG_TYPE_SOCK_OPS program now commonly needs a large switch on sock_ops->op, and the shared bpf_sock_ops_kern context has become harder to extend because different callbacks have different locking, argument, skb, and helper requirements. The existing 'union { u32 args[4]; u32 replylong[4]; }' is also not reliable in passing args to bpf prog when there are multiple progs attached to a cgroup. The above has already been solved in struct_ops. Add a TCP-specific struct_ops type, bpf_tcp_ops, and support attaching it to cgroups. This allows each callback have its own func signature and allows the verifier to select kfuncs/helpers based on the specific struct_ops member being implemented. This patch wires up the following existing sock_ops callbacks: - BPF_SOCK_OPS_TIMEOUT_INIT - BPF_SOCK_OPS_RWND_INIT - BPF_SOCK_OPS_RTT_CB - BPF_SOCK_OPS_STATE_CB - BPF_SOCK_OPS_RETRANS_CB - BPF_SOCK_OPS_TCP_CONNECT_CB - BPF_SOCK_OPS_TCP_LISTEN_CB - BPF_SOCK_OPS_RTO_CB - BPF_SOCK_OPS_ACTIVE_ESTABLISHED_CB - BPF_SOCK_OPS_PASSIVE_ESTABLISHED_CB BASE_RTT is ignored as it is not particularly useful. NEEDS_ECN should be done in bpf-tcp-cc instead. The tstamp ones should be a separate struct_ops (e.g. "bpf_sock_ops") that can work in both TCP and UDP. timeout_init and rwnd_init could have a request_sock pointer. This patch tries a different API and directly passes the request_sock pointer as an arg. Two other approaches were considered before settling on having bpf_get_retval() read the dispatcher's run_ctx via saved_run_ctx. The first was to inherit the retval in the trampoline itself: add a helper in the four __bpf_prog_enter*() paths that, for struct_ops programs, copies the chained value from the caller's run_ctx (now saved_run_ctx) into the program's own run_ctx. It works but puts a per-enter program-type check on the generic trampoline fast path, taxing all fentry/fexit/lsm callers for a cgroup-struct_ops-only feature. The second was to do that same inherit only for the int-returning members via a gen_prologue that emits a hidden kfunc at the start of timeout_init/rwnd_init; this keeps the cost off the generic path and scoped to bpf_tcp_ops, but needs a kfunc + BTF_ID + prologue-emission machinery. The chosen approach avoids both: it touches neither the trampoline nor the program, since saved_run_ctx already points at the dispatcher's run_ctx that carries the value. [1], page 13: https://drive.google.com/file/d/1wjKZth6T0llLJ_ONPAL_6Q_jbxbAjByp/view?usp=sharing Signed-off-by: Martin KaFai Lau <martin.lau@kernel.org> Signed-off-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com> Link: https://patch.msgid.link/20260917200542.3689605-12-ameryhung@gmail.com
13 daysbpf: Make struct_ops tasks_rcu grace period optionalMartin KaFai Lau
bpf_struct_ops_map_free() currently waits for both a regular RCU grace period and a tasks RCU grace period for every struct_ops map through synchronize_rcu_mult(call_rcu, call_rcu_tasks). A regular RCU grace period is still required for all struct_ops maps because the struct_ops trampoline ksyms requires a rcu grace period (take a look at the list_del_rcu in __bpf_ksym_del). Add a map_free_pre_rcu() callback so the struct_ops map can remove ksyms before bpf_map_put() wait for the regular rcu grace period. The tasks RCU grace period is only needed by tcp_congestion_ops. Add free_after_tasks_rcu_gp only to struct bpf_struct_ops instead of the bpf_map. When CONFIG_TASKS_RCU=n, synchronize_rcu_tasks() is the same as synchronize_rcu(). Since all struct_ops maps now complete a regular RCU grace period before bpf_struct_ops_map_free() runs, skip the extra synchronize_rcu_tasks() call in this case. This cleanup prepares for a later patch that needs to support free_after_mult_rcu_gp. Signed-off-by: Martin KaFai Lau <martin.lau@kernel.org> Signed-off-by: Amery Hung <ameryhung@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com> Reviewed-by: Eduard Zingerman <eddyz87@gmail.com> Link: https://patch.msgid.link/20260917200542.3689605-3-ameryhung@gmail.com
13 daysnet: allow IFLA_INET_CONF messages when NLA_F_NESTED unsetQuentin Armitage
Commit fa8fca88714c ("ipv4: validate IPV4_DEVCONF attributes properly") added validation of IFLA_INET_CONF attributes, and in the process changed the call of nla_for_each_nested() to nla_parse_nested(). A side effect of this change is that the IFLA_INET_CONF option is now tested for NLA_F_NESTED being set, and fails if it is not. Prior to the commit there was no check of NLA_F_NESTED. Change nla_parse_nested() to nla_parse(). This restores the previous functionality of not checking NLA_F_NESTED, thereby allowing code that (incorrectly) doesn't set NLA_F_NESTED to continue to work. This issue was identified because keepalived started logging errors when it was configuring macvlans that it created. Fixes: fa8fca88714c ("ipv4: validate IPV4_DEVCONF attributes properly") Signed-off-by: Quentin Armitage <quentin@armitage.org.uk> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Link: https://patch.msgid.link/20260915213320.1527029-2-quentin@armitage.org.uk Signed-off-by: Jakub Kicinski <kuba@kernel.org>
13 daystcp: Set unhashed_state in inet_twsk_hashdance_schedule().Kuniyuki Iwashima
inet_unhash() sets inet_csk(sk)->unhashed_state only when the socket is hashed because tcp_set_state(sk, TCP_CLOSE) could be called multiple times, e.g. tcp_abort() calls it directly and tcp_done_with_error(). However, inet_twsk_hashdance_schedule() also unhashes a socket when replacing it with twsk, allowing the socket to bypass checks for inet_csk(sk)->unhashed_state. Let's update inet_csk(sk)->unhashed_state there as well. Fixes: 8cc3aef0cb19 ("tcp: Do not allow buggy transitions between ehash and lhash2.") Reported-by: Daniel Zahka <daniel.zahka@gmail.com> Closes: https://lore.kernel.org/netdev/DLHLRA8GVI5B.2Q1IRQG5BVJNZ@gmail.com/ Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com> Reviewed-by: Daniel Zahka <daniel.zahka@gmail.com> Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org> Link: https://patch.msgid.link/20260917191554.1600494-1-kuniyu@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
13 daysipv4: igmp: convert ip_mc_msfget() to sockopt_tBreno Leitao
IP_MSFILTER reads its reply through ip_mc_msfget(), reached from do_ip_getsockopt() and from nowhere else. Convert it, and build the sockopt_t at the call site for as long as the caller still carries a sockptr_t pair. optlen here only has to cover the header, and the real reply size comes from the imsf_numsrc field inside it. This is nasty, but userspace relies on it, so sockopt_expand_out() preserves the same mechanism: it grows optval only for a user address, and assumes the caller left room for the size its own header asked for. The *optlen store moves out of ip_mc_msfget() and into the call site, guarded by !err so the -EINVAL, -ENODEV and -EADDRNOTAVAIL returns still leave the caller's optlen word untouched. The source list also moves from copy_to_sockptr_offset() to a sequential copy_to_iter(). IP_MSFILTER_SIZE(0) and offsetof(struct ip_msfilter, imsf_slist_flex) are both 16, so the bytes land where they did. Signed-off-by: Breno Leitao <leitao@debian.org> Acked-by: Stanislav Fomichev <sdf@fomichev.me> Link: https://patch.msgid.link/20260914-getsockopt_phase6-v2-2-e48befc9602e@debian.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17ipv4: fib: fix data-race and stale genid check around nh->nh_saddrLinkui Xiao
fib_select_multipath() compares nexthop_nh->nh_saddr against the flow source address with no lock held, while fib_info_update_nhc_saddr() stores a new value from another CPU as soon as the preferred source address of the egress device changes. Commit 195374d89368 ("ipv4: fib: annotate races around nh->nh_saddr_genid and nh->nh_saddr") added WRITE_ONCE() on the store side and READ_ONCE() in fib_result_prefsrc() after syzbot reported BUG: KCSAN: data-race in fib_select_path / fib_select_path but it only covered that reader. fib_select_multipath(), reached from fib_select_path(), is a second lockless reader of nh->nh_saddr and was left bare. Moreover, nh_saddr is only meaningful when nh_saddr_genid matches dev_addr_genid, as established by commit 436c3b66ec98 ("ipv4: Invalidate nexthop cache nh_saddr more correctly."). fib_select_multipath() skips that validation, so it can score a nexthop using a stale source address and skew the ECMP selection. Annotate both reads with READ_ONCE() and refresh the cached source address via fib_info_update_nhc_saddr() when the genid does not match, mirroring fib_result_prefsrc(). Fixes: 32607a332cfe ("ipv4: prefer multipath nexthop that matches source address") Signed-off-by: Linkui Xiao <xiaolinkui@kylinos.cn> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Reviewed-by: Eric Dumazet <edumazet@google.com> Link: https://patch.msgid.link/20260916125316.988044-1-xiaolinkui@126.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17Merge git://git.kernel.org/pub/scm/linux/kernel/git/netdev/netJakub Kicinski
Cross-merge networking fixes after downstream PR (net-7.3-rc4). Conflicts: net/core/neighbour.c 979aabdad8dd0 ("neighbour: Skip default parms when resumed in neightbl_dump_info().") 7b430fcfc972f ("neighbour: Don't render blackhole_netdev via RTM_GETNEIGHTBL.") fae1c59810b86 ("neighbour: Remove unnecessary net_eq().") https://lore.kernel.org/20260911173056.44ec06e0@kernel.org https://lore.kernel.org/aqfbJi7nAX4IbmnR@sirena.co.uk Adjacent changes: net/netlink/af_netlink.c ceac0de741bf ("netlink: do not free nlk->groups while lockless readers can use it") 7c0ec6288b49 ("net: Replace %pK output with 0") net/bridge/br_vlan.c 2842ce397dd0 ("net: bridge: vlan: fix bugs caused by switchdev deletion errors") 5bec8f861114 ("net: bridge: vlan: annotate lockless use of num_vlans") 2b1f8fd3118c ("net: bridge: vlan: annotate lockless vlan flags use") net/bridge/br_mst.c 18a6fe05fb6e ("net: bridge: mst: move switchdev call outside rcu") 120207a08fc0 ("net: bridge: vlan: annotate lockless use of msti") drivers/net/ethernet/stmicro/stmmac/hwif.h 90e4b849dfa6 ("net: stmmac: propagate FPE preemption-class mapping errors") 85ca3292d7a3 ("net: stmmac: Remove ARP offload code") Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17tcp: exclude old ACKs from tcp fast pathInbal Schussheim
Exclude old ACKs before SND.UNA from the tcp fast path as well as ACKs after SND.NXT. Such ACKs will fall through to the slow path, where tcp_ack() performs the appropriate validation and challenge ACK handling according to RFC5961 and Commit 3d501dd326fb1c7 ("tcp: do not accept ACK of bytes we never sent"). This prevents old ACKs from being accepted or modifying connection state as part of the fast path before appropriate ACK validation is applied. In particular, this prevents payload carried by a segment with an excessively old ACK from advancing RCV.NXT before the ACK is rejected. Fixes: 31770e34e43d ("tcp: Revert "tcp: remove header prediction"") Reported-by: Amit Klein <amit.klein@mail.huji.ac.il> Reported-by: Tamir Shahar <tamir.shahar1@mail.huji.ac.il> Reported-by: Inbal Schussheim <inbal.lipshtat@mail.huji.ac.il> Suggested-by: Eric Dumazet <edumazet@google.com> Cc: stable@vger.kernel.org Signed-off-by: Inbal Schussheim <inbal.lipshtat@mail.huji.ac.il> Reviewed-by: Eric Dumazet <edumazet@google.com> Link: https://patch.msgid.link/20260914090408.1435080-2-inbal.lipshtat@mail.huji.ac.il Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-16net: psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()Daniel Zahka
PSP conflicts with TLS ULP in its usage of both skb->decrypted and sk->sk_validate_xmit_skb(). Make PSP mutually exclusive with TLS ULP, the only other user of either of these. As other users of skb->decrypted come along, they can be added to sk_has_decrypt_user(). It would make sense to also assert that sk->sk_validate_xmit_skb() is also NULL in both of these setup paths for similar future proofing, but the PSP listener/sk_clone() path is still broken and it could be seen as a regression to not allow rx assoc to run on a child of a listener socket with PSP tx assoc state. Include all TCP ULPs in the sk_has_decrypt_user() check, even though TLS is the only one that conflicts with PSP via the decrypted bit. This is intentional because PSP was not designed to be used with ULPs. It is best to close off surface area that may make bugs reachable, until someone wishes to design and test an actual user of PSP with ULPs. Fixes: 6b46ca260e22 ("net: psp: add socket security association code") Signed-off-by: Daniel Zahka <daniel.zahka@gmail.com> Reviewed-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/20260915-psp-ktls-fix-v2-1-0eedc3b148ec@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16Merge tag 'ipsec-2026-09-16' of ↵Jakub Kicinski
git://git.kernel.org/pub/scm/linux/kernel/git/klassert/ipsec Steffen Klassert says: ==================== pull request (net): ipsec 2026-09-16 1) xfrm: iptfs: fix stack OOB read in iptfs_skb_reset_frag_walk() Add the up-front nr_frags guard iptfs_skb_add_frags() already has, so an out-of-range offset can't walk past the on-stack frags[] array. 2) xfrm: serialize state GC with device state flush Serialize xfrm_state destruction against the deferred-device pass with a dedicated mutex, since the device GC list doesn't hold a state reference and the two paths could free the same state. 3) xfrm: add missing RCU read lock in xfrm_send_migrate_state() Hold the RCU read lock around xfrm_nlmsg_multicast() so the rcu_dereference() of net->xfrm.nlsk doesn't warn. 4) xfrm: iptfs: fix runt reassembly panic from short inner tot_len Require the runt length to cover at least the minimum IP header, so a tot_len in [6, 19] (IPv4) can't write past the declared length and trip skb_over_panic(). 5) ipv6: xfrm: use full sockets in local error paths Use skb_to_full_sk() in xfrm6_local_rxpmtu() and xfrm6_local_error() and bail out without a full socket, so a TCP_NEW_SYN_RECV request_sock isn't miscast as a full inet/IPv6 socket. 6) xfrm: fix compat ALLOCSPI request use-after-free Drop the redundant alloc_compat() in xfrm_alloc_userspi() so the compat translator no longer reads past the payload and publishes a child a multicast clone can still see after xfrm_user_rcv_msg() frees. 7) xfrm: add missing rcu_read_lock(), skb_dst_force() and dev_hold() for xfrm_trans_reinject() Force the dst before queuing, hold dev across the workqueue deferral, and take rcu_read_lock() around the finish() loop, so transport-mode reinjection doesn't deref non-refcounted dst/dev under workqueue. 8) xfrm: use hlist_del_init_rcu for state_cache and state_cache_input Switch to hlist_del_init_rcu() so a second __xfrm_state_delete() is a no-op instead of writing through LIST_POISON2, closing the UAFs. 9) esp: downgrade zerocopy managed frags before mutating skb frags Call skb_zcopy_downgrade_managed() before ESP rewrites the skb frag array, so per-frag unrefs in esp_ssg_unref() and skb_release_data() stay balanced for ubuf-owned managed frags. 10) xfrm: hold net_device reference under RCU in bundle creation Read dst->dev via dst_dev_rcu() and keep RCU active through xfrm_fill_dst(), so a concurrent RTM_DELLINK can't free dev under bundle creation. 11) xfrm: save input state data before secpath resets Save the state protocol on the stack while it's still valid and use the saved address family for transport_finish(), so post-reset dereferences (VTI, XFRM if, MAX_DEPTH error) can't UAF the state. 12) net: xfrm: reject unrepresentable espintcp transport headers Use the careful transport-header helper and drop the skb through the XFRM error path when the offset can't be represented, instead of silently truncating it. * tag 'ipsec-2026-09-16' of git://git.kernel.org/pub/scm/linux/kernel/git/klassert/ipsec: net: xfrm: reject unrepresentable espintcp transport headers xfrm: save input state data before secpath resets xfrm: hold net_device reference under RCU in bundle creation esp: downgrade zerocopy managed frags before mutating skb frags xfrm: use hlist_del_init_rcu for state_cache and state_cache_input xfrm: add missing rcu_read_lock(), skb_dst_force() and dev_hold() for xfrm_trans_reinject() xfrm: fix compat ALLOCSPI request use-after-free ipv6: xfrm: use full sockets in local error paths xfrm: iptfs: fix runt reassembly panic from short inner tot_len xfrm: add missing RCU read lock in xfrm_send_migrate_state() xfrm: serialize state GC with device state flush xfrm: iptfs: fix stack OOB read in iptfs_skb_reset_frag_walk() ==================== Link: https://patch.msgid.link/20260916101938.118628-1-steffen.klassert@secunet.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15net: ip_tunnel: initialize `options_len` before referencing optionsGris Ge
The following command triggers a kernel panic: ip link add d0 type dummy; ip link set d0 up ip route add 10.30.0.0/16 \ encap ip id 300 geneve_opts 4660:66:11223344 dev d0 memcpy: detected buffer overflow: 4 byte write of buffer size 0 kernel BUG at lib/string_helpers.c:1044! ... ip_tun_parse_opts.part.0.cold+0x10/0x10 ip_tun_build_state+0x116/0x2a0 On kernels built with GCC 15+ and `CONFIG_FORTIFY_SOURCE`, the fortified `memcpy()` got 0 sized destination with request of 4 bytes length: static int ip_tun_parse_opts_geneve(...) { ... attr = tb[LWTUNNEL_IP_OPT_GENEVE_DATA]; data_len = nla_len(attr); /* == 4 */ struct geneve_opt *opt = ip_tunnel_info_opts(info) + opts_len; memcpy(opt->opt_data, nla_data(attr), data_len); /* ^^^^^^^^^^^^^ 0 since options_len is assigned afterwards */ Fixed by initializing the counter before the options are referenced. Matching what `tunnel_key_opts_set()` already does. Fixes: bb5e62f2d547 ("net: Add options as a flexible array to struct ip_tunnel_info") Cc: stable@vger.kernel.org Signed-off-by: Gris Ge <cnfourt@gmail.com> Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn> Reviewed-by: Gustavo A. R. Silva <gustavoars@kernel.org> Link: https://patch.msgid.link/20260913090851.468216-1-cnfourt@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15tcp: do not let tcp_rmem be set below 4096Eric Dumazet
We can hit a division by zero crash in tcp_rcvbuf_grow() and tcp_rcv_space_adjust(): divide error: 0000 [#1] PREEMPT SMP RIP: 0010:tcp_rcvbuf_grow+0x187/0x450 net/ipv4/tcp_input.c:939 ... grow = div_u64(((u64)rcvwin << 1) * (newval - oldval), oldval); The division uses oldval = tp->rcvq_space.space as divisor. When tp->rcvq_space.space is zero, this leads to a divide-by-zero exception. tp->rcvq_space.space is initialized in tcp_init_buffer_space(): tp->rcvq_space.space = min3(tp->rcv_ssthresh, tp->rcv_wnd, (u32)TCP_INIT_CWND * tp->advmss); If tcp_rmem[1] is configured to very small values (such as 1), sk->sk_rcvbuf is initialized to 1. Then tcp_full_space(sk), which computes (sk->sk_rcvbuf * scaling_ratio) >> 8, truncates to 0. This sets tp->window_clamp = 0, tp->rcv_ssthresh = 0, and tp->rcvq_space.space = 0. Later, when data arrives and DRS is invoked, tcp_rcvbuf_grow() divides by oldval == 0. Back in 2015, commit b1cb59cf2efe ("net: sysctl_net_core: check SNDBUF and RCVBUF for min length") ensured that net.core.rmem_default and net.core.rmem_max cannot be set below SOCK_MIN_RCVBUF. Similarly, SO_RCVBUF setsockopt enforces max_t(int, val * 2, SOCK_MIN_RCVBUF). However, net.ipv4.tcp_rmem still had .extra1 = SYSCTL_ONE, allowing arbitrarily small values. Because SOCK_MIN_RCVBUF depends on sizeof(struct sk_buff) and cacheline alignment, its value varies across architectures and configuration options. Using a fixed constant of 4096 ensures a predictable, architecture- independent lower bound that is safely above SOCK_MIN_RCVBUF everywhere and matches the documented 4K default. Fix this by setting tcp_rmem.extra1 to 4096 and updating the documentation. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Signed-off-by: Eric Dumazet <edumazet@google.com> Reviewed-by: Simon Horman <horms@kernel.org> Link: https://patch.msgid.link/20260912144848.3448026-1-edumazet@google.com Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15tcp: make smp_rmb() conditional in tcp_poll()Eric Dumazet
Commit a4d258036ed9 ("tcp: Fix race in tcp_poll") added smp_rmb() in tcp_poll() and smp_wmb() in tcp_reset() (now tcp_done_with_error()) to ensure that if tcp_poll() observed socket closure, it would also observe sk->sk_err. Currently, tcp_poll() unconditionally executes smp_rmb() at the end of every invocation, which on weakly-ordered architectures such as ARM64 emits a memory barrier instruction (dmb ishld) on the poll fast path, even for healthy, active sockets. However, tcp_poll() only needs this barrier if socket closure has been observed, to ensure that the error code set by tcp_done_with_error() before socket closure is visible before returning EPOLLERR. Move smp_rmb() inside the conditional block handling socket closure (shutdown == SHUTDOWN_MASK || state == TCP_CLOSE). For healthy connected sockets in epoll, tcp_poll() avoids the barrier entirely. Signed-off-by: Eric Dumazet <edumazet@google.com> Reviewed-by: Simon Horman <horms@kernel.org> Link: https://patch.msgid.link/20260913123224.762935-1-edumazet@google.com Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15ip_tunnel: Support per-netns device unregistration.Kuniyuki Iwashima
ip_tunnel_delete_net() iterates ip_tunnel devices whose link_net is dying and queues them for destruction. The devices may reside in different netns. Let's use unregister_netdevice_queue_net() to support per-netns device unregistration. Even after ip_tunnel_delete_net() queues a cross-netns ip_tunnel device, ip_tunnel_changelink(), ip_tunnel_dellink(), and ip_tunnel_ctl() could be called concurrently for it (once RTNL is removed). In such a case, __rtnl_net_unlock() will perform the unregistration. Also, ip_tunnel_ctl() needs to check check_net(t->net), otherwise it could create a new dev in dying netns after ip_tunnel_delete_net(). In the example below, we can see the fallback tunnel device (gre0) and the cross-netns device (gre1) are unregistered by different processes: # bpftrace -e '#include <linux/netdevice.h> kprobe:ip_tunnel_uninit { $dev = (struct net_device *)arg0; printf("PID: %d | DEV: %s%s\n", pid, $dev->name, kstack()); } kprobe:ipgre_exit_rtnl { printf("PID: %d%s\n", pid, kstack()); }' & # ip netns add ns1 # ip netns add ns2 # ip -n ns1 link add name gre1 link-netns ns2 \ type gre local 192.168.0.1 remote 192.168.1.1 # ip netns del ns2 PID: 12 ipgre_exit_rtnl+5 ops_undo_list+702 cleanup_net+1122 process_scheduled_works+2538 ... PID: 12 | DEV: gre0 <------ fallback device (itn->fb_tunnel_dev). ip_tunnel_uninit+5 unregister_netdevice_many_notify+7129 unregister_netdevice_many_net+1050 __rtnl_net_unlock+37 ops_undo_list+754 cleanup_net+1122 process_scheduled_works+2538 ... PID: 10 | DEV: gre1 ip_tunnel_uninit+5 unregister_netdevice_many_notify+7129 unregister_netdevice_many_net+1050 rtnl_net_work_func+136 process_scheduled_works+2538 Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com> Link: https://patch.msgid.link/20260912230043.2586313-8-kuniyu@google.com Reviewed-by: Ido Schimmel <idosch@nvidia.com> Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15ip_tunnel: Protect ip_tunnel_net.tunnels[] with mutex.Kuniyuki Iwashima
struct ip_tunnel.net is the netns where encapsulated packets flow into. struct ip_tunnel is linked to ip_tunnel_net.tunnels[] of netns. During netns dismantle or module unload, ip_tunnel_delete_net() iterates the list and queues devices for destruction regardless of the devices' netns. Thus, once RTNL is removed, the list can be modified concurrently from different netns due to device removal. Let's protect it with per-netns mutex. Note that dev_siocdevprivate() calls netdev_lock_ops() but it must be NOP for tunnel devices to avoid AB-BA deadlock. DEBUG_NET_WARN_ON_ONCE() is added to annotate the locking explicitly. Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com> Link: https://patch.msgid.link/20260912230043.2586313-7-kuniyu@google.com Reviewed-by: Ido Schimmel <idosch@nvidia.com> Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15ip_tunnel: Unify error paths in ip_tunnel_newlink() and ip_tunnel_changelink().Kuniyuki Iwashima
The next patch will introduce per-netns mutex and acquire it in ip_tunnel_newlink() and ip_tunnel_changelink(). To make the diff cleaner, let's unify the error paths. Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com> Link: https://patch.msgid.link/20260912230043.2586313-6-kuniyu@google.com Reviewed-by: Ido Schimmel <idosch@nvidia.com> Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15ip_tunnel: Centralise ip_tunnel_del() to ip_tunnel_dellink().Kuniyuki Iwashima
With the previous patch, itn->fb_tunnel_dev can be removed via ->dellink(). However, ioctl(SIOCDELTUNNEL) still uses unregister_netdevice(), which requires ip_tunnel_del() in ip_tunnel_uninit(). Let's use ip_tunnel_dellink() everywhere to remove ip_tunnel device and remove ip_tunnel_del() in ip_tunnel_uninit(). Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com> Link: https://patch.msgid.link/20260912230043.2586313-5-kuniyu@google.com Reviewed-by: Ido Schimmel <idosch@nvidia.com> Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15ip_tunnel: Don't pass rtnl_link_ops to ip_tunnel_delete_net().Kuniyuki Iwashima
ip_tunnel_delete_net() no longer uses the 3rd argument, struct rtnl_link_ops *ops. Let's remove it. Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com> Link: https://patch.msgid.link/20260912230043.2586313-4-kuniyu@google.com Reviewed-by: Ido Schimmel <idosch@nvidia.com> Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15ip_tunnel: Set itn->fb_tunnel_dev to NULL in ip_tunnel_delete_net().Kuniyuki Iwashima
ip_tunnel_dellink() ignores itn->fb_tunnel_dev, so the per-netns fallback tunnel device cannot be removed by userspace. This also makes default_device_exit_batch() impossible to remove the device since it calls ->dellink(). So, ip_tunnel_delete_net() has to iterate devices in the dying netns and call unregister_netdevice_queue() directly. But then, this duplicates ip_tunnel_del() in ip_tunnel_dellink() and ip_tunnel_uninit(). Let's set itn->fb_tunnel_dev to NULL in ip_tunnel_delete_net() and remove for_each_netdev_safe() in ip_tunnel_delete_net(). Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com> Link: https://patch.msgid.link/20260912230043.2586313-3-kuniyu@google.com Reviewed-by: Ido Schimmel <idosch@nvidia.com> Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15ipmr: Call ->dellink() to remove DVMRP tunnel device.Kuniyuki Iwashima
ipmr.c uses unregister_netdevice() to remove DVMRP tunnel devices created in ipmr_new_tunnel(). This is fine because currently ip_tunnel_uninit() also calls ip_tunnel_del() to unlink the device from the hash table. However, we will move ip_tunnel_del() from ip_tunnel_uninit() to ip_tunnel_dellink(). Removing DVMRP tunnel devices by unregister_netdevice() would leave them in the hash table. Let's call ->dellink for DVMRP tunnel devices. Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com> Link: https://patch.msgid.link/20260912230043.2586313-2-kuniyu@google.com Reviewed-by: Ido Schimmel <idosch@nvidia.com> Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15ipv4: icmp: reject RTN_UNREACHABLE input routes in icmp_route_lookupDong Chenchen
When the forward output route cannot be used in icmp_route_lookup(), it enters the "reverse path" and calls ip_route_input() on fl4_dec.daddr, the original packet's source address. ip_route_input() only returns an error for truly invalid packets. For unreachable addresses it will succeed and return an input route whose dst.output is set to ip_rt_bug(). The existing check only rejects RTN_LOCAL routes, so the RTN_UNREACHABLE route types can still be returned and later used for output, syzkaller triggering a WARN_ON_ONCE() in ip_rt_bug() as bellow: ------------[ cut here ]------------ WARNING: net/ipv4/route.c:1273 at ip_rt_bug+0x14/0x20 RIP: 0010:ip_rt_bug+0x14/0x20 Call Trace: ip_push_pending_frames+0xfa/0x100 __icmp_send+0x905/0xf10 ip_options_compile+0xc0/0xd0 ip_rcv_finish_core+0x321/0xae0 ip_rcv+0x1de/0x260 __netif_receive_skb_one_core+0x11a/0x130 netif_receive_skb+0x7b/0x260 tun_get_user+0x11bf/0x1c10 ------------[ cut here ]------------ Reject input route that is RTN_UNREACHABLE to fix it. The net warning is only printed for RTN_LOCAL, as RTN_UNREACHABLE is not the result of a race condition. Fixes: 8b7817f3a959 ("[IPSEC]: Add ICMP host relookup support") Suggested-by: Ido Schimmel <idosch@nvidia.com> Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Signed-off-by: Dong Chenchen <dongchenchen2@huawei.com> Link: https://patch.msgid.link/20260910140042.1880242-1-dongchenchen2@huawei.com Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-14ipmr, ip6mr: annotate data-races in vif_seq_show()Linkui Xiao
ipmr_vif_seq_show() and ip6mr_vif_seq_show() read vif->bytes_in, vif->pkt_in, vif->bytes_out and vif->pkt_out with only rcu_read_lock() held: since commit b96ef16d2f83 ("ipmr: convert /proc handlers to rcu_read_lock()") both seq_start helpers are annotated __acquires(RCU) and no longer take mrt_lock. Those counters are updated from softirq context and the writers already use WRITE_ONCE(): ipmr_prepare_xmit() and ip_mr_forward() on the IPv4 side, ip6mr_prepare_xmit() and ip6_mr_forward() on the IPv6 side. The other lockless readers use READ_ONCE() as well - ipmr_ioctl(), ipmr_compat_ioctl(), ipmr_fill_vif(), ip6mr_ioctl() and ip6mr_compat_ioctl(). The two vif_seq_show() helpers are the only remaining bare readers, so KCSAN flags them and the compiler is free to tear or reload the values while the /proc/net/ip_mr_vif and /proc/net/ip6_mr_vif lines are being formatted. Annotate them like the other readers; these are plain statistics, no locking is needed. Signed-off-by: Linkui Xiao <xiaolinkui@kylinos.cn> Link: https://patch.msgid.link/20260910093452.2070079-1-xiaolinkui@126.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-11net: dropreason: add SKB_DROP_REASON_IP_TTL_EXCEEDEDJunjie Cao
The forwarding paths report an expired TTL or hop limit as SKB_DROP_REASON_IP_INHDR, the reason otherwise used for a header that is malformed (ip_input.c, exthdrs.c, br_netfilter). Nothing else in the drop path separates the two: IPSTATS_MIB_INHDRERRORS covers both, and the TTL check runs before NF_INET_FORWARD, so netfilter tracing stops at PREROUTING and never sees the drop. The Fedora bug linked below shows how that reads in practice. The reporter took kfree_skb(reason=IP_INHDR, loc=ip_forward) to mean the software header checksum check had failed, and worked through RX checksum offload, tc csum actions and both libvirt firewall backends before the drops turned out to be replies arriving with TTL 1. ip_forward() never verifies the header checksum; that runs earlier, in ip_rcv_core(), and reports IP_CSUM. TTL expiry is not a corner case -- every traceroute through a Linux router goes through too_many_hops. The three loopback hop limit checks in exthdrs.c drop with no reason at all; give them the new one. IPSTATS_MIB_INHDRERRORS stays as it is: RFC 1213 counts time-to-live exceeded under ipInHdrErrors. The drop reason has no such constraint. Link: https://bugzilla.redhat.com/show_bug.cgi?id=2517131 Signed-off-by: Junjie Cao <junjie.cao@intel.com> Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev> Reviewed-by: Fernando Fernandez Mancera <fmancera@suse.de> Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn> Link: https://patch.msgid.link/20260910094937.536150-1-junjie.cao@intel.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-10ip_tunnel: use WRITE_ONCE in ip_tunnel_encap_setupEric Dumazet
Update ip_tunnel_encap_setup() to use WRITE_ONCE() when writing to encap fields (type, sport, dport, flags) and hlen fields. This ensures that concurrent lockless readers (like fill_info) do not see torn writes. Also remove the unsafe memset() on t->encap which could cause concurrent readers to transiently see zeroed fields. Removing it also fixes a bug where t->encap was left cleared even if ip_encap_hlen() failed, resulting in partial configuration. Fixes: 56328486539d ("net: Changes to ip_tunnel to support foo-over-udp encapsulation") Signed-off-by: Eric Dumazet <edumazet@google.com> Acked-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com> Link: https://patch.msgid.link/20260907075846.2913645-4-edumazet@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-10tcp: Skip cond_resched() in inet_csk_listen_stop() under BPF contextJiayuan Chen
bpf_sock_destroy() runs from the tcp iterator, under rcu_read_lock(). If the sock is a listener that still has children in its accept queue, tcp_abort() ends up in inet_csk_listen_stop() and the cond_resched() there trips the debug check: BUG: sleeping function called from invalid context at net/ipv4/inet_connection_sock.c:1523 in_atomic(): 0, irqs_disabled(): 0, non_block: 0, pid: 628, name: test_progs preempt_count: 0, expected: 0 RCU nest depth: 1, expected: 0 locks held by test_progs/628: 3, last CPU#3: #0: ffff8881158cee18 (&p->lock){+.+.}-{4:4}, at: bpf_seq_read+0x56/0x1210 #1: ffff8881106bb858 (sk_lock-AF_INET6){+.+.}-{0:0}, at: bpf_iter_tcp_seq_show+0x32b/0x4b0 #2: ffffffffb435af20 (rcu_read_lock){....}-{1:3}, at: bpf_iter_run_prog+0x46b/0xde0 CPU: 3 UID: 0 PID: 628 Comm: test_progs Tainted: G W 7.2.0+ #65 PREEMPT Tainted: [W]=WARN Call Trace: <TASK> dump_stack_lvl+0xc1/0xf0 dump_stack+0x10/0x20 __might_resched+0x3d2/0x610 inet_csk_listen_stop+0x7b/0xbf0 tcp_abort+0x23b/0x3b0 bpf_sock_destroy+0xfc/0x140 bpf_prog_448133d24601754f_iter_tcp6_server+0x81/0x8a bpf_iter_run_prog+0x538/0xde0 bpf_iter_tcp_seq_show+0x26b/0x4b0 bpf_seq_read+0x424/0x1210 vfs_read+0x197/0xe40 ksys_read+0x119/0x240 __x64_sys_read+0x72/0xc0 x64_sys_call+0x647/0x27e0 do_syscall_64+0xe5/0x610 entry_SYSCALL_64_after_hwframe+0x76/0x7e RIP: 0033:0x7fad39b28aca RSP: 002b:00007ffc381c61c0 EFLAGS: 00000246 ORIG_RAX: 0000000000000000 RAX: ffffffffffffffda RBX: 00007ffc381c6a88 RCX: 00007fad39b28aca RDX: 0000000000000032 RSI: 00007ffc381c6250 RDI: 0000000000000014 RBP: 00007ffc381c61e0 R08: 0000000000000000 R09: 0000000000000000 R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000003 R13: 0000000000000000 R14: 000055f077c1bbb0 R15: 00007fad3a0f3000 </TASK> The commit that added the kfunc already guards lock_sock() in tcp_abort() and udp_abort() with has_current_bpf_ctx(), but missed the listener path. Do the same for the cond_resched(). The loop runs inside the iterator's rcu_read_lock(), it must not reschedule or report a quiescent state there. Fixes: 4ddbcb886268 ("bpf: Add bpf_sock_destroy kfunc") Signed-off-by: Jiayuan Chen <jiayuan.chen@linux.dev> Link: https://lore.kernel.org/r/20260910112736.153710-1-jiayuan.chen@linux.dev Signed-off-by: Alexei Starovoitov <ast@kernel.org>
2026-09-10Merge git://git.kernel.org/pub/scm/linux/kernel/git/netdev/netJakub Kicinski
Cross-merge networking fixes after downstream PR (net-7.3-rc3). Conflicts: drivers/net/dsa/mt7530.c 3c18e3c9a54e ("net: dsa: mt7530: populate lpi_interfaces to fix EEE support") 10d9d8328e8a ("net: dsa: mt7530: replace mt7530_read with regmap_read") Adjacent changes: drivers/net/bonding/bond_alb.c 1746ef2e2df2 ("bonding: use skb_cow_head() in bond_do_alb_xmit() and rlb_arp_xmit()") 4cef95f72bbd ("bonding: fix u32 overflow in compute_gap()") Signed-off-by: Jakub Kicinski <kuba@kernel.org>