summaryrefslogtreecommitdiff
path: root/net
AgeCommit message (Collapse)Author
11 hoursMerge branch 'headers' of git://git.infradead.org/users/willy/pagecache.gitMark Brown
# Conflicts: # drivers/gpu/drm/amd/amdkfd/kfd_migrate.c # net/ceph/osd_client.c
13 hoursMerge branch 'master' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git # Conflicts: # Documentation/scheduler/index.rst # arch/arm64/configs/defconfig
13 hoursMerge branch 'modules-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/modules/linux.git
14 hoursMerge branch 'master' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth-next.git
14 hoursMerge branch 'for-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next.git # Conflicts: # arch/arm64/net/bpf_jit_comp.c # arch/x86/net/bpf_jit_comp.c # mm/internal.h
14 hoursMerge branch 'main' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next.git # Conflicts: # drivers/net/ethernet/realtek/r8169_main.c # net/mac80211/ieee80211_i.h # net/mac80211/tx.c
14 hoursMerge branch 'fs-next' of linux-nextMark Brown
# Conflicts: # fs/coredump.c # fs/f2fs/f2fs.h # fs/fuse/dax.c # fs/xfs/libxfs/xfs_btree.c
15 hoursMerge branch 'for-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/mm/linux.git # Conflicts: # arch/arm64/kvm/mmu.c
15 hoursMerge branch 'master' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth.git
15 hoursMerge branch 'master' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/klassert/ipsec.git
15 hoursMerge branch 'vfs.all' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs.git # Conflicts: # fs/smb/server/smb2pdu.c # fs/smb/server/vfs.c # fs/smb/server/vfs.h
15 hoursMerge branch 'nfsd-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux
15 hoursMerge branch 'linux-next' of git://git.linux-nfs.org/projects/anna/linux-nfs.gitMark Brown
16 hoursseg6: fix HMAC validation when an extension header precedes the SRHYuya Kusakabe
seg6_hmac_validate_skb() derived the SRH from skb_transport_header(). That only holds while the two coincide, which is not true on the seg6_local input path. ip6_rcv_core() leaves the transport header just past the IPv6 header. A Hop-by-Hop options header is consumed before the route lookup and advances it, but a Destination Options header is not: the seg6_local lwtunnel is entered through an input redirect from the route lookup, which bypasses the extension header handlers. The transport header then still points at the Destination Options header while seg6_get_srh() has located the real SRH further down the chain. The HMAC is therefore computed over the Destination Options header, and a packet carrying a valid HMAC TLV is dropped when seg6_require_hmac is set. Such a packet is legitimate: RFC 8200 allows Destination Options before a routing header, and get_srh() has walked the header chain since commit 5829d70b0b6c ("ipv6: sr: fix get_srh() to comply with IPv6 standard "RFC 8200""). With a Fragment or an Authentication header in front of the SRH, the same mistake also reads past the data pulled by seg6_get_srh(). This happens before seg6_require_hmac is read, so the default configuration is affected. Reproduce by giving a node a seg6local End SID with net.ipv6.conf.<dev>.seg6_require_hmac=1 and a key installed with "ip sr hmac set <keyid> sha1", then sending IPv6 -> Destination Options -> SRH (carrying a valid HMAC TLV) -> payload to that SID: it is dropped, while the same packet without the Destination Options header passes. Fixes: 5829d70b0b6c ("ipv6: sr: fix get_srh() to comply with IPv6 standard "RFC 8200"") Assisted-by: LLM Signed-off-by: Yuya Kusakabe <yuya.kusakabe@gmail.com> Reviewed-by: Andrea Mayer <andrea.mayer@uniroma2.it> Link: https://patch.msgid.link/20260926-b4-seg6-hmac-transport-header-v2-1-8087b76f6375@gmail.com Signed-off-by: Paolo Abeni <pabeni@redhat.com>
17 hoursnet: only give queue leasing devices a separate instance lock classJakub Kicinski
Commit b6f74dff6d26 ("net: use two lockdep classes for the netdev instance lock") put every device without a parent in the virtual class. It also restricted the locking order for the virtual class, because queue leasing has hard requirements on the exact order. This bites us back on bond, which is "virtual" and needs to be taken before taking the locks of the lowers. NIPA hit the following on the new test I recently posted for XDP+bond: WARNING: possible circular locking dependency detected ------------------------------------------------------ python3/18635 is trying to acquire lock: ff11000120b8ce30 (&dev->lock){+.+.}-{4:4}, at: netdev_put_lock+0x2d/0x1a0 but task is already holding lock: ff110001ef77ae30 (&netdev_virt_instance_lock_key){+.+.}-{4:4}, at: netdev_put_lock+0x2d/0x1a0 which lock already depends on the new lock. the existing dependency chain (in reverse order) is: -> #1 (&netdev_virt_instance_lock_key){+.+.}-{4:4}: __mutex_lock+0x1ae/0x1f10 xdp_set_features_flag+0x2b/0x50 bond_xdp_set_features+0x1eb/0x360 bond_netdev_event+0x13f/0x300 notifier_call_chain+0xae/0x300 call_netdevice_notifiers+0x70/0xa0 bnxt_xdp_set+0x2f6/0x620 netif_xdp_propagate+0x503/0xc60 dev_xdp_propagate+0xa1/0x230 bond_xdp_set+0x234/0x700 dev_xdp_install+0x592/0xd70 dev_xdp_attach+0x355/0xf50 dev_change_xdp_fd+0x176/0x210 do_setlink.isra.0+0x220d/0x2b20 rtnl_newlink+0x9f1/0x11b0 -> #0 (&dev->lock){+.+.}-{4:4}: __mutex_lock+0x1ae/0x1f10 netdev_put_lock+0x2d/0x1a0 netdev_nl_queue_create_doit+0x801/0x1a70 genl_family_rcv_msg_doit+0x206/0x300 Possible unsafe locking scenario: CPU0 CPU1 ---- ---- lock(&netdev_virt_instance_lock_key); lock(&dev->lock); lock(&netdev_virt_instance_lock_key); lock(&dev->lock); Let's narrow down the "virtual" class to only the devices which can actually create a queue. More LoC and complexity, but that is what we actually care about here. The rest needs to nest under rtnl_lock, which bond does (famous last words?) Take the instance locks in two passes, first the netkits then the rest (matching the queue leasing order). An alternative would be to make sure the close list is sorted correctly from the start (queue head/tail appropriately in unregister_netdevice_queue()). I think it works but feels a little more fragile. Happy to change, tho. netdev_can_create_queue() will now be used on paths where we genuinely handle non-netkit, so we can't always set the extack. Unfortunately, the (recently) added tracepoint in extack fires even when extack is NULL. Fixes: b6f74dff6d26 ("net: use two lockdep classes for the netdev instance lock") Signed-off-by: Jakub Kicinski <kuba@kernel.org> Acked-by: Stanislav Fomichev <sdf@fomichev.me> Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org> Acked-by: Daniel Borkmann <daniel@iogearbox.net> Link: https://patch.msgid.link/20260929191543.3295633-1-kuba@kernel.org Signed-off-by: Paolo Abeni <pabeni@redhat.com>
17 hoursMerge tag 'nf-26-09-30' of ↵Paolo Abeni
git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf Pablo Neira Ayuso says: ==================== Netfilter/IPVS fixes for net The following batch contains Netfilter fixes for net. This batch fixes crashes as recent feature regression, one of the due to a dependency that has been pulled into -stable: 1) Expand existing ipset fix for bitmap sets to disallow comments updates from kernel-side adds, from Florian Westphal. 2) Drop flowtable reference if nf_ct_netns_get() fails, otherwise flowtable cannot ever be removed, from Aohan Mei. 3) nft_rbtree GC should collect end elements that contained in this transaction batch, new or deleted elements are never expired. From Weiming Shi. 4) Restrict nf_nat_bpf so it does not set unknown NF_NAT_MANIP_* values, from Fernando F. Mancera. 5) Flowtable GC must skip flows that are pending hardware updates, generalize the PENDING flag and use it to inhibit GC. 6) Restore flowtable with ieee80211 which broke due to a relatively recent commit, which was pulled in by -stable, causing a regression in 6.18 kernels. And the following IPVS fixes: 1) Fix accounting of cache entries in IPVS LBLC for destinations, which eventually fills up the table and trigger recurrent resizing, from Julian Anastasov. 2) Limit IPVS cache growth for LBLCR and LBLC schedulers, from Zhiling Zou. 3) Restrict IP_VS_CONN_F_ONE_PACKET for normal connections, do not allow to use it with templates. Also from Julian. 4) Sanitize flags in IPVS sync messages received in the backup. From Julian Anastasov. netfilter pull request 26-09-30 * tag 'nf-26-09-30' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf: netfilter: flowtable: restore ieee80211 forward path netfilter: flowtable: generalize pending status bit netfilter: bpf: reject invalid NAT manipulation types netfilter: nft_set_rbtree: skip transaction elements during GC ipvs: filter some flags received in the backup server ipvs: do not create invisible templates ipvs: bound LBLCR and LBLC cache growth ipvs: fix missing counter decrement in lblc netfilter: nft_flow_offload: drop flowtable reference on init error path netfilter: ipset: do not update comments from kernel-side adds ==================== Link: https://patch.msgid.link/20260930074142.298353-1-pablo@netfilter.org Signed-off-by: Paolo Abeni <pabeni@redhat.com>
18 hoursmptcp: annotate lockless access to sk->sk_errQuanye Yang
sock_error() can clear sk_err with xchg() without the socket lock. On the no-data msk recv and splice paths, call sock_error() once and only stop when it returns a non-zero error. Annotate the remaining msk send and subflow error-report peeks with READ_ONCE(). Signed-off-by: Quanye Yang <quanyeyang@proton.me> Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org> Link: https://patch.msgid.link/20260925-mptcp-sk-err-net-v5-2-0cac04d6ea48@proton.me Signed-off-by: Paolo Abeni <pabeni@redhat.com>
18 hourstcp: annotate lockless access to sk->sk_errQuanye Yang
BUG: KCSAN: data-race in do_recvmmsg / mptcp_recvmsg read-write (marked) to 0xffff8880134d391c of 4 bytes by task 2619 on cpu 1: instrument_atomic_read_write include/linux/instrumented.h:113 [inline] sock_error include/net/sock.h:2565 [inline] do_recvmmsg+0x50c/0x580 net/socket.c:3049 __sys_recvmmsg net/socket.c:3144 [inline] __do_sys_recvmmsg net/socket.c:3167 [inline] __se_sys_recvmmsg net/socket.c:3160 [inline] __x64_sys_recvmmsg+0x161/0x180 net/socket.c:3160 x64_sys_call+0x19c7/0x1ca0 arch/x86/include/generated/asm/syscalls_64.h:300 do_syscall_x64 arch/x86/entry/syscall_64.c:61 [inline] do_syscall_64+0xde/0x3d0 arch/x86/entry/syscall_64.c:84 entry_SYSCALL_64_after_hwframe+0x77/0x7f read to 0xffff8880134d391c of 4 bytes by task 2620 on cpu 0: tcp_recv_should_stop include/net/tcp.h:3086 [inline] mptcp_recvmsg+0x54d/0xd50 net/mptcp/protocol.c:2466 inet_recvmsg+0x204/0x210 net/ipv4/af_inet.c:894 sock_recvmsg_nosec net/socket.c:1151 [inline] sock_recvmsg+0x11a/0x140 net/socket.c:1173 ____sys_recvmsg+0x14b/0x3c0 net/socket.c:2933 ___sys_recvmsg+0x116/0x160 net/socket.c:2975 __sys_recvmsg net/socket.c:3008 [inline] __do_sys_recvmsg net/socket.c:3014 [inline] __se_sys_recvmsg net/socket.c:3011 [inline] __x64_sys_recvmsg+0xeb/0x160 net/socket.c:3011 x64_sys_call+0x1319/0x1ca0 arch/x86/include/generated/asm/syscalls_64.h:48 do_syscall_x64 arch/x86/entry/syscall_64.c:61 [inline] do_syscall_64+0xde/0x3d0 arch/x86/entry/syscall_64.c:84 entry_SYSCALL_64_after_hwframe+0x77/0x7f value changed: 0x0000006b -> 0x00000000 Reported by Kernel Concurrency Sanitizer on: CPU: 0 UID: 0 PID: 2620 Comm: syz.2.33 Not tainted 7.2.0-g39d4f32c5d53 #76 PREEMPT(full) Hardware name: QEMU Ubuntu 26.04 PC (i440FX + PIIX, 1996), BIOS 1.17.0-debian-1.17.0-1ubuntu1 04/01/2014 do_recvmmsg() and getsockopt(SO_ERROR) call sock_error() without the socket lock. sock_error() clears sk_err with xchg(), which races with unmarked loads of the same field. KCSAN reported the unmarked peek in tcp_recv_should_stop(). Annotate that helper and the other send-side peeks with READ_ONCE(). On the no-data recv and splice paths, if (sk_err) followed by sock_error() and an unconditional break can return 0 after another thread consumes the error. Call sock_error() once and only stop when it returns a non-zero error. tcp_bpf_sendmsg() read sk_err twice; fold those unmarked loads into one READ_ONCE() and use that value as the returned errno. The field is still not consumed. MPTCP is handled in the next patch. Suggested-by: Jiayuan Chen <jiayuan.chen@linux.dev> Link: https://lore.kernel.org/netdev/8bbee583-6f21-4817-bfeb-2d60057380a3@linux.dev/ Reported-by: Matthieu Baerts (NGI0) <matttbe@kernel.org> Closes: https://github.com/multipath-tcp/mptcp_net-next/issues/632 Signed-off-by: Quanye Yang <quanyeyang@proton.me> Reviewed-by: Eric Dumazet <edumazet@kernel.org> Link: https://patch.msgid.link/20260925-mptcp-sk-err-net-v5-1-0cac04d6ea48@proton.me Signed-off-by: Paolo Abeni <pabeni@redhat.com>
19 hoursnet: drop unused LLC-related includesJakub Kicinski
The bridge and the openvswitch code include <net/llc.h>, <net/llc_pdu.h> and <linux/llc.h> without using anything out of them. parse_ethertype() is the one exception: it wants LLC_SAP_SNAP, which comes from the uAPI header <linux/llc.h> pulls in, so flow.c keeps that one. The LLC core has leftovers of its own - slab, string, interrupt and net_namespace - which it has not used in a long time. None of this is new, the type 2 removal has just shrunk the LLC headers enough to make it easy to see. Signed-off-by: Jakub Kicinski <kuba@kernel.org> Acked-by: Aaron Conole <aconole@redhat.com> Link: https://patch.msgid.link/20260928190800.2521749-4-kuba@kernel.org Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org> Signed-off-by: Paolo Abeni <pabeni@redhat.com>
19 hoursllc: strip the leftovers of the llc2 removalJakub Kicinski
Nothing in tree registers the type 1 / type 2 packet handlers or the station handler any more, and nothing looks at the socket hashes hanging off struct llc_sap. The two SAPs opened in tree - SNAP, and STP for the bridge's BPDUs and GARP's PDUs - receive through the per-SAP rcv_func(), so llc_rcv() boils down to a SAP lookup and a call. llc_sap_list becomes static and five exports go away with all this; the llc2 module out of tree has been reworked to open its own SAPs. Frames with a NULL DSAP used to be handed to the station handler before the SAP lookup. They take the normal path now, meaning they get dropped unless something registers SAP 0. struct llc_sap is down to 48 bytes on 64-bit from over a kilobyte, which also moves its GFP_ATOMIC allocation from kmalloc-2k to kmalloc-64. The two 64-entry socket hashes are the bulk of it, and the address it carried is now just the SAP number - nothing has read the MAC half since the socket layer left. llc_pdu.h keeps only what its remaining users need, plus LLC_PDU_RSP to document the one argument which can take it. That takes out the type 2 (I and S format, FRMR) definitions, the XID and TEST builders with their constants - the kernel neither sends nor answers either any more - the SAP address defines, the field accessors, and the prototypes of the llc_pdu.c helpers which went out of tree with the rest of LLC2. With the I and S formats gone llc_pdu_header_init() has one PDU type left, so drop the argument. Signed-off-by: Jakub Kicinski <kuba@kernel.org> Link: https://patch.msgid.link/20260928190800.2521749-3-kuba@kernel.org Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org> Signed-off-by: Paolo Abeni <pabeni@redhat.com>
19 hoursllc: move the type 2 sockets out of treeJakub Kicinski
802.2 LLC is only used in tree by protocols which need the connectionless type 1 subset: STP, which carries the bridge's BPDUs, GARP, which sends its own UI PDUs, and SNAP. The llc2 module on top of the core - the type 2 connection state machine, the type 1 SAP state machine, the station component and the PF_LLC socket family - has no in-kernel users and nobody who can test it; what we get instead is a slow trickle of drive-by fixes. There was a recent patch from Ernestas Kulik indicating potential real life use, but it was new/experimental and that person is not responding to off-list pings. Let LLC2 follow AX.25, hamradio and AppleTalk out of the Linux tree. We will maintain the code at: github.com/linux-netdev/mod-orphan for anyone interested in playing with it. PF_LLC goes in full, both the class two SOCK_STREAM and the class one SOCK_DGRAM half, and so do /proc/net/llc/ and /proc/sys/net/llc/. Note that the kernel also stops answering XID and TEST commands, addressed to a SAP or to the station - those are type 1, but they lived in the module, and they got answered whether or not any socket was open. Nothing in tree asks for them; what the core keeps is SAP registration and the UI path the in-tree users need. Retain the uAPI for now, like we did for AppleTalk. Only the socket ABI half of it is vestigial: STP, GARP, the bridge and openvswitch use the SAP numbers it defines. Cleaning up what the core no longer needs follows in the next patch. Signed-off-by: Jakub Kicinski <kuba@kernel.org> Link: https://patch.msgid.link/20260928190800.2521749-2-kuba@kernel.org Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org> Signed-off-by: Paolo Abeni <pabeni@redhat.com>
27 hoursnet/smc: Hold a socket reference for transmit workChengfeng Ye
SMC transmit work is queued without holding a socket reference. After an active close has cancelled tx_work, a received CDC message can queue it again while the socket is in SMC_PEERCLOSEWAIT1. Passive close can then reach SMC_CLOSED, call smc_conn_free() and drop the last socket reference before smc_tx_work() acquires the socket lock. The worker then accesses the freed socket. KASAN reported: BUG: KASAN: slab-use-after-free in lock_sock_nested+0x97/0x180 Write of size 8 by task kworker/0:1/11 Workqueue: smc_tx_wq-00000000 smc_tx_work Call Trace: lock_sock_nested+0x97/0x180 smc_tx_work+0x5d/0x170 process_one_work+0x5ce/0xeb0 worker_thread+0x45b/0xd10 Allocated by task 24: sk_prot_alloc+0x56/0x210 sk_alloc+0x2b/0x6f0 smc_tcp_listen_work+0x16d/0xfc0 Freed by task 182: slab_free_after_rcu_debug+0xa6/0x1e0 rcu_core+0x509/0x1850 Hold a socket reference for each successful enqueue and release it when the worker finishes or a pending invocation is cancelled. Cover all three queue sites and all cancellation sites, including the socket options. Check conn->freed in smc_tx_pending() under the socket lock: retaining the socket does not retain connection resources released by smc_conn_free(). This also covers pending transmit processing from smc_release_cb(). Use queue_delayed_work() for the busy-slot retry as well. All tx_work delays are zero, so an already queued invocation needs no timer update. Unlike mod_delayed_work(), its return value distinguishes a successful enqueue from work disabled temporarily by cancel_delayed_work_sync(), allowing the extra reference to be returned when no work was queued. Cc: stable+noautosel@kernel.org # LLM report + LLM fix, not seen in real life Fixes: e6727f39004b ("smc: send data (through RDMA)") Link: https://lists.openwall.net/netdev/2026/09/09/105 Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com> Link: https://patch.msgid.link/20260927075120.3695060-1-nicoyip.dev@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
27 hoursipv6: sr: use skb_get_hash_net() in seg6_make_flowlabel()Eric Dumazet
Since commit d58e468b1112 ("flow_dissector: implements flow dissector BPF hook") __skb_flow_dissect() needs a net pointer, either from skb->dev, skb->sk, or since commit 3cbf4ffba5ee ("net: plumb network namespace into __skb_flow_dissect") a caller provided pointer. syzbot was able to reach seg6_make_flowlabel() with an skb having neither skb->dev nor skb->sk set: a TIPC UDP bearer sends a discovery message through an IPv4 route using seg6 encap, while net.ipv6.seg6_flowlabel is set to 1. seg6_make_flowlabel() already has a net pointer, use skb_get_hash_net(). WARNING: net/core/flow_dissector.c:1131 at __skb_flow_dissect+0x910/0x5368 net/core/flow_dissector.c:1126, CPU#0: syz.0.17/4930 Call trace: __skb_flow_dissect+0x910/0x5368 net/core/flow_dissector.c:1126 (P) __skb_get_hash_net+0xe0/0x29c net/core/flow_dissector.c:1903 skb_get_hash include/linux/skbuff.h:1663 [inline] seg6_make_flowlabel+0xcc/0x1ec net/ipv6/seg6_iptunnel.c:132 __seg6_do_srh_encap+0x320/0xbf4 net/ipv6/seg6_iptunnel.c:161 seg6_do_srh+0x4b4/0xa44 net/ipv6/seg6_iptunnel.c:431 seg6_output_core+0x164/0x688 net/ipv6/seg6_iptunnel.c:681 seg6_output+0x44/0x1ac net/ipv6/seg6_iptunnel.c:744 lwtunnel_output+0x3d0/0x664 net/core/lwtunnel.c:356 dst_output include/net/dst.h:470 [inline] ip_local_out+0x110/0x148 net/ipv4/ip_output.c:131 iptunnel_xmit+0x50c/0xd38 net/ipv4/ip_tunnel_core.c:97 udp_tunnel_xmit_skb+0x220/0x348 net/ipv4/udp_tunnel_core.c:187 tipc_udp_xmit+0x75c/0x9c0 net/tipc/udp_media.c:202 tipc_udp_send_msg+0x214/0x374 net/tipc/udp_media.c:274 tipc_bearer_xmit_skb+0x260/0x3b0 net/tipc/bearer.c:576 tipc_enable_bearer net/tipc/bearer.c:366 [inline] __tipc_nl_bearer_enable+0xc90/0xfb0 net/tipc/bearer.c:1048 tipc_nl_bearer_enable+0x2c/0x48 net/tipc/bearer.c:1057 genl_family_rcv_msg_doit+0x1e4/0x2d4 net/netlink/genetlink.c:1114 Fixes: d58e468b1112 ("flow_dissector: implements flow dissector BPF hook") Reported-by: syzbot+9408fbe0e6452a12e9ab@syzkaller.appspotmail.com Closes: https://lore.kernel.org/netdev/6abac2eb.3654fce1.1bec97.0000.GAE@google.com/ Signed-off-by: Eric Dumazet <edumazet@kernel.org> Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Link: https://patch.msgid.link/20260928194524.3617299-1-edumazet@kernel.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
27 hoursnet: Revalidate queue config for ringparam changesBjörn Töpel
Memory-provider queue configuration is validated when the provider is bound. A later ethtool ring change may invalidate it because drivers can size queue memory from both ring depth and RX page size. The fbnic consumer is added in the following patch. Keep configured RX ring depths in netdev_config and stage proposed values in cfg_pending. Validate every RX queue before calling the driver. Each check validates the device defaults, then any queue memory-provider override. Commit the values only after the driver accepts them. Drivers which consume stored ring depths through queue configuration must initialize every RX depth before registering the netdev. Stored values override callback defaults, including when zero. The callback receives a rendered configuration rather than a queue ID. Validation should depend on the configuration, not queue identity. Checking defaults also covers the case where every queue has a memory-provider override. Drivers may normalize ring depths when applying them. Require the validation callback to use the same normalization. Drivers must report the applied depths through the ethtool_ringparam argument so the core records the result. Use the same transaction for ioctl and netlink. Drivers without ndo_validate_qcfg skip the new validation. Link: https://lore.kernel.org/all/20250421222827.283737-14-kuba@kernel.org/ Signed-off-by: Björn Töpel <bjorn@kernel.org> Reviewed-by: Simon Horman <horms@kernel.org> Reviewed-by: Joe Damato <joe@dama.to> Link: https://patch.msgid.link/20260925104417.2325213-4-bjorn@kernel.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
27 hoursnet: Add netdev_config helpersJakub Kicinski
netdev_config manipulation will become slightly more complicated soon and will be used by both ethtool and the queue API. Encapsulate the logic in helper functions. Signed-off-by: Björn Töpel <bjorn@kernel.org> Reviewed-by: Breno Leitao <leitao@debian.org> Reviewed-by: Simon Horman <horms@kernel.org> Reviewed-by: Joe Damato <joe@dama.to> Link: https://patch.msgid.link/20260925104417.2325213-2-bjorn@kernel.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
27 hourssctp: check RCV_SHUTDOWN after the sendmsg connect waitJun Yang
sctp_wait_for_connect() drops the socket lock while it sleeps. An out-of-the-blue ABORT can then be processed from the socket backlog and unlink the association. If a concurrent shutdown(fd, SHUT_RD) sets RCV_SHUTDOWN, the waiter breaks with err == 0 before checking asoc->base.dead. Its final sctp_association_put() can then free the association, leaving sctp_sendmsg_to_asoc() to continue with a dangling pointer. Check RCV_SHUTDOWN along with the wait error in sctp_sendmsg_to_asoc() before using the association again. The check only accesses the socket, so it needs no additional association reference. Return the existing -ESRCH so that sctp_sendmsg() skips freeing a new association that may already have been destroyed. Keep sctp_wait_for_connect() unchanged to preserve its behavior for the connect() caller. Fixes: 668c9beb9020 ("sctp: implement assign_number for sctp_stream_interleave") Cc: stable@vger.kernel.org Reported-by: TencentOS Corvus AI <corvus@tencent.com> Signed-off-by: Jun Yang <junvyyang@tencent.com> Acked-by: Xin Long <lucien.xin@gmail.com> Link: https://patch.msgid.link/20260926095606.68601-1-juny24602@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
27 hourstcp: remove mmap_lock fallback pathDave Hansen
Previously, the per-VMA locking could fail in the face of writers which necessitates a fallback to mmap_lock. The new vma_start_read_unlocked() will wait for writers instead of failing. Use the new helper. Wait for writers. Remove the fallback to mmap_lock. The fallback removal does not affect NOMMU case because TCP_ZEROCOPY is gated on CONFIG_MMU. This really is a nice cleanup. It removes the need to pass the lock state back and forth to find_tcp_vma(). Link: https://lore.kernel.org/20260831203056.838265-6-surenb@google.com Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com> Signed-off-by: Suren Baghdasaryan <surenb@google.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Acked-by: Lorenzo Stoakes <ljs@kernel.org> Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org> Tested-by: syzbot@syzkaller.appspotmail.com Cc: Liam R. Howlett <liam@infradead.org> Cc: Shakeel Butt <shakeel.butt@linux.dev> Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org> Cc: Arve Hjønnevåg <arve@android.com> Cc: Todd Kjos <tkjos@android.com> Cc: Christian Brauner <christian@brauner.io> Cc: Carlos Llamas <cmllamas@google.com> Cc: Alice Ryhl <aliceryhl@google.com> Cc: David S. Miller <davem@davemloft.net> Cc: David Ahern <dsahern@kernel.org> Cc: David Hildenbrand (Arm) <david@kernel.org>
27 hoursnet/smc: Serialize link activation with teardownChengfeng Ye
smc_llc_link_active() marks a link active before scheduling its testlink work. The first-link confirmation paths call it without llc_conf_mutex, allowing link-down processing to clear the link between these operations: Connection setup Link-down worker smc_llc_link_active() link->state = SMC_LNK_ACTIVE smcr_link_clear() link->clearing = 1 smc_llc_link_clear() cancel_delayed_work_sync() schedule_delayed_work() The connection reference keeps the link alive during activation, but the newly queued work outlives that reference. Later cleanup skips clearing an already-clearing link, leaving its timer armed when the link group is freed. KASAN reported: BUG: KASAN: use-after-free in __run_timers+0x723/0x8d0 Write of size 8 at addr ffff88810f108810 by task swapper/1/0 Call Trace: __run_timers+0x723/0x8d0 timer_expire_remote+0xd3/0x120 tmigr_handle_remote_up+0x4f4/0xab0 __walk_groups_from+0x40/0x150 tmigr_handle_remote+0x229/0x2c0 run_timer_softirq+0x1f5/0x250 Hold llc_conf_mutex around first-link activation on both sides so that link-down processing cannot cancel the work before it is queued. Also protect the client's initial optional add-link processing, which can activate a second link without the lock. The other add-link paths already hold llc_conf_mutex. This serializes every activation with link teardown without changing the handshake sequence or error handling. Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com> Link: https://patch.msgid.link/20260927074547.3694742-1-nicoyip.dev@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
28 hoursnetpoll: bound the deferred transmit queueZack Gomez
__netpoll_send_skb() parks an skb on npinfo->txq whenever the device cannot take it at once, and once the queue is non-empty every later skb goes straight there to keep ordering. queue_process() drains it from a workqueue and, unlike the direct path, never polls the device for completions: when the ring is stopped it backs off HZ/10. Nothing limits the queue length. A producer that outruns that drain therefore grows the queue until the host is out of memory. Observed with netconsole forwarding a GPU driver that logged one line at ~1e5/s after a firmware hang. The NIC was moving ~17k packets/s: completions for each burst surfaced tens of ms later, outside the one-tick window, so queue_process() slept HZ/10 per ring while ~1e5 lines/s kept arriving. The queue grew at ~170 MB/s, unreclaimable slab reached 51 GiB in five minutes and the OOM killer ran from kswapd with 341 MiB of anonymous memory on the whole box. What the queue held was the flood itself; the OOM report never left the host. Reproduced on the same host (netconsole over a 10G ConnectX-4 Lx) under the same slow-completion condition: 200k lines to /dev/kmsg in 0.12 s grew unreclaimable slab by 173 MiB, about 188k skbs, draining at ~8-10k packets/s. With prompt completions the same burst drains at line rate; a stall on the link while lines keep arriving faster than the drain reproduces the growth. Until the 2006 netpoll rework [1] the deferred path drained through dev_queue_xmit(), with the stack's own backpressure, and was capped at 16 skbs (MAX_QUEUE_DEPTH). That series moved it to a direct hard_start_xmit() with the HZ/10 back-off and made the queue per-device, dropping the cap on the way. Cap it at 1024 skbs per device and drop new skbs beyond that. The drop is counted in tx_dropped of the device whose queue is full and freed with SKB_DROP_REASON_FULL_RING. A netconsole target bound directly to that device also gets NET_XMIT_DROP and, with CONFIG_NETCONSOLE_DYNAMIC, counts it in xmit_drop_count. When a stacked device (bond, bridge, team, vlan, macvlan) passes the skb down and the lower device's queue is the one that fills, the return value does not reach netconsole and the lower device's tx_dropped is the record. The bound holds either way. Nothing is logged on the drop path because that would recurse into the console being drained. [1] https://lore.kernel.org/netdev/20061026225645.482978803@osdl.org/ Signed-off-by: Zack Gomez <zack.gomez@gmail.com> Reviewed-by: Breno Leitao <leitao@debian.org> Link: https://patch.msgid.link/20260925204537.2664119-1-zack.gomez@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
28 hoursseg6: reallocate the skb head on L2 encapsulation only when neededYuya Kusakabe
The L2 encapsulation modes of the seg6 lwtunnel reallocate the skb head on every packet, where the IPv6 encapsulation modes reallocate only when they have to. Ask for the whole encapsulation up front instead, so that the reallocation happens at most once and only when the headroom really is too small: skb->mac_len + sizeof(struct ipv6hdr) + ipv6_optlen(tinfo->srh) + dst_dev_overhead(cache_dst, skb) __seg6_do_srh_encap() then finds the room it needs and its own skb_cow_head() becomes a no-op. Drivers reserve more than that on the forwarding path, so the reallocation usually disappears altogether. A single-segment policy on ixgbe needs 14 (mac_len) + 40 (ipv6hdr) + 24 (SRH) + 16 (LL_RESERVED_SPACE) = 94 against the 206 bytes the driver leaves. Where the headroom is smaller, as on a veth pair, pskb_expand_head() is called once per forwarded packet instead of twice. Asking only for skb->mac_len would still take two whenever the skb is header-cloned, because the cow that unclones it does not also make room for the outer header. The cost is amplified by CONFIG_INIT_ON_ALLOC_DEFAULT_ON, which many distributions enable: every new head is zeroed in full, and that memset alone accounts for 16% of the datapath profile. Throughput at 0.5% packet loss, 64-byte frames forwarded through one 2.30 GHz core (Xeon E5-2650 v3, ixgbe 82599ES), offered by TRex and binary-searched over 10 runs of 10 s: Before: 654.6 kpps After: 965.7 kpps Signed-off-by: Yuya Kusakabe <yuya.kusakabe@gmail.com> Reviewed-by: Eric Dumazet <edumazet@google.com> Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn> Link: https://patch.msgid.link/20260925-seg6-l2cow-v3-1-fc83821542a7@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
30 hoursMerge tag 'wireless-next-2026-09-30' of ↵Jakub Kicinski
https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless-next Johannes Berg says: ==================== More features: - ath10k: NVMEM device tree bindings - ath12k: QMI firmware alignments - mm81x: AP improvements - mac80211: - CIP (control frame integrity) support - NAN improvements - cfg80211: - improvements for AP regulatory checks * tag 'wireless-next-2026-09-30' of https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless-next: (67 commits) wifi: mac80211: Gracefully deauthenticate on association timeout wifi: libertas: fix RX OOB access from device-controlled pkt_ptr wifi: mwifiex: Reattach interfaces on suspend failure wifi: b43legacy: work around stack frame size warning wifi: cw1200: fix link_id_db OOB access via device-controlled link ID wifi: mwifiex: bound SDIO fw dump count by memory table size wifi: mac80211: start next ROC after purging an interface wifi: mac80211: fix potential ack-skb leak on error path wifi: mac80211: mesh: don't send peering close in listen wifi: mac80211_hwsim: Support NAN deferred schedule completion wifi: mac80211: Restrict probe request rates for minimal content wifi: nl80211: allow a NAN peer schedule with 2 channels in the same slot wifi: cfg80211: use sysfs_emit_at() in addresses_show wifi: cfg80211: require zero terminator in valid_regdb country table wifi: cfg80211: fix NAN local schedule update ordering and allocation wifi: mac80211: Fix a race when expiring a mesh path wifi: cfg80211: validate monitor channel set against radio usage wifi: nl80211: reject color-change requests that change 6 GHz power type wifi: nl80211: defer AP beacon regulatory check to start_ap wifi: radiotap: add definitions for UHR U-SIG ... ==================== Link: https://patch.msgid.link/20260930124547.228697-29-johannes@sipsolutions.net Signed-off-by: Jakub Kicinski <kuba@kernel.org>
30 hoursMerge tag 'wireless-2026-09-30' of ↵Jakub Kicinski
https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless Johannes Berg says: ==================== Still more fixes coming in, notably: - ath11k: avoid running out of stations on HW restart - mac80211: - drop too large fragmented MPDUs - mesh path handling fixes - validation improvements - reject CSA with bad 320 MHz bandwidth - cfg80211: fix RTS for single radio devices * tag 'wireless-2026-09-30' of https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless: (27 commits) wifi: mac80211: fix slab-out-of-bounds read in ieee80211_monitor_select_queue() wifi: mac80211: reject invalid 320 MHz CSA bandwidth wifi: mac80211: set info->band for 802.3 encap offload frames wifi: mac80211: prevent AP VLAN tx from other interfaces wifi: cfg80211: preserve hidden-group beacon IE ownership wifi: mac80211: keep fallback association elements alive wifi: mac80211: shut down RX BA session timer on teardown wifi: mac80211: validate TX status rate metadata wifi: ath9k_htc: bound TX aggregation to MAX_TX_BUF_SIZE wifi: ath9k: reject short WMI command responses wifi: ath9k: Clean up device initialisation guards wifi: ath11k: reset ar->num_stations on hardware start wifi: cfg80211: fix RTS threshold setting for single-radio PHY wifi: mac80211: handle empty FILS association request payload wifi: mac80211: minstrel_ht: validate fixed rate index wifi: p54: validate firmware record lengths wifi: mac80211: fix mesh fast xmit path deletion UAF wifi: mac80211: drop oversized fragments to avoid extra_len overflow wifi: wlcore: Fix runtime PM leak in wlcore_remove() wifi: mac80211: drain PS delivery work during station teardown ... ==================== Link: https://patch.msgid.link/20260930124440.224799-3-johannes@sipsolutions.net Signed-off-by: Jakub Kicinski <kuba@kernel.org>
30 hourstcp: preserve timestamps across receive queue collapseJason Xing
When tcp_collapse() rebuilds skbs under memory pressure, the copy process doesn't include the right tstamp and hwtstamp from the old skb. And memcpy(nskb->cb, skb->cb, ...) copies has_rxtstamp, but nskb->tstamp and hwtstamps are left at zero, so tcp_recv_timestamp() ends up emitting no cmsg at all. In net timestamping case, if such an skb happens to be the last one consumed in a recvmsg() call, the application receives no RX timestamp for that call. Fix this by copying both tstamp and hwtstamp of the last skb to the new skb, matching tcp_try_coalesce()/tcp_add_backlog(). Note that the has_rxtstamp flag can still be inherited through the cb memcpy from an skb that contributes no bytes (fully covered skb left in the ofo tree by the tcp_ooo_try_coalesce() -> coalesce_done path), so set TCP_SKB_CB(nskb)->has_rxtstamp to false which makes the new block the only place setting it. Signed-off-by: Jason Xing <kerneljasonxing@gmail.com> Reviewed-by: Eric Dumazet <edumazet@kernel.org> Link: https://patch.msgid.link/20260924152529.5689-1-kerneljasonxing@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
44 hoursnetfilter: flowtable: restore ieee80211 forward pathPablo Neira Ayuso
Before commit 871df5007eda ("netfilter: flowtable: bail out if forward path cannot be discovered"), there was a fallback to set up a forward path in case .ndo_fill_forward_path fails or DEV_PATH_MTK_WDMA was used. Such fallback was used by commit d787a3e38f01 ("mac80211: add support for .ndo_fill_forward_path"). One possibility is to handle DEV_PATH_MTK_WDMA from the flowtable forward path discovery. However, this is only used internally by drivers to retrieve mtk_wdma information to set up hardware offload. Felix decided to use the .fill_forward_path interface for this purpose due to the lack of a better interface at that time. Add a new DEV_PATH_IEEE80211 path which is offered if the new ieee80211 flag is set on in the struct net_device_path_ctx to restore the flowtable with a ieee80211 netdevice. Handle this new DEV_PATH_IEEE80211 path just like DEV_PATH_ETHERNET and DEV_PATH_DSA, ie. this is the last netdevice in the stack. This new ieee80211 flag is implicitly unset for mtk_ppe and airoha which call dev_fill_forward_path() to retrieve a DEV_PATH_MTK_WDMA path. Fixes: 871df5007eda ("netfilter: flowtable: bail out if forward path cannot be discovered") Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
44 hoursnetfilter: flowtable: generalize pending status bitPablo Neira Ayuso
Rename NF_FLOW_HW_PENDING to NF_FLOW_PENDING and use it to inhibit the flowtable GC worker until pending hw offload work has been completed. Apparently, nf_flow_offload_stats() can schedule work to retrieve stats while the flow is being removed by GC. And this bit can also be used in a follow up patch to disable GC until the flow has been fully added in both directions. Revert the reordering done in commit d644b23afe1e ("netfilter: flowtable: publish HW_DEAD after worker is done") to prevent a race between GC and hw offload handler. Fixes: 2c8897953f3b ("netfilter: flowtable: Add pending bit for offload work") Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
44 hoursnetfilter: bpf: reject invalid NAT manipulation typesFernando Fernandez Mancera
As bpf_ct_set_nat_info() is not validating the NAT manipulation type a wrong value can be passed directly to nf_nat_setup_info(). This triggers the WARN_ON() at nf_nat_setup_info() and if panic_on_warn isn't set, then IPS_SRC_NAT_DONE is set without adding nat_bysource and conntrack cleanup tries to unlink an uninitialized hlist node. Fix this by checking that NAT manipulation type is correct before calling nf_nat_setup_info(). In addition, if the WARN_ON is hit, return NF_DROP instead of continuing with the processing to avoid similar situations in the future. Reported-by: VEGA <vega@nebusec.ai> Fixes: 0fabd2aa199f ("net: netfilter: add bpf_ct_set_nat_info kfunc helper") Signed-off-by: Fernando Fernandez Mancera <fmancera@suse.de> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
44 hoursnetfilter: nft_set_rbtree: skip transaction elements during GCWeiming Shi
Since nft_set_commit_update() runs set commit callbacks before processing NEWSETELEM transactions, nft_rbtree_gc_scan() can observe elements added by the transaction being committed. The scan records an interval end in rbe_end without checking the element's transaction state. A later, unrelated expired start then moves both elements to the expired list. The synchronous GC queue can free the new end element before the transaction subsequently activates it, causing a use-after-free. Only consider elements that are fully active in both generations. This keeps transaction-state elements out of the GC scan and preserves interval pairing across skipped elements. KASAN reports: BUG: KASAN: slab-use-after-free in nft_setelem_activate nft_setelem_activate net/netfilter/nf_tables_api.c:7047 nf_tables_commit net/netfilter/nf_tables_api.c:11137 Allocated by task 130: nft_set_elem_init net/netfilter/nf_tables_api.c:6794 nft_add_set_elem net/netfilter/nf_tables_api.c:7523 Freed by task 130: nft_trans_gc_trans_free net/netfilter/nf_tables_api.c:10506 rcu_core kernel/rcu/tree.c:2919 Fixes: 1e3b9e1c77fe ("netfilter: nf_tables: call set ops .commit when building new ruleset blob") Reported-by: <co+ee5e50ef2670e5f4@bugs.sh> Assisted-by: LLM Signed-off-by: Weiming Shi <bestswngs@gmail.com> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
44 hoursipvs: filter some flags received in the backup serverJulian Anastasov
While the IPVS SYNC protocol is not secure by design we can still protect the backup server from messages that can wreak havoc. This commit addresses problems from received connection flags or their combinations. We now drop messages as follows: 1. the NO_CPORT+TEMPLATE combination allows lookups for normal connections to hit template which can break in many ways. While the master does not sync connections with NO_CPORT flag, i.e. before they are established, we still accept NO_CPORT without TEMPLATE. 2. ONE_PACKET: it is not sent by master, so we do not expect it in backup. Before now it was ignored by IP_VS_CONN_F_BACKUP_MASK for protocol v1 while protocol v0 created connections that are not hashed and dropped immediately. Better to apply the IP_VS_CONN_F_BACKUP_MASK also to the flags from v0 messages for consistency with v1. Fixes: 87375ab47cd0 ("[IPVS]: ip_vs_ftp breaks connections using persistence") Signed-off-by: Julian Anastasov <ja@ssi.bg> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
44 hoursipvs: do not create invisible templatesJulian Anastasov
The IP_VS_CONN_F_ONE_PACKET flag was implemented for normal connections. When conn template inherits this flag from dest->conn_flags it will not be hashed. As result, we will create new template for every new normal connection. Fix it to allow one template to be used by many normal connections. Fixes: 26ec037f9841 ("IPVS: one-packet scheduling") Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260916231652.127456-1-pablo%40netfilter.org Signed-off-by: Julian Anastasov <ja@ssi.bg> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
44 hoursipvs: bound LBLCR and LBLC cache growthZhiling Zou
ip_vs_lblcr_new() and ip_vs_lblc_new() create cache entries for every previously unseen destination address. The table max_size only tells the periodic collector to reclaim entries after the cache has already exceeded the limit. It does not reclaim entries that the attacker continues to use. Reject new cache entries once either table reaches max_size * 3 / 2. The extra headroom lets the periodic collector catch up while the existing scheduler fallback continues to use the selected destination when cache creation fails. New traffic therefore stays serviceable without growing the tables further. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Suggested-by: Julian Anastasov <ja@ssi.bg> Signed-off-by: Zhiling Zou <zhilinz@nebusec.ai> Acked-by: Julian Anastasov <ja@ssi.bg> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
44 hoursipvs: fix missing counter decrement in lblcJulian Anastasov
LBLC may delete cache entries for destinations that are removed or overloaded and replace them with available ones. But ip_vs_lblc_new() forgets to decrement the tbl->entries counter after calling ip_vs_lblc_del(). This can lead to increased shrinking of the cache with every new garbage collection. Fixes: 2f3d771a35fe ("ipvs: do not use dest after ip_vs_dest_put in LBLC") Link: https://sashiko.dev/#/patchset/0bdd5abe9968ded7ca2b9cb6844ba83d94cc8d53.1787318053.git.zhilinz%40nebusec.ai Signed-off-by: Julian Anastasov <ja@ssi.bg> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
44 hoursnetfilter: nft_flow_offload: drop flowtable reference on init error pathAohan Mei
nft_flow_offload_init() bumps the flowtable use count with nft_use_inc() before calling nf_ct_netns_get(). When the latter fails, the error is returned as-is and the reference is leaked. The upper layers do not balance it either: nf_tables_newexpr() clears expr->ops when the expression init callback fails, so the nft_expr_more() iteration in nft_rule_expr_deactivate() and nf_tables_rule_destroy() stops right before the failed expression and its ->destroy callback, which would drop the reference, never runs. Each failed rule addition therefore leaks one flowtable reference and the flowtable can no longer be removed: NFT_MSG_DELFLOWTABLE keeps reporting -EBUSY even though no rule references it. Save the nf_ct_netns_get() return value and undo the nft_use_inc() when it fails, restoring the inc/dec pairing within nft_flow_offload_init() itself. Fixes: a3c90f7a2323 ("netfilter: nf_tables: flow offload expression") Reported-by: TencentOS Corvus AI <corvus@tencent.com> Cc: stable@vger.kernel.org Assisted-by: CodeBuddy:Kimi-K3 Signed-off-by: Aohan Mei <henrymei@tencent.com> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
44 hoursMerge branch into tip/master: 'timers/core'Ingo Molnar
# New commits in timers/core: 348f54c435bf ("selftests/timers: clocksource-switch: Fix unchecked open()/read()") b03638013add ("selftests: timers: Measure the CPU timers on the clock they count") a252cb93e156 ("selftests: timers: Count what tick is worth in the drift estimate") 3945c4a3ea15 ("time/kunit: Add time64_to_tm() case beyond 32-bit day count") 2927f7ca7844 ("time: Prevent time64_to_tm() day truncation on 32-bit") bc5b66c300b8 ("timekeeping: Use READ_ONCE/WRITE_ONCE() for ktime_sec to prevent tearing") bb41ece16463 ("timers/migration: Mark racy updates to tmigr_event::ignore field") 1159ad0a6aaa ("timers: Mark racy updates to hlist_node::pprev field") 5dffe33bfdaf ("hrtimer: Apply READ_ONCE() to lockless base->running loads") ad80926d7803 ("hrtimer: Mark the hrtimer_sleeper structure's ->task field __private") 15f84398330c ("hrtimer: Update hrtimer_resolution only if value changes") 7eed1771a6e8 ("futex: Use accessor for hrtimer_sleeper ->task field in requeue") 9cd2f4e304d9 ("rtmutex: Use accessor for hrtimer_sleeper ->task field") 3990d196954a ("net: pktgen: Use accessor for hrtimer_sleeper ->task field") 9b7bcd671e9c ("timers: Use accessor for hrtimer_sleeper ->task field in sleep_timeout.c") fa3d486086d6 ("futex: Use accessor for hrtimer_sleeper ->task field in waitwake.c") 851c30277918 ("io-uring/rw: Use accessor for hrtimer_sleeper ->task field") 66b29025b676 ("wait: Use accessor for hrtimer_sleeper ->task field") cefc1a24ace5 ("aio: Use accessor for hrtimer_sleeper ->task field") d166a1cd9016 ("hrtimer: Mark data-racy accesses to hrtimer_sleeper ->task field") 7b07d15ed1d7 ("posix-timers: Handle exit in do_exit() completely") 54ad1e0ea42c ("posix-cpu-timers: Prevent enqueueing when PF_EXITING is set") 760ad335d600 ("posix-cpu-timers: Use PF_EXITING to indicate exit") 8e9ed3b66f40 ("posix-cpu-timers: Move inlines out of public header") 24adce0b86c9 ("posix-timers: Move POSIX timer group exit related code out of do_exit()") 7117334f022b ("posix-timers: Move posixtimer_exec_cleanup() out of exec.c") Signed-off-by: Ingo Molnar <mingo@kernel.org>
44 hourstcp: refresh TS.Recent for accepted old ACKsJeff Jo
A TCP packet can carry new data while acknowledging traffic in the opposite direction. With overlapping traffic in both directions, a delayed packet's acknowledgment can be older than one Linux has already accepted, even when that packet fills a gap in the received data. Linux accepts the data, but tcp_ack() takes the old_ack path and skips updating TS.Recent, the timestamp saved for outgoing acknowledgments. The reply therefore echoes an older timestamp. If the sender uses this echo to measure round-trip time after a long idle period, its estimate includes the idle time and can reduce its sending rate. Update TS.Recent in old_ack using tcp_replace_ts_recent(), before SACK processing can trigger a transmission. This reuses the existing timestamp and sequence checks, including PAWS protection against old duplicate packets. ACK validation already rejects old ACKs in SYN_RECV before this path, so no additional state check is needed. Echoing the timestamp of the packet that fills the receive gap follows RFC 7323 section 4.3. In a socket reproduction with 300 seconds idle, controlled reordering and retransmission to exercise timestamp-based RTT sampling, the sender's smoothed round-trip time was 37.5 seconds without the fix and 15.5 ms with it. Fixes: 12fb3dd9dc3c ("tcp: call tcp_replace_ts_recent() from tcp_ack()") Assisted-by: LLM sparse Signed-off-by: Jeff Jo <jeffjo@openai.com> Reviewed-by: Eric Dumazet <edumazet@google.com> Signed-off-by: David S. Miller <davem@davemloft.net>
2 daysmptcp: shrink struct mptcp_options_receivedQuanye Yang
struct mptcp_options_received is allocated on the stack while parsing incoming MPTCP options. Several suboptions are mutually exclusive, as enforced by mptcp_parse_option(), so their payloads can overlap. Group fields by suboption and place the mutually exclusive payloads in an anonymous union. Keep DSS and rm_list outside the union: they can be combined with other suboptions. Move join_id into the MP_JOIN group, and overlap token, thmac and hmac inside that group. Further shrinking would require changing the parser so currently coexisting fields (DSS mapping vs ACK, rm_list, status flags) can overlap. That adds complexity for little gain, since the outer union is already dominated by the MP_JOIN / ADD_ADDR members. This reduces the structure size from 136 to 72 bytes on x86_64. Closes: https://github.com/multipath-tcp/mptcp_net-next/issues/625 Signed-off-by: Quanye Yang <quanyeyang@proton.me> Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org> Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org> Link: https://patch.msgid.link/20260926-net-next-mptcp-misc-feat-7-4-v1-3-67af4ab37406@kernel.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2 daysmptcp: split FASTCLOSE key from rcvr_keyQuanye Yang
MP_CAPABLE and MP_FASTCLOSE both stored their key in rcvr_key. Give FASTCLOSE a dedicated fc_recv_key overlapped with rcvr_key in a union. This helps shrink struct mptcp_options_received. Signed-off-by: Quanye Yang <quanyeyang@proton.me> Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org> Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org> Link: https://patch.msgid.link/20260926-net-next-mptcp-misc-feat-7-4-v1-2-67af4ab37406@kernel.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2 daysmptcp: remove thmac from subflow ctxMatthieu Baerts (NGI0)
This entry is only used in subflow_finish_connect(). Instead, use the original value from mp_opt, and pass it to subflow_thmac_valid() to do the validation with the given truncated hmac. While at it, rename the variables in subflow_thmac_valid() to avoid confusions about the received one vs the expected one. Reviewed-by: Geliang Tang <geliang@kernel.org> Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org> Link: https://patch.msgid.link/20260926-net-next-mptcp-misc-feat-7-4-v1-1-67af4ab37406@kernel.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2 daysnet: extend IPv6 exthdr detection of tunneled packetsWillem de Bruijn
Commit c4336a07eb6b ("net: correctly handle tunneled traffic on IPV6_CSUM GSO fallback") split skb_gso_has_extension_hdr() into mutually exclusive branches on skb->encapsulation. This did not yet address all paths: 1. With skb->encapsulation set, the outer header is not checked. IPv6 tunnels such as ip6_gre and ip6_tunnel add an outer Destination Options header by default (encap_limit). Their GSO packets skip software GSO, then skb_csum_hwoffload_help() sees the outer extension header and calls skb_checksum_help() on the GSO skb, which warns and drops it. 2. Tunnels over IPv6 without an outer transport header, such as ip6_tunnel, leave skb->transport_header at the inner transport header. skb_network_header_len() then spans the outer IPv6, tunnel and inner IP headers, a false positive. 3. UDP tunnels without an inner network header, such as SCTP-in-UDP or PSP, have no inner IPv6 header to check. Decide on the outer header alone. 4. Directly dereferencing inner_ip_hdr(skb)->version without skb_header_pointer() is unsafe. Instead, check ipv6_ext_hdr(nexthdr) on the outer IPv6 header and, if set, on the inner IPv6 header. Read the headers with skb_header_pointer(). Keep the skb_network_header_len() check when !skb->encapsulation to also catch encapsulated packets without skb->encapsulation (e.g., virtio). Use the same helper in skb_csum_hwoffload_help(). Its open coded test has the false positive of (2) and ignores the inner header. Background: checksum offload of tunneled packets invariants: Non-GSO skb: - If the inner packet is CHECKSUM_PARTIAL, Local Checksum Offload computes the outer checksum in software and the device offloads only the inner L4 checksum. - If the inner packet is CHECKSUM_NONE (e.g., SCTP-in-UDP, ESP-in-UDP, Remote Checksum Offload), the device offloads the outer UDP or GRE checksum instead. GSO skb: - A device with NETIF_F_GSO_UDP_TUNNEL_CSUM or NETIF_F_GSO_GRE_CSUM computes both inner and outer checksums per segment. The outer checksum is seeded with only the pseudo-header checksum. A NETIF_F_IPV6_CSUM device must parse through the outer headers to reach the inner ones. Fixes: c4336a07eb6b ("net: correctly handle tunneled traffic on IPV6_CSUM GSO fallback") Cc: stable@vger.kernel.org Signed-off-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/20260926140506.2335137-1-willemdebruijn.kernel@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2 daysnet: cap skb->queue_mapping when the tx queue is pickedJamal Hadi Salim
skbedit can set skb->queue_mapping and raise the per-CPU skip_txqueue flag so __dev_queue_xmit() honours the mapping. __dev_queue_xmit() cleared the flag before sch_handle_egress() and only read it afterwards, so the flag was not confined to the xmit that set it: a nested xmit (mirred redirect or mirror, or a drop after skbedit) could set the flag and the outer xmit would consume it for an skb that never went through skbedit. A forwarded packet still carries the ingress NIC's rx_queue + 1 in skb->queue_mapping, so the outer device then indexes its tx queue state with that stale value. Taprio's child array q->qdiscs[] is sized to the device's queue count, so taprio_enqueue() indexes past its allocation and dereferences the result as a struct Qdisc *. We (ab)use the skb->nf_skip_egress which means "skip netfilter egress for this packet" to tag to "am I in tc egress?". Despite the overload I dont see it as a conflict since the marker is set only around the single sch_handle_egress() call and ingress path is guarded by tc_at_ingress. I will send a followup(net-next) patch once this hits net-next to rename the skb->nf_skip_egress bit/flag to skb->skip_egress Arm the flag only from the egress classifier that can use it: raise skip_txqueue from tcf_skbedit_act() only when it runs inside sch_handle_egress(), thanks to skb->nf_skip_egress. An egress qdisc classifier runs in q->enqueue(), after the tx queue has been picked, so a mapping it sets cannot affect the current packet; arming the flag there only pollutes it for a later xmit. Then own the flag for the xmit frame the egress hook runs in: save the incoming value and clear it just before sch_handle_egress(), and restore it after the hook - on the consumed (drop) path, or, in the same call that reads it, on the surviving path. The save and the restores stay inside the egress_needed_key static branch, so a packet pays for them only when egress hooks are active (2f1e85b1aee4). Store the value netdev_cap_txqueue() selected back into skb->queue_mapping in netdev_tx_queue_mapping(), as netdev_core_pick_tx() already does, so the skip_txqueue path never hands a later reader on the xmit path a mapping the device cannot serve. A store made still later in the same frame, by a tc BPF program attached to a transmit qdisc, is outside this path and is not re-capped; a separate followup will resolve that path. netdev_xmit_skip_txqueue() returns the previous flag value so the save-and-clear is one call, and a no-op stub is provided when CONFIG_NET_EGRESS is disabled. skb->nf_skip_egress is compiled under CONFIG_NET_EGRESS rather than CONFIG_NETFILTER_SKIP_EGRESS, so skb_at_tc_egress() is valid whenever the egress path is built. A local user in a network namespace can redirect a packet from a device with more TX queues to one with fewer after setting a mapping valid only on the larger device. That reaches these reads and, under KASAN, faults with "slab-out-of-bounds in taprio_enqueue". Conditions to recreate the bug: the report's own trigger is a local user with CAP_NET_ADMIN in a network namespace, so no eBPF program is needed. With CONFIG_NET_SCH_TAPRIO=y, CONFIG_NET_ACT_SKBEDIT=y, CONFIG_NET_ACT_MIRRED=y, CONFIG_NET_CLS_MATCHALL=y, CONFIG_NET_SCH_PRIO=y and KASAN enabled, create qa (3 queues), qb (2 queues) and qc (1 queue) as dummy devices; put a taprio root on qb (num_tc 1, queues 2@0) and clsact on all three; then add an egress matchall filter on every device. On qa: "action skbedit queue_mapping 2 pipe action mirred egress redirect dev qb". On qb: "action mirred egress mirror dev qc". On qc: "action skbedit queue_mapping 0 pipe". Send one packet out qa. qc's skbedit sets the flag while qb's outer xmit is in flight; without the fix qb consumes it and reads its two-entry taprio child array with the forwarded packet's stale mapping. A qc whose skbedit is instead installed in a transmit-qdisc classifier (a matchall filter on the qc root qdisc) reaches the same code path the same way without the fix. Testing: on a KASAN build with panic_on_warn=1 the unfixed kernel panics with "BUG: KASAN: slab-out-of-bounds in taprio_enqueue", a read 0 bytes past a 16-byte taprio_init() allocation, for the clsact-setter and the transmit-qdisc-classifier reproducers and for a clsact skbedit-then-tc-BPF store; the fixed kernel runs all three with no report, and the BPF store variant additionally shows the expected "selects TX queue" clamp notice from the write-back. Fixes: 2f1e85b1aee4 ("net: sched: use queue_mapping to pick tx queue") Reported-by: Zero Day Initiative <zdi-disclosures@trendmicro.com> Link: https://lore.kernel.org/netdev/CANn89iLwYx8nCVf0pCEk_MmEiyC6kQaMwCQT9WkQVeeNzNQHqQ@mail.gmail.com/ Link: https://lore.kernel.org/netdev/179008581937.2160803.7117814290574262942@kernel.org/ Link: https://lore.kernel.org/netdev/179033713973.2160803.4914570693994398206@kernel.org/ Link: https://lore.kernel.org/netdev/20260925180407.63647514@kernel.org/ Link: https://lore.kernel.org/netdev/CANn89i+k-mZKDQVtvws_MEXeuMTAdaCcOXFZE-RfhcGTu90sjA@mail.gmail.com/ Suggested-by: Eric Dumazet <edumazet@google.com> Suggested-by: Jakub Kicinski <kuba@kernel.org> Tested-by: hybris <hybris@mojatatu.ai> Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com> Reviewed-by: Eric Dumazet <edumazet@google.com> Link: https://patch.msgid.link/QDISC-9R8V.v4.20260928081529@mojatatu.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2 daysipv6: update NUD_FAILED neighbors from NA messagesLawrence Lee
Transition a FAILED neighbor entry to STALE upon receipt of an NA message on routers when accept_untracked_na is enabled. This extends the RFC 9131 accept_untracked_na behavior so that FAILED entries are treated the same as non-existent entries. RFC 4861 section 7.3.3 says that an entry should be deleted when address resolution fails. Linux instead retains the entry in NUD_FAILED, so treating it as untracked is consistent with the protocol model. Trying to resolve FAILED neighbors via periodic probing (e.g. using NTF_EXT_MANAGED) is more work compared to this approach which uses information in NAs that the kernel may already be receiving. Note that because this behavior in IPv6 is dependent on the accept_untracked_na sysctl setting, this approach is more conservative than IPv4 which transitions FAILED neighbors to STALE by default upon receiving GARPs. Link: https://lore.kernel.org/r/20260813233344.445265-1-lfqlee314@gmail.com Signed-off-by: Lawrence Lee <lfqlee314@gmail.com> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Link: https://patch.msgid.link/6be400e601f4ffa59e47ddb5daa33ecdd85dd88b.1790127207.git.lfqlee314@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>