| Age | Commit message (Collapse) | Author |
|
# Conflicts:
# drivers/gpu/drm/amd/amdkfd/kfd_migrate.c
# net/ceph/osd_client.c
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git
# Conflicts:
# Documentation/scheduler/index.rst
# arch/arm64/configs/defconfig
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/modules/linux.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth-next.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next.git
# Conflicts:
# arch/arm64/net/bpf_jit_comp.c
# arch/x86/net/bpf_jit_comp.c
# mm/internal.h
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next.git
# Conflicts:
# drivers/net/ethernet/realtek/r8169_main.c
# net/mac80211/ieee80211_i.h
# net/mac80211/tx.c
|
|
# Conflicts:
# fs/coredump.c
# fs/f2fs/f2fs.h
# fs/fuse/dax.c
# fs/xfs/libxfs/xfs_btree.c
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/mm/linux.git
# Conflicts:
# arch/arm64/kvm/mmu.c
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/klassert/ipsec.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs.git
# Conflicts:
# fs/smb/server/smb2pdu.c
# fs/smb/server/vfs.c
# fs/smb/server/vfs.h
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux
|
|
|
|
seg6_hmac_validate_skb() derived the SRH from skb_transport_header().
That only holds while the two coincide, which is not true on the
seg6_local input path.
ip6_rcv_core() leaves the transport header just past the IPv6 header.
A Hop-by-Hop options header is consumed before the route lookup and
advances it, but a Destination Options header is not: the seg6_local
lwtunnel is entered through an input redirect from the route lookup,
which bypasses the extension header handlers. The transport header
then still points at the Destination Options header while
seg6_get_srh() has located the real SRH further down the chain.
The HMAC is therefore computed over the Destination Options header,
and a packet carrying a valid HMAC TLV is dropped when
seg6_require_hmac is set. Such a packet is legitimate: RFC 8200 allows
Destination Options before a routing header, and get_srh() has walked
the header chain since commit 5829d70b0b6c ("ipv6: sr: fix get_srh() to
comply with IPv6 standard "RFC 8200"").
With a Fragment or an Authentication header in front of the SRH, the
same mistake also reads past the data pulled by seg6_get_srh(). This
happens before seg6_require_hmac is read, so the default configuration
is affected.
Reproduce by giving a node a seg6local End SID with
net.ipv6.conf.<dev>.seg6_require_hmac=1 and a key installed with
"ip sr hmac set <keyid> sha1", then sending
IPv6 -> Destination Options -> SRH (carrying a valid HMAC TLV) -> payload
to that SID: it is dropped, while the same packet without the
Destination Options header passes.
Fixes: 5829d70b0b6c ("ipv6: sr: fix get_srh() to comply with IPv6 standard "RFC 8200"")
Assisted-by: LLM
Signed-off-by: Yuya Kusakabe <yuya.kusakabe@gmail.com>
Reviewed-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Link: https://patch.msgid.link/20260926-b4-seg6-hmac-transport-header-v2-1-8087b76f6375@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
Commit b6f74dff6d26 ("net: use two lockdep classes for the netdev
instance lock") put every device without a parent in the virtual class.
It also restricted the locking order for the virtual class, because
queue leasing has hard requirements on the exact order.
This bites us back on bond, which is "virtual" and needs to be taken
before taking the locks of the lowers. NIPA hit the following on the
new test I recently posted for XDP+bond:
WARNING: possible circular locking dependency detected
------------------------------------------------------
python3/18635 is trying to acquire lock:
ff11000120b8ce30 (&dev->lock){+.+.}-{4:4}, at:
netdev_put_lock+0x2d/0x1a0
but task is already holding lock:
ff110001ef77ae30 (&netdev_virt_instance_lock_key){+.+.}-{4:4}, at:
netdev_put_lock+0x2d/0x1a0
which lock already depends on the new lock.
the existing dependency chain (in reverse order) is:
-> #1 (&netdev_virt_instance_lock_key){+.+.}-{4:4}:
__mutex_lock+0x1ae/0x1f10
xdp_set_features_flag+0x2b/0x50
bond_xdp_set_features+0x1eb/0x360
bond_netdev_event+0x13f/0x300
notifier_call_chain+0xae/0x300
call_netdevice_notifiers+0x70/0xa0
bnxt_xdp_set+0x2f6/0x620
netif_xdp_propagate+0x503/0xc60
dev_xdp_propagate+0xa1/0x230
bond_xdp_set+0x234/0x700
dev_xdp_install+0x592/0xd70
dev_xdp_attach+0x355/0xf50
dev_change_xdp_fd+0x176/0x210
do_setlink.isra.0+0x220d/0x2b20
rtnl_newlink+0x9f1/0x11b0
-> #0 (&dev->lock){+.+.}-{4:4}:
__mutex_lock+0x1ae/0x1f10
netdev_put_lock+0x2d/0x1a0
netdev_nl_queue_create_doit+0x801/0x1a70
genl_family_rcv_msg_doit+0x206/0x300
Possible unsafe locking scenario:
CPU0 CPU1
---- ----
lock(&netdev_virt_instance_lock_key);
lock(&dev->lock);
lock(&netdev_virt_instance_lock_key);
lock(&dev->lock);
Let's narrow down the "virtual" class to only the devices which
can actually create a queue. More LoC and complexity, but that
is what we actually care about here. The rest needs to nest
under rtnl_lock, which bond does (famous last words?)
Take the instance locks in two passes, first the netkits then
the rest (matching the queue leasing order).
An alternative would be to make sure the close list is sorted
correctly from the start (queue head/tail appropriately in
unregister_netdevice_queue()). I think it works but feels
a little more fragile. Happy to change, tho.
netdev_can_create_queue() will now be used on paths where we
genuinely handle non-netkit, so we can't always set the extack.
Unfortunately, the (recently) added tracepoint in extack fires
even when extack is NULL.
Fixes: b6f74dff6d26 ("net: use two lockdep classes for the netdev instance lock")
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org>
Acked-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://patch.msgid.link/20260929191543.3295633-1-kuba@kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf
Pablo Neira Ayuso says:
====================
Netfilter/IPVS fixes for net
The following batch contains Netfilter fixes for net. This batch
fixes crashes as recent feature regression, one of the due to a
dependency that has been pulled into -stable:
1) Expand existing ipset fix for bitmap sets to disallow comments
updates from kernel-side adds, from Florian Westphal.
2) Drop flowtable reference if nf_ct_netns_get() fails, otherwise
flowtable cannot ever be removed, from Aohan Mei.
3) nft_rbtree GC should collect end elements that contained in
this transaction batch, new or deleted elements are never
expired. From Weiming Shi.
4) Restrict nf_nat_bpf so it does not set unknown NF_NAT_MANIP_*
values, from Fernando F. Mancera.
5) Flowtable GC must skip flows that are pending hardware updates,
generalize the PENDING flag and use it to inhibit GC.
6) Restore flowtable with ieee80211 which broke due to a relatively
recent commit, which was pulled in by -stable, causing a regression
in 6.18 kernels.
And the following IPVS fixes:
1) Fix accounting of cache entries in IPVS LBLC for destinations,
which eventually fills up the table and trigger recurrent
resizing, from Julian Anastasov.
2) Limit IPVS cache growth for LBLCR and LBLC schedulers,
from Zhiling Zou.
3) Restrict IP_VS_CONN_F_ONE_PACKET for normal connections,
do not allow to use it with templates. Also from Julian.
4) Sanitize flags in IPVS sync messages received in the backup.
From Julian Anastasov.
netfilter pull request 26-09-30
* tag 'nf-26-09-30' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf:
netfilter: flowtable: restore ieee80211 forward path
netfilter: flowtable: generalize pending status bit
netfilter: bpf: reject invalid NAT manipulation types
netfilter: nft_set_rbtree: skip transaction elements during GC
ipvs: filter some flags received in the backup server
ipvs: do not create invisible templates
ipvs: bound LBLCR and LBLC cache growth
ipvs: fix missing counter decrement in lblc
netfilter: nft_flow_offload: drop flowtable reference on init error path
netfilter: ipset: do not update comments from kernel-side adds
====================
Link: https://patch.msgid.link/20260930074142.298353-1-pablo@netfilter.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
sock_error() can clear sk_err with xchg() without the socket lock.
On the no-data msk recv and splice paths, call sock_error() once and
only stop when it returns a non-zero error. Annotate the remaining
msk send and subflow error-report peeks with READ_ONCE().
Signed-off-by: Quanye Yang <quanyeyang@proton.me>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Link: https://patch.msgid.link/20260925-mptcp-sk-err-net-v5-2-0cac04d6ea48@proton.me
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
BUG: KCSAN: data-race in do_recvmmsg / mptcp_recvmsg
read-write (marked) to 0xffff8880134d391c of 4 bytes by task 2619 on cpu 1:
instrument_atomic_read_write include/linux/instrumented.h:113 [inline]
sock_error include/net/sock.h:2565 [inline]
do_recvmmsg+0x50c/0x580 net/socket.c:3049
__sys_recvmmsg net/socket.c:3144 [inline]
__do_sys_recvmmsg net/socket.c:3167 [inline]
__se_sys_recvmmsg net/socket.c:3160 [inline]
__x64_sys_recvmmsg+0x161/0x180 net/socket.c:3160
x64_sys_call+0x19c7/0x1ca0 arch/x86/include/generated/asm/syscalls_64.h:300
do_syscall_x64 arch/x86/entry/syscall_64.c:61 [inline]
do_syscall_64+0xde/0x3d0 arch/x86/entry/syscall_64.c:84
entry_SYSCALL_64_after_hwframe+0x77/0x7f
read to 0xffff8880134d391c of 4 bytes by task 2620 on cpu 0:
tcp_recv_should_stop include/net/tcp.h:3086 [inline]
mptcp_recvmsg+0x54d/0xd50 net/mptcp/protocol.c:2466
inet_recvmsg+0x204/0x210 net/ipv4/af_inet.c:894
sock_recvmsg_nosec net/socket.c:1151 [inline]
sock_recvmsg+0x11a/0x140 net/socket.c:1173
____sys_recvmsg+0x14b/0x3c0 net/socket.c:2933
___sys_recvmsg+0x116/0x160 net/socket.c:2975
__sys_recvmsg net/socket.c:3008 [inline]
__do_sys_recvmsg net/socket.c:3014 [inline]
__se_sys_recvmsg net/socket.c:3011 [inline]
__x64_sys_recvmsg+0xeb/0x160 net/socket.c:3011
x64_sys_call+0x1319/0x1ca0 arch/x86/include/generated/asm/syscalls_64.h:48
do_syscall_x64 arch/x86/entry/syscall_64.c:61 [inline]
do_syscall_64+0xde/0x3d0 arch/x86/entry/syscall_64.c:84
entry_SYSCALL_64_after_hwframe+0x77/0x7f
value changed: 0x0000006b -> 0x00000000
Reported by Kernel Concurrency Sanitizer on:
CPU: 0 UID: 0 PID: 2620 Comm: syz.2.33 Not tainted 7.2.0-g39d4f32c5d53 #76 PREEMPT(full)
Hardware name: QEMU Ubuntu 26.04 PC (i440FX + PIIX, 1996), BIOS 1.17.0-debian-1.17.0-1ubuntu1 04/01/2014
do_recvmmsg() and getsockopt(SO_ERROR) call sock_error() without the
socket lock. sock_error() clears sk_err with xchg(), which races with
unmarked loads of the same field.
KCSAN reported the unmarked peek in tcp_recv_should_stop(). Annotate
that helper and the other send-side peeks with READ_ONCE().
On the no-data recv and splice paths, if (sk_err) followed by
sock_error() and an unconditional break can return 0 after another
thread consumes the error. Call sock_error() once and only stop when
it returns a non-zero error.
tcp_bpf_sendmsg() read sk_err twice; fold those unmarked loads into
one READ_ONCE() and use that value as the returned errno. The field
is still not consumed.
MPTCP is handled in the next patch.
Suggested-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Link: https://lore.kernel.org/netdev/8bbee583-6f21-4817-bfeb-2d60057380a3@linux.dev/
Reported-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Closes: https://github.com/multipath-tcp/mptcp_net-next/issues/632
Signed-off-by: Quanye Yang <quanyeyang@proton.me>
Reviewed-by: Eric Dumazet <edumazet@kernel.org>
Link: https://patch.msgid.link/20260925-mptcp-sk-err-net-v5-1-0cac04d6ea48@proton.me
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
The bridge and the openvswitch code include <net/llc.h>,
<net/llc_pdu.h> and <linux/llc.h> without using anything out of them.
parse_ethertype() is the one exception: it wants LLC_SAP_SNAP, which
comes from the uAPI header <linux/llc.h> pulls in, so flow.c keeps
that one. The LLC core has leftovers of its own - slab, string,
interrupt and net_namespace - which it has not used in a long time.
None of this is new, the type 2 removal has just shrunk the LLC
headers enough to make it easy to see.
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Acked-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260928190800.2521749-4-kuba@kernel.org
Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
Nothing in tree registers the type 1 / type 2 packet handlers or the
station handler any more, and nothing looks at the socket hashes
hanging off struct llc_sap. The two SAPs opened in tree - SNAP, and STP
for the bridge's BPDUs and GARP's PDUs - receive through the per-SAP
rcv_func(), so llc_rcv() boils down to a SAP lookup and a call.
llc_sap_list becomes static and five exports go away with all this; the
llc2 module out of tree has been reworked to open its own SAPs.
Frames with a NULL DSAP used to be handed to the station handler
before the SAP lookup. They take the normal path now, meaning they get
dropped unless something registers SAP 0.
struct llc_sap is down to 48 bytes on 64-bit from over a kilobyte,
which also moves its GFP_ATOMIC allocation from kmalloc-2k to
kmalloc-64. The two 64-entry socket hashes are the bulk of it, and the
address it carried is now just the SAP number - nothing has read the
MAC half since the socket layer left.
llc_pdu.h keeps only what its remaining users need, plus LLC_PDU_RSP to
document the one argument which can take it. That takes out the type 2
(I and S format, FRMR) definitions, the XID and TEST builders with
their constants - the kernel neither sends nor answers either any more
- the SAP address defines, the field accessors, and the prototypes of
the llc_pdu.c helpers which went out of tree with the rest of LLC2.
With the I and S formats gone llc_pdu_header_init() has one PDU type
left, so drop the argument.
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Link: https://patch.msgid.link/20260928190800.2521749-3-kuba@kernel.org
Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
802.2 LLC is only used in tree by protocols which need the
connectionless type 1 subset: STP, which carries the bridge's BPDUs,
GARP, which sends its own UI PDUs, and SNAP. The llc2 module on top of
the core - the type 2 connection state machine, the type 1 SAP state
machine, the station component and the PF_LLC socket family - has no
in-kernel users and nobody who can test it; what we get instead is a
slow trickle of drive-by fixes.
There was a recent patch from Ernestas Kulik indicating potential
real life use, but it was new/experimental and that person is
not responding to off-list pings.
Let LLC2 follow AX.25, hamradio and AppleTalk out of the Linux tree.
We will maintain the code at: github.com/linux-netdev/mod-orphan
for anyone interested in playing with it.
PF_LLC goes in full, both the class two SOCK_STREAM and the class one
SOCK_DGRAM half, and so do /proc/net/llc/ and /proc/sys/net/llc/. Note
that the kernel also stops answering XID and TEST commands, addressed
to a SAP or to the station - those are type 1, but they lived in the
module, and they got answered whether or not any socket was open.
Nothing in tree asks for them; what the core keeps is SAP registration
and the UI path the in-tree users need.
Retain the uAPI for now, like we did for AppleTalk. Only the socket ABI
half of it is vestigial: STP, GARP, the bridge and openvswitch use the
SAP numbers it defines. Cleaning up what the core no longer needs
follows in the next patch.
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Link: https://patch.msgid.link/20260928190800.2521749-2-kuba@kernel.org
Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
SMC transmit work is queued without holding a socket reference. After
an active close has cancelled tx_work, a received CDC message can queue
it again while the socket is in SMC_PEERCLOSEWAIT1. Passive close can
then reach SMC_CLOSED, call smc_conn_free() and drop the last socket
reference before smc_tx_work() acquires the socket lock. The worker
then accesses the freed socket.
KASAN reported:
BUG: KASAN: slab-use-after-free in lock_sock_nested+0x97/0x180
Write of size 8 by task kworker/0:1/11
Workqueue: smc_tx_wq-00000000 smc_tx_work
Call Trace:
lock_sock_nested+0x97/0x180
smc_tx_work+0x5d/0x170
process_one_work+0x5ce/0xeb0
worker_thread+0x45b/0xd10
Allocated by task 24:
sk_prot_alloc+0x56/0x210
sk_alloc+0x2b/0x6f0
smc_tcp_listen_work+0x16d/0xfc0
Freed by task 182:
slab_free_after_rcu_debug+0xa6/0x1e0
rcu_core+0x509/0x1850
Hold a socket reference for each successful enqueue and release it when
the worker finishes or a pending invocation is cancelled. Cover all three
queue sites and all cancellation sites, including the socket options.
Check conn->freed in smc_tx_pending() under the socket lock: retaining the
socket does not retain connection resources released by smc_conn_free().
This also covers pending transmit processing from smc_release_cb().
Use queue_delayed_work() for the busy-slot retry as well. All tx_work
delays are zero, so an already queued invocation needs no timer update.
Unlike mod_delayed_work(), its return value distinguishes a successful
enqueue from work disabled temporarily by cancel_delayed_work_sync(),
allowing the extra reference to be returned when no work was queued.
Cc: stable+noautosel@kernel.org # LLM report + LLM fix, not seen in real life
Fixes: e6727f39004b ("smc: send data (through RDMA)")
Link: https://lists.openwall.net/netdev/2026/09/09/105
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Link: https://patch.msgid.link/20260927075120.3695060-1-nicoyip.dev@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Since commit d58e468b1112 ("flow_dissector: implements flow dissector
BPF hook") __skb_flow_dissect() needs a net pointer, either from
skb->dev, skb->sk, or since commit 3cbf4ffba5ee ("net: plumb network
namespace into __skb_flow_dissect") a caller provided pointer.
syzbot was able to reach seg6_make_flowlabel() with an skb having
neither skb->dev nor skb->sk set: a TIPC UDP bearer sends a discovery
message through an IPv4 route using seg6 encap, while
net.ipv6.seg6_flowlabel is set to 1.
seg6_make_flowlabel() already has a net pointer, use skb_get_hash_net().
WARNING: net/core/flow_dissector.c:1131 at __skb_flow_dissect+0x910/0x5368 net/core/flow_dissector.c:1126, CPU#0: syz.0.17/4930
Call trace:
__skb_flow_dissect+0x910/0x5368 net/core/flow_dissector.c:1126 (P)
__skb_get_hash_net+0xe0/0x29c net/core/flow_dissector.c:1903
skb_get_hash include/linux/skbuff.h:1663 [inline]
seg6_make_flowlabel+0xcc/0x1ec net/ipv6/seg6_iptunnel.c:132
__seg6_do_srh_encap+0x320/0xbf4 net/ipv6/seg6_iptunnel.c:161
seg6_do_srh+0x4b4/0xa44 net/ipv6/seg6_iptunnel.c:431
seg6_output_core+0x164/0x688 net/ipv6/seg6_iptunnel.c:681
seg6_output+0x44/0x1ac net/ipv6/seg6_iptunnel.c:744
lwtunnel_output+0x3d0/0x664 net/core/lwtunnel.c:356
dst_output include/net/dst.h:470 [inline]
ip_local_out+0x110/0x148 net/ipv4/ip_output.c:131
iptunnel_xmit+0x50c/0xd38 net/ipv4/ip_tunnel_core.c:97
udp_tunnel_xmit_skb+0x220/0x348 net/ipv4/udp_tunnel_core.c:187
tipc_udp_xmit+0x75c/0x9c0 net/tipc/udp_media.c:202
tipc_udp_send_msg+0x214/0x374 net/tipc/udp_media.c:274
tipc_bearer_xmit_skb+0x260/0x3b0 net/tipc/bearer.c:576
tipc_enable_bearer net/tipc/bearer.c:366 [inline]
__tipc_nl_bearer_enable+0xc90/0xfb0 net/tipc/bearer.c:1048
tipc_nl_bearer_enable+0x2c/0x48 net/tipc/bearer.c:1057
genl_family_rcv_msg_doit+0x1e4/0x2d4 net/netlink/genetlink.c:1114
Fixes: d58e468b1112 ("flow_dissector: implements flow dissector BPF hook")
Reported-by: syzbot+9408fbe0e6452a12e9ab@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/netdev/6abac2eb.3654fce1.1bec97.0000.GAE@google.com/
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260928194524.3617299-1-edumazet@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Memory-provider queue configuration is validated when the provider is
bound. A later ethtool ring change may invalidate it because drivers
can size queue memory from both ring depth and RX page size. The fbnic
consumer is added in the following patch.
Keep configured RX ring depths in netdev_config and stage proposed
values in cfg_pending. Validate every RX queue before calling the
driver. Each check validates the device defaults, then any queue
memory-provider override. Commit the values only after the driver
accepts them.
Drivers which consume stored ring depths through queue configuration
must initialize every RX depth before registering the netdev. Stored
values override callback defaults, including when zero.
The callback receives a rendered configuration rather than a queue ID.
Validation should depend on the configuration, not queue identity.
Checking defaults also covers the case where every queue has a
memory-provider override.
Drivers may normalize ring depths when applying them. Require the
validation callback to use the same normalization. Drivers must report
the applied depths through the ethtool_ringparam argument so the core
records the result.
Use the same transaction for ioctl and netlink. Drivers without
ndo_validate_qcfg skip the new validation.
Link: https://lore.kernel.org/all/20250421222827.283737-14-kuba@kernel.org/
Signed-off-by: Björn Töpel <bjorn@kernel.org>
Reviewed-by: Simon Horman <horms@kernel.org>
Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260925104417.2325213-4-bjorn@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
netdev_config manipulation will become slightly more complicated
soon and will be used by both ethtool and the queue API.
Encapsulate the logic in helper functions.
Signed-off-by: Björn Töpel <bjorn@kernel.org>
Reviewed-by: Breno Leitao <leitao@debian.org>
Reviewed-by: Simon Horman <horms@kernel.org>
Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260925104417.2325213-2-bjorn@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
sctp_wait_for_connect() drops the socket lock while it sleeps. An
out-of-the-blue ABORT can then be processed from the socket backlog and
unlink the association. If a concurrent shutdown(fd, SHUT_RD) sets
RCV_SHUTDOWN, the waiter breaks with err == 0 before checking
asoc->base.dead. Its final sctp_association_put() can then free the
association, leaving sctp_sendmsg_to_asoc() to continue with a dangling
pointer.
Check RCV_SHUTDOWN along with the wait error in sctp_sendmsg_to_asoc()
before using the association again. The check only accesses the socket,
so it needs no additional association reference. Return the existing
-ESRCH so that sctp_sendmsg() skips freeing a new association that may
already have been destroyed.
Keep sctp_wait_for_connect() unchanged to preserve its behavior for the
connect() caller.
Fixes: 668c9beb9020 ("sctp: implement assign_number for sctp_stream_interleave")
Cc: stable@vger.kernel.org
Reported-by: TencentOS Corvus AI <corvus@tencent.com>
Signed-off-by: Jun Yang <junvyyang@tencent.com>
Acked-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/20260926095606.68601-1-juny24602@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Previously, the per-VMA locking could fail in the face of writers
which necessitates a fallback to mmap_lock. The new
vma_start_read_unlocked() will wait for writers instead of failing.
Use the new helper. Wait for writers. Remove the fallback to mmap_lock.
The fallback removal does not affect NOMMU case because TCP_ZEROCOPY
is gated on CONFIG_MMU.
This really is a nice cleanup. It removes the need to pass the lock
state back and forth to find_tcp_vma().
Link: https://lore.kernel.org/20260831203056.838265-6-surenb@google.com
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Signed-off-by: Suren Baghdasaryan <surenb@google.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Acked-by: Lorenzo Stoakes <ljs@kernel.org>
Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Tested-by: syzbot@syzkaller.appspotmail.com
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Arve Hjønnevåg <arve@android.com>
Cc: Todd Kjos <tkjos@android.com>
Cc: Christian Brauner <christian@brauner.io>
Cc: Carlos Llamas <cmllamas@google.com>
Cc: Alice Ryhl <aliceryhl@google.com>
Cc: David S. Miller <davem@davemloft.net>
Cc: David Ahern <dsahern@kernel.org>
Cc: David Hildenbrand (Arm) <david@kernel.org>
|
|
smc_llc_link_active() marks a link active before scheduling its testlink
work. The first-link confirmation paths call it without llc_conf_mutex,
allowing link-down processing to clear the link between these operations:
Connection setup Link-down worker
smc_llc_link_active()
link->state = SMC_LNK_ACTIVE
smcr_link_clear()
link->clearing = 1
smc_llc_link_clear()
cancel_delayed_work_sync()
schedule_delayed_work()
The connection reference keeps the link alive during activation, but the
newly queued work outlives that reference. Later cleanup skips clearing
an already-clearing link, leaving its timer armed when the link group is
freed. KASAN reported:
BUG: KASAN: use-after-free in __run_timers+0x723/0x8d0
Write of size 8 at addr ffff88810f108810 by task swapper/1/0
Call Trace:
__run_timers+0x723/0x8d0
timer_expire_remote+0xd3/0x120
tmigr_handle_remote_up+0x4f4/0xab0
__walk_groups_from+0x40/0x150
tmigr_handle_remote+0x229/0x2c0
run_timer_softirq+0x1f5/0x250
Hold llc_conf_mutex around first-link activation on both sides so that
link-down processing cannot cancel the work before it is queued. Also
protect the client's initial optional add-link processing, which can
activate a second link without the lock. The other add-link paths already
hold llc_conf_mutex. This serializes every activation with link teardown
without changing the handshake sequence or error handling.
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Link: https://patch.msgid.link/20260927074547.3694742-1-nicoyip.dev@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
__netpoll_send_skb() parks an skb on npinfo->txq whenever the device
cannot take it at once, and once the queue is non-empty every later skb
goes straight there to keep ordering. queue_process() drains it from a
workqueue and, unlike the direct path, never polls the device for
completions: when the ring is stopped it backs off HZ/10. Nothing limits
the queue length.
A producer that outruns that drain therefore grows the queue until the
host is out of memory. Observed with netconsole forwarding a GPU driver
that logged one line at ~1e5/s after a firmware hang. The NIC was
moving ~17k packets/s: completions for each burst surfaced tens of ms
later, outside the one-tick window, so queue_process() slept HZ/10 per
ring while ~1e5 lines/s kept arriving. The queue grew at ~170 MB/s,
unreclaimable slab reached 51 GiB in five minutes and the OOM killer
ran from kswapd with 341 MiB of anonymous memory on the whole box. What
the queue held was the flood itself; the OOM report never left the
host.
Reproduced on the same host (netconsole over a 10G ConnectX-4 Lx) under
the same slow-completion condition: 200k lines to /dev/kmsg in 0.12 s
grew unreclaimable slab by 173 MiB, about 188k skbs, draining at
~8-10k packets/s. With prompt completions the same burst drains at line
rate; a stall on the link while lines keep arriving faster than the
drain reproduces the growth.
Until the 2006 netpoll rework [1] the deferred path drained through
dev_queue_xmit(), with the stack's own backpressure, and was capped at
16 skbs (MAX_QUEUE_DEPTH). That series moved it to a direct
hard_start_xmit() with the HZ/10 back-off and made the queue
per-device, dropping the cap on the way.
Cap it at 1024 skbs per device and drop new skbs beyond that. The drop
is counted in tx_dropped of the device whose queue is full and freed
with SKB_DROP_REASON_FULL_RING. A netconsole target bound directly to
that device also gets NET_XMIT_DROP and, with CONFIG_NETCONSOLE_DYNAMIC,
counts it in xmit_drop_count. When a stacked device (bond, bridge,
team, vlan, macvlan) passes the skb down and the lower device's queue
is the one that fills, the return value does not reach netconsole and
the lower device's tx_dropped is the record. The bound holds either
way. Nothing is logged on the drop path because that would recurse
into the console being drained.
[1] https://lore.kernel.org/netdev/20061026225645.482978803@osdl.org/
Signed-off-by: Zack Gomez <zack.gomez@gmail.com>
Reviewed-by: Breno Leitao <leitao@debian.org>
Link: https://patch.msgid.link/20260925204537.2664119-1-zack.gomez@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The L2 encapsulation modes of the seg6 lwtunnel reallocate the skb head
on every packet, where the IPv6 encapsulation modes reallocate only when
they have to. Ask for the whole encapsulation up front instead, so that
the reallocation happens at most once and only when the headroom really
is too small:
skb->mac_len + sizeof(struct ipv6hdr) + ipv6_optlen(tinfo->srh)
+ dst_dev_overhead(cache_dst, skb)
__seg6_do_srh_encap() then finds the room it needs and its own
skb_cow_head() becomes a no-op.
Drivers reserve more than that on the forwarding path, so the
reallocation usually disappears altogether. A single-segment policy
on ixgbe needs
14 (mac_len) + 40 (ipv6hdr) + 24 (SRH) + 16 (LL_RESERVED_SPACE) = 94
against the 206 bytes the driver leaves. Where the headroom is
smaller, as on a veth pair, pskb_expand_head() is called once per
forwarded packet instead of twice. Asking only for skb->mac_len would
still take two whenever the skb is header-cloned, because the cow that
unclones it does not also make room for the outer header.
The cost is amplified by CONFIG_INIT_ON_ALLOC_DEFAULT_ON, which many
distributions enable: every new head is zeroed in full, and that memset
alone accounts for 16% of the datapath profile.
Throughput at 0.5% packet loss, 64-byte frames forwarded through one
2.30 GHz core (Xeon E5-2650 v3, ixgbe 82599ES), offered by TRex and
binary-searched over 10 runs of 10 s:
Before: 654.6 kpps
After: 965.7 kpps
Signed-off-by: Yuya Kusakabe <yuya.kusakabe@gmail.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260925-seg6-l2cow-v3-1-fc83821542a7@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless-next
Johannes Berg says:
====================
More features:
- ath10k: NVMEM device tree bindings
- ath12k: QMI firmware alignments
- mm81x: AP improvements
- mac80211:
- CIP (control frame integrity) support
- NAN improvements
- cfg80211:
- improvements for AP regulatory checks
* tag 'wireless-next-2026-09-30' of https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless-next: (67 commits)
wifi: mac80211: Gracefully deauthenticate on association timeout
wifi: libertas: fix RX OOB access from device-controlled pkt_ptr
wifi: mwifiex: Reattach interfaces on suspend failure
wifi: b43legacy: work around stack frame size warning
wifi: cw1200: fix link_id_db OOB access via device-controlled link ID
wifi: mwifiex: bound SDIO fw dump count by memory table size
wifi: mac80211: start next ROC after purging an interface
wifi: mac80211: fix potential ack-skb leak on error path
wifi: mac80211: mesh: don't send peering close in listen
wifi: mac80211_hwsim: Support NAN deferred schedule completion
wifi: mac80211: Restrict probe request rates for minimal content
wifi: nl80211: allow a NAN peer schedule with 2 channels in the same slot
wifi: cfg80211: use sysfs_emit_at() in addresses_show
wifi: cfg80211: require zero terminator in valid_regdb country table
wifi: cfg80211: fix NAN local schedule update ordering and allocation
wifi: mac80211: Fix a race when expiring a mesh path
wifi: cfg80211: validate monitor channel set against radio usage
wifi: nl80211: reject color-change requests that change 6 GHz power type
wifi: nl80211: defer AP beacon regulatory check to start_ap
wifi: radiotap: add definitions for UHR U-SIG
...
====================
Link: https://patch.msgid.link/20260930124547.228697-29-johannes@sipsolutions.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless
Johannes Berg says:
====================
Still more fixes coming in, notably:
- ath11k: avoid running out of stations on HW restart
- mac80211:
- drop too large fragmented MPDUs
- mesh path handling fixes
- validation improvements
- reject CSA with bad 320 MHz bandwidth
- cfg80211: fix RTS for single radio devices
* tag 'wireless-2026-09-30' of https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless: (27 commits)
wifi: mac80211: fix slab-out-of-bounds read in ieee80211_monitor_select_queue()
wifi: mac80211: reject invalid 320 MHz CSA bandwidth
wifi: mac80211: set info->band for 802.3 encap offload frames
wifi: mac80211: prevent AP VLAN tx from other interfaces
wifi: cfg80211: preserve hidden-group beacon IE ownership
wifi: mac80211: keep fallback association elements alive
wifi: mac80211: shut down RX BA session timer on teardown
wifi: mac80211: validate TX status rate metadata
wifi: ath9k_htc: bound TX aggregation to MAX_TX_BUF_SIZE
wifi: ath9k: reject short WMI command responses
wifi: ath9k: Clean up device initialisation guards
wifi: ath11k: reset ar->num_stations on hardware start
wifi: cfg80211: fix RTS threshold setting for single-radio PHY
wifi: mac80211: handle empty FILS association request payload
wifi: mac80211: minstrel_ht: validate fixed rate index
wifi: p54: validate firmware record lengths
wifi: mac80211: fix mesh fast xmit path deletion UAF
wifi: mac80211: drop oversized fragments to avoid extra_len overflow
wifi: wlcore: Fix runtime PM leak in wlcore_remove()
wifi: mac80211: drain PS delivery work during station teardown
...
====================
Link: https://patch.msgid.link/20260930124440.224799-3-johannes@sipsolutions.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
When tcp_collapse() rebuilds skbs under memory pressure, the copy
process doesn't include the right tstamp and hwtstamp from the
old skb. And memcpy(nskb->cb, skb->cb, ...) copies has_rxtstamp,
but nskb->tstamp and hwtstamps are left at zero, so
tcp_recv_timestamp() ends up emitting no cmsg at all.
In net timestamping case, if such an skb happens to be the last
one consumed in a recvmsg() call, the application receives no RX
timestamp for that call.
Fix this by copying both tstamp and hwtstamp of the last skb to
the new skb, matching tcp_try_coalesce()/tcp_add_backlog().
Note that the has_rxtstamp flag can still be inherited through
the cb memcpy from an skb that contributes no bytes (fully covered
skb left in the ofo tree by the tcp_ooo_try_coalesce() ->
coalesce_done path), so set TCP_SKB_CB(nskb)->has_rxtstamp to false
which makes the new block the only place setting it.
Signed-off-by: Jason Xing <kerneljasonxing@gmail.com>
Reviewed-by: Eric Dumazet <edumazet@kernel.org>
Link: https://patch.msgid.link/20260924152529.5689-1-kerneljasonxing@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Before commit 871df5007eda ("netfilter: flowtable: bail out if forward
path cannot be discovered"), there was a fallback to set up a forward
path in case .ndo_fill_forward_path fails or DEV_PATH_MTK_WDMA was used.
Such fallback was used by commit d787a3e38f01 ("mac80211: add support
for .ndo_fill_forward_path").
One possibility is to handle DEV_PATH_MTK_WDMA from the flowtable
forward path discovery. However, this is only used internally by drivers
to retrieve mtk_wdma information to set up hardware offload. Felix
decided to use the .fill_forward_path interface for this purpose due to
the lack of a better interface at that time.
Add a new DEV_PATH_IEEE80211 path which is offered if the new ieee80211
flag is set on in the struct net_device_path_ctx to restore the
flowtable with a ieee80211 netdevice. Handle this new DEV_PATH_IEEE80211
path just like DEV_PATH_ETHERNET and DEV_PATH_DSA, ie. this is the last
netdevice in the stack.
This new ieee80211 flag is implicitly unset for mtk_ppe and airoha which
call dev_fill_forward_path() to retrieve a DEV_PATH_MTK_WDMA path.
Fixes: 871df5007eda ("netfilter: flowtable: bail out if forward path cannot be discovered")
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
Rename NF_FLOW_HW_PENDING to NF_FLOW_PENDING and use it to inhibit the
flowtable GC worker until pending hw offload work has been completed.
Apparently, nf_flow_offload_stats() can schedule work to retrieve stats
while the flow is being removed by GC.
And this bit can also be used in a follow up patch to disable GC until
the flow has been fully added in both directions.
Revert the reordering done in commit d644b23afe1e ("netfilter:
flowtable: publish HW_DEAD after worker is done") to prevent a race
between GC and hw offload handler.
Fixes: 2c8897953f3b ("netfilter: flowtable: Add pending bit for offload work")
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
As bpf_ct_set_nat_info() is not validating the NAT manipulation type a
wrong value can be passed directly to nf_nat_setup_info(). This triggers
the WARN_ON() at nf_nat_setup_info() and if panic_on_warn isn't set,
then IPS_SRC_NAT_DONE is set without adding nat_bysource and conntrack
cleanup tries to unlink an uninitialized hlist node.
Fix this by checking that NAT manipulation type is correct before
calling nf_nat_setup_info(). In addition, if the WARN_ON is hit, return
NF_DROP instead of continuing with the processing to avoid similar
situations in the future.
Reported-by: VEGA <vega@nebusec.ai>
Fixes: 0fabd2aa199f ("net: netfilter: add bpf_ct_set_nat_info kfunc helper")
Signed-off-by: Fernando Fernandez Mancera <fmancera@suse.de>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
Since nft_set_commit_update() runs set commit callbacks before processing
NEWSETELEM transactions, nft_rbtree_gc_scan() can observe elements added by
the transaction being committed.
The scan records an interval end in rbe_end without checking the element's
transaction state. A later, unrelated expired start then moves both
elements to the expired list. The synchronous GC queue can free the new end
element before the transaction subsequently activates it, causing a
use-after-free.
Only consider elements that are fully active in both generations. This
keeps transaction-state elements out of the GC scan and preserves interval
pairing across skipped elements.
KASAN reports:
BUG: KASAN: slab-use-after-free in nft_setelem_activate
nft_setelem_activate net/netfilter/nf_tables_api.c:7047
nf_tables_commit net/netfilter/nf_tables_api.c:11137
Allocated by task 130:
nft_set_elem_init net/netfilter/nf_tables_api.c:6794
nft_add_set_elem net/netfilter/nf_tables_api.c:7523
Freed by task 130:
nft_trans_gc_trans_free net/netfilter/nf_tables_api.c:10506
rcu_core kernel/rcu/tree.c:2919
Fixes: 1e3b9e1c77fe ("netfilter: nf_tables: call set ops .commit when building new ruleset blob")
Reported-by: <co+ee5e50ef2670e5f4@bugs.sh>
Assisted-by: LLM
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
While the IPVS SYNC protocol is not secure by design
we can still protect the backup server from messages that
can wreak havoc.
This commit addresses problems from received connection flags
or their combinations. We now drop messages as follows:
1. the NO_CPORT+TEMPLATE combination allows lookups for normal
connections to hit template which can break in many ways.
While the master does not sync connections with NO_CPORT flag,
i.e. before they are established, we still accept NO_CPORT
without TEMPLATE.
2. ONE_PACKET: it is not sent by master, so we do not
expect it in backup. Before now it was ignored by
IP_VS_CONN_F_BACKUP_MASK for protocol v1 while protocol
v0 created connections that are not hashed and dropped
immediately. Better to apply the IP_VS_CONN_F_BACKUP_MASK
also to the flags from v0 messages for consistency with v1.
Fixes: 87375ab47cd0 ("[IPVS]: ip_vs_ftp breaks connections using persistence")
Signed-off-by: Julian Anastasov <ja@ssi.bg>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
The IP_VS_CONN_F_ONE_PACKET flag was implemented for normal
connections. When conn template inherits this flag from
dest->conn_flags it will not be hashed. As result, we will
create new template for every new normal connection.
Fix it to allow one template to be used by many normal
connections.
Fixes: 26ec037f9841 ("IPVS: one-packet scheduling")
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260916231652.127456-1-pablo%40netfilter.org
Signed-off-by: Julian Anastasov <ja@ssi.bg>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
ip_vs_lblcr_new() and ip_vs_lblc_new() create cache entries for
every previously unseen destination address. The table max_size only
tells the periodic collector to reclaim entries after the cache has
already exceeded the limit. It does not reclaim entries that the
attacker continues to use.
Reject new cache entries once either table reaches max_size * 3 / 2.
The extra headroom lets the periodic collector catch up while the
existing scheduler fallback continues to use the selected destination
when cache creation fails. New traffic therefore stays serviceable
without growing the tables further.
Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Suggested-by: Julian Anastasov <ja@ssi.bg>
Signed-off-by: Zhiling Zou <zhilinz@nebusec.ai>
Acked-by: Julian Anastasov <ja@ssi.bg>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
LBLC may delete cache entries for destinations that are
removed or overloaded and replace them with available ones.
But ip_vs_lblc_new() forgets to decrement the tbl->entries
counter after calling ip_vs_lblc_del(). This can lead to
increased shrinking of the cache with every new garbage
collection.
Fixes: 2f3d771a35fe ("ipvs: do not use dest after ip_vs_dest_put in LBLC")
Link: https://sashiko.dev/#/patchset/0bdd5abe9968ded7ca2b9cb6844ba83d94cc8d53.1787318053.git.zhilinz%40nebusec.ai
Signed-off-by: Julian Anastasov <ja@ssi.bg>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
nft_flow_offload_init() bumps the flowtable use count with
nft_use_inc() before calling nf_ct_netns_get(). When the latter
fails, the error is returned as-is and the reference is leaked.
The upper layers do not balance it either: nf_tables_newexpr()
clears expr->ops when the expression init callback fails, so the
nft_expr_more() iteration in nft_rule_expr_deactivate() and
nf_tables_rule_destroy() stops right before the failed expression
and its ->destroy callback, which would drop the reference, never
runs.
Each failed rule addition therefore leaks one flowtable reference
and the flowtable can no longer be removed: NFT_MSG_DELFLOWTABLE
keeps reporting -EBUSY even though no rule references it.
Save the nf_ct_netns_get() return value and undo the nft_use_inc()
when it fails, restoring the inc/dec pairing within
nft_flow_offload_init() itself.
Fixes: a3c90f7a2323 ("netfilter: nf_tables: flow offload expression")
Reported-by: TencentOS Corvus AI <corvus@tencent.com>
Cc: stable@vger.kernel.org
Assisted-by: CodeBuddy:Kimi-K3
Signed-off-by: Aohan Mei <henrymei@tencent.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
# New commits in timers/core:
348f54c435bf ("selftests/timers: clocksource-switch: Fix unchecked open()/read()")
b03638013add ("selftests: timers: Measure the CPU timers on the clock they count")
a252cb93e156 ("selftests: timers: Count what tick is worth in the drift estimate")
3945c4a3ea15 ("time/kunit: Add time64_to_tm() case beyond 32-bit day count")
2927f7ca7844 ("time: Prevent time64_to_tm() day truncation on 32-bit")
bc5b66c300b8 ("timekeeping: Use READ_ONCE/WRITE_ONCE() for ktime_sec to prevent tearing")
bb41ece16463 ("timers/migration: Mark racy updates to tmigr_event::ignore field")
1159ad0a6aaa ("timers: Mark racy updates to hlist_node::pprev field")
5dffe33bfdaf ("hrtimer: Apply READ_ONCE() to lockless base->running loads")
ad80926d7803 ("hrtimer: Mark the hrtimer_sleeper structure's ->task field __private")
15f84398330c ("hrtimer: Update hrtimer_resolution only if value changes")
7eed1771a6e8 ("futex: Use accessor for hrtimer_sleeper ->task field in requeue")
9cd2f4e304d9 ("rtmutex: Use accessor for hrtimer_sleeper ->task field")
3990d196954a ("net: pktgen: Use accessor for hrtimer_sleeper ->task field")
9b7bcd671e9c ("timers: Use accessor for hrtimer_sleeper ->task field in sleep_timeout.c")
fa3d486086d6 ("futex: Use accessor for hrtimer_sleeper ->task field in waitwake.c")
851c30277918 ("io-uring/rw: Use accessor for hrtimer_sleeper ->task field")
66b29025b676 ("wait: Use accessor for hrtimer_sleeper ->task field")
cefc1a24ace5 ("aio: Use accessor for hrtimer_sleeper ->task field")
d166a1cd9016 ("hrtimer: Mark data-racy accesses to hrtimer_sleeper ->task field")
7b07d15ed1d7 ("posix-timers: Handle exit in do_exit() completely")
54ad1e0ea42c ("posix-cpu-timers: Prevent enqueueing when PF_EXITING is set")
760ad335d600 ("posix-cpu-timers: Use PF_EXITING to indicate exit")
8e9ed3b66f40 ("posix-cpu-timers: Move inlines out of public header")
24adce0b86c9 ("posix-timers: Move POSIX timer group exit related code out of do_exit()")
7117334f022b ("posix-timers: Move posixtimer_exec_cleanup() out of exec.c")
Signed-off-by: Ingo Molnar <mingo@kernel.org>
|
|
A TCP packet can carry new data while acknowledging traffic in the
opposite direction. With overlapping traffic in both directions, a
delayed packet's acknowledgment can be older than one Linux has already
accepted, even when that packet fills a gap in the received data.
Linux accepts the data, but tcp_ack() takes the old_ack path and skips
updating TS.Recent, the timestamp saved for outgoing acknowledgments.
The reply therefore echoes an older timestamp. If the sender uses this
echo to measure round-trip time after a long idle period, its estimate
includes the idle time and can reduce its sending rate.
Update TS.Recent in old_ack using tcp_replace_ts_recent(), before SACK
processing can trigger a transmission. This reuses the existing timestamp
and sequence checks, including PAWS protection against old duplicate
packets. ACK validation already rejects old ACKs in SYN_RECV before this
path, so no additional state check is needed.
Echoing the timestamp of the packet that fills the receive gap follows
RFC 7323 section 4.3. In a socket reproduction with 300 seconds idle,
controlled reordering and retransmission to exercise timestamp-based RTT
sampling, the sender's smoothed round-trip time was 37.5 seconds without
the fix and 15.5 ms with it.
Fixes: 12fb3dd9dc3c ("tcp: call tcp_replace_ts_recent() from tcp_ack()")
Assisted-by: LLM sparse
Signed-off-by: Jeff Jo <jeffjo@openai.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
struct mptcp_options_received is allocated on the stack while parsing
incoming MPTCP options. Several suboptions are mutually exclusive, as
enforced by mptcp_parse_option(), so their payloads can overlap.
Group fields by suboption and place the mutually exclusive payloads in
an anonymous union. Keep DSS and rm_list outside the union: they can
be combined with other suboptions. Move join_id into the MP_JOIN
group, and overlap token, thmac and hmac inside that group.
Further shrinking would require changing the parser so currently
coexisting fields (DSS mapping vs ACK, rm_list, status flags) can
overlap. That adds complexity for little gain, since the outer union is
already dominated by the MP_JOIN / ADD_ADDR members.
This reduces the structure size from 136 to 72 bytes on x86_64.
Closes: https://github.com/multipath-tcp/mptcp_net-next/issues/625
Signed-off-by: Quanye Yang <quanyeyang@proton.me>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Link: https://patch.msgid.link/20260926-net-next-mptcp-misc-feat-7-4-v1-3-67af4ab37406@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
MP_CAPABLE and MP_FASTCLOSE both stored their key in rcvr_key. Give
FASTCLOSE a dedicated fc_recv_key overlapped with rcvr_key in a union.
This helps shrink struct mptcp_options_received.
Signed-off-by: Quanye Yang <quanyeyang@proton.me>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Link: https://patch.msgid.link/20260926-net-next-mptcp-misc-feat-7-4-v1-2-67af4ab37406@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
This entry is only used in subflow_finish_connect().
Instead, use the original value from mp_opt, and pass it to
subflow_thmac_valid() to do the validation with the given truncated
hmac.
While at it, rename the variables in subflow_thmac_valid() to avoid
confusions about the received one vs the expected one.
Reviewed-by: Geliang Tang <geliang@kernel.org>
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Link: https://patch.msgid.link/20260926-net-next-mptcp-misc-feat-7-4-v1-1-67af4ab37406@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Commit c4336a07eb6b ("net: correctly handle tunneled traffic on IPV6_CSUM
GSO fallback") split skb_gso_has_extension_hdr() into mutually exclusive
branches on skb->encapsulation. This did not yet address all paths:
1. With skb->encapsulation set, the outer header is not checked. IPv6
tunnels such as ip6_gre and ip6_tunnel add an outer Destination
Options header by default (encap_limit). Their GSO packets skip
software GSO, then skb_csum_hwoffload_help() sees the outer extension
header and calls skb_checksum_help() on the GSO skb, which warns and
drops it.
2. Tunnels over IPv6 without an outer transport header, such as
ip6_tunnel, leave skb->transport_header at the inner transport
header. skb_network_header_len() then spans the outer IPv6, tunnel
and inner IP headers, a false positive.
3. UDP tunnels without an inner network header, such as SCTP-in-UDP or
PSP, have no inner IPv6 header to check. Decide on the outer header
alone.
4. Directly dereferencing inner_ip_hdr(skb)->version without
skb_header_pointer() is unsafe.
Instead, check ipv6_ext_hdr(nexthdr) on the outer IPv6 header and, if
set, on the inner IPv6 header. Read the headers with skb_header_pointer().
Keep the skb_network_header_len() check when !skb->encapsulation to also
catch encapsulated packets without skb->encapsulation (e.g., virtio).
Use the same helper in skb_csum_hwoffload_help(). Its open coded test
has the false positive of (2) and ignores the inner header.
Background: checksum offload of tunneled packets invariants:
Non-GSO skb:
- If the inner packet is CHECKSUM_PARTIAL, Local Checksum Offload computes
the outer checksum in software and the device offloads only the inner L4
checksum.
- If the inner packet is CHECKSUM_NONE (e.g., SCTP-in-UDP, ESP-in-UDP,
Remote Checksum Offload), the device offloads the outer UDP or GRE
checksum instead.
GSO skb:
- A device with NETIF_F_GSO_UDP_TUNNEL_CSUM or NETIF_F_GSO_GRE_CSUM
computes both inner and outer checksums per segment.
The outer checksum is seeded with only the pseudo-header checksum.
A NETIF_F_IPV6_CSUM device must parse through the outer headers to
reach the inner ones.
Fixes: c4336a07eb6b ("net: correctly handle tunneled traffic on IPV6_CSUM GSO fallback")
Cc: stable@vger.kernel.org
Signed-off-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260926140506.2335137-1-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
skbedit can set skb->queue_mapping and raise the per-CPU skip_txqueue
flag so __dev_queue_xmit() honours the mapping. __dev_queue_xmit()
cleared the flag before sch_handle_egress() and only read it afterwards,
so the flag was not confined to the xmit that set it: a nested xmit
(mirred redirect or mirror, or a drop after skbedit) could set the flag
and the outer xmit would consume it for an skb that never went through
skbedit.
A forwarded packet still carries the ingress NIC's rx_queue + 1 in
skb->queue_mapping, so the outer device then indexes its tx queue state
with that stale value. Taprio's child array q->qdiscs[] is sized to the
device's queue count, so taprio_enqueue() indexes past its allocation
and dereferences the result as a struct Qdisc *.
We (ab)use the skb->nf_skip_egress which means "skip netfilter egress
for this packet" to tag to "am I in tc egress?". Despite the overload
I dont see it as a conflict since the marker is set only around the
single sch_handle_egress() call and ingress path is guarded by
tc_at_ingress.
I will send a followup(net-next) patch once this hits net-next to
rename the skb->nf_skip_egress bit/flag to skb->skip_egress
Arm the flag only from the egress classifier that can use it: raise
skip_txqueue from tcf_skbedit_act() only when it runs inside
sch_handle_egress(), thanks to skb->nf_skip_egress. An egress qdisc
classifier runs in q->enqueue(), after the tx queue has been picked,
so a mapping it sets cannot affect the current packet; arming the flag
there only pollutes it for a later xmit. Then own the flag for the xmit
frame the egress hook runs in: save the incoming value and clear it just
before sch_handle_egress(), and restore it after the hook - on the
consumed (drop) path, or, in the same call that reads it, on the
surviving path. The save and the restores stay inside the
egress_needed_key static branch, so a packet pays for them only when
egress hooks are active (2f1e85b1aee4).
Store the value netdev_cap_txqueue() selected back into skb->queue_mapping
in netdev_tx_queue_mapping(), as netdev_core_pick_tx() already does, so
the skip_txqueue path never hands a later reader on the xmit path a
mapping the device cannot serve. A store made still later in the same
frame, by a tc BPF program attached to a transmit qdisc, is outside this
path and is not re-capped; a separate followup will resolve that path.
netdev_xmit_skip_txqueue() returns the previous flag value so the
save-and-clear is one call, and a no-op stub is provided when
CONFIG_NET_EGRESS is disabled. skb->nf_skip_egress is compiled under
CONFIG_NET_EGRESS rather than CONFIG_NETFILTER_SKIP_EGRESS, so
skb_at_tc_egress() is valid whenever the egress path is built.
A local user in a network namespace can redirect a packet from a device
with more TX queues to one with fewer after setting a mapping valid only
on the larger device. That reaches these reads and, under KASAN, faults
with "slab-out-of-bounds in taprio_enqueue".
Conditions to recreate the bug: the report's own trigger is a local user
with CAP_NET_ADMIN in a network namespace, so no eBPF program is needed.
With CONFIG_NET_SCH_TAPRIO=y, CONFIG_NET_ACT_SKBEDIT=y,
CONFIG_NET_ACT_MIRRED=y, CONFIG_NET_CLS_MATCHALL=y,
CONFIG_NET_SCH_PRIO=y and KASAN enabled, create qa (3 queues), qb
(2 queues) and qc (1 queue) as dummy devices; put a taprio root on qb
(num_tc 1, queues 2@0) and clsact on all three; then add an egress
matchall filter on every device. On qa: "action skbedit queue_mapping 2
pipe action mirred egress redirect dev qb". On qb: "action mirred egress
mirror dev qc". On qc: "action skbedit queue_mapping 0 pipe". Send one
packet out qa. qc's skbedit sets the flag while qb's outer xmit is in
flight; without the fix qb consumes it and reads its two-entry taprio
child array with the forwarded packet's stale mapping. A qc whose
skbedit is instead installed in a transmit-qdisc classifier (a matchall
filter on the qc root qdisc) reaches the same code path the same way
without the fix.
Testing: on a KASAN build with panic_on_warn=1 the unfixed kernel panics
with "BUG: KASAN: slab-out-of-bounds in taprio_enqueue", a read 0 bytes
past a 16-byte taprio_init() allocation, for the clsact-setter and the
transmit-qdisc-classifier reproducers and for a clsact skbedit-then-tc-BPF
store; the fixed kernel runs all three with no report, and the BPF store
variant additionally shows the expected "selects TX queue" clamp notice
from the write-back.
Fixes: 2f1e85b1aee4 ("net: sched: use queue_mapping to pick tx queue")
Reported-by: Zero Day Initiative <zdi-disclosures@trendmicro.com>
Link: https://lore.kernel.org/netdev/CANn89iLwYx8nCVf0pCEk_MmEiyC6kQaMwCQT9WkQVeeNzNQHqQ@mail.gmail.com/
Link: https://lore.kernel.org/netdev/179008581937.2160803.7117814290574262942@kernel.org/
Link: https://lore.kernel.org/netdev/179033713973.2160803.4914570693994398206@kernel.org/
Link: https://lore.kernel.org/netdev/20260925180407.63647514@kernel.org/
Link: https://lore.kernel.org/netdev/CANn89i+k-mZKDQVtvws_MEXeuMTAdaCcOXFZE-RfhcGTu90sjA@mail.gmail.com/
Suggested-by: Eric Dumazet <edumazet@google.com>
Suggested-by: Jakub Kicinski <kuba@kernel.org>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/QDISC-9R8V.v4.20260928081529@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Transition a FAILED neighbor entry to STALE upon receipt of an NA
message on routers when accept_untracked_na is enabled. This extends the
RFC 9131 accept_untracked_na behavior so that FAILED entries are treated
the same as non-existent entries.
RFC 4861 section 7.3.3 says that an entry should be deleted when address
resolution fails. Linux instead retains the entry in NUD_FAILED, so
treating it as untracked is consistent with the protocol model.
Trying to resolve FAILED neighbors via periodic probing (e.g. using
NTF_EXT_MANAGED) is more work compared to this approach which uses
information in NAs that the kernel may already be receiving. Note that
because this behavior in IPv6 is dependent on the accept_untracked_na
sysctl setting, this approach is more conservative than IPv4 which
transitions FAILED neighbors to STALE by default upon receiving GARPs.
Link: https://lore.kernel.org/r/20260813233344.445265-1-lfqlee314@gmail.com
Signed-off-by: Lawrence Lee <lfqlee314@gmail.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/6be400e601f4ffa59e47ddb5daa33ecdd85dd88b.1790127207.git.lfqlee314@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|