| Age | Commit message (Collapse) | Author |
|
https://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next.git
# Conflicts:
# arch/arm64/net/bpf_jit_comp.c
# arch/x86/net/bpf_jit_comp.c
# mm/internal.h
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next.git
# Conflicts:
# drivers/net/ethernet/realtek/r8169_main.c
# net/mac80211/ieee80211_i.h
# net/mac80211/tx.c
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/mm/linux.git
# Conflicts:
# arch/arm64/kvm/mmu.c
|
|
BUG: KCSAN: data-race in do_recvmmsg / mptcp_recvmsg
read-write (marked) to 0xffff8880134d391c of 4 bytes by task 2619 on cpu 1:
instrument_atomic_read_write include/linux/instrumented.h:113 [inline]
sock_error include/net/sock.h:2565 [inline]
do_recvmmsg+0x50c/0x580 net/socket.c:3049
__sys_recvmmsg net/socket.c:3144 [inline]
__do_sys_recvmmsg net/socket.c:3167 [inline]
__se_sys_recvmmsg net/socket.c:3160 [inline]
__x64_sys_recvmmsg+0x161/0x180 net/socket.c:3160
x64_sys_call+0x19c7/0x1ca0 arch/x86/include/generated/asm/syscalls_64.h:300
do_syscall_x64 arch/x86/entry/syscall_64.c:61 [inline]
do_syscall_64+0xde/0x3d0 arch/x86/entry/syscall_64.c:84
entry_SYSCALL_64_after_hwframe+0x77/0x7f
read to 0xffff8880134d391c of 4 bytes by task 2620 on cpu 0:
tcp_recv_should_stop include/net/tcp.h:3086 [inline]
mptcp_recvmsg+0x54d/0xd50 net/mptcp/protocol.c:2466
inet_recvmsg+0x204/0x210 net/ipv4/af_inet.c:894
sock_recvmsg_nosec net/socket.c:1151 [inline]
sock_recvmsg+0x11a/0x140 net/socket.c:1173
____sys_recvmsg+0x14b/0x3c0 net/socket.c:2933
___sys_recvmsg+0x116/0x160 net/socket.c:2975
__sys_recvmsg net/socket.c:3008 [inline]
__do_sys_recvmsg net/socket.c:3014 [inline]
__se_sys_recvmsg net/socket.c:3011 [inline]
__x64_sys_recvmsg+0xeb/0x160 net/socket.c:3011
x64_sys_call+0x1319/0x1ca0 arch/x86/include/generated/asm/syscalls_64.h:48
do_syscall_x64 arch/x86/entry/syscall_64.c:61 [inline]
do_syscall_64+0xde/0x3d0 arch/x86/entry/syscall_64.c:84
entry_SYSCALL_64_after_hwframe+0x77/0x7f
value changed: 0x0000006b -> 0x00000000
Reported by Kernel Concurrency Sanitizer on:
CPU: 0 UID: 0 PID: 2620 Comm: syz.2.33 Not tainted 7.2.0-g39d4f32c5d53 #76 PREEMPT(full)
Hardware name: QEMU Ubuntu 26.04 PC (i440FX + PIIX, 1996), BIOS 1.17.0-debian-1.17.0-1ubuntu1 04/01/2014
do_recvmmsg() and getsockopt(SO_ERROR) call sock_error() without the
socket lock. sock_error() clears sk_err with xchg(), which races with
unmarked loads of the same field.
KCSAN reported the unmarked peek in tcp_recv_should_stop(). Annotate
that helper and the other send-side peeks with READ_ONCE().
On the no-data recv and splice paths, if (sk_err) followed by
sock_error() and an unconditional break can return 0 after another
thread consumes the error. Call sock_error() once and only stop when
it returns a non-zero error.
tcp_bpf_sendmsg() read sk_err twice; fold those unmarked loads into
one READ_ONCE() and use that value as the returned errno. The field
is still not consumed.
MPTCP is handled in the next patch.
Suggested-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Link: https://lore.kernel.org/netdev/8bbee583-6f21-4817-bfeb-2d60057380a3@linux.dev/
Reported-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Closes: https://github.com/multipath-tcp/mptcp_net-next/issues/632
Signed-off-by: Quanye Yang <quanyeyang@proton.me>
Reviewed-by: Eric Dumazet <edumazet@kernel.org>
Link: https://patch.msgid.link/20260925-mptcp-sk-err-net-v5-1-0cac04d6ea48@proton.me
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
Previously, the per-VMA locking could fail in the face of writers
which necessitates a fallback to mmap_lock. The new
vma_start_read_unlocked() will wait for writers instead of failing.
Use the new helper. Wait for writers. Remove the fallback to mmap_lock.
The fallback removal does not affect NOMMU case because TCP_ZEROCOPY
is gated on CONFIG_MMU.
This really is a nice cleanup. It removes the need to pass the lock
state back and forth to find_tcp_vma().
Link: https://lore.kernel.org/20260831203056.838265-6-surenb@google.com
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Signed-off-by: Suren Baghdasaryan <surenb@google.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Acked-by: Lorenzo Stoakes <ljs@kernel.org>
Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Tested-by: syzbot@syzkaller.appspotmail.com
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Arve Hjønnevåg <arve@android.com>
Cc: Todd Kjos <tkjos@android.com>
Cc: Christian Brauner <christian@brauner.io>
Cc: Carlos Llamas <cmllamas@google.com>
Cc: Alice Ryhl <aliceryhl@google.com>
Cc: David S. Miller <davem@davemloft.net>
Cc: David Ahern <dsahern@kernel.org>
Cc: David Hildenbrand (Arm) <david@kernel.org>
|
|
When tcp_collapse() rebuilds skbs under memory pressure, the copy
process doesn't include the right tstamp and hwtstamp from the
old skb. And memcpy(nskb->cb, skb->cb, ...) copies has_rxtstamp,
but nskb->tstamp and hwtstamps are left at zero, so
tcp_recv_timestamp() ends up emitting no cmsg at all.
In net timestamping case, if such an skb happens to be the last
one consumed in a recvmsg() call, the application receives no RX
timestamp for that call.
Fix this by copying both tstamp and hwtstamp of the last skb to
the new skb, matching tcp_try_coalesce()/tcp_add_backlog().
Note that the has_rxtstamp flag can still be inherited through
the cb memcpy from an skb that contributes no bytes (fully covered
skb left in the ofo tree by the tcp_ooo_try_coalesce() ->
coalesce_done path), so set TCP_SKB_CB(nskb)->has_rxtstamp to false
which makes the new block the only place setting it.
Signed-off-by: Jason Xing <kerneljasonxing@gmail.com>
Reviewed-by: Eric Dumazet <edumazet@kernel.org>
Link: https://patch.msgid.link/20260924152529.5689-1-kerneljasonxing@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
A TCP packet can carry new data while acknowledging traffic in the
opposite direction. With overlapping traffic in both directions, a
delayed packet's acknowledgment can be older than one Linux has already
accepted, even when that packet fills a gap in the received data.
Linux accepts the data, but tcp_ack() takes the old_ack path and skips
updating TS.Recent, the timestamp saved for outgoing acknowledgments.
The reply therefore echoes an older timestamp. If the sender uses this
echo to measure round-trip time after a long idle period, its estimate
includes the idle time and can reduce its sending rate.
Update TS.Recent in old_ack using tcp_replace_ts_recent(), before SACK
processing can trigger a transmission. This reuses the existing timestamp
and sequence checks, including PAWS protection against old duplicate
packets. ACK validation already rejects old ACKs in SYN_RECV before this
path, so no additional state check is needed.
Echoing the timestamp of the packet that fills the receive gap follows
RFC 7323 section 4.3. In a socket reproduction with 300 seconds idle,
controlled reordering and retransmission to exercise timestamp-based RTT
sampling, the sender's smoothed round-trip time was 37.5 seconds without
the fix and 15.5 ms with it.
Fixes: 12fb3dd9dc3c ("tcp: call tcp_replace_ts_recent() from tcp_ack()")
Assisted-by: LLM sparse
Signed-off-by: Jeff Jo <jeffjo@openai.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
Observed with two independent servers in the same process:
fd1 = socket(AF_INET, SOCK_DGRAM, 0);
setsockopt(fd1, SOL_SOCKET, SO_REUSEPORT, ...);
bind(fd1, port 0); /* got 40000 */
fd2 = socket(AF_INET, SOCK_DGRAM, 0);
setsockopt(fd2, SOL_SOCKET, SO_REUSEPORT, ...);
bind(fd2, port 0); /* got 40000 as well */
Both sockets end up on the same port and join the same reuseport
group, so each of them takes part of the other's datagrams. The same
port sharing happens with SO_REUSEADDR, without the reuseport group.
udp_lib_lport_inuse() keeps the reuse rules when it scans for a free
port: a socket with SO_REUSEADDR set, or one with SO_REUSEPORT and
the same uid, is not a conflict, so its port is never marked in the
bitmap and the scan can hand it out again. Those rules only make
sense when the user asks for a specific port. bind(0) wants a free
port and has no way to know whose port it lands on.
TCP already does this: inet_csk_find_open_port() passes relax=false
and reuseport_ok=false, so the scan treats every port in use as a
conflict whatever options the sockets have. Commit aacd9289af8b
("tcp: bind() use stronger condition for bind_conflict") did it for
SO_REUSEADDR and commit 0643ee4fd1b7 ("inet: Fix get port to handle
zero port number with soreuseport set") for SO_REUSEPORT.
Do the same for UDP: ignore both options during a scan, keep them
for an explicit port.
This changes bind(0) for SO_REUSEADDR sockets once the port range is
full: it used to share a port and now fails with EADDRINUSE, like
TCP. Sharing a port on purpose when the range is full is a feature,
TCP has net.ipv4.ip_autobind_reuse for it, off by default. UDP can
get the same knob later, this patch only fixes the scan.
Fixes: ba418fa357a7 ("soreuseport: UDP/IPv4 implementation")
Cc: stable+noautosel@kernel.org # bind() behaviour change
Signed-off-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Link: https://patch.msgid.link/20260928023145.301855-1-jiayuan.chen@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
IP-in-IP GSO can re-enter inet_gso_segment() or ipv6_gso_segment()
for each nested IP header. encap_level tracks header bytes, not callback
depth, so a deep chain can exhaust the kernel stack. Making
inet_gso_segment() stackable introduced unbounded IPv4 nesting; IPIP
GSO/TSO later made the path reachable. The IPv6 stackable path was
introduced separately and uses the same guard.
Count IPv4 and IPv6 GSO handler entries in skb_gso_cb, initialized once
per top-level GSO operation and preserved across GRE/UDP context changes.
Use the existing IP_TUNNEL_RECURSION_LIMIT for both handlers. The first
five entries pass, and the sixth returns -EINVAL before dispatching
another GSO callback.
Fixes: 3347c9602955 ("ipv4: gso: make inet_gso_segment() stackable")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Closes: https://lore.kernel.org/all/cover.1790157745.git.zihanx@nebusec.ai/
Assisted-by: LLM
Co-developed-by: Luxing Yin <root@tr0jan.top>
Signed-off-by: Luxing Yin <root@tr0jan.top>
Signed-off-by: Zihan Xi <zihanx@nebusec.ai>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260924051521.32568-2-zihanx@nebusec.ai
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
Fix typos reported by scripts/checkpatch.pl using the misspelling list
in scripts/spelling.txt. Only touches comments, no code changes.
Assisted-by: Cursor:claude-opus-5
Signed-off-by: Hemanth Selam <hemanth.selam@gmail.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
Cross-merge networking fixes after downstream PR (net-7.3-rc5).
Conflicts:
drivers/net/mdio/mdio-realtek-rtl9300.c
89a8a1eef2d44 ("net: mdio: realtek-rtl9300: fix RTL931x C22 extended page selection")
cc3cb8db1eef9 ("net: mdio: realtek-rtl9300: Add page tracking")
https://lore.kernel.org/arUZOqy73bE2pp0w@sirena.org.uk
net/8021q/vlan_dev.c
cd5dd68267c4 ("vlan: ensure sufficient headroom in vlan_dev_hard_header()")
ca6ff8dd70eb ("vlan: annotate data-races in vlan_dev_priv fields")
Adjacent changes:
net/ipv6/ip6_gre.c
dd47bcf279f1 ("ip6_gre: Call ip6erspan_tunnel_unlink_md() in ip6erspan_changelink().")
cce829e2aa1d ("ip6_gre: Protect ip6gre_net.tunnels[][] with mutex.")
drivers/net/ethernet/meta/fbnic/fbnic_txrx.c
b5d9e9d4d0c1 ("eth: fbnic: use the Rx queue napi pointer to find the napi vector")
c0aca269ec07 ("eth: fbnic: Make Rx completion coalescing configurable")
drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c
d68acbf93531 ("net: stmmac: selftests: Support running selftests on DSA conduits")
85ca3292d7a3 ("net: stmmac: Remove ARP offload code")
drivers/net/ethernet/wangxun/libwx/wx_hw.c
3173cba11701 ("net: libwx: fix races in Tx timestamp handling")
7042c8c193e5 ("net: libwx: rename wx_pf_flags to wx_flags")
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Jakub Kicinski:
"Including fixes from Bluetooth, NFC and Netfilter.
Every week in this release is record-setting for number of posted
patches. It doesn't seem like we're creating any regressions with all
these fixes, three 'Fixes' tags here point to 7.2 commits but none are
true regression fixes. We're trying to keep the count down,
nonetheless.
Previous releases - regressions:
- net: don't require the hwtstamp NDOs when a PHY provides
timestamping
- ipv6: fix dst leak for uncached routes
- vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE
packets
Previous releases - always broken:
- packet: use ubuf_info completion for TX_RING packets
- arp: terminate device name before lookup
- ipv6: do not let ipv6_find_hdr() return an offset past the packet
end
- udp: remove a disconnected socket from the 4-tuple hash table
- sctp: discard the rest of the packet on a stale-cookie error
- eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged
eswitch"
[ And lots of other random network driver fixes ]
* tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (189 commits)
tcp: prevent collapsing skbs across boundary in rtx queue
vlan: ensure sufficient headroom in vlan_dev_hard_header()
net/sched: sch_teql: fix shadowed err in __teql_resolve()
bridge: check llc_mac_hdr_init() return value in br_send_bpdu()
llc: fix skb UAF and leaks on llc_mac_hdr_init() failure
llc: reserve device headroom for allocated frames
gve: DQO: reject TSO packets with an out of range MSS
gve: fix TX drop when GSO MSS is too small for hw
gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO
net: flush skb_defer_nodes in dev_cpu_dead()
net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails
af_packet: fix integer overflow in prb_calc_retire_blk_tmo()
tipc: Fix a data race on mon->peer_cnt in mon_timeout()
net: phy: intel-xway: workaround 100BASE-TX Link-Up issue
net/smc: fix UAF on lgr list traversal in smcr_port_err()
net/rds: size a connection's path set by the transport it ends up with
nfp: hold IPsec RX state under the XArray lock
net: ena: fix MMIO read buffer leak on probe failure
net: ena: fix PHC cleanup on probe failure
net/sched: act_ct: fix helper UAF due to extensions realloc
...
|
|
When tcp_send_synack() replaces the cloned SYN skb at the head of the
retransmit queue with a copy, it frees the original with
tcp_rtx_queue_unlink_and_free() and only repairs tp->highest_sack.
tp->retransmit_skb_hint keeps pointing at the freed
skbuff_fclone_cache object.
The dangling hint is read in tcp_verify_retransmit_hint() and used as
the root of the rbtree walk in tcp_xmit_retransmit_queue(). An
unprivileged TFO client (sendmsg(MSG_FASTOPEN)) can arm the hint with
an attacker-supplied ICMP fragmentation-needed message, after which a
simultaneous open frees the armed SYN skb:
BUG: KASAN: slab-use-after-free in tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316)
Read of size 4 at addr ffff88800604d928 by task swapper/1/0
Call Trace:
tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316)
tcp_simple_retransmit (net/ipv4/tcp_input.c:3158)
tcp_v4_err (net/ipv4/tcp_ipv4.c:587)
Sync the hint to the copy.
Fixes: c31b70c9968f ("tcp: Add logic to check for SYN w/ data in tcp_simple_retransmit")
Reported-by: Kimi Security Team <bug-report@moonshot.ai>
Tested-by: Weiming Shi <shiweiming@moonshot.ai>
Signed-off-by: Yilin Zhang <yilinzhang@moonshot.ai>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/8a9dff4063a2745653b7e88ceb745d75efa16e68.1790224474.git.yilinzhang@moonshot.ai
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Cross-merge BPF and other fixes after downstream PR.
Conflicts:
kernel/bpf/helpers.c
tools/testing/selftests/bpf/prog_tests/cb_refs.c
tools/testing/selftests/bpf/prog_tests/verifier.c
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
|
|
Pull bpf fixes from Alexei Starovoitov:
- Fix bpf_skb_change_tail() to drop the checksum offload instead of
rejecting the trim of CHECKSUM_PARTIAL skbs (Daniel Borkmann)
- Add KF_PERFMON kfunc flag and require CAP_PERFMON for kfuncs that
read arbitrary memory and for untrusted read-only memory reads
(Daniel Borkmann)
- Clear scalar delta on narrowing stack spill (Daniel Borkmann)
- Set up the frame pointer for the exception callback in arm64 JIT, and
zero-fill other CPUs when BPF_F_CPU update creates a per-cpu hash
element (Donggeun Yoo)
- Various fixes (Emil Tsalapatis):
- Fix bounds check underflow for skb-backed dynptrs
- Fix rx_queue_mapping context access code generation in bpf_sock
- Reject packet pointer arguments to subprogs that may mutate the
packet
- Reject ALU instructions that see arena and non-arena operands on
different code paths
- Fix copied_seq double-counting on sockmap self-redirect
(Geliang Tang)
- Fix divide-by-zero in btf_struct_walk() on a flexible array of
zero-sized elements, fix out-of-bounds read of rtt_min in sock_ops
(Jiayuan Chen)
- Fix bpf_sock_destroy() out-of-bounds read of sk_protocol on TIME_WAIT
and request socks, and sleeping under RCU when destroying a listener
with pending children (Jiayuan Chen)
- Fix JEQ/JNE with immediate operand in MIPS32 JIT and missing zero
extension of BSWAP 16/32 in MIPS64 JIT (Johan Almbladh)
- Avoid soft lockup in htab lookup[_and_delete] batch operations on
large maps (Jose Fernandez)
- Various fixes (Kumar Kartikeya Dwivedi):
- Verify global subprogs in each sleepability context they are
called from
- Make post-verification instruction rewrites killable
- Preserve packet pointer displacement in regsafe()
- Apply CO-RE relocations before subprogram validation, restrict
CO-RE poisoning to relocatable instructions, and reject truncated
ldimm64 CO-RE relocations in libbpf
- Assign lock identity to callback map values
- Compare stack frames in regs_exact()
- Bound ownership depth through local kptrs and graph roots
- Fix u32 overflow in map batch operations when the map size exceeds
4GB (Masoud Aghasi)
- Fix UAF in bpf memalloc due to concurrent consumption of ttrace lists
in alloc_bulk() (Pu Lehui)
- Allow gotox as the terminal instruction of a program or a subprogram
(Siddharth Chintamaneni)
- Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL, skip unsettled links
in link iterator, and reject dev-bound-only programs on other devices
(Weiming Shi)
- Reject non-negative stack offsets in stack_slot_obj_get_spi()
(Xu Yunxiang)
- Check params size before reading reserved fields in
bpf_crypto_ctx_create() (Yuqi Xu)
- Reject max_entries > INT_MAX in sock_map_alloc() (Zhao Gongyi)
- Use a 32-bit compare in xsk_map_gen_lookup() (Zhiling Zou)
- Use kvfree() in xdp_test_run_teardown() (Zhixing Chen)
* tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: (58 commits)
selftests/bpf: Test per-cpu initialization of a BPF_F_CPU created element
bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element
bpf: Fix BSWAP 32 and 16 on MIPS64
bpf: Fix immediate JMP JEQ/JNE on MIPS32
bpf: Reject dev-bound-only programs on other devices
bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc
selftests/bpf: Test for mixed arena/nonarena code paths
bpf: Prevent variable arena/non-arena register contents
selftests/bpf: Test rejection of pkt args to mutating subprogs
bpf: Reject pkt arguments in mutating subprogs
selftests/bpf: Add selftests for rx_queue_mapping context access
bpf: Fix bpf_sock context code generation
selftests/bpf: Test dynptr slices past end of skb
bpf: Fix bounds check for skb-backed dynptrs
selftests/bpf: Reject iterator destruction through fp+0
bpf: Reject non-negative offsets in stack_slot_obj_get_spi()
bpf: Check params size before reading reserved fields
selftests/bpf: Check local object ownership depth
bpf: Bound ownership depth through local kptrs and graph roots
selftests/bpf: Cover frame changes in bounded loops
...
|
|
ic_dhcp_init_options() appends the hostname (option 12), vendor-class
(option 60) and client-ID (option 61) options into the fixed 312-byte
bootp_pkt.exten[] buffer. Only the client-ID branch checked the
remaining space; the hostname and vendor-class writes were unbounded.
A 64-byte hostname together with the maximum 252-byte dhcpclass=
identifier needs 18 + (2 + 64) + (2 + 252) = 338 of the 312 available
bytes even before the terminating END marker, so the vendor-class memcpy
runs past the end of exten[]. With CONFIG_FORTIFY_SOURCE this is
reported as a field-spanning write and, when the kernel is booted with
panic_on_warn=1, aborts boot with a panic.
Route the optional options through a common helper that makes sure the
option, its 2-byte header and the END marker all fit and drops an option
that would not. Configurations with short options keep sending exactly
the same bytes as before.
Fixes: 130c0f47fdf9 ("ipconfig: send host-name in DHCP requests")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Assisted-by: LLM
Signed-off-by: Yuqi Xu <xuyuqiabc@gmail.com>
Reviewed-by: Ren Wei <weir@nebusec.ai>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/7808dfbfa2162dfd0b19f59aff5742d6e0db2abb.1789798023.git.xuyuqiabc@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
ipgre_netlink_parms() can enable collect_md on an existing GRE, GRETAP
or ERSPAN device. Unlike newlink, changelink does not enforce metadata
tunnel uniqueness. Converting a non-metadata device can therefore
replace the metadata receive entry for another device of the same type
in the same netns. Deleting either device then clears the shared entry,
breaking metadata receive lookup for the surviving device.
If parameter validation fails after collect_md is set, deleting the
modified device can also clear an entry it never owned.
Reject enabling metadata mode in both changelink callbacks before any
encapsulation or tunnel parameters are modified. Allow requests that
repeat the metadata attribute on an existing metadata device.
Fixes: 2e15ea390e6f ("ip_gre: Add support to collect tunnel metadata.")
Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260921031859.9283-1-xuanqiang.luo@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The ARP ioctl copies a user-provided struct arpreq into a stack object. Its
arp_dev field may contain IFNAMSIZ bytes without a NUL terminator.
Such input is passed to dev_get_by_name_rcu() or __dev_get_by_name(), where
strcmp() can read past the end of the stack object when a matching
alternative interface name exists.
Terminate the field before the lookup to prevent the out-of-bounds read.
Fixes: 36fbf1e52bd3 ("net: rtnetlink: add linkprop commands to add and delete alternative ifnames")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Signed-off-by: Zijie Huang <milkory@outlook.com>
Signed-off-by: Ren Wei <weir@nebusec.ai>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/fabf02a70787d17299e4b3153eadffaf20d154b3.1789910973.git.milkory@outlook.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Commit 7a9bc9e3f423 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.") added
NLA_POLICY_MIN(NLA_U8, 1) to fou_nl_policy[FOU_ATTR_IPPROTO], which
rejects an explicitly supplied FOU_ATTR_IPPROTO == 0 attribute with
-ERANGE.
However, FOU_ATTR_IPPROTO is an optional netlink attribute. When a user
sends FOU_CMD_ADD with FOU_ATTR_TYPE set to FOU_ENCAP_DIRECT and omits
FOU_ATTR_IPPROTO entirely, nla_policy validation succeeds and
parse_nl_config() leaves cfg->protocol as 0 (from memset(cfg, 0,
sizeof(*cfg))). fou_create() then creates a FOU_ENCAP_DIRECT socket with
fou->protocol == 0.
In fou_udp_recv(), returning -fou->protocol to udp_queue_rcv_one_skb()
triggers IP protocol resubmission when fou->protocol > 0, whereas
returning 0 tells the UDP tunnel layer that the skb was consumed without
freeing it. When fou->protocol == 0, every packet received on the socket
returns 0 from fou_udp_recv() and leaks the sk_buff.
Reject FOU_ENCAP_DIRECT when !cfg->protocol in fou_create() so that
creating a direct encapsulation port without FOU_ATTR_IPPROTO fails with
-EINVAL while leaving FOU_CMD_DEL and FOU_CMD_GET (which share
parse_nl_config()) unaffected.
Tested in QEMU against Linux 7.3.0-rc3 by sending a FOU_CMD_ADD Generic
Netlink request with FOU_ATTR_PORT = 5555 and FOU_ATTR_TYPE =
FOU_ENCAP_DIRECT while omitting FOU_ATTR_IPPROTO. On the unfixed kernel,
FOU_CMD_ADD succeeds (err = 0), FOU_CMD_GET reports fou->type = 1 and
fou->protocol = 0, and sending 4000 UDP packets to 127.0.0.1:5555 leaks
all 4000 sk_buffs (SUnreclaim in /proc/meminfo grows from 41456 kB to
59008 kB, +17552 kB); with this patch applied, FOU_CMD_ADD is rejected
with -EINVAL (-22).
Fixes: 23461551c006 ("fou: Support for foo-over-udp RX path")
Fixes: 7a9bc9e3f423 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.")
Cc: stable@vger.kernel.org
Signed-off-by: Hui Peng <benquike@gmail.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260921045920.1613098-1-benquike@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Commit c5c37af6ecad9 ("tcp: Convert do_tcp_sendpages() to use
MSG_SPLICE_PAGES") moved tcp_rate_check_app_limited() inside
do_tcp_sendpages(), turning it into a wrapper around tcp_sendmsg_locked().
Later, commit ebf2e8860eea ("tcp_bpf: Inline do_tcp_sendpages as it's now
a wrapper around tcp_sendmsg") inlined the wrapper in tcp_bpf_push() with
direct tcp_sendmsg_locked() calls, which perform the check on every path
that queues data, but kept the outer tcp_rate_check_app_limited() that
was previously needed to cover do_tcp_sendpages(). The outer call is now
redundant.
The site changed here, tcp_bpf_push(), holds the socket lock and invokes
tcp_sendmsg_locked() on every iteration. The early-return paths in
tcp_sendmsg_locked() that skip tcp_rate_check_app_limited() - the
MSG_ZEROCOPY allocation failure and MSG_FASTOPEN branches - return without
queueing any MSG_SPLICE_PAGES data, so there is no functional consequence
from omitting the outer check.
A potential benefit of this change is that it facilitates future reuse of
tcp_bpf_push() for sockmap support in protocols beyond TCP, such as MPTCP.
Since tcp_rate_check_app_limited() is TCP-specific while sendmsg_locked()
is a generic interface in struct proto_ops, this change allows us to switch
to different protocols via sk->sk_socket->ops->sendmsg_locked() without
carrying protocol-specific assumptions.
Signed-off-by: Geliang Tang <tanggeliang@kylinos.cn>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Link: https://patch.msgid.link/1f7dc605b16fc0590f7cbf5a27d57271926c01ce.1790074764.git.tanggeliang@kylinos.cn
|
|
Since commit 3d501dd326fb ("tcp: do not accept ACK of bytes we never
sent"), tcp_ack() bounds the acceptable old ACK window by
min(tp->max_window, tp->bytes_acked).
When sk->sk_state == TCP_SYN_RECV, tp->bytes_acked is always 0, so
any segment with before(ack, prior_snd_una) immediately returns
-SKB_DROP_REASON_TCP_TOO_OLD_ACK and never reaches the old_ack label
(which returns 0).
Therefore, tcp_ack() can only return 0 in closing states (where old
ACKs are accepted), and can never return 0 in TCP_SYN_RECV.
Simplify the tcp_ack() return value check in tcp_rcv_state_process()
to only check for negative return values and remove the unreachable
!reason branch.
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Link: https://patch.msgid.link/20260922010627.2291980-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
A UDP socket bound to a specific address and port keeps its entry in the
4-tuple hash table after it is disconnected:
sk binds to 127.0.0.1:21001
sk connects to 127.0.0.2:20001 // filed in the 4-tuple table
sk disconnects, connect(AF_UNSPEC) // still filed, peer now 0.0.0.0:0
__udp_disconnect() takes a socket out of that table only as a side effect
of ->rehash() or ->unhash(), and it skips ->rehash() when
SOCK_BINDADDR_LOCK is set and ->unhash() when SOCK_BINDPORT_LOCK is set.
commit 6996a2d2d0a6 ("udp: Unhash auto-bound connected sk from 4-tuple hash
table when disconnected.") fixed the same end state for a wildcard-bound
socket, by a path this one does not take.
The entry is counted whether or not anything hits it. hash4_cnt on the
hash2 slot stays raised for as long as the socket lives, so udp_has_hash4()
keeps sending every packet for that address and port through the 4-tuple
lookup first.
On IPv6 it can also be hit. __udp_disconnect() does not clear sk_v6_daddr,
so udp_v6_rehash() files the entry under the peer the socket was connected
to with a zero dport, and inet6_match() compares that same
field: a datagram from the former peer with a zero source port matches,
and source port zero is accepted on receive. On IPv4 the peer is cleared,
so a match would need a zero source address as well, which the routing
layer rejects as martian. The stale sk_v6_daddr is a separate defect, not
addressed here; removing the entry closes this path either way.
The entry can also be relocated. __udp_disconnect() clears sk_bound_dev_if,
so a subsequent SO_BINDTODEVICE calls ->rehash(), and because the receive
address is still specific udp_lib_rehash() moves the entry instead of
removing it, into the bucket that (rcv_saddr, num, 0, 0) hashes to -- a
pure function of the address and port, so every socket reaching this state
on one address and port collects in one bucket. The bucket cannot be chosen
from outside, as udp_ehashfn() is seeded with a per-boot secret. This last
one became reachable only with commit 644f9108f3a5 ("udp: Make rehash4
independent in udp_lib_rehash()"), which moved the hash4 handling out of a
branch a disconnected socket does not take; the stale entry itself dates
from the commit in Fixes.
Take the socket out of the table before __udp_disconnect() runs, while it
still matches how it was filed. This also reaches the wildcard case ahead
of udp_lib_rehash()'s udp_unhash4() branch, leaving that branch unreachable
from udp_disconnect(); removing it belongs in net-next. udp_disconnect()
and udp_abort() are the only UDP entries into __udp_disconnect(), which is
shared with raw, ping and l2tp sockets that are not struct udp_sock:
ping_prot.obj_size is sizeof(struct inet_sock), so udp_hashed4() on one
would read past the allocation.
Fixes: 78c91ae2c6de ("ipv4/udp: Add 4-tuple hash for connected socket")
Assisted-by: LLM
Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-2-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
A connected UDP socket that connects again to a different peer is not
re-filed in the 4-tuple hash table:
sk binds to 127.0.0.1:21001
sk connects to 127.0.0.2:20001 // filed under hash(sk, peer1)
sk connects to 127.0.0.3:20002 // still filed under hash(sk, peer1)
packet from 127.0.0.3:20002 // hash(sk, peer2) misses, so the
// lookup falls back to scoring the
// hash2 chain for this address
// and port
udp_lib_hash4() returns early when the socket is already hashed, assuming
->rehash() relocates it. ->rehash() runs from __ip{4,6}_datagram_connect()
only while the receive address is unset, which a second connect never is:
the first connect assigns it, whether the socket was bound to a specific
address or to the wildcard. commit 644f9108f3a5 ("udp: Make rehash4
independent in udp_lib_rehash()") added that early return and named
connect(AF_UNSPEC) as the way around it. That workaround does not help a
socket with both SOCK_BINDADDR_LOCK and SOCK_BINDPORT_LOCK set, because
__udp_disconnect() skips ->rehash() for the first and ->unhash() for the
second.
Delivery is correct either way.
Relocate the socket when the hash it is filed under differs from the one
requested, which is what commit 78c91ae2c6de ("ipv4/udp: Add 4-tuple hash
for connected socket") did before the early return became unconditional. It
is done here under hslot->lock, which that version did not take, to match
udp_lib_rehash() and udp_lib_unhash(). hslot2 is unchanged, so hash4_cnt
needs no adjustment, as in udp_lib_rehash(). A first connect is unaffected,
and IPv6 shares the code.
With 500 sockets on the port, a re-connected socket measured 522,553 pps
without this change and 2,055,078 with it. The UDP side was noted as
remaining work in [1].
Link: https://lore.kernel.org/netdev/apnHqmYZQ4yzOP4N@v4bel/ [1]
Fixes: 644f9108f3a5 ("udp: Make rehash4 independent in udp_lib_rehash()")
Assisted-by: LLM
Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-1-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
Add the TCP header option callbacks to the bpf_tcp_ops struct_ops type:
parse_hdr - parse the options of an incoming skb on an established
connection
hdr_opt_len - reserve space in the TCP header for bpf options
write_hdr_opt - write the reserved bpf options
These mirror the BPF_SOCK_OPS_PARSE_HDR_OPT_CB, _HDR_OPT_LEN_CB and
_WRITE_HDR_OPT_CB legacy sockops callbacks, but are exposed as struct_ops
members so a program can implement them with normal function signatures
and per-member helper sets.
The reserved header window is shared between the legacy sockops and
bpf_tcp_ops paths. tcp_{syn,synack,established}_options() first run the
legacy BPF_SOCK_OPS_HDR_OPT_LEN_CB and then call hdr_opt_len, so both
sources accumulate into opts->bpf_opt_len; at write time the legacy
options are emitted first and bpf_tcp_ops writes after them.
API design
bpf_tcp_ops overloads the sock_ops header-option helpers rather than
introducing a new API: bpf_reserve_hdr_opt(), bpf_store_hdr_opt() and
bpf_load_hdr_opt() are exposed per-member (reserve for hdr_opt_len,
store/load for write_hdr_opt, load for parse_hdr) and share the existing
kernel option-walking core via _bpf_sock_ops{store,load}hdr_opt(), with
the bpf_tcp_ops wrappers synthesizing a temporary bpf_sock_ops_kern from
the program ctx. This keeps a port from the legacy
BPF_SOCK_OPS*_HDR_OPT_CB callbacks mechanical (same helper calls) and
adds no new UAPI helper/kfunc surface.
An alternative considered was to drop the option helpers entirely: have
hdr_opt_len reserve space purely through its return value, and introduce
a dedicated TCP-header-option dynptr used for both reading and writing.
That is a cleaner, more self-contained interface, but it is a larger
change and does not reuse the legacy helpers, making a port from sockops
less mechanical. It can be pursued as a follow-up; the helper-based
interface here keeps this series focused on moving the hooks to
struct_ops.
The hdr_opt_len fast path in tcp_established_options() is gated by
cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS). Note this is a global,
per-attach-type static branch: it is enabled whenever any bpf_tcp_ops is
attached, even one that does not implement hdr_opt_len or that is attached
to a different cgroup. In those cases the block still runs but
bpf_tcp_ops_hdr_opt_len() no-ops via the per-member check in the dispatch
macro. A per-member/per-cgroup gate could be added later if the extra
fast-path work proves measurable.
Signed-off-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Link: https://patch.msgid.link/20260917200542.3689605-13-ameryhung@gmail.com
|
|
In LSFMMBPF 2025, I have talked about moving the BPF_PROG_TYPE_SOCK_OPS
to a struct_ops interface [1].
The BPF_SOCK_OPS_*_CB enum interface has grown over time as new TCP
callback points were added. A BPF_PROG_TYPE_SOCK_OPS program now
commonly needs a large switch on sock_ops->op, and the shared
bpf_sock_ops_kern context has become harder to extend because different
callbacks have different locking, argument, skb, and helper
requirements. The existing 'union { u32 args[4]; u32 replylong[4]; }' is
also not reliable in passing args to bpf prog when there are multiple
progs attached to a cgroup.
The above has already been solved in struct_ops. Add a TCP-specific
struct_ops type, bpf_tcp_ops, and support attaching it to cgroups.
This allows each callback have its own func signature and allows
the verifier to select kfuncs/helpers based on the specific
struct_ops member being implemented.
This patch wires up the following existing sock_ops callbacks:
- BPF_SOCK_OPS_TIMEOUT_INIT
- BPF_SOCK_OPS_RWND_INIT
- BPF_SOCK_OPS_RTT_CB
- BPF_SOCK_OPS_STATE_CB
- BPF_SOCK_OPS_RETRANS_CB
- BPF_SOCK_OPS_TCP_CONNECT_CB
- BPF_SOCK_OPS_TCP_LISTEN_CB
- BPF_SOCK_OPS_RTO_CB
- BPF_SOCK_OPS_ACTIVE_ESTABLISHED_CB
- BPF_SOCK_OPS_PASSIVE_ESTABLISHED_CB
BASE_RTT is ignored as it is not particularly useful. NEEDS_ECN should
be done in bpf-tcp-cc instead. The tstamp ones should be a separate
struct_ops (e.g. "bpf_sock_ops") that can work in both TCP and UDP.
timeout_init and rwnd_init could have a request_sock pointer. This patch
tries a different API and directly passes the request_sock pointer as
an arg.
Two other approaches were considered before settling on having
bpf_get_retval() read the dispatcher's run_ctx via saved_run_ctx. The
first was to inherit the retval in the trampoline itself: add a helper
in the four __bpf_prog_enter*() paths that, for struct_ops programs,
copies the chained value from the caller's run_ctx (now saved_run_ctx)
into the program's own run_ctx. It works but puts a per-enter
program-type check on the generic trampoline fast path, taxing all
fentry/fexit/lsm callers for a cgroup-struct_ops-only feature. The
second was to do that same inherit only for the int-returning members
via a gen_prologue that emits a hidden kfunc at the start of
timeout_init/rwnd_init; this keeps the cost off the generic path and
scoped to bpf_tcp_ops, but needs a kfunc + BTF_ID + prologue-emission
machinery. The chosen approach avoids both: it touches neither the
trampoline nor the program, since saved_run_ctx already points at the
dispatcher's run_ctx that carries the value.
[1], page 13: https://drive.google.com/file/d/1wjKZth6T0llLJ_ONPAL_6Q_jbxbAjByp/view?usp=sharing
Signed-off-by: Martin KaFai Lau <martin.lau@kernel.org>
Signed-off-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Link: https://patch.msgid.link/20260917200542.3689605-12-ameryhung@gmail.com
|
|
bpf_struct_ops_map_free() currently waits for both a regular RCU grace
period and a tasks RCU grace period for every struct_ops map through
synchronize_rcu_mult(call_rcu, call_rcu_tasks).
A regular RCU grace period is still required for all struct_ops maps
because the struct_ops trampoline ksyms requires a rcu grace period
(take a look at the list_del_rcu in __bpf_ksym_del).
Add a map_free_pre_rcu() callback so the struct_ops map can remove
ksyms before bpf_map_put() wait for the regular rcu grace period.
The tasks RCU grace period is only needed by tcp_congestion_ops.
Add free_after_tasks_rcu_gp only to struct bpf_struct_ops instead
of the bpf_map.
When CONFIG_TASKS_RCU=n, synchronize_rcu_tasks() is the same as
synchronize_rcu(). Since all struct_ops maps now complete a regular RCU
grace period before bpf_struct_ops_map_free() runs, skip the extra
synchronize_rcu_tasks() call in this case.
This cleanup prepares for a later patch that needs to support
free_after_mult_rcu_gp.
Signed-off-by: Martin KaFai Lau <martin.lau@kernel.org>
Signed-off-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Reviewed-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260917200542.3689605-3-ameryhung@gmail.com
|
|
Commit fa8fca88714c ("ipv4: validate IPV4_DEVCONF attributes properly")
added validation of IFLA_INET_CONF attributes, and in the process
changed the call of nla_for_each_nested() to nla_parse_nested(). A
side effect of this change is that the IFLA_INET_CONF option is now
tested for NLA_F_NESTED being set, and fails if it is not. Prior to the
commit there was no check of NLA_F_NESTED.
Change nla_parse_nested() to nla_parse(). This restores the previous
functionality of not checking NLA_F_NESTED, thereby allowing code that
(incorrectly) doesn't set NLA_F_NESTED to continue to work.
This issue was identified because keepalived started logging errors when
it was configuring macvlans that it created.
Fixes: fa8fca88714c ("ipv4: validate IPV4_DEVCONF attributes properly")
Signed-off-by: Quentin Armitage <quentin@armitage.org.uk>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260915213320.1527029-2-quentin@armitage.org.uk
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
inet_unhash() sets inet_csk(sk)->unhashed_state only when the
socket is hashed because tcp_set_state(sk, TCP_CLOSE) could be
called multiple times, e.g. tcp_abort() calls it directly and
tcp_done_with_error().
However, inet_twsk_hashdance_schedule() also unhashes a socket
when replacing it with twsk, allowing the socket to bypass
checks for inet_csk(sk)->unhashed_state.
Let's update inet_csk(sk)->unhashed_state there as well.
Fixes: 8cc3aef0cb19 ("tcp: Do not allow buggy transitions between ehash and lhash2.")
Reported-by: Daniel Zahka <daniel.zahka@gmail.com>
Closes: https://lore.kernel.org/netdev/DLHLRA8GVI5B.2Q1IRQG5BVJNZ@gmail.com/
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Daniel Zahka <daniel.zahka@gmail.com>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Link: https://patch.msgid.link/20260917191554.1600494-1-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
IP_MSFILTER reads its reply through ip_mc_msfget(), reached from
do_ip_getsockopt() and from nowhere else. Convert it, and build the
sockopt_t at the call site for as long as the caller still carries a
sockptr_t pair.
optlen here only has to cover the header, and the real reply size comes
from the imsf_numsrc field inside it. This is nasty, but userspace
relies on it, so sockopt_expand_out() preserves the same mechanism: it
grows optval only for a user address, and assumes the caller left room
for the size its own header asked for.
The *optlen store moves out of ip_mc_msfget() and into the call site,
guarded by !err so the -EINVAL, -ENODEV and -EADDRNOTAVAIL returns still
leave the caller's optlen word untouched.
The source list also moves from copy_to_sockptr_offset() to a sequential
copy_to_iter(). IP_MSFILTER_SIZE(0) and offsetof(struct ip_msfilter,
imsf_slist_flex) are both 16, so the bytes land where they did.
Signed-off-by: Breno Leitao <leitao@debian.org>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Link: https://patch.msgid.link/20260914-getsockopt_phase6-v2-2-e48befc9602e@debian.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
fib_select_multipath() compares nexthop_nh->nh_saddr against the flow
source address with no lock held, while fib_info_update_nhc_saddr()
stores a new value from another CPU as soon as the preferred source
address of the egress device changes.
Commit 195374d89368 ("ipv4: fib: annotate races around nh->nh_saddr_genid
and nh->nh_saddr") added WRITE_ONCE() on the store side and READ_ONCE()
in fib_result_prefsrc() after syzbot reported
BUG: KCSAN: data-race in fib_select_path / fib_select_path
but it only covered that reader. fib_select_multipath(), reached from
fib_select_path(), is a second lockless reader of nh->nh_saddr and was
left bare.
Moreover, nh_saddr is only meaningful when nh_saddr_genid matches
dev_addr_genid, as established by commit 436c3b66ec98 ("ipv4: Invalidate
nexthop cache nh_saddr more correctly."). fib_select_multipath()
skips that validation, so it can score a nexthop using a stale source
address and skew the ECMP selection.
Annotate both reads with READ_ONCE() and refresh the cached source
address via fib_info_update_nhc_saddr() when the genid does not match,
mirroring fib_result_prefsrc().
Fixes: 32607a332cfe ("ipv4: prefer multipath nexthop that matches source address")
Signed-off-by: Linkui Xiao <xiaolinkui@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260916125316.988044-1-xiaolinkui@126.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Cross-merge networking fixes after downstream PR (net-7.3-rc4).
Conflicts:
net/core/neighbour.c
979aabdad8dd0 ("neighbour: Skip default parms when resumed in neightbl_dump_info().")
7b430fcfc972f ("neighbour: Don't render blackhole_netdev via RTM_GETNEIGHTBL.")
fae1c59810b86 ("neighbour: Remove unnecessary net_eq().")
https://lore.kernel.org/20260911173056.44ec06e0@kernel.org
https://lore.kernel.org/aqfbJi7nAX4IbmnR@sirena.co.uk
Adjacent changes:
net/netlink/af_netlink.c
ceac0de741bf ("netlink: do not free nlk->groups while lockless readers can use it")
7c0ec6288b49 ("net: Replace %pK output with 0")
net/bridge/br_vlan.c
2842ce397dd0 ("net: bridge: vlan: fix bugs caused by switchdev deletion errors")
5bec8f861114 ("net: bridge: vlan: annotate lockless use of num_vlans")
2b1f8fd3118c ("net: bridge: vlan: annotate lockless vlan flags use")
net/bridge/br_mst.c
18a6fe05fb6e ("net: bridge: mst: move switchdev call outside rcu")
120207a08fc0 ("net: bridge: vlan: annotate lockless use of msti")
drivers/net/ethernet/stmicro/stmmac/hwif.h
90e4b849dfa6 ("net: stmmac: propagate FPE preemption-class mapping errors")
85ca3292d7a3 ("net: stmmac: Remove ARP offload code")
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Exclude old ACKs before SND.UNA from the tcp fast path
as well as ACKs after SND.NXT.
Such ACKs will fall through to the slow path, where tcp_ack()
performs the appropriate validation and challenge ACK handling
according to RFC5961 and Commit 3d501dd326fb1c7 ("tcp: do not
accept ACK of bytes we never sent").
This prevents old ACKs from being accepted
or modifying connection state as part of the fast path before
appropriate ACK validation is applied.
In particular, this prevents payload carried by a segment with
an excessively old ACK from advancing RCV.NXT before the ACK
is rejected.
Fixes: 31770e34e43d ("tcp: Revert "tcp: remove header prediction"")
Reported-by: Amit Klein <amit.klein@mail.huji.ac.il>
Reported-by: Tamir Shahar <tamir.shahar1@mail.huji.ac.il>
Reported-by: Inbal Schussheim <inbal.lipshtat@mail.huji.ac.il>
Suggested-by: Eric Dumazet <edumazet@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Inbal Schussheim <inbal.lipshtat@mail.huji.ac.il>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260914090408.1435080-2-inbal.lipshtat@mail.huji.ac.il
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
PSP conflicts with TLS ULP in its usage of both skb->decrypted and
sk->sk_validate_xmit_skb().
Make PSP mutually exclusive with TLS ULP, the only other user of either
of these. As other users of skb->decrypted come along, they can be added
to sk_has_decrypt_user(). It would make sense to also assert that
sk->sk_validate_xmit_skb() is also NULL in both of these setup paths for
similar future proofing, but the PSP listener/sk_clone() path is still
broken and it could be seen as a regression to not allow rx assoc to run
on a child of a listener socket with PSP tx assoc state.
Include all TCP ULPs in the sk_has_decrypt_user() check, even though TLS
is the only one that conflicts with PSP via the decrypted bit. This is
intentional because PSP was not designed to be used with ULPs. It is
best to close off surface area that may make bugs reachable, until
someone wishes to design and test an actual user of PSP with ULPs.
Fixes: 6b46ca260e22 ("net: psp: add socket security association code")
Signed-off-by: Daniel Zahka <daniel.zahka@gmail.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260915-psp-ktls-fix-v2-1-0eedc3b148ec@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/klassert/ipsec
Steffen Klassert says:
====================
pull request (net): ipsec 2026-09-16
1) xfrm: iptfs: fix stack OOB read in iptfs_skb_reset_frag_walk()
Add the up-front nr_frags guard iptfs_skb_add_frags() already has,
so an out-of-range offset can't walk past the on-stack frags[] array.
2) xfrm: serialize state GC with device state flush
Serialize xfrm_state destruction against the deferred-device pass
with a dedicated mutex, since the device GC list doesn't hold a state
reference and the two paths could free the same state.
3) xfrm: add missing RCU read lock in xfrm_send_migrate_state()
Hold the RCU read lock around xfrm_nlmsg_multicast() so the
rcu_dereference() of net->xfrm.nlsk doesn't warn.
4) xfrm: iptfs: fix runt reassembly panic from short inner tot_len
Require the runt length to cover at least the minimum IP header,
so a tot_len in [6, 19] (IPv4) can't write past the declared length
and trip skb_over_panic().
5) ipv6: xfrm: use full sockets in local error paths
Use skb_to_full_sk() in xfrm6_local_rxpmtu() and xfrm6_local_error()
and bail out without a full socket, so a TCP_NEW_SYN_RECV request_sock
isn't miscast as a full inet/IPv6 socket.
6) xfrm: fix compat ALLOCSPI request use-after-free
Drop the redundant alloc_compat() in xfrm_alloc_userspi() so the
compat translator no longer reads past the payload and publishes a
child a multicast clone can still see after xfrm_user_rcv_msg() frees.
7) xfrm: add missing rcu_read_lock(), skb_dst_force() and dev_hold() for xfrm_trans_reinject()
Force the dst before queuing, hold dev across the workqueue deferral,
and take rcu_read_lock() around the finish() loop, so transport-mode
reinjection doesn't deref non-refcounted dst/dev under workqueue.
8) xfrm: use hlist_del_init_rcu for state_cache and state_cache_input
Switch to hlist_del_init_rcu() so a second __xfrm_state_delete() is
a no-op instead of writing through LIST_POISON2, closing the UAFs.
9) esp: downgrade zerocopy managed frags before mutating skb frags
Call skb_zcopy_downgrade_managed() before ESP rewrites the skb frag
array, so per-frag unrefs in esp_ssg_unref() and skb_release_data()
stay balanced for ubuf-owned managed frags.
10) xfrm: hold net_device reference under RCU in bundle creation
Read dst->dev via dst_dev_rcu() and keep RCU active through
xfrm_fill_dst(), so a concurrent RTM_DELLINK can't free dev
under bundle creation.
11) xfrm: save input state data before secpath resets
Save the state protocol on the stack while it's still valid and
use the saved address family for transport_finish(), so post-reset
dereferences (VTI, XFRM if, MAX_DEPTH error) can't UAF the state.
12) net: xfrm: reject unrepresentable espintcp transport headers
Use the careful transport-header helper and drop the skb through
the XFRM error path when the offset can't be represented, instead
of silently truncating it.
* tag 'ipsec-2026-09-16' of git://git.kernel.org/pub/scm/linux/kernel/git/klassert/ipsec:
net: xfrm: reject unrepresentable espintcp transport headers
xfrm: save input state data before secpath resets
xfrm: hold net_device reference under RCU in bundle creation
esp: downgrade zerocopy managed frags before mutating skb frags
xfrm: use hlist_del_init_rcu for state_cache and state_cache_input
xfrm: add missing rcu_read_lock(), skb_dst_force() and dev_hold() for xfrm_trans_reinject()
xfrm: fix compat ALLOCSPI request use-after-free
ipv6: xfrm: use full sockets in local error paths
xfrm: iptfs: fix runt reassembly panic from short inner tot_len
xfrm: add missing RCU read lock in xfrm_send_migrate_state()
xfrm: serialize state GC with device state flush
xfrm: iptfs: fix stack OOB read in iptfs_skb_reset_frag_walk()
====================
Link: https://patch.msgid.link/20260916101938.118628-1-steffen.klassert@secunet.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The following command triggers a kernel panic:
ip link add d0 type dummy; ip link set d0 up
ip route add 10.30.0.0/16 \
encap ip id 300 geneve_opts 4660:66:11223344 dev d0
memcpy: detected buffer overflow: 4 byte write of buffer size 0
kernel BUG at lib/string_helpers.c:1044!
...
ip_tun_parse_opts.part.0.cold+0x10/0x10
ip_tun_build_state+0x116/0x2a0
On kernels built with GCC 15+ and `CONFIG_FORTIFY_SOURCE`, the fortified
`memcpy()` got 0 sized destination with request of 4 bytes length:
static int ip_tun_parse_opts_geneve(...)
{
...
attr = tb[LWTUNNEL_IP_OPT_GENEVE_DATA];
data_len = nla_len(attr); /* == 4 */
struct geneve_opt *opt = ip_tunnel_info_opts(info) + opts_len;
memcpy(opt->opt_data, nla_data(attr), data_len);
/* ^^^^^^^^^^^^^ 0 since options_len is assigned afterwards */
Fixed by initializing the counter before the options are referenced.
Matching what `tunnel_key_opts_set()` already does.
Fixes: bb5e62f2d547 ("net: Add options as a flexible array to struct ip_tunnel_info")
Cc: stable@vger.kernel.org
Signed-off-by: Gris Ge <cnfourt@gmail.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Reviewed-by: Gustavo A. R. Silva <gustavoars@kernel.org>
Link: https://patch.msgid.link/20260913090851.468216-1-cnfourt@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
We can hit a division by zero crash in tcp_rcvbuf_grow()
and tcp_rcv_space_adjust():
divide error: 0000 [#1] PREEMPT SMP
RIP: 0010:tcp_rcvbuf_grow+0x187/0x450 net/ipv4/tcp_input.c:939
...
grow = div_u64(((u64)rcvwin << 1) * (newval - oldval), oldval);
The division uses oldval = tp->rcvq_space.space as divisor.
When tp->rcvq_space.space is zero, this leads to a divide-by-zero
exception.
tp->rcvq_space.space is initialized in tcp_init_buffer_space():
tp->rcvq_space.space = min3(tp->rcv_ssthresh, tp->rcv_wnd,
(u32)TCP_INIT_CWND * tp->advmss);
If tcp_rmem[1] is configured to very small values (such as 1),
sk->sk_rcvbuf is initialized to 1. Then tcp_full_space(sk), which
computes (sk->sk_rcvbuf * scaling_ratio) >> 8, truncates to 0.
This sets tp->window_clamp = 0, tp->rcv_ssthresh = 0, and
tp->rcvq_space.space = 0. Later, when data arrives and DRS is invoked,
tcp_rcvbuf_grow() divides by oldval == 0.
Back in 2015, commit b1cb59cf2efe ("net: sysctl_net_core: check SNDBUF
and RCVBUF for min length") ensured that net.core.rmem_default and
net.core.rmem_max cannot be set below SOCK_MIN_RCVBUF. Similarly,
SO_RCVBUF setsockopt enforces max_t(int, val * 2, SOCK_MIN_RCVBUF).
However, net.ipv4.tcp_rmem still had .extra1 = SYSCTL_ONE, allowing
arbitrarily small values.
Because SOCK_MIN_RCVBUF depends on sizeof(struct sk_buff) and cacheline
alignment, its value varies across architectures and configuration options.
Using a fixed constant of 4096 ensures a predictable, architecture-
independent lower bound that is safely above SOCK_MIN_RCVBUF everywhere
and matches the documented 4K default.
Fix this by setting tcp_rmem.extra1 to 4096 and updating the documentation.
Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260912144848.3448026-1-edumazet@google.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
Commit a4d258036ed9 ("tcp: Fix race in tcp_poll") added smp_rmb() in
tcp_poll() and smp_wmb() in tcp_reset() (now tcp_done_with_error())
to ensure that if tcp_poll() observed socket closure, it would also
observe sk->sk_err.
Currently, tcp_poll() unconditionally executes smp_rmb() at the end
of every invocation, which on weakly-ordered architectures such as ARM64
emits a memory barrier instruction (dmb ishld) on the poll fast path,
even for healthy, active sockets.
However, tcp_poll() only needs this barrier if socket closure has been
observed, to ensure that the error code set by tcp_done_with_error()
before socket closure is visible before returning EPOLLERR.
Move smp_rmb() inside the conditional block handling socket closure
(shutdown == SHUTDOWN_MASK || state == TCP_CLOSE). For healthy
connected sockets in epoll, tcp_poll() avoids the barrier entirely.
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260913123224.762935-1-edumazet@google.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
ip_tunnel_delete_net() iterates ip_tunnel devices whose link_net
is dying and queues them for destruction.
The devices may reside in different netns.
Let's use unregister_netdevice_queue_net() to support per-netns
device unregistration.
Even after ip_tunnel_delete_net() queues a cross-netns ip_tunnel
device, ip_tunnel_changelink(), ip_tunnel_dellink(), and
ip_tunnel_ctl() could be called concurrently for it (once RTNL is
removed). In such a case, __rtnl_net_unlock() will perform the
unregistration.
Also, ip_tunnel_ctl() needs to check check_net(t->net), otherwise
it could create a new dev in dying netns after ip_tunnel_delete_net().
In the example below, we can see the fallback tunnel device (gre0)
and the cross-netns device (gre1) are unregistered by different
processes:
# bpftrace -e '#include <linux/netdevice.h>
kprobe:ip_tunnel_uninit {
$dev = (struct net_device *)arg0;
printf("PID: %d | DEV: %s%s\n", pid, $dev->name, kstack());
}
kprobe:ipgre_exit_rtnl {
printf("PID: %d%s\n", pid, kstack());
}' &
# ip netns add ns1
# ip netns add ns2
# ip -n ns1 link add name gre1 link-netns ns2 \
type gre local 192.168.0.1 remote 192.168.1.1
# ip netns del ns2
PID: 12
ipgre_exit_rtnl+5
ops_undo_list+702
cleanup_net+1122
process_scheduled_works+2538
...
PID: 12 | DEV: gre0 <------ fallback device (itn->fb_tunnel_dev).
ip_tunnel_uninit+5
unregister_netdevice_many_notify+7129
unregister_netdevice_many_net+1050
__rtnl_net_unlock+37
ops_undo_list+754
cleanup_net+1122
process_scheduled_works+2538
...
PID: 10 | DEV: gre1
ip_tunnel_uninit+5
unregister_netdevice_many_notify+7129
unregister_netdevice_many_net+1050
rtnl_net_work_func+136
process_scheduled_works+2538
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260912230043.2586313-8-kuniyu@google.com
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
struct ip_tunnel.net is the netns where encapsulated packets
flow into.
struct ip_tunnel is linked to ip_tunnel_net.tunnels[] of netns.
During netns dismantle or module unload, ip_tunnel_delete_net()
iterates the list and queues devices for destruction regardless
of the devices' netns.
Thus, once RTNL is removed, the list can be modified concurrently
from different netns due to device removal.
Let's protect it with per-netns mutex.
Note that dev_siocdevprivate() calls netdev_lock_ops() but
it must be NOP for tunnel devices to avoid AB-BA deadlock.
DEBUG_NET_WARN_ON_ONCE() is added to annotate the locking
explicitly.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260912230043.2586313-7-kuniyu@google.com
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
The next patch will introduce per-netns mutex and acquire it
in ip_tunnel_newlink() and ip_tunnel_changelink().
To make the diff cleaner, let's unify the error paths.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260912230043.2586313-6-kuniyu@google.com
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
With the previous patch, itn->fb_tunnel_dev can be removed
via ->dellink().
However, ioctl(SIOCDELTUNNEL) still uses unregister_netdevice(),
which requires ip_tunnel_del() in ip_tunnel_uninit().
Let's use ip_tunnel_dellink() everywhere to remove ip_tunnel device
and remove ip_tunnel_del() in ip_tunnel_uninit().
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260912230043.2586313-5-kuniyu@google.com
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
ip_tunnel_delete_net() no longer uses the 3rd argument,
struct rtnl_link_ops *ops.
Let's remove it.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260912230043.2586313-4-kuniyu@google.com
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
ip_tunnel_dellink() ignores itn->fb_tunnel_dev, so the per-netns
fallback tunnel device cannot be removed by userspace.
This also makes default_device_exit_batch() impossible to remove
the device since it calls ->dellink().
So, ip_tunnel_delete_net() has to iterate devices in the dying netns
and call unregister_netdevice_queue() directly.
But then, this duplicates ip_tunnel_del() in ip_tunnel_dellink()
and ip_tunnel_uninit().
Let's set itn->fb_tunnel_dev to NULL in ip_tunnel_delete_net() and
remove for_each_netdev_safe() in ip_tunnel_delete_net().
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260912230043.2586313-3-kuniyu@google.com
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
ipmr.c uses unregister_netdevice() to remove DVMRP tunnel devices
created in ipmr_new_tunnel().
This is fine because currently ip_tunnel_uninit() also calls
ip_tunnel_del() to unlink the device from the hash table.
However, we will move ip_tunnel_del() from ip_tunnel_uninit() to
ip_tunnel_dellink().
Removing DVMRP tunnel devices by unregister_netdevice() would leave
them in the hash table.
Let's call ->dellink for DVMRP tunnel devices.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260912230043.2586313-2-kuniyu@google.com
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
When the forward output route cannot be used in icmp_route_lookup(),
it enters the "reverse path" and calls ip_route_input() on fl4_dec.daddr,
the original packet's source address.
ip_route_input() only returns an error for truly invalid packets. For
unreachable addresses it will succeed and return an input route whose
dst.output is set to ip_rt_bug(). The existing check only rejects
RTN_LOCAL routes, so the RTN_UNREACHABLE route types can still be returned
and later used for output, syzkaller triggering a WARN_ON_ONCE()
in ip_rt_bug() as bellow:
------------[ cut here ]------------
WARNING: net/ipv4/route.c:1273 at ip_rt_bug+0x14/0x20
RIP: 0010:ip_rt_bug+0x14/0x20
Call Trace:
ip_push_pending_frames+0xfa/0x100
__icmp_send+0x905/0xf10
ip_options_compile+0xc0/0xd0
ip_rcv_finish_core+0x321/0xae0
ip_rcv+0x1de/0x260
__netif_receive_skb_one_core+0x11a/0x130
netif_receive_skb+0x7b/0x260
tun_get_user+0x11bf/0x1c10
------------[ cut here ]------------
Reject input route that is RTN_UNREACHABLE to fix it. The net warning
is only printed for RTN_LOCAL, as RTN_UNREACHABLE is not the result of
a race condition.
Fixes: 8b7817f3a959 ("[IPSEC]: Add ICMP host relookup support")
Suggested-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Dong Chenchen <dongchenchen2@huawei.com>
Link: https://patch.msgid.link/20260910140042.1880242-1-dongchenchen2@huawei.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
ipmr_vif_seq_show() and ip6mr_vif_seq_show() read vif->bytes_in,
vif->pkt_in, vif->bytes_out and vif->pkt_out with only rcu_read_lock()
held: since commit b96ef16d2f83 ("ipmr: convert /proc handlers to
rcu_read_lock()") both seq_start helpers are annotated __acquires(RCU)
and no longer take mrt_lock.
Those counters are updated from softirq context and the writers already
use WRITE_ONCE(): ipmr_prepare_xmit() and ip_mr_forward() on the IPv4
side, ip6mr_prepare_xmit() and ip6_mr_forward() on the IPv6 side. The
other lockless readers use READ_ONCE() as well - ipmr_ioctl(),
ipmr_compat_ioctl(), ipmr_fill_vif(), ip6mr_ioctl() and
ip6mr_compat_ioctl().
The two vif_seq_show() helpers are the only remaining bare readers, so
KCSAN flags them and the compiler is free to tear or reload the values
while the /proc/net/ip_mr_vif and /proc/net/ip6_mr_vif lines are being
formatted. Annotate them like the other readers; these are plain
statistics, no locking is needed.
Signed-off-by: Linkui Xiao <xiaolinkui@kylinos.cn>
Link: https://patch.msgid.link/20260910093452.2070079-1-xiaolinkui@126.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The forwarding paths report an expired TTL or hop limit as
SKB_DROP_REASON_IP_INHDR, the reason otherwise used for a header that is
malformed (ip_input.c, exthdrs.c, br_netfilter). Nothing else in the drop
path separates the two: IPSTATS_MIB_INHDRERRORS covers both, and the TTL
check runs before NF_INET_FORWARD, so netfilter tracing stops at
PREROUTING and never sees the drop.
The Fedora bug linked below shows how that reads in practice. The
reporter took kfree_skb(reason=IP_INHDR, loc=ip_forward) to mean the
software header checksum check had failed, and worked through RX checksum
offload, tc csum actions and both libvirt firewall backends before the
drops turned out to be replies arriving with TTL 1. ip_forward() never
verifies the header checksum; that runs earlier, in ip_rcv_core(), and
reports IP_CSUM.
TTL expiry is not a corner case -- every traceroute through a Linux
router goes through too_many_hops.
The three loopback hop limit checks in exthdrs.c drop with no reason at
all; give them the new one.
IPSTATS_MIB_INHDRERRORS stays as it is: RFC 1213 counts time-to-live
exceeded under ipInHdrErrors. The drop reason has no such constraint.
Link: https://bugzilla.redhat.com/show_bug.cgi?id=2517131
Signed-off-by: Junjie Cao <junjie.cao@intel.com>
Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Reviewed-by: Fernando Fernandez Mancera <fmancera@suse.de>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260910094937.536150-1-junjie.cao@intel.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Update ip_tunnel_encap_setup() to use WRITE_ONCE() when writing
to encap fields (type, sport, dport, flags) and hlen fields.
This ensures that concurrent lockless readers (like fill_info)
do not see torn writes.
Also remove the unsafe memset() on t->encap which could cause
concurrent readers to transiently see zeroed fields.
Removing it also fixes a bug where t->encap was left cleared
even if ip_encap_hlen() failed, resulting in partial configuration.
Fixes: 56328486539d ("net: Changes to ip_tunnel to support foo-over-udp encapsulation")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Acked-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Link: https://patch.msgid.link/20260907075846.2913645-4-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
bpf_sock_destroy() runs from the tcp iterator, under rcu_read_lock(). If
the sock is a listener that still has children in its accept queue,
tcp_abort() ends up in inet_csk_listen_stop() and the cond_resched()
there trips the debug check:
BUG: sleeping function called from invalid context at net/ipv4/inet_connection_sock.c:1523
in_atomic(): 0, irqs_disabled(): 0, non_block: 0, pid: 628, name: test_progs
preempt_count: 0, expected: 0
RCU nest depth: 1, expected: 0
locks held by test_progs/628: 3, last CPU#3:
#0: ffff8881158cee18 (&p->lock){+.+.}-{4:4}, at: bpf_seq_read+0x56/0x1210
#1: ffff8881106bb858 (sk_lock-AF_INET6){+.+.}-{0:0}, at: bpf_iter_tcp_seq_show+0x32b/0x4b0
#2: ffffffffb435af20 (rcu_read_lock){....}-{1:3}, at: bpf_iter_run_prog+0x46b/0xde0
CPU: 3 UID: 0 PID: 628 Comm: test_progs Tainted: G W 7.2.0+ #65 PREEMPT
Tainted: [W]=WARN
Call Trace:
<TASK>
dump_stack_lvl+0xc1/0xf0
dump_stack+0x10/0x20
__might_resched+0x3d2/0x610
inet_csk_listen_stop+0x7b/0xbf0
tcp_abort+0x23b/0x3b0
bpf_sock_destroy+0xfc/0x140
bpf_prog_448133d24601754f_iter_tcp6_server+0x81/0x8a
bpf_iter_run_prog+0x538/0xde0
bpf_iter_tcp_seq_show+0x26b/0x4b0
bpf_seq_read+0x424/0x1210
vfs_read+0x197/0xe40
ksys_read+0x119/0x240
__x64_sys_read+0x72/0xc0
x64_sys_call+0x647/0x27e0
do_syscall_64+0xe5/0x610
entry_SYSCALL_64_after_hwframe+0x76/0x7e
RIP: 0033:0x7fad39b28aca
RSP: 002b:00007ffc381c61c0 EFLAGS: 00000246 ORIG_RAX: 0000000000000000
RAX: ffffffffffffffda RBX: 00007ffc381c6a88 RCX: 00007fad39b28aca
RDX: 0000000000000032 RSI: 00007ffc381c6250 RDI: 0000000000000014
RBP: 00007ffc381c61e0 R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000003
R13: 0000000000000000 R14: 000055f077c1bbb0 R15: 00007fad3a0f3000
</TASK>
The commit that added the kfunc already guards lock_sock() in tcp_abort()
and udp_abort() with has_current_bpf_ctx(), but missed the listener path.
Do the same for the cond_resched(). The loop runs inside the iterator's
rcu_read_lock(), it must not reschedule or report a quiescent state there.
Fixes: 4ddbcb886268 ("bpf: Add bpf_sock_destroy kfunc")
Signed-off-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Link: https://lore.kernel.org/r/20260910112736.153710-1-jiayuan.chen@linux.dev
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
|
|
Cross-merge networking fixes after downstream PR (net-7.3-rc3).
Conflicts:
drivers/net/dsa/mt7530.c
3c18e3c9a54e ("net: dsa: mt7530: populate lpi_interfaces to fix EEE support")
10d9d8328e8a ("net: dsa: mt7530: replace mt7530_read with regmap_read")
Adjacent changes:
drivers/net/bonding/bond_alb.c
1746ef2e2df2 ("bonding: use skb_cow_head() in bond_do_alb_xmit() and rlb_arp_xmit()")
4cef95f72bbd ("bonding: fix u32 overflow in compute_gap()")
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|