summaryrefslogtreecommitdiff
path: root/include
AgeCommit message (Collapse)Author
8 hoursMerge tag 'pm-7.3-rc6' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm Pull power management fix from Rafael Wysocki: "Restore the previous behavior on systems where the cpufreq pressure was not visible in the scheduler and is not expected to be visible there. It became visible after a change made during the 7.2 development cycle that had gone too far" * tag 'pm-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm: cpufreq: intel_pstate: Fix max_freq fallback in cpufreq_update_pressure()
9 hoursMerge tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvmLinus Torvalds
Pull kvm fixes from Paolo Bonzini: "The most intrusive change is reverting a commit from 7.3-rc1 that made struct kvm a bit too large, and fixing the same issue otherwise. There are again a lot of selftests lines; the sheer number of commits is not small but I don't expect much more for 7.3 due to people travelling to Plumbers next week. ARM: - Take a reference on the last IRQ loaded into an LR to prevent it from being freed while running the guest (Marc Zyngier) - Ensure that the ITS MOVALL command only affects LPIs that were previously affined to the source redistributor (Marc Zyngier) - Fix + test for honoring the host's trap configuration when running non-protected VMs while KVM is in protected mode (Fuad Tabba) - Use the host stage-1 mapping granularity for VM_PFNMAP mappings at stage-2 (Mostafa Saleh) x86: Various bugfixes where the guest could do stupid things on purpose to cause problems in the host: - Failed VMRUNs can cause pending TLB flushes to be dropped, and in general some actions done through VMCB control fields have to be redone if VMRUN fails - Toggling MSR interceptions or eVMCS execution controls can cause the host to use a stale MSR permission bitmap - Bad page tables can cause a WARN. Also fix issues in last week's pull request (my fault, for changing email workflow and thus missing feedback sent to kvm@ but not LKML). Generic: - Take kvm_lock when creating vCPUs. For almost two decades everybody thought it was not done for some unspecified performance reasons, but in reality it was only done because kvm_lock was originally a spinlock. This is a better fix than 97d65b544f48 ("KVM: Check for duplicate vcpu_id as early as possible", from the 7.3 merge window), and does not waste 2K per VM, hence its inclusion here" * tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (29 commits) KVM: arm64: Use stage-1 leaf size for VM_PFNMAP KVM: arm64: selftests: Check a feature hidden in an ID register is UNDEF KVM: arm64: Use the host's HCR_EL2 for non-protected VMs in pKVM KVM: arm64: Clear HCR_EL2.RW for 32-bit non-protected vCPUs KVM: arm64: Apply the fine-grained UNDEFs without FEAT_FGT KVM: arm64: vgic-its: Fix MOVALL handling of source redistributor KVM: arm64: vgic: Take a refcount on IRQs referenced by last_lr_irq KVM: arm64: vgic: Allow last_lr_irq to be NULL when LRs are not overflowing KVM: SEV: Do cache maintenance on the source VM *before* clearing SEV state KVM: SEV: Nullify "have run CPUs" mask pointer when freeing it KVM: selftests: Extend nested x2APIC test to validate using eVMCS for vmcs12 KVM: selftests: Extend nested x2APIC test to validate disabling x2APIC virt KVM: selftests: Verify that L0's TPR doesn't get clobbered KVM: selftests: Run the nested x2APIC with and without APICv being inhibited in L2 KVM: selftests: Add x2APIC MSR test for inhibiting APICv while nested KVM: nVMX: Force MSR bitmap refresh if runtime eVMCS controls are modified KVM: SVM: Use the active VMCB's MSR bitmap when checking if MSR is intercepted KVM: SVM: Sync guest's PERF_CNTR_GLOBAL_CTL from h/w only on successful VMRUN KVM: SVM: Don't mark ASID fields as dirty when setting control.tlb_ctl KVM: SVM: Update control fields on #VMEXIT if and only if VMRUN succeeded ...
13 hoursMerge tag 'net-7.3-rc6' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net Pull networking fixes from Paolo Abeni: "Including fixes from Bluetooth, WiFi and netfilter. We are actively retargeting several non-urgent fixes towards next, but the traffic on the ML looks ever-increasing, and propagating the push-back towards subsystems is not immediate. No known outstanding regressions. Current release - regressions: - netfilter: nft_set_rbtree: skip transaction elements during GC Previous releases - regressions: - sched: cls_api: reclaim an empty proto on the error path - core: - fix checksum offsets in skb_splice_from_iter() - cap skb->queue_mapping when the tx queue is picked - page_pool: fix use-after-free in page_pool_recycle_ring_bulk() - wifi: - mac80211: fix slab-out-of-bounds read in ieee80211_monitor_select_queue() - mac80211: drop oversized fragments to avoid extra_len overflow - netfilter: - flowtable: restore ieee80211 forward path - bluetooth: hci_conn: Lock parent access during enhanced SCO setup - eth: - bcmgenet: allocate RX buffers as page fragments - stmmac: fix rx Scatter-Gather support - octeontx2-pf: fix aura BPID assignment when CONFIG_DCB is enabled - gve: DQO: accept TSO packets with non-protocol gso_type bits - r8169: disable EEE on RTL8168h/8111h Previous releases - always broken: - tcp: refresh TS.Recent for accepted old ACKs - wifi: - ath11k: reset ar->num_stations on hardware start - cfg80211: fix RTS threshold setting for single-radio PHY - bluetooth: btintel_pcie: fix plen overflow in btintel_pcie_recv_frame() - eth: bcmgenet: fix NULL dereference in set_coalesce before first open Misc: - Eric is retiring from google and updating his contact info" * tag 'net-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (96 commits) net: phy: aquantia: fix system interface type not updated in forced mode net: usb: qmi_wwan: add Rolling Wireless RN947R net: mvneta: clear XDP pfmemalloc flag between frames ipv6: sr: use skb_get_hash_net() in seg6_make_flowlabel() net/mlx5e: Fix AF_XDP TX timestamp teardown NULL dereference r8169: disable EEE on RTL8168h/8111h octeontx2-pf: Fix RSS indirection table size sctp: check RCV_SHUTDOWN after the sendmsg connect wait net: sparx5: make ports inherit the switch base mac address type net: microchip: vcap: stop scanning after deleting key field netfilter: flowtable: restore ieee80211 forward path netfilter: flowtable: generalize pending status bit netfilter: bpf: reject invalid NAT manipulation types netfilter: nft_set_rbtree: skip transaction elements during GC ipvs: filter some flags received in the backup server ipvs: do not create invisible templates ipvs: bound LBLCR and LBLC cache growth ipvs: fix missing counter decrement in lblc netfilter: nft_flow_offload: drop flowtable reference on init error path selftests: net: check timestamp echo after an old ACK ...
17 hoursMerge tag 'nf-26-09-30' of ↵Paolo Abeni
git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf Pablo Neira Ayuso says: ==================== Netfilter/IPVS fixes for net The following batch contains Netfilter fixes for net. This batch fixes crashes as recent feature regression, one of the due to a dependency that has been pulled into -stable: 1) Expand existing ipset fix for bitmap sets to disallow comments updates from kernel-side adds, from Florian Westphal. 2) Drop flowtable reference if nf_ct_netns_get() fails, otherwise flowtable cannot ever be removed, from Aohan Mei. 3) nft_rbtree GC should collect end elements that contained in this transaction batch, new or deleted elements are never expired. From Weiming Shi. 4) Restrict nf_nat_bpf so it does not set unknown NF_NAT_MANIP_* values, from Fernando F. Mancera. 5) Flowtable GC must skip flows that are pending hardware updates, generalize the PENDING flag and use it to inhibit GC. 6) Restore flowtable with ieee80211 which broke due to a relatively recent commit, which was pulled in by -stable, causing a regression in 6.18 kernels. And the following IPVS fixes: 1) Fix accounting of cache entries in IPVS LBLC for destinations, which eventually fills up the table and trigger recurrent resizing, from Julian Anastasov. 2) Limit IPVS cache growth for LBLCR and LBLC schedulers, from Zhiling Zou. 3) Restrict IP_VS_CONN_F_ONE_PACKET for normal connections, do not allow to use it with templates. Also from Julian. 4) Sanitize flags in IPVS sync messages received in the backup. From Julian Anastasov. netfilter pull request 26-09-30 * tag 'nf-26-09-30' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf: netfilter: flowtable: restore ieee80211 forward path netfilter: flowtable: generalize pending status bit netfilter: bpf: reject invalid NAT manipulation types netfilter: nft_set_rbtree: skip transaction elements during GC ipvs: filter some flags received in the backup server ipvs: do not create invisible templates ipvs: bound LBLCR and LBLC cache growth ipvs: fix missing counter decrement in lblc netfilter: nft_flow_offload: drop flowtable reference on init error path netfilter: ipset: do not update comments from kernel-side adds ==================== Link: https://patch.msgid.link/20260930074142.298353-1-pablo@netfilter.org Signed-off-by: Paolo Abeni <pabeni@redhat.com>
26 hoursnet: mvneta: clear XDP pfmemalloc flag between framesLorenzo Bianconi
mvneta_swbm_add_rx_fragment() sets XDP_FLAGS_FRAGS_PF_MEMALLOC on the xdp_buff when a fragment page is a pfmemalloc one (page under memory pressure). The xdp_buff is reused for the next frame, but only the XDP_FLAGS_HAS_FRAGS bit was cleared at frame start, so the pfmemalloc bit leaked from one frame into the following ones. mvneta_swbm_build_skb() propagates the flag to skb->pfmemalloc through xdp_update_skb_frags_info(), so the skb of a subsequent fragmented frame could be wrongly marked as pfmemalloc even if none of its pages are under pressure. Clear all the xdp_buff flags in mvneta_swbm_rx_frame(), which is invoked for each new frame, instead of just the XDP_FLAGS_HAS_FRAGS bit. Fixes: ed7a58cb40bd ("net: marvell: rely on xdp_update_skb_shared_info utility routine") Reviewed-by: Simon Horman <horms@kernel.org> Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com> Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com> Link: https://patch.msgid.link/20260929-mvneta-xdp-clear-frag-fix-v4-1-1e63b25eeed8@oss.qualcomm.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
44 hoursnetfilter: flowtable: restore ieee80211 forward pathPablo Neira Ayuso
Before commit 871df5007eda ("netfilter: flowtable: bail out if forward path cannot be discovered"), there was a fallback to set up a forward path in case .ndo_fill_forward_path fails or DEV_PATH_MTK_WDMA was used. Such fallback was used by commit d787a3e38f01 ("mac80211: add support for .ndo_fill_forward_path"). One possibility is to handle DEV_PATH_MTK_WDMA from the flowtable forward path discovery. However, this is only used internally by drivers to retrieve mtk_wdma information to set up hardware offload. Felix decided to use the .fill_forward_path interface for this purpose due to the lack of a better interface at that time. Add a new DEV_PATH_IEEE80211 path which is offered if the new ieee80211 flag is set on in the struct net_device_path_ctx to restore the flowtable with a ieee80211 netdevice. Handle this new DEV_PATH_IEEE80211 path just like DEV_PATH_ETHERNET and DEV_PATH_DSA, ie. this is the last netdevice in the stack. This new ieee80211 flag is implicitly unset for mtk_ppe and airoha which call dev_fill_forward_path() to retrieve a DEV_PATH_MTK_WDMA path. Fixes: 871df5007eda ("netfilter: flowtable: bail out if forward path cannot be discovered") Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
44 hoursnetfilter: flowtable: generalize pending status bitPablo Neira Ayuso
Rename NF_FLOW_HW_PENDING to NF_FLOW_PENDING and use it to inhibit the flowtable GC worker until pending hw offload work has been completed. Apparently, nf_flow_offload_stats() can schedule work to retrieve stats while the flow is being removed by GC. And this bit can also be used in a follow up patch to disable GC until the flow has been fully added in both directions. Revert the reordering done in commit d644b23afe1e ("netfilter: flowtable: publish HW_DEAD after worker is done") to prevent a race between GC and hw offload handler. Fixes: 2c8897953f3b ("netfilter: flowtable: Add pending bit for offload work") Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
2 daysnet: cap skb->queue_mapping when the tx queue is pickedJamal Hadi Salim
skbedit can set skb->queue_mapping and raise the per-CPU skip_txqueue flag so __dev_queue_xmit() honours the mapping. __dev_queue_xmit() cleared the flag before sch_handle_egress() and only read it afterwards, so the flag was not confined to the xmit that set it: a nested xmit (mirred redirect or mirror, or a drop after skbedit) could set the flag and the outer xmit would consume it for an skb that never went through skbedit. A forwarded packet still carries the ingress NIC's rx_queue + 1 in skb->queue_mapping, so the outer device then indexes its tx queue state with that stale value. Taprio's child array q->qdiscs[] is sized to the device's queue count, so taprio_enqueue() indexes past its allocation and dereferences the result as a struct Qdisc *. We (ab)use the skb->nf_skip_egress which means "skip netfilter egress for this packet" to tag to "am I in tc egress?". Despite the overload I dont see it as a conflict since the marker is set only around the single sch_handle_egress() call and ingress path is guarded by tc_at_ingress. I will send a followup(net-next) patch once this hits net-next to rename the skb->nf_skip_egress bit/flag to skb->skip_egress Arm the flag only from the egress classifier that can use it: raise skip_txqueue from tcf_skbedit_act() only when it runs inside sch_handle_egress(), thanks to skb->nf_skip_egress. An egress qdisc classifier runs in q->enqueue(), after the tx queue has been picked, so a mapping it sets cannot affect the current packet; arming the flag there only pollutes it for a later xmit. Then own the flag for the xmit frame the egress hook runs in: save the incoming value and clear it just before sch_handle_egress(), and restore it after the hook - on the consumed (drop) path, or, in the same call that reads it, on the surviving path. The save and the restores stay inside the egress_needed_key static branch, so a packet pays for them only when egress hooks are active (2f1e85b1aee4). Store the value netdev_cap_txqueue() selected back into skb->queue_mapping in netdev_tx_queue_mapping(), as netdev_core_pick_tx() already does, so the skip_txqueue path never hands a later reader on the xmit path a mapping the device cannot serve. A store made still later in the same frame, by a tc BPF program attached to a transmit qdisc, is outside this path and is not re-capped; a separate followup will resolve that path. netdev_xmit_skip_txqueue() returns the previous flag value so the save-and-clear is one call, and a no-op stub is provided when CONFIG_NET_EGRESS is disabled. skb->nf_skip_egress is compiled under CONFIG_NET_EGRESS rather than CONFIG_NETFILTER_SKIP_EGRESS, so skb_at_tc_egress() is valid whenever the egress path is built. A local user in a network namespace can redirect a packet from a device with more TX queues to one with fewer after setting a mapping valid only on the larger device. That reaches these reads and, under KASAN, faults with "slab-out-of-bounds in taprio_enqueue". Conditions to recreate the bug: the report's own trigger is a local user with CAP_NET_ADMIN in a network namespace, so no eBPF program is needed. With CONFIG_NET_SCH_TAPRIO=y, CONFIG_NET_ACT_SKBEDIT=y, CONFIG_NET_ACT_MIRRED=y, CONFIG_NET_CLS_MATCHALL=y, CONFIG_NET_SCH_PRIO=y and KASAN enabled, create qa (3 queues), qb (2 queues) and qc (1 queue) as dummy devices; put a taprio root on qb (num_tc 1, queues 2@0) and clsact on all three; then add an egress matchall filter on every device. On qa: "action skbedit queue_mapping 2 pipe action mirred egress redirect dev qb". On qb: "action mirred egress mirror dev qc". On qc: "action skbedit queue_mapping 0 pipe". Send one packet out qa. qc's skbedit sets the flag while qb's outer xmit is in flight; without the fix qb consumes it and reads its two-entry taprio child array with the forwarded packet's stale mapping. A qc whose skbedit is instead installed in a transmit-qdisc classifier (a matchall filter on the qc root qdisc) reaches the same code path the same way without the fix. Testing: on a KASAN build with panic_on_warn=1 the unfixed kernel panics with "BUG: KASAN: slab-out-of-bounds in taprio_enqueue", a read 0 bytes past a 16-byte taprio_init() allocation, for the clsact-setter and the transmit-qdisc-classifier reproducers and for a clsact skbedit-then-tc-BPF store; the fixed kernel runs all three with no report, and the BPF store variant additionally shows the expected "selects TX queue" clamp notice from the write-back. Fixes: 2f1e85b1aee4 ("net: sched: use queue_mapping to pick tx queue") Reported-by: Zero Day Initiative <zdi-disclosures@trendmicro.com> Link: https://lore.kernel.org/netdev/CANn89iLwYx8nCVf0pCEk_MmEiyC6kQaMwCQT9WkQVeeNzNQHqQ@mail.gmail.com/ Link: https://lore.kernel.org/netdev/179008581937.2160803.7117814290574262942@kernel.org/ Link: https://lore.kernel.org/netdev/179033713973.2160803.4914570693994398206@kernel.org/ Link: https://lore.kernel.org/netdev/20260925180407.63647514@kernel.org/ Link: https://lore.kernel.org/netdev/CANn89i+k-mZKDQVtvws_MEXeuMTAdaCcOXFZE-RfhcGTu90sjA@mail.gmail.com/ Suggested-by: Eric Dumazet <edumazet@google.com> Suggested-by: Jakub Kicinski <kuba@kernel.org> Tested-by: hybris <hybris@mojatatu.ai> Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com> Reviewed-by: Eric Dumazet <edumazet@google.com> Link: https://patch.msgid.link/QDISC-9R8V.v4.20260928081529@mojatatu.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2 daysMerge tag 'mtd/fixes-for-7.3-rc6' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/mtd/linux Pull MTD fixes from Miquel Raynal: "The most important set of fixes are around the handling of the QE bit in SPI NAND. There are also a couple of behavioral fixes (mutex issue in SPI-NOR, spurious bitflips on vf610_nfc, OOB bytes count in SPI NAND and cfi_cmdset stack usage). The rest is mostly AI fuzzing results" * tag 'mtd/fixes-for-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/mtd/linux: mtd: spinand: Do not update the QE bit on devices without one mtd: spi-nor: core: Fix mutex leak in spi_nor_rww_start_exclusive() mtd: rawnand: cadence: Initialize IRQ state before requesting IRQ mtd: rawnand: vf610_nfc: fix false bitflips on reads of erased pages mtd: rawnand: vf610_nfc: fix reads on chips with more than 64 bytes of OOB mtd: spinand: fix zero oobavail when no ECC engine is used mtd: spinand: fix NULL pointer dereference with no ECC engine mtd: mtd_intel_dg: reset poll counter for each erase mtd: cfi_cmdset_0001: shrink do_write_buffer() stack frame mtd: core: call _get_device() with the master MTD mtd: core: avoid double-free of OTP NVMEM device mtd: spinand: Enable QE on all dies mtd: block2mtd: Fix divide error when erase_size is zero
3 dayscpufreq: intel_pstate: Fix max_freq fallback in cpufreq_update_pressure()Rafael J. Wysocki
After commit d2d5c129d07e ("cpufreq: Make cpufreq_update_pressure() fall back to cpuinfo.max_freq"), cpufreq pressure appears in the CPU load balancer unexpectedly in some cases in which it was not present before, leading to confusion and uncertainty. Clearly, the scheduler assumes that cpufreq pressure will not be set unless the capacity reference frequency of the CPU is known, and the commit mentioned above violates that assumption. However, in some cases the capacity reference frequency of the CPU is in fact known even though arch_scale_freq_ref() returns 0 and in those cases it should be possible to set cpufreq pressure as appropriate. For this purpose, introduce a new cpufreq driver callback returning the CPU capacity reference frequency, .scale_freq_ref(), and make cpufreq_update_pressure() invoke it, if present, instead of falling back to cpuinfo.max_freq unconditionally. Add that callback to the intel_pstate driver and make it return 0 unless the scale-invariant capacity of the given CPU has been explicitly set, in which cases its reference frequency is always cpuinfo.max_freq. Fixes: d2d5c129d07e ("cpufreq: Make cpufreq_update_pressure() fall back to cpuinfo.max_freq") Reported-by: Jianyong Wu <wujianyong@hygon.cn> Closes: https://lore.kernel.org/linux-pm/20260915065747.1671965-1-wujianyong@hygon.cn/ Tested-by: Jianyong Wu <wujianyong@hygon.cn> Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com> Tested-by: Ricardo Neri <ricardo.neri-calderon@linux.intel.com> # Intel hybrid parts Tested-by: Chen Yu <yu.c.chen@intel.com> [ rjw: Add READ_ONCE() around a capacity_perf read ] Link: https://patch.msgid.link/12975163.O9o76ZdvQC@rafael.j.wysocki Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
3 daysnet: gso: limit recursive IP-in-IP segmentationZihan Xi
IP-in-IP GSO can re-enter inet_gso_segment() or ipv6_gso_segment() for each nested IP header. encap_level tracks header bytes, not callback depth, so a deep chain can exhaust the kernel stack. Making inet_gso_segment() stackable introduced unbounded IPv4 nesting; IPIP GSO/TSO later made the path reachable. The IPv6 stackable path was introduced separately and uses the same guard. Count IPv4 and IPv6 GSO handler entries in skb_gso_cb, initialized once per top-level GSO operation and preserved across GRE/UDP context changes. Use the existing IP_TUNNEL_RECURSION_LIMIT for both handlers. The first five entries pass, and the sixth returns -EINVAL before dispatching another GSO callback. Fixes: 3347c9602955 ("ipv4: gso: make inet_gso_segment() stackable") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Closes: https://lore.kernel.org/all/cover.1790157745.git.zihanx@nebusec.ai/ Assisted-by: LLM Co-developed-by: Luxing Yin <root@tr0jan.top> Signed-off-by: Luxing Yin <root@tr0jan.top> Signed-off-by: Zihan Xi <zihanx@nebusec.ai> Reviewed-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/20260924051521.32568-2-zihanx@nebusec.ai Signed-off-by: Paolo Abeni <pabeni@redhat.com>
3 daysMerge tag 'mm-hotfixes-stable-2026-09-27-19-12' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm Pull MM fixes from Andrew Morton: - Fix module loading incorrectly returning success after alloc_tag codetag setup failed - Restore MADV_COLLAPSE semantics for shmem so forced collapse ignores the shmem THP/mTHP sysfs settings - Make alloc_tag UAPI structure padding explicit - Fix DAMON schemes unexpectedly stopping after quotas are disabled through the online parameter update interface - Fix a false SW_TAGS KASAN invalid-access report when freeing vmapped task stacks - Split the MEMORY MANAGEMENT - MEMORY POLICY AND MIGRATION MAINTAINERS entry into separate MIGRATION and NUMA PLACEMENT entries - Move memory tiering maintenance under NUMA PLACEMENT - Add Gregory Price as a NUMA PLACEMENT co-maintainer - Add Heming Zhao as an ocfs2 reviewer - Fix mmap_prepare() state being copied onto a merged VMA rather than only onto a newly allocated VMA * tag 'mm-hotfixes-stable-2026-09-27-19-12' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm: mm/vma: predicate setting mmap_prepare VMA fields on new vma alloc MAINTAINERS: add Heming Zhao as ocfs2 reviewer MAINTAINERS: make Gregory a co-maintainer of MEMORY MANAGEMENT - NUMA PLACEMENT MAINTAINERS: move memory tiering under MEMORY MANAGEMENT - NUMA PLACEMENT MAINTAINERS: split up MEMORY MANAGEMENT - MEMORY POLICY AND MIGRATION kasan: unpoison task stack below watermark only in generic mode mm/damon/core: don't skip damos_adjust_quota() while esz is not zero alloc_tag: avoid implicit padding in uapi mm: shmem: ignore sysfs configs for shmem forced collapse module: fix lost error code from codetag_load_module()
3 daysMerge tag 'for-net-2026-09-28' of ↵Jakub Kicinski
git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth Luiz Augusto von Dentz says: ==================== bluetooth pull request for net: Core: - hci_core: Serialize ACL scheduling with channel deletion - hci_core: Serialize SCO and ISO scheduling with teardown - hci_core: Serialize fragmented ISO packet queueing - hci_core: Fix inquiry cache timestamps on 64-bit systems - hci_core: free the HCI ID if naming fails - hci_conn: Lock parent access during enhanced SCO setup - hci_sync: Fix inquiry cache use-after-free - hci_sync: don't drain cmd_sync backlog on unregister - RFCOMM: Fix initial port reference race - SMP: Serialize SMP remote OOB data access Drivers: - btintel: fix buffer over-read in btintel_hw_error() - btintel: validate DDC record lengths - btintel_pcie: fix plen overflow in btintel_pcie_recv_frame() - btintel_pcie: reject oversized TX packets in send_frame() * tag 'for-net-2026-09-28' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth: Bluetooth: hci_core: Serialize SCO and ISO scheduling with teardown Bluetooth: hci_core: Serialize ACL scheduling with channel deletion Bluetooth: Serialize SMP remote OOB data access Bluetooth: RFCOMM: Fix initial port reference race Bluetooth: hci_sync: Fix inquiry cache use-after-free Bluetooth: hci_sync: don't drain cmd_sync backlog on unregister Bluetooth: hci_core: Serialize fragmented ISO packet queueing Bluetooth: hci_conn: Lock parent access during enhanced SCO setup Bluetooth: btintel: validate DDC record lengths Bluetooth: hci_core: free the HCI ID if naming fails Bluetooth: hci_core: Fix inquiry cache timestamps on 64-bit systems Bluetooth: btintel_pcie: reject oversized TX packets in send_frame() Bluetooth: btintel_pcie: fix plen overflow in btintel_pcie_recv_frame() Bluetooth: btintel: fix buffer over-read in btintel_hw_error() ==================== Link: https://patch.msgid.link/20260928152919.942973-1-luiz.dentz@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
4 daysBluetooth: Serialize SMP remote OOB data accessChengfeng Ye
build_pairing_cmd() looks up remote OOB data and copies its contents without holding hdev->lock, which serializes the list's writers. After SMP finds an entry, a concurrent management Remove Remote OOB Data command can unlink and free it before SMP reads its present flag or copies its random and confirmation values. Removal can also invalidate an entry while the lookup is still traversing the list. KASAN reported: BUG: KASAN: slab-use-after-free in build_pairing_cmd+0x948/0x9b0 Call Trace: build_pairing_cmd+0x948/0x9b0 smp_recv_cb+0x459f/0x8110 l2cap_recv_frame+0xf14/0x9190 l2cap_recv_acldata+0xa64/0xd40 hci_rx_work+0x4ca/0x730 Allocated by task 87: hci_add_remote_oob_data+0x11d/0x530 add_remote_oob_data+0x282/0x400 hci_sock_sendmsg+0x1033/0x1ea0 Freed by task 93: hci_remote_oob_data_clear+0x108/0x1c0 remove_remote_oob_data+0x198/0x220 hci_sock_sendmsg+0x1033/0x1ea0 Taking hdev->lock in build_pairing_cmd() would recurse for callers that already hold it and invert the device-to-L2CAP lock order on the receive path. Add a per-device remote_oob_lock instead, held across the SMP lookup and copies and by the add, remove and clear helpers. Cover initialization and in-place updates as well, so SMP cannot read partially initialized or updated OOB values. Release the mutex on allocation failure, preserving the existing error return. The new critical sections acquire no device, connection or channel locks. Writers retain their existing hdev->lock protection, which continues to serialize the other readers without changing their locking or behavior. Link: https://lore.kernel.org/r/00660cd3-7d71-13a4-f617-229e6defb701@gmail.com Fixes: 02b05bd8b0a6 ("Bluetooth: Set SMP OOB flag if OOB data is available") Cc: stable@vger.kernel.org Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com> Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
4 daysMerge tag 'sched-urgent-2026-09-27' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull scheduler fixes from Ingo Molnar: - Fix LLC mis-scheduling bugs (Tim Chen, Lu Wang) - Fix cache-grouping related scheduling statistics UAF bugs (Tim Chen) - Skip kernel threads for cache aware scheduling to rubustify the code (Chen Yu) - Refresh LLC capacity across CPU hotplug, to fix capacity underestimation bug (Davi Chaves Azevedo) - Account PSI IRQ time to the execution context, not the scheduling context, to fix proxy scheduling accounting bug (Zhan Xusheng) * tag 'sched-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: sched/core: Account PSI IRQ time to the execution context, not the scheduling context sched/cache: Refresh LLC capacity across CPU hotplug, to fix capacity underestimation bug sched/cache: Skip kernel threads for cache aware scheduling to rubustify the code sched/cache: Introduce task_struct->sched_cache_grp to fix UAF sched/cache: Decouple sched_cache_group from mm to fix UAF sched/cache: Honor migrate_llc_task semantics in active load balance, to fix LLC mis-scheduling bug sched/cache: Keep nr_pref_llc_running in the runnable domain, to fix LLC mis-scheduling bug
4 daysMerge tag 'perf-urgent-2026-09-27' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull perf events fixes from Ingo Molnar: - Fixes for KVM guest PEBS virtualization (Sean Christopherson) - Fixes for various Intel PMUs related to PEBS data-source (Dapeng Mi) - Fix Intel Panther Cove event scheduling constraints (Dapeng Mi) - Fix Intel DMR/NVL OMR extra registers event scheduling (Dapeng Mi) - Rename two confusingly named PMU attributes (Dapeng Mi) - Fix a refcount leak in attach_perf_ctx_data() (Namhyung Kim) - Fix NULL pointer dereference crash in __perf_pmu_sched_task() (Puranjay Mohan) - Fix CPU-wide event scheduling (Puranjay Mohan) - Fix x86 LBR branch entry generation (Puranjay Mohan) * tag 'perf-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: perf/core: Fill branch entries with a single assignment perf/core: Run sched_task() for PMUs with only CPU-wide events perf/core: Fix NULL pmu_ctx passed to pmu->sched_task() perf/core: Fix a refcount leak in attach_perf_ctx_data() perf/x86/intel: Rename NVL offcore_rsp attribute to offmodule_rsp perf/x86/intel: Rename DMR offcore_rsp attribute to offmodule_rsp perf/x86/intel: Fix precise OMR event scheduling for DMR/NVL perf/x86/intel: Constrain Panther Cove UOPS_DISPATCHED events to PMCs 0-3 perf/x86/intel: Delete dead NVL PEBS data-source initcall perf/x86/intel: Fix Panther Cove PEBS data-source snoop states perf/x86/intel: Remove incorrect Panther Cove PEBS data-source constraints perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints perf/x86/intel: Update arw_latency_data() mem-op direction handling perf/x86/intel: Fix DKT PEBS load/store direction for latency events, to fix sample classification perf/x86/intel: Fix CMT PEBS load/store direction for latency events, to fix sample classification perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs() perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) if PEBS is unused perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has PEBS isolation, to fix stuck PEBS_ENABLED perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted PERF_GLOBAL_CTRL bits
5 daysalloc_tag: avoid implicit padding in uapiArnd Bergmann
The implied padding causes a harmless warning when testing the uapi headers with -Wpadded that could in theory indicate incompatibilities or data leaks: ./usr/include/linux/alloc_tag.h:41:1: error: padding struct size to alignment boundary with 7 bytes [-Werror=padded] The code here is fine, but it's better to make the padding explicit and avoid the warning here. Link: https://lore.kernel.org/20260916065830.1619425-1-arnd@kernel.org Link: https://lore.kernel.org/20260915202404.3568029-1-arnd@kernel.org Fixes: 1d581ab2348c ("alloc_tag: add ioctl to /proc/allocinfo") Signed-off-by: Arnd Bergmann <arnd@arndb.de> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Acked-by: Suren Baghdasaryan <surenb@google.com> Acked-by: SJ Park <sj@kernel.org> Acked-by: Hao Ge <hao.ge@linux.dev>
5 daysMerge tag 'ata-7.3-rc5' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux Pull ata fixes from Niklas Cassel: - Extend the quirk "no LPM on ATI" quirk, that is currently only applied for Samsung drives, to include AMD controllers as well. The AMD AHCI controllers are newer versions of the ATI AHCI controllers, and these controllers still have LPM issues with Samsung drives - LPM works with drives from other vendors (me) - Fix errors in the libata.force parameter documentation (me) - Verify the sense data descriptor lengths for ATA PASS-THROUGH command, so that a malicious device cannot write past the buffer length (Matthias) - Mention the libata for-next branch in MAINTAINERS such that the git ls-remote command done by get_maintainer.pl --self-test=scm can verify it (Matthias) * tag 'ata-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux: MAINTAINERS: name the libata/linux for-next branch ata: libata-scsi: bound the ATA passthru sense descriptor writes ata: libata: Correct libata.force parameter documentation ata: libata-core: Extend Samsung LPM quirk to AMD controllers
5 daysMerge tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvmLinus Torvalds
Pull kvm fixes from Paolo Bonzini: "Arm: - Invalidate the ITS translation cache when the guest changes the base address of the ITS tables (Fuad Tabba) - Skip saving ITS devices with device IDs that are out-of-bounds rather than failing the entire ITS save ioctl (Fuad Tabba) - Close race between VM teardown and invalidations of nested MMUs when handling MMU operations that are allowed to block (Lorenzo Stoakes) - Various fixes for the handling of the host's untrusted SVE configuration in pKVM (Fuad Tabba) - Make sure that empty SMCCC ranges based at 0 are rejected by the kvm_smccc_set_filter() (Karl Mehltretter) - Revoke the host mapping for pKVM's private stack pages, along with a new sanity check that all mappings in the hyp's private VA range have been correctly marked as hyp-owned (Fuad Tabba) - Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that concurrent vCPU initialization cannot relocate in-use MMUs. Defer the freeing of shadow stage-2 MMUs to the point that no other users (e.g. MMU notifier) could reference them (Marc Zyngier) - Drop useless WARN when rejecting an unsupported ioctl for pKVM (Fuad Tabba) - Fix the steal_time selftest to install correctly-sized mappings for non-4K hosts (Sebastian Ott) - Correct mapping of fine-grained trap for GCSPOPX instruction (Mark Brown) - Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit guests (Karl Mehltretter) RISC-V: - Synchronize hrtimer during VCPU teardown - Fix the conversion between vsip and hvip values - Serialize IMSIC attributes with vCPU migration - Release unused page after MMU invalidation - Propagate interrupted G-stage faults to KVM user-space as EINTR - Fix nested acceleration hfence entry update order - Fix sdata leak and stale snapshot_addr in snapshot_set_shmem - Preserve firmware counter value across PMU counter stop/start - Report PMU snapshot write failure to the guest - Fix perf-backed counter accounting across PMU stop and read - Correctly propagate error of a hart status SBI call s390: - Ensure that accesses through kvm_arch_set_irq_inatomic mark as dirty the pages that contain indicator and summary bits - Fix compile warning for kvm_s390_update_cmma_dirty() - Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to userspace - Move s390_kvm_mmu_commit_memory_region() into s390_kvm_mmu_prepare_memory_region() so that it can fail instead of WARN - Add missing srcu in kvm_s390_set_irq_state() - Fix potential races in storage functions - Fix race in _destroy_pages_crste() - Fix issues in the handling of KVM interrupt and page resources, when a queue that is assigned to a mediated device (mdev) is removed from the host's AP configuration - Fix loop condition in uv_find_secrets - Prevent potential out-of-bounds read x86: - Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs as valid on AMD - Fix a regression in the hardware disable selftest where it checked the wrong macro when detecting glibc support (breaks at least musl) - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially important for KVM_BUG_ON() flows, which often guard more dangerous bugs - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug where KVM would let userspace run a broken setup with stale vmcs12 pages - Fix a class of bugs where KVM would fail to fill kvm_run exit fields if getting nested pages failed - Treat reserved entries in the memory attributes xarray as "no attributes", to fix false positives when checking for mixed attributes - Fix memcg accounting for the memory attributes xarray (the xarray library subtly requires the xarray to be configured for accounting upfront; the gfp flags taken at runtime are used only rarely) - Don't pre-reserve xarray entries when storing empty attributes, as storing NULL must not require memory allocation (KVM and other subsystems heavily rely on this behavior) - Fix a memory leak and a cache maintenance issue related to doing intra-host migration on an SEV guest" * tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (54 commits) KVM: SEV: Do cache maintenance on the source VM during intra-host migration KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV KVM: Don't pre-reserve xarray entries when storing empty/NULL attributes KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg KVM: Don't treat reserved xarray entries as having memory attributes KVM: x86: Fill kvm_run exit fields in common get_nested_state_pages() error paths KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails KVM: arm64: Fix AArch32 DBGBXVR<n> handling KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP KVM: selftests: fix steal_time for arm64 with host page size > 4K KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction KVM: arm64: nv: Fix life cycle of the nested_mmus array KVM: arm64: Check every private mapping is hyp-owned at pKVM init KVM: arm64: Move the private VA allocation cursor to __io_map_next KVM: arm64: Match hyp text by physical address in fix_host_ownership() KVM: arm64: Transfer the hyp stack pages out of the host stage-2 KVM: arm64: selftests: Test empty SMCCC filter range at base 0 KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0 KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2 ...
6 daysRevert "KVM: Check for duplicate vcpu_id as early as possible"Sean Christopherson
Now that KVM uses kvm_get_vcpu_by_id() to check for an existing vCPU ID before doing any meaningful work, which was made possible by holding kvm->lock for the entirety of vCPU creation, revert the now-redundant "early" vCPU ID tracking. The claims about the impact of kvm->vcpu_ids on the memory footprint were a wee bit wrong: the worst case scenario isn't 256 bytes per VM, it's 256 "unsigned longs" per VM, i.e. 2048 bytes per VM. Increasing the size of "struct kvm" by 2048 nearly doubled the total size on many architectures, and tripped x86's KVM_SANITY_CHECK_VM_STRUCT_SIZE, which was added to detect this *exact* scenario, where a single change significantly increased the size of "struct kvm". I.e. attempting to build KVM with CONFIG_DEBUG_KERNEL=n fails on x86 (the build failures got missed because all build bots apparently test only CONFIG_DEBUG_KERNEL=y kernels, and maintainers' test flows were similarly lacking). This reverts commit 97d65b544f48b2ee49f6aea32145e3e7969955dc. Fixes: 97d65b544f48 ("KVM: Check for duplicate vcpu_id as early as possible") Reported-by: Jean-Christophe Guillain <jean-christophe@guillain.net> Closes: https://lore.kernel.org/all/56a4bc35ee605588b7cc36c8e45c12b5f3b506cb.camel@guillain.net Reported-by: Paweł S <spawel523@gmail.com> Closes: https://lore.kernel.org/all/CABD%3DWFOS4j4hDv%2BpW-eEM9HAM2q2GY_iYdAG%2BqvYcUEinUrcQQ@mail.gmail.com Tested-by: Jean-Christophe Guillain <jean-christophe@guillain.net> Signed-off-by: Sean Christopherson <seanjc@google.com> Tested-by: Naveen N Rao (AMD) <naveen@kernel.org> Message-ID: <20260921174445.911676-7-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
6 daysMerge tag 'kvm-x86-fixes-7.3-rc5' of https://github.com/kvm-x86/linux into HEADPaolo Bonzini
KVM fixes for 7.3-rcN - Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs as valid on AMD. - Fix a regression in the hardware disable selftest where it checked the wrong macro when detecting glibc support (breaks at least musl). - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially important for KVM_BUG_ON() flows, which often guard more dangerous bugs. - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug where KVM would let userspace run a broken setup with stale vmcs12 pages. - Fix a class of bugs where KVM would fail to fill kvm_run exit fields if getting nested pages failed. - Treat reserved entries in the memory attributes xarray as "no attributes", to fix false positives when checking for mixed attributes. - Fix memcg accounting for the memory attributes xarray (the xarray library subtly requires the xarray to be configured for accounting upfront; the gfp flags taken at runtime are used only rarely). - Don't pre-reserve xarray entries when storing empty attributes, as storing NULL must not require memory allocation (KVM and other subsystems heavily rely on this behavior).
6 daysMerge tag 'vfs-7.3-rc5.fixes' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs Pull vfs fixes from Christian Brauner: - Revert "put_mnt_ns(): leave mounts connected". This allows the creation of reference count cycles in a very trivial way. We can't bring this in until we have fixed the underlying cause - vfs: Don't create the private nullfs instance for kthreads under namespace_sem to avoid false lockdeps complaints - binfmt_misc: - Copy the name into a stack buffer and look up the copy in bpf_binprm_select_interp() - bpf_binprm_set_interp() and bpf_binprm_set_interp_arg(): Check the private copy instead so the string that gets staged is the kstring that was checked - netfs: - Make netfs_read_gaps() use separate sink folios rather than one reused sink folio to discard unwanted data so that cifs checksum checking sees all the data that was fetched - Trim reads down to i_size so afs symlinks read correctly from the cache - Wrap the direct mempool ->alloc() calls the GFP_KERNEL paths make in alloc_hooks() via a new mempool_alloc_noreserve() helper - iov_iter: Use iov_iter_alignment() for the start and length check added to iov_iter_extract_bvecs() this cycle. It used iter_iov_addr() and iter_iov_len() which are only valid for ITER_UBUF and ITER_IOVEC iterators - super: Make iterate_supers_type() deletion-safe - inode: Stop evict_inodes() from rescanning the same inodes - writeback: Bound the cleanup_offline_cgwb() rescans - ntfs3: Use d_instantiate_new() in ntfs_create_inode() - ovl: Fix a use-after-free in the ovl_do_mkdir() debug print - dcache: Unpoison the inline name buffer in __d_alloc() for KMSAN - autofs: Fix a pipe file reference leak in autofs_kill_sb() - bpf: Drop the path_unlink and path_rmdir hooks from the list of hooks for which the verifier rewrites bpf_{set,remove}_dentry_xattr() to the _locked variants - squashfs: Range check the xz dictionary size before shifting by it - selftests: Add the missing eventfd, open_tree_ns, openat2 and xattr filesystems selftests to TARGETS and drop the stale openat2 entry left behind when those tests moved * tag 'vfs-7.3-rc5.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs: netfs: Fix missing alloc tagging of direct mempool allocations bpf: fs/xattr: don't assume the inode is locked in path_unlink/path_rmdir autofs: fix sbi->pipe file reference leak in autofs_kill_sb() dcache: unpoison the inline name buffer in __d_alloc() ovl: fix UAF in ovl_do_mkdir() debug print super: make iterate_supers_type() deletion-safe Revert "put_mnt_ns(): leave mounts connected" Revert "selftests/filesystems: add mntns cleanup test" binfmt_misc: fix racy checks in bpf set_interp kfuncs binfmt_misc: fix OOB read in bpf_binprm_select_interp() fs: don't create the private nullfs mount under namespace_sem writeback: bound cleanup_offline_cgwb() rescans by rotating scanned inodes fs: avoid repeated scans in evict_inodes() netfs, afs: Fix symlink reading netfs: Fix netfs_read_gaps() to use separate sink folios squashfs: Add dictionary size range check to prevent shift-out-of-bounds fs/ntfs3: use d_instantiate_new() in ntfs_create_inode() and murder syzbot's "WARNING in do_new_mount" saga selftests/filesystems: fix missing and stale TARGETS entries block: Fix start and length check added to iov_iter_extract_bvecs()
6 daysnetfs: Fix missing alloc tagging of direct mempool allocationsHao Ge
Commit 1d78d56c43ef ("netfs: Fix folio_queue ENOMEM in writeback by adding a mempool") added a mempool for the folio_queues and made the request, subrequest and folio_queue allocations distinguish between writeback and everything else. Writeback is part of memory reclaim and must not fail due to ENOMEM, so it allocates under GFP_NOFS through mempool_alloc(), which may dip into the pool's reserve and, if that runs empty, wait for elements to be returned. The GFP_KERNEL paths, which can return -ENOMEM to their callers, invoke the pool's ->alloc() callback directly instead. The direct call, however, skips the alloc_hooks() wrapper that the mempool_alloc() macro provides. The pool callbacks, mempool_alloc_slab() and mempool_kmalloc(), call kmem_cache_alloc_noprof() and kmalloc_noprof() and rely on current->alloc_tag having been set by the caller. With CONFIG_MEM_ALLOC_PROFILING_DEBUG=y this leads to current->alloc_tag not set WARNING: ./include/linux/alloc_tag.h:161 at __alloc_tagging_slab_alloc_hook alloc_tag was not set WARNING: ./include/linux/alloc_tag.h:166 at __alloc_tagging_slab_free_hook at allocation and free time respectively, as reported when reading files on a CIFS mount. The allocations are also missing from /proc/allocinfo. Wrap the direct ->alloc() invocations in alloc_hooks() with a new mempool_alloc_noreserve() helper in include/linux/mempool.h, next to the other alloc_hooks()-wrapped macros such as mempool_alloc(). The GFP_KERNEL paths keep their failable allocation semantics, they just get tagged now. Fixes: 1d78d56c43ef ("netfs: Fix folio_queue ENOMEM in writeback by adding a mempool") Reported-by: Erhard Furtner <erhard_f@mailbox.org> Closes: https://lore.kernel.org/all/0b004319-9ef7-437c-a4dd-174d6a9a83db@mailbox.org/ Tested-by: Erhard Furtner <erhard_f@mailbox.org> Suggested-by: Suren Baghdasaryan <surenb@google.com> Cc: stable@vger.kernel.org Signed-off-by: Hao Ge <hao.ge@linux.dev> Link: https://patch.msgid.link/20260923063759.34667-1-hao.ge@linux.dev Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
7 daysMerge tag 'net-7.3-rc5' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net Pull networking fixes from Jakub Kicinski: "Including fixes from Bluetooth, NFC and Netfilter. Every week in this release is record-setting for number of posted patches. It doesn't seem like we're creating any regressions with all these fixes, three 'Fixes' tags here point to 7.2 commits but none are true regression fixes. We're trying to keep the count down, nonetheless. Previous releases - regressions: - net: don't require the hwtstamp NDOs when a PHY provides timestamping - ipv6: fix dst leak for uncached routes - vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE packets Previous releases - always broken: - packet: use ubuf_info completion for TX_RING packets - arp: terminate device name before lookup - ipv6: do not let ipv6_find_hdr() return an offset past the packet end - udp: remove a disconnected socket from the 4-tuple hash table - sctp: discard the rest of the packet on a stale-cookie error - eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged eswitch" [ And lots of other random network driver fixes ] * tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (189 commits) tcp: prevent collapsing skbs across boundary in rtx queue vlan: ensure sufficient headroom in vlan_dev_hard_header() net/sched: sch_teql: fix shadowed err in __teql_resolve() bridge: check llc_mac_hdr_init() return value in br_send_bpdu() llc: fix skb UAF and leaks on llc_mac_hdr_init() failure llc: reserve device headroom for allocated frames gve: DQO: reject TSO packets with an out of range MSS gve: fix TX drop when GSO MSS is too small for hw gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO net: flush skb_defer_nodes in dev_cpu_dead() net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails af_packet: fix integer overflow in prb_calc_retire_blk_tmo() tipc: Fix a data race on mon->peer_cnt in mon_timeout() net: phy: intel-xway: workaround 100BASE-TX Link-Up issue net/smc: fix UAF on lgr list traversal in smcr_port_err() net/rds: size a connection's path set by the transport it ends up with nfp: hold IPsec RX state under the XArray lock net: ena: fix MMIO read buffer leak on probe failure net: ena: fix PHC cleanup on probe failure net/sched: act_ct: fix helper UAF due to extensions realloc ...
7 daystcp: prevent collapsing skbs across boundary in rtx queueWillem de Bruijn
tcp_write_collapse_fence() sets TCP_SKB_CB(skb)->eor = 1 on tcp_write_queue_tail(sk) to prevent skbs queued after a switch to device encryption from being collapsed into earlier skbs. The fence is a no-op if all earlier data has already been transmitted when the switch happens: sk->sk_write_queue is empty. The not yet acknowledged earlier skbs wait in sk->tcp_rtx_queue with eor 0. On a subsequent retransmit or SACK shift, tcp_retrans_try_collapse() or tcp_shift_skb_data() can then merge an skb queued after the switch into one queued before it. Both users of the fence are affected: - psp: devices only encrypt skbs with skb->decrypted set. The merged skb keeps decrypted = 0 from the earlier skb, so merged data sent after psp_sock_assoc_set_tx() is retransmitted in cleartext. - tls device offload: the merged skb straddles the start marker set in tls_set_device_offload(). The software fallback (fill_sg_in() returns -EINVAL) and the mlx5, nfp and funeth drivers cannot handle such an skb and drop it. Every retransmit rebuilds the same skb, so the connection stalls. Fix this in two places, for defense in depth: 1. Fall back to tcp_rtx_queue_tail(sk) in tcp_write_collapse_fence() when tcp_write_queue_tail(sk) is NULL. 2. Check !skb_cmp_decrypted(to, from) in tcp_skb_can_collapse(), as tcp_skb_can_collapse_rx() does on receive. skb_shift(), which both collapse paths call, already has a DEBUG_NET_WARN_ON_ONCE() for this condition. Fixes: e8f69799810c ("net/tls: Add generic NIC offload infrastructure") Cc: stable@vger.kernel.org Signed-off-by: Willem de Bruijn <willemb@google.com> Reviewed-by: Eric Dumazet <edumazet@google.com> Reviewed-by: Daniel Zahka <daniel.zahka@gmail.com> Link: https://patch.msgid.link/20260924154427.953800-1-willemdebruijn.kernel@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
7 daysMerge tag 'landlock-7.3-rc5' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux Pull Landlock fixes from Mickaël Salaün: "This mainly fixes the Landlock tracepoint support merged this cycle so that denial and rule events report the intended policy context, whether through tracefs or BTF-visible callbacks. The size of this all is mainly from propagating the corrected contract through event definitions and producers, adding new tests for the reported context, and updating the documentation. Also improve annotation and fix a GCC 16 build warning" * tag 'landlock-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux: landlock: Widen ruleset versions to 64 bits landlock: Add counted_by in landlock_domain landlock: Fix tracepoint contract documentation selftests/landlock: Test network denial context selftests/landlock: Test filesystem denial blockers landlock: Report the effective signal number landlock: Report the actual ptrace tracer landlock: Fix network denial trace context landlock: Fix rule tracepoint context landlock: Fix filesystem denial blocker reporting landlock: Fix tracepoint fixed-width type names landlock: Work around gcc-16 -Wuninitialized warning
7 daysnet: openvswitch: conntrack: avoid modifying shared unconfirmed ct entryIlya Maximets
In a case where skb with an unconfirmed ct entry gets cloned, we may end up committing both but with different sets of extensions. The series of events: 1. The first clone wants to commit and runs the helpers wiring up the extension pointer into the expectation list. 2. Then it looses the confirmation keeping the entry unconfirmed. 3. Second clone now wants to commit labels and adds the new extension for that breaking the pointer in the expectation list causing UAF on the destruction path later. While this is possible to trigger, there should be no practical network pipeline where committing both clones without modifications into the same zone is needed. So, let's just reset the entry in case for some reason we got an skb with a shared one during commit. This doesn't affect any known use cases, but avoids any potential problems with sharing and modification of the unconfirmed ct entry. The fixes tag points to the introduction of helpers, since that's the main UAF trigger for the sharing. Fixes: cae3a2627520 ("openvswitch: Allow attaching helpers to ct action") Cc: stable@vger.kernel.org Reported-by: Axel Mierczuk <axel.mierczuk@1password.com> Signed-off-by: Ilya Maximets <i.maximets@ovn.org> Reviewed-by: Aaron Conole <aconole@redhat.com> Link: https://patch.msgid.link/20260921145655.3167436-2-i.maximets@ovn.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
7 daysMerge tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpfLinus Torvalds
Pull bpf fixes from Alexei Starovoitov: - Fix bpf_skb_change_tail() to drop the checksum offload instead of rejecting the trim of CHECKSUM_PARTIAL skbs (Daniel Borkmann) - Add KF_PERFMON kfunc flag and require CAP_PERFMON for kfuncs that read arbitrary memory and for untrusted read-only memory reads (Daniel Borkmann) - Clear scalar delta on narrowing stack spill (Daniel Borkmann) - Set up the frame pointer for the exception callback in arm64 JIT, and zero-fill other CPUs when BPF_F_CPU update creates a per-cpu hash element (Donggeun Yoo) - Various fixes (Emil Tsalapatis): - Fix bounds check underflow for skb-backed dynptrs - Fix rx_queue_mapping context access code generation in bpf_sock - Reject packet pointer arguments to subprogs that may mutate the packet - Reject ALU instructions that see arena and non-arena operands on different code paths - Fix copied_seq double-counting on sockmap self-redirect (Geliang Tang) - Fix divide-by-zero in btf_struct_walk() on a flexible array of zero-sized elements, fix out-of-bounds read of rtt_min in sock_ops (Jiayuan Chen) - Fix bpf_sock_destroy() out-of-bounds read of sk_protocol on TIME_WAIT and request socks, and sleeping under RCU when destroying a listener with pending children (Jiayuan Chen) - Fix JEQ/JNE with immediate operand in MIPS32 JIT and missing zero extension of BSWAP 16/32 in MIPS64 JIT (Johan Almbladh) - Avoid soft lockup in htab lookup[_and_delete] batch operations on large maps (Jose Fernandez) - Various fixes (Kumar Kartikeya Dwivedi): - Verify global subprogs in each sleepability context they are called from - Make post-verification instruction rewrites killable - Preserve packet pointer displacement in regsafe() - Apply CO-RE relocations before subprogram validation, restrict CO-RE poisoning to relocatable instructions, and reject truncated ldimm64 CO-RE relocations in libbpf - Assign lock identity to callback map values - Compare stack frames in regs_exact() - Bound ownership depth through local kptrs and graph roots - Fix u32 overflow in map batch operations when the map size exceeds 4GB (Masoud Aghasi) - Fix UAF in bpf memalloc due to concurrent consumption of ttrace lists in alloc_bulk() (Pu Lehui) - Allow gotox as the terminal instruction of a program or a subprogram (Siddharth Chintamaneni) - Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL, skip unsettled links in link iterator, and reject dev-bound-only programs on other devices (Weiming Shi) - Reject non-negative stack offsets in stack_slot_obj_get_spi() (Xu Yunxiang) - Check params size before reading reserved fields in bpf_crypto_ctx_create() (Yuqi Xu) - Reject max_entries > INT_MAX in sock_map_alloc() (Zhao Gongyi) - Use a 32-bit compare in xsk_map_gen_lookup() (Zhiling Zou) - Use kvfree() in xdp_test_run_teardown() (Zhixing Chen) * tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: (58 commits) selftests/bpf: Test per-cpu initialization of a BPF_F_CPU created element bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element bpf: Fix BSWAP 32 and 16 on MIPS64 bpf: Fix immediate JMP JEQ/JNE on MIPS32 bpf: Reject dev-bound-only programs on other devices bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc selftests/bpf: Test for mixed arena/nonarena code paths bpf: Prevent variable arena/non-arena register contents selftests/bpf: Test rejection of pkt args to mutating subprogs bpf: Reject pkt arguments in mutating subprogs selftests/bpf: Add selftests for rx_queue_mapping context access bpf: Fix bpf_sock context code generation selftests/bpf: Test dynptr slices past end of skb bpf: Fix bounds check for skb-backed dynptrs selftests/bpf: Reject iterator destruction through fp+0 bpf: Reject non-negative offsets in stack_slot_obj_get_spi() bpf: Check params size before reading reserved fields selftests/bpf: Check local object ownership depth bpf: Bound ownership depth through local kptrs and graph roots selftests/bpf: Cover frame changes in bounded loops ...
8 daysMerge tag 'nfc-7.3-rc5' of https://codeberg.org/linux-nfc/linuxJakub Kicinski
David Heidelberg says: ==================== NFC fixes for net 7.3-rc5 * tag 'nfc-7.3-rc5' of https://codeberg.org/linux-nfc/linux: nfc: pn533: fix OOB read in pn533_acr122_is_rx_frame_valid() nfc: llcp: fix slab-out-of-bounds reads when logging service names nfc: llcp: fix WKS SAP hijacking via prefix match in nfc_llcp_wks_sap() nfc: llcp: fix -ENOMEM on connect with zero-length service name nfc: st21nfca: validate ISO15693 inventory length nfc: fix use-after-free in nfc_get_local_general_bytes nfc: trf7970a: power down on startup RX gain failure nfc: llcp: fix sdreq TLV list leak on parse/alloc/send failure nfc: llcp: drop truncated I/RR/RNR PDUs in nfc_llcp_recv_hdlc() nfc: virtual_ncidev: Add missing ioctl compat handler selftests/nci: Fix out-of-bounds store on thread join selftests: nci: Fix uninitialized family ID on missing attribute nfc: llcp: Fix race condition in accept_queue lifecycle selftests: nci: Correct pthread_create return value check nfc: port100: reject frames whose declared length exceeds the received data nfc: llcp: Fix list corruption / refcount desync in nfc_llcp_recv_dm() nfc: st21nfca: validate received frame size nfc: nfcmrvl: validate helper command length before pull ==================== Link: https://patch.msgid.link/adeaccc1-cc04-4bb9-a28a-61a61d75ba14@ixit.cz Signed-off-by: Jakub Kicinski <kuba@kernel.org>
8 daysnet: ipv6: keep room for the mac header in dst_dev_overhead()Yuya Kusakabe
The seg6, ioam6 and rpl lwtunnels size their skb_cow_head() request as the length they are about to push plus dst_dev_overhead(), then push the new headers and rebuild the mac header below them with skb_mac_header_rebuild(). That rebuild needs skb->mac_len of headroom, but dst_dev_overhead() leaves LL_RESERVED_SPACE() of the egress device, 16 bytes for plain Ethernet. Where the mac header is longer than that, as it is on ingress through a VLAN device with reorder_hdr off, the rebuild runs out of room: skb_set_mac_header(skb, -skb->mac_len) computes a negative offset, stores it unchecked in the u16 skb->mac_header, and the memmove that follows writes skb->mac_len bytes about 64 KB past skb->head. Forwarding plain ping6 traffic through such a device reproduces it on all five seg6 encapsulation modes and on the rpl and ioam6 inline paths; skb->mac_header comes back as 65534 on a 704-byte head. Return the larger of the two. The helper already returns skb->mac_len when it has no dst, so this only makes the other branch agree, and it covers every caller rather than each call site in turn. Fixes: 40475b63761a ("net: ipv6: seg6_iptunnel: mitigate 2-realloc issue") Fixes: dce525185bc9 ("net: ipv6: ioam6_iptunnel: mitigate 2-realloc issue") Fixes: 985ec6f5e623 ("net: ipv6: rpl_iptunnel: mitigate 2-realloc issue") Suggested-by: Andrea Mayer <andrea.mayer@uniroma2.it> Signed-off-by: Yuya Kusakabe <yuya.kusakabe@gmail.com> Reviewed-by: Justin Iurman <justin.iurman@gmail.com> Reviewed-by: Gabriel Goller <g.goller@proxmox.com> Reviewed-by: Eric Dumazet <edumazet@google.com> Reviewed-by: Andrea Mayer <andrea.mayer@uniroma2.it> Link: https://patch.msgid.link/20260922-seg6-maclen-headroom-v3-1-7b2f982ef79d@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
8 daysnfc: fix use-after-free in nfc_get_local_general_bytesLuxiao Xu
Commit 6709d4b7bc2e ("net: nfc: Fix use-after-free caused by nfc_llcp_find_local") attempted to fix a use-after-free (UAF) issue by invoking nfc_llcp_local_put(local) after accessing local->gb. However, if the reference count drops to zero, local is freed immediately, leading to a use-after-free when callers access the returned pointer. Alternative approaches using dynamic allocation (e.g. kmemdup) introduced memory leaks because callers consistently treat the returned pointer as borrowed memory. Fix this properly by refactoring nfc_llcp_general_bytes() and nfc_get_local_general_bytes() to accept a caller-provided output buffer (out_gb) and its maximum length (gb_max_len). The general bytes are safely copied into out_gb before calling nfc_llcp_local_put(local), ensuring safe lifetime management without ownership transfer complications. Update all callers across drivers (microread, pn533, pn544, st21nfca, digital_dep, and nci) to provide their own destination buffers and pass them to nfc_get_local_general_bytes(). Fixes: 6709d4b7bc2e ("net: nfc: Fix use-after-free caused by nfc_llcp_find_local") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Assisted-by: LLM Signed-off-by: Luxiao Xu <rakukuip@gmail.com> Signed-off-by: Ren Wei <weir@nebusec.ai> Reviewed-by: Simon Horman <horms@kernel.org> Link: https://patch.msgid.link/3cbaac3bee23f8ff3a3284ed32d347696eb1d208.1788841683.git.rakukuip@gmail.com Signed-off-by: David Heidelberg <david@ixit.cz>
9 daysBluetooth: hci_core: Fix inquiry cache timestamps on 64-bit systemsLinmao Li
On 64-bit systems, outgoing BR/EDR connections always fall back to page scan repetition mode R2 with no clock offset once the system has been up for more than five minutes, even when inquiry found the peer only seconds earlier. This lengthens paging and can increase the risk of a Page Timeout. The inquiry cache stores jiffies in __u32 timestamps, but its age helpers subtract them from unsigned long jiffies. INITIAL_JIFFIES casts -300 * HZ through unsigned int, so jiffies crosses 2^32 five minutes after boot on 64-bit systems. Assigning it to __u32 then drops the upper 32 bits. In one trace, a 7.6-second-old entry (HZ=1000) was reported as 2^32 + 7620 ticks old and rejected by hci_acl_create_conn_sync(). hci_inquiry() is affected by the same truncation when checking the whole cache. 32-bit systems are unaffected because unsigned long is 32 bits wide there. Use unsigned long for both timestamps so they have the same width as jiffies on 32-bit and 64-bit systems, and update the debugfs format specifier accordingly. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Cc: stable@vger.kernel.org Signed-off-by: Linmao Li <lilinmao@kylinos.cn> Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
9 daysperf/core: Fill branch entries with a single assignmentPuranjay Mohan
perf_clear_branch_entry_bitfields() clears the bitfields of struct perf_branch_entry one by one and leaves from/to alone, since callers overwrite those straight away. The list has to be kept in sync with the struct by hand and has already fallen behind: new_type and priv were added to perf_branch_entry and never added here. Only BRBE writes those two, and neither for every record. brbe_set_perf_entry_type() leaves new_type alone for a branch type it does not recognise, and priv is not set for source-only records. arm_pmuv3.c allocates the per-CPU branch stack with kmalloc(), so such a record reaches userspace with whatever the slot held: uninitialised kmalloc() data on the first pass over the buffer, the previous record's values after that. Nothing under arch/x86/events/ writes either field, so only arm64 is affected. Assign the whole entry at each site instead. Everything not named is then zero, and there is no list to keep in sync. The bitfields add up to exactly 64 bits, so the struct has no padding to leave undefined. perf_clear_branch_entry_bitfields() has no callers left, so remove it. perf_entry_from_brbe_regset() assigns an empty literal instead, since it fills from/to conditionally. PERF_BR_SPEC_NA is 0, so dropping the explicit spec assignment changes nothing. Fixes: b190bc4ac9e6 ("perf: Extend branch type classification") Fixes: 5402d25aa571 ("perf: Capture branch privilege information") Suggested-by: Peter Zijlstra <peterz@infradead.org> Signed-off-by: Puranjay Mohan <puranjay@kernel.org> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Tested-by: Yifan Wu <wuyifan50@huawei.com> Link: https://patch.msgid.link/20260810133540.1947118-4-puranjay@kernel.org
9 dayspacket: use ubuf_info completion for TX_RING packetsWillem de Bruijn
tpacket_snd sends skbs with frags pointing into its ring slots. Slots are released when skb->destructor is called. A call to skb_orphan calls skb->destructor before the skb is freed. This can cause the slot to be reused while still linked into the skb. Switch to standard zerocopy completion (ubuf_info) so the slot is only released once all references to the payload are freed or copied. Restore skb->destructor to standard sock_wfree. The ubuf_info completion callback can be called with a NULL skb, but only from net_zcopy_put and related API, used by zerocopy implementations that hold their own reference on the uarg, such as MSG_ZEROCOPY. This uarg is only ever completed from skb_zcopy_clear, so skb is always set. To prevent userspace from aliasing in-flight state on shared ring slots, allocate tpacket_uarg per packet, rather than per slot. This adds a small allocation to the transmit path. Use standard kmalloc to allow backporting to stable kernels. The uarg holds an sk_wmem_alloc reference, rather than an sk_refcnt reference. packet_free_tx_ring waits on sk_wmem_alloc before freeing the ring pages. Always allocate vec->deferred for tx_ring so page-backed rings also wait on sk_wmem_alloc when skb_copy_ubufs drops page refs before calling tpacket_ubuf_complete. Drop the tx_ring.pg_vec test that tpacket_destruct_skb performed before accessing the slot. The sk_wmem_alloc reference now guarantees that the slot is valid. The test is also not sufficient by itself, as it reads pg_vec without pg_vec_lock, so it can race with packet_set_ring. As a result a slot is released when its payload is copied, which can be before transmission (e.g., in skb_orphan_frags_rx). If copied before skb_tx_timestamp() is called, no slot timestamp is recorded, similar to when skb_orphan() was called early in the datapath before this patch. Revert the now unused previous skb_zcopy_.._nouarg infra. Depends on commit 992cc9f94ca9 ("net/packet: defer vmalloc TX_RING free until skbs finish"). Reported-by: Katherine Leaver <kleaver@janestreet.com> Reported-by: Bjoern Doebel <doebel@amazon.de> Closes: https://lore.kernel.org/netdev/20260909085542.3370986-1-doebel@amazon.de/ Fixes: 5cd8d46ea156 ("packet: copy user buffers before orphan or clone") Cc: stable@vger.kernel.org Signed-off-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/20260919004748.1463985-3-willemdebruijn.kernel@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
9 daysbpf: Prevent variable arena/non-arena register contentsEmil Tsalapatis
The verifier marks ALU instructions that include at least one arena operand with needs_zext: These instructions are fixed up after verification to be ALU32 instructions to ensure that the result is a valid offset into an arena. However, different code paths may provide two non-arena 64-bit arguments to the same instruction. The result of the operation in that code path is wrong, since it is now unexpectedly truncated to 32 bits and zero-extended. Add logic to the verifier to ensure every instruction either always has at least one PTR_TO_ARENA argument, or never does. Since needs_zext already tracks the first scenario, add a prevent_zext field in bpf_insn_aux to track the latter. Reject instructions that use arena arguments and have prevent_zext set, or do not have arena arguments and have needs_zext set. Fixes: 6082b6c328b5 ("bpf: Recognize addr_space_cast instruction in the verifier.") Reported-by: Nicholas Carlini <nicholas@carlini.com> Suggested-by: Nicholas Carlini <nicholas@carlini.com> Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260922172028.6269-8-emil@etsalapatis.com
9 daysbpf: Fix bounds check for skb-backed dynptrsEmil Tsalapatis
The skb_pointer_if_linear() function checks whether a memory region of length len starting at offset off into the skb is in the linear area, and returns a pointer to the region if so. The check currently subtracts between skb_headlen and offset of the check, and since skb_headlen is unsigned the subtraction can underflow. This causes the bounds check to spuriously pass and generate an arbitrary pointer of the form *(skb->data + off). The only user of this helper is currently skb-backed BPF dynptr code. Returning the wrong pointer leads to the dynptr erroneously being backed with invalid memory. Ensure the subtraction cannot underflow, and fail the check if it would. Use u64 arithmetic to also prevent overflow when calculating (skb_headlen(skb) - off) since off is unsigned. Fixes: 6f5a630d7c57 ("bpf, net: Introduce skb_pointer_if_linear().") Reported-by: Nicholas Carlini <nicholas@carlini.com> Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev> Link: https://patch.msgid.link/20260922172028.6269-2-emil@etsalapatis.com
9 dayslandlock: Widen ruleset versions to 64 bitsMickaël Salaün
Tracepoint consumers use a ruleset ID and version to identify the successful landlock_add_rule(2) call prefix used to create a domain. LANDLOCK_MAX_NUM_RULES bounds distinct stored rules, not successful calls: re-adding already-present rights for an object or port succeeds without increasing num_rules. Because every successful call increments the version, these calls can wrap the 32-bit counter and give different prefixes the same trace identity. Widen the counter and its trace fields to 64 bits so the counter cannot wrap in practice, while preserving the successful-call semantics. Saturating would alias all subsequent histories, while rejecting a call at the limit would change otherwise valid syscall behavior solely for trace metadata. Cc: Günther Noack <gnoack@google.com> Cc: Steven Rostedt <rostedt@goodmis.org> Fixes: 63747c94774d ("landlock: Add landlock_add_rule_fs and landlock_add_rule_net tracepoints") Reviewed-by: Günther Noack <gnoack@google.com> Link: https://patch.msgid.link/20260922132615.1025945-1-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>
10 dayssched/cache: Refresh LLC capacity across CPU hotplug, to fix capacity ↵Davi Chaves Azevedo
underestimation bug The scheduler scales LLC capacity by the fraction of cache-sharing CPUs covered by a domain: llc_bytes = cache_size * span_weight / shared_weight During CPU teardown, sched_cpu_deactivate() rebuilds scheduler domains before cacheinfo_cpu_pre_down() removes the CPU from shared_cpu_map. The new domains therefore use the old sharing weight. The later call to sched_update_llc_bytes() looks up the departing CPU's sd_llc, which has already been detached, and returns without correcting the surviving CPUs. On a Ryzen 5 7535U with twelve logical CPUs sharing a 16 MiB LLC, offlining one SMT sibling left the remaining CPUs with: llc_bytes = floor(16777216 * 11 / 12) = 15379114 bytes The correct capacity is still 16777216 bytes. On systems with active cache-aware scheduling, an underestimated capacity can cause exceed_llc_capacity() to reject aggregation for a process whose footprint would fit. Unchanged cpuset partitions sharing the physical cache can also retain stale capacity when a CPU comes online in another partition. Pass the cache-sharing mask already retained by cacheinfo to the scheduler update. Refresh every surviving CPU using its own LLC domain so that each partition receives the correct share. This also preserves the correction needed as cache-sharing maps grow during boot. Keep the existing CPU-hotplug and scheduler-domain synchronization. The update remains on the hotplug path; no steady-state scheduling operation or persistent allocation is added. Fixes: 7030513a0877 ("sched/cache: Calculate the LLC size and store it in sched_domain") Signed-off-by: Davi Chaves Azevedo <davichazbh@gmail.com> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Reviewed-by: Chen Yu <yu.c.chen@intel.com> Reviewed-by: Tim Chen <tim.c.chen@linux.intel.com> Reviewed-by: K Prateek Nayak <kprateek.nayak@amd.com> Tested-by: Chen Yu <yu.c.chen@intel.com> Tested-by: K Prateek Nayak <kprateek.nayak@amd.com> Cc: <stable@kernel.org> # v7.2.x Link: https://patch.msgid.link/6751d93e15889e624796c74db0bfe66603d60b1b.1790035273.git.tim.c.chen@linux.intel.com
10 dayssched/cache: Introduce task_struct->sched_cache_grp to fix UAFTim Chen
Add a sched_cache_grp pointer to task_struct so that scheduler code can access the cache group directly via the task, without going through mm->sched_cache_grp. This decouples the scheduler's hot-path accesses from the mm_struct. Each task holds its own refcount on the sched_cache_group, separate from the reference held by its mm_struct. The reference is acquired in copy_mm() (fork) and exec_mmap() (exec), and released in exit_mm(). This fixes the use-after-free when account_mm_sched() reaches the group through a task whose mm is being switched, as reported by Hyunwoo: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Convert all scheduler code in fair.c and exit.c to use p->sched_cache_grp instead of p->mm->sched_cache_grp. Keep the fork/exec/exit reference management out of the generic mm paths: add sched_cache_fork(), sched_cache_fork_cleanup(), sched_cache_exec_mmap() and sched_cache_exit_mm() in kernel/sched/cache_sched.c (with empty stubs for !CONFIG_SCHED_CACHE), so fs/exec.c, kernel/fork.c and kernel/exit.c each call one helper instead of open-coding the refcounting under #ifdef. Also add sched_cache_group_get() and task_cache_group_get(). Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing") Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Closes: https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@linux.dev/ Reported-by: Hyunwoo Kim <imv4bel@gmail.com> Reported-by: Zenghui Yu (Huawei) <zenghui.yu@linux.dev> Co-developed-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Cc: <stable@kernel.org> #7.2.x Link: https://patch.msgid.link/ae7081dc54736bf115215f9867abb2711a7403fb.1790035273.git.tim.c.chen@linux.intel.com
10 dayssched/cache: Decouple sched_cache_group from mm to fix UAFTim Chen
Currently the sched cache grouping is by mm and the scheduling statistics sched_cache_stat lives in the mm structure. This ties the life cycle of scheduling stats with mm. In account_mm_sched(), the scheduling stats are accessed by task->mm->sc_stat. However, a task may be switching mm on one CPU when another CPU is running account_mm_sched(), and possibly accessing the old mm that was freed. This problem was found when running tests with KASAN by Hyunwoo: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Instead of serializing the mm access by introducing extra acquisition of rq lock in the mm free path, extract sched_cache_stat from mm_struct, rename it as sched_cache_group and manage its life cycle apart from mm_struct with its own ref counting. This allows us in the next patch access sched_cache_group directly from task, and add a refcount on sched_cache_group when a task links to it. This prevents the use after free issue when accessing stale and released old mm and its sched cache stat a task switches to a new mm while account_mm_sched() is done elsewhere. The other benefit of this restructure is in the future, the grouping of tasks to a LLC would have the flexibility to be associated with a user defined grouping, or cgroup, cookie group, numa_group or others instead of just with a single mm address space. Rename sched_cache_stat to sched_cache_group and turn it into a refcounted object allocated from mm_struct. The mm_struct now holds a pointer (sched_cache_grp) to this object instead of embedding it. Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing") Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Closes: https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@linux.dev/ Reported-by: Hyunwoo Kim <imv4bel@gmail.com> Reported-by: Zenghui Yu (Huawei) <zenghui.yu@linux.dev> Co-developed-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Cc: <stable@kernel.org> #7.2.x Link: https://patch.msgid.link/91fd1e3266707c865bc9abecfb3e17bc676712df.1790035273.git.tim.c.chen@linux.intel.com
10 daysipv6: Fix dst leak for uncached routes.Kuniyuki Iwashima
ip6_route_output_flags(), ip6_rt_put_flags(), and ip6_dst_check() detect an uncached route by list_empty(&rt->dst.rt_uncached), which replaced the static DST_NOCACHE flag check in commit a4c2fd7f7891 ("net: remove DST_NOCACHE flag"). When a device is unregistered, rt6_uncached_list_flush_dev() unlinks uncached routes tied to the device from rt6_uncached_list. Previously, they were moved to another list with list_move() (__list_del_entry() + list_add()), and since commit 98aa546af5e4 ("inet: remove (struct uncached_list)->quarantine"), the routes are just unlinked with list_del_init(). If list_del_init() runs concurrently, list_empty() evaluates to true; ip6_route_output_flags() calls dst_hold_safe() incorrectly and ip6_rt_put_flags() skips ip6_rt_put(), leaking dst, and thus dev tied via rt->from as well. The same race is partially fixed by commit 9a6f0c4d5796 ("dst: fix races in rt6_uncached_list_del() and rt_del_uncached_list()"). Let's check rt6->dst.rt_uncached_list instead. Note that IPv4 does not have the same issue. Fixes: 98aa546af5e4 ("inet: remove (struct uncached_list)->quarantine") Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com> Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn> Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Reviewed-by: Eric Dumazet <edumazet@google.com> Link: https://patch.msgid.link/20260920191558.2990636-1-kuniyu@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
11 daysata: libata-core: Extend Samsung LPM quirk to AMD controllersNiklas Cassel
A Samsung SSD 870 QVO 8TB connected to an AMD 600 Series chipset SATA controller is reported to time out on STANDBY IMMEDIATE during system suspend with med_power_with_dipm enabled. The command completes when using max_performance instead. The existing Samsung LPM quirk only matches ATI controllers, leaving AMD controllers unaffected. Rename it to ATA_QUIRK_NO_LPM_ON_ATI_AND_AMD and extend the vendor check to AMD for the same Samsung SSD model patterns. Keep LPM behavior unchanged for other controller vendors, including Intel. Leave ATA_QUIRK_NO_NCQ_ON_ATI restricted to ATI, since the reported AMD issue concerns LPM rather than NCQ. Link: https://bugzilla.kernel.org/show_bug.cgi?id=221986 Reviewed-by: Damien Le Moal <dlemoal@kernel.org> Reviewed-by: Mario Limonciello (AMD) <superm1@kernel.org>> --- Link: https://lore.kernel.org/r/20260918124030.1962773-5-cassel@kernel.org Signed-off-by: Niklas Cassel <cassel@kernel.org>
12 daysMerge tag 'input-for-v7.3-rc3' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/dtor/input Pull input fixes from Dmitry Torokhov: - Fixes for evdev and input compat handling to zero-initialize on-stack absinfo and force-feedback effect structures before partial or compat copies from userspace, preventing kernel stack memory disclosure - Fixes for the Synaptics RMI4 driver to prevent an out-of-bounds read when writing multi-chunk blocks over SMBus and to avoid a NULL pointer dereference during suspend/resume when the RMI device is unbound - Fixes for the soc_button_array driver to propagate -EPROBE_DEFER on non-Bay Trail/Cherry Trail platforms (fixing broken power and volume buttons on the Microsoft Surface Pro 11) and to validate the ACPI package element count before dereferencing - A fix for the adp5588-keys driver to cache the initial GPIO hardware state before registering the gpiochip so pre-configured pin states are not clobbered by GPIO hogs during registration - A fix for the cyttsp5 touchscreen driver to clamp the device-supplied HID report size before copying into the response buffer, preventing a buffer overflow - A fix for the HP SDC serio driver to use timer_shutdown_sync() on module exit so the periodic kicker timer cannot rearm itself during teardown - A fix for the eeti_ts touchscreen driver to export its OF module alias so the module autoloads on Device Tree platforms - Updates to the xpad joystick driver adding support for the Victrix Pro BFG controller and Azeron devices, and fixing the device type classification for the PDP Marvel Xbox 360 controller - Quirks for the i8042 and atkbd drivers to keep the built-in keyboards functional on the Acer Aspire Go 15 AG15-42P and Xiaomi Redmi Book Pro 16 2026 - A quirk for the Synaptics PS/2 touchpad driver disabling SMBus InterTouch on the Lenovo ThinkPad T440p (board ID 2722) so the touchpad and TrackPoint respond immediately at boot - Other minor updates and documentation fixes, including reading the "ti,poll-period" property as u32 in tsc2007, adding the mt6572 compatible to the MediaTek keypad Device Tree binding, fixing an attribute name typo in the trackpoint sysfs ABI documentation, and documenting that no new LED codes should be added to the input subsystem * tag 'input-for-v7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/dtor/input: Input: hp_sdc - shut down kicker timer on module exit Input: xpad - add support for Victrix Pro BFG Controller Input: tsc2007 - read "ti,poll-period" as u32 Input: trackpoint - fix the inertia attribute name in the ABI document Input: eeti_ts - publish the OF module alias Input: xpad - add support for Azeron devices Input: xpad - fix PDP Marvel Xbox 360 controller Input: document that no new LED codes should be added Input: soc_button_array - check btns_desc->package.count Input: soc_button_array - fix MS Surface Pro 11 probe failure Input: i8042 - add quirk for Acer Aspire Go 15 AG15-42P Input: synaptics - disable InterTouch on ThinkPad T440p (board id 2722) Input: cyttsp5 - clamp the HID report size before memcpy Input: zero ff_effect before compat copy in input_ff_effect_from_user Input: evdev - zero absinfo before partial copy in EVIOCSABS Input: synaptics-rmi4 - fix GPF in suspend and resume when unbound Input: rmi_smbus - fix out-of-bounds read in rmi_smb_write_block() Input: atkbd - skip deactivate for Xiaomi Redmi Book Pro 16 2026 dt-bindings: input: mediatek,mt6779-keypad: add mt6572 Input: adp5588-keys - cache GPIO state before registering the gpiochip
12 dayslandlock: Fix tracepoint contract documentationMickaël Salaün
The tracepoint documentation claims that denial and lifecycle events expose every input needed to reproduce a verdict. Instead document how denial, ruleset, and domain events identify the denying policy, checked operation and object, and reason for denial. Direct consumers to generic tracepoints for additional operational context. State the reconstruction limits: IDs are boot-local, rule checks have no request ID, and exported records may be lost or cross-CPU reordered. Also replace the incorrect BPF_RAW_TRACEPOINT guidance with libbpf SEC("tp_btf/...") attachment and refer consumers to the event prototypes for callback argument layouts. Cc: Günther Noack <gnoack@google.com> Cc: Steven Rostedt <rostedt@goodmis.org> Link: https://patch.msgid.link/20260918185036.608651-10-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>
12 dayslandlock: Report the effective signal numberMickaël Salaün
The signal-scope denial callback identifies its target but not the effective signal. This loses permission-probe signal zero and makes the file-owner hook's zero sentinel ambiguous. Append an int signal argument to the typed-BPF callback. Preserve sig, including zero, in hook_task_kill(). In hook_file_send_sigiotask(), translate signum zero to SIGIO at the producer, where its meaning is known. Carry the effective signal and target domain ID in a private, stack-backed context consumed synchronously. This requires no allocation or task reference in the interrupt-capable file-owner path. Gate this context and the remaining scope-only domain IDs with CONFIG_TRACEPOINTS. Keep the tracefs record and audit output unchanged. Cc: Günther Noack <gnoack@google.com> Cc: Steven Rostedt <rostedt@goodmis.org> Fixes: bb91730f16c0 ("landlock: Add tracepoints for ptrace and scope denials") Link: https://patch.msgid.link/20260918185036.608651-7-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>
12 dayslandlock: Report the actual ptrace tracerMickaël Salaün
The ptrace denial callback identifies only the tracee. Current is the tracer during hook_ptrace_access_check(), but it is the tracee during PTRACE_TRACEME, where the parent is the actual tracer. A consumer therefore cannot infer both parties from the existing arguments. Append the actual tracer task to the typed-BPF callback: current for hook_ptrace_access_check() and parent for hook_ptrace_traceme(). Carry it with the tracee domain ID in a private ptrace context. Both hooks keep the selected tasks alive through synchronous dispatch, so no extra task reference is needed. Keep the tracefs record unchanged. The new context is available only to typed BPF, while same_exec continues to describe the tracer that owns the denying policy. Cc: Günther Noack <gnoack@google.com> Cc: Steven Rostedt <rostedt@goodmis.org> Fixes: bb91730f16c0 ("landlock: Add tracepoints for ptrace and scope denials") Link: https://patch.msgid.link/20260918185036.608651-6-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>
12 dayslandlock: Fix network denial trace contextMickaël Salaün
Network denial events report source and destination ports reconstructed from audit data. Their zero values are ambiguous, and neither identifies the complete endpoint that Landlock checked. Carry the checked sockaddr and its signed length in a private trace-only context. For an enabled event, validate the length and copy only the initialized prefix into zeroed local storage. This prevents a typed BPF program from reading uninitialized bytes while exposing the socket family, socket, address, and length. Replace the source and destination trace-record fields with one signed port derived from the checked address. A value of -1 means that no port was checked, zero is a valid port, and positive values use host endianness. Bind blockers select the bind address; connect and send blockers select the destination. Cc: Günther Noack <gnoack@google.com> Cc: Steven Rostedt <rostedt@goodmis.org> Fixes: 01ce260f5ccf ("landlock: Add landlock_deny_access_fs and landlock_deny_access_net") Link: https://patch.msgid.link/20260918185036.608651-5-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>
12 dayslandlock: Fix rule tracepoint contextMickaël Salaün
Name each event after the identity it reports. Add-rule events describe UAPI rule insertion, so rename them after LANDLOCK_RULE_PATH_BENEATH and LANDLOCK_RULE_NET_PORT. Check-rule events describe matches in internal rule trees, so rename them after LANDLOCK_KEY_INODE and LANDLOCK_KEY_NET_PORT. This remains accurate if multiple UAPI rule types share one lookup and stored rule. Keep denial event names based on filesystem and network families because they describe final access decisions. Use u64 for growable access masks passed by value to add-rule and check-rule typed BTF callbacks. CO-RE can relocate pointer-reached fields, but it cannot widen a scalar callback slot declared by a BPF program. Keep native access_mask_t for internal state and trace records. For add-rule callbacks, report the normalized per-call contribution passed to landlock_insert_rule() and expose the complete validated flags value. Put the ruleset and flags first as a common invocation prefix. This distinguishes duplicate and effective-zero additions without recovering arguments from saved syscall registers. Cc: Günther Noack <gnoack@google.com> Cc: Steven Rostedt <rostedt@goodmis.org> Fixes: 63747c94774d ("landlock: Add landlock_add_rule_fs and landlock_add_rule_net tracepoints") Fixes: 3f1f106e4c14 ("landlock: Add tracepoints for rule checking") Link: https://patch.msgid.link/20260918185036.608651-4-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>
12 dayslandlock: Fix filesystem denial blocker reportingMickaël Salaün
Filesystem topology denials are rendered with an empty blockers value because their blocker is identified by the request type instead of an access mask. Introduce the private struct landlock_blockers to carry the request type and final missing access mask to filesystem and network denial tracepoints. Copy both members into named trace-record fields, then use the type to print change_topology for topology denials while preserving symbolic access masks for ordinary denials. The request type lets typed BPF consumers distinguish topology denials from access denials. Keeping the native access mask in a pointer-reached field also lets CO-RE adjust existing programs' load width if access_mask_t grows. Cc: Günther Noack <gnoack@google.com> Cc: Steven Rostedt <rostedt@goodmis.org> Fixes: 01ce260f5ccf ("landlock: Add landlock_deny_access_fs and landlock_deny_access_net") Link: https://patch.msgid.link/20260918185036.608651-3-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>
12 dayslandlock: Fix tracepoint fixed-width type namesMickaël Salaün
The new Landlock tracepoints use UAPI-prefixed __u32 and __u64 names for callback arguments and record fields, including internal IDs that are not Landlock UAPI values. Typed BPF consumers see callback typedef names through BTF. Use the kernel u32 and u64 aliases before release so the tracepoint contract does not present internal values as Landlock UAPI types. This changes BTF-visible typedef spelling but not integer widths, calling conventions, tracefs formats, or record layouts. Cc: Günther Noack <gnoack@google.com> Cc: Steven Rostedt <rostedt@goodmis.org> Link: https://patch.msgid.link/20260918185036.608651-2-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>