| Age | Commit message (Collapse) | Author |
|
git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm
Pull power management fix from Rafael Wysocki:
"Restore the previous behavior on systems where the cpufreq pressure
was not visible in the scheduler and is not expected to be visible
there.
It became visible after a change made during the 7.2 development cycle
that had gone too far"
* tag 'pm-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm:
cpufreq: intel_pstate: Fix max_freq fallback in cpufreq_update_pressure()
|
|
Pull kvm fixes from Paolo Bonzini:
"The most intrusive change is reverting a commit from 7.3-rc1 that made
struct kvm a bit too large, and fixing the same issue otherwise.
There are again a lot of selftests lines; the sheer number of commits
is not small but I don't expect much more for 7.3 due to people
travelling to Plumbers next week.
ARM:
- Take a reference on the last IRQ loaded into an LR to prevent it
from being freed while running the guest (Marc Zyngier)
- Ensure that the ITS MOVALL command only affects LPIs that were
previously affined to the source redistributor (Marc Zyngier)
- Fix + test for honoring the host's trap configuration when running
non-protected VMs while KVM is in protected mode (Fuad Tabba)
- Use the host stage-1 mapping granularity for VM_PFNMAP mappings at
stage-2 (Mostafa Saleh)
x86:
Various bugfixes where the guest could do stupid things on purpose to
cause problems in the host:
- Failed VMRUNs can cause pending TLB flushes to be dropped, and in
general some actions done through VMCB control fields have to be
redone if VMRUN fails
- Toggling MSR interceptions or eVMCS execution controls can cause
the host to use a stale MSR permission bitmap
- Bad page tables can cause a WARN.
Also fix issues in last week's pull request (my fault, for changing
email workflow and thus missing feedback sent to kvm@ but not LKML).
Generic:
- Take kvm_lock when creating vCPUs. For almost two decades everybody
thought it was not done for some unspecified performance reasons,
but in reality it was only done because kvm_lock was originally a
spinlock.
This is a better fix than 97d65b544f48 ("KVM: Check for duplicate
vcpu_id as early as possible", from the 7.3 merge window), and does
not waste 2K per VM, hence its inclusion here"
* tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (29 commits)
KVM: arm64: Use stage-1 leaf size for VM_PFNMAP
KVM: arm64: selftests: Check a feature hidden in an ID register is UNDEF
KVM: arm64: Use the host's HCR_EL2 for non-protected VMs in pKVM
KVM: arm64: Clear HCR_EL2.RW for 32-bit non-protected vCPUs
KVM: arm64: Apply the fine-grained UNDEFs without FEAT_FGT
KVM: arm64: vgic-its: Fix MOVALL handling of source redistributor
KVM: arm64: vgic: Take a refcount on IRQs referenced by last_lr_irq
KVM: arm64: vgic: Allow last_lr_irq to be NULL when LRs are not overflowing
KVM: SEV: Do cache maintenance on the source VM *before* clearing SEV state
KVM: SEV: Nullify "have run CPUs" mask pointer when freeing it
KVM: selftests: Extend nested x2APIC test to validate using eVMCS for vmcs12
KVM: selftests: Extend nested x2APIC test to validate disabling x2APIC virt
KVM: selftests: Verify that L0's TPR doesn't get clobbered
KVM: selftests: Run the nested x2APIC with and without APICv being inhibited in L2
KVM: selftests: Add x2APIC MSR test for inhibiting APICv while nested
KVM: nVMX: Force MSR bitmap refresh if runtime eVMCS controls are modified
KVM: SVM: Use the active VMCB's MSR bitmap when checking if MSR is intercepted
KVM: SVM: Sync guest's PERF_CNTR_GLOBAL_CTL from h/w only on successful VMRUN
KVM: SVM: Don't mark ASID fields as dirty when setting control.tlb_ctl
KVM: SVM: Update control fields on #VMEXIT if and only if VMRUN succeeded
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Paolo Abeni:
"Including fixes from Bluetooth, WiFi and netfilter.
We are actively retargeting several non-urgent fixes towards next,
but the traffic on the ML looks ever-increasing, and propagating the
push-back towards subsystems is not immediate.
No known outstanding regressions.
Current release - regressions:
- netfilter: nft_set_rbtree: skip transaction elements during GC
Previous releases - regressions:
- sched: cls_api: reclaim an empty proto on the error path
- core:
- fix checksum offsets in skb_splice_from_iter()
- cap skb->queue_mapping when the tx queue is picked
- page_pool: fix use-after-free in page_pool_recycle_ring_bulk()
- wifi:
- mac80211: fix slab-out-of-bounds read in ieee80211_monitor_select_queue()
- mac80211: drop oversized fragments to avoid extra_len overflow
- netfilter:
- flowtable: restore ieee80211 forward path
- bluetooth: hci_conn: Lock parent access during enhanced SCO setup
- eth:
- bcmgenet: allocate RX buffers as page fragments
- stmmac: fix rx Scatter-Gather support
- octeontx2-pf: fix aura BPID assignment when CONFIG_DCB is enabled
- gve: DQO: accept TSO packets with non-protocol gso_type bits
- r8169: disable EEE on RTL8168h/8111h
Previous releases - always broken:
- tcp: refresh TS.Recent for accepted old ACKs
- wifi:
- ath11k: reset ar->num_stations on hardware start
- cfg80211: fix RTS threshold setting for single-radio PHY
- bluetooth: btintel_pcie: fix plen overflow in btintel_pcie_recv_frame()
- eth: bcmgenet: fix NULL dereference in set_coalesce before first open
Misc:
- Eric is retiring from google and updating his contact info"
* tag 'net-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (96 commits)
net: phy: aquantia: fix system interface type not updated in forced mode
net: usb: qmi_wwan: add Rolling Wireless RN947R
net: mvneta: clear XDP pfmemalloc flag between frames
ipv6: sr: use skb_get_hash_net() in seg6_make_flowlabel()
net/mlx5e: Fix AF_XDP TX timestamp teardown NULL dereference
r8169: disable EEE on RTL8168h/8111h
octeontx2-pf: Fix RSS indirection table size
sctp: check RCV_SHUTDOWN after the sendmsg connect wait
net: sparx5: make ports inherit the switch base mac address type
net: microchip: vcap: stop scanning after deleting key field
netfilter: flowtable: restore ieee80211 forward path
netfilter: flowtable: generalize pending status bit
netfilter: bpf: reject invalid NAT manipulation types
netfilter: nft_set_rbtree: skip transaction elements during GC
ipvs: filter some flags received in the backup server
ipvs: do not create invisible templates
ipvs: bound LBLCR and LBLC cache growth
ipvs: fix missing counter decrement in lblc
netfilter: nft_flow_offload: drop flowtable reference on init error path
selftests: net: check timestamp echo after an old ACK
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf
Pablo Neira Ayuso says:
====================
Netfilter/IPVS fixes for net
The following batch contains Netfilter fixes for net. This batch
fixes crashes as recent feature regression, one of the due to a
dependency that has been pulled into -stable:
1) Expand existing ipset fix for bitmap sets to disallow comments
updates from kernel-side adds, from Florian Westphal.
2) Drop flowtable reference if nf_ct_netns_get() fails, otherwise
flowtable cannot ever be removed, from Aohan Mei.
3) nft_rbtree GC should collect end elements that contained in
this transaction batch, new or deleted elements are never
expired. From Weiming Shi.
4) Restrict nf_nat_bpf so it does not set unknown NF_NAT_MANIP_*
values, from Fernando F. Mancera.
5) Flowtable GC must skip flows that are pending hardware updates,
generalize the PENDING flag and use it to inhibit GC.
6) Restore flowtable with ieee80211 which broke due to a relatively
recent commit, which was pulled in by -stable, causing a regression
in 6.18 kernels.
And the following IPVS fixes:
1) Fix accounting of cache entries in IPVS LBLC for destinations,
which eventually fills up the table and trigger recurrent
resizing, from Julian Anastasov.
2) Limit IPVS cache growth for LBLCR and LBLC schedulers,
from Zhiling Zou.
3) Restrict IP_VS_CONN_F_ONE_PACKET for normal connections,
do not allow to use it with templates. Also from Julian.
4) Sanitize flags in IPVS sync messages received in the backup.
From Julian Anastasov.
netfilter pull request 26-09-30
* tag 'nf-26-09-30' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf:
netfilter: flowtable: restore ieee80211 forward path
netfilter: flowtable: generalize pending status bit
netfilter: bpf: reject invalid NAT manipulation types
netfilter: nft_set_rbtree: skip transaction elements during GC
ipvs: filter some flags received in the backup server
ipvs: do not create invisible templates
ipvs: bound LBLCR and LBLC cache growth
ipvs: fix missing counter decrement in lblc
netfilter: nft_flow_offload: drop flowtable reference on init error path
netfilter: ipset: do not update comments from kernel-side adds
====================
Link: https://patch.msgid.link/20260930074142.298353-1-pablo@netfilter.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
Before commit 871df5007eda ("netfilter: flowtable: bail out if forward
path cannot be discovered"), there was a fallback to set up a forward
path in case .ndo_fill_forward_path fails or DEV_PATH_MTK_WDMA was used.
Such fallback was used by commit d787a3e38f01 ("mac80211: add support
for .ndo_fill_forward_path").
One possibility is to handle DEV_PATH_MTK_WDMA from the flowtable
forward path discovery. However, this is only used internally by drivers
to retrieve mtk_wdma information to set up hardware offload. Felix
decided to use the .fill_forward_path interface for this purpose due to
the lack of a better interface at that time.
Add a new DEV_PATH_IEEE80211 path which is offered if the new ieee80211
flag is set on in the struct net_device_path_ctx to restore the
flowtable with a ieee80211 netdevice. Handle this new DEV_PATH_IEEE80211
path just like DEV_PATH_ETHERNET and DEV_PATH_DSA, ie. this is the last
netdevice in the stack.
This new ieee80211 flag is implicitly unset for mtk_ppe and airoha which
call dev_fill_forward_path() to retrieve a DEV_PATH_MTK_WDMA path.
Fixes: 871df5007eda ("netfilter: flowtable: bail out if forward path cannot be discovered")
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
skbedit can set skb->queue_mapping and raise the per-CPU skip_txqueue
flag so __dev_queue_xmit() honours the mapping. __dev_queue_xmit()
cleared the flag before sch_handle_egress() and only read it afterwards,
so the flag was not confined to the xmit that set it: a nested xmit
(mirred redirect or mirror, or a drop after skbedit) could set the flag
and the outer xmit would consume it for an skb that never went through
skbedit.
A forwarded packet still carries the ingress NIC's rx_queue + 1 in
skb->queue_mapping, so the outer device then indexes its tx queue state
with that stale value. Taprio's child array q->qdiscs[] is sized to the
device's queue count, so taprio_enqueue() indexes past its allocation
and dereferences the result as a struct Qdisc *.
We (ab)use the skb->nf_skip_egress which means "skip netfilter egress
for this packet" to tag to "am I in tc egress?". Despite the overload
I dont see it as a conflict since the marker is set only around the
single sch_handle_egress() call and ingress path is guarded by
tc_at_ingress.
I will send a followup(net-next) patch once this hits net-next to
rename the skb->nf_skip_egress bit/flag to skb->skip_egress
Arm the flag only from the egress classifier that can use it: raise
skip_txqueue from tcf_skbedit_act() only when it runs inside
sch_handle_egress(), thanks to skb->nf_skip_egress. An egress qdisc
classifier runs in q->enqueue(), after the tx queue has been picked,
so a mapping it sets cannot affect the current packet; arming the flag
there only pollutes it for a later xmit. Then own the flag for the xmit
frame the egress hook runs in: save the incoming value and clear it just
before sch_handle_egress(), and restore it after the hook - on the
consumed (drop) path, or, in the same call that reads it, on the
surviving path. The save and the restores stay inside the
egress_needed_key static branch, so a packet pays for them only when
egress hooks are active (2f1e85b1aee4).
Store the value netdev_cap_txqueue() selected back into skb->queue_mapping
in netdev_tx_queue_mapping(), as netdev_core_pick_tx() already does, so
the skip_txqueue path never hands a later reader on the xmit path a
mapping the device cannot serve. A store made still later in the same
frame, by a tc BPF program attached to a transmit qdisc, is outside this
path and is not re-capped; a separate followup will resolve that path.
netdev_xmit_skip_txqueue() returns the previous flag value so the
save-and-clear is one call, and a no-op stub is provided when
CONFIG_NET_EGRESS is disabled. skb->nf_skip_egress is compiled under
CONFIG_NET_EGRESS rather than CONFIG_NETFILTER_SKIP_EGRESS, so
skb_at_tc_egress() is valid whenever the egress path is built.
A local user in a network namespace can redirect a packet from a device
with more TX queues to one with fewer after setting a mapping valid only
on the larger device. That reaches these reads and, under KASAN, faults
with "slab-out-of-bounds in taprio_enqueue".
Conditions to recreate the bug: the report's own trigger is a local user
with CAP_NET_ADMIN in a network namespace, so no eBPF program is needed.
With CONFIG_NET_SCH_TAPRIO=y, CONFIG_NET_ACT_SKBEDIT=y,
CONFIG_NET_ACT_MIRRED=y, CONFIG_NET_CLS_MATCHALL=y,
CONFIG_NET_SCH_PRIO=y and KASAN enabled, create qa (3 queues), qb
(2 queues) and qc (1 queue) as dummy devices; put a taprio root on qb
(num_tc 1, queues 2@0) and clsact on all three; then add an egress
matchall filter on every device. On qa: "action skbedit queue_mapping 2
pipe action mirred egress redirect dev qb". On qb: "action mirred egress
mirror dev qc". On qc: "action skbedit queue_mapping 0 pipe". Send one
packet out qa. qc's skbedit sets the flag while qb's outer xmit is in
flight; without the fix qb consumes it and reads its two-entry taprio
child array with the forwarded packet's stale mapping. A qc whose
skbedit is instead installed in a transmit-qdisc classifier (a matchall
filter on the qc root qdisc) reaches the same code path the same way
without the fix.
Testing: on a KASAN build with panic_on_warn=1 the unfixed kernel panics
with "BUG: KASAN: slab-out-of-bounds in taprio_enqueue", a read 0 bytes
past a 16-byte taprio_init() allocation, for the clsact-setter and the
transmit-qdisc-classifier reproducers and for a clsact skbedit-then-tc-BPF
store; the fixed kernel runs all three with no report, and the BPF store
variant additionally shows the expected "selects TX queue" clamp notice
from the write-back.
Fixes: 2f1e85b1aee4 ("net: sched: use queue_mapping to pick tx queue")
Reported-by: Zero Day Initiative <zdi-disclosures@trendmicro.com>
Link: https://lore.kernel.org/netdev/CANn89iLwYx8nCVf0pCEk_MmEiyC6kQaMwCQT9WkQVeeNzNQHqQ@mail.gmail.com/
Link: https://lore.kernel.org/netdev/179008581937.2160803.7117814290574262942@kernel.org/
Link: https://lore.kernel.org/netdev/179033713973.2160803.4914570693994398206@kernel.org/
Link: https://lore.kernel.org/netdev/20260925180407.63647514@kernel.org/
Link: https://lore.kernel.org/netdev/CANn89i+k-mZKDQVtvws_MEXeuMTAdaCcOXFZE-RfhcGTu90sjA@mail.gmail.com/
Suggested-by: Eric Dumazet <edumazet@google.com>
Suggested-by: Jakub Kicinski <kuba@kernel.org>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/QDISC-9R8V.v4.20260928081529@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/mtd/linux
Pull MTD fixes from Miquel Raynal:
"The most important set of fixes are around the handling of the QE bit
in SPI NAND.
There are also a couple of behavioral fixes (mutex issue in SPI-NOR,
spurious bitflips on vf610_nfc, OOB bytes count in SPI NAND and
cfi_cmdset stack usage).
The rest is mostly AI fuzzing results"
* tag 'mtd/fixes-for-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/mtd/linux:
mtd: spinand: Do not update the QE bit on devices without one
mtd: spi-nor: core: Fix mutex leak in spi_nor_rww_start_exclusive()
mtd: rawnand: cadence: Initialize IRQ state before requesting IRQ
mtd: rawnand: vf610_nfc: fix false bitflips on reads of erased pages
mtd: rawnand: vf610_nfc: fix reads on chips with more than 64 bytes of OOB
mtd: spinand: fix zero oobavail when no ECC engine is used
mtd: spinand: fix NULL pointer dereference with no ECC engine
mtd: mtd_intel_dg: reset poll counter for each erase
mtd: cfi_cmdset_0001: shrink do_write_buffer() stack frame
mtd: core: call _get_device() with the master MTD
mtd: core: avoid double-free of OTP NVMEM device
mtd: spinand: Enable QE on all dies
mtd: block2mtd: Fix divide error when erase_size is zero
|
|
After commit d2d5c129d07e ("cpufreq: Make cpufreq_update_pressure() fall
back to cpuinfo.max_freq"), cpufreq pressure appears in the CPU load
balancer unexpectedly in some cases in which it was not present before,
leading to confusion and uncertainty.
Clearly, the scheduler assumes that cpufreq pressure will not be set
unless the capacity reference frequency of the CPU is known, and the
commit mentioned above violates that assumption.
However, in some cases the capacity reference frequency of the CPU is
in fact known even though arch_scale_freq_ref() returns 0 and in those
cases it should be possible to set cpufreq pressure as appropriate.
For this purpose, introduce a new cpufreq driver callback returning
the CPU capacity reference frequency, .scale_freq_ref(), and make
cpufreq_update_pressure() invoke it, if present, instead of falling
back to cpuinfo.max_freq unconditionally.
Add that callback to the intel_pstate driver and make it return 0 unless
the scale-invariant capacity of the given CPU has been explicitly set,
in which cases its reference frequency is always cpuinfo.max_freq.
Fixes: d2d5c129d07e ("cpufreq: Make cpufreq_update_pressure() fall back to cpuinfo.max_freq")
Reported-by: Jianyong Wu <wujianyong@hygon.cn>
Closes: https://lore.kernel.org/linux-pm/20260915065747.1671965-1-wujianyong@hygon.cn/
Tested-by: Jianyong Wu <wujianyong@hygon.cn>
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
Tested-by: Ricardo Neri <ricardo.neri-calderon@linux.intel.com> # Intel hybrid parts
Tested-by: Chen Yu <yu.c.chen@intel.com>
[ rjw: Add READ_ONCE() around a capacity_perf read ]
Link: https://patch.msgid.link/12975163.O9o76ZdvQC@rafael.j.wysocki
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull scheduler fixes from Ingo Molnar:
- Fix LLC mis-scheduling bugs (Tim Chen, Lu Wang)
- Fix cache-grouping related scheduling statistics UAF bugs (Tim Chen)
- Skip kernel threads for cache aware scheduling to rubustify the code
(Chen Yu)
- Refresh LLC capacity across CPU hotplug, to fix capacity
underestimation bug (Davi Chaves Azevedo)
- Account PSI IRQ time to the execution context, not the scheduling
context, to fix proxy scheduling accounting bug (Zhan Xusheng)
* tag 'sched-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
sched/core: Account PSI IRQ time to the execution context, not the scheduling context
sched/cache: Refresh LLC capacity across CPU hotplug, to fix capacity underestimation bug
sched/cache: Skip kernel threads for cache aware scheduling to rubustify the code
sched/cache: Introduce task_struct->sched_cache_grp to fix UAF
sched/cache: Decouple sched_cache_group from mm to fix UAF
sched/cache: Honor migrate_llc_task semantics in active load balance, to fix LLC mis-scheduling bug
sched/cache: Keep nr_pref_llc_running in the runnable domain, to fix LLC mis-scheduling bug
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull perf events fixes from Ingo Molnar:
- Fixes for KVM guest PEBS virtualization (Sean Christopherson)
- Fixes for various Intel PMUs related to PEBS data-source (Dapeng Mi)
- Fix Intel Panther Cove event scheduling constraints (Dapeng Mi)
- Fix Intel DMR/NVL OMR extra registers event scheduling (Dapeng Mi)
- Rename two confusingly named PMU attributes (Dapeng Mi)
- Fix a refcount leak in attach_perf_ctx_data() (Namhyung Kim)
- Fix NULL pointer dereference crash in __perf_pmu_sched_task()
(Puranjay Mohan)
- Fix CPU-wide event scheduling (Puranjay Mohan)
- Fix x86 LBR branch entry generation (Puranjay Mohan)
* tag 'perf-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
perf/core: Fill branch entries with a single assignment
perf/core: Run sched_task() for PMUs with only CPU-wide events
perf/core: Fix NULL pmu_ctx passed to pmu->sched_task()
perf/core: Fix a refcount leak in attach_perf_ctx_data()
perf/x86/intel: Rename NVL offcore_rsp attribute to offmodule_rsp
perf/x86/intel: Rename DMR offcore_rsp attribute to offmodule_rsp
perf/x86/intel: Fix precise OMR event scheduling for DMR/NVL
perf/x86/intel: Constrain Panther Cove UOPS_DISPATCHED events to PMCs 0-3
perf/x86/intel: Delete dead NVL PEBS data-source initcall
perf/x86/intel: Fix Panther Cove PEBS data-source snoop states
perf/x86/intel: Remove incorrect Panther Cove PEBS data-source constraints
perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints
perf/x86/intel: Update arw_latency_data() mem-op direction handling
perf/x86/intel: Fix DKT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Fix CMT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs()
perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) if PEBS is unused
perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has PEBS isolation, to fix stuck PEBS_ENABLED
perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted PERF_GLOBAL_CTRL bits
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux
Pull ata fixes from Niklas Cassel:
- Extend the quirk "no LPM on ATI" quirk, that is currently only
applied for Samsung drives, to include AMD controllers as well.
The AMD AHCI controllers are newer versions of the ATI AHCI
controllers, and these controllers still have LPM issues with
Samsung drives - LPM works with drives from other vendors (me)
- Fix errors in the libata.force parameter documentation (me)
- Verify the sense data descriptor lengths for ATA PASS-THROUGH
command, so that a malicious device cannot write past the buffer
length (Matthias)
- Mention the libata for-next branch in MAINTAINERS such that the
git ls-remote command done by get_maintainer.pl --self-test=scm
can verify it (Matthias)
* tag 'ata-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux:
MAINTAINERS: name the libata/linux for-next branch
ata: libata-scsi: bound the ATA passthru sense descriptor writes
ata: libata: Correct libata.force parameter documentation
ata: libata-core: Extend Samsung LPM quirk to AMD controllers
|
|
Pull kvm fixes from Paolo Bonzini:
"Arm:
- Invalidate the ITS translation cache when the guest changes the
base address of the ITS tables (Fuad Tabba)
- Skip saving ITS devices with device IDs that are out-of-bounds
rather than failing the entire ITS save ioctl (Fuad Tabba)
- Close race between VM teardown and invalidations of nested MMUs
when handling MMU operations that are allowed to block (Lorenzo
Stoakes)
- Various fixes for the handling of the host's untrusted SVE
configuration in pKVM (Fuad Tabba)
- Make sure that empty SMCCC ranges based at 0 are rejected by the
kvm_smccc_set_filter() (Karl Mehltretter)
- Revoke the host mapping for pKVM's private stack pages, along with
a new sanity check that all mappings in the hyp's private VA range
have been correctly marked as hyp-owned (Fuad Tabba)
- Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that
concurrent vCPU initialization cannot relocate in-use MMUs. Defer
the freeing of shadow stage-2 MMUs to the point that no other users
(e.g. MMU notifier) could reference them (Marc Zyngier)
- Drop useless WARN when rejecting an unsupported ioctl for pKVM
(Fuad Tabba)
- Fix the steal_time selftest to install correctly-sized mappings for
non-4K hosts (Sebastian Ott)
- Correct mapping of fine-grained trap for GCSPOPX instruction (Mark
Brown)
- Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit
guests (Karl Mehltretter)
RISC-V:
- Synchronize hrtimer during VCPU teardown
- Fix the conversion between vsip and hvip values
- Serialize IMSIC attributes with vCPU migration
- Release unused page after MMU invalidation
- Propagate interrupted G-stage faults to KVM user-space as EINTR
- Fix nested acceleration hfence entry update order
- Fix sdata leak and stale snapshot_addr in snapshot_set_shmem
- Preserve firmware counter value across PMU counter stop/start
- Report PMU snapshot write failure to the guest
- Fix perf-backed counter accounting across PMU stop and read
- Correctly propagate error of a hart status SBI call
s390:
- Ensure that accesses through kvm_arch_set_irq_inatomic mark as
dirty the pages that contain indicator and summary bits
- Fix compile warning for kvm_s390_update_cmma_dirty()
- Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to
userspace
- Move s390_kvm_mmu_commit_memory_region() into
s390_kvm_mmu_prepare_memory_region() so that it can fail instead of
WARN
- Add missing srcu in kvm_s390_set_irq_state()
- Fix potential races in storage functions
- Fix race in _destroy_pages_crste()
- Fix issues in the handling of KVM interrupt and page resources,
when a queue that is assigned to a mediated device (mdev) is
removed from the host's AP configuration
- Fix loop condition in uv_find_secrets
- Prevent potential out-of-bounds read
x86:
- Fix a brown paper bag bug where KVM would incorrectly treat Intel
PMU MSRs as valid on AMD
- Fix a regression in the hardware disable selftest where it checked
the wrong macro when detecting glibc support (breaks at least musl)
- Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is
especially important for KVM_BUG_ON() flows, which often guard more
dangerous bugs
- Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a
bug where KVM would let userspace run a broken setup with stale
vmcs12 pages
- Fix a class of bugs where KVM would fail to fill kvm_run exit
fields if getting nested pages failed
- Treat reserved entries in the memory attributes xarray as "no
attributes", to fix false positives when checking for mixed
attributes
- Fix memcg accounting for the memory attributes xarray (the xarray
library subtly requires the xarray to be configured for accounting
upfront; the gfp flags taken at runtime are used only rarely)
- Don't pre-reserve xarray entries when storing empty attributes, as
storing NULL must not require memory allocation (KVM and other
subsystems heavily rely on this behavior)
- Fix a memory leak and a cache maintenance issue related to doing
intra-host migration on an SEV guest"
* tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (54 commits)
KVM: SEV: Do cache maintenance on the source VM during intra-host migration
KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV
KVM: Don't pre-reserve xarray entries when storing empty/NULL attributes
KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg
KVM: Don't treat reserved xarray entries as having memory attributes
KVM: x86: Fill kvm_run exit fields in common get_nested_state_pages() error paths
KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails
KVM: arm64: Fix AArch32 DBGBXVR<n> handling
KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP
KVM: selftests: fix steal_time for arm64 with host page size > 4K
KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode
KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction
KVM: arm64: nv: Fix life cycle of the nested_mmus array
KVM: arm64: Check every private mapping is hyp-owned at pKVM init
KVM: arm64: Move the private VA allocation cursor to __io_map_next
KVM: arm64: Match hyp text by physical address in fix_host_ownership()
KVM: arm64: Transfer the hyp stack pages out of the host stage-2
KVM: arm64: selftests: Test empty SMCCC filter range at base 0
KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0
KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2
...
|
|
Now that KVM uses kvm_get_vcpu_by_id() to check for an existing vCPU ID
before doing any meaningful work, which was made possible by holding
kvm->lock for the entirety of vCPU creation, revert the now-redundant
"early" vCPU ID tracking. The claims about the impact of kvm->vcpu_ids on
the memory footprint were a wee bit wrong: the worst case scenario isn't
256 bytes per VM, it's 256 "unsigned longs" per VM, i.e. 2048 bytes per VM.
Increasing the size of "struct kvm" by 2048 nearly doubled the total size
on many architectures, and tripped x86's KVM_SANITY_CHECK_VM_STRUCT_SIZE,
which was added to detect this *exact* scenario, where a single change
significantly increased the size of "struct kvm". I.e. attempting to build
KVM with CONFIG_DEBUG_KERNEL=n fails on x86 (the build failures got missed
because all build bots apparently test only CONFIG_DEBUG_KERNEL=y kernels,
and maintainers' test flows were similarly lacking).
This reverts commit 97d65b544f48b2ee49f6aea32145e3e7969955dc.
Fixes: 97d65b544f48 ("KVM: Check for duplicate vcpu_id as early as possible")
Reported-by: Jean-Christophe Guillain <jean-christophe@guillain.net>
Closes: https://lore.kernel.org/all/56a4bc35ee605588b7cc36c8e45c12b5f3b506cb.camel@guillain.net
Reported-by: Paweł S <spawel523@gmail.com>
Closes: https://lore.kernel.org/all/CABD%3DWFOS4j4hDv%2BpW-eEM9HAM2q2GY_iYdAG%2BqvYcUEinUrcQQ@mail.gmail.com
Tested-by: Jean-Christophe Guillain <jean-christophe@guillain.net>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Tested-by: Naveen N Rao (AMD) <naveen@kernel.org>
Message-ID: <20260921174445.911676-7-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
KVM fixes for 7.3-rcN
- Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs
as valid on AMD.
- Fix a regression in the hardware disable selftest where it checked the wrong
macro when detecting glibc support (breaks at least musl).
- Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially
important for KVM_BUG_ON() flows, which often guard more dangerous bugs.
- Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug
where KVM would let userspace run a broken setup with stale vmcs12 pages.
- Fix a class of bugs where KVM would fail to fill kvm_run exit fields if
getting nested pages failed.
- Treat reserved entries in the memory attributes xarray as "no attributes",
to fix false positives when checking for mixed attributes.
- Fix memcg accounting for the memory attributes xarray (the xarray library
subtly requires the xarray to be configured for accounting upfront; the gfp
flags taken at runtime are used only rarely).
- Don't pre-reserve xarray entries when storing empty attributes, as storing
NULL must not require memory allocation (KVM and other subsystems heavily
rely on this behavior).
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs
Pull vfs fixes from Christian Brauner:
- Revert "put_mnt_ns(): leave mounts connected". This allows the
creation of reference count cycles in a very trivial way. We can't
bring this in until we have fixed the underlying cause
- vfs: Don't create the private nullfs instance for kthreads under
namespace_sem to avoid false lockdeps complaints
- binfmt_misc:
- Copy the name into a stack buffer and look up the copy in
bpf_binprm_select_interp()
- bpf_binprm_set_interp() and bpf_binprm_set_interp_arg(): Check
the private copy instead so the string that gets staged is the
kstring that was checked
- netfs:
- Make netfs_read_gaps() use separate sink folios rather than one
reused sink folio to discard unwanted data so that cifs checksum
checking sees all the data that was fetched
- Trim reads down to i_size so afs symlinks read correctly from the
cache
- Wrap the direct mempool ->alloc() calls the GFP_KERNEL paths make
in alloc_hooks() via a new mempool_alloc_noreserve() helper
- iov_iter: Use iov_iter_alignment() for the start and length check
added to iov_iter_extract_bvecs() this cycle. It used iter_iov_addr()
and iter_iov_len() which are only valid for ITER_UBUF and ITER_IOVEC
iterators
- super: Make iterate_supers_type() deletion-safe
- inode: Stop evict_inodes() from rescanning the same inodes
- writeback: Bound the cleanup_offline_cgwb() rescans
- ntfs3: Use d_instantiate_new() in ntfs_create_inode()
- ovl: Fix a use-after-free in the ovl_do_mkdir() debug print
- dcache: Unpoison the inline name buffer in __d_alloc() for KMSAN
- autofs: Fix a pipe file reference leak in autofs_kill_sb()
- bpf: Drop the path_unlink and path_rmdir hooks from the list of hooks
for which the verifier rewrites bpf_{set,remove}_dentry_xattr() to
the _locked variants
- squashfs: Range check the xz dictionary size before shifting by it
- selftests: Add the missing eventfd, open_tree_ns, openat2 and xattr
filesystems selftests to TARGETS and drop the stale openat2 entry
left behind when those tests moved
* tag 'vfs-7.3-rc5.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs:
netfs: Fix missing alloc tagging of direct mempool allocations
bpf: fs/xattr: don't assume the inode is locked in path_unlink/path_rmdir
autofs: fix sbi->pipe file reference leak in autofs_kill_sb()
dcache: unpoison the inline name buffer in __d_alloc()
ovl: fix UAF in ovl_do_mkdir() debug print
super: make iterate_supers_type() deletion-safe
Revert "put_mnt_ns(): leave mounts connected"
Revert "selftests/filesystems: add mntns cleanup test"
binfmt_misc: fix racy checks in bpf set_interp kfuncs
binfmt_misc: fix OOB read in bpf_binprm_select_interp()
fs: don't create the private nullfs mount under namespace_sem
writeback: bound cleanup_offline_cgwb() rescans by rotating scanned inodes
fs: avoid repeated scans in evict_inodes()
netfs, afs: Fix symlink reading
netfs: Fix netfs_read_gaps() to use separate sink folios
squashfs: Add dictionary size range check to prevent shift-out-of-bounds
fs/ntfs3: use d_instantiate_new() in ntfs_create_inode() and murder syzbot's "WARNING in do_new_mount" saga
selftests/filesystems: fix missing and stale TARGETS entries
block: Fix start and length check added to iov_iter_extract_bvecs()
|
|
Commit 1d78d56c43ef ("netfs: Fix folio_queue ENOMEM in writeback by
adding a mempool") added a mempool for the folio_queues and made the
request, subrequest and folio_queue allocations distinguish between
writeback and everything else. Writeback is part of memory reclaim
and must not fail due to ENOMEM, so it allocates under GFP_NOFS
through mempool_alloc(), which may dip into the pool's reserve and,
if that runs empty, wait for elements to be returned. The
GFP_KERNEL paths, which can return -ENOMEM to their callers, invoke
the pool's ->alloc() callback directly instead.
The direct call, however, skips the alloc_hooks() wrapper that the
mempool_alloc() macro provides. The pool callbacks, mempool_alloc_slab()
and mempool_kmalloc(), call kmem_cache_alloc_noprof() and kmalloc_noprof()
and rely on current->alloc_tag having been set by the caller. With
CONFIG_MEM_ALLOC_PROFILING_DEBUG=y this leads to
current->alloc_tag not set
WARNING: ./include/linux/alloc_tag.h:161 at __alloc_tagging_slab_alloc_hook
alloc_tag was not set
WARNING: ./include/linux/alloc_tag.h:166 at __alloc_tagging_slab_free_hook
at allocation and free time respectively, as reported when reading
files on a CIFS mount. The allocations are also missing from
/proc/allocinfo.
Wrap the direct ->alloc() invocations in alloc_hooks() with a new
mempool_alloc_noreserve() helper in include/linux/mempool.h, next to
the other alloc_hooks()-wrapped macros such as mempool_alloc(). The
GFP_KERNEL paths keep their failable allocation semantics, they just
get tagged now.
Fixes: 1d78d56c43ef ("netfs: Fix folio_queue ENOMEM in writeback by adding a mempool")
Reported-by: Erhard Furtner <erhard_f@mailbox.org>
Closes: https://lore.kernel.org/all/0b004319-9ef7-437c-a4dd-174d6a9a83db@mailbox.org/
Tested-by: Erhard Furtner <erhard_f@mailbox.org>
Suggested-by: Suren Baghdasaryan <surenb@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Hao Ge <hao.ge@linux.dev>
Link: https://patch.msgid.link/20260923063759.34667-1-hao.ge@linux.dev
Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Jakub Kicinski:
"Including fixes from Bluetooth, NFC and Netfilter.
Every week in this release is record-setting for number of posted
patches. It doesn't seem like we're creating any regressions with all
these fixes, three 'Fixes' tags here point to 7.2 commits but none are
true regression fixes. We're trying to keep the count down,
nonetheless.
Previous releases - regressions:
- net: don't require the hwtstamp NDOs when a PHY provides
timestamping
- ipv6: fix dst leak for uncached routes
- vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE
packets
Previous releases - always broken:
- packet: use ubuf_info completion for TX_RING packets
- arp: terminate device name before lookup
- ipv6: do not let ipv6_find_hdr() return an offset past the packet
end
- udp: remove a disconnected socket from the 4-tuple hash table
- sctp: discard the rest of the packet on a stale-cookie error
- eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged
eswitch"
[ And lots of other random network driver fixes ]
* tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (189 commits)
tcp: prevent collapsing skbs across boundary in rtx queue
vlan: ensure sufficient headroom in vlan_dev_hard_header()
net/sched: sch_teql: fix shadowed err in __teql_resolve()
bridge: check llc_mac_hdr_init() return value in br_send_bpdu()
llc: fix skb UAF and leaks on llc_mac_hdr_init() failure
llc: reserve device headroom for allocated frames
gve: DQO: reject TSO packets with an out of range MSS
gve: fix TX drop when GSO MSS is too small for hw
gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO
net: flush skb_defer_nodes in dev_cpu_dead()
net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails
af_packet: fix integer overflow in prb_calc_retire_blk_tmo()
tipc: Fix a data race on mon->peer_cnt in mon_timeout()
net: phy: intel-xway: workaround 100BASE-TX Link-Up issue
net/smc: fix UAF on lgr list traversal in smcr_port_err()
net/rds: size a connection's path set by the transport it ends up with
nfp: hold IPsec RX state under the XArray lock
net: ena: fix MMIO read buffer leak on probe failure
net: ena: fix PHC cleanup on probe failure
net/sched: act_ct: fix helper UAF due to extensions realloc
...
|
|
Pull bpf fixes from Alexei Starovoitov:
- Fix bpf_skb_change_tail() to drop the checksum offload instead of
rejecting the trim of CHECKSUM_PARTIAL skbs (Daniel Borkmann)
- Add KF_PERFMON kfunc flag and require CAP_PERFMON for kfuncs that
read arbitrary memory and for untrusted read-only memory reads
(Daniel Borkmann)
- Clear scalar delta on narrowing stack spill (Daniel Borkmann)
- Set up the frame pointer for the exception callback in arm64 JIT, and
zero-fill other CPUs when BPF_F_CPU update creates a per-cpu hash
element (Donggeun Yoo)
- Various fixes (Emil Tsalapatis):
- Fix bounds check underflow for skb-backed dynptrs
- Fix rx_queue_mapping context access code generation in bpf_sock
- Reject packet pointer arguments to subprogs that may mutate the
packet
- Reject ALU instructions that see arena and non-arena operands on
different code paths
- Fix copied_seq double-counting on sockmap self-redirect
(Geliang Tang)
- Fix divide-by-zero in btf_struct_walk() on a flexible array of
zero-sized elements, fix out-of-bounds read of rtt_min in sock_ops
(Jiayuan Chen)
- Fix bpf_sock_destroy() out-of-bounds read of sk_protocol on TIME_WAIT
and request socks, and sleeping under RCU when destroying a listener
with pending children (Jiayuan Chen)
- Fix JEQ/JNE with immediate operand in MIPS32 JIT and missing zero
extension of BSWAP 16/32 in MIPS64 JIT (Johan Almbladh)
- Avoid soft lockup in htab lookup[_and_delete] batch operations on
large maps (Jose Fernandez)
- Various fixes (Kumar Kartikeya Dwivedi):
- Verify global subprogs in each sleepability context they are
called from
- Make post-verification instruction rewrites killable
- Preserve packet pointer displacement in regsafe()
- Apply CO-RE relocations before subprogram validation, restrict
CO-RE poisoning to relocatable instructions, and reject truncated
ldimm64 CO-RE relocations in libbpf
- Assign lock identity to callback map values
- Compare stack frames in regs_exact()
- Bound ownership depth through local kptrs and graph roots
- Fix u32 overflow in map batch operations when the map size exceeds
4GB (Masoud Aghasi)
- Fix UAF in bpf memalloc due to concurrent consumption of ttrace lists
in alloc_bulk() (Pu Lehui)
- Allow gotox as the terminal instruction of a program or a subprogram
(Siddharth Chintamaneni)
- Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL, skip unsettled links
in link iterator, and reject dev-bound-only programs on other devices
(Weiming Shi)
- Reject non-negative stack offsets in stack_slot_obj_get_spi()
(Xu Yunxiang)
- Check params size before reading reserved fields in
bpf_crypto_ctx_create() (Yuqi Xu)
- Reject max_entries > INT_MAX in sock_map_alloc() (Zhao Gongyi)
- Use a 32-bit compare in xsk_map_gen_lookup() (Zhiling Zou)
- Use kvfree() in xdp_test_run_teardown() (Zhixing Chen)
* tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: (58 commits)
selftests/bpf: Test per-cpu initialization of a BPF_F_CPU created element
bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element
bpf: Fix BSWAP 32 and 16 on MIPS64
bpf: Fix immediate JMP JEQ/JNE on MIPS32
bpf: Reject dev-bound-only programs on other devices
bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc
selftests/bpf: Test for mixed arena/nonarena code paths
bpf: Prevent variable arena/non-arena register contents
selftests/bpf: Test rejection of pkt args to mutating subprogs
bpf: Reject pkt arguments in mutating subprogs
selftests/bpf: Add selftests for rx_queue_mapping context access
bpf: Fix bpf_sock context code generation
selftests/bpf: Test dynptr slices past end of skb
bpf: Fix bounds check for skb-backed dynptrs
selftests/bpf: Reject iterator destruction through fp+0
bpf: Reject non-negative offsets in stack_slot_obj_get_spi()
bpf: Check params size before reading reserved fields
selftests/bpf: Check local object ownership depth
bpf: Bound ownership depth through local kptrs and graph roots
selftests/bpf: Cover frame changes in bounded loops
...
|
|
perf_clear_branch_entry_bitfields() clears the bitfields of struct
perf_branch_entry one by one and leaves from/to alone, since callers
overwrite those straight away. The list has to be kept in sync with the
struct by hand and has already fallen behind: new_type and priv were
added to perf_branch_entry and never added here.
Only BRBE writes those two, and neither for every record.
brbe_set_perf_entry_type() leaves new_type alone for a branch type it
does not recognise, and priv is not set for source-only records.
arm_pmuv3.c allocates the per-CPU branch stack with kmalloc(), so such a
record reaches userspace with whatever the slot held: uninitialised
kmalloc() data on the first pass over the buffer, the previous record's
values after that. Nothing under arch/x86/events/ writes either field,
so only arm64 is affected.
Assign the whole entry at each site instead. Everything not named is
then zero, and there is no list to keep in sync. The bitfields add up to
exactly 64 bits, so the struct has no padding to leave undefined.
perf_clear_branch_entry_bitfields() has no callers left, so remove it.
perf_entry_from_brbe_regset() assigns an empty literal instead, since it
fills from/to conditionally. PERF_BR_SPEC_NA is 0, so dropping the
explicit spec assignment changes nothing.
Fixes: b190bc4ac9e6 ("perf: Extend branch type classification")
Fixes: 5402d25aa571 ("perf: Capture branch privilege information")
Suggested-by: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Tested-by: Yifan Wu <wuyifan50@huawei.com>
Link: https://patch.msgid.link/20260810133540.1947118-4-puranjay@kernel.org
|
|
tpacket_snd sends skbs with frags pointing into its ring slots. Slots
are released when skb->destructor is called.
A call to skb_orphan calls skb->destructor before the skb is freed.
This can cause the slot to be reused while still linked into the skb.
Switch to standard zerocopy completion (ubuf_info) so the slot is only
released once all references to the payload are freed or copied.
Restore skb->destructor to standard sock_wfree.
The ubuf_info completion callback can be called with a NULL skb, but
only from net_zcopy_put and related API, used by zerocopy implementations
that hold their own reference on the uarg, such as MSG_ZEROCOPY. This
uarg is only ever completed from skb_zcopy_clear, so skb is always set.
To prevent userspace from aliasing in-flight state on shared ring
slots, allocate tpacket_uarg per packet, rather than per slot. This
adds a small allocation to the transmit path. Use standard kmalloc to
allow backporting to stable kernels.
The uarg holds an sk_wmem_alloc reference, rather than an sk_refcnt
reference. packet_free_tx_ring waits on sk_wmem_alloc before freeing
the ring pages. Always allocate vec->deferred for tx_ring so page-backed
rings also wait on sk_wmem_alloc when skb_copy_ubufs drops page refs
before calling tpacket_ubuf_complete.
Drop the tx_ring.pg_vec test that tpacket_destruct_skb performed before
accessing the slot. The sk_wmem_alloc reference now guarantees that the
slot is valid. The test is also not sufficient by itself, as it reads
pg_vec without pg_vec_lock, so it can race with packet_set_ring.
As a result a slot is released when its payload is copied, which can
be before transmission (e.g., in skb_orphan_frags_rx). If copied
before skb_tx_timestamp() is called, no slot timestamp is recorded,
similar to when skb_orphan() was called early in the datapath before
this patch.
Revert the now unused previous skb_zcopy_.._nouarg infra.
Depends on commit 992cc9f94ca9 ("net/packet: defer vmalloc TX_RING
free until skbs finish").
Reported-by: Katherine Leaver <kleaver@janestreet.com>
Reported-by: Bjoern Doebel <doebel@amazon.de>
Closes: https://lore.kernel.org/netdev/20260909085542.3370986-1-doebel@amazon.de/
Fixes: 5cd8d46ea156 ("packet: copy user buffers before orphan or clone")
Cc: stable@vger.kernel.org
Signed-off-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260919004748.1463985-3-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The verifier marks ALU instructions that include at least
one arena operand with needs_zext: These instructions are
fixed up after verification to be ALU32 instructions to
ensure that the result is a valid offset into an arena.
However, different code paths may provide two non-arena
64-bit arguments to the same instruction. The result of
the operation in that code path is wrong, since it is
now unexpectedly truncated to 32 bits and zero-extended.
Add logic to the verifier to ensure every instruction either
always has at least one PTR_TO_ARENA argument, or never does.
Since needs_zext already tracks the first scenario, add a
prevent_zext field in bpf_insn_aux to track the latter.
Reject instructions that use arena arguments and have prevent_zext
set, or do not have arena arguments and have needs_zext set.
Fixes: 6082b6c328b5 ("bpf: Recognize addr_space_cast instruction in the verifier.")
Reported-by: Nicholas Carlini <nicholas@carlini.com>
Suggested-by: Nicholas Carlini <nicholas@carlini.com>
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260922172028.6269-8-emil@etsalapatis.com
|
|
The skb_pointer_if_linear() function checks whether a
memory region of length len starting at offset off into
the skb is in the linear area, and returns a pointer to
the region if so. The check currently subtracts between
skb_headlen and offset of the check, and since skb_headlen
is unsigned the subtraction can underflow. This causes the
bounds check to spuriously pass and generate an arbitrary
pointer of the form *(skb->data + off).
The only user of this helper is currently skb-backed BPF
dynptr code. Returning the wrong pointer leads to the
dynptr erroneously being backed with invalid memory.
Ensure the subtraction cannot underflow, and fail the check if
it would. Use u64 arithmetic to also prevent overflow when
calculating (skb_headlen(skb) - off) since off is unsigned.
Fixes: 6f5a630d7c57 ("bpf, net: Introduce skb_pointer_if_linear().")
Reported-by: Nicholas Carlini <nicholas@carlini.com>
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Link: https://patch.msgid.link/20260922172028.6269-2-emil@etsalapatis.com
|
|
underestimation bug
The scheduler scales LLC capacity by the fraction of cache-sharing CPUs
covered by a domain:
llc_bytes = cache_size * span_weight / shared_weight
During CPU teardown, sched_cpu_deactivate() rebuilds scheduler domains
before cacheinfo_cpu_pre_down() removes the CPU from shared_cpu_map. The
new domains therefore use the old sharing weight. The later call to
sched_update_llc_bytes() looks up the departing CPU's sd_llc, which has
already been detached, and returns without correcting the surviving CPUs.
On a Ryzen 5 7535U with twelve logical CPUs sharing a 16 MiB LLC,
offlining one SMT sibling left the remaining CPUs with:
llc_bytes = floor(16777216 * 11 / 12) = 15379114 bytes
The correct capacity is still 16777216 bytes. On systems with active
cache-aware scheduling, an underestimated capacity can cause
exceed_llc_capacity() to reject aggregation for a process whose footprint
would fit. Unchanged cpuset partitions sharing the physical cache can
also retain stale capacity when a CPU comes online in another partition.
Pass the cache-sharing mask already retained by cacheinfo to the
scheduler update. Refresh every surviving CPU using its own LLC domain
so that each partition receives the correct share. This also preserves
the correction needed as cache-sharing maps grow during boot.
Keep the existing CPU-hotplug and scheduler-domain synchronization. The
update remains on the hotplug path; no steady-state scheduling operation
or persistent allocation is added.
Fixes: 7030513a0877 ("sched/cache: Calculate the LLC size and store it in sched_domain")
Signed-off-by: Davi Chaves Azevedo <davichazbh@gmail.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Reviewed-by: Chen Yu <yu.c.chen@intel.com>
Reviewed-by: Tim Chen <tim.c.chen@linux.intel.com>
Reviewed-by: K Prateek Nayak <kprateek.nayak@amd.com>
Tested-by: Chen Yu <yu.c.chen@intel.com>
Tested-by: K Prateek Nayak <kprateek.nayak@amd.com>
Cc: <stable@kernel.org> # v7.2.x
Link: https://patch.msgid.link/6751d93e15889e624796c74db0bfe66603d60b1b.1790035273.git.tim.c.chen@linux.intel.com
|
|
Add a sched_cache_grp pointer to task_struct so that scheduler code
can access the cache group directly via the task, without going
through mm->sched_cache_grp. This decouples the scheduler's hot-path
accesses from the mm_struct.
Each task holds its own refcount on the sched_cache_group, separate
from the reference held by its mm_struct. The reference is acquired
in copy_mm() (fork) and exec_mmap() (exec), and released in exit_mm().
This fixes the use-after-free when account_mm_sched() reaches the group
through a task whose mm is being switched, as reported by Hyunwoo:
https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Convert all scheduler code in fair.c and exit.c to use
p->sched_cache_grp instead of p->mm->sched_cache_grp.
Keep the fork/exec/exit reference management out of the generic mm
paths: add sched_cache_fork(), sched_cache_fork_cleanup(),
sched_cache_exec_mmap() and sched_cache_exit_mm() in
kernel/sched/cache_sched.c (with empty stubs for !CONFIG_SCHED_CACHE),
so fs/exec.c, kernel/fork.c and kernel/exit.c each call one helper
instead of open-coding the refcounting under #ifdef. Also add
sched_cache_group_get() and task_cache_group_get().
Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing")
Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Closes: https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@linux.dev/
Reported-by: Hyunwoo Kim <imv4bel@gmail.com>
Reported-by: Zenghui Yu (Huawei) <zenghui.yu@linux.dev>
Co-developed-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Cc: <stable@kernel.org> #7.2.x
Link: https://patch.msgid.link/ae7081dc54736bf115215f9867abb2711a7403fb.1790035273.git.tim.c.chen@linux.intel.com
|
|
Currently the sched cache grouping is by mm and the scheduling statistics
sched_cache_stat lives in the mm structure. This ties the life cycle
of scheduling stats with mm.
In account_mm_sched(), the scheduling stats are accessed by
task->mm->sc_stat. However, a task may be switching mm on one CPU when
another CPU is running account_mm_sched(), and possibly accessing the
old mm that was freed. This problem was found when running tests with
KASAN by Hyunwoo:
https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Instead of serializing the mm access by introducing extra acquisition of
rq lock in the mm free path, extract sched_cache_stat from mm_struct,
rename it as sched_cache_group and manage its life cycle apart from
mm_struct with its own ref counting. This allows us in the next patch
access sched_cache_group directly from task, and add a refcount
on sched_cache_group when a task links to it. This prevents the use
after free issue when accessing stale and released old mm and its
sched cache stat a task switches to a new mm while account_mm_sched()
is done elsewhere.
The other benefit of this restructure is in the future, the grouping of
tasks to a LLC would have the flexibility to be associated with a user
defined grouping, or cgroup, cookie group, numa_group or others instead
of just with a single mm address space.
Rename sched_cache_stat to sched_cache_group and turn it into a refcounted
object allocated from mm_struct. The mm_struct now holds a pointer
(sched_cache_grp) to this object instead of embedding it.
Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing")
Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Closes: https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@linux.dev/
Reported-by: Hyunwoo Kim <imv4bel@gmail.com>
Reported-by: Zenghui Yu (Huawei) <zenghui.yu@linux.dev>
Co-developed-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Cc: <stable@kernel.org> #7.2.x
Link: https://patch.msgid.link/91fd1e3266707c865bc9abecfb3e17bc676712df.1790035273.git.tim.c.chen@linux.intel.com
|
|
A Samsung SSD 870 QVO 8TB connected to an AMD 600 Series chipset SATA
controller is reported to time out on STANDBY IMMEDIATE during system
suspend with med_power_with_dipm enabled. The command completes when
using max_performance instead.
The existing Samsung LPM quirk only matches ATI controllers, leaving
AMD controllers unaffected. Rename it to
ATA_QUIRK_NO_LPM_ON_ATI_AND_AMD and extend the vendor check to AMD for
the same Samsung SSD model patterns. Keep LPM behavior unchanged for
other controller vendors, including Intel.
Leave ATA_QUIRK_NO_NCQ_ON_ATI restricted to ATI, since the reported AMD
issue concerns LPM rather than NCQ.
Link: https://bugzilla.kernel.org/show_bug.cgi?id=221986
Reviewed-by: Damien Le Moal <dlemoal@kernel.org>
Reviewed-by: Mario Limonciello (AMD) <superm1@kernel.org>> ---
Link: https://lore.kernel.org/r/20260918124030.1962773-5-cassel@kernel.org
Signed-off-by: Niklas Cassel <cassel@kernel.org>
|
|
__vlan_insert_inner_tag() only guarantees head room via skb_cow_head(),
never that mac_len bytes of MAC header are present. Its ETH_HLEN
wrappers - __vlan_insert_tag() under skb_vlan_push(), and
vlan_insert_tag() under validate_xmit_vlan() on the generic transmit
path - therefore rewrite the first 16 bytes at skb->data: a 12-byte
memmove plus two 2-byte stores at +12 and +14. No caller supplies the
bound, while the pop helpers use skb_ensure_writable()/pskb_may_pull().
An IFF_TUN device has hard_header_len == 0, so packet_snd() accepts a
one-byte AF_PACKET/SOCK_RAW frame. The first vlan push only sets a
hwaccel tag; the next - clsact "action vlan push" or
bpf_skb_vlan_push() - enters the helper with skb->len still 1. The
head comes from skbuff_small_head without __GFP_ZERO, so each push
drags bytes from beyond skb->tail into the frame. After three the
one-byte send leaves as 13 bytes carrying 11 bytes of uninitialised
slab:
0000: 5a b3 62 12 80 88 ff ff 00 b3 62 12 81
`------------------------------'
only 0x5a was sent; the rest is slab, here the top 56 bits of a
linear-map address
Require the MAC header the helper rewrites to be present, so such a
frame is dropped rather than transmitted.
Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Cc: stable@vger.kernel.org
Reported-by: co+0ea1ac045375cf05@bugs.sh
Signed-off-by: Xiang Mei <xmei5@asu.edu>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260915083152.705309-1-xmei5@asu.edu
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
regs_exact() compares the register state up to id, followed by the ID
mappings, but does not compare frameno. The PTR_TO_STACK case in regsafe()
checks frameno separately, which is bypassed when exact comparison is
requested. Consequently, infinite-loop detection can treat pointers to
different stack frames as the same pointer and reject a finite loop.
For example, initialize fp-8 to zero in the caller and to one in the
callee, then pass the caller's fp-8 to the callee as r1:
loop:
r0 = *(u64 *)(r1 + 0);
if r0 != 0 goto done;
r1 = r10;
r1 += -8;
goto loop;
done:
exit;
The loop terminates after reading the callee's slot on its second
iteration. At the loop header, however, the only relevant difference is
r1's frameno, so exact comparison incorrectly reports an infinite loop.
The same problem occurs when the pointer is spilled to the stack.
Move frameno into the type-specific metadata union, ahead of id, so the
existing prefix comparison in regs_exact() covers it. Ordinary stack
pointers do not use another union member. Iterator and IRQ stack-slot
states use their dedicated union views and do not need a frame lookup.
This also keeps bpf_reg_state at 80 bytes.
Since frameno now shares storage with other pointer metadata, it is only
meaningful for PTR_TO_STACK registers. Return NULL from bpf_func() for
other register types. process_iter_arg(), get_constant_map_key() and
is_dynptr_reg_valid_init() look up the frame before checking the register
type and would otherwise index frame[] with a byte of the register's map
or BTF pointer. They dereference the frame only after their type check.
Move the states_maybe_looping() boundary from frameno to precise after the
field relocation. Its prefix comparison continues to cover the complete
value state and now includes frameno.
Continue to ignore precise. Precision marks control whether pruning may
ignore scalar ranges; they do not change the represented values, and exact
comparison already compares those ranges unconditionally. Marks can also
change through backtracking while an ancestor state is still being
explored.
Fixes: d5b892fd607a ("bpf: make infinite loop detection in is_state_visited() exact")
Reported-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260919014213.1840880-2-memxor@gmail.com
|
|
An offline self test that brings the interface down and back up with
netif_close() / netif_open() requires rtnl_lock for both. Since the
ethtool IOCTL path became rtnl-optional for ops-locked drivers, the
ETHTOOL_TEST ioctl runs holding only the netdev instance lock, so on an
ops-locked driver the self test now tears the device down without
rtnl_lock.
With lockdep this reproduces deterministically on every offline self
test on such a driver; note the sole lock held is the instance lock, not
rtnl:
WARNING: suspicious RCU usage
net/core/netpoll.c:207 suspicious rcu_dereference_protected() usage!
1 lock held by ethtool/107:
#0: (&dev->lock){+.+.}, at: dev_ethtool
Call Trace:
netpoll_poll_disable
__dev_close_many
netif_close_many
netif_close
fbnic_self_test
dev_ethtool_locked
dev_ethtool
dev_ioctl
sock_ioctl
__x64_sys_ioctl
Without lockdep the same condition trips ASSERT_RTNL() in
__dev_close_many() / __dev_open(); that check only samples the global
rtnl state, so it can be masked by a concurrent rtnl holder, but the
device is still being reconfigured without the lock it requires.
The ethtool self_test is a legacy ioctl-only command, so an ETHTOOL_TEST
case is only needed on the ioctl path. Add an opt-in bit for drivers whose
self test needs rtnl_lock and set it on the ops-locked drivers whose
offline self test tears the interface down and up:
- fbnic (ops-locked via queue_mgmt_ops): fbnic_self_test() offline path
uses netif_close() / netif_open().
- bnxt (ops-locked via queue_mgmt_ops): bnxt_self_test() offline path
goes through bnxt_close_nic() / bnxt_half_open_nic() /
bnxt_half_close_nic() / bnxt_open_nic(), which close and reopen the
device.
Fixes: f994752b1127 ("net: ethtool: optionally skip rtnl_lock on IOCTL path")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942019771.7700.338431553546884773.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Pull drm fixes from Dave Airlie:
"Things have picked back up a bit this week, mostly amdgpu, xe and msm
this time. There are a bunch of scattered changes across the rest of
drivers and core stuff, nouveau, i915.
core:
- fix vblank pending event leak
ttm:
- swapout fixes
dma-buf:
- scattergather fixes
- enable dma-buf debug on debug kernels
dma-fence:
- fix signaling bit checks
sched:
- fix virtual runtime race
msm:
- DT:
- Corrected indentation
- Core:
- Marked fbdev as system memory
- GPU:
- Fixed autosuspend cleanup on teardown
- a750: fix timestamps
- Increase GMU fw init timeout
- Misc fixes/cleanups
- DPU:
- Fixed clock rounding, unbreaking newest platforms
- Cleared pending flush state
- DP:
- Skip PUSH_IDLE when link was never enabled
- Fixed bandwidth checks
- HDMI:
- Fixed runtime PM cleanup on probe failure
xe:
- shrinker related fixes
- xe_mmio_gem fault handler and destroy fixes
- xe disable i2c irq on unbind
i915:
- Revert a commit touching registers that don't necessarily exist
- Check for negative numbers before passing to BIT()
amdgpu:
- SMU 14.x fix
- DC IRQ fix
- Runtime PM fix for P2P
- RAS fix
- PCIe reporting fix
- DCN 6 fix
- Device removal fix
- DC MALL fix
amdkfd:
- GC 12.x fixes
- Boundary checks
- Mapping clear fix
nouveau:
- suspend/resume fixes
gud:
- out of bounds access fix
- ignore damage clips in full update
vc4:
- use-after-free fix
versilicon:
- plane format fix
longsoon:
- blend mode property fix"
* tag 'drm-fixes-2026-09-19' of https://gitlab.freedesktop.org/drm/kernel: (59 commits)
drm/amd/display: fix MALL hysteresis timer underflow at high refresh rates
drm/amdgpu: fix rmmio iounmap skipped on device removal
drm/amdgpu: Skip KFD mapping clear before initialization
drm/amd/display: Fix NULL dereference in dcn50/dcn60 init_hw
drm/amdkfd: Avoid integer underflow in EOP ring size calculation.
drm/amdkfd: Avoid integer underflow with ffs in EOP ring size calc
drm/amdgpu: Fix GPU PCIe link capability reporting
drm/amdgpu: check ras and obj before dereference
drm/amdgpu: hold a runtime PM reference for P2P dma-buf attachments
drm/amdkfd: implement restore_mqd callbacks for GFX12/12.1
drm/amd/display: Atomize IRQ register read/modify/write ops
drm/amd/pm: report energy accumulator for smu 14.0.3
drm/loongson: Create blend mode property for cursor plane
drm/xe/i2c: Disable IRQ on unbind
Revert "drm/i915/display: Clear SEL_FETCH_PLANE_CTL on plane disable"
drm/verisilicon: remove ARGB formats from primary plane
drm/verisilicon: add primary modifier for format tables
drm/verisilicon: set blend mode for the cursor plane
drm/sched: Fix virtual runtime race
drm/i915/display: check configuration index before shifting
...
|
|
A nested bpf_for_each_map_elem() callback can unlock a different element
of the same map:
static long inner(void *map, int *key, struct value *v,
struct value **outer_value)
{
bpf_spin_lock(&v->lock);
bpf_spin_unlock(&(*outer_value)->lock);
return 0;
}
static long outer(void *map, int *key, struct value *v, void *ctx)
{
bpf_for_each_map_elem(map, inner, &v, 0);
return 0;
}
Both callback values currently have ID zero and the same map_ptr.
process_spin_lock() compares those two fields, so it accepts the unlock
even though the two callbacks can receive different map elements.
Assign a fresh ID to every callback map value in the for-each,
timer/workqueue, and task-work constructors. Copies of one callback
argument retain its ID, so locking and unlocking through that argument
continues to work. Distinct callbacks also get distinct IDs for
single-element arrays, including inner arrays sharing inner_map_meta.
Preserve map_uid for every inner-map lookup and compare it through
check_ids() during state pruning. This preserves relationships between
maps, keys, and values while allowing equivalent states with different
lookup IDs to match. It avoids field-specific rules for when an inner map
needs an identity.
Move map_uid out of the metadata union and next to the other IDs, so
register comparisons can use the existing memcmp() ranges and remap the
IDs separately. Clear it when resetting a register or converting a map
lookup result to a socket pointer. Shrink frameno to u8, which is enough
for MAX_CALL_FRAMES, to make room without growing bpf_reg_state.
Fixes: d0d78c1df9b1 ("bpf: Allow locking bpf_spin_lock global variables")
Reported-by: Nicholas Carlini <npc@anthropic.com>
Suggested-by: Nicholas Carlini <npc@anthropic.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260917233222.2542500-9-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
|
|
check_subprogs() verifies that each subprogram ends in an exit or an
unconditional jump before in-kernel CO-RE relocations are applied. An
unresolved relocation can then replace that terminal instruction with an
invalid helper call. The resulting fall-through into another subprogram
breaks the CFG invariant used by postorder and stack liveness analysis,
which can write past their per-subprogram arrays.
Apply CO-RE relocations immediately after preparing the program BTF, before
subprogram discovery and validation. Keep func_info and line_info validation
after subprogram discovery because those records depend on the complete
subprogram layout.
Reject an ldimm64 first slot at the end of the instruction stream before
CO-RE can inspect its missing second slot. check_subprogs() previously
rejected this form before relocation processing because it is not a valid
subprogram terminator. Moving CO-RE ahead of check_subprogs() removes that
implicit protection, so perform an explicit check before applying
relocations.
Include core_relo_cnt when deciding whether to prepare program BTF. A load
that supplied only CO-RE relocation metadata previously skipped both BTF
setup and relocation processing.
Fixes: fbd94c7afcf9 ("bpf: Pass a set of bpf_core_relo-s to prog_load command.")
Suggested-by: Andrii Nakryiko <andrii@kernel.org>
Suggested-by: Alexei Starovoitov <ast@kernel.org>
Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260917233222.2542500-5-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
|
|
Global subprograms are verified independently with a fresh verifier root.
do_check_common() currently seeds that root's in_sleepable state from the
program, even though a global subprogram can also run from callbacks whose
execution context differs from the program's main entry point.
In particular, workqueue and task-work callbacks are sleepable even when
the containing program is not. A global subprogram of that program is
therefore verified as non-sleepable, making in_rcu_cs() true and allowing
loads of RCU-protected kptrs to produce trusted MEM_RCU pointers. The same
subprogram can then be called from a sleepable callback without a classic
RCU reader. It can retain such a pointer while the object is freed and use
it after free.
The verifier's execution-context predicates are complementary. A state is
sleepable only when in_sleepable is set and no RCU, preemption, IRQ, or lock
region is active. Each condition which prevents sleeping also provides RCU
protection, while in_rcu_cs() treats a non-sleepable state as implicitly
protected.
Use this relationship to represent a global subprogram caller with only the
result of in_sleepable_context(). A protected sleepable caller is normalized
to in_sleepable=false at the independent verification root. This both
prevents sleepable operations and makes in_rcu_cs() true without copying
caller-owned lock state.
Track only the contexts in which each global subprogram is actually
reached. Verify it once if all reachable calls use the same context, and
twice only if both sleepable and non-sleepable calls reach it. Calls found
while verifying globals or asynchronous callbacks mark further contexts
for checking. Repeat the existing subprogram walk until all called
contexts have been verified; unreachable global calls remain unchecked.
Accumulate instruction counts over those verification passes. Preserve
the total recorded before each pass, since path accounting has already
added this pass's synchronous instructions and its root total must also
include asynchronous subprograms.
This makes an unprotected callback verify the global subprogram as
sleepable, turning its RCU-protected kptr load into an untrusted pointer.
Protected callers and global subprograms which do not depend on implicit RCU
protection remain valid.
Fixes: 81f1d7a583fa ("bpf: wq: add bpf_wq_set_callback_impl")
Fixes: 38aa7003e369 ("bpf: task work scheduling kfuncs")
Reported-by: Nicholas Carlini <npc@anthropic.com>
Suggested-by: Nicholas Carlini <npc@anthropic.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260914131923.2544250-2-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Paolo Abeni:
"Including fixes from Netfilter, Bluetooth, IPSec and WiFi.
Previous releases - regressions:
- netfilter: hold reference on ct until flow is released
- bridge:
- move switchdev call outside rcu
- vlan: fix bugs caused by switchdev deletion errors
- wifi:
- mac80211: reset state when starting AP fails
- cfg80211: don't free driver-owned scan requests
- tcp: don't call skb_clone_and_charge_r() for close()d listener in
tcp_v6_do_rcv()
- mptcp: return sk_wait_data() errors from recvmsg()
- xfrm: serialize state GC with device state flush
- drop_monitor: synchronize tracepoint unregistration on error path
- bluetooth:
- eir: validate service data length before reading UUID
- hci_sync: serialize local codec list cleanup
- RFCOMM: avoid socket lock inversion in listener cleanup
- eth:
- lan743x: fix RX checksum use-after-free
- mvpp2: prevent buffer overflow in page_pool allocation
Previous releases - always broken:
- core: lock the socket in sock_gettstamp()
- neighbour: enforce min/max to NDTPA_INTERVAL_PROBE_TIME_MS.
- sched: codel: bound the dropping loop per dequeue call
- wifi: mac80211: include TIM bitmap control for buffered S1G mcast
traffic
- psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()
- xfrm: fix stack OOB read in iptfs_skb_reset_frag_walk()
- bluetooth: hci_qca: do not write to the serial port after it is
closed
- dsa: mxl862xx: disable the stats poll on teardown
- eth:
- stmmac: fix TSO header length truncation
- ip_tunnel: initialize `options_len` before referencing options"
* tag 'net-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (159 commits)
mptcp: fix bad accounting in __mptcp_subflow_push_pending()
mptcp: close race between scheduler and state change
mptcp: avoid unneeded actions on subflow reset
net: skbuff: do not leave stale header offsets after pskb_carve()
selftests: net: packetdrill: test exclusion of old ACK from TCP fast path
tcp: exclude old ACKs from tcp fast path
dpll: reject a reference sync pin which is not on the pin's dpll
net: mvpp2: prevent buffer overflow in page_pool allocation
net: macb: fix ordering around PTP timestamp read
selftests: drv-net: psp: test PSP and TCP ULP mutual exclusion
net: psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()
net: stmmac: preserve real_num_tx_queues on mqprio setup failure
net: stmmac: propagate FPE preemption-class mapping errors
net: wwan: t7xx: validate the netif index in t7xx_ccmni_recv_skb()
net: wwan: mhi_wwan_mbim: check skb_copy_bits() return value
net: wwan: mhi_wwan_mbim: guard against a cyclic NDP chain
net: ethernet: cortina: Ack RX overrun interrupt correctly
net: lock the socket in sock_gettstamp()
eth: fbnic: ring the doorbell if a burst ends in a drop
net: netsec: fix device_node reference leak on phy_np
...
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless
Johannes Berg says:
====================
Many fixes:
- mac80211: S1G TIM bitmap fix
- ath12k: remove undocumented DT ABI implementation
- various firmware API and over-the-air hardening changes
- fixes for most cfg80211/mac80211 syzbot reports
* tag 'wireless-2026-09-16' of https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless: (67 commits)
wifi: brcmsmac: fix UAF in brcms_free_timer()
wifi: brcmfmac: fix lost 802.1x TX completion wakeup
wifi: ath11k: cleanup arsta in ath11k_mac_peer_cleanup_all()
wifi: wcn36xx: Fix potential use-after-free in TX ack timer teardown
wifi: ath12k: ahb: Revert undocumented ABI and dead code
wifi: mac80211: refuse to make a monitor active when it has no queue
wifi: libipw: reject TKIP frames without a full MIC
wifi: virt_wifi: don't transfer operstate before register
wifi: cfg80211: check if AP has been started or joined a mesh before adding new station
wifi: cfg80211: move link_id validation earlier in nl80211_new_station()
wifi: cfg80211: do not support direct add of station to AP_VLAN interfaces
wifi: cfg80211: verify if AP_VLAN belongs to the correct AP
wifi: mac80211: set up the TX info early to fix failure paths
wifi: mac80211: mesh: release the channel if start fails
wifi: mac80211: mesh: reset the CSA state when leaving
wifi: mac80211: add HE 6 GHz capability in the scan elems len
wifi: mac80211: don't access the TSF of a down interface
wifi: mac80211: don't RCU-dereference the mesh CSA settings we just set
wifi: mac80211: don't allow link changes when iface is down
wifi: mac80211: require a peer station for TDLS setup confirm
...
====================
Link: https://patch.msgid.link/20260916083642.110609-3-johannes@sipsolutions.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Backmerging to get drm-misc-fixes up to v7.3-rc3.
Signed-off-by: Maxime Ripard <mripard@kernel.org>
|
|
The patch "dma-buf: dma-fence: Fix potential NULL pointer dereference"
changed the check to test for the ops pointer instead of the signaled
bit to avoid a potential NULL dereference when the ops pointer has been
cleared.
The problem is now that the ops pointer is cleared only when neither the
release nor the wait callback is implemented and this isn't true for a lot
of dma_fence implementations yet. So those implementations lost the RCU
protection after signaling of the returned string resulting in potential
use after free.
Add the signaling check additional to the ops pointer check so that we
have both the protection against NULL dereference as well as the RCU
protection after signaling for the returned string.
v2: improve comments to note RCU protection and explain why we check
both signaling state and ops pointer
v3: some comment improvements suggested by Philip
Signed-off-by: Christian König <christian.koenig@amd.com>
Fixes: 035219a760ed ("dma-buf: dma-fence: Fix potential NULL pointer dereference")
CC: stable@vger.kernel.org # 7.2+
Reported-by: Jonghyuk Kim(MalHyuk) <malhyuk97@gmail.com>
Tested-by: Jonghyuk Kim(MalHyuk) <malhyuk97@gmail.com>
Reviewed-by: Philipp Stanner <phasta@kernel.org>
Link: https://lore.kernel.org/r/20260914182740.1587-1-christian.koenig@amd.com
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull x86 fixes from Dave Hansen:
"The most notable fix is THP not silently losing user data and having
been around for a couple of years. The main explanation I'd have for
its longevity is that it requires a few different things to align at
the same time: MADV_FREE, THP and heavy reclaim.
- Fix user-space data loss with THP
- Fix set_memory oopses
- Fix addition of large constants in mul_u64_add_u64_div_u64()
- Fix FineIBT hash offset in cfi_get_func_hash()
- Fix PCI device reference counting in amd_smn_init()"
* tag 'x86_urgent_for_7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
x86/amd_node: Fix PCI device reference counting in amd_smn_init()
x86/div64: Fix addition of large constants in mul_u64_add_u64_div_u64()
x86/cfi: Fix FineIBT hash offset in cfi_get_func_hash()
x86/mm: Fix user-space data loss with MADV_FREE and THP
x86/mm/pat: Allocate split page tables as kernel page tables
x86/alternatives: Exclude text poking against change_page_attr()
x86/mm/pat: Acquire init_mm read lock on attribute changes to avoid UAF
x86/mm/pat: Acquire init_mm write lock on collapse to avoid UAF
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace
Pull tracing fixes from Steven Rostedt:
- Don't destroy user event fields when removal fails
User event fields are destroyed before the event is removed from
visibility. But that can fail leaving the still visible event with no
fields. Move the destroying of the fields to after the event is
successfully removed from visibility.
- Initialize function graph state is fork before calling
copy_exec_state()
For non-CLONE_VM forks, copy_exec_state() allocates a new
task_exec_state. If that allocation fails, ftrace_graph_exit_task()
will free the tasks ret_stack pointer. Since that pointer is still
using the parent's ret_stack, it mistakenly frees the parent's
pointer too.
Call ftrace_graph_init() on the task first which will NULL out the
new tasks's ret_stack and if the copy fails, it will not free
anything.
- Remove FGRAPH_MAX_INDEX
The macro FGRAPH_MAX_INDEX was added but never used. Remove it.
- Save ent_size in function graph printing of nested functions
The function graph tracer needs to look at the next event to see if
the next event is the return of the current function entry. If it is,
it prints a single line:
ktime_get();
Otherwise it prints it like a nested function:
tick_nohz_irq_exit() {
ktime_get();
kcpustat_irq_exit();
}
In order to look at the next event, it must save the current event so
that it has the information to print from it. It saves the event in
the iterator descriptor called "ent". What it doesn't save is the
ent_size of the event which is now used to know if the function graph
arguments are to be printed. The peek doesn't save the size so the
size used happens to be that of the size of the last event that was
seen.
Save the entry event size in the iterator descriptor so that the
correct size is used.
- Fix several errors with freeing data in the histogram code
The histogram code had a lot of leaked or or incorrect accounting
when failures happen. Correct them.
- Fix histogram regression of .percent and .graph modifiers
Up until 6.3 histogram values could have "percent" or "graph"
modifiers that changed how they were printed. But a change that added
restricting histograms values from being strings, stack traces and
other modifiers inadvertently prevented them from using the percent
and graph modifiers, which were legal use cases for values.
Put back the percent and graph modifiers.
- Fix various typos in the comments
- Set the trace_clock before initializing a histogram with clock
argument
The histogram API allows the user to specific which trace clock to
use via a "clock=" string. The histogram is set up first before the
clock is checked. If the passed in clock is not valid, it exits
without fully fixing up the histogram leaving it on the list and a
use-after-free can trigger.
Update the clock argument first and if it fails then exit gracefully
before the histogram trigger is placed on any lists.
- Restore :mod: trailer after parsing in ftrace_set_clr_event
The function ftrace_set_clr_event() modifies the parse string and
needs to put it back to what was passed in. It searches for ":mod:"
via a strsep() but fails to put back the first ':' in the string.
Add back the ':' in the passed in string.
- Take trace_array reference when opening a tracer options file
The options files are dynamically created and some tracers add their
own options. When a tracer adds their own list of options, the
trace_array holding them has an array to hold the list of options for
each tracer. This array increases in size via a krealloc(), and the
new entry gets a newly allocated array to hold the options of the new
tracer being added.
The element in each entry of the tracer's option array holds a
pointer back to the trace_array, a pointer to the tracer it is
associated to, a pointer to the flags of the option.
The issue is that these arrays are freed when the trace_array is
freed when its instance it represents is removed from the instances
directory. There's a race that an open of one of these options files
can happen when the instance is being removed.
Add a new helper function to be called by the open function of the
options file to iterate all existing trace_arrays under a lock and
find the one that has the given option element in one of it's tracer
arrays. If found, then update the associated trace_array's reference
counter to keep it from being freed. If not found, have the open call
return -ENODEV.
- Disable interrupts when acquiring the lock in rb_wake_up_waiters()
The function rb_wake_up_waiters() assumes it will be called in
interrupt context and does not disable irqs when taking
cpu_buffer->reader_lock, which can be called in hard interrupt
context. The issue is in PREEMPT_RT, this function is called in
thread context leaving this lock open to a deadlock.
Take the lock with interrupts disabled.
- Use rcu_assign_pointer() for tmp_ops filter hash
The tmp_ops used in update_ftrace_direct_mod() assigns its
filter_hash field directly, but that field is annotated as __rcu and
sparse complains. Assign it with rcu_assign_pointer()
- Fix use-after-free in enable_trigger_private_data_free()
The trace_event_call is accessed through the event_trigger_data's
trace_event_file pointer to put the trace_event_call on freeing. The
issue is that the trace_event_file data may have been freed already
causing a use-after-free. Add a field to the event_trigger_data that
points directly to the trace_event_call so that it can decrement its
reference directly without needing to go through the
trace_event_file.
- Fix accounting of buffer data remote headers
trace_buffer_desc_size() and trace_remote_alloc_buffer() undercount
the number of pages is needed for the asked for size as it doesn't
take into account the meta data on each page. Add a helper function
to do the calculation properly and use that in these functions.
- Catch nr_page_va overflow in ring_buffer_desc sizing
The number of pages per remote ring buffer is capped by
ring_buffer_desc::nr_page_va (32 bits). A buffer_size large enough to
overflow that field would silently allocate a descriptor smaller than
what was asked for.
- Do not resize the subbuf order if any per_cpu buffer is disabled
The mmapping of ring buffers disables resizing the subbuffers, but it
is done per-cpu whereas the subbuf size change is done for all the
per_cpu buffers under the buffer->mutex. It could change the size of
some while the mapping is happening on others. Have the resize of the
subbuf order check all the per_cpu buffers under the lock to see if
any of them is disabled before starting and causing an inconsistency
between buffers that are being mapped.
* tag 'trace-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace: (25 commits)
ring-buffer: Check resize_disabled before publishing the new subbuf order
tracing/remotes: Catch nr_page_va overflow in ring_buffer_desc sizing
tracing/remotes: Account for ring buffer page header in size calculation
tracing: Don't dereference trace_event_file in deferred trigger free
ftrace: Use rcu_assign_pointer() for tmp_ops filter hash
ring-buffer: Acquire the lock with irqsave in rb_wake_up_waiters()
tracing: Take trace_array reference when opening a tracer options file
tracing: Fix ring_buffer_read_page_size() kernel-doc
tracing: Restore :mod: trailer after parsing in ftrace_set_clr_event()
tracing: Fix memory corruption from a "STACKTRACE" histogram key
tracing: Fix memory corruption from the histogram stacktrace modifier
tracing: Undo the registration when enabling the histogram trigger fails
tracing: Take the reference before publishing the named histogram trigger
tracing: Set the trace clock before registering the histogram trigger
tracing: Fix typo "preceeded" in comment
tracing: Fix typo "availabe" in comment
tracing: Let histogram values keep the percent and graph modifiers
tracing: Keep the entry count when the histogram stats allocation fails
tracing: Free histogram the field rejected for a bad modifier
tracing: Free histogram the var ref when its initialization fails
...
|
|
The number of pages per remote ring buffer is capped by
ring_buffer_desc::nr_page_va (32 bits). A buffer_size large enough to
overflow that field would silently allocate a descriptor smaller than
what was asked for.
Return SIZE_MAX from trace_buffer_desc_size() on nr_page_va overflow.
Link: https://patch.msgid.link/20260911193937.602202-3-vdonnefort@google.com
Fixes: 2e67fabd8b77 ("ring-buffer: Introduce ring-buffer remotes")
Signed-off-by: Vincent Donnefort <vdonnefort@google.com>
Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
|
|
trace_buffer_desc_size() and trace_remote_alloc_buffer() undercount the
required pages because every ring buffer page contains a header
(BUF_PAGE_HDR_SIZE). Account for that header to ensure allocated remote
ring buffers aren't smaller than requested by the user.
The newly introduced helper __calc_nr_pages_ring_buffer_desc() can
return a value that overflows the descriptor nr_pages field (32 bits).
Link: https://patch.msgid.link/20260911193937.602202-2-vdonnefort@google.com
Fixes: 2e67fabd8b77 ("ring-buffer: Introduce ring-buffer remotes")
Signed-off-by: Vincent Donnefort <vdonnefort@google.com>
Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull scheduler fixes from Ingo Molnar:
- Fix EEVDF se->max_slice value on enqueueing (Vincent Guittot)
- Fix EEVDF augmented rb-trees re-balancing with multiple
fields (Vincent Guittot)
- In proxy scheduling, account cgroup CPU time to the execution
context, not the scheduling context (Hui Su)
- Likewise, call wq_worker_tick() for the execution context,
not the scheduling context (Hui Su)
* tag 'sched-urgent-2026-09-13' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
sched/core: Call wq_worker_tick() for the execution context
sched: Account cgroup CPU time to the execution context
sched/eevdf: Fix rb augmented with multi fields
sched/eevdf: Fix augmented max_slice
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull entry code fix from Ingo Molnar:
- Fix generic entry code cross-build failure on
!CONFIG_AUDITSYSCALL kernels using older
RISCV64 and S390 cross-compilers (Thomas Gleixner)
* tag 'core-urgent-2026-09-13' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
entry: Guard syscall_enter_audit() invocation with CONFIG_AUDITSYSCALL
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/riscv/linux
Pull RISC-V fixes from Paul Walmsley:
"From a RISC-V point of view, there's one notable fix here, reverting
an earlier bogus fix to the pointer masking code. Fortunately the
practical impact appears to be small.
- Revert a bad fix, likely LLM-generated, in the pointer masking code
that confused the RISC-V hardware pointer masking implementation
with the Linux kernel tagged address feature
- Fix unexpected faults caused by kprobe instruction slot writes when
!CONFIG_STRICT_MODULE_RWX
- Fix unexpected faults on minimal configurations during runtime code
patching on !CONFIG_STRICT_MODULE_RWX systems
- Fix a misplaced variable clear causing incorrect reuse of previous
values in the RISC-V hardware feature probing code
- Fix two bugs in the PMU SBI perf code on rv32: use BIT_ULL rather
than BIT on 64-bit masks; and use a bitmap rather than an unsigned
long on a quantity that can exceed 32 bits
And a few miscellaneous cleanups:
- Avoid a potential dereference-before-NULL-pointer-check bug in the
PMU SBI perf driver
- Use CONFIG_GENERIC_BUG_RELATIVE_POINTERS to simplify the rv32 bug
table code (like x86 and PPC)
- Report the RISC-V standard ISA extensions Z[v]fhmin when support is
claimed for the superset RISC-V standard ISA extensions Z[v]fh; and
simplify our FPU test code to only check for the presence of the D
extension
- Use an existing kernel string helper in place of some open-coded
code in kernel/usercfi.c
- Fix some yamllint issues in the RISC-V DT bindings for CPUs
- Convert one use of __ASSEMBLY__ to __ASSEMBLER__ that snuck into
the RISC-V CFI selftest code
- Update the translation for the simplified Chinese translation of
the RISC-V kernel patch acceptance policy"
* tag 'riscv-for-linus-7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/riscv/linux:
riscv: skip software algning code for HAVE_EFFICIENT_UNALIGNED_ACCESS
kselftest/riscv: Replace __ASSEMBLY__ with __ASSEMBLER__
docs/zh_CN: Update arch/riscv/patch-acceptance.rst translation
dt-bindings: riscv: cpus: Fix yamllint style issues
riscv: hwprobe: simplify has_fpu() to check D extension only
perf: RISC-V: check cpu_hw_evt before dereference in overflow IRQ
riscv: report Zfhmin/Zvfhmin when Zfh/Zvfh are present
perf: RISC-V: store available counter mask as bitmap
perf: RISC-V: use BIT_ULL for u64 overflow masks
riscv: bug: Make RV32 use GENERIC_BUG_RELATIVE_POINTERS
riscv: hwprobe: initialize pair->value in hwprobe_one_pair()
riscv: use string helper in setup_global_riscv_enable()
Revert "riscv: Reset pmm when PR_TAGGED_ADDR_ENABLE is not set"
riscv: patch: skip fixmap mapping when kernel text is already writable
riscv: mm: make EXECMEM_KPROBES writable without CONFIG_STRICT_MODULE_RWX
|
|
Use kvm_test_request() instead of kvm_check_request() when querying
KVM_REQ_VM_DEAD, i.e. don't clear KVM_REQ_VM_DEAD, as the entire purpose
of KVM_REQ_VM_DEAD is to prevent the vCPU from enterring the guest ever
again, even if userspace insists on redoing KVM_RUN.
Ensuring KVM_REQ_VM_DEAD is never cleared will allow relaxing KVM's rule
that ioctls can't be invoked on dead VMs, to only disallow ioctls if the
VM is bugged, i.e. if KVM hit a KVM_BUG_ON().
Opportunistically add compile-time assertions to guard against clearing
KVM_REQ_VM_DEAD through the standard APIs.
Reviewed-by: Kai Huang <kai.huang@intel.com>
Acked-by: Marc Zyngier <maz@kernel.org>
Link: https://patch.msgid.link/20260806214618.82180-1-seanjc@google.com
Signed-off-by: Sean Christopherson <seanjc@google.com>
|
|
Tracing related BPF helpers e.g. under bpf_base_func_proto() are gated
behind CAP_PERFMON. However, the same is currently not true for kfuncs
and they are accessible via plain CAP_BPF. Add a new KF_PERFMON flag
which can be used such that check_kfunc_call() ensures env->allow_ptr_leaks
is permitted. This follows similar pattern to existing KF_DESTRUCTIVE flag.
The rejection returns -EPERM to match the other CAP_PERFMON gates in the
verifier, that is, check_ptr_to_btf_access() and check_ptr_to_map_access(),
which report the very same policy to user space.
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://lore.kernel.org/r/20260910213510.49358-1-daniel@iogearbox.net
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Jakub Kicinski:
"Nothing too exciting, usual stream of fixes. Including fixes from
Netfilter, Bluetooth and WPAN.
Current release - new code bugs:
- Bluetooth: hci_sync: fix not setting CE length properly
- eth: enic: match mailbox replies to request numbers
Previous releases - regressions:
- tunnels: drop stale dst when building an ICMP error for PMTUD
- ipv6: null-check fib6_node before accessing in __ip6_del_rt_siblings()
(bug in the rtnl_lock -> RCU conversion)
- eth: bnxt_en:
- fix crashes on Thor2 due to OOB coalescing buffer accesses
- prevent queue stop with deferred completions
Previous releases - always broken:
- eth:
- ice: don't dereference pointers from TP_printk()
- fix OOB writes on ethtool flow rule dump in 3 drivers
- mlx5: fix FEC configuration with RS_544_514_INTERLEAVED_QUAD
- dsa: tag_brcm: legacy FCS: request needed tailroom
Misc:
- net: cap tx_queue_len at S16_MAX to prevent oversized ring alloc
- ipv6: flowlabel: cap duplicate leases per socket"
* tag 'net-7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (164 commits)
selftests: tc-testing: test action batch failure cleanup
net/sched: act_api: release all action references on NEWACTION failure
openvswitch: fix wrong flag value in get_ipv6_ext_hdrs()
ipmr: account multicast table and route memory
net: phy: dp83td510: handle the active-high LED polarity mode
net: macb: initialize PTP state before registering clock
net: hsr: enable promiscuous mode on interlink port with fwd offload
ipv6: fix fib6 walker UAF on seq stop
net: stmmac: fix TX descriptor availability check for TSO traffic
net/rds: fix tcp stream corruption with large pages
net: mana: restore the XDP program pointer when pre-allocation fails
net: phy: dp83867: handle the active-high LED polarity mode
octeontx2-af: fix PF/CGX debugfs PCI bus lookup
net: net_failover: Fix the deadlock in net_failover_slave_name_change()
net: phy: mediatek-ge: disable EEE on the MT7530 PHY
tcp: reject non zerocopy devmem tx
net: ethernet: mtk_eth_soc: populate lpi_interfaces to fix EEE support
net: dsa: mt7530: populate lpi_interfaces to fix EEE support
net: hinic: fix mailbox segment buffer overflow
net: sun4i-emac: fix missing of_node_put() for phy_node
...
|
|
The eevdf rb tree maintains 3 augmented fields but only one is currently
copied when balancing the tree.
Add a more generic define that can be used when there are several augmented
fields. In this case, we provide a function that takes care of copying all
fields.
Fixes: aef6987d8954 ("sched/eevdf: Propagate min_slice up the cgroup hierarchy")
Signed-off-by: Vincent Guittot <vincent.guittot@linaro.org>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Reviewed-by: K Prateek Nayak <kprateek.nayak@amd.com>
Tested-by: K Prateek Nayak <kprateek.nayak@amd.com>
Link: https://patch.msgid.link/20260909150522.858312-1-vincent.guittot@linaro.org
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs
Pull vfs fixes from Christian Brauner:
- netfs:
- Fix an uninitialized return value in netfs_unbuffered_write()
when preparing the first subrequest fails
- For partial unbuffered/DIO writes return the amount transferred
rather than an error
- Update i_size with the amount actually written when a partial
transfer ends in an error
- Fix a subrequest reference leak when the io_iter ends up empty
- Handle netfs_alloc_subrequest() failure during unbuffered writes
- Load all readahead folios into the rolling buffer upfront and
drop the readahead references once the first subrequest is
dispatched
- Mark folios for copy-to-cache while issuing subrequests
- Fix read progress reporting
- afs:
- Add the missing kunmap in the error path of afs_dir_search_bucket()
- Fix a double kunmap in afs_edit_dir_remove()
- Don't free an existing server's endpoint state when cleaning up a
candidate server in afs_lookup_server()
- Unbind peers removed from a server's address list
- ufs:
- Load the cylinder group metadata before creating the root dentry
- Validate the cylinder group index and rotor positions before
caching them
- Treat an unreadable directory block as not empty
- exec:
- Close the close-on-exec files before taking exec_update_lock
Closing a file can block on the filesystem, so a hung filesystem
blocked everything that takes exec_update_lock and a FUSE server
inspecting the calling process could deadlock
- Drop the bprm loader before closing bprm->file in free_bprm()
- exit: Hold a reference to thread_pid across proc_flush_pid()
- reboot: Fix a use-after-free on cad_pid
- nsfs: Keep the namespace tree fields out of the rcu_head used by
kfree_rcu()
- nstree: Check listing permission before taking a namespace
reference in listns()
- super: Return 0 when a nested thaw drops its hold while other
freezers remain
- ext4: Don't set I_METADATA_WRITEBACK during fastcommit replay
- adfs: Free s_fs_info in ->kill_sb()
- autofs: Free the inode info allocated in autofs_fill_super() when
the root inode allocation fails
- ovl: Return EINVAL instead of EIO on a user namespace mismatch now
that it's a plain refusal and not an internal error
- cachefiles: Don't cast the variable-length coherency data to a
__be64 in the coherency tracepoint
* tag 'vfs-7.3-rc3.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs: (28 commits)
nstree: check listing permission before taking a namespace reference
exec: do_close_on_exec() before taking exec_update_lock
exit: hold a reference to thread_pid across proc_flush_pid
fs: autofs: fix memory leak in autofs_fill_super()
exec: Drop bprm loader before closing bprm->file
afs: Clear stale peer app data after address list changes
afs: Fix incorrect free in candidate cleanup in afs_lookup_server()
afs: Fix double-unmap of directory block
afs: Fix missing kunmap in afs_dir_search_bucket()
ovl: return EINVAL instead of EIO in case of mismatched user_ns
reboot: fix cad_pid use-after-free race
cachefiles: Fix potential UAF/KASAN warning
netfs: Fix read progress reporting
netfs: Mark folios with COPY_TO_CACHE whilst issuing subreqs
netfs: Fix readahead synchronisation issues by loading all folios upfront
netfs: break unbuffered write when netfs_alloc_subrequest() fails
netfs: Fix subreq ref leak
netfs: Fix i_size update for partial transfer
netfs: Fix error vs transferred passed to ->ki_complete()
netfs: Fix unbuffered/DIO write partial transfer error return
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/printk/linux
Pull printk fixes from Petr Mladek:
- Use lazy irq_work for waking printk kthreads
- Flush pending irq_work before destroying printk kthreads
- Remove redundant WARN() when a printk kthread can't be created
- Typo fix
* tag 'printk-for-7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/printk/linux:
printk/nbcon: Change nbcon_irq_work to IRQ_WORK_LAZY
printk/nbcon: Flush nbcon_irq_work in nbcon_free()
console: fix /dev/kmsg reference in flags kernel doc
printk: Don't WARN on kthread_run failure.
|