summaryrefslogtreecommitdiff
path: root/arch/x86/kernel/cpu
AgeCommit message (Collapse)Author
18 hoursMerge branch 'driver-core-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/driver-core/driver-core.git
18 hoursMerge branch 'next' of https://github.com/kvm-x86/linux.gitMark Brown
18 hoursMerge branch 'master' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git # Conflicts: # Documentation/scheduler/index.rst # arch/arm64/configs/defconfig
19 hoursMerge branch 'modules-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/modules/linux.git
35 hoursMerge branch 'misc'Sean Christopherson
* misc: (44 commits) KVM: selftests: Add coverage for the EFER_LMSLE_MBZ defeature KVM: selftests: Add module param API to check if nested virtualization is enabled KVM: selftests: Rename svm_nested_clear_efer_svme to svm_nested_efer_test KVM: x86: Honor the guest's EFER_LMSLE_MBZ KVM: x86: Advertise EFER_LMSLE_MBZ when KVM disallows EFER.LMSLE KVM: SVM: Add support for virtualizating Bus Lock Detect KVM: SVM: Add a vCPU-aware helper to get supported DEBUGCTL bits KVM: nSVM: Don't assume all active-low bits DR6 are fixed-1 KVM: nSVM: Open code check on LBR virtualization being enabled in vmcb12 KVM: nSVM: Disable LBRV in nested control cache when unsupported KVM: SVM: Add helper to query if LBR virtualization needs to be enabled KVM: x86: Kill off DR6_FIXED_1 to prevent future misuse KVM: x86: Set guest's fixed-1 DR6 bits when delivering #DB payload KVM: x86: Force fixed-1 bits in DR6 after synchronizing with hardware KVM: Return an "unsigned long", not "u64" for the fixed-1 DR6 bits KVM: x86: Rename kvm_dr6_fixed() => kvm_get_dr6_fixed_1() KVM: x86: Preserve DR6.BLD (Bus Lock Detect) when delivering #DB payload KVM: x86: Restrict saved GPA writes to hardware write faults KVM: x86/pmu: Don't retry a counter whose config was rejected KVM: SVM: Add Page modification logging support ...
2 daysMerge branch into tip/master: 'x86/sgx'Ingo Molnar
# New commits in x86/sgx: 993b65a2fc6f ("x86/sgx: Report RCU-Tasks quiescent state in EPC sanitization loop") Signed-off-by: Ingo Molnar <mingo@kernel.org>
2 daysMerge branch into tip/master: 'x86/sev'Ingo Molnar
# New commits in x86/sev: 250ee734115b ("x86/sev: Report MSR_AMD64_SEV in sysfs") edb1b49e724e ("x86/sev: Do RMP optimizations on SNP guest shutdown") c6c6f156d192 ("x86/sev: Perform RMP optimizations asynchronously") 15d974be0a3b ("x86/sev: Initialize RMPOPT configuration MSRs") 674bb17a32c7 ("x86/sev: Disable CPU hotplug while SNP is active") 1ee97d0f7b20 ("x86/cpufeatures: Add X86_FEATURE_RMPOPT feature flag") Signed-off-by: Ingo Molnar <mingo@kernel.org>
2 daysMerge branch into tip/master: 'x86/microcode'Ingo Molnar
# New commits in x86/microcode: cf086537af4e ("x86/microcode/intel: Refresh old_microcode defines with the May 2026 release") Signed-off-by: Ingo Molnar <mingo@kernel.org>
2 daysMerge branch into tip/master: 'x86/cpu'Ingo Molnar
# New commits in x86/cpu: 2711d67bc3a7 ("x86/cpu: Don't transiently clear the boot CPU's capabilities") db349783ca7a ("x86/cpu: Move 32-bit SEP setup into identify_cpu()") ad86fe2134cc ("x86/cpu: Inline generic_identify() into identify_cpu()") 7a762b51a149 ("x86/cpu: Initialize boot CPU cpuinfo defaults early") 54a2882ab9f7 ("x86/cpu: Factor init_cpu_info() out of identify_cpu()") 228200f695c0 ("x86/CPU/AMD: Fix Zen5 TLB sizes reporting") 9025f3bf73f0 ("x86/shstk: Return the correct error value when user shadow stacks are disabled") Signed-off-by: Ingo Molnar <mingo@kernel.org>
2 daysMerge branch into tip/master: 'x86/cleanups'Ingo Molnar
# New commits in x86/cleanups: 2c6a75adb15f ("x86/um: Remove unused <asm/required-features.h> header") c7243accb48a ("x86/cpu: Constify struct x86_cpu_id") 975809caf4b0 ("x86/asm/string_64: Remove unused linux/jump_label.h include") c8beda268a6b ("x86/mtrr: Fix kernel-doc notation of amd_set_mtrr()") Signed-off-by: Ingo Molnar <mingo@kernel.org>
2 daysMerge branch into tip/master: 'x86/cache'Ingo Molnar
# New commits in x86/cache: a17d099481bd ("x86/resctrl: Update documented unit for the "activity" event") a3f5e6eba418 ("fs/resctrl: Simplify pseudo_lock_measure_trigger()") 4b31656d917c ("fs/resctrl: Avoid extra call to strlen() in schemata_list_add()") 942135446059 ("fs/resctrl: Factor MBA parse-time conversion to be per-arch") 255f92316ac1 ("arm_mpam: resctrl: Add pass-through resctrl_arch_preconvert_bw()") 57a354d20a01 ("x86,fs/resctrl: Add resctrl_arch_preconvert_bw()") Signed-off-by: Ingo Molnar <mingo@kernel.org>
2 daysMerge branch into tip/master: 'x86/bugs'Ingo Molnar
# New commits in x86/bugs: ae1d2082d93b ("x86/bugs: Adapt SRSO mitigation to Zen6") Signed-off-by: Ingo Molnar <mingo@kernel.org>
2 daysMerge branch into tip/master: 'sched/core'Ingo Molnar
# New commits in sched/core: 1fb28c664a19 ("virt/steal_governor: Enable the driver") 27d47ebce4d6 ("virt/steal_governor: Implement steal_governor policy loop") 4b9302d494ff ("virt/steal_governor: Add control knobs for handling steal values") 9a8e740ee9f6 ("virt: Introduce steal governor driver") 68957caaa9c0 ("sched/debug: Add migration stats due to non preferred CPUs") 74699f56ebcf ("sched/core: Push current task from non preferred CPU") 4ee29b029058 ("sched/fair: Load balance only among preferred CPUs") d8a3da0de843 ("sched/core: Try to use a preferred CPU in is_cpu_allowed") 620824516557 ("sysfs: Add preferred CPU file") 518b32bd5bb3 ("cpumask: Introduce cpu_preferred_mask") 06a49ef784ac ("sched/docs: Document cpu_preferred_mask and Preferred CPU concept") cfb463b7172d ("cpumask: Introduce cpumask_intersects_and") a8d0854a76a8 ("sched/cputime: Add kcpustat_field_total helper") be100c77178e ("sched: Add sched_ext hooks for proxy execution") 57c75e3ae38c ("sched: Add helper to block retained proxy donors") a49653d0abeb ("sched/core: Mark wakeups completed through ttwu_runnable()") 8f8c0417e973 ("sched/core: Dequeue waking proxy donors before reset") 313b652837d0 ("sched/core: Drop mutex locks before proxy rescheduling") 627ea30aca3b ("sched/wait: Clarify WF_SYNC wakeup semantics") d2e010082757 ("sched/eevdf: Handle more short slice waking cases") 4bf32ec3327d ("sched/eevdf: Align update_protect_slice to set_protect_slice") aae2a33ea662 ("sched/eevdf: Ensure that vprot will never go above a min slice") c9ce69fc43bd ("sched/fair: Randomize equally shallow slow-path candidates") abe440b3770f ("sched/fair: Drop idle recency from slow-path CPU selection") fbbc63fed0b0 ("sched/core: Remove redundant core_sched_seq") 819224e506bc ("sched/fair: Remove dead code on enqueue_task_fair()") c72945693b90 ("sched: Restart fair hrtick after same-task repicks") a9b3c7570564 ("sched/headers: Replace __ASSEMBLY__ with __ASSEMBLER__ in the <uapi/linux/sched.h> header") e81ee0630837 ("sched/fair: Reset NUMA fault locality after scan period update") ef9293b3b797 ("sched: dynamic: Fix preemption model strings") 879eaa76e608 ("sched: Remove unneeded function type cast in do_balance_callbacks()") f549101187c8 ("sched/deadline: check start_dl_timer expiry with ktime_before()") 2a672daa4b27 ("sched/feat: Use the new static key API for sched_feat") a5576ebce920 ("sched: Convert paravirt_steal to new static key APIs") 9650ce11f2e3 ("sched: dynamic: Simplify preempt model accessors") 5b9a28eeed37 ("sched: dynamic: Remove HAVE_PREEMPT_DYNAMIC_{CALL,KEY}") aa4178f63847 ("sched: dynamic: Simplify irqentry_exit_cond_resched()") b9d267b9d632 ("sched: dynamic: Simplify preempt_schedule{,_notrace}()") 88e0b3bb9930 ("sched: dynamic: Simplify {cond,might}_resched()") d3d16750693b ("sched: dynamic: Make PREEMPT_DYNAMIC depend on ARCH_HAS_PREEMPT_LAZY") 772d9ffbfd26 ("sched: Migrate whole chain in proxy_migrate_task()") 6b73a09e943f ("sched: Break out core of attach_tasks() helper into sched.h") 1f8805138593 ("sched: Switch rq->next_class in proxy_reset_donor()") 09351db90a28 ("sched/core: Don't proxy-exec unmatched cookie lock owners") 9be817f991e2 ("sched/core: Avoid migrating blocked_on tasks") 3dd95f077371 ("sched/core: Don't steal a proxy-exec donor") Signed-off-by: Ingo Molnar <mingo@kernel.org>
3 daysx86/cpufeatures: Add Page modification loggingNikunj A Dadhania
Page modification logging(PML) is a hardware feature designed to track guest modified memory pages. PML enables the hypervisor to identify which pages in a guest's memory have been changed since the last checkpoint or during live migration. The PML feature is advertised via CPUID leaf 0x8000000A ECX[4] bit. Acked-by: Borislav Petkov (AMD) <bp@alien8.de> Signed-off-by: Nikunj A Dadhania <nikunj@amd.com> Link: https://patch.msgid.link/20260907063906.1964557-7-nikunj@amd.com Signed-off-by: Sean Christopherson <seanjc@google.com>
3 daysx86/microcode/intel: Refresh old_microcode defines with the May 2026 releaseSohil Mehta
Update the minimum expected revisions of Intel microcode based on the microcode-20260512 (May 2026) release. Note, the three new entries are for INTEL_PANTHERLAKE_L steppings. Signed-off-by: Sohil Mehta <sohil.mehta@intel.com> Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com> Link: https://patch.msgid.link/20260928222607.3593973-1-sohil.mehta@intel.com
4 daysMerge tag 'v7.3-rc5' into driver-core-nextDanilo Krummrich
We need the driver-core fixes in here as well to build on top of. Signed-off-by: Danilo Krummrich <dakr@kernel.org>
4 daysx86/cpu: Constify struct x86_cpu_idChristophe JAILLET
'struct x86_cpu_id' is not modified in this compilation unit. Constifying this structure moves some data to a read-only section, so increases overall security. It is only used in cpu_has_old_microcode() which is an __init function. So, using __initconst is safe and will save about 6 kB of memory at runtime. On a x86_64, with allmodconfig: text data bss dec hex filename 55725 28774 512 85011 14c13 arch/x86/kernel/cpu/common.o.before 61464 22918 512 84894 14b9e arch/x86/kernel/cpu/common.o.after [ bp: Massage commit message. ] Signed-off-by: Christophe JAILLET <christophe.jaillet@wanadoo.fr> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Link: https://patch.msgid.link/f08d3a0e7aefe3bad66251927cbd03b8cbf9df73.1786311119.git.christophe.jaillet@wanadoo.fr
9 daysx86/mce: Fix hardware debug register corruption on task migrationMasami Hiramatsu (Google)
In exc_machine_check_user(), local_db_save() and local_db_restore() are invoked in the outer entry stubs (DEFINE_IDTENTRY_MCE_USER, DEFINE_FREDENTRY_MCE, and DEFINE_IDTENTRY_RAW), surrounding exc_machine_check_user(). However, exc_machine_check_user() calls irqentry_exit_to_user_mode(), which handles pending thread work and may schedule() if TIF_NEED_RESCHED is set. If the task migrates to another CPU during schedule(), local_db_restore() runs on the new CPU with the dr7 state saved from the old CPU. This corrupts the new CPU's DR7 hardware debug register and leaves the old CPU's DR7 disabled. In short, local_db_save() and local_db_restore() pair must be run on the same CPU. To fix this, move local_db_save() and local_db_restore() inside exc_machine_check_user() and exc_machine_check_kernel(). In exc_machine_check_user(), DR7 is saved and restored strictly around do_machine_check() to avoid schedule() during migration. In exc_machine_check_kernel(), local_db_save() is called at the entry point to prevent early memory accesses from triggering nested #DB exceptions, and restored on all exits. Fixes: cd840e424f27 ("x86/entry, mce: Disallow #DB during #MC") Assisted-by: LLM Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org> Cc: <stable@kernel.org> Link: https://patch.msgid.link/179005109564.388919.3937970081044095776.stgit@devnote2
10 daysumh, treewide: Explicitly include linux/umh.h where neededPetr Pavlu
The usermode helper declarations were previously provided by linux/kmod.h but commit c1f3fa2a4fde ("kmod: split off umh headers into its own file") moved them to linux/umh.h in 2017. Add explicit includes of linux/umh.h to files that use usermode helpers and remove linux/kmod.h where it is no longer needed. Acked-by: Alex Elder <elder@riscstar.com> # for greybus Signed-off-by: Petr Pavlu <petr.pavlu@suse.com>
13 daysx86/cpufeatures: Add X86_FEATURE_RMPOPT feature flagAshish Kalra
Add a flag indicating whether RMPOPT instruction is supported. RMPOPT is a new instruction that reduces the performance overhead of RMP checks for the hypervisor and non-SNP guests by allowing those checks to be skipped when 1-GB memory regions are known to contain no SEV-SNP guest memory. For more information on the RMPOPT instruction, see the AMD64 RMPOPT technical documentation. [ bp: Zap respective tools/ change. ] Suggested-by: Borislav Petkov (AMD) <bp@alien8.de> Signed-off-by: Ashish Kalra <ashish.kalra@amd.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Reviewed-by: Dave Hansen <dave.hansen@linux.intel.com> Reviewed-by: Ackerley Tng <ackerleytng@google.com> Reviewed-by: Tom Lendacky <thomas.lendacky@amd.com> Link: https://patch.msgid.link/39e9ee269a572c516a3f4e937bfe12d00697d5e6.1782841284.git.ashish.kalra@amd.com
2026-09-17x86/microcode/intel: Reject problematic loading on Granite Rapids systemsChang S. Bae
Microcode updates can usually jump revisions. However, there is an erratum on Granite Rapids systems. If they "jump over" revision 0x1000405, they result in an #MC. Avoid it. Signed-off-by: Chang S. Bae <chang.seok.bae@intel.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Reviewed-by: Dave Hansen <dave.hansen@linux.intel.com> Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260916225939.1144524-1-chang.seok.bae@intel.com
2026-09-17x86/cpu: Don't transiently clear the boot CPU's capabilitiesIhor Solodrai
On the boot CPU, identify_cpu() runs from arch_cpu_finalize_init(), with interrupts enabled and before alternatives are patched. So cpu_feature_enabled() still evaluates against boot_cpu_data. identify_cpu() rebuilds c->x86_capability from scratch: the reset zeroes the array and the CPUID rescan fills it in again. An interrupt delivered in that window finds X86_FEATURE_LA57 clear in boot_cpu_data, so pgtable_l5_enabled() is false and KASAN checks a 5-level address against the 4-level addressability limit. The result is a bogus "wild-memory-access" report, and under kasan_multi_shot a report storm that wedges the boot. The boot CPU has already been scanned by early_identify_cpu(), with interrupts disabled, and its capabilities cannot have changed since. Reset only the CPUs which have not been scanned yet. The window is as old as identify_cpu() rebuilding the capabilities. Commit 39b9552281ab ("x86/mm: Optimize boot-time paging mode switching cost") merely let KASAN notice it by making pgtable_l5_enabled() read the feature bit. So no Fixes: tag. Closes: https://lore.kernel.org/bpf/20260610175651.647515-1-ihor.solodrai@linux.dev/ Signed-off-by: Ihor Solodrai <ihor.solodrai@linux.dev> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Link: https://patch.msgid.link/20260916195203.1099646-6-ihor.solodrai@linux.dev
2026-09-17x86/cpu: Move 32-bit SEP setup into identify_cpu()Ihor Solodrai
identify_boot_cpu() and identify_secondary_cpu() both call enable_sep_cpu() under CONFIG_X86_32 immediately after identify_cpu(). Do it once and drop the ifdefs while at it. No functional changes. Signed-off-by: Ihor Solodrai <ihor.solodrai@linux.dev> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Reviewed-by: Nikolay Borisov <nik.borisov@suse.com> Link: https://patch.msgid.link/20260916195203.1099646-5-ihor.solodrai@linux.dev
2026-09-17x86/cpu: Inline generic_identify() into identify_cpu()Ihor Solodrai
generic_identify() has exactly one call site: at the top of identify_cpu(). Fold it into identify_cpu() so that a single function does the job for both the boot CPU and the secondary CPUs. While at it, fix up both copies of the Cyrix comment. No functional changes. Signed-off-by: Ihor Solodrai <ihor.solodrai@linux.dev> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Link: https://patch.msgid.link/20260916195203.1099646-4-ihor.solodrai@linux.dev
2026-09-17x86/cpu: Initialize boot CPU cpuinfo defaults earlyIhor Solodrai
early_identify_cpu() clears the capability array, the CPUID table and extended_cpuid_level, but the architectural defaults for the rest of struct cpuinfo_x86 are set only later, in identify_cpu(). Use the same defaults from the start, so that the boot CPU does not depend on a later reset to end up with the right ones. Signed-off-by: Ihor Solodrai <ihor.solodrai@linux.dev> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Link: https://patch.msgid.link/20260916195203.1099646-3-ihor.solodrai@linux.dev
2026-09-17x86/cpu: Factor init_cpu_info() out of identify_cpu()Ihor Solodrai
identify_cpu() unconditionally resets the struct cpuinfo_x86 fields to their default values and clears the capability arrays with memset() before rescanning the CPU to fill it in again. However the boot CPU capabilities have already been scanned by early_identify_cpu(), with interrupts disabled. Introduce init_cpu_info() helper in preparation for letting the callers decide whether the reset is needed. No functional changes. Signed-off-by: Ihor Solodrai <ihor.solodrai@linux.dev> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Link: https://patch.msgid.link/20260916195203.1099646-2-ihor.solodrai@linux.dev
2026-09-16x86/mce: Use __DEVICE_ATTR() macro to initialize dev_ext_attributeThomas Weißschuh
The upcoming constification of the device_show_int() and device_show_bool() signatures requires the users to handle the transition automatically. Switch to the __DEVICE_ATTR() macro which can do this. Signed-off-by: Thomas Weißschuh <linux@weissschuh.net> Link: https://patch.msgid.link/20260907-sysfs-const-attr-dev_ext_attr-v2-1-bf53afe57071@weissschuh.net Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2026-09-15x86/mtrr: Fix kernel-doc notation of amd_set_mtrr()Manuel Ebner
Add ':' to arguments in kernel-doc comment to fix: $ ./scripts/kernel-doc -none arch/x86/kernel/cpu/mtrr/amd.c Warning: arch/x86/kernel/cpu/mtrr/amd.c:63 function parameter 'reg' not described in 'amd_set_mtrr' Warning: arch/x86/kernel/cpu/mtrr/amd.c:63 function parameter 'base' not described in 'amd_set_mtrr' Warning: arch/x86/kernel/cpu/mtrr/amd.c:63 function parameter 'size' not described in 'amd_set_mtrr' Warning: arch/x86/kernel/cpu/mtrr/amd.c:63 function parameter 'type' not described in 'amd_set_mtrr' [ bp: Massage commit message. ] Signed-off-by: Manuel Ebner <manuelebnerli@mailbox.org> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Link: https://patch.msgid.link/20260915103136.171924-3-manuelebnerli@mailbox.org
2026-09-14x86/CPU/AMD: Fix Zen5 TLB sizes reportingBorislav Petkov (AMD)
Starting with Zen5, TLB sizes in CPUID_Fn80000006_E[AB]X are reported as multiples of 32. There's a CPUID bit which determines that: CPUID_Fn80000021_EAX [Extended Feature 2 EAX] (Core::X86::Cpuid::FeatureExt2Eax) ... 14: L2TlbSizeX32. Read-only. Reset: 1. Indicates that L2TLB sizes are encoded as multiples of 32. Update the places which report that information. With it, the numbers look correct now: -Last level iTLB entries: 4KB 64, 2MB 64, 4MB 32 -Last level dTLB entries: 4KB 128, 2MB 128, 4MB 64, 1GB 0 +Last level iTLB entries: 4KB 2048, 2MB 2048, 4MB 1024 +Last level dTLB entries: 4KB 4096, 2MB 4096, 4MB 2048, 1GB 0 /proc/cpuinfo -TLB size : 192 4K pages +TLB size : 6144 4K pages Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Link: https://lore.kernel.org/r/20260821022031.946311-1-bp@kernel.org
2026-09-14x86,fs/resctrl: Add resctrl_arch_preconvert_bw()Dave Martin
On MPAM systems the rounding behaviour of the MBA control would be improved if the rounding in the fs/resctrl code is removed but this is not the case for x86. To allow any rounding or conversion of the bandwidth value provided by the user to be specified by the arch code a new arch hook is required. Introduce resctrl_arch_preconvert_bw(), and add its x86 implementation. This is currently unused in resctrl but when plumbed in it will replace the call to roundup() in bw_validate(). Signed-off-by: Dave Martin <dave.martin@arm.com> Signed-off-by: Ben Horgan <ben.horgan@arm.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Reviewed-by: Reinette Chatre <reinette.chatre@intel.com> Reviewed-by: Gavin Shan <gshan@redhat.com> Link: https://patch.msgid.link/20260911163613.1131447-2-ben.horgan@arm.com
2026-09-08x86/MCE/AMD: Fix inverted interrupt enablement during storm handlingJasjeet Rangi
mce_amd_handle_storm() currently does the opposite of what storm handling needs: it enables thresholding interrupts when a storm is detected and disables them when the storm subsides. Flip the "on" function argument before passing it to threshold_restart_bank() as it should have been done. To clarify: "on" to mce_handle_storm() means, the storm is on now when "on" is true, and off when "on" is false. [ bp: Simplify. ] Fixes: 5c4663ed1eac ("x86/mce: Handle AMD threshold interrupt storms") Signed-off-by: Jasjeet Rangi <jrangi@purestorage.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Signed-off-by: Ingo Molnar <mingo@kernel.org> Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260812221514.598842-2-jrangi@purestorage.com
2026-09-02x86/sgx: Report RCU-Tasks quiescent state in EPC sanitization loopJun Miao
When the kernel boots from kexec, the EPC pages may have a stale state. The kernel sanitizes all EPC pages to reset them to a clean state before their first use in any enclave. The EPC size could be several GBs and resetting them could take a significant amount of time. Because of that, the kernel performs the reset in a loop through a kernel thread ksgxd() at early boot, and there's a cond_resched() after resetting each EPC page. This is fine in most cases, but becomes a problem when there's other kernel code waiting for an RCU-Tasks grace period but the cond_resched() in ksgxd() never triggers rescheduling. Because cond_resched() doesn't report a quiescent state when it doesn't trigger rescheduling, the thread that is waiting for an RCU-Tasks grace period will wait until all EPC pages are reset. For instance, BPF LSM subsystem can invoke synchronize_rcu_tasks() at kernel boot time. A VM with a large EPC assigned and BPF LSM enabled can take a long time to boot, with a call trace triggered: rcu_tasks_wait_gp: rcu_tasks grace period number 1 (since boot) is 130631 jiffies old. INFO: task systemd:1 blocked for more than 122 seconds. ... task:systemd state:D stack:0 pid:1 tpid:1 ppid:0 flags:0x00000002 Call Trace: ... schedule_timeout+0x157/0x170 wait_for_completion+0x88/0x150 __wait_rcu_gp+0x17e/0x190 synchronize_rcu_tasks_generic+0x64/0x60 ... synchronize_rcu_tasks+0x15/0x20 register_ftrace_direct+0x31f/0x350 ... bpf_trampoline_link_prog+0x33/0x60 bpf_tracing_prog_attach+0x3c5/0x5f0 Replace cond_resched() with cond_resched_tasks_rcu_qs() which explicitly reports quiescent state regardless of whether actual rescheduling is triggered. Resetting all EPC pages in ksgxd() isn't performance critical so the extra cost of cond_resched_tasks_rcu_qs() isn't a problem. Tests showed this reduced the VM kernel boot time from ~50s to ~700ms. Co-developed-by: Fan Du <fan.du@intel.com> Fixes: e7e0545299d8 ("x86/sgx: Initialize metadata for Enclave Page Cache (EPC) sections") Suggested-by: Kai Huang <kai.huang@intel.com> Signed-off-by: Fan Du <fan.du@intel.com> Signed-off-by: Jun Miao <jun.miao@intel.com> Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com> Reviewed-by: Kai Huang <kai.huang@intel.com> Reviewed-by: Jarkko Sakkinen <jarkko@kernel.org> Tested-by: Challvy Tee <challvy.tee@gmail.com> Link: https://github.com/systemd/systemd/issues/40423 Link: https://patch.msgid.link/20260902013653.690506-1-jun.miao@intel.com
2026-09-02sched: Convert paravirt_steal to new static key APIsHongyan Xia
paravirt_steal_rq_enabled and paravirt_steal_enabled use raw static_key APIs which are now deprecated. Use the new API instead. No functional change. Signed-off-by: Hongyan Xia <hongyan.xia@transsion.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Acked-by: Juergen Gross <jgross@suse.com> Link: https://patch.msgid.link/20260819081207.12150-1-hongyan.xia@transsion.com
2026-09-01x86/bugs: Adapt SRSO mitigation to Zen6Borislav Petkov (AMD)
Zen6 has BTB protection which isolates the different contexts (user/kernel, guest/host) from one another. This makes the SafeRET mitigation there unnecessary leaving the user/user and guest/guest attack vectors open, whose protection is handled by the Spectre v2 mitigation setting to do IBPB on a context switch. Detect that setting and report it with a new mitigation string. Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Link: https://patch.msgid.link/20260822013231.1109255-1-bp@kernel.org
2026-08-26Merge tag 'hyperv-next-signed-20260826' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/hyperv/linux Pull hyperv updates from Wei Liu: - Decrypt netvsc buffer on contiguous direct-map addresses (Kameron Carr) - Drop WS2012/2012R2 & Win8/8.1 Hyper-V support (Michael Kelley) - Use more meaningful errnos for hypercall status code (Hardik Garg) - Fix lost interrupts on CPU hot-unplug for Hyper-V PCI/MSI (Naman Jain) - Reserve more MSHV vectors for Linux root partition (Wei Liu) * tag 'hyperv-next-signed-20260826' of git://git.kernel.org/pub/scm/linux/kernel/git/hyperv/linux: clocksource: hyper-v: Remove support for stimer interrupts in message mode scsi: storvsc: Remove support for storvsc protocol of old Hyper-V hosts hv_netvsc: Remove GPADL teardown special case for old Hyper-V hosts hv_sock: Remove check for old Hyper-V hosts Drivers: hv: Remove support for WS2012/2012R2 & Win8/8.1 version of Hyper-V hv_netvsc: Allocate send/receive buffers using vmbus_alloc_buffer() Drivers: hv: vmbus: Add vmbus_alloc_buffer()/vmbus_free_buffer() for CoCo VMs Drivers: hv: vmbus: add vmbus_establish_gpadl_caller_decrypted() Drivers: hv: vmbus: Skip VMBus module cleanup for non-nested root partition x86/hyperv: reserve more vectors PCI: hv: Set irq_retrigger callback for the Hyper-V PCI MSI irqchip Drivers: hv: Use meaningful errnos for hypercall status codes
2026-08-24clocksource: hyper-v: Remove support for stimer interrupts in message modeMichael Kelley
In Hyper-V versions prior to WS2016/Win10, Hyper-V synthetic timers interrupt the guest by delivering a message that is initially handled by the Linux VMBus driver. Starting with WS2016/Win10, Hyper-V can deliver stimer interrupts directly to an assigned interrupt vector without involving the VMBus driver. This is called "Direct Mode". With the overall removal of Linux support for running on Hyper-V hosts earlier than WS2016 and Windows 10, it's no longer necessary to support the legacy message-based delivery. Remove that delivery mechanism and always use Direct Mode. If for some reason, the Hyper-V host does not enumerate Direct Mode, output an error message but continue to run using the LAPIC timer instead of an stimer. With these changes, the VMBus driver no longer calls the stimer interrupt service routine. This removal has a broader benefit in unblocking the disentangling of VMBus code and stimer code, as they should be independent of each other. The final disentangling will come as a follow-on patch set. Signed-off-by: Michael Kelley <mhklinux@outlook.com> Signed-off-by: Wei Liu <wei.liu@kernel.org>
2026-08-24x86/hyperv: reserve more vectorsWei Liu
Microsoft Hypervisor delivers three vectors to the NT HAL running in the root partition and refuses to map a device interrupt to any of them when interrupt remapping is not available in the system. As of writing, the nested MSHV setup has no interrupt remapping capability. The three vectors are: HAL_NT_APC_VECTOR 0x1F HAL_NT_DPC_VECTOR 0x2F HAL_NT_CLOCK_IPI_VECTOR 0xD2 0x1F is below FIRST_EXTERNAL_VECTOR so the vector allocator never hands it out, but 0x2F and 0xD2 are both inside the allocatable range and are handed out once enough vectors are in use. Mapping such an interrupt then fails with HV_STATUS_INVALID_PARAMETER, and the interrupt is never delivered. Reserve all three next to the hypervisor debug vectors that are already kept out of the allocator's hands. Reviewed-by: Michael Kelley <mhklinux@outlook.com> Signed-off-by: Wei Liu <wei.liu@kernel.org>
2026-08-20Merge tag 'mm-stable-2026-08-18-18-39' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm Pull MM updates from Andrew Morton: - "mm: drop "sub" prefix from various places" (Dev Jain) page->folio conversion and a naming cleanup - "mm/kasan: remove redundant initialization for kasan_flag_write_only" (Igor Putko) KASAN cleanup work - "mm/filemap: reduce unnecessary xarray lookups" (Chi Zhiling) Small speedup in the pagecaache read code - "mm/percpu: Fix possible NOFS/NOIO reclaim recursion" (Kaitao Cheng) Improve the vmalloc code - mainly the avoidance of GFP_KERNEL allocations when the caller asked for GFP_NOFS or GFP_NOIO - "mm/kmemleak: avoid soft lockup when scanning task stacks" (Breno Leitao) Avoid a soft lockup watchdog trigger from the kmemleak scanning code in extreme situations - "mm/page_owner: misc cleanups" (Ye Liu) Cleanups to the page_owner code. For some reason lots of people have been working on the page_owner code this cycle. - "mm: convert to walk_page_range_vma() to eliminate find_vma()" (Kefeng Wang) Simplify and accelerate the page walking library function - "mm/migrate: preparatory cleanups for batch copy and offload" (Shivank Garg) Cleanups in the migration code - "mm/page_owner: add per-fd filter infrastructure for print_mode and NUMA filtering" (Zhen Ni) Per-fd filtering to page_owner in order to reduce the sometimes vast amount of output it can produce - "mm: Refactor bootmem gigantic hugepage allocation" (Muchun Song) Fixes and preparatory cleanups around bootmem HugeTLB handling, sparse initialization ordering, and related vmemmap setup - "mm/zsmalloc: reduce lock contention in zs_free()" (Wenchao Hao) Reduce lock contention in zs_free(), which dominates the unmap path under memory pressure on Android (LMK kills) and on x86 servers running zswap-heavy workloads. Up to 1.83x improvement in microbenchmarking. - "move alloc_tag.c file under mm/" (Suren Baghdasaryan) - "samples/damon: handle damon_{start,stop}() failures" (SJ Park) Fix improper handling of damon_start(), damon_stop(), and damon_call() failures across DAMON sample modules to prevent potential memory leaks, operation disruptions and use-after-free bugs - "mm/damon/sysfs: kobject_del() directories that users can create/remove" (SJ Park) Fix delayed sysfs directory removal under DEBUG_KOBJECT_RELEASE causeing creation failures due to duplicate directory names by adding missing kobject_del() calls before creating new directories - "mm: cleanup clear_not_present_full_ptes()" (David Hildenbrand) Clean up the core pte handling code - "selftests/damon: misc fixes for test bugs" (Kunwu Chan) Fix several bugs in the DAMON selftests - "selftests/damon: fix memcg_path staging handling" (Cheng Nie) Fix a bug in _damon_sysfs.py for damos_filter memcg_path setup, and add a test case for it in sysfs.py. - "selftests/damon: test kdamond refresh_ms" (Ruslan Valiyev) Selftest coverage for DAMON's refresh_ms sysfs feature by updating the test control module and verifying that scheme stats update automatically without manual intervention - "mm/damon: five misc fixups" (Akinobu Mita) Miscellaneous DAMON fixups. - "mm/damon/core: detect internal variation above max_nr_regions/2" (Jiayuan Chen) Fix DAMON's region splitting behavior when region counts exceed half the maximum budget by dynamically scaling down the split fraction as the limit approaches, preventing large regions from staying un-split, and add corresponding KUnit test coverage - "mm: preparatory patches for PMD level swap entries" (Usama Arif) Refactor and clean up PMD softleaf helpers, call sites, and architecture flags to lay the groundwork for a follow-up series that introduces PMD page table swap entries - "mm/damon: update, optimize, and clean up doc, tests, and code" (SJ Park) Update DAMON design and ABI documentation, expands unit and selftest coverage, optimize damon_commit_target_regions(), and clean up recently added sysfs interface code for better readability - "mm/vmpressure: reduce CPU, memory and code overhead on cgroup v2" (Usama Arif) Optimize vmpressure() by skipping unnecessary work on cgroup v2 for userspace event notifications and refactor v1-only eventfd handling into mm/memcontrol-v1.c to reduce memory overhead and code complexity - "selftests/mm: refactor pkey helpers and fix mmap error handling" (Hongfu Li) Refactor pkeys shared tracing and assertion helpers into a common file, unify protection key selftests to use consistent diagnostic logging and assertions, and enforce standardized MAP_FAILED return checks for mmap() calls across the tests - "mm/damon: optimize out nr_accesses_bp" (SJ Park) Replace the error-prone, continuously updated nr_accesses_bp field in damon_region with an on-demand moving sum function, reducing structure memory overhead and avoiding state corruption bugs - "Open HugeTLB allocation routine for more generic use" (Ackerley Tng) Decouple HugeTLB folio allocation from VMA dependencies by introducing hugetlb_alloc_folio(), enabling subsystems like guest_memfd to allocate HugeTLB folios without standard VMA reservations or pseudo-VMAs - "mm/damon: provide pseudo moving sum probe_hits" (SJ Park) Integrate DAMON's probe_hits attribute counter into the pseudo moving sum infrastructure, enabling real-time, online monitoring without waiting for full aggregation intervals - "mm: Some cleanups for page allocator APIs" (Brendan Jackman) Simplify and refactor the page allocator entry points and flags by unifying allocation paths, adding internal alloc_flags arguments, and eliminating redundant __ prefixed alloc_pages variants. - "Fix incorrect access of hugetlb pte entries" (Dev Jain) Enforce the consistent use of huge_ptep_get() instead of ptep_get() for HugeTLB entries and fixes an unaligned address issue in arm64's huge_ptep_get() implementation - "mm/damon: validate all parameters in the core" (SJ Park) Consolidate parameter validation into the DAMON core specifically within damon_start() and damon_commit_ctx() to centralize error checking, eliminate caller-side redundant checks and to improve maintenance efficiency - "tools/mm/page_owner_sort: fix filtering and cleanup issues" (Yichong Chen) Rename is_need() to filter_record() for clearer return semantics, fix per-record allocation memory leaks and bound output copies in search_pattern() to address an existing buffer issue - "memcg: bail out reclaim when memcg is dying" (Jiayuan Chen) Mitigate a system-wide stall which occurs when a cgroup is removed while one of its memory control files is doing synchronous reclaim - "mm/memory-failure: add panic option for unrecoverable pages" (Breno Leitao) Introduce an opt-in vm.panic_on_unrecoverable_memory_failure sysctl that immediately panics the kernel on unrecoverable memory errors in kernel-owned pages to preserve error context and prevent delayed, silent data corruption - "mm/damon: refactor damon_{start,stop,commit}() for simple error handling" (SJ Park) Refactor the DAMON core API functions to guarantee that all contexts are fully stopped when damon_start(), damon_stop(), or damon_commit() fail, eliminating the need for complex and error-prone caller-side cleanup code - "Keep tail page private zero at free and folio split" (Zi Yan) Add checks to ensure tail_page->private is zero when freeing compound or high-order pages and when promoting tail pages during large folio splits. By validating these fields at free and split time, it allows the removal of redundant private field clearing inside prep_compound_tail() - "mm: drop redundant lru_add_drain in anon folio reuse paths" (Barry Song) Eliminate redundant lru_add_drain() calls in wp_can_reuse_anon_folio() and do_swap_page() to reduce LRU lock contention and system overhead By validating folio refcounts against the LRU cache before draining and removing unnecessary drains in the swap path, it achieves up to a 30.5% reduction in drain calls during heavy swap workloads - "mm: clean up folio LRU and swap declarations" (Jianyue Wu) Reorganize folio LRU and swap code by relocating page-cluster state to mm/swap_state.c, renaming mm/swap.c to mm/folio.c, and moving MM-internal reclaim declarations into mm/internal.h. - "userfaultfd: working set tracking for VM guest memory" (Kiryl Shutsemau) Add userfaultfd support for tracking the working set of VM guest memory, so a VMM can identify hot pages and reclaim cold ones to tiered or remote storage - "mm: remove CONFIG_HAVE_BOOTMEM_INFO_NODE (Part 2)" (David Hildenbrand) Remove the remaining pieces of CONFIG_HAVE_BOOTMEM_INFO_NODE, performing some smaller cleanups around freeing of reserved vmemmap pages on the way. - "mm/damon: update probe hits for runtime parameter commits" (SJ Park) Ensure that DAMON's probe_hits attribute counter is properly updated when monitoring intervals are changed at runtime, matching the behavior of nr_accesses. To achieve this, it refactors and renames existing helper functions for shared use, applies the updates to probe_hits, and handles edge cases in damon_probe_hits_mvsum() to maintain measurement accuracy. - "KSM: performance optimizations for rmap_walk_ksm" (xu xin) Resolve a severe KSM reverse-mapping performance bottleneck where thousands of split VMAs sharing a single anon_vma cause extended lock contention. By adding an interval-filtering check during the rmap walk, it reduces worst-case anon_vma lock hold times from over 500ms down to under 2ms, preventing application freezes and latency spikes under memory pressure. - "mm: split a couple of headers from internal.h" (Mike Rapoport) Split declarations related to mm_init, memblock, vmalloc and sparse into new headers - "KSM: use linear_page_index in collect_procs_ksm()" (xu xin) Apply the interval tree optimization from rmap_walk_ksm() to collect_procs_ksm() to avoid iterating over non-matching VMAs during KSM memory error handling. It hoists loop-invariant address initialization and restricts the anon_vma_interval_tree_foreach walk to a targeted page offset range, reducing redundant checks and improving lookup efficiency. - "selftests/mm: avoid false failures in hugetlb and KSM tests" (Sayali Patil) Fix issues in the hugetlb and KSM MM selftest categories that can report failures when the prerequisites for the tests are not satisfied - "mm/damon: introduce data attributes only monitoring" (SJ Park) Introduce attribute-weighted region management in DAMON, allowing users to prioritize specific data attributes (such as page sizes or cgroups) over or instead of access monitoring. By assigning weights to attribute probes, DAMON can completely disable access tracking and adjust monitoring regions based on weighted probe-hit counters to optimize monitoring quality for attribute-focused workloads. - "mm/hmm: Add mmap lock-drop support for userfaultfd-backed mappings" (Stanislav Kinsburskii) Extend hmm_range_fault() to support userfaultfd-backed regions by allowing the mmap lock to be dropped during fault handling via a new hmm_range_fault_locked() helper. By accepting a locked pointer and signaling retry status when lock release occurs, it enables page fault resolution in userfaultfd regions while preserving backward compatibility for existing callers. - "mm: make VMA page offset handling more consistent" (Lorenzo Stoakes) Clean up and standardize how vma->vm_pgoff is accessed and manipulated across file-backed and anonymous mappings in the kernel It introduces dedicated helper functions such as vma_start_pgoff(), vma_end_pgoff(), vma_set_pgoff() and linear_page_delta() while renaming rmap interval tree helpers to better reflect their functionality. These changes establish a cleaner foundation for future work that will unify virtual page offset indexing for all anonymous and CoW'd folios. - "mm: handle device-private PMDs in walk callbacks" (Usama Arif) Address kernel panics and state corruption caused by MM walk callbacks reaching non-present device-private PMD swap entries created during HMM migrations It ensures that functions which acquire pmd_trans_huge_lock() properly recognize device-private PMDs instead of assuming a present THP or a standard migration entry. - "mm/rmap: Refactor try_to_unmap_one" (Dev Jain) Refactor try_to_unmap_one by modularizing Hugetlb, anonymous-lazyfree, and anonymous-swapbacked logic into dedicated functions, laying the structural groundwork for batched anonymous large folio unmapping. - "Docs/ABI/damon: sysfs ABI document fixes and additions" (Song Hu) Fix typos and fills in missing entries in the DAMON sysfs ABI document - "dax/kmem: atomic whole-device hotplug via sysfs" (Gregory Price) Introduce an atomic sysfs state attribute and supporting DAX/MM infrastructure to prevent userland races when offlining and removing entire memory regions By adding an unplugged state alongside standard online modes, it enables whole-device atomic hotplug control while preserving backward compatibility. - "mm: convert more vm_flags_t users to vma_flags_t" (Lorenzo Stoakes) Continue transitioning the kernel from the deprecated vm_flags_t type to vma_flags_t across core memory management infrastructure. It replaces legacy type usage in core functions such as do_mmap(), unmapped area allocation, mm->def_vma_flags, and VMA operations like mlock, mprotect, and mremap. - "Two small patches to clean up mm/mm_slot.h" (xu xin) Refactor mm_slot.h by introducing mm_slot_remove() to unify duplicate slot deletion sequences in khugepaged and KSM. It also adds code documentation explaining why mm_slot_lookup and mm_slot_insert must remain as preprocessor macros rather than static inline functions. - "mm/damon/core: hide core-private struct fields" (SJ Park) Clean up DAMON core structures by consistently marking internal-only fields with private: comment tags to prevent improper direct access from outer layers. It enforces encapsulation across core structures including damon_region, damon_target, and damon_ctx and updates DAMON_SYSFS to interact through approved access APIs instead of exposing raw struct members. - "mm/damon: unurgent fixes for infinite loop, NULL de-ref and races" (SJ Park) Address potential infinite loops, NULL dereferences, and race conditions identified in DAMON It fixes an infinite loop triggered by extreme user configurations, a NULL pointer dereference within unit tests and minor monitoring accuracy degradation caused by subtle runtime races. - "mm/page_alloc: fixes for free_pages_nolock() on RT/UP" (Brendan Jackman) Fix an NMI safety flaw in __free_frozen_pages() where freeing pages on non-SMP or PREEMPT_RT kernels can bypass can_spin_trylock() checks via non-PCP or isolated migration paths. It also resolves potential kernel crashes and privilege escalation risks triggered when BPF tracing runs in NMI context alongside memory hotplug or large allocation frees. - "mm/page_alloc: couple of followups for recent cleanups" (Brendan Jackman) Clean up and update page allocator nomenclature, documentation, and debug assertions. It aligns internal FPI_ flags with the public "nolock" naming convention, removes outdated internal implementation details from high-level page allocator comments, and eliminates obsolete VM_BUG_ON() assertions in allocation paths. - "mm/mseal: further cleanups" (Lorenzo Stoakes) Refactor and simplify the mseal implementation by clarifying API boundaries and removing unnecessary code complexity. It replaces generic do_mseal() usage outside the syscall with a dedicated mseal_mmap_page_zero() helper for MMAP_PAGE_ZERO, eliminates mm_struct parameters to enforce that sealing applies only to current->mm, and streamlines overall logic and comments with no functional changes intended. - "mm/vmscan: fix swappiness=max and clean up per-node proactive reclaim" (Ridong Chen) Resolve reclaim behavior bugs and clean up function parameters across memory reclaim paths It fixes swappiness=max in both standard reclaim and MGLRU so unswappable anonymous memory no longer falls back to evicting page cache, ensures reclaim_store() returns accurate error codes instead of collapsing all failures into -EAGAIN, and removes the obsolete gfp_mask parameter from __node_reclaim(). - "mm: mincore: misc cleanups" (Kefeng Wang) Clean up and simplifies the mincore code. Most importantly, it removes the historical special behavior that always reports VM_PFNMAP pages as non-resident. - "mm/huge_memory: drop dead split helper variants" (Kiryl Shutsemau) Two trivial cleanups in the folio split API - "mm/damon: fix uninitialized DAMOS field and kunit exec expectation bugs" (SJ Park) Resolve minor operational and testing bugs in DAMON identified by Sashiko. It initializes the damos->last_applied field to prevent occasional efficiency degradation and fixes invalid memory accesses in DAMON KUnit tests during test failure handling. - "cleanup for stable_page_flags()" (Jinjiang Tu) Clean up and refactor stable_page_flags() used by /proc/kpageflags without altering functionality. It uses BIT_ULL() to prevent shift-overflow warnings on 64-bit flag bits, converts folio-specific flag checks to standard folio_test_*() helpers, and removes redundant CONFIG_PAGE_IDLE_FLAG handling. - "Batch unmap of uffd-wp file folios" (Dev Jain) Extend batched folio unmapping support to file folios within userfaultfd write-protect (uffd-wp) VMAs by adding batching capabilities to pte_install_uffd_wp_if_needed(). This removes special-case restrictions on uffd-wp VMAs in try_to_unmap_one(), significantly simplifying the function's control flow and complexity. - "mm/early_ioremap: clarify and clean up early_ioremap_reset()" (Sang-Heon Jeon) Clarify and clean up the architecture-specific usage of __late_set_fixmap() and __late_clear_fixmap() after early_ioremap_reset() It adds explicit documentation regarding when early_ioremap_reset() must be called and removes redundant macro definitions and reset calls in the RISC-V and ARM64 architectures. - "mm: fix reclaim storms in defrag_mode" (Johannes Weiner) Address severe performance regressions, swap storms, and spurious OOMs caused by vm.defrag_mode=1 under high memory pressure in Meta production It updates the page allocator slowpath so non-movable allocation requests actively trigger direct reclaim and direct compaction at pageblock_order scale, allowing them to claim whole pageblocks rather than spinning unproductively. - "zram: lockmap tweaks" (Sebastian Siewior) Optimize and fix lockdep tracking for zram devices by consolidating per-entry lockmaps and isolate lock classes across multiple instances This reduces memory overhead by replacing per-entry lockdep_map instances with a single map per struct zram, and assigns a dynamic lock_class_key to each instance to prevent false deadlock reports when different zram devices are backed by distinct filesystems. * tag 'mm-stable-2026-08-18-18-39' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm: (501 commits) selftests/mm: thuge-gen: fix test_shmget() for PAGE_SIZE check selftests/mm: unpoison pages in memory-failure teardown mm/shmem: downgrade final i_blocks check in shmem_evict_inode() to pr_warn() mm/khugepaged: replace mutex_lock/mutex_unlock usage with guard macro mm/zsmalloc: fix release order of locks in zs_page_migrate() Documentation: zram: remove sections numbering ksm: stop iterating VMAs when ksm_test_exit returns true mm: fold userfaultfd_rwp() to false without CONFIG_ARCH_HAS_PTE_PROTNONE mm/migrate: report RCU-tasks quiescent states in migrate_pages_batch() zram: use a custom key for each zram object zram: move lockmap to be per-zram instead per table selftests/mm: fix gup_longterm EINVAL error message mm: page_alloc: fix non-movable reclaim storm in defrag_mode mm: page_alloc: move capture_control to the page allocator mm: compaction: support non-movable compaction for pageblock requests mm: page_alloc: __GFP_FS lockdep annotation for direct compaction hugetlb: evaluate subpool free state while locked mm/damon: remove trailing semicolons after function definitions mm/damon/ops-common: prevent migration fallback to non-target nodes mm/damon: update outdated comment about DAMOS filter handling ...
2026-08-20Merge tag 'sysctl-7.03-rc1' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/sysctl/sysctl Pull sysctl updates from Joel Granados: - Fix kernel-doc warnings by adjusting in file documentation - Consolidate do_proc_* function into do_proc_vec Consolidate three slightly different implementations of applying a converter on all elements of a vector. Fixes to this function now propagate to the three types. - Replace CONFIG_PROC_SYSCTL with CONFIG_SYSCTL (they were the same) and restrict cad_pid modifications to global root (GLOBAL_ROOT_UID) * tag 'sysctl-7.03-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/sysctl/sysctl: sysctl: remove CONFIG_PROC_SYSCTL, it just mirrors CONFIG_SYSCTL sysctl: move the "cad_pid" entry from pid_table[] to kern_reboot_table[] sysctl: repair some kernel-doc comments sysctl: add Returns: kernel-doc for all functions sysctl: Update API function documentation sysctl: Rename proc_doulongvec_minmax_conv to proc_doulongvec_conv sysctl: Group proc_handler declarations and document sysctl: Replace do_proc_do{int,ulong,uint}vec with do_proc_vec sysctl: Add negp parameter to douintvec converter functions sysctl: Move default converter assignment out of do_proc_dointvec
2026-08-18Merge tag 'x86_cpu_for_v7.3_rc1' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull x86 cpuid updates from Borislav Petkov: - Get rid of static_cpu_has() - one less API to care about testing CPU features - Unify the handling of CPU core types (performance, efficient, etc) by mapping the vendor-specific types to Linux ones - Continuation of the work of Ahmed Darwish to centralize CPUID leaf representation * tag 'x86_cpu_for_v7.3_rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: x86/CPU: Rename struct cpuid_read_output to struct cpuid_output x86/cpu/scattered: Sort it properly x86/cpu: Use parsed CPUID(0x1) x86/lib: Add CPUID(0x1) family and model calculation x86/cpu: Use parsed CPUID(0x0) x86/cpu/transmeta: Rescan CPUID(0x1) after modifying capabilities x86/topology: Add TOPO_CPU_TYPE_LOW_POWER x86/topology: Name the AMD core-type values x86/topo: Map vendor CPU types to generic Linux such types x86/bugs: Don't use cpu-type matching in cpu_vuln_blacklist x86/cpu: Hide and rename static_cpu_has()
2026-08-18Merge tag 'x86_cleanups_for_v7.3_rc1' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull x86 cleanups from Borislav Petkov: - The usual pile of smallish cleanups and fixlets all over the place * tag 'x86_cleanups_for_v7.3_rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: x86/cpu: Remove unnecessary __maybe_unused annotations x86/msr: Document the I/O-like write semantics in the msr driver x86/apic: Ensure ICR register write value is handled as 32 bits x86/boot/compressed/head_64.S: Clean up SEV-related comments Documentation/arch/x86/amd-memory-encryption.rst: Fix typo x86/platform/quark: Fix kernel-doc warnings in imr.c x86/ras: Move contents from arch/x86/ras/Kconfig into drivers/ras/Kconfig x86/cpu: Move intel_get_platform_id() to cpu/intel.c x86/mm: Fix typo in comment x86/fpu: Fix kernel-doc formatting above fpu_enable_guest_xfd_features() x86/cfi: Use symmetric SYM_START and SYM_END in __CFI_TYPE() x86/cfi: Add __init_or_module annotations for fineibt
2026-08-18Merge tag 'x86_cache_for_v7.3_rc1' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull x86 resource control updates from Borislav Petkov: - How refreshing: no new features but a whole pile of fixes to more or less serious issues reported by Sashiko along with miscellaneous cleanups all over the place. All except one by Reinette Chatre, the one by Tony Luck. * tag 'x86_cache_for_v7.3_rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: fs/resctrl: Inform user space when status buffer overflowed fs/resctrl: Communicate resource group deleted error via last_cmd_status fs/resctrl: Add last_cmd_status support for writes to max_threshold_occupancy fs/resctrl: Change last_cmd_status custom during input parsing fs/resctrl: Use accurate and symmetric exit flows fs/resctrl: Pass error reading event through to user space fs/resctrl: Use accurate type for rdt_resource::rid fs/resctrl: Change pattern used to track number of entries in enum resctrl_conf_type x86/resctrl: Protect against bad shift fs/resctrl: Use correct format specifier for printing error pointers fs/resctrl: Fix UAF from worker threads when domains are removed x86/resctrl: Ensure domain fully initialized before placed on RCU list fs/resctrl: Prevent deadlock and use-after-free in info file handlers fs/resctrl: Prevent use-after-free in rdtgroup_kn_put() fs/resctrl: Fix deadlock on errors during mount fs/resctrl: Move functions to avoid forward references in subsequent fixes x86,fs/resctrl: Document safe RCU list traversal
2026-08-18Merge tag 'x86-msr-2026-08-17' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull x86 MSR updates from Ingo Molnar: - Streamline the x86 MSR handling APIs along the 64-bit variants, simplifying the interfaces. Removal of the old APIs is planned for the next cycle, to reduce churn & integration pain (Juergen Gross) * tag 'x86-msr-2026-08-17' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: (21 commits) x86/mce: Work around build warning after MSR-interface switch cpufreq: Stop using 32-bit MSR interfaces x86/featctl: Stop using 32-bit MSR interfaces KVM/x86: Stop using 32-bit MSR interfaces x86/mtrr: Stop using 32-bit MSR interfaces acpi: Stop using 32-bit MSR interfaces powercap: Stop using 32-bit MSR interfaces thermal/intel: Stop using 32-bit MSR interfaces x86/olpc: Stop using 32-bit MSR interfaces x86/hyperv: Stop using 32-bit MSR interfaces hwmon: Stop using 32-bit MSR interfaces EDAC: Stop using 32-bit MSR interfaces x86/cpu: Stop using 32-bit MSR interfaces x86/apic: Stop using 32-bit MSR interfaces x86/resctrl: Stop using 32-bit MSR interfaces x86/tsc: Stop using 32-bit MSR interfaces x86/amd: Stop using 32-bit MSR interfaces x86/pci: Stop using 32-bit MSR interfaces x86/hygon: Stop using 32-bit MSR interfaces x86/mce: Stop using 32-bit MSR interfaces ...
2026-08-18Merge tag 'locking-core-2026-08-17' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull locking updates from Ingo Molnar: "Futexes: - Use runtime constants for futex_hash computation (K Prateek Nayak, Peter Zijlstra) - Optimise the size check get_futex_key() (Sebastian Andrzej Siewior) - Avoid private hash use-after-free on final put (Felix Hoffmann) - Tell kmemleak we're not leaking __futex_queues (Peter Zijlstra) Rust integration updates: - Implement refcounted interrupt disable and SpinLockIrq for Rust (Boqun Feng, Heiko Carstens, Joel Fernandes, Lyude Paul) - Rust sync: add helpers for mb, dma_mb and friends; add generic memory barriers and use LKMM atomics instead of Rust atomics in the revocable code (Gary Guo) - Add abstraction and integrate synchronize_rcu() (Philipp Stanner) Lock debugging: - Add qspinlock contended_release tracepoint (Dmitry Ilvokhin, Peter Zijlstra) - Enable the printing of held locks of remote running tasks and print task CPU (Ingo Molnar) - percpu-rwsem: Annotate intentional data race in readers_active_check() (Sun Shaojie) Misc fixes and updates by Boqun Feng, Peter Zijlstra, Fangrui Song, Naveen Kumar Chaudhary and Thomas Huth" * tag 'locking-core-2026-08-17' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: (44 commits) rust: sync: Introduce SpinLockIrq::lock_with() and friends rust: sync: Add SpinLockIrq rust: sync: Use super::* in spinlock.rs rust: helper: Add spin_{un,}lock_irq_{enable,disable}() helpers rust: Introduce interrupt module s390/preempt: Enable HAS_SEPARATE_PREEMPT_RESCHED_BITS arm64: sched/preempt: Enable HAS_SEPARATE_PREEMPT_RESCHED_BITS preempt: Introduce HAS_SEPARATE_PREEMPT_RESCHED_BITS sched: Avoid signed comparison of preempt_count() in __cant_migrate() sched: Remove the unused preempt_offset parameter of __cant_sleep() locking: Switch to _irq_{disable,enable}() variants in cleanup guards irq: Add KUnit test for refcounted interrupt enable/disable irq,spin_lock: Add counted interrupt disabling/enabling openrisc: Include <linux/cpumask.h> in smp.h preempt: Introduce __preempt_count_{sub,add}_return() preempt: Introduce HARDIRQ_DISABLE_BITS preempt: Track NMI nesting to separate per-CPU counter futex: Tell kmemleak we're not leaking __futex_queues x86/paravirt: Trace contended_release on unlock tracing/lock: Use TRACE_EVENT_FN() for contended_release ...
2026-08-18Merge tag 'for-linus-7.3-rc1-tag' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/xen/tip Pull xen updates from Juergen Gross: - Small cleanups for the Xen ACPI pad driver and the gnttab driver - Fix an issue with Xen PV device initialization seen with QubesOS tests - Fixes for the Xen balloon driver and the xenbus driver - Simplify Xen related kernel configuration * tag 'for-linus-7.3-rc1-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/xen/tip: xenbus: Unregister reboot notifier on init failure x86/xen: Drop CONFIG_XEN_PVHVM_SMP xen: Drop CONFIG_XEN_AUTO_XLATE xen: Drop CONFIG_XEN_PVHVM x86/xen: Remove redundant config dependency on X86_LOCAL_APIC x86/xen: fix init of balloon stats again xen/xenbus: check otherend_id only after it has been initialized xen/xenbus: log more information when device state got reset Xen/gnttab: adjust two uses of sizeof() ACPI: PAD: xen: Stop setting acpi_device_name/class()
2026-08-10x86/CPU: Add a tlbi= cmdline switchRik van Riel
With the recently found INVLPGB / TLBSYNC issue, there has been some interest in disabling INVLPGB-based TLB flushing, in order to rule out that CPU issue as a cause of userspace crashes. Add a kernel command line option to control the TLB flushing behavior. If the need arises, we will add a "tlbi=broadcast" for the case when TLB invalidation broadcasts need to be explicitly selected, but this is not needed now yet. [ bp: Rewrite commit message, move to cpu/common.c, add documentation. ] Fixes: 767ae437a32d ("x86/mm: Add INVLPGB feature and Kconfig entry") Suggested-by: Borislav Petkov <bp@alien8.de> Signed-off-by: Rik van Riel <riel@surriel.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Cc: <stable@kernel.org> Link: https://patch.msgid.link/20260729204341.3eb0b5ea@fangorn
2026-08-10preempt: Introduce HAS_SEPARATE_PREEMPT_RESCHED_BITSBoqun Feng
With the changes that enable preempt count to track IRQ disabling nesting, we don't have enough bits in 32-bit preempt count implementation, as a result we move NMI nesting bits out of the 32-bit preempt count. However on the architectures that can support 64-bit preempt count implementation, we can keep the NMI nesting bits in the 32-bit preempt count and avoid maintaining NMI nesting bits outside of the same cache line. Therefore HAS_SEPARATE_PREEMPT_RESCHED_BITS is introduced to allow architectures to select this. Note that under this Kconfig, preempt count is maintained in a 64-bit word however preempt_count() still remains as an int because all the effective bits still fit in (previously we mask out NEED_RESCHED bit in preempt_count()). This should make no functional changes for existing preempt_count() users. Enable this for x86_64 along with the introduction of the Kconfig. [boqun: Undo the __preempt_count_{add,sub}() optimization in 32-bit preempt count since it may introduce {over,under}flow] Originally-by: Peter Zijlstra <peterz@infradead.org> Signed-off-by: Boqun Feng <boqun@kernel.org> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://patch.msgid.link/20260804161447.84806-11-boqun@kernel.org
2026-08-08Merge tag 'x86-urgent-2026-08-08' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull x86 fix from Ingo Molnar: - Fix MCE CMCI discovery initialization ordering bug (Breno Leitao) * tag 'x86-urgent-2026-08-08' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: x86/mce: Set up the polling timer before CMCI discovery
2026-08-06xen: Drop CONFIG_XEN_PVHVMJuergen Gross
On x86 CONFIG_XEN_PVHVM is now a synonym of CONFIG_XEN. In Xen specific x86 code it can be just dropped, in non-Xen specific x86 code it can be replaced with CONFIG_XEN. In architecture independent code it is used only where CONFIG_XEN is defined, so it can be replaced with CONFIG_X86 there. Reviewed-by: Stefano Stabellini <sstabellini@kernel.org> Signed-off-by: Juergen Gross <jgross@suse.com> Message-ID: <20260805082137.1214967-3-jgross@suse.com>
2026-08-05Merge tag 'x86_bugs_saferet' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip - Add a mitigation for the attack vector of interrupting the saferet sequence used in the SRSO mitigation and still poisoning the RSB. Do that by emulating the saferet sequence and thus avoiding executing a RET instruction. * tag 'x86_bugs_saferet' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: x86/bugs: Make Safe-RET robust against interrupt injection