summaryrefslogtreecommitdiff
path: root/arch/x86/kernel
AgeCommit message (Collapse)Author
11 hoursMerge branch 'driver-core-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/driver-core/driver-core.git
11 hoursMerge branch 'next' of https://github.com/kvm-x86/linux.gitMark Brown
11 hoursMerge branch 'non-rcu/next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/paulmck/linux-rcu.git
11 hoursMerge branch 'next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/liveupdate/linux.git # Conflicts: # mm/memblock.c # mm/mm_init.c
11 hoursMerge branch 'master' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git # Conflicts: # Documentation/scheduler/index.rst # arch/arm64/configs/defconfig
11 hoursMerge branch 'modules-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/modules/linux.git
12 hoursMerge branch 'linux-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm.git
13 hoursMerge branch 'for-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/mm/linux.git # Conflicts: # arch/arm64/kvm/mmu.c
25 hoursuprobes: remove VM_IO, set VM_MIXEDMAP for mapped kernel pagesLorenzo Stoakes (ARM)
These are not MMIO pages so VMA_IO_BIT is an inappropriate flag to set. Instead, set them VMA_MIXEDMAP_BIT as they are kernel mappings and this is the appropriate flag to set for those. This provides the semantics required - no VMA merging is permitted, but does not prevent GUP. However this has no meaningful impact as these are refcounted and thus can be GUPed. A previous commit already prevented __mm_populate() from being invoked on XOL areas, which VMA_IO_BIT was previously relied upon to do, so that is no longer required. Both VMAs set a VMA name, so always_dump_vma() returns true before vma_dump_size() reaches its VMA_IO_BIT check, and thus there is no change in core dump behaviour. Change this for both the core xol_add_vma() function and the x86-specific get_uprobe_trampoline() function. Link: https://lore.kernel.org/20260917-b4-mmap-prepare-vma-flag-sanify-v3-22-4583d8a23bca@kernel.org Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Cc: Albert Ou <aou@eecs.berkeley.edu> Cc: Alexander Gordeev <agordeev@linux.ibm.com> Cc: Alexei Starovoitov <ast@kernel.org> Cc: Alistair Popple <apopple@nvidia.com> Cc: Al Viro <viro@zeniv.linux.org.uk> Cc: Andreas Larsson <andreas@gaisler.com> Cc: Andrii Nakryiko <andrii@kernel.org> Cc: "Aneesh Kumar K.V" <aneesh.kumar@kernel.org> Cc: Anup Patel <anup@brainfault.org> Cc: Arnaldo Carvalho de Melo <acme@kernel.org> Cc: Arnd Bergmann <arnd@arndb.de> Cc: Axel Rasmussen <axelrasmussen@google.com> Cc: Baolin Wang <baolin.wang@linux.alibaba.com> Cc: Baoquan He <baoquan.he@linux.dev> Cc: Barry Song <baohua@kernel.org> Cc: "Borislav Petkov (AMD)" <bp@alien8.de> Cc: Byungchul Park <byungchul@sk.com> Cc: Catalin Marinas <catalin.marinas@arm.com> Cc: Chengming Zhou <chengming.zhou@linux.dev> Cc: Chris Li <chrisl@kernel.org> Cc: Christian Borntraeger <borntraeger@linux.ibm.com> Cc: Christian Brauner <brauner@kernel.org> Cc: Claudio Imbrenda <imbrenda@linux.ibm.com> Cc: Dave Airlie <airlied@gmail.com> Cc: Dave Hansen <dave.hansen@linux.intel.com> Cc: David Hildenbrand <david@kernel.org> Cc: David S. Miller <davem@davemloft.net> Cc: Dennis Dalessandro <dennis.dalessandro@cornelisnetworks.com> Cc: Dev Jain <dev.jain@arm.com> Cc: Doug Gilbert <dgilbert@interlog.com> Cc: Eduard Zingerman <eddyz87@gmail.com> Cc: Emil Tsalapatis <emil@etsalapatis.com> Cc: Gerald Schaefer <gerald.schaefer@linux.ibm.com> Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org> Cc: Gregory Price <gourry@gourry.net> Cc: Harry Yoo <harry@kernel.org> Cc: Heiko Carstens <hca@linux.ibm.com> Cc: Helge Deller <deller@gmx.de> Cc: "Huang, Ying" <ying.huang@linux.alibaba.com> Cc: Ingo Molnar <mingo@redhat.com> Cc: James Bottomley <james.bottomley@HansenPartnership.com> Cc: Jan Kara <jack@suse.cz> Cc: Jann Horn <jannh@google.com> Cc: Janosch Frank <frankja@linux.ibm.com> Cc: Jaroslav Kysela <perex@perex.cz> Cc: Jason Gunthorpe <jgg@ziepe.ca> Cc: Jaya Kumar <jayalk@intworks.biz> Cc: Johannes Weiner <hannes@cmpxchg.org> Cc: John Hubbard <jhubbard@nvidia.com> Cc: Jonathan Corbet <corbet@lwn.net> Cc: Joshua Hahn <joshua.hahnjy@gmail.com> Cc: Juri Lelli <juri.lelli@redhat.com> Cc: Kairui Song <kasong@tencent.com> Cc: Kemeng Shi <shikemeng@huaweicloud.com> Cc: Kiryl Shutsemau <kas@kernel.org> Cc: Kumar Kartikeya Dwivedi <memxor@gmail.com> Cc: Lance Yang <lance.yang@linux.dev> Cc: Leon Romanovsky <leon@kernel.org> Cc: Liam R. Howlett <liam@infradead.org> Cc: Maarten Lankhorst <maarten.lankhorst@linux.intel.com> Cc: Madhavan Srinivasan <maddy@linux.ibm.com> Cc: Marc Rutland <mark.rutland@arm.com> Cc: Marc Zyngier <maz@kernel.org> Cc: "Masami Hiramatsu (Google)" <mhiramat@kernel.org> Cc: Matthew Brost <matthew.brost@intel.com> Cc: Matthew Wilcox (Oracle) <willy@infradead.org> Cc: Maxime Ripard <mripard@kernel.org> Cc: Michal Hocko <mhocko@kernel.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Mike Rapoport <rppt@kernel.org> Cc: Miklos Szeredi <miklos@szeredi.hu> Cc: Muchun Song <muchun.song@linux.dev> Cc: Namhyung kim <namhyung@kernel.org> Cc: Nhat Pham <nphamcs@gmail.com> Cc: Nicholas Piggin <npiggin@gmail.com> Cc: Oleg Nesterov <oleg@redhat.com> Cc: Oscar Salvador <osalvador@suse.de> Cc: Palmer Dabbelt <palmer@dabbelt.com> Cc: Paul Moore <paul@paul-moore.com> Cc: Pedro Falcato <pfalcato@suse.de> Cc: Peter Xu <peterx@redhat.com> Cc: Peter Zijlstra <peterz@infradead.org> Cc: Rakie Kim <rakie.kim@sk.com> Cc: Rik van Riel <riel@surriel.com> Cc: Ryan Roberts <ryan.roberts@arm.com> Cc: Sebastian Reichel <sre@kernel.org> Cc: Shakeel Butt <shakeel.butt@linux.dev> Cc: Stephen Smalley <stephen.smalley.work@gmail.com> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Takashi Iwai (SUSE) <tiwai@suse.de> Cc: Takashi Iwai <tiwai@suse.com> Cc: Thomas Zimemrmann <tzimmermann@suse.de> Cc: Vasily Gorbik <gor@linux.ibm.com> Cc: Vincent Guittot <vincent.guittot@linaro.org> Cc: Vlastimil Babka <vbabka@kernel.org> Cc: Wei Xu <weixugc@google.com> Cc: Will Deacon <will@kernel.org> Cc: Yuanchu Xie <yuanchu@google.com> Cc: Zi Yan <ziy@nvidia.com>
28 hoursMerge branch 'misc'Sean Christopherson
* misc: (44 commits) KVM: selftests: Add coverage for the EFER_LMSLE_MBZ defeature KVM: selftests: Add module param API to check if nested virtualization is enabled KVM: selftests: Rename svm_nested_clear_efer_svme to svm_nested_efer_test KVM: x86: Honor the guest's EFER_LMSLE_MBZ KVM: x86: Advertise EFER_LMSLE_MBZ when KVM disallows EFER.LMSLE KVM: SVM: Add support for virtualizating Bus Lock Detect KVM: SVM: Add a vCPU-aware helper to get supported DEBUGCTL bits KVM: nSVM: Don't assume all active-low bits DR6 are fixed-1 KVM: nSVM: Open code check on LBR virtualization being enabled in vmcb12 KVM: nSVM: Disable LBRV in nested control cache when unsupported KVM: SVM: Add helper to query if LBR virtualization needs to be enabled KVM: x86: Kill off DR6_FIXED_1 to prevent future misuse KVM: x86: Set guest's fixed-1 DR6 bits when delivering #DB payload KVM: x86: Force fixed-1 bits in DR6 after synchronizing with hardware KVM: Return an "unsigned long", not "u64" for the fixed-1 DR6 bits KVM: x86: Rename kvm_dr6_fixed() => kvm_get_dr6_fixed_1() KVM: x86: Preserve DR6.BLD (Bus Lock Detect) when delivering #DB payload KVM: x86: Restrict saved GPA writes to hardware write faults KVM: x86/pmu: Don't retry a counter whose config was rejected KVM: SVM: Add Page modification logging support ...
42 hoursMerge branch into tip/master: 'x86/tdx'Ingo Molnar
# New commits in x86/tdx: a49e2d257594 ("x86/tdx: Restrict attestation exports to the tdx-guest driver") d48b2d347105 ("virt: tdx-guest: Remove unused and confusing function argument") c9fc85f1e44e ("x86/virt/tdx: Formalize SEAMCALL leaf version encoding support") e61c837817ac ("Documentation/x86: Add documentation for TDX's Dynamic PAMT") 1e3d79135327 ("x86/virt/tdx: Enable Dynamic PAMT") d7fe610e6a66 ("KVM/TDX: Get/put PAMT pages when (un)mapping private memory") 2c8ca2ba9bb6 ("x86/virt/tdx: Add APIs to support Dynamic PAMT ops from KVM's fault path") 86305da521d1 ("KVM/TDX: Allocate PAMT memory for TD and vCPU control structures") eca9d4a3f1f8 ("x86/virt/tdx: Handle multiple callers in tdx_pamt_get/put()") 5386c5288001 ("x86/virt/tdx: Allocate refcounts for Dynamic PAMT memory") ad52ce8389d9 ("x86/virt/tdx: Add __tdx_pamt_get/put() helpers") eadde434aa9d ("x86/virt/tdx: Allocate page bitmap for Dynamic PAMT") 01a43c55ae5f ("x86/virt/tdx: Simplify PAMT layout calculation") Signed-off-by: Ingo Molnar <mingo@kernel.org>
42 hoursMerge branch into tip/master: 'x86/sgx'Ingo Molnar
# New commits in x86/sgx: 993b65a2fc6f ("x86/sgx: Report RCU-Tasks quiescent state in EPC sanitization loop") Signed-off-by: Ingo Molnar <mingo@kernel.org>
42 hoursMerge branch into tip/master: 'x86/sev'Ingo Molnar
# New commits in x86/sev: 250ee734115b ("x86/sev: Report MSR_AMD64_SEV in sysfs") edb1b49e724e ("x86/sev: Do RMP optimizations on SNP guest shutdown") c6c6f156d192 ("x86/sev: Perform RMP optimizations asynchronously") 15d974be0a3b ("x86/sev: Initialize RMPOPT configuration MSRs") 674bb17a32c7 ("x86/sev: Disable CPU hotplug while SNP is active") 1ee97d0f7b20 ("x86/cpufeatures: Add X86_FEATURE_RMPOPT feature flag") Signed-off-by: Ingo Molnar <mingo@kernel.org>
42 hoursMerge branch into tip/master: 'x86/microcode'Ingo Molnar
# New commits in x86/microcode: cf086537af4e ("x86/microcode/intel: Refresh old_microcode defines with the May 2026 release") Signed-off-by: Ingo Molnar <mingo@kernel.org>
42 hoursMerge branch into tip/master: 'x86/kdump'Ingo Molnar
# New commits in x86/kdump: d949fa7b1ec5 ("crash: Update stale NR_CPUS_DEFAULT references in elfcorehdr sizing docs") 6664ad1026b5 ("x86/crash: Reserve elfcorehdr for CONFIG_NR_CPUS, not CONFIG_NR_CPUS_DEFAULT") Signed-off-by: Ingo Molnar <mingo@kernel.org>
42 hoursMerge branch into tip/master: 'x86/cpu'Ingo Molnar
# New commits in x86/cpu: 2711d67bc3a7 ("x86/cpu: Don't transiently clear the boot CPU's capabilities") db349783ca7a ("x86/cpu: Move 32-bit SEP setup into identify_cpu()") ad86fe2134cc ("x86/cpu: Inline generic_identify() into identify_cpu()") 7a762b51a149 ("x86/cpu: Initialize boot CPU cpuinfo defaults early") 54a2882ab9f7 ("x86/cpu: Factor init_cpu_info() out of identify_cpu()") 228200f695c0 ("x86/CPU/AMD: Fix Zen5 TLB sizes reporting") 9025f3bf73f0 ("x86/shstk: Return the correct error value when user shadow stacks are disabled") Signed-off-by: Ingo Molnar <mingo@kernel.org>
42 hoursMerge branch into tip/master: 'x86/cleanups'Ingo Molnar
# New commits in x86/cleanups: 2c6a75adb15f ("x86/um: Remove unused <asm/required-features.h> header") c7243accb48a ("x86/cpu: Constify struct x86_cpu_id") 975809caf4b0 ("x86/asm/string_64: Remove unused linux/jump_label.h include") c8beda268a6b ("x86/mtrr: Fix kernel-doc notation of amd_set_mtrr()") Signed-off-by: Ingo Molnar <mingo@kernel.org>
42 hoursMerge branch into tip/master: 'x86/cache'Ingo Molnar
# New commits in x86/cache: a17d099481bd ("x86/resctrl: Update documented unit for the "activity" event") a3f5e6eba418 ("fs/resctrl: Simplify pseudo_lock_measure_trigger()") 4b31656d917c ("fs/resctrl: Avoid extra call to strlen() in schemata_list_add()") 942135446059 ("fs/resctrl: Factor MBA parse-time conversion to be per-arch") 255f92316ac1 ("arm_mpam: resctrl: Add pass-through resctrl_arch_preconvert_bw()") 57a354d20a01 ("x86,fs/resctrl: Add resctrl_arch_preconvert_bw()") Signed-off-by: Ingo Molnar <mingo@kernel.org>
42 hoursMerge branch into tip/master: 'x86/bugs'Ingo Molnar
# New commits in x86/bugs: ae1d2082d93b ("x86/bugs: Adapt SRSO mitigation to Zen6") Signed-off-by: Ingo Molnar <mingo@kernel.org>
42 hoursMerge branch into tip/master: 'sched/core'Ingo Molnar
# New commits in sched/core: 1fb28c664a19 ("virt/steal_governor: Enable the driver") 27d47ebce4d6 ("virt/steal_governor: Implement steal_governor policy loop") 4b9302d494ff ("virt/steal_governor: Add control knobs for handling steal values") 9a8e740ee9f6 ("virt: Introduce steal governor driver") 68957caaa9c0 ("sched/debug: Add migration stats due to non preferred CPUs") 74699f56ebcf ("sched/core: Push current task from non preferred CPU") 4ee29b029058 ("sched/fair: Load balance only among preferred CPUs") d8a3da0de843 ("sched/core: Try to use a preferred CPU in is_cpu_allowed") 620824516557 ("sysfs: Add preferred CPU file") 518b32bd5bb3 ("cpumask: Introduce cpu_preferred_mask") 06a49ef784ac ("sched/docs: Document cpu_preferred_mask and Preferred CPU concept") cfb463b7172d ("cpumask: Introduce cpumask_intersects_and") a8d0854a76a8 ("sched/cputime: Add kcpustat_field_total helper") be100c77178e ("sched: Add sched_ext hooks for proxy execution") 57c75e3ae38c ("sched: Add helper to block retained proxy donors") a49653d0abeb ("sched/core: Mark wakeups completed through ttwu_runnable()") 8f8c0417e973 ("sched/core: Dequeue waking proxy donors before reset") 313b652837d0 ("sched/core: Drop mutex locks before proxy rescheduling") 627ea30aca3b ("sched/wait: Clarify WF_SYNC wakeup semantics") d2e010082757 ("sched/eevdf: Handle more short slice waking cases") 4bf32ec3327d ("sched/eevdf: Align update_protect_slice to set_protect_slice") aae2a33ea662 ("sched/eevdf: Ensure that vprot will never go above a min slice") c9ce69fc43bd ("sched/fair: Randomize equally shallow slow-path candidates") abe440b3770f ("sched/fair: Drop idle recency from slow-path CPU selection") fbbc63fed0b0 ("sched/core: Remove redundant core_sched_seq") 819224e506bc ("sched/fair: Remove dead code on enqueue_task_fair()") c72945693b90 ("sched: Restart fair hrtick after same-task repicks") a9b3c7570564 ("sched/headers: Replace __ASSEMBLY__ with __ASSEMBLER__ in the <uapi/linux/sched.h> header") e81ee0630837 ("sched/fair: Reset NUMA fault locality after scan period update") ef9293b3b797 ("sched: dynamic: Fix preemption model strings") 879eaa76e608 ("sched: Remove unneeded function type cast in do_balance_callbacks()") f549101187c8 ("sched/deadline: check start_dl_timer expiry with ktime_before()") 2a672daa4b27 ("sched/feat: Use the new static key API for sched_feat") a5576ebce920 ("sched: Convert paravirt_steal to new static key APIs") 9650ce11f2e3 ("sched: dynamic: Simplify preempt model accessors") 5b9a28eeed37 ("sched: dynamic: Remove HAVE_PREEMPT_DYNAMIC_{CALL,KEY}") aa4178f63847 ("sched: dynamic: Simplify irqentry_exit_cond_resched()") b9d267b9d632 ("sched: dynamic: Simplify preempt_schedule{,_notrace}()") 88e0b3bb9930 ("sched: dynamic: Simplify {cond,might}_resched()") d3d16750693b ("sched: dynamic: Make PREEMPT_DYNAMIC depend on ARCH_HAS_PREEMPT_LAZY") 772d9ffbfd26 ("sched: Migrate whole chain in proxy_migrate_task()") 6b73a09e943f ("sched: Break out core of attach_tasks() helper into sched.h") 1f8805138593 ("sched: Switch rq->next_class in proxy_reset_donor()") 09351db90a28 ("sched/core: Don't proxy-exec unmatched cookie lock owners") 9be817f991e2 ("sched/core: Avoid migrating blocked_on tasks") 3dd95f077371 ("sched/core: Don't steal a proxy-exec donor") Signed-off-by: Ingo Molnar <mingo@kernel.org>
42 hoursMerge branch into tip/master: 'perf/merge'Ingo Molnar
# New commits in perf/merge: 7606d01b963f ("x86/fpu: Remove additional remnants of math emulation") 67639ee98521 ("x86/fpu: Remove unnecessary checks for X86_FEATURE_FPU") 391cac638edd ("selftests/x86: Add tests for signal frame FPU portability") fd14edd82077 ("x86/fpu: Pre-fault only required size of xstate buffer") f1751b220bc8 ("x86/fpu: Fix potential underflow in xstate_calculate_size()") 731300df4f7d ("x86/fpu: Document reasoning of FX-only fallback") 6916dec7a02f ("x86/fpu: Extract restore_from_ia32_fxstate() and clean up fpu__restore_sig()") 344907fab5fd ("x86/fpu: Clean up and rename variables in signal frame handling") 3d69596c684b ("x86/fpu: Document signal frame layout and portability") 6350de8671b9 ("perf/x86/amd/uncore: Free counter slot by index") 4a8557a4e5d2 ("perf/x86/amd/uncore: Remove redundant event slot scan") ba29babd881e ("perf/x86/intel: Check only PMC bits in PEBS_ENABLED when detecting host PEBS usage") 193ef44e3261 ("KVM: VMX: Only tell perf to enable PEBS counters for fully enabled PMCs") 57c764783330 ("KVM: VMX: Drop a redundant pmu->global_ctrl check when processing pebs_enable") 31f7cf337ce7 ("perf/x86/intel: KVM: Handle cross-mapped PEBS PMCs entirely within KVM") a634e4536ec2 ("perf/x86: KVM: Have perf define a dedicated struct for getting guest PEBS data") b7b84fff2bd6 ("perf/x86/intel: Invert names of intel_ctrl_{guest,host}_mask") 52a6457f59ad ("perf/x86/intel: Annotate x86_pmu::hybrid_pmu with __counted_by_ptr") df935e26ca9a ("perf/x86/amd/uncore: Turn amd_uncore_ctx events into a flexible array") c0a44d1745d3 ("xor: Add AVX-512 optimized xor_gen()") 2405d15e5ae4 ("xor: Remove redundant X86_FEATURE_OSXSAVE check") b93f05df9c66 ("lib/crc: x86: Stop using cpu_has_xfeatures()") ffe52429f870 ("lib/crypto: x86: Stop using cpu_has_xfeatures()") 9256e0496c14 ("crypto: x86 - Stop using cpu_has_xfeatures()") 19c5d90800a2 ("um: Check for missing AVX and AVX-512 xstate bits") 518aa9948e4f ("x86/fpu: Check for missing AVX and AVX-512 xstate bits") 68aca309e49c ("perf/x86/intel: Add sanity check for PEBS record/fragment size") c34db2f95094 ("perf/x86: Activate back-to-back NMI detection for arch-PEBS induced NMIs") 00cf8daabe2b ("perf/x86/intel: Advertise PERF_PMU_CAP_SIMD_REGS capability") c4fabb67a47f ("perf/x86/intel: Support arch-PEBS based SIMD/eGPRs sampling") 098cdd582a9b ("perf/x86: Support SSP sampling using sample_regs_* fields") 578460d9af8d ("perf/x86: Support eGPRs sampling using sample_regs_* fields") a2c64c74c029 ("perf: Enhance perf_reg_validate() with simd_enabled argument") 74d55a827e31 ("perf/x86: Support OPMASK sampling using sample_simd_pred_reg_* fields") 3807f6996a0b ("perf/x86: Support ZMM sampling using sample_simd_vec_reg_* fields") b76210d32147 ("perf/x86: Support YMM sampling using sample_simd_vec_reg_* fields") e9d76ada769c ("perf/x86: Support XMM sampling using sample_simd_vec_reg_* fields") 918b7d6d1729 ("perf: Add sampling support for SIMD registers") edd9aec51213 ("perf/x86: Enable XMM register sampling for REGS_USER case") b84c96283684 ("perf/x86: Enable XMM register sampling for non-PEBS events") c09466807467 ("perf/x86/intel: Centralize PERF_PMU_CAP_EXTENDED_REGS updates") 05fe8825796e ("perf: Move and enhance has_extended_regs() for arch-specific use") 450d73dc5f91 ("x86/fpu: Add update_fpu_state_and_flag() helper") 03b89c9f202e ("x86/fpu/xstate: Add xsaves_nmi() helper") 9c05620b4971 ("perf/x86: Use x86_perf_regs in NMI handlers") cb388bd530cc ("perf: Eliminate duplicate arch-specific function definitions") bbdf84fcedfa ("perf/x86/intel: Convert x86_perf_regs to per-cpu variables") 62e074f55f41 ("perf/x86/intel: Enable large PEBS sampling for XMMs") 9106892e27ca ("perf/x86: Move hybrid PMU initialization before x86_pmu_starting_cpu()") Signed-off-by: Ingo Molnar <mingo@kernel.org>
3 daysMerge branch 'acpi-cppc' into linux-nextRafael J. Wysocki
* acpi-cppc: cpufreq: CPPC: Create the FIE worker before enabling PCC callbacks cpufreq: CPPC: Select the frequency-invariance callback per CPU ACPI: CPPC: Accept requests to retain immutable autonomous selection ACPI: CPPC: Propagate errors from cross-CPU FFH calls ACPI: CPPC: Validate FFH register fields before hardware access ACPI: CPPC: Keep Performance Limited clearable on NVIDIA T41 ACPI: CPPC: Clear Performance Limited without a stale read ACPI: CPPC: Validate SystemIO overlaps across processors ACPI: CPPC: Validate PCC overlaps across processors ACPI: CPPC: Validate SystemIO register layouts ACPI: CPPC: Validate and access PCC register layouts ACPI: CPPC: Reject direct reads of write-only controls ACPI: CPPC: Reject unsafe cross-CPU SystemMemory RMW ACPI: CPPC: Release PCC data after probe failures ACPI: CPPC: Release CPC descriptors through kobject ACPI: CPPC: Serialize PCC EPP payload updates ACPI: CPPC: Serialize PCC single-register payload updates ACPI: CPPC: Propagate performance-control write errors ACPI: CPPC: Validate _CPC entry and control semantics ACPI: CPPC: Validate the _CPC package header
3 daysx86/cpufeatures: Add Page modification loggingNikunj A Dadhania
Page modification logging(PML) is a hardware feature designed to track guest modified memory pages. PML enables the hypervisor to identify which pages in a guest's memory have been changed since the last checkpoint or during live migration. The PML feature is advertised via CPUID leaf 0x8000000A ECX[4] bit. Acked-by: Borislav Petkov (AMD) <bp@alien8.de> Signed-off-by: Nikunj A Dadhania <nikunj@amd.com> Link: https://patch.msgid.link/20260907063906.1964557-7-nikunj@amd.com Signed-off-by: Sean Christopherson <seanjc@google.com>
3 daysx86/microcode/intel: Refresh old_microcode defines with the May 2026 releaseSohil Mehta
Update the minimum expected revisions of Intel microcode based on the microcode-20260512 (May 2026) release. Note, the three new entries are for INTEL_PANTHERLAKE_L steppings. Signed-off-by: Sohil Mehta <sohil.mehta@intel.com> Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com> Link: https://patch.msgid.link/20260928222607.3593973-1-sohil.mehta@intel.com
4 daysx86/mm: Don't apply va_align to hugetlb mappings on AMD F15hLaurent Wandrebeck
get_align_mask() returns huge_page_mask_align() for hugetlbfs, but get_align_bits() adds va_align.bits regardless, so vm_unmapped_area() returns an address off the huge page boundary and __unmap_hugepage_range() hits BUG_ON(start & ~huge_page_mask(h)) at teardown. This can be triggered on Carrizo and FX-8370E, both hstates. Pass the file to get_align_bits() and skip the randomisation for hugetlbfs. [ bp: Massage commit message. ] Fixes: 1317a5e7f7b1 ("arch/x86: teach arch_get_unmapped_area_vmflags to handle hugetlb mappings") Suggested-by: Dave Hansen <dave.hansen@intel.com> Acked-by: Dave Hansen <dave.hansen@intel.com> Signed-off-by: Laurent Wandrebeck <l.wandrebeck@quelquesmots.fr> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Cc: stable@vger.kernel.org # 6.13+ Link: https://patch.msgid.link/20260922085032.46144-1-l.wandrebeck@quelquesmots.fr
4 daysMerge tag 'v7.3-rc5' into driver-core-nextDanilo Krummrich
We need the driver-core fixes in here as well to build on top of. Signed-off-by: Danilo Krummrich <dakr@kernel.org>
4 daysx86/cpu: Constify struct x86_cpu_idChristophe JAILLET
'struct x86_cpu_id' is not modified in this compilation unit. Constifying this structure moves some data to a read-only section, so increases overall security. It is only used in cpu_has_old_microcode() which is an __init function. So, using __initconst is safe and will save about 6 kB of memory at runtime. On a x86_64, with allmodconfig: text data bss dec hex filename 55725 28774 512 85011 14c13 arch/x86/kernel/cpu/common.o.before 61464 22918 512 84894 14b9e arch/x86/kernel/cpu/common.o.after [ bp: Massage commit message. ] Signed-off-by: Christophe JAILLET <christophe.jaillet@wanadoo.fr> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Link: https://patch.msgid.link/f08d3a0e7aefe3bad66251927cbd03b8cbf9df73.1786311119.git.christophe.jaillet@wanadoo.fr
5 daysMerge branch 'x86/fpu' into perf/merge, to resolve semantic conflictIngo Molnar
Remove duplicate xstate_calculate_size() prototype. Signed-off-by: Ingo Molnar <mingo@kernel.org>
6 daysx86/fpu: Remove unnecessary checks for X86_FEATURE_FPUEric Biggers
Since ab05214025ee ("x86/fpu: Remove MATH_EMULATION and related glue code"), X86_FEATURE_FPU is mandatory. If it's absent, the kernel halts execution in fpu__init_system_early_generic(). Therefore, remove unnecessary checks for X86_FEATURE_FPU that occur later in x86/fpu code. This makes struct swregs_state and the extern declarations of fpregs_soft_get and fpregs_soft_set all unused. So remove those too. Signed-off-by: Eric Biggers <ebiggers@kernel.org> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Link: https://patch.msgid.link/20260924033405.57613-1-ebiggers@kernel.org
6 daysx86/fpu: Pre-fault only required size of xstate bufferAndrei Vagin
The kernel previously used the default task FPU state size (user_size) to fault in the user buffer when restoring FPU registers from a signal frame. This can lead to attempting to fault in memory past the end of the actual frame if the frame was smaller than the default size. Introduce consistency checks to calculate the actual required size for the features enabled in the xfeatures mask, ensure that the provided xstate_size is sufficient, and shrink it to the actual required size. Use this validated size to fault in the user buffer. Keep the strict check that the provided xstate_size does not exceed the default user_size for now. Signed-off-by: Andrei Vagin <avagin@google.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Reviewed-by: Alexander Mikhalitsyn <alexander@mihalicyn.com> Reviewed-by: Chang S. Bae <chang.seok.bae@intel.com> Link: https://patch.msgid.link/20260925162454.1403405-7-avagin@google.com
6 daysx86/fpu: Fix potential underflow in xstate_calculate_size()Andrei Vagin
xstate_calculate_size() calculates the size required for a given set of xfeatures. It determines the topmost feature by finding the most significant bit in xfeatures using fls64(xfeatures) - 1. If xfeatures is 0, fls64(0) returns 0, and topmost becomes -1. Previously, topmost was unsigned int, so -1 underflowed to UINT_MAX. This caused the subsequent check `topmost <= XFEATURE_SSE` to fail, and the code proceeded to access xstate arrays using topmost (UINT_MAX) as an index, leading to an out-of-bounds access. [ bp: Remove text explaining what the patch does. ] Fixes: d6d6d50f1e80 ("x86/fpu/xstate: Consolidate size calculations") Signed-off-by: Andrei Vagin <avagin@google.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Reviewed-by: Alexander Mikhalitsyn <alexander@mihalicyn.com> Reviewed-by: Chang S. Bae <chang.seok.bae@intel.com> Link: https://patch.msgid.link/20260925162454.1403405-6-avagin@google.com
6 daysx86/fpu: Document reasoning of FX-only fallbackAndrei Vagin
Add a comment to check_xstate_in_sigframe() to explain the reasoning behind falling back to the FX-only state when signal frame metadata is inconsistent. The fallback is intended to preserve backward compatibility with legacy user-space processes that are not aware of XSAVE states and might only fill or copy the legacy FP state. However, this fallback should be avoided whenever possible. If a process was actively using extended features, falling back to the FX-only state silently resets those extended registers to their initial state, which can lead to silent user-space state corruption. Signed-off-by: Andrei Vagin <avagin@google.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Reviewed-by: Alexander Mikhalitsyn <alexander@mihalicyn.com> Reviewed-by: Chang S. Bae <chang.seok.bae@intel.com> Link: https://patch.msgid.link/20260925162454.1403405-5-avagin@google.com
6 daysx86/fpu: Extract restore_from_ia32_fxstate() and clean up fpu__restore_sig()Andrei Vagin
When restoring an FPU signal frame for a 32-bit or IA32-compat task on FXSR-enabled systems, a legacy 32-bit FP frame is present alongside the FX/XSAVE frame. Because the legacy FP frame duplicates the FP state portion of the FX/XSAVE frame, for backward compatibility it is treated as the source of truth, and its state is folded into the FX/XSAVE state before restoring the registers. Currently, most of __fpu_restore_sig() is dedicated to handling this 32-bit legacy/compat fpstate, while the native direct restore path lives in restore_fpregs_from_user(). Having the compat handling intermixed with the main signal restoration flow makes it tricky to quickly see what code is doing what. Extract the 32-bit legacy/compat FPU restore handling into a separate helper function, restore_from_ia32_fxstate(), and inline the remainder of __fpu_restore_sig() directly into fpu__restore_sig(). Signed-off-by: Andrei Vagin <avagin@google.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Reviewed-by: Alexander Mikhalitsyn <alexander@mihalicyn.com> Reviewed-by: Chang S. Bae <chang.seok.bae@intel.com> Link: https://patch.msgid.link/20260925162454.1403405-4-avagin@google.com
6 daysx86/fpu: Clean up and rename variables in signal frame handlingAndrei Vagin
Clean up signal frame handling code by renaming several variables for clarity and consistency, and moving masking logic closer to its usage. - Rename 'fxbuf' to 'buf_fx' in check_xstate_in_sigframe() for consistency. - Rename label 'setfx' to 'err_setfx' in check_xstate_in_sigframe() to indicate it is an error path. - In __restore_fpregs_from_user(), rename 'ufeatures' to 'task_xfeatures' and 'xrestore' to 'xrestore_mask'. - Move the masking logic 'xrestore_mask &= task_xfeatures' from restore_fpregs_from_user() into __restore_fpregs_from_user(). - Rename 'xrestore' to 'xrestore_mask' in restore_fpregs_from_user() to match the name in __restore_fpregs_from_user() and __fpu_restore_sig(). - In __fpu_restore_sig(), rename 'buf' to 'buf_f' to distinguish it from 'buf_fx', and 'user_xfeatures' to 'xrestore_mask'. No functional changes. Suggested-by: Ingo Molnar <mingo@kernel.org> Signed-off-by: Andrei Vagin <avagin@google.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Reviewed-by: Alexander Mikhalitsyn <alexander@mihalicyn.com> Link: https://patch.msgid.link/20260925162454.1403405-3-avagin@google.com
8 dayskho: rename KHO scratch to KHO bootmemPratyush Yadav (Google)
The term "KHO scratch" is vague and overloaded. For one, it does not accurately describe what the memory is for. From KHO's perspective, it is memory passed by the previous kernel that is guaranteed to have no allocations and hence it is safe to allocate from. For another, the term is overloaded. KHO also knows to discover other ranges of memory that don't have any preservations and are safe to allocate from. This is done by kho_extend_scratch(). These areas, while called "scratch", are completely different from the chunk of memory passed by the previous kernel for early boot allocations. Rename "KHO scratch" to "KHO boot memory", or "KHO bootmem" in short. Update all function names, variable names, comments, and documentation to use this. Call the discovered areas "noprsrv", matching what memblock would see them as. Rename kho_extend_scratch() to reflect this. This patch largely has no functional changes. The only functional changes are renames of debugfs files from "scratch_phys" and "scratch_len" to "bootmem_phys" and "bootmem_len" respectively. Signed-off-by: Pratyush Yadav (Google) <pratyush@kernel.org> Link: https://patch.msgid.link/20260922041321.233986-4-pratyush@kernel.org Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
8 daysmemblock: rename KHO_SCRATCH to KHO_NOPRSRVPratyush Yadav (Google)
The term "KHO scratch" is vague. It does not accurately describe what the memory is for. Generally, scratch is for things that don't contain anything useful. This is not true. For memblock, KHO_SCRATCH is the _only_ useful memory during a KHO boot. It is memory passed by the previous kernel that is guaranteed to have no preserved pages, and so is the only thing safe to allocate from early in boot. In addition, KHO now has the ability to discover other "scratch" areas. These aren't directly passed by the previous kernel. Instead, they are found by searching the KHO preserved memory map. Memblock does not care about the difference between the two. From memblock's perspective, both are memory areas with no preservations and so they are safe to allocate from before KHO fully initializes. Rename MEMBLOCK_KHO_SCRATCH to MEMBLOCK_KHO_NOPRSRV. This name clearly shows what the memory is for from memblock's perspective. Update memblock functions and comments that refer to scratch to use noprsrv. There are no functional changes apart from the flag name change in debugfs. Signed-off-by: Pratyush Yadav (Google) <pratyush@kernel.org> Link: https://patch.msgid.link/20260922041321.233986-3-pratyush@kernel.org Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
8 daysx86/mce: Fix hardware debug register corruption on task migrationMasami Hiramatsu (Google)
In exc_machine_check_user(), local_db_save() and local_db_restore() are invoked in the outer entry stubs (DEFINE_IDTENTRY_MCE_USER, DEFINE_FREDENTRY_MCE, and DEFINE_IDTENTRY_RAW), surrounding exc_machine_check_user(). However, exc_machine_check_user() calls irqentry_exit_to_user_mode(), which handles pending thread work and may schedule() if TIF_NEED_RESCHED is set. If the task migrates to another CPU during schedule(), local_db_restore() runs on the new CPU with the dr7 state saved from the old CPU. This corrupts the new CPU's DR7 hardware debug register and leaves the old CPU's DR7 disabled. In short, local_db_save() and local_db_restore() pair must be run on the same CPU. To fix this, move local_db_save() and local_db_restore() inside exc_machine_check_user() and exc_machine_check_kernel(). In exc_machine_check_user(), DR7 is saved and restored strictly around do_machine_check() to avoid schedule() during migration. In exc_machine_check_kernel(), local_db_save() is called at the entry point to prevent early memory accesses from triggering nested #DB exceptions, and restored on all exits. Fixes: cd840e424f27 ("x86/entry, mce: Disallow #DB during #MC") Assisted-by: LLM Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org> Cc: <stable@kernel.org> Link: https://patch.msgid.link/179005109564.388919.3937970081044095776.stgit@devnote2
9 daysMerge branch 'perf/urgent' into perf/core, to resolve conflictIngo Molnar
Conflicts: include/linux/perf_event.h Signed-off-by: Ingo Molnar <mingo@kernel.org>
10 daysumh, treewide: Explicitly include linux/umh.h where neededPetr Pavlu
The usermode helper declarations were previously provided by linux/kmod.h but commit c1f3fa2a4fde ("kmod: split off umh headers into its own file") moved them to linux/umh.h in 2017. Add explicit includes of linux/umh.h to files that use usermode helpers and remove linux/kmod.h where it is no longer needed. Acked-by: Alex Elder <elder@riscstar.com> # for greybus Signed-off-by: Petr Pavlu <petr.pavlu@suse.com>
10 daysx86/virt/tdx: Formalize SEAMCALL leaf version encoding supportXu Yilun
The TDX architecture includes a syscall-like ABI for OSes to communicate with SEAM mode software. This ABI has the concept of a "leaf". Each leaf is roughly analogous to a Linux syscall: it has a set of register arguments and does one logical thing like adding a page of memory to a VM or running a VM. The TDX architecture refers to this interface function as a "SEAMCALL leaf". Just like syscalls, the TDX architecture wants to evolve the ABI to extend functionality while keeping compatibility. But unlike syscalls, which do this by picking a totally new syscall number, the ABI encodes the "version" number directly into the bits of a register argument. So instead of openatN being whatever the next free number is, the SEAMCALL leaf number and version number are encoded into certain bits in how the RAX register is defined in the ABI. In Linux, several seamcall*() wrappers have been introduced to invoke SEAMCALL leafs. They all take a u64 "fn" argument for the leaf number which eventually gets set in the register. As above, the version number lives in the same register as the leaf number. So callers that want to select a specific version of the leaf, can jam it in the right place in the register by passing it in the "fn" argument. Today only the caller of TDH.VP.INIT does this hack, but future kernel changes will need to select versions for more SEAMCALL leafs. So a less hacky solution is needed. Explicitly define the arguments for seamcall*() wrappers: - Add a build-time assertion to ensure "fn" only contains valid SEAMCALL leaf number bits. Note the P-SEAMLDR selector bit (bit 63) is part of the leaf number for SEAM loader calls, so it is allowed. This also means the "fn" can't be a narrower type, such as u16, to exclude unrelated bits. - Add a "version" field in struct tdx_module_args [1], so most existing callers get a default "version == 0" behavior without code churn. Update the TDH.VP.INIT caller to specify the version descriptively. Encode the leaf number and tdx_module_args.version into RAX in assembly, because this is the place where the "fn" and struct tdx_module_args fields are marshaled into registers. AI was used under supervision to review code and workshop logs. In particular, it helped evaluate the assembly changes. Signed-off-by: Xu Yilun <yilun.xu@linux.intel.com> Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com> Reviewed-by: Kiryl Shutsemau (Meta) <kas@kernel.org> Reviewed-by: Nikolay Borisov <nik.borisov@suse.com> Reviewed-by: Tony Lindgren <tony.lindgren@linux.intel.com> Reviewed-by: Rick Edgecombe <rick.p.edgecombe@intel.com> Reviewed-by: Kishen Maloor <kishen.maloor@intel.com> Link: https://lore.kernel.org/kvm/4f4b0f29-424b-45ed-8cfd-c77da2ea390f@intel.com/ # [1] Link: https://patch.msgid.link/20260921-seamcall-version-v7-1-cf05fe76b467@linux.intel.com
11 daysMerge tag 'x86-urgent-2026-09-20' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull x86 fixes from Ingo Molnar: - Reject the loading of a potentially problematic microcode version on Intel Granite Rapids systems (Chang S. Bae) - On FRED, reconstruct the proper #GP context for rejected INT instructions, to fix a signal ABI regression (Matthew Schwartz) - Add a test for this signal ABI regression the x86 self-test suite (Matthew Schwartz) - Don't emit the new and not yet properly supported EGPR instructions (%r16-%r31) on CONFIG_X86_NATIVE_CPU=y builds (Chang S. Bae) * tag 'x86-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: x86/build/64: Prevent native builds from generating EGPR use selftests/x86: Check signal state for rejected software interrupts x86/fred: Reconstruct the #GP context for rejected INT instructions x86/microcode/intel: Reject problematic loading on Granite Rapids systems
13 daysx86/cpufeatures: Add X86_FEATURE_RMPOPT feature flagAshish Kalra
Add a flag indicating whether RMPOPT instruction is supported. RMPOPT is a new instruction that reduces the performance overhead of RMP checks for the hypervisor and non-SNP guests by allowing those checks to be skipped when 1-GB memory regions are known to contain no SEV-SNP guest memory. For more information on the RMPOPT instruction, see the AMD64 RMPOPT technical documentation. [ bp: Zap respective tools/ change. ] Suggested-by: Borislav Petkov (AMD) <bp@alien8.de> Signed-off-by: Ashish Kalra <ashish.kalra@amd.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Reviewed-by: Dave Hansen <dave.hansen@linux.intel.com> Reviewed-by: Ackerley Tng <ackerleytng@google.com> Reviewed-by: Tom Lendacky <thomas.lendacky@amd.com> Link: https://patch.msgid.link/39e9ee269a572c516a3f4e937bfe12d00697d5e6.1782841284.git.ashish.kalra@amd.com
13 daysACPI: CPPC: Validate FFH register fields before hardware accessChristian Loehle
The x86 CPPC FFH accessors construct masks and shift values using the firmware's Bit Width and Bit Offset without checking that the field fits in an MSR. A zero-width field or a field extending beyond bit 63 can therefore cause an invalid shift. The safe MSR accessors only handle an access fault, not invalid field arithmetic after a successful read. Also, the 64-bit GAS address is implicitly narrowed to the 32-bit MSR number. A descriptor with nonzero upper address bits can access a different MSR from the one described by firmware. The arm64 AMU counter readers have the same unchecked field arithmetic. Validate the field bounds before reading a counter, including both descriptors in the paired counter path. Dispatch FFH accesses before decoding the GAS Access Size in the common read and write paths. Otherwise, a large access_width can trigger an invalid shift in GET_BIT_WIDTH() before the architecture validates its register. That field has architecture-specific FFH semantics and is not needed by the generic memory/port accessor on this path. Validate x86 MSR addresses and share the field-bounds checks between related accessors within each architecture. Fixes: a6cbcdd5ab5f ("ACPI / CPPC: Add support for functional fixed hardware address") Fixes: 68c5debcc06d ("arm64: implement CPPC FFH support using AMUs") Fixes: f489c948028b ("ACPI: CPPC: Fix access width used for PCC registers") Cc: All applicable <stable@vger.kernel.org> Signed-off-by: Christian Loehle <christian.loehle@arm.com> Tested-by: Sumit Gupta <sumitg@nvidia.com> Link: https://patch.msgid.link/20260916162805.1039247-17-christian.loehle@arm.com Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
2026-09-17x86/microcode/intel: Reject problematic loading on Granite Rapids systemsChang S. Bae
Microcode updates can usually jump revisions. However, there is an erratum on Granite Rapids systems. If they "jump over" revision 0x1000405, they result in an #MC. Avoid it. Signed-off-by: Chang S. Bae <chang.seok.bae@intel.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Reviewed-by: Dave Hansen <dave.hansen@linux.intel.com> Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260916225939.1144524-1-chang.seok.bae@intel.com
2026-09-17x86/nmi: Enable panic() when NMI handler runs too longPaul E. McKenney
Firmware, device-driver, and configuration issues can result in long-running NMI handlers, for example, nmi_cpu_backtrace_handler() running for more than two seconds, and with no console issues. In this case, this is not the fault of nmi_cpu_backtrace_handler(), but it is the first symptom of the underlying problem. These NMI handlers result in various other problems, including CSD-lock warnings, clocksource watchdog false positives, scheduler frequency invariance wobbliness, systemd excessive-CPU complaints, and even the occasional RCU CPU stall warning. Eventually, the automation will take the system out of service, but with quite a bit of noise obscuring the underlying issue, and thus confusing both automation and human beings. Therefore, provide a toolong_nmi_panic __setup() parameter that makes nmi_check_duration() invoke nmi_panic() if the just-completed NMI handler ran for more than the specified number of milliseconds. This parameter defaults to zero, which disables the nmi_panic(). As in, by default you get the same behavior that you get today. [ paulmck: Apply kernel test robot feedback. ] Signed-off-by: Paul E. McKenney <paulmck@kernel.org> Cc: Thomas Gleixner <tglx@kernel.org> Cc: Ingo Molnar <mingo@redhat.com> Cc: Borislav Petkov <bp@alien8.de> Cc: Dave Hansen <dave.hansen@linux.intel.com> Cc: "H. Peter Anvin" <hpa@zytor.com> Cc: <x86@kernel.org>
2026-09-17x86/kprobes: Fix crash when probing CS CALL instructionsJinke Han
When using eBPF to probe CS CALL instructions within a function, a crash can be triggered. The eBPF tool probes offset 257 of the __hrtimer_run_queues() function: <__hrtimer_run_queues+249>: nopl 0x0(%rax,%rax,1) <__hrtimer_run_queues+254>: mov %r14,%rdi <__hrtimer_run_queues+257>: cs call <__x86_indirect_thunk_r12> <__hrtimer_run_queues+263>: mov %eax,%r12d <__hrtimer_run_queues+266>: xchg %ax,%ax <__hrtimer_run_queues+268>: mov %r13,%rdi Which triggers this crash: BUG: unable to handle page fault for address: 00000000000f41c9 #PF: supervisor write access in kernel mode #PF: error_code(0x0002) - not-present page PGD 0 P4D 0 Oops: 0002 [#1] SMP NOPTI CPU: 1 PID: 0 Comm: swapper/1 Kdump: loaded Tainted: P RIP: 0010:__hrtimer_run_queues+0x106/0x230 Note that __hrtimer_run_queues+0x106 is __hrtimer_run_queues+262, which is at the 6th byte of the above CS CALL instruction. Since the CS CALL instruction occupies 6 bytes, the exception occurred in the middle of that call instruction. The root cause is that when using eBPF tools to probe in the middle of a function, a kprobe with INT3 is used as the underlying implementation. During single-step emulation of the original CALL instruction, int3_emulate_call() assumes that the probed CALL instruction is 5 bytes long. However, the actual CS-prefixed CALL instruction occupies 6 bytes, so it constructs an incorrect exception return address. When the CPU returns from the kprobe handler, the next instruction to be executed is at the address of the last byte of that CS CALL instruction. Coincidentally, starting from that address, the CPU fetches and decodes a completely different instruction, which ultimately triggers a kernel crash. Fix the issue by using the actual instruction length obtained from the instruction decoder when constructing the exception return address, rather than relying on the hardcoded CALL_INSN_SIZE macro. [ mingo: Refined the changelog ] Fixes: 6256e668b7af ("x86/kprobes: Use int3 instead of debug trap for single-step") Suggested-by: Masami Hiramatsu (Google) <mhiramat@kernel.org> Signed-off-by: Jinke Han <jinkehan@didiglobal.com> Signed-off-by: Ingo Molnar <mingo@kernel.org> Reviewed-by: Masami Hiramatsu (Google) <mhiramat@kernel.org> Acked-by: Yafang Shao <laoar.shao@gmail.com> Acked-by: Borislav Petkov <bp@alien8.de> Cc: Peter Zijlstra <peterz@infradead.org> Link: https://patch.msgid.link/20260908073742.GA10517@didi-ThinkCentre-M920t-N000
2026-09-17x86/cpu: Don't transiently clear the boot CPU's capabilitiesIhor Solodrai
On the boot CPU, identify_cpu() runs from arch_cpu_finalize_init(), with interrupts enabled and before alternatives are patched. So cpu_feature_enabled() still evaluates against boot_cpu_data. identify_cpu() rebuilds c->x86_capability from scratch: the reset zeroes the array and the CPUID rescan fills it in again. An interrupt delivered in that window finds X86_FEATURE_LA57 clear in boot_cpu_data, so pgtable_l5_enabled() is false and KASAN checks a 5-level address against the 4-level addressability limit. The result is a bogus "wild-memory-access" report, and under kasan_multi_shot a report storm that wedges the boot. The boot CPU has already been scanned by early_identify_cpu(), with interrupts disabled, and its capabilities cannot have changed since. Reset only the CPUs which have not been scanned yet. The window is as old as identify_cpu() rebuilding the capabilities. Commit 39b9552281ab ("x86/mm: Optimize boot-time paging mode switching cost") merely let KASAN notice it by making pgtable_l5_enabled() read the feature bit. So no Fixes: tag. Closes: https://lore.kernel.org/bpf/20260610175651.647515-1-ihor.solodrai@linux.dev/ Signed-off-by: Ihor Solodrai <ihor.solodrai@linux.dev> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Link: https://patch.msgid.link/20260916195203.1099646-6-ihor.solodrai@linux.dev
2026-09-17x86/cpu: Move 32-bit SEP setup into identify_cpu()Ihor Solodrai
identify_boot_cpu() and identify_secondary_cpu() both call enable_sep_cpu() under CONFIG_X86_32 immediately after identify_cpu(). Do it once and drop the ifdefs while at it. No functional changes. Signed-off-by: Ihor Solodrai <ihor.solodrai@linux.dev> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Reviewed-by: Nikolay Borisov <nik.borisov@suse.com> Link: https://patch.msgid.link/20260916195203.1099646-5-ihor.solodrai@linux.dev
2026-09-17x86/cpu: Inline generic_identify() into identify_cpu()Ihor Solodrai
generic_identify() has exactly one call site: at the top of identify_cpu(). Fold it into identify_cpu() so that a single function does the job for both the boot CPU and the secondary CPUs. While at it, fix up both copies of the Cyrix comment. No functional changes. Signed-off-by: Ihor Solodrai <ihor.solodrai@linux.dev> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Link: https://patch.msgid.link/20260916195203.1099646-4-ihor.solodrai@linux.dev
2026-09-17x86/cpu: Initialize boot CPU cpuinfo defaults earlyIhor Solodrai
early_identify_cpu() clears the capability array, the CPUID table and extended_cpuid_level, but the architectural defaults for the rest of struct cpuinfo_x86 are set only later, in identify_cpu(). Use the same defaults from the start, so that the boot CPU does not depend on a later reset to end up with the right ones. Signed-off-by: Ihor Solodrai <ihor.solodrai@linux.dev> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Link: https://patch.msgid.link/20260916195203.1099646-3-ihor.solodrai@linux.dev