| Age | Commit message (Collapse) | Author |
|
https://git.kernel.org/pub/scm/linux/kernel/git/driver-core/driver-core.git
|
|
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git
# Conflicts:
# Documentation/scheduler/index.rst
# arch/arm64/configs/defconfig
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/modules/linux.git
|
|
* misc: (44 commits)
KVM: selftests: Add coverage for the EFER_LMSLE_MBZ defeature
KVM: selftests: Add module param API to check if nested virtualization is enabled
KVM: selftests: Rename svm_nested_clear_efer_svme to svm_nested_efer_test
KVM: x86: Honor the guest's EFER_LMSLE_MBZ
KVM: x86: Advertise EFER_LMSLE_MBZ when KVM disallows EFER.LMSLE
KVM: SVM: Add support for virtualizating Bus Lock Detect
KVM: SVM: Add a vCPU-aware helper to get supported DEBUGCTL bits
KVM: nSVM: Don't assume all active-low bits DR6 are fixed-1
KVM: nSVM: Open code check on LBR virtualization being enabled in vmcb12
KVM: nSVM: Disable LBRV in nested control cache when unsupported
KVM: SVM: Add helper to query if LBR virtualization needs to be enabled
KVM: x86: Kill off DR6_FIXED_1 to prevent future misuse
KVM: x86: Set guest's fixed-1 DR6 bits when delivering #DB payload
KVM: x86: Force fixed-1 bits in DR6 after synchronizing with hardware
KVM: Return an "unsigned long", not "u64" for the fixed-1 DR6 bits
KVM: x86: Rename kvm_dr6_fixed() => kvm_get_dr6_fixed_1()
KVM: x86: Preserve DR6.BLD (Bus Lock Detect) when delivering #DB payload
KVM: x86: Restrict saved GPA writes to hardware write faults
KVM: x86/pmu: Don't retry a counter whose config was rejected
KVM: SVM: Add Page modification logging support
...
|
|
# New commits in x86/sgx:
993b65a2fc6f ("x86/sgx: Report RCU-Tasks quiescent state in EPC sanitization loop")
Signed-off-by: Ingo Molnar <mingo@kernel.org>
|
|
# New commits in x86/sev:
250ee734115b ("x86/sev: Report MSR_AMD64_SEV in sysfs")
edb1b49e724e ("x86/sev: Do RMP optimizations on SNP guest shutdown")
c6c6f156d192 ("x86/sev: Perform RMP optimizations asynchronously")
15d974be0a3b ("x86/sev: Initialize RMPOPT configuration MSRs")
674bb17a32c7 ("x86/sev: Disable CPU hotplug while SNP is active")
1ee97d0f7b20 ("x86/cpufeatures: Add X86_FEATURE_RMPOPT feature flag")
Signed-off-by: Ingo Molnar <mingo@kernel.org>
|
|
# New commits in x86/microcode:
cf086537af4e ("x86/microcode/intel: Refresh old_microcode defines with the May 2026 release")
Signed-off-by: Ingo Molnar <mingo@kernel.org>
|
|
# New commits in x86/cpu:
2711d67bc3a7 ("x86/cpu: Don't transiently clear the boot CPU's capabilities")
db349783ca7a ("x86/cpu: Move 32-bit SEP setup into identify_cpu()")
ad86fe2134cc ("x86/cpu: Inline generic_identify() into identify_cpu()")
7a762b51a149 ("x86/cpu: Initialize boot CPU cpuinfo defaults early")
54a2882ab9f7 ("x86/cpu: Factor init_cpu_info() out of identify_cpu()")
228200f695c0 ("x86/CPU/AMD: Fix Zen5 TLB sizes reporting")
9025f3bf73f0 ("x86/shstk: Return the correct error value when user shadow stacks are disabled")
Signed-off-by: Ingo Molnar <mingo@kernel.org>
|
|
# New commits in x86/cleanups:
2c6a75adb15f ("x86/um: Remove unused <asm/required-features.h> header")
c7243accb48a ("x86/cpu: Constify struct x86_cpu_id")
975809caf4b0 ("x86/asm/string_64: Remove unused linux/jump_label.h include")
c8beda268a6b ("x86/mtrr: Fix kernel-doc notation of amd_set_mtrr()")
Signed-off-by: Ingo Molnar <mingo@kernel.org>
|
|
# New commits in x86/cache:
a17d099481bd ("x86/resctrl: Update documented unit for the "activity" event")
a3f5e6eba418 ("fs/resctrl: Simplify pseudo_lock_measure_trigger()")
4b31656d917c ("fs/resctrl: Avoid extra call to strlen() in schemata_list_add()")
942135446059 ("fs/resctrl: Factor MBA parse-time conversion to be per-arch")
255f92316ac1 ("arm_mpam: resctrl: Add pass-through resctrl_arch_preconvert_bw()")
57a354d20a01 ("x86,fs/resctrl: Add resctrl_arch_preconvert_bw()")
Signed-off-by: Ingo Molnar <mingo@kernel.org>
|
|
# New commits in x86/bugs:
ae1d2082d93b ("x86/bugs: Adapt SRSO mitigation to Zen6")
Signed-off-by: Ingo Molnar <mingo@kernel.org>
|
|
# New commits in sched/core:
1fb28c664a19 ("virt/steal_governor: Enable the driver")
27d47ebce4d6 ("virt/steal_governor: Implement steal_governor policy loop")
4b9302d494ff ("virt/steal_governor: Add control knobs for handling steal values")
9a8e740ee9f6 ("virt: Introduce steal governor driver")
68957caaa9c0 ("sched/debug: Add migration stats due to non preferred CPUs")
74699f56ebcf ("sched/core: Push current task from non preferred CPU")
4ee29b029058 ("sched/fair: Load balance only among preferred CPUs")
d8a3da0de843 ("sched/core: Try to use a preferred CPU in is_cpu_allowed")
620824516557 ("sysfs: Add preferred CPU file")
518b32bd5bb3 ("cpumask: Introduce cpu_preferred_mask")
06a49ef784ac ("sched/docs: Document cpu_preferred_mask and Preferred CPU concept")
cfb463b7172d ("cpumask: Introduce cpumask_intersects_and")
a8d0854a76a8 ("sched/cputime: Add kcpustat_field_total helper")
be100c77178e ("sched: Add sched_ext hooks for proxy execution")
57c75e3ae38c ("sched: Add helper to block retained proxy donors")
a49653d0abeb ("sched/core: Mark wakeups completed through ttwu_runnable()")
8f8c0417e973 ("sched/core: Dequeue waking proxy donors before reset")
313b652837d0 ("sched/core: Drop mutex locks before proxy rescheduling")
627ea30aca3b ("sched/wait: Clarify WF_SYNC wakeup semantics")
d2e010082757 ("sched/eevdf: Handle more short slice waking cases")
4bf32ec3327d ("sched/eevdf: Align update_protect_slice to set_protect_slice")
aae2a33ea662 ("sched/eevdf: Ensure that vprot will never go above a min slice")
c9ce69fc43bd ("sched/fair: Randomize equally shallow slow-path candidates")
abe440b3770f ("sched/fair: Drop idle recency from slow-path CPU selection")
fbbc63fed0b0 ("sched/core: Remove redundant core_sched_seq")
819224e506bc ("sched/fair: Remove dead code on enqueue_task_fair()")
c72945693b90 ("sched: Restart fair hrtick after same-task repicks")
a9b3c7570564 ("sched/headers: Replace __ASSEMBLY__ with __ASSEMBLER__ in the <uapi/linux/sched.h> header")
e81ee0630837 ("sched/fair: Reset NUMA fault locality after scan period update")
ef9293b3b797 ("sched: dynamic: Fix preemption model strings")
879eaa76e608 ("sched: Remove unneeded function type cast in do_balance_callbacks()")
f549101187c8 ("sched/deadline: check start_dl_timer expiry with ktime_before()")
2a672daa4b27 ("sched/feat: Use the new static key API for sched_feat")
a5576ebce920 ("sched: Convert paravirt_steal to new static key APIs")
9650ce11f2e3 ("sched: dynamic: Simplify preempt model accessors")
5b9a28eeed37 ("sched: dynamic: Remove HAVE_PREEMPT_DYNAMIC_{CALL,KEY}")
aa4178f63847 ("sched: dynamic: Simplify irqentry_exit_cond_resched()")
b9d267b9d632 ("sched: dynamic: Simplify preempt_schedule{,_notrace}()")
88e0b3bb9930 ("sched: dynamic: Simplify {cond,might}_resched()")
d3d16750693b ("sched: dynamic: Make PREEMPT_DYNAMIC depend on ARCH_HAS_PREEMPT_LAZY")
772d9ffbfd26 ("sched: Migrate whole chain in proxy_migrate_task()")
6b73a09e943f ("sched: Break out core of attach_tasks() helper into sched.h")
1f8805138593 ("sched: Switch rq->next_class in proxy_reset_donor()")
09351db90a28 ("sched/core: Don't proxy-exec unmatched cookie lock owners")
9be817f991e2 ("sched/core: Avoid migrating blocked_on tasks")
3dd95f077371 ("sched/core: Don't steal a proxy-exec donor")
Signed-off-by: Ingo Molnar <mingo@kernel.org>
|
|
Page modification logging(PML) is a hardware feature designed to track
guest modified memory pages. PML enables the hypervisor to identify which
pages in a guest's memory have been changed since the last checkpoint or
during live migration.
The PML feature is advertised via CPUID leaf 0x8000000A ECX[4] bit.
Acked-by: Borislav Petkov (AMD) <bp@alien8.de>
Signed-off-by: Nikunj A Dadhania <nikunj@amd.com>
Link: https://patch.msgid.link/20260907063906.1964557-7-nikunj@amd.com
Signed-off-by: Sean Christopherson <seanjc@google.com>
|
|
Update the minimum expected revisions of Intel microcode based on the
microcode-20260512 (May 2026) release.
Note, the three new entries are for INTEL_PANTHERLAKE_L steppings.
Signed-off-by: Sohil Mehta <sohil.mehta@intel.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Link: https://patch.msgid.link/20260928222607.3593973-1-sohil.mehta@intel.com
|
|
We need the driver-core fixes in here as well to build on top of.
Signed-off-by: Danilo Krummrich <dakr@kernel.org>
|
|
'struct x86_cpu_id' is not modified in this compilation unit.
Constifying this structure moves some data to a read-only section, so
increases overall security.
It is only used in cpu_has_old_microcode() which is an __init function. So,
using __initconst is safe and will save about 6 kB of memory at runtime.
On a x86_64, with allmodconfig:
text data bss dec hex filename
55725 28774 512 85011 14c13 arch/x86/kernel/cpu/common.o.before
61464 22918 512 84894 14b9e arch/x86/kernel/cpu/common.o.after
[ bp: Massage commit message. ]
Signed-off-by: Christophe JAILLET <christophe.jaillet@wanadoo.fr>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Link: https://patch.msgid.link/f08d3a0e7aefe3bad66251927cbd03b8cbf9df73.1786311119.git.christophe.jaillet@wanadoo.fr
|
|
In exc_machine_check_user(), local_db_save() and local_db_restore() are
invoked in the outer entry stubs (DEFINE_IDTENTRY_MCE_USER,
DEFINE_FREDENTRY_MCE, and DEFINE_IDTENTRY_RAW), surrounding
exc_machine_check_user().
However, exc_machine_check_user() calls irqentry_exit_to_user_mode(), which
handles pending thread work and may schedule() if TIF_NEED_RESCHED is set. If
the task migrates to another CPU during schedule(), local_db_restore() runs on
the new CPU with the dr7 state saved from the old CPU. This corrupts the new
CPU's DR7 hardware debug register and leaves the old CPU's DR7 disabled. In
short, local_db_save() and local_db_restore() pair must be run on the same
CPU.
To fix this, move local_db_save() and local_db_restore() inside
exc_machine_check_user() and exc_machine_check_kernel(). In
exc_machine_check_user(), DR7 is saved and restored strictly around
do_machine_check() to avoid schedule() during migration. In
exc_machine_check_kernel(), local_db_save() is called at the entry point to
prevent early memory accesses from triggering nested #DB exceptions, and
restored on all exits.
Fixes: cd840e424f27 ("x86/entry, mce: Disallow #DB during #MC")
Assisted-by: LLM
Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Cc: <stable@kernel.org>
Link: https://patch.msgid.link/179005109564.388919.3937970081044095776.stgit@devnote2
|
|
The usermode helper declarations were previously provided by linux/kmod.h
but commit c1f3fa2a4fde ("kmod: split off umh headers into its own file")
moved them to linux/umh.h in 2017.
Add explicit includes of linux/umh.h to files that use usermode helpers and
remove linux/kmod.h where it is no longer needed.
Acked-by: Alex Elder <elder@riscstar.com> # for greybus
Signed-off-by: Petr Pavlu <petr.pavlu@suse.com>
|
|
Add a flag indicating whether RMPOPT instruction is supported.
RMPOPT is a new instruction that reduces the performance overhead of RMP
checks for the hypervisor and non-SNP guests by allowing those checks to be
skipped when 1-GB memory regions are known to contain no SEV-SNP guest memory.
For more information on the RMPOPT instruction, see the AMD64 RMPOPT
technical documentation.
[ bp: Zap respective tools/ change. ]
Suggested-by: Borislav Petkov (AMD) <bp@alien8.de>
Signed-off-by: Ashish Kalra <ashish.kalra@amd.com>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Reviewed-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Ackerley Tng <ackerleytng@google.com>
Reviewed-by: Tom Lendacky <thomas.lendacky@amd.com>
Link: https://patch.msgid.link/39e9ee269a572c516a3f4e937bfe12d00697d5e6.1782841284.git.ashish.kalra@amd.com
|
|
Microcode updates can usually jump revisions. However, there is an erratum on
Granite Rapids systems. If they "jump over" revision 0x1000405, they result in
an #MC. Avoid it.
Signed-off-by: Chang S. Bae <chang.seok.bae@intel.com>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Reviewed-by: Dave Hansen <dave.hansen@linux.intel.com>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260916225939.1144524-1-chang.seok.bae@intel.com
|
|
On the boot CPU, identify_cpu() runs from arch_cpu_finalize_init(), with
interrupts enabled and before alternatives are patched. So
cpu_feature_enabled() still evaluates against boot_cpu_data.
identify_cpu() rebuilds c->x86_capability from scratch: the reset zeroes the
array and the CPUID rescan fills it in again. An interrupt delivered in that
window finds X86_FEATURE_LA57 clear in boot_cpu_data, so pgtable_l5_enabled()
is false and KASAN checks a 5-level address against the 4-level addressability
limit. The result is a bogus "wild-memory-access" report, and under
kasan_multi_shot a report storm that wedges the boot.
The boot CPU has already been scanned by early_identify_cpu(), with interrupts
disabled, and its capabilities cannot have changed since. Reset only the CPUs
which have not been scanned yet.
The window is as old as identify_cpu() rebuilding the capabilities. Commit
39b9552281ab ("x86/mm: Optimize boot-time paging mode switching cost")
merely let KASAN notice it by making pgtable_l5_enabled() read the feature
bit. So no Fixes: tag.
Closes: https://lore.kernel.org/bpf/20260610175651.647515-1-ihor.solodrai@linux.dev/
Signed-off-by: Ihor Solodrai <ihor.solodrai@linux.dev>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Link: https://patch.msgid.link/20260916195203.1099646-6-ihor.solodrai@linux.dev
|
|
identify_boot_cpu() and identify_secondary_cpu() both call
enable_sep_cpu() under CONFIG_X86_32 immediately after identify_cpu().
Do it once and drop the ifdefs while at it.
No functional changes.
Signed-off-by: Ihor Solodrai <ihor.solodrai@linux.dev>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Reviewed-by: Nikolay Borisov <nik.borisov@suse.com>
Link: https://patch.msgid.link/20260916195203.1099646-5-ihor.solodrai@linux.dev
|
|
generic_identify() has exactly one call site: at the top of
identify_cpu(). Fold it into identify_cpu() so that a single function
does the job for both the boot CPU and the secondary CPUs.
While at it, fix up both copies of the Cyrix comment.
No functional changes.
Signed-off-by: Ihor Solodrai <ihor.solodrai@linux.dev>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Link: https://patch.msgid.link/20260916195203.1099646-4-ihor.solodrai@linux.dev
|
|
early_identify_cpu() clears the capability array, the CPUID table and
extended_cpuid_level, but the architectural defaults for the rest of struct
cpuinfo_x86 are set only later, in identify_cpu().
Use the same defaults from the start, so that the boot CPU does not
depend on a later reset to end up with the right ones.
Signed-off-by: Ihor Solodrai <ihor.solodrai@linux.dev>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Link: https://patch.msgid.link/20260916195203.1099646-3-ihor.solodrai@linux.dev
|
|
identify_cpu() unconditionally resets the struct cpuinfo_x86 fields to
their default values and clears the capability arrays with memset()
before rescanning the CPU to fill it in again.
However the boot CPU capabilities have already been scanned by
early_identify_cpu(), with interrupts disabled.
Introduce init_cpu_info() helper in preparation for letting the
callers decide whether the reset is needed.
No functional changes.
Signed-off-by: Ihor Solodrai <ihor.solodrai@linux.dev>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Link: https://patch.msgid.link/20260916195203.1099646-2-ihor.solodrai@linux.dev
|
|
The upcoming constification of the device_show_int() and
device_show_bool() signatures requires the users to handle the
transition automatically.
Switch to the __DEVICE_ATTR() macro which can do this.
Signed-off-by: Thomas Weißschuh <linux@weissschuh.net>
Link: https://patch.msgid.link/20260907-sysfs-const-attr-dev_ext_attr-v2-1-bf53afe57071@weissschuh.net
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
Add ':' to arguments in kernel-doc comment to fix:
$ ./scripts/kernel-doc -none arch/x86/kernel/cpu/mtrr/amd.c
Warning: arch/x86/kernel/cpu/mtrr/amd.c:63 function parameter 'reg' not described in 'amd_set_mtrr'
Warning: arch/x86/kernel/cpu/mtrr/amd.c:63 function parameter 'base' not described in 'amd_set_mtrr'
Warning: arch/x86/kernel/cpu/mtrr/amd.c:63 function parameter 'size' not described in 'amd_set_mtrr'
Warning: arch/x86/kernel/cpu/mtrr/amd.c:63 function parameter 'type' not described in 'amd_set_mtrr'
[ bp: Massage commit message. ]
Signed-off-by: Manuel Ebner <manuelebnerli@mailbox.org>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Link: https://patch.msgid.link/20260915103136.171924-3-manuelebnerli@mailbox.org
|
|
Starting with Zen5, TLB sizes in CPUID_Fn80000006_E[AB]X are reported
as multiples of 32. There's a CPUID bit which determines that:
CPUID_Fn80000021_EAX [Extended Feature 2 EAX] (Core::X86::Cpuid::FeatureExt2Eax)
...
14: L2TlbSizeX32. Read-only. Reset: 1. Indicates that L2TLB sizes are encoded as multiples of 32.
Update the places which report that information.
With it, the numbers look correct now:
-Last level iTLB entries: 4KB 64, 2MB 64, 4MB 32
-Last level dTLB entries: 4KB 128, 2MB 128, 4MB 64, 1GB 0
+Last level iTLB entries: 4KB 2048, 2MB 2048, 4MB 1024
+Last level dTLB entries: 4KB 4096, 2MB 4096, 4MB 2048, 1GB 0
/proc/cpuinfo
-TLB size : 192 4K pages
+TLB size : 6144 4K pages
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Link: https://lore.kernel.org/r/20260821022031.946311-1-bp@kernel.org
|
|
On MPAM systems the rounding behaviour of the MBA control would be improved if
the rounding in the fs/resctrl code is removed but this is not the case for
x86. To allow any rounding or conversion of the bandwidth value provided by
the user to be specified by the arch code a new arch hook is required.
Introduce resctrl_arch_preconvert_bw(), and add its x86 implementation.
This is currently unused in resctrl but when plumbed in it will replace the
call to roundup() in bw_validate().
Signed-off-by: Dave Martin <dave.martin@arm.com>
Signed-off-by: Ben Horgan <ben.horgan@arm.com>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Reviewed-by: Reinette Chatre <reinette.chatre@intel.com>
Reviewed-by: Gavin Shan <gshan@redhat.com>
Link: https://patch.msgid.link/20260911163613.1131447-2-ben.horgan@arm.com
|
|
mce_amd_handle_storm() currently does the opposite of what storm
handling needs: it enables thresholding interrupts when a storm is
detected and disables them when the storm subsides.
Flip the "on" function argument before passing it to threshold_restart_bank()
as it should have been done.
To clarify: "on" to mce_handle_storm() means, the storm is on now when
"on" is true, and off when "on" is false.
[ bp: Simplify. ]
Fixes: 5c4663ed1eac ("x86/mce: Handle AMD threshold interrupt storms")
Signed-off-by: Jasjeet Rangi <jrangi@purestorage.com>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260812221514.598842-2-jrangi@purestorage.com
|
|
When the kernel boots from kexec, the EPC pages may have a stale state.
The kernel sanitizes all EPC pages to reset them to a clean state before
their first use in any enclave. The EPC size could be several GBs and
resetting them could take a significant amount of time. Because of that,
the kernel performs the reset in a loop through a kernel thread ksgxd() at
early boot, and there's a cond_resched() after resetting each EPC page.
This is fine in most cases, but becomes a problem when there's other kernel
code waiting for an RCU-Tasks grace period but the cond_resched() in
ksgxd() never triggers rescheduling. Because cond_resched() doesn't report
a quiescent state when it doesn't trigger rescheduling, the thread that is
waiting for an RCU-Tasks grace period will wait until all EPC pages are
reset.
For instance, BPF LSM subsystem can invoke synchronize_rcu_tasks() at
kernel boot time. A VM with a large EPC assigned and BPF LSM enabled can
take a long time to boot, with a call trace triggered:
rcu_tasks_wait_gp: rcu_tasks grace period number 1 (since boot) is
130631 jiffies old.
INFO: task systemd:1 blocked for more than 122 seconds.
...
task:systemd state:D stack:0 pid:1 tpid:1 ppid:0 flags:0x00000002
Call Trace:
...
schedule_timeout+0x157/0x170
wait_for_completion+0x88/0x150
__wait_rcu_gp+0x17e/0x190
synchronize_rcu_tasks_generic+0x64/0x60
...
synchronize_rcu_tasks+0x15/0x20
register_ftrace_direct+0x31f/0x350
...
bpf_trampoline_link_prog+0x33/0x60
bpf_tracing_prog_attach+0x3c5/0x5f0
Replace cond_resched() with cond_resched_tasks_rcu_qs() which explicitly
reports quiescent state regardless of whether actual rescheduling is
triggered. Resetting all EPC pages in ksgxd() isn't performance critical
so the extra cost of cond_resched_tasks_rcu_qs() isn't a problem.
Tests showed this reduced the VM kernel boot time from ~50s to ~700ms.
Co-developed-by: Fan Du <fan.du@intel.com>
Fixes: e7e0545299d8 ("x86/sgx: Initialize metadata for Enclave Page Cache (EPC) sections")
Suggested-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Fan Du <fan.du@intel.com>
Signed-off-by: Jun Miao <jun.miao@intel.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Reviewed-by: Jarkko Sakkinen <jarkko@kernel.org>
Tested-by: Challvy Tee <challvy.tee@gmail.com>
Link: https://github.com/systemd/systemd/issues/40423
Link: https://patch.msgid.link/20260902013653.690506-1-jun.miao@intel.com
|
|
paravirt_steal_rq_enabled and paravirt_steal_enabled use raw static_key
APIs which are now deprecated. Use the new API instead.
No functional change.
Signed-off-by: Hongyan Xia <hongyan.xia@transsion.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Acked-by: Juergen Gross <jgross@suse.com>
Link: https://patch.msgid.link/20260819081207.12150-1-hongyan.xia@transsion.com
|
|
Zen6 has BTB protection which isolates the different contexts
(user/kernel, guest/host) from one another. This makes the SafeRET
mitigation there unnecessary leaving the user/user and guest/guest
attack vectors open, whose protection is handled by the Spectre v2
mitigation setting to do IBPB on a context switch.
Detect that setting and report it with a new mitigation string.
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Link: https://patch.msgid.link/20260822013231.1109255-1-bp@kernel.org
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/hyperv/linux
Pull hyperv updates from Wei Liu:
- Decrypt netvsc buffer on contiguous direct-map addresses (Kameron
Carr)
- Drop WS2012/2012R2 & Win8/8.1 Hyper-V support (Michael Kelley)
- Use more meaningful errnos for hypercall status code (Hardik Garg)
- Fix lost interrupts on CPU hot-unplug for Hyper-V PCI/MSI (Naman
Jain)
- Reserve more MSHV vectors for Linux root partition (Wei Liu)
* tag 'hyperv-next-signed-20260826' of git://git.kernel.org/pub/scm/linux/kernel/git/hyperv/linux:
clocksource: hyper-v: Remove support for stimer interrupts in message mode
scsi: storvsc: Remove support for storvsc protocol of old Hyper-V hosts
hv_netvsc: Remove GPADL teardown special case for old Hyper-V hosts
hv_sock: Remove check for old Hyper-V hosts
Drivers: hv: Remove support for WS2012/2012R2 & Win8/8.1 version of Hyper-V
hv_netvsc: Allocate send/receive buffers using vmbus_alloc_buffer()
Drivers: hv: vmbus: Add vmbus_alloc_buffer()/vmbus_free_buffer() for CoCo VMs
Drivers: hv: vmbus: add vmbus_establish_gpadl_caller_decrypted()
Drivers: hv: vmbus: Skip VMBus module cleanup for non-nested root partition
x86/hyperv: reserve more vectors
PCI: hv: Set irq_retrigger callback for the Hyper-V PCI MSI irqchip
Drivers: hv: Use meaningful errnos for hypercall status codes
|
|
In Hyper-V versions prior to WS2016/Win10, Hyper-V synthetic timers
interrupt the guest by delivering a message that is initially handled
by the Linux VMBus driver. Starting with WS2016/Win10, Hyper-V can
deliver stimer interrupts directly to an assigned interrupt vector
without involving the VMBus driver. This is called "Direct Mode".
With the overall removal of Linux support for running on Hyper-V
hosts earlier than WS2016 and Windows 10, it's no longer necessary
to support the legacy message-based delivery. Remove that delivery
mechanism and always use Direct Mode. If for some reason, the
Hyper-V host does not enumerate Direct Mode, output an error
message but continue to run using the LAPIC timer instead of an
stimer.
With these changes, the VMBus driver no longer calls the stimer
interrupt service routine. This removal has a broader benefit in
unblocking the disentangling of VMBus code and stimer code, as
they should be independent of each other. The final disentangling
will come as a follow-on patch set.
Signed-off-by: Michael Kelley <mhklinux@outlook.com>
Signed-off-by: Wei Liu <wei.liu@kernel.org>
|
|
Microsoft Hypervisor delivers three vectors to the NT HAL running in the
root partition and refuses to map a device interrupt to any of them when
interrupt remapping is not available in the system. As of writing, the
nested MSHV setup has no interrupt remapping capability.
The three vectors are:
HAL_NT_APC_VECTOR 0x1F
HAL_NT_DPC_VECTOR 0x2F
HAL_NT_CLOCK_IPI_VECTOR 0xD2
0x1F is below FIRST_EXTERNAL_VECTOR so the vector allocator never hands
it out, but 0x2F and 0xD2 are both inside the allocatable range and are
handed out once enough vectors are in use. Mapping such an interrupt
then fails with HV_STATUS_INVALID_PARAMETER, and the interrupt is never
delivered.
Reserve all three next to the hypervisor debug vectors that are already
kept out of the allocator's hands.
Reviewed-by: Michael Kelley <mhklinux@outlook.com>
Signed-off-by: Wei Liu <wei.liu@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
Pull MM updates from Andrew Morton:
- "mm: drop "sub" prefix from various places" (Dev Jain)
page->folio conversion and a naming cleanup
- "mm/kasan: remove redundant initialization for kasan_flag_write_only"
(Igor Putko)
KASAN cleanup work
- "mm/filemap: reduce unnecessary xarray lookups" (Chi Zhiling)
Small speedup in the pagecaache read code
- "mm/percpu: Fix possible NOFS/NOIO reclaim recursion" (Kaitao Cheng)
Improve the vmalloc code - mainly the avoidance of GFP_KERNEL
allocations when the caller asked for GFP_NOFS or GFP_NOIO
- "mm/kmemleak: avoid soft lockup when scanning task stacks" (Breno
Leitao)
Avoid a soft lockup watchdog trigger from the kmemleak scanning code
in extreme situations
- "mm/page_owner: misc cleanups" (Ye Liu)
Cleanups to the page_owner code. For some reason lots of people have
been working on the page_owner code this cycle.
- "mm: convert to walk_page_range_vma() to eliminate find_vma()"
(Kefeng Wang)
Simplify and accelerate the page walking library function
- "mm/migrate: preparatory cleanups for batch copy and offload"
(Shivank Garg)
Cleanups in the migration code
- "mm/page_owner: add per-fd filter infrastructure for print_mode and
NUMA filtering" (Zhen Ni)
Per-fd filtering to page_owner in order to reduce the sometimes vast
amount of output it can produce
- "mm: Refactor bootmem gigantic hugepage allocation" (Muchun Song)
Fixes and preparatory cleanups around bootmem HugeTLB handling,
sparse initialization ordering, and related vmemmap setup
- "mm/zsmalloc: reduce lock contention in zs_free()" (Wenchao Hao)
Reduce lock contention in zs_free(), which dominates the unmap path
under memory pressure on Android (LMK kills) and on x86 servers
running zswap-heavy workloads.
Up to 1.83x improvement in microbenchmarking.
- "move alloc_tag.c file under mm/" (Suren Baghdasaryan)
- "samples/damon: handle damon_{start,stop}() failures" (SJ Park)
Fix improper handling of damon_start(), damon_stop(), and
damon_call() failures across DAMON sample modules to prevent
potential memory leaks, operation disruptions and use-after-free
bugs
- "mm/damon/sysfs: kobject_del() directories that users can
create/remove" (SJ Park)
Fix delayed sysfs directory removal under DEBUG_KOBJECT_RELEASE
causeing creation failures due to duplicate directory names by adding
missing kobject_del() calls before creating new directories
- "mm: cleanup clear_not_present_full_ptes()" (David Hildenbrand)
Clean up the core pte handling code
- "selftests/damon: misc fixes for test bugs" (Kunwu Chan)
Fix several bugs in the DAMON selftests
- "selftests/damon: fix memcg_path staging handling" (Cheng Nie)
Fix a bug in _damon_sysfs.py for damos_filter memcg_path setup, and
add a test case for it in sysfs.py.
- "selftests/damon: test kdamond refresh_ms" (Ruslan Valiyev)
Selftest coverage for DAMON's refresh_ms sysfs feature by updating
the test control module and verifying that scheme stats update
automatically without manual intervention
- "mm/damon: five misc fixups" (Akinobu Mita)
Miscellaneous DAMON fixups.
- "mm/damon/core: detect internal variation above max_nr_regions/2"
(Jiayuan Chen)
Fix DAMON's region splitting behavior when region counts exceed half
the maximum budget by dynamically scaling down the split fraction as
the limit approaches, preventing large regions from staying un-split,
and add corresponding KUnit test coverage
- "mm: preparatory patches for PMD level swap entries" (Usama Arif)
Refactor and clean up PMD softleaf helpers, call sites, and
architecture flags to lay the groundwork for a follow-up series that
introduces PMD page table swap entries
- "mm/damon: update, optimize, and clean up doc, tests, and code" (SJ
Park)
Update DAMON design and ABI documentation, expands unit and selftest
coverage, optimize damon_commit_target_regions(), and clean up
recently added sysfs interface code for better readability
- "mm/vmpressure: reduce CPU, memory and code overhead on cgroup v2"
(Usama Arif)
Optimize vmpressure() by skipping unnecessary work on cgroup v2 for
userspace event notifications and refactor v1-only eventfd handling
into mm/memcontrol-v1.c to reduce memory overhead and code complexity
- "selftests/mm: refactor pkey helpers and fix mmap error handling"
(Hongfu Li)
Refactor pkeys shared tracing and assertion helpers into a common
file, unify protection key selftests to use consistent diagnostic
logging and assertions, and enforce standardized MAP_FAILED return
checks for mmap() calls across the tests
- "mm/damon: optimize out nr_accesses_bp" (SJ Park)
Replace the error-prone, continuously updated nr_accesses_bp field in
damon_region with an on-demand moving sum function, reducing
structure memory overhead and avoiding state corruption bugs
- "Open HugeTLB allocation routine for more generic use" (Ackerley Tng)
Decouple HugeTLB folio allocation from VMA dependencies by
introducing hugetlb_alloc_folio(), enabling subsystems like
guest_memfd to allocate HugeTLB folios without standard VMA
reservations or pseudo-VMAs
- "mm/damon: provide pseudo moving sum probe_hits" (SJ Park)
Integrate DAMON's probe_hits attribute counter into the pseudo moving
sum infrastructure, enabling real-time, online monitoring without
waiting for full aggregation intervals
- "mm: Some cleanups for page allocator APIs" (Brendan Jackman)
Simplify and refactor the page allocator entry points and flags by
unifying allocation paths, adding internal alloc_flags arguments, and
eliminating redundant __ prefixed alloc_pages variants.
- "Fix incorrect access of hugetlb pte entries" (Dev Jain)
Enforce the consistent use of huge_ptep_get() instead of ptep_get()
for HugeTLB entries and fixes an unaligned address issue in arm64's
huge_ptep_get() implementation
- "mm/damon: validate all parameters in the core" (SJ Park)
Consolidate parameter validation into the DAMON core specifically
within damon_start() and damon_commit_ctx() to centralize error
checking, eliminate caller-side redundant checks and to improve
maintenance efficiency
- "tools/mm/page_owner_sort: fix filtering and cleanup issues" (Yichong
Chen)
Rename is_need() to filter_record() for clearer return semantics, fix
per-record allocation memory leaks and bound output copies in
search_pattern() to address an existing buffer issue
- "memcg: bail out reclaim when memcg is dying" (Jiayuan Chen)
Mitigate a system-wide stall which occurs when a cgroup is removed
while one of its memory control files is doing synchronous reclaim
- "mm/memory-failure: add panic option for unrecoverable pages" (Breno
Leitao)
Introduce an opt-in vm.panic_on_unrecoverable_memory_failure sysctl
that immediately panics the kernel on unrecoverable memory errors in
kernel-owned pages to preserve error context and prevent delayed,
silent data corruption
- "mm/damon: refactor damon_{start,stop,commit}() for simple error
handling" (SJ Park)
Refactor the DAMON core API functions to guarantee that all contexts
are fully stopped when damon_start(), damon_stop(), or damon_commit()
fail, eliminating the need for complex and error-prone caller-side
cleanup code
- "Keep tail page private zero at free and folio split" (Zi Yan)
Add checks to ensure tail_page->private is zero when freeing compound
or high-order pages and when promoting tail pages during large folio
splits. By validating these fields at free and split time, it allows
the removal of redundant private field clearing inside
prep_compound_tail()
- "mm: drop redundant lru_add_drain in anon folio reuse paths" (Barry
Song)
Eliminate redundant lru_add_drain() calls in
wp_can_reuse_anon_folio() and do_swap_page() to reduce LRU lock
contention and system overhead
By validating folio refcounts against the LRU cache before draining
and removing unnecessary drains in the swap path, it achieves up to a
30.5% reduction in drain calls during heavy swap workloads
- "mm: clean up folio LRU and swap declarations" (Jianyue Wu)
Reorganize folio LRU and swap code by relocating page-cluster state
to mm/swap_state.c, renaming mm/swap.c to mm/folio.c, and moving
MM-internal reclaim declarations into mm/internal.h.
- "userfaultfd: working set tracking for VM guest memory" (Kiryl
Shutsemau)
Add userfaultfd support for tracking the working set of VM guest
memory, so a VMM can identify hot pages and reclaim cold ones to
tiered or remote storage
- "mm: remove CONFIG_HAVE_BOOTMEM_INFO_NODE (Part 2)" (David
Hildenbrand)
Remove the remaining pieces of CONFIG_HAVE_BOOTMEM_INFO_NODE,
performing some smaller cleanups around freeing of reserved vmemmap
pages on the way.
- "mm/damon: update probe hits for runtime parameter commits" (SJ Park)
Ensure that DAMON's probe_hits attribute counter is properly updated
when monitoring intervals are changed at runtime, matching the
behavior of nr_accesses. To achieve this, it refactors and renames
existing helper functions for shared use, applies the updates to
probe_hits, and handles edge cases in damon_probe_hits_mvsum() to
maintain measurement accuracy.
- "KSM: performance optimizations for rmap_walk_ksm" (xu xin)
Resolve a severe KSM reverse-mapping performance bottleneck where
thousands of split VMAs sharing a single anon_vma cause extended lock
contention.
By adding an interval-filtering check during the rmap walk, it
reduces worst-case anon_vma lock hold times from over 500ms down to
under 2ms, preventing application freezes and latency spikes under
memory pressure.
- "mm: split a couple of headers from internal.h" (Mike Rapoport)
Split declarations related to mm_init, memblock, vmalloc and sparse
into new headers
- "KSM: use linear_page_index in collect_procs_ksm()" (xu xin)
Apply the interval tree optimization from rmap_walk_ksm() to
collect_procs_ksm() to avoid iterating over non-matching VMAs during
KSM memory error handling.
It hoists loop-invariant address initialization and restricts the
anon_vma_interval_tree_foreach walk to a targeted page offset range,
reducing redundant checks and improving lookup efficiency.
- "selftests/mm: avoid false failures in hugetlb and KSM tests" (Sayali
Patil)
Fix issues in the hugetlb and KSM MM selftest categories that can
report failures when the prerequisites for the tests are not
satisfied
- "mm/damon: introduce data attributes only monitoring" (SJ Park)
Introduce attribute-weighted region management in DAMON, allowing
users to prioritize specific data attributes (such as page sizes or
cgroups) over or instead of access monitoring.
By assigning weights to attribute probes, DAMON can completely
disable access tracking and adjust monitoring regions based on
weighted probe-hit counters to optimize monitoring quality for
attribute-focused workloads.
- "mm/hmm: Add mmap lock-drop support for userfaultfd-backed mappings"
(Stanislav Kinsburskii)
Extend hmm_range_fault() to support userfaultfd-backed regions by
allowing the mmap lock to be dropped during fault handling via a new
hmm_range_fault_locked() helper.
By accepting a locked pointer and signaling retry status when lock
release occurs, it enables page fault resolution in userfaultfd
regions while preserving backward compatibility for existing callers.
- "mm: make VMA page offset handling more consistent" (Lorenzo Stoakes)
Clean up and standardize how vma->vm_pgoff is accessed and
manipulated across file-backed and anonymous mappings in the kernel
It introduces dedicated helper functions such as vma_start_pgoff(),
vma_end_pgoff(), vma_set_pgoff() and linear_page_delta() while
renaming rmap interval tree helpers to better reflect their
functionality.
These changes establish a cleaner foundation for future work that
will unify virtual page offset indexing for all anonymous and CoW'd
folios.
- "mm: handle device-private PMDs in walk callbacks" (Usama Arif)
Address kernel panics and state corruption caused by MM walk
callbacks reaching non-present device-private PMD swap entries
created during HMM migrations
It ensures that functions which acquire pmd_trans_huge_lock()
properly recognize device-private PMDs instead of assuming a present
THP or a standard migration entry.
- "mm/rmap: Refactor try_to_unmap_one" (Dev Jain)
Refactor try_to_unmap_one by modularizing Hugetlb,
anonymous-lazyfree, and anonymous-swapbacked logic into dedicated
functions, laying the structural groundwork for batched anonymous
large folio unmapping.
- "Docs/ABI/damon: sysfs ABI document fixes and additions" (Song Hu)
Fix typos and fills in missing entries in the DAMON sysfs ABI
document
- "dax/kmem: atomic whole-device hotplug via sysfs" (Gregory Price)
Introduce an atomic sysfs state attribute and supporting DAX/MM
infrastructure to prevent userland races when offlining and removing
entire memory regions
By adding an unplugged state alongside standard online modes, it
enables whole-device atomic hotplug control while preserving backward
compatibility.
- "mm: convert more vm_flags_t users to vma_flags_t" (Lorenzo Stoakes)
Continue transitioning the kernel from the deprecated vm_flags_t type
to vma_flags_t across core memory management infrastructure.
It replaces legacy type usage in core functions such as do_mmap(),
unmapped area allocation, mm->def_vma_flags, and VMA operations like
mlock, mprotect, and mremap.
- "Two small patches to clean up mm/mm_slot.h" (xu xin)
Refactor mm_slot.h by introducing mm_slot_remove() to unify duplicate
slot deletion sequences in khugepaged and KSM. It also adds code
documentation explaining why mm_slot_lookup and mm_slot_insert must
remain as preprocessor macros rather than static inline functions.
- "mm/damon/core: hide core-private struct fields" (SJ Park)
Clean up DAMON core structures by consistently marking internal-only
fields with private: comment tags to prevent improper direct access
from outer layers.
It enforces encapsulation across core structures including
damon_region, damon_target, and damon_ctx and updates DAMON_SYSFS to
interact through approved access APIs instead of exposing raw struct
members.
- "mm/damon: unurgent fixes for infinite loop, NULL de-ref and races"
(SJ Park)
Address potential infinite loops, NULL dereferences, and race
conditions identified in DAMON
It fixes an infinite loop triggered by extreme user configurations, a
NULL pointer dereference within unit tests and minor monitoring
accuracy degradation caused by subtle runtime races.
- "mm/page_alloc: fixes for free_pages_nolock() on RT/UP" (Brendan
Jackman)
Fix an NMI safety flaw in __free_frozen_pages() where freeing pages
on non-SMP or PREEMPT_RT kernels can bypass can_spin_trylock() checks
via non-PCP or isolated migration paths.
It also resolves potential kernel crashes and privilege escalation
risks triggered when BPF tracing runs in NMI context alongside memory
hotplug or large allocation frees.
- "mm/page_alloc: couple of followups for recent cleanups" (Brendan
Jackman)
Clean up and update page allocator nomenclature, documentation, and
debug assertions.
It aligns internal FPI_ flags with the public "nolock" naming
convention, removes outdated internal implementation details from
high-level page allocator comments, and eliminates obsolete
VM_BUG_ON() assertions in allocation paths.
- "mm/mseal: further cleanups" (Lorenzo Stoakes)
Refactor and simplify the mseal implementation by clarifying API
boundaries and removing unnecessary code complexity.
It replaces generic do_mseal() usage outside the syscall with a
dedicated mseal_mmap_page_zero() helper for MMAP_PAGE_ZERO,
eliminates mm_struct parameters to enforce that sealing applies only
to current->mm, and streamlines overall logic and comments with no
functional changes intended.
- "mm/vmscan: fix swappiness=max and clean up per-node proactive
reclaim" (Ridong Chen)
Resolve reclaim behavior bugs and clean up function parameters across
memory reclaim paths
It fixes swappiness=max in both standard reclaim and MGLRU so
unswappable anonymous memory no longer falls back to evicting page
cache, ensures reclaim_store() returns accurate error codes instead
of collapsing all failures into -EAGAIN, and removes the obsolete
gfp_mask parameter from __node_reclaim().
- "mm: mincore: misc cleanups" (Kefeng Wang)
Clean up and simplifies the mincore code. Most importantly, it
removes the historical special behavior that always reports VM_PFNMAP
pages as non-resident.
- "mm/huge_memory: drop dead split helper variants" (Kiryl Shutsemau)
Two trivial cleanups in the folio split API
- "mm/damon: fix uninitialized DAMOS field and kunit exec expectation
bugs" (SJ Park)
Resolve minor operational and testing bugs in DAMON identified by
Sashiko. It initializes the damos->last_applied field to prevent
occasional efficiency degradation and fixes invalid memory accesses
in DAMON KUnit tests during test failure handling.
- "cleanup for stable_page_flags()" (Jinjiang Tu)
Clean up and refactor stable_page_flags() used by /proc/kpageflags
without altering functionality.
It uses BIT_ULL() to prevent shift-overflow warnings on 64-bit flag
bits, converts folio-specific flag checks to standard folio_test_*()
helpers, and removes redundant CONFIG_PAGE_IDLE_FLAG handling.
- "Batch unmap of uffd-wp file folios" (Dev Jain)
Extend batched folio unmapping support to file folios within
userfaultfd write-protect (uffd-wp) VMAs by adding batching
capabilities to pte_install_uffd_wp_if_needed().
This removes special-case restrictions on uffd-wp VMAs in
try_to_unmap_one(), significantly simplifying the function's control
flow and complexity.
- "mm/early_ioremap: clarify and clean up early_ioremap_reset()"
(Sang-Heon Jeon)
Clarify and clean up the architecture-specific usage of
__late_set_fixmap() and __late_clear_fixmap() after
early_ioremap_reset()
It adds explicit documentation regarding when early_ioremap_reset()
must be called and removes redundant macro definitions and reset
calls in the RISC-V and ARM64 architectures.
- "mm: fix reclaim storms in defrag_mode" (Johannes Weiner)
Address severe performance regressions, swap storms, and spurious
OOMs caused by vm.defrag_mode=1 under high memory pressure in Meta
production
It updates the page allocator slowpath so non-movable allocation
requests actively trigger direct reclaim and direct compaction at
pageblock_order scale, allowing them to claim whole pageblocks rather
than spinning unproductively.
- "zram: lockmap tweaks" (Sebastian Siewior)
Optimize and fix lockdep tracking for zram devices by consolidating
per-entry lockmaps and isolate lock classes across multiple instances
This reduces memory overhead by replacing per-entry lockdep_map
instances with a single map per struct zram, and assigns a dynamic
lock_class_key to each instance to prevent false deadlock reports
when different zram devices are backed by distinct filesystems.
* tag 'mm-stable-2026-08-18-18-39' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm: (501 commits)
selftests/mm: thuge-gen: fix test_shmget() for PAGE_SIZE check
selftests/mm: unpoison pages in memory-failure teardown
mm/shmem: downgrade final i_blocks check in shmem_evict_inode() to pr_warn()
mm/khugepaged: replace mutex_lock/mutex_unlock usage with guard macro
mm/zsmalloc: fix release order of locks in zs_page_migrate()
Documentation: zram: remove sections numbering
ksm: stop iterating VMAs when ksm_test_exit returns true
mm: fold userfaultfd_rwp() to false without CONFIG_ARCH_HAS_PTE_PROTNONE
mm/migrate: report RCU-tasks quiescent states in migrate_pages_batch()
zram: use a custom key for each zram object
zram: move lockmap to be per-zram instead per table
selftests/mm: fix gup_longterm EINVAL error message
mm: page_alloc: fix non-movable reclaim storm in defrag_mode
mm: page_alloc: move capture_control to the page allocator
mm: compaction: support non-movable compaction for pageblock requests
mm: page_alloc: __GFP_FS lockdep annotation for direct compaction
hugetlb: evaluate subpool free state while locked
mm/damon: remove trailing semicolons after function definitions
mm/damon/ops-common: prevent migration fallback to non-target nodes
mm/damon: update outdated comment about DAMOS filter handling
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/sysctl/sysctl
Pull sysctl updates from Joel Granados:
- Fix kernel-doc warnings by adjusting in file documentation
- Consolidate do_proc_* function into do_proc_vec
Consolidate three slightly different implementations of applying a
converter on all elements of a vector. Fixes to this function now
propagate to the three types.
- Replace CONFIG_PROC_SYSCTL with CONFIG_SYSCTL (they were the same)
and restrict cad_pid modifications to global root (GLOBAL_ROOT_UID)
* tag 'sysctl-7.03-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/sysctl/sysctl:
sysctl: remove CONFIG_PROC_SYSCTL, it just mirrors CONFIG_SYSCTL
sysctl: move the "cad_pid" entry from pid_table[] to kern_reboot_table[]
sysctl: repair some kernel-doc comments
sysctl: add Returns: kernel-doc for all functions
sysctl: Update API function documentation
sysctl: Rename proc_doulongvec_minmax_conv to proc_doulongvec_conv
sysctl: Group proc_handler declarations and document
sysctl: Replace do_proc_do{int,ulong,uint}vec with do_proc_vec
sysctl: Add negp parameter to douintvec converter functions
sysctl: Move default converter assignment out of do_proc_dointvec
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull x86 cpuid updates from Borislav Petkov:
- Get rid of static_cpu_has() - one less API to care about testing CPU
features
- Unify the handling of CPU core types (performance, efficient, etc) by
mapping the vendor-specific types to Linux ones
- Continuation of the work of Ahmed Darwish to centralize CPUID leaf
representation
* tag 'x86_cpu_for_v7.3_rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
x86/CPU: Rename struct cpuid_read_output to struct cpuid_output
x86/cpu/scattered: Sort it properly
x86/cpu: Use parsed CPUID(0x1)
x86/lib: Add CPUID(0x1) family and model calculation
x86/cpu: Use parsed CPUID(0x0)
x86/cpu/transmeta: Rescan CPUID(0x1) after modifying capabilities
x86/topology: Add TOPO_CPU_TYPE_LOW_POWER
x86/topology: Name the AMD core-type values
x86/topo: Map vendor CPU types to generic Linux such types
x86/bugs: Don't use cpu-type matching in cpu_vuln_blacklist
x86/cpu: Hide and rename static_cpu_has()
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull x86 cleanups from Borislav Petkov:
- The usual pile of smallish cleanups and fixlets all over the place
* tag 'x86_cleanups_for_v7.3_rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
x86/cpu: Remove unnecessary __maybe_unused annotations
x86/msr: Document the I/O-like write semantics in the msr driver
x86/apic: Ensure ICR register write value is handled as 32 bits
x86/boot/compressed/head_64.S: Clean up SEV-related comments
Documentation/arch/x86/amd-memory-encryption.rst: Fix typo
x86/platform/quark: Fix kernel-doc warnings in imr.c
x86/ras: Move contents from arch/x86/ras/Kconfig into drivers/ras/Kconfig
x86/cpu: Move intel_get_platform_id() to cpu/intel.c
x86/mm: Fix typo in comment
x86/fpu: Fix kernel-doc formatting above fpu_enable_guest_xfd_features()
x86/cfi: Use symmetric SYM_START and SYM_END in __CFI_TYPE()
x86/cfi: Add __init_or_module annotations for fineibt
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull x86 resource control updates from Borislav Petkov:
- How refreshing: no new features but a whole pile of fixes to more or
less serious issues reported by Sashiko along with miscellaneous
cleanups all over the place. All except one by Reinette Chatre, the
one by Tony Luck.
* tag 'x86_cache_for_v7.3_rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
fs/resctrl: Inform user space when status buffer overflowed
fs/resctrl: Communicate resource group deleted error via last_cmd_status
fs/resctrl: Add last_cmd_status support for writes to max_threshold_occupancy
fs/resctrl: Change last_cmd_status custom during input parsing
fs/resctrl: Use accurate and symmetric exit flows
fs/resctrl: Pass error reading event through to user space
fs/resctrl: Use accurate type for rdt_resource::rid
fs/resctrl: Change pattern used to track number of entries in enum resctrl_conf_type
x86/resctrl: Protect against bad shift
fs/resctrl: Use correct format specifier for printing error pointers
fs/resctrl: Fix UAF from worker threads when domains are removed
x86/resctrl: Ensure domain fully initialized before placed on RCU list
fs/resctrl: Prevent deadlock and use-after-free in info file handlers
fs/resctrl: Prevent use-after-free in rdtgroup_kn_put()
fs/resctrl: Fix deadlock on errors during mount
fs/resctrl: Move functions to avoid forward references in subsequent fixes
x86,fs/resctrl: Document safe RCU list traversal
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull x86 MSR updates from Ingo Molnar:
- Streamline the x86 MSR handling APIs along the 64-bit variants,
simplifying the interfaces.
Removal of the old APIs is planned for the next cycle, to reduce
churn & integration pain (Juergen Gross)
* tag 'x86-msr-2026-08-17' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: (21 commits)
x86/mce: Work around build warning after MSR-interface switch
cpufreq: Stop using 32-bit MSR interfaces
x86/featctl: Stop using 32-bit MSR interfaces
KVM/x86: Stop using 32-bit MSR interfaces
x86/mtrr: Stop using 32-bit MSR interfaces
acpi: Stop using 32-bit MSR interfaces
powercap: Stop using 32-bit MSR interfaces
thermal/intel: Stop using 32-bit MSR interfaces
x86/olpc: Stop using 32-bit MSR interfaces
x86/hyperv: Stop using 32-bit MSR interfaces
hwmon: Stop using 32-bit MSR interfaces
EDAC: Stop using 32-bit MSR interfaces
x86/cpu: Stop using 32-bit MSR interfaces
x86/apic: Stop using 32-bit MSR interfaces
x86/resctrl: Stop using 32-bit MSR interfaces
x86/tsc: Stop using 32-bit MSR interfaces
x86/amd: Stop using 32-bit MSR interfaces
x86/pci: Stop using 32-bit MSR interfaces
x86/hygon: Stop using 32-bit MSR interfaces
x86/mce: Stop using 32-bit MSR interfaces
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull locking updates from Ingo Molnar:
"Futexes:
- Use runtime constants for futex_hash computation (K Prateek Nayak,
Peter Zijlstra)
- Optimise the size check get_futex_key() (Sebastian Andrzej Siewior)
- Avoid private hash use-after-free on final put (Felix Hoffmann)
- Tell kmemleak we're not leaking __futex_queues (Peter Zijlstra)
Rust integration updates:
- Implement refcounted interrupt disable and SpinLockIrq for Rust
(Boqun Feng, Heiko Carstens, Joel Fernandes, Lyude Paul)
- Rust sync: add helpers for mb, dma_mb and friends; add generic
memory barriers and use LKMM atomics instead of Rust atomics in the
revocable code (Gary Guo)
- Add abstraction and integrate synchronize_rcu() (Philipp Stanner)
Lock debugging:
- Add qspinlock contended_release tracepoint (Dmitry Ilvokhin, Peter
Zijlstra)
- Enable the printing of held locks of remote running tasks and print
task CPU (Ingo Molnar)
- percpu-rwsem: Annotate intentional data race in readers_active_check()
(Sun Shaojie)
Misc fixes and updates by Boqun Feng, Peter Zijlstra, Fangrui Song,
Naveen Kumar Chaudhary and Thomas Huth"
* tag 'locking-core-2026-08-17' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: (44 commits)
rust: sync: Introduce SpinLockIrq::lock_with() and friends
rust: sync: Add SpinLockIrq
rust: sync: Use super::* in spinlock.rs
rust: helper: Add spin_{un,}lock_irq_{enable,disable}() helpers
rust: Introduce interrupt module
s390/preempt: Enable HAS_SEPARATE_PREEMPT_RESCHED_BITS
arm64: sched/preempt: Enable HAS_SEPARATE_PREEMPT_RESCHED_BITS
preempt: Introduce HAS_SEPARATE_PREEMPT_RESCHED_BITS
sched: Avoid signed comparison of preempt_count() in __cant_migrate()
sched: Remove the unused preempt_offset parameter of __cant_sleep()
locking: Switch to _irq_{disable,enable}() variants in cleanup guards
irq: Add KUnit test for refcounted interrupt enable/disable
irq,spin_lock: Add counted interrupt disabling/enabling
openrisc: Include <linux/cpumask.h> in smp.h
preempt: Introduce __preempt_count_{sub,add}_return()
preempt: Introduce HARDIRQ_DISABLE_BITS
preempt: Track NMI nesting to separate per-CPU counter
futex: Tell kmemleak we're not leaking __futex_queues
x86/paravirt: Trace contended_release on unlock
tracing/lock: Use TRACE_EVENT_FN() for contended_release
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/xen/tip
Pull xen updates from Juergen Gross:
- Small cleanups for the Xen ACPI pad driver and the gnttab driver
- Fix an issue with Xen PV device initialization seen with QubesOS
tests
- Fixes for the Xen balloon driver and the xenbus driver
- Simplify Xen related kernel configuration
* tag 'for-linus-7.3-rc1-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/xen/tip:
xenbus: Unregister reboot notifier on init failure
x86/xen: Drop CONFIG_XEN_PVHVM_SMP
xen: Drop CONFIG_XEN_AUTO_XLATE
xen: Drop CONFIG_XEN_PVHVM
x86/xen: Remove redundant config dependency on X86_LOCAL_APIC
x86/xen: fix init of balloon stats again
xen/xenbus: check otherend_id only after it has been initialized
xen/xenbus: log more information when device state got reset
Xen/gnttab: adjust two uses of sizeof()
ACPI: PAD: xen: Stop setting acpi_device_name/class()
|
|
With the recently found INVLPGB / TLBSYNC issue, there has been some
interest in disabling INVLPGB-based TLB flushing, in order to rule out
that CPU issue as a cause of userspace crashes.
Add a kernel command line option to control the TLB flushing behavior.
If the need arises, we will add a "tlbi=broadcast" for the case when TLB
invalidation broadcasts need to be explicitly selected, but this is not
needed now yet.
[ bp: Rewrite commit message, move to cpu/common.c, add documentation. ]
Fixes: 767ae437a32d ("x86/mm: Add INVLPGB feature and Kconfig entry")
Suggested-by: Borislav Petkov <bp@alien8.de>
Signed-off-by: Rik van Riel <riel@surriel.com>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Cc: <stable@kernel.org>
Link: https://patch.msgid.link/20260729204341.3eb0b5ea@fangorn
|
|
With the changes that enable preempt count to track IRQ disabling
nesting, we don't have enough bits in 32-bit preempt count
implementation, as a result we move NMI nesting bits out of the 32-bit
preempt count. However on the architectures that can support 64-bit
preempt count implementation, we can keep the NMI nesting bits in the
32-bit preempt count and avoid maintaining NMI nesting bits outside of
the same cache line.
Therefore HAS_SEPARATE_PREEMPT_RESCHED_BITS is introduced to allow
architectures to select this. Note that under this Kconfig, preempt
count is maintained in a 64-bit word however preempt_count() still
remains as an int because all the effective bits still fit in
(previously we mask out NEED_RESCHED bit in preempt_count()). This
should make no functional changes for existing preempt_count() users.
Enable this for x86_64 along with the introduction of the Kconfig.
[boqun: Undo the __preempt_count_{add,sub}() optimization in 32-bit
preempt count since it may introduce {over,under}flow]
Originally-by: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Boqun Feng <boqun@kernel.org>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://patch.msgid.link/20260804161447.84806-11-boqun@kernel.org
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull x86 fix from Ingo Molnar:
- Fix MCE CMCI discovery initialization ordering bug (Breno Leitao)
* tag 'x86-urgent-2026-08-08' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
x86/mce: Set up the polling timer before CMCI discovery
|
|
On x86 CONFIG_XEN_PVHVM is now a synonym of CONFIG_XEN.
In Xen specific x86 code it can be just dropped, in non-Xen specific
x86 code it can be replaced with CONFIG_XEN.
In architecture independent code it is used only where CONFIG_XEN is
defined, so it can be replaced with CONFIG_X86 there.
Reviewed-by: Stefano Stabellini <sstabellini@kernel.org>
Signed-off-by: Juergen Gross <jgross@suse.com>
Message-ID: <20260805082137.1214967-3-jgross@suse.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
- Add a mitigation for the attack vector of interrupting the saferet
sequence used in the SRSO mitigation and still poisoning the RSB.
Do that by emulating the saferet sequence and thus avoiding executing
a RET instruction.
* tag 'x86_bugs_saferet' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
x86/bugs: Make Safe-RET robust against interrupt injection
|