| Age | Commit message (Collapse) | Author |
|
https://git.kernel.org/pub/scm/linux/kernel/git/driver-core/driver-core.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/ras/ras.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git
# Conflicts:
# Documentation/scheduler/index.rst
# arch/arm64/configs/defconfig
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/broonie/regmap.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm.git
|
|
# Conflicts:
# fs/coredump.c
# fs/f2fs/f2fs.h
# fs/fuse/dax.c
# fs/xfs/libxfs/xfs_btree.c
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/mm/linux.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/broonie/regmap.git
|
|
# New commits in sched/core:
4a3b51aab6e2 ("smpboot: Don't park the thread if work is pending")
40dcc9bdbef3 ("irq_work: Flush lazy work CPU down on PREEMPT_RT")
791b1760accd ("irq_work: Update a comment regarding CPU hotplug invocation")
648d44bda731 ("sched/topology: Add asymmetric SMT packing override")
c8fc4136fd3c ("sched/fair: Honor asymmetric SMT priority in idle selection")
53bc5c556b82 ("sched: Set TIF_NEED_RESCHED before calling __trace_set_need_resched()")
4b1f75be23c4 ("sched/core: Fix context analysis errors in non-preferred CPU push")
1fb28c664a19 ("virt/steal_governor: Enable the driver")
27d47ebce4d6 ("virt/steal_governor: Implement steal_governor policy loop")
4b9302d494ff ("virt/steal_governor: Add control knobs for handling steal values")
9a8e740ee9f6 ("virt: Introduce steal governor driver")
68957caaa9c0 ("sched/debug: Add migration stats due to non preferred CPUs")
74699f56ebcf ("sched/core: Push current task from non preferred CPU")
4ee29b029058 ("sched/fair: Load balance only among preferred CPUs")
d8a3da0de843 ("sched/core: Try to use a preferred CPU in is_cpu_allowed")
620824516557 ("sysfs: Add preferred CPU file")
518b32bd5bb3 ("cpumask: Introduce cpu_preferred_mask")
06a49ef784ac ("sched/docs: Document cpu_preferred_mask and Preferred CPU concept")
cfb463b7172d ("cpumask: Introduce cpumask_intersects_and")
a8d0854a76a8 ("sched/cputime: Add kcpustat_field_total helper")
be100c77178e ("sched: Add sched_ext hooks for proxy execution")
57c75e3ae38c ("sched: Add helper to block retained proxy donors")
a49653d0abeb ("sched/core: Mark wakeups completed through ttwu_runnable()")
8f8c0417e973 ("sched/core: Dequeue waking proxy donors before reset")
313b652837d0 ("sched/core: Drop mutex locks before proxy rescheduling")
627ea30aca3b ("sched/wait: Clarify WF_SYNC wakeup semantics")
d2e010082757 ("sched/eevdf: Handle more short slice waking cases")
4bf32ec3327d ("sched/eevdf: Align update_protect_slice to set_protect_slice")
aae2a33ea662 ("sched/eevdf: Ensure that vprot will never go above a min slice")
c9ce69fc43bd ("sched/fair: Randomize equally shallow slow-path candidates")
abe440b3770f ("sched/fair: Drop idle recency from slow-path CPU selection")
fbbc63fed0b0 ("sched/core: Remove redundant core_sched_seq")
819224e506bc ("sched/fair: Remove dead code on enqueue_task_fair()")
c72945693b90 ("sched: Restart fair hrtick after same-task repicks")
a9b3c7570564 ("sched/headers: Replace __ASSEMBLY__ with __ASSEMBLER__ in the <uapi/linux/sched.h> header")
e81ee0630837 ("sched/fair: Reset NUMA fault locality after scan period update")
ef9293b3b797 ("sched: dynamic: Fix preemption model strings")
879eaa76e608 ("sched: Remove unneeded function type cast in do_balance_callbacks()")
f549101187c8 ("sched/deadline: check start_dl_timer expiry with ktime_before()")
2a672daa4b27 ("sched/feat: Use the new static key API for sched_feat")
a5576ebce920 ("sched: Convert paravirt_steal to new static key APIs")
9650ce11f2e3 ("sched: dynamic: Simplify preempt model accessors")
5b9a28eeed37 ("sched: dynamic: Remove HAVE_PREEMPT_DYNAMIC_{CALL,KEY}")
aa4178f63847 ("sched: dynamic: Simplify irqentry_exit_cond_resched()")
b9d267b9d632 ("sched: dynamic: Simplify preempt_schedule{,_notrace}()")
88e0b3bb9930 ("sched: dynamic: Simplify {cond,might}_resched()")
d3d16750693b ("sched: dynamic: Make PREEMPT_DYNAMIC depend on ARCH_HAS_PREEMPT_LAZY")
772d9ffbfd26 ("sched: Migrate whole chain in proxy_migrate_task()")
6b73a09e943f ("sched: Break out core of attach_tasks() helper into sched.h")
1f8805138593 ("sched: Switch rq->next_class in proxy_reset_donor()")
09351db90a28 ("sched/core: Don't proxy-exec unmatched cookie lock owners")
9be817f991e2 ("sched/core: Avoid migrating blocked_on tasks")
3dd95f077371 ("sched/core: Don't steal a proxy-exec donor")
Signed-off-by: Ingo Molnar <mingo@kernel.org>
|
|
mm-unstable into for-next
Signed-off-by: David Hildenbrand (Arm) <david@kernel.org>
|
|
into for-next
Signed-off-by: David Hildenbrand (Arm) <david@kernel.org>
|
|
All present section iterators run before memory hotplug added any further
memory sections, Therefore, we can simply use the SECTION_IS_EARLY flag by
setting that flag earlier in sparse_sections_init().
Get rid of SECTION_MARKED_PRESENT entirely and rename
for_each_present_section_nr() to for_each_early_section_nr().
Take care of the .clang-format for_each_present_section_nr() handling.
No functional change intended.
Link: https://lore.kernel.org/20260921-b4-sparsemem_cleanups-v2-10-54d81d65e125@kernel.org
Signed-off-by: David Hildenbrand (Arm) <david@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Acked-by: Oscar Salvador <osalvador@suse.de>
Acked-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Reviewed-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Kairui Song <kasong@tencent.com>
Cc: Qi Zheng <qi.zheng@linux.dev>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Barry Song <baohua@kernel.org>
Cc: Axel Rasmussen <axelrasmussen@google.com>
Cc: Yuanchu Xie <yuanchu@google.com>
Cc: Wei Xu <weixugc@google.com>
Cc: Baoquan He <baoquan.he@linux.dev>
Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: Brendan Jackman <brendan.jackman@linux.dev>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: Zi Yan <ziy@nvidia.com>
Cc: Jan Kiszka <jan.kiszka@siemens.com>
Cc: Kieran Bingham <kbingham@kernel.org>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Rafael J. Wysocki <rafael@kernel.org>
Cc: Danilo Krummrich <dakr@kernel.org>
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
|
|
* ras/edac-urgent: (1107 commits)
EDAC/amd64: Mask UMC chip select to the four implemented selects
Linux 7.3-rc6
Input: atkbd - skip deactivate for Lenovo IdeaPad Slim 3 15IWC11
PCI: Accept AtomicOps already enabled by the hypervisor
bpf: Fix missing migration protection in __rhtab_map_lookup_and_delete_batch()
random: vDSO: avoid call to memset() when zeroing reserved parameter
futex: Fix private hash use-after-free on resize
irq: Make refcount_interrupt kunit test selectable
tty: add missing driver flag kernel-doc colon
kprobes: Skip disarmed probes when checking optkprobe overlap
drm/mediatek: Fix ovl adaptor platform device leak
drm/mediatek: Fix runtime PM leak in mtk_hdmi_ddc_v2_probe()
drm/mediatek: Fix pdev reference leak in mtk_drm_bind()
drm/amdgpu: reset VI ASIC on MacBookPro14,3
drm/amd/pm/si: Fix updating clock limits on AC/DC
drm/amd/display: Fix stale replay_events after mod_power stream removal
drm/radeon: Read the VRAM VBIOS signature with readb()
drm/amd/display: guard dc_sink dereferences in MST mode validation
drm/amd/display: Try RGB before YCbCr 4:4:4 in stream validation
drm/amd/display: Fix sanitizer check for the DML frame size limit
...
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
|
|
* pm-runtime:
PM: runtime: Clarify driver callback expectations and structure Section 2
PM: runtime: Clarify ->runtime_idle() callback return value handling
PM: runtime: Add "Section" hyperlinks
PM: runtime: Misc improvements to runtime_pm.rst
PM: runtime: More kerneldoc formatting
PM: runtime: Correct pm_runtime_autosuspend_expiration() doc
* pm-sleep:
x86/hibernate: Ignore page-zero RAM in E820 checksum
PM: sleep: Add resume event mapping for PMSG_POWEROFF
|
|
PM_EVENT_POWEROFF (0x0800) and PMSG_POWEROFF were already defined in
include/linux/pm.h, and pm_op(), pm_late_early_op() and pm_noirq_op()
already dispatch PM_EVENT_POWEROFF to the same ops->poweroff*() callbacks
as PM_EVENT_HIBERNATE. Add the matching entry to resume_event() so that
when a poweroff suspend sequence is aborted (for example a device's
->poweroff_noirq() callback fails and the PM core rolls the
already-suspended devices back), those devices are resumed with
PMSG_RESTORE and their ->restore*() callbacks run.
Map it to PMSG_RESTORE, mirroring the existing PM_EVENT_HIBERNATE
handling. This is correct because both events suspend via
ops->poweroff*(), whose counterpart is ops->restore*(). Mapping to
PMSG_ON would be a no-op: it is what the default return already yields,
and pm_op()/pm_late_early_op()/pm_noirq_op() have no PM_EVENT_ON case,
so they would return NULL and silently skip every resume callback on the
abort path.
This completes the PMSG_POWEROFF infrastructure needed to support
differentiating between hibernation and shutdown sequences when
re-using callbacks for common code.
Hibernation is started by writing a hibernation method (such as 'platform'
'shutdown', or 'reboot') to use into /sys/power/disk and writing 'disk' to
/sys/power/state.
Shutdown is initiated with the reboot() syscall with arguments on whether
to halt the system or power it off.
Tested-by: Eric Naim <dnaim@cachyos.org>
Signed-off-by: Mario Limonciello (AMD) <superm1@kernel.org>
[ rjw: Subject tweak ]
Link: https://patch.msgid.link/20260926225555.2916525-1-mario.limonciello@amd.com
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
|
|
In taking another pass at these docs, I found some more inconsistencies.
Signed-off-by: Brian Norris <briannorris@chromium.org>
Link: https://patch.msgid.link/20260929194956.v3.2.Ia1cae8a40a24662df11553c225f4438eab65b7db@changeid
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
|
|
The "adjusted to be nonzero" comment may be a relic from when this API
previously used jiffies, although I'm not quite sure about that either.
In any case, it doesn't seem correct today.
Signed-off-by: Brian Norris <briannorris@chromium.org>
Link: https://patch.msgid.link/20260929194956.v3.1.Ifc8b2d1742729cc82f46220d85c6382fdee1fc8b@changeid
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
|
|
* pm-misc:
PM: clk: fix typo "sucessfully" in comment
* pm-em:
PM: EM: Fix use-after-free of perf domain in netlink doit handlers
|
|
* pm-runtime:
PM: runtime: call pm_runtime_dont_use_autosuspend() on reinit()
PM: core: Document struct dev_pm_info with kerneldoc
PM: runtime: Pull API docs from kerneldoc
PM: runtime: kerneldoc wording improvements
PM: runtime: Improve set_{status,active,suspended} docs
PM: runtime: kerneldoc fixes
PM: runtime: Add kunit test for supplier idle/suspend
PM: runtime: Only queue an idle check for RPM-linked suppliers
* pm-sleep:
PM: sleep: Add DPM watchdog to prepare/late/early/noirq/complete phases
PM: hibernate: docs: update swsusp.txt references to swsusp.rst
* pm-tools:
tools: power: pm-graph: fix typo "hierachy" in comment
|
|
linux-next
* acpi-scan:
ACPI: scan: Combine two conditionals in acpi_bus_attach()
ACPI: PM: Move acpi_bus_init_power() declaration to internal header file
ACPI: scan: Stop calling acpi_bus_init_power() early
ACPI: PM: Drop parent state update from acpi_device_get_power()
ACPI: scan: Drop useless and noisy debug statement
* acpi-glue:
ACPI: glue: Skip devices with no type in acpi_device_notify()
ACPI: glue: Fix up and adjust acpi_unbind_one()
ACPI: glue: Rearrange acpi_bind_one() to avoid breakage
ACPI: glue: Carry out companion lookup under bus_type_sem
ACPI: glue: Rework the success message in acpi_device_notify()
ACPI: glue: Reduce debug noise from acpi_device_notify()
* acpi-bus:
driver core/ACPI: Introduce companion_bus_register()
ACPI: bus: Reduce runtime memory footprint of struct acpi_device
* acpi-sysfs:
ACPI: sysfs: use strscpy() instead of strcpy()
|
|
|
|
Peng Fan (OSS) <peng.fan@oss.nxp.com> says:
The regmap lock is taken by calling the lock and unlock callbacks
directly at every call site:
map->lock(map->lock_arg);
...
map->unlock(map->lock_arg);
The open-coded pattern forces every error path to unlock by hand,
which spreads goto out_unlock chains and duplicated unlock statements
throughout regmap.c and regcache.c and makes it easy to leak the lock
on a newly added return path.
Define a scoped guard for the regmap lock (DEFINE_GUARD) in internal.h
and convert the users to it. The lock and unlock callbacks are chosen
at init time (mutex, spinlock, raw spinlock, hwspinlock or none) and
return void, so an unconditional guard is sufficient. Function-scope
critical sections use guard(regmap)(); sites that must run work after
the lock is dropped - for example regmap_register_patch() calling
regmap_async_complete() and the debugfs cache_only handler calling
regcache_sync() (which takes the lock itself) - use
scoped_guard(regmap, ...) so the trailing work stays outside the
guarded region.
regcache_sync() and regcache_sync_region() are deliberately left
unconverted: they already use a single goto out unlock path, so a
guard would save nothing while forcing either a goto inside a
scoped_guard scope or a control-flow rewrite, neither of which is an
improvement.
This patchset removes many manual unlock statements together with the
associated goto out_unlock labels.
The series is split to keep the regcache: and regmap: changes on their
own commits and to preserve independent revertibility:
1. define the guard and convert regmap.c
2. convert regcache.c (except the sync helpers, see above)
3. convert the regcache rbtree debugfs dump
4. convert the regmap debugfs write handlers
No functional change.
Tested with the regmap KUnit suite (drivers/base/regmap/regmap-kunit.c)
under ARCH=um: 551/551 tests pass, and 551/551 again with lockdep
(PROVE_LOCKING, DEBUG_LOCK_ALLOC, DEBUG_ATOMIC_SLEEP) enabled with no
splats. Built clean with sparse (C=1) showing no lock-context
imbalance warnings.
Link: https://patch.msgid.link/20260924-regmap-lock-guard-v3-0-8a6127223c16@nxp.com
|
|
Convert the open-coded map->lock()/map->unlock() users in the debugfs
cache_only and cache_bypass write handlers to the regmap scoped guard.
regmap_cache_bypass_write_file() uses guard(regmap)() since the locked
region runs to the return. regmap_cache_only_write_file() uses
scoped_guard(regmap, ...) because the subsequent regcache_sync() must
run with the lock released - it takes the regmap lock itself - so it
stays outside the guarded scope exactly as before. No functional
change.
Assisted-by: LLM
Signed-off-by: Peng Fan <peng.fan@nxp.com>
Link: https://patch.msgid.link/20260924-regmap-lock-guard-v3-4-8a6127223c16@nxp.com
Signed-off-by: Mark Brown <broonie@kernel.org>
|
|
Convert the open-coded map->lock()/map->unlock() pair in rbtree_show()
to guard(regmap)(). The locked region spans the whole function body up
to the single return, so a function-scope guard drops the manual
unlock with no functional change.
Assisted-by: LLM
Signed-off-by: Peng Fan <peng.fan@nxp.com>
Link: https://patch.msgid.link/20260924-regmap-lock-guard-v3-3-8a6127223c16@nxp.com
Signed-off-by: Mark Brown <broonie@kernel.org>
|
|
Convert the open-coded map->lock()/map->unlock() users in regcache.c
to the regmap scoped guard introduced for regmap.c. Use
scoped_guard(regmap, ...) in regcache_exit(), where the locked region
is a subsection of the function, and guard(regmap)() for the
function-scope critical sections.
regcache_init() keeps explicit map->lock()/map->unlock() calls: it
has a goto err_* cleanup ladder and mixing goto with cleanup helpers
in the same function is not allowed by cleanup.h.
regcache_sync() and regcache_sync_region() are left as-is: they
already use a single goto out unlock path, so converting them would
require either mixing a goto with a scoped_guard scope or restructuring
their control flow, neither of which is an improvement.
No functional change.
Assisted-by: LLM
Signed-off-by: Peng Fan <peng.fan@nxp.com>
Link: https://patch.msgid.link/20260924-regmap-lock-guard-v3-2-8a6127223c16@nxp.com
Signed-off-by: Mark Brown <broonie@kernel.org>
|
|
The regmap lock and unlock callbacks are invoked directly as
map->lock(map->lock_arg) / map->unlock(map->lock_arg) at every call
site. This open-coded pattern requires manual unlocking on every
return path, which spreads goto out_unlock chains and duplicated
unlock statements throughout the code and is an easy place to leak
the lock on an error path.
Define a scoped guard for the regmap lock in internal.h. The lock and
unlock callbacks are chosen at init time (mutex, spinlock, raw
spinlock, hwspinlock or none) and return void, so an unconditional
DEFINE_GUARD() is sufficient.
Convert all lock/unlock users in regmap.c to guard(regmap)() for
function-scope critical sections and scoped_guard(regmap, ...) where
work must run outside the lock (e.g. regmap_register_patch() calling
regmap_async_complete()). This drops every manual unlock and the
associated goto out_unlock labels with no functional change.
Assisted-by: LLM
Signed-off-by: Peng Fan <peng.fan@nxp.com>
Link: https://patch.msgid.link/20260924-regmap-lock-guard-v3-1-8a6127223c16@nxp.com
Signed-off-by: Mark Brown <broonie@kernel.org>
|
|
Add a "preferred" file in /sys/devices/system/cpu/ when kernel is
built with CONFIG_PREFERRED_CPU=y.
This would help
- Users to quickly check which CPUs are marked as preferred.
- Userspace daemons such as irqbalance to use this mask to
route irqs into preferred CPUs.
For example:
cat /sys/devices/system/cpu/online
0-719
cat /sys/devices/system/cpu/preferred
0-599 <<< Implies 0-599 are preferred for workloads and 600-719
should be avoided at this moment.
cat /sys/devices/system/cpu/preferred
0-719 <<< All CPUs are usable. There is no preference.
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://patch.msgid.link/20260928053728.797539-6-sshegde@linux.ibm.com
|
|
We need the driver-core fixes in here as well to build on top of.
Signed-off-by: Danilo Krummrich <dakr@kernel.org>
|
|
There is no need to double check the 'pa' object.
Reviewed-by: Dawid Osuchowski <dawid.osuchowski@linux.intel.com>
Signed-off-by: Cezary Rojewski <cezary.rojewski@intel.com>
Reviewed-by: Vishal Moola (Fractile) <vishal.moola@gmail.com>
Link: https://patch.msgid.link/20260925093928.3765486-1-cezary.rojewski@intel.com
Signed-off-by: Danilo Krummrich <dakr@kernel.org>
|
|
Add a pm_runtime_dont_use_autosuspend() call to pm_runtime_reinit() to
ensure autosuspend is disabled on driver teardown.
Suggested-by: Frank Li <Frank.Li@nxp.com>
Signed-off-by: Joshua Crofts <joshua.crofts1@gmail.com>
Link: https://patch.msgid.link/20260919-move-dont-use-autosuspend-v1-1-f6e2d1315c23@gmail.com
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
|
|
Correct "sucessfully" to "successfully", reported by scripts/checkpatch.pl
using the misspelling list in scripts/spelling.txt.
Only touches comments, no code changes.
Signed-off-by: Hemanth Selam <hemanth.selam@gmail.com>
Reviewed-by: Randy Dunlap <rdunlap@infradead.org>
[ rjw: Subject adjustment ]
Link: https://patch.msgid.link/20260907064725.10686-2-hemanth.selam@gmail.com
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
|
|
Extend the DPM watchdog to wrap device_prepare, device_suspend_late,
device_suspend_noirq, device_resume_noirq, device_resume_early, and
device_complete callbacks. If a driver hangs during these transitions,
the watchdog will fire and dump a stack trace to help identify the
offending driver.
To prevent false-positive timeouts, the watchdog is set only after
waiting for subordinate (during suspend) and superior (during resume)
devices. With this change, DPM watchdog coverage is extended across all
phases of system sleep transitions.
Signed-off-by: Mayank Rungta <mrungta@google.com>
Reviewed-by: Douglas Anderson <dianders@chromium.org>
Reviewed-by: Tzung-Bi Shih <tzungbi@kernel.org>
Link: https://patch.msgid.link/20260909224818.2177311-1-mrungta@google.com
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
|
|
The ACPI bus type does not allow drivers to be matched to devices, so
the sysfs attributes related to drivers created for it are useless and
their existence is confusing.
Moreover, it is better to prevent drivers from being registered and
looked up for a bus like that.
To allow skipping the creation of those sysfs attributes and preventing
driver registration and lookup for the ACPI bus type, introduce a
"companion" bus type concept and add a special registration function
for registering "companion" bus types, companion_bus_register().
The "drivers" directory under the ACPI bus type is still needed because
there are versions of systemd that depend on it [1].
Link: https://lore.kernel.org/linux-acpi/SN6PR02MB41575266A4580339E186E5D9D4812@SN6PR02MB4157.namprd02.prod.outlook.com/ [1]
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
Acked-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Reviewed-by: Danilo Krummrich <dakr@kernel.org>
Tested-by: Michael Kelley <mhklinux@outlook.com>
Reviewed-by: Michael Kelley <mhklinux@outlook.com>
[ rjw: Add comment regarding drivers_kset creation in bus_register_internal() ]
Link: https://patch.msgid.link/12985239.O9o76ZdvQC@rafael.j.wysocki
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
|
|
When move_pfn_range_to_zone() or remove_pfn_range_from_zone() updates a
zone, set_zone_contiguous() rescans the entire zone pageblock-by-pageblock
to rebuild zone->contiguous. For large zones this is a significant cost
during memory hotplug and hot-unplug.
Add a new zone member, pages_with_online_memmap, that tracks the
number of pages within the zone span that have an online memory map,
including present pages and memory holes whose memory map has been
initialized and for which pfn_to_online_page() succeeds.
For early boot memory, pages_with_online_memmap is calculated in
memmap_init_zone_range(). PFNs initialized by memmap_init_range() are
included in pages_with_online_memmap, and hole PFNs for which
pfn_to_online_page() succeeds are also counted in
init_unavailable_range(). For hotplugged memory,
pages_with_online_memmap is updated through adjust_present_page_count(),
which is called during memory online and offline operations. When
spanned_pages == pages_with_online_memmap, every PFN in the zone span
has a valid memmap entry, so pfn_to_page() can be called for any PFN
within the zone span without an additional pfn_valid() check.
Note: this counter may temporarily undercount when pages with an
online memory map exist outside the current zone span. Such pages
are only created during boot, when initializing the memory map of
pages that do not fall into any zone span. The undercount itself
can only happen after boot, during memory hotplug, when growing
the zone to cover such pages and later shrinking it back, which
may result in a "too small" value. This is safe: it merely
prevents detecting a contiguous zone.
Here is an example (page numbers are just for illustration purposes):
after boot:
[ zone span ]
[ zone pages ]
spanned=10, initialized=15, online=10
online == spanned -> contiguous
growing after hotplug (hotplug 5):
[ zone span ]
[ zone pages ] [ zone pages ]
spanned=30, initialized=20, online=15
online != spanned -> not contiguous
shrinking after hotunplug (hotunplug 5 again):
[ zone span ]
[ zone pages ]
spanned=15, initialized=15, online=10
online != spanned -> not contiguous although contiguous
The contiguity check using pages_with_online_memmap is stricter than
the old pageblock-by-pageblock scan. The old set_zone_contiguous()
iterated at pageblock granularity via pageblock_pfn_to_page(), so a
zone could be marked contiguous even if a subsection-sized hole
existed within a pageblock. The new check requires
spanned_pages == pages_with_online_memmap, meaning every PFN in the
zone span must satisfy pfn_to_online_page().
The following test cases of memory hotplug for a VM [1], tested in the
environment [2], show that this optimization can significantly reduce the
memory hotplug time [3].
+----------------+------+---------------+--------------+----------------+
| | Size | Time (before) | Time (after) | Time Reduction |
| +------+---------------+--------------+----------------+
| Plug Memory | 256G | 10s | 3s | 70% |
| +------+---------------+--------------+----------------+
| | 512G | 36s | 7s | 81% |
+----------------+------+---------------+--------------+----------------+
+----------------+------+---------------+--------------+----------------+
| | Size | Time (before) | Time (after) | Time Reduction |
| +------+---------------+--------------+----------------+
| Unplug Memory | 256G | 11s | 4s | 64% |
| +------+---------------+--------------+----------------+
| | 512G | 36s | 9s | 75% |
+----------------+------+---------------+--------------+----------------+
[1] Qemu commands to hotplug 256G/512G memory for a VM:
object_add memory-backend-ram,id=hotmem0,size=256G/512G,share=on
device_add virtio-mem-pci,id=vmem1,memdev=hotmem0,bus=port1
qom-set vmem1 requested-size 256G/512G (Plug Memory)
qom-set vmem1 requested-size 0G (Unplug Memory)
[2] Hardware : Intel Icelake server
Guest Kernel : v7.3-rc3
Qemu : v9.0.0
Launch VM :
qemu-system-x86_64 -accel kvm -cpu host \
-drive file=./Centos10_cloud.qcow2,format=qcow2,if=virtio \
-drive file=./seed.img,format=raw,if=virtio \
-smp 3,cores=3,threads=1,sockets=1,maxcpus=3 \
-m 2G,slots=10,maxmem=2052472M \
-device pcie-root-port,id=port1,bus=pcie.0,slot=1,multifunction=on \
-device pcie-root-port,id=port2,bus=pcie.0,slot=2 \
-nographic -machine q35 \
-nic user,hostfwd=tcp::3000-:22
Guest kernel auto-onlines newly added memory blocks:
echo online > /sys/devices/system/memory/auto_online_blocks
[3] The time from typing the QEMU commands in [1] to when the output of
'grep MemTotal /proc/meminfo' on Guest reflects that all hotplugged
memory is recognized.
Reported-by: Nanhai Zou <nanhai.zou@intel.com>
Reported-by: Chen Zhang <zhangchen.kidd@jd.com>
Tested-by: Yuan Liu <yuan1.liu@intel.com>
Reviewed-by: Jason Zeng <jason.zeng@intel.com>
Reviewed-by: Chen Yu <yu.c.chen@intel.com>
Reviewed-by: Pan Deng <pan.deng@intel.com>
Co-developed-by: Tianyou Li <tianyou.li@intel.com>
Signed-off-by: Tianyou Li <tianyou.li@intel.com>
Signed-off-by: Yuan Liu <yuan1.liu@intel.com>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
Reviewed-by: Wei Yang <richard.weiyang@gmail.com>
Link: https://patch.msgid.link/20260920084946.3266279-3-yuan1.liu@intel.com
Signed-off-by: David Hildenbrand (Arm) <david@kernel.org>
|
|
underestimation bug
The scheduler scales LLC capacity by the fraction of cache-sharing CPUs
covered by a domain:
llc_bytes = cache_size * span_weight / shared_weight
During CPU teardown, sched_cpu_deactivate() rebuilds scheduler domains
before cacheinfo_cpu_pre_down() removes the CPU from shared_cpu_map. The
new domains therefore use the old sharing weight. The later call to
sched_update_llc_bytes() looks up the departing CPU's sd_llc, which has
already been detached, and returns without correcting the surviving CPUs.
On a Ryzen 5 7535U with twelve logical CPUs sharing a 16 MiB LLC,
offlining one SMT sibling left the remaining CPUs with:
llc_bytes = floor(16777216 * 11 / 12) = 15379114 bytes
The correct capacity is still 16777216 bytes. On systems with active
cache-aware scheduling, an underestimated capacity can cause
exceed_llc_capacity() to reject aggregation for a process whose footprint
would fit. Unchanged cpuset partitions sharing the physical cache can
also retain stale capacity when a CPU comes online in another partition.
Pass the cache-sharing mask already retained by cacheinfo to the
scheduler update. Refresh every surviving CPU using its own LLC domain
so that each partition receives the correct share. This also preserves
the correction needed as cache-sharing maps grow during boot.
Keep the existing CPU-hotplug and scheduler-domain synchronization. The
update remains on the hotplug path; no steady-state scheduling operation
or persistent allocation is added.
Fixes: 7030513a0877 ("sched/cache: Calculate the LLC size and store it in sched_domain")
Signed-off-by: Davi Chaves Azevedo <davichazbh@gmail.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Reviewed-by: Chen Yu <yu.c.chen@intel.com>
Reviewed-by: Tim Chen <tim.c.chen@linux.intel.com>
Reviewed-by: K Prateek Nayak <kprateek.nayak@amd.com>
Tested-by: Chen Yu <yu.c.chen@intel.com>
Tested-by: K Prateek Nayak <kprateek.nayak@amd.com>
Cc: <stable@kernel.org> # v7.2.x
Link: https://patch.msgid.link/6751d93e15889e624796c74db0bfe66603d60b1b.1790035273.git.tim.c.chen@linux.intel.com
|
|
devices and drivers
The ability to add and remove devices from a driver through the sysfs
"bind" and "unbind" files was created all those decades ago as a way
that kernel developers can iterate faster, and provide a debugging way
for users to attempt to add a new device to a driver without having to
rebuild their kernel.
This api over the years has been abused and recently come under a major
fuzzing "attack" through tools like syzbot which decided that it would
attempt to just randomly bind any driver to any type of device, causing
loads of unneeded errors and pointless kernel patches to be generated by
unsuspecting new developers.
Handle all of this by adding a new taint flag, TAINT_FORCED_BIND, which
will be set on the driver if the bind/unbind sysfs files are ever
written to. This lets kernel developers "know" that a user is
attempting to do something that is not normal, and as such, if the
kernel breaks they get to keep the shiny pieces laying around on the
floor.
The flag is 'Y' which was unused, and can remembered as the user is
"yeeting" the device being operated on here (thrown with force without
regard for the thing being thrown).
Note, the taint flag gets set _BEFORE_ the bind/unbind callback happens,
as many times crashes/oops/warnings/failures happen within the callback,
and the taint flag needs to be there to show what was being attempted.
If it were to be set after the callback happens, the oops report would
not properly reflect what foolishness was being attempted.
Fuzzing tools like syzbot, that doesn't have hand-crafted rules to keep
the tool from hitting bind/unbind, should be run with panic_on_taint
enabled so that they fall over and don't continue on, thinking that they
actually found a real issue.
Userspace operations that rely on the bind/unbind files to work around
the lack of will to upgrade a kernel image to a newer version with
proper support for new devices, or the lack of will to submit valid
device ids to driver authors, will still work properly, but now the
kernel will be flagged in a way that will show that perhaps those users
should reconsider their behavior and work to have the drivers properly
support these devices in a "native" manner.
Finally, the bind/unbind files can find real use-after-free issues with
some drivers by forcing the process to happen virtually without having
to rely on manual removal processes. Those real bugs should still be
worked on, but by adding this taint flag, developers can more easily
determine bug reports that are actually worth looking at.
Reviewed-by: Johan Hovold <johan@kernel.org>
Tested-by: Johan Hovold <johan@kernel.org>
Reviewed-by: Aaron Tomlin <atomlin@atomlin.com>
Reviewed-by: Bradley Morgan <brads@mainlining.org>
Acked-by: Danilo Krummrich <dakr@kernel.org>
Link: https://patch.msgid.link/20260914-bind_taint-v4-3-eadf8a090903@linuxfoundation.org
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
In preparation for removing duplicate documentation from
Documentation/power/runtime_pm.rst, borrow some of the useful wording
from runtime_pm.rst, and update other language for clarity, ease of
reading, and completeness.
Other guiding principles in this change:
* Try to highlight "core", as in, "functions that are not for driver
use but are exported because the real entrypoints are inline
functions"
* Rework pm_runtime_barrier() docs significantly. More below.
* Include some clarifying cross-references and recommendations for
pm_runtime_put_sync{,_suspend,_autosuspend}()
* Attempt to deemphasize some of the implementation details (e.g.,
"asynchronous" instead of "queue")
* Try for more clear user-facing language. For example, "set up
autosuspend" isn't quite clear whether we're configuring autosuspend,
or if we're initiating an attempt to autosuspend (i.e., setting a
timer).
pm_runtime_barrier(): currently, we speak a lot about implementation
details and sequences of events, but obscure the key point that it
treats "pending resume" and "pending suspend" very differently -- I try
to improve that.
Signed-off-by: Brian Norris <briannorris@chromium.org>
Link: https://patch.msgid.link/20260904141215.3.I3bdec7dcda46e1bee6d4f91530518fbe9c9096fd@changeid
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
|
|
The set_active()/set_suspended() docs don't mention that they also clear
the 'runtime_error' field. This is a very important note, since that's
one key purpose for using them.
Fix a typo in __pm_runtime_set_status() while we're at it.
Signed-off-by: Brian Norris <briannorris@chromium.org>
Link: https://patch.msgid.link/20260904141215.2.If02a538531ea004987c9a28a8d030ee9f71221cd@changeid
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
|
|
When included into a Documentation/.../*.rst file, `make htmldocs`
complains:
./include/linux/pm_runtime.h:359: WARNING: Bullet list ends without a blank line; unexpected unindent. [docutils]
[... more ...]
We should fix up the list format here to look nicer in HTML form, and
avoid warnings. Adjust to a few other kerneldoc-isms (formatting,
"Return:") while we're at it too.
The result now passes 'make htmldocs' without warning, once these files
are included in Documentation/.../*.rst.
Signed-off-by: Brian Norris <briannorris@chromium.org>
Link: https://patch.msgid.link/20260904141215.1.Ib31d6be7e93f6c02321d03c78ea410e02d7380fc@changeid
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
|
|
Prepare for the constification of 'struct dev_ext_attribute' by changing
the signature of the standard callback functions.
Migrate the hv-24x7 driver in the same commit. It is the only user to
manually assign one of the standard callbacks.
Signed-off-by: Thomas Weißschuh <linux@weissschuh.net>
Link: https://patch.msgid.link/20260907-sysfs-const-attr-dev_ext_attr-v2-3-bf53afe57071@weissschuh.net
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
Add a unit test that covers the bugfix/optimization in the previous
patch, "PM: runtime: Only queue an idle check for RPM-linked suppliers".
This test ensures:
1) device links with DL_FLAG_PM_RUNTIME can RPM-resume/suspend properly;
and
2) device links without DL_FLAG_PM_RUNTIME are not accidentally
resumed/suspended along with the consumer.
Notably, without patch 1 ("PM: runtime: Only queue an idle check for
RPM-linked suppliers"), the no-RPM supplier will fail 2 of its last 3
test expectations:
$ tools/testing/kunit/kunit.py run --kconfig_add CONFIG_PM=y 'pm_runtime*'
...
## Fails because rpm_suspend_suppliers() suspended the wrong links
[17:01:44] # pm_runtime_supplier_suspend_test: EXPECTATION FAILED at drivers/base/power/runtime-test.c:279
[17:01:44] Expected pm_runtime_active(norpm_supplier) to be true, but is false
## Fails because the no-RPM supplier is already suspended unexpectedly
[17:01:44] # pm_runtime_supplier_suspend_test: EXPECTATION FAILED at drivers/base/power/runtime-test.c:282
[17:01:44] Expected 0 == pm_runtime_suspend(norpm_supplier), but
[17:01:44] pm_runtime_suspend(norpm_supplier) == 1 (0x1)
...
Signed-off-by: Brian Norris <briannorris@chromium.org>
Link: https://patch.msgid.link/20260901013341.1417232-2-briannorris@chromium.org
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
|
|
When a device RPM-suspends, it queues up an idle check for all of its
suppliers, in case that device was the last consumer, and those
suppliers are now able to suspend. Today, we do this for all suppliers,
and not only for those suppliers that are marked for use by runtime PM.
This isn't directly harmful, but it is a bit wasteful, and also may
induce unexpected suspend attempts (e.g., if a device was purposely last
touched with pm_runtime_put_noidle()).
To avoid excess idle checks, only call pm_request_idle() on linked
suppliers that opted into runtime PM.
Noticed by inspection of trace logs.
This is not expected to have a functional impact on systems that are
using device links properly, and should only be considered an
optimization.
Fixes: 5244f5e2d801 ("PM: runtime: Defer suspending suppliers")
Signed-off-by: Brian Norris <briannorris@chromium.org>
Link: https://patch.msgid.link/20260901013341.1417232-1-briannorris@chromium.org
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
|
|
get_cpu_cacheinfo_id() is a static inline that requires get_cpu_cacheinfo(),
which is not exported. Modules that need to identify which cache instance
a CPU belongs to therefore cannot use it.
Move it out of line and export it.
Signed-off-by: Qiuxu Zhuo <qiuxu.zhuo@intel.com>
Signed-off-by: Tony Luck <tony.luck@intel.com>
Link: https://patch.msgid.link/20260828153002.10290-2-tony.luck@intel.com
|
|
Fix typos in comments, reported by scripts/checkpatch.pl using the
misspelling list in scripts/spelling.txt. Only touches comments, no code
changes.
Assisted-by: LLM
Signed-off-by: Hemanth Selam <hemanth.selam@gmail.com>
Reviewed-by: Andy Shevchenko <andriy.shevchenko@linux.intel.com>
Acked-by: Randy Dunlap <rdunlap@infradead.org>
Link: https://patch.msgid.link/20260904111816.31254-1-hemanth.selam@gmail.com
Signed-off-by: Danilo Krummrich <dakr@kernel.org>
|
|
node_init_node_access() frees the access node with kfree() if
device_register() fails. After device_register() the embedded device is
initialized and must be released with put_device() so that
node_access_release() can free it.
Fixes: 08d9dbe72b1f ("node: Link memory nodes to their compute nodes")
Signed-off-by: Linkai Gong <gonglinkai@kylinos.cn>
Link: https://patch.msgid.link/20260907024732.1452228-1-gonglinkai@kylinos.cn
Signed-off-by: Danilo Krummrich <dakr@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/kees/linux
Pull kmalloc_obj conversions from Kees Cook:
"Another run of the Coccinelle script for converting kmalloc()
family of allocations to kmalloc_obj() via the existing rules
in scripts/coccinelle/api/kmalloc_objs.cocci"
* tag 'kmalloc_obj-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/kees/linux:
treewide: refresh kmalloc_obj() conversions
drm/amd/display: Fix harmless type mismatch in allocation
|
|
This is another run of the Coccinelle script for converting kmalloc()
family of allocations to kmalloc_obj() via the existing rules in
scripts/coccinelle/api/kmalloc_objs.cocci
This catches both the set of kmalloc() uses added since the first
kmalloc_obj() conversions in v7.0 and adds a large group missed in the
first pass due to Coccinelle not interacting well with the cleanup.h
scoped_...() family of macros[1]. I worked around this with spatch's
"--macro-file" argument to a file with all the scoped_...() macros mapped
to Coccinelle's YACFE_ITERATOR[2] as that was the closest viable control
flow indicator I could find.
Build tested allmodconfig on x86, arm64, arm, loongarch, mips, powerpc,
riscv, and s390 with no new warnings.
Link: https://lore.kernel.org/lkml/202609021314.8A9C0B8@keescook/ [1]
Link: https://github.com/coccinelle/coccinelle/blob/master/standard.h [2]
Signed-off-by: Kees Cook <kees+treewide@kernel.org>
|
|
Correct "asynchrnous" to "asynchronous", reported by scripts/checkpatch.pl
using the misspelling list in scripts/spelling.txt. Only touches comments,
no code changes.
Assisted-by: Cursor:claude-opus-5
Signed-off-by: Hemanth Selam <hemanth.selam@gmail.com>
Link: https://patch.msgid.link/20260904110110.8823-1-hemanth.selam@gmail.com
Signed-off-by: Mark Brown <broonie@kernel.org>
|