diff options
| author | Mark Brown <broonie@kernel.org> | 2026-10-01 14:23:33 +0100 |
|---|---|---|
| committer | Mark Brown <broonie@kernel.org> | 2026-10-01 14:23:33 +0100 |
| commit | e273534c5e5f259fb43cda0228ce56b45ac0c93f (patch) | |
| tree | 64509f18eed07ad7bcc0766a9d01dadf747a5b52 | |
| parent | caab1372065c106f4baf18d11863e498a88a8940 (diff) | |
| parent | 114920a3e82922869f8c27916e852b4ba0d560c4 (diff) | |
| download | linux-next-e273534c5e5f259fb43cda0228ce56b45ac0c93f.tar.gz linux-next-e273534c5e5f259fb43cda0228ce56b45ac0c93f.zip | |
Merge branch 'linux-next' of https://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm.git
113 files changed, 3827 insertions, 1757 deletions
diff --git a/Documentation/ABI/testing/sysfs-devices-system-cpu b/Documentation/ABI/testing/sysfs-devices-system-cpu index 82d10d556cc8..73c8c204820d 100644 --- a/Documentation/ABI/testing/sysfs-devices-system-cpu +++ b/Documentation/ABI/testing/sysfs-devices-system-cpu @@ -335,12 +335,16 @@ Description: Performance Limited Read to check if platform throttling (thermal/power/current limits) caused delivered performance to fall below the requested level. A non-zero value indicates throttling occurred. - - Write the bitmask of bits to clear: - - - 0x1 = clear bit 0 (desired performance excursion) - - 0x2 = clear bit 1 (minimum performance excursion) - - 0x3 = clear both bits + If firmware provides the register but Linux cannot access it + safely, reads return the literal string "<unsupported>". + + Write 0x3 to clear both bits. Selectively clearing one bit is + not supported because ACPI does not define the effect of writing + one to these status bits and an interlocked read-modify-write is + not generally available. + Writing 0x3 returns -EOPNOTSUPP when firmware describes the + register using an access which Linux cannot clear safely; reads + may remain available in that case. The platform sets these bits; OSPM can only clear them. diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt index c31c4285ab96..568923dfc842 100644 --- a/Documentation/admin-guide/kernel-parameters.txt +++ b/Documentation/admin-guide/kernel-parameters.txt @@ -501,13 +501,6 @@ Kernel parameters disable Disable amd-pstate preferred core. - amd_dynamic_epp= - [X86] - disable - Disable amd-pstate dynamic EPP. - enable - Enable amd-pstate dynamic EPP. - amijoy.map= [HW,JOY] Amiga joystick support Map of devices attached to JOY0DAT and JOY1DAT Format: <a>,<b> diff --git a/Documentation/admin-guide/pm/amd-pstate.rst b/Documentation/admin-guide/pm/amd-pstate.rst index 6fd72ba52df2..4c5d3d8acf64 100644 --- a/Documentation/admin-guide/pm/amd-pstate.rst +++ b/Documentation/admin-guide/pm/amd-pstate.rst @@ -481,12 +481,6 @@ For systems that support ``amd-pstate`` preferred core, the core rankings will always be advertised by the platform. But OS can choose to ignore that via the kernel parameter ``amd_prefcore=disable``. -``amd_dynamic_epp`` - -When AMD pstate is in auto mode, dynamic EPP will control whether the kernel -autonomously changes the EPP mode. The default is disabled. It can be enabled -with the kernel parameter ``amd_dynamic_epp=enable``. - User Space Interface in ``sysfs`` - General =========================================== diff --git a/Documentation/admin-guide/pm/cpuidle.rst b/Documentation/admin-guide/pm/cpuidle.rst index be4c1120e3f0..e91ebd432f6f 100644 --- a/Documentation/admin-guide/pm/cpuidle.rst +++ b/Documentation/admin-guide/pm/cpuidle.rst @@ -316,7 +316,7 @@ interval" value. Now, the governor is ready to walk the list of idle states and choose one of them. For this purpose, it compares the target residency of each state with -the predicted idle duration and the exit latency of it with the with the latency +the predicted idle duration and the exit latency of it with the latency limit coming from the power management quality of service, or `PM QoS <cpu-pm-qos_>`_, framework. It selects the state with the target residency closest to the predicted idle duration, but still below it, and exit latency that does not exceed the diff --git a/Documentation/driver-api/thermal/cpu-idle-cooling.rst b/Documentation/driver-api/thermal/cpu-idle-cooling.rst index c2a7ca676853..b9d4c91a7f15 100644 --- a/Documentation/driver-api/thermal/cpu-idle-cooling.rst +++ b/Documentation/driver-api/thermal/cpu-idle-cooling.rst @@ -161,7 +161,7 @@ tree. So with the idle injection mechanism, we want an average power specific OPP and idle another amount of time. That could be put in a equation:: - P(opp)target = ((Trunning x (P(opp)running) + (Tidle x P(opp)idle)) / + P(opp)target = (Trunning x (P(opp)running) + (Tidle x P(opp)idle)) / (Trunning + Tidle) ... diff --git a/Documentation/firmware-guide/acpi/apei/einj.rst b/Documentation/firmware-guide/acpi/apei/einj.rst index 7d8435d35a18..70fa59bd97d8 100644 --- a/Documentation/firmware-guide/acpi/apei/einj.rst +++ b/Documentation/firmware-guide/acpi/apei/einj.rst @@ -180,8 +180,10 @@ function are specified using param1:: | segment | bus | device | function | reserved | +-------------------------------------------------+ -Anyway, you get the idea, if there's doubt just take a look at the code -in drivers/acpi/apei/einj.c. +Anyway, you get the idea, if there's doubt just take a look at the code: +drivers/acpi/apei/einj-core.c holds the core logic, while CXL-specific +handling lives in drivers/acpi/apei/einj-cxl.c and the shared plumbing +in drivers/acpi/apei/apei-internal.h. An ACPI 5.0 BIOS may also allow vendor-specific errors to be injected. In this case a file named vendor will contain identifying information diff --git a/Documentation/power/runtime_pm.rst b/Documentation/power/runtime_pm.rst index a53ab09c37d5..39fdeeda7a1e 100644 --- a/Documentation/power/runtime_pm.rst +++ b/Documentation/power/runtime_pm.rst @@ -203,103 +203,12 @@ rules: 3. Runtime PM Device Fields =========================== -The following device runtime PM fields are present in 'struct dev_pm_info', as -defined in include/linux/pm.h: +Device PM fields are found in 'struct dev_pm_info', as defined in +include/linux/pm.h. Many of those fields track runtime PM configuration and +state. - `struct timer_list suspend_timer;` - - timer used for scheduling (delayed) suspend and autosuspend requests - - `unsigned long timer_expires;` - - timer expiration time, in jiffies (if this is different from zero, the - timer is running and will expire at that time, otherwise the timer is not - running) - - `struct work_struct work;` - - work structure used for queuing up requests (i.e. work items in pm_wq) - - `wait_queue_head_t wait_queue;` - - wait queue used if any of the helper functions needs to wait for another - one to complete - - `spinlock_t lock;` - - lock used for synchronization - - `atomic_t usage_count;` - - the usage counter of the device - - `atomic_t child_count;` - - the count of 'active' children of the device - - `unsigned int ignore_children;` - - if set, the value of child_count is ignored (but still updated) - - `unsigned int disable_depth;` - - used for disabling the helper functions (they work normally if this is - equal to zero); the initial value of it is 1 (i.e. runtime PM is - initially disabled for all devices) - - `int runtime_error;` - - if set, there was a fatal error (one of the callbacks returned error code - as described in Section 2), so the helper functions will not work until - this flag is cleared; this is the error code returned by the failing - callback - - `unsigned int idle_notification;` - - if set, ->runtime_idle() is being executed - - `unsigned int request_pending;` - - if set, there's a pending request (i.e. a work item queued up into pm_wq) - - `enum rpm_request request;` - - type of request that's pending (valid if request_pending is set) - - `unsigned int deferred_resume;` - - set if ->runtime_resume() is about to be run while ->runtime_suspend() is - being executed for that device and it is not practical to wait for the - suspend to complete; means "start a resume as soon as you've suspended" - - `enum rpm_status runtime_status;` - - the runtime PM status of the device; this field's initial value is - RPM_SUSPENDED, which means that each device is initially regarded by the - PM core as 'suspended', regardless of its real hardware status - - `enum rpm_status last_status;` - - the last runtime PM status of the device captured before disabling runtime - PM for it (invalid initially and when disable_depth is 0) - - `unsigned int runtime_auto;` - - if set, indicates that the user space has allowed the device driver to - power manage the device at run time via the /sys/devices/.../power/control - `interface;` it may only be modified with the help of the - pm_runtime_allow() and pm_runtime_forbid() helper functions - - `unsigned int no_callbacks;` - - indicates that the device does not use the runtime PM callbacks (see - Section 8); it may be modified only by the pm_runtime_no_callbacks() - helper function - - `unsigned int irq_safe;` - - indicates that the ->runtime_suspend() and ->runtime_resume() callbacks - will be invoked with the spinlock held and interrupts disabled - - `unsigned int use_autosuspend;` - - indicates that the device's driver supports delayed autosuspend (see - Section 9); it may be modified only by the - pm_runtime{_dont}_use_autosuspend() helper functions - - `unsigned int timer_autosuspends;` - - indicates that the PM core should attempt to carry out an autosuspend - when the timer expires rather than a normal suspend - - `int autosuspend_delay;` - - the delay time (in milliseconds) to be used for autosuspend - - `unsigned long last_busy;` - - the time (in jiffies) when the pm_runtime_mark_last_busy() helper - function was last called for this device; used in calculating inactivity - periods for autosuspend - -All of the above fields are members of the 'power' member of 'struct device'. +.. kernel-doc:: include/linux/pm.h + :identifiers: dev_pm_info 4. Runtime PM Device Helper Functions ===================================== @@ -307,219 +216,10 @@ All of the above fields are members of the 'power' member of 'struct device'. The following runtime PM helper functions are defined in drivers/base/power/runtime.c and include/linux/pm_runtime.h: - `void pm_runtime_init(struct device *dev);` - - initialize the device runtime PM fields in 'struct dev_pm_info' - - `void pm_runtime_remove(struct device *dev);` - - make sure that the runtime PM of the device will be disabled after - removing the device from device hierarchy - - `int pm_runtime_idle(struct device *dev);` - - execute the subsystem-level idle callback for the device; returns an - error code on failure, where -EINPROGRESS means that ->runtime_idle() is - already being executed; if there is no callback or the callback returns 0 - then run pm_runtime_autosuspend(dev) and return its result - - `int pm_runtime_suspend(struct device *dev);` - - execute the subsystem-level suspend callback for the device; returns 0 on - success, 1 if the device's runtime PM status was already 'suspended', or - error code on failure, where -EAGAIN or -EBUSY means it is safe to attempt - to suspend the device again in future and -EACCES means that - 'power.disable_depth' is different from 0 - - `int pm_runtime_autosuspend(struct device *dev);` - - same as pm_runtime_suspend() except that a call to - pm_runtime_mark_last_busy() is made and an autosuspend is scheduled for - the appropriate time and 0 is returned - - `int pm_runtime_resume(struct device *dev);` - - execute the subsystem-level resume callback for the device; returns 0 on - success, 1 if the device's runtime PM status is already 'active' (also if - 'power.disable_depth' is nonzero, but the status was 'active' when it was - changing from 0 to 1) or error code on failure, where -EAGAIN means it may - be safe to attempt to resume the device again in future, but - 'power.runtime_error' should be checked additionally, and -EACCES means - that the callback could not be run, because 'power.disable_depth' was - different from 0 - - `int pm_runtime_resume_and_get(struct device *dev);` - - run pm_runtime_resume(dev) and if successful, increment the device's - usage counter; returns 0 on success (whether or not the device's - runtime PM status was already 'active') or the error code from - pm_runtime_resume() on failure. - - `int pm_request_idle(struct device *dev);` - - submit a request to execute the subsystem-level idle callback for the - device (the request is represented by a work item in pm_wq); returns 0 on - success or error code if the request has not been queued up - - `int pm_request_autosuspend(struct device *dev);` - - Call pm_runtime_mark_last_busy() and schedule the execution of the - subsystem-level suspend callback for the device when the autosuspend delay - expires - - `int pm_schedule_suspend(struct device *dev, unsigned int delay);` - - schedule the execution of the subsystem-level suspend callback for the - device in future, where 'delay' is the time to wait before queuing up a - suspend work item in pm_wq, in milliseconds (if 'delay' is zero, the work - item is queued up immediately); returns 0 on success, 1 if the device's PM - runtime status was already 'suspended', or error code if the request - hasn't been scheduled (or queued up if 'delay' is 0); if the execution of - ->runtime_suspend() is already scheduled and not yet expired, the new - value of 'delay' will be used as the time to wait - - `int pm_request_resume(struct device *dev);` - - submit a request to execute the subsystem-level resume callback for the - device (the request is represented by a work item in pm_wq); returns 0 on - success, 1 if the device's runtime PM status was already 'active', or - error code if the request hasn't been queued up - - `void pm_runtime_get_noresume(struct device *dev);` - - increment the device's usage counter - - `int pm_runtime_get(struct device *dev);` - - increment the device's usage counter, run pm_request_resume(dev) and - return its result - - `int pm_runtime_get_sync(struct device *dev);` - - increment the device's usage counter, run pm_runtime_resume(dev) and - return its result; - note that it does not drop the device's usage counter on errors, so - consider using pm_runtime_resume_and_get() instead of it, especially - if its return value is checked by the caller, as this is likely to - result in cleaner code. - - `int pm_runtime_get_if_in_use(struct device *dev);` - - return -EINVAL if 'power.disable_depth' is nonzero; otherwise, if the - runtime PM status is RPM_ACTIVE and the runtime PM usage counter is - nonzero, increment the counter and return 1; otherwise return 0 without - changing the counter - - `int pm_runtime_get_if_active(struct device *dev);` - - return -EINVAL if 'power.disable_depth' is nonzero; otherwise, if the - runtime PM status is RPM_ACTIVE, increment the counter and - return 1; otherwise return 0 without changing the counter - - `void pm_runtime_put_noidle(struct device *dev);` - - decrement the device's usage counter - - `int pm_runtime_put(struct device *dev);` - - decrement the device's usage counter; if the result is 0 then run - pm_request_idle(dev) and return its result - - `int pm_runtime_put_autosuspend(struct device *dev);` - - set the power.last_busy field to the current time and decrement the - device's usage counter; if the result is 0 then run - pm_request_autosuspend(dev) and return its result - - `int __pm_runtime_put_autosuspend(struct device *dev);` - - decrement the device's usage counter; if the result is 0 then run - pm_request_autosuspend(dev) and return its result - - `int pm_runtime_put_sync(struct device *dev);` - - decrement the device's usage counter; if the result is 0 then run - pm_runtime_idle(dev) and return its result - - `int pm_runtime_put_sync_suspend(struct device *dev);` - - decrement the device's usage counter; if the result is 0 then run - pm_runtime_suspend(dev) and return its result - - `int pm_runtime_put_sync_autosuspend(struct device *dev);` - - set the power.last_busy field to the current time and decrement the - device's usage counter; if the result is 0 then run - pm_runtime_autosuspend(dev) and return its result - - `void pm_runtime_enable(struct device *dev);` - - decrement the device's 'power.disable_depth' field; if that field is equal - to zero, the runtime PM helper functions can execute subsystem-level - callbacks described in Section 2 for the device - - `int pm_runtime_disable(struct device *dev);` - - increment the device's 'power.disable_depth' field (if the value of that - field was previously zero, this prevents subsystem-level runtime PM - callbacks from being run for the device), make sure that all of the - pending runtime PM operations on the device are either completed or - canceled; returns 1 if there was a resume request pending and it was - necessary to execute the subsystem-level resume callback for the device - to satisfy that request, otherwise 0 is returned - - `void pm_runtime_barrier(struct device *dev);` - - check if there's a resume request pending for the device and resume it - (synchronously) in that case, cancel any other pending runtime PM requests - regarding it and wait for all runtime PM operations on it in progress to - complete - - `void pm_suspend_ignore_children(struct device *dev, bool enable);` - - set/unset the power.ignore_children flag of the device - - `int pm_runtime_set_active(struct device *dev);` - - clear the device's 'power.runtime_error' flag, set the device's runtime - PM status to 'active' and update its parent's counter of 'active' - children as appropriate (it is only valid to use this function if - 'power.runtime_error' is set or 'power.disable_depth' is greater than - zero); it will fail and return error code if the device has a parent - which is not active and the 'power.ignore_children' flag of which is unset - - `void pm_runtime_set_suspended(struct device *dev);` - - clear the device's 'power.runtime_error' flag, set the device's runtime - PM status to 'suspended' and update its parent's counter of 'active' - children as appropriate (it is only valid to use this function if - 'power.runtime_error' is set or 'power.disable_depth' is greater than - zero) - - `bool pm_runtime_active(struct device *dev);` - - return true if the device's runtime PM status is 'active' or its - 'power.disable_depth' field is not equal to zero, or false otherwise - - `bool pm_runtime_suspended(struct device *dev);` - - return true if the device's runtime PM status is 'suspended' and its - 'power.disable_depth' field is equal to zero, or false otherwise - - `bool pm_runtime_status_suspended(struct device *dev);` - - return true if the device's runtime PM status is 'suspended' - - `void pm_runtime_no_callbacks(struct device *dev);` - - set the power.no_callbacks flag for the device and remove the runtime - PM attributes from /sys/devices/.../power (or prevent them from being - added when the device is registered) - - `void pm_runtime_irq_safe(struct device *dev);` - - set the power.irq_safe flag for the device, causing the runtime-PM - callbacks to be invoked with interrupts off - - `bool pm_runtime_is_irq_safe(struct device *dev);` - - return true if power.irq_safe flag was set for the device, causing - the runtime-PM callbacks to be invoked with interrupts off - - `void pm_runtime_mark_last_busy(struct device *dev);` - - set the power.last_busy field to the current time - - `void pm_runtime_use_autosuspend(struct device *dev);` - - set the power.use_autosuspend flag, enabling autosuspend delays; call - pm_runtime_get_sync if the flag was previously cleared and - power.autosuspend_delay is negative - - `void pm_runtime_dont_use_autosuspend(struct device *dev);` - - clear the power.use_autosuspend flag, disabling autosuspend delays; - decrement the device's usage counter if the flag was previously set and - power.autosuspend_delay is negative; call pm_runtime_idle - - `void pm_runtime_set_autosuspend_delay(struct device *dev, int delay);` - - set the power.autosuspend_delay value to 'delay' (expressed in - milliseconds); if 'delay' is negative then runtime suspends are - prevented; if power.use_autosuspend is set, pm_runtime_get_sync may be - called or the device's usage counter may be decremented and - pm_runtime_idle called depending on if power.autosuspend_delay is - changed to or from a negative value; if power.use_autosuspend is clear, - pm_runtime_idle is called - - `unsigned long pm_runtime_autosuspend_expiration(struct device *dev);` - - calculate the time when the current autosuspend delay period will expire, - based on power.last_busy and power.autosuspend_delay; if the delay time - is 1000 ms or larger then the expiration time is rounded up to the - nearest second; returns 0 if the delay period has already expired or - power.use_autosuspend isn't set, otherwise returns the expiration time - in jiffies +.. kernel-doc:: drivers/base/power/runtime.c + :export: + +.. kernel-doc:: include/linux/pm_runtime.h It is safe to execute the following helper functions from interrupt context: @@ -728,67 +428,8 @@ Subsystems may wish to conserve code space by using the set of generic power management callbacks provided by the PM core, defined in driver/base/power/generic_ops.c: - `int pm_generic_runtime_suspend(struct device *dev);` - - invoke the ->runtime_suspend() callback provided by the driver of this - device and return its result, or return 0 if not defined - - `int pm_generic_runtime_resume(struct device *dev);` - - invoke the ->runtime_resume() callback provided by the driver of this - device and return its result, or return 0 if not defined - - `int pm_generic_suspend(struct device *dev);` - - if the device has not been suspended at run time, invoke the ->suspend() - callback provided by its driver and return its result, or return 0 if not - defined - - `int pm_generic_suspend_noirq(struct device *dev);` - - if pm_runtime_suspended(dev) returns "false", invoke the ->suspend_noirq() - callback provided by the device's driver and return its result, or return - 0 if not defined - - `int pm_generic_resume(struct device *dev);` - - invoke the ->resume() callback provided by the driver of this device and, - if successful, change the device's runtime PM status to 'active' - - `int pm_generic_resume_noirq(struct device *dev);` - - invoke the ->resume_noirq() callback provided by the driver of this device - - `int pm_generic_freeze(struct device *dev);` - - if the device has not been suspended at run time, invoke the ->freeze() - callback provided by its driver and return its result, or return 0 if not - defined - - `int pm_generic_freeze_noirq(struct device *dev);` - - if pm_runtime_suspended(dev) returns "false", invoke the ->freeze_noirq() - callback provided by the device's driver and return its result, or return - 0 if not defined - - `int pm_generic_thaw(struct device *dev);` - - if the device has not been suspended at run time, invoke the ->thaw() - callback provided by its driver and return its result, or return 0 if not - defined - - `int pm_generic_thaw_noirq(struct device *dev);` - - if pm_runtime_suspended(dev) returns "false", invoke the ->thaw_noirq() - callback provided by the device's driver and return its result, or return - 0 if not defined - - `int pm_generic_poweroff(struct device *dev);` - - if the device has not been suspended at run time, invoke the ->poweroff() - callback provided by its driver and return its result, or return 0 if not - defined - - `int pm_generic_poweroff_noirq(struct device *dev);` - - if pm_runtime_suspended(dev) returns "false", run the ->poweroff_noirq() - callback provided by the device's driver and return its result, or return - 0 if not defined - - `int pm_generic_restore(struct device *dev);` - - invoke the ->restore() callback provided by the driver of this device and, - if successful, change the device's runtime PM status to 'active' - - `int pm_generic_restore_noirq(struct device *dev);` - - invoke the ->restore_noirq() callback provided by the device's driver +.. kernel-doc:: drivers/base/power/generic_ops.c + :export: These functions are the defaults used by the PM core if a subsystem doesn't provide its own callbacks for ->runtime_idle(), ->runtime_suspend(), diff --git a/Documentation/power/userland-swsusp.rst b/Documentation/power/userland-swsusp.rst index 1cf62d80a9ca..b00b850148b2 100644 --- a/Documentation/power/userland-swsusp.rst +++ b/Documentation/power/userland-swsusp.rst @@ -4,9 +4,9 @@ Documentation for userland software suspend interface (C) 2006 Rafael J. Wysocki <rjw@sisk.pl> -First, the warnings at the beginning of swsusp.txt still apply. +First, the warnings at the beginning of swsusp.rst still apply. -Second, you should read the FAQ in swsusp.txt _now_ if you have not +Second, you should read the FAQ in swsusp.rst _now_ if you have not done it already. Now, to use the userland interface for software suspend you need special diff --git a/arch/arm64/kernel/topology.c b/arch/arm64/kernel/topology.c index 1773dac6fdf8..928079332fe6 100644 --- a/arch/arm64/kernel/topology.c +++ b/arch/arm64/kernel/topology.c @@ -425,7 +425,7 @@ int counters_read_on_cpu(int cpu, smp_call_func_t func, void *val) return -EPERM; func(val); } else { - smp_call_function_single(cpu, func, val, 1); + return smp_call_function_single(cpu, func, val, 1); } return 0; @@ -468,6 +468,12 @@ static void amu_read_core_const_ctrs(void *val) cpu_read_corecnt(&ctrs->corecnt); } +static bool cpc_ffh_reg_valid(const struct cpc_reg *reg) +{ + return reg->bit_width && reg->bit_width <= 64 && + reg->bit_offset <= 64 - reg->bit_width; +} + static u64 cpc_ffh_extract_bits(const struct cpc_reg *reg, u64 val) { val &= GENMASK_ULL(reg->bit_offset + reg->bit_width - 1, @@ -504,7 +510,8 @@ int cpc_read_ffh_fb_ctrs(int cpu, struct cpc_reg *reg1, u64 *val1, struct amu_ffh_ctrs ctrs; int ret; - if (!is_amu_ctr_reg(reg1) || !is_amu_ctr_reg(reg2)) + if (!is_amu_ctr_reg(reg1) || !is_amu_ctr_reg(reg2) || + !cpc_ffh_reg_valid(reg1) || !cpc_ffh_reg_valid(reg2)) return -EINVAL; ret = counters_read_on_cpu(cpu, amu_read_core_const_ctrs, &ctrs); @@ -528,6 +535,9 @@ int cpc_read_ffh(int cpu, struct cpc_reg *reg, u64 *val) { int ret = -EOPNOTSUPP; + if (!cpc_ffh_reg_valid(reg)) + return -EINVAL; + switch ((u64)reg->address) { case CPC_FFH_CTR_CORE: ret = counters_read_on_cpu(cpu, cpu_read_corecnt, val); diff --git a/arch/x86/kernel/acpi/cppc.c b/arch/x86/kernel/acpi/cppc.c index bbade0da5130..d2185b7a2030 100644 --- a/arch/x86/kernel/acpi/cppc.c +++ b/arch/x86/kernel/acpi/cppc.c @@ -5,6 +5,7 @@ */ #include <linux/bitfield.h> +#include <linux/limits.h> #include <acpi/cppc_acpi.h> #include <asm/msr.h> @@ -45,10 +46,20 @@ bool cpc_ffh_supported(void) return true; } +static bool cpc_ffh_reg_valid(const struct cpc_reg *reg) +{ + return reg->address <= U32_MAX && reg->bit_width && + reg->bit_width <= 64 && + reg->bit_offset <= 64 - reg->bit_width; +} + int cpc_read_ffh(int cpunum, struct cpc_reg *reg, u64 *val) { int err; + if (!cpc_ffh_reg_valid(reg)) + return -EINVAL; + err = rdmsrq_safe_on_cpu(cpunum, reg->address, val); if (!err) { u64 mask = GENMASK_ULL(reg->bit_offset + reg->bit_width - 1, @@ -65,6 +76,9 @@ int cpc_write_ffh(int cpunum, struct cpc_reg *reg, u64 val) u64 rd_val; int err; + if (!cpc_ffh_reg_valid(reg)) + return -EINVAL; + err = rdmsrq_safe_on_cpu(cpunum, reg->address, &rd_val); if (!err) { u64 mask = GENMASK_ULL(reg->bit_offset + reg->bit_width - 1, diff --git a/drivers/acpi/acpi_extlog.c b/drivers/acpi/acpi_extlog.c index 7ad3b36013cc..9e61354a807b 100644 --- a/drivers/acpi/acpi_extlog.c +++ b/drivers/acpi/acpi_extlog.c @@ -134,22 +134,42 @@ static int print_extlog_rcd(const char *pfx, } static void extlog_print_pcie(struct cper_sec_pcie *pcie_err, - int severity) + int severity, u32 len) { -#ifdef ACPI_APEI_PCIEAER - struct aer_capability_regs *aer; +#ifdef CONFIG_ACPI_APEI_PCIEAER + struct aer_capability_regs aer_regs = {}; struct pci_dev *pdev; unsigned int devfn; unsigned int bus; int aer_severity; int domain; + if (len < sizeof(*pcie_err)) { + pr_warn_ratelimited(FW_WARN + "PCIe error section too small (%u)\n", len); + return; + } + if (!(pcie_err->validation_bits & CPER_PCIE_VALID_DEVICE_ID && pcie_err->validation_bits & CPER_PCIE_VALID_AER_INFO)) return; aer_severity = cper_severity_to_aer(severity); - aer = (struct aer_capability_regs *)pcie_err->aer_info; + + /* + * struct pcie_tlp_log is larger than the hardware layout, so aer_info + * only maps onto the struct up to the four Header Log DWORDs. Copy that + * much, then place the TLP Prefix Log from where the hardware keeps it. + * Everything else stays zero: nothing reads root_command, root_status or + * the error source IDs, and header_len and flit are software-only. + */ + memcpy(&aer_regs, pcie_err->aer_info, + offsetof(struct aer_capability_regs, header_log) + + PCIE_STD_NUM_TLP_HEADERLOG * sizeof(u32)); + memcpy(aer_regs.header_log.prefix, + pcie_err->aer_info + PCI_ERR_PREFIX_LOG, + sizeof(aer_regs.header_log.prefix)); + domain = pcie_err->device_id.segment; bus = pcie_err->device_id.bus; devfn = PCI_DEVFN(pcie_err->device_id.device, @@ -158,28 +178,11 @@ static void extlog_print_pcie(struct cper_sec_pcie *pcie_err, if (!pdev) return; - pci_print_aer(pdev, aer_severity, aer); + pci_print_aer(pdev, aer_severity, &aer_regs); pci_dev_put(pdev); #endif } -static void -extlog_cxl_cper_handle_prot_err(struct cxl_cper_sec_prot_err *prot_err, - int severity) -{ -#ifdef ACPI_APEI_PCIEAER - struct cxl_cper_prot_err_work_data wd; - - if (cxl_cper_sec_prot_err_valid(prot_err)) - return; - - if (cxl_cper_setup_prot_err_work_data(&wd, prot_err, severity)) - return; - - cxl_cper_handle_prot_err(&wd); -#endif -} - static int extlog_print(struct notifier_block *nb, unsigned long val, void *data) { @@ -208,6 +211,15 @@ static int extlog_print(struct notifier_block *nb, unsigned long val, tmp = (struct acpi_hest_generic_status *)elog_buf; + /* + * Bound the length before cper_estatus_check() walks the sections: it + * iterates over data_length, which is not yet known to fit elog_buf. + * cper_estatus_check_header() then rejects a length that wrapped, which + * the bound cannot see. + */ + if (cper_estatus_len(tmp) > ELOG_ENTRY_LEN || cper_estatus_check(tmp)) + return NOTIFY_DONE; + if (!ras_userspace_consumers()) { print_extlog_rcd(NULL, tmp, cpu); goto out; @@ -235,12 +247,14 @@ static int extlog_print(struct notifier_block *nb, unsigned long val, struct cxl_cper_sec_prot_err *prot_err = acpi_hest_get_payload(gdata); - extlog_cxl_cper_handle_prot_err(prot_err, - gdata->error_severity); + cxl_cper_post_prot_err(prot_err, + gdata->error_severity, + gdata->error_data_length); } else if (guid_equal(sec_type, &CPER_SEC_PCIE)) { struct cper_sec_pcie *pcie_err = acpi_hest_get_payload(gdata); - extlog_print_pcie(pcie_err, gdata->error_severity); + extlog_print_pcie(pcie_err, gdata->error_severity, + gdata->error_data_length); } else { void *err = acpi_hest_get_payload(gdata); diff --git a/drivers/acpi/acpi_mrrm.c b/drivers/acpi/acpi_mrrm.c index e99cbda83452..83934e2aa0d4 100644 --- a/drivers/acpi/acpi_mrrm.c +++ b/drivers/acpi/acpi_mrrm.c @@ -36,17 +36,11 @@ static u32 mrrm_mem_entry_num; static int get_node_num(struct mrrm_mem_range_entry *e) { - unsigned int nid; + struct zone *zone; - for_each_online_node(nid) { - for (int z = 0; z < MAX_NR_ZONES; z++) { - struct zone *zone = NODE_DATA(nid)->node_zones + z; - - if (!populated_zone(zone)) - continue; - if (zone_intersects(zone, PHYS_PFN(e->base), PHYS_PFN(e->length))) - return zone_to_nid(zone); - } + for_each_populated_zone(zone) { + if (zone_intersects(zone, PHYS_PFN(e->base), PHYS_PFN(e->length))) + return zone_to_nid(zone); } return -ENOENT; diff --git a/drivers/acpi/acpi_pcc.c b/drivers/acpi/acpi_pcc.c index 438c67189511..345f233d77cd 100644 --- a/drivers/acpi/acpi_pcc.c +++ b/drivers/acpi/acpi_pcc.c @@ -28,6 +28,7 @@ * to PCC commands */ #define PCC_CMD_WAIT_RETRIES_NUM 500ULL +#define PCC_SIGNATURE_SIZE sizeof(u32) struct pcc_data { struct pcc_mbox_chan *pcc_chan; @@ -49,10 +50,24 @@ static acpi_status acpi_pcc_address_space_setup(acpi_handle region_handle, u32 function, void *handler_context, void **region_context) { - struct pcc_data *data; struct acpi_pcc_info *ctx = handler_context; struct pcc_mbox_chan *pcc_chan; + struct pcc_data *data; acpi_status ret; + u64 usecs_lat; + + if (function == ACPI_REGION_DEACTIVATE) { + data = *region_context; + if (data) { + pcc_mbox_free_channel(data->pcc_chan); + kfree(data); + *region_context = NULL; + } + return AE_OK; + } + + if (function != ACPI_REGION_ACTIVATE) + return AE_BAD_PARAMETER; data = kzalloc_obj(*data); if (!data) @@ -74,6 +89,14 @@ acpi_pcc_address_space_setup(acpi_handle region_handle, u32 function, } pcc_chan = data->pcc_chan; + if (pcc_chan->shmem_size < PCC_SIGNATURE_SIZE || + ctx->length > pcc_chan->shmem_size - PCC_SIGNATURE_SIZE) { + pr_err("PCC channel-%d shared memory is too small.\n", + ctx->subspace_id); + ret = AE_AML_REGION_LIMIT; + goto err_free_channel; + } + if (!pcc_chan->mchan->mbox->txdone_irq) { pr_err("This channel-%d does not support interrupt.\n", ctx->subspace_id); @@ -81,6 +104,16 @@ acpi_pcc_address_space_setup(acpi_handle region_handle, u32 function, goto err_free_channel; } + /* + * pcc_chan->latency is just a Nominal value. In reality the remote + * processor could be much slower to reply. So add an arbitrary + * amount of wait on top of Nominal. + */ + usecs_lat = PCC_CMD_WAIT_RETRIES_NUM * pcc_chan->latency; + data->cl.tx_tout = DIV_ROUND_UP_ULL(usecs_lat, 1000); + if (!data->cl.tx_tout) + data->cl.tx_tout = 1; + *region_context = data; return AE_OK; @@ -97,27 +130,23 @@ acpi_pcc_address_space_handler(u32 function, acpi_physical_address addr, u32 bits, acpi_integer *value, void *handler_context, void *region_context) { - int ret; struct pcc_data *data = region_context; - u64 usecs_lat; + void __iomem *pcc_opregion; + int ret; + + pcc_opregion = data->pcc_chan->shmem + PCC_SIGNATURE_SIZE; reinit_completion(&data->done); - /* Write to Shared Memory */ - memcpy_toio(data->pcc_chan->shmem, (void *)value, data->ctx.length); + /* Write to the PCC OperationRegion after the shared memory signature. */ + memcpy_toio(pcc_opregion, (void *)value, data->ctx.length); ret = mbox_send_message(data->pcc_chan->mchan, NULL); if (ret < 0) return AE_ERROR; - /* - * pcc_chan->latency is just a Nominal value. In reality the remote - * processor could be much slower to reply. So add an arbitrary - * amount of wait on top of Nominal. - */ - usecs_lat = PCC_CMD_WAIT_RETRIES_NUM * data->pcc_chan->latency; ret = wait_for_completion_timeout(&data->done, - usecs_to_jiffies(usecs_lat)); + msecs_to_jiffies(data->cl.tx_tout)); if (ret == 0) { pr_err("PCC command executed timeout!\n"); return AE_TIME; @@ -125,7 +154,7 @@ acpi_pcc_address_space_handler(u32 function, acpi_physical_address addr, mbox_chan_txdone(data->pcc_chan->mchan, ret); - memcpy_fromio(value, data->pcc_chan->shmem, data->ctx.length); + memcpy_fromio(value, pcc_opregion, data->ctx.length); return AE_OK; } diff --git a/drivers/acpi/acpi_video.c b/drivers/acpi/acpi_video.c index 4d6fd9f6e9ad..2a2822516665 100644 --- a/drivers/acpi/acpi_video.c +++ b/drivers/acpi/acpi_video.c @@ -1700,9 +1700,9 @@ static int acpi_video_resume(struct notifier_block *nb, static void acpi_video_dev_register_backlight(struct acpi_video_device *device) { + struct device *phys_dev, *parent = NULL; struct backlight_properties props; struct pci_dev *pdev; - struct device *parent = NULL; int result; static int count; char *name; @@ -1729,11 +1729,10 @@ static void acpi_video_dev_register_backlight(struct acpi_video_device *device) device, &acpi_backlight_ops, &props); - put_device(parent); kfree(name); if (IS_ERR(device->backlight)) { device->backlight = NULL; - return; + goto put_parent; } /* @@ -1743,8 +1742,12 @@ static void acpi_video_dev_register_backlight(struct acpi_video_device *device) device->backlight->props.brightness = acpi_video_get_brightness(device->backlight); - device->cooling_dev = thermal_cooling_device_register("LCD", device, - &video_cooling_ops); + phys_dev = acpi_bus_get_primary_device(device->dev); + if (!phys_dev) + phys_dev = get_device(parent); + + device->cooling_dev = thermal_cooling_device_create(phys_dev, "LCD", device, + &video_cooling_ops); if (IS_ERR(device->cooling_dev)) { /* * Set cooling_dev to NULL so we don't crash trying to free it. @@ -1753,21 +1756,17 @@ static void acpi_video_dev_register_backlight(struct acpi_video_device *device) * -- dtor */ device->cooling_dev = NULL; - return; + goto put_phys_dev; } - dev_info(&device->dev->dev, "registered as cooling_device%d\n", - device->cooling_dev->id); - result = sysfs_create_link(&device->dev->dev.kobj, - &device->cooling_dev->device.kobj, - "thermal_cooling"); - if (result) - pr_info("sysfs link creation failed\n"); + dev_info(&device->cooling_dev->device, "Using ACPI device %s\n", + acpi_dev_name(device->dev)); - result = sysfs_create_link(&device->cooling_dev->device.kobj, - &device->dev->dev.kobj, "device"); - if (result) - pr_info("Reverse sysfs link creation failed\n"); +put_phys_dev: + put_device(phys_dev); + +put_parent: + put_device(parent); } static void acpi_video_run_bcl_for_osi(struct acpi_video_bus *video) @@ -1825,6 +1824,10 @@ static int acpi_video_bus_register_backlight(struct acpi_video_bus *video) static void acpi_video_dev_unregister_backlight(struct acpi_video_device *device) { + if (device->cooling_dev) { + thermal_cooling_device_unregister(device->cooling_dev); + device->cooling_dev = NULL; + } if (device->backlight) { backlight_device_unregister(device->backlight); device->backlight = NULL; @@ -1834,12 +1837,6 @@ static void acpi_video_dev_unregister_backlight(struct acpi_video_device *device kfree(device->brightness); device->brightness = NULL; } - if (device->cooling_dev) { - sysfs_remove_link(&device->dev->dev.kobj, "thermal_cooling"); - sysfs_remove_link(&device->cooling_dev->device.kobj, "device"); - thermal_cooling_device_unregister(device->cooling_dev); - device->cooling_dev = NULL; - } } static int acpi_video_bus_unregister_backlight(struct acpi_video_bus *video) diff --git a/drivers/acpi/acpica/dsfield.c b/drivers/acpi/acpica/dsfield.c index d7d56cc600d3..f7e503bf57c8 100644 --- a/drivers/acpi/acpica/dsfield.c +++ b/drivers/acpi/acpica/dsfield.c @@ -522,7 +522,8 @@ acpi_ds_create_field(union acpi_parse_object *op, } if (info.region_node->object->region.space_id == - ACPI_ADR_SPACE_PLATFORM_COMM) { + ACPI_ADR_SPACE_PLATFORM_COMM && + !region_node->object->field.internal_pcc_buffer) { region_node->object->field.internal_pcc_buffer = ACPI_ALLOCATE_ZEROED(info.region_node->object->region. length); diff --git a/drivers/acpi/acpica/exfield.c b/drivers/acpi/acpica/exfield.c index 9a55524ed8f4..50317fe0cbd5 100644 --- a/drivers/acpi/acpica/exfield.c +++ b/drivers/acpi/acpica/exfield.c @@ -45,12 +45,13 @@ static const u8 acpi_protocol_lengths[] = { /* * The following macros determine a given offset is a COMD field. - * According to the specification, generic subspaces (types 0-2) contains a - * 2-byte COMD field at offset 4 and master subspaces (type 3) contains a 4-byte - * COMD field starting at offset 12. + * According to the specification, the PCC OperationRegion begins after + * the PCC signature. The raw shared memory COMD offsets of 4 for generic + * subspaces (types 0-2) and 12 for master subspaces (type 3) therefore + * appear at OperationRegion offsets 0 and 8. */ -#define GENERIC_SUBSPACE_COMMAND(a) (4 == a || a == 5) -#define MASTER_SUBSPACE_COMMAND(a) (12 <= a && a <= 15) +#define GENERIC_SUBSPACE_COMMAND(a) (0 == a || a == 1) +#define MASTER_SUBSPACE_COMMAND(a) (8 <= a && a <= 11) /******************************************************************************* * diff --git a/drivers/acpi/apei/ghes.c b/drivers/acpi/apei/ghes.c index fe10ab0e02f6..08c985e729d6 100644 --- a/drivers/acpi/apei/ghes.c +++ b/drivers/acpi/apei/ghes.c @@ -642,11 +642,14 @@ static void ghes_handle_aer(struct acpi_hest_generic_data *gdata) #ifdef CONFIG_ACPI_APEI_PCIEAER struct cper_sec_pcie *pcie_err = acpi_hest_get_payload(gdata); + if (gdata->error_data_length < sizeof(*pcie_err)) + return; + if (pcie_err->validation_bits & CPER_PCIE_VALID_DEVICE_ID && pcie_err->validation_bits & CPER_PCIE_VALID_AER_INFO) { + struct aer_capability_regs *aer_info; unsigned int devfn; int aer_severity; - u8 *aer_info; devfn = PCI_DEVFN(pcie_err->device_id.device, pcie_err->device_id.function); @@ -664,13 +667,25 @@ static void ghes_handle_aer(struct acpi_hest_generic_data *gdata) sizeof(struct aer_capability_regs)); if (!aer_info) return; - memcpy(aer_info, pcie_err->aer_info, sizeof(struct aer_capability_regs)); + + /* + * Map aer_info onto the struct as extlog_print_pcie() does: + * copy up to the four Header Log DWORDs, then place the TLP + * Prefix Log from where the hardware keeps it. The rest stays + * zero, so firmware cannot drive the pcie_print_tlp_log() loop + * over dw[] out of bounds. + */ + memset(aer_info, 0, sizeof(struct aer_capability_regs)); + memcpy(aer_info, pcie_err->aer_info, + offsetof(struct aer_capability_regs, header_log) + + PCIE_STD_NUM_TLP_HEADERLOG * sizeof(u32)); + memcpy(aer_info->header_log.prefix, + pcie_err->aer_info + PCI_ERR_PREFIX_LOG, + sizeof(aer_info->header_log.prefix)); aer_recover_queue(pcie_err->device_id.segment, pcie_err->device_id.bus, - devfn, aer_severity, - (struct aer_capability_regs *) - aer_info); + devfn, aer_severity, aer_info); } #endif } @@ -752,13 +767,13 @@ static DEFINE_KFIFO(cxl_cper_prot_err_fifo, struct cxl_cper_prot_err_work_data, static DEFINE_RAW_SPINLOCK(cxl_cper_prot_err_work_lock); struct work_struct *cxl_cper_prot_err_work; -static void cxl_cper_post_prot_err(struct cxl_cper_sec_prot_err *prot_err, - int severity) +void cxl_cper_post_prot_err(struct cxl_cper_sec_prot_err *prot_err, + int severity, u32 len) { #ifdef CONFIG_ACPI_APEI_PCIEAER struct cxl_cper_prot_err_work_data wd; - if (cxl_cper_sec_prot_err_valid(prot_err)) + if (cxl_cper_sec_prot_err_valid(prot_err, len)) return; guard(raw_spinlock_irqsave)(&cxl_cper_prot_err_work_lock); @@ -777,6 +792,7 @@ static void cxl_cper_post_prot_err(struct cxl_cper_sec_prot_err *prot_err, schedule_work(cxl_cper_prot_err_work); #endif } +EXPORT_SYMBOL_FOR_MODULES(cxl_cper_post_prot_err, "acpi_extlog"); void cxl_cper_register_prot_err_work(struct work_struct *work) { @@ -823,10 +839,15 @@ static DEFINE_RAW_SPINLOCK(cxl_cper_work_lock); struct work_struct *cxl_cper_work; static void cxl_cper_post_event(enum cxl_event_type event_type, - struct cxl_cper_event_rec *rec) + struct cxl_cper_event_rec *rec, u32 len) { struct cxl_cper_work_data wd; + if (len < sizeof(*rec)) { + pr_err(FW_WARN "CXL CPER section too small (%u)\n", len); + return; + } + if (rec->hdr.length <= sizeof(rec->hdr) || rec->hdr.length > sizeof(*rec)) { pr_err(FW_WARN "CXL CPER Invalid section length (%u)\n", @@ -923,6 +944,28 @@ static void ghes_log_hwerr(int sev, guid_t *sec_type) hwerr_log_error_type(HWERR_RECOV_OTHERS); } +/* + * The fields from "extended" on are absent from the 73-byte UEFI 2.1/2.2 + * layout that older firmware still emits. Return the length needed for the + * fields the validation bits claim, so over-claiming is rejected without + * rejecting an honest short record. + */ +static u32 ghes_mem_err_min_len(u64 validation_bits) +{ + u32 len = sizeof(struct cper_sec_mem_err_old); + + if (validation_bits & (CPER_MEM_VALID_ROW_EXT | CPER_MEM_VALID_CHIP_ID)) + len = offsetof(struct cper_sec_mem_err, rank); + if (validation_bits & CPER_MEM_VALID_RANK_NUMBER) + len = offsetof(struct cper_sec_mem_err, mem_array_handle); + if (validation_bits & CPER_MEM_VALID_CARD_HANDLE) + len = offsetof(struct cper_sec_mem_err, mem_dev_handle); + if (validation_bits & CPER_MEM_VALID_MODULE_HANDLE) + len = sizeof(struct cper_sec_mem_err); + + return len; +} + static void ghes_do_proc(struct ghes *ghes, const struct acpi_hest_generic_status *estatus) { @@ -948,6 +991,25 @@ static void ghes_do_proc(struct ghes *ghes, if (guid_equal(sec_type, &CPER_SEC_PLATFORM_MEM)) { struct cper_sec_mem_err *mem_err = acpi_hest_get_payload(gdata); + /* + * Check once for all three consumers below. The 73-byte + * UEFI 2.1/2.2 layout is the floor, matching + * cper_estatus_print_section() and making + * validation_bits safe to read. + */ + if (gdata->error_data_length < + sizeof(struct cper_sec_mem_err_old)) + continue; + + /* Then require what the claimed fields actually need. */ + if (gdata->error_data_length < + ghes_mem_err_min_len(mem_err->validation_bits)) { + pr_warn_ratelimited(FW_WARN GHES_PFX + "memory error section too small (%u) for the fields it claims\n", + gdata->error_data_length); + continue; + } + atomic_notifier_call_chain(&ghes_report_chain, sev, mem_err); arch_apei_report_mem_error(sev, mem_err); @@ -959,19 +1021,23 @@ static void ghes_do_proc(struct ghes *ghes, } else if (guid_equal(sec_type, &CPER_SEC_CXL_PROT_ERR)) { struct cxl_cper_sec_prot_err *prot_err = acpi_hest_get_payload(gdata); - cxl_cper_post_prot_err(prot_err, gdata->error_severity); + cxl_cper_post_prot_err(prot_err, gdata->error_severity, + gdata->error_data_length); } else if (guid_equal(sec_type, &CPER_SEC_CXL_GEN_MEDIA_GUID)) { struct cxl_cper_event_rec *rec = acpi_hest_get_payload(gdata); - cxl_cper_post_event(CXL_CPER_EVENT_GEN_MEDIA, rec); + cxl_cper_post_event(CXL_CPER_EVENT_GEN_MEDIA, rec, + gdata->error_data_length); } else if (guid_equal(sec_type, &CPER_SEC_CXL_DRAM_GUID)) { struct cxl_cper_event_rec *rec = acpi_hest_get_payload(gdata); - cxl_cper_post_event(CXL_CPER_EVENT_DRAM, rec); + cxl_cper_post_event(CXL_CPER_EVENT_DRAM, rec, + gdata->error_data_length); } else if (guid_equal(sec_type, &CPER_SEC_CXL_MEM_MODULE_GUID)) { struct cxl_cper_event_rec *rec = acpi_hest_get_payload(gdata); - cxl_cper_post_event(CXL_CPER_EVENT_MEM_MODULE, rec); + cxl_cper_post_event(CXL_CPER_EVENT_MEM_MODULE, rec, + gdata->error_data_length); } else { void *err = acpi_hest_get_payload(gdata); diff --git a/drivers/acpi/apei/ghes_helpers.c b/drivers/acpi/apei/ghes_helpers.c index bc7111b740af..df41b993f413 100644 --- a/drivers/acpi/apei/ghes_helpers.c +++ b/drivers/acpi/apei/ghes_helpers.c @@ -5,8 +5,15 @@ #include <linux/aer.h> #include <cxl/event.h> -int cxl_cper_sec_prot_err_valid(struct cxl_cper_sec_prot_err *prot_err) +int cxl_cper_sec_prot_err_valid(struct cxl_cper_sec_prot_err *prot_err, u32 len) { + if (len < sizeof(*prot_err)) { + pr_err_ratelimited(FW_WARN + "CXL CPER prot err section too small (%u)\n", + len); + return -EINVAL; + } + if (!(prot_err->valid_bits & PROT_ERR_VALID_AGENT_ADDRESS)) { pr_err_ratelimited("CXL CPER invalid agent type\n"); return -EINVAL; @@ -23,6 +30,15 @@ int cxl_cper_sec_prot_err_valid(struct cxl_cper_sec_prot_err *prot_err) return -EINVAL; } + /* The RAS Capability block sits after a firmware-sized DVSEC. */ + if (sizeof(*prot_err) + prot_err->dvsec_len + + sizeof(struct cxl_ras_capability_regs) > len) { + pr_err_ratelimited(FW_WARN + "CXL CPER prot err DVSEC (%u) overruns section (%u)\n", + prot_err->dvsec_len, len); + return -EINVAL; + } + if ((prot_err->agent_type == RCD || prot_err->agent_type == DEVICE || prot_err->agent_type == LD || prot_err->agent_type == FMLD) && !(prot_err->valid_bits & PROT_ERR_VALID_SERIAL_NUMBER)) diff --git a/drivers/acpi/arm64/amba.c b/drivers/acpi/arm64/amba.c index 1350083bce5f..7ee431d9464a 100644 --- a/drivers/acpi/arm64/amba.c +++ b/drivers/acpi/arm64/amba.c @@ -87,11 +87,12 @@ static int amba_handler_attach(struct acpi_device *adev, * the amba device we are about to create. */ if (parent) - dev->dev.parent = acpi_get_first_physical_node(parent); + dev->dev.parent = acpi_bus_get_primary_device(parent); device_set_node(&dev->dev, acpi_fwnode_handle(adev)); ret = amba_device_add(dev, &iomem_resource); + put_device(dev->dev.parent); if (ret) { dev_err(&adev->dev, "%s(): amba_device_add() failed (%d)\n", __func__, ret); diff --git a/drivers/acpi/battery.c b/drivers/acpi/battery.c index 670853ec3a4d..8599949f8786 100644 --- a/drivers/acpi/battery.c +++ b/drivers/acpi/battery.c @@ -341,7 +341,9 @@ static int acpi_battery_get_property(struct power_supply *psy, return ret; } -static const enum power_supply_property charge_battery_props[] = { +/* For devices supporting the _BIX ACPI control method */ + +static const enum power_supply_property charge_battery_extended_props[] = { POWER_SUPPLY_PROP_STATUS, POWER_SUPPLY_PROP_PRESENT, POWER_SUPPLY_PROP_TECHNOLOGY, @@ -359,7 +361,7 @@ static const enum power_supply_property charge_battery_props[] = { POWER_SUPPLY_PROP_SERIAL_NUMBER, }; -static const enum power_supply_property charge_battery_full_cap_broken_props[] = { +static const enum power_supply_property charge_battery_full_cap_broken_extended_props[] = { POWER_SUPPLY_PROP_STATUS, POWER_SUPPLY_PROP_PRESENT, POWER_SUPPLY_PROP_TECHNOLOGY, @@ -373,7 +375,7 @@ static const enum power_supply_property charge_battery_full_cap_broken_props[] = POWER_SUPPLY_PROP_SERIAL_NUMBER, }; -static const enum power_supply_property energy_battery_props[] = { +static const enum power_supply_property energy_battery_extended_props[] = { POWER_SUPPLY_PROP_STATUS, POWER_SUPPLY_PROP_PRESENT, POWER_SUPPLY_PROP_TECHNOLOGY, @@ -391,7 +393,7 @@ static const enum power_supply_property energy_battery_props[] = { POWER_SUPPLY_PROP_SERIAL_NUMBER, }; -static const enum power_supply_property energy_battery_full_cap_broken_props[] = { +static const enum power_supply_property energy_battery_full_cap_broken_extended_props[] = { POWER_SUPPLY_PROP_STATUS, POWER_SUPPLY_PROP_PRESENT, POWER_SUPPLY_PROP_TECHNOLOGY, @@ -405,6 +407,68 @@ static const enum power_supply_property energy_battery_full_cap_broken_props[] = POWER_SUPPLY_PROP_SERIAL_NUMBER, }; +/* For devices supporting only the _BIF ACPI control method */ + +static const enum power_supply_property charge_battery_props[] = { + POWER_SUPPLY_PROP_STATUS, + POWER_SUPPLY_PROP_PRESENT, + POWER_SUPPLY_PROP_TECHNOLOGY, + POWER_SUPPLY_PROP_VOLTAGE_MIN_DESIGN, + POWER_SUPPLY_PROP_VOLTAGE_NOW, + POWER_SUPPLY_PROP_CURRENT_NOW, + POWER_SUPPLY_PROP_CHARGE_FULL_DESIGN, + POWER_SUPPLY_PROP_CHARGE_FULL, + POWER_SUPPLY_PROP_CHARGE_NOW, + POWER_SUPPLY_PROP_CAPACITY, + POWER_SUPPLY_PROP_CAPACITY_LEVEL, + POWER_SUPPLY_PROP_MODEL_NAME, + POWER_SUPPLY_PROP_MANUFACTURER, + POWER_SUPPLY_PROP_SERIAL_NUMBER, +}; + +static const enum power_supply_property charge_battery_full_cap_broken_props[] = { + POWER_SUPPLY_PROP_STATUS, + POWER_SUPPLY_PROP_PRESENT, + POWER_SUPPLY_PROP_TECHNOLOGY, + POWER_SUPPLY_PROP_VOLTAGE_MIN_DESIGN, + POWER_SUPPLY_PROP_VOLTAGE_NOW, + POWER_SUPPLY_PROP_CURRENT_NOW, + POWER_SUPPLY_PROP_CHARGE_NOW, + POWER_SUPPLY_PROP_MODEL_NAME, + POWER_SUPPLY_PROP_MANUFACTURER, + POWER_SUPPLY_PROP_SERIAL_NUMBER, +}; + +static const enum power_supply_property energy_battery_props[] = { + POWER_SUPPLY_PROP_STATUS, + POWER_SUPPLY_PROP_PRESENT, + POWER_SUPPLY_PROP_TECHNOLOGY, + POWER_SUPPLY_PROP_VOLTAGE_MIN_DESIGN, + POWER_SUPPLY_PROP_VOLTAGE_NOW, + POWER_SUPPLY_PROP_POWER_NOW, + POWER_SUPPLY_PROP_ENERGY_FULL_DESIGN, + POWER_SUPPLY_PROP_ENERGY_FULL, + POWER_SUPPLY_PROP_ENERGY_NOW, + POWER_SUPPLY_PROP_CAPACITY, + POWER_SUPPLY_PROP_CAPACITY_LEVEL, + POWER_SUPPLY_PROP_MODEL_NAME, + POWER_SUPPLY_PROP_MANUFACTURER, + POWER_SUPPLY_PROP_SERIAL_NUMBER, +}; + +static const enum power_supply_property energy_battery_full_cap_broken_props[] = { + POWER_SUPPLY_PROP_STATUS, + POWER_SUPPLY_PROP_PRESENT, + POWER_SUPPLY_PROP_TECHNOLOGY, + POWER_SUPPLY_PROP_VOLTAGE_MIN_DESIGN, + POWER_SUPPLY_PROP_VOLTAGE_NOW, + POWER_SUPPLY_PROP_POWER_NOW, + POWER_SUPPLY_PROP_ENERGY_NOW, + POWER_SUPPLY_PROP_MODEL_NAME, + POWER_SUPPLY_PROP_MANUFACTURER, + POWER_SUPPLY_PROP_SERIAL_NUMBER, +}; + /* Battery Management */ struct acpi_offsets { size_t offset; /* offset inside struct acpi_sbs_battery */ @@ -904,6 +968,7 @@ static void __exit battery_hook_exit(void) static int sysfs_add_battery(struct acpi_battery *battery) { + bool extended_info_available = test_bit(ACPI_BATTERY_XINFO_PRESENT, &battery->flags); struct power_supply_config psy_cfg = { .drv_data = battery, .attr_grp = acpi_battery_groups, @@ -922,25 +987,55 @@ static int sysfs_add_battery(struct acpi_battery *battery) if (power_unit == ACPI_BATTERY_POWER_UNIT_MA) { if (full_cap_broken) { - battery->bat_desc.properties = - charge_battery_full_cap_broken_props; - battery->bat_desc.num_properties = - ARRAY_SIZE(charge_battery_full_cap_broken_props); + if (extended_info_available) { + battery->bat_desc.properties = + charge_battery_full_cap_broken_extended_props; + battery->bat_desc.num_properties = + ARRAY_SIZE(charge_battery_full_cap_broken_extended_props); + } else { + battery->bat_desc.properties = + charge_battery_full_cap_broken_props; + battery->bat_desc.num_properties = + ARRAY_SIZE(charge_battery_full_cap_broken_props); + } } else { - battery->bat_desc.properties = charge_battery_props; - battery->bat_desc.num_properties = - ARRAY_SIZE(charge_battery_props); + if (extended_info_available) { + battery->bat_desc.properties = + charge_battery_extended_props; + battery->bat_desc.num_properties = + ARRAY_SIZE(charge_battery_extended_props); + } else { + battery->bat_desc.properties = + charge_battery_props; + battery->bat_desc.num_properties = + ARRAY_SIZE(charge_battery_props); + } } } else { if (full_cap_broken) { - battery->bat_desc.properties = - energy_battery_full_cap_broken_props; - battery->bat_desc.num_properties = - ARRAY_SIZE(energy_battery_full_cap_broken_props); + if (extended_info_available) { + battery->bat_desc.properties = + energy_battery_full_cap_broken_extended_props; + battery->bat_desc.num_properties = + ARRAY_SIZE(energy_battery_full_cap_broken_extended_props); + } else { + battery->bat_desc.properties = + energy_battery_full_cap_broken_props; + battery->bat_desc.num_properties = + ARRAY_SIZE(energy_battery_full_cap_broken_props); + } } else { - battery->bat_desc.properties = energy_battery_props; - battery->bat_desc.num_properties = - ARRAY_SIZE(energy_battery_props); + if (extended_info_available) { + battery->bat_desc.properties = + energy_battery_extended_props; + battery->bat_desc.num_properties = + ARRAY_SIZE(energy_battery_extended_props); + } else { + battery->bat_desc.properties = + energy_battery_props; + battery->bat_desc.num_properties = + ARRAY_SIZE(energy_battery_props); + } } } diff --git a/drivers/acpi/bus.c b/drivers/acpi/bus.c index 808c6746be14..48f0b4809f94 100644 --- a/drivers/acpi/bus.c +++ b/drivers/acpi/bus.c @@ -1120,11 +1120,6 @@ EXPORT_SYMBOL_GPL(acpi_driver_match_device); ACPI Bus operations -------------------------------------------------------------------------- */ -static int acpi_bus_match(struct device *dev, const struct device_driver *drv) -{ - return 0; -} - static int acpi_device_uevent(const struct device *dev, struct kobj_uevent_env *env) { return __acpi_device_uevent_modalias(to_acpi_device(dev), env); @@ -1132,7 +1127,6 @@ static int acpi_device_uevent(const struct device *dev, struct kobj_uevent_env * const struct bus_type acpi_bus_type = { .name = "acpi", - .match = acpi_bus_match, .uevent = acpi_device_uevent, }; @@ -1451,7 +1445,7 @@ static int __init acpi_bus_init(void) */ acpi_root_dir = proc_mkdir(ACPI_BUS_FILE_ROOT, NULL); - result = bus_register(&acpi_bus_type); + result = companion_bus_register(&acpi_bus_type); if (!result) return 0; diff --git a/drivers/acpi/button.c b/drivers/acpi/button.c index cdbb1023a8ee..34c7173b5a6c 100644 --- a/drivers/acpi/button.c +++ b/drivers/acpi/button.c @@ -146,6 +146,18 @@ static const struct dmi_system_id dmi_lid_quirks[] = { }, { /* + * Razer Blade Stealth 13 early 2020, when the lid is opened + * while suspended the open notification is lost and _LID keeps + * returning closed after resume, causing spurious re-suspends. + */ + .matches = { + DMI_MATCH(DMI_SYS_VENDOR, "Razer"), + DMI_MATCH(DMI_PRODUCT_NAME, "Blade Stealth 13 (Early 2020) - RZ09-0310"), + }, + .driver_data = (void *)(long)ACPI_BUTTON_LID_INIT_OPEN, + }, + { + /* * Samsung galaxybook2 ,initial _LID device notification returns * lid closed. */ diff --git a/drivers/acpi/cppc_acpi.c b/drivers/acpi/cppc_acpi.c index fef54fcd00b7..80e2e6b32ce3 100644 --- a/drivers/acpi/cppc_acpi.c +++ b/drivers/acpi/cppc_acpi.c @@ -34,8 +34,12 @@ #define pr_fmt(fmt) "ACPI CPPC: " fmt #include <linux/delay.h> +#include <linux/interval_tree_generic.h> #include <linux/iopoll.h> #include <linux/ktime.h> +#include <linux/list.h> +#include <linux/mutex.h> +#include <linux/rbtree.h> #include <linux/rwsem.h> #include <linux/wait.h> #include <linux/topology.h> @@ -70,6 +74,8 @@ struct cppc_pcc_data { * Take write_lock for all purposes which gives exclusive access */ struct rw_semaphore pcc_lock; + /* Serialize byte-oriented accesses to aliased PCC payload fields. */ + raw_spinlock_t payload_lock; /* Wait queue for CPUs whose requests were batched */ wait_queue_head_t pcc_write_wait_q; @@ -81,6 +87,7 @@ struct cppc_pcc_data { /* Array to represent the PCC channel per subspace ID */ static struct cppc_pcc_data *pcc_data[MAX_PCC_SUBSPACES]; +static DEFINE_MUTEX(pcc_data_lock); /* The cpu_pcc_subspace_idx contains per CPU subspace ID */ static DEFINE_PER_CPU(int, cpu_pcc_subspace_idx); @@ -93,9 +100,77 @@ static DEFINE_PER_CPU(int, cpu_pcc_subspace_idx); */ static DEFINE_PER_CPU(struct cpc_desc *, cpc_desc_ptr); +/* Protect immutable capability queries against descriptor removal. */ +static DEFINE_MUTEX(cpc_desc_lock); + +static void cpc_set_desc(unsigned int cpu, struct cpc_desc *desc) +{ + guard(mutex)(&cpc_desc_lock); + per_cpu(cpc_desc_ptr, cpu) = desc; +} + +struct cpc_sysmem_node { + struct rb_node rb; + u64 subtree_last; + u64 start; + u64 last; + struct cpc_desc *desc; + unsigned int reg_idx; + struct list_head aliases; + struct list_head alias_node; + struct cpc_sysmem_node *alias_of; + bool registered; +}; + +struct cpc_non_mmio_node { + struct rb_node rb; + u64 subtree_last; + u64 start; + u64 last; + struct cpc_desc *desc; + unsigned int reg_idx; + u8 space_id; + u8 pcc_ss_id; + bool registered; +}; + +#define CPC_SYSMEM_START(node) ((node)->start) +#define CPC_SYSMEM_LAST(node) ((node)->last) + +INTERVAL_TREE_DEFINE(struct cpc_sysmem_node, rb, u64, subtree_last, + CPC_SYSMEM_START, CPC_SYSMEM_LAST, static inline, + cpc_sysmem_itree) + +static struct rb_root_cached cpc_sysmem_tree = RB_ROOT_CACHED; +static DEFINE_MUTEX(cpc_sysmem_lock); + +#define CPC_NON_MMIO_START(node) ((node)->start) +#define CPC_NON_MMIO_LAST(node) ((node)->last) + +INTERVAL_TREE_DEFINE(struct cpc_non_mmio_node, rb, u64, subtree_last, + CPC_NON_MMIO_START, CPC_NON_MMIO_LAST, static inline, + cpc_non_mmio_itree) + +static struct rb_root_cached cpc_pcc_trees[MAX_PCC_SUBSPACES]; +static struct rb_root_cached cpc_sysio_tree = RB_ROOT_CACHED; +static DEFINE_MUTEX(cpc_non_mmio_lock); + +static struct cpc_sysmem_node *cpc_sysmem_first(u64 start, u64 last) +{ + return cpc_sysmem_itree_iter_first(&cpc_sysmem_tree, start, last); +} + +static struct cpc_sysmem_node *cpc_sysmem_next(struct cpc_sysmem_node *node, + u64 start, u64 last) +{ + return cpc_sysmem_itree_iter_next(node, start, last); +} + +#define CPC_PCC_HEADER_SIZE 0x8 + /* pcc mapped address + header size + offset within PCC subspace */ #define GET_PCC_VADDR(offs, pcc_ss_id) (pcc_data[pcc_ss_id]->pcc_channel->shmem + \ - 0x8 + (offs)) + CPC_PCC_HEADER_SIZE + (offs)) /* Check if a CPC register is in PCC */ #define CPC_IN_PCC(cpc) ((cpc)->type == ACPI_TYPE_BUFFER && \ @@ -129,6 +204,28 @@ static DEFINE_PER_CPU(struct cpc_desc *, cpc_desc_ptr); !!(cpc)->cpc_entry.int_value : \ !IS_NULL_REG(&(cpc)->cpc_entry.reg)) +static bool cpc_is_writable(const struct cpc_register_resource *cpc) +{ + return cpc->type == ACPI_TYPE_BUFFER && + !IS_NULL_REG(&cpc->cpc_entry.reg) && + !cpc->cpc_entry.write_unsupported; +} + +static bool cpc_is_readable(const struct cpc_register_resource *cpc) +{ + return cpc->type != ACPI_TYPE_BUFFER || + !cpc->cpc_entry.read_unsupported; +} + +static bool cpc_entry_present(const struct cpc_register_resource *cpc) +{ + if (cpc->type == ACPI_TYPE_INTEGER) + return true; + + return cpc->type == ACPI_TYPE_BUFFER && + !IS_NULL_REG(&cpc->cpc_entry.reg); +} + /* * Each bit indicates the optionality of the register in per-cpu * cpc_regs[] with the corresponding index. 0 means mandatory and 1 @@ -142,6 +239,36 @@ static DEFINE_PER_CPU(struct cpc_desc *, cpc_desc_ptr); */ #define IS_OPTIONAL_CPC_REG(reg_idx) (REG_OPTIONAL & (1U << (reg_idx))) +static bool cpc_integer_entry_valid(unsigned int reg_idx, u64 value, + bool *legacy_null) +{ + *legacy_null = false; + + switch (reg_idx) { + case HIGHEST_PERF: + case NOMINAL_PERF: + case LOW_NON_LINEAR_PERF: + case LOWEST_PERF: + case REFERENCE_PERF: + case LOWEST_FREQ: + case NOMINAL_FREQ: + return value <= U32_MAX; + case CTR_WRAP_TIME: + /* AML Integers and the kernel interface are both 64-bit. */ + return true; + case AUTO_SEL_ENABLE: + return value <= 1; + case DESIRED_PERF: + /* Validated against Autonomous Selection after parsing. */ + *legacy_null = value == 0; + return *legacy_null; + default: + /* Tolerate legacy Integer 0 placeholders for absent options. */ + *legacy_null = value == 0 && IS_OPTIONAL_CPC_REG(reg_idx); + return *legacy_null; + } +} + /* * Arbitrary Retries in case the remote processor is slow to respond * to PCC commands. Keeping it high enough to cover emulators where @@ -149,7 +276,8 @@ static DEFINE_PER_CPU(struct cpc_desc *, cpc_desc_ptr); */ #define NUM_RETRIES 500ULL -#define OVER_16BTS_MASK ~0xFFFFULL +#define CPC_GENERIC_REGISTER_DESCRIPTOR 0x82 +#define CPC_GENERIC_REGISTER_LENGTH (sizeof(struct cpc_reg) - 3) #define define_one_cppc_ro(_name) \ static struct kobj_attribute _name = \ @@ -200,15 +328,82 @@ show_cppc_data(cppc_get_perf_ctrs, cppc_perf_fb_ctrs, wraparound_time); ((((val) & GENMASK(((reg)->bit_width) - 1, 0)) << (reg)->bit_offset) | \ ((prev_val) & ~(GENMASK(((reg)->bit_width) - 1, 0) << (reg)->bit_offset))) \ -static u64 cpc_sysmem_access_size(const struct cpc_register_resource *reg) +static unsigned int cpc_reg_access_width(const struct cpc_reg *reg) { - const struct cpc_reg *gas = ®->cpc_entry.reg; - unsigned int width; - - if (gas->access_width > 4) + if (reg->access_width > 4) return 0; - width = GET_BIT_WIDTH(gas); + if (reg->access_width) + return 8U << (reg->access_width - 1); + + return reg->bit_width; +} + +enum cpc_platform_quirk { + CPC_QUIRK_PERF_LIMITED_OWNS_UNIT = BIT(0), +}; + +static const struct acpi_platform_list cpc_platform_quirk_list[] = { + { + .oem_id = "NVIDIA", + .oem_table_id = "T41", + .table = ACPI_SIG_DSDT, + .pred = all_versions, + .reason = "Performance Limited owns its access unit", + .data = CPC_QUIRK_PERF_LIMITED_OWNS_UNIT, + }, + { } +}; + +static DEFINE_MUTEX(cpc_platform_quirk_lock); +static bool cpc_platform_quirks_initialized; +static u32 cpc_platform_quirks; + +static int cpc_get_platform_quirks(u32 *quirks) +{ + int idx, ret = 0; + + mutex_lock(&cpc_platform_quirk_lock); + if (!cpc_platform_quirks_initialized) { + idx = acpi_match_platform_list(cpc_platform_quirk_list); + if (idx < 0 && idx != -ENODEV) { + ret = idx; + goto out; + } + if (idx >= 0) + cpc_platform_quirks = cpc_platform_quirk_list[idx].data; + cpc_platform_quirks_initialized = true; + } + *quirks = cpc_platform_quirks; +out: + mutex_unlock(&cpc_platform_quirk_lock); + + return ret; +} + +static void cpc_apply_platform_quirks(struct cpc_reg *reg, + unsigned int reg_idx, u32 quirks) +{ + unsigned int access_width; + + if (!(quirks & CPC_QUIRK_PERF_LIMITED_OWNS_UNIT) || + reg_idx != PERF_LIMITED || + reg->space_id != ACPI_ADR_SPACE_SYSTEM_MEMORY || + reg->bit_width != 2 || reg->bit_offset) + return; + + access_width = cpc_reg_access_width(reg); + if (access_width != 32) + return; + + reg->bit_width = access_width; + pr_info_once("firmware quirk: Performance Limited owns its access unit, using Bit Width %u\n", + access_width); +} + +static u64 cpc_sysmem_access_size(const struct cpc_register_resource *reg) +{ + unsigned int width = cpc_reg_access_width(®->cpc_entry.reg); if (width != 8 && width != 16 && width != 32 && width != 64) return 0; @@ -216,13 +411,35 @@ static u64 cpc_sysmem_access_size(const struct cpc_register_resource *reg) return width / 8; } +static u64 cpc_sysmem_field_size(const struct cpc_reg *gas) +{ + return DIV_ROUND_UP((u64)gas->bit_offset + gas->bit_width, 8); +} + +static u64 cpc_sysmem_claim_size(const struct cpc_register_resource *reg) +{ + const struct cpc_reg *gas = ®->cpc_entry.reg; + u64 access_size = cpc_sysmem_access_size(reg); + + if (!gas->bit_width) + return access_size; + + return max(access_size, cpc_sysmem_field_size(gas)); +} + +static bool cpc_reg_access_aligned(const struct cpc_reg *reg, u64 access_size) +{ + /* x86 MMIO and port-I/O accessors support unaligned addresses. */ + return IS_ENABLED(CONFIG_X86) || IS_ALIGNED(reg->address, access_size); +} + static bool cpc_sysmem_access_units_overlap(const struct cpc_register_resource *a, const struct cpc_register_resource *b) { const struct cpc_reg *a_gas = &a->cpc_entry.reg; const struct cpc_reg *b_gas = &b->cpc_entry.reg; - u64 a_size = cpc_sysmem_access_size(a); - u64 b_size = cpc_sysmem_access_size(b); + u64 a_size = cpc_sysmem_claim_size(a); + u64 b_size = cpc_sysmem_claim_size(b); /* Keep the conservative locking path for malformed access widths. */ if (!a_size || !b_size) @@ -234,36 +451,883 @@ static bool cpc_sysmem_access_units_overlap(const struct cpc_register_resource * return a_gas->address - b_gas->address < b_size; } -static void cpc_mark_rmw_lock_users(struct cpc_desc *cpc_desc) +static bool cpc_reg_is_writable(unsigned int reg_idx) { - int i, j; + /* Only controls written by this driver can be competing writers. */ + switch (reg_idx) { + case DESIRED_PERF: + case MIN_PERF: + case MAX_PERF: + case PERF_LIMITED: + case ENABLE: + case AUTO_SEL_ENABLE: + case AUTO_ACT_WINDOW: + case ENERGY_PERF: + return true; + default: + return false; + } +} + +static bool cpc_reg_is_write_only(const struct cpc_desc *cpc_desc, + unsigned int reg_idx) +{ + return cpc_desc->version >= CPPC_V4_REV && + (reg_idx == DESIRED_PERF || reg_idx == OSPM_NOMINAL_PERF); +} +static void cpc_disable_reg(struct cpc_desc *cpc_desc, unsigned int reg_idx) +{ + struct cpc_register_resource *reg = &cpc_desc->cpc_regs[reg_idx]; + + reg->type = ACPI_TYPE_INTEGER; + reg->cpc_entry.int_value = 0; +} + +static bool cpc_optional_writer_can_be_disabled(unsigned int reg_idx) +{ + if (!IS_OPTIONAL_CPC_REG(reg_idx) || !cpc_reg_is_writable(reg_idx) || + reg_idx == MIN_PERF || reg_idx == MAX_PERF || reg_idx == ENABLE || + reg_idx == AUTO_SEL_ENABLE) + return false; + + return true; +} + +static bool cpc_sysmem_reg_needs_rmw(const struct cpc_register_resource *reg) +{ + const struct cpc_reg *gas = ®->cpc_entry.reg; + u64 access_size = cpc_sysmem_access_size(reg); + + return gas->bit_offset || gas->bit_width != access_size * 8; +} + +static int cpc_validate_sysmem_reg(struct cpc_desc *cpc_desc, + const struct cpc_reg *gas, + unsigned int reg_idx) +{ + unsigned int access_width = cpc_reg_access_width(gas); + u64 access_size; + + if (access_width != 8 && access_width != 16 && + access_width != 32 && access_width != 64) + goto invalid; + + if (!gas->bit_width || gas->bit_width > access_width || + gas->bit_offset >= access_width || + gas->bit_width > access_width - gas->bit_offset) + goto invalid; + + access_size = access_width / 8; + if (!gas->address || gas->address > U64_MAX - (access_size - 1)) + goto invalid; + if (!cpc_reg_access_aligned(gas, access_size)) + goto invalid; + + if (reg_idx == PERF_LIMITED) { + if (access_width == 64 && !IS_ENABLED(CONFIG_64BIT)) { + pr_warn_once("CPU%d: Performance Limited register cannot be accessed atomically; keeping its range reserved\n", + cpc_desc->cpu_id); + cpc_desc->cpc_regs[reg_idx].cpc_entry.read_unsupported = true; + cpc_desc->cpc_regs[reg_idx].cpc_entry.write_unsupported = true; + return 0; + } + + if (gas->bit_offset || gas->bit_width != access_width) { + pr_warn_once("CPU%d: Performance Limited register cannot be cleared safely; keeping it readable\n", + cpc_desc->cpu_id); + cpc_desc->cpc_regs[reg_idx].cpc_entry.write_unsupported = true; + } + } + + return 0; + +invalid: + access_size = 0; + if (access_width == 8 || access_width == 16 || + access_width == 32 || access_width == 64) + access_size = access_width / 8; + if (gas->bit_width) + access_size = max(access_size, cpc_sysmem_field_size(gas)); + if ((cpc_reg_is_write_only(cpc_desc, reg_idx) || + reg_idx == PERF_LIMITED) && gas->address && access_size && + gas->address <= U64_MAX - (access_size - 1)) { + struct cpc_register_resource *reg = &cpc_desc->cpc_regs[reg_idx]; + + if (reg_idx == PERF_LIMITED) + pr_warn_once("CPU%d: _CPC v%d register %u is inaccessible; keeping its range reserved\n", + cpc_desc->cpu_id, cpc_desc->version, reg_idx); + else + pr_warn("CPU%d: _CPC v%d register %u is inaccessible; keeping its range reserved\n", + cpc_desc->cpu_id, cpc_desc->version, reg_idx); + reg->cpc_entry.read_unsupported = true; + reg->cpc_entry.write_unsupported = true; + return 0; + } + + pr_debug("CPU:%d invalid SystemMemory GAS for _CPC register %u\n", + cpc_desc->cpu_id, reg_idx); + return -EINVAL; +} + +static bool cpc_immutable_autonomous(const struct cpc_desc *cpc_desc) +{ + const struct cpc_register_resource *reg; + + reg = &cpc_desc->cpc_regs[AUTO_SEL_ENABLE]; + return osc_sb_cppc2_support_acked && reg->type == ACPI_TYPE_INTEGER && + reg->cpc_entry.int_value == 1; +} + +static bool cpc_retain_pcc_status(struct cpc_desc *cpc_desc, + unsigned int reg_idx); + +static int cpc_resolve_unsupported(struct cpc_desc *cpc_desc, + u32 unsupported) +{ + unsigned int i; + u32 bounds = BIT(MIN_PERF) | BIT(MAX_PERF); + bool min_unusable, max_unusable; + + if (unsupported & bounds) { + min_unusable = (unsupported & BIT(MIN_PERF)) || + !cpc_is_writable(&cpc_desc->cpc_regs[MIN_PERF]); + max_unusable = (unsupported & BIT(MAX_PERF)) || + !cpc_is_writable(&cpc_desc->cpc_regs[MAX_PERF]); + if (min_unusable && max_unusable) { + pr_warn("CPU%d: ignoring inaccessible Minimum and Maximum Performance registers\n", + cpc_desc->cpu_id); + cpc_disable_reg(cpc_desc, MIN_PERF); + cpc_disable_reg(cpc_desc, MAX_PERF); + unsupported &= ~bounds; + } + } + + for (i = 0; i < cpc_desc->num_entries - 2; i++) { + if (!(unsupported & BIT(i))) + continue; + + /* CPPC control does not depend on Performance Limited status. */ + if (i == PERF_LIMITED) { + if (CPC_IN_PCC(&cpc_desc->cpc_regs[i]) && + cpc_retain_pcc_status(cpc_desc, i)) + continue; + + pr_warn_once("CPU%d: ignoring inaccessible Performance Limited register\n", + cpc_desc->cpu_id); + cpc_disable_reg(cpc_desc, i); + continue; + } + + if (i == DESIRED_PERF && cpc_immutable_autonomous(cpc_desc)) { + pr_warn("CPU%d: ignoring inaccessible Desired Performance register in autonomous mode\n", + cpc_desc->cpu_id); + cpc_disable_reg(cpc_desc, i); + continue; + } + + /* + * A present Enable or Autonomous Selection control must remain + * usable. Disabling the latter could leave autonomous selection + * enabled while OSPM believes that it has disabled it. + */ + if (i == ENABLE || + (i == AUTO_SEL_ENABLE && cpc_entry_present(&cpc_desc->cpc_regs[i])) || + i == MIN_PERF || i == MAX_PERF || + !IS_OPTIONAL_CPC_REG(i)) { + pr_err("CPU%d: cannot access _CPC register %u\n", + cpc_desc->cpu_id, i); + return -EINVAL; + } + + pr_warn("CPU%d: ignoring inaccessible optional _CPC register %u\n", + cpc_desc->cpu_id, i); + cpc_disable_reg(cpc_desc, i); + } + + return 0; +} + +static int cpc_validate_required_controls(struct cpc_desc *cpc_desc) +{ + unsigned int i; + + /* + * Performance Limited is required by the specification, but tolerate a + * NULL descriptor used by firmware which cannot report limiting events. + * CPPC control does not depend on this status. + */ for (i = 0; i < cpc_desc->num_entries - 2; i++) { - struct cpc_register_resource *a = &cpc_desc->cpc_regs[i]; + if (i != DESIRED_PERF && i != PERF_LIMITED && + !IS_OPTIONAL_CPC_REG(i) && + !cpc_entry_present(&cpc_desc->cpc_regs[i])) { + pr_debug("CPU:%d lacks mandatory _CPC register %u\n", + cpc_desc->cpu_id, i); + return -EINVAL; + } + } + + /* Desired may be absent only for immutable autonomous operation. */ + if (!cpc_is_writable(&cpc_desc->cpc_regs[DESIRED_PERF]) && + !cpc_immutable_autonomous(cpc_desc)) { + pr_debug("CPU:%d lacks a writable Desired Performance register\n", + cpc_desc->cpu_id); + return -EINVAL; + } + + return 0; +} + +static int cpc_validate_bound_controls(struct cpc_desc *cpc_desc) +{ + bool have_min, have_max; + + have_min = cpc_is_writable(&cpc_desc->cpc_regs[MIN_PERF]); + have_max = cpc_is_writable(&cpc_desc->cpc_regs[MAX_PERF]); + if (have_min != have_max) { + pr_err("CPU%d: _CPC must provide both Minimum and Maximum Performance or neither\n", + cpc_desc->cpu_id); + return -EINVAL; + } + + return 0; +} + +static bool +cpc_retain_pcc_status(struct cpc_desc *cpc_desc, unsigned int reg_idx) +{ + struct cpc_register_resource *reg = &cpc_desc->cpc_regs[reg_idx]; + const struct cpc_reg *gas = ®->cpc_entry.reg; + u64 size; + + if (reg_idx != PERF_LIMITED || !gas->bit_width) + return false; + + size = DIV_ROUND_UP((u64)gas->bit_offset + gas->bit_width, 8); + if (!size || gas->address > U64_MAX - (size - 1)) + return false; + + reg->cpc_entry.read_unsupported = true; + reg->cpc_entry.write_unsupported = true; + pr_warn_once("CPU%d: Performance Limited register cannot be accessed; keeping its PCC range reserved\n", + cpc_desc->cpu_id); + return true; +} + +static bool +cpc_retain_sysio_status(struct cpc_desc *cpc_desc, unsigned int reg_idx) +{ + struct cpc_register_resource *reg = &cpc_desc->cpc_regs[reg_idx]; + const struct cpc_reg *gas = ®->cpc_entry.reg; + + if (reg_idx != PERF_LIMITED || !gas->bit_width) + return false; + + /* Retain any in-range portion for overlap validation only. */ + if (gas->address > U16_MAX) + return false; + + pr_warn_once("CPU%d: Performance Limited register cannot be accessed; keeping its SystemIO range reserved\n", + cpc_desc->cpu_id); + reg->cpc_entry.read_unsupported = true; + reg->cpc_entry.write_unsupported = true; + return true; +} + +static u64 cpc_non_mmio_access_size(const struct cpc_register_resource *reg) +{ + const struct cpc_reg *gas = ®->cpc_entry.reg; + + if (gas->space_id == ACPI_ADR_SPACE_PLATFORM_COMM) + return DIV_ROUND_UP((u64)gas->bit_offset + gas->bit_width, 8); + + return max((u64)cpc_reg_access_width(gas) / 8, + DIV_ROUND_UP((u64)gas->bit_offset + gas->bit_width, 8)); +} + +static void cpc_validate_pcc_bounds(struct cpc_desc *cpc_desc, + int pcc_ss_id, struct cppc_pcc_data *data, + u32 *unsupported) +{ + u64 shmem_size = data->pcc_channel->shmem_size; + unsigned int i; + + for (i = 0; i < cpc_desc->num_entries - 2; i++) { + struct cpc_register_resource *reg = &cpc_desc->cpc_regs[i]; struct cpc_reg *gas; u64 access_size; - if (!CPC_SUPPORTED(a) || !CPC_IN_SYSTEM_MEMORY(a)) + if ((*unsupported & BIT(i)) || !CPC_SUPPORTED(reg) || + !CPC_IN_PCC(reg)) continue; - gas = &a->cpc_entry.reg; - access_size = cpc_sysmem_access_size(a); - if (gas->bit_offset || !access_size || - gas->bit_width != access_size * 8) - a->cpc_entry.use_rmw_lock = true; + gas = ®->cpc_entry.reg; + if (gas->access_width != pcc_ss_id) + continue; + access_size = cpc_non_mmio_access_size(reg); + if (shmem_size >= CPC_PCC_HEADER_SIZE && + gas->address <= shmem_size - CPC_PCC_HEADER_SIZE && + access_size <= shmem_size - CPC_PCC_HEADER_SIZE - gas->address) + continue; - for (j = i + 1; j < cpc_desc->num_entries - 2; j++) { - struct cpc_register_resource *b = &cpc_desc->cpc_regs[j]; + pr_debug("CPU%d: _CPC register %u exceeds the PCC shared region\n", + cpc_desc->cpu_id, i); + *unsupported |= BIT(i); + } +} - if (!CPC_SUPPORTED(b) || !CPC_IN_SYSTEM_MEMORY(b)) - continue; - if (!cpc_sysmem_access_units_overlap(a, b)) - continue; +static bool cpc_pcc_access_needed(const struct cpc_desc *cpc_desc) +{ + unsigned int i; + + for (i = 0; i < cpc_desc->num_entries - 2; i++) { + const struct cpc_register_resource *reg = &cpc_desc->cpc_regs[i]; + + if (CPC_SUPPORTED(reg) && CPC_IN_PCC(reg) && + (cpc_is_readable(reg) || cpc_is_writable(reg))) + return true; + } + + return false; +} + +static bool cpc_non_mmio_overlap_conflicts(u8 space_id, bool a_writable, + bool b_writable, bool a_write_only, + bool b_write_only) +{ + /* Only a write-only control can use a separate read-side port alias. */ + if (space_id == ACPI_ADR_SPACE_SYSTEM_IO && a_writable != b_writable) + return a_writable ? !a_write_only : !b_write_only; + + return a_writable || b_writable; +} + +static bool +cpc_sysio_perf_limited_conflicts(unsigned int a_idx, bool a_writable, + unsigned int b_idx, bool b_writable) +{ + return (a_idx == PERF_LIMITED && b_writable) || + (b_idx == PERF_LIMITED && a_writable); +} + +static struct rb_root_cached *cpc_non_mmio_tree(u8 space_id, u8 pcc_ss_id) +{ + if (space_id == ACPI_ADR_SPACE_PLATFORM_COMM) + return &cpc_pcc_trees[pcc_ss_id]; + if (space_id == ACPI_ADR_SPACE_SYSTEM_IO) + return &cpc_sysio_tree; + return NULL; +} + +static bool cpc_same_non_mmio_register(const struct cpc_non_mmio_node *a, + const struct cpc_non_mmio_node *b) +{ + const struct cpc_reg *a_gas = + &a->desc->cpc_regs[a->reg_idx].cpc_entry.reg; + const struct cpc_reg *b_gas = + &b->desc->cpc_regs[b->reg_idx].cpc_entry.reg; + + return a->space_id == b->space_id && a->pcc_ss_id == b->pcc_ss_id && + a->reg_idx == b->reg_idx && a->start == b->start && + a->last == b->last && a_gas->bit_offset == b_gas->bit_offset && + a_gas->bit_width == b_gas->bit_width && + (a->space_id == ACPI_ADR_SPACE_PLATFORM_COMM || + cpc_reg_access_width(a_gas) == cpc_reg_access_width(b_gas)); +} + +static int cpc_validate_non_mmio_pair(const struct cpc_non_mmio_node *a, + const struct cpc_non_mmio_node *b) +{ + const struct cpc_register_resource *a_reg; + const struct cpc_register_resource *b_reg; + bool a_writable, b_writable; + const char *name; + + a_reg = &a->desc->cpc_regs[a->reg_idx]; + b_reg = &b->desc->cpc_regs[b->reg_idx]; + a_writable = cpc_reg_is_writable(a->reg_idx) && cpc_is_writable(a_reg); + b_writable = cpc_reg_is_writable(b->reg_idx) && cpc_is_writable(b_reg); + + if (!cpc_non_mmio_overlap_conflicts(a->space_id, a_writable, + b_writable, + cpc_reg_is_write_only(a->desc, a->reg_idx), + cpc_reg_is_write_only(b->desc, b->reg_idx)) && + !(a->space_id == ACPI_ADR_SPACE_SYSTEM_IO && + cpc_sysio_perf_limited_conflicts(a->reg_idx, a_writable, + b->reg_idx, b_writable))) + return 0; + + if (cpc_same_non_mmio_register(a, b)) + return 0; + + name = a->space_id == ACPI_ADR_SPACE_PLATFORM_COMM ? + "PCC" : "SystemIO"; + pr_err("CPU%d: %s _CPC register %u conflicts with CPU%d register %u\n", + a->desc->cpu_id, name, a->reg_idx, b->desc->cpu_id, + b->reg_idx); + return -EINVAL; +} + +static void cpc_unregister_non_mmio_desc_locked(struct cpc_desc *cpc_desc) +{ + unsigned int i; + + if (!cpc_desc->non_mmio_nodes) + return; + + for (i = 0; i < cpc_desc->num_entries - 2; i++) { + struct cpc_non_mmio_node *node = &cpc_desc->non_mmio_nodes[i]; + struct rb_root_cached *tree; + + if (!node->registered) + continue; + + tree = cpc_non_mmio_tree(node->space_id, node->pcc_ss_id); + cpc_non_mmio_itree_remove(node, tree); + } + + kfree(cpc_desc->non_mmio_nodes); + cpc_desc->non_mmio_nodes = NULL; +} + +static int cpc_register_non_mmio_desc(struct cpc_desc *cpc_desc) +{ + unsigned int nr_regs = cpc_desc->num_entries - 2; + unsigned int i; + int ret = 0; + bool found = false; + + for (i = 0; i < nr_regs; i++) { + struct cpc_register_resource *reg = &cpc_desc->cpc_regs[i]; + u8 space_id; + + if (!CPC_SUPPORTED(reg) || reg->type != ACPI_TYPE_BUFFER) + continue; + space_id = reg->cpc_entry.reg.space_id; + if (space_id == ACPI_ADR_SPACE_PLATFORM_COMM || + space_id == ACPI_ADR_SPACE_SYSTEM_IO) { + found = true; + break; + } + } + if (!found) + return 0; + + cpc_desc->non_mmio_nodes = kcalloc(nr_regs, + sizeof(*cpc_desc->non_mmio_nodes), + GFP_KERNEL); + if (!cpc_desc->non_mmio_nodes) + return -ENOMEM; + + mutex_lock(&cpc_non_mmio_lock); + + for (i = 0; i < nr_regs; i++) { + struct cpc_register_resource *reg = &cpc_desc->cpc_regs[i]; + struct cpc_non_mmio_node *match, *node; + struct rb_root_cached *tree; + u8 space_id; + u64 size; + + if (!CPC_SUPPORTED(reg) || reg->type != ACPI_TYPE_BUFFER) + continue; + + space_id = reg->cpc_entry.reg.space_id; + if (space_id != ACPI_ADR_SPACE_PLATFORM_COMM && + space_id != ACPI_ADR_SPACE_SYSTEM_IO) + continue; + + node = &cpc_desc->non_mmio_nodes[i]; + size = cpc_non_mmio_access_size(reg); + node->start = reg->cpc_entry.reg.address; + node->last = node->start + size - 1; + node->desc = cpc_desc; + node->reg_idx = i; + node->space_id = space_id; + node->pcc_ss_id = space_id == ACPI_ADR_SPACE_PLATFORM_COMM ? + reg->cpc_entry.reg.access_width : 0; + tree = cpc_non_mmio_tree(space_id, node->pcc_ss_id); + + match = cpc_non_mmio_itree_iter_first(tree, node->start, + node->last); + while (match) { + ret = cpc_validate_non_mmio_pair(node, match); + if (ret) + goto out_unregister; + + match = cpc_non_mmio_itree_iter_next(match, node->start, + node->last); + } + cpc_non_mmio_itree_insert(node, tree); + node->registered = true; + } + + mutex_unlock(&cpc_non_mmio_lock); + return 0; + +out_unregister: + cpc_unregister_non_mmio_desc_locked(cpc_desc); + mutex_unlock(&cpc_non_mmio_lock); + return ret; +} + +static void cpc_unregister_non_mmio_desc(struct cpc_desc *cpc_desc) +{ + if (!cpc_desc->non_mmio_nodes) + return; + + mutex_lock(&cpc_non_mmio_lock); + cpc_unregister_non_mmio_desc_locked(cpc_desc); + mutex_unlock(&cpc_non_mmio_lock); +} + +static void cpc_mark_rmw_lock_users(struct cpc_desc *cpc_desc) +{ + int i; + + for (i = 0; i < cpc_desc->num_entries - 2; i++) { + struct cpc_register_resource *reg = &cpc_desc->cpc_regs[i]; + + if (CPC_SUPPORTED(reg) && CPC_IN_SYSTEM_MEMORY(reg) && + cpc_is_writable(reg)) + reg->cpc_entry.use_rmw_lock = + cpc_sysmem_reg_needs_rmw(reg); + } +} + +struct cpc_bit_position { + u64 byte; + u8 bit; +}; + +static bool cpc_bit_position_before(const struct cpc_bit_position *a, + const struct cpc_bit_position *b) +{ + return a->byte < b->byte || (a->byte == b->byte && a->bit < b->bit); +} + +static bool cpc_sysmem_fields_overlap(const struct cpc_register_resource *a, + const struct cpc_register_resource *b) +{ + const struct cpc_reg *a_gas = &a->cpc_entry.reg; + const struct cpc_reg *b_gas = &b->cpc_entry.reg; + unsigned int a_last_bit = a_gas->bit_offset + a_gas->bit_width - 1; + unsigned int b_last_bit = b_gas->bit_offset + b_gas->bit_width - 1; + struct cpc_bit_position a_start = { + .byte = a_gas->address + a_gas->bit_offset / 8, + .bit = a_gas->bit_offset % 8, + }; + struct cpc_bit_position a_end = { + .byte = a_gas->address + a_last_bit / 8, + .bit = a_last_bit % 8, + }; + struct cpc_bit_position b_start = { + .byte = b_gas->address + b_gas->bit_offset / 8, + .bit = b_gas->bit_offset % 8, + }; + struct cpc_bit_position b_end = { + .byte = b_gas->address + b_last_bit / 8, + .bit = b_last_bit % 8, + }; + + return !cpc_bit_position_before(&a_end, &b_start) && + !cpc_bit_position_before(&b_end, &a_start); +} + +static bool cpc_sysmem_access_overlaps_field(const struct cpc_register_resource *access, + const struct cpc_register_resource *field) +{ + const struct cpc_reg *access_gas = &access->cpc_entry.reg; + const struct cpc_reg *field_gas = &field->cpc_entry.reg; + u64 access_last; + u64 field_start; + u64 field_last; + + if (!field_gas->bit_width) + return cpc_sysmem_access_units_overlap(access, field); + + access_last = access_gas->address + + cpc_sysmem_access_size(access) - 1; + field_start = field_gas->address + field_gas->bit_offset / 8; + field_last = field_gas->address + + (field_gas->bit_offset + field_gas->bit_width - 1) / 8; + + return access_gas->address <= field_last && field_start <= access_last; +} + +static bool cpc_same_sysmem_register(unsigned int a_idx, + const struct cpc_register_resource *a, + unsigned int b_idx, + const struct cpc_register_resource *b) +{ + const struct cpc_reg *a_gas = &a->cpc_entry.reg; + const struct cpc_reg *b_gas = &b->cpc_entry.reg; + + return a_idx == b_idx && + a_gas->address == b_gas->address && + a_gas->bit_width == b_gas->bit_width && + a_gas->bit_offset == b_gas->bit_offset && + cpc_reg_access_width(a_gas) == cpc_reg_access_width(b_gas); +} + +static int cpc_validate_sysmem_pair(const struct cpc_desc *a_desc, + unsigned int a_idx, + const struct cpc_desc *b_desc, + unsigned int b_idx) +{ + const struct cpc_register_resource *a = &a_desc->cpc_regs[a_idx]; + const struct cpc_register_resource *b = &b_desc->cpc_regs[b_idx]; + bool a_write_only, b_write_only; + bool a_writable, b_writable; + bool fields_overlap; + + /* The overlap helper includes each descriptor's conservative claim. */ + if (!CPC_SUPPORTED(a) || !CPC_IN_SYSTEM_MEMORY(a) || + !CPC_SUPPORTED(b) || !CPC_IN_SYSTEM_MEMORY(b) || + !cpc_sysmem_access_units_overlap(a, b)) + return 0; + + a_write_only = cpc_reg_is_write_only(a_desc, a_idx); + b_write_only = cpc_reg_is_write_only(b_desc, b_idx); + fields_overlap = !a->cpc_entry.reg.bit_width || + !b->cpc_entry.reg.bit_width || + cpc_sysmem_fields_overlap(a, b); + /* A readable field must not expose another field's undefined bits. */ + if (a_write_only != b_write_only && + cpc_is_readable(a_write_only ? b : a) && + fields_overlap) + goto conflict; + + a_writable = cpc_reg_is_writable(a_idx) && cpc_is_writable(a); + b_writable = cpc_reg_is_writable(b_idx) && cpc_is_writable(b); + if (!a_writable && !b_writable) + return 0; + + if (cpc_same_sysmem_register(a_idx, a, b_idx, b)) { + u64 access_size = cpc_sysmem_access_size(a); + + /* + * Exact partial aliases update the same field and retain + * last-writer-wins semantics when the complete access is one native + * transaction. A 64-bit MMIO write may be split on 32-bit kernels, + * and an unaligned x86 access is not guaranteed to be one device + * transaction. + */ + if (!a_writable || + (IS_ALIGNED(a->cpc_entry.reg.address, access_size) && + (access_size < sizeof(u64) || + IS_ENABLED(CONFIG_64BIT)))) + return 0; + goto conflict; + } + + /* + * The platform may set Performance Limited asynchronously. A write to + * another field in the same access unit could write back stale status + * bits, which an OSPM lock cannot prevent. + */ + if ((a_idx == PERF_LIMITED && b_writable) || + (b_idx == PERF_LIMITED && a_writable)) + goto conflict; + + /* A full-width writer must not overwrite another logical field. */ + if (fields_overlap && + ((a_writable && b_writable) || + (a_writable && !cpc_sysmem_reg_needs_rmw(a)) || + (b_writable && !cpc_sysmem_reg_needs_rmw(b)))) + goto conflict; + + /* Different descriptors do not share their partial-write locks. */ + if (a_desc != b_desc && a_writable && b_writable) + goto conflict; + + /* + * RMW of either writer preserves the other field. If that field is + * write-only, its readback is undefined and cannot safely be replayed. + */ + if ((a_write_only && b_writable && + cpc_sysmem_reg_needs_rmw(b) && + cpc_sysmem_access_overlaps_field(b, a)) || + (b_write_only && a_writable && + cpc_sysmem_reg_needs_rmw(a) && + cpc_sysmem_access_overlaps_field(a, b))) + goto conflict; - a->cpc_entry.use_rmw_lock = true; - b->cpc_entry.use_rmw_lock = true; + return 0; + +conflict: + pr_err("CPU%d: SystemMemory _CPC register %u conflicts with CPU%d register %u\n", + a_desc->cpu_id, a_idx, b_desc->cpu_id, b_idx); + return -EINVAL; +} + +static bool cpc_disable_new_sysmem_writer(struct cpc_desc *cpc_desc, + unsigned int reg_idx, + const struct cpc_sysmem_node *node) +{ + struct cpc_register_resource *reg = &cpc_desc->cpc_regs[reg_idx]; + struct cpc_sysmem_node *match; + unsigned int status_cpu = 0; + bool found = false; + + if (!cpc_optional_writer_can_be_disabled(reg_idx) || + !cpc_is_writable(reg)) + return false; + + match = cpc_sysmem_first(node->start, node->last); + while (match) { + const struct cpc_register_resource *status; + + if (match->reg_idx == PERF_LIMITED) { + status = &match->desc->cpc_regs[PERF_LIMITED]; + if (!status->cpc_entry.reg.bit_width || + cpc_sysmem_fields_overlap(reg, status)) + return false; + status_cpu = match->desc->cpu_id; + found = true; + } + + match = cpc_sysmem_next(match, node->start, node->last); + } + if (!found) + return false; + + pr_warn_once("CPU%d: ignoring optional _CPC register %u sharing CPU%d Performance Limited access unit\n", + cpc_desc->cpu_id, reg_idx, status_cpu); + cpc_disable_reg(cpc_desc, reg_idx); + return true; +} + +static void cpc_unregister_sysmem_desc_locked(struct cpc_desc *cpc_desc) +{ + unsigned int i; + + if (!cpc_desc->sysmem_nodes) + return; + + for (i = 0; i < cpc_desc->num_entries - 2; i++) { + struct cpc_sysmem_node *node = &cpc_desc->sysmem_nodes[i]; + struct cpc_sysmem_node *alias, *child; + + if (node->alias_of) { + list_del(&node->alias_node); + continue; } + if (!node->registered) + continue; + + cpc_sysmem_itree_remove(node, &cpc_sysmem_tree); + node->registered = false; + if (list_empty(&node->aliases)) + continue; + + /* Keep one representative for aliases owned by live descriptors. */ + alias = list_first_entry(&node->aliases, + struct cpc_sysmem_node, alias_node); + list_del_init(&alias->alias_node); + alias->alias_of = NULL; + alias->registered = true; + list_splice_init(&node->aliases, &alias->aliases); + list_for_each_entry(child, &alias->aliases, alias_node) + child->alias_of = alias; + cpc_sysmem_itree_insert(alias, &cpc_sysmem_tree); } + + kfree(cpc_desc->sysmem_nodes); + cpc_desc->sysmem_nodes = NULL; +} + +static int cpc_register_sysmem_desc(struct cpc_desc *cpc_desc) +{ + unsigned int nr_regs = cpc_desc->num_entries - 2; + unsigned int i; + int ret = 0; + bool found = false; + + for (i = 0; i < nr_regs; i++) { + struct cpc_register_resource *reg = &cpc_desc->cpc_regs[i]; + + if (CPC_SUPPORTED(reg) && CPC_IN_SYSTEM_MEMORY(reg)) { + found = true; + break; + } + } + if (!found) + return 0; + + cpc_desc->sysmem_nodes = kcalloc(nr_regs, + sizeof(*cpc_desc->sysmem_nodes), + GFP_KERNEL); + if (!cpc_desc->sysmem_nodes) + return -ENOMEM; + + mutex_lock(&cpc_sysmem_lock); + + for (i = 0; i < nr_regs; i++) { + struct cpc_register_resource *reg = &cpc_desc->cpc_regs[i]; + struct cpc_sysmem_node *alias = NULL, *match, *node; + u64 size; + + if (!CPC_SUPPORTED(reg) || !CPC_IN_SYSTEM_MEMORY(reg)) + continue; + + node = &cpc_desc->sysmem_nodes[i]; + size = cpc_sysmem_claim_size(reg); + node->start = reg->cpc_entry.reg.address; + node->last = node->start + size - 1; + node->desc = cpc_desc; + node->reg_idx = i; + INIT_LIST_HEAD(&node->aliases); + INIT_LIST_HEAD(&node->alias_node); + + /* Performance Limited precedes every optional writer we may disable. */ + if (cpc_disable_new_sysmem_writer(cpc_desc, i, node)) + continue; + + match = cpc_sysmem_first(node->start, node->last); + while (match) { + struct cpc_register_resource *match_reg; + + match_reg = &match->desc->cpc_regs[match->reg_idx]; + ret = cpc_validate_sysmem_pair(cpc_desc, i, match->desc, + match->reg_idx); + if (ret) + goto out_unregister; + if (cpc_desc == match->desc) { + reg->cpc_entry.use_rmw_lock = true; + match_reg->cpc_entry.use_rmw_lock = true; + } + if (cpc_same_sysmem_register(i, reg, match->reg_idx, match_reg)) + alias = match; + + match = cpc_sysmem_next(match, node->start, node->last); + } + if (alias) { + node->alias_of = alias; + list_add_tail(&node->alias_node, &alias->aliases); + continue; + } + + cpc_sysmem_itree_insert(node, &cpc_sysmem_tree); + node->registered = true; + } + + mutex_unlock(&cpc_sysmem_lock); + return 0; + +out_unregister: + cpc_unregister_sysmem_desc_locked(cpc_desc); + mutex_unlock(&cpc_sysmem_lock); + return ret; +} + +static void cpc_unregister_sysmem_desc(struct cpc_desc *cpc_desc) +{ + if (!cpc_desc->sysmem_nodes) + return; + + mutex_lock(&cpc_sysmem_lock); + cpc_unregister_sysmem_desc_locked(cpc_desc); + mutex_unlock(&cpc_sysmem_lock); } static ssize_t show_feedback_ctrs(struct kobject *kobj, @@ -297,7 +1361,30 @@ static struct attribute *cppc_attrs[] = { }; ATTRIBUTE_GROUPS(cppc); +static void cppc_free_desc(struct cpc_desc *cpc_ptr) +{ + unsigned int i; + + cpc_unregister_non_mmio_desc(cpc_ptr); + cpc_unregister_sysmem_desc(cpc_ptr); + + for (i = 2; i < cpc_ptr->num_entries; i++) { + void __iomem *addr = cpc_ptr->cpc_regs[i - 2].sys_mem_vaddr; + + if (addr) + iounmap(addr); + } + + kfree(cpc_ptr); +} + +static void cppc_kobj_release(struct kobject *kobj) +{ + cppc_free_desc(to_cpc_desc(kobj)); +} + static const struct kobj_type cppc_ktype = { + .release = cppc_kobj_release, .sysfs_ops = &kobj_sysfs_ops, .default_groups = cppc_groups, }; @@ -321,6 +1408,8 @@ static int check_pcc_chan(int pcc_ss_id, bool chk_err_bit) pcc_ss_data->deadline_us); if (likely(!ret)) { + /* Order completion status before reading the returned payload. */ + rmb(); pcc_ss_data->platform_owns_pcc = false; if (chk_err_bit && (status & PCC_ERROR_MASK)) ret = -EIO; @@ -333,13 +1422,47 @@ static int check_pcc_chan(int pcc_ss_id, bool chk_err_bit) return ret; } +static void cppc_complete_pcc_write(int pcc_ss_id, + struct cppc_pcc_data *pcc_ss_data, int ret) +{ + int i; + + if (unlikely(ret)) { + for_each_possible_cpu(i) { + struct cpc_desc *desc = per_cpu(cpc_desc_ptr, i); + + if (!desc || + per_cpu(cpu_pcc_subspace_idx, i) != pcc_ss_id) + continue; + + if (desc->write_cmd_id == pcc_ss_data->pcc_write_cnt) + desc->write_cmd_status = ret; + } + } + + pcc_ss_data->pcc_write_cnt++; + wake_up_all(&pcc_ss_data->pcc_write_wait_q); +} + +/* The caller must hold pcc_lock for write. */ +static void cppc_abort_pending_pcc_write(int pcc_ss_id, + struct cppc_pcc_data *pcc_ss_data, + int ret) +{ + if (!pcc_ss_data->pending_pcc_write_cmd) + return; + + pcc_ss_data->pending_pcc_write_cmd = false; + cppc_complete_pcc_write(pcc_ss_id, pcc_ss_data, ret); +} + /* * This function transfers the ownership of the PCC to the platform * So it must be called while holding write_lock(pcc_lock) */ static int send_pcc_cmd(int pcc_ss_id, u16 cmd) { - int ret = -EIO, i; + int ret = -EIO; struct cppc_pcc_data *pcc_ss_data = pcc_data[pcc_ss_id]; struct acpi_pcct_shared_memory __iomem *generic_comm_base = pcc_ss_data->pcc_channel->shmem; @@ -431,21 +1554,8 @@ static int send_pcc_cmd(int pcc_ss_id, u16 cmd) mbox_client_txdone(pcc_ss_data->pcc_channel->mchan, ret); end: - if (cmd == CMD_WRITE) { - if (unlikely(ret)) { - for_each_possible_cpu(i) { - struct cpc_desc *desc = per_cpu(cpc_desc_ptr, i); - - if (!desc) - continue; - - if (desc->write_cmd_id == pcc_ss_data->pcc_write_cnt) - desc->write_cmd_status = ret; - } - } - pcc_ss_data->pcc_write_cnt++; - wake_up_all(&pcc_ss_data->pcc_write_wait_q); - } + if (cmd == CMD_WRITE) + cppc_complete_pcc_write(pcc_ss_id, pcc_ss_data, ret); return ret; } @@ -555,7 +1665,7 @@ bool cppc_allow_fast_switch(const struct cpumask *cpus) min_reg = &cpc_ptr->cpc_regs[MIN_PERF]; max_reg = &cpc_ptr->cpc_regs[MAX_PERF]; - if (!CPC_SUPPORTED(desired_reg) || + if (!cpc_is_writable(desired_reg) || (!CPC_IN_SYSTEM_MEMORY(desired_reg) && !CPC_IN_SYSTEM_IO(desired_reg)) || (CPC_SUPPORTED(min_reg) && @@ -643,35 +1753,49 @@ EXPORT_SYMBOL_GPL(acpi_get_psd_map); static int register_pcc_channel(int pcc_ss_idx) { + struct cppc_pcc_data *data; struct pcc_mbox_chan *pcc_chan; u64 usecs_lat; + int ret = 0; - if (pcc_ss_idx >= 0) { - pcc_chan = pcc_mbox_request_channel(&cppc_mbox_cl, pcc_ss_idx); - - if (IS_ERR(pcc_chan)) { - pr_err("Failed to find PCC channel for subspace %d\n", - pcc_ss_idx); - return -ENODEV; - } + if (pcc_ss_idx < 0 || pcc_ss_idx >= MAX_PCC_SUBSPACES) + return -EINVAL; - pcc_data[pcc_ss_idx]->pcc_channel = pcc_chan; - /* - * cppc_ss->latency is just a Nominal value. In reality - * the remote processor could be much slower to reply. - * So add an arbitrary amount of wait on top of Nominal. - */ - usecs_lat = NUM_RETRIES * pcc_chan->latency; - pcc_data[pcc_ss_idx]->deadline_us = usecs_lat; - pcc_data[pcc_ss_idx]->pcc_mrtt = pcc_chan->min_turnaround_time; - pcc_data[pcc_ss_idx]->pcc_mpar = pcc_chan->max_access_rate; - pcc_data[pcc_ss_idx]->pcc_nominal = pcc_chan->latency; + mutex_lock(&pcc_data_lock); + data = pcc_data[pcc_ss_idx]; + if (!data) { + ret = -ENODEV; + goto out_unlock; + } + if (data->pcc_channel_acquired) + goto out_unlock; - /* Set flag so that we don't come here for each CPU. */ - pcc_data[pcc_ss_idx]->pcc_channel_acquired = true; + pcc_chan = pcc_mbox_request_channel(&cppc_mbox_cl, pcc_ss_idx); + if (IS_ERR(pcc_chan)) { + ret = -ENODEV; + goto out_unlock; } - return 0; + data->pcc_channel = pcc_chan; + /* + * cppc_ss->latency is just a Nominal value. In reality + * the remote processor could be much slower to reply. + * So add an arbitrary amount of wait on top of Nominal. + */ + usecs_lat = NUM_RETRIES * pcc_chan->latency; + data->deadline_us = usecs_lat; + data->pcc_mrtt = pcc_chan->min_turnaround_time; + data->pcc_mpar = pcc_chan->max_access_rate; + data->pcc_nominal = pcc_chan->latency; + init_rwsem(&data->pcc_lock); + init_waitqueue_head(&data->pcc_write_wait_q); + + /* Reuse this channel when another CPU references the same subspace. */ + data->pcc_channel_acquired = true; + +out_unlock: + mutex_unlock(&pcc_data_lock); + return ret; } /** @@ -713,19 +1837,50 @@ bool __weak cpc_supported_by_cpu(void) */ static int pcc_data_alloc(int pcc_ss_id) { + struct cppc_pcc_data *data; + int ret = 0; + if (pcc_ss_id < 0 || pcc_ss_id >= MAX_PCC_SUBSPACES) return -EINVAL; - if (pcc_data[pcc_ss_id]) { - pcc_data[pcc_ss_id]->refcount++; - } else { - pcc_data[pcc_ss_id] = kzalloc_obj(struct cppc_pcc_data); - if (!pcc_data[pcc_ss_id]) - return -ENOMEM; - pcc_data[pcc_ss_id]->refcount++; + mutex_lock(&pcc_data_lock); + data = pcc_data[pcc_ss_id]; + if (!data) { + data = kzalloc_obj(struct cppc_pcc_data); + if (!data) { + ret = -ENOMEM; + goto out_unlock; + } + raw_spin_lock_init(&data->payload_lock); + pcc_data[pcc_ss_id] = data; } + data->refcount++; - return 0; +out_unlock: + mutex_unlock(&pcc_data_lock); + return ret; +} + +static void pcc_data_put(int pcc_ss_id) +{ + struct cppc_pcc_data *data; + + if (pcc_ss_id < 0 || pcc_ss_id >= MAX_PCC_SUBSPACES) + return; + + mutex_lock(&pcc_data_lock); + data = pcc_data[pcc_ss_id]; + if (!data || --data->refcount) + goto out_unlock; + + pcc_data[pcc_ss_id] = NULL; + if (data->pcc_channel_acquired) + pcc_mbox_free_channel(data->pcc_channel); + + kfree(data); + +out_unlock: + mutex_unlock(&pcc_data_lock); } /* @@ -772,9 +1927,24 @@ int acpi_cppc_processor_probe(struct acpi_processor *pr) struct device *cpu_dev; acpi_handle handle = pr->handle; unsigned int num_ent, i, cpc_rev; + u32 unsupported_regs = 0; + u32 platform_quirks; int pcc_subspace_id = -1; + bool pcc_data_ref = false; + bool cpc_present = false; acpi_status status; - int ret = -ENODATA; + int ret = -EINVAL; + int err; + + if (per_cpu(cpc_desc_ptr, pr->id)) + return 0; + ret = cpc_get_platform_quirks(&platform_quirks); + if (ret) { + pr_err("CPU%d: failed to match CPPC platform quirks: %d\n", + pr->id, ret); + return ret; + } + per_cpu(cpu_pcc_subspace_idx, pr->id) = -1; if (!osc_sb_cppc2_support_acked) { pr_debug("CPPC v2 _OSC not acked\n"); @@ -791,24 +1961,35 @@ int acpi_cppc_processor_probe(struct acpi_processor *pr) ret = -ENODEV; goto out_buf_free; } + cpc_present = true; + ret = -EINVAL; out_obj = (union acpi_object *) output.pointer; + if (out_obj->package.count < 2) { + pr_debug("Unexpected _CPC package count (%u) for CPU:%d\n", + out_obj->package.count, pr->id); + goto out_buf_free; + } cpc_ptr = kzalloc_obj(struct cpc_desc); if (!cpc_ptr) { ret = -ENOMEM; goto out_buf_free; } + cpc_ptr->cpu_id = pr->id; /* First entry is NumEntries. */ cpc_obj = &out_obj->package.elements[0]; if (cpc_obj->type == ACPI_TYPE_INTEGER) { - num_ent = cpc_obj->integer.value; - if (num_ent <= 1) { - pr_debug("Unexpected _CPC NumEntries value (%d) for CPU:%d\n", - num_ent, pr->id); + if (cpc_obj->integer.value < 2 || + cpc_obj->integer.value > out_obj->package.count) { + pr_debug("Invalid _CPC NumEntries (%llu) for package count (%u) on CPU:%d\n", + cpc_obj->integer.value, out_obj->package.count, + pr->id); goto out_free; } + + num_ent = cpc_obj->integer.value; } else { pr_debug("Unexpected _CPC NumEntries entry type (%d) for CPU:%d\n", cpc_obj->type, pr->id); @@ -818,6 +1999,12 @@ int acpi_cppc_processor_probe(struct acpi_processor *pr) /* Second entry should be revision. */ cpc_obj = &out_obj->package.elements[1]; if (cpc_obj->type == ACPI_TYPE_INTEGER) { + if (cpc_obj->integer.value > U8_MAX) { + pr_debug("Invalid _CPC Revision (%llu) for CPU:%d\n", + cpc_obj->integer.value, pr->id); + ret = -EINVAL; + goto out_free; + } cpc_rev = cpc_obj->integer.value; } else { pr_debug("Unexpected _CPC Revision entry type (%d) for CPU:%d\n", @@ -857,11 +2044,45 @@ int acpi_cppc_processor_probe(struct acpi_processor *pr) cpc_obj = &out_obj->package.elements[i]; if (cpc_obj->type == ACPI_TYPE_INTEGER) { - cpc_ptr->cpc_regs[i-2].type = ACPI_TYPE_INTEGER; - cpc_ptr->cpc_regs[i-2].cpc_entry.int_value = cpc_obj->integer.value; + bool legacy_null; + + if (!cpc_integer_entry_valid(i - 2, + cpc_obj->integer.value, + &legacy_null)) { + pr_debug("Invalid Integer _CPC register %u for CPU:%d\n", + i - 2, pr->id); + ret = -EINVAL; + goto out_free; + } + if (legacy_null) + pr_warn_once(FW_BUG "_CPC register %u uses Integer 0 for an absent Buffer\n", + i - 2); + cpc_ptr->cpc_regs[i - 2].type = ACPI_TYPE_INTEGER; + cpc_ptr->cpc_regs[i - 2].cpc_entry.int_value = cpc_obj->integer.value; } else if (cpc_obj->type == ACPI_TYPE_BUFFER) { + if (cpc_obj->buffer.length < sizeof(*gas_t)) { + pr_debug("Invalid register descriptor for CPU:%d\n", + pr->id); + ret = -EINVAL; + goto out_free; + } + gas_t = (struct cpc_reg *) cpc_obj->buffer.pointer; + if (gas_t->descriptor != CPC_GENERIC_REGISTER_DESCRIPTOR || + gas_t->length != CPC_GENERIC_REGISTER_LENGTH) { + pr_debug("Invalid register resource for CPU:%d\n", + pr->id); + ret = -EINVAL; + goto out_free; + } + + cpc_ptr->cpc_regs[i - 2].type = ACPI_TYPE_BUFFER; + memcpy(&cpc_ptr->cpc_regs[i - 2].cpc_entry.reg, gas_t, + sizeof(*gas_t)); + gas_t = &cpc_ptr->cpc_regs[i - 2].cpc_entry.reg; + cpc_apply_platform_quirks(gas_t, i - 2, + platform_quirks); /* * The PCC Subspace index is encoded inside @@ -870,65 +2091,142 @@ int acpi_cppc_processor_probe(struct acpi_processor *pr) * so extract it only once. */ if (gas_t->space_id == ACPI_ADR_SPACE_PLATFORM_COMM) { + /* These registers have no specified 32-bit upper bound. */ + bool wide_write = i - 2 == PERF_LIMITED || + i - 2 == ENABLE || + i - 2 == AUTO_SEL_ENABLE; + bool write_width_supported = gas_t->bit_width == 8 || + gas_t->bit_width == 16 || + gas_t->bit_width == 32 || + gas_t->bit_width == 64; + bool unsupported; + + unsupported = !gas_t->bit_width || + gas_t->bit_width > 64 || + gas_t->bit_offset || + gas_t->bit_width % 8 || + (cpc_reg_is_writable(i - 2) && + (!write_width_supported || + (!wide_write && gas_t->bit_width > 32))); + if (unsupported) { + if (!cpc_retain_pcc_status(cpc_ptr, i - 2)) + unsupported_regs |= BIT(i - 2); + continue; + } + if (pcc_subspace_id < 0) { pcc_subspace_id = gas_t->access_width; - if (pcc_data_alloc(pcc_subspace_id)) - goto out_free; } else if (pcc_subspace_id != gas_t->access_width) { pr_debug("Mismatched PCC ids in _CPC for CPU:%d\n", pr->id); + ret = -EINVAL; goto out_free; } + + if (!pcc_data_ref) { + err = pcc_data_alloc(pcc_subspace_id); + if (err) { + ret = err; + goto out_free; + } + pcc_data_ref = true; + } } else if (gas_t->space_id == ACPI_ADR_SPACE_SYSTEM_MEMORY) { - if (gas_t->address) { + if (!IS_NULL_REG(gas_t)) { void __iomem *addr; size_t access_width; + err = cpc_validate_sysmem_reg(cpc_ptr, gas_t, + i - 2); + if (err) { + unsupported_regs |= BIT(i - 2); + continue; + } + if (!cpc_is_readable(&cpc_ptr->cpc_regs[i - 2]) && + !cpc_is_writable(&cpc_ptr->cpc_regs[i - 2])) + continue; + if (!osc_cpc_flexible_adr_space_confirmed) { pr_debug("Flexible address space capability not supported\n"); + ret = -EOPNOTSUPP; if (!cpc_supported_by_cpu()) goto out_free; + ret = -EINVAL; } - access_width = GET_BIT_WIDTH(gas_t) / 8; + access_width = cpc_reg_access_width(gas_t); + access_width /= 8; addr = ioremap(gas_t->address, access_width); - if (!addr) + if (!addr) { + ret = -ENOMEM; goto out_free; - cpc_ptr->cpc_regs[i-2].sys_mem_vaddr = addr; + } + cpc_ptr->cpc_regs[i - 2].sys_mem_vaddr = addr; } } else if (gas_t->space_id == ACPI_ADR_SPACE_SYSTEM_IO) { - if (gas_t->access_width < 1 || gas_t->access_width > 3) { - /* - * 1 = 8-bit, 2 = 16-bit, and 3 = 32-bit. - * SystemIO doesn't implement 64-bit - * registers. - */ - pr_debug("Invalid access width %d for SystemIO register in _CPC\n", - gas_t->access_width); - goto out_free; + u64 access_size; + const char *reason = "uses unsupported SystemIO geometry"; + unsigned int access_width; + bool partial = false; + bool unsupported; + + access_width = cpc_reg_access_width(gas_t); + unsupported = !IS_ENABLED(CONFIG_HAS_IOPORT); + if (unsupported) + reason = "requires unavailable SystemIO support"; + else + unsupported = access_width != 8 && + access_width != 16 && + access_width != 32; + if (!unsupported) { + access_size = access_width / 8; + unsupported = !gas_t->bit_width || + gas_t->bit_width > access_width || + gas_t->bit_offset >= access_width || + gas_t->bit_width > access_width - + gas_t->bit_offset; + partial = gas_t->bit_offset || + gas_t->bit_width != access_width; } - if (gas_t->address & OVER_16BTS_MASK) { - /* SystemIO registers use 16-bit integer addresses */ - pr_debug("Invalid IO port %llu for SystemIO register in _CPC\n", - gas_t->address); - goto out_free; + if (!unsupported) { + unsupported = (cpc_reg_is_writable(i - 2) && + i - 2 != PERF_LIMITED && + cpc_is_writable(&cpc_ptr->cpc_regs[i - 2]) && + partial) || + !cpc_reg_access_aligned(gas_t, + access_size) || + gas_t->address > + U16_MAX - (access_size - 1); + } + if (unsupported) { + if (cpc_retain_sysio_status(cpc_ptr, i - 2)) + continue; + pr_debug("CPU%d: _CPC register %u %s\n", + pr->id, i - 2, reason); + unsupported_regs |= BIT(i - 2); + continue; + } + if (i - 2 == PERF_LIMITED && partial) { + pr_warn_once("CPU%d: Performance Limited register cannot be cleared safely; keeping it readable\n", + cpc_ptr->cpu_id); + cpc_ptr->cpc_regs[i - 2].cpc_entry.write_unsupported = true; } if (!osc_cpc_flexible_adr_space_confirmed) { pr_debug("Flexible address space capability not supported\n"); + ret = -EOPNOTSUPP; if (!cpc_supported_by_cpu()) goto out_free; + ret = -EINVAL; } } else { if (gas_t->space_id != ACPI_ADR_SPACE_FIXED_HARDWARE || !cpc_ffh_supported()) { /* Support only PCC, SystemMemory, SystemIO, and FFH type regs. */ pr_debug("Unsupported register type (%d) in _CPC\n", gas_t->space_id); + ret = -EOPNOTSUPP; goto out_free; } } - - cpc_ptr->cpc_regs[i-2].type = ACPI_TYPE_BUFFER; - memcpy(&cpc_ptr->cpc_regs[i-2].cpc_entry.reg, gas_t, sizeof(*gas_t)); } else if (cpc_obj->type == ACPI_TYPE_PACKAGE && (i - 2) == RESOURCE_PRIORITY) { /* * ACPI 6.6, s8.4.6.1.2.7 defines Resource Priority as a @@ -945,17 +2243,18 @@ int acpi_cppc_processor_probe(struct acpi_processor *pr) goto out_free; } } - per_cpu(cpu_pcc_subspace_idx, pr->id) = pcc_subspace_id; - /* - * In CPPC v1, DESIRED_PERF is mandatory. In CPPC v2, it is optional - * only when AUTO_SEL_ENABLE is supported. - */ - if (!CPC_SUPPORTED(&cpc_ptr->cpc_regs[DESIRED_PERF]) && - (!osc_sb_cppc2_support_acked || - !CPC_SUPPORTED(&cpc_ptr->cpc_regs[AUTO_SEL_ENABLE]))) - pr_warn("Desired perf. register is mandatory if CPPC v2 is not supported " - "or autonomous selection is disabled\n"); + per_cpu(cpu_pcc_subspace_idx, pr->id) = pcc_data_ref ? + pcc_subspace_id : -1; + + ret = cpc_resolve_unsupported(cpc_ptr, unsupported_regs); + if (ret) + goto out_free; + unsupported_regs = 0; + + ret = cpc_validate_required_controls(cpc_ptr); + if (ret) + goto out_free; /* * Initialize the remaining cpc_regs as unsupported. @@ -968,8 +2267,6 @@ int acpi_cppc_processor_probe(struct acpi_processor *pr) } - /* Store CPU Logical ID */ - cpc_ptr->cpu_id = pr->id; cpc_mark_rmw_lock_users(cpc_ptr); raw_spin_lock_init(&cpc_ptr->rmw_lock); @@ -978,16 +2275,43 @@ int acpi_cppc_processor_probe(struct acpi_processor *pr) if (ret) goto out_free; + ret = cpc_register_sysmem_desc(cpc_ptr); + if (ret) + goto out_free; + /* Register PCC channel once for all PCC subspace ID. */ - if (pcc_subspace_id >= 0 && !pcc_data[pcc_subspace_id]->pcc_channel_acquired) { + if (pcc_data_ref) { ret = register_pcc_channel(pcc_subspace_id); + if (ret) { + pr_err("Failed to find PCC channel for subspace %d\n", + pcc_subspace_id); + goto out_free; + } + + cpc_validate_pcc_bounds(cpc_ptr, pcc_subspace_id, + pcc_data[pcc_subspace_id], + &unsupported_regs); + + ret = cpc_resolve_unsupported(cpc_ptr, unsupported_regs); if (ret) goto out_free; - init_rwsem(&pcc_data[pcc_subspace_id]->pcc_lock); - init_waitqueue_head(&pcc_data[pcc_subspace_id]->pcc_write_wait_q); + /* A range-only status entry needs the channel only for bounds. */ + if (!cpc_pcc_access_needed(cpc_ptr)) { + pcc_data_put(pcc_subspace_id); + pcc_data_ref = false; + per_cpu(cpu_pcc_subspace_idx, pr->id) = -1; + } } + ret = cpc_validate_bound_controls(cpc_ptr); + if (ret) + goto out_free; + + ret = cpc_register_non_mmio_desc(cpc_ptr); + if (ret) + goto out_free; + /* Everything looks okay */ pr_debug("Parsed CPC struct for CPU: %d\n", pr->id); @@ -999,30 +2323,32 @@ int acpi_cppc_processor_probe(struct acpi_processor *pr) } /* Plug PSD data into this CPU's CPC descriptor. */ - per_cpu(cpc_desc_ptr, pr->id) = cpc_ptr; + cpc_set_desc(pr->id, cpc_ptr); ret = kobject_init_and_add(&cpc_ptr->kobj, &cppc_ktype, &cpu_dev->kobj, "acpi_cppc"); if (ret) { - per_cpu(cpc_desc_ptr, pr->id) = NULL; + cpc_set_desc(pr->id, NULL); + cpc_unregister_non_mmio_desc(cpc_ptr); + cpc_unregister_sysmem_desc(cpc_ptr); kobject_put(&cpc_ptr->kobj); - goto out_free; + goto out_pcc_put; } kfree(output.pointer); return 0; out_free: - /* Free all the mapped sys mem areas for this CPU */ - for (i = 2; i < cpc_ptr->num_entries; i++) { - void __iomem *addr = cpc_ptr->cpc_regs[i-2].sys_mem_vaddr; + cppc_free_desc(cpc_ptr); - if (addr) - iounmap(addr); - } - kfree(cpc_ptr); +out_pcc_put: + if (pcc_data_ref) + pcc_data_put(pcc_subspace_id); + per_cpu(cpu_pcc_subspace_idx, pr->id) = -1; out_buf_free: + if (cpc_present) + pr_err("CPU%d: failed to initialize _CPC: %d\n", pr->id, ret); kfree(output.pointer); return ret; } @@ -1037,34 +2363,24 @@ EXPORT_SYMBOL_GPL(acpi_cppc_processor_probe); void acpi_cppc_processor_exit(struct acpi_processor *pr) { struct cpc_desc *cpc_ptr; - unsigned int i; - void __iomem *addr; - int pcc_ss_id = per_cpu(cpu_pcc_subspace_idx, pr->id); - - if (pcc_ss_id >= 0 && pcc_data[pcc_ss_id]) { - if (pcc_data[pcc_ss_id]->pcc_channel_acquired) { - pcc_data[pcc_ss_id]->refcount--; - if (!pcc_data[pcc_ss_id]->refcount) { - pcc_mbox_free_channel(pcc_data[pcc_ss_id]->pcc_channel); - kfree(pcc_data[pcc_ss_id]); - pcc_data[pcc_ss_id] = NULL; - } - } - } + int pcc_ss_id; cpc_ptr = per_cpu(cpc_desc_ptr, pr->id); - if (!cpc_ptr) + if (!cpc_ptr) { + per_cpu(cpu_pcc_subspace_idx, pr->id) = -1; return; - - /* Free all the mapped sys mem areas for this CPU */ - for (i = 2; i < cpc_ptr->num_entries; i++) { - addr = cpc_ptr->cpc_regs[i-2].sys_mem_vaddr; - if (addr) - iounmap(addr); } + pcc_ss_id = per_cpu(cpu_pcc_subspace_idx, pr->id); + cpc_set_desc(pr->id, NULL); + kobject_del(&cpc_ptr->kobj); + cpc_unregister_non_mmio_desc(cpc_ptr); + cpc_unregister_sysmem_desc(cpc_ptr); + + pcc_data_put(pcc_ss_id); + per_cpu(cpu_pcc_subspace_idx, pr->id) = -1; + kobject_put(&cpc_ptr->kobj); - kfree(cpc_ptr); } EXPORT_SYMBOL_GPL(acpi_cppc_processor_exit); @@ -1123,23 +2439,34 @@ int __weak cpc_write_ffh(int cpunum, struct cpc_reg *reg, u64 val) static int cpc_read(int cpu, struct cpc_register_resource *reg_res, u64 *val) { void __iomem *vaddr = NULL; + unsigned long flags; + u8 buf[sizeof(*val)]; + unsigned int i; int size; int pcc_ss_id = per_cpu(cpu_pcc_subspace_idx, cpu); struct cpc_reg *reg = ®_res->cpc_entry.reg; + if (!cpc_is_readable(reg_res)) + return -EOPNOTSUPP; + if (reg_res->type == ACPI_TYPE_INTEGER) { *val = reg_res->cpc_entry.int_value; return 0; } *val = 0; + if (reg->space_id == ACPI_ADR_SPACE_FIXED_HARDWARE) + return cpc_read_ffh(cpu, reg, val); + size = GET_BIT_WIDTH(reg); - if (IS_ENABLED(CONFIG_HAS_IOPORT) && - reg->space_id == ACPI_ADR_SPACE_SYSTEM_IO) { + if (reg->space_id == ACPI_ADR_SPACE_SYSTEM_IO) { u32 val_u32; acpi_status status; + if (!IS_ENABLED(CONFIG_HAS_IOPORT)) + return -EOPNOTSUPP; + status = acpi_os_read_port((acpi_io_address)reg->address, &val_u32, size); if (ACPI_FAILURE(status)) { @@ -1148,20 +2475,33 @@ static int cpc_read(int cpu, struct cpc_register_resource *reg_res, u64 *val) return -EFAULT; } - *val = val_u32; + *val = MASK_VAL_READ(reg, val_u32); return 0; - } else if (reg->space_id == ACPI_ADR_SPACE_PLATFORM_COMM && pcc_ss_id >= 0) { + } else if (reg->space_id == ACPI_ADR_SPACE_PLATFORM_COMM) { + if (pcc_ss_id < 0 || !pcc_data[pcc_ss_id]) + return -ENODEV; + /* * For registers in PCC space, the register size is determined * by the bit width field; the access size is used to indicate * the PCC subspace id. */ vaddr = GET_PCC_VADDR(reg->address, pcc_ss_id); - } - else if (reg->space_id == ACPI_ADR_SPACE_SYSTEM_MEMORY) + size = reg->bit_width / 8; + if (!size || size > sizeof(buf) || reg->bit_width % 8) + return -EFAULT; + + raw_spin_lock_irqsave(&pcc_data[pcc_ss_id]->payload_lock, flags); + memcpy_fromio(buf, vaddr, size); + raw_spin_unlock_irqrestore(&pcc_data[pcc_ss_id]->payload_lock, + flags); + + *val = 0; + for (i = 0; i < size; i++) + *val |= (u64)buf[i] << (i * 8); + return 0; + } else if (reg->space_id == ACPI_ADR_SPACE_SYSTEM_MEMORY) vaddr = reg_res->sys_mem_vaddr; - else if (reg->space_id == ACPI_ADR_SPACE_FIXED_HARDWARE) - return cpc_read_ffh(cpu, reg, val); else return acpi_os_read_memory((acpi_physical_address)reg->address, val, size); @@ -1180,18 +2520,12 @@ static int cpc_read(int cpu, struct cpc_register_resource *reg_res, u64 *val) *val = readq_relaxed(vaddr); break; default: - if (reg->space_id == ACPI_ADR_SPACE_SYSTEM_MEMORY) { - pr_debug("Error: Cannot read %u bit width from system memory: 0x%llx\n", - size, reg->address); - } else if (reg->space_id == ACPI_ADR_SPACE_PLATFORM_COMM) { - pr_debug("Error: Cannot read %u bit width from PCC for ss: %d\n", - size, pcc_ss_id); - } + pr_debug("Error: Cannot read %u bit width from system memory: 0x%llx\n", + size, reg->address); return -EFAULT; } - if (reg->space_id == ACPI_ADR_SPACE_SYSTEM_MEMORY) - *val = MASK_VAL_READ(reg, *val); + *val = MASK_VAL_READ(reg, *val); return 0; } @@ -1203,17 +2537,28 @@ static int cpc_write(int cpu, struct cpc_register_resource *reg_res, u64 val) u64 prev_val; void __iomem *vaddr = NULL; int pcc_ss_id = per_cpu(cpu_pcc_subspace_idx, cpu); - struct cpc_reg *reg = ®_res->cpc_entry.reg; + struct cpc_reg *reg; struct cpc_desc *cpc_desc; unsigned long flags; + u8 buf[sizeof(val)]; + unsigned int i; bool locked = false; + if (!cpc_is_writable(reg_res)) + return -EOPNOTSUPP; + + reg = ®_res->cpc_entry.reg; + if (reg->space_id == ACPI_ADR_SPACE_FIXED_HARDWARE) + return cpc_write_ffh(cpu, reg, val); + size = GET_BIT_WIDTH(reg); - if (IS_ENABLED(CONFIG_HAS_IOPORT) && - reg->space_id == ACPI_ADR_SPACE_SYSTEM_IO) { + if (reg->space_id == ACPI_ADR_SPACE_SYSTEM_IO) { acpi_status status; + if (!IS_ENABLED(CONFIG_HAS_IOPORT)) + return -EOPNOTSUPP; + status = acpi_os_write_port((acpi_io_address)reg->address, (u32)val, size); if (ACPI_FAILURE(status)) { @@ -1223,60 +2568,72 @@ static int cpc_write(int cpu, struct cpc_register_resource *reg_res, u64 val) } return 0; - } else if (reg->space_id == ACPI_ADR_SPACE_PLATFORM_COMM && pcc_ss_id >= 0) { + } else if (reg->space_id == ACPI_ADR_SPACE_PLATFORM_COMM) { + if (pcc_ss_id < 0 || !pcc_data[pcc_ss_id]) + return -ENODEV; + /* * For registers in PCC space, the register size is determined * by the bit width field; the access size is used to indicate * the PCC subspace id. */ vaddr = GET_PCC_VADDR(reg->address, pcc_ss_id); - } - else if (reg->space_id == ACPI_ADR_SPACE_SYSTEM_MEMORY) + size = reg->bit_width / 8; + if (!size || size > sizeof(buf) || reg->bit_width % 8) + return -EFAULT; + + for (i = 0; i < size; i++) + buf[i] = val >> (i * 8); + + raw_spin_lock_irqsave(&pcc_data[pcc_ss_id]->payload_lock, flags); + memcpy_toio(vaddr, buf, size); + /* Publish every payload byte before another CPU can ring the doorbell. */ + wmb(); + raw_spin_unlock_irqrestore(&pcc_data[pcc_ss_id]->payload_lock, + flags); + return 0; + } else if (reg->space_id == ACPI_ADR_SPACE_SYSTEM_MEMORY) vaddr = reg_res->sys_mem_vaddr; - else if (reg->space_id == ACPI_ADR_SPACE_FIXED_HARDWARE) - return cpc_write_ffh(cpu, reg, val); else return acpi_os_write_memory((acpi_physical_address)reg->address, val, size); - if (reg->space_id == ACPI_ADR_SPACE_SYSTEM_MEMORY) { - /* - * The _CPC layout is immutable after probe. The precomputed flag - * retains serialization for partial fields or overlapping access - * units; standalone full-width registers avoid the lock. - */ - locked = reg_res->cpc_entry.use_rmw_lock; - if (locked) { - cpc_desc = per_cpu(cpc_desc_ptr, cpu); - if (!cpc_desc) { - pr_debug("No CPC descriptor for CPU:%d\n", cpu); - return -ENODEV; - } - raw_spin_lock_irqsave(&cpc_desc->rmw_lock, flags); + /* Partial fields and local overlaps use the descriptor lock. */ + locked = reg_res->cpc_entry.use_rmw_lock; + if (locked) { + cpc_desc = per_cpu(cpc_desc_ptr, cpu); + if (!cpc_desc) { + pr_debug("No CPC descriptor for CPU:%d\n", cpu); + return -ENODEV; } + raw_spin_lock_irqsave(&cpc_desc->rmw_lock, flags); + } - if (reg->bit_offset || reg->bit_width != size) { - switch (size) { - case 8: - prev_val = readb_relaxed(vaddr); - break; - case 16: - prev_val = readw_relaxed(vaddr); - break; - case 32: - prev_val = readl_relaxed(vaddr); - break; - case 64: - prev_val = readq_relaxed(vaddr); - break; - default: - if (locked) - raw_spin_unlock_irqrestore(&cpc_desc->rmw_lock, - flags); - return -EFAULT; - } - val = MASK_VAL_WRITE(reg, prev_val, val); + if (reg->bit_offset || reg->bit_width != size) { + /* + * MASK_VAL_WRITE() discards the field's old bits, so undefined + * readback from a write-only field is not propagated. + */ + switch (size) { + case 8: + prev_val = readb_relaxed(vaddr); + break; + case 16: + prev_val = readw_relaxed(vaddr); + break; + case 32: + prev_val = readl_relaxed(vaddr); + break; + case 64: + prev_val = readq_relaxed(vaddr); + break; + default: + if (locked) + raw_spin_unlock_irqrestore(&cpc_desc->rmw_lock, + flags); + return -EFAULT; } + val = MASK_VAL_WRITE(reg, prev_val, val); } switch (size) { @@ -1293,19 +2650,17 @@ static int cpc_write(int cpu, struct cpc_register_resource *reg_res, u64 val) writeq_relaxed(val, vaddr); break; default: - if (reg->space_id == ACPI_ADR_SPACE_SYSTEM_MEMORY) { - pr_debug("Error: Cannot write %u bit width to system memory: 0x%llx\n", - size, reg->address); - } else if (reg->space_id == ACPI_ADR_SPACE_PLATFORM_COMM) { - pr_debug("Error: Cannot write %u bit width to PCC for ss: %d\n", - size, pcc_ss_id); - } + pr_debug("Error: Cannot write %u bit width to system memory: 0x%llx\n", + size, reg->address); ret_val = -EFAULT; break; } - if (locked) + if (locked) { + if (!ret_val) + mmiowb_set_pending(); raw_spin_unlock_irqrestore(&cpc_desc->rmw_lock, flags); + } return ret_val; } @@ -1347,15 +2702,25 @@ static int cppc_get_reg_val(int cpu, enum cppc_regs reg_idx, u64 *val) pr_debug("No CPC descriptor for CPU:%d\n", cpu); return -ENODEV; } + if (cpc_reg_is_write_only(cpc_desc, reg_idx)) + return -EOPNOTSUPP; reg = &cpc_desc->cpc_regs[reg_idx]; - if ((reg->type == ACPI_TYPE_INTEGER && IS_OPTIONAL_CPC_REG(reg_idx) && + /* + * Desired and Performance Limited may be disabled despite not being + * generally optional. + */ + if ((reg->type == ACPI_TYPE_INTEGER && + (IS_OPTIONAL_CPC_REG(reg_idx) || reg_idx == DESIRED_PERF || + reg_idx == PERF_LIMITED) && !reg->cpc_entry.int_value) || (reg->type != ACPI_TYPE_INTEGER && IS_NULL_REG(®->cpc_entry.reg))) { pr_debug("CPC register is not supported\n"); return -EOPNOTSUPP; } + if (!cpc_is_readable(reg)) + return -EOPNOTSUPP; if (CPC_IN_PCC(reg)) return cppc_get_reg_val_in_pcc(cpu, reg, val); @@ -1366,23 +2731,33 @@ static int cppc_get_reg_val(int cpu, enum cppc_regs reg_idx, u64 *val) static int cppc_set_reg_val_in_pcc(int cpu, struct cpc_register_resource *reg, u64 val) { int pcc_ss_id = per_cpu(cpu_pcc_subspace_idx, cpu); - struct cppc_pcc_data *pcc_ss_data = NULL; + struct cppc_pcc_data *pcc_ss_data; int ret; if (pcc_ss_id < 0) { pr_debug("Invalid pcc_ss_id\n"); return -ENODEV; } + pcc_ss_data = pcc_data[pcc_ss_id]; + if (!pcc_ss_data) + return -ENODEV; - ret = cpc_write(cpu, reg, val); + down_write(&pcc_ss_data->pcc_lock); + + ret = check_pcc_chan(pcc_ss_id, false); if (ret) - return ret; + goto out; - pcc_ss_data = pcc_data[pcc_ss_id]; + ret = cpc_write(cpu, reg, val); + if (ret) + goto out; - down_write(&pcc_ss_data->pcc_lock); /* after writing CPC, transfer the ownership of PCC to platform */ ret = send_pcc_cmd(pcc_ss_id, CMD_WRITE); + +out: + if (ret) + cppc_abort_pending_pcc_write(pcc_ss_id, pcc_ss_data, ret); up_write(&pcc_ss_data->pcc_lock); return ret; @@ -1400,8 +2775,13 @@ static int cppc_set_reg_val(int cpu, enum cppc_regs reg_idx, u64 val) reg = &cpc_desc->cpc_regs[reg_idx]; + /* Integer 1 describes autonomous selection that is always enabled. */ + if (reg_idx == AUTO_SEL_ENABLE && reg->type == ACPI_TYPE_INTEGER && + reg->cpc_entry.int_value == 1) + return val == 1 ? 0 : -EOPNOTSUPP; + /* if a register is writeable, it must be a buffer and not null */ - if ((reg->type != ACPI_TYPE_BUFFER) || IS_NULL_REG(®->cpc_entry.reg)) { + if (!cpc_is_writable(reg)) { pr_debug("CPC register is not supported\n"); return -EOPNOTSUPP; } @@ -1412,11 +2792,6 @@ static int cppc_set_reg_val(int cpu, enum cppc_regs reg_idx, u64 val) return cpc_write(cpu, reg, val); } -static bool cppc_desired_perf_readable(const struct cpc_desc *cpc_desc) -{ - return cpc_desc->version < CPPC_V4_REV; -} - /** * cppc_get_desired_perf - Get the desired performance register value. * @cpunum: CPU from which to get desired performance. @@ -1427,15 +2802,6 @@ static bool cppc_desired_perf_readable(const struct cpc_desc *cpc_desc) */ int cppc_get_desired_perf(int cpunum, u64 *desired_perf) { - struct cpc_desc *cpc_desc = per_cpu(cpc_desc_ptr, cpunum); - - if (!cpc_desc) - return -ENODEV; - - /* _CPC revision 4 no longer specifies Desired Performance as readable. */ - if (!cppc_desired_perf_readable(cpc_desc)) - return -EOPNOTSUPP; - return cppc_get_reg_val(cpunum, DESIRED_PERF, desired_perf); } EXPORT_SYMBOL_GPL(cppc_get_desired_perf); @@ -1491,7 +2857,7 @@ int cppc_get_perf_caps(int cpunum, struct cppc_perf_caps *perf_caps) struct cpc_register_resource *highest_reg, *lowest_reg, *lowest_non_linear_reg, *nominal_reg, *reference_reg, *guaranteed_reg, *low_freq_reg = NULL, *nom_freq_reg = NULL; - u64 high, low, guaranteed, nom, ref, min_nonlinear, + u64 high, low, guaranteed = 0, nom, ref, min_nonlinear, low_f = 0, nom_f = 0; int pcc_ss_id = per_cpu(cpu_pcc_subspace_idx, cpunum); struct cppc_pcc_data *pcc_ss_data = NULL; @@ -1574,7 +2940,12 @@ int cppc_get_perf_caps(int cpunum, struct cppc_perf_caps *perf_caps) goto out_err; perf_caps->lowest_nonlinear_perf = min_nonlinear; - if (!high || !low || !nom || !ref || !min_nonlinear) { + if (!high || !low || !nom || !ref || !min_nonlinear || + high > U32_MAX || low > U32_MAX || guaranteed > U32_MAX || + nom > U32_MAX || ref > U32_MAX || min_nonlinear > U32_MAX || + high < nom || nom < min_nonlinear || min_nonlinear < low || + (CPC_SUPPORTED(guaranteed_reg) && + (guaranteed < low || guaranteed > nom))) { ret = -EFAULT; goto out_err; } @@ -1591,6 +2962,14 @@ int cppc_get_perf_caps(int cpunum, struct cppc_perf_caps *perf_caps) if (ret) goto out_err; } + /* Require ordered anchors and a nonzero slope when frequencies differ. */ + if (low_f > U32_MAX || nom_f > U32_MAX || + (low_f && nom_f && + (nom_f < low_f || nom < low || + (nom_f != low_f && nom == low)))) { + ret = -EFAULT; + goto out_err; + } perf_caps->lowest_freq = low_f; perf_caps->nominal_freq = nom_f; @@ -1613,6 +2992,9 @@ bool cppc_perf_ctrs_in_pcc_cpu(unsigned int cpu) { struct cpc_desc *cpc_desc = per_cpu(cpc_desc_ptr, cpu); + if (!cpc_desc) + return false; + return CPC_IN_PCC(&cpc_desc->cpc_regs[DELIVERED_CTR]) || CPC_IN_PCC(&cpc_desc->cpc_regs[REFERENCE_CTR]) || CPC_IN_PCC(&cpc_desc->cpc_regs[CTR_WRAP_TIME]); @@ -1754,8 +3136,10 @@ int cppc_set_epp_perf(int cpu, struct cppc_perf_ctrls *perf_ctrls, bool enable) struct cpc_register_resource *auto_sel_reg; struct cpc_desc *cpc_desc = per_cpu(cpc_desc_ptr, cpu); struct cppc_pcc_data *pcc_ss_data = NULL; - bool autosel_ffh_sysmem; - bool epp_ffh_sysmem; + bool auto_sel_pcc; + bool auto_sel_non_pcc; + bool epp_pcc; + bool epp_non_pcc; int ret; if (!cpc_desc) { @@ -1765,53 +3149,69 @@ int cppc_set_epp_perf(int cpu, struct cppc_perf_ctrls *perf_ctrls, bool enable) auto_sel_reg = &cpc_desc->cpc_regs[AUTO_SEL_ENABLE]; epp_set_reg = &cpc_desc->cpc_regs[ENERGY_PERF]; + if (!enable && auto_sel_reg->type == ACPI_TYPE_INTEGER && + auto_sel_reg->cpc_entry.int_value == 1) + return -EOPNOTSUPP; + + auto_sel_pcc = cpc_is_writable(auto_sel_reg) && + CPC_IN_PCC(auto_sel_reg); + epp_pcc = cpc_is_writable(epp_set_reg) && CPC_IN_PCC(epp_set_reg); - epp_ffh_sysmem = CPC_SUPPORTED(epp_set_reg) && - (CPC_IN_FFH(epp_set_reg) || CPC_IN_SYSTEM_MEMORY(epp_set_reg)); - autosel_ffh_sysmem = CPC_SUPPORTED(auto_sel_reg) && - (CPC_IN_FFH(auto_sel_reg) || CPC_IN_SYSTEM_MEMORY(auto_sel_reg)); + auto_sel_non_pcc = cpc_is_writable(auto_sel_reg) && !auto_sel_pcc; + epp_non_pcc = cpc_is_writable(epp_set_reg) && !epp_pcc; - if (CPC_IN_PCC(epp_set_reg) || CPC_IN_PCC(auto_sel_reg)) { + /* Complete fallible non-PCC writes before staging PCC data. */ + if (auto_sel_non_pcc) { + ret = cpc_write(cpu, auto_sel_reg, enable); + if (ret) + return ret; + } + if (epp_non_pcc) { + ret = cpc_write(cpu, epp_set_reg, perf_ctrls->energy_perf); + if (ret) + return ret; + } + + if (epp_pcc || auto_sel_pcc) { if (pcc_ss_id < 0) { pr_debug("Invalid pcc_ss_id for CPU:%d\n", cpu); return -ENODEV; } - if (CPC_SUPPORTED(auto_sel_reg)) { + pcc_ss_data = pcc_data[pcc_ss_id]; + if (!pcc_ss_data) + return -ENODEV; + + down_write(&pcc_ss_data->pcc_lock); + + ret = check_pcc_chan(pcc_ss_id, false); + if (ret) + goto out_unlock; + + if (auto_sel_pcc) { ret = cpc_write(cpu, auto_sel_reg, enable); if (ret) - return ret; + goto out_unlock; } - if (CPC_SUPPORTED(epp_set_reg)) { + if (epp_pcc) { ret = cpc_write(cpu, epp_set_reg, perf_ctrls->energy_perf); if (ret) - return ret; + goto out_unlock; } - pcc_ss_data = pcc_data[pcc_ss_id]; - - down_write(&pcc_ss_data->pcc_lock); /* after writing CPC, transfer the ownership of PCC to platform */ ret = send_pcc_cmd(pcc_ss_id, CMD_WRITE); - up_write(&pcc_ss_data->pcc_lock); - } else if (osc_cpc_flexible_adr_space_confirmed && - (epp_ffh_sysmem || autosel_ffh_sysmem)) { - if (autosel_ffh_sysmem) { - ret = cpc_write(cpu, auto_sel_reg, enable); - if (ret) - return ret; - } - if (epp_ffh_sysmem) { - ret = cpc_write(cpu, epp_set_reg, - perf_ctrls->energy_perf); - if (ret) - return ret; - } +out_unlock: + if (ret) + cppc_abort_pending_pcc_write(pcc_ss_id, pcc_ss_data, ret); + up_write(&pcc_ss_data->pcc_lock); + } else if (epp_non_pcc || auto_sel_non_pcc) { + ret = 0; } else { - ret = -ENOTSUPP; - pr_debug("_CPC in PCC/FFH/SystemMemory are not supported\n"); + ret = -EOPNOTSUPP; + pr_debug("No writable EPP controls for CPU:%d\n", cpu); } return ret; @@ -1925,6 +3325,28 @@ int cppc_get_auto_sel(int cpu, bool *enable) EXPORT_SYMBOL_GPL(cppc_get_auto_sel); /** + * cppc_auto_sel_is_immutable - Check for always-enabled autonomous selection. + * @cpu: CPU whose _CPC descriptor to check. + * + * Context: Process context. + * Return: true for Integer 1, false for a register or an absent descriptor. + */ +bool cppc_auto_sel_is_immutable(int cpu) +{ + struct cpc_desc *cpc_desc; + struct cpc_register_resource *reg; + + guard(mutex)(&cpc_desc_lock); + cpc_desc = per_cpu(cpc_desc_ptr, cpu); + if (!cpc_desc) + return false; + + reg = &cpc_desc->cpc_regs[AUTO_SEL_ENABLE]; + return reg->type == ACPI_TYPE_INTEGER && reg->cpc_entry.int_value == 1; +} +EXPORT_SYMBOL_GPL(cppc_auto_sel_is_immutable); + +/** * cppc_set_auto_sel - Write autonomous selection register. * @cpu : CPU to which to write register. * @enable : the desired value of autonomous selection resiter to be updated. @@ -1982,6 +3404,7 @@ int cppc_get_perf(int cpu, struct cppc_perf_ctrls *perf_ctrls) max_perf_reg = &cpc_desc->cpc_regs[MAX_PERF]; energy_perf_reg = &cpc_desc->cpc_regs[ENERGY_PERF]; auto_sel_reg = &cpc_desc->cpc_regs[AUTO_SEL_ENABLE]; + perf_ctrls->min_perf_valid = false; /* Are any of the regs PCC ?*/ if (CPC_IN_PCC(min_perf_reg) || CPC_IN_PCC(max_perf_reg) || @@ -2006,6 +3429,10 @@ int cppc_get_perf(int cpu, struct cppc_perf_ctrls *perf_ctrls) ret = cpc_read(cpu, max_perf_reg, &max); if (ret) goto out_err; + if (max > U32_MAX) { + ret = -EFAULT; + goto out_err; + } } perf_ctrls->max_perf = max; @@ -2013,6 +3440,11 @@ int cppc_get_perf(int cpu, struct cppc_perf_ctrls *perf_ctrls) ret = cpc_read(cpu, min_perf_reg, &min); if (ret) goto out_err; + if (min > U32_MAX) { + ret = -EFAULT; + goto out_err; + } + perf_ctrls->min_perf_valid = true; } perf_ctrls->min_perf = min; @@ -2052,7 +3484,9 @@ int cppc_set_perf(int cpu, struct cppc_perf_ctrls *perf_ctrls) struct cpc_register_resource *desired_reg, *min_perf_reg, *max_perf_reg; int pcc_ss_id = per_cpu(cpu_pcc_subspace_idx, cpu); struct cppc_pcc_data *pcc_ss_data = NULL; - bool regs_in_pcc; + bool desired_update, min_update, max_update; + bool desired_pcc, min_pcc, max_pcc, pcc_update; + bool pcc_layout, direct_layout, mixed_layout; int ret = 0; if (!cpc_desc) { @@ -2063,54 +3497,162 @@ int cppc_set_perf(int cpu, struct cppc_perf_ctrls *perf_ctrls) desired_reg = &cpc_desc->cpc_regs[DESIRED_PERF]; min_perf_reg = &cpc_desc->cpc_regs[MIN_PERF]; max_perf_reg = &cpc_desc->cpc_regs[MAX_PERF]; - regs_in_pcc = CPC_IN_PCC(desired_reg) || CPC_IN_PCC(min_perf_reg) || - CPC_IN_PCC(max_perf_reg); - - /* - * This is Phase-I where we want to write to CPC registers - * -> We want all CPUs to be able to execute this phase in parallel - * - * Since read_lock can be acquired by multiple CPUs simultaneously we - * achieve that goal here - */ - if (regs_in_pcc) { + desired_update = cpc_is_writable(desired_reg); + min_update = cpc_is_writable(min_perf_reg) && + (perf_ctrls->min_perf || perf_ctrls->min_perf_valid); + max_update = cpc_is_writable(max_perf_reg) && + perf_ctrls->max_perf; + desired_pcc = desired_update && CPC_IN_PCC(desired_reg); + min_pcc = min_update && CPC_IN_PCC(min_perf_reg); + max_pcc = max_update && CPC_IN_PCC(max_perf_reg); + pcc_update = desired_pcc || min_pcc || max_pcc; + pcc_layout = (cpc_is_writable(desired_reg) && CPC_IN_PCC(desired_reg)) || + (cpc_is_writable(min_perf_reg) && CPC_IN_PCC(min_perf_reg)) || + (cpc_is_writable(max_perf_reg) && CPC_IN_PCC(max_perf_reg)); + direct_layout = (cpc_is_writable(desired_reg) && + !CPC_IN_PCC(desired_reg)) || + (cpc_is_writable(min_perf_reg) && + !CPC_IN_PCC(min_perf_reg)) || + (cpc_is_writable(max_perf_reg) && + !CPC_IN_PCC(max_perf_reg)); + mixed_layout = pcc_layout && direct_layout; + + if (mixed_layout || pcc_update) { if (pcc_ss_id < 0) { pr_debug("Invalid pcc_ss_id\n"); return -ENODEV; } pcc_ss_data = pcc_data[pcc_ss_id]; - down_read(&pcc_ss_data->pcc_lock); /* BEGIN Phase-I */ + if (!pcc_ss_data) + return -ENODEV; + } + + /* + * A mixed layout cannot batch fallible direct writes safely: another + * CPU's staged PCC values may no longer match if a direct write fails. + * Serialize the complete mixed transaction and drain an older batch + * before changing a direct control. + */ + if (mixed_layout) { + down_write(&pcc_ss_data->pcc_lock); + if (pcc_ss_data->pending_pcc_write_cmd) { + ret = send_pcc_cmd(pcc_ss_id, CMD_WRITE); + if (ret) + goto out_mixed_unlock; + } + if (pcc_ss_data->platform_owns_pcc) { ret = check_pcc_chan(pcc_ss_id, false); - if (ret) { - up_read(&pcc_ss_data->pcc_lock); + if (ret) + goto out_mixed_unlock; + } + + if (desired_update && !desired_pcc) { + ret = cpc_write(cpu, desired_reg, + perf_ctrls->desired_perf); + if (ret) + goto out_mixed_unlock; + } + if (min_update && !min_pcc) { + ret = cpc_write(cpu, min_perf_reg, + perf_ctrls->min_perf); + if (ret) + goto out_mixed_unlock; + } + if (max_update && !max_pcc) { + ret = cpc_write(cpu, max_perf_reg, + perf_ctrls->max_perf); + if (ret) + goto out_mixed_unlock; + } + + if (desired_pcc) { + ret = cpc_write(cpu, desired_reg, + perf_ctrls->desired_perf); + if (ret) + goto out_mixed_unlock; + } + if (min_pcc) { + ret = cpc_write(cpu, min_perf_reg, + perf_ctrls->min_perf); + if (ret) + goto out_mixed_unlock; + } + if (max_pcc) { + ret = cpc_write(cpu, max_perf_reg, + perf_ctrls->max_perf); + if (ret) + goto out_mixed_unlock; + } + + if (pcc_update) { + WRITE_ONCE(pcc_ss_data->pending_pcc_write_cmd, true); + cpc_desc->write_cmd_id = pcc_ss_data->pcc_write_cnt; + cpc_desc->write_cmd_status = 0; + ret = send_pcc_cmd(pcc_ss_id, CMD_WRITE); + } + +out_mixed_unlock: + up_write(&pcc_ss_data->pcc_lock); + return ret; + } + + /* A request without PCC updates has no payload to coordinate. */ + if (!pcc_update) { + if (desired_update) { + ret = cpc_write(cpu, desired_reg, + perf_ctrls->desired_perf); + if (ret) return ret; - } } - /* - * Update the pending_write to make sure a PCC CMD_READ will not - * arrive and steal the channel during the switch to write lock - */ - pcc_ss_data->pending_pcc_write_cmd = true; - cpc_desc->write_cmd_id = pcc_ss_data->pcc_write_cnt; - cpc_desc->write_cmd_status = 0; + if (min_update) { + ret = cpc_write(cpu, min_perf_reg, + perf_ctrls->min_perf); + if (ret) + return ret; + } + if (max_update) + ret = cpc_write(cpu, max_perf_reg, + perf_ctrls->max_perf); + return ret; } - if (CPC_SUPPORTED(desired_reg)) - cpc_write(cpu, desired_reg, perf_ctrls->desired_perf); + down_read(&pcc_ss_data->pcc_lock); /* BEGIN Phase-I */ + if (pcc_ss_data->platform_owns_pcc) { + ret = check_pcc_chan(pcc_ss_id, false); + if (ret) + goto out_pcc_read_unlock; + } /* - * Only write if min_perf and max_perf not zero. Some drivers pass zero - * value to min and max perf, but they don't mean to set the zero value, - * they just don't want to write to those registers. + * This is Phase-I where we want to write to CPC registers + * -> We want all CPUs to be able to execute this phase in parallel + * + * Since read_lock can be acquired by multiple CPUs simultaneously we + * achieve that goal here. */ - if (perf_ctrls->min_perf && CPC_SUPPORTED(min_perf_reg)) - cpc_write(cpu, min_perf_reg, perf_ctrls->min_perf); - if (perf_ctrls->max_perf && CPC_SUPPORTED(max_perf_reg)) - cpc_write(cpu, max_perf_reg, perf_ctrls->max_perf); + if (desired_pcc) { + ret = cpc_write(cpu, desired_reg, perf_ctrls->desired_perf); + if (ret) + goto out_pcc_read_unlock; + } - if (regs_in_pcc) - up_read(&pcc_ss_data->pcc_lock); /* END Phase-I */ + if (min_pcc) { + ret = cpc_write(cpu, min_perf_reg, perf_ctrls->min_perf); + if (ret) + goto out_pcc_read_unlock; + } + if (max_pcc) { + ret = cpc_write(cpu, max_perf_reg, perf_ctrls->max_perf); + if (ret) + goto out_pcc_read_unlock; + } + + /* Block a PCC read until the staged payload has been submitted. */ + WRITE_ONCE(pcc_ss_data->pending_pcc_write_cmd, true); + cpc_desc->write_cmd_id = pcc_ss_data->pcc_write_cnt; + cpc_desc->write_cmd_status = 0; + up_read(&pcc_ss_data->pcc_lock); /* END Phase-I */ /* * This is Phase-II where we transfer the ownership of PCC to Platform * @@ -2157,20 +3699,22 @@ int cppc_set_perf(int cpu, struct cppc_perf_ctrls *perf_ctrls) * case during a CMD_READ and if there are pending writes it delivers * the write command before servicing the read command */ - if (regs_in_pcc) { - if (down_write_trylock(&pcc_ss_data->pcc_lock)) {/* BEGIN Phase-II */ - /* Update only if there are pending write commands */ - if (pcc_ss_data->pending_pcc_write_cmd) - send_pcc_cmd(pcc_ss_id, CMD_WRITE); - up_write(&pcc_ss_data->pcc_lock); /* END Phase-II */ - } else - /* Wait until pcc_write_cnt is updated by send_pcc_cmd */ - wait_event(pcc_ss_data->pcc_write_wait_q, - cpc_desc->write_cmd_id != pcc_ss_data->pcc_write_cnt); - - /* send_pcc_cmd updates the status in case of failure */ - ret = cpc_desc->write_cmd_status; + if (down_write_trylock(&pcc_ss_data->pcc_lock)) {/* BEGIN Phase-II */ + /* Update only if there are pending write commands */ + if (pcc_ss_data->pending_pcc_write_cmd) + send_pcc_cmd(pcc_ss_id, CMD_WRITE); + up_write(&pcc_ss_data->pcc_lock); /* END Phase-II */ + } else { + /* Wait until pcc_write_cnt is updated by send_pcc_cmd */ + wait_event(pcc_ss_data->pcc_write_wait_q, + cpc_desc->write_cmd_id != pcc_ss_data->pcc_write_cnt); } + + /* send_pcc_cmd updates the status in case of failure */ + return cpc_desc->write_cmd_status; + +out_pcc_read_unlock: + up_read(&pcc_ss_data->pcc_lock); return ret; } EXPORT_SYMBOL_GPL(cppc_set_perf); @@ -2194,7 +3738,7 @@ EXPORT_SYMBOL_GPL(cppc_get_perf_limited); /** * cppc_set_perf_limited() - Clear bits in the Performance Limited register. * @cpu: CPU on which to write register. - * @bits_to_clear: Bitmask of bits to clear in the perf_limited register. + * @bits_to_clear: Zero for no-op or CPPC_PERF_LIMITED_MASK to clear both bits. * * The Performance Limited register contains two sticky bits set by platform: * - Bit 0 (Desired_Excursion): Set when delivered performance is constrained @@ -2203,31 +3747,29 @@ EXPORT_SYMBOL_GPL(cppc_get_perf_limited); * below minimum performance. * * These bits are sticky and remain set until OSPM explicitly clears them. - * This function only allows clearing bits (the platform sets them). + * Selective clears are unsupported because they require an interlocked RMW. * * Return: 0 for success, -EINVAL for invalid bits, -EIO on register * access failure, -EOPNOTSUPP if not supported. */ int cppc_set_perf_limited(int cpu, u64 bits_to_clear) { - u64 current_val, new_val; - int ret; - /* Only bits 0 and 1 are valid */ - if (bits_to_clear & ~CPPC_PERF_LIMITED_MASK) + if (bits_to_clear & ~(u64)CPPC_PERF_LIMITED_MASK) return -EINVAL; if (!bits_to_clear) return 0; - ret = cppc_get_perf_limited(cpu, ¤t_val); - if (ret) - return ret; - - /* Clear the specified bits */ - new_val = current_val & ~bits_to_clear; + /* + * Writing zero clears both bits without depending on how a platform + * treats written ones. ACPI does not define the effect of writing one, + * so a selective clear cannot be implemented without an interlocked RMW. + */ + if (bits_to_clear != CPPC_PERF_LIMITED_MASK) + return -EOPNOTSUPP; - return cppc_set_reg_val(cpu, PERF_LIMITED, new_val); + return cppc_set_reg_val(cpu, PERF_LIMITED, 0); } EXPORT_SYMBOL_GPL(cppc_set_perf_limited); @@ -2267,6 +3809,9 @@ int cppc_get_transition_latency(int cpu_num) return -ENODATA; desired_reg = &cpc_desc->cpc_regs[DESIRED_PERF]; + if (!cpc_is_writable(desired_reg)) + return -ENODATA; + if (CPC_IN_SYSTEM_MEMORY(desired_reg) || CPC_IN_SYSTEM_IO(desired_reg)) return 0; diff --git a/drivers/acpi/device_pm.c b/drivers/acpi/device_pm.c index aa55ecfc2923..76104f715efa 100644 --- a/drivers/acpi/device_pm.c +++ b/drivers/acpi/device_pm.c @@ -23,6 +23,8 @@ #include "fan.h" #include "internal.h" +#define ACPI_D_STATE_INVALID ACPI_D_STATE_COUNT + /** * acpi_power_state_string - String representation of ACPI device power state. * @state: ACPI device power state to return the string representation of. @@ -75,15 +77,14 @@ static int acpi_dev_pm_explicit_get(struct acpi_device *device, int *state) int acpi_device_get_power(struct acpi_device *device, int *state) { int result = ACPI_STATE_UNKNOWN; - struct acpi_device *parent; int error; if (!device || !state) return -EINVAL; - parent = acpi_dev_parent(device); - if (!device->flags.power_manageable) { + struct acpi_device *parent = acpi_dev_parent(device); + /* TBD: Non-recursive algorithm for walking up hierarchy. */ *state = parent ? parent->power.state : ACPI_STATE_D0; goto out; @@ -119,16 +120,6 @@ int acpi_device_get_power(struct acpi_device *device, int *state) result = psc > ACPI_STATE_D2 ? ACPI_STATE_D3_HOT : psc; } - /* - * If we were unsure about the device parent's power state up to this - * point, the fact that the device is in D0 implies that the parent has - * to be in D0 too, except if ignore_parent is set. - */ - if (!device->power.flags.ignore_parent && parent && - parent->power.state == ACPI_STATE_UNKNOWN && - result == ACPI_STATE_D0) - parent->power.state = ACPI_STATE_D0; - *state = result; out: @@ -304,24 +295,28 @@ int acpi_bus_set_power(acpi_handle handle, int state) } EXPORT_SYMBOL(acpi_bus_set_power); -int acpi_bus_init_power(struct acpi_device *device) +static int acpi_device_init_power(struct acpi_device *device) { int state; int result; - if (!device) - return -EINVAL; - - device->power.state = ACPI_STATE_UNKNOWN; - if (!acpi_device_is_present(device)) { - device->flags.initialized = false; - return -ENXIO; - } - result = acpi_device_get_power(device, &state); if (result) return result; + /* + * If the current power state of the device is D0 and it has a parent + * whose power state is not ignored, and the parent's power state + * initialization has failed, the parent's power state can be updated to + * D0 for consistency. + */ + if (!device->power.flags.ignore_parent && state == ACPI_STATE_D0) { + struct acpi_device *parent = acpi_dev_parent(device); + + if (parent && parent->power.state == ACPI_D_STATE_INVALID) + parent->power.state = ACPI_STATE_D0; + } + if (state < ACPI_STATE_D3_COLD && device->power.flags.power_resources) { /* Reference count the power resources. */ result = acpi_power_on_resources(device, state); @@ -351,9 +346,36 @@ int acpi_bus_init_power(struct acpi_device *device) state = ACPI_STATE_D0; } device->power.state = state; + + acpi_handle_debug(device->handle, "Initial power state: %s\n", + acpi_power_state_string(state)); + return 0; } +int acpi_bus_init_power(struct acpi_device *device) +{ + int result; + + if (device->power.state != ACPI_STATE_UNKNOWN) + return 0; + + /* + * The ACPI device power state can be only initialized once. If this + * fails, ACPI power management will not be used for the device going + * forward. + */ + result = acpi_device_init_power(device); + if (result) { + device->flags.power_manageable = 0; + device->power.state = ACPI_D_STATE_INVALID; + acpi_handle_info(device->handle, + "Initial power state undetermined, ACPI PM disabled\n"); + } + + return result; +} + /** * acpi_device_fix_up_power - Force device with missing _PSC into D0. * @device: Device object whose power state is to be fixed up. @@ -475,8 +497,13 @@ static int acpi_power_up_if_adr_present(struct acpi_device *adev, void *not_used if (!(adev->flags.power_manageable && adev->pnp.type.bus_address)) return 0; - acpi_handle_debug(adev->handle, "Power state: %s\n", - acpi_power_state_string(adev->power.state)); + /* + * This is done during the PCI root initialization which occurs before + * acpi_bus_attach() is called for the device, so the ACPI power state + * of the device needs to be initialized here. + */ + if (acpi_bus_init_power(adev)) + return 0; if (adev->power.state == ACPI_STATE_D3_COLD) return acpi_device_set_power(adev, ACPI_STATE_D0); diff --git a/drivers/acpi/fan.h b/drivers/acpi/fan.h index e20d6ad9df80..3faa247af514 100644 --- a/drivers/acpi/fan.h +++ b/drivers/acpi/fan.h @@ -52,7 +52,7 @@ struct acpi_fan_fst { }; struct acpi_fan { - acpi_handle handle; + struct acpi_device *adev; bool acpi4; bool has_fst; struct acpi_fan_fif fif; diff --git a/drivers/acpi/fan_core.c b/drivers/acpi/fan_core.c index 624d0736b581..3674f9613d40 100644 --- a/drivers/acpi/fan_core.c +++ b/drivers/acpi/fan_core.c @@ -54,8 +54,7 @@ MODULE_DEVICE_TABLE(acpi, fan_device_ids); static int fan_get_max_state(struct thermal_cooling_device *cdev, unsigned long *state) { - struct acpi_device *device = cdev->devdata; - struct acpi_fan *fan = acpi_driver_data(device); + struct acpi_fan *fan = cdev->devdata; if (fan->acpi4) { if (fan->fif.fine_grain_ctrl) @@ -72,42 +71,34 @@ static int fan_get_max_state(struct thermal_cooling_device *cdev, unsigned long int acpi_fan_get_fst(acpi_handle handle, struct acpi_fan_fst *fst) { struct acpi_buffer buffer = { ACPI_ALLOCATE_BUFFER, NULL }; - union acpi_object *obj; acpi_status status; - int ret = 0; status = acpi_evaluate_object(handle, "_FST", NULL, &buffer); if (ACPI_FAILURE(status)) - return -EIO; + return -ENXIO; - obj = buffer.pointer; + union acpi_object *obj __free(acpi_object_free) = buffer.pointer; if (!obj) return -ENODATA; - if (obj->type != ACPI_TYPE_PACKAGE || obj->package.count != 3) { - ret = -EPROTO; - goto err; - } + if (obj->type != ACPI_TYPE_PACKAGE || obj->package.count != 3) + return -EPROTO; if (obj->package.elements[0].type != ACPI_TYPE_INTEGER || obj->package.elements[1].type != ACPI_TYPE_INTEGER || - obj->package.elements[2].type != ACPI_TYPE_INTEGER) { - ret = -EPROTO; - goto err; - } + obj->package.elements[2].type != ACPI_TYPE_INTEGER) + return -EPROTO; fst->revision = obj->package.elements[0].integer.value; fst->control = obj->package.elements[1].integer.value; fst->speed = obj->package.elements[2].integer.value; -err: - kfree(obj); - return ret; + return 0; } -static int fan_get_state_acpi4(struct acpi_device *device, unsigned long *state) +static int fan_get_state_acpi4(struct acpi_fan *fan, unsigned long *state) { - struct acpi_fan *fan = acpi_driver_data(device); + struct acpi_device *device = fan->adev; struct acpi_fan_fst fst; int status, i; @@ -159,13 +150,12 @@ static int fan_get_state(struct acpi_device *device, unsigned long *state) static int fan_get_cur_state(struct thermal_cooling_device *cdev, unsigned long *state) { - struct acpi_device *device = cdev->devdata; - struct acpi_fan *fan = acpi_driver_data(device); + struct acpi_fan *fan = cdev->devdata; if (fan->acpi4) - return fan_get_state_acpi4(device, state); + return fan_get_state_acpi4(fan, state); else - return fan_get_state(device, state); + return fan_get_state(fan->adev, state); } static int fan_set_state(struct acpi_device *device, unsigned long state) @@ -177,9 +167,9 @@ static int fan_set_state(struct acpi_device *device, unsigned long state) state ? ACPI_STATE_D0 : ACPI_STATE_D3_COLD); } -static int fan_set_state_acpi4(struct acpi_device *device, unsigned long state) +static int fan_set_state_acpi4(struct acpi_fan *fan, unsigned long state) { - struct acpi_fan *fan = acpi_driver_data(device); + struct acpi_device *device = fan->adev; acpi_status status; u64 value = state; int max_state; @@ -213,13 +203,12 @@ static int fan_set_state_acpi4(struct acpi_device *device, unsigned long state) static int fan_set_cur_state(struct thermal_cooling_device *cdev, unsigned long state) { - struct acpi_device *device = cdev->devdata; - struct acpi_fan *fan = acpi_driver_data(device); + struct acpi_fan *fan = cdev->devdata; if (fan->acpi4) - return fan_set_state_acpi4(device, state); + return fan_set_state_acpi4(fan, state); else - return fan_set_state(device, state); + return fan_set_state(fan->adev, state); } static const struct thermal_cooling_device_ops fan_cooling_ops = { @@ -240,25 +229,22 @@ static int acpi_fan_get_fif(struct acpi_device *device) struct acpi_buffer format = { sizeof("NNNN"), "NNNN" }; u64 fields[4]; struct acpi_buffer fif = { sizeof(fields), fields }; - union acpi_object *obj; acpi_status status; status = acpi_evaluate_object(device->handle, "_FIF", NULL, &buffer); if (ACPI_FAILURE(status)) - return status; + return -ENXIO; - obj = buffer.pointer; + union acpi_object *obj __free(acpi_object_free) = buffer.pointer; if (!obj || obj->type != ACPI_TYPE_PACKAGE) { dev_err(&device->dev, "Invalid _FIF data\n"); - status = -EINVAL; - goto err; + return -ENODATA; } status = acpi_extract_package(obj, &format, &fif); if (ACPI_FAILURE(status)) { dev_err(&device->dev, "Invalid _FIF element\n"); - status = -EINVAL; - goto err; + return -ENODATA; } fan->fif.revision = fields[0]; @@ -272,9 +258,8 @@ static int acpi_fan_get_fif(struct acpi_device *device) /* If step size > 9, change to 9 (by spec valid values 1-9) */ else if (fan->fif.step_size > 9) fan->fif.step_size = 9; -err: - kfree(obj); - return status; + + return 0; } static int acpi_fan_speed_cmp(const void *a, const void *b) @@ -284,34 +269,28 @@ static int acpi_fan_speed_cmp(const void *a, const void *b) return fps1->speed - fps2->speed; } -static int acpi_fan_get_fps(struct acpi_device *device) +static int acpi_fan_get_fps(struct device *dev, struct acpi_device *device) { struct acpi_fan *fan = acpi_driver_data(device); struct acpi_buffer buffer = { ACPI_ALLOCATE_BUFFER, NULL }; - union acpi_object *obj; acpi_status status; int i; status = acpi_evaluate_object(device->handle, "_FPS", NULL, &buffer); if (ACPI_FAILURE(status)) - return status; + return -ENXIO; - obj = buffer.pointer; + union acpi_object *obj __free(acpi_object_free) = buffer.pointer; if (!obj || obj->type != ACPI_TYPE_PACKAGE || obj->package.count < 2) { dev_err(&device->dev, "Invalid _FPS data\n"); - status = -EINVAL; - goto err; + return -ENODATA; } fan->fps_count = obj->package.count - 1; /* minus revision field */ - fan->fps = devm_kcalloc(&device->dev, - fan->fps_count, sizeof(struct acpi_fan_fps), - GFP_KERNEL); - if (!fan->fps) { - dev_err(&device->dev, "Not enough memory\n"); - status = -ENOMEM; - goto err; - } + fan->fps = devm_kcalloc(dev, fan->fps_count, sizeof(*fan->fps), GFP_KERNEL); + if (!fan->fps) + return -ENOMEM; + for (i = 0; i < fan->fps_count; i++) { struct acpi_buffer format = { sizeof("NNNNN"), "NNNNN" }; struct acpi_buffer fps = { offsetof(struct acpi_fan_fps, name), @@ -320,7 +299,7 @@ static int acpi_fan_get_fps(struct acpi_device *device) &format, &fps); if (ACPI_FAILURE(status)) { dev_err(&device->dev, "Invalid _FPS element\n"); - goto err; + return -ENODATA; } } @@ -328,9 +307,7 @@ static int acpi_fan_get_fps(struct acpi_device *device) sort(fan->fps, fan->fps_count, sizeof(*fan->fps), acpi_fan_speed_cmp, NULL); -err: - kfree(obj); - return status; + return 0; } static int acpi_fan_dsm_init(struct device *dev) @@ -343,30 +320,28 @@ static int acpi_fan_dsm_init(struct device *dev) }, }; struct acpi_fan *fan = dev_get_drvdata(dev); - union acpi_object *obj; - int ret = 0; + acpi_handle fan_handle = fan->adev->handle; - if (!acpi_check_dsm(fan->handle, &acpi_fan_microsoft_guid, 0, + if (!acpi_check_dsm(fan_handle, &acpi_fan_microsoft_guid, 0, BIT(ACPI_FAN_DSM_GET_TRIP_POINT_GRANULARITY) | BIT(ACPI_FAN_DSM_SET_TRIP_POINTS))) return 0; dev_info(dev, "Using Microsoft fan extensions\n"); - obj = acpi_evaluate_dsm_typed(fan->handle, &acpi_fan_microsoft_guid, 0, - ACPI_FAN_DSM_GET_TRIP_POINT_GRANULARITY, &dummy, - ACPI_TYPE_INTEGER); + union acpi_object *obj __free(acpi_object_free) = + acpi_evaluate_dsm_typed(fan_handle, &acpi_fan_microsoft_guid, 0, + ACPI_FAN_DSM_GET_TRIP_POINT_GRANULARITY, + &dummy, ACPI_TYPE_INTEGER); if (!obj) - return -EIO; + return -ENXIO; if (obj->integer.value > U32_MAX) - ret = -EOVERFLOW; - else - fan->fan_trip_granularity = obj->integer.value; + return -EOVERFLOW; - kfree(obj); + fan->fan_trip_granularity = obj->integer.value; - return ret; + return 0; } static int acpi_fan_dsm_set_trip_points(struct device *dev, u64 upper, u64 lower) @@ -395,9 +370,9 @@ static int acpi_fan_dsm_set_trip_points(struct device *dev, u64 upper, u64 lower }; union acpi_object *obj; - obj = acpi_evaluate_dsm(fan->handle, &acpi_fan_microsoft_guid, 0, - ACPI_FAN_DSM_SET_TRIP_POINTS, &in); - kfree(obj); + obj = acpi_evaluate_dsm(fan->adev->handle, &acpi_fan_microsoft_guid, + 0, ACPI_FAN_DSM_SET_TRIP_POINTS, &in); + ACPI_FREE(obj); return 0; } @@ -506,7 +481,7 @@ static int acpi_fan_probe(struct platform_device *pdev) return -ENOMEM; } - fan->handle = device->handle; + fan->adev = device; device->driver_data = fan; platform_set_drvdata(pdev, fan); @@ -522,7 +497,7 @@ static int acpi_fan_probe(struct platform_device *pdev) if (result) return result; - result = acpi_fan_get_fps(device); + result = acpi_fan_get_fps(&pdev->dev, device); if (result) return result; } @@ -567,8 +542,7 @@ static int acpi_fan_probe(struct platform_device *pdev) else name = acpi_device_bid(device); - cdev = thermal_cooling_device_register(name, device, - &fan_cooling_ops); + cdev = thermal_cooling_device_create(&pdev->dev, name, fan, &fan_cooling_ops); if (IS_ERR(cdev)) { result = PTR_ERR(cdev); goto err_end; @@ -577,28 +551,9 @@ static int acpi_fan_probe(struct platform_device *pdev) dev_dbg(&pdev->dev, "registered as cooling_device%d\n", cdev->id); fan->cdev = cdev; - result = sysfs_create_link(&pdev->dev.kobj, - &cdev->device.kobj, - "thermal_cooling"); - if (result) { - dev_err(&pdev->dev, "Failed to create sysfs link 'thermal_cooling'\n"); - goto err_unregister; - } - - result = sysfs_create_link(&cdev->device.kobj, - &pdev->dev.kobj, - "device"); - if (result) { - dev_err(&pdev->dev, "Failed to create sysfs link 'device'\n"); - goto err_remove_link; - } return 0; -err_remove_link: - sysfs_remove_link(&pdev->dev.kobj, "thermal_cooling"); -err_unregister: - thermal_cooling_device_unregister(cdev); err_end: if (fan->has_fst) acpi_fan_delete_attributes(device); @@ -615,8 +570,6 @@ static void acpi_fan_remove(struct platform_device *pdev) acpi_fan_delete_attributes(device); } - sysfs_remove_link(&pdev->dev.kobj, "thermal_cooling"); - sysfs_remove_link(&fan->cdev->device.kobj, "device"); thermal_cooling_device_unregister(fan->cdev); } diff --git a/drivers/acpi/fan_hwmon.c b/drivers/acpi/fan_hwmon.c index d3374f8f524b..c5d8419f41c5 100644 --- a/drivers/acpi/fan_hwmon.c +++ b/drivers/acpi/fan_hwmon.c @@ -94,7 +94,7 @@ static int acpi_fan_hwmon_read(struct device *dev, enum hwmon_sensor_types type, struct acpi_fan_fst fst; int ret; - ret = acpi_fan_get_fst(fan->handle, &fst); + ret = acpi_fan_get_fst(fan->adev->handle, &fst); if (ret < 0) return ret; diff --git a/drivers/acpi/glue.c b/drivers/acpi/glue.c index b1776809279d..0959f58cb77c 100644 --- a/drivers/acpi/glue.c +++ b/drivers/acpi/glue.c @@ -59,19 +59,37 @@ int unregister_acpi_bus_type(struct acpi_bus_type *type) } EXPORT_SYMBOL_GPL(unregister_acpi_bus_type); -static struct acpi_bus_type *acpi_get_bus_type(struct device *dev) +static struct acpi_device *acpi_companion_lookup(struct device *dev) { - struct acpi_bus_type *tmp, *ret = NULL; + struct acpi_bus_type *type; - down_read(&bus_type_sem); - list_for_each_entry(tmp, &bus_type_list, list) { - if (tmp->match(dev)) { - ret = tmp; - break; + if (!dev->type) + return NULL; + + guard(rwsem_read)(&bus_type_sem); + + list_for_each_entry(type, &bus_type_list, list) { + struct acpi_device *adev; + + if (!type->match(dev)) + continue; + + adev = type->find_companion(dev); + if (!adev) { + dev_dbg(dev, "ACPI companion not found\n"); + return NULL; } + if (acpi_bind_one(dev, adev)) { + dev_dbg(dev, "Binding to ACPI companion failed\n"); + return NULL; + } + if (type->setup) + type->setup(dev); + + return adev; } - up_read(&bus_type_sem); - return ret; + + return NULL; } #define FIND_CHILD_MIN_SCORE 1 @@ -228,31 +246,25 @@ static void acpi_physnode_link_name(char *buf, unsigned int node_id) int acpi_bind_one(struct device *dev, struct acpi_device *acpi_dev) { struct acpi_device_physical_node *physical_node, *pn; + struct acpi_device *comp_dev = ACPI_COMPANION(dev); char physical_node_name[PHYSICAL_NODE_NAME_SIZE]; struct list_head *physnode_list; unsigned int node_id; int retval = -EINVAL; - if (has_acpi_companion(dev)) { - if (acpi_dev) { - dev_warn(dev, "ACPI companion already set\n"); + if (!acpi_dev) { + if (!comp_dev) return -EINVAL; - } else { - acpi_dev = ACPI_COMPANION(dev); - } - } - if (!acpi_dev) - return -EINVAL; - acpi_dev_get(acpi_dev); - get_device(dev); - physical_node = kzalloc_obj(*physical_node); - if (!physical_node) { - retval = -ENOMEM; - goto err; + /* If the companion has been set upfront, pick it up. */ + acpi_dev = comp_dev; + } else if (comp_dev && acpi_dev != comp_dev) { + dev_warn(dev, "ACPI companion already set to %s which is not %s\n", + acpi_dev_name(comp_dev), acpi_dev_name(acpi_dev)); + return -EEXIST; } - mutex_lock(&acpi_dev->physical_node_lock); + guard(mutex)(&acpi_dev->physical_node_lock); /* * Keep the list sorted by node_id so that the IDs of removed nodes can @@ -263,15 +275,12 @@ int acpi_bind_one(struct device *dev, struct acpi_device *acpi_dev) list_for_each_entry(pn, &acpi_dev->physical_node_list, node) { /* Sanity check. */ if (pn->dev == dev) { - mutex_unlock(&acpi_dev->physical_node_lock); - - dev_warn(dev, "Already associated with ACPI node\n"); - kfree(physical_node); - if (ACPI_COMPANION(dev) != acpi_dev) - goto err; - - put_device(dev); - acpi_dev_put(acpi_dev); + if (!comp_dev) { + /* Really unexpected. */ + ACPI_COMPANION_SET(dev, acpi_dev); + dev_warn(&acpi_dev->dev, + "Physical device list corruption fixed up\n"); + } return 0; } if (pn->node_id == node_id) { @@ -280,12 +289,19 @@ int acpi_bind_one(struct device *dev, struct acpi_device *acpi_dev) } } + physical_node = kzalloc_obj(*physical_node); + if (!physical_node) + return -ENOMEM; + + acpi_dev_get(acpi_dev); + get_device(dev); + physical_node->node_id = node_id; physical_node->dev = dev; list_add(&physical_node->node, physnode_list); acpi_dev->physical_node_count++; - if (!has_acpi_companion(dev)) + if (!comp_dev) ACPI_COMPANION_SET(dev, acpi_dev); acpi_physnode_link_name(physical_node_name, node_id); @@ -301,28 +317,20 @@ int acpi_bind_one(struct device *dev, struct acpi_device *acpi_dev) dev_err(dev, "Failed to create link firmware_node (%d)\n", retval); - mutex_unlock(&acpi_dev->physical_node_lock); - if (acpi_dev->wakeup.flags.valid) device_set_wakeup_capable(dev, true); return 0; - - err: - ACPI_COMPANION_SET(dev, NULL); - put_device(dev); - acpi_dev_put(acpi_dev); - return retval; } EXPORT_SYMBOL_GPL(acpi_bind_one); -int acpi_unbind_one(struct device *dev) +void acpi_unbind_one(struct device *dev) { struct acpi_device *acpi_dev = ACPI_COMPANION(dev); struct acpi_device_physical_node *entry; if (!acpi_dev) - return 0; + return; mutex_lock(&acpi_dev->physical_node_lock); @@ -337,15 +345,17 @@ int acpi_unbind_one(struct device *dev) sysfs_remove_link(&acpi_dev->dev.kobj, physnode_name); sysfs_remove_link(&dev->kobj, "firmware_node"); ACPI_COMPANION_SET(dev, NULL); + + mutex_unlock(&acpi_dev->physical_node_lock); + /* Drop references taken by acpi_bind_one(). */ put_device(dev); acpi_dev_put(acpi_dev); kfree(entry); - break; + return; } mutex_unlock(&acpi_dev->physical_node_lock); - return 0; } EXPORT_SYMBOL_GPL(acpi_unbind_one); @@ -354,48 +364,29 @@ void acpi_device_notify(struct device *dev) struct acpi_device *adev; int ret; + /* ACPI devices have no ACPI companions. */ + if (dev->bus == &acpi_bus_type) + return; + ret = acpi_bind_one(dev, NULL); if (ret) { - struct acpi_bus_type *type = acpi_get_bus_type(dev); - - if (!type) - goto err; - - adev = type->find_companion(dev); - if (!adev) { - dev_dbg(dev, "ACPI companion not found\n"); - goto err; - } - ret = acpi_bind_one(dev, adev); - if (ret) - goto err; - - if (type->setup) { - type->setup(dev); - goto done; - } + adev = acpi_companion_lookup(dev); + if (!adev) + return; } else { adev = ACPI_COMPANION(dev); if (dev_is_pci(dev)) { pci_acpi_setup(dev, adev); - goto done; } else if (dev_is_platform(dev)) { acpi_configure_pmsi_domain(dev); + + if (adev->handler && adev->handler->bind) + adev->handler->bind(dev); } } - if (adev->handler && adev->handler->bind) - adev->handler->bind(dev); - -done: - acpi_handle_debug(ACPI_HANDLE(dev), "Bound to device %s\n", - dev_name(dev)); - - return; - -err: - dev_dbg(dev, "No ACPI support\n"); + dev_dbg(dev, "Bound to ACPI device %s\n", acpi_dev_name(adev)); } void acpi_device_notify_remove(struct device *dev) diff --git a/drivers/acpi/internal.h b/drivers/acpi/internal.h index 40f875b265a9..011da4ed6a80 100644 --- a/drivers/acpi/internal.h +++ b/drivers/acpi/internal.h @@ -157,6 +157,7 @@ void acpi_turn_off_unused_power_resources(void); Device Power Management -------------------------------------------------------------------------- */ int acpi_device_get_power(struct acpi_device *device, int *state); +int acpi_bus_init_power(struct acpi_device *device); int acpi_wakeup_device_init(void); /* -------------------------------------------------------------------------- diff --git a/drivers/acpi/numa/hmat.c b/drivers/acpi/numa/hmat.c index 9792dc394756..3006789ae31f 100644 --- a/drivers/acpi/numa/hmat.c +++ b/drivers/acpi/numa/hmat.c @@ -995,7 +995,7 @@ static int hmat_calculate_adistance(struct notifier_block *self, return NOTIFY_STOP; } -static struct notifier_block hmat_adist_nb __meminitdata = { +static struct notifier_block hmat_adist_nb = { .notifier_call = hmat_calculate_adistance, .priority = 100, }; diff --git a/drivers/acpi/numa/srat.c b/drivers/acpi/numa/srat.c index 5c407dc6401e..f5616830f5a9 100644 --- a/drivers/acpi/numa/srat.c +++ b/drivers/acpi/numa/srat.c @@ -111,7 +111,7 @@ int __init fix_pxm_node_maps(int max_nid) for (j = 0; j <= max_nid; j++) { if ((emu_nid_to_phys[j] == i) && WARN(node_to_pxm_map_copy[j] != PXM_INVAL, - "Node %d is already binded to PXM %d\n", + "Node %d is already bound to PXM %d\n", j, node_to_pxm_map_copy[j])) return -1; if (emu_nid_to_phys[j] == i) { diff --git a/drivers/acpi/osl.c b/drivers/acpi/osl.c index ed2162a97bd2..5b7aefb48eeb 100644 --- a/drivers/acpi/osl.c +++ b/drivers/acpi/osl.c @@ -159,7 +159,7 @@ void __printf(1, 0) acpi_os_vprintf(const char *fmt, va_list args) { static char buffer[512]; - vsprintf(buffer, fmt, args); + vsnprintf(buffer, sizeof(buffer), fmt, args); #ifdef ENABLE_DEBUGGER if (acpi_in_debugger) { diff --git a/drivers/acpi/pfr_update.c b/drivers/acpi/pfr_update.c index 9afd2c52fdbd..98ace679601b 100644 --- a/drivers/acpi/pfr_update.c +++ b/drivers/acpi/pfr_update.c @@ -422,7 +422,7 @@ free_acpi_buffer: static long pfru_ioctl(struct file *file, unsigned int cmd, unsigned long arg) { - struct pfru_update_cap_info cap_hdr; + struct pfru_update_cap_info cap_hdr = {}; struct pfru_device *pfru_dev = to_pfru_dev(file); void __user *p = (void __user *)arg; u32 rev; diff --git a/drivers/acpi/power.c b/drivers/acpi/power.c index 23a4e207a01e..c6955d1ad29d 100644 --- a/drivers/acpi/power.c +++ b/drivers/acpi/power.c @@ -1121,6 +1121,19 @@ static const struct dmi_system_id dmi_leave_unused_power_resources_on[] = { } }, + { + /* + * Acer Swift 3 SF314-56G with NVIDIA MX250 suffers from the + * same ACPI power resource regression as the Thunderobot ZERO. + * The power resource controlling the dGPU is incorrectly turned off + * during initialization, causing the GPU to fall off the bus (D3cold). + */ + .matches = { + DMI_MATCH(DMI_SYS_VENDOR, "Acer"), + DMI_MATCH(DMI_PRODUCT_NAME, "Swift SF314-56G"), + } + + }, {} }; diff --git a/drivers/acpi/processor_driver.c b/drivers/acpi/processor_driver.c index cdc2ae1632b2..bdd1529d79f0 100644 --- a/drivers/acpi/processor_driver.c +++ b/drivers/acpi/processor_driver.c @@ -164,7 +164,7 @@ static int __acpi_processor_start(struct acpi_device *device) acpi_pss_perf_init(pr); - result = acpi_processor_thermal_init(pr, device); + result = acpi_processor_thermal_init(pr); if (result) goto err_power_exit; @@ -179,7 +179,7 @@ static int __acpi_processor_start(struct acpi_device *device) return 0; err_thermal_exit: - acpi_processor_thermal_exit(pr, device); + acpi_processor_thermal_exit(pr); err_power_exit: acpi_processor_power_exit(pr); return result; @@ -203,7 +203,7 @@ static int acpi_processor_stop(struct device *dev) acpi_cppc_processor_exit(pr); - acpi_processor_thermal_exit(pr, device); + acpi_processor_thermal_exit(pr); return 0; } diff --git a/drivers/acpi/processor_thermal.c b/drivers/acpi/processor_thermal.c index c7b1dc5687ec..11036f0d2b94 100644 --- a/drivers/acpi/processor_thermal.c +++ b/drivers/acpi/processor_thermal.c @@ -235,15 +235,7 @@ static int processor_get_max_state(struct thermal_cooling_device *cdev, unsigned long *state) { - struct acpi_device *device = cdev->devdata; - struct acpi_processor *pr; - - if (!device) - return -EINVAL; - - pr = acpi_driver_data(device); - if (!pr) - return -EINVAL; + struct acpi_processor *pr = cdev->devdata; *state = acpi_processor_max_state(pr); return 0; @@ -253,15 +245,7 @@ static int processor_get_cur_state(struct thermal_cooling_device *cdev, unsigned long *cur_state) { - struct acpi_device *device = cdev->devdata; - struct acpi_processor *pr; - - if (!device) - return -EINVAL; - - pr = acpi_driver_data(device); - if (!pr) - return -EINVAL; + struct acpi_processor *pr = cdev->devdata; *cur_state = cpufreq_get_cur_state(pr->id); if (pr->flags.throttling) @@ -273,18 +257,10 @@ static int processor_set_cur_state(struct thermal_cooling_device *cdev, unsigned long state) { - struct acpi_device *device = cdev->devdata; - struct acpi_processor *pr; + struct acpi_processor *pr = cdev->devdata; int result = 0; int max_pstate; - if (!device) - return -EINVAL; - - pr = acpi_driver_data(device); - if (!pr) - return -EINVAL; - max_pstate = cpufreq_get_max_state(pr->id); if (state > acpi_processor_max_state(pr)) @@ -308,55 +284,21 @@ const struct thermal_cooling_device_ops processor_cooling_ops = { .set_cur_state = processor_set_cur_state, }; -int acpi_processor_thermal_init(struct acpi_processor *pr, - struct acpi_device *device) +int acpi_processor_thermal_init(struct acpi_processor *pr) { - int result = 0; - - pr->cdev = thermal_cooling_device_register("Processor", device, - &processor_cooling_ops); - if (IS_ERR(pr->cdev)) { - result = PTR_ERR(pr->cdev); - return result; - } - - dev_dbg(&device->dev, "registered as cooling_device%d\n", - pr->cdev->id); + pr->cdev = thermal_cooling_device_create(pr->dev, "Processor", pr, + &processor_cooling_ops); + if (IS_ERR(pr->cdev)) + return PTR_ERR(pr->cdev); - result = sysfs_create_link(&device->dev.kobj, - &pr->cdev->device.kobj, - "thermal_cooling"); - if (result) { - dev_err(&device->dev, - "Failed to create sysfs link 'thermal_cooling'\n"); - goto err_thermal_unregister; - } - - result = sysfs_create_link(&pr->cdev->device.kobj, - &device->dev.kobj, - "device"); - if (result) { - dev_err(&pr->cdev->device, - "Failed to create sysfs link 'device'\n"); - goto err_remove_sysfs_thermal; - } + dev_dbg(pr->dev, "registered as cooling_device%d\n", pr->cdev->id); return 0; - -err_remove_sysfs_thermal: - sysfs_remove_link(&device->dev.kobj, "thermal_cooling"); -err_thermal_unregister: - thermal_cooling_device_unregister(pr->cdev); - - return result; } -void acpi_processor_thermal_exit(struct acpi_processor *pr, - struct acpi_device *device) +void acpi_processor_thermal_exit(struct acpi_processor *pr) { - if (pr->cdev) { - sysfs_remove_link(&device->dev.kobj, "thermal_cooling"); - sysfs_remove_link(&pr->cdev->device.kobj, "device"); + if (!IS_ERR_OR_NULL(pr->cdev)) { thermal_cooling_device_unregister(pr->cdev); pr->cdev = NULL; } diff --git a/drivers/acpi/riscv/cppc.c b/drivers/acpi/riscv/cppc.c index 42c1a9052470..4ea4ddedd91f 100644 --- a/drivers/acpi/riscv/cppc.c +++ b/drivers/acpi/riscv/cppc.c @@ -97,6 +97,7 @@ bool cpc_ffh_supported(void) int cpc_read_ffh(int cpu, struct cpc_reg *reg, u64 *val) { struct sbi_cppc_data data; + int ret; if (WARN_ON_ONCE(irqs_disabled())) return -EPERM; @@ -107,19 +108,27 @@ int cpc_read_ffh(int cpu, struct cpc_reg *reg, u64 *val) data.reg = FFH_CPPC_SBI_REG(reg->address); - smp_call_function_single(cpu, sbi_cppc_read, &data, 1); + ret = smp_call_function_single(cpu, sbi_cppc_read, &data, 1); + if (ret) + return ret; + if (data.ret.error) + return sbi_err_map_linux_errno(data.ret.error); *val = data.ret.value; - return (data.ret.error) ? sbi_err_map_linux_errno(data.ret.error) : 0; + return 0; } else if (FFH_CPPC_TYPE(reg->address) == FFH_CPPC_CSR) { data.reg = FFH_CPPC_CSR_NUM(reg->address); - smp_call_function_single(cpu, cppc_ffh_csr_read, &data, 1); + ret = smp_call_function_single(cpu, cppc_ffh_csr_read, &data, 1); + if (ret) + return ret; + if (data.ret.error) + return data.ret.error; *val = data.ret.value; - return data.ret.error; + return 0; } return -EINVAL; @@ -128,6 +137,7 @@ int cpc_read_ffh(int cpu, struct cpc_reg *reg, u64 *val) int cpc_write_ffh(int cpu, struct cpc_reg *reg, u64 val) { struct sbi_cppc_data data; + int ret; if (WARN_ON_ONCE(irqs_disabled())) return -EPERM; @@ -139,14 +149,18 @@ int cpc_write_ffh(int cpu, struct cpc_reg *reg, u64 val) data.reg = FFH_CPPC_SBI_REG(reg->address); data.val = val; - smp_call_function_single(cpu, sbi_cppc_write, &data, 1); + ret = smp_call_function_single(cpu, sbi_cppc_write, &data, 1); + if (ret) + return ret; return (data.ret.error) ? sbi_err_map_linux_errno(data.ret.error) : 0; } else if (FFH_CPPC_TYPE(reg->address) == FFH_CPPC_CSR) { data.reg = FFH_CPPC_CSR_NUM(reg->address); data.val = val; - smp_call_function_single(cpu, cppc_ffh_csr_write, &data, 1); + ret = smp_call_function_single(cpu, cppc_ffh_csr_write, &data, 1); + if (ret) + return ret; return data.ret.error; } diff --git a/drivers/acpi/sbs.c b/drivers/acpi/sbs.c index 86b7c7975852..f10bbf13c242 100644 --- a/drivers/acpi/sbs.c +++ b/drivers/acpi/sbs.c @@ -318,7 +318,7 @@ static struct acpi_battery_reader state_readers[] = { {0x0a, SMBUS_READ_WORD, offsetof(struct acpi_battery, rate_now)}, {0x0b, SMBUS_READ_WORD, offsetof(struct acpi_battery, rate_avg)}, {0x0f, SMBUS_READ_WORD, offsetof(struct acpi_battery, capacity_now)}, - {0x0e, SMBUS_READ_WORD, offsetof(struct acpi_battery, state_of_charge)}, + {0x0d, SMBUS_READ_WORD, offsetof(struct acpi_battery, state_of_charge)}, {0x16, SMBUS_READ_WORD, offsetof(struct acpi_battery, state)}, }; @@ -639,10 +639,8 @@ static int acpi_sbs_probe(struct platform_device *pdev) return -ENODEV; sbs = kzalloc_obj(struct acpi_sbs); - if (!sbs) { - result = -ENOMEM; - goto end; - } + if (!sbs) + return -ENOMEM; mutex_init(&sbs->lock); diff --git a/drivers/acpi/scan.c b/drivers/acpi/scan.c index 163a3cccf197..30bfef67cbaa 100644 --- a/drivers/acpi/scan.c +++ b/drivers/acpi/scan.c @@ -20,6 +20,7 @@ #include <linux/kthread.h> #include <linux/dmi.h> #include <linux/dma-map-ops.h> +#include <linux/pci.h> #include <linux/platform_data/x86/apple.h> #include <linux/pgtable.h> #include <linux/crc32.h> @@ -637,10 +638,9 @@ static struct acpi_device *handle_to_device(acpi_handle handle, status = acpi_get_data_full(handle, acpi_scan_drop_device, (void **)&adev, callback); - if (ACPI_FAILURE(status) || !adev) { - acpi_handle_debug(handle, "No context!\n"); + if (ACPI_FAILURE(status) || !adev) return NULL; - } + return adev; } @@ -1143,8 +1143,7 @@ static void acpi_bus_get_power_flags(struct acpi_device *device) device->power.states[ACPI_STATE_D3_COLD].flags.valid = 1; } - if (acpi_bus_init_power(device)) - device->flags.power_manageable = 0; + device->power.state = ACPI_STATE_UNKNOWN; } static void acpi_bus_get_flags(struct acpi_device *device) @@ -2338,50 +2337,57 @@ static int acpi_scan_attach_handler(struct acpi_device *device) static int acpi_bus_attach(struct acpi_device *device, void *first_pass) { bool skip = !first_pass && device->flags.visited; + struct pci_dev *pci; acpi_handle ejd; int ret; if (skip) goto ok; + device->flags.initialized = true; + if (ACPI_SUCCESS(acpi_bus_get_ejd(device->handle, &ejd))) register_dock_dependent_device(device, ejd); acpi_bus_get_status(device); - /* Skip devices that are not ready for enumeration (e.g. not present) */ - if (!acpi_dev_ready_for_enumeration(device)) { - device->flags.initialized = false; + /* + * If the given ACPI device object has been already associated with a + * PCI device found on the bus, its status is effectively "present and + * enabled", and dependencies are not tracked for PCI devices, so it is + * not necessary or even useful to check the device's readiness in that + * case. + */ + pci = acpi_dev_get_pci_dev(device); + if (pci) { + acpi_handle_debug(device->handle, "PCI companion %s found\n", + pci_name(pci)); + + if (!acpi_device_is_present(device)) + pci_info(pci, FW_BUG "ACPI status differs from reality\n"); + + pci_dev_put(pci); + } else if (!acpi_dev_ready_for_enumeration(device)) { + /* The device is not ready (e.g. not present), so skip it. */ acpi_device_clear_enumerated(device); - device->flags.power_manageable = 0; return 0; } + if (device->handler) goto ok; acpi_ec_register_opregions(device); - if (!device->flags.initialized) { - device->flags.power_manageable = - device->power.states[ACPI_STATE_D0].flags.valid; - if (acpi_bus_init_power(device)) - device->flags.power_manageable = 0; + acpi_bus_init_power(device); - device->flags.initialized = true; - } else if (device->flags.visited) { + if (device->flags.visited) goto ok; - } ret = acpi_scan_attach_handler(device); if (ret < 0) return 0; - if (ret > 0 && !device->flags.enumeration_by_parent) { - acpi_device_set_enumerated(device); - goto ok; - } - - if (device->pnp.type.platform_id || device->pnp.type.backlight || - device->flags.enumeration_by_parent) + if (device->flags.enumeration_by_parent || + (!ret && (device->pnp.type.platform_id || device->pnp.type.backlight))) acpi_default_enumeration(device); else acpi_device_set_enumerated(device); diff --git a/drivers/acpi/sysfs.c b/drivers/acpi/sysfs.c index dd4a99f09efe..a99373547126 100644 --- a/drivers/acpi/sysfs.c +++ b/drivers/acpi/sysfs.c @@ -11,6 +11,7 @@ #include <linux/kernel.h> #include <linux/kstrtox.h> #include <linux/moduleparam.h> +#include <linux/string.h> #include "internal.h" @@ -181,10 +182,10 @@ static int param_set_trace_method_name(const char *val, /* This is a hack. We can't kmalloc in early boot. */ if (is_abs_path) - strcpy(trace_method_name, val); + strscpy(trace_method_name, val, sizeof(trace_method_name)); else { trace_method_name[0] = '\\'; - strcpy(trace_method_name+1, val); + strscpy(trace_method_name + 1, val, sizeof(trace_method_name) - 1); } /* Restore the original tracer state */ diff --git a/drivers/acpi/tables.c b/drivers/acpi/tables.c index eb93c060426f..aaa1de79bccf 100644 --- a/drivers/acpi/tables.c +++ b/drivers/acpi/tables.c @@ -560,6 +560,9 @@ acpi_table_initrd_override(struct acpi_table_header *existing_table, while (table_offset + ACPI_HEADER_SIZE <= all_tables_size) { table = acpi_os_map_memory(acpi_tables_addr + table_offset, ACPI_HEADER_SIZE); + if (!table) + return AE_NO_MEMORY; + if (table_offset + table->length > all_tables_size) { acpi_os_unmap_memory(table, ACPI_HEADER_SIZE); WARN_ON(1); diff --git a/drivers/acpi/thermal.c b/drivers/acpi/thermal.c index dd7666c176a0..13e9a3a79d92 100644 --- a/drivers/acpi/thermal.c +++ b/drivers/acpi/thermal.c @@ -564,17 +564,20 @@ static bool acpi_thermal_should_bind_cdev(struct thermal_zone_device *thermal, struct cooling_spec *c) { struct acpi_thermal_trip *acpi_trip = trip->priv; - struct acpi_device *cdev_adev = cdev->devdata; + struct device *parent = cdev->device.parent; + acpi_handle parent_handle; int i; - /* Skip critical and hot trips. */ - if (!acpi_trip) + /* Skip critical and hot trips and parentless cooling devices. */ + if (!acpi_trip || !parent) return false; - for (i = 0; i < acpi_trip->devices.count; i++) { - acpi_handle handle = acpi_trip->devices.handles[i]; + parent_handle = ACPI_HANDLE(parent); + if (!parent_handle) + return false; - if (acpi_fetch_acpi_dev(handle) == cdev_adev) + for (i = 0; i < acpi_trip->devices.count; i++) { + if (acpi_trip->devices.handles[i] == parent_handle) return true; } @@ -738,7 +741,9 @@ static void acpi_thermal_aml_dependency_fix(struct acpi_thermal *tz) */ static void acpi_thermal_guess_offset(struct acpi_thermal *tz, long crit_temp) { - if (crit_temp != THERMAL_TEMP_INVALID && crit_temp % 5 == 1) + if (crit_temp != THERMAL_TEMP_INVALID && crit_temp % 5 == 0) + tz->kelvin_offset = 273000; + else if (crit_temp != THERMAL_TEMP_INVALID && crit_temp % 5 == 1) tz->kelvin_offset = 273100; else tz->kelvin_offset = 273200; diff --git a/drivers/acpi/utils.c b/drivers/acpi/utils.c index d499b72574ab..7c584d490d6a 100644 --- a/drivers/acpi/utils.c +++ b/drivers/acpi/utils.c @@ -1069,19 +1069,26 @@ EXPORT_SYMBOL(acpi_dev_is_video_device); * @plat: pointer to acpi_platform_list table terminated by a NULL entry * * Return the matched index if the system is found in the platform list. - * Otherwise, return a negative error code. + * Return -ENODEV for no match, or another negative error code if a table + * header could not be read and no entry matched. */ int acpi_match_platform_list(const struct acpi_platform_list *plat) { struct acpi_table_header hdr; + acpi_status status; + int ret = -ENODEV; int idx = 0; if (acpi_disabled) return -ENODEV; for (; plat->oem_id[0]; plat++, idx++) { - if (ACPI_FAILURE(acpi_get_table_header(plat->table, 0, &hdr))) + status = acpi_get_table_header(plat->table, 0, &hdr); + if (ACPI_FAILURE(status)) { + if (status != AE_NOT_FOUND) + ret = status == AE_NO_MEMORY ? -ENOMEM : -EIO; continue; + } if (strncmp(plat->oem_id, hdr.oem_id, ACPI_OEM_ID_SIZE)) continue; @@ -1096,6 +1103,6 @@ int acpi_match_platform_list(const struct acpi_platform_list *plat) return idx; } - return -ENODEV; + return ret; } EXPORT_SYMBOL(acpi_match_platform_list); diff --git a/drivers/acpi/video_detect.c b/drivers/acpi/video_detect.c index e3b69f876d24..3037fd18d420 100644 --- a/drivers/acpi/video_detect.c +++ b/drivers/acpi/video_detect.c @@ -929,6 +929,14 @@ static const struct dmi_system_id video_detect_dmi_table[] = { DMI_MATCH(DMI_PRODUCT_NAME, "Nitro AN515-46"), }, }, + { + .callback = video_detect_force_native, + /* Acer Nitro AN515-58 */ + .matches = { + DMI_MATCH(DMI_SYS_VENDOR, "Acer"), + DMI_MATCH(DMI_PRODUCT_NAME, "Nitro AN515-58"), + }, + }, /* * x86 android tablets which directly control the backlight through diff --git a/drivers/acpi/x86/s2idle.c b/drivers/acpi/x86/s2idle.c index b6b1dd76a06b..c278ea14c0b3 100644 --- a/drivers/acpi/x86/s2idle.c +++ b/drivers/acpi/x86/s2idle.c @@ -24,6 +24,34 @@ #ifdef CONFIG_SUSPEND +static void acpi_setup_ixx_gpes(void) +{ + acpi_handle gpe_root; + unsigned int i; + char gpe_nr_str[5]; + + if (ACPI_FAILURE(acpi_get_handle(NULL, "\\_GPE", &gpe_root))) + return; + + for (i = 0; i <= 0xff; i++) { + scnprintf(gpe_nr_str, sizeof(gpe_nr_str), "_I%02X", i); + if (!acpi_has_method(gpe_root, gpe_nr_str)) + continue; + + /* + * Enable the GPE if it has a handler method because marking it + * as wake-capable causes acpi_update_all_gpes() to skip it. + */ + if (ACPI_FAILURE(acpi_enable_gpe_cond(NULL, i, ACPI_GPE_DISPATCH_METHOD))) + continue; + + acpi_mark_gpe_for_wake(NULL, i); + acpi_set_gpe_wake_mask(NULL, i, ACPI_GPE_ENABLE); + + pm_pr_dbg("ACPI: GPE0x%02x armed for wake via %s\n", i, gpe_nr_str); + } +} + static bool sleep_no_lps0 __read_mostly; module_param(sleep_no_lps0, bool, 0644); MODULE_PARM_DESC(sleep_no_lps0, "Do not use the special LPS0 device interface"); @@ -649,6 +677,7 @@ void __init acpi_s2idle_setup(void) { acpi_scan_add_handler(&lps0_handler); s2idle_set_ops(&acpi_s2idle_ops_lps0); + acpi_setup_ixx_gpes(); } int acpi_register_lps0_dev(struct acpi_s2idle_dev_ops *arg) diff --git a/drivers/base/base.h b/drivers/base/base.h index a5b7abc10ff0..21c7d244f8aa 100644 --- a/drivers/base/base.h +++ b/drivers/base/base.h @@ -27,6 +27,7 @@ * @drivers_autoprobe: gate whether new devices are automatically attached to * registered drivers, or new drivers automatically attach * to existing devices. + * @no_drivers: gate whether drivers can be registered. * @bus: pointer back to the struct bus_type that this structure is associated * with. * @dev_root: Default device to use as the parent. @@ -51,6 +52,7 @@ struct subsys_private { struct klist klist_drivers; struct blocking_notifier_head bus_notifier; unsigned int drivers_autoprobe:1; + unsigned int no_drivers:1; const struct bus_type *bus; struct device *dev_root; diff --git a/drivers/base/bus.c b/drivers/base/bus.c index d17bd91490ee..d89333806a4e 100644 --- a/drivers/base/bus.c +++ b/drivers/base/bus.c @@ -738,6 +738,11 @@ int bus_add_driver(struct device_driver *drv) if (!sp) return -EINVAL; + if (sp->no_drivers) { + error = -ENXIO; + goto out_put_bus; + } + /* * Reference in sp is now incremented and will be dropped when * the driver is removed from the bus @@ -930,15 +935,7 @@ static ssize_t bus_uevent_store(const struct bus_type *bus, static struct bus_attribute bus_attr_uevent = __ATTR(uevent, 0200, NULL, bus_uevent_store); -/** - * bus_register - register a driver-core subsystem - * @bus: bus to register - * - * Once we have that, we register the bus with the kobject - * infrastructure, then register the children subsystems it has: - * the devices and drivers that belong to the subsystem. - */ -int bus_register(const struct bus_type *bus) +static int bus_register_internal(const struct bus_type *bus, bool use_drivers) { int retval; struct subsys_private *priv; @@ -960,7 +957,8 @@ int bus_register(const struct bus_type *bus) bus_kobj->kset = bus_kset; bus_kobj->ktype = &bus_ktype; - priv->drivers_autoprobe = 1; + priv->drivers_autoprobe = use_drivers; + priv->no_drivers = !use_drivers; retval = kset_register(&priv->subsys); if (retval) @@ -976,6 +974,10 @@ int bus_register(const struct bus_type *bus) goto bus_devices_fail; } + /* + * Create the drivers kset even if use_drivers is unset to prevent some + * older versions of systemd from failing. + */ priv->drivers_kset = kset_create_and_add("drivers", NULL, bus_kobj); if (!priv->drivers_kset) { retval = -ENOMEM; @@ -989,9 +991,11 @@ int bus_register(const struct bus_type *bus) klist_init(&priv->klist_devices, klist_devices_get, klist_devices_put); klist_init(&priv->klist_drivers, NULL, NULL); - retval = add_probe_files(bus); - if (retval) - goto bus_probe_files_fail; + if (use_drivers) { + retval = add_probe_files(bus); + if (retval) + goto bus_probe_files_fail; + } retval = sysfs_create_groups(bus_kobj, bus->bus_groups); if (retval) @@ -1016,9 +1020,41 @@ out: kfree(priv); return retval; } + +/** + * bus_register - register a driver-core subsystem + * @bus: bus to register + * + * Once we have that, we register the bus with the kobject + * infrastructure, then register the children subsystems it has: + * the devices and drivers that belong to the subsystem. + */ +int bus_register(const struct bus_type *bus) +{ + return bus_register_internal(bus, true); +} EXPORT_SYMBOL_GPL(bus_register); /** + * companion_bus_register - register a companion bus type + * @bus: companion bus to register + * + * A companion bus is a bus without drivers. Devices that belong to it can be + * bound to other devices as their "companions" and represent interfaces that + * can be used by the drivers of those other devices. They may also be used for + * the enumeration of those other devices. + * + * The ACPI bus is a specific example of a companion bus. + * + * Registering a companion bus is like registering a regular bus except that it + * skips the creation of sysfs interfaces related to drivers for @bus. + */ +int companion_bus_register(const struct bus_type *bus) +{ + return bus_register_internal(bus, false); +} + +/** * bus_unregister - remove a bus from the system * @bus: bus. * @@ -1415,6 +1451,11 @@ struct device_driver *driver_find(const char *name, const struct bus_type *bus) if (!sp) return NULL; + if (sp->no_drivers) { + subsys_put(sp); + return NULL; + } + k = kset_find_obj(sp->drivers_kset, name); subsys_put(sp); if (!k) diff --git a/drivers/base/power/clock_ops.c b/drivers/base/power/clock_ops.c index 59bb37e8244c..3b8ba27802a8 100644 --- a/drivers/base/power/clock_ops.c +++ b/drivers/base/power/clock_ops.c @@ -250,7 +250,7 @@ EXPORT_SYMBOL_GPL(pm_clk_add); * * Add the clock to the list of clocks used for the power management of @dev. * The power-management code will take control of the clock reference, so - * callers should not call clk_put() on @clk after this function sucessfully + * callers should not call clk_put() on @clk after this function successfully * returned. */ int pm_clk_add_clk(struct device *dev, struct clk *clk) diff --git a/drivers/base/power/main.c b/drivers/base/power/main.c index e130da428141..4e0d6357a37c 100644 --- a/drivers/base/power/main.c +++ b/drivers/base/power/main.c @@ -802,6 +802,7 @@ static void device_resume_noirq(struct device *dev, pm_message_t state, bool asy const char *info = NULL; bool skip_resume; int error = 0; + DECLARE_DPM_WATCHDOG_ON_STACK(wd); TRACE_DEVICE(dev); TRACE_RESUME(0); @@ -827,6 +828,7 @@ static void device_resume_noirq(struct device *dev, pm_message_t state, bool asy if (!dpm_wait_for_superior(dev, async)) goto Out; + dpm_watchdog_set(&wd, dev); skip_resume = dev_pm_skip_resume(dev); /* * If the driver callback is skipped below or by the middle layer @@ -871,6 +873,7 @@ Run: error = dpm_run_callback(callback, dev, state, info); Skip: + dpm_watchdog_clear(&wd); dev->power.is_noirq_suspended = false; Out: @@ -971,6 +974,7 @@ static void device_resume_early(struct device *dev, pm_message_t state, bool asy pm_callback_t callback = NULL; const char *info = NULL; int error = 0; + DECLARE_DPM_WATCHDOG_ON_STACK(wd); TRACE_DEVICE(dev); TRACE_RESUME(0); @@ -987,6 +991,7 @@ static void device_resume_early(struct device *dev, pm_message_t state, bool asy if (!dpm_wait_for_superior(dev, async)) goto Out; + dpm_watchdog_set(&wd, dev); if (dev->pm_domain) { info = "early power domain "; callback = pm_late_early_op(&dev->pm_domain->ops, state); @@ -1004,7 +1009,7 @@ static void device_resume_early(struct device *dev, pm_message_t state, bool asy goto Run; if (dev_pm_skip_resume(dev)) - goto Skip; + goto End; if (dev->driver && dev->driver->pm) { info = "early driver "; @@ -1014,6 +1019,9 @@ static void device_resume_early(struct device *dev, pm_message_t state, bool asy Run: error = dpm_run_callback(callback, dev, state, info); +End: + dpm_watchdog_clear(&wd); + Skip: dev->power.is_late_suspended = false; pm_runtime_enable(dev); @@ -1281,10 +1289,12 @@ static void device_complete(struct device *dev, pm_message_t state) { void (*callback)(struct device *) = NULL; const char *info = NULL; + DECLARE_DPM_WATCHDOG_ON_STACK(wd); if (dev->power.syscore) goto out; + dpm_watchdog_set(&wd, dev); device_lock(dev); if (dev->pm_domain) { @@ -1312,6 +1322,7 @@ static void device_complete(struct device *dev, pm_message_t state) } device_unlock(dev); + dpm_watchdog_clear(&wd); out: /* If enabling runtime PM for the device is blocked, unblock it. */ @@ -1508,6 +1519,7 @@ static void device_suspend_noirq(struct device *dev, pm_message_t state, bool as pm_callback_t callback = NULL; const char *info = NULL; int error = 0; + DECLARE_DPM_WATCHDOG_ON_STACK(wd); TRACE_DEVICE(dev); TRACE_SUSPEND(0); @@ -1520,6 +1532,7 @@ static void device_suspend_noirq(struct device *dev, pm_message_t state, bool as if (dev->power.syscore || dev->power.direct_complete) goto Complete; + dpm_watchdog_set(&wd, dev); if (dev->pm_domain) { info = "noirq power domain "; callback = pm_noirq_op(&dev->pm_domain->ops, state); @@ -1550,7 +1563,7 @@ Run: WRITE_ONCE(async_error, error); dpm_save_failed_dev(dev_name(dev)); pm_dev_err(dev, state, async ? " async noirq" : " noirq", error); - goto Complete; + goto End; } Skip: @@ -1569,6 +1582,9 @@ Skip: if (dev->power.must_resume) dpm_superior_set_must_resume(dev); +End: + dpm_watchdog_clear(&wd); + Complete: complete_all(&dev->power.completion); TRACE_SUSPEND(error); @@ -1703,6 +1719,7 @@ static void device_suspend_late(struct device *dev, pm_message_t state, bool asy pm_callback_t callback = NULL; const char *info = NULL; int error = 0; + DECLARE_DPM_WATCHDOG_ON_STACK(wd); TRACE_DEVICE(dev); TRACE_SUSPEND(0); @@ -1720,6 +1737,8 @@ static void device_suspend_late(struct device *dev, pm_message_t state, bool asy if (dev->power.direct_complete) goto Complete; + dpm_watchdog_set(&wd, dev); + /* * After this point, any runtime PM operations targeting the device * will fail until the corresponding pm_runtime_enable() call in @@ -1761,13 +1780,16 @@ Run: dpm_save_failed_dev(dev_name(dev)); pm_dev_err(dev, state, async ? " async late" : " late", error); pm_runtime_enable(dev); - goto Complete; + goto End; } dpm_propagate_wakeup_to_parent(dev); Skip: dev->power.is_late_suspended = true; +End: + dpm_watchdog_clear(&wd); + Complete: TRACE_SUSPEND(error); complete_all(&dev->power.completion); @@ -2201,6 +2223,7 @@ static int device_prepare(struct device *dev, pm_message_t state) int (*callback)(struct device *) = NULL; bool smart_suspend; int ret = 0; + DECLARE_DPM_WATCHDOG_ON_STACK(wd); /* * If a device's parent goes into runtime suspend at the wrong time, @@ -2220,6 +2243,7 @@ static int device_prepare(struct device *dev, pm_message_t state) if (dev->power.syscore) return 0; + dpm_watchdog_set(&wd, dev); device_lock(dev); dev->power.wakeup_path = false; @@ -2245,6 +2269,7 @@ static int device_prepare(struct device *dev, pm_message_t state) unlock: device_unlock(dev); + dpm_watchdog_clear(&wd); if (ret < 0) { suspend_report_result(dev, callback, ret); diff --git a/drivers/base/power/runtime-test.c b/drivers/base/power/runtime-test.c index 1535ad2b0264..24865ce844fc 100644 --- a/drivers/base/power/runtime-test.c +++ b/drivers/base/power/runtime-test.c @@ -5,6 +5,7 @@ #include <linux/cleanup.h> #include <linux/pm_runtime.h> +#include <linux/workqueue.h> #include <kunit/device.h> #include <kunit/test.h> @@ -229,6 +230,61 @@ static void pm_runtime_probe_active_test(struct kunit *test) KUNIT_EXPECT_TRUE(test, pm_runtime_suspended(dev)); } +static void pm_runtime_supplier_suspend_test(struct kunit *test) +{ + struct device *dev = kunit_device_register(test, DEVICE_NAME); + struct device *supplier = kunit_device_register(test, DEVICE_NAME "_supplier"); + struct device *norpm_supplier = kunit_device_register(test, DEVICE_NAME "_norpm_supplier"); + + KUNIT_ASSERT_NOT_ERR_OR_NULL(test, dev); + KUNIT_ASSERT_NOT_ERR_OR_NULL(test, supplier); + KUNIT_ASSERT_NOT_ERR_OR_NULL(test, norpm_supplier); + + KUNIT_ASSERT_NOT_NULL(test, device_link_add(dev, supplier, DL_FLAG_PM_RUNTIME)); + KUNIT_ASSERT_NOT_NULL(test, device_link_add(dev, norpm_supplier, 0)); + + pm_runtime_enable(dev); + pm_runtime_enable(supplier); + pm_runtime_enable(norpm_supplier); + + KUNIT_EXPECT_TRUE(test, pm_runtime_suspended(dev)); + KUNIT_EXPECT_TRUE(test, pm_runtime_suspended(supplier)); + KUNIT_EXPECT_TRUE(test, pm_runtime_suspended(norpm_supplier)); + + /* + * Resume the device and RPM-linked supplier; non-RPM-linked supplier + * stays suspended. + */ + KUNIT_EXPECT_EQ(test, 0, pm_runtime_get_sync(dev)); + KUNIT_EXPECT_TRUE(test, pm_runtime_active(dev)); + KUNIT_EXPECT_TRUE(test, pm_runtime_active(supplier)); + KUNIT_EXPECT_TRUE(test, pm_runtime_suspended(norpm_supplier)); + + /* Also resume the non-RPM-linked supplier. */ + KUNIT_EXPECT_EQ(test, 0, pm_runtime_get_sync(norpm_supplier)); + KUNIT_EXPECT_TRUE(test, pm_runtime_active(norpm_supplier)); + + pm_runtime_put_noidle(norpm_supplier); + KUNIT_EXPECT_TRUE(test, pm_runtime_active(norpm_supplier)); + + /* + * Suspend device. This should only suspend the device and its + * RPM-linked supplier, not the non-RPM-linked supplier. + */ + KUNIT_EXPECT_EQ(test, 0, pm_runtime_put_sync(dev)); + /* Note: supplier suspend is async, so we have to flush the queue. */ + flush_workqueue(pm_wq); + KUNIT_EXPECT_TRUE(test, pm_runtime_suspended(dev)); + KUNIT_EXPECT_TRUE(test, pm_runtime_suspended(supplier)); + KUNIT_EXPECT_TRUE(test, pm_runtime_active(norpm_supplier)); + + /* Now suspend non-RPM-linked supplier. */ + KUNIT_EXPECT_EQ(test, 0, pm_runtime_suspend(norpm_supplier)); + KUNIT_EXPECT_TRUE(test, pm_runtime_suspended(dev)); + KUNIT_EXPECT_TRUE(test, pm_runtime_suspended(supplier)); + KUNIT_EXPECT_TRUE(test, pm_runtime_suspended(norpm_supplier)); +} + static struct kunit_case pm_runtime_test_cases[] = { KUNIT_CASE(pm_runtime_depth_test), KUNIT_CASE(pm_runtime_already_suspended_test), @@ -236,6 +292,7 @@ static struct kunit_case pm_runtime_test_cases[] = { KUNIT_CASE(pm_runtime_disabled_test), KUNIT_CASE(pm_runtime_error_test), KUNIT_CASE(pm_runtime_probe_active_test), + KUNIT_CASE(pm_runtime_supplier_suspend_test), {} }; diff --git a/drivers/base/power/runtime.c b/drivers/base/power/runtime.c index fab38bc98113..655c0b8af095 100644 --- a/drivers/base/power/runtime.c +++ b/drivers/base/power/runtime.c @@ -162,7 +162,7 @@ static void pm_runtime_cancel_pending(struct device *dev) dev->power.request = RPM_REQ_NONE; } -/* +/** * pm_runtime_autosuspend_expiration - Get a device's autosuspend-delay expiration time. * @dev: Device to handle. * @@ -200,7 +200,7 @@ static int dev_memalloc_noio(struct device *dev, void *data) return dev->power.memalloc_noio; } -/* +/** * pm_runtime_set_memalloc_noio - Set a device's memalloc_noio flag. * @dev: Device to handle. * @enable: True for setting the flag and False for clearing the flag. @@ -361,7 +361,8 @@ static void rpm_suspend_suppliers(struct device *dev) list_for_each_entry_rcu(link, &dev->links.suppliers, c_node, device_links_read_lock_held()) - pm_request_idle(link->supplier); + if (device_link_test(link, DL_FLAG_PM_RUNTIME)) + pm_request_idle(link->supplier); device_links_read_unlock(idx); } @@ -1043,6 +1044,11 @@ static enum hrtimer_restart pm_suspend_timer_fn(struct hrtimer *timer) * pm_schedule_suspend - Set up a timer to submit a suspend request in future. * @dev: Device to suspend. * @delay: Time to wait before submitting a suspend request, in milliseconds. + * + * Return: + * * %1: Success; @dev is already %RPM_SUSPENDED. + * * %0: Success. + * * Error code on failure. */ int pm_schedule_suspend(struct device *dev, unsigned int delay) { @@ -1096,16 +1102,15 @@ static int rpm_drop_usage_count(struct device *dev) } /** - * __pm_runtime_idle - Entry point for runtime idle operations. + * __pm_runtime_idle - Core entry point for runtime idle operations. * @dev: Device to send idle notification for. * @rpmflags: Flag bits. * - * If the RPM_GET_PUT flag is set, decrement the device's usage count and - * return immediately if it is larger than zero (if it becomes negative, log a - * warning, increment it, and return an error). Then carry out an idle - * notification, either synchronous or asynchronous. + * Carry out an idle check for @dev, either synchronous or asynchronous. + * If %RPM_GET_PUT is set in @rpmflags, decrement the device's usage count + * first, proceeding with idle notification only if the counter drops to zero. * - * This routine may be called in atomic context if the RPM_ASYNC flag is set, + * This routine may be called in atomic context if the %RPM_ASYNC flag is set, * or if pm_runtime_irq_safe() has been called. */ int __pm_runtime_idle(struct device *dev, int rpmflags) @@ -1134,16 +1139,15 @@ int __pm_runtime_idle(struct device *dev, int rpmflags) EXPORT_SYMBOL_GPL(__pm_runtime_idle); /** - * __pm_runtime_suspend - Entry point for runtime put/suspend operations. + * __pm_runtime_suspend - Core entry point for runtime put/suspend operations. * @dev: Device to suspend. * @rpmflags: Flag bits. * - * If the RPM_GET_PUT flag is set, decrement the device's usage count and - * return immediately if it is larger than zero (if it becomes negative, log a - * warning, increment it, and return an error). Then carry out a suspend, - * either synchronous or asynchronous. + * Carry out a suspend operation for @dev, either synchronous or asynchronous. + * If %RPM_GET_PUT is set in @rpmflags, decrement the device's usage count + * first, proceeding with suspend only if the counter drops to zero. * - * This routine may be called in atomic context if the RPM_ASYNC flag is set, + * This routine may be called in atomic context if the %RPM_ASYNC flag is set, * or if pm_runtime_irq_safe() has been called. */ int __pm_runtime_suspend(struct device *dev, int rpmflags) @@ -1172,14 +1176,15 @@ int __pm_runtime_suspend(struct device *dev, int rpmflags) EXPORT_SYMBOL_GPL(__pm_runtime_suspend); /** - * __pm_runtime_resume - Entry point for runtime resume operations. + * __pm_runtime_resume - Core entry point for runtime resume operations. * @dev: Device to resume. * @rpmflags: Flag bits. * - * If the RPM_GET_PUT flag is set, increment the device's usage count. Then - * carry out a resume, either synchronous or asynchronous. + * Carry out a runtime resume operation for @dev, either synchronous or + * asynchronous. If %RPM_GET_PUT is set in @rpmflags, increment the device's + * usage count first, then bring the device to %RPM_ACTIVE state. * - * This routine may be called in atomic context if the RPM_ASYNC flag is set, + * This routine may be called in atomic context if the %RPM_ASYNC flag is set, * or if pm_runtime_irq_safe() has been called. */ int __pm_runtime_resume(struct device *dev, int rpmflags) @@ -1254,9 +1259,12 @@ static int pm_runtime_get_conditional(struct device *dev, bool ign_usage_count) * @dev: Target device. * * Increment the runtime PM usage counter of @dev if its runtime PM status is - * %RPM_ACTIVE, in which case it returns 1. If the device is in a different - * state, 0 is returned. -EINVAL is returned if runtime PM is disabled for the - * device, in which case also the usage_count will remain unmodified. + * already %RPM_ACTIVE + * + * Return: + * * %-EINVAL: Runtime PM is disabled for @dev. The usage counter is not incremented. + * * %1: Success; usage counter is incremented. + * * %0: @dev was not active. */ int pm_runtime_get_if_active(struct device *dev) { @@ -1268,17 +1276,15 @@ EXPORT_SYMBOL_GPL(pm_runtime_get_if_active); * pm_runtime_get_if_in_use - Conditionally bump up runtime PM usage counter. * @dev: Target device. * - * Increment the runtime PM usage counter of @dev if its runtime PM status is - * %RPM_ACTIVE and its runtime PM usage counter is greater than 0 or it is not - * ignoring children and its active child count is nonzero. 1 is returned in - * this case. - * - * If @dev is in a different state or it is not in use (that is, its usage - * counter is 0, or it is ignoring children, or its active child count is 0), - * 0 is returned. + * Increment the runtime PM usage counter of @dev if it is "in use." A device + * is considered in use if its runtime PM status is %RPM_ACTIVE and its runtime + * PM usage counter is greater than 0, or if it is not ignoring children and + * its active child count is nonzero. * - * -EINVAL is returned if runtime PM is disabled for the device, in which case - * also the usage counter of @dev is not updated. + * Return: + * * %-EINVAL: Runtime PM is disabled for @dev. The usage counter is not incremented. + * * %1: Success; usage counter is incremented. + * * %0: @dev was not in use; usage counter is not incremented. */ int pm_runtime_get_if_in_use(struct device *dev) { @@ -1287,7 +1293,7 @@ int pm_runtime_get_if_in_use(struct device *dev) EXPORT_SYMBOL_GPL(pm_runtime_get_if_in_use); /** - * __pm_runtime_set_status - Set runtime PM status of a device. + * __pm_runtime_set_status - Set runtime PM status of a device and clear errors. * @dev: Device to handle. * @status: New runtime PM status of the device. * @@ -1306,7 +1312,7 @@ EXPORT_SYMBOL_GPL(pm_runtime_get_if_in_use); * If @dev has any suppliers (as reflected by device links to them), and @status * is RPM_ACTIVE, they will be activated upfront and if the activation of one * of them fails, the status of @dev will be changed to RPM_SUSPENDED (instead - * of the @status value) and the suppliers will be deacticated on exit. The + * of the @status value) and the suppliers will be deactivated on exit. The * error returned by the failing supplier activation will be returned in that * case. */ @@ -1462,11 +1468,13 @@ static void __pm_runtime_barrier(struct device *dev) * pm_runtime_barrier - Flush pending requests and wait for completions. * @dev: Device to handle. * - * Prevent the device from being suspended by incrementing its usage counter and - * if there's a pending resume request for the device, wake the device up. - * Next, make sure that all pending requests for the device have been flushed - * from pm_wq and wait for all runtime PM operations involving the device in - * progress to complete. + * If the device has a pending resume request, resume it synchronously. For all + * other request types, cancel any queued request, and wait for running + * operations to complete. + * + * Note that this is intentionally asymmetric, as it guarantees any queued + * asynchronous resume request will complete, but it may cancel asynchronous + * suspend requests. */ void pm_runtime_barrier(struct device *dev) { @@ -1550,8 +1558,16 @@ void __pm_runtime_disable(struct device *dev, bool check_resume) EXPORT_SYMBOL_GPL(__pm_runtime_disable); /** - * pm_runtime_enable - Enable runtime PM of a device. + * pm_runtime_enable - Enable runtime PM for a device. * @dev: Device to handle. + * + * Enable runtime PM transitions for @dev by decrementing its disable counter. + * Once the counter reaches zero, the PM core is permitted to execute runtime + * PM callbacks for @dev as power conditions change. + * + * Callers should ensure that the device's runtime PM status accurately reflects + * its physical hardware state (via pm_runtime_set_active() or + * pm_runtime_set_suspended()) before enabling runtime PM. */ void pm_runtime_enable(struct device *dev) { @@ -1874,6 +1890,9 @@ void pm_runtime_reinit(struct device *dev) if (dev->power.runtime_status == RPM_ACTIVE) pm_runtime_set_suspended(dev); + if (dev->power.use_autosuspend) + pm_runtime_dont_use_autosuspend(dev); + if (dev->power.irq_safe) { spin_lock_irq(&dev->power.lock); dev->power.irq_safe = 0; diff --git a/drivers/clocksource/timer-ti-dm.c b/drivers/clocksource/timer-ti-dm.c index 6787acac9a43..922a8f14562c 100644 --- a/drivers/clocksource/timer-ti-dm.c +++ b/drivers/clocksource/timer-ti-dm.c @@ -1528,7 +1528,7 @@ err_disable: */ static void omap_dm_timer_remove(struct platform_device *pdev) { - struct dmtimer *timer; + struct dmtimer *timer, *found = NULL; unsigned long flags; int ret = -EINVAL; @@ -1536,14 +1536,17 @@ static void omap_dm_timer_remove(struct platform_device *pdev) list_for_each_entry(timer, &omap_timer_list, node) if (!strcmp(dev_name(&timer->pdev->dev), dev_name(&pdev->dev))) { - if (!(timer->capability & OMAP_TIMER_ALWON)) - cpu_pm_unregister_notifier(&timer->nb); list_del(&timer->node); + found = timer; ret = 0; break; } spin_unlock_irqrestore(&dm_timer_lock, flags); + /* Unregister outside the lock: cpu_pm_unregister_notifier() may sleep. */ + if (found && !(found->capability & OMAP_TIMER_ALWON)) + cpu_pm_unregister_notifier(&found->nb); + pm_runtime_disable(&pdev->dev); if (ret) diff --git a/drivers/cpufreq/amd-pstate-ut.c b/drivers/cpufreq/amd-pstate-ut.c index e23773680e05..c2c1a166b3a9 100644 --- a/drivers/cpufreq/amd-pstate-ut.c +++ b/drivers/cpufreq/amd-pstate-ut.c @@ -226,11 +226,14 @@ static int amd_pstate_ut_check_freq(u32 index) for_each_online_cpu(cpu) { struct cpufreq_policy *policy __free(put_cpufreq_policy) = NULL; struct amd_cpudata *cpudata; + union perf_cached perf; policy = cpufreq_cpu_get(cpu); if (!policy) continue; + cpudata = policy->driver_data; + perf = READ_ONCE(cpudata->perf); if (!((policy->cpuinfo.max_freq >= cpudata->nominal_freq) && (cpudata->nominal_freq > cpudata->lowest_nonlinear_freq) && @@ -242,7 +245,25 @@ static int amd_pstate_ut_check_freq(u32 index) return -EINVAL; } - if (cpudata->lowest_nonlinear_freq != policy->min) { + if (perf.bios_min_perf) { + u32 bios_min_freq = perf_to_freq(perf, cpudata->nominal_freq, + perf.bios_min_perf); + + /* + * User set bios_min_freq cannot be trusted to + * be within the driver limits. Clamp it similar + * to cpufreq_verify_within_cpu_limits(). + */ + bios_min_freq = clamp_t(u32, bios_min_freq, + policy->cpuinfo.min_freq, + policy->cpuinfo.max_freq); + + if (bios_min_freq != policy->min) { + pr_err("%s cpu%d bios_min_freq=%d policy_min=%d, they should be equal!\n", + __func__, cpu, bios_min_freq, policy->min); + return -EINVAL; + } + } else if (cpudata->lowest_nonlinear_freq != policy->min) { pr_err("%s cpu%d cpudata_lowest_nonlinear_freq=%d policy_min=%d, they should be equal!\n", __func__, cpu, cpudata->lowest_nonlinear_freq, policy->min); return -EINVAL; diff --git a/drivers/cpufreq/amd-pstate.c b/drivers/cpufreq/amd-pstate.c index 8bfd46d60843..fbccbe86c94d 100644 --- a/drivers/cpufreq/amd-pstate.c +++ b/drivers/cpufreq/amd-pstate.c @@ -55,10 +55,10 @@ #define AMD_PSTATE_TRANSITION_DELAY 1000 #define AMD_PSTATE_FAST_CPPC_TRANSITION_DELAY 600 -#define AMD_CPPC_EPP_PERFORMANCE 0x00 -#define AMD_CPPC_EPP_BALANCE_PERFORMANCE 0x80 -#define AMD_CPPC_EPP_BALANCE_POWERSAVE 0xBF -#define AMD_CPPC_EPP_POWERSAVE 0xFF +#define AMD_CPPC_EPP_LEGACY_PERFORMANCE 0x00 +#define AMD_CPPC_EPP_LEGACY_BALANCE_PERFORMANCE 0x80 +#define AMD_CPPC_EPP_LEGACY_BALANCE_POWERSAVE 0xBF +#define AMD_CPPC_EPP_LEGACY_POWERSAVE 0xFF static const char * const amd_pstate_mode_string[] = { [AMD_PSTATE_UNDEFINED] = "undefined", @@ -129,14 +129,112 @@ static const char * const energy_perf_strings[] = { }; static_assert(ARRAY_SIZE(energy_perf_strings) == EPP_INDEX_MAX); -static unsigned int epp_values[] = { - [EPP_INDEX_DEFAULT] = 0, - [EPP_INDEX_PERFORMANCE] = AMD_CPPC_EPP_PERFORMANCE, - [EPP_INDEX_BALANCE_PERFORMANCE] = AMD_CPPC_EPP_BALANCE_PERFORMANCE, - [EPP_INDEX_BALANCE_POWERSAVE] = AMD_CPPC_EPP_BALANCE_POWERSAVE, - [EPP_INDEX_POWERSAVE] = AMD_CPPC_EPP_POWERSAVE, +/* + * The numeric EPP value programmed for each named preference. First dimension + * is CPU type (TOPO_CPU_TYPE_ANY for non-hybrid, TOPO_CPU_TYPE_PERFORMANCE/ + * EFFICIENCY/LOW_POWER for hybrid). The initializer holds the legacy values + * used as the fallback on any platform not listed in amd_pstate_epp_soc_ids[]; + * amd_pstate_init_epp_values() overwrites slots at boot when the running SoC + * has a per-SoC (and potentially per-CPU-type) override. + */ +static u8 epp_values[][EPP_INDEX_MAX] = { + /* + * Initialize all CPU types to legacy defaults. + * amd_pstate_init_epp_values() will fix these up + * based on the platform during boot. + */ + [TOPO_CPU_TYPE_ANY ... TOPO_CPU_TYPE_LOW_POWER] = { + [EPP_INDEX_DEFAULT] = 0, + [EPP_INDEX_PERFORMANCE] = AMD_CPPC_EPP_LEGACY_PERFORMANCE, + [EPP_INDEX_BALANCE_PERFORMANCE] = AMD_CPPC_EPP_LEGACY_BALANCE_PERFORMANCE, + [EPP_INDEX_BALANCE_POWERSAVE] = AMD_CPPC_EPP_LEGACY_BALANCE_POWERSAVE, + [EPP_INDEX_POWERSAVE] = AMD_CPPC_EPP_LEGACY_POWERSAVE, + }, +}; +static_assert(ARRAY_SIZE(epp_values) == TOPO_CPU_TYPE_LOW_POWER + 1, + "epp_values must have entries for all CPU types up to TOPO_CPU_TYPE_LOW_POWER"); + +/* + * Get the EPP value row for a given CPU, accounting for hybrid CPU types. + * Non-hybrid systems use TOPO_CPU_TYPE_ANY; hybrid systems use the CPU's + * actual type (PERFORMANCE/EFFICIENCY/LOW_POWER). + */ +static inline u8 *amd_pstate_cpu_epp_values(enum x86_topology_cpu_type cpu_type) +{ + switch (cpu_type) { + case TOPO_CPU_TYPE_PERFORMANCE: + case TOPO_CPU_TYPE_EFFICIENCY: + case TOPO_CPU_TYPE_LOW_POWER: + return epp_values[cpu_type]; + default: + return epp_values[TOPO_CPU_TYPE_ANY]; + } +} + +/** + * struct amd_pstate_epp_values - EPP values for the four named preferences + * @performance: value for the "performance" preference + * @balance_performance: value for the "balance_performance" preference + * @balance_power: value for the "balance_power" preference + * @power: value for the "power" preference + */ +struct amd_pstate_epp_values { + u8 performance; + u8 balance_performance; + u8 balance_power; + u8 power; +}; + +/** + * struct amd_pstate_epp_soc - per-CPU-type EPP overrides for hybrid systems + * @performance_core: values for TOPO_CPU_TYPE_PERFORMANCE cores + * @efficiency_core: values for TOPO_CPU_TYPE_EFFICIENCY cores + * @low_power_core: values for TOPO_CPU_TYPE_LOW_POWER cores + * + * Referenced from amd_pstate_epp_soc_ids[] to give a hybrid platform its own + * numeric EPP values for the four named preferences, with distinct values per + * CPU type. Non-hybrid systems are not listed in the table and always use the + * legacy defaults. + */ +struct amd_pstate_epp_soc { + struct amd_pstate_epp_values performance_core; + struct amd_pstate_epp_values efficiency_core; + struct amd_pstate_epp_values low_power_core; +}; + +/* + * Per-CPU-type EPP overrides for hybrid systems. Only hybrid SoCs should be + * listed here; non-hybrid systems always use the legacy defaults. + */ +static const struct amd_pstate_epp_soc epp_soc_zen6_client __initconst = { + .performance_core = { + .performance = 25, + .balance_performance = 51, + .balance_power = 64, + .power = 64, + }, + .efficiency_core = { + .performance = 25, + .balance_performance = 51, + .balance_power = 64, + .power = 115, + }, + .low_power_core = { + .performance = 25, + .balance_performance = 51, + .balance_power = 64, + .power = 115, + }, +}; + +static const struct x86_cpu_id amd_pstate_epp_soc_ids[] __initconst = { + X86_MATCH_VENDOR_FAM_MODEL(AMD, 0x1A, 0x80, &epp_soc_zen6_client), + X86_MATCH_VENDOR_FAM_MODEL(AMD, 0x1A, 0x81, &epp_soc_zen6_client), + X86_MATCH_VENDOR_FAM_MODEL(AMD, 0x1A, 0x84, &epp_soc_zen6_client), + X86_MATCH_VENDOR_FAM_MODEL(AMD, 0x1A, 0x85, &epp_soc_zen6_client), + X86_MATCH_VENDOR_FAM_MODEL(AMD, 0x1A, 0xe0, &epp_soc_zen6_client), + {} }; -static_assert(ARRAY_SIZE(epp_values) == EPP_INDEX_MAX - 2); typedef int (*cppc_mode_transition_fn)(int); @@ -145,19 +243,6 @@ static struct quirk_entry quirk_amd_7k62 = { .lowest_freq = 550, }; -static inline u8 freq_to_perf(union perf_cached perf, u32 nominal_freq, unsigned int freq_val) -{ - u32 perf_val = DIV_ROUND_UP_ULL((u64)freq_val * perf.nominal_perf, nominal_freq); - - return (u8)clamp(perf_val, perf.lowest_perf, perf.highest_perf); -} - -static inline u32 perf_to_freq(union perf_cached perf, u32 nominal_freq, u8 perf_val) -{ - return DIV_ROUND_UP_ULL((u64)nominal_freq * perf_val, - perf.nominal_perf); -} - static int __init dmi_matched_7k62_bios_bug(const struct dmi_system_id *dmi) { /** @@ -544,7 +629,11 @@ static int shmem_update_perf(struct cpufreq_policy *policy, u8 min_perf, u8 des_perf, u8 max_perf, u8 epp, bool fast_switch) { struct amd_cpudata *cpudata = policy->driver_data; - struct cppc_perf_ctrls perf_ctrls; + struct cppc_perf_ctrls perf_ctrls = { + .max_perf = max_perf, + .min_perf = min_perf, + .desired_perf = des_perf, + }; u64 value, prev; int ret; @@ -577,10 +666,6 @@ static int shmem_update_perf(struct cpufreq_policy *policy, u8 min_perf, if (value == prev) return 0; - perf_ctrls.max_perf = max_perf; - perf_ctrls.min_perf = min_perf; - perf_ctrls.desired_perf = des_perf; - ret = cppc_set_perf(cpudata->cpu, &perf_ctrls); if (ret) return ret; @@ -1068,6 +1153,7 @@ static int amd_pstate_cpu_init(struct cpufreq_policy *policy) return -ENOMEM; cpudata->cpu = policy->cpu; + cpudata->cpu_type = cpu_data(policy->cpu).topo.cpu_type; ret = amd_pstate_init_perf(cpudata); if (ret) @@ -1195,13 +1281,16 @@ static int amd_pstate_power_supply_notifier(struct notifier_block *nb, static int amd_pstate_get_epp_from_platform_profile(struct cpufreq_policy *policy, enum platform_profile_option profile) { + struct amd_cpudata *cpudata = policy->driver_data; + u8 *values = amd_pstate_cpu_epp_values(cpudata->cpu_type); + switch (profile) { case PLATFORM_PROFILE_PERFORMANCE: - return AMD_CPPC_EPP_PERFORMANCE; + return values[EPP_INDEX_PERFORMANCE]; case PLATFORM_PROFILE_BALANCED: return amd_pstate_get_balanced_epp(policy); case PLATFORM_PROFILE_LOW_POWER: - return AMD_CPPC_EPP_POWERSAVE; + return values[EPP_INDEX_POWERSAVE]; default: break; } @@ -1411,6 +1500,7 @@ ssize_t store_energy_performance_preference(struct cpufreq_policy *policy, const char *buf, size_t count) { struct amd_cpudata *cpudata = policy->driver_data; + u8 *values = amd_pstate_cpu_epp_values(cpudata->cpu_type); ssize_t ret; bool raw_epp = false; u8 epp; @@ -1445,12 +1535,13 @@ ssize_t store_energy_performance_preference(struct cpufreq_policy *policy, } if (ret) - epp = epp_values[ret]; + epp = values[ret]; else epp = cpudata->epp_default_dc; } - if (epp > 0 && cpudata->policy == CPUFREQ_POLICY_PERFORMANCE) { + if (epp > 0 && epp != values[EPP_INDEX_PERFORMANCE] && + cpudata->policy == CPUFREQ_POLICY_PERFORMANCE) { pr_debug("EPP cannot be set under performance policy\n"); return -EBUSY; } @@ -1475,34 +1566,31 @@ EXPORT_SYMBOL_FOR_PSTATE_UT(store_energy_performance_preference); ssize_t show_energy_performance_preference(struct cpufreq_policy *policy, char *buf) { struct amd_cpudata *cpudata = policy->driver_data; - u8 preference, epp; + u8 *values = amd_pstate_cpu_epp_values(cpudata->cpu_type); + u8 epp; + int i; epp = FIELD_GET(AMD_CPPC_EPP_PERF_MASK, cpudata->cppc_req_cached); if (!cpudata->dynamic_epp && cpudata->raw_epp) return sysfs_emit(buf, "%u\n", epp); - switch (epp) { - case AMD_CPPC_EPP_PERFORMANCE: - preference = EPP_INDEX_PERFORMANCE; - break; - case AMD_CPPC_EPP_BALANCE_PERFORMANCE: - preference = EPP_INDEX_BALANCE_PERFORMANCE; - break; - case AMD_CPPC_EPP_BALANCE_POWERSAVE: - preference = EPP_INDEX_BALANCE_POWERSAVE; - break; - case AMD_CPPC_EPP_POWERSAVE: - preference = EPP_INDEX_POWERSAVE; - break; - default: - return -EINVAL; - } + /* + * Map the cached EPP value back to a named preference. Skip the + * "default" slot (index 0) so an EPP of 0 reports as "performance". + * Stop at POWERSAVE; CUSTOM and DYNAMIC are not initialized in epp_values. + */ + for (i = EPP_INDEX_PERFORMANCE; i <= EPP_INDEX_POWERSAVE; i++) { + const char *name = energy_perf_strings[i]; - if (cpudata->dynamic_epp) - return sysfs_emit(buf, "dynamic(profile:%s)\n", energy_perf_strings[preference]); + if (epp == values[i]) { + if (cpudata->dynamic_epp) + return sysfs_emit(buf, "dynamic(profile:%s)\n", name); + return sysfs_emit(buf, "%s\n", name); + } + } - return sysfs_emit(buf, "%s\n", energy_perf_strings[preference]); + return sysfs_emit(buf, "%u\n", epp); } EXPORT_SYMBOL_FOR_PSTATE_UT(show_energy_performance_preference); @@ -1794,7 +1882,7 @@ EXPORT_SYMBOL_FOR_PSTATE_UT(amd_pstate_get_status); int amd_pstate_update_status(const char *buf, size_t size) { - int mode_idx; + int cpu, mode_idx; if (size > strlen("passive") || size < strlen("active")) return -EINVAL; @@ -1803,12 +1891,23 @@ int amd_pstate_update_status(const char *buf, size_t size) if (mode_idx < 0) return mode_idx; - if (mode_state_machine[cppc_state][mode_idx]) { - guard(mutex)(&amd_pstate_driver_lock); - return mode_state_machine[cppc_state][mode_idx](mode_idx); + guard(mutex)(&amd_pstate_driver_lock); + + if (!mode_state_machine[cppc_state][mode_idx]) + return 0; + + if (mode_idx == AMD_PSTATE_PASSIVE && + !cpu_feature_enabled(X86_FEATURE_CPPC)) { + guard(cpus_read_lock)(); + + /* Check offline CPUs too, before changing or removing the driver. */ + for_each_present_cpu(cpu) { + if (cppc_auto_sel_is_immutable(cpu)) + return -EOPNOTSUPP; + } } - return 0; + return mode_state_machine[cppc_state][mode_idx](mode_idx); } EXPORT_SYMBOL_FOR_PSTATE_UT(amd_pstate_update_status); @@ -1894,6 +1993,7 @@ static int amd_pstate_epp_cpu_init(struct cpufreq_policy *policy) return -ENOMEM; cpudata->cpu = policy->cpu; + cpudata->cpu_type = cpu_data(policy->cpu).topo.cpu_type; ret = amd_pstate_init_perf(cpudata); if (ret) @@ -1944,9 +2044,11 @@ static int amd_pstate_epp_cpu_init(struct cpufreq_policy *policy) cpudata->epp_default_ac = cpudata->epp_default_dc = default_epp; cpudata->current_profile = PLATFORM_PROFILE_PERFORMANCE; } else { + u8 *values = amd_pstate_cpu_epp_values(cpudata->cpu_type); + policy->policy = CPUFREQ_POLICY_POWERSAVE; - cpudata->epp_default_ac = AMD_CPPC_EPP_PERFORMANCE; - cpudata->epp_default_dc = AMD_CPPC_EPP_BALANCE_PERFORMANCE; + cpudata->epp_default_ac = values[EPP_INDEX_PERFORMANCE]; + cpudata->epp_default_dc = values[EPP_INDEX_BALANCE_PERFORMANCE]; cpudata->current_profile = PLATFORM_PROFILE_BALANCED; } @@ -2245,6 +2347,38 @@ static bool amd_cppc_supported(void) return true; } +/* + * Resolve the numeric EPP values for hybrid systems. Only hybrid SoCs are listed + * in amd_pstate_epp_soc_ids[]; non-hybrid systems always use the legacy defaults. + */ +static inline void __init amd_pstate_set_epp_values(enum x86_topology_cpu_type type, + const struct amd_pstate_epp_values *core) +{ + epp_values[type][EPP_INDEX_PERFORMANCE] = core->performance; + epp_values[type][EPP_INDEX_BALANCE_PERFORMANCE] = core->balance_performance; + epp_values[type][EPP_INDEX_BALANCE_POWERSAVE] = core->balance_power; + epp_values[type][EPP_INDEX_POWERSAVE] = core->power; +} + +static void __init amd_pstate_init_epp_values(void) +{ + const struct x86_cpu_id *id = x86_match_cpu(amd_pstate_epp_soc_ids); + const struct amd_pstate_epp_soc *soc; + + if (!id || !id->driver_data) { + if (cpu_feature_enabled(X86_FEATURE_ZEN6) && + cpu_feature_enabled(X86_FEATURE_AMD_HTR_CORES)) + pr_warn_once("No EPP tunings found for platform\n"); + return; + } + + soc = (const struct amd_pstate_epp_soc *)id->driver_data; + + amd_pstate_set_epp_values(TOPO_CPU_TYPE_PERFORMANCE, &soc->performance_core); + amd_pstate_set_epp_values(TOPO_CPU_TYPE_EFFICIENCY, &soc->efficiency_core); + amd_pstate_set_epp_values(TOPO_CPU_TYPE_LOW_POWER, &soc->low_power_core); +} + static int __init amd_pstate_init(void) { struct device *dev_root; @@ -2272,6 +2406,9 @@ static int __init amd_pstate_init(void) /* check if this machine need CPPC quirks */ dmi_check_system(amd_pstate_quirks_table); + /* resolve per-SoC EPP values for the named preferences */ + amd_pstate_init_epp_values(); + /* * determine the driver mode from the command line or kernel config. * If no command line input is provided, cppc_state will be AMD_PSTATE_UNDEFINED. diff --git a/drivers/cpufreq/amd-pstate.h b/drivers/cpufreq/amd-pstate.h index 9f5a81976eae..c7189d177b81 100644 --- a/drivers/cpufreq/amd-pstate.h +++ b/drivers/cpufreq/amd-pstate.h @@ -146,6 +146,8 @@ struct amd_cpudata { enum platform_profile_option current_profile; struct device *ppdev; char *profile_name; + + enum x86_topology_cpu_type cpu_type; }; /* @@ -159,6 +161,20 @@ enum amd_pstate_mode { AMD_PSTATE_GUIDED, AMD_PSTATE_MAX, }; + +static inline u8 freq_to_perf(union perf_cached perf, u32 nominal_freq, unsigned int freq_val) +{ + u32 perf_val = DIV_ROUND_UP_ULL((u64)freq_val * perf.nominal_perf, nominal_freq); + + return (u8)clamp(perf_val, perf.lowest_perf, perf.highest_perf); +} + +static inline u32 perf_to_freq(union perf_cached perf, u32 nominal_freq, u8 perf_val) +{ + return DIV_ROUND_UP_ULL((u64)nominal_freq * perf_val, + perf.nominal_perf); +} + const char *amd_pstate_get_mode_string(enum amd_pstate_mode mode); int amd_pstate_get_status(void); int amd_pstate_update_status(const char *buf, size_t size); diff --git a/drivers/cpufreq/cppc_cpufreq.c b/drivers/cpufreq/cppc_cpufreq.c index 80893844353c..09e55bfdca88 100644 --- a/drivers/cpufreq/cppc_cpufreq.c +++ b/drivers/cpufreq/cppc_cpufreq.c @@ -18,6 +18,7 @@ #include <linux/cpufreq.h> #include <linux/irq_work.h> #include <linux/kthread.h> +#include <linux/mutex.h> #include <linux/time.h> #include <linux/vmalloc.h> #include <uapi/linux/sched/types.h> @@ -41,6 +42,7 @@ MODULE_PARM_DESC(fie_disabled, "Disable Frequency Invariance Engine (FIE)"); /* Frequency invariance support */ struct cppc_freq_invariance { int cpu; + bool pcc_work_initialized; struct irq_work irq_work; struct kthread_work work; struct cppc_perf_fb_ctrs prev_perf_fb_ctrs; @@ -49,10 +51,12 @@ struct cppc_freq_invariance { static DEFINE_PER_CPU(struct cppc_freq_invariance, cppc_freq_inv); static struct kthread_worker *kworker_fie; +static DEFINE_MUTEX(cppc_fie_lock); static int cppc_perf_from_fbctrs(u64 reference_perf, struct cppc_perf_fb_ctrs *fb_ctrs_t0, struct cppc_perf_fb_ctrs *fb_ctrs_t1); +static int cppc_fie_kworker_init(void); /** * __cppc_scale_freq_tick - CPPC arch_freq_scale updater for frequency invariance @@ -149,7 +153,7 @@ static struct scale_freq_data cppc_sftd_pcc = { static void cppc_cpufreq_cpu_fie_init(struct cpufreq_policy *policy) { - struct scale_freq_data *sftd = &cppc_sftd; + struct scale_freq_data *sftd; struct cppc_freq_invariance *cppc_fi; int cpu, ret; @@ -161,9 +165,12 @@ static void cppc_cpufreq_cpu_fie_init(struct cpufreq_policy *policy) cppc_fi->cpu = cpu; cppc_fi->cpu_data = policy->driver_data; if (cppc_perf_ctrs_in_pcc_cpu(cpu)) { + if (cppc_fie_kworker_init()) + return; + kthread_init_work(&cppc_fi->work, cppc_scale_freq_workfn); init_irq_work(&cppc_fi->irq_work, cppc_irq_work); - sftd = &cppc_sftd_pcc; + cppc_fi->pcc_work_initialized = true; } ret = cppc_get_perf_ctrs(cpu, &cppc_fi->prev_perf_fb_ctrs); @@ -179,17 +186,20 @@ static void cppc_cpufreq_cpu_fie_init(struct cpufreq_policy *policy) } } - /* Register for freq-invariance */ - topology_set_scale_freq_source(sftd, policy->cpus); + /* A shared policy may contain both PCC and non-PCC counters. */ + for_each_cpu(cpu, policy->cpus) { + cppc_fi = &per_cpu(cppc_freq_inv, cpu); + if (cppc_fi->pcc_work_initialized) + sftd = &cppc_sftd_pcc; + else + sftd = &cppc_sftd; + topology_set_scale_freq_source(sftd, cpumask_of(cpu)); + } } /* - * We free all the resources on policy's removal and not on CPU removal as the - * irq-work are per-cpu and the hotplug core takes care of flushing the pending - * irq-works (hint: smpcfd_dying_cpu()) on CPU hotplug. Even if the kthread-work - * fires on another CPU after the concerned CPU is removed, it won't harm. - * - * We just need to make sure to remove them all on policy->exit(). + * Drain work initialized by this policy even if processor removal has + * already unpublished the CPU's CPC descriptor. */ static void cppc_cpufreq_cpu_fie_exit(struct cpufreq_policy *policy) { @@ -203,16 +213,18 @@ static void cppc_cpufreq_cpu_fie_exit(struct cpufreq_policy *policy) topology_clear_scale_freq_source(SCALE_FREQ_SOURCE_CPPC, policy->related_cpus); for_each_cpu(cpu, policy->related_cpus) { - if (!cppc_perf_ctrs_in_pcc_cpu(cpu)) - continue; cppc_fi = &per_cpu(cppc_freq_inv, cpu); + if (!cppc_fi->pcc_work_initialized) + continue; irq_work_sync(&cppc_fi->irq_work); kthread_cancel_work_sync(&cppc_fi->work); + cppc_fi->pcc_work_initialized = false; } } -static void cppc_fie_kworker_init(void) +static int cppc_fie_kworker_init(void) { + struct kthread_worker *worker; struct sched_attr attr = { .size = sizeof(struct sched_attr), .sched_policy = SCHED_DEADLINE, @@ -228,23 +240,28 @@ static void cppc_fie_kworker_init(void) }; int ret; - kworker_fie = kthread_run_worker(0, "cppc_fie"); - if (IS_ERR(kworker_fie)) { + guard(mutex)(&cppc_fie_lock); + + if (kworker_fie) + return 0; + + worker = kthread_run_worker(0, "cppc_fie"); + if (IS_ERR(worker)) { pr_warn("%s: failed to create kworker_fie: %ld\n", __func__, - PTR_ERR(kworker_fie)); - fie_disabled = FIE_DISABLED; - kworker_fie = NULL; - return; + PTR_ERR(worker)); + return PTR_ERR(worker); } - ret = sched_setattr_nocheck(kworker_fie->task, &attr); + ret = sched_setattr_nocheck(worker->task, &attr); if (ret) { pr_warn("%s: failed to set SCHED_DEADLINE: %d\n", __func__, ret); - kthread_destroy_worker(kworker_fie); - fie_disabled = FIE_DISABLED; - kworker_fie = NULL; + kthread_destroy_worker(worker); + return ret; } + + kworker_fie = worker; + return 0; } static void __init cppc_freq_invariance_init(void) @@ -259,11 +276,6 @@ static void __init cppc_freq_invariance_init(void) fie_disabled = FIE_ENABLED; } } - - if (fie_disabled || !perf_ctrs_in_pcc) - return; - - cppc_fie_kworker_init(); } static void cppc_freq_invariance_exit(void) @@ -883,6 +895,7 @@ static ssize_t store_auto_select(struct cpufreq_policy *policy, const char *buf, size_t count) { struct cppc_cpudata *cpu_data = policy->driver_data; + bool old_auto_sel = cpu_data->perf_ctrls.auto_sel; bool val; int ret; @@ -911,8 +924,8 @@ static ssize_t store_auto_select(struct cpufreq_policy *policy, if (ret) { cpu_data->perf_ctrls.min_perf = old_min_perf; cpu_data->perf_ctrls.max_perf = old_max_perf; - cppc_set_auto_sel(policy->cpu, false); - cpu_data->perf_ctrls.auto_sel = false; + cppc_set_auto_sel(policy->cpu, old_auto_sel); + cpu_data->perf_ctrls.auto_sel = old_auto_sel; return ret; } } diff --git a/drivers/cpufreq/cpufreq.c b/drivers/cpufreq/cpufreq.c index 96515880b4ac..54dde8419bdc 100644 --- a/drivers/cpufreq/cpufreq.c +++ b/drivers/cpufreq/cpufreq.c @@ -2590,8 +2590,8 @@ static void cpufreq_update_pressure(struct cpufreq_policy *policy) cpu = cpumask_first(policy->related_cpus); max_freq = arch_scale_freq_ref(cpu); - if (!max_freq) - max_freq = policy->cpuinfo.max_freq; + if (!max_freq && cpufreq_driver->scale_freq_ref) + max_freq = cpufreq_driver->scale_freq_ref(policy); capped_freq = policy->max; diff --git a/drivers/cpufreq/cpufreq_conservative.c b/drivers/cpufreq/cpufreq_conservative.c index 0b32ae28ec85..02bfd46543e9 100644 --- a/drivers/cpufreq/cpufreq_conservative.c +++ b/drivers/cpufreq/cpufreq_conservative.c @@ -85,10 +85,12 @@ static unsigned int cs_dbs_update(struct cpufreq_policy *policy) freq_step = get_freq_step(cs_tuners, policy); /* - * Decrease requested_freq one freq_step for each idle period that - * we didn't update the frequency. + * Apply deferred down steps only when the policy sample load is + * below down_threshold. Otherwise, multiple deferred down steps may + * cause a net frequency decrease outside the downscaling region. */ - if (policy_dbs->idle_periods < UINT_MAX) { + if (policy_dbs->max_sample_load < cs_tuners->down_threshold && + policy_dbs->idle_periods < UINT_MAX) { unsigned int freq_steps = policy_dbs->idle_periods * freq_step; if (requested_freq > policy->min + freq_steps) diff --git a/drivers/cpufreq/cpufreq_governor.c b/drivers/cpufreq/cpufreq_governor.c index 710d93ec89b5..f3bd3dbe7c41 100644 --- a/drivers/cpufreq/cpufreq_governor.c +++ b/drivers/cpufreq/cpufreq_governor.c @@ -124,7 +124,8 @@ unsigned int dbs_update(struct cpufreq_policy *policy) struct policy_dbs_info *policy_dbs = policy->governor_data; struct dbs_data *dbs_data = policy_dbs->dbs_data; unsigned int ignore_nice = dbs_data->ignore_nice_load; - unsigned int max_load = 0, idle_periods = UINT_MAX; + unsigned int max_load = 0, max_sample_load = 0; + unsigned int idle_periods = UINT_MAX; unsigned int sampling_rate, io_busy, j; u64 cur_nice; @@ -147,7 +148,7 @@ unsigned int dbs_update(struct cpufreq_policy *policy) struct cpu_dbs_info *j_cdbs = &per_cpu(cpu_dbs, j); u64 update_time, cur_idle_time; unsigned int idle_time, time_elapsed; - unsigned int load; + unsigned int load, sample_load; cur_idle_time = get_cpu_idle_time(j, &update_time, io_busy); @@ -186,6 +187,20 @@ unsigned int dbs_update(struct cpufreq_policy *policy) j_cdbs->prev_cpu_nice = cur_nice; + /* + * Compute the sample load separately from the prev_load value + * that may be reused after a long idle interval. The conservative + * governor uses it to decide whether to apply deferred down steps. + * If no time has elapsed, retain the existing behavior and use + * prev_load. + */ + if (unlikely(!time_elapsed)) + sample_load = j_cdbs->prev_load; + else if (time_elapsed > idle_time) + sample_load = 100 * (time_elapsed - idle_time) / time_elapsed; + else + sample_load = 0; + if (unlikely(!time_elapsed)) { /* * That can only happen when this function is called @@ -220,11 +235,7 @@ unsigned int dbs_update(struct cpufreq_policy *policy) load = j_cdbs->prev_load; j_cdbs->prev_load = 0; } else { - if (time_elapsed > idle_time) - load = 100 * (time_elapsed - idle_time) / time_elapsed; - else - load = 0; - + load = sample_load; j_cdbs->prev_load = load; } @@ -237,9 +248,13 @@ unsigned int dbs_update(struct cpufreq_policy *policy) if (load > max_load) max_load = load; + + if (sample_load > max_sample_load) + max_sample_load = sample_load; } policy_dbs->idle_periods = idle_periods; + policy_dbs->max_sample_load = max_sample_load; return max_load; } diff --git a/drivers/cpufreq/cpufreq_governor.h b/drivers/cpufreq/cpufreq_governor.h index 73b8ed7cfaae..806d8fb4dff1 100644 --- a/drivers/cpufreq/cpufreq_governor.h +++ b/drivers/cpufreq/cpufreq_governor.h @@ -95,6 +95,8 @@ struct policy_dbs_info { /* Multiplier for increasing sample delay temporarily. */ unsigned int rate_mult; unsigned int idle_periods; /* For conservative */ + /* Maximum load from the current policy sample. */ + unsigned int max_sample_load; /* Status indicators */ bool is_shared; /* This object is used by multiple CPUs */ bool work_in_progress; /* Work is being queued up or in progress */ diff --git a/drivers/cpufreq/intel_pstate.c b/drivers/cpufreq/intel_pstate.c index ceb340f7a110..17158cb25dd8 100644 --- a/drivers/cpufreq/intel_pstate.c +++ b/drivers/cpufreq/intel_pstate.c @@ -1135,6 +1135,14 @@ static bool hybrid_clear_max_perf_cpu(void) return ret; } +static unsigned int intel_pstate_scale_freq_ref(struct cpufreq_policy *policy) +{ + if (READ_ONCE(all_cpu_data[policy->cpu]->capacity_perf)) + return policy->cpuinfo.max_freq; + + return 0; +} + static void intel_pstate_update_freq_limits(struct cpudata *cpu) { int scaling = cpu->pstate.scaling; @@ -3088,6 +3096,7 @@ static struct cpufreq_driver intel_pstate = { .offline = intel_pstate_cpu_offline, .online = intel_pstate_cpu_online, .update_limits = intel_pstate_update_limits, + .scale_freq_ref = intel_pstate_scale_freq_ref, .name = "intel_pstate", }; @@ -3411,6 +3420,7 @@ static struct cpufreq_driver intel_cpufreq = { .suspend = intel_cpufreq_suspend, .resume = intel_pstate_resume, .update_limits = intel_pstate_update_limits, + .scale_freq_ref = intel_pstate_scale_freq_ref, .name = "intel_cpufreq", }; diff --git a/drivers/cpuidle/cpuidle-tegra.c b/drivers/cpuidle/cpuidle-tegra.c index aca907a62bb5..d2fc216c38fc 100644 --- a/drivers/cpuidle/cpuidle-tegra.c +++ b/drivers/cpuidle/cpuidle-tegra.c @@ -281,7 +281,7 @@ static int tegra114_enter_s2idle(struct cpuidle_device *dev, * LP2 | C7 (CPU core power gating) * LP2 | CC6 (CPU cluster power gating) * - * Note that that the older CPUIDLE driver versions didn't explicitly + * Note that the older CPUIDLE driver versions didn't explicitly * differentiate the LP2 states because these states either used the same * code path or because CC6 wasn't supported. */ diff --git a/drivers/cpuidle/governors/menu.c b/drivers/cpuidle/governors/menu.c index 544a5d593007..eb529d73d4b7 100644 --- a/drivers/cpuidle/governors/menu.c +++ b/drivers/cpuidle/governors/menu.c @@ -284,10 +284,10 @@ static int menu_select(struct cpuidle_driver *drv, struct cpuidle_device *dev, data->bucket = BUCKETS - 1; } - if (latency_req == 0 || - ((data->next_timer_ns < drv->states[1].target_residency_ns || - latency_req < drv->states[1].exit_latency_ns) && - !dev->states_usage[0].disable)) { + if (!dev->states_usage[0].disable && + (latency_req == 0 || + data->next_timer_ns < drv->states[1].target_residency_ns || + latency_req < drv->states[1].exit_latency_ns)) { /* * In this case state[0] will be used no matter what, so return * it right away and keep the tick running if state[0] is a diff --git a/drivers/cpuidle/governors/teo.c b/drivers/cpuidle/governors/teo.c index ac43b9b013b3..906f4f89e365 100644 --- a/drivers/cpuidle/governors/teo.c +++ b/drivers/cpuidle/governors/teo.c @@ -427,6 +427,17 @@ static int teo_select(struct cpuidle_driver *drv, struct cpuidle_device *dev, } /* + * If the latency constraint does not allow any of the enabled idle + * states to be used, the candidate state index will be capped to 0 + * by the check below, but state 0 may be disabled. To prevent the + * selection of a disabled state in that case, ensure that + * constraint_idx is at least equal to the index of the first + * enabled idle state. + */ + if (constraint_idx < idx0) + constraint_idx = idx0; + + /* * If there is a latency constraint, it may be necessary to select an * idle state shallower than the current candidate one. */ diff --git a/drivers/cxl/core/ras.c b/drivers/cxl/core/ras.c index e307361bb39e..c91db125a971 100644 --- a/drivers/cxl/core/ras.c +++ b/drivers/cxl/core/ras.c @@ -77,7 +77,7 @@ static int match_memdev_by_parent(struct device *dev, const void *uport) return 0; } -void cxl_cper_handle_prot_err(struct cxl_cper_prot_err_work_data *data) +static void cxl_cper_handle_prot_err(struct cxl_cper_prot_err_work_data *data) { unsigned int devfn = PCI_DEVFN(data->prot_err.agent_addr.device, data->prot_err.agent_addr.function); @@ -118,7 +118,6 @@ void cxl_cper_handle_prot_err(struct cxl_cper_prot_err_work_data *data) else cxl_cper_trace_uncorr_prot_err(cxlmd, data->ras_cap); } -EXPORT_SYMBOL_GPL(cxl_cper_handle_prot_err); static void cxl_cper_prot_err_work_fn(struct work_struct *work) { diff --git a/drivers/firmware/efi/cper.c b/drivers/firmware/efi/cper.c index 06b4fdb59917..13b3e72c3b62 100644 --- a/drivers/firmware/efi/cper.c +++ b/drivers/firmware/efi/cper.c @@ -389,10 +389,28 @@ void cper_mem_err_pack(const struct cper_sec_mem_err *mem, cmem->requestor_id = mem->requestor_id; cmem->responder_id = mem->responder_id; cmem->target_id = mem->target_id; - cmem->extended = mem->extended; - cmem->rank = mem->rank; - cmem->mem_array_handle = mem->mem_array_handle; - cmem->mem_dev_handle = mem->mem_dev_handle; + + /* + * These four sit past the end of the UEFI 2.1/2.2 layout, which older + * firmware still emits, so reading them unconditionally runs off a + * short record. Every consumer of the compact record gates them on the + * same validation bits, so leave them zero when firmware does not + * claim them. + */ + cmem->extended = 0; + cmem->rank = 0; + cmem->mem_array_handle = 0; + cmem->mem_dev_handle = 0; + + if (mem->validation_bits & + (CPER_MEM_VALID_ROW_EXT | CPER_MEM_VALID_CHIP_ID)) + cmem->extended = mem->extended; + if (mem->validation_bits & CPER_MEM_VALID_RANK_NUMBER) + cmem->rank = mem->rank; + if (mem->validation_bits & CPER_MEM_VALID_CARD_HANDLE) + cmem->mem_array_handle = mem->mem_array_handle; + if (mem->validation_bits & CPER_MEM_VALID_MODULE_HANDLE) + cmem->mem_dev_handle = mem->mem_dev_handle; } EXPORT_SYMBOL_GPL(cper_mem_err_pack); @@ -745,6 +763,17 @@ int cper_estatus_check_header(const struct acpi_hest_generic_status *estatus) estatus->raw_data_offset < sizeof(*estatus) + estatus->data_length) return -EINVAL; + /* + * cper_estatus_len() sums these into a u32, and a wrapped sum reads + * back smaller than the record. Reject a length that cannot be + * expressed so no caller is handed the short value. + */ + if ((u64)sizeof(*estatus) + estatus->data_length > U32_MAX) + return -EINVAL; + if (estatus->raw_data_length && + (u64)estatus->raw_data_offset + estatus->raw_data_length > U32_MAX) + return -EINVAL; + return 0; } EXPORT_SYMBOL_GPL(cper_estatus_check_header); @@ -752,7 +781,7 @@ EXPORT_SYMBOL_GPL(cper_estatus_check_header); int cper_estatus_check(const struct acpi_hest_generic_status *estatus) { struct acpi_hest_generic_data *gdata; - unsigned int data_len, record_size; + unsigned int data_len; int rc; rc = cper_estatus_check_header(estatus); @@ -762,10 +791,18 @@ int cper_estatus_check(const struct acpi_hest_generic_status *estatus) data_len = estatus->data_length; apei_estatus_for_each_section(estatus, gdata) { - if (acpi_hest_get_size(gdata) > data_len) + int record_size; + + /* + * The <acpi/ghes.h> helpers sum these as a signed int, so a + * huge error_data_length wraps small rather than large and the + * walk then advances by that wrapped value. Reject a size an + * int cannot carry. + */ + if (check_add_overflow(acpi_hest_get_size(gdata), + gdata->error_data_length, &record_size)) return -EINVAL; - record_size = acpi_hest_get_record_size(gdata); if (record_size > data_len) return -EINVAL; diff --git a/drivers/idle/intel_idle.c b/drivers/idle/intel_idle.c index 651408df9c24..a045a384b949 100644 --- a/drivers/idle/intel_idle.c +++ b/drivers/idle/intel_idle.c @@ -1792,9 +1792,7 @@ static bool __init intel_idle_cst_usable(void) { int cstate, limit; - limit = min_t(int, min_t(int, CPUIDLE_STATE_MAX, max_cstate + 1), - acpi_state_table.count); - + limit = min3(CPUIDLE_STATE_MAX, max_cstate + 1, acpi_state_table.count); for (cstate = 1; cstate < limit; cstate++) { struct acpi_processor_cx *cx = &acpi_state_table.states[cstate]; @@ -1906,7 +1904,7 @@ static void __init intel_idle_init_cstates_acpi_lpi(struct cpuidle_driver *drv) state = &drv->states[drv->state_count++]; scnprintf(state->name, CPUIDLE_NAME_LEN, "C%d_LPI", index + 1); - strscpy(state->desc, lpi_state->desc, CPUIDLE_DESC_LEN); + strscpy(state->desc, lpi_state->desc); state->exit_latency = lpi_state->wake_latency; state->target_residency = lpi_state->min_residency; state->flags = MWAIT2flg(lpi_state->address); @@ -1938,7 +1936,7 @@ static void __init intel_idle_init_cstates_acpi_lpi(struct cpuidle_driver *drv) static void __init intel_idle_init_cstates_acpi_cst(struct cpuidle_driver *drv) { - int cstate, limit = min_t(int, CPUIDLE_STATE_MAX, acpi_state_table.count); + int cstate, limit = min(CPUIDLE_STATE_MAX, acpi_state_table.count); /* * If limit > 0, intel_idle_cst_usable() has returned 'true', so all of @@ -1956,7 +1954,7 @@ static void __init intel_idle_init_cstates_acpi_cst(struct cpuidle_driver *drv) state = &drv->states[drv->state_count++]; snprintf(state->name, CPUIDLE_NAME_LEN, "C%d_ACPI", cstate); - strscpy(state->desc, cx->desc, CPUIDLE_DESC_LEN); + strscpy(state->desc, cx->desc); state->exit_latency = cx->latency; /* * For C1-type C-states use the same number for both the exit @@ -2018,7 +2016,7 @@ static bool __init intel_idle_off_by_default_lpi(unsigned int flags, u32 mwait_h static bool __init intel_idle_off_by_default_cst(unsigned int flags, u32 mwait_hint) { - int cstate, limit = min_t(int, CPUIDLE_STATE_MAX, acpi_state_table.count); + int cstate, limit = min(CPUIDLE_STATE_MAX, acpi_state_table.count); /* * If limit > 0, intel_idle_cst_usable() has returned 'true', so all of diff --git a/drivers/platform/x86/lenovo/yogabook.c b/drivers/platform/x86/lenovo/yogabook.c index 1a4b2ab1f35d..8a91d440724b 100644 --- a/drivers/platform/x86/lenovo/yogabook.c +++ b/drivers/platform/x86/lenovo/yogabook.c @@ -353,13 +353,13 @@ static int yogabook_wmi_probe(struct wmi_device *wdev, const void *context) goto error_put_devs; } - data->kbd_dev = get_device(acpi_get_first_physical_node(data->kbd_adev)); + data->kbd_dev = acpi_bus_get_primary_device(data->kbd_adev); if (!data->kbd_dev || !data->kbd_dev->driver) { r = -EPROBE_DEFER; goto error_put_devs; } - data->dig_dev = get_device(acpi_get_first_physical_node(data->dig_adev)); + data->dig_dev = acpi_bus_get_primary_device(data->dig_adev); if (!data->dig_dev || !data->dig_dev->driver) { r = -EPROBE_DEFER; goto error_put_devs; diff --git a/drivers/platform/x86/serdev_helpers.h b/drivers/platform/x86/serdev_helpers.h index 57eac75805e2..20267e1941f1 100644 --- a/drivers/platform/x86/serdev_helpers.h +++ b/drivers/platform/x86/serdev_helpers.h @@ -74,8 +74,7 @@ get_serdev_controller(const char *serial_ctrl_hid, return ERR_PTR(-ENODEV); } - /* get_first_physical_node() returns a weak ref */ - parent = get_device(acpi_get_first_physical_node(adev)); + parent = acpi_bus_get_primary_device(adev); acpi_dev_put(adev); if (!parent) { pr_err("error could not get %s/%s serial-ctrl physical node\n", diff --git a/drivers/platform/x86/x86-android-tablets/core.c b/drivers/platform/x86/x86-android-tablets/core.c index cfff7f5eac5d..2c262f117235 100644 --- a/drivers/platform/x86/x86-android-tablets/core.c +++ b/drivers/platform/x86/x86-android-tablets/core.c @@ -390,7 +390,6 @@ static int gpio_secondary_fwnode_init(struct device *parent, { const struct software_node *const *swnode; struct fwnode_handle *fwnode; - struct device *phys_dev; int ret; if (!node_group) @@ -418,7 +417,8 @@ static int gpio_secondary_fwnode_init(struct device *parent, if (WARN_ON(!fwnode)) return -ENOENT; - phys_dev = acpi_get_first_physical_node(to_acpi_device(dev)); + struct device *phys_dev __free(put_device) = + acpi_bus_get_primary_device(to_acpi_device(dev)); if (!phys_dev) return dev_err_probe(parent, -ENODEV, "No physical device for ACPI GPIO dev: %pfwP\n", diff --git a/drivers/pnp/driver.c b/drivers/pnp/driver.c index 6d58b1465081..80c2a0f15b80 100644 --- a/drivers/pnp/driver.c +++ b/drivers/pnp/driver.c @@ -96,13 +96,13 @@ static int pnp_device_probe(struct device *dev) if (!(pnp_drv->flags & PNP_DRIVER_RES_DO_NOT_CHANGE)) { error = pnp_activate_dev(pnp_dev); if (error < 0) - return error; + goto fail; } } else if ((pnp_drv->flags & PNP_DRIVER_RES_DISABLE) == PNP_DRIVER_RES_DISABLE) { error = pnp_disable_dev(pnp_dev); if (error < 0) - return error; + goto fail; } error = 0; if (pnp_drv->probe) { diff --git a/drivers/powercap/intel_rapl_msr.c b/drivers/powercap/intel_rapl_msr.c index a34543e66446..d6686d03056a 100644 --- a/drivers/powercap/intel_rapl_msr.c +++ b/drivers/powercap/intel_rapl_msr.c @@ -456,6 +456,7 @@ static const struct x86_cpu_id rapl_ids[] = { X86_MATCH_VFM(INTEL_EMERALDRAPIDS_X, &rapl_defaults_spr_server), X86_MATCH_VFM(INTEL_LUNARLAKE_M, &rapl_defaults_core), X86_MATCH_VFM(INTEL_PANTHERLAKE_L, &rapl_defaults_core_pl4_pmu), + X86_MATCH_VFM(INTEL_PANTHERLAKE_R, &rapl_defaults_core_pl4_pmu), X86_MATCH_VFM(INTEL_WILDCATLAKE_L, &rapl_defaults_core_pl4_pmu), X86_MATCH_VFM(INTEL_NOVALAKE, &rapl_defaults_core_pl4), X86_MATCH_VFM(INTEL_NOVALAKE_L, &rapl_defaults_core_pl4), @@ -582,8 +583,8 @@ static int intel_rapl_msr_init(void) rapl_msr_platdev = platform_device_register_data(NULL, "intel_rapl_msr", 0, def, sizeof(*def)); if (IS_ERR(rapl_msr_platdev)) - pr_debug("intel_rapl_msr device register failed, ret:%ld\n", - PTR_ERR(rapl_msr_platdev)); + pr_debug("intel_rapl_msr device register failed, ret:%pe\n", + rapl_msr_platdev); return 0; } diff --git a/drivers/thermal/cpufreq_cooling.c b/drivers/thermal/cpufreq_cooling.c index 768859a7aed0..d75b42fb955e 100644 --- a/drivers/thermal/cpufreq_cooling.c +++ b/drivers/thermal/cpufreq_cooling.c @@ -488,12 +488,12 @@ static int cpufreq_set_cur_state(struct thermal_cooling_device *cdev, frequency = get_state_freq(cpufreq_cdev, state); ret = freq_qos_update_request(&cpufreq_cdev->qos_req, frequency); - if (ret >= 0) { - cpufreq_cdev->cpufreq_state = state; - ret = 0; - } + if (ret < 0) + return ret; + + cpufreq_cdev->cpufreq_state = state; - return ret; + return 0; } /** @@ -661,8 +661,8 @@ of_cpufreq_cooling_register(struct cpufreq_policy *policy) cdev = __cpufreq_cooling_register(np, policy, em); if (IS_ERR(cdev)) { - pr_err("cpufreq_cooling: cpu%d failed to register as cooling device: %ld\n", - policy->cpu, PTR_ERR(cdev)); + pr_err("cpufreq_cooling: cpu%d failed to register as cooling device: %pe\n", + policy->cpu, cdev); cdev = NULL; } } diff --git a/drivers/thermal/devfreq_cooling.c b/drivers/thermal/devfreq_cooling.c index 0330a8112832..9ada52cbf7dd 100644 --- a/drivers/thermal/devfreq_cooling.c +++ b/drivers/thermal/devfreq_cooling.c @@ -156,8 +156,8 @@ static unsigned long get_voltage(struct devfreq *df, unsigned long freq) opp = dev_pm_opp_find_freq_exact(dev, freq, false); if (IS_ERR(opp)) { - dev_err_ratelimited(dev, "Failed to find OPP for frequency %lu: %ld\n", - freq, PTR_ERR(opp)); + dev_err_ratelimited(dev, "Failed to find OPP for frequency %lu: %pe\n", + freq, opp); return 0; } diff --git a/drivers/thermal/gov_power_allocator.c b/drivers/thermal/gov_power_allocator.c index 37f2e22a999e..b5c254187628 100644 --- a/drivers/thermal/gov_power_allocator.c +++ b/drivers/thermal/gov_power_allocator.c @@ -660,10 +660,15 @@ static void power_allocator_update_tz(struct thermal_zone_device *tz, enum thermal_notify_event reason) { struct power_allocator_params *params = tz->governor_data; - const struct thermal_trip_desc *td = trip_to_trip_desc(params->trip_max); + const struct thermal_trip_desc *td; struct thermal_instance *instance; int num_actors = 0; + if (!params->trip_max) + return; + + td = trip_to_trip_desc(params->trip_max); + switch (reason) { case THERMAL_TZ_BIND_CDEV: case THERMAL_TZ_UNBIND_CDEV: diff --git a/drivers/thermal/intel/int340x_thermal/processor_thermal_device.c b/drivers/thermal/intel/int340x_thermal/processor_thermal_device.c index f80dbe2ca7e4..850bfe7c1ede 100644 --- a/drivers/thermal/intel/int340x_thermal/processor_thermal_device.c +++ b/drivers/thermal/intel/int340x_thermal/processor_thermal_device.c @@ -179,17 +179,21 @@ static int proc_thermal_get_zone_temp(struct thermal_zone_device *zone, { int cpu; int curr_temp, ret; - - *temp = 0; + bool temp_valid = false; for_each_online_cpu(cpu) { ret = intel_tcc_get_temp(cpu, &curr_temp, false); if (ret < 0) return ret; - if (!*temp || curr_temp > *temp) + if (!temp_valid || curr_temp > *temp) { *temp = curr_temp; + temp_valid = true; + } } + if (!temp_valid) + return -ENODATA; + *temp *= 1000; return 0; @@ -291,10 +295,8 @@ int proc_thermal_add(struct device *dev, struct proc_thermal_device *proc_priv) } proc_priv->int340x_zone = int340x_thermal_zone_add(adev, get_temp); - if (IS_ERR(proc_priv->int340x_zone)) { + if (IS_ERR(proc_priv->int340x_zone)) return PTR_ERR(proc_priv->int340x_zone); - } else - ret = 0; ret = acpi_install_notify_handler(adev->handle, ACPI_DEVICE_NOTIFY, proc_thermal_notify, diff --git a/drivers/thermal/intel/int340x_thermal/processor_thermal_soc_slider.c b/drivers/thermal/intel/int340x_thermal/processor_thermal_soc_slider.c index 91f291627132..e161b95828e8 100644 --- a/drivers/thermal/intel/int340x_thermal/processor_thermal_soc_slider.c +++ b/drivers/thermal/intel/int340x_thermal/processor_thermal_soc_slider.c @@ -66,15 +66,16 @@ static int slider_def_balance_set(const char *arg, const struct kernel_param *kp guard(mutex)(&slider_param_lock); ret = kstrtou8(arg, 16, &slider_val); - if (!ret) { - if (slider_val <= slider_values[SOC_POWER_SLIDER_PERFORMANCE] || - slider_val >= slider_values[SOC_POWER_SLIDER_POWERSAVE]) - return -EINVAL; + if (ret) + return ret; - slider_balanced_param = slider_val; - } + if (slider_val <= slider_values[SOC_POWER_SLIDER_PERFORMANCE] || + slider_val >= slider_values[SOC_POWER_SLIDER_POWERSAVE]) + return -EINVAL; - return ret; + slider_balanced_param = slider_val; + + return 0; } static int slider_def_balance_get(char *buf, const struct kernel_param *kp) @@ -101,14 +102,15 @@ static int slider_def_offset_set(const char *arg, const struct kernel_param *kp) guard(mutex)(&slider_param_lock); ret = kstrtou8(arg, 16, &offset); - if (!ret) { - if (offset > SOC_SLIDER_VALUE_MAXIMUM) - return -EINVAL; + if (ret) + return ret; - slider_offset = offset; - } + if (offset > SOC_SLIDER_VALUE_MAXIMUM) + return -EINVAL; - return ret; + slider_offset = offset; + + return 0; } static int slider_def_offset_get(char *buf, const struct kernel_param *kp) diff --git a/drivers/thermal/intel/intel_powerclamp.c b/drivers/thermal/intel/intel_powerclamp.c index bd7fd98dc310..1ffcb0a699e5 100644 --- a/drivers/thermal/intel/intel_powerclamp.c +++ b/drivers/thermal/intel/intel_powerclamp.c @@ -94,7 +94,7 @@ static int duration_set(const char *arg, const struct kernel_param *kp) } mutex_lock(&powerclamp_lock); - duration = clamp(new_duration, 6ul, 25ul) * 1000; + duration = new_duration * 1000; mutex_unlock(&powerclamp_lock); exit: @@ -103,13 +103,9 @@ exit: static int duration_get(char *buf, const struct kernel_param *kp) { - int ret; - - mutex_lock(&powerclamp_lock); - ret = sysfs_emit(buf, "%d\n", duration / 1000); - mutex_unlock(&powerclamp_lock); + guard(mutex)(&powerclamp_lock); - return ret; + return sysfs_emit(buf, "%u\n", duration / 1000); } static const struct kernel_param_ops duration_ops = { @@ -143,12 +139,9 @@ copy_mask: } /* Return true if the cpumask and idle percent combination is invalid */ -static bool check_invalid(cpumask_var_t mask, u8 idle) +static bool check_invalid(const struct cpumask *mask, u8 idle) { - if (cpumask_equal(cpu_present_mask, mask) && idle > MAX_ALL_CPU_IDLE) - return true; - - return false; + return cpumask_equal(cpu_present_mask, mask) && idle > MAX_ALL_CPU_IDLE; } static int cpumask_set(const char *arg, const struct kernel_param *kp) @@ -214,41 +207,32 @@ MODULE_PARM_DESC(cpumask, "Mask of CPUs to use for idle injection."); static int max_idle_set(const char *arg, const struct kernel_param *kp) { u8 new_max_idle; - int ret = 0; + int ret; - mutex_lock(&powerclamp_lock); + guard(mutex)(&powerclamp_lock); /* Can't set mask when cooling device is in use */ - if (powerclamp_data.clamping) { - ret = -EAGAIN; - goto skip_limit_set; - } + if (powerclamp_data.clamping) + return -EAGAIN; ret = kstrtou8(arg, 10, &new_max_idle); if (ret) - goto skip_limit_set; + return ret; - if (new_max_idle > MAX_TARGET_RATIO) { - ret = -EINVAL; - goto skip_limit_set; - } + if (new_max_idle > MAX_TARGET_RATIO) + return -EINVAL; if (!cpumask_available(idle_injection_cpu_mask)) { ret = allocate_copy_idle_injection_mask(cpu_present_mask); if (ret) - goto skip_limit_set; + return ret; } - if (check_invalid(idle_injection_cpu_mask, new_max_idle)) { - ret = -EINVAL; - goto skip_limit_set; - } + if (check_invalid(idle_injection_cpu_mask, new_max_idle)) + return -EINVAL; max_idle = new_max_idle; -skip_limit_set: - mutex_unlock(&powerclamp_lock); - return ret; } @@ -289,9 +273,10 @@ static int window_size_set(const char *arg, const struct kernel_param *kp) pr_err("Out of recommended window size %lu, between 2-10\n", new_window_size); ret = -EINVAL; + goto exit_win; } - window_size = clamp(new_window_size, 2ul, 10ul); + window_size = new_window_size; smp_mb(); exit_win: @@ -536,23 +521,17 @@ static struct idle_inject_device *ii_dev; */ static bool idle_inject_update(void) { - bool update = false; - /* We can't sleep in this callback */ if (!mutex_trylock(&powerclamp_lock)) return true; if (!(powerclamp_data.count % powerclamp_data.window_size_now)) { + unsigned int runtime; should_skip = powerclamp_adjust_controls(powerclamp_data.target_ratio, powerclamp_data.guard, powerclamp_data.window_size_now); - update = true; - } - - if (update) { - unsigned int runtime = get_run_time(); - + runtime = get_run_time(); idle_inject_set_duration(ii_dev, runtime, duration); } @@ -560,10 +539,7 @@ static bool idle_inject_update(void) mutex_unlock(&powerclamp_lock); - if (should_skip) - return false; - - return true; + return !should_skip; } /* This function starts idle injection by calling idle_inject_start() */ diff --git a/drivers/thermal/intel/intel_tcc_cooling.c b/drivers/thermal/intel/intel_tcc_cooling.c index 52ea4f7fc9f9..75f6fbec2530 100644 --- a/drivers/thermal/intel/intel_tcc_cooling.c +++ b/drivers/thermal/intel/intel_tcc_cooling.c @@ -69,6 +69,7 @@ static const struct x86_cpu_id tcc_ids[] __initconst = { X86_MATCH_VFM(INTEL_ARROWLAKE_U, NULL), X86_MATCH_VFM(INTEL_ARROWLAKE_H, NULL), X86_MATCH_VFM(INTEL_PANTHERLAKE_L, NULL), + X86_MATCH_VFM(INTEL_PANTHERLAKE_R, NULL), X86_MATCH_VFM(INTEL_WILDCATLAKE_L, NULL), X86_MATCH_VFM(INTEL_NOVALAKE, NULL), X86_MATCH_VFM(INTEL_NOVALAKE_L, NULL), diff --git a/drivers/thermal/qcom/qcom-spmi-adc-tm5.c b/drivers/thermal/qcom/qcom-spmi-adc-tm5.c index af72db6299cd..4554a6fa55d1 100644 --- a/drivers/thermal/qcom/qcom-spmi-adc-tm5.c +++ b/drivers/thermal/qcom/qcom-spmi-adc-tm5.c @@ -680,8 +680,8 @@ static int adc_tm5_register_tzd(struct adc_tm5_chip *adc_tm) continue; } - dev_err(adc_tm->dev, "Error registering TZ zone for channel %d: %ld\n", - adc_tm->channels[i].channel, PTR_ERR(tzd)); + dev_err(adc_tm->dev, "Error registering TZ zone for channel %d: %pe\n", + adc_tm->channels[i].channel, tzd); return PTR_ERR(tzd); } adc_tm->channels[i].tzd = tzd; diff --git a/drivers/thermal/thermal_core.c b/drivers/thermal/thermal_core.c index 82e2f0d8a26d..e8d4da0aeac3 100644 --- a/drivers/thermal/thermal_core.c +++ b/drivers/thermal/thermal_core.c @@ -1005,7 +1005,8 @@ out_kfree_cdev: return ERR_PTR(ret); } -int thermal_cooling_device_add(struct thermal_cooling_device *cdev, void *devdata) +int thermal_cooling_device_add(struct thermal_cooling_device *cdev, + struct device *parent, void *devdata) { unsigned long current_state; int ret; @@ -1013,6 +1014,7 @@ int thermal_cooling_device_add(struct thermal_cooling_device *cdev, void *devdat mutex_init(&cdev->lock); INIT_LIST_HEAD(&cdev->thermal_instances); cdev->updated = false; + cdev->device.parent = parent; cdev->device.class = &thermal_class; cdev->device.release = thermal_cdev_release; device_initialize(&cdev->device); @@ -1062,21 +1064,23 @@ out_put_device: } /** - * thermal_cooling_device_register() - register a new thermal cooling device + * thermal_cooling_device_create() - register a new thermal cooling device + * @parent: parent device (optional). * @type: the thermal cooling device type. * @devdata: device private data. * @ops: standard thermal cooling devices callbacks. * - * This interface function adds a new thermal cooling device (fan/processor/...) - * to /sys/class/thermal/ folder as cooling_device[0-*]. It tries to bind itself - * to all the thermal zone devices registered at the same time. + * Allocate and register a new thermal cooling device under the given parent (if + * not NULL) and with the given type, device data, and operations. During the + * registration, it will be matched against all of the registered thermal zones + * and it will be bound to the matching ones. * - * Return: a pointer to the created struct thermal_cooling_device or an - * ERR_PTR. Caller must check return value with IS_ERR*() helpers. + * Return: A pointer to the created struct thermal_cooling_device or an ERR_PTR. + * Callers must use IS_ERR*() helpers to check the return value. */ -struct thermal_cooling_device * -thermal_cooling_device_register(const char *type, void *devdata, - const struct thermal_cooling_device_ops *ops) +struct thermal_cooling_device *thermal_cooling_device_create( + struct device *parent, const char *type, void *devdata, + const struct thermal_cooling_device_ops *ops) { struct thermal_cooling_device *cdev; int ret; @@ -1085,13 +1089,13 @@ thermal_cooling_device_register(const char *type, void *devdata, if (IS_ERR(cdev)) return cdev; - ret = thermal_cooling_device_add(cdev, devdata); + ret = thermal_cooling_device_add(cdev, parent, devdata); if (ret) return ERR_PTR(ret); return cdev; } -EXPORT_SYMBOL_GPL(thermal_cooling_device_register); +EXPORT_SYMBOL_GPL(thermal_cooling_device_create); static void thermal_cooling_device_release(void *data) { @@ -1662,7 +1666,7 @@ EXPORT_SYMBOL_GPL(thermal_zone_device_unregister); * * Return: On success returns a reference to an unique thermal zone with * matching name equals to @name, an ERR_PTR otherwise (-EINVAL for invalid - * paramenters, -ENODEV for not found and -EEXIST for multiple matches). + * parameters, -ENODEV for not found and -EEXIST for multiple matches). */ struct thermal_zone_device *thermal_zone_get_zone_by_name(const char *name) { diff --git a/drivers/thermal/thermal_core.h b/drivers/thermal/thermal_core.h index e98b0aa5aacc..7ef7c6cab437 100644 --- a/drivers/thermal/thermal_core.h +++ b/drivers/thermal/thermal_core.h @@ -270,7 +270,8 @@ void thermal_governor_update_tz(struct thermal_zone_device *tz, struct thermal_cooling_device * thermal_cooling_device_alloc(const char *type, const struct thermal_cooling_device_ops *ops); -int thermal_cooling_device_add(struct thermal_cooling_device *cdev, void *devdata); +int thermal_cooling_device_add(struct thermal_cooling_device *cdev, + struct device *parent, void *devdata); /* Helpers */ #define for_each_trip_desc(__tz, __td) \ diff --git a/drivers/thermal/thermal_of.c b/drivers/thermal/thermal_of.c index 0217a49b08ae..9e96fb1b033e 100644 --- a/drivers/thermal/thermal_of.c +++ b/drivers/thermal/thermal_of.c @@ -63,22 +63,23 @@ static int thermal_of_get_trip_type(struct device_node *np, static int thermal_of_populate_trip(struct device_node *np, struct thermal_trip *trip) { - int prop; + u32 hysteresis; + s32 temperature; int ret; - ret = of_property_read_u32(np, "temperature", &prop); + ret = of_property_read_s32(np, "temperature", &temperature); if (ret < 0) { pr_err("missing temperature property\n"); return ret; } - trip->temperature = prop; + trip->temperature = temperature; - ret = of_property_read_u32(np, "hysteresis", &prop); + ret = of_property_read_u32(np, "hysteresis", &hysteresis); if (ret < 0) { pr_err("missing hysteresis property\n"); return ret; } - trip->hysteresis = prop; + trip->hysteresis = hysteresis; ret = thermal_of_get_trip_type(np, &trip->type); if (ret < 0) { @@ -349,10 +350,10 @@ out: } /** - * thermal_of_zone_unregister - Cleanup the specific allocated ressources + * thermal_of_zone_unregister - Cleanup the specific allocated resources * * This function disables the thermal zone and frees the different - * ressources allocated specific to the thermal OF. + * resources allocated specific to the thermal OF. * * @tz: a pointer to the thermal zone structure */ @@ -561,7 +562,7 @@ thermal_of_cooling_device_register(struct device_node *np, u32 cdev_id, cdev->np = np; cdev->cdev_id = cdev_id; - ret = thermal_cooling_device_add(cdev, devdata); + ret = thermal_cooling_device_add(cdev, NULL, devdata); if (ret) return ERR_PTR(ret); diff --git a/drivers/thermal/ti-soc-thermal/ti-bandgap.c b/drivers/thermal/ti-soc-thermal/ti-bandgap.c index ba43399d0b38..5c028d25a565 100644 --- a/drivers/thermal/ti-soc-thermal/ti-bandgap.c +++ b/drivers/thermal/ti-soc-thermal/ti-bandgap.c @@ -332,7 +332,7 @@ static inline int ti_bandgap_validate(struct ti_bandgap *bgp, int id) * ti_bandgap_read_counter() - read the sensor counter * @bgp: pointer to bandgap instance * @id: sensor id - * @interval: resulting update interval in miliseconds + * @interval: resulting update interval in milliseconds */ static void ti_bandgap_read_counter(struct ti_bandgap *bgp, int id, int *interval) @@ -352,7 +352,7 @@ static void ti_bandgap_read_counter(struct ti_bandgap *bgp, int id, * ti_bandgap_read_counter_delay() - read the sensor counter delay * @bgp: pointer to bandgap instance * @id: sensor id - * @interval: resulting update interval in miliseconds + * @interval: resulting update interval in milliseconds */ static void ti_bandgap_read_counter_delay(struct ti_bandgap *bgp, int id, int *interval) @@ -394,7 +394,7 @@ static void ti_bandgap_read_counter_delay(struct ti_bandgap *bgp, int id, * ti_bandgap_read_update_interval() - read the sensor update interval * @bgp: pointer to bandgap instance * @id: sensor id - * @interval: resulting update interval in miliseconds + * @interval: resulting update interval in milliseconds * * Return: 0 on success or the proper error code */ @@ -427,7 +427,7 @@ exit: * ti_bandgap_write_counter_delay() - set the counter_delay * @bgp: pointer to bandgap instance * @id: sensor id - * @interval: desired update interval in miliseconds + * @interval: desired update interval in milliseconds * * Return: 0 on success or the proper error code */ @@ -471,7 +471,7 @@ static int ti_bandgap_write_counter_delay(struct ti_bandgap *bgp, int id, * ti_bandgap_write_counter() - set the bandgap sensor counter * @bgp: pointer to bandgap instance * @id: sensor id - * @interval: desired update interval in miliseconds + * @interval: desired update interval in milliseconds */ static void ti_bandgap_write_counter(struct ti_bandgap *bgp, int id, u32 interval) @@ -486,7 +486,7 @@ static void ti_bandgap_write_counter(struct ti_bandgap *bgp, int id, * ti_bandgap_write_update_interval() - set the update interval * @bgp: pointer to bandgap instance * @id: sensor id - * @interval: desired update interval in miliseconds + * @interval: desired update interval in milliseconds * * Return: 0 on success or the proper error code */ diff --git a/drivers/thunderbolt/acpi.c b/drivers/thunderbolt/acpi.c index 53546bc477a5..be91bc604900 100644 --- a/drivers/thunderbolt/acpi.c +++ b/drivers/thunderbolt/acpi.c @@ -17,8 +17,8 @@ static acpi_status tb_acpi_add_link(acpi_handle handle, u32 level, void *data, struct acpi_device *adev = acpi_fetch_acpi_dev(handle); struct fwnode_handle *fwnode; struct tb_nhi *nhi = data; + struct device *dev = NULL; struct pci_dev *pdev; - struct device *dev; if (!adev) return AE_OK; @@ -37,7 +37,7 @@ static acpi_status tb_acpi_add_link(acpi_handle handle, u32 level, void *data, * USB3 ports might not even have a physical device yet if xHCI driver * isn't bound yet. */ - dev = acpi_get_first_physical_node(adev); + dev = acpi_bus_get_primary_device(adev); if (!dev || !dev_is_pci(dev)) goto out_put; @@ -74,6 +74,7 @@ static acpi_status tb_acpi_add_link(acpi_handle handle, u32 level, void *data, } out_put: + put_device(dev); fwnode_handle_put(fwnode); return AE_OK; } diff --git a/include/acpi/acpi_bus.h b/include/acpi/acpi_bus.h index a10a591c18b2..a755b15af9db 100644 --- a/include/acpi/acpi_bus.h +++ b/include/acpi/acpi_bus.h @@ -9,6 +9,8 @@ #ifndef __ACPI_BUS_H__ #define __ACPI_BUS_H__ +#ifdef CONFIG_ACPI + #include <linux/completion.h> #include <linux/container_of.h> #include <linux/device.h> @@ -58,7 +60,6 @@ bool acpi_dock_match(acpi_handle handle); bool acpi_check_dsm(acpi_handle handle, const guid_t *guid, u64 rev, u64 funcs); union acpi_object *acpi_evaluate_dsm(acpi_handle handle, const guid_t *guid, u64 rev, u64 func, union acpi_object *argv4); -#ifdef CONFIG_ACPI bool acpi_get_physical_device_location(acpi_handle handle, struct acpi_pld_info **pld); @@ -77,7 +78,6 @@ acpi_evaluate_dsm_typed(acpi_handle handle, const guid_t *guid, u64 rev, return obj; } -#endif #define ACPI_INIT_DSM_ARGV4(cnt, eles) \ { \ @@ -90,8 +90,6 @@ bool acpi_dev_found(const char *hid); bool acpi_dev_present(const char *hid, const char *uid, s64 hrv); bool acpi_reduced_hardware(void); -#ifdef CONFIG_ACPI - struct proc_dir_entry; #define ACPI_BUS_FILE_ROOT "acpi" @@ -435,9 +433,9 @@ struct acpi_device_software_nodes { /* Device */ struct acpi_device { + acpi_handle handle; /* no handle for fixed hardware */ u32 pld_crc; int device_type; - acpi_handle handle; /* no handle for fixed hardware */ struct fwnode_handle fwnode; struct list_head wakeup_list; struct list_head del_list; @@ -613,7 +611,6 @@ int acpi_bus_get_status(struct acpi_device *device); int acpi_bus_set_power(acpi_handle handle, int state); const char *acpi_power_state_string(int state); int acpi_device_set_power(struct acpi_device *device, int state); -int acpi_bus_init_power(struct acpi_device *device); int acpi_device_fix_up_power(struct acpi_device *device); void acpi_device_fix_up_power_extended(struct acpi_device *adev); void acpi_device_fix_up_power_children(struct acpi_device *adev); @@ -665,7 +662,7 @@ struct acpi_bus_type { int register_acpi_bus_type(struct acpi_bus_type *); int unregister_acpi_bus_type(struct acpi_bus_type *); int acpi_bind_one(struct device *dev, struct acpi_device *adev); -int acpi_unbind_one(struct device *dev); +void acpi_unbind_one(struct device *dev); enum acpi_bridge_type { ACPI_BRIDGE_TYPE_PCIE = 1, @@ -830,7 +827,15 @@ static inline bool acpi_str_uid_match(struct acpi_device *adev, const char *uid2 { const char *uid1 = acpi_device_uid(adev); - return uid1 && uid2 && !strcmp(uid1, uid2); + if (!uid1 || !uid2) + return false; + + if (*uid1 == '\\' && uid1[1]) + uid1++; + if (*uid2 == '\\' && uid2[1]) + uid2++; + + return !strcmp(uid1, uid2); } static inline bool acpi_int_uid_match(struct acpi_device *adev, u64 uid2) @@ -857,6 +862,10 @@ static inline bool acpi_int_uid_match(struct acpi_device *adev, u64 uid2) * * Matches UID in @adev with given @uid2. * + * If both the UID in @adev and @uid2 are strings, they are compared + * after optionally skipping a leading backslash ('\') if the given + * string contains additional characters. + * * Returns: %true if matches, %false otherwise. */ #define acpi_dev_uid_match(adev, uid2) \ @@ -942,8 +951,21 @@ int acpi_wait_for_acpi_ipmi(void); int acpi_scan_add_dep(acpi_handle handle, struct acpi_handle_list *dep_devices); u32 arch_acpi_add_auto_dep(acpi_handle handle); + #else /* CONFIG_ACPI */ +static inline struct acpi_device * +acpi_find_child_device(struct acpi_device *parent, u64 address, + bool check_children) +{ + return NULL; +} + +static inline bool acpi_has_method(acpi_handle handle, char *name) +{ + return false; +} + static inline struct device *acpi_bus_get_primary_device(struct acpi_device *adev) { return NULL; @@ -959,6 +981,11 @@ static inline const char *acpi_device_hid(struct acpi_device *device) return ""; } +static inline char *acpi_device_uid(struct acpi_device *device) +{ + return NULL; +} + static inline bool acpi_get_physical_device_location(acpi_handle handle, struct acpi_pld_info **pld) { diff --git a/include/acpi/cppc_acpi.h b/include/acpi/cppc_acpi.h index 94a6277edab2..3f0005abac64 100644 --- a/include/acpi/cppc_acpi.h +++ b/include/acpi/cppc_acpi.h @@ -72,11 +72,16 @@ struct cpc_register_resource { struct { struct cpc_reg reg; bool use_rmw_lock; + bool read_unsupported; + bool write_unsupported; }; u64 int_value; } cpc_entry; }; +struct cpc_sysmem_node; +struct cpc_non_mmio_node; + /* Container to hold the CPC details for each CPU */ struct cpc_desc { int num_entries; @@ -84,10 +89,12 @@ struct cpc_desc { int cpu_id; int write_cmd_status; int write_cmd_id; - /* Lock used for RMW operations in cpc_write() */ + /* Serialize partial SystemMemory writes within this descriptor. */ raw_spinlock_t rmw_lock; struct cpc_register_resource cpc_regs[MAX_CPC_REG_ENT]; struct acpi_psd_package domain_info; + struct cpc_sysmem_node *sysmem_nodes; + struct cpc_non_mmio_node *non_mmio_nodes; struct kobject kobj; }; @@ -141,6 +148,8 @@ struct cppc_perf_ctrls { u32 desired_perf; u32 energy_perf; bool auto_sel; + /* Allow an explicit zero minimum; otherwise zero omits the update. */ + bool min_perf_valid; }; struct cppc_perf_fb_ctrs { @@ -188,6 +197,7 @@ extern int cppc_set_epp(int cpu, u64 epp_val); extern int cppc_get_auto_act_window(int cpu, u64 *auto_act_window); extern int cppc_set_auto_act_window(int cpu, u64 auto_act_window); extern int cppc_get_auto_sel(int cpu, bool *enable); +bool cppc_auto_sel_is_immutable(int cpu); extern int cppc_set_auto_sel(int cpu, bool enable); extern int cppc_get_perf_limited(int cpu, u64 *perf_limited); extern int cppc_set_perf_limited(int cpu, u64 bits_to_clear); @@ -289,6 +299,12 @@ static inline int cppc_get_auto_sel(int cpu, bool *enable) { return -EOPNOTSUPP; } + +static inline bool cppc_auto_sel_is_immutable(int cpu) +{ + return false; +} + static inline int cppc_set_auto_sel(int cpu, bool enable) { return -EOPNOTSUPP; diff --git a/include/acpi/ghes.h b/include/acpi/ghes.h index 8d7e5caef3f1..7acf209061ea 100644 --- a/include/acpi/ghes.h +++ b/include/acpi/ghes.h @@ -85,6 +85,10 @@ int devm_ghes_register_vendor_record_notifier(struct device *dev, struct list_head *ghes_get_devices(void); void ghes_estatus_pool_region_free(unsigned long addr, u32 size); + +struct cxl_cper_sec_prot_err; +void cxl_cper_post_prot_err(struct cxl_cper_sec_prot_err *prot_err, + int severity, u32 len); #else static inline struct list_head *ghes_get_devices(void) { return NULL; } diff --git a/include/acpi/processor.h b/include/acpi/processor.h index 554be224ce76..b5447af4d40b 100644 --- a/include/acpi/processor.h +++ b/include/acpi/processor.h @@ -427,10 +427,8 @@ int acpi_processor_ffh_lpi_enter(struct acpi_lpi_state *lpi); #endif /* CONFIG_ACPI_PROCESSOR_IDLE */ /* in processor_thermal.c */ -int acpi_processor_thermal_init(struct acpi_processor *pr, - struct acpi_device *device); -void acpi_processor_thermal_exit(struct acpi_processor *pr, - struct acpi_device *device); +int acpi_processor_thermal_init(struct acpi_processor *pr); +void acpi_processor_thermal_exit(struct acpi_processor *pr); extern const struct thermal_cooling_device_ops processor_cooling_ops; #ifdef CONFIG_CPU_FREQ void acpi_thermal_cpufreq_init(struct cpufreq_policy *policy); diff --git a/include/cxl/event.h b/include/cxl/event.h index b5673384d930..a9f5c2381cc9 100644 --- a/include/cxl/event.h +++ b/include/cxl/event.h @@ -312,13 +312,13 @@ static inline int cxl_cper_prot_err_kfifo_get(struct cxl_cper_prot_err_work_data #endif #ifdef CONFIG_ACPI_APEI_PCIEAER -int cxl_cper_sec_prot_err_valid(struct cxl_cper_sec_prot_err *prot_err); +int cxl_cper_sec_prot_err_valid(struct cxl_cper_sec_prot_err *prot_err, u32 len); int cxl_cper_setup_prot_err_work_data(struct cxl_cper_prot_err_work_data *wd, struct cxl_cper_sec_prot_err *prot_err, int severity); #else static inline int -cxl_cper_sec_prot_err_valid(struct cxl_cper_sec_prot_err *prot_err) +cxl_cper_sec_prot_err_valid(struct cxl_cper_sec_prot_err *prot_err, u32 len) { return -EOPNOTSUPP; } @@ -331,6 +331,4 @@ cxl_cper_setup_prot_err_work_data(struct cxl_cper_prot_err_work_data *wd, } #endif -void cxl_cper_handle_prot_err(struct cxl_cper_prot_err_work_data *wd); - #endif /* _LINUX_CXL_EVENT_H */ diff --git a/include/linux/acpi.h b/include/linux/acpi.h index ddacac812094..aef0a01f4cfa 100644 --- a/include/linux/acpi.h +++ b/include/linux/acpi.h @@ -25,17 +25,19 @@ struct irq_domain_ops; #define _LINUX #endif #include <acpi/acpi.h> +#include <acpi/acpi_bus.h> #include <acpi/acpi_numa.h> #ifdef CONFIG_ACPI +DEFINE_FREE(acpi_object_free, union acpi_object *, if (_T) ACPI_FREE(_T)); + #include <linux/list.h> #include <linux/dynamic_debug.h> #include <linux/module.h> #include <linux/mutex.h> #include <linux/fw_table.h> -#include <acpi/acpi_bus.h> #include <acpi/acpi_drivers.h> #include <acpi/acpi_io.h> #include <asm/acpi.h> @@ -434,7 +436,7 @@ extern acpi_handle ec_get_handle(void); extern bool acpi_is_pnp_device(struct acpi_device *); -#if defined(CONFIG_ACPI_WMI) || defined(CONFIG_ACPI_WMI_MODULE) +#if IS_ENABLED(CONFIG_ACPI_WMI) typedef void (*wmi_notify_handler) (union acpi_object *data, void *context); @@ -1305,34 +1307,34 @@ void __acpi_handle_debug(struct _ddebug *descriptor, acpi_handle handle, const c #endif /* - * acpi_handle_<level>: Print message with ACPI prefix and object path + * acpi_handle_<level> - Print a message with ACPI prefix and object path * - * These interfaces acquire the global namespace mutex to obtain an object - * path. In interrupt context, it shows the object path as <n/a>. + * In thread context, the global namespace mutex is acquired to obtain the + * object path. In interrupt context, the object path is shown as <n/a>. */ #define acpi_handle_emerg(handle, fmt, ...) \ - acpi_handle_printk(KERN_EMERG, handle, fmt, ##__VA_ARGS__) + acpi_handle_printk(KERN_EMERG, handle, dev_fmt(fmt), ##__VA_ARGS__) #define acpi_handle_alert(handle, fmt, ...) \ - acpi_handle_printk(KERN_ALERT, handle, fmt, ##__VA_ARGS__) + acpi_handle_printk(KERN_ALERT, handle, dev_fmt(fmt), ##__VA_ARGS__) #define acpi_handle_crit(handle, fmt, ...) \ - acpi_handle_printk(KERN_CRIT, handle, fmt, ##__VA_ARGS__) + acpi_handle_printk(KERN_CRIT, handle, dev_fmt(fmt), ##__VA_ARGS__) #define acpi_handle_err(handle, fmt, ...) \ - acpi_handle_printk(KERN_ERR, handle, fmt, ##__VA_ARGS__) + acpi_handle_printk(KERN_ERR, handle, dev_fmt(fmt), ##__VA_ARGS__) #define acpi_handle_warn(handle, fmt, ...) \ - acpi_handle_printk(KERN_WARNING, handle, fmt, ##__VA_ARGS__) + acpi_handle_printk(KERN_WARNING, handle, dev_fmt(fmt), ##__VA_ARGS__) #define acpi_handle_notice(handle, fmt, ...) \ - acpi_handle_printk(KERN_NOTICE, handle, fmt, ##__VA_ARGS__) + acpi_handle_printk(KERN_NOTICE, handle, dev_fmt(fmt), ##__VA_ARGS__) #define acpi_handle_info(handle, fmt, ...) \ - acpi_handle_printk(KERN_INFO, handle, fmt, ##__VA_ARGS__) + acpi_handle_printk(KERN_INFO, handle, dev_fmt(fmt), ##__VA_ARGS__) #if defined(DEBUG) #define acpi_handle_debug(handle, fmt, ...) \ - acpi_handle_printk(KERN_DEBUG, handle, fmt, ##__VA_ARGS__) + acpi_handle_printk(KERN_DEBUG, handle, dev_fmt(fmt), ##__VA_ARGS__) #else #if defined(CONFIG_DYNAMIC_DEBUG) #define acpi_handle_debug(handle, fmt, ...) \ _dynamic_func_call(fmt, __acpi_handle_debug, \ - handle, pr_fmt(fmt), ##__VA_ARGS__) + handle, dev_fmt(fmt), ##__VA_ARGS__) #else #define acpi_handle_debug(handle, fmt, ...) \ ({ \ diff --git a/include/linux/apple-gmux.h b/include/linux/apple-gmux.h index 206d97ffda79..d7939c3a08fd 100644 --- a/include/linux/apple-gmux.h +++ b/include/linux/apple-gmux.h @@ -109,7 +109,7 @@ static inline bool apple_gmux_detect(struct pnp_dev *pnp_dev, enum apple_gmux_ty if (!adev) return false; - dev = get_device(acpi_get_first_physical_node(adev)); + dev = acpi_bus_get_primary_device(adev); acpi_dev_put(adev); if (!dev) return false; diff --git a/include/linux/cpufreq.h b/include/linux/cpufreq.h index 35ce665edfd8..d3d0d9d02aa4 100644 --- a/include/linux/cpufreq.h +++ b/include/linux/cpufreq.h @@ -420,6 +420,9 @@ struct cpufreq_driver { /* Will be called after the driver is fully initialized */ void (*ready)(struct cpufreq_policy *policy); + /* Return the capacity reference frequency for policy. */ + unsigned int (*scale_freq_ref)(struct cpufreq_policy *policy); + struct freq_attr **attr; /* platform specific boost support code */ diff --git a/include/linux/device/bus.h b/include/linux/device/bus.h index a38f7229b8f4..89e156b9acb5 100644 --- a/include/linux/device/bus.h +++ b/include/linux/device/bus.h @@ -113,6 +113,7 @@ struct bus_type { bool need_parent_lock; }; +int __must_check companion_bus_register(const struct bus_type *bus); int __must_check bus_register(const struct bus_type *bus); void bus_unregister(const struct bus_type *bus); diff --git a/include/linux/pm.h b/include/linux/pm.h index afcaaa37a812..ef3f1310e749 100644 --- a/include/linux/pm.h +++ b/include/linux/pm.h @@ -663,6 +663,99 @@ struct pm_subsys_data { #define DPM_FLAG_SMART_SUSPEND BIT(2) #define DPM_FLAG_MAY_SKIP_RESUME BIT(3) +/** + * struct dev_pm_info - Device power management information. + * + * @power_state: Legacy power state (mostly unused in modern kernels). + * @can_wakeup: Device is capable of generating wakeup signals. + * @async_suspend: Device can be suspended and resumed asynchronously. + * @in_dpm_list: Device is on the dpm_list. + * @is_prepared: Device's ->prepare() callback has run successfully. + * @is_suspended: Device is suspended during a system sleep transition. + * @is_noirq_suspended: Device's noirq suspend callback has run successfully. + * @is_late_suspended: Device's late suspend callback has run successfully. + * @no_pm: Device does not participate in power management transitions. + * @early_init: Device was initialized before standard PM initialization. + * @direct_complete: Device can skip suspend/resume callbacks and remain + * runtime-suspended during system sleep. + * @driver_flags: Driver flags (e.g. %DPM_FLAG_SMART_SUSPEND) set at probe time. + * @lock: Spinlock used for synchronizing PM state transitions and runtime PM + * operations. + * @entry: List head for device power management lists. + * @completion: Completion for synchronization during asynchronous system + * suspend/resume. + * @wakeup: Wakeup source object associated with the device. + * @work_in_progress: Asynchronous PM operation in progress. + * @wakeup_path: Device is in the wakeup path or can wake the system up. + * @syscore: Device participates in syscore power management operations. + * @no_pm_callbacks: Device has no PM callbacks; handled by parent or subsystem. + * @smart_suspend: Driver requested smart-suspend behavior. + * @must_resume: Device must be resumed during system resume. + * @may_skip_resume: Set by subsystems to indicate driver resume callbacks may + * be skipped. + * @out_band_wakeup: Out-of-band wakeup is supported. + * @strict_midlayer: Middle layer code does not want callbacks invoked via + * pm_runtime_force_suspend() / pm_runtime_force_resume(). + * @should_wakeup: Wakeup flag when system sleep is not enabled. + * @suspend_timer: High-resolution timer used for scheduling delayed runtime + * suspend and autosuspend requests. + * @timer_expires: Timer expiration time in nanoseconds monotonic time + * (runtime PM). + * @work: Work structure used for queuing up requests into pm_wq (runtime PM). + * @wait_queue: Wait queue used if any helper functions need to wait for another + * state change to complete (runtime PM). + * @wakeirq: Dedicated wakeup interrupt for the device. + * @usage_count: Device runtime PM usage counter. + * @child_count: Count of active children of the device (runtime PM). + * @disable_depth: Disable counter for runtime PM (runtime PM is enabled when + * this is 0; initial value is 1). + * @idle_notification: Set if ->runtime_idle() is being executed. + * @request_pending: Set if a work item is queued into pm_wq (runtime PM). + * @deferred_resume: Set if ->runtime_resume() should run as soon as + * ->runtime_suspend() completes. + * @needs_force_resume: Indicates the device was forced into suspend by + * pm_runtime_force_suspend() and must be resumed by + * pm_runtime_force_resume(). + * @runtime_auto: User space has allowed the driver to power manage the device + * at runtime via sysfs control attribute; also can be set by + * pm_runtime_allow() or pm_runtime_forbid(). + * @ignore_children: If set, the value of child_count is ignored for runtime + * suspend and idle decisions. + * @no_callbacks: Indicates the device does not use runtime PM callbacks. + * @irq_safe: Indicates runtime PM callbacks will be invoked with the spinlock + * held and interrupts disabled. + * @use_autosuspend: Indicates the device driver supports delayed runtime + * autosuspend. + * @timer_autosuspends: Indicates the runtime PM core should attempt an + * autosuspend rather than a normal suspend when the timer expires. + * @memalloc_noio: Indicates memory allocation during runtime PM transitions + * must avoid I/O (GFP_NOIO). + * @links_count: Number of device links that require runtime PM coordination. + * @request: Type of pending runtime PM request (valid if request_pending is + * set). + * @runtime_status: Runtime PM status of the device. + * @last_status: Last status captured before disabling runtime PM, or + * %RPM_BLOCKED / %RPM_INVALID. + * @runtime_error: Fatal error code returned by a failing callback, blocking + * helpers until cleared. + * @autosuspend_delay: Delay time in milliseconds to be used for runtime + * autosuspend. + * @last_busy: Timestamp in nanoseconds when pm_runtime_mark_last_busy() was + * last called. Used in calculating inactivity periods for autosuspend. + * @active_time: Accumulated time in nanoseconds spent in %RPM_ACTIVE state. + * @suspended_time: Accumulated time in nanoseconds spent in %RPM_SUSPENDED + * state. + * @accounting_timestamp: Timestamp in nanoseconds of the last runtime PM state + * accounting update. + * @subsys_data: Subsystem-specific power management data. + * @set_latency_tolerance: Callback for setting latency tolerance. + * @qos: Per-device PM Quality of Service (QoS) constraints. + * @detach_power_off: Indicates device should be detached from PM domain on + * power off. + * + * Device power management information stored in the "power" member of struct + * device. + */ struct dev_pm_info { pm_message_t power_state; bool can_wakeup:1; diff --git a/include/linux/pm_runtime.h b/include/linux/pm_runtime.h index 64921b10ac74..322e3b17f987 100644 --- a/include/linux/pm_runtime.h +++ b/include/linux/pm_runtime.h @@ -137,13 +137,14 @@ static inline void pm_runtime_put_noidle(struct device *dev) * pm_runtime_suspended - Check whether or not a device is runtime-suspended. * @dev: Target device. * - * Return %true if runtime PM is enabled for @dev and its runtime PM status is - * %RPM_SUSPENDED, or %false otherwise. - * * Note that the return value of this function can only be trusted if it is * called under the runtime PM lock of @dev or under conditions in which * runtime PM cannot be either disabled or enabled for @dev and its runtime PM * status cannot change. + * + * Return: + * * %true: @dev has runtime PM enabled and its status is %RPM_SUSPENDED. + * * %false: Otherwise. */ static inline bool pm_runtime_suspended(struct device *dev) { @@ -155,13 +156,14 @@ static inline bool pm_runtime_suspended(struct device *dev) * pm_runtime_active - Check whether or not a device is runtime-active. * @dev: Target device. * - * Return %true if runtime PM is disabled for @dev or its runtime PM status is - * %RPM_ACTIVE, or %false otherwise. - * * Note that the return value of this function can only be trusted if it is * called under the runtime PM lock of @dev or under conditions in which * runtime PM cannot be either disabled or enabled for @dev and its runtime PM * status cannot change. + * + * Return: + * * %true: Runtime PM is disabled for @dev or its status is %RPM_ACTIVE. + * * %false: Otherwise. */ static inline bool pm_runtime_active(struct device *dev) { @@ -173,12 +175,13 @@ static inline bool pm_runtime_active(struct device *dev) * pm_runtime_status_suspended - Check if runtime PM status is "suspended". * @dev: Target device. * - * Return %true if the runtime PM status of @dev is %RPM_SUSPENDED, or %false - * otherwise, regardless of whether or not runtime PM has been enabled for @dev. - * * Note that the return value of this function can only be trusted if it is * called under the runtime PM lock of @dev or under conditions in which the * runtime PM status of @dev cannot change. + * + * Return: + * * %true: Runtime PM status of @dev is %RPM_SUSPENDED. + * * %false: Otherwise. */ static inline bool pm_runtime_status_suspended(struct device *dev) { @@ -189,11 +192,13 @@ static inline bool pm_runtime_status_suspended(struct device *dev) * pm_runtime_enabled - Check if runtime PM is enabled. * @dev: Target device. * - * Return %true if runtime PM is enabled for @dev or %false otherwise. - * * Note that the return value of this function can only be trusted if it is * called under the runtime PM lock of @dev or under conditions in which * runtime PM cannot be either disabled or enabled for @dev. + * + * Return: + * * %true: Runtime PM is enabled for @dev. + * * %false: Otherwise. */ static inline bool pm_runtime_enabled(struct device *dev) { @@ -205,6 +210,10 @@ static inline bool pm_runtime_enabled(struct device *dev) * @dev: Target device. * * Do not call this function outside system suspend/resume code paths. + * + * Return: + * * %true: Runtime PM enabling is blocked for @dev. + * * %false: Otherwise. */ static inline bool pm_runtime_blocked(struct device *dev) { @@ -215,8 +224,9 @@ static inline bool pm_runtime_blocked(struct device *dev) * pm_runtime_has_no_callbacks - Check if runtime PM callbacks may be present. * @dev: Target device. * - * Return %true if @dev is a special device without runtime PM callbacks or - * %false otherwise. + * Return: + * * %true: @dev is marked as having no runtime PM callbacks. + * * %false: Otherwise. */ static inline bool pm_runtime_has_no_callbacks(struct device *dev) { @@ -239,9 +249,11 @@ static inline void pm_runtime_mark_last_busy(struct device *dev) * pm_runtime_is_irq_safe - Check if runtime PM can work in interrupt context. * @dev: Target device. * - * Return %true if @dev has been marked as an "IRQ-safe" device (with respect - * to runtime PM), in which case its runtime PM callabcks can be expected to - * work correctly when invoked from interrupt handlers. + * Return: + * * %true: @dev has been marked as an "IRQ-safe" device, in which case its + * runtime PM callbacks can be expected to work correctly from interrupt + * handlers. + * * %false: Otherwise. */ static inline bool pm_runtime_is_irq_safe(struct device *dev) { @@ -340,25 +352,25 @@ static inline int pm_runtime_force_resume(struct device *dev) { return -ENXIO; } #endif /* CONFIG_PM_SLEEP */ /** - * pm_runtime_idle - Conditionally set up autosuspend of a device or suspend it. + * pm_runtime_idle - Conditionally initiate autosuspend of a device or suspend it. * @dev: Target device. * * Invoke the "idle check" callback of @dev and, depending on its return value, - * set up autosuspend of @dev or suspend it (depending on whether or not + * initiate autosuspend of @dev or suspend it (depending on whether or not * autosuspend has been enabled for it). * * Return: - * * 0: Success. - * * -EINVAL: Runtime PM error. - * * -EACCES: Runtime PM disabled. - * * -EAGAIN: Runtime PM usage counter non-zero, Runtime PM status change - * ongoing or device not in %RPM_ACTIVE state. - * * -EBUSY: Runtime PM child_count non-zero. - * * -EPERM: Device PM QoS resume latency 0. - * * -EINPROGRESS: Suspend already in progress. - * * -ENOSYS: CONFIG_PM not enabled. - * Other values and conditions for the above values are possible as returned by - * Runtime PM idle and suspend callbacks. + * * %0: Success. + * * %-EINVAL: Runtime PM error. + * * %-EACCES: Runtime PM disabled. + * * %-EAGAIN: Runtime PM usage counter non-zero, Runtime PM status change + * ongoing or device not in %RPM_ACTIVE state. + * * %-EBUSY: Runtime PM child_count non-zero. + * * %-EPERM: Device PM QoS resume latency 0. + * * %-EINPROGRESS: Suspend already in progress. + * * %-ENOSYS: %CONFIG_PM not enabled. + * * Other values and conditions for the above values are possible as returned + * by Runtime PM idle and suspend callbacks. */ static inline int pm_runtime_idle(struct device *dev) { @@ -370,17 +382,17 @@ static inline int pm_runtime_idle(struct device *dev) * @dev: Target device. * * Return: - * * 1: Success; device was already suspended. - * * 0: Success. - * * -EINVAL: Runtime PM error. - * * -EACCES: Runtime PM disabled. - * * -EAGAIN: Runtime PM usage counter non-zero or Runtime PM status change - * ongoing. - * * -EBUSY: Runtime PM child_count non-zero. - * * -EPERM: Device PM QoS resume latency 0. - * * -ENOSYS: CONFIG_PM not enabled. - * Other values and conditions for the above values are possible as returned by - * Runtime PM suspend callbacks. + * * %1: Success; device was already suspended. + * * %0: Success. + * * %-EINVAL: Runtime PM error. + * * %-EACCES: Runtime PM disabled. + * * %-EAGAIN: Runtime PM usage counter non-zero or Runtime PM status change + * ongoing. + * * %-EBUSY: Runtime PM child_count non-zero. + * * %-EPERM: Device PM QoS resume latency 0. + * * %-ENOSYS: %CONFIG_PM not enabled. + * * Other values and conditions for the above values are possible as returned + * by Runtime PM suspend callbacks. */ static inline int pm_runtime_suspend(struct device *dev) { @@ -388,26 +400,26 @@ static inline int pm_runtime_suspend(struct device *dev) } /** - * pm_runtime_autosuspend - Update the last access time and set up autosuspend + * pm_runtime_autosuspend - Update the last access time and initiate autosuspend * of a device. * @dev: Target device. * - * First update the last access time, then set up autosuspend of @dev or suspend - * it (depending on whether or not autosuspend is enabled for it) without - * engaging its "idle check" callback. + * First update the last access time, then initiate autosuspend of @dev or + * suspend it (depending on whether or not autosuspend is enabled for it) + * without engaging its "idle check" callback. * * Return: - * * 1: Success; device was already suspended. - * * 0: Success. - * * -EINVAL: Runtime PM error. - * * -EACCES: Runtime PM disabled. - * * -EAGAIN: Runtime PM usage counter non-zero or Runtime PM status change - * ongoing. - * * -EBUSY: Runtime PM child_count non-zero. - * * -EPERM: Device PM QoS resume latency 0. - * * -ENOSYS: CONFIG_PM not enabled. - * Other values and conditions for the above values are possible as returned by - * Runtime PM suspend callbacks. + * * %1: Success; device was already suspended. + * * %0: Success. + * * %-EINVAL: Runtime PM error. + * * %-EACCES: Runtime PM disabled. + * * %-EAGAIN: Runtime PM usage counter non-zero or Runtime PM status change + * ongoing. + * * %-EBUSY: Runtime PM child_count non-zero. + * * %-EPERM: Device PM QoS resume latency 0. + * * %-ENOSYS: %CONFIG_PM not enabled. + * * Other values and conditions for the above values are possible as returned + * by Runtime PM suspend callbacks. */ static inline int pm_runtime_autosuspend(struct device *dev) { @@ -418,6 +430,11 @@ static inline int pm_runtime_autosuspend(struct device *dev) /** * pm_runtime_resume - Resume a device synchronously. * @dev: Target device. + * + * Return: + * * %1: Success; @dev is already %RPM_ACTIVE. + * * %0: Success. + * * Error code on failure. */ static inline int pm_runtime_resume(struct device *dev) { @@ -425,22 +442,22 @@ static inline int pm_runtime_resume(struct device *dev) } /** - * pm_request_idle - Queue up "idle check" execution for a device. + * pm_request_idle - Request an asynchronous idle check for a device. * @dev: Target device. * - * Queue up a work item to run an equivalent of pm_runtime_idle() for @dev - * asynchronously. + * Asynchronously request the PM core to evaluate whether @dev can be idled + * or suspended, invoking its ->runtime_idle() callback if provided. * * Return: - * * 0: Success. - * * -EINVAL: Runtime PM error. - * * -EACCES: Runtime PM disabled. - * * -EAGAIN: Runtime PM usage counter non-zero, Runtime PM status change - * ongoing or device not in %RPM_ACTIVE state. - * * -EBUSY: Runtime PM child_count non-zero. - * * -EPERM: Device PM QoS resume latency 0. - * * -EINPROGRESS: Suspend already in progress. - * * -ENOSYS: CONFIG_PM not enabled. + * * %0: Success. + * * %-EINVAL: Runtime PM error. + * * %-EACCES: Runtime PM disabled. + * * %-EAGAIN: Runtime PM usage counter non-zero, Runtime PM status change + * ongoing or device not in %RPM_ACTIVE state. + * * %-EBUSY: Runtime PM child_count non-zero. + * * %-EPERM: Device PM QoS resume latency 0. + * * %-EINPROGRESS: Suspend already in progress. + * * %-ENOSYS: %CONFIG_PM not enabled. */ static inline int pm_request_idle(struct device *dev) { @@ -448,8 +465,16 @@ static inline int pm_request_idle(struct device *dev) } /** - * pm_request_resume - Queue up runtime-resume of a device. + * pm_request_resume - Request an asynchronous runtime resume for a device. * @dev: Target device. + * + * Asynchronously request the PM core to resume @dev to %RPM_ACTIVE state + * without modifying its usage counter. + * + * Return: + * * %1: Success; @dev is already %RPM_ACTIVE. + * * %0: Success. + * * Error code on failure. */ static inline int pm_request_resume(struct device *dev) { @@ -457,24 +482,23 @@ static inline int pm_request_resume(struct device *dev) } /** - * pm_request_autosuspend - Update the last access time and queue up autosuspend - * of a device. + * pm_request_autosuspend - Update access time and request delayed suspension. * @dev: Target device. * - * Update the last access time of a device and queue up a work item to run an - * equivalent pm_runtime_autosuspend() for @dev asynchronously. + * Update the last access time of @dev and asynchronously request the PM core + * to suspend it after the autosuspend delay has elapsed. * * Return: - * * 1: Success; device was already suspended. - * * 0: Success. - * * -EINVAL: Runtime PM error. - * * -EACCES: Runtime PM disabled. - * * -EAGAIN: Runtime PM usage counter non-zero or Runtime PM status change - * ongoing. - * * -EBUSY: Runtime PM child_count non-zero. - * * -EPERM: Device PM QoS resume latency 0. - * * -EINPROGRESS: Suspend already in progress. - * * -ENOSYS: CONFIG_PM not enabled. + * * %1: Success; device was already suspended. + * * %0: Success. + * * %-EINVAL: Runtime PM error. + * * %-EACCES: Runtime PM disabled. + * * %-EAGAIN: Runtime PM usage counter non-zero or Runtime PM status change + * ongoing. + * * %-EBUSY: Runtime PM child_count non-zero. + * * %-EPERM: Device PM QoS resume latency 0. + * * %-EINPROGRESS: Suspend already in progress. + * * %-ENOSYS: %CONFIG_PM not enabled. */ static inline int pm_request_autosuspend(struct device *dev) { @@ -483,11 +507,16 @@ static inline int pm_request_autosuspend(struct device *dev) } /** - * pm_runtime_get - Bump up usage counter and queue up resume of a device. + * pm_runtime_get - Increment usage counter and request asynchronous resume. * @dev: Target device. * - * Bump up the runtime PM usage counter of @dev and queue up a work item to - * carry out runtime-resume of it. + * Increment the runtime PM usage counter of @dev and, if the device is + * currently suspended, asynchronously request the PM core to resume it. + * + * Return: + * * %1: Success; @dev is already %RPM_ACTIVE. + * * %0: Success; runtime-resume was queued. + * * Error code on failure. */ static inline int pm_runtime_get(struct device *dev) { @@ -501,12 +530,15 @@ static inline int pm_runtime_get(struct device *dev) * Bump up the runtime PM usage counter of @dev and carry out runtime-resume of * it synchronously. * - * The possible return values of this function are the same as for - * pm_runtime_resume() and the runtime PM usage counter of @dev remains - * incremented in all cases, even if it returns an error code. - * Consider using pm_runtime_resume_and_get() instead of it, especially - * if its return value is checked by the caller, as this is likely to result - * in cleaner code. + * Note that the runtime PM usage counter of @dev remains incremented in all + * cases, even if it returns an error code. Consider using + * pm_runtime_resume_and_get() instead, especially if the return value is + * checked by the caller, as this is likely to result in cleaner code. + * + * Return: + * * %1: Success; @dev is already %RPM_ACTIVE. + * * %0: Success. + * * Error code on failure. */ static inline int pm_runtime_get_sync(struct device *dev) { @@ -531,8 +563,11 @@ static inline int pm_runtime_get_active(struct device *dev, int rpmflags) * @dev: Target device. * * Resume @dev synchronously and if that is successful, increment its runtime - * PM usage counter. Return 0 if the runtime PM usage counter of @dev has been - * incremented or a negative error code otherwise. + * PM usage counter. + * + * Return: + * * %0: Success; @dev is active and its usage counter has been incremented. + * * Negative error code on failure; usage counter is unchanged. */ static inline int pm_runtime_resume_and_get(struct device *dev) { @@ -540,11 +575,12 @@ static inline int pm_runtime_resume_and_get(struct device *dev) } /** - * pm_runtime_put - Drop device usage counter and queue up "idle check" if 0. + * pm_runtime_put - Drop device usage counter and request asynchronous idle check. * @dev: Target device. * - * Decrement the runtime PM usage counter of @dev and if it turns out to be - * equal to 0, queue up a work item for @dev like in pm_request_idle(). + * Decrement the runtime PM usage counter of @dev. If the counter reaches zero + * and the device has no active child dependencies, asynchronously request the + * PM core to idle or suspend the device. */ static inline void pm_runtime_put(struct device *dev) { @@ -559,16 +595,16 @@ static inline void pm_runtime_put(struct device *dev) * equal to 0, queue up a work item for @dev like in pm_request_autosuspend(). * * Return: - * * 1: Success. Usage counter dropped to zero, but device was already suspended. - * * 0: Success. - * * -EINVAL: Runtime PM error. - * * -EACCES: Runtime PM disabled. - * * -EAGAIN: Runtime PM usage counter became non-zero or Runtime PM status - * change ongoing. - * * -EBUSY: Runtime PM child_count non-zero. - * * -EPERM: Device PM QoS resume latency 0. - * * -EINPROGRESS: Suspend already in progress. - * * -ENOSYS: CONFIG_PM not enabled. + * * %1: Success. Usage counter dropped to zero, but device was already suspended. + * * %0: Success. + * * %-EINVAL: Runtime PM error. + * * %-EACCES: Runtime PM disabled. + * * %-EAGAIN: Runtime PM usage counter became non-zero or Runtime PM status + * change ongoing. + * * %-EBUSY: Runtime PM child_count non-zero. + * * %-EPERM: Device PM QoS resume latency 0. + * * %-EINPROGRESS: Suspend already in progress. + * * %-ENOSYS: %CONFIG_PM not enabled. */ static inline int __pm_runtime_put_autosuspend(struct device *dev) { @@ -576,25 +612,25 @@ static inline int __pm_runtime_put_autosuspend(struct device *dev) } /** - * pm_runtime_put_autosuspend - Update the last access time of a device, drop - * its usage counter and queue autosuspend if the usage counter becomes 0. + * pm_runtime_put_autosuspend - Update the last access time, drop usage counter + * and request autosuspend. * @dev: Target device. * - * Update the last access time of @dev, decrement runtime PM usage counter of - * @dev and if it turns out to be equal to 0, queue up a work item for @dev like - * in pm_request_autosuspend(). + * Update the last access time of @dev and decrement its runtime PM usage + * counter. If the counter drops to zero, asynchronously request the PM core to + * suspend the device once its autosuspend delay has elapsed. * * Return: - * * 1: Success. Usage counter dropped to zero, but device was already suspended. - * * 0: Success. - * * -EINVAL: Runtime PM error. - * * -EACCES: Runtime PM disabled. - * * -EAGAIN: Runtime PM usage counter became non-zero or Runtime PM status - * change ongoing. - * * -EBUSY: Runtime PM child_count non-zero. - * * -EPERM: Device PM QoS resume latency 0. - * * -EINPROGRESS: Suspend already in progress. - * * -ENOSYS: CONFIG_PM not enabled. + * * %1: Success. Usage counter dropped to zero, but device was already suspended. + * * %0: Success. + * * %-EINVAL: Runtime PM error. + * * %-EACCES: Runtime PM disabled. + * * %-EAGAIN: Runtime PM usage counter became non-zero or Runtime PM status + * change ongoing. + * * %-EBUSY: Runtime PM child_count non-zero. + * * %-EPERM: Device PM QoS resume latency 0. + * * %-EINPROGRESS: Suspend already in progress. + * * %-ENOSYS: %CONFIG_PM not enabled. */ static inline int pm_runtime_put_autosuspend(struct device *dev) { @@ -653,26 +689,28 @@ DEFINE_GUARD_COND(pm_runtime_active_auto, _try_enabled, * pm_runtime_put_sync - Drop device usage counter and run "idle check" if 0. * @dev: Target device. * - * Decrement the runtime PM usage counter of @dev and if it turns out to be - * equal to 0, invoke the "idle check" callback of @dev and, depending on its - * return value, set up autosuspend of @dev or suspend it (depending on whether - * or not autosuspend has been enabled for it). + * Decrement the runtime PM usage counter of @dev. If the counter drops to zero, + * synchronously evaluate and trigger idle/suspend handling. + * + * Note that this does not update the last access time, but it does respect + * existing autosuspend timers. If @dev uses autosuspend, consider using + * pm_runtime_put_sync_autosuspend() or pm_runtime_put_sync_suspend() instead. * * The runtime PM usage counter of @dev remains decremented in all cases, even * if it returns an error code. * * Return: - * * 1: Success. Usage counter dropped to zero, but device was already suspended. - * * 0: Success. - * * -EINVAL: Runtime PM error. - * * -EACCES: Runtime PM disabled. - * * -EAGAIN: Runtime PM usage counter became non-zero or Runtime PM status - * change ongoing. - * * -EBUSY: Runtime PM child_count non-zero. - * * -EPERM: Device PM QoS resume latency 0. - * * -ENOSYS: CONFIG_PM not enabled. - * Other values and conditions for the above values are possible as returned by - * Runtime PM suspend callbacks. + * * %1: Success. Usage counter dropped to zero, but device was already suspended. + * * %0: Success. + * * %-EINVAL: Runtime PM error. + * * %-EACCES: Runtime PM disabled. + * * %-EAGAIN: Runtime PM usage counter became non-zero or Runtime PM status + * change ongoing. + * * %-EBUSY: Runtime PM child_count non-zero. + * * %-EPERM: Device PM QoS resume latency 0. + * * %-ENOSYS: %CONFIG_PM not enabled. + * * Other values and conditions for the above values are possible as returned + * by Runtime PM suspend callbacks. */ static inline int pm_runtime_put_sync(struct device *dev) { @@ -683,24 +721,28 @@ static inline int pm_runtime_put_sync(struct device *dev) * pm_runtime_put_sync_suspend - Drop device usage counter and suspend if 0. * @dev: Target device. * - * Decrement the runtime PM usage counter of @dev and if it turns out to be - * equal to 0, carry out runtime-suspend of @dev synchronously. + * Decrement the runtime PM usage counter of @dev. If the counter drops to zero, + * suspend the device synchronously. + * + * This API differs from pm_runtime_put_sync() and + * pm_runtime_put_sync_autosuspend() in that it ignores any outstanding + * autosuspend delays. * * The runtime PM usage counter of @dev remains decremented in all cases, even * if it returns an error code. * * Return: - * * 1: Success. Usage counter dropped to zero, but device was already suspended. - * * 0: Success. - * * -EINVAL: Runtime PM error. - * * -EACCES: Runtime PM disabled. - * * -EAGAIN: Runtime PM usage counter became non-zero or Runtime PM status - * change ongoing. - * * -EBUSY: Runtime PM child_count non-zero. - * * -EPERM: Device PM QoS resume latency 0. - * * -ENOSYS: CONFIG_PM not enabled. - * Other values and conditions for the above values are possible as returned by - * Runtime PM suspend callbacks. + * * %1: Success. Usage counter dropped to zero, but device was already suspended. + * * %0: Success. + * * %-EINVAL: Runtime PM error. + * * %-EACCES: Runtime PM disabled. + * * %-EAGAIN: Runtime PM usage counter became non-zero or Runtime PM status + * change ongoing. + * * %-EBUSY: Runtime PM child_count non-zero. + * * %-EPERM: Device PM QoS resume latency 0. + * * %-ENOSYS: %CONFIG_PM not enabled. + * * Other values and conditions for the above values are possible as returned + * by Runtime PM suspend callbacks. */ static inline int pm_runtime_put_sync_suspend(struct device *dev) { @@ -712,27 +754,28 @@ static inline int pm_runtime_put_sync_suspend(struct device *dev) * drop device usage counter and autosuspend if 0. * @dev: Target device. * - * Update the last access time of @dev, decrement the runtime PM usage counter - * of @dev and if it turns out to be equal to 0, set up autosuspend of @dev or - * suspend it synchronously (depending on whether or not autosuspend has been - * enabled for it). + * Update the last access time of @dev and decrement its runtime PM usage + * counter. If the counter drops to zero, synchronously suspend the device (or + * schedule autosuspend if the delay has not elapsed). + * + * Prefer this API over pm_runtime_put_sync() for devices that use autosuspend. * * The runtime PM usage counter of @dev remains decremented in all cases, even * if it returns an error code. * * Return: - * * 1: Success. Usage counter dropped to zero, but device was already suspended. - * * 0: Success. - * * -EINVAL: Runtime PM error. - * * -EACCES: Runtime PM disabled. - * * -EAGAIN: Runtime PM usage counter became non-zero or Runtime PM status - * change ongoing. - * * -EBUSY: Runtime PM child_count non-zero. - * * -EPERM: Device PM QoS resume latency 0. - * * -EINPROGRESS: Suspend already in progress. - * * -ENOSYS: CONFIG_PM not enabled. - * Other values and conditions for the above values are possible as returned by - * Runtime PM suspend callbacks. + * * %1: Success. Usage counter dropped to zero, but device was already suspended. + * * %0: Success. + * * %-EINVAL: Runtime PM error. + * * %-EACCES: Runtime PM disabled. + * * %-EAGAIN: Runtime PM usage counter became non-zero or Runtime PM status + * change ongoing. + * * %-EBUSY: Runtime PM child_count non-zero. + * * %-EPERM: Device PM QoS resume latency 0. + * * %-EINPROGRESS: Suspend already in progress. + * * %-ENOSYS: %CONFIG_PM not enabled. + * * Other values and conditions for the above values are possible as returned + * by Runtime PM suspend callbacks. */ static inline int pm_runtime_put_sync_autosuspend(struct device *dev) { @@ -741,13 +784,22 @@ static inline int pm_runtime_put_sync_autosuspend(struct device *dev) } /** - * pm_runtime_set_active - Set runtime PM status to "active". + * pm_runtime_set_active - Set runtime PM status to "active" and clear errors. * @dev: Target device. * - * Set the runtime PM status of @dev to %RPM_ACTIVE and ensure that dependencies - * of it will be taken into account. + * Set the runtime PM status of @dev to %RPM_ACTIVE and ensure that its + * dependencies will be taken into account. Also clear the device's error + * status (@dev->power.runtime_error). + * + * It is only valid to call this function if runtime PM is disabled or if + * @dev->power.runtime_error is set. * - * It is not valid to call this function for devices with runtime PM enabled. + * This will fail if suppliers cannot be resumed, or if the parent is not in + * the correct state. + * + * Return: + * * %0: Success. + * * Error code on failure. */ static inline int pm_runtime_set_active(struct device *dev) { @@ -755,13 +807,19 @@ static inline int pm_runtime_set_active(struct device *dev) } /** - * pm_runtime_set_suspended - Set runtime PM status to "suspended". + * pm_runtime_set_suspended - Set runtime PM status to "suspended" and clear errors. * @dev: Target device. * - * Set the runtime PM status of @dev to %RPM_SUSPENDED and ensure that - * dependencies of it will be taken into account. + * Set the runtime PM status of @dev to %RPM_SUSPENDED and ensure that its + * dependencies will be taken into account. Also clear the device's error + * status (@dev->power.runtime_error). * - * It is not valid to call this function for devices with runtime PM enabled. + * It is only valid to call this function if runtime PM is disabled or if + * @dev->power.runtime_error is set. + * + * Return: + * * %0: Success. + * * Error code on failure. */ static inline int pm_runtime_set_suspended(struct device *dev) { @@ -777,9 +835,8 @@ static inline int pm_runtime_set_suspended(struct device *dev) * * If the counter is zero when this function runs and there is a pending runtime * resume request for @dev, it will be resumed. If the counter is still zero at - * that point, all of the pending runtime PM requests for @dev will be canceled - * and all runtime PM operations in progress involving it will be waited for to - * complete. + * that point, this function cancels all pending runtime PM requests for @dev + * and waits for its runtime PM operations to complete (if any). * * For each invocation of this function for @dev, there must be a matching * pm_runtime_enable() call, so that runtime PM is eventually enabled for it diff --git a/include/linux/thermal.h b/include/linux/thermal.h index 083b4f533933..306ad17aed89 100644 --- a/include/linux/thermal.h +++ b/include/linux/thermal.h @@ -293,8 +293,9 @@ struct device *thermal_zone_device(struct thermal_zone_device *tzd); void thermal_zone_device_update(struct thermal_zone_device *, enum thermal_notify_event); -struct thermal_cooling_device *thermal_cooling_device_register(const char *, - void *, const struct thermal_cooling_device_ops *); +struct thermal_cooling_device *thermal_cooling_device_create( + struct device *parent, const char *type, void *devdata, + const struct thermal_cooling_device_ops *ops); struct thermal_cooling_device * devm_thermal_cooling_device_register(struct device *dev, const char *type, void *devdata, @@ -340,9 +341,9 @@ static inline void thermal_zone_device_update(struct thermal_zone_device *tz, enum thermal_notify_event event) { } -static inline struct thermal_cooling_device * -thermal_cooling_device_register(const char *type, void *devdata, - const struct thermal_cooling_device_ops *ops) +static inline struct thermal_cooling_device *thermal_cooling_device_create( + struct device *parent, const char *type, void *devdata, + const struct thermal_cooling_device_ops *ops) { return ERR_PTR(-ENODEV); } static inline struct thermal_cooling_device * @@ -391,4 +392,12 @@ static inline void thermal_pm_prepare(void) {} static inline void thermal_pm_complete(void) {} #endif /* CONFIG_THERMAL */ +static inline struct thermal_cooling_device *thermal_cooling_device_register( + const char *type, void *devdata, + const struct thermal_cooling_device_ops *ops) +{ + return thermal_cooling_device_create(NULL, type, devdata, ops); +} + + #endif /* __THERMAL_H__ */ diff --git a/include/uapi/linux/thermal.h b/include/uapi/linux/thermal.h index 46a2633d33aa..36f3da2d6a7c 100644 --- a/include/uapi/linux/thermal.h +++ b/include/uapi/linux/thermal.h @@ -81,11 +81,11 @@ enum thermal_genl_event { THERMAL_GENL_EVENT_CDEV_STATE_UPDATE, /* Cdev state updated */ THERMAL_GENL_EVENT_TZ_GOV_CHANGE, /* Governor policy changed */ THERMAL_GENL_EVENT_CPU_CAPABILITY_CHANGE, /* CPU capability changed */ - THERMAL_GENL_EVENT_THRESHOLD_ADD, /* A thresold has been added */ - THERMAL_GENL_EVENT_THRESHOLD_DELETE, /* A thresold has been deleted */ + THERMAL_GENL_EVENT_THRESHOLD_ADD, /* A threshold has been added */ + THERMAL_GENL_EVENT_THRESHOLD_DELETE, /* A threshold has been deleted */ THERMAL_GENL_EVENT_THRESHOLD_FLUSH, /* All thresolds have been deleted */ - THERMAL_GENL_EVENT_THRESHOLD_UP, /* A thresold has been crossed the way up */ - THERMAL_GENL_EVENT_THRESHOLD_DOWN, /* A thresold has been crossed the way down */ + THERMAL_GENL_EVENT_THRESHOLD_UP, /* A threshold has been crossed the way up */ + THERMAL_GENL_EVENT_THRESHOLD_DOWN, /* A threshold has been crossed the way down */ __THERMAL_GENL_EVENT_MAX, }; #define THERMAL_GENL_EVENT_MAX (__THERMAL_GENL_EVENT_MAX - 1) diff --git a/kernel/cpu_pm.c b/kernel/cpu_pm.c index 7481fbb947d3..a2a598ad6e6d 100644 --- a/kernel/cpu_pm.c +++ b/kernel/cpu_pm.c @@ -10,6 +10,7 @@ #include <linux/cpu_pm.h> #include <linux/module.h> #include <linux/notifier.h> +#include <linux/rcupdate.h> #include <linux/spinlock.h> #include <linux/syscore_ops.h> @@ -76,7 +77,8 @@ EXPORT_SYMBOL_GPL(cpu_pm_register_notifier); * * Remove a driver from the CPU PM notifier list. * - * This function has the same return conditions as raw_notifier_chain_unregister. + * This function may sleep, and has the same return conditions as + * raw_notifier_chain_unregister. */ int cpu_pm_unregister_notifier(struct notifier_block *nb) { @@ -86,6 +88,11 @@ int cpu_pm_unregister_notifier(struct notifier_block *nb) raw_spin_lock_irqsave(&cpu_pm_notifier.lock, flags); ret = raw_notifier_chain_unregister(&cpu_pm_notifier.chain, nb); raw_spin_unlock_irqrestore(&cpu_pm_notifier.lock, flags); + + /* Wait for the rcu_read_lock() walkers in cpu_pm_notify(). */ + if (!ret) + synchronize_rcu(); + return ret; } EXPORT_SYMBOL_GPL(cpu_pm_unregister_notifier); diff --git a/kernel/power/em_netlink.c b/kernel/power/em_netlink.c index 4d4fd29bd2be..1f68d456a713 100644 --- a/kernel/power/em_netlink.c +++ b/kernel/power/em_netlink.c @@ -95,41 +95,62 @@ static int __em_nl_get_pd_for_dump(struct em_perf_domain *pd, void *data) return ret; } +struct em_nl_doit_ctx { + struct genl_info *info; + int cmd; + struct sk_buff *msg; +}; + +static int __em_nl_get_pd_doit_fill(struct em_perf_domain *pd, void *data) +{ + struct em_nl_doit_ctx *ctx = data; + void *hdr; + + hdr = genlmsg_put_reply(ctx->msg, ctx->info, &dev_energymodel_nl_family, 0, + ctx->cmd); + if (!hdr) + return -EMSGSIZE; + + if (__em_nl_get_pd(pd, ctx->msg)) { + genlmsg_cancel(ctx->msg, hdr); + return -EMSGSIZE; + } + + genlmsg_end(ctx->msg, hdr); + return 0; +} + int dev_energymodel_nl_get_perf_domains_doit(struct sk_buff *skb, - struct genl_info *info) + struct genl_info *info) { - int id, ret = -EMSGSIZE, msg_sz = 0; - int cmd = info->genlhdr->cmd; - struct em_perf_domain *pd; + struct em_nl_doit_ctx ctx = { + .info = info, + .cmd = info->genlhdr->cmd, + }; struct sk_buff *msg; - void *hdr; + int id, ret, msg_sz = 0; if (!info->attrs[DEV_ENERGYMODEL_A_PERF_DOMAIN_PERF_DOMAIN_ID]) return -EINVAL; id = nla_get_u32(info->attrs[DEV_ENERGYMODEL_A_PERF_DOMAIN_PERF_DOMAIN_ID]); - pd = em_perf_domain_get_by_id(id); - if (!pd) - return -EINVAL; - __em_nl_get_pd_size(pd, &msg_sz); + /* Encode under em_pd_list_mutex, like the dumpit path. */ + ret = em_perf_domain_for_id(id, __em_nl_get_pd_size, &msg_sz); + if (ret) + return ret; + msg = genlmsg_new(msg_sz, GFP_KERNEL); if (!msg) return -ENOMEM; - hdr = genlmsg_put_reply(msg, info, &dev_energymodel_nl_family, 0, cmd); - if (!hdr) - goto out_free_msg; - - ret = __em_nl_get_pd(pd, msg); + ctx.msg = msg; + ret = em_perf_domain_for_id(id, __em_nl_get_pd_doit_fill, &ctx); if (ret) - goto out_cancel_msg; - genlmsg_end(msg, hdr); + goto out_free_msg; return genlmsg_reply(msg, info); -out_cancel_msg: - genlmsg_cancel(msg, hdr); out_free_msg: nlmsg_free(msg); return ret; @@ -148,19 +169,6 @@ int dev_energymodel_nl_get_perf_domains_dumpit(struct sk_buff *skb, return for_each_em_perf_domain(__em_nl_get_pd_for_dump, &ctx); } -static struct em_perf_domain *__em_nl_get_pd_table_id(struct nlattr **attrs) -{ - struct em_perf_domain *pd; - int id; - - if (!attrs[DEV_ENERGYMODEL_A_PERF_TABLE_PERF_DOMAIN_ID]) - return NULL; - - id = nla_get_u32(attrs[DEV_ENERGYMODEL_A_PERF_TABLE_PERF_DOMAIN_ID]); - pd = em_perf_domain_get_by_id(id); - return pd; -} - static int __em_nl_get_pd_table_size(const struct em_perf_domain *pd) { int id_sz, ps_sz; @@ -245,34 +253,58 @@ out_err: return -EMSGSIZE; } +static int __em_nl_get_pd_table_size_cb(struct em_perf_domain *pd, void *data) +{ + *(int *)data = __em_nl_get_pd_table_size(pd); + return 0; +} + +static int __em_nl_get_pd_table_doit_fill(struct em_perf_domain *pd, + void *data) +{ + struct em_nl_doit_ctx *ctx = data; + void *hdr; + + hdr = genlmsg_put_reply(ctx->msg, ctx->info, &dev_energymodel_nl_family, 0, + ctx->cmd); + if (!hdr) + return -EMSGSIZE; + + if (__em_nl_get_pd_table(ctx->msg, pd)) + return -EMSGSIZE; + + genlmsg_end(ctx->msg, hdr); + return 0; +} + int dev_energymodel_nl_get_perf_table_doit(struct sk_buff *skb, - struct genl_info *info) + struct genl_info *info) { - int cmd = info->genlhdr->cmd; - int msg_sz, ret = -EMSGSIZE; - struct em_perf_domain *pd; + struct em_nl_doit_ctx ctx = { + .info = info, + .cmd = info->genlhdr->cmd, + }; struct sk_buff *msg; - void *hdr; + int id, ret, msg_sz; - pd = __em_nl_get_pd_table_id(info->attrs); - if (!pd) + if (!info->attrs[DEV_ENERGYMODEL_A_PERF_TABLE_PERF_DOMAIN_ID]) return -EINVAL; - msg_sz = __em_nl_get_pd_table_size(pd); + id = nla_get_u32(info->attrs[DEV_ENERGYMODEL_A_PERF_TABLE_PERF_DOMAIN_ID]); + + ret = em_perf_domain_for_id(id, __em_nl_get_pd_table_size_cb, &msg_sz); + if (ret) + return ret; msg = genlmsg_new(msg_sz, GFP_KERNEL); if (!msg) return -ENOMEM; - hdr = genlmsg_put_reply(msg, info, &dev_energymodel_nl_family, 0, cmd); - if (!hdr) - goto out_free_msg; - - ret = __em_nl_get_pd_table(msg, pd); + ctx.msg = msg; + ret = em_perf_domain_for_id(id, __em_nl_get_pd_table_doit_fill, &ctx); if (ret) goto out_free_msg; - genlmsg_end(msg, hdr); return genlmsg_reply(msg, info); out_free_msg: diff --git a/kernel/power/em_netlink.h b/kernel/power/em_netlink.h index 583d7f1c3939..bc98a3c278b9 100644 --- a/kernel/power/em_netlink.h +++ b/kernel/power/em_netlink.h @@ -12,7 +12,8 @@ #if defined(CONFIG_ENERGY_MODEL) && defined(CONFIG_NET) int for_each_em_perf_domain(int (*cb)(struct em_perf_domain*, void *), void *data); -struct em_perf_domain *em_perf_domain_get_by_id(int id); +int em_perf_domain_for_id(int id, int (*cb)(struct em_perf_domain *, void *), + void *data); void em_notify_pd_created(const struct em_perf_domain *pd); void em_notify_pd_deleted(const struct em_perf_domain *pd); void em_notify_pd_updated(const struct em_perf_domain *pd); @@ -24,9 +25,10 @@ int for_each_em_perf_domain(int (*cb)(struct em_perf_domain*, void *), return -EINVAL; } static inline -struct em_perf_domain *em_perf_domain_get_by_id(int id) +int em_perf_domain_for_id(int id, int (*cb)(struct em_perf_domain *, void *), + void *data) { - return NULL; + return -EINVAL; } static inline void em_notify_pd_created(const struct em_perf_domain *pd) {} diff --git a/kernel/power/energy_model.c b/kernel/power/energy_model.c index e610cf8e9a06..a76089e2ce8c 100644 --- a/kernel/power/energy_model.c +++ b/kernel/power/energy_model.c @@ -1031,7 +1031,9 @@ int for_each_em_perf_domain(int (*cb)(struct em_perf_domain*, void *), return 0; } -struct em_perf_domain *em_perf_domain_get_by_id(int id) +/* Run @cb on the matching domain with em_pd_list_mutex held. */ +int em_perf_domain_for_id(int id, int (*cb)(struct em_perf_domain *, void *), + void *data) { struct em_perf_domain *pd; @@ -1040,9 +1042,9 @@ struct em_perf_domain *em_perf_domain_get_by_id(int id) list_for_each_entry(pd, &em_pd_list, node) { if (pd->id == id) - return pd; + return cb(pd, data); } - return NULL; + return -EINVAL; } #endif diff --git a/sound/hda/codecs/side-codecs/aw88399_hda.c b/sound/hda/codecs/side-codecs/aw88399_hda.c index 11cef4923024..439ad45557a4 100644 --- a/sound/hda/codecs/side-codecs/aw88399_hda.c +++ b/sound/hda/codecs/side-codecs/aw88399_hda.c @@ -209,8 +209,7 @@ static int aw88399_hda_acpi_probe(struct aw88399_hda *aw88399) return -ENODEV; } - struct device *physdev __free(put_device) = - get_device(acpi_get_first_physical_node(adev)); + struct device *physdev __free(put_device) = acpi_bus_get_primary_device(adev); acpi_dev_put(adev); if (!physdev) return -ENODEV; diff --git a/sound/soc/amd/acp-es8336.c b/sound/soc/amd/acp-es8336.c index 9f3f11256788..34e953d1a3a7 100644 --- a/sound/soc/amd/acp-es8336.c +++ b/sound/soc/amd/acp-es8336.c @@ -198,7 +198,7 @@ static int st_es8336_late_probe(struct snd_soc_card *card) if (!adev) return -ENODEV; - codec_dev = acpi_get_first_physical_node(adev); + codec_dev = acpi_bus_get_primary_device(adev); acpi_dev_put(adev); if (!codec_dev) { dev_err(card->dev, "can not find codec dev\n"); diff --git a/sound/soc/amd/acp/acp3x-es83xx/acp3x-es83xx.c b/sound/soc/amd/acp/acp3x-es83xx/acp3x-es83xx.c index 3a640e652314..c522b421e1ec 100644 --- a/sound/soc/amd/acp/acp3x-es83xx/acp3x-es83xx.c +++ b/sound/soc/amd/acp/acp3x-es83xx/acp3x-es83xx.c @@ -428,7 +428,7 @@ static int acp3x_es83xx_probe(struct snd_soc_card *card) return -ENXIO; } - codec_dev = acpi_get_first_physical_node(adev); + codec_dev = acpi_bus_get_primary_device(adev); acpi_dev_put(adev); if (!codec_dev) { dev_warn(dev, "Error cannot find codec device, will defer probe\n"); diff --git a/sound/soc/intel/boards/bytcht_es8316.c b/sound/soc/intel/boards/bytcht_es8316.c index ea387dc74273..4d8d1d176a49 100644 --- a/sound/soc/intel/boards/bytcht_es8316.c +++ b/sound/soc/intel/boards/bytcht_es8316.c @@ -601,11 +601,11 @@ static int snd_byt_cht_es8316_mc_probe(struct platform_device *pdev) return -ENOENT; } - codec_dev = acpi_get_first_physical_node(adev); + codec_dev = acpi_bus_get_primary_device(adev); acpi_dev_put(adev); if (!codec_dev) return -EPROBE_DEFER; - priv->codec_dev = get_device(codec_dev); + priv->codec_dev = codec_dev; /* override platform name, if required */ byt_cht_es8316_card.dev = dev; diff --git a/sound/soc/intel/boards/bytcr_rt5640.c b/sound/soc/intel/boards/bytcr_rt5640.c index 40da3eea5fa7..c84b9a0fe65a 100644 --- a/sound/soc/intel/boards/bytcr_rt5640.c +++ b/sound/soc/intel/boards/bytcr_rt5640.c @@ -1744,11 +1744,11 @@ static int snd_byt_rt5640_mc_probe(struct platform_device *pdev) return -ENOENT; } - codec_dev = acpi_get_first_physical_node(adev); + codec_dev = acpi_bus_get_primary_device(adev); acpi_dev_put(adev); if (codec_dev) { - priv->codec_dev = get_device(codec_dev); + priv->codec_dev = codec_dev; } else { /* * Special case for Android tablets where the codec i2c_client diff --git a/sound/soc/intel/boards/bytcr_rt5651.c b/sound/soc/intel/boards/bytcr_rt5651.c index c9b1053205de..fbfaf0b351c9 100644 --- a/sound/soc/intel/boards/bytcr_rt5651.c +++ b/sound/soc/intel/boards/bytcr_rt5651.c @@ -960,11 +960,11 @@ static int snd_byt_rt5651_mc_probe(struct platform_device *pdev) return -ENOENT; } - codec_dev = acpi_get_first_physical_node(adev); + codec_dev = acpi_bus_get_primary_device(adev); acpi_dev_put(adev); if (!codec_dev) return -EPROBE_DEFER; - priv->codec_dev = get_device(codec_dev); + priv->codec_dev = codec_dev; /* * swap SSP0 if bytcr is detected diff --git a/sound/soc/intel/boards/cht_bsw_rt5645.c b/sound/soc/intel/boards/cht_bsw_rt5645.c index 249be121be15..4f20928a3604 100644 --- a/sound/soc/intel/boards/cht_bsw_rt5645.c +++ b/sound/soc/intel/boards/cht_bsw_rt5645.c @@ -529,7 +529,6 @@ static int snd_cht_mc_probe(struct platform_device *pdev) const char *platform_name; struct cht_mc_private *drv; struct acpi_device *adev; - struct device *codec_dev; bool sof_parent; bool found = false; bool is_bytcr = false; @@ -583,8 +582,8 @@ static int snd_cht_mc_probe(struct platform_device *pdev) return -ENOENT; } - /* acpi_get_first_physical_node() returns a borrowed ref, no need to deref */ - codec_dev = acpi_get_first_physical_node(adev); + struct device *codec_dev __free(put_device) = acpi_bus_get_primary_device(adev); + acpi_dev_put(adev); if (!codec_dev) return -EPROBE_DEFER; diff --git a/sound/soc/intel/boards/sof_cirrus_common.c b/sound/soc/intel/boards/sof_cirrus_common.c index 88fc6cb2bfd4..a4ecb1f9a615 100644 --- a/sound/soc/intel/boards/sof_cirrus_common.c +++ b/sound/soc/intel/boards/sof_cirrus_common.c @@ -168,7 +168,7 @@ static int cs35l41_compute_codec_conf(void) cs35l41_name_prefixes[uid]); continue; } - physdev = get_device(acpi_get_first_physical_node(adev)); + physdev = acpi_bus_get_primary_device(adev); acpi_dev_put(adev); if (!physdev) { pr_devel("Cannot find physical node for HID %s UID %u (%s)\n", CS35L41_HID, diff --git a/sound/soc/intel/boards/sof_es8336.c b/sound/soc/intel/boards/sof_es8336.c index f1e62c2e79fe..d40f10c6ae54 100644 --- a/sound/soc/intel/boards/sof_es8336.c +++ b/sound/soc/intel/boards/sof_es8336.c @@ -702,11 +702,11 @@ static int sof_es8336_probe(struct platform_device *pdev) return -ENOENT; } - codec_dev = acpi_get_first_physical_node(adev); + codec_dev = acpi_bus_get_primary_device(adev); acpi_dev_put(adev); if (!codec_dev) return -EPROBE_DEFER; - priv->codec_dev = get_device(codec_dev); + priv->codec_dev = codec_dev; ret = snd_soc_fixup_dai_links_platform_name(&sof_es8336_card, mach->mach_params.platform); diff --git a/sound/soc/loongson/loongson_card.c b/sound/soc/loongson/loongson_card.c index 6422cc1703b6..9f3907278462 100644 --- a/sound/soc/loongson/loongson_card.c +++ b/sound/soc/loongson/loongson_card.c @@ -196,7 +196,6 @@ static int loongson_card_parse_acpi(struct loongson_card_data *data) struct snd_soc_card *card = &data->snd_card; const char *codec_dai_name; struct acpi_device *adev; - struct device *phy_dev; int i, ret; /* fixup platform name based on reference node */ @@ -204,7 +203,7 @@ static int loongson_card_parse_acpi(struct loongson_card_data *data) if (!adev) return -ENOENT; - phy_dev = acpi_get_first_physical_node(adev); + struct device *phy_dev __free(put_device) = acpi_bus_get_primary_device(adev); if (!phy_dev) return -EPROBE_DEFER; diff --git a/tools/power/pm-graph/sleepgraph.py b/tools/power/pm-graph/sleepgraph.py index f6d172254829..c8bc3bed3f5a 100755 --- a/tools/power/pm-graph/sleepgraph.py +++ b/tools/power/pm-graph/sleepgraph.py @@ -1448,7 +1448,7 @@ class DevProps: # Class: DeviceNode # Description: -# A container used to create a device hierachy, with a single root node +# A container used to create a device hierarchy, with a single root node # and a tree of child nodes. Used by Data.deviceTopology() class DeviceNode: def __init__(self, nodename, nodedepth): |
