| Age | Commit message (Collapse) | Author |
|
https://gitlab.freedesktop.org/drm/amdgpu/kernel into drm-fixes
amd-drm-fixes-7.3-2026-10-01:
amdgpu:
- dc_state_create_copy() fix
- HDMI RGB limited range fix
- eDP ASSR fix
- DCE 6 fixes
- DCE 8.1 fix
- SI DPM fixes
- Reset fixes
- Workaround for multiple SDMA entities with DCC
- PWM backlight fix
- GPUVM fixes
- Switcheroo fix
- Error handling leak fixes
- GC 6 unload FW leak fix
- SDMA 7.1 fix
- DML frame size limit fix
- RGB vs YCbCr 4:4:4 fix
- DP MST fix
- DC Power module fixes
- MacBookPro14,3 fix
amdkfd:
- SVM fix
radeon:
- Sparc64 fix
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Alex Deucher <alexander.deucher@amd.com>
Link: https://patch.msgid.link/20261001230115.1319089-1-alexander.deucher@amd.com
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/chunkuang.hu/linux into drm-fixes
Mediatek DRM Fixes - 20261002
1. Add missing IS_ERR check for ovl_adaptor platform device
2. Fix VID_DOWNSAMPLE_CONFIG register offset
3. Fix pdev reference leak in mtk_drm_bind()
4. Fix runtime PM leak in mtk_hdmi_ddc_v2_probe()
5. Fix ovl adaptor platform device leak
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Chun-Kuang Hu <chunkuang.hu@kernel.org>
Link: https://patch.msgid.link/20261001234906.14940-1-chunkuang.hu@kernel.org
|
|
https://gitlab.freedesktop.org/drm/i915/kernel into drm-fixes
drm/i915 fixes for v7.3-rc6:
- Disable VRR DC balance by default to fix timing issues
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Jani Nikula <jani.nikula@intel.com>
Link: https://patch.msgid.link/5a9a11777f605b59550783923add4568e974a6ea@intel.com
|
|
mtk_drm_probe() creates an OVL adaptor platform device with
platform_device_register_data() when the display pipeline requires the
OVL adaptor.
If a later initialization step fails, the probe error path releases
the DRM resources without unregistering the already registered OVL
adaptor device. The normal remove path likewise leaves the device
registered after the DRM driver is unbound.
Keep track of whether the OVL adaptor was successfully registered and
unregister it on probe failure. Also recover the platform device from
the stored DDP component device and unregister it during normal
removal.
The issue was identified by a static analysis tool I developed and
confirmed by manual review.
Fixes: 0d9eee9118b7 ("drm/mediatek: Add drm ovl_adaptor sub driver for MT8195")
Cc: stable@vger.kernel.org
Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com>
Link: https://patchwork.kernel.org/project/linux-mediatek/patch/20260921131156.403652-1-lgs201920130244@gmail.com/
Signed-off-by: Chun-Kuang Hu <chunkuang.hu@kernel.org>
|
|
pm_runtime_get_sync() unconditionally bumps the usage counter, but
nothing puts it back when devm_i2c_add_adapter() fails, leaving the DDC
runtime-resumed after a failed probe. Drop the reference on that error
path.
Fixes: 8d0f79886273 ("drm/mediatek: Introduce HDMI/DDC v2 for MT8195/MT8188")
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Link: https://patchwork.kernel.org/project/linux-mediatek/patch/20260916174229.2088824-1-vulab@iscas.ac.cn/
Signed-off-by: Chun-Kuang Hu <chunkuang.hu@kernel.org>
|
|
When the device is not the mmsys master, mtk_drm_bind() returns early
after taking a reference on the disp-mutex device via
of_find_device_by_node(), without ever dropping it: mtk_drm_unbind()
only puts mutex_dev for the master. Drop the reference before returning
from the non-master path.
Fixes: 1ef7ed48356c ("drm/mediatek: Modify mediatek-drm for mt8195 multi mmsys support")
Cc: stable@vger.kernel.org
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Link: https://patchwork.kernel.org/project/linux-mediatek/patch/20260916174058.2088709-1-vulab@iscas.ac.cn/
Signed-off-by: Chun-Kuang Hu <chunkuang.hu@kernel.org>
|
|
On a MacBookPro14,3 with a Radeon Pro 555 (Polaris11), the framebuffer
is at MC address 0 when amdgpu loads after a cold boot, as the firmware
leaves it (MC_VM_FB_LOCATION = 0x007f0000), while the VBIOS ASIC_Init
table places it at 0xF4_0000_0000 (0xf47ff400). amdgpu reads the
location once, at init, so after a re-POST (S3 resume or GPU reset)
the framebuffer has moved and the driver keeps programming the old
one: the SMU is handed a table that was never written and the GPU
does not come back, which leaves the internal panel black.
Resetting the ASIC on load makes ASIC_Init run before the driver reads
the location, so the driver uses the VBIOS placement from the start and
every later re-POST puts the framebuffer back where it already is.
Add the Radeon Pro 555 used in this machine to the existing VI reset
quirk table.
Tested on a MacBookPro14,3 on 6.18.49 with the quirk table backported
(the kernel also carries unrelated local PCI and ACPI patches for this
machine). The framebuffer is at 0x000000F400000000 after both cold and
warm boot, and the GPU survived 9 S3 cycles (lid close and rtcwake, one
of them with the lid closed for about 7.5 minutes and a USB-C disk
attached), each followed by a few minutes of 3D load; no ring timeouts
or VM faults were reported. The reset adds about 0.23 s to amdgpu init.
Suggested-by: Christian König <christian.koenig@amd.com>
Suggested-by: Alex Deucher <alexander.deucher@amd.com>
Link: https://lore.kernel.org/all/20260924132952.25054-1-fbeltranmillalen@gmail.com/
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Francisco Beltrán Millalén <fbeltranmillalen@gmail.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit b6b1d97218518003107d4446fac278fa3e0a19a0)
Cc: stable@vger.kernel.org
|
|
Assume that the AC limits are the maximum of all power states,
and the DC limits are the maximum of battery power states.
This shouldn't make any difference in practice, but is
cleaner and more robust against bogus information in the VBIOS.
Fixes: e6c5d36756e7 ("drm/amd/pm/si: Fix updating clock limits from power states")
Signed-off-by: Timur Kristóf <timur.kristof@gmail.com>
Reviewed-by: Mario Limonciello <mario.limonciello@amd.com>
Link: https://patch.msgid.link/20260923120354.1027996-2-timur.kristof@gmail.com
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit ef8cb9dcc7db7b7abcbdbe870a080cc81e4f0bf3)
Cc: stable@vger.kernel.org
|
|
[Why]
mod_power_remove_stream() shifts the remaining power_entity slots down
but does not move replay_events, and mod_power_add_stream() does not
initialize it. replay_events therefore stay bound to the map slot
instead of the stream.
When several streams are disabled in one atomic commit,
amdgpu_dm_mod_power_update_streams() removes them one after another.
The eDP stream can then be looked up in a slot whose stale
replay_events already have replay_event_hw_programming set, so
amdgpu_dm_replay_set_event() returns early ("already in desired state")
without calling mod_power_set_replay_event(). Replay is not disabled
before the eDP panel is powered off. After DPMS on, the sink reports
neither replay state nor frame lock (DPCD 0x378 = 0x00, no error bits),
so the HPD IRQ recovery does not trigger and the panel stays black
until a full modeset.
Seen with an eDP panel using FreeSync Replay plus two DP-MST displays:
DPMS off/on of all outputs leaves eDP black, while DPMS of eDP alone
works. Doing an eDP-only DPMS first makes the next all-output DPMS
fail reliably.
[How]
Shift replay_events together with the PSR cached fields in
mod_power_remove_stream() and initialize it to replay_event_vsync in
mod_power_add_stream(), matching the psr_event_vsync initial value used
for PSR (both vsync events are driven together by
amdgpu_dm_crtc_set_static_screen_optimze()).
Tested on 7.3.0-rc3 (238650ef6c7c): the reproducer above now recovers
reliably, and Replay still engages when the screen is idle.
The issue was debugged with help from an AI assistant (Claude), which
analysed ftrace/kprobe traces and the driver source, pointed to the
missing replay_events handling and suggested this change. I collected
the traces and built and tested the fix on the affected hardware.
Fixes: 4cef2ac4c795 ("drm/amd/display: Introduce power module on Linux")
Assisted-by: Claude:claude-opus-5
Signed-off-by: Simon Polack <spolack+git@mailbox.org>
Reviewed-by: Ray Wu <ray.wu@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 3d47e38e271195435e125527b224d11cfdb777b1)
Cc: stable@vger.kernel.org
|
|
igp_read_bios_from_vram() checked bios[0]/bios[1] with a plain
__iomem load, which faults on sparc64 before the copy runs at all.
radeon_read_bios() already reads its two signature bytes with
readb() ahead of its own copy; use the same accessor here, keeping
the check before the allocation.
Fixes: b442962a9e82 ("drm/radeon/kms: add support for "Surround View"")
Signed-off-by: Imre Kaloz <kaloz@kernel.org>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 09155b8932e5013dbce0f5cdff7d75264a279ee5)
Cc: stable@vger.kernel.org
|
|
dm_dp_mst_is_port_support_mode() reads
aconnector->dc_sink->dsc_caps... for the DSC branch-throughput check,
and get_conv_frl_bw()'s HDMI-PCON FRL-bandwidth path reads
aconnector->dc_sink->edid_caps.max_frl_rate, both without a NULL
check. dc_sink is cleared asynchronously on MST unplug, and both
functions run from paths that the driver's own comments document as
racing that teardown: the connector probe worker's ->mode_valid
callback and a compositor's atomic check, neither of which holds the
MST manager lock that the teardown path uses. The former does have an
existing dsc_aux NULL check, but dsc_aux isn't reliably cleared in
every path that clears dc_sink, so it doesn't cover this.
Fail the port-support check and skip the FRL conversion path when the
sink is already gone.
Fixes: f04d275d94e1 ("drm/amd/display: add mst port output bw check")
Fixes: 5c9b8b27a883 ("drm/amd/display: Tie FRL support into amdgpu_dm")
Assisted-by: gkh_clanker_t1000
Signed-off-by: Hari Mishal <harimishal1@gmail.com>
Reviewed-by: Fangzhi Zuo <jerry.zuo@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 7c3db8da4e039698ae198c870712428644e965c7)
Cc: stable@vger.kernel.org
|
|
amdgpu_dm_create_validate_stream_for_sink() walks encoding_order[] and
uses the first encoding that validates. YCbCr 4:4:4 is listed before
RGB, so an HDMI sink that advertises 4:4:4 gets YCbCr 4:4:4 whenever the
"color format" property is left at AUTO, even though RGB fits the same
link.
That contradicts the documented AUTO behaviour for HDMI in enum
drm_connector_color_format (RGB, falling back to YCbCr 4:2:0 only when
the bandwidth is not available or the mode is 4:2:0-only), which the
amdgpu implementation of the property also describes. It also leaves
the "Broadcast RGB" property without effect on such sinks, since the
quantization range it selects only applies to RGB output.
Try RGB first. The mask still holds every encoding the sink supports,
so a mode that cannot carry RGB falls back exactly as before.
For reference, v7.2 picked RGB here unless YCbCr 4:4:4 was forced
through debugfs, while earlier kernels picked YCbCr 4:4:4 for any HDMI
sink that advertised it.
Fixes: 0b0ff65d3ca1 ("drm/amd/display: Refactor stream validation")
Suggested-by: Adolfo Rodrigues <adolfotregosa@gmail.com>
Assisted-by: Claude Code:claude-fable-5-1
Signed-off-by: Adrian Betschart <adrian.betschart@cinemaone.ch>
Reviewed-by: Fangzhi Zuo <jerry.zuo@amd.com>
Tested-by: Adolfo Rodrigues <adolfotregosa@gmail.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 7de432bc753133764dc40d37f9388da56de251e7)
|
|
Building x86_64 allmodconfig with clang fails:
.../dml2_0/dml21/src/dml2_core/dml2_core_dcn4_calcs.c:10491:13:
error: stack frame size (3128) exceeds limit (3072) in
'dml_core_mode_programming' [-Werror,-Wframe-larger-than]
That config enables both KASAN and UBSAN, and should get the 4096
byte limit meant for clang COMPILE_TEST sanitizer builds. The check
in the dml and dml2_0 Makefiles concatenates the three symbols with
no separator:
ifeq ($(filter y,$(CONFIG_KASAN)$(CONFIG_KCSAN)$(CONFIG_UBSAN)),y)
KASAN and KCSAN cannot be enabled together, so before UBSAN was
added the string was at most "y". UBSAN can be enabled alongside
either of them, and with both set the string becomes "yy",
$(filter y,yy) is empty, and the build falls through to the
non-sanitizer limit of 3072. Adding UBSAN to the check therefore
lowered the limit for the most heavily instrumented builds instead
of raising it.
Separate the symbols with spaces so that $(filter) sees individual
words, and test for a non-empty result. Only the configs with two
sanitizers enabled change behaviour.
Found by KernelCI builds of the linus-next tree.
Fixes: ebf8b0fd8508 ("drm/amd/display: Relax DML frame limit with UBSAN")
Assisted-by: LLM
Signed-off-by: Sasha Levin <sashal@kernel.org>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 0400d53fd03e7d4beb09fdabda42a7b2c4ca465d)
Cc: stable@vger.kernel.org
|
|
In gmc_v12_1_flush_gpu_tlb(), flush_type was not properly
passed to gmc_v12_1_flush_vm_hub(). This only affected
the MMIO path. In most cases the flush would go through
MES.
Reviewed-by: Mukul Joshi <mukul.joshi@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 1f0f66fab18dfa8f14db82b8a6b82dcc21926248)
Cc: stable@vger.kernel.org
|
|
Not enabled in 7.1 so drop the assignment.
Cc: Horatio Zhang <hongkun.zhang@amd.com>
Cc: Hawking.Zhang@amd.com
Fixes: 4ed5116aacf6 ("drm/amdgpu: Add sdma v7_1_0 support")
Reviewed-by: Hawking Zhang <Hawking.Zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 106e19f69a066fb6354dac04e5f5ed43a06727f7)
Cc: stable@vger.kernel.org
|
|
The firmware requested in gfx_v6_0_init_microcode() is not released
when the GFX block is torn down, leaking the firmware resources.
Add gfx_v6_0_free_microcode() and call it from gfx_v6_0_sw_fini()
to release the PFP, ME, CE and RLC firmware.
Signed-off-by: Willian Oliveira <williandossantosdeoliveira287@gmail.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 386e81346f95ab0be5a82c5535de932e7a456924)
Cc: stable@vger.kernel.org
|
|
acp_hw_init() registers ACP child devices with mfd_add_devices() before
attaching them to the ACP power domain and initializing the hardware.
If attaching a child to the power domain fails, or if the ACP reset or
clock enable operation times out, the failure path frees only the
source cell, resource, and platform-data allocations. The child
platform devices already registered by mfd_add_devices() remain
registered and are never released.
Remove the children from the power domain and unregister the MFD
devices on failures that occur after mfd_add_devices() succeeds. Keep
mfd_add_devices() failures on the existing cleanup path since the MFD
core already rolls back partially registered children itself.
The issue was identified by a static analysis tool I developed and
confirmed by manual review.
Fixes: 25030321ba28 ("drm/amd: add pm domain for ACP IP sub blocks")
Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 63da68d6a2279ec945581c990c5a59a5ae8b5560)
Cc: stable@vger.kernel.org
|
|
amdgpu_discovery_sysfs_ips() allocates ip_hw_instance with
kzalloc_flex() and initializes its embedded kobject before calling
kobject_add().
If kobject_add() fails, the return value is ignored and execution
continues without dropping the initial kobject reference. The failed
kobject is not retained in the kset list, so the normal sysfs teardown
path cannot find it. As a result, ip_hw_instance_release() is never
called and the ip_hw_instance allocation is leaked.
Call kobject_put() when kobject_add() fails so the initial reference is
dropped and ip_hw_instance_release() can free the allocation. Keep the
existing best-effort sysfs behavior by continuing with the remaining IP
entries after the failed registration.
The issue was identified by a static analysis tool I developed and
confirmed by manual review.
Fixes: a6c40b178092 ("drm/amdgpu: Show IP discovery in sysfs")
Reviewed-by: Lijo Lazar <lijo.lazar@amd.com>
Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 5c75fbcca8af1096d8d9415b91d8ea6212c22ee1)
Cc: stable@vger.kernel.org
|
|
amdgpu_pmops_suspend_noirq() resets the ASIC unconditionally. When the GPU
has been parked by vga_switcheroo it has neither power nor a PCIe link, so
the reset cannot reach it: pci_set_power_state() reports the device as
inaccessible and amdgpu_asic_reset() returns -EINVAL. A failure there
aborts the entire noirq suspend phase, and with it the system suspend, so
the machine cannot sleep at all while the GPU is switched off.
amdgpu_device_prepare(), amdgpu_device_suspend() and amdgpu_device_resume()
all bail out early on DRM_SWITCH_POWER_OFF. This callback was added later,
for an unrelated reason, and did not inherit the check. nouveau guards
every one of its PM callbacks the same way.
Bail out the same way here. On a single-GPU system switch_power_state is
never DRM_SWITCH_POWER_OFF, so this is a no-op there.
Found on a MacBookPro11,5, where the Radeon is powered down through
apple-gmux so that the internal panel can be driven by the iGPU instead.
Every suspend failed in amdgpu_pmops_suspend_noirq() while the card was
off; with this check a full S3 cycle completes.
Fixes: 9e051720f9d3 ("drm/amdgpu: Ensure HDA function is suspended before ASIC reset")
Signed-off-by: Theo Andersen Carton <andersen.theo@gmail.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 539d0558638a6dcc066f1cf592b6b9a52adc8ea4)
Cc: stable@vger.kernel.org
|
|
In gmc_v12_0_flush_gpu_tlb(), flush_type was not properly
passed to gmc_v12_0_flush_vm_hub(). This only affected
the MMIO path. In most cases the flush would go through
MES.
Reviewed-by: Mukul Joshi <mukul.joshi@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 6498e33f22a12fdc759a0c78bfe0547819277a36)
Cc: stable@vger.kernel.org
|
|
[Why/How]
On the PWM path, convert the userspace brightness to an input signal and
derive the target luminance from the custom backlight curve, then pass
the resulting millipercent to the power module.
This keeps PWM programming and backlight curve mapping in the power
module, aligning the Linux path with the Windows behavior.
Closes: https://gitlab.freedesktop.org/drm/amd/-/issues/5723
Fixes: 3c108046e1d6 ("drm/amd/display: Add power module on Linux")
Assisted-by: Cursor:Claude-Opus-4.8
Signed-off-by: Ray Wu <ray.wu@amd.com>
Signed-off-by: Fangzhi Zuo <jerry.zuo@amd.com>
Reviewed-by: ChiaHsuan (Tom) Chung <chiahsuan.chung@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit bc4e099caa85e8d4c3555851a0feb07058c4a1fe)
Cc: stable@vger.kernel.org
|
|
For unknown reasons, on gfx12 using multiple entities can causes
random corruption of BOs with DCC.
This workaround seems to prevent the issue until the root cause
is understood and fixed.
Link: https://gitlab.freedesktop.org/drm/amd/-/work_items/5663
Fixes: 3a6f6eeb3db5 ("drm/amdgpu: give ttm entities access to all the sdma scheds")
Signed-off-by: Pierre-Eric Pelloux-Prayer <pierre-eric.pelloux-prayer@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Reviewed-by: Christian König <christian.koenig@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit cd647796ac1334acec0468fae49a0dacbba2a34c)
Cc: stable@vger.kernel.org
|
|
CWSR ranges use the always mapped flag and only need a mapping on the
GPU running the queue. However, svm_range_validate_and_map now updates
the mapping on every GPU with the ACCESS attribute, which is all
supported GPUs when XNACK is on. After the range is prefetched to VRAM,
those GPUs may span different XGMI hives, so the page owner mismatches
and is set to NULL. hmm_range_fault then tries to migrate the VRAM
pages back to system memory, but svm_migrate_to_ram skips it because
migrating pages back to system memory is not allowed while the GPU
mapping is being updated, so the CPU page fault is never resolved.
Only map to the ACCESS GPUs if the range is not mapped to any GPU and
is not in VRAM.
Fixes: b96f26104946 ("drm/amdkfd: Unmap svm range from GPU set to no-access")
Signed-off-by: Andrew Martin <andrew.martin@amd.com>
Assisted-by: Claude:claude-opus-5
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Reviewed-by: Philip Yang <philip.yang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 093cf8842ace35ce212858d5059fee007ec8e778)
Cc: stable@vger.kernel.org
|
|
Set KMD_QUEUE in the GFX MQD and kernel ring control registers so CP
preserves kernel IB VMIDs instead of replacing them with the queue VMID.
Use prop->kernel_queue to distinguish kernel and user MQDs.
Signed-off-by: Jesse Zhang <Jesse.Zhang@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 893642e0e3b2f5aa818e0ee4a49e2d401d541e3a)
|
|
Set KMD_QUEUE in the GFX MQD and kernel ring control registers so CP
preserves kernel IB VMIDs instead of replacing them with the queue VMID.
Use prop->kernel_queue to distinguish kernel and user MQDs.
Signed-off-by: Jesse Zhang <Jesse.Zhang@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 14741c01680a803050e949744089b57fa53e02f7)
|
|
amdgpu_ring_set_fence_errors_and_reemit() and
amdgpu_ring_backup_unprocessed_commands() masked last_seq and sync_seq to
fence-array indices before the loop, then walked a do/while until the two
indices met. When hardware completed every command before the driver
signalled the fences, last_seq == sync_seq, but the masked-index do/while
could not tell that empty interval from a full table and walked the whole
fence array, backing up and replaying already-completed commands.
Read the raw sequence numbers, return early through amdgpu_fence_process()
when they are equal, and mask only at the array index. Guard the
force-completion against a NULL guilty_fence so a non-guilty backup does
not dereference it.
Signed-off-by: Jesse Zhang <Jesse.Zhang@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 78fbe156637a11456fcd29615a56e098ec8a6990)
|
|
amdgpu_ring_backup_unprocessed_commands() is also used to preserve a
ring that carries no guilty command - for example a collateral kernel
gfx ring that shares the faulting pipe during a gfx user-queue pipe
reset. Such a ring is backed up with a NULL guilty_fence and all of its
unprocessed commands must be saved so they can be replayed after the
reset rebuilds the ring's MQD.
The early-return that skips an already-seen guilty fence compared
"ring->guilty_fence == guilty_fence", which is also true when both are
NULL. A non-guilty save therefore hit the early return and left
ring_backup_entries_to_copy at 0, so the pending commands were dropped
when the ring was rebuilt.
Guard the early return with a non-NULL guilty_fence so a NULL fence
never takes it and every unprocessed command is backed up.
Signed-off-by: Jesse Zhang <Jesse.Zhang@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit b3f73daa1c2498273981b68b7534e540614069e3)
|
|
Unfortunately, some desktop boards eg. the FirePro D500
have the HARDWAREDC platform flag set and also have
a battery power state configured in their VBIOS.
This makes no sense for a desktop GPU and results in
incorrect behaviour: the clocks are stuck at lowest.
We observed that the kernel driver can work around
the issue in two possible ways:
1. Set PPSMC_SWSTATE_FLAG_DC on all power states
2. Clear PPSMC_SYSTEMFLAG_GPIO_DC
We think that the PPSMC_SYSTEMFLAG_GPIO_DC flag
makes the SMC assume it's running on battery even
though this is a desktop machine with no battery,
and that's why it doesn't do DPM on power states
without PPSMC_SYSTEMFLAG_GPIO_DC.
Issue was uncovered by "Fix updating clock limits from
power states" because previously the limits for the
battery state were not tracked separately. However,
battery power state has lower frequencies and voltages,
so the kernel doesn't set the DC flag on the current
power state anymore. That causes the SMC to be stuck
on the lowest clocks.
Let's clear PPSMC_SYSTEMFLAG_GPIO_DC on desktop GPUs.
We can use the AMD_IS_MOBILITY flag to determine that.
Fixes: e6c5d36756e7 ("drm/amd/pm/si: Fix updating clock limits from power states")
Closes: https://gitlab.freedesktop.org/mesa/mesa/-/work_items/16352
Signed-off-by: Timur Kristóf <timur.kristof@gmail.com>
Reviewed-by: Mario Limonciello <mario.limonciello@amd.com>
Link: https://patch.msgid.link/20260923120354.1027996-1-timur.kristof@gmail.com
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 58473c7c49ce9f5f5a4a1c24ffb9aa1d3c8771f4)
Cc: stable@vger.kernel.org
|
|
DCE 6.4 is found in Oland chips, which were advertised
as low-end gaming GPUs in 2013~2015 and then have been
sold as low-end workstation GPUs until 2017~2019.
These are not APUs.
Fixes: 7c15fd86aaec ("drm/amd/display: dc/dce: add initial DCE6 support (v10)")
Signed-off-by: Timur Kristóf <timur.kristof@gmail.com>
Reviewed-by: Mario Limonciello <mario.limonciello@amd.com>
Link: https://patch.msgid.link/20260923122033.1046651-3-timur.kristof@gmail.com
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 1072aa811ea1e61d5f763e4ec1bc6d24bb61f14a)
Cc: stable@vger.kernel.org
|
|
DCE8 should use legacy fast updates.
This was already set for DCE 8.0 and 8.3, but it seems
that DCE 8.1 (found in Kaveri chips) was forgotten.
This fix should be backported to all kernel versions that
have "Refactor fast update to use new HWSS build sequence",
but this patch won't apply cleanly to older kernels due to
another refactor in "Remove dc param from check_update".
Fixes: 0baae6246307 ("drm/amd/display: Refactor fast update to use new HWSS build sequence")
Fixes: 9ec11bb842b6 ("drm/amd/display: Remove dc param from check_update")
Signed-off-by: Timur Kristóf <timur.kristof@gmail.com>
Reviewed-by: Mario Limonciello <mario.limonciello@amd.com>
Link: https://patch.msgid.link/20260923122033.1046651-2-timur.kristof@gmail.com
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 57605379559e3b5044c8274a1010ec658e5e73af)
Cc: stable@vger.kernel.org
|
|
DCE6 should use legacy fast updates, just like all DCE8 - DCN2.
Set that flag for DCE 6.0, 6.1 and 6.4 (all DCE6 versions).
Technically DCE6 should have been included when the fast update
code was refactored, but it just happened to work until now.
A recent commit "Attach only plane updates that actually changed"
has exposed this issue because with that commit, DC accidentally
took the new fast update code path for DCE6 as well, which caused
it to boot into a black screen with a flip done timeout.
This fix should be backported to all kernel versions that
have "Refactor fast update to use new HWSS build sequence",
but this patch won't apply cleanly to older kernels due to
another refactor in "Remove dc param from check_update".
Fixes: 0baae6246307 ("drm/amd/display: Refactor fast update to use new HWSS build sequence")
Fixes: 9ec11bb842b6 ("drm/amd/display: Remove dc param from check_update")
Signed-off-by: Timur Kristóf <timur.kristof@gmail.com>
Reviewed-by: Mario Limonciello <mario.limonciello@amd.com>
Link: https://patch.msgid.link/20260923122033.1046651-1-timur.kristof@gmail.com
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 698fda6e360625850413a2d1a90b72ae5f03073d)
Cc: stable@vger.kernel.org
|
|
Commit 56d8ce9d8c17 ("drm/amd/display: Apply correct panel mode when
reinitializing hardware") made ASSR failure fall back to the default panel
mode unless eDP mode had previously been applied.
However, dc_link is zero-initialized and DP_PANEL_MODE_DEFAULT is zero.
Before the first call to dp_set_panel_mode(), panel_mode therefore looks
like a previously applied default mode. If ASSR setup fails during the
first link training attempt, the driver incorrectly trains the eDP link
using the default scrambling mode.
On a Google Vilboz Chromebook running self-built coreboot firmware and
PSP verstage, this leaves the panel mostly black with corrupted output
along the top edge.
Track whether panel_mode has actually been initialized, and only consult
the saved mode after it has been applied. This retains the recovery
behavior while preserving eDP mode during initial link training.
Fixes: 56d8ce9d8c17 ("drm/amd/display: Apply correct panel mode when reinitializing hardware")
Assisted-by: Pi:gpt-5.6-sol
Reviewed-by: George Zhang <george.zhang@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Arthur Heymans <arthur@aheymans.xyz>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit ccb6c2702033f1002b45807fb3cc938760f26f78)
Cc: stable@vger.kernel.org # 6.4.x
|
|
With the Broadcast RGB property at Automatic,
amdgpu_dm_get_output_color_space() selects COLOR_SPACE_SRGB for RGB
output, so DC neither compresses the pixels nor, without a QS-capable
sink, signals full range in the AVI InfoFrame. A sink that follows
CTA-861 treats the default quantization of a CTA video format as limited
and expands 16-235 to 0-255: everything below 16 is crushed to black,
everything above 235 clips.
Follow the CTA-861 default for Automatic instead, as the DRM HDMI state
helper (hdmi_is_limited_range()) and i915 do: limited range on an HDMI
sink for CTA modes other than 640x480, full range elsewhere. Full and
Limited keep their explicit meaning. The same rule applies to the
BT.2020 RGB branch. With commit 892659399f64 ("drm/amd/display:
Propagate HDMI RGB quantization selectability") the AVI InfoFrame then
carries the matching Q value on sinks that support selection.
Measured on a Radeon RX 7600 (DCN 3.2.1) driving a JVC DLA-RS4100 at
1920x1080p24 RGB 12 bpc: the sink reports the signal as limited range
while the picture shows crushed shadows; rendering limited range in the
client makes it match a reference source.
The KUnit fixtures that reach amdgpu_dm_get_output_color_space() now
carry a connector, since Automatic reads its display_info, and five
cases cover the new rule.
Closes: https://gitlab.freedesktop.org/drm/amd/-/issues/5796
Reviewed-by: George Zhang <george.zhang@amd.com>
Signed-off-by: Adrian Betschart <adrian.betschart@cinemaone.ch>
Assisted-by: Claude Code:claude-fable-5-1
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit d23f735d0eac52c673d91e21d146b09588537efc)
|
|
dc_state_create_copy() can return NULL on allocation failure.
dm_suspend() only conditionally skips dm_gpureset_toggle_interrupts()
and continues execution, returning success. dm_resume() then
dereferences the NULL cached_dc_state in link_enc_cfg_copy() and the
following dc_state->stream_count loop, crashing during GPU reset
recovery.
Return -ENOMEM immediately if the copy fails, so the caller aborts
suspend instead of leaving a NULL cached state for resume.
Fixes: 8092aa3ab8f7 ("drm/amd/display: Add null checker before passing variables")
Reviewed-by: George Zhang <george.zhang@amd.com>
Signed-off-by: Jiangshan Yi <yijiangshan@kylinos.cn>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 84b6d7932fdd2857795a257f9f258562a3cddcab)
Cc: stable@vger.kernel.org
|
|
According to the datasheet, VID_DOWNSAMPLE_CONFIG is at offset 0x8f0;
0x8d0 is the VID_CSC_COEFF_0 register.
Fixes: 8d0f79886273 ("drm/mediatek: Introduce HDMI/DDC v2 for MT8195/MT8188")
Signed-off-by: Julien Stephan <jstephan@baylibre.com>
Reviewed-by: Louis-Alexis Eyraud <louisalexis.eyraud@collabora.com>
Reviewed-by: AngeloGioacchino Del Regno <angelogioacchino.delregno@collabora.com>
Link: https://patchwork.kernel.org/project/linux-mediatek/patch/20260828-mtk-hdmi-v2-fix-register-offset-v1-1-118ad5d7ebe3@baylibre.com/
Signed-off-by: Chun-Kuang Hu <chunkuang.hu@kernel.org>
|
|
platform_device_register_data() can fail and return an ERR_PTR, but the
return value is used without checking, leading to an invalid pointer
being stored in ddp_comp[].dev and passed to component_match_add() and
mtk_ddp_comp_init(), which could result in a kernel crash.
Add an IS_ERR() check to jump to the error handling path on failure.
Fixes: 0d9eee9118b7 ("drm/mediatek: Add drm ovl_adaptor sub driver for MT8195")
Cc: stable@vger.kernel.org
Signed-off-by: Haojie Li <lihaojie@kylinos.cn>
Reviewed-by: CK Hu <ck.hu@mediatek.com>
Link: https://patchwork.kernel.org/project/linux-mediatek/patch/20260825100845.438893-1-lihaojie@kylinos.cn/
Signed-off-by: Chun-Kuang Hu <chunkuang.hu@kernel.org>
|
|
When VFs are enabled on dGFX the driver resizes the PF VF_LMEM_BAR to
fit the requested layout. After VFs are disabled the PF VF BAR
size is left as-is. On platforms with tight MMIO apertures a
subsequent unplug/rescan followed by another enable may fail with:
"VF BAR …: can't assign; no space"
because the PCI core reserves address space based on the (now large) VF
template, often multiplied by totalvfs.
Closes: https://gitlab.freedesktop.org/drm/xe/kernel/-/issues/5937
Fixes: 94eae6ee4c2d ("drm/xe/pf: Set VF LMEM BAR size")
Signed-off-by: Marcin Bernatowicz <marcin.bernatowicz@linux.intel.com>
Cc: Michał Wajdeczko <michal.wajdeczko@intel.com>
Cc: Michał Winiarski <michal.winiarski@intel.com>
Reviewed-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
Link: https://patch.msgid.link/20260918110130.700332-1-marcin.bernatowicz@linux.intel.com
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
(cherry picked from commit 0646547a67d25c407f5a4ac71b4eefe8b941202b)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
|
|
Improve and move diagnostics messages to the helper function to
keep the caller function tidy.
Signed-off-by: Michal Wajdeczko <michal.wajdeczko@intel.com>
Reviewed-by: Michał Winiarski <michal.winiarski@intel.com>
Link: https://patch.msgid.link/20260911182306.14973-1-michal.wajdeczko@intel.com
(cherry picked from commit 10628c52a3732a10426499a3d462cc2e6bc371ae)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
|
|
Disable VRR DC balance by default due to timing issues observed on some
panel/TCON combinations.
Keep the module parameter to enable DC balance during debugging and to
isolate DC balance effects from underlying VRR/display timing issues.
--v2:
- Make enable_dc_balance a bool and keep it disabled by default; fix the
parameter type/value mismatch and correct the description (Chaitanya
Kumar Borah, Jani Nikula)
- Explain in the commit message why the feature is gated and why a
module parameter is used (Jani Nikula)
--v3:
- Commit message update (Jani Nikula)
Fixes: 555819270707 ("drm/i915/vrr: Enable DC Balance")
Cc: <stable@vger.kernel.org> # v7.0+
Signed-off-by: Mitul Golani <mitulkumar.ajitkumar.golani@intel.com>
Reviewed-by: Ankit Nautiyal <ankit.k.nautiyal@intel.com>
Signed-off-by: Ankit Nautiyal <ankit.k.nautiyal@intel.com>
Link: https://patch.msgid.link/20260917074322.2606738-1-mitulkumar.ajitkumar.golani@intel.com
(cherry picked from commit d1ef78f0581e856c4238c751e9ae2884ce58c275)
Signed-off-by: Jani Nikula <jani.nikula@intel.com>
|
|
https://gitlab.freedesktop.org/drm/misc/kernel into drm-fixes
A number of fixes:
- bridge:
- samsung-dsim: fix GPIO lifetime
- client: Null pointer dereference fix
- imagination: error handling fix, page handling fix
- nouveau: fix reference leaks, double-frees, out-of-bounds accesses,
use-after-frees, don't reject config without SCDC, a number of
workarounds
- virtio: fix memory leak, reference leaks, null pointer dereference,
add pixel blend mode, cache coherency fix
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Maxime Ripard <self@mripard.dev>
Link: https://patch.msgid.link/arU22zzqUGDEco1y@houat
|
|
https://gitlab.freedesktop.org/drm/amdgpu/kernel into drm-fixes
amd-drm-fixes-7.3-2026-09-24:
amdgpu:
- Display ref count fix
- Userq fixes
- VCN 4, 5 reset fixes
- Fixes for various error paths
- Stack frame size fixes for various combinations of compilers and configs
amdkfd:
- Possible UAF fix
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Alex Deucher <alexander.deucher@amd.com>
Link: https://patch.msgid.link/20260924172938.634777-1-alexander.deucher@amd.com
|
|
https://gitlab.freedesktop.org/drm/xe/kernel into drm-fixes
Fixes in:
- CRI throttle reasons report (Sk)
- TLB invalidation at wedge (Shuicheng)
- SVM eviction and VM close (Brost)
- Display corruption on LNL on Xen PV (Szymon)
- W/a fix and addition (Tilak)
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Rodrigo Vivi <rodrigo.vivi@intel.com>
Link: https://patch.msgid.link/arUoUf9LpsJpJouN@intel.com
|
|
Some configs with gcc are also now affected.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 76b9706e7fa1a6e374703170128b1be2f590cda7)
|
|
The firmware reconstruction count controls accesses to the request's
fixed freelist ID array and the copy into the fixed response array.
Neither access currently bounds the count to those protocol arrays.
Clamp the count to the request capacity, which is shared by the response
layout, and use that count consistently for reconstruction and response
publication. Keep the firmware recovery exchange instead of dropping an
oversized request without a response, as discussed with the firmware
maintainer.
The issue was found by our static-analysis tool.
Fixes: 6eedddab733b ("drm/imagination: Implement free list and HWRT create and destroy ioctls")
Assisted-by: gpt 5
Signed-off-by: Pengpeng Hou <hppiscas@163.com>
Reviewed-by: Alessio Belle <alessio.belle@imgtec.com>
Link: https://patch.msgid.link/20260920034329.16614-1-hppiscas@163.com
Signed-off-by: Brajesh Gupta <brajesh.gupta@imgtec.com>
|
|
The GPU virtual start address wasn't included in the calculation for the
amount of page tables required for mapping a BO object in map() interface.
It resulted in map failure later due to not enough pages at L0/L1 level.
Update pvr_mmu_op_context_create() interface to pass device address as well
to allow correct calculation for page table memory.
If L0 tables cover 2MB (0x200000), the range defined by device address
0x80001ff000 (general heap at 2MB - 4KB) and size 0x2000 (two 4KB pages)
requires two L0 pages to be mapped, but without the base
address a range of 0x2000 computes to a single L0 page which is not enough.
Fixes: ff5f643de0bf ("drm/imagination: Add GEM and VM related code")
Reviewed-by: Alexandru Dadu <alexandru.dadu@imgtec.com>
Reviewed-by: Alessio Belle <alessio.belle@imgtec.com>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260922-mmu_fix-v4-2-12f1a871456a@imgtec.com
Signed-off-by: Brajesh Gupta <brajesh.gupta@imgtec.com>
|
|
Map failure from pvr_mmu_map_sgl() interface was not returned correctly
to pvr_mmu_map() interface. This resulted in pvr_mmu_map() interface to
continue instead of returning an error to caller.
Fix it by returning a proper error code from pvr_mmu_map_sgl() interface.
Call stack for crash:
[ 1179.286237] Unable to handle kernel NULL pointer dereference at virtual address 0000000000000008
[ 1179.295067] Mem abort info:
[ 1179.297877] ESR = 0x0000000096000004
[ 1179.301656] EC = 0x25: DABT (current EL), IL = 32 bits
[ 1179.306987] SET = 0, FnV = 0
[ 1179.310048] EA = 0, S1PTW = 0
[ 1179.313198] FSC = 0x04: level 0 translation fault
[ 1179.318096] Data abort info:
[ 1179.320993] ISV = 0, ISS = 0x00000004, ISS2 = 0x00000000
[ 1179.326483] CM = 0, WnR = 0, TnD = 0, TagAccess = 0
[ 1179.331546] GCS = 0, Overlay = 0, DirtyBit = 0, Xs = 0
[ 1179.336895] user pgtable: 4k pages, 48-bit VAs, pgdp=000000009822a000
[ 1179.343402] [0000000000000008] pgd=0000000000000000, p4d=0000000000000000
[ 1179.350243] Internal error: Oops: 0000000096000004 [#2] SMP
[ 1179.355908] Modules linked in: powervr gpu_sched drm_shmem_helper drm_gpuvm drm_exec xhci_plat_hcd xhci_hcd dwc3 usbcore usb_common snd_soc_simple_card snd_soc_simple_card_utils dwc3_am62 at24 sa2ul sha512 libsha512 sha256 authenc sch_fq_codel fuse dm_mod ipv6
[ 1179.378992] CPU: 1 UID: 1000 PID: 680 Comm: deqp-vk Tainted: G D 6.17.0 #1 PREEMPT
[ 1179.388120] Tainted: [D]=DIE
[ 1179.390994] Hardware name: Texas Instruments AM625 SK (DT)
[ 1179.396467] pstate: 00000005 (nzcv daif -PAN -UAO -TCO -DIT -SSBS BTYPE=--)
[ 1179.403415] pc : pvr_mmu_op_context_unmap_curr_page+0x6c/0x134 [powervr]
[ 1179.410140] lr : pvr_mmu_op_context_unmap_curr_page+0x58/0x134 [powervr]
[ 1179.416848] sp : ffff8000839ab8c0
[ 1179.420153] x29: ffff8000839ab8c0 x28: 0000000000000001 x27: 000000008f386000
[ 1179.427283] x26: ffff000016d1df98 x25: 0000000000247000 x24: 00000000000001e6
[ 1179.434413] x23: 0000000000000002 x22: 000000000000ffff x21: 0000000000000247
[ 1179.441540] x20: 0000000000000245 x19: ffff000016d1df60 x18: 0000000000000002
[ 1179.448668] x17: 0000000000000000 x16: 0000000000000000 x15: 0000000000000001
[ 1179.455793] x14: 0000000000060810 x13: ffff80007fffffff x12: ffff000004190480
[ 1179.462921] x11: ffff8000853f7000 x10: ffff8000811ae000 x9 : ffff0000041900b8
[ 1179.470051] x8 : 0000000000000000 x7 : 00000000990c4001 x6 : 0000000000000007
[ 1179.477177] x5 : ffff000016d1df60 x4 : 0000000000000000 x3 : ffff00000a7d8000
[ 1179.484306] x2 : 00000000000001ff x1 : 0000000000000000 x0 : 0000000000000000
[ 1179.491433] Call trace:
[ 1179.493872] pvr_mmu_op_context_unmap_curr_page+0x6c/0x134 [powervr] (P)
[ 1179.500582] pvr_mmu_map+0x31c/0x388 [powervr]
[ 1179.505027] pvr_vm_gpuva_map+0x40/0x88 [powervr]
[ 1179.509732] __drm_gpuvm_sm_map+0x250/0x44c [drm_gpuvm]
[ 1179.514952] drm_gpuvm_sm_map+0x48/0x5c [drm_gpuvm]
[ 1179.519822] pvr_vm_bind_op_exec+0x64/0x70 [powervr]
[ 1179.524785] pvr_vm_map+0x1f8/0x2a8 [powervr]
[ 1179.529142] pvr_ioctl_vm_map+0x12c/0x188 [powervr]
[ 1179.534018] drm_ioctl_kernel+0xb8/0x128
[ 1179.537941] drm_ioctl+0x21c/0x4ec
[ 1179.541337] __arm64_sys_ioctl+0xac/0x108
[ 1179.545344] invoke_syscall+0x44/0x100
[ 1179.549091] el0_svc_common.constprop.0+0x40/0xe0
[ 1179.553790] do_el0_svc+0x1c/0x28
[ 1179.557106] el0_svc+0x34/0xf0
[ 1179.560159] el0t_64_sync_handler+0xd0/0xe4
[ 1179.564334] el0t_64_sync+0x198/0x19c
[ 1179.567996] Code: 54000300 35000360 f9402261 79409a62 (f9400421)
[ 1179.574081] ---[ end trace 0000000000000000 ]---
Fixes: ff5f643de0bf ("drm/imagination: Add GEM and VM related code")
Reviewed-by: Alexandru Dadu <alexandru.dadu@imgtec.com>
Reviewed-by: Alessio Belle <alessio.belle@imgtec.com>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260922-mmu_fix-v4-1-12f1a871456a@imgtec.com
Signed-off-by: Brajesh Gupta <brajesh.gupta@imgtec.com>
|
|
[Why&How]
When building the DML files with clang without any sanitizer or LTO,
the following -Wframe-larger-than errors break the build under
CONFIG_WERROR:
display_mode_vba_30.c: error: stack frame size (2512) exceeds limit
(2048) in 'dml30_ModeSupportAndSystemConfigurationFull'
display_mode_vba_31.c: error: stack frame size (2416) exceeds limit
(2048) in 'dml31_ModeSupportAndSystemConfigurationFull'
display_mode_vba_314.c: error: stack frame size (2392) exceeds limit
(2048) in 'dml314_ModeSupportAndSystemConfigurationFull'
Clang consistently spills more than gcc, pushing the frame past the 2048
byte limit.
Apply an existing approach of increasing the warn stack size to the
non-sanitizer path so plain clang builds use a 3072 byte limit.
Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5642
Signed-off-by: Ivan Lipski <ivan.lipski@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 21711b6e66bb7b41b1aec67b2d99aafe768c8fcb)
Cc: stable@vger.kernel.org
|
|
[WHY]
UBSAN instrumentation adds checks and handler calls and increases
stack usage in the large DML calculation functions, similar to
KASAN and KCSAN. With UBSAN enabled these files exceed the default
-Wframe-larger-than limit and fail to build when -Werror is in effect.
Reproduced with LLVM (make LLVM=1, clang 19.1.1), CONFIG_UBSAN=y,
CONFIG_GCOV_PROFILE_ALL=y and CONFIG_DRM_AMDGPU_WERROR=y on x86_64.
[HOW]
Include CONFIG_UBSAN in the sanitizer check that selects the
higher per-file frame warning limit in the dml and dml2_0
Makefiles.
Suggested-by: Leo Li <sunpeng.li@amd.com>
Assisted-by: Copilot:Claude-Opus-5.5
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit ebf8b0fd8508b744f85a8eee82b745b1d3502dd0)
Cc: stable@vger.kernel.org
|
|
amdgpu_debugfs_test_ib_show() resumes the device with
pm_runtime_get_sync() before taking the reset domain semaphore with
down_write_killable(). If the write lock acquisition is interrupted,
the function returns without calling pm_runtime_put_autosuspend(),
leaking the runtime PM reference acquired for the device and keeping
the GPU awake.
Drop the runtime PM reference on the interrupted down_write_killable()
error path before returning.
Fixes: 6049db43d6dd ("drm/amdgpu: change reset lock from mutex to rw_semaphore")
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit ec30a576c2d4c0364549e6c04218f50704ef56c8)
Cc: stable@vger.kernel.org
|
|
amdgpu_vm_init() initializes vm->last_update, vm->last_unlocked and
vm->last_tlb_flush with references to the stub fence taken via
dma_fence_get_stub(). The error label at the end of the function
releases the last_unlocked and last_tlb_flush references with
dma_fence_put(), but the reference stored in vm->last_update is never
dropped, so whenever the page table root creation, the reservation of
the root BO or the PASID registration fails, the stub fence reference
leaks.
Drop the vm->last_update reference together with the other stub fence
references on the error path.
Fixes: 187916e6ed9d ("drm/amdgpu: install stub fence into potential unused fence pointers")
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit e7979c84fc05a176bdf855ee664871b1648404c9)
Cc: stable@vger.kernel.org
|