| Age | Commit message (Collapse) | Author |
|
https://gitlab.freedesktop.org/drm/misc/kernel into drm-fixes
A number of fixes:
- bridge:
- samsung-dsim: fix GPIO lifetime
- client: Null pointer dereference fix
- imagination: error handling fix, page handling fix
- nouveau: fix reference leaks, double-frees, out-of-bounds accesses,
use-after-frees, don't reject config without SCDC, a number of
workarounds
- virtio: fix memory leak, reference leaks, null pointer dereference,
add pixel blend mode, cache coherency fix
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Maxime Ripard <self@mripard.dev>
Link: https://patch.msgid.link/arU22zzqUGDEco1y@houat
|
|
https://gitlab.freedesktop.org/drm/amdgpu/kernel into drm-fixes
amd-drm-fixes-7.3-2026-09-24:
amdgpu:
- Display ref count fix
- Userq fixes
- VCN 4, 5 reset fixes
- Fixes for various error paths
- Stack frame size fixes for various combinations of compilers and configs
amdkfd:
- Possible UAF fix
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Alex Deucher <alexander.deucher@amd.com>
Link: https://patch.msgid.link/20260924172938.634777-1-alexander.deucher@amd.com
|
|
https://gitlab.freedesktop.org/drm/xe/kernel into drm-fixes
Fixes in:
- CRI throttle reasons report (Sk)
- TLB invalidation at wedge (Shuicheng)
- SVM eviction and VM close (Brost)
- Display corruption on LNL on Xen PV (Szymon)
- W/a fix and addition (Tilak)
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Rodrigo Vivi <rodrigo.vivi@intel.com>
Link: https://patch.msgid.link/arUoUf9LpsJpJouN@intel.com
|
|
Some configs with gcc are also now affected.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 76b9706e7fa1a6e374703170128b1be2f590cda7)
|
|
The firmware reconstruction count controls accesses to the request's
fixed freelist ID array and the copy into the fixed response array.
Neither access currently bounds the count to those protocol arrays.
Clamp the count to the request capacity, which is shared by the response
layout, and use that count consistently for reconstruction and response
publication. Keep the firmware recovery exchange instead of dropping an
oversized request without a response, as discussed with the firmware
maintainer.
The issue was found by our static-analysis tool.
Fixes: 6eedddab733b ("drm/imagination: Implement free list and HWRT create and destroy ioctls")
Assisted-by: gpt 5
Signed-off-by: Pengpeng Hou <hppiscas@163.com>
Reviewed-by: Alessio Belle <alessio.belle@imgtec.com>
Link: https://patch.msgid.link/20260920034329.16614-1-hppiscas@163.com
Signed-off-by: Brajesh Gupta <brajesh.gupta@imgtec.com>
|
|
The GPU virtual start address wasn't included in the calculation for the
amount of page tables required for mapping a BO object in map() interface.
It resulted in map failure later due to not enough pages at L0/L1 level.
Update pvr_mmu_op_context_create() interface to pass device address as well
to allow correct calculation for page table memory.
If L0 tables cover 2MB (0x200000), the range defined by device address
0x80001ff000 (general heap at 2MB - 4KB) and size 0x2000 (two 4KB pages)
requires two L0 pages to be mapped, but without the base
address a range of 0x2000 computes to a single L0 page which is not enough.
Fixes: ff5f643de0bf ("drm/imagination: Add GEM and VM related code")
Reviewed-by: Alexandru Dadu <alexandru.dadu@imgtec.com>
Reviewed-by: Alessio Belle <alessio.belle@imgtec.com>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260922-mmu_fix-v4-2-12f1a871456a@imgtec.com
Signed-off-by: Brajesh Gupta <brajesh.gupta@imgtec.com>
|
|
Map failure from pvr_mmu_map_sgl() interface was not returned correctly
to pvr_mmu_map() interface. This resulted in pvr_mmu_map() interface to
continue instead of returning an error to caller.
Fix it by returning a proper error code from pvr_mmu_map_sgl() interface.
Call stack for crash:
[ 1179.286237] Unable to handle kernel NULL pointer dereference at virtual address 0000000000000008
[ 1179.295067] Mem abort info:
[ 1179.297877] ESR = 0x0000000096000004
[ 1179.301656] EC = 0x25: DABT (current EL), IL = 32 bits
[ 1179.306987] SET = 0, FnV = 0
[ 1179.310048] EA = 0, S1PTW = 0
[ 1179.313198] FSC = 0x04: level 0 translation fault
[ 1179.318096] Data abort info:
[ 1179.320993] ISV = 0, ISS = 0x00000004, ISS2 = 0x00000000
[ 1179.326483] CM = 0, WnR = 0, TnD = 0, TagAccess = 0
[ 1179.331546] GCS = 0, Overlay = 0, DirtyBit = 0, Xs = 0
[ 1179.336895] user pgtable: 4k pages, 48-bit VAs, pgdp=000000009822a000
[ 1179.343402] [0000000000000008] pgd=0000000000000000, p4d=0000000000000000
[ 1179.350243] Internal error: Oops: 0000000096000004 [#2] SMP
[ 1179.355908] Modules linked in: powervr gpu_sched drm_shmem_helper drm_gpuvm drm_exec xhci_plat_hcd xhci_hcd dwc3 usbcore usb_common snd_soc_simple_card snd_soc_simple_card_utils dwc3_am62 at24 sa2ul sha512 libsha512 sha256 authenc sch_fq_codel fuse dm_mod ipv6
[ 1179.378992] CPU: 1 UID: 1000 PID: 680 Comm: deqp-vk Tainted: G D 6.17.0 #1 PREEMPT
[ 1179.388120] Tainted: [D]=DIE
[ 1179.390994] Hardware name: Texas Instruments AM625 SK (DT)
[ 1179.396467] pstate: 00000005 (nzcv daif -PAN -UAO -TCO -DIT -SSBS BTYPE=--)
[ 1179.403415] pc : pvr_mmu_op_context_unmap_curr_page+0x6c/0x134 [powervr]
[ 1179.410140] lr : pvr_mmu_op_context_unmap_curr_page+0x58/0x134 [powervr]
[ 1179.416848] sp : ffff8000839ab8c0
[ 1179.420153] x29: ffff8000839ab8c0 x28: 0000000000000001 x27: 000000008f386000
[ 1179.427283] x26: ffff000016d1df98 x25: 0000000000247000 x24: 00000000000001e6
[ 1179.434413] x23: 0000000000000002 x22: 000000000000ffff x21: 0000000000000247
[ 1179.441540] x20: 0000000000000245 x19: ffff000016d1df60 x18: 0000000000000002
[ 1179.448668] x17: 0000000000000000 x16: 0000000000000000 x15: 0000000000000001
[ 1179.455793] x14: 0000000000060810 x13: ffff80007fffffff x12: ffff000004190480
[ 1179.462921] x11: ffff8000853f7000 x10: ffff8000811ae000 x9 : ffff0000041900b8
[ 1179.470051] x8 : 0000000000000000 x7 : 00000000990c4001 x6 : 0000000000000007
[ 1179.477177] x5 : ffff000016d1df60 x4 : 0000000000000000 x3 : ffff00000a7d8000
[ 1179.484306] x2 : 00000000000001ff x1 : 0000000000000000 x0 : 0000000000000000
[ 1179.491433] Call trace:
[ 1179.493872] pvr_mmu_op_context_unmap_curr_page+0x6c/0x134 [powervr] (P)
[ 1179.500582] pvr_mmu_map+0x31c/0x388 [powervr]
[ 1179.505027] pvr_vm_gpuva_map+0x40/0x88 [powervr]
[ 1179.509732] __drm_gpuvm_sm_map+0x250/0x44c [drm_gpuvm]
[ 1179.514952] drm_gpuvm_sm_map+0x48/0x5c [drm_gpuvm]
[ 1179.519822] pvr_vm_bind_op_exec+0x64/0x70 [powervr]
[ 1179.524785] pvr_vm_map+0x1f8/0x2a8 [powervr]
[ 1179.529142] pvr_ioctl_vm_map+0x12c/0x188 [powervr]
[ 1179.534018] drm_ioctl_kernel+0xb8/0x128
[ 1179.537941] drm_ioctl+0x21c/0x4ec
[ 1179.541337] __arm64_sys_ioctl+0xac/0x108
[ 1179.545344] invoke_syscall+0x44/0x100
[ 1179.549091] el0_svc_common.constprop.0+0x40/0xe0
[ 1179.553790] do_el0_svc+0x1c/0x28
[ 1179.557106] el0_svc+0x34/0xf0
[ 1179.560159] el0t_64_sync_handler+0xd0/0xe4
[ 1179.564334] el0t_64_sync+0x198/0x19c
[ 1179.567996] Code: 54000300 35000360 f9402261 79409a62 (f9400421)
[ 1179.574081] ---[ end trace 0000000000000000 ]---
Fixes: ff5f643de0bf ("drm/imagination: Add GEM and VM related code")
Reviewed-by: Alexandru Dadu <alexandru.dadu@imgtec.com>
Reviewed-by: Alessio Belle <alessio.belle@imgtec.com>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260922-mmu_fix-v4-1-12f1a871456a@imgtec.com
Signed-off-by: Brajesh Gupta <brajesh.gupta@imgtec.com>
|
|
[Why&How]
When building the DML files with clang without any sanitizer or LTO,
the following -Wframe-larger-than errors break the build under
CONFIG_WERROR:
display_mode_vba_30.c: error: stack frame size (2512) exceeds limit
(2048) in 'dml30_ModeSupportAndSystemConfigurationFull'
display_mode_vba_31.c: error: stack frame size (2416) exceeds limit
(2048) in 'dml31_ModeSupportAndSystemConfigurationFull'
display_mode_vba_314.c: error: stack frame size (2392) exceeds limit
(2048) in 'dml314_ModeSupportAndSystemConfigurationFull'
Clang consistently spills more than gcc, pushing the frame past the 2048
byte limit.
Apply an existing approach of increasing the warn stack size to the
non-sanitizer path so plain clang builds use a 3072 byte limit.
Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5642
Signed-off-by: Ivan Lipski <ivan.lipski@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 21711b6e66bb7b41b1aec67b2d99aafe768c8fcb)
Cc: stable@vger.kernel.org
|
|
[WHY]
UBSAN instrumentation adds checks and handler calls and increases
stack usage in the large DML calculation functions, similar to
KASAN and KCSAN. With UBSAN enabled these files exceed the default
-Wframe-larger-than limit and fail to build when -Werror is in effect.
Reproduced with LLVM (make LLVM=1, clang 19.1.1), CONFIG_UBSAN=y,
CONFIG_GCOV_PROFILE_ALL=y and CONFIG_DRM_AMDGPU_WERROR=y on x86_64.
[HOW]
Include CONFIG_UBSAN in the sanitizer check that selects the
higher per-file frame warning limit in the dml and dml2_0
Makefiles.
Suggested-by: Leo Li <sunpeng.li@amd.com>
Assisted-by: Copilot:Claude-Opus-5.5
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit ebf8b0fd8508b744f85a8eee82b745b1d3502dd0)
Cc: stable@vger.kernel.org
|
|
amdgpu_debugfs_test_ib_show() resumes the device with
pm_runtime_get_sync() before taking the reset domain semaphore with
down_write_killable(). If the write lock acquisition is interrupted,
the function returns without calling pm_runtime_put_autosuspend(),
leaking the runtime PM reference acquired for the device and keeping
the GPU awake.
Drop the runtime PM reference on the interrupted down_write_killable()
error path before returning.
Fixes: 6049db43d6dd ("drm/amdgpu: change reset lock from mutex to rw_semaphore")
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit ec30a576c2d4c0364549e6c04218f50704ef56c8)
Cc: stable@vger.kernel.org
|
|
amdgpu_vm_init() initializes vm->last_update, vm->last_unlocked and
vm->last_tlb_flush with references to the stub fence taken via
dma_fence_get_stub(). The error label at the end of the function
releases the last_unlocked and last_tlb_flush references with
dma_fence_put(), but the reference stored in vm->last_update is never
dropped, so whenever the page table root creation, the reservation of
the root BO or the PASID registration fails, the stub fence reference
leaks.
Drop the vm->last_update reference together with the other stub fence
references on the error path.
Fixes: 187916e6ed9d ("drm/amdgpu: install stub fence into potential unused fence pointers")
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit e7979c84fc05a176bdf855ee664871b1648404c9)
Cc: stable@vger.kernel.org
|
|
amdgpu_acpi_enumerate_xcc() looks up each XCC ACPI device with
acpi_dev_get_first_match_dev(), which takes a reference to the device.
The reference is dropped with acpi_dev_put() after the XCC info is
initialized, but if the kzalloc_obj() allocation of the XCC info fails
the function returns -ENOMEM without releasing the reference, leaking
the last reference to the ACPI device.
Drop the ACPI device reference on the allocation failure path before
returning.
Fixes: 4d5275ab0b18 ("drm/amdgpu: Add parsing of acpi xcc objects")
Reviewed-by: Lijo Lazar <lijo.lazar@amd.com>
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 9211ef48b31ec66999cf55e04d0cbc60cd855fd5)
Cc: stable@vger.kernel.org
|
|
amdgpu_ring_init() initializes ring->vmid_wait with a reference to the
stub fence taken via dma_fence_get_stub(). When a later step of the
initialization fails, e.g. amdgpu_fence_driver_init_ring(), a writeback
slot allocation or the ring buffer allocation, the function returns an
error without releasing the stub fence reference and the reference is
leaked if the ring is torn down without amdgpu_ring_fini().
Move the stub fence assignment to the end of the initialization, right
before the ring is registered with the GPU scheduler, where no further
failure is possible. The stub fence is only consumed by command
submission handling in amdgpu_ids.c once the ring is up and running, so
nothing reads it during the error-prone part of the initialization.
Fixes: 48e9fbd1a284 ("drm/amdgpu: initialize the vmid_wait with the stub fence")
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit f2b96986851203e9c50ca0d13aaa3581ca3e8ebd)
Cc: stable@vger.kernel.org
|
|
kfd_dev_mapping caches the address_space of the first /dev/kfd opener
so that the GPU reset path can call unmap_mapping_range() to zap all
userspace mappings of doorbell and MMIO ranges. This design has two
bugs that both manifest under SRIOV with multiple containers:
1. Use-after-free / rwsem deadlock. The cached pointer refers to an
inode owned by the first opener's container. When that container
exits and its inode is released, kfd_dev_mapping becomes a dangling
pointer. A subsequent GPU reset dereferences it inside
unmap_mapping_range(), which takes i_mmap_rwsem on the freed inode,
causing a hard hang observable as an uninterruptible rwsem wait.
2. Multi-container gap. Only the first opener's address_space is cached;
VMAs created by later openers live in a different address_space and
are never reached by unmap_mapping_range(). After a GPU reset those
stale mappings keep doorbell and MMIO pages accessible to guest
userspace with no GPU behind them, risking PCIe transaction timeouts
and NMI panics.
Fix both bugs with the same approach used by DRM core (drm_drv.c):
create a private pseudo-filesystem at module init time and allocate one
anonymous inode from it. In kfd_open() redirect every opener's
file->f_mapping to that inode's address_space. The inode is
module-owned, lives exactly as long as the amdgpu module, and collects
VMAs from all openers in one address_space. A single
unmap_mapping_range() call in the reset path then correctly reaches
every container's mappings with no dangling pointer risk.
The hang manifests as an NMI backtrace on the GPU reset workqueue stuck
spinning in rwsem_down_read_slowpath() with a corrupted i_mmap_rwsem:
Workqueue: amdgpu-reset-dev xgpu_ai_mailbox_flr_work [amdgpu]
Call Trace:
<TASK>
kvm_wait+0x1f/0x40
__pv_queued_spin_lock_slowpath+0x31d/0x3a0
_raw_spin_lock_irq+0x51/0x80
rwsem_down_read_slowpath+0xb3/0x550
down_read+0x48/0xd0
unmap_mapping_range+0x71/0x140
kfd_dev_unmap_mapping_range+0x5b/0x140 [amdgpu]
amdgpu_amdkfd_clear_kfd_mapping+0xd8/0x190 [amdgpu]
amdgpu_device_gpu_recover+0x232/0x450 [amdgpu]
xgpu_ai_mailbox_flr_work+0xb5/0xc0 [amdgpu]
process_one_work+0x18e/0x3e0
worker_thread+0x2e3/0x420
kthread+0x10a/0x230
Fixes: 70cadefcc616 ("drm/amdgpu: unmap all user mappings of framebuffer and doorbell before mode1 reset")
Signed-off-by: Asad Kamal <asad.kamal@amd.com>
Reviewed-by: Lijo Lazar <lijo.lazar@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 1128b4a52de1572e87431de837fd9850cb99542c)
Cc: stable@vger.kernel.org
|
|
vcn_v4_0_3_reset_jpeg_pre_helper() passes adev->video_timeout directly
to amdgpu_fence_wait_polling(), whose timeout parameter is documented
and implemented in usecs (busy-wait loop decrementing by udelay(2)).
adev->video_timeout is set in jiffies by
amdgpu_device_get_job_timeout_settings(), via msecs_to_jiffies().
Passing it unconverted means the intended ~2s wait for outstanding
JPEG fences to complete before the JPEG queue is torn down actually
lasts only a couple of microseconds (HZ jiffies interpreted as usecs),
so pending jobs are almost never given a real chance to finish before
the reset path forces completion in the following helper.
Convert the jiffies value to usecs with jiffies_to_usecs() before
passing it to amdgpu_fence_wait_polling().
Fixes: d25c67fd9d6f ("drm/amdgpu/vcn4.0.3: rework reset handling")
Cc: Jesse.Zhang <Jesse.Zhang@amd.com>
Assisted-by: Claude:claude-sonnet-5
Signed-off-by: Sunil Khatri <sunil.khatri@amd.com>
Reviewed-by: Christian König <christian.koenig@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 5feabbd673c10ebee22b880e4d812f08974d2ef7)
Cc: stable@vger.kernel.org
|
|
vcn_v5_0_1_reset_jpeg_pre_helper() passes adev->video_timeout directly
to amdgpu_fence_wait_polling(), whose timeout parameter is documented
and implemented in usecs (busy-wait loop decrementing by udelay(2)).
adev->video_timeout is set in jiffies by
amdgpu_device_get_job_timeout_settings(), via msecs_to_jiffies().
Passing it unconverted means the intended ~2s wait for outstanding
JPEG fences to complete before the JPEG queue is torn down actually
lasts only a couple of microseconds (HZ jiffies interpreted as usecs),
so pending jobs are almost never given a real chance to finish before
the reset path forces completion in the following helper.
Convert the jiffies value to usecs with jiffies_to_usecs() before
passing it to amdgpu_fence_wait_polling().
Fixes: fab47d2db5ca ("drm/amdgpu/vcn5.0.1: rework reset handling")
Cc: Jesse.Zhang <Jesse.Zhang@amd.com>
Assisted-by: Claude:claude-sonnet-5
Signed-off-by: Sunil Khatri <sunil.khatri@amd.com>
Reviewed-by: Christian König <christian.koenig@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit b8334fec8b90ebffcaa01001a23edca9f29a05e9)
Cc: stable@vger.kernel.org
|
|
Function amdgpu_userq_start_hang_detect_work() calls msecs_to_jiffies()
on adev->gfx_timeout/compute_timeout/sdma_timeout before arming
hang_detect_work. These timeout values already hold jiffies values from
amdgpu_device_get_job_timeout_settings() at device init.
This silently shrinks the real hang-detect deadline to (2 * HZ) ms
instead of the intended timeout. e.g. 500ms instead of the 2000ms
default on a CONFIG_HZ=250 kernel, only coincidentally correct at
HZ=1000. The shortened window is easily exceeded by ordinary
fence-completion latency, causing hang_detect_work to fire and
trigger a per-queue or full GPU reset for queues that are not
actually hung.
Pass the jiffies value directly to queue_delayed_work() instead of
converting it a second time.
Fixes: fc3336be9c62 ("drm/amd/amdgpu: Add independent hang detect work for user queue fence")
Signed-off-by: Sunil Khatri <sunil.khatri@amd.com>
Reviewed-by: Christian König <christian.koenig@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 13d44ca033cb74756c2aef0ade54a75cdf2f6271)
Cc: stable@vger.kernel.org
|
|
The eviction fence suspend worker waits for every pending userq fence
from inside a dma_fence_begin_signalling() critical section. Waiting on
another DMA fence while responsible for signalling one violates the
cross-driver fence contract and is reported by lockdep as a
dma_fence_map dependency.
Move the wait before dma_fence_begin_signalling(). Keep userq_mutex held
so queue lifetime remains stable while inspecting last_fence.
Fixes: fc61df151617 ("drm/amdgpu: annotate eviction fence signaling path")
Signed-off-by: Prike Liang <Prike.Liang@amd.com>
Reviewed-by: Vitaly Prosyak <vitaly.prosyak@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 3bd4fbc5ed89621340b5cd249869092691a9c81f)
Cc: stable@vger.kernel.org
|
|
In dm_update_crtc_state(), when a modeset is required the newly created
stream is stored in dm_new_crtc_state->stream and an extra reference is
taken with dc_stream_retain(). The reference returned by
create_validate_stream_for_sink() is released as an extra reference at
the skip_modeset label, leaving the stream owned by the new CRTC state.
If amdgpu_dm_check_crtc_color_mgmt() fails afterwards, the code jumps
to the fail label which releases new_stream again. Since the extra
reference was already released at skip_modeset, this drops the
reference owned by dm_new_crtc_state->stream and the stream is
released while the atomic state still points to it, leading to a
premature free of the dc stream.
Set new_stream to NULL after releasing the extra reference at the
skip_modeset label so that a later goto fail cannot release the
reference owned by the new CRTC state.
Fixes: 7cd4b70091a5 ("drm/amd/display: Rework CRTC color management")
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 102a47065a62dc8f6bbbb47cf082a2934282eb08)
Cc: stable@vger.kernel.org
|
|
Avoid programming the IDLEDLY timer to less than 5 microseconds.
Apply wa_14025941587 to Graphics Versions 20.01 to 35.11
and Media Versions 13.01 to 35.03
v2: Use xe_rtp_match_not_sriov_vf, move to local variable
Remove warn and other knits - Matt R
v3: Add verbose comment - Tejas
v4: Restore IDLE_DLY register on engine reset.
Add it to GUC save-restore list. -Vivek
v5: Extend WA to Media Versions 13.01 to 35.03 - Vinay
v6: Avoid clearing inhibit switch - Bala
Refactor code accordingly by adding idle_reg_val.
v7: Rebased with the divide-by-zero/overflow guards living in
a separate hardening patch.
v8: Preserve the Wa_16023105232 floor (DIV_ROUND_DOWN_ULL) and the
maxcnt == 0 guard from the hardening patch. Round up
(DIV_ROUND_UP_ULL) the Wa_14025941587 minimum conversion instead,
so the tick-quantized delay cannot round back below 5 us.
v9: Evaluate the Wa_16023105232 xe_gt_WARN_ON() against the value
read from hardware instead of the Wa_14025941587-bumped value,
so it no longer fires on the driver's own floor. Re-check the
rounded-up tick value against maxcnt and floor it if tick
quantization pushed it back to/above maxcnt, logging via
xe_gt_dbg since this is the driver's own value, not a hardware
anomaly.
Assisted-by: GitHub_Copilot:claude-opus-4.8
Signed-off-by: Tangudu Tilak Tirumalesh <tilak.tirumalesh.tangudu@intel.com>
Reviewed-by: Vinay Belgaumkar <vinay.belgaumkar@intel.com>
Link: https://patch.msgid.link/20260916100545.779894-3-tilak.tirumalesh.tangudu@intel.com
Signed-off-by: Matt Roper <matthew.d.roper@intel.com>
(cherry picked from commit 9453c528fc909076468ff10df1c2e334ca5a9b00)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
|
|
adjust_idledly() has several corner-case issues flagged during review:
1. If xe_gt_clock_init() failed to recognise the crystal clock,
gt->info.timestamp_base is 0, which makes idledly_units_ps also 0.
The subsequent DIV_ROUND_CLOSEST(..., idledly_units_ps) is then a
divide-by-zero and panics the kernel.
2. The tick-to-ns conversions are done in u32:
idledly * idledly_units_ps, (maxcnt - 1) * 1000
Both overflow u32 before DIV_ROUND_CLOSEST() sees them.
3. If IDLE_WAIT_TIME reads back as 0, maxcnt evaluates to 0 and
the maxcnt - 1 clamp wraps to 0xFFFFFFFF in u32.
4. The register only stores whole ticks, so the clamped ns value has
to be converted to ticks and back. DIV_ROUND_CLOSEST() can round
that conversion up past maxcnt:
maxcnt = 640 ns, one tick = 666664 ps
clamp: maxcnt - 1 = 639 ns
ns -> ticks: 639000 / 666664 = 0.958 -> rounds to 1 tick
tick -> ns: 1 * 666664 / 1000 = 667 ns
667 ns is programmed into RING_IDLEDLY, but 667 >= maxcnt (640),
so xe_gt_WARN_ON() fires again on every subsequent init.
Return early if timestamp_base is 0 (the unknown-crystal path).
Do the conversions in u64 via the *_ULL() helpers so they cannot wrap.
Clamp with a floor (DIV_ROUND_DOWN_ULL) so the programmed delay stays
strictly below maxcnt, and guard the maxcnt == 0 case with a zero delay
while still writing RING_IDLEDLY so INHIBIT_SWITCH_UNTIL_PREEMPTED is
cleared.
v2: Drop the redundant warn on the timestamp_base == 0 path;
xe_gt_clock_init() already warns on an unrecognised crystal clock.
Keep the early return to avoid the divide-by-zero. - Vinay
v3: Field-mask the RING_IDLEDLY write with REG_FIELD_PREP(IDLE_DELAY, ...)
instead of writing the raw tick count, which could clobber
INHIBIT_SWITCH_UNTIL_PREEMPTED and reserved bits. Split the
inhibit-switch clear from the maxcnt clamp so a set inhibit bit no
longer forces a needless delay overwrite when the delay itself is
already valid. Use gt_to_xe(gt) instead of gt_to_xe(hwe->gt).
Fixes: d2de4410a88f ("drm/xe: Apply Wa_16023105232")
Cc: stable@vger.kernel.org
Assisted-by: GitHub_Copilot:claude-opus-4.8
Signed-off-by: Tangudu Tilak Tirumalesh <tilak.tirumalesh.tangudu@intel.com>
Reviewed-by: Vinay Belgaumkar <vinay.belgaumkar@intel.com>
Link: https://patch.msgid.link/20260916100545.779894-2-tilak.tirumalesh.tangudu@intel.com
Signed-off-by: Matt Roper <matthew.d.roper@intel.com>
(cherry picked from commit d864065ea25e9d12897c175de9176ce46677e176)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
|
|
Fix display corruption on Xen PV dom0, where DMA buffers are not
guaranteed machine-contiguous, in which case bounce buffering kicks
in, breaking xe's memory coherency assumptions.
Apply the same workaround i915 carries in i915_sg_segment_size() since
commit 78a07fe777c4 ("drm/i915: stop abusing swiotlb_max_segment").
Fixes: dd08ebf6c352 ("drm/xe: Introduce a new DRM driver for Intel GPUs")
Reported-by: Marek Marczykowski-Górecki <marmarek@invisiblethingslab.com>
Closes: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8382
Link: https://lore.kernel.org/xen-devel/aYtznP_tT6xNPwf-@mail-itl/
Link: https://lore.kernel.org/all/20221020110308.1582518-1-hch@lst.de/ # i915 counterpart
Cc: Christoph Hellwig <hch@lst.de>
Cc: Robert Beckett <bob.beckett@collabora.com>
Cc: stable@vger.kernel.org # v6.8+
Signed-off-by: Szymon Acedański <accek@invisiblethingslab.com>
Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Signed-off-by: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Link: https://patch.msgid.link/20260916173030.3223833-1-accek@invisiblethingslab.com
(cherry picked from commit 77f704158f099b952681f207478a22d5b8218edb)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
|
|
i915_gem_busy_ioctl uses dma_resv_for_each_fence_unlocked() to iterate
over the fences in an GEM object without holding a reference but only
the RCU read side lock.
What can happen here is that the GEM object is destroyed concurrently
while i915_gem_busy_ioctl is still running. This won't free the GEM
objects memory, but still drops all the dma_fence references.
Now when dma_resv_for_each_fence_unlocked() sees a destroyed dma_fence it
assumes that a new fence list was installed and re-starts the loop.
But in the case of a destroyed GEM object a new fence list is never
installed, only the old one freed and therefore the iteration never
finishes resulting in an endless loop.
The solution is to drop the fence references only after the RCU grace
period.
The fixes tag is not necessary the patch introducing the problem, but the
one making it so worse that we need to address it.
This problem was pointed out by Sashiko-bot.
Signed-off-by: Christian König <christian.koenig@amd.com>
Fixes: 912ff2ebd695 ("drm/i915: use the new iterator in i915_gem_busy_ioctl v2")
CC: stable@vger.kernel.org
Reviewed-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
Signed-off-by: Tvrtko Ursulin <tursulin@ursulin.net>
Link: https://lore.kernel.org/r/20260903113621.54660-1-christian.koenig@amd.com
(cherry picked from commit 5113479556025093bf8133bb2dcaa33be2d50921)
Signed-off-by: Jani Nikula <jani.nikula@intel.com>
|
|
Use EXPORT_SYMBOL_IF_KUNIT() instead of the regular EXPORT_SYMBOL() to
export the symbols to the kunit namespace. Otherwise, the symbols get
exported for all the kernel to see, and the corresponding
MODULE_IMPORT_NS("EXPORTED_FOR_KUNIT_TESTING") in the tests is
meaningless.
Fixes: 2eb9982ff179 ("drm/i915/kunit: Export link training and caps funcs for testing")
Cc: Imre Deak <imre.deak@intel.com>
Reviewed-by: Imre Deak <imre.deak@intel.com>
Link: https://patch.msgid.link/20260915160620.779372-1-jani.nikula@intel.com
Signed-off-by: Jani Nikula <jani.nikula@intel.com>
(cherry picked from commit 4ffdb772716e4279d63dfaaadf965da73e799aeb)
|
|
During an atomic commit after all the MST stream CRTC state is computed
the driver ensures that the sum of TUs of all the streams on a given MST
topology link is within limits (63 for 8b10 and 64 for 128b132b). For a
disconnected stream the DRM MST core's BW verification doesn't ensure
this, because the topology state it uses for this is destroyed as soon
as the stream (i.e. MST connector/port) is disconnected. The driver
should keep the link state valid even for such disconnected streams, as
userspace may disable them one-by-one only in a deferred way. Ensure the
link's sum of TUs stays within limits in this case by simply reusing the
maximum link BPP limit from the stream's (i.e. CRTC's) old state.
The disconnection can happen either via the whole topology getting
disconnected or via only the given stream's port getting disconnected.
Check for both of these conditions separately, as a connector gets
unregistered after a link disconnect event only in a deferred way.
Cc: stable@vger.kernel.org # v6.10+
Link: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16073
Link: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16384
Reviewed-by: Luca Coelho <luciano.coelho@intel.com>
Signed-off-by: Imre Deak <imre.deak@intel.com>
Link: https://patch.msgid.link/20260907174413.741851-2-imre.deak@intel.com
(cherry picked from commit ee00f8fbb2b202002ab90834e02e9ba372773a36)
Signed-off-by: Jani Nikula <jani.nikula@intel.com>
|
|
During an atomic commit after all the MST stream CRTC state is computed
the driver ensures that the FEC is configured the same way (enabled or
disabled) for all the streams on a given MST topology's link.
drm_dp_mst_port_downstream_of_parent() used to determine if a stream is
downstream of an MST port will return false if the whole topology is
disconnected, since in that case it can't verify that the port/
parent_port passed to it is in the given MST topology. This is a problem
during the above FEC configuration check, since
intel_dp_mst_check_dsc_change()->get_pipes_downstream_of_mst_ports()
will not return all the stream CRTCs/pipes for the topology as expected.
Since passing parent_port==NULL to get_pipes_downstream_of_mst_port()
is meant to return all the streams for the given topology (i.e. mst_mgr)
skip checking if an MST port is downstream of a parent port in this
case.
This fixes a problem where the FEC configuration check explained above
failed to ensure that all streams' FEC is configured the same way if the
topology was disconnected, leading to a FEC state mismatch error.
Cc: stable@vger.kernel.org # v6.10+
Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16073
Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16384
Reviewed-by: Luca Coelho <luciano.coelho@intel.com>
Signed-off-by: Imre Deak <imre.deak@intel.com>
Link: https://patch.msgid.link/20260907174413.741851-1-imre.deak@intel.com
(cherry picked from commit 270681fbffbba2b6ccf5b7e3c34b8b563b36167f)
Signed-off-by: Jani Nikula <jani.nikula@intel.com>
|
|
The eDP panel on the HP Pavilion Plus Laptop 14-ew1xxx advertises HBR3
while leaving the TPS4 support bit clear. The output however flickers, once
link is trained with HBR3.
Until commit 8c9006283e4b ("Revert "drm/i915/dp: Reject HBR3 when sink
doesn't support TPS4"") such sinks were capped at HBR2 by the TPS4 check
which incidentally kept this panel stable. That check was reverted because
other panels legitimately need HBR3 without advertising TPS4, and the
per-machine QUIRK_EDP_LIMIT_RATE_HBR2 was introduced to handle the affected
machines instead.
Add the machine to the list of devices that need the
QUIRK_EDP_LIMIT_RATE_HBR2.
Fixes: 8c9006283e4b ("Revert "drm/i915/dp: Reject HBR3 when sink doesn't support TPS4"")
Reported-by: Annoy Cc <annoycc@gmail.com>
Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16743
Cc: <stable@vger.kernel.org> # v6.18+
Tested-by: Annoy Cc <annoycc@gmail.com>
Signed-off-by: Ankit Nautiyal <ankit.k.nautiyal@intel.com>
Reviewed-by: Nemesa Garg <nemesa.garg@intel.com>
Link: https://patch.msgid.link/20260907034555.2753846-1-ankit.k.nautiyal@intel.com
(cherry picked from commit 550b703fdbb2a2022faa75b4b11ab135241afbd9)
Signed-off-by: Jani Nikula <jani.nikula@intel.com>
|
|
Selective fetch is dropped while pipe CRC is active, and the planes keep
their SEL_FETCH_PLANE_CTL / SEL_FETCH_CUR_CTL enable bit set in hardware
over that. A plane disabled while selective fetch is off never gets the
bit cleared, as the disable path is guarded by enable_psr2_sel_fetch.
Once selective fetch comes back the hardware resumes fetching for a
plane that is no longer enabled and keeps its DDB range reserved.
Clear the bits as selective fetch is turned off instead. Atomic check
has both the old and the new crtc state, so record the transition there
and let the plane and cursor arm paths write the registers to 0 for that
commit.
v2: Drop the old_crtc_state->hw.active check. [Jouni]
Fixes: b1f5279b5981 ("drm/i915/psr: Move plane sel fetch configuration into plane source files")
Closes: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8739
Assisted-by: Copilot:Claude-Opus-5
Signed-off-by: Nemesa Garg <nemesa.garg@intel.com>
Reviewed-by: Jouni Högander <jouni.hogander@intel.com>
Signed-off-by: Suraj Kandpal <suraj.kandpal@intel.com>
Link: https://patch.msgid.link/20260909110332.3528029-3-nemesa.garg@intel.com
(cherry picked from commit a4c0e7f80429eda6990960971aebd4e4b9533cc6)
Signed-off-by: Jani Nikula <jani.nikula@intel.com>
|
|
In xe_vm_close_and_put(), external-BO VMAs are queued on the contested
list for deferred destruction via xe_vma_destroy_unlocked(). However,
xe_vm_pt_destroy() was previously invoked before processing contested
VMAs, destroying vm->pt_root while those VMAs were still linked to their
respective buffer objects (vm_bo->list.gpuva).
If a concurrent thread evicts one of those shared buffer objects,
xe_bo_trigger_rebind() holding only bo->resv walks the BO's VMAs and, in
fault mode, calls xe_vm_invalidate_vma() -> xe_pt_zap_ptes(). Because
vm->pt_root[tile->id] is already NULL, dereferencing pt->level causes a
NULL ptr deref.
Fix this by deferring xe_vm_free_scratch() and xe_vm_pt_destroy() until
after all contested VMAs have been unlinked and destroyed.
User is reporting hitting a NULL ptr deref in xe_pt_zap_ptes(), which
could be explained by this race.
Assisted-by: LLM
Fixes: b06d47be7c83 ("drm/xe: Port Xe to GPUVA")
Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/9290
Signed-off-by: Matthew Auld <matthew.auld@intel.com>
Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: <stable@vger.kernel.org> # v6.12+
Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Reviewed-by: Matthew Brost <matthew.brost@intel.com>
Link: https://patch.msgid.link/20260918131034.598078-2-matthew.auld@intel.com
(cherry picked from commit c2863648959489767f08892fd6e90577d2ea0b6a)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
|
|
The desired behavior for SVM eviction failures, which can occur due to
various uncontrollable races, is for TTM to continue walking the LRU
list and look for another eviction candidate. This is expressed by
returning -ENOSPC from the ->move() callback.
Adjust the SVM eviction failure path because of races in ->move() to
return -ENOSPC so that TTM continues searching for another buffer to
evict.
Fixes: 3ca608dc7561 ("drm/xe: Basic SVM BO eviction")
Cc: stable@vger.kernel.org
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
Link: https://patch.msgid.link/20260917203158.292823-1-matthew.brost@intel.com
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
(cherry picked from commit 36a86c23588b8f57c9d20feb4cf5a2ab27e3baba)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
|
|
A TLB invalidation issued on a wedged device fails with
-ENOTRECOVERABLE. xe_tlb_inval_issue() squashes only -ECANCELED, so the
error reaches callers that treat it as unexpected and WARN, tainting the
kernel on a wedge that was deliberately caused:
ggtt_invalidate_gt_tlb() drivers/gpu/drm/xe/xe_ggtt.c
xe_svm_invalidate() drivers/gpu/drm/xe/xe_svm.c
xe_bo_trigger_rebind() drivers/gpu/drm/xe/xe_bo.c
xe_vma_userptr_do_inval() drivers/gpu/drm/xe/xe_userptr.c
igt@xe_exec_reset@gt-reset-fault-injection hits the GGTT one, turning an
otherwise passing run into an abort:
*ERROR* SIGID=102 FATAL (-EIO) WEDGED: Device declared wedged!
Tile0: GT1: Failed to invalidate GGTT (-ENOTRECOVERABLE)
WARNING: drivers/gpu/drm/xe/xe_ggtt.c:588 at ggtt_invalidate_gt_tlb
Workqueue: xe-guc-destroy-wq __guc_exec_queue_destroy_async [xe]
ggtt_node_remove+0xe3/0x100 [xe]
xe_ggtt_remove_bo+0x89/0x2c0 [xe]
xe_ttm_bo_destroy+0xcb/0x330 [xe]
...
xe_lrc_destroy+0x74/0x90 [xe]
xe_exec_queue_fini+0x2d/0x60 [xe]
-ECANCELED and -ENOTRECOVERABLE mean the same thing at this layer: the
message was dropped rather than delivered, and the fence has already
been signalled before the error is returned, so there is nothing left to
wait for. Squash both.
A wedged device is only recovered by a fresh initialisation, so the
error return in xe_bo_trigger_rebind() becomes unreachable by design.
v2: fix all invalidation paths. (Sashiko)
Fixes: 50fa9acac26f ("drm/xe/guc: distinguish wedged from recoverable cancellation")
Assisted-by: Claude:claude-opus-5
Cc: Sk Anirban <sk.anirban@intel.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Signed-off-by: Shuicheng Lin <shuicheng.lin@intel.com>
Reviewed-by: Matthew Brost <matthew.brost@intel.com>
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Link: https://patch.msgid.link/20260914215318.200603-1-shuicheng.lin@intel.com
(cherry picked from commit af14e3705345cb57c53b237169873233052a5c16)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
|
|
CRI defines bit 5 of the perf limit reasons register as a power brake
(PWRBRK) indicator. Add PWRBRK_MASK and a reason_pwrbrk sysfs attribute
for CRI in place of reason_ratl.
Signed-off-by: Sk Anirban <sk.anirban@intel.com>
Fixes: 8578e6d0546c ("drm/xe/gt_throttle: Drop individual show functions")
Reviewed-by: Raag Jadav <raag.jadav@intel.com>
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Link: https://patch.msgid.link/20260909114931.1039331-2-sk.anirban@intel.com
(cherry picked from commit e199c851c0ab608a0ca89e7be1756a1461b41a2c)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
|
|
When the Exynos DSI driver was generalized into samsung-dsim, the TE
GPIO acquisition was switched from gpiod_get_optional() to
devm_gpiod_get_optional() while keeping the matching gpiod_put() calls.
That combination is wrong for a managed descriptor.
However, dropping the puts and keeping the managed get is also wrong:
samsung_dsim_register_te_irq() runs from the DSI host attach callback,
and host detach/reattach can happen without destroying the device that
owns the managed action. A second attach would then request the GPIO
again without having released it.
Switch back to a non-managed gpiod_get_optional() and keep the explicit
gpiod_put() on the request_irq() error path and in
samsung_dsim_unregister_te_irq().
Fixes: e7447128ca4a ("drm: bridge: Generalize Exynos-DSI driver into a Samsung DSIM bridge")
Suggested-by: Luca Ceresoli <luca.ceresoli@bootlin.com>
Signed-off-by: Li Youhong <liyouhong@kylinos.cn>
Reviewed-by: Luca Ceresoli <luca.ceresoli@bootlin.com>
Tested-by: Luca Ceresoli <luca.ceresoli@bootlin.com>
Link: https://patch.msgid.link/20260904014958.1572918-1-dayou5941@163.com
[Luca: remove unnecessary comment]
Signed-off-by: Luca Ceresoli <luca.ceresoli@bootlin.com>
|
|
virtio_gpu_cmd_transfer_to_host_{2d,3d}() sync the shmem backing for the
device before the transfer, but nothing syncs for the CPU when a transfer
runs the other way. That breaks two ways. Where the DMA layer bounces, the
device writes into the bounce buffer while the guest keeps reading the
original pages. Where DMA is not coherent, the device writes memory while
the CPU keeps stale cache lines, because nothing reaches
arch_sync_dma_for_cpu(). Either way DRM_IOCTL_VIRTGPU_TRANSFER_FROM_HOST
hands back stale data.
Sashiko originally found this in
https://lore.kernel.org/dri-devel/20260806231002.27B4D1F000E9@smtp.kernel.org
but the suggestion there to fix this with dma_sync_sgtable_for_cpu()
isn't a sufficient fix, for two reasons.
- The transfer is asynchronous. virtio_gpu_cmd_transfer_from_host_3d() only
queues the command, so a sync there would run before the device had written
anything. It belongs on completion, and ahead of any fence signalling.
A waiter woken by the fence would otherwise race the sync and read the
backing pages regardless. It needs its own pass over the reclaim list
rather than a step inside the existing one, because
virtio_gpu_fence_event_process() also signals every earlier fence in the
same context, so any entry in that loop may signal an earlier entry's
fence.
- The transfer is also partial, carrying an offset, a level and a box.
Where the mapping bounces, a sync for the CPU copies the whole mapping
back, so unless the mapping is primed first the regions the device did not
write come back holding whatever the bounce buffer contained, discarding
data the guest owned.
So the fix: Prime the mapping before queueing, tag the vbuffer, and sync
for the CPU on completion before the fence is signalled.
A second transfer must not snapshot the mapping while an earlier one is
still in flight, or the snapshot would predate whatever the CPU wrote once
the earlier fence signalled and the later sync would discard it.
To mitigate this, wait for outstanding fences under the reservation before
priming.
Neither sync copies anything unless the mapping genuinely bounces:
swiotlb_sync_single_for_cpu() and its Xen counterpart look the address up
in the bounce pool and return early when it is absent. On a platform with
non-coherent DMA they still perform the necessary cache maintenance.
The range cannot be narrowed to the box, since for a non-blob resource
virtio_gpu_transfer_from_host_ioctl() rejects a caller-supplied stride and
layer_stride, leaving the layout to the host and the guest with no way to
work out which bytes the device writes. A host3d guest blob does carry
both, so its extent could be bounded, but the sync is left whole there too
rather than special-cased: priming makes the untouched regions round-trip
unchanged either way.
Behaviour changes worth noting:
- TRANSFER_FROM_HOST can now block, where before it returned as soon as the
command was queued. Repeated readbacks of one resource serialise, and a
readback can wait behind an earlier queued command that touched it, since
virtio_gpu_array_add_fence() tags uploads, execbufs and plane flushes
alike with DMA_RESV_USAGE_WRITE. -ERESTARTSYS was already possible here
via dma_resv_lock_interruptible().
- A CPU write racing an in-flight transfer to the same resource is now
lost, where before it survived and the transfer was lost instead. Priming
captures the pages as of queueing, so a write landing before completion is
overwritten by the sync.
- TRANSFER_TO_HOST can also block now, but only while a guest-bound
transfer on the same resource is outstanding, which happens only for
callers that issue both without waiting.
- Where a batch of completions contains a guest-bound transfer, the sync
pass delays fence signalling for the whole batch. Only bounced pages are
copied and the swiotlb pool bounds it. A batch with no such transfer is
unaffected.
Tested under QEMU on x86 with swiotlb=force and virtio-vga-gl
iommu_platform=on, which forces both preconditions required to hit the
original bug.
Fixes: a3b815f09bb8 ("drm/virtio: add iommu support.")
Reported-by: Sashiko AI review <sashiko-bot@kernel.org>
Closes: https://lore.kernel.org/dri-devel/20260806231002.27B4D1F000E9@smtp.kernel.org/
Signed-off-by: Benjamin Leggett <benjamin@edera.io>
Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com>
Link: https://patch.msgid.link/20260814-virtgpu-from-host-sync-v4-1-64dd736b1779@edera.io
|
|
The cursor plane exposes a format with an alpha channel
(DRM_FORMAT_ARGB8888) without a pixel blend mode property. Since
commit 860e748bddcc ("drm: ensure blend mode supported if pixel
format with alpha exposed") this triggers a warning during
drm_mode_config_validate():
[ 0.649020] ------------[ cut here ]------------
[ 0.649040] [PLANE:36:plane-1] pixel format with alpha exposed but blend mode not setup
[ 0.649081] WARNING: drivers/gpu/drm/drm_mode_config.c:872 at drm_mode_config_validate
......
[ 0.649761] Call trace:
[ 0.649764] drm_mode_config_validate+0x398/0x558 [drm] (P)
[ 0.649912] drm_dev_register+0x1cc/0x2a0 [drm]
[ 0.650058] virtio_gpu_probe+0xd4/0x1c0 [virtio_gpu]
[ 0.650088] virtio_dev_probe+0x1c8/0x310
......
[ 0.650261] ---[ end trace 0000000000000000 ]---
Create the property with the only supported blend mode,
DRM_MODE_BLEND_PREMULTI, which is also the property's default and
matches what userspace had to assume before the property existed.
Fixes: 860e748bddcc ("drm: ensure blend mode supported if pixel format with alpha exposed")
Reported-by: Ye Liu <liuye@kylinos.cn>
Signed-off-by: Shixiong Ou <oushixiong@kylinos.cn>
Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com>
Link: https://patch.msgid.link/20260828090123.784944-1-oushixiong1025@163.com
|
|
virtio_gpu_fence_alloc() can fail due to memory pressure and return NULL,
but its caller like virtio_gpu_init_submit() never checks it. Later,
virtio_gpu_init_submit() passes the NULL fence to
virtio_gpu_fence_event_create(), which unconditionally dereferences it.
Found when fuzzing the virtio driver with Syzkaller:
Oops: general protection fault, probably for non-canonical address 0xdffffc0000000012: 0000 [#1] SMP KASAN NOPTI
KASAN: null-ptr-deref in range [0x0000000000000090-0x0000000000000097]
CPU: 1 UID: 0 PID: 9991 Comm: syz.0.121 Not tainted 7.2.0 #4 PREEMPT(full)
Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS rel-1.17.0-0-gb52ca86e094d-prebuilt.qemu.org 04/01/2014
RIP: 0010:virtio_gpu_fence_event_create drivers/gpu/drm/virtio/virtgpu_submit.c:295 [inline]
RIP: 0010:virtio_gpu_init_submit drivers/gpu/drm/virtio/virtgpu_submit.c:398 [inline]
RIP: 0010:virtio_gpu_execbuffer_ioctl+0xc78/0x1aa0 drivers/gpu/drm/virtio/virtgpu_submit.c:505
Code: 85 ed 0f 85 21 09 00 00 e8 05 5a c9 fb 48 8b 44 24 10 48 8d b8 90 00 00 00 48 b8 00 00 00 00 00 fc ff df 48 89 fa 48 c1 ea 03 <80> 3c 02 00 0f 85 9a 0d 00 00 48 8b 44 24 10 4c 89 b0 90 00 00 00
RSP: 0018:ffffc900039dfad0 EFLAGS: 00010216
RAX: dffffc0000000000 RBX: ffffc900039dfdd8 RCX: ffffffff85f6fd3d
RDX: 0000000000000012 RSI: ffffffff85f6fd4b RDI: 0000000000000090
RBP: 0000000000000000 R0virtio_gpu_virgl_process_cmd: ctrl 0x102, error 0x1203
R10: 0000000000000000 R11: 0000000000000000 R12: ffff8880132c4000
R13: 0000000000000000 R14: ffff888073b6c700 R15: 000000000000003b
FS: 00007fab480b96c0(0000) GS:ffff8880eb6e9000(0000) knlGS:0000000000000000
CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 00007effbf5e55a8 CR3: 0000000048d19000 CR4: 0000000000350ef0
Call Trace:
<TASK>
drm_ioctl_kernel+0x1f4/0x3e0 drivers/gpu/drm/drm_ioctl.c:817
drm_ioctl+0x5f4/0xc70 drivers/gpu/drm/drm_ioctl.c:914
vfs_ioctl fs/ioctl.c:51 [inline]
__do_sys_ioctl fs/ioctl.c:597 [inline]
__se_sys_ioctl fs/ioctl.c:583 [inline]
__x64_sys_ioctl+0x18e/0x210 fs/ioctl.c:583
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x116/0x800 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7fab471a82bd
Code: ff c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 c7 c1 b0 ff ff ff f7 d8 64 89 01 48
RSP: 002b:00007fab480b9018 EFLAGS: 00000246 ORIG_RAX: 0000000000000010
RAX: ffffffffffffffda RBX: 00007fab47435fa0 RCX: 00007fab471a82bd
RDX: 00002000000000c0 RSI: 00000000c0406442 RDI: 0000000000000003
RBP: 00007fab480b9080 R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000001
R13: 00007fab47436038 R14: 00007fab47435fa0 R15: 00007ffe85ab0740
</TASK>
Modules linked in:
---[ end trace 0000000000000000 ]---
RIP: 0010:virtio_gpu_fence_event_create drivers/gpu/drm/virtio/virtgpu_submit.c:295 [inline]
RIP: 0010:virtio_gpu_init_submit drivers/gpu/drm/virtio/virtgpu_submit.c:398 [inline]
RIP: 0010:virtio_gpu_execbuffer_ioctl+0xc78/0x1aa0 drivers/gpu/drm/virtio/virtgpu_submit.c:505
Code: 85 ed 0f 85 21 09 00 00 e8 05 5a c9 fb 48 8b 44 24 10 48 8d b8 90 00 00 00 48 b8 00 00 00 00 00 fc ff df 48 89 fa 48 c1 ea 03 <80> 3c 02 00 0f 85 9a 0d 00 00 48 8b 44 24 10 4c 89 b0 90 00 00 00
RSP: 0018:ffffc900039dfad0 EFLAGS: 00010216
RAX: dffffc0000000000 RBX: ffffc900039dfdd8 RCX: ffffffff85f6fd3d
RDX: 0000000000000012 RSI: ffffffff85f6fd4b RDI: 0000000000000090
RBP: 0000000000000000 R08: 0000000000000005 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000000 R12: ffff8880132c4000
R13: 0000000000000000 R14: ffff888073b6c700 R15: 000000000000003b
FS: 00007fab480b96c0(0000) GS:ffff888098ae9000(0000) knlGS:0000000000000000
CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 00007f24c3759000 CR3: 0000000048d19000 CR4: 0000000000350ef0
----------------
Code disassembly (best guess):
0: 85 ed test %ebp,%ebp
2: 0f 85 21 09 00 00 jne 0x929
8: e8 05 5a c9 fb call 0xfbc95a12
d: 48 8b 44 24 10 mov 0x10(%rsp),%rax
12: 48 8d b8 90 00 00 00 lea 0x90(%rax),%rdi
19: 48 b8 00 00 00 00 00 movabs $0xdffffc0000000000,%rax
20: fc ff df
23: 48 89 fa mov %rdi,%rdx
26: 48 c1 ea 03 shr $0x3,%rdx
* 2a: 80 3c 02 00 cmpb $0x0,(%rdx,%rax,1) <-- trapping instruction
2e: 0f 85 9a 0d 00 00 jne 0xdce
34: 48 8b 44 24 10 mov 0x10(%rsp),%rax
39: 4c 89 b0 90 00 00 00 mov %r14,0x90(%rax)
Fix by checking virtio_gpu_fence_alloc() in virtio_gpu_init_submit() and
returning -ENOMEM before any later code can dereference the NULL fence.
Fixes: 70d1ace56db6 ("drm/virtio: Conditionally allocate virtio_gpu_fence")
Cc: stable@vger.kernel.org
Signed-off-by: Peiyang He <peiyang_he@smail.nju.edu.cn>
Assisted-by: Codex:gpt-5.5
Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com>
Link: https://patch.msgid.link/00EFE4BA92889B14+20260909091114.2622550-1-peiyang_he@smail.nju.edu.cn
|
|
Guest userspace may import udmabuf to vrend. Vrend doesn't support guest
blobs, and thus, further 3d operations with the imported blob are failing.
Typical scenario of the problem shown with mouse cursor RGBA image imported
into virtio-gpu, which previously was rejected by virtio-gpu driver.
Revert enabling guest blobs importing into vrend to fix the regression.
Link: https://gitlab.freedesktop.org/virgl/virglrenderer/-/work_items/674
Fixes: df4dc947c46b ("drm/virtio: Allow importing prime buffers when 3D is enabled")
Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com>
Reviewed-by: Val Packett <val@invisiblethingslab.com>
Link: https://patch.msgid.link/20260911144204.2089401-1-dmitry.osipenko@collabora.com
|
|
virtio_gpu_vram_create() frees the object with a bare kfree(vram) on
both error paths after drm_gem_private_object_init() has run, and on the
second one after drm_gem_create_mmap_offset() has linked obj->vma_node
into the device's VMA offset manager. The freed object stays in that
interval tree, so a later lookup or insertion walks freed memory, and
the dma_resv and gpuva lock are never destroyed.
Call drm_gem_object_release() before kfree() on both paths.
Fixes: 16845c5d5409 ("drm/virtio: implement blob resources: implement vram object")
Assisted-by: Claude:claude-opus-5
Signed-off-by: Junrui Luo <moonafterrain@outlook.com>
Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com>
Link: https://patch.msgid.link/20260915-fixes-v2-4-a0d799e4db66@outlook.com
|
|
virtio_gpu_resource_create_blob_ioctl() calls drm_gem_object_release() on
both the virtio_gpu_resource_assign_uuid() and drm_gem_handle_create()
error paths instead of dropping the reference it owns, so
obj->funcs->free() never runs and the virtio_gpu_object, the resource id
and the host-side resource are leaked.
Use drm_gem_object_put() instead.
Fixes: 897b4d1acaf5 ("drm/virtio: implement blob resources: resource create blob ioctl")
Assisted-by: Claude:claude-opus-5
Signed-off-by: Junrui Luo <moonafterrain@outlook.com>
Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com>
Link: https://patch.msgid.link/20260915-fixes-v2-3-a0d799e4db66@outlook.com
|
|
virtio_gpu_resource_create_ioctl() calls drm_gem_object_release() on the
drm_gem_handle_create() error path instead of dropping the reference it
owns, so obj->funcs->free() never runs and the virtio_gpu_object, its
pages and sg table, the resource id and the host-side resource are
leaked.
Use drm_gem_object_put() instead.
Fixes: 62fb7a5e1096 ("virtio-gpu: add 3d/virgl support")
Assisted-by: Claude:claude-opus-5
Signed-off-by: Junrui Luo <moonafterrain@outlook.com>
Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com>
Link: https://patch.msgid.link/20260915-fixes-v2-2-a0d799e4db66@outlook.com
|
|
virtio_gpu_gem_create() owns the reference taken by
virtio_gpu_object_create(). On the drm_gem_handle_create() error path it
calls drm_gem_object_release() instead of dropping that reference.
drm_gem_object_release() is the inverse of drm_gem_object_init() and does
not touch the reference count or call obj->funcs->free(), so it is only
correct as the last step of a destructor, as in
virtio_gpu_cleanup_object(). Using it here leaves the bo at refcount 1
with no remaining reference, so virtio_gpu_free_object() never runs and
the shmem pages, sg table and virtio_gpu_object are leaked. Since
virtio_gpu_object_create() has already set bo->created,
VIRTIO_GPU_CMD_RESOURCE_UNREF is not queued either, leaking the host-side
resource and the resource id.
drm_gem_handle_create_tail() drops the handle reference on all of its
internal error paths, so the caller only has to drop its own. Use
drm_gem_object_put(), matching the success path below.
Fixes: dc5698e80cf7 ("Add virtio gpu driver.")
Reported-by: Yuhao Jiang <danisjiang@gmail.com>
Assisted-by: Claude:claude-opus-5
Signed-off-by: Junrui Luo <moonafterrain@outlook.com>
Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com>
Link: https://patch.msgid.link/20260915-fixes-v2-1-a0d799e4db66@outlook.com
|
|
virtio_gpu_execbuffer_ioctl() reserves a DRM event with
drm_event_reserve_init() when VIRTGPU_EXECBUF_RING_IDX selects a ring that
userspace has enabled polling for. virtio_gpu_init_submit() does this
before the BO handles, the command buffer, the syncobj arrays and the
in-fence are processed, so every later error path runs with the event
already pending, including plain argument validation failures such as an
invalid bo_handle or an in-syncobj that carries no fence.
On those paths, virtio_gpu_cleanup_submit() drops the out-fence without
cancelling the event. The fence is freed without ever having been emitted,
taking the only driver-side pointer to the event with it. Closing the DRM
file does not help. drm_events_release() unlinks pending events but
deliberately leaves the freeing to the driver's later drm_send_event(),
which never runs for an orphaned event, so the allocation is leaked
permanently.
Found when fuzzing the virtio driver with Syzkaller:
BUG: memory leak
unreferenced object 0xffff88802c176e80 (size 96):
comm "syz.1.367", pid 10561, jiffies 4294960122
hex dump (first 32 bytes):
00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 ................
c8 6e 17 2c 80 88 ff ff 00 00 00 00 00 00 00 00 .n.,............
backtrace (crc e1973c6b):
kmemleak_alloc_recursive include/linux/kmemleak.h:44 [inline]
slab_post_alloc_hook mm/slub.c:4597 [inline]
slab_alloc_node mm/slub.c:4917 [inline]
__kmalloc_cache_noprof+0x49d/0x6f0 mm/slub.c:5485
_kmalloc_noprof include/linux/slab.h:988 [inline]
_kzalloc_noprof include/linux/slab.h:1309 [inline]
virtio_gpu_fence_event_create drivers/gpu/drm/virtio/virtgpu_submit.c:282 [inline]
virtio_gpu_init_submit drivers/gpu/drm/virtio/virtgpu_submit.c:398 [inline]
virtio_gpu_execbuffer_ioctl+0xbbf/0x1aa0 drivers/gpu/drm/virtio/virtgpu_submit.c:505
drm_ioctl_kernel+0x1f4/0x3e0 drivers/gpu/drm/drm_ioctl.c:817
drm_ioctl+0x5f4/0xc70 drivers/gpu/drm/drm_ioctl.c:914
vfs_ioctl fs/ioctl.c:51 [inline]
__do_sys_ioctl fs/ioctl.c:597 [inline]
__se_sys_ioctl fs/ioctl.c:583 [inline]
__x64_sys_ioctl+0x18e/0x210 fs/ioctl.c:583
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x116/0x800 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
Fix by cancelling and freeing the DRM event on the execbuffer error path
before dropping the fence. Clear the fence's event pointer after
cancellation so it does not retain a dangling pointer.
Fixes: cd7f5ca33585 ("drm/virtio: implement context init: add virtio_gpu_fence_event")
Cc: stable@vger.kernel.org
Signed-off-by: Peiyang He <peiyang_he@smail.nju.edu.cn>
Assisted-by: Codex:gpt-5.6-luna
Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com>
Link: https://patch.msgid.link/D320EAB5680C1411+20260908121323.2405044-1-peiyang_he@smail.nju.edu.cn
|
|
https://gitlab.freedesktop.org/drm/misc/kernel into drm-fixes
Two ttm fixes for ttm_tt_swapout(), one page-alignment and one overflow
fix for dma-buf, a drm_pending_vblank_event leak fix for drm,
suspend/resume fixes for nouveau, one out-of-bounds access fix for gud,
a use-after-free fix for vc4, a fence signaling fix, a race condition
fix for sched, planes formats fixes for verisilicon, and add the blend
mode property for loongson
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Maxime Ripard <self@mripard.dev>
Link: https://patch.msgid.link/aqvVENQ4ksJEIcdb@houat
|
|
nv50_hdmi_enable() passes the sink's SCDC capability from its EDID
straight through to nvif_outp_hdmi(). On pre-Maxwell-2 cards there is no
hdmi->scdc callback, so nvkm_uoutp_mthd_hdmi() rejects the whole
configuration with -EINVAL, and nv50_hdmi_enable() returns before
hdmi->ctrl() runs and before the AVI and VSI infoframes are sent.
The result on such a card driving an SCDC-capable HDMI 2.0 sink is that
HDMI audio silently stops working. Video is unaffected, and nothing is
logged, which makes the failure hard to attribute.
SCDC is optional, and the hdmi->scdc() call further down is already
guarded against a missing callback. Requesting it on a card that cannot
do it need not invalidate the rest of the HDMI configuration, so drop
that term from the condition and let the existing guard skip SCDC alone.
Fixes: 6c6abab20b99 ("drm/nouveau/disp: add output hdmi config method")
Signed-off-by: Giuseppe Ranieri <giuseppe@ranieri.dev>
Co-Authored-By: Tano Dzhinski <tano.dzhinski@gmail.com>
Signed-off-by: Tano Dzhinski <tano.dzhinski@gmail.com>
Tested-by: Tano Dzhinski <tano.dzhinski@gmail.com>
Reviewed-by: Lyude Paul <lyude@redhat.com>
Signed-off-by: Lyude Paul <lyude@redhat.com>
Link: https://patch.msgid.link/20260917215114.1136715-1-tano.dzhinski@gmail.com
|
|
nouveau_bo_pin_locked() checks whether an already pinned BO is in a
memory domain compatible with a new pin request. When the domains are
incompatible, it sets -EBUSY but still calls ttm_bo_pin() before
returning.
Callers treat a failed nouveau_bo_pin() as not having acquired a new pin,
so the extra pin count is never decreased by a matching unpin.
This triggers the warning in ttm_bo_release():
WARN_ON_ONCE(bo->pin_count);
Found when fuzzing the nouveau driver with a modified Syzkaller:
WARNING: drivers/gpu/drm/ttm/ttm_bo.c:256 at ttm_bo_release+0x827/0x9e0 drivers/gpu/drm/ttm/ttm_bo.c:256, CPU#1: syz.3.24/2212
Modules linked in:
CPU: 1 UID: 0 PID: 2212 Comm: syz.3.24 Not tainted 7.2.0 #24 PREEMPT(lazy)
nouveau 0000:01:00.0: gsp:msg fn:103 len:0x40/0x20 res:0x19 resp:0x19
Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
RIP: 0010:ttm_bo_release+0x827/0x9e0 drivers/gpu/drm/ttm/ttm_bo.c:256
Code: 02 00 0f 85 51 01 00 00 48 8b 7b 08 e8 d2 20 01 00 e9 80 fd ff ff e8 d8 15 c0 fe 90 0f 0b 90 e9 e1 f8 ff ff e8 ca 15 c0 fe 90 <0f> 0b 90 e9 a4 f8 ff ff e8 bc 15 c0 fe be 03 00 00 00 4c 89 e7 e8
msg: 00000000: 05 00 d0 c1 04 00 f0 f1 01 30 00 00 2d 90 00 00 .........0..-...
RSP: 0018:ffffc9000f5cf710 EFLAGS: 00010293
RAX: 0000000000000000 RBX: ffff888018e5d2a8 RCX: ffffffff82bb1b36
RDX: ffff888017b68000 RSI: 0000000000000004 RDI: ffff888018e5d2a8
msg: 00000010: 19 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 ................
RBP: ffff88801261c720 R08: 0000000000000001 R09: ffffed10031cba55
R10: ffff888018e5d2ab R11: 00000000000000f3 R12: ffff888018e5d290
R13: ffff888018e5d2d4 R14: ffff88801b219c18 R15: dffffc0000000000
FS: 0000000000000000(0000) GS:ffff8880e0f6f000(0000) knlGS:0000000000000000
CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 0000001b31223ffc CR3: 0000000028e00005 CR4: 0000000000770ef0
PKRU: 80000000
Call Trace:
<TASK>
kref_put include/linux/kref.h:65 [inline]
ttm_bo_put drivers/gpu/drm/ttm/ttm_bo.c:325 [inline]
ttm_bo_fini+0x55/0x80 drivers/gpu/drm/ttm/ttm_bo.c:330
nouveau_gem_object_del+0xb2/0x1b0 drivers/gpu/drm/nouveau/nouveau_gem.c:90
drm_gem_object_free+0x5f/0x90 drivers/gpu/drm/drm_gem.c:1165
kref_put include/linux/kref.h:65 [inline]
__drm_gem_object_put include/drm/drm_gem.h:562 [inline]
drm_gem_object_put include/drm/drm_gem.h:575 [inline]
nouveau_abi16_chan_fini.constprop.0+0x44f/0x5a0 drivers/gpu/drm/nouveau/nouveau_abi16.c:195
nouveau 0000:01:00.0: syz.2.23[2209]: Unknown handle 0x00000000
nouveau_abi16_fini+0x1d0/0x340 drivers/gpu/drm/nouveau/nouveau_abi16.c:225
nouveau_drm_postclose+0x18b/0x3e0 drivers/gpu/drm/nouveau/nouveau_drm.c:1284
nouveau 0000:01:00.0: syz.2.23[2209]: validate_init
drm_file_free.part.0+0x6d6/0xb60 drivers/gpu/drm/drm_file.c:267
drm_file_free drivers/gpu/drm/drm_file.c:237 [inline]
drm_close_helper.isra.0+0x11a/0x160 drivers/gpu/drm/drm_file.c:290
drm_release+0x1ab/0x330 drivers/gpu/drm/drm_file.c:438
__fput+0x39c/0xa60 fs/file_table.c:512
nouveau 0000:01:00.0: syz.2.23[2209]: validate: -2
task_work_run+0x15a/0x230 kernel/task_work.c:233
exit_task_work include/linux/task_work.h:40 [inline]
do_exit+0x82b/0x25a0 kernel/exit.c:1009
do_group_exit+0xc2/0x280 kernel/exit.c:1152
get_signal+0x1d6e/0x1f30 kernel/signal.c:3046
arch_do_signal_or_restart+0x7d/0x6e0 arch/x86/kernel/signal.c:337
__exit_to_user_mode_loop kernel/entry/common.c:66 [inline]
exit_to_user_mode_loop+0xdf/0x440 kernel/entry/common.c:101
__exit_to_user_mode_prepare include/linux/irq-entry-common.h:207 [inline]
syscall_exit_to_user_mode_prepare include/linux/irq-entry-common.h:230 [inline]
syscall_exit_to_user_mode include/linux/entry-common.h:318 [inline]
do_syscall_64+0x4f8/0x690 arch/x86/entry/syscall_64.c:100
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7f12bac8594d
Code: Unable to access opcode bytes at 0x7f12bac85923.
RSP: 002b:00007f12b96e70d8 EFLAGS: 00000246 ORIG_RAX: 00000000000000ca
RAX: 0000000000000001 RBX: 00007f12baf15fa8 RCX: 00007f12bac8594d
RDX: 00000000000f4240 RSI: 0000000000000081 RDI: 00007f12baf15fac
RBP: 00007f12baf15fa0 R08: 00007f12baee8000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000
R13: 00007f12baf16038 R14: 0000000000000006 R15: 00007ffe2ed394b0
</TASK>
irq event stamp: 47867
hardirqs last enabled at (47883): [<ffffffff815cafc6>] __up_console_sem+0x66/0x70 kernel/printk/printk.c:347
hardirqs last disabled at (47892): [<ffffffff815cafab>] __up_console_sem+0x4b/0x70 kernel/printk/printk.c:345
softirqs last enabled at (47880): [<ffffffff81434277>] __do_softirq kernel/softirq.c:656 [inline]
softirqs last enabled at (47880): [<ffffffff81434277>] invoke_softirq kernel/softirq.c:496 [inline]
softirqs last enabled at (47880): [<ffffffff81434277>] __irq_exit_rcu+0x137/0x1c0 kernel/softirq.c:735
softirqs last disabled at (47875): [<ffffffff81434277>] __do_softirq kernel/softirq.c:656 [inline]
softirqs last disabled at (47875): [<ffffffff81434277>] invoke_softirq kernel/softirq.c:496 [inline]
softirqs last disabled at (47875): [<ffffffff81434277>] __irq_exit_rcu+0x137/0x1c0 kernel/softirq.c:735
Fix by calling ttm_bo_pin() only when the existing placement is compatible
with the new pin request. This matches the correct behavior in other DRM
drivers such as amdgpu_bo_pin() in amdgpu.
Cc: stable@vger.kernel.org
Fixes: ad76b3f7c7a0 ("drm/nouveau: teach nouveau_bo_pin() how to force a contig vram allocation")
Signed-off-by: Peiyang He <peiyang_he@smail.nju.edu.cn>
Assisted-by: LLM
Reviewed-by: Lyude Paul <lyude@redhat.com>
Signed-off-by: Lyude Paul <lyude@redhat.com>
Link: https://patch.msgid.link/EACEF2F4E098413F+20260918025312.2814889-1-peiyang_he@smail.nju.edu.cn
|
|
nvkm_cstate_prog() reuses 'ret' for the voltage and fan-speed restore
calls it makes after reprogramming the clocks. Those calls almost always
succeed, so the status of the reclock itself is overwritten and the
function reports success even when clk->func->calc() or clk->func->prog()
failed. The converse is also true: a successful reclock is reported as an
error if the final restore call fails, even though that failure is only
logged and otherwise ignored.
The only consumer of the return value is the error message in
nvkm_pstate_work(), so in practice a failing reclock is simply never
reported. Nothing else changes, but a function that returns success on
failure is a trap for the next caller.
Keep the calc/prog status in 'ret' and use a separate local for the
restore calls.
Fixes: 3eca809b3c05 ("drm/nouveau/clk: cosmetic changes")
Signed-off-by: Francesco Magazzu <postadelmaga@gmail.com>
Reviewed-by: Lyude Paul <lyude@redhat.com>
Signed-off-by: Lyude Paul <lyude@redhat.com>
Link: https://patch.msgid.link/20260918131620.405133-5-postadelmaga@gmail.com
|
|
nvkm_control_mthd_pstate_attr() looks up the pstate at the index supplied
by userspace by walking clk->states, and then keeps using the
list_for_each_entry cursor after the loop. This is not triggerable today:
the function already rejects args->v0.state >= clk->state_nr before the
loop, and clk->state_nr is kept in sync with the number of entries on
clk->states, so the lookup always breaks on a real entry.
Should the loop ever run to completion, the cursor would point at the list
head rather than at a pstate, and the pstate->base.domain[] read and the
walk of pstate->list that follow would read past it. Rather than leave
that trap in place, track whether the entry was found and return -EINVAL if
it was not, like the other lookup failures in this function.
No functional change.
Signed-off-by: Francesco Magazzu <postadelmaga@gmail.com>
Reviewed-by: Lyude Paul <lyude@redhat.com>
Signed-off-by: Lyude Paul <lyude@redhat.com>
Link: https://patch.msgid.link/20260918131620.405133-4-postadelmaga@gmail.com
|
|
nvkm_pstate_prog() walks clk->states looking for the entry at index
'pstatei' and then keeps using the list_for_each_entry cursor after the
loop. This is not triggerable today: every caller clamps the index
against clk->state_nr before calling, so the loop always breaks on a real
entry. It is safe by virtue of what the callers happen to do, not by
anything the function itself checks.
Should a caller ever pass an index that is not on the list, the cursor
would point at the list head rather than at a pstate, and the
pstate->base.domain[] and pstate->fanspeed accesses that follow would read
past it. Rather than leave that trap in place for the next caller, track
whether the entry was found and return -EINVAL if it was not.
No functional change.
Signed-off-by: Francesco Magazzu <postadelmaga@gmail.com>
Reviewed-by: Lyude Paul <lyude@redhat.com>
Signed-off-by: Lyude Paul <lyude@redhat.com>
Link: https://patch.msgid.link/20260918131620.405133-3-postadelmaga@gmail.com
|
|
If list_for_each_entry() exits without hitting a break then "pstate" is
not a valid pstate pointer. Introduce a "found" variable instead.
The check is reachable from userspace: nvkm_clk_ustate_update() takes the
pstate id straight from the 'pstate' debugfs file, so requesting an id
that is not in clk->states - or any id at all when the perf tables are
broken and the list is empty - makes the pstate->pstate != req test
dereference the list head cast to a struct nvkm_pstate, which is an
out-of-bounds read.
Fixes: 7c8565220697 ("drm/nouveau/clk: implement power state and engine clock control in core")
Signed-off-by: Dan Carpenter <dan.carpenter@oracle.com>
[Francesco: rebased on drm-misc-next, expanded the commit message]
Signed-off-by: Francesco Magazzu <postadelmaga@gmail.com>
Reviewed-by: Lyude Paul <lyude@redhat.com>
Signed-off-by: Lyude Paul <lyude@redhat.com>
Link: https://patch.msgid.link/20260918131620.405133-2-postadelmaga@gmail.com
|
|
https://gitlab.freedesktop.org/drm/amdgpu/kernel into drm-fixes
amd-drm-fixes-7.3-2026-09-17:
amdgpu:
- SMU 14.x fix
- DC IRQ fix
- Runtime PM fix for P2P
- RAS fix
- PCIe reporting fix
- DCN 6 fix
- Device removal fix
- DC MALL fix
amdkfd:
- GC 12.x fixes
- Boundary checks
- Mapping clear fix
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Alex Deucher <alexander.deucher@amd.com>
Link: https://patch.msgid.link/20260917201213.3880863-1-alexander.deucher@amd.com
|