summaryrefslogtreecommitdiff
path: root/drivers/gpu/drm
AgeCommit message (Collapse)Author
5 daysMerge tag 'drm-misc-fixes-2026-09-24' of ↵Dave Airlie
https://gitlab.freedesktop.org/drm/misc/kernel into drm-fixes A number of fixes: - bridge: - samsung-dsim: fix GPIO lifetime - client: Null pointer dereference fix - imagination: error handling fix, page handling fix - nouveau: fix reference leaks, double-frees, out-of-bounds accesses, use-after-frees, don't reject config without SCDC, a number of workarounds - virtio: fix memory leak, reference leaks, null pointer dereference, add pixel blend mode, cache coherency fix Signed-off-by: Dave Airlie <airlied@redhat.com> From: Maxime Ripard <self@mripard.dev> Link: https://patch.msgid.link/arU22zzqUGDEco1y@houat
6 daysMerge tag 'amd-drm-fixes-7.3-2026-09-24' of ↵Dave Airlie
https://gitlab.freedesktop.org/drm/amdgpu/kernel into drm-fixes amd-drm-fixes-7.3-2026-09-24: amdgpu: - Display ref count fix - Userq fixes - VCN 4, 5 reset fixes - Fixes for various error paths - Stack frame size fixes for various combinations of compilers and configs amdkfd: - Possible UAF fix Signed-off-by: Dave Airlie <airlied@redhat.com> From: Alex Deucher <alexander.deucher@amd.com> Link: https://patch.msgid.link/20260924172938.634777-1-alexander.deucher@amd.com
6 daysMerge tag 'drm-xe-fixes-2026-09-24' of ↵Dave Airlie
https://gitlab.freedesktop.org/drm/xe/kernel into drm-fixes Fixes in: - CRI throttle reasons report (Sk) - TLB invalidation at wedge (Shuicheng) - SVM eviction and VM close (Brost) - Display corruption on LNL on Xen PV (Szymon) - W/a fix and addition (Tilak) Signed-off-by: Dave Airlie <airlied@redhat.com> From: Rodrigo Vivi <rodrigo.vivi@intel.com> Link: https://patch.msgid.link/arUoUf9LpsJpJouN@intel.com
6 daysdrm/amd/display: Bump frame warning limit for all builds of dmlAlex Deucher
Some configs with gcc are also now affected. Signed-off-by: Alex Deucher <alexander.deucher@amd.com> (cherry picked from commit 76b9706e7fa1a6e374703170128b1be2f590cda7)
7 daysdrm/imagination: clamp freelist reconstruction requestsPengpeng Hou
The firmware reconstruction count controls accesses to the request's fixed freelist ID array and the copy into the fixed response array. Neither access currently bounds the count to those protocol arrays. Clamp the count to the request capacity, which is shared by the response layout, and use that count consistently for reconstruction and response publication. Keep the firmware recovery exchange instead of dropping an oversized request without a response, as discussed with the firmware maintainer. The issue was found by our static-analysis tool. Fixes: 6eedddab733b ("drm/imagination: Implement free list and HWRT create and destroy ioctls") Assisted-by: gpt 5 Signed-off-by: Pengpeng Hou <hppiscas@163.com> Reviewed-by: Alessio Belle <alessio.belle@imgtec.com> Link: https://patch.msgid.link/20260920034329.16614-1-hppiscas@163.com Signed-off-by: Brajesh Gupta <brajesh.gupta@imgtec.com>
7 daysdrm/imagination: Fix page count for page table for map() interfaceBrajesh Gupta
The GPU virtual start address wasn't included in the calculation for the amount of page tables required for mapping a BO object in map() interface. It resulted in map failure later due to not enough pages at L0/L1 level. Update pvr_mmu_op_context_create() interface to pass device address as well to allow correct calculation for page table memory. If L0 tables cover 2MB (0x200000), the range defined by device address 0x80001ff000 (general heap at 2MB - 4KB) and size 0x2000 (two 4KB pages) requires two L0 pages to be mapped, but without the base address a range of 0x2000 computes to a single L0 page which is not enough. Fixes: ff5f643de0bf ("drm/imagination: Add GEM and VM related code") Reviewed-by: Alexandru Dadu <alexandru.dadu@imgtec.com> Reviewed-by: Alessio Belle <alessio.belle@imgtec.com> Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260922-mmu_fix-v4-2-12f1a871456a@imgtec.com Signed-off-by: Brajesh Gupta <brajesh.gupta@imgtec.com>
7 daysdrm/imagination: Propagate map failures correctly from pvr_mmu_map_sgl()Brajesh Gupta
Map failure from pvr_mmu_map_sgl() interface was not returned correctly to pvr_mmu_map() interface. This resulted in pvr_mmu_map() interface to continue instead of returning an error to caller. Fix it by returning a proper error code from pvr_mmu_map_sgl() interface. Call stack for crash: [ 1179.286237] Unable to handle kernel NULL pointer dereference at virtual address 0000000000000008 [ 1179.295067] Mem abort info: [ 1179.297877] ESR = 0x0000000096000004 [ 1179.301656] EC = 0x25: DABT (current EL), IL = 32 bits [ 1179.306987] SET = 0, FnV = 0 [ 1179.310048] EA = 0, S1PTW = 0 [ 1179.313198] FSC = 0x04: level 0 translation fault [ 1179.318096] Data abort info: [ 1179.320993] ISV = 0, ISS = 0x00000004, ISS2 = 0x00000000 [ 1179.326483] CM = 0, WnR = 0, TnD = 0, TagAccess = 0 [ 1179.331546] GCS = 0, Overlay = 0, DirtyBit = 0, Xs = 0 [ 1179.336895] user pgtable: 4k pages, 48-bit VAs, pgdp=000000009822a000 [ 1179.343402] [0000000000000008] pgd=0000000000000000, p4d=0000000000000000 [ 1179.350243] Internal error: Oops: 0000000096000004 [#2] SMP [ 1179.355908] Modules linked in: powervr gpu_sched drm_shmem_helper drm_gpuvm drm_exec xhci_plat_hcd xhci_hcd dwc3 usbcore usb_common snd_soc_simple_card snd_soc_simple_card_utils dwc3_am62 at24 sa2ul sha512 libsha512 sha256 authenc sch_fq_codel fuse dm_mod ipv6 [ 1179.378992] CPU: 1 UID: 1000 PID: 680 Comm: deqp-vk Tainted: G D 6.17.0 #1 PREEMPT [ 1179.388120] Tainted: [D]=DIE [ 1179.390994] Hardware name: Texas Instruments AM625 SK (DT) [ 1179.396467] pstate: 00000005 (nzcv daif -PAN -UAO -TCO -DIT -SSBS BTYPE=--) [ 1179.403415] pc : pvr_mmu_op_context_unmap_curr_page+0x6c/0x134 [powervr] [ 1179.410140] lr : pvr_mmu_op_context_unmap_curr_page+0x58/0x134 [powervr] [ 1179.416848] sp : ffff8000839ab8c0 [ 1179.420153] x29: ffff8000839ab8c0 x28: 0000000000000001 x27: 000000008f386000 [ 1179.427283] x26: ffff000016d1df98 x25: 0000000000247000 x24: 00000000000001e6 [ 1179.434413] x23: 0000000000000002 x22: 000000000000ffff x21: 0000000000000247 [ 1179.441540] x20: 0000000000000245 x19: ffff000016d1df60 x18: 0000000000000002 [ 1179.448668] x17: 0000000000000000 x16: 0000000000000000 x15: 0000000000000001 [ 1179.455793] x14: 0000000000060810 x13: ffff80007fffffff x12: ffff000004190480 [ 1179.462921] x11: ffff8000853f7000 x10: ffff8000811ae000 x9 : ffff0000041900b8 [ 1179.470051] x8 : 0000000000000000 x7 : 00000000990c4001 x6 : 0000000000000007 [ 1179.477177] x5 : ffff000016d1df60 x4 : 0000000000000000 x3 : ffff00000a7d8000 [ 1179.484306] x2 : 00000000000001ff x1 : 0000000000000000 x0 : 0000000000000000 [ 1179.491433] Call trace: [ 1179.493872] pvr_mmu_op_context_unmap_curr_page+0x6c/0x134 [powervr] (P) [ 1179.500582] pvr_mmu_map+0x31c/0x388 [powervr] [ 1179.505027] pvr_vm_gpuva_map+0x40/0x88 [powervr] [ 1179.509732] __drm_gpuvm_sm_map+0x250/0x44c [drm_gpuvm] [ 1179.514952] drm_gpuvm_sm_map+0x48/0x5c [drm_gpuvm] [ 1179.519822] pvr_vm_bind_op_exec+0x64/0x70 [powervr] [ 1179.524785] pvr_vm_map+0x1f8/0x2a8 [powervr] [ 1179.529142] pvr_ioctl_vm_map+0x12c/0x188 [powervr] [ 1179.534018] drm_ioctl_kernel+0xb8/0x128 [ 1179.537941] drm_ioctl+0x21c/0x4ec [ 1179.541337] __arm64_sys_ioctl+0xac/0x108 [ 1179.545344] invoke_syscall+0x44/0x100 [ 1179.549091] el0_svc_common.constprop.0+0x40/0xe0 [ 1179.553790] do_el0_svc+0x1c/0x28 [ 1179.557106] el0_svc+0x34/0xf0 [ 1179.560159] el0t_64_sync_handler+0xd0/0xe4 [ 1179.564334] el0t_64_sync+0x198/0x19c [ 1179.567996] Code: 54000300 35000360 f9402261 79409a62 (f9400421) [ 1179.574081] ---[ end trace 0000000000000000 ]--- Fixes: ff5f643de0bf ("drm/imagination: Add GEM and VM related code") Reviewed-by: Alexandru Dadu <alexandru.dadu@imgtec.com> Reviewed-by: Alessio Belle <alessio.belle@imgtec.com> Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260922-mmu_fix-v4-1-12f1a871456a@imgtec.com Signed-off-by: Brajesh Gupta <brajesh.gupta@imgtec.com>
7 daysdrm/amd/display: Bump frame warning limit for clang builds of dmlIvan Lipski
[Why&How] When building the DML files with clang without any sanitizer or LTO, the following -Wframe-larger-than errors break the build under CONFIG_WERROR: display_mode_vba_30.c: error: stack frame size (2512) exceeds limit (2048) in 'dml30_ModeSupportAndSystemConfigurationFull' display_mode_vba_31.c: error: stack frame size (2416) exceeds limit (2048) in 'dml31_ModeSupportAndSystemConfigurationFull' display_mode_vba_314.c: error: stack frame size (2392) exceeds limit (2048) in 'dml314_ModeSupportAndSystemConfigurationFull' Clang consistently spills more than gcc, pushing the frame past the 2048 byte limit. Apply an existing approach of increasing the warn stack size to the non-sanitizer path so plain clang builds use a 3072 byte limit. Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5642 Signed-off-by: Ivan Lipski <ivan.lipski@amd.com> Signed-off-by: Alex Deucher <alexander.deucher@amd.com> (cherry picked from commit 21711b6e66bb7b41b1aec67b2d99aafe768c8fcb) Cc: stable@vger.kernel.org
7 daysdrm/amd/display: Relax DML frame limit with UBSANAlex Hung
[WHY] UBSAN instrumentation adds checks and handler calls and increases stack usage in the large DML calculation functions, similar to KASAN and KCSAN. With UBSAN enabled these files exceed the default -Wframe-larger-than limit and fail to build when -Werror is in effect. Reproduced with LLVM (make LLVM=1, clang 19.1.1), CONFIG_UBSAN=y, CONFIG_GCOV_PROFILE_ALL=y and CONFIG_DRM_AMDGPU_WERROR=y on x86_64. [HOW] Include CONFIG_UBSAN in the sanitizer check that selects the higher per-file frame warning limit in the dml and dml2_0 Makefiles. Suggested-by: Leo Li <sunpeng.li@amd.com> Assisted-by: Copilot:Claude-Opus-5.5 Signed-off-by: Alex Hung <alex.hung@amd.com> Signed-off-by: Alex Deucher <alexander.deucher@amd.com> (cherry picked from commit ebf8b0fd8508b744f85a8eee82b745b1d3502dd0) Cc: stable@vger.kernel.org
7 daysdrm/amdgpu: Fix runtime PM leak in amdgpu_debugfs_test_ib_show()Wentao Liang
amdgpu_debugfs_test_ib_show() resumes the device with pm_runtime_get_sync() before taking the reset domain semaphore with down_write_killable(). If the write lock acquisition is interrupted, the function returns without calling pm_runtime_put_autosuspend(), leaking the runtime PM reference acquired for the device and keeping the GPU awake. Drop the runtime PM reference on the interrupted down_write_killable() error path before returning. Fixes: 6049db43d6dd ("drm/amdgpu: change reset lock from mutex to rw_semaphore") Signed-off-by: Wentao Liang <vulab@iscas.ac.cn> Signed-off-by: Alex Deucher <alexander.deucher@amd.com> (cherry picked from commit ec30a576c2d4c0364549e6c04218f50704ef56c8) Cc: stable@vger.kernel.org
7 daysdrm/amdgpu: Fix last_update fence leak in amdgpu_vm_init()Wentao Liang
amdgpu_vm_init() initializes vm->last_update, vm->last_unlocked and vm->last_tlb_flush with references to the stub fence taken via dma_fence_get_stub(). The error label at the end of the function releases the last_unlocked and last_tlb_flush references with dma_fence_put(), but the reference stored in vm->last_update is never dropped, so whenever the page table root creation, the reservation of the root BO or the PASID registration fails, the stub fence reference leaks. Drop the vm->last_update reference together with the other stub fence references on the error path. Fixes: 187916e6ed9d ("drm/amdgpu: install stub fence into potential unused fence pointers") Signed-off-by: Wentao Liang <vulab@iscas.ac.cn> Signed-off-by: Alex Deucher <alexander.deucher@amd.com> (cherry picked from commit e7979c84fc05a176bdf855ee664871b1648404c9) Cc: stable@vger.kernel.org
7 daysdrm/amdgpu: Fix acpi device leak in amdgpu_acpi_enumerate_xcc()Wentao Liang
amdgpu_acpi_enumerate_xcc() looks up each XCC ACPI device with acpi_dev_get_first_match_dev(), which takes a reference to the device. The reference is dropped with acpi_dev_put() after the XCC info is initialized, but if the kzalloc_obj() allocation of the XCC info fails the function returns -ENOMEM without releasing the reference, leaking the last reference to the ACPI device. Drop the ACPI device reference on the allocation failure path before returning. Fixes: 4d5275ab0b18 ("drm/amdgpu: Add parsing of acpi xcc objects") Reviewed-by: Lijo Lazar <lijo.lazar@amd.com> Signed-off-by: Wentao Liang <vulab@iscas.ac.cn> Signed-off-by: Alex Deucher <alexander.deucher@amd.com> (cherry picked from commit 9211ef48b31ec66999cf55e04d0cbc60cd855fd5) Cc: stable@vger.kernel.org
7 daysdrm/amdgpu: Fix vmid_wait fence leak in amdgpu_ring_init()Wentao Liang
amdgpu_ring_init() initializes ring->vmid_wait with a reference to the stub fence taken via dma_fence_get_stub(). When a later step of the initialization fails, e.g. amdgpu_fence_driver_init_ring(), a writeback slot allocation or the ring buffer allocation, the function returns an error without releasing the stub fence reference and the reference is leaked if the ring is torn down without amdgpu_ring_fini(). Move the stub fence assignment to the end of the initialization, right before the ring is registered with the GPU scheduler, where no further failure is possible. The stub fence is only consumed by command submission handling in amdgpu_ids.c once the ring is up and running, so nothing reads it during the error-prone part of the initialization. Fixes: 48e9fbd1a284 ("drm/amdgpu: initialize the vmid_wait with the stub fence") Signed-off-by: Wentao Liang <vulab@iscas.ac.cn> Signed-off-by: Alex Deucher <alexander.deucher@amd.com> (cherry picked from commit f2b96986851203e9c50ca0d13aaa3581ca3e8ebd) Cc: stable@vger.kernel.org
7 daysdrm/amdkfd: fix use-after-free and multi-container gap in kfd_dev_mappingAsad Kamal
kfd_dev_mapping caches the address_space of the first /dev/kfd opener so that the GPU reset path can call unmap_mapping_range() to zap all userspace mappings of doorbell and MMIO ranges. This design has two bugs that both manifest under SRIOV with multiple containers: 1. Use-after-free / rwsem deadlock. The cached pointer refers to an inode owned by the first opener's container. When that container exits and its inode is released, kfd_dev_mapping becomes a dangling pointer. A subsequent GPU reset dereferences it inside unmap_mapping_range(), which takes i_mmap_rwsem on the freed inode, causing a hard hang observable as an uninterruptible rwsem wait. 2. Multi-container gap. Only the first opener's address_space is cached; VMAs created by later openers live in a different address_space and are never reached by unmap_mapping_range(). After a GPU reset those stale mappings keep doorbell and MMIO pages accessible to guest userspace with no GPU behind them, risking PCIe transaction timeouts and NMI panics. Fix both bugs with the same approach used by DRM core (drm_drv.c): create a private pseudo-filesystem at module init time and allocate one anonymous inode from it. In kfd_open() redirect every opener's file->f_mapping to that inode's address_space. The inode is module-owned, lives exactly as long as the amdgpu module, and collects VMAs from all openers in one address_space. A single unmap_mapping_range() call in the reset path then correctly reaches every container's mappings with no dangling pointer risk. The hang manifests as an NMI backtrace on the GPU reset workqueue stuck spinning in rwsem_down_read_slowpath() with a corrupted i_mmap_rwsem: Workqueue: amdgpu-reset-dev xgpu_ai_mailbox_flr_work [amdgpu] Call Trace: <TASK> kvm_wait+0x1f/0x40 __pv_queued_spin_lock_slowpath+0x31d/0x3a0 _raw_spin_lock_irq+0x51/0x80 rwsem_down_read_slowpath+0xb3/0x550 down_read+0x48/0xd0 unmap_mapping_range+0x71/0x140 kfd_dev_unmap_mapping_range+0x5b/0x140 [amdgpu] amdgpu_amdkfd_clear_kfd_mapping+0xd8/0x190 [amdgpu] amdgpu_device_gpu_recover+0x232/0x450 [amdgpu] xgpu_ai_mailbox_flr_work+0xb5/0xc0 [amdgpu] process_one_work+0x18e/0x3e0 worker_thread+0x2e3/0x420 kthread+0x10a/0x230 Fixes: 70cadefcc616 ("drm/amdgpu: unmap all user mappings of framebuffer and doorbell before mode1 reset") Signed-off-by: Asad Kamal <asad.kamal@amd.com> Reviewed-by: Lijo Lazar <lijo.lazar@amd.com> Signed-off-by: Alex Deucher <alexander.deucher@amd.com> (cherry picked from commit 1128b4a52de1572e87431de837fd9850cb99542c) Cc: stable@vger.kernel.org
7 daysdrm/amdgpu/vcn4.0.3: fix video_timeout unit mismatch in jpeg reset waitSunil Khatri
vcn_v4_0_3_reset_jpeg_pre_helper() passes adev->video_timeout directly to amdgpu_fence_wait_polling(), whose timeout parameter is documented and implemented in usecs (busy-wait loop decrementing by udelay(2)). adev->video_timeout is set in jiffies by amdgpu_device_get_job_timeout_settings(), via msecs_to_jiffies(). Passing it unconverted means the intended ~2s wait for outstanding JPEG fences to complete before the JPEG queue is torn down actually lasts only a couple of microseconds (HZ jiffies interpreted as usecs), so pending jobs are almost never given a real chance to finish before the reset path forces completion in the following helper. Convert the jiffies value to usecs with jiffies_to_usecs() before passing it to amdgpu_fence_wait_polling(). Fixes: d25c67fd9d6f ("drm/amdgpu/vcn4.0.3: rework reset handling") Cc: Jesse.Zhang <Jesse.Zhang@amd.com> Assisted-by: Claude:claude-sonnet-5 Signed-off-by: Sunil Khatri <sunil.khatri@amd.com> Reviewed-by: Christian König <christian.koenig@amd.com> Signed-off-by: Alex Deucher <alexander.deucher@amd.com> (cherry picked from commit 5feabbd673c10ebee22b880e4d812f08974d2ef7) Cc: stable@vger.kernel.org
7 daysdrm/amdgpu/vcn5.0.1: fix video_timeout unit mismatch in jpeg reset waitSunil Khatri
vcn_v5_0_1_reset_jpeg_pre_helper() passes adev->video_timeout directly to amdgpu_fence_wait_polling(), whose timeout parameter is documented and implemented in usecs (busy-wait loop decrementing by udelay(2)). adev->video_timeout is set in jiffies by amdgpu_device_get_job_timeout_settings(), via msecs_to_jiffies(). Passing it unconverted means the intended ~2s wait for outstanding JPEG fences to complete before the JPEG queue is torn down actually lasts only a couple of microseconds (HZ jiffies interpreted as usecs), so pending jobs are almost never given a real chance to finish before the reset path forces completion in the following helper. Convert the jiffies value to usecs with jiffies_to_usecs() before passing it to amdgpu_fence_wait_polling(). Fixes: fab47d2db5ca ("drm/amdgpu/vcn5.0.1: rework reset handling") Cc: Jesse.Zhang <Jesse.Zhang@amd.com> Assisted-by: Claude:claude-sonnet-5 Signed-off-by: Sunil Khatri <sunil.khatri@amd.com> Reviewed-by: Christian König <christian.koenig@amd.com> Signed-off-by: Alex Deucher <alexander.deucher@amd.com> (cherry picked from commit b8334fec8b90ebffcaa01001a23edca9f29a05e9) Cc: stable@vger.kernel.org
7 daysdrm/amdgpu/userq: fix double jiffies conversion in hang detect timeoutSunil Khatri
Function amdgpu_userq_start_hang_detect_work() calls msecs_to_jiffies() on adev->gfx_timeout/compute_timeout/sdma_timeout before arming hang_detect_work. These timeout values already hold jiffies values from amdgpu_device_get_job_timeout_settings() at device init. This silently shrinks the real hang-detect deadline to (2 * HZ) ms instead of the intended timeout. e.g. 500ms instead of the 2000ms default on a CONFIG_HZ=250 kernel, only coincidentally correct at HZ=1000. The shortened window is easily exceeded by ordinary fence-completion latency, causing hang_detect_work to fire and trigger a per-queue or full GPU reset for queues that are not actually hung. Pass the jiffies value directly to queue_delayed_work() instead of converting it a second time. Fixes: fc3336be9c62 ("drm/amd/amdgpu: Add independent hang detect work for user queue fence") Signed-off-by: Sunil Khatri <sunil.khatri@amd.com> Reviewed-by: Christian König <christian.koenig@amd.com> Signed-off-by: Alex Deucher <alexander.deucher@amd.com> (cherry picked from commit 13d44ca033cb74756c2aef0ade54a75cdf2f6271) Cc: stable@vger.kernel.org
7 daysdrm/amdgpu: move userq fence wait out of signalling sectionPrike Liang
The eviction fence suspend worker waits for every pending userq fence from inside a dma_fence_begin_signalling() critical section. Waiting on another DMA fence while responsible for signalling one violates the cross-driver fence contract and is reported by lockdep as a dma_fence_map dependency. Move the wait before dma_fence_begin_signalling(). Keep userq_mutex held so queue lifetime remains stable while inspecting last_fence. Fixes: fc61df151617 ("drm/amdgpu: annotate eviction fence signaling path") Signed-off-by: Prike Liang <Prike.Liang@amd.com> Reviewed-by: Vitaly Prosyak <vitaly.prosyak@amd.com> Signed-off-by: Alex Deucher <alexander.deucher@amd.com> (cherry picked from commit 3bd4fbc5ed89621340b5cd249869092691a9c81f) Cc: stable@vger.kernel.org
7 daysdrm/amd/display: Fix dc stream excess put in dm_update_crtc_state()Wentao Liang
In dm_update_crtc_state(), when a modeset is required the newly created stream is stored in dm_new_crtc_state->stream and an extra reference is taken with dc_stream_retain(). The reference returned by create_validate_stream_for_sink() is released as an extra reference at the skip_modeset label, leaving the stream owned by the new CRTC state. If amdgpu_dm_check_crtc_color_mgmt() fails afterwards, the code jumps to the fail label which releases new_stream again. Since the extra reference was already released at skip_modeset, this drops the reference owned by dm_new_crtc_state->stream and the stream is released while the atomic state still points to it, leading to a premature free of the dc stream. Set new_stream to NULL after releasing the extra reference at the skip_modeset label so that a later goto fail cannot release the reference owned by the new CRTC state. Fixes: 7cd4b70091a5 ("drm/amd/display: Rework CRTC color management") Signed-off-by: Wentao Liang <vulab@iscas.ac.cn> Signed-off-by: Alex Deucher <alexander.deucher@amd.com> (cherry picked from commit 102a47065a62dc8f6bbbb47cf082a2934282eb08) Cc: stable@vger.kernel.org
8 daysdrm/xe: Add wa_14025941587 to xe2, xe3 and xe3p platformsTangudu Tilak Tirumalesh
Avoid programming the IDLEDLY timer to less than 5 microseconds. Apply wa_14025941587 to Graphics Versions 20.01 to 35.11 and Media Versions 13.01 to 35.03 v2: Use xe_rtp_match_not_sriov_vf, move to local variable Remove warn and other knits - Matt R v3: Add verbose comment - Tejas v4: Restore IDLE_DLY register on engine reset. Add it to GUC save-restore list. -Vivek v5: Extend WA to Media Versions 13.01 to 35.03 - Vinay v6: Avoid clearing inhibit switch - Bala Refactor code accordingly by adding idle_reg_val. v7: Rebased with the divide-by-zero/overflow guards living in a separate hardening patch. v8: Preserve the Wa_16023105232 floor (DIV_ROUND_DOWN_ULL) and the maxcnt == 0 guard from the hardening patch. Round up (DIV_ROUND_UP_ULL) the Wa_14025941587 minimum conversion instead, so the tick-quantized delay cannot round back below 5 us. v9: Evaluate the Wa_16023105232 xe_gt_WARN_ON() against the value read from hardware instead of the Wa_14025941587-bumped value, so it no longer fires on the driver's own floor. Re-check the rounded-up tick value against maxcnt and floor it if tick quantization pushed it back to/above maxcnt, logging via xe_gt_dbg since this is the driver's own value, not a hardware anomaly. Assisted-by: GitHub_Copilot:claude-opus-4.8 Signed-off-by: Tangudu Tilak Tirumalesh <tilak.tirumalesh.tangudu@intel.com> Reviewed-by: Vinay Belgaumkar <vinay.belgaumkar@intel.com> Link: https://patch.msgid.link/20260916100545.779894-3-tilak.tirumalesh.tangudu@intel.com Signed-off-by: Matt Roper <matthew.d.roper@intel.com> (cherry picked from commit 9453c528fc909076468ff10df1c2e334ca5a9b00) Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
8 daysdrm/xe: harden adjust_idledly() against divide-by-zero and overflowTangudu Tilak Tirumalesh
adjust_idledly() has several corner-case issues flagged during review: 1. If xe_gt_clock_init() failed to recognise the crystal clock, gt->info.timestamp_base is 0, which makes idledly_units_ps also 0. The subsequent DIV_ROUND_CLOSEST(..., idledly_units_ps) is then a divide-by-zero and panics the kernel. 2. The tick-to-ns conversions are done in u32: idledly * idledly_units_ps, (maxcnt - 1) * 1000 Both overflow u32 before DIV_ROUND_CLOSEST() sees them. 3. If IDLE_WAIT_TIME reads back as 0, maxcnt evaluates to 0 and the maxcnt - 1 clamp wraps to 0xFFFFFFFF in u32. 4. The register only stores whole ticks, so the clamped ns value has to be converted to ticks and back. DIV_ROUND_CLOSEST() can round that conversion up past maxcnt: maxcnt = 640 ns, one tick = 666664 ps clamp: maxcnt - 1 = 639 ns ns -> ticks: 639000 / 666664 = 0.958 -> rounds to 1 tick tick -> ns: 1 * 666664 / 1000 = 667 ns 667 ns is programmed into RING_IDLEDLY, but 667 >= maxcnt (640), so xe_gt_WARN_ON() fires again on every subsequent init. Return early if timestamp_base is 0 (the unknown-crystal path). Do the conversions in u64 via the *_ULL() helpers so they cannot wrap. Clamp with a floor (DIV_ROUND_DOWN_ULL) so the programmed delay stays strictly below maxcnt, and guard the maxcnt == 0 case with a zero delay while still writing RING_IDLEDLY so INHIBIT_SWITCH_UNTIL_PREEMPTED is cleared. v2: Drop the redundant warn on the timestamp_base == 0 path; xe_gt_clock_init() already warns on an unrecognised crystal clock. Keep the early return to avoid the divide-by-zero. - Vinay v3: Field-mask the RING_IDLEDLY write with REG_FIELD_PREP(IDLE_DELAY, ...) instead of writing the raw tick count, which could clobber INHIBIT_SWITCH_UNTIL_PREEMPTED and reserved bits. Split the inhibit-switch clear from the maxcnt clamp so a set inhibit bit no longer forces a needless delay overwrite when the delay itself is already valid. Use gt_to_xe(gt) instead of gt_to_xe(hwe->gt). Fixes: d2de4410a88f ("drm/xe: Apply Wa_16023105232") Cc: stable@vger.kernel.org Assisted-by: GitHub_Copilot:claude-opus-4.8 Signed-off-by: Tangudu Tilak Tirumalesh <tilak.tirumalesh.tangudu@intel.com> Reviewed-by: Vinay Belgaumkar <vinay.belgaumkar@intel.com> Link: https://patch.msgid.link/20260916100545.779894-2-tilak.tirumalesh.tangudu@intel.com Signed-off-by: Matt Roper <matthew.d.roper@intel.com> (cherry picked from commit d864065ea25e9d12897c175de9176ce46677e176) Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
8 daysdrm/xe: Limit sg segment size to PAGE_SIZE on Xen PVSzymon Acedański
Fix display corruption on Xen PV dom0, where DMA buffers are not guaranteed machine-contiguous, in which case bounce buffering kicks in, breaking xe's memory coherency assumptions. Apply the same workaround i915 carries in i915_sg_segment_size() since commit 78a07fe777c4 ("drm/i915: stop abusing swiotlb_max_segment"). Fixes: dd08ebf6c352 ("drm/xe: Introduce a new DRM driver for Intel GPUs") Reported-by: Marek Marczykowski-Górecki <marmarek@invisiblethingslab.com> Closes: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8382 Link: https://lore.kernel.org/xen-devel/aYtznP_tT6xNPwf-@mail-itl/ Link: https://lore.kernel.org/all/20221020110308.1582518-1-hch@lst.de/ # i915 counterpart Cc: Christoph Hellwig <hch@lst.de> Cc: Robert Beckett <bob.beckett@collabora.com> Cc: stable@vger.kernel.org # v6.8+ Signed-off-by: Szymon Acedański <accek@invisiblethingslab.com> Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com> Signed-off-by: Thomas Hellström <thomas.hellstrom@linux.intel.com> Link: https://patch.msgid.link/20260916173030.3223833-1-accek@invisiblethingslab.com (cherry picked from commit 77f704158f099b952681f207478a22d5b8218edb) Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
9 daysdrm/i915: fix incorrect RCU teardown orderChristian König
i915_gem_busy_ioctl uses dma_resv_for_each_fence_unlocked() to iterate over the fences in an GEM object without holding a reference but only the RCU read side lock. What can happen here is that the GEM object is destroyed concurrently while i915_gem_busy_ioctl is still running. This won't free the GEM objects memory, but still drops all the dma_fence references. Now when dma_resv_for_each_fence_unlocked() sees a destroyed dma_fence it assumes that a new fence list was installed and re-starts the loop. But in the case of a destroyed GEM object a new fence list is never installed, only the old one freed and therefore the iteration never finishes resulting in an endless loop. The solution is to drop the fence references only after the RCU grace period. The fixes tag is not necessary the patch introducing the problem, but the one making it so worse that we need to address it. This problem was pointed out by Sashiko-bot. Signed-off-by: Christian König <christian.koenig@amd.com> Fixes: 912ff2ebd695 ("drm/i915: use the new iterator in i915_gem_busy_ioctl v2") CC: stable@vger.kernel.org Reviewed-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com> Signed-off-by: Tvrtko Ursulin <tursulin@ursulin.net> Link: https://lore.kernel.org/r/20260903113621.54660-1-christian.koenig@amd.com (cherry picked from commit 5113479556025093bf8133bb2dcaa33be2d50921) Signed-off-by: Jani Nikula <jani.nikula@intel.com>
9 daysdrm/i915/dp: use EXPORT_SYMBOL_IF_KUNIT() for kunit helpersJani Nikula
Use EXPORT_SYMBOL_IF_KUNIT() instead of the regular EXPORT_SYMBOL() to export the symbols to the kunit namespace. Otherwise, the symbols get exported for all the kernel to see, and the corresponding MODULE_IMPORT_NS("EXPORTED_FOR_KUNIT_TESTING") in the tests is meaningless. Fixes: 2eb9982ff179 ("drm/i915/kunit: Export link training and caps funcs for testing") Cc: Imre Deak <imre.deak@intel.com> Reviewed-by: Imre Deak <imre.deak@intel.com> Link: https://patch.msgid.link/20260915160620.779372-1-jani.nikula@intel.com Signed-off-by: Jani Nikula <jani.nikula@intel.com> (cherry picked from commit 4ffdb772716e4279d63dfaaadf965da73e799aeb)
9 daysdrm/i915/dp_mst: Fix configuring TUs for a disconnected streamImre Deak
During an atomic commit after all the MST stream CRTC state is computed the driver ensures that the sum of TUs of all the streams on a given MST topology link is within limits (63 for 8b10 and 64 for 128b132b). For a disconnected stream the DRM MST core's BW verification doesn't ensure this, because the topology state it uses for this is destroyed as soon as the stream (i.e. MST connector/port) is disconnected. The driver should keep the link state valid even for such disconnected streams, as userspace may disable them one-by-one only in a deferred way. Ensure the link's sum of TUs stays within limits in this case by simply reusing the maximum link BPP limit from the stream's (i.e. CRTC's) old state. The disconnection can happen either via the whole topology getting disconnected or via only the given stream's port getting disconnected. Check for both of these conditions separately, as a connector gets unregistered after a link disconnect event only in a deferred way. Cc: stable@vger.kernel.org # v6.10+ Link: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16073 Link: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16384 Reviewed-by: Luca Coelho <luciano.coelho@intel.com> Signed-off-by: Imre Deak <imre.deak@intel.com> Link: https://patch.msgid.link/20260907174413.741851-2-imre.deak@intel.com (cherry picked from commit ee00f8fbb2b202002ab90834e02e9ba372773a36) Signed-off-by: Jani Nikula <jani.nikula@intel.com>
9 daysdrm/i915/dp_mst: Fix configuring FEC for a disconnected streamImre Deak
During an atomic commit after all the MST stream CRTC state is computed the driver ensures that the FEC is configured the same way (enabled or disabled) for all the streams on a given MST topology's link. drm_dp_mst_port_downstream_of_parent() used to determine if a stream is downstream of an MST port will return false if the whole topology is disconnected, since in that case it can't verify that the port/ parent_port passed to it is in the given MST topology. This is a problem during the above FEC configuration check, since intel_dp_mst_check_dsc_change()->get_pipes_downstream_of_mst_ports() will not return all the stream CRTCs/pipes for the topology as expected. Since passing parent_port==NULL to get_pipes_downstream_of_mst_port() is meant to return all the streams for the given topology (i.e. mst_mgr) skip checking if an MST port is downstream of a parent port in this case. This fixes a problem where the FEC configuration check explained above failed to ensure that all streams' FEC is configured the same way if the topology was disconnected, leading to a FEC state mismatch error. Cc: stable@vger.kernel.org # v6.10+ Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16073 Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16384 Reviewed-by: Luca Coelho <luciano.coelho@intel.com> Signed-off-by: Imre Deak <imre.deak@intel.com> Link: https://patch.msgid.link/20260907174413.741851-1-imre.deak@intel.com (cherry picked from commit 270681fbffbba2b6ccf5b7e3c34b8b563b36167f) Signed-off-by: Jani Nikula <jani.nikula@intel.com>
9 daysdrm/i915/quirks: Limit eDP rate to HBR2 on HP Pavilion Plus 14-ew1Ankit Nautiyal
The eDP panel on the HP Pavilion Plus Laptop 14-ew1xxx advertises HBR3 while leaving the TPS4 support bit clear. The output however flickers, once link is trained with HBR3. Until commit 8c9006283e4b ("Revert "drm/i915/dp: Reject HBR3 when sink doesn't support TPS4"") such sinks were capped at HBR2 by the TPS4 check which incidentally kept this panel stable. That check was reverted because other panels legitimately need HBR3 without advertising TPS4, and the per-machine QUIRK_EDP_LIMIT_RATE_HBR2 was introduced to handle the affected machines instead. Add the machine to the list of devices that need the QUIRK_EDP_LIMIT_RATE_HBR2. Fixes: 8c9006283e4b ("Revert "drm/i915/dp: Reject HBR3 when sink doesn't support TPS4"") Reported-by: Annoy Cc <annoycc@gmail.com> Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16743 Cc: <stable@vger.kernel.org> # v6.18+ Tested-by: Annoy Cc <annoycc@gmail.com> Signed-off-by: Ankit Nautiyal <ankit.k.nautiyal@intel.com> Reviewed-by: Nemesa Garg <nemesa.garg@intel.com> Link: https://patch.msgid.link/20260907034555.2753846-1-ankit.k.nautiyal@intel.com (cherry picked from commit 550b703fdbb2a2022faa75b4b11ab135241afbd9) Signed-off-by: Jani Nikula <jani.nikula@intel.com>
9 daysdrm/i915/psr: Clear stale sel fetch enable bits on sel fetch disableNemesa Garg
Selective fetch is dropped while pipe CRC is active, and the planes keep their SEL_FETCH_PLANE_CTL / SEL_FETCH_CUR_CTL enable bit set in hardware over that. A plane disabled while selective fetch is off never gets the bit cleared, as the disable path is guarded by enable_psr2_sel_fetch. Once selective fetch comes back the hardware resumes fetching for a plane that is no longer enabled and keeps its DDB range reserved. Clear the bits as selective fetch is turned off instead. Atomic check has both the old and the new crtc state, so record the transition there and let the plane and cursor arm paths write the registers to 0 for that commit. v2: Drop the old_crtc_state->hw.active check. [Jouni] Fixes: b1f5279b5981 ("drm/i915/psr: Move plane sel fetch configuration into plane source files") Closes: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8739 Assisted-by: Copilot:Claude-Opus-5 Signed-off-by: Nemesa Garg <nemesa.garg@intel.com> Reviewed-by: Jouni Högander <jouni.hogander@intel.com> Signed-off-by: Suraj Kandpal <suraj.kandpal@intel.com> Link: https://patch.msgid.link/20260909110332.3528029-3-nemesa.garg@intel.com (cherry picked from commit a4c0e7f80429eda6990960971aebd4e4b9533cc6) Signed-off-by: Jani Nikula <jani.nikula@intel.com>
9 daysdrm/xe/vm: nuke PTs only after unlinking contested VMAsMatthew Auld
In xe_vm_close_and_put(), external-BO VMAs are queued on the contested list for deferred destruction via xe_vma_destroy_unlocked(). However, xe_vm_pt_destroy() was previously invoked before processing contested VMAs, destroying vm->pt_root while those VMAs were still linked to their respective buffer objects (vm_bo->list.gpuva). If a concurrent thread evicts one of those shared buffer objects, xe_bo_trigger_rebind() holding only bo->resv walks the BO's VMAs and, in fault mode, calls xe_vm_invalidate_vma() -> xe_pt_zap_ptes(). Because vm->pt_root[tile->id] is already NULL, dereferencing pt->level causes a NULL ptr deref. Fix this by deferring xe_vm_free_scratch() and xe_vm_pt_destroy() until after all contested VMAs have been unlinked and destroyed. User is reporting hitting a NULL ptr deref in xe_pt_zap_ptes(), which could be explained by this race. Assisted-by: LLM Fixes: b06d47be7c83 ("drm/xe: Port Xe to GPUVA") Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/9290 Signed-off-by: Matthew Auld <matthew.auld@intel.com> Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com> Cc: Matthew Brost <matthew.brost@intel.com> Cc: <stable@vger.kernel.org> # v6.12+ Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com> Reviewed-by: Matthew Brost <matthew.brost@intel.com> Link: https://patch.msgid.link/20260918131034.598078-2-matthew.auld@intel.com (cherry picked from commit c2863648959489767f08892fd6e90577d2ea0b6a) Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
9 daysdrm/xe: Keep walking on SVM eviction failureMatthew Brost
The desired behavior for SVM eviction failures, which can occur due to various uncontrollable races, is for TTM to continue walking the LRU list and look for another eviction candidate. This is expressed by returning -ENOSPC from the ->move() callback. Adjust the SVM eviction failure path because of races in ->move() to return -ENOSPC so that TTM continues searching for another buffer to evict. Fixes: 3ca608dc7561 ("drm/xe: Basic SVM BO eviction") Cc: stable@vger.kernel.org Signed-off-by: Matthew Brost <matthew.brost@intel.com> Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com> Link: https://patch.msgid.link/20260917203158.292823-1-matthew.brost@intel.com Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com> (cherry picked from commit 36a86c23588b8f57c9d20feb4cf5a2ab27e3baba) Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
9 daysdrm/xe/tlb_inval: Treat wedged-device invalidations as completeShuicheng Lin
A TLB invalidation issued on a wedged device fails with -ENOTRECOVERABLE. xe_tlb_inval_issue() squashes only -ECANCELED, so the error reaches callers that treat it as unexpected and WARN, tainting the kernel on a wedge that was deliberately caused: ggtt_invalidate_gt_tlb() drivers/gpu/drm/xe/xe_ggtt.c xe_svm_invalidate() drivers/gpu/drm/xe/xe_svm.c xe_bo_trigger_rebind() drivers/gpu/drm/xe/xe_bo.c xe_vma_userptr_do_inval() drivers/gpu/drm/xe/xe_userptr.c igt@xe_exec_reset@gt-reset-fault-injection hits the GGTT one, turning an otherwise passing run into an abort: *ERROR* SIGID=102 FATAL (-EIO) WEDGED: Device declared wedged! Tile0: GT1: Failed to invalidate GGTT (-ENOTRECOVERABLE) WARNING: drivers/gpu/drm/xe/xe_ggtt.c:588 at ggtt_invalidate_gt_tlb Workqueue: xe-guc-destroy-wq __guc_exec_queue_destroy_async [xe] ggtt_node_remove+0xe3/0x100 [xe] xe_ggtt_remove_bo+0x89/0x2c0 [xe] xe_ttm_bo_destroy+0xcb/0x330 [xe] ... xe_lrc_destroy+0x74/0x90 [xe] xe_exec_queue_fini+0x2d/0x60 [xe] -ECANCELED and -ENOTRECOVERABLE mean the same thing at this layer: the message was dropped rather than delivered, and the fence has already been signalled before the error is returned, so there is nothing left to wait for. Squash both. A wedged device is only recovered by a fresh initialisation, so the error return in xe_bo_trigger_rebind() becomes unreachable by design. v2: fix all invalidation paths. (Sashiko) Fixes: 50fa9acac26f ("drm/xe/guc: distinguish wedged from recoverable cancellation") Assisted-by: Claude:claude-opus-5 Cc: Sk Anirban <sk.anirban@intel.com> Cc: Matthew Brost <matthew.brost@intel.com> Signed-off-by: Shuicheng Lin <shuicheng.lin@intel.com> Reviewed-by: Matthew Brost <matthew.brost@intel.com> Signed-off-by: Matthew Brost <matthew.brost@intel.com> Link: https://patch.msgid.link/20260914215318.200603-1-shuicheng.lin@intel.com (cherry picked from commit af14e3705345cb57c53b237169873233052a5c16) Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
9 daysdrm/xe/gt_throttle: Report power brake as a throttle reason on CRISk Anirban
CRI defines bit 5 of the perf limit reasons register as a power brake (PWRBRK) indicator. Add PWRBRK_MASK and a reason_pwrbrk sysfs attribute for CRI in place of reason_ratl. Signed-off-by: Sk Anirban <sk.anirban@intel.com> Fixes: 8578e6d0546c ("drm/xe/gt_throttle: Drop individual show functions") Reviewed-by: Raag Jadav <raag.jadav@intel.com> Signed-off-by: Matthew Brost <matthew.brost@intel.com> Link: https://patch.msgid.link/20260909114931.1039331-2-sk.anirban@intel.com (cherry picked from commit e199c851c0ab608a0ca89e7be1756a1461b41a2c) Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
10 daysdrm/bridge: samsung-dsim: fix TE GPIO lifetime for host attachLi Youhong
When the Exynos DSI driver was generalized into samsung-dsim, the TE GPIO acquisition was switched from gpiod_get_optional() to devm_gpiod_get_optional() while keeping the matching gpiod_put() calls. That combination is wrong for a managed descriptor. However, dropping the puts and keeping the managed get is also wrong: samsung_dsim_register_te_irq() runs from the DSI host attach callback, and host detach/reattach can happen without destroying the device that owns the managed action. A second attach would then request the GPIO again without having released it. Switch back to a non-managed gpiod_get_optional() and keep the explicit gpiod_put() on the request_irq() error path and in samsung_dsim_unregister_te_irq(). Fixes: e7447128ca4a ("drm: bridge: Generalize Exynos-DSI driver into a Samsung DSIM bridge") Suggested-by: Luca Ceresoli <luca.ceresoli@bootlin.com> Signed-off-by: Li Youhong <liyouhong@kylinos.cn> Reviewed-by: Luca Ceresoli <luca.ceresoli@bootlin.com> Tested-by: Luca Ceresoli <luca.ceresoli@bootlin.com> Link: https://patch.msgid.link/20260904014958.1572918-1-dayou5941@163.com [Luca: remove unnecessary comment] Signed-off-by: Luca Ceresoli <luca.ceresoli@bootlin.com>
10 daysdrm/virtio: sync shmem backing on guest-bound transfersBenjamin Leggett
virtio_gpu_cmd_transfer_to_host_{2d,3d}() sync the shmem backing for the device before the transfer, but nothing syncs for the CPU when a transfer runs the other way. That breaks two ways. Where the DMA layer bounces, the device writes into the bounce buffer while the guest keeps reading the original pages. Where DMA is not coherent, the device writes memory while the CPU keeps stale cache lines, because nothing reaches arch_sync_dma_for_cpu(). Either way DRM_IOCTL_VIRTGPU_TRANSFER_FROM_HOST hands back stale data. Sashiko originally found this in https://lore.kernel.org/dri-devel/20260806231002.27B4D1F000E9@smtp.kernel.org but the suggestion there to fix this with dma_sync_sgtable_for_cpu() isn't a sufficient fix, for two reasons. - The transfer is asynchronous. virtio_gpu_cmd_transfer_from_host_3d() only queues the command, so a sync there would run before the device had written anything. It belongs on completion, and ahead of any fence signalling. A waiter woken by the fence would otherwise race the sync and read the backing pages regardless. It needs its own pass over the reclaim list rather than a step inside the existing one, because virtio_gpu_fence_event_process() also signals every earlier fence in the same context, so any entry in that loop may signal an earlier entry's fence. - The transfer is also partial, carrying an offset, a level and a box. Where the mapping bounces, a sync for the CPU copies the whole mapping back, so unless the mapping is primed first the regions the device did not write come back holding whatever the bounce buffer contained, discarding data the guest owned. So the fix: Prime the mapping before queueing, tag the vbuffer, and sync for the CPU on completion before the fence is signalled. A second transfer must not snapshot the mapping while an earlier one is still in flight, or the snapshot would predate whatever the CPU wrote once the earlier fence signalled and the later sync would discard it. To mitigate this, wait for outstanding fences under the reservation before priming. Neither sync copies anything unless the mapping genuinely bounces: swiotlb_sync_single_for_cpu() and its Xen counterpart look the address up in the bounce pool and return early when it is absent. On a platform with non-coherent DMA they still perform the necessary cache maintenance. The range cannot be narrowed to the box, since for a non-blob resource virtio_gpu_transfer_from_host_ioctl() rejects a caller-supplied stride and layer_stride, leaving the layout to the host and the guest with no way to work out which bytes the device writes. A host3d guest blob does carry both, so its extent could be bounded, but the sync is left whole there too rather than special-cased: priming makes the untouched regions round-trip unchanged either way. Behaviour changes worth noting: - TRANSFER_FROM_HOST can now block, where before it returned as soon as the command was queued. Repeated readbacks of one resource serialise, and a readback can wait behind an earlier queued command that touched it, since virtio_gpu_array_add_fence() tags uploads, execbufs and plane flushes alike with DMA_RESV_USAGE_WRITE. -ERESTARTSYS was already possible here via dma_resv_lock_interruptible(). - A CPU write racing an in-flight transfer to the same resource is now lost, where before it survived and the transfer was lost instead. Priming captures the pages as of queueing, so a write landing before completion is overwritten by the sync. - TRANSFER_TO_HOST can also block now, but only while a guest-bound transfer on the same resource is outstanding, which happens only for callers that issue both without waiting. - Where a batch of completions contains a guest-bound transfer, the sync pass delays fence signalling for the whole batch. Only bounced pages are copied and the swiotlb pool bounds it. A batch with no such transfer is unaffected. Tested under QEMU on x86 with swiotlb=force and virtio-vga-gl iommu_platform=on, which forces both preconditions required to hit the original bug. Fixes: a3b815f09bb8 ("drm/virtio: add iommu support.") Reported-by: Sashiko AI review <sashiko-bot@kernel.org> Closes: https://lore.kernel.org/dri-devel/20260806231002.27B4D1F000E9@smtp.kernel.org/ Signed-off-by: Benjamin Leggett <benjamin@edera.io> Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com> Link: https://patch.msgid.link/20260814-virtgpu-from-host-sync-v4-1-64dd736b1779@edera.io
10 daysdrm/virtio: Add pixel blend mode property to cursor planeShixiong Ou
The cursor plane exposes a format with an alpha channel (DRM_FORMAT_ARGB8888) without a pixel blend mode property. Since commit 860e748bddcc ("drm: ensure blend mode supported if pixel format with alpha exposed") this triggers a warning during drm_mode_config_validate(): [ 0.649020] ------------[ cut here ]------------ [ 0.649040] [PLANE:36:plane-1] pixel format with alpha exposed but blend mode not setup [ 0.649081] WARNING: drivers/gpu/drm/drm_mode_config.c:872 at drm_mode_config_validate ...... [ 0.649761] Call trace: [ 0.649764] drm_mode_config_validate+0x398/0x558 [drm] (P) [ 0.649912] drm_dev_register+0x1cc/0x2a0 [drm] [ 0.650058] virtio_gpu_probe+0xd4/0x1c0 [virtio_gpu] [ 0.650088] virtio_dev_probe+0x1c8/0x310 ...... [ 0.650261] ---[ end trace 0000000000000000 ]--- Create the property with the only supported blend mode, DRM_MODE_BLEND_PREMULTI, which is also the property's default and matches what userspace had to assume before the property existed. Fixes: 860e748bddcc ("drm: ensure blend mode supported if pixel format with alpha exposed") Reported-by: Ye Liu <liuye@kylinos.cn> Signed-off-by: Shixiong Ou <oushixiong@kylinos.cn> Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com> Link: https://patch.msgid.link/20260828090123.784944-1-oushixiong1025@163.com
10 daysdrm/virtio: fix NULL pointer dereference on fence allocation failurePeiyang He
virtio_gpu_fence_alloc() can fail due to memory pressure and return NULL, but its caller like virtio_gpu_init_submit() never checks it. Later, virtio_gpu_init_submit() passes the NULL fence to virtio_gpu_fence_event_create(), which unconditionally dereferences it. Found when fuzzing the virtio driver with Syzkaller: Oops: general protection fault, probably for non-canonical address 0xdffffc0000000012: 0000 [#1] SMP KASAN NOPTI KASAN: null-ptr-deref in range [0x0000000000000090-0x0000000000000097] CPU: 1 UID: 0 PID: 9991 Comm: syz.0.121 Not tainted 7.2.0 #4 PREEMPT(full) Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS rel-1.17.0-0-gb52ca86e094d-prebuilt.qemu.org 04/01/2014 RIP: 0010:virtio_gpu_fence_event_create drivers/gpu/drm/virtio/virtgpu_submit.c:295 [inline] RIP: 0010:virtio_gpu_init_submit drivers/gpu/drm/virtio/virtgpu_submit.c:398 [inline] RIP: 0010:virtio_gpu_execbuffer_ioctl+0xc78/0x1aa0 drivers/gpu/drm/virtio/virtgpu_submit.c:505 Code: 85 ed 0f 85 21 09 00 00 e8 05 5a c9 fb 48 8b 44 24 10 48 8d b8 90 00 00 00 48 b8 00 00 00 00 00 fc ff df 48 89 fa 48 c1 ea 03 <80> 3c 02 00 0f 85 9a 0d 00 00 48 8b 44 24 10 4c 89 b0 90 00 00 00 RSP: 0018:ffffc900039dfad0 EFLAGS: 00010216 RAX: dffffc0000000000 RBX: ffffc900039dfdd8 RCX: ffffffff85f6fd3d RDX: 0000000000000012 RSI: ffffffff85f6fd4b RDI: 0000000000000090 RBP: 0000000000000000 R0virtio_gpu_virgl_process_cmd: ctrl 0x102, error 0x1203 R10: 0000000000000000 R11: 0000000000000000 R12: ffff8880132c4000 R13: 0000000000000000 R14: ffff888073b6c700 R15: 000000000000003b FS: 00007fab480b96c0(0000) GS:ffff8880eb6e9000(0000) knlGS:0000000000000000 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 CR2: 00007effbf5e55a8 CR3: 0000000048d19000 CR4: 0000000000350ef0 Call Trace: <TASK> drm_ioctl_kernel+0x1f4/0x3e0 drivers/gpu/drm/drm_ioctl.c:817 drm_ioctl+0x5f4/0xc70 drivers/gpu/drm/drm_ioctl.c:914 vfs_ioctl fs/ioctl.c:51 [inline] __do_sys_ioctl fs/ioctl.c:597 [inline] __se_sys_ioctl fs/ioctl.c:583 [inline] __x64_sys_ioctl+0x18e/0x210 fs/ioctl.c:583 do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline] do_syscall_64+0x116/0x800 arch/x86/entry/syscall_64.c:94 entry_SYSCALL_64_after_hwframe+0x77/0x7f RIP: 0033:0x7fab471a82bd Code: ff c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 c7 c1 b0 ff ff ff f7 d8 64 89 01 48 RSP: 002b:00007fab480b9018 EFLAGS: 00000246 ORIG_RAX: 0000000000000010 RAX: ffffffffffffffda RBX: 00007fab47435fa0 RCX: 00007fab471a82bd RDX: 00002000000000c0 RSI: 00000000c0406442 RDI: 0000000000000003 RBP: 00007fab480b9080 R08: 0000000000000000 R09: 0000000000000000 R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000001 R13: 00007fab47436038 R14: 00007fab47435fa0 R15: 00007ffe85ab0740 </TASK> Modules linked in: ---[ end trace 0000000000000000 ]--- RIP: 0010:virtio_gpu_fence_event_create drivers/gpu/drm/virtio/virtgpu_submit.c:295 [inline] RIP: 0010:virtio_gpu_init_submit drivers/gpu/drm/virtio/virtgpu_submit.c:398 [inline] RIP: 0010:virtio_gpu_execbuffer_ioctl+0xc78/0x1aa0 drivers/gpu/drm/virtio/virtgpu_submit.c:505 Code: 85 ed 0f 85 21 09 00 00 e8 05 5a c9 fb 48 8b 44 24 10 48 8d b8 90 00 00 00 48 b8 00 00 00 00 00 fc ff df 48 89 fa 48 c1 ea 03 <80> 3c 02 00 0f 85 9a 0d 00 00 48 8b 44 24 10 4c 89 b0 90 00 00 00 RSP: 0018:ffffc900039dfad0 EFLAGS: 00010216 RAX: dffffc0000000000 RBX: ffffc900039dfdd8 RCX: ffffffff85f6fd3d RDX: 0000000000000012 RSI: ffffffff85f6fd4b RDI: 0000000000000090 RBP: 0000000000000000 R08: 0000000000000005 R09: 0000000000000000 R10: 0000000000000000 R11: 0000000000000000 R12: ffff8880132c4000 R13: 0000000000000000 R14: ffff888073b6c700 R15: 000000000000003b FS: 00007fab480b96c0(0000) GS:ffff888098ae9000(0000) knlGS:0000000000000000 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 CR2: 00007f24c3759000 CR3: 0000000048d19000 CR4: 0000000000350ef0 ---------------- Code disassembly (best guess): 0: 85 ed test %ebp,%ebp 2: 0f 85 21 09 00 00 jne 0x929 8: e8 05 5a c9 fb call 0xfbc95a12 d: 48 8b 44 24 10 mov 0x10(%rsp),%rax 12: 48 8d b8 90 00 00 00 lea 0x90(%rax),%rdi 19: 48 b8 00 00 00 00 00 movabs $0xdffffc0000000000,%rax 20: fc ff df 23: 48 89 fa mov %rdi,%rdx 26: 48 c1 ea 03 shr $0x3,%rdx * 2a: 80 3c 02 00 cmpb $0x0,(%rdx,%rax,1) <-- trapping instruction 2e: 0f 85 9a 0d 00 00 jne 0xdce 34: 48 8b 44 24 10 mov 0x10(%rsp),%rax 39: 4c 89 b0 90 00 00 00 mov %r14,0x90(%rax) Fix by checking virtio_gpu_fence_alloc() in virtio_gpu_init_submit() and returning -ENOMEM before any later code can dereference the NULL fence. Fixes: 70d1ace56db6 ("drm/virtio: Conditionally allocate virtio_gpu_fence") Cc: stable@vger.kernel.org Signed-off-by: Peiyang He <peiyang_he@smail.nju.edu.cn> Assisted-by: Codex:gpt-5.5 Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com> Link: https://patch.msgid.link/00EFE4BA92889B14+20260909091114.2622550-1-peiyang_he@smail.nju.edu.cn
10 daysRevert "drm/virtio: Allow importing prime buffers when 3D is enabled"Dmitry Osipenko
Guest userspace may import udmabuf to vrend. Vrend doesn't support guest blobs, and thus, further 3d operations with the imported blob are failing. Typical scenario of the problem shown with mouse cursor RGBA image imported into virtio-gpu, which previously was rejected by virtio-gpu driver. Revert enabling guest blobs importing into vrend to fix the regression. Link: https://gitlab.freedesktop.org/virgl/virglrenderer/-/work_items/674 Fixes: df4dc947c46b ("drm/virtio: Allow importing prime buffers when 3D is enabled") Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com> Reviewed-by: Val Packett <val@invisiblethingslab.com> Link: https://patch.msgid.link/20260911144204.2089401-1-dmitry.osipenko@collabora.com
10 daysdrm/virtio: release the GEM object on virtio_gpu_vram_create() errorsJunrui Luo
virtio_gpu_vram_create() frees the object with a bare kfree(vram) on both error paths after drm_gem_private_object_init() has run, and on the second one after drm_gem_create_mmap_offset() has linked obj->vma_node into the device's VMA offset manager. The freed object stays in that interval tree, so a later lookup or insertion walks freed memory, and the dma_resv and gpuva lock are never destroyed. Call drm_gem_object_release() before kfree() on both paths. Fixes: 16845c5d5409 ("drm/virtio: implement blob resources: implement vram object") Assisted-by: Claude:claude-opus-5 Signed-off-by: Junrui Luo <moonafterrain@outlook.com> Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com> Link: https://patch.msgid.link/20260915-fixes-v2-4-a0d799e4db66@outlook.com
10 daysdrm/virtio: fix object leaks in virtio_gpu_resource_create_blob_ioctl()Junrui Luo
virtio_gpu_resource_create_blob_ioctl() calls drm_gem_object_release() on both the virtio_gpu_resource_assign_uuid() and drm_gem_handle_create() error paths instead of dropping the reference it owns, so obj->funcs->free() never runs and the virtio_gpu_object, the resource id and the host-side resource are leaked. Use drm_gem_object_put() instead. Fixes: 897b4d1acaf5 ("drm/virtio: implement blob resources: resource create blob ioctl") Assisted-by: Claude:claude-opus-5 Signed-off-by: Junrui Luo <moonafterrain@outlook.com> Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com> Link: https://patch.msgid.link/20260915-fixes-v2-3-a0d799e4db66@outlook.com
10 daysdrm/virtio: fix object leak in virtio_gpu_resource_create_ioctl()Junrui Luo
virtio_gpu_resource_create_ioctl() calls drm_gem_object_release() on the drm_gem_handle_create() error path instead of dropping the reference it owns, so obj->funcs->free() never runs and the virtio_gpu_object, its pages and sg table, the resource id and the host-side resource are leaked. Use drm_gem_object_put() instead. Fixes: 62fb7a5e1096 ("virtio-gpu: add 3d/virgl support") Assisted-by: Claude:claude-opus-5 Signed-off-by: Junrui Luo <moonafterrain@outlook.com> Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com> Link: https://patch.msgid.link/20260915-fixes-v2-2-a0d799e4db66@outlook.com
10 daysdrm/virtio: fix object leak when drm_gem_handle_create() failsJunrui Luo
virtio_gpu_gem_create() owns the reference taken by virtio_gpu_object_create(). On the drm_gem_handle_create() error path it calls drm_gem_object_release() instead of dropping that reference. drm_gem_object_release() is the inverse of drm_gem_object_init() and does not touch the reference count or call obj->funcs->free(), so it is only correct as the last step of a destructor, as in virtio_gpu_cleanup_object(). Using it here leaves the bo at refcount 1 with no remaining reference, so virtio_gpu_free_object() never runs and the shmem pages, sg table and virtio_gpu_object are leaked. Since virtio_gpu_object_create() has already set bo->created, VIRTIO_GPU_CMD_RESOURCE_UNREF is not queued either, leaking the host-side resource and the resource id. drm_gem_handle_create_tail() drops the handle reference on all of its internal error paths, so the caller only has to drop its own. Use drm_gem_object_put(), matching the success path below. Fixes: dc5698e80cf7 ("Add virtio gpu driver.") Reported-by: Yuhao Jiang <danisjiang@gmail.com> Assisted-by: Claude:claude-opus-5 Signed-off-by: Junrui Luo <moonafterrain@outlook.com> Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com> Link: https://patch.msgid.link/20260915-fixes-v2-1-a0d799e4db66@outlook.com
10 daysdrm/virtio: fix memory leak of fence event on execbuffer failurePeiyang He
virtio_gpu_execbuffer_ioctl() reserves a DRM event with drm_event_reserve_init() when VIRTGPU_EXECBUF_RING_IDX selects a ring that userspace has enabled polling for. virtio_gpu_init_submit() does this before the BO handles, the command buffer, the syncobj arrays and the in-fence are processed, so every later error path runs with the event already pending, including plain argument validation failures such as an invalid bo_handle or an in-syncobj that carries no fence. On those paths, virtio_gpu_cleanup_submit() drops the out-fence without cancelling the event. The fence is freed without ever having been emitted, taking the only driver-side pointer to the event with it. Closing the DRM file does not help. drm_events_release() unlinks pending events but deliberately leaves the freeing to the driver's later drm_send_event(), which never runs for an orphaned event, so the allocation is leaked permanently. Found when fuzzing the virtio driver with Syzkaller: BUG: memory leak unreferenced object 0xffff88802c176e80 (size 96): comm "syz.1.367", pid 10561, jiffies 4294960122 hex dump (first 32 bytes): 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 ................ c8 6e 17 2c 80 88 ff ff 00 00 00 00 00 00 00 00 .n.,............ backtrace (crc e1973c6b): kmemleak_alloc_recursive include/linux/kmemleak.h:44 [inline] slab_post_alloc_hook mm/slub.c:4597 [inline] slab_alloc_node mm/slub.c:4917 [inline] __kmalloc_cache_noprof+0x49d/0x6f0 mm/slub.c:5485 _kmalloc_noprof include/linux/slab.h:988 [inline] _kzalloc_noprof include/linux/slab.h:1309 [inline] virtio_gpu_fence_event_create drivers/gpu/drm/virtio/virtgpu_submit.c:282 [inline] virtio_gpu_init_submit drivers/gpu/drm/virtio/virtgpu_submit.c:398 [inline] virtio_gpu_execbuffer_ioctl+0xbbf/0x1aa0 drivers/gpu/drm/virtio/virtgpu_submit.c:505 drm_ioctl_kernel+0x1f4/0x3e0 drivers/gpu/drm/drm_ioctl.c:817 drm_ioctl+0x5f4/0xc70 drivers/gpu/drm/drm_ioctl.c:914 vfs_ioctl fs/ioctl.c:51 [inline] __do_sys_ioctl fs/ioctl.c:597 [inline] __se_sys_ioctl fs/ioctl.c:583 [inline] __x64_sys_ioctl+0x18e/0x210 fs/ioctl.c:583 do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline] do_syscall_64+0x116/0x800 arch/x86/entry/syscall_64.c:94 entry_SYSCALL_64_after_hwframe+0x77/0x7f Fix by cancelling and freeing the DRM event on the execbuffer error path before dropping the fence. Clear the fence's event pointer after cancellation so it does not retain a dangling pointer. Fixes: cd7f5ca33585 ("drm/virtio: implement context init: add virtio_gpu_fence_event") Cc: stable@vger.kernel.org Signed-off-by: Peiyang He <peiyang_he@smail.nju.edu.cn> Assisted-by: Codex:gpt-5.6-luna Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com> Link: https://patch.msgid.link/D320EAB5680C1411+20260908121323.2405044-1-peiyang_he@smail.nju.edu.cn
12 daysMerge tag 'drm-misc-fixes-2026-09-17' of ↵Dave Airlie
https://gitlab.freedesktop.org/drm/misc/kernel into drm-fixes Two ttm fixes for ttm_tt_swapout(), one page-alignment and one overflow fix for dma-buf, a drm_pending_vblank_event leak fix for drm, suspend/resume fixes for nouveau, one out-of-bounds access fix for gud, a use-after-free fix for vc4, a fence signaling fix, a race condition fix for sched, planes formats fixes for verisilicon, and add the blend mode property for loongson Signed-off-by: Dave Airlie <airlied@redhat.com> From: Maxime Ripard <self@mripard.dev> Link: https://patch.msgid.link/aqvVENQ4ksJEIcdb@houat
12 daysdrm/nouveau/disp: don't reject HDMI config on cards without SCDCGiuseppe Ranieri
nv50_hdmi_enable() passes the sink's SCDC capability from its EDID straight through to nvif_outp_hdmi(). On pre-Maxwell-2 cards there is no hdmi->scdc callback, so nvkm_uoutp_mthd_hdmi() rejects the whole configuration with -EINVAL, and nv50_hdmi_enable() returns before hdmi->ctrl() runs and before the AVI and VSI infoframes are sent. The result on such a card driving an SCDC-capable HDMI 2.0 sink is that HDMI audio silently stops working. Video is unaffected, and nothing is logged, which makes the failure hard to attribute. SCDC is optional, and the hdmi->scdc() call further down is already guarded against a missing callback. Requesting it on a card that cannot do it need not invalidate the rest of the HDMI configuration, so drop that term from the condition and let the existing guard skip SCDC alone. Fixes: 6c6abab20b99 ("drm/nouveau/disp: add output hdmi config method") Signed-off-by: Giuseppe Ranieri <giuseppe@ranieri.dev> Co-Authored-By: Tano Dzhinski <tano.dzhinski@gmail.com> Signed-off-by: Tano Dzhinski <tano.dzhinski@gmail.com> Tested-by: Tano Dzhinski <tano.dzhinski@gmail.com> Reviewed-by: Lyude Paul <lyude@redhat.com> Signed-off-by: Lyude Paul <lyude@redhat.com> Link: https://patch.msgid.link/20260917215114.1136715-1-tano.dzhinski@gmail.com
12 daysdrm/nouveau: don't bump pin count on failed re-pin in nouveau_bo_pin_locked()Peiyang He
nouveau_bo_pin_locked() checks whether an already pinned BO is in a memory domain compatible with a new pin request. When the domains are incompatible, it sets -EBUSY but still calls ttm_bo_pin() before returning. Callers treat a failed nouveau_bo_pin() as not having acquired a new pin, so the extra pin count is never decreased by a matching unpin. This triggers the warning in ttm_bo_release(): WARN_ON_ONCE(bo->pin_count); Found when fuzzing the nouveau driver with a modified Syzkaller: WARNING: drivers/gpu/drm/ttm/ttm_bo.c:256 at ttm_bo_release+0x827/0x9e0 drivers/gpu/drm/ttm/ttm_bo.c:256, CPU#1: syz.3.24/2212 Modules linked in: CPU: 1 UID: 0 PID: 2212 Comm: syz.3.24 Not tainted 7.2.0 #24 PREEMPT(lazy) nouveau 0000:01:00.0: gsp:msg fn:103 len:0x40/0x20 res:0x19 resp:0x19 Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014 RIP: 0010:ttm_bo_release+0x827/0x9e0 drivers/gpu/drm/ttm/ttm_bo.c:256 Code: 02 00 0f 85 51 01 00 00 48 8b 7b 08 e8 d2 20 01 00 e9 80 fd ff ff e8 d8 15 c0 fe 90 0f 0b 90 e9 e1 f8 ff ff e8 ca 15 c0 fe 90 <0f> 0b 90 e9 a4 f8 ff ff e8 bc 15 c0 fe be 03 00 00 00 4c 89 e7 e8 msg: 00000000: 05 00 d0 c1 04 00 f0 f1 01 30 00 00 2d 90 00 00 .........0..-... RSP: 0018:ffffc9000f5cf710 EFLAGS: 00010293 RAX: 0000000000000000 RBX: ffff888018e5d2a8 RCX: ffffffff82bb1b36 RDX: ffff888017b68000 RSI: 0000000000000004 RDI: ffff888018e5d2a8 msg: 00000010: 19 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 ................ RBP: ffff88801261c720 R08: 0000000000000001 R09: ffffed10031cba55 R10: ffff888018e5d2ab R11: 00000000000000f3 R12: ffff888018e5d290 R13: ffff888018e5d2d4 R14: ffff88801b219c18 R15: dffffc0000000000 FS: 0000000000000000(0000) GS:ffff8880e0f6f000(0000) knlGS:0000000000000000 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 CR2: 0000001b31223ffc CR3: 0000000028e00005 CR4: 0000000000770ef0 PKRU: 80000000 Call Trace: <TASK> kref_put include/linux/kref.h:65 [inline] ttm_bo_put drivers/gpu/drm/ttm/ttm_bo.c:325 [inline] ttm_bo_fini+0x55/0x80 drivers/gpu/drm/ttm/ttm_bo.c:330 nouveau_gem_object_del+0xb2/0x1b0 drivers/gpu/drm/nouveau/nouveau_gem.c:90 drm_gem_object_free+0x5f/0x90 drivers/gpu/drm/drm_gem.c:1165 kref_put include/linux/kref.h:65 [inline] __drm_gem_object_put include/drm/drm_gem.h:562 [inline] drm_gem_object_put include/drm/drm_gem.h:575 [inline] nouveau_abi16_chan_fini.constprop.0+0x44f/0x5a0 drivers/gpu/drm/nouveau/nouveau_abi16.c:195 nouveau 0000:01:00.0: syz.2.23[2209]: Unknown handle 0x00000000 nouveau_abi16_fini+0x1d0/0x340 drivers/gpu/drm/nouveau/nouveau_abi16.c:225 nouveau_drm_postclose+0x18b/0x3e0 drivers/gpu/drm/nouveau/nouveau_drm.c:1284 nouveau 0000:01:00.0: syz.2.23[2209]: validate_init drm_file_free.part.0+0x6d6/0xb60 drivers/gpu/drm/drm_file.c:267 drm_file_free drivers/gpu/drm/drm_file.c:237 [inline] drm_close_helper.isra.0+0x11a/0x160 drivers/gpu/drm/drm_file.c:290 drm_release+0x1ab/0x330 drivers/gpu/drm/drm_file.c:438 __fput+0x39c/0xa60 fs/file_table.c:512 nouveau 0000:01:00.0: syz.2.23[2209]: validate: -2 task_work_run+0x15a/0x230 kernel/task_work.c:233 exit_task_work include/linux/task_work.h:40 [inline] do_exit+0x82b/0x25a0 kernel/exit.c:1009 do_group_exit+0xc2/0x280 kernel/exit.c:1152 get_signal+0x1d6e/0x1f30 kernel/signal.c:3046 arch_do_signal_or_restart+0x7d/0x6e0 arch/x86/kernel/signal.c:337 __exit_to_user_mode_loop kernel/entry/common.c:66 [inline] exit_to_user_mode_loop+0xdf/0x440 kernel/entry/common.c:101 __exit_to_user_mode_prepare include/linux/irq-entry-common.h:207 [inline] syscall_exit_to_user_mode_prepare include/linux/irq-entry-common.h:230 [inline] syscall_exit_to_user_mode include/linux/entry-common.h:318 [inline] do_syscall_64+0x4f8/0x690 arch/x86/entry/syscall_64.c:100 entry_SYSCALL_64_after_hwframe+0x77/0x7f RIP: 0033:0x7f12bac8594d Code: Unable to access opcode bytes at 0x7f12bac85923. RSP: 002b:00007f12b96e70d8 EFLAGS: 00000246 ORIG_RAX: 00000000000000ca RAX: 0000000000000001 RBX: 00007f12baf15fa8 RCX: 00007f12bac8594d RDX: 00000000000f4240 RSI: 0000000000000081 RDI: 00007f12baf15fac RBP: 00007f12baf15fa0 R08: 00007f12baee8000 R09: 0000000000000000 R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000 R13: 00007f12baf16038 R14: 0000000000000006 R15: 00007ffe2ed394b0 </TASK> irq event stamp: 47867 hardirqs last enabled at (47883): [<ffffffff815cafc6>] __up_console_sem+0x66/0x70 kernel/printk/printk.c:347 hardirqs last disabled at (47892): [<ffffffff815cafab>] __up_console_sem+0x4b/0x70 kernel/printk/printk.c:345 softirqs last enabled at (47880): [<ffffffff81434277>] __do_softirq kernel/softirq.c:656 [inline] softirqs last enabled at (47880): [<ffffffff81434277>] invoke_softirq kernel/softirq.c:496 [inline] softirqs last enabled at (47880): [<ffffffff81434277>] __irq_exit_rcu+0x137/0x1c0 kernel/softirq.c:735 softirqs last disabled at (47875): [<ffffffff81434277>] __do_softirq kernel/softirq.c:656 [inline] softirqs last disabled at (47875): [<ffffffff81434277>] invoke_softirq kernel/softirq.c:496 [inline] softirqs last disabled at (47875): [<ffffffff81434277>] __irq_exit_rcu+0x137/0x1c0 kernel/softirq.c:735 Fix by calling ttm_bo_pin() only when the existing placement is compatible with the new pin request. This matches the correct behavior in other DRM drivers such as amdgpu_bo_pin() in amdgpu. Cc: stable@vger.kernel.org Fixes: ad76b3f7c7a0 ("drm/nouveau: teach nouveau_bo_pin() how to force a contig vram allocation") Signed-off-by: Peiyang He <peiyang_he@smail.nju.edu.cn> Assisted-by: LLM Reviewed-by: Lyude Paul <lyude@redhat.com> Signed-off-by: Lyude Paul <lyude@redhat.com> Link: https://patch.msgid.link/EACEF2F4E098413F+20260918025312.2814889-1-peiyang_he@smail.nju.edu.cn
12 daysdrm/nouveau/clk: don't clobber reclock status when restoring volt/fanFrancesco Magazzu
nvkm_cstate_prog() reuses 'ret' for the voltage and fan-speed restore calls it makes after reprogramming the clocks. Those calls almost always succeed, so the status of the reclock itself is overwritten and the function reports success even when clk->func->calc() or clk->func->prog() failed. The converse is also true: a successful reclock is reported as an error if the final restore call fails, even though that failure is only logged and otherwise ignored. The only consumer of the return value is the error message in nvkm_pstate_work(), so in practice a failing reclock is simply never reported. Nothing else changes, but a function that returns success on failure is a trap for the next caller. Keep the calc/prog status in 'ret' and use a separate local for the restore calls. Fixes: 3eca809b3c05 ("drm/nouveau/clk: cosmetic changes") Signed-off-by: Francesco Magazzu <postadelmaga@gmail.com> Reviewed-by: Lyude Paul <lyude@redhat.com> Signed-off-by: Lyude Paul <lyude@redhat.com> Link: https://patch.msgid.link/20260918131620.405133-5-postadelmaga@gmail.com
12 daysdrm/nouveau/device: don't use the pstate cursor after the loopFrancesco Magazzu
nvkm_control_mthd_pstate_attr() looks up the pstate at the index supplied by userspace by walking clk->states, and then keeps using the list_for_each_entry cursor after the loop. This is not triggerable today: the function already rejects args->v0.state >= clk->state_nr before the loop, and clk->state_nr is kept in sync with the number of entries on clk->states, so the lookup always breaks on a real entry. Should the loop ever run to completion, the cursor would point at the list head rather than at a pstate, and the pstate->base.domain[] read and the walk of pstate->list that follow would read past it. Rather than leave that trap in place, track whether the entry was found and return -EINVAL if it was not, like the other lookup failures in this function. No functional change. Signed-off-by: Francesco Magazzu <postadelmaga@gmail.com> Reviewed-by: Lyude Paul <lyude@redhat.com> Signed-off-by: Lyude Paul <lyude@redhat.com> Link: https://patch.msgid.link/20260918131620.405133-4-postadelmaga@gmail.com
12 daysdrm/nouveau/clk: don't use the pstate cursor after the loopFrancesco Magazzu
nvkm_pstate_prog() walks clk->states looking for the entry at index 'pstatei' and then keeps using the list_for_each_entry cursor after the loop. This is not triggerable today: every caller clamps the index against clk->state_nr before calling, so the loop always breaks on a real entry. It is safe by virtue of what the callers happen to do, not by anything the function itself checks. Should a caller ever pass an index that is not on the list, the cursor would point at the list head rather than at a pstate, and the pstate->base.domain[] and pstate->fanspeed accesses that follow would read past it. Rather than leave that trap in place for the next caller, track whether the entry was found and return -EINVAL if it was not. No functional change. Signed-off-by: Francesco Magazzu <postadelmaga@gmail.com> Reviewed-by: Lyude Paul <lyude@redhat.com> Signed-off-by: Lyude Paul <lyude@redhat.com> Link: https://patch.msgid.link/20260918131620.405133-3-postadelmaga@gmail.com
12 daysdrm/nouveau/clk: fix list cursor use after loop in nvkm_clk_ustate_updateDan Carpenter
If list_for_each_entry() exits without hitting a break then "pstate" is not a valid pstate pointer. Introduce a "found" variable instead. The check is reachable from userspace: nvkm_clk_ustate_update() takes the pstate id straight from the 'pstate' debugfs file, so requesting an id that is not in clk->states - or any id at all when the perf tables are broken and the list is empty - makes the pstate->pstate != req test dereference the list head cast to a struct nvkm_pstate, which is an out-of-bounds read. Fixes: 7c8565220697 ("drm/nouveau/clk: implement power state and engine clock control in core") Signed-off-by: Dan Carpenter <dan.carpenter@oracle.com> [Francesco: rebased on drm-misc-next, expanded the commit message] Signed-off-by: Francesco Magazzu <postadelmaga@gmail.com> Reviewed-by: Lyude Paul <lyude@redhat.com> Signed-off-by: Lyude Paul <lyude@redhat.com> Link: https://patch.msgid.link/20260918131620.405133-2-postadelmaga@gmail.com
13 daysMerge tag 'amd-drm-fixes-7.3-2026-09-17' of ↵Dave Airlie
https://gitlab.freedesktop.org/drm/amdgpu/kernel into drm-fixes amd-drm-fixes-7.3-2026-09-17: amdgpu: - SMU 14.x fix - DC IRQ fix - Runtime PM fix for P2P - RAS fix - PCIe reporting fix - DCN 6 fix - Device removal fix - DC MALL fix amdkfd: - GC 12.x fixes - Boundary checks - Mapping clear fix Signed-off-by: Dave Airlie <airlied@redhat.com> From: Alex Deucher <alexander.deucher@amd.com> Link: https://patch.msgid.link/20260917201213.3880863-1-alexander.deucher@amd.com