summaryrefslogtreecommitdiff
path: root/drivers/infiniband
AgeCommit message (Collapse)Author
11 hoursMerge branch 'main' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next.git # Conflicts: # drivers/net/ethernet/realtek/r8169_main.c # net/mac80211/ieee80211_i.h # net/mac80211/tx.c
12 hoursMerge branch 'for-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/rdma/rdma.git # Conflicts: # drivers/infiniband/sw/rxe/rxe_verbs.c
12 hoursMerge branch 'for-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/mm/linux.git # Conflicts: # arch/arm64/kvm/mmu.c
25 hoursfault-inject: fix dentry leakMichael Liang
fault_create_debugfs_attr() has always taken an extra dentry reference on the created directory (attr->dname = dget(dir)) so that fail_dump() could print the name via %pd from any context. Nothing anywhere in the tree ever calls dput() on attr->dname. For callers with a matching teardown, that unmatched reference causes one dentry plus its attached inode to leak per fault_create_debugfs_attr / debugfs_remove_recursive cycle. simple_recursive_removal() drops debugfs's own +1 ref on the child dentry, but the dget()'d ref keeps its refcount at 1: the dentry ends up unhashed but pinned, and its inode is never freed. Boot-once callers (mm/failslab, block/blk-core, etc.) leak exactly once at init and never destroy the tree, so the impact there is bounded. But per-lifecycle callers (drivers/nvme, drivers/infiniband/hw/hfi1, drivers/mmc, drivers/iommu/iommufd, drivers/media, drivers/misc, drivers/gpu/drm/msm, drivers/crypto, net/sunrpc) leak on every create/destroy cycle. We observed this in production: an NVMe/RDMA host repeatedly reconnecting to a target that rejected the CRTO Property Get went through ~50 nvme controller create/destroy cycles per second, and dentry and inode_cache grew by ~13k pinned objects per 240 s -- unrecoverable through drop_caches. Byte math matched a per-cycle 1-dentry / 1-inode leak from the "fault_inject" directory dentry. Fix this by not holding any external reference in fault_attr. Embed the directory name as a fixed-size char array (FAULT_ATTR_DNAME_LEN, 64 bytes) inside struct fault_attr, copied by strscpy() at fault_create_debugfs_attr() time. fail_dump() prints it via %s. Advantages of an embedded array over kstrdup() + kfree() paired with a new destroy API: - Zero API footprint. No new export and no caller changes required: callers already own their fault_attr's memory and free it when they are done, and now that suffices. - No allocation on the create path. - fault_create_debugfs_attr() cannot fail from the name-copy step. - No lifetime coupling between attr->dname and debugfs; the string is valid for exactly as long as the containing struct. The 64-byte length accommodates every in-tree caller with generous headroom (the longest current name is "fail_dma_array_full", 19 chars). The user-visible fail_dump() format changes from "name %pd" to "name %s", but the printed content is identical -- %pd on the created directory renders the same string that was passed in as @name. drivers/infiniband/hw/hfi1/fault.c drops a now-invalid "attr.dname = NULL" statement; the surrounding kzalloc() already zero-initialises the array. Link: https://lore.kernel.org/20260821181527.3271414-1-mliang@purestorage.com Fixes: 6adc4a22f20b ("fault-inject: add ratelimit option") Signed-off-by: Michael Liang <mliang@purestorage.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Reviewed-by: Andrew Morton <akpm@linux-foundation.org> Cc: Akinbou Mita <akinobu.mita@gmail.com> Cc: Dennis Dalessandro <dennis.dalessandro@cornelisnetworks.com> Cc: Jason Gunthorpe <jgg@ziepe.ca> Cc: Leon Romanovsky <leon@kernel.org> Cc: Vlastimil Babka <vbabka@kernel.org> Cc: <stable@vger.kernel.org>
25 hoursinfiniband: update hfi1 to use remap_vmalloc_range()Lorenzo Stoakes (ARM)
In cases which map chip memory from vmalloc()'d ranges, the hfi1 infiniband drivers currently installs a fault handler, and then smuggles the kernel virtual address of this range in vma->vm_pgoff. This is exposing KASLR-sensitive internal kernel state in the VMA, and is entirely unnecessary. Instead, use remap_vmalloc_range() to remap the VMA to the span, and eliminate the fault handler altogether. remap_vmalloc_range() checks that the VMA does not extend beyond the vmalloc area, and the driver already requires the VMA to exactly match the span of the memory being mapped, so this has no impact. The memory is all preallocated so not having a fault handler has no impact either, other than pre-mapping the ranges which is beneficial. We also remove the VM_IO flag as it's not appropriate here, and the VM_DONTEXPAND flag as remap_vmalloc_range() will set it (and also mark the range correctly as a mixed map). We also update the vmalloc paths to place the virtual kernel address in memvirt, rather than overloading the physical address memaddr. We predicate the vmalloc handling on the vmalloc flag before we check memvirt for the virtual address-derived PFN remap path, so this works fine. remap_vmalloc_range() requires that the vmalloc()'d areas were all allocated using vmalloc_user() - each of cq->comps, uctxt->subctxt_rcvegrbuf, uctxt->subctxt_rcvhdr_base, uctxt->subctxt_uregbase and dd->events were allocated this way, so that requirement is satisfied. We also remove VM_IO and VM_DONTEXPAND from the STATUS command, as these are both set on remap. Finally, we remove VM_DONTEXPAND from the PIO_BUFS, PIO_BUFS_SOP and UREGS commands, as these are also all set on remap. PIO_CRED retains it, as dma_mmap_coherent() may map via vm_insert_page() on the IOMMU-DMA path, which sets only VM_MIXEDMAP. The RCV_HDRQ, RCV_EGRBUF and RTAIL commands also map via dma_mmap_coherent() and never set VM_DONTEXPAND, so set it for them for the same reason. Note that we retain expected behaviour throughout - the vmalloc remapped ranges set VM_MIXEDMAP | VM_DONTDUMP | VM_DONTEXPAND for each range. VM_IO was never appropriate as the ranges are explicitly not MMIO, and the reference to the v3.7 VM_RESERVED semantics map on to VM_MIXEDMAP | VM_DONTDUMP | VM_DONTEXPAND correctly - no core dump, unmergeable, no normal vm page for purposes of reclaim/migration/etc. There is a change in behaviour in that pages mapped using remap_vmalloc_range() will now have normal GUP-able pages, however this should have no impact as there is no reason not to allow this. Link: https://lore.kernel.org/20260917-b4-mmap-prepare-vma-flag-sanify-v3-11-4583d8a23bca@kernel.org Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Cc: Albert Ou <aou@eecs.berkeley.edu> Cc: Alexander Gordeev <agordeev@linux.ibm.com> Cc: Alexei Starovoitov <ast@kernel.org> Cc: Alistair Popple <apopple@nvidia.com> Cc: Al Viro <viro@zeniv.linux.org.uk> Cc: Andreas Larsson <andreas@gaisler.com> Cc: Andrii Nakryiko <andrii@kernel.org> Cc: "Aneesh Kumar K.V" <aneesh.kumar@kernel.org> Cc: Anup Patel <anup@brainfault.org> Cc: Arnaldo Carvalho de Melo <acme@kernel.org> Cc: Arnd Bergmann <arnd@arndb.de> Cc: Axel Rasmussen <axelrasmussen@google.com> Cc: Baolin Wang <baolin.wang@linux.alibaba.com> Cc: Baoquan He <baoquan.he@linux.dev> Cc: Barry Song <baohua@kernel.org> Cc: "Borislav Petkov (AMD)" <bp@alien8.de> Cc: Byungchul Park <byungchul@sk.com> Cc: Catalin Marinas <catalin.marinas@arm.com> Cc: Chengming Zhou <chengming.zhou@linux.dev> Cc: Chris Li <chrisl@kernel.org> Cc: Christian Borntraeger <borntraeger@linux.ibm.com> Cc: Christian Brauner <brauner@kernel.org> Cc: Claudio Imbrenda <imbrenda@linux.ibm.com> Cc: Dave Airlie <airlied@gmail.com> Cc: Dave Hansen <dave.hansen@linux.intel.com> Cc: David Hildenbrand <david@kernel.org> Cc: David S. Miller <davem@davemloft.net> Cc: Dennis Dalessandro <dennis.dalessandro@cornelisnetworks.com> Cc: Dev Jain <dev.jain@arm.com> Cc: Doug Gilbert <dgilbert@interlog.com> Cc: Eduard Zingerman <eddyz87@gmail.com> Cc: Emil Tsalapatis <emil@etsalapatis.com> Cc: Gerald Schaefer <gerald.schaefer@linux.ibm.com> Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org> Cc: Gregory Price <gourry@gourry.net> Cc: Harry Yoo <harry@kernel.org> Cc: Heiko Carstens <hca@linux.ibm.com> Cc: Helge Deller <deller@gmx.de> Cc: "Huang, Ying" <ying.huang@linux.alibaba.com> Cc: Ingo Molnar <mingo@redhat.com> Cc: James Bottomley <james.bottomley@HansenPartnership.com> Cc: Jan Kara <jack@suse.cz> Cc: Jann Horn <jannh@google.com> Cc: Janosch Frank <frankja@linux.ibm.com> Cc: Jaroslav Kysela <perex@perex.cz> Cc: Jason Gunthorpe <jgg@ziepe.ca> Cc: Jaya Kumar <jayalk@intworks.biz> Cc: Johannes Weiner <hannes@cmpxchg.org> Cc: John Hubbard <jhubbard@nvidia.com> Cc: Jonathan Corbet <corbet@lwn.net> Cc: Joshua Hahn <joshua.hahnjy@gmail.com> Cc: Juri Lelli <juri.lelli@redhat.com> Cc: Kairui Song <kasong@tencent.com> Cc: Kemeng Shi <shikemeng@huaweicloud.com> Cc: Kiryl Shutsemau <kas@kernel.org> Cc: Kumar Kartikeya Dwivedi <memxor@gmail.com> Cc: Lance Yang <lance.yang@linux.dev> Cc: Leon Romanovsky <leon@kernel.org> Cc: Liam R. Howlett <liam@infradead.org> Cc: Maarten Lankhorst <maarten.lankhorst@linux.intel.com> Cc: Madhavan Srinivasan <maddy@linux.ibm.com> Cc: Marc Rutland <mark.rutland@arm.com> Cc: Marc Zyngier <maz@kernel.org> Cc: "Masami Hiramatsu (Google)" <mhiramat@kernel.org> Cc: Matthew Brost <matthew.brost@intel.com> Cc: Matthew Wilcox (Oracle) <willy@infradead.org> Cc: Maxime Ripard <mripard@kernel.org> Cc: Michal Hocko <mhocko@kernel.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Mike Rapoport <rppt@kernel.org> Cc: Miklos Szeredi <miklos@szeredi.hu> Cc: Muchun Song <muchun.song@linux.dev> Cc: Namhyung kim <namhyung@kernel.org> Cc: Nhat Pham <nphamcs@gmail.com> Cc: Nicholas Piggin <npiggin@gmail.com> Cc: Oleg Nesterov <oleg@redhat.com> Cc: Oscar Salvador <osalvador@suse.de> Cc: Palmer Dabbelt <palmer@dabbelt.com> Cc: Paul Moore <paul@paul-moore.com> Cc: Pedro Falcato <pfalcato@suse.de> Cc: Peter Xu <peterx@redhat.com> Cc: Peter Zijlstra <peterz@infradead.org> Cc: Rakie Kim <rakie.kim@sk.com> Cc: Rik van Riel <riel@surriel.com> Cc: Ryan Roberts <ryan.roberts@arm.com> Cc: Sebastian Reichel <sre@kernel.org> Cc: Shakeel Butt <shakeel.butt@linux.dev> Cc: Stephen Smalley <stephen.smalley.work@gmail.com> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Takashi Iwai (SUSE) <tiwai@suse.de> Cc: Takashi Iwai <tiwai@suse.com> Cc: Thomas Zimemrmann <tzimmermann@suse.de> Cc: Vasily Gorbik <gor@linux.ibm.com> Cc: Vincent Guittot <vincent.guittot@linaro.org> Cc: Vlastimil Babka <vbabka@kernel.org> Cc: Wei Xu <weixugc@google.com> Cc: Will Deacon <will@kernel.org> Cc: Yuanchu Xie <yuanchu@google.com> Cc: Zi Yan <ziy@nvidia.com>
2 daysRDMA/ionic: Fix double removal of CMB mmap entries in create QP error pathAbhijit Gangurde
The error unwind in ionic_create_qp() removes the SQ and RQ CMB mmap entries explicitly, and then falls through to ionic_qp_rq_destroy() and ionic_qp_sq_destroy(), whose ionic_qp_{rq,sq}_destroy_cmb() helpers remove the very same entries a second time. Drop the explicit removals and leave ionic_qp_{rq,sq}_destroy_cmb() as the single owner on the ionic_destroy_qp() path. Fixes: e8521822c733 ("RDMA/ionic: Register device ops for control path") Signed-off-by: Abhijit Gangurde <abhijit.gangurde@amd.com> Link: https://patch.msgid.link/20260925093638.3032149-2-abhijit.gangurde@amd.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2 daysRDMA/ionic: Fix XArray initialization to use XArray flagsAbhijit Gangurde
xa_init_flags() takes XArray flags (XA_FLAGS_*), not gfp flags, even though the parameter is typed gfp_t. Passing GFP_ATOMIC leaves both lock-type bits clear, so xa_lock_type() resolves to XA_LOCK_NORMAL for qp_tbl and cq_tbl. Both tables are used with the IRQ-safe accessors: xa_store_irq() with GFP_KERNEL from process context, and xa_lock_irqsave() from the event queue path. When a store needs to grow the tree, __xas_nomem() drops the lock using the recorded lock type before allocating: if (gfpflags_allow_blocking(gfp)) { xas_unlock_type(xas, lock_type); xas->xa_alloc = kmem_cache_alloc_lru(...); xas_lock_type(xas, lock_type); } With XA_LOCK_NORMAL that is a plain spin_unlock(), so the sleeping GFP_KERNEL allocation runs with interrupts still disabled by xa_store_irq(), and lockdep annotates the lock with the wrong class. GFP_ATOMIC also happens to set __GFP_HIGH (0x20), which collides with XA_FLAGS_ACCOUNT (32U) and silently enables memcg accounting of the XArray nodes. Use XA_FLAGS_LOCK_IRQ so the lock type matches how the tables are actually accessed. Fixes: e8521822c733 ("RDMA/ionic: Register device ops for control path") Signed-off-by: Abhijit Gangurde <abhijit.gangurde@amd.com> Link: https://patch.msgid.link/20260925093638.3032149-1-abhijit.gangurde@amd.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2 daysRDMA/siw: drop duplicate check_app_limited in siw_tcp_sendpagesGeliang Tang
After do_tcp_sendpages() was converted to use MSG_SPLICE_PAGES and later inlined as direct tcp_sendmsg_locked() calls, callers that previously needed an explicit tcp_rate_check_app_limited(sk) to cover do_tcp_sendpages() no longer need it. tcp_sendmsg_locked() performs the check on every path that queues data, making the outer call redundant. The site changed here, siw_tcp_sendpages() in siw_qp_tx.c, holds the socket lock and invokes tcp_sendmsg_locked() on every iteration. The early-return paths in tcp_sendmsg_locked() that skip tcp_rate_check_app_limited() - the MSG_ZEROCOPY allocation failure and MSG_FASTOPEN branches - return without queueing any MSG_SPLICE_PAGES data, so there is no functional consequence from omitting the outer check. Signed-off-by: Geliang Tang <tanggeliang@kylinos.cn> Link: https://patch.msgid.link/32ef40cc81486adb8f2d43585224ab44d7e77f73.1790234916.git.tanggeliang@kylinos.cn Acked-by: Bernard Metzler <bernard.metzler@linux.dev> Signed-off-by: Leon Romanovsky <leon@kernel.org>
2 daysRDMA/rxe: Reject prefetch of a non-ODP MRNorbert Szetei
rxe_ib_advise_mr_prefetch() and rxe_ib_prefetch_sg_list() look up the MR by lkey and hand it to rxe_odp_do_pagefault_and_lock() without checking that it is an ODP MR. That path runs to_ib_umem_odp() on mr->umem, and for a non-ODP MR mr->umem is a plain struct ib_umem from ib_umem_get(), so the container_of() in to_ib_umem_odp() lands past the end of the object and ib_umem_odp_map_dma_and_lock() reads its ib_umem_odp fields out of bounds. lookup_mr() validates the lkey, PD, access and state but not the MR type, and IB_UVERBS_ADVISE_MR_ADVICE_PREFETCH is accepted for any MR. BUG: KASAN: slab-out-of-bounds in ib_umem_odp_map_dma_and_lock+0x884/0x8a0 Read of size 8 at addr ffff88810a3ebcf0 by task advi/921 ib_umem_odp_map_dma_and_lock+0x884/0x8a0 rxe_ib_advise_mr+0x543/0xad0 ib_uverbs_handler_UVERBS_METHOD_ADVISE_MR+0x446/0x530 ib_uverbs_cmd_verbs+0x2b3c/0x3b20 ib_uverbs_ioctl+0x1e3/0x310 Allocated by task 921: __ib_umem_get_va+0x13e/0xae0 rxe_mr_init_user+0x2ae/0xb00 rxe_reg_user_mr+0x337/0x510 The buggy address belongs to the object at ffff88810a3ebc80 which belongs to the cache kmalloc-96 of size 96 Ask lookup_mr() for IB_ACCESS_ON_DEMAND in both the synchronous and the asynchronous prefetch arm. mr->access carries that flag only for an MR registered as ODP, so the existing (access & mr->access) != access test rejects a plain MR and the prefetch fails with -EINVAL. Fixes: 3576b0df1588 ("RDMA/rxe: Implement synchronous prefetch for ODP MRs") Signed-off-by: Norbert Szetei <norbert@doyensec.com> Link: https://patch.msgid.link/20260923-rxe-advise-mr-v3-v3-2-95e1a4077e4d@doyensec.com Reviewed-by: Zhu Yanjun <yanjun.zhu@linux.dev> Signed-off-by: Leon Romanovsky <leon@kernel.org>
2 daysRDMA/rxe: Reject IB_ACCESS_ON_DEMAND changes after MR creationNorbert Szetei
Whether an MR is an ODP MR is decided once, at registration time: rxe_reg_user_mr() picks rxe_odp_mr_init_user() over rxe_mr_init_user() based on IB_ACCESS_ON_DEMAND, and only the former builds an ib_umem_odp via ib_umem_odp_get(). The umem cannot change type afterwards, and is_odp_mr() reads mr->umem->is_odp. Two paths assign mr->access after that point and can leave it describing an MR type the umem does not have: rxe_rereg_user_mr() with IB_MR_REREG_ACCESS overwrites mr->access with the caller's value, and IB_ACCESS_ON_DEMAND is part of RXE_ACCESS_SUPPORTED_MR, so userspace can set the flag on a plain MR or clear it on an ODP MR while the umem stays what it was. rxe_reg_fast_mr() takes mr->access from the REG_MR work request unmasked and moves the MR to RXE_MR_STATE_VALID, on an MR that rxe_mr_init_fast() left with a NULL umem. Reject IB_ACCESS_ON_DEMAND in both, so mr->access carries the flag only for an MR that has an ODP umem and the flag can be used to identify one. Fixes: 544c7f62cf32 ("RDMA/rxe: Implement rereg_user_mr") Signed-off-by: Norbert Szetei <norbert@doyensec.com> Link: https://patch.msgid.link/20260923-rxe-advise-mr-v3-v3-1-95e1a4077e4d@doyensec.com Reviewed-by: Zhu Yanjun <yanjun.zhu@linux.dev> Signed-off-by: Leon Romanovsky <leon@kernel.org>
2 daysRDMA/rxe: unpublish the per-net tunnel socket beforeBinbin Deng
KASAN reports a slab-use-after-free in ip6_route_output_flags() reached from rxe_find_route(), with the free in __sk_destruct() after rxe_sock_put(), and a user with CAP_NET_ADMIN can remove the device with "rdma link del rxe0" while RoCE v2 over IPv6 traffic keeps looking the socket up. BUG: KASAN: slab-use-after-free in ip6_route_output_flags+0x300/0x360 Read of size 4 at addr ffff888115ddd794 by task kworker/u32:6/309 Call Trace: <TASK> dump_stack_lvl+0x53/0x70 print_report+0xd0/0x630 ? __pfx__raw_spin_lock_irqsave+0x10/0x10 ? ip6_route_output_flags+0x300/0x360 kasan_report+0xce/0x100 ? ip6_route_output_flags+0x300/0x360 ip6_route_output_flags+0x300/0x360 ip6_dst_lookup_tail.constprop.0+0x76c/0xcc0 ? ct_nmi_exit+0xc3/0xf0 ip6_dst_lookup_flow+0xf5/0x1e0 ? __pfx_ip6_dst_lookup_flow+0x10/0x10 rxe_find_route+0x426/0xa30 ? __kasan_slab_alloc+0x6e/0x70 ? __pfx_rxe_find_route+0x10/0x10 ? kmem_cache_alloc_node_noprof+0x141/0x370 ? kmalloc_reserve+0x103/0x2b0 ? rxe_icrc_generate+0x229/0x330 ? __pfx___alloc_skb+0x10/0x10 rxe_prepare+0x9e8/0x18b0 ? rxe_init_packet+0x3c7/0x4f0 rxe_requester+0x1a0f/0x51f0 ? rxe_completer+0x1de9/0x38c0 ? __pfx_rxe_completer+0x10/0x10 ? __queue_work+0x43e/0x11f0 ? __pfx_rxe_requester+0x10/0x10 ? irqentry_exit+0xd2/0x640 ? _raw_spin_lock_irqsave+0x85/0xe0 ? __pfx__raw_spin_lock_irqsave+0x10/0x10 ? __pfx_rxe_sender+0x10/0x10 rxe_sender+0xe/0x30 do_work+0x144/0x470 process_one_work+0x633/0x1030 ? assign_work+0x11d/0x370 worker_thread+0x45b/0xd10 ? __pfx_worker_thread+0x10/0x10 kthread+0x2c6/0x3b0 ? recalc_sigpending+0x15c/0x1e0 ? __pfx_kthread+0x10/0x10 ret_from_fork+0x36e/0x5a0 ? __pfx_ret_from_fork+0x10/0x10 ? __switch_to+0x572/0xdd0 ? __pfx_kthread+0x10/0x10 ret_from_fork_asm+0x1a/0x30 </TASK> Allocated by task 146020: kasan_save_stack+0x33/0x60 kasan_save_track+0x14/0x30 __kasan_slab_alloc+0x6e/0x70 kmem_cache_alloc_noprof+0x130/0x360 sk_prot_alloc+0x56/0x210 Fix by clearing the per-net pointer before the last reference is dropped. Fixes: f1327abd6abed ("RDMA/rxe: Support RDMA link creation and destruction per net namespace") Signed-off-by: Binbin Deng <18983559317@163.com> Link: https://patch.msgid.link/20260923051800.244740-1-18983559317@163.com Reviewed-by: Zhu Yanjun <yanjun.zhu@linux.dev> Signed-off-by: Leon Romanovsky <leon@kernel.org>
2 daysRDMA/rtrs-clt: Reject invalid queue_depth from the peerQuanye Yang
queue_depth is taken from the server's CM private_data and later used as a session invariant. A peer can send 0, which the client stores and then trips WARN_ON() in create_con_cq_qp() when the next connection is set up. Reject a zero queue_depth on ESTABLISHED, same as a bad magic or version, and fail the handshake with -ECONNRESET. Fixes: 6a98d71daea1 ("RDMA/rtrs: client: main functionality") Reported-by: Farhad Alemi <farhad.alemi@berkeley.edu> Closes: https://lore.kernel.org/linux-rdma/CA+0ovCg_sG8gaZXXj9zmOrpxRt=8+-x9wvycVNUKwWmMbr6-Yg@mail.gmail.com Signed-off-by: Quanye Yang <quanyeyang@proton.me> Link: https://patch.msgid.link/20260922-rtrs-fix-warning-rtrs-clt-rdma-cm-handler-v1-1-2076c0ee1455@proton.me Acked-by: Jack Wang <jinpu.wang@cloud.ionos.com> Signed-off-by: Leon Romanovsky <leon@kernel.org>
2 daysRDMA/nldev: Fix NULL deref in dumps of objects abandoned by DRIVER_FAILUREYili Zhang
fill_res_cq_entry() dereferences cq->uobject->uevent.uobject.context->res.id unconditionally for user resources. This is normally safe by ordering: destroy_hw() removes the resource from the restrack before uverbs_destroy_uobject() clears ->context, and the XA_ZERO_ENTRY marker set by rdma_restrack_begin_del() hides the entry from concurrent netlink dumps. The ordering breaks when a driver keeps failing to destroy an object during ucontext teardown. uverbs_destroy_ufile_hw() then falls back to __uverbs_cleanup_ufile(RDMA_REMOVE_DRIVER_FAILURE), which abandons the object in place: the HW object and its restrack entry are intentionally leaked, while the uobject bookkeeping is torn down and ->context is explicitly set to NULL by uverbs_destroy_uobject(). rdma_restrack_del() is never reached, so the orphaned entry stays in the device restrack with a valid kref, reachable by any subsequent netlink dump. Observed on 6.1.52 with MLNX_OFED 24.10, after an mlx5 FW failure left a process unable to tear down its CQ (destroy_cq FW command failing during FD close): WARNING: ... uverbs_destroy_ufile_hw+0xe3/0x100 BUG: kernel NULL pointer dereference, address: 0000000000000058 RIP: 0010:fill_res_cq_entry+0x15e/0x180 [ib_core] res_get_common_dumpit+0x304/0x530 [ib_core] nldev_res_get_cq_dumpit+0x1a/0x20 [ib_core] The faulting chain maps to the source (CR2 = 0x58): cq->uobject (struct ib_cq +0x08, res at +0x98) uobject->context (struct ib_uobject +0x10, NULL after abandon) context->res.id (struct ib_ucontext +0x58) The recent restrack rework (8d186210677c and its series) moved the restrack deletion to the start of the destroy flow and thus fences concurrent dumps from an object being destroyed, but it does not cover this case: when destroy fails, rdma_restrack_abort_del() restores the entry, and the RDMA_REMOVE_DRIVER_FAILURE sweep still never removes it from the restrack. From the fallback until the device is unregistered, any "rdma res show cq" deterministically takes the NULL pointer dereference; this is a long-lived state, NOT A RACE. The dump path holds neither the restrack lock (dropped before the fill callback runs) nor ufile->hw_destroy_rwsem, and rdma_restrack_get() only guarantees that the res memory stays alive, not that ->context is still valid. Fix the dump side: check the uobject context directly instead of the rdma_is_kernel_res() marker. A non-NULL uobject already implies a user resource, so kernel resources keep being skipped, and the additional ->context check skips the RES_CTXN attribute for orphaned entries instead of crashing the dump. The orphaned resource itself remains visible in "rdma res show" (cqn/cqe/usecnt/pid), which is what an operator needs after the accompanying uverbs WARN to diagnose the driver destroy failure. fill_res_pd_entry() has the same pattern (pd->uobject->context->res.id) and is fixed the same way; a PD can even reach the fallback without its own driver callback failing, e.g. uverbs_free_pd() returns -EBUSY while another object that failed to destroy still holds the PD usecnt. Fixes: c3d02788b45a ("RDMA/nldev: Provide parent IDs for PD, MR and QP objects") Signed-off-by: Yili Zhang <zhangyili01@baidu.com> Link: https://patch.msgid.link/20260923131239.6521-1-zhangyili01@baidu.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2 daysRDMA/efa: Pass relaxed ordering flag to deviceDana Malachi
Accept the relaxed ordering access flag during memory region registration and forward it to the device firmware via the admin command path. Reviewed-by: Chen Brasch <cbrasch@amazon.com> Reviewed-by: Yonatan Nachum <ynachum@amazon.com> Signed-off-by: Dana Malachi <danamala@amazon.com> Signed-off-by: Michael Margolin <mrgolin@amazon.com> Link: https://patch.msgid.link/20260922160118.990-3-mrgolin@amazon.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2 daysRDMA/efa: Use device ABI MR permissions instead of verbs flagsDana Malachi
Set MR permissions explicitly using device interface definitions rather than copying raw verbs access flags. This is needed for the next commit, to allow access flag bits that are not in the permissions field. Reviewed-by: Chen Brasch <cbrasch@amazon.com> Reviewed-by: Yonatan Nachum <ynachum@amazon.com> Signed-off-by: Dana Malachi <danamala@amazon.com> Signed-off-by: Michael Margolin <mrgolin@amazon.com> Link: https://patch.msgid.link/20260922160118.990-2-mrgolin@amazon.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2 daysIB/mad: fix UAF and double-free in RMPP receive reassemblyDairui Zhang
mad_rmpp_recv.state is protected by two locks that do not exclude each other: recv_timeout_handler() (port workqueue) uses agent->lock, continue_rmpp() (CQ softirq) uses rmpp_recv->lock. Nothing serializes the check-then-set sequence on the state word: T1 (CQ softirq) T2 (port workqueue) continue_rmpp() lock(rmpp_recv->lock) state == ACTIVE [passes] recv_timeout_handler() lock(agent->lock) state = TIMEOUT list_del(&rmpp_recv->list) unlock(agent->lock) destroy_rmpp_recv() wait_for_completion() [blocks on T1's ref] state = COMPLETE unlock(rmpp_recv->lock) complete_rmpp() queue cleanup_work (+10s) deref() [unblocks] kfree(rmpp_recv) ib_free_recv_mad(done_wc) +10s: recv_cleanup_handler() runs lock/list_del/destroy on the freed rmpp_recv -> UAF write and double free; nothing can cancel that cleanup_work anymore, the entry is already off agent->rmpp_list. done_wc is also returned to ib_mad_complete_recv() and freed again -> UAF read and double free. Reaching this requires sending MADs to a kernel RMPP agent from the InfiniBand/RoCE fabric: the per-port sa_query agent accepts subnet manager responses whose SLID/GID/TID are not authenticated on the fabric, and legacy user_mad agents registered without IB_USER_MAD_USER_RMPP start reassembly on any incoming RMPP DATA MAD with an attacker-chosen TID, no in-flight send required. The attacker controls segment timing and can run many transactions in parallel (agent->rmpp_list has no cap), so landing the microseconds-wide receive critical section on the 40-second timeout boundary is a matter of patience, not luck. Unify state transitions under rmpp_recv->lock in recv_timeout_handler(), keeping list manipulation under agent->lock (which preserves the agent->lock -> rmpp_recv->lock ordering used by continue_rmpp()). State check/set in both paths is then serialized, so the loser of the race always observes the final state and bails out. recv_cleanup_handler() needs no change: with transitions serialized, a cleanup_work can only be pending after COMPLETE, in which case the timeout handler exits early. The race is as old as the RMPP implementation itself. Fixes: fa619a77046b ("[PATCH] IB: Add RMPP implementation") Assisted-by: LLM Signed-off-by: Dairui Zhang <zhangdairui@gmail.com> Link: https://patch.msgid.link/20260922152342.1474845-1-zhangdairui@gmail.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
3 daysRDMA/bnxt_re: Use bnxt_ext_stats_supported for counter count selectionSelvin Xavier
Selecting between BNXT_RE_NUM_EXT_COUNTERS and BNXT_RE_NUM_STD_COUNTERS relied only on bnxt_qplib_is_chip_gen_p5_p7(), which doesn't account for the extended stats capability flag or the PF/VF restriction. Use bnxt_ext_stats_supported() instead, matching the check already used to populate the extended stats, so the counter count stays consistent with what gets filled in. Fixes: 8238c7bd8420 ("RDMA/bnxt_re: Fix the statistics for Gen P7 VF") Reviewed-by: Kalesh AP <kalesh-anakkur.purayil@broadcom.com> Signed-off-by: Selvin Xavier <selvin.xavier@broadcom.com> Signed-off-by: Leon Romanovsky <leon@kernel.org>
3 daysRDMA/bnxt_re: Fix the PD and DPI table sizeSelvin Xavier
The PD and DPI bitmaps were sized as max >> 3 (bytes), but bitmap ops (set_bit(), clear_bit(), find_first_bit(), test_and_set_bit()) operate on whole unsigned long words, so whenever max isn't a multiple of BITS_PER_LONG, the buffer under-allocates and the top word's bitops read/write past the end of the kmalloc()'d buffer. Most exposed on the DPI table, since dpit->max comes from the firmware-reported dev_attr->max_dpi with no alignment guarantee. Fix the size of both allocations with BITS_TO_LONGS(max) * sizeof(unsigned long). Fixes: 1ac5a4047975 ("RDMA/bnxt_re: Add bnxt_re RoCE driver") Signed-off-by: Selvin Xavier <selvin.xavier@broadcom.com> Signed-off-by: Leon Romanovsky <leon@kernel.org>
3 daysRDMA/bnxt_re: Serialize dcb_wq access against async notifierSelvin Xavier
bnxt_re_async_notifier() runs in NAPI/softirq context via bnxt_en's RCU-protected ULP ops dispatch, never under rtnl_lock, so it isn't serialized against bnxt_re_uninit_dcb_wq() destroying rdev->dcb_wq. A concurrent notifier call can queue_work() on a workqueue that is being, or has just been, destroyed. Add a spinlock scoped to dcb_wq: the notifier takes it before checking dcb_wq and queuing work, and bnxt_re_uninit_dcb_wq() takes it to atomically clear dcb_wq before destroying it. Initialize the lock at rdev allocation so it is valid on every teardown path, including bnxt_re_dev_init()'s early failure labels. Fixes: 51dc5312dcd9 ("RDMA/bnxt_re: Add support to handle DCB_CONFIG_CHANGE event") Signed-off-by: Selvin Xavier <selvin.xavier@broadcom.com> Signed-off-by: Leon Romanovsky <leon@kernel.org>
3 daysRDMA/bnxt_re: Initialize wqe.flags in bnxt_re_post_srq_recv()Selvin Xavier
wqe is declared once above the WR loop and never zeroed, so wqe.flags holds indeterminate stack contents on the first WR and the previous WR's stale value thereafter. This propagates straight into the hardware-visible srqe->flags in bnxt_qplib_post_srq_recv(). Most values of the wqe struct is set except flags. So initialize flags also. Fixes: 37cb11acf1f7 ("RDMA/bnxt_re: Add SRQ support for Broadcom adapters") Signed-off-by: Selvin Xavier <selvin.xavier@broadcom.com> Signed-off-by: Leon Romanovsky <leon@kernel.org>
3 daysRDMA/bnxt_re: Validate SRQ max_sge at create timeSelvin Xavier
bnxt_re_create_srq() stored attr.max_sge into srq->qplib_srq.max_sge unvalidated, which defeats the num_sge check in bnxt_re_post_srq_recv() since that check compares against this same attacker-chosen value. Reject max_sge > dev_attr->max_srq_sges at create time, and clamp dev_attr->max_srq_sges to BNXT_STATIC_MAX_SGE since SRQ WQEs use the fixed-size SGE array. Fixes: 37cb11acf1f7 ("RDMA/bnxt_re: Add SRQ support for Broadcom adapters") Signed-off-by: Selvin Xavier <selvin.xavier@broadcom.com> Signed-off-by: Leon Romanovsky <leon@kernel.org>
3 daysRDMA/bnxt_re: Detect wrong sge_len passed for inlineSelvin Xavier
Avoid handling wrong sge_len by adding extra check to see if the passed length is more than the inline size supported. Fixes: 1ac5a4047975 ("RDMA/bnxt_re: Add bnxt_re RoCE driver") Reviewed-by: Kalesh AP <kalesh-anakkur.purayil@broadcom.com> Signed-off-by: Selvin Xavier <selvin.xavier@broadcom.com> Signed-off-by: Leon Romanovsky <leon@kernel.org>
3 daysRDMA/bnxt_re: Fix integer overflow in send payload size computationSelvin Xavier
bnxt_re_build_sgl() summed SGE lengths into a signed int, which could overflow. Make it return u32, and have bnxt_re_copy_wr_payload() report the size via a u32 out-parameter with the return value carrying only the error status, instead of overloading a signed int with both. Fixes: 1ac5a4047975 ("RDMA/bnxt_re: Add bnxt_re RoCE driver") Signed-off-by: Selvin Xavier <selvin.xavier@broadcom.com> Signed-off-by: Leon Romanovsky <leon@kernel.org>
3 daysRDMA/mana_ib: Poll UD completions and flush software error QPsKonstantin Taranov
Build WCs directly from CQEs and their corresponding published shadow entries. This avoids iterating over all QPs that share a CQ. Track only QPs in the error state instead of every QP associated with the CQ. Keep the explicit GSI send-drain path responsible for adding the QP to the send error list and invoking its CQ handler. To invoke the handler for WQEs posted after GSI draining, check the QP state after posting WQEs. This is required because MANA hardware ignores doorbells for QPs in the error state. Introduce a WC-budgeted polling context that caches CQEs representing multiple WCs. Before consuming a shadow entry, match the CQE against the low 24 bits of the posted WQE offset to filter out stale CQEs. Implement WR signaling to skip WC generation for unsignaled WQEs. Signed-off-by: Konstantin Taranov <kotaranov@microsoft.com> Link: https://patch.msgid.link/20260923131155.4055875-6-kotaranov@linux.microsoft.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
3 daysRDMA/mana_ib: Make kernel CQ arming robustKonstantin Taranov
Track CQ polling credits and ring the doorbell promptly to advance the hardware CQ tail pointer. The distance between the new and previously armed tails must remain below four times the CQ length. Expose CQ arming to mana_ib through mana_gd_wq_ring_doorbell_ext(). RC QPs will also use this helper to ring special doorbell offsets. MANA hardware ignores duplicate doorbells. Calculate the arming offset from the previously armed offset and poll_credit to avoid duplicates. Do not arm the CQ if no CQE was produced, because the CQ remains armed. Signed-off-by: Konstantin Taranov <kotaranov@microsoft.com> Link: https://patch.msgid.link/20260923131155.4055875-5-kotaranov@linux.microsoft.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
3 daysRDMA/mana_ib: Revise UD receive posting with GDMA_WR_IB_SGLKonstantin Taranov
Decouple RQ WQE posting from the QP structure so it can be reused by other QP types. Batch multiple receive WQEs under a single doorbell. Introduce GDMA_WR_IB_SGL to pass SGEs directly to GDMA instead of copying them into a small stack array. Support receive WQEs with num_sge set to zero. Signed-off-by: Konstantin Taranov <kotaranov@microsoft.com> Link: https://patch.msgid.link/20260923131155.4055875-4-kotaranov@linux.microsoft.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
3 daysRDMA/mana_ib: Revise UD send posting and WQE definitionsKonstantin Taranov
Batch multiple send requests under a single doorbell. Move the send PSN to the common QP structure so it can be shared by all QP types. Define new WQE formats for UD and RC QPs. Signed-off-by: Konstantin Taranov <kotaranov@microsoft.com> Link: https://patch.msgid.link/20260923131155.4055875-3-kotaranov@linux.microsoft.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
3 daysRDMA/mana_ib: Optimize shadow queue bookkeepingKonstantin Taranov
Pack the posted WQE size and send opcode into compact bitfields. Move the UD receive shadow fields into the common shadow_wqe_header. This simplifies the code and lets all QP and WQ types share one structure. Use release and acquire operations so posting and polling can safely run on different CPUs. Signed-off-by: Konstantin Taranov <kotaranov@microsoft.com> Link: https://patch.msgid.link/20260923131155.4055875-2-kotaranov@linux.microsoft.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
3 daysIB/isert: Don't share drop_cmd_list across concurrent connection teardownsLi RongQing
isert_put_unsol_pending_cmds() declares drop_cmd_list as a static LIST_HEAD, so the list is shared across all invocations of the function. It is called from isert_wait_conn(), the per-connection .iscsit_wait_conn transport callback, which runs independently for each connection during teardown. conn->cmd_lock only protects each connection's own conn_cmd_list, not the shared drop_cmd_list. If two connections tear down at the same time, both threads list_move_tail() commands into the same drop_cmd_list without any mutual exclusion, and then both iterate and list_del/put it concurrently. This corrupts the linked list and can double-free iscsit_cmd structures. Make drop_cmd_list a stack-local list, matching the pattern already used in isert_free_np(). Fixes: 3e03c4b01da3 ("iser-target: Put the reference on commands waiting for unsol data") Signed-off-by: Li RongQing <lirongqing@baidu.com> Link: https://patch.msgid.link/20260921031100.2199-1-lirongqing@baidu.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
3 daysRDMA/mlx5: Restore mr->pi_mr after temporary stack-local assignmentLi RongQing
In handle_reg_mr_integrity, when mr->pi_mr is NULL (physical-address optimization path), the code points mr->pi_mr at a stack-local pa_pi_mr so set_pi_umr_wr can read it during this call. The pointer is never restored, so after the function returns mr->pi_mr dangles at a freed stack frame. A second IB_WR_REG_MR_INTEGRITY posted on the same MR without remapping reads the stale pointer and dereferences it. Save the original pi_mr (captured at function entry) and restore it on the out path. Fixes: 2563e2f30acb ("RDMA/mlx5: Use PA mapping for PI handover") Signed-off-by: Li RongQing <lirongqing@baidu.com> Link: https://patch.msgid.link/20260919100848.2442-1-lirongqing@baidu.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
4 daysRDMA/mlx5: Fix signed integer overflow in EQE qp_srq type shiftYishai Hadas
eqe->data.qp_srq.type is a u8. Shifting it left by MLX5_USER_INDEX_LEN (24) without a cast promotes it to signed int, causing undefined behaviour for values >= 128. All currently used type values are well below 128, so this does not manifest with current hardware. Cast to u32 before the shift to make the code well-defined across the full range. Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-16-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
4 daysRDMA/mlx5: Initialize QP/RQ/SQ resource refcount before publishing itYishai Hadas
create_resource_common() inserted the resource into the radix tree before its initialization was complete. A firmware EQE in that window finds a zero refcount_t and calls refcount_inc() on it; the refcount API treats this as a use-after-free, saturates the counter, and permanently breaks the resource's lifecycle tracking. Move the initialization before the radix_tree_insert(). Fixes: e126ba97dba9 ("mlx5: Add driver for Mellanox Connect-IB adapters") Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-15-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
4 daysRDMA/mlx5: Set QP event handler before firmware QPC insertionYishai Hadas
Four QP creation paths assigned base->container_mibqp and base->mqp.event after the firmware create call that inserts the QP into dev->qp_table.tree. A firmware error EQE in that window reaches qp->event() with a NULL pointer. Move both assignments before the firmware call. Also set ibqp.qp_num inside mlx5_qpc_create_qp() before create_resource_common() so an EQE arriving early does not observe qp_num==0. In addition, add a WARN_ON_ONCE(!qp->event) guard in rsc_event_notifier() as a safeguard for future regressions. Fixes: e126ba97dba9 ("mlx5: Add driver for Mellanox Connect-IB adapters") Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-14-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
4 daysRDMA/mlx5: Set RQ event handler for raw-packet QPYishai Hadas
create_raw_packet_qp() creates a separate firmware RQ object but never assigned rq->base.mqp.event. Any firmware event on this resource (e.g. WQ_CATAS_ERROR, WQ_ACCESS_ERROR) calls through a NULL pointer. Set the event pointer before the RQ is inserted into the resource table. Fixes: 0fb2ed66a14c ("IB/mlx5: Add create and destroy functionality for Raw Packet QP") Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-13-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
4 daysRDMA/mlx5: Set SRQ event handler before xarray insertionYishai Hadas
mlx5_cmd_create_srq() stores the SRQ into the xarray before srq->msrq.event was assigned. A firmware SRQ event arriving in that window calls through a NULL pointer. Move the assignment before mlx5_cmd_create_srq(). Fixes: e126ba97dba9 ("mlx5: Add driver for Mellanox Connect-IB adapters") Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-12-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
4 daysRDMA/mlx5: Set WQ event handler before firmware RQ insertionYishai Hadas
create_rq() inserts the RQ into dev->qp_table.tree, making it visible to rsc_event_notifier() before rwq->core_qp.event was assigned. A firmware WQ_CATAS_ERROR event in that window finds a NULL event pointer. Move the assignment before create_rq(). Fixes: 350d0e4c7e4b ("IB/mlx5: Track asynchronous events on a receive work queue") Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-11-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
4 daysRDMA/mlx5: Fix mlx5_ib_dev_res_init() failure when XRC cap is absentYishai Hadas
mlx5_ib_dev_res_init() returns -EOPNOTSUPP when the firmware XRC capability is absent, failing driver probe on XRC-less devices. Make XRC optional throughout the devr resource path: initialize cq_lock and srq_lock unconditionally (needed for every QP creation), skip xrcd allocation and the XRC-type placeholder SRQ (s0) when xrc=0, and guard their cleanup paths accordingly. Fixes: f4375443b786 ("RDMA/mlx5: Get XRCD number directly for the internal use") Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-9-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
4 daysRDMA/umem: Derive DMA direction from IB access flagsYishai Hadas
Map non-writable user memory DMA_TO_DEVICE: when no write access is granted, the NIC only reads the pages and DMA_BIDIRECTIONAL is unnecessarily broad. Where the IOMMU enforces direction this prevents the device writing to read-only buffers; where it does not the mapping is still semantically correct. Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-8-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
4 daysRDMA/umem: Map CQ buffers DMA_FROM_DEVICEYishai Hadas
For a CQ ring the device writes CQEs and the CPU only reads. DMA_FROM_DEVICE is the correct direction. Use it explicitly in ib_umem_get_cq_buf() and ib_umem_get_cq_buf_or_va() rather than deriving it from the access flags: CQ callers pass IB_ACCESS_LOCAL_WRITE so that get_user_pages() pins a writable private page (FOLL_WRITE), which the device needs to DMA-write CQEs into. That flag would yield DMA_BIDIRECTIONAL from the access-flag derivation, which is wrong for the bus direction. Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-7-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
4 daysRDMA/umem: Support an explicit DMA direction other than DMA_BIDIRECTIONALYishai Hadas
Add a dma_dir field to struct ib_umem to record the direction chosen at map time and reuse it at unmap. Thread an explicit dma_data_direction parameter through the internal pinning helpers so each caller can supply the direction that matches the device's actual access pattern. All existing callers still pass DMA_BIDIRECTIONAL, so this is a pure mechanism addition with no behavioural change. Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-6-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
4 daysRDMA/umem: Reuse ib_umem_get_cq_buf_or_va() for VA-only CQ pinningYishai Hadas
Several drivers pin their CQ ring buffer via ib_umem_get_va() rather than the CQ-specific helper. Switch them to ib_umem_get_cq_buf_or_va() so the subsequent DMA_FROM_DEVICE change covers all CQ paths at once. For shared helpers used by both CQ and non-CQ callers (qedr, mana, erdma, hns) a new is_cq parameter selects the appropriate pinning function; non-CQ callers pass false. vmw_pvrdma's CQ is excluded: it embeds a ring-state header that the driver CPU-writes after polling, making it genuinely bidirectional. This is a pure refactor with no functional change. Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-5-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
4 daysRDMA/vmw_pvrdma: Pin QP and SRQ rings writable to match device DMA write accessYishai Hadas
pvrdma embeds a pvrdma_ring_state header directly inside the QP and SRQ ring buffers, and the hypervisor writes cons_head into that header. Pinning with access=0 does not force a private page copy (no FOLL_WRITE), so the hypervisor may write into a shared page. Pass IB_ACCESS_LOCAL_WRITE for all three rings. The SRQ ring_state is not currently used by the driver but is pinned writable for consistency. Fixes: 29c8d9eba550 ("IB: Add vmw_pvrdma driver") Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-4-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
4 daysRDMA/hns: Pin CQ buffer writable to match device DMA write accessYishai Hadas
alloc_cq_buf() leaves user_access zero-initialized, but the device writes CQEs into the CQ buffer. Set IB_ACCESS_LOCAL_WRITE so the buffer is pinned writable, matching the device's write access. Fixes: 744b7bdfa79e ("RDMA/hns: Support 0 hop addressing for CQE buffer") Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-3-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
4 daysRDMA/erdma: Pin CQ buffer writable to match device DMA write accessYishai Hadas
erdma_init_user_cq() pins the CQ ring with access=0, but the device writes CQEs into it. Pass IB_ACCESS_LOCAL_WRITE so the buffer is pinned writable and marked dirty on unpin. Fixes: 155055771704 ("RDMA/erdma: Add verbs implementation") Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-2-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
4 daysRDMA/irdma: Free IRQ when CEQ vector mapping failsSerhat Kumral
irdma_cfg_ceq_vector() registers an IRQ handler before requesting the CEQ-to-vector mapping through the virtual channel. If the mapping fails, the function returns without releasing the IRQ resources. The caller destroys the CEQ, leaving the IRQ handler registered and potentially referencing freed memory. Call irdma_destroy_irq() on mapping failure to release the IRQ resources and drain the associated tasklet before the caller destroys the CEQ. Use the same dev_id passed to request_irq(), accounting for CEQ0 sharing its interrupt vector with the AEQ. Fixes: b800e82feba7 ("RDMA/irdma: Add GEN3 support for AEQ and CEQ") Assisted-by: Claude:claude-opus-5 Signed-off-by: Serhat Kumral <serhatkumral1@gmail.com> Link: https://patch.msgid.link/20260918110428.36586-1-serhatkumral1@gmail.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
9 daysRDMA/mlx5: Use DEFINE_RAW_FLEX for leftovers flow attributesLi RongQing
Commit e91fb8b9d0ed ("RDMA/mlx5: Avoid flexible array warning") split the leftovers_specs[] array into two separate structs to silence the sparse "array of flexible structures" warning, but in doing so reversed the member order: eth_flow was placed before flow_attr. _create_flow_rule() locates the flow specs with: ib_flow = (void *)flow_attr + sizeof(*flow_attr); which expects the spec data to immediately follow the ib_flow_attr. With the reversed order, flow_attr is the last member, so ib_flow points past the end of the struct and the num_of_specs = 1 loop reads out-of-bounds memory (whatever static data follows in the .data section), producing incorrect flow rules or a crash. Rebuild the leftovers attributes with DEFINE_RAW_FLEX(), which declares an on-stack struct ib_flow_attr with a properly sized trailing flows[] array. The ETH spec lives in flows[0].eth at exactly the offset _create_flow_rule() expects, so no layout-dependent composite struct is needed. This also avoids embedding a structure that contains a flexible array member as a non-last member, which would trigger -Wflex-array-member-not-at-end once the warning is enabled. Fixes: e91fb8b9d0ed ("RDMA/mlx5: Avoid flexible array warning") Signed-off-by: Li RongQing <lirongqing@baidu.com> Link: https://patch.msgid.link/20260918092337.2338-1-lirongqing@baidu.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
9 daysRDMA/erdma: Move kernel QP helpers after memory helpersCheng Xu
Move the kernel QP helpers after the shared memory helpers to remove forward declarations. Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com> Link: https://patch.msgid.link/20260917120459.13286-6-chengyou@linux.alibaba.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
9 daysRDMA/erdma: Unify userspace and kernel queue buffer managementCheng Xu
Userspace and kernel space queue buffers use separate helpers despite sharing MTT metadata and lifetime rules. Manage both through erdma_mem_init() and erdma_mem_uninit(), while keeping backing allocation and release type-specific. Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com> Link: https://patch.msgid.link/20260917120459.13286-5-chengyou@linux.alibaba.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
9 daysRDMA/erdma: Support non-contiguous kernel CQ buffersCheng Xu
A single coherent allocation for a kernel CQ can fail when memory is fragmented. Use page-sized coherent buffers and describe them with the existing MTT. Keep the userspace CQ path and doorbell allocation unchanged. Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com> Link: https://patch.msgid.link/20260917120459.13286-4-chengyou@linux.alibaba.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
9 daysRDMA/erdma: Support non-contiguous kernel QP buffersCheng Xu
A single coherent allocation for kernel QP queues can fail for large queues when memory is fragmented. Allocate page-sized coherent buffers and describe them with the existing MTT. Keep the userspace QP path unchanged. Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com> Link: https://patch.msgid.link/20260917120459.13286-3-chengyou@linux.alibaba.com Signed-off-by: Leon Romanovsky <leon@kernel.org>