| Age | Commit message (Collapse) | Author |
|
https://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next.git
# Conflicts:
# drivers/net/ethernet/realtek/r8169_main.c
# net/mac80211/ieee80211_i.h
# net/mac80211/tx.c
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/rdma/rdma.git
# Conflicts:
# drivers/infiniband/sw/rxe/rxe_verbs.c
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/mm/linux.git
# Conflicts:
# arch/arm64/kvm/mmu.c
|
|
fault_create_debugfs_attr() has always taken an extra dentry reference on
the created directory (attr->dname = dget(dir)) so that fail_dump() could
print the name via %pd from any context. Nothing anywhere in the tree
ever calls dput() on attr->dname.
For callers with a matching teardown, that unmatched reference causes one
dentry plus its attached inode to leak per fault_create_debugfs_attr /
debugfs_remove_recursive cycle. simple_recursive_removal() drops
debugfs's own +1 ref on the child dentry, but the dget()'d ref keeps its
refcount at 1: the dentry ends up unhashed but pinned, and its inode is
never freed.
Boot-once callers (mm/failslab, block/blk-core, etc.) leak exactly once at
init and never destroy the tree, so the impact there is bounded. But
per-lifecycle callers (drivers/nvme, drivers/infiniband/hw/hfi1,
drivers/mmc, drivers/iommu/iommufd, drivers/media, drivers/misc,
drivers/gpu/drm/msm, drivers/crypto, net/sunrpc) leak on every
create/destroy cycle.
We observed this in production: an NVMe/RDMA host repeatedly reconnecting
to a target that rejected the CRTO Property Get went through ~50 nvme
controller create/destroy cycles per second, and dentry and inode_cache
grew by ~13k pinned objects per 240 s -- unrecoverable through
drop_caches. Byte math matched a per-cycle 1-dentry / 1-inode leak from
the "fault_inject" directory dentry.
Fix this by not holding any external reference in fault_attr. Embed the
directory name as a fixed-size char array (FAULT_ATTR_DNAME_LEN, 64 bytes)
inside struct fault_attr, copied by strscpy() at
fault_create_debugfs_attr() time. fail_dump() prints it via %s.
Advantages of an embedded array over kstrdup() + kfree() paired with a new
destroy API:
- Zero API footprint. No new export and no caller changes required:
callers already own their fault_attr's memory and free it when
they are done, and now that suffices.
- No allocation on the create path.
- fault_create_debugfs_attr() cannot fail from the name-copy step.
- No lifetime coupling between attr->dname and debugfs; the string
is valid for exactly as long as the containing struct.
The 64-byte length accommodates every in-tree caller with generous
headroom (the longest current name is "fail_dma_array_full", 19 chars).
The user-visible fail_dump() format changes from "name %pd" to "name %s",
but the printed content is identical -- %pd on the created directory
renders the same string that was passed in as @name.
drivers/infiniband/hw/hfi1/fault.c drops a now-invalid "attr.dname = NULL"
statement; the surrounding kzalloc() already zero-initialises the array.
Link: https://lore.kernel.org/20260821181527.3271414-1-mliang@purestorage.com
Fixes: 6adc4a22f20b ("fault-inject: add ratelimit option")
Signed-off-by: Michael Liang <mliang@purestorage.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reviewed-by: Andrew Morton <akpm@linux-foundation.org>
Cc: Akinbou Mita <akinobu.mita@gmail.com>
Cc: Dennis Dalessandro <dennis.dalessandro@cornelisnetworks.com>
Cc: Jason Gunthorpe <jgg@ziepe.ca>
Cc: Leon Romanovsky <leon@kernel.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: <stable@vger.kernel.org>
|
|
In cases which map chip memory from vmalloc()'d ranges, the hfi1
infiniband drivers currently installs a fault handler, and then smuggles
the kernel virtual address of this range in vma->vm_pgoff.
This is exposing KASLR-sensitive internal kernel state in the VMA, and is
entirely unnecessary.
Instead, use remap_vmalloc_range() to remap the VMA to the span, and
eliminate the fault handler altogether.
remap_vmalloc_range() checks that the VMA does not extend beyond the
vmalloc area, and the driver already requires the VMA to exactly match the
span of the memory being mapped, so this has no impact.
The memory is all preallocated so not having a fault handler has no impact
either, other than pre-mapping the ranges which is beneficial.
We also remove the VM_IO flag as it's not appropriate here, and the
VM_DONTEXPAND flag as remap_vmalloc_range() will set it (and also mark the
range correctly as a mixed map).
We also update the vmalloc paths to place the virtual kernel address in
memvirt, rather than overloading the physical address memaddr. We
predicate the vmalloc handling on the vmalloc flag before we check memvirt
for the virtual address-derived PFN remap path, so this works fine.
remap_vmalloc_range() requires that the vmalloc()'d areas were all
allocated using vmalloc_user() - each of cq->comps,
uctxt->subctxt_rcvegrbuf, uctxt->subctxt_rcvhdr_base,
uctxt->subctxt_uregbase and dd->events were allocated this way, so that
requirement is satisfied.
We also remove VM_IO and VM_DONTEXPAND from the STATUS command, as these
are both set on remap.
Finally, we remove VM_DONTEXPAND from the PIO_BUFS, PIO_BUFS_SOP and UREGS
commands, as these are also all set on remap. PIO_CRED retains it, as
dma_mmap_coherent() may map via vm_insert_page() on the IOMMU-DMA path,
which sets only VM_MIXEDMAP. The RCV_HDRQ, RCV_EGRBUF and RTAIL commands
also map via dma_mmap_coherent() and never set VM_DONTEXPAND, so set it
for them for the same reason.
Note that we retain expected behaviour throughout - the vmalloc remapped
ranges set VM_MIXEDMAP | VM_DONTDUMP | VM_DONTEXPAND for each range.
VM_IO was never appropriate as the ranges are explicitly not MMIO, and the
reference to the v3.7 VM_RESERVED semantics map on to VM_MIXEDMAP |
VM_DONTDUMP | VM_DONTEXPAND correctly - no core dump, unmergeable, no
normal vm page for purposes of reclaim/migration/etc.
There is a change in behaviour in that pages mapped using
remap_vmalloc_range() will now have normal GUP-able pages, however this
should have no impact as there is no reason not to allow this.
Link: https://lore.kernel.org/20260917-b4-mmap-prepare-vma-flag-sanify-v3-11-4583d8a23bca@kernel.org
Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Cc: Albert Ou <aou@eecs.berkeley.edu>
Cc: Alexander Gordeev <agordeev@linux.ibm.com>
Cc: Alexei Starovoitov <ast@kernel.org>
Cc: Alistair Popple <apopple@nvidia.com>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: Andreas Larsson <andreas@gaisler.com>
Cc: Andrii Nakryiko <andrii@kernel.org>
Cc: "Aneesh Kumar K.V" <aneesh.kumar@kernel.org>
Cc: Anup Patel <anup@brainfault.org>
Cc: Arnaldo Carvalho de Melo <acme@kernel.org>
Cc: Arnd Bergmann <arnd@arndb.de>
Cc: Axel Rasmussen <axelrasmussen@google.com>
Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: Baoquan He <baoquan.he@linux.dev>
Cc: Barry Song <baohua@kernel.org>
Cc: "Borislav Petkov (AMD)" <bp@alien8.de>
Cc: Byungchul Park <byungchul@sk.com>
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Chengming Zhou <chengming.zhou@linux.dev>
Cc: Chris Li <chrisl@kernel.org>
Cc: Christian Borntraeger <borntraeger@linux.ibm.com>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Claudio Imbrenda <imbrenda@linux.ibm.com>
Cc: Dave Airlie <airlied@gmail.com>
Cc: Dave Hansen <dave.hansen@linux.intel.com>
Cc: David Hildenbrand <david@kernel.org>
Cc: David S. Miller <davem@davemloft.net>
Cc: Dennis Dalessandro <dennis.dalessandro@cornelisnetworks.com>
Cc: Dev Jain <dev.jain@arm.com>
Cc: Doug Gilbert <dgilbert@interlog.com>
Cc: Eduard Zingerman <eddyz87@gmail.com>
Cc: Emil Tsalapatis <emil@etsalapatis.com>
Cc: Gerald Schaefer <gerald.schaefer@linux.ibm.com>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Gregory Price <gourry@gourry.net>
Cc: Harry Yoo <harry@kernel.org>
Cc: Heiko Carstens <hca@linux.ibm.com>
Cc: Helge Deller <deller@gmx.de>
Cc: "Huang, Ying" <ying.huang@linux.alibaba.com>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: James Bottomley <james.bottomley@HansenPartnership.com>
Cc: Jan Kara <jack@suse.cz>
Cc: Jann Horn <jannh@google.com>
Cc: Janosch Frank <frankja@linux.ibm.com>
Cc: Jaroslav Kysela <perex@perex.cz>
Cc: Jason Gunthorpe <jgg@ziepe.ca>
Cc: Jaya Kumar <jayalk@intworks.biz>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: John Hubbard <jhubbard@nvidia.com>
Cc: Jonathan Corbet <corbet@lwn.net>
Cc: Joshua Hahn <joshua.hahnjy@gmail.com>
Cc: Juri Lelli <juri.lelli@redhat.com>
Cc: Kairui Song <kasong@tencent.com>
Cc: Kemeng Shi <shikemeng@huaweicloud.com>
Cc: Kiryl Shutsemau <kas@kernel.org>
Cc: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Cc: Lance Yang <lance.yang@linux.dev>
Cc: Leon Romanovsky <leon@kernel.org>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Maarten Lankhorst <maarten.lankhorst@linux.intel.com>
Cc: Madhavan Srinivasan <maddy@linux.ibm.com>
Cc: Marc Rutland <mark.rutland@arm.com>
Cc: Marc Zyngier <maz@kernel.org>
Cc: "Masami Hiramatsu (Google)" <mhiramat@kernel.org>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
Cc: Maxime Ripard <mripard@kernel.org>
Cc: Michal Hocko <mhocko@kernel.org>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Miklos Szeredi <miklos@szeredi.hu>
Cc: Muchun Song <muchun.song@linux.dev>
Cc: Namhyung kim <namhyung@kernel.org>
Cc: Nhat Pham <nphamcs@gmail.com>
Cc: Nicholas Piggin <npiggin@gmail.com>
Cc: Oleg Nesterov <oleg@redhat.com>
Cc: Oscar Salvador <osalvador@suse.de>
Cc: Palmer Dabbelt <palmer@dabbelt.com>
Cc: Paul Moore <paul@paul-moore.com>
Cc: Pedro Falcato <pfalcato@suse.de>
Cc: Peter Xu <peterx@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Rakie Kim <rakie.kim@sk.com>
Cc: Rik van Riel <riel@surriel.com>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: Sebastian Reichel <sre@kernel.org>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Stephen Smalley <stephen.smalley.work@gmail.com>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Takashi Iwai (SUSE) <tiwai@suse.de>
Cc: Takashi Iwai <tiwai@suse.com>
Cc: Thomas Zimemrmann <tzimmermann@suse.de>
Cc: Vasily Gorbik <gor@linux.ibm.com>
Cc: Vincent Guittot <vincent.guittot@linaro.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Wei Xu <weixugc@google.com>
Cc: Will Deacon <will@kernel.org>
Cc: Yuanchu Xie <yuanchu@google.com>
Cc: Zi Yan <ziy@nvidia.com>
|
|
The error unwind in ionic_create_qp() removes the SQ and RQ CMB mmap
entries explicitly, and then falls through to ionic_qp_rq_destroy() and
ionic_qp_sq_destroy(), whose ionic_qp_{rq,sq}_destroy_cmb() helpers
remove the very same entries a second time.
Drop the explicit removals and leave ionic_qp_{rq,sq}_destroy_cmb() as
the single owner on the ionic_destroy_qp() path.
Fixes: e8521822c733 ("RDMA/ionic: Register device ops for control path")
Signed-off-by: Abhijit Gangurde <abhijit.gangurde@amd.com>
Link: https://patch.msgid.link/20260925093638.3032149-2-abhijit.gangurde@amd.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
xa_init_flags() takes XArray flags (XA_FLAGS_*), not gfp flags, even
though the parameter is typed gfp_t. Passing GFP_ATOMIC leaves both
lock-type bits clear, so xa_lock_type() resolves to XA_LOCK_NORMAL for
qp_tbl and cq_tbl.
Both tables are used with the IRQ-safe accessors: xa_store_irq() with
GFP_KERNEL from process context, and xa_lock_irqsave() from the event
queue path. When a store needs to grow the tree, __xas_nomem() drops
the lock using the recorded lock type before allocating:
if (gfpflags_allow_blocking(gfp)) {
xas_unlock_type(xas, lock_type);
xas->xa_alloc = kmem_cache_alloc_lru(...);
xas_lock_type(xas, lock_type);
}
With XA_LOCK_NORMAL that is a plain spin_unlock(), so the sleeping
GFP_KERNEL allocation runs with interrupts still disabled by
xa_store_irq(), and lockdep annotates the lock with the wrong class.
GFP_ATOMIC also happens to set __GFP_HIGH (0x20), which collides with
XA_FLAGS_ACCOUNT (32U) and silently enables memcg accounting of the
XArray nodes.
Use XA_FLAGS_LOCK_IRQ so the lock type matches how the tables are
actually accessed.
Fixes: e8521822c733 ("RDMA/ionic: Register device ops for control path")
Signed-off-by: Abhijit Gangurde <abhijit.gangurde@amd.com>
Link: https://patch.msgid.link/20260925093638.3032149-1-abhijit.gangurde@amd.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
After do_tcp_sendpages() was converted to use MSG_SPLICE_PAGES and
later inlined as direct tcp_sendmsg_locked() calls, callers that
previously needed an explicit tcp_rate_check_app_limited(sk) to cover
do_tcp_sendpages() no longer need it. tcp_sendmsg_locked() performs the
check on every path that queues data, making the outer call redundant.
The site changed here, siw_tcp_sendpages() in siw_qp_tx.c, holds the socket
lock and invokes tcp_sendmsg_locked() on every iteration. The early-return
paths in tcp_sendmsg_locked() that skip tcp_rate_check_app_limited() - the
MSG_ZEROCOPY allocation failure and MSG_FASTOPEN branches - return without
queueing any MSG_SPLICE_PAGES data, so there is no functional consequence
from omitting the outer check.
Signed-off-by: Geliang Tang <tanggeliang@kylinos.cn>
Link: https://patch.msgid.link/32ef40cc81486adb8f2d43585224ab44d7e77f73.1790234916.git.tanggeliang@kylinos.cn
Acked-by: Bernard Metzler <bernard.metzler@linux.dev>
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
rxe_ib_advise_mr_prefetch() and rxe_ib_prefetch_sg_list() look up the MR
by lkey and hand it to rxe_odp_do_pagefault_and_lock() without checking
that it is an ODP MR. That path runs to_ib_umem_odp() on mr->umem, and
for a non-ODP MR mr->umem is a plain struct ib_umem from ib_umem_get(),
so the container_of() in to_ib_umem_odp() lands past the end of the
object and ib_umem_odp_map_dma_and_lock() reads its ib_umem_odp fields
out of bounds.
lookup_mr() validates the lkey, PD, access and state but not the MR
type, and IB_UVERBS_ADVISE_MR_ADVICE_PREFETCH is accepted for any MR.
BUG: KASAN: slab-out-of-bounds in ib_umem_odp_map_dma_and_lock+0x884/0x8a0
Read of size 8 at addr ffff88810a3ebcf0 by task advi/921
ib_umem_odp_map_dma_and_lock+0x884/0x8a0
rxe_ib_advise_mr+0x543/0xad0
ib_uverbs_handler_UVERBS_METHOD_ADVISE_MR+0x446/0x530
ib_uverbs_cmd_verbs+0x2b3c/0x3b20
ib_uverbs_ioctl+0x1e3/0x310
Allocated by task 921:
__ib_umem_get_va+0x13e/0xae0
rxe_mr_init_user+0x2ae/0xb00
rxe_reg_user_mr+0x337/0x510
The buggy address belongs to the object at ffff88810a3ebc80
which belongs to the cache kmalloc-96 of size 96
Ask lookup_mr() for IB_ACCESS_ON_DEMAND in both the synchronous and
the asynchronous prefetch arm. mr->access carries that flag only for
an MR registered as ODP, so the existing
(access & mr->access) != access test rejects a plain MR and the
prefetch fails with -EINVAL.
Fixes: 3576b0df1588 ("RDMA/rxe: Implement synchronous prefetch for ODP MRs")
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
Link: https://patch.msgid.link/20260923-rxe-advise-mr-v3-v3-2-95e1a4077e4d@doyensec.com
Reviewed-by: Zhu Yanjun <yanjun.zhu@linux.dev>
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Whether an MR is an ODP MR is decided once, at registration time:
rxe_reg_user_mr() picks rxe_odp_mr_init_user() over rxe_mr_init_user()
based on IB_ACCESS_ON_DEMAND, and only the former builds an ib_umem_odp
via ib_umem_odp_get(). The umem cannot change type afterwards, and
is_odp_mr() reads mr->umem->is_odp.
Two paths assign mr->access after that point and can leave it
describing an MR type the umem does not have:
rxe_rereg_user_mr() with IB_MR_REREG_ACCESS overwrites mr->access with
the caller's value, and IB_ACCESS_ON_DEMAND is part of
RXE_ACCESS_SUPPORTED_MR, so userspace can set the flag on a plain MR or
clear it on an ODP MR while the umem stays what it was.
rxe_reg_fast_mr() takes mr->access from the REG_MR work request
unmasked and moves the MR to RXE_MR_STATE_VALID, on an MR that
rxe_mr_init_fast() left with a NULL umem.
Reject IB_ACCESS_ON_DEMAND in both, so mr->access carries the flag only
for an MR that has an ODP umem and the flag can be used to identify one.
Fixes: 544c7f62cf32 ("RDMA/rxe: Implement rereg_user_mr")
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
Link: https://patch.msgid.link/20260923-rxe-advise-mr-v3-v3-1-95e1a4077e4d@doyensec.com
Reviewed-by: Zhu Yanjun <yanjun.zhu@linux.dev>
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
KASAN reports a slab-use-after-free in ip6_route_output_flags()
reached from rxe_find_route(), with the free in __sk_destruct() after
rxe_sock_put(), and a user with CAP_NET_ADMIN can remove the device with
"rdma link del rxe0" while RoCE v2 over IPv6 traffic keeps looking the
socket up.
BUG: KASAN: slab-use-after-free in ip6_route_output_flags+0x300/0x360
Read of size 4 at addr ffff888115ddd794 by task kworker/u32:6/309
Call Trace:
<TASK>
dump_stack_lvl+0x53/0x70
print_report+0xd0/0x630
? __pfx__raw_spin_lock_irqsave+0x10/0x10
? ip6_route_output_flags+0x300/0x360
kasan_report+0xce/0x100
? ip6_route_output_flags+0x300/0x360
ip6_route_output_flags+0x300/0x360
ip6_dst_lookup_tail.constprop.0+0x76c/0xcc0
? ct_nmi_exit+0xc3/0xf0
ip6_dst_lookup_flow+0xf5/0x1e0
? __pfx_ip6_dst_lookup_flow+0x10/0x10
rxe_find_route+0x426/0xa30
? __kasan_slab_alloc+0x6e/0x70
? __pfx_rxe_find_route+0x10/0x10
? kmem_cache_alloc_node_noprof+0x141/0x370
? kmalloc_reserve+0x103/0x2b0
? rxe_icrc_generate+0x229/0x330
? __pfx___alloc_skb+0x10/0x10
rxe_prepare+0x9e8/0x18b0
? rxe_init_packet+0x3c7/0x4f0
rxe_requester+0x1a0f/0x51f0
? rxe_completer+0x1de9/0x38c0
? __pfx_rxe_completer+0x10/0x10
? __queue_work+0x43e/0x11f0
? __pfx_rxe_requester+0x10/0x10
? irqentry_exit+0xd2/0x640
? _raw_spin_lock_irqsave+0x85/0xe0
? __pfx__raw_spin_lock_irqsave+0x10/0x10
? __pfx_rxe_sender+0x10/0x10
rxe_sender+0xe/0x30
do_work+0x144/0x470
process_one_work+0x633/0x1030
? assign_work+0x11d/0x370
worker_thread+0x45b/0xd10
? __pfx_worker_thread+0x10/0x10
kthread+0x2c6/0x3b0
? recalc_sigpending+0x15c/0x1e0
? __pfx_kthread+0x10/0x10
ret_from_fork+0x36e/0x5a0
? __pfx_ret_from_fork+0x10/0x10
? __switch_to+0x572/0xdd0
? __pfx_kthread+0x10/0x10
ret_from_fork_asm+0x1a/0x30
</TASK>
Allocated by task 146020:
kasan_save_stack+0x33/0x60
kasan_save_track+0x14/0x30
__kasan_slab_alloc+0x6e/0x70
kmem_cache_alloc_noprof+0x130/0x360
sk_prot_alloc+0x56/0x210
Fix by clearing the per-net pointer before the last reference is dropped.
Fixes: f1327abd6abed ("RDMA/rxe: Support RDMA link creation and destruction per net namespace")
Signed-off-by: Binbin Deng <18983559317@163.com>
Link: https://patch.msgid.link/20260923051800.244740-1-18983559317@163.com
Reviewed-by: Zhu Yanjun <yanjun.zhu@linux.dev>
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
queue_depth is taken from the server's CM private_data and later
used as a session invariant. A peer can send 0, which the client
stores and then trips WARN_ON() in create_con_cq_qp() when the
next connection is set up.
Reject a zero queue_depth on ESTABLISHED, same as a bad
magic or version, and fail the handshake with -ECONNRESET.
Fixes: 6a98d71daea1 ("RDMA/rtrs: client: main functionality")
Reported-by: Farhad Alemi <farhad.alemi@berkeley.edu>
Closes: https://lore.kernel.org/linux-rdma/CA+0ovCg_sG8gaZXXj9zmOrpxRt=8+-x9wvycVNUKwWmMbr6-Yg@mail.gmail.com
Signed-off-by: Quanye Yang <quanyeyang@proton.me>
Link: https://patch.msgid.link/20260922-rtrs-fix-warning-rtrs-clt-rdma-cm-handler-v1-1-2076c0ee1455@proton.me
Acked-by: Jack Wang <jinpu.wang@cloud.ionos.com>
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
fill_res_cq_entry() dereferences cq->uobject->uevent.uobject.context->res.id
unconditionally for user resources.
This is normally safe by ordering: destroy_hw() removes the resource
from the restrack before uverbs_destroy_uobject() clears ->context,
and the XA_ZERO_ENTRY marker set by rdma_restrack_begin_del() hides
the entry from concurrent netlink dumps.
The ordering breaks when a driver keeps failing to destroy an object
during ucontext teardown. uverbs_destroy_ufile_hw() then falls back to
__uverbs_cleanup_ufile(RDMA_REMOVE_DRIVER_FAILURE), which
abandons the object in place: the HW object and its restrack entry
are intentionally leaked, while the uobject bookkeeping is torn down
and ->context is explicitly set to NULL by uverbs_destroy_uobject().
rdma_restrack_del() is never reached, so the orphaned entry stays in
the device restrack with a valid kref, reachable by any subsequent
netlink dump.
Observed on 6.1.52 with MLNX_OFED 24.10, after an mlx5 FW failure
left a process unable to tear down its CQ (destroy_cq FW command
failing during FD close):
WARNING: ... uverbs_destroy_ufile_hw+0xe3/0x100
BUG: kernel NULL pointer dereference, address: 0000000000000058
RIP: 0010:fill_res_cq_entry+0x15e/0x180 [ib_core]
res_get_common_dumpit+0x304/0x530 [ib_core]
nldev_res_get_cq_dumpit+0x1a/0x20 [ib_core]
The faulting chain maps to the source (CR2 = 0x58):
cq->uobject (struct ib_cq +0x08, res at +0x98)
uobject->context (struct ib_uobject +0x10, NULL after abandon)
context->res.id (struct ib_ucontext +0x58)
The recent restrack rework (8d186210677c and its series) moved the
restrack deletion to the start of the destroy flow and thus fences
concurrent dumps from an object being destroyed, but it does not
cover this case: when destroy fails, rdma_restrack_abort_del()
restores the entry, and the RDMA_REMOVE_DRIVER_FAILURE sweep still
never removes it from the restrack.
From the fallback until the device is unregistered, any "rdma res
show cq" deterministically takes the NULL pointer dereference; this
is a long-lived state, NOT A RACE. The dump path holds neither the
restrack lock (dropped before the fill callback runs) nor
ufile->hw_destroy_rwsem, and rdma_restrack_get() only guarantees
that the res memory stays alive, not that ->context is still valid.
Fix the dump side: check the uobject context directly instead of the
rdma_is_kernel_res() marker. A non-NULL uobject already implies a
user resource, so kernel resources keep being skipped, and the
additional ->context check skips the RES_CTXN attribute for
orphaned entries instead of crashing the dump. The orphaned resource
itself remains visible in "rdma res show" (cqn/cqe/usecnt/pid),
which is what an operator needs after the accompanying uverbs WARN
to diagnose the driver destroy failure.
fill_res_pd_entry() has the same pattern (pd->uobject->context->res.id)
and is fixed the same way; a PD can even reach the fallback without
its own driver callback failing, e.g. uverbs_free_pd() returns -EBUSY
while another object that failed to destroy still holds the PD usecnt.
Fixes: c3d02788b45a ("RDMA/nldev: Provide parent IDs for PD, MR and QP objects")
Signed-off-by: Yili Zhang <zhangyili01@baidu.com>
Link: https://patch.msgid.link/20260923131239.6521-1-zhangyili01@baidu.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Accept the relaxed ordering access flag during memory region
registration and forward it to the device firmware via the
admin command path.
Reviewed-by: Chen Brasch <cbrasch@amazon.com>
Reviewed-by: Yonatan Nachum <ynachum@amazon.com>
Signed-off-by: Dana Malachi <danamala@amazon.com>
Signed-off-by: Michael Margolin <mrgolin@amazon.com>
Link: https://patch.msgid.link/20260922160118.990-3-mrgolin@amazon.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Set MR permissions explicitly using device interface
definitions rather than copying raw verbs access flags.
This is needed for the next commit, to allow access flag bits
that are not in the permissions field.
Reviewed-by: Chen Brasch <cbrasch@amazon.com>
Reviewed-by: Yonatan Nachum <ynachum@amazon.com>
Signed-off-by: Dana Malachi <danamala@amazon.com>
Signed-off-by: Michael Margolin <mrgolin@amazon.com>
Link: https://patch.msgid.link/20260922160118.990-2-mrgolin@amazon.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
mad_rmpp_recv.state is protected by two locks that do not exclude
each other: recv_timeout_handler() (port workqueue) uses agent->lock,
continue_rmpp() (CQ softirq) uses rmpp_recv->lock. Nothing serializes
the check-then-set sequence on the state word:
T1 (CQ softirq) T2 (port workqueue)
continue_rmpp()
lock(rmpp_recv->lock)
state == ACTIVE [passes]
recv_timeout_handler()
lock(agent->lock)
state = TIMEOUT
list_del(&rmpp_recv->list)
unlock(agent->lock)
destroy_rmpp_recv()
wait_for_completion()
[blocks on T1's ref]
state = COMPLETE
unlock(rmpp_recv->lock)
complete_rmpp()
queue cleanup_work (+10s)
deref()
[unblocks]
kfree(rmpp_recv)
ib_free_recv_mad(done_wc)
+10s: recv_cleanup_handler() runs lock/list_del/destroy on the
freed rmpp_recv -> UAF write and double free; nothing can
cancel that cleanup_work anymore, the entry is already off
agent->rmpp_list. done_wc is also returned to
ib_mad_complete_recv() and freed again -> UAF read and
double free.
Reaching this requires sending MADs to a kernel RMPP agent from the
InfiniBand/RoCE fabric: the per-port sa_query agent accepts subnet
manager responses whose SLID/GID/TID are not authenticated on the
fabric, and legacy user_mad agents registered without
IB_USER_MAD_USER_RMPP start reassembly on any incoming RMPP DATA
MAD with an attacker-chosen TID, no in-flight send required. The
attacker controls segment timing and can run many transactions in
parallel (agent->rmpp_list has no cap), so landing the
microseconds-wide receive critical section on the 40-second timeout
boundary is a matter of patience, not luck.
Unify state transitions under rmpp_recv->lock in
recv_timeout_handler(), keeping list manipulation under agent->lock
(which preserves the agent->lock -> rmpp_recv->lock ordering used
by continue_rmpp()). State check/set in both paths is then
serialized, so the loser of the race always observes the final
state and bails out. recv_cleanup_handler() needs no change: with
transitions serialized, a cleanup_work can only be pending after
COMPLETE, in which case the timeout handler exits early.
The race is as old as the RMPP implementation itself.
Fixes: fa619a77046b ("[PATCH] IB: Add RMPP implementation")
Assisted-by: LLM
Signed-off-by: Dairui Zhang <zhangdairui@gmail.com>
Link: https://patch.msgid.link/20260922152342.1474845-1-zhangdairui@gmail.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Selecting between BNXT_RE_NUM_EXT_COUNTERS and
BNXT_RE_NUM_STD_COUNTERS relied only on
bnxt_qplib_is_chip_gen_p5_p7(), which doesn't account for the
extended stats capability flag or the PF/VF restriction. Use
bnxt_ext_stats_supported() instead, matching the check already
used to populate the extended stats, so the counter count stays
consistent with what gets filled in.
Fixes: 8238c7bd8420 ("RDMA/bnxt_re: Fix the statistics for Gen P7 VF")
Reviewed-by: Kalesh AP <kalesh-anakkur.purayil@broadcom.com>
Signed-off-by: Selvin Xavier <selvin.xavier@broadcom.com>
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
The PD and DPI bitmaps were sized as max >> 3 (bytes),
but bitmap ops (set_bit(), clear_bit(), find_first_bit(),
test_and_set_bit()) operate on whole unsigned long words,
so whenever max isn't a multiple of BITS_PER_LONG,
the buffer under-allocates and the top word's bitops
read/write past the end of the kmalloc()'d buffer.
Most exposed on the DPI table, since dpit->max comes from the
firmware-reported dev_attr->max_dpi with no alignment guarantee.
Fix the size of both allocations with
BITS_TO_LONGS(max) * sizeof(unsigned long).
Fixes: 1ac5a4047975 ("RDMA/bnxt_re: Add bnxt_re RoCE driver")
Signed-off-by: Selvin Xavier <selvin.xavier@broadcom.com>
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
bnxt_re_async_notifier() runs in NAPI/softirq context via bnxt_en's
RCU-protected ULP ops dispatch, never under rtnl_lock, so it isn't
serialized against bnxt_re_uninit_dcb_wq() destroying rdev->dcb_wq.
A concurrent notifier call can queue_work() on a workqueue that is
being, or has just been, destroyed.
Add a spinlock scoped to dcb_wq: the notifier takes it before
checking dcb_wq and queuing work, and bnxt_re_uninit_dcb_wq() takes
it to atomically clear dcb_wq before destroying it. Initialize the
lock at rdev allocation so it is valid on every teardown path,
including bnxt_re_dev_init()'s early failure labels.
Fixes: 51dc5312dcd9 ("RDMA/bnxt_re: Add support to handle DCB_CONFIG_CHANGE event")
Signed-off-by: Selvin Xavier <selvin.xavier@broadcom.com>
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
wqe is declared once above the WR loop and never zeroed, so
wqe.flags holds indeterminate stack contents on the first WR and the
previous WR's stale value thereafter. This propagates straight into
the hardware-visible srqe->flags in bnxt_qplib_post_srq_recv().
Most values of the wqe struct is set except flags. So initialize
flags also.
Fixes: 37cb11acf1f7 ("RDMA/bnxt_re: Add SRQ support for Broadcom adapters")
Signed-off-by: Selvin Xavier <selvin.xavier@broadcom.com>
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
bnxt_re_create_srq() stored attr.max_sge into srq->qplib_srq.max_sge
unvalidated, which defeats the num_sge check in
bnxt_re_post_srq_recv() since that check compares against this same
attacker-chosen value. Reject max_sge > dev_attr->max_srq_sges at
create time, and clamp dev_attr->max_srq_sges to BNXT_STATIC_MAX_SGE
since SRQ WQEs use the fixed-size SGE array.
Fixes: 37cb11acf1f7 ("RDMA/bnxt_re: Add SRQ support for Broadcom adapters")
Signed-off-by: Selvin Xavier <selvin.xavier@broadcom.com>
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Avoid handling wrong sge_len by adding extra check
to see if the passed length is more than the inline
size supported.
Fixes: 1ac5a4047975 ("RDMA/bnxt_re: Add bnxt_re RoCE driver")
Reviewed-by: Kalesh AP <kalesh-anakkur.purayil@broadcom.com>
Signed-off-by: Selvin Xavier <selvin.xavier@broadcom.com>
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
bnxt_re_build_sgl() summed SGE lengths into a signed int, which could
overflow. Make it return u32, and have bnxt_re_copy_wr_payload() report
the size via a u32 out-parameter with the return value carrying only
the error status, instead of overloading a signed int with both.
Fixes: 1ac5a4047975 ("RDMA/bnxt_re: Add bnxt_re RoCE driver")
Signed-off-by: Selvin Xavier <selvin.xavier@broadcom.com>
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Build WCs directly from CQEs and their corresponding published shadow
entries. This avoids iterating over all QPs that share a CQ. Track only
QPs in the error state instead of every QP associated with the CQ.
Keep the explicit GSI send-drain path responsible for adding the QP to
the send error list and invoking its CQ handler. To invoke the handler
for WQEs posted after GSI draining, check the QP state after posting
WQEs. This is required because MANA hardware ignores doorbells for QPs
in the error state.
Introduce a WC-budgeted polling context that caches CQEs representing
multiple WCs.
Before consuming a shadow entry, match the CQE against the low 24 bits
of the posted WQE offset to filter out stale CQEs. Implement WR signaling
to skip WC generation for unsignaled WQEs.
Signed-off-by: Konstantin Taranov <kotaranov@microsoft.com>
Link: https://patch.msgid.link/20260923131155.4055875-6-kotaranov@linux.microsoft.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Track CQ polling credits and ring the doorbell promptly to advance the
hardware CQ tail pointer. The distance between the new and previously
armed tails must remain below four times the CQ length.
Expose CQ arming to mana_ib through mana_gd_wq_ring_doorbell_ext().
RC QPs will also use this helper to ring special doorbell offsets.
MANA hardware ignores duplicate doorbells. Calculate the arming offset
from the previously armed offset and poll_credit to avoid duplicates.
Do not arm the CQ if no CQE was produced, because the CQ remains armed.
Signed-off-by: Konstantin Taranov <kotaranov@microsoft.com>
Link: https://patch.msgid.link/20260923131155.4055875-5-kotaranov@linux.microsoft.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Decouple RQ WQE posting from the QP structure so it can be reused by
other QP types. Batch multiple receive WQEs under a single doorbell.
Introduce GDMA_WR_IB_SGL to pass SGEs directly to GDMA instead of copying
them into a small stack array. Support receive WQEs with num_sge set to
zero.
Signed-off-by: Konstantin Taranov <kotaranov@microsoft.com>
Link: https://patch.msgid.link/20260923131155.4055875-4-kotaranov@linux.microsoft.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Batch multiple send requests under a single doorbell. Move the send PSN
to the common QP structure so it can be shared by all QP types. Define
new WQE formats for UD and RC QPs.
Signed-off-by: Konstantin Taranov <kotaranov@microsoft.com>
Link: https://patch.msgid.link/20260923131155.4055875-3-kotaranov@linux.microsoft.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Pack the posted WQE size and send opcode into compact bitfields. Move the
UD receive shadow fields into the common shadow_wqe_header. This
simplifies the code and lets all QP and WQ types share one structure.
Use release and acquire operations so posting and polling can safely run
on different CPUs.
Signed-off-by: Konstantin Taranov <kotaranov@microsoft.com>
Link: https://patch.msgid.link/20260923131155.4055875-2-kotaranov@linux.microsoft.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
isert_put_unsol_pending_cmds() declares drop_cmd_list as a static
LIST_HEAD, so the list is shared across all invocations of the function.
It is called from isert_wait_conn(), the per-connection .iscsit_wait_conn
transport callback, which runs independently for each connection during
teardown.
conn->cmd_lock only protects each connection's own conn_cmd_list, not the
shared drop_cmd_list. If two connections tear down at the same time, both
threads list_move_tail() commands into the same drop_cmd_list without any
mutual exclusion, and then both iterate and list_del/put it concurrently.
This corrupts the linked list and can double-free iscsit_cmd structures.
Make drop_cmd_list a stack-local list, matching the pattern already used
in isert_free_np().
Fixes: 3e03c4b01da3 ("iser-target: Put the reference on commands waiting for unsol data")
Signed-off-by: Li RongQing <lirongqing@baidu.com>
Link: https://patch.msgid.link/20260921031100.2199-1-lirongqing@baidu.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
In handle_reg_mr_integrity, when mr->pi_mr is NULL (physical-address
optimization path), the code points mr->pi_mr at a stack-local pa_pi_mr
so set_pi_umr_wr can read it during this call. The pointer is never
restored, so after the function returns mr->pi_mr dangles at a freed stack
frame. A second IB_WR_REG_MR_INTEGRITY posted on the same MR without
remapping reads the stale pointer and dereferences it.
Save the original pi_mr (captured at function entry) and restore it on
the out path.
Fixes: 2563e2f30acb ("RDMA/mlx5: Use PA mapping for PI handover")
Signed-off-by: Li RongQing <lirongqing@baidu.com>
Link: https://patch.msgid.link/20260919100848.2442-1-lirongqing@baidu.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
eqe->data.qp_srq.type is a u8. Shifting it left by MLX5_USER_INDEX_LEN
(24) without a cast promotes it to signed int, causing undefined
behaviour for values >= 128.
All currently used type values are well below 128, so this does not
manifest with current hardware. Cast to u32 before the shift to make the
code well-defined across the full range.
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-16-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
create_resource_common() inserted the resource into the radix tree
before its initialization was complete. A firmware EQE in that window
finds a zero refcount_t and calls refcount_inc() on it; the refcount API
treats this as a use-after-free, saturates the counter, and permanently
breaks the resource's lifecycle tracking.
Move the initialization before the radix_tree_insert().
Fixes: e126ba97dba9 ("mlx5: Add driver for Mellanox Connect-IB adapters")
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-15-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Four QP creation paths assigned base->container_mibqp and
base->mqp.event after the firmware create call that inserts the QP into
dev->qp_table.tree. A firmware error EQE in that window reaches
qp->event() with a NULL pointer.
Move both assignments before the firmware call.
Also set ibqp.qp_num inside mlx5_qpc_create_qp() before
create_resource_common() so an EQE arriving early does not observe
qp_num==0.
In addition, add a WARN_ON_ONCE(!qp->event) guard in
rsc_event_notifier() as a safeguard for future regressions.
Fixes: e126ba97dba9 ("mlx5: Add driver for Mellanox Connect-IB adapters")
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-14-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
create_raw_packet_qp() creates a separate firmware RQ object but never
assigned rq->base.mqp.event. Any firmware event on this resource (e.g.
WQ_CATAS_ERROR, WQ_ACCESS_ERROR) calls through a NULL pointer.
Set the event pointer before the RQ is inserted into the resource table.
Fixes: 0fb2ed66a14c ("IB/mlx5: Add create and destroy functionality for Raw Packet QP")
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-13-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
mlx5_cmd_create_srq() stores the SRQ into the xarray before
srq->msrq.event was assigned. A firmware SRQ event arriving in that
window calls through a NULL pointer.
Move the assignment before mlx5_cmd_create_srq().
Fixes: e126ba97dba9 ("mlx5: Add driver for Mellanox Connect-IB adapters")
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-12-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
create_rq() inserts the RQ into dev->qp_table.tree, making it visible to
rsc_event_notifier() before rwq->core_qp.event was assigned. A firmware
WQ_CATAS_ERROR event in that window finds a NULL event pointer.
Move the assignment before create_rq().
Fixes: 350d0e4c7e4b ("IB/mlx5: Track asynchronous events on a receive work queue")
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-11-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
mlx5_ib_dev_res_init() returns -EOPNOTSUPP when the firmware XRC
capability is absent, failing driver probe on XRC-less devices.
Make XRC optional throughout the devr resource path: initialize cq_lock
and srq_lock unconditionally (needed for every QP creation), skip xrcd
allocation and the XRC-type placeholder SRQ (s0) when xrc=0, and guard
their cleanup paths accordingly.
Fixes: f4375443b786 ("RDMA/mlx5: Get XRCD number directly for the internal use")
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-9-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Map non-writable user memory DMA_TO_DEVICE: when no write access is
granted, the NIC only reads the pages and DMA_BIDIRECTIONAL is
unnecessarily broad. Where the IOMMU enforces direction this prevents
the device writing to read-only buffers; where it does not the mapping
is still semantically correct.
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-8-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
For a CQ ring the device writes CQEs and the CPU only reads.
DMA_FROM_DEVICE is the correct direction. Use it explicitly in
ib_umem_get_cq_buf() and ib_umem_get_cq_buf_or_va() rather than deriving
it from the access flags: CQ callers pass IB_ACCESS_LOCAL_WRITE so that
get_user_pages() pins a writable private page (FOLL_WRITE), which the
device needs to DMA-write CQEs into. That flag would yield
DMA_BIDIRECTIONAL from the access-flag derivation, which is wrong for
the bus direction.
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-7-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Add a dma_dir field to struct ib_umem to record the direction chosen at
map time and reuse it at unmap. Thread an explicit dma_data_direction
parameter through the internal pinning helpers so each caller can supply
the direction that matches the device's actual access pattern.
All existing callers still pass DMA_BIDIRECTIONAL, so this is a pure
mechanism addition with no behavioural change.
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-6-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Several drivers pin their CQ ring buffer via ib_umem_get_va() rather
than the CQ-specific helper. Switch them to ib_umem_get_cq_buf_or_va()
so the subsequent DMA_FROM_DEVICE change covers all CQ paths at once.
For shared helpers used by both CQ and non-CQ callers (qedr, mana,
erdma, hns) a new is_cq parameter selects the appropriate pinning
function; non-CQ callers pass false.
vmw_pvrdma's CQ is excluded: it embeds a ring-state header that the
driver CPU-writes after polling, making it genuinely bidirectional.
This is a pure refactor with no functional change.
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-5-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
pvrdma embeds a pvrdma_ring_state header directly inside the QP and SRQ
ring buffers, and the hypervisor writes cons_head into that header.
Pinning with access=0 does not force a private page copy (no
FOLL_WRITE), so the hypervisor may write into a shared page.
Pass IB_ACCESS_LOCAL_WRITE for all three rings. The SRQ ring_state is
not currently used by the driver but is pinned writable for consistency.
Fixes: 29c8d9eba550 ("IB: Add vmw_pvrdma driver")
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-4-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
alloc_cq_buf() leaves user_access zero-initialized, but the device
writes CQEs into the CQ buffer. Set IB_ACCESS_LOCAL_WRITE so the buffer
is pinned writable, matching the device's write access.
Fixes: 744b7bdfa79e ("RDMA/hns: Support 0 hop addressing for CQE buffer")
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-3-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
erdma_init_user_cq() pins the CQ ring with access=0, but the device
writes CQEs into it. Pass IB_ACCESS_LOCAL_WRITE so the buffer is pinned
writable and marked dirty on unpin.
Fixes: 155055771704 ("RDMA/erdma: Add verbs implementation")
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-2-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
irdma_cfg_ceq_vector() registers an IRQ handler before requesting the
CEQ-to-vector mapping through the virtual channel. If the mapping
fails, the function returns without releasing the IRQ resources.
The caller destroys the CEQ, leaving the IRQ handler registered
and potentially referencing freed memory.
Call irdma_destroy_irq() on mapping failure to release
the IRQ resources and drain the associated tasklet
before the caller destroys the CEQ. Use the same dev_id
passed to request_irq(), accounting for CEQ0 sharing
its interrupt vector with the AEQ.
Fixes: b800e82feba7 ("RDMA/irdma: Add GEN3 support for AEQ and CEQ")
Assisted-by: Claude:claude-opus-5
Signed-off-by: Serhat Kumral <serhatkumral1@gmail.com>
Link: https://patch.msgid.link/20260918110428.36586-1-serhatkumral1@gmail.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Commit e91fb8b9d0ed ("RDMA/mlx5: Avoid flexible array warning") split
the leftovers_specs[] array into two separate structs to silence the
sparse "array of flexible structures" warning, but in doing so reversed
the member order: eth_flow was placed before flow_attr.
_create_flow_rule() locates the flow specs with:
ib_flow = (void *)flow_attr + sizeof(*flow_attr);
which expects the spec data to immediately follow the ib_flow_attr.
With the reversed order, flow_attr is the last member, so ib_flow
points past the end of the struct and the num_of_specs = 1 loop reads
out-of-bounds memory (whatever static data follows in the .data
section), producing incorrect flow rules or a crash.
Rebuild the leftovers attributes with DEFINE_RAW_FLEX(), which declares
an on-stack struct ib_flow_attr with a properly sized trailing flows[]
array. The ETH spec lives in flows[0].eth at exactly the offset
_create_flow_rule() expects, so no layout-dependent composite struct is
needed. This also avoids embedding a structure that contains a flexible
array member as a non-last member, which would trigger
-Wflex-array-member-not-at-end once the warning is enabled.
Fixes: e91fb8b9d0ed ("RDMA/mlx5: Avoid flexible array warning")
Signed-off-by: Li RongQing <lirongqing@baidu.com>
Link: https://patch.msgid.link/20260918092337.2338-1-lirongqing@baidu.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Move the kernel QP helpers after the shared memory helpers to remove
forward declarations.
Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com>
Link: https://patch.msgid.link/20260917120459.13286-6-chengyou@linux.alibaba.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Userspace and kernel space queue buffers use separate helpers despite
sharing MTT metadata and lifetime rules. Manage both through
erdma_mem_init() and erdma_mem_uninit(), while keeping backing allocation
and release type-specific.
Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com>
Link: https://patch.msgid.link/20260917120459.13286-5-chengyou@linux.alibaba.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
A single coherent allocation for a kernel CQ can fail when memory is
fragmented.
Use page-sized coherent buffers and describe them with the existing MTT.
Keep the userspace CQ path and doorbell allocation unchanged.
Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com>
Link: https://patch.msgid.link/20260917120459.13286-4-chengyou@linux.alibaba.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
A single coherent allocation for kernel QP queues can fail for large
queues when memory is fragmented.
Allocate page-sized coherent buffers and describe them with the existing
MTT. Keep the userspace QP path unchanged.
Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com>
Link: https://patch.msgid.link/20260917120459.13286-3-chengyou@linux.alibaba.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|