summaryrefslogtreecommitdiff
path: root/drivers/infiniband
AgeCommit message (Collapse)Author
6 daysRDMA/umem: Support an explicit DMA direction other than DMA_BIDIRECTIONALYishai Hadas
Add a dma_dir field to struct ib_umem to record the direction chosen at map time and reuse it at unmap. Thread an explicit dma_data_direction parameter through the internal pinning helpers so each caller can supply the direction that matches the device's actual access pattern. All existing callers still pass DMA_BIDIRECTIONAL, so this is a pure mechanism addition with no behavioural change. Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-6-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
6 daysRDMA/umem: Reuse ib_umem_get_cq_buf_or_va() for VA-only CQ pinningYishai Hadas
Several drivers pin their CQ ring buffer via ib_umem_get_va() rather than the CQ-specific helper. Switch them to ib_umem_get_cq_buf_or_va() so the subsequent DMA_FROM_DEVICE change covers all CQ paths at once. For shared helpers used by both CQ and non-CQ callers (qedr, mana, erdma, hns) a new is_cq parameter selects the appropriate pinning function; non-CQ callers pass false. vmw_pvrdma's CQ is excluded: it embeds a ring-state header that the driver CPU-writes after polling, making it genuinely bidirectional. This is a pure refactor with no functional change. Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-5-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
6 daysRDMA/vmw_pvrdma: Pin QP and SRQ rings writable to match device DMA write accessYishai Hadas
pvrdma embeds a pvrdma_ring_state header directly inside the QP and SRQ ring buffers, and the hypervisor writes cons_head into that header. Pinning with access=0 does not force a private page copy (no FOLL_WRITE), so the hypervisor may write into a shared page. Pass IB_ACCESS_LOCAL_WRITE for all three rings. The SRQ ring_state is not currently used by the driver but is pinned writable for consistency. Fixes: 29c8d9eba550 ("IB: Add vmw_pvrdma driver") Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-4-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
6 daysRDMA/hns: Pin CQ buffer writable to match device DMA write accessYishai Hadas
alloc_cq_buf() leaves user_access zero-initialized, but the device writes CQEs into the CQ buffer. Set IB_ACCESS_LOCAL_WRITE so the buffer is pinned writable, matching the device's write access. Fixes: 744b7bdfa79e ("RDMA/hns: Support 0 hop addressing for CQE buffer") Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-3-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
6 daysRDMA/erdma: Pin CQ buffer writable to match device DMA write accessYishai Hadas
erdma_init_user_cq() pins the CQ ring with access=0, but the device writes CQEs into it. Pass IB_ACCESS_LOCAL_WRITE so the buffer is pinned writable and marked dirty on unpin. Fixes: 155055771704 ("RDMA/erdma: Add verbs implementation") Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Link: https://patch.msgid.link/20260915140933.40580-2-yishaih@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
6 daysRDMA/irdma: Free IRQ when CEQ vector mapping failsSerhat Kumral
irdma_cfg_ceq_vector() registers an IRQ handler before requesting the CEQ-to-vector mapping through the virtual channel. If the mapping fails, the function returns without releasing the IRQ resources. The caller destroys the CEQ, leaving the IRQ handler registered and potentially referencing freed memory. Call irdma_destroy_irq() on mapping failure to release the IRQ resources and drain the associated tasklet before the caller destroys the CEQ. Use the same dev_id passed to request_irq(), accounting for CEQ0 sharing its interrupt vector with the AEQ. Fixes: b800e82feba7 ("RDMA/irdma: Add GEN3 support for AEQ and CEQ") Assisted-by: Claude:claude-opus-5 Signed-off-by: Serhat Kumral <serhatkumral1@gmail.com> Link: https://patch.msgid.link/20260918110428.36586-1-serhatkumral1@gmail.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
11 daysRDMA/mlx5: Use DEFINE_RAW_FLEX for leftovers flow attributesLi RongQing
Commit e91fb8b9d0ed ("RDMA/mlx5: Avoid flexible array warning") split the leftovers_specs[] array into two separate structs to silence the sparse "array of flexible structures" warning, but in doing so reversed the member order: eth_flow was placed before flow_attr. _create_flow_rule() locates the flow specs with: ib_flow = (void *)flow_attr + sizeof(*flow_attr); which expects the spec data to immediately follow the ib_flow_attr. With the reversed order, flow_attr is the last member, so ib_flow points past the end of the struct and the num_of_specs = 1 loop reads out-of-bounds memory (whatever static data follows in the .data section), producing incorrect flow rules or a crash. Rebuild the leftovers attributes with DEFINE_RAW_FLEX(), which declares an on-stack struct ib_flow_attr with a properly sized trailing flows[] array. The ETH spec lives in flows[0].eth at exactly the offset _create_flow_rule() expects, so no layout-dependent composite struct is needed. This also avoids embedding a structure that contains a flexible array member as a non-last member, which would trigger -Wflex-array-member-not-at-end once the warning is enabled. Fixes: e91fb8b9d0ed ("RDMA/mlx5: Avoid flexible array warning") Signed-off-by: Li RongQing <lirongqing@baidu.com> Link: https://patch.msgid.link/20260918092337.2338-1-lirongqing@baidu.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
11 daysRDMA/erdma: Move kernel QP helpers after memory helpersCheng Xu
Move the kernel QP helpers after the shared memory helpers to remove forward declarations. Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com> Link: https://patch.msgid.link/20260917120459.13286-6-chengyou@linux.alibaba.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
11 daysRDMA/erdma: Unify userspace and kernel queue buffer managementCheng Xu
Userspace and kernel space queue buffers use separate helpers despite sharing MTT metadata and lifetime rules. Manage both through erdma_mem_init() and erdma_mem_uninit(), while keeping backing allocation and release type-specific. Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com> Link: https://patch.msgid.link/20260917120459.13286-5-chengyou@linux.alibaba.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
11 daysRDMA/erdma: Support non-contiguous kernel CQ buffersCheng Xu
A single coherent allocation for a kernel CQ can fail when memory is fragmented. Use page-sized coherent buffers and describe them with the existing MTT. Keep the userspace CQ path and doorbell allocation unchanged. Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com> Link: https://patch.msgid.link/20260917120459.13286-4-chengyou@linux.alibaba.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
11 daysRDMA/erdma: Support non-contiguous kernel QP buffersCheng Xu
A single coherent allocation for kernel QP queues can fail for large queues when memory is fragmented. Allocate page-sized coherent buffers and describe them with the existing MTT. Keep the userspace QP path unchanged. Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com> Link: https://patch.msgid.link/20260917120459.13286-3-chengyou@linux.alibaba.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
11 daysRDMA/erdma: Unwind kernel QP initialization failuresCheng Xu
Allocation failures currently use the normal kernel QP teardown path. Unwind failures directly so free_kernel_qp() only handles complete QPs. Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com> Link: https://patch.msgid.link/20260917120459.13286-2-chengyou@linux.alibaba.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
11 daysRDMA/mlx5: Guard steering anchor resource destruction on create errorLi RongQing
The steering anchor resources are shared between anchors with the same priority and are protected by rule_goto_table_ref. When steering anchor creation succeeds but uverbs_copy_to() fails, the error path drops the reference acquired for the new anchor. However, it unconditionally calls mlx5_steering_anchor_destroy_res(), even when other anchors are still using the shared resources. This can destroy the shared flow table, flow groups, and rules while existing anchors still hold references to them, resulting in a use-after-free when those anchors are subsequently cleaned up. Only destroy the shared resources when dropping the reference makes rule_goto_table_ref zero, consistent with the existing cleanup path. Fixes: e1f4a52ac171 ("RDMA/mlx5: Create an indirect flow table for steering anchor") Signed-off-by: Li RongQing <lirongqing@baidu.com> Link: https://patch.msgid.link/20260917085419.1792-1-lirongqing@baidu.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
11 daysRDMA/multicast: Fix rb-tree corruption on MGID reassignment in join_handlerLi RongQing
When the SA assigns a wildcard (mgid0) join a concrete MGID that already has a group in the port table, join_handler() erases the group from the tree and calls mcast_insert() with allow_duplicates=false. mcast_insert() then finds the existing group and returns it without re-inserting, but join_handler() discards the return value, leaving the group erased yet still alive. When the orphaned group's refcount later drops to zero, release_group() calls rb_erase() on the node that is no longer in the tree, corrupting the rb-tree (a stale parent pointer causes a valid subtree link to be overwritten with NULL). Clear the node with RB_CLEAR_NODE() after the erase in join_handler() so release_group() can detect via RB_EMPTY_NODE() that it is already out of the tree and skip the second erase. A successful re-insert by mcast_insert() restores the in-tree state via rb_link_node(). Signed-off-by: Li RongQing <lirongqing@baidu.com> Link: https://patch.msgid.link/20260917052012.2181-1-lirongqing@baidu.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
11 daysRDMA/siw: Fix qp use-after-free in siw_get_base_qp()Wentao Liang
The reference taken by siw_qp_id2obj() is dropped before the address of the embedded ib_qp is taken from the QP. If that reference was the last one the QP is freed and the pointer returned to the iwarp core is dangling; the core only pins the QP with iw_add_ref() after the call has returned. Copy the pointer while the reference is still held. Fixes: 6c52fdc244b5 ("rdma/siw: connection management") Signed-off-by: Wentao Liang <vulab@iscas.ac.cn> Link: https://patch.msgid.link/20260916184400.2093106-1-vulab@iscas.ac.cn Signed-off-by: Leon Romanovsky <leon@kernel.org>
11 daysRDMA/mlx5: Fix RQ resource reference leak in mlx5_ib_wq_event()Wentao Liang
Since commit 312b8f79eb05 ("RDMA/mlx: Calling qp event handler in workqueue context") rsc_event_notifier() expects the event handler to drop the resource reference taken by mlx5_get_rsc(). mlx5_ib_wq_event() was missed by that conversion, so every RQ event leaks one reference, which keeps destroy_resource_common() waiting on the completion. Put the resource reference on all exits of the handler. Fixes: 312b8f79eb05 ("RDMA/mlx: Calling qp event handler in workqueue context") Signed-off-by: Wentao Liang <vulab@iscas.ac.cn> Link: https://patch.msgid.link/20260916183921.2092821-1-vulab@iscas.ac.cn Signed-off-by: Leon Romanovsky <leon@kernel.org>
11 daysRDMA/cxgb4: Fix skb reference leak in rx_pkt()Wentao Liang
rx_pkt() takes an extra reference on the received skb so it can be handed to the firmware completion path via req->cookie. When the work request skb allocation in send_fw_pass_open_req() fails the function returns without releasing that extra reference, leaking one skb reference. Free the packet skb before returning when the allocation fails, matching the cxgb4_ofld_send() failure path. Fixes: 1cab775c3e75 ("RDMA/cxgb4: Fix LE hash collision bug for passive open connection") Signed-off-by: Wentao Liang <vulab@iscas.ac.cn> Link: https://patch.msgid.link/20260916183715.2092663-1-vulab@iscas.ac.cn Signed-off-by: Leon Romanovsky <leon@kernel.org>
11 daysRDMA/cxgb4: Fix neigh reference leak in rx_pkt()Wentao Liang
If ip_dev_find() fails on the loopback path the code jumps to free_dst without releasing the neighbour reference taken by dst_neigh_lookup_skb(), leaking one neigh reference. Release the neighbour before jumping to free_dst. Fixes: ef42520240aa ("RDMA/cxgb4: add null-ptr-check after ip_dev_find()") Signed-off-by: Wentao Liang <vulab@iscas.ac.cn> Link: https://patch.msgid.link/20260916183530.2092534-1-vulab@iscas.ac.cn Signed-off-by: Leon Romanovsky <leon@kernel.org>
11 daysRDMA/irdma: check vport device allocationSlavin Liu
An existing RDMA function does not guarantee allocation of the new vport device. Return -ENOMEM before initializing a NULL iwdev. Detected by static analysis and reviewed with AI-assisted source auditing. Fixes: 2ad49ae7330b ("RDMA/irdma: Introduce GEN3 vPort driver support") Assisted-by: LLM Signed-off-by: Slavin Liu <bolin.liu@seu.edu.cn> Link: https://patch.msgid.link/20260911060909.94219-1-bolin.liu@seu.edu.cn Signed-off-by: Leon Romanovsky <leon@kernel.org>
11 daysRDMA/irdma: Use kvzalloc for lvl2 leafmem allocationJeremy Bongio
During VM live migration or when registering large memory regions (e.g., guest memory >= 256 GB backed by 4 KiB pages), get_lvl2_pble() calculates total leaves exceeding ~131,072. The memory required for lvl2->leafmem.va exceeds 4 MiB (sizeof(*leaf) * total > 5.24 MiB). Because kzalloc() requires physically contiguous pages, requests larger than MAX_PAGE_ORDER (4 MiB on x86_64) fail and trigger a warning: WARNING: CPU: 123 PID: 119906 at mm/page_alloc.c:5997 __alloc_frozen_pages_noprof+0x50c/0x1050 Call Trace: <TASK> alloc_frozen_pages_noprof+0xc0/0x240 ___kmalloc_large_node+0x3d/0xd0 __kmalloc_large_node_noprof+0x1b/0x90 __kmalloc_noprof+0x254/0x660 get_lvl1_lvl2_pble+0xbc/0x270 [irdma] irdma_get_pble+0x44/0xc0 [irdma] irdma_setup_pbles+0x46/0x1d0 [irdma] irdma_reg_user_mr_type_mem+0x41/0x230 [irdma] irdma_reg_user_mr+0x194/0x1c0 [irdma] ib_uverbs_reg_mr+0x1b0/0x2c0 [ib_uverbs] The leafmem array is purely host software tracking structures (struct irdma_pble_info) and does not require physically contiguous memory. Switch the allocation from kzalloc() to kvzalloc() and free it with kvfree() so that requests exceeding MAX_PAGE_ORDER fall back to virtually contiguous memory via vmalloc. Fixes: e8c4dbc2fcac ("RDMA/irdma: Add PBLE resource manager") Assisted-by: LLM Signed-off-by: Jeremy Bongio <jbongio@google.com> Link: https://patch.msgid.link/20260916173203.2555458-1-jbongio@google.com Tested-by: Jacob Moroni <jmoroni@google.com> Signed-off-by: Leon Romanovsky <leon@kernel.org>
12 daysRDMA/irdma: Fix erroneous -ENOMEM for large MR registrationJacob Moroni
Commit cc8997c94bf3 ("RDMA/irdma: Refactor PBLE functions") replaced the explicit level1_only boolean with a bit mask but did not properly convert the "if (level1_only)" check in irdma_get_pble() to be "lvl == PBLE_LEVEL_1" and instead left it as "if (lvl)". User MR registration passes "lvl = PBLE_LEVEL_1 | PBLE_LEVEL_2" so this check always evaluates to true and causes the loop to bail after grabbing a single SD even if the region actually requires more than one SD (as is the case for regions larger than 1 GiB with 4k pages). As a result, these registrations can fail with -ENOMEM even if there are sufficient PBLE resources. Fixes: cc8997c94bf3 ("RDMA/irdma: Refactor PBLE functions") Signed-off-by: Jacob Moroni <jmoroni@google.com> Link: https://patch.msgid.link/20260915170317.2901672-1-jmoroni@google.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
12 daysRDMA/core: Force disconnect if DREP and DREQ failAmmar Ratnani
When we call `rdma_disconnect`, it's possible that we fail to send both a DREQ and a DREP. In that case, we currently either do nothing or move into the `IB_CM_TIMEWAIT` state. Either way, we remain in the `RDMA_CM_CONNECT` state. Anyone waiting for us to disconnect hangs. Instead, while we're connected, when we call `rdma_disconnect`, and fail to send both a DREQ and a DREP, mark the connection as disconnected. Call the upper layer's handler for the disconnect, and destroy the connection if instructed. Cc: Leon Romanovsky <leon@kernel.org> Cc: Jason Gunthorpe <jgg@ziepe.ca> Cc: Namjae Jeon <linkinjeon@kernel.org> Cc: Paulo Alcantara <pc@manguebit.org> Cc: Stefan Metzmacher <metze@samba.org> Cc: Tom Talpey <tom@talpey.com> Cc: linux-rdma@vger.kernel.org Cc: linux-cifs@vger.kernel.org Cc: samba-technical@lists.samba.org Cc: linux-kernel@vger.kernel.org Signed-off-by: Ammar Ratnani <ammrat13@gmail.com> Link: https://patch.msgid.link/20260908154454.10966-2-ammrat13@gmail.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
12 daysRDMA/hns: Use unsigned comparison in the CQ cleanup loopYishai Hadas
__hns_roce_v2_cq_clean() sweeps the CQ backwards from the producer index down to the consumer index: while ((int) --prod_index - (int) hr_cq->cons_index >= 0) Both indexes are free running u32 counters, so the comparison has to be done modulo 2^32. Casting each operand to int and subtracting does not do that: the subtraction overflows whenever the two indexes straddle 2^31, which is undefined behaviour, and a compiler that assumes signed overflow cannot occur is free to discard the subtraction and fold the expression into a plain signed comparison. That comparison is not wraparound safe. The kernel is built with -fno-strict-overflow, so gcc and clang both retain the subtraction today and the generated code is unaffected; there is no known user-visible impact from the current code. Still, correctness here shouldn't depend on that build flag. Replace the loop with a plain unsigned equality check instead: while (prod_index != hr_cq->cons_index) { --prod_index; ... The preceding forward scan starts prod_index at cons_index and advances it only while get_sw_cqe_v2() still reports a software owned entry, so the two indexes stay within one CQ's worth of entries. Decrementing prod_index reaches cons_index in exactly that many iterations regardless of whether either counter has wrapped, because the loop no longer compares magnitudes at all. Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Signed-off-by: Edward Srouji <edwards@nvidia.com> Link: https://patch.msgid.link/20260915-fix-cq-cleanup-v1-4-e991944cf898@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
12 daysRDMA/mthca: Use unsigned comparison in the CQ cleanup loopYishai Hadas
mthca_cq_clean() sweeps the CQ backwards from the producer index down to the consumer index: while ((int) --prod_index - (int) cq->cons_index >= 0) Both indexes are free running u32 counters, so the comparison has to be done modulo 2^32. Casting each operand to int and subtracting does not do that: the subtraction overflows whenever the two indexes straddle 2^31, which is undefined behaviour, and a compiler that assumes signed overflow cannot occur is free to discard the subtraction and fold the expression into a plain signed comparison. That comparison is not wraparound safe. The kernel is built with -fno-strict-overflow, so gcc and clang both retain the subtraction today and the generated code is unaffected; there is no known user-visible impact from the current code. Still, correctness here shouldn't depend on that build flag. Replace the loop with a plain unsigned equality check instead: while (prod_index != cq->cons_index) { --prod_index; ... The preceding forward scan starts prod_index at cons_index and stops no later than cons_index + cq->ibcq.cqe, so the two indexes are at most cq->ibcq.cqe apart. Decrementing prod_index reaches cons_index in exactly that many iterations regardless of whether either counter has wrapped, because the loop no longer compares magnitudes at all. Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Signed-off-by: Edward Srouji <edwards@nvidia.com> Link: https://patch.msgid.link/20260915-fix-cq-cleanup-v1-3-e991944cf898@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
12 daysRDMA/mlx4: Use unsigned comparison in the CQ cleanup loopYishai Hadas
__mlx4_ib_cq_clean() sweeps the CQ backwards from the producer index down to the consumer index: while ((int) --prod_index - (int) cq->mcq.cons_index >= 0) Both indexes are free running u32 counters, so the comparison has to be done modulo 2^32. Casting each operand to int and subtracting does not do that: the subtraction overflows whenever the two indexes straddle 2^31, which is undefined behaviour, and a compiler that assumes signed overflow cannot occur is free to discard the subtraction and fold the expression into a plain signed comparison. That comparison is not wraparound safe. The kernel is built with -fno-strict-overflow, so gcc and clang both retain the subtraction today and the generated code is unaffected; there is no known user-visible impact from the current code. Still, correctness here shouldn't depend on that build flag. Replace the loop with a plain unsigned equality check instead: while (prod_index != cq->mcq.cons_index) { --prod_index; ... The preceding forward scan starts prod_index at cons_index and stops no later than cons_index + cq->ibcq.cqe, so the two indexes are at most cq->ibcq.cqe apart. Decrementing prod_index reaches cons_index in exactly that many iterations regardless of whether either counter has wrapped, because the loop no longer compares magnitudes at all. Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Signed-off-by: Edward Srouji <edwards@nvidia.com> Link: https://patch.msgid.link/20260915-fix-cq-cleanup-v1-2-e991944cf898@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
12 daysRDMA/mlx5: Use unsigned comparison in the CQ cleanup loopYishai Hadas
__mlx5_ib_cq_clean() sweeps the CQ backwards from the producer index down to the consumer index: while ((int) --prod_index - (int) cq->mcq.cons_index >= 0) Both indexes are free running u32 counters, so the comparison has to be done modulo 2^32. Casting each operand to int and subtracting does not do that: the subtraction overflows whenever the two indexes straddle 2^31, which is undefined behaviour, and a compiler that assumes signed overflow cannot occur is free to discard the subtraction and fold the expression into a plain signed comparison. That comparison is not wraparound safe. The kernel is built with -fno-strict-overflow, so gcc and clang both retain the subtraction today and the generated code is unaffected; there is no known user-visible impact from the current code. Still, correctness here shouldn't depend on that build flag. Replace the loop with a plain unsigned equality check instead. while (prod_index != cq->mcq.cons_index) { --prod_index; ... The preceding forward scan starts prod_index at cons_index and stops no later than cons_index + cq->ibcq.cqe, so the two indexes are at most cq->ibcq.cqe apart. Decrementing prod_index reaches cons_index in exactly that many iterations regardless of whether either counter has wrapped, because the loop no longer compares magnitudes at all. Signed-off-by: Yishai Hadas <yishaih@nvidia.com> Signed-off-by: Edward Srouji <edwards@nvidia.com> Link: https://patch.msgid.link/20260915-fix-cq-cleanup-v1-1-e991944cf898@nvidia.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
12 daysRDMA/irdma: avoid use-after-free in icrdma_removeGuangshuo Li
icrdma_remove() dereferences iwdev after calling irdma_ib_unregister_device(): irdma_ib_unregister_device(iwdev) -> ib_unregister_device(&iwdev->ibdev) -> __ib_unregister_device() -> ib_dealloc_device() irdma registers irdma_ib_dealloc_device() as the dealloc_driver callback in its ib_device_ops. The RDMA core explicitly documents ib_unregister_device() that when ops.dealloc_driver is used, ib_dev will be freed upon return from the function. The ib_device is embedded in struct irdma_device and iwdev itself was allocated by ib_alloc_device(irdma_device, ibdev). Consequently, once irdma_ib_unregister_device() returns, iwdev must no longer be dereferenced. However, icrdma_remove() currently accesses iwdev->rf afterwards when deinitializing interrupts, destroying ah_tbl_lock, and freeing rf. Although the final ib_device storage is released with kfree_rcu(), the object lifetime has already ended and callers must not rely on the RCU grace period to continue dereferencing iwdev. Save iwdev->rf before unregistering the RDMA device and use the saved pointer for the remaining cleanup. This issue was found by manual code inspection. Fixes: 8498a30e1b94 ("RDMA/irdma: Register auxiliary driver and implement private channel OPs") Cc: stable@vger.kernel.org Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com> Link: https://patch.msgid.link/20260914122840.1709743-1-lgs201920130244@gmail.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
12 daysRDMA/core: Fix partial copy of IPv6 flow_lbl in LAG hash skbLi RongQing
rdma_build_skb() constructs a synthetic IPv6 header for LAG slave selection and copies the flow label from the AH attribute: memcpy(&ip6h->flow_lbl, &ah_attr->grh.flow_label, sizeof(*ip6h->flow_lbl)); ipv6hdr.flow_lbl is __u8[3], so sizeof(*ip6h->flow_lbl) dereferences the array to a single __u8 and yields 1. The memcpy therefore copies only the first byte of the flow label, leaving flow_lbl[1] and flow_lbl[2] uninitialised in the skb buffer (alloc_skb() does not zero the data). The resulting LAG hash is computed over one byte of real flow-label data and two bytes of kmalloc residue, so IPv6 RoCEv2 traffic on a bonded interface can be steered to the wrong slave. Use sizeof(ip6h->flow_lbl) (the array, 3 bytes) so the whole flow label is copied. Fixes: bd3920eac1331 ("RDMA/core: Add LAG functionality") Signed-off-by: Li RongQing <lirongqing@baidu.com> Link: https://patch.msgid.link/20260916120926.1761-1-lirongqing@baidu.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2026-09-18RDMA/uverbs: Drop restrack ref on ib_init_ucontext() failure in GET_CONTEXTLi RongQing
The UVERBS_METHOD_GET_CONTEXT handler allocates the ucontext via ib_alloc_ucontext(), which calls rdma_restrack_new() (initialising the kref to 1) and rdma_restrack_set_name(NULL), the latter attaching the current task and taking a task_struct reference. When ib_init_ucontext() subsequently fails, the error path only does: kfree(attrs->context); attrs->context = NULL; without first calling rdma_restrack_put() on &attrs->context->res. The kref therefore never reaches zero, restrack_release() is never invoked, and put_task_struct() is never called, so the task_struct reference taken during set_name is leaked permanently. A local user with access to an RDMA device can repeat this path (e.g. by hitting an RDMA cgroup limit or supplying invalid ucaps) and leak one task_struct reference per attempt, eventually preventing those processes from being reaped or exhausting memory. The legacy ib_uverbs_get_context() path already handles this correctly at its err_ucontext label, where rdma_restrack_put() precedes kfree(). Mirror that ordering here. Fixes: a1123418ba10 ("RDMA/uverbs: Add ioctl command to get a device context") Signed-off-by: Li RongQing <lirongqing@baidu.com> Link: https://patch.msgid.link/20260916085529.1838-1-lirongqing@baidu.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2026-09-18RDMA/restrack: Don't set RESTRACK_DD mark after a failed xa_insert()Li RongQing
When adding a QP to the restrack xarray, rdma_restrack_add() does: ret = xa_insert(&rt->xa, res->id, res, GFP_KERNEL); if (ret) res->id = 0; if (qp->qp_type >= IB_QPT_DRIVER) xa_set_mark(&rt->xa, res->id, RESTRACK_DD); The xa_set_mark() call is not guarded by "!ret". When xa_insert() fails (possible on a qp_num collision or memory pressure), res->id is reset to 0, yet the mark is still applied to index 0. If a legitimate QP whose qp_num is 0 already occupies that slot - e.g. a normal QP or the SMI QP - it is spuriously tagged as driver-private (RESTRACK_DD), and the netlink dump path res_get_common_dumpit() then hides it from "rdma res" output whenever the caller does not request driver details. Only set the mark when the insertion actually succeeded by folding the mark into the non-failure branch. Fixes: e18fa0bbcedf8 ("RDMA/core: Add an option to display driver-specific QPs in the rdmatool") Signed-off-by: Li RongQing <lirongqing@baidu.com> Link: https://patch.msgid.link/20260916073042.1746-1-lirongqing@baidu.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2026-09-18RDMA/core: Fix swapped list_add_tail() arguments in ib_add_sub_device()Li RongQing
ib_add_sub_device() links a new sub-device into its parent's sub-device list with: list_add_tail(&parent->subdev_list_head, &sub->subdev_list); list_add_tail(new, head) expects the node to insert as the first argument and the list head as the second, so the call above does the opposite of what was intended: it treats the parent's list head as the new node and the sub-device's node as the list head. Because _ib_alloc_device() initialises subdev_list to a self-referencing empty head, the misplaced insertion does not crash, but it corrupts the parent->subdev_list_head list. After adding two or more sub-devices, only the last one is reachable through the parent's head, and ib_del_sub_device_and_put()'s list_del() on &sub->subdev_list rewrites the head pointers, potentially severing earlier sub-devices from the chain. During parent teardown, the reverse iteration over subdev_list_head then skips the orphaned sub-devices, so their del_sub_dev() callbacks are never invoked and ib_device_put() on the parent is never paired, leaking both the HW sub-devices and a refcount. Swap the arguments to match the documented intent: insert the sub-device node into the parent's list head. Fixes: bca51197620a ("RDMA/core: Support IB sub device with type "SMI"") Signed-off-by: Li RongQing <lirongqing@baidu.com> Link: https://patch.msgid.link/20260916064435.2180-1-lirongqing@baidu.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2026-09-17Merge git://git.kernel.org/pub/scm/linux/kernel/git/netdev/netJakub Kicinski
Cross-merge networking fixes after downstream PR (net-7.3-rc4). Conflicts: net/core/neighbour.c 979aabdad8dd0 ("neighbour: Skip default parms when resumed in neightbl_dump_info().") 7b430fcfc972f ("neighbour: Don't render blackhole_netdev via RTM_GETNEIGHTBL.") fae1c59810b86 ("neighbour: Remove unnecessary net_eq().") https://lore.kernel.org/20260911173056.44ec06e0@kernel.org https://lore.kernel.org/aqfbJi7nAX4IbmnR@sirena.co.uk Adjacent changes: net/netlink/af_netlink.c ceac0de741bf ("netlink: do not free nlk->groups while lockless readers can use it") 7c0ec6288b49 ("net: Replace %pK output with 0") net/bridge/br_vlan.c 2842ce397dd0 ("net: bridge: vlan: fix bugs caused by switchdev deletion errors") 5bec8f861114 ("net: bridge: vlan: annotate lockless use of num_vlans") 2b1f8fd3118c ("net: bridge: vlan: annotate lockless vlan flags use") net/bridge/br_mst.c 18a6fe05fb6e ("net: bridge: mst: move switchdev call outside rcu") 120207a08fc0 ("net: bridge: vlan: annotate lockless use of msti") drivers/net/ethernet/stmicro/stmmac/hwif.h 90e4b849dfa6 ("net: stmmac: propagate FPE preemption-class mapping errors") 85ca3292d7a3 ("net: stmmac: Remove ARP offload code") Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15RDMA/hns: Fix the use of uninitialized variableJunxian Huang
The variable 'i' is used without being initialized in active-backup mode. Fixes: d9023e461b73 ("RDMA/hns: Implement bonding init/uninit process") Signed-off-by: Junxian Huang <huangjunxian6@hisilicon.com> Link: https://patch.msgid.link/20260914024800.132429-1-huangjunxian6@hisilicon.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2026-09-15RDMA/srpt: Fix srp_sq_size documentationSerhat Kumral
Replace "Shared receive queue (SRQ)" with "Send queue". Signed-off-by: Serhat Kumral <serhatkumral1@gmail.com> Link: https://patch.msgid.link/20260914130536.19769-1-serhatkumral1@gmail.com Reviewed-by: Bart Van Assche <bvanassche@acm.org> Signed-off-by: Leon Romanovsky <leon@kernel.org>
2026-09-15RDMA/mlx5: Print err code when create_qp failsLi RongQing
Add the actual error code to the "Create QP type %d failed" log line in create_qp(), so failures can be diagnosed from dmesg alone without guessing the underlying error. Signed-off-by: Li RongQing <lirongqing@baidu.com> Link: https://patch.msgid.link/20260914121253.2193-1-lirongqing@baidu.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2026-09-15RDMA/uverbs: Make CQ handle mandatory for WQ creationLi RongQing
UVERBS_ATTR_CREATE_WQ_CQ_HANDLE is declared UA_OPTIONAL in the ioctl method definition, so the mandatory attribute bitmap does not enforce its presence. When userspace omits it, uverbs_attr_get_obj() returns ERR_PTR(-ENOENT) and the handler stores that error pointer into wq_init_attr.cq without validation. The bogus cq pointer is then passed to the driver's create_wq callback. In the mlx5 case, create_rq() calls to_mcq(init_attr->cq) which applies container_of to the ERR_PTR value, producing a near-NULL pointer. The subsequent access in get_rq_ts_format() triggers a kernel NULL pointer dereference: BUG: kernel NULL pointer dereference, address: 0000000000000296 RIP: 0010:create_rq+0x32/0x550 [mlx5_ib] Call Trace: mlx5_ib_create_wq+0x14a/0x210 [mlx5_ib] ib_uverbs_handler_UVERBS_METHOD_WQ_CREATE+0x1f0/0x320 [ib_uverbs] ib_uverbs_run_method+0x296/0x320 [ib_uverbs] ib_uverbs_cmd_verbs+0x1a0/0x260 [ib_uverbs] ib_uverbs_ioctl+0xa8/0x120 [ib_uverbs] A WQ without a CQ was never valid; the legacy write path always required one via uobj_get_obj_read() in ib_uverbs_ex_create_wq(). Declare the attribute UA_MANDATORY so the uverbs framework rejects the ioctl early when the CQ handle is missing, before the handler ever runs. Fixes: ef3bc084a8ed ("IB/uverbs: Introduce create/destroy WQ commands over ioctl") Signed-off-by: Li RongQing <lirongqing@baidu.com> Link: https://patch.msgid.link/20260911021557.2113-1-lirongqing@baidu.com Reviewed-by: Kalesh AP <kalesh-anakkur.purayil@broadcom.com> Signed-off-by: Leon Romanovsky <leon@kernel.org>
2026-09-15RDMA/rxe: Use validated num_sge in local bufferNicolas Morey
For both SRQ and non-SRQ receive paths, the WQE is copied into a local buffer to provide a kernel-owned, validated copy. While calculating the memcpy size from the validated num_sge prevents overflow during the copy, memcpy() itself still copies num_sge from shared memory. A concurrent userspace modification before or during memcpy() leaves an unvalidated num_sge in the local buffer, leading to potential out-of-bounds reads in rxe_resp_check_length() and copy_data() causing: BUG: KASAN: slab-out-of-bounds in rxe_receiver+0x8109/0x9ec0 [rdma_rxe] Read of size 4 at addr ffff88812c4867f8 by task kworker/u9:6/361 Workqueue: rxe_wq do_work [rdma_rxe] Call Trace: rxe_receiver+0x8109/0x9ec0 [rdma_rxe] do_work+0x149/0x610 [rdma_rxe] process_one_work+0x726/0x10a0 The buggy address belongs to the object at ffff88812c486000 which belongs to the cache kmalloc-part-13-2k of size 2048 The buggy address is located 0 bytes to the right of allocated 2040-byte region [ffff88812c486000, ffff88812c4867f8) Explicitly assign the validated num_sge to the local buffer after the copy to prevent this race. Fixes: 22b8fbded65b ("RDMA/rxe: Fix TOCTOU heap overflow in get_srq_wqe") Fixes: d6ab440240a0 ("RDMA/rxe: Copy WQE to local buffer in non-SRQ receive path") Reviewed-by: Zhu Yanjun <yanjun.zhu@linux.dev> Signed-off-by: Nicolas Morey <nmorey@suse.com> Link: https://patch.msgid.link/20260909160132.1491248-1-nmorey@suse.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2026-09-14Merge tag 'for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/rdma/rdmaLinus Torvalds
Pull rdma fixes from Jason Gunthorpe: "Lots of bug fixes from the last weeks: - Various error unwind bugs - Several more races and bugs in siw and rxe, including remote triggerable - HFI1 corruption with its credit scheme - Remove a bogus user triggerable dev_warn - Lock __ethtool_get_link_ksettings() properly - Fix a lockdep loop with diassociation - Several storage related bugs, some triggerable remotely - Do no leak physical addresses to userspace in bnxt_re - Fix wrong irq context for the xarrays in erdma - User triggerable race in ucma with multicast - Race in ipoib with multicast flushing and destruction" * tag 'for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/rdma/rdma: (28 commits) RDMA/siw: Bound fragmented header copies by the remaining length RDMA/efa: Keep EQ resources alive while IRQ is registered RDMA/efa: Keep admin queues alive while IRQ is registered RDMA/core: fix refcount bug in iwpm_get_nlmsg_request() IB/IPoIB: Avoid restoring OPER_UP after multicast flush RDMA/ucma: Serialize join and leave on copy_to_user failure RDMA/rtrs-clt: Fix CQ pool leak when connect is interrupted RDMA/irdma: Enforce local fence for IB_WR_REG_MR RDMA/erdma: Use IRQ-safe XArray helpers for QP and CQ tables RDMA/mad: Fix receive buffer leak when PKey enforcement fails RDMA/uverbs: Fix potential leak of resources->collection in flow_resources_alloc() RDMA/bnxt_re: Avoid exposing umdbr to userspace RDMA/rtrs: guard against null kobj name RDMA/bnxt_re: check create_singlethread_workqueue() in DCB setup IB/isert: wait for deferred control PDU completions before releasing the connection IB/iser: reject a remote invalidation of an unregistered direction RDMA/srp: Fix srp_remove_target() IB/mlx4: Fix use-after-free on pkey sysfs registration failure RDMA/uverbs: Fix mmap_lock/disassociation_lock circular dependency RDMA/core: Reject unregistering netdevs in ib_get_eth_speed ...
2026-09-10Merge git://git.kernel.org/pub/scm/linux/kernel/git/netdev/netJakub Kicinski
Cross-merge networking fixes after downstream PR (net-7.3-rc3). Conflicts: drivers/net/dsa/mt7530.c 3c18e3c9a54e ("net: dsa: mt7530: populate lpi_interfaces to fix EEE support") 10d9d8328e8a ("net: dsa: mt7530: replace mt7530_read with regmap_read") Adjacent changes: drivers/net/bonding/bond_alb.c 1746ef2e2df2 ("bonding: use skb_cow_head() in bond_do_alb_xmit() and rlb_arp_xmit()") 4cef95f72bbd ("bonding: fix u32 overflow in compute_gap()") Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-10RDMA/siw: Bound fragmented header copies by the remaining lengthJérémy Jean
siw_get_hdr() can receive an extended DDP/RDMAP header across more than one TCP callback. The first callback may receive most of the header, while the next one still limits the copy to hdrlen - MIN_DDP_HDR instead of the number of missing bytes. This makes the destination move past the end of the header and overwrite the receive state, including fpdu_part_rcvd. A later callback can then use a negative fpdu_part_rcvd value as a copy offset, which creates an OOB write. Use the number of header bytes already received when calculating the next copy length. Fixes: 754209850df8 ("RDMA/siw: Always consume all skbuf data in sk_data_ready() upcall.") Signed-off-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr> Link: https://patch.msgid.link/20260908085520.1746329-1-Jeremy.Jean@oss.cyber.gouv.fr Assisted-by: Codex:gpt-6 Acked-by: Bernard Metzler <bernard.metzler@linux.dev> Signed-off-by: Leon Romanovsky <leon@kernel.org>
2026-09-10RDMA/efa: Keep EQ resources alive while IRQ is registeredLeon Romanovsky
The completion IRQ handler accesses the EQ state and DMA buffer. Its IRQ was registered before that state was initialized, while teardown released the buffer before free_irq() synchronized the handler. Initialize the EQ without arming it, register the IRQ, and then arm it. Reverse the resource order during teardown by freeing the IRQ before destroying the EQ. Fixes: 2a152512a155 ("RDMA/efa: CQ notifications") Link: https://patch.msgid.link/20260907-use-after-free-of-admin-queue-struct-v1-2-dd9d9267fbf4@nvidia.com Reviewed-by: Michael Margolin <mrgolin@amazon.com> Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
2026-09-10RDMA/efa: Keep admin queues alive while IRQ is registeredLeon Romanovsky
The management IRQ handler accesses both the admin completion queue and the async event queue. The driver registered the IRQ before constructing these queues and destroyed them before freeing the IRQ, so the handler's lifetime was not contained by the resources it accesses. Initialize the queues with interrupts masked, request the IRQ, and then switch to interrupt mode. On removal, reset the device and free the IRQ before destroying the queues. Also reset the device before destroying the queues if IRQ registration fails, because the device already has their DMA addresses. Fixes: b7f5e880f377 ("RDMA/efa: Add the efa module") Link: https://patch.msgid.link/20260907-use-after-free-of-admin-queue-struct-v1-1-dd9d9267fbf4@nvidia.com Reviewed-by: Michael Margolin <mrgolin@amazon.com> Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
2026-09-10RDMA: fix repeated words in commentsHemanth Selam
Drop words accidentally written twice, reported by checkpatch.pl as a possible repeated word. Only touches comments, no code changes. Assisted-by: Cursor:claude-opus-5 Signed-off-by: Hemanth Selam <hemanth.selam@gmail.com> Link: https://patch.msgid.link/20260907065047.26773-3-hemanth.selam@gmail.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2026-09-10RDMA: fix typos in commentsHemanth Selam
Fix typos in comments, reported by scripts/checkpatch.pl using the misspelling list in scripts/spelling.txt. Only touches comments, no code changes. Assisted-by: Cursor:claude-opus-5 Signed-off-by: Hemanth Selam <hemanth.selam@gmail.com> Link: https://patch.msgid.link/20260907065047.26773-2-hemanth.selam@gmail.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2026-09-10RDMA/hns: Fix GID capacity loss in 64K systemJunxian Huang
Our HW always handles GMV BT pages with a fixed 4K size, but the driver calculates the needed GMV BT pages number with PAGE_SIZE, which is 64K in 64K system. Only the first 4K (GMV index 0-127) can be reached by HW, causing GID capacity loss and memory waste. Split a single 64K BT page into multiple 4K pages and register them to HW per 4K block so that HW can correctly reach all GMV entries. Fixes: 32053e584e4a ("RDMA/hns: Add support for filling GMV table") Signed-off-by: Junxian Huang <huangjunxian6@hisilicon.com> Link: https://patch.msgid.link/20260907084901.2420703-4-huangjunxian6@hisilicon.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2026-09-10RDMA/hns: Limit gmv_entry_num to avoid memory wasteJunxian Huang
The GMV entry is a HW object corresponding to a GID. Since gid_table_len is already limited to a maximum of 256, there is no need to allocate memory for those extra GMV entries as they will never be touched. Fixes: 7243396aaf12 ("RDMA/hns: Add a max length of gid table") Signed-off-by: Junxian Huang <huangjunxian6@hisilicon.com> Link: https://patch.msgid.link/20260907084901.2420703-3-huangjunxian6@hisilicon.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2026-09-10RDMA/hns: Use u32 for gid_table_lenJunxian Huang
The gid_table_len is always read from HW registers or derived from u32 arithmetic, and it is never negative. Change its type from int to u32 to avoid signed/unsigned mixed-type operations with other u32 fields such as gmv_entry_num. Signed-off-by: Junxian Huang <huangjunxian6@hisilicon.com> Link: https://patch.msgid.link/20260907084901.2420703-2-huangjunxian6@hisilicon.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2026-09-06RDMA/core: fix refcount bug in iwpm_get_nlmsg_request()Jeffin Philip
iwpm_get_nlmsg_request() initializes refcount _after_ list_add_tail() making it accessible to global list where another CPU can kref_get() on nlmsg_request causing a refcount "addition on 0" bug. Fix this by initializing kref _before_ list_add_tail() so refcount for nlmsg_request can be incremented/decremented normally. In addition, also initialize every field before list_add_tail(). Reported-by: syzbot+bd317784d628820741b5@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=bd317784d628820741b5 Fixes: 30dc5e63d6a5 ("RDMA/core: Add support for iWARP Port Mapper user space service") Cc: stable@vger.kernel.org Signed-off-by: Jeffin Philip <jeffinphilip14@gmail.com> Link: https://patch.msgid.link/20260904131437.12917-1-jeffinphilip14@gmail.com Signed-off-by: Leon Romanovsky <leon@kernel.org>
2026-09-06RDMA/irdma: Remove unused post_sq argumentsLeon Romanovsky
The only callers of irdma_sc_qp_flush_wqes() and irdma_sc_mr_fast_register() always request their respective send queues to be posted, so their post_sq false branches are unreachable. Remove both arguments and post the queues unconditionally. Drop the redundant QP flush request assignments as well. Link: https://patch.msgid.link/20260903-64-bit-iova-is-silently-truncated-to-v1-2-96e878ea6873@nvidia.com Reviewed-by: Kalesh AP <kalesh-anakkur.purayil@broadcom.com> Reviewed-by: Jacob Moroni <jmoroni@google.com> Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
2026-09-06RDMA/irdma: Preserve fast-registration IOVA on 32-bitLeon Romanovsky
ib_mr::iova is u64, but the fast-registration path passes it through void * and uintptr_t. These conversions truncate the upper 32 bits on 32-bit kernels before the WQE is built. Store the IOVA as u64 and write it directly to the WQE. The sole caller always uses VA-based addressing, so remove the unused FBO selection and set the VA-based bit unconditionally. Fixes: b48c24c2d710 ("RDMA/irdma: Implement device supported verb APIs") Link: https://patch.msgid.link/20260903-64-bit-iova-is-silently-truncated-to-v1-1-96e878ea6873@nvidia.com Reviewed-by: Jacob Moroni <jmoroni@google.com> Signed-off-by: Leon Romanovsky <leonro@nvidia.com>