| Age | Commit message (Collapse) | Author |
|
Add a dma_dir field to struct ib_umem to record the direction chosen at
map time and reuse it at unmap. Thread an explicit dma_data_direction
parameter through the internal pinning helpers so each caller can supply
the direction that matches the device's actual access pattern.
All existing callers still pass DMA_BIDIRECTIONAL, so this is a pure
mechanism addition with no behavioural change.
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-6-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Several drivers pin their CQ ring buffer via ib_umem_get_va() rather
than the CQ-specific helper. Switch them to ib_umem_get_cq_buf_or_va()
so the subsequent DMA_FROM_DEVICE change covers all CQ paths at once.
For shared helpers used by both CQ and non-CQ callers (qedr, mana,
erdma, hns) a new is_cq parameter selects the appropriate pinning
function; non-CQ callers pass false.
vmw_pvrdma's CQ is excluded: it embeds a ring-state header that the
driver CPU-writes after polling, making it genuinely bidirectional.
This is a pure refactor with no functional change.
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-5-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
pvrdma embeds a pvrdma_ring_state header directly inside the QP and SRQ
ring buffers, and the hypervisor writes cons_head into that header.
Pinning with access=0 does not force a private page copy (no
FOLL_WRITE), so the hypervisor may write into a shared page.
Pass IB_ACCESS_LOCAL_WRITE for all three rings. The SRQ ring_state is
not currently used by the driver but is pinned writable for consistency.
Fixes: 29c8d9eba550 ("IB: Add vmw_pvrdma driver")
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-4-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
alloc_cq_buf() leaves user_access zero-initialized, but the device
writes CQEs into the CQ buffer. Set IB_ACCESS_LOCAL_WRITE so the buffer
is pinned writable, matching the device's write access.
Fixes: 744b7bdfa79e ("RDMA/hns: Support 0 hop addressing for CQE buffer")
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-3-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
erdma_init_user_cq() pins the CQ ring with access=0, but the device
writes CQEs into it. Pass IB_ACCESS_LOCAL_WRITE so the buffer is pinned
writable and marked dirty on unpin.
Fixes: 155055771704 ("RDMA/erdma: Add verbs implementation")
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Link: https://patch.msgid.link/20260915140933.40580-2-yishaih@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
irdma_cfg_ceq_vector() registers an IRQ handler before requesting the
CEQ-to-vector mapping through the virtual channel. If the mapping
fails, the function returns without releasing the IRQ resources.
The caller destroys the CEQ, leaving the IRQ handler registered
and potentially referencing freed memory.
Call irdma_destroy_irq() on mapping failure to release
the IRQ resources and drain the associated tasklet
before the caller destroys the CEQ. Use the same dev_id
passed to request_irq(), accounting for CEQ0 sharing
its interrupt vector with the AEQ.
Fixes: b800e82feba7 ("RDMA/irdma: Add GEN3 support for AEQ and CEQ")
Assisted-by: Claude:claude-opus-5
Signed-off-by: Serhat Kumral <serhatkumral1@gmail.com>
Link: https://patch.msgid.link/20260918110428.36586-1-serhatkumral1@gmail.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Commit e91fb8b9d0ed ("RDMA/mlx5: Avoid flexible array warning") split
the leftovers_specs[] array into two separate structs to silence the
sparse "array of flexible structures" warning, but in doing so reversed
the member order: eth_flow was placed before flow_attr.
_create_flow_rule() locates the flow specs with:
ib_flow = (void *)flow_attr + sizeof(*flow_attr);
which expects the spec data to immediately follow the ib_flow_attr.
With the reversed order, flow_attr is the last member, so ib_flow
points past the end of the struct and the num_of_specs = 1 loop reads
out-of-bounds memory (whatever static data follows in the .data
section), producing incorrect flow rules or a crash.
Rebuild the leftovers attributes with DEFINE_RAW_FLEX(), which declares
an on-stack struct ib_flow_attr with a properly sized trailing flows[]
array. The ETH spec lives in flows[0].eth at exactly the offset
_create_flow_rule() expects, so no layout-dependent composite struct is
needed. This also avoids embedding a structure that contains a flexible
array member as a non-last member, which would trigger
-Wflex-array-member-not-at-end once the warning is enabled.
Fixes: e91fb8b9d0ed ("RDMA/mlx5: Avoid flexible array warning")
Signed-off-by: Li RongQing <lirongqing@baidu.com>
Link: https://patch.msgid.link/20260918092337.2338-1-lirongqing@baidu.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Move the kernel QP helpers after the shared memory helpers to remove
forward declarations.
Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com>
Link: https://patch.msgid.link/20260917120459.13286-6-chengyou@linux.alibaba.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Userspace and kernel space queue buffers use separate helpers despite
sharing MTT metadata and lifetime rules. Manage both through
erdma_mem_init() and erdma_mem_uninit(), while keeping backing allocation
and release type-specific.
Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com>
Link: https://patch.msgid.link/20260917120459.13286-5-chengyou@linux.alibaba.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
A single coherent allocation for a kernel CQ can fail when memory is
fragmented.
Use page-sized coherent buffers and describe them with the existing MTT.
Keep the userspace CQ path and doorbell allocation unchanged.
Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com>
Link: https://patch.msgid.link/20260917120459.13286-4-chengyou@linux.alibaba.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
A single coherent allocation for kernel QP queues can fail for large
queues when memory is fragmented.
Allocate page-sized coherent buffers and describe them with the existing
MTT. Keep the userspace QP path unchanged.
Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com>
Link: https://patch.msgid.link/20260917120459.13286-3-chengyou@linux.alibaba.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Allocation failures currently use the normal kernel QP teardown path.
Unwind failures directly so free_kernel_qp() only handles complete QPs.
Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com>
Link: https://patch.msgid.link/20260917120459.13286-2-chengyou@linux.alibaba.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
The steering anchor resources are shared between anchors with the same
priority and are protected by rule_goto_table_ref.
When steering anchor creation succeeds but uverbs_copy_to() fails, the
error path drops the reference acquired for the new anchor. However, it
unconditionally calls mlx5_steering_anchor_destroy_res(), even when
other anchors are still using the shared resources.
This can destroy the shared flow table, flow groups, and rules while
existing anchors still hold references to them, resulting in a
use-after-free when those anchors are subsequently cleaned up.
Only destroy the shared resources when dropping the reference makes
rule_goto_table_ref zero, consistent with the existing cleanup path.
Fixes: e1f4a52ac171 ("RDMA/mlx5: Create an indirect flow table for steering anchor")
Signed-off-by: Li RongQing <lirongqing@baidu.com>
Link: https://patch.msgid.link/20260917085419.1792-1-lirongqing@baidu.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
When the SA assigns a wildcard (mgid0) join a concrete MGID that already
has a group in the port table, join_handler() erases the group from the
tree and calls mcast_insert() with allow_duplicates=false. mcast_insert()
then finds the existing group and returns it without re-inserting, but
join_handler() discards the return value, leaving the group erased yet
still alive.
When the orphaned group's refcount later drops to zero, release_group()
calls rb_erase() on the node that is no longer in the tree, corrupting
the rb-tree (a stale parent pointer causes a valid subtree link to be
overwritten with NULL).
Clear the node with RB_CLEAR_NODE() after the erase in join_handler() so
release_group() can detect via RB_EMPTY_NODE() that it is already out of
the tree and skip the second erase. A successful re-insert by
mcast_insert() restores the in-tree state via rb_link_node().
Signed-off-by: Li RongQing <lirongqing@baidu.com>
Link: https://patch.msgid.link/20260917052012.2181-1-lirongqing@baidu.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
The reference taken by siw_qp_id2obj() is dropped before the address
of the embedded ib_qp is taken from the QP. If that reference was
the last one the QP is freed and the pointer returned to the iwarp
core is dangling; the core only pins the QP with iw_add_ref() after
the call has returned.
Copy the pointer while the reference is still held.
Fixes: 6c52fdc244b5 ("rdma/siw: connection management")
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Link: https://patch.msgid.link/20260916184400.2093106-1-vulab@iscas.ac.cn
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Since commit 312b8f79eb05 ("RDMA/mlx: Calling qp event handler in
workqueue context") rsc_event_notifier() expects the event handler to
drop the resource reference taken by mlx5_get_rsc(). mlx5_ib_wq_event()
was missed by that conversion, so every RQ event leaks one reference,
which keeps destroy_resource_common() waiting on the completion.
Put the resource reference on all exits of the handler.
Fixes: 312b8f79eb05 ("RDMA/mlx: Calling qp event handler in workqueue context")
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Link: https://patch.msgid.link/20260916183921.2092821-1-vulab@iscas.ac.cn
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
rx_pkt() takes an extra reference on the received skb so it can be
handed to the firmware completion path via req->cookie. When the
work request skb allocation in send_fw_pass_open_req() fails the
function returns without releasing that extra reference, leaking
one skb reference.
Free the packet skb before returning when the allocation fails,
matching the cxgb4_ofld_send() failure path.
Fixes: 1cab775c3e75 ("RDMA/cxgb4: Fix LE hash collision bug for passive open connection")
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Link: https://patch.msgid.link/20260916183715.2092663-1-vulab@iscas.ac.cn
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
If ip_dev_find() fails on the loopback path the code jumps to
free_dst without releasing the neighbour reference taken by
dst_neigh_lookup_skb(), leaking one neigh reference.
Release the neighbour before jumping to free_dst.
Fixes: ef42520240aa ("RDMA/cxgb4: add null-ptr-check after ip_dev_find()")
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Link: https://patch.msgid.link/20260916183530.2092534-1-vulab@iscas.ac.cn
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
An existing RDMA function does not guarantee allocation of the new
vport device. Return -ENOMEM before initializing a NULL iwdev.
Detected by static analysis and reviewed with AI-assisted source auditing.
Fixes: 2ad49ae7330b ("RDMA/irdma: Introduce GEN3 vPort driver support")
Assisted-by: LLM
Signed-off-by: Slavin Liu <bolin.liu@seu.edu.cn>
Link: https://patch.msgid.link/20260911060909.94219-1-bolin.liu@seu.edu.cn
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
During VM live migration or when registering large memory regions
(e.g., guest memory >= 256 GB backed by 4 KiB pages), get_lvl2_pble()
calculates total leaves exceeding ~131,072. The memory required for
lvl2->leafmem.va exceeds 4 MiB (sizeof(*leaf) * total > 5.24 MiB).
Because kzalloc() requires physically contiguous pages, requests larger
than MAX_PAGE_ORDER (4 MiB on x86_64) fail and trigger a warning:
WARNING: CPU: 123 PID: 119906 at mm/page_alloc.c:5997 __alloc_frozen_pages_noprof+0x50c/0x1050
Call Trace:
<TASK>
alloc_frozen_pages_noprof+0xc0/0x240
___kmalloc_large_node+0x3d/0xd0
__kmalloc_large_node_noprof+0x1b/0x90
__kmalloc_noprof+0x254/0x660
get_lvl1_lvl2_pble+0xbc/0x270 [irdma]
irdma_get_pble+0x44/0xc0 [irdma]
irdma_setup_pbles+0x46/0x1d0 [irdma]
irdma_reg_user_mr_type_mem+0x41/0x230 [irdma]
irdma_reg_user_mr+0x194/0x1c0 [irdma]
ib_uverbs_reg_mr+0x1b0/0x2c0 [ib_uverbs]
The leafmem array is purely host software tracking structures (struct
irdma_pble_info) and does not require physically contiguous memory.
Switch the allocation from kzalloc() to kvzalloc() and free it with
kvfree() so that requests exceeding MAX_PAGE_ORDER fall back to
virtually contiguous memory via vmalloc.
Fixes: e8c4dbc2fcac ("RDMA/irdma: Add PBLE resource manager")
Assisted-by: LLM
Signed-off-by: Jeremy Bongio <jbongio@google.com>
Link: https://patch.msgid.link/20260916173203.2555458-1-jbongio@google.com
Tested-by: Jacob Moroni <jmoroni@google.com>
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Commit cc8997c94bf3 ("RDMA/irdma: Refactor PBLE functions")
replaced the explicit level1_only boolean with a bit mask
but did not properly convert the "if (level1_only)" check
in irdma_get_pble() to be "lvl == PBLE_LEVEL_1" and instead
left it as "if (lvl)".
User MR registration passes "lvl = PBLE_LEVEL_1 | PBLE_LEVEL_2"
so this check always evaluates to true and causes the
loop to bail after grabbing a single SD even if the region
actually requires more than one SD (as is the case for
regions larger than 1 GiB with 4k pages). As a result, these
registrations can fail with -ENOMEM even if there are
sufficient PBLE resources.
Fixes: cc8997c94bf3 ("RDMA/irdma: Refactor PBLE functions")
Signed-off-by: Jacob Moroni <jmoroni@google.com>
Link: https://patch.msgid.link/20260915170317.2901672-1-jmoroni@google.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
When we call `rdma_disconnect`, it's possible that we fail to send both
a DREQ and a DREP. In that case, we currently either do nothing or move
into the `IB_CM_TIMEWAIT` state. Either way, we remain in the
`RDMA_CM_CONNECT` state. Anyone waiting for us to disconnect hangs.
Instead, while we're connected, when we call `rdma_disconnect`, and fail
to send both a DREQ and a DREP, mark the connection as disconnected.
Call the upper layer's handler for the disconnect, and destroy the
connection if instructed.
Cc: Leon Romanovsky <leon@kernel.org>
Cc: Jason Gunthorpe <jgg@ziepe.ca>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Paulo Alcantara <pc@manguebit.org>
Cc: Stefan Metzmacher <metze@samba.org>
Cc: Tom Talpey <tom@talpey.com>
Cc: linux-rdma@vger.kernel.org
Cc: linux-cifs@vger.kernel.org
Cc: samba-technical@lists.samba.org
Cc: linux-kernel@vger.kernel.org
Signed-off-by: Ammar Ratnani <ammrat13@gmail.com>
Link: https://patch.msgid.link/20260908154454.10966-2-ammrat13@gmail.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
__hns_roce_v2_cq_clean() sweeps the CQ backwards from the producer
index down to the consumer index:
while ((int) --prod_index - (int) hr_cq->cons_index >= 0)
Both indexes are free running u32 counters, so the comparison has to
be done modulo 2^32. Casting each operand to int and subtracting
does not do that: the subtraction overflows whenever the two indexes
straddle 2^31, which is undefined behaviour, and a compiler that
assumes signed overflow cannot occur is free to discard the
subtraction and fold the expression into a plain signed comparison.
That comparison is not wraparound safe.
The kernel is built with -fno-strict-overflow, so gcc and clang both
retain the subtraction today and the generated code is unaffected;
there is no known user-visible impact from the current code. Still,
correctness here shouldn't depend on that build flag.
Replace the loop with a plain unsigned equality check instead:
while (prod_index != hr_cq->cons_index) {
--prod_index;
...
The preceding forward scan starts prod_index at cons_index and advances
it only while get_sw_cqe_v2() still reports a software owned entry, so
the two indexes stay within one CQ's worth of entries. Decrementing
prod_index reaches cons_index in exactly that many iterations regardless
of whether either counter has wrapped, because the loop no longer
compares magnitudes at all.
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Signed-off-by: Edward Srouji <edwards@nvidia.com>
Link: https://patch.msgid.link/20260915-fix-cq-cleanup-v1-4-e991944cf898@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
mthca_cq_clean() sweeps the CQ backwards from the producer index
down to the consumer index:
while ((int) --prod_index - (int) cq->cons_index >= 0)
Both indexes are free running u32 counters, so the comparison has to
be done modulo 2^32. Casting each operand to int and subtracting
does not do that: the subtraction overflows whenever the two indexes
straddle 2^31, which is undefined behaviour, and a compiler that
assumes signed overflow cannot occur is free to discard the
subtraction and fold the expression into a plain signed comparison.
That comparison is not wraparound safe.
The kernel is built with -fno-strict-overflow, so gcc and clang both
retain the subtraction today and the generated code is unaffected;
there is no known user-visible impact from the current code. Still,
correctness here shouldn't depend on that build flag.
Replace the loop with a plain unsigned equality check instead:
while (prod_index != cq->cons_index) {
--prod_index;
...
The preceding forward scan starts prod_index at cons_index and stops no
later than cons_index + cq->ibcq.cqe, so the two indexes are at most
cq->ibcq.cqe apart. Decrementing prod_index reaches cons_index in
exactly that many iterations regardless of whether either counter has
wrapped, because the loop no longer compares magnitudes at all.
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Signed-off-by: Edward Srouji <edwards@nvidia.com>
Link: https://patch.msgid.link/20260915-fix-cq-cleanup-v1-3-e991944cf898@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
__mlx4_ib_cq_clean() sweeps the CQ backwards from the producer index
down to the consumer index:
while ((int) --prod_index - (int) cq->mcq.cons_index >= 0)
Both indexes are free running u32 counters, so the comparison has to
be done modulo 2^32. Casting each operand to int and subtracting
does not do that: the subtraction overflows whenever the two indexes
straddle 2^31, which is undefined behaviour, and a compiler that
assumes signed overflow cannot occur is free to discard the
subtraction and fold the expression into a plain signed comparison.
That comparison is not wraparound safe.
The kernel is built with -fno-strict-overflow, so gcc and clang both
retain the subtraction today and the generated code is unaffected;
there is no known user-visible impact from the current code. Still,
correctness here shouldn't depend on that build flag.
Replace the loop with a plain unsigned equality check instead:
while (prod_index != cq->mcq.cons_index) {
--prod_index;
...
The preceding forward scan starts prod_index at cons_index and stops no
later than cons_index + cq->ibcq.cqe, so the two indexes are at most
cq->ibcq.cqe apart. Decrementing prod_index reaches cons_index in
exactly that many iterations regardless of whether either counter has
wrapped, because the loop no longer compares magnitudes at all.
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Signed-off-by: Edward Srouji <edwards@nvidia.com>
Link: https://patch.msgid.link/20260915-fix-cq-cleanup-v1-2-e991944cf898@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
__mlx5_ib_cq_clean() sweeps the CQ backwards from the producer index
down to the consumer index:
while ((int) --prod_index - (int) cq->mcq.cons_index >= 0)
Both indexes are free running u32 counters, so the comparison has to
be done modulo 2^32. Casting each operand to int and subtracting
does not do that: the subtraction overflows whenever the two indexes
straddle 2^31, which is undefined behaviour, and a compiler that
assumes signed overflow cannot occur is free to discard the
subtraction and fold the expression into a plain signed comparison.
That comparison is not wraparound safe.
The kernel is built with -fno-strict-overflow, so gcc and clang both
retain the subtraction today and the generated code is unaffected;
there is no known user-visible impact from the current code. Still,
correctness here shouldn't depend on that build flag.
Replace the loop with a plain unsigned equality check instead.
while (prod_index != cq->mcq.cons_index) {
--prod_index;
...
The preceding forward scan starts prod_index at cons_index and stops no
later than cons_index + cq->ibcq.cqe, so the two indexes are at most
cq->ibcq.cqe apart. Decrementing prod_index reaches cons_index in
exactly that many iterations regardless of whether either counter has
wrapped, because the loop no longer compares magnitudes at all.
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
Signed-off-by: Edward Srouji <edwards@nvidia.com>
Link: https://patch.msgid.link/20260915-fix-cq-cleanup-v1-1-e991944cf898@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
icrdma_remove() dereferences iwdev after calling
irdma_ib_unregister_device():
irdma_ib_unregister_device(iwdev)
-> ib_unregister_device(&iwdev->ibdev)
-> __ib_unregister_device()
-> ib_dealloc_device()
irdma registers irdma_ib_dealloc_device() as the dealloc_driver
callback in its ib_device_ops. The RDMA core explicitly documents
ib_unregister_device() that when ops.dealloc_driver is used, ib_dev
will be freed upon return from the function.
The ib_device is embedded in struct irdma_device and iwdev itself was
allocated by ib_alloc_device(irdma_device, ibdev). Consequently, once
irdma_ib_unregister_device() returns, iwdev must no longer be
dereferenced.
However, icrdma_remove() currently accesses iwdev->rf afterwards when
deinitializing interrupts, destroying ah_tbl_lock, and freeing rf.
Although the final ib_device storage is released with kfree_rcu(), the
object lifetime has already ended and callers must not rely on the RCU
grace period to continue dereferencing iwdev.
Save iwdev->rf before unregistering the RDMA device and use the saved
pointer for the remaining cleanup.
This issue was found by manual code inspection.
Fixes: 8498a30e1b94 ("RDMA/irdma: Register auxiliary driver and implement private channel OPs")
Cc: stable@vger.kernel.org
Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com>
Link: https://patch.msgid.link/20260914122840.1709743-1-lgs201920130244@gmail.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
rdma_build_skb() constructs a synthetic IPv6 header for LAG slave
selection and copies the flow label from the AH attribute:
memcpy(&ip6h->flow_lbl, &ah_attr->grh.flow_label,
sizeof(*ip6h->flow_lbl));
ipv6hdr.flow_lbl is __u8[3], so sizeof(*ip6h->flow_lbl) dereferences the
array to a single __u8 and yields 1. The memcpy therefore copies only
the first byte of the flow label, leaving flow_lbl[1] and flow_lbl[2]
uninitialised in the skb buffer (alloc_skb() does not zero the data).
The resulting LAG hash is computed over one byte of real flow-label
data and two bytes of kmalloc residue, so IPv6 RoCEv2 traffic on a
bonded interface can be steered to the wrong slave.
Use sizeof(ip6h->flow_lbl) (the array, 3 bytes) so the whole flow label
is copied.
Fixes: bd3920eac1331 ("RDMA/core: Add LAG functionality")
Signed-off-by: Li RongQing <lirongqing@baidu.com>
Link: https://patch.msgid.link/20260916120926.1761-1-lirongqing@baidu.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
The UVERBS_METHOD_GET_CONTEXT handler allocates the ucontext via
ib_alloc_ucontext(), which calls rdma_restrack_new() (initialising the
kref to 1) and rdma_restrack_set_name(NULL), the latter attaching the
current task and taking a task_struct reference. When ib_init_ucontext()
subsequently fails, the error path only does:
kfree(attrs->context);
attrs->context = NULL;
without first calling rdma_restrack_put() on &attrs->context->res. The
kref therefore never reaches zero, restrack_release() is never invoked,
and put_task_struct() is never called, so the task_struct reference taken
during set_name is leaked permanently.
A local user with access to an RDMA device can repeat this path (e.g. by
hitting an RDMA cgroup limit or supplying invalid ucaps) and leak one
task_struct reference per attempt, eventually preventing those processes
from being reaped or exhausting memory.
The legacy ib_uverbs_get_context() path already handles this correctly
at its err_ucontext label, where rdma_restrack_put() precedes kfree().
Mirror that ordering here.
Fixes: a1123418ba10 ("RDMA/uverbs: Add ioctl command to get a device context")
Signed-off-by: Li RongQing <lirongqing@baidu.com>
Link: https://patch.msgid.link/20260916085529.1838-1-lirongqing@baidu.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
When adding a QP to the restrack xarray, rdma_restrack_add() does:
ret = xa_insert(&rt->xa, res->id, res, GFP_KERNEL);
if (ret)
res->id = 0;
if (qp->qp_type >= IB_QPT_DRIVER)
xa_set_mark(&rt->xa, res->id, RESTRACK_DD);
The xa_set_mark() call is not guarded by "!ret". When xa_insert() fails
(possible on a qp_num collision or memory pressure), res->id is reset to
0, yet the mark is still applied to index 0. If a legitimate QP whose
qp_num is 0 already occupies that slot - e.g. a normal QP or the SMI QP
- it is spuriously tagged as driver-private (RESTRACK_DD), and the
netlink dump path res_get_common_dumpit() then hides it from "rdma res"
output whenever the caller does not request driver details.
Only set the mark when the insertion actually succeeded by folding the
mark into the non-failure branch.
Fixes: e18fa0bbcedf8 ("RDMA/core: Add an option to display driver-specific QPs in the rdmatool")
Signed-off-by: Li RongQing <lirongqing@baidu.com>
Link: https://patch.msgid.link/20260916073042.1746-1-lirongqing@baidu.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
ib_add_sub_device() links a new sub-device into its parent's sub-device
list with:
list_add_tail(&parent->subdev_list_head, &sub->subdev_list);
list_add_tail(new, head) expects the node to insert as the first
argument and the list head as the second, so the call above does the
opposite of what was intended: it treats the parent's list head as the
new node and the sub-device's node as the list head.
Because _ib_alloc_device() initialises subdev_list to a self-referencing
empty head, the misplaced insertion does not crash, but it corrupts the
parent->subdev_list_head list. After adding two or more sub-devices,
only the last one is reachable through the parent's head, and
ib_del_sub_device_and_put()'s list_del() on &sub->subdev_list rewrites
the head pointers, potentially severing earlier sub-devices from the
chain. During parent teardown, the reverse iteration over
subdev_list_head then skips the orphaned sub-devices, so their
del_sub_dev() callbacks are never invoked and ib_device_put() on the
parent is never paired, leaking both the HW sub-devices and a refcount.
Swap the arguments to match the documented intent: insert the
sub-device node into the parent's list head.
Fixes: bca51197620a ("RDMA/core: Support IB sub device with type "SMI"")
Signed-off-by: Li RongQing <lirongqing@baidu.com>
Link: https://patch.msgid.link/20260916064435.2180-1-lirongqing@baidu.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Cross-merge networking fixes after downstream PR (net-7.3-rc4).
Conflicts:
net/core/neighbour.c
979aabdad8dd0 ("neighbour: Skip default parms when resumed in neightbl_dump_info().")
7b430fcfc972f ("neighbour: Don't render blackhole_netdev via RTM_GETNEIGHTBL.")
fae1c59810b86 ("neighbour: Remove unnecessary net_eq().")
https://lore.kernel.org/20260911173056.44ec06e0@kernel.org
https://lore.kernel.org/aqfbJi7nAX4IbmnR@sirena.co.uk
Adjacent changes:
net/netlink/af_netlink.c
ceac0de741bf ("netlink: do not free nlk->groups while lockless readers can use it")
7c0ec6288b49 ("net: Replace %pK output with 0")
net/bridge/br_vlan.c
2842ce397dd0 ("net: bridge: vlan: fix bugs caused by switchdev deletion errors")
5bec8f861114 ("net: bridge: vlan: annotate lockless use of num_vlans")
2b1f8fd3118c ("net: bridge: vlan: annotate lockless vlan flags use")
net/bridge/br_mst.c
18a6fe05fb6e ("net: bridge: mst: move switchdev call outside rcu")
120207a08fc0 ("net: bridge: vlan: annotate lockless use of msti")
drivers/net/ethernet/stmicro/stmmac/hwif.h
90e4b849dfa6 ("net: stmmac: propagate FPE preemption-class mapping errors")
85ca3292d7a3 ("net: stmmac: Remove ARP offload code")
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The variable 'i' is used without being initialized in active-backup
mode.
Fixes: d9023e461b73 ("RDMA/hns: Implement bonding init/uninit process")
Signed-off-by: Junxian Huang <huangjunxian6@hisilicon.com>
Link: https://patch.msgid.link/20260914024800.132429-1-huangjunxian6@hisilicon.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Replace "Shared receive queue (SRQ)" with "Send queue".
Signed-off-by: Serhat Kumral <serhatkumral1@gmail.com>
Link: https://patch.msgid.link/20260914130536.19769-1-serhatkumral1@gmail.com
Reviewed-by: Bart Van Assche <bvanassche@acm.org>
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Add the actual error code to the "Create QP type %d failed" log
line in create_qp(), so failures can be diagnosed from dmesg
alone without guessing the underlying error.
Signed-off-by: Li RongQing <lirongqing@baidu.com>
Link: https://patch.msgid.link/20260914121253.2193-1-lirongqing@baidu.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
UVERBS_ATTR_CREATE_WQ_CQ_HANDLE is declared UA_OPTIONAL in the ioctl
method definition, so the mandatory attribute bitmap does not enforce
its presence. When userspace omits it, uverbs_attr_get_obj() returns
ERR_PTR(-ENOENT) and the handler stores that error pointer into
wq_init_attr.cq without validation.
The bogus cq pointer is then passed to the driver's create_wq callback.
In the mlx5 case, create_rq() calls to_mcq(init_attr->cq) which applies
container_of to the ERR_PTR value, producing a near-NULL pointer. The
subsequent access in get_rq_ts_format() triggers a kernel NULL pointer
dereference:
BUG: kernel NULL pointer dereference, address: 0000000000000296
RIP: 0010:create_rq+0x32/0x550 [mlx5_ib]
Call Trace:
mlx5_ib_create_wq+0x14a/0x210 [mlx5_ib]
ib_uverbs_handler_UVERBS_METHOD_WQ_CREATE+0x1f0/0x320 [ib_uverbs]
ib_uverbs_run_method+0x296/0x320 [ib_uverbs]
ib_uverbs_cmd_verbs+0x1a0/0x260 [ib_uverbs]
ib_uverbs_ioctl+0xa8/0x120 [ib_uverbs]
A WQ without a CQ was never valid; the legacy write path always required
one via uobj_get_obj_read() in ib_uverbs_ex_create_wq(). Declare the
attribute UA_MANDATORY so the uverbs framework rejects the ioctl early
when the CQ handle is missing, before the handler ever runs.
Fixes: ef3bc084a8ed ("IB/uverbs: Introduce create/destroy WQ commands over ioctl")
Signed-off-by: Li RongQing <lirongqing@baidu.com>
Link: https://patch.msgid.link/20260911021557.2113-1-lirongqing@baidu.com
Reviewed-by: Kalesh AP <kalesh-anakkur.purayil@broadcom.com>
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
For both SRQ and non-SRQ receive paths, the WQE is copied into a local
buffer to provide a kernel-owned, validated copy. While calculating the
memcpy size from the validated num_sge prevents overflow during the
copy, memcpy() itself still copies num_sge from shared memory.
A concurrent userspace modification before or during memcpy() leaves
an unvalidated num_sge in the local buffer, leading to potential
out-of-bounds reads in rxe_resp_check_length() and copy_data() causing:
BUG: KASAN: slab-out-of-bounds in rxe_receiver+0x8109/0x9ec0 [rdma_rxe]
Read of size 4 at addr ffff88812c4867f8 by task kworker/u9:6/361
Workqueue: rxe_wq do_work [rdma_rxe]
Call Trace:
rxe_receiver+0x8109/0x9ec0 [rdma_rxe]
do_work+0x149/0x610 [rdma_rxe]
process_one_work+0x726/0x10a0
The buggy address belongs to the object at ffff88812c486000
which belongs to the cache kmalloc-part-13-2k of size 2048
The buggy address is located 0 bytes to the right of
allocated 2040-byte region [ffff88812c486000, ffff88812c4867f8)
Explicitly assign the validated num_sge to the local buffer after the
copy to prevent this race.
Fixes: 22b8fbded65b ("RDMA/rxe: Fix TOCTOU heap overflow in get_srq_wqe")
Fixes: d6ab440240a0 ("RDMA/rxe: Copy WQE to local buffer in non-SRQ receive path")
Reviewed-by: Zhu Yanjun <yanjun.zhu@linux.dev>
Signed-off-by: Nicolas Morey <nmorey@suse.com>
Link: https://patch.msgid.link/20260909160132.1491248-1-nmorey@suse.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Pull rdma fixes from Jason Gunthorpe:
"Lots of bug fixes from the last weeks:
- Various error unwind bugs
- Several more races and bugs in siw and rxe, including remote
triggerable
- HFI1 corruption with its credit scheme
- Remove a bogus user triggerable dev_warn
- Lock __ethtool_get_link_ksettings() properly
- Fix a lockdep loop with diassociation
- Several storage related bugs, some triggerable remotely
- Do no leak physical addresses to userspace in bnxt_re
- Fix wrong irq context for the xarrays in erdma
- User triggerable race in ucma with multicast
- Race in ipoib with multicast flushing and destruction"
* tag 'for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/rdma/rdma: (28 commits)
RDMA/siw: Bound fragmented header copies by the remaining length
RDMA/efa: Keep EQ resources alive while IRQ is registered
RDMA/efa: Keep admin queues alive while IRQ is registered
RDMA/core: fix refcount bug in iwpm_get_nlmsg_request()
IB/IPoIB: Avoid restoring OPER_UP after multicast flush
RDMA/ucma: Serialize join and leave on copy_to_user failure
RDMA/rtrs-clt: Fix CQ pool leak when connect is interrupted
RDMA/irdma: Enforce local fence for IB_WR_REG_MR
RDMA/erdma: Use IRQ-safe XArray helpers for QP and CQ tables
RDMA/mad: Fix receive buffer leak when PKey enforcement fails
RDMA/uverbs: Fix potential leak of resources->collection in flow_resources_alloc()
RDMA/bnxt_re: Avoid exposing umdbr to userspace
RDMA/rtrs: guard against null kobj name
RDMA/bnxt_re: check create_singlethread_workqueue() in DCB setup
IB/isert: wait for deferred control PDU completions before releasing the connection
IB/iser: reject a remote invalidation of an unregistered direction
RDMA/srp: Fix srp_remove_target()
IB/mlx4: Fix use-after-free on pkey sysfs registration failure
RDMA/uverbs: Fix mmap_lock/disassociation_lock circular dependency
RDMA/core: Reject unregistering netdevs in ib_get_eth_speed
...
|
|
Cross-merge networking fixes after downstream PR (net-7.3-rc3).
Conflicts:
drivers/net/dsa/mt7530.c
3c18e3c9a54e ("net: dsa: mt7530: populate lpi_interfaces to fix EEE support")
10d9d8328e8a ("net: dsa: mt7530: replace mt7530_read with regmap_read")
Adjacent changes:
drivers/net/bonding/bond_alb.c
1746ef2e2df2 ("bonding: use skb_cow_head() in bond_do_alb_xmit() and rlb_arp_xmit()")
4cef95f72bbd ("bonding: fix u32 overflow in compute_gap()")
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
siw_get_hdr() can receive an extended DDP/RDMAP header across more than
one TCP callback. The first callback may receive most of the header,
while the next one still limits the copy to hdrlen - MIN_DDP_HDR instead
of the number of missing bytes. This makes the destination move past the
end of the header and overwrite the receive state, including
fpdu_part_rcvd. A later callback can then use a negative fpdu_part_rcvd
value as a copy offset, which creates an OOB write.
Use the number of header bytes already received when calculating the
next copy length.
Fixes: 754209850df8 ("RDMA/siw: Always consume all skbuf data in sk_data_ready() upcall.")
Signed-off-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr>
Link: https://patch.msgid.link/20260908085520.1746329-1-Jeremy.Jean@oss.cyber.gouv.fr
Assisted-by: Codex:gpt-6
Acked-by: Bernard Metzler <bernard.metzler@linux.dev>
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
The completion IRQ handler accesses the EQ state and DMA buffer. Its IRQ was
registered before that state was initialized, while teardown released the
buffer before free_irq() synchronized the handler.
Initialize the EQ without arming it, register the IRQ, and then arm it.
Reverse the resource order during teardown by freeing the IRQ before
destroying the EQ.
Fixes: 2a152512a155 ("RDMA/efa: CQ notifications")
Link: https://patch.msgid.link/20260907-use-after-free-of-admin-queue-struct-v1-2-dd9d9267fbf4@nvidia.com
Reviewed-by: Michael Margolin <mrgolin@amazon.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
|
|
The management IRQ handler accesses both the admin completion queue and the
async event queue. The driver registered the IRQ before constructing these
queues and destroyed them before freeing the IRQ, so the handler's lifetime
was not contained by the resources it accesses.
Initialize the queues with interrupts masked, request the IRQ, and then
switch to interrupt mode. On removal, reset the device and free the IRQ
before destroying the queues. Also reset the device before destroying the
queues if IRQ registration fails, because the device already has their DMA
addresses.
Fixes: b7f5e880f377 ("RDMA/efa: Add the efa module")
Link: https://patch.msgid.link/20260907-use-after-free-of-admin-queue-struct-v1-1-dd9d9267fbf4@nvidia.com
Reviewed-by: Michael Margolin <mrgolin@amazon.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
|
|
Drop words accidentally written twice, reported by checkpatch.pl as a
possible repeated word. Only touches comments, no code changes.
Assisted-by: Cursor:claude-opus-5
Signed-off-by: Hemanth Selam <hemanth.selam@gmail.com>
Link: https://patch.msgid.link/20260907065047.26773-3-hemanth.selam@gmail.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Fix typos in comments, reported by scripts/checkpatch.pl using the
misspelling list in scripts/spelling.txt. Only touches comments, no code
changes.
Assisted-by: Cursor:claude-opus-5
Signed-off-by: Hemanth Selam <hemanth.selam@gmail.com>
Link: https://patch.msgid.link/20260907065047.26773-2-hemanth.selam@gmail.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
Our HW always handles GMV BT pages with a fixed 4K size, but the driver
calculates the needed GMV BT pages number with PAGE_SIZE, which is 64K
in 64K system. Only the first 4K (GMV index 0-127) can be reached by HW,
causing GID capacity loss and memory waste.
Split a single 64K BT page into multiple 4K pages and register them to
HW per 4K block so that HW can correctly reach all GMV entries.
Fixes: 32053e584e4a ("RDMA/hns: Add support for filling GMV table")
Signed-off-by: Junxian Huang <huangjunxian6@hisilicon.com>
Link: https://patch.msgid.link/20260907084901.2420703-4-huangjunxian6@hisilicon.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
The GMV entry is a HW object corresponding to a GID. Since gid_table_len
is already limited to a maximum of 256, there is no need to allocate
memory for those extra GMV entries as they will never be touched.
Fixes: 7243396aaf12 ("RDMA/hns: Add a max length of gid table")
Signed-off-by: Junxian Huang <huangjunxian6@hisilicon.com>
Link: https://patch.msgid.link/20260907084901.2420703-3-huangjunxian6@hisilicon.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
The gid_table_len is always read from HW registers or derived from
u32 arithmetic, and it is never negative. Change its type from int
to u32 to avoid signed/unsigned mixed-type operations with other
u32 fields such as gmv_entry_num.
Signed-off-by: Junxian Huang <huangjunxian6@hisilicon.com>
Link: https://patch.msgid.link/20260907084901.2420703-2-huangjunxian6@hisilicon.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
iwpm_get_nlmsg_request() initializes refcount _after_ list_add_tail()
making it accessible to global list where another CPU can kref_get()
on nlmsg_request causing a refcount "addition on 0" bug. Fix this
by initializing kref _before_ list_add_tail() so refcount for
nlmsg_request can be incremented/decremented normally. In addition,
also initialize every field before list_add_tail().
Reported-by: syzbot+bd317784d628820741b5@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=bd317784d628820741b5
Fixes: 30dc5e63d6a5 ("RDMA/core: Add support for iWARP Port Mapper user space service")
Cc: stable@vger.kernel.org
Signed-off-by: Jeffin Philip <jeffinphilip14@gmail.com>
Link: https://patch.msgid.link/20260904131437.12917-1-jeffinphilip14@gmail.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
The only callers of irdma_sc_qp_flush_wqes() and
irdma_sc_mr_fast_register() always request their respective send queues to
be posted, so their post_sq false branches are unreachable.
Remove both arguments and post the queues unconditionally. Drop the
redundant QP flush request assignments as well.
Link: https://patch.msgid.link/20260903-64-bit-iova-is-silently-truncated-to-v1-2-96e878ea6873@nvidia.com
Reviewed-by: Kalesh AP <kalesh-anakkur.purayil@broadcom.com>
Reviewed-by: Jacob Moroni <jmoroni@google.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
|
|
ib_mr::iova is u64, but the fast-registration path passes it through
void * and uintptr_t. These conversions truncate the upper 32 bits on
32-bit kernels before the WQE is built.
Store the IOVA as u64 and write it directly to the WQE. The sole caller
always uses VA-based addressing, so remove the unused FBO selection and
set the VA-based bit unconditionally.
Fixes: b48c24c2d710 ("RDMA/irdma: Implement device supported verb APIs")
Link: https://patch.msgid.link/20260903-64-bit-iova-is-silently-truncated-to-v1-1-96e878ea6873@nvidia.com
Reviewed-by: Jacob Moroni <jmoroni@google.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
|