| Age | Commit message (Collapse) | Author |
|
https://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux.git
|
|
* for-7.4/io_uring:
io_uring: wait for in-flight requests on ring release
io_uring: drop registered files and buffers at release time
io_uring: run cancelations synchronously on ring release
io_uring/cancel: cancel and wait for all requests on process exit
io_uring/notif: count pending zerocopy notifications per ring
io_uring/uring_cmd: only cancel requests of the given task
io_uring: put request files before posting the completions
io_uring/rw: don't reap io-wq IOPOLL completions while io-wq has a reference
io_uring: post io-wq completions from the last request reference
io_uring/io-wq: put the request file before posting a completion
io_uring/io-wq: release raw spinlock before calling wake_up()
io_uring/rsrc: fix accounting for cloned buffers with huge pages
io_uring: check for room before posting the dummy skip CQE
io_uring/fdinfo: ignore IORING_CQE_F_32 in last CQ array slot
nvme/ioctl: call io_uring_cmd_set_res32() in ->end_io()
io_uring/cmd: split io_uring_cmd_set_res() from io_uring_cmd_done()
io_uring: move req_set_*() to public header
|
|
* block-7.3: (22 commits)
nvme-multipath: set BLK_FEAT_ZONED only after the zone info is known
nvme: fix command effects log lifetime for multipath heads
nvmet: don't allow I/O admission after percpu ns reference is killed
nvmet: defer setting ns->enabled to false in nvmet_ns_disable()
nvmet: copy the hostid into the ctrl before creating PR pc_refs
nvmet-auth: fix out-of-bounds write in nvmet_auth_challenge()
nvmet: pci-epf: reject too-short SGL segments
nvme-multipath: fix underflow in ANA log bounds checks
nvme: work around all -Wformat-security warnings
nvme: work around -Wformat-security warning
nvme: do not reset controllers in NVME_CTRL_NEW state
nvme-tcp: delay nvme_tcp_reclassify_socket()
Revert "nvme-tcp: lockdep: use dynamic lockdep keys per socket instance"
drbd: remove unused drbd_nl_mcgrps[] array
blk-mq: allow cached requests to be used for flush operations
blk-mq: set RQF_USE_SCHED when the operation is known
block: reject polled dio with user integrity metadata
selftests: ublk: fix unused_result error
blk-cgroup: save IRQ state in blkg_tryget_closest()
nvmet: preserve device path on allocation failure
...
|
|
The namespace head is marked zoned at allocation time based only on the
command set identifier. If the zone info query reports a zero zone size,
the path namespace is registered without zoned limits while the head
still advertises the zoned capability with a zone size of zero, and
reporting zones or writing to the head then shifts by ilog2(0).
Drop the zoned feature from the head allocation and let the head limits
refresh stack it in from the path namespace instead.
Fixes: 28982ad73d6a ("nvme: set BLK_FEAT_ZONED for ZNS multipath disks")
Fixes: 3838e80fcfb3 ("nvme: skip the zoned limits update if the zone info query failed")
Cc: stable@vger.kernel.org
Signed-off-by: Guixin Liu <kanie@linux.alibaba.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
KASAN reported a use-after-free when an I/O passthrough command was sent
through a multipath namespace head after the controller path that first
created the head had been removed:
BUG: KASAN: slab-use-after-free in nvme_command_effects+0x192/0x200 [nvme_core]
Read of size 4 at addr ffff888141b14400 by task nvme/19811
nvme_command_effects+0x192/0x200 [nvme_core]
nvme_cmd_allowed+0x7e/0x1b0 [nvme_core]
nvme_user_cmd.constprop.0+0x1b5/0x450 [nvme_core]
nvme_ns_head_chr_ioctl+0xf4/0x2a0 [nvme_core]
Move the log cache to the subsystem, with one entry per command set. Both
controllers and namespace heads hold subsystem references, keeping their
log pointers valid until subsystem release.
Fixes: be93e87e7802 ("nvme: support for multiple Command Sets Supported and Effects log pages")
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Yao Sang <sangyao@kylinos.cn>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
nvmet_req_find_ns() uses percpu_ref_get() to obtain a reference to
the namespace. However, percpu_ref_get() can acquire a reference even
after the namespace reference has been killed (or marked DEAD).
This is undesirable during namespace disable because nvmet_ns_disable()
kills the namespace reference and then waits for all outstanding
references to drain. Acquiring a new reference after the reference is
killed can therefore extend the namespace drain period.
Replace percpu_ref_get() in nvmet_req_find_ns() with
percpu_ref_tryget_live_rcu(), which only acquires a reference while
the namespace reference is still live. This handles the race where
nvmet_req_find_ns() finds ns is admitting I/O (or it's live) but before
it acquires the reference to ns, its reference is killed in
nvmet_ns_disable(). For instance check this race:
CPU0 CPU1
nvmet_req_find_ns(): nvmet_ns_disable():
xa_load() -> ns
IO_LIVE == set
xa_clear_mark()
percpu_ref_kill() // DEAD
percpu_ref_get() synchronize_rcu()
| wait_for_completion()
+-- succeed
Replacing percpu_ref_get() with percpu_ref_tryget_live_rcu() prevents
the I/O request from acquiring a namespace reference once the
reference has been marked DEAD.
Perform the namespace lookup and reference acquisition in
nvmet_req_find_ns() within an RCU read-side critical section.
nvmet_ns_disable() uses synchronize_rcu() before draining and exiting
the namespace reference, ensuring that RCU readers which may be
acquiring the namespace reference have completed before the reference
is exited.
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Nilay Shroff <nilay@linux.ibm.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
nvmet_ns_disable() currently clears ns->enabled before draining
in-flight I/O references. This allows namespace configuration to be
changed while existing I/O requests can still hold a reference to the
namespace.
This can race with configuration of namespace attributes such as the
device path, UUID, NGUID etc. These attributes can be accessed by I/O
requests without holding subsys->lock and must not be modified while
such requests are still using the namespace.
In nvmet_ns_disable(), keep ns->enabled set while existing namespace
references are being drained, so namespace configuration remains blocked
until all in-flight I/O has completed. Set ns->enabled to false only
after the namespace references have been drained and the namespace
device has been disabled.
Introduce the NVMET_NS_IO_LIVE flag, which is set after the namespace
is successfully enabled in nvmet_ns_enable(). When nvmet_ns_disable()
starts, clear NVMET_NS_IO_LIVE so that the I/O path stops admitting new
I/O once the flag is cleared. Using test_and_clear_bit() in
nvmet_ns_disable() also prevents concurrent callers from starting a
second disable operation.
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Nilay Shroff <nilay@linux.ibm.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
Commit 6202783184bf ("nvmet: Improve nvmet_alloc_ctrl() interface and
implementation") added a second uuid_copy() of args->hostid near the end
of nvmet_alloc_ctrl(), and commit 7b658153f1b8 ("nvmet: Remove duplicate
uuid_copy") removed the original copy that sat before
nvmet_ctrl_init_pr() instead of the new one.
Since then nvmet_ctrl_init_pr() snapshots ctrl->hostid into the
per-controller per-namespace reservation refs while the uuid_copy()
from the connect data runs later, after the controller is published.
The ctrl is allocated with kzalloc(), so every pc_ref created on this
path stores the nil UUID.
pc_ref->hostid has a single consumer: nvmet_pr_set_ctrl_to_abort()
matches it against the preempted registrant's hostid to kill and drain
the victim's in-flight I/O for Preempt and Abort. The match can never
hit with the nil UUID, so whenever a namespace with reservations
enabled exists before a host connects, which includes every reconnect,
Preempt and Abort silently degrades into a plain Preempt: the
preempting host sees success while the victim's in-flight I/O is still
in the air.
Copy the hostid where the rest of the connect data is consumed, before
the controller is published and before nvmet_ctrl_init_pr() takes its
snapshot.
Fixes: 7b658153f1b8 ("nvmet: Remove duplicate uuid_copy")
Cc: stable@vger.kernel.org
Reviewed-by: Nilay Shroff <nilay@linux.ibm.com>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Reviewed-by: Hannes Reinecke <hare@suse.de>
Signed-off-by: Guixin Liu <kanie@linux.alibaba.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
nvmet_auth_challenge() receives its output buffer as a void pointer.
sizeof(*d) therefore evaluates to one with GCC and undercounts the fixed
challenge header by 15 bytes. A short AUTH_RECEIVE buffer can pass the
check before the challenge is copied past the end of its allocation.
Use the typed challenge pointer when calculating the response size.
BUG: KASAN: slab-out-of-bounds in nvmet_execute_auth_receive (drivers/nvme/target/fabrics-cmd-auth.c:442 drivers/nvme/target/fabrics-cmd-auth.c:569)
Write of size 32 by task kworker/1:1H/64
Workqueue: nvmet_tcp_wq nvmet_tcp_io_work
Call Trace:
nvmet_execute_auth_receive (drivers/nvme/target/fabrics-cmd-auth.c:442 drivers/nvme/target/fabrics-cmd-auth.c:569)
nvmet_tcp_try_recv_pdu (drivers/nvme/target/tcp.c:1119 drivers/nvme/target/tcp.c:1243)
nvmet_tcp_io_work (drivers/nvme/target/tcp.c:1350 drivers/nvme/target/tcp.c:1382 drivers/nvme/target/tcp.c:1445)
process_one_work (kernel/workqueue.c:3322)
worker_thread (kernel/workqueue.c:3405 kernel/workqueue.c:3486)
kthread (kernel/kthread.c:436)
ret_from_fork (arch/x86/kernel/process.c:158)
ret_from_fork_asm (arch/x86/entry/entry_64.S:245)
The buggy address is located 16 bytes inside of
allocated 33-byte region [ffff888000554e80, ffff888000554ea1)
Kernel panic - not syncing: KASAN: panic_on_warn set ...
Cc: stable@vger.kernel.org
Fixes: db1312dd9548 ("nvmet: implement basic In-Band Authentication")
Reported-by: co+77553a7fc66ac133@bugs.sh
Closes: https://lore.kernel.org/all/Ah6QavQXwvshtqyUcYD1P1R7XE0ND9fB7lbV@bugs.sh/
Assisted-by: Codex:gpt-5.6
Reviewed-by: Christoph Hellwig <hch@lst.de>
Reviewed-by: Hannes Reinecke <hare@kernel.org>
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
The host controls an SGL segment's length. A nonzero length smaller than
one SGL descriptor makes nr_descs zero, so
nvmet_pci_epf_get_sgl_segment() reads the type byte of sgls[-1] outside
the allocated buffer. If that byte contains a segment descriptor type,
the function also copies the preceding 16 bytes and returns a negative
descriptor count.
The read was reproduced in 3/3 KASAN boots with only the host transfer
mocked by a KUnit test:
BUG: KASAN: slab-out-of-bounds in nvmet_pci_epf_get_sgl_segment
Read of size 1 by task kunit_try_catch
Call Trace:
nvmet_pci_epf_get_sgl_segment
nvmet_pci_epf_short_sgl_test
Kernel panic - not syncing: KASAN: panic_on_warn set ...
Reject segments that cannot hold one descriptor before allocating or
transferring their contents.
Fixes: 0faa0fe6f90e ("nvmet: New NVMe PCI endpoint function target driver")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
The number of ANA groups advertised through Identify Controller sizes the
ANA log buffer. The ANA log header and group descriptors independently
supply the number of groups and namespace IDs to parse.
Both bounds checks subtract an untrusted object size from ana_log_size
before comparing the current offset. If the object is larger than the
buffer, the size_t subtraction underflows and lets the parser read beyond
ana_log_buf.
Check the current offset before the first subtraction and compare each
object size with the remaining buffer instead. This rejects inconsistent
ANA data before dereferencing a truncated group descriptor or walking an
oversized namespace ID array.
Fixes: 0d0b660f214d ("nvme: add ANA support")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
The sysfs code passes two string variables into sysfs_emit(), which is safe
in this instance but causes the compiler to warn when -Wformat-security
is enabled:
host/sysfs.c: In function ‘cntrltype_show’:
host/sysfs.c:682:9: error: format not a string literal and no format arguments [-Werror=format-security]
682 | return sysfs_emit(buf, type[ctrl->cntrltype]);
Print these using a "%s" format like all other instances in the same file.
Reviewed-by: Nilay Shroff <nilay@linux.ibm.com>
Signed-off-by: Arnd Bergmann <arnd@arndb.de>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
Passing a string variable into dev_set_name() causes a warning
when building with -Wformat-security enabled:
drivers/nvme/host/core.c: In function 'nvme_cdev_add':
drivers/nvme/host/core.c:3911:9: error: format not a string literal and no format arguments [-Werror=format-security]
3911 | ret = dev_set_name(cdev_device, name);
Remove the temporary strings and let dev_set_name() do the
same thing internally.
Fixes: 26acdaa357cd ("nvme: fix crash and memory leak during invalid cdev teardown")
Reviewed-by: Nilay Shroff <nilay@linux.ibm.com>
Reviewed-by: John Garry <john.garry@linux.dev>
Reviewed-by: Damien Le Moal <dlemoal@kernel.org>
Signed-off-by: Arnd Bergmann <arnd@arndb.de>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
During NVMe controller creation, the controller is exposed to sysfs via
nvme_add_ctrl() while it is still in the NVME_CTRL_NEW state.
This creates a narrow race window where a userspace process can write to
the reset_controller sysfs node before the initialization thread
transitions the state to NVME_CTRL_CONNECTING.
If a reset is triggered during this window, the state machine allows the
transition from NVME_CTRL_NEW to NVME_CTRL_RESETTING, and the reset work
is queued. However, the original creation thread continues its execution,
subsequently moving the state to NVME_CTRL_CONNECTING and finally to
NVME_CTRL_LIVE.
When the delayed reset work finally executes, it attempts to tear down
the controller and transition the state to NVME_CTRL_CONNECTING. Because
the state is now NVME_CTRL_LIVE, this transition fails, triggering a
WARN_ON in the reset work.
Fix this by removing NVME_CTRL_NEW from the allowed prior states for
the NVME_CTRL_RESETTING transition.
Fixes: 8bfc3b4c6f9d ("nvmet: switch loopback target state to connecting when resetting")
Reported-by: syzbot+4a1d521d19d6321f5aad@syzkaller.appspotmail.com
Reviewed-by: Daniel Wagner <dwagner@suse.de>
Signed-off-by: Maurizio Lombardi <mlombard@redhat.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
Commit 841aee4d75f1 ("nvme-tcp: lockdep: annotate in-kernel sockets")
introduced nvme_tcp_reclassify_socket() to resolve a lockdep WARN by
distinguishing userspace sockets from nvme-tcp sockets. After that, TLS
encryption support was introduced to nvme-tcp and it added another lock
dependency on userspace socket handling by tlshd, which resulted in a
new lockdep WARN. This WARN was hidden by the commit 19bdb70c77d3
("nvme-tcp: lockdep: use dynamic lockdep keys per socket instance").
However, it turned out the commit 19bdb70c77d3 has a bug of lockdep key
lifetime management, and it is to be reverted by another patch in this
series. After the revert, the WARN was unveiled and observed at the
blktests test case nvme/062:
WARNING: possible circular locking dependency detected
7.3.0-rc2+ #455 Tainted: G W
------------------------------------------------------
tlshd/75867 is trying to acquire lock:
ffffffff91ada5e0 (fs_reclaim){+.+.}-{0:0}, at: __kmalloc_cache_noprof+0x64/0x6d0
but task is already holding lock:
ffff888197808258 (sk_lock-AF_INET-NVME){+.+.}-{0:0}, at: do_tcp_setsockopt+0x499/0x26a0
which lock already depends on the new lock.
To fix the lockdep WARN caused by the TLS encryption support, delay the
reclassification of the nvme-tcp sockets. Currently the reclassification
happens before the TLS handshake starts, which pulls in the additional
lock dependencies related to tlshd. Reclassify the sockets after the TLS
handshake instead, to cut that dependency.
Link: https://lore.kernel.org/lkml/CANn89i+wnTLC==UnXCpjsS4YxvEfhe5oK0N7fttbqr1zKyqdug@mail.gmail.com/
Reviewed-by: Nilay Shroff <nilay@linux.ibm.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: Shin'ichiro Kawasaki <shinichiro.kawasaki@wdc.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
This reverts commit 19bdb70c77d3b24239a453291299b64040bdba86.
The commit 19bdb70c77d3 ("nvme-tcp: lockdep: use dynamic lockdep keys
per socket instance") addressed the lockdep WARN caused by the circular
lock dependency among six locks:
set->srcu -> sk_lock -> cpu_hotplug_lock -> fs_reclaim -> q_usage_counter -> elevator_lock -> set->srcu
As its title says, the commit cut the dependency by introducing the
dynamic lockdep keys per socket instance. However, as described in the
Link tag URL, the commit made a wrong assumption: it assumed that
__fput_sync(queue->sock->file) in nvme_tcp_free_queue() would
synchronously destroy the socket. This is wrong: when in-flight packets
cause delayed free of the socket, the prematurely freed lockdep key is
referred to and causes another WARN. The commit is an imperfect fix.
Hence revert it.
To address the circular dependency among the six locks, another solution
is required. It is provided by the commit 0ba6912f7e97 ("Revert "once:
don't use a work queue to reset sleepable static key""). It cuts the
dependency between sk_lock and cpu_hotplug_lock. This solution is
simpler, and reduces the complexity in nvme-tcp.
Link: https://lore.kernel.org/lkml/CANn89i+wnTLC==UnXCpjsS4YxvEfhe5oK0N7fttbqr1zKyqdug@mail.gmail.com/
Reviewed-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Nilay Shroff <nilay@linux.ibm.com>
Signed-off-by: Shin'ichiro Kawasaki <shinichiro.kawasaki@wdc.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
nvmet_ns_device_path_store() frees the old path before allocating its
replacement, losing the existing configuration if allocation fails.
Allocate the new path before freeing the old one.
Fixes: a07b4970f464 ("nvmet: add a generic NVMe target")
Signed-off-by: Julian Sun <sunjunchao@bytedance.com>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
In nvme_fc_reconnect_or_delete(), the system attempts to remove
the controller when reconnect limits are reached by calling
nvme_delete_ctrl(). This call is currently wrapped in a WARN_ON().
If a user concurrently deletes the controller via sysfs, the
controller's state may have already transitioned out of the
CONNECTING state. When this happens, nvme_delete_ctrl() correctly
rejects the operation and returns -EBUSY. Because of the WARN_ON()
wrapper, this race condition triggers an unnecessary kernel warning
stack trace.
Remove the WARN_ON() macro around nvme_delete_ctrl();
this allows the system to safely ignore the -EBUSY return value
if the controller is already being deleted.
Signed-off-by: Maurizio Lombardi <mlombard@redhat.com>
Reviewed-by: Daniel Wagner <dwagner@suse.de>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
The Micron 4100AT fetches PRP Lists in 512 byte strides, but the driver
allocates small PRP List descriptors in 256 byte strides. A descriptor
in the last 256 bytes of a dmapool page makes the controller read past
the page: the access lands on an adjacent IOVA, the IOMMU faults the
command. That last block only became reachable with commit da9619a30e73
("dmapool: link blocks across pages"), so this affects v6.4 and later.
Align the small descriptor pool to 512 bytes, as commit ebefac564796
("nvme-pci: 512 byte aligned dma pool segment quirk") already does for
another controller with the same erratum. The last block then starts
at offset 0xE00 and the fetch stays inside the page.
Cc: <stable@vger.kernel.org> # 6.4.x: ebefac564796: nvme-pci: 512 byte aligned dma pool segment quirk
Cc: <stable@vger.kernel.org> # 6.4.x
Signed-off-by: Gaurav Sinha <gsinha@micron.com>
Signed-off-by: Bean Huo <beanhuo@micron.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
io_uring_cmd_set_res32() only performs loads and stores to the io_uring
request state, so it's safe to call in interrupt context. Move the call
from the nvme_uring_task_cb() task work to nvme_uring_cmd_end_io(). This
unifies the 2 places setting the NVMe status and result on the uring_cmd
and removes the need to pass them through struct nvme_uring_cmd_pdu,
saving 16 bytes and a couple memory accesses.
Signed-off-by: Caleb Sander Mateos <csander@purestorage.com>
Link: https://patch.msgid.link/20260909155848.2069290-4-csander@purestorage.com
Signed-off-by: Jens Axboe <axboe@kernel.dk>
|
|
Function dma_opt_mapping_size() implies from its name that it returns a
target or sweet spot DMA mapping size. However, it is just an upper limit
optimal DMA mapping size. Above this size, DMA mapping performance may
significantly degrade.
Rename to dma_max_opt_mapping_size() to reflect the real behaviour. Also
rename the internal DMA mapping symbols to align with this.
The DMA API documentation already described this behaviour properly (so
there is nothing to update).
Signed-off-by: John Garry <john.garry@linux.dev>
Link: https://lore.kernel.org/r/20260831093620.3481337-1-john.g.garry@oracle.com
Signed-off-by: Marek Szyprowski <m.szyprowski@samsung.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/kees/linux
Pull kmalloc_obj conversions from Kees Cook:
"Another run of the Coccinelle script for converting kmalloc()
family of allocations to kmalloc_obj() via the existing rules
in scripts/coccinelle/api/kmalloc_objs.cocci"
* tag 'kmalloc_obj-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/kees/linux:
treewide: refresh kmalloc_obj() conversions
drm/amd/display: Fix harmless type mismatch in allocation
|
|
This is another run of the Coccinelle script for converting kmalloc()
family of allocations to kmalloc_obj() via the existing rules in
scripts/coccinelle/api/kmalloc_objs.cocci
This catches both the set of kmalloc() uses added since the first
kmalloc_obj() conversions in v7.0 and adds a large group missed in the
first pass due to Coccinelle not interacting well with the cleanup.h
scoped_...() family of macros[1]. I worked around this with spatch's
"--macro-file" argument to a file with all the scoped_...() macros mapped
to Coccinelle's YACFE_ITERATOR[2] as that was the closest viable control
flow indicator I could find.
Build tested allmodconfig on x86, arm64, arm, loongarch, mips, powerpc,
riscv, and s390 with no new warnings.
Link: https://lore.kernel.org/lkml/202609021314.8A9C0B8@keescook/ [1]
Link: https://github.com/coccinelle/coccinelle/blob/master/standard.h [2]
Signed-off-by: Kees Cook <kees+treewide@kernel.org>
|
|
Pull NVMe fixes from Keith:
"- Harden the tcp host and target against malformed PDUs: reject C2HData
for a non-read command, bound an over-long PDU before copying it, and
reject unsolicited H2CData (Yehyeong, Shivam)
- Fix circular locking on TLS queues (Xixin)
- Fix a soft lockup when scanning sparse namespace ID space (Mohamed)
- Fix racy access to the FDP placement id array (Kanchan)
- RDMA host and target fixes for a double cleanup on the queue_rq
error path and a queue leak when the connect backlog is exceeded
(Xixin)
- Authentication fixes: drain the target's expiry work before the SQ
is freed, and release the DH-CHAP secret when parsing fails (Kazuki,
Xu Rao)
- Fix nvme-fc options double free when nvme_add_ctrl() fails (Niklas)
- Add missing SRCU grace period to nvme_alloc_ns() error path (Tristan)
- Skip zoned limits update when the zone info query failed (Chao)
- Reject enabling a target namespace with no device path (Seokgyu)
- Add opcode filtering for fault injection (Mohamed)
- Drop the kernel-doc comments from nvme-tcp.h (Randy)"
* tag 'nvme-7.3-2026-09-03' of git://git.infradead.org/nvme: (21 commits)
nvme-tcp.h: drop kernel-doc comments, fix a few descriptions
nvme-fc: fix double free of fabrics options when nvme_add_ctrl() fails
nvmet: reject namespace enable without device path
nvmet-auth: Synchronize timeout work during SQ teardown
MAINTAINERS: update nvme entry
nvmet-tcp: reject unsolicited H2CData PDUs
nvme-tcp: defer TLS inline send to io_work
nvmet-tcp: fix out-of-bounds write when receiving an over-long PDU
nvme-tcp: return -EPROTO for a C2HData on a write
nvmet: print namespace IDs as unsigned 32bit value
nvme: print namespace IDs as unsigned 32bit value
nvme: remove stale namespaces by NSID range during scan
nvme: add missing SRCU grace period in error path
nvme-fabrics: fix DHCHAP secret leak on parse failure
nvmet-rdma: fix queue leak when connect backlog is exceeded
nvme: add opcode filtering for fault injection
nvme: fix racy access to FDP placement id array
nvme: set ns->head in nvme_alloc_ns_head
nvme-rdma: fix -EIO cleanup order in queue_rq
nvme: skip the zoned limits update if the zone info query failed
...
|
|
nvmf_create_ctrl() owns the fabrics options and frees them whenever
->create_ctrl() returns an error, so a transport must not free them on
its own error paths. nvme-fc tracks this by testing ctrl->ctrl.opts in
nvme_fc_ctrl_free(), which requires nvme_fc_init_ctrl() to clear that
pointer on every error exit.
The coupling is implicit, and commit 1a9e218195a5 ("nvme: split device
add from initialization") broke it by adding a second error exit. When
nvme_add_ctrl() fails, nvme_fc_init_ctrl() jumps to out_put_ctrl:, past
the "ctrl->ctrl.opts = NULL" that only sits on the fail_ctrl: path, so
nvme_fc_ctrl_free() frees the options and nvmf_create_ctrl() frees them
a second time:
BUG: KASAN: slab-use-after-free in nvmf_free_options+0x30/0x190
nvmf_free_options+0x30/0x190 drivers/nvme/host/fabrics.c:1284
nvmf_create_ctrl drivers/nvme/host/fabrics.c:1374 [inline]
Freed by task 5534:
nvme_fc_ctrl_free drivers/nvme/host/fc.c:2374 [inline]
nvme_fc_init_ctrl+0xe17/0x1450 drivers/nvme/host/fc.c:3605
nvme_add_ctrl() fails when dev_set_name() cannot allocate, so this is
reachable under memory pressure or fault injection. Without KASAN the
options are freed twice.
Rather than clear the pointer on the second exit as well, derive
ownership the way nvme-tcp, nvme-rdma and nvme-loop do, from list
membership: their free_ctrl leaves the options alone unless the
controller made it onto the transport list.
The list cannot simply be populated on the success path as it is there.
nvme-fc runs the initial connect synchronously via flush_delayed_work(),
and the controller has to be reachable on rport->ctrl_list for the whole
of it: nvme_fc_unregister_remoteport() needs to find it to signal
connectivity loss, nvme_fc_match_disconn_ls() matches an incoming
Disconnect Association LS against ctrl->association_id, which is only
assigned during that window, nvme_fc_resume_controller() needs it on
remoteport re-registration, and nvme_fc_existing_controller() uses it to
reject a duplicate connect racing the one in flight.
Keep the insertion where it is and add a fail_unlist: label, falling
into fail_ctrl:, for the error paths that run after it. The earlier
error paths never reach the insertion and keep using fail_ctrl:
directly, so the list is only touched where the controller is actually
on it.
nvme_fc_ctrl_free() cannot use the plain "goto free_ctrl" the other
transports use, because it still has to put_device(), release the rport
reference and free the ida entry for resources taken before the
insertion. Sample list_empty() under rport->lock instead.
ctrl->ctrl.opts also stays valid for the whole teardown now. That is
not the bug being fixed, but it removes some fragility around the old
idiom: nvme_free_ctrl() calls nvme_auth_free() before ->free_ctrl(), and
ctrl_max_dhchaps() dereferences ctrl->opts without a NULL check when
ctrl->dhchap_ctxs is set, which nvme-fc permits since NVMF_ALLOWED_OPTS
allows the dhchap options. The nvme sysfs attributes that dereference
ctrl->opts, such as hostnqn and address, evaluate their is_visible()
test once at device_add() time and stay readable until
cdev_device_del().
Fixes: 1a9e218195a5 ("nvme: split device add from initialization")
Cc: stable@vger.kernel.org
Reported-by: syzbot+f58e57380a6083c4041d@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=f58e57380a6083c4041d
Signed-off-by: Niklas Cassel <cassel@kernel.org>
Tested-by: Rihyeon Kim <rihyeon8648@gmail.com>
Reviewed-by: Hannes Reinecke <hare@kernel.org>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
A newly allocated namespace has a NULL device_path until userspace
configures the device_path attribute.
If buffered_io is enabled before device_path is configured,
nvmet_bdev_ns_enable() returns -ENOTBLK and nvmet_ns_enable() falls
back to nvmet_file_ns_enable(). The latter passes the NULL
device_path to filp_open(), causing a NULL pointer dereference in
getname_kernel().
Reject namespace enable when device_path has not been configured.
Reported-by: syzbot+f613f9f010ec98eb9d86@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=f613f9f010ec98eb9d86
Signed-off-by: Seokgyu Choi <tjrrb0313@gmail.com>
Reviewed-by: Sagi Grimberg <sagi@grimberg.me>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
nvmet_auth_sq_free() cancels auth_expired_work with
cancel_delayed_work(). If the work has already started, cancellation does
not wait for the callback. Transport teardown can consequently free or
reuse the queue containing struct nvmet_sq while
nvmet_auth_expired_work() still accesses that SQ.
Add a teardown-specific helper that synchronously drains the delayed work
before freeing authentication state, and use it from nvmet_sq_destroy().
Keep the non-synchronous helper for in-band authentication state cleanup,
where the SQ owner remains alive.
Fixes: 1a70200f404a ("nvmet-auth: expire authentication sessions")
Cc: stable@vger.kernel.org
Signed-off-by: Kazuki Hanai <hnkz.64@gmail.com>
Reviewed-by: Sagi Grimberg <sagi@grimberg.me>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
nvmet_tcp_handle_h2c_data_pdu() accepts an H2CData PDU after only checking
that its TTAG is a valid in-range command index and that the command's
data buffers are mapped. It never checks that the target has actually
solicited that data by sending an R2T for the command.
A remote host can abuse this. It submits a write command that takes the
R2T path and, before the target transmits the R2T, sends an H2CData PDU
for that command's tag. The data completes the command early, and when
the command then fails synchronously (e.g. a length mismatch caught by
nvmet_check_transfer_len()), it is completed a second time. Each
completion calls nvmet_tcp_queue_response(), so the same command is added
to queue->resp_list twice while it is still linked; the second llist_add()
makes the node point to itself (lentry->next == lentry).
nvmet_tcp_process_resp_list() then walks that self-referential node and
adds the command to resp_send_list twice. With CONFIG_DEBUG_LIST this
trips the "list_add double add" check (kernel BUG); without it the loop
never terminates and the nvmet_tcp workqueue wedges (soft-lockup). It is
remotely triggerable and needs no authentication on an allow_any_host
subsystem.
Track whether an R2T has been transmitted for a command and reject an
H2CData PDU that arrives before it. The flag is cleared on command reuse
(nvmet_tcp_get_cmd() zeroes cmd->flags) and stays set across the multiple
H2CData PDUs of a single solicited transfer.
Fixes: 872d26a391da ("nvmet-tcp: add NVMe over TCP target driver")
Cc: stable@vger.kernel.org
Reviewed-by: Sagi Grimberg <sagi@grimberg.me>
Signed-off-by: Shivam Kumar <kumar.shivam43666@gmail.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
blk_mq holds set->srcu while queuing and running requests. The kTLS
software send path takes ctx->tx_lock. lockdep knows that tx_lock
nests under elevator_lock which then waits on srcu, so an inline
send from that path under TLS triggers circular locking.
Skip the inline send optimization for TLS queues so the send runs
from the workqueue instead. The same workqueue already retries TLS
sends on write-space notifications. Plain TCP keeps the inline path.
Fixes: be8e82caa685 ("nvme-tcp: enable TLS handshake upcall")
Reviewed-by: Hannes Reinecke <hare@kernel.org>
Signed-off-by: Xixin Liu <liuxixin@kylinos.cn>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
nvmet_tcp_try_recv_pdu() reads a PDU header into the fixed 128-byte
queue->pdu union, then computes the remaining payload length as
queue->left = hdr->hlen - queue->offset + hdgst;
and reads that many more bytes into &queue->pdu + queue->offset, without
ever bounding the result against sizeof(queue->pdu).
A struct nvme_tcp_icreq_pdu is itself 128 bytes, exactly the size of the
union. Once a header digest has been negotiated (hdgst = 4), a second
ICReq passes the hlen == nvmet_tcp_pdu_size() check but yields
queue->left = 128 - 8 + 4 = 124, so bytes 8..132 are written into the
128-byte buffer -- 4 bytes past its end, over queue->hdr_digest and
queue->data_digest. Those bytes are attacker-controlled (an ICReq
carries no digest), and the duplicate ICReq is only rejected later,
after the overflow. A remote unauthenticated host can thus corrupt
kernel memory adjacent to the receive buffer.
Reject any PDU whose declared length would read past the end of
queue->pdu before the second recv.
Fixes: 872d26a391da ("nvmet-tcp: add NVMe over TCP target driver")
Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Shivam Kumar <kumar.shivam43666@gmail.com>
Cc: stable@vger.kernel.org
Reviewed-by: Sagi Grimberg <sagi@grimberg.me>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
The direction check in nvme_tcp_handle_c2h_data() returns -EIO. A
C2HData PDU naming a command that did not ask for data is a protocol
violation, and the check that rejects a PDU on those grounds a few
lines below it - SUCCESS set without LAST - returns -EPROTO.
No caller distinguishes the two, so this changes the error code alone.
Suggested-by: Sagi Grimberg <sagi@grimberg.me>
Signed-off-by: Yehyeong Lee <yhlee@isslab.korea.ac.kr>
Reviewed-by: Sagi Grimberg <sagi@grimberg.me>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
struct nvmet_ns.nsid is a u32, but a few messages print it with %d.
An NSID larger than 0x7fffffff is rendered as a negative number, which
is misleading in general and particularly so for the configfs messages
that echo back the NSID the user just asked for.
For example:
[ T200] nvmet: adding nsid -16 to subsystem mysubsystem
Print them with %u. The invalid-NSID error in nvmet_ns_make() keeps its
%#x because the two values it rejects, 0 and NVME_NSID_ALL, are more
readable in hex format. No functional change other than how the NSID is
formatted.
Fixes: a07b4970f464 ("nvmet: add a generic NVMe target")
Fixes: c6925093d0b2 ("nvmet: Optionally use PCI P2P memory")
Fixes: 5a47c2080a73 ("nvmet: support reservation feature")
Signed-off-by: Mohamed Khalfella <mkhalfella@purestorage.com>
Reviewed-by: Sagi Grimberg <sagi@grimberg.me>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
NSIDs are 32-bit unsigned values, but a number of log messages print
them with %d. An NSID larger than 0x7fffffff is rendered as a negative
number, which is confusing in the kernel log and makes the message hard
to correlate with the namespace it talks about. Sparse NSID spaces
where high NSIDs are common are the most likely to hit this.
The nsid sysfs attribute has the same problem, and there it is worse
because userspace parses the value.
For example:
$ grep . /sys/class/block/nvme0*/nsid
/sys/class/block/nvme0c0n1/nsid:10
/sys/class/block/nvme0c0n2/nsid:-16
/sys/class/block/nvme0c0n3/nsid:11
/sys/class/block/nvme0c0n4/nsid:-2000000016
/sys/class/block/nvme0n1/nsid:10
/sys/class/block/nvme0n2/nsid:-16
/sys/class/block/nvme0n3/nsid:11
/sys/class/block/nvme0n4/nsid:-2000000016
$
Print all of them with %u. Several messages in these files, including
two in zns.c right next to the ones being changed, already use %u, so
this only makes the rest consistent with them. No functional change
other than how the NSID is formatted.
Fixes: 2b9b6e86bca7 ("NVMe: Export namespace attributes to sysfs")
Fixes: 1d5df6af8c74 ("nvme: don't blindly overwrite identifiers on disk revalidate")
Fixes: ed754e5deeb1 ("nvme: track shared namespaces")
Fixes: 9ad1927a3bc2 ("nvme: always search for namespace head")
Fixes: 71010c309454 ("nvme: implement multiple I/O Command Set support")
Fixes: 2f4c9ba23b88 ("nvme: export zoned namespaces without Zone Append support read-only")
Fixes: 0ec84df4953b ("nvme-core: check ctrl css before setting up zns")
Fixes: 2079f41ec6ff ("nvme: check that EUI/GUID/UUID are globally unique")
Fixes: ce8d78616a6b ("nvme: warn about shared namespaces without CONFIG_NVME_MULTIPATH")
Fixes: ac522fc6c316 ("nvme: don't reject probe due to duplicate IDs for single-ported PCIe devices")
Signed-off-by: Mohamed Khalfella <mkhalfella@purestorage.com>
Reviewed-by: Sagi Grimberg <sagi@grimberg.me>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
nvme_scan_ns_list() drops the stale namespaces in each gap in the
reported NSID list one NSID at a time. Every iteration calls
nvme_find_get_ns() to look the namespace up and removes it if it is
present. The loop runs once per NSID in the gap rather than once per
namespace actually present.
NSIDs are 32-bit, so a target with a sparse NSID space can make a
single gap spin the loop billions of times with nothing to remove.
watchdog: BUG: soft lockup - CPU#4 stuck for 26s!
Workqueue: nvme-wq nvme_scan_work [nvme_core]
RIP: 0010:__srcu_read_unlock+0xb/0x20
Call Trace:
nvme_find_get_ns+0x7d/0xb0 [nvme_core]
nvme_scan_ns_list+0xe8/0x280 [nvme_core]
nvme_scan_work+0x18a/0x280 [nvme_core]
process_one_work+0x197/0x380
worker_thread+0x2fe/0x410
kthread+0xe0/0x100
Rename nvme_remove_invalid_namespaces() to nvme_remove_nsid_range()
and give it an open (start, end) NSID range. ctrl->namespaces is
sorted by NSID, so the whole gap is dropped in a single walk that
stops once end is reached. This bounds the work by the namespaces
that are present instead of by the size of the gap.
Fixes: 540c801c65eb ("NVMe: Implement namespace list scanning")
Signed-off-by: Mohamed Khalfella <mkhalfella@purestorage.com>
Reviewed-by: Sagi Grimberg <sagi@grimberg.me>
Reviewed-by: Randy Jennings <randyj@purestorage.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
nvme_alloc_ns() error path at out_unlink_ns removes ns from the
namespace head siblings list with list_del_rcu(&ns->siblings) but
does not wait for SRCU readers before freeing the namespace struct.
Multipath code iterates the head->list under srcu_read_lock() in
nvme_find_path() and nvme_mpath_revalidate_paths(), so a concurrent
reader can still hold a reference to ns when kfree(ns) runs.
The normal removal path in nvme_ns_remove() correctly calls
synchronize_srcu(&ns->head->srcu) after list_del_rcu() to wait for
in-progress readers. Add the same grace period in the error path.
Fixes: ed754e5deeb1 ("nvme: track shared namespaces")
Cc: stable@vger.kernel.org
Signed-off-by: Tristan Madani <tristan@talencesecurity.com>
Reviewed-by: Sagi Grimberg <sagi@grimberg.me>
Reviewed-by: John Garry <john.g.garry@oracle.com>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
nvmf_parse_options() duplicates dhchap_secret and dhchap_ctrl_secret
with match_strdup() before validating the DHHC-1: representation.
If validation fails, the parser returns -EINVAL before the temporary
string in p is assigned to opts->dhchap_secret or
opts->dhchap_ctrl_secret. nvmf_create_ctrl() subsequently frees opts,
but nvmf_free_options() cannot release the unassigned temporary string.
Each rejected option therefore leaks one allocation.
This is easy to miss because valid secrets transfer ownership to opts
and are freed normally, while the malformed-secret path still returns
the expected -EINVAL to userspace.
With CONFIG_NVME_HOST_AUTH enabled, the leak is reachable before the
required-option checks and transport lookup. No NVMe-oF target or
working transport connection is required; for example, repeatedly
writing
dhchap_secret=BAD
or
dhchap_ctrl_secret=BAD
to /dev/nvme-fabrics deterministically takes the leaking parse path.
Free the temporary string before leaving both validation error paths.
Use kfree_sensitive() because the copied option may contain secret
material even when its representation is rejected, matching the
sensitive cleanup used for stored DHCHAP secrets.
Fixes: f50fff73d620 ("nvme: implement In-Band authentication")
Cc: stable@vger.kernel.org
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Xu Rao <raoxu@uniontech.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/vkoul/dmaengine
Pull dmaengine updates from Vinod Koul:
"Core:
- New API to combine configuration and preparation and users
New hardware support:
- Mediatek MT8189 SoC uart dma support
Updates:
- Designware dma driver flatten desc structures and simplify code,
interrupt-path groundwork changes, first part of PCI EP DMA support
- Updates to zynqmp_dma with runtime PM and device removal
improvments
- Xilinx dma optimizations for AXIDMA and MCDMA channel management"
* tag 'dmaengine-7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/vkoul/dmaengine: (73 commits)
dmaengine: dw-edma: Mark emulated IRQ as level-triggered
dmaengine: idxd: assign all engines to group 0 in IAA defaults
dmaengine: qcom_hidma: remove conditional return with no effect
dmaengine: qcom-bam-dma: fix autosuspend cleanup during removal
dmaengine: fsl-edma: tracing: no ptr dereference during log output
dmaengine: dw-edma: Program endpoint function numbers
dmaengine: dw-edma-pcie: Add chip flags to match data
dmaengine: dw-edma-pcie: Handle optional data blocks
dmaengine: dw-edma-pcie: Factor out descriptor block address lookup
dmaengine: dw-edma-pcie: Add register offset match flag
dmaengine: dw-edma-pcie: Add platform ops to match data
dmaengine: dw-edma-pcie: Rename vsec_data to dma_data
dmaengine: dw-edma-pcie: Add capability match data
dmaengine: dw-edma-pcie: Track non-LL mode in DMA data
dmaengine: dw-edma: Add partial channel ownership mode
dmaengine: dw-edma: Initialize IRQ data before requesting IRQs
dmaengine: dw-edma: Add core quiesce operations
dmaengine: dw-edma: Add per-channel interrupt routing control
dmaengine: dw-edma: Factor out HDMA interrupt setup helper
dmaengine: dw-edma: Defer channel IRQ handling to workqueue
...
|
|
Pull RDMA updates from Jason Gunthorpe:
"About the normal size, still a lot of AI bug fixes and so on, but some
interesting new functionality too:
- Assorted locking, bounds-checking, cleanup, and error-path fixes
across UCMA/CMA, bng_re, bnxt_re, cxgb4, EFA, ERDMA, HFI1, HNS,
ionic, iRDMA, mlx4/mlx5, RXE, SIW, SRP/SRPT, and iSER target.
- netlink report for max # of supported resources
- get_zeroed_page()/etc removal
- Robust udata for ionic
- Allow unique RDMA device names per network namespace
- Completion counters and v2 admit queue support for EFA
- UC QP support for MANA
- Completion timestamps for ionic
- Harden uverbs data validation and resource lifetime handling,
fixing several core use-after-free conditions.
- bnxt_re toggle-page ownership and lifetime bug fixes
- dmabuf SRQ support for mlx5"
* tag 'for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/rdma/rdma: (160 commits)
RDMA/ucma: Allow path records to exactly fit the output buffer
RDMA/uverbs: Guard legacy bundles without method_elm
RDMA/efa: Add support for 128B admin v2 SQ entry
RDMA/efa: Generalize the admin SQ
RDMA/efa: Decouple admin command payload from admin header
RDMA/rxe: Fix OOB in free_rd_atomic_resources()
RDMA/cma: Fix WARNING in res_to_rt
RDMA/cxgb4: Free debugfs on registration failure
RDMA/cxgb4: Cancel reg_work before freeing device on remove
RDMA/ucma: Lock the handler in ucma_set_ib_path()
RDMA/ucma: Lock the handler in ucma_write_cm_event()
RDMA/erdma: restrict the driver to little-endian systems
RDMA/ionic: Embed counter driver data in rdma_counter allocation
RDMA/ionic: Cap eq_count to the eth driver's interrupt vector budget
RDMA/siw: Fix use-after-free in siw_accept()
IB/isert: post the full-feature receive buffers after session registration
IB/isert: delay the final Login Response until the session is registered
RDMA/srp: fix heap information leak on a truncated SRP_CRED_REQ
RDMA/erdma: Hold QP references for AE and CM processing
RDMA/erdma: Hold CQ references when processing EQ events
...
|
|
When pending disconnecting queues exceed the backlog limit, the
connect path only drops the device reference and leaks the newly
allocated queue and its IB resources.
Fixes: badc53620fe8 ("nvme: target: rdma: fix ndev refcount leak on queue connect")
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Xixin Liu <liuxixin@kylinos.cn>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
Currently NVMe fault injection applies to every command routed through
nvme_should_fail(), which makes it hard to target a specific command
type when reproducing an issue in error-handling paths.
Add an "opcode" debugfs attribute alongside the existing "status" and
"dont_retry" knobs. It defaults to 0xffff, meaning "match any opcode"
and preserving the previous behavior. When set to a valid opcode
(<= 0xff), fault injection is only considered for commands whose opcode
matches.
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Mohamed Khalfella <mkhalfella@purestorage.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
nvme_query_fdp_info() is called per-path and therefore prone to races.
It populates head->nr_plids/head->plids for fdp registration.
But nothing protects that pair from concurrent access - two paths scanning
the same namespace can race to populate it.
Avoid the race by moving this initialization work to nvme_alloc_ns_head()
which is called once per shared namespace.
Fixes: 30b5f20bb2dd ("nvme: register fdp parameters with the block layer")
Reported-by: Hari Mishal <harimishal1@gmail.com>
Link: https://lore.kernel.org/linux-nvme/20260725135111.14041-2-harimishal1@gmail.com/
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
so that it becomes possible to submit non-admin commands.
This is a prep patch with no functional changes.
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
On -EIO, the RDMA queue_rq path reports a host path error and then
still cleans up the command and unmaps the SQE DMA. The path error
helper completes the request, so that is double cleanup and DMA unmap
after the request is already complete.
Unmap the SQE first, then report the host path error. Skip the outer
command cleanup on that path.
Fixes: 62eca39722fd ("nvme-rdma: handle nvme_rdma_post_send failures better")
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Xixin Liu <liuxixin@kylinos.cn>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
nvme_query_zone_info() returns either a negative errno or a positive
NVMe status code, but nvme_update_ns_info_block() only tests for the
negative case:
ret = nvme_query_zone_info(ns, lbaf, &zi);
if (ret < 0)
goto out;
If the device fails the Identify Namespace (I/O Command Set specific)
command, or the Identify Controller command issued by
nvme_set_max_append(), the positive status falls through and setup
continues with the zero-initialized zone info. nvme_update_zone_info()
then marks the queue zoned with chunk_sectors and ns->head->zsze set to
zero.
blk_validate_zoned_limits() does not check chunk_sectors, so the limits
commit succeeds. blk_revalidate_disk_zones() does reject the zero zone
size, but by then the limits are live and nothing rolls them back, so
I/O keeps being submitted to a zoned queue with a zero zone size and
disk_zone_no() shifts by ilog2(0):
nvme0n1: Invalid non power of two zone size (0)
UBSAN: shift-out-of-bounds in include/linux/blkdev.h:747:16
shift exponent -1 is negative
disk_zone_no include/linux/blkdev.h:747 [inline]
bio_straddles_zones include/linux/blkdev.h:1058 [inline]
blk_zone_wplug_handle_write block/blk-zoned.c:1423 [inline]
blk_zone_plug_bio.cold+0x25/0x1c8 block/blk-zoned.c:1605
blk_mq_submit_bio+0x18fb/0x2870 block/blk-mq.c:3196
submit_bh_wbc+0x575/0x740 fs/buffer.c:2824
__block_write_full_folio+0x728/0xdd0 fs/buffer.c:1933
Any device, firmware or NVMe-oF target that fails this one command
reaches this.
Skip the zoned limits update in that case, and log which of the two
things happened: during a revalidation the queue keeps the zone
geometry it was last validated with, and on a first scan the namespace
is registered without zoned limits, so that it is still available as a
handle for admin commands. Neither of the paths in
nvme_query_zone_info() that return a positive status logs anything, so
the failure would otherwise be silent.
zi.zone_size is an exact indicator: every path that returns a positive
status returns before it is assigned, and after that the only failure
left is -ENODEV, which the caller already handles.
Found by FuzzNvme.
Fixes: c85c9ab926a5 ("nvme: split nvme_update_zone_info")
Cc: stable@vger.kernel.org
Cc: Weidong Zhu <weizhu@fiu.edu>
Suggested-by: Keith Busch <kbusch@kernel.org>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Chao Shi <coshi036@gmail.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
nvme_tcp_handle_c2h_data() finds the request by command id and checks
that it has a payload, but it does not check that the command asked for
data to be read. A controller that answers a write command with C2HData
therefore reaches nvme_tcp_recv_data(), where _copy_to_iter() hits
WARN_ON_ONCE(i->data_source) and returns 0. The receive path turns that
into -EFAULT and resets the controller.
No data is copied, so this is not memory corruption. What a controller
gets is a kernel warning it can raise at will, which is fatal on a host
booted with panic_on_warn.
The send path already knows the direction - it consults rq_data_dir()
when it builds a command - and nvme_tcp_handle_r2t() checks the length
and the offset of the request it names. The C2HData path does not check
the direction at all.
Reject a C2HData PDU whose command is not a read. Rejecting it fails
the command and resets the controller, as the neighbouring check in this
function does; what goes away is the warning.
[ 6.885580] ------------[ cut here ]------------
[ 6.886457] WARNING: lib/iov_iter.c:193 at _copy_to_iter+0x289/0x1330, CPU#0: kworker/0:1H/71
[ 6.888137] CPU: 0 UID: 0 PID: 71 Comm: kworker/0:1H Not tainted 7.2.0-rc5-NVMETCP-gf5098b6bae76 #1 PREEMPT(lazy)
[ 6.891165] Workqueue: nvme_tcp_wq nvme_tcp_io_work
[ 6.891875] RIP: 0010:_copy_to_iter+0x289/0x1330
[ 6.903739] Call Trace:
[ 6.904085] <TASK>
[ 6.909254] __skb_datagram_iter+0x433/0x820
[ 6.911026] skb_copy_datagram_iter+0x37/0x120
[ 6.911622] nvme_tcp_recv_skb+0xa07/0x4320
[ 6.913378] __tcp_read_sock+0x1ab/0x810
[ 6.915788] nvme_tcp_try_recv+0x152/0x1e0
[ 6.918222] nvme_tcp_io_work+0x1e4/0x6c0
[ 6.926906] </TASK>
[ 6.927226] ---[ end trace 0000000000000000 ]---
[ 6.927878] nvme nvme0: queue 1 failed to copy request 0x71 data
[ 6.928709] nvme nvme0: receive failed: -14
Fixes: 3f2304f8c6d6 ("nvme-tcp: add NVMe over TCP host driver")
Cc: stable@vger.kernel.org
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Yehyeong Lee <yhlee@isslab.korea.ac.kr>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
for-7.3/block
Pull NVMe updates from Keith:
"- Enable context analysis for the nvme host driver, annotating the
subsystem's locks, along with the LIST_HEAD_GUARDED support it needs
(Nilay, Marco)
- Harden the tcp host and target against malformed PDUs and out of
range SGL lengths (Yehyeong, Ibrahim, Greg)
- Fix unserialized page_frag_cache use in nvme-tcp request setup
(Dmitry)
- Bound identify, FDP and passthrough descriptor parsing to the
allocated buffers (Hari, Guixin)
- Zoned namespace fixes for host and the target (Xixin, Guixin, Yao)
- Apple controller fixes: page aligned admin queue buffers, NVMMU TCB
setup, DMA direction and admin queue teardown (Sven, Gui-Dong)
- Add a namespace level debugfs directory exposing reservation state,
and ABI documentation for the host sysfs and target configfs
interfaces (Guixin)
- Fix cdev and namespace lifetimes (John)
- Parallelize nvme-rdma I/O queue allocation and startup (Surabhi)
- Fix nvmet-rdma response resource leak on queue teardown (Shin'ichiro)
- Authentication fixes: AUTH_RECEIVE buffer and an out of bounds read
in negotiate (Xixin, Bryam, Guixin, Eric)
- Fix pci-epf use-after-free and CQ reference leak (Shin'ichiro, Yifei)
- Reject passthrough of driver managed Set Features (Chao)
- Various error path and teardown fixes across the host and target
addressing issues with use-after-free and leaking resources (Guixin,
Maurizio, Ewan, Zhengrong, Jiang HongHui, Myeonghun, Yang, Geliang,
Yehyeong)
- Various cleanups and typo fixes (Nilay, Guixin, Pan Chuang)"
* tag 'nvme-7.3-2026-08-13' of git://git.infradead.org/nvme: (81 commits)
nvmet: fix max_qid race between configfs and controller allocation
nvme: nvme-fc: Fix nvme_fc_create_hw_io_queues() queue deletion in error path
nvme: ratelimit the completion-path messages driven by device data
nvme-tcp: fix host memory disclosure on R2T for a read command
nvme-tcp: do not accept C2HData based on blk_rq_payload_bytes() alone
nvme-tcp: reject a read that transferred too few bytes
nvmet: zns: reject full zone report when buffer is too small
nvme-tcp: fix usage of page_frag_cache
nvme: reject passthrough of driver-managed Set Features
nvmet: fix NULL pointer dereference in nvmet_execute_identify_ns_zns()
nvmet: pci-epf: fix use-after-free in nvmet_pci_epf_exec_iod_work()
nvmet: pci-epf: put CQ ref on create_cq mapping failure
nvme-apple: Drop the PRP null check chicken bit
nvme-apple: Require page aligned buffers on the admin queue
nvme: Add a quirk for page aligned admin queue buffers
nvme-apple: Never set the opcode in the NVMMU TCB
nvme-apple: Don't set a DMA direction for commands without a data transfer
nvme-apple: Destroy the admin queue on removal
nvmet: fix heap out-of-bounds read in nvmet_auth_negotiate()
nvme: raise FDP placement handle cap to U8_MAX and warn on overflow
...
|
|
The function nvmet_subsys_attr_qid_max_store() can race against
nvmet_alloc_ctrl() when a subsystem's max_qid limit is modified.
Suppose max_qid is currently 64. If nvmet_alloc_ctrl() executes:
ctrl->sqs = kzalloc_objs(struct nvmet_sq *, subsys->max_qid + 1);
and at this exact point, a userspace process changes max_qid to 128,
nvmet_subsys_attr_qid_max_store() will set the new max_qid value. It
attempts to delete active controllers to force a reconnect, but the
new controller won't be deleted because it hasn't been added to the
subsys->ctrls list yet.
nvmet_alloc_ctrl() then proceeds and adds the new controller to the
subsys->ctrls list. Later, when nvmet_install_queue() is called, it
will see max_qid set to 128, but the memory allocated for sqs is only
sized for 64 entries. This results in a KASAN out-of-bounds warning
and potential memory corruptions.
Fix this by protecting the queue allocations and list insertion in
nvmet_alloc_ctrl() with down_read(&nvmet_config_sem). Because
nvmet_subsys_attr_qid_max_store() acquires down_write(&nvmet_config_sem)
to modify the attribute, this safely prevents the configfs writer from
modifying max_qid during controller creation.
Copy the max_qid from the subsystem to the controller's structure
during the allocation; ctrl->max_qid never changes as long as the
controller remains in LIVE state, so this will prevent similar race
conditions.
Fixes: 3e980f5995e0 ("nvmet: expose max queues to configfs")
Reported-by: syzbot+2626e846cd2585c9aa67@syzkaller.appspotmail.com
Signed-off-by: Maurizio Lombardi <mlombard@redhat.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
nvme_fc_create_hw_io_queues() will call __nvme_fc_delete_hw_queue() for the
last queue on which __nvme_fc_create_hw_queue() reported an error when deleting
all the io queues if they cannot all be created. This is incorrect since the
last queue did not actually get created.
The most recent change to this code was commit 17a1ec08ce70 ("nvme/fc: simplify
error handling of nvme_fc_create_hw_io_queues") which moved the cleanup to the
delete_queues: label and changed the loop bounds, however the code was not
correct prior to this change in a different way. The original commit
e399441de911 ("nvme-fabrics: Add host support for FC transport") had a
different error which called __nvme_fc_delete_hw_queue() on queue index 0 which
is used for the admin queue.
Fix this by correcting the initial loop index when deleting the io queues.
Fixes: 17a1ec08ce70 ("nvme/fc: simplify error handling of nvme_fc_create_hw_io_queues")
Fixes: e399441de911 ("nvme-fabrics: Add host support for FC transport")
Cc: stable@vger.kernel.org
Assisted-by: Claude:claude-opus-4-6
Reviewed-by: Maurizio Lombardi <mlombard@redhat.com>
Reviewed-by: Laurence Oberman <loberman@redhat.com>
Reviewed-by: Justin Tee <justin.tee@broadcom.com>
Signed-off-by: Ewan D. Milne <emilne@redhat.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|
|
nvme_find_rq() and nvme_handle_cqe() print an unratelimited message for
every completion queue entry whose command id does not resolve to an
in-flight request. Both are reached from the completion interrupt path
(nvme_irq() -> nvme_poll_cq() -> nvme_handle_cqe()) and the decision to
print is made entirely from device-supplied data, so a controller that
posts a stream of bogus command ids drives unbounded printk from hard
interrupt context.
This is not hypothetical. A single boot under an emulated controller
that posts invalid completions produced 846 "could not locate request
for tag 0x0", 846 "invalid id 0 completed on queue 2" and 123 "genctr
mismatch" lines. Once the tag set has been torn down every subsequent
completion resolves to nothing, so the print rate is bounded only by how
fast the device can post entries.
Ratelimit the three messages. The information they carry is diagnostic
and repeats, so the suppression count printed by the ratelimit helpers
is enough to tell that the condition persists. This matches how the
other device-driven error prints in the driver are already handled, for
example the status messages in nvme_log_error() and
nvme_log_err_passthru().
nvme_find_rq() lives in nvme.h and is shared by pci, tcp, rdma, apple and
target-loop, so all transports are covered.
Found by FuzzNvme.
Signed-off-by: Chao Shi <coshi036@gmail.com>
Signed-off-by: Keith Busch <kbusch@kernel.org>
|