summaryrefslogtreecommitdiff
AgeCommit message (Collapse)Author
2026-09-14Merge branch 'for-next' of ↵fs-nextMark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs.git
2026-09-14Merge branch 'vfs.all' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs.git
2026-09-14Merge branch 'for-next' of https://git.kernel.org/pub/scm/fs/xfs/xfs-linux.gitMark Brown
2026-09-14Merge branch '9p-next' of https://github.com/martinetd/linuxMark Brown
2026-09-14Merge branch 'master' of ↵Mark Brown
https://github.com/Paragon-Software-Group/linux-ntfs3.git
2026-09-14Merge branch 'ntfs-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/ntfs.git
2026-09-14Merge branch 'nfsd-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux
2026-09-14Merge branch 'ksmbd-for-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/smb.git
2026-09-14Merge branch 'for-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/gfs2/linux-gfs2.git
2026-09-14Merge branch 'for-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/mszeredi/fuse.git
2026-09-14Merge branch 'dev' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/jaegeuk/f2fs.git
2026-09-14Merge branch 'for_next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/jack/linux-fs.git
2026-09-14Merge branch 'dev' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/exfat.git
2026-09-14Merge branch 'next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/teigland/linux-dlm.git
2026-09-14Merge branch 'configfs-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/leitao/linux.git
2026-09-14Merge branch 'cifs-next' of https://git.manguebit.org/linux.gitMark Brown
2026-09-14Merge branch 'for-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux.git
2026-09-14Merge branch 'nfsd-fixes' of ↵fs-currentMark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux
2026-09-14Merge branch 'fixes' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs.git
2026-09-14Merge branch 'next-fixes' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux.git
2026-09-14Merge branch 'vfs.fixes' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs.git
2026-09-14fs/ntfs3: zero frame pages on read errorWeiming Shi
ni_read_frame() marks every frame page uptodate at the 'out:' label regardless of the return value. The pages come from ntfs_lock_new_page() and are not zeroed, and several error paths reach 'out:' before anything is written to them (e.g. a failed decompress_lznt()/decompress_lzx_xpress() on a corrupted chunk, or an allocation failure). A page marked uptodate is served directly from the page cache, so a later read() of the file returns the uninitialized page contents to userspace. On a crafted compressed image this leaks kernel memory, including pointers. Zero each page on the error path before marking it uptodate. The success path is unchanged. Fixes: 4342306f0f0d ("fs/ntfs3: Add file operations and implementation") Reported-by: Xiang Mei <xmei5@asu.edu> Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Weiming Shi <bestswngs@gmail.com> Signed-off-by: Konstantin Komarov <almaz.alexandrovich@paragon-software.com>
2026-09-14fs/ntfs3: update ctime only on successful setxattrBaolin Liu
ntfs_setxattr() updates ctime and marks the inode dirty even when the xattr operation fails. Do that only on success. Fixes: 2d44667c306e ("fs/ntfs3: Update i_ctime when xattr is added") Signed-off-by: Baolin Liu <liubaolin@kylinos.cn> Signed-off-by: Konstantin Komarov <almaz.alexandrovich@paragon-software.com>
2026-09-14fs/ntfs3: take run_lock when freeing offs_folio in ni_decompress_fileWeiming Shi
attr_wof_frame_info() allocates, caches and dereferences ni->file.offs_folio under down_write(&ni->file.run_lock). ni_decompress_file() frees the same folio with folio_put() without taking run_lock. When a WOF externally-compressed file is opened for write, ni_decompress_file() can drop the last reference while a concurrent O_RDONLY reader is dereferencing the folio in attr_wof_frame_info(), leading to a use-after-free: BUG: KASAN: use-after-free in attr_wof_frame_info (fs/ntfs3/attrib.c:1619) Read of size 4 at addr ffff8880120c9000 by task exploit Call Trace: attr_wof_frame_info (fs/ntfs3/attrib.c:1619) ni_read_frame (fs/ntfs3/frecord.c:2421) ni_read_folio_cmpr (fs/ntfs3/frecord.c:1916) ntfs_read_folio (fs/ntfs3/inode.c:648) read_pages (mm/readahead.c:181) page_cache_ra_unbounded (mm/readahead.c:292) force_page_cache_ra (mm/readahead.c:364) page_cache_sync_ra (mm/readahead.c:573) filemap_get_pages (mm/filemap.c:2688) filemap_read (mm/filemap.c:2806) generic_file_read_iter (mm/filemap.c:2994) ntfs_file_read_iter (fs/ntfs3/file.c:842) vfs_read (fs/read_write.c:574) __x64_sys_pread64 (fs/read_write.c:773) Take run_lock around the folio_put(), as is done for every other access to ni->file.offs_folio. Fixes: 4342306f0f0d ("fs/ntfs3: Add file operations and implementation") Reported-by: Xiang Mei <xmei5@asu.edu> Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Weiming Shi <bestswngs@gmail.com> Signed-off-by: Konstantin Komarov <almaz.alexandrovich@paragon-software.com>
2026-09-13Merge tag 'x86_urgent_for_7.3-rc4' of ↵stableLinus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull x86 fixes from Dave Hansen: "The most notable fix is THP not silently losing user data and having been around for a couple of years. The main explanation I'd have for its longevity is that it requires a few different things to align at the same time: MADV_FREE, THP and heavy reclaim. - Fix user-space data loss with THP - Fix set_memory oopses - Fix addition of large constants in mul_u64_add_u64_div_u64() - Fix FineIBT hash offset in cfi_get_func_hash() - Fix PCI device reference counting in amd_smn_init()" * tag 'x86_urgent_for_7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: x86/amd_node: Fix PCI device reference counting in amd_smn_init() x86/div64: Fix addition of large constants in mul_u64_add_u64_div_u64() x86/cfi: Fix FineIBT hash offset in cfi_get_func_hash() x86/mm: Fix user-space data loss with MADV_FREE and THP x86/mm/pat: Allocate split page tables as kernel page tables x86/alternatives: Exclude text poking against change_page_attr() x86/mm/pat: Acquire init_mm read lock on attribute changes to avoid UAF x86/mm/pat: Acquire init_mm write lock on collapse to avoid UAF
2026-09-13lockd: Serialize block retries against host teardownChuck Lever
nlmsvc_grant_blocked() unlinks a block from nlm_blocked before it retries the lock, then re-inserts it. nlmsvc_traverse_blocks() skips a block that is not on nlm_blocked, so a teardown scan that runs during a retry passes it by and the retry puts it back. The surviving block pins its host. lockd warns that it could not shut down the host module, and the host outlives its network namespace. Hold the file's f_mutex across the retry, and extend the scan's hold across its unlink, so a scan and a retry of the same file can no longer interleave. Drop the mutex before releasing a block reference, since the last put takes f_mutex. A retry that waited out a scan re-checks under nlm_blocked_lock that its block is still queued and due. Reported-by: sashiko-bot <sashiko-bot@kernel.org> Closes: https://sashiko.dev/#/patchset/20260819162247.2970703-1-cel@kernel.org?part=1 Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260822-lockd-retry-blocked-uaf-v3-2-761661eae60c@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13lockd: Fix use-after-free in nlmsvc_retry_blockedChuck Lever
nlmsvc_retry_blocked() examines the block at the head of nlm_blocked under nlm_blocked_lock, then releases the lock before calling nlmsvc_grant_blocked() or retry_deferred_block(). The nlm_blocked list reference is all that keeps the block alive across that window. nlmsvc_grant_blocked() does take one of its own, but not until after the lock has been dropped. Unmounting the nfsd filesystem while a lock request is still blocked reaches nlmsvc_traverse_blocks(), which drops the list reference and frees the block along with the nlm_rqst hanging off it. BUG: KASAN: slab-use-after-free in nlm_async_call+0xd6/0x230 Read of size 8 at addr ffff88811b04c808 by task lockd/8377 nlm_async_call+0xd6/0x230 nlmsvc_retry_blocked+0x61c/0x800 lockd+0x144/0x1c0 Freed by task 8392: nlmsvc_release_block+0x231/0x290 nlmsvc_traverse_blocks+0x139/0x1b0 nlm_traverse_files+0x1aa/0xa00 nlmsvc_free_host_resources+0x12/0x60 nlm_shutdown_hosts_net+0x127/0x280 lockd_down+0xd5/0x1c0 Take a reference before releasing nlm_blocked_lock and drop it once the retry has run. Fixes: 0e4ac9d93515 ("lockd: handle fl_grant callbacks") Cc: stable@vger.kernel.org Reported-by: Shuangpeng Bai <shuangpeng.kernel@gmail.com> Closes: https://lore.kernel.org/linux-nfs/20260818235808.3458075-1-shuangpeng.kernel@gmail.com/ Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260822-lockd-retry-blocked-uaf-v3-1-761661eae60c@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: docs: Fix pNFS SCSI Kconfig symbolKarl Mehltretter
The pNFS SCSI layout server is controlled by NFSD_SCSILAYOUT. NFSD_SCSI has never existed. Fixes: f99d4fbdae67 ("nfsd: add SCSI layout support") Assisted-by: Codex:gpt-5.6-sol Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com> Acked-by: Randy Dunlap <rdunlap@infradead.org> Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260822081033.82098-1-kmehltretter@gmail.com Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13SUNRPC: Bypass sock_recvmsg() for the TLS control-record receiveChuck Lever
svc_tcp_recvfrom() parses the RPC record stream with ->read_sock, which calls neither security_socket_recvmsg() nor the sock:sock_recv_length tracepoint. svc_tcp_recv_cmsg() still goes through sock_recvmsg(), so an LSM mediates only the TLS control records on a server socket, and sock:sock_recv_length reports only those. Partial coverage is worse than none. It makes the RPC stream look mediated and observed when it is not. Until the record stream moved to ->read_sock, an LSM saw every octet NFSD read from a TCP socket. An SELinux policy that denies SOCKET__READ to NFSD blocked the receive. After this change no call on the server's TCP receive path consults an LSM, so that denial has no effect. Dispatch ->recvmsg directly so the whole receive path behaves one way. sock_recvmsg_nosec() reaches ->recvmsg through INDIRECT_CALL_INET(), so on a retpoline build the direct dispatch costs one indirect call per control record. Control records are rare on an established connection. Link: https://patch.msgid.link/20260821-tls-read-sock-2-v1-5-7ffce164eb45@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13SUNRPC: Receive RPC records with ->read_sockChuck Lever
svc_tcp_recvfrom() reads an RPC record with two recvmsg() calls, one for the four-octet fragment marker and one for the body. Neither supplies a control-message buffer, so kTLS reports a TLS control record on either one by raising MSG_CTRUNC and delivering nothing. Both receives have to recognize that outcome and hand off to a recovery path that re-reads the record. Commit bee47cb026e7 ("sunrpc: fix handling of server side tls alerts") built that recovery path. It stays on both receives for as long as one recvmsg() serves data and control alike. Instead, parse the record stream in a ->read_sock actor. The data path then carries no control-message buffer at all. A control record at the head stops ->read_sock short, so svc_tcp_recv_ctrl_record() classifies the head of the receive queue with MSG_PEEK before consuming anything. skb_copy_bits() walks an skb from its head to reach the copy offset, so an actor that copies a page at a time rewalks a large GRO or kTLS skb once per page. Copy each callback's share of the record into rq_bvec with one skb_copy_datagram_iter() instead. A non-final fragment carries four octets of marker and may carry no payload. sk_datalen advances only by the payload, so neither a run of empty fragments nor a run of one-octet fragments trips the sv_max_mesg check before desc->count runs out. Only desc->count bounds such a run, four or five octets at a time. Cap the fragments per socket-lock hold. recvmsg() skips an urgent octet and clears the condition, so the old receive path advanced past it. ->read_sock consumes what precedes the urgent octet and then stops there for good. The data path calls no recvmsg() now, so the record stream cannot advance again. An RPC stream carries no urgent data, so close the connection. The read that first reaches the mark returns a positive count, so a zero-length read does not mark the stop. Key the close on an incomplete record. Link: https://patch.msgid.link/20260821-tls-read-sock-2-v1-4-7ffce164eb45@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13SUNRPC: Flush a received record's pages once it is completeChuck Lever
svc_tcp_read_msg() flushes the destination pages after every receive. A record that arrives in several pieces is therefore flushed once per piece. A page spanning two pieces is flushed twice. Nothing reads the message body before the record is complete, so no reader needs the intermediate flushes. Flush every page of the record in one pass from svc_tcp_recvfrom(), once the last fragment has arrived. Remove svc_flush_bvec() and its ARCH_IMPLEMENTS_FLUSH_DCACHE_PAGE guard. The guard avoided setting up a bvec iterator, and a plain walk of rq_pages compiles away on its own where flush_dcache_page() is an empty inline. Link: https://patch.msgid.link/20260821-tls-read-sock-2-v1-3-7ffce164eb45@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13SUNRPC: Close the transport on an unhandled TLS record typeChuck Lever
svc_tcp_sock_recv_cmsg() drains a TLS record that is neither an alert nor application data, then returns -EAGAIN so the receive loop retries. It runs only on an established TLS session. A handshake record there carries a post-handshake message, and the server has no handler for one. Draining a KeyUpdate only delays the close. kTLS sets key_update_pending when it decrypts that record, and a later receive returns -EKEYEXPIRED. kTLS flags only a KeyUpdate, so any other post-handshake message disappears and the connection keeps running. Return -EPROTO for an unhandled record type. Any error but -EAGAIN closes the transport. An application data record keeps its -EAGAIN return. No NFS client is known to send a handshake record on an established connection. Link: https://patch.msgid.link/20260821-tls-read-sock-2-v1-2-7ffce164eb45@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13SUNRPC: Separate the TLS control-record receive from its policyChuck Lever
svc_tcp_sock_recv_cmsg() receives the record at the head of the kTLS receive queue and decides what its content type means for the transport. The receive and the decision are one step, so a caller cannot learn a record's type without also acting on it. A later caller classifies the head of the queue with MSG_PEEK before it decides whether to consume the record. It needs the receive without the decision. Move the receive into svc_tcp_recv_cmsg(), which takes the recvmsg() flags and reports the octet count, the record type, and the message flags. The alert policy stays in svc_tcp_sock_recv_cmsg(), the helper's only caller in this patch. The zeroed control buffer replaces the msg_controllen check. An unfilled buffer reports record type zero, so a positive octet count with no record type now returns -EBADMSG instead of a count. The alert parse runs on a second msghdr over the alert buffer, so the copy and the parse share no iterator state. The rewind that commit bee47cb026e7 ("sunrpc: fix handling of server side tls alerts") added by hand is no longer needed. Link: https://patch.msgid.link/20260821-tls-read-sock-2-v1-1-7ffce164eb45@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Resolve the recall-any mask names in the trace formatChuck Lever
RCA4_TYPE_MASK_* are enum constants, so the preprocessor cannot fold them into the print format that show_rca_mask() builds for the nfsd_cb_recall_any event. Nothing declares an eval map for them either, so trace_event_eval_update() has no substitution to apply at module load, and the event's format file ships the enumerator names verbatim. trace-cmd and perf cannot decode the bmval0 field. Declare the eval maps for the nine mask bits show_rca_mask() decodes. The format then carries the shift counts as integers, the same shape the SUNRPC trace points already emit from their BIT() flag decoders. Fixes: 638593be55c0 ("NFSD: add CB_RECALL_ANY tracepoints") Cc: stable@vger.kernel.org Reviewed-by: Jeff Layton <jlayton@kernel.org> Reviewed-by: Christoph Hellwig <hch@lst.de> Link: https://patch.msgid.link/20260818184151.31180-1-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Clean up header guards in fs/nfsd/xdr.hChuck Lever
Make the header guards less ambiguous about their provenance. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260818140035.12740-3-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Include <linux/nfs_fh.h> where struct nfs_fh is usedChuck Lever
struct nfsd4_copy embeds a struct nfs_fh, and nlm_fopen() reads the size and data fields of one. Neither fs/nfsd/xdr4.h nor fs/nfsd/lockd.c includes the header that defines the type; both reach it by way of nfsd.h, which pulls in <linux/nfs.h>. Add the direct include to both files, so nfsd.h can later drop the <linux/nfs.h> it carries for no use of its own. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260818140035.12740-2-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Move the nfs3.h include out of nfsd.hChuck Lever
fs/nfsd/nfsd.h is included throughout the server, yet it uses no NFSv3 protocol definition of its own. The <linux/nfs3.h> include there served only to make those definitions reach the few source files that need them, by way of nfsd.h itself or the xdr.h chain that pulls it in. Give each consumer its own include and drop the one in nfsd.h, so the header no longer carries a dependency unrelated to its contents. nfsfh.c, nfsctl.c, nfs3xdr.c, nfs3proc.c, and nfs2acl.c reference NFS3_* definitions directly; add <linux/nfs3.h> to each. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260818140035.12740-1-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Apportion CB_RECALL_ANY recalls among clientsChuck Lever
deleg_reaper() asks each eligible client to return one delegation whenever it runs, whether or not anything needs the memory. A delegation returned before it is needed costs the client an OPEN when it next touches the file. Nothing sizes the request either. The delegation scan callback discards nr_to_scan, which is reclaim's statement of how many objects it wants back. Record each delegation scan request in nfsd_deleg_backlog and pass the accumulated total to deleg_reaper(). Handing that total to every client would ask for it once per client, so scale it by each client's share of the delegations this sweep can reach. The count callback reports what is left after the outstanding requests, so concurrent reclaimers do not each ask for the same delegations. cl_ra_time keeps the next sweep from returning to the clients this one reached. Nothing is recalled until a scan arrives. nfs4_laundromat() is the exception. It has no scan request to pass, so it computes what must go for num_delegations to fall below max_delegations, and passes only this namespace's share. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-8-3b2cffce701e@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Pace the state shrinker's scan requestsChuck Lever
Currently, neither of the scan callback functions records anything before returning SHRINK_STOP, so the size of each scan request is discarded. That size is the only real measure NFSD gets of reclaim pressure. Both count callbacks report their population whether or not the reaper is already queued to reclaim it, so reclaim asks again for work that is pending. Accumulate each courtesy scan request in nfsd_shrink_backlog and subtract the backlog from what that count callback reports. The worker retires the backlog once courtesy_client_reaper() has run. That reaper expires the clients synchronously, so the discount covers exactly the interval the work is pending. Delegations need a different bound. This is because deleg_reaper() only sends CB_RECALL_ANY and does not track how many delegations were actually returned by the targeted client. Report the delegations only once NFSD_RECALL_ANY_COOLDOWN_SECS have passed since the last sweep. deleg_reaper() skips any client it recalled from within that window, so an earlier scan request cannot produce another recall. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-7-3b2cffce701e@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Give delegations their own state shrinkerChuck Lever
Since commit 44df6f439a17 ("NFSD: add delegation reaper to react to low memory condition"), nfsd_client_shrinker has managed two unrelated populations of objects. One population is courtesy clients. Shrinking that population can be done synchronously and without risk of deadlock. The shrinker callback could return a precise count of the number of objects that were released. The other population is delegations. Shrinking that population requires sending a CB_RECALL_ANY; clients are not obligated to return any delegation. The shrinker callback is structurally unable to report progress. What's more, the single shrinker callback falls back to delegation reaping only when there are no courtesy clients left to reclaim. A single courtesy client is enough to keep a namespace's delegations out of the count it reports. To begin to resolve these issues, refactor the existing state shrinker into two: one for courtesy clients and one for reaping delegations. Each manages the size of its own population, and the shrinker names become namespace-specific. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-6-3b2cffce701e@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Count delegations per network namespaceChuck Lever
The state shrinker is allocated per network namespace, but nfsd4_state_shrinker_count() reports num_delegations, which counts the delegations held by the whole host. Every namespace therefore reports every delegation on the server. Reclaim sees the population multiplied by the number of namespaces running NFSD. A namespace holding no delegations of its own still reports a nonzero count and queues its reaper, which then finds nothing to recall. Count the delegations in each namespace and report that instead. num_delegations stays for the admission check in __alloc_init_deleg() and the ceiling check in nfs4_laundromat(). Both compare against max_delegations, which is sized from host memory and so remains a host-wide limit. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-5-3b2cffce701e@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Send a meaningful CB_RECALL_ANY keep countChuck Lever
deleg_reaper() sets craa_objects_to_keep to zero on every CB_RECALL_ANY. RFC 8881 Section 20.6.3 defines that field as the number of objects the client may keep, leaving the client to choose which of the excess to return, because the server cannot read lack of recent use as lack of usefulness. Zero asks for every delegation the client holds, including the ones backing files an application still has open. There is also no reason NFSD has to reclaim the entire delegation working set on the first sign of memory pressure. Derive the keep count from cl_deleg_count so that each callback asks for one delegation. Both the shrinker and the laundromat re-arm while their condition lasts, so a client with more to give is asked again on the next pass. The Linux client ignores craa_objects_to_keep and returns unused delegations selected from the type mask alone, so the count changes nothing for it. Fixes: 44df6f439a17 ("NFSD: add delegation reaper to react to low memory condition") Cc: stable@vger.kernel.org Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-4-3b2cffce701e@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Name directory delegations in the CB_RECALL_ANY type maskChuck Lever
RFC 8881 Section 20.6.3 distinguishes an NFSv4.1 server implementation that shares one pool among all classes of recallable objects from one that keeps separate pools per class. NFSD falls in the former category. The CB_RECALL_ANY operation's craa_type_mask argument names the types of objects in the recallable resource pool, but NFSD's implementation does not name directory delegations, even though they are allocated through __alloc_init_deleg(), they are counted against the max_delegations budget, and the state shrinker reclaims them. Add RCA4_TYPE_MASK_DIR_DLG to craa_type_mask so clients that implement directory delegations consider them when choosing which delegations to return. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-3-3b2cffce701e@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Count the delegations held by each clientChuck Lever
struct nfs4_client records the delegations it holds on cl_delegations but keeps no count of them. deleg_reaper() walks nn->client_lru under nn->client_lock, but cl_delegations is serialized by nn->deleg_lock, which nests outside nn->client_lock. A caller there cannot take nn->deleg_lock to count the list. The cost tells against the walk as well: an O(n) count per client, on a pass that already visits every client. Add cl_deleg_count, maintained at the two sites that mutate cl_delegations. Both hold nn->deleg_lock, so the counter is already serialized against itself and needs no atomic of its own. The decrement sits below the delegation_hashed() test, next to the list_del_init it pairs with, so it runs only when the delegation really leaves the list. A reader that holds only nn->client_lock is not synchronized against either update site, so it can see a count that does not match the list. Such a reader marks the access with data_race() and may not depend on the value for correctness. No functional change. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-2-3b2cffce701e@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Do not send CB_RECALL_ANY to NFSv4.0 clientsChuck Lever
deleg_reaper() sends CB_RECALL_ANY to every ACTIVE client holding delegations, but CB_RECALL_ANY is an NFSv4.1 operation. An NFSv4.0 client's callback service accepts only CB_GETATTR and CB_RECALL, so it replies OP_ILLEGAL. The decoder maps the unexpected opnum to -EIO, and nfsd4_cb_done() marks the client's callback channel down. Nothing brings the channel back. nfsd4_run_cb_work() sets NFSD4_CB_UP only for a minor version above zero, and the only nfsd4_probe_callback() call site an NFSv4.0 client reaches is nfsd4_setclientid_confirm(). One visit from the reaper therefore leaves the channel marked down until the client re-establishes its clientid. RENEW then returns NFS4ERR_CB_PATH_DOWN for as long as the client holds delegations. nfsd4_cb_channel_good() stops returning true, so the client is granted no further delegations. Skip clients at minor version zero. Fixes: 44df6f439a17 ("NFSD: add delegation reaper to react to low memory condition") Cc: stable@vger.kernel.org Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-1-3b2cffce701e@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13nfsd: set op->status when an operation's header cannot be encodedChuck Lever
nfsd4_encode_operation() leaves op->status alone when the reply buffer has no room for the operation's opcode and status word. nfsd4_proc_compound() reads the unchanged nfs_ok as success and goes on to the next operation, so the reply counts an operation whose result was never encoded. Report the failure through nfsd4_check_resp_size(), which the rest of the function already uses. It returns NFS4ERR_REP_TOO_BIG, or NFS4ERR_REP_TOO_BIG_TO_CACHE on a session, and the COMPOUND ends at that operation. Two paths narrow the reply buffer: nfsd4_sequence(), which rejects a SEQUENCE result that does not fit, and nfsd4_encode_splice_read(), which can leave a single XDR word in the head page. Whether a COMPOUND reaches that boundary is unproven, so this is a guard rather than a fix. Link: https://patch.msgid.link/20260817-jean-v1-2-9e356596ab85@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13nfsd: preflight SEQUENCE replies before accepting a slotJérémy Jean
nfsd4_sequence() narrows the reply buffer to the session's cached reply limit before it accepts the slot seqid. A client may negotiate ca_maxresponsesize_cached down to NFSD_MIN_HDR_SEQ_SZ, and nfsd4_alloc_slot() then gives every slot a zero-length sl_data[]. A COMPOUND tag can fill that narrowed buffer until it holds the SEQUENCE opcode but not the status word that follows. nfsd4_encode_operation() returns without running nfsd4_encode_sequence(), so cstate.data_offset stays zero. It leaves op->status at nfs_ok as well, so the COMPOUND is treated as having succeeded. nfsd4_store_cache_entry() declines to cache a lone SEQUENCE that returned an error. That test reads the status the operation reported, so it passes here. The copy starts at offset zero and takes the whole reply, RPC and COMPOUND headers included, into the zero-length sl_data[]. The COMPOUND tag is copied along with it, so the client picks most of the bytes written past the end of the slot: BUG: KASAN: slab-out-of-bounds in read_bytes_from_xdr_buf+0x1bc/0x390 Write of size 80 at addr ffff888003a549cd by task kunit_try_catch/24 __asan_memcpy+0x38/0x60 read_bytes_from_xdr_buf+0x1bc/0x390 nfsd4_sequence_done+0x5b0/0x810 nfs4svc_encode_compoundres+0x1bf/0x240 Check that the fixed-size SEQUENCE result, plus room for a following operation's error status, fits the negotiated limit before narrowing the buffer and consuming the slot seqid. The slot and its reply cache are left unchanged, as RFC 8881 Section 2.10.6.1.2 requires of an error returned from SEQUENCE. Fixes: 47ee52986472 ("nfsd4: adjust buflen to session channel limit") Cc: stable@vger.kernel.org Assisted-by: Codex:gpt-5 Signed-off-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr> Link: https://patch.msgid.link/20260817-jean-v1-1-9e356596ab85@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Fail a pool_threads read whose reply does not fitChuck Lever
The reply to a pool_threads read is the list of per-pool thread counts, formatted into a buffer of SIMPLE_TRANSACTION_LIMIT bytes. snprintf() truncates its last write and strlen() measures only what fit, so a reply too long for that buffer ends mid-number with no terminating newline. A pool running 4096 threads is reported as 40. Nothing marks the reply as incomplete, so an administrator reads a plausible but wrong count. Take snprintf()'s return value, which reports the truncation strlen() cannot see, and fail the read with -ENAMETOOLONG when the list does not fit. That is the errno svc_one_xprt_name() already returns for the same condition. Suggested-by: David Laight <david.laight.linux@gmail.com> Fixes: eed2965af1ba ("[PATCH] knfsd: allow admin to set nthreads per node") Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260812193349.13347-1-david.laight.linux@gmail.com Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13SUNRPC: Reject a socket that already has an svc_sock attachedChuck Lever
Writing the same socket descriptor to /proc/fs/nfsd/portlist twice attaches a second svc_sock to one socket. svc_setup_socket() saves the socket's callbacks before installing its own, so the second attach records svc_write_space() as the old write_space callback. svc_udp_init() invokes that callback by way of svc_sock_setbufsize(), and svc_write_space() then calls itself until the kernel stack is exhausted: BUG: TASK stack guard page was hit at ffffc900037d7ff8 svc_write_space+0x90/0x2b0 net/sunrpc/svcsock.c:429 svc_write_space+0xe6/0x2b0 net/sunrpc/svcsock.c:430 ... 700 more ... svc_sock_setbufsize+0x18d/0x220 net/sunrpc/svcsock.c:386 svc_udp_init net/sunrpc/svcsock.c:854 [inline] svc_setup_socket+0xb2f/0x1090 net/sunrpc/svcsock.c:1498 svc_addsock+0x2fd/0x760 net/sunrpc/svcsock.c:1547 __write_ports_addfd fs/nfsd/nfsctl.c:742 [inline] write_ports+0xa5b/0xcc0 fs/nfsd/nfsctl.c:861 nfsctl_transaction_write+0x106/0x1a0 fs/nfsd/nfsctl.c:112 svc_data_ready() and svc_tcp_state_change() chain through their saved callbacks the same way, so a TCP descriptor added twice recurses on the next incoming segment instead. Reaching any of this takes a writer on portlist, and the nfsd filesystem sets no FS_USERNS_MOUNT, so the reproducer needs CAP_SYS_ADMIN in the initial user namespace. Reject a socket that already carries sk_user_data. svc_setup_socket() overwrites that field unconditionally, so a socket some other consumer has claimed is one NFSD would corrupt whether or not the callbacks recurse. Fixes: b41b66d63c73 ("[PATCH] knfsd: allow sockets to be passed to nfsd via 'portlist'") Cc: stable@vger.kernel.org Reported-by: syzbot+54cdc566f64abf51b7f1@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=54cdc566f64abf51b7f1 Link: https://patch.msgid.link/20260815162844.8219-1-cel@kernel.org Reviewed-by: Jeff Layton <jlayton@kernel.org> Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13sunrpc: honor the netlink unix_gid NEGATIVE flagAmeer Hamza
The unix_gid netlink upcall protocol has an explicit SUNRPC_A_UNIX_GID_NEGATIVE attribute for failed group lookups, and mountd sends it, but sunrpc_nl_parse_one_unix_gid() only allocates an empty group list for a flagged reply without propagating the flag into the entry, so it installs a valid positive entry with zero groups. Such an entry strips the uid of all supplementary groups until it is refreshed or expires: the same defect the previous patch fixes on the classic channel, on a transport that can say "lookup failed" explicitly. Set CACHE_NEGATIVE for flagged replies, as the netlink ip_map path already does for its negative flag. An empty GIDS list without the flag remains a positive entry. Fixes: 0850e8603cd7 ("sunrpc: add netlink upcall for the auth.unix.gid cache") Cc: stable@vger.kernel.org Assisted-by: Claude:claude-fable-5 Signed-off-by: Ameer Hamza <ameer.hamza@truenas.com> Link: https://patch.msgid.link/20260814221953.108837-3-ameer.hamza@truenas.com Signed-off-by: Chuck Lever <cel@kernel.org>