summaryrefslogtreecommitdiff
path: root/net/sunrpc
AgeCommit message (Collapse)Author
9 hoursMerge branch 'headers' of git://git.infradead.org/users/willy/pagecache.gitMark Brown
# Conflicts: # drivers/gpu/drm/amd/amdkfd/kfd_migrate.c # net/ceph/osd_client.c
13 hoursMerge branch 'nfsd-next' of ↵Mark Brown
https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux
9 daysSUNRPC: Skip xpt_reserved accounting for non-UDP transportsChuck Lever
The xpt_reserved counter exists for UDP socket-buffer back-pressure. svc_udp_has_wspace() is the only has_wspace implementation that consults it, so on TCP and RDMA the counter is maintained and never read. svc_handle_xprt() adds to it once per RPC. svc_reserve() shrinks it again on each call from svc_process_common(), from svc_xprt_release(), and from each proc function that calls svc_reserve_auth(). Every shrinking call also runs svc_xprt_resource_released(), which issues an smp_mb() and can enqueue the transport. Add an xcl_flags field to svc_xprt_class and set SVC_XPRT_FLAG_WSPACE_RESERVE on the UDP class. Gate the xpt_reserved accounting on that flag. After the change, svc_reserve() no longer calls svc_xprt_resource_released() on TCP and RDMA. Two paths still cover that enqueue. svc_xprt_release() reaches the helper through svc_xprt_release_slot(), and svc_xprt_received() enqueues a transport whose XPT_DATA remains set. Link: https://patch.msgid.link/20260828135036.796842-4-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
9 dayssvcrdma: Clear XPT_DATA when the last receive context is consumedChuck Lever
svc_rdma_wc_receive() and svc_rdma_wc_read_done() set XPT_DATA after adding a completed context to sc_rq_dto_q or sc_read_complete_q. svc_rdma_recvfrom() dequeues one context and leaves XPT_DATA set, so the svc_xprt_received() that follows re-enqueues the transport and svc_xprt_enqueue() dispatches a second thread. That thread finds both queues empty and returns zero. Recheck the receive queues after each dequeue and clear XPT_DATA when the last context is taken, rather than only when a dequeue finds nothing. svc_xprt_received()'s kernel-doc no longer describes every transport, so relax its note about when XPT_DATA is cleared. Measured on one NFSv4.2 connection over 100GbE RoCE. A 4KB random read at queue depth 1 falls from 2.997 transport dequeues per RPC to 1.998, and from 96,629 to 82,398 server cycles per RPC. A 256KB random write falls from 3.270 dequeues to 2.004, and from 259,936 to 246,059 cycles. Each dispatch removed is worth about 10,000 cycles. The gain shrinks as the receive queues fill, since a leftover XPT_DATA then dispatches a thread that finds real work. An 8KB random write at queue depth 512 already runs at the two dequeues an RPC with a Read chunk requires, and shows no change. Throughput moves only where the server has no idle CPU to absorb the saving, so only the queue depth 1 read gains, by 1.8%. One dispatch per RPC remains. svc_rdma_send_ctxt_put() sets XPT_DATA to schedule a drain of sc_send_release_list, and svc_rdma_recvfrom() does not service that list. Link: https://patch.msgid.link/20260828135036.796842-3-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
9 dayssvcrdma: Grant credits from the clamped sc_max_requestsChuck Lever
svc_rdma_accept() computes sc_fc_credits from sc_max_requests before the Receive Queue depth is checked against the device's max_qp_wr. When that check lowers sc_max_requests, the credit grant keeps the original value, so the server advertises more credits than it has Receives posted. A client that uses the full grant overruns the Receive Queue, and the connection is lost with an RNR error. Set sc_fc_credits after the clamp so the grant matches the number of Receives the server posts. Fixes: fc2e69db82c1 ("svcrdma: Clean up comment in svc_rdma_accept()") Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260828135036.796842-2-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
9 daysSUNRPC: Bypass sock_recvmsg() for the TLS control-record receiveChuck Lever
svc_tcp_recvfrom() parses the RPC record stream with ->read_sock, which calls neither security_socket_recvmsg() nor the sock:sock_recv_length tracepoint. svc_tcp_recv_cmsg() still goes through sock_recvmsg(), so an LSM mediates only the TLS control records on a server socket, and sock:sock_recv_length reports only those. Partial coverage is worse than none. It makes the RPC stream look mediated and observed when it is not. Until the record stream moved to ->read_sock, an LSM saw every octet NFSD read from a TCP socket. An SELinux policy that denies SOCKET__READ to NFSD blocked the receive. After this change no call on the server's TCP receive path consults an LSM, so that denial has no effect. Dispatch ->recvmsg directly so the whole receive path behaves one way. sock_recvmsg_nosec() reaches ->recvmsg through INDIRECT_CALL_INET(), so on a retpoline build the direct dispatch costs one indirect call per control record. Control records are rare on an established connection. Link: https://patch.msgid.link/20260821-tls-read-sock-2-v1-5-7ffce164eb45@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
9 daysSUNRPC: Receive RPC records with ->read_sockChuck Lever
svc_tcp_recvfrom() reads an RPC record with two recvmsg() calls, one for the four-octet fragment marker and one for the body. Neither supplies a control-message buffer, so kTLS reports a TLS control record on either one by raising MSG_CTRUNC and delivering nothing. Both receives have to recognize that outcome and hand off to a recovery path that re-reads the record. Commit bee47cb026e7 ("sunrpc: fix handling of server side tls alerts") built that recovery path. It stays on both receives for as long as one recvmsg() serves data and control alike. Instead, parse the record stream in a ->read_sock actor. The data path then carries no control-message buffer at all. A control record at the head stops ->read_sock short, so svc_tcp_recv_ctrl_record() classifies the head of the receive queue with MSG_PEEK before consuming anything. skb_copy_bits() walks an skb from its head to reach the copy offset, so an actor that copies a page at a time rewalks a large GRO or kTLS skb once per page. Copy each callback's share of the record into rq_bvec with one skb_copy_datagram_iter() instead. A non-final fragment carries four octets of marker and may carry no payload. sk_datalen advances only by the payload, so neither a run of empty fragments nor a run of one-octet fragments trips the sv_max_mesg check before desc->count runs out. Only desc->count bounds such a run, four or five octets at a time. Cap the fragments per socket-lock hold. recvmsg() skips an urgent octet and clears the condition, so the old receive path advanced past it. ->read_sock consumes what precedes the urgent octet and then stops there for good. The data path calls no recvmsg() now, so the record stream cannot advance again. An RPC stream carries no urgent data, so close the connection. The read that first reaches the mark returns a positive count, so a zero-length read does not mark the stop. Key the close on an incomplete record. Link: https://patch.msgid.link/20260821-tls-read-sock-2-v1-4-7ffce164eb45@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
9 daysSUNRPC: Flush a received record's pages once it is completeChuck Lever
svc_tcp_read_msg() flushes the destination pages after every receive. A record that arrives in several pieces is therefore flushed once per piece. A page spanning two pieces is flushed twice. Nothing reads the message body before the record is complete, so no reader needs the intermediate flushes. Flush every page of the record in one pass from svc_tcp_recvfrom(), once the last fragment has arrived. Remove svc_flush_bvec() and its ARCH_IMPLEMENTS_FLUSH_DCACHE_PAGE guard. The guard avoided setting up a bvec iterator, and a plain walk of rq_pages compiles away on its own where flush_dcache_page() is an empty inline. Link: https://patch.msgid.link/20260821-tls-read-sock-2-v1-3-7ffce164eb45@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
9 daysSUNRPC: Close the transport on an unhandled TLS record typeChuck Lever
svc_tcp_sock_recv_cmsg() drains a TLS record that is neither an alert nor application data, then returns -EAGAIN so the receive loop retries. It runs only on an established TLS session. A handshake record there carries a post-handshake message, and the server has no handler for one. Draining a KeyUpdate only delays the close. kTLS sets key_update_pending when it decrypts that record, and a later receive returns -EKEYEXPIRED. kTLS flags only a KeyUpdate, so any other post-handshake message disappears and the connection keeps running. Return -EPROTO for an unhandled record type. Any error but -EAGAIN closes the transport. An application data record keeps its -EAGAIN return. No NFS client is known to send a handshake record on an established connection. Link: https://patch.msgid.link/20260821-tls-read-sock-2-v1-2-7ffce164eb45@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
9 daysSUNRPC: Separate the TLS control-record receive from its policyChuck Lever
svc_tcp_sock_recv_cmsg() receives the record at the head of the kTLS receive queue and decides what its content type means for the transport. The receive and the decision are one step, so a caller cannot learn a record's type without also acting on it. A later caller classifies the head of the queue with MSG_PEEK before it decides whether to consume the record. It needs the receive without the decision. Move the receive into svc_tcp_recv_cmsg(), which takes the recvmsg() flags and reports the octet count, the record type, and the message flags. The alert policy stays in svc_tcp_sock_recv_cmsg(), the helper's only caller in this patch. The zeroed control buffer replaces the msg_controllen check. An unfilled buffer reports record type zero, so a positive octet count with no record type now returns -EBADMSG instead of a count. The alert parse runs on a second msghdr over the alert buffer, so the copy and the parse share no iterator state. The rewind that commit bee47cb026e7 ("sunrpc: fix handling of server side tls alerts") added by hand is no longer needed. Link: https://patch.msgid.link/20260821-tls-read-sock-2-v1-1-7ffce164eb45@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
9 daysSUNRPC: fold xs_sock_process_cmsg() into its only callerChuck Lever
xs_sock_process_cmsg() switches on the TLS record type, and every arm but TLS_RECORD_TYPE_ALERT returns the -EAGAIN its caller passed in. The DATA arm clears MSG_EOR in the caller's msghdr, but xs_sock_recvmsg() has already cleared that flag before the call. Deriving the record type a second time inside the helper also fires trace_tls_contenttype() twice for every alert. Move the alert handling into xs_sock_recv_cmsg() and delete the helper. Every other record type still returns -EAGAIN. The DATA arm's account of MSG_EOR moves to xs_sock_recvmsg(), where the flag is cleared. Signed-off-by: Chuck Lever <cel@kernel.org> Signed-off-by: Anna Schumaker <anna.schumaker@hammerspace.com>
9 daysSUNRPC: treat every client-side TLS error alert as fatalChuck Lever
xs_sock_process_cmsg() decides whether an alert ends the session by reading the alert's level octet. RFC 8446 Section 6 retired that field. The severity is implicit in the description, and a receiver treats every alert listed in Section 6.2 as an error alert "regardless of the AlertLevel in the message". A peer that aborts with unexpected_message but leaves the legacy octet set to warning makes the client return -EAGAIN. xs_stream_data_receive() wakes no pending task for that error, so RPC Calls queued on a dead TLS session wait for their timeouts to expire. Decide from the alert description instead. close_notify and user_canceled are the closure alerts (RFC 8446 Section 6.1). Every other description ends the session, including one this kernel does not recognize. Fixes: 39067dda1d86 ("SUNRPC: Use new helpers to handle TLS Alerts") Signed-off-by: Chuck Lever <cel@kernel.org> Signed-off-by: Anna Schumaker <anna.schumaker@hammerspace.com>
9 daysSUNRPC: reject a client-side TLS alert record that is not two octetsChuck Lever
tls_alert_recv() reads two octets from the kvec it is handed and does not check the length (net/handshake/alert.c). xs_sock_process_cmsg() calls it for any alert record, and the alert[] buffer that xs_sock_recv_cmsg() supplies carries no initializer. A one-octet alert body leaves the description read from uninitialized stack and reported through trace_tls_alert_recv(). The peer controls that length. Neither tls_rx_msg_size() nor tls_rx_one_record() enforces the two-octet Alert payload. A TLS 1.3 record carrying only the inner content-type octet decrypts to a zero-length payload. RFC 8446 Section 5.1 requires a record with an Alert type to carry exactly one message, so any other length is malformed. RFC 9289 Section 5 bars RPC-with-TLS from negotiating a version below TLS 1.3, so no other alert framing applies. Require exactly two octets before parsing and return -EACCES otherwise. xs_stream_data_receive() already treats -EACCES as a fatal alert and reports it to the pending tasks. Gate the path on a control message rather than a positive count so that a zero-length record reaches the check. Fixes: cc5d59081fa2 ("sunrpc: fix client side handling of tls alerts") Signed-off-by: Chuck Lever <cel@kernel.org> Signed-off-by: Anna Schumaker <anna.schumaker@hammerspace.com>
9 daysSUNRPC: Reject a socket that already has an svc_sock attachedChuck Lever
Writing the same socket descriptor to /proc/fs/nfsd/portlist twice attaches a second svc_sock to one socket. svc_setup_socket() saves the socket's callbacks before installing its own, so the second attach records svc_write_space() as the old write_space callback. svc_udp_init() invokes that callback by way of svc_sock_setbufsize(), and svc_write_space() then calls itself until the kernel stack is exhausted: BUG: TASK stack guard page was hit at ffffc900037d7ff8 svc_write_space+0x90/0x2b0 net/sunrpc/svcsock.c:429 svc_write_space+0xe6/0x2b0 net/sunrpc/svcsock.c:430 ... 700 more ... svc_sock_setbufsize+0x18d/0x220 net/sunrpc/svcsock.c:386 svc_udp_init net/sunrpc/svcsock.c:854 [inline] svc_setup_socket+0xb2f/0x1090 net/sunrpc/svcsock.c:1498 svc_addsock+0x2fd/0x760 net/sunrpc/svcsock.c:1547 __write_ports_addfd fs/nfsd/nfsctl.c:742 [inline] write_ports+0xa5b/0xcc0 fs/nfsd/nfsctl.c:861 nfsctl_transaction_write+0x106/0x1a0 fs/nfsd/nfsctl.c:112 svc_data_ready() and svc_tcp_state_change() chain through their saved callbacks the same way, so a TCP descriptor added twice recurses on the next incoming segment instead. Reaching any of this takes a writer on portlist, and the nfsd filesystem sets no FS_USERNS_MOUNT, so the reproducer needs CAP_SYS_ADMIN in the initial user namespace. Reject a socket that already carries sk_user_data. svc_setup_socket() overwrites that field unconditionally, so a socket some other consumer has claimed is one NFSD would corrupt whether or not the callbacks recurse. Fixes: b41b66d63c73 ("[PATCH] knfsd: allow sockets to be passed to nfsd via 'portlist'") Cc: stable@vger.kernel.org Reported-by: syzbot+54cdc566f64abf51b7f1@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=54cdc566f64abf51b7f1 Link: https://patch.msgid.link/20260815162844.8219-1-cel@kernel.org Reviewed-by: Jeff Layton <jlayton@kernel.org> Signed-off-by: Chuck Lever <cel@kernel.org>
9 dayssunrpc: honor the netlink unix_gid NEGATIVE flagAmeer Hamza
The unix_gid netlink upcall protocol has an explicit SUNRPC_A_UNIX_GID_NEGATIVE attribute for failed group lookups, and mountd sends it, but sunrpc_nl_parse_one_unix_gid() only allocates an empty group list for a flagged reply without propagating the flag into the entry, so it installs a valid positive entry with zero groups. Such an entry strips the uid of all supplementary groups until it is refreshed or expires: the same defect the previous patch fixes on the classic channel, on a transport that can say "lookup failed" explicitly. Set CACHE_NEGATIVE for flagged replies, as the netlink ip_map path already does for its negative flag. An empty GIDS list without the flag remains a positive entry. Fixes: 0850e8603cd7 ("sunrpc: add netlink upcall for the auth.unix.gid cache") Cc: stable@vger.kernel.org Assisted-by: Claude:claude-fable-5 Signed-off-by: Ameer Hamza <ameer.hamza@truenas.com> Link: https://patch.msgid.link/20260814221953.108837-3-ameer.hamza@truenas.com Signed-off-by: Chuck Lever <cel@kernel.org>
9 dayssunrpc: treat empty auth.unix.gid replies as negative entriesAmeer Hamza
When rpc.mountd cannot resolve a uid (getpwuid() or getgrouplist() failure, e.g. while winbind or sssd is briefly unreachable), it answers the auth.unix.gid upcall with zero groups. unix_gid_parse() installs that as a valid positive entry, and svcauth_unix_set_client() then replaces the credential's group list with the empty one on every request, RPCSEC_GSS included via svcauth_gss_set_client(). One failed lookup strips that uid of all supplementary groups on every export for up to mountd's configured TTL (30 minutes by default), long after the NSS backend has recovered. mountd cannot send an empty list for a successful lookup, since getgrouplist(3) always includes at least the user's primary group, so a zero-group reply can only mean the lookup failed. Record it as a negative entry: unix_gid_find() then returns -ENOENT and svcauth_unix_set_client() keeps the groups the RPC credential already carries. This is the fallback that commit 3fc605a2aa38 ("[PATCH] knfsd: allow the server to provide a gid list when using AUTH_UNIX authentication") promised when no answer is available, and the same state try_to_negate_entry() already creates when no listener holds the channel open. Fixes: 3fc605a2aa38 ("[PATCH] knfsd: allow the server to provide a gid list when using AUTH_UNIX authentication") Cc: stable@vger.kernel.org Assisted-by: Claude:claude-fable-5 Signed-off-by: Ameer Hamza <ameer.hamza@truenas.com> Link: https://patch.msgid.link/20260814221953.108837-2-ameer.hamza@truenas.com Signed-off-by: Chuck Lever <cel@kernel.org>
11 daysSUNRPC: Resume receiving after a TLS control recordChuck Lever
A TLS control record delivers no payload to the RPC layer. svc_tcp_recvfrom() clears XPT_DATA before the receive, and svc_tcp_sock_recv_cmsg() returns -EAGAIN for the record it consumed. Nothing marks the transport ready again. kTLS raises data_ready for arriving TCP segments, not for records it has already decrypted. An RPC Call queued behind an alert or a KeyUpdate waits until the client sends more. The client blocks until its RPC timeout expires. The receive takes only the first two octets of the record. kTLS holds the remainder on its receive list, where each later receive takes two octets more. Drain a record that is not an alert, then mark the transport ready once a control record has been consumed. Fixes: 5e052dda121e ("SUNRPC: Recognize control messages in server-side TCP socket code") Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-4-62d9a631c880@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
11 daysSUNRPC: Treat every TLS error alert as fatalChuck Lever
svc_tcp_sock_process_cmsg() decides whether an alert ends the session by reading the alert's level octet. RFC 8446 Section 6 retired that field. The severity is implicit in the description, and a receiver treats every alert listed in Section 6.2 as an error alert "regardless of the AlertLevel in the message". A peer that aborts with unexpected_message but leaves the legacy octet set to warning makes the server return -EAGAIN. svc_tcp_recvfrom() then leaves a dead TLS session attached to an open transport. NFSD keeps polling it. Decide from the alert description instead. close_notify and user_canceled are the closure alerts (RFC 8446 Section 6.1). Every other description ends the session, including one this kernel does not recognize. Fixes: 39067dda1d86 ("SUNRPC: Use new helpers to handle TLS Alerts") Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-3-62d9a631c880@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
11 daysSUNRPC: Reject a TLS alert record that is not two octetsChuck Lever
tls_alert_recv() reads two octets from the kvec it is handed and does not check the length (net/handshake/alert.c). svc_tcp_sock_recv_cmsg() calls it for any positive receive, and the alert[] buffer it supplies carries no initializer. A one-octet alert body leaves the description read from uninitialized stack and reported through trace_tls_alert_recv(). The peer controls that length. Neither tls_rx_msg_size() nor tls_rx_one_record() enforces the two-octet Alert payload. A TLS 1.3 record carrying only the inner content-type octet decrypts to a zero-length payload. RFC 8446 Section 5.1 requires a record with an Alert type to carry exactly one message, so any other length is malformed. Require exactly two octets before parsing and return -EBADMSG otherwise. That closes the transport rather than acting on a partly uninitialized alert. Gate the path on a control message rather than a positive count so that a zero-length record reaches the check. Fixes: bee47cb026e7 ("sunrpc: fix handling of server side tls alerts") Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-2-62d9a631c880@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
11 daysSUNRPC: Do not credit control-record octets to the RPC streamChuck Lever
svc_tcp_sock_recv_cmsg() receives up to two octets into a local buffer, and returns that count for any record type other than TLS_RECORD_TYPE_ALERT. Nothing reached the caller's buffer, but svc_tcp_read_marker() adds the count to sk_tcplen and svc_tcp_read_msg()'s caller adds it to sk_datalen. The RPC stream advances over octets it never received. The fragment marker is assembled from stale sk_marker octets. The message body comes from pages nothing wrote. A conforming client reaches this. RFC 8446 Section 4.6.3 lets either peer send KeyUpdate once it has sent its Finished, and svcsock has no rekey path. kTLS leaves the partially consumed record on ctx->rx_list, so the body drains two octets per svc_tcp_recvfrom() call. Each pair is credited the same way. Return -EAGAIN for a record that is not an alert. That is what svc_tcp_sock_process_cmsg()'s default arm returned before the receive moved into a local buffer. Fixes: bee47cb026e7 ("sunrpc: fix handling of server side tls alerts") Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-1-62d9a631c880@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-04treewide: refresh kmalloc_obj() conversionsKees Cook
This is another run of the Coccinelle script for converting kmalloc() family of allocations to kmalloc_obj() via the existing rules in scripts/coccinelle/api/kmalloc_objs.cocci This catches both the set of kmalloc() uses added since the first kmalloc_obj() conversions in v7.0 and adds a large group missed in the first pass due to Coccinelle not interacting well with the cleanup.h scoped_...() family of macros[1]. I worked around this with spatch's "--macro-file" argument to a file with all the scoped_...() macros mapped to Coccinelle's YACFE_ITERATOR[2] as that was the closest viable control flow indicator I could find. Build tested allmodconfig on x86, arm64, arm, loongarch, mips, powerpc, riscv, and s390 with no new warnings. Link: https://lore.kernel.org/lkml/202609021314.8A9C0B8@keescook/ [1] Link: https://github.com/coccinelle/coccinelle/blob/master/standard.h [2] Signed-off-by: Kees Cook <kees+treewide@kernel.org>
2026-08-26Merge tag 'nfs-for-7.3-1' of git://git.linux-nfs.org/projects/trondmy/linux-nfsLinus Torvalds
Pull NFS client updates from Trond Myklebust: "Highlights include: Stable fixes: - Use-after-free fixes for the sunrpc client code - Delegation hash table leak - NULL dereference on lockowner allocation failure - Fix a handshake completion race in the TLS code - Fix an error sign checking issue when deciding whether the pNFS layout is still in use, or can be returned - Fix a layout segment leak in pnfs_layout_process() Other bugfixes: - Fix a missing NULL check in the rpcbind client - annotate shared socket callbacks with READ_ONCE/WRITE_ONCE - nfs_inode_set_delegation() error paths should return the delegation - Use clear_and_wake_up_bit() in nfs_clear_invalid_mapping() and the pNFS code. - Fix the nfs4_alloc_client() error paths to free the IDR allocation - fix folio dereference before NULL check in nfs_inode_remove_request() - Fix delayed delegation return - Fix another state manager race with umount - Fix device leaks on parse failure - Avoid cancelling in-flight I/O during a layout recall if the server doesn't require it - flexfiles: report cancelled I/O as a layout error - flexfiles: fix NULL dereference for NFSv4.0 data servers - Fix incorrect argument passed to nfs4_delete_lease() - Fix several symlink issues resulting from nfs_atomic_open_v23() - Fix an uninitialised variable issue in the NFSv4.1 callback code - fix LAYOUTSTATS send buffer exhaustion Features and cleanups: - NFSv4.2: Allow the server to specify that file data may not be cached - localio: optimise I/O submission when when not doing memory reclaim - localio: Remove duplicate wait code in nfs_local_commit - flexfiles: support loosely coupled NFSv4.x data servers - pNFS: key the data server cache on the NFS version" * tag 'nfs-for-7.3-1' of git://git.linux-nfs.org/projects/trondmy/linux-nfs: (33 commits) NFSv4.1: fix layout segment leak on the pnfs_layout_process() forget path NFSv4/pnfs: key the data server cache on the NFS version NFSv4.2: fix LAYOUTSTATS send buffer exhaustion pNFS: Fix EBUSY check in pnfs_layout_need_return NFSv4.1: zero referring call lists before decoding nfs: fix ENXIO on O_CREAT open of existing symlink over NFSv3 SUNRPC: wait for in-flight client TLS handshake callback NFSv4: Fix incorrect argument passed to nfs4_delete_lease() in nfs4_add_lease() lockd: fix NULL dereference on lockowner allocation failure NFS: fix delegation_hash_table leak when nfs4_server_common_setup() fails NFSv4/flexfiles: support loosely coupled data servers NFSv4/flexfiles: fix NULL dereference for NFSv4.0 data servers NFSv4: pin the superblock for active state owners sunrpc: fix use-after-free in __rpc_clnt_handle_event and __rpc_clnt_remove_pipedir NFS/localio: issue commit inline when not in a memory-reclaim context NFS/localio: remove dead FLUSH_SYNC handling from nfs_local_commit NFS/localio: issue IO inline when not in a memory-reclaim context NFS: Fix delayed delegation return list handling NFS: Verify symlink inode before caching target NFS: fix folio dereference before NULL check in nfs_inode_remove_request() ...
2026-08-25sunrpc: Remove pagemap.hMatthew Wilcox (Oracle)
No file in net/sunrpc needs pagemap.h, nor depends on it bringing in any of its dependencies. After this patch, no file in net/sunrpc depends on pagemap.h, even transitively. Signed-off-by: Matthew Wilcox (Oracle) <willy@infradead.org>
2026-08-21Merge tag 'for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/rdma/rdmaLinus Torvalds
Pull RDMA updates from Jason Gunthorpe: "About the normal size, still a lot of AI bug fixes and so on, but some interesting new functionality too: - Assorted locking, bounds-checking, cleanup, and error-path fixes across UCMA/CMA, bng_re, bnxt_re, cxgb4, EFA, ERDMA, HFI1, HNS, ionic, iRDMA, mlx4/mlx5, RXE, SIW, SRP/SRPT, and iSER target. - netlink report for max # of supported resources - get_zeroed_page()/etc removal - Robust udata for ionic - Allow unique RDMA device names per network namespace - Completion counters and v2 admit queue support for EFA - UC QP support for MANA - Completion timestamps for ionic - Harden uverbs data validation and resource lifetime handling, fixing several core use-after-free conditions. - bnxt_re toggle-page ownership and lifetime bug fixes - dmabuf SRQ support for mlx5" * tag 'for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/rdma/rdma: (160 commits) RDMA/ucma: Allow path records to exactly fit the output buffer RDMA/uverbs: Guard legacy bundles without method_elm RDMA/efa: Add support for 128B admin v2 SQ entry RDMA/efa: Generalize the admin SQ RDMA/efa: Decouple admin command payload from admin header RDMA/rxe: Fix OOB in free_rd_atomic_resources() RDMA/cma: Fix WARNING in res_to_rt RDMA/cxgb4: Free debugfs on registration failure RDMA/cxgb4: Cancel reg_work before freeing device on remove RDMA/ucma: Lock the handler in ucma_set_ib_path() RDMA/ucma: Lock the handler in ucma_write_cm_event() RDMA/erdma: restrict the driver to little-endian systems RDMA/ionic: Embed counter driver data in rdma_counter allocation RDMA/ionic: Cap eq_count to the eth driver's interrupt vector budget RDMA/siw: Fix use-after-free in siw_accept() IB/isert: post the full-feature receive buffers after session registration IB/isert: delay the final Login Response until the session is registered RDMA/srp: fix heap information leak on a truncated SRP_CRED_REQ RDMA/erdma: Hold QP references for AE and CM processing RDMA/erdma: Hold CQ references when processing EQ events ...
2026-08-17SUNRPC: wait for in-flight client TLS handshake callbackJérémy Jean
xs_tls_handshake_sync() gives xs_tls_handshake_done() a reference to the lower transport before submitting the handshake request. On timeout or signal, the synchronous waiter drops that reference after calling tls_handshake_cancel(). handshake_req_cancel() returns false when handshake_complete() has already marked the request complete. In that case the completion callback can still be running, so dropping the callback-owned reference in the waiter can free the lower transport before xs_tls_handshake_done() stores xprt_err or drops its own reference. If cancellation loses to completion, wait until xs_tls_handshake_done() signals handshake_done and let the callback release its reference. This mirrors the server-side handshake lifetime handling and keeps the timeout or signal return value unchanged. Fixes: 75eb6af7acdf ("SUNRPC: Add a TCP-with-TLS RPC transport class") Cc: stable@vger.kernel.org Assisted-by: Codex:gpt-5 Signed-off-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr> Reviewed-by: Chuck Lever <cel@kernel.org> Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>
2026-08-17sunrpc: fix use-after-free in __rpc_clnt_handle_event and ↵Luxiao Xu
__rpc_clnt_remove_pipedir Normal client creation goes through rpc_setup_pipedir(), which records clnt->pipefs_sb, but the mount-event path in __rpc_clnt_handle_event() calls rpc_setup_pipedir_sb() directly and never refreshes that field. The umount path also removes the directory without clearing clnt->pipefs_sb. After a late pipefs mount or any remount, rpc_clnt_remove_pipedir() compares the current superblock against a stale pipefs_sb pointer and skips cleanup, leaving pipefs dentries whose inode private data still points at a freed rpc_clnt, leading to a potential use-after-free during subsequent rpc_info_open() or rpc_show_info() calls. Fix this by properly updating clnt->pipefs_sb upon mount events and clearing it during unmount or failure paths. Fixes: bfca5fb4e97c ("SUNRPC: Fix RPC client cleaned up the freed pipefs dentries") Cc: stable@vger.kernel.org Reported-by: Yuan Tan <yuantan098@gmail.com> Reported-by: Xin Liu <dstsmallbird@foxmail.com> Reviewed-by: Ren Wei <enjou1224z@gmail.com> Assisted-by: Codex:gpt-5.4 Signed-off-by: Luxiao Xu <rakukuip@gmail.com> Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>
2026-08-10sunrpc: xprtsock: annotate shared socket callbacks with READ_ONCE/WRITE_ONCERunyu Xiao
xprtsock replaces and restores sk->sk_data_ready and sk->sk_write_space on live sockets with plain stores, and xs_udp_do_set_buffer_size() invokes sk->sk_write_space via a plain load. These callback pointers are shared with generic socket and protocol paths that may read or invoke them concurrently, so xprtsock needs the same READ_ONCE()/WRITE_ONCE() callback visibility contract that the validated 4022 family applied elsewhere. When SUNRPC takes over an AF_LOCAL, UDP, or TCP socket and later restores the lower-socket callbacks during teardown, another CPU may still hold an earlier callback snapshot. The plain replace/restore pattern leaves the same visibility hole as the validated 4022 family, so a stale snapshot can still invoke xs_data_ready() or xs_udp_write_space() after the live callback fields have already been restored to the lower-socket handlers. Use WRITE_ONCE() for the shared sk_data_ready and sk_write_space stores in xs_local_finish_connecting(), xs_udp_finish_connecting(), xs_tcp_finish_connecting(), and xs_restore_old_callbacks(). Use READ_ONCE() for the direct sk_write_space invocation in xs_udp_do_set_buffer_size(). This matches the required callback visibility contract while leaving adjacent sk_state_change and sk_error_report handling unchanged. Fixes: a246b0105bbd ("[PATCH] RPC: introduce client-side transport switch") Signed-off-by: Runyu Xiao <runyu.xiao@seu.edu.cn> Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>
2026-08-10SUNRPC: check rpc_sockaddr2uaddr() return value in rpcb_register_inet4/6Weiming Shi
rpcb_register_inet4() and rpcb_register_inet6() store the result of rpc_sockaddr2uaddr() into map->r_addr without checking it for NULL. rpc_sockaddr2uaddr() returns NULL when its final kstrdup() fails, and the unchecked NULL is then carried into the synchronous RPCBPROC_SET encode path: rpcb_register_call() -> rpc_call_sync() -> rpcb_enc_getaddr() -> encode_rpcb_string(), whose first statement is strlen(string), dereferencing NULL and oopsing the kernel. The crash reproduces under failslab on v6.12; with KASAN the NULL dereference surfaces as a fault on the shadow of address zero: Oops: general protection fault, probably for non-canonical address 0xdffffc0000000000 [#1] PREEMPT SMP KASAN RIP: 0010:strlen (lib/string.c:409) Call Trace: encode_rpcb_string (net/sunrpc/rpcb_clnt.c:890) rpcb_enc_getaddr (net/sunrpc/rpcb_clnt.c:910) rpcauth_wrap_req_encode (net/sunrpc/auth.c:745) call_encode (net/sunrpc/clnt.c:1966) __rpc_execute (net/sunrpc/sched.c:952) rpc_run_task (net/sunrpc/clnt.c:1243) rpc_call_sync (net/sunrpc/clnt.c:1272) rpcb_v4_register (net/sunrpc/rpcb_clnt.c:500) svc_generic_rpcbind_set nfsd_rpcbind_set svc_register svc_setup_socket svc_addsock write_ports nfsctl_transaction_write vfs_write The crash is reachable when an in-kernel RPC service (nfsd, lockd, nfs-callback) registers with the local rpcbind under enough memory pressure for the small GFP_KERNEL kstrdup() in rpc_sockaddr2uaddr() to fail. The asynchronous getport path already handles this exact failure mode by returning -ENOMEM; only the two register helpers omit the check. Mirror that handling: bail out with -ENOMEM when rpc_sockaddr2uaddr() returns NULL, before the address is fed into the encoder. Fixes: d77385f23830 ("SUNRPC: Fix rpc_sockaddr2uaddr") Reported-by: Xiang Mei <xmei5@asu.edu> Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Weiming Shi <bestswngs@gmail.com> Reviewed-by: Jeff Layton <jlayton@kernel.org> Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>
2026-08-10sunrpc: remove unused svc_version vs_count fieldJeff Layton
Now that svc_seq_show() and the nfsd netlink stats handler both use the per-netns svc_stat vs_count arrays, the global per-version vs_count percpu counters are no longer read by anything. Remove the vs_count field from struct svc_version and all the associated DEFINE_PER_CPU_ALIGNED arrays and initializers across nfsd, lockd, and the NFS client callback service. Assisted-by: LLM Signed-off-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260717-exportd-netlink-v7-4-b7ce17b83b60@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-08-10sunrpc: use per-net counts in svc_seq_show()Jeff Layton
Update svc_seq_show() to read from the per-netns statp->vs_count[] arrays instead of the global svc_version->vs_count[]. The only caller is nfsd, which always allocates vs_count via svc_stat_alloc_counts() in nfsd_net_init(), so the per-netns arrays are always available. This makes /proc/net/rpc/nfsd report per-network-namespace procedure call counts. Assisted-by: LLM Signed-off-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260717-exportd-netlink-v7-2-b7ce17b83b60@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-08-10sunrpc: add per-netns per-procedure call counts to svc_statJeff Layton
The existing per-procedure call counts live in global svc_version->vs_count[] arrays which are not network-namespace-aware. Add per-netns equivalents in struct svc_stat so the upcoming netlink stats interface can return namespace-scoped statistics. Add a vs_count pointer array to struct svc_stat, along with svc_stat_alloc_counts() and svc_stat_free_counts() helpers to manage per-version percpu call count arrays. Increment the per-net counter alongside the global one in svc_generic_init_request(). Call the alloc/free helpers from nfsd_net_init() and nfsd_net_exit(). Assisted-by: LLM Signed-off-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260717-exportd-netlink-v7-1-b7ce17b83b60@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-08-10sunrpc: derive the pool count instead of caching it in sv_nrpoolsJeff Layton
Now that the pool mode is always pernode, svc_serv.sv_nrpools is redundant with sv_is_pooled: an unpooled service always has a single pool, and a pooled service has svc_pool_map.npools pools (which is one on a single-node host). sv_nrpools cannot distinguish an unpooled service from a pooled service that happens to have one pool, so it is sv_nrpools, not sv_is_pooled, that carries no unique information. Replace the cached field with a svc_serv_nrpools() helper that derives the count from sv_is_pooled and the pool map, and convert all readers to it. svc_pool_map is file-local to svc.c, so export the helper for the svc_xprt.c and nfsd callers. Reading svc_pool_map.npools without svc_pool_map_mutex is safe: the mutex protects only svc_pool_map.count, and npools is already read locklessly in svc_pool_for_cpu(). A pooled service holds a map reference for its whole lifetime, so npools is stable while any reader could observe it. The hot path (svc_pool_for_cpu()) already dereferences svc_pool_map for to_pool, and npools shares that cacheline, so there is no new locking or coherence cost. __svc_create() keeps using its local npools argument for the sv_pools[] allocation, since sv_is_pooled is not set until svc_create_pooled() has returned from it. Doing this also removes a modulus operation from svc_pool_for_cpu(), which should make for more efficient RPC queueing. Assisted-by: Claude:claude-opus-4-8 Suggested-by: NeilBrown <neilb@ownmail.net> Signed-off-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260706-sunrpc-pool-mode-v5-5-6c4ee7cd89aa@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-08-10sunrpc: tear down pool counters before dropping the pool map referenceJeff Layton
svc_destroy() drops the service's reference to the global svc_pool_map before iterating serv->sv_pools[] to destroy each pool's percpu counters. That ordering happens to be fine today because the loop is bounded by the per-service sv_nrpools field. A following patch removes sv_nrpools and derives the pool count from the pool map instead. svc_pool_map_put() zeroes svc_pool_map.npools when the last reference is dropped, so a derived loop bound would read as zero for the last pooled service and skip svc_pool_destroy_counters() entirely, leaking the percpu counters (which remain linked on the global percpu_counters list while the svc_serv is freed). Reorder svc_destroy() to destroy the pool counters while the map is still referenced, then drop the reference. No functional change. Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260706-sunrpc-pool-mode-v5-4-6c4ee7cd89aa@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-08-10sunrpc: guarantee a thread per pool when auto-distributingJeff Layton
svc_set_num_threads() spreads the requested thread count evenly across the service's pools. In pernode mode each pool maps to a NUMA node, and svc_pool_for_cpu() steers an incoming transport to the pool for the node it arrived on. When fewer threads than pools are requested, even distribution leaves some pools empty, and a transport steered to an empty pool has no thread to service it. Floor each pool at one thread when auto-distributing a non-zero count, so no pool is left empty. Every pool maps to a node that had CPUs when the pool map was built (svc_pool_map_init_pernode() only creates pools for nodes returned by for_each_node_with_cpus()), so there is no pool that should be left threadless. The resulting total may exceed the requested count. This only affects the auto-distribute path (a single-value array, i.e. svc_set_num_threads()); callers that set per-pool counts explicitly via svc_set_pool_threads() are unchanged and may still set a pool to zero. Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Jeff Layton <jlayton@kernel.org> Reviewed-by: NeilBrown <neil@brown.name> Link: https://patch.msgid.link/20260706-sunrpc-pool-mode-v5-3-6c4ee7cd89aa@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-08-10sunrpc: hardcode pool_mode to pernode, remove other modesJeff Layton
The SVC_POOL_AUTO/GLOBAL/PERCPU/PERNODE pool mode selection machinery was added when NUMA was new and the right default was unclear. The default has always been "global" (a single pool for the whole service); the other modes were only used when an admin explicitly set the pool_mode parameter or asked for "auto", which then picked a mode from the host topology. Today, pernode is the right choice everywhere: - On multi-NUMA hosts, it gives one pool per node with proper thread affinity and NUMA-local memory allocation. - On single-node hosts, pernode degenerates to exactly one pool, identical to the old "global" mode -- svc_pool_for_cpu() short- circuits when sv_nrpools <= 1, no CPU affinity is set, and memory is allocated from the single node. The percpu mode (one pool per CPU) created excessive pools relative to the number of threads most deployments run, and was only auto-selected in a narrow case (single node, >2 CPUs). Note that this changes the default behaviour on multi-NUMA hosts: a service that previously ran with a single global pool now gets one pool per NUMA node by default. This in turn means a host running fewer threads than it has NUMA nodes can end up with pools that have no threads. svc_pool_for_cpu() already falls back to a populated pool in that case, so transports are still serviced. Remove the SVC_POOL_* enum, mode selection heuristic, svc_pool_map_init_percpu(), and all mode-based switch statements. Simplify pool map functions to always use the pernode path. If pool map allocation fails, svc_pool_map_get() now returns 0 and service creation fails, rather than silently falling back to a single global pool. With the mode check gone, svc_pool_map_get_node() would dereference the shared pool_to[] for every service that starts a thread. Only services created via svc_create_pooled() hold a map reference that keeps that array allocated, so gate the lookup in svc_new_thread() on sv_is_pooled: unpooled services (e.g. lockd, the NFS callback) use NUMA_NO_NODE and never consult the map. The kmalloc_node() callers in svc_prepare_thread() already accept NUMA_NO_NODE, but __folio_alloc_node() requires a valid node id, so resolve NUMA_NO_NODE to numa_mem_id() for the scratch folio allocation. The module parameter and netlink interfaces are preserved for backward compatibility: - Writing any of the four documented mode names still succeeds silently - Reading always returns "pernode" - Writing to the module parameter emits a deprecation notice Update Documentation/admin-guide/kernel-parameters.txt to mark the pool_mode parameter deprecated and describe the new behaviour. Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260706-sunrpc-pool-mode-v5-2-6c4ee7cd89aa@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-08-10sunrpc: route to a populated pool in svc_pool_for_cpu()Jeff Layton
svc_set_num_threads() spreads the requested threads evenly across the service's pools (base = nrservs / sv_nrpools). When a service runs fewer threads than it has pools -- e.g. an nfsd configured with fewer threads than the host has NUMA nodes while running in "pernode" or "percpu" mode -- the trailing pools are left with no threads at all. svc_xprt_enqueue() selects a pool from the CPU servicing the transport, queues the transport on that pool's sp_xprts, and only wakes a thread from the same pool. Each thread services exclusively its own pool, so a transport that lands on a threadless pool is enqueued on sp_xprts and never picked up: the connection hangs indefinitely. Have svc_pool_for_cpu() skip pools that currently have no threads, falling back to the next populated pool. This trades NUMA locality for a guarantee that the work is actually serviced. sp_nrthreads is only updated under the service mutex; the lockless read here is a best-effort routing hint, so annotate it with data_race(). Fixes: bfd241600a3b ("[PATCH] knfsd: make rpc threads pools numa aware") Cc: stable@vger.kernel.org Signed-off-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260706-sunrpc-pool-mode-v5-1-6c4ee7cd89aa@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-08-10SUNRPC: Restore NUMA_NO_NODE for svc thread allocations in global modeAmeer Hamza
Commit d57e43b72bf2 ("SUNRPC: Update svcxdr_init_decode() to call xdr_set_scratch_folio()") changed svc_pool_map_get_node() to return numa_mem_id() instead of NUMA_NO_NODE, because __folio_alloc_node() cannot accept NUMA_NO_NODE. That return value is not equivalent: it is evaluated in the context of the task creating the nfsd threads, once per thread created, and it is passed to kthread_create_on_node() and to the per-thread allocations in svc_prepare_thread(). Since commit d1a89197589c ("kthread: Default affine kthread to its preferred NUMA node"), the node argument of kthread_create_on_node() no longer only places the task structure and stack: a kthread created with a real node id normally affines itself to that node's CPUs when it is first woken to run its thread function. All nfsd threads are typically started together, by one task writing to /proc/fs/nfsd/threads, so under the default pool_mode=global each nfsd thread is now affined to the local-memory node of the CPU its creating iteration happened to run on - typically the same node for every thread. The CPUs of the other nodes are then unable to run nfsd at all, and the threads' allocations - svc_rqst structures, page pointer arrays, newly allocated task stacks, and the per-RPC pages allocated at run time - all prefer that one node. Restore the NUMA_NO_NODE behaviour that global mode has had since commit 11fd165c68b7 ("sunrpc: use better NUMA affinities"), and handle NUMA_NO_NODE at the one call site that cannot take it by resolving it to numa_mem_id() there, exactly as alloc_pages_node() did for the scratch page before the conversion. The mapped percpu and pernode branches are unchanged. Unpooled services such as lockd and the NFS client callback service also take this fallback when no percpu or pernode map is active, restoring their thread placement in that case. A bisect of a 2x NFS READ throughput regression between v6.17 and v6.18 converged on d57e43b72bf2. On the affected 4-node server every nfsd thread comes up with its CPU affinity restricted to the CPUs of a single node; with this change the threads are runnable on all CPUs again and the observed regression is resolved. Fixes: d57e43b72bf2 ("SUNRPC: Update svcxdr_init_decode() to call xdr_set_scratch_folio()") Cc: stable@vger.kernel.org Signed-off-by: Ameer Hamza <ameer.hamza@truenas.com> Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260722182012.2063936-1-ameer.hamza@truenas.com Signed-off-by: Chuck Lever <cel@kernel.org>
2026-08-10svcrdma: Reject inline replies that overflow the pull-up bufferChuck Lever
An RPC-over-RDMA client can request a reply, such as an NFS READ payload, without providing a Write list or a Reply chunk to carry it. When such a reply needs more scatter/gather entries than the device's Send Queue supports, svc_rdma_pull_up_needed() selects pull-up and svc_rdma_pull_up_reply_msg() linearizes the whole reply into sctxt->sc_xprt_buf. That buffer is only sc_max_req_size bytes, while the reply on this path is bounded only by the client's request, so svc_rdma_xb_linearize() copies past the end of the buffer and corrupts adjacent slab memory. The oversized length is then stored in sc_sges[0].length and posted, so the device also reads beyond the mapped region. The SGE-exhaustion branch is the only pull-up path that can exceed the buffer: the threshold branch pulls up only replies smaller than RPCRDMA_PULLUP_THRESH, and replies that fit the device's SGE budget are sent directly without linearization. Make svc_rdma_pull_up_needed() report -E2BIG when the reply it would pull up cannot fit sc_max_req_size, and fail the request with ERR_CHUNK as RFC 8166 Section 4.5.3 directs rather than dropping the connection. The helper no longer answers a simple yes/no question: it now reports pull-up, no pull-up, or -E2BIG for a reply too large to linearize. Rename svc_rdma_pull_up_needed() to svc_rdma_check_pull_up() so its name no longer implies a boolean predicate. Fixes: e248aa7be86e ("svcrdma: Remove max_sge check at connect time") Cc: stable@vger.kernel.org Reported-by: Chris Mason <clm@meta.com> Assisted-by: kres:claude-opus-4-7 Link: https://patch.msgid.link/20260623014728.826032-1-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-08-10sunrpc: defer rq_argp and rq_resp free until after RCU grace periodJeff Layton
svc_rqst_free() frees rqstp->rq_argp and rqstp->rq_resp synchronously via kfree(), but defers the rqstp struct free via kfree_rcu(). After svc_exit_thread() calls list_del_rcu() and svc_rqst_free(), there is a window where RCU readers that started before list_del_rcu() can still traverse the thread list and find the rqstp. These readers (e.g. nfsd_nl_rpc_status_get_dumpit()) dereference rqstp->rq_argp, which has already been freed — a use-after-free. Fix this by moving the kfree of rq_argp and rq_resp into an explicit call_rcu() callback alongside the struct free. Resources not accessed by RCU readers (bvec, buffer pages, scratch folio, auth_data) remain synchronously freed. Fixes: 812443865c5f ("sunrpc: add a rcu_head to svc_rqst and use kfree_rcu to free it") Cc: stable@vger.kernel.org Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260611-nfsd-testing-v2-4-5b90e276f2d9@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-08-10SUNRPC: Add svc_serv_maxthreads() to report the thread ceilingChuck Lever
A pooled RPC service sizes its threads dynamically, growing and shrinking each pool between its minimum and maximum bounds as load varies. The count of running threads therefore reflects recent demand, not the service's capacity. A consumer that sizes a data structure against the concurrency the service can sustain -- NFSD's NFSv4 session slot tables, for one -- needs that stable ceiling, and computing it means summing sp_nrthrmax across every pool. Add svc_serv_maxthreads() so the summation, and its dependence on the layout of struct svc_serv and struct svc_pool, stays within sunrpc. The read is lock-free: pool maxima change only when a service is reconfigured, a path callers already serialize against startup and shutdown, so a racing reader observes at worst a transient value. This is acceptable for the sizing heuristics that will consume it. nfsd_nrthreads() already sums sp_nrthrmax across pools by hand; convert it to svc_serv_maxthreads(), giving the new export an in-tree consumer and removing a copy of the dependence on svc_serv internals. Reviewed-by: NeilBrown <neil@brown.name> Reviewed-by: Jeff Layton <jlayton@kernel.org> Reviewed-by: Benjamin Coddington <bcodding@hammerspace.com> Link: https://patch.msgid.link/20260610-nfsd-slot-growth-clamp-v1-1-7b966700df0b@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-08-10net/sunrpc/svcauth_unix: Use strscpy() to copy strings into arraysDavid Laight
Replacing strcpy() with strscpy() ensures that overflow of the target buffer cannot happen. Signed-off-by: David Laight <david.laight.linux@gmail.com> Link: https://patch.msgid.link/20260608095523.2606-16-david.laight.linux@gmail.com Signed-off-by: Chuck Lever <chuck.lever@oracle.com>
2026-08-10rpcrdma: arm rn_done before publishing the notificationChuck Lever
rpcrdma_rn_register() inserts @rn into rd_xa with xa_alloc() before storing the caller's callback in rn->rn_done. The xarray makes @rn reachable to rpcrdma_remove_one(), which walks rd_xa and invokes rn->rn_done(rn) for every registered notification. A device removal that races a fresh registration can therefore observe @rn with rn_done still NULL, because the notification objects are zero allocated by their owners, and call through a NULL function pointer. Store rn->rn_done before xa_alloc() publishes @rn. The xarray's store-side and load-side ordering then guarantees that any CPU which finds @rn in rd_xa also observes the armed callback. rpcrdma_rn_unregister() treats a non-NULL rn_done as the sentinel for a completed registration, so the early store must not survive a failed registration. Clear rn_done again when xa_alloc() fails. Were it left set, the failed-accept cleanup path would call rpcrdma_rn_unregister() on an @rn that was never inserted, erasing an unrelated rd_xa slot and underflowing rd_kref. Fixes: 7e86845a0346 ("rpcrdma: Implement generic device removal") Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260601201703.46078-1-cel@kernel.org Signed-off-by: Chuck Lever <chuck.lever@oracle.com>
2026-08-10SUNRPC: Check svc pool percpu counter allocationChuck Lever
__svc_create() initializes three per-pool percpu_counter stats and ignores every return value. On SMP, percpu_counter_init() fails when __alloc_percpu_gfp() cannot satisfy the allocation, leaving the failed counter with fbc->counters == NULL and its embedded raw_spinlock_t, list_head, and count never initialized. __svc_create() returns the half-constructed svc_serv to nfsd, lockd, or the NFS callback service anyway. Once that service is live, the hot-path increments in svc_xprt_enqueue(), svc_handle_xprt(), and svc_pool_wake_idle_thread() reach a counter whose backing pointer is NULL. The pointer is a per-cpu offset, so the access does not fault: it resolves to offset zero of the current CPU's per-cpu area and silently corrupts whatever variable lives there. A /proc/fs/nfsd/pool_stats read walks the same NULL per-cpu storage and returns garbage, and on CONFIG_DEBUG_SPINLOCK or lockdep it splats on the never-initialized lock. Creating the broken service requires a percpu allocation failure during RPC server startup, so it is reachable only by a local administrator under memory pressure or fault injection; a remote peer cannot induce the bad state on its own. Check each percpu_counter_init() return value in __svc_create() and fail when an allocation fails, unwinding the counters already set up in the current pool and in every pool initialized before it. A discrete percpu_counter_destroy() per counter at teardown frees each per-cpu allocation exactly once. Fixes: ccf08bed6e7a ("SUNRPC: Replace pool stats with per-CPU variables") Cc: stable@vger.kernel.org Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260530-tier2-local-v2-2-5a0fd532db57@oracle.com Signed-off-by: Chuck Lever <chuck.lever@oracle.com>
2026-08-10sunrpc: init gssp_lock before publishing proc entryChris Mason
create_use_gss_proxy_proc_entry() publishes /proc/net/rpc/use-gss-proxy via proc_create_data() before init_gssp_clnt() runs mutex_init() on sn->gssp_lock. Once the dentry is linked under proc_subdir_lock it is immediately reachable from userspace, so a write that lands in the window drives set_gssp_clnt() into mutex_lock() on a zero-initialized struct mutex. create_use_gss_proxy_proc_entry(net) proc_create_data("use-gss-proxy", ...) /* dentry live */ init_gssp_clnt(sn) mutex_init(&sn->gssp_lock) /* too late */ write_gssp() set_gssp_clnt(net) mutex_lock(&sn->gssp_lock) /* uninitialized */ gssp_rpc_create(...) sn->gssp_clnt = clnt mutex_unlock(&sn->gssp_lock) The window spans only the two statements between proc_create_data() returning and init_gssp_clnt(), so a writer reaches it only if the registering thread is preempted there while another task is already opening the freshly published file. register_pernet_subsys() runs in preemptible context under pernet_ops_rwsem, so that preemption is possible, and the window widens on auth_rpcgss module load, when the proc entry is created for every live net namespace whose tasks are already running. A writer that wins the race locks a zero-filled struct mutex. On CONFIG_DEBUG_MUTEXES the missing magic value trips a "lock used without init" splat; on a production kernel the fast path acquires the lock via CMPXCHG(owner, 0, current). In the latter case a second writer that arrives before init_gssp_clnt() re-zeroes owner can enter set_gssp_clnt() concurrently, shut down the first writer's clnt while it is still in use, and leak the loser's clnt. Fix by initializing sn->gssp_lock in sunrpc_init_net() so its lifetime matches the sunrpc_net it lives in. sn->gssp_clnt is already NULL from the kzalloc that backs net_generic storage, so the lazy helper is no longer needed; drop init_gssp_clnt(), its prototype, and the call from create_use_gss_proxy_proc_entry(). sunrpc.ko is a build-time dependency of auth_rpcgss.ko, so sunrpc_init_net() has always run on every netns before any auth_gss pernet init can publish the proc entry. Fixes: 030d794bf498 ("SUNRPC: Use gssproxy upcall for server RPCGSS authentication.") Cc: stable@vger.kernel.org Assisted-by: kres:claude-opus-4-7 Signed-off-by: Chris Mason <clm@meta.com> Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260530-tier2-local-v2-1-5a0fd532db57@oracle.com Signed-off-by: Chuck Lever <chuck.lever@oracle.com>
2026-08-10SUNRPC: close backchannel before destroying callback serviceChuck Lever
A backchannel receive can complete a request while the NFS callback service is being torn down. xprt_complete_bc_request() removes the request from bc_pa_list, drops bc_alloc_count, marks the request in use, and then asks xprt_enqueue_bc_request() to hand it to the callback service. If teardown has already cleared xprt->bc_serv, xprt_enqueue_bc_request() currently returns without enqueueing or freeing the committed request. The xprt_get() taken on entry is leaked as well. If the producer wins the race before bc_serv is cleared, it can also enqueue onto sv_cb_list after nfs_callback_down() has stopped the callback threads, leaving the request linked to a svc_serv that is about to be freed. Close the producer side before callback threads are stopped. Add xprt_svc_shutdown_bc() to clear xprt->bc_serv under bc_pa_lock, and call it on callback shutdown and callback-start failure before stopping the service threads. Requests that lose the NULL transition in xprt_enqueue_bc_request() are released through the normal backchannel free path after balancing bc_slot_count. Finally, drain any remaining sv_cb_list requests after the callback threads have stopped and before svc_destroy() frees the service. Fixes: 441244d4273a ("SUNRPC: cleanup common code in backchannel request") Fixes: 9e9fdd0ad0fb ("NFSv4.1: protect destroying and nullifying bc_serv structure") Cc: stable@vger.kernel.org Signed-off-by: Chris Mason <clm@meta.com> Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260528-tier2-v1-6-d026a1415e0b@oracle.com Signed-off-by: Chuck Lever <chuck.lever@oracle.com>
2026-08-10SUNRPC: Zero rpc_gss_wire_cred at svcauth_gss_decode_credbody() entryChris Mason
svcauth_gss_decode_credbody() writes the caller's rpc_gss_wire_cred field by field and assigns gc_ctx.len only on the success tail. The caller storage is svcdata->clcred, which lives in the per-svc_rqst gss_svc_data and is reused across requests. Early decode failures leave partially decoded state mixed with residue from the prior request. The trailing body_len tightness check is the sharpest case: xdr_stream_decode_opaque_inline() has already written gc_ctx.data with a borrowed inline pointer into the current request's XDR pages, but gc_ctx.len retains its prior value. Once the request pages are released the pooled clcred carries a dangling pointer paired with a stale length. Zero the caller's rpc_gss_wire_cred at function entry so that every early-return path leaves a deterministic all-zero cred. On the trailing tightness-check path, gc_ctx.len is now zero instead of stale, which neuters length-driven consumers such as gss_svc_searchbyctx() that would otherwise walk the dangling data pointer. Fixes: b0bc53470d1a ("SUNRPC: Convert the svcauth_gss_accept() pre-amble to use xdr_stream") Cc: stable@vger.kernel.org Signed-off-by: Chris Mason <clm@meta.com> Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260528-tier2-v1-5-d026a1415e0b@oracle.com Signed-off-by: Chuck Lever <chuck.lever@oracle.com>
2026-08-10SUNRPC: Guard svcauth_gss_release() dispatch on rq_auth_statChris Mason
svcauth_gss_release() reads gc_proc and switches on gc_svc before consulting rq_auth_stat. On the SVC_DENIED path after a failed svcauth_gss_accept(), those fields may hold stale values from a prior request or uninitialized slab residue: svcauth_gss_accept() allocates gss_svc_data with non-zeroing kmalloc and clears only gsd_databody_offset and rsci per request, not clcred. Because RPC_GSS_PROC_DATA is zero, a zeroed or stale-zero gc_proc passes the existing guard and falls through into the gc_svc switch, which can dispatch to svcauth_gss_wrap_integ() or svcauth_gss_wrap_priv(). Both wrap helpers call svcauth_gss_prepare_to_wrap() before any rsci->mechctx dereference, and that helper already returns early when rq_auth_stat is not rpc_auth_ok, so the downstream NULL dereference is blocked. The dispatch itself remains structurally wrong: it reads scalars that the caller has no contract to have initialized after a failed authentication. Mirror the existing rq_auth_stat gate in svcauth_gss_prepare_to_wrap() one frame up, so svcauth_gss_release() skips the clcred dispatch entirely when authentication has not succeeded. The cleanup tail that releases rq_client, rq_gssclient, cr_group_info, and rsci still runs. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Cc: stable@vger.kernel.org Signed-off-by: Chris Mason <clm@meta.com> Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260528-tier2-v1-4-d026a1415e0b@oracle.com Signed-off-by: Chuck Lever <chuck.lever@oracle.com>
2026-08-10SUNRPC: reject duplicate CREDS_VALUE optionsChris Mason
gssx_dec_option_array() walks the wire-supplied option array and, for every entry whose name matches CREDS_VALUE, calls gssx_dec_linux_creds() on the same struct svc_cred. That helper unconditionally installs a fresh groups_alloc() result into creds->cr_group_info without releasing whatever pointer was already there: for (i = 0; i < count; i++) { ... decode name ... if (length == sizeof(CREDS_VALUE) && memcmp(p, CREDS_VALUE, sizeof(CREDS_VALUE)) == 0) { err = gssx_dec_linux_creds(xdr, creds); ... } } A reply that carries two CREDS_VALUE entries therefore overwrites cr_group_info on the second iteration and orphans the group_info allocated by the first call. The earlier free_creds path only releases the last cr_group_info via free_svc_cred(), so the first allocation's refcount stays at one and its kvmalloc-backed storage is leaked. No in-tree caller of gssp_accept_sec_context_upcall() expects more than one CREDS_VALUE per reply. Fix by tracking whether a CREDS_VALUE option has already been decoded and returning -EINVAL on any subsequent match, so the free_creds path releases the single group_info that was installed. Fixes: 1d658336b05f ("SUNRPC: Add RPC based upcall mechanism for RPCGSS auth") Cc: stable@vger.kernel.org Assisted-by: kres (claude-opus-4-7) Signed-off-by: Chris Mason <clm@meta.com> Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260528-tier2-v1-3-d026a1415e0b@oracle.com Signed-off-by: Chuck Lever <chuck.lever@oracle.com>
2026-08-10SUNRPC: fix gssx_dec_option_array error path bugsChris Mason
Four coupled defects in the gssx XDR option-array decoder make the error paths unsafe: a NULL deref in the caller, a refcount leak on the decoded group_info, and a latent use-after-free that the leak fix would otherwise expose. gssx_dec_option_array() sets oa->count = 1 before allocating oa->data. If that allocation fails, -ENOMEM is returned with oa->count == 1 and oa->data == NULL. All other error paths jump to free_oa: which frees oa->data and NULLs it but also leaves oa->count == 1. The caller trusts the count: gssp_accept_sec_context_upcall() gssx_dec_accept_sec_context() gssx_dec_option_array() /* fails, count=1 data=NULL */ data = res.options.data[0].value /* NULL deref */ Independently, free_creds: releases the partially decoded svc_cred with a bare kfree(creds). gssx_dec_linux_creds() installs a groups_alloc() result into creds->cr_group_info; that object is kvmalloc-backed and refcounted, and only put_group_info() reaches kvfree(). A plain kfree(creds) drops the wrapper and leaks the group_info allocation. The natural fix for the leak is to call free_svc_cred(creds) before kfree(creds), but free_svc_cred() invokes put_group_info() on creds->cr_group_info unconditionally when non-NULL. The existing out_free_groups: path in gssx_dec_linux_creds() already called groups_free() on that pointer without clearing it, so once free_svc_cred() is wired in, the subsequent put_group_info() would touch freed memory. Fix all four together: - Move the oa->count = 1 assignment below the oa->data allocation so it is never set when oa->data is NULL. - Reset oa->count to 0 at free_oa: so count and data stay coherent and the caller sees an empty option array. - Call free_svc_cred(creds) before kfree(creds) at free_creds: so the refcounted cr_group_info is released. free_svc_cred() either NULL-guards each field explicitly (cr_group_info has an if() check) or delegates to a helper that is NULL-safe itself (kfree for the string fields, gss_mech_put() which guards with if(gm) at gss_mech_switch.c:342), so it is safe to call on a partially decoded svc_cred where only cr_uid/cr_gid/cr_group_info have been written and everything else is zero from kzalloc. - In gssx_dec_linux_creds()'s out_free_groups: path, release cr_group_info with put_group_info() rather than groups_free() so the teardown matches free_svc_cred()'s refcount-aware path, and clear the pointer so a later free_svc_cred() on the same creds does not release it a second time. Fixes: 3cfcfc102a5e ("SUNRPC: fix some memleaks in gssx_dec_option_array") Cc: stable@vger.kernel.org Assisted-by: kres (claude-opus-4-7) Signed-off-by: Chris Mason <clm@meta.com> Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260528-tier2-v1-2-d026a1415e0b@oracle.com Signed-off-by: Chuck Lever <chuck.lever@oracle.com>
2026-08-10SUNRPC: Reject krb5 v2 wrap tokens with oversized ec fieldChuck Lever
gss_krb5_unwrap_v2() sets buf->len to a logical length, which can be much smaller than head[0].iov_len (the allocated receive-page capacity). It then calls xdr_buf_trim() with a trim length derived from the 16-bit "extra count" (ec) field in the Kerberos v2 token header. The ec field is authenticated by the post-decrypt memcmp() against the encrypted header copy, so a randomly-mutated value is rejected. However, any peer holding a valid GSS context can legitimately encrypt a token whose ec exceeds the plaintext length. Per RFC 4121, such a token is structurally malformed. Although xdr_buf_trim() now clamps the buf->len subtraction to avoid unsigned underflow, the buffer is still left in a semantically invalid state (zero length, inconsistent iov lengths) when ec is oversized. Reject these tokens before calling xdr_buf_trim(), giving callers a well-defined GSS_S_DEFECTIVE_TOKEN error and keeping the xdr_buf internally consistent. The wrapped blob begins at a nonzero offset -- both callers pass len as offset + opaque_len -- so buf->len still counts the offset bytes that precede the blob. Compare the trim length against the remaining wrapped segment, buf->len - offset, rather than the whole buffer; comparing against buf->len alone leaves an offset-wide window in which an oversized ec passes the test and xdr_buf_trim() cuts into the bytes ahead of the blob. Fixes: cf4c024b9083 ("sunrpc: trim off EC bytes in GSSAPI v2 unwrap") Cc: stable@vger.kernel.org Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260528-tier2-v1-1-d026a1415e0b@oracle.com Signed-off-by: Chuck Lever <chuck.lever@oracle.com>