summaryrefslogtreecommitdiff
AgeCommit message (Collapse)Author
2026-09-13sunrpc: treat empty auth.unix.gid replies as negative entriesAmeer Hamza
When rpc.mountd cannot resolve a uid (getpwuid() or getgrouplist() failure, e.g. while winbind or sssd is briefly unreachable), it answers the auth.unix.gid upcall with zero groups. unix_gid_parse() installs that as a valid positive entry, and svcauth_unix_set_client() then replaces the credential's group list with the empty one on every request, RPCSEC_GSS included via svcauth_gss_set_client(). One failed lookup strips that uid of all supplementary groups on every export for up to mountd's configured TTL (30 minutes by default), long after the NSS backend has recovered. mountd cannot send an empty list for a successful lookup, since getgrouplist(3) always includes at least the user's primary group, so a zero-group reply can only mean the lookup failed. Record it as a negative entry: unix_gid_find() then returns -ENOENT and svcauth_unix_set_client() keeps the groups the RPC credential already carries. This is the fallback that commit 3fc605a2aa38 ("[PATCH] knfsd: allow the server to provide a gid list when using AUTH_UNIX authentication") promised when no answer is available, and the same state try_to_negate_entry() already creates when no listener holds the channel open. Fixes: 3fc605a2aa38 ("[PATCH] knfsd: allow the server to provide a gid list when using AUTH_UNIX authentication") Cc: stable@vger.kernel.org Assisted-by: Claude:claude-fable-5 Signed-off-by: Ameer Hamza <ameer.hamza@truenas.com> Link: https://patch.msgid.link/20260814221953.108837-2-ameer.hamza@truenas.com Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13nfsd: back CB_NOTIFY notify_mask words with per-delegation storageJeff Layton
nfsd4_cb_notify_prepare() reserved a word from the encoding xdr stream for each notify4's host-order notify_mask, storing the pointer in ncn_nf[].notify_mask.element. That element must survive until the RPC encode re-reads ncn_nf, so the host-endian word lived inside the XDR staging buffer for the whole callback lifetime - the same fragile pattern as the attrmask, and one that keeps host-order bytes in a buffer meant to hold big-endian XDR. ncn_nf is a bounded per-delegation array reused across every CB_NOTIFY, so give it a parallel ncn_masks array with the same lifetime: - allocate/free ncn_masks alongside ncn_nf in alloc_init_dir_deleg() / nfs4_free_dir_deleg() - point notify_mask.element at &ncn_masks[i] (events) and &ncn_masks[count] (dir attr change) The mask backing is now pre-allocated, so the per-word NULL checks in prepare go away. The staging stream holds only encoded XDR. Assisted-by: LLM Signed-off-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260812-dir-deleg-v1-2-411faa713068@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13nfsd: pass caller-provided attrmask storage into nfsd4_setup_notify_entry4()Jeff Layton
nfsd4_setup_notify_entry4() stole 3 words from the xdr stream via xdr_reserve_space() to hold the host-order bmval[3] attrmask that nfsd4_encode_attr_vals() consumes and ne_attrs.attrmask.element points at. Stashing host-endian scratch in an XDR stream buffer is fragile: the buffer layout is not guaranteed by sunrpc, and it blocks moving the encoder to pages or xdrgen. The attrmask only needs to live until the enclosing encode call serializes the notify_entry4, so hand it caller-provided stack storage instead. The two callers keep the words on the stack: up to three concurrent entries for a rename in nfsd4_encode_notify_event(), one in nfsd4_encode_dir_attr_change(). No wire change: the reserved words were never emitted; attr_vals.data/len are still captured relative to xdr->p. Assisted-by: LLM Signed-off-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260812-dir-deleg-v1-1-411faa713068@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Remove xdr-related headers from fs/nfsd/vfs.cChuck Lever
Clean up: Nothing in fs/nfsd/vfs.c references an NFSv3 XDR definition. The preceding patch moved the last NFSv4 reference, a struct nfsd4_compoundres dereference that recovered a pair of file handles for a tracepoint, into fs/nfsd/nfs4proc.c. Remove both includes so that the VFS layer no longer names on-the-wire types. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260812142436.35042-4-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Move NFSv4-specific CLONE logic into nfsd4_clone()Chuck Lever
nfsd4_clone_file_range() lives in fs/nfsd/vfs.c but reaches into the NFSv4 compound reply buffer: nfsd4_get_cstate() casts rq_resp to a struct nfsd4_compoundres to recover the saved and current file handles a tracepoint wants. That is the only reference to the NFSv4 XDR definitions left in vfs.c, and it puts knowledge of the compound reply layout in the VFS layer. Refactor nfsd4_clone_file_range() to remove NFSv4-specific componentry from fs/nfsd/vfs.c. Splitting nfsd_clone_file_range() and nfsd_clone_sync_range() lets nfsd4_clone() distinguish a clone failure from a sync failure, which it has to do because only the latter invalidates the write verifier. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260812142436.35042-3-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Make the write verifier reset helper available outside vfs.cChuck Lever
A subsequent patch moves the NFSv4-specific portions of nfsd4_clone_file_range() out of fs/nfsd/vfs.c and into its caller in fs/nfsd/nfs4proc.c. One of those portions resets the write verifier when the post-clone sync fails. netns.h already exposes nfsd_reset_write_verifier(), but that is the unconditional reset. commit_reset_write_verifier() wraps it with the policy that decides which errors warrant a reset: -EAGAIN and -ESTALE do not indicate a problem with durable storage, so they leave the verifier alone. A caller outside vfs.c has to apply the same policy, so make the wrapper visible rather than duplicate its switch. Rename it to nfsd_maybe_reset_write_verifier() on the way out. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260812142436.35042-2-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Move version-specific ACCESS maps into per-version codeChuck Lever
nfsd_access() owns three static tables that map on-the-wire ACCESS bits to NFSD_MAY flags, and every NFS version shares them. That puts protocol-version specifics in the version-agnostic VFS layer. The NFSv4.2 extended-attribute bits are wedged into the NFSv3 tables under CONFIG_NFSD_V4. Give each version its own tables in its proc code and pass the matching set into nfsd_access() as a new argument. Splitting the tables also stops the NFSv2-ACL and NFSv3 ACCESS paths from answering for the NFSv4.2 extended-attribute bits. Both use the NFSv4-augmented table whenever CONFIG_NFSD_V4 is set, so an NFSv3 request that sets an xattr bit has it echoed back in the reply even though those bits are undefined for v3. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260810141900.33846-2-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Replace NFS3_ACCESS_FULL in nfsd4_access()Chuck Lever
Clean up: Remove an NFSv3 constant (NFS3_ACCESS_FULL) used inside an NFSv4 code path. After this patch is applied, the nfs3.h header is no longer an implicit dependency of nfs4proc.c for this value. The definition itself stays in nfs3.h to avoid a kernel-user space API regression. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260810141900.33846-1-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13SUNRPC: Fold xs_sock_process_cmsg() into its only callerChuck Lever
xs_sock_process_cmsg() switches on the TLS record type, and every arm but TLS_RECORD_TYPE_ALERT returns the -EAGAIN its caller passed in. The DATA arm clears MSG_EOR in the caller's msghdr, but xs_sock_recvmsg() has already cleared that flag before the call. Deriving the record type a second time inside the helper also fires trace_tls_contenttype() twice for every alert. Move the alert handling into xs_sock_recv_cmsg() and delete the helper. Every other record type still returns -EAGAIN. The DATA arm's account of MSG_EOR moves to xs_sock_recvmsg(), where the flag is cleared. Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-7-62d9a631c880@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13SUNRPC: Treat every client-side TLS error alert as fatalChuck Lever
xs_sock_process_cmsg() decides whether an alert ends the session by reading the alert's level octet. RFC 8446 Section 6 retired that field. The severity is implicit in the description, and a receiver treats every alert listed in Section 6.2 as an error alert "regardless of the AlertLevel in the message". A peer that aborts with unexpected_message but leaves the legacy octet set to warning makes the client return -EAGAIN. xs_stream_data_receive() wakes no pending task for that error, so RPC Calls queued on a dead TLS session wait for their timeouts to expire. Decide from the alert description instead. close_notify and user_canceled are the closure alerts (RFC 8446 Section 6.1). Every other description ends the session, including one this kernel does not recognize. Fixes: 39067dda1d86 ("SUNRPC: Use new helpers to handle TLS Alerts") Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-6-62d9a631c880@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13SUNRPC: Reject a client-side TLS alert record that is not two octetsChuck Lever
tls_alert_recv() reads two octets from the kvec it is handed and does not check the length (net/handshake/alert.c). xs_sock_process_cmsg() calls it for any alert record, and the alert[] buffer that xs_sock_recv_cmsg() supplies carries no initializer. A one-octet alert body leaves the description read from uninitialized stack and reported through trace_tls_alert_recv(). The peer controls that length. Neither tls_rx_msg_size() nor tls_rx_one_record() enforces the two-octet Alert payload. A TLS 1.3 record carrying only the inner content-type octet decrypts to a zero-length payload. RFC 8446 Section 5.1 requires a record with an Alert type to carry exactly one message, so any other length is malformed. RFC 9289 Section 5 bars RPC-with-TLS from negotiating a version below TLS 1.3, so no other alert framing applies. Require exactly two octets before parsing and return -EACCES otherwise. xs_stream_data_receive() already treats -EACCES as a fatal alert and reports it to the pending tasks. Gate the path on a control message rather than a positive count so that a zero-length record reaches the check. Fixes: cc5d59081fa2 ("sunrpc: fix client side handling of tls alerts") Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-5-62d9a631c880@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13SUNRPC: Resume receiving after a TLS control recordChuck Lever
A TLS control record delivers no payload to the RPC layer. svc_tcp_recvfrom() clears XPT_DATA before the receive, and svc_tcp_sock_recv_cmsg() returns -EAGAIN for the record it consumed. Nothing marks the transport ready again. kTLS raises data_ready for arriving TCP segments, not for records it has already decrypted. An RPC Call queued behind an alert or a KeyUpdate waits until the client sends more. The client blocks until its RPC timeout expires. The receive takes only the first two octets of the record. kTLS holds the remainder on its receive list, where each later receive takes two octets more. Drain a record that is not an alert, then mark the transport ready once a control record has been consumed. Fixes: 5e052dda121e ("SUNRPC: Recognize control messages in server-side TCP socket code") Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-4-62d9a631c880@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13SUNRPC: Treat every TLS error alert as fatalChuck Lever
svc_tcp_sock_process_cmsg() decides whether an alert ends the session by reading the alert's level octet. RFC 8446 Section 6 retired that field. The severity is implicit in the description, and a receiver treats every alert listed in Section 6.2 as an error alert "regardless of the AlertLevel in the message". A peer that aborts with unexpected_message but leaves the legacy octet set to warning makes the server return -EAGAIN. svc_tcp_recvfrom() then leaves a dead TLS session attached to an open transport. NFSD keeps polling it. Decide from the alert description instead. close_notify and user_canceled are the closure alerts (RFC 8446 Section 6.1). Every other description ends the session, including one this kernel does not recognize. Fixes: 39067dda1d86 ("SUNRPC: Use new helpers to handle TLS Alerts") Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-3-62d9a631c880@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13SUNRPC: Reject a TLS alert record that is not two octetsChuck Lever
tls_alert_recv() reads two octets from the kvec it is handed and does not check the length (net/handshake/alert.c). svc_tcp_sock_recv_cmsg() calls it for any positive receive, and the alert[] buffer it supplies carries no initializer. A one-octet alert body leaves the description read from uninitialized stack and reported through trace_tls_alert_recv(). The peer controls that length. Neither tls_rx_msg_size() nor tls_rx_one_record() enforces the two-octet Alert payload. A TLS 1.3 record carrying only the inner content-type octet decrypts to a zero-length payload. RFC 8446 Section 5.1 requires a record with an Alert type to carry exactly one message, so any other length is malformed. Require exactly two octets before parsing and return -EBADMSG otherwise. That closes the transport rather than acting on a partly uninitialized alert. Gate the path on a control message rather than a positive count so that a zero-length record reaches the check. Fixes: bee47cb026e7 ("sunrpc: fix handling of server side tls alerts") Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-2-62d9a631c880@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13SUNRPC: Do not credit control-record octets to the RPC streamChuck Lever
svc_tcp_sock_recv_cmsg() receives up to two octets into a local buffer, and returns that count for any record type other than TLS_RECORD_TYPE_ALERT. Nothing reached the caller's buffer, but svc_tcp_read_marker() adds the count to sk_tcplen and svc_tcp_read_msg()'s caller adds it to sk_datalen. The RPC stream advances over octets it never received. The fragment marker is assembled from stale sk_marker octets. The message body comes from pages nothing wrote. A conforming client reaches this. RFC 8446 Section 4.6.3 lets either peer send KeyUpdate once it has sent its Finished, and svcsock has no rekey path. kTLS leaves the partially consumed record on ctx->rx_list, so the body drains two octets per svc_tcp_recvfrom() call. Each pair is credited the same way. Return -EAGAIN for a record that is not an alert. That is what svc_tcp_sock_process_cmsg()'s default arm returned before the receive moved into a local buffer. Fixes: bee47cb026e7 ("sunrpc: fix handling of server side tls alerts") Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-1-62d9a631c880@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Point contributors and sashiko.dev to the nfsd-testing branchChuck Lever
Scripting and automation is sensitive to branch names in the subsystem entries in MAINTAINERS. Rather than pulling from cel.git/master, we really want CI to pull from cel.git/nfsd-testing. Link: https://patch.msgid.link/20260804184630.1395002-1-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Send referring calls with CB_RECALLChuck Lever
When CB_RECALL races ahead of the reply that granted the delegation, the client has not yet recorded the delegation stateid and responds NFS4ERR_BADHANDLE or NFS4ERR_BAD_STATEID. The slot that carried the grant has not retired at that point, so NFSD cannot read the rejection as proof that the client never held the delegation. It retries the recall and, once the retries lapse, revokes a delegation the client is by then able to return. Remove the ambiguity with the referring call mechanism of RFC 8881 Section 2.10.6.3: until the slot that carried the grant retires, name that request as a referring call in the CB_SEQUENCE of each recall. A client that finds it still outstanding may respond NFS4ERR_DELAY, and the recall is retried until the client has processed the grant. A recall reuses one callback context across its retries, and ->prepare does not run on every send. The granting request does not change, so a send that inherits the previous list sends the right one. Retirement of the granting slot drops the list, and nfs4_free_deleg() releases what is left. Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-9-323aa7196055@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Destroy a recalled delegation the client does not holdChuck Lever
A client that answers CB_RECALL with NFS4ERR_BADHANDLE or NFS4ERR_BAD_STATEID has no record of the delegation, so the FREE_STATEID that clears it from cl_revoked never arrives. Every later SEQUENCE reply carries SEQ4_STATUS_RECALLABLE_STATE_REVOKED, and the client loops issuing TEST_STATEID. Destroy such a delegation when it is reaped rather than revoking it onto cl_revoked. RFC 8881 Section 20.2.4 completes the recall at the reply when its status is neither NFS4_OK nor NFS4ERR_DELAY, so a rejected recall leaves nothing to revoke. An administrative revoke keeps that path, since NFS4ERR_ADMIN_REVOKED reports it. A destroyed stateid returns NFS4ERR_BAD_STATEID instead of NFS4ERR_DELEG_REVOKED. A client that rejects the recall but still holds the delegation gets no notice that its state was revoked. CB_RECALL can outrun the reply that granted the delegation, so honor a rejection only once the client has seen that grant. Per RFC 8881 Section 2.10.6.3, retirement of the slot that carried the grant is that proof; retry until then, and revoke when the retries lapse. Fixes: 3bd64a5ba171 ("nfsd4: implement SEQ4_STATUS_RECALLABLE_STATE_REVOKED") Cc: stable@vger.kernel.org # 6.14.x Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-8-323aa7196055@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Correct locking documentation for delegation sc_statusChuck Lever
The comment above the SC_STATUS_ flags states that nn->deleg_lock protects sc_status for delegation stateids, but only the transitions made while a delegation is hashed are taken under that lock. This comment was accurate until commit c88c150a467f ("nfsd: fix possible badness in FREE_STATEID") set SC_STATUS_CLOSED under ->cl_lock. Commit 8dd91e8d31fe ("nfsd: fix race between laundromat and free_stateid") added the other two sites. Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-7-323aa7196055@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Budget the CB_RECALL_ANY opcodeChuck Lever
NFS4_enc_cb_recall_any_sz counts the objects-to-keep field, the bitmap array length, and the bitmap word. encode_cb_recallany4args() emits an opcode ahead of all three, so the macro falls one XDR word short. This macro sizes p_arglen and nothing else. rq_callsize pads that with two credential slacks, so the shortfall has never reached the send buffer. No backport is needed. Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-6-323aa7196055@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Budget the CB_NOTIFY_LOCK opcodeChuck Lever
NFS4_enc_cb_notify_lock_sz counts the lock owner and the file handle. nfs4_xdr_enc_cb_notify_lock() emits an opcode ahead of both, so the macro falls one XDR word short. This macro sizes p_arglen and nothing else. rq_callsize pads that with two credential slacks, so the shortfall has never reached the send buffer. No backport is needed. Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-5-323aa7196055@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Budget the CB_OFFLOAD opcodeChuck Lever
NFS4_enc_cb_offload_sz counts the file handle, the stateid, and the offload information. encode_cb_offload4args() emits an opcode ahead of all three, so the macro falls one XDR word short. This macro sizes p_arglen and nothing else. rq_callsize pads that with two credential slacks, so the shortfall has never reached the send buffer. No backport is needed. Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-4-323aa7196055@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Budget the CB_LAYOUTRECALL recall stateidChuck Lever
NFS4_enc_cb_layout_sz counts the opcode, the three scalar fields, the file handle, and the offset and length hypers. encode_cb_layout4args() also emits the layoutrecall4 discriminator and the recall stateid, so the macro falls five XDR words short. This macro sizes p_arglen and nothing else. rq_callsize pads that with two credential slacks, so the shortfall has never reached the send buffer. No backport is needed. Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-3-323aa7196055@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Budget the CB_RECALL truncate fieldChuck Lever
NFS4_enc_cb_recall_sz counts the CB_RECALL opcode, the stateid, and the file handle. encode_cb_recall4args() also emits the truncate field, so the macro falls one XDR word short. NFSD_CB_MAX_REQ_SZ derives from this macro, so the minimum ca_maxrequestsize a client must advertise rises by four bytes. The field has been unbudgeted since the macro was written. Neither consumer of the macro justifies a backport. rq_callsize covers p_arglen with two credential slacks. The only client affected is one whose ca_maxrequestsize falls inside those four bytes. Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-2-323aa7196055@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Budget the CB_SEQUENCE opcode and referring call array countChuck Lever
cb_sequence_enc_sz counts the session ID, the four scalar fields, and one referring call list. encode_cb_sequence4args() also emits the CB_SEQUENCE opcode and the csa_referring_call_lists array count, so the macro falls two XDR words short. Every NFS4_enc_cb_*_sz built on it is short by the same two words. NFSD_CB_MAX_REQ_SZ derives from NFS4_enc_cb_recall_sz, so the two missing CB_SEQUENCE words shrink the ca_maxrequestsize that check_backchannel_attrs() accepts by eight bytes. Count both words. The minimum a client must advertise rises by those eight bytes. The short count cannot overrun the send buffer. The macro sizes p_arglen, and rq_callsize adds two credential slacks on top of that. The only client affected is one whose ca_maxrequestsize falls inside those eight bytes. No backport is needed. Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-1-323aa7196055@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Don't apply NFS version-specific behavior to LOCALIO requestsChuck Lever
LOCALIO serves NFS clients of every version through one entry point, so no protocol version is associated with such a request. nfsd_set_fh_dentry() selects version-specific behavior anyway: its switch keys off fh_maxsize, and nfsd_open_local_fh() passes NFS4_FHSIZE because that is the size of the buffer it copies into, so LOCALIO lands in the NFSv4 arm. fh_getattr() keys off fh_maxsize too and does run on a LOCALIO open, adding STATX_BTIME and STATX_CHANGE_COOKIE to the mask it requests: work on filesystems that compute them for a caller that never reads them. nfsd_open_local_fh() only verifies a handle it received, so it has no maximum size to state. Pass NFSD_FHSIZE_UNSPEC as nlm_fopen() already does, which selects the switch arm that applies no version-specific behavior, and state the bound on the copy out of struct nfs_fh as NFS_MAXFHSIZE. Suggested-by: NeilBrown <neil@brown.name> Reviewed-by: NeilBrown <neil@brown.name> Reviewed-by: Mike Snitzer <snitzer@kernel.org> Link: https://patch.msgid.link/20260728165911.462534-6-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Name the fh_maxsize value that carries no NFS versionChuck Lever
nfsd_set_fh_dentry() selects behavior specific to an NFS protocol version by matching fh_maxsize against NFS_FHSIZE, NFS3_FHSIZE, or NFS4_FHSIZE. A filehandle that reaches NFSD outside an NFS request has no such version. nlm_fopen() opts out of the switch by passing a bare 0, which matches no arm, and the literal says nothing about why, so an adjacent comment has to carry it. Reviewed-by: NeilBrown <neil@brown.name> Reviewed-by: Mike Snitzer <snitzer@kernel.org> Link: https://patch.msgid.link/20260728165911.462534-5-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Tighten header includes in localio.cChuck Lever
As a prerequisite to converting NFSD to use xdrgen more broadly, NFSD source files should not depend on NFS client headers. fs/nfsd/localio.c is server-side LOCALIO code, yet it pulled in three of them: <linux/nfs_fs.h>, the client inode header (struct nfs_inode, NFS_I(), writeback helpers), which server code never uses; <linux/nfs_xdr.h>, whose only referenced symbol is decode_opaque_fixed(), a static inline that exists to remap the error return to -EIO for client call sites; and the catch-all <linux/nfs.h>. Convert the UUID decoder to call the canonical SUNRPC primitive xdr_stream_decode_opaque_fixed() directly. It is shared by client and server, performs the identical bounds check, and is already reachable through <linux/sunrpc/clnt.h>. With the wrapper gone, localio.c references no symbol from <linux/nfs_xdr.h>, and with that header gone, none of the NFSv3 definitions its structs embed are needed here. Drop all three client includes and add what the file actually uses: struct nfs_fh comes from <linux/nfs_fh.h>, included directly rather than through nfslocalio.h's conditional re-export, and NFS4_FHSIZE from <linux/nfs4.h>. enum nfs_stat and nfs_stat_to_errno continue to come from the already-included <linux/nfs_common.h>. Reviewed-by: NeilBrown <neil@brown.name> Reviewed-by: Mike Snitzer <snitzer@kernel.org> Link: https://patch.msgid.link/20260728165911.462534-4-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13nfs_common: Remove "#include <linux/nfs.h>" from linux/nfslocalio.hChuck Lever
Clean up: linux/nfslocalio.h pulls in linux/nfs.h only for the definition of struct nfs_fh, which now lives in linux/nfs_fh.h. Replace linux/nfs.h with linux/nfs_fh.h so that nfslocalio.h no longer carries uapi/linux/nfs.h into its consumers. Reviewed-by: NeilBrown <neil@brown.name> Reviewed-by: Mike Snitzer <snitzer@kernel.org> Link: https://patch.msgid.link/20260728165911.462534-3-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Move the RPC program definition for LOCALIOChuck Lever
Clean up: The definitions for the LOCALIO program are not needed by most files that include linux/nfs.h. Following the convention used by most other in-kernel RPC program implementations, relocate the LOCALIO program definitions to a localio-specific header. Reviewed-by: NeilBrown <neil@brown.name> Reviewed-by: Mike Snitzer <snitzer@kernel.org> Link: https://patch.msgid.link/20260728165911.462534-2-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13nfsd: fix race between client_info_show() and free_client()Ameer Hamza
client_info_show() renders /proc/fs/nfsd/clients/<id>/info and walks clp->cl_sessions under clp->cl_lock to print each session's slot counts. free_client() tears down the same list without taking cl_lock, and is the only unlocked mutator of cl_sessions. A reader can observe a client mid-teardown because get_nfsdfs_clp() pins the nfs4_client but not its sessions: free_client() frees every session before calling nfsd_client_rmdir(), so an in-flight seq_file reader can follow a list_del()'d node whose ->next now holds LIST_POISON1 and take a general protection fault: Oops: general protection fault, probably for non-canonical address 0xdead00000000014c CPU: 1 UID: 0 PID: 132488 Comm: cat RIP: 0010:client_info_show+0x2bf/0x3d0 RAX: dead000000000100 Call Trace: seq_read_iter+0x12a/0x4b0 seq_read+0xf1/0x130 vfs_read+0xbf/0x350 ksys_read+0x6f/0xf0 do_syscall_64+0x8b/0xcb0 entry_SYSCALL_64_after_hwframe+0x76/0x7e Kernel panic - not syncing: Fatal exception Detach cl_sessions onto a local reaplist under cl_lock, then free the sessions after dropping the lock. Removing entries from cl_sessions under cl_lock matches unhash_session(), and the detach-then-reap shape matches how __destroy_client() reaps cl_delegations. The sessions cannot be freed while cl_lock is held, since free_session() calls nfsd4_del_conns(), which re-acquires it. Reported-by: Nicholas Wolff <nicholas.wolff@truenas.com> Fixes: 601c8cb349c2 ("nfsd: add session slot count to /proc/fs/nfsd/clients/*/info") Cc: stable@vger.kernel.org Signed-off-by: Ameer Hamza <ameer.hamza@truenas.com> Link: https://patch.msgid.link/20260726124658.1715711-1-ameer.hamza@truenas.com Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Replace nfsd_write()'s "stable" argument with "iocb_flags"Chuck Lever
The current nfsd_write() API is not NFS version-agnostic, as it relies on callers to pass an NFSv3 stable_how value to determine the persistence of the requested WRITE. NFSv2 does not use a stable-how value on the wire, and NFSv4 has its own stable_how4 (though stable_how and stable_how4 happen to share the same numeric values). To remove the dependence on NFSv3-specific XDR values from NFSD's generic VFS APIs, replace nfsd_write()'s stable argument with an argument that passes a set of IOCB flags instead of an XDR-defined value. The NFSv4 WRITE and COPY paths had been borrowing the NFSv3 stable_how constants for their own on-the-wire stable values, relying on the numeric coincidence noted above. Convert those sites to the stable_how4 enumerators so the v4 code expresses its own protocol's values directly, with no change in behavior. While here, bound-check the decoded NFSv3 WRITE stable value, as the NFSv4 WRITE decoder already does, and make the nfsd3_writeargs stable field unsigned to suit. The larger benefit is one less NFSv4 dependency on nfs3.h. Link: https://patch.msgid.link/20260723182043.990391-3-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFS: Move definition of enum nfs3_stable_howChuck Lever
Clean up: enum nfs3_stable_how was introduced in NFSv3. NFSv2 has no stable_how on the wire; its write path passes NFS_FILE_SYNC only as a placeholder that the protocol ignores. The stable_how constants describe an NFSv3 wire value, so they belong in linux/nfs3.h. Link: https://patch.msgid.link/20260723182043.990391-2-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Split linux/nfs_ssc.hChuck Lever
The nfs_ssc.h header contains both client- and server-side data structures, which means each of those implementations has to pull in headers from the other. Create a linux/nfsd_ssc.h for the server side APIs which no longer includes uapi/linux/nfs.h either directly or indirectly. Because nfsd_ssc.h drops the transitive include of the NFS client headers, fs/nfsd/nfs4proc.c now includes <linux/pagemap.h> directly for filemap_check_wb_err(). struct nfsd4_ssc_umount_item is private to nfsd. Move it into fs/nfsd/xdr4.h alongside its only consumers rather than into the exported nfsd_ssc.h. As an added clean-up, add missing header guard macros and the struct file and struct vfsmount forward declarations the server prototypes need. Cc: Olga Kornievskaia <okorniev@redhat.com> Cc: Dai Ngo <dai.ngo@oracle.com> Link: https://patch.msgid.link/20260721162306.894558-5-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13nfs_common: Synchronize access to the SSC client ops tableChuck Lever
nfsd42_ssc_open() and nfsd42_ssc_close() load ssc_nfs4_ops without synchronization while nfs42_ssc_register() and nfs42_ssc_unregister() store to it. Those reads are safe today only through a non-obvious invariant: an inter-server copy holds an active vers=4.2 mount of the source across both calls, the mount pins the nfsv4 module through the nfs_client's cl_nfs_mod reference, and unregister runs only at nfsv4 module exit, so it cannot run while a call is in flight. Replace that implicit contract with synchronization local to the broker, so its safety no longer rests on a caller in another subsystem. Read the pointer under RCU so a reader observes it atomically as a valid table or NULL. nfs42_ssc_unregister() stores NULL and then calls synchronize_rcu(), so it cannot return while a reader still holds the pointer. The two readers need different handling because one sleeps and the other does not. sco_close() does not sleep, so nfsd42_ssc_close() runs it to completion inside the RCU read-side section and the synchronize_rcu() in unregister waits for it. __nfs42_ssc_open() does sleep -- it issues a GETATTR RPC to the source server and allocates with GFP_KERNEL -- so it must not run inside an RCU read-side section. Pin the provider module with try_module_get() while still under rcu_read_lock(), drop the lock, invoke the open, then release the module. The reference keeps the provider mapped across the sleep without relying on the caller's mount. If the table has already been torn down the copy gets -EIO. Cc: Olga Kornievskaia <okorniev@redhat.com> Cc: Dai Ngo <dai.ngo@oracle.com> Link: https://patch.msgid.link/20260721162306.894558-4-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Hoist nfs42_ssc_open() into fs/nfs_common/nfs_ssc.cChuck Lever
Refactor: The infrastructure and details for calling the client's ssc_open method can be hidden in nfs_ssc.c. This reduces the SSC footprint in fs/nfsd/nfs4proc.c, a step toward removing that file's dependency on <linux/nfs_fs.h>, which indirectly includes <uapi/linux/nfs.h>. The open and close functions are named "nfsd42_" since they are meant to be invoked only by NFSD. Cc: Olga Kornievskaia <okorniev@redhat.com> Cc: Dai Ngo <dai.ngo@oracle.com> Link: https://patch.msgid.link/20260721162306.894558-3-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13nfs_common: Remove unused nfs_ssc_client_ops infrastructureChuck Lever
Clean up: Commit 75333d48f922 ("NFSD: fix use-after-free in __nfs42_ssc_open()") addressed a use-after-free bug by removing the nfsd4_interssc_disconnect() function. Post-copy clean-up was then delegated to NFSD's laundromat. Since that commit, the nfs_do_sb_deactive() wrapper function and the entire nfs_ssc_client_ops infrastructure no longer have any consumers. This includes nfs_do_sb_deactive(), struct nfs_ssc_client_ops, nfs_ssc_register(), nfs_ssc_unregister(), and related registrations in the NFS client. Cc: Olga Kornievskaia <okorniev@redhat.com> Cc: Dai Ngo <dai.ngo@oracle.com> Link: https://patch.msgid.link/20260721162306.894558-2-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Use struct knfsd_fh in struct pnfs_ff_layoutChuck Lever
The file handle held in struct pnfs_ff_layout is copied directly out of a struct svc_fh, whose fh_handle member is a struct knfsd_fh. Storing the layout's copy as struct nfs_fh instead forced an open-coded field-by-field copy between two unrelated structures. Hold the layout's file handle in struct knfsd_fh so the copy uses the canonical fh_copy_shallow() helper and server code no longer reaches into a separate file handle representation. Cc: Thomas Haynes <loghyr@hammerspace.com> Link: https://patch.msgid.link/20260720141442.783935-4-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13lockd: Switch linux/nfs.h to linux/nfs_fh.hChuck Lever
lockd references "struct nfs_fh" but none of the other definitions in linux/nfs.h. That header also pulls in cred.h, several sunrpc headers, and uapi/linux/nfs.h, none of which lockd needs. A new linux/nfs_fh.h provides "struct nfs_fh" and its helpers without the rest of that surface. Switch lockd's xdr.h to linux/nfs_fh.h, and drop the now-redundant linux/nfs.h includes from svc.c and trace.h. Link: https://patch.msgid.link/20260720141442.783935-3-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFS: Add linux/nfs_fh.hChuck Lever
Plenty of spots around the kernel need the full definition of struct nfs_fh but not the cred, sunrpc, and uapi dependencies that linux/nfs.h pulls in along with it. Relocate struct nfs_fh to its own header, and include that header in linux/nfs.h so existing consumers keep building. Over time, consumers can then replace #include <linux/nfs.h> with #include <linux/nfs_fh.h> While relocating the code, add kernel-doc comments for the FH operations and convert nfs_compare_fh() to return bool. Link: https://patch.msgid.link/20260720141442.783935-2-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Map flex file layout IDs through the request's user namespaceChuck Lever
nfsd4_ff_proc_layoutget() and nfsd4_ff_encode_layoutget() translate the file's owner and group with init_user_ns, but every other identity nfsd places on the wire goes through nfsd_user_namespace(). When the transport carries a credential from a non-initial user namespace, the flex file layout reports host-global IDs. The client copies those IDs into the AUTH_SYS credential it presents to the data server, and svcauth_unix_accept() resolves that credential in the transport's namespace, so data server I/O runs under an identity unrelated to the file's owner. Switching to the request's namespace introduces a second hazard. from_kuid() returns (uid_t)-1 when the target namespace has no mapping for the owner, and the IOMODE_READ arm adds one to that result to derive an identity for which the data server denies writes. The addition would wrap to zero, handing the client uid 0 instead of an identity distinct from the owner. Translate both IDs in nfsd4_ff_proc_layoutget(), which has the svc_rqst, and carry the wire values in struct pnfs_ff_layout. from_kuid_munged() substitutes overflowuid for an unmapped owner and thus never returns (uid_t)-1, so the increment cannot wrap to zero. Fixes: 9b9960a0ca47 ("nfsd: Add a super simple flex file server") Cc: stable@vger.kernel.org Reported-by: sashiko-bot <sashiko-bot@kernel.org> Closes: https://sashiko.dev/#/patchset/20260720141442.783935-1-cel@kernel.org?part=3 Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260720163724.810227-1-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Move nfsd_v4client() out of nfsd.hChuck Lever
nfsd_v4client() is the last user in nfsd.h of XDR-defined item references. Once this helper moves out, <linux/nfs.h> and the other XDR-specific headers nfsd.h includes have nothing left to provide, and can be dropped. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260717184112.507548-7-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Evacuate NFSv4 entry-point prototypes from nfsd.hChuck Lever
nfsd.h is included by nearly every NFSD translation unit, yet the two blocks of NFSv4 lifecycle and control prototypes it carries are referenced by only seven of them (out of over two dozen). The remaining consumers, including the NFSv2 and NFSv3 ACL and XDR paths, parse these declarations for no benefit. The declarations cannot simply move into an NFSv4-only header such as state.h: nfssvc.c, nfsctl.c, vfs.c, and export.c call the routines unconditionally and rely on the CONFIG_NFSD_V4=n stubs, and pulling the heavy state.h types into those lean translation units to obtain a handful of prototypes would trade one form of coupling for a worse one. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260717184112.507548-6-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Relocate NFSv4-internal constants to state.hChuck Lever
The COMPOUND encode-slack sizes and the state-management timeouts at the tail of nfsd.h are evaluated only by NFSv4 code (nfs4state.c, nfs4proc.c, and nfs4xdr.c). They nonetheless sit in nfsd.h, where every NFSD translation unit, including the NFSv2 and NFSv3 paths that have no use for them, has to parse them. All three consumers already reach state.h through xdr4.h, so move the block there. nfsd.h keeps the NFSv4 prototypes for now. Only the pure constants move. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260717184112.507548-5-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Remove two unused NFSv4 constantsChuck Lever
Neither COMPOUND_SLACK_SPACE nor NFSD_COURTESY_CLIENT_TIMEOUT has a remaining user. COMPOUND_SLACK_SPACE lost its last reference in commit ea8d7720b274 ("nfsd4: remove redundant encode buffer size checking"), which deleted the encode buffer-space check the macro fed; the comment above it still describes that departed check. NFSD_COURTESY_CLIENT_TIMEOUT is likewise unreferenced: the courteous-server code expires clients through the laundromat's reaper and conflict paths, never a fixed 24-hour timer. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260717184112.507548-4-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Move pre-xdr'ed status codes out of nfsd.hChuck Lever
nfsd.h is included by nearly every NFSD translation unit, so its include of <linux/nfs4.h> reaches all of them, whether or not they touch NFSv4. That include existed solely for the block of pre-xdr'ed nfserr_* values at the end of the file: several of those values, such as nfserr_delay and nfserr_admin_revoked, are built from NFS4ERR_* constants defined in nfs4.h. The NFSD-internal error enum that follows the block (NFSERR_EOF and friends) is anchored at an impossible nfsstat4 value, thus it also needs nothing from nfs4.h. But these codes are used by all NFS versions, so their new home must be version-neutral. Move the pre-xdr'ed value block and the internal error enum into a new fs/nfsd/nfserr.h, which includes nfs4.h itself, and drop the nfs4.h include from nfsd.h. Include nfserr.h directly from each translation unit that references the pre-xdr'ed values or the internal error codes, rather than carrying it in a widely-included header. A translation unit that includes nfsd.h without using the error block no longer pulls in nfs4.h. The ones that reference the block can still reach it through nfserr.h. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260717184112.507548-3-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13NFSD: Move XDR encoding helpers out of xdr4.hChuck Lever
These static inline helpers use the nfserr_resource macro, which pulls in the whole rack of NFS status codes. Move those helpers into the only file that uses them, to get rid of the nfserr macro dependency globally. These helper were originally placed in xdr4.h because I thought they would be utilized in the rest of the NFSv4 XDR code, but XDR translation is eventually to be subsumed by xdrgen instead. I'm not converting them now because that is much more churn than this patch is. Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260717184112.507548-2-cel@kernel.org Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13nfsd: separate out VFS-specific code from nfsd4_create_file()NeilBrown
All the code in nfsd4_create_file() that is VFS manipulation, with now NFS-specific knowledge, has been localised. Now we split that out into a separate function: do_lookup_open(). It is planned to provide a vfs_lookup_open() in vfs code which provides this functionality. This will share more code with the syscall open path, and make it easier to modify locking at the VFS level. Signed-off-by: NeilBrown <neil@brown.name> Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260717093001.1972119-18-neilb@ownmail.net Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13nfsd: move v0 checking out of nfsd_check_obj_isreg()NeilBrown
A future patch will use nfsd_check_obj_isreg() in a context where the protocol version is not easily available. So move the version check out and put it at the end of do_open_lookup(). Also change to return errno error code and use nfserrno() to convert to nfs error codes. Use -ELOOP for nfserr_symlink, which is an error indication a problem with symlinks. -EFTYPE is a good match for nfserr_wrong_type. Signed-off-by: NeilBrown <neil@brown.name> Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260717093001.1972119-17-neilb@ownmail.net Signed-off-by: Chuck Lever <cel@kernel.org>
2026-09-13nfsd: reduce want-write range in nfsd4_create_file()NeilBrown
nfsd4_create_file() needs write access to the mount for two purposes: 1/ to create/open the file. 2/ to set attributes on the newly created (or pre-existing) file. Currently this is all handled by holding the write access across the open and the setattr. A subsequent patch will necessarily change how write access is gained for the open. So we reduce the range for the first want_write, and add another one to cover setattr. If we failed to get write access, it is only fatal if there were attrs to set. We call nfsd_create_setattr() if at all possible, even when no attrs, as it also calls commit_metadata and we need to be certain that the file creation has been synced. If the mount became read-only since the creation happened, we can safely assume that the sync happened as part of that. Signed-off-by: NeilBrown <neil@brown.name> Reviewed-by: Jeff Layton <jlayton@kernel.org> Link: https://patch.msgid.link/20260717093001.1972119-16-neilb@ownmail.net Signed-off-by: Chuck Lever <cel@kernel.org>