| Age | Commit message (Collapse) | Author |
|
When rpc.mountd cannot resolve a uid (getpwuid() or getgrouplist()
failure, e.g. while winbind or sssd is briefly unreachable), it
answers the auth.unix.gid upcall with zero groups. unix_gid_parse()
installs that as a valid positive entry, and svcauth_unix_set_client()
then replaces the credential's group list with the empty one on
every request, RPCSEC_GSS included via svcauth_gss_set_client().
One failed lookup strips that uid of all supplementary groups on
every export for up to mountd's configured TTL (30 minutes by
default), long after the NSS backend has recovered.
mountd cannot send an empty list for a successful lookup, since
getgrouplist(3) always includes at least the user's primary group,
so a zero-group reply can only mean the lookup failed. Record it as
a negative entry: unix_gid_find() then returns -ENOENT and
svcauth_unix_set_client() keeps the groups the RPC credential
already carries. This is the fallback that
commit 3fc605a2aa38 ("[PATCH] knfsd: allow the server to provide a
gid list when using AUTH_UNIX authentication") promised when no
answer is available, and the same state try_to_negate_entry()
already creates when no listener holds the channel open.
Fixes: 3fc605a2aa38 ("[PATCH] knfsd: allow the server to provide a gid list when using AUTH_UNIX authentication")
Cc: stable@vger.kernel.org
Assisted-by: Claude:claude-fable-5
Signed-off-by: Ameer Hamza <ameer.hamza@truenas.com>
Link: https://patch.msgid.link/20260814221953.108837-2-ameer.hamza@truenas.com
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd4_cb_notify_prepare() reserved a word from the encoding xdr stream
for each notify4's host-order notify_mask, storing the pointer in
ncn_nf[].notify_mask.element. That element must survive until the RPC
encode re-reads ncn_nf, so the host-endian word lived inside the XDR
staging buffer for the whole callback lifetime - the same fragile
pattern as the attrmask, and one that keeps host-order bytes in a
buffer meant to hold big-endian XDR.
ncn_nf is a bounded per-delegation array reused across every CB_NOTIFY,
so give it a parallel ncn_masks array with the same lifetime:
- allocate/free ncn_masks alongside ncn_nf in alloc_init_dir_deleg() /
nfs4_free_dir_deleg()
- point notify_mask.element at &ncn_masks[i] (events) and
&ncn_masks[count] (dir attr change)
The mask backing is now pre-allocated, so the per-word NULL checks in
prepare go away. The staging stream holds only encoded XDR.
Assisted-by: LLM
Signed-off-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260812-dir-deleg-v1-2-411faa713068@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd4_setup_notify_entry4() stole 3 words from the xdr stream via
xdr_reserve_space() to hold the host-order bmval[3] attrmask that
nfsd4_encode_attr_vals() consumes and ne_attrs.attrmask.element points
at. Stashing host-endian scratch in an XDR stream buffer is fragile:
the buffer layout is not guaranteed by sunrpc, and it blocks moving the
encoder to pages or xdrgen.
The attrmask only needs to live until the enclosing encode call
serializes the notify_entry4, so hand it caller-provided stack storage
instead. The two callers keep the words on the stack: up to three
concurrent entries for a rename in nfsd4_encode_notify_event(), one in
nfsd4_encode_dir_attr_change().
No wire change: the reserved words were never emitted; attr_vals.data/len
are still captured relative to xdr->p.
Assisted-by: LLM
Signed-off-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260812-dir-deleg-v1-1-411faa713068@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Clean up: Nothing in fs/nfsd/vfs.c references an NFSv3 XDR definition.
The preceding patch moved the last NFSv4 reference, a struct
nfsd4_compoundres dereference that recovered a pair of file handles for
a tracepoint, into fs/nfsd/nfs4proc.c. Remove both includes so that the
VFS layer no longer names on-the-wire types.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260812142436.35042-4-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd4_clone_file_range() lives in fs/nfsd/vfs.c but reaches into the
NFSv4 compound reply buffer: nfsd4_get_cstate() casts rq_resp to a
struct nfsd4_compoundres to recover the saved and current file
handles a tracepoint wants. That is the only reference to the NFSv4
XDR definitions left in vfs.c, and it puts knowledge of the compound
reply layout in the VFS layer.
Refactor nfsd4_clone_file_range() to remove NFSv4-specific
componentry from fs/nfsd/vfs.c.
Splitting nfsd_clone_file_range() and nfsd_clone_sync_range() lets
nfsd4_clone() distinguish a clone failure from a sync failure, which
it has to do because only the latter invalidates the write verifier.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260812142436.35042-3-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
A subsequent patch moves the NFSv4-specific portions of
nfsd4_clone_file_range() out of fs/nfsd/vfs.c and into its caller in
fs/nfsd/nfs4proc.c. One of those portions resets the write verifier
when the post-clone sync fails.
netns.h already exposes nfsd_reset_write_verifier(), but that is the
unconditional reset. commit_reset_write_verifier() wraps it with the
policy that decides which errors warrant a reset: -EAGAIN and -ESTALE
do not indicate a problem with durable storage, so they leave the
verifier alone. A caller outside vfs.c has to apply the same policy,
so make the wrapper visible rather than duplicate its switch.
Rename it to nfsd_maybe_reset_write_verifier() on the way out.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260812142436.35042-2-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd_access() owns three static tables that map on-the-wire ACCESS
bits to NFSD_MAY flags, and every NFS version shares them. That puts
protocol-version specifics in the version-agnostic VFS layer. The
NFSv4.2 extended-attribute bits are wedged into the NFSv3 tables
under CONFIG_NFSD_V4.
Give each version its own tables in its proc code and pass the
matching set into nfsd_access() as a new argument.
Splitting the tables also stops the NFSv2-ACL and NFSv3 ACCESS paths
from answering for the NFSv4.2 extended-attribute bits. Both use the
NFSv4-augmented table whenever CONFIG_NFSD_V4 is set, so an NFSv3
request that sets an xattr bit has it echoed back in the reply even
though those bits are undefined for v3.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260810141900.33846-2-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Clean up: Remove an NFSv3 constant (NFS3_ACCESS_FULL) used inside an
NFSv4 code path.
After this patch is applied, the nfs3.h header is no longer an
implicit dependency of nfs4proc.c for this value. The definition
itself stays in nfs3.h to avoid a kernel-user space API regression.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260810141900.33846-1-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
xs_sock_process_cmsg() switches on the TLS record type, and every arm
but TLS_RECORD_TYPE_ALERT returns the -EAGAIN its caller passed in.
The DATA arm clears MSG_EOR in the caller's msghdr, but
xs_sock_recvmsg() has already cleared that flag before the call.
Deriving the record type a second time inside the helper also fires
trace_tls_contenttype() twice for every alert.
Move the alert handling into xs_sock_recv_cmsg() and delete the
helper. Every other record type still returns -EAGAIN. The DATA arm's
account of MSG_EOR moves to xs_sock_recvmsg(), where the flag is
cleared.
Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-7-62d9a631c880@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
xs_sock_process_cmsg() decides whether an alert ends the session by
reading the alert's level octet. RFC 8446 Section 6 retired that
field. The severity is implicit in the description, and a receiver
treats every alert listed in Section 6.2 as an error alert
"regardless of the AlertLevel in the message". A peer that aborts
with unexpected_message but leaves the legacy octet set to warning
makes the client return -EAGAIN. xs_stream_data_receive() wakes no
pending task for that error, so RPC Calls queued on a dead TLS
session wait for their timeouts to expire.
Decide from the alert description instead. close_notify and
user_canceled are the closure alerts (RFC 8446 Section 6.1). Every
other description ends the session, including one this kernel does
not recognize.
Fixes: 39067dda1d86 ("SUNRPC: Use new helpers to handle TLS Alerts")
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-6-62d9a631c880@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
tls_alert_recv() reads two octets from the kvec it is handed and does
not check the length (net/handshake/alert.c). xs_sock_process_cmsg()
calls it for any alert record, and the alert[] buffer that
xs_sock_recv_cmsg() supplies carries no initializer. A one-octet alert
body leaves the description read from uninitialized stack and reported
through trace_tls_alert_recv().
The peer controls that length. Neither tls_rx_msg_size() nor
tls_rx_one_record() enforces the two-octet Alert payload. A TLS 1.3
record carrying only the inner content-type octet decrypts to a
zero-length payload. RFC 8446 Section 5.1 requires a record with an
Alert type to carry exactly one message, so any other length is
malformed. RFC 9289 Section 5 bars RPC-with-TLS from negotiating a
version below TLS 1.3, so no other alert framing applies.
Require exactly two octets before parsing and return -EACCES
otherwise. xs_stream_data_receive() already treats -EACCES as a fatal
alert and reports it to the pending tasks. Gate the path on a control
message rather than a positive count so that a zero-length record
reaches the check.
Fixes: cc5d59081fa2 ("sunrpc: fix client side handling of tls alerts")
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-5-62d9a631c880@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
A TLS control record delivers no payload to the RPC layer.
svc_tcp_recvfrom() clears XPT_DATA before the receive, and
svc_tcp_sock_recv_cmsg() returns -EAGAIN for the record it consumed.
Nothing marks the transport ready again. kTLS raises data_ready for
arriving TCP segments, not for records it has already decrypted. An
RPC Call queued behind an alert or a KeyUpdate waits until the client
sends more. The client blocks until its RPC timeout expires.
The receive takes only the first two octets of the record. kTLS holds
the remainder on its receive list, where each later receive takes two
octets more.
Drain a record that is not an alert, then mark the transport ready
once a control record has been consumed.
Fixes: 5e052dda121e ("SUNRPC: Recognize control messages in server-side TCP socket code")
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-4-62d9a631c880@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
svc_tcp_sock_process_cmsg() decides whether an alert ends the session
by reading the alert's level octet. RFC 8446 Section 6 retired that
field. The severity is implicit in the description, and a receiver
treats every alert listed in Section 6.2 as an error alert
"regardless of the AlertLevel in the message". A peer that aborts
with unexpected_message but leaves the legacy octet set to warning
makes the server return -EAGAIN. svc_tcp_recvfrom() then leaves a
dead TLS session attached to an open transport. NFSD keeps polling
it.
Decide from the alert description instead. close_notify and
user_canceled are the closure alerts (RFC 8446 Section 6.1). Every
other description ends the session, including one this kernel does
not recognize.
Fixes: 39067dda1d86 ("SUNRPC: Use new helpers to handle TLS Alerts")
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-3-62d9a631c880@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
tls_alert_recv() reads two octets from the kvec it is handed and does
not check the length (net/handshake/alert.c). svc_tcp_sock_recv_cmsg()
calls it for any positive receive, and the alert[] buffer it supplies
carries no initializer. A one-octet alert body leaves the description
read from uninitialized stack and reported through
trace_tls_alert_recv().
The peer controls that length. Neither tls_rx_msg_size() nor
tls_rx_one_record() enforces the two-octet Alert payload. A TLS 1.3
record carrying only the inner content-type octet decrypts to a
zero-length payload. RFC 8446 Section 5.1 requires a record with an
Alert type to carry exactly one message, so any other length is
malformed.
Require exactly two octets before parsing and return -EBADMSG
otherwise. That closes the transport rather than acting on a partly
uninitialized alert. Gate the path on a control message rather than a
positive count so that a zero-length record reaches the check.
Fixes: bee47cb026e7 ("sunrpc: fix handling of server side tls alerts")
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-2-62d9a631c880@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
svc_tcp_sock_recv_cmsg() receives up to two octets into a local
buffer, and returns that count for any record type other than
TLS_RECORD_TYPE_ALERT. Nothing reached the caller's buffer, but
svc_tcp_read_marker() adds the count to sk_tcplen and
svc_tcp_read_msg()'s caller adds it to sk_datalen. The RPC stream
advances over octets it never received. The fragment marker is
assembled from stale sk_marker octets. The message body comes from
pages nothing wrote.
A conforming client reaches this. RFC 8446 Section 4.6.3 lets either
peer send KeyUpdate once it has sent its Finished, and svcsock has no
rekey path. kTLS leaves the partially consumed record on ctx->rx_list,
so the body drains two octets per svc_tcp_recvfrom() call. Each pair
is credited the same way.
Return -EAGAIN for a record that is not an alert. That is what
svc_tcp_sock_process_cmsg()'s default arm returned before the receive
moved into a local buffer.
Fixes: bee47cb026e7 ("sunrpc: fix handling of server side tls alerts")
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260808-svcsock-cmsg-fixes-v3-1-62d9a631c880@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Scripting and automation is sensitive to branch names in the
subsystem entries in MAINTAINERS. Rather than pulling from
cel.git/master, we really want CI to pull from
cel.git/nfsd-testing.
Link: https://patch.msgid.link/20260804184630.1395002-1-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
When CB_RECALL races ahead of the reply that granted the delegation,
the client has not yet recorded the delegation stateid and responds
NFS4ERR_BADHANDLE or NFS4ERR_BAD_STATEID. The slot that carried the
grant has not retired at that point, so NFSD cannot read the rejection
as proof that the client never held the delegation. It retries the
recall and, once the retries lapse, revokes a delegation the client is
by then able to return.
Remove the ambiguity with the referring call mechanism of RFC 8881
Section 2.10.6.3: until the slot that carried the grant retires, name
that request as a referring call in the CB_SEQUENCE of each recall. A
client that finds it still outstanding may respond NFS4ERR_DELAY, and
the recall is retried until the client has processed the grant.
A recall reuses one callback context across its retries, and ->prepare
does not run on every send. The granting request does not change, so a
send that inherits the previous list sends the right one. Retirement
of the granting slot drops the list, and nfs4_free_deleg() releases
what is left.
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-9-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
A client that answers CB_RECALL with NFS4ERR_BADHANDLE or
NFS4ERR_BAD_STATEID has no record of the delegation, so the
FREE_STATEID that clears it from cl_revoked never arrives. Every later
SEQUENCE reply carries SEQ4_STATUS_RECALLABLE_STATE_REVOKED, and the
client loops issuing TEST_STATEID.
Destroy such a delegation when it is reaped rather than revoking it
onto cl_revoked. RFC 8881 Section 20.2.4 completes the recall at the
reply when its status is neither NFS4_OK nor NFS4ERR_DELAY, so a
rejected recall leaves nothing to revoke. An administrative revoke
keeps that path, since NFS4ERR_ADMIN_REVOKED reports it. A destroyed
stateid returns NFS4ERR_BAD_STATEID instead of NFS4ERR_DELEG_REVOKED. A
client that rejects the recall but still holds the delegation gets no
notice that its state was revoked.
CB_RECALL can outrun the reply that granted the delegation, so honor
a rejection only once the client has seen that grant. Per RFC 8881
Section 2.10.6.3, retirement of the slot that carried the grant is that
proof; retry until then, and revoke when the retries lapse.
Fixes: 3bd64a5ba171 ("nfsd4: implement SEQ4_STATUS_RECALLABLE_STATE_REVOKED")
Cc: stable@vger.kernel.org # 6.14.x
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-8-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
The comment above the SC_STATUS_ flags states that nn->deleg_lock
protects sc_status for delegation stateids, but only the transitions
made while a delegation is hashed are taken under that lock. This
comment was accurate until commit c88c150a467f ("nfsd: fix possible
badness in FREE_STATEID") set SC_STATUS_CLOSED under ->cl_lock.
Commit 8dd91e8d31fe ("nfsd: fix race between laundromat and
free_stateid") added the other two sites.
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-7-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
NFS4_enc_cb_recall_any_sz counts the objects-to-keep field, the
bitmap array length, and the bitmap word. encode_cb_recallany4args()
emits an opcode ahead of all three, so the macro falls one XDR word
short.
This macro sizes p_arglen and nothing else. rq_callsize pads that
with two credential slacks, so the shortfall has never reached
the send buffer. No backport is needed.
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-6-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
NFS4_enc_cb_notify_lock_sz counts the lock owner and the file
handle. nfs4_xdr_enc_cb_notify_lock() emits an opcode ahead of both,
so the macro falls one XDR word short.
This macro sizes p_arglen and nothing else. rq_callsize pads that
with two credential slacks, so the shortfall has never reached
the send buffer. No backport is needed.
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-5-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
NFS4_enc_cb_offload_sz counts the file handle, the stateid, and the
offload information. encode_cb_offload4args() emits an opcode ahead
of all three, so the macro falls one XDR word short.
This macro sizes p_arglen and nothing else. rq_callsize pads that
with two credential slacks, so the shortfall has never reached
the send buffer. No backport is needed.
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-4-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
NFS4_enc_cb_layout_sz counts the opcode, the three scalar fields,
the file handle, and the offset and length hypers.
encode_cb_layout4args() also emits the layoutrecall4 discriminator
and the recall stateid, so the macro falls five XDR words short.
This macro sizes p_arglen and nothing else. rq_callsize pads that
with two credential slacks, so the shortfall has never reached
the send buffer. No backport is needed.
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-3-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
NFS4_enc_cb_recall_sz counts the CB_RECALL opcode, the stateid, and the
file handle. encode_cb_recall4args() also emits the truncate field, so
the macro falls one XDR word short.
NFSD_CB_MAX_REQ_SZ derives from this macro, so the minimum
ca_maxrequestsize a client must advertise rises by four bytes.
The field has been unbudgeted since the macro was written. Neither
consumer of the macro justifies a backport. rq_callsize covers
p_arglen with two credential slacks. The only client affected is one
whose ca_maxrequestsize falls inside those four bytes.
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-2-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
cb_sequence_enc_sz counts the session ID, the four scalar fields, and
one referring call list. encode_cb_sequence4args() also emits the
CB_SEQUENCE opcode and the csa_referring_call_lists array count, so
the macro falls two XDR words short. Every NFS4_enc_cb_*_sz built on
it is short by the same two words.
NFSD_CB_MAX_REQ_SZ derives from NFS4_enc_cb_recall_sz, so the two
missing CB_SEQUENCE words shrink the ca_maxrequestsize that
check_backchannel_attrs() accepts by eight bytes. Count both words.
The minimum a client must advertise rises by those eight bytes.
The short count cannot overrun the send buffer. The macro sizes
p_arglen, and rq_callsize adds two credential slacks on top of that.
The only client affected is one whose ca_maxrequestsize falls inside
those eight bytes. No backport is needed.
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-1-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
LOCALIO serves NFS clients of every version through one entry point,
so no protocol version is associated with such a request.
nfsd_set_fh_dentry() selects version-specific behavior anyway: its
switch keys off fh_maxsize, and nfsd_open_local_fh() passes
NFS4_FHSIZE because that is the size of the buffer it copies into, so
LOCALIO lands in the NFSv4 arm. fh_getattr() keys off fh_maxsize too
and does run on a LOCALIO open, adding STATX_BTIME and
STATX_CHANGE_COOKIE to the mask it requests: work on filesystems that
compute them for a caller that never reads them.
nfsd_open_local_fh() only verifies a handle it received, so it has no
maximum size to state. Pass NFSD_FHSIZE_UNSPEC as nlm_fopen() already
does, which selects the switch arm that applies no version-specific
behavior, and state the bound on the copy out of struct nfs_fh as
NFS_MAXFHSIZE.
Suggested-by: NeilBrown <neil@brown.name>
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Mike Snitzer <snitzer@kernel.org>
Link: https://patch.msgid.link/20260728165911.462534-6-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd_set_fh_dentry() selects behavior specific to an NFS protocol
version by matching fh_maxsize against NFS_FHSIZE, NFS3_FHSIZE, or
NFS4_FHSIZE. A filehandle that reaches NFSD outside an NFS request has
no such version. nlm_fopen() opts out of the switch by passing a bare
0, which matches no arm, and the literal says nothing about why, so an
adjacent comment has to carry it.
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Mike Snitzer <snitzer@kernel.org>
Link: https://patch.msgid.link/20260728165911.462534-5-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
As a prerequisite to converting NFSD to use xdrgen more broadly,
NFSD source files should not depend on NFS client headers.
fs/nfsd/localio.c is server-side LOCALIO code, yet it pulled in
three of them: <linux/nfs_fs.h>, the client inode header (struct
nfs_inode, NFS_I(), writeback helpers), which server code never
uses; <linux/nfs_xdr.h>, whose only referenced symbol is
decode_opaque_fixed(), a static inline that exists to remap the
error return to -EIO for client call sites; and the catch-all
<linux/nfs.h>.
Convert the UUID decoder to call the canonical SUNRPC primitive
xdr_stream_decode_opaque_fixed() directly. It is shared by client
and server, performs the identical bounds check, and is already
reachable through <linux/sunrpc/clnt.h>. With the wrapper gone,
localio.c references no symbol from <linux/nfs_xdr.h>, and with
that header gone, none of the NFSv3 definitions its structs embed
are needed here.
Drop all three client includes and add what the file actually
uses: struct nfs_fh comes from <linux/nfs_fh.h>, included directly
rather than through nfslocalio.h's conditional re-export, and
NFS4_FHSIZE from <linux/nfs4.h>. enum nfs_stat and
nfs_stat_to_errno continue to come from the already-included
<linux/nfs_common.h>.
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Mike Snitzer <snitzer@kernel.org>
Link: https://patch.msgid.link/20260728165911.462534-4-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Clean up: linux/nfslocalio.h pulls in linux/nfs.h only for the
definition of struct nfs_fh, which now lives in linux/nfs_fh.h.
Replace linux/nfs.h with linux/nfs_fh.h so that nfslocalio.h no longer
carries uapi/linux/nfs.h into its consumers.
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Mike Snitzer <snitzer@kernel.org>
Link: https://patch.msgid.link/20260728165911.462534-3-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Clean up: The definitions for the LOCALIO program are not needed by
most files that include linux/nfs.h. Following the convention used
by most other in-kernel RPC program implementations, relocate the
LOCALIO program definitions to a localio-specific header.
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Mike Snitzer <snitzer@kernel.org>
Link: https://patch.msgid.link/20260728165911.462534-2-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
client_info_show() renders /proc/fs/nfsd/clients/<id>/info and walks
clp->cl_sessions under clp->cl_lock to print each session's slot
counts. free_client() tears down the same list without taking
cl_lock, and is the only unlocked mutator of cl_sessions. A reader
can observe a client mid-teardown because get_nfsdfs_clp() pins the
nfs4_client but not its sessions: free_client() frees every session
before calling nfsd_client_rmdir(), so an in-flight seq_file reader
can follow a list_del()'d node whose ->next now holds LIST_POISON1
and take a general protection fault:
Oops: general protection fault, probably for non-canonical address
0xdead00000000014c
CPU: 1 UID: 0 PID: 132488 Comm: cat
RIP: 0010:client_info_show+0x2bf/0x3d0
RAX: dead000000000100
Call Trace:
seq_read_iter+0x12a/0x4b0
seq_read+0xf1/0x130
vfs_read+0xbf/0x350
ksys_read+0x6f/0xf0
do_syscall_64+0x8b/0xcb0
entry_SYSCALL_64_after_hwframe+0x76/0x7e
Kernel panic - not syncing: Fatal exception
Detach cl_sessions onto a local reaplist under cl_lock, then free the
sessions after dropping the lock. Removing entries from cl_sessions
under cl_lock matches unhash_session(), and the detach-then-reap shape
matches how __destroy_client() reaps cl_delegations. The sessions
cannot be freed while cl_lock is held, since free_session() calls
nfsd4_del_conns(), which re-acquires it.
Reported-by: Nicholas Wolff <nicholas.wolff@truenas.com>
Fixes: 601c8cb349c2 ("nfsd: add session slot count to /proc/fs/nfsd/clients/*/info")
Cc: stable@vger.kernel.org
Signed-off-by: Ameer Hamza <ameer.hamza@truenas.com>
Link: https://patch.msgid.link/20260726124658.1715711-1-ameer.hamza@truenas.com
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
The current nfsd_write() API is not NFS version-agnostic, as it
relies on callers to pass an NFSv3 stable_how value to determine
the persistence of the requested WRITE. NFSv2 does not use a
stable-how value on the wire, and NFSv4 has its own stable_how4
(though stable_how and stable_how4 happen to share the same
numeric values).
To remove the dependence on NFSv3-specific XDR values from NFSD's
generic VFS APIs, replace nfsd_write()'s stable argument with an
argument that passes a set of IOCB flags instead of an XDR-defined
value.
The NFSv4 WRITE and COPY paths had been borrowing the NFSv3
stable_how constants for their own on-the-wire stable values,
relying on the numeric coincidence noted above. Convert those
sites to the stable_how4 enumerators so the v4 code expresses
its own protocol's values directly, with no change in behavior.
While here, bound-check the decoded NFSv3 WRITE stable value, as
the NFSv4 WRITE decoder already does, and make the nfsd3_writeargs
stable field unsigned to suit.
The larger benefit is one less NFSv4 dependency on nfs3.h.
Link: https://patch.msgid.link/20260723182043.990391-3-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Clean up: enum nfs3_stable_how was introduced in NFSv3. NFSv2 has no
stable_how on the wire; its write path passes NFS_FILE_SYNC only as
a placeholder that the protocol ignores. The stable_how constants
describe an NFSv3 wire value, so they belong in linux/nfs3.h.
Link: https://patch.msgid.link/20260723182043.990391-2-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
The nfs_ssc.h header contains both client- and server-side data
structures, which means each of those implementations has to pull in
headers from the other.
Create a linux/nfsd_ssc.h for the server side APIs which no longer
includes uapi/linux/nfs.h either directly or indirectly. Because
nfsd_ssc.h drops the transitive include of the NFS client headers,
fs/nfsd/nfs4proc.c now includes <linux/pagemap.h> directly for
filemap_check_wb_err().
struct nfsd4_ssc_umount_item is private to nfsd. Move it into
fs/nfsd/xdr4.h alongside its only consumers rather than into the
exported nfsd_ssc.h.
As an added clean-up, add missing header guard macros and the
struct file and struct vfsmount forward declarations the server
prototypes need.
Cc: Olga Kornievskaia <okorniev@redhat.com>
Cc: Dai Ngo <dai.ngo@oracle.com>
Link: https://patch.msgid.link/20260721162306.894558-5-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd42_ssc_open() and nfsd42_ssc_close() load ssc_nfs4_ops without
synchronization while nfs42_ssc_register() and nfs42_ssc_unregister()
store to it. Those reads are safe today only through a non-obvious
invariant: an inter-server copy holds an active vers=4.2 mount of the
source across both calls, the mount pins the nfsv4 module through the
nfs_client's cl_nfs_mod reference, and unregister runs only at nfsv4
module exit, so it cannot run while a call is in flight. Replace that
implicit contract with synchronization local to the broker, so its
safety no longer rests on a caller in another subsystem.
Read the pointer under RCU so a reader observes it atomically as a
valid table or NULL. nfs42_ssc_unregister() stores NULL and then calls
synchronize_rcu(), so it cannot return while a reader still holds the
pointer.
The two readers need different handling because one sleeps and the
other does not. sco_close() does not sleep, so nfsd42_ssc_close() runs
it to completion inside the RCU read-side section and the
synchronize_rcu() in unregister waits for it. __nfs42_ssc_open() does
sleep -- it issues a GETATTR RPC to the source server and allocates
with GFP_KERNEL -- so it must not run inside an RCU read-side section.
Pin the provider module with try_module_get() while still under
rcu_read_lock(), drop the lock, invoke the open, then release the
module. The reference keeps the provider mapped across the sleep
without relying on the caller's mount. If the table has already been
torn down the copy gets -EIO.
Cc: Olga Kornievskaia <okorniev@redhat.com>
Cc: Dai Ngo <dai.ngo@oracle.com>
Link: https://patch.msgid.link/20260721162306.894558-4-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Refactor: The infrastructure and details for calling the client's
ssc_open method can be hidden in nfs_ssc.c. This reduces the SSC
footprint in fs/nfsd/nfs4proc.c, a step toward removing that
file's dependency on <linux/nfs_fs.h>, which indirectly includes
<uapi/linux/nfs.h>.
The open and close functions are named "nfsd42_" since they are
meant to be invoked only by NFSD.
Cc: Olga Kornievskaia <okorniev@redhat.com>
Cc: Dai Ngo <dai.ngo@oracle.com>
Link: https://patch.msgid.link/20260721162306.894558-3-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Clean up: Commit 75333d48f922 ("NFSD: fix use-after-free in
__nfs42_ssc_open()") addressed a use-after-free bug by removing the
nfsd4_interssc_disconnect() function. Post-copy clean-up was then
delegated to NFSD's laundromat.
Since that commit, the nfs_do_sb_deactive() wrapper function and the
entire nfs_ssc_client_ops infrastructure no longer have any
consumers. This includes nfs_do_sb_deactive(), struct
nfs_ssc_client_ops, nfs_ssc_register(), nfs_ssc_unregister(), and
related registrations in the NFS client.
Cc: Olga Kornievskaia <okorniev@redhat.com>
Cc: Dai Ngo <dai.ngo@oracle.com>
Link: https://patch.msgid.link/20260721162306.894558-2-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
The file handle held in struct pnfs_ff_layout is copied directly out
of a struct svc_fh, whose fh_handle member is a struct knfsd_fh.
Storing the layout's copy as struct nfs_fh instead forced an
open-coded field-by-field copy between two unrelated structures.
Hold the layout's file handle in struct knfsd_fh so the copy uses the
canonical fh_copy_shallow() helper and server code no longer reaches
into a separate file handle representation.
Cc: Thomas Haynes <loghyr@hammerspace.com>
Link: https://patch.msgid.link/20260720141442.783935-4-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
lockd references "struct nfs_fh" but none of the other definitions
in linux/nfs.h. That header also pulls in cred.h, several sunrpc
headers, and uapi/linux/nfs.h, none of which lockd needs. A new
linux/nfs_fh.h provides "struct nfs_fh" and its helpers without the
rest of that surface.
Switch lockd's xdr.h to linux/nfs_fh.h, and drop the now-redundant
linux/nfs.h includes from svc.c and trace.h.
Link: https://patch.msgid.link/20260720141442.783935-3-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Plenty of spots around the kernel need the full definition of struct
nfs_fh but not the cred, sunrpc, and uapi dependencies that
linux/nfs.h pulls in along with it.
Relocate struct nfs_fh to its own header, and include that header in
linux/nfs.h so existing consumers keep building. Over time, consumers
can then replace
#include <linux/nfs.h>
with
#include <linux/nfs_fh.h>
While relocating the code, add kernel-doc comments for the FH
operations and convert nfs_compare_fh() to return bool.
Link: https://patch.msgid.link/20260720141442.783935-2-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd4_ff_proc_layoutget() and nfsd4_ff_encode_layoutget() translate
the file's owner and group with init_user_ns, but every other identity
nfsd places on the wire goes through nfsd_user_namespace(). When the
transport carries a credential from a non-initial user namespace, the
flex file layout reports host-global IDs. The client copies those IDs
into the AUTH_SYS credential it presents to the data server, and
svcauth_unix_accept() resolves that credential in the transport's
namespace, so data server I/O runs under an identity unrelated to the
file's owner.
Switching to the request's namespace introduces a second hazard.
from_kuid() returns (uid_t)-1 when the target namespace has no
mapping for the owner, and the IOMODE_READ arm adds one to that
result to derive an identity for which the data server denies
writes. The addition would wrap to zero, handing the client uid 0
instead of an identity distinct from the owner.
Translate both IDs in nfsd4_ff_proc_layoutget(), which has the
svc_rqst, and carry the wire values in struct pnfs_ff_layout.
from_kuid_munged() substitutes overflowuid for an unmapped owner and
thus never returns (uid_t)-1, so the increment cannot wrap to zero.
Fixes: 9b9960a0ca47 ("nfsd: Add a super simple flex file server")
Cc: stable@vger.kernel.org
Reported-by: sashiko-bot <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260720141442.783935-1-cel@kernel.org?part=3
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260720163724.810227-1-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd_v4client() is the last user in nfsd.h of XDR-defined item
references. Once this helper moves out, <linux/nfs.h> and the other
XDR-specific headers nfsd.h includes have nothing left to provide,
and can be dropped.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260717184112.507548-7-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd.h is included by nearly every NFSD translation unit, yet the
two blocks of NFSv4 lifecycle and control prototypes it carries are
referenced by only seven of them (out of over two dozen). The
remaining consumers, including the NFSv2 and NFSv3 ACL and XDR
paths, parse these declarations for no benefit.
The declarations cannot simply move into an NFSv4-only header such as
state.h: nfssvc.c, nfsctl.c, vfs.c, and export.c call the routines
unconditionally and rely on the CONFIG_NFSD_V4=n stubs, and pulling
the heavy state.h types into those lean translation units to obtain a
handful of prototypes would trade one form of coupling for a worse
one.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260717184112.507548-6-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
The COMPOUND encode-slack sizes and the state-management timeouts at
the tail of nfsd.h are evaluated only by NFSv4 code (nfs4state.c,
nfs4proc.c, and nfs4xdr.c). They nonetheless sit in nfsd.h, where
every NFSD translation unit, including the NFSv2 and NFSv3 paths
that have no use for them, has to parse them.
All three consumers already reach state.h through xdr4.h, so move
the block there. nfsd.h keeps the NFSv4 prototypes for now. Only the
pure constants move.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260717184112.507548-5-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Neither COMPOUND_SLACK_SPACE nor NFSD_COURTESY_CLIENT_TIMEOUT has
a remaining user. COMPOUND_SLACK_SPACE lost its last reference in
commit ea8d7720b274 ("nfsd4: remove redundant encode buffer size
checking"), which deleted the encode buffer-space check the macro
fed; the comment above it still describes that departed check.
NFSD_COURTESY_CLIENT_TIMEOUT is likewise unreferenced: the
courteous-server code expires clients through the laundromat's
reaper and conflict paths, never a fixed 24-hour timer.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260717184112.507548-4-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd.h is included by nearly every NFSD translation unit, so its
include of <linux/nfs4.h> reaches all of them, whether or not they
touch NFSv4. That include existed solely for the block of pre-xdr'ed
nfserr_* values at the end of the file: several of those values,
such as nfserr_delay and nfserr_admin_revoked, are built from
NFS4ERR_* constants defined in nfs4.h. The NFSD-internal error enum
that follows the block (NFSERR_EOF and friends) is anchored at an
impossible nfsstat4 value, thus it also needs nothing from nfs4.h.
But these codes are used by all NFS versions, so their new home must
be version-neutral.
Move the pre-xdr'ed value block and the internal error enum into a
new fs/nfsd/nfserr.h, which includes nfs4.h itself, and drop the
nfs4.h include from nfsd.h. Include nfserr.h directly from each
translation unit that references the pre-xdr'ed values or the
internal error codes, rather than carrying it in a widely-included
header.
A translation unit that includes nfsd.h without using the error
block no longer pulls in nfs4.h. The ones that reference the block
can still reach it through nfserr.h.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260717184112.507548-3-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
These static inline helpers use the nfserr_resource macro, which
pulls in the whole rack of NFS status codes. Move those helpers
into the only file that uses them, to get rid of the nfserr macro
dependency globally.
These helper were originally placed in xdr4.h because I thought
they would be utilized in the rest of the NFSv4 XDR code, but
XDR translation is eventually to be subsumed by xdrgen instead.
I'm not converting them now because that is much more churn than
this patch is.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260717184112.507548-2-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
All the code in nfsd4_create_file() that is VFS manipulation, with now
NFS-specific knowledge, has been localised. Now we split that out into
a separate function: do_lookup_open().
It is planned to provide a vfs_lookup_open() in vfs code which provides
this functionality. This will share more code with the syscall open
path, and make it easier to modify locking at the VFS level.
Signed-off-by: NeilBrown <neil@brown.name>
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260717093001.1972119-18-neilb@ownmail.net
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
A future patch will use nfsd_check_obj_isreg() in a context where the
protocol version is not easily available. So move the version check out
and put it at the end of do_open_lookup().
Also change to return errno error code and use nfserrno() to convert to
nfs error codes. Use -ELOOP for nfserr_symlink, which is an error
indication a problem with symlinks. -EFTYPE is a good match for
nfserr_wrong_type.
Signed-off-by: NeilBrown <neil@brown.name>
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260717093001.1972119-17-neilb@ownmail.net
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd4_create_file() needs write access to the mount for two purposes:
1/ to create/open the file.
2/ to set attributes on the newly created (or pre-existing) file.
Currently this is all handled by holding the write access across the
open and the setattr. A subsequent patch will necessarily change how
write access is gained for the open. So we reduce the range for the
first want_write, and add another one to cover setattr. If we failed to
get write access, it is only fatal if there were attrs to set.
We call nfsd_create_setattr() if at all possible, even when no attrs, as
it also calls commit_metadata and we need to be certain that the file
creation has been synced. If the mount became read-only since the
creation happened, we can safely assume that the sync happened as part
of that.
Signed-off-by: NeilBrown <neil@brown.name>
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260717093001.1972119-16-neilb@ownmail.net
Signed-off-by: Chuck Lever <cel@kernel.org>
|