| Age | Commit message (Collapse) | Author |
|
# Conflicts:
# drivers/gpu/drm/amd/amdkfd/kfd_migrate.c
# net/ceph/osd_client.c
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/modules/linux.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs.git
# Conflicts:
# fs/smb/server/smb2pdu.c
# fs/smb/server/vfs.c
# fs/smb/server/vfs.h
|
|
|
|
Commit 0ac903d1bfdc ("NFS: NFSERR_INVAL is not defined by NFSv2")
changed the status NFSD returns for an NFSv2 READ or WRITE of a
symlink or other non-regular file from NFSERR_INVAL to NFSERR_IO.
RFC 1094 defines neither an INVAL nor a SYMLINK status.
U-Boot's NFS client follows a symlink only when READ fails with
NFSERR_ISDIR or NFSERR_INVAL. Since that commit it cannot load a
boot image through a symlink exported by a Linux NFS server: the
load completes with zero bytes transferred.
Solaris returns NFSERR_ISDIR when the target of an NFSv2 READ or
WRITE is not a regular file. Return the same status from NFSD.
Other NFSv2 procedures continue to return NFSERR_IO.
Fixes: 0ac903d1bfdc ("NFS: NFSERR_INVAL is not defined by NFSv2")
Cc: stable@vger.kernel.org
Reported-by: Jörg Sommer <joerg@jo-so.de>
Closes: https://lore.kernel.org/linux-nfs/apE90ZRZB8IW_AiS@jo-so.de/
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260828231841.282003-1-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
client_has_openowners() walks clp->cl_openowners and reads so_stateids
without clp->cl_lock. nfs4_put_stateowner() unhashes that openowner under
cl_lock and then frees it, so a concurrent EXCHANGE_ID with mismatched
creds can use-after-free the nfs4_openowner.
BUG: KASAN: slab-use-after-free in client_has_state+0x10a/0x140
fs/nfsd/nfs4state.c:3718 client_has_openowners()
nfsd4_exchange_id
nfsd4_proc_compound
nfsd_dispatch
svc_process
Take clp->cl_lock while client_has_openowners() walks the openowner
list.
Fixes: 4eaea1342507 ("nfsd: improve client_has_state to check for unused openowners")
Cc: stable@vger.kernel.org
Reported-by: Xiang Mei (Microsoft) <xmei5@asu.edu>
Cc: AutonomousCodeSecurity@microsoft.com
Signed-off-by: Cen Zhang (Microsoft Security FORGE Labs) <blbllhy@gmail.com>
Link: https://patch.msgid.link/20260828041925.36758-1-blbllhy@gmail.com
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
fs/nfsd/trace.h includes include/trace/misc/nfs.h for
show_nfs4_seq4_status() and show_rca_mask(). That header pulls in
linux/nfs.h, so every NFSD translation unit that reads trace.h also
sees the NFS_OK, NFSERR_*, and file-type enumerators.
I'm about to switch fs/nfsd/nfserr.h to an xdrgen-generated header,
which defines enum nfsstat and enum ftype with those same names. Any
translation unit that includes both headers then fails to build with
enumerator redefinition errors.
To address this, stop including trace/misc/nfs.h in fs/nfsd/trace.h.
Link: https://patch.msgid.link/20260827185134.197322-1-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
fs/nfsd/nfserr.h builds NFSD's internal __be32 nfserr_* values
by wrapping the version-agnostic NFSERR_* status codes from
uapi/linux/nfs.h in cpu_to_be32(), which it reaches through its
include of linux/nfs.h.
A subsequent patch replaces that include with the xdrgen-generated
linux/sunrpc/xdrgen/nfs2.h. The generated header re-supplies
only the eighteen NFSv2 status codes; the NFSv3 and NFSv4 codes
that nfserr.h also uses (NFSERR_INVAL, NFSERR_JUKEBOX,
NFSERR_RESOURCE, and the rest) would be left undeclared.
Respell those definitions in terms of the version-specific
NFS3ERR_* and NFS4ERR_* codes from linux/nfs3.h and linux/nfs4.h,
and switch the one bare NFSERR_MOVED in nfs4xdr.c to NFS4ERR_MOVED.
The wire values are unchanged; NFSD no longer depends on the
version-agnostic NFSv3 and NFSv4 names.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260826194444.148243-2-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd4_sequence() claims a session slot by setting NFSD4_SLOT_INUSE
under nn->client_lock. A reply served from the slot's reply cache does
not claim it, and that distinction lived only in cstate->status, which
nfsd4_sequence() set to nfserr_replay_cache. Commit cc028a10a48c
("NFSD: Hoist status code encoding into XDR encoder functions") moved
nfsd4_proc_compound()'s cstate->status assignment below the out:
label, and the replay path's goto out was the one path that relied on
skipping it. The test in nfsd4_sequence_done() therefore no longer
identifies a replay, and every replay now stores its reply and clears
NFSD4_SLOT_INUSE as though it owned the slot.
NFSD4_SLOT_INUSE is the interlock: check_slot_seqid() rejects every
sequence id for a slot that is in use, before replay_matches_cache()
runs. A replay never sets it, so two replays can be in flight on the
same slot at once. One can read slot->sl_cred under nn->client_lock
while the other's completion frees and rebuilds it from the XDR
encoder without that lock. Two completions can also collide with each
other and release the same group_info twice. free_svc_cred() leaves
cr_uid, cr_gid and cr_flavor intact, so the reader passes every
earlier test in same_creds() and dereferences a NULL cr_group_info.
From a 6.12.91 production server:
BUG: kernel NULL pointer dereference, address: 0000000000000004
CPU: 63 UID: 0 PID: 39015 Comm: nfsd
RIP: 0010:same_creds+0x38/0xa0 [nfsd]
RDX: 0000000000000000
Call Trace:
nfsd4_sequence+0x6a8/0x910 [nfsd]
nfsd4_proc_compound+0x345/0x670 [nfsd]
nfsd_dispatch+0x100/0x220 [nfsd]
svc_process_common+0x311/0x700 [sunrpc]
svc_process+0x131/0x1c0 [sunrpc]
svc_recv+0x7ef/0x9c0 [sunrpc]
nfsd+0xa3/0x100 [nfsd]
Kernel panic - not syncing: Fatal exception
Since v6.14 a live session's slot table can shrink, and the unowned
store becomes a use-after-free write. nfsd4_sequence() defers the
shrink while a slot is in use, but it tests NFSD4_SLOT_INUSE, which a
replay does not set, so free_session_slots() can kfree() the slot the
replay still holds in cstate->slot.
Record the claim in cstate->slot_owned when nfsd4_sequence() accepts a
request and test that in nfsd4_sequence_done(), so only the request
that claimed the slot updates its cached reply and clears
NFSD4_SLOT_INUSE. svc_generic_init_request() zeroes the compound
response before each request, so the flag starts clear.
Fixes: cc028a10a48c ("NFSD: Hoist status code encoding into XDR encoder functions")
Cc: stable@vger.kernel.org
Assisted-by: Claude:claude-opus-5
Signed-off-by: Ameer Hamza <ameer.hamza@truenas.com>
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260824173641.3274260-1-ameer.hamza@truenas.com
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd4_decode_fattr4() allocates POSIX ACLs while decoding OP_OPEN,
OP_CREATE, and OP_SETATTR, leaving the only reference to these ACLs
in the operation's argument structure. Executing the operation hands
that reference to struct nfsd_attrs, which drops it. However, if the
operation is decoded but never executes, those ACLs are leaked. A
client can repeat an aborting compound to force the server to leak
memory.
Give struct nfsd_attrs its own reference with posix_acl_dup() so
nfsd_attrs_free() still balances the reference the operation took.
Release the ACLs when the compound completes. Have the decoder record
each ACL on the compound's temporary allocation chain, and give each
chained item an optional release callback. The chain holds a reference
for the life of the compound, so OP_OPEN no longer needs an op_release
method.
Fixes: 5fc51dfc2eb1 ("NFSD: Add support for XDR decoding POSIX draft ACLs")
Cc: stable+noautosel@kernel.org # experimental, disabled by default
Reported-by: Prabhakar Pujeri <prabhakar.pujeri@dell.com>
Closes: https://lore.kernel.org/linux-nfs/20260823113255.3417-1-prabhakar.pujeri@dell.com/
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260824-nfsd-posix-acl-ownership-v1-2-090fffc608ec@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd_nl_rpc_status_get_dumpit() loads args->opcnt and
args->ops separately, then walks ops[] before rechecking
rq_status_counter. A COMPOUND that completes between the two loads
runs nfsd4_release_compoundargs(), which zeroes opcnt and points ops
back at the eight-entry inline array. A dump that already sampled an
opcnt of 200 clamps it to the sixteen slots in rq_opnum, then indexes
iops[0..15]. iops is the last member of struct nfsd4_compoundargs and
rq_argp is allocated at exactly that size, so the walk runs off the
end of the allocation. The trailing recheck discards the sampled data,
but the read has already happened. NFSD_CMD_RPC_STATUS_GET carries
no GENL_ADMIN_PERM, so an unprivileged local user can repeat the dump
against a busy server until it lands in the window.
Sample opcnt and ops into locals, then finish the counter recheck
before dereferencing ops. An unchanged counter means both came from the
same COMPOUND, where opcnt cannot exceed what ops holds.
Fixes: bd9d6a3efa97 ("NFSD: add rpc_status netlink support")
Cc: stable@vger.kernel.org
Reported-by: Prabhakar Pujeri <prabhakar.pujeri@dell.com>
Closes: https://lore.kernel.org/linux-nfs/20260823113255.3417-1-prabhakar.pujeri@dell.com/
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260824-nfsd-posix-acl-ownership-v1-1-090fffc608ec@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Make the header guards less ambiguous about their provenance.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260818140035.12740-3-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
struct nfsd4_copy embeds a struct nfs_fh, and nlm_fopen() reads the
size and data fields of one. Neither fs/nfsd/xdr4.h nor
fs/nfsd/lockd.c includes the header that defines the type; both
reach it by way of nfsd.h, which pulls in <linux/nfs.h>.
Add the direct include to both files, so nfsd.h can later drop the
<linux/nfs.h> it carries for no use of its own.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260818140035.12740-2-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
fs/nfsd/nfsd.h is included throughout the server, yet it uses no
NFSv3 protocol definition of its own. The <linux/nfs3.h> include
there served only to make those definitions reach the few source
files that need them, by way of nfsd.h itself or the xdr.h chain
that pulls it in.
Give each consumer its own include and drop the one in nfsd.h, so
the header no longer carries a dependency unrelated to its
contents. nfsfh.c, nfsctl.c, nfs3xdr.c, nfs3proc.c, and nfs2acl.c
reference NFS3_* definitions directly; add <linux/nfs3.h> to each.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260818140035.12740-1-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
deleg_reaper() asks each eligible client to return one delegation
whenever it runs, whether or not anything needs the memory. A
delegation returned before it is needed costs the client an OPEN
when it next touches the file. Nothing sizes the request either.
The delegation scan callback discards nr_to_scan, which is reclaim's
statement of how many objects it wants back.
Record each delegation scan request in nfsd_deleg_backlog and pass
the accumulated total to deleg_reaper(). Handing that total to
every client would ask for it once per client, so scale it by each
client's share of the delegations this sweep can reach. The count
callback reports what is left after the outstanding requests, so
concurrent reclaimers do not each ask for the same delegations.
cl_ra_time keeps the next sweep from returning to the clients this
one reached.
Nothing is recalled until a scan arrives. nfs4_laundromat() is the
exception. It has no scan request to pass, so it computes what must
go for num_delegations to fall below max_delegations, and passes
only this namespace's share.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-8-3b2cffce701e@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Currently, neither of the scan callback functions records anything
before returning SHRINK_STOP, so the size of each scan request is
discarded. That size is the only real measure NFSD gets of reclaim
pressure. Both count callbacks report their population whether or
not the reaper is already queued to reclaim it, so reclaim asks
again for work that is pending.
Accumulate each courtesy scan request in nfsd_shrink_backlog and
subtract the backlog from what that count callback reports. The
worker retires the backlog once courtesy_client_reaper() has run.
That reaper expires the clients synchronously, so the discount
covers exactly the interval the work is pending.
Delegations need a different bound. This is because deleg_reaper()
only sends CB_RECALL_ANY and does not track how many delegations
were actually returned by the targeted client.
Report the delegations only once NFSD_RECALL_ANY_COOLDOWN_SECS have
passed since the last sweep. deleg_reaper() skips any client it
recalled from within that window, so an earlier scan request cannot
produce another recall.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-7-3b2cffce701e@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Since commit 44df6f439a17 ("NFSD: add delegation reaper to react to
low memory condition"), nfsd_client_shrinker has managed two
unrelated populations of objects.
One population is courtesy clients. Shrinking that population can
be done synchronously and without risk of deadlock. The shrinker
callback could return a precise count of the number of objects
that were released.
The other population is delegations. Shrinking that population
requires sending a CB_RECALL_ANY; clients are not obligated to
return any delegation. The shrinker callback is structurally
unable to report progress.
What's more, the single shrinker callback falls back to
delegation reaping only when there are no courtesy clients left to
reclaim. A single courtesy client is enough to keep a namespace's
delegations out of the count it reports.
To begin to resolve these issues, refactor the existing state
shrinker into two: one for courtesy clients and one for reaping
delegations. Each manages the size of its own population, and the
shrinker names become namespace-specific.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-6-3b2cffce701e@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
The state shrinker is allocated per network namespace, but
nfsd4_state_shrinker_count() reports num_delegations, which counts
the delegations held by the whole host. Every namespace therefore
reports every delegation on the server. Reclaim sees the
population multiplied by the number of namespaces running NFSD. A
namespace holding no delegations of its own still reports a
nonzero count and queues its reaper, which then finds nothing to
recall.
Count the delegations in each namespace and report that instead.
num_delegations stays for the admission check in
__alloc_init_deleg() and the ceiling check in nfs4_laundromat().
Both compare against max_delegations, which is sized from host
memory and so remains a host-wide limit.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-5-3b2cffce701e@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
deleg_reaper() sets craa_objects_to_keep to zero on every
CB_RECALL_ANY. RFC 8881 Section 20.6.3 defines that field as the
number of objects the client may keep, leaving the client to choose
which of the excess to return, because the server cannot read lack
of recent use as lack of usefulness. Zero asks for every delegation
the client holds, including the ones backing files an application
still has open.
There is also no reason NFSD has to reclaim the entire delegation
working set on the first sign of memory pressure.
Derive the keep count from cl_deleg_count so that each callback
asks for one delegation. Both the shrinker and the laundromat re-arm
while their condition lasts, so a client with more to give is asked
again on the next pass.
The Linux client ignores craa_objects_to_keep and returns unused
delegations selected from the type mask alone, so the count changes
nothing for it.
Fixes: 44df6f439a17 ("NFSD: add delegation reaper to react to low memory condition")
Cc: stable@vger.kernel.org
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-4-3b2cffce701e@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
RFC 8881 Section 20.6.3 distinguishes an NFSv4.1 server
implementation that shares one pool among all classes of recallable
objects from one that keeps separate pools per class. NFSD falls
in the former category.
The CB_RECALL_ANY operation's craa_type_mask argument names the
types of objects in the recallable resource pool, but NFSD's
implementation does not name directory delegations, even though
they are allocated through __alloc_init_deleg(), they are counted
against the max_delegations budget, and the state shrinker reclaims
them.
Add RCA4_TYPE_MASK_DIR_DLG to craa_type_mask so clients that
implement directory delegations consider them when choosing which
delegations to return.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-3-3b2cffce701e@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
struct nfs4_client records the delegations it holds on cl_delegations
but keeps no count of them. deleg_reaper() walks nn->client_lru under
nn->client_lock, but cl_delegations is serialized by nn->deleg_lock,
which nests outside nn->client_lock. A caller there cannot take
nn->deleg_lock to count the list. The cost tells against the walk as
well: an O(n) count per client, on a pass that already visits every
client.
Add cl_deleg_count, maintained at the two sites that mutate
cl_delegations. Both hold nn->deleg_lock, so the counter is already
serialized against itself and needs no atomic of its own. The decrement
sits below the delegation_hashed() test, next to the list_del_init it
pairs with, so it runs only when the delegation really leaves the list.
A reader that holds only nn->client_lock is not synchronized against
either update site, so it can see a count that does not match the
list. Such a reader marks the access with data_race() and may not
depend on the value for correctness.
No functional change.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-2-3b2cffce701e@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
deleg_reaper() sends CB_RECALL_ANY to every ACTIVE client holding
delegations, but CB_RECALL_ANY is an NFSv4.1 operation. An NFSv4.0
client's callback service accepts only CB_GETATTR and CB_RECALL, so it
replies OP_ILLEGAL. The decoder maps the unexpected opnum to -EIO, and
nfsd4_cb_done() marks the client's callback channel down.
Nothing brings the channel back. nfsd4_run_cb_work() sets NFSD4_CB_UP
only for a minor version above zero, and the only nfsd4_probe_callback()
call site an NFSv4.0 client reaches is nfsd4_setclientid_confirm(). One
visit from the reaper therefore leaves the channel marked down until the
client re-establishes its clientid. RENEW then returns
NFS4ERR_CB_PATH_DOWN for as long as the client holds delegations.
nfsd4_cb_channel_good() stops returning true, so the client is granted
no further delegations.
Skip clients at minor version zero.
Fixes: 44df6f439a17 ("NFSD: add delegation reaper to react to low memory condition")
Cc: stable@vger.kernel.org
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-1-3b2cffce701e@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd4_encode_operation() leaves op->status alone when the reply
buffer has no room for the operation's opcode and status word.
nfsd4_proc_compound() reads the unchanged nfs_ok as success and
goes on to the next operation, so the reply counts an operation
whose result was never encoded.
Report the failure through nfsd4_check_resp_size(), which the rest
of the function already uses. It returns NFS4ERR_REP_TOO_BIG, or
NFS4ERR_REP_TOO_BIG_TO_CACHE on a session, and the COMPOUND ends
at that operation.
Two paths narrow the reply buffer: nfsd4_sequence(), which rejects
a SEQUENCE result that does not fit, and nfsd4_encode_splice_read(),
which can leave a single XDR word in the head page. Whether a
COMPOUND reaches that boundary is unproven, so this is a guard
rather than a fix.
Link: https://patch.msgid.link/20260817-jean-v1-2-9e356596ab85@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd4_sequence() narrows the reply buffer to the session's cached
reply limit before it accepts the slot seqid. A client may negotiate
ca_maxresponsesize_cached down to NFSD_MIN_HDR_SEQ_SZ, and
nfsd4_alloc_slot() then gives every slot a zero-length sl_data[]. A
COMPOUND tag can fill that narrowed buffer until it holds the
SEQUENCE opcode but not the status word that follows.
nfsd4_encode_operation() returns without running
nfsd4_encode_sequence(), so cstate.data_offset stays zero. It leaves
op->status at nfs_ok as well, so the COMPOUND is treated as having
succeeded.
nfsd4_store_cache_entry() declines to cache a lone SEQUENCE that
returned an error. That test reads the status the operation
reported, so it passes here. The copy starts at offset zero and
takes the whole reply, RPC and COMPOUND headers included, into the
zero-length sl_data[]. The COMPOUND tag is copied along with it, so
the client picks most of the bytes written past the end of the slot:
BUG: KASAN: slab-out-of-bounds in read_bytes_from_xdr_buf+0x1bc/0x390
Write of size 80 at addr ffff888003a549cd by task kunit_try_catch/24
__asan_memcpy+0x38/0x60
read_bytes_from_xdr_buf+0x1bc/0x390
nfsd4_sequence_done+0x5b0/0x810
nfs4svc_encode_compoundres+0x1bf/0x240
Check that the fixed-size SEQUENCE result, plus room for a following
operation's error status, fits the negotiated limit before narrowing
the buffer and consuming the slot seqid. The slot and its reply
cache are left unchanged, as RFC 8881 Section 2.10.6.1.2 requires of
an error returned from SEQUENCE.
Fixes: 47ee52986472 ("nfsd4: adjust buflen to session channel limit")
Cc: stable@vger.kernel.org
Assisted-by: Codex:gpt-5
Signed-off-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr>
Tested-by: Mayank Jangid (OpenSec Intelligence) <mayank.jangid.moon@gmail.com>
Link: https://patch.msgid.link/20260817-jean-v1-1-9e356596ab85@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
The reply to a pool_threads read is the list of per-pool thread counts,
formatted into a buffer of SIMPLE_TRANSACTION_LIMIT bytes. snprintf()
truncates its last write and strlen() measures only what fit, so a reply
too long for that buffer ends mid-number with no terminating newline. A
pool running 4096 threads is reported as 40. Nothing marks the reply as
incomplete, so an administrator reads a plausible but wrong count.
Take snprintf()'s return value, which reports the truncation strlen()
cannot see, and fail the read with -ENAMETOOLONG when the list does
not fit. That is the errno svc_one_xprt_name() already returns for the
same condition.
Suggested-by: David Laight <david.laight.linux@gmail.com>
Fixes: eed2965af1ba ("[PATCH] knfsd: allow admin to set nthreads per node")
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260812193349.13347-1-david.laight.linux@gmail.com
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd4_cb_notify_prepare() reserved a word from the encoding xdr stream
for each notify4's host-order notify_mask, storing the pointer in
ncn_nf[].notify_mask.element. That element must survive until the RPC
encode re-reads ncn_nf, so the host-endian word lived inside the XDR
staging buffer for the whole callback lifetime - the same fragile
pattern as the attrmask, and one that keeps host-order bytes in a
buffer meant to hold big-endian XDR.
ncn_nf is a bounded per-delegation array reused across every CB_NOTIFY,
so give it a parallel ncn_masks array with the same lifetime:
- allocate/free ncn_masks alongside ncn_nf in alloc_init_dir_deleg() /
nfs4_free_dir_deleg()
- point notify_mask.element at &ncn_masks[i] (events) and
&ncn_masks[count] (dir attr change)
The mask backing is now pre-allocated, so the per-word NULL checks in
prepare go away. The staging stream holds only encoded XDR.
Assisted-by: LLM
Signed-off-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260812-dir-deleg-v1-2-411faa713068@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd4_setup_notify_entry4() stole 3 words from the xdr stream via
xdr_reserve_space() to hold the host-order bmval[3] attrmask that
nfsd4_encode_attr_vals() consumes and ne_attrs.attrmask.element points
at. Stashing host-endian scratch in an XDR stream buffer is fragile:
the buffer layout is not guaranteed by sunrpc, and it blocks moving the
encoder to pages or xdrgen.
The attrmask only needs to live until the enclosing encode call
serializes the notify_entry4, so hand it caller-provided stack storage
instead. The two callers keep the words on the stack: up to three
concurrent entries for a rename in nfsd4_encode_notify_event(), one in
nfsd4_encode_dir_attr_change().
No wire change: the reserved words were never emitted; attr_vals.data/len
are still captured relative to xdr->p.
Assisted-by: LLM
Signed-off-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260812-dir-deleg-v1-1-411faa713068@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Clean up: Nothing in fs/nfsd/vfs.c references an NFSv3 XDR definition.
The preceding patch moved the last NFSv4 reference, a struct
nfsd4_compoundres dereference that recovered a pair of file handles for
a tracepoint, into fs/nfsd/nfs4proc.c. Remove both includes so that the
VFS layer no longer names on-the-wire types.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260812142436.35042-4-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd4_clone_file_range() lives in fs/nfsd/vfs.c but reaches into the
NFSv4 compound reply buffer: nfsd4_get_cstate() casts rq_resp to a
struct nfsd4_compoundres to recover the saved and current file
handles a tracepoint wants. That is the only reference to the NFSv4
XDR definitions left in vfs.c, and it puts knowledge of the compound
reply layout in the VFS layer.
Refactor nfsd4_clone_file_range() to remove NFSv4-specific
componentry from fs/nfsd/vfs.c.
Splitting nfsd_clone_file_range() and nfsd_clone_sync_range() lets
nfsd4_clone() distinguish a clone failure from a sync failure, which
it has to do because only the latter invalidates the write verifier.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260812142436.35042-3-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
A subsequent patch moves the NFSv4-specific portions of
nfsd4_clone_file_range() out of fs/nfsd/vfs.c and into its caller in
fs/nfsd/nfs4proc.c. One of those portions resets the write verifier
when the post-clone sync fails.
netns.h already exposes nfsd_reset_write_verifier(), but that is the
unconditional reset. commit_reset_write_verifier() wraps it with the
policy that decides which errors warrant a reset: -EAGAIN and -ESTALE
do not indicate a problem with durable storage, so they leave the
verifier alone. A caller outside vfs.c has to apply the same policy,
so make the wrapper visible rather than duplicate its switch.
Rename it to nfsd_maybe_reset_write_verifier() on the way out.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260812142436.35042-2-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd_access() owns three static tables that map on-the-wire ACCESS
bits to NFSD_MAY flags, and every NFS version shares them. That puts
protocol-version specifics in the version-agnostic VFS layer. The
NFSv4.2 extended-attribute bits are wedged into the NFSv3 tables
under CONFIG_NFSD_V4.
Give each version its own tables in its proc code and pass the
matching set into nfsd_access() as a new argument.
Splitting the tables also stops the NFSv2-ACL and NFSv3 ACCESS paths
from answering for the NFSv4.2 extended-attribute bits. Both use the
NFSv4-augmented table whenever CONFIG_NFSD_V4 is set, so an NFSv3
request that sets an xattr bit has it echoed back in the reply even
though those bits are undefined for v3.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260810141900.33846-2-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Clean up: Remove an NFSv3 constant (NFS3_ACCESS_FULL) used inside an
NFSv4 code path.
After this patch is applied, the nfs3.h header is no longer an
implicit dependency of nfs4proc.c for this value. The definition
itself stays in nfs3.h to avoid a kernel-user space API regression.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260810141900.33846-1-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
The usermode helper declarations were previously provided by linux/kmod.h
but commit c1f3fa2a4fde ("kmod: split off umh headers into its own file")
moved them to linux/umh.h in 2017.
Add explicit includes of linux/umh.h to files that use usermode helpers and
remove linux/kmod.h where it is no longer needed.
Acked-by: Alex Elder <elder@riscstar.com> # for greybus
Signed-off-by: Petr Pavlu <petr.pavlu@suse.com>
|
|
When CB_RECALL races ahead of the reply that granted the delegation,
the client has not yet recorded the delegation stateid and responds
NFS4ERR_BADHANDLE or NFS4ERR_BAD_STATEID. The slot that carried the
grant has not retired at that point, so NFSD cannot read the rejection
as proof that the client never held the delegation. It retries the
recall and, once the retries lapse, revokes a delegation the client is
by then able to return.
Remove the ambiguity with the referring call mechanism of RFC 8881
Section 2.10.6.3: until the slot that carried the grant retires, name
that request as a referring call in the CB_SEQUENCE of each recall. A
client that finds it still outstanding may respond NFS4ERR_DELAY, and
the recall is retried until the client has processed the grant.
A recall reuses one callback context across its retries, and ->prepare
does not run on every send. The granting request does not change, so a
send that inherits the previous list sends the right one. Retirement
of the granting slot drops the list, and nfs4_free_deleg() releases
what is left.
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-9-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
A client that answers CB_RECALL with NFS4ERR_BADHANDLE or
NFS4ERR_BAD_STATEID has no record of the delegation, so the
FREE_STATEID that clears it from cl_revoked never arrives. Every later
SEQUENCE reply carries SEQ4_STATUS_RECALLABLE_STATE_REVOKED, and the
client loops issuing TEST_STATEID.
Destroy such a delegation when it is reaped rather than revoking it
onto cl_revoked. RFC 8881 Section 20.2.4 completes the recall at the
reply when its status is neither NFS4_OK nor NFS4ERR_DELAY, so a
rejected recall leaves nothing to revoke. An administrative revoke
keeps that path, since NFS4ERR_ADMIN_REVOKED reports it. A destroyed
stateid returns NFS4ERR_BAD_STATEID instead of NFS4ERR_DELEG_REVOKED. A
client that rejects the recall but still holds the delegation gets no
notice that its state was revoked.
CB_RECALL can outrun the reply that granted the delegation, so honor
a rejection only once the client has seen that grant. Per RFC 8881
Section 2.10.6.3, retirement of the slot that carried the grant is that
proof; retry until then, and revoke when the retries lapse.
Fixes: 3bd64a5ba171 ("nfsd4: implement SEQ4_STATUS_RECALLABLE_STATE_REVOKED")
Cc: stable@vger.kernel.org # 6.14.x
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-8-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
The comment above the SC_STATUS_ flags states that nn->deleg_lock
protects sc_status for delegation stateids, but only the transitions
made while a delegation is hashed are taken under that lock. This
comment was accurate until commit c88c150a467f ("nfsd: fix possible
badness in FREE_STATEID") set SC_STATUS_CLOSED under ->cl_lock.
Commit 8dd91e8d31fe ("nfsd: fix race between laundromat and
free_stateid") added the other two sites.
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-7-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
NFS4_enc_cb_recall_any_sz counts the objects-to-keep field, the
bitmap array length, and the bitmap word. encode_cb_recallany4args()
emits an opcode ahead of all three, so the macro falls one XDR word
short.
This macro sizes p_arglen and nothing else. rq_callsize pads that
with two credential slacks, so the shortfall has never reached
the send buffer. No backport is needed.
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-6-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
NFS4_enc_cb_notify_lock_sz counts the lock owner and the file
handle. nfs4_xdr_enc_cb_notify_lock() emits an opcode ahead of both,
so the macro falls one XDR word short.
This macro sizes p_arglen and nothing else. rq_callsize pads that
with two credential slacks, so the shortfall has never reached
the send buffer. No backport is needed.
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-5-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
NFS4_enc_cb_offload_sz counts the file handle, the stateid, and the
offload information. encode_cb_offload4args() emits an opcode ahead
of all three, so the macro falls one XDR word short.
This macro sizes p_arglen and nothing else. rq_callsize pads that
with two credential slacks, so the shortfall has never reached
the send buffer. No backport is needed.
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-4-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
NFS4_enc_cb_layout_sz counts the opcode, the three scalar fields,
the file handle, and the offset and length hypers.
encode_cb_layout4args() also emits the layoutrecall4 discriminator
and the recall stateid, so the macro falls five XDR words short.
This macro sizes p_arglen and nothing else. rq_callsize pads that
with two credential slacks, so the shortfall has never reached
the send buffer. No backport is needed.
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-3-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
NFS4_enc_cb_recall_sz counts the CB_RECALL opcode, the stateid, and the
file handle. encode_cb_recall4args() also emits the truncate field, so
the macro falls one XDR word short.
NFSD_CB_MAX_REQ_SZ derives from this macro, so the minimum
ca_maxrequestsize a client must advertise rises by four bytes.
The field has been unbudgeted since the macro was written. Neither
consumer of the macro justifies a backport. rq_callsize covers
p_arglen with two credential slacks. The only client affected is one
whose ca_maxrequestsize falls inside those four bytes.
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-2-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
cb_sequence_enc_sz counts the session ID, the four scalar fields, and
one referring call list. encode_cb_sequence4args() also emits the
CB_SEQUENCE opcode and the csa_referring_call_lists array count, so
the macro falls two XDR words short. Every NFS4_enc_cb_*_sz built on
it is short by the same two words.
NFSD_CB_MAX_REQ_SZ derives from NFS4_enc_cb_recall_sz, so the two
missing CB_SEQUENCE words shrink the ca_maxrequestsize that
check_backchannel_attrs() accepts by eight bytes. Count both words.
The minimum a client must advertise rises by those eight bytes.
The short count cannot overrun the send buffer. The macro sizes
p_arglen, and rq_callsize adds two credential slacks on top of that.
The only client affected is one whose ca_maxrequestsize falls inside
those eight bytes. No backport is needed.
Link: https://patch.msgid.link/20260802-nfsd-deleg-destroy-badhandle-v1-1-323aa7196055@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
LOCALIO serves NFS clients of every version through one entry point,
so no protocol version is associated with such a request.
nfsd_set_fh_dentry() selects version-specific behavior anyway: its
switch keys off fh_maxsize, and nfsd_open_local_fh() passes
NFS4_FHSIZE because that is the size of the buffer it copies into, so
LOCALIO lands in the NFSv4 arm. fh_getattr() keys off fh_maxsize too
and does run on a LOCALIO open, adding STATX_BTIME and
STATX_CHANGE_COOKIE to the mask it requests: work on filesystems that
compute them for a caller that never reads them.
nfsd_open_local_fh() only verifies a handle it received, so it has no
maximum size to state. Pass NFSD_FHSIZE_UNSPEC as nlm_fopen() already
does, which selects the switch arm that applies no version-specific
behavior, and state the bound on the copy out of struct nfs_fh as
NFS_MAXFHSIZE.
Suggested-by: NeilBrown <neil@brown.name>
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Mike Snitzer <snitzer@kernel.org>
Link: https://patch.msgid.link/20260728165911.462534-6-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd_set_fh_dentry() selects behavior specific to an NFS protocol
version by matching fh_maxsize against NFS_FHSIZE, NFS3_FHSIZE, or
NFS4_FHSIZE. A filehandle that reaches NFSD outside an NFS request has
no such version. nlm_fopen() opts out of the switch by passing a bare
0, which matches no arm, and the literal says nothing about why, so an
adjacent comment has to carry it.
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Mike Snitzer <snitzer@kernel.org>
Link: https://patch.msgid.link/20260728165911.462534-5-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
As a prerequisite to converting NFSD to use xdrgen more broadly,
NFSD source files should not depend on NFS client headers.
fs/nfsd/localio.c is server-side LOCALIO code, yet it pulled in
three of them: <linux/nfs_fs.h>, the client inode header (struct
nfs_inode, NFS_I(), writeback helpers), which server code never
uses; <linux/nfs_xdr.h>, whose only referenced symbol is
decode_opaque_fixed(), a static inline that exists to remap the
error return to -EIO for client call sites; and the catch-all
<linux/nfs.h>.
Convert the UUID decoder to call the canonical SUNRPC primitive
xdr_stream_decode_opaque_fixed() directly. It is shared by client
and server, performs the identical bounds check, and is already
reachable through <linux/sunrpc/clnt.h>. With the wrapper gone,
localio.c references no symbol from <linux/nfs_xdr.h>, and with
that header gone, none of the NFSv3 definitions its structs embed
are needed here.
Drop all three client includes and add what the file actually
uses: struct nfs_fh comes from <linux/nfs_fh.h>, included directly
rather than through nfslocalio.h's conditional re-export, and
NFS4_FHSIZE from <linux/nfs4.h>. enum nfs_stat and
nfs_stat_to_errno continue to come from the already-included
<linux/nfs_common.h>.
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Mike Snitzer <snitzer@kernel.org>
Link: https://patch.msgid.link/20260728165911.462534-4-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
client_info_show() renders /proc/fs/nfsd/clients/<id>/info and walks
clp->cl_sessions under clp->cl_lock to print each session's slot
counts. free_client() tears down the same list without taking
cl_lock, and is the only unlocked mutator of cl_sessions. A reader
can observe a client mid-teardown because get_nfsdfs_clp() pins the
nfs4_client but not its sessions: free_client() frees every session
before calling nfsd_client_rmdir(), so an in-flight seq_file reader
can follow a list_del()'d node whose ->next now holds LIST_POISON1
and take a general protection fault:
Oops: general protection fault, probably for non-canonical address
0xdead00000000014c
CPU: 1 UID: 0 PID: 132488 Comm: cat
RIP: 0010:client_info_show+0x2bf/0x3d0
RAX: dead000000000100
Call Trace:
seq_read_iter+0x12a/0x4b0
seq_read+0xf1/0x130
vfs_read+0xbf/0x350
ksys_read+0x6f/0xf0
do_syscall_64+0x8b/0xcb0
entry_SYSCALL_64_after_hwframe+0x76/0x7e
Kernel panic - not syncing: Fatal exception
Detach cl_sessions onto a local reaplist under cl_lock, then free the
sessions after dropping the lock. Removing entries from cl_sessions
under cl_lock matches unhash_session(), and the detach-then-reap shape
matches how __destroy_client() reaps cl_delegations. The sessions
cannot be freed while cl_lock is held, since free_session() calls
nfsd4_del_conns(), which re-acquires it.
Reported-by: Nicholas Wolff <nicholas.wolff@truenas.com>
Fixes: 601c8cb349c2 ("nfsd: add session slot count to /proc/fs/nfsd/clients/*/info")
Cc: stable@vger.kernel.org
Signed-off-by: Ameer Hamza <ameer.hamza@truenas.com>
Link: https://patch.msgid.link/20260726124658.1715711-1-ameer.hamza@truenas.com
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
The current nfsd_write() API is not NFS version-agnostic, as it
relies on callers to pass an NFSv3 stable_how value to determine
the persistence of the requested WRITE. NFSv2 does not use a
stable-how value on the wire, and NFSv4 has its own stable_how4
(though stable_how and stable_how4 happen to share the same
numeric values).
To remove the dependence on NFSv3-specific XDR values from NFSD's
generic VFS APIs, replace nfsd_write()'s stable argument with an
argument that passes a set of IOCB flags instead of an XDR-defined
value.
The NFSv4 WRITE and COPY paths had been borrowing the NFSv3
stable_how constants for their own on-the-wire stable values,
relying on the numeric coincidence noted above. Convert those
sites to the stable_how4 enumerators so the v4 code expresses
its own protocol's values directly, with no change in behavior.
While here, bound-check the decoded NFSv3 WRITE stable value, as
the NFSv4 WRITE decoder already does, and make the nfsd3_writeargs
stable field unsigned to suit.
The larger benefit is one less NFSv4 dependency on nfs3.h.
Link: https://patch.msgid.link/20260723182043.990391-3-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
The nfs_ssc.h header contains both client- and server-side data
structures, which means each of those implementations has to pull in
headers from the other.
Create a linux/nfsd_ssc.h for the server side APIs which no longer
includes uapi/linux/nfs.h either directly or indirectly. Because
nfsd_ssc.h drops the transitive include of the NFS client headers,
fs/nfsd/nfs4proc.c now includes <linux/pagemap.h> directly for
filemap_check_wb_err().
struct nfsd4_ssc_umount_item is private to nfsd. Move it into
fs/nfsd/xdr4.h alongside its only consumers rather than into the
exported nfsd_ssc.h.
As an added clean-up, add missing header guard macros and the
struct file and struct vfsmount forward declarations the server
prototypes need.
Cc: Olga Kornievskaia <okorniev@redhat.com>
Cc: Dai Ngo <dai.ngo@oracle.com>
Link: https://patch.msgid.link/20260721162306.894558-5-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Refactor: The infrastructure and details for calling the client's
ssc_open method can be hidden in nfs_ssc.c. This reduces the SSC
footprint in fs/nfsd/nfs4proc.c, a step toward removing that
file's dependency on <linux/nfs_fs.h>, which indirectly includes
<uapi/linux/nfs.h>.
The open and close functions are named "nfsd42_" since they are
meant to be invoked only by NFSD.
Cc: Olga Kornievskaia <okorniev@redhat.com>
Cc: Dai Ngo <dai.ngo@oracle.com>
Link: https://patch.msgid.link/20260721162306.894558-3-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Clean up: Commit 75333d48f922 ("NFSD: fix use-after-free in
__nfs42_ssc_open()") addressed a use-after-free bug by removing the
nfsd4_interssc_disconnect() function. Post-copy clean-up was then
delegated to NFSD's laundromat.
Since that commit, the nfs_do_sb_deactive() wrapper function and the
entire nfs_ssc_client_ops infrastructure no longer have any
consumers. This includes nfs_do_sb_deactive(), struct
nfs_ssc_client_ops, nfs_ssc_register(), nfs_ssc_unregister(), and
related registrations in the NFS client.
Cc: Olga Kornievskaia <okorniev@redhat.com>
Cc: Dai Ngo <dai.ngo@oracle.com>
Link: https://patch.msgid.link/20260721162306.894558-2-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|