| Age | Commit message (Collapse) | Author |
|
https://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs.git
|
|
|
|
|
|
https://github.com/Paragon-Software-Group/linux-ntfs3.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/ntfs.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/smb.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/gfs2/linux-gfs2.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/mszeredi/fuse.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/jaegeuk/f2fs.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/jack/linux-fs.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/exfat.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/teigland/linux-dlm.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/leitao/linux.git
|
|
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs.git
|
|
ni_read_frame() marks every frame page uptodate at the 'out:' label
regardless of the return value. The pages come from ntfs_lock_new_page()
and are not zeroed, and several error paths reach 'out:' before anything
is written to them (e.g. a failed decompress_lznt()/decompress_lzx_xpress()
on a corrupted chunk, or an allocation failure).
A page marked uptodate is served directly from the page cache, so a later
read() of the file returns the uninitialized page contents to userspace.
On a crafted compressed image this leaks kernel memory, including pointers.
Zero each page on the error path before marking it uptodate. The success
path is unchanged.
Fixes: 4342306f0f0d ("fs/ntfs3: Add file operations and implementation")
Reported-by: Xiang Mei <xmei5@asu.edu>
Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Signed-off-by: Konstantin Komarov <almaz.alexandrovich@paragon-software.com>
|
|
ntfs_setxattr() updates ctime and marks the inode dirty even when the
xattr operation fails.
Do that only on success.
Fixes: 2d44667c306e ("fs/ntfs3: Update i_ctime when xattr is added")
Signed-off-by: Baolin Liu <liubaolin@kylinos.cn>
Signed-off-by: Konstantin Komarov <almaz.alexandrovich@paragon-software.com>
|
|
attr_wof_frame_info() allocates, caches and dereferences
ni->file.offs_folio under down_write(&ni->file.run_lock).
ni_decompress_file() frees the same folio with folio_put() without
taking run_lock. When a WOF externally-compressed file is opened for
write, ni_decompress_file() can drop the last reference while a
concurrent O_RDONLY reader is dereferencing the folio in
attr_wof_frame_info(), leading to a use-after-free:
BUG: KASAN: use-after-free in attr_wof_frame_info (fs/ntfs3/attrib.c:1619)
Read of size 4 at addr ffff8880120c9000 by task exploit
Call Trace:
attr_wof_frame_info (fs/ntfs3/attrib.c:1619)
ni_read_frame (fs/ntfs3/frecord.c:2421)
ni_read_folio_cmpr (fs/ntfs3/frecord.c:1916)
ntfs_read_folio (fs/ntfs3/inode.c:648)
read_pages (mm/readahead.c:181)
page_cache_ra_unbounded (mm/readahead.c:292)
force_page_cache_ra (mm/readahead.c:364)
page_cache_sync_ra (mm/readahead.c:573)
filemap_get_pages (mm/filemap.c:2688)
filemap_read (mm/filemap.c:2806)
generic_file_read_iter (mm/filemap.c:2994)
ntfs_file_read_iter (fs/ntfs3/file.c:842)
vfs_read (fs/read_write.c:574)
__x64_sys_pread64 (fs/read_write.c:773)
Take run_lock around the folio_put(), as is done for every other
access to ni->file.offs_folio.
Fixes: 4342306f0f0d ("fs/ntfs3: Add file operations and implementation")
Reported-by: Xiang Mei <xmei5@asu.edu>
Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Signed-off-by: Konstantin Komarov <almaz.alexandrovich@paragon-software.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull x86 fixes from Dave Hansen:
"The most notable fix is THP not silently losing user data and having
been around for a couple of years. The main explanation I'd have for
its longevity is that it requires a few different things to align at
the same time: MADV_FREE, THP and heavy reclaim.
- Fix user-space data loss with THP
- Fix set_memory oopses
- Fix addition of large constants in mul_u64_add_u64_div_u64()
- Fix FineIBT hash offset in cfi_get_func_hash()
- Fix PCI device reference counting in amd_smn_init()"
* tag 'x86_urgent_for_7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
x86/amd_node: Fix PCI device reference counting in amd_smn_init()
x86/div64: Fix addition of large constants in mul_u64_add_u64_div_u64()
x86/cfi: Fix FineIBT hash offset in cfi_get_func_hash()
x86/mm: Fix user-space data loss with MADV_FREE and THP
x86/mm/pat: Allocate split page tables as kernel page tables
x86/alternatives: Exclude text poking against change_page_attr()
x86/mm/pat: Acquire init_mm read lock on attribute changes to avoid UAF
x86/mm/pat: Acquire init_mm write lock on collapse to avoid UAF
|
|
nlmsvc_grant_blocked() unlinks a block from nlm_blocked before it
retries the lock, then re-inserts it. nlmsvc_traverse_blocks() skips a
block that is not on nlm_blocked, so a teardown scan that runs during a
retry passes it by and the retry puts it back. The surviving block pins
its host. lockd warns that it could not shut down the host module, and
the host outlives its network namespace.
Hold the file's f_mutex across the retry, and extend the scan's hold
across its unlink, so a scan and a retry of the same file can no longer
interleave. Drop the mutex before releasing a block reference, since
the last put takes f_mutex. A retry that waited out a scan re-checks
under nlm_blocked_lock that its block is still queued and due.
Reported-by: sashiko-bot <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260819162247.2970703-1-cel@kernel.org?part=1
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260822-lockd-retry-blocked-uaf-v3-2-761661eae60c@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nlmsvc_retry_blocked() examines the block at the head of nlm_blocked
under nlm_blocked_lock, then releases the lock before calling
nlmsvc_grant_blocked() or retry_deferred_block(). The nlm_blocked list
reference is all that keeps the block alive across that window.
nlmsvc_grant_blocked() does take one of its own, but not until after
the lock has been dropped. Unmounting the nfsd filesystem while a lock
request is still blocked reaches nlmsvc_traverse_blocks(), which drops
the list reference and frees the block along with the nlm_rqst hanging
off it.
BUG: KASAN: slab-use-after-free in nlm_async_call+0xd6/0x230
Read of size 8 at addr ffff88811b04c808 by task lockd/8377
nlm_async_call+0xd6/0x230
nlmsvc_retry_blocked+0x61c/0x800
lockd+0x144/0x1c0
Freed by task 8392:
nlmsvc_release_block+0x231/0x290
nlmsvc_traverse_blocks+0x139/0x1b0
nlm_traverse_files+0x1aa/0xa00
nlmsvc_free_host_resources+0x12/0x60
nlm_shutdown_hosts_net+0x127/0x280
lockd_down+0xd5/0x1c0
Take a reference before releasing nlm_blocked_lock and drop it once
the retry has run.
Fixes: 0e4ac9d93515 ("lockd: handle fl_grant callbacks")
Cc: stable@vger.kernel.org
Reported-by: Shuangpeng Bai <shuangpeng.kernel@gmail.com>
Closes: https://lore.kernel.org/linux-nfs/20260818235808.3458075-1-shuangpeng.kernel@gmail.com/
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260822-lockd-retry-blocked-uaf-v3-1-761661eae60c@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
The pNFS SCSI layout server is controlled by NFSD_SCSILAYOUT.
NFSD_SCSI has never existed.
Fixes: f99d4fbdae67 ("nfsd: add SCSI layout support")
Assisted-by: Codex:gpt-5.6-sol
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
Acked-by: Randy Dunlap <rdunlap@infradead.org>
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260822081033.82098-1-kmehltretter@gmail.com
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
svc_tcp_recvfrom() parses the RPC record stream with ->read_sock,
which calls neither security_socket_recvmsg() nor the
sock:sock_recv_length tracepoint. svc_tcp_recv_cmsg() still goes
through sock_recvmsg(), so an LSM mediates only the TLS control
records on a server socket, and sock:sock_recv_length reports only
those. Partial coverage is worse than none. It makes the RPC stream
look mediated and observed when it is not.
Until the record stream moved to ->read_sock, an LSM saw every octet
NFSD read from a TCP socket. An SELinux policy that denies
SOCKET__READ to NFSD blocked the receive. After this change no call
on the server's TCP receive path consults an LSM, so that denial has
no effect.
Dispatch ->recvmsg directly so the whole receive path behaves one
way. sock_recvmsg_nosec() reaches ->recvmsg through
INDIRECT_CALL_INET(), so on a retpoline build the direct dispatch
costs one indirect call per control record. Control records are rare
on an established connection.
Link: https://patch.msgid.link/20260821-tls-read-sock-2-v1-5-7ffce164eb45@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
svc_tcp_recvfrom() reads an RPC record with two recvmsg() calls, one
for the four-octet fragment marker and one for the body. Neither
supplies a control-message buffer, so kTLS reports a TLS control
record on either one by raising MSG_CTRUNC and delivering nothing.
Both receives have to recognize that outcome and hand off to a
recovery path that re-reads the record. Commit bee47cb026e7 ("sunrpc:
fix handling of server side tls alerts") built that recovery path. It
stays on both receives for as long as one recvmsg() serves data and
control alike.
Instead, parse the record stream in a ->read_sock actor. The data
path then carries no control-message buffer at all. A control record
at the head stops ->read_sock short, so svc_tcp_recv_ctrl_record()
classifies the head of the receive queue with MSG_PEEK before
consuming anything.
skb_copy_bits() walks an skb from its head to reach the copy offset,
so an actor that copies a page at a time rewalks a large GRO or kTLS
skb once per page. Copy each callback's share of the record into
rq_bvec with one skb_copy_datagram_iter() instead.
A non-final fragment carries four octets of marker and may carry no
payload. sk_datalen advances only by the payload, so neither a run
of empty fragments nor a run of one-octet fragments trips the
sv_max_mesg check before desc->count runs out. Only desc->count
bounds such a run, four or five octets at a time. Cap the fragments
per socket-lock hold.
recvmsg() skips an urgent octet and clears the condition, so the old
receive path advanced past it. ->read_sock consumes what precedes
the urgent octet and then stops there for good. The data path calls
no recvmsg() now, so the record stream cannot advance again. An RPC
stream carries no urgent data, so close the connection. The read
that first reaches the mark returns a positive count, so a
zero-length read does not mark the stop. Key the close on an
incomplete record.
Link: https://patch.msgid.link/20260821-tls-read-sock-2-v1-4-7ffce164eb45@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
svc_tcp_read_msg() flushes the destination pages after every
receive. A record that arrives in several pieces is therefore
flushed once per piece. A page spanning two pieces is flushed
twice. Nothing reads the message body before the record is
complete, so no reader needs the intermediate flushes.
Flush every page of the record in one pass from svc_tcp_recvfrom(),
once the last fragment has arrived. Remove svc_flush_bvec() and its
ARCH_IMPLEMENTS_FLUSH_DCACHE_PAGE guard. The guard avoided setting
up a bvec iterator, and a plain walk of rq_pages compiles away on
its own where flush_dcache_page() is an empty inline.
Link: https://patch.msgid.link/20260821-tls-read-sock-2-v1-3-7ffce164eb45@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
svc_tcp_sock_recv_cmsg() drains a TLS record that is neither an alert
nor application data, then returns -EAGAIN so the receive loop retries.
It runs only on an established TLS session. A handshake record there
carries a post-handshake message, and the server has no handler for
one. Draining a KeyUpdate only delays the close. kTLS sets
key_update_pending when it decrypts that record, and a later receive
returns -EKEYEXPIRED. kTLS flags only a KeyUpdate, so any other
post-handshake message disappears and the connection keeps running.
Return -EPROTO for an unhandled record type. Any error but -EAGAIN
closes the transport. An application data record keeps its -EAGAIN
return. No NFS client is known to send a handshake record on an
established connection.
Link: https://patch.msgid.link/20260821-tls-read-sock-2-v1-2-7ffce164eb45@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
svc_tcp_sock_recv_cmsg() receives the record at the head of the kTLS
receive queue and decides what its content type means for the
transport. The receive and the decision are one step, so a caller
cannot learn a record's type without also acting on it. A later
caller classifies the head of the queue with MSG_PEEK before it
decides whether to consume the record. It needs the receive without
the decision.
Move the receive into svc_tcp_recv_cmsg(), which takes the recvmsg()
flags and reports the octet count, the record type, and the message
flags. The alert policy stays in svc_tcp_sock_recv_cmsg(), the
helper's only caller in this patch.
The zeroed control buffer replaces the msg_controllen check. An
unfilled buffer reports record type zero, so a positive octet count
with no record type now returns -EBADMSG instead of a count. The
alert parse runs on a second msghdr over the alert buffer, so the
copy and the parse share no iterator state. The rewind that commit
bee47cb026e7 ("sunrpc: fix handling of server side tls alerts")
added by hand is no longer needed.
Link: https://patch.msgid.link/20260821-tls-read-sock-2-v1-1-7ffce164eb45@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
RCA4_TYPE_MASK_* are enum constants, so the preprocessor cannot fold
them into the print format that show_rca_mask() builds for the
nfsd_cb_recall_any event. Nothing declares an eval map for them
either, so trace_event_eval_update() has no substitution to apply at
module load, and the event's format file ships the enumerator names
verbatim. trace-cmd and perf cannot decode the bmval0 field.
Declare the eval maps for the nine mask bits show_rca_mask() decodes.
The format then carries the shift counts as integers, the same shape
the SUNRPC trace points already emit from their BIT() flag decoders.
Fixes: 638593be55c0 ("NFSD: add CB_RECALL_ANY tracepoints")
Cc: stable@vger.kernel.org
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Link: https://patch.msgid.link/20260818184151.31180-1-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Make the header guards less ambiguous about their provenance.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260818140035.12740-3-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
struct nfsd4_copy embeds a struct nfs_fh, and nlm_fopen() reads the
size and data fields of one. Neither fs/nfsd/xdr4.h nor
fs/nfsd/lockd.c includes the header that defines the type; both
reach it by way of nfsd.h, which pulls in <linux/nfs.h>.
Add the direct include to both files, so nfsd.h can later drop the
<linux/nfs.h> it carries for no use of its own.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260818140035.12740-2-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
fs/nfsd/nfsd.h is included throughout the server, yet it uses no
NFSv3 protocol definition of its own. The <linux/nfs3.h> include
there served only to make those definitions reach the few source
files that need them, by way of nfsd.h itself or the xdr.h chain
that pulls it in.
Give each consumer its own include and drop the one in nfsd.h, so
the header no longer carries a dependency unrelated to its
contents. nfsfh.c, nfsctl.c, nfs3xdr.c, nfs3proc.c, and nfs2acl.c
reference NFS3_* definitions directly; add <linux/nfs3.h> to each.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260818140035.12740-1-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
deleg_reaper() asks each eligible client to return one delegation
whenever it runs, whether or not anything needs the memory. A
delegation returned before it is needed costs the client an OPEN
when it next touches the file. Nothing sizes the request either.
The delegation scan callback discards nr_to_scan, which is reclaim's
statement of how many objects it wants back.
Record each delegation scan request in nfsd_deleg_backlog and pass
the accumulated total to deleg_reaper(). Handing that total to
every client would ask for it once per client, so scale it by each
client's share of the delegations this sweep can reach. The count
callback reports what is left after the outstanding requests, so
concurrent reclaimers do not each ask for the same delegations.
cl_ra_time keeps the next sweep from returning to the clients this
one reached.
Nothing is recalled until a scan arrives. nfs4_laundromat() is the
exception. It has no scan request to pass, so it computes what must
go for num_delegations to fall below max_delegations, and passes
only this namespace's share.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-8-3b2cffce701e@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Currently, neither of the scan callback functions records anything
before returning SHRINK_STOP, so the size of each scan request is
discarded. That size is the only real measure NFSD gets of reclaim
pressure. Both count callbacks report their population whether or
not the reaper is already queued to reclaim it, so reclaim asks
again for work that is pending.
Accumulate each courtesy scan request in nfsd_shrink_backlog and
subtract the backlog from what that count callback reports. The
worker retires the backlog once courtesy_client_reaper() has run.
That reaper expires the clients synchronously, so the discount
covers exactly the interval the work is pending.
Delegations need a different bound. This is because deleg_reaper()
only sends CB_RECALL_ANY and does not track how many delegations
were actually returned by the targeted client.
Report the delegations only once NFSD_RECALL_ANY_COOLDOWN_SECS have
passed since the last sweep. deleg_reaper() skips any client it
recalled from within that window, so an earlier scan request cannot
produce another recall.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-7-3b2cffce701e@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Since commit 44df6f439a17 ("NFSD: add delegation reaper to react to
low memory condition"), nfsd_client_shrinker has managed two
unrelated populations of objects.
One population is courtesy clients. Shrinking that population can
be done synchronously and without risk of deadlock. The shrinker
callback could return a precise count of the number of objects
that were released.
The other population is delegations. Shrinking that population
requires sending a CB_RECALL_ANY; clients are not obligated to
return any delegation. The shrinker callback is structurally
unable to report progress.
What's more, the single shrinker callback falls back to
delegation reaping only when there are no courtesy clients left to
reclaim. A single courtesy client is enough to keep a namespace's
delegations out of the count it reports.
To begin to resolve these issues, refactor the existing state
shrinker into two: one for courtesy clients and one for reaping
delegations. Each manages the size of its own population, and the
shrinker names become namespace-specific.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-6-3b2cffce701e@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
The state shrinker is allocated per network namespace, but
nfsd4_state_shrinker_count() reports num_delegations, which counts
the delegations held by the whole host. Every namespace therefore
reports every delegation on the server. Reclaim sees the
population multiplied by the number of namespaces running NFSD. A
namespace holding no delegations of its own still reports a
nonzero count and queues its reaper, which then finds nothing to
recall.
Count the delegations in each namespace and report that instead.
num_delegations stays for the admission check in
__alloc_init_deleg() and the ceiling check in nfs4_laundromat().
Both compare against max_delegations, which is sized from host
memory and so remains a host-wide limit.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-5-3b2cffce701e@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
deleg_reaper() sets craa_objects_to_keep to zero on every
CB_RECALL_ANY. RFC 8881 Section 20.6.3 defines that field as the
number of objects the client may keep, leaving the client to choose
which of the excess to return, because the server cannot read lack
of recent use as lack of usefulness. Zero asks for every delegation
the client holds, including the ones backing files an application
still has open.
There is also no reason NFSD has to reclaim the entire delegation
working set on the first sign of memory pressure.
Derive the keep count from cl_deleg_count so that each callback
asks for one delegation. Both the shrinker and the laundromat re-arm
while their condition lasts, so a client with more to give is asked
again on the next pass.
The Linux client ignores craa_objects_to_keep and returns unused
delegations selected from the type mask alone, so the count changes
nothing for it.
Fixes: 44df6f439a17 ("NFSD: add delegation reaper to react to low memory condition")
Cc: stable@vger.kernel.org
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-4-3b2cffce701e@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
RFC 8881 Section 20.6.3 distinguishes an NFSv4.1 server
implementation that shares one pool among all classes of recallable
objects from one that keeps separate pools per class. NFSD falls
in the former category.
The CB_RECALL_ANY operation's craa_type_mask argument names the
types of objects in the recallable resource pool, but NFSD's
implementation does not name directory delegations, even though
they are allocated through __alloc_init_deleg(), they are counted
against the max_delegations budget, and the state shrinker reclaims
them.
Add RCA4_TYPE_MASK_DIR_DLG to craa_type_mask so clients that
implement directory delegations consider them when choosing which
delegations to return.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-3-3b2cffce701e@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
struct nfs4_client records the delegations it holds on cl_delegations
but keeps no count of them. deleg_reaper() walks nn->client_lru under
nn->client_lock, but cl_delegations is serialized by nn->deleg_lock,
which nests outside nn->client_lock. A caller there cannot take
nn->deleg_lock to count the list. The cost tells against the walk as
well: an O(n) count per client, on a pass that already visits every
client.
Add cl_deleg_count, maintained at the two sites that mutate
cl_delegations. Both hold nn->deleg_lock, so the counter is already
serialized against itself and needs no atomic of its own. The decrement
sits below the delegation_hashed() test, next to the list_del_init it
pairs with, so it runs only when the delegation really leaves the list.
A reader that holds only nn->client_lock is not synchronized against
either update site, so it can see a count that does not match the
list. Such a reader marks the access with data_race() and may not
depend on the value for correctness.
No functional change.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-2-3b2cffce701e@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
deleg_reaper() sends CB_RECALL_ANY to every ACTIVE client holding
delegations, but CB_RECALL_ANY is an NFSv4.1 operation. An NFSv4.0
client's callback service accepts only CB_GETATTR and CB_RECALL, so it
replies OP_ILLEGAL. The decoder maps the unexpected opnum to -EIO, and
nfsd4_cb_done() marks the client's callback channel down.
Nothing brings the channel back. nfsd4_run_cb_work() sets NFSD4_CB_UP
only for a minor version above zero, and the only nfsd4_probe_callback()
call site an NFSv4.0 client reaches is nfsd4_setclientid_confirm(). One
visit from the reaper therefore leaves the channel marked down until the
client re-establishes its clientid. RENEW then returns
NFS4ERR_CB_PATH_DOWN for as long as the client holds delegations.
nfsd4_cb_channel_good() stops returning true, so the client is granted
no further delegations.
Skip clients at minor version zero.
Fixes: 44df6f439a17 ("NFSD: add delegation reaper to react to low memory condition")
Cc: stable@vger.kernel.org
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260817-recall-any-keep-count-v5-1-3b2cffce701e@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd4_encode_operation() leaves op->status alone when the reply
buffer has no room for the operation's opcode and status word.
nfsd4_proc_compound() reads the unchanged nfs_ok as success and
goes on to the next operation, so the reply counts an operation
whose result was never encoded.
Report the failure through nfsd4_check_resp_size(), which the rest
of the function already uses. It returns NFS4ERR_REP_TOO_BIG, or
NFS4ERR_REP_TOO_BIG_TO_CACHE on a session, and the COMPOUND ends
at that operation.
Two paths narrow the reply buffer: nfsd4_sequence(), which rejects
a SEQUENCE result that does not fit, and nfsd4_encode_splice_read(),
which can leave a single XDR word in the head page. Whether a
COMPOUND reaches that boundary is unproven, so this is a guard
rather than a fix.
Link: https://patch.msgid.link/20260817-jean-v1-2-9e356596ab85@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
nfsd4_sequence() narrows the reply buffer to the session's cached
reply limit before it accepts the slot seqid. A client may negotiate
ca_maxresponsesize_cached down to NFSD_MIN_HDR_SEQ_SZ, and
nfsd4_alloc_slot() then gives every slot a zero-length sl_data[]. A
COMPOUND tag can fill that narrowed buffer until it holds the
SEQUENCE opcode but not the status word that follows.
nfsd4_encode_operation() returns without running
nfsd4_encode_sequence(), so cstate.data_offset stays zero. It leaves
op->status at nfs_ok as well, so the COMPOUND is treated as having
succeeded.
nfsd4_store_cache_entry() declines to cache a lone SEQUENCE that
returned an error. That test reads the status the operation
reported, so it passes here. The copy starts at offset zero and
takes the whole reply, RPC and COMPOUND headers included, into the
zero-length sl_data[]. The COMPOUND tag is copied along with it, so
the client picks most of the bytes written past the end of the slot:
BUG: KASAN: slab-out-of-bounds in read_bytes_from_xdr_buf+0x1bc/0x390
Write of size 80 at addr ffff888003a549cd by task kunit_try_catch/24
__asan_memcpy+0x38/0x60
read_bytes_from_xdr_buf+0x1bc/0x390
nfsd4_sequence_done+0x5b0/0x810
nfs4svc_encode_compoundres+0x1bf/0x240
Check that the fixed-size SEQUENCE result, plus room for a following
operation's error status, fits the negotiated limit before narrowing
the buffer and consuming the slot seqid. The slot and its reply
cache are left unchanged, as RFC 8881 Section 2.10.6.1.2 requires of
an error returned from SEQUENCE.
Fixes: 47ee52986472 ("nfsd4: adjust buflen to session channel limit")
Cc: stable@vger.kernel.org
Assisted-by: Codex:gpt-5
Signed-off-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr>
Link: https://patch.msgid.link/20260817-jean-v1-1-9e356596ab85@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
The reply to a pool_threads read is the list of per-pool thread counts,
formatted into a buffer of SIMPLE_TRANSACTION_LIMIT bytes. snprintf()
truncates its last write and strlen() measures only what fit, so a reply
too long for that buffer ends mid-number with no terminating newline. A
pool running 4096 threads is reported as 40. Nothing marks the reply as
incomplete, so an administrator reads a plausible but wrong count.
Take snprintf()'s return value, which reports the truncation strlen()
cannot see, and fail the read with -ENAMETOOLONG when the list does
not fit. That is the errno svc_one_xprt_name() already returns for the
same condition.
Suggested-by: David Laight <david.laight.linux@gmail.com>
Fixes: eed2965af1ba ("[PATCH] knfsd: allow admin to set nthreads per node")
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260812193349.13347-1-david.laight.linux@gmail.com
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
Writing the same socket descriptor to /proc/fs/nfsd/portlist twice
attaches a second svc_sock to one socket. svc_setup_socket() saves
the socket's callbacks before installing its own, so the second
attach records svc_write_space() as the old write_space callback.
svc_udp_init() invokes that callback by way of svc_sock_setbufsize(),
and svc_write_space() then calls itself until the kernel stack is
exhausted:
BUG: TASK stack guard page was hit at ffffc900037d7ff8
svc_write_space+0x90/0x2b0 net/sunrpc/svcsock.c:429
svc_write_space+0xe6/0x2b0 net/sunrpc/svcsock.c:430
... 700 more ...
svc_sock_setbufsize+0x18d/0x220 net/sunrpc/svcsock.c:386
svc_udp_init net/sunrpc/svcsock.c:854 [inline]
svc_setup_socket+0xb2f/0x1090 net/sunrpc/svcsock.c:1498
svc_addsock+0x2fd/0x760 net/sunrpc/svcsock.c:1547
__write_ports_addfd fs/nfsd/nfsctl.c:742 [inline]
write_ports+0xa5b/0xcc0 fs/nfsd/nfsctl.c:861
nfsctl_transaction_write+0x106/0x1a0 fs/nfsd/nfsctl.c:112
svc_data_ready() and svc_tcp_state_change() chain through their saved
callbacks the same way, so a TCP descriptor added twice recurses on
the next incoming segment instead. Reaching any of this takes a
writer on portlist, and the nfsd filesystem sets no FS_USERNS_MOUNT,
so the reproducer needs CAP_SYS_ADMIN in the initial user namespace.
Reject a socket that already carries sk_user_data. svc_setup_socket()
overwrites that field unconditionally, so a socket some other
consumer has claimed is one NFSD would corrupt whether or not the
callbacks recurse.
Fixes: b41b66d63c73 ("[PATCH] knfsd: allow sockets to be passed to nfsd via 'portlist'")
Cc: stable@vger.kernel.org
Reported-by: syzbot+54cdc566f64abf51b7f1@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=54cdc566f64abf51b7f1
Link: https://patch.msgid.link/20260815162844.8219-1-cel@kernel.org
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
The unix_gid netlink upcall protocol has an explicit
SUNRPC_A_UNIX_GID_NEGATIVE attribute for failed group lookups, and
mountd sends it, but sunrpc_nl_parse_one_unix_gid() only allocates
an empty group list for a flagged reply without propagating the
flag into the entry, so it installs a valid positive entry with
zero groups. Such an entry strips the uid of all supplementary
groups until it is refreshed or expires: the same defect the
previous patch fixes on the classic channel, on a transport that
can say "lookup failed" explicitly.
Set CACHE_NEGATIVE for flagged replies, as the netlink ip_map
path already does for its negative flag. An empty GIDS list
without the flag remains a positive entry.
Fixes: 0850e8603cd7 ("sunrpc: add netlink upcall for the auth.unix.gid cache")
Cc: stable@vger.kernel.org
Assisted-by: Claude:claude-fable-5
Signed-off-by: Ameer Hamza <ameer.hamza@truenas.com>
Link: https://patch.msgid.link/20260814221953.108837-3-ameer.hamza@truenas.com
Signed-off-by: Chuck Lever <cel@kernel.org>
|