| Age | Commit message (Collapse) | Author |
|
https://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs.git
# Conflicts:
# fs/smb/server/smb2pdu.c
# fs/smb/server/vfs.c
# fs/smb/server/vfs.h
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/cel/linux
|
|
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/jaegeuk/f2fs.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/jack/linux-fs.git
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/leitao/linux.git
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/mtd/linux
Pull MTD fixes from Miquel Raynal:
"The most important set of fixes are around the handling of the QE bit
in SPI NAND.
There are also a couple of behavioral fixes (mutex issue in SPI-NOR,
spurious bitflips on vf610_nfc, OOB bytes count in SPI NAND and
cfi_cmdset stack usage).
The rest is mostly AI fuzzing results"
* tag 'mtd/fixes-for-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/mtd/linux:
mtd: spinand: Do not update the QE bit on devices without one
mtd: spi-nor: core: Fix mutex leak in spi_nor_rww_start_exclusive()
mtd: rawnand: cadence: Initialize IRQ state before requesting IRQ
mtd: rawnand: vf610_nfc: fix false bitflips on reads of erased pages
mtd: rawnand: vf610_nfc: fix reads on chips with more than 64 bytes of OOB
mtd: spinand: fix zero oobavail when no ECC engine is used
mtd: spinand: fix NULL pointer dereference with no ECC engine
mtd: mtd_intel_dg: reset poll counter for each erase
mtd: cfi_cmdset_0001: shrink do_write_buffer() stack frame
mtd: core: call _get_device() with the master MTD
mtd: core: avoid double-free of OTP NVMEM device
mtd: spinand: Enable QE on all dies
mtd: block2mtd: Fix divide error when erase_size is zero
|
|
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull scheduler fixes from Ingo Molnar:
- Fix LLC mis-scheduling bugs (Tim Chen, Lu Wang)
- Fix cache-grouping related scheduling statistics UAF bugs (Tim Chen)
- Skip kernel threads for cache aware scheduling to rubustify the code
(Chen Yu)
- Refresh LLC capacity across CPU hotplug, to fix capacity
underestimation bug (Davi Chaves Azevedo)
- Account PSI IRQ time to the execution context, not the scheduling
context, to fix proxy scheduling accounting bug (Zhan Xusheng)
* tag 'sched-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
sched/core: Account PSI IRQ time to the execution context, not the scheduling context
sched/cache: Refresh LLC capacity across CPU hotplug, to fix capacity underestimation bug
sched/cache: Skip kernel threads for cache aware scheduling to rubustify the code
sched/cache: Introduce task_struct->sched_cache_grp to fix UAF
sched/cache: Decouple sched_cache_group from mm to fix UAF
sched/cache: Honor migrate_llc_task semantics in active load balance, to fix LLC mis-scheduling bug
sched/cache: Keep nr_pref_llc_running in the runnable domain, to fix LLC mis-scheduling bug
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull perf events fixes from Ingo Molnar:
- Fixes for KVM guest PEBS virtualization (Sean Christopherson)
- Fixes for various Intel PMUs related to PEBS data-source (Dapeng Mi)
- Fix Intel Panther Cove event scheduling constraints (Dapeng Mi)
- Fix Intel DMR/NVL OMR extra registers event scheduling (Dapeng Mi)
- Rename two confusingly named PMU attributes (Dapeng Mi)
- Fix a refcount leak in attach_perf_ctx_data() (Namhyung Kim)
- Fix NULL pointer dereference crash in __perf_pmu_sched_task()
(Puranjay Mohan)
- Fix CPU-wide event scheduling (Puranjay Mohan)
- Fix x86 LBR branch entry generation (Puranjay Mohan)
* tag 'perf-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
perf/core: Fill branch entries with a single assignment
perf/core: Run sched_task() for PMUs with only CPU-wide events
perf/core: Fix NULL pmu_ctx passed to pmu->sched_task()
perf/core: Fix a refcount leak in attach_perf_ctx_data()
perf/x86/intel: Rename NVL offcore_rsp attribute to offmodule_rsp
perf/x86/intel: Rename DMR offcore_rsp attribute to offmodule_rsp
perf/x86/intel: Fix precise OMR event scheduling for DMR/NVL
perf/x86/intel: Constrain Panther Cove UOPS_DISPATCHED events to PMCs 0-3
perf/x86/intel: Delete dead NVL PEBS data-source initcall
perf/x86/intel: Fix Panther Cove PEBS data-source snoop states
perf/x86/intel: Remove incorrect Panther Cove PEBS data-source constraints
perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints
perf/x86/intel: Update arw_latency_data() mem-op direction handling
perf/x86/intel: Fix DKT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Fix CMT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs()
perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) if PEBS is unused
perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has PEBS isolation, to fix stuck PEBS_ENABLED
perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted PERF_GLOBAL_CTRL bits
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux
Pull ata fixes from Niklas Cassel:
- Extend the quirk "no LPM on ATI" quirk, that is currently only
applied for Samsung drives, to include AMD controllers as well.
The AMD AHCI controllers are newer versions of the ATI AHCI
controllers, and these controllers still have LPM issues with
Samsung drives - LPM works with drives from other vendors (me)
- Fix errors in the libata.force parameter documentation (me)
- Verify the sense data descriptor lengths for ATA PASS-THROUGH
command, so that a malicious device cannot write past the buffer
length (Matthias)
- Mention the libata for-next branch in MAINTAINERS such that the
git ls-remote command done by get_maintainer.pl --self-test=scm
can verify it (Matthias)
* tag 'ata-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux:
MAINTAINERS: name the libata/linux for-next branch
ata: libata-scsi: bound the ATA passthru sense descriptor writes
ata: libata: Correct libata.force parameter documentation
ata: libata-core: Extend Samsung LPM quirk to AMD controllers
|
|
Pull kvm fixes from Paolo Bonzini:
"Arm:
- Invalidate the ITS translation cache when the guest changes the
base address of the ITS tables (Fuad Tabba)
- Skip saving ITS devices with device IDs that are out-of-bounds
rather than failing the entire ITS save ioctl (Fuad Tabba)
- Close race between VM teardown and invalidations of nested MMUs
when handling MMU operations that are allowed to block (Lorenzo
Stoakes)
- Various fixes for the handling of the host's untrusted SVE
configuration in pKVM (Fuad Tabba)
- Make sure that empty SMCCC ranges based at 0 are rejected by the
kvm_smccc_set_filter() (Karl Mehltretter)
- Revoke the host mapping for pKVM's private stack pages, along with
a new sanity check that all mappings in the hyp's private VA range
have been correctly marked as hyp-owned (Fuad Tabba)
- Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that
concurrent vCPU initialization cannot relocate in-use MMUs. Defer
the freeing of shadow stage-2 MMUs to the point that no other users
(e.g. MMU notifier) could reference them (Marc Zyngier)
- Drop useless WARN when rejecting an unsupported ioctl for pKVM
(Fuad Tabba)
- Fix the steal_time selftest to install correctly-sized mappings for
non-4K hosts (Sebastian Ott)
- Correct mapping of fine-grained trap for GCSPOPX instruction (Mark
Brown)
- Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit
guests (Karl Mehltretter)
RISC-V:
- Synchronize hrtimer during VCPU teardown
- Fix the conversion between vsip and hvip values
- Serialize IMSIC attributes with vCPU migration
- Release unused page after MMU invalidation
- Propagate interrupted G-stage faults to KVM user-space as EINTR
- Fix nested acceleration hfence entry update order
- Fix sdata leak and stale snapshot_addr in snapshot_set_shmem
- Preserve firmware counter value across PMU counter stop/start
- Report PMU snapshot write failure to the guest
- Fix perf-backed counter accounting across PMU stop and read
- Correctly propagate error of a hart status SBI call
s390:
- Ensure that accesses through kvm_arch_set_irq_inatomic mark as
dirty the pages that contain indicator and summary bits
- Fix compile warning for kvm_s390_update_cmma_dirty()
- Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to
userspace
- Move s390_kvm_mmu_commit_memory_region() into
s390_kvm_mmu_prepare_memory_region() so that it can fail instead of
WARN
- Add missing srcu in kvm_s390_set_irq_state()
- Fix potential races in storage functions
- Fix race in _destroy_pages_crste()
- Fix issues in the handling of KVM interrupt and page resources,
when a queue that is assigned to a mediated device (mdev) is
removed from the host's AP configuration
- Fix loop condition in uv_find_secrets
- Prevent potential out-of-bounds read
x86:
- Fix a brown paper bag bug where KVM would incorrectly treat Intel
PMU MSRs as valid on AMD
- Fix a regression in the hardware disable selftest where it checked
the wrong macro when detecting glibc support (breaks at least musl)
- Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is
especially important for KVM_BUG_ON() flows, which often guard more
dangerous bugs
- Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a
bug where KVM would let userspace run a broken setup with stale
vmcs12 pages
- Fix a class of bugs where KVM would fail to fill kvm_run exit
fields if getting nested pages failed
- Treat reserved entries in the memory attributes xarray as "no
attributes", to fix false positives when checking for mixed
attributes
- Fix memcg accounting for the memory attributes xarray (the xarray
library subtly requires the xarray to be configured for accounting
upfront; the gfp flags taken at runtime are used only rarely)
- Don't pre-reserve xarray entries when storing empty attributes, as
storing NULL must not require memory allocation (KVM and other
subsystems heavily rely on this behavior)
- Fix a memory leak and a cache maintenance issue related to doing
intra-host migration on an SEV guest"
* tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (54 commits)
KVM: SEV: Do cache maintenance on the source VM during intra-host migration
KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV
KVM: Don't pre-reserve xarray entries when storing empty/NULL attributes
KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg
KVM: Don't treat reserved xarray entries as having memory attributes
KVM: x86: Fill kvm_run exit fields in common get_nested_state_pages() error paths
KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails
KVM: arm64: Fix AArch32 DBGBXVR<n> handling
KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP
KVM: selftests: fix steal_time for arm64 with host page size > 4K
KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode
KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction
KVM: arm64: nv: Fix life cycle of the nested_mmus array
KVM: arm64: Check every private mapping is hyp-owned at pKVM init
KVM: arm64: Move the private VA allocation cursor to __io_map_next
KVM: arm64: Match hyp text by physical address in fix_host_ownership()
KVM: arm64: Transfer the hyp stack pages out of the host stage-2
KVM: arm64: selftests: Test empty SMCCC filter range at base 0
KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0
KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2
...
|
|
KVM fixes for 7.3-rcN
- Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs
as valid on AMD.
- Fix a regression in the hardware disable selftest where it checked the wrong
macro when detecting glibc support (breaks at least musl).
- Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially
important for KVM_BUG_ON() flows, which often guard more dangerous bugs.
- Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug
where KVM would let userspace run a broken setup with stale vmcs12 pages.
- Fix a class of bugs where KVM would fail to fill kvm_run exit fields if
getting nested pages failed.
- Treat reserved entries in the memory attributes xarray as "no attributes",
to fix false positives when checking for mixed attributes.
- Fix memcg accounting for the memory attributes xarray (the xarray library
subtly requires the xarray to be configured for accounting upfront; the gfp
flags taken at runtime are used only rarely).
- Don't pre-reserve xarray entries when storing empty attributes, as storing
NULL must not require memory allocation (KVM and other subsystems heavily
rely on this behavior).
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs
Pull vfs fixes from Christian Brauner:
- Revert "put_mnt_ns(): leave mounts connected". This allows the
creation of reference count cycles in a very trivial way. We can't
bring this in until we have fixed the underlying cause
- vfs: Don't create the private nullfs instance for kthreads under
namespace_sem to avoid false lockdeps complaints
- binfmt_misc:
- Copy the name into a stack buffer and look up the copy in
bpf_binprm_select_interp()
- bpf_binprm_set_interp() and bpf_binprm_set_interp_arg(): Check
the private copy instead so the string that gets staged is the
kstring that was checked
- netfs:
- Make netfs_read_gaps() use separate sink folios rather than one
reused sink folio to discard unwanted data so that cifs checksum
checking sees all the data that was fetched
- Trim reads down to i_size so afs symlinks read correctly from the
cache
- Wrap the direct mempool ->alloc() calls the GFP_KERNEL paths make
in alloc_hooks() via a new mempool_alloc_noreserve() helper
- iov_iter: Use iov_iter_alignment() for the start and length check
added to iov_iter_extract_bvecs() this cycle. It used iter_iov_addr()
and iter_iov_len() which are only valid for ITER_UBUF and ITER_IOVEC
iterators
- super: Make iterate_supers_type() deletion-safe
- inode: Stop evict_inodes() from rescanning the same inodes
- writeback: Bound the cleanup_offline_cgwb() rescans
- ntfs3: Use d_instantiate_new() in ntfs_create_inode()
- ovl: Fix a use-after-free in the ovl_do_mkdir() debug print
- dcache: Unpoison the inline name buffer in __d_alloc() for KMSAN
- autofs: Fix a pipe file reference leak in autofs_kill_sb()
- bpf: Drop the path_unlink and path_rmdir hooks from the list of hooks
for which the verifier rewrites bpf_{set,remove}_dentry_xattr() to
the _locked variants
- squashfs: Range check the xz dictionary size before shifting by it
- selftests: Add the missing eventfd, open_tree_ns, openat2 and xattr
filesystems selftests to TARGETS and drop the stale openat2 entry
left behind when those tests moved
* tag 'vfs-7.3-rc5.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs:
netfs: Fix missing alloc tagging of direct mempool allocations
bpf: fs/xattr: don't assume the inode is locked in path_unlink/path_rmdir
autofs: fix sbi->pipe file reference leak in autofs_kill_sb()
dcache: unpoison the inline name buffer in __d_alloc()
ovl: fix UAF in ovl_do_mkdir() debug print
super: make iterate_supers_type() deletion-safe
Revert "put_mnt_ns(): leave mounts connected"
Revert "selftests/filesystems: add mntns cleanup test"
binfmt_misc: fix racy checks in bpf set_interp kfuncs
binfmt_misc: fix OOB read in bpf_binprm_select_interp()
fs: don't create the private nullfs mount under namespace_sem
writeback: bound cleanup_offline_cgwb() rescans by rotating scanned inodes
fs: avoid repeated scans in evict_inodes()
netfs, afs: Fix symlink reading
netfs: Fix netfs_read_gaps() to use separate sink folios
squashfs: Add dictionary size range check to prevent shift-out-of-bounds
fs/ntfs3: use d_instantiate_new() in ntfs_create_inode() and murder syzbot's "WARNING in do_new_mount" saga
selftests/filesystems: fix missing and stale TARGETS entries
block: Fix start and length check added to iov_iter_extract_bvecs()
|
|
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
|
|
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
Commit 1d78d56c43ef ("netfs: Fix folio_queue ENOMEM in writeback by
adding a mempool") added a mempool for the folio_queues and made the
request, subrequest and folio_queue allocations distinguish between
writeback and everything else. Writeback is part of memory reclaim
and must not fail due to ENOMEM, so it allocates under GFP_NOFS
through mempool_alloc(), which may dip into the pool's reserve and,
if that runs empty, wait for elements to be returned. The
GFP_KERNEL paths, which can return -ENOMEM to their callers, invoke
the pool's ->alloc() callback directly instead.
The direct call, however, skips the alloc_hooks() wrapper that the
mempool_alloc() macro provides. The pool callbacks, mempool_alloc_slab()
and mempool_kmalloc(), call kmem_cache_alloc_noprof() and kmalloc_noprof()
and rely on current->alloc_tag having been set by the caller. With
CONFIG_MEM_ALLOC_PROFILING_DEBUG=y this leads to
current->alloc_tag not set
WARNING: ./include/linux/alloc_tag.h:161 at __alloc_tagging_slab_alloc_hook
alloc_tag was not set
WARNING: ./include/linux/alloc_tag.h:166 at __alloc_tagging_slab_free_hook
at allocation and free time respectively, as reported when reading
files on a CIFS mount. The allocations are also missing from
/proc/allocinfo.
Wrap the direct ->alloc() invocations in alloc_hooks() with a new
mempool_alloc_noreserve() helper in include/linux/mempool.h, next to
the other alloc_hooks()-wrapped macros such as mempool_alloc(). The
GFP_KERNEL paths keep their failable allocation semantics, they just
get tagged now.
Fixes: 1d78d56c43ef ("netfs: Fix folio_queue ENOMEM in writeback by adding a mempool")
Reported-by: Erhard Furtner <erhard_f@mailbox.org>
Closes: https://lore.kernel.org/all/0b004319-9ef7-437c-a4dd-174d6a9a83db@mailbox.org/
Tested-by: Erhard Furtner <erhard_f@mailbox.org>
Suggested-by: Suren Baghdasaryan <surenb@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Hao Ge <hao.ge@linux.dev>
Link: https://patch.msgid.link/20260923063759.34667-1-hao.ge@linux.dev
Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Rename and export it so filesystems that need a custom superblock
matching policy can reuse it directly instead of open-coding sget_fc()
+ fill_super(). No functional change.
This is a preparatory fix for the next patch.
Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com>
Link: https://patch.msgid.link/20260812142907.1010046-2-gscrivan@redhat.com
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Since commit e896474fe485 ("getname_maybe_null() - the third variant of
pathname copy-in"), vfs_empty_path() has no callers.
So remove it.
No functional change.
Signed-off-by: Sang-Heon Jeon <ekffu200098@gmail.com>
Link: https://patch.msgid.link/20260918165105.1013792-1-ekffu200098@gmail.com
Reviewed-by: Jan Kara <jack@suse.cz>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
coredump_wait() sets core_state->nr_threads to the number of tasks
killed and waits for the last thread to enter coredump_task_exit() to
signal completion. Let's just wait on the count directly. The exiting
tasks can use atomic_dec_and_wake_up() and the dumping task sleeps in
wait_var_event_state().
The dumping task must remain freezable since commit f5d39b020809
("freezer,sched: Rewrite core freezer logic"). So keep the wait
TASK_UNINTERRUPTIBLE|TASK_FREEZABLE.
Drop the completion and rename nr_threads to threads_remaining.
No functional changes.
Suggested-by: NeilBrown <neilb@ownmail.net>
Link: https://lore.kernel.org/178899497961.207413.10554121774377911612@noble.neil.brown.name
Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-12-a5c1800dc930@kernel.org
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
All wait_var_event() sleep in a fixed task state. For coredumps we need
a variant that takes the state from the caller the way
wait_event_state() does. This allows us to continue sleeping with
TASK_FREEZABLE. That's certainly also a useful addition for other places.
Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-11-a5c1800dc930@kernel.org
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
The core_state->dumper field isn't used anymore. Only its ->next pointer
is. The current task is always the dumping thread and the ->task pointer
is never read. Replace it with a plain pointer to the list of parked
threads.
Historically, core_state->dumper was used. Its ->task pointer was read.
by fill_note_info() started at &core_state->dumper to ensure that the
dumping thread came first in the ELF thread notes. That changed in
commit 4b0e21d64253 ("[elf][regset] simplify thread list handling in
fill_note_info()"). The first iteration was taken out of the loop. So
it's been unused ever since.
No functional changes.
Suggested-by: NeilBrown <neilb@ownmail.net>
Link: https://lore.kernel.org/178900159210.207413.8292125177519817528@noble.neil.brown.name
Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-10-a5c1800dc930@kernel.org
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Rename the helper and align it with close_files().
No functional changes.
Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-8-a5c1800dc930@kernel.org
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
exec is the only caller left since commit 433967cab51e ("coredump: stop
unsharing the file descriptor table"). All it does is call unshare_fd()
with CLONE_FILES and install the copy. Kill the pointless helper and
open-code it.
No functional changes.
Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-4-a5c1800dc930@kernel.org
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Move unshare_fd() where the rest of the descriptor table lifecycle
helpers live.
No functional changes.
Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-3-a5c1800dc930@kernel.org
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Add switch_files_struct() to install another table on a task. It
consumes the reference to the new table and puts the old one. Convert
every place that switches a descriptor table except unshare_files().
No functional changes.
Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-2-a5c1800dc930@kernel.org
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
DCACHE_PRIVATE may be used by any filesystem for its own purposes, much
like d_fsdata and d_time.
I plan to use this in a similar manner the way nfs stores
NFS_FSDATA_BLOCKED in d_fsdata.
Signed-off-by: NeilBrown <neil@brown.name>
Link: https://patch.msgid.link/20260904215142.1060510-8-neilb@ownmail.net
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
DCACHE_PAR_LOOKUP acts like a lock in that threads can block waiting for
it to clear. As we plan to make changes to lock order for this lock,
teach lockdep to monitor it so as to help detect bugs early.
As NFS allocates an in-lookup dentry to unlink a silly-renamed file, and
completes the lookup in a different thread, we need interfaces to
release and the acquire ownership of the lock. This avoids lockdep
complaining that a lock is still held on return to user-space.
Signed-off-by: NeilBrown <neil@brown.name>
Link: https://patch.msgid.link/20260904215142.1060510-7-neilb@ownmail.net
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Some ->lookup handlers will need to drop and retake the parent lock, so
they can safely use d_alloc_parallel().
->lookup can be called with the parent lock either exclusive or shared.
A new flag, LOOKUP_SHARED, tells ->lookup how the parent is locked.
This is rather ugly, but will be gone soon after we move
d_alloc_parallel() out of the directory lock as ->lookup() will *always*
called with a shared lock on the parent.
Signed-off-by: NeilBrown <neil@brown.name>
Link: https://patch.msgid.link/20260904215142.1060510-6-neilb@ownmail.net
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Occasionally a single operation can require two sub-operations on the
same name, and it is important that a d_alloc_parallel() (once that can
be run unlocked) does not create another dentry with the same name
between the operations.
Two examples:
1/ rename where the target name (a positive dentry) needs to be
"silly-renamed" to a temporary name so it will remain available on the
server (NFS and AFS). Here the same name needs to be the subject
of one rename, and the target of another.
2/ rename where the subject needs to be replaced with a white-out
(shmemfs). Here the same name need to be the subject of a rename
and the target of a mknod()
In both cases the original dentry is renamed to something else, and a
replacement is instantiated, possibly as the target of d_move(), possibly
by d_instantiate().
Currently d_alloc() is used to create the dentry and the exclusive lock
on the parent ensures no other dentry is created. When
d_alloc_parallel() is moved out of the parent lock, this will no longer
be sufficient. In particular if the original is renamed away before the
new is instantiated, there is a window where d_alloc_parallel() could
create another name. "silly-rename" does work in this order. shmemfs
whiteout doesn't open this hole but is essentially the same pattern and
should use the same approach.
The new d_duplicate() creates an in-lookup dentry with the same name as
the original dentry, which must be hashed. There is no need to check if
an in-lookup dentry exists with the same name as d_alloc_parallel() will
never try add one while the hashed dentry exists. Once the new
in-lookup is created, d_alloc_parallel() will find it and wait for it to
complete, then use it.
Signed-off-by: NeilBrown <neil@brown.name>
Link: https://patch.msgid.link/20260904215142.1060510-5-neilb@ownmail.net
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Several filesystems use the results of readdir to prime the dcache.
These filesystems use d_alloc_parallel() which can block if there is a
concurrent lookup. Blocking in that case is pointless as the lookup
will add info to the dcache and there is no value in the readdir waiting
to see if it should add the info too.
Also these calls to d_alloc_parallel() are made while the parent
directory is locked. A proposed change to locking will lock the parent
later, after d_alloc_parallel(). This means it won't be safe to wait in
d_alloc_parallel() while holding the directory lock.
So this patch introduces d_alloc_trylock() which doesn't block but
instead returns ERR_PTR(-EWOULDBLOCK). Filesystems that prime the
dcache (smb/client, nfs, fuse, cephfs) can now use that and ignore
-EWOULDBLOCK errors as harmless.
Unlike d_alloc_parallel(), d_alloc_trylock() calculates the hash and
performs a lookup before an allocation, as that is what all callers
want. This is done using try_lookup_noperm(), necessitating the
inclusion of namei.h in dcache.c.
Signed-off-by: NeilBrown <neil@brown.name>
Link: https://patch.msgid.link/20260904215142.1060510-4-neilb@ownmail.net
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Cachefiles currently uses the backing filesystem's idea of what data is
held in a backing file and queries this by means of SEEK_DATA and
SEEK_HOLE. However, this means it does two seek operations on the backing
file for each individual read call it wants to prepare (unless the first
returns -ENXIO). Worse, the backing filesystem is at liberty to insert or
remove blocks of zeros in order to optimise its layout which may cause
false positives and false negatives.
The problem is that keeping track of what is dirty is tricky (if storing
info in xattrs, which may have limited capacity and must be read and
written as one piece) and expensive (in terms of diskspace at least) and is
basically duplicating what a filesystem does.
However, the most common write case, in which the application does {
open(O_TRUNC); write(); write(); ... write(); close(); } where each write
follows directly on from the previous and leaves no gaps in the file is
reasonably easy to detect and can be noted in the primary xattr as
CACHEFILES_CONTENT_ALL, indicating we have everything up to the object size
stored.
In this specific case, given that it is known that there are no holes in
the file, there's no need to call SEEK_DATA/HOLE or use any other mechanism
to track the contents. That speeds things up enormously.
Even when it is necessary to use SEEK_DATA/HOLE, it may not be necessary to
call it for each cache read subrequest generated.
Implement this by adding support for the CACHEFILES_CONTENT_ALL content
type (which is defined, but currently unused), which requires a slight
adjustment in how backing files are managed. Specifically, the driver
needs to know how much of the tail block is data and whether storing more
data will create a hole.
To this end, the way that the size of a backing file is managed is changed.
Currently, the backing file is expanded to strictly match the size of the
network file, but this can be changed to carry more useful information.
This makes two pieces of metadata available: xattr.object_size and the
backing file's i_size. Apply the following schema:
(a) i_size is always a multiple of the DIO block size.
(b) i_size is only updated to the end of the highest write stored. This
is used to work out if we are following on without leaving a hole.
(c) xattr.object_size is the size of the network filesystem file cached
in this backing file.
(d) xattr.object_size must point after the start of the last block
(unless both are 0).
(e) If xattr.object_size is at or after the block at the current end of
the backing file (ie. i_size), then we have all the contents of the
block (if xattr.content == CACHEFILES_CONTENT_ALL).
(f) If xattr.object_size is somewhere in the middle of the last block,
then the data following it is invalid and must be ignored.
(g) If data is added to the last block, then that block must be fetched,
modified and rewritten (it must be a buffered write through the
pagecache and not DIO).
(h) Writes to cache are rounded out to blocks on both sides and the
folios used as sources must contain data for any lower gap and must
have been cleared for any upper gap, and so will rewrite any
non-data area in the tail block.
To implement this, the following changes are made:
(1) cookie->object_size is no longer updated when writes are copied into
the pagecache, but rather only updated when a write request completes.
This prevents object size miscomparison when checking the xattr
causing the backing file to be invalidated (opening and marking the
backing file and modifying the pagecache run in parallel).
(2) The cache's current idea of the amount of data that should be stored
in the backing file is kept track of in object->object_size.
Possibly this is redundant with cookie->object_size, but the latter
gets updated in some addition circumstances.
(3) The size of the backing file at the start of a request is now tracked
in struct netfs_cache_resources so that the partial EOF block can be
located and cleaned.
(4) The cache block size is now used consistently rather than using
CACHEFILES_DIO_BLOCK_SIZE (4096).
(5) The backing file size is no longer adjusted when looking up an object.
(6) When shortening a file, if the new size is not block aligned, the part
beyond the new size is cleared. If the file is truncated to zero, the
content_info gets reset to CACHEFILES_CONTENT_NO_DATA.
(7) A new struct, fscache_occupancy, is instituted to track the region
being read. Netfslib allocates it and fills in the start and end of
the region to be read then calls the ->query_occupancy() method to
find and fill in the extents. It also indicates whether a recorded
extent contains data or just contains a region that's all zeros
(FSCACHE_EXTENT_DATA or FSCACHE_EXTENT_ZERO).
(8) The ->prepare_read() cache method is changed such that, if given, it
just limits the amount that can be read from the cache in one go. It
no longer indicates what source of read should be done; that
information is now obtained from ->query_occupancy().
(9) A new cache method, ->collect_write(), is added that is called when a
contiguous series of writes have completed and a discontiguity or the
end of the request has been hit. It it supplied with the start and
length of the write made to the backing file and can use this
information to update the cache metadata.
(10) cachefiles_query_occupancy() is altered to find the next two "extents"
of data stored in the backing file by doing SEEK_DATA/HOLE between the
bounds set - unless it is known that there are no holes, in which case
a whole-file first extent can be set.
(11) cachefiles_collect_write() is implemented to take the collated write
completion information and use this to update the cache metadata, in
particular working out whether there's now a hole in the backing file
requiring future use of SEEK_DATA/HOLE instead of just assuming the
data is all present.
It also uses fallocate(FALLOC_FL_ZERO_RANGE) to clean the part of a
partial block that extended beyond the old object size. It might be
better to perform a synchronous DIO write for this purpose, but that
would mandate an RMW cycle. Ideally, it should be all zeros anyway,
but, unfortunately, shared-writable mmap can interfere.
(12) cachefiles_begin_operation() is updated to note the current backing
file size and the cache DIO size.
(13) cachefiles_create_tmpfile() no longer expands the backing file when it
creates it.
(14) cachefiles_set_object_xattr() is changed to use object->object_size
rather than cookie->object_size.
(15) cachefiles_check_auxdata() is altered to actually store the content
type and to also set object->object_size. The cachefiles_coherency
tracepoint is also modified to display xattr.object_size.
(16) netfs_read_to_pagecache() is reworked. The cache ->prepare_read()
method is replaced with ->query_occupancy() as the arbiter of what
region of the file is read from where, and that retrieves up to two
occupied extents of the backing file at once.
The cache ->prepare_read() method is now repurposed to be the same as
the equivalent network filesystem method and allows the cache to limit
the size of the read before the iterator is prepared.
netfs_single_dispatch_read() is similarly modified.
(17) netfs_update_i_size() and afs_update_i_size() no longer call
fscache_update_cookie() to update cookie->object_size.
(18) Write collection now collates contiguous sequences of writes to the
cache and calls the cache ->collect_write() method.
Signed-off-by: David Howells <dhowells@redhat.com>
Link: https://patch.msgid.link/20260910220242.2165023-5-dhowells@redhat.com
Reviewed-by: Paulo Alcantara <pc@manguebit.org>
cc: Matthew Wilcox <willy@infradead.org>
cc: linux-cifs@vger.kernel.org
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Jakub Kicinski:
"Including fixes from Bluetooth, NFC and Netfilter.
Every week in this release is record-setting for number of posted
patches. It doesn't seem like we're creating any regressions with all
these fixes, three 'Fixes' tags here point to 7.2 commits but none are
true regression fixes. We're trying to keep the count down,
nonetheless.
Previous releases - regressions:
- net: don't require the hwtstamp NDOs when a PHY provides
timestamping
- ipv6: fix dst leak for uncached routes
- vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE
packets
Previous releases - always broken:
- packet: use ubuf_info completion for TX_RING packets
- arp: terminate device name before lookup
- ipv6: do not let ipv6_find_hdr() return an offset past the packet
end
- udp: remove a disconnected socket from the 4-tuple hash table
- sctp: discard the rest of the packet on a stale-cookie error
- eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged
eswitch"
[ And lots of other random network driver fixes ]
* tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (189 commits)
tcp: prevent collapsing skbs across boundary in rtx queue
vlan: ensure sufficient headroom in vlan_dev_hard_header()
net/sched: sch_teql: fix shadowed err in __teql_resolve()
bridge: check llc_mac_hdr_init() return value in br_send_bpdu()
llc: fix skb UAF and leaks on llc_mac_hdr_init() failure
llc: reserve device headroom for allocated frames
gve: DQO: reject TSO packets with an out of range MSS
gve: fix TX drop when GSO MSS is too small for hw
gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO
net: flush skb_defer_nodes in dev_cpu_dead()
net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails
af_packet: fix integer overflow in prb_calc_retire_blk_tmo()
tipc: Fix a data race on mon->peer_cnt in mon_timeout()
net: phy: intel-xway: workaround 100BASE-TX Link-Up issue
net/smc: fix UAF on lgr list traversal in smcr_port_err()
net/rds: size a connection's path set by the transport it ends up with
nfp: hold IPsec RX state under the XArray lock
net: ena: fix MMIO read buffer leak on probe failure
net: ena: fix PHC cleanup on probe failure
net/sched: act_ct: fix helper UAF due to extensions realloc
...
|
|
Pull bpf fixes from Alexei Starovoitov:
- Fix bpf_skb_change_tail() to drop the checksum offload instead of
rejecting the trim of CHECKSUM_PARTIAL skbs (Daniel Borkmann)
- Add KF_PERFMON kfunc flag and require CAP_PERFMON for kfuncs that
read arbitrary memory and for untrusted read-only memory reads
(Daniel Borkmann)
- Clear scalar delta on narrowing stack spill (Daniel Borkmann)
- Set up the frame pointer for the exception callback in arm64 JIT, and
zero-fill other CPUs when BPF_F_CPU update creates a per-cpu hash
element (Donggeun Yoo)
- Various fixes (Emil Tsalapatis):
- Fix bounds check underflow for skb-backed dynptrs
- Fix rx_queue_mapping context access code generation in bpf_sock
- Reject packet pointer arguments to subprogs that may mutate the
packet
- Reject ALU instructions that see arena and non-arena operands on
different code paths
- Fix copied_seq double-counting on sockmap self-redirect
(Geliang Tang)
- Fix divide-by-zero in btf_struct_walk() on a flexible array of
zero-sized elements, fix out-of-bounds read of rtt_min in sock_ops
(Jiayuan Chen)
- Fix bpf_sock_destroy() out-of-bounds read of sk_protocol on TIME_WAIT
and request socks, and sleeping under RCU when destroying a listener
with pending children (Jiayuan Chen)
- Fix JEQ/JNE with immediate operand in MIPS32 JIT and missing zero
extension of BSWAP 16/32 in MIPS64 JIT (Johan Almbladh)
- Avoid soft lockup in htab lookup[_and_delete] batch operations on
large maps (Jose Fernandez)
- Various fixes (Kumar Kartikeya Dwivedi):
- Verify global subprogs in each sleepability context they are
called from
- Make post-verification instruction rewrites killable
- Preserve packet pointer displacement in regsafe()
- Apply CO-RE relocations before subprogram validation, restrict
CO-RE poisoning to relocatable instructions, and reject truncated
ldimm64 CO-RE relocations in libbpf
- Assign lock identity to callback map values
- Compare stack frames in regs_exact()
- Bound ownership depth through local kptrs and graph roots
- Fix u32 overflow in map batch operations when the map size exceeds
4GB (Masoud Aghasi)
- Fix UAF in bpf memalloc due to concurrent consumption of ttrace lists
in alloc_bulk() (Pu Lehui)
- Allow gotox as the terminal instruction of a program or a subprogram
(Siddharth Chintamaneni)
- Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL, skip unsettled links
in link iterator, and reject dev-bound-only programs on other devices
(Weiming Shi)
- Reject non-negative stack offsets in stack_slot_obj_get_spi()
(Xu Yunxiang)
- Check params size before reading reserved fields in
bpf_crypto_ctx_create() (Yuqi Xu)
- Reject max_entries > INT_MAX in sock_map_alloc() (Zhao Gongyi)
- Use a 32-bit compare in xsk_map_gen_lookup() (Zhiling Zou)
- Use kvfree() in xdp_test_run_teardown() (Zhixing Chen)
* tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: (58 commits)
selftests/bpf: Test per-cpu initialization of a BPF_F_CPU created element
bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element
bpf: Fix BSWAP 32 and 16 on MIPS64
bpf: Fix immediate JMP JEQ/JNE on MIPS32
bpf: Reject dev-bound-only programs on other devices
bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc
selftests/bpf: Test for mixed arena/nonarena code paths
bpf: Prevent variable arena/non-arena register contents
selftests/bpf: Test rejection of pkt args to mutating subprogs
bpf: Reject pkt arguments in mutating subprogs
selftests/bpf: Add selftests for rx_queue_mapping context access
bpf: Fix bpf_sock context code generation
selftests/bpf: Test dynptr slices past end of skb
bpf: Fix bounds check for skb-backed dynptrs
selftests/bpf: Reject iterator destruction through fp+0
bpf: Reject non-negative offsets in stack_slot_obj_get_spi()
bpf: Check params size before reading reserved fields
selftests/bpf: Check local object ownership depth
bpf: Bound ownership depth through local kptrs and graph roots
selftests/bpf: Cover frame changes in bounded loops
...
|
|
|
|
The xpt_reserved counter exists for UDP socket-buffer back-pressure.
svc_udp_has_wspace() is the only has_wspace implementation that
consults it, so on TCP and RDMA the counter is maintained and never
read. svc_handle_xprt() adds to it once per RPC. svc_reserve()
shrinks it again on each call from svc_process_common(), from
svc_xprt_release(), and from each proc function that calls
svc_reserve_auth(). Every shrinking call also runs
svc_xprt_resource_released(), which issues an smp_mb() and can
enqueue the transport.
Add an xcl_flags field to svc_xprt_class and set
SVC_XPRT_FLAG_WSPACE_RESERVE on the UDP class. Gate the xpt_reserved
accounting on that flag.
After the change, svc_reserve() no longer calls
svc_xprt_resource_released() on TCP and RDMA. Two paths still cover
that enqueue. svc_xprt_release() reaches the helper through
svc_xprt_release_slot(), and svc_xprt_received() enqueues a transport
whose XPT_DATA remains set.
Link: https://patch.msgid.link/20260828135036.796842-4-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
enum nfs_stat in uapi/linux/nfs.h collects the on-the-wire status
codes for every NFS version under version-agnostic NFSERR_* names.
The in-kernel NFS client and server reference these names
internally, which drags uapi/linux/nfs.h into many translation
units that need nothing else from that header.
xdrgen conversion will provide spec-based definitions of the NFS
status enum constants. Replacing the existing internal references
with version-specific names requires a version-specific spelling
for each code that NFSD uses.
linux/nfs4.h already supplies the NFSv4 codes, but NFSv3 has no
counterpart header, and not every NFSv3 code has an NFS4ERR_*
spelling: NFSERR_NOT_SYNC (10002) and NFSERR_REMOTE (71) have no
NFSv4 equivalent at all, and NFSERR_JUKEBOX (10008) appears in
NFSv4 as NFS4ERR_DELAY.
Introduce the NFSv3 status codes as NFS3ERR_* in linux/nfs3.h so a
subsequent patch can respell NFSD's status definitions without
relying on uapi/linux/nfs.h.
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Link: https://patch.msgid.link/20260826194444.148243-1-cel@kernel.org
Signed-off-by: Chuck Lever <cel@kernel.org>
|
|
perf_clear_branch_entry_bitfields() clears the bitfields of struct
perf_branch_entry one by one and leaves from/to alone, since callers
overwrite those straight away. The list has to be kept in sync with the
struct by hand and has already fallen behind: new_type and priv were
added to perf_branch_entry and never added here.
Only BRBE writes those two, and neither for every record.
brbe_set_perf_entry_type() leaves new_type alone for a branch type it
does not recognise, and priv is not set for source-only records.
arm_pmuv3.c allocates the per-CPU branch stack with kmalloc(), so such a
record reaches userspace with whatever the slot held: uninitialised
kmalloc() data on the first pass over the buffer, the previous record's
values after that. Nothing under arch/x86/events/ writes either field,
so only arm64 is affected.
Assign the whole entry at each site instead. Everything not named is
then zero, and there is no list to keep in sync. The bitfields add up to
exactly 64 bits, so the struct has no padding to leave undefined.
perf_clear_branch_entry_bitfields() has no callers left, so remove it.
perf_entry_from_brbe_regset() assigns an empty literal instead, since it
fills from/to conditionally. PERF_BR_SPEC_NA is 0, so dropping the
explicit spec assignment changes nothing.
Fixes: b190bc4ac9e6 ("perf: Extend branch type classification")
Fixes: 5402d25aa571 ("perf: Capture branch privilege information")
Suggested-by: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Tested-by: Yifan Wu <wuyifan50@huawei.com>
Link: https://patch.msgid.link/20260810133540.1947118-4-puranjay@kernel.org
|