summaryrefslogtreecommitdiff
path: root/include/linux
AgeCommit message (Collapse)Author
5 dayssched/debug: Add migration stats due to non preferred CPUsShrikanth Hegde
Add a new per-task stat, - nr_migrations_cpu_non_preferred: number of push migrations while the CPU is non-preferred. Since this new stat is per-task, it changes only /proc/<pid>/sched. It doesn't update /proc/schedstat. Hence increasing the schedstat version is not necessary. Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://patch.msgid.link/20260928053728.797539-10-sshegde@linux.ibm.com
5 dayscpumask: Introduce cpu_preferred_maskShrikanth Hegde
Provide the preferred CPU infrastructure. Define get/set macros which could be used to get/set CPU state as preferred. CONFIG_PREFERRED_CPU will be selected by the driver which handles steal time values. It is going to set/clear preferred CPU state. This driver will be called steal_governor and it is introduced in subsequent patches. It periodically computes the steal ratio and decides on preferred CPU state. A CPU is set to preferred when it becomes active. Later it may be marked as non-preferred depending on steal ratio by the steal_governor. Always maintain design construct of preferred is subset of active. i.e. preferred ⊆ active ⊆ online ⊆ present ⊆ possible With CONFIG_PREFERRED_CPU=n, ensure set_cpu_preferred is a nop and get method returns the active state in that case. Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://patch.msgid.link/20260928053728.797539-5-sshegde@linux.ibm.com
5 dayscpumask: Introduce cpumask_intersects_andShrikanth Hegde
Introduce bitmap_intersects_and() to determine whether the intersection of three bitmaps is non-empty. Unlike cpumask_first_and_and(), this returns immediately when an intersecting word is found and does not calculate the first matching bit. Add cpumask_intersects_and() as the corresponding cpumask wrapper. A subsequent patch uses the helper to determine whether a task has a CPU that is present in its affinity mask, the preferred CPU mask, and task possible CPU mask. Suggested-by: Yury Norov <yury.norov@gmail.com> Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com> Reviwed-by: Yury Norov <yury.norov@gmail.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Reviewed-by: Yury Norov <yury.norov@gmail.com> Link: https://patch.msgid.link/20260928053728.797539-3-sshegde@linux.ibm.com
5 dayssched/cputime: Add kcpustat_field_total helperShrikanth Hegde
Provide a new helper function which sums up a given type of cpustat over a specified cpumask. This allows the caller's code to be simpler and avoids duplication. For example, subsequent patch in the steal governor use this exact same pattern when calculating steal time. Number of cpus can be derived from cpumask_weight() where necessary. Suggested-by: Yury Norov <yury.norov@gmail.com> Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Reviewed-by: Yury Norov <ynorov@nvidia.com> Reviewed-by: Mete Durlu <meted@linux.ibm.com> Acked-by: Frederic Weisbecker <frederic@kernel.org> Link: https://patch.msgid.link/20260928053728.797539-2-sshegde@linux.ibm.com
6 daysMerge branch 'i2c/i2c-fixes' into i2c/i2c-nextAndi Shyti
6 daysMerge tag 'v7.3-rc5' into driver-core-nextDanilo Krummrich
We need the driver-core fixes in here as well to build on top of. Signed-off-by: Danilo Krummrich <dakr@kernel.org>
6 daysMerge tag 'sched-urgent-2026-09-27' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull scheduler fixes from Ingo Molnar: - Fix LLC mis-scheduling bugs (Tim Chen, Lu Wang) - Fix cache-grouping related scheduling statistics UAF bugs (Tim Chen) - Skip kernel threads for cache aware scheduling to rubustify the code (Chen Yu) - Refresh LLC capacity across CPU hotplug, to fix capacity underestimation bug (Davi Chaves Azevedo) - Account PSI IRQ time to the execution context, not the scheduling context, to fix proxy scheduling accounting bug (Zhan Xusheng) * tag 'sched-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: sched/core: Account PSI IRQ time to the execution context, not the scheduling context sched/cache: Refresh LLC capacity across CPU hotplug, to fix capacity underestimation bug sched/cache: Skip kernel threads for cache aware scheduling to rubustify the code sched/cache: Introduce task_struct->sched_cache_grp to fix UAF sched/cache: Decouple sched_cache_group from mm to fix UAF sched/cache: Honor migrate_llc_task semantics in active load balance, to fix LLC mis-scheduling bug sched/cache: Keep nr_pref_llc_running in the runnable domain, to fix LLC mis-scheduling bug
6 daysMerge tag 'perf-urgent-2026-09-27' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull perf events fixes from Ingo Molnar: - Fixes for KVM guest PEBS virtualization (Sean Christopherson) - Fixes for various Intel PMUs related to PEBS data-source (Dapeng Mi) - Fix Intel Panther Cove event scheduling constraints (Dapeng Mi) - Fix Intel DMR/NVL OMR extra registers event scheduling (Dapeng Mi) - Rename two confusingly named PMU attributes (Dapeng Mi) - Fix a refcount leak in attach_perf_ctx_data() (Namhyung Kim) - Fix NULL pointer dereference crash in __perf_pmu_sched_task() (Puranjay Mohan) - Fix CPU-wide event scheduling (Puranjay Mohan) - Fix x86 LBR branch entry generation (Puranjay Mohan) * tag 'perf-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: perf/core: Fill branch entries with a single assignment perf/core: Run sched_task() for PMUs with only CPU-wide events perf/core: Fix NULL pmu_ctx passed to pmu->sched_task() perf/core: Fix a refcount leak in attach_perf_ctx_data() perf/x86/intel: Rename NVL offcore_rsp attribute to offmodule_rsp perf/x86/intel: Rename DMR offcore_rsp attribute to offmodule_rsp perf/x86/intel: Fix precise OMR event scheduling for DMR/NVL perf/x86/intel: Constrain Panther Cove UOPS_DISPATCHED events to PMCs 0-3 perf/x86/intel: Delete dead NVL PEBS data-source initcall perf/x86/intel: Fix Panther Cove PEBS data-source snoop states perf/x86/intel: Remove incorrect Panther Cove PEBS data-source constraints perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints perf/x86/intel: Update arw_latency_data() mem-op direction handling perf/x86/intel: Fix DKT PEBS load/store direction for latency events, to fix sample classification perf/x86/intel: Fix CMT PEBS load/store direction for latency events, to fix sample classification perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs() perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) if PEBS is unused perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has PEBS isolation, to fix stuck PEBS_ENABLED perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted PERF_GLOBAL_CTRL bits
6 daysMerge branch 'perf/core' into perf/merge, to assist CI testingIngo Molnar
Signed-off-by: Ingo Molnar <mingo@kernel.org>
7 daysMerge tag 'ata-7.3-rc5' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux Pull ata fixes from Niklas Cassel: - Extend the quirk "no LPM on ATI" quirk, that is currently only applied for Samsung drives, to include AMD controllers as well. The AMD AHCI controllers are newer versions of the ATI AHCI controllers, and these controllers still have LPM issues with Samsung drives - LPM works with drives from other vendors (me) - Fix errors in the libata.force parameter documentation (me) - Verify the sense data descriptor lengths for ATA PASS-THROUGH command, so that a malicious device cannot write past the buffer length (Matthias) - Mention the libata for-next branch in MAINTAINERS such that the git ls-remote command done by get_maintainer.pl --self-test=scm can verify it (Matthias) * tag 'ata-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux: MAINTAINERS: name the libata/linux for-next branch ata: libata-scsi: bound the ATA passthru sense descriptor writes ata: libata: Correct libata.force parameter documentation ata: libata-core: Extend Samsung LPM quirk to AMD controllers
7 daysMerge tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvmLinus Torvalds
Pull kvm fixes from Paolo Bonzini: "Arm: - Invalidate the ITS translation cache when the guest changes the base address of the ITS tables (Fuad Tabba) - Skip saving ITS devices with device IDs that are out-of-bounds rather than failing the entire ITS save ioctl (Fuad Tabba) - Close race between VM teardown and invalidations of nested MMUs when handling MMU operations that are allowed to block (Lorenzo Stoakes) - Various fixes for the handling of the host's untrusted SVE configuration in pKVM (Fuad Tabba) - Make sure that empty SMCCC ranges based at 0 are rejected by the kvm_smccc_set_filter() (Karl Mehltretter) - Revoke the host mapping for pKVM's private stack pages, along with a new sanity check that all mappings in the hyp's private VA range have been correctly marked as hyp-owned (Fuad Tabba) - Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that concurrent vCPU initialization cannot relocate in-use MMUs. Defer the freeing of shadow stage-2 MMUs to the point that no other users (e.g. MMU notifier) could reference them (Marc Zyngier) - Drop useless WARN when rejecting an unsupported ioctl for pKVM (Fuad Tabba) - Fix the steal_time selftest to install correctly-sized mappings for non-4K hosts (Sebastian Ott) - Correct mapping of fine-grained trap for GCSPOPX instruction (Mark Brown) - Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit guests (Karl Mehltretter) RISC-V: - Synchronize hrtimer during VCPU teardown - Fix the conversion between vsip and hvip values - Serialize IMSIC attributes with vCPU migration - Release unused page after MMU invalidation - Propagate interrupted G-stage faults to KVM user-space as EINTR - Fix nested acceleration hfence entry update order - Fix sdata leak and stale snapshot_addr in snapshot_set_shmem - Preserve firmware counter value across PMU counter stop/start - Report PMU snapshot write failure to the guest - Fix perf-backed counter accounting across PMU stop and read - Correctly propagate error of a hart status SBI call s390: - Ensure that accesses through kvm_arch_set_irq_inatomic mark as dirty the pages that contain indicator and summary bits - Fix compile warning for kvm_s390_update_cmma_dirty() - Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to userspace - Move s390_kvm_mmu_commit_memory_region() into s390_kvm_mmu_prepare_memory_region() so that it can fail instead of WARN - Add missing srcu in kvm_s390_set_irq_state() - Fix potential races in storage functions - Fix race in _destroy_pages_crste() - Fix issues in the handling of KVM interrupt and page resources, when a queue that is assigned to a mediated device (mdev) is removed from the host's AP configuration - Fix loop condition in uv_find_secrets - Prevent potential out-of-bounds read x86: - Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs as valid on AMD - Fix a regression in the hardware disable selftest where it checked the wrong macro when detecting glibc support (breaks at least musl) - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially important for KVM_BUG_ON() flows, which often guard more dangerous bugs - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug where KVM would let userspace run a broken setup with stale vmcs12 pages - Fix a class of bugs where KVM would fail to fill kvm_run exit fields if getting nested pages failed - Treat reserved entries in the memory attributes xarray as "no attributes", to fix false positives when checking for mixed attributes - Fix memcg accounting for the memory attributes xarray (the xarray library subtly requires the xarray to be configured for accounting upfront; the gfp flags taken at runtime are used only rarely) - Don't pre-reserve xarray entries when storing empty attributes, as storing NULL must not require memory allocation (KVM and other subsystems heavily rely on this behavior) - Fix a memory leak and a cache maintenance issue related to doing intra-host migration on an SEV guest" * tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (54 commits) KVM: SEV: Do cache maintenance on the source VM during intra-host migration KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV KVM: Don't pre-reserve xarray entries when storing empty/NULL attributes KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg KVM: Don't treat reserved xarray entries as having memory attributes KVM: x86: Fill kvm_run exit fields in common get_nested_state_pages() error paths KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails KVM: arm64: Fix AArch32 DBGBXVR<n> handling KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP KVM: selftests: fix steal_time for arm64 with host page size > 4K KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction KVM: arm64: nv: Fix life cycle of the nested_mmus array KVM: arm64: Check every private mapping is hyp-owned at pKVM init KVM: arm64: Move the private VA allocation cursor to __io_map_next KVM: arm64: Match hyp text by physical address in fix_host_ownership() KVM: arm64: Transfer the hyp stack pages out of the host stage-2 KVM: arm64: selftests: Test empty SMCCC filter range at base 0 KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0 KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2 ...
7 daysbus: mhi: host: Fix typo "intented" in commentHemanth Selam
Correct "intented" to "intended", reported by scripts/checkpatch.pl using the misspelling list in scripts/spelling.txt. Only touches comments, no code changes. Assisted-by: LLM Signed-off-by: Hemanth Selam <hemanth.selam@gmail.com> Signed-off-by: Manivannan Sadhasivam <manivannan.sadhasivam@oss.qualcomm.com> Link: https://patch.msgid.link/20260904124731.9011-1-hemanth.selam@gmail.com
7 daysRevert "KVM: Check for duplicate vcpu_id as early as possible"Sean Christopherson
Now that KVM uses kvm_get_vcpu_by_id() to check for an existing vCPU ID before doing any meaningful work, which was made possible by holding kvm->lock for the entirety of vCPU creation, revert the now-redundant "early" vCPU ID tracking. The claims about the impact of kvm->vcpu_ids on the memory footprint were a wee bit wrong: the worst case scenario isn't 256 bytes per VM, it's 256 "unsigned longs" per VM, i.e. 2048 bytes per VM. Increasing the size of "struct kvm" by 2048 nearly doubled the total size on many architectures, and tripped x86's KVM_SANITY_CHECK_VM_STRUCT_SIZE, which was added to detect this *exact* scenario, where a single change significantly increased the size of "struct kvm". I.e. attempting to build KVM with CONFIG_DEBUG_KERNEL=n fails on x86 (the build failures got missed because all build bots apparently test only CONFIG_DEBUG_KERNEL=y kernels, and maintainers' test flows were similarly lacking). This reverts commit 97d65b544f48b2ee49f6aea32145e3e7969955dc. Fixes: 97d65b544f48 ("KVM: Check for duplicate vcpu_id as early as possible") Reported-by: Jean-Christophe Guillain <jean-christophe@guillain.net> Closes: https://lore.kernel.org/all/56a4bc35ee605588b7cc36c8e45c12b5f3b506cb.camel@guillain.net Reported-by: Paweł S <spawel523@gmail.com> Closes: https://lore.kernel.org/all/CABD%3DWFOS4j4hDv%2BpW-eEM9HAM2q2GY_iYdAG%2BqvYcUEinUrcQQ@mail.gmail.com Tested-by: Jean-Christophe Guillain <jean-christophe@guillain.net> Signed-off-by: Sean Christopherson <seanjc@google.com> Tested-by: Naveen N Rao (AMD) <naveen@kernel.org> Message-ID: <20260921174445.911676-7-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
7 daysMerge tag 'kvm-x86-fixes-7.3-rc5' of https://github.com/kvm-x86/linux into HEADPaolo Bonzini
KVM fixes for 7.3-rcN - Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs as valid on AMD. - Fix a regression in the hardware disable selftest where it checked the wrong macro when detecting glibc support (breaks at least musl). - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially important for KVM_BUG_ON() flows, which often guard more dangerous bugs. - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug where KVM would let userspace run a broken setup with stale vmcs12 pages. - Fix a class of bugs where KVM would fail to fill kvm_run exit fields if getting nested pages failed. - Treat reserved entries in the memory attributes xarray as "no attributes", to fix false positives when checking for mixed attributes. - Fix memcg accounting for the memory attributes xarray (the xarray library subtly requires the xarray to be configured for accounting upfront; the gfp flags taken at runtime are used only rarely). - Don't pre-reserve xarray entries when storing empty attributes, as storing NULL must not require memory allocation (KVM and other subsystems heavily rely on this behavior).
8 daysMerge ras/edac-urgent into for-nextBorislav Petkov (AMD)
* ras/edac-urgent: (1178 commits) EDAC/altera: Fix use-after-free in error paths EDAC/altera: Fix memory leak on dci allocation failure EDAC/altera: Drop __init from ECC setup paths for re-probe safety EDAC/altera: Do not allow driver unbinding Linux 7.3-rc4 net: qrtr: resend HELLO on MHI resume i2c: qcom-cci: fix device_node refcount leak in cci_probe()/cci_remove() i2c: qcom-geni: release DMA channels on probe error i2c: imx: release DMA channels on probe error i2c: at91: release DMA channels on remove and probe error posix-cpu-timers: Prevent freeing a timer which is queued on the expiry list x86/build/64: Prevent native builds from generating EGPR use watchdog: da9063: fix suspend/resume handling of HW_RUNNING watchdog soc: samsung: exynos-pmu: fix use-after-free of interrupt generator node selftests/x86: Check signal state for rejected software interrupts x86/fred: Reconstruct the #GP context for rejected INT instructions sched/core: Avoid false migration warning for proxy donors perf: Fix null pointer access in is_include_guest_event() x86/microcode/intel: Reject problematic loading on Granite Rapids systems cifs: Fix server use-after-free in cifs_chan_skip_or_disable() ... Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
8 daysnet: ethtool: generate RSS keys that spread flows over all queuesEric Dumazet
netdev_rss_key_fill() returns a key made of uniformly random bytes. That is not enough, because the Toeplitz hash is linear over GF(2). Walking the hash input MSB first, each set bit contributes a 32-bit sliding window of the key, and hardware indexes the indirection table with the low order bits of the result. Only the tail of each window therefore reaches the queue index: v(i) = key bits [i + 32 - q .. i + 31] with q = log2(number of RX queues). Consecutive input bits give windows overlapping in q - 1 positions, so the q vectors belonging to the q lowest bits of a header field form a Toeplitz matrix built from 2 * q - 1 key bits, rather than q * q independent ones. Over GF(2) a random Toeplitz matrix is singular with probability exactly 1/2, whatever its size. When it is singular, flows differing only in the low order bits of that field cannot reach all the queues. This is not theoretical: a burst of connections draws ephemeral ports from a narrow range, and on one affected host only 4 of the 16 RX queues received any traffic at all, until its key was replaced. Keep drawing the key at random, since it is a secret that stops a remote attacker from steering flows onto a single queue, but force the handful of bits that decide this. Writing d[t] for key bit (lsb + 31 - t), where lsb is the position of the least significant bit of a field in the hash input, the matrices of all the q values up to 8 are non singular if and only if d[2 * i] = 1 ^ d[i] ^ d[i + 1] ^ ... ^ d[2 * i - 1] The odd positions stay free, so this is a one pass fixup rather than a search. It is in fact a bijection from those free positions onto the set of the values having the property, so the key stays uniformly distributed over that set and the whole cost is 8 bits of entropy per position. Searching for such a key by rejection would not have been an option: a freshly drawn one has the property everywhere with probability 2^-1008. Apply this at every 16-bit aligned position of the key, rather than at the offsets of the 2-tuple and 4-tuple layouts only. The core does not get to know what a given NIC hashes. Hardware may select the bytes it feeds to Toeplitz out of a header window with a bitmap, and hash an encapsulated header: for PSP over UDP over IPv6 it can pick the outer addresses and the inner TCP ports, which sit 78 bytes into the frame and read key bits well past the 40 bytes an IPv6 4-tuple needs. Hashed fields are 16 bits wide at the smallest and are not expected to straddle that grid, so covering the grid covers the layouts that were never written down, at no cost in code. This spends 8 bits of entropy per position, 1008 bits out of the 2048 bits of netdev_rss_key, leaving 1040 bits. 8 is also the largest usable bound, as each q constrains 2 * q - 1 bits and anything larger would make the ranges of two adjacent positions overlap. What the fixup leaves behind is visible structure: 8 of every 16 bits are derived from the 8 others, so a 16-bit aligned word of the key takes only 2^8 values and about 27 of the 128 words of netdev_rss_key duplicate another one. That much is forced rather than an artefact of this implementation, 1040 bits spread over 128 words being a little over 8 bits each, but it has one consequence worth removing. Two 32-bit windows a whole number of words apart now collide with probability 2^-16 instead of 2^-32, and two input bits reading the same window are indistinguishable to the hash, since flipping both of them leaves it unchanged. Over 500 keys, 23% of them had such a pair, where a uniformly random key has one with probability 2^-11. So draw another key when that happens. Four out of five pass. Which distances to look at follows from where the structure is: covering the multiples of 16 leaves 1.2e-3 expected colliding pairs per key, still 2.6 times the 4.6e-4 of a plain random key, and almost all of that excess sits at a distance of 8 modulo 16. Covering every multiple of 8 brings the total down to 4.2e-4, below what a plain random key gives over all distances, and within a few percent of the 4.0e-4 it gives over the distances that are left. Two windows a multiple of 8 bits apart are two windows at the same offset modulo 8, so this is a handful of pairwise sweeps rather than one pass over the key per distance. DO_ONCE() runs the generator under a spinlock with hard IRQs disabled, so keep the windows of a class in an array rather than recomputing both sides of every pair: 116 us instead of 254 us for the worst case, a 2048-bit key with no collision anywhere, at a cost of 512 bytes of stack. The shared key is fixed up once and every driver prefix inherits both properties. Move netdev_rss_key and netdev_rss_key_fill() from net/ethtool/ioctl.c to net/ethtool/common.c, and move the netdev_rss_key declaration out of include/linux/netdevice.h into net/core/dev.h. Checked against an independent Toeplitz implementation: for every 16-bit aligned position and every q in 1..8, an aligned block of 2^q consecutive values of a field ending there lands on the 2^q queues exactly once each. Over 20 random keys that is 20160 checks, which 50.1% of plain random keys fail and none of the generated keys do, for an average of 502 rewritten bits out of 2048. Signed-off-by: Eric Dumazet <edumazet@google.com> Link: https://patch.msgid.link/20260922163458.3900996-3-edumazet@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
8 daysfirewire: core: add __counted_by_ptr attribute to struct fw_deviceBill Wendling
The GCC and Clang attribute '__counted_by_ptr' associates a pointer field with an integer field holding its element count. This enables runtime bounds checking by KASAN and UBSAN to prevent out-of-bounds accesses to the pointer. In 'struct fw_device' (defined in 'include/linux/firewire.h'), the field 'config_rom' is a pointer to the device's Configuration ROM data, and its associated element count is stored in 'config_rom_length'. Additionally, update the existing KUnit test in 'drivers/firewire/device-attribute-test.c' where 'config_rom_length' was incorrectly initialized with the byte size ('sizeof(simple_avc_config_rom)') instead of the element count ('ARRAY_SIZE(simple_avc_config_rom)'). This ensures the '__counted_by_ptr' bounds-checking annotation does not trigger any false-positive panics or compile/runtime checks. Cc: codemender-patching+linux@google.com Assisted-by: LLM Signed-off-by: Bill Wendling <morbo@google.com> Link: https://lore.kernel.org/r/20260925212552.125652-1-morbo@google.com Signed-off-by: Takashi Sakamoto <o-takashi@sakamocchi.jp>
8 daysfirewire: core: add __counted_by_ptr to struct fw_packetBill Wendling
In 'struct fw_packet', the 'payload' field points to a buffer of size 'payload_length' bytes. To enable compiler bounds checking (via KASAN and __builtin_dynamic_object_size), add the '__counted_by_ptr' attribute to the 'payload' pointer, associated with the 'payload_length' count. A thorough analysis of all allocation and initialization points for 'struct fw_packet' was conducted. The pointer and its length are assigned together. Because the count field is always correctly initialized before the pointer is accessed, the '__counted_by_ptr' attribute is fully safe and will not cause runtime false positives or panics. Cc: codemender-patching+linux@google.com Assisted-by: LLM Signed-off-by: Bill Wendling <morbo@google.com> Link: https://lore.kernel.org/r/20260925210912.113052-1-morbo@google.com Signed-off-by: Takashi Sakamoto <o-takashi@sakamocchi.jp>
8 daysbpf: Add file descriptor interface for program streamsKumar Kartikeya Dwivedi
The existing BPF_PROG_STREAM_READ_BY_FD command only supports polling a program stream through repeated bpf() calls. It cannot block for new data or integrate with poll-based event loops. Add BPF_PROG_STREAM_OPEN to return a read-only, close-on-exec file descriptor for a selected program stream. Reads block by default and BPF_F_STREAM_NONBLOCK, the only accepted flag, provides non-blocking behavior. poll reports readable data and reports hangup once the program has been freed. Like pipes and sockets, the descriptor is not seekable and lseek fails with ESPIPE. A stream descriptor deliberately does not retain the program. Move each stream into a separately refcounted allocation so program teardown can mark it dead and wake descriptor users while outstanding descriptors drain buffered data safely. Readers sample the dead flag before looking for data, so EOF is reported only when the stream was already dead before it was found empty; data published right before teardown is never skipped. Only programs loaded through BPF_PROG_LOAD get streams. Classic BPF filters, JIT subprograms and shim programs never write to one, and kernel-side writers already resolve a subprogram to its main program, so those programs no longer carry stream state. Readiness needs its own counter. Stream capacity is charged before allocation and before an element is published to the stream log, so using that reservation as the read and poll condition can report readable data while no element exists: a blocking reader retries instead of sleeping and a lone non-blocking reader can see POLLIN followed by EAGAIN. Publish bytes with release ordering after adding elements to the lockless log, use acquire loads before consuming them or reporting readiness, limit each read to its readable snapshot and subtract only bytes actually copied. This keeps the aggregate count correct even when concurrent publishers update it out of publication order. With several readers on one stream, readiness remains advisory, as it is for pipes. The capacity counter is kept solely for enforcing the stream size limit. Wakeups are always deferred through irq_work. Stream writers run in whatever context the program runs in: NMI context for perf_event programs, sections with interrupts disabled inside bpf_spin_lock or rqspinlock critical sections since bpf_stream_vprintk() is KF_SPINLOCK_SAFE, and tracing programs attached anywhere in the kernel, including inside the wait queue and epoll code itself. Waking waiters directly from there can deadlock, and no cheap context check covers every case: on PREEMPT_RT, spinlock_t sections do not disable interrupts, so in_nmi() or irqs_disabled() cannot tell such a program apart from a benign one. Queue an irq_work item instead, as bpf_ringbuf does. Queue it only when a publication turns an empty stream readable. Readers block and pollers wait only after finding the stream empty, and the readable count never drops below zero because each read is bounded by its snapshot, so the first publication after such an observation is the one that makes the count positive, and it is the one that queues the wakeup. Publications into a stream that already holds data raise no interrupt, so a program that prints while nobody drains its stream pays for a single irq_work until the stream is emptied again. This matches bpf_ringbuf, which notifies only once the consumer has caught up. Blocking readers and level-triggered pollers re-check the readable count before waiting, so they cannot miss data, and edge-triggered epoll consumers drain until EAGAIN before waiting again, as epoll(7) requires. Synchronize pending work before releasing the final stream reference so the callback cannot outlive the stream, but only when the work was ever queued: irq_work_sync() waits for an RCU grace period on PREEMPT_RT and on architectures without an irq_work interrupt. Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260925045536.1480933-3-memxor@gmail.com
8 daysbpf: Retarget indirect jump targets across prologue prependsDaniel Borkmann
A few patches do not expand in place, they prepend: bpf_convert_ctx_accesses() puts the ctx save for gen_epilogue and the instructions of gen_prologue in front of insn 0, and bpf_do_misc_fixups() prepends the may_goto counter init at the start of every subprog that uses may_goto. They copy the original instruction of 'tgt_idx' into the last slot of the patch buffer, so it now lives at tgt_idx + delta, and call adjust_jmp_off() to move direct branches from tgt_idx to tgt_idx + delta. Nothing does the same for indirect branches, so a BPF_MAP_TYPE_INSN_ARRAY slot that named tgt_idx keeps naming tgt_idx, which is now the first prepended instruction, and insn_aux_data[tgt_idx].indirect_target makes the JIT emit the landing pad there. Add a __bpf_patch_insn_data() variant and pass BPF_PATCH_MOVE_TARGET at the affected call-sites to fix up the delta. The poke descriptors of direct tail calls are shifted the same way, and the original instruction is not searched for by content in that mode, as it is the last slot by construction. Fixes: 493d9e0d6083 ("bpf, x86: add support for indirect jumps") Fixes: 07ae6c130b46 ("bpf: Add helper to detect indirect jump targets") Reported-by: Nicholas Carlini <npc@anthropic.com> Suggested-by: Nicholas Carlini <npc@anthropic.com> Signed-off-by: Daniel Borkmann <daniel@iogearbox.net> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Acked-by: Anton Protopopov <a.s.protopopov@gmail.com> Acked-by: Eduard Zingerman <eddyz87@gmail.com> Link: https://patch.msgid.link/20260925175244.1136329-11-daniel@iogearbox.net
8 daysbpf: Cache the jump table of a subprogram during CFG discoveryDaniel Borkmann
create_jt() builds the jump table of the subprogram containing a gotox by copying out and sorting every insn_array map of the program, and it does so once per gotox instruction. The cost is therefore the number of gotox instructions times the number of entries in all of the maps. A program of 4003 instructions with 2000 gotox and one 500k entry map holding two distinct targets has 4000 indirect jump edges, 0.4% of the limit, and takes 343s to be rejected. The map costs next to nothing to prepare, as an unset entry is already a valid target. At the insn limit, with a single 1M entry map, the same shape extrapolates to 41 hours. All gotox instructions of a subprogram share the same jump table, so build the table of every subprogram in a single pass over the maps and let every gotox use it in place: bpf_insn_successors() resolves a gotox through its containing subprogram, and the tables are kept until the verifier environment is torn down, so nothing is copied into the instruction aux data. The edge accounting in visit_gotox_insn() moves to the first visit of the gotox, marked by the BRANCH bit of its CFG state, and a revisit returns right away since the first one already pushed every target that still needed exploring. The only user left of insn_aux_data[].jt is then the table which visit_abnormal_return_insn() allocates for tail_call and ld_{abs,ind} insns, so that bpf_insn_successors() reports the hidden exit from their subprogram. Both are recognisable by their opcode, so derive that edge in bpf_insn_successors() from the exit_idx of the containing subprogram instead, again leaving it out when the subprogram has no exit, and drop the field along with bpf_clear_insn_aux_data(). The instruction aux data then owns no allocation, thus nothing needs to be freed when insns are removed or the verifier environment is torn down. Subprograms removed as dead code free their table in adjust_subprog_starts_after_remove(), which also clears the slots that the compaction of subprog_info vacates, as they still hold copies of the moved entries and with them their table pointers. check_cfg() is then linear in the number of map entries plus the number of indirect jump edges, so what still scales now with the program is what BPF_MAX_GOTOX_EDGES bounds: gotox map entries edges before after ---------------------------------------------- 500 250000 1000 35.34s 0.07s 1000 250000 2000 82.00s 0.07s 2000 250000 4000 148.62s 0.07s 2000 125000 4000 74.54s 0.05s 2000 500000 4000 342.71s 0.19s Fixes: 493d9e0d6083 ("bpf, x86: add support for indirect jumps") Suggested-by: Eduard Zingerman <eddyz87@gmail.com> Signed-off-by: Daniel Borkmann <daniel@iogearbox.net> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260925175244.1136329-6-daniel@iogearbox.net
8 daysbpf: Leave out the hidden exit edge of a subprogram without an exitDaniel Borkmann
visit_abnormal_return_insn() gives tail_call and ld_{abs,ind} insns a second successor, the exit_idx of their subprogram, so that the hidden exit from the subprogram is part of the CFG. check_subprogs() only assigns exit_idx when it walks over a BPF_EXIT, and a subprogram whose last insn jumps back into itself never has one, so exit_idx stays zero and the edge points at insn 0 of the program: 0: r0 = 0 1: r0 = *(u8 *)skb[0] 2: goto -2 For a subprogram other than main, update_insn() then reads the liveness masks at a negative relative index. Such a subprogram can still load when it leaves through a bpf_throw(), so mark exit_idx as unset in that case and leave the hidden edge out, as there is no exit for it to reach. Fixes: e40f5a6bf88a ("bpf: correct stack liveness for tail calls") Fixes: ee861486e377 ("bpf: Fix ld_{abs,ind} failure path analysis in subprogs") Signed-off-by: Daniel Borkmann <daniel@iogearbox.net> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260925175244.1136329-5-daniel@iogearbox.net
8 daysbpf: Bound the number of indirect jump edges in a programDaniel Borkmann
Every gotox instruction gets its own copy of the jump table of the subprog containing it, and each distinct target in that table is a CFG successor of the instruction. The number of such edges is therefore the number of gotox instructions times the number of distinct targets, and neither factor is bounded by anything except the instruction limit. What is expensive is a BPF prog whose gotox instructions are themselves the targets, which makes the edge count quadratic. 1024 such gotox are already ~1e6 edges and about 4s of CPU to load. Bound the total across the program at BPF_COMPLEXITY_LIMIT_INSNS, aka the limit on the number of instructions the verifier processes. Progs with real switch statements are orders of magnitude below this. This bounds the edges only. Building the jump table of a gotox costs the number of entries of all insn array maps regardless of how many distinct targets they hold, which a later patch addresses separately. Fixes: 493d9e0d6083 ("bpf, x86: add support for indirect jumps") Reported-by: STAR Labs SG <info@starlabs.sg> Signed-off-by: Daniel Borkmann <daniel@iogearbox.net> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Acked-by: Anton Protopopov <a.s.protopopov@gmail.com> Link: https://patch.msgid.link/20260925175244.1136329-3-daniel@iogearbox.net
8 daysMerge tag 'vfs-7.3-rc5.fixes' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs Pull vfs fixes from Christian Brauner: - Revert "put_mnt_ns(): leave mounts connected". This allows the creation of reference count cycles in a very trivial way. We can't bring this in until we have fixed the underlying cause - vfs: Don't create the private nullfs instance for kthreads under namespace_sem to avoid false lockdeps complaints - binfmt_misc: - Copy the name into a stack buffer and look up the copy in bpf_binprm_select_interp() - bpf_binprm_set_interp() and bpf_binprm_set_interp_arg(): Check the private copy instead so the string that gets staged is the kstring that was checked - netfs: - Make netfs_read_gaps() use separate sink folios rather than one reused sink folio to discard unwanted data so that cifs checksum checking sees all the data that was fetched - Trim reads down to i_size so afs symlinks read correctly from the cache - Wrap the direct mempool ->alloc() calls the GFP_KERNEL paths make in alloc_hooks() via a new mempool_alloc_noreserve() helper - iov_iter: Use iov_iter_alignment() for the start and length check added to iov_iter_extract_bvecs() this cycle. It used iter_iov_addr() and iter_iov_len() which are only valid for ITER_UBUF and ITER_IOVEC iterators - super: Make iterate_supers_type() deletion-safe - inode: Stop evict_inodes() from rescanning the same inodes - writeback: Bound the cleanup_offline_cgwb() rescans - ntfs3: Use d_instantiate_new() in ntfs_create_inode() - ovl: Fix a use-after-free in the ovl_do_mkdir() debug print - dcache: Unpoison the inline name buffer in __d_alloc() for KMSAN - autofs: Fix a pipe file reference leak in autofs_kill_sb() - bpf: Drop the path_unlink and path_rmdir hooks from the list of hooks for which the verifier rewrites bpf_{set,remove}_dentry_xattr() to the _locked variants - squashfs: Range check the xz dictionary size before shifting by it - selftests: Add the missing eventfd, open_tree_ns, openat2 and xattr filesystems selftests to TARGETS and drop the stale openat2 entry left behind when those tests moved * tag 'vfs-7.3-rc5.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs: netfs: Fix missing alloc tagging of direct mempool allocations bpf: fs/xattr: don't assume the inode is locked in path_unlink/path_rmdir autofs: fix sbi->pipe file reference leak in autofs_kill_sb() dcache: unpoison the inline name buffer in __d_alloc() ovl: fix UAF in ovl_do_mkdir() debug print super: make iterate_supers_type() deletion-safe Revert "put_mnt_ns(): leave mounts connected" Revert "selftests/filesystems: add mntns cleanup test" binfmt_misc: fix racy checks in bpf set_interp kfuncs binfmt_misc: fix OOB read in bpf_binprm_select_interp() fs: don't create the private nullfs mount under namespace_sem writeback: bound cleanup_offline_cgwb() rescans by rotating scanned inodes fs: avoid repeated scans in evict_inodes() netfs, afs: Fix symlink reading netfs: Fix netfs_read_gaps() to use separate sink folios squashfs: Add dictionary size range check to prevent shift-out-of-bounds fs/ntfs3: use d_instantiate_new() in ntfs_create_inode() and murder syzbot's "WARNING in do_new_mount" saga selftests/filesystems: fix missing and stale TARGETS entries block: Fix start and length check added to iov_iter_extract_bvecs()
8 daysregulator: 88pm886: Add Vbus regulatorDuje Mihanović
Add support for the PMIC's Vbus regulator. This regulator is mandatory for USB OTG support on boards using the PMIC. Reviewed-by: Karel Balej <balejk@matfyz.cz> Signed-off-by: Duje Mihanović <duje@dujemihanovic.xyz> Link: https://patch.msgid.link/20260613-88pm886-vbus-v2-3-021dfb02c6bb@dujemihanovic.xyz Signed-off-by: Mark Brown <broonie@kernel.org>
8 daysdriver core/ACPI: Introduce companion_bus_register()Rafael J. Wysocki
The ACPI bus type does not allow drivers to be matched to devices, so the sysfs attributes related to drivers created for it are useless and their existence is confusing. Moreover, it is better to prevent drivers from being registered and looked up for a bus like that. To allow skipping the creation of those sysfs attributes and preventing driver registration and lookup for the ACPI bus type, introduce a "companion" bus type concept and add a special registration function for registering "companion" bus types, companion_bus_register(). The "drivers" directory under the ACPI bus type is still needed because there are versions of systemd that depend on it [1]. Link: https://lore.kernel.org/linux-acpi/SN6PR02MB41575266A4580339E186E5D9D4812@SN6PR02MB4157.namprd02.prod.outlook.com/ [1] Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com> Acked-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org> Reviewed-by: Danilo Krummrich <dakr@kernel.org> Tested-by: Michael Kelley <mhklinux@outlook.com> Reviewed-by: Michael Kelley <mhklinux@outlook.com> [ rjw: Add comment regarding drivers_kset creation in bus_register_internal() ] Link: https://patch.msgid.link/12985239.O9o76ZdvQC@rafael.j.wysocki Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
8 daysMerge branch 'vfs-7.4.sync_close' into vfs.allChristian Brauner
8 daysMerge branch 'vfs-7.4.shared.super' into vfs.allChristian Brauner
Signed-off-by: Christian Brauner <brauner@kernel.org>
8 daysMerge branch 'vfs-7.4.shared.lsm.foll_force' into vfs.allChristian Brauner
Signed-off-by: Christian Brauner <brauner@kernel.org>
8 daysMerge branch 'vfs-7.4.netfs' into vfs.allChristian Brauner
Signed-off-by: Christian Brauner <brauner@kernel.org>
8 daysMerge branch 'vfs-7.4.misc' into vfs.allChristian Brauner
8 daysMerge branch 'vfs-7.4.lookup' into vfs.allChristian Brauner
8 daysMerge branch 'vfs-7.4.iomap' into vfs.allChristian Brauner
Signed-off-by: Christian Brauner <brauner@kernel.org>
8 daysMerge branch 'vfs-7.4.idmap' into vfs.allChristian Brauner
8 daysMerge branch 'vfs-7.4.file' into vfs.allChristian Brauner
Signed-off-by: Christian Brauner <brauner@kernel.org>
8 daysMerge branch 'vfs-7.4.coredump' into vfs.allChristian Brauner
8 daysMerge branch 'vfs-7.4.close_range' into vfs.allChristian Brauner
Signed-off-by: Christian Brauner <brauner@kernel.org>
8 daysMerge branch 'vfs-7.4.bh' into vfs.allChristian Brauner
Signed-off-by: Christian Brauner <brauner@kernel.org>
8 daysMerge branch 'namespace-7.4.misc' into vfs.allChristian Brauner
Signed-off-by: Christian Brauner <brauner@kernel.org>
8 daysMerge branch 'kernel-7.4.signal' into vfs.allChristian Brauner
Signed-off-by: Christian Brauner <brauner@kernel.org>
8 daysnetfs: Fix missing alloc tagging of direct mempool allocationsHao Ge
Commit 1d78d56c43ef ("netfs: Fix folio_queue ENOMEM in writeback by adding a mempool") added a mempool for the folio_queues and made the request, subrequest and folio_queue allocations distinguish between writeback and everything else. Writeback is part of memory reclaim and must not fail due to ENOMEM, so it allocates under GFP_NOFS through mempool_alloc(), which may dip into the pool's reserve and, if that runs empty, wait for elements to be returned. The GFP_KERNEL paths, which can return -ENOMEM to their callers, invoke the pool's ->alloc() callback directly instead. The direct call, however, skips the alloc_hooks() wrapper that the mempool_alloc() macro provides. The pool callbacks, mempool_alloc_slab() and mempool_kmalloc(), call kmem_cache_alloc_noprof() and kmalloc_noprof() and rely on current->alloc_tag having been set by the caller. With CONFIG_MEM_ALLOC_PROFILING_DEBUG=y this leads to current->alloc_tag not set WARNING: ./include/linux/alloc_tag.h:161 at __alloc_tagging_slab_alloc_hook alloc_tag was not set WARNING: ./include/linux/alloc_tag.h:166 at __alloc_tagging_slab_free_hook at allocation and free time respectively, as reported when reading files on a CIFS mount. The allocations are also missing from /proc/allocinfo. Wrap the direct ->alloc() invocations in alloc_hooks() with a new mempool_alloc_noreserve() helper in include/linux/mempool.h, next to the other alloc_hooks()-wrapped macros such as mempool_alloc(). The GFP_KERNEL paths keep their failable allocation semantics, they just get tagged now. Fixes: 1d78d56c43ef ("netfs: Fix folio_queue ENOMEM in writeback by adding a mempool") Reported-by: Erhard Furtner <erhard_f@mailbox.org> Closes: https://lore.kernel.org/all/0b004319-9ef7-437c-a4dd-174d6a9a83db@mailbox.org/ Tested-by: Erhard Furtner <erhard_f@mailbox.org> Suggested-by: Suren Baghdasaryan <surenb@google.com> Cc: stable@vger.kernel.org Signed-off-by: Hao Ge <hao.ge@linux.dev> Link: https://patch.msgid.link/20260923063759.34667-1-hao.ge@linux.dev Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 daysfs: rename vfs_get_super() to get_tree_super() and export itGiuseppe Scrivano
Rename and export it so filesystems that need a custom superblock matching policy can reuse it directly instead of open-coding sget_fc() + fill_super(). No functional change. This is a preparatory fix for the next patch. Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com> Link: https://patch.msgid.link/20260812142907.1010046-2-gscrivan@redhat.com Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 daysfs: remove unused vfs_empty_path()Sang-Heon Jeon
Since commit e896474fe485 ("getname_maybe_null() - the third variant of pathname copy-in"), vfs_empty_path() has no callers. So remove it. No functional change. Signed-off-by: Sang-Heon Jeon <ekffu200098@gmail.com> Link: https://patch.msgid.link/20260918165105.1013792-1-ekffu200098@gmail.com Reviewed-by: Jan Kara <jack@suse.cz> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 dayscoredump: replace the startup completion with a thread countChristian Brauner
coredump_wait() sets core_state->nr_threads to the number of tasks killed and waits for the last thread to enter coredump_task_exit() to signal completion. Let's just wait on the count directly. The exiting tasks can use atomic_dec_and_wake_up() and the dumping task sleeps in wait_var_event_state(). The dumping task must remain freezable since commit f5d39b020809 ("freezer,sched: Rewrite core freezer logic"). So keep the wait TASK_UNINTERRUPTIBLE|TASK_FREEZABLE. Drop the completion and rename nr_threads to threads_remaining. No functional changes. Suggested-by: NeilBrown <neilb@ownmail.net> Link: https://lore.kernel.org/178899497961.207413.10554121774377911612@noble.neil.brown.name Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-12-a5c1800dc930@kernel.org Reviewed-by: NeilBrown <neil@brown.name> Reviewed-by: Oleg Nesterov <oleg@redhat.com> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 dayssched: add wait_var_event_state()Christian Brauner
All wait_var_event() sleep in a fixed task state. For coredumps we need a variant that takes the state from the caller the way wait_event_state() does. This allows us to continue sleeping with TASK_FREEZABLE. That's certainly also a useful addition for other places. Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-11-a5c1800dc930@kernel.org Reviewed-by: NeilBrown <neil@brown.name> Reviewed-by: Oleg Nesterov <oleg@redhat.com> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 dayscoredump: drop core_state->dumperChristian Brauner
The core_state->dumper field isn't used anymore. Only its ->next pointer is. The current task is always the dumping thread and the ->task pointer is never read. Replace it with a plain pointer to the list of parked threads. Historically, core_state->dumper was used. Its ->task pointer was read. by fill_note_info() started at &core_state->dumper to ensure that the dumping thread came first in the ELF thread notes. That changed in commit 4b0e21d64253 ("[elf][regset] simplify thread list handling in fill_note_info()"). The first iteration was taken out of the loop. So it's been unused ever since. No functional changes. Suggested-by: NeilBrown <neilb@ownmail.net> Link: https://lore.kernel.org/178900159210.207413.8292125177519817528@noble.neil.brown.name Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-10-a5c1800dc930@kernel.org Reviewed-by: NeilBrown <neil@brown.name> Reviewed-by: Oleg Nesterov <oleg@redhat.com> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 daysfs: rename do_close_on_exec() to close_cloexec_files()Christian Brauner
Rename the helper and align it with close_files(). No functional changes. Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-8-a5c1800dc930@kernel.org Reviewed-by: NeilBrown <neil@brown.name> Reviewed-by: Oleg Nesterov <oleg@redhat.com> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 daysfs: remove unshare_files()Christian Brauner
exec is the only caller left since commit 433967cab51e ("coredump: stop unsharing the file descriptor table"). All it does is call unshare_fd() with CLONE_FILES and install the copy. Kill the pointless helper and open-code it. No functional changes. Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-4-a5c1800dc930@kernel.org Reviewed-by: NeilBrown <neil@brown.name> Reviewed-by: Oleg Nesterov <oleg@redhat.com> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 daysfs: move unshare_fd() to fs/file.cChristian Brauner
Move unshare_fd() where the rest of the descriptor table lifecycle helpers live. No functional changes. Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-3-a5c1800dc930@kernel.org Reviewed-by: NeilBrown <neil@brown.name> Reviewed-by: Oleg Nesterov <oleg@redhat.com> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 daysfs: add switch_files_struct()Christian Brauner
Add switch_files_struct() to install another table on a task. It consumes the reference to the new table and puts the old one. Convert every place that switches a descriptor table except unshare_files(). No functional changes. Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-2-a5c1800dc930@kernel.org Reviewed-by: NeilBrown <neil@brown.name> Reviewed-by: Oleg Nesterov <oleg@redhat.com> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>