| Age | Commit message (Collapse) | Author |
|
Add a new per-task stat,
- nr_migrations_cpu_non_preferred: number of push migrations while the
CPU is non-preferred.
Since this new stat is per-task, it changes only /proc/<pid>/sched.
It doesn't update /proc/schedstat. Hence increasing the schedstat version
is not necessary.
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://patch.msgid.link/20260928053728.797539-10-sshegde@linux.ibm.com
|
|
Provide the preferred CPU infrastructure. Define get/set macros
which could be used to get/set CPU state as preferred.
CONFIG_PREFERRED_CPU will be selected by the driver which handles
steal time values. It is going to set/clear preferred CPU state.
This driver will be called steal_governor and it is introduced in
subsequent patches. It periodically computes the steal ratio and
decides on preferred CPU state.
A CPU is set to preferred when it becomes active. Later it may be
marked as non-preferred depending on steal ratio by the steal_governor.
Always maintain design construct of preferred is subset of active.
i.e. preferred ⊆ active ⊆ online ⊆ present ⊆ possible
With CONFIG_PREFERRED_CPU=n, ensure set_cpu_preferred is a nop and get
method returns the active state in that case.
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://patch.msgid.link/20260928053728.797539-5-sshegde@linux.ibm.com
|
|
Introduce bitmap_intersects_and() to determine whether the intersection
of three bitmaps is non-empty. Unlike cpumask_first_and_and(), this
returns immediately when an intersecting word is found and does not
calculate the first matching bit.
Add cpumask_intersects_and() as the corresponding cpumask wrapper.
A subsequent patch uses the helper to determine whether a task
has a CPU that is present in its affinity mask, the preferred CPU mask,
and task possible CPU mask.
Suggested-by: Yury Norov <yury.norov@gmail.com>
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Reviwed-by: Yury Norov <yury.norov@gmail.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Reviewed-by: Yury Norov <yury.norov@gmail.com>
Link: https://patch.msgid.link/20260928053728.797539-3-sshegde@linux.ibm.com
|
|
Provide a new helper function which sums up a given type of cpustat
over a specified cpumask.
This allows the caller's code to be simpler and avoids duplication.
For example, subsequent patch in the steal governor use this exact
same pattern when calculating steal time.
Number of cpus can be derived from cpumask_weight() where necessary.
Suggested-by: Yury Norov <yury.norov@gmail.com>
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Reviewed-by: Yury Norov <ynorov@nvidia.com>
Reviewed-by: Mete Durlu <meted@linux.ibm.com>
Acked-by: Frederic Weisbecker <frederic@kernel.org>
Link: https://patch.msgid.link/20260928053728.797539-2-sshegde@linux.ibm.com
|
|
|
|
We need the driver-core fixes in here as well to build on top of.
Signed-off-by: Danilo Krummrich <dakr@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull scheduler fixes from Ingo Molnar:
- Fix LLC mis-scheduling bugs (Tim Chen, Lu Wang)
- Fix cache-grouping related scheduling statistics UAF bugs (Tim Chen)
- Skip kernel threads for cache aware scheduling to rubustify the code
(Chen Yu)
- Refresh LLC capacity across CPU hotplug, to fix capacity
underestimation bug (Davi Chaves Azevedo)
- Account PSI IRQ time to the execution context, not the scheduling
context, to fix proxy scheduling accounting bug (Zhan Xusheng)
* tag 'sched-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
sched/core: Account PSI IRQ time to the execution context, not the scheduling context
sched/cache: Refresh LLC capacity across CPU hotplug, to fix capacity underestimation bug
sched/cache: Skip kernel threads for cache aware scheduling to rubustify the code
sched/cache: Introduce task_struct->sched_cache_grp to fix UAF
sched/cache: Decouple sched_cache_group from mm to fix UAF
sched/cache: Honor migrate_llc_task semantics in active load balance, to fix LLC mis-scheduling bug
sched/cache: Keep nr_pref_llc_running in the runnable domain, to fix LLC mis-scheduling bug
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull perf events fixes from Ingo Molnar:
- Fixes for KVM guest PEBS virtualization (Sean Christopherson)
- Fixes for various Intel PMUs related to PEBS data-source (Dapeng Mi)
- Fix Intel Panther Cove event scheduling constraints (Dapeng Mi)
- Fix Intel DMR/NVL OMR extra registers event scheduling (Dapeng Mi)
- Rename two confusingly named PMU attributes (Dapeng Mi)
- Fix a refcount leak in attach_perf_ctx_data() (Namhyung Kim)
- Fix NULL pointer dereference crash in __perf_pmu_sched_task()
(Puranjay Mohan)
- Fix CPU-wide event scheduling (Puranjay Mohan)
- Fix x86 LBR branch entry generation (Puranjay Mohan)
* tag 'perf-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
perf/core: Fill branch entries with a single assignment
perf/core: Run sched_task() for PMUs with only CPU-wide events
perf/core: Fix NULL pmu_ctx passed to pmu->sched_task()
perf/core: Fix a refcount leak in attach_perf_ctx_data()
perf/x86/intel: Rename NVL offcore_rsp attribute to offmodule_rsp
perf/x86/intel: Rename DMR offcore_rsp attribute to offmodule_rsp
perf/x86/intel: Fix precise OMR event scheduling for DMR/NVL
perf/x86/intel: Constrain Panther Cove UOPS_DISPATCHED events to PMCs 0-3
perf/x86/intel: Delete dead NVL PEBS data-source initcall
perf/x86/intel: Fix Panther Cove PEBS data-source snoop states
perf/x86/intel: Remove incorrect Panther Cove PEBS data-source constraints
perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints
perf/x86/intel: Update arw_latency_data() mem-op direction handling
perf/x86/intel: Fix DKT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Fix CMT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs()
perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) if PEBS is unused
perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has PEBS isolation, to fix stuck PEBS_ENABLED
perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted PERF_GLOBAL_CTRL bits
|
|
Signed-off-by: Ingo Molnar <mingo@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux
Pull ata fixes from Niklas Cassel:
- Extend the quirk "no LPM on ATI" quirk, that is currently only
applied for Samsung drives, to include AMD controllers as well.
The AMD AHCI controllers are newer versions of the ATI AHCI
controllers, and these controllers still have LPM issues with
Samsung drives - LPM works with drives from other vendors (me)
- Fix errors in the libata.force parameter documentation (me)
- Verify the sense data descriptor lengths for ATA PASS-THROUGH
command, so that a malicious device cannot write past the buffer
length (Matthias)
- Mention the libata for-next branch in MAINTAINERS such that the
git ls-remote command done by get_maintainer.pl --self-test=scm
can verify it (Matthias)
* tag 'ata-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux:
MAINTAINERS: name the libata/linux for-next branch
ata: libata-scsi: bound the ATA passthru sense descriptor writes
ata: libata: Correct libata.force parameter documentation
ata: libata-core: Extend Samsung LPM quirk to AMD controllers
|
|
Pull kvm fixes from Paolo Bonzini:
"Arm:
- Invalidate the ITS translation cache when the guest changes the
base address of the ITS tables (Fuad Tabba)
- Skip saving ITS devices with device IDs that are out-of-bounds
rather than failing the entire ITS save ioctl (Fuad Tabba)
- Close race between VM teardown and invalidations of nested MMUs
when handling MMU operations that are allowed to block (Lorenzo
Stoakes)
- Various fixes for the handling of the host's untrusted SVE
configuration in pKVM (Fuad Tabba)
- Make sure that empty SMCCC ranges based at 0 are rejected by the
kvm_smccc_set_filter() (Karl Mehltretter)
- Revoke the host mapping for pKVM's private stack pages, along with
a new sanity check that all mappings in the hyp's private VA range
have been correctly marked as hyp-owned (Fuad Tabba)
- Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that
concurrent vCPU initialization cannot relocate in-use MMUs. Defer
the freeing of shadow stage-2 MMUs to the point that no other users
(e.g. MMU notifier) could reference them (Marc Zyngier)
- Drop useless WARN when rejecting an unsupported ioctl for pKVM
(Fuad Tabba)
- Fix the steal_time selftest to install correctly-sized mappings for
non-4K hosts (Sebastian Ott)
- Correct mapping of fine-grained trap for GCSPOPX instruction (Mark
Brown)
- Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit
guests (Karl Mehltretter)
RISC-V:
- Synchronize hrtimer during VCPU teardown
- Fix the conversion between vsip and hvip values
- Serialize IMSIC attributes with vCPU migration
- Release unused page after MMU invalidation
- Propagate interrupted G-stage faults to KVM user-space as EINTR
- Fix nested acceleration hfence entry update order
- Fix sdata leak and stale snapshot_addr in snapshot_set_shmem
- Preserve firmware counter value across PMU counter stop/start
- Report PMU snapshot write failure to the guest
- Fix perf-backed counter accounting across PMU stop and read
- Correctly propagate error of a hart status SBI call
s390:
- Ensure that accesses through kvm_arch_set_irq_inatomic mark as
dirty the pages that contain indicator and summary bits
- Fix compile warning for kvm_s390_update_cmma_dirty()
- Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to
userspace
- Move s390_kvm_mmu_commit_memory_region() into
s390_kvm_mmu_prepare_memory_region() so that it can fail instead of
WARN
- Add missing srcu in kvm_s390_set_irq_state()
- Fix potential races in storage functions
- Fix race in _destroy_pages_crste()
- Fix issues in the handling of KVM interrupt and page resources,
when a queue that is assigned to a mediated device (mdev) is
removed from the host's AP configuration
- Fix loop condition in uv_find_secrets
- Prevent potential out-of-bounds read
x86:
- Fix a brown paper bag bug where KVM would incorrectly treat Intel
PMU MSRs as valid on AMD
- Fix a regression in the hardware disable selftest where it checked
the wrong macro when detecting glibc support (breaks at least musl)
- Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is
especially important for KVM_BUG_ON() flows, which often guard more
dangerous bugs
- Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a
bug where KVM would let userspace run a broken setup with stale
vmcs12 pages
- Fix a class of bugs where KVM would fail to fill kvm_run exit
fields if getting nested pages failed
- Treat reserved entries in the memory attributes xarray as "no
attributes", to fix false positives when checking for mixed
attributes
- Fix memcg accounting for the memory attributes xarray (the xarray
library subtly requires the xarray to be configured for accounting
upfront; the gfp flags taken at runtime are used only rarely)
- Don't pre-reserve xarray entries when storing empty attributes, as
storing NULL must not require memory allocation (KVM and other
subsystems heavily rely on this behavior)
- Fix a memory leak and a cache maintenance issue related to doing
intra-host migration on an SEV guest"
* tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (54 commits)
KVM: SEV: Do cache maintenance on the source VM during intra-host migration
KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV
KVM: Don't pre-reserve xarray entries when storing empty/NULL attributes
KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg
KVM: Don't treat reserved xarray entries as having memory attributes
KVM: x86: Fill kvm_run exit fields in common get_nested_state_pages() error paths
KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails
KVM: arm64: Fix AArch32 DBGBXVR<n> handling
KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP
KVM: selftests: fix steal_time for arm64 with host page size > 4K
KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode
KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction
KVM: arm64: nv: Fix life cycle of the nested_mmus array
KVM: arm64: Check every private mapping is hyp-owned at pKVM init
KVM: arm64: Move the private VA allocation cursor to __io_map_next
KVM: arm64: Match hyp text by physical address in fix_host_ownership()
KVM: arm64: Transfer the hyp stack pages out of the host stage-2
KVM: arm64: selftests: Test empty SMCCC filter range at base 0
KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0
KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2
...
|
|
Correct "intented" to "intended", reported by scripts/checkpatch.pl using
the misspelling list in scripts/spelling.txt. Only touches comments, no
code changes.
Assisted-by: LLM
Signed-off-by: Hemanth Selam <hemanth.selam@gmail.com>
Signed-off-by: Manivannan Sadhasivam <manivannan.sadhasivam@oss.qualcomm.com>
Link: https://patch.msgid.link/20260904124731.9011-1-hemanth.selam@gmail.com
|
|
Now that KVM uses kvm_get_vcpu_by_id() to check for an existing vCPU ID
before doing any meaningful work, which was made possible by holding
kvm->lock for the entirety of vCPU creation, revert the now-redundant
"early" vCPU ID tracking. The claims about the impact of kvm->vcpu_ids on
the memory footprint were a wee bit wrong: the worst case scenario isn't
256 bytes per VM, it's 256 "unsigned longs" per VM, i.e. 2048 bytes per VM.
Increasing the size of "struct kvm" by 2048 nearly doubled the total size
on many architectures, and tripped x86's KVM_SANITY_CHECK_VM_STRUCT_SIZE,
which was added to detect this *exact* scenario, where a single change
significantly increased the size of "struct kvm". I.e. attempting to build
KVM with CONFIG_DEBUG_KERNEL=n fails on x86 (the build failures got missed
because all build bots apparently test only CONFIG_DEBUG_KERNEL=y kernels,
and maintainers' test flows were similarly lacking).
This reverts commit 97d65b544f48b2ee49f6aea32145e3e7969955dc.
Fixes: 97d65b544f48 ("KVM: Check for duplicate vcpu_id as early as possible")
Reported-by: Jean-Christophe Guillain <jean-christophe@guillain.net>
Closes: https://lore.kernel.org/all/56a4bc35ee605588b7cc36c8e45c12b5f3b506cb.camel@guillain.net
Reported-by: Paweł S <spawel523@gmail.com>
Closes: https://lore.kernel.org/all/CABD%3DWFOS4j4hDv%2BpW-eEM9HAM2q2GY_iYdAG%2BqvYcUEinUrcQQ@mail.gmail.com
Tested-by: Jean-Christophe Guillain <jean-christophe@guillain.net>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Tested-by: Naveen N Rao (AMD) <naveen@kernel.org>
Message-ID: <20260921174445.911676-7-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
KVM fixes for 7.3-rcN
- Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs
as valid on AMD.
- Fix a regression in the hardware disable selftest where it checked the wrong
macro when detecting glibc support (breaks at least musl).
- Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially
important for KVM_BUG_ON() flows, which often guard more dangerous bugs.
- Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug
where KVM would let userspace run a broken setup with stale vmcs12 pages.
- Fix a class of bugs where KVM would fail to fill kvm_run exit fields if
getting nested pages failed.
- Treat reserved entries in the memory attributes xarray as "no attributes",
to fix false positives when checking for mixed attributes.
- Fix memcg accounting for the memory attributes xarray (the xarray library
subtly requires the xarray to be configured for accounting upfront; the gfp
flags taken at runtime are used only rarely).
- Don't pre-reserve xarray entries when storing empty attributes, as storing
NULL must not require memory allocation (KVM and other subsystems heavily
rely on this behavior).
|
|
* ras/edac-urgent: (1178 commits)
EDAC/altera: Fix use-after-free in error paths
EDAC/altera: Fix memory leak on dci allocation failure
EDAC/altera: Drop __init from ECC setup paths for re-probe safety
EDAC/altera: Do not allow driver unbinding
Linux 7.3-rc4
net: qrtr: resend HELLO on MHI resume
i2c: qcom-cci: fix device_node refcount leak in cci_probe()/cci_remove()
i2c: qcom-geni: release DMA channels on probe error
i2c: imx: release DMA channels on probe error
i2c: at91: release DMA channels on remove and probe error
posix-cpu-timers: Prevent freeing a timer which is queued on the expiry list
x86/build/64: Prevent native builds from generating EGPR use
watchdog: da9063: fix suspend/resume handling of HW_RUNNING watchdog
soc: samsung: exynos-pmu: fix use-after-free of interrupt generator node
selftests/x86: Check signal state for rejected software interrupts
x86/fred: Reconstruct the #GP context for rejected INT instructions
sched/core: Avoid false migration warning for proxy donors
perf: Fix null pointer access in is_include_guest_event()
x86/microcode/intel: Reject problematic loading on Granite Rapids systems
cifs: Fix server use-after-free in cifs_chan_skip_or_disable()
...
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
|
|
netdev_rss_key_fill() returns a key made of uniformly random bytes. That
is not enough, because the Toeplitz hash is linear over GF(2).
Walking the hash input MSB first, each set bit contributes a 32-bit
sliding window of the key, and hardware indexes the indirection table with
the low order bits of the result. Only the tail of each window therefore
reaches the queue index:
v(i) = key bits [i + 32 - q .. i + 31]
with q = log2(number of RX queues). Consecutive input bits give windows
overlapping in q - 1 positions, so the q vectors belonging to the q lowest
bits of a header field form a Toeplitz matrix built from 2 * q - 1 key
bits, rather than q * q independent ones. Over GF(2) a random Toeplitz
matrix is singular with probability exactly 1/2, whatever its size.
When it is singular, flows differing only in the low order bits of that
field cannot reach all the queues. This is not theoretical: a burst of
connections draws ephemeral ports from a narrow range, and on one affected
host only 4 of the 16 RX queues received any traffic at all, until its key
was replaced.
Keep drawing the key at random, since it is a secret that stops a remote
attacker from steering flows onto a single queue, but force the handful of
bits that decide this. Writing d[t] for key bit (lsb + 31 - t), where lsb
is the position of the least significant bit of a field in the hash input,
the matrices of all the q values up to 8 are non singular if and only if
d[2 * i] = 1 ^ d[i] ^ d[i + 1] ^ ... ^ d[2 * i - 1]
The odd positions stay free, so this is a one pass fixup rather than a
search. It is in fact a bijection from those free positions onto the set
of the values having the property, so the key stays uniformly distributed
over that set and the whole cost is 8 bits of entropy per position.
Searching for such a key by rejection would not have been an option: a
freshly drawn one has the property everywhere with probability 2^-1008.
Apply this at every 16-bit aligned position of the key, rather than at the
offsets of the 2-tuple and 4-tuple layouts only. The core does not get to
know what a given NIC hashes. Hardware may select the bytes it feeds to
Toeplitz out of a header window with a bitmap, and hash an encapsulated
header: for PSP over UDP over IPv6 it can pick the outer addresses and the
inner TCP ports, which sit 78 bytes into the frame and read key bits well
past the 40 bytes an IPv6 4-tuple needs. Hashed fields are 16 bits wide at
the smallest and are not expected to straddle that grid, so covering the
grid covers the layouts that were never written down, at no cost in code.
This spends 8 bits of entropy per position, 1008 bits out of the 2048 bits
of netdev_rss_key, leaving 1040 bits. 8 is also the largest usable bound,
as each q constrains 2 * q - 1 bits and anything larger would make the
ranges of two adjacent positions overlap.
What the fixup leaves behind is visible structure: 8 of every 16 bits are
derived from the 8 others, so a 16-bit aligned word of the key takes only
2^8 values and about 27 of the 128 words of netdev_rss_key duplicate
another one. That much is forced rather than an artefact of this
implementation, 1040 bits spread over 128 words being a little over 8 bits
each, but it has one consequence worth removing. Two 32-bit windows a
whole number of words apart now collide with probability 2^-16 instead of
2^-32, and two input bits reading the same window are indistinguishable to
the hash, since flipping both of them leaves it unchanged. Over 500 keys,
23% of them had such a pair, where a uniformly random key has one with
probability 2^-11.
So draw another key when that happens. Four out of five pass. Which
distances to look at follows from where the structure is: covering the
multiples of 16 leaves 1.2e-3 expected colliding pairs per key, still 2.6
times the 4.6e-4 of a plain random key, and almost all of that excess sits
at a distance of 8 modulo 16. Covering every multiple of 8 brings the total
down to 4.2e-4, below what a plain random key gives over all distances, and
within a few percent of the 4.0e-4 it gives over the distances that are
left.
Two windows a multiple of 8 bits apart are two windows at the same offset
modulo 8, so this is a handful of pairwise sweeps rather than one pass over
the key per distance. DO_ONCE() runs the generator under a spinlock with
hard IRQs disabled, so keep the windows of a class in an array rather than
recomputing both sides of every pair: 116 us instead of 254 us for the
worst case, a 2048-bit key with no collision anywhere, at a cost of 512
bytes of stack.
The shared key is fixed up once and every driver prefix inherits both
properties. Move netdev_rss_key and netdev_rss_key_fill() from
net/ethtool/ioctl.c to net/ethtool/common.c, and move the netdev_rss_key
declaration out of include/linux/netdevice.h into net/core/dev.h.
Checked against an independent Toeplitz implementation: for every 16-bit
aligned position and every q in 1..8, an aligned block of 2^q consecutive
values of a field ending there lands on the 2^q queues exactly once each.
Over 20 random keys that is 20160 checks, which 50.1% of plain random keys
fail and none of the generated keys do, for an average of 502 rewritten
bits out of 2048.
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260922163458.3900996-3-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The GCC and Clang attribute '__counted_by_ptr' associates a pointer
field with an integer field holding its element count. This enables
runtime bounds checking by KASAN and UBSAN to prevent out-of-bounds
accesses to the pointer.
In 'struct fw_device' (defined in 'include/linux/firewire.h'), the
field 'config_rom' is a pointer to the device's Configuration ROM data,
and its associated element count is stored in 'config_rom_length'.
Additionally, update the existing KUnit test in
'drivers/firewire/device-attribute-test.c' where 'config_rom_length'
was incorrectly initialized with the byte size
('sizeof(simple_avc_config_rom)') instead of the element count
('ARRAY_SIZE(simple_avc_config_rom)'). This ensures the
'__counted_by_ptr' bounds-checking annotation does not trigger any
false-positive panics or compile/runtime checks.
Cc: codemender-patching+linux@google.com
Assisted-by: LLM
Signed-off-by: Bill Wendling <morbo@google.com>
Link: https://lore.kernel.org/r/20260925212552.125652-1-morbo@google.com
Signed-off-by: Takashi Sakamoto <o-takashi@sakamocchi.jp>
|
|
In 'struct fw_packet', the 'payload' field points to a buffer of size
'payload_length' bytes. To enable compiler bounds checking (via KASAN
and __builtin_dynamic_object_size), add the '__counted_by_ptr' attribute
to the 'payload' pointer, associated with the 'payload_length' count.
A thorough analysis of all allocation and initialization points for
'struct fw_packet' was conducted. The pointer and its length are
assigned together.
Because the count field is always correctly initialized before the
pointer is accessed, the '__counted_by_ptr' attribute is fully safe and
will not cause runtime false positives or panics.
Cc: codemender-patching+linux@google.com
Assisted-by: LLM
Signed-off-by: Bill Wendling <morbo@google.com>
Link: https://lore.kernel.org/r/20260925210912.113052-1-morbo@google.com
Signed-off-by: Takashi Sakamoto <o-takashi@sakamocchi.jp>
|
|
The existing BPF_PROG_STREAM_READ_BY_FD command only supports polling a
program stream through repeated bpf() calls. It cannot block for new data
or integrate with poll-based event loops.
Add BPF_PROG_STREAM_OPEN to return a read-only, close-on-exec file
descriptor for a selected program stream. Reads block by default and
BPF_F_STREAM_NONBLOCK, the only accepted flag, provides non-blocking
behavior. poll reports readable data and reports hangup once the program
has been freed. Like pipes and sockets, the descriptor is not seekable and
lseek fails with ESPIPE.
A stream descriptor deliberately does not retain the program. Move each
stream into a separately refcounted allocation so program teardown can mark
it dead and wake descriptor users while outstanding descriptors drain
buffered data safely. Readers sample the dead flag before looking for data,
so EOF is reported only when the stream was already dead before it was
found empty; data published right before teardown is never skipped.
Only programs loaded through BPF_PROG_LOAD get streams. Classic BPF
filters, JIT subprograms and shim programs never write to one, and
kernel-side writers already resolve a subprogram to its main program, so
those programs no longer carry stream state.
Readiness needs its own counter. Stream capacity is charged before
allocation and before an element is published to the stream log, so using
that reservation as the read and poll condition can report readable data
while no element exists: a blocking reader retries instead of sleeping and
a lone non-blocking reader can see POLLIN followed by EAGAIN. Publish bytes
with release ordering after adding elements to the lockless log, use
acquire loads before consuming them or reporting readiness, limit each read
to its readable snapshot and subtract only bytes actually copied. This
keeps the aggregate count correct even when concurrent publishers update it
out of publication order. With several readers on one stream, readiness
remains advisory, as it is for pipes. The capacity counter is kept solely
for enforcing the stream size limit.
Wakeups are always deferred through irq_work. Stream writers run in
whatever context the program runs in: NMI context for perf_event programs,
sections with interrupts disabled inside bpf_spin_lock or rqspinlock
critical sections since bpf_stream_vprintk() is KF_SPINLOCK_SAFE, and
tracing programs attached anywhere in the kernel, including inside the wait
queue and epoll code itself. Waking waiters directly from there can
deadlock, and no cheap context check covers every case: on PREEMPT_RT,
spinlock_t sections do not disable interrupts, so in_nmi() or
irqs_disabled() cannot tell such a program apart from a benign one. Queue
an irq_work item instead, as bpf_ringbuf does.
Queue it only when a publication turns an empty stream readable. Readers
block and pollers wait only after finding the stream empty, and the
readable count never drops below zero because each read is bounded by its
snapshot, so the first publication after such an observation is the one
that makes the count positive, and it is the one that queues the wakeup.
Publications into a stream that already holds data raise no interrupt, so
a program that prints while nobody drains its stream pays for a single
irq_work until the stream is emptied again. This matches bpf_ringbuf, which
notifies only once the consumer has caught up. Blocking readers and
level-triggered pollers re-check the readable count before waiting, so they
cannot miss data, and edge-triggered epoll consumers drain until EAGAIN
before waiting again, as epoll(7) requires.
Synchronize pending work before releasing the final stream reference so
the callback cannot outlive the stream, but only when the work was ever
queued: irq_work_sync() waits for an RCU grace period on PREEMPT_RT and on
architectures without an irq_work interrupt.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260925045536.1480933-3-memxor@gmail.com
|
|
A few patches do not expand in place, they prepend: bpf_convert_ctx_accesses()
puts the ctx save for gen_epilogue and the instructions of gen_prologue in
front of insn 0, and bpf_do_misc_fixups() prepends the may_goto counter init
at the start of every subprog that uses may_goto.
They copy the original instruction of 'tgt_idx' into the last slot of the
patch buffer, so it now lives at tgt_idx + delta, and call adjust_jmp_off()
to move direct branches from tgt_idx to tgt_idx + delta.
Nothing does the same for indirect branches, so a BPF_MAP_TYPE_INSN_ARRAY
slot that named tgt_idx keeps naming tgt_idx, which is now the first prepended
instruction, and insn_aux_data[tgt_idx].indirect_target makes the JIT emit
the landing pad there. Add a __bpf_patch_insn_data() variant and pass
BPF_PATCH_MOVE_TARGET at the affected call-sites to fix up the delta. The
poke descriptors of direct tail calls are shifted the same way, and the
original instruction is not searched for by content in that mode, as it
is the last slot by construction.
Fixes: 493d9e0d6083 ("bpf, x86: add support for indirect jumps")
Fixes: 07ae6c130b46 ("bpf: Add helper to detect indirect jump targets")
Reported-by: Nicholas Carlini <npc@anthropic.com>
Suggested-by: Nicholas Carlini <npc@anthropic.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Anton Protopopov <a.s.protopopov@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260925175244.1136329-11-daniel@iogearbox.net
|
|
create_jt() builds the jump table of the subprogram containing a gotox by
copying out and sorting every insn_array map of the program, and it does
so once per gotox instruction. The cost is therefore the number of gotox
instructions times the number of entries in all of the maps. A program of
4003 instructions with 2000 gotox and one 500k entry map holding two
distinct targets has 4000 indirect jump edges, 0.4% of the limit, and
takes 343s to be rejected. The map costs next to nothing to prepare, as
an unset entry is already a valid target. At the insn limit, with a single
1M entry map, the same shape extrapolates to 41 hours.
All gotox instructions of a subprogram share the same jump table, so
build the table of every subprogram in a single pass over the maps and
let every gotox use it in place: bpf_insn_successors() resolves a gotox
through its containing subprogram, and the tables are kept until the
verifier environment is torn down, so nothing is copied into the
instruction aux data. The edge accounting in visit_gotox_insn() moves
to the first visit of the gotox, marked by the BRANCH bit of its CFG
state, and a revisit returns right away since the first one already
pushed every target that still needed exploring.
The only user left of insn_aux_data[].jt is then the table which
visit_abnormal_return_insn() allocates for tail_call and
ld_{abs,ind} insns, so that bpf_insn_successors() reports the hidden
exit from their subprogram. Both are recognisable by their opcode, so
derive that edge in bpf_insn_successors() from the exit_idx of the
containing subprogram instead, again leaving it out when the subprogram
has no exit, and drop the field along with bpf_clear_insn_aux_data().
The instruction aux data then owns no allocation, thus nothing needs to
be freed when insns are removed or the verifier environment is torn
down.
Subprograms removed as dead code free their table in
adjust_subprog_starts_after_remove(), which also clears the slots that
the compaction of subprog_info vacates, as they still hold copies of the
moved entries and with them their table pointers.
check_cfg() is then linear in the number of map entries plus the number
of indirect jump edges, so what still scales now with the program is what
BPF_MAX_GOTOX_EDGES bounds:
gotox map entries edges before after
----------------------------------------------
500 250000 1000 35.34s 0.07s
1000 250000 2000 82.00s 0.07s
2000 250000 4000 148.62s 0.07s
2000 125000 4000 74.54s 0.05s
2000 500000 4000 342.71s 0.19s
Fixes: 493d9e0d6083 ("bpf, x86: add support for indirect jumps")
Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260925175244.1136329-6-daniel@iogearbox.net
|
|
visit_abnormal_return_insn() gives tail_call and ld_{abs,ind} insns a
second successor, the exit_idx of their subprogram, so that the hidden
exit from the subprogram is part of the CFG. check_subprogs() only
assigns exit_idx when it walks over a BPF_EXIT, and a subprogram whose
last insn jumps back into itself never has one, so exit_idx stays zero
and the edge points at insn 0 of the program:
0: r0 = 0
1: r0 = *(u8 *)skb[0]
2: goto -2
For a subprogram other than main, update_insn() then reads the liveness
masks at a negative relative index. Such a subprogram can still load
when it leaves through a bpf_throw(), so mark exit_idx as unset in that
case and leave the hidden edge out, as there is no exit for it to reach.
Fixes: e40f5a6bf88a ("bpf: correct stack liveness for tail calls")
Fixes: ee861486e377 ("bpf: Fix ld_{abs,ind} failure path analysis in subprogs")
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260925175244.1136329-5-daniel@iogearbox.net
|
|
Every gotox instruction gets its own copy of the jump table of the subprog
containing it, and each distinct target in that table is a CFG successor
of the instruction. The number of such edges is therefore the number of
gotox instructions times the number of distinct targets, and neither
factor is bounded by anything except the instruction limit.
What is expensive is a BPF prog whose gotox instructions are themselves
the targets, which makes the edge count quadratic. 1024 such gotox are
already ~1e6 edges and about 4s of CPU to load.
Bound the total across the program at BPF_COMPLEXITY_LIMIT_INSNS, aka
the limit on the number of instructions the verifier processes. Progs
with real switch statements are orders of magnitude below this.
This bounds the edges only. Building the jump table of a gotox costs the
number of entries of all insn array maps regardless of how many distinct
targets they hold, which a later patch addresses separately.
Fixes: 493d9e0d6083 ("bpf, x86: add support for indirect jumps")
Reported-by: STAR Labs SG <info@starlabs.sg>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Anton Protopopov <a.s.protopopov@gmail.com>
Link: https://patch.msgid.link/20260925175244.1136329-3-daniel@iogearbox.net
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs
Pull vfs fixes from Christian Brauner:
- Revert "put_mnt_ns(): leave mounts connected". This allows the
creation of reference count cycles in a very trivial way. We can't
bring this in until we have fixed the underlying cause
- vfs: Don't create the private nullfs instance for kthreads under
namespace_sem to avoid false lockdeps complaints
- binfmt_misc:
- Copy the name into a stack buffer and look up the copy in
bpf_binprm_select_interp()
- bpf_binprm_set_interp() and bpf_binprm_set_interp_arg(): Check
the private copy instead so the string that gets staged is the
kstring that was checked
- netfs:
- Make netfs_read_gaps() use separate sink folios rather than one
reused sink folio to discard unwanted data so that cifs checksum
checking sees all the data that was fetched
- Trim reads down to i_size so afs symlinks read correctly from the
cache
- Wrap the direct mempool ->alloc() calls the GFP_KERNEL paths make
in alloc_hooks() via a new mempool_alloc_noreserve() helper
- iov_iter: Use iov_iter_alignment() for the start and length check
added to iov_iter_extract_bvecs() this cycle. It used iter_iov_addr()
and iter_iov_len() which are only valid for ITER_UBUF and ITER_IOVEC
iterators
- super: Make iterate_supers_type() deletion-safe
- inode: Stop evict_inodes() from rescanning the same inodes
- writeback: Bound the cleanup_offline_cgwb() rescans
- ntfs3: Use d_instantiate_new() in ntfs_create_inode()
- ovl: Fix a use-after-free in the ovl_do_mkdir() debug print
- dcache: Unpoison the inline name buffer in __d_alloc() for KMSAN
- autofs: Fix a pipe file reference leak in autofs_kill_sb()
- bpf: Drop the path_unlink and path_rmdir hooks from the list of hooks
for which the verifier rewrites bpf_{set,remove}_dentry_xattr() to
the _locked variants
- squashfs: Range check the xz dictionary size before shifting by it
- selftests: Add the missing eventfd, open_tree_ns, openat2 and xattr
filesystems selftests to TARGETS and drop the stale openat2 entry
left behind when those tests moved
* tag 'vfs-7.3-rc5.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs:
netfs: Fix missing alloc tagging of direct mempool allocations
bpf: fs/xattr: don't assume the inode is locked in path_unlink/path_rmdir
autofs: fix sbi->pipe file reference leak in autofs_kill_sb()
dcache: unpoison the inline name buffer in __d_alloc()
ovl: fix UAF in ovl_do_mkdir() debug print
super: make iterate_supers_type() deletion-safe
Revert "put_mnt_ns(): leave mounts connected"
Revert "selftests/filesystems: add mntns cleanup test"
binfmt_misc: fix racy checks in bpf set_interp kfuncs
binfmt_misc: fix OOB read in bpf_binprm_select_interp()
fs: don't create the private nullfs mount under namespace_sem
writeback: bound cleanup_offline_cgwb() rescans by rotating scanned inodes
fs: avoid repeated scans in evict_inodes()
netfs, afs: Fix symlink reading
netfs: Fix netfs_read_gaps() to use separate sink folios
squashfs: Add dictionary size range check to prevent shift-out-of-bounds
fs/ntfs3: use d_instantiate_new() in ntfs_create_inode() and murder syzbot's "WARNING in do_new_mount" saga
selftests/filesystems: fix missing and stale TARGETS entries
block: Fix start and length check added to iov_iter_extract_bvecs()
|
|
Add support for the PMIC's Vbus regulator. This regulator is mandatory
for USB OTG support on boards using the PMIC.
Reviewed-by: Karel Balej <balejk@matfyz.cz>
Signed-off-by: Duje Mihanović <duje@dujemihanovic.xyz>
Link: https://patch.msgid.link/20260613-88pm886-vbus-v2-3-021dfb02c6bb@dujemihanovic.xyz
Signed-off-by: Mark Brown <broonie@kernel.org>
|
|
The ACPI bus type does not allow drivers to be matched to devices, so
the sysfs attributes related to drivers created for it are useless and
their existence is confusing.
Moreover, it is better to prevent drivers from being registered and
looked up for a bus like that.
To allow skipping the creation of those sysfs attributes and preventing
driver registration and lookup for the ACPI bus type, introduce a
"companion" bus type concept and add a special registration function
for registering "companion" bus types, companion_bus_register().
The "drivers" directory under the ACPI bus type is still needed because
there are versions of systemd that depend on it [1].
Link: https://lore.kernel.org/linux-acpi/SN6PR02MB41575266A4580339E186E5D9D4812@SN6PR02MB4157.namprd02.prod.outlook.com/ [1]
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
Acked-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Reviewed-by: Danilo Krummrich <dakr@kernel.org>
Tested-by: Michael Kelley <mhklinux@outlook.com>
Reviewed-by: Michael Kelley <mhklinux@outlook.com>
[ rjw: Add comment regarding drivers_kset creation in bus_register_internal() ]
Link: https://patch.msgid.link/12985239.O9o76ZdvQC@rafael.j.wysocki
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
|
|
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
|
|
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
Signed-off-by: Christian Brauner <brauner@kernel.org>
|
|
Commit 1d78d56c43ef ("netfs: Fix folio_queue ENOMEM in writeback by
adding a mempool") added a mempool for the folio_queues and made the
request, subrequest and folio_queue allocations distinguish between
writeback and everything else. Writeback is part of memory reclaim
and must not fail due to ENOMEM, so it allocates under GFP_NOFS
through mempool_alloc(), which may dip into the pool's reserve and,
if that runs empty, wait for elements to be returned. The
GFP_KERNEL paths, which can return -ENOMEM to their callers, invoke
the pool's ->alloc() callback directly instead.
The direct call, however, skips the alloc_hooks() wrapper that the
mempool_alloc() macro provides. The pool callbacks, mempool_alloc_slab()
and mempool_kmalloc(), call kmem_cache_alloc_noprof() and kmalloc_noprof()
and rely on current->alloc_tag having been set by the caller. With
CONFIG_MEM_ALLOC_PROFILING_DEBUG=y this leads to
current->alloc_tag not set
WARNING: ./include/linux/alloc_tag.h:161 at __alloc_tagging_slab_alloc_hook
alloc_tag was not set
WARNING: ./include/linux/alloc_tag.h:166 at __alloc_tagging_slab_free_hook
at allocation and free time respectively, as reported when reading
files on a CIFS mount. The allocations are also missing from
/proc/allocinfo.
Wrap the direct ->alloc() invocations in alloc_hooks() with a new
mempool_alloc_noreserve() helper in include/linux/mempool.h, next to
the other alloc_hooks()-wrapped macros such as mempool_alloc(). The
GFP_KERNEL paths keep their failable allocation semantics, they just
get tagged now.
Fixes: 1d78d56c43ef ("netfs: Fix folio_queue ENOMEM in writeback by adding a mempool")
Reported-by: Erhard Furtner <erhard_f@mailbox.org>
Closes: https://lore.kernel.org/all/0b004319-9ef7-437c-a4dd-174d6a9a83db@mailbox.org/
Tested-by: Erhard Furtner <erhard_f@mailbox.org>
Suggested-by: Suren Baghdasaryan <surenb@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Hao Ge <hao.ge@linux.dev>
Link: https://patch.msgid.link/20260923063759.34667-1-hao.ge@linux.dev
Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Rename and export it so filesystems that need a custom superblock
matching policy can reuse it directly instead of open-coding sget_fc()
+ fill_super(). No functional change.
This is a preparatory fix for the next patch.
Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com>
Link: https://patch.msgid.link/20260812142907.1010046-2-gscrivan@redhat.com
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Since commit e896474fe485 ("getname_maybe_null() - the third variant of
pathname copy-in"), vfs_empty_path() has no callers.
So remove it.
No functional change.
Signed-off-by: Sang-Heon Jeon <ekffu200098@gmail.com>
Link: https://patch.msgid.link/20260918165105.1013792-1-ekffu200098@gmail.com
Reviewed-by: Jan Kara <jack@suse.cz>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
coredump_wait() sets core_state->nr_threads to the number of tasks
killed and waits for the last thread to enter coredump_task_exit() to
signal completion. Let's just wait on the count directly. The exiting
tasks can use atomic_dec_and_wake_up() and the dumping task sleeps in
wait_var_event_state().
The dumping task must remain freezable since commit f5d39b020809
("freezer,sched: Rewrite core freezer logic"). So keep the wait
TASK_UNINTERRUPTIBLE|TASK_FREEZABLE.
Drop the completion and rename nr_threads to threads_remaining.
No functional changes.
Suggested-by: NeilBrown <neilb@ownmail.net>
Link: https://lore.kernel.org/178899497961.207413.10554121774377911612@noble.neil.brown.name
Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-12-a5c1800dc930@kernel.org
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
All wait_var_event() sleep in a fixed task state. For coredumps we need
a variant that takes the state from the caller the way
wait_event_state() does. This allows us to continue sleeping with
TASK_FREEZABLE. That's certainly also a useful addition for other places.
Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-11-a5c1800dc930@kernel.org
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
The core_state->dumper field isn't used anymore. Only its ->next pointer
is. The current task is always the dumping thread and the ->task pointer
is never read. Replace it with a plain pointer to the list of parked
threads.
Historically, core_state->dumper was used. Its ->task pointer was read.
by fill_note_info() started at &core_state->dumper to ensure that the
dumping thread came first in the ELF thread notes. That changed in
commit 4b0e21d64253 ("[elf][regset] simplify thread list handling in
fill_note_info()"). The first iteration was taken out of the loop. So
it's been unused ever since.
No functional changes.
Suggested-by: NeilBrown <neilb@ownmail.net>
Link: https://lore.kernel.org/178900159210.207413.8292125177519817528@noble.neil.brown.name
Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-10-a5c1800dc930@kernel.org
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Rename the helper and align it with close_files().
No functional changes.
Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-8-a5c1800dc930@kernel.org
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
exec is the only caller left since commit 433967cab51e ("coredump: stop
unsharing the file descriptor table"). All it does is call unshare_fd()
with CLONE_FILES and install the copy. Kill the pointless helper and
open-code it.
No functional changes.
Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-4-a5c1800dc930@kernel.org
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Move unshare_fd() where the rest of the descriptor table lifecycle
helpers live.
No functional changes.
Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-3-a5c1800dc930@kernel.org
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Add switch_files_struct() to install another table on a task. It
consumes the reference to the new table and puts the old one. Convert
every place that switches a descriptor table except unshare_files().
No functional changes.
Link: https://patch.msgid.link/20260910-work-coredump-unlock-self-v4-2-a5c1800dc930@kernel.org
Reviewed-by: NeilBrown <neil@brown.name>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|