| Age | Commit message (Collapse) | Author |
|
DCACHE_PRIVATE may be used by any filesystem for its own purposes, much
like d_fsdata and d_time.
I plan to use this in a similar manner the way nfs stores
NFS_FSDATA_BLOCKED in d_fsdata.
Signed-off-by: NeilBrown <neil@brown.name>
Link: https://patch.msgid.link/20260904215142.1060510-8-neilb@ownmail.net
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
DCACHE_PAR_LOOKUP acts like a lock in that threads can block waiting for
it to clear. As we plan to make changes to lock order for this lock,
teach lockdep to monitor it so as to help detect bugs early.
As NFS allocates an in-lookup dentry to unlink a silly-renamed file, and
completes the lookup in a different thread, we need interfaces to
release and the acquire ownership of the lock. This avoids lockdep
complaining that a lock is still held on return to user-space.
Signed-off-by: NeilBrown <neil@brown.name>
Link: https://patch.msgid.link/20260904215142.1060510-7-neilb@ownmail.net
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Some ->lookup handlers will need to drop and retake the parent lock, so
they can safely use d_alloc_parallel().
->lookup can be called with the parent lock either exclusive or shared.
A new flag, LOOKUP_SHARED, tells ->lookup how the parent is locked.
This is rather ugly, but will be gone soon after we move
d_alloc_parallel() out of the directory lock as ->lookup() will *always*
called with a shared lock on the parent.
Signed-off-by: NeilBrown <neil@brown.name>
Link: https://patch.msgid.link/20260904215142.1060510-6-neilb@ownmail.net
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Occasionally a single operation can require two sub-operations on the
same name, and it is important that a d_alloc_parallel() (once that can
be run unlocked) does not create another dentry with the same name
between the operations.
Two examples:
1/ rename where the target name (a positive dentry) needs to be
"silly-renamed" to a temporary name so it will remain available on the
server (NFS and AFS). Here the same name needs to be the subject
of one rename, and the target of another.
2/ rename where the subject needs to be replaced with a white-out
(shmemfs). Here the same name need to be the subject of a rename
and the target of a mknod()
In both cases the original dentry is renamed to something else, and a
replacement is instantiated, possibly as the target of d_move(), possibly
by d_instantiate().
Currently d_alloc() is used to create the dentry and the exclusive lock
on the parent ensures no other dentry is created. When
d_alloc_parallel() is moved out of the parent lock, this will no longer
be sufficient. In particular if the original is renamed away before the
new is instantiated, there is a window where d_alloc_parallel() could
create another name. "silly-rename" does work in this order. shmemfs
whiteout doesn't open this hole but is essentially the same pattern and
should use the same approach.
The new d_duplicate() creates an in-lookup dentry with the same name as
the original dentry, which must be hashed. There is no need to check if
an in-lookup dentry exists with the same name as d_alloc_parallel() will
never try add one while the hashed dentry exists. Once the new
in-lookup is created, d_alloc_parallel() will find it and wait for it to
complete, then use it.
Signed-off-by: NeilBrown <neil@brown.name>
Link: https://patch.msgid.link/20260904215142.1060510-5-neilb@ownmail.net
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Several filesystems use the results of readdir to prime the dcache.
These filesystems use d_alloc_parallel() which can block if there is a
concurrent lookup. Blocking in that case is pointless as the lookup
will add info to the dcache and there is no value in the readdir waiting
to see if it should add the info too.
Also these calls to d_alloc_parallel() are made while the parent
directory is locked. A proposed change to locking will lock the parent
later, after d_alloc_parallel(). This means it won't be safe to wait in
d_alloc_parallel() while holding the directory lock.
So this patch introduces d_alloc_trylock() which doesn't block but
instead returns ERR_PTR(-EWOULDBLOCK). Filesystems that prime the
dcache (smb/client, nfs, fuse, cephfs) can now use that and ignore
-EWOULDBLOCK errors as harmless.
Unlike d_alloc_parallel(), d_alloc_trylock() calculates the hash and
performs a lookup before an allocation, as that is what all callers
want. This is done using try_lookup_noperm(), necessitating the
inclusion of namei.h in dcache.c.
Signed-off-by: NeilBrown <neil@brown.name>
Link: https://patch.msgid.link/20260904215142.1060510-4-neilb@ownmail.net
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Cachefiles currently uses the backing filesystem's idea of what data is
held in a backing file and queries this by means of SEEK_DATA and
SEEK_HOLE. However, this means it does two seek operations on the backing
file for each individual read call it wants to prepare (unless the first
returns -ENXIO). Worse, the backing filesystem is at liberty to insert or
remove blocks of zeros in order to optimise its layout which may cause
false positives and false negatives.
The problem is that keeping track of what is dirty is tricky (if storing
info in xattrs, which may have limited capacity and must be read and
written as one piece) and expensive (in terms of diskspace at least) and is
basically duplicating what a filesystem does.
However, the most common write case, in which the application does {
open(O_TRUNC); write(); write(); ... write(); close(); } where each write
follows directly on from the previous and leaves no gaps in the file is
reasonably easy to detect and can be noted in the primary xattr as
CACHEFILES_CONTENT_ALL, indicating we have everything up to the object size
stored.
In this specific case, given that it is known that there are no holes in
the file, there's no need to call SEEK_DATA/HOLE or use any other mechanism
to track the contents. That speeds things up enormously.
Even when it is necessary to use SEEK_DATA/HOLE, it may not be necessary to
call it for each cache read subrequest generated.
Implement this by adding support for the CACHEFILES_CONTENT_ALL content
type (which is defined, but currently unused), which requires a slight
adjustment in how backing files are managed. Specifically, the driver
needs to know how much of the tail block is data and whether storing more
data will create a hole.
To this end, the way that the size of a backing file is managed is changed.
Currently, the backing file is expanded to strictly match the size of the
network file, but this can be changed to carry more useful information.
This makes two pieces of metadata available: xattr.object_size and the
backing file's i_size. Apply the following schema:
(a) i_size is always a multiple of the DIO block size.
(b) i_size is only updated to the end of the highest write stored. This
is used to work out if we are following on without leaving a hole.
(c) xattr.object_size is the size of the network filesystem file cached
in this backing file.
(d) xattr.object_size must point after the start of the last block
(unless both are 0).
(e) If xattr.object_size is at or after the block at the current end of
the backing file (ie. i_size), then we have all the contents of the
block (if xattr.content == CACHEFILES_CONTENT_ALL).
(f) If xattr.object_size is somewhere in the middle of the last block,
then the data following it is invalid and must be ignored.
(g) If data is added to the last block, then that block must be fetched,
modified and rewritten (it must be a buffered write through the
pagecache and not DIO).
(h) Writes to cache are rounded out to blocks on both sides and the
folios used as sources must contain data for any lower gap and must
have been cleared for any upper gap, and so will rewrite any
non-data area in the tail block.
To implement this, the following changes are made:
(1) cookie->object_size is no longer updated when writes are copied into
the pagecache, but rather only updated when a write request completes.
This prevents object size miscomparison when checking the xattr
causing the backing file to be invalidated (opening and marking the
backing file and modifying the pagecache run in parallel).
(2) The cache's current idea of the amount of data that should be stored
in the backing file is kept track of in object->object_size.
Possibly this is redundant with cookie->object_size, but the latter
gets updated in some addition circumstances.
(3) The size of the backing file at the start of a request is now tracked
in struct netfs_cache_resources so that the partial EOF block can be
located and cleaned.
(4) The cache block size is now used consistently rather than using
CACHEFILES_DIO_BLOCK_SIZE (4096).
(5) The backing file size is no longer adjusted when looking up an object.
(6) When shortening a file, if the new size is not block aligned, the part
beyond the new size is cleared. If the file is truncated to zero, the
content_info gets reset to CACHEFILES_CONTENT_NO_DATA.
(7) A new struct, fscache_occupancy, is instituted to track the region
being read. Netfslib allocates it and fills in the start and end of
the region to be read then calls the ->query_occupancy() method to
find and fill in the extents. It also indicates whether a recorded
extent contains data or just contains a region that's all zeros
(FSCACHE_EXTENT_DATA or FSCACHE_EXTENT_ZERO).
(8) The ->prepare_read() cache method is changed such that, if given, it
just limits the amount that can be read from the cache in one go. It
no longer indicates what source of read should be done; that
information is now obtained from ->query_occupancy().
(9) A new cache method, ->collect_write(), is added that is called when a
contiguous series of writes have completed and a discontiguity or the
end of the request has been hit. It it supplied with the start and
length of the write made to the backing file and can use this
information to update the cache metadata.
(10) cachefiles_query_occupancy() is altered to find the next two "extents"
of data stored in the backing file by doing SEEK_DATA/HOLE between the
bounds set - unless it is known that there are no holes, in which case
a whole-file first extent can be set.
(11) cachefiles_collect_write() is implemented to take the collated write
completion information and use this to update the cache metadata, in
particular working out whether there's now a hole in the backing file
requiring future use of SEEK_DATA/HOLE instead of just assuming the
data is all present.
It also uses fallocate(FALLOC_FL_ZERO_RANGE) to clean the part of a
partial block that extended beyond the old object size. It might be
better to perform a synchronous DIO write for this purpose, but that
would mandate an RMW cycle. Ideally, it should be all zeros anyway,
but, unfortunately, shared-writable mmap can interfere.
(12) cachefiles_begin_operation() is updated to note the current backing
file size and the cache DIO size.
(13) cachefiles_create_tmpfile() no longer expands the backing file when it
creates it.
(14) cachefiles_set_object_xattr() is changed to use object->object_size
rather than cookie->object_size.
(15) cachefiles_check_auxdata() is altered to actually store the content
type and to also set object->object_size. The cachefiles_coherency
tracepoint is also modified to display xattr.object_size.
(16) netfs_read_to_pagecache() is reworked. The cache ->prepare_read()
method is replaced with ->query_occupancy() as the arbiter of what
region of the file is read from where, and that retrieves up to two
occupied extents of the backing file at once.
The cache ->prepare_read() method is now repurposed to be the same as
the equivalent network filesystem method and allows the cache to limit
the size of the read before the iterator is prepared.
netfs_single_dispatch_read() is similarly modified.
(17) netfs_update_i_size() and afs_update_i_size() no longer call
fscache_update_cookie() to update cookie->object_size.
(18) Write collection now collates contiguous sequences of writes to the
cache and calls the cache ->collect_write() method.
Signed-off-by: David Howells <dhowells@redhat.com>
Link: https://patch.msgid.link/20260910220242.2165023-5-dhowells@redhat.com
Reviewed-by: Paulo Alcantara <pc@manguebit.org>
cc: Matthew Wilcox <willy@infradead.org>
cc: linux-cifs@vger.kernel.org
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Add support for Qualcomm Messaging Interface (QMI) based Thermal Mitigation
Device (TMD) cooling devices provided by remote subsystems.
On Qualcomm platforms where remote processors expose mitigation controls
through the TMD QMI service, client drivers need support to discover the
service, register cooling devices for available mitigation endpoints,
and forward cooling state updates to remote subsystems.
Signed-off-by: Casey Connolly <casey.connolly@linaro.org>
Co-developed-by: Daniel Lezcano <daniel.lezcano@oss.qualcomm.com>
Signed-off-by: Daniel Lezcano <daniel.lezcano@oss.qualcomm.com>
Co-developed-by: Gaurav Kohli <gaurav.kohli@oss.qualcomm.com>
Signed-off-by: Gaurav Kohli <gaurav.kohli@oss.qualcomm.com>
Link: https://lore.kernel.org/r/20260809-b4-qmi-tmd-v8-2-b15d47adc379@oss.qualcomm.com
Signed-off-by: Bjorn Andersson <andersson@kernel.org>
|
|
Cortex-R52 cores can boot from DDR. In such case, elf's boot address
will be in the ddr, and it should be passed to the PLM via EEMI call.
Add ddr boot config EEMI call before starting RPU to configure the boot
address. Also make sure TCM-Boot is support for old platforms as well.
Remove HIVEC & LOVEC enum, as now PLM supports boot address directly.
Signed-off-by: Tanmay Shah <tanmay.shah@amd.com>
Link: https://patch.msgid.link/20260803200502.209361-1-tanmay.shah@amd.com
Signed-off-by: Michal Simek <michal.simek@amd.com>
|
|
The Versal Gen 2 UFS driver polls the firmware for M-PHY TX/RX
configuration readiness and SRAM initialisation completion with two
open-coded do/while loops. Each iteration is a full firmware round-trip
(PM_IOCTL/IOCTL_READ_REG of a protected PMC_IOU_SLCR register), so the
loop can issue up to a million EEMI calls, and it hard-codes the wait
policy inside the controller driver.
Introduce coarse blocking helpers, zynqmp_pm_wait_mphy_tx_rx_config_ready()
and zynqmp_pm_wait_sram_init_done(), that take a caller-supplied timeout
budget and contain the poll loop. The loop is EEMI-specific (legacy
firmware only exposes the per-read status primitive) so it lives in the
firmware driver, keeping the UFS driver backend-agnostic: a future
backend can offload the wait to the platform in a single call without
touching the controller driver again. The existing per-read primitives stay
exported, so the current EEMI interface is unchanged.
The timeout budget remains owned by the UFS driver (the consumer that
knows the hardware) and is passed down, so EEMI and any future backend
stay consistent.
Reviewed-by: Sai Krishna Potthuri <sai.krishna.potthuri@amd.com>
Link: https://patch.msgid.link/eaeaaea8ed76069943e0706a0917139ff6569929.1785855749.git.michal.simek@amd.com
Signed-off-by: Michal Simek <michal.simek@amd.com>
|
|
Drop a superfluous empty line from the end of the 'bam_positions'
structure group within the 'bam_transaction' structure.
No functional changes.
Signed-off-by: Gabor Juhos <j4g8y7@gmail.com>
Signed-off-by: Miquel Raynal <miquel.raynal@bootlin.com>
|
|
Add the sub-cache table for Nord, derived from the Nord AU (Safe IVI)
SCT table and limited to the POR use cases that have a unique SCID and a
usecase ID in llcc-qcom.h. Nord has the CFG0/CFG1 pairs and the full
ALGO_ALLOC0..3 set, so it uses the v6 register offsets.
num_banks is set explicitly because LLCC_LB_CNT in LLCC_COMMON_STATUS0
is four bits wide and cannot report sixteen banks; a wrong count makes
the driver map a bank aperture in place of the broadcast region and abort
on the first broadcast read. no_edac is set because the bootloader
leaves the LLCC EDAC registers locked, as on Glymur.
Signed-off-by: Anurag Pateriya <anurag.pateriya@oss.qualcomm.com>
Tested-by: Zhangfei Gao <zhangfei.gao@oss.qualcomm.com>
Reviewed-by: Shawn Guo <shengchao.guo@oss.qualcomm.com>
Reviewed-by: Abel Vesa <abel.vesa@oss.qualcomm.com>
Reviewed-by: Komal Bajaj <komal.bajaj@oss.qualcomm.com>
Reviewed-by: Konrad Dybcio <konrad.dybcio@oss.qualcomm.com>
Link: https://patch.msgid.link/20260910-nord_llcc-v1-2-30fd1cca65fe@oss.qualcomm.com
Signed-off-by: Bjorn Andersson <bjorn.andersson@oss.qualcomm.com>
|
|
Add the missing kernel-doc description for the new break_wait field in
struct tty_struct.
Fixes: 845a2737e0f5 ("tty: abort break signalling on hangup")
Reported-by: Randy Dunlap <rdunlap@infradead.org>
Link: https://lore.kernel.org/20d75a7a-2e7a-484f-a267-c21db82dd922@infradead.org
Signed-off-by: Johan Hovold <johan@kernel.org>
Link: https://patch.msgid.link/20260925094222.18303-1-johan@kernel.org
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
Add support for the AMD IOMMU Performance Optimization (PerfOpt)
feature as defined in the AMD I/O Virtualization Technology (IOMMU)
Specification, Section 3.4.9 (MMIO Offset 016Ch).
This feature allows privileged integrated I/O devices (GPUs) to bypass
the IOMMU when directly accessing system memory. The IOMMU only
enforces the IR/IW permission bits without GPA->SPA translations.
amd_iommu_enable_perfopt() performs a detach/reattach cycle to rehome
devices already on the identity domain with ATS/PRI/PASID/GCR3
disabled (skip_caps path). amd_iommu_disable_perfopt() restores those
capabilities. The per-device dev_data->perfopt flag tracks state.
PERF_OPT_EN is a single control bit per IOMMU, shared by every device
behind that IOMMU, while enablement is requested per device. It is
therefore reference counted (amd_iommu->perfopt_refcount): armed on the
first requesting device and cleared on the last, so one device's
teardown never clears the bit while a peer behind the same IOMMU still
needs it.
The per-device flag is cleared on every teardown path
(blocked_domain_attach, release_device, and amd_iommu_disable_perfopt),
dropping the reference with it, so a reused dev_data never carries stale
PerfOpt state onto its next bind.
On suspend/resume the hardware is reprogrammed from scratch:
amd_iommu_perfopt_clear() forces the bit off without touching the
reference count, and amd_iommu_perfopt_restore() re-asserts it from the
count after early_enable_iommu(), so armed devices keep the optimization
across resume without relying on each consumer driver to re-arm.
The exported amd_iommu_enable_perfopt()/amd_iommu_disable_perfopt() run
only from a consumer driver's bind/unbind path. group->mutex is not
exposed to drivers, but a device bound to its native driver cannot have
its IOMMU domain changed concurrently by the core, which serializes the
detach/attach pair against core-driven attach.
PerfOpt is opt-in -- only enabled when explicitly requested by a driver.
Tested-by: Derek J. Clark <derekjohn.clark@gmail.com>
Co-developed-by: Jatin Kataria <jkataria@netflix.com>
Signed-off-by: Jatin Kataria <jkataria@netflix.com>
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
Tested-by: Boqun Feng (Netflix) <boqun@kernel.org>
Acked-by: Alex Deucher <alexander.deucher@amd.com>
[joerg.roedel@amd.com: Fixed build failure on i386]
Signed-off-by: Joerg Roedel <joerg.roedel@amd.com>
|
|
All existing architectures which implement KVM pre-fault (x86, s390) have
mechanisms for disallowing pre-faulting.
Currently these are open coded as part of kvm_arch_vcpu_pre_fault_memory().
Formalise this by moving them into a new kvm_arch_pre_fault_allowed() hook,
which every architecture selecting CONFIG_KVM_GENERIC_PRE_FAULT_MEMORY must
implement, returning an error code if the operation is disallowed or 0
otherwise.
The hook is called early in the generic code, allowing architectures to
disallow the operation prior to vCPU load.
This is important, as kvm_vcpu_pre_fault_memory() is the only place where
generic code can call vcpu_load() on a vCPU that has not yet been
initialised.
This lays the foundation for a future change which implements pre-faulting
for arm64 which needs to disallow the mechanism for uninitialised vCPUs.
Suggested-by: Oliver Upton <oupton@kernel.org>
Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Acked-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Tested-by: Fuad Tabba <fuad.tabba@linux.dev>
Link: https://patch.msgid.link/20260923-kvm-arm-prefault-v4-1-d4b0b4dfa8c3@kernel.org
Signed-off-by: Marc Zyngier <maz@kernel.org>
|
|
|
|
|
|
The transfer timeout for an I2C controller should reflect the actual
message length and bus frequency rather than a static 1-second value.
A static timeout causes unnecessary delays on error paths for short
messages, and may be insufficient for very long transfers.
Add i2c_update_timeout() to i2c-core which computes a transfer-specific
timeout and stores it directly in the standard adap->timeout field. The
formula accounts for 9 bits per byte (8 data + 1 ACK) at the configured
bus frequency. The caller supplies a safety multiplier and a minimum
floor so that each driver retains full control over its timing policy
without those values becoming public API.
Storing the result in adap->timeout makes it visible to all consumers of
that field, including the arbitration-loss retry loop in __i2c_transfer().
The function is gated by CONFIG_I2C_DYNAMIC_TIMEOUT. When the config is
disabled, i2c_update_timeout() compiles to a no-op inline stub so drivers
that call it build cleanly and the existing static 1-second default is
preserved unchanged.
A timeout explicitly configured by userspace via the I2C_TIMEOUT ioctl is
stored in a new adap->user_timeout field and always takes precedence over
the kernel-computed value. When userspace has not configured a timeout,
the computed value is used. The ioctl keeps writing adap->timeout as well,
so adapters that never call i2c_update_timeout() continue to honour it
exactly as before.
As i2c_update_timeout() is an exported helper, guard against a zero bus
frequency from a misbehaving caller with WARN_ON_ONCE() and return early,
leaving the existing timeout untouched as a safe fallback rather than
dividing by zero.
Signed-off-by: Aniket Randive <aniket.randive@oss.qualcomm.com>
Reviewed-by: Mukesh Kumar Savaliya <mukesh.savaliya@oss.qualcomm.com>
Signed-off-by: Andi Shyti <andi.shyti@kernel.org>
Link: https://patch.msgid.link/20260911-master-v9-1-77ac458344e2@oss.qualcomm.com
|
|
Commit 9f3d900a69be ("bpf: Grow the verifier id scratch on demand") turned
the id map comparing the ids of two states, which also serves as the id
stack of release_reference() and as the id set of
bpf_clear_singular_ids(), into arrays grown on demand, so that the env
would not embed a map sized for 2 KiB frames. That gave the map a way to
fail, and with it release_reference() and the iterator destroy path an
-ENOMEM to hand up, states_equal() a case where states compare unequal
for lack of memory, and bpf_clear_singular_ids() a flag to keep every id
when the set could not record them all.
The env is allocated once per program load, so the map can afford to be
sized for the largest state again: MAX_BPF_STACK_SLOTS in each of
MAX_CALL_FRAMES frames, which a state on a private stack can reach since
every frame there may use the whole budget. That makes the map 34 KiB
instead of 10 KiB, a small price per load for having no growth that can
fail. Make the map and the set a fixed union in the env again and drop
the growth and its failure handling.
Suggested-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/DLNQW1IRYK94.1DP0M5VKZSY32@gmail.com
Link: https://patch.msgid.link/20260924193403.3032037-1-memxor@gmail.com
|
|
Cross-merge networking fixes after downstream PR (net-7.3-rc5).
Conflicts:
drivers/net/mdio/mdio-realtek-rtl9300.c
89a8a1eef2d44 ("net: mdio: realtek-rtl9300: fix RTL931x C22 extended page selection")
cc3cb8db1eef9 ("net: mdio: realtek-rtl9300: Add page tracking")
https://lore.kernel.org/arUZOqy73bE2pp0w@sirena.org.uk
net/8021q/vlan_dev.c
cd5dd68267c4 ("vlan: ensure sufficient headroom in vlan_dev_hard_header()")
ca6ff8dd70eb ("vlan: annotate data-races in vlan_dev_priv fields")
Adjacent changes:
net/ipv6/ip6_gre.c
dd47bcf279f1 ("ip6_gre: Call ip6erspan_tunnel_unlink_md() in ip6erspan_changelink().")
cce829e2aa1d ("ip6_gre: Protect ip6gre_net.tunnels[][] with mutex.")
drivers/net/ethernet/meta/fbnic/fbnic_txrx.c
b5d9e9d4d0c1 ("eth: fbnic: use the Rx queue napi pointer to find the napi vector")
c0aca269ec07 ("eth: fbnic: Make Rx completion coalescing configurable")
drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c
d68acbf93531 ("net: stmmac: selftests: Support running selftests on DSA conduits")
85ca3292d7a3 ("net: stmmac: Remove ARP offload code")
drivers/net/ethernet/wangxun/libwx/wx_hw.c
3173cba11701 ("net: libwx: fix races in Tx timestamp handling")
7042c8c193e5 ("net: libwx: rename wx_pf_flags to wx_flags")
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Introduce acpi_object_free for freeing union acpi_object objects
allocated by AML and use it for simplifying AML error handling in
the ACPI fan driver.
While at it, update the driver to use consistent error values across
all function using the union acpi_object data type.
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
Reviewed-by: Andy Shevchenko <andriy.shevchenko@linux.intel.com>
Reviewed-by: Armin Wolf <W_Armin@gmx.de>
[ rjw: Put the new free definition under CONFIG_ACPI ]
Link: https://patch.msgid.link/6029013.DvuYhMxLoT@rafael.j.wysocki
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
|
|
BPF crypto was implemented using the lskcipher API, which doesn't seem
to be going anywhere. lskcipher supports only ARC4 and block ciphers in
CBC and ECB mode, and only with unoptimized implementations.
Library APIs also have been found to be a much better approach, for a
variety of reasons, including reduced overhead, greater flexibility, and
having to be explicit about the crypto algorithms that are supported.
We can safely ignore the theoretical ARC4 and non-AES block cipher
support in BPF crypto as unused, which leaves AES-CBC and AES-ECB.
AES-CBC was stated to be needed for decrypting packets using a homebrew
UDP-based protocol
(https://lore.kernel.org/r/d1cdfc23-b336-49a9-8833-29f05b5b9fec@linux.dev/).
AES-ECB is used by the BPF self-tests, and it was stated to maybe be
used in the future for QUIC-LB
(https://lore.kernel.org/r/5f9c3aab-5339-463c-a86d-edac297e1e95@linux.dev/).
Those reasons don't make much sense either, especially AES-ECB which
isn't appropriate in new systems and should be dropped. Regardless,
let's assume that both AES-CBC and AES-ECB need to be kept for now.
There are library APIs for both of these now, which are much easier to
use and more efficient. Reimplement BPF crypto on top of them, greatly
simplifying the code. As part of this, the bpf_crypto_type abstraction
layer is removed, as it's not useful.
This removes the only user of the lskcipher API, so follow-up patches
will be able to remove that code as well.
Signed-off-by: Eric Biggers <ebiggers@kernel.org>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924184254.142720-1-ebiggers@kernel.org
|
|
Annotate the "configs" pointer field of "struct
tegra210_clk_emc_provider" with the "__counted_by_ptr" attribute,
allowing the compiler to perform runtime bounds checking on accesses to
"configs" based on the value of "num_configs".
The "emc->provider.configs = devm_kcalloc(...)" call uses
"emc->num_timings" for the number of elements, which is then assigned to
"emc->provider.num_configs" before any accesses to "configs". The
"num_configs" field isn't modified after assignment.
Cc: codemender-patching+linux@google.com
Assisted-by: LLM
Signed-off-by: Bill Wendling <morbo@google.com>
Reviewed-by: Gustavo A. R. Silva <gustavoars@kernel.org>
Reviewed-by: Kees Cook <kees@kernel.org>
Signed-off-by: Brian Masney <bmasney@redhat.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Jakub Kicinski:
"Including fixes from Bluetooth, NFC and Netfilter.
Every week in this release is record-setting for number of posted
patches. It doesn't seem like we're creating any regressions with all
these fixes, three 'Fixes' tags here point to 7.2 commits but none are
true regression fixes. We're trying to keep the count down,
nonetheless.
Previous releases - regressions:
- net: don't require the hwtstamp NDOs when a PHY provides
timestamping
- ipv6: fix dst leak for uncached routes
- vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE
packets
Previous releases - always broken:
- packet: use ubuf_info completion for TX_RING packets
- arp: terminate device name before lookup
- ipv6: do not let ipv6_find_hdr() return an offset past the packet
end
- udp: remove a disconnected socket from the 4-tuple hash table
- sctp: discard the rest of the packet on a stale-cookie error
- eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged
eswitch"
[ And lots of other random network driver fixes ]
* tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (189 commits)
tcp: prevent collapsing skbs across boundary in rtx queue
vlan: ensure sufficient headroom in vlan_dev_hard_header()
net/sched: sch_teql: fix shadowed err in __teql_resolve()
bridge: check llc_mac_hdr_init() return value in br_send_bpdu()
llc: fix skb UAF and leaks on llc_mac_hdr_init() failure
llc: reserve device headroom for allocated frames
gve: DQO: reject TSO packets with an out of range MSS
gve: fix TX drop when GSO MSS is too small for hw
gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO
net: flush skb_defer_nodes in dev_cpu_dead()
net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails
af_packet: fix integer overflow in prb_calc_retire_blk_tmo()
tipc: Fix a data race on mon->peer_cnt in mon_timeout()
net: phy: intel-xway: workaround 100BASE-TX Link-Up issue
net/smc: fix UAF on lgr list traversal in smcr_port_err()
net/rds: size a connection's path set by the transport it ends up with
nfp: hold IPsec RX state under the XArray lock
net: ena: fix MMIO read buffer leak on probe failure
net: ena: fix PHC cleanup on probe failure
net/sched: act_ct: fix helper UAF due to extensions realloc
...
|
|
The verifier checks every stack access and the combined depth of a call
chain against MAX_BPF_STACK, which is also the frame size of the
interpreter and the frame that JITs without subprogram tail call
support set up for tail-call targets. A JIT that lays out frames of any
size and lets a tail-called program set up its own frame does not need
that limit; it only needs the verifier to bound how much stack a
program uses in total.
Add bpf_jit_supports_large_stack() for a JIT to claim that, and give
each program its budget through bpf_prog_stack_limit(): MAX_BPF_STACK_JIT
when the JIT is requested, the program is not offloaded and the JIT
supports large stacks as well as tail calls from subprograms,
MAX_BPF_STACK otherwise. The latter is what lets a tail-called program
set up its own frame: without it, do_misc_fixups() gives every program
with tail calls a MAX_BPF_STACK frame, which a deeper frame verified
against the larger budget would overrun. The verifier keeps the budget
in env->stack_limit and uses it for the bounds of fixed and variable
offset stack accesses, for unprivileged stack pointer arithmetic and its
speculation limit, and for the combined and private stack depth checks.
A frame may use any part of its program's budget. The interpreter paths
keep MAX_BPF_STACK: a program whose main frame is deeper falls back to
the JIT-required path of bpf_prog_select_runtime() and one with deeper
subprogram frames is rejected when patching calls for the interpreter.
The extra stack that may_goto and the timed may_goto instrumentation
add below a frame is, as before, not counted against the budget of a
JITed program and rejected past MAX_BPF_STACK for an interpreted one.
Stack liveness treats a read through a pointer of unknown offset, or a
call passing a frame pointer to a subprogram, as reaching the whole
frame, and widens the masks of that frame to the deepest half-slot such
a read can cover. Bound that by the program's budget too: no access
past it is accepted, so a program kept at MAX_BPF_STACK carries masks
of two words for such frames, as before, instead of the eight that
MAX_BPF_STACK_JIT needs. The three selftests matching a whole-frame
read in the liveness log accept either depth.
bpf_clone_redirect() transmits from inside the program, and a tc egress
or lwt_xmit program that redirects to its own device runs again on top
of its own frame until the datapath's recursion limit drops the packet,
ten frames deep. Ten MAX_BPF_STACK frames fit the kernel stack as they
always did; ten MAX_BPF_STACK_JIT frames would not, so a program that
calls bpf_clone_redirect() keeps MAX_BPF_STACK. The redirect helpers
that transmit after the program has returned leave no frame behind and
do not affect the budget. Nesting through other attach points is not
accounted, as before.
The spill tracker of the liveness analysis follows no slot past the
budget either.
The capability is a boolean and the budget a single constant, in the
style of the other bpf_jit_supports_*() queries, rather than a per JIT
size: the budget is meant to be the same everywhere it is raised, so
that programs verify identically across those architectures.
No JIT declares support yet, so every program keeps its 512-byte budget.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924165740.2146806-14-memxor@gmail.com
|
|
The verifier keeps a few structures whose size follows the deepest
frame a program may have: the backtracking and scratched-slot bitmaps,
the jump history slot index and the clamp of the liveness masks. They
are all expressed through MAX_BPF_STACK_SLOTS, which derives from
MAX_BPF_STACK, the frame size of the interpreter.
Introduce MAX_BPF_STACK_JIT, the stack budget a program may get on a
JIT that can lay out frames of any size, and derive those structures
from it so that a frame may be as deep as that budget. Nothing grants
the budget yet, so no program verifies differently; the only visible
change is that the liveness log prints a whole-frame read up to the new
depth, so the three selftests matching such reads are updated.
The spill tracker of the liveness analysis keeps a table entry per
instruction and tracked slot, so it follows more than the 64 slots of a
MAX_BPF_STACK frame only while that table stays within what 64 slots
need for the largest program; a subprog of a million instructions keeps
64, one of a quarter million may track all 256. This bounds the table
at its old worst case of 640 MiB instead of letting a single deep store
push it past what kvmalloc() serves.
The backtracking and scratched-slot bitmaps grow from one to four
words per frame, a fixed few hundred bytes per verifier environment.
tmp_str_buf, which formats a frame's slot list for the log, grows from
320 to 1408 bytes so that all 256 slots still fit, and the log's line
buffer from 1 to 2 KiB so that a line built from it is not cut; the
environment stays within its 64 KiB allocation. The liveness masks are
only as wide as the stack a frame uses, so most frames cost the same as
before; a frame that is read as a whole, through a pointer of unknown
offset or by bpf_loop() with two callbacks, now carries masks of eight
words, 192 bytes per instruction per frame instead of 48. Measured over
the 5075 selftest programs, that is 0.2% of the total peak verifier
memory: strobemeta_bpf_loop and pyperf600_bpf_loop grow by 11% (1.2 MiB
and 0.6 MiB), a few dozen small programs by 40 to 100 KiB each,
everything else is unchanged. The next patch bounds such reads by the
program's budget, so this cost is only paid once a JIT grants it.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924165740.2146806-13-memxor@gmail.com
|
|
The id map used to compare the ids of two states, which also serves as
the id stack of release_reference() and as the id set of
bpf_clear_singular_ids(), is a fixed array embedded in struct
bpf_verifier_env and sized for the most registers and stack slots a
state can possibly hold: 1312 entries, 10 KiB, for 512-byte frames, and
four times that once frames may reach 2 KiB, which pushes the env
allocation from 64 KiB to 128 KiB for every program verified. States
compare a few dozen ids in practice.
Turn the map and the set into arrays grown on demand, starting at 64
entries and doubling, and free them with the env. A map that cannot
grow treats the states as different, an id set that cannot grow keeps
every id of the state, since its counts are then incomplete, and the id
stack reports -ENOMEM, which release_reference()
hands to its callers. The iterator destroy path used to warn on any
failure of release_reference(), which could only come from a bug while
the id stack could not run out of room; it now returns -ENOMEM and
keeps warning about anything else. An allocation failure thus fails
the load and is never unsafe. This removes the last structure whose
size scaled with the stack bound and shrinks the env by 10 KiB.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924165740.2146806-9-memxor@gmail.com
|
|
Stack liveness tracks three masks of 4-byte stack slots - may_read,
must_write and live_before - for every instruction of every frame of
every function instance. Each mask used to be a fixed 128-bit spis_t,
wide enough for the deepest frame MAX_BPF_STACK allows, so an
instruction paid 48 bytes per frame no matter how little stack the frame
actually touched. These per instruction arrays are the largest liveness
allocation: a 40k instruction program spends ~2 MiB on each frame array.
Sizing them for a larger stack budget would multiply that cost for every
program, while most frames use a fraction of the budget.
Replace the fixed masks with a variable width bitmap array per
(instance, frame). struct frame_masks holds the width shared by all of
its masks in @words plus a flat unsigned long array, where mask @kind of
the instruction at relative index @i lives at
&bits[(i * FM_MASK_CNT + kind) * words]. An array starts at the width
the first recorded half-slot needs (minimum one word, that is 256 bytes
of stack with 64-bit words) and is reallocated into a wider stride once
a deeper half-slot shows up. Only the marking functions widen an array,
and they all run during the static analysis before do_check(), so no
verification time path can move one.
The marking API now takes an inclusive half-slot range, which is what
record_stack_access_off() computed anyway, and clamps the upper end to
the deepest half-slot a frame can have. An access below the stack bound
is rejected by the main verifier pass later; liveness only has to stay
inside the masks. A range that covers no whole half-slot, as a byte
store at fp-1 produces, marks nothing.
A "read everything" mark, used for calls that are neither helpers nor
kfuncs and for pointers whose offset or frame identity was lost, widens
the array to the maximum width and sets every bit: recorded at a
narrower width it would lose the half-slots a later widening adds, and
no write can cancel it as no write covers the whole frame. A half-slot
past the width of the masks was never read by the frame and is never
live.
merge_instances() may see a different width on each side, so it widens
dst to cover src and counts a word only dst has as zero on the src
side, which unions into may_read as a no-op and intersects must_write
to empty, never claiming a write that did not happen.
The rest is mechanical: update_insn() propagates live_before word by
word over the array's width (all instructions of one array share it,
hence so do an instruction's successors), is_live_before() answers
false past the width, and the use:/def: log printing iterates up to the
width.
Liveness results are unchanged. A half-slot only ever gets a bit from a
recorded access and recording one widens the array to cover it, so every
bit the fixed masks could hold still fits. Frames whose deepest recorded
half-slot fits in a single word now cost 24 bytes per instruction per
frame instead of 48.
The spill tracker that feeds the analysis had the same limit built in: it
followed 64 slots, fp-8 to fp-512, per subprog, and a fill from a deeper
slot returned an imprecise pointer, so every later access through it
counted as a read of the whole stack of every frame. Let it follow the
deepest 8-byte access a subprog makes directly through R10 when that
reaches past those 64 slots, which is how spills and fills are compiled,
and keep the count with the per-callsite snapshot a callee uses to fill
from its caller's frame. Programs within 512 bytes are tracked exactly as
before; a program spilling deeper stays precise.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924165740.2146806-8-memxor@gmail.com
|
|
The set of stack slots touched since the verifier state was last
printed lives in a u64, so the log could only ever mark the first 64
slots of a frame as scratched. Make it a bitmap sized by
MAX_BPF_STACK_SLOTS and go through the bitmap helpers for setting,
testing, clearing and filling it.
No functional change.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924165740.2146806-6-memxor@gmail.com
|
|
Precision backtracking keeps the stack slots that still need a precise
mark in one u64 per frame, which ties it to frames of at most 64 slots.
Turn the per-frame masks into bitmaps sized by MAX_BPF_STACK_SLOTS and
use the bitmap helpers for setting, clearing, testing and iterating
them. The formatting helper takes a bitmap and the leftover-slot bug
reports print the formatted slot list instead of a hex mask.
mark_reg_stack_read() collected zero spills in a u64 of its own before
handing it to the backtracker; it now counts them and revisits the
range to mark each slot, which drops the only remaining mask-typed
entry point.
No functional change.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924165740.2146806-5-memxor@gmail.com
|
|
Linked scalar registers are recorded in the jump history packed into a
u64 as five 11-bit entries, each naming a frame and a register or stack
slot. The 6-bit slot field covers exactly the 64 slots of a 512-byte
frame, so a spilled scalar in a deeper slot could not be linked and
larger frames were ruled out by construction.
Store the linked registers as an array of five u16 entries plus a count
instead, each entry holding the frame number, a register-or-slot bit
and an 11-bit register or slot index, which covers frames of up to
16 KiB. Callers that record no linked registers pass NULL.
The history entry grows from 16 to 20 bytes, and the history is the
one verifier structure whose size follows the number of instructions a
loop iterates over rather than the state count. Measured over the 5075
selftest programs, peak verifier memory is unchanged for all but the
loop-heavy ones, which grow by 8 to 16%: loop1/nested_loops from 17.6 to
19.2 MiB, verifier_loops1/jumps_out_rather_than_in from 4.4 to 5.1 MiB,
strobemeta by 0.3% and pyperf600_nounroll by 0.8%. Verdicts, processed
instructions and state counts stay the same everywhere.
No functional change.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924165740.2146806-4-memxor@gmail.com
|
|
Jump history entries record the stack slot touched by a spill or fill in
a 6-bit field, which only fits the 64 slots of a 512-byte frame and is
pinned to that size by a static_assert on MAX_BPF_STACK. Move the flags
into the first word and give the slot index 12 bits of the second word
instead, so the entry stays 16 bytes while frames of up to 32 KiB can
be recorded.
Introduce MAX_BPF_STACK_SLOTS for the number of 8-byte slots a frame can
have and use it for the static_assert and for BPF_ID_MAP_SIZE, so the
verifier expresses per-frame capacity through one constant.
No functional change.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924165740.2146806-3-memxor@gmail.com
|
|
The verifier indexes a frame's stack state directly through
state->stack[spi] and computes the number of tracked slots as
allocated_stack / BPF_REG_SIZE in every file that touches stack slots,
and so does the nfp offload driver. Route all of these through two
helpers, bpf_stack_slot() and bpf_stack_nr_slots(), so the layout of the
per-frame stack state is visible in one place. The one lookup that goes
the other way, reg_to_target() in the diagnostics code, maps a register
pointer back to its slot by address and now says that it relies on the
slots forming one contiguous array. Both helpers take a const frame: the
slot accessor returns the slot through the frame's stack pointer, so
read-only code such as the state printer can use it without giving up
its qualifiers. Functions that look up the same slot repeatedly now
fetch it once.
No functional change.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924165740.2146806-2-memxor@gmail.com
|
|
scx_dispatch_enqueue() calls rb_add() for a vtime-ordered DSQ while
holding dsq->lock. At each level of the descent, scx_dsq_priq_less()
loads the visited task's dsq_vtime, and that comparison selects rb_left
or rb_right from the same task's dsq_priq. In the x86-64 benchmark
configuration, these accesses fell on different 64-byte cache lines
before this change, so a cold visited task could require a second cache
line fill on the dependent descent path.
Move dsq_vtime immediately before dsq_priq, keeping dsq_seq and
dsq_flags next to dsq_list. In the two x86-64 layouts inspected, where
scx starts at offsets 776 and 840 in task_struct, the fields from
dsq_list through dsq_priq occupy one cache line. The exact cache line
placement depends on the containing task_struct layout and is not
guaranteed for every configuration or architecture. Where they share a
line, avoiding the dependent cache line fill helps reduce dsq->lock hold
time during vtime insertion and therefore contention on the lock.
In a 16-vCPU KVM guest, the instrumented enqueue interval was 4-6%
shorter with scx_lavd and scx_layered. scx_mitosis and end-to-end
workload time showed no consistent change.
This only reorders fields; no scheduling behavior change is intended.
Signed-off-by: Usama Arif <usama.arif@linux.dev>
Signed-off-by: Tejun Heo <tj@kernel.org>
|
|
Cross-merge BPF and other fixes after downstream PR.
Conflicts:
kernel/bpf/helpers.c
tools/testing/selftests/bpf/prog_tests/cb_refs.c
tools/testing/selftests/bpf/prog_tests/verifier.c
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
|
|
Add BTF_KIND_LOC_PARAM, BTF_KIND_LOC_PROTO and BTF_KIND_LOCSEC
to help represent location information for functions.
BTF_KIND_LOC_PARAM is used to represent how we retrieve data at a
location; either via register(s), or register+offset, a dereference
of a register+offset or a constant value.
BTF_KIND_LOC_PROTO represents location information about a location
with multiple BTF_KIND_LOC_PARAMs.
And finally BTF_KIND_LOCSEC is a set of location sites, each
of which has
- a BTF_KIND_FUNC function associated with the inline site
- a location prototype specifying where to find the function
parameters
- an address offset relative to the kernel base address
This can be used to support representing
- a fully-inlined function at potentially multiple inline sites
with potentially different parameter availability
- a partially-inlined function where some _LOC_PROTOs represent
inlined sites as above and others have normal _FUNC representations
Also BTF_KIND_LOCSEC struct btf_loc will have two type id
references; one for the associated func, the other for the loc_proto.
Accordingly increase the number of m_offs references in btf_field_desc
to 2.
Signed-off-by: Alan Maguire <alan.maguire@oracle.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260924111428.75957-2-alan.maguire@oracle.com
|
|
The Arm CCA guest TSM provider currently binds through the arm-cca-dev
platform device. Like arm-smccc-trng, this device is not an independent
platform resource; it is a software representation of the RSI firmware
service discovered through SMCCC.
Move RSI discovery into the SMCCC firmware driver. When the SMCCC
conduit is SMC and if RSI ABI version call is supported, create an
arm-rsi SMCCC device. Convert the Arm CCA guest TSM provider to an SMCCC
driver so it binds to that discovered RSI service and keeps module
autoloading through the SMCCC device id table.
Keep the old arm-cca-dev platform-device registration for now. Userspace
has used that device as a Realm-guest indicator, so removing it is left to
a follow-up patch that adds a replacement sysfs ABI.
Reviewed-by: Jason Gunthorpe <jgg@nvidia.com>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Reviewed-by: Catalin Marinas <catalin.marinas@arm.com>
Reviewed-by: Sudeep Holla <sudeep.holla@kernel.org>
Reviewed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Cc: Will Deacon <will@kernel.org>
Cc: Mark Rutland <mark.rutland@arm.com>
Cc: Lorenzo Pieralisi <lpieralisi@kernel.org>
Signed-off-by: Aneesh Kumar K.V (Arm) <aneesh.kumar@kernel.org>
Signed-off-by: Catalin Marinas <catalin.marinas@arm.com>
|
|
The RSI SMCCC function IDs describe a firmware ABI and are not arm64
architecture specific definitions. Follow-up changes need to use them from
non-arch code, including drivers/firmware/smccc and the Arm CCA guest
driver.
Move the complete Realm Service Interface (RSI) implementation from
arch/arm64 to drivers/firmware/arm_rmm. The RSI SMCCC definitions and
command helpers are also moved to include/linux so they can be shared by
architecture code and firmware or driver code. This also keeps the
firmware interface outside architecture code, as requested [1].
[1] https://lore.kernel.org/all/agsNO9cc7H-b0H8L@willie-the-truck
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Reviewed-by: Catalin Marinas <catalin.marinas@arm.com>
Reviewed-by: Jason Gunthorpe <jgg@nvidia.com>
Acked-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Cc: Will Deacon <will@kernel.org>
Cc: Mark Rutland <mark.rutland@arm.com>
Signed-off-by: Aneesh Kumar K.V (Arm) <aneesh.kumar@kernel.org>
Signed-off-by: Catalin Marinas <catalin.marinas@arm.com>
|
|
Pull bpf fixes from Alexei Starovoitov:
- Fix bpf_skb_change_tail() to drop the checksum offload instead of
rejecting the trim of CHECKSUM_PARTIAL skbs (Daniel Borkmann)
- Add KF_PERFMON kfunc flag and require CAP_PERFMON for kfuncs that
read arbitrary memory and for untrusted read-only memory reads
(Daniel Borkmann)
- Clear scalar delta on narrowing stack spill (Daniel Borkmann)
- Set up the frame pointer for the exception callback in arm64 JIT, and
zero-fill other CPUs when BPF_F_CPU update creates a per-cpu hash
element (Donggeun Yoo)
- Various fixes (Emil Tsalapatis):
- Fix bounds check underflow for skb-backed dynptrs
- Fix rx_queue_mapping context access code generation in bpf_sock
- Reject packet pointer arguments to subprogs that may mutate the
packet
- Reject ALU instructions that see arena and non-arena operands on
different code paths
- Fix copied_seq double-counting on sockmap self-redirect
(Geliang Tang)
- Fix divide-by-zero in btf_struct_walk() on a flexible array of
zero-sized elements, fix out-of-bounds read of rtt_min in sock_ops
(Jiayuan Chen)
- Fix bpf_sock_destroy() out-of-bounds read of sk_protocol on TIME_WAIT
and request socks, and sleeping under RCU when destroying a listener
with pending children (Jiayuan Chen)
- Fix JEQ/JNE with immediate operand in MIPS32 JIT and missing zero
extension of BSWAP 16/32 in MIPS64 JIT (Johan Almbladh)
- Avoid soft lockup in htab lookup[_and_delete] batch operations on
large maps (Jose Fernandez)
- Various fixes (Kumar Kartikeya Dwivedi):
- Verify global subprogs in each sleepability context they are
called from
- Make post-verification instruction rewrites killable
- Preserve packet pointer displacement in regsafe()
- Apply CO-RE relocations before subprogram validation, restrict
CO-RE poisoning to relocatable instructions, and reject truncated
ldimm64 CO-RE relocations in libbpf
- Assign lock identity to callback map values
- Compare stack frames in regs_exact()
- Bound ownership depth through local kptrs and graph roots
- Fix u32 overflow in map batch operations when the map size exceeds
4GB (Masoud Aghasi)
- Fix UAF in bpf memalloc due to concurrent consumption of ttrace lists
in alloc_bulk() (Pu Lehui)
- Allow gotox as the terminal instruction of a program or a subprogram
(Siddharth Chintamaneni)
- Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL, skip unsettled links
in link iterator, and reject dev-bound-only programs on other devices
(Weiming Shi)
- Reject non-negative stack offsets in stack_slot_obj_get_spi()
(Xu Yunxiang)
- Check params size before reading reserved fields in
bpf_crypto_ctx_create() (Yuqi Xu)
- Reject max_entries > INT_MAX in sock_map_alloc() (Zhao Gongyi)
- Use a 32-bit compare in xsk_map_gen_lookup() (Zhiling Zou)
- Use kvfree() in xdp_test_run_teardown() (Zhixing Chen)
* tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: (58 commits)
selftests/bpf: Test per-cpu initialization of a BPF_F_CPU created element
bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element
bpf: Fix BSWAP 32 and 16 on MIPS64
bpf: Fix immediate JMP JEQ/JNE on MIPS32
bpf: Reject dev-bound-only programs on other devices
bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc
selftests/bpf: Test for mixed arena/nonarena code paths
bpf: Prevent variable arena/non-arena register contents
selftests/bpf: Test rejection of pkt args to mutating subprogs
bpf: Reject pkt arguments in mutating subprogs
selftests/bpf: Add selftests for rx_queue_mapping context access
bpf: Fix bpf_sock context code generation
selftests/bpf: Test dynptr slices past end of skb
bpf: Fix bounds check for skb-backed dynptrs
selftests/bpf: Reject iterator destruction through fp+0
bpf: Reject non-negative offsets in stack_slot_obj_get_spi()
bpf: Check params size before reading reserved fields
selftests/bpf: Check local object ownership depth
bpf: Bound ownership depth through local kptrs and graph roots
selftests/bpf: Cover frame changes in bounded loops
...
|
|
DA9150 enables IRQ wake after registering its regmap IRQ chip, but does not
disable it on a later probe failure or driver removal. This leaves the wake
depth elevated after the handler is removed, and repeated bind attempts can
accumulate the imbalance.
Remember whether enabling IRQ wake succeeded and balance only a successful
call. Remove MFD children first so their nested IRQ users are gone, then
disable wake before removing the regmap IRQ chip. Preserve the existing
non-fatal behavior when IRQ wake cannot be enabled.
This issue was identified during our ongoing static-analysis research
while reviewing kernel code.
Fixes: b8fce55c09d3 ("mfd: Add support for DA9150 combined charger & fuel-gauge device")
Assisted-by: LLM
Co-developed-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Myeonghun Pak <mhun512@gmail.com>
Link: https://patch.msgid.link/20260923021936.957065-1-mhun512@gmail.com
Signed-off-by: Lee Jones <lee@kernel.org>
|
|
Add a MFD subdevice composition used in Tegra20 based Mot board
(Motorola Atrix 4G and Droid X2).
Signed-off-by: Svyatoslav Ryhel <clamor95@gmail.com>
Link: https://patch.msgid.link/20260721095654.429346-7-clamor95@gmail.com
Signed-off-by: Lee Jones <lee@kernel.org>
|
|
MFD have rigid subdevice structure which does not allow flexible dynamic
subdevice linking. Address this by diverging CPCAP subdevice composition
to take into account board specific configuration.
Create a common and default subdevice composition, rename edit existing
subdevice composition into cpcap_mapphone_devices since it targets mainly
Mapphone board.
Removed st,6556002 as it is no longer applicable to all cases and
duplicates motorola,cpcap, which is used as the default composition.
Signed-off-by: Svyatoslav Ryhel <clamor95@gmail.com>
Link: https://patch.msgid.link/20260721095654.429346-6-clamor95@gmail.com
Signed-off-by: Lee Jones <lee@kernel.org>
|
|
The cs42l44 is a cost optimised variant of cs42l43b. Add basic support
for this new device.
Signed-off-by: Charles Keepax <ckeepax@opensource.cirrus.com>
Link: https://patch.msgid.link/20260901151417.2546618-1-ckeepax@opensource.cirrus.com
Signed-off-by: Lee Jones <lee@kernel.org>
|
|
'ib-mfd-backlight-iio-leds-7.4' into ibs-for-mfd-merged
|
|
Some hardware can autonomously activate/deactivate hardware control.
After that, the LED hardware notifies the LED driver. Currently, there
is no mechanism for LED drivers to notify the LED core about such events
and initiate a trigger transition to reflect the hardware state.
Add a new interface called led_trigger_notify_hw_control_changed(), so
that LED drivers can call it to notify the LED core about the
transition.
The interface only allows two transitions:
1. "none" => private trigger
2. private trigger => "none"
If the current trigger is neither the private trigger nor "none", no
transition will be made. This protects the currently selected software
trigger.
Note that LED_OFF won't be emitted during the #2 transition, as some
hardware may have selected a new brightness level during its hardware
state transition (e.g., laptop keyboards with a shortcut cycling through
different backlight brightnesses and auto mode).
The interface is designed as a void function as any failure should be
non-fatal and the result of transition should not have any impact on the
LED drivers' event handling procedures.
To use the interface, the config LEDS_TRIGGERS_HW_CHANGED must be
enabled, and the LED driver must set the LED_TRIG_HW_CHANGED flag for
the classdev.
By default, the config is enabled when LEDS_BRIGHTNESS_HW_CHANGED is
enabled.
Acked-by: Ike Panhc <ikepanhc@gmail.com>
Signed-off-by: Rong Zhang <i@rong.moe>
Link: https://patch.msgid.link/20260921-leds-trigger-hw-changed-v7-10-fe3cdb6dec51@rong.moe
Signed-off-by: Lee Jones <lee@kernel.org>
|
|
trigger_may_offload_to_hw attribute
There are multiple triggers implementing hardware control. However, the
LED trigger core doesn't really know the hardware control (offloaded)
state since the coordination is done directly between the trigger and
the LED driver. It can only assume private triggers as offloaded and
generic ones as not offloaded.
Add an hw_offloaded() callback so that triggers can report their
offloaded states to the LED trigger core. When unimplemented, it
defaults to true for private triggers and false for generic ones to keep
the current behavior unchanged.
With that, provide a new attribute "trigger_may_offload_to_hw", so that
userspace can determine:
- if the LED device supports hardware control (supported => visible)
- which trigger is the hardware control trigger selected by the LED
device and offloaded to hardware ("[foo_trigger]")
Note: the documentation describes the attribute as "returning a list"
despite the LED core currently only supports one hardware control
trigger per LED device. This is intentional to make the attribute
extensible in the future without breaking userspace. IOW, userspace
should parse the new attribute in the same way as
/sys/class/leds/<led>/trigger.
Acked-by: Ike Panhc <ikepanhc@gmail.com>
Signed-off-by: Rong Zhang <i@rong.moe>
Link: https://patch.msgid.link/20260921-leds-trigger-hw-changed-v7-5-fe3cdb6dec51@rong.moe
Signed-off-by: Lee Jones <lee@kernel.org>
|
|
Since commit b819dc7d8fb2 ("leds: core: Report ENODATA for brightness of
hardware controlled LED"), the brightness attribute becomes write-only
when the LED is controlled fully by the hardware. A write-only attribute
is very confusing.
Moreover, most LED drivers set hardware brightness innocently with the
side effect of disabling hardware control, but the hardware control
trigger remains active, resulting in the software and hardware being out
of sync.
Fix it by removing the hardware control trigger when writing the
brightness attribute.
This should also match the semantics of hardware control:
When the LED is in hw control, no software blink is possible and
doing so will effectively disable hw control.
Fixes: b819dc7d8fb2 ("leds: core: Report ENODATA for brightness of hardware controlled LED")
Acked-by: Ike Panhc <ikepanhc@gmail.com>
Signed-off-by: Rong Zhang <i@rong.moe>
Link: https://patch.msgid.link/20260921-leds-trigger-hw-changed-v7-3-fe3cdb6dec51@rong.moe
Signed-off-by: Lee Jones <lee@kernel.org>
|
|
Currently led_trigger_is_hw_controlled() is placed at led-class.c, which
is not an right place as it falls into the triggers namespace and does
triggers stuff.
Move it into led-triggers.c, and split it into locked and unlocked
variant for convenience.
Fixes: b819dc7d8fb2 ("leds: core: Report ENODATA for brightness of hardware controlled LED")
Acked-by: Ike Panhc <ikepanhc@gmail.com>
Signed-off-by: Rong Zhang <i@rong.moe>
Link: https://patch.msgid.link/20260921-leds-trigger-hw-changed-v7-2-fe3cdb6dec51@rong.moe
Signed-off-by: Lee Jones <lee@kernel.org>
|
|
The CONFIG_PINCTRL=n stubs of pinctrl_gpio_get_config() and
pinctrl_gpio_set_config() return -ENOTSUPP, but the header only pulls in
linux/err.h, which provides asm/errno.h and not the kernel-internal
ENOTSUPP from linux/errno.h. A translation unit that includes
pinctrl/consumer.h before anything else fails to build:
include/linux/pinctrl/consumer.h:110:10: error: use of undeclared identifier 'ENOTSUPP'
Include linux/errno.h explicitly.
Fixes: b44aec87658e ("pinctrl: make the CONFIG_PINCTRL=n gpio config stubs return -ENOTSUPP")
Reported-by: kernel test robot <lkp@intel.com>
Closes: https://lore.kernel.org/oe-kbuild-all/202609240855.U59akHDO-lkp@intel.com/
Closes: https://lore.kernel.org/oe-kbuild-all/202609241240.61Hz5rix-lkp@intel.com/
Signed-off-by: Mehmet Fide <mehmet.fide@screeningeagle.com>
Reviewed-by: Bartosz Golaszewski <bartosz.golaszewski@oss.qualcomm.com>
Signed-off-by: Linus Walleij <linusw@kernel.org>
|
|
SMCCC-discovered firmware services are currently represented by separate
platform devices, such as smccc_trng and arm-cca-dev. Those devices do not
represent independent DT/ACPI-described platform resources; they are
features of the SMCCC firmware interface.
Add an Arm SMCCC bus for services discovered through the SMCCC firmware
interface. The bus provides SMCCC device and driver registration
helpers, function ID based matching, modalias generation, and a sysfs
modalias attribute so SMCCC service drivers can bind to discovered
firmware services and autoload as modules.
Follow-up changes can then register SMCCC firmware services as arm-smccc
devices instead of creating independent per-feature platform devices.
Based on arm_ffa code
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Reviewed-by: Catalin Marinas <catalin.marinas@arm.com>
Reviewed-by: Jason Gunthorpe <jgg@nvidia.com>
Reviewed-by: Sudeep Holla <sudeep.holla@kernel.org>
Cc: Mark Rutland <mark.rutland@arm.com>
Cc: Lorenzo Pieralisi <lpieralisi@kernel.org>
Cc: Sudeep Holla <sudeep.holla@kernel.org>
Cc: Nathan Chancellor <nathan@kernel.org>
Cc: Nicolas Schier <nsc@kernel.org>
Signed-off-by: Aneesh Kumar K.V (Arm) <aneesh.kumar@kernel.org>
Signed-off-by: Catalin Marinas <catalin.marinas@arm.com>
|
|
The is_extended / prog_array_member_cnt protocol introduced by commit
d6083f040d5d ("bpf: Prevent tailcall infinite loop caused by freplace")
keeps a prog extended by a freplace program out of prog_array maps, and
vice versa: once a tail call re-enters an extended subprogram, its
tail_call_cnt resets on every execution and the loop never terminates.
But is_extended is a plain boolean, while one target prog can carry
several freplace links at the same time, one on its entry and one on a
global subprogram. __bpf_trampoline_unlink_prog() cleared is_extended
whenever *any* freplace link detached, so detaching one of two links
re-armed the unbounded loop through the remaining one.
Replace the is_extended boolean with a count of the freplace links
attached to each target prog, so the target stays extended until its
last link detaches. Also rename bpf_freplace_check_tgt_prog() to
bpf_freplace_link_tgt_prog(), as the helper has never been a pure
check: it reserves the target prog on success.
Fixes: d6083f040d5d ("bpf: Prevent tailcall infinite loop caused by freplace")
Signed-off-by: Yuan Chen <chenyuan@kylinos.cn>
Acked-by: Jiri Olsa <jolsa@kernel.org>
Link: https://lore.kernel.org/bpf/20260924023737.1140521-2-chenyuan_fl@163.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
|