summaryrefslogtreecommitdiff
path: root/include/linux
AgeCommit message (Collapse)Author
8 daysVFS: reserve a d_flags bit for fs-specific usageNeilBrown
DCACHE_PRIVATE may be used by any filesystem for its own purposes, much like d_fsdata and d_time. I plan to use this in a similar manner the way nfs stores NFS_FSDATA_BLOCKED in d_fsdata. Signed-off-by: NeilBrown <neil@brown.name> Link: https://patch.msgid.link/20260904215142.1060510-8-neilb@ownmail.net Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 daysVFS: add lockdep monitoring of DCACHE_PAR_LOOKUP lock.NeilBrown
DCACHE_PAR_LOOKUP acts like a lock in that threads can block waiting for it to clear. As we plan to make changes to lock order for this lock, teach lockdep to monitor it so as to help detect bugs early. As NFS allocates an in-lookup dentry to unlink a silly-renamed file, and completes the lookup in a different thread, we need interfaces to release and the acquire ownership of the lock. This avoids lockdep complaining that a lock is still held on return to user-space. Signed-off-by: NeilBrown <neil@brown.name> Link: https://patch.msgid.link/20260904215142.1060510-7-neilb@ownmail.net Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 daysVFS: Add LOOKUP_SHARED flag.NeilBrown
Some ->lookup handlers will need to drop and retake the parent lock, so they can safely use d_alloc_parallel(). ->lookup can be called with the parent lock either exclusive or shared. A new flag, LOOKUP_SHARED, tells ->lookup how the parent is locked. This is rather ugly, but will be gone soon after we move d_alloc_parallel() out of the directory lock as ->lookup() will *always* called with a shared lock on the parent. Signed-off-by: NeilBrown <neil@brown.name> Link: https://patch.msgid.link/20260904215142.1060510-6-neilb@ownmail.net Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 daysVFS: add d_duplicate()NeilBrown
Occasionally a single operation can require two sub-operations on the same name, and it is important that a d_alloc_parallel() (once that can be run unlocked) does not create another dentry with the same name between the operations. Two examples: 1/ rename where the target name (a positive dentry) needs to be "silly-renamed" to a temporary name so it will remain available on the server (NFS and AFS). Here the same name needs to be the subject of one rename, and the target of another. 2/ rename where the subject needs to be replaced with a white-out (shmemfs). Here the same name need to be the subject of a rename and the target of a mknod() In both cases the original dentry is renamed to something else, and a replacement is instantiated, possibly as the target of d_move(), possibly by d_instantiate(). Currently d_alloc() is used to create the dentry and the exclusive lock on the parent ensures no other dentry is created. When d_alloc_parallel() is moved out of the parent lock, this will no longer be sufficient. In particular if the original is renamed away before the new is instantiated, there is a window where d_alloc_parallel() could create another name. "silly-rename" does work in this order. shmemfs whiteout doesn't open this hole but is essentially the same pattern and should use the same approach. The new d_duplicate() creates an in-lookup dentry with the same name as the original dentry, which must be hashed. There is no need to check if an in-lookup dentry exists with the same name as d_alloc_parallel() will never try add one while the hashed dentry exists. Once the new in-lookup is created, d_alloc_parallel() will find it and wait for it to complete, then use it. Signed-off-by: NeilBrown <neil@brown.name> Link: https://patch.msgid.link/20260904215142.1060510-5-neilb@ownmail.net Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 daysVFS: introduce d_alloc_trylock()NeilBrown
Several filesystems use the results of readdir to prime the dcache. These filesystems use d_alloc_parallel() which can block if there is a concurrent lookup. Blocking in that case is pointless as the lookup will add info to the dcache and there is no value in the readdir waiting to see if it should add the info too. Also these calls to d_alloc_parallel() are made while the parent directory is locked. A proposed change to locking will lock the parent later, after d_alloc_parallel(). This means it won't be safe to wait in d_alloc_parallel() while holding the directory lock. So this patch introduces d_alloc_trylock() which doesn't block but instead returns ERR_PTR(-EWOULDBLOCK). Filesystems that prime the dcache (smb/client, nfs, fuse, cephfs) can now use that and ignore -EWOULDBLOCK errors as harmless. Unlike d_alloc_parallel(), d_alloc_trylock() calculates the hash and performs a lookup before an allocation, as that is what all callers want. This is done using try_lookup_noperm(), necessitating the inclusion of namei.h in dcache.c. Signed-off-by: NeilBrown <neil@brown.name> Link: https://patch.msgid.link/20260904215142.1060510-4-neilb@ownmail.net Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 dayscachefiles: Don't rely on backing fs storage map for most use casesDavid Howells
Cachefiles currently uses the backing filesystem's idea of what data is held in a backing file and queries this by means of SEEK_DATA and SEEK_HOLE. However, this means it does two seek operations on the backing file for each individual read call it wants to prepare (unless the first returns -ENXIO). Worse, the backing filesystem is at liberty to insert or remove blocks of zeros in order to optimise its layout which may cause false positives and false negatives. The problem is that keeping track of what is dirty is tricky (if storing info in xattrs, which may have limited capacity and must be read and written as one piece) and expensive (in terms of diskspace at least) and is basically duplicating what a filesystem does. However, the most common write case, in which the application does { open(O_TRUNC); write(); write(); ... write(); close(); } where each write follows directly on from the previous and leaves no gaps in the file is reasonably easy to detect and can be noted in the primary xattr as CACHEFILES_CONTENT_ALL, indicating we have everything up to the object size stored. In this specific case, given that it is known that there are no holes in the file, there's no need to call SEEK_DATA/HOLE or use any other mechanism to track the contents. That speeds things up enormously. Even when it is necessary to use SEEK_DATA/HOLE, it may not be necessary to call it for each cache read subrequest generated. Implement this by adding support for the CACHEFILES_CONTENT_ALL content type (which is defined, but currently unused), which requires a slight adjustment in how backing files are managed. Specifically, the driver needs to know how much of the tail block is data and whether storing more data will create a hole. To this end, the way that the size of a backing file is managed is changed. Currently, the backing file is expanded to strictly match the size of the network file, but this can be changed to carry more useful information. This makes two pieces of metadata available: xattr.object_size and the backing file's i_size. Apply the following schema: (a) i_size is always a multiple of the DIO block size. (b) i_size is only updated to the end of the highest write stored. This is used to work out if we are following on without leaving a hole. (c) xattr.object_size is the size of the network filesystem file cached in this backing file. (d) xattr.object_size must point after the start of the last block (unless both are 0). (e) If xattr.object_size is at or after the block at the current end of the backing file (ie. i_size), then we have all the contents of the block (if xattr.content == CACHEFILES_CONTENT_ALL). (f) If xattr.object_size is somewhere in the middle of the last block, then the data following it is invalid and must be ignored. (g) If data is added to the last block, then that block must be fetched, modified and rewritten (it must be a buffered write through the pagecache and not DIO). (h) Writes to cache are rounded out to blocks on both sides and the folios used as sources must contain data for any lower gap and must have been cleared for any upper gap, and so will rewrite any non-data area in the tail block. To implement this, the following changes are made: (1) cookie->object_size is no longer updated when writes are copied into the pagecache, but rather only updated when a write request completes. This prevents object size miscomparison when checking the xattr causing the backing file to be invalidated (opening and marking the backing file and modifying the pagecache run in parallel). (2) The cache's current idea of the amount of data that should be stored in the backing file is kept track of in object->object_size. Possibly this is redundant with cookie->object_size, but the latter gets updated in some addition circumstances. (3) The size of the backing file at the start of a request is now tracked in struct netfs_cache_resources so that the partial EOF block can be located and cleaned. (4) The cache block size is now used consistently rather than using CACHEFILES_DIO_BLOCK_SIZE (4096). (5) The backing file size is no longer adjusted when looking up an object. (6) When shortening a file, if the new size is not block aligned, the part beyond the new size is cleared. If the file is truncated to zero, the content_info gets reset to CACHEFILES_CONTENT_NO_DATA. (7) A new struct, fscache_occupancy, is instituted to track the region being read. Netfslib allocates it and fills in the start and end of the region to be read then calls the ->query_occupancy() method to find and fill in the extents. It also indicates whether a recorded extent contains data or just contains a region that's all zeros (FSCACHE_EXTENT_DATA or FSCACHE_EXTENT_ZERO). (8) The ->prepare_read() cache method is changed such that, if given, it just limits the amount that can be read from the cache in one go. It no longer indicates what source of read should be done; that information is now obtained from ->query_occupancy(). (9) A new cache method, ->collect_write(), is added that is called when a contiguous series of writes have completed and a discontiguity or the end of the request has been hit. It it supplied with the start and length of the write made to the backing file and can use this information to update the cache metadata. (10) cachefiles_query_occupancy() is altered to find the next two "extents" of data stored in the backing file by doing SEEK_DATA/HOLE between the bounds set - unless it is known that there are no holes, in which case a whole-file first extent can be set. (11) cachefiles_collect_write() is implemented to take the collated write completion information and use this to update the cache metadata, in particular working out whether there's now a hole in the backing file requiring future use of SEEK_DATA/HOLE instead of just assuming the data is all present. It also uses fallocate(FALLOC_FL_ZERO_RANGE) to clean the part of a partial block that extended beyond the old object size. It might be better to perform a synchronous DIO write for this purpose, but that would mandate an RMW cycle. Ideally, it should be all zeros anyway, but, unfortunately, shared-writable mmap can interfere. (12) cachefiles_begin_operation() is updated to note the current backing file size and the cache DIO size. (13) cachefiles_create_tmpfile() no longer expands the backing file when it creates it. (14) cachefiles_set_object_xattr() is changed to use object->object_size rather than cookie->object_size. (15) cachefiles_check_auxdata() is altered to actually store the content type and to also set object->object_size. The cachefiles_coherency tracepoint is also modified to display xattr.object_size. (16) netfs_read_to_pagecache() is reworked. The cache ->prepare_read() method is replaced with ->query_occupancy() as the arbiter of what region of the file is read from where, and that retrieves up to two occupied extents of the backing file at once. The cache ->prepare_read() method is now repurposed to be the same as the equivalent network filesystem method and allows the cache to limit the size of the read before the iterator is prepared. netfs_single_dispatch_read() is similarly modified. (17) netfs_update_i_size() and afs_update_i_size() no longer call fscache_update_cookie() to update cookie->object_size. (18) Write collection now collates contiguous sequences of writes to the cache and calls the cache ->collect_write() method. Signed-off-by: David Howells <dhowells@redhat.com> Link: https://patch.msgid.link/20260910220242.2165023-5-dhowells@redhat.com Reviewed-by: Paulo Alcantara <pc@manguebit.org> cc: Matthew Wilcox <willy@infradead.org> cc: linux-cifs@vger.kernel.org cc: netfs@lists.linux.dev cc: linux-fsdevel@vger.kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 dayssoc: qcom: Add QMI TMD support for remote thermal mitigationCasey Connolly
Add support for Qualcomm Messaging Interface (QMI) based Thermal Mitigation Device (TMD) cooling devices provided by remote subsystems. On Qualcomm platforms where remote processors expose mitigation controls through the TMD QMI service, client drivers need support to discover the service, register cooling devices for available mitigation endpoints, and forward cooling state updates to remote subsystems. Signed-off-by: Casey Connolly <casey.connolly@linaro.org> Co-developed-by: Daniel Lezcano <daniel.lezcano@oss.qualcomm.com> Signed-off-by: Daniel Lezcano <daniel.lezcano@oss.qualcomm.com> Co-developed-by: Gaurav Kohli <gaurav.kohli@oss.qualcomm.com> Signed-off-by: Gaurav Kohli <gaurav.kohli@oss.qualcomm.com> Link: https://lore.kernel.org/r/20260809-b4-qmi-tmd-v8-2-b15d47adc379@oss.qualcomm.com Signed-off-by: Bjorn Andersson <andersson@kernel.org>
8 daysfirmware: xilinx: support RPU boot from DDRTanmay Shah
Cortex-R52 cores can boot from DDR. In such case, elf's boot address will be in the ddr, and it should be passed to the PLM via EEMI call. Add ddr boot config EEMI call before starting RPU to configure the boot address. Also make sure TCM-Boot is support for old platforms as well. Remove HIVEC & LOVEC enum, as now PLM supports boot address directly. Signed-off-by: Tanmay Shah <tanmay.shah@amd.com> Link: https://patch.msgid.link/20260803200502.209361-1-tanmay.shah@amd.com Signed-off-by: Michal Simek <michal.simek@amd.com>
8 daysfirmware: xilinx: ufs: move PHY/SRAM ready polling into the firmware backendMichal Simek
The Versal Gen 2 UFS driver polls the firmware for M-PHY TX/RX configuration readiness and SRAM initialisation completion with two open-coded do/while loops. Each iteration is a full firmware round-trip (PM_IOCTL/IOCTL_READ_REG of a protected PMC_IOU_SLCR register), so the loop can issue up to a million EEMI calls, and it hard-codes the wait policy inside the controller driver. Introduce coarse blocking helpers, zynqmp_pm_wait_mphy_tx_rx_config_ready() and zynqmp_pm_wait_sram_init_done(), that take a caller-supplied timeout budget and contain the poll loop. The loop is EEMI-specific (legacy firmware only exposes the per-read status primitive) so it lives in the firmware driver, keeping the UFS driver backend-agnostic: a future backend can offload the wait to the platform in a single call without touching the controller driver again. The existing per-read primitives stay exported, so the current EEMI interface is unchanged. The timeout budget remains owned by the UFS driver (the consumer that knows the hardware) and is passed down, so EEMI and any future backend stay consistent. Reviewed-by: Sai Krishna Potthuri <sai.krishna.potthuri@amd.com> Link: https://patch.msgid.link/eaeaaea8ed76069943e0706a0917139ff6569929.1785855749.git.michal.simek@amd.com Signed-off-by: Michal Simek <michal.simek@amd.com>
8 daysmtd: nand: qpic_common: drop stray empty line from struct 'bam_transaction'Gabor Juhos
Drop a superfluous empty line from the end of the 'bam_positions' structure group within the 'bam_transaction' structure. No functional changes. Signed-off-by: Gabor Juhos <j4g8y7@gmail.com> Signed-off-by: Miquel Raynal <miquel.raynal@bootlin.com>
8 dayssoc: qcom: llcc: Add LLCC configuration for NordAnurag Pateriya
Add the sub-cache table for Nord, derived from the Nord AU (Safe IVI) SCT table and limited to the POR use cases that have a unique SCID and a usecase ID in llcc-qcom.h. Nord has the CFG0/CFG1 pairs and the full ALGO_ALLOC0..3 set, so it uses the v6 register offsets. num_banks is set explicitly because LLCC_LB_CNT in LLCC_COMMON_STATUS0 is four bits wide and cannot report sixteen banks; a wrong count makes the driver map a bank aperture in place of the broadcast region and abort on the first broadcast read. no_edac is set because the bootloader leaves the LLCC EDAC registers locked, as on Glymur. Signed-off-by: Anurag Pateriya <anurag.pateriya@oss.qualcomm.com> Tested-by: Zhangfei Gao <zhangfei.gao@oss.qualcomm.com> Reviewed-by: Shawn Guo <shengchao.guo@oss.qualcomm.com> Reviewed-by: Abel Vesa <abel.vesa@oss.qualcomm.com> Reviewed-by: Komal Bajaj <komal.bajaj@oss.qualcomm.com> Reviewed-by: Konrad Dybcio <konrad.dybcio@oss.qualcomm.com> Link: https://patch.msgid.link/20260910-nord_llcc-v1-2-30fd1cca65fe@oss.qualcomm.com Signed-off-by: Bjorn Andersson <bjorn.andersson@oss.qualcomm.com>
8 daystty: add break_wait kernel-docJohan Hovold
Add the missing kernel-doc description for the new break_wait field in struct tty_struct. Fixes: 845a2737e0f5 ("tty: abort break signalling on hangup") Reported-by: Randy Dunlap <rdunlap@infradead.org> Link: https://lore.kernel.org/20d75a7a-2e7a-484f-a267-c21db82dd922@infradead.org Signed-off-by: Johan Hovold <johan@kernel.org> Link: https://patch.msgid.link/20260925094222.18303-1-johan@kernel.org Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
8 daysiommu/amd: Add PerfOpt IOMMU performance optimization supportMario Limonciello
Add support for the AMD IOMMU Performance Optimization (PerfOpt) feature as defined in the AMD I/O Virtualization Technology (IOMMU) Specification, Section 3.4.9 (MMIO Offset 016Ch). This feature allows privileged integrated I/O devices (GPUs) to bypass the IOMMU when directly accessing system memory. The IOMMU only enforces the IR/IW permission bits without GPA->SPA translations. amd_iommu_enable_perfopt() performs a detach/reattach cycle to rehome devices already on the identity domain with ATS/PRI/PASID/GCR3 disabled (skip_caps path). amd_iommu_disable_perfopt() restores those capabilities. The per-device dev_data->perfopt flag tracks state. PERF_OPT_EN is a single control bit per IOMMU, shared by every device behind that IOMMU, while enablement is requested per device. It is therefore reference counted (amd_iommu->perfopt_refcount): armed on the first requesting device and cleared on the last, so one device's teardown never clears the bit while a peer behind the same IOMMU still needs it. The per-device flag is cleared on every teardown path (blocked_domain_attach, release_device, and amd_iommu_disable_perfopt), dropping the reference with it, so a reused dev_data never carries stale PerfOpt state onto its next bind. On suspend/resume the hardware is reprogrammed from scratch: amd_iommu_perfopt_clear() forces the bit off without touching the reference count, and amd_iommu_perfopt_restore() re-asserts it from the count after early_enable_iommu(), so armed devices keep the optimization across resume without relying on each consumer driver to re-arm. The exported amd_iommu_enable_perfopt()/amd_iommu_disable_perfopt() run only from a consumer driver's bind/unbind path. group->mutex is not exposed to drivers, but a device bound to its native driver cannot have its IOMMU domain changed concurrently by the core, which serializes the detach/attach pair against core-driven attach. PerfOpt is opt-in -- only enabled when explicitly requested by a driver. Tested-by: Derek J. Clark <derekjohn.clark@gmail.com> Co-developed-by: Jatin Kataria <jkataria@netflix.com> Signed-off-by: Jatin Kataria <jkataria@netflix.com> Signed-off-by: Mario Limonciello <mario.limonciello@amd.com> Tested-by: Boqun Feng (Netflix) <boqun@kernel.org> Acked-by: Alex Deucher <alexander.deucher@amd.com> [joerg.roedel@amd.com: Fixed build failure on i386] Signed-off-by: Joerg Roedel <joerg.roedel@amd.com>
9 daysKVM: Allow architectures to disallow pre-faultLorenzo Stoakes (ARM)
All existing architectures which implement KVM pre-fault (x86, s390) have mechanisms for disallowing pre-faulting. Currently these are open coded as part of kvm_arch_vcpu_pre_fault_memory(). Formalise this by moving them into a new kvm_arch_pre_fault_allowed() hook, which every architecture selecting CONFIG_KVM_GENERIC_PRE_FAULT_MEMORY must implement, returning an error code if the operation is disallowed or 0 otherwise. The hook is called early in the generic code, allowing architectures to disallow the operation prior to vCPU load. This is important, as kvm_vcpu_pre_fault_memory() is the only place where generic code can call vcpu_load() on a vCPU that has not yet been initialised. This lays the foundation for a future change which implements pre-faulting for arm64 which needs to disallow the mechanism for uninitialised vCPUs. Suggested-by: Oliver Upton <oupton@kernel.org> Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org> Acked-by: Sean Christopherson <seanjc@google.com> Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev> Tested-by: Fuad Tabba <fuad.tabba@linux.dev> Link: https://patch.msgid.link/20260923-kvm-arm-prefault-v4-1-d4b0b4dfa8c3@kernel.org Signed-off-by: Marc Zyngier <maz@kernel.org>
9 daysMerge branch 'i2c/i2c-2' into i2c/i2c-nextAndi Shyti
9 daysMerge branch 'i2c/i2c-fixes' into i2c/i2c-nextAndi Shyti
9 daysi2c: core: Add i2c_update_timeout() helper for dynamic transfer timeoutsAniket Randive
The transfer timeout for an I2C controller should reflect the actual message length and bus frequency rather than a static 1-second value. A static timeout causes unnecessary delays on error paths for short messages, and may be insufficient for very long transfers. Add i2c_update_timeout() to i2c-core which computes a transfer-specific timeout and stores it directly in the standard adap->timeout field. The formula accounts for 9 bits per byte (8 data + 1 ACK) at the configured bus frequency. The caller supplies a safety multiplier and a minimum floor so that each driver retains full control over its timing policy without those values becoming public API. Storing the result in adap->timeout makes it visible to all consumers of that field, including the arbitration-loss retry loop in __i2c_transfer(). The function is gated by CONFIG_I2C_DYNAMIC_TIMEOUT. When the config is disabled, i2c_update_timeout() compiles to a no-op inline stub so drivers that call it build cleanly and the existing static 1-second default is preserved unchanged. A timeout explicitly configured by userspace via the I2C_TIMEOUT ioctl is stored in a new adap->user_timeout field and always takes precedence over the kernel-computed value. When userspace has not configured a timeout, the computed value is used. The ioctl keeps writing adap->timeout as well, so adapters that never call i2c_update_timeout() continue to honour it exactly as before. As i2c_update_timeout() is an exported helper, guard against a zero bus frequency from a misbehaving caller with WARN_ON_ONCE() and return early, leaving the existing timeout untouched as a safe fallback rather than dividing by zero. Signed-off-by: Aniket Randive <aniket.randive@oss.qualcomm.com> Reviewed-by: Mukesh Kumar Savaliya <mukesh.savaliya@oss.qualcomm.com> Signed-off-by: Andi Shyti <andi.shyti@kernel.org> Link: https://patch.msgid.link/20260911-master-v9-1-77ac458344e2@oss.qualcomm.com
9 daysbpf: Return the id scratch to a fixed array sized for the stack budgetKumar Kartikeya Dwivedi
Commit 9f3d900a69be ("bpf: Grow the verifier id scratch on demand") turned the id map comparing the ids of two states, which also serves as the id stack of release_reference() and as the id set of bpf_clear_singular_ids(), into arrays grown on demand, so that the env would not embed a map sized for 2 KiB frames. That gave the map a way to fail, and with it release_reference() and the iterator destroy path an -ENOMEM to hand up, states_equal() a case where states compare unequal for lack of memory, and bpf_clear_singular_ids() a flag to keep every id when the set could not record them all. The env is allocated once per program load, so the map can afford to be sized for the largest state again: MAX_BPF_STACK_SLOTS in each of MAX_CALL_FRAMES frames, which a state on a private stack can reach since every frame there may use the whole budget. That makes the map 34 KiB instead of 10 KiB, a small price per load for having no growth that can fail. Make the map and the set a fixed union in the env again and drop the growth and its failure handling. Suggested-by: Alexei Starovoitov <ast@kernel.org> Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://lore.kernel.org/bpf/DLNQW1IRYK94.1DP0M5VKZSY32@gmail.com Link: https://patch.msgid.link/20260924193403.3032037-1-memxor@gmail.com
9 daysMerge git://git.kernel.org/pub/scm/linux/kernel/git/netdev/netJakub Kicinski
Cross-merge networking fixes after downstream PR (net-7.3-rc5). Conflicts: drivers/net/mdio/mdio-realtek-rtl9300.c 89a8a1eef2d44 ("net: mdio: realtek-rtl9300: fix RTL931x C22 extended page selection") cc3cb8db1eef9 ("net: mdio: realtek-rtl9300: Add page tracking") https://lore.kernel.org/arUZOqy73bE2pp0w@sirena.org.uk net/8021q/vlan_dev.c cd5dd68267c4 ("vlan: ensure sufficient headroom in vlan_dev_hard_header()") ca6ff8dd70eb ("vlan: annotate data-races in vlan_dev_priv fields") Adjacent changes: net/ipv6/ip6_gre.c dd47bcf279f1 ("ip6_gre: Call ip6erspan_tunnel_unlink_md() in ip6erspan_changelink().") cce829e2aa1d ("ip6_gre: Protect ip6gre_net.tunnels[][] with mutex.") drivers/net/ethernet/meta/fbnic/fbnic_txrx.c b5d9e9d4d0c1 ("eth: fbnic: use the Rx queue napi pointer to find the napi vector") c0aca269ec07 ("eth: fbnic: Make Rx completion coalescing configurable") drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c d68acbf93531 ("net: stmmac: selftests: Support running selftests on DSA conduits") 85ca3292d7a3 ("net: stmmac: Remove ARP offload code") drivers/net/ethernet/wangxun/libwx/wx_hw.c 3173cba11701 ("net: libwx: fix races in Tx timestamp handling") 7042c8c193e5 ("net: libwx: rename wx_pf_flags to wx_flags") Signed-off-by: Jakub Kicinski <kuba@kernel.org>
9 daysACPI: fan: Use __free() to simplify AML error handlingRafael J. Wysocki
Introduce acpi_object_free for freeing union acpi_object objects allocated by AML and use it for simplifying AML error handling in the ACPI fan driver. While at it, update the driver to use consistent error values across all function using the union acpi_object data type. Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com> Reviewed-by: Andy Shevchenko <andriy.shevchenko@linux.intel.com> Reviewed-by: Armin Wolf <W_Armin@gmx.de> [ rjw: Put the new free definition under CONFIG_ACPI ] Link: https://patch.msgid.link/6029013.DvuYhMxLoT@rafael.j.wysocki Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
9 daysbpf: crypto: Use AES-CBC and AES-ECB librariesEric Biggers
BPF crypto was implemented using the lskcipher API, which doesn't seem to be going anywhere. lskcipher supports only ARC4 and block ciphers in CBC and ECB mode, and only with unoptimized implementations. Library APIs also have been found to be a much better approach, for a variety of reasons, including reduced overhead, greater flexibility, and having to be explicit about the crypto algorithms that are supported. We can safely ignore the theoretical ARC4 and non-AES block cipher support in BPF crypto as unused, which leaves AES-CBC and AES-ECB. AES-CBC was stated to be needed for decrypting packets using a homebrew UDP-based protocol (https://lore.kernel.org/r/d1cdfc23-b336-49a9-8833-29f05b5b9fec@linux.dev/). AES-ECB is used by the BPF self-tests, and it was stated to maybe be used in the future for QUIC-LB (https://lore.kernel.org/r/5f9c3aab-5339-463c-a86d-edac297e1e95@linux.dev/). Those reasons don't make much sense either, especially AES-ECB which isn't appropriate in new systems and should be dropped. Regardless, let's assume that both AES-CBC and AES-ECB need to be kept for now. There are library APIs for both of these now, which are much easier to use and more efficient. Reimplement BPF crypto on top of them, greatly simplifying the code. As part of this, the bpf_crypto_type abstraction layer is removed, as it's not useful. This removes the only user of the lskcipher API, so follow-up patches will be able to remove that code as well. Signed-off-by: Eric Biggers <ebiggers@kernel.org> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260924184254.142720-1-ebiggers@kernel.org
9 daysclk: tegra: annotate struct tegra210_clk_emc_provider with __counted_by_ptrBill Wendling
Annotate the "configs" pointer field of "struct tegra210_clk_emc_provider" with the "__counted_by_ptr" attribute, allowing the compiler to perform runtime bounds checking on accesses to "configs" based on the value of "num_configs". The "emc->provider.configs = devm_kcalloc(...)" call uses "emc->num_timings" for the number of elements, which is then assigned to "emc->provider.num_configs" before any accesses to "configs". The "num_configs" field isn't modified after assignment. Cc: codemender-patching+linux@google.com Assisted-by: LLM Signed-off-by: Bill Wendling <morbo@google.com> Reviewed-by: Gustavo A. R. Silva <gustavoars@kernel.org> Reviewed-by: Kees Cook <kees@kernel.org> Signed-off-by: Brian Masney <bmasney@redhat.com>
9 daysMerge tag 'net-7.3-rc5' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net Pull networking fixes from Jakub Kicinski: "Including fixes from Bluetooth, NFC and Netfilter. Every week in this release is record-setting for number of posted patches. It doesn't seem like we're creating any regressions with all these fixes, three 'Fixes' tags here point to 7.2 commits but none are true regression fixes. We're trying to keep the count down, nonetheless. Previous releases - regressions: - net: don't require the hwtstamp NDOs when a PHY provides timestamping - ipv6: fix dst leak for uncached routes - vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE packets Previous releases - always broken: - packet: use ubuf_info completion for TX_RING packets - arp: terminate device name before lookup - ipv6: do not let ipv6_find_hdr() return an offset past the packet end - udp: remove a disconnected socket from the 4-tuple hash table - sctp: discard the rest of the packet on a stale-cookie error - eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged eswitch" [ And lots of other random network driver fixes ] * tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (189 commits) tcp: prevent collapsing skbs across boundary in rtx queue vlan: ensure sufficient headroom in vlan_dev_hard_header() net/sched: sch_teql: fix shadowed err in __teql_resolve() bridge: check llc_mac_hdr_init() return value in br_send_bpdu() llc: fix skb UAF and leaks on llc_mac_hdr_init() failure llc: reserve device headroom for allocated frames gve: DQO: reject TSO packets with an out of range MSS gve: fix TX drop when GSO MSS is too small for hw gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO net: flush skb_defer_nodes in dev_cpu_dead() net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails af_packet: fix integer overflow in prb_calc_retire_blk_tmo() tipc: Fix a data race on mon->peer_cnt in mon_timeout() net: phy: intel-xway: workaround 100BASE-TX Link-Up issue net/smc: fix UAF on lgr list traversal in smcr_port_err() net/rds: size a connection's path set by the transport it ends up with nfp: hold IPsec RX state under the XArray lock net: ena: fix MMIO read buffer leak on probe failure net: ena: fix PHC cleanup on probe failure net/sched: act_ct: fix helper UAF due to extensions realloc ...
9 daysbpf: Bound program stack use by a per-program limitKumar Kartikeya Dwivedi
The verifier checks every stack access and the combined depth of a call chain against MAX_BPF_STACK, which is also the frame size of the interpreter and the frame that JITs without subprogram tail call support set up for tail-call targets. A JIT that lays out frames of any size and lets a tail-called program set up its own frame does not need that limit; it only needs the verifier to bound how much stack a program uses in total. Add bpf_jit_supports_large_stack() for a JIT to claim that, and give each program its budget through bpf_prog_stack_limit(): MAX_BPF_STACK_JIT when the JIT is requested, the program is not offloaded and the JIT supports large stacks as well as tail calls from subprograms, MAX_BPF_STACK otherwise. The latter is what lets a tail-called program set up its own frame: without it, do_misc_fixups() gives every program with tail calls a MAX_BPF_STACK frame, which a deeper frame verified against the larger budget would overrun. The verifier keeps the budget in env->stack_limit and uses it for the bounds of fixed and variable offset stack accesses, for unprivileged stack pointer arithmetic and its speculation limit, and for the combined and private stack depth checks. A frame may use any part of its program's budget. The interpreter paths keep MAX_BPF_STACK: a program whose main frame is deeper falls back to the JIT-required path of bpf_prog_select_runtime() and one with deeper subprogram frames is rejected when patching calls for the interpreter. The extra stack that may_goto and the timed may_goto instrumentation add below a frame is, as before, not counted against the budget of a JITed program and rejected past MAX_BPF_STACK for an interpreted one. Stack liveness treats a read through a pointer of unknown offset, or a call passing a frame pointer to a subprogram, as reaching the whole frame, and widens the masks of that frame to the deepest half-slot such a read can cover. Bound that by the program's budget too: no access past it is accepted, so a program kept at MAX_BPF_STACK carries masks of two words for such frames, as before, instead of the eight that MAX_BPF_STACK_JIT needs. The three selftests matching a whole-frame read in the liveness log accept either depth. bpf_clone_redirect() transmits from inside the program, and a tc egress or lwt_xmit program that redirects to its own device runs again on top of its own frame until the datapath's recursion limit drops the packet, ten frames deep. Ten MAX_BPF_STACK frames fit the kernel stack as they always did; ten MAX_BPF_STACK_JIT frames would not, so a program that calls bpf_clone_redirect() keeps MAX_BPF_STACK. The redirect helpers that transmit after the program has returned leave no frame behind and do not affect the budget. Nesting through other attach points is not accounted, as before. The spill tracker of the liveness analysis follows no slot past the budget either. The capability is a boolean and the budget a single constant, in the style of the other bpf_jit_supports_*() queries, rather than a per JIT size: the budget is meant to be the same everywhere it is raised, so that programs verify identically across those architectures. No JIT declares support yet, so every program keeps its 512-byte budget. Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260924165740.2146806-14-memxor@gmail.com
9 daysbpf: Size the per-frame verifier structures for a 2 KiB stackKumar Kartikeya Dwivedi
The verifier keeps a few structures whose size follows the deepest frame a program may have: the backtracking and scratched-slot bitmaps, the jump history slot index and the clamp of the liveness masks. They are all expressed through MAX_BPF_STACK_SLOTS, which derives from MAX_BPF_STACK, the frame size of the interpreter. Introduce MAX_BPF_STACK_JIT, the stack budget a program may get on a JIT that can lay out frames of any size, and derive those structures from it so that a frame may be as deep as that budget. Nothing grants the budget yet, so no program verifies differently; the only visible change is that the liveness log prints a whole-frame read up to the new depth, so the three selftests matching such reads are updated. The spill tracker of the liveness analysis keeps a table entry per instruction and tracked slot, so it follows more than the 64 slots of a MAX_BPF_STACK frame only while that table stays within what 64 slots need for the largest program; a subprog of a million instructions keeps 64, one of a quarter million may track all 256. This bounds the table at its old worst case of 640 MiB instead of letting a single deep store push it past what kvmalloc() serves. The backtracking and scratched-slot bitmaps grow from one to four words per frame, a fixed few hundred bytes per verifier environment. tmp_str_buf, which formats a frame's slot list for the log, grows from 320 to 1408 bytes so that all 256 slots still fit, and the log's line buffer from 1 to 2 KiB so that a line built from it is not cut; the environment stays within its 64 KiB allocation. The liveness masks are only as wide as the stack a frame uses, so most frames cost the same as before; a frame that is read as a whole, through a pointer of unknown offset or by bpf_loop() with two callbacks, now carries masks of eight words, 192 bytes per instruction per frame instead of 48. Measured over the 5075 selftest programs, that is 0.2% of the total peak verifier memory: strobemeta_bpf_loop and pyperf600_bpf_loop grow by 11% (1.2 MiB and 0.6 MiB), a few dozen small programs by 40 to 100 KiB each, everything else is unchanged. The next patch bounds such reads by the program's budget, so this cost is only paid once a JIT grants it. Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260924165740.2146806-13-memxor@gmail.com
9 daysbpf: Grow the verifier id scratch on demandKumar Kartikeya Dwivedi
The id map used to compare the ids of two states, which also serves as the id stack of release_reference() and as the id set of bpf_clear_singular_ids(), is a fixed array embedded in struct bpf_verifier_env and sized for the most registers and stack slots a state can possibly hold: 1312 entries, 10 KiB, for 512-byte frames, and four times that once frames may reach 2 KiB, which pushes the env allocation from 64 KiB to 128 KiB for every program verified. States compare a few dozen ids in practice. Turn the map and the set into arrays grown on demand, starting at 64 entries and doubling, and free them with the env. A map that cannot grow treats the states as different, an id set that cannot grow keeps every id of the state, since its counts are then incomplete, and the id stack reports -ENOMEM, which release_reference() hands to its callers. The iterator destroy path used to warn on any failure of release_reference(), which could only come from a bug while the id stack could not run out of room; it now returns -ENOMEM and keeps warning about anything else. An allocation failure thus fails the load and is never unsafe. This removes the last structure whose size scaled with the stack bound and shrinks the env by 10 KiB. Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260924165740.2146806-9-memxor@gmail.com
9 daysbpf: Size liveness stack masks by the stack each frame usesKumar Kartikeya Dwivedi
Stack liveness tracks three masks of 4-byte stack slots - may_read, must_write and live_before - for every instruction of every frame of every function instance. Each mask used to be a fixed 128-bit spis_t, wide enough for the deepest frame MAX_BPF_STACK allows, so an instruction paid 48 bytes per frame no matter how little stack the frame actually touched. These per instruction arrays are the largest liveness allocation: a 40k instruction program spends ~2 MiB on each frame array. Sizing them for a larger stack budget would multiply that cost for every program, while most frames use a fraction of the budget. Replace the fixed masks with a variable width bitmap array per (instance, frame). struct frame_masks holds the width shared by all of its masks in @words plus a flat unsigned long array, where mask @kind of the instruction at relative index @i lives at &bits[(i * FM_MASK_CNT + kind) * words]. An array starts at the width the first recorded half-slot needs (minimum one word, that is 256 bytes of stack with 64-bit words) and is reallocated into a wider stride once a deeper half-slot shows up. Only the marking functions widen an array, and they all run during the static analysis before do_check(), so no verification time path can move one. The marking API now takes an inclusive half-slot range, which is what record_stack_access_off() computed anyway, and clamps the upper end to the deepest half-slot a frame can have. An access below the stack bound is rejected by the main verifier pass later; liveness only has to stay inside the masks. A range that covers no whole half-slot, as a byte store at fp-1 produces, marks nothing. A "read everything" mark, used for calls that are neither helpers nor kfuncs and for pointers whose offset or frame identity was lost, widens the array to the maximum width and sets every bit: recorded at a narrower width it would lose the half-slots a later widening adds, and no write can cancel it as no write covers the whole frame. A half-slot past the width of the masks was never read by the frame and is never live. merge_instances() may see a different width on each side, so it widens dst to cover src and counts a word only dst has as zero on the src side, which unions into may_read as a no-op and intersects must_write to empty, never claiming a write that did not happen. The rest is mechanical: update_insn() propagates live_before word by word over the array's width (all instructions of one array share it, hence so do an instruction's successors), is_live_before() answers false past the width, and the use:/def: log printing iterates up to the width. Liveness results are unchanged. A half-slot only ever gets a bit from a recorded access and recording one widens the array to cover it, so every bit the fixed masks could hold still fits. Frames whose deepest recorded half-slot fits in a single word now cost 24 bytes per instruction per frame instead of 48. The spill tracker that feeds the analysis had the same limit built in: it followed 64 slots, fp-8 to fp-512, per subprog, and a fill from a deeper slot returned an imprecise pointer, so every later access through it counted as a read of the whole stack of every frame. Let it follow the deepest 8-byte access a subprog makes directly through R10 when that reaches past those 64 slots, which is how spills and fills are compiled, and keep the count with the per-callsite snapshot a callee uses to fill from its caller's frame. Programs within 512 bytes are tracked exactly as before; a program spilling deeper stays precise. Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260924165740.2146806-8-memxor@gmail.com
9 daysbpf: Track scratched stack slots with a bitmapKumar Kartikeya Dwivedi
The set of stack slots touched since the verifier state was last printed lives in a u64, so the log could only ever mark the first 64 slots of a frame as scratched. Make it a bitmap sized by MAX_BPF_STACK_SLOTS and go through the bitmap helpers for setting, testing, clearing and filling it. No functional change. Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260924165740.2146806-6-memxor@gmail.com
9 daysbpf: Track backtracking stack slots with bitmapsKumar Kartikeya Dwivedi
Precision backtracking keeps the stack slots that still need a precise mark in one u64 per frame, which ties it to frames of at most 64 slots. Turn the per-frame masks into bitmaps sized by MAX_BPF_STACK_SLOTS and use the bitmap helpers for setting, clearing, testing and iterating them. The formatting helper takes a bitmap and the leftover-slot bug reports print the formatted slot list instead of a hex mask. mark_reg_stack_read() collected zero spills in a u64 of its own before handing it to the backtracker; it now counts them and revisits the range to mark each slot, which drops the only remaining mask-typed entry point. No functional change. Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260924165740.2146806-5-memxor@gmail.com
9 daysbpf: Store linked registers in the jump history as an arrayKumar Kartikeya Dwivedi
Linked scalar registers are recorded in the jump history packed into a u64 as five 11-bit entries, each naming a frame and a register or stack slot. The 6-bit slot field covers exactly the 64 slots of a 512-byte frame, so a spilled scalar in a deeper slot could not be linked and larger frames were ruled out by construction. Store the linked registers as an array of five u16 entries plus a count instead, each entry holding the frame number, a register-or-slot bit and an 11-bit register or slot index, which covers frames of up to 16 KiB. Callers that record no linked registers pass NULL. The history entry grows from 16 to 20 bytes, and the history is the one verifier structure whose size follows the number of instructions a loop iterates over rather than the state count. Measured over the 5075 selftest programs, peak verifier memory is unchanged for all but the loop-heavy ones, which grow by 8 to 16%: loop1/nested_loops from 17.6 to 19.2 MiB, verifier_loops1/jumps_out_rather_than_in from 4.4 to 5.1 MiB, strobemeta by 0.3% and pyperf600_nounroll by 0.8%. Verdicts, processed instructions and state counts stay the same everywhere. No functional change. Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260924165740.2146806-4-memxor@gmail.com
9 daysbpf: Widen the stack slot index in the jump historyKumar Kartikeya Dwivedi
Jump history entries record the stack slot touched by a spill or fill in a 6-bit field, which only fits the 64 slots of a 512-byte frame and is pinned to that size by a static_assert on MAX_BPF_STACK. Move the flags into the first word and give the slot index 12 bits of the second word instead, so the entry stays 16 bytes while frames of up to 32 KiB can be recorded. Introduce MAX_BPF_STACK_SLOTS for the number of 8-byte slots a frame can have and use it for the static_assert and for BPF_ID_MAP_SIZE, so the verifier expresses per-frame capacity through one constant. No functional change. Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260924165740.2146806-3-memxor@gmail.com
9 daysbpf: Add accessors for verifier stack slotsKumar Kartikeya Dwivedi
The verifier indexes a frame's stack state directly through state->stack[spi] and computes the number of tracked slots as allocated_stack / BPF_REG_SIZE in every file that touches stack slots, and so does the nfp offload driver. Route all of these through two helpers, bpf_stack_slot() and bpf_stack_nr_slots(), so the layout of the per-frame stack state is visible in one place. The one lookup that goes the other way, reg_to_target() in the diagnostics code, maps a register pointer back to its slot by address and now says that it relies on the slots forming one contiguous array. Both helpers take a const frame: the slot accessor returns the slot through the frame's stack pointer, so read-only code such as the state printer can use it without giving up its qualifiers. Functions that look up the same slot repeatedly now fetch it once. No functional change. Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260924165740.2146806-2-memxor@gmail.com
9 dayssched_ext: Place dsq_vtime next to dsq_priqUsama Arif
scx_dispatch_enqueue() calls rb_add() for a vtime-ordered DSQ while holding dsq->lock. At each level of the descent, scx_dsq_priq_less() loads the visited task's dsq_vtime, and that comparison selects rb_left or rb_right from the same task's dsq_priq. In the x86-64 benchmark configuration, these accesses fell on different 64-byte cache lines before this change, so a cold visited task could require a second cache line fill on the dependent descent path. Move dsq_vtime immediately before dsq_priq, keeping dsq_seq and dsq_flags next to dsq_list. In the two x86-64 layouts inspected, where scx starts at offsets 776 and 840 in task_struct, the fields from dsq_list through dsq_priq occupy one cache line. The exact cache line placement depends on the containing task_struct layout and is not guaranteed for every configuration or architecture. Where they share a line, avoiding the dependent cache line fill helps reduce dsq->lock hold time during vtime insertion and therefore contention on the lock. In a 16-vCPU KVM guest, the instrumented enqueue interval was 4-6% shorter with scx_lavd and scx_layered. scx_mitosis and end-to-end workload time showed no consistent change. This only reorders fields; no scheduling behavior change is intended. Signed-off-by: Usama Arif <usama.arif@linux.dev> Signed-off-by: Tejun Heo <tj@kernel.org>
9 daysMerge git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf 7.3-rc4Alexei Starovoitov
Cross-merge BPF and other fixes after downstream PR. Conflicts: kernel/bpf/helpers.c tools/testing/selftests/bpf/prog_tests/cb_refs.c tools/testing/selftests/bpf/prog_tests/verifier.c Signed-off-by: Alexei Starovoitov <ast@kernel.org>
9 daysbtf: Extend UAPI to support BTF location (inline site) infoAlan Maguire
Add BTF_KIND_LOC_PARAM, BTF_KIND_LOC_PROTO and BTF_KIND_LOCSEC to help represent location information for functions. BTF_KIND_LOC_PARAM is used to represent how we retrieve data at a location; either via register(s), or register+offset, a dereference of a register+offset or a constant value. BTF_KIND_LOC_PROTO represents location information about a location with multiple BTF_KIND_LOC_PARAMs. And finally BTF_KIND_LOCSEC is a set of location sites, each of which has - a BTF_KIND_FUNC function associated with the inline site - a location prototype specifying where to find the function parameters - an address offset relative to the kernel base address This can be used to support representing - a fully-inlined function at potentially multiple inline sites with potentially different parameter availability - a partially-inlined function where some _LOC_PROTOs represent inlined sites as above and others have normal _FUNC representations Also BTF_KIND_LOCSEC struct btf_loc will have two type id references; one for the associated func, the other for the loc_proto. Accordingly increase the number of m_offs references in btf_field_desc to 2. Signed-off-by: Alan Maguire <alan.maguire@oracle.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Acked-by: Eduard Zingerman <eddyz87@gmail.com> Link: https://patch.msgid.link/20260924111428.75957-2-alan.maguire@oracle.com
9 daysfirmware: smccc: arm-cca-guest: Bind the TSM provider to an SMCCC deviceAneesh Kumar K.V (Arm)
The Arm CCA guest TSM provider currently binds through the arm-cca-dev platform device. Like arm-smccc-trng, this device is not an independent platform resource; it is a software representation of the RSI firmware service discovered through SMCCC. Move RSI discovery into the SMCCC firmware driver. When the SMCCC conduit is SMC and if RSI ABI version call is supported, create an arm-rsi SMCCC device. Convert the Arm CCA guest TSM provider to an SMCCC driver so it binds to that discovered RSI service and keeps module autoloading through the SMCCC device id table. Keep the old arm-cca-dev platform-device registration for now. Userspace has used that device as a Realm-guest indicator, so removing it is left to a follow-up patch that adds a replacement sysfs ABI. Reviewed-by: Jason Gunthorpe <jgg@nvidia.com> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com> Reviewed-by: Catalin Marinas <catalin.marinas@arm.com> Reviewed-by: Sudeep Holla <sudeep.holla@kernel.org> Reviewed-by: Suzuki K Poulose <suzuki.poulose@arm.com> Cc: Will Deacon <will@kernel.org> Cc: Mark Rutland <mark.rutland@arm.com> Cc: Lorenzo Pieralisi <lpieralisi@kernel.org> Signed-off-by: Aneesh Kumar K.V (Arm) <aneesh.kumar@kernel.org> Signed-off-by: Catalin Marinas <catalin.marinas@arm.com>
9 daysfirmware: arm_rmm: Move RSI support out of arch/arm64Aneesh Kumar K.V (Arm)
The RSI SMCCC function IDs describe a firmware ABI and are not arm64 architecture specific definitions. Follow-up changes need to use them from non-arch code, including drivers/firmware/smccc and the Arm CCA guest driver. Move the complete Realm Service Interface (RSI) implementation from arch/arm64 to drivers/firmware/arm_rmm. The RSI SMCCC definitions and command helpers are also moved to include/linux so they can be shared by architecture code and firmware or driver code. This also keeps the firmware interface outside architecture code, as requested [1]. [1] https://lore.kernel.org/all/agsNO9cc7H-b0H8L@willie-the-truck Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com> Reviewed-by: Catalin Marinas <catalin.marinas@arm.com> Reviewed-by: Jason Gunthorpe <jgg@nvidia.com> Acked-by: Suzuki K Poulose <suzuki.poulose@arm.com> Cc: Will Deacon <will@kernel.org> Cc: Mark Rutland <mark.rutland@arm.com> Signed-off-by: Aneesh Kumar K.V (Arm) <aneesh.kumar@kernel.org> Signed-off-by: Catalin Marinas <catalin.marinas@arm.com>
9 daysMerge tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpfLinus Torvalds
Pull bpf fixes from Alexei Starovoitov: - Fix bpf_skb_change_tail() to drop the checksum offload instead of rejecting the trim of CHECKSUM_PARTIAL skbs (Daniel Borkmann) - Add KF_PERFMON kfunc flag and require CAP_PERFMON for kfuncs that read arbitrary memory and for untrusted read-only memory reads (Daniel Borkmann) - Clear scalar delta on narrowing stack spill (Daniel Borkmann) - Set up the frame pointer for the exception callback in arm64 JIT, and zero-fill other CPUs when BPF_F_CPU update creates a per-cpu hash element (Donggeun Yoo) - Various fixes (Emil Tsalapatis): - Fix bounds check underflow for skb-backed dynptrs - Fix rx_queue_mapping context access code generation in bpf_sock - Reject packet pointer arguments to subprogs that may mutate the packet - Reject ALU instructions that see arena and non-arena operands on different code paths - Fix copied_seq double-counting on sockmap self-redirect (Geliang Tang) - Fix divide-by-zero in btf_struct_walk() on a flexible array of zero-sized elements, fix out-of-bounds read of rtt_min in sock_ops (Jiayuan Chen) - Fix bpf_sock_destroy() out-of-bounds read of sk_protocol on TIME_WAIT and request socks, and sleeping under RCU when destroying a listener with pending children (Jiayuan Chen) - Fix JEQ/JNE with immediate operand in MIPS32 JIT and missing zero extension of BSWAP 16/32 in MIPS64 JIT (Johan Almbladh) - Avoid soft lockup in htab lookup[_and_delete] batch operations on large maps (Jose Fernandez) - Various fixes (Kumar Kartikeya Dwivedi): - Verify global subprogs in each sleepability context they are called from - Make post-verification instruction rewrites killable - Preserve packet pointer displacement in regsafe() - Apply CO-RE relocations before subprogram validation, restrict CO-RE poisoning to relocatable instructions, and reject truncated ldimm64 CO-RE relocations in libbpf - Assign lock identity to callback map values - Compare stack frames in regs_exact() - Bound ownership depth through local kptrs and graph roots - Fix u32 overflow in map batch operations when the map size exceeds 4GB (Masoud Aghasi) - Fix UAF in bpf memalloc due to concurrent consumption of ttrace lists in alloc_bulk() (Pu Lehui) - Allow gotox as the terminal instruction of a program or a subprogram (Siddharth Chintamaneni) - Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL, skip unsettled links in link iterator, and reject dev-bound-only programs on other devices (Weiming Shi) - Reject non-negative stack offsets in stack_slot_obj_get_spi() (Xu Yunxiang) - Check params size before reading reserved fields in bpf_crypto_ctx_create() (Yuqi Xu) - Reject max_entries > INT_MAX in sock_map_alloc() (Zhao Gongyi) - Use a 32-bit compare in xsk_map_gen_lookup() (Zhiling Zou) - Use kvfree() in xdp_test_run_teardown() (Zhixing Chen) * tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: (58 commits) selftests/bpf: Test per-cpu initialization of a BPF_F_CPU created element bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element bpf: Fix BSWAP 32 and 16 on MIPS64 bpf: Fix immediate JMP JEQ/JNE on MIPS32 bpf: Reject dev-bound-only programs on other devices bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc selftests/bpf: Test for mixed arena/nonarena code paths bpf: Prevent variable arena/non-arena register contents selftests/bpf: Test rejection of pkt args to mutating subprogs bpf: Reject pkt arguments in mutating subprogs selftests/bpf: Add selftests for rx_queue_mapping context access bpf: Fix bpf_sock context code generation selftests/bpf: Test dynptr slices past end of skb bpf: Fix bounds check for skb-backed dynptrs selftests/bpf: Reject iterator destruction through fp+0 bpf: Reject non-negative offsets in stack_slot_obj_get_spi() bpf: Check params size before reading reserved fields selftests/bpf: Check local object ownership depth bpf: Bound ownership depth through local kptrs and graph roots selftests/bpf: Cover frame changes in bounded loops ...
9 daysmfd: da9150: Balance IRQ wake on teardownMyeonghun Pak
DA9150 enables IRQ wake after registering its regmap IRQ chip, but does not disable it on a later probe failure or driver removal. This leaves the wake depth elevated after the handler is removed, and repeated bind attempts can accumulate the imbalance. Remember whether enabling IRQ wake succeeded and balance only a successful call. Remove MFD children first so their nested IRQ users are gone, then disable wake before removing the regmap IRQ chip. Preserve the existing non-fatal behavior when IRQ wake cannot be enabled. This issue was identified during our ongoing static-analysis research while reviewing kernel code. Fixes: b8fce55c09d3 ("mfd: Add support for DA9150 combined charger & fuel-gauge device") Assisted-by: LLM Co-developed-by: Ijae Kim <ae878000@gmail.com> Signed-off-by: Ijae Kim <ae878000@gmail.com> Signed-off-by: Myeonghun Pak <mhun512@gmail.com> Link: https://patch.msgid.link/20260923021936.957065-1-mhun512@gmail.com Signed-off-by: Lee Jones <lee@kernel.org>
9 daysmfd: motorola-cpcap: Add support for Mot CPCAP compositionSvyatoslav Ryhel
Add a MFD subdevice composition used in Tegra20 based Mot board (Motorola Atrix 4G and Droid X2). Signed-off-by: Svyatoslav Ryhel <clamor95@gmail.com> Link: https://patch.msgid.link/20260721095654.429346-7-clamor95@gmail.com Signed-off-by: Lee Jones <lee@kernel.org>
9 daysmfd: motorola-cpcap: Diverge configuration per-boardSvyatoslav Ryhel
MFD have rigid subdevice structure which does not allow flexible dynamic subdevice linking. Address this by diverging CPCAP subdevice composition to take into account board specific configuration. Create a common and default subdevice composition, rename edit existing subdevice composition into cpcap_mapphone_devices since it targets mainly Mapphone board. Removed st,6556002 as it is no longer applicable to all cases and duplicates motorola,cpcap, which is used as the default composition. Signed-off-by: Svyatoslav Ryhel <clamor95@gmail.com> Link: https://patch.msgid.link/20260721095654.429346-6-clamor95@gmail.com Signed-off-by: Lee Jones <lee@kernel.org>
9 daysmfd: cs42l43: Add support for new cs42l44 variantCharles Keepax
The cs42l44 is a cost optimised variant of cs42l43b. Add basic support for this new device. Signed-off-by: Charles Keepax <ckeepax@opensource.cirrus.com> Link: https://patch.msgid.link/20260901151417.2546618-1-ckeepax@opensource.cirrus.com Signed-off-by: Lee Jones <lee@kernel.org>
9 daysMerge branches 'ib-mfd-thermal-7.4', 'ib-mfd-clk-gpio-regulator-rtc-7.4' and ↵Lee Jones
'ib-mfd-backlight-iio-leds-7.4' into ibs-for-mfd-merged
9 daysleds: trigger: Add led_trigger_notify_hw_control_changed() interfaceRong Zhang
Some hardware can autonomously activate/deactivate hardware control. After that, the LED hardware notifies the LED driver. Currently, there is no mechanism for LED drivers to notify the LED core about such events and initiate a trigger transition to reflect the hardware state. Add a new interface called led_trigger_notify_hw_control_changed(), so that LED drivers can call it to notify the LED core about the transition. The interface only allows two transitions: 1. "none" => private trigger 2. private trigger => "none" If the current trigger is neither the private trigger nor "none", no transition will be made. This protects the currently selected software trigger. Note that LED_OFF won't be emitted during the #2 transition, as some hardware may have selected a new brightness level during its hardware state transition (e.g., laptop keyboards with a shortcut cycling through different backlight brightnesses and auto mode). The interface is designed as a void function as any failure should be non-fatal and the result of transition should not have any impact on the LED drivers' event handling procedures. To use the interface, the config LEDS_TRIGGERS_HW_CHANGED must be enabled, and the LED driver must set the LED_TRIG_HW_CHANGED flag for the classdev. By default, the config is enabled when LEDS_BRIGHTNESS_HW_CHANGED is enabled. Acked-by: Ike Panhc <ikepanhc@gmail.com> Signed-off-by: Rong Zhang <i@rong.moe> Link: https://patch.msgid.link/20260921-leds-trigger-hw-changed-v7-10-fe3cdb6dec51@rong.moe Signed-off-by: Lee Jones <lee@kernel.org>
9 daysleds: trigger: Add hw_offloaded() callback and provide ↵Rong Zhang
trigger_may_offload_to_hw attribute There are multiple triggers implementing hardware control. However, the LED trigger core doesn't really know the hardware control (offloaded) state since the coordination is done directly between the trigger and the LED driver. It can only assume private triggers as offloaded and generic ones as not offloaded. Add an hw_offloaded() callback so that triggers can report their offloaded states to the LED trigger core. When unimplemented, it defaults to true for private triggers and false for generic ones to keep the current behavior unchanged. With that, provide a new attribute "trigger_may_offload_to_hw", so that userspace can determine: - if the LED device supports hardware control (supported => visible) - which trigger is the hardware control trigger selected by the LED device and offloaded to hardware ("[foo_trigger]") Note: the documentation describes the attribute as "returning a list" despite the LED core currently only supports one hardware control trigger per LED device. This is intentional to make the attribute extensible in the future without breaking userspace. IOW, userspace should parse the new attribute in the same way as /sys/class/leds/<led>/trigger. Acked-by: Ike Panhc <ikepanhc@gmail.com> Signed-off-by: Rong Zhang <i@rong.moe> Link: https://patch.msgid.link/20260921-leds-trigger-hw-changed-v7-5-fe3cdb6dec51@rong.moe Signed-off-by: Lee Jones <lee@kernel.org>
9 daysleds: class: Remove hardware control trigger when writing brightnessRong Zhang
Since commit b819dc7d8fb2 ("leds: core: Report ENODATA for brightness of hardware controlled LED"), the brightness attribute becomes write-only when the LED is controlled fully by the hardware. A write-only attribute is very confusing. Moreover, most LED drivers set hardware brightness innocently with the side effect of disabling hardware control, but the hardware control trigger remains active, resulting in the software and hardware being out of sync. Fix it by removing the hardware control trigger when writing the brightness attribute. This should also match the semantics of hardware control: When the LED is in hw control, no software blink is possible and doing so will effectively disable hw control. Fixes: b819dc7d8fb2 ("leds: core: Report ENODATA for brightness of hardware controlled LED") Acked-by: Ike Panhc <ikepanhc@gmail.com> Signed-off-by: Rong Zhang <i@rong.moe> Link: https://patch.msgid.link/20260921-leds-trigger-hw-changed-v7-3-fe3cdb6dec51@rong.moe Signed-off-by: Lee Jones <lee@kernel.org>
9 daysleds: trigger: Move led_trigger_is_hw_controlled() to the right placeRong Zhang
Currently led_trigger_is_hw_controlled() is placed at led-class.c, which is not an right place as it falls into the triggers namespace and does triggers stuff. Move it into led-triggers.c, and split it into locked and unlocked variant for convenience. Fixes: b819dc7d8fb2 ("leds: core: Report ENODATA for brightness of hardware controlled LED") Acked-by: Ike Panhc <ikepanhc@gmail.com> Signed-off-by: Rong Zhang <i@rong.moe> Link: https://patch.msgid.link/20260921-leds-trigger-hw-changed-v7-2-fe3cdb6dec51@rong.moe Signed-off-by: Lee Jones <lee@kernel.org>
9 dayspinctrl: include linux/errno.h in consumer.h for ENOTSUPPMehmet Fide
The CONFIG_PINCTRL=n stubs of pinctrl_gpio_get_config() and pinctrl_gpio_set_config() return -ENOTSUPP, but the header only pulls in linux/err.h, which provides asm/errno.h and not the kernel-internal ENOTSUPP from linux/errno.h. A translation unit that includes pinctrl/consumer.h before anything else fails to build: include/linux/pinctrl/consumer.h:110:10: error: use of undeclared identifier 'ENOTSUPP' Include linux/errno.h explicitly. Fixes: b44aec87658e ("pinctrl: make the CONFIG_PINCTRL=n gpio config stubs return -ENOTSUPP") Reported-by: kernel test robot <lkp@intel.com> Closes: https://lore.kernel.org/oe-kbuild-all/202609240855.U59akHDO-lkp@intel.com/ Closes: https://lore.kernel.org/oe-kbuild-all/202609241240.61Hz5rix-lkp@intel.com/ Signed-off-by: Mehmet Fide <mehmet.fide@screeningeagle.com> Reviewed-by: Bartosz Golaszewski <bartosz.golaszewski@oss.qualcomm.com> Signed-off-by: Linus Walleij <linusw@kernel.org>
9 daysfirmware: smccc: Add an Arm SMCCC busAneesh Kumar K.V (Arm)
SMCCC-discovered firmware services are currently represented by separate platform devices, such as smccc_trng and arm-cca-dev. Those devices do not represent independent DT/ACPI-described platform resources; they are features of the SMCCC firmware interface. Add an Arm SMCCC bus for services discovered through the SMCCC firmware interface. The bus provides SMCCC device and driver registration helpers, function ID based matching, modalias generation, and a sysfs modalias attribute so SMCCC service drivers can bind to discovered firmware services and autoload as modules. Follow-up changes can then register SMCCC firmware services as arm-smccc devices instead of creating independent per-feature platform devices. Based on arm_ffa code Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com> Reviewed-by: Catalin Marinas <catalin.marinas@arm.com> Reviewed-by: Jason Gunthorpe <jgg@nvidia.com> Reviewed-by: Sudeep Holla <sudeep.holla@kernel.org> Cc: Mark Rutland <mark.rutland@arm.com> Cc: Lorenzo Pieralisi <lpieralisi@kernel.org> Cc: Sudeep Holla <sudeep.holla@kernel.org> Cc: Nathan Chancellor <nathan@kernel.org> Cc: Nicolas Schier <nsc@kernel.org> Signed-off-by: Aneesh Kumar K.V (Arm) <aneesh.kumar@kernel.org> Signed-off-by: Catalin Marinas <catalin.marinas@arm.com>
9 daysbpf: Keep target extended until its last freplace link detachesYuan Chen
The is_extended / prog_array_member_cnt protocol introduced by commit d6083f040d5d ("bpf: Prevent tailcall infinite loop caused by freplace") keeps a prog extended by a freplace program out of prog_array maps, and vice versa: once a tail call re-enters an extended subprogram, its tail_call_cnt resets on every execution and the loop never terminates. But is_extended is a plain boolean, while one target prog can carry several freplace links at the same time, one on its entry and one on a global subprogram. __bpf_trampoline_unlink_prog() cleared is_extended whenever *any* freplace link detached, so detaching one of two links re-armed the unbounded loop through the remaining one. Replace the is_extended boolean with a count of the freplace links attached to each target prog, so the target stays extended until its last link detaches. Also rename bpf_freplace_check_tgt_prog() to bpf_freplace_link_tgt_prog(), as the helper has never been a pure check: it reserves the target prog on success. Fixes: d6083f040d5d ("bpf: Prevent tailcall infinite loop caused by freplace") Signed-off-by: Yuan Chen <chenyuan@kylinos.cn> Acked-by: Jiri Olsa <jolsa@kernel.org> Link: https://lore.kernel.org/bpf/20260924023737.1140521-2-chenyuan_fl@163.com Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>