| Age | Commit message (Collapse) | Author |
|
The data direct capability FW checks are duplicated inline in
mlx5_ib_data_direct_init() and mlx5_ib_data_direct_cleanup().
Wrap it in a mlx5_data_direct_supported() helper and use it in both
places, in preparation for moving the data direct matching code to
mlx5_core. For the same reason put it in driver.h instead of
data_direct.h.
No functional change.
Signed-off-by: Dragos Tatulea <dtatulea@nvidia.com>
Reviewed-by: Leon Romanovsky <leonro@nvidia.com>
Reviewed-by: Cosmin Ratiu <cratiu@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260930114227.139274-3-tariqt@nvidia.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
|
|
linux/module.h appears in roughly 15k #include directives across the
kernel. This makes it a "hot" header, so it should avoid pulling in
unnecessary definitions.
The header currently includes linux/error-injection.h to obtain the
definition of `struct error_injection_entry`. However, this is unnecessary
because the type is only referenced in the file as a pointer, for which an
incomplete type is sufficient.
Remove the linux/error-injection.h include from linux/module.h and add it
to kernel/module/main.c instead, where
`sizeof(struct error_injection_entry)` is actually needed.
Reviewed-by: Aaron Tomlin <atomlin@atomlin.com>
Signed-off-by: Petr Pavlu <petr.pavlu@suse.com>
|
|
include/linux/compat.h and include/linux/syscalls.h use
ALLOW_ERROR_INJECTION(), which is defined in asm-generic/error-injection.h.
They currently rely on that header being included indirectly through other
files, typically via linux/module.h -> linux/error-injection.h.
Add the missing include in preparation for removing the
linux/error-injection.h include from linux/module.h.
Signed-off-by: Petr Pavlu <petr.pavlu@suse.com>
|
|
* kvm-arm64/pre-faulting:
: \
: Enable stage-2 pre-faulting in the canonical IPA space, implementing the
: KVM_PRE_FAULT_MEMORY API. Patches courtesy of Lorenzo Stoakes, and based
: on an initial work by Jack Thomson.
: /
KVM: selftests: Add nested pre-fault test for arm64
KVM: selftests: Add option for different backing in pre-fault tests
KVM: selftests: Enable pre_fault_memory_test for arm64
Documentation: KVM: document arm64 KVM_PRE_FAULT_MEMORY
KVM: arm64: Implement KVM_PRE_FAULT_MEMORY
KVM: arm64: Pass walk flags to kvm_pgtable_get_leaf()
KVM: arm64: Propagate EHWPOISON in kvm_s2_fault_pin_pfn()
KVM: arm64: Size the stage-2 memcache from the fault MMU
KVM: arm64: Propagate and use kvm_s2_fault_result on S2 fault
KVM: arm64: Propagate and use mmu in s2fd when handling guest aborts
KVM: arm64: Propagate and use esr in s2fd when handling guest aborts
KVM: arm64: Use ESR helpers in guest abort handling
arm64: Add ESR fault helpers
KVM: Allow architectures to disallow pre-fault
Signed-off-by: Marc Zyngier <maz@kernel.org>
|
|
* kvm-arm64/gicv5-7.4: (52 commits)
: \
: GICv5 IRS support, courtesy of Sascha Bischoff. From the cover letter:
:
: "This series builds on the initial vGICv5 support and adds support
: for the GICv5 IRS, as described by the GICv5 (EAC0) specification.
: With this, a GICv5 guest is no longer restricted to PPIs, and can make
: use of SPIs and LPIs as well."
: /
KVM: arm64: vgic-v5: Correctly handle host ISTE __le32 conversion
KVM: arm64: vgic-v5: Tidy-up programming of vpe descriptor address
KVM: arm64: vgic-v5: Correctly reset h_lpi_ist to NULL
KVM: arm64: vgic-v5: Drop __iomem attribute from {vmd,vpet}_base
KVM: selftests: Add VGICv5 sparse vCPU IDs test
KVM: selftests: Add VGICv5 IST save/restore coverage
KVM: selftests: Add VGICv5 LPI delivery tests
KVM: selftests: Add VGICv5 SPI injection tests
KVM: selftests: Add VGICv5 CPU sysreg attribute tests
KVM: selftests: Add VGICv5 USERSPACE_PPIS tests
KVM: selftests: Add VGICv5 IST attribute tests
KVM: selftests: Add VGICv5 IRS_REGS attribute tests
KVM: selftests: Add VGICv5 NR_IRQS attribute tests
KVM: selftests: Add VGICv5 IRS address attribute tests
Documentation: KVM: Add the VGICv5 IRS save/restore sequences
Documentation: KVM: Add docs for KVM_DEV_ARM_VGIC_GRP_IST
Documentation: KVM: Add KVM_DEV_ARM_VGIC_GRP_IRS_REGS to VGICv5 docs
Documentation: KVM: Document KVM_DEV_ARM_VGIC_GRP_CPU_SYSREGS for VGICv5
KVM: arm64: gic-v5: Implement save/restore mechanisms for ISTs
KVM: arm64: gic-v5: Add VGICv5 IST save/restore UAPI
...
Signed-off-by: Marc Zyngier <maz@kernel.org>
|
|
'arm64-fixes-for-7.3', 'arm64-for-7.4', 'clk-fixes-for-7.3', 'clk-for-7.4', 'drivers-fixes-for-7.3' and 'drivers-for-7.4' into for-next
|
|
TRBE doesn't support sysfs mode, but the enable_sink file can still be
successfully written to enable the device, and only attempting to enable
the source would later fail.
Avoid misleading users by adding a flag that devices can use to hide
either the enable_sink or enable_source files, and set it for TRBE.
Don't set it for ETE as it's possible that ETE could appear on the
legacy bus and work with sysfs, and writing to enable_source already
reports EINVAL if the device doesn't support sysfs mode.
Signed-off-by: James Clark <james.clark@linaro.org>
Reviewed-by: Leo Yan <leo.yan@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Link: https://lore.kernel.org/r/20260807-james-cs-hide-trbe-enable-v2-1-0b2af223feed@linaro.org
|
|
A file that was opened before pps_gen_unregister_source() keeps
pps_gen, and with it the pointer to the driver's pps_gen_source_info.
PPS_GEN_SETENABLE and PPS_GEN_USESYSTEMCLOCK keep using that pointer
after the driver has gone:
- pps_gen_tio allocates its info with devm_kzalloc(), so after an
unbind both ioctls read freed memory, and SETENABLE calls the
driver's enable() with its freed private data.
- pps_gen-dummy does not set info->owner, so it can be unloaded while
the file is open, and the ioctls then read the unloaded module's
data and call its code.
TIO can't probe in a VM, as it needs ART. A test driver that registers
its info in devm memory like TIO does, unbound while /dev/pps-gen0 is
open, gives:
BUG: KASAN: slab-use-after-free in pps_gen_cdev_ioctl+0x4c1/0x580
Read of size 1 at addr ffff88800e647f38 by task ppsgen64/145
...
Freed by task 146:
kfree+0x25a/0x6d0
release_nodes+0xd1/0x140
devres_release_all+0x10e/0x1a0
device_unbind_cleanup+0x71/0x250
device_release_driver_internal+0x41b/0x570
unbind_store+0xd9/0x100
and unloading pps_gen-dummy while its file is open oopses:
BUG: unable to handle page fault for address: ffffffffa0203040
Oops: Oops: 0000 [#1] SMP KASAN NOPTI
RIP: 0010:pps_gen_cdev_ioctl+0x372/0x580
Protect info in the ioctl handler with a mutex, and clear it under the
mutex on unregister, so that unregister waits for the ioctls using it
and later ones fail with -ENODEV. PPS_KC_BIND does the same for a
removed PPS device since commit 3649f9a6b897 ("pps: don't allow
PPS_KC_BIND on removed devices"). The sysfs attributes need nothing
new: they are removed, and their callbacks drained, before info is
cleared.
Fixes: 86b525bed275 ("drivers pps: add PPS generators support")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Danish Khateeb <danishkhateeb03@gmail.com>
Acked-by: Rodolfo Giometti <giometti@enneenne.com>
Link: https://patch.msgid.link/20260929124937.51114-3-danishkhateeb03@gmail.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
The cdev of a PPS generator is embedded in struct pps_gen_device, but
nothing ties the lifetime of that structure to the cdev: pps_gen is
freed by the release function of its device, and an open file holds a
device reference only until pps_gen_cdev_release() drops it.
When the generator is unregistered while /dev/pps-genN is open, that
put_device() drops the last reference and frees pps_gen, and __fput()
then calls cdev_put() on the freed cdev:
BUG: KASAN: slab-use-after-free in cdev_put+0x53/0x60
Read of size 8 at addr ffff88801383e138 by task ppsgen64/149
Call Trace:
cdev_put+0x53/0x60
__fput+0x745/0xad0
fput_close_sync+0xd9/0x1b0
__x64_sys_close+0x86/0xf0
...
Freed by task 149:
kfree+0x25a/0x6d0
device_release+0xca/0x3c0
kobject_put+0x169/0x320
pps_gen_cdev_release+0x51/0x80
__fput+0x36a/0xad0
pps.c had the same bug, fixed in commit c79a39dc8d06 ("pps: Fix a
use-after-free").
Fix it the usual way: embed the struct device in pps_gen_device and
register both with cdev_device_add(). This makes the device the parent
of the cdev, so the cdev holds a device reference until the last file
is closed.
Fixes: 86b525bed275 ("drivers pps: add PPS generators support")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Danish Khateeb <danishkhateeb03@gmail.com>
Acked-by: Rodolfo Giometti <giometti@enneenne.com>
Link: https://patch.msgid.link/20260929124937.51114-2-danishkhateeb03@gmail.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Paolo Abeni:
"Including fixes from Bluetooth, WiFi and netfilter.
We are actively retargeting several non-urgent fixes towards next,
but the traffic on the ML looks ever-increasing, and propagating the
push-back towards subsystems is not immediate.
No known outstanding regressions.
Current release - regressions:
- netfilter: nft_set_rbtree: skip transaction elements during GC
Previous releases - regressions:
- sched: cls_api: reclaim an empty proto on the error path
- core:
- fix checksum offsets in skb_splice_from_iter()
- cap skb->queue_mapping when the tx queue is picked
- page_pool: fix use-after-free in page_pool_recycle_ring_bulk()
- wifi:
- mac80211: fix slab-out-of-bounds read in ieee80211_monitor_select_queue()
- mac80211: drop oversized fragments to avoid extra_len overflow
- netfilter:
- flowtable: restore ieee80211 forward path
- bluetooth: hci_conn: Lock parent access during enhanced SCO setup
- eth:
- bcmgenet: allocate RX buffers as page fragments
- stmmac: fix rx Scatter-Gather support
- octeontx2-pf: fix aura BPID assignment when CONFIG_DCB is enabled
- gve: DQO: accept TSO packets with non-protocol gso_type bits
- r8169: disable EEE on RTL8168h/8111h
Previous releases - always broken:
- tcp: refresh TS.Recent for accepted old ACKs
- wifi:
- ath11k: reset ar->num_stations on hardware start
- cfg80211: fix RTS threshold setting for single-radio PHY
- bluetooth: btintel_pcie: fix plen overflow in btintel_pcie_recv_frame()
- eth: bcmgenet: fix NULL dereference in set_coalesce before first open
Misc:
- Eric is retiring from google and updating his contact info"
* tag 'net-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (96 commits)
net: phy: aquantia: fix system interface type not updated in forced mode
net: usb: qmi_wwan: add Rolling Wireless RN947R
net: mvneta: clear XDP pfmemalloc flag between frames
ipv6: sr: use skb_get_hash_net() in seg6_make_flowlabel()
net/mlx5e: Fix AF_XDP TX timestamp teardown NULL dereference
r8169: disable EEE on RTL8168h/8111h
octeontx2-pf: Fix RSS indirection table size
sctp: check RCV_SHUTDOWN after the sendmsg connect wait
net: sparx5: make ports inherit the switch base mac address type
net: microchip: vcap: stop scanning after deleting key field
netfilter: flowtable: restore ieee80211 forward path
netfilter: flowtable: generalize pending status bit
netfilter: bpf: reject invalid NAT manipulation types
netfilter: nft_set_rbtree: skip transaction elements during GC
ipvs: filter some flags received in the backup server
ipvs: do not create invisible templates
ipvs: bound LBLCR and LBLC cache growth
ipvs: fix missing counter decrement in lblc
netfilter: nft_flow_offload: drop flowtable reference on init error path
selftests: net: check timestamp echo after an old ACK
...
|
|
|
|
The omap24xx SoC platform was used in the Nokia N800 and N810 tablets
released on 2007 and was among the first ARMv6 CPUs in products that
made it into products.
Unfortunately the early ARM1136r0 CPU cores had a number of quirks that
caused a disproportionate amount of work to keep them compatible with
later ARMv6K/v7/v8 implementations.
Drop the support for this SoC in order to allow removing ARM1136r0.
Cc: Andreas Kemnade <andreas@kemnade.info>
Cc: Kevin Hilman <khilman@baylibre.com>
Cc: Roger Quadros <rogerq@kernel.org>
Cc: Tony Lindgren <tony@atomide.com>
Cc: Paul Walmsley <paul@pwsan.com>
Cc: Herbert Xu <herbert@gondor.apana.org.au>
Cc: "David S. Miller" <davem@davemloft.net>
Cc: Mauro Carvalho Chehab <mchehab@kernel.org>
Cc: Kyungmin Park <kyungmin.park@samsung.com>
Cc: Helge Deller <deller@gmx.de>
Cc: Tero Kristo <kristo@kernel.org>
Cc: Stephen Boyd <sboyd@kernel.org>
Cc: linux-kernel@vger.kernel.org
Cc: linux-omap@vger.kernel.org
Cc: devicetree@vger.kernel.org
Cc: linux-media@vger.kernel.org
Cc: linux-mtd@lists.infradead.org
Cc: linux-fbdev@vger.kernel.org
Cc: dri-devel@lists.freedesktop.org
Cc: linux-clk@vger.kernel.org
Reviewed-by: Thomas Zimmermann <tzimmermann@suse.de> # omapfb
Acked-by: Aaro Koskinen <aaro.koskinen@iki.fi>
Signed-off-by: Arnd Bergmann <arnd@arndb.de>
Acked-by: Kevin Hilman <khilman@baylibre.com>
Link: https://patch.msgid.link/20260914140821.1805449-7-arnd@kernel.org
Signed-off-by: Krzysztof Kozlowski <krzk@kernel.org>
|
|
|
|
ssh://gitolite.kernel.org/pub/scm/linux/kernel/git/jic23/iio into char-misc-next
Jonathan writes:
IIO: 1st set of device support, features and cleanup for 7.4
2 merges
- v7.3-rc4 to get a x86 build fix needed by the ad9910 driver.
- devm_notifier_chain_register to pick up the IIO specific users of this
new infrastructure.
New device support
------------------
Major refactors or new drivers needed:
adi,ad3530r
- Support for the AD5710R and AD5711R 8 channel IDAC / VDAC parts.
Needed substantial driver rework to support channel type controls.
adi,ad5529r
- New driver for this 16 channel DAC with 12 and 16 bit variants.
adi,ad7768
- New driver for the 8 (AD7768) and 4 (AD7768-4) channel variants of this
24-bit simultaneous sampling ADC family. They are very different from
the 1 channel variant that has as separate driver.
adi,ad9910
- New driver for this DDS (Direct Digital Synthesizer). This was a large
undertaking is it is a complex chip requiring quite a lot of new IIO ABI.
Delayed from last cycle by an x86 bug that was fixed in rc4.
adi,ade9000
- Add support for the ADE9078 polyphase energy metering device.
Required significant changes to introduce multiple part support into
this driver.
adi,adis16201
- Add support for the ADIS16203 device, replacing the driver in staging.
axiado,saradc
- New driver for the SARADC found on the AX3000 and ADX3005 SoCs.
liteon,ltr501
- Add support for the LTR-329ALS-01 ambient light sensor. Included
various bits of driver modernization.
microchip,mcp48fev02
- Split previously I2C specific mcp47feb02 into core and I2C specific
parts.
- Add an SPI specific part for the many parts in the MCP48FxBy1/2/4/8
series of SPI DACS 24 different part numbers covering:
- 8, 10 and 12 bit resolutions
- 1, 2, 4, and 8 channels
- EEPROM and no non volatile storage variants.
- Substantial driver modification and cleanup was needed prior to the new
support.
maxim,max40080
- New driver for this current-sense amplifier with integrated ADC.
renesas,rzg2l_adc
- Support the RZ/G3L ADC 1 that is dedicate to the on-chip thermal
sensor unit. The thermal parts of this support will be going through
that tree.
sensortek,stk3310
- Add support for the STK36C61 Ambient Light, proximity and RGB color
sensor. Handling had to be added for RGB elements and the driver in
general had to be made ready for support multiple parts.
ti,ads1100
- Add support for the ADS1110 with faster data rate and 2.048V reference.
ti,ads112c04
- New driver to support this 16bit sigma-delta ADC.
vishay,veml6031x00
- New driver for this ambient light sensor that was built in in stages.
- Add triggered buffer data capture.
- Add support for events and triggers.
- Enable mysterious 'internal calibration' setting.
Minor additions such as IDs, compatibles or device specific data.
adi,ad3530r
- Support the AD3536R which is low resolution equivalent of the already
supported AD3532R.
adi,ad4080
- Support the AD4885 - only needed chip specific data.
aosong,am2315
- Add support for the AM2320, fully compatible wth the AM2315.
invensense,icm42600
- Minor tweaks + addition of compatible for the ICM42630 IMU.
mediatek,mt2701-adc
- Compatible for the MT6572
rockchip,saradc
- Support the RV1106.
Features
--------
IIO Core
- Add a helper to discover if timestamps are enabled. In a few cases
(typically when there is batching involved) we may have to refuse
to start the buffer if timestamps are enabled. Provided a helper
to allow that to be queried.
IIO backend
- CRC support.
infineon,dps310
- Add triggered buffer support.
linear,ltc2497
- Add 2x conversion speed mode.
- Add the internal temperature channel.
renesas,r9a09g077
- Add missing DMA properties to DT and enable use of DMA
- Expose the sampling frequency as per channel attributes.
ti,ad1015
- Add per channel labels from DT support.
ti,ads112c14
- Add dataready trigger using continuous mode if only one channel is
enabled.
- Add burnout current control including new IIO ABI.
- Add filter control.
- Add new ABI for settling time.
Cleanup and minor or late breaking fixes
----------------------------------------
Only calling out the more significant changes or ones that were carried
on out on several drivers.
multiple drivers
- Add sanity check for spi_device_get_match_data() returning NULL due
to use of driver_override.
- Add some useful __counted_by_ptr markings.
- Add some missing MODULE_DEVICE_TABLE() calls.
- Use dmaenginge_get_dma_device() instead of chan->device->dev
- A few precursor cleanups for planed devm_ object allocators.
- Disable autosuspend on exit. Note there is work ongoing to make
this unnecessary.
- Header reorders and IWYU being applied to what was included.
- Use devm_blocking_notifier_chain_register() to replace open coded
equivalent.
- Ensure some clk related structures are always fully initialized.
- Use the size checked iio_push_to_buffers_with_ts() replacement for
iio_push_to_buffers_with_timestamp().
- Avoid shadowing various error codes from store() callbacks.
adi,axi-adc
- Fix missing mutex_init()
adi,ad4080
- Ensure autoincrement is set for multi register reads.
adi,ad5758
- Fix wrong offset calculation
- Reject out of range values rather than truncating values.
- Make sure buffer used for bulk acceses is DMA safe.
adi,ad7816
- Local lock to protect state by serializing accesses. A few extra locations
were added much later in the cycle.
- Stop using non DMA safe buffers by using spi_write_then_read()
to ensure the data was bounce buffered.
adi,adxl367
- Ensure interrupt mapping is resolved before device setup reducing work
done before potential deferral.
- Support INT2 pin.
asahi-kasei,ak8975
- Use BIT() and GENMASK() to improve readability of defines.
- Add an enum for scan mask elements.
- Switch to devm_ for all resources.
avago,apds9306
- Fix the default sampling frequency.
awinc,aw96103
- Fix an early return that stopped all channels being checked.
broadcom,iproc-adc
- General modernization and conversion to devm managed cleanup.
- Use dev_err_probe(), local dev pointer and drop some duplicate error
messages.
capella,cm32181
- Add device specific dt-binding as trivial-devices doesn't provide
vdd-supply or interrupt properties.
freescale,mma8452
- Modernize driver via guard(), IIO specific cleanup helpers and
bulk regulator handling.
infineon,dps310
- Add context analysis lock markings.
- Single local acquire / release per raw read.
- Replace opencoded get_unaligned_be24().
- Fix incorrect CFG_REG bit definitions.
invensense,icm42600
- Ensure ODR is updated when switching power mode to paper over a hardware
bug.
- Fix detecting ODR change in invalid data packets.
- Optimize high frequency data readback by not bothering to read the
FIFO count when the watermark is known to have been passed.
- Various other minor fixes.
isil,isl29028
- Fix a runtime pm reference leak.
liteon,ltr501
- Make the dt-binding property proximit-near-level only available
on parts that support it.
kionix,kx022a
- Use as structure to make the channel layout explicit and use that for
all cases (previously there were two similar buffers).
pulsed-light,lidar-lite-v2
- Move binding from trivial-devices to it's own file as there are power
enable and mode control pins that needs describing.
renesas,r9a09g077
- Drop some interrupt descriptions for interrupts that aren't connected to
the interrupt controller.
- Avoid direct reads when buffer enabled.
- Store the device PA and IRQ in state ready for the DMA support.
st,hts221
- Fix a potential division by zero if calibration data corrupted.
- Allow an unknown Who Am I value to support fallback dt-compatibles.
- Replace opencoded oversampling_ration_available handling with read_avail
callback.
st,stm32-adc
- Fix up some wrong minimum sample times.
ti,ads1100
- Fix readings after datarate change by ensuring a long enough wait for
a new sample.
ti,tmp117
- Fix case where cached state was updated before we know if the write
to the device succeeds.
vishay,veml3328
- Reshape the scale array to improve readability as it is made up
of value pairs.
* tag 'iio-for-7.4a' of ssh://gitolite.kernel.org/pub/scm/linux/kernel/git/jic23/iio: (270 commits)
iio: dac: ad5758: Fix the offset calculation
iio: dac: ad5758: Reject out-of-range raw values
iio: dac: ad5758: Fix alignment for DMA safety
iio: pressure: dps310: check the lock markings with context analysis
iio: core: add an accessor for scan_timestamp
iio: pressure: dps310: add triggered buffer support
iio: pressure: dps310: take the lock once per raw read
iio: pressure: dps310: use get_unaligned_be24() for the 24-bit results
iio: pressure: dps310: use a local device pointer in probe
iio: pressure: dps310: fix CFG_REG bit definitions
staging: iio: adc: ad7816: Protect sysfs attributes with mutex
dt-bindings: iio: Add dedicated CM32181 schema
iio: adc: stm32: fix minimum sampling time values
iio: light: veml6031x00: Enable internal calibration
iio: dac: add support for Microchip MCP48FEB02
dt-bindings: iio: dac: add support for MCP48FEB02 SPI
iio: dac: mcp47feb02: refactor MCP47FEB02 I2C driver into two modules
iio: dac: mcp47feb02: protect EEPROM store sequence with mutex
iio: dac: mcp47feb02: Avoid unjustified probe error on missing label
iio: adc: rzg2l_adc: Add RZ/G3L ADC support for TSU
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf
Pablo Neira Ayuso says:
====================
Netfilter/IPVS fixes for net
The following batch contains Netfilter fixes for net. This batch
fixes crashes as recent feature regression, one of the due to a
dependency that has been pulled into -stable:
1) Expand existing ipset fix for bitmap sets to disallow comments
updates from kernel-side adds, from Florian Westphal.
2) Drop flowtable reference if nf_ct_netns_get() fails, otherwise
flowtable cannot ever be removed, from Aohan Mei.
3) nft_rbtree GC should collect end elements that contained in
this transaction batch, new or deleted elements are never
expired. From Weiming Shi.
4) Restrict nf_nat_bpf so it does not set unknown NF_NAT_MANIP_*
values, from Fernando F. Mancera.
5) Flowtable GC must skip flows that are pending hardware updates,
generalize the PENDING flag and use it to inhibit GC.
6) Restore flowtable with ieee80211 which broke due to a relatively
recent commit, which was pulled in by -stable, causing a regression
in 6.18 kernels.
And the following IPVS fixes:
1) Fix accounting of cache entries in IPVS LBLC for destinations,
which eventually fills up the table and trigger recurrent
resizing, from Julian Anastasov.
2) Limit IPVS cache growth for LBLCR and LBLC schedulers,
from Zhiling Zou.
3) Restrict IP_VS_CONN_F_ONE_PACKET for normal connections,
do not allow to use it with templates. Also from Julian.
4) Sanitize flags in IPVS sync messages received in the backup.
From Julian Anastasov.
netfilter pull request 26-09-30
* tag 'nf-26-09-30' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf:
netfilter: flowtable: restore ieee80211 forward path
netfilter: flowtable: generalize pending status bit
netfilter: bpf: reject invalid NAT manipulation types
netfilter: nft_set_rbtree: skip transaction elements during GC
ipvs: filter some flags received in the backup server
ipvs: do not create invisible templates
ipvs: bound LBLCR and LBLC cache growth
ipvs: fix missing counter decrement in lblc
netfilter: nft_flow_offload: drop flowtable reference on init error path
netfilter: ipset: do not update comments from kernel-side adds
====================
Link: https://patch.msgid.link/20260930074142.298353-1-pablo@netfilter.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
Resetting saved termios state on device registration is needed where a
minor number can be reused for an entirely different device and where
the old settings may prevent the port from even being opened (e.g. when
CLOCAL is not set).
Not all TTY drivers guarantee that the minor number is no longer in use
when registering devices however, something which can lead to a
use-after-free when closing a TTY (and saving its termios) races with
re-registration.
Add a new TTY_DRIVER_RESET_SAVED_TERMIOS flag to request that any saved
termios state is reset on registration and only set it for drivers that
make sure that the minor number is no longer in use.
Fixes: 93857edd9829 ("tty: reset termios state on device registration")
Reported-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Link: https://lore.kernel.org/20260926184154.3017929-1-nicoyip.dev@gmail.com
Cc: stable@kernel.org # 4.12
Signed-off-by: Johan Hovold <johan@kernel.org>
Link: https://patch.msgid.link/20260930131845.1809256-1-johan@kernel.org
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
802.2 LLC is only used in tree by protocols which need the
connectionless type 1 subset: STP, which carries the bridge's BPDUs,
GARP, which sends its own UI PDUs, and SNAP. The llc2 module on top of
the core - the type 2 connection state machine, the type 1 SAP state
machine, the station component and the PF_LLC socket family - has no
in-kernel users and nobody who can test it; what we get instead is a
slow trickle of drive-by fixes.
There was a recent patch from Ernestas Kulik indicating potential
real life use, but it was new/experimental and that person is
not responding to off-list pings.
Let LLC2 follow AX.25, hamradio and AppleTalk out of the Linux tree.
We will maintain the code at: github.com/linux-netdev/mod-orphan
for anyone interested in playing with it.
PF_LLC goes in full, both the class two SOCK_STREAM and the class one
SOCK_DGRAM half, and so do /proc/net/llc/ and /proc/sys/net/llc/. Note
that the kernel also stops answering XID and TEST commands, addressed
to a SAP or to the station - those are type 1, but they lived in the
module, and they got answered whether or not any socket was open.
Nothing in tree asks for them; what the core keeps is SAP registration
and the UI path the in-tree users need.
Retain the uAPI for now, like we did for AppleTalk. Only the socket ABI
half of it is vestigial: STP, GARP, the bridge and openvswitch use the
SAP numbers it defines. Cleaning up what the core no longer needs
follows in the next patch.
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Link: https://patch.msgid.link/20260928190800.2521749-2-kuba@kernel.org
Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
Intel Diamond Rapids exposes Application Energy Telemetry through two new PMT
aggregator GUIDs: one covering the "energy" feature, and one covering the
"perf" feature, both documented in the Intel-PMT XML repository at
https://github.com/intel/Intel-PMT
The "perf" aggregator references two events not previously enumerated by
resctrl: retired instructions (INST_RETIRED), and a count of unhalted core
clock cycles perceived by hardware as contributing to instruction execution
(PCNT).
Add definitions for both aggregators, and add the two new events to the
resctrl filesystem so they can be exposed to user space.
Signed-off-by: Tony Luck <tony.luck@intel.com>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Reviewed-by: Reinette Chatre <reinette.chatre@intel.com>
Link: https://patch.msgid.link/20260930171554.12794-1-tony.luck@intel.com
|
|
geni_icc_get() takes an icc_ddr string argument but never uses it: the
DDR interconnect path is looked up with the "qup-memory" string literal
directly, and the sole caller passes that same literal. The parameter
carries no information.
Drop the icc_ddr parameter and update the prototype and call site
accordingly.
Signed-off-by: Somesh Dey <somesh.dey@oss.qualcomm.com>
Reviewed-by: Konrad Dybcio <konrad.dybcio@oss.qualcomm.com>
Link: https://patch.msgid.link/20260930-icc-ddr-param-remove-v1-1-61223714c9bc@oss.qualcomm.com
Signed-off-by: Bjorn Andersson <andersson@kernel.org>
|
|
https://gitlab.freedesktop.org/drm/kernel into drm-next
pci/vgaarb rework for drm/vfio
This reworks the vgaarb API to be a bit more sane.
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Dave Airlie <airlied@gmail.com>
Link: https://patch.msgid.link/CAPM=9tzd3LteJG-Au=na-_X5Rff4tacFZ3X_cp5g0bCcQkV_jA@mail.gmail.com
|
|
Sashiko reported [1] that 'Missing __GFP_ZERO in folio_alloc() causes
uninitialized kernel memory to be exposed in the receive message buffer'.
An ism dmb is receive-only, so the data is not leaked to a remote peer.
In general the smc kernel module (dibs client) will only push newly
received data to userspace. We still should not have uninitialized data
in a receive buffer.
Since
commit 750afb08ca71 ("cross-tree: phase out dma_zalloc_coherent()")
dma_alloc_coherent no longer required the __GFP_ZERO flag, but when
commit 83781384a96b ("s390/ism: Properly fix receive message buffer allocation")
switched to folio_alloc(), it should have added back the __GFP_ZERO flag.
Add __GFP_ZERO flag and state in dibs.h that register_dbm() provides
a zerorized buffer (dibs_lo already does).
Link: https://lore.kernel.org/linux-s390/20260903143746.A5CC41F00A3A@smtp.kernel.org/ [1]
Fixes: 83781384a96b ("s390/ism: Properly fix receive message buffer allocation")
Signed-off-by: Alexandra Winter <wintera@linux.ibm.com>
Reviewed-by: Julian Ruess <julianr@linux.ibm.com>
Link: https://patch.msgid.link/20260928151420.383105-1-wintera@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Memory-provider queue configuration is validated when the provider is
bound. A later ethtool ring change may invalidate it because drivers
can size queue memory from both ring depth and RX page size. The fbnic
consumer is added in the following patch.
Keep configured RX ring depths in netdev_config and stage proposed
values in cfg_pending. Validate every RX queue before calling the
driver. Each check validates the device defaults, then any queue
memory-provider override. Commit the values only after the driver
accepts them.
Drivers which consume stored ring depths through queue configuration
must initialize every RX depth before registering the netdev. Stored
values override callback defaults, including when zero.
The callback receives a rendered configuration rather than a queue ID.
Validation should depend on the configuration, not queue identity.
Checking defaults also covers the case where every queue has a
memory-provider override.
Drivers may normalize ring depths when applying them. Require the
validation callback to use the same normalization. Drivers must report
the applied depths through the ethtool_ringparam argument so the core
records the result.
Use the same transaction for ioctl and netlink. Drivers without
ndo_validate_qcfg skip the new validation.
Link: https://lore.kernel.org/all/20250421222827.283737-14-kuba@kernel.org/
Signed-off-by: Björn Töpel <bjorn@kernel.org>
Reviewed-by: Simon Horman <horms@kernel.org>
Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260925104417.2325213-4-bjorn@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
- Pass the correct physical and virtual endpoint function numbers when
setting translation of inbound memory windows (Koichiro Den)
- Make embedded doorbell IRQ non-shared, since currently only the first EPF
can allocate doorbells (Koichiro Den)
- MSI doorbells currently target the first EPF attached to an EPC; let
other EPFs try the embedded doorbell instead of failing when they can't
use the MSI doorbell (Koichiro Den)
- Serialize scan of vNTB virtual PCI bus with other PCI topology changes
(Koichiro Den)
- Manage vNTB lifetimes to avoid leaking virtual devices and buses and
allow removal (Koichiro Den)
* pci/endpoint:
PCI: endpoint: pci-epf-vntb: Manage virtual NTB and PCI bus lifetime
PCI: endpoint: pci-epf-vntb: Serialize virtual PCI bus scan
PCI: endpoint: pci-ep-msi: Let non-first EPFs use embedded doorbells
PCI: endpoint: pci-ep-msi: Make embedded doorbell IRQ exclusive
PCI: endpoint: pci-epf-vntb: Pass PF/VF number when BAR programming
|
|
- Use regmap APIs for tc9563 I2C accesses instead of open-coding I2C
transfers (Lorenzo Bianconi)
- Use devm-managed tc9563 I2C dummy device allocation (Lorenzo Bianconi)
- Allow tc9563 RESX reset assertion to sleep (Abel Vesa)
- Add DT documentation and driver support for TC9563 PCIe switch embedded
GPIO controller (Lorenzo Bianconi)
- TC9563 aux device
- TC9563 per-port reset
* pci/pwrctrl:
arm64: dts: qcom: qcs6490-rb3gen2: Enable TC9563 embedded GPIO controller
PCI/pwrctrl: tc9563: Switch per-port reset to GPIO descriptor API
PCI/pwrctrl: tc9563: Add GPIO auxiliary device support
gpio: tc9563: Add support for the embedded GPIO controller
dt-bindings: PCI: toshiba,tc9563: Document embedded GPIO controller
PCI/pwrctrl: tc9563: Allow RESX reset assertion to sleep
PCI/pwrctrl: tc9563: Use devm-managed I2C dummy device allocation
PCI/pwrctrl: tc9563: Rely on regmap APIs
|
|
Add 'input-debounce-ns' to the generic parameters used for parsing DT
files, along with the corresponding configuration parameter
PIN_CONFIG_INPUT_DEBOUNCE_NS. This allows debounce time to be specified
in nanoseconds as an alternative to the existing 'input-debounce'
property which uses microseconds
Signed-off-by: Changhuang Liang <changhuang.liang@starfivetech.com>
Signed-off-by: Linus Walleij <linusw@kernel.org>
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless-next
Johannes Berg says:
====================
More features:
- ath10k: NVMEM device tree bindings
- ath12k: QMI firmware alignments
- mm81x: AP improvements
- mac80211:
- CIP (control frame integrity) support
- NAN improvements
- cfg80211:
- improvements for AP regulatory checks
* tag 'wireless-next-2026-09-30' of https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless-next: (67 commits)
wifi: mac80211: Gracefully deauthenticate on association timeout
wifi: libertas: fix RX OOB access from device-controlled pkt_ptr
wifi: mwifiex: Reattach interfaces on suspend failure
wifi: b43legacy: work around stack frame size warning
wifi: cw1200: fix link_id_db OOB access via device-controlled link ID
wifi: mwifiex: bound SDIO fw dump count by memory table size
wifi: mac80211: start next ROC after purging an interface
wifi: mac80211: fix potential ack-skb leak on error path
wifi: mac80211: mesh: don't send peering close in listen
wifi: mac80211_hwsim: Support NAN deferred schedule completion
wifi: mac80211: Restrict probe request rates for minimal content
wifi: nl80211: allow a NAN peer schedule with 2 channels in the same slot
wifi: cfg80211: use sysfs_emit_at() in addresses_show
wifi: cfg80211: require zero terminator in valid_regdb country table
wifi: cfg80211: fix NAN local schedule update ordering and allocation
wifi: mac80211: Fix a race when expiring a mesh path
wifi: cfg80211: validate monitor channel set against radio usage
wifi: nl80211: reject color-change requests that change 6 GHz power type
wifi: nl80211: defer AP beacon regulatory check to start_ap
wifi: radiotap: add definitions for UHR U-SIG
...
====================
Link: https://patch.msgid.link/20260930124547.228697-29-johannes@sipsolutions.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
|
|
Fix netfs to erase the contents of a hole created after the EOF by an
ordinary write if dirty data has been previously left there by writes
through an mmapped region. Neither the buffered nor the unbuffered/DIO
write path clears that stale pagecache.
Zero the tail of the folio straddling the EOF before an extending
write. That is the only folio that can hold data written past the EOF
through an mmap, as pages wholly beyond the EOF can't be faulted in.
The folio is zeroed rather than dropped so a concurrent extending write
can't lose data.
Both write paths downgrade the i_rwsem to shared, so extending writes
can run concurrently and the i_size read by the caller may be stale by
the time the folio is locked. Re-read i_size under the folio lock and
clamp the zeroed range up to it, so a racing write that already put
data into the folio isn't clobbered.
Wait for any writeback on the folio to finish before zeroing it so that
the pagecache isn't modified while it may still be read by the transport
during transmission. Honour IOCB_NOWAIT by returning -EAGAIN rather
than blocking on the folio lock, on writeback, or in folio_mkclean()'s
rmap walk when the folio is mapped.
truncate_pagecache() can't be used here: it must be called with the
i_rwsem held exclusively, but these write paths only hold it shared,
and it would block unconditionally, breaking IOCB_NOWAIT.
Callers that hold i_rwsem exclusively for the whole resize (truncate,
setattr, fallocate, clone) exclude any genuine concurrent buffered
writer, so staleness can instead be decided from the folio's dirty
state, as pagecache_isize_extended() already does for filesystems that
serialise writes against truncate/setattr via a single i_rwsem.
Export netfs_clear_stale_post_isize() helper to handle such case.
The helper is required by the CIFS client to fix generic/363.
Closes: https://sashiko.dev/#/patchset/20260921230755.1133425-1-pc%40manguebit.org
Fixes: 938e13a73b24 ("netfs: Implement buffered write API")
Fixes: 153a9961b551 ("netfs: Implement unbuffered/DIO write support")
Reviewed-by: David Howells <dhowells@redhat.com>
Reviewed-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
|
|
Yu-Chun Lin <eleanor.lin@realtek.com> says:
This patch series introduces support for the Realtek RTD1625 SPI NOR Flash
controller.
Link: https://patch.msgid.link/20260930065945.88008-1-eleanor.lin@realtek.com
|
|
Each alarm field lives in the low bits of its register and the driver
masks its writes accordingly, so the high byte of four of them is storage
the RTC never touches. MediaTek names these RTC_NEW_SPARE0 to
RTC_NEW_SPARE3 and gives the first to a fuel gauge, which is how its PMIC
battery drivers carry a state of charge over a reboot.
Offer all four as a battery-backed nvmem provider, so that a consumer does
not have to reach into this block behind the driver's back. Doing it here
is what makes it safe: a write lands under the same lock the alarm paths
take, so it can neither be lost inside mtk_rtc_set_alarm()'s
read-modify-write nor fire the write trigger in the middle of one.
The nvmem core does not range check a cell against the provider size, so
the callbacks check the offset themselves.
Tested on an MT6397; MediaTek's spare map for mt6323 matches, and the
alarm field masks are common to every compatible this driver binds.
Assisted-by: LLM
Signed-off-by: Ryan Brue <ryanbrue.dev@gmail.com>
Reviewed-by: AngeloGioacchino Del Regno <angelogioacchino.delregno@collabora.com>
Link: https://patch.msgid.link/20260917-rbrue-suez-upstreaming-mt6397-rtc-nvmem-v1-2-558b8f95cfbc@gmail.com
Signed-off-by: Alexandre Belloni <alexandre.belloni@bootlin.com>
|
|
omap_rtc_power_off_program takes one argument that is never used. One call
passes the struct device of the rtc but the other one passes its parent.
To avoid confusion, stop taking any argument
Reviewed-by: Nishanth Menon <nm@ti.com>
Link: https://patch.msgid.link/20260917142654.2576889-1-alexandre.belloni@bootlin.com
Signed-off-by: Alexandre Belloni <alexandre.belloni@bootlin.com>
|
|
Add virtual I3C bus support for the hub and provide interface to enable
or disable downstream ports.
Signed-off-by: Aman Kumar Pandey <aman.kumarpandey@nxp.com>
Signed-off-by: Vikash Bansal <vikash.bansal@nxp.com>
Signed-off-by: Lakshay Piplani <lakshay.piplani@nxp.com>
Reviewed-by: Frank Li <Frank.Li@nxp.com>
Link: https://patch.msgid.link/20260922103551.2754613-7-lakshay.piplani@nxp.com
Signed-off-by: Alexandre Belloni <alexandre.belloni@bootlin.com>
|
|
Add CCC helpers to check CCC support and send CCC commands, address slot
helpers to query and update I3C bus address slot state, registering virtual
masters with an explicit firmware node, and exposing the bus maintenance
lock helpers.
These additions prepare for I3C hub support. A hub driver needs to reserve
and query parent bus address slots, forward CCC commands, register virtual
target port controllers using the target-port firmware node, and serialize
operations against the parent bus maintenance lock.
The hub also forwards private transfers via i3c_dev_do_xfers_locked() and
serializes its IBI and private-transfer paths against the shared lock, so
the normal-use lock/unlock pair is exposed alongside the maintenance-lock
helpers.
i3c_master_register_fwnode() allows virtual I3C masters to register using a
firmware node different from their parent device node without temporarily
modifying parent->of_node.
The new helpers are:
1) i3c_master_send_ccc_cmd()
2) i3c_master_supports_ccc_cmd()
3) i3c_bus_get_addr_slot_status()
4) i3c_bus_set_addr_slot_status()
5) i3c_bus_maintenance_lock()
6) i3c_bus_maintenance_unlock()
7) i3c_master_register_fwnode()
8) i3c_bus_normaluse_lock()
9) i3c_bus_normaluse_unlock()
10) i3c_dev_do_xfers_locked()
Signed-off-by: Aman Kumar Pandey <aman.kumarpandey@nxp.com>
Signed-off-by: Lakshay Piplani <lakshay.piplani@nxp.com>
Signed-off-by: Vikash Bansal <vikash.bansal@nxp.com>
Reviewed-by: Frank Li <Frank.Li@nxp.com>
Link: https://patch.msgid.link/20260922103551.2754613-2-lakshay.piplani@nxp.com
Signed-off-by: Alexandre Belloni <alexandre.belloni@bootlin.com>
|
|
Move the firmware attributes class helper from drivers/platform/x86 to
drivers/firmware and expose its class declaration through a public Linux
header.
The helper is not x86-specific. Keeping it in drivers/firmware lets
coreboot firmware drivers use the standard firmware-attributes ABI without
living under platform/x86.
Replace the affected drivers' relative helper includes directly with the
new public header.
Suggested-by: Derek J. Clark <derekjohn.clark@gmail.com>
Reviewed-by: Mark Pearson <mpearson-lenovo@squebb.ca>
Reviewed-by: Derek J. Clark <derekjohn.clark@gmail.com>
Tested-by: Oliver Lin <oliver@liuxiaozhen.dev>
Signed-off-by: Sean Rhodes <sean@starlabs.systems>
Acked-by: Ilpo Järvinen <ilpo.jarvinen@linux.intel.com>
Link: https://lore.kernel.org/r/5682c2228fa4a784d3953664b56a06e0dd9ccdef.1788284852.git.sean@starlabs.systems
Signed-off-by: Tzung-Bi Shih <tzungbi@kernel.org>
|
|
Before commit 871df5007eda ("netfilter: flowtable: bail out if forward
path cannot be discovered"), there was a fallback to set up a forward
path in case .ndo_fill_forward_path fails or DEV_PATH_MTK_WDMA was used.
Such fallback was used by commit d787a3e38f01 ("mac80211: add support
for .ndo_fill_forward_path").
One possibility is to handle DEV_PATH_MTK_WDMA from the flowtable
forward path discovery. However, this is only used internally by drivers
to retrieve mtk_wdma information to set up hardware offload. Felix
decided to use the .fill_forward_path interface for this purpose due to
the lack of a better interface at that time.
Add a new DEV_PATH_IEEE80211 path which is offered if the new ieee80211
flag is set on in the struct net_device_path_ctx to restore the
flowtable with a ieee80211 netdevice. Handle this new DEV_PATH_IEEE80211
path just like DEV_PATH_ETHERNET and DEV_PATH_DSA, ie. this is the last
netdevice in the stack.
This new ieee80211 flag is implicitly unset for mtk_ppe and airoha which
call dev_fill_forward_path() to retrieve a DEV_PATH_MTK_WDMA path.
Fixes: 871df5007eda ("netfilter: flowtable: bail out if forward path cannot be discovered")
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
This changes the vgaarb client API so that the user can pass a
private data pointer into the register that will get used in
the decode callback.
This allows a bunch of pdev conversions in the drivers, and lets
some future vfio cleanups be nicer.
I'd like to merge this via the drm next tree but also fine with
it going via pci.
Signed-off-by: Dave Airlie <airlied@redhat.com>
Reviewed-by: Jani Nikula <jani.nikula@intel.com>
Acked-by: Alex Williamson <alex@shazbot.org>
Acked-by: Bjorn Helgaas <bhelgaas@google.com>
Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Link: https://lore.kernel.org/dri-devel/20260922071807.2533884-1-airlied@gmail.com/
|
|
skbedit can set skb->queue_mapping and raise the per-CPU skip_txqueue
flag so __dev_queue_xmit() honours the mapping. __dev_queue_xmit()
cleared the flag before sch_handle_egress() and only read it afterwards,
so the flag was not confined to the xmit that set it: a nested xmit
(mirred redirect or mirror, or a drop after skbedit) could set the flag
and the outer xmit would consume it for an skb that never went through
skbedit.
A forwarded packet still carries the ingress NIC's rx_queue + 1 in
skb->queue_mapping, so the outer device then indexes its tx queue state
with that stale value. Taprio's child array q->qdiscs[] is sized to the
device's queue count, so taprio_enqueue() indexes past its allocation
and dereferences the result as a struct Qdisc *.
We (ab)use the skb->nf_skip_egress which means "skip netfilter egress
for this packet" to tag to "am I in tc egress?". Despite the overload
I dont see it as a conflict since the marker is set only around the
single sch_handle_egress() call and ingress path is guarded by
tc_at_ingress.
I will send a followup(net-next) patch once this hits net-next to
rename the skb->nf_skip_egress bit/flag to skb->skip_egress
Arm the flag only from the egress classifier that can use it: raise
skip_txqueue from tcf_skbedit_act() only when it runs inside
sch_handle_egress(), thanks to skb->nf_skip_egress. An egress qdisc
classifier runs in q->enqueue(), after the tx queue has been picked,
so a mapping it sets cannot affect the current packet; arming the flag
there only pollutes it for a later xmit. Then own the flag for the xmit
frame the egress hook runs in: save the incoming value and clear it just
before sch_handle_egress(), and restore it after the hook - on the
consumed (drop) path, or, in the same call that reads it, on the
surviving path. The save and the restores stay inside the
egress_needed_key static branch, so a packet pays for them only when
egress hooks are active (2f1e85b1aee4).
Store the value netdev_cap_txqueue() selected back into skb->queue_mapping
in netdev_tx_queue_mapping(), as netdev_core_pick_tx() already does, so
the skip_txqueue path never hands a later reader on the xmit path a
mapping the device cannot serve. A store made still later in the same
frame, by a tc BPF program attached to a transmit qdisc, is outside this
path and is not re-capped; a separate followup will resolve that path.
netdev_xmit_skip_txqueue() returns the previous flag value so the
save-and-clear is one call, and a no-op stub is provided when
CONFIG_NET_EGRESS is disabled. skb->nf_skip_egress is compiled under
CONFIG_NET_EGRESS rather than CONFIG_NETFILTER_SKIP_EGRESS, so
skb_at_tc_egress() is valid whenever the egress path is built.
A local user in a network namespace can redirect a packet from a device
with more TX queues to one with fewer after setting a mapping valid only
on the larger device. That reaches these reads and, under KASAN, faults
with "slab-out-of-bounds in taprio_enqueue".
Conditions to recreate the bug: the report's own trigger is a local user
with CAP_NET_ADMIN in a network namespace, so no eBPF program is needed.
With CONFIG_NET_SCH_TAPRIO=y, CONFIG_NET_ACT_SKBEDIT=y,
CONFIG_NET_ACT_MIRRED=y, CONFIG_NET_CLS_MATCHALL=y,
CONFIG_NET_SCH_PRIO=y and KASAN enabled, create qa (3 queues), qb
(2 queues) and qc (1 queue) as dummy devices; put a taprio root on qb
(num_tc 1, queues 2@0) and clsact on all three; then add an egress
matchall filter on every device. On qa: "action skbedit queue_mapping 2
pipe action mirred egress redirect dev qb". On qb: "action mirred egress
mirror dev qc". On qc: "action skbedit queue_mapping 0 pipe". Send one
packet out qa. qc's skbedit sets the flag while qb's outer xmit is in
flight; without the fix qb consumes it and reads its two-entry taprio
child array with the forwarded packet's stale mapping. A qc whose
skbedit is instead installed in a transmit-qdisc classifier (a matchall
filter on the qc root qdisc) reaches the same code path the same way
without the fix.
Testing: on a KASAN build with panic_on_warn=1 the unfixed kernel panics
with "BUG: KASAN: slab-out-of-bounds in taprio_enqueue", a read 0 bytes
past a 16-byte taprio_init() allocation, for the clsact-setter and the
transmit-qdisc-classifier reproducers and for a clsact skbedit-then-tc-BPF
store; the fixed kernel runs all three with no report, and the BPF store
variant additionally shows the expected "selects TX queue" clamp notice
from the write-back.
Fixes: 2f1e85b1aee4 ("net: sched: use queue_mapping to pick tx queue")
Reported-by: Zero Day Initiative <zdi-disclosures@trendmicro.com>
Link: https://lore.kernel.org/netdev/CANn89iLwYx8nCVf0pCEk_MmEiyC6kQaMwCQT9WkQVeeNzNQHqQ@mail.gmail.com/
Link: https://lore.kernel.org/netdev/179008581937.2160803.7117814290574262942@kernel.org/
Link: https://lore.kernel.org/netdev/179033713973.2160803.4914570693994398206@kernel.org/
Link: https://lore.kernel.org/netdev/20260925180407.63647514@kernel.org/
Link: https://lore.kernel.org/netdev/CANn89i+k-mZKDQVtvws_MEXeuMTAdaCcOXFZE-RfhcGTu90sjA@mail.gmail.com/
Suggested-by: Eric Dumazet <edumazet@google.com>
Suggested-by: Jakub Kicinski <kuba@kernel.org>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/QDISC-9R8V.v4.20260928081529@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
|
|
i2c_smbus_read_block_data() is dangerous to use because it may deliver
up to I2C_SMBUS_BLOCK_MAX (32) bytes, which may be surprising to the
caller. Callers tend to allocate buffers of sizes big enough to hold
data from a well-behaving device and do not expect that
i2c_smbus_read_block_data() may attempt to write more data than
expected.
To make i2c_smbus_read_block_data() safer to use, change it so that
it accepts size of the supplied buffer as another argument and ensure
that it will not copy more data than the size of the buffer. Signal
oversized responses with -EMSGSIZE.
To allow users to gradually transition to the new API employ some
macro trickery allowing calling i2c_smbus_read_block_data() with either
3 or 4 arguments. When called with 3 arguments it is assumed that
the buffer size is I2C_SMBUS_BLOCK_MAX bytes. Once everyone is
transitioned to the 4 argument form the macros should be removed.
Signed-off-by: Dmitry Torokhov <dmitry.torokhov@gmail.com>
Signed-off-by: Andi Shyti <andi.shyti@kernel.org>
Link: https://patch.msgid.link/ammWUROtUGeriY8G@google.com
|
|
hrtimer_rearm_deferred_user_irq() removes the rearm bit from the local
copy of the TIF work with:
*tif_work &= ~TIF_HRTIMER_REARM;
TIF_HRTIMER_REARM is the bit number (12), not the mask. That clears
TIF_NOTIFY_SIGNAL (bit 2) and TIF_MEMDIE (bit 3) from the copy and
leaves bit 12 set.
So the function never returns true, and the copy handed to
exit_to_user_mode_loop() lacks TIF_NOTIFY_SIGNAL. If that was the only
work for the loop, the task returns to user space with the task work
still pending. hrtimer_interrupt() sets TIF_HRTIMER_REARM every time,
so the next tick repeats that. Task work queued with TWA_SIGNAL from a
hrtimer callback stays pending until the task does a syscall, takes an
interrupt without deferred rearm or has to reschedule.
A task which polls the io_uring completion ring in user space sees a
timeout completion after 30ms (median) instead of 70us. 2 CPU QEMU
guest, HZ=1000.
Use the mask.
Fixes: 15dd3a948855 ("hrtimer: Push reprogramming timers into the interrupt return path")
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Assisted-by: LLM
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260919144027.70250-1-kmehltretter@gmail.com
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/mtd/linux
Pull MTD fixes from Miquel Raynal:
"The most important set of fixes are around the handling of the QE bit
in SPI NAND.
There are also a couple of behavioral fixes (mutex issue in SPI-NOR,
spurious bitflips on vf610_nfc, OOB bytes count in SPI NAND and
cfi_cmdset stack usage).
The rest is mostly AI fuzzing results"
* tag 'mtd/fixes-for-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/mtd/linux:
mtd: spinand: Do not update the QE bit on devices without one
mtd: spi-nor: core: Fix mutex leak in spi_nor_rww_start_exclusive()
mtd: rawnand: cadence: Initialize IRQ state before requesting IRQ
mtd: rawnand: vf610_nfc: fix false bitflips on reads of erased pages
mtd: rawnand: vf610_nfc: fix reads on chips with more than 64 bytes of OOB
mtd: spinand: fix zero oobavail when no ECC engine is used
mtd: spinand: fix NULL pointer dereference with no ECC engine
mtd: mtd_intel_dg: reset poll counter for each erase
mtd: cfi_cmdset_0001: shrink do_write_buffer() stack frame
mtd: core: call _get_device() with the master MTD
mtd: core: avoid double-free of OTP NVMEM device
mtd: spinand: Enable QE on all dies
mtd: block2mtd: Fix divide error when erase_size is zero
|
|
hlist_unhashed_lockless() locklessly samples the hlist_node structure's
pprev field, but detach_timer(), hlist_move_list(), and hlist_splice_init()
all use plain C-language stores to update this field, despite the "The
READ_ONCE() is paired with the various WRITE_ONCE() in hlist helpers that
are defined below" in the hlist_unhashed_lockless() header comment.
Therefore use WRITE_ONCE() for these hlist_node::pprev updates.
KCSAN located this issue.
Signed-off-by: Paul E. McKenney <paulmck@kernel.org>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Link: https://patch.msgid.link/20260919002523.3133928-1-paulmck@kernel.org
|
|
The hrtimer_sleeper structure's ->task field is now used only by the
hrtimer_sleeper_task_get() and hrtimer_sleeper_task_set() functions,
and there is no reason for it to be directly accessed anywhere else.
Therefore, mark this field __private and use ACCESS_PRIVATE() in
hrtimer_sleeper_task_get() and hrtimer_sleeper_task_set().
Suggested-by: Thomas Gleixner <tglx@kernel.org>
Signed-off-by: Paul E. McKenney <paulmck@kernel.org>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Link: https://patch.msgid.link/20260919001428.3133388-11-paulmck@kernel.org
|
|
The hrtimer_sleeper structure's ->task field is used as a flag to indicate
that the associated hrtimer has expired. This means that the hrtimer
handler can be storing to this field while other code is loading from it
to check for expiry. Note that additional races appear for hrtimers that
can be restarted, which could be argued to be a user error. However, that
is no reason to let the compiler introduce additional confusion, and to
this end, the hrtimer_sleeper_task_get() was introduced, use of which also
has the benefit of avoiding open-code access to hrtimer_sleeper innards.
Therefore, apply this accessor to the __wait_event_hrtimeout() macro.
KCSAN located this issue.
Signed-off-by: Paul E. McKenney <paulmck@kernel.org>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Link: https://patch.msgid.link/20260919001428.3133388-3-paulmck@kernel.org
|
|
The hrtimer_sleeper structure's ->task field is used as a flag to indicate
that the associated hrtimer has expired. This means that the hrtimer
handler can be storing to this field while other code is loading from it
to check for expiry. Note that additional races appear for hrtimers that
can be restarted, which could be argued to be a user error. However,
that is no reason to let the compiler introduce additional confusion.
Therefore, mark data-racy accesses to the hrtimer_sleeper ->task field
using READ_ONCE() (using a new hrtimer_sleeper_task_get() access function)
and WRITE_ONCE() (using a new hrtimer_sleeper_task_set() access function).
KCSAN located this issue.
Signed-off-by: Paul E. McKenney <paulmck@kernel.org>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Reviewed-by: Dmitry Ilvokhin <d@ilvokhin.com>
Link: https://patch.msgid.link/20260919001428.3133388-1-paulmck@kernel.org
|
|
The macro clk_div_mask() currently wraps to zero when width is 32 due to
1 << 32 being undefined behavior. This leads to incorrect mask generation
and prevents correct retrieval of register field values for 32-bit-wide
dividers.
Although it is unlikely to exhaust all U32_MAX div, some clock IPs may rely
on a 32-bit val entry in their div_table to match a div, so providing a
full 32-bit mask is necessary.
Fix this by using the standard GENMASK() macro. This safely resolves the
undefined behavior on both 32-bit and 64-bit architectures, while also
benefiting from the built-in compile-time type and bounds checking
provided by the GENMASK() macro.
Cc: Troy Mitchell <troy.mitchell@linux.spacemit.com>
Cc: Brian Masney <bmasney@redhat.com>
Signed-off-by: Junhui Liu <junhui.liu@pigmoral.tech>
Reviewed-by: Troy Mitchell <troy.mitchell@linux.spacemit.com>
Reviewed-by: Brian Masney <bmasney@redhat.com>
Reviewed-by: Jerome Brunet <jbrunet@baylibre.com>
Signed-off-by: Brian Masney <bmasney@redhat.com>
|
|
On 32-bit architectures where dma_addr_t is wider than unsigned long,
page_pool_set_dma_addr_netmem() stores page-aligned DMA addresses shifted
by PAGE_SHIFT. The net_iov branch of __skb_frag_dma_map() adds byte offsets
to the encoded value, so the NIC is programmed with an invalid DMA address.
This can trigger an IOMMU fault or DMA from unintended memory.
Consolidate DMA address encoding, decoding, and representability checks in
netmem helpers. Use the common decoder from the page pool and net_iov TX
paths so both interpret stored addresses consistently.
Reviewed-by: Mina Almasry <almasrymina@google.com>
Signed-off-by: Stanislav Fomichev <sdf@fomichev.me>
Link: https://patch.msgid.link/20260925201522.254717-4-sdf@fomichev.me
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
|
|
* fixes:
cpufreq: intel_pstate: Fix max_freq fallback in cpufreq_update_pressure()
|
|
* pm-runtime:
PM: runtime: call pm_runtime_dont_use_autosuspend() on reinit()
PM: core: Document struct dev_pm_info with kerneldoc
PM: runtime: Pull API docs from kerneldoc
PM: runtime: kerneldoc wording improvements
PM: runtime: Improve set_{status,active,suspended} docs
PM: runtime: kerneldoc fixes
PM: runtime: Add kunit test for supplier idle/suspend
PM: runtime: Only queue an idle check for RPM-linked suppliers
* pm-sleep:
PM: sleep: Add DPM watchdog to prepare/late/early/noirq/complete phases
PM: hibernate: docs: update swsusp.txt references to swsusp.rst
* pm-tools:
tools: power: pm-graph: fix typo "hierachy" in comment
|