| Age | Commit message (Collapse) | Author |
|
git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext
Pull sched_ext fixes from Tejun Heo:
- Taking a CPU offline could hang, or stall until the watchdog ejected
the BPF scheduler, when tasks on the dying CPU were still held by the
scheduler or sitting on a user dispatch queue. Re-enqueue them onto
the local queue when the runqueue goes offline so that the CPU pushes
them off like the other sched classes.
- The sequence number guarding against stale dispatches was per
runqueue, so a task re-enqueued on another CPU could get the same
number and a dispatch meant for its earlier instance was applied to
the new one. Use a per-task counter.
- A task dispatched to another CPU's local queue got its ops.dequeue()
only when picked to run and flagged as a core-sched pick. Call it at
insertion like for same-CPU dispatches.
* tag 'sched_ext-for-7.3-rc6-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext:
sched_ext: Generate qseq from a per-task counter
selftests/sched_ext: Add a test for ops.dequeue() on remote local DSQ moves
sched_ext: Call ops.dequeue() when a task arrives on a remote local DSQ
sched_ext: Fix CPU hotplug hang when a dying CPU's tasks sit in the BPF scheduler
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/printk/linux
Pull printk fix from Petr Mladek:
- Allow using Braille console with a serial console driver converted
to NBCON API
* tag 'printk-for-7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/printk/linux:
braille: nbcon: Allow to use a serial console with NBCON API as Braille console
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull timer fix from Ingo Molnar:
- Fix task work flags management regression in the hrtimer
rearming code that can leave task work items unprocessed
(Karl Mehltretter)
* tag 'timers-urgent-2026-10-04' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
hrtimer: Use the mask to clear TIF_HRTIMER_REARM from the exit work
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull perf events fixes from Ingo Molnar:
- Fix race between perf_event_exit_task() and perf_pending_task()
(Luo Gengkun)
- Fix perf header output management regressions (Ian Rogers)
- Require kernel access for text poke events (Zhengchuan Liang)
* tag 'perf-urgent-2026-10-04' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
perf: Require kernel access for text poke events
perf: Replace perf_event_header__init_id with full header init
perf: Fix race between perf_event_exit_task() and perf_pending_task()
|
|
finish_dispatch() uses the qseq embedded in p->scx.ops_state to tell
whether the QUEUED instance of a task it's about to claim is the one
scx_bpf_dsq_insert() saw. qseq is generated from rq->scx.ops_qseq, but
the counters of different rqs are independent, so if a task is dequeued
and re-enqueued on a different rq between scx_bpf_dsq_insert() and
finish_dispatch(), the new QUEUED instance can end up with the same
qseq as the old one:
CPU X CPU Z
----- -----
enqueue p on rq A, qseq = N
ops.dispatch()
scx_bpf_dsq_insert(p)
records qseq N
sched_setaffinity(p)
dequeue p from rq A
enqueue p on rq B, qseq = N
finish_dispatch(p, N)
qseq matches, p is claimed
The claim itself is still atomic so the core stays consistent, but an
insert issued for a previous QUEUED instance gets applied to a new one
which the BPF scheduler has just received through ops.enqueue(). This
breaks the guarantee that dispatches targeting a stale instance are
ignored.
Generate qseq from a per-task counter, p->scx.ops_qseq, instead so that
consecutive QUEUED instances of a task never share a qseq regardless of
which rq they're on. The counter is only updated in
scx_do_enqueue_task() with the task's rq locked, so no additional
synchronization is needed, and it fits in an existing hole in struct
sched_ext_entity on 64bit. Remove the now unused rq->scx.ops_qseq.
Never generate qseq 0. NONE and DISPATCHING don't carry a qseq, so
scx_bpf_dsq_insert() on a task in either state records 0. With a
per-task counter, every task's first QUEUED instance would otherwise get
qseq 0 and could be claimed by such an insert. Wrap the counter where
the QSEQ field wraps so that it can't reach a value that shifts to 0 on
32bit either.
Fixes: f0e1a0643a59 ("sched_ext: Implement BPF extensible scheduler class")
Cc: stable@vger.kernel.org # v6.12+
Assisted-by: Claude:claude-opus-5.5
Signed-off-by: Kuba Piecuch <jpiecuch@google.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/gregkh/tty
Pull tty/serial fixes from Greg KH:
"Here are some small tty/serial driver fixes for 7.3-rc6. Nothing major
here, just lots of small fixes for reported issues, some of them very
long-standing:
- tty hangup fixes that have been there since the BKL days and kept
tripping people up over time.
- vt selection bugfix
- other vt bugfixes (memory leaks and screen update fixes)
- n_gsm bugfix
- qcom-geni serial driver bugfix
- 8250 serial driver bugfixes
- other tiny serial driver fixes
All of these have been in linux-next, the last few only a few days but
testing here seems solid (this pull request was generated on that
tree)"
* tag 'tty-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/gregkh/tty: (37 commits)
tty: add missing driver flag kernel-doc colon
vt: selection: Fix unsigned underflow and slab-out-of-bounds read in paste_selection()
vt: skip screen update for DEC alignment test on backgroup consoles
vc_screen: reload vc pointer before if (ret) in vcs_write() to avoid UAF
serial: sc16is7xx: reduce TX refill rate with half-FIFO trigger
serial: sc16is7xx: refill TX FIFO below trigger using fresh TXLVL
serial: tegra: don't clear the Tx FIFO on an Rx-only reset
serial: sc16is7xx: fix TX gap caused by kfifo circular buffer wrap-around
tty: fix saved termios reset race
tty: serial: mpc52xx_uart: move static declarations up.
tty: serial: max3100: shut down timer before freeing port
tty: add break_wait kernel-doc
serial: qcom-geni: keep registered console runtime active
serial: qcom-geni: Fix unbalanced runtime PM resume for no_console_suspend
serial: qcom-geni: avoid unused-function warning
tty: serial: qcom_geni_serial: Keep console RX functional after deep idle
soc: qcom: geni-se: Correct QUP Core ICC vote constants
serial: 8250_bcm7271: fix use-after-free in brcmuart_remove()
serial: vt8500: Fix clock reference leak in vt8500_serial_probe()
kgdboc: Fix tty driver reference leak in configure_kgdboc()
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/gregkh/usb
Pull USB/Thunderbolt fixes from Greg KH:
"Here is a big set of USB and Thunderbolt driver fixes for 7.3-rc6.
They were delayed on my side due to conference travel, not the fault
of the submitters at all. Included in here are:
- lots of small thunderbolt fixes for reported issues due to more
testing and devices and a few reverts as well based on that work
- more usb-serial device ids added
- usb-serial and cdc-acm driver hangup and other fixes
- dwc3 driver fixes for reported problems
- lots of usb gadget driver fixes as people again fuzz these drivers
and send in fixes, which is nice to finally see
- typec driver fixes for reported problems
- octeon-hcd driver fixes for reported problems
- more usb-storage quirks added
- other small USB driver bugs resolved for reported problems
All of these have been in linux-next without any reported issues"
* tag 'usb-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/gregkh/usb: (63 commits)
usb: dwc3: gadget: fix IRQ storm on invalid event buffer count
Revert "usb: dwc3: gadget: fix IRQ storm on invalid event buffer count"
USB: gadget: dummy-hcd: Fix wait for outstanding request completions
usb: typec: port-mapper: Only match USB4 port if host interface is available
usb: cdns3: Fix NULL pointer dereference in cdns3_pci_probe
usb: dwc3: gadget: fix IRQ storm on invalid event buffer count
usb: typec: ucsi: Get the connector fwnode based on reg value
usb: gadget: f_uac1_legacy: validate bRequest index in generic_{set,get}_cmd
usb: core: clear both ep_in and ep_out for non-ep0 control endpoints
USB: cdc-acm: skip URB restart in port_shutdown if disconnected
usb: gadget: f_fs: Fix NULL pointer dereference in FUNCTIONFS_ENDPOINT_DESC
usb: gadget: aspeed-vhub: cancel wake work on device removal
thunderbolt: Disable CL states for the Anker Prime TB5 dock
thunderbolt: stream: Announce support for FMODE_NOWAIT
usb: typec: ucsi: displayport: Current CAM OOB index fixup
usb: ohci-st: disable controller wakeup on removal
usb: ohci-spear: disable controller wakeup on removal
usb: ohci-s3c2410: disable controller wakeup on removal
usb: ohci-da8xx: disable controller wakeup on cleanup
usb: cdns3: fix use-after-free in cdns3_gadget_exit()
...
|
|
Pull smb client fixes from Paulo Alcantara:
"Fix a series of data corruption and I/O error bugs found by running
generic/363 (fsx) in a loop against Windows Server 2022 and Samba.
- Stop data dirtied past EOF through an mmap from reappearing as file
content once the file is extended by a write, truncate, zero range,
copy range or clone range
- Flush dirty data and drain in-flight I/O before operations that
assume the pagecache and the server agree on the file: querying
allocated ranges, the O_TRUNC open, interior zero range, and
server-side copy/clone
- Stop a genuine size-extending zero range or preallocate from being
refused with -EOPNOTSUPP when the inode is not read caching, by
querying the server's authoritative EOF instead of trusting a stale
cached i_size
- Zero the untransferred tail of a short read, both in the netfs
read-gaps path (where stale folio content could otherwise be
written back to the server) and in the DIO/unbuffered read
collector, and tell a real EOF apart from a stale cached
remote_i_size after a lease downgrade
- Require stable pages on signed connections so a buffered write
can't modify a folio whose signature has already been computed and
is in flight, which the server rejected with STATUS_ACCESS_DENIED
and the client surfaced as -EIO
- Split several cifsFileInfo flags out of a shared bitfield byte so
concurrent updates taken under different locks no longer clobber
each other through a byte-level RMW"
* tag 'cifs-fixes-7.3-rc6' of https://git.manguebit.org/linux:
smb: client: split cifsFileInfo bitfields to avoid shared-byte RMW races
smb: client: require stable pages for signed connections
smb: client: distinguish real EOF from a stale remote_i_size on read
netfs: zero the tail of a short DIO/unbuffered read
smb: client: only require read lease for size-extending preallocate
netfs: zero gaps in read-gaps folio to avoid writing back stale data
smb: client: only require read lease for size-extending zero range
smb: client: drain and invalidate before server-side copy/clone
smb: client: flush dirty data before zeroing a range
smb: client: drain outstanding I/O before truncating on O_TRUNC open
smb: client: flush and commit data before querying allocated ranges
smb: client: discard post-EOF pagecache when extending a file via clone range
smb: client: discard post-EOF pagecache when extending a file via copy range
smb: client: discard post-EOF pagecache when extending a file via zero range
smb: client: clear post-EOF pagecache when extending a file via truncate
netfs: clear post-EOF pagecache when extending a file via write
|
|
Pull bpf fixes from Alexei Starovoitov:
- Fix overflow of backward jump offset in constant blinding
(Alexei Starovoitov)
- Fix packet range of packet pointers sharing an id when var_off
tightens umax of one pointer and not the other (Alexei Starovoitov)
- Fix objects stuck in free_by_rcu_ttrace list of bpf memalloc
(Alexei Starovoitov)
- Fix use-after-free of progs detached from busy trampolines: wait for
an RCU tasks grace period before freeing trampoline progs, and patch
detached progs out of trampoline images that are still in use
(Florent Revest)
- Hold map BTF for the memory allocator destructor record to fix UAF in
deferred bpf_mem_alloc destruction (Kumar Kartikeya Dwivedi)
- Fix missing migration protection in resizable hashtab
lookup_and_delete batch operation (Ömer Mete Kaya)
* tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf:
bpf: Fix missing migration protection in __rhtab_map_lookup_and_delete_batch()
selftests/bpf: Add a test for objects stuck in free_by_rcu_ttrace
bpf: Fix objects stuck in free_by_rcu_ttrace
bpf: Factor out __do_call_rcu_ttrace()
selftests/bpf: Test packet range of pointers sharing an id
bpf: Fix packet range of pointers sharing an id
selftests/bpf: Detach a trampoline prog while a task sleeps before it
bpf: Skip detached progs in trampoline images that are still in use
bpf: Wait for an RCU tasks grace period before freeing trampoline progs
bpf: Hold map BTF for the memory allocator destructor record
bpf: Fix overflow of jump offset in constant blinding
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound
Pull sound fixes from Takashi Iwai:
"A fair amount of small fixes, which became much larger as a pile of
pending homework during my vacation in the last weeks.
The majority of changes are device-specific quirks and ASoC updates,
along with a few ALSA core fixes and USB-audio hardening as well as a
few regression fixes.
ALSA Core:
- Serialize ALSA sequencer compat port-info ioctls
HD-audio:
- Fix ALC235 codec headset handling
- Fix regression on Tegra194 controller support
- Quirks / fixes for Lenovo, ASUS, Acer, Dell, Higole, HP, and IPASON
laptops
USB-audio:
- Hardening for line6 and ua101 drivers to deal with malformed
packets
- Fix input device leak on probe error in caiaq driver
- Apply quirks generically for Audient and Behringer devices
- Quirks for ASUS SupremeFX Hi-Fi, HiBy FC4 (native DSD), and MSI MAG
B850M MORTAR WIFI
ASoC:
- Runtime PM fixes for SDCA class function driver
- Fix uninitialized stream configs in RT1017 and MAX98363 SDW codecs
- AMD ACP SoundWire support for ASUS and Lenovo models
- Quirks for Intel machines (Lenovo Yoga Slim 7, Chuwi Hi8)
- Qualcomm LPASS fixes for playback constraints and WSA macro volume
- Codec / platform fixes for rt274/rt286, cs42l43, wm8903, Tegra ASRC
Misc:
- Regression fix for silent playback on Creative Titanium HD
- Error handling fix at probe via auto-cleanup in ice1712"
* tag 'sound-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound: (53 commits)
spi: cs42l43: Workaround for wrong speaker ID on Dell XPS 13 DX13260
ALSA: ice1712: Fix the error handling via auto-cleanup at probe
ALSA: core: Define auto-cleanup for snd_card_free()
ALSA: hda/realtek: Add ALC235 support for headset mode
ALSA: hda/realtek: Add micmute LED quirk for Acer Swift SFG16-72
ALSA: hda/realtek: Fix mute and micmute LEDs on HP ENVY x360 15-ey0xxx
ALSA: hda/realtek: Add mute LED quirk for HP Pavilion 14-dv1001TU
ALSA: ctxfi: Fix silent playback on Titanium HD (SB1270)
ALSA: usb-audio: Apply IGNORE_CTL_ERROR quirk to all Audient devices
ALSA: usb-audio: Apply boot quirk for Behringer models generically
ASoC: codecs: lpass-wsa-macro: rewrite the interpolator volume after enabling clocks
ALSA: hda/realtek: Fix speakers on ASUS PM3406CHA
ALSA: hda/realtek: Add quirk for HP Pavilion AiO 24 speaker
ALSA: hda/realtek: Add mute LED quirk for HP Laptop 15-fd2xxx
ASoC: amd: acp-config: Add ACP70 DMI quirk for ASUS UM5406GA
ALSA: hda/realtek: Add mute LED quirks for HP Pavilion 15-eh3174nw and HP Laptop 15-dw1xxx
ASoC: Intel: Add quirk to block match table for Lenovo Yoga Slim 7
ALSA: hda/realtek: Add mute LED quirk for HP Victus 15-fb2xxx
ALSA: hda/realtek: Add quirk for Acer Nitro AN515-45
ALSA: hda/realtek: Add quirk for Lenovo ThinkBook G8+
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/crng/random
Pull random number generator fixes from Jason Donenfeld:
- VMGENID memory needs to be mapped with the decrypted tag, so that
SEV-SNP machines can boot
- A fix for an initialization race in VMGENID, followed by a cleanup
- Trivial kernel doc cleanups in siphash and random.c
- A fix for a new compilation failure with recent clang on PPC and
RISC-V, due to generating an out-of-line memset in the vDSO
* tag 'random-7.3-rc6-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/crng/random:
random: vDSO: avoid call to memset() when zeroing reserved parameter
random: fix vgetrandom_opaque_params kernel-doc
random: vDSO: fix repeated word 'to' in comment
siphash: clean up kernel-doc comments
virt: vmgenid: move to using dev_set/get_drvdata
virt: vmgenid: set driver_data before registering notification handlers
virt: vmgenid: remap memory as decrypted
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/broonie/sound into for-linus
ASoC: Fixes for v7.3
A bigger collection of fixes than usual due to your vacation but nothing
hugely remarkable here, just fairly standard quirks and driver specific
bugfixes.
|
|
A recent change adding a new TTY driver flag left out a colon required
for well-formed kernel-doc.
Fixes: 8df07fe93573 ("tty: fix saved termios reset race")
Reported-by: Randy Dunlap <rdunlap@infradead.org>
Link: https://lore.kernel.org/ec3a09d6-4ba4-41b8-abb1-b59e265f4532@infradead.org
Signed-off-by: Johan Hovold <johan@kernel.org>
Link: https://patch.msgid.link/20261002072736.2063004-1-johan@kernel.org
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
The Braille console is not registered in console_list. Instead, it is
integrated with the virtual terminal (VT) and shows what is displayed
on the terminal. It writes the data using con->write*() callback
of the associated serial console driver.
braille_write() is called from the VT code under console_lock().
The associated serial console driver can be converted to the NBCON API
though. The situation is similar to flushing nbcon consoles in the legacy
loop when some boot consoles are still registered, see
nbcon_legacy_emit_next_record().
But there is a big difference though. braille_write() is not directly
called from the code paths flushing registered consoles. The VT code
expects that braille_write() succeeds. It does not replay the message
when it can't acquire the ownership. As a result, braille_write():
+ must try harder to get the ownership.
+ has to be synchronized only against non-printk serial console
which depend nbcon_device_try_acquire() using NBCON_PRIO_NORMAL.
Let's look at it from another side and try to simulate the original
locking using the NBCON API:
1. Take con->device_lock(), aka the port->lock in the legacy serial
console driver.
2. Acquire nbcon context to provide some synchronization for a panic()
context. Use NBCON_PRIO_NORMAL because it contends only with
nbcon_device_try_acquire() users. Do it in a busy loop. It should
always succeed when con->device_lock() succeeded because all other
users do the same. The only exception is when the context get acquired
by a CPU handling panic.
3. In panic, disable interrupts and try to acquire the nbcon context.
Use NBCON_PRIO_PANIC. And try even an unsafe takeover because
otherwise the Braille console won't see the text shown during panic().
It is similar to the oops_in_progress/trylock handling in the legacy
serial console driver.
Finally, avoid the newline prepending logic in the existing serial console
drivers when they are used as a Braille console. As explained above,
the Braille console shows the last modified line on the terminal (VT).
braille_write() is called when single characters are added. Most
messages are not ended by newline. Anyway, the VT code does not have
logic to reply partially printed messages.
Fixes: 13189fa73afa ("printk: nbcon: Rely on kthreads for normal operation")
Reviewed-by: John Ogness <john.ogness@linutronix.de>
[pmladek@suse.com: Remove white space changes in braille_register_console(). Typo fix.]
Link: https://patch.msgid.link/20261001140727.124398-2-pmladek@suse.com
Signed-off-by: Petr Mladek <pmladek@suse.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm
Pull power management fix from Rafael Wysocki:
"Restore the previous behavior on systems where the cpufreq pressure
was not visible in the scheduler and is not expected to be visible
there.
It became visible after a change made during the 7.2 development cycle
that had gone too far"
* tag 'pm-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm:
cpufreq: intel_pstate: Fix max_freq fallback in cpufreq_update_pressure()
|
|
Pull kvm fixes from Paolo Bonzini:
"The most intrusive change is reverting a commit from 7.3-rc1 that made
struct kvm a bit too large, and fixing the same issue otherwise.
There are again a lot of selftests lines; the sheer number of commits
is not small but I don't expect much more for 7.3 due to people
travelling to Plumbers next week.
ARM:
- Take a reference on the last IRQ loaded into an LR to prevent it
from being freed while running the guest (Marc Zyngier)
- Ensure that the ITS MOVALL command only affects LPIs that were
previously affined to the source redistributor (Marc Zyngier)
- Fix + test for honoring the host's trap configuration when running
non-protected VMs while KVM is in protected mode (Fuad Tabba)
- Use the host stage-1 mapping granularity for VM_PFNMAP mappings at
stage-2 (Mostafa Saleh)
x86:
Various bugfixes where the guest could do stupid things on purpose to
cause problems in the host:
- Failed VMRUNs can cause pending TLB flushes to be dropped, and in
general some actions done through VMCB control fields have to be
redone if VMRUN fails
- Toggling MSR interceptions or eVMCS execution controls can cause
the host to use a stale MSR permission bitmap
- Bad page tables can cause a WARN.
Also fix issues in last week's pull request (my fault, for changing
email workflow and thus missing feedback sent to kvm@ but not LKML).
Generic:
- Take kvm_lock when creating vCPUs. For almost two decades everybody
thought it was not done for some unspecified performance reasons,
but in reality it was only done because kvm_lock was originally a
spinlock.
This is a better fix than 97d65b544f48 ("KVM: Check for duplicate
vcpu_id as early as possible", from the 7.3 merge window), and does
not waste 2K per VM, hence its inclusion here"
* tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (29 commits)
KVM: arm64: Use stage-1 leaf size for VM_PFNMAP
KVM: arm64: selftests: Check a feature hidden in an ID register is UNDEF
KVM: arm64: Use the host's HCR_EL2 for non-protected VMs in pKVM
KVM: arm64: Clear HCR_EL2.RW for 32-bit non-protected vCPUs
KVM: arm64: Apply the fine-grained UNDEFs without FEAT_FGT
KVM: arm64: vgic-its: Fix MOVALL handling of source redistributor
KVM: arm64: vgic: Take a refcount on IRQs referenced by last_lr_irq
KVM: arm64: vgic: Allow last_lr_irq to be NULL when LRs are not overflowing
KVM: SEV: Do cache maintenance on the source VM *before* clearing SEV state
KVM: SEV: Nullify "have run CPUs" mask pointer when freeing it
KVM: selftests: Extend nested x2APIC test to validate using eVMCS for vmcs12
KVM: selftests: Extend nested x2APIC test to validate disabling x2APIC virt
KVM: selftests: Verify that L0's TPR doesn't get clobbered
KVM: selftests: Run the nested x2APIC with and without APICv being inhibited in L2
KVM: selftests: Add x2APIC MSR test for inhibiting APICv while nested
KVM: nVMX: Force MSR bitmap refresh if runtime eVMCS controls are modified
KVM: SVM: Use the active VMCB's MSR bitmap when checking if MSR is intercepted
KVM: SVM: Sync guest's PERF_CNTR_GLOBAL_CTL from h/w only on successful VMRUN
KVM: SVM: Don't mark ASID fields as dirty when setting control.tlb_ctl
KVM: SVM: Update control fields on #VMEXIT if and only if VMRUN succeeded
...
|
|
For the errors at the probe time, we'd need to call rather
snd_card_free() instead of the snd_card_unref() -- the former calls
explicitly snd_card_disconnect() that cleans up the registered
devices, etc, while the latter may leak in corner cases.
For convenience, define a new auto-clean with snd_card_free.
Link: https://patch.msgid.link/20261001151039.592033-1-tiwai@suse.de
Signed-off-by: Takashi Iwai <tiwai@suse.de>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Paolo Abeni:
"Including fixes from Bluetooth, WiFi and netfilter.
We are actively retargeting several non-urgent fixes towards next,
but the traffic on the ML looks ever-increasing, and propagating the
push-back towards subsystems is not immediate.
No known outstanding regressions.
Current release - regressions:
- netfilter: nft_set_rbtree: skip transaction elements during GC
Previous releases - regressions:
- sched: cls_api: reclaim an empty proto on the error path
- core:
- fix checksum offsets in skb_splice_from_iter()
- cap skb->queue_mapping when the tx queue is picked
- page_pool: fix use-after-free in page_pool_recycle_ring_bulk()
- wifi:
- mac80211: fix slab-out-of-bounds read in ieee80211_monitor_select_queue()
- mac80211: drop oversized fragments to avoid extra_len overflow
- netfilter:
- flowtable: restore ieee80211 forward path
- bluetooth: hci_conn: Lock parent access during enhanced SCO setup
- eth:
- bcmgenet: allocate RX buffers as page fragments
- stmmac: fix rx Scatter-Gather support
- octeontx2-pf: fix aura BPID assignment when CONFIG_DCB is enabled
- gve: DQO: accept TSO packets with non-protocol gso_type bits
- r8169: disable EEE on RTL8168h/8111h
Previous releases - always broken:
- tcp: refresh TS.Recent for accepted old ACKs
- wifi:
- ath11k: reset ar->num_stations on hardware start
- cfg80211: fix RTS threshold setting for single-radio PHY
- bluetooth: btintel_pcie: fix plen overflow in btintel_pcie_recv_frame()
- eth: bcmgenet: fix NULL dereference in set_coalesce before first open
Misc:
- Eric is retiring from google and updating his contact info"
* tag 'net-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (96 commits)
net: phy: aquantia: fix system interface type not updated in forced mode
net: usb: qmi_wwan: add Rolling Wireless RN947R
net: mvneta: clear XDP pfmemalloc flag between frames
ipv6: sr: use skb_get_hash_net() in seg6_make_flowlabel()
net/mlx5e: Fix AF_XDP TX timestamp teardown NULL dereference
r8169: disable EEE on RTL8168h/8111h
octeontx2-pf: Fix RSS indirection table size
sctp: check RCV_SHUTDOWN after the sendmsg connect wait
net: sparx5: make ports inherit the switch base mac address type
net: microchip: vcap: stop scanning after deleting key field
netfilter: flowtable: restore ieee80211 forward path
netfilter: flowtable: generalize pending status bit
netfilter: bpf: reject invalid NAT manipulation types
netfilter: nft_set_rbtree: skip transaction elements during GC
ipvs: filter some flags received in the backup server
ipvs: do not create invisible templates
ipvs: bound LBLCR and LBLC cache growth
ipvs: fix missing counter decrement in lblc
netfilter: nft_flow_offload: drop flowtable reference on init error path
selftests: net: check timestamp echo after an old ACK
...
|
|
perf_iterate_sb() invokes its callback for each matching perf_event on
the CPU and task context, passing a shared caller-allocated event
structure.
perf_event_header__init_id() mutated header->size in place by adding
event->id_header_size, requiring sideband output callbacks to save and
restore header fields across iterations. Three sideband callbacks failed
to save and restore header.size around perf_event_header__init_id():
- perf_event_ksymbol_output()
- perf_event_bpf_output()
- perf_event_text_poke_output()
When multiple events with attr.ksymbol, attr.bpf_event, or
attr.text_poke and sample_id_all are active on the same CPU, each
subsequent event receives a record whose header.size is inflated by all
preceding events' id_header_size values while only a single id_sample is
written, leaving uninitialized ring-buffer bytes at the end of the
record and causing userspace perf to fail with -EFAULT ("Bad address")
when parsing the sample_id trailer.
Similarly, perf_event_mmap_output() set PERF_RECORD_MISC_MMAP_BUILD_ID
in mmap_event->event_id.header.misc when event->attr.build_id was
enabled, but only saved and restored header.size and header.type. If an
event with attr.build_id was followed by an event with attr.mmap2 and
!attr.build_id, the second event received PERF_RECORD_MISC_MMAP_BUILD_ID
in header.misc while its payload contained maj/min/ino/ino_generation
instead of a build ID.
Rather than splitting header initialization between callers and output
callbacks and saving/restoring mutated header fields, replace
perf_event_header__init_id() with perf_event_header__init(), which
initializes header->type, header->misc, and header->size alongside the
sample_id fields on each invocation.
Fixes: 76193a94522f ("perf, bpf: Introduce PERF_RECORD_KSYMBOL")
Fixes: 6ee52e2a3fe4 ("perf, bpf: Introduce PERF_RECORD_BPF_EVENT")
Fixes: e17d43b93e54 ("perf: Add perf text poke event")
Fixes: 88a16a130933 ("perf: Add build id data in mmap2 event")
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://patch.msgid.link/20260929222332.973435-1-irogers@google.com
Cc: stable@vger.kernel.org
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf
Pablo Neira Ayuso says:
====================
Netfilter/IPVS fixes for net
The following batch contains Netfilter fixes for net. This batch
fixes crashes as recent feature regression, one of the due to a
dependency that has been pulled into -stable:
1) Expand existing ipset fix for bitmap sets to disallow comments
updates from kernel-side adds, from Florian Westphal.
2) Drop flowtable reference if nf_ct_netns_get() fails, otherwise
flowtable cannot ever be removed, from Aohan Mei.
3) nft_rbtree GC should collect end elements that contained in
this transaction batch, new or deleted elements are never
expired. From Weiming Shi.
4) Restrict nf_nat_bpf so it does not set unknown NF_NAT_MANIP_*
values, from Fernando F. Mancera.
5) Flowtable GC must skip flows that are pending hardware updates,
generalize the PENDING flag and use it to inhibit GC.
6) Restore flowtable with ieee80211 which broke due to a relatively
recent commit, which was pulled in by -stable, causing a regression
in 6.18 kernels.
And the following IPVS fixes:
1) Fix accounting of cache entries in IPVS LBLC for destinations,
which eventually fills up the table and trigger recurrent
resizing, from Julian Anastasov.
2) Limit IPVS cache growth for LBLCR and LBLC schedulers,
from Zhiling Zou.
3) Restrict IP_VS_CONN_F_ONE_PACKET for normal connections,
do not allow to use it with templates. Also from Julian.
4) Sanitize flags in IPVS sync messages received in the backup.
From Julian Anastasov.
netfilter pull request 26-09-30
* tag 'nf-26-09-30' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf:
netfilter: flowtable: restore ieee80211 forward path
netfilter: flowtable: generalize pending status bit
netfilter: bpf: reject invalid NAT manipulation types
netfilter: nft_set_rbtree: skip transaction elements during GC
ipvs: filter some flags received in the backup server
ipvs: do not create invisible templates
ipvs: bound LBLCR and LBLC cache growth
ipvs: fix missing counter decrement in lblc
netfilter: nft_flow_offload: drop flowtable reference on init error path
netfilter: ipset: do not update comments from kernel-side adds
====================
Link: https://patch.msgid.link/20260930074142.298353-1-pablo@netfilter.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
Resetting saved termios state on device registration is needed where a
minor number can be reused for an entirely different device and where
the old settings may prevent the port from even being opened (e.g. when
CLOCAL is not set).
Not all TTY drivers guarantee that the minor number is no longer in use
when registering devices however, something which can lead to a
use-after-free when closing a TTY (and saving its termios) races with
re-registration.
Add a new TTY_DRIVER_RESET_SAVED_TERMIOS flag to request that any saved
termios state is reset on registration and only set it for drivers that
make sure that the minor number is no longer in use.
Fixes: 93857edd9829 ("tty: reset termios state on device registration")
Reported-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Link: https://lore.kernel.org/20260926184154.3017929-1-nicoyip.dev@gmail.com
Cc: stable@kernel.org # 4.12
Signed-off-by: Johan Hovold <johan@kernel.org>
Link: https://patch.msgid.link/20260930131845.1809256-1-johan@kernel.org
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
mvneta_swbm_add_rx_fragment() sets XDP_FLAGS_FRAGS_PF_MEMALLOC on the
xdp_buff when a fragment page is a pfmemalloc one (page under memory
pressure). The xdp_buff is reused for the next frame, but only the
XDP_FLAGS_HAS_FRAGS bit was cleared at frame start, so the pfmemalloc
bit leaked from one frame into the following ones. mvneta_swbm_build_skb()
propagates the flag to skb->pfmemalloc through xdp_update_skb_frags_info(),
so the skb of a subsequent fragmented frame could be wrongly marked as
pfmemalloc even if none of its pages are under pressure.
Clear all the xdp_buff flags in mvneta_swbm_rx_frame(), which is invoked
for each new frame, instead of just the XDP_FLAGS_HAS_FRAGS bit.
Fixes: ed7a58cb40bd ("net: marvell: rely on xdp_update_skb_shared_info utility routine")
Reviewed-by: Simon Horman <horms@kernel.org>
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Link: https://patch.msgid.link/20260929-mvneta-xdp-clear-frag-fix-v4-1-1e63b25eeed8@oss.qualcomm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Fix netfs to erase the contents of a hole created after the EOF by an
ordinary write if dirty data has been previously left there by writes
through an mmapped region. Neither the buffered nor the unbuffered/DIO
write path clears that stale pagecache.
Zero the tail of the folio straddling the EOF before an extending
write. That is the only folio that can hold data written past the EOF
through an mmap, as pages wholly beyond the EOF can't be faulted in.
The folio is zeroed rather than dropped so a concurrent extending write
can't lose data.
Both write paths downgrade the i_rwsem to shared, so extending writes
can run concurrently and the i_size read by the caller may be stale by
the time the folio is locked. Re-read i_size under the folio lock and
clamp the zeroed range up to it, so a racing write that already put
data into the folio isn't clobbered.
Wait for any writeback on the folio to finish before zeroing it so that
the pagecache isn't modified while it may still be read by the transport
during transmission. Honour IOCB_NOWAIT by returning -EAGAIN rather
than blocking on the folio lock, on writeback, or in folio_mkclean()'s
rmap walk when the folio is mapped.
truncate_pagecache() can't be used here: it must be called with the
i_rwsem held exclusively, but these write paths only hold it shared,
and it would block unconditionally, breaking IOCB_NOWAIT.
Callers that hold i_rwsem exclusively for the whole resize (truncate,
setattr, fallocate, clone) exclude any genuine concurrent buffered
writer, so staleness can instead be decided from the folio's dirty
state, as pagecache_isize_extended() already does for filesystems that
serialise writes against truncate/setattr via a single i_rwsem.
Export netfs_clear_stale_post_isize() helper to handle such case.
The helper is required by the CIFS client to fix generic/363.
Closes: https://sashiko.dev/#/patchset/20260921230755.1133425-1-pc%40manguebit.org
Fixes: 938e13a73b24 ("netfs: Implement buffered write API")
Fixes: 153a9961b551 ("netfs: Implement unbuffered/DIO write support")
Reviewed-by: David Howells <dhowells@redhat.com>
Reviewed-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
|
|
Before commit 871df5007eda ("netfilter: flowtable: bail out if forward
path cannot be discovered"), there was a fallback to set up a forward
path in case .ndo_fill_forward_path fails or DEV_PATH_MTK_WDMA was used.
Such fallback was used by commit d787a3e38f01 ("mac80211: add support
for .ndo_fill_forward_path").
One possibility is to handle DEV_PATH_MTK_WDMA from the flowtable
forward path discovery. However, this is only used internally by drivers
to retrieve mtk_wdma information to set up hardware offload. Felix
decided to use the .fill_forward_path interface for this purpose due to
the lack of a better interface at that time.
Add a new DEV_PATH_IEEE80211 path which is offered if the new ieee80211
flag is set on in the struct net_device_path_ctx to restore the
flowtable with a ieee80211 netdevice. Handle this new DEV_PATH_IEEE80211
path just like DEV_PATH_ETHERNET and DEV_PATH_DSA, ie. this is the last
netdevice in the stack.
This new ieee80211 flag is implicitly unset for mtk_ppe and airoha which
call dev_fill_forward_path() to retrieve a DEV_PATH_MTK_WDMA path.
Fixes: 871df5007eda ("netfilter: flowtable: bail out if forward path cannot be discovered")
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
Rename NF_FLOW_HW_PENDING to NF_FLOW_PENDING and use it to inhibit the
flowtable GC worker until pending hw offload work has been completed.
Apparently, nf_flow_offload_stats() can schedule work to retrieve stats
while the flow is being removed by GC.
And this bit can also be used in a follow up patch to disable GC until
the flow has been fully added in both directions.
Revert the reordering done in commit d644b23afe1e ("netfilter:
flowtable: publish HW_DEAD after worker is done") to prevent a race
between GC and hw offload handler.
Fixes: 2c8897953f3b ("netfilter: flowtable: Add pending bit for offload work")
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
skbedit can set skb->queue_mapping and raise the per-CPU skip_txqueue
flag so __dev_queue_xmit() honours the mapping. __dev_queue_xmit()
cleared the flag before sch_handle_egress() and only read it afterwards,
so the flag was not confined to the xmit that set it: a nested xmit
(mirred redirect or mirror, or a drop after skbedit) could set the flag
and the outer xmit would consume it for an skb that never went through
skbedit.
A forwarded packet still carries the ingress NIC's rx_queue + 1 in
skb->queue_mapping, so the outer device then indexes its tx queue state
with that stale value. Taprio's child array q->qdiscs[] is sized to the
device's queue count, so taprio_enqueue() indexes past its allocation
and dereferences the result as a struct Qdisc *.
We (ab)use the skb->nf_skip_egress which means "skip netfilter egress
for this packet" to tag to "am I in tc egress?". Despite the overload
I dont see it as a conflict since the marker is set only around the
single sch_handle_egress() call and ingress path is guarded by
tc_at_ingress.
I will send a followup(net-next) patch once this hits net-next to
rename the skb->nf_skip_egress bit/flag to skb->skip_egress
Arm the flag only from the egress classifier that can use it: raise
skip_txqueue from tcf_skbedit_act() only when it runs inside
sch_handle_egress(), thanks to skb->nf_skip_egress. An egress qdisc
classifier runs in q->enqueue(), after the tx queue has been picked,
so a mapping it sets cannot affect the current packet; arming the flag
there only pollutes it for a later xmit. Then own the flag for the xmit
frame the egress hook runs in: save the incoming value and clear it just
before sch_handle_egress(), and restore it after the hook - on the
consumed (drop) path, or, in the same call that reads it, on the
surviving path. The save and the restores stay inside the
egress_needed_key static branch, so a packet pays for them only when
egress hooks are active (2f1e85b1aee4).
Store the value netdev_cap_txqueue() selected back into skb->queue_mapping
in netdev_tx_queue_mapping(), as netdev_core_pick_tx() already does, so
the skip_txqueue path never hands a later reader on the xmit path a
mapping the device cannot serve. A store made still later in the same
frame, by a tc BPF program attached to a transmit qdisc, is outside this
path and is not re-capped; a separate followup will resolve that path.
netdev_xmit_skip_txqueue() returns the previous flag value so the
save-and-clear is one call, and a no-op stub is provided when
CONFIG_NET_EGRESS is disabled. skb->nf_skip_egress is compiled under
CONFIG_NET_EGRESS rather than CONFIG_NETFILTER_SKIP_EGRESS, so
skb_at_tc_egress() is valid whenever the egress path is built.
A local user in a network namespace can redirect a packet from a device
with more TX queues to one with fewer after setting a mapping valid only
on the larger device. That reaches these reads and, under KASAN, faults
with "slab-out-of-bounds in taprio_enqueue".
Conditions to recreate the bug: the report's own trigger is a local user
with CAP_NET_ADMIN in a network namespace, so no eBPF program is needed.
With CONFIG_NET_SCH_TAPRIO=y, CONFIG_NET_ACT_SKBEDIT=y,
CONFIG_NET_ACT_MIRRED=y, CONFIG_NET_CLS_MATCHALL=y,
CONFIG_NET_SCH_PRIO=y and KASAN enabled, create qa (3 queues), qb
(2 queues) and qc (1 queue) as dummy devices; put a taprio root on qb
(num_tc 1, queues 2@0) and clsact on all three; then add an egress
matchall filter on every device. On qa: "action skbedit queue_mapping 2
pipe action mirred egress redirect dev qb". On qb: "action mirred egress
mirror dev qc". On qc: "action skbedit queue_mapping 0 pipe". Send one
packet out qa. qc's skbedit sets the flag while qb's outer xmit is in
flight; without the fix qb consumes it and reads its two-entry taprio
child array with the forwarded packet's stale mapping. A qc whose
skbedit is instead installed in a transmit-qdisc classifier (a matchall
filter on the qc root qdisc) reaches the same code path the same way
without the fix.
Testing: on a KASAN build with panic_on_warn=1 the unfixed kernel panics
with "BUG: KASAN: slab-out-of-bounds in taprio_enqueue", a read 0 bytes
past a 16-byte taprio_init() allocation, for the clsact-setter and the
transmit-qdisc-classifier reproducers and for a clsact skbedit-then-tc-BPF
store; the fixed kernel runs all three with no report, and the BPF store
variant additionally shows the expected "selects TX queue" clamp notice
from the write-back.
Fixes: 2f1e85b1aee4 ("net: sched: use queue_mapping to pick tx queue")
Reported-by: Zero Day Initiative <zdi-disclosures@trendmicro.com>
Link: https://lore.kernel.org/netdev/CANn89iLwYx8nCVf0pCEk_MmEiyC6kQaMwCQT9WkQVeeNzNQHqQ@mail.gmail.com/
Link: https://lore.kernel.org/netdev/179008581937.2160803.7117814290574262942@kernel.org/
Link: https://lore.kernel.org/netdev/179033713973.2160803.4914570693994398206@kernel.org/
Link: https://lore.kernel.org/netdev/20260925180407.63647514@kernel.org/
Link: https://lore.kernel.org/netdev/CANn89i+k-mZKDQVtvws_MEXeuMTAdaCcOXFZE-RfhcGTu90sjA@mail.gmail.com/
Suggested-by: Eric Dumazet <edumazet@google.com>
Suggested-by: Jakub Kicinski <kuba@kernel.org>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/QDISC-9R8V.v4.20260928081529@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
hrtimer_rearm_deferred_user_irq() removes the rearm bit from the local
copy of the TIF work with:
*tif_work &= ~TIF_HRTIMER_REARM;
TIF_HRTIMER_REARM is the bit number (12), not the mask. That clears
TIF_NOTIFY_SIGNAL (bit 2) and TIF_MEMDIE (bit 3) from the copy and
leaves bit 12 set.
So the function never returns true, and the copy handed to
exit_to_user_mode_loop() lacks TIF_NOTIFY_SIGNAL. If that was the only
work for the loop, the task returns to user space with the task work
still pending. hrtimer_interrupt() sets TIF_HRTIMER_REARM every time,
so the next tick repeats that. Task work queued with TWA_SIGNAL from a
hrtimer callback stays pending until the task does a syscall, takes an
interrupt without deferred rearm or has to reschedule.
A task which polls the io_uring completion ring in user space sees a
timeout completion after 30ms (median) instead of 70us. 2 CPU QEMU
guest, HZ=1000.
Use the mask.
Fixes: 15dd3a948855 ("hrtimer: Push reprogramming timers into the interrupt return path")
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Assisted-by: LLM
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260919144027.70250-1-kmehltretter@gmail.com
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/mtd/linux
Pull MTD fixes from Miquel Raynal:
"The most important set of fixes are around the handling of the QE bit
in SPI NAND.
There are also a couple of behavioral fixes (mutex issue in SPI-NOR,
spurious bitflips on vf610_nfc, OOB bytes count in SPI NAND and
cfi_cmdset stack usage).
The rest is mostly AI fuzzing results"
* tag 'mtd/fixes-for-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/mtd/linux:
mtd: spinand: Do not update the QE bit on devices without one
mtd: spi-nor: core: Fix mutex leak in spi_nor_rww_start_exclusive()
mtd: rawnand: cadence: Initialize IRQ state before requesting IRQ
mtd: rawnand: vf610_nfc: fix false bitflips on reads of erased pages
mtd: rawnand: vf610_nfc: fix reads on chips with more than 64 bytes of OOB
mtd: spinand: fix zero oobavail when no ECC engine is used
mtd: spinand: fix NULL pointer dereference with no ECC engine
mtd: mtd_intel_dg: reset poll counter for each erase
mtd: cfi_cmdset_0001: shrink do_write_buffer() stack frame
mtd: core: call _get_device() with the master MTD
mtd: core: avoid double-free of OTP NVMEM device
mtd: spinand: Enable QE on all dies
mtd: block2mtd: Fix divide error when erase_size is zero
|
|
After commit d2d5c129d07e ("cpufreq: Make cpufreq_update_pressure() fall
back to cpuinfo.max_freq"), cpufreq pressure appears in the CPU load
balancer unexpectedly in some cases in which it was not present before,
leading to confusion and uncertainty.
Clearly, the scheduler assumes that cpufreq pressure will not be set
unless the capacity reference frequency of the CPU is known, and the
commit mentioned above violates that assumption.
However, in some cases the capacity reference frequency of the CPU is
in fact known even though arch_scale_freq_ref() returns 0 and in those
cases it should be possible to set cpufreq pressure as appropriate.
For this purpose, introduce a new cpufreq driver callback returning
the CPU capacity reference frequency, .scale_freq_ref(), and make
cpufreq_update_pressure() invoke it, if present, instead of falling
back to cpuinfo.max_freq unconditionally.
Add that callback to the intel_pstate driver and make it return 0 unless
the scale-invariant capacity of the given CPU has been explicitly set,
in which cases its reference frequency is always cpuinfo.max_freq.
Fixes: d2d5c129d07e ("cpufreq: Make cpufreq_update_pressure() fall back to cpuinfo.max_freq")
Reported-by: Jianyong Wu <wujianyong@hygon.cn>
Closes: https://lore.kernel.org/linux-pm/20260915065747.1671965-1-wujianyong@hygon.cn/
Tested-by: Jianyong Wu <wujianyong@hygon.cn>
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
Tested-by: Ricardo Neri <ricardo.neri-calderon@linux.intel.com> # Intel hybrid parts
Tested-by: Chen Yu <yu.c.chen@intel.com>
[ rjw: Add READ_ONCE() around a capacity_perf read ]
Link: https://patch.msgid.link/12975163.O9o76ZdvQC@rafael.j.wysocki
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
|
|
IP-in-IP GSO can re-enter inet_gso_segment() or ipv6_gso_segment()
for each nested IP header. encap_level tracks header bytes, not callback
depth, so a deep chain can exhaust the kernel stack. Making
inet_gso_segment() stackable introduced unbounded IPv4 nesting; IPIP
GSO/TSO later made the path reachable. The IPv6 stackable path was
introduced separately and uses the same guard.
Count IPv4 and IPv6 GSO handler entries in skb_gso_cb, initialized once
per top-level GSO operation and preserved across GRE/UDP context changes.
Use the existing IP_TUNNEL_RECURSION_LIMIT for both handlers. The first
five entries pass, and the sixth returns -EINVAL before dispatching
another GSO callback.
Fixes: 3347c9602955 ("ipv4: gso: make inet_gso_segment() stackable")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Closes: https://lore.kernel.org/all/cover.1790157745.git.zihanx@nebusec.ai/
Assisted-by: LLM
Co-developed-by: Luxing Yin <root@tr0jan.top>
Signed-off-by: Luxing Yin <root@tr0jan.top>
Signed-off-by: Zihan Xi <zihanx@nebusec.ai>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260924051521.32568-2-zihanx@nebusec.ai
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
bpf_tramp_image_put() makes sure a trampoline image is not freed while
a task may still be running in it, but nothing similar is done for the
progs called by that image. Detach drops the last prog reference right
away and the prog is freed after grace periods, on the basis that a
task still in the traced function skips the fexit progs once the nop at
ip_after_call is patched to a jump.
That leaves out a task sleeping in a sleepable prog that runs before
the detached one, which no grace period waits for:
CPU 0 CPU 1
in image I, sleeping in prog S
detach P from I's trampoline
-> new image, bpf_tramp_image_put(I)
bpf_prog_put(P), last ref
grace periods, P freed
back from S
__bpf_prog_enter(P)
call P->bpf_func
If S and P are fexit progs the task is already past the patched jump,
and fentry only images don't have one. On x86 this is an int3 in
poisoned bpf_prog_pack memory:
Oops: int3: 0000 [#1] SMP NOPTI
CPU: 18 UID: 0 PID: 94573 Comm: x169 Not tainted 6.18.44 #1 PREEMPT(lazy)
RIP: 0010:0xffffffffc0601d8d
Call Trace:
<TASK>
? bpf_trampoline_6442515411+0x1a4/0x21b
bpf_lsm_bprm_committed_creds+0x5/0x10
security_bprm_committed_creds+0x5f/0x70
begin_new_exec+0x2d6/0x410
...
We hit this in production when progs attached through trampolines got
detached while their hooks were busy, and it was independently found
with a fuzzer and KASAN.
Have the JITs emit a patchable nop in front of each prog call sequence
and record it in the image. When a prog is detached, patch its nop to a
jump over the call sequence. Progs that stay attached keep running for
the tasks that are in the image, and ip_after_call isn't needed anymore.
The task can be in any image that isn't freed yet, not only in the
current one. It sleeps in image I1 that calls S, P and Q, then P is
detached and the trampoline moves to image I2, then Q is detached and
its call is still in I1. So the trampoline keeps a list of its images
until they are freed, and detaching a prog patches its nop in all of
them. Images hold a reference on the trampoline for that long.
On riscv and loongarch a jump of any range takes several instructions,
and a task preempted in the middle of them could resume into half of the
new sequence. The nop is a single instruction there, patched to a near
branch through arch_bpf_trampoline_skip().
With the extra nops, BPF_MAX_TRAMP_LINKS progs no longer fit in a page
on x86 and arm64 (and already didn't on powerpc), so lower the limit
there like s390 does.
Fixes: e21aa341785c ("bpf: Fix fexit trampoline.")
Reported-by: Sechang Lim <rhkrqnwk98@gmail.com>
Suggested-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Florent Revest (Anthropic) <florent.revest@linux.dev>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260926135605.1217928-3-florent.revest@linux.dev
Closes: https://lore.kernel.org/bpf/20260815071927.147049-1-zirajs7@gmail.com/
|
|
When a prog is detached from a trampoline, it is freed after an RCU
grace period, or an RCU tasks trace one if it is sleepable. This covers
the tasks that are running the prog, since the prog's enter helper takes
the matching read lock before the prog is called. On a preemptible
kernel, it doesn't cover a task that was preempted in the trampoline
just before the enter helper. That task holds no lock yet, and it calls
the prog after it was freed:
BUG: KASAN: vmalloc-out-of-bounds in __bpf_prog_enter_recur+0x3a5/0x3f0
Read of size 8 at addr ffffc90000055040 by task candidate/110
CPU: 1 UID: 0 PID: 110 Comm: candidate Not tainted 7.3.0-rc2-00014-g15071f2a1263-dirty #2 PREEMPT(full)
Call Trace:
<TASK>
__bpf_prog_enter_recur+0x3a5/0x3f0
bpf_trampoline_6442509193+0x37/0xf1
__x64_sys_futex+0x9/0x410
do_syscall_64+0xb0/0x530
...
Wait for an RCU tasks grace period before the existing one when freeing
a prog that was linked to a trampoline. An RCU tasks grace period only
ends once the tasks that were preempted have run again, and
bpf_tramp_image_put() already relies on it to free the image. It doesn't
wait for tasks that sleep in the trampoline, the next commit takes care
of those.
Only progs that were linked to a trampoline can be called this way, so
bpf_trampoline_add_prog() marks them and other progs are still freed as
before.
Fixes: e21aa341785c ("bpf: Fix fexit trampoline.")
Reported-by: Junseo Lim <zirajs7@gmail.com>
Signed-off-by: Florent Revest (Anthropic) <florent.revest@linux.dev>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260926135605.1217928-2-florent.revest@linux.dev
Closes: https://lore.kernel.org/bpf/aqdrwVpanH3WGurX@omen-arch/
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
Pull MM fixes from Andrew Morton:
- Fix module loading incorrectly returning success after alloc_tag
codetag setup failed
- Restore MADV_COLLAPSE semantics for shmem so forced collapse ignores
the shmem THP/mTHP sysfs settings
- Make alloc_tag UAPI structure padding explicit
- Fix DAMON schemes unexpectedly stopping after quotas are disabled
through the online parameter update interface
- Fix a false SW_TAGS KASAN invalid-access report when freeing vmapped
task stacks
- Split the MEMORY MANAGEMENT - MEMORY POLICY AND MIGRATION MAINTAINERS
entry into separate MIGRATION and NUMA PLACEMENT entries
- Move memory tiering maintenance under NUMA PLACEMENT
- Add Gregory Price as a NUMA PLACEMENT co-maintainer
- Add Heming Zhao as an ocfs2 reviewer
- Fix mmap_prepare() state being copied onto a merged VMA rather than
only onto a newly allocated VMA
* tag 'mm-hotfixes-stable-2026-09-27-19-12' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm:
mm/vma: predicate setting mmap_prepare VMA fields on new vma alloc
MAINTAINERS: add Heming Zhao as ocfs2 reviewer
MAINTAINERS: make Gregory a co-maintainer of MEMORY MANAGEMENT - NUMA PLACEMENT
MAINTAINERS: move memory tiering under MEMORY MANAGEMENT - NUMA PLACEMENT
MAINTAINERS: split up MEMORY MANAGEMENT - MEMORY POLICY AND MIGRATION
kasan: unpoison task stack below watermark only in generic mode
mm/damon/core: don't skip damos_adjust_quota() while esz is not zero
alloc_tag: avoid implicit padding in uapi
mm: shmem: ignore sysfs configs for shmem forced collapse
module: fix lost error code from codetag_load_module()
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth
Luiz Augusto von Dentz says:
====================
bluetooth pull request for net:
Core:
- hci_core: Serialize ACL scheduling with channel deletion
- hci_core: Serialize SCO and ISO scheduling with teardown
- hci_core: Serialize fragmented ISO packet queueing
- hci_core: Fix inquiry cache timestamps on 64-bit systems
- hci_core: free the HCI ID if naming fails
- hci_conn: Lock parent access during enhanced SCO setup
- hci_sync: Fix inquiry cache use-after-free
- hci_sync: don't drain cmd_sync backlog on unregister
- RFCOMM: Fix initial port reference race
- SMP: Serialize SMP remote OOB data access
Drivers:
- btintel: fix buffer over-read in btintel_hw_error()
- btintel: validate DDC record lengths
- btintel_pcie: fix plen overflow in btintel_pcie_recv_frame()
- btintel_pcie: reject oversized TX packets in send_frame()
* tag 'for-net-2026-09-28' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth:
Bluetooth: hci_core: Serialize SCO and ISO scheduling with teardown
Bluetooth: hci_core: Serialize ACL scheduling with channel deletion
Bluetooth: Serialize SMP remote OOB data access
Bluetooth: RFCOMM: Fix initial port reference race
Bluetooth: hci_sync: Fix inquiry cache use-after-free
Bluetooth: hci_sync: don't drain cmd_sync backlog on unregister
Bluetooth: hci_core: Serialize fragmented ISO packet queueing
Bluetooth: hci_conn: Lock parent access during enhanced SCO setup
Bluetooth: btintel: validate DDC record lengths
Bluetooth: hci_core: free the HCI ID if naming fails
Bluetooth: hci_core: Fix inquiry cache timestamps on 64-bit systems
Bluetooth: btintel_pcie: reject oversized TX packets in send_frame()
Bluetooth: btintel_pcie: fix plen overflow in btintel_pcie_recv_frame()
Bluetooth: btintel: fix buffer over-read in btintel_hw_error()
====================
Link: https://patch.msgid.link/20260928152919.942973-1-luiz.dentz@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
build_pairing_cmd() looks up remote OOB data and copies its contents
without holding hdev->lock, which serializes the list's writers. After
SMP finds an entry, a concurrent management Remove Remote OOB Data
command can unlink and free it before SMP reads its present flag or
copies its random and confirmation values. Removal can also invalidate
an entry while the lookup is still traversing the list.
KASAN reported:
BUG: KASAN: slab-use-after-free in build_pairing_cmd+0x948/0x9b0
Call Trace:
build_pairing_cmd+0x948/0x9b0
smp_recv_cb+0x459f/0x8110
l2cap_recv_frame+0xf14/0x9190
l2cap_recv_acldata+0xa64/0xd40
hci_rx_work+0x4ca/0x730
Allocated by task 87:
hci_add_remote_oob_data+0x11d/0x530
add_remote_oob_data+0x282/0x400
hci_sock_sendmsg+0x1033/0x1ea0
Freed by task 93:
hci_remote_oob_data_clear+0x108/0x1c0
remove_remote_oob_data+0x198/0x220
hci_sock_sendmsg+0x1033/0x1ea0
Taking hdev->lock in build_pairing_cmd() would recurse for callers that
already hold it and invert the device-to-L2CAP lock order on the receive
path. Add a per-device remote_oob_lock instead, held across the SMP
lookup and copies and by the add, remove and clear helpers. Cover
initialization and in-place updates as well, so SMP cannot read partially
initialized or updated OOB values. Release the mutex on allocation
failure, preserving the existing error return.
The new critical sections acquire no device, connection or channel locks.
Writers retain their existing hdev->lock protection, which continues to
serialize the other readers without changing their locking or behavior.
Link: https://lore.kernel.org/r/00660cd3-7d71-13a4-f617-229e6defb701@gmail.com
Fixes: 02b05bd8b0a6 ("Bluetooth: Set SMP OOB flag if OOB data is available")
Cc: stable@vger.kernel.org
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull scheduler fixes from Ingo Molnar:
- Fix LLC mis-scheduling bugs (Tim Chen, Lu Wang)
- Fix cache-grouping related scheduling statistics UAF bugs (Tim Chen)
- Skip kernel threads for cache aware scheduling to rubustify the code
(Chen Yu)
- Refresh LLC capacity across CPU hotplug, to fix capacity
underestimation bug (Davi Chaves Azevedo)
- Account PSI IRQ time to the execution context, not the scheduling
context, to fix proxy scheduling accounting bug (Zhan Xusheng)
* tag 'sched-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
sched/core: Account PSI IRQ time to the execution context, not the scheduling context
sched/cache: Refresh LLC capacity across CPU hotplug, to fix capacity underestimation bug
sched/cache: Skip kernel threads for cache aware scheduling to rubustify the code
sched/cache: Introduce task_struct->sched_cache_grp to fix UAF
sched/cache: Decouple sched_cache_group from mm to fix UAF
sched/cache: Honor migrate_llc_task semantics in active load balance, to fix LLC mis-scheduling bug
sched/cache: Keep nr_pref_llc_running in the runnable domain, to fix LLC mis-scheduling bug
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull perf events fixes from Ingo Molnar:
- Fixes for KVM guest PEBS virtualization (Sean Christopherson)
- Fixes for various Intel PMUs related to PEBS data-source (Dapeng Mi)
- Fix Intel Panther Cove event scheduling constraints (Dapeng Mi)
- Fix Intel DMR/NVL OMR extra registers event scheduling (Dapeng Mi)
- Rename two confusingly named PMU attributes (Dapeng Mi)
- Fix a refcount leak in attach_perf_ctx_data() (Namhyung Kim)
- Fix NULL pointer dereference crash in __perf_pmu_sched_task()
(Puranjay Mohan)
- Fix CPU-wide event scheduling (Puranjay Mohan)
- Fix x86 LBR branch entry generation (Puranjay Mohan)
* tag 'perf-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
perf/core: Fill branch entries with a single assignment
perf/core: Run sched_task() for PMUs with only CPU-wide events
perf/core: Fix NULL pmu_ctx passed to pmu->sched_task()
perf/core: Fix a refcount leak in attach_perf_ctx_data()
perf/x86/intel: Rename NVL offcore_rsp attribute to offmodule_rsp
perf/x86/intel: Rename DMR offcore_rsp attribute to offmodule_rsp
perf/x86/intel: Fix precise OMR event scheduling for DMR/NVL
perf/x86/intel: Constrain Panther Cove UOPS_DISPATCHED events to PMCs 0-3
perf/x86/intel: Delete dead NVL PEBS data-source initcall
perf/x86/intel: Fix Panther Cove PEBS data-source snoop states
perf/x86/intel: Remove incorrect Panther Cove PEBS data-source constraints
perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints
perf/x86/intel: Update arw_latency_data() mem-op direction handling
perf/x86/intel: Fix DKT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Fix CMT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs()
perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) if PEBS is unused
perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has PEBS isolation, to fix stuck PEBS_ENABLED
perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted PERF_GLOBAL_CTRL bits
|
|
The implied padding causes a harmless warning when testing the uapi
headers with -Wpadded that could in theory indicate incompatibilities
or data leaks:
./usr/include/linux/alloc_tag.h:41:1: error: padding struct size to alignment boundary with 7 bytes [-Werror=padded]
The code here is fine, but it's better to make the padding explicit and
avoid the warning here.
Link: https://lore.kernel.org/20260916065830.1619425-1-arnd@kernel.org
Link: https://lore.kernel.org/20260915202404.3568029-1-arnd@kernel.org
Fixes: 1d581ab2348c ("alloc_tag: add ioctl to /proc/allocinfo")
Signed-off-by: Arnd Bergmann <arnd@arndb.de>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Acked-by: Suren Baghdasaryan <surenb@google.com>
Acked-by: SJ Park <sj@kernel.org>
Acked-by: Hao Ge <hao.ge@linux.dev>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux
Pull ata fixes from Niklas Cassel:
- Extend the quirk "no LPM on ATI" quirk, that is currently only
applied for Samsung drives, to include AMD controllers as well.
The AMD AHCI controllers are newer versions of the ATI AHCI
controllers, and these controllers still have LPM issues with
Samsung drives - LPM works with drives from other vendors (me)
- Fix errors in the libata.force parameter documentation (me)
- Verify the sense data descriptor lengths for ATA PASS-THROUGH
command, so that a malicious device cannot write past the buffer
length (Matthias)
- Mention the libata for-next branch in MAINTAINERS such that the
git ls-remote command done by get_maintainer.pl --self-test=scm
can verify it (Matthias)
* tag 'ata-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux:
MAINTAINERS: name the libata/linux for-next branch
ata: libata-scsi: bound the ATA passthru sense descriptor writes
ata: libata: Correct libata.force parameter documentation
ata: libata-core: Extend Samsung LPM quirk to AMD controllers
|
|
Pull kvm fixes from Paolo Bonzini:
"Arm:
- Invalidate the ITS translation cache when the guest changes the
base address of the ITS tables (Fuad Tabba)
- Skip saving ITS devices with device IDs that are out-of-bounds
rather than failing the entire ITS save ioctl (Fuad Tabba)
- Close race between VM teardown and invalidations of nested MMUs
when handling MMU operations that are allowed to block (Lorenzo
Stoakes)
- Various fixes for the handling of the host's untrusted SVE
configuration in pKVM (Fuad Tabba)
- Make sure that empty SMCCC ranges based at 0 are rejected by the
kvm_smccc_set_filter() (Karl Mehltretter)
- Revoke the host mapping for pKVM's private stack pages, along with
a new sanity check that all mappings in the hyp's private VA range
have been correctly marked as hyp-owned (Fuad Tabba)
- Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that
concurrent vCPU initialization cannot relocate in-use MMUs. Defer
the freeing of shadow stage-2 MMUs to the point that no other users
(e.g. MMU notifier) could reference them (Marc Zyngier)
- Drop useless WARN when rejecting an unsupported ioctl for pKVM
(Fuad Tabba)
- Fix the steal_time selftest to install correctly-sized mappings for
non-4K hosts (Sebastian Ott)
- Correct mapping of fine-grained trap for GCSPOPX instruction (Mark
Brown)
- Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit
guests (Karl Mehltretter)
RISC-V:
- Synchronize hrtimer during VCPU teardown
- Fix the conversion between vsip and hvip values
- Serialize IMSIC attributes with vCPU migration
- Release unused page after MMU invalidation
- Propagate interrupted G-stage faults to KVM user-space as EINTR
- Fix nested acceleration hfence entry update order
- Fix sdata leak and stale snapshot_addr in snapshot_set_shmem
- Preserve firmware counter value across PMU counter stop/start
- Report PMU snapshot write failure to the guest
- Fix perf-backed counter accounting across PMU stop and read
- Correctly propagate error of a hart status SBI call
s390:
- Ensure that accesses through kvm_arch_set_irq_inatomic mark as
dirty the pages that contain indicator and summary bits
- Fix compile warning for kvm_s390_update_cmma_dirty()
- Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to
userspace
- Move s390_kvm_mmu_commit_memory_region() into
s390_kvm_mmu_prepare_memory_region() so that it can fail instead of
WARN
- Add missing srcu in kvm_s390_set_irq_state()
- Fix potential races in storage functions
- Fix race in _destroy_pages_crste()
- Fix issues in the handling of KVM interrupt and page resources,
when a queue that is assigned to a mediated device (mdev) is
removed from the host's AP configuration
- Fix loop condition in uv_find_secrets
- Prevent potential out-of-bounds read
x86:
- Fix a brown paper bag bug where KVM would incorrectly treat Intel
PMU MSRs as valid on AMD
- Fix a regression in the hardware disable selftest where it checked
the wrong macro when detecting glibc support (breaks at least musl)
- Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is
especially important for KVM_BUG_ON() flows, which often guard more
dangerous bugs
- Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a
bug where KVM would let userspace run a broken setup with stale
vmcs12 pages
- Fix a class of bugs where KVM would fail to fill kvm_run exit
fields if getting nested pages failed
- Treat reserved entries in the memory attributes xarray as "no
attributes", to fix false positives when checking for mixed
attributes
- Fix memcg accounting for the memory attributes xarray (the xarray
library subtly requires the xarray to be configured for accounting
upfront; the gfp flags taken at runtime are used only rarely)
- Don't pre-reserve xarray entries when storing empty attributes, as
storing NULL must not require memory allocation (KVM and other
subsystems heavily rely on this behavior)
- Fix a memory leak and a cache maintenance issue related to doing
intra-host migration on an SEV guest"
* tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (54 commits)
KVM: SEV: Do cache maintenance on the source VM during intra-host migration
KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV
KVM: Don't pre-reserve xarray entries when storing empty/NULL attributes
KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg
KVM: Don't treat reserved xarray entries as having memory attributes
KVM: x86: Fill kvm_run exit fields in common get_nested_state_pages() error paths
KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails
KVM: arm64: Fix AArch32 DBGBXVR<n> handling
KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP
KVM: selftests: fix steal_time for arm64 with host page size > 4K
KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode
KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction
KVM: arm64: nv: Fix life cycle of the nested_mmus array
KVM: arm64: Check every private mapping is hyp-owned at pKVM init
KVM: arm64: Move the private VA allocation cursor to __io_map_next
KVM: arm64: Match hyp text by physical address in fix_host_ownership()
KVM: arm64: Transfer the hyp stack pages out of the host stage-2
KVM: arm64: selftests: Test empty SMCCC filter range at base 0
KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0
KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2
...
|
|
Now that KVM uses kvm_get_vcpu_by_id() to check for an existing vCPU ID
before doing any meaningful work, which was made possible by holding
kvm->lock for the entirety of vCPU creation, revert the now-redundant
"early" vCPU ID tracking. The claims about the impact of kvm->vcpu_ids on
the memory footprint were a wee bit wrong: the worst case scenario isn't
256 bytes per VM, it's 256 "unsigned longs" per VM, i.e. 2048 bytes per VM.
Increasing the size of "struct kvm" by 2048 nearly doubled the total size
on many architectures, and tripped x86's KVM_SANITY_CHECK_VM_STRUCT_SIZE,
which was added to detect this *exact* scenario, where a single change
significantly increased the size of "struct kvm". I.e. attempting to build
KVM with CONFIG_DEBUG_KERNEL=n fails on x86 (the build failures got missed
because all build bots apparently test only CONFIG_DEBUG_KERNEL=y kernels,
and maintainers' test flows were similarly lacking).
This reverts commit 97d65b544f48b2ee49f6aea32145e3e7969955dc.
Fixes: 97d65b544f48 ("KVM: Check for duplicate vcpu_id as early as possible")
Reported-by: Jean-Christophe Guillain <jean-christophe@guillain.net>
Closes: https://lore.kernel.org/all/56a4bc35ee605588b7cc36c8e45c12b5f3b506cb.camel@guillain.net
Reported-by: Paweł S <spawel523@gmail.com>
Closes: https://lore.kernel.org/all/CABD%3DWFOS4j4hDv%2BpW-eEM9HAM2q2GY_iYdAG%2BqvYcUEinUrcQQ@mail.gmail.com
Tested-by: Jean-Christophe Guillain <jean-christophe@guillain.net>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Tested-by: Naveen N Rao (AMD) <naveen@kernel.org>
Message-ID: <20260921174445.911676-7-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
KVM fixes for 7.3-rcN
- Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs
as valid on AMD.
- Fix a regression in the hardware disable selftest where it checked the wrong
macro when detecting glibc support (breaks at least musl).
- Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially
important for KVM_BUG_ON() flows, which often guard more dangerous bugs.
- Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug
where KVM would let userspace run a broken setup with stale vmcs12 pages.
- Fix a class of bugs where KVM would fail to fill kvm_run exit fields if
getting nested pages failed.
- Treat reserved entries in the memory attributes xarray as "no attributes",
to fix false positives when checking for mixed attributes.
- Fix memcg accounting for the memory attributes xarray (the xarray library
subtly requires the xarray to be configured for accounting upfront; the gfp
flags taken at runtime are used only rarely).
- Don't pre-reserve xarray entries when storing empty attributes, as storing
NULL must not require memory allocation (KVM and other subsystems heavily
rely on this behavior).
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs
Pull vfs fixes from Christian Brauner:
- Revert "put_mnt_ns(): leave mounts connected". This allows the
creation of reference count cycles in a very trivial way. We can't
bring this in until we have fixed the underlying cause
- vfs: Don't create the private nullfs instance for kthreads under
namespace_sem to avoid false lockdeps complaints
- binfmt_misc:
- Copy the name into a stack buffer and look up the copy in
bpf_binprm_select_interp()
- bpf_binprm_set_interp() and bpf_binprm_set_interp_arg(): Check
the private copy instead so the string that gets staged is the
kstring that was checked
- netfs:
- Make netfs_read_gaps() use separate sink folios rather than one
reused sink folio to discard unwanted data so that cifs checksum
checking sees all the data that was fetched
- Trim reads down to i_size so afs symlinks read correctly from the
cache
- Wrap the direct mempool ->alloc() calls the GFP_KERNEL paths make
in alloc_hooks() via a new mempool_alloc_noreserve() helper
- iov_iter: Use iov_iter_alignment() for the start and length check
added to iov_iter_extract_bvecs() this cycle. It used iter_iov_addr()
and iter_iov_len() which are only valid for ITER_UBUF and ITER_IOVEC
iterators
- super: Make iterate_supers_type() deletion-safe
- inode: Stop evict_inodes() from rescanning the same inodes
- writeback: Bound the cleanup_offline_cgwb() rescans
- ntfs3: Use d_instantiate_new() in ntfs_create_inode()
- ovl: Fix a use-after-free in the ovl_do_mkdir() debug print
- dcache: Unpoison the inline name buffer in __d_alloc() for KMSAN
- autofs: Fix a pipe file reference leak in autofs_kill_sb()
- bpf: Drop the path_unlink and path_rmdir hooks from the list of hooks
for which the verifier rewrites bpf_{set,remove}_dentry_xattr() to
the _locked variants
- squashfs: Range check the xz dictionary size before shifting by it
- selftests: Add the missing eventfd, open_tree_ns, openat2 and xattr
filesystems selftests to TARGETS and drop the stale openat2 entry
left behind when those tests moved
* tag 'vfs-7.3-rc5.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs:
netfs: Fix missing alloc tagging of direct mempool allocations
bpf: fs/xattr: don't assume the inode is locked in path_unlink/path_rmdir
autofs: fix sbi->pipe file reference leak in autofs_kill_sb()
dcache: unpoison the inline name buffer in __d_alloc()
ovl: fix UAF in ovl_do_mkdir() debug print
super: make iterate_supers_type() deletion-safe
Revert "put_mnt_ns(): leave mounts connected"
Revert "selftests/filesystems: add mntns cleanup test"
binfmt_misc: fix racy checks in bpf set_interp kfuncs
binfmt_misc: fix OOB read in bpf_binprm_select_interp()
fs: don't create the private nullfs mount under namespace_sem
writeback: bound cleanup_offline_cgwb() rescans by rotating scanned inodes
fs: avoid repeated scans in evict_inodes()
netfs, afs: Fix symlink reading
netfs: Fix netfs_read_gaps() to use separate sink folios
squashfs: Add dictionary size range check to prevent shift-out-of-bounds
fs/ntfs3: use d_instantiate_new() in ntfs_create_inode() and murder syzbot's "WARNING in do_new_mount" saga
selftests/filesystems: fix missing and stale TARGETS entries
block: Fix start and length check added to iov_iter_extract_bvecs()
|
|
Commit 1d78d56c43ef ("netfs: Fix folio_queue ENOMEM in writeback by
adding a mempool") added a mempool for the folio_queues and made the
request, subrequest and folio_queue allocations distinguish between
writeback and everything else. Writeback is part of memory reclaim
and must not fail due to ENOMEM, so it allocates under GFP_NOFS
through mempool_alloc(), which may dip into the pool's reserve and,
if that runs empty, wait for elements to be returned. The
GFP_KERNEL paths, which can return -ENOMEM to their callers, invoke
the pool's ->alloc() callback directly instead.
The direct call, however, skips the alloc_hooks() wrapper that the
mempool_alloc() macro provides. The pool callbacks, mempool_alloc_slab()
and mempool_kmalloc(), call kmem_cache_alloc_noprof() and kmalloc_noprof()
and rely on current->alloc_tag having been set by the caller. With
CONFIG_MEM_ALLOC_PROFILING_DEBUG=y this leads to
current->alloc_tag not set
WARNING: ./include/linux/alloc_tag.h:161 at __alloc_tagging_slab_alloc_hook
alloc_tag was not set
WARNING: ./include/linux/alloc_tag.h:166 at __alloc_tagging_slab_free_hook
at allocation and free time respectively, as reported when reading
files on a CIFS mount. The allocations are also missing from
/proc/allocinfo.
Wrap the direct ->alloc() invocations in alloc_hooks() with a new
mempool_alloc_noreserve() helper in include/linux/mempool.h, next to
the other alloc_hooks()-wrapped macros such as mempool_alloc(). The
GFP_KERNEL paths keep their failable allocation semantics, they just
get tagged now.
Fixes: 1d78d56c43ef ("netfs: Fix folio_queue ENOMEM in writeback by adding a mempool")
Reported-by: Erhard Furtner <erhard_f@mailbox.org>
Closes: https://lore.kernel.org/all/0b004319-9ef7-437c-a4dd-174d6a9a83db@mailbox.org/
Tested-by: Erhard Furtner <erhard_f@mailbox.org>
Suggested-by: Suren Baghdasaryan <surenb@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Hao Ge <hao.ge@linux.dev>
Link: https://patch.msgid.link/20260923063759.34667-1-hao.ge@linux.dev
Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Add the missing kernel-doc description for the new break_wait field in
struct tty_struct.
Fixes: 845a2737e0f5 ("tty: abort break signalling on hangup")
Reported-by: Randy Dunlap <rdunlap@infradead.org>
Link: https://lore.kernel.org/20d75a7a-2e7a-484f-a267-c21db82dd922@infradead.org
Signed-off-by: Johan Hovold <johan@kernel.org>
Link: https://patch.msgid.link/20260925094222.18303-1-johan@kernel.org
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Jakub Kicinski:
"Including fixes from Bluetooth, NFC and Netfilter.
Every week in this release is record-setting for number of posted
patches. It doesn't seem like we're creating any regressions with all
these fixes, three 'Fixes' tags here point to 7.2 commits but none are
true regression fixes. We're trying to keep the count down,
nonetheless.
Previous releases - regressions:
- net: don't require the hwtstamp NDOs when a PHY provides
timestamping
- ipv6: fix dst leak for uncached routes
- vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE
packets
Previous releases - always broken:
- packet: use ubuf_info completion for TX_RING packets
- arp: terminate device name before lookup
- ipv6: do not let ipv6_find_hdr() return an offset past the packet
end
- udp: remove a disconnected socket from the 4-tuple hash table
- sctp: discard the rest of the packet on a stale-cookie error
- eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged
eswitch"
[ And lots of other random network driver fixes ]
* tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (189 commits)
tcp: prevent collapsing skbs across boundary in rtx queue
vlan: ensure sufficient headroom in vlan_dev_hard_header()
net/sched: sch_teql: fix shadowed err in __teql_resolve()
bridge: check llc_mac_hdr_init() return value in br_send_bpdu()
llc: fix skb UAF and leaks on llc_mac_hdr_init() failure
llc: reserve device headroom for allocated frames
gve: DQO: reject TSO packets with an out of range MSS
gve: fix TX drop when GSO MSS is too small for hw
gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO
net: flush skb_defer_nodes in dev_cpu_dead()
net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails
af_packet: fix integer overflow in prb_calc_retire_blk_tmo()
tipc: Fix a data race on mon->peer_cnt in mon_timeout()
net: phy: intel-xway: workaround 100BASE-TX Link-Up issue
net/smc: fix UAF on lgr list traversal in smcr_port_err()
net/rds: size a connection's path set by the transport it ends up with
nfp: hold IPsec RX state under the XArray lock
net: ena: fix MMIO read buffer leak on probe failure
net: ena: fix PHC cleanup on probe failure
net/sched: act_ct: fix helper UAF due to extensions realloc
...
|
|
tcp_write_collapse_fence() sets TCP_SKB_CB(skb)->eor = 1 on
tcp_write_queue_tail(sk) to prevent skbs queued after a switch to
device encryption from being collapsed into earlier skbs.
The fence is a no-op if all earlier data has already been transmitted
when the switch happens: sk->sk_write_queue is empty. The not yet
acknowledged earlier skbs wait in sk->tcp_rtx_queue with eor 0.
On a subsequent retransmit or SACK shift, tcp_retrans_try_collapse() or
tcp_shift_skb_data() can then merge an skb queued after the switch into
one queued before it.
Both users of the fence are affected:
- psp: devices only encrypt skbs with skb->decrypted set. The merged skb
keeps decrypted = 0 from the earlier skb, so merged data sent after
psp_sock_assoc_set_tx() is retransmitted in cleartext.
- tls device offload: the merged skb straddles the start marker set in
tls_set_device_offload(). The software fallback (fill_sg_in() returns
-EINVAL) and the mlx5, nfp and funeth drivers cannot handle such an
skb and drop it. Every retransmit rebuilds the same skb, so the
connection stalls.
Fix this in two places, for defense in depth:
1. Fall back to tcp_rtx_queue_tail(sk) in tcp_write_collapse_fence()
when tcp_write_queue_tail(sk) is NULL.
2. Check !skb_cmp_decrypted(to, from) in tcp_skb_can_collapse(), as
tcp_skb_can_collapse_rx() does on receive. skb_shift(), which both
collapse paths call, already has a DEBUG_NET_WARN_ON_ONCE() for this
condition.
Fixes: e8f69799810c ("net/tls: Add generic NIC offload infrastructure")
Cc: stable@vger.kernel.org
Signed-off-by: Willem de Bruijn <willemb@google.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Daniel Zahka <daniel.zahka@gmail.com>
Link: https://patch.msgid.link/20260924154427.953800-1-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux
Pull Landlock fixes from Mickaël Salaün:
"This mainly fixes the Landlock tracepoint support merged this cycle so
that denial and rule events report the intended policy context,
whether through tracefs or BTF-visible callbacks.
The size of this all is mainly from propagating the corrected contract
through event definitions and producers, adding new tests for the
reported context, and updating the documentation.
Also improve annotation and fix a GCC 16 build warning"
* tag 'landlock-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux:
landlock: Widen ruleset versions to 64 bits
landlock: Add counted_by in landlock_domain
landlock: Fix tracepoint contract documentation
selftests/landlock: Test network denial context
selftests/landlock: Test filesystem denial blockers
landlock: Report the effective signal number
landlock: Report the actual ptrace tracer
landlock: Fix network denial trace context
landlock: Fix rule tracepoint context
landlock: Fix filesystem denial blocker reporting
landlock: Fix tracepoint fixed-width type names
landlock: Work around gcc-16 -Wuninitialized warning
|
|
In a case where skb with an unconfirmed ct entry gets cloned, we may
end up committing both but with different sets of extensions.
The series of events:
1. The first clone wants to commit and runs the helpers wiring up
the extension pointer into the expectation list.
2. Then it looses the confirmation keeping the entry unconfirmed.
3. Second clone now wants to commit labels and adds the new extension
for that breaking the pointer in the expectation list causing
UAF on the destruction path later.
While this is possible to trigger, there should be no practical
network pipeline where committing both clones without modifications
into the same zone is needed. So, let's just reset the entry in case
for some reason we got an skb with a shared one during commit. This
doesn't affect any known use cases, but avoids any potential problems
with sharing and modification of the unconfirmed ct entry.
The fixes tag points to the introduction of helpers, since that's the
main UAF trigger for the sharing.
Fixes: cae3a2627520 ("openvswitch: Allow attaching helpers to ct action")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-2-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Pull bpf fixes from Alexei Starovoitov:
- Fix bpf_skb_change_tail() to drop the checksum offload instead of
rejecting the trim of CHECKSUM_PARTIAL skbs (Daniel Borkmann)
- Add KF_PERFMON kfunc flag and require CAP_PERFMON for kfuncs that
read arbitrary memory and for untrusted read-only memory reads
(Daniel Borkmann)
- Clear scalar delta on narrowing stack spill (Daniel Borkmann)
- Set up the frame pointer for the exception callback in arm64 JIT, and
zero-fill other CPUs when BPF_F_CPU update creates a per-cpu hash
element (Donggeun Yoo)
- Various fixes (Emil Tsalapatis):
- Fix bounds check underflow for skb-backed dynptrs
- Fix rx_queue_mapping context access code generation in bpf_sock
- Reject packet pointer arguments to subprogs that may mutate the
packet
- Reject ALU instructions that see arena and non-arena operands on
different code paths
- Fix copied_seq double-counting on sockmap self-redirect
(Geliang Tang)
- Fix divide-by-zero in btf_struct_walk() on a flexible array of
zero-sized elements, fix out-of-bounds read of rtt_min in sock_ops
(Jiayuan Chen)
- Fix bpf_sock_destroy() out-of-bounds read of sk_protocol on TIME_WAIT
and request socks, and sleeping under RCU when destroying a listener
with pending children (Jiayuan Chen)
- Fix JEQ/JNE with immediate operand in MIPS32 JIT and missing zero
extension of BSWAP 16/32 in MIPS64 JIT (Johan Almbladh)
- Avoid soft lockup in htab lookup[_and_delete] batch operations on
large maps (Jose Fernandez)
- Various fixes (Kumar Kartikeya Dwivedi):
- Verify global subprogs in each sleepability context they are
called from
- Make post-verification instruction rewrites killable
- Preserve packet pointer displacement in regsafe()
- Apply CO-RE relocations before subprogram validation, restrict
CO-RE poisoning to relocatable instructions, and reject truncated
ldimm64 CO-RE relocations in libbpf
- Assign lock identity to callback map values
- Compare stack frames in regs_exact()
- Bound ownership depth through local kptrs and graph roots
- Fix u32 overflow in map batch operations when the map size exceeds
4GB (Masoud Aghasi)
- Fix UAF in bpf memalloc due to concurrent consumption of ttrace lists
in alloc_bulk() (Pu Lehui)
- Allow gotox as the terminal instruction of a program or a subprogram
(Siddharth Chintamaneni)
- Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL, skip unsettled links
in link iterator, and reject dev-bound-only programs on other devices
(Weiming Shi)
- Reject non-negative stack offsets in stack_slot_obj_get_spi()
(Xu Yunxiang)
- Check params size before reading reserved fields in
bpf_crypto_ctx_create() (Yuqi Xu)
- Reject max_entries > INT_MAX in sock_map_alloc() (Zhao Gongyi)
- Use a 32-bit compare in xsk_map_gen_lookup() (Zhiling Zou)
- Use kvfree() in xdp_test_run_teardown() (Zhixing Chen)
* tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: (58 commits)
selftests/bpf: Test per-cpu initialization of a BPF_F_CPU created element
bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element
bpf: Fix BSWAP 32 and 16 on MIPS64
bpf: Fix immediate JMP JEQ/JNE on MIPS32
bpf: Reject dev-bound-only programs on other devices
bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc
selftests/bpf: Test for mixed arena/nonarena code paths
bpf: Prevent variable arena/non-arena register contents
selftests/bpf: Test rejection of pkt args to mutating subprogs
bpf: Reject pkt arguments in mutating subprogs
selftests/bpf: Add selftests for rx_queue_mapping context access
bpf: Fix bpf_sock context code generation
selftests/bpf: Test dynptr slices past end of skb
bpf: Fix bounds check for skb-backed dynptrs
selftests/bpf: Reject iterator destruction through fp+0
bpf: Reject non-negative offsets in stack_slot_obj_get_spi()
bpf: Check params size before reading reserved fields
selftests/bpf: Check local object ownership depth
bpf: Bound ownership depth through local kptrs and graph roots
selftests/bpf: Cover frame changes in bounded loops
...
|