| Age | Commit message (Collapse) | Author |
|
https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git
# Conflicts:
# Documentation/scheduler/index.rst
# arch/arm64/configs/defconfig
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/modules/linux.git
|
|
# New commits in sched/core:
4a3b51aab6e2 ("smpboot: Don't park the thread if work is pending")
40dcc9bdbef3 ("irq_work: Flush lazy work CPU down on PREEMPT_RT")
791b1760accd ("irq_work: Update a comment regarding CPU hotplug invocation")
648d44bda731 ("sched/topology: Add asymmetric SMT packing override")
c8fc4136fd3c ("sched/fair: Honor asymmetric SMT priority in idle selection")
53bc5c556b82 ("sched: Set TIF_NEED_RESCHED before calling __trace_set_need_resched()")
4b1f75be23c4 ("sched/core: Fix context analysis errors in non-preferred CPU push")
1fb28c664a19 ("virt/steal_governor: Enable the driver")
27d47ebce4d6 ("virt/steal_governor: Implement steal_governor policy loop")
4b9302d494ff ("virt/steal_governor: Add control knobs for handling steal values")
9a8e740ee9f6 ("virt: Introduce steal governor driver")
68957caaa9c0 ("sched/debug: Add migration stats due to non preferred CPUs")
74699f56ebcf ("sched/core: Push current task from non preferred CPU")
4ee29b029058 ("sched/fair: Load balance only among preferred CPUs")
d8a3da0de843 ("sched/core: Try to use a preferred CPU in is_cpu_allowed")
620824516557 ("sysfs: Add preferred CPU file")
518b32bd5bb3 ("cpumask: Introduce cpu_preferred_mask")
06a49ef784ac ("sched/docs: Document cpu_preferred_mask and Preferred CPU concept")
cfb463b7172d ("cpumask: Introduce cpumask_intersects_and")
a8d0854a76a8 ("sched/cputime: Add kcpustat_field_total helper")
be100c77178e ("sched: Add sched_ext hooks for proxy execution")
57c75e3ae38c ("sched: Add helper to block retained proxy donors")
a49653d0abeb ("sched/core: Mark wakeups completed through ttwu_runnable()")
8f8c0417e973 ("sched/core: Dequeue waking proxy donors before reset")
313b652837d0 ("sched/core: Drop mutex locks before proxy rescheduling")
627ea30aca3b ("sched/wait: Clarify WF_SYNC wakeup semantics")
d2e010082757 ("sched/eevdf: Handle more short slice waking cases")
4bf32ec3327d ("sched/eevdf: Align update_protect_slice to set_protect_slice")
aae2a33ea662 ("sched/eevdf: Ensure that vprot will never go above a min slice")
c9ce69fc43bd ("sched/fair: Randomize equally shallow slow-path candidates")
abe440b3770f ("sched/fair: Drop idle recency from slow-path CPU selection")
fbbc63fed0b0 ("sched/core: Remove redundant core_sched_seq")
819224e506bc ("sched/fair: Remove dead code on enqueue_task_fair()")
c72945693b90 ("sched: Restart fair hrtick after same-task repicks")
a9b3c7570564 ("sched/headers: Replace __ASSEMBLY__ with __ASSEMBLER__ in the <uapi/linux/sched.h> header")
e81ee0630837 ("sched/fair: Reset NUMA fault locality after scan period update")
ef9293b3b797 ("sched: dynamic: Fix preemption model strings")
879eaa76e608 ("sched: Remove unneeded function type cast in do_balance_callbacks()")
f549101187c8 ("sched/deadline: check start_dl_timer expiry with ktime_before()")
2a672daa4b27 ("sched/feat: Use the new static key API for sched_feat")
a5576ebce920 ("sched: Convert paravirt_steal to new static key APIs")
9650ce11f2e3 ("sched: dynamic: Simplify preempt model accessors")
5b9a28eeed37 ("sched: dynamic: Remove HAVE_PREEMPT_DYNAMIC_{CALL,KEY}")
aa4178f63847 ("sched: dynamic: Simplify irqentry_exit_cond_resched()")
b9d267b9d632 ("sched: dynamic: Simplify preempt_schedule{,_notrace}()")
88e0b3bb9930 ("sched: dynamic: Simplify {cond,might}_resched()")
d3d16750693b ("sched: dynamic: Make PREEMPT_DYNAMIC depend on ARCH_HAS_PREEMPT_LAZY")
772d9ffbfd26 ("sched: Migrate whole chain in proxy_migrate_task()")
6b73a09e943f ("sched: Break out core of attach_tasks() helper into sched.h")
1f8805138593 ("sched: Switch rq->next_class in proxy_reset_donor()")
09351db90a28 ("sched/core: Don't proxy-exec unmatched cookie lock owners")
9be817f991e2 ("sched/core: Avoid migrating blocked_on tasks")
3dd95f077371 ("sched/core: Don't steal a proxy-exec donor")
Signed-off-by: Ingo Molnar <mingo@kernel.org>
|
|
Add a 'perf test -w false_sharing' workload that hammers one shared
struct from several CPUs, shaped as a TCP connection: a read-mostly
identity (five-tuple) shares a cacheline with per-packet rx counters
(the false-sharing line), a second line has packet-path private tx and
congestion control counters, and a third the connection config.
The packet path runs in the main thread and up to four lookup threads
sum the five-tuple and pull the config, reading one volatile shared
instance directly so the accesses are PC-relative and resolvable by the
data type profiler.
Data type profiling of the workload with cacheline info:
$ perf mem record -- perf test -r5 -w false_sharing
[ perf record: Woken up 19 times to write data ]
[ perf record: Captured and wrote 5.612 MB perf.data (72454 samples) ]
$ perf report -s type,typecln -H --group --stdio
# Total Lost Samples: 0
#
# Samples: 72K of events 'cpu/mem-loads,ldlat=30/P, cpu/mem-stores/P'
# Event count (approx.): 1261814616
#
# Overhead Data Type / Data Type Cacheline
# ................... ...............................
#
88.99% 35.88% struct net_conn
56.43% 17.50% struct net_conn: cache-line 0
32.55% 0.00% struct net_conn: cache-line 2
0.01% 18.38% struct net_conn: cache-line 1
10.08% 63.03% (unknown)
10.08% 63.03% (unknown): cache-line 0
0.82% 0.13% int
0.82% 0.13% int: cache-line 0
0.01% 0.02% struct folio
0.01% 0.02% struct folio: cache-line 0
0.01% 0.03% Elf64_Addr
0.01% 0.03% Elf64_Addr: cache-line 0
0.01% 0.03% struct sched_entity
0.01% 0.03% struct sched_entity: cache-line 1
0.00% 0.00% struct sched_entity: cache-line 4
0.00% 0.00% struct sched_entity: cache-line 2
0.00% 0.00% struct sched_entity: cache-line 3
0.00% 0.00% struct sched_entity: cache-line 0
0.01% 0.00% struct task_group
0.01% 0.00% struct task_group: cache-line 4
0.00% 0.00% struct task_group: cache-line 5
0.00% 0.00% struct css_rstat_cpu
0.00% 0.00% struct css_rstat_cpu: cache-line 0
$ perf report -s type,typecln,typeoff -H --group --stdio
# Total Lost Samples: 0
#
# Samples: 72K of events 'cpu/mem-loads,ldlat=30/P, cpu/mem-stores/P'
# Event count (approx.): 1261814616
#
# Overhead Data Type / Data Type Cacheline / Data Type Offset
# ...................... ..................................................
#
88.99% 35.88% struct net_conn
56.43% 17.50% struct net_conn: cache-line 0
9.81% 0.00% struct net_conn +0xa (dport)
9.54% 0.00% struct net_conn +0x8 (sport)
9.53% 0.00% struct net_conn +0xc (state)
9.48% 0.00% struct net_conn +0xd (protocol)
9.27% 0.00% struct net_conn +0x4 (daddr)
8.79% 0.00% struct net_conn +0 (saddr)
0.00% 12.08% struct net_conn +0x10 (bytes_rx)
0.00% 1.71% struct net_conn +0x18 (packets_rx)
0.00% 2.80% struct net_conn +0x20 (rx_queue)
0.00% 0.91% struct net_conn +0x3c (last_ack)
32.55% 0.00% struct net_conn: cache-line 2
5.70% 0.00% struct net_conn +0x83 (rcv_wscale)
5.46% 0.00% struct net_conn +0x84 (keepalive_int)
5.43% 0.00% struct net_conn +0x82 (snd_wscale)
5.39% 0.00% struct net_conn +0x80 (mss)
5.34% 0.00% struct net_conn +0x88 (mark)
5.23% 0.00% struct net_conn +0x8c (priority)
0.01% 18.38% struct net_conn: cache-line 1
0.01% 0.05% struct net_conn +0x5c (retrans)
0.00% 4.83% struct net_conn +0x40 (bytes_tx)
0.00% 10.63% struct net_conn +0x48 (packets_tx)
0.00% 1.54% struct net_conn +0x58 (rtt_us)
0.00% 1.07% struct net_conn +0x50 (cwnd)
0.00% 0.25% struct net_conn +0x54 (ssthresh)
10.08% 63.03% (unknown)
10.08% 63.03% (unknown): cache-line 0
10.08% 63.03% (unknown)
0.82% 0.13% int
0.82% 0.13% int: cache-line 0
0.82% 0.13% int +0 (no field)
0.01% 0.02% struct folio
0.01% 0.02% struct folio: cache-line 0
0.01% 0.00% struct folio +0 (flags.f)
0.00% 0.01% struct folio +0x34 (_refcount.counter)
0.00% 0.00% struct folio +0x18 (mapping)
Assisted-by: LLM
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Now that --progress is being added for the stdio case, wire up its
counterpart for the browsers: the TUI and GTK ones present progress of
their own and there is no way to turn it off. Install the no-op
ui_progress ops, the ones already used until a backend sets theirs,
after setup_browser() installed the ones of the browser in use. The
phases are still counted, nothing is shown for them, and no second
option is needed for it: parse-options provides --no-progress as the
negation of --progress, report.progress_set saying that it was asked
for, report.progress being false both when nothing was asked for and
when --no-progress was.
Suggested-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: LLM
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Processing a large session with stdio output gives no feedback about
which phase perf is in or how far along it is: ui_progress updates are
only shown by the TUI. Add --progress, installing a stdio backend
(ui/stdio/progress.c) that prints the phase title, percentage and
counts:
Processing events... [ 42.3%] 317M / 746M
Phases can be nested, so the backend tracks the ones started to
complete the right one on ui_progress__finish(); that requires
init()/finish() pairs, fixed in ordered-events.c and the pipe and
directory event processing.
With a pager both stdout and stderr lead to it, so isatty(stderr)
turns false and the updates end up printed one per line in the
pager's output: use /dev/tty in that case, so that the updates keep
updating in place, untouched by the pager.
Assisted-by: LLM
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Config file access shares the static parser state and, since
perf_config__set_variable() became an entry point for background
features, can run on more than one thread: a feature writing the
config while perf top's display thread reads it. Serialize parsing
and rewriting with config_mutex and the whole read-modify-write of
perf_config__set_variable() with config_update_mutex; config_set_mutex
comes before it, as building the set parses the config files.
perf has its own mutex type, util/mutex.h; the file scope mutexes
here use the DEFINE_MUTEX() static initializer added in the previous
patch. The two lazy inits that system_path() and home_perfconfig()
do would leak all but one of the strings racing on two threads, so
they move to DO_ONCE(), added in the previous patch as well.
bad_config() runs on the dispatching thread, outside config_mutex, so
it stops reading config_file_name, which the parsing thread owns.
Assisted-by: LLM
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
For lazy init that more than one thread can race into, without
pthread_once() so that it stays within perf's locking primitives and
clang's -Wthread-safety keeps an eye on the mutexes involved.
Mirrors the kernel's include/linux/once.h DO_ONCE(), dropping its
static key fast path, unnecessary at these cold init paths: a per call
site static bool plus a statically initialized mutex, double checked.
As there, code reachable from more than one call site must go through
a common helper.
The lockless fast path reads ___done with an acquire, paired with a
release store after fn(), so that fn()'s side effects are visible to
threads taking the fast path, which the dropped static key used to
guarantee.
The __atomic builtins are used rather than smp_load_acquire()/
smp_store_release() because ThreadSanitizer understands them,
avoiding reports on ___done and on what fn() initialized.
Assisted-by: LLM
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
For file scope mutexes, so that they don't need a constructor function
to initialize them at startup, mirroring PTHREAD_MUTEX_INITIALIZER
while keeping the clang -Wthread-safety annotations of struct mutex.
DEBUG=1 builds have mutex_init() set PTHREAD_MUTEX_ERRORCHECK, making
the CHECK_ERR() paths in mutex_lock()/mutex_unlock() report relocking a
held mutex and unlocking one that isn't held. PTHREAD_MUTEX_INITIALIZER
gives a default type mutex, so statically initialized ones would silently
lose that: use PTHREAD_ERRORCHECK_MUTEX_INITIALIZER_NP, a glibc
extension, falling back to the default type in libcs that lack it.
Assisted-by: LLM
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
perf_etc_perfconfig() returns what system_path() gives, and that
allocates, so on failure every caller dereferenced NULL, the
pre-existing ones in perf_config_set__init() and in the daemon
included. For the usual absolute ETC_PERFCONFIG system_path() just
returns a strdup of it, so use ETC_PERFCONFIG itself when that fails.
Assisted-by: LLM
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Move perf_config__set_variable() out of the 'perf config' builtin so
that opt-in features can persist their choice from outside it, e.g.
util/debuginfo.c writing core.debuginfod=false when the user disables
debuginfod for the rest of the session. The set_config() body becomes
perf_config_set__write(), with the system_config choice as an argument,
as the builtin's use_system_config/use_user_config statics are not
available outside it.
perf_config_set__write() checked fopen() but none of the fprintf()s or
fclose(), so a write failure after truncating the file was reported as
success. Harmless for the interactive 'perf config' this came from,
but this makes it an entry point a background feature can call with no
other feedback, so propagate those errors too.
Assisted-by: LLM
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Although config4 and tee sample_simd_* fields has been added into
'struct perf_event_attr', the size expectation for the struct hasn't
been adjusted. This hasn't been noticed, since the testcase's return
value has been ignored after rewriting the testcase to python until
previous commit.
Fix that.
Fixes: eb89aef367e47010 ("perf headers: Sync perf_event.h/perf_regs.h with the kernel headers")
Fixes: 80cdf208117a36de ("tools headers UAPI: Sync linux/perf_event.h with the kernel sources")
Signed-off-by: Michael Petlan <mpetlan@redhat.com>
Acked-by: Namhyung Kim <namhyung@kernel.org>
Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
[ Updated it to take into account 80cdf208117a36de (sample_simd_* fields) 144 -> 176 ]
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Currently it does not matter what the python script actually returns,
the test always passes, as the $err variable is always 0.
Fix that.
Fixes: 8519e4f44c2af722 ("perf test: Add a shell wrapper for "Setup struct perf_event_attr"")
Signed-off-by: Michael Petlan <mpetlan@redhat.com>
Acked-by: Namhyung Kim <namhyung@kernel.org>
Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Test launching ttimechart through 'perf timechart --tui', the error when
there are no scheduler events, the --dump output with the -P, -T and -p
options, and drive the textual app headless exercising the views and
key bindings.
Committer testing:
root@x2:/home/acme/git/perf-tools-next# perf test "perf timechart --tui (ttimechart.py) test"
167: perf timechart --tui (ttimechart.py) test : Ok
=== Test Summary ===
Passed main tests : 1
Passed subtests : 0
Skipped tests : 0
Failed tests : 0
root@x2:/home/acme/git/perf-tools-next# perf test -vv "perf timechart --tui (ttimechart.py) test"
167: perf timechart --tui (ttimechart.py) test:
---- start ----
test child forked, pid 2833475
perf timechart --tui plumbing test
perf timechart --tui plumbing test [Success]
ttimechart no events test
ttimechart no events test [Success]
ttimechart dump test
ttimechart dump test [Success]
ttimechart headless UI test
ttimechart headless UI test [Success]
---- end(0) ----
167: perf timechart --tui (ttimechart.py) test : Ok
=== Test Summary ===
Passed main tests : 1
Passed subtests : 0
Skipped tests : 0
Failed tests : 0
root@x2:/home/acme/git/perf-tools-next#
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Alice Rogers <alice.mei.rogers@gmail.com>
Co-developed-by: Ian Rogers <irogers@google.com>
Signed-off-by: Ian Rogers <irogers@google.com>
Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Add ttimechart.py, a textual based interactive timechart, and a --tui
option to perf timechart that launches it through perf script. Rather
than writing an SVG file, per-CPU (busy, frequency and idle state),
per-task (running, runnable and blocked) and I/O timelines are shown in
the terminal along with a summary table. The timelines can be zoomed,
panned and the state of the selected row at the cursor is described,
including the task that woke it. The -i, -P, -T and -p options are
passed to the script.
The data is loaded in a background thread and the timelines are shown
and updated as it loads. A --dump option prints a text summary without
the UI.
Committer testing:
Fist record a short session with:
root@x2:/home/acme/git/perf-tools-next# perf timechart record
^C[ perf record: Woken up 9 times to write data ]
[ perf record: Captured and wrote 4.489 MB perf.data (29719 samples) ]
root@x2:/home/acme/git/perf-tools-next#
Then try the TUI with:
root@x2:/home/acme/git/perf-tools-next# tools/perf/python/ttimechart.py
or:
root@x2:/home/acme/git/perf-tools-next# perf timechart --tui
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Alice Rogers <alice.mei.rogers@gmail.com>
Co-developed-by: Ian Rogers <irogers@google.com>
Signed-off-by: Ian Rogers <irogers@google.com>
Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Test launching treport through 'perf script', the error for a file that
isn't a perf.data file, drive the textual app headless checking the
profile matches one built without the UI and exercising the key
bindings, and that quitting while loading stops the background load.
Committer testing:
root@x2:/home/acme/git/perf-tools-next# perf test "perf script treport test"
166: perf script treport test : Ok
=== Test Summary ===
Passed main tests : 1
Passed subtests : 0
Skipped tests : 0
Failed tests : 0
root@x2:/home/acme/git/perf-tools-next# perf test -vv "perf script treport test"
166: perf script treport test:
---- start ----
test child forked, pid 2829708
treport plumbing test
treport plumbing test [Success]
treport bad file test
treport bad file test [Success]
treport headless UI test
treport headless UI test [Success]
treport cancel test
treport cancel test [Success]
---- end(0) ----
166: perf script treport test : Ok
=== Test Summary ===
Passed main tests : 1
Passed subtests : 0
Skipped tests : 0
Failed tests : 0
root@x2:/home/acme/git/perf-tools-next#
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Alice Rogers <alice.mei.rogers@gmail.com>
Co-developed-by: Ian Rogers <irogers@google.com>
Signed-off-by: Ian Rogers <irogers@google.com>
Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Reading a large perf.data file caused a long stall before the treport
app started. Start the app immediately and build the profile in a
background thread, showing the report tree and flame graph as they are
built with the progress in the header. The time between updates grows
with the load time as updating the views costs more as the profile
grows. A lock guards the profile shared between the threads. Quitting
while loading cancels the load.
Tree nodes are now created lazily as they are expanded, expanded nodes
and the cursor are kept across updates, and events are kept in the order
they are first seen as their values aren't comparable. Node names are
escaped so they aren't interpreted as markup.
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Alice Rogers <alice.mei.rogers@gmail.com>
Co-developed-by: Ian Rogers <irogers@google.com>
Signed-off-by: Ian Rogers <irogers@google.com>
Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Test launching ilist through perf script and perf list --tui, and drive
the textual app headless: check the tree of PMUs and metrics, search for
and select the software task-clock event checking its counters are
shown, check a failed search shows an error and select a metric. Failing
to open counters, for example for want of permissions, is tolerated as
an error dialog is shown.
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Add a --tui option to perf list that launches the textual based ilist
script through perf script, similar to perf timechart --tui. Rather
than printing the events, the PMUs, events and metrics can be browsed
interactively and the selected event or metric is counted on each CPU.
Arguments after '--' are passed to the script.
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
When a python session reads from a pipe, perf_session__new() and
perf_session__process_events() block waiting for data with the GIL held,
freezing any other python threads such as a TUI thread. Release the GIL
around both calls and acquire it with PyGILState_Ensure() in the tool and
call-return callbacks that invoke python.
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Previously, pyrf_event__new() eagerly resolved the sample's address
location and callchain (including DWARF/FP unwinding and symbol lookup)
for every sample with a callchain, even when the Python callback never
accessed event.callchain (such as in ttimechart.py).
Move callchain resolution to pyrf_sample_event__get_callchain() so that
callchains are only unwound and resolved when event.callchain is
accessed, reusing the sample's cached pevent->al via
pyrf_sample_event__resolve_al().
Because process_events() borrows the event and sample during the
callback, callchains are never copied in the common case. Only when a
callback retains a reference to the event beyond its return
(Py_REFCNT(pyevent) > 1), the callchain was not already resolved during
the callback, and the callchain is synthesized or LBR-merged (not backed
by event_copy) does pyrf_event__copy() allocate a copy of
sample->callchain.
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
When perf.session.process_events() delivers an event to a Python
callback, pyrf_event__new() previously copied the raw perf_event into a
4,160-byte union perf_event embedded in struct pyrf_event and called
evsel__parse_sample() a second time to re-parse the sample against the
copied buffer.
In the common case, the Python callback inspects the event and returns
without storing a reference to it. Avoid the event memcpy, the duplicate
evsel__parse_sample() call, and the large PyObject allocation by:
1. Storing pointers to the underlying union perf_event and
struct perf_sample in struct pyrf_event (using PyGetSetDef descriptors
backed by PyMember_GetOne() to read fields through those pointers),
shrinking struct pyrf_event so it fits in CPython's small-object
allocator pool.
2. Borrowing the event and pre-parsed sample pointers from
process_events() for the duration of the Python callback.
3. Checking Py_REFCNT(pyevent) > 1 after the callback returns and only
allocating event_copy and re-parsing into sample_storage when the
Python callback retained a reference to the event object (or immediately
in pyrf_evlist__read_on_cpu() where no pre-parsed sample is passed).
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
When a session callback raises a python exception the callback returned
-1 causing reader__read_event() to print a "processing failed for event
of type: ..." error. This is redundant with the python exception that is
raised from process_events(), and corrupts the display of a TUI that
raises from a callback to deliberately stop processing, such as when
cancelling a background load.
Instead set session_done so that processing stops and return success,
the pending exception is then raised by process_events(). As a single
event may still cause further callbacks, for example multiple
call_return callbacks followed by a sample callback, skip calling python
while an exception is pending.
The session whose callback raised resets session_done after processing
so later sessions process all their events. As session_done is a global,
a session processed concurrently in another thread may be stopped early
too. Count the stop requests so that such a session raises an error
rather than silently dropping events.
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
When session_done() is set, for example by SIGINT or by a python
callback requesting processing stop, the reader loops stop but the
deferred samples and auxtrace data are still flushed afterwards. For
large traces decoding and delivering the remaining auxtrace data can be
slow, and the events are delivered to tools that asked to stop.
Discard the remaining deferred samples, and skip
auxtrace__flush_events(), when session_done() is set. The
ordered_events flush already honors session_done().
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Since the Python extension switched to linking perf libraries, setup.py
only lists util/python.c as a source. The libraries are supplied through
LDFLAGS, which setuptools does not consider when deciding whether the
extension needs to be rebuilt.
When a library changes, make invokes setup.py, but setuptools can skip
the build and the recipe copies the stale cached extension back into
python/. This leaves perf and its Python module running different
versions of the same code.
Pass the linked library paths to setup.py and declare them as Extension
dependencies so that library changes trigger a rebuild. Include
EXTRA_PERFLIBS in the shared library list so that both make and
setuptools track those inputs as well, preserving the existing linker
order.
Fixes: 9dabf4003423c8d3 ("perf python: Switch module to linking libraries from building source")
Reviewed-by: Ian Rogers <irogers@google.com>
Assisted-by: Codex:GPT-6-Astra
Signed-off-by: James Clark <james.clark@linaro.org>
[ Applied manually to address some fuzz ]
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
With sockaddr arguments copied as buffers via beauty_map_enter, the
32-byte TRACE_AUG_MAX_BUF limit only held 2 bytes of sa_family plus 30
bytes of AF_LOCAL sun_path, truncating longer socket paths such as
"/var/run/.heim_org.h5l.kcm-socket".
Since struct augmented_arg already reserves PATH_MAX (4096) bytes,
increase TRACE_AUG_MAX_BUF to 128 (SS_MAXSIZE, the size of struct
sockaddr_storage), which covers all 110 bytes of struct sockaddr_un
without changing map sizes. Also document the 1-based negative -(arg_idx
+ 1) encoding used in beauty_array for paired length arguments.
Committer testing:
root@x2:~# perf trace -e connect,bind --max-events=5
0.000 ( 0.021 ms): DNS Res~r #244/2760970 bind(fd: 202, umyaddr: { .family: NETLINK }, addrlen: 12) = 0
18446744073709.250 ( 0.031 ms): DNS Res~r #245/2765812 bind(fd: 202, umyaddr: { .family: NETLINK }, addrlen: 12) = 0
0.306 ( 0.086 ms): DNS Res~r #245/2765812 connect(fd: 202, uservaddr: { .family: LOCAL, path: /run/systemd/resolve/io.systemd.Resolve }, addrlen: 42) = 0
0.551 ( 0.056 ms): DNS Res~r #244/2760970 connect(fd: 204, uservaddr: { .family: LOCAL, path: /run/systemd/resolve/io.systemd.Resolve }, addrlen: 42) = 0
2.074 ( 0.069 ms): systemd-resolv/103685 connect(fd: 27, uservaddr: { .family: INET, port: 53, addr: 192.168.30.1 }, addrlen: 16) = 0
root@x2:~#
Suggested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Cc: Howard Chu <howardchu95@gmail.com>
Cc: Jakub Brnak <jbrnak@redhat.com>
Cc: Namhyung Kim <namhyung@kernel.org>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
In task_traced(), bpf_get_current_pid_tgid() was called before checking
has_pids_to_trace. Because BPF helper calls are not marked pure, the
compiler cannot sink the call past the has_pids_to_trace check, so the
helper was called on every syscall enter and exit even when tracing
system-wide.
Check has_pids_to_trace before calling bpf_get_current_pid_tgid().
Suggested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Cc: Howard Chu <howardchu95@gmail.com>
Cc: Jakub Brnak <jbrnak@redhat.com>
Cc: Namhyung Kim <namhyung@kernel.org>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Fix several missing dependencies and ordering issues for clean and
install targets:
1. In tools/perf/Makefile, order 'install*' goals (excluding
'install-build-deps', which is ordered before build/install goals)
after 'all' when both are passed on the command line (e.g.
'make -j clean all install') so two concurrent Makefile.perf
sub-makes do not race in the same build directory, and mark 'all' and
'clean' as .PHONY.
2. In tools/perf/Makefile.perf, when 'clean' is passed alongside build
or install goals (e.g. 'make -f Makefile.perf clean install'), run
'clean' sequentially before 'fixdep' and 'sub-make' rather than
running 'clean' in parallel with the build inside 'sub-make'.
3. Ensure '$(OUTPUT)python' is created inside the recipe for
'$(OUTPUT)python/perf$(PYTHON_EXTENSION_SUFFIX)' rather than only at
Makefile parse time, in case 'clean' removed the directory.
4. Add '$(LANG_BINDINGS)' as a prerequisite of 'install-python_ext' so
the Python C extension is built with the proper compiler/linker flags
and perf libraries before 'setup.py install' runs.
5. Add 'install-bin', 'install-tools', 'install-tests',
'install-python_ext', '$(DOC_TARGETS)', and '$(INSTALL_DOC_TARGETS)'
to .PHONY in Makefile.perf.
6. Prefix targets emitted by Documentation/build-docdep.perl with
'$(OUTPUT)' so '$(OUTPUT)doc.dep' matches out-of-tree documentation
targets when 'O=' is specified.
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Commit 4751bddd3f983af2 ("perf tools: Make GTK2 support opt-in")
switched Makefile.config and Makefile.perf from checking 'ifndef
NO_GTK2' to 'ifdef GTK2' (now 'ifdef GTK4'), leaving 'NO_GTK4 := 1' in
Makefile.config unused.
As a result, when building with GTK4=1 on a system without gtk4
development headers, Makefile.perf still attempted to build and install
libperf-gtk.so under 'ifdef GTK4'.
Replace 'NO_GTK4 := 1' with 'override GTK4 :=' so that 'ifdef GTK4' in
Makefile.perf evaluates to false when the gtk4 feature check fails.
Fixes: 4751bddd3f983af2 ("perf tools: Make GTK2 support opt-in")
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Following the approach used for pylint on standalone scripts and tests,
move the shellcheck and mypy checks in tools/perf/Build and
tools/perf/tests/Build, as well as the mypy and pylint checks in
tools/perf/util/Build and tools/perf/pmu-events/Build, into dedicated
phony targets invoked as top-level sub-makes from Makefile.perf.
Also add 'perf' to pylint's --ignored-modules (leaving perf module
type-checking to mypy via perf.pyi), since astroid's ImportlibFinder
only resolves .pyi stubs for package directories (__init__.pyi) rather
than single-file stubs like perf.pyi and does not introspect C
extensions by default. This removes the need for pylint to wait on
building $(LANG_BINDINGS).
This avoids blocking jevents code generation, archive creation
(libperf-util.a, libperf-test.a, libpmu-events.a), and final linking of
the perf binary and Python extension on linter execution.
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
optparse is deprecated since Python 3.2 and triggers a pylint
deprecated-module warning on newer versions of pylint.
Replace optparse.OptionParser in tools/perf/tests/shell/lib/attr.py with
argparse.ArgumentParser.
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
'perf probe' names the event after the function and binary, so concurrent
runs, as with 'perf test -r3', all add probe_testfile:foo and all but one
fail with 'event "foo" already exists'. Name the event foo_$$.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Now it's safe to run multiple 'perf trace' commands at the same time.
Let's make them non-exclusive so that they can run in parallel.
$ sudo perf test 'perf trace'
113: Check open filename arg using perf trace + vfs_getname : Skip
114: perf trace enum augmentation tests : Ok
115: perf trace BTF general tests : Ok
116: perf trace exit race : Ok
117: perf trace record and replay : Ok
118: perf trace summary : Ok
[ irogers: Keep trace+probe_vfs_getname.sh exclusive, as without BPF every
'perf trace' opens its probe:vfs_getname* probe. ]
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
Link: https://lore.kernel.org/r/20250814071754.193265-6-namhyung@kernel.org
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
With --max-events=1 an unrelated event can end perf trace before the
command's syscall is seen. Trace the whole command and grep for the
expected line, as trace_btf_enum.sh does, and remove the exclusive tag.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
On a mismatch print the command, the match count, the matching lines and
the end of the output.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Scope the uprobe name to the pid, retry adding and deleting it as
concurrent uprobe_events writes can fail with EBUSY, and delete it from
an exit trap. Create the temporary files with mktemp rather than reserving
names with mktemp -u. Remove the exclusive tag.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
The fixed vfs_getname probe name collides between parallel tests, and the
cleanup deletes every probe:vfs_getname* probe. Name the probe
getname_flags_$$, match it exactly, and remove it from an exit trap. Not
starting with vfs_getname also stops perf trace, which opens every
probe:vfs_getname* event, from pinning it.
Remove the exclusive tag from probe_vfs_getname.sh and
record+script_probe_vfs_getname.sh. trace+probe_vfs_getname.sh needs perf
trace to find its probe, so it uses vfs_getname_$$ and stays exclusive.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Writing 0 to events/enable disables every tracepoint, breaking perf
sessions running in parallel. Disable just the kprobes and uprobes, an
enabled one would make clearing kprobe_events or uprobe_events fail with
EBUSY.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
'perf trace' falls back to the raw_syscalls tracepoints, filtering in user
space, when the augmentation BPF programs can't be loaded.
Add --no-syscall-augment to use that path even when they can, so both
can be tested.
Suggest it when tasks are lost, as the tracepoints' per-task events
don't limit how many tasks are traced.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Now syscall init for augmented arguments is simplified. Let's get rid
of dead code.
Reviewed-by: Aaron Tomlin <atomlin@atomlin.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
Link: https://lore.kernel.org/r/20250814071754.193265-5-namhyung@kernel.org
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Howard reported that returning 0 from the BPF resulted in affecting
global syscall tracepoint handling. What we want to do is just to drop
syscall output in the current perf session. So we need a different
approach.
Currently perf trace uses bpf-output event for augmented arguments and
raw_syscalls:sys_{enter,exit} tracepoint events for normal arguments.
But I think we can just use bpf-output in both cases and drop the trace
point events.
Then it needs to distinguish bpf-output data if it's for enter or exit.
Repurpose struct trace_entry.type which is common in both syscall entry
and exit tracepoints.
[ irogers: Use the local syscall_{enter,exit}_args, always return 1 from
augmented__output(), output the unaugmented args when the perf_event_open
augmenter fails, dispatch from a trace__bpf_output() handler, add a
per-task tracking event and keep excluding kernel callchains. ]
Closes: https://lore.kernel.org/r/20250529065537.529937-1-howardchu95@gmail.com
Suggested-by: Howard Chu <howardchu95@gmail.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
Link: https://lore.kernel.org/r/20250814071754.193265-4-namhyung@kernel.org
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
The bpf-output event is opened per task without inherit, and
__augmented_syscalls__ holds each CPU's event for the first thread in
the thread map. bpf_perf_event_output() only writes to an event active
on the current CPU, so the target's other threads and its children
aren't augmented.
As BPF now picks the target's tasks, open the event system wide. A CPU
event isn't enabled by exec, so for a workload enable it when opened,
BPF waiting for the exec instead.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
The augmented syscalls BPF programs are attached to raw_syscalls system
wide, so they run for every task. To make the bpf-output event system
wide, and so augment the target's children, they need to know which
tasks are the target's.
Add a pids_to_trace map, seeded from the target's thread map, that
sys_enter and sys_exit check, leaving other tasks' events alone. tp_btf
programs add children on fork when inheriting, remove exited tasks and
follow a thread that exec gives the leader's pid. Tasks that don't fit
in the map, including when seeding, are counted and reported as lost.
Other forks drop any entry for the child's pid, seeded for a task that
had exited, such as a zombie. A workload waits for its exec, like
enable_on_exec.
When inheriting, unless given threads with -t, key the map by tgid,
like off-cpu's task_filter, so a process's threads share its entry. New
threads are then traced with their process, which is forgotten when its
last thread exits, and exec keeps the key.
The sched programs attach when the skeleton loads and the map is seeded
before the events are opened, so the forks and exits that follow are
seen. sys_enter and sys_exit now attach once the maps are populated,
rather than at load where the empty prog arrays made them veto every
raw_syscalls event.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
bpf-filter reads the tgid of each thread in a thread map from /proc, to
key its BPF map by process, and perf trace is about to do the same. Move
convert_to_tgid() to thread_map.c as thread_map__tgid(), taking the
thread map and index both users have.
While moving it, fix two bugs. Match Tgid: at the start of a line, as a
task can set its comm, shown on the Name: line before it, to contain
"Tgid: <n>". And check the character after the number before freeing the
buffer it points into, rather than after.
Fixes: eb1693b1150d4b99 ("perf bpf-filter: Split per-task filter use case")
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
If augmented_syscalls__prepare() fails to load the skeleton, 'perf trace'
falls back to unaugmented tracing but leaves the skeleton in place.
augmented_syscalls__set_filter_pids() then writes to maps that were
never created and fails with -ENOENT, so a system wide session, or one
using --filter-pids, exits with "Not enough memory to run!".
Destroy the skeleton on failure and clear the pointer, so the setters
do nothing.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
augmented_syscalls__setup_bpf_output() looks up the bpf-output event's
file descriptor using the CPU number, but the fd xyarray is indexed by
position in the CPU map. With -C, or with offline CPUs, the two differ
and the lookup finds another CPU's fd or NULL, leaving those CPUs
unaugmented.
Use the CPU map index.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
We want to handle syscall exit path differently so let's split the
unaugmented exit BPF program. Currently it does nothing (same as
sys_enter).
[ irogers: Use struct syscall_{enter,exit}_args and make
trace__find_syscall_bpf_prog() fall back to the exit program for
exits ]
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
Link: https://lore.kernel.org/r/20250814071754.193265-3-namhyung@kernel.org
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
syscall__augmented_args() treats any data after the syscall arguments as
augmented arguments. trace__sys_enter() avoids this for
raw_syscalls:sys_enter, but trace__fprintf_sys_enter() also handles
syscalls:sys_enter_* tracepoint events, whose trailing data since v6.19
holds the __data_loc strings of internal fields. Only the BPF output
event carries augmented arguments, so only look for them there.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
syscall__read_info() sizes sc->arg_fmt by the tracepoint's field count,
which includes the internal fields that follow the arguments, and only
then subtracts them from sc->nr_args.
syscall__alloc_arg_fmts() copies sc->fmt->arg[] for every entry, so with
more than 6 fields, for example renameat2 with its 2 path strings, it
reads beyond the end of that RAW_SYSCALL_ARGS_NUM sized array.
Count the arguments before allocating, and pass the array's size to
syscall_arg_fmt__init_array() so its walk stops before the internal
fields, rather than stepping beyond the array.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Since v6.19 syscall tracepoints may end with __data_loc char[] fields
holding user space strings, these aren't syscall arguments.
syscall__scnprintf_args() walks them anyway, printing the extra raw
syscall arguments at those indices under the internal field's name, and
trace__bpf_sys_enter_beauty_map() considers them when building the BPF
beauty map. The internal fields come after the arguments, so stop at the
first one.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|