| Age | Commit message (Collapse) | Author |
|
Update perf-script documentation to reflect standalone Python script
execution and the removal of embedded Python and Perl scripting:
- Remove documentation for the removed -g and -s options and legacy
record/report script wrapper modes in perf-script.txt.
- Remove references to perf-script-perl and delete obsolete
perf-script-perl.txt.
- Rewrite perf-script-python.txt to document writing and running
standalone Python scripts using the perf module.
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Refactor 'perf script' to launch standalone scripts directly via fork()
and execvp() and remove the legacy embedded Perl and Python scripting
engines:
- Remove the embedded Perl scripting engine
(util/scripting-engines/trace-event-perl.c), Perl scripts, bin
wrappers, and Trace-Util library
(scripts/perl/Perf-Trace-Util/), and script_perl.sh test.
- Remove libperl feature checks from Makefile.config, Makefile.perf,
builtin-check.c, and Documentation/perf-check.txt.
- Remove -g / --gen-script option and scripting_ops dispatch table from
builtin-script.c and trace-event-scripting.c.
- Hide the legacy -s / --script option and update script discovery in
find_script() and list_available_scripts() to prioritize the system
'python' directory over bare filenames in the current directory.
- Update the Python script shell tests (test_arm_coresight_disasm.sh,
test_task_analyzer.sh, and test_*_python.sh) to invoke scripts via
'perf script <script>' rather than running the Python interpreter
directly on the script file path.
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Replace the libpython feature test with a python-module feature test
checking for Python C extension build capability, and update feature
test references accordingly.
Remove references to the legacy scripts/python directory and install
standalone Python scripts (python/*.py), type stubs (python/perf.pyi),
and the compiled perf extension module (python/perf*.so, when built)
directly under the python directory in libexec. Update the TUI script
browser (ui/browsers/scripts.c) to discover standalone scripts from the
updated installation path.
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Remove embedded Python interpreter support (libpython) from perf, as
all Python scripts have been migrated to standalone scripts using the
perf Python extension module.
Changes include:
- Remove libpython detection and build flags from Makefile.config.
- Remove legacy Python script installation rules from Makefile.perf.
- Delete tools/perf/util/scripting-engines/trace-event-python.c and
tools/perf/scripts/python/Perf-Trace-Util/Context.c.
- Remove Python scripting engine registration from
trace-event-scripting.c.
- Remove libpython from the supported features list in builtin-check.c
and Documentation/perf-check.txt.
- Delete the legacy Python scripts and bin wrappers in
tools/perf/scripts/python/ and update shell tests.
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Port export-to-postgresql.py to a standalone script in
tools/perf/python/ using the perf module and libpq via ctypes.
Improvements compared to the legacy script:
- Remove the dependency on PySide/QtSql for database creation and DDL
execution by driving libpq directly via ctypes (PQconnectdb, PQexec,
PQputCopyData, PQputCopyEnd) and streaming binary PostgreSQL COPY
files in PostgresExporter, enabling export on headless servers without
Qt installed.
- Harden database connection and SQL identifier handling by rejecting
URI ('://') and connection-parameter ('=') injection strings in
setup_db() and connect(), escaping single quotes and backslashes in
connection strings, and quoting SQL identifiers (quote_ident).
- Support Intel PT and instruction trace export via
perf.session(itrace=...) and perf.call_return callbacks, reconstructing
relational call_paths and calls tables (including id=0 placeholder
rows and callfk/returnfk foreign keys) and unpacking synthesized PT
payloads (ptwrite, cbr, mwait, pwre, exstop, pwrx).
- Use a dedicated context_switch ID counter so context-switch exports
do not advance sample database IDs out of sync with call_return
references.
Update Documentation/db-export.txt and add a shell test
(test_export_to_postgresql_python.sh) to verify the standalone exporter.
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Port export-to-sqlite.py to a standalone script in tools/perf/python/
using the perf module and Python's standard library sqlite3 module.
Improvements compared to the legacy script:
- Remove the dependency on PySide/QtSql by using Python's built-in
sqlite3 module in DatabaseExporter, allowing SQLite export on headless
and minimal systems without Qt installed.
- Support Intel PT and hardware instruction trace export via
perf.session(itrace=...) and perf.call_return callbacks, reconstructing
relational call_paths and calls tables and decoding synthesized PT
payloads (ptwrite, cbr, mwait, pwre, exstop, pwrx).
- Export context_switches via perf.session's context_switch callback
and map thread and comm IDs to their relational database keys.
- Manage temporary staging files inside an isolated tempfile.mkdtemp()
directory with guaranteed cleanup in a finally block.
Update Documentation/db-export.txt and add a shell test
(test_export_to_sqlite_python.sh) to verify the standalone exporter.
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
The data type profiling per-sample stream keys cross-CPU contention on
sample->cpu, which without PERF_SAMPLE_CPU is the "no CPU info"
sentinel, making same-instance accesses from different cores
indistinguishable from same-CPU traffic. 'perf mem record' already
passes -d and -W to the record parser, add --sample-cpu and document it
in perf-mem(1).
Also grow rec_argv: the new entry overflows it on PMUs with separate
load and store events, as the space for the fixed arguments was not
reserved.
Reviewed-by: Ian Rogers <irogers@google.com>
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
The function view is currently TUI-only, so it cannot be used by builds
without SLANG support, when output is piped, or from a script. Add a
--function option that prints the fully expanded three-level hierarchy to
stdout.
Keep the stdio renderer in builtin-c2c.c and reuse the common function-view
model introduced by the merged series. Export only the coalescing-field
capability check from the model, preserving the util/UI boundary and
leaving the TUI object in libperf-ui.a.
Stop padding the final identity column in symbol_view_entry(). The generic
formatter pads non-final columns but deliberately leaves the final column
unpadded, avoiding trailing whitespace in function-view table rows. The TUI
remains unchanged because its browser clears the rest of each rendered row.
--function implies --stdio and is rejected together with --stats. Validate
the iaddr requirement before processing events. Return function-view build
failures from the report command, and preserve TUI browser errors when
converting the display helpers to return a status.
Committer notes:
root@x2:~# uname -r
7.2.5-200.fc44.x86_64
root@x2:~# rpm -q kernel-debuginfo
kernel-debuginfo-7.2.5-200.fc44.x86_64
root@x2:~# perf c2c record perf bench futex hash -t 4 -r 1 -s > /dev/null 2>&1
root@x2:~# perf c2c report --stdio --function | grep "Functions Table" -A25
Shared Data Functions Table
=================================================
#
# Cycles Store
# % count Function / Contending function / Cacheline
# ......... ....... ..............................
#
- 47.89% 703 - [k] _raw_spin_lock
322 - [k] _raw_spin_lock
24 0xffff8b6dcbadf040
23 0xffff8b6dcbadf2c0
23 0xffff8b6dcbadf080
23 0xffff8b6dcbadf380
22 0xffff8b6dcbadf280
22 0xffff8b6dcbadf140
22 0xffff8b6dcbadf340
21 0xffff8b6dcbadf300
21 0xffff8b6dcbadf400
20 0xffff8b6dcbadf1c0
19 0xffff8b6dcbadf3c0
19 0xffff8b6dcbadf180
19 0xffff8b6dcbadf0c0
17 0xffff8b6dcbadf100
16 0xffff8b6dcbadf240
11 0xffff8b6dcbadf200
205 - [k] futex_q_unlock
root@x2:~#
Reviewed-by: Tianyou Li <tianyou.li@intel.com>
Reviewed-by: Wangyang Guo <wangyang.guo@intel.com>
Signed-off-by: Jiebin Sun <jiebin.sun@intel.com>
Acked-by: Namhyung Kim <namhyung@kernel.org>
Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Cc: Dapeng Mi <dapeng1.mi@linux.intel.com>
Cc: Ian Rogers <irogers@google.com>
Cc: James Clark <james.clark@linaro.org>
Cc: Thomas Falcon <thomas.falcon@intel.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
The default coalesce fields dropped pid in favor of iaddr, but the man
page still documents the old pid,iaddr default. Update it to match the
command.
Fixes: 423701a0c8d754d9 ("perf c2c: Change the default coalesce setup")
Reviewed-by: Tianyou Li <tianyou.li@intel.com>
Reviewed-by: Wangyang Guo <wangyang.guo@intel.com>
Signed-off-by: Jiebin Sun <jiebin.sun@intel.com>
Acked-by: Namhyung Kim <namhyung@kernel.org>
Cc: Dapeng Mi <dapeng1.mi@linux.intel.com>
Cc: Ian Rogers <irogers@google.com>
Cc: James Clark <james.clark@linaro.org>
Cc: Jiri Olsa <jolsa@kernel.org>
Cc: Thomas Falcon <thomas.falcon@intel.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
On s390 the kernel uses s390 back chain instead of frame pointers for
stack tracing of user space since v6.7 commit aa44433ac4ee ("s390: add
USER_STACKTRACE support"). This is because frame pointers on s390
cannot be used for stack tracing. [1]
This requires user space to maintain a s390 back chain. For instance
user space to be built with compiler option '-mbackchain' (instead of
'-fno-omit-frame-pointer' used on other architectures, which should
better not be used on s390 [1]). Only few distributions and users build
user space with '-mbackchain'. Therefore '--call-graph fp' may in
practice not produce the expected results.
Commit ca76fb67ebdd ("perf evlist: Improve default event for s390")
changed the default '-g' option to 'dwarf' on s390 and added a warning
for s390 that wrongly claimed that "Framepointer unwinding lacks kernel
support". This warning resulted from a misinterpretation of comments
in s390 cpumsf_pmu_event_init().
The restriction in cpumsf_pmu_event_init() applies to callchain sampling
with the s390 CPU Measurement Sampling Facility (CPUMSF) hardware PMU.
CPUMSF events, such as 'cycles', do not support callchain sampling.
This is independent of whether perf uses the 'fp' or 'dwarf' call-graph
mode. The CPUMSF PMU provides samples collected asynchronously, making
it impossible to associate a callchain to the historic IPs. Therefore
callchains can only be used with software events on s390.
Remove the incorrect warning. The kernel supports the 'fp' call-graph
mode on s390 by walking the back chain. Do not replace it with a
hint to use 'dwarf' when the resulting callchain is incomplete. Such a
suggestion could imply that 'fp' is inherently inferior. The same
general limitation exists on other architectures when user space is not
built with frame pointers (i.e. '-fno-omit-frame-pointers').
Instead, warn on s390 when callchain sampling is requested with an event
provided by the CPUMCF or CPUMSF hardware PMUs, because that combination
is not supported.
Update the perf record documentation to state that 'dwarf' is the
default call-graph mode on s390 and that 'fp' uses back chain instead
of frame pointers on s390.
Note that '--call-graph fp' may also be useful for other applications,
such as OpenJDK maintaining a s390 back chain (does not require JVM
option '-XX:+PreserveFramePointer' on s390):
$ perf record --call-graph fp ... -- \
java -XX:+UnlockDiagnosticVMOptions -XX:+DumpPerfMapAtExit ...
[1]: s390: Stack tracing using Frame Pointer, Back Chain, and SFrame,
https://conf.gnu-tools-cauldron.org/opo25/talk/Y3CVHY/
Fixes: ca76fb67ebdd5e1a ("perf evlist: Improve default event for s390")
Reviewed-by: Thomas Richter <tmricht@linux.ibm.com>
Signed-off-by: Jens Remus <jremus@linux.ibm.com>
Acked-by: Ian Rogers <irogers@google.com>
Cc: Namhyung Kim <namhyung@kernel.org>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
task-clock/cpu-clock appear in perf stat's own example output but were
never explained in perf-stat.txt, unlike time elapsed/user/sys. This
came up unanswered on the list in 2013 and 2015 without docs ever being
added.
Reviewed-by: Ian Rogers <irogers@google.com>
Signed-off-by: Christian Melendez <chrismelnu@gmail.com>
Link: https://lore.kernel.org/all/arWi_dY4y15J2rqh@google.com ]
[ Removed the :u modifier from the task-clock example, per Namhyung request ]
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Memory ranges were created to track different types of memory, such as
persistent, high bandwidth, or CXL-attached, that may be present on a
system for performance monitoring and resource control purposes.
Memory range data are parsed from the ACPI MRRM table and exposed to
userspace tools via sysfs [1]. Memory range data is read from:
/sys/firmware/acpi/memory_ranges/rangeX
With the following attributes:
u64 base;
u64 length;
int node;
u8 local_region_id;
u8 remote_region_id;
Read memory range data from sysfs if present and save it in the header
of the perf data file under a new feature bit, HEADER_MEMORY_RANGES
(35). Memory range data can be viewed with the --header or --header-only
options of perf-report and perf-script.
Example output:
# memory ranges (nr 5):
# range0: [0x0000000000000000-0x00000000bfffffff], node = 0, local_region_id = 0, remote_region_id = 255
# range1: [0x0000000100000000-0x000000203fffffff], node = 0, local_region_id = 0, remote_region_id = 255
# range2: [0x0000008000000000-0x0000027fffffffff], node = -2, local_region_id = 1, remote_region_id = 255
# range3: [0x0000028000000000-0x0000047fffffffff], node = -2, local_region_id = 1, remote_region_id = 255
# range4: [0x0000048000000000-0x000004ffffffffff], node = -2, local_region_id = 1, remote_region_id = 255
[1]: https://lore.kernel.org/lkml/20250505173819.419271-1-tony.luck@intel.com/
Reviewed-by: Ian Rogers <irogers@google.com>
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: GitHub-Copilot:claude-opus-4-8
Assisted-by: Sashiko:gemini-3.1-pro-preview
Signed-off-by: Thomas Falcon <thomas.falcon@intel.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Add support for printing PERF_MEM_LVLNUM_L0 in perf mem report.
The L0 cache is newly added a small piece of cache which is the closest,
lowest-latency memory cache tied directly to the execution pipeline.
The table "Table 9-4. Data Source Field Encodings for Panther Cove and
Coyote Cove Microarchitectures" in the ISE doc chapter "9.2.1 Panther
Cove and Coyote Cove Microarchitectures Memory Auxiliary Field
Layout"[1] indicates the code "01H" means the "L0 Hit - Minimal latency
core cache hit. This request was satisfied by the L0 data cache."
[1]: https://www.intel.com/content/www/us/en/content-details/922690/intel-architecture-instruction-set-extensions-programming-reference.html
Assisted-by: Sashiko:gemini-3.1-pro-preview
Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
Reviewed-by: Ian Rogers <irogers@google.com>
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
Co-developed-by: Thomas Falcon <thomas.falcon@intel.com>
Signed-off-by: Thomas Falcon <thomas.falcon@intel.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Provide a core.hybrid-merge configuration option in .perfconfig to allow
enabling hybrid event aggregation by default, avoiding the need to pass
--hybrid-merge explicitly on every invocation.
The config value is a default rather than an explicit request, so with
--hierarchy it is ignored with a warning, while giving both --hierarchy
and --hybrid-merge on the command line remains an error. For the same
reason it is silently ignored when there are no events to merge, as
would be the case on any machine with a single core PMU, whereas an
explicit --hybrid-merge still warns. Record whether the option came from
the command line in symbol_conf so both can tell the two apart.
'perf stat' has its own merging options for counting, the sampling
core.hybrid-merge deliberately doesn't alter it and its --hybrid-merge
must still be given explicitly.
Signed-off-by: Ian Rogers <irogers@google.com>
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Add dynamic TUI hints prompting users to restart with --hybrid-merge
when heterogeneous core PMU events are populated in the sample view.
Signed-off-by: Ian Rogers <irogers@google.com>
Assisted-by: Antigravity:gemini-3.1-pro
Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Add informative text outlining the IPC imbalances associated with
merging cross-hybrid core events like cycles.
Signed-off-by: Ian Rogers <irogers@google.com>
Assisted-by: Antigravity:gemini-3.1-pro
Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Add --hybrid-merge to perf report and perf top. It merges the events a
wildcard expanded to across the core PMUs, and their histograms, so
that a symbol which ran on more than one kind of core is reported once
with the total rather than once per PMU.
evlist__merge_hybrid() links the events and evlist__merge_hists_hybrid()
links the histograms. There is nothing to merge on a machine with a
single core PMU, or when the events didn't come from a wildcard, so
warn in that case rather than quietly producing an unmerged report.
Merging collapses the per-PMU entries into one set, which doesn't
combine with the per-level breakdown of --hierarchy, so asking for both
is an error.
Signed-off-by: Ian Rogers <irogers@google.com>
Assisted-by: Antigravity:gemini-3.1-pro
Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Add support for the newly introduced SIMD register sampling format by
adding the following 5 functions:
uint64_t perf_intr_simd_reg_class_mask(uint16_t e_machine, bool pred);
uint64_t perf_user_simd_reg_class_mask(uint16_t e_machine, bool pred);
uint64_t perf_intr_simd_reg_class_bitmap_qwords(uint16_t e_machine, int reg_c,
uint16_t *qwords, bool pred);
uint64_t perf_user_simd_reg_class_bitmap_qwords(uint16_t e_machine, int reg_c,
uint16_t *qwords, bool pred);
const char *perf_simd_reg_class_name(uint16_t e_machine, int id, bool pred);
The perf_{intr|user}_simd_reg_class_mask() functions retrieve the bitmap
of kernel supported SIMD/PRED register classes on current platform for
intr-regs and user-regs sampling, such as OPMASK/XMM/YMM/ZMM on
x86 platforms.
The perf_{intr|user}_simd_reg_class_bitmap_qwords() functions retrieve
the bitmap and qwords length of a certain class of SIMD/PRED register
on current platform for intr-regs and user-regs sampling. For example,
for the XMM registers on x86 platforms, the returned bitmap is 0xffff
(XMM0 ~ XMM15) and the qwords length is 2 (128 bits for each XMM
register).
The perf_simd_reg_class_name() function gets the register class name for
a certain register class index.
Additionally, the function __parse_regs() is enhanced to support parsing
these newly introduced SIMD/PRED registers. Currently, each class of
register can only be sampled collectively; sampling a specific SIMD
register is not supported. For example, all XMM registers are sampled
together rather than sampling only XMM0.
When multiple overlapping register types, such as XMM and YMM, are
sampled simultaneously, only the superset (YMM registers) is sampled.
With this patch, all supported sampling registers on x86 platforms are
displayed as follows.
$perf record --intr-regs=?
available registers: AX BX CX DX SI DI BP SP IP FLAGS CS SS R8 R9 R10
R11 R12 R13 R14 R15 R16 R17 R18 R19 R20 R21 R22 R23 R24 R25 R26 R27 R28
R29 R30 R31 SSP XMM0-15 YMM0-15 ZMM0-31 OPMASK0-7
$perf record --user-regs=?
available registers: AX BX CX DX SI DI BP SP IP FLAGS CS SS R8 R9 R10
R11 R12 R13 R14 R15 R16 R17 R18 R19 R20 R21 R22 R23 R24 R25 R26 R27 R28
R29 R30 R31 SSP XMM0-15 YMM0-15 ZMM0-31 OPMASK0-7
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
Reviewed-by: Ian Rogers <irogers@google.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Port straight to GTK 4 rather than GTK 3, since GTK 4 is where new
development happens and GTK 3 is old itself now.
GTK 4 drops GtkContainer, GdkScreen, and the gtk_main()/
gtk_dialog_run() family perf's GTK UI relied on. Containers get
per-widget setters (gtk_box_append() and friends), monitor geometry
comes from GdkMonitor instead of GdkScreen, and the main and
error-dialog loops become explicit GMainLoops quit from the
"close-request" and "response" signals. Widgets are visible by default
now, so gtk_widget_show_all()/set_no_show_all() go away, and the
remaining gtk_widget_show()/gtk_widget_hide() calls become
gtk_widget_set_visible() (with a small wrapper where "response" needs
to pass gtk_widget_hide() as a callback, since it no longer exists as
a plain function).
gtk_ui_progress__finish() skips destroying a progress dialog that was
never created, since gtk_window_destroy() asserts on NULL where the old
widget destroy tolerated it. Two spots the GTK 2 to GTK 3 port had
missed (builtin-annotate.c, ui/gtk/setup.c still using
HAVE_GTK2_SUPPORT and gtk_main_quit()) are fixed to match.
Runtime fallout from the new signal-driven loops: the error dialog's
nested loop hung if the parent window closed
(GTK_DIALOG_DESTROY_WITH_PARENT destroys without emitting "response");
gtk_info_bar_get_content_area() is gone, breaking GTK_INFO_BAR_SUPPORT;
the progress dialog's static widget pointers dangled after a manual
close; perf_gtk__error() and the warning functions reused an exhausted
va_list when vasprintf() failed.
The error loop is tracked in a list instead of a single pointer, since
perf_gtk__error() can be called re-entrantly (the dialog isn't modal)
and a lone global leaked the outer loop when that happened. The list
is only ever touched from the main thread: perf_gtk__error() updates
it while handling a dialog, and SIGINT/SIGQUIT/SIGTERM are deferred to
a GLib source via g_unix_signal_add() rather than calling
perf_gtk__exit() straight out of a real signal handler, so quitting on
those signals is serialized with the list update instead of racing it
from signal-handler context. SIGSEGV/SIGFPE keep a real handler, since
they're synchronous faults with no "later" to defer to, but it's pared
down to reporting and reraising the default disposition
(perf_gtk__fatal_signal()): there's no safe way to run GTK/GLib code
from the faulting context. stdarg.h, stdio.h, and string.h are now
included explicitly where used (util.c, hists.c, annotate.c) rather
than relying on transitive includes, which musl doesn't guarantee.
The gtk4-infobar feature check is dropped: GtkInfoBar has existed
unconditionally since GTK 3.10, so the check can only ever pass, and it
was failing outright here anyway since gtk_info_bar_new() is deprecated
and the check treats deprecation warnings as errors.
HAVE_GTK_INFO_BAR_SUPPORT and its statusbar-only fallback go away; the
info bar is now built unconditionally.
Signed-off-by: Matt Turner <mattst88@gmail.com>
Link: https://lore.kernel.org/r/20260908-perf-gtk2-v8-1-e90d5d155f0d@gmail.com
[ Fixed up some patch fuzz ]
[ Removed gtk2 from FEATURES_DISPLAY, it is opt-in use 'make VF=1' to see if it was detected ]
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Add support for showing all the three possible per IP weights in
annotate. The weights are shown by defaults if any are non zero. This
is useful, especially with the new insn lat statistics, but also
for all the existing weights.
Add a hotkey to the interactive browser to turn them off (w), as well
as a perf annotate command line option.
The weights are stored unconditionally in the sym_hist_entry, which
will increase memory consumption somewhat.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: omp:GPT-5.6-Luna
Signed-off-by: Andi Kleen <ak@linux.intel.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Add a -W/--weight option to perf top to collect weights too. Useful with
follow on patches.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: omp:GPT-5.6-Luna
Signed-off-by: Andi Kleen <ak@linux.intel.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Modernize the -W / --weight description in the manpage to cover more
cases that are supported now.
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Signed-off-by: Andi Kleen <ak@linux.intel.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Describe the function view hierarchy (read-side function -> contending
writer function -> shared cachelines), the per-level indentation, and the
keys, with a worked example.
Document that reliable function attribution requires `iaddr` in
`--coalesce`, that the reader and writer may be the same function, and why
the coalesced function view cannot distinguish same-thread from
different-thread accesses in that case. Also document that verbose mode
includes code addresses in function rows.
Signed-off-by: Jiebin Sun <jiebin.sun@intel.com>
Reviewed-by: Tianyou Li <tianyou.li@intel.com>
Reviewed-by: Wangyang Guo <wangyang.guo@intel.com>
Reviewed-by: Ian Rogers <irogers@google.com>
Cc: Dapeng Mi <dapeng1.mi@linux.intel.com>
Cc: James Clark <james.clark@linaro.org>
Cc: Thomas Falcon <thomas.falcon@intel.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
While 'perf sched latency' reports task runtime and delay statistics
(average and maximum delay), it does not provide a visual representation
of how task wait times are distributed across latency ranges between
snapshots (start and finish of the analysis window).
The --histogram option collects CPU wait latencies (time between when
a task becomes runnable and when it gets scheduled onto a CPU) into 22
latency buckets, displaying an ASCII bar chart distribution.
The --hist-mode option configures the bucketing scheme:
- log (default). Logarithmic latency buckets ranging from
sub-microsecond (< 1 us) up to >= 1.05 seconds
- linear. Equal-width linear latency buckets
(i.e., 100 us steps up to >= 2.1 ms)
The --time option allows filtering trace event processing to a
specific time interval [start,stop].
Example histogram output excerpt:
❯ sudo perf sched latency --histogram --CPU 0
CPU Wait Latency Distribution Histogram (between snapshots) (total samples: 36114)
-------------------------------------------------------------------
Latency Range | Count | Pct | Histogram Graph
-------------------------------------------------------------------
< 1 us | 17 | 0.0% | #
2 - 4 us | 673 | 1.9% | #
4 - 8 us | 6237 | 17.3% | ######
8 - 16 us | 3224 | 8.9% | ###
16 - 32 us | 1388 | 3.8% | #
32 - 64 us | 709 | 2.0% | #
64 - 128 us | 690 | 1.9% | #
128 - 256 us | 789 | 2.2% | #
256 - 512 us | 541 | 1.5% | #
512 - 1024 us | 2256 | 6.2% | ##
1 - 2 ms | 3577 | 9.9% | ###
2 - 4 ms | 13259 | 36.7% | ##############
4 - 8 ms | 2523 | 7.0% | ##
8 - 16 ms | 222 | 0.6% | #
16 - 32 ms | 10 | 0.0% | #
>= 1.05 s | 3 | 0.0% | #
-------------------------------------------------------------------
Reviewed-by: Ian Rogers <irogers@google.com>
Signed-off-by: Aaron Tomlin <atomlin@atomlin.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
Introduce a new '--bitmask-list' command-line option for 'perf trace'.
When this option is specified, the formatting of cpumasks is delegated
to bitmap_scnprintf(), enabling cpumasks to be displayed as a condensed,
human-readable list (e.g., "0,2-5,7") instead of the default hexadecimal
representation. An example is provided below:
❯ sudo ./perf trace --show-cpu --bitmask-list --event ipi:ipi_send_cpumask --max-event 5
0.000 [000] Xorg/1434 ipi:ipi_send_cpumask(cpumask: 2-3,6, callsite: 0xffffffff9994f8e4, callback: 0xffffffff9994fdd0)
694.527 [002] chrome/2894 ipi:ipi_send_cpumask(cpumask: 1,3-5, callsite: 0xffffffff9994f8e4, callback: 0xffffffff9994fdd0)
2666.608 [003] Chrome_ChildIO/2948 ipi:ipi_send_cpumask(cpumask: 4,7, callsite: 0xffffffff9994f8e4, callback: 0xffffffff9994fdd0)
2673.638 [000] Chrome_IOThrea/2920 ipi:ipi_send_cpumask(cpumask: 2-5, callsite: 0xffffffff9994f8e4, callback: 0xffffffff9994fdd0)
2714.228 [005] chrome/3375 ipi:ipi_send_cpumask(cpumask: 0-4,6-7, callsite: 0xffffffff9994f8e4, callback: 0xffffffff9994fdd0)
Signed-off-by: Aaron Tomlin <atomlin@atomlin.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
When monitoring a large number of events (e.g., with wildcards such as
--event 'syscalls:sys_enter_*'), many matched events will return a count
of zero. This clutters the output, making it difficult to spot the
active events.
Add a new option --hide-zero-events to suppress printing events that
have a count of zero.
To prevent formatting and diagnostic issues, the zero-skipping logic
implements the following rules:
1. In metric-only mode (i.e., --metric-only), columns must remain
aligned in the output grid. We evaluate config->metric_only first
to avoid skipping zero-valued columns, preventing values from
shifting left and aligning under incorrect headers
2. For explicitly requested events, we ensure they are not silently
hidden if they are unsupported. We only hide a zero-count event
if counter->supported is true, ensuring that unsupported explicit
events still report "<not supported>"
Signed-off-by: Aaron Tomlin <atomlin@atomlin.com>
Reviewed-by: Ian Rogers <irogers@google.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
Use MAP_FAILED instead of NULL to detect mmap errors, and fix the
slots_p variable name typo in the sample code.
Signed-off-by: Hongfu Li <lihongfu@kylinos.cn>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
Add a CoreSight shell test for synthesized callchains.
The test uses the new callchain workload to generate trace and decodes
it with synthesis callchain. It then verifies that the instruction
samples show the expected callchain push and pop.
Use control FIFOs so tracing starts only around the workload, which
keeps the trace data small. The test is limited to with the cs_etm
event available and root permission.
After:
perf test 138 -vvv
138: CoreSight synthesized callchain:
---- start ----
test child forked, pid 35581
Callchain flow matched:
l1=4642868 l2=4642880 l3=4642895 l4=4642919 l5=4670494 l6=4670500 l7=4670520
---- end(0) ----
138: CoreSight synthesized callchain : Ok
Assisted-by: Codex:GPT-5.5
Reviewed-by: James Clark <james.clark@linaro.org>
Signed-off-by: Leo Yan <leo.yan@arm.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
Add documentation for recently added HEADER_E_MACHINE and
HEADER_CLN_SIZE data to the perf.data file. Also fix a typo
at the end of the header section.
Reviewed-by: Ian Rogers <irogers@google.com>
Signed-off-by: Thomas Falcon <thomas.falcon@intel.com>
Cc: Adrian Hunter <adrian.hunter@intel.com>
Cc: Alexander Shishkin <alexander.shishkin@linux.intel.com>
Cc: Dapeng Mi <dapeng1.mi@linux.intel.com>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: James Clark <james.clark@linaro.org>
Cc: Jiri Olsa <jolsa@kernel.org>
Cc: Mark Rutland <mark.rutland@arm.com>
Cc: Namhyung Kim <namhyung@kernel.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Add a workload that runs X threads that run a unique function named
"named_threads_thread[x]" which performs a multiplication in a loop for
Y loops. Each thread sets its name to "thread[x]".
This can be used to test that processor trace decoding handles
concurrent threads correctly and the correct symbols and thread names
are assigned to samples.
Signed-off-by: James Clark <james.clark@linaro.org>
Cc: Amir Ayupov <aaupov@meta.com>
Cc: Ian Rogers <irogers@google.com>
Cc: Jiri Olsa <jolsa@kernel.org>
Cc: Jonathan Corbet <corbet@lwn.net>
Cc: Leo Yan <leo.yan@arm.com>
Cc: Mike Leach <mike.leach@arm.com>
Cc: Namhyung Kim <namhyung@kernel.org>
Cc: Paschalis Mpeis <Paschalis.Mpeis@arm.com>
Cc: Shuah Khan <skhan@linuxfoundation.org>
Cc: Suzuki Poulouse <suzuki.poulose@arm.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Add a workload that does the same thing every time for testing CPU trace
decoding.
Reviewed-by: Leo Yan <leo.yan@arm.com>
Signed-off-by: James Clark <james.clark@linaro.org>
Cc: Amir Ayupov <aaupov@meta.com>
Cc: Ian Rogers <irogers@google.com>
Cc: Jiri Olsa <jolsa@kernel.org>
Cc: Jonathan Corbet <corbet@lwn.net>
Cc: Mike Leach <mike.leach@arm.com>
Cc: Namhyung Kim <namhyung@kernel.org>
Cc: Paschalis Mpeis <Paschalis.Mpeis@arm.com>
Cc: Shuah Khan <skhan@linuxfoundation.org>
Cc: Suzuki Poulouse <suzuki.poulose@arm.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
This workload launches two processes that block when reading and writing
to each other forcing the other process to be scheduled for each
read/write pair.
Signed-off-by: James Clark <james.clark@linaro.org>
Cc: Amir Ayupov <aaupov@meta.com>
Cc: Ian Rogers <irogers@google.com>
Cc: Jiri Olsa <jolsa@kernel.org>
Cc: Jonathan Corbet <corbet@lwn.net>
Cc: Leo Yan <leo.yan@arm.com>
Cc: Mike Leach <mike.leach@arm.com>
Cc: Namhyung Kim <namhyung@kernel.org>
Cc: Paschalis Mpeis <Paschalis.Mpeis@arm.com>
Cc: Shuah Khan <skhan@linuxfoundation.org>
Cc: Suzuki Poulouse <suzuki.poulose@arm.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Add a --workload-ctl=fifo:ctl-fifo[,ack-fifo] option for 'perf test
-w'. When set, run_workload() opens the named FIFO, writes enable before
invoking the builtin workload, writes disable before returning, and
waits for ack responses when an ack FIFO is provided to ensure that the
workload doesn't run until the events are enabled.
This can be used to limit the scope of the recording to only the
workload execution and avoid recording Perf setup and teardown code if
Perf record is started with events disabled (-D 1).
Assisted-by: Codex:GPT-5.5
Signed-off-by: James Clark <james.clark@linaro.org>
Cc: Amir Ayupov <aaupov@meta.com>
Cc: Ian Rogers <irogers@google.com>
Cc: Jiri Olsa <jolsa@kernel.org>
Cc: Jonathan Corbet <corbet@lwn.net>
Cc: Leo Yan <leo.yan@arm.com>
Cc: Mike Leach <mike.leach@arm.com>
Cc: Namhyung Kim <namhyung@kernel.org>
Cc: Paschalis Mpeis <Paschalis.Mpeis@arm.com>
Cc: Shuah Khan <skhan@linuxfoundation.org>
Cc: Suzuki Poulouse <suzuki.poulose@arm.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Include examples of:
o Privilege filter with Fetch and Op PMUs, including swfilt approach on
Zen5 and older platforms and hardware assisted filter on Zen6 and newer
platforms
o Streaming store filter with Op PMU
o Fetch latency filter with Fetch PMU
Signed-off-by: Ravi Bangoria <ravi.bangoria@amd.com>
Acked-by: Namhyung Kim <namhyung@kernel.org>
Cc: Ananth Narayan <ananth.narayan@amd.com>
Cc: Dapeng Mi <dapeng1.mi@linux.intel.com>
Cc: Ian Rogers <irogers@google.com>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: James Clark <james.clark@linaro.org>
Cc: Manali Shukla <manali.shukla@amd.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Sandipan Das <sandipan.das@amd.com>
Cc: Santosh Shukla <santosh.shukla@amd.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
When tracing system-wide workloads or specific events, it is highly
valuable to know exactly which CPU executed a specific event. Currently,
perf trace output defaults to omitting CPU information.
Introduce a new "--show-cpu" command-line option. When provided, this
flag extracts the CPU from the perf sample and prints it in a "[000]"
format immediately following the timestamp. This mirrors the behaviour of
other tracing tools like ftrace and perf script. For example:
# perf trace -e sched:sched_switch --max-events 5 --show-cpu
0.000 [002] :0/0 sched:sched_switch(prev_comm: "swapper/2", prev_prio: 120, next_comm: "rcu_preempt", next_pid: 16 (rcu_preempt), next_prio: 120)
0.009 [002] rcu_preempt/16 sched:sched_switch(prev_comm: "rcu_preempt", prev_pid: 16 (rcu_preempt), prev_prio: 120, prev_state: 128, next_comm: "swapper/2", next_prio: 120)
0.033 [002] :0/0 sched:sched_switch(prev_comm: "swapper/2", prev_prio: 120, next_comm: "kworker/u32:48", next_pid: 35840 (kworker/u32:48-), next_prio: 120)
0.041 [002] kworker/u32:48/35840 sched:sched_switch(prev_comm: "kworker/u32:48", prev_pid: 35840 (kworker/u32:48-), prev_prio: 120, prev_state: 128, next_comm: "swapper/2", next_prio: 120)
0.045 [002] :0/0 sched:sched_switch(prev_comm: "swapper/2", prev_prio: 120, next_comm: "kworker/u32:48", next_pid: 35840 (kworker/u32:48-), next_prio: 120)
The feature is implemented strictly as an opt-in toggle to prevent
cluttering the standard output and to preserve backwards compatibility
for scripts parsing the default output format.
Signed-off-by: Aaron Tomlin <atomlin@atomlin.com>
Cc: Adrian Hunter <adrian.hunter@intel.com>
Cc: Alexander Shishkin <alexander.shishkin@linux.intel.com>
Cc: Daniel Vacek <neelx@suse.com>
Cc: Howard Chu <howardchu95@gmail.com>
Cc: Ian Rogers <irogers@google.com>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: James Clark <james.clark@linaro.org>
Cc: Jiri Olsa <jolsa@kernel.org>
Cc: Mark Rutland <mark.rutland@arm.com>
Cc: Namhyung Kim <namhyung@kernel.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Sean Ashe <sean@ashe.io>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Update SIMD architecture and predicate flags.
Reviewed-by: James Clark <james.clark@linaro.org>
Reviewed-by: Ian Rogers <irogers@google.com>
Signed-off-by: Leo Yan <leo.yan@arm.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
Make symbol_conf::addr2line_disable_warn configurable by reading
the perfconfig file.
Use section core and addr2line-disable-warn = value.
Update documentation.
Example:
# perf config -l
core.addr2line-timeout=5000
core.addr2line-disable-warn=1
#
Signed-off-by: Thomas Richter <tmricht@linux.ibm.com>
Reviewed-by: Ian Rogers <irogers@google.com>
Suggested-by: Namhyung Kim <namhyung@kernel.org>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
This patch adds a new --pmu-filter option to perf-stat command to allow
filtering events on specific PMUs. This is useful when there are
multiple PMUs with same type (e.g. hisi_sicl2_cpa0 and hisi_sicl0_cpa0).
[root@localhost tmp]# perf stat -M cpa_p0_avg_bw
Performance counter stats for 'system wide':
19,417,779,115 hisi_sicl0_cpa0/cpa_cycles/ # 0.00 cpa_p0_avg_bw
0 hisi_sicl0_cpa0/cpa_p0_wr_dat/
0 hisi_sicl0_cpa0/cpa_p0_rd_dat_64b/
0 hisi_sicl0_cpa0/cpa_p0_rd_dat_32b/
19,417,751,103 hisi_sicl10_cpa0/cpa_cycles/ # 0.00 cpa_p0_avg_bw
0 hisi_sicl10_cpa0/cpa_p0_wr_dat/
0 hisi_sicl10_cpa0/cpa_p0_rd_dat_64b/
0 hisi_sicl10_cpa0/cpa_p0_rd_dat_32b/
19,417,730,679 hisi_sicl2_cpa0/cpa_cycles/ # 0.31 cpa_p0_avg_bw
75,635,749 hisi_sicl2_cpa0/cpa_p0_wr_dat/
18,520,640 hisi_sicl2_cpa0/cpa_p0_rd_dat_64b/
0 hisi_sicl2_cpa0/cpa_p0_rd_dat_32b/
19,417,674,227 hisi_sicl8_cpa0/cpa_cycles/ # 0.00 cpa_p0_avg_bw
0 hisi_sicl8_cpa0/cpa_p0_wr_dat/
0 hisi_sicl8_cpa0/cpa_p0_rd_dat_64b/
0 hisi_sicl8_cpa0/cpa_p0_rd_dat_32b/
19.417734480 seconds time elapsed
[root@localhost tmp]# perf stat --pmu-filter hisi_sicl2_cpa0 -M cpa_p0_avg_bw
Performance counter stats for 'system wide':
6,234,093,559 cpa_cycles # 0.60 cpa_p0_avg_bw
50,548,465 cpa_p0_wr_dat
7,552,182 cpa_p0_rd_dat_64b
0 cpa_p0_rd_dat_32b
6.234139320 seconds time elapsed
Signed-off-by: Qinxin Xia <xiaqinxin@huawei.com>
Reviewed-by: Ian Rogers <irogers@google.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
The "comm" column allows grouping events by the process command. It is
intended to group like programs, despite having different PIDs. But some
workloads may adjust their own command, so that a unique identifier
(e.g. a PID or some other numeric value) is part of the command name.
This destroys the utility of "comm", forcing perf to place each unique
process name into its own bucket, which can contribute to a
combinatorial explosion of memory use in perf report.
Create a less strict version of this column, which ignores digits when
comparing command names. Commands whose names are the same (ignoring
digits) are sorted into the same histogram buckets, and displayed with
the placeholder value "<N>" in the place of digits. For example,
hypothetical command names "kworker/1" "kworker/2" "kworker/3" would
sort into the same bucket and be represented as "kworker/<N>".
Committer testing:
$ perf report -s comm,comm_nodigit | grep -F "<N>"
0.01% CPU 6/TCG CPU <N>/TCG
0.01% kworker/53:2-mm kworker/<N>:<N>-mm
0.01% migration/24 migration/<N>
0.01% kworker/24:1-ev kworker/<N>:<N>-ev
0.01% llvmpipe-8 llvmpipe-<N>
Signed-off-by: Stephen Brennan <stephen.s.brennan@oracle.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
Add support for parsing an optional layout parameter in the --symfs
command line option. The format is:
--symfs <directory[,layout]>
Where layout can be:
- 'hierarchy': matches full path (default)
- 'flat': only matches base name
When debugging symbol files from a copy of the filesystem (e.g., from a
container or remote machine), the debug files are often stored in a
flat directory structure with only filenames, not the full original
paths. In this case, using 'flat' layout allows perf to find debug
symbols by matching only the filename rather than the full path.
For example, given a binary path like:
/build/output/lib/foo.so
With 'perf report --symfs /debug/files,flat', perf will look for:
/debug/files/foo.so
Instead of:
/debug/files/build/output/lib/foo.so
This is particularly useful when:
- Extracting debug files from containers with different directory layouts
- Working with build systems that flatten directory structures
Signed-off-by: Changbin Du <changbin.du@huawei.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
So that it can measure overhead of mmap_lock and/or per-VMA lock
contention.
$ perf bench mem mmap -f demand -l 1000 -t 1
# Running 'mem/mmap' benchmark:
# function 'demand' (Demand loaded mmap())
# Copying 1MB bytes ...
2.786858 GB/sec
$ perf bench mem mmap -f demand -l 1000 -t 2
# Running 'mem/mmap' benchmark:
# function 'demand' (Demand loaded mmap())
# Copying 1MB bytes ...
1.624468 GB/sec/thread ( +- 0.30% )
$ perf bench mem mmap -f demand -l 1000 -t 3
# Running 'mem/mmap' benchmark:
# function 'demand' (Demand loaded mmap())
# Copying 1MB bytes ...
1.493068 GB/sec/thread ( +- 0.15% )
$ perf bench mem mmap -f demand -l 1000 -t 4
# Running 'mem/mmap' benchmark:
# function 'demand' (Demand loaded mmap())
# Copying 1MB bytes ...
1.006087 GB/sec/thread ( +- 0.41% )
Reviewed-by: Ankur Arora <ankur.a.arora@oracle.com>
Reviewed-by: James Clark <james.clark@linaro.org>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/perf/perf-tools
Pull perf tools updates from Arnaldo Carvalho de Melo:
- Introduce 'perf sched stats' tool with record/report/diff workflows
using schedstat counters
- Add a faster libdw based addr2line implementation and allow selecting
it or its alternatives via 'perf config addr2line.style='
- Data-type profiling fixes and improvements including the ability to
select fields using 'perf report''s -F/-fields, e.g.:
'perf report --fields overhead,type'
- Add 'perf test' regression tests for Data-type profiling with C and
Rust workloads
- Fix srcline printing with inlines in callchains, make sure this has
coverage in 'perf test'
- Fix printing of leaf IP in LBR callchains
- Fix display of metrics without sufficient permission in 'perf stat'
- Print all machines in 'perf kvm report -vvv', not just the host
- Switch from SHA-1 to BLAKE2s for build ID generation, remove SHA-1
code
- Fix 'perf report's histogram entry collapsing with '-F' option
- Use system's cacheline size instead of a hardcoded value in 'perf
report'
- Allow filtering conversion by time range in 'perf data'
- Cover conversion to CTF using 'perf data' in 'perf test'
- Address newer glibc const-correctness (-Werror=discarded-qualifiers)
issues
- Fixes and improvements for ARM's CoreSight support, simplify ARM SPE
event config in 'perf mem', update docs for 'perf c2c' including the
ARM events it can be used with
- Build support for generating metrics from arch specific python
script, add extra AMD, Intel, ARM64 metrics using it
- Add AMD Zen 6 events and metrics
- Add JSON file with OpenHW Risc-V CVA6 hardware counters
- Add 'perf kvm' stats live testing
- Add more 'perf stat' tests to 'perf test'
- Fix segfault in `perf lock contention -b/--use-bpf`
- Fix various 'perf test' cases for s390
- Build system cleanups, bump minimum shellcheck version to 0.7.2
- Support building the capstone based annotation routines as a plugin
- Allow passing extra Clang flags via EXTRA_BPF_FLAGS
* tag 'perf-tools-for-v7.0-1-2026-02-21' of git://git.kernel.org/pub/scm/linux/kernel/git/perf/perf-tools: (255 commits)
perf test script: Add python script testing support
perf test script: Add perl script testing support
perf script: Allow the generated script to be a path
perf test: perf data --to-ctf testing
perf test: Test pipe mode with data conversion --to-json
perf json: Pipe mode --to-ctf support
perf json: Pipe mode --to-json support
perf check: Add libbabeltrace to the listed features
perf build: Allow passing extra Clang flags via EXTRA_BPF_FLAGS
perf test data_type_profiling.sh: Skip just the Rust tests if code_with_type workload is missing
tools build: Fix feature test for rust compiler
perf libunwind: Fix calls to thread__e_machine()
perf stat: Add no-affinity flag
perf evlist: Reduce affinity use and move into iterator, fix no affinity
perf evlist: Missing TPEBS close in evlist__close()
perf evlist: Special map propagation for tool events that read on 1 CPU
perf stat-shadow: In prepare_metric fix guard on reading NULL perf_stat_evsel
Revert "perf tool_pmu: More accurately set the cpus for tool events"
tools build: Emit dependencies file for test-rust.bin
tools build: Make test-rust.bin be removed by the 'clean' target
...
|
|
Allow the script generated by "perf script -g <language>" to be a file
path and the language determined by the file extension.
This is useful in testing so that the generated script file can be
written to a test directory.
Committer testing:
$ perf record ls a.a
ls: cannot access 'a.a': No such file or directory
[ perf record: Woken up 2 times to write data ]
[ perf record: Captured and wrote 0.003 MB perf.data (7 samples) ]
$ perf script -g python
generated Python script: perf-script.py
$ perf script -g myscript.py
generated Python script: myscript.py
$ diff -u perf-script.py myscript.py
$ tail myscript.py
def trace_unhandled(event_name, context, event_fields_dict, perf_sample_dict):
print(get_dict_as_string(event_fields_dict))
print('Sample: {'+get_dict_as_string(perf_sample_dict['sample'], ', ')+'}')
def print_header(event_name, cpu, secs, nsecs, pid, comm):
print("%-20s %5u %05u.%09u %8u %-20s " % \
(event_name, cpu, secs, nsecs, pid, comm), end="")
def get_dict_as_string(a_dict, delimiter=' '):
return delimiter.join(['%s=%s'%(k,str(v))for k,v in sorted(a_dict.items())])
$
Signed-off-by: Ian Rogers <irogers@google.com>
Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Cc: Adrian Hunter <adrian.hunter@intel.com>
Cc: Alexander Shishkin <alexander.shishkin@linux.intel.com>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: James Clark <james.clark@linaro.org>
Cc: Jiri Olsa <jolsa@kernel.org>
Cc: Leo Yan <leo.yan@arm.com>
Cc: Namhyung Kim <namhyung@kernel.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Sandipan Das <sandipan.das@amd.com>
Cc: Yujie Liu <yujie.liu@intel.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Add flag that disables affinity behavior.
Using sched_setaffinity() to place a perf thread on a CPU can avoid
certain interprocessor interrupts but may introduce a delay due to the
scheduling, particularly on loaded machines.
Add a command line option to disable the behavior.
This behavior is less present in other tools like `perf record`, as it
uses a ring buffer and doesn't make repeated system calls.
Signed-off-by: Ian Rogers <irogers@google.com>
Cc: Adrian Hunter <adrian.hunter@intel.com>
Cc: Alexander Shishkin <alexander.shishkin@linux.intel.com>
Cc: Andi Kleen <ak@linux.intel.com>
Cc: Andres Freund <andres@anarazel.de>
Cc: Dapeng Mi <dapeng1.mi@linux.intel.com>
Cc: Dr. David Alan Gilbert <linux@treblig.org>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: James Clark <james.clark@linaro.org>
Cc: Jiri Olsa <jolsa@kernel.org>
Cc: Namhyung Kim <namhyung@kernel.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Thomas Falcon <thomas.falcon@intel.com>
Cc: Thomas Richter <tmricht@linux.ibm.com>
Cc: Yang Li <yang.lee@linux.alibaba.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
FEAT_TME has been dropped from the architecture. Retrospectively.
I'm sure someone is crying somewhere, but most of us won't.
Clean-up time.
Reviewed-by: Fuad Tabba <tabba@google.com>
Tested-by: Fuad Tabba <tabba@google.com>
Link: https://patch.msgid.link/20260202184329.2724080-18-maz@kernel.org
Signed-off-by: Marc Zyngier <maz@kernel.org>
|
|
Fix the incorrect description of the schedstats report. Also fix the
spelling errors in man page.
Fixes: 800af362d68945e5 ("perf sched stats: Add details in man page")
Reviewed-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Reported-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Signed-off-by: Swapnil Sapkal <swapnil.sapkal@amd.com>
Cc: Adrian Hunter <adrian.hunter@intel.com>
Cc: Alexander Shishkin <alexander.shishkin@linux.intel.com>
Cc: Anubhav Shelat <ashelat@redhat.com>
Cc: Chen Yu <yu.c.chen@intel.com>
Cc: Gautham Shenoy <gautham.shenoy@amd.com>
Cc: Ian Rogers <irogers@google.com>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: James Clark <james.clark@arm.com>
Cc: Jiri Olsa <jolsa@kernel.org>
Cc: Mark Rutland <mark.rutland@arm.com>
Cc: Namhyung Kim <namhyung@kernel.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Ravi Bangoria <ravi.bangoria@amd.com>
Cc: Thomas Falcon <thomas.falcon@intel.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Document 'perf sched stats' purpose, usage examples and guide on how to
interpret the report data in the perf-sched man page.
Signed-off-by: Ravi Bangoria <ravi.bangoria@amd.com>
Signed-off-by: Swapnil Sapkal <swapnil.sapkal@amd.com>
Tested-by: Chen Yu <yu.c.chen@intel.com>
Acked-by: Ian Rogers <irogers@google.com>
Acked-by: Peter Zijlstra <peterz@infradead.org>
Cc: Adrian Hunter <adrian.hunter@intel.com>
Cc: Alexander Shishkin <alexander.shishkin@linux.intel.com>
Cc: Andi Kleen <ak@linux.intel.com>
Cc: Anubhav Shelat <ashelat@redhat.com>
Cc: Ben Gainey <ben.gainey@arm.com>
Cc: Blake Jones <blakejones@google.com>
Cc: Chun-Tse Shao <ctshao@google.com>
Cc: David Vernet <void@manifault.com>
Cc: Dmitriy Vyukov <dvyukov@google.com>
Cc: Dr. David Alan Gilbert <linux@treblig.org>
Cc: Gautham Shenoy <gautham.shenoy@amd.com>
Cc: Graham Woodward <graham.woodward@arm.com>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: James Clark <james.clark@arm.com>
Cc: Jiri Olsa <jolsa@kernel.org>
Cc: Juri Lelli <juri.lelli@redhat.com>
Cc: K Prateek Nayak <kprateek.nayak@amd.com>
Cc: Kan Liang <kan.liang@linux.intel.com>
Cc: Leo Yan <leo.yan@arm.com>
Cc: Madadi Vineeth Reddy <vineethr@linux.ibm.com>
Cc: Mark Rutland <mark.rutland@arm.com>
Cc: Namhyung Kim <namhyung@kernel.org>
Cc: Sandipan Das <sandipan.das@amd.com>
Cc: Santosh Shukla <santosh.shukla@amd.com>
Cc: Shrikanth Hegde <sshegde@linux.ibm.com>
Cc: Steven Rostedt (VMware) <rostedt@goodmis.org>
Cc: Tejun Heo <tj@kernel.org>
Cc: Thomas Falcon <thomas.falcon@intel.com>
Cc: Tim Chen <tim.c.chen@linux.intel.com>
Cc: Vincent Guittot <vincent.guittot@linaro.org>
Cc: Yang Jihong <yangjihong@bytedance.com>
Cc: Yujie Liu <yujie.liu@intel.com>
Cc: Zhongqiu Han <quic_zhonhan@quicinc.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
The '/proc/schedstat' file gives info about load balancing statistics
within a given domain.
It also contains the cpu_mask giving information about the sibling cpus
and domain names after schedstat version 17.
Storing this information in perf header will help tools like `perf sched
stats` for better analysis.
Signed-off-by: Swapnil Sapkal <swapnil.sapkal@amd.com>
Tested-by: Chen Yu <yu.c.chen@intel.com>
Acked-by: Ian Rogers <irogers@google.com>
Acked-by: Namhyung Kim <namhyung@kernel.org>
Acked-by: Peter Zijlstra <peterz@infradead.org>
Cc: Adrian Hunter <adrian.hunter@intel.com>
Cc: Alexander Shishkin <alexander.shishkin@linux.intel.com>
Cc: Andi Kleen <ak@linux.intel.com>
Cc: Anubhav Shelat <ashelat@redhat.com>
Cc: Ben Gainey <ben.gainey@arm.com>
Cc: Blake Jones <blakejones@google.com>
Cc: Chun-Tse Shao <ctshao@google.com>
Cc: David Vernet <void@manifault.com>
Cc: Dmitriy Vyukov <dvyukov@google.com>
Cc: Dr. David Alan Gilbert <linux@treblig.org>
Cc: Gautham Shenoy <gautham.shenoy@amd.com>
Cc: Graham Woodward <graham.woodward@arm.com>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: James Clark <james.clark@arm.com>
Cc: Jiri Olsa <jolsa@kernel.org>
Cc: Juri Lelli <juri.lelli@redhat.com>
Cc: K Prateek Nayak <kprateek.nayak@amd.com>
Cc: Kan Liang <kan.liang@linux.intel.com>
Cc: Leo Yan <leo.yan@arm.com>
Cc: Madadi Vineeth Reddy <vineethr@linux.ibm.com>
Cc: Mark Rutland <mark.rutland@arm.com>
Cc: Ravi Bangoria <ravi.bangoria@amd.com>
Cc: Sandipan Das <sandipan.das@amd.com>
Cc: Santosh Shukla <santosh.shukla@amd.com>
Cc: Shrikanth Hegde <sshegde@linux.ibm.com>
Cc: Steven Rostedt (VMware) <rostedt@goodmis.org>
Cc: Tejun Heo <tj@kernel.org>
Cc: Thomas Falcon <thomas.falcon@intel.com>
Cc: Tim Chen <tim.c.chen@linux.intel.com>
Cc: Vincent Guittot <vincent.guittot@linaro.org>
Cc: Yang Jihong <yangjihong@bytedance.com>
Cc: Yujie Liu <yujie.liu@intel.com>
Cc: Zhongqiu Han <quic_zhonhan@quicinc.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
Users may occasionally need to see which options are applied to memory
events.
This helps to understand the behavior of "perf c2c" and "perf mem", and
provides guidance for configuring memory event options directly.
Add a table to track memory events and their corresponding options, and
include the Arm SPE events in it.
Suggested-by: Al Grant <al.grant@arm.com>
Reviewed-by: James Clark <james.clark@linaro.org>
Signed-off-by: Leo Yan <leo.yan@arm.com>
Cc: Adrian Hunter <adrian.hunter@intel.com>
Cc: Ian Rogers <irogers@google.com>
Cc: Jiri Olsa <jolsa@kernel.org>
Cc: Mark Rutland <mark.rutland@arm.com>
Cc: Mike Leach <mike.leach@linaro.org>
Cc: Namhyung Kim <namhyung@kernel.org>
Cc: Will Deacon <will@kernel.org>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|
|
There are applications not built with frame pointers, so DWARF is needed
to get the stack traces.
`perf record --call-graph dwarf` saves the stack and register data for
each sample to get the stacktrace offline. But sometimes this data may
have sensitive information and we don't want to keep them in the file.
This new 'perf inject --convert-callchain' option creates the callchains
and discards the stack and register after that.
This saves storage space and processing time for the new data file.
Of course, users should remove the original data file to not keep
sensitive data around. :)
The down side is that it cannot handle inlined callchain entries as they
all have the same IPs.
Maybe we can add an option to 'perf report' to look up inlined functions
using DWARF - IIUC it doesn't require stack and register data.
This is an example.
$ perf record --call-graph dwarf -- perf test -w noploop
$ perf report --stdio --no-children --percent-limit=0 > output-prev
$ perf inject -i perf.data --convert-callchain -o perf.data.out
$ perf report --stdio --no-children --percent-limit=0 -i perf.data.out > output-next
$ diff -u output-prev output-next
...
0.23% perf ld-linux-x86-64.so.2 [.] _dl_relocate_object_no_relro
|
- ---elf_dynamic_do_Rela (inlined)
- _dl_relocate_object_no_relro
+ ---_dl_relocate_object_no_relro
_dl_relocate_object
dl_main
_dl_sysdep_start
- _dl_start_final (inlined)
_dl_start
_start
Reviewed-by: Ian Rogers <irogers@google.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
Cc: Adrian Hunter <adrian.hunter@intel.com>
Cc: Ingo Molnar <mingo@kernel.org>
Cc: James Clark <james.clark@linaro.org>
Cc: Jiri Olsa <jolsa@kernel.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
|