summaryrefslogtreecommitdiff
path: root/tools/perf/Documentation
AgeCommit message (Collapse)Author
5 daysperf Documentation: Update for standalone Python scriptsIan Rogers
Update perf-script documentation to reflect standalone Python script execution and the removal of embedded Python and Perl scripting: - Remove documentation for the removed -g and -s options and legacy record/report script wrapper modes in perf-script.txt. - Remove references to perf-script-perl and delete obsolete perf-script-perl.txt. - Rewrite perf-script-python.txt to document writing and running standalone Python scripts using the perf module. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers <irogers@google.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
5 daysperf script: Support standalone scripts and remove embedded scriptingIan Rogers
Refactor 'perf script' to launch standalone scripts directly via fork() and execvp() and remove the legacy embedded Perl and Python scripting engines: - Remove the embedded Perl scripting engine (util/scripting-engines/trace-event-perl.c), Perl scripts, bin wrappers, and Trace-Util library (scripts/perl/Perf-Trace-Util/), and script_perl.sh test. - Remove libperl feature checks from Makefile.config, Makefile.perf, builtin-check.c, and Documentation/perf-check.txt. - Remove -g / --gen-script option and scripting_ops dispatch table from builtin-script.c and trace-event-scripting.c. - Hide the legacy -s / --script option and update script discovery in find_script() and list_available_scripts() to prioritize the system 'python' directory over bare filenames in the current directory. - Update the Python script shell tests (test_arm_coresight_disasm.sh, test_task_analyzer.sh, and test_*_python.sh) to invoke scripts via 'perf script <script>' rather than running the Python interpreter directly on the script file path. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers <irogers@google.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
5 daysperf Makefile: Update Python script installation pathIan Rogers
Replace the libpython feature test with a python-module feature test checking for Python C extension build capability, and update feature test references accordingly. Remove references to the legacy scripts/python directory and install standalone Python scripts (python/*.py), type stubs (python/perf.pyi), and the compiled perf extension module (python/perf*.so, when built) directly under the python directory in libexec. Update the TUI script browser (ui/browsers/scripts.c) to discover standalone scripts from the updated installation path. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers <irogers@google.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
5 daysperf: Remove libpython support and legacy Python scriptsIan Rogers
Remove embedded Python interpreter support (libpython) from perf, as all Python scripts have been migrated to standalone scripts using the perf Python extension module. Changes include: - Remove libpython detection and build flags from Makefile.config. - Remove legacy Python script installation rules from Makefile.perf. - Delete tools/perf/util/scripting-engines/trace-event-python.c and tools/perf/scripts/python/Perf-Trace-Util/Context.c. - Remove Python scripting engine registration from trace-event-scripting.c. - Remove libpython from the supported features list in builtin-check.c and Documentation/perf-check.txt. - Delete the legacy Python scripts and bin wrappers in tools/perf/scripts/python/ and update shell tests. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers <irogers@google.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
5 daysperf python: Port export-to-postgresql to perf moduleIan Rogers
Port export-to-postgresql.py to a standalone script in tools/perf/python/ using the perf module and libpq via ctypes. Improvements compared to the legacy script: - Remove the dependency on PySide/QtSql for database creation and DDL execution by driving libpq directly via ctypes (PQconnectdb, PQexec, PQputCopyData, PQputCopyEnd) and streaming binary PostgreSQL COPY files in PostgresExporter, enabling export on headless servers without Qt installed. - Harden database connection and SQL identifier handling by rejecting URI ('://') and connection-parameter ('=') injection strings in setup_db() and connect(), escaping single quotes and backslashes in connection strings, and quoting SQL identifiers (quote_ident). - Support Intel PT and instruction trace export via perf.session(itrace=...) and perf.call_return callbacks, reconstructing relational call_paths and calls tables (including id=0 placeholder rows and callfk/returnfk foreign keys) and unpacking synthesized PT payloads (ptwrite, cbr, mwait, pwre, exstop, pwrx). - Use a dedicated context_switch ID counter so context-switch exports do not advance sample database IDs out of sync with call_return references. Update Documentation/db-export.txt and add a shell test (test_export_to_postgresql_python.sh) to verify the standalone exporter. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers <irogers@google.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
5 daysperf python: Port export-to-sqlite to perf moduleIan Rogers
Port export-to-sqlite.py to a standalone script in tools/perf/python/ using the perf module and Python's standard library sqlite3 module. Improvements compared to the legacy script: - Remove the dependency on PySide/QtSql by using Python's built-in sqlite3 module in DatabaseExporter, allowing SQLite export on headless and minimal systems without Qt installed. - Support Intel PT and hardware instruction trace export via perf.session(itrace=...) and perf.call_return callbacks, reconstructing relational call_paths and calls tables and decoding synthesized PT payloads (ptwrite, cbr, mwait, pwre, exstop, pwrx). - Export context_switches via perf.session's context_switch callback and map thread and comm IDs to their relational database keys. - Manage temporary staging files inside an isolated tempfile.mkdtemp() directory with guaranteed cleanup in a finally block. Update Documentation/db-export.txt and add a shell test (test_export_to_sqlite_python.sh) to verify the standalone exporter. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers <irogers@google.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
6 daysperf mem record: Request PERF_SAMPLE_CPU by defaultArnaldo Carvalho de Melo
The data type profiling per-sample stream keys cross-CPU contention on sample->cpu, which without PERF_SAMPLE_CPU is the "no CPU info" sentinel, making same-instance accesses from different cores indistinguishable from same-CPU traffic. 'perf mem record' already passes -d and -W to the record parser, add --sample-cpu and document it in perf-mem(1). Also grow rec_argv: the new entry overflows it on PMUs with separate load and store events, as the space for the fixed arguments was not reserved. Reviewed-by: Ian Rogers <irogers@google.com> Reviewed-by: Namhyung Kim <namhyung@kernel.org> Assisted-by: LLM Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
6 daysperf c2c: Add stdio support for the function viewJiebin Sun
The function view is currently TUI-only, so it cannot be used by builds without SLANG support, when output is piped, or from a script. Add a --function option that prints the fully expanded three-level hierarchy to stdout. Keep the stdio renderer in builtin-c2c.c and reuse the common function-view model introduced by the merged series. Export only the coalescing-field capability check from the model, preserving the util/UI boundary and leaving the TUI object in libperf-ui.a. Stop padding the final identity column in symbol_view_entry(). The generic formatter pads non-final columns but deliberately leaves the final column unpadded, avoiding trailing whitespace in function-view table rows. The TUI remains unchanged because its browser clears the rest of each rendered row. --function implies --stdio and is rejected together with --stats. Validate the iaddr requirement before processing events. Return function-view build failures from the report command, and preserve TUI browser errors when converting the display helpers to return a status. Committer notes: root@x2:~# uname -r 7.2.5-200.fc44.x86_64 root@x2:~# rpm -q kernel-debuginfo kernel-debuginfo-7.2.5-200.fc44.x86_64 root@x2:~# perf c2c record perf bench futex hash -t 4 -r 1 -s > /dev/null 2>&1 root@x2:~# perf c2c report --stdio --function | grep "Functions Table" -A25 Shared Data Functions Table ================================================= # # Cycles Store # % count Function / Contending function / Cacheline # ......... ....... .............................. # - 47.89% 703 - [k] _raw_spin_lock 322 - [k] _raw_spin_lock 24 0xffff8b6dcbadf040 23 0xffff8b6dcbadf2c0 23 0xffff8b6dcbadf080 23 0xffff8b6dcbadf380 22 0xffff8b6dcbadf280 22 0xffff8b6dcbadf140 22 0xffff8b6dcbadf340 21 0xffff8b6dcbadf300 21 0xffff8b6dcbadf400 20 0xffff8b6dcbadf1c0 19 0xffff8b6dcbadf3c0 19 0xffff8b6dcbadf180 19 0xffff8b6dcbadf0c0 17 0xffff8b6dcbadf100 16 0xffff8b6dcbadf240 11 0xffff8b6dcbadf200 205 - [k] futex_q_unlock root@x2:~# Reviewed-by: Tianyou Li <tianyou.li@intel.com> Reviewed-by: Wangyang Guo <wangyang.guo@intel.com> Signed-off-by: Jiebin Sun <jiebin.sun@intel.com> Acked-by: Namhyung Kim <namhyung@kernel.org> Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com> Cc: Dapeng Mi <dapeng1.mi@linux.intel.com> Cc: Ian Rogers <irogers@google.com> Cc: James Clark <james.clark@linaro.org> Cc: Thomas Falcon <thomas.falcon@intel.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
6 daysperf c2c: Fix documented default coalesce fieldsJiebin Sun
The default coalesce fields dropped pid in favor of iaddr, but the man page still documents the old pid,iaddr default. Update it to match the command. Fixes: 423701a0c8d754d9 ("perf c2c: Change the default coalesce setup") Reviewed-by: Tianyou Li <tianyou.li@intel.com> Reviewed-by: Wangyang Guo <wangyang.guo@intel.com> Signed-off-by: Jiebin Sun <jiebin.sun@intel.com> Acked-by: Namhyung Kim <namhyung@kernel.org> Cc: Dapeng Mi <dapeng1.mi@linux.intel.com> Cc: Ian Rogers <irogers@google.com> Cc: James Clark <james.clark@linaro.org> Cc: Jiri Olsa <jolsa@kernel.org> Cc: Thomas Falcon <thomas.falcon@intel.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
7 daysperf evsel: Improve callchain warning for s390Jens Remus
On s390 the kernel uses s390 back chain instead of frame pointers for stack tracing of user space since v6.7 commit aa44433ac4ee ("s390: add USER_STACKTRACE support"). This is because frame pointers on s390 cannot be used for stack tracing. [1] This requires user space to maintain a s390 back chain. For instance user space to be built with compiler option '-mbackchain' (instead of '-fno-omit-frame-pointer' used on other architectures, which should better not be used on s390 [1]). Only few distributions and users build user space with '-mbackchain'. Therefore '--call-graph fp' may in practice not produce the expected results. Commit ca76fb67ebdd ("perf evlist: Improve default event for s390") changed the default '-g' option to 'dwarf' on s390 and added a warning for s390 that wrongly claimed that "Framepointer unwinding lacks kernel support". This warning resulted from a misinterpretation of comments in s390 cpumsf_pmu_event_init(). The restriction in cpumsf_pmu_event_init() applies to callchain sampling with the s390 CPU Measurement Sampling Facility (CPUMSF) hardware PMU. CPUMSF events, such as 'cycles', do not support callchain sampling. This is independent of whether perf uses the 'fp' or 'dwarf' call-graph mode. The CPUMSF PMU provides samples collected asynchronously, making it impossible to associate a callchain to the historic IPs. Therefore callchains can only be used with software events on s390. Remove the incorrect warning. The kernel supports the 'fp' call-graph mode on s390 by walking the back chain. Do not replace it with a hint to use 'dwarf' when the resulting callchain is incomplete. Such a suggestion could imply that 'fp' is inherently inferior. The same general limitation exists on other architectures when user space is not built with frame pointers (i.e. '-fno-omit-frame-pointers'). Instead, warn on s390 when callchain sampling is requested with an event provided by the CPUMCF or CPUMSF hardware PMUs, because that combination is not supported. Update the perf record documentation to state that 'dwarf' is the default call-graph mode on s390 and that 'fp' uses back chain instead of frame pointers on s390. Note that '--call-graph fp' may also be useful for other applications, such as OpenJDK maintaining a s390 back chain (does not require JVM option '-XX:+PreserveFramePointer' on s390): $ perf record --call-graph fp ... -- \ java -XX:+UnlockDiagnosticVMOptions -XX:+DumpPerfMapAtExit ... [1]: s390: Stack tracing using Frame Pointer, Back Chain, and SFrame, https://conf.gnu-tools-cauldron.org/opo25/talk/Y3CVHY/ Fixes: ca76fb67ebdd5e1a ("perf evlist: Improve default event for s390") Reviewed-by: Thomas Richter <tmricht@linux.ibm.com> Signed-off-by: Jens Remus <jremus@linux.ibm.com> Acked-by: Ian Rogers <irogers@google.com> Cc: Namhyung Kim <namhyung@kernel.org> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
7 daysperf stat: Document task-clock/cpu-clock in TIMINGSChristian Melendez
task-clock/cpu-clock appear in perf stat's own example output but were never explained in perf-stat.txt, unlike time elapsed/user/sys. This came up unanswered on the list in 2013 and 2015 without docs ever being added. Reviewed-by: Ian Rogers <irogers@google.com> Signed-off-by: Christian Melendez <chrismelnu@gmail.com> Link: https://lore.kernel.org/all/arWi_dY4y15J2rqh@google.com ] [ Removed the :u modifier from the task-clock example, per Namhyung request ] Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
7 daysperf header: Support memory rangesThomas Falcon
Memory ranges were created to track different types of memory, such as persistent, high bandwidth, or CXL-attached, that may be present on a system for performance monitoring and resource control purposes. Memory range data are parsed from the ACPI MRRM table and exposed to userspace tools via sysfs [1]. Memory range data is read from: /sys/firmware/acpi/memory_ranges/rangeX With the following attributes: u64 base; u64 length; int node; u8 local_region_id; u8 remote_region_id; Read memory range data from sysfs if present and save it in the header of the perf data file under a new feature bit, HEADER_MEMORY_RANGES (35). Memory range data can be viewed with the --header or --header-only options of perf-report and perf-script. Example output: # memory ranges (nr 5): # range0: [0x0000000000000000-0x00000000bfffffff], node = 0, local_region_id = 0, remote_region_id = 255 # range1: [0x0000000100000000-0x000000203fffffff], node = 0, local_region_id = 0, remote_region_id = 255 # range2: [0x0000008000000000-0x0000027fffffffff], node = -2, local_region_id = 1, remote_region_id = 255 # range3: [0x0000028000000000-0x0000047fffffffff], node = -2, local_region_id = 1, remote_region_id = 255 # range4: [0x0000048000000000-0x000004ffffffffff], node = -2, local_region_id = 1, remote_region_id = 255 [1]: https://lore.kernel.org/lkml/20250505173819.419271-1-tony.luck@intel.com/ Reviewed-by: Ian Rogers <irogers@google.com> Reviewed-by: Namhyung Kim <namhyung@kernel.org> Assisted-by: GitHub-Copilot:claude-opus-4-8 Assisted-by: Sashiko:gemini-3.1-pro-preview Signed-off-by: Thomas Falcon <thomas.falcon@intel.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
7 daysperf mem: Add support for printing PERF_MEM_LVLNUM_L0Dapeng Mi
Add support for printing PERF_MEM_LVLNUM_L0 in perf mem report. The L0 cache is newly added a small piece of cache which is the closest, lowest-latency memory cache tied directly to the execution pipeline. The table "Table 9-4. Data Source Field Encodings for Panther Cove and Coyote Cove Microarchitectures" in the ISE doc chapter "9.2.1 Panther Cove and Coyote Cove Microarchitectures Memory Auxiliary Field Layout"[1] indicates the code "01H" means the "L0 Hit - Minimal latency core cache hit. This request was satisfied by the L0 data cache." [1]: https://www.intel.com/content/www/us/en/content-details/922690/intel-architecture-instruction-set-extensions-programming-reference.html Assisted-by: Sashiko:gemini-3.1-pro-preview Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Reviewed-by: Ian Rogers <irogers@google.com> Reviewed-by: Namhyung Kim <namhyung@kernel.org> Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Co-developed-by: Thomas Falcon <thomas.falcon@intel.com> Signed-off-by: Thomas Falcon <thomas.falcon@intel.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
13 daysperf config: Add core.hybrid-merge to configure event mergingIan Rogers
Provide a core.hybrid-merge configuration option in .perfconfig to allow enabling hybrid event aggregation by default, avoiding the need to pass --hybrid-merge explicitly on every invocation. The config value is a default rather than an explicit request, so with --hierarchy it is ignored with a warning, while giving both --hierarchy and --hybrid-merge on the command line remains an error. For the same reason it is silently ignored when there are no events to merge, as would be the case on any machine with a single core PMU, whereas an explicit --hybrid-merge still warns. Record whether the option came from the command line in symbol_conf so both can tell the two apart. 'perf stat' has its own merging options for counting, the sampling core.hybrid-merge deliberately doesn't alter it and its --hybrid-merge must still be given explicitly. Signed-off-by: Ian Rogers <irogers@google.com> Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
13 daysperf tools: Add TUI hints for --hybrid-mergeIan Rogers
Add dynamic TUI hints prompting users to restart with --hybrid-merge when heterogeneous core PMU events are populated in the sample view. Signed-off-by: Ian Rogers <irogers@google.com> Assisted-by: Antigravity:gemini-3.1-pro Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
13 daysperf Documentation: Add tip for hybrid event mergingIan Rogers
Add informative text outlining the IPC imbalances associated with merging cross-hybrid core events like cycles. Signed-off-by: Ian Rogers <irogers@google.com> Assisted-by: Antigravity:gemini-3.1-pro Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
13 daysperf tools: Expose opt-in --hybrid-mergeIan Rogers
Add --hybrid-merge to perf report and perf top. It merges the events a wildcard expanded to across the core PMUs, and their histograms, so that a symbol which ran on more than one kind of core is reported once with the total rather than once per PMU. evlist__merge_hybrid() links the events and evlist__merge_hists_hybrid() links the histograms. There is nothing to merge on a machine with a single core PMU, or when the events didn't come from a wildcard, so warn in that case rather than quietly producing an unmerged report. Merging collapses the per-PMU entries into one set, which doesn't combine with the per-level breakdown of --hierarchy, so asking for both is an error. Signed-off-by: Ian Rogers <irogers@google.com> Assisted-by: Antigravity:gemini-3.1-pro Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-09-14perf regs: Support x86 SIMD registers samplingDapeng Mi
Add support for the newly introduced SIMD register sampling format by adding the following 5 functions: uint64_t perf_intr_simd_reg_class_mask(uint16_t e_machine, bool pred); uint64_t perf_user_simd_reg_class_mask(uint16_t e_machine, bool pred); uint64_t perf_intr_simd_reg_class_bitmap_qwords(uint16_t e_machine, int reg_c, uint16_t *qwords, bool pred); uint64_t perf_user_simd_reg_class_bitmap_qwords(uint16_t e_machine, int reg_c, uint16_t *qwords, bool pred); const char *perf_simd_reg_class_name(uint16_t e_machine, int id, bool pred); The perf_{intr|user}_simd_reg_class_mask() functions retrieve the bitmap of kernel supported SIMD/PRED register classes on current platform for intr-regs and user-regs sampling, such as OPMASK/XMM/YMM/ZMM on x86 platforms. The perf_{intr|user}_simd_reg_class_bitmap_qwords() functions retrieve the bitmap and qwords length of a certain class of SIMD/PRED register on current platform for intr-regs and user-regs sampling. For example, for the XMM registers on x86 platforms, the returned bitmap is 0xffff (XMM0 ~ XMM15) and the qwords length is 2 (128 bits for each XMM register). The perf_simd_reg_class_name() function gets the register class name for a certain register class index. Additionally, the function __parse_regs() is enhanced to support parsing these newly introduced SIMD/PRED registers. Currently, each class of register can only be sampled collectively; sampling a specific SIMD register is not supported. For example, all XMM registers are sampled together rather than sampling only XMM0. When multiple overlapping register types, such as XMM and YMM, are sampled simultaneously, only the superset (YMM registers) is sampled. With this patch, all supported sampling registers on x86 platforms are displayed as follows. $perf record --intr-regs=? available registers: AX BX CX DX SI DI BP SP IP FLAGS CS SS R8 R9 R10 R11 R12 R13 R14 R15 R16 R17 R18 R19 R20 R21 R22 R23 R24 R25 R26 R27 R28 R29 R30 R31 SSP XMM0-15 YMM0-15 ZMM0-31 OPMASK0-7 $perf record --user-regs=? available registers: AX BX CX DX SI DI BP SP IP FLAGS CS SS R8 R9 R10 R11 R12 R13 R14 R15 R16 R17 R18 R19 R20 R21 R22 R23 R24 R25 R26 R27 R28 R29 R30 R31 SSP XMM0-15 YMM0-15 ZMM0-31 OPMASK0-7 Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Reviewed-by: Ian Rogers <irogers@google.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-09-09tools: Port perf ui from GTK 2 to GTK 4Matt Turner
Port straight to GTK 4 rather than GTK 3, since GTK 4 is where new development happens and GTK 3 is old itself now. GTK 4 drops GtkContainer, GdkScreen, and the gtk_main()/ gtk_dialog_run() family perf's GTK UI relied on. Containers get per-widget setters (gtk_box_append() and friends), monitor geometry comes from GdkMonitor instead of GdkScreen, and the main and error-dialog loops become explicit GMainLoops quit from the "close-request" and "response" signals. Widgets are visible by default now, so gtk_widget_show_all()/set_no_show_all() go away, and the remaining gtk_widget_show()/gtk_widget_hide() calls become gtk_widget_set_visible() (with a small wrapper where "response" needs to pass gtk_widget_hide() as a callback, since it no longer exists as a plain function). gtk_ui_progress__finish() skips destroying a progress dialog that was never created, since gtk_window_destroy() asserts on NULL where the old widget destroy tolerated it. Two spots the GTK 2 to GTK 3 port had missed (builtin-annotate.c, ui/gtk/setup.c still using HAVE_GTK2_SUPPORT and gtk_main_quit()) are fixed to match. Runtime fallout from the new signal-driven loops: the error dialog's nested loop hung if the parent window closed (GTK_DIALOG_DESTROY_WITH_PARENT destroys without emitting "response"); gtk_info_bar_get_content_area() is gone, breaking GTK_INFO_BAR_SUPPORT; the progress dialog's static widget pointers dangled after a manual close; perf_gtk__error() and the warning functions reused an exhausted va_list when vasprintf() failed. The error loop is tracked in a list instead of a single pointer, since perf_gtk__error() can be called re-entrantly (the dialog isn't modal) and a lone global leaked the outer loop when that happened. The list is only ever touched from the main thread: perf_gtk__error() updates it while handling a dialog, and SIGINT/SIGQUIT/SIGTERM are deferred to a GLib source via g_unix_signal_add() rather than calling perf_gtk__exit() straight out of a real signal handler, so quitting on those signals is serialized with the list update instead of racing it from signal-handler context. SIGSEGV/SIGFPE keep a real handler, since they're synchronous faults with no "later" to defer to, but it's pared down to reporting and reraising the default disposition (perf_gtk__fatal_signal()): there's no safe way to run GTK/GLib code from the faulting context. stdarg.h, stdio.h, and string.h are now included explicitly where used (util.c, hists.c, annotate.c) rather than relying on transitive includes, which musl doesn't guarantee. The gtk4-infobar feature check is dropped: GtkInfoBar has existed unconditionally since GTK 3.10, so the check can only ever pass, and it was failing outright here anyway since gtk_info_bar_new() is deprecated and the check treats deprecation warnings as errors. HAVE_GTK_INFO_BAR_SUPPORT and its statusbar-only fallback go away; the info bar is now built unconditionally. Signed-off-by: Matt Turner <mattst88@gmail.com> Link: https://lore.kernel.org/r/20260908-perf-gtk2-v8-1-e90d5d155f0d@gmail.com [ Fixed up some patch fuzz ] [ Removed gtk2 from FEATURES_DISPLAY, it is opt-in use 'make VF=1' to see if it was detected ] Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-09-04perf tools: Add support for displaying weights in annotateAndi Kleen
Add support for showing all the three possible per IP weights in annotate. The weights are shown by defaults if any are non zero. This is useful, especially with the new insn lat statistics, but also for all the existing weights. Add a hotkey to the interactive browser to turn them off (w), as well as a perf annotate command line option. The weights are stored unconditionally in the sym_hist_entry, which will increase memory consumption somewhat. Reviewed-by: Namhyung Kim <namhyung@kernel.org> Assisted-by: omp:GPT-5.6-Luna Signed-off-by: Andi Kleen <ak@linux.intel.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-09-04perf tools top: Add --weight optionAndi Kleen
Add a -W/--weight option to perf top to collect weights too. Useful with follow on patches. Reviewed-by: Namhyung Kim <namhyung@kernel.org> Assisted-by: omp:GPT-5.6-Luna Signed-off-by: Andi Kleen <ak@linux.intel.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-09-04perf tools record: Modernize -W man pageAndi Kleen
Modernize the -W / --weight description in the manpage to cover more cases that are supported now. Reviewed-by: Namhyung Kim <namhyung@kernel.org> Signed-off-by: Andi Kleen <ak@linux.intel.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-08-17perf c2c: document function view in perf-c2c man pageJiebin Sun
Describe the function view hierarchy (read-side function -> contending writer function -> shared cachelines), the per-level indentation, and the keys, with a worked example. Document that reliable function attribution requires `iaddr` in `--coalesce`, that the reader and writer may be the same function, and why the coalesced function view cannot distinguish same-thread from different-thread accesses in that case. Also document that verbose mode includes code addresses in function rows. Signed-off-by: Jiebin Sun <jiebin.sun@intel.com> Reviewed-by: Tianyou Li <tianyou.li@intel.com> Reviewed-by: Wangyang Guo <wangyang.guo@intel.com> Reviewed-by: Ian Rogers <irogers@google.com> Cc: Dapeng Mi <dapeng1.mi@linux.intel.com> Cc: James Clark <james.clark@linaro.org> Cc: Thomas Falcon <thomas.falcon@intel.com> Signed-off-by: Namhyung Kim <namhyung@kernel.org>
2026-08-07perf sched latency: Add histogram and time interval optionsAaron Tomlin
While 'perf sched latency' reports task runtime and delay statistics (average and maximum delay), it does not provide a visual representation of how task wait times are distributed across latency ranges between snapshots (start and finish of the analysis window). The --histogram option collects CPU wait latencies (time between when a task becomes runnable and when it gets scheduled onto a CPU) into 22 latency buckets, displaying an ASCII bar chart distribution. The --hist-mode option configures the bucketing scheme: - log (default). Logarithmic latency buckets ranging from sub-microsecond (< 1 us) up to >= 1.05 seconds - linear. Equal-width linear latency buckets (i.e., 100 us steps up to >= 2.1 ms) The --time option allows filtering trace event processing to a specific time interval [start,stop]. Example histogram output excerpt: ❯ sudo perf sched latency --histogram --CPU 0 CPU Wait Latency Distribution Histogram (between snapshots) (total samples: 36114) ------------------------------------------------------------------- Latency Range | Count | Pct | Histogram Graph ------------------------------------------------------------------- < 1 us | 17 | 0.0% | # 2 - 4 us | 673 | 1.9% | # 4 - 8 us | 6237 | 17.3% | ###### 8 - 16 us | 3224 | 8.9% | ### 16 - 32 us | 1388 | 3.8% | # 32 - 64 us | 709 | 2.0% | # 64 - 128 us | 690 | 1.9% | # 128 - 256 us | 789 | 2.2% | # 256 - 512 us | 541 | 1.5% | # 512 - 1024 us | 2256 | 6.2% | ## 1 - 2 ms | 3577 | 9.9% | ### 2 - 4 ms | 13259 | 36.7% | ############## 4 - 8 ms | 2523 | 7.0% | ## 8 - 16 ms | 222 | 0.6% | # 16 - 32 ms | 10 | 0.0% | # >= 1.05 s | 3 | 0.0% | # ------------------------------------------------------------------- Reviewed-by: Ian Rogers <irogers@google.com> Signed-off-by: Aaron Tomlin <atomlin@atomlin.com> Signed-off-by: Namhyung Kim <namhyung@kernel.org>
2026-07-23perf trace: Add --bitmask-list command-line optionAaron Tomlin
Introduce a new '--bitmask-list' command-line option for 'perf trace'. When this option is specified, the formatting of cpumasks is delegated to bitmap_scnprintf(), enabling cpumasks to be displayed as a condensed, human-readable list (e.g., "0,2-5,7") instead of the default hexadecimal representation. An example is provided below: ❯ sudo ./perf trace --show-cpu --bitmask-list --event ipi:ipi_send_cpumask --max-event 5 0.000 [000] Xorg/1434 ipi:ipi_send_cpumask(cpumask: 2-3,6, callsite: 0xffffffff9994f8e4, callback: 0xffffffff9994fdd0) 694.527 [002] chrome/2894 ipi:ipi_send_cpumask(cpumask: 1,3-5, callsite: 0xffffffff9994f8e4, callback: 0xffffffff9994fdd0) 2666.608 [003] Chrome_ChildIO/2948 ipi:ipi_send_cpumask(cpumask: 4,7, callsite: 0xffffffff9994f8e4, callback: 0xffffffff9994fdd0) 2673.638 [000] Chrome_IOThrea/2920 ipi:ipi_send_cpumask(cpumask: 2-5, callsite: 0xffffffff9994f8e4, callback: 0xffffffff9994fdd0) 2714.228 [005] chrome/3375 ipi:ipi_send_cpumask(cpumask: 0-4,6-7, callsite: 0xffffffff9994f8e4, callback: 0xffffffff9994fdd0) Signed-off-by: Aaron Tomlin <atomlin@atomlin.com> Signed-off-by: Namhyung Kim <namhyung@kernel.org>
2026-07-19perf stat: Add --hide-zero-events option to suppress zero-count eventsAaron Tomlin
When monitoring a large number of events (e.g., with wildcards such as --event 'syscalls:sys_enter_*'), many matched events will return a count of zero. This clutters the output, making it difficult to spot the active events. Add a new option --hide-zero-events to suppress printing events that have a count of zero. To prevent formatting and diagnostic issues, the zero-skipping logic implements the following rules: 1. In metric-only mode (i.e., --metric-only), columns must remain aligned in the output grid. We evaluate config->metric_only first to avoid skipping zero-valued columns, preventing values from shifting left and aligning under incorrect headers 2. For explicitly requested events, we ensure they are not silently hidden if they are unsupported. We only hide a zero-count event if counter->supported is true, ensuring that unsupported explicit events still report "<not supported>" Signed-off-by: Aaron Tomlin <atomlin@atomlin.com> Reviewed-by: Ian Rogers <irogers@google.com> Signed-off-by: Namhyung Kim <namhyung@kernel.org>
2026-07-15perf doc: Fix mmap failure checks in topdown exampleHongfu Li
Use MAP_FAILED instead of NULL to detect mmap errors, and fix the slots_p variable name typo in the sample code. Signed-off-by: Hongfu Li <lihongfu@kylinos.cn> Signed-off-by: Namhyung Kim <namhyung@kernel.org>
2026-07-03perf test: Add Arm CoreSight callchain testLeo Yan
Add a CoreSight shell test for synthesized callchains. The test uses the new callchain workload to generate trace and decodes it with synthesis callchain. It then verifies that the instruction samples show the expected callchain push and pop. Use control FIFOs so tracing starts only around the workload, which keeps the trace data small. The test is limited to with the cs_etm event available and root permission. After: perf test 138 -vvv 138: CoreSight synthesized callchain: ---- start ---- test child forked, pid 35581 Callchain flow matched: l1=4642868 l2=4642880 l3=4642895 l4=4642919 l5=4670494 l6=4670500 l7=4670520 ---- end(0) ---- 138: CoreSight synthesized callchain : Ok Assisted-by: Codex:GPT-5.5 Reviewed-by: James Clark <james.clark@linaro.org> Signed-off-by: Leo Yan <leo.yan@arm.com> Signed-off-by: Namhyung Kim <namhyung@kernel.org>
2026-06-15perf tools: Document recent additions to the perf.data file headerThomas Falcon
Add documentation for recently added HEADER_E_MACHINE and HEADER_CLN_SIZE data to the perf.data file. Also fix a typo at the end of the header section. Reviewed-by: Ian Rogers <irogers@google.com> Signed-off-by: Thomas Falcon <thomas.falcon@intel.com> Cc: Adrian Hunter <adrian.hunter@intel.com> Cc: Alexander Shishkin <alexander.shishkin@linux.intel.com> Cc: Dapeng Mi <dapeng1.mi@linux.intel.com> Cc: Ingo Molnar <mingo@redhat.com> Cc: James Clark <james.clark@linaro.org> Cc: Jiri Olsa <jolsa@kernel.org> Cc: Mark Rutland <mark.rutland@arm.com> Cc: Namhyung Kim <namhyung@kernel.org> Cc: Peter Zijlstra <peterz@infradead.org> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-06-10perf test: Add named_threads workloadJames Clark
Add a workload that runs X threads that run a unique function named "named_threads_thread[x]" which performs a multiplication in a loop for Y loops. Each thread sets its name to "thread[x]". This can be used to test that processor trace decoding handles concurrent threads correctly and the correct symbols and thread names are assigned to samples. Signed-off-by: James Clark <james.clark@linaro.org> Cc: Amir Ayupov <aaupov@meta.com> Cc: Ian Rogers <irogers@google.com> Cc: Jiri Olsa <jolsa@kernel.org> Cc: Jonathan Corbet <corbet@lwn.net> Cc: Leo Yan <leo.yan@arm.com> Cc: Mike Leach <mike.leach@arm.com> Cc: Namhyung Kim <namhyung@kernel.org> Cc: Paschalis Mpeis <Paschalis.Mpeis@arm.com> Cc: Shuah Khan <skhan@linuxfoundation.org> Cc: Suzuki Poulouse <suzuki.poulose@arm.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-06-10perf test: Add deterministic workloadJames Clark
Add a workload that does the same thing every time for testing CPU trace decoding. Reviewed-by: Leo Yan <leo.yan@arm.com> Signed-off-by: James Clark <james.clark@linaro.org> Cc: Amir Ayupov <aaupov@meta.com> Cc: Ian Rogers <irogers@google.com> Cc: Jiri Olsa <jolsa@kernel.org> Cc: Jonathan Corbet <corbet@lwn.net> Cc: Mike Leach <mike.leach@arm.com> Cc: Namhyung Kim <namhyung@kernel.org> Cc: Paschalis Mpeis <Paschalis.Mpeis@arm.com> Cc: Shuah Khan <skhan@linuxfoundation.org> Cc: Suzuki Poulouse <suzuki.poulose@arm.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-06-10perf test: Add a workload that forces context switchesJames Clark
This workload launches two processes that block when reading and writing to each other forcing the other process to be scheduled for each read/write pair. Signed-off-by: James Clark <james.clark@linaro.org> Cc: Amir Ayupov <aaupov@meta.com> Cc: Ian Rogers <irogers@google.com> Cc: Jiri Olsa <jolsa@kernel.org> Cc: Jonathan Corbet <corbet@lwn.net> Cc: Leo Yan <leo.yan@arm.com> Cc: Mike Leach <mike.leach@arm.com> Cc: Namhyung Kim <namhyung@kernel.org> Cc: Paschalis Mpeis <Paschalis.Mpeis@arm.com> Cc: Shuah Khan <skhan@linuxfoundation.org> Cc: Suzuki Poulouse <suzuki.poulose@arm.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-06-10perf test: Add workload-ctl optionJames Clark
Add a --workload-ctl=fifo:ctl-fifo[,ack-fifo] option for 'perf test -w'. When set, run_workload() opens the named FIFO, writes enable before invoking the builtin workload, writes disable before returning, and waits for ack responses when an ack FIFO is provided to ensure that the workload doesn't run until the events are enabled. This can be used to limit the scope of the recording to only the workload execution and avoid recording Perf setup and teardown code if Perf record is started with events disabled (-D 1). Assisted-by: Codex:GPT-5.5 Signed-off-by: James Clark <james.clark@linaro.org> Cc: Amir Ayupov <aaupov@meta.com> Cc: Ian Rogers <irogers@google.com> Cc: Jiri Olsa <jolsa@kernel.org> Cc: Jonathan Corbet <corbet@lwn.net> Cc: Leo Yan <leo.yan@arm.com> Cc: Mike Leach <mike.leach@arm.com> Cc: Namhyung Kim <namhyung@kernel.org> Cc: Paschalis Mpeis <Paschalis.Mpeis@arm.com> Cc: Shuah Khan <skhan@linuxfoundation.org> Cc: Suzuki Poulouse <suzuki.poulose@arm.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-05-22perf doc: Document new IBS capabilities in man pageRavi Bangoria
Include examples of: o Privilege filter with Fetch and Op PMUs, including swfilt approach on Zen5 and older platforms and hardware assisted filter on Zen6 and newer platforms o Streaming store filter with Op PMU o Fetch latency filter with Fetch PMU Signed-off-by: Ravi Bangoria <ravi.bangoria@amd.com> Acked-by: Namhyung Kim <namhyung@kernel.org> Cc: Ananth Narayan <ananth.narayan@amd.com> Cc: Dapeng Mi <dapeng1.mi@linux.intel.com> Cc: Ian Rogers <irogers@google.com> Cc: Ingo Molnar <mingo@redhat.com> Cc: James Clark <james.clark@linaro.org> Cc: Manali Shukla <manali.shukla@amd.com> Cc: Peter Zijlstra <peterz@infradead.org> Cc: Sandipan Das <sandipan.das@amd.com> Cc: Santosh Shukla <santosh.shukla@amd.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-05-16perf trace: Introduce --show-cpu option to display cpu idAaron Tomlin
When tracing system-wide workloads or specific events, it is highly valuable to know exactly which CPU executed a specific event. Currently, perf trace output defaults to omitting CPU information. Introduce a new "--show-cpu" command-line option. When provided, this flag extracts the CPU from the perf sample and prints it in a "[000]" format immediately following the timestamp. This mirrors the behaviour of other tracing tools like ftrace and perf script. For example: # perf trace -e sched:sched_switch --max-events 5 --show-cpu 0.000 [002] :0/0 sched:sched_switch(prev_comm: "swapper/2", prev_prio: 120, next_comm: "rcu_preempt", next_pid: 16 (rcu_preempt), next_prio: 120) 0.009 [002] rcu_preempt/16 sched:sched_switch(prev_comm: "rcu_preempt", prev_pid: 16 (rcu_preempt), prev_prio: 120, prev_state: 128, next_comm: "swapper/2", next_prio: 120) 0.033 [002] :0/0 sched:sched_switch(prev_comm: "swapper/2", prev_prio: 120, next_comm: "kworker/u32:48", next_pid: 35840 (kworker/u32:48-), next_prio: 120) 0.041 [002] kworker/u32:48/35840 sched:sched_switch(prev_comm: "kworker/u32:48", prev_pid: 35840 (kworker/u32:48-), prev_prio: 120, prev_state: 128, next_comm: "swapper/2", next_prio: 120) 0.045 [002] :0/0 sched:sched_switch(prev_comm: "swapper/2", prev_prio: 120, next_comm: "kworker/u32:48", next_pid: 35840 (kworker/u32:48-), next_prio: 120) The feature is implemented strictly as an opt-in toggle to prevent cluttering the standard output and to preserve backwards compatibility for scripts parsing the default output format. Signed-off-by: Aaron Tomlin <atomlin@atomlin.com> Cc: Adrian Hunter <adrian.hunter@intel.com> Cc: Alexander Shishkin <alexander.shishkin@linux.intel.com> Cc: Daniel Vacek <neelx@suse.com> Cc: Howard Chu <howardchu95@gmail.com> Cc: Ian Rogers <irogers@google.com> Cc: Ingo Molnar <mingo@redhat.com> Cc: James Clark <james.clark@linaro.org> Cc: Jiri Olsa <jolsa@kernel.org> Cc: Mark Rutland <mark.rutland@arm.com> Cc: Namhyung Kim <namhyung@kernel.org> Cc: Peter Zijlstra <peterz@infradead.org> Cc: Sean Ashe <sean@ashe.io> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-04-10perf report: Update document for SIMD flagsLeo Yan
Update SIMD architecture and predicate flags. Reviewed-by: James Clark <james.clark@linaro.org> Reviewed-by: Ian Rogers <irogers@google.com> Signed-off-by: Leo Yan <leo.yan@arm.com> Signed-off-by: Namhyung Kim <namhyung@kernel.org>
2026-04-08perf config: Make symbol_conf::addr2line_disable_warn configurableThomas Richter
Make symbol_conf::addr2line_disable_warn configurable by reading the perfconfig file. Use section core and addr2line-disable-warn = value. Update documentation. Example: # perf config -l core.addr2line-timeout=5000 core.addr2line-disable-warn=1 # Signed-off-by: Thomas Richter <tmricht@linux.ibm.com> Reviewed-by: Ian Rogers <irogers@google.com> Suggested-by: Namhyung Kim <namhyung@kernel.org> Signed-off-by: Namhyung Kim <namhyung@kernel.org>
2026-03-27perf tools: Add --pmu-filter option for filtering PMUsQinxin Xia
This patch adds a new --pmu-filter option to perf-stat command to allow filtering events on specific PMUs. This is useful when there are multiple PMUs with same type (e.g. hisi_sicl2_cpa0 and hisi_sicl0_cpa0). [root@localhost tmp]# perf stat -M cpa_p0_avg_bw Performance counter stats for 'system wide': 19,417,779,115 hisi_sicl0_cpa0/cpa_cycles/ # 0.00 cpa_p0_avg_bw 0 hisi_sicl0_cpa0/cpa_p0_wr_dat/ 0 hisi_sicl0_cpa0/cpa_p0_rd_dat_64b/ 0 hisi_sicl0_cpa0/cpa_p0_rd_dat_32b/ 19,417,751,103 hisi_sicl10_cpa0/cpa_cycles/ # 0.00 cpa_p0_avg_bw 0 hisi_sicl10_cpa0/cpa_p0_wr_dat/ 0 hisi_sicl10_cpa0/cpa_p0_rd_dat_64b/ 0 hisi_sicl10_cpa0/cpa_p0_rd_dat_32b/ 19,417,730,679 hisi_sicl2_cpa0/cpa_cycles/ # 0.31 cpa_p0_avg_bw 75,635,749 hisi_sicl2_cpa0/cpa_p0_wr_dat/ 18,520,640 hisi_sicl2_cpa0/cpa_p0_rd_dat_64b/ 0 hisi_sicl2_cpa0/cpa_p0_rd_dat_32b/ 19,417,674,227 hisi_sicl8_cpa0/cpa_cycles/ # 0.00 cpa_p0_avg_bw 0 hisi_sicl8_cpa0/cpa_p0_wr_dat/ 0 hisi_sicl8_cpa0/cpa_p0_rd_dat_64b/ 0 hisi_sicl8_cpa0/cpa_p0_rd_dat_32b/ 19.417734480 seconds time elapsed [root@localhost tmp]# perf stat --pmu-filter hisi_sicl2_cpa0 -M cpa_p0_avg_bw Performance counter stats for 'system wide': 6,234,093,559 cpa_cycles # 0.60 cpa_p0_avg_bw 50,548,465 cpa_p0_wr_dat 7,552,182 cpa_p0_rd_dat_64b 0 cpa_p0_rd_dat_32b 6.234139320 seconds time elapsed Signed-off-by: Qinxin Xia <xiaqinxin@huawei.com> Reviewed-by: Ian Rogers <irogers@google.com> Signed-off-by: Namhyung Kim <namhyung@kernel.org>
2026-03-26perf report: Add comm_nodigit sort keyStephen Brennan
The "comm" column allows grouping events by the process command. It is intended to group like programs, despite having different PIDs. But some workloads may adjust their own command, so that a unique identifier (e.g. a PID or some other numeric value) is part of the command name. This destroys the utility of "comm", forcing perf to place each unique process name into its own bucket, which can contribute to a combinatorial explosion of memory use in perf report. Create a less strict version of this column, which ignores digits when comparing command names. Commands whose names are the same (ignoring digits) are sorted into the same histogram buckets, and displayed with the placeholder value "<N>" in the place of digits. For example, hypothetical command names "kworker/1" "kworker/2" "kworker/3" would sort into the same bucket and be represented as "kworker/<N>". Committer testing: $ perf report -s comm,comm_nodigit | grep -F "<N>" 0.01% CPU 6/TCG CPU <N>/TCG 0.01% kworker/53:2-mm kworker/<N>:<N>-mm 0.01% migration/24 migration/<N> 0.01% kworker/24:1-ev kworker/<N>:<N>-ev 0.01% llvmpipe-8 llvmpipe-<N> Signed-off-by: Stephen Brennan <stephen.s.brennan@oracle.com> Signed-off-by: Namhyung Kim <namhyung@kernel.org>
2026-03-10perf tools: Add layout support for --symfs optionChangbin Du
Add support for parsing an optional layout parameter in the --symfs command line option. The format is: --symfs <directory[,layout]> Where layout can be: - 'hierarchy': matches full path (default) - 'flat': only matches base name When debugging symbol files from a copy of the filesystem (e.g., from a container or remote machine), the debug files are often stored in a flat directory structure with only filenames, not the full original paths. In this case, using 'flat' layout allows perf to find debug symbols by matching only the filename rather than the full path. For example, given a binary path like: /build/output/lib/foo.so With 'perf report --symfs /debug/files,flat', perf will look for: /debug/files/foo.so Instead of: /debug/files/build/output/lib/foo.so This is particularly useful when: - Extracting debug files from containers with different directory layouts - Working with build systems that flatten directory structures Signed-off-by: Changbin Du <changbin.du@huawei.com> Signed-off-by: Namhyung Kim <namhyung@kernel.org>
2026-02-26perf bench: Add -t/--threads option to perf bench mem mmapNamhyung Kim
So that it can measure overhead of mmap_lock and/or per-VMA lock contention. $ perf bench mem mmap -f demand -l 1000 -t 1 # Running 'mem/mmap' benchmark: # function 'demand' (Demand loaded mmap()) # Copying 1MB bytes ... 2.786858 GB/sec $ perf bench mem mmap -f demand -l 1000 -t 2 # Running 'mem/mmap' benchmark: # function 'demand' (Demand loaded mmap()) # Copying 1MB bytes ... 1.624468 GB/sec/thread ( +- 0.30% ) $ perf bench mem mmap -f demand -l 1000 -t 3 # Running 'mem/mmap' benchmark: # function 'demand' (Demand loaded mmap()) # Copying 1MB bytes ... 1.493068 GB/sec/thread ( +- 0.15% ) $ perf bench mem mmap -f demand -l 1000 -t 4 # Running 'mem/mmap' benchmark: # function 'demand' (Demand loaded mmap()) # Copying 1MB bytes ... 1.006087 GB/sec/thread ( +- 0.41% ) Reviewed-by: Ankur Arora <ankur.a.arora@oracle.com> Reviewed-by: James Clark <james.clark@linaro.org> Signed-off-by: Namhyung Kim <namhyung@kernel.org>
2026-02-21Merge tag 'perf-tools-for-v7.0-1-2026-02-21' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/perf/perf-tools Pull perf tools updates from Arnaldo Carvalho de Melo: - Introduce 'perf sched stats' tool with record/report/diff workflows using schedstat counters - Add a faster libdw based addr2line implementation and allow selecting it or its alternatives via 'perf config addr2line.style=' - Data-type profiling fixes and improvements including the ability to select fields using 'perf report''s -F/-fields, e.g.: 'perf report --fields overhead,type' - Add 'perf test' regression tests for Data-type profiling with C and Rust workloads - Fix srcline printing with inlines in callchains, make sure this has coverage in 'perf test' - Fix printing of leaf IP in LBR callchains - Fix display of metrics without sufficient permission in 'perf stat' - Print all machines in 'perf kvm report -vvv', not just the host - Switch from SHA-1 to BLAKE2s for build ID generation, remove SHA-1 code - Fix 'perf report's histogram entry collapsing with '-F' option - Use system's cacheline size instead of a hardcoded value in 'perf report' - Allow filtering conversion by time range in 'perf data' - Cover conversion to CTF using 'perf data' in 'perf test' - Address newer glibc const-correctness (-Werror=discarded-qualifiers) issues - Fixes and improvements for ARM's CoreSight support, simplify ARM SPE event config in 'perf mem', update docs for 'perf c2c' including the ARM events it can be used with - Build support for generating metrics from arch specific python script, add extra AMD, Intel, ARM64 metrics using it - Add AMD Zen 6 events and metrics - Add JSON file with OpenHW Risc-V CVA6 hardware counters - Add 'perf kvm' stats live testing - Add more 'perf stat' tests to 'perf test' - Fix segfault in `perf lock contention -b/--use-bpf` - Fix various 'perf test' cases for s390 - Build system cleanups, bump minimum shellcheck version to 0.7.2 - Support building the capstone based annotation routines as a plugin - Allow passing extra Clang flags via EXTRA_BPF_FLAGS * tag 'perf-tools-for-v7.0-1-2026-02-21' of git://git.kernel.org/pub/scm/linux/kernel/git/perf/perf-tools: (255 commits) perf test script: Add python script testing support perf test script: Add perl script testing support perf script: Allow the generated script to be a path perf test: perf data --to-ctf testing perf test: Test pipe mode with data conversion --to-json perf json: Pipe mode --to-ctf support perf json: Pipe mode --to-json support perf check: Add libbabeltrace to the listed features perf build: Allow passing extra Clang flags via EXTRA_BPF_FLAGS perf test data_type_profiling.sh: Skip just the Rust tests if code_with_type workload is missing tools build: Fix feature test for rust compiler perf libunwind: Fix calls to thread__e_machine() perf stat: Add no-affinity flag perf evlist: Reduce affinity use and move into iterator, fix no affinity perf evlist: Missing TPEBS close in evlist__close() perf evlist: Special map propagation for tool events that read on 1 CPU perf stat-shadow: In prepare_metric fix guard on reading NULL perf_stat_evsel Revert "perf tool_pmu: More accurately set the cpus for tool events" tools build: Emit dependencies file for test-rust.bin tools build: Make test-rust.bin be removed by the 'clean' target ...
2026-02-12perf script: Allow the generated script to be a pathIan Rogers
Allow the script generated by "perf script -g <language>" to be a file path and the language determined by the file extension. This is useful in testing so that the generated script file can be written to a test directory. Committer testing: $ perf record ls a.a ls: cannot access 'a.a': No such file or directory [ perf record: Woken up 2 times to write data ] [ perf record: Captured and wrote 0.003 MB perf.data (7 samples) ] $ perf script -g python generated Python script: perf-script.py $ perf script -g myscript.py generated Python script: myscript.py $ diff -u perf-script.py myscript.py $ tail myscript.py def trace_unhandled(event_name, context, event_fields_dict, perf_sample_dict): print(get_dict_as_string(event_fields_dict)) print('Sample: {'+get_dict_as_string(perf_sample_dict['sample'], ', ')+'}') def print_header(event_name, cpu, secs, nsecs, pid, comm): print("%-20s %5u %05u.%09u %8u %-20s " % \ (event_name, cpu, secs, nsecs, pid, comm), end="") def get_dict_as_string(a_dict, delimiter=' '): return delimiter.join(['%s=%s'%(k,str(v))for k,v in sorted(a_dict.items())]) $ Signed-off-by: Ian Rogers <irogers@google.com> Tested-by: Arnaldo Carvalho de Melo <acme@redhat.com> Cc: Adrian Hunter <adrian.hunter@intel.com> Cc: Alexander Shishkin <alexander.shishkin@linux.intel.com> Cc: Ingo Molnar <mingo@redhat.com> Cc: James Clark <james.clark@linaro.org> Cc: Jiri Olsa <jolsa@kernel.org> Cc: Leo Yan <leo.yan@arm.com> Cc: Namhyung Kim <namhyung@kernel.org> Cc: Peter Zijlstra <peterz@infradead.org> Cc: Sandipan Das <sandipan.das@amd.com> Cc: Yujie Liu <yujie.liu@intel.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-02-10perf stat: Add no-affinity flagIan Rogers
Add flag that disables affinity behavior. Using sched_setaffinity() to place a perf thread on a CPU can avoid certain interprocessor interrupts but may introduce a delay due to the scheduling, particularly on loaded machines. Add a command line option to disable the behavior. This behavior is less present in other tools like `perf record`, as it uses a ring buffer and doesn't make repeated system calls. Signed-off-by: Ian Rogers <irogers@google.com> Cc: Adrian Hunter <adrian.hunter@intel.com> Cc: Alexander Shishkin <alexander.shishkin@linux.intel.com> Cc: Andi Kleen <ak@linux.intel.com> Cc: Andres Freund <andres@anarazel.de> Cc: Dapeng Mi <dapeng1.mi@linux.intel.com> Cc: Dr. David Alan Gilbert <linux@treblig.org> Cc: Ingo Molnar <mingo@redhat.com> Cc: James Clark <james.clark@linaro.org> Cc: Jiri Olsa <jolsa@kernel.org> Cc: Namhyung Kim <namhyung@kernel.org> Cc: Peter Zijlstra <peterz@infradead.org> Cc: Thomas Falcon <thomas.falcon@intel.com> Cc: Thomas Richter <tmricht@linux.ibm.com> Cc: Yang Li <yang.lee@linux.alibaba.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-02-05KVM: arm64: Remove all traces of FEAT_TMEMarc Zyngier
FEAT_TME has been dropped from the architecture. Retrospectively. I'm sure someone is crying somewhere, but most of us won't. Clean-up time. Reviewed-by: Fuad Tabba <tabba@google.com> Tested-by: Fuad Tabba <tabba@google.com> Link: https://patch.msgid.link/20260202184329.2724080-18-maz@kernel.org Signed-off-by: Marc Zyngier <maz@kernel.org>
2026-01-28perf sched stats: Fixes in man pageSwapnil Sapkal
Fix the incorrect description of the schedstats report. Also fix the spelling errors in man page. Fixes: 800af362d68945e5 ("perf sched stats: Add details in man page") Reviewed-by: Shrikanth Hegde <sshegde@linux.ibm.com> Reported-by: Shrikanth Hegde <sshegde@linux.ibm.com> Signed-off-by: Swapnil Sapkal <swapnil.sapkal@amd.com> Cc: Adrian Hunter <adrian.hunter@intel.com> Cc: Alexander Shishkin <alexander.shishkin@linux.intel.com> Cc: Anubhav Shelat <ashelat@redhat.com> Cc: Chen Yu <yu.c.chen@intel.com> Cc: Gautham Shenoy <gautham.shenoy@amd.com> Cc: Ian Rogers <irogers@google.com> Cc: Ingo Molnar <mingo@redhat.com> Cc: James Clark <james.clark@arm.com> Cc: Jiri Olsa <jolsa@kernel.org> Cc: Mark Rutland <mark.rutland@arm.com> Cc: Namhyung Kim <namhyung@kernel.org> Cc: Peter Zijlstra <peterz@infradead.org> Cc: Ravi Bangoria <ravi.bangoria@amd.com> Cc: Thomas Falcon <thomas.falcon@intel.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-01-22perf sched stats: Add details in man pageSwapnil Sapkal
Document 'perf sched stats' purpose, usage examples and guide on how to interpret the report data in the perf-sched man page. Signed-off-by: Ravi Bangoria <ravi.bangoria@amd.com> Signed-off-by: Swapnil Sapkal <swapnil.sapkal@amd.com> Tested-by: Chen Yu <yu.c.chen@intel.com> Acked-by: Ian Rogers <irogers@google.com> Acked-by: Peter Zijlstra <peterz@infradead.org> Cc: Adrian Hunter <adrian.hunter@intel.com> Cc: Alexander Shishkin <alexander.shishkin@linux.intel.com> Cc: Andi Kleen <ak@linux.intel.com> Cc: Anubhav Shelat <ashelat@redhat.com> Cc: Ben Gainey <ben.gainey@arm.com> Cc: Blake Jones <blakejones@google.com> Cc: Chun-Tse Shao <ctshao@google.com> Cc: David Vernet <void@manifault.com> Cc: Dmitriy Vyukov <dvyukov@google.com> Cc: Dr. David Alan Gilbert <linux@treblig.org> Cc: Gautham Shenoy <gautham.shenoy@amd.com> Cc: Graham Woodward <graham.woodward@arm.com> Cc: Ingo Molnar <mingo@redhat.com> Cc: James Clark <james.clark@arm.com> Cc: Jiri Olsa <jolsa@kernel.org> Cc: Juri Lelli <juri.lelli@redhat.com> Cc: K Prateek Nayak <kprateek.nayak@amd.com> Cc: Kan Liang <kan.liang@linux.intel.com> Cc: Leo Yan <leo.yan@arm.com> Cc: Madadi Vineeth Reddy <vineethr@linux.ibm.com> Cc: Mark Rutland <mark.rutland@arm.com> Cc: Namhyung Kim <namhyung@kernel.org> Cc: Sandipan Das <sandipan.das@amd.com> Cc: Santosh Shukla <santosh.shukla@amd.com> Cc: Shrikanth Hegde <sshegde@linux.ibm.com> Cc: Steven Rostedt (VMware) <rostedt@goodmis.org> Cc: Tejun Heo <tj@kernel.org> Cc: Thomas Falcon <thomas.falcon@intel.com> Cc: Tim Chen <tim.c.chen@linux.intel.com> Cc: Vincent Guittot <vincent.guittot@linaro.org> Cc: Yang Jihong <yangjihong@bytedance.com> Cc: Yujie Liu <yujie.liu@intel.com> Cc: Zhongqiu Han <quic_zhonhan@quicinc.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-01-21perf header: Support CPU DOMAIN relation infoSwapnil Sapkal
The '/proc/schedstat' file gives info about load balancing statistics within a given domain. It also contains the cpu_mask giving information about the sibling cpus and domain names after schedstat version 17. Storing this information in perf header will help tools like `perf sched stats` for better analysis. Signed-off-by: Swapnil Sapkal <swapnil.sapkal@amd.com> Tested-by: Chen Yu <yu.c.chen@intel.com> Acked-by: Ian Rogers <irogers@google.com> Acked-by: Namhyung Kim <namhyung@kernel.org> Acked-by: Peter Zijlstra <peterz@infradead.org> Cc: Adrian Hunter <adrian.hunter@intel.com> Cc: Alexander Shishkin <alexander.shishkin@linux.intel.com> Cc: Andi Kleen <ak@linux.intel.com> Cc: Anubhav Shelat <ashelat@redhat.com> Cc: Ben Gainey <ben.gainey@arm.com> Cc: Blake Jones <blakejones@google.com> Cc: Chun-Tse Shao <ctshao@google.com> Cc: David Vernet <void@manifault.com> Cc: Dmitriy Vyukov <dvyukov@google.com> Cc: Dr. David Alan Gilbert <linux@treblig.org> Cc: Gautham Shenoy <gautham.shenoy@amd.com> Cc: Graham Woodward <graham.woodward@arm.com> Cc: Ingo Molnar <mingo@redhat.com> Cc: James Clark <james.clark@arm.com> Cc: Jiri Olsa <jolsa@kernel.org> Cc: Juri Lelli <juri.lelli@redhat.com> Cc: K Prateek Nayak <kprateek.nayak@amd.com> Cc: Kan Liang <kan.liang@linux.intel.com> Cc: Leo Yan <leo.yan@arm.com> Cc: Madadi Vineeth Reddy <vineethr@linux.ibm.com> Cc: Mark Rutland <mark.rutland@arm.com> Cc: Ravi Bangoria <ravi.bangoria@amd.com> Cc: Sandipan Das <sandipan.das@amd.com> Cc: Santosh Shukla <santosh.shukla@amd.com> Cc: Shrikanth Hegde <sshegde@linux.ibm.com> Cc: Steven Rostedt (VMware) <rostedt@goodmis.org> Cc: Tejun Heo <tj@kernel.org> Cc: Thomas Falcon <thomas.falcon@intel.com> Cc: Tim Chen <tim.c.chen@linux.intel.com> Cc: Vincent Guittot <vincent.guittot@linaro.org> Cc: Yang Jihong <yangjihong@bytedance.com> Cc: Yujie Liu <yujie.liu@intel.com> Cc: Zhongqiu Han <quic_zhonhan@quicinc.com> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-01-20perf c2c: Update documentation for adding memory event tableLeo Yan
Users may occasionally need to see which options are applied to memory events. This helps to understand the behavior of "perf c2c" and "perf mem", and provides guidance for configuring memory event options directly. Add a table to track memory events and their corresponding options, and include the Arm SPE events in it. Suggested-by: Al Grant <al.grant@arm.com> Reviewed-by: James Clark <james.clark@linaro.org> Signed-off-by: Leo Yan <leo.yan@arm.com> Cc: Adrian Hunter <adrian.hunter@intel.com> Cc: Ian Rogers <irogers@google.com> Cc: Jiri Olsa <jolsa@kernel.org> Cc: Mark Rutland <mark.rutland@arm.com> Cc: Mike Leach <mike.leach@linaro.org> Cc: Namhyung Kim <namhyung@kernel.org> Cc: Will Deacon <will@kernel.org> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2026-01-20perf inject: Add --convert-callchain optionNamhyung Kim
There are applications not built with frame pointers, so DWARF is needed to get the stack traces. `perf record --call-graph dwarf` saves the stack and register data for each sample to get the stacktrace offline. But sometimes this data may have sensitive information and we don't want to keep them in the file. This new 'perf inject --convert-callchain' option creates the callchains and discards the stack and register after that. This saves storage space and processing time for the new data file. Of course, users should remove the original data file to not keep sensitive data around. :) The down side is that it cannot handle inlined callchain entries as they all have the same IPs. Maybe we can add an option to 'perf report' to look up inlined functions using DWARF - IIUC it doesn't require stack and register data. This is an example. $ perf record --call-graph dwarf -- perf test -w noploop $ perf report --stdio --no-children --percent-limit=0 > output-prev $ perf inject -i perf.data --convert-callchain -o perf.data.out $ perf report --stdio --no-children --percent-limit=0 -i perf.data.out > output-next $ diff -u output-prev output-next ... 0.23% perf ld-linux-x86-64.so.2 [.] _dl_relocate_object_no_relro | - ---elf_dynamic_do_Rela (inlined) - _dl_relocate_object_no_relro + ---_dl_relocate_object_no_relro _dl_relocate_object dl_main _dl_sysdep_start - _dl_start_final (inlined) _dl_start _start Reviewed-by: Ian Rogers <irogers@google.com> Signed-off-by: Namhyung Kim <namhyung@kernel.org> Cc: Adrian Hunter <adrian.hunter@intel.com> Cc: Ingo Molnar <mingo@kernel.org> Cc: James Clark <james.clark@linaro.org> Cc: Jiri Olsa <jolsa@kernel.org> Cc: Peter Zijlstra <peterz@infradead.org> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>