diff options
| author | Yunzhao Li <yunzhao@cloudflare.com> | 2026-07-02 11:07:35 -0700 |
|---|---|---|
| committer | Andrew Morton <akpm@linux-foundation.org> | 2026-07-22 21:12:09 -0700 |
| commit | 06507d9bb8eea65cd09ec45c14c4dba5739f7936 (patch) | |
| tree | ebd84d38a10489292e325e09a318e2f7044f1b02 /include/acpi/actbl.h | |
| parent | c66b5233c512a6e4ad3e7678c2dfe26a9479172d (diff) | |
| download | linux-next-06507d9bb8eea65cd09ec45c14c4dba5739f7936.tar.gz linux-next-06507d9bb8eea65cd09ec45c14c4dba5739f7936.zip | |
mm/zswap: use ratelimited stats flush in zswap_shrinker_count()
zswap_shrinker_count() calls mem_cgroup_flush_stats(), which takes the
global cgroup rstat lock synchronously. On machines with many CPUs and
NUMA nodes, this creates severe lock contention in the kswapd reclaim
path:
- Multiple kswapd threads (one per NUMA node) run concurrently.
- do_shrink_slab() invokes zswap_shrinker_count() for each
memcg-aware shrinker pass.
- Each call flushes the full cgroup rstat hierarchy under the global
lock.
On AMD EPYC 9684X machines (96 cores, 192 threads, 12 NUMA nodes) running
production workloads with zswap enabled, perf shows 2.88% of kernel cycles
in osq_lock contention from this path:
2.88% [k] osq_lock
--__mutex_lock.constprop.0
--__cgroup_rstat_lock
--cgroup_rstat_flush_locked
--cgroup_rstat_flush
--zswap_shrinker_count
do_shrink_slab
shrink_slab
shrink_node
balance_pgdat
kswapd
84% of kswapd kernel cycles are spent in
shrink_slab -> zswap_shrinker_count -> cgroup_rstat_flush, not in actual
page reclaim (shrink_lruvec).
Controlled A/B on identical hardware and workload:
shrinker=Y: 2.88% osq_lock, memory PSI 1.58%
shrinker=N: 0.00% osq_lock, memory PSI 0.57%
eBPF-based rstat lock wait measurement across 8 production metals
confirms the contention splits cleanly along shrinker enablement:
shrinker=Y: 50-250x more contended lock acquisitions (248/s vs 1.1/s)
shrinker=N: baseline lock wait (0.0017 s/s vs 1.04 s/s)
zswap_shrinker_count() only produces a heuristic estimate, scaled by
compression ratio via mult_frac(). The actual writeback happens in
zswap_shrinker_scan(). Slightly stale stats are acceptable here.
Switch to mem_cgroup_flush_stats_ratelimited(), which only flushes if
the periodic 2-second flusher is one full cycle late. This matches the
approach already used in prepare_scan_control() (mm/vmscan.c) for the
same reclaim path.
After applying this patch, rstat flush latency and lock wait time on
shrinker=Y machines dropped to the same level as shrinker=N controls,
while the zswap shrinker continues to function (pool size remains
bounded under the max_pool_percent cap).
Previously discussed:
- Chengming Zhou (Dec 2023): rstat contention from
zswap_shrinker_count [1]
- Shakeel Butt (Aug 2024): zswap_shrinker_count still uses sync
flush [2]
- Yosry Ahmed (Aug 2024): suggested eliminating in-kernel
flushers [3]
- Jesper Dangaard Brouer (Sep 2024): cgroup/rstat V11 patch [4]
Link: https://lore.kernel.org/20260702180908.150136-1-yunzhao@cloudflare.com
Link: https://lore.kernel.org/linux-mm/20231206103935.3440502-1-zhouchengming@bytedance.com/ [1]
Link: https://lore.kernel.org/linux-mm/CALvZod7LFxLCxVpOFH8b2Ppm8T40HPGMKQwX_=NPCWB_mFW+oQ@mail.gmail.com/ [2]
Link: https://lore.kernel.org/linux-mm/CAJD7tkYvFyOSX+rP_FKGBhxvZiCDxtpsNp-c5CGOA-4Bq9oXSg@mail.gmail.com/ [3]
Link: https://lore.kernel.org/linux-mm/172616070094.2055617.17676042522679701515.stgit@firesoul/ [4]
Suggested-by: Jesper Dangaard Brouer <hawk@kernel.org>
Signed-off-by: Jesper Dangaard Brouer <hawk@kernel.org>
Signed-off-by: Yunzhao Li <yunzhao@cloudflare.com>
Tested-by: Yunzhao Li <yunzhao@cloudflare.com>
Acked-by: Johannes Weiner <hannes@cmpxchg.org>
Acked-by: Jesper Dangaard Brouer <hawk@kernel.org>
Acked-by: Nhat Pham <nphamcs@gmail.com>
Cc: Chengming Zhou <chengming.zhou@linux.dev>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Yosry Ahmed <yosry@kernel.org>
Cc: Yunzhao Li <yunzhao@cloudflare.com>
Cc: Sourav Panda <souravpanda@google.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Diffstat (limited to 'include/acpi/actbl.h')
0 files changed, 0 insertions, 0 deletions
