From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 5642CC001DF for ; Fri, 20 Oct 2023 19:14:49 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id CBDB3800AB; Fri, 20 Oct 2023 15:14:48 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id C48328008B; Fri, 20 Oct 2023 15:14:48 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id AC0FA800AB; Fri, 20 Oct 2023 15:14:48 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0015.hostedemail.com [216.40.44.15]) by kanga.kvack.org (Postfix) with ESMTP id 9378B8008B for ; Fri, 20 Oct 2023 15:14:48 -0400 (EDT) Received: from smtpin10.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay02.hostedemail.com (Postfix) with ESMTP id 517BE12062B for ; Fri, 20 Oct 2023 19:14:48 +0000 (UTC) X-FDA: 81366791856.10.01643AF Received: from mail-io1-f47.google.com (mail-io1-f47.google.com [209.85.166.47]) by imf18.hostedemail.com (Postfix) with ESMTP id 6E4F71C0003 for ; Fri, 20 Oct 2023 19:14:46 +0000 (UTC) Authentication-Results: imf18.hostedemail.com; dkim=pass header.d=gmail.com header.s=20230601 header.b="dFnnGul/"; spf=pass (imf18.hostedemail.com: domain of nphamcs@gmail.com designates 209.85.166.47 as permitted sender) smtp.mailfrom=nphamcs@gmail.com; dmarc=pass (policy=none) header.from=gmail.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1697829286; a=rsa-sha256; cv=none; b=dtIKks+VpsLJxjvErnkU8Vvhn02tfvkWlSRJpKEQ/69TiyMvOztfRWmhpY8l7ZSqrvxxhv +0xhcORZQXACcXX5ELp0CdYmdfMV4GbFWlzn5ksLqMTufWXl+xjHs1FpJHT26JzjFz9Z0W uLZa1k0/lz5elLsiC65GvOaRe3qGcd0= ARC-Authentication-Results: i=1; imf18.hostedemail.com; dkim=pass header.d=gmail.com header.s=20230601 header.b="dFnnGul/"; spf=pass (imf18.hostedemail.com: domain of nphamcs@gmail.com designates 209.85.166.47 as permitted sender) smtp.mailfrom=nphamcs@gmail.com; dmarc=pass (policy=none) header.from=gmail.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1697829286; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=JyKq5yoM9pUoUw1oiKGaKb63cLqQKx4ooc1oYR34zEw=; b=fl8wAx5FE9Z9N6CdmqdSfsMayIi9umZlgOPbPkUqtJIAVf/sAaqguuMO4diGZE7y2Ylb1O sNt5jF22hYUh2HAdUGFPOC4HeBZ32tkBrB+W+6yrdANZlxvcV+6mg4lVq+YLK+ozNR8TzG MbhBj77X/dkSyTE4zuwsjkZOqj5Ste0= Received: by mail-io1-f47.google.com with SMTP id ca18e2360f4ac-7a68b87b265so40387939f.2 for ; Fri, 20 Oct 2023 12:14:46 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1697829285; x=1698434085; darn=kvack.org; h=content-transfer-encoding:cc:to:subject:message-id:date:from :in-reply-to:references:mime-version:from:to:cc:subject:date :message-id:reply-to; bh=JyKq5yoM9pUoUw1oiKGaKb63cLqQKx4ooc1oYR34zEw=; b=dFnnGul/lyCRnsYDPKHkY5FYMbUKWdOyr8nALSmqKUCYmEXYY0nchsaNldh3Hbx9+I KfaPgEBUOiSrsCT/FUecMg4lWrqlUsJUjuWzk2HM07+QuYe2pSgwKicLZGgOPa088IeA UCCKf75dj3BJFwBZdEdpb6nyC/ZREQuc7L7mXUR1yyrIvp1WXh/2aaTPfTXafDu6fggb 0ysMpCIJWEwQkLmUTrYZhS8HQnkZHXs8UfTRW+fDDtB5l8Zxoi9xTweQGe8UHiN5ij08 2PEGDmovYU050OdLDwbRyFd/fCXX6x5R6dmePMd+CGsSw7eE4dmCITLUiXYIcwrEs5fB Flfg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1697829285; x=1698434085; h=content-transfer-encoding:cc:to:subject:message-id:date:from :in-reply-to:references:mime-version:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=JyKq5yoM9pUoUw1oiKGaKb63cLqQKx4ooc1oYR34zEw=; b=U34D/PbmQZWxAK/wKtSZbwT8WyRblEUvyJm85J/roveZzvq0zlXfQU5r6uP2nhxshf bqPGnPBpdOZ7hE4h3hGDEy0M6fMLyb+2ADnjTjMpz0NEo/fknF80SrvlvANAb9VN+RtD 9iDyqgOsVb1GlDlF59zIq60suKQuS8nxnw9rrWBNJV6k5gpmiIYHL8xHoG7psw7zv4wF ubM5y9c0pn3HXpt6BSuhI4n1OaAe1sSHpXPFe0atfgCuH6BA6NekTA85qkxoIaWw3YOV xhxchMoLHU2MqmFtjH58vMOudSpHqB0z6JosHnZFgtCUsnmMQO4Q73wEw1ZTok7JHnAf HYhg== X-Gm-Message-State: AOJu0YzBjpEPRdY+5fACWiAgKnPgPykaVUJ9duvZlwdg6MJdaPOM4+WZ SE1dHpAIszVOCeJLe3ijcXdf61Q+pgZ86bo2qOc= X-Google-Smtp-Source: AGHT+IEpMSGBe5RBnSQgcHDd5jjdz3/GpEHn6PknJAIAFApi6kSkYEgAfluRuEfFX8u/R2RQ+orIvlJMmmsYkJjKZdw= X-Received: by 2002:a05:6602:1696:b0:794:da97:d194 with SMTP id s22-20020a056602169600b00794da97d194mr3329373iow.19.1697829285265; Fri, 20 Oct 2023 12:14:45 -0700 (PDT) MIME-Version: 1.0 References: <20231017232152.2605440-1-nphamcs@gmail.com> <20231017232152.2605440-6-nphamcs@gmail.com> In-Reply-To: From: Nhat Pham Date: Fri, 20 Oct 2023 12:14:34 -0700 Message-ID: Subject: Re: [PATCH v3 5/5] zswap: shrinks zswap pool based on memory pressure To: Yosry Ahmed Cc: akpm@linux-foundation.org, hannes@cmpxchg.org, cerasuolodomenico@gmail.com, sjenning@redhat.com, ddstreet@ieee.org, vitaly.wool@konsulko.com, mhocko@kernel.org, roman.gushchin@linux.dev, shakeelb@google.com, muchun.song@linux.dev, linux-mm@kvack.org, kernel-team@meta.com, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, shuah@kernel.org Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Rspamd-Server: rspam08 X-Rspamd-Queue-Id: 6E4F71C0003 X-Stat-Signature: gt598njfj1myw5g5tg7ahxioxq7yrunc X-Rspam-User: X-HE-Tag: 1697829286-727711 X-HE-Meta: U2FsdGVkX1+IUXnilxGNfCSsn8peX4Rqcqm4TEWcSO5DTh3hTF1abvbAgS+rRzj7ttR2sMXFIYpzHr0X+tCPLdgkaoitst+uVRJIkJf5yocUXwar2FAssVB7N8lhlJY7hpjkISBVsf+gNFzdvP0V8VSfY7YswIx8xF+mOjhi5Y6Re5taOks+yjyPTkVym7KDbPlACGuI+F+AdKh52mmMPt3SpHE2EqooMEFJdZpPdcqFR0fSgauI7fKVOKbuwa7veATwump/206C4Md8Y4iasvV6LC53UoRpFhm6qJZ6rnOK8B6E5YQq1wVCG0nMfNgydWDp7iF6MaptWLWHmJrHNyvLHAJEEGC15DyZBvhNaK+YJZJ3EJxalse4Y9ct6fhdR37jMzKF3HHqs+6bFCwUU2pGuVssDJE98cjAHrHyMsj4TM/ck0raDN1zVff3X+HC3fefzIdObjad0FJVvu4rHVjjXDTN5yH0ze6YOa2+kQ29F/RZrtE4iSiV36RfrHNT83nZgzjb6ET58uMbE4989EI3FsIXXQCd8MMZq1zBABBwkQr6SdFnICUHG2jy/p5nHa5tET3A5x+rj7kM+b57QH80sRcqs1YwxMYY28xin1ddjMxRRzwDLaNMOl9LDx2ncWdajVuq8jkZMc7b/ZHQ5Wmo0+GRceogN2ixSZo42Lg5hxos2p9Iog0CfnnN/OXieZHBqEckCSstaoYNvOGKYLXU8lhysF+FosmGDlzCpDAVPfNN3iG8WHgmR/tzEiuEvO32wvnVqg+VGXIVPVT/SrAheRXvHP8QkAxXCmNBIxtds1BWiTMZNZhoz23AYTeRirphPftGXGN3qCulH5z4/zPHtXBM4qOiYog5UY4Yqap0sNCyyzmANktzfy6NCI5fQFiCWOwXwIqB1CEb217eRs2e2M1lXg6Kx134MJrQUvKL2qBFk3gQkTOpxdAPLWURQAuBpV3rj8Oca1KT4yR SIfh/w/y Cc+wDd4HDXZDtiCcG2EU3F3IEpthzsGZVByL0zNeA6WKh0vUIiQHvDR2LuYwAOAYnxZcPIvX3L0zdiE0yPCXZbuowCMk2r25bflyI5QPJSj4zs9B3pGoJwvZyvypSpOv0oAErdf3LykNdNWk1L1ag6ihCkLbeC2hJd7EXlrHxrdMnul+lZO0RWMb+MwjawItybCM6i80jN4R38/Wv/tOoRmRmN6d2oxbtF/qVnEGnj1Tu2dy+gTJkIlnoLQ25lo0BqHCOuhWvR5sQK/TGzp2Im2S5CZZWstYlK9qqx9cpRnQK2U7OCVSPK1m7WmoeQHjAw7UCnk9LHUzf1L8zTegtb3XgLFokSeV5shRwWzoNo+doH7FL8D/pAZz0g26o+wmJAXV5vtpAq8uq7eorlXFWXmbMAw== X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: On Wed, Oct 18, 2023 at 4:36=E2=80=AFPM Yosry Ahmed = wrote: > > On Tue, Oct 17, 2023 at 4:21=E2=80=AFPM Nhat Pham wro= te: > > > > Currently, we only shrink the zswap pool when the user-defined limit is > > hit. This means that if we set the limit too high, cold data that are > > unlikely to be used again will reside in the pool, wasting precious > > memory. It is hard to predict how much zswap space will be needed ahead > > of time, as this depends on the workload (specifically, on factors such > > as memory access patterns and compressibility of the memory pages). > > > > This patch implements a memcg- and NUMA-aware shrinker for zswap, that > > is initiated when there is memory pressure. The shrinker does not > > have any parameter that must be tuned by the user, and can be opted in > > or out on a per-memcg basis. > > > > Furthermore, to make it more robust for many workloads and prevent > > overshrinking (i.e evicting warm pages that might be refaulted into > > memory), we build in the following heuristics: > > > > * Estimate the number of warm pages residing in zswap, and attempt to > > protect this region of the zswap LRU. > > * Scale the number of freeable objects by an estimate of the memory > > saving factor. The better zswap compresses the data, the fewer pages > > we will evict to swap (as we will otherwise incur IO for relatively > > small memory saving). > > * During reclaim, if the shrinker encounters a page that is also being > > brought into memory, the shrinker will cautiously terminate its > > shrinking action, as this is a sign that it is touching the warmer > > region of the zswap LRU. > > I really hope someone with more familiarity with reclaim heuristics > makes sure this makes sense. > > > > > As a proof of concept, we ran the following synthetic benchmark: > > build the linux kernel in a memory-limited cgroup, and allocate some > > cold data in tmpfs to see if the shrinker could write them out and > > improved the overall performance. Depending on the amount of cold data > > generated, we observe from 14% to 35% reduction in kernel CPU time used > > in the kernel builds. > > > > Signed-off-by: Nhat Pham > > --- > > Documentation/admin-guide/mm/zswap.rst | 7 ++ > > include/linux/mmzone.h | 14 +++ > > mm/mmzone.c | 3 + > > mm/swap_state.c | 21 +++- > > mm/zswap.c | 161 +++++++++++++++++++++++-- > > 5 files changed, 196 insertions(+), 10 deletions(-) > > > > diff --git a/Documentation/admin-guide/mm/zswap.rst b/Documentation/adm= in-guide/mm/zswap.rst > > index 45b98390e938..522ae22ccb84 100644 > > --- a/Documentation/admin-guide/mm/zswap.rst > > +++ b/Documentation/admin-guide/mm/zswap.rst > > @@ -153,6 +153,13 @@ attribute, e. g.:: > > > > Setting this parameter to 100 will disable the hysteresis. > > > > +When there is a sizable amount of cold memory residing in the zswap po= ol, it > > +can be advantageous to proactively write these cold pages to swap and = reclaim > > +the memory for other use cases. By default, the zswap shrinker is disa= bled. > > +User can enable it as follows: > > + > > + echo Y > /sys/module/zswap/parameters/shrinker_enabled > > + > > A debugfs interface is provided for various statistic about pool size,= number > > of pages stored, same-value filled pages and various counters for the = reasons > > pages are rejected. > > diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h > > index 486587fcd27f..8947a1bfbe9c 100644 > > --- a/include/linux/mmzone.h > > +++ b/include/linux/mmzone.h > > @@ -637,6 +637,20 @@ struct lruvec { > > #ifdef CONFIG_MEMCG > > struct pglist_data *pgdat; > > #endif > > +#ifdef CONFIG_ZSWAP > > + /* > > + * Number of pages in zswap that should be protected from the s= hrinker. > > + * This number is an estimate of the following counts: > > + * > > + * a) Recent page faults. > > + * b) Recent insertion to the zswap LRU. This includes new zswa= p stores, > > + * as well as recent zswap LRU rotations. > > + * > > + * These pages are likely to be warm, and might incur IO if the= are written > > + * to swap. > > + */ > > + atomic_long_t nr_zswap_protected; > > +#endif > > Instead of the ifdef's all over the code, we can define > nr_zswap_protected inside a struct and helpers that > increment/initialize nr_zswap_protected in zswap.h, and have a single > ifdef there. All the code would be oblivious to the existence of > nr_zswap_protected. > > Something like: > > #ifdef CONFIG_ZSWAP > > struct zswap_lruvec_state { > /* insert large comment */ > atomic_long_t nr_zswap_protected; > }; > > static inline void zswap_lruvec_init(..) > { > atomic_long_set(&lruvec->nr_zswap_protected, 0); > } > > static inline void zswap_lruvec_swapin(..) > { > if (page) { > struct lruvec *lruvec =3D folio_lruvec(page_folio(page)); > > atomic_long_inc(&lruvec->nr_zswap_protected); > } > } > > #else > > /* empty struct and functions > > #endif > > > }; > > > > /* Isolate for asynchronous migration */ > > diff --git a/mm/mmzone.c b/mm/mmzone.c > > index 68e1511be12d..4137f3ac42cd 100644 > > --- a/mm/mmzone.c > > +++ b/mm/mmzone.c > > @@ -78,6 +78,9 @@ void lruvec_init(struct lruvec *lruvec) > > > > memset(lruvec, 0, sizeof(struct lruvec)); > > spin_lock_init(&lruvec->lru_lock); > > +#ifdef CONFIG_ZSWAP > > + atomic_long_set(&lruvec->nr_zswap_protected, 0); > > +#endif > > > > for_each_lru(lru) > > INIT_LIST_HEAD(&lruvec->lists[lru]); > > diff --git a/mm/swap_state.c b/mm/swap_state.c > > index 0356df52b06a..a60197b55a28 100644 > > --- a/mm/swap_state.c > > +++ b/mm/swap_state.c > > @@ -676,7 +676,15 @@ struct page *swap_cluster_readahead(swp_entry_t en= try, gfp_t gfp_mask, > > lru_add_drain(); /* Push any new pages onto the LRU now = */ > > skip: > > /* The page was likely read above, so no need for plugging here= */ > > - return read_swap_cache_async(entry, gfp_mask, vma, addr, NULL); > > + page =3D read_swap_cache_async(entry, gfp_mask, vma, addr, NULL= ); > > +#ifdef CONFIG_ZSWAP > > + if (page) { > > + struct lruvec *lruvec =3D folio_lruvec(page_folio(page)= ); > > + > > + atomic_long_inc(&lruvec->nr_zswap_protected); > > + } > > +#endif > > + return page; > > } > > > > int init_swap_address_space(unsigned int type, unsigned long nr_pages) > > @@ -843,8 +851,15 @@ static struct page *swap_vma_readahead(swp_entry_t= fentry, gfp_t gfp_mask, > > lru_add_drain(); > > skip: > > /* The page was likely read above, so no need for plugging here= */ > > - return read_swap_cache_async(fentry, gfp_mask, vma, vmf->addres= s, > > - NULL); > > + page =3D read_swap_cache_async(fentry, gfp_mask, vma, vmf->addr= ess, NULL); > > +#ifdef CONFIG_ZSWAP > > + if (page) { > > + struct lruvec *lruvec =3D folio_lruvec(page_folio(page)= ); > > + > > + atomic_long_inc(&lruvec->nr_zswap_protected); > > + } > > +#endif > > + return page; > > } > > > > /** > > diff --git a/mm/zswap.c b/mm/zswap.c > > index 15485427e3fa..1d1fe75a5237 100644 > > --- a/mm/zswap.c > > +++ b/mm/zswap.c > > @@ -145,6 +145,10 @@ module_param_named(exclusive_loads, zswap_exclusiv= e_loads_enabled, bool, 0644); > > /* Number of zpools in zswap_pool (empirically determined for scalabil= ity) */ > > #define ZSWAP_NR_ZPOOLS 32 > > > > +/* Enable/disable memory pressure-based shrinker. */ > > +static bool zswap_shrinker_enabled; > > +module_param_named(shrinker_enabled, zswap_shrinker_enabled, bool, 064= 4); > > + > > /********************************* > > * data structures > > **********************************/ > > @@ -174,6 +178,8 @@ struct zswap_pool { > > char tfm_name[CRYPTO_MAX_ALG_NAME]; > > struct list_lru list_lru; > > struct mem_cgroup *next_shrink; > > + struct shrinker *shrinker; > > + atomic_t nr_stored; > > }; > > > > /* > > @@ -272,17 +278,26 @@ static bool zswap_can_accept(void) > > DIV_ROUND_UP(zswap_pool_total_size, PAGE_SIZE); > > } > > > > +static u64 get_zswap_pool_size(struct zswap_pool *pool) > > +{ > > + u64 pool_size =3D 0; > > + int i; > > + > > + for (i =3D 0; i < ZSWAP_NR_ZPOOLS; i++) > > + pool_size +=3D zpool_get_total_size(pool->zpools[i]); > > + > > + return pool_size; > > +} > > + > > static void zswap_update_total_size(void) > > { > > struct zswap_pool *pool; > > u64 total =3D 0; > > - int i; > > > > rcu_read_lock(); > > > > list_for_each_entry_rcu(pool, &zswap_pools, list) > > - for (i =3D 0; i < ZSWAP_NR_ZPOOLS; i++) > > - total +=3D zpool_get_total_size(pool->zpools[i]= ); > > + total +=3D get_zswap_pool_size(pool); > > > > rcu_read_unlock(); > > > > @@ -326,8 +341,24 @@ static void zswap_entry_cache_free(struct zswap_en= try *entry) > > static bool zswap_lru_add(struct list_lru *list_lru, struct zswap_entr= y *entry) > > { > > struct mem_cgroup *memcg =3D get_mem_cgroup_from_entry(entry); > > - bool added =3D __list_lru_add(list_lru, &entry->lru, entry_to_n= id(entry), memcg); > > - > > + int nid =3D entry_to_nid(entry); > > + struct lruvec *lruvec =3D mem_cgroup_lruvec(memcg, NODE_DATA(ni= d)); > > + bool added =3D __list_lru_add(list_lru, &entry->lru, nid, memcg= ); > > + unsigned long lru_size, old, new; > > + > > + if (added) { > > + lru_size =3D list_lru_count_one(list_lru, entry_to_nid(= entry), memcg); > > + old =3D atomic_long_inc_return(&lruvec->nr_zswap_protec= ted); > > + > > + /* > > + * Decay to avoid overflow and adapt to changing worklo= ads. > > + * This is based on LRU reclaim cost decaying heuristic= s. > > + */ > > + do { > > + new =3D old > lru_size / 4 ? old / 2 : old; > > + } while ( > > + !atomic_long_try_cmpxchg(&lruvec->nr_zswap_prot= ected, &old, new)); > > + } > > mem_cgroup_put(memcg); > > return added; > > } > > @@ -427,6 +458,7 @@ static void zswap_free_entry(struct zswap_entry *en= try) > > else { > > zswap_lru_del(&entry->pool->list_lru, entry); > > zpool_free(zswap_find_zpool(entry), entry->handle); > > + atomic_dec(&entry->pool->nr_stored); > > zswap_pool_put(entry->pool); > > } > > zswap_entry_cache_free(entry); > > @@ -468,6 +500,93 @@ static struct zswap_entry *zswap_entry_find_get(st= ruct rb_root *root, > > return entry; > > } > > > > +/********************************* > > +* shrinker functions > > +**********************************/ > > +static enum lru_status shrink_memcg_cb(struct list_head *item, struct = list_lru_one *l, > > + spinlock_t *lock, void *arg); > > + > > +static unsigned long zswap_shrinker_scan(struct shrinker *shrinker, > > + struct shrink_control *sc) > > +{ > > + struct lruvec *lruvec =3D mem_cgroup_lruvec(sc->memcg, NODE_DAT= A(sc->nid)); > > + unsigned long shrink_ret, nr_protected, lru_size; > > + struct zswap_pool *pool =3D shrinker->private_data; > > + bool encountered_page_in_swapcache =3D false; > > + > > + nr_protected =3D atomic_long_read(&lruvec->nr_zswap_protected); > > + lru_size =3D list_lru_shrink_count(&pool->list_lru, sc); > > + > > + /* > > + * Abort if the shrinker is disabled or if we are shrinking int= o the > > + * protected region. > > + */ > > + if (!zswap_shrinker_enabled || nr_protected >=3D lru_size - sc-= >nr_to_scan) { > > + sc->nr_scanned =3D 0; > > + return SHRINK_STOP; > > + } > > + > > + shrink_ret =3D list_lru_shrink_walk(&pool->list_lru, sc, &shrin= k_memcg_cb, > > + &encountered_page_in_swapcache); > > + > > + if (encountered_page_in_swapcache) > > + return SHRINK_STOP; > > + > > + return shrink_ret ? shrink_ret : SHRINK_STOP; > > +} > > + > > +static unsigned long zswap_shrinker_count(struct shrinker *shrinker, > > + struct shrink_control *sc) > > +{ > > + struct zswap_pool *pool =3D shrinker->private_data; > > + struct mem_cgroup *memcg =3D sc->memcg; > > + struct lruvec *lruvec =3D mem_cgroup_lruvec(memcg, NODE_DATA(sc= ->nid)); > > + unsigned long nr_backing, nr_stored, nr_freeable, nr_protected; > > + > > +#ifdef CONFIG_MEMCG_KMEM > > + cgroup_rstat_flush(memcg->css.cgroup); > > + nr_backing =3D memcg_page_state(memcg, MEMCG_ZSWAP_B) >> PAGE_S= HIFT; > > + nr_stored =3D memcg_page_state(memcg, MEMCG_ZSWAPPED); > > +#else > > + /* use pool stats instead of memcg stats */ > > + nr_backing =3D get_zswap_pool_size(pool) >> PAGE_SHIFT; > > + nr_stored =3D atomic_read(&pool->nr_stored); > > +#endif > > + > > + if (!zswap_shrinker_enabled || !nr_stored) > > + return 0; > > + > > + nr_protected =3D atomic_long_read(&lruvec->nr_zswap_protected); > > + nr_freeable =3D list_lru_shrink_count(&pool->list_lru, sc); > > + /* > > + * Subtract the lru size by an estimate of the number of pages > > + * that should be protected. > > + */ > > + nr_freeable =3D nr_freeable > nr_protected ? nr_freeable - nr_p= rotected : 0; > > + > > + /* > > + * Scale the number of freeable pages by the memory saving fact= or. > > + * This ensures that the better zswap compresses memory, the fe= wer > > + * pages we will evict to swap (as it will otherwise incur IO f= or > > + * relatively small memory saving). > > + */ > > + return mult_frac(nr_freeable, nr_backing, nr_stored); > > +} > > + > > +static void zswap_alloc_shrinker(struct zswap_pool *pool) > > +{ > > + pool->shrinker =3D > > + shrinker_alloc(SHRINKER_NUMA_AWARE | SHRINKER_MEMCG_AWA= RE, "mm-zswap"); > > + if (!pool->shrinker) > > + return; > > + > > + pool->shrinker->private_data =3D pool; > > + pool->shrinker->scan_objects =3D zswap_shrinker_scan; > > + pool->shrinker->count_objects =3D zswap_shrinker_count; > > + pool->shrinker->batch =3D 0; > > + pool->shrinker->seeks =3D DEFAULT_SEEKS; > > +} > > + > > /********************************* > > * per-cpu code > > **********************************/ > > @@ -663,8 +782,10 @@ static enum lru_status shrink_memcg_cb(struct list= _head *item, struct list_lru_o > > spinlock_t *lock, void *arg) > > { > > struct zswap_entry *entry =3D container_of(item, struct zswap_e= ntry, lru); > > + bool *encountered_page_in_swapcache =3D (bool *)arg; > > struct mem_cgroup *memcg; > > struct zswap_tree *tree; > > + struct lruvec *lruvec; > > pgoff_t swpoffset; > > enum lru_status ret =3D LRU_REMOVED_RETRY; > > int writeback_result; > > @@ -698,8 +819,22 @@ static enum lru_status shrink_memcg_cb(struct list= _head *item, struct list_lru_o > > /* we cannot use zswap_lru_add here, because it increme= nts node's lru count */ > > list_lru_putback(&entry->pool->list_lru, item, entry_to= _nid(entry), memcg); > > spin_unlock(lock); > > - mem_cgroup_put(memcg); > > ret =3D LRU_RETRY; > > + > > + /* > > + * Encountering a page already in swap cache is a sign = that we are shrinking > > + * into the warmer region. We should terminate shrinkin= g (if we're in the dynamic > > + * shrinker context). > > + */ > > + if (writeback_result =3D=3D -EEXIST && encountered_page= _in_swapcache) { > > + ret =3D LRU_SKIP; > > + *encountered_page_in_swapcache =3D true; > > + } > > + lruvec =3D mem_cgroup_lruvec(memcg, NODE_DATA(entry_to_= nid(entry))); > > + /* Increment the protection area to account for the LRU= rotation. */ > > + atomic_long_inc(&lruvec->nr_zswap_protected); > > + > > + mem_cgroup_put(memcg); > > goto put_unlock; > > } > > zswap_written_back_pages++; > > @@ -822,6 +957,11 @@ static struct zswap_pool *zswap_pool_create(char *= type, char *compressor) > > &pool->node); > > if (ret) > > goto error; > > + > > + zswap_alloc_shrinker(pool); > > + if (!pool->shrinker) > > + goto error; > > + > > pr_debug("using %s compressor\n", pool->tfm_name); > > > > /* being the current pool takes 1 ref; this func expects the > > @@ -829,13 +969,18 @@ static struct zswap_pool *zswap_pool_create(char = *type, char *compressor) > > */ > > kref_init(&pool->kref); > > INIT_LIST_HEAD(&pool->list); > > - list_lru_init_memcg(&pool->list_lru, NULL); > > + if (list_lru_init_memcg(&pool->list_lru, pool->shrinker)) > > + goto lru_fail; > > + shrinker_register(pool->shrinker); > > INIT_WORK(&pool->shrink_work, shrink_worker); > > > > zswap_pool_debug("created", pool); > > > > return pool; > > > > +lru_fail: > > + list_lru_destroy(&pool->list_lru); > > + shrinker_free(pool->shrinker); > > error: > > if (pool->acomp_ctx) > > free_percpu(pool->acomp_ctx); > > @@ -893,6 +1038,7 @@ static void zswap_pool_destroy(struct zswap_pool *= pool) > > > > zswap_pool_debug("destroying", pool); > > > > + shrinker_free(pool->shrinker); > > cpuhp_state_remove_instance(CPUHP_MM_ZSWP_POOL_PREPARE, &pool->= node); > > free_percpu(pool->acomp_ctx); > > list_lru_destroy(&pool->list_lru); > > @@ -1440,6 +1586,7 @@ bool zswap_store(struct folio *folio) > > if (entry->length) { > > INIT_LIST_HEAD(&entry->lru); > > zswap_lru_add(&pool->list_lru, entry); > > + atomic_inc(&pool->nr_stored); > > } > > spin_unlock(&tree->lock); > > > > -- > > 2.34.1 I like this. And FWIW, if we have more states to store (i.e if the shrinker heuristics needs to change), we just need to update things in zswap.h. Will be present in v4 of the patch series. Thanks for the suggestion, Yosry!