From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id E8C0FC7EE23 for ; Wed, 7 Jun 2023 21:39:22 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 65D638E0003; Wed, 7 Jun 2023 17:39:22 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 60E528E0001; Wed, 7 Jun 2023 17:39:22 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 4D53C8E0003; Wed, 7 Jun 2023 17:39:22 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0010.hostedemail.com [216.40.44.10]) by kanga.kvack.org (Postfix) with ESMTP id 3BE088E0001 for ; Wed, 7 Jun 2023 17:39:22 -0400 (EDT) Received: from smtpin12.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay09.hostedemail.com (Postfix) with ESMTP id E4FCF803E7 for ; Wed, 7 Jun 2023 21:39:21 +0000 (UTC) X-FDA: 80877268122.12.5121ABE Received: from mail-qv1-f43.google.com (mail-qv1-f43.google.com [209.85.219.43]) by imf12.hostedemail.com (Postfix) with ESMTP id 2136C4000A for ; Wed, 7 Jun 2023 21:39:19 +0000 (UTC) Authentication-Results: imf12.hostedemail.com; dkim=pass header.d=gmail.com header.s=20221208 header.b=mJEU7MyY; dmarc=pass (policy=none) header.from=gmail.com; spf=pass (imf12.hostedemail.com: domain of nphamcs@gmail.com designates 209.85.219.43 as permitted sender) smtp.mailfrom=nphamcs@gmail.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1686173960; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=pGDCkbdGN1MhlLFnaI8S08OX/yFHe0s8kyjhWCk8H7w=; b=uc3XM+tXvIZGZSB3HF8bJSKiJ3mA/nJSjWEDX4tPTgNdKI2he5v27EDf9ct9k8e/4hXS0j e8QKeTEnh4IeQ4k2/TVPhz4WaOHOegNFMMIie4lJrx/VS0zuuRpMw1m439V348r6yYVK9l FU00rGLR4ID7pJSZUm2SUL5H3uKHvKU= ARC-Authentication-Results: i=1; imf12.hostedemail.com; dkim=pass header.d=gmail.com header.s=20221208 header.b=mJEU7MyY; dmarc=pass (policy=none) header.from=gmail.com; spf=pass (imf12.hostedemail.com: domain of nphamcs@gmail.com designates 209.85.219.43 as permitted sender) smtp.mailfrom=nphamcs@gmail.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1686173960; a=rsa-sha256; cv=none; b=mweQszSXe9x/40QGmunvNj0RQe45PWPgsapJ0IAkEyTKImUmIZVTyH+CZ8RndAemxVNtUx N54Ckp/4XiUZ3emSotDBhJt6TaGqbTvVzDctFb2+eBwTZLJe+0By1IYHeaBcDlZldgICqT u9MWKcT9v9ONqr25frM/o43PsN7lUxs= Received: by mail-qv1-f43.google.com with SMTP id 6a1803df08f44-6261a61b38cso59203586d6.1 for ; Wed, 07 Jun 2023 14:39:19 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20221208; t=1686173959; x=1688765959; h=content-transfer-encoding:cc:to:subject:message-id:date:from :in-reply-to:references:mime-version:from:to:cc:subject:date :message-id:reply-to; bh=pGDCkbdGN1MhlLFnaI8S08OX/yFHe0s8kyjhWCk8H7w=; b=mJEU7MyYWGTtC0MSZzP22m9UhOH+6qglKGonrIho8evgn8Qlzatxa8kwZz4sQM/s4V vwKmprb4LMvOaorr1m+q+Wx3QPX97wad3ALtAY0Ba0BQnm1UFkiEuhougJltFW7QYx9j qKFdFnXlq3uvYMg/qGJ6YU+g9PyEIr69nRWOlKV7LrIY4lybQ4iLaiufQw05ro+5I/BU Ee6nzHBNsDYP1X16N6aiEg6082syMfOExN5woE0YPCBIxGmsFSzNjvM0wo+CWKF2yAUx uPO5y02tYwKsbIavKH8zD3/T3iuKRDKLCZfITNWNAMT1zUC3/U6l+qDSICjfPZJ2JEO+ P1xg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20221208; t=1686173959; x=1688765959; h=content-transfer-encoding:cc:to:subject:message-id:date:from :in-reply-to:references:mime-version:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=pGDCkbdGN1MhlLFnaI8S08OX/yFHe0s8kyjhWCk8H7w=; b=CqHAPxuxWa63rxr9ZnYc3R6AsTboYvBKMMlD2grR9wNTmIbFYvqBxF1z1o389vapyp IkjvG/7CgliJ0+KdNxzF8Ba+mlrK/Nvo+tJaciz6JqOy0Cz4ZksASLjvQfjisjiVAFgi Zby0/7by4k1/dSjbMO5NJYZWaPS2xhSaeTiOxg+bk0fKhzZKzTJ3ztEUC2SEirmqkn9Q cHa+jJrh1dK9ZEJmbWEJIO7nf6Q48VPwsVUc6ZkyiJc5tfcDcKSGYe2w5BgNB7FUaMJt C9qdqpCdH599ukMqjjR4mOTWuxgDe6YlWAlXCKJnfzAPCKM35QR+RUzQiRq9gT5z6i3a eNoA== X-Gm-Message-State: AC+VfDyvZ6Ag9Jkd0w02gny6wGxJEHiMZWpJ58G3tNydZDMBOY2rR8hO JwRzms8jMNKh0Cn1t/ZYCjOf5HVqLFyLxiWjJDA= X-Google-Smtp-Source: ACHHUZ6VTV6Ci31JbJ0RpEI0E3uasHmPWgr1Tpkojx8LJQVD4ywRNEC2r16bGrwsje/ImzEDs/qdz+Cmn7NcSY6gmOw= X-Received: by 2002:ad4:5747:0:b0:61a:197b:605 with SMTP id q7-20020ad45747000000b0061a197b0605mr5575505qvx.1.1686173958985; Wed, 07 Jun 2023 14:39:18 -0700 (PDT) MIME-Version: 1.0 References: <20230606145611.704392-1-cerasuolodomenico@gmail.com> <20230606145611.704392-2-cerasuolodomenico@gmail.com> In-Reply-To: <20230606145611.704392-2-cerasuolodomenico@gmail.com> From: Nhat Pham Date: Wed, 7 Jun 2023 14:39:08 -0700 Message-ID: Subject: Re: [RFC PATCH v2 1/7] mm: zswap: add pool shrinking mechanism To: Domenico Cerasuolo Cc: vitaly.wool@konsulko.com, minchan@kernel.org, senozhatsky@chromium.org, yosryahmed@google.com, linux-mm@kvack.org, ddstreet@ieee.org, sjenning@redhat.com, hannes@cmpxchg.org, akpm@linux-foundation.org, linux-kernel@vger.kernel.org, kernel-team@meta.com Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Rspamd-Queue-Id: 2136C4000A X-Rspam-User: X-Rspamd-Server: rspam05 X-Stat-Signature: eope4g8d8qrr5ag3p6gjjzinwigmkrr5 X-HE-Tag: 1686173959-721147 X-HE-Meta: U2FsdGVkX19Y5t327aFuVvVHO6B88E/PMHsAmTvOG8GHBPF1zSfQIi26b7nMsxTxBuxdMZ9fP6h1m7xfmeEa/zDEhWQYYgtM49b7zBCsObgQiBemtd2ZlveTltd1AJKFm+J1nG2myi5BajwMp/PTX5DFMV4QVpGUjPSggLTXn53oRBoDLdPt4wjFyxX0Ew2vte0DQjWZr9Qj1py1GvMs93h6aYj6iMm2n70Nsv9X9lw6wuBAMDIQ1bkUFzG/3JT9wG1Ed7mPTDrD0tTYSLdabwiVxSc/tYou0gAIwoKVqvW4fDC7/stbUvM4gTawxN8K5yo/ENu88LgI5aurWBnO5WIHuVlJmmw55HJ2kD0ALc2/XeCWunyAGdbiFtyFpmFcefgSQugTXH0VyLq3GGdiUcPeTaKZNNkXj8Fb/v9uqKGnBtJyHPT9fUTab1a2wRx+4Dh1aQ3ogLvz7fQ87Y/2AF6OmIUJwH90SgoPmcH3c/Qo8fc2NBSSVpFDMxA4L3RZZIttu5MdWBspNIhsx+LZINNrv6Zn0oRSEVARepyh68KqaAKrNGAS16m1wn+U+HMfKmDVmZ/O/mUvsDgTEK4kji8NlNwNFUleD5xkIS2rsfWLOhXZ8UWMc/C6INufmDilvKPlAgO+ZuYfuo0ovDerUvRVYqp7c0hAkcUqU5C4vqzbwaQy9uOYUTvdFXbvVDGOurZajd/qttaUNk+ygBEq3k5gkSq6u/5rAr6xJXJpxPtMtnkBbHKK3571T1uyfFNhkSFGqsYhBxttwHDigj7gKH1L+th6pUtXT3xlvr0EOD/UWocrQvMS/f4jwe2tQJznnep+Dz1MGObYk1dRsQEcdx376NYoFvLJQ0au7foBeADOWhTxFSc3dGUBr96daLC5Vhl75GievSb5eyU+XxgxiYn/NGFgfaCcLrvyj25Ht4o0opeEAFuA/EgBZjMSqO2kbZZGFuUAGUG0osdFW5x kTu4lqdq bb2xkNmed2VAqDgtq1VDBeUnoVbL+m1YxZ2xd1XkYHwNlPzg/hKgj+FfD43tlWiVp6GHKdFKwS261/Lx056QLDgXbzQW/3fOCzKYcoDZvNDr7/i2mdHdzsk/uon+AdRg6KXqGXC8SVmZ3WHK/mmfdGGIQ2wP3e4Lw9eQC219qZu0TYUjgw1t0WpMijGQaLDqGXHpXSciygmpxVBVZ3IjmoVLytJJCor1ZRQTfRQv41cjWLyzdOzehXd7zxZcN3JLeG/QMFgfeyOShSQbW0dwWpoiR0rSDevyYrJ2X6QhxD9+q/F1srd8DB7pufquhzq9pek6jNG75gFmg56JsvougMm0+1Ft3Uw2HHA+4GSXqnrql/1pjpcJ3os90kg== X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: On Tue, Jun 6, 2023 at 7:56=E2=80=AFAM Domenico Cerasuolo wrote: > > Each zpool driver (zbud, z3fold and zsmalloc) implements its own shrink > function, which is called from zpool_shrink. However, with this commit, > a unified shrink function is added to zswap. The ultimate goal is to > eliminate the need for zpool_shrink once all zpool implementations have > dropped their shrink code. > > To ensure the functionality of each commit, this change focuses solely > on adding the mechanism itself. No modifications are made to > the backends, meaning that functionally, there are no immediate changes. > The zswap mechanism will only come into effect once the backends have > removed their shrink code. The subsequent commits will address the > modifications needed in the backends. > > Signed-off-by: Domenico Cerasuolo > --- > mm/zswap.c | 96 +++++++++++++++++++++++++++++++++++++++++++++++++++--- > 1 file changed, 91 insertions(+), 5 deletions(-) > > diff --git a/mm/zswap.c b/mm/zswap.c > index bcb82e09eb64..c99bafcefecf 100644 > --- a/mm/zswap.c > +++ b/mm/zswap.c > @@ -150,6 +150,12 @@ struct crypto_acomp_ctx { > struct mutex *mutex; > }; > > +/* > + * The lock ordering is zswap_tree.lock -> zswap_pool.lru_lock. > + * The only case where lru_lock is not acquired while holding tree.lock = is > + * when a zswap_entry is taken off the lru for writeback, in that case i= t > + * needs to be verified that it's still valid in the tree. > + */ > struct zswap_pool { > struct zpool *zpool; > struct crypto_acomp_ctx __percpu *acomp_ctx; > @@ -159,6 +165,8 @@ struct zswap_pool { > struct work_struct shrink_work; > struct hlist_node node; > char tfm_name[CRYPTO_MAX_ALG_NAME]; > + struct list_head lru; > + spinlock_t lru_lock; > }; > > /* > @@ -176,10 +184,12 @@ struct zswap_pool { > * be held while changing the refcount. Since the lock must > * be held, there is no reason to also make refcount atomic. > * length - the length in bytes of the compressed page data. Needed dur= ing > - * decompression. For a same value filled page length is 0. > + * decompression. For a same value filled page length is 0, and= both > + * pool and lru are invalid and must be ignored. > * pool - the zswap_pool the entry's data is in > * handle - zpool allocation handle that stores the compressed page data > * value - value of the same-value filled pages which have same content > + * lru - handle to the pool's lru used to evict pages. > */ > struct zswap_entry { > struct rb_node rbnode; > @@ -192,6 +202,7 @@ struct zswap_entry { > unsigned long value; > }; > struct obj_cgroup *objcg; > + struct list_head lru; > }; > > struct zswap_header { > @@ -364,6 +375,12 @@ static void zswap_free_entry(struct zswap_entry *ent= ry) > if (!entry->length) > atomic_dec(&zswap_same_filled_pages); > else { > + /* zpool_evictable will be removed once all 3 backends have migra= ted */ > + if (!zpool_evictable(entry->pool->zpool)) { > + spin_lock(&entry->pool->lru_lock); > + list_del(&entry->lru); > + spin_unlock(&entry->pool->lru_lock); > + } > zpool_free(entry->pool->zpool, entry->handle); > zswap_pool_put(entry->pool); > } > @@ -584,14 +601,70 @@ static struct zswap_pool *zswap_pool_find_get(char = *type, char *compressor) > return NULL; > } > > +static int zswap_shrink(struct zswap_pool *pool) > +{ > + struct zswap_entry *lru_entry, *tree_entry =3D NULL; > + struct zswap_header *zhdr; > + struct zswap_tree *tree; > + int swpoffset; > + int ret; > + > + /* get a reclaimable entry from LRU */ > + spin_lock(&pool->lru_lock); > + if (list_empty(&pool->lru)) { > + spin_unlock(&pool->lru_lock); > + return -EINVAL; > + } > + lru_entry =3D list_last_entry(&pool->lru, struct zswap_entry, lru= ); > + list_del_init(&lru_entry->lru); > + zhdr =3D zpool_map_handle(pool->zpool, lru_entry->handle, ZPOOL_M= M_RO); > + tree =3D zswap_trees[swp_type(zhdr->swpentry)]; > + zpool_unmap_handle(pool->zpool, lru_entry->handle); > + /* > + * Once the pool lock is dropped, the lru_entry might get freed. = The > + * swpoffset is copied to the stack, and lru_entry isn't deref'd = again > + * until the entry is verified to still be alive in the tree. > + */ > + swpoffset =3D swp_offset(zhdr->swpentry); > + spin_unlock(&pool->lru_lock); > + > + /* hold a reference from tree so it won't be freed during writeba= ck */ > + spin_lock(&tree->lock); > + tree_entry =3D zswap_entry_find_get(&tree->rbroot, swpoffset); > + if (tree_entry !=3D lru_entry) { > + if (tree_entry) > + zswap_entry_put(tree, tree_entry); > + spin_unlock(&tree->lock); > + return -EAGAIN; > + } > + spin_unlock(&tree->lock); > + > + ret =3D zswap_writeback_entry(pool->zpool, lru_entry->handle); > + > + spin_lock(&tree->lock); > + if (ret) { > + spin_lock(&pool->lru_lock); > + list_move(&lru_entry->lru, &pool->lru); > + spin_unlock(&pool->lru_lock); > + } > + zswap_entry_put(tree, tree_entry); > + spin_unlock(&tree->lock); > + > + return ret ? -EAGAIN : 0; > +} > + > static void shrink_worker(struct work_struct *w) > { > struct zswap_pool *pool =3D container_of(w, typeof(*pool), > shrink_work); > int ret, failures =3D 0; > > + /* zpool_evictable will be removed once all 3 backends have migra= ted */ > do { > - ret =3D zpool_shrink(pool->zpool, 1, NULL); > + if (zpool_evictable(pool->zpool)) > + ret =3D zpool_shrink(pool->zpool, 1, NULL); > + else > + ret =3D zswap_shrink(pool); > if (ret) { > zswap_reject_reclaim_fail++; > if (ret !=3D -EAGAIN) > @@ -655,6 +728,8 @@ static struct zswap_pool *zswap_pool_create(char *typ= e, char *compressor) > */ > kref_init(&pool->kref); > INIT_LIST_HEAD(&pool->list); > + INIT_LIST_HEAD(&pool->lru); > + spin_lock_init(&pool->lru_lock); > INIT_WORK(&pool->shrink_work, shrink_worker); > > zswap_pool_debug("created", pool); > @@ -1270,7 +1345,7 @@ static int zswap_frontswap_store(unsigned type, pgo= ff_t offset, > } > > /* store */ > - hlen =3D zpool_evictable(entry->pool->zpool) ? sizeof(zhdr) : 0; > + hlen =3D sizeof(zhdr); > gfp =3D __GFP_NORETRY | __GFP_NOWARN | __GFP_KSWAPD_RECLAIM; > if (zpool_malloc_support_movable(entry->pool->zpool)) > gfp |=3D __GFP_HIGHMEM | __GFP_MOVABLE; > @@ -1313,6 +1388,12 @@ static int zswap_frontswap_store(unsigned type, pg= off_t offset, > zswap_entry_put(tree, dupentry); > } > } while (ret =3D=3D -EEXIST); > + /* zpool_evictable will be removed once all 3 backends have migra= ted */ > + if (entry->length && !zpool_evictable(entry->pool->zpool)) { > + spin_lock(&entry->pool->lru_lock); > + list_add(&entry->lru, &entry->pool->lru); > + spin_unlock(&entry->pool->lru_lock); > + } > spin_unlock(&tree->lock); > > /* update stats */ > @@ -1384,8 +1465,7 @@ static int zswap_frontswap_load(unsigned type, pgof= f_t offset, > /* decompress */ > dlen =3D PAGE_SIZE; > src =3D zpool_map_handle(entry->pool->zpool, entry->handle, ZPOOL= _MM_RO); > - if (zpool_evictable(entry->pool->zpool)) > - src +=3D sizeof(struct zswap_header); > + src +=3D sizeof(struct zswap_header); > > if (!zpool_can_sleep_mapped(entry->pool->zpool)) { > memcpy(tmp, src, entry->length); > @@ -1415,6 +1495,12 @@ static int zswap_frontswap_load(unsigned type, pgo= ff_t offset, > freeentry: > spin_lock(&tree->lock); > zswap_entry_put(tree, entry); > + /* zpool_evictable will be removed once all 3 backends have migra= ted */ > + if (entry->length && !zpool_evictable(entry->pool->zpool)) { > + spin_lock(&entry->pool->lru_lock); > + list_move(&entry->lru, &entry->pool->lru); > + spin_unlock(&entry->pool->lru_lock); > + } > spin_unlock(&tree->lock); > > return ret; > -- > 2.34.1 > Looks real solid to me! Thanks, Domenico. Acked-by: Nhat Pham