From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 2E9ECC77B73 for ; Fri, 26 May 2023 18:53:47 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 7CE17900003; Fri, 26 May 2023 14:53:46 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 75825900002; Fri, 26 May 2023 14:53:46 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 5D364900003; Fri, 26 May 2023 14:53:46 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0010.hostedemail.com [216.40.44.10]) by kanga.kvack.org (Postfix) with ESMTP id 4F895900002 for ; Fri, 26 May 2023 14:53:46 -0400 (EDT) Received: from smtpin09.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay04.hostedemail.com (Postfix) with ESMTP id 236941A0F50 for ; Fri, 26 May 2023 18:53:46 +0000 (UTC) X-FDA: 80833305252.09.A9C59DB Received: from mail-ej1-f48.google.com (mail-ej1-f48.google.com [209.85.218.48]) by imf26.hostedemail.com (Postfix) with ESMTP id 451E714001C for ; Fri, 26 May 2023 18:53:43 +0000 (UTC) Authentication-Results: imf26.hostedemail.com; dkim=pass header.d=google.com header.s=20221208 header.b=JlR2KKEc; dmarc=pass (policy=reject) header.from=google.com; spf=pass (imf26.hostedemail.com: domain of yosryahmed@google.com designates 209.85.218.48 as permitted sender) smtp.mailfrom=yosryahmed@google.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1685127224; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=/dhx1N1TGbdvEAeyn+I2yRhyKZQXdgWDNWhZQJTXxzk=; b=EJHZ4STvk/NvFUhTbP+T38/f8v44p8B1H+FOAUleOE82oC1U5owRUrZr5PxV1T7ir6vqhc bSUysXOPQsb6cy1ey6//2olmyI2DyCIdlIvATrGd9QpSwg/9PIkLfsqQEwJLyBSjDvhpCg zYMr4p9Nb4D43djqsaMIHVSW5hgTa+Y= ARC-Authentication-Results: i=1; imf26.hostedemail.com; dkim=pass header.d=google.com header.s=20221208 header.b=JlR2KKEc; dmarc=pass (policy=reject) header.from=google.com; spf=pass (imf26.hostedemail.com: domain of yosryahmed@google.com designates 209.85.218.48 as permitted sender) smtp.mailfrom=yosryahmed@google.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1685127224; a=rsa-sha256; cv=none; b=nbNaaAvC1ZIMSExuGAmIfpU4rn6D/Kt4kpHAKolz0UPa8iaTsSFgKAhTpPY8dneHy/E1fS K6hssw6F9xwcuQ8AIr11UjzxKDfDGQ+swe5BSto5ErhbfWHUhClYFF4cWAOFbf3bgxVWVQ JvKliipoZmI41zZihSKQ4p1MAb+BVvA= Received: by mail-ej1-f48.google.com with SMTP id a640c23a62f3a-96fbc74fbf1so188357766b.1 for ; Fri, 26 May 2023 11:53:43 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20221208; t=1685127223; x=1687719223; h=content-transfer-encoding:cc:to:subject:message-id:date:from :in-reply-to:references:mime-version:from:to:cc:subject:date :message-id:reply-to; bh=/dhx1N1TGbdvEAeyn+I2yRhyKZQXdgWDNWhZQJTXxzk=; b=JlR2KKEcvO4TYDGY/qCKJnnEN1I6+TVVKBVzDxAwUXL5Szv6TJ2BVB1WoekEuG3xwd 92bYotG3atN7k+AOlirMP/bjRx41mEGDm6Hs29/0Ni5KVTd3uObR9WwN58p+4YlsFNPF iSFsghph/aeCwCTO0MPEqjeK03a0YWsV0fCQnUUr/IxQ9q+rk9UdsFCt64+IQwXg3wCN DtKNyXOlFQ9vcgiVhkyRz2Gzlm6yacd40kCKPe14XLLqdZw5V+hGqPw6v4wyVfl+jmVH e9i2L4aHCUwiK5pNxBNRpY1AuI6W5YhV/mHV/ZrrHCBrA6Epu3tohkQI/A7PqcIMJ/C/ ZM+g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20221208; t=1685127223; x=1687719223; h=content-transfer-encoding:cc:to:subject:message-id:date:from :in-reply-to:references:mime-version:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=/dhx1N1TGbdvEAeyn+I2yRhyKZQXdgWDNWhZQJTXxzk=; b=JcosMgwcc0+dhagyZbYBBqGRYGzdgAIDxDb3P5e0LLu9D7MPDlR4GfzAmRAfSh0o9R EoPCNDzvuaLyiXmX78MAxO6G+KfvIIMoFUc1WEe2076/TTNFfYs//nGA4PxlodzSlY5A Rpis8bYDvwYaHmGejbXt2tcJygPa8fhmharn+WTTJB2XRO9BwrA+H0EhGnP7Df19ti5C UYbUUSgsktTPrEn2L/04gNbX7tK/Ea5fxkOBRSx5eDA2fwT6lzx73P7+NjqskFfJ6Pab ig/n9/FKFWOyDx/nyn8kIrVaJv9/5hEWC9gPZehflvAKwlwcGwv37fv6KNyEE4ERBWmB 0hUA== X-Gm-Message-State: AC+VfDx/l/fsFGV1FZxAWdIU4RMvFB1l04jnjkwXnPZR4rMrRWJU+EcQ brSZnwidEXzi8dNYr9GPIJt/ANv2EfAUVPoNm+/Z3A== X-Google-Smtp-Source: ACHHUZ6IVKQsBrbRNLeTcBFPaNMVlPmDaJp7qf/a111fJNpbkC18GhU7zCZA2OXBXfrqmOqbLk+7iZoOJaWCQeF+uqo= X-Received: by 2002:a17:906:5d0c:b0:965:7fba:6bcf with SMTP id g12-20020a1709065d0c00b009657fba6bcfmr3288218ejt.67.1685127222651; Fri, 26 May 2023 11:53:42 -0700 (PDT) MIME-Version: 1.0 References: <20230526183227.793977-1-cerasuolodomenico@gmail.com> In-Reply-To: <20230526183227.793977-1-cerasuolodomenico@gmail.com> From: Yosry Ahmed Date: Fri, 26 May 2023 11:53:05 -0700 Message-ID: Subject: Re: [PATCH v3] mm: zswap: shrink until can accept To: Domenico Cerasuolo Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, sjenning@redhat.com, ddstreet@ieee.org, vitaly.wool@konsulko.com, hannes@cmpxchg.org, kernel-team@fb.com Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Rspam-User: X-Stat-Signature: e6z64p34hhcx4cd1zcxw99zmajh58yy3 X-Rspamd-Server: rspam07 X-Rspamd-Queue-Id: 451E714001C X-HE-Tag: 1685127223-863314 X-HE-Meta: U2FsdGVkX1+KLO7QlpIY1P3WqA1eyH5DvuBt0iqgiti+QwBo5T+xfGSsLr+7Mvxbu1UOWgxKcU9f6hQFvtOO2RnYXBC3Hlh4tCPOSaaGZksrB1BEurAwqowg57jSj0QW3KMx44iXp7oVIQka4NC9zkDAprJre8MT6XqEjs/C1E3n5wM9fBim0cjZJjnQE5cKrhVWK6A/nTdNTQ5Wi8FJgcSFSwU0/IR1VAx9GGMIww5te3UuwnPrEs1kXIof2UTTvpCqK3xa4cVh94AYIU+VKJ2rdG4/LblfmB3zSJOSPLnfSaPXPOt81LmSv51Jj6Aanxk1QmNdeuP5VIrXJ0HQpRfPkpjbmc31NsRFwvvQULEoRG17voqujN3oa/7JabNAQAJFKFxeX8iy+SdUoNMRzdYU8tzOpZfvmWq0K6MynzjTlqSRHu+v6ud31PppkDwTrmwmRngz37VTYTVaFMq0PNSw3MqFT7ObJ9DYtqPQpEuPZcYaVRrQmAepA8PFBvdtnXoP1FLx09jDOHk4WnTlL2EWtx+vP3BOMw+RM3RylH+0Wv4o5g0chV4uqmSQPGHZPviHUHHUVkLzNBf0KrXQemt0RimV/p31Ha2/ZSPU8+XXsu+gIAl5O8bix4d6FvrtIaH4NBV3AG24sTiWwXhgeIGAf4XCwxgNJfFsB6066ylwaE2X+ENHkY3jyrAQ/yQAnJFSgrYAJmUi6BhBVHDP/7ypAhR6Yt0VmFaSe0Lcr5COrWTg5gJ6qaGSu2H9lstDooxQxSnC+1UL3QHnAc0nyKvufiaoRSlcdHQgy+R5mpuZ05iiGjymRSX9uTzYoIyYyYnyZpqc/n9LcQb3ZsA6W8aUXlgacM/IDiB/ZqOZArTNZwwri0QqQHJpWpmplgTnekukyLiOxinPefoqH+PmkshHs69o4M8ThMcbspL55h6f/BGCCLBEGCfYaZrqgyZi9DQJv8yf42qPfitMHsO 18ml4UVM KaGugQUhwyq4DYCt6HtYmIKz2+2ccdJCqprcteUm1YHcGtfsPxyn3cAe7r6loAth3KW17L6BS9wvG3cHAyMpF2o92HcSrNEu3nj2OciZmdJO3fD6zxDHqEoqLzsvn7SROu4gFTetb0D57NvfwcEYPXsrRstVVBkbyqxDrpagY5dlgCyz0CLgBYXD5qc4lV66zC/Zx7IxMwkcGcN838DEQrbCarB9n58ofmqqPmn3LtUBQiT2yuAxnwtmOhUca3ltfKT+v34hAv41yigUvdSCNriMN2g== X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: On Fri, May 26, 2023 at 11:32=E2=80=AFAM Domenico Cerasuolo wrote: > > This update addresses an issue with the zswap reclaim mechanism, which > hinders the efficient offloading of cold pages to disk, thereby > compromising the preservation of the LRU order and consequently > diminishing, if not inverting, its performance benefits. > > The functioning of the zswap shrink worker was found to be inadequate, > as shown by basic benchmark test. For the test, a kernel build was > utilized as a reference, with its memory confined to 1G via a cgroup and > a 5G swap file provided. The results are presented below, these are > averages of three runs without the use of zswap: > > real 46m26s > user 35m4s > sys 7m37s > > With zswap (zbud) enabled and max_pool_percent set to 1 (in a 32G > system), the results changed to: > > real 56m4s > user 35m13s > sys 8m43s > > written_back_pages: 18 > reject_reclaim_fail: 0 > pool_limit_hit:1478 > > Besides the evident regression, one thing to notice from this data is > the extremely low number of written_back_pages and pool_limit_hit. > > The pool_limit_hit counter, which is increased in zswap_frontswap_store > when zswap is completely full, doesn't account for a particular > scenario: once zswap hits his limit, zswap_pool_reached_full is set to > true; with this flag on, zswap_frontswap_store rejects pages if zswap is > still above the acceptance threshold. Once we include the rejections due > to zswap_pool_reached_full && !zswap_can_accept(), the number goes from > 1478 to a significant 21578266. > > Zswap is stuck in an undesirable state where it rejects pages because > it's above the acceptance threshold, yet fails to attempt memory > reclaimation. This happens because the shrink work is only queued when > zswap_frontswap_store detects that it's full and the work itself only > reclaims one page per run. > > This state results in hot pages getting written directly to disk, > while cold ones remain memory, waiting only to be invalidated. The LRU > order is completely broken and zswap ends up being just an overhead > without providing any benefits. > > This commit applies 2 changes: a) the shrink worker is set to reclaim > pages until the acceptance threshold is met and b) the task is also > enqueued when zswap is not full but still above the threshold. > > Testing this suggested update showed much better numbers: > > real 36m37s > user 35m8s > sys 9m32s > > written_back_pages: 10459423 > reject_reclaim_fail: 12896 > pool_limit_hit: 75653 > > V2: > - loop against =3D=3D -EAGAIN rather than !=3D -EINVAL and also break the= loop > on MAX_RECLAIM_RETRIES (thanks Yosry) > - cond_resched() to ensure that the loop doesn't burn the cpu (thanks > Vitaly) > > V3: > - fix wrong loop break, should continue on !ret (thanks Johannes) > > Fixes: 45190f01dd40 ("mm/zswap.c: add allocation hysteresis if pool limit= is hit") > Signed-off-by: Domenico Cerasuolo Reviewed-by: Yosry Ahmed > --- > mm/zswap.c | 17 ++++++++++++++--- > 1 file changed, 14 insertions(+), 3 deletions(-) > > diff --git a/mm/zswap.c b/mm/zswap.c > index 59da2a415fbb..bcb82e09eb64 100644 > --- a/mm/zswap.c > +++ b/mm/zswap.c > @@ -37,6 +37,7 @@ > #include > > #include "swap.h" > +#include "internal.h" > > /********************************* > * statistics > @@ -587,9 +588,19 @@ static void shrink_worker(struct work_struct *w) > { > struct zswap_pool *pool =3D container_of(w, typeof(*pool), > shrink_work); > + int ret, failures =3D 0; > > - if (zpool_shrink(pool->zpool, 1, NULL)) > - zswap_reject_reclaim_fail++; > + do { > + ret =3D zpool_shrink(pool->zpool, 1, NULL); > + if (ret) { > + zswap_reject_reclaim_fail++; > + if (ret !=3D -EAGAIN) > + break; > + if (++failures =3D=3D MAX_RECLAIM_RETRIES) > + break; > + } > + cond_resched(); > + } while (!zswap_can_accept()); > zswap_pool_put(pool); > } > > @@ -1188,7 +1199,7 @@ static int zswap_frontswap_store(unsigned type, pgo= ff_t offset, > if (zswap_pool_reached_full) { > if (!zswap_can_accept()) { > ret =3D -ENOMEM; > - goto reject; > + goto shrink; > } else > zswap_pool_reached_full =3D false; > } > -- > 2.34.1 >