From mboxrd@z Thu Jan  1 00:00:00 1970
Return-Path: <owner-linux-mm@kvack.org>
X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on
	aws-us-west-2-korg-lkml-1.web.codeaurora.org
Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17])
	by smtp.lore.kernel.org (Postfix) with ESMTP id DA05AC02194
	for <linux-mm@archiver.kernel.org>; Thu,  6 Feb 2025 19:15:25 +0000 (UTC)
Received: by kanga.kvack.org (Postfix)
	id 3121D280002; Thu,  6 Feb 2025 14:15:25 -0500 (EST)
Received: by kanga.kvack.org (Postfix, from userid 40)
	id 2BEE3280004; Thu,  6 Feb 2025 14:15:25 -0500 (EST)
X-Delivered-To: int-list-linux-mm@kvack.org
Received: by kanga.kvack.org (Postfix, from userid 63042)
	id 15FE8280002; Thu,  6 Feb 2025 14:15:25 -0500 (EST)
X-Delivered-To: linux-mm@kvack.org
Received: from relay.hostedemail.com (smtprelay0010.hostedemail.com [216.40.44.10])
	by kanga.kvack.org (Postfix) with ESMTP id E33F76B008A
	for <linux-mm@kvack.org>; Thu,  6 Feb 2025 14:15:24 -0500 (EST)
Received: from smtpin12.hostedemail.com (a10.router.float.18 [10.200.18.1])
	by unirelay07.hostedemail.com (Postfix) with ESMTP id 84BD11613D1
	for <linux-mm@kvack.org>; Thu,  6 Feb 2025 19:15:24 +0000 (UTC)
X-FDA: 83090473368.12.312C63C
Received: from out-178.mta0.migadu.com (out-178.mta0.migadu.com [91.218.175.178])
	by imf04.hostedemail.com (Postfix) with ESMTP id AF0714000A
	for <linux-mm@kvack.org>; Thu,  6 Feb 2025 19:15:22 +0000 (UTC)
Authentication-Results: imf04.hostedemail.com;
	dkim=pass header.d=linux.dev header.s=key1 header.b=rhOuGx60;
	spf=pass (imf04.hostedemail.com: domain of yosry.ahmed@linux.dev designates 91.218.175.178 as permitted sender) smtp.mailfrom=yosry.ahmed@linux.dev;
	dmarc=pass (policy=none) header.from=linux.dev
ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com;
	s=arc-20220608; t=1738869322;
	h=from:from:sender:reply-to:subject:subject:date:date:
	 message-id:message-id:to:to:cc:cc:mime-version:mime-version:
	 content-type:content-type:content-transfer-encoding:
	 in-reply-to:in-reply-to:references:references:dkim-signature;
	bh=3ZQyY8MCXQQhnq9axd4H2VgxvWzK3oBbYgT4qwXq3zg=;
	b=jTxYolLNDRUMmSEIXBBibx2ilzQHZo0ax0eE3MpjQHxtE8u9Qs/5gbBt/LtwYx0qsBGC3h
	ItjU2GkhHAKOO7b+1W4pRNqPr2gQYsCuJl9nNZxML0E/JQQttq6mlP8/HY+quFNWfCCGlH
	KgoJgbWYUjbUsMednFEeV+wl0o2dehM=
ARC-Authentication-Results: i=1;
	imf04.hostedemail.com;
	dkim=pass header.d=linux.dev header.s=key1 header.b=rhOuGx60;
	spf=pass (imf04.hostedemail.com: domain of yosry.ahmed@linux.dev designates 91.218.175.178 as permitted sender) smtp.mailfrom=yosry.ahmed@linux.dev;
	dmarc=pass (policy=none) header.from=linux.dev
ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1738869322; a=rsa-sha256;
	cv=none;
	b=NXSWvs7RH1OhlDcNW8EACI+j6yz2tm1Tdiqe73ZqjUIuRTUzzMw9V8ltC72oSuUCuoo9y6
	GkSxb11pzb5jbD2Km0TKNGEA/lgtcZDZid4wPOFNmsl7IcOYv7miWr7xfrSo4rdu3uOGOZ
	fhkINWwtx2kr7ctlMu3uHSjO8hiLvEA=
Date: Thu, 6 Feb 2025 19:15:15 +0000
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1;
	t=1738869321;
	h=from:from:reply-to:subject:subject:date:date:message-id:message-id:
	 to:to:cc:cc:mime-version:mime-version:content-type:content-type:
	 in-reply-to:in-reply-to:references:references;
	bh=3ZQyY8MCXQQhnq9axd4H2VgxvWzK3oBbYgT4qwXq3zg=;
	b=rhOuGx60ql/eguU89eWhDhiQ14Rd0VDaXkm3ijXCqFwnFvMnvbHa4UP14hib7Zx9M8PrAo
	Z0xBySVvFJhRYkQRycVd10kFmK+M4Nam4F4sKSOyYPB9LX3jf2SCihncVmkM5ZQYylibMs
	iTPIdx38nl3NZLHnzhSnDV2FcjkbeJE=
X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers.
From: Yosry Ahmed <yosry.ahmed@linux.dev>
To: Kanchana P Sridhar <kanchana.p.sridhar@intel.com>
Cc: linux-kernel@vger.kernel.org, linux-mm@kvack.org, hannes@cmpxchg.org,
	nphamcs@gmail.com, chengming.zhou@linux.dev, usamaarif642@gmail.com,
	ryan.roberts@arm.com, 21cnbao@gmail.com, akpm@linux-foundation.org,
	linux-crypto@vger.kernel.org, herbert@gondor.apana.org.au,
	davem@davemloft.net, clabbe@baylibre.com, ardb@kernel.org,
	ebiggers@google.com, surenb@google.com, kristen.c.accardi@intel.com,
	wajdi.k.feghali@intel.com, vinodh.gopal@intel.com
Subject: Re: [PATCH v6 16/16] mm: zswap: Fix for zstd performance regression
 with 2M folios.
Message-ID: <Z6UKQ04ClABSePLZ@google.com>
References: <20250206072102.29045-1-kanchana.p.sridhar@intel.com>
 <20250206072102.29045-17-kanchana.p.sridhar@intel.com>
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <20250206072102.29045-17-kanchana.p.sridhar@intel.com>
X-Migadu-Flow: FLOW_OUT
X-Rspam-User: 
X-Rspamd-Server: rspam09
X-Rspamd-Queue-Id: AF0714000A
X-Stat-Signature: u8tocbxosp4c8e6r3e1p3x3t9j4rb96g
X-HE-Tag: 1738869322-713471
X-HE-Meta: U2FsdGVkX1+rnh5LWksiJ5fCaH8ga8jci5R8+QGGBbHirkomZJDi7m2SpA6IIJ8iL5Lqs1Bu4t9Kdnwzv4yOWJHTges2cjQARCwpnXi2phiYMnjTl1AacXWn+UuSeG+fRAN3EpZ6pSYIVlF8ciAWD1KJEu/yFIqa6OF0BT1WdaM1ASDeJNOlnlWELsYSNDa2nbMFKI7uTBFz4RP1R/YedInf52McsvqBL5OMiFy324dXtT+5tUHVWZvFKxOEzQb86EpRnPFKTQDWHvun8qaTi8+6kElCvQHFu5xeEOEjn8aqDVt7HOVDppuKulDLS1+0qSavXzCiJQ7mHP0tvA83L07zFgg/UQ8oFW/5Hd/qQw1gMImjs0VlBqgJmx+7cDj2FRP1APPMPVLq6dEf+Q537pMkScpoan+2PMTow4MhO869a/EJimxDUCgluQkRrixpMrd4Fni9rm/igN+tKSn3lCSgATaiitGBNemZNYiaaLd49JeX2EfXm9rSthEE+WHsZrr0y5FfCJIDlutb+2ArnbAar5aMIXQcAYbUggNOjNoV/SYd3PMV9KOH7Uz9N4wGAhcd9zNzPyU0BGky4YTYSDxpjy5wEp+eZsdreAVhC7oXRQ478y8RF65ZqpkGlsy/3+ZW0EpK39HeVaCUcMA4p3jrrJfoOrx/Jwud5cPf6fjf6M7+QhPfMKDdaxX2whOf3NnRWLHU5xty9EIjQ2Mwhk5Rhht+7vKZ5031//QJ0rO6nkRn1+YL1jIvGdyib8SOQOT9owJmJn1BPnwg03xLC1x9sJ7LvVuYgw7D+SY4gy2ug5kXB5ficM0bjjCF6Vrt6YXlWbPs27wSUAU4PzfstGTj2WafWufCTj1GKPbMpnQGvTiqV6YMWJgZGHUWauEIIp/cNSSQcZGV5wNj+9HTBjO4dwrlznlb6n6/QUPxCX3Dx2TAskWgHtgg0NMWcJAy2bvGumYu8MNmj4zwogO
 KLbMZaSX
 MNd/ykzFkfK8Dk/WD7kp8nua4kRIlBWzW38l+SR+fDDFRPccl6pLUU1KvziJshOzgttWBR0VUZHbjpilKZJnzSQM0MYWUZlyOzxd3crQY3+kcq8xuI8WucyLF3Z4tDwYoOsMWhfCIQIwfvIViwVh4AzMmEP76LwmM1UeRydZvEl2y+tXQH8HDNjNtJoMKIuaxuNWEgT6+C+CpyEQimTGQuqLg5g==
X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4
Sender: owner-linux-mm@kvack.org
Precedence: bulk
X-Loop: owner-majordomo@kvack.org
List-ID: <linux-mm.kvack.org>
List-Subscribe: <mailto:majordomo@kvack.org>
List-Unsubscribe: <mailto:majordomo@kvack.org>

On Wed, Feb 05, 2025 at 11:21:02PM -0800, Kanchana P Sridhar wrote:
> With the previous patch that enables support for batch compressions in
> zswap_compress_folio(), a 6.2% throughput regression was seen with zstd and
> 2M folios, using vm-scalability/usemem.
> 
> For compressors that don't support batching, this was root-caused to the
> following zswap_store_folio() structure:
> 
>  Batched stores:
>  ---------------
>  - Allocate all entries,
>  - Compress all entries,
>  - Store all entries in xarray/LRU.
> 
> Hence, the above structure is maintained only for batched stores, and the
> following structure is implemented for sequential stores of large folio pages,
> that fixes the zstd regression, while preserving common code paths for batched
> and sequential stores of a folio:
> 
>  Sequential stores:
>  ------------------
>  For each page in folio:
>   - allocate an entry,
>   - compress the page,
>   - store the entry in xarray/LRU.
> 
> This is submitted as a separate patch only for code review purposes. I will
> squash this with the previous commit in subsequent versions of this
> patch-series.

Could it be the cache locality?

I wonder if we should do what Chengming initially suggested and batch
everything at ZSWAP_MAX_BATCH_SIZE instead. Instead of
zswap_compress_folio() operating on the entire folio, we can operate on
batches of size ZSWAP_MAX_BATCH_SIZE, regardless of whether the
underlying compressor supports batching.

If we do this, instead of:
- Allocate all entries
- Compress all entries
- Store all entries

We can do:
  - For each batch (8 entries)
  	- Allocate all entries
	- Compress all entries
	- Store all entries

This should help unify the code, and I suspect it may also fix the zstd
regression. We can also skip the entries array allocation and use one on
the stack.