From mboxrd@z Thu Jan  1 00:00:00 1970
Return-Path: <owner-linux-mm@kvack.org>
X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on
	aws-us-west-2-korg-lkml-1.web.codeaurora.org
Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17])
	by smtp.lore.kernel.org (Postfix) with ESMTP id E9E5AC54E66
	for <linux-mm@archiver.kernel.org>; Tue, 12 Mar 2024 08:03:16 +0000 (UTC)
Received: by kanga.kvack.org (Postfix)
	id 57E806B01A3; Tue, 12 Mar 2024 04:03:16 -0400 (EDT)
Received: by kanga.kvack.org (Postfix, from userid 40)
	id 52E9E8D0017; Tue, 12 Mar 2024 04:03:16 -0400 (EDT)
X-Delivered-To: int-list-linux-mm@kvack.org
Received: by kanga.kvack.org (Postfix, from userid 63042)
	id 3F6556B01A5; Tue, 12 Mar 2024 04:03:16 -0400 (EDT)
X-Delivered-To: linux-mm@kvack.org
Received: from relay.hostedemail.com (smtprelay0011.hostedemail.com [216.40.44.11])
	by kanga.kvack.org (Postfix) with ESMTP id 306586B01A3
	for <linux-mm@kvack.org>; Tue, 12 Mar 2024 04:03:16 -0400 (EDT)
Received: from smtpin23.hostedemail.com (a10.router.float.18 [10.200.18.1])
	by unirelay01.hostedemail.com (Postfix) with ESMTP id 022261C0285
	for <linux-mm@kvack.org>; Tue, 12 Mar 2024 08:03:15 +0000 (UTC)
X-FDA: 81887646792.23.4D72741
Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.8])
	by imf21.hostedemail.com (Postfix) with ESMTP id BB6231C0010
	for <linux-mm@kvack.org>; Tue, 12 Mar 2024 08:03:13 +0000 (UTC)
Authentication-Results: imf21.hostedemail.com;
	dkim=pass header.d=intel.com header.s=Intel header.b=kzon94xU;
	spf=pass (imf21.hostedemail.com: domain of ying.huang@intel.com designates 192.198.163.8 as permitted sender) smtp.mailfrom=ying.huang@intel.com;
	dmarc=pass (policy=none) header.from=intel.com
ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com;
	s=arc-20220608; t=1710230594;
	h=from:from:sender:reply-to:subject:subject:date:date:
	 message-id:message-id:to:to:cc:cc:mime-version:mime-version:
	 content-type:content-type:content-transfer-encoding:
	 in-reply-to:in-reply-to:references:references:dkim-signature;
	bh=m8O5AJ+7PrrH1vJisjcZNXDCqKKymath4j6nbieQ894=;
	b=ZfQmHtIbuAGVN8H6BuU2LRdZmGfwVSL6J1uQSlKEuH9J/cWOaPW6Wy2X5jytiAc6d1Kg7Q
	ZMaK6QbqF9Z+SqnUdjCkkaI6atUg4WwkMZBlWigAGh7ZvyJO+D7pzREWMRs5vyR4CYrk5l
	Eu59SfuGULw1Dq5/qjtzfmgGuJ4RvkQ=
ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1710230594; a=rsa-sha256;
	cv=none;
	b=6bxaXDGZcBBnfAj5Mvs0Z5mKhJcAQCnuohXb55Y30dTRxwNWoErRd9Qv6cBd736iE6v87A
	cRCYaEREQWgMc8kUvlw7IxkCdB/5fC2VyrQ3TkZT4CGoACVJ22PLZ0KZRi1ivlMn4YGpQL
	683Ovro5jYt/r0dP5byDnAJ/9LwDGkI=
ARC-Authentication-Results: i=1;
	imf21.hostedemail.com;
	dkim=pass header.d=intel.com header.s=Intel header.b=kzon94xU;
	spf=pass (imf21.hostedemail.com: domain of ying.huang@intel.com designates 192.198.163.8 as permitted sender) smtp.mailfrom=ying.huang@intel.com;
	dmarc=pass (policy=none) header.from=intel.com
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple;
  d=intel.com; i=@intel.com; q=dns/txt; s=Intel;
  t=1710230594; x=1741766594;
  h=from:to:cc:subject:in-reply-to:references:date:
   message-id:mime-version;
  bh=hUuRCZbJAfCL6ahmXMw92X3zLMAmqys0a5FggAtZv+g=;
  b=kzon94xU25N4Oim78A158foFaHWpxLO4W7u/7dMoLParoHjE7/MHq0Rj
   hYqAdl6yRDlNIwdQYdMaTbgrPuOjoELjyczMdzbYPDRXsF9AZNNmVdjVt
   gPz+gsGVXLkIR7vufFcobhoZ3/ah7WN3p0V8vxuSsSqanveEIG7N4WKS8
   FMYbgj9tLoJSAIXwOHuTLroelOEF1LRSi8l+xNH70017uho3JlPzeTbvl
   5ioVrdIDUa8K618U2/uhVh4BBMIouKzBIUxsd7nCAf/Yh5A7r7p6RbJzB
   FSAm5Z3WqHpvKkZcZbZ2qfS8mB6+o6bOFvLdfwx/7mx95CVhXrZpqVUQA
   Q==;
X-IronPort-AV: E=McAfee;i="6600,9927,11010"; a="22444292"
X-IronPort-AV: E=Sophos;i="6.07,118,1708416000"; 
   d="scan'208";a="22444292"
Received: from fmviesa010.fm.intel.com ([10.60.135.150])
  by fmvoesa102.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 12 Mar 2024 01:03:12 -0700
X-ExtLoop1: 1
X-IronPort-AV: E=Sophos;i="6.07,118,1708416000"; 
   d="scan'208";a="11364384"
Received: from yhuang6-desk2.sh.intel.com (HELO yhuang6-desk2.ccr.corp.intel.com) ([10.238.208.55])
  by fmviesa010-auth.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 12 Mar 2024 01:03:08 -0700
From: "Huang, Ying" <ying.huang@intel.com>
To: Ryan Roberts <ryan.roberts@arm.com>
Cc: Andrew Morton <akpm@linux-foundation.org>,  David Hildenbrand
 <david@redhat.com>,  Matthew Wilcox <willy@infradead.org>,  Gao Xiang
 <xiang@kernel.org>,  Yu Zhao <yuzhao@google.com>,  Yang Shi
 <shy828301@gmail.com>,  Michal Hocko <mhocko@suse.com>,  Kefeng Wang
 <wangkefeng.wang@huawei.com>,  Barry Song <21cnbao@gmail.com>,  Chris Li
 <chrisl@kernel.org>,  <linux-mm@kvack.org>,
  <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH v4 0/6] Swap-out mTHP without splitting
In-Reply-To: <20240311150058.1122862-1-ryan.roberts@arm.com> (Ryan Roberts's
	message of "Mon, 11 Mar 2024 15:00:52 +0000")
References: <20240311150058.1122862-1-ryan.roberts@arm.com>
Date: Tue, 12 Mar 2024 16:01:15 +0800
Message-ID: <878r2n516c.fsf@yhuang6-desk2.ccr.corp.intel.com>
User-Agent: Gnus/5.13 (Gnus v5.13)
MIME-Version: 1.0
Content-Type: text/plain; charset=ascii
X-Rspamd-Queue-Id: BB6231C0010
X-Rspam-User: 
X-Stat-Signature: riuacmsaguzecewndj1dmkdcyexgcdtf
X-Rspamd-Server: rspam03
X-HE-Tag: 1710230593-426183
X-HE-Meta: U2FsdGVkX1919KFnOBbpv5zgYp+vrAxgnh3x4PDQH8dW+sOeQT8cJi1R5GCp6TY6H88eK+tUCRUHqzY1Ke9p/qgp/CfZtaoFZuQgRDw0yjUumdqxfBP0ybOQRuXnaG2GZKnydVqrHXDrYz/KgPudnvPA4PN8fYUgbh7p191rATqFC1150e/XP2WJU+2vonRR2zgc8RZj0Hg3QT9gA9YC3tax/3vGgzhfU+LEuJIk2EiArffK6sVrVItjad7aJ6wwAkJ+5xkJWgrFWUWZKqW597/RGEvj9EkGpzE1rAY+sXIoysB3F1gEml1NVWU973Vpx+zOxWkwTZKP83c4m1+zasnVC9EuI/d03AiS68dDOwol5P6zONJlot7pQdCDVuFLUwgr+u5uRBBLhF+f+zV3QQPSOls83fquLdN+cG5y7AA7R58d/DhZFX8wo9wjOBTOERwvG6n3Yy7LBnb9/eaojPSvmShmmBanvg1yb4jyGqldShamgYn0n0YDHETvWpPs9jGg7xDDedUbFe2zFzJL9kU6tFxA890BiiIc2ObrGei66O/X3DI6LEjhU1Kf8TsQuxrVLHQ2CAGh92Fsp8ll5wCggBUp21fwu/MiHYl8Z6Vc2wChecMPtbxoxN2GOsvNMGHW0wO62SyAu6bwVWZg6SPguWKLzdlqMT3Wtvhg0BtAb5o0T6+jJGuW+088TxkdOP/BQi8OdP0QAGO9RoQFwV4Slf2Uiqbio6YCrPk4fNc8oyE/PUx02hOFwpDLsc2bST90CvH9tP6KXmLy7Sqdh+dzChXbUdtkkf0cAPKmUXX/ToNUtFzXOzItk9TMH7d8yMUZO3791mLh1V74NVS4teh5f6ifM4UY
X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4
Sender: owner-linux-mm@kvack.org
Precedence: bulk
X-Loop: owner-majordomo@kvack.org
List-ID: <linux-mm.kvack.org>
List-Subscribe: <mailto:majordomo@kvack.org>
List-Unsubscribe: <mailto:majordomo@kvack.org>

Ryan Roberts <ryan.roberts@arm.com> writes:

> Hi All,
>
> This series adds support for swapping out multi-size THP (mTHP) without needing
> to first split the large folio via split_huge_page_to_list_to_order(). It
> closely follows the approach already used to swap-out PMD-sized THP.
>
> There are a couple of reasons for swapping out mTHP without splitting:
>
>   - Performance: It is expensive to split a large folio and under extreme memory
>     pressure some workloads regressed performance when using 64K mTHP vs 4K
>     small folios because of this extra cost in the swap-out path. This series
>     not only eliminates the regression but makes it faster to swap out 64K mTHP
>     vs 4K small folios.
>
>   - Memory fragmentation avoidance: If we can avoid splitting a large folio
>     memory is less likely to become fragmented, making it easier to re-allocate
>     a large folio in future.
>
>   - Performance: Enables a separate series [4] to swap-in whole mTHPs, which
>     means we won't lose the TLB-efficiency benefits of mTHP once the memory has
>     been through a swap cycle.
>
> I've done what I thought was the smallest change possible, and as a result, this
> approach is only employed when the swap is backed by a non-rotating block device
> (just as PMD-sized THP is supported today). Discussion against the RFC concluded
> that this is sufficient.
>
>
> Performance Testing
> ===================
>
> I've run some swap performance tests on Ampere Altra VM (arm64) with 8 CPUs. The
> VM is set up with a 35G block ram device as the swap device and the test is run
> from inside a memcg limited to 40G memory. I've then run `usemem` from
> vm-scalability with 70 processes, each allocating and writing 1G of memory. I've
> repeated everything 6 times and taken the mean performance improvement relative
> to 4K page baseline:
>
> | alloc size |            baseline |       + this series |
> |            |  v6.6-rc4+anonfolio |                     |
> |:-----------|--------------------:|--------------------:|
> | 4K Page    |                0.0% |                1.4% |
> | 64K THP    |              -14.6% |               44.2% |
> | 2M THP     |               87.4% |               97.7% |
>
> So with this change, the 64K swap performance goes from a 15% regression to a
> 44% improvement. 4K and 2M swap improves slightly too.

I don't understand why the performance of 2M THP improves.  The swap
entry allocation becomes a little slower.  Can you provide some
perf-profile to root cause it?

--
Best Regards,
Huang, Ying

> This test also acts as a good stress test for swap and, more generally mm. A
> couple of existing bugs were found as a result [5] [6].
>
>
> ---
> The series applies against mm-unstable (d7182786dd0a). Although I've
> additionally been running with a couple of extra fixes to avoid the issues at
> [6].
>
>
> Changes since v3 [3]
> ====================
>
>  - Renamed SWAP_NEXT_NULL -> SWAP_NEXT_INVALID (per Huang, Ying)
>  - Simplified max offset calculation (per Huang, Ying)
>  - Reinstated struct percpu_cluster to contain per-cluster, per-order `next`
>    offset (per Huang, Ying)
>  - Removed swap_alloc_large() and merged its functionality into
>    scan_swap_map_slots() (per Huang, Ying)
>  - Avoid extra cost of folio ref and lock due to removal of CLUSTER_FLAG_HUGE
>    by freeing swap entries in batches (see patch 2) (per DavidH)
>  - vmscan splits folio if its partially mapped (per Barry Song, DavidH)
>  - Avoid splitting in MADV_PAGEOUT path (per Barry Song)
>  - Dropped "mm: swap: Simplify ssd behavior when scanner steals entry" patch
>    since it's not actually a problem for THP as I first thought.
>
>
> Changes since v2 [2]
> ====================
>
>  - Reuse scan_swap_map_try_ssd_cluster() between order-0 and order > 0
>    allocation. This required some refactoring to make everything work nicely
>    (new patches 2 and 3).
>  - Fix bug where nr_swap_pages would say there are pages available but the
>    scanner would not be able to allocate them because they were reserved for the
>    per-cpu allocator. We now allow stealing of order-0 entries from the high
>    order per-cpu clusters (in addition to exisiting stealing from order-0
>    per-cpu clusters).
>
>
> Changes since v1 [1]
> ====================
>
>  - patch 1:
>     - Use cluster_set_count() instead of cluster_set_count_flag() in
>       swap_alloc_cluster() since we no longer have any flag to set. I was unable
>       to kill cluster_set_count_flag() as proposed against v1 as other call
>       sites depend explicitly setting flags to 0.
>  - patch 2:
>     - Moved large_next[] array into percpu_cluster to make it per-cpu
>       (recommended by Huang, Ying).
>     - large_next[] array is dynamically allocated because PMD_ORDER is not
>       compile-time constant for powerpc (fixes build error).
>
>
> [1] https://lore.kernel.org/linux-mm/20231010142111.3997780-1-ryan.roberts@arm.com/
> [2] https://lore.kernel.org/linux-mm/20231017161302.2518826-1-ryan.roberts@arm.com/
> [3] https://lore.kernel.org/linux-mm/20231025144546.577640-1-ryan.roberts@arm.com/
> [4] https://lore.kernel.org/linux-mm/20240304081348.197341-1-21cnbao@gmail.com/
> [5] https://lore.kernel.org/linux-mm/20240311084426.447164-1-ying.huang@intel.com/
> [6] https://lore.kernel.org/linux-mm/79dad067-1d26-4867-8eb1-941277b9a77b@arm.com/
>
> Thanks,
> Ryan
>
>
> Ryan Roberts (6):
>   mm: swap: Remove CLUSTER_FLAG_HUGE from swap_cluster_info:flags
>   mm: swap: free_swap_and_cache_nr() as batched free_swap_and_cache()
>   mm: swap: Simplify struct percpu_cluster
>   mm: swap: Allow storage of all mTHP orders
>   mm: vmscan: Avoid split during shrink_folio_list()
>   mm: madvise: Avoid split during MADV_PAGEOUT and MADV_COLD
>
>  include/linux/pgtable.h |  28 ++++
>  include/linux/swap.h    |  33 +++--
>  mm/huge_memory.c        |   3 -
>  mm/internal.h           |  48 +++++++
>  mm/madvise.c            | 101 ++++++++------
>  mm/memory.c             |  13 +-
>  mm/swapfile.c           | 298 ++++++++++++++++++++++------------------
>  mm/vmscan.c             |   9 +-
>  8 files changed, 332 insertions(+), 201 deletions(-)
>
> --
> 2.25.1