From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 4D567C54E58 for ; Tue, 12 Mar 2024 06:18:01 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 7722C6B0190; Tue, 12 Mar 2024 02:18:00 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 722576B0191; Tue, 12 Mar 2024 02:18:00 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 5C2EC6B0192; Tue, 12 Mar 2024 02:18:00 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0012.hostedemail.com [216.40.44.12]) by kanga.kvack.org (Postfix) with ESMTP id 49C556B0190 for ; Tue, 12 Mar 2024 02:18:00 -0400 (EDT) Received: from smtpin28.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay04.hostedemail.com (Postfix) with ESMTP id CAFD91A0FB7 for ; Tue, 12 Mar 2024 06:17:59 +0000 (UTC) X-FDA: 81887381478.28.1348E0A Received: from mail-qk1-f172.google.com (mail-qk1-f172.google.com [209.85.222.172]) by imf21.hostedemail.com (Postfix) with ESMTP id 283F91C0008 for ; Tue, 12 Mar 2024 06:17:56 +0000 (UTC) Authentication-Results: imf21.hostedemail.com; dkim=pass header.d=bytedance.com header.s=google header.b="APsq/FPD"; spf=pass (imf21.hostedemail.com: domain of horenchuang@bytedance.com designates 209.85.222.172 as permitted sender) smtp.mailfrom=horenchuang@bytedance.com; dmarc=pass (policy=quarantine) header.from=bytedance.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1710224278; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-transfer-encoding:content-transfer-encoding: in-reply-to:references:dkim-signature; bh=fgGl/GezxVhRISLyx8tZ9PTH8m4h7yDB3m04Y7JMT5Y=; b=B7Pu9Z4V4UeviZyPED8I9sfkFMiajzn5I5Ow3enpZVTUXWMgUfe+AiA+DOOXHUJ/i/WFEM cSPrJcmNGrkkfdXhTp2eO14bR4EjXETdwkSGn8zOV/7IKHbmUxvN8GJQbO33T03RJ60M61 HHxtgt0aqPhCokIwIcpkhw7pEE43YqY= ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1710224278; a=rsa-sha256; cv=none; b=nDJ2YFn36bC/sLZYc3gDMQ3W7xVYy907D+e4LawciLYC26I6Brjo5QZW4qEJKwWiCLcbmQ 7vJOiSP3WZ40akcKy6YsL2qASKkcdLSCz6arM31p+B+q4lglcTOKGMZPs2jzPZ8Ay9sZTJ 4SGgQ352Yu5+y2wH2VbDgd+X8Baqw5I= ARC-Authentication-Results: i=1; imf21.hostedemail.com; dkim=pass header.d=bytedance.com header.s=google header.b="APsq/FPD"; spf=pass (imf21.hostedemail.com: domain of horenchuang@bytedance.com designates 209.85.222.172 as permitted sender) smtp.mailfrom=horenchuang@bytedance.com; dmarc=pass (policy=quarantine) header.from=bytedance.com Received: by mail-qk1-f172.google.com with SMTP id af79cd13be357-783045e88a6so318255685a.0 for ; Mon, 11 Mar 2024 23:17:56 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1710224276; x=1710829076; darn=kvack.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to; bh=fgGl/GezxVhRISLyx8tZ9PTH8m4h7yDB3m04Y7JMT5Y=; b=APsq/FPDVHdsWzNUaipSRTOQdpkjVB9A+GOxZNo02iyPhack9OT2u9iGYreU+g8JJO QA98tqSRHY0aEPUqCRx3G3iO7rmnAia6O79uGTxR6nHYcSvarbXAQQr6m2P7jKVRIHIM 3wqffrbIlNYdf6TttR2uO+HY2fA1gy3Yd1Jwj9JffYfDBKkrJqJ53s1tQABlwNVVw88R +iQglNs26BIasDo6edbEgZsRr0QCdogon3xzW57nRgWBD1851Ynlf/QN2yJtHS9ym7kt guuN53T84RICWCP5uNIPMd/Db8X07XwkPpATRyVr9zVdsRoQh9xFTGj33m+xRGUj/606 kXQw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1710224276; x=1710829076; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=fgGl/GezxVhRISLyx8tZ9PTH8m4h7yDB3m04Y7JMT5Y=; b=WRaF/Z5AyiRvdCUCnd2sRacNHWGiCDxZkrBWnt3HAqireEBqDqUk5lKVXcFa25WTy4 bmvjebkRUPxL79g7ch6327FxpPXPtdh8HTP2zQZ8adZu6OwmIVjbQAHs33/F8F/+CBw6 vu6ciMKQFve0wItm2QlD9gcMCqjhyBdWY3LuVykDP+ABdqzg3k25kvE2couYElJXVM1A EEHcLuf+ITrPLV4NBnGK7y1a/6ItAxfT3chG+7Dkxh/3x07rtOu230liSk+YzIMGOVyt efi1PDDlKUTtHCb0XDP4+5jwLaLy00wuTH1q9qmYeTQ9WjaXL1ZcKd1Kn0+PQKMXTsdt B4Bw== X-Forwarded-Encrypted: i=1; AJvYcCWFzWnSN4g3RSsuOHOfyiGvDdffnNyGWiwyi0Zyo9Fcs7uGvFlyWuiCi1jgS50ZzLtB+l7eoDSNWC8sg9a9fdbjKg4= X-Gm-Message-State: AOJu0YzGLmwdsUF6BjrXqEXXWJmg1ZHIf6dZ9M1B5avNjE4DYDoz8/F5 eCOIgn+qhlovoj78OjltShzptBRSruplC6tVnXfSyXgh6w35PfhmopiXTuU5OuU= X-Google-Smtp-Source: AGHT+IGQVjf55XaSaA7QRv11XERHDa0ikizCRigRPLUcQWIpU3ODAYNJ7IKwN9aaMW4bZrAoONZoqw== X-Received: by 2002:a05:620a:4494:b0:788:7dc5:cf8f with SMTP id x20-20020a05620a449400b007887dc5cf8fmr499319qkp.35.1710224276179; Mon, 11 Mar 2024 23:17:56 -0700 (PDT) Received: from n231-228-171.byted.org ([147.160.184.133]) by smtp.gmail.com with ESMTPSA id m18-20020a05620a221200b00787b93d8df1sm3394396qkh.99.2024.03.11.23.17.55 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 11 Mar 2024 23:17:55 -0700 (PDT) From: "Ho-Ren (Jack) Chuang" To: "Gregory Price" , aneesh.kumar@linux.ibm.com, mhocko@suse.com, tj@kernel.org, john@jagalactic.com, "Eishan Mirakhur" , "Vinicius Tavares Petrucci" , "Ravis OpenSrc" , "Alistair Popple" , "Rafael J. Wysocki" , Len Brown , Dan Williams , Vishal Verma , Dave Jiang , Andrew Morton , Jonathan Cameron , Huang Ying , "Ho-Ren (Jack) Chuang" , linux-acpi@vger.kernel.org, linux-kernel@vger.kernel.org, nvdimm@lists.linux.dev, linux-cxl@vger.kernel.org, linux-mm@kvack.org Cc: "Ho-Ren (Jack) Chuang" , "Ho-Ren (Jack) Chuang" , qemu-devel@nongnu.org Subject: [PATCH v2 0/1] Improved Memory Tier Creation for CPUless NUMA Nodes Date: Tue, 12 Mar 2024 06:17:26 +0000 Message-Id: <20240312061729.1997111-1-horenchuang@bytedance.com> X-Mailer: git-send-email 2.20.1 MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Rspamd-Queue-Id: 283F91C0008 X-Rspam-User: X-Stat-Signature: orohqkana84yctgpfm99pwx1d4inr1t4 X-Rspamd-Server: rspam03 X-HE-Tag: 1710224276-762389 X-HE-Meta: U2FsdGVkX1/TozZEPJLI9GIpdUTxQbuBSedFx7qFuFubjn1geEFTkNGPIjmE7eeRkPxrnp/F+Y/y545O3IGtZ68I9S/xtmVl8qNrfyh9W9l3/VBLO6uBudv7VkCNDWbVCX12IfSB5rUX+abSbmiGWxEit3Kz09zNT6hZU6saldzg4y6xAcbnqNsv3ZEi05XrTLXoJb19rmXhOJjbk7QNR/+p2oRvbI93qRkEBIzGEKE54eaDS3ZXk6nq6QtiHP3gqu/jAIWpc8OaF6Kp2ZMflL+yNp9+42jZqTjmZUuf+QnE4vM696dYU0oyk8LrHSYfXUdFfkEQ7BZV1TyDcJzqggBsXJvGMv6GYuy/quIrBZlocseHPxOt8NTo8ISB5d3HbVOklEJg3wm9FK6Nh/Fvn4G3B4zZl+NaSlGP6STSWJ5nDHIaGD+OnM2qhgyEwMbU3nnzjuMX2nZ4UGYNfy4S6YrnRlRVWlCnzBADKDswXQ2JZtlIELsKigRbk6qsRqjj/tGDD8SsVUI3nvC977os/lwC/LPSyXQ7hYy9cSdpDMRIYmmCowq1ga0b2Xb816xAHNrAQhaqSzAlTvY9IHD8E4tnsvQ4+olpWewzzHxOi+zHRm/eeLW26LpOIXgX8d9c6f/ud3Lr0/xoKJsC6hVhk/joctZCh28tkCeeqLbotHASgPvhJYN45No9Xt6EVCR24ua7flLmcDntwfKfAAd/LbfW+cckjNgMfpg8tW86pZxh1+FiVCB1GDZPxLW7+XMf4Tu8WZ7/eDHVx73Pm2DOgWakAWFgHF3WPLMznDZ2LgLDas30TSD+84IfbqAI/ftcCD7Mj8YSRhky/DVu9k2Tk4K2K0lphGSfZ7R/LuXMdObElcsoRDAM/h9xHje0UiS9Yh0+W/H0daz6ThWdS+97pX5rydjprXP5YngVi2SJQvqTsXkFTgyxTz+ujKhD3H0Ynh/2KdHjOuen4c3J5Dq A0Mlohbq ibNLK9BvuuU0rYAJi/0ePoUGQko58yG/bivgj8nWuTrKTnqby13ln/d27Svq7OsEE2aEtq1nd+gyAa0eFVFoqjALqcsMT+nKDk2xUXjD7Hcdd+F0g1tGjzazwS6z3Wm85+nOKHFMchU119ARJbcWHiqS3qemjCiqYEpwpGUNbigYSdbShUncgxVpKdLRZ0cDp9BGAxzE67tBfvu3CsW9JqqoclubE7MwwUbAACyGI6L5R6unf2ohjj+FRSS3hElWr/Ckpxhom4uKDtsOdYWauMNWJZiPK4OCJUvwybLFmnnwFbvEGWFtqsGYQfOVAYO+05/VKgfaDsQiZjcsGZXjBjJILPiktGAYN7pxgQnZJKAUosGYdpVN24PnbyGal8UWmbpBhy+xE7HF+lGZs1EftnJQOYDWKo+ZyyzvcjMJlfgvmYDYD2ibS0H7h6n50hAKtskMJ6lpS7us/ChQyNS5xKeqA+4n3Um2EIIgoX85+USf0c0axVSqH8e9PPpZZxAI255P1pOz3I9lVqIDab+2y1JeJC+5eAnAu/BzSM5N+GCYnWOXVx6DWZy0D29C7HJs2EfIM X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: When a memory device, such as CXL1.1 type3 memory, is emulated as normal memory (E820_TYPE_RAM), the memory device is indistinguishable from normal DRAM in terms of memory tiering with the current implementation. The current memory tiering assigns all detected normal memory nodes to the same DRAM tier. This results in normal memory devices with different attributions being unable to be assigned to the correct memory tier, leading to the inability to migrate pages between different types of memory. https://lore.kernel.org/linux-mm/PH0PR08MB7955E9F08CCB64F23963B5C3A860A@PH0PR08MB7955.namprd08.prod.outlook.com/T/ This patchset automatically resolves the issues. It delays the initialization of memory tiers for CPUless NUMA nodes until they obtain HMAT information at boot time, eliminating the need for user intervention. If no HMAT is specified, it falls back to using `default_dram_type`. Example usecase: We have CXL memory on the host, and we create VMs with a new system memory device backed by host CXL memory. We inject CXL memory performance attributes through QEMU, and the guest now sees memory nodes with performance attributes in HMAT. With this change, we enable the guest kernel to construct the correct memory tiering for the memory nodes. -v2: Thanks to Ying's comments, * Rewrite cover letter & patch description * Rename functions, don't use _hmat * Abstract common functions into find_alloc_memory_type() * Use the expected way to use set_node_memory_tier instead of modifying it -v1: * https://lore.kernel.org/linux-mm/20240301082248.3456086-1-horenchuang@bytedance.com/T/ Ho-Ren (Jack) Chuang (1): memory tier: acpi/hmat: create CPUless memory tiers after obtaining HMAT info drivers/acpi/numa/hmat.c | 11 ++++++ drivers/dax/kmem.c | 13 +------ include/linux/acpi.h | 6 ++++ include/linux/memory-tiers.h | 8 +++++ mm/memory-tiers.c | 70 +++++++++++++++++++++++++++++++++--- 5 files changed, 92 insertions(+), 16 deletions(-) -- Ho-Ren (Jack) Chuang