From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id BCF14C54E71 for ; Fri, 22 Mar 2024 07:04:13 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 392106B0087; Fri, 22 Mar 2024 03:04:12 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 2ACFA6B008A; Fri, 22 Mar 2024 03:04:12 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id EDB1E6B0088; Fri, 22 Mar 2024 03:04:11 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0015.hostedemail.com [216.40.44.15]) by kanga.kvack.org (Postfix) with ESMTP id CB37E6B0083 for ; Fri, 22 Mar 2024 03:04:11 -0400 (EDT) Received: from smtpin06.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay08.hostedemail.com (Postfix) with ESMTP id 8CF5C140BE5 for ; Fri, 22 Mar 2024 07:04:11 +0000 (UTC) X-FDA: 81923785902.06.9BBC6F9 Received: from mail-ot1-f47.google.com (mail-ot1-f47.google.com [209.85.210.47]) by imf28.hostedemail.com (Postfix) with ESMTP id 73546C0017 for ; Fri, 22 Mar 2024 07:04:08 +0000 (UTC) Authentication-Results: imf28.hostedemail.com; dkim=pass header.d=bytedance.com header.s=google header.b=jQ0uuaiE; spf=pass (imf28.hostedemail.com: domain of horenchuang@bytedance.com designates 209.85.210.47 as permitted sender) smtp.mailfrom=horenchuang@bytedance.com; dmarc=pass (policy=quarantine) header.from=bytedance.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1711091049; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-transfer-encoding:content-transfer-encoding: in-reply-to:references:dkim-signature; bh=q7neNI4OXUjIQ5rL9BmGxx5XZjsprNGi1CqpvhfOJ5A=; b=Q014eGyQQQq8j7bRIb5itl1PXhSplelCdh0OMUGxy7Rt5Sl9cJJL7xZBvy4ALwUuJm7tZS a/BRutqYXoh0rsaeX7xTxoz7cW0/m3i+qKHdDOuXq5e53LlIyCI2DMosu1Rj2b/TAIm4h+ E1qKNdRTjtkh5sanuwn2a8L/yXgzn4w= ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1711091049; a=rsa-sha256; cv=none; b=2MBqd3hEKP7VRi3KV/gs7/YgAzLRQY+WTIGSQ4ZbdYUzMiKF0XFizTMEOtPKeDLOTSks/3 +3cxslvS9dCvmT80f3BH8luqRA3Fd64EPH5UGr3Sy0bsZjw7ebwmmWoCjwceYNoaFjR7tA CawlSVkOBAGKy+RNAX//ih5bh/TA85I= ARC-Authentication-Results: i=1; imf28.hostedemail.com; dkim=pass header.d=bytedance.com header.s=google header.b=jQ0uuaiE; spf=pass (imf28.hostedemail.com: domain of horenchuang@bytedance.com designates 209.85.210.47 as permitted sender) smtp.mailfrom=horenchuang@bytedance.com; dmarc=pass (policy=quarantine) header.from=bytedance.com Received: by mail-ot1-f47.google.com with SMTP id 46e09a7af769-6e4f7975121so385424a34.1 for ; Fri, 22 Mar 2024 00:04:07 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1711091047; x=1711695847; darn=kvack.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to; bh=q7neNI4OXUjIQ5rL9BmGxx5XZjsprNGi1CqpvhfOJ5A=; b=jQ0uuaiEBqBPlzn98E7h4VVRuP2T70wooYh/9+q3QCzZpN70wPH3uxwKTigbz9LyTu g0iUqJcs8AUtpwTkmaw7NejjafhXgCedGvPvy08bzAbeCu8tpnHD4OqxNpnxEX79Z2X3 5CS3psTTHCXHAM2L+1gmFPjEUzV8gCMAXnmoKAZ9PJ9yWDFHDoUF1NtaxxN/AUZIaQp2 kso6C3BCAp4znFmCZi95CBPI0jyk44BmoTuJNLFwxORrrqmBzhAW43JAabg1D2phT459 iWNnPbPTAlJhOATYXQYDxrbY/b4u9l+pmPYP//aNTHekp3JUTUdxQWgvaNuArvEVeHGX mIyw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1711091047; x=1711695847; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=q7neNI4OXUjIQ5rL9BmGxx5XZjsprNGi1CqpvhfOJ5A=; b=B3B3cStMsydwHJqj4CkrxFcqLmBm7s6BVb250Y2pNPauFHa08FEzkY3HI3COBopqeS EO8c8Mm7Nv8o7A+7FaWjQMZIz9oeCCwdQwON6gSk+zmyVCJSU7JVWnZW1lUFYwdswaP3 zLEEh1Ncz4Wt9vrRwvBIGIgsQbzlhkyn6RcAWKte2Y8ccePwdLIoy7ULXnrPeyqTfXUu MTGtUPCUZhTHi3lrmbJLEGICEAaRj8txdCskdoz8Xksn8fCtMNFglhqM6uOItM3WmtGP +5cfjAWW72CQgT3N1pZm/BCeu+zDHhNg0G3yJGE9UlbBD416D5C160RcWU+emTimar7X fgBQ== X-Forwarded-Encrypted: i=1; AJvYcCVgGxl71+Yj/mj3qDC+ye8ogE7FB+HqWhgI+mBQ5cJwCeM0MamkkiTGLUluRFRt+0A60MDvm2G+dXum+ir2dMACx9M= X-Gm-Message-State: AOJu0YzGhflntuzQTLIpp136texFKyB7M4VFwjcAKVsnRIBG8w+TZnfI flv/Y41drsc+h1s42dx7xDe2wSVmiwOR18qTEUnSMG/0Qe0FMQeUaF6rcOWIqtw= X-Google-Smtp-Source: AGHT+IGDL07edZqKGbxe2OTitulABo89PW1+afEDf8VZk/5xbSmUkIHDwnIlESs+qhqPqacblb6/Pg== X-Received: by 2002:a05:6830:87:b0:6e6:7b48:b801 with SMTP id a7-20020a056830008700b006e67b48b801mr1669775oto.16.1711091047143; Fri, 22 Mar 2024 00:04:07 -0700 (PDT) Received: from n231-228-171.byted.org ([130.44.215.108]) by smtp.gmail.com with ESMTPSA id k1-20020ae9f101000000b00789fc794dbesm553974qkg.45.2024.03.22.00.04.06 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 22 Mar 2024 00:04:06 -0700 (PDT) From: "Ho-Ren (Jack) Chuang" To: "Huang, Ying" , "Gregory Price" , aneesh.kumar@linux.ibm.com, mhocko@suse.com, tj@kernel.org, john@jagalactic.com, "Eishan Mirakhur" , "Vinicius Tavares Petrucci" , "Ravis OpenSrc" , "Alistair Popple" , "Srinivasulu Thanneeru" , Dan Williams , Vishal Verma , Dave Jiang , Andrew Morton , nvdimm@lists.linux.dev, linux-cxl@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org Cc: "Ho-Ren (Jack) Chuang" , "Ho-Ren (Jack) Chuang" , "Ho-Ren (Jack) Chuang" , qemu-devel@nongnu.org Subject: [PATCH v4 0/2] Improved Memory Tier Creation for CPUless NUMA Nodes Date: Fri, 22 Mar 2024 07:03:53 +0000 Message-Id: <20240322070356.315922-1-horenchuang@bytedance.com> X-Mailer: git-send-email 2.20.1 MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Rspamd-Queue-Id: 73546C0017 X-Rspam-User: X-Rspamd-Server: rspam11 X-Stat-Signature: ucw99okuzkt9dqe3h6jwsfqujogm75kc X-HE-Tag: 1711091048-632259 X-HE-Meta: U2FsdGVkX19RqeBBCrEujClAUiAHvocLyxF1VkQ3RCiX0XSDGqA9YwnDOB9EKakGHoTg9SEqo85pQJAoY6UuQGksiZjP+yWtQcFjpUrUlMnPffRxZsaGMqa6IFH1biXd9CSTSZzTu7mIp2qVWdaUS+MHnpCG+qhyQpyDcrKDFNWDaATHBRdBVkYRdRRBE2KyjfU4dAtkX//+rWMznuq+9N3ZMeyUzUJBdDjQ1lxlY3EgmCEVU94WfH2ykvP8eEl4t2rnc+jJvplfl1o73X9hhbx1ybOcY8ectWii5hvm2UqeBb5HW1hh+BU52u0gH0UraDFuoaWP1XaF+kRHnb1zpOfMg5pfJilQ93Z5SQSHx5WKdgvv135kmqEe8+lW8Tu780d4H7WtBpfNQJB1aXrnBgV35sKUx/r9WV4JfAzHh5OvEN4VB/J0GeaGqW6lFqiBe05L3+aR/CgrjisXvO3DcryGuJXZOc7QMIrb3tR1y2eSBFZw8El1NrZM6DRY4RLEbZwSgEfBoo119ISCxK/ljdq6sokDBQPs7BUDA5UMxy4MALZ0uHBSnRhQYNqtAzvqzFYB69YGLbMgNsbtw0NnJYYZ3vU7j3jICEyOBphNvx+cQAkTEBSuWQ8bKVRiP1IL/HkmN9mt2S2FnQzyRnEXO+7bFj2mw4qDdf27K8r+gkKIxqe0v/Mbncz5cQgrsqa0zNoRFoFuK14z4BgnC1wscWelj+SpZdIShmHlqUYrqCt7SpxxCp2EbXbYeBgDiK0GHpQ07UnJeKTbaJrKmY0KbyJAXzDFtO+dW17HkHk+SO/S5kdNt366VhwwkGuMzwa0bwzkkzpiJIVxRtHkCAZyI1AG54rVOSdyxVZiEEs53PLyo3s4zw1dPrhOfBgWW6+tNRHtlKgSHZr1ZLWI2or1sr4Aavxb2I3EIsBUEO8hfNV41yCUR6EOy7PGANq11SXx+iR4JdpDCqUrQpogOV1 EeLT6RsC gwnAg8mf+nh9eQJe1vX1bgw2roivx1mtzE/xtSvST7NpxYyJxD/GlQh0kybFNMJDKZpxQgRXjcLfPCDhi4T8eXvcElslKsvY5/59XxpSTjROVNWagOwKg+AGga9FwkfJ5XndUYBca20GGndDCP6jtvUfQ/DD+SNfJDGXJhA3qZHz18hMaR8oog+Ww8msH5DP1KVyHfoI4ch39elFoGpwwv5zy+RioyINv0HyBAtHkD1AaNVa2OFAuykq5/2JnNHDZ/u8cGSdSqWAX6bc5gd9AkEQ2efr+Vm3nmjq6hK1NDezBRwfDfsDFQBAFNhwX2cvJ+H/dnWGHY8mmMzP/DUa4ESDhJFWsrk7g3hIkuwr0GmYg+vptBGAH9voWvq/6zAc6Oew96coTrVCOhCgdnxWYD4fj7MWNpYgv+0ZxlO9eUx+EUC+Ur7CDkc04IRTSwH8nn+Fuy1fmEIgCgLiRJ5arxhgHLab9b+FWkAtktR++UwONMsxsARQPhJw8suP6A6hcXGnVs2efOprSZ1pal4Itlu4kjU9dhMBqZMnROmRvqib/frgSJyC8ErW6x76SRjGmJdUw X-Bogosity: Ham, tests=bogofilter, spamicity=0.000001, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: When a memory device, such as CXL1.1 type3 memory, is emulated as normal memory (E820_TYPE_RAM), the memory device is indistinguishable from normal DRAM in terms of memory tiering with the current implementation. The current memory tiering assigns all detected normal memory nodes to the same DRAM tier. This results in normal memory devices with different attributions being unable to be assigned to the correct memory tier, leading to the inability to migrate pages between different types of memory. https://lore.kernel.org/linux-mm/PH0PR08MB7955E9F08CCB64F23963B5C3A860A@PH0PR08MB7955.namprd08.prod.outlook.com/T/ This patchset automatically resolves the issues. It delays the initialization of memory tiers for CPUless NUMA nodes until they obtain HMAT information and after all devices are initialized at boot time, eliminating the need for user intervention. If no HMAT is specified, it falls back to using `default_dram_type`. Example usecase: We have CXL memory on the host, and we create VMs with a new system memory device backed by host CXL memory. We inject CXL memory performance attributes through QEMU, and the guest now sees memory nodes with performance attributes in HMAT. With this change, we enable the guest kernel to construct the correct memory tiering for the memory nodes. -v4: Thanks to Ying's comments, * Remove redundant code * Reorganize patches accordingly -v3: Thanks to Ying's comments, * Make the newly added code independent of HMAT * Upgrade set_node_memory_tier to support more cases * Put all non-driver-initialized memory types into default_memory_types instead of using hmat_memory_types * find_alloc_memory_type -> mt_find_alloc_memory_type * https://lore.kernel.org/lkml/20240320061041.3246828-1-horenchuang@bytedance.com/T/#u -v2: Thanks to Ying's comments, * Rewrite cover letter & patch description * Rename functions, don't use _hmat * Abstract common functions into find_alloc_memory_type() * Use the expected way to use set_node_memory_tier instead of modifying it * https://lore.kernel.org/lkml/20240312061729.1997111-1-horenchuang@bytedance.com/T/#u -v1: * https://lore.kernel.org/lkml/20240301082248.3456086-1-horenchuang@bytedance.com/T/#u Ho-Ren (Jack) Chuang (2): memory tier: dax/kmem: introduce an abstract layer for finding, allocating, and putting memory types memory tier: create CPUless memory tiers after obtaining HMAT info drivers/dax/kmem.c | 20 +------ include/linux/memory-tiers.h | 13 +++++ mm/memory-tiers.c | 105 +++++++++++++++++++++++++++++++---- 3 files changed, 110 insertions(+), 28 deletions(-) -- Ho-Ren (Jack) Chuang