From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id B0B16C3DA7D for ; Thu, 5 Jan 2023 10:19:09 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 4D798940007; Thu, 5 Jan 2023 05:19:09 -0500 (EST) Received: by kanga.kvack.org (Postfix, from userid 40) id 486738E0001; Thu, 5 Jan 2023 05:19:09 -0500 (EST) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 301BB940007; Thu, 5 Jan 2023 05:19:09 -0500 (EST) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0015.hostedemail.com [216.40.44.15]) by kanga.kvack.org (Postfix) with ESMTP id 166808E0001 for ; Thu, 5 Jan 2023 05:19:09 -0500 (EST) Received: from smtpin20.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay08.hostedemail.com (Postfix) with ESMTP id E995A140217 for ; Thu, 5 Jan 2023 10:19:08 +0000 (UTC) X-FDA: 80320347576.20.6CBAF7B Received: from mail-vs1-f73.google.com (mail-vs1-f73.google.com [209.85.217.73]) by imf08.hostedemail.com (Postfix) with ESMTP id 5897516000A for ; Thu, 5 Jan 2023 10:19:07 +0000 (UTC) Authentication-Results: imf08.hostedemail.com; dkim=pass header.d=google.com header.s=20210112 header.b="mO/o/8j5"; dmarc=pass (policy=reject) header.from=google.com; spf=pass (imf08.hostedemail.com: domain of 3GqS2YwoKCGEISGNTFGSNMFNNFKD.BNLKHMTW-LLJU9BJ.NQF@flex--jthoughton.bounces.google.com designates 209.85.217.73 as permitted sender) smtp.mailfrom=3GqS2YwoKCGEISGNTFGSNMFNNFKD.BNLKHMTW-LLJU9BJ.NQF@flex--jthoughton.bounces.google.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1672913947; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=wyTgKoUPaqLvPlKgkL6dENYZlNeSHJ5X+BJXc4mR0WE=; b=sLs05vGye4+n1qa40z7+TnoQyP+0ILQUmNu7QOsyNBp6YcfeSoXg8l2gsM508scuzujb69 d84t6s3wz0BoosCSg7OxgeviONd1aIKP+yXlVqHjS+Y116LG8ZrWt+GnEy6R229t8NaaXw 7451PIhSLnaGwPitF8jsm9MIE/WcHTQ= ARC-Authentication-Results: i=1; imf08.hostedemail.com; dkim=pass header.d=google.com header.s=20210112 header.b="mO/o/8j5"; dmarc=pass (policy=reject) header.from=google.com; spf=pass (imf08.hostedemail.com: domain of 3GqS2YwoKCGEISGNTFGSNMFNNFKD.BNLKHMTW-LLJU9BJ.NQF@flex--jthoughton.bounces.google.com designates 209.85.217.73 as permitted sender) smtp.mailfrom=3GqS2YwoKCGEISGNTFGSNMFNNFKD.BNLKHMTW-LLJU9BJ.NQF@flex--jthoughton.bounces.google.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1672913947; a=rsa-sha256; cv=none; b=wosTFv7wqvtrMo1VUyksZUDnRPJljZyuHZ3JbNecQzzXuqsRMqR/ULOSSeawklb7h4Xzl/ 2dbFA83MoFQSIKYrZucV5eja7QVz5LtYbEXkx025fbfo7feK9b2FDcmwY8Hbjzsxh7qOc5 bn2sB8qrulQs8cZbL+RgZahjSkLYcPI= Received: by mail-vs1-f73.google.com with SMTP id v127-20020a676185000000b003c95257942bso5707545vsb.20 for ; Thu, 05 Jan 2023 02:19:07 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20210112; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:from:to:cc:subject:date:message-id:reply-to; bh=wyTgKoUPaqLvPlKgkL6dENYZlNeSHJ5X+BJXc4mR0WE=; b=mO/o/8j5uzBbSG/A9cnee+je9u5JCkoj3aSd2P7vs7zhonYorgkMtjaT6Camo16MRO HNtN+6XL9Mv6MTsGSvNsS0ThZoOWReEOx+mlNlqsi2BHHzTVLxaSyppMdsTCWcfSc5jK 6+ZKrfvxX1DRZ4xkYZFJxsbNBzLnTyhaq61sRVQmunxV1fBNLuDxQ/deFSB6JHPEpTUt 5/tku/ASVRacTDjSteTyVtqapAitjwzsKEVznkEuAhKi0Dyl3YHpk9+wTziB0iAhpapX LWhXUWqa8fiTZGxfOcvuaDuvdtQuOtDLPsve4S6HaeJ5gJdlA8bacDMpeFwBpswodRai BbcA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=wyTgKoUPaqLvPlKgkL6dENYZlNeSHJ5X+BJXc4mR0WE=; b=r2A7mSmJgHfLr2GxapKcfl6OBsXpNgZ02edhflZ3uDYq2V8tH1zUQwraxOV8sWfAbL qXWmoxBYDtK/O9hwqtinusmA5s1AJMD6HDdfsr9417XmVAXkctIKupuabChPMZj0U8jn d+XtmfV0HZSrbgoQXYHT8ioC4EDqSCP94U1HpHBEIEc0nz4dIsSpHult85Bj7k3TGypc JhwXKyJg+UJ/HR8UR2P4YOSCj4xFQ97es4dfx7ty0dBZZe7xjGrHhdv8hm9+jb6+lweL 1W0oewwYF7044IHWDkcX2Zbomkn3/TRntMAdefwyxVzgYFvZg3Un19ZDqwKojoy1dUv1 OudA== X-Gm-Message-State: AFqh2kqKzg1UCfjFLJV1xBlR3E9u2ZD8JXtx/DdlDjDEJdmYOF8Y5PJT euAeRivzWIELXuP5hEFnN8UePjhbZsKJZQPD X-Google-Smtp-Source: AMrXdXvO4uC3f87rB5JMezqUX+x18qYnUlhtx11jU+WhbWs+qlfQayBtMoiabh7QOd7CQ701U3D2KnBikpjvwFJS X-Received: from jthoughton.c.googlers.com ([fda3:e722:ac3:cc00:14:4d90:c0a8:2a4f]) (user=jthoughton job=sendgmr) by 2002:a67:eb15:0:b0:3c6:5f9:689d with SMTP id a21-20020a67eb15000000b003c605f9689dmr4653326vso.8.1672913946620; Thu, 05 Jan 2023 02:19:06 -0800 (PST) Date: Thu, 5 Jan 2023 10:18:07 +0000 In-Reply-To: <20230105101844.1893104-1-jthoughton@google.com> Mime-Version: 1.0 References: <20230105101844.1893104-1-jthoughton@google.com> X-Mailer: git-send-email 2.39.0.314.g84b9a713c41-goog Message-ID: <20230105101844.1893104-10-jthoughton@google.com> Subject: [PATCH 09/46] mm: add MADV_SPLIT to enable HugeTLB HGM From: James Houghton To: Mike Kravetz , Muchun Song , Peter Xu Cc: David Hildenbrand , David Rientjes , Axel Rasmussen , Mina Almasry , "Zach O'Keefe" , Manish Mishra , Naoya Horiguchi , "Dr . David Alan Gilbert" , "Matthew Wilcox (Oracle)" , Vlastimil Babka , Baolin Wang , Miaohe Lin , Yang Shi , Andrew Morton , linux-mm@kvack.org, linux-kernel@vger.kernel.org, James Houghton Content-Type: text/plain; charset="UTF-8" X-Rspam-User: X-Rspamd-Server: rspam02 X-Rspamd-Queue-Id: 5897516000A X-Stat-Signature: r6h81aa7q1t7797yfda7kc5aacozk3es X-HE-Tag: 1672913947-645801 X-HE-Meta: U2FsdGVkX18k/yrqFfjC57J9hpy+D9eU5MCX3zlt9ugqr7mCpDCo/EDMxXhExR28Dlhi9Hrfl/iuC4B09a76yKpR6Yhi5tt0GcFsu4FOop4rQeK7WoQFIa460UeI80t10JdD1qulFEHeRhtV2IV1g/SITqohDlkn7T3X8feymdtcL1rF6Dm4sV9757fLDvAjE8m1lAJdsa0uS29zJ6PgilmnBxBk5F57tST/fUC6bFrW9Sp/0heWdojGrME3SIwuEvbZ1hC4BjSa/iyNZzTHvC9OnLh3xol4/VH1rhxJ9UpRA77Jl2ZGpNY0Jt1ZGxtv15HCC72S2BQDloEXWGdmrAMDyqfk4Q0eWoHyloojG6I2c8rKpskKm393UQ3oSDjs0Ql57oj/V00B52jVCA0jx6cHUVaqZOY4RGWHtfI3Tqx3i2yjPm6eLyE5xcCrDIGpEHhI7tFwgeoPa25lLo8Yv6+8zsagYx6j/R52CRaIjH9QoWlhaGZGCAMqi+SR1p0nPMIN0BpnIhaUc7TfUrWtWXQt5bHNyfVSeinv77TT2+CyBUiLHoritQ7NxYmjfqp9QfTt1mIv1nsHNX6TZSdh0LK3QGWpDJR2DVfB6ujlVsE84l+GehkHfReHSbrdbB4vH1tNLXpRV2uK+E08eRK1s54kdY13t1aLg2lFSaAoh/NrYkW2dUsvOSDI9xcfYvQFhgoGPPFpZQMfZH8PtvIkqMI1L4GS3OJfP5vMRe+XtUq5O4Merc9GWeGzFKAucMG0QL2P3f2aYnD8Mkgl0wgaQlQfb2KMuqspCbcSlcqmkS9bYEPE7AtFZuHneQmvOQWVTGzArV7rQCls/K4TUeuoF3ycgFce11ytBkRzBqzSuZj1HSfSP9lIBgAxM9W80OUalmzFQ+WVdPrEpeKxJQ8DIEyj82s+K5PjYiSkon8qE8QFhfRzxfE+5S/9YMHpJQRPdHAP0SirwAQhpFtNcl5 +zcUtMqI xWRlVUIBz8UUrS0vk5XFz5ax78d/RWZR9UGDQ3YPTVbsovE/ofM6G0I167hCv/Tni46jD+6/xfDVPASH8LWREhWhblqd2+aovPe8GURcCStO9VgNlVZIFWt6/qDrNf3dHOOEXHcrCTBdLYvfssGbEKNJVf9JyctX7jzf1yVZTL7RgfR41UqySjJ4/5M2rS3D+j5VMqQiXiv0MnixagriUKlk3eWM5XriiWZZsJr7x+H42JXpKogPZYa8Ulg7VUrC9MMdKTIltAxkHYIRhyGBSudfQpHnmxLuYvmBbGRUaMhWGWSbdzBnk0R37SYgCBgd0a2hYOrY4JseHsJUz82h234c3aYcMfDk+uhNQThJCOyTfDlsfz5YM9PXSLHFGjARsM+Qh5a8RXwSevGMyQgPyZwLJ8sw1Fwp7nu5LxOVRdssLdiiPijVVHpATWsolHCRr3km/ X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: Issuing ioctl(MADV_SPLIT) on a HugeTLB address range will enable HugeTLB HGM. MADV_SPLIT was chosen for the name so that this API can be applied to non-HugeTLB memory in the future, if such an application is to arise. MADV_SPLIT provides several API changes for some syscalls on HugeTLB address ranges: 1. UFFDIO_CONTINUE is allowed for MAP_SHARED VMAs at PAGE_SIZE alignment. 2. read()ing a page fault event from a userfaultfd will yield a PAGE_SIZE-rounded address, instead of a huge-page-size-rounded address (unless UFFD_FEATURE_EXACT_ADDRESS is used). There is no way to disable the API changes that come with issuing MADV_SPLIT. MADV_COLLAPSE can be used to collapse high-granularity page table mappings that come from the extended functionality that comes with using MADV_SPLIT. For post-copy live migration, the expected use-case is: 1. mmap(MAP_SHARED, some_fd) primary mapping 2. mmap(MAP_SHARED, some_fd) alias mapping 3. MADV_SPLIT the primary mapping 4. UFFDIO_REGISTER/etc. the primary mapping 5. Copy memory contents into alias mapping and UFFDIO_CONTINUE the corresponding PAGE_SIZE sections in the primary mapping. More API changes may be added in the future. Signed-off-by: James Houghton --- arch/alpha/include/uapi/asm/mman.h | 2 ++ arch/mips/include/uapi/asm/mman.h | 2 ++ arch/parisc/include/uapi/asm/mman.h | 2 ++ arch/xtensa/include/uapi/asm/mman.h | 2 ++ include/linux/hugetlb.h | 2 ++ include/uapi/asm-generic/mman-common.h | 2 ++ mm/hugetlb.c | 3 +-- mm/madvise.c | 26 ++++++++++++++++++++++++++ 8 files changed, 39 insertions(+), 2 deletions(-) diff --git a/arch/alpha/include/uapi/asm/mman.h b/arch/alpha/include/uapi/asm/mman.h index 763929e814e9..7a26f3648b90 100644 --- a/arch/alpha/include/uapi/asm/mman.h +++ b/arch/alpha/include/uapi/asm/mman.h @@ -78,6 +78,8 @@ #define MADV_COLLAPSE 25 /* Synchronous hugepage collapse */ +#define MADV_SPLIT 26 /* Enable hugepage high-granularity APIs */ + /* compatibility flags */ #define MAP_FILE 0 diff --git a/arch/mips/include/uapi/asm/mman.h b/arch/mips/include/uapi/asm/mman.h index c6e1fc77c996..f8a74a3a0928 100644 --- a/arch/mips/include/uapi/asm/mman.h +++ b/arch/mips/include/uapi/asm/mman.h @@ -105,6 +105,8 @@ #define MADV_COLLAPSE 25 /* Synchronous hugepage collapse */ +#define MADV_SPLIT 26 /* Enable hugepage high-granularity APIs */ + /* compatibility flags */ #define MAP_FILE 0 diff --git a/arch/parisc/include/uapi/asm/mman.h b/arch/parisc/include/uapi/asm/mman.h index 68c44f99bc93..a6dc6a56c941 100644 --- a/arch/parisc/include/uapi/asm/mman.h +++ b/arch/parisc/include/uapi/asm/mman.h @@ -72,6 +72,8 @@ #define MADV_COLLAPSE 25 /* Synchronous hugepage collapse */ +#define MADV_SPLIT 74 /* Enable hugepage high-granularity APIs */ + #define MADV_HWPOISON 100 /* poison a page for testing */ #define MADV_SOFT_OFFLINE 101 /* soft offline page for testing */ diff --git a/arch/xtensa/include/uapi/asm/mman.h b/arch/xtensa/include/uapi/asm/mman.h index 1ff0c858544f..f98a77c430a9 100644 --- a/arch/xtensa/include/uapi/asm/mman.h +++ b/arch/xtensa/include/uapi/asm/mman.h @@ -113,6 +113,8 @@ #define MADV_COLLAPSE 25 /* Synchronous hugepage collapse */ +#define MADV_SPLIT 26 /* Enable hugepage high-granularity APIs */ + /* compatibility flags */ #define MAP_FILE 0 diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h index 8713d9c4f86c..16fc3e381801 100644 --- a/include/linux/hugetlb.h +++ b/include/linux/hugetlb.h @@ -109,6 +109,8 @@ struct hugetlb_vma_lock { struct vm_area_struct *vma; }; +void hugetlb_vma_lock_alloc(struct vm_area_struct *vma); + extern struct resv_map *resv_map_alloc(void); void resv_map_release(struct kref *ref); diff --git a/include/uapi/asm-generic/mman-common.h b/include/uapi/asm-generic/mman-common.h index 6ce1f1ceb432..996e8ded092f 100644 --- a/include/uapi/asm-generic/mman-common.h +++ b/include/uapi/asm-generic/mman-common.h @@ -79,6 +79,8 @@ #define MADV_COLLAPSE 25 /* Synchronous hugepage collapse */ +#define MADV_SPLIT 26 /* Enable hugepage high-granularity APIs */ + /* compatibility flags */ #define MAP_FILE 0 diff --git a/mm/hugetlb.c b/mm/hugetlb.c index d27fe05d5ef6..5bd53ae8ca4b 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -92,7 +92,6 @@ struct mutex *hugetlb_fault_mutex_table ____cacheline_aligned_in_smp; /* Forward declaration */ static int hugetlb_acct_memory(struct hstate *h, long delta); static void hugetlb_vma_lock_free(struct vm_area_struct *vma); -static void hugetlb_vma_lock_alloc(struct vm_area_struct *vma); static void __hugetlb_vma_unlock_write_free(struct vm_area_struct *vma); static inline bool subpool_is_free(struct hugepage_subpool *spool) @@ -361,7 +360,7 @@ static void hugetlb_vma_lock_free(struct vm_area_struct *vma) } } -static void hugetlb_vma_lock_alloc(struct vm_area_struct *vma) +void hugetlb_vma_lock_alloc(struct vm_area_struct *vma) { struct hugetlb_vma_lock *vma_lock; diff --git a/mm/madvise.c b/mm/madvise.c index 025be3517af1..04ee28992e52 100644 --- a/mm/madvise.c +++ b/mm/madvise.c @@ -1011,6 +1011,24 @@ static long madvise_remove(struct vm_area_struct *vma, return error; } +static int madvise_split(struct vm_area_struct *vma, + unsigned long *new_flags) +{ + if (!is_vm_hugetlb_page(vma) || !hugetlb_hgm_eligible(vma)) + return -EINVAL; + /* + * Attempt to allocate the VMA lock again. If it isn't allocated, + * MADV_COLLAPSE won't work. + */ + hugetlb_vma_lock_alloc(vma); + + /* PMD sharing doesn't work with HGM. */ + hugetlb_unshare_all_pmds(vma); + + *new_flags |= VM_HUGETLB_HGM; + return 0; +} + /* * Apply an madvise behavior to a region of a vma. madvise_update_vma * will handle splitting a vm area into separate areas, each area with its own @@ -1089,6 +1107,11 @@ static int madvise_vma_behavior(struct vm_area_struct *vma, break; case MADV_COLLAPSE: return madvise_collapse(vma, prev, start, end); + case MADV_SPLIT: + error = madvise_split(vma, &new_flags); + if (error) + goto out; + break; } anon_name = anon_vma_name(vma); @@ -1183,6 +1206,9 @@ madvise_behavior_valid(int behavior) case MADV_HUGEPAGE: case MADV_NOHUGEPAGE: case MADV_COLLAPSE: +#endif +#ifdef CONFIG_HUGETLB_HIGH_GRANULARITY_MAPPING + case MADV_SPLIT: #endif case MADV_DONTDUMP: case MADV_DODUMP: -- 2.39.0.314.g84b9a713c41-goog