From: Christian Borntraeger <borntraeger@de.ibm.com>
To: Andrew Morton <akpm@linux-foundation.org>
Cc: linux-mm@kvack.org, linux-s390@vger.kernel.org,
kvm@vger.kernel.org, Janosch Frank <frankja@linux.ibm.com>,
David Hildenbrand <david@redhat.com>,
Cornelia Huck <cohuck@redhat.com>,
linux-kernel@vger.kernel.org,
Martin Schwidefsky <schwidefsky@de.ibm.com>,
Andrea Arcangeli <aarcange@redhat.com>,
Mike Rapoport <rppt@linux.vnet.ibm.com>
Subject: Re: [PATCHi v2] mm: do not drop unused pages when userfaultd is running
Date: Tue, 3 Jul 2018 07:23:12 +0200 [thread overview]
Message-ID: <f40c57df-d8ea-d317-891b-89959ebf6353@de.ibm.com> (raw)
In-Reply-To: <20180702140638.eb3edfaa611ba9fa018f92eb@linux-foundation.org>
On 07/02/2018 11:06 PM, Andrew Morton wrote:
> On Mon, 2 Jul 2018 09:50:49 +0200 Christian Borntraeger <borntraeger@de.ibm.com> wrote:
>
>> KVM guests on s390 can notify the host of unused pages. This can result
>> in pte_unused callbacks to be true for KVM guest memory.
>>
>> If a page is unused (checked with pte_unused) we might drop this page
>> instead of paging it. This can have side-effects on userfaultd, when the
>> page in question was already migrated:
>>
>> The next access of that page will trigger a fault and a user fault
>> instead of faulting in a new and empty zero page. As QEMU does not
>> expect a userfault on an already migrated page this migration will fail.
>>
>> The most straightforward solution is to ignore the pte_unused hint if a
>> userfault context is active for this VMA.
>>
>> ...
>>
>> --- a/mm/rmap.c
>> +++ b/mm/rmap.c
>> @@ -64,6 +64,7 @@
>> #include <linux/backing-dev.h>
>> #include <linux/page_idle.h>
>> #include <linux/memremap.h>
>> +#include <linux/userfaultfd_k.h>
>>
>> #include <asm/tlbflush.h>
>>
>> @@ -1481,7 +1482,7 @@ static bool try_to_unmap_one(struct page *page, struct vm_area_struct *vma,
>> set_pte_at(mm, address, pvmw.pte, pteval);
>> }
>>
>> - } else if (pte_unused(pteval)) {
>> + } else if (pte_unused(pteval) && !userfaultfd_armed(vma)) {
>> /*
>> * The guest indicated that the page content is of no
>> * interest anymore. Simply discard the pte, vmscan
>
> A reader of this code will wonder why we're checking
> userfaultfd_armed(). So the writer of this code should add a comment
> which explains this to them ;) Please.
>
Something like: /*
* The guest indicated that the page content is of no
* interest anymore. Simply discard the pte, vmscan
* will take care of the rest.
* A future reference will then fault in a new zero
* page. When userfaultfd is active, we must not drop
* this page though, as its main user (postcopy
* migration) will not expect userfaults on already
* copied pages.
*/
?
prev parent reply other threads:[~2018-07-03 5:23 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2018-07-02 7:50 Christian Borntraeger
2018-07-02 21:06 ` Andrew Morton
2018-07-03 5:23 ` Christian Borntraeger [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=f40c57df-d8ea-d317-891b-89959ebf6353@de.ibm.com \
--to=borntraeger@de.ibm.com \
--cc=aarcange@redhat.com \
--cc=akpm@linux-foundation.org \
--cc=cohuck@redhat.com \
--cc=david@redhat.com \
--cc=frankja@linux.ibm.com \
--cc=kvm@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-s390@vger.kernel.org \
--cc=rppt@linux.vnet.ibm.com \
--cc=schwidefsky@de.ibm.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox