From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-10.8 required=3.0 tests=BAYES_00,DKIM_SIGNED, DKIM_VALID,DKIM_VALID_AU,HEADER_FROM_DIFFERENT_DOMAINS,MAILING_LIST_MULTI, MENTIONS_GIT_HOSTING,SPF_HELO_NONE,SPF_PASS,URIBL_BLOCKED autolearn=ham autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id 6F183C433DB for ; Fri, 8 Jan 2021 12:04:26 +0000 (UTC) Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by mail.kernel.org (Postfix) with ESMTP id DC3D4238EA for ; Fri, 8 Jan 2021 12:04:25 +0000 (UTC) DMARC-Filter: OpenDMARC Filter v1.3.2 mail.kernel.org DC3D4238EA Authentication-Results: mail.kernel.org; dmarc=fail (p=quarantine dis=none) header.from=suse.com Authentication-Results: mail.kernel.org; spf=pass smtp.mailfrom=owner-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix) id 065D76B038E; Fri, 8 Jan 2021 07:04:25 -0500 (EST) Received: by kanga.kvack.org (Postfix, from userid 40) id F30EF6B038F; Fri, 8 Jan 2021 07:04:24 -0500 (EST) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id E47446B0390; Fri, 8 Jan 2021 07:04:24 -0500 (EST) X-Delivered-To: linux-mm@kvack.org Received: from forelay.hostedemail.com (smtprelay0113.hostedemail.com [216.40.44.113]) by kanga.kvack.org (Postfix) with ESMTP id CDDF06B038E for ; Fri, 8 Jan 2021 07:04:24 -0500 (EST) Received: from smtpin11.hostedemail.com (10.5.19.251.rfc1918.com [10.5.19.251]) by forelay03.hostedemail.com (Postfix) with ESMTP id 9BBF7824556B for ; Fri, 8 Jan 2021 12:04:24 +0000 (UTC) X-FDA: 77682475248.11.spade21_1910433274f2 Received: from filter.hostedemail.com (10.5.16.251.rfc1918.com [10.5.16.251]) by smtpin11.hostedemail.com (Postfix) with ESMTP id 7E437180F8B80 for ; Fri, 8 Jan 2021 12:04:24 +0000 (UTC) X-HE-Tag: spade21_1910433274f2 X-Filterd-Recvd-Size: 5330 Received: from mx2.suse.de (mx2.suse.de [195.135.220.15]) by imf17.hostedemail.com (Postfix) with ESMTP for ; Fri, 8 Jan 2021 12:04:23 +0000 (UTC) X-Virus-Scanned: by amavisd-new at test-mx.suse.de DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=suse.com; s=susede1; t=1610107462; h=from:from:reply-to:date:date:message-id:message-id:to:to:cc:cc: mime-version:mime-version:content-type:content-type: in-reply-to:in-reply-to:references:references; bh=3cxt4elwiduxYGwhS5P7rMAog2VRT9vUyx01asODV6c=; b=Xr0nQSIXNNiSaOOzj3iFonNGtiCnASx4aursLYuELx2gm7kL62QjX9P9yImQx8l3Zw4l9z Q+oN7UBQfhRo4uBuG6SSWyAPKFsVEem5sM2f5ThTe+0CeR3B0jKNgvd0Vd45GRtd9K7rC2 sfsOkjQIIjMNTaKAzx5b6S5ItJ07pg8= Received: from relay2.suse.de (unknown [195.135.221.27]) by mx2.suse.de (Postfix) with ESMTP id A9C5DACAF; Fri, 8 Jan 2021 12:04:22 +0000 (UTC) Date: Fri, 8 Jan 2021 13:04:21 +0100 From: Michal Hocko To: Muchun Song Cc: Mike Kravetz , Andrew Morton , Naoya Horiguchi , Andi Kleen , Linux Memory Management List , LKML Subject: Re: [External] Re: [PATCH v2 3/6] mm: hugetlb: fix a race between freeing and dissolving the page Message-ID: <20210108120421.GC13207@dhcp22.suse.cz> References: <20210107123854.GJ13207@dhcp22.suse.cz> <20210107141130.GL13207@dhcp22.suse.cz> <20210108084330.GW13207@dhcp22.suse.cz> <20210108093136.GY13207@dhcp22.suse.cz> <20210108114411.GZ13207@dhcp22.suse.cz> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: On Fri 08-01-21 19:52:54, Muchun Song wrote: > On Fri, Jan 8, 2021 at 7:44 PM Michal Hocko wrote: > > > > On Fri 08-01-21 18:08:57, Muchun Song wrote: > > > On Fri, Jan 8, 2021 at 5:31 PM Michal Hocko wrote: > > > > > > > > On Fri 08-01-21 17:01:03, Muchun Song wrote: > > > > > On Fri, Jan 8, 2021 at 4:43 PM Michal Hocko wrote: > > > > > > > > > > > > On Thu 07-01-21 23:11:22, Muchun Song wrote: > > > > [..] > > > > > > > But I find a tricky problem to solve. See free_huge_page(). > > > > > > > If we are in non-task context, we should schedule a work > > > > > > > to free the page. We reuse the page->mapping. If the page > > > > > > > is already freed by the dissolve path. We should not touch > > > > > > > the page->mapping. So we need to check PageHuge(). > > > > > > > The check and llist_add() should be protected by > > > > > > > hugetlb_lock. But we cannot do that. Right? If dissolve > > > > > > > happens after it is linked to the list. We also should > > > > > > > remove it from the list (hpage_freelist). It seems to make > > > > > > > the thing more complex. > > > > > > > > > > > > I am not sure I follow you here but yes PageHuge under hugetlb_lock > > > > > > should be the reliable way to check for the race. I am not sure why we > > > > > > really need to care about mapping or other state. > > > > > > > > > > CPU0: CPU1: > > > > > free_huge_page(page) > > > > > if (PageHuge(page)) > > > > > dissolve_free_huge_page(page) > > > > > spin_lock(&hugetlb_lock) > > > > > update_and_free_page(page) > > > > > spin_unlock(&hugetlb_lock) > > > > > llist_add(page->mapping) > > > > > // the mapping is corrupted > > > > > > > > > > The PageHuge(page) and llist_add() should be protected by > > > > > hugetlb_lock. Right? If so, we cannot hold hugetlb_lock > > > > > in free_huge_page() path. > > > > > > > > OK, I see. I completely forgot about this snowflake. I thought that > > > > free_huge_page was a typo missing initial __. Anyway you are right that > > > > this path needs a check as well. But I don't see why we couldn't use the > > > > lock here. The lock can be held only inside the !in_task branch. > > > > > > Because we hold the hugetlb_lock without disable irq. So if an interrupt > > > occurs after we hold the lock. And we also free a HugeTLB page. Then > > > it leads to deadlock. > > > > There is nothing really to prevent making hugetlb_lock irq safe, isn't > > it? > > Yeah. We can make the hugetlb_lock irq safe. But why have we not > done this? Maybe the commit changelog can provide more information. > > See https://github.com/torvalds/linux/commit/c77c0a8ac4c522638a8242fcb9de9496e3cdbb2d Dang! Maybe it is the time to finally stack one workaround on top of the other and put this code into the shape. The amount of hackery and subtle details has just grown beyond healthy! -- Michal Hocko SUSE Labs