From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-9.8 required=3.0 tests=BAYES_00,DKIM_SIGNED, DKIM_VALID,HEADER_FROM_DIFFERENT_DOMAINS,INCLUDES_PATCH,MAILING_LIST_MULTI, SIGNED_OFF_BY,SPF_HELO_NONE,SPF_PASS,URIBL_BLOCKED autolearn=ham autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id EADA8C433DF for ; Tue, 25 Aug 2020 02:43:40 +0000 (UTC) Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by mail.kernel.org (Postfix) with ESMTP id A5161205CB for ; Tue, 25 Aug 2020 02:43:40 +0000 (UTC) Authentication-Results: mail.kernel.org; dkim=pass (2048-bit key) header.d=bytedance-com.20150623.gappssmtp.com header.i=@bytedance-com.20150623.gappssmtp.com header.b="Xik35GAU" DMARC-Filter: OpenDMARC Filter v1.3.2 mail.kernel.org A5161205CB Authentication-Results: mail.kernel.org; dmarc=none (p=none dis=none) header.from=bytedance.com Authentication-Results: mail.kernel.org; spf=pass smtp.mailfrom=owner-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix) id 0C6856B00C6; Mon, 24 Aug 2020 22:43:40 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 07881900029; Mon, 24 Aug 2020 22:43:40 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id EA8846B00C8; Mon, 24 Aug 2020 22:43:39 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from forelay.hostedemail.com (smtprelay0195.hostedemail.com [216.40.44.195]) by kanga.kvack.org (Postfix) with ESMTP id D2CBC6B00C6 for ; Mon, 24 Aug 2020 22:43:39 -0400 (EDT) Received: from smtpin21.hostedemail.com (10.5.19.251.rfc1918.com [10.5.19.251]) by forelay02.hostedemail.com (Postfix) with ESMTP id 98922362B for ; Tue, 25 Aug 2020 02:43:39 +0000 (UTC) X-FDA: 77187545358.21.quilt58_01054de27058 Received: from filter.hostedemail.com (10.5.16.251.rfc1918.com [10.5.16.251]) by smtpin21.hostedemail.com (Postfix) with ESMTP id 6B2DD180442C2 for ; Tue, 25 Aug 2020 02:43:39 +0000 (UTC) X-HE-Tag: quilt58_01054de27058 X-Filterd-Recvd-Size: 7338 Received: from mail-pj1-f67.google.com (mail-pj1-f67.google.com [209.85.216.67]) by imf39.hostedemail.com (Postfix) with ESMTP for ; Tue, 25 Aug 2020 02:43:38 +0000 (UTC) Received: by mail-pj1-f67.google.com with SMTP id nv17so451515pjb.3 for ; Mon, 24 Aug 2020 19:43:38 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance-com.20150623.gappssmtp.com; s=20150623; h=mime-version:references:in-reply-to:from:date:message-id:subject:to :cc; bh=PXXvfJjv+PfyYEP8RerkMouFzdtCWZ46cy+YpsBryXM=; b=Xik35GAUruB2YLQx720bmys7UuUJGGrYPGuEWgm4bPqBkTXHUDAkLWkGaggEzipwWv AO56TZIgMYaCIqml9oATDFk1HJx0zki2DTurK1g3tetjO3zxIj9xZX6PUaY0kO1NN7iJ P2ujzWCdfsHgZykDkWcT9h9qBjq+OEiH6URRIRnV857aBeXNpEifOJ9ls7EpWEPq2Uq6 ZiTuauFw9sGVSMnlxJcWm+nu/ERr+0zYUJRmq3KhEc5BcyoXUFjbOsq0rrS8jINzvFjY o+79Cg6qYv2B0CfhkwFXCroS4zXtL3/8PmR60S407BuJI5DwNjh0wKh/BAeGgUNiGOs1 SKbA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-gm-message-state:mime-version:references:in-reply-to:from:date :message-id:subject:to:cc; bh=PXXvfJjv+PfyYEP8RerkMouFzdtCWZ46cy+YpsBryXM=; b=HJ8q8R0xDUXkqvNz+9BdYLj2WSXlsDHVS2YI9syrGSW50PkDqB4YdH4NJc1Dl+4jiu ljYaJPWvIAv4Qu59NlEE1f9BRbtPULrGkPORUJpQDy3zhMEbEIIo5AAE+2x0riB4zYsW PvFVr1Zv6laVqjlp6Udvh/S+N46mTeniJrYsSErkqByVdfpy9YHLnaK0kQD6O9F6xr4/ Ok7ElmjB1kI1hyaUOEOxxRWqqqt+N1BuvsYXEcIxwpLZobVE8AHkwpu/lR/FMdtpkh/N zPt++RcjScjPi/5govjtKhZtHMk7BtPtT8alAfkWkl+IyIRN2OkwItohE4ZycI1i03Pn aS2A== X-Gm-Message-State: AOAM533lHu2NTxvePoyTV9fDpo9818VNoZWC7N/psZLyTW5w/dwo04dz g8Q9gFtRvZnbUTEiSMwf07abu6ymyS4pQ+t0u8RFxw== X-Google-Smtp-Source: ABdhPJxbM2UTrkjNNbuYhIYu/UiCbPfU46l9LiniAtioxQL1QQTA6YdZMibG5etJIacAOqOjPvawN2MFfmH81tyM/6w= X-Received: by 2002:a17:90a:bd0e:: with SMTP id y14mr1828365pjr.13.1598323417348; Mon, 24 Aug 2020 19:43:37 -0700 (PDT) MIME-Version: 1.0 References: <20200822095328.61306-1-songmuchun@bytedance.com> In-Reply-To: <20200822095328.61306-1-songmuchun@bytedance.com> From: Muchun Song Date: Tue, 25 Aug 2020 10:42:58 +0800 Message-ID: Subject: Re: [PATCH] mm/hugetlb: Fix a race between hugetlb sysctl handlers To: mike.kravetz@oracle.com, Andrew Morton Cc: ak@linux.intel.com, Linux Memory Management List , LKML Content-Type: text/plain; charset="UTF-8" X-Rspamd-Queue-Id: 6B2DD180442C2 X-Spamd-Result: default: False [0.00 / 100.00] X-Rspamd-Server: rspam02 X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: Hi Andrew and Mike, On Sat, Aug 22, 2020 at 5:53 PM Muchun Song wrote: > > There is a race between the assignment of `table->data` and write value > to the pointer of `table->data` in the __do_proc_doulongvec_minmax(). > Fix this by duplicating the `table`, and only update the duplicate of > it. And introduce a helper of proc_hugetlb_doulongvec_minmax() to > simplify the code. I am sorry, I didn't expose more details about how the race happened. CPU0: CPU1: proc_sys_write hugetlb_sysctl_handler proc_sys_call_handler hugetlb_sysctl_handler_common hugetlb_sysctl_handler table->data = &tmp; hugetlb_sysctl_handler_common table->data = &tmp; proc_doulongvec_minmax do_proc_doulongvec_minmax sysctl_head_finish __do_proc_doulongvec_minmax i = table->data; *i = val; // corrupt CPU1 stack > > The following oops was seen: > > BUG: kernel NULL pointer dereference, address: 0000000000000000 > #PF: supervisor instruction fetch in kernel mode > #PF: error_code(0x0010) - not-present page > Code: Bad RIP value. Here we can see the "Bad RIP value", so the stack frame is corrupted by others. > ... > Call Trace: > ? set_max_huge_pages+0x3da/0x4f0 > ? alloc_pool_huge_page+0x150/0x150 > ? proc_doulongvec_minmax+0x46/0x60 > ? hugetlb_sysctl_handler_common+0x1c7/0x200 > ? nr_hugepages_store+0x20/0x20 > ? copy_fd_bitmaps+0x170/0x170 > ? hugetlb_sysctl_handler+0x1e/0x20 > ? proc_sys_call_handler+0x2f1/0x300 > ? unregister_sysctl_table+0xb0/0xb0 > ? __fd_install+0x78/0x100 > ? proc_sys_write+0x14/0x20 > ? __vfs_write+0x4d/0x90 > ? vfs_write+0xef/0x240 > ? ksys_write+0xc0/0x160 > ? __ia32_sys_read+0x50/0x50 > ? __close_fd+0x129/0x150 > ? __x64_sys_write+0x43/0x50 > ? do_syscall_64+0x6c/0x200 > ? entry_SYSCALL_64_after_hwframe+0x44/0xa9 > > Fixes: e5ff215941d5 ("hugetlb: multiple hstates for multiple page sizes") > Signed-off-by: Muchun Song > --- > mm/hugetlb.c | 27 +++++++++++++++++++++------ > 1 file changed, 21 insertions(+), 6 deletions(-) > > diff --git a/mm/hugetlb.c b/mm/hugetlb.c > index a301c2d672bf..818d6125af49 100644 > --- a/mm/hugetlb.c > +++ b/mm/hugetlb.c > @@ -3454,6 +3454,23 @@ static unsigned int allowed_mems_nr(struct hstate *h) > } > > #ifdef CONFIG_SYSCTL > +static int proc_hugetlb_doulongvec_minmax(struct ctl_table *table, int write, > + void *buffer, size_t *length, > + loff_t *ppos, unsigned long *out) > +{ > + struct ctl_table dup_table; > + > + /* > + * In order to avoid races with __do_proc_doulongvec_minmax(), we > + * can duplicate the @table and alter the duplicate of it. > + */ > + dup_table = *table; > + dup_table.data = out; > + dup_table.maxlen = sizeof(unsigned long); > + > + return proc_doulongvec_minmax(&dup_table, write, buffer, length, ppos); > +} > + > static int hugetlb_sysctl_handler_common(bool obey_mempolicy, > struct ctl_table *table, int write, > void *buffer, size_t *length, loff_t *ppos) > @@ -3465,9 +3482,8 @@ static int hugetlb_sysctl_handler_common(bool obey_mempolicy, > if (!hugepages_supported()) > return -EOPNOTSUPP; > > - table->data = &tmp; > - table->maxlen = sizeof(unsigned long); > - ret = proc_doulongvec_minmax(table, write, buffer, length, ppos); > + ret = proc_hugetlb_doulongvec_minmax(table, write, buffer, length, ppos, > + &tmp); > if (ret) > goto out; > > @@ -3510,9 +3526,8 @@ int hugetlb_overcommit_handler(struct ctl_table *table, int write, > if (write && hstate_is_gigantic(h)) > return -EINVAL; > > - table->data = &tmp; > - table->maxlen = sizeof(unsigned long); > - ret = proc_doulongvec_minmax(table, write, buffer, length, ppos); > + ret = proc_hugetlb_doulongvec_minmax(table, write, buffer, length, ppos, > + &tmp); > if (ret) > goto out; > > -- > 2.11.0 > -- Yours, Muchun