From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 20C2AC28B20 for ; Wed, 2 Apr 2025 12:24:51 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 085F0280006; Wed, 2 Apr 2025 08:24:50 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 03463280001; Wed, 2 Apr 2025 08:24:49 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id E183C280006; Wed, 2 Apr 2025 08:24:49 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0011.hostedemail.com [216.40.44.11]) by kanga.kvack.org (Postfix) with ESMTP id C42BF280001 for ; Wed, 2 Apr 2025 08:24:49 -0400 (EDT) Received: from smtpin13.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay03.hostedemail.com (Postfix) with ESMTP id 0177BAD532 for ; Wed, 2 Apr 2025 12:24:49 +0000 (UTC) X-FDA: 83289022740.13.7F6B0DE Received: from mail-wr1-f45.google.com (mail-wr1-f45.google.com [209.85.221.45]) by imf02.hostedemail.com (Postfix) with ESMTP id EEE1A80011 for ; Wed, 2 Apr 2025 12:24:47 +0000 (UTC) Authentication-Results: imf02.hostedemail.com; dkim=pass header.d=suse.com header.s=google header.b=AvY8GB6T; spf=pass (imf02.hostedemail.com: domain of mhocko@suse.com designates 209.85.221.45 as permitted sender) smtp.mailfrom=mhocko@suse.com; dmarc=pass (policy=quarantine) header.from=suse.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1743596688; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=KRFn/kYndNELtCs9bm4l6TqR46CF9PJ02S18fDZOVDM=; b=LII857E/PJu2ppnuHMaDUf+fVsONk0XiGmpiHFxhazZnQsoXZv9lBQ2D4uMt/YfO8J3l6/ /LRSOcFRcVf96uXkGIT7C3GdD6XqQaii0gcqSBDgKmm6IPaQeEWR1guqwWwQ39+4Fitp/D OK/Ct1ELYVNDpfl32wVSrX7g+PUBRxE= ARC-Authentication-Results: i=1; imf02.hostedemail.com; dkim=pass header.d=suse.com header.s=google header.b=AvY8GB6T; spf=pass (imf02.hostedemail.com: domain of mhocko@suse.com designates 209.85.221.45 as permitted sender) smtp.mailfrom=mhocko@suse.com; dmarc=pass (policy=quarantine) header.from=suse.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1743596688; a=rsa-sha256; cv=none; b=OtF9jwj0F7ifeUlcuTxDZpzCO52eCvxzXo9rhx/hszSrnimCpT7x99dYHCHsarQcpD8v08 geuAiwehCHV8NeDRLFVpSkw3eF75hx91fCyetVMX6gvemJ+FhEnoAjfWOlrKEGDJNeLi+r x+jWCrhT6lsmfxjQb3XQ6X+xTDlugCw= Received: by mail-wr1-f45.google.com with SMTP id ffacd0b85a97d-3914a5def6bso3820137f8f.1 for ; Wed, 02 Apr 2025 05:24:47 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=suse.com; s=google; t=1743596686; x=1744201486; darn=kvack.org; h=in-reply-to:content-transfer-encoding:content-disposition :mime-version:references:message-id:subject:cc:to:from:date:from:to :cc:subject:date:message-id:reply-to; bh=KRFn/kYndNELtCs9bm4l6TqR46CF9PJ02S18fDZOVDM=; b=AvY8GB6TvdYY3EJIbahOC94xA4S0gsf4ohcrzV6XUuwxheinF5hu7HQvkBwbbqwrG9 0hBa6QsFQBoq9aTojsPg/A4a59WARmltq0DkZTTYZk3/mRqLGvyyAU/cc0iITiPKDn5k O3G5EMyxX43zcRcYqG8kiIudtard4rr/GRi1ABQyD0VOp4dX0EqXhGLswoC6dVY2BKFI rPUrJykw559S+tPZU9FcbVC7N0Q/s13FaiFjb86Bv0rdJZ9A34QRolihxfOTcgJ3QUap 0GU8AjhyVVoHhkjE7UfPzvasGtg5mbWm5mp1VB15sxpcFceQazO3Fhh0BjiKzi4U9oIO cnCw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1743596686; x=1744201486; h=in-reply-to:content-transfer-encoding:content-disposition :mime-version:references:message-id:subject:cc:to:from:date :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=KRFn/kYndNELtCs9bm4l6TqR46CF9PJ02S18fDZOVDM=; b=egwyfyEuEhNL2CTgr10c+T8JdfpriSsP1WmQeLWY6u93/ufdZmXmecxvOq0aDow2b3 0nz1z66sjtY8uOLlKFUah8GpQOAa1bvtgHA7oIIQQ1EEpOeqdpplz+OTRpqHyZEIoeU6 7UQJYYEmgh7SquKPTk6ZXMb6RpUYG3XOf7ejORhLO5GaTXRss3XGDiHR0oQOhnjhOmol fPU0pa94FtDayGQ7KTVdPSY0K9L3sClbsa1MYt1mP7kRe+P1fweZApHJ/F9RM3rF+C81 sECOPkNgxHkOkPIFDSdOkKAPx3Ug8jpEbyzQrvckiFzfBrTQNYcHmcn+FJYMRCYNTMSO ZTGg== X-Forwarded-Encrypted: i=1; AJvYcCUy4lQY7LQfYUr4+ftvFJHfV2UNaUCKkJsacPmZ0SGQHAJ9xXbuGonPkWCymA/FsJaAzovdwmwMCg==@kvack.org X-Gm-Message-State: AOJu0Yy7lh2Axbe7aVapZYBtxoxFB4t6fn76z0JdXHif/bm6N8QMnvcY aKFpU0/zSSz9CVAu7IFCq2oqih2hWcAgUj/B0wrVcsZM63K3cYDAoshVxjY9xjc= X-Gm-Gg: ASbGnctT53zUZDQYSzvrxZ7C/5qbFVsDwR9J2QjA6jEA/MqygaKQj8T3tKboDuAR/4e QHfMcvJvi/YYP0o+vTG8iD2prKvzDgGUsHHJOfi3FRz1Fa/i9fIVBTV3MMXVnzEhwcs3tl7P1kR pRniXhHTXHlflf1z6cVlZbyNB3KPRvlhsk81NhbPhPzy90lJU/5xIfVsoUHNsLMhMGw7gocKY3X NX5oyn2V8rVbmwI0SV5/wI/jPcew+v/gwCoBH0QyI7/R0YimNGTnDM/5xiZj48Rv0Bz88sXSkPh aAOusTqpnNh+ospEacMIQES4Gz74nAPHhuQtCnIJBTqK1FmBaWC33FQs8Yfiogq6baZc X-Google-Smtp-Source: AGHT+IH94G8NMZggEmYTqBe2slfssz4AV4QqXUWOJ4XU66EOBQZvQTodDOoaZK39wpDvzUvmqsjg2Q== X-Received: by 2002:a5d:59a2:0:b0:399:71d4:a2 with SMTP id ffacd0b85a97d-39c120de28cmr14089922f8f.14.1743596686324; Wed, 02 Apr 2025 05:24:46 -0700 (PDT) Received: from localhost (109-81-92-185.rct.o2.cz. [109.81.92.185]) by smtp.gmail.com with UTF8SMTPSA id 5b1f17b1804b1-43eb6135dc4sm19179155e9.33.2025.04.02.05.24.45 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 02 Apr 2025 05:24:46 -0700 (PDT) Date: Wed, 2 Apr 2025 14:24:45 +0200 From: Michal Hocko To: Dave Chinner Cc: Yafang Shao , Harry Yoo , Kees Cook , joel.granados@kernel.org, linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org, Josef Bacik , linux-mm@kvack.org, Vlastimil Babka Subject: Re: [PATCH] proc: Avoid costly high-order page allocations when reading proc files Message-ID: References: <20250401073046.51121-1-laoar.shao@gmail.com> <3315D21B-0772-4312-BCFB-402F408B0EF6@kernel.org> MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: X-Rspamd-Queue-Id: EEE1A80011 X-Stat-Signature: kqnohgnu5qrqsaxc97uw43riahfsnjyh X-Rspam-User: X-Rspamd-Server: rspam12 X-HE-Tag: 1743596687-324945 X-HE-Meta: U2FsdGVkX18/JePQknNhMZ4S+BG7MtAZweHcALONG3oCT5YSbQzyw9XMIUdUwqF8V6j7aKZuXeJPGWkfP0Qqj0suMynnY17ZX/N3MGKjf93bULP8TOmrBPIBSGZS1oD0YLmhnSocWPOTDL/lPjceuPxBK/UDA7ZTRttHR8Jhjdr3mrxKMhpwBB1BDGpngwq5Hqw2AEjrfQmyQT3Kay6mndIfuN1d82CGf6WvmeiJNoyYOZ/TbvJAAl2o3vH2EOmXhYPdZOEQwd/DqHHAu26kxdZLOA7r6XHj0NTSZQgTTqjS/bpDwRdNAaIr+FRqtH/Kjn0rC+j4jNLLWze3yzA/E7VZBvyl/E989Xu0lZ2UoLOOciL4gaiP86DMLpEfeXO95kZQbCHvSlKJyhpOtbkV2xKzPG1/McVtvWSpKlWn3YChmDyDRjjSyGH6/+RbjOzNAdlQaKyWXXaBu4Nqhq+PUIBkmgEhkVmlU+8hPLNVmfVV3GdMn9diqI2zliWnE9KPXpzoafW/Y15wIe55AMFB/w2D6KA3e3YFvHLVxDK0hfDv88MdO1VwSsho8dJXO4MmT+Fdx1WeukCZcFLD600xM6LeQKTz5Yi7YN5M3CtXFOlvEJTUD7zNdU87sgd6pes37/NFFnIn0c3scAE8cJKQwVamHru/3AmcFCT2chrAf1WQRHeI5PjNz2iApbkNU3B7DjB/gvosn0K6KPEwxtMSqzOpOHT9ykjIK/5ToFFx39vrdoTN0AL+ghe+ypiQLkSeAe5QrVXuuioxhAFv4Z0INphdjmo9j+UZc9wsw1w48vWC/coV1i4TlVoyo7clY35gwC+WnfrGCV/9lMiQZCiefTtjlKrYvRrupr+YLHk6AqGTuRAeLkfTY7UEPoWImUi42TJ0D1GHVb4cobaPcKcLI9mN5KpL0Ht3mkPelMVHo7zFsX/r2/kU7fidJjYbp5JElBuRpApkhSdtHhZDNum swW2U8dl vcQqmLPeaNbUB4MKq9JAPxP3iJoBCAhEO2lNRFAEVr6//7ati8yENtw3bU5Co6n6jLbiCwyyUeg2YtCWnlKBuppWqAGpcUbUsUYZQzowejeNz+iCRUDzokWELtsngbt59kUpLYbaRwz0B1zwlf0lgzZB73FwBE1RzmNASZMKw80cWMkbytAKlOIpmwvRoDnAVcQchTrJvwexu2B9atYijpRDiuecyhowbAF3rXfx+b0P485xZFrmzuOzfmgUpVGS9FebmcWSBIgJk2orgzVfxvAOiVq4q6FnKqNlDAMVcAN4PwcBb1g7IVHjeYCLLA2GvsZK5nBxaPmVRj7Kl0r8FwUFGrQVDnfRmZT2SHt1Qvs1gGHEYAIaoCtfh0Jyj05kwEtuUwRnF+EueFgDew49TQYBYDZ1u1MYCAKPVCwB65lBH7tQYZtaTXJTmmeD7HfRJNk8+hi83/iCEUvaWyAFQllvHrjtXqcct0OT8oM9M6+MXi1XXQV2JTs3QxgNIxEWEcTzugd606GrjIKQ= X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: On Wed 02-04-25 22:32:14, Dave Chinner wrote: > On Wed, Apr 02, 2025 at 04:42:06PM +0800, Yafang Shao wrote: > > On Wed, Apr 2, 2025 at 12:15 PM Harry Yoo wrote: > > > > > > On Tue, Apr 01, 2025 at 07:01:04AM -0700, Kees Cook wrote: > > > > > > > > > > > > On April 1, 2025 12:30:46 AM PDT, Yafang Shao wrote: > > > > >While investigating a kcompactd 100% CPU utilization issue in production, I > > > > >observed frequent costly high-order (order-6) page allocations triggered by > > > > >proc file reads from monitoring tools. This can be reproduced with a simple > > > > >test case: > > > > > > > > > > fd = open(PROC_FILE, O_RDONLY); > > > > > size = read(fd, buff, 256KB); > > > > > close(fd); > > > > > > > > > >Although we should modify the monitoring tools to use smaller buffer sizes, > > > > >we should also enhance the kernel to prevent these expensive high-order > > > > >allocations. > > > > > > > > > >Signed-off-by: Yafang Shao > > > > >Cc: Josef Bacik > > > > >--- > > > > > fs/proc/proc_sysctl.c | 10 +++++++++- > > > > > 1 file changed, 9 insertions(+), 1 deletion(-) > > > > > > > > > >diff --git a/fs/proc/proc_sysctl.c b/fs/proc/proc_sysctl.c > > > > >index cc9d74a06ff0..c53ba733bda5 100644 > > > > >--- a/fs/proc/proc_sysctl.c > > > > >+++ b/fs/proc/proc_sysctl.c > > > > >@@ -581,7 +581,15 @@ static ssize_t proc_sys_call_handler(struct kiocb *iocb, struct iov_iter *iter, > > > > > error = -ENOMEM; > > > > > if (count >= KMALLOC_MAX_SIZE) > > > > > goto out; > > > > >- kbuf = kvzalloc(count + 1, GFP_KERNEL); > > > > >+ > > > > >+ /* > > > > >+ * Use vmalloc if the count is too large to avoid costly high-order page > > > > >+ * allocations. > > > > >+ */ > > > > >+ if (count < (PAGE_SIZE << PAGE_ALLOC_COSTLY_ORDER)) > > > > >+ kbuf = kvzalloc(count + 1, GFP_KERNEL); > > > > > > > > Why not move this check into kvmalloc family? > > > > > > Hmm should this check really be in kvmalloc family? > > > > Modifying the existing kvmalloc functions risks performance regressions. > > Could we instead introduce a new variant like vkmalloc() (favoring > > vmalloc over kmalloc) or kvmalloc_costless()? > > We should fix kvmalloc() instead of continuing to force > subsystems to work around the limitations of kvmalloc(). Agreed! > Have a look at xlog_kvmalloc() in XFS. It implements a basic > fast-fail, no retry high order kmalloc before it falls back to > vmalloc by turning off direct reclaim for the kmalloc() call. > Hence if the there isn't a high-order page on the free lists ready > to allocate, it falls back to vmalloc() immediately. > > For XFS, using xlog_kvmalloc() reduced the high-order per-allocation > overhead by around 80% when compared to a standard kvmalloc() > call. Numbers and profiles were documented in the commit message > (reproduced in whole below)... Btw. it would be really great to have such concerns to be posted to the linux-mm ML so that we are aware of that. kvmalloc currently doesn't support GFP_NOWAIT semantic but it does allow to express - I prefer SLAB allocator over vmalloc. I think we could make the default kvmalloc slab path weaker by default as those who really want slab already have means to achieve that. There is a risk of long term fragmentation but I think this is worth trying diff --git a/mm/util.c b/mm/util.c index 60aa40f612b8..8386f6976d7d 100644 --- a/mm/util.c +++ b/mm/util.c @@ -601,14 +601,18 @@ static gfp_t kmalloc_gfp_adjust(gfp_t flags, size_t size) * We want to attempt a large physically contiguous block first because * it is less likely to fragment multiple larger blocks and therefore * contribute to a long term fragmentation less than vmalloc fallback. - * However make sure that larger requests are not too disruptive - no - * OOM killer and no allocation failure warnings as we have a fallback. + * However make sure that larger requests are not too disruptive - i.e. + * do not direct reclaim unless physically continuous memory is preferred + * (__GFP_RETRY_MAYFAIL mode). We still kick in kswapd/kcompactd to start + * working in the background but the allocation itself. */ if (size > PAGE_SIZE) { flags |= __GFP_NOWARN; if (!(flags & __GFP_RETRY_MAYFAIL)) flags |= __GFP_NORETRY; + else + flags &= ~__GFP_DIRECT_RECLAIM; /* nofail semantic is implemented by the vmalloc fallback */ flags &= ~__GFP_NOFAIL; -- Michal Hocko SUSE Labs