From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 94ECFC6FD35 for ; Thu, 29 Aug 2024 08:31:19 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 2E9F86B00BE; Thu, 29 Aug 2024 04:31:19 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 29A076B00C0; Thu, 29 Aug 2024 04:31:19 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 162606B00C2; Thu, 29 Aug 2024 04:31:19 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0010.hostedemail.com [216.40.44.10]) by kanga.kvack.org (Postfix) with ESMTP id ED72A6B00BE for ; Thu, 29 Aug 2024 04:31:18 -0400 (EDT) Received: from smtpin01.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay01.hostedemail.com (Postfix) with ESMTP id A881B1C52EC for ; Thu, 29 Aug 2024 08:31:18 +0000 (UTC) X-FDA: 82504613436.01.0C303EE Received: from mail-yb1-f170.google.com (mail-yb1-f170.google.com [209.85.219.170]) by imf08.hostedemail.com (Postfix) with ESMTP id 1D809160020 for ; Thu, 29 Aug 2024 08:31:16 +0000 (UTC) Authentication-Results: imf08.hostedemail.com; dkim=pass header.d=gmail.com header.s=20230601 header.b=NB3w3iNN; dmarc=pass (policy=none) header.from=gmail.com; spf=pass (imf08.hostedemail.com: domain of jingxiangzeng.cas@gmail.com designates 209.85.219.170 as permitted sender) smtp.mailfrom=jingxiangzeng.cas@gmail.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1724920178; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-transfer-encoding:in-reply-to: references:dkim-signature; bh=XP/0pqR7fTz/XMxG/X4ZE62yRuJu0s/9oxHpT/dTS3A=; b=FilbiACR7OZuoDfJSfiC/1sOpzEr09FW0KDQuQysn7XA3MuZWfGsJaubGcMRxJhZk7nDJ6 JiqxvS+NB+6tAe8N7MpbaSy04xWnO2IPXVcxVLBK0Mc1nEcT4u6Q8AXGTkY1SAw1z06k40 jKvo4o9aIgR+SNRiKegSx61MdVVjxu4= ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1724920178; a=rsa-sha256; cv=none; b=PTKHc/hoa7KEMBju020HP/R8sLe4/MpXUirS4bj6J4eGQlkvDQq1GmzBUrypSn7SiETfN6 q0Ob/bTnRHk4d9xn5+KhcWZHtHcXwL/fJb2e/6Xw9GQpXys+NF80bpdLi++/FrWsN9eEaK 2uvUZTIDwx8QMmxByjxettkTIdBrgls= ARC-Authentication-Results: i=1; imf08.hostedemail.com; dkim=pass header.d=gmail.com header.s=20230601 header.b=NB3w3iNN; dmarc=pass (policy=none) header.from=gmail.com; spf=pass (imf08.hostedemail.com: domain of jingxiangzeng.cas@gmail.com designates 209.85.219.170 as permitted sender) smtp.mailfrom=jingxiangzeng.cas@gmail.com Received: by mail-yb1-f170.google.com with SMTP id 3f1490d57ef6-e116d2f5f7fso1204194276.1 for ; Thu, 29 Aug 2024 01:31:16 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1724920276; x=1725525076; darn=kvack.org; h=cc:to:subject:message-id:date:from:mime-version:from:to:cc:subject :date:message-id:reply-to; bh=XP/0pqR7fTz/XMxG/X4ZE62yRuJu0s/9oxHpT/dTS3A=; b=NB3w3iNNVF6zCSijVb4u4/OgW50dScDvd7u5rGki0bgg1LNkmqbFWBcooUGlX3358g 32AfM9LYtxRf/QGzrfhaoATr/Ue2aDsziys3vI4oeVtYcqCMbmYVnPTaEAt6QGxg9gvx D7x/0cj+nw64Q2jWQ6Z5BCZQ6eUBx8z0d7h0UkqMM/nSWhNEksvjKr2/N2kkoj47+VWk NhVhfC4j9rqyUD1cPqbVlpvBy3s9xCyGALPHR1IZ7ez2yxv5MUR2hoTu8JzPRcdZFyny hvIANLTVcTrnpPOoyhHArlcShYMH/5iDzAYwRW2w4ssiORbJENqk3uB7zjEZgbw4XHgF D7wg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1724920276; x=1725525076; h=cc:to:subject:message-id:date:from:mime-version:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=XP/0pqR7fTz/XMxG/X4ZE62yRuJu0s/9oxHpT/dTS3A=; b=ZsGe9DjWZnrc3lrkXKHoQnMBcgP9DdAVPAw/1uFPGvom95ZynIFzL18rzSc7/zpQ9o /m2eyVr2AIvCJ34Pac/+HvxU7ECuefHeLtnLXrOShd4AELWvtqgvVF9PUetnR/Cc8jd6 JxKeMb7c/yZIIXn9fqQYNzg3VheEch0xNNXgoVFKL24qqdUXqVATBVSkqrm+BuGRJtl5 UoR/ydBImt93JLVUCCOu6ObINX2a6h7vqnGvDBJPJMD/akcVqfpoc9sQwWW29ln/cTBS OdLMxwTsCkO45eOnCTiwjYCsv3n62aEHVyh1IRYb+JbDKwDKmXBbUoXt6P6NpWxCvOcQ cWew== X-Gm-Message-State: AOJu0Yxdf7dTQaxmxlafZZdr1RfGmzYpf3BFX4C6fF1ZV6AbsfSNATN4 cdJa4xoQQ4NjAg2LE5Gd7yfV/eIDVxKTYz30eL9tyKmRu1JLxeRqgIBUiJsIJqubL1Te9Zr51bV CT9KjJb7KigTqZyJa1d7GsX3c+A+DIlnK X-Google-Smtp-Source: AGHT+IFsipChTxpOmT6QECK99xmwbADi0euatcOpmvtwhMzDsBFiTTH6IDVzNcc9s1+zbsRgQR1XSJ98F2SacSDdvCI= X-Received: by 2002:a05:6902:1244:b0:e0e:7fb3:cf89 with SMTP id 3f1490d57ef6-e1a5c665c67mr1386023276.13.1724920275920; Thu, 29 Aug 2024 01:31:15 -0700 (PDT) MIME-Version: 1.0 From: jingxiang zeng Date: Thu, 29 Aug 2024 16:31:05 +0800 Message-ID: Subject: [PATCH] mm/mglru: wake up flushers when legacy cgroups run out of clean caches To: Andrew Morton Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, Yu Zhao , Alexander Motin , Kairui Song Content-Type: multipart/alternative; boundary="0000000000004ff0d40620ce4c96" X-Rspamd-Server: rspam07 X-Rspamd-Queue-Id: 1D809160020 X-Stat-Signature: skofjwr88mqziqn5eqzepq5opdz5n3et X-Rspam-User: X-HE-Tag: 1724920276-120616 X-HE-Meta: U2FsdGVkX1+04um4djP+k7eAOrIS5sW9VzuanSC4OVrSdxtCX64ra92TwEZ7vx/Ox0mLd81aVgFOQDcVhz9wCJnQ6eeQokBcivitoC/lztajozHtjUqCM6eJdc8pqL7S900mhEIc1NAsoMZecHEmcIAWOgcSJb7kzU6P5ZmJahVNpAplcft01ivySrhu15QyhHBkfhO4XbAg9Q6qzbK2Q42QWgvfc1RpQFgbenMZLUWowHrCvNejHHNPHfq/VuPOt+Rgi35pmzbryh0TB8hkCseDj2kr0NJmRW46h2ywukrjLLIkVsOHL6pDG3GEmXJcHdk9xgAvHbDJLiyiWiwHHwYCeXMbmnSEpsqNSIxNyDzV7VVdt4zeocICKGrIWWb0j0SvN0r/OaPC/IG1fyIIAXUSQGsTHbk+oEK2DW+Q0OgNLO7hhirGEd/V8BuvfFi/yjbZolRl2dJTYimoPrFZ6SVi6DEprQzHNkvFQsN6PUefqmiwqjc7ZgEXZxKxaqCuQzR9xrDxGT59p7uw8vrl44/FbMUAZ9v0dHX05J9uOKoqJvfdfth8q7fW2SRV0G/1WZPNkDC0Qzief4BvFV+9kaslLs468SV0Tz+IJb1/qmgAy3uTcgZHoe9ZBNC4pQ7CtDY3759Jpi+cRGeHvRELlYNOdWfq222sqc1R3D9NP94DnfeUutbibldQl+1oliuiYxrsXhgzGzhIjxSEhEqqqz7V+CvF2thyYkBg4zf6y56eSFzFdrXecrx77lmz+1uDvkg6y0MTKP5kjTHxWtmXSwijJAns0craWhHL2qNXNVagXpKHqgIjv3zKYSX0EyTrpZg4Gz4RhTmdvYu0V+dAJED7wOw4+YLN2vE1mkM/hBfslWWHo3bShPOTF8u9/SXQ7rYE/hkFX1U7i8gnVibPdwL+DAHS8jgPVukZKKv3kSt2yuR8tT0gUUXq8hKzaempnz/wyjwTzfAAYIy83px 7ftuVCUk HVjktDCYfZNPnJ520eVNa3XJnHlq9OJK2YbUBRN+aVoxR3yfnEQCPH4mJwFdbgpomAQw5It2mPR/rBnRN3p5D/2SHUIrRGBuajISvlVBlkF5FFF8kBDpwmhENoRUN97kH4cqGVeGVG52oEFh5/ndYEKShRhD+icK0YPgWOrAh8Um4sbYf/5Mmj1URrEstkPWIxkS2g8x50Uz+zYww3S0KlURtUOt5UqA9X2QRzk25rpHAMR5VNgph2n1UbyRikxwIYSRcVYEIJ5lrQwHub9VaRQz1ydzJl+WDQ1iSZD3WiA9YGRJEe+wPx1A9kiPvXaUNn2ZngmMnUGB5oXgdcFGlta7qy84Jskfe7YkOiUBCl90aFWmL1wplV0WH9mqlDzrzYr8DIUX8kF+B7sVymLjXU6Bk4jLE61HFn+tIhC0iZ35rb2cqDB+q9HpV3Tn+Xn0oOlhiSqRuo3CGh6TSeUcZscKRBsh6nFXzzQ9iMSB32KVmDvDG9SjbUflP7WdUssh8Jj4I X-Bogosity: Ham, tests=bogofilter, spamicity=0.006510, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: --0000000000004ff0d40620ce4c96 Content-Type: text/plain; charset="UTF-8" Commit 14aa8b2d5c2e ("mm/mglru: don't sync disk for each aging cycle") removed the opportunity to wake up flushers during the MGLRU page reclamation process can lead to an increased likelihood of triggering OOM when encountering many dirty pages during reclamation on MGLRU. This leads to premature OOM if there are too many dirty pages in cgroup: Killed dd invoked oom-killer: gfp_mask=0x101cca(GFP_HIGHUSER_MOVABLE|__GFP_WRITE), order=0, oom_score_adj=0 Call Trace: dump_stack_lvl+0x5f/0x80 dump_stack+0x14/0x20 dump_header+0x46/0x1b0 oom_kill_process+0x104/0x220 out_of_memory+0x112/0x5a0 mem_cgroup_out_of_memory+0x13b/0x150 try_charge_memcg+0x44f/0x5c0 charge_memcg+0x34/0x50 __mem_cgroup_charge+0x31/0x90 filemap_add_folio+0x4b/0xf0 __filemap_get_folio+0x1a4/0x5b0 ? srso_return_thunk+0x5/0x5f ? __block_commit_write+0x82/0xb0 ext4_da_write_begin+0xe5/0x270 generic_perform_write+0x134/0x2b0 ext4_buffered_write_iter+0x57/0xd0 ext4_file_write_iter+0x76/0x7d0 ? selinux_file_permission+0x119/0x150 ? srso_return_thunk+0x5/0x5f ? srso_return_thunk+0x5/0x5f vfs_write+0x30c/0x440 ksys_write+0x65/0xe0 __x64_sys_write+0x1e/0x30 x64_sys_call+0x11c2/0x1d50 do_syscall_64+0x47/0x110 entry_SYSCALL_64_after_hwframe+0x76/0x7e memory: usage 308224kB, limit 308224kB, failcnt 2589 swap: usage 0kB, limit 9007199254740988kB, failcnt 0 ... file_dirty 303247360 file_writeback 0 ... oom-kill:constraint=CONSTRAINT_MEMCG,nodemask=(null),cpuset=test, mems_allowed=0,oom_memcg=/test,task_memcg=/test,task=dd,pid=4404,uid=0 Memory cgroup out of memory: Killed process 4404 (dd) total-vm:10512kB, anon-rss:1152kB, file-rss:1824kB, shmem-rss:0kB, UID:0 pgtables:76kB oom_score_adj:0 Wake up flushers when legacy cgroups run out of clean caches. Fixes: 14aa8b2d5c2e ("mm/mglru: don't sync disk for each aging cycle") Signed-off-by: Zeng Jingxiang Signed-off-by: kasong --0000000000004ff0d40620ce4c96 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable
Commit 14aa8b2d5c2e ("mm/mglru: don't sync disk f= or each aging cycle")
removed the opportunity to wake up flushers d= uring the MGLRU page
reclamation process can lead to an increased likeli= hood of triggering
OOM when encountering many dirty pages during reclama= tion on MGLRU.

This leads to premature OOM if there are too many dir= ty pages in cgroup:
Killed

dd invoked oom-killer: gfp_mask=3D0x10= 1cca(GFP_HIGHUSER_MOVABLE|__GFP_WRITE),
order=3D0, oom_score_adj=3D0
=
Call Trace:
=C2=A0 <TASK>
=C2=A0 dump_stack_lvl+0x5f/0x80=C2=A0 dump_stack+0x14/0x20
=C2=A0 dump_header+0x46/0x1b0
=C2=A0 oo= m_kill_process+0x104/0x220
=C2=A0 out_of_memory+0x112/0x5a0
=C2=A0 me= m_cgroup_out_of_memory+0x13b/0x150
=C2=A0 try_charge_memcg+0x44f/0x5c0=C2=A0 charge_memcg+0x34/0x50
=C2=A0 __mem_cgroup_charge+0x31/0x90
= =C2=A0 filemap_add_folio+0x4b/0xf0
=C2=A0 __filemap_get_folio+0x1a4/0x5b= 0
=C2=A0 ? srso_return_thunk+0x5/0x5f
=C2=A0 ? __block_commit_write+0= x82/0xb0
=C2=A0 ext4_da_write_begin+0xe5/0x270
=C2=A0 generic_perform= _write+0x134/0x2b0
=C2=A0 ext4_buffered_write_iter+0x57/0xd0
=C2=A0 e= xt4_file_write_iter+0x76/0x7d0
=C2=A0 ? selinux_file_permission+0x119/0x= 150
=C2=A0 ? srso_return_thunk+0x5/0x5f
=C2=A0 ? srso_return_thunk+0x= 5/0x5f
=C2=A0 vfs_write+0x30c/0x440
=C2=A0 ksys_write+0x65/0xe0
= =C2=A0 __x64_sys_write+0x1e/0x30
=C2=A0 x64_sys_call+0x11c2/0x1d50
= =C2=A0 do_syscall_64+0x47/0x110
=C2=A0 entry_SYSCALL_64_after_hwframe+0x= 76/0x7e

=C2=A0memory: usage 308224kB, limit 308224kB, failcnt 2589=C2=A0swap: usage 0kB, limit 9007199254740988kB, failcnt 0

=C2=A0 = ...
=C2=A0 file_dirty 303247360
=C2=A0 file_writeback 0
=C2=A0 ...=

oom-kill:constraint=3DCONSTRAINT_MEMCG,nodemask=3D(null),cpuset=3Dt= est,
mems_allowed=3D0,oom_memcg=3D/test,task_memcg=3D/test,task=3Ddd,pid= =3D4404,uid=3D0
Memory cgroup out of memory: Killed process 4404 (dd) to= tal-vm:10512kB,
anon-rss:1152kB, file-rss:1824kB, shmem-rss:0kB, UID:0 p= gtables:76kB
oom_score_adj:0

Wake up flushers when legacy cgroups= run out of clean caches.

Fixes: 14aa8b2d5c2e ("mm/mglru: don&#= 39;t sync disk for each aging cycle")
Signed-off-by: Zeng Jingxiang= <linuszeng@tencent.com>=
Signed-off-by: kasong <kasong@= tencent.com>
--0000000000004ff0d40620ce4c96--