From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 932EBD68B33 for ; Thu, 14 Nov 2024 16:19:02 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id DFCC66B009A; Thu, 14 Nov 2024 11:19:01 -0500 (EST) Received: by kanga.kvack.org (Postfix, from userid 40) id DAD9C6B00AB; Thu, 14 Nov 2024 11:19:01 -0500 (EST) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id C26EC6B00AC; Thu, 14 Nov 2024 11:19:01 -0500 (EST) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0017.hostedemail.com [216.40.44.17]) by kanga.kvack.org (Postfix) with ESMTP id A28A96B009A for ; Thu, 14 Nov 2024 11:19:01 -0500 (EST) Received: from smtpin04.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay01.hostedemail.com (Postfix) with ESMTP id 5766B1C7060 for ; Thu, 14 Nov 2024 16:19:01 +0000 (UTC) X-FDA: 82785208800.04.B902150 Received: from mail-qt1-f176.google.com (mail-qt1-f176.google.com [209.85.160.176]) by imf13.hostedemail.com (Postfix) with ESMTP id D1C0920006 for ; Thu, 14 Nov 2024 16:18:13 +0000 (UTC) Authentication-Results: imf13.hostedemail.com; dkim=pass header.d=google.com header.s=20230601 header.b=hVlgpzt9; dmarc=pass (policy=reject) header.from=google.com; spf=pass (imf13.hostedemail.com: domain of surenb@google.com designates 209.85.160.176 as permitted sender) smtp.mailfrom=surenb@google.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1731601008; a=rsa-sha256; cv=none; b=vCv8yGA8vnqytGqoVECOk5HoJmAIfUPguRzX54xvpCcJqKVz/4L95G4XDWSPFXX+aczl85 rP8fsaYH6Fxc+oFzk8LrnbqbTlh61etSyjRmvOCv/eS52Mrc487kEFZ783B9vsvmRHFuIp 8TzwyYNX67a0Ble1sZIVKHT/0viizVQ= ARC-Authentication-Results: i=1; imf13.hostedemail.com; dkim=pass header.d=google.com header.s=20230601 header.b=hVlgpzt9; dmarc=pass (policy=reject) header.from=google.com; spf=pass (imf13.hostedemail.com: domain of surenb@google.com designates 209.85.160.176 as permitted sender) smtp.mailfrom=surenb@google.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1731601008; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=qKBFPyDpG0DZLpYX58E37SKTdIlh5Yx8PAsqBjBLq3c=; b=HSCwyHakmD05EgEshVUuku3e4GNJn1WFGREwgHndujwY6tUxYsNltQ/FStLTzXMMsymAER 3KGQh+TOAvwm75foIN40BIB03ld9IWI7XkT1xWrjsUCu0YqL16etO4H+CTtDbMOCFAdKza JV93g258t1zwMUa7dgnfBdbPRna93OQ= Received: by mail-qt1-f176.google.com with SMTP id d75a77b69052e-4608dddaa35so338521cf.0 for ; Thu, 14 Nov 2024 08:18:59 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20230601; t=1731601138; x=1732205938; darn=kvack.org; h=content-transfer-encoding:to:subject:message-id:date:from :in-reply-to:references:mime-version:from:to:cc:subject:date :message-id:reply-to; bh=qKBFPyDpG0DZLpYX58E37SKTdIlh5Yx8PAsqBjBLq3c=; b=hVlgpzt9pgMXB0v52i8BjygPwihz4kwexU+zt/JS1p72IKG6ZRddYxsVKdtxUbP9Ms e6WOJMIZ1e/AgZSeHFD0QdNp9k/4/4qiOwg45B+OC+5gvkmM3+xkfLy2nuQKUXLM/DJt HAyzv3b1AeBkXCxWnW+pkKTFL0kgHySM52wOLjIQGNuzOijIUCvK0ukcilJpZ0xdXW+n gGACHadba2vEI0JclBrd8Od67QFf/sLTCGt5mOOBVEDyDlvBdemqqvGl48iaNjkMIcU/ pbw1FBrhoAX+9UWDmTvg2wxJAeEYNblYYFCjfQkRiY+mERKAmOGmnzaDPgzqGbNaUhQi 0lUA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1731601138; x=1732205938; h=content-transfer-encoding:to:subject:message-id:date:from :in-reply-to:references:mime-version:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=qKBFPyDpG0DZLpYX58E37SKTdIlh5Yx8PAsqBjBLq3c=; b=pmEoUDpiDR9jXDq26zJ9QOT8mM3X91UCcbXr2sjMY6osTgZaB3JV94flkd/6SKtwq9 hHxn6pZPHTjRlxe0ZPPzq8lz5YjHlnS5ysDhoSXt4RcuzK1u9nekOVhgP4SAdyyK4kPz 6DWdxr8Q05AFPQ5htf2axdMr/okTukwFMCjRSa73hEFARcbn57kEsX0SLPPPn8UuJsUK 0brsI9V79prR+fe0lLa3P/h3/DjZ9qZODbMK1d2dp51rhtxgaQWlLOCHSuMGiVSdfux6 vPoD7GSvZQ78RKdafDHCnHkhLLf7+k1HnKI1zeG1VGUp6elAEJAR680jfDqHEEB7YHUu Nlsw== X-Forwarded-Encrypted: i=1; AJvYcCU7qvGZSu7OY7gJGLnKzQ2WV38R6vDvchKqdI6Fdnbs37tUDN2B7/AU/gktSP0wazZxqAKWzE3GGA==@kvack.org X-Gm-Message-State: AOJu0YwLQz2L7SUp1Qlpl2rvzEWB1QUSqhxTTBbdH3DywPqW5iM9pqz2 Pgio71lcuDp+zbBYEky+5i3p1UIez8OnW4GXHDB8SUQ9X1QRpi/JPkS5RJ6tyQOwXWx/FwXvlY9 Y8o/uIdRN5F4nfndLwUc+CeVW3qTvnP4fNSW+ X-Gm-Gg: ASbGncsnzFIBwofFac1KzRLwGmxSHcH2Yuxman3BcwIwbO3JyM6y4a57E1AmqQcITBX 8Qds8MXOwY8y5GMR2mq0vHR0jH7UflEi3ptFANtyFUkrCh17FiwOSvWSgct4Ibw== X-Google-Smtp-Source: AGHT+IHqYIg7U+CrWVjFtXE+tQkUDF5NR0AG+Em2XDri9p3SsXmXob/NyUrItbaDN5cKZUD/VTFF5tH6/FiAskeGJ+I= X-Received: by 2002:a05:622a:3ca:b0:460:f093:f259 with SMTP id d75a77b69052e-463572ad5c8mr3719181cf.22.1731601138301; Thu, 14 Nov 2024 08:18:58 -0800 (PST) MIME-Version: 1.0 References: <20241112194635.444146-1-surenb@google.com> <20241112194635.444146-5-surenb@google.com> <54b8d0b9-a1c7-4c1b-a588-2e5308a977fb@suse.cz> In-Reply-To: From: Suren Baghdasaryan Date: Thu, 14 Nov 2024 08:18:47 -0800 Message-ID: Subject: Re: [PATCH v2 4/5] mm: make vma cache SLAB_TYPESAFE_BY_RCU To: "Liam R. Howlett" , Suren Baghdasaryan , Matthew Wilcox , Vlastimil Babka , akpm@linux-foundation.org, lorenzo.stoakes@oracle.com, mhocko@suse.com, hannes@cmpxchg.org, mjguzik@gmail.com, oliver.sang@intel.com, mgorman@techsingularity.net, david@redhat.com, peterx@redhat.com, oleg@redhat.com, dave@stgolabs.net, paulmck@kernel.org, brauner@kernel.org, dhowells@redhat.com, hdanton@sina.com, hughd@google.com, minchan@google.com, jannh@google.com, shakeel.butt@linux.dev, souravpanda@google.com, pasha.tatashin@soleen.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, kernel-team@android.com Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Rspamd-Queue-Id: D1C0920006 X-Stat-Signature: hh5yuds9juco8e5kxnshdtyqiwogn6u3 X-Rspam-User: X-Rspamd-Server: rspam05 X-HE-Tag: 1731601093-871770 X-HE-Meta: U2FsdGVkX19FJo5S+y+8FTlEJJCV6MqpnrvZEEhgBqxxzGo5BLUNPrMPH1gorUhZiIrJdid6y0NSV+oGer0Sl4Wh4dWBEvBTUyyp0xkit5TP5Lg33IUDWUsLzx8FvHjSq7Vin4vWtQERy2FUteoPJLjqG1mjyd9xq1sK5U5hYU+bNG8GLwWxzfffDp+yngkdeNLiIq8Cbpyt3S/xKsaQRFtsUUOOT7uL1byCU5/s7CgkJiezHUkw3T/BG4rVrtDxzEMFjMfeFgOx9/CNtK4zMASHqofv9rzcTfdCC5jNqzZt8GrOGdmz/TDefeb04vVex7jHX2wxEQ+sg2eKPOskiSfhPmEjEyckeHlU55YQ/WhbaEWZqsNKEAVm3/4SefrxsUK9tTSdjUjQL5HkCIgiriUahgdcv6xs1eu/8uP2bJTj3iD+UVnYkHuZqrnAw8MHQOxZIPqcpkgS4C+ca4+6ncn1c/a7el1+HfG4cWe1jy8O2u/hq7WHWA89/McjVID0+R+qcoZhAumgSbY/hPfK5pijRdZhIBnXBia55FvU7YyJ8KLXLB7ylIa+yvegXJ5eJldw6v7w9mJsoExCkNM+dLrAQ3SZCsRpwMMu5XV94Wu/C1fWjzvDZvONGuTvGfptZOBVgZNWnXTEQCpimX/XC+ZwwpxqV0Hjxp73adDeQRSWDk4BCBhTdRYQZ821UQVIct+D4vPsjCJqL8pHe0rmir1M3mrDPETRHR3IdieBMuccSXF7YIjI2dk5Zl4qDRC4/4cQoNymJsLWtZVrKmowjpxmgCwnaso+c3PaMRiyuinDw2Unh6GCmpr3LOPefLKdq1UKydjnXFv44VSB4R9+h3nHvdt76/1ITVw3uBzjqVSQRCJUCKksJIhEWFzt01n7h/5SDnDkg+ZmmGEQxg3DP9naejmBDP1XAa6zKrL2BNjUy4Qpl7Z3vwOZ1ufT8MvODO7U8YBmBh+W3MxSNrm estTXuKM n65D6Jn1tS0EdhC7lE8/Tb8blZnT36QczdEkZIkLT1rCoZvV22MAWfE+6GdJ35DbpBpIaPGd0igvFsuMoHD7VmGN7FPjCVR67wLI0IZUUi+tkoE3MYOIGyaLoWOaNx8THxFRbS2ZPRK92TFdqhsh13kes3+UlrW3sIcy7HkTJj9H1O69wBqoj53uz8wW7Kr/TEqoOQxqTlzJ+1TseG1me8CM5208yzh7+GyfwI974nMHo5ibOap5tjYII3inCfOvkFz+bpEJBl8U+r0PATjK7se6ldDOug24WrBypihWuY1GlIeQbJOwZRqxQnLo+Ta9iMFVYs6AwMrzSEXvolprRjIlzMvibB5AuTU2kY0LoNVqmYFDvx/qIshr5uQ== X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: On Wed, Nov 13, 2024 at 11:05=E2=80=AFAM Suren Baghdasaryan wrote: > > On Wed, Nov 13, 2024 at 7:47=E2=80=AFAM Suren Baghdasaryan wrote: > > > > On Wed, Nov 13, 2024 at 7:29=E2=80=AFAM Liam R. Howlett wrote: > > > > > > * Suren Baghdasaryan [241113 10:25]: > > > > On Wed, Nov 13, 2024 at 7:23=E2=80=AFAM 'Liam R. Howlett' via kerne= l-team > > > > wrote: > > > > > > > > > > * Matthew Wilcox [241113 08:57]: > > > > > > On Wed, Nov 13, 2024 at 07:38:02AM -0500, Liam R. Howlett wrote= : > > > > > > > > Hi, I was wondering if we actually need the detached flag. = Couldn't > > > > > > > > "detached" simply mean vma->vm_mm =3D=3D NULL and we save 4= bytes? Do we ever > > > > > > > > need a vma that's detached but still has a mm pointer? I'd = hope the places > > > > > > > > that set detached to false have the mm pointer around so it= 's not inconvenient. > > > > > > > > > > > > > > I think the gate vmas ruin this plan. > > > > > > > > > > > > But the gate VMAs aren't to be found in the VMA tree. Used to = be that > > > > > > was because the VMA tree was the injective RB tree and so VMAs = could > > > > > > only be in one tree at a time. We could change that now! > > > > > > > > > > \o/ > > > > > > > > > > > > > > > > > Anyway, we could use (void *)1 instead of NULL to indicate a "d= etached" > > > > > > VMA if we need to distinguish between a detached VMA and a gate= VMA. > > > > > > > > > > I was thinking a pointer to itself vma->vm_mm =3D vma, then a che= ck for > > > > > this, instead of null like we do today. > > > > > > > > The motivation for having a separate detached flag was that vma->vm= _mm > > > > is used when read/write locking the vma, so it has to stay valid ev= en > > > > when vma gets detached. Maybe we can be more cautious in > > > > vma_start_read()/vma_start_write() about it but I don't recall if > > > > those were the only places that was an issue. > > > > > > We have the mm form the callers though, so it could be passed in? > > > > Let me try and see if something else blows up. When I was implementing > > per-vma locks I thought about using vma->vm_mm to indicate detached > > state but there were some issues that caused me reconsider. > > Yeah, a quick change reveals the first mine explosion: > > [ 2.838900] BUG: kernel NULL pointer dereference, address: 00000000000= 00480 > [ 2.840671] #PF: supervisor read access in kernel mode > [ 2.841958] #PF: error_code(0x0000) - not-present page > [ 2.843248] PGD 800000010835a067 P4D 800000010835a067 PUD 10835b067 PM= D 0 > [ 2.844920] Oops: Oops: 0000 [#1] PREEMPT SMP PTI > [ 2.846078] CPU: 2 UID: 0 PID: 1 Comm: init Not tainted > 6.12.0-rc6-00258-ga587fcd91b06-dirty #111 > [ 2.848277] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), > BIOS 1.16.3-debian-1.16.3-2 04/01/2014 > [ 2.850673] RIP: 0010:unmap_vmas+0x84/0x190 > [ 2.851717] Code: 00 00 00 00 48 c7 44 24 48 00 00 00 00 48 c7 44 > 24 18 00 00 00 00 48 89 44 24 28 4c 89 44 24 38 e8 b1 c0 d1 00 48 8b > 44 24 28 <48> 83 b8 80 04 00 00 00 0f 85 dd 00 00 00 45 0f b6 ed 49 83 > ec 01 > [ 2.856287] RSP: 0000:ffffa298c0017a18 EFLAGS: 00010246 > [ 2.857599] RAX: 0000000000000000 RBX: 00007f48ccbb4000 RCX: 00007f48c= cbb4000 > [ 2.859382] RDX: ffff8918c26401e0 RSI: ffffa298c0017b98 RDI: ffffa298c= 0017ab0 > [ 2.861156] RBP: 00007f48ccdb6000 R08: 00007f48ccdb6000 R09: 000000000= 0000001 > [ 2.862941] R10: 0000000000000040 R11: ffff8918c2637108 R12: 000000000= 0000001 > [ 2.864719] R13: 0000000000000001 R14: ffff8918c26401e0 R15: ffffa298c= 0017b98 > [ 2.866472] FS: 0000000000000000(0000) GS:ffff8927bf080000(0000) > knlGS:0000000000000000 > [ 2.868439] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 > [ 2.869877] CR2: 0000000000000480 CR3: 000000010263e000 CR4: 000000000= 0750ef0 > [ 2.871661] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 000000000= 0000000 > [ 2.873419] DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7: 000000000= 0000400 > [ 2.875185] PKRU: 55555554 > [ 2.875871] Call Trace: > [ 2.876503] > [ 2.877047] ? __die+0x1e/0x60 > [ 2.877824] ? page_fault_oops+0x17b/0x4a0 > [ 2.878857] ? exc_page_fault+0x6b/0x150 > [ 2.879841] ? asm_exc_page_fault+0x26/0x30 > [ 2.880886] ? unmap_vmas+0x84/0x190 > [ 2.881783] ? unmap_vmas+0x7f/0x190 > [ 2.882680] vms_clear_ptes+0x106/0x160 > [ 2.883621] vms_complete_munmap_vmas+0x53/0x170 > [ 2.884762] do_vmi_align_munmap+0x15e/0x1d0 > [ 2.885838] do_vmi_munmap+0xcb/0x160 > [ 2.886760] __vm_munmap+0xa4/0x150 > [ 2.887637] elf_load+0x1c4/0x250 > [ 2.888473] load_elf_binary+0xabb/0x1680 > [ 2.889476] ? __kernel_read+0x111/0x320 > [ 2.890458] ? load_misc_binary+0x1bc/0x2c0 > [ 2.891510] bprm_execve+0x23e/0x5e0 > [ 2.892408] kernel_execve+0xf3/0x140 > [ 2.893331] ? __pfx_kernel_init+0x10/0x10 > [ 2.894356] kernel_init+0xe5/0x1c0 > [ 2.895241] ret_from_fork+0x2c/0x50 > [ 2.896141] ? __pfx_kernel_init+0x10/0x10 > [ 2.897164] ret_from_fork_asm+0x1a/0x30 > [ 2.898148] > > Looks like we are detaching VMAs and then unmapping them, where > vms_clear_ptes() uses vms->vma->vm_mm. I'll try to clean up this and > other paths and will see how many changes are required to make this > work. Ok, my vma->detached deprecation effort got to the point that my QEMU boots. The change is not pretty and I'm quite sure I did not cover all cases yet (like hugepages): arch/arm/kernel/process.c | 2 +- arch/arm64/kernel/vdso.c | 4 +- arch/loongarch/kernel/vdso.c | 2 +- arch/powerpc/kernel/vdso.c | 2 +- arch/powerpc/platforms/pseries/vas.c | 2 +- arch/riscv/kernel/vdso.c | 4 +- arch/s390/kernel/vdso.c | 2 +- arch/s390/mm/gmap.c | 2 +- arch/x86/entry/vdso/vma.c | 2 +- arch/x86/kernel/cpu/sgx/encl.c | 2 +- arch/x86/um/mem_32.c | 2 +- drivers/android/binder_alloc.c | 2 +- drivers/gpu/drm/i915/i915_mm.c | 4 +- drivers/infiniband/core/uverbs_main.c | 4 +- drivers/misc/sgi-gru/grumain.c | 2 +- fs/exec.c | 2 +- fs/hugetlbfs/inode.c | 3 +- include/linux/mm.h | 111 +++++++++++++++++--------- include/linux/mm_types.h | 6 -- kernel/bpf/arena.c | 2 +- kernel/fork.c | 6 +- mm/debug_vm_pgtable.c | 2 +- mm/internal.h | 2 +- mm/madvise.c | 4 +- mm/memory.c | 39 ++++----- mm/mmap.c | 9 +-- mm/nommu.c | 6 +- mm/oom_kill.c | 2 +- mm/vma.c | 62 +++++++------- mm/vma.h | 2 +- net/ipv4/tcp.c | 4 +- 31 files changed, 164 insertions(+), 136 deletions(-) Many of the unmap_* and zap_* functions should get an `mm` parameter to make this work. So, if we take this route, it should definitely be a separate patch, which will likely cause some instability issues for some time until all the edge cases are ironed out. I would like to proceed with this patch series first before attempting to deprecate vma->detached. Let me know if you have objections to this plan. > > > > > > > > > > > > > > > > > > > > Either way, we should make it a function so it's easier to reuse = for > > > > > whatever we need in the future, wdyt? > > > > > > > > > > To unsubscribe from this group and stop receiving emails from it,= send an email to kernel-team+unsubscribe@android.com. > > > > >