From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 0647BC83F09 for ; Fri, 4 Jul 2025 06:26:34 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 7A0AD6B800D; Fri, 4 Jul 2025 02:26:34 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 750F56B800A; Fri, 4 Jul 2025 02:26:34 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 6676D6B800D; Fri, 4 Jul 2025 02:26:34 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0013.hostedemail.com [216.40.44.13]) by kanga.kvack.org (Postfix) with ESMTP id 51F3C6B800A for ; Fri, 4 Jul 2025 02:26:34 -0400 (EDT) Received: from smtpin08.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay04.hostedemail.com (Postfix) with ESMTP id 0309B1A1DA2 for ; Fri, 4 Jul 2025 06:26:33 +0000 (UTC) X-FDA: 83625598308.08.A18FA6B Received: from mail-pf1-f176.google.com (mail-pf1-f176.google.com [209.85.210.176]) by imf25.hostedemail.com (Postfix) with ESMTP id 44C00A0002 for ; Fri, 4 Jul 2025 06:26:31 +0000 (UTC) Authentication-Results: imf25.hostedemail.com; dkim=pass header.d=bytedance.com header.s=google header.b="Njl3hG/N"; spf=pass (imf25.hostedemail.com: domain of lizhe.67@bytedance.com designates 209.85.210.176 as permitted sender) smtp.mailfrom=lizhe.67@bytedance.com; dmarc=pass (policy=quarantine) header.from=bytedance.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1751610392; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-transfer-encoding:content-transfer-encoding: in-reply-to:references:dkim-signature; bh=vtAMrec9bYb31s5/Huy8ScC44Y0MeGNii8Ny4beAYaA=; b=UcJ5LvgrJmf+sWnuPtvz9ZQDDfut5i9tcg1c0xXbwwZBkRKF0WlmFNiuY6g88G2j6BZ1VV l/Rg+v4UcWQRR4DCbLh9es6SiPooQ3e7x3g1Cvu5QYqo8rS9cieOL2bgYyVwLCLG20/kqm vCsk+bLX6pr8SMyHFBeMbL5/irC3hu8= ARC-Authentication-Results: i=1; imf25.hostedemail.com; dkim=pass header.d=bytedance.com header.s=google header.b="Njl3hG/N"; spf=pass (imf25.hostedemail.com: domain of lizhe.67@bytedance.com designates 209.85.210.176 as permitted sender) smtp.mailfrom=lizhe.67@bytedance.com; dmarc=pass (policy=quarantine) header.from=bytedance.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1751610392; a=rsa-sha256; cv=none; b=R0UsFsDv1ictvxfDzKvTMPZ1iDENnRaxJJorGT+9Y/BJQ1xock8PxSa3Y35phkGT16v/Co Su9jexhP1TLyASI9lYKKHdksR+JxHbRy1G3qXwKzsBO3uOXprUo/7eBJIqgPNwXw3Fi8q5 XLV1njEvvMNneTvYa3eBFim54ZTNNAU= Received: by mail-pf1-f176.google.com with SMTP id d2e1a72fcca58-742c3d06de3so863274b3a.0 for ; Thu, 03 Jul 2025 23:26:30 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1751610390; x=1752215190; darn=kvack.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to; bh=vtAMrec9bYb31s5/Huy8ScC44Y0MeGNii8Ny4beAYaA=; b=Njl3hG/NkZBQWg2rknBZ+uOL+bCMj162ej0kKD0SeFBXbZRldUGyvZQHwuL0BpsDHn cxOFnMwSPnkAvq34HVK+5Q/WojiOiqkhRCTBwhFIWg8MuYVCOwPUpSh3EDboak8sG7ew FnGwiy1yHXQywjrxacpHEAdLwoho13bAndGys0wFeHe5BaJeL5RG+aTxIJL/DuM1Z+SQ cqUZ8YjDpg9RfM6Ve8z6C6ttqzpJQZiAm4oSt6dCPFgH+oSaFnolFYea3ua3+wvRRA7O pCFoGIvMDwm3qmXUWYIJWA7TPnrdHV5PKnu3v2/sgyK9v7gu68UR5q1RQ6Ctpt2QRC4F otug== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1751610390; x=1752215190; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=vtAMrec9bYb31s5/Huy8ScC44Y0MeGNii8Ny4beAYaA=; b=BlqnhZbqjsWjv8TRYKk9PEQGWg41ki+C+xIbjJOsEeygy4gPe/kCfI2GZUuGPQB0z5 OIB76puzGJCYiT/qEIQkBvnHFLqR4fHHut21xBdaeUysLGrN3+naENQWEOr7yu39CzId PPm2BWDWJ2xbqtej8wchHRxXmyUOF+t5gWZn1mEqyII/953+qEE6NjKsl54LY3/BZ+oE z3JPnWF1VE5BhUsstYrCBvKd13mgur2QDJ3u4ljzhF5e/2Fx3hLgYLHiut7piB2JEjeS iq3GGTT2GMgE8/U6aAWE6IvPccWe6PABUHfYCxiBaBAkf7qGf/O/uJhrQkkpJ1HtLjTv XaCA== X-Forwarded-Encrypted: i=1; AJvYcCV+u2r1+qX1Hc6WtIbwoEhin3nnR7GbA+aH4TukNo7Z2AiM8XGVBMA6mYj4/vfoof72+kQ/J4It4g==@kvack.org X-Gm-Message-State: AOJu0YyJmdBUG/4Pl5264Q+iHEdVveRghIX8tljXfMhWnW+wYcXf1xqT 34w3E4MNyBDRdpx2zNTnnJl/RHiP5X2qhKP1bgbfqtAdgO7WaF7PQV5qQgsHxOcZBWU= X-Gm-Gg: ASbGnctzJYPhbBh0JzMovCG/JFg32rZNpI6ywG3Rueq3+t1jAGOthEkZ1qeHNqam+3t al4J/jCqLFlILh8GmalzZTLfX4cOyu0TaDz3t7w2z28E//dX6apvjQ/cbuy1BldYz4APTqUi8YF dtjd9CL/+8KaQYmwpK3UenLes8BMBjlcR7WQ5UegMYr89uRrW1RTL9jzWNSCXjVQbyVC0DkK5na KHJxh9YmColSXFUS0U9JyxZKckBVxk/Xezx2vyrYhHqFtXP6cu9X12/aHolWayqTljHg2/ttTb9 7dWsoxHsJZt+iTqWzgZzpDVAIM7yBO2e4e+KlV1YhT/W2X9EQaoBZvgirVVT7P4yk+c4VtBSPnF 402ig09sWcLKY X-Google-Smtp-Source: AGHT+IFTOzGu3zFWSHTrf4KF59AoRGl/bkuEswL36beEW28ePztb4UqI8rnrlZ/0ccy+thk4y06H6Q== X-Received: by 2002:a05:6a21:694:b0:201:b65:81ab with SMTP id adf61e73a8af0-225b9f64162mr2679259637.23.1751610389786; Thu, 03 Jul 2025 23:26:29 -0700 (PDT) Received: from localhost.localdomain ([203.208.189.8]) by smtp.gmail.com with ESMTPSA id 41be03b00d2f7-b38ee5f643dsm1183240a12.37.2025.07.03.23.26.26 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Thu, 03 Jul 2025 23:26:29 -0700 (PDT) From: lizhe.67@bytedance.com To: alex.williamson@redhat.com, akpm@linux-foundation.org, david@redhat.com, peterx@redhat.com, jgg@ziepe.ca Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, lizhe.67@bytedance.com Subject: [PATCH v2 0/5] vfio/type1: optimize vfio_pin_pages_remote() and vfio_unpin_pages_remote() Date: Fri, 4 Jul 2025 14:25:57 +0800 Message-ID: <20250704062602.33500-1-lizhe.67@bytedance.com> X-Mailer: git-send-email 2.45.2 MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Stat-Signature: 85yx16ccgfzfj1po5diwh5cgxfcykaeh X-Rspamd-Queue-Id: 44C00A0002 X-Rspam-User: X-Rspamd-Server: rspam07 X-HE-Tag: 1751610391-475748 X-HE-Meta: U2FsdGVkX1+jyQdLT85HigryJpCDyl0O3wvWRGLqLV8xKvgX1ghuh3gsIom5P/mhQ5wCRloIoT8JF0dvVLMfkAeaWzcz+DsNQ75KXBTsgv+RGvm9pZc492W63iKZHMXONzmAVFcU2i50fhJho/9+awfSCxmUUm/DqLVPis73fnJ5fjg4GnFTsT2ABN6oTEXxDKWa+xv18bzerrcKkNuNNwaFOxvFsZcxBZygGfRgjBEfcl4Kmo5qe/JH+DXdgXXcl5It8F9ZcEFOwno0Vdi8irXbItQpg/QhhNc2xMTwenZN9w7V3WM4l/SIebkdL537KWZhBqg+M2OOzu3b2QU7PxCcRsQZId07Ep+25Iz83jlHz2hmRYGZ1VAuIpmxo3TEtzspYPNR7dTuxCChTNLEBKAVrGIy642cIIAtkVOZWtYToXtm2llInR6mpSSFCt/twfHvFzw5+EzCWXzFAesJtC+JdjjTcImOWnnLCDMKWBqAW9srSdi8PkHNhcsbsEt7ZUqGrdvxa2DSo1wjaCz0i2k//MUmeBQ3Zed1CjkI6p7P3UwBKH9bEP7Osu1BlqIB9WhtZcZdFpJHo8y0bMJ91mBFSBNxVBs6h0YXZEVhvBcblvOzANfWRlP590esXmO7iur1cBVeKW0VlaQG71GsEJbpTIrBPEgKnrKKyDarQRpBLlYsSQvMnnqUDDvm6ZD3YcuR0x/dPGhe2dZMmZS7Pp3JE4+hqCoxlyAdvDghspzm9vk8t9SkpkR2zTGY3KaYcDUBzrc0shNTBDwg4evCDF2nlDa3NwkxxquxMGWFk2aCua/LFFY0h1k/xSqBjoTqg4mddr4SQE+WFyLl4cJGyN5TO9MS9Q0/VlteXzlxmSdDBhYzNcc12EkItc+LxR4eYy52AC2Ao/ESCgmfk3dK9WrXkv8Ci95YTZZGXcUcrZmp2MlD6HvDI43C44cPL/VrjpWuks9oCetjR3uJwY4 zXTMXr+o XfAiFXmNqW1u7AdtykbAPSqYCqII++L4QKo3hWuqfINjX62uaBHbsqSSM/HQaEfcfr5hc+oNLZxFeoGaj1UTEuDhLlw+1tdJ3Q3zPNSEvhKlN65fPTbzdaKXRd53P4bfcb2I34S9wsakFzkhtyzNGe90ElQsHwFFhqFjDVJBvfZQvnrfEvbMNVQKLS+jTNIoUY1fvmp3vORfrzUxg48Rn94a0fIHVfErOObW9DLtIK15RL02/EAz41JF5fYS6Kz5j63+yqSIO/eMC7zcT1HCDtrok2Z+3UUqF74P2Tk1qOGDwKY5eoQWsMGRkt4ghxL2Tx9p4X2aHpPUuIemKnpBT3duw4a/P1Z9GVmpx2hbbz3miByM= X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: From: Li Zhe This patchset is an integration of the two previous patchsets[1][2]. When vfio_pin_pages_remote() is called with a range of addresses that includes large folios, the function currently performs individual statistics counting operations for each page. This can lead to significant performance overheads, especially when dealing with large ranges of pages. The function vfio_unpin_pages_remote() has a similar issue, where executing put_pfn() for each pfn brings considerable consumption. This patchset primarily optimizes the performance of the relevant functions by batching the less efficient operations mentioned before. The first two patch optimizes the performance of the function vfio_pin_pages_remote(), while the remaining patches optimize the performance of the function vfio_unpin_pages_remote(). The performance test results, based on v6.16-rc4, for completing the 16G VFIO MAP/UNMAP DMA, obtained through unit test[3] with slight modifications[4], are as follows. Base(6.16-rc4): ./vfio-pci-mem-dma-map 0000:03:00.0 16 ------- AVERAGE (MADV_HUGEPAGE) -------- VFIO MAP DMA in 0.047 s (340.2 GB/s) VFIO UNMAP DMA in 0.135 s (118.6 GB/s) ------- AVERAGE (MAP_POPULATE) -------- VFIO MAP DMA in 0.280 s (57.2 GB/s) VFIO UNMAP DMA in 0.312 s (51.3 GB/s) ------- AVERAGE (HUGETLBFS) -------- VFIO MAP DMA in 0.052 s (310.5 GB/s) VFIO UNMAP DMA in 0.136 s (117.3 GB/s) With this patchset: ------- AVERAGE (MADV_HUGEPAGE) -------- VFIO MAP DMA in 0.027 s (600.7 GB/s) VFIO UNMAP DMA in 0.045 s (357.0 GB/s) ------- AVERAGE (MAP_POPULATE) -------- VFIO MAP DMA in 0.261 s (61.4 GB/s) VFIO UNMAP DMA in 0.288 s (55.6 GB/s) ------- AVERAGE (HUGETLBFS) -------- VFIO MAP DMA in 0.031 s (516.4 GB/s) VFIO UNMAP DMA in 0.045 s (353.9 GB/s) For large folio, we achieve an over 40% performance improvement for VFIO MAP DMA and an over 66% performance improvement for VFIO DMA UNMAP. For small folios, the performance test results show a slight improvement with the performance before optimization. [1]: https://lore.kernel.org/all/20250529064947.38433-1-lizhe.67@bytedance.com/ [2]: https://lore.kernel.org/all/20250620032344.13382-1-lizhe.67@bytedance.com/#t [3]: https://github.com/awilliam/tests/blob/vfio-pci-mem-dma-map/vfio-pci-mem-dma-map.c [4]: https://lore.kernel.org/all/20250610031013.98556-1-lizhe.67@bytedance.com/ Li Zhe (5): mm: introduce num_pages_contiguous() vfio/type1: optimize vfio_pin_pages_remote() vfio/type1: batch vfio_find_vpfn() in function vfio_unpin_pages_remote() vfio/type1: introduce a new member has_rsvd for struct vfio_dma vfio/type1: optimize vfio_unpin_pages_remote() drivers/vfio/vfio_iommu_type1.c | 111 ++++++++++++++++++++++++++------ include/linux/mm.h | 20 ++++++ 2 files changed, 110 insertions(+), 21 deletions(-) --- Changelogs: v1->v2: - Update the performance test results. - The function num_pages_contiguous() is extracted and placed in a separate commit. - The phrase 'for large folio' has been removed from the patchset title. v1: https://lore.kernel.org/all/20250630072518.31846-1-lizhe.67@bytedance.com/ -- 2.20.1