From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from mail-pg1-f197.google.com (mail-pg1-f197.google.com [209.85.215.197]) by kanga.kvack.org (Postfix) with ESMTP id 1ED188E0001 for ; Thu, 10 Jan 2019 20:47:33 -0500 (EST) Received: by mail-pg1-f197.google.com with SMTP id a2so7470603pgt.11 for ; Thu, 10 Jan 2019 17:47:33 -0800 (PST) Received: from ipmail06.adl6.internode.on.net (ipmail06.adl6.internode.on.net. [150.101.137.145]) by mx.google.com with ESMTP id b12si5932067pls.32.2019.01.10.17.47.30 for ; Thu, 10 Jan 2019 17:47:31 -0800 (PST) Date: Fri, 11 Jan 2019 12:47:28 +1100 From: Dave Chinner Subject: Re: [PATCH] mm/mincore: allow for making sys_mincore() privileged Message-ID: <20190111014728.GL27534@dastard> References: <20190108044336.GB27534@dastard> <20190109022430.GE27534@dastard> <20190109043906.GF27534@dastard> <20190110004424.GH27534@dastard> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: Sender: owner-linux-mm@kvack.org List-ID: To: Andy Lutomirski Cc: Linus Torvalds , Jiri Kosina , Matthew Wilcox , Jann Horn , Andrew Morton , Greg KH , Peter Zijlstra , Michal Hocko , Linux-MM , kernel list , Linux API On Wed, Jan 09, 2019 at 09:26:41PM -0800, Andy Lutomirski wrote: > Since direct IO has been brought up, I have a question. I've wondered > for years why direct IO works the way it does. If I were implementing > it from scratch, my first inclination would be to use the page cache > instead of fighting it. To do a single-page direct read, I would look > that page up in the page cache (i.e. i_pages these days). If the page > is there, I would do a normal buffered read. If the page is not Therein lies the problem. Copying data is prohibitively expensive, and that's the primary reason for O_DIRECT existing. i.e. O_DIRECT is a low-overhead, zero-copy data movement interface. The moment we switch from using CPU to dispatch IO to copying data, performance goes down because we will be unable to keep storage pipelines full. IOWs, any rework of O_DIRECT that involves copying data is a non-starter. But let's bring this back to the issue at hand - observability of page cache residency of file pages. If th epage is caceh resident, then it will have a latency of copying that data out of the page (i.e. very low latency). If the page is not resident, then it will do IO and take much, much longer to complete. i.e. we have clear timing differences between cachce hit and cache miss IO. This is exactly the timing information needed for observing page cache residency. We need to work out how to make page cache residency less observable, not add new, near perfect observation mechanisms that third parties can easily exploit... Cheers, Dave. -- Dave Chinner david@fromorbit.com