From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id A2843CE79AC for ; Wed, 20 Sep 2023 08:50:59 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 2CDC66B0135; Wed, 20 Sep 2023 04:50:59 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 27D9B6B0136; Wed, 20 Sep 2023 04:50:59 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 146036B0137; Wed, 20 Sep 2023 04:50:59 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0015.hostedemail.com [216.40.44.15]) by kanga.kvack.org (Postfix) with ESMTP id 04E4C6B0135 for ; Wed, 20 Sep 2023 04:50:59 -0400 (EDT) Received: from smtpin12.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay10.hostedemail.com (Postfix) with ESMTP id B3D4BC0E73 for ; Wed, 20 Sep 2023 08:50:58 +0000 (UTC) X-FDA: 81256355796.12.0046A4D Received: from rivendell.linuxfromscratch.org (rivendell.linuxfromscratch.org [208.118.68.85]) by imf21.hostedemail.com (Postfix) with ESMTP id ACD541C000A for ; Wed, 20 Sep 2023 08:50:56 +0000 (UTC) Authentication-Results: imf21.hostedemail.com; dkim=pass header.d=linuxfromscratch.org header.s=cert4 header.b=eObk1KcH; spf=pass (imf21.hostedemail.com: domain of xry111@linuxfromscratch.org designates 208.118.68.85 as permitted sender) smtp.mailfrom=xry111@linuxfromscratch.org; dmarc=pass (policy=quarantine) header.from=linuxfromscratch.org ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1695199856; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=aOrmEQVHlGP4NIc/KsyybAf3h59NYTPKZ1pCQKmMgow=; b=DXJC3Ch3I9WtnVGkb7H8CClTB6NMVvwAGNQpKIO39Mrc6iROlg9MPsKUbzCQAK+jrDP/vZ 53EICB6nN7t1yn0Wjc9fyulVW/Cxhq8dqJw2DG1BrMxDTmxa0fN5GD7Qfqnxw7n7pyoWfo gs4vjbgsDSfFdN0/NeDqwQZa+lo+pWU= ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1695199856; a=rsa-sha256; cv=none; b=znFITAT1wdpTgvtMy+wzyFiyzzbJukBrCuwwYA5kuTZEp3VXm2HJu63rOJP4HSrJKyxhrA KXSzlvZl0GNrr49/dESiPzhbjwgfiMOdA9EuW7bjQ3YdCQh0HfmJSiwe2aNDzdKZTfCxM9 Vpc6ncEGTTyiUvH0IvPxC74XysJmHRw= ARC-Authentication-Results: i=1; imf21.hostedemail.com; dkim=pass header.d=linuxfromscratch.org header.s=cert4 header.b=eObk1KcH; spf=pass (imf21.hostedemail.com: domain of xry111@linuxfromscratch.org designates 208.118.68.85 as permitted sender) smtp.mailfrom=xry111@linuxfromscratch.org; dmarc=pass (policy=quarantine) header.from=linuxfromscratch.org Received: from [192.168.3.211] (unknown [36.44.137.238]) by rivendell.linuxfromscratch.org (Postfix) with ESMTPSA id 1EF0B1C1DCD; Wed, 20 Sep 2023 08:50:31 +0000 (GMT) X-Virus-Status: Clean X-Virus-Scanned: clamav-milter 1.0.0 at rivendell.linuxfromscratch.org DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=linuxfromscratch.org; s=cert4; t=1695199854; bh=3Z1+GidcJMavvwBcREoqn7oXqIoMxeJ9etaya6Fu9M0=; h=Subject:From:To:Cc:Date:In-Reply-To:References; b=eObk1KcHIFgyOq0PM0GHhNP2/hul2wPtxGcaRxEcqFDCkKZIPj2t/9TxEWDmhO6YZ aezBQiQYv28rq2gmvKDO1vLC0cjJ4lkZqX+AWhOX80av065LcuDU/dQQlCUdAz+F5v EMnFKjI7RQ2SB6tMDzdbLxzYFo10Nqo0SoW1H1a/NoxdCJZXpxyBoREnlKsc2adPNi tbCaBepclPdYSRgGsku+r9JOsLdUStV6efcHtTldH9x3uUyad5uMYo/g4osBB/tLJ8 12slT1IuoRsY1p0CjzdaDH1JpjFGWIS5aueo9MwBbmOsZYKMRjHoGNDe8fOb9NSvgA Jjo932cwggiGw== Message-ID: <34d45270efccc44b64af835e73c1d1e111ce5098.camel@linuxfromscratch.org> Subject: Re: [PATCH v7 12/13] ext4: switch to multigrain timestamps From: Xi Ruoyao To: Christian Brauner , Jeff Layton Cc: Bruno Haible , Jan Kara , bug-gnulib@gnu.org, Alexander Viro , Eric Van Hensbergen , Latchesar Ionkov , Dominique Martinet , Christian Schoenebeck , David Howells , Marc Dionne , Chris Mason , Josef Bacik , David Sterba , Xiubo Li , Ilya Dryomov , Jan Harkes , coda@cs.cmu.edu, Tyler Hicks , Gao Xiang , Chao Yu , Yue Hu , Jeffle Xu , Namjae Jeon , Sungjong Seo , Jan Kara , Theodore Ts'o , Andreas Dilger , Jaegeuk Kim , OGAWA Hirofumi , Miklos Szeredi , Bo b Peterson , Andreas Gruenbacher , Greg Kroah-Hartman , Tejun Heo , Trond Myklebust , Anna Schumaker , Konstantin Komarov , Mark Fasheh , Joel Becker , Joseph Qi , Mike Marshall , Martin Brandenburg , Luis Chamberlain , Kees Cook , Iurii Zaikin , Steve French , Paulo Alcantara , Ronnie Sahlberg , Shyam Prasad N , Tom Talpey , Sergey Senozhatsky , Richard Weinberger , Hans de Goede , Hugh Dickins , Andrew Morton , Amir Goldstein , "Darrick J. Wong" , Benjamin Coddington , linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org, v9fs@lists.linux.dev, linux-afs@lists.infradead.org, linux-btrfs@vger.kernel.org, ceph-devel@vger.kernel.org, codalist@coda.cs.cmu.edu, ecryptfs@vger.kernel.org, linux-erofs@lists.ozlabs.org, linux-ext4@vger.kernel.org, linux-f2fs-devel@lists.sourceforge.net, cluster-devel@redhat.com, linux-nfs@vger.kernel.org, ntfs3@lists.linux.dev, ocfs2-devel@lists.linux.dev, devel@lists.orangefs.org, linux-cifs@vger.kernel.org, samba-technical@lists.samba.org, linux-mtd@lists.infradead.org, linux-mm@kvack.org, linux-unionfs@vger.kernel.org, linux-xfs@vger.kernel.org Date: Wed, 20 Sep 2023 16:50:26 +0800 In-Reply-To: <20230920-leerung-krokodil-52ec6cb44707@brauner> References: <20230807-mgctime-v7-0-d1dec143a704@kernel.org> <20230919110457.7fnmzo4nqsi43yqq@quack3> <1f29102c09c60661758c5376018eac43f774c462.camel@kernel.org> <4511209.uG2h0Jr0uP@nimes> <08b5c6fd3b08b87fa564bb562d89381dd4e05b6a.camel@kernel.org> <20230920-leerung-krokodil-52ec6cb44707@brauner> Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable User-Agent: Evolution 3.50.0 MIME-Version: 1.0 X-Rspamd-Queue-Id: ACD541C000A X-Rspam-User: X-Stat-Signature: i5mwdo9pb4trijkwmqadu8m33geqfr8z X-Rspamd-Server: rspam03 X-HE-Tag: 1695199856-854834 X-HE-Meta: U2FsdGVkX18wyISfyg1NxqyDA5Z6yiHVHJgZ9XS56p3q1yMFwj02fHRb7i4voH6Yw8iQwvz8cCLlfCObmhR+tADkBGlVbP7Yev2DUNVQFSPhhnLRlmDA787saF1yIRguiVMXTkwvbCBfQQr7w1By0WDGizH26flbK3oFs0cJwpm4X7EfiMsipbaNqASw8WQSevo+PoDM2wxLE5Bxg7mrJRqfQzCB/+bIpuh2she8aRxmT2pSDeEKkceBhgn1RWJ5W13eJJpD0IzRPU74CrVCYWE91UzC95sH86b0qxRYdkBQ0A5ELBTswHQEr6AFv/sirsLNJo89r5XaZM8FRDLqUhpwMl42MHFGnP1nJyR16JqRlzMZwsq7HKX658EabIefJ3bZklgTPic1as0ojRcKv8bgEA0miWrFa8M1wq814qWD0/gwEzeDBbpI58NZZCPE5Rhil0kGB//2BHn/SXyF/aWr4y0B4Nq2ldDF1RMz+G/5rSQq3M49Odf9/Lh+Xo8F+COZeEVSkalee/NjgZOyqhDmZ/CFsxn9BBM51dibfMUezWewQxXZVroV6/y1lDoydz/LBWHJwwn6hY+P5eMuvLd49U1yoEtWPU54k6TwPLLWJdHbMbTPu4j7qKFsFLemQLuVxppzUsima2FKsz6oFfLMqdt9ngPBuc5IoMKuhPDj9nqyt5YuK/FN38NfSsLbZjrUqzpknyUnmWx3bhvpFTY/c1oQuOXOslVBelKEBZ8D7mdeBw3O7lbhOReOvQ0y+nGj8YwyU6QhSyhTj+8w/i9wesoD0KYgMvbLa1hIradcVBDEwuXxVwlbw2yvTHbyGaWFwR4nH1aEFkPEKN4ICcFT+3whI5czWM8t0uJ/0Tg8oVufqn7GLsLunGvTNnitetmlRqXLrTpFQQP5uN5IkPLOfRcmc5vaQAC6MrSyQM8gV6AKmbBoqGtuQmOZKaKjdVZd0AUMhHSMR/WMwJd qT8l+zDo PCkZPgIoxchuAM6moMic2goSw1cE6FuEi9rIrVFcZm2mUOAzP0Q+Z+YZj6e04Ue2K3V26oQQtm/kJ4NFLwP4Ar+tUCwhbHJuHAYVg+dJYttmrp1hMVk8vXCMNOqvsoRDCY8AfizP0Q76KX6YDoJuZu9tZ/Ec7/vn7lZ84N1gT553y+1b/JUBSHP+4/dG+IjTj79B62n6X3R0pveG0C9Mu1ZDp5vV5RsRQK6s1wC22jVw/6BvHd2s2w0fxxVOAQ9tep5UOH8gHqwOg5AdW8eCtPKT8wEBzNbEPsWIVnP3dBbHbRT6IbgmzCbqlR/hEPgvH5eK3I5JqZbsWdX3X7uspHA5Tv5c944MYF8WT X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: On Wed, 2023-09-20 at 10:41 +0200, Christian Brauner wrote: > > > f1 was last written to *after* f2 was last written to. If the timesta= mp of f1 > > > is then lower than the timestamp of f2, timestamps are fundamentally = broken. > > >=20 > > > Many things in user-space depend on timestamps, such as build system > > > centered around 'make', but also 'find ... -newer ...'. > > >=20 > >=20 > >=20 > > What does breakage with make look like in this situation? The "fuzz" > > here is going to be on the order of a jiffy. The typical case for make > > timestamp comparisons is comparing source files vs. a build target. If > > those are being written nearly simultaneously, then that could be an > > issue, but is that a typical behavior? It seems like it would be hard t= o > > rely on that anyway, esp. given filesystems like NFS that can do lazy > > writeback. > >=20 > > One of the operating principles with this series is that timestamps can > > be of varying granularity between different files. Note that Linux > > already violates this assumption when you're working across filesystems > > of different types. > >=20 > > As to potential fixes if this is a real problem: > >=20 > > I don't really want to put this behind a mount or mkfs option (a'la > > relatime, etc.), but that is one possibility. > >=20 > > I wonder if it would be feasible to just advance the coarse-grained > > current_time whenever we end up updating a ctime with a fine-grained > > timestamp? It might produce some inode write amplification. Files that >=20 > Less than ideal imho. >=20 > If this risks breaking existing workloads by enabling it unconditionally > and there isn't a clear way to detect and handle these situations > without risk of regression then we should move this behind a mount > option. >=20 > So how about the following: >=20 > From cb14add421967f6e374eb77c36cc4a0526b10d17 Mon Sep 17 00:00:00 2001 > From: Christian Brauner > Date: Wed, 20 Sep 2023 10:00:08 +0200 > Subject: [PATCH] vfs: move multi-grain timestamps behind a mount option >=20 > While we initially thought we can do this unconditionally it turns out > that this might break existing workloads that rely on timestamps in very > specific ways and we always knew this was a possibility. Move > multi-grain timestamps behind a vfs mount option. I agree with this solution. You can add some metainfo: Reported-by: Ken Moffat Bisected-by: Xi Ruoyao Link: https://lists.linuxfromscratch.org/sympa/arc/lfs-dev/2023-09/msg00036= .html > Signed-off-by: Christian Brauner > --- > =C2=A0fs/fs_context.c=C2=A0=C2=A0=C2=A0=C2=A0 | 18 ++++++++++++++++++ > =C2=A0fs/inode.c=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 |= =C2=A0 4 ++-- > =C2=A0fs/proc_namespace.c |=C2=A0 1 + > =C2=A0fs/stat.c=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2= =A0 |=C2=A0 2 +- > =C2=A0include/linux/fs.h=C2=A0 |=C2=A0 4 +++- > =C2=A05 files changed, 25 insertions(+), 4 deletions(-) >=20 > diff --git a/fs/fs_context.c b/fs/fs_context.c > index a0ad7a0c4680..dd4dade0bb9e 100644 > --- a/fs/fs_context.c > +++ b/fs/fs_context.c > @@ -44,6 +44,7 @@ static const struct constant_table common_set_sb_flag[]= =3D { > =C2=A0 { "mand", SB_MANDLOCK }, > =C2=A0 { "ro", SB_RDONLY }, > =C2=A0 { "sync", SB_SYNCHRONOUS }, > + { "mgtime", SB_MGTIME }, > =C2=A0 { }, > =C2=A0}; > =C2=A0 > @@ -52,18 +53,32 @@ static const struct constant_table common_clear_sb_fl= ag[] =3D { > =C2=A0 { "nolazytime", SB_LAZYTIME }, > =C2=A0 { "nomand", SB_MANDLOCK }, > =C2=A0 { "rw", SB_RDONLY }, > + { "nomgtime", SB_MGTIME }, > =C2=A0 { }, > =C2=A0}; > =C2=A0 > +static inline int check_mgtime(unsigned int token, const struct fs_conte= xt *fc) > +{ > + if (token !=3D SB_MGTIME) > + return 0; > + if (!(fc->fs_type->fs_flags & FS_MGTIME)) > + return invalf(fc, "Filesystem doesn't support multi-grain timestamps")= ; > + return 0; > +} > + > =C2=A0/* > =C2=A0 * Check for a common mount option that manipulates s_flags. > =C2=A0 */ > =C2=A0static int vfs_parse_sb_flag(struct fs_context *fc, const char *key= ) > =C2=A0{ > =C2=A0 unsigned int token; > + int ret; > =C2=A0 > =C2=A0 token =3D lookup_constant(common_set_sb_flag, key, 0); > =C2=A0 if (token) { > + ret =3D check_mgtime(token, fc); > + if (ret) > + return ret; > =C2=A0 fc->sb_flags |=3D token; > =C2=A0 fc->sb_flags_mask |=3D token; > =C2=A0 return 0; > @@ -71,6 +86,9 @@ static int vfs_parse_sb_flag(struct fs_context *fc, con= st char *key) > =C2=A0 > =C2=A0 token =3D lookup_constant(common_clear_sb_flag, key, 0); > =C2=A0 if (token) { > + ret =3D check_mgtime(token, fc); > + if (ret) > + return ret; > =C2=A0 fc->sb_flags &=3D ~token; > =C2=A0 fc->sb_flags_mask |=3D token; > =C2=A0 return 0; > diff --git a/fs/inode.c b/fs/inode.c > index 54237f4242ff..fd1a2390aaa3 100644 > --- a/fs/inode.c > +++ b/fs/inode.c > @@ -2141,7 +2141,7 @@ EXPORT_SYMBOL(current_mgtime); > =C2=A0 > =C2=A0static struct timespec64 current_ctime(struct inode *inode) > =C2=A0{ > - if (is_mgtime(inode)) > + if (IS_MGTIME(inode)) > =C2=A0 return current_mgtime(inode); > =C2=A0 return current_time(inode); > =C2=A0} > @@ -2588,7 +2588,7 @@ struct timespec64 inode_set_ctime_current(struct in= ode *inode) > =C2=A0 now =3D current_time(inode); > =C2=A0 > =C2=A0 /* Just copy it into place if it's not multigrain */ > - if (!is_mgtime(inode)) { > + if (!IS_MGTIME(inode)) { > =C2=A0 inode_set_ctime_to_ts(inode, now); > =C2=A0 return now; > =C2=A0 } > diff --git a/fs/proc_namespace.c b/fs/proc_namespace.c > index 250eb5bf7b52..08f5bf4d2c6c 100644 > --- a/fs/proc_namespace.c > +++ b/fs/proc_namespace.c > @@ -49,6 +49,7 @@ static int show_sb_opts(struct seq_file *m, struct supe= r_block *sb) > =C2=A0 { SB_DIRSYNC, ",dirsync" }, > =C2=A0 { SB_MANDLOCK, ",mand" }, > =C2=A0 { SB_LAZYTIME, ",lazytime" }, > + { SB_MGTIME, ",mgtime" }, > =C2=A0 { 0, NULL } > =C2=A0 }; > =C2=A0 const struct proc_fs_opts *fs_infop; > diff --git a/fs/stat.c b/fs/stat.c > index 6e60389d6a15..2f18dd5de18b 100644 > --- a/fs/stat.c > +++ b/fs/stat.c > @@ -90,7 +90,7 @@ void generic_fillattr(struct mnt_idmap *idmap, u32 requ= est_mask, > =C2=A0 stat->size =3D i_size_read(inode); > =C2=A0 stat->atime =3D inode->i_atime; > =C2=A0 > - if (is_mgtime(inode)) { > + if (IS_MGTIME(inode)) { > =C2=A0 fill_mg_cmtime(stat, request_mask, inode); > =C2=A0 } else { > =C2=A0 stat->mtime =3D inode->i_mtime; > diff --git a/include/linux/fs.h b/include/linux/fs.h > index 4aeb3fa11927..03e415fb3a7c 100644 > --- a/include/linux/fs.h > +++ b/include/linux/fs.h > @@ -1114,6 +1114,7 @@ extern int send_sigurg(struct fown_struct *fown); > =C2=A0#define SB_NODEV=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 BIT(2) /= * Disallow access to device special files */ > =C2=A0#define SB_NOEXEC=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 BIT(3) /* Dis= allow program execution */ > =C2=A0#define SB_SYNCHRONOUS=C2=A0 BIT(4) /* Writes are synced at once */ > +#define SB_MGTIME BIT(5) /* Use multi-grain timestamps */ > =C2=A0#define SB_MANDLOCK=C2=A0=C2=A0=C2=A0=C2=A0 BIT(6) /* Allow mandato= ry locks on an FS */ > =C2=A0#define SB_DIRSYNC=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 BIT(7) /* Director= y modifications are synchronous */ > =C2=A0#define SB_NOATIME=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 BIT(10) /* Do not = update access times. */ > @@ -2105,6 +2106,7 @@ static inline bool sb_rdonly(const struct super_blo= ck *sb) { return sb->s_flags > =C2=A0 ((inode)->i_flags & (S_SYNC|S_DIRSYNC))) > =C2=A0#define IS_MANDLOCK(inode) __IS_FLG(inode, SB_MANDLOCK) > =C2=A0#define IS_NOATIME(inode) __IS_FLG(inode, SB_RDONLY|SB_NOATIME) > +#define IS_MGTIME(inode) __IS_FLG(inode, SB_MGTIME) > =C2=A0#define IS_I_VERSION(inode) __IS_FLG(inode, SB_I_VERSION) > =C2=A0 > =C2=A0#define IS_NOQUOTA(inode) ((inode)->i_flags & S_NOQUOTA) > @@ -2366,7 +2368,7 @@ struct file_system_type { > =C2=A0 */ > =C2=A0static inline bool is_mgtime(const struct inode *inode) > =C2=A0{ > - return inode->i_sb->s_type->fs_flags & FS_MGTIME; > + return inode->i_sb->s_flags & SB_MGTIME; > =C2=A0} > =C2=A0 > =C2=A0extern struct dentry *mount_bdev(struct file_system_type *fs_type,