Commits · bbd6851a3213a525128473e978b692ab6ac11aba · Kirill Smelkov / linux

12 Jun, 2009 40 commits

Push lock_super() into the ->remount_fs() of filesystems that care about it · bbd6851a

Al Viro authored May 06, 2009

Note that since we can't run into contention between remount_fs and write_super
(due to exclusion on s_umount), we have to care only about filesystems that
touch lock_super() on their own.  Out of those ext3, ext4, hpfs, sysv and ufs
do need it; fat doesn't since its ->remount_fs() only accesses assign-once
data (basically, it's "we have no atime on directories and only have atime on
files for vfat; force nodiratime and possibly noatime into *flags").

[folded a build fix from hch]
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

bbd6851a

push BKL down into ->put_super · 6cfd0148

Christoph Hellwig authored May 05, 2009

Move BKL into ->put_super from the only caller.  A couple of
filesystems had trivial enough ->put_super (only kfree and NULLing of
s_fs_info + stuff in there) to not get any locking: coda, cramfs, efs,
hugetlbfs, omfs, qnx4, shmem, all others got the full treatment.  Most
of them probably don't need it, but I'd rather sort that out individually.
Preferably after all the other BKL pushdowns in that area.

[AV: original used to move lock_super() down as well; these changes are
removed since we don't do lock_super() at all in generic_shutdown_super()
now]
[AV: fuse, btrfs and xfs are known to need no damn BKL, exempt]
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

6cfd0148

No need to do lock_super() for exclusion in generic_shutdown_super() · a9e220f8

Al Viro authored May 05, 2009

We can't run into contention on it.  All other callers of lock_super()
either hold s_umount (and we have it exclusive) or hold an active
reference to superblock in question, which prevents the call of
generic_shutdown_super() while the reference is held.  So we can
replace lock_super(s) with get_fs_excl() in generic_shutdown_super()
(and corresponding change for unlock_super(), of course).

Since ext4 expects s_lock held for its put_super, take lock_super()
into it.  The rest of filesystems do not care at all.
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

a9e220f8

Trim a bit of crap from fs.h · 62c6943b

Al Viro authored May 07, 2009

do_remount_sb() is fs/internal.h fodder, fsync_no_super() is long gone.
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

62c6943b

Make sure that all callers of remount hold s_umount exclusive · 443b94ba
Al Viro authored May 05, 2009
```
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
```
443b94ba

enforce ->sync_fs is only called for rw superblock · 5af7926f

Christoph Hellwig authored May 05, 2009

Make sure a superblock really is writeable by checking MS_RDONLY
under s_umount.  sync_filesystems needed some re-arragement for
that, but all but one sync_filesystem caller had the correct locking
already so that we could add that check there.  cachefiles grew
s_umount locking.

I've also added a WARN_ON to sync_filesystem to assert this for
future callers.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

5af7926f

cleanup sync_supers · e5004753

Christoph Hellwig authored May 05, 2009

Merge the write_super helper into sync_super and move the check for
->write_super earlier so that we can avoid grabbing a reference to
a superblock that doesn't have it.

While we're at it also add a little comment documenting sync_supers.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

e5004753

dcache: extrace and use d_unlinked() · f3da392e

Alexey Dobriyan authored May 04, 2009

d_unlinked() will be used in middle-term to ban checkpointing when opened
but unlinked file is detected, and in long term, to detect such situation
and special case on it.
Signed-off-by: Alexey Dobriyan <adobriyan@gmail.com>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

f3da392e

remove ->write_super call in generic_shutdown_super · 8c85e125

Christoph Hellwig authored Apr 28, 2009

We just did a full fs writeout using sync_filesystem before, and if
that's not enough for the filesystem it can perform it's own writeout
in ->put_super, which many filesystems already do.

Move a call to foofs_write_super into every foofs_put_super for now to
guarantee identical behaviour until it's cleaned up by the individual
filesystem maintainers.

Exceptions:

 - affs already has identical copy & pasted code at the beginning of
   affs_put_super so no need to do it twice.
 - xfs does the right thing without it and I have changes pending for
   the xfs tree touching this are so I don't really need conflicts
   here..
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

8c85e125

qnx4: remove ->write_super · 517bfae2

Christoph Hellwig authored Apr 27, 2009

Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

517bfae2

ocfs2: remove ->write_super and stop maintaining ->s_dirt · 94cb993f

Christoph Hellwig authored Apr 27, 2009

Signed-off-by: Christoph Hellwig <hch@lst.de>
Acked-by: Joel Becker <joel.becker@oracle.com>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

94cb993f

gfs2: remove ->write_super and stop maintaining ->s_dirt · b7d245de

Christoph Hellwig authored Apr 27, 2009

Signed-off-by: Christoph Hellwig <hch@lst.de>
Acked-by: Steven Whitehouse <swhiteho@redhat.com>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

b7d245de

ext3: remove ->write_super and stop maintaining ->s_dirt · ca41f7b9
Christoph Hellwig authored Apr 27, 2009
```
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
```
ca41f7b9
btrfs: remove ->write_super and stop maintaining ->s_dirt · 59d697b7
Christoph Hellwig authored Apr 27, 2009
```
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
```
59d697b7

quota: Introduce writeout_quota_sb() (version 4) · c3f8a40c

Jan Kara authored Apr 27, 2009

Introduce this function which just writes all the quota structures but
avoids all the syncing and cache pruning work to expose quota structures
to userspace. Use this function from __sync_filesystem when wait == 0.
Signed-off-by: Jan Kara <jack@suse.cz>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

c3f8a40c

quota: cleanup dquota sync functions (version 4) · 850b201b

Christoph Hellwig authored Apr 27, 2009

Currently the VFS calls vfs_dq_sync to sync out disk quotas for a given
superblock.  This is a small wrapper around sync_dquots which for the
case of a non-NULL superblock is a small wrapper around quota_sync_sb.

Just make quota_sync_sb global (rename it to sync_quota_sb) and call it
directly.  Also call it directly for those cases in quota.c that have a
superblock and leave sync_dquots purely an iterator over sync_quota_sb and
remove it's superblock argument.

To make this nicer move the check for the lack of a quota_sync method
from the callers into sync_quota_sb.

[folded build fix from Alexander Beregalov <a.beregalov@gmail.com>]
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jan Kara <jack@suse.cz>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

850b201b

vfs: Rename fsync_super() to sync_filesystem() (version 4) · 60b0680f

Jan Kara authored Apr 27, 2009

Rename the function so that it better describe what it really does. Also
remove the unnecessary include of buffer_head.h.
Signed-off-by: Jan Kara <jack@suse.cz>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

60b0680f

vfs: Move syncing code from super.c to sync.c (version 4) · c15c54f5

Jan Kara authored Apr 27, 2009

Move sync_filesystems(), __fsync_super(), fsync_super() from
super.c to sync.c where it fits better.

[build fixes folded]
Signed-off-by: Jan Kara <jack@suse.cz>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

c15c54f5

vfs: Make sys_sync() use fsync_super() (version 4) · 5cee5815

Jan Kara authored Apr 27, 2009

It is unnecessarily fragile to have two places (fsync_super() and do_sync())
doing data integrity sync of the filesystem. Alter __fsync_super() to
accommodate needs of both callers and use it. So after this patch
__fsync_super() is the only place where we gather all the calls needed to
properly send all data on a filesystem to disk.

Nice bonus is that we get a complete livelock avoidance and write_supers()
is now only used for periodic writeback of superblocks.

sync_blockdevs() introduced a couple of patches ago is gone now.

[build fixes folded]
Signed-off-by: Jan Kara <jack@suse.cz>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

5cee5815

vfs: Make __fsync_super() a static function (version 4) · 429479f0

Jan Kara authored Apr 27, 2009

__fsync_super() does the same thing as fsync_super(). So change the only
caller to use fsync_super() and make __fsync_super() static. This removes
unnecessarily duplicated call to sync_blockdev() and prepares ground
for the changes to __fsync_super() in the following patches.
Signed-off-by: Jan Kara <jack@suse.cz>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

429479f0

vfs: Call ->sync_fs() even if s_dirt is 0 (version 4) · bfe88125

Jan Kara authored Apr 27, 2009

sync_filesystems() has a condition that if wait == 0 and s_dirt == 0, then
->sync_fs() isn't called. This does not really make much sence since s_dirt is
generally used by a filesystem to mean that ->write_super() needs to be called.
But ->sync_fs() does different things. I even suspect that some filesystems
(btrfs?) sets s_dirt just to fool this logic.
Signed-off-by: Jan Kara <jack@suse.cz>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

bfe88125

vfs: Fix sys_sync() and fsync_super() reliability (version 4) · 5a3e5cb8

Jan Kara authored Apr 27, 2009

So far, do_sync() called:
  sync_inodes(0);
  sync_supers();
  sync_filesystems(0);
  sync_filesystems(1);
  sync_inodes(1);

This ordering makes it kind of hard for filesystems as sync_inodes(0) need not
submit all the IO (for example it skips inodes with I_SYNC set) so e.g. forcing
transaction to disk in ->sync_fs() is not really enough. Therefore sys_sync has
not been completely reliable on some filesystems (ext3, ext4, reiserfs, ocfs2
and others are hit by this) when racing e.g. with background writeback. A
similar problem hits also other filesystems (e.g. ext2) because of
write_supers() being called before the sync_inodes(1).

Change the ordering of calls in do_sync() - this requires a new function
sync_blockdevs() to preserve the property that block devices are always synced
after write_super() / sync_fs() call.

The same issue is fixed in __fsync_super() function used on umount /
remount read-only.

[AV: build fixes]
Signed-off-by: Jan Kara <jack@suse.cz>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

5a3e5cb8

remove s_async_list · 876a9f76

Christoph Hellwig authored Apr 28, 2009

Remove the unused s_async_list in the superblock, a leftover of the
broken async inode deletion code that leaked into mainline.  Having this
in the middle of the sync/unmount path is not helpful for the following
cleanups.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

876a9f76

fs: move mark_files_ro into file_table.c · 864d7c4c

npiggin@suse.de authored Apr 26, 2009

This function walks the s_files lock, and operates primarily on the
files in a superblock, so it better belongs here (eg. see also
fs_may_remount_ro).

[AV: ... and it shouldn't be static after that move]
Signed-off-by: Nick Piggin <npiggin@suse.de>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

864d7c4c

fs: introduce mnt_clone_write · 96029c4e

npiggin@suse.de authored Apr 26, 2009

This patch speeds up lmbench lat_mmap test by about another 2% after the
first patch.

Before:
 avg = 462.286
 std = 5.46106

After:
 avg = 453.12
 std = 9.58257

(50 runs of each, stddev gives a reasonable confidence)

It does this by introducing mnt_clone_write, which avoids some heavyweight
operations of mnt_want_write if called on a vfsmount which we know already
has a write count; and mnt_want_write_file, which can call mnt_clone_write
if the file is open for write.

After these two patches, mnt_want_write and mnt_drop_write go from 7% on
the profile down to 1.3% (including mnt_clone_write).

[AV: mnt_want_write_file() should take file alone and derive mnt from it;
not only all callers have that form, but that's the only mnt about which
we know that it's already held for write if file is opened for write]

Cc: Dave Hansen <haveblue@us.ibm.com>
Signed-off-by: Nick Piggin <npiggin@suse.de>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

96029c4e

fs: mnt_want_write speedup · d3ef3d73

npiggin@suse.de authored Apr 26, 2009

This patch speeds up lmbench lat_mmap test by about 8%. lat_mmap is set up
basically to mmap a 64MB file on tmpfs, fault in its pages, then unmap it.
A microbenchmark yes, but it exercises some important paths in the mm.

Before:
 avg = 501.9
 std = 14.7773

After:
 avg = 462.286
 std = 5.46106

(50 runs of each, stddev gives a reasonable confidence, but there is quite
a bit of variation there still)

It does this by removing the complex per-cpu locking and counter-cache and
replaces it with a percpu counter in struct vfsmount. This makes the code
much simpler, and avoids spinlocks (although the msync is still pretty
costly, unfortunately). It results in about 900 bytes smaller code too. It
does increase the size of a vfsmount, however.

It should also give a speedup on large systems if CPUs are frequently operating
on different mounts (because the existing scheme has to operate on an atomic in
the struct vfsmount when switching between mounts). But I'm most interested in
the single threaded path performance for the moment.

[AV: minor cleanup]

Cc: Dave Hansen <haveblue@us.ibm.com>
Signed-off-by: Nick Piggin <npiggin@suse.de>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

d3ef3d73

Move junk from proc_fs.h to fs/proc/internal.h · 3174c21b
Al Viro authored Apr 07, 2009
```
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
```
3174c21b
switch lookup_mnt() · 1c755af4
Al Viro authored Apr 18, 2009
```
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
```
1c755af4
switch follow_mount() · 79ed0226
Al Viro authored Apr 18, 2009
```
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
```
79ed0226
switch follow_down() · 9393bd07
Al Viro authored Apr 18, 2009
```
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
```
9393bd07
Switch collect_mounts() to struct path · 589ff870
Al Viro authored Apr 18, 2009
```
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
```
589ff870
switch follow_up() to struct path · bab77ebf
Al Viro authored Apr 18, 2009
```
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
```
bab77ebf
switch rqst_exp_parent() · e64c390c
Al Viro authored Apr 18, 2009
```
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
```
e64c390c
switch rqst_exp_get_by_name() · 91c9fa8f
Al Viro authored Apr 18, 2009
```
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
```
91c9fa8f

switch exp_parent() to struct path · 5bf3bd2b

Al Viro authored Apr 18, 2009

... and lose the always-NULL last argument (non-NULL case had been
split off a while ago).
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

5bf3bd2b

nfsd struct path use: exp_get_by_name() · 55430e2e
Al Viro authored Apr 18, 2009
```
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
```
55430e2e

Don't bother with check_mnt() in do_add_mount() on shrinkable ones · dd5cae6e

Al Viro authored Apr 07, 2009

These guys are what we add as submounts; checks for "is that attached in
our namespace" are simply irrelevant for those and counterproductive for
use of private vfsmount trees a-la what NFS folks want.
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

dd5cae6e

Make vfs_path_lookup() use starting point as root · 5b857119
Al Viro authored Apr 07, 2009
```
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
```
5b857119

Cache root in nameidata · 2a737871

Al Viro authored Apr 07, 2009

New field: nd->root. When pathname resolution wants to know the root,
check if nd->root.mnt is non-NULL; use nd->root if it is, otherwise
copy current->fs->root there. After path_walk() is finished, we check
if we'd got a cached value in nd->root and drop it. Before calling
path_walk() we should either set nd->root.mnt to NULL *or* copy (and
pin down) some path to nd->root. In the latter case we won't be
looking at current->fs->root at all.
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

2a737871

Preparations to caching root in path_walk() · 9b4a9b14

Al Viro authored Apr 07, 2009

Split do_path_lookup(), opencode the call from do_filp_open()
do_filp_open() is the only caller of do_path_lookup() that
cares about root afterwards (it keeps resolving symlinks on
O_CREAT path after it'd done LOOKUP_PARENT walk).  So when
we start caching fs->root in path_walk(), it'll need a different
treatment.
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>

9b4a9b14