Commits · 671415b7db49f62896f0b6d50fc4f312a0512983 · nexedi / linux

25 Oct, 2012 7 commits

Btrfs: fix deadlock caused by the nested chunk allocation · 671415b7

Miao Xie authored Oct 16, 2012

Steps to reproduce:
 # mkfs.btrfs -m raid1 <disk1> <disk2>
 # btrfstune -S 1 <disk1>
 # mount <disk1> <mnt>
 # btrfs device add <disk3> <disk4> <mnt>
 # mount -o remount,rw <mnt>
 # dd if=/dev/zero of=<mnt>/tmpfile bs=1M count=1
 Deadlock happened.

It is because of the nested chunk allocation. When we wrote the data
into the filesystem, we would allocate the data chunk because there was
no data chunk in the filesystem. At the end of the data chunk allocation,
we should insert the metadata of the data chunk into the extent tree, but
there was no raid1 chunk, so we tried to lock the chunk allocation mutex to
allocate the new chunk, but we had held the mutex, the deadlock happened.

By rights, we would allocate the raid1 chunk when we added the second device
because the profile of the seed filesystem is raid1 and we had two devices.
But we didn't do that in fact. It is because the last step of the first device
insertion didn't commit the transaction. So when we added the second device,
we didn't cow the tree, and just inserted the relative metadata into the leaves
which were generated by the first device insertion, and its profile was dup.

So, I fix this problem by commiting the transaction at the end of the first
device insertion.
Signed-off-by: Miao Xie <miaox@cn.fujitsu.com>

671415b7

btrfs: Return EINVAL when length to trim is less than FSB · e515c18b

Lukas Czerner authored Oct 16, 2012

Currently if len argument in btrfs_ioctl_fitrim() is smaller than
one FSB we will continue and finally return 0 bytes discarded.
However if the length to discard is smaller then file system block
we should really return EINVAL.
Signed-off-by: Lukas Czerner <lczerner@redhat.com>

e515c18b

Btrfs: fix memory leak in btrfs_quota_enable() · 5b7ff5b3

Tsutomu Itoh authored Oct 16, 2012

We should free quota_root before returning from the error
handling code.
Signed-off-by: Tsutomu Itoh <t-itoh@jp.fujitsu.com>

5b7ff5b3

Btrfs: send correct rdev and mode in btrfs-send · d79e5043

Arne Jansen authored Oct 15, 2012

When sending a device file, the stream was missing the mode. Also the
rdev was encoded wrongly.
Signed-off-by: Arne Jansen <sensille@gmx.net>

d79e5043

Btrfs: extended inode refs support for send mechanism · 96b5bd77

Jan Schmidt authored Oct 15, 2012

This adds support for the new extended inode refs to btrfs send.
Signed-off-by: Jan Schmidt <list.btrfs@jan-o-sch.net>

96b5bd77

Btrfs: Fix wrong error handling code · 84167d19

Stefan Behrens authored Oct 11, 2012

gcc says "warning: comparison of unsigned expression >= 0 is always
true" because i is an unsigned long. And gcc is right this time.
Signed-off-by: Stefan Behrens <sbehrens@giantdisaster.de>

84167d19

Fix a sign bug causing invalid memory access in the ino_paths ioctl. · 661bec6b

Gabriel de Perthuis authored Oct 10, 2012

To see the problem, create many hardlinks to the same file (120 should do it),
then look up paths by inode with:

  ls -i
  btrfs inspect inode-resolve -v $ino /mnt/btrfs

I noticed the memory layout of the fspath->val data had some irregularities
(some unnecessary gaps that stop appearing about halfway),
so I'm not sure there aren't any bugs left in it.

661bec6b

09 Oct, 2012 33 commits

btrfs: init ref_index to zero in add_inode_ref · f46dbe3d
Chris Mason authored Oct 09, 2012
```
Signed-off-by: Chris Mason <chris.mason@fusionio.com>
```
f46dbe3d

Btrfs: remove repeated eb->pages check in, disk-io.c/csum_dirty_buffer · 1037a5af

Wang Sheng-Hui authored Oct 08, 2012

In csum_dirty_buffer, we first get eb from page->private.
Then we check if the page is the first page of eb. Later
we check it again. Remove the repeated check here.
Signed-off-by: Wang Sheng-Hui <shhuiw@gmail.com>

1037a5af

Btrfs: fix page leakage · f60b1b49

Josef Bacik authored Oct 05, 2012

Alloc_dummy_extent_buffer will not free the first page in the eb array if we
fail to allocate a page, fix this.  Thanks,
Reported-by: David Sterba <dave@jikos.cz>
Signed-off-by: Josef Bacik <jbacik@fusionio.com>

f60b1b49

Btrfs: do not warn_on when we cannot alloc a page for an extent buffer · 4804b382

Josef Bacik authored Oct 05, 2012

It's just annoying and the user will have gotten a nice OOM killer message
so they are already fully aware they are screwed :).  Thanks,
Reported-by: Jérôme Poulin <jeromepoulin@gmail.com>
Signed-off-by: Josef Bacik <jbacik@fusionio.com>

4804b382

Btrfs: don't bug on enomem in readpage · edd33c99

Josef Bacik authored Oct 05, 2012

Get rid of the BUG_ON(ret == -ENOMEM) in __extent_read_full_page.  Thanks,
Reported-by: Jérôme Poulin <jeromepoulin@gmail.com>
Signed-off-by: Josef Bacik <jbacik@fusionio.com>

edd33c99

Btrfs: cleanup pages properly when ENOMEM in compression · 15e3004a

Josef Bacik authored Oct 05, 2012

We were freeing non-existent pages which was causing a panic for a user who
was suffering from ENOMEM.  This patch fixes the problem.  Thanks,
Reported-by: Jérôme Poulin <jeromepoulin@gmail.com>
Signed-off-by: Josef Bacik <jbacik@fusionio.com>

15e3004a

Btrfs: make filesystem read-only when submitting barrier fails · 5af3e8cc

Stefan Behrens authored Aug 01, 2012

So far the return code of barrier_all_devices() is ignored, which
means that errors are ignored. The result can be a corrupt
filesystem which is not consistent.
This commit adds code to evaluate the return code of
barrier_all_devices(). The normal btrfs_error() mechanism is used to
switch the filesystem into read-only mode when errors are detected.

In order to decide whether barrier_all_devices() should return
error or success, the number of disks that are allowed to fail the
barrier submission is calculated. This calculation accounts for the
worst RAID level of metadata, system and data. If single, dup or
RAID0 is in use, a single disk error is already considered to be
fatal. Otherwise a single disk error is tolerated.

The calculation of the number of disks that are tolerated to fail
the barrier operation is performed when the filesystem gets mounted,
when a balance operation is started and finished, and when devices
are added or removed.
Signed-off-by: Stefan Behrens <sbehrens@giantdisaster.de>

5af3e8cc

Btrfs: detect corrupted filesystem after write I/O errors · 62856a9b

Stefan Behrens authored Jul 31, 2012

In check-integrity, detect when a superblock is written that points
to blocks that have not been written to disk due to I/O write errors.
Signed-off-by: Stefan Behrens <sbehrens@giantdisaster.de>

62856a9b

Btrfs: make compress and nodatacow mount options mutually exclusive · bedb2cca

Andrei Popa authored Sep 20, 2012

If a filesystem is mounted with compression and then remounted by adding nodatacow,
the compression is disabled but the compress flag is still visible.
Also, if a filesystem is mounted with nodatacow and then remounted with compression,
nodatacow flag is still present but it's not active.
This patch:
- removes compress flags and notifies that the compression has been disabled if the
  filesystem is mounted with nodatacow
- removes nodatacow and nodatasum flags if mounted with compress.
Signed-off-by: Andrei Popa <andrei.popa@i-neo.ro>

bedb2cca

btrfs: fix message printing · 48940662

Daniel J Blueman authored May 07, 2012

Fix various messages to include newline and module prefix.
Signed-off-by: Daniel J Blueman <daniel@quora.org>

48940662

Btrfs: don't bother committing delayed inode updates when fsyncing · 94edf4ae

Josef Bacik authored Sep 25, 2012

We can just copy the in memory inode into the tree log directly, no sense in
updating the fs tree so we can copy it into the tree log tree. Thanks,
Signed-off-by: Josef Bacik <jbacik@fusionio.com>

94edf4ae

btrfs: move inline function code to header file · 479ed9ab

Robin Dong authored Sep 29, 2012

When building btrfs from kernel code, it will report:

fs/btrfs/extent_io.h:281: warning: 'extent_buffer_page' declared inline after being called
fs/btrfs/extent_io.h:281: warning: previous declaration of 'extent_buffer_page' was here
fs/btrfs/extent_io.h:280: warning: 'num_extent_pages' declared inline after being called
fs/btrfs/extent_io.h:280: warning: previous declaration of 'num_extent_pages' was here

because of the wrong declaration of inline functions.
Signed-off-by: Robin Dong <sanbai@taobao.com>

479ed9ab

Btrfs: remove unnecessary IS_ERR in bio_readpage_error() · 7a2d6a64

Tsutomu Itoh authored Oct 01, 2012

Because the value of extent_map is only a correct value or NULL,
so IS_ERR is unnecessary.
Signed-off-by: Tsutomu Itoh <t-itoh@jp.fujitsu.com>

7a2d6a64

btrfs: remove unused function btrfs_insert_some_items() · 8d1a1317

Robin Dong authored Sep 29, 2012

The function btrfs_insert_some_items() would not be called by any other functions,
so remove it.
Signed-off-by: Robin Dong <sanbai@taobao.com>

8d1a1317

Btrfs: don't commit instead of overcommitting · 44734ed1

Josef Bacik authored Sep 28, 2012

I don't think we have the same problem that this was supposed to fix
originally since we can allocate chunks in the enospc path now. This code
is causing us to constantly commit the transaction as we get close to using
all of our available space in our currently allocated chunks, instead of
allocating another chunk and carrying on with life, which is not nice for
performance. Thanks,
Signed-off-by: Josef Bacik <jbacik@fusionio.com>

44734ed1

Btrfs: confirmation of value is added before trace_btrfs_get_extent() is called · f0bd95ea

Tsutomu Itoh authored Oct 01, 2012

We should confirm the value of extent_map before calling
trace_btrfs_get_extent() because the value of extent_map has the
possibility of NULL.
Signed-off-by: Tsutomu Itoh <t-itoh@jp.fujitsu.com>

f0bd95ea

Btrfs: be smarter about dropping things from the tree log · 18ec90d6

Josef Bacik authored Sep 28, 2012

When we truncate existing items in the tree log we've been searching for
each individual item and removing them. This is unnecessary churn and
searching, just keep track of the slot we are on and how many items we need
to delete and delete them all at once. Thanks,
Signed-off-by: Josef Bacik <jbacik@fusionio.com>

18ec90d6

Btrfs: don't lookup csums for prealloc extents · 6f1fed77

Josef Bacik authored Sep 26, 2012

The tree logging stuff was looking up csums to copy over for prealloc
extents which is just work we don't need to be doing.  Thanks,
Signed-off-by: Josef Bacik <jbacik@fusionio.com>

6f1fed77

Btrfs: cache extent state when writing out dirty metadata pages · e6138876

Josef Bacik authored Sep 27, 2012

Everytime we write out dirty pages we search for an offset in the tree,
convert the bits in the state, and then when we wait we search for the
offset again and clear the bits. So for every dirty range in the io tree we
are doing 4 rb searches, which is suboptimal. With this patch we are only
doing 2 searches for every cycle (modulo weird things happening). Thanks,
Signed-off-by: Josef Bacik <jbacik@fusionio.com>

e6138876

Btrfs: do not hold the file extent leaf locked when adding extent item · ce195332

Josef Bacik authored Sep 25, 2012

For some reason we unlock everything except the leaf we are on, set the path
blocking and then add the extent item for the extent we just finished
writing. I can't for the life of me figure out why we would want to do
this, and the history doesn't really indicate that there was a real reason
for it, so just remove it. This will reduce our tree lock contention on
heavy writes. Thanks,
Signed-off-by: Josef Bacik <jbacik@fusionio.com>

ce195332

Btrfs: do not async metadata csumming in certain situations · de0022b9

Josef Bacik authored Sep 25, 2012

There are a coule scenarios where farming metadata csumming off to an async
thread doesn't help. The first is if our processor supports crc32c, in
which case the csumming will be fast and so the overhead of the async model
is not worth the cost. The other case is for our tree log. We will be
making that stuff dirty and writing it out and waiting for it immediately.
Even with software crc32c this gives me a ~15% increase in speed with O_SYNC
workloads. Thanks,
Signed-off-by: Josef Bacik <jbacik@fusionio.com>

de0022b9

btrfs: fix min csum item size warnings in 32bit · 221b8318

Zach Brown authored Sep 20, 2012

commit 7ca4be45 limited csum items to
PAGE_CACHE_SIZE.  It used min() with incompatible types in 32bit which
generates warnings:

fs/btrfs/file-item.c: In function ‘btrfs_csum_file_blocks’:
fs/btrfs/file-item.c:717: warning: comparison of distinct pointer types lacks a cast

This uses min_t(u32,) to fix the warnings.  u32 seemed reasonable
because btrfs_root->leafsize is u32 and PAGE_CACHE_SIZE is unsigned
long.
Signed-off-by: Zach Brown <zab@zabbo.net>

221b8318

Btrfs: run delayed refs first when out of space · 67b0fd63

Josef Bacik authored Sep 24, 2012

Running delayed refs is faster than running delalloc, so lets do that first
to try and reclaim space.  This makes my fs_mark test about 20% faster.
Thanks,
Signed-off-by: Josef Bacik <jbacik@fusionio.com>

67b0fd63

Btrfs: fix orphan transaction on the freezed filesystem · 354aa0fb

Miao Xie authored Sep 20, 2012

With the following debug patch:

 static int btrfs_freeze(struct super_block *sb)
 {
+ 	struct btrfs_fs_info *fs_info = btrfs_sb(sb);
+	struct btrfs_transaction *trans;
+
+	spin_lock(&fs_info->trans_lock);
+	trans = fs_info->running_transaction;
+	if (trans) {
+		printk("Transid %llu, use_count %d, num_writer %d\n",
+			trans->transid, atomic_read(&trans->use_count),
+			atomic_read(&trans->num_writers));
+	}
+	spin_unlock(&fs_info->trans_lock);
 	return 0;
 }

I found there was a orphan transaction after the freeze operation was done.

It is because the transaction may not be committed when the transaction handle
end even though it is the last handle of the current transaction. This design
avoid committing the transaction frequently, but also introduce the above
problem.

So I add btrfs_attach_transaction() which can catch the current transaction
and commit it. If there is no transaction, it will return ENOENT, and do not
anything.

This function also can be used to instead of btrfs_join_transaction_freeze()
because it don't increase the writer counter and don't start a new transaction,
so it also can fix the deadlock between sync and freeze.

Besides that, it is used to instead of btrfs_join_transaction() in
transaction_kthread(), because if there is no transaction, the transaction
kthread needn't anything.
Signed-off-by: Miao Xie <miaox@cn.fujitsu.com>

354aa0fb

Btrfs: add a type field for the transaction handle · a698d075

Miao Xie authored Sep 20, 2012

This patch add a type field into the transaction handle structure,
in this way, we needn't implement various end-transaction functions
and can make the code more simple and readable.
Signed-off-by: Miao Xie <miaox@cn.fujitsu.com>

a698d075

Btrfs: fix memory leak in start_transaction() · e8830e60

Miao Xie authored Sep 19, 2012

This patch fixes memory leak of the transaction handle which happened
when starting transaction failed on a freezed fs.
Signed-off-by: Miao Xie <miaox@cn.fujitsu.com>

e8830e60

btrfs: extended inode ref iteration · d24bec3a

Mark Fasheh authored Aug 08, 2012

The iterate_irefs in backref.c is used to build path components from inode
refs. This patch adds code to iterate extended refs as well.

I had modify the callback function signature to abstract out some of the
differences between ref structures. iref_to_path() also needed similar
changes.
Signed-off-by: Mark Fasheh <mfasheh@suse.de>

d24bec3a

btrfs: extended inode refs · f186373f

Mark Fasheh authored Aug 08, 2012

This patch adds basic support for extended inode refs. This includes support
for link and unlink of the refs, which basically gets us support for rename
as well.

Inode creation does not need changing - extended refs are only added after
the ref array is full.
Signed-off-by: Mark Fasheh <mfasheh@suse.de>

f186373f

btrfs: improved readablity for add_inode_ref · 5a1d7843

Jan Schmidt authored Aug 17, 2012

Moved part of the code into a sub function and replaced most of the gotos
by ifs, hoping that it will be easier to read now.
Signed-off-by: Jan Schmidt <list.btrfs@jan-o-sch.net>
Signed-off-by: Mark Fasheh <mfasheh@suse.de>

5a1d7843

Btrfs: handle not finding the extent exactly when logging changed extents · 0aa4a17d

Josef Bacik authored Sep 19, 2012

I started hitting warnings when running xfstest 68 in a loop because there
were EM's that were not lined up properly with the physical extents. This
is ok, if we do something like punch a hole or write to a preallocated space
or something like that we can have an EM that doesn't cover the entire
physical extent. So fix the tree logging stuff to cope with this case so we
don't just commit the transaction. With this patch I no longer see the
warnings from the tree logging code. Thanks,
Signed-off-by: Josef Bacik <jbacik@fusionio.com>

0aa4a17d

btrfs: move transaction aborts to the point of failure · 005d6427

David Sterba authored Sep 18, 2012

Call btrfs_abort_transaction as early as possible when an error
condition is detected, that way the line number reported is useful
and we're not clueless anymore which error path led to the abort.
Signed-off-by: David Sterba <dsterba@suse.cz>

005d6427

Btrfs: fix the missing error information in create_pending_snapshot() · 8732d44f

Miao Xie authored Sep 17, 2012

The macro btrfs_abort_transaction() can get the line number of the code
where the problem happens, so we should invoke it in the place that the
error occurs, or we will lose the line number.
Reported-by: David Sterba <dave@jikos.cz>
Signed-off-by: Miao Xie <miaox@cn.fujitsu.com>

8732d44f

Btrfs: fix off-by-one in file clone · aa42ffd9

Liu Bo authored Sep 18, 2012

Btrfs uses inclusive range end for lock_extent(), unlock_extent() and
related functions, so we made off-by-one errors in file clone.

This fixes it and also fixes some style problems.
Signed-off-by: Liu Bo <bo.li.liu@oracle.com>

aa42ffd9