Commits · 537c49d6afadb4be54be03c9a8cb1f1ade07b104 · Kirill Smelkov / linux

22 Oct, 2023 40 commits

bcachefs: Fix btree node merge -> split operations · 537c49d6

Kent Overstreet authored Dec 11, 2020

If a btree node merger is followed by a split or compact of the parent
node, we could end up with the parent btree node iterator pointing to
the whiteout inserted by the btree node merge operation - the fix is to
ensure that interior btree node iterators always point to the first non
whiteout.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

537c49d6

bcachefs: Always check if we need disk res in extent update path · 5b9bf43c

Kent Overstreet authored Dec 10, 2020

With erasure coding, we now have processes in the background that
compact data, causing it to take up less space on disk than when it was
written, or potentially when it was read.

This means that we can't trust the page cache when it says "we have data
on disk taking up x amount of space here" - there's always the potential
to race with background compaction.

To fix this, just check if we need to add to our disk reservation in the
bch2_extent_update() path, in the transaction that will do the btree
update.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

5b9bf43c

bcachefs: Update transactional triggers interface to pass old & new keys · 719fe7fb

Kent Overstreet authored Dec 10, 2020

This is needed to fix a bug where we're overflowing iterators within a
btree transaction, because we're updating the stripes btree (to update
block counts) and the stripes btree trigger is unnecessarily updating
the alloc btree - it doesn't need to update the alloc btree when the
pointers within a stripe aren't changing.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

719fe7fb

bcachefs: Only try to get existing stripe once in stripe create path · 66bddc6c

Kent Overstreet authored Dec 09, 2020

The stripe creation path was too state-machiney: it would always run the
full state machine until it had succesfully created a new stripe.

But if we tried to get and reuse an existing stripe after we'd already
allocated some buckets, the buckets we'd allocated might have conflicted
with the blocks in the existing stripe we need to keep - oops.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

66bddc6c

bcachefs: Fix __btree_iter_next() when all iters are in use_next() when all iters are in use · cc578a36

Kent Overstreet authored Dec 09, 2020

Also, print out more information on btree transaction iterator overflow.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

cc578a36

bcachefs: Fix rand_delete() test · d5b98fe2

Kent Overstreet authored Dec 07, 2020

When we didn't find a key to delete we were getting a null ptr deref.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

d5b98fe2

bcachefs: Try to print full btree error message · a2bfc841

Kent Overstreet authored Dec 06, 2020

Metadata corruption bugs are hard to debug if we can't see exactly what
went wrong - try to allocate a bigger buffer so we can print out
everything we have.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

a2bfc841

bcachefs: Prevent journal reclaim from spinning · b18df768

Kent Overstreet authored Dec 06, 2020

Without checking if we actually flushed anything, journal reclaim could
still go into an infinite loop while trying ot shut down.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

b18df768

bcachefs: Fix btree key cache dirty checks · f51e84fe

Kent Overstreet authored Dec 05, 2020

Had a type that meant we were triggering journal reclaim _much_ more
aggressively than needed. Also, fix a potential integer overflow.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

f51e84fe

bcachefs: Be more conservation about journal pre-reservations · 5d32c5bb

Kent Overstreet authored Dec 05, 2020

 - Try to always keep 1/8th of the journal free, on top of
   pre-reservations
 - Move the check for whether the journal is stuck to
   bch2_journal_space_available, and make it only fire when there aren't
   any journal writes in flight (that might free up space by updating
   last_seq)
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

5d32c5bb

bcachefs: Don't require flush/fua on every journal write · adbcada4

Kent Overstreet authored Nov 14, 2020

This patch adds a flag to journal entries which, if set, indicates that
they weren't done as flush/fua writes.

 - non flush/fua journal writes don't update last_seq (i.e. they don't
   free up space in the journal), thus the journal free space
   calculations now check whether nonflush journal writes are currently
   allowed (i.e. are we low on free space, or would doing a flush write
   free up a lot of space in the journal)

 - write_delay_ms, the user configurable option for when open journal
   entries are automatically written, is now interpreted as the max
   delay between flush journal writes (default 1 second).

 - bch2_journal_flush_seq_async is changed to ensure a flush write >=
   the requested sequence number has happened

 - journal read/replay must now ignore, and blacklist, any journal
   entries newer than the most recent flush entry in the journal. Also,
   the way the read_entire_journal option is handled has been improved;
   struct journal_replay now has an entry, 'ignore', for entries that
   were read but should not be used.

 - assorted refactoring and improvements related to journal read in
   journal_io.c and recovery.c

Previously, we'd have to issue a flush/fua write every time we
accumulated a full journal entry - typically the bucket size. Now we
need to issue them much less frequently: when an fsync is requested, or
it's been more than write_delay_ms since the last flush, or when we need
to free up space in the journal. This is a significant performance
improvement on many write heavy workloads.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

adbcada4

bcachefs: Improve journal free space calculations · b6df4325

Kent Overstreet authored Nov 14, 2020

Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

b6df4325

bcachefs: Increase journal pipelining · ebb84d09

Kent Overstreet authored Nov 13, 2020

This patch increases the maximum journal buffers in flight from 2 to 4 -
this will be particularly helpful when in the future we stop requiring
flush+fua for every journal write.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

ebb84d09

bcachefs: Don't issue btree writes that weren't journalled · 5db43418

Kent Overstreet authored Dec 03, 2020

If we have an error in the btree interior update path that prevents us
from journalling the update, we can't issue the corresponding btree node
write - we didn't get a journal sequence number that would cause it to
be ignored in recovery.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

5db43418

bcachefs: Check for errors in bch2_journal_reclaim() · afa7cb0c

Kent Overstreet authored Dec 03, 2020

If the journal is halted, journal reclaim won't necessarily be able to
make any forward progress, and won't accomplish anything anyways - we
should bail out so that we don't get stuck looping in reclaim when the
caches are too dirty and we should be shutting down.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

afa7cb0c

bcachefs: Flag inodes that had btree update errors · 33c74e41

Kent Overstreet authored Dec 03, 2020

On write error, the vfs inode's i_size may be inconsistent with the
btree inode's i_size - flag this so we don't have spurious assertions.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

33c74e41

bcachefs: Improve some IO error messages · 0fefe8d8

Kent Overstreet authored Dec 03, 2020

it's useful to know whether an error was for a read or a write - this
also standardizes error messages a bit more.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

0fefe8d8

bcachefs: Refactor filesystem usage accounting · f299d573

Kent Overstreet authored Nov 13, 2020

Various filesystem usage counters are kept in percpu counters, with one
set per in flight journal buffer. Right now all the code that deals with
it assumes that there's only two buffers/sets of counters, but the
number of journal bufs is getting increased to 4 in the next patch - so
refactor that code to not assume a constant.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

f299d573

bcachefs: Fix spurious alloc errors on forced shutdown · 7bfbbd88

Kent Overstreet authored Dec 02, 2020

Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

7bfbbd88

bcachefs: Fix some spurious gcc warnings · b206df6e

Kent Overstreet authored Dec 03, 2020

These only come up when building in userspace, for some reason.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

b206df6e

bcachefs: Fix journal_flush_seq() · c5bb1690

Kent Overstreet authored Dec 02, 2020

The error check was inverted - leading fsyncs to get stuck and hang,
oops.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

c5bb1690

bcachefs: bch2_trans_get_iter() no longer returns errors · 3eb26d01

Kent Overstreet authored Dec 01, 2020

Since we now always preallocate the maximum number of iterators when we
initialize a btree transaction, getting an iterator never fails - we can
delete a fair amount of error path code.

This patch also simplifies the iterator allocation code a bit.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

3eb26d01

bcachefs: Add error handling to unit & perf tests · ec3d21a9

Kent Overstreet authored Dec 01, 2020

This way, these tests can be used with tests that inject IO errors and
shut down the filesystem.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

ec3d21a9

bcachefs: Journal pin refactoring · 231db03c

Kent Overstreet authored Dec 01, 2020

This deletes some duplicated code.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

231db03c

bcachefs: Fix for fsck spuriously finding duplicate extents · 34c1cd6a

Kent Overstreet authored Dec 01, 2020

Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

34c1cd6a

bcachefs: Use BTREE_ITER_PREFETCH in journal+btree iter · 2e9f3b88

Kent Overstreet authored Dec 01, 2020

Introducing the journal+btree iter introduced a regression where we
stopped using BTREE_ITER_PREFETCH - this is a performance regression on
rotating disks.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

2e9f3b88

bcachefs: Ensure we always have a journal pin in interior update path · 04e23a56

Kent Overstreet authored Nov 30, 2020

For the new nodes an interior btree update makes reachable, updates to
those nodes may be journalled after the btree update starts but before
the transactional part - where we make those nodes reachable. Those
updates need to be kept in the journal until after the btree update
completes, hence we should always get a journal pin at the start of the
interior update.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

04e23a56

bcachefs: Change a BUG_ON() to a fatal error · d7b04163

Kent Overstreet authored Nov 30, 2020

In the btree key cache code, failing to flush a dirty key is a serious
error, but it doesn't need to be a BUG_ON(), we can stop the filesystem
instead.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

d7b04163

bcachefs: Fix error in filesystem initialization · d0022290

Kent Overstreet authored Nov 29, 2020

The rhashtable code doesn't like when we destroy an rhashtable that was
never initialized
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

d0022290

bcachefs: Fix journal reclaim spinning in recovery · 5731cf01

Kent Overstreet authored Nov 29, 2020

We can't run journal reclaim until we've finished replaying updates to
interior btree nodes - the check for this was in the wrong place though,
leading to journal reclaim spinning before it was allowed to proceed.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

5731cf01

bcachefs: Fix for __readahead_batch getting partial batch · 89931472

Kent Overstreet authored Nov 29, 2020

We were incorrectly ignoring the return value of __readahead_batch,
leading to a null ptr deref in __bch2_page_state_create().
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

89931472

bcachefs: Optimize bch2_journal_flush_seq_async() · 33b3b1dc

Kent Overstreet authored Nov 20, 2020

Avoid taking the journal lock if we don't have to.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

33b3b1dc

bcachefs: Delete dead code · 7b489207

Kent Overstreet authored Nov 20, 2020

The interior btree node update path has changed, this is no longer
needed.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

7b489207

bcachefs: bch2_btree_delete_range_trans() · 087c2019

Kent Overstreet authored Nov 20, 2020

This helps reduce stack usage by avoiding multiple btree_trans on the
stack.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

087c2019

bcachefs: Don't use bkey cache for inode update in fsck · 6584e84a

Kent Overstreet authored Nov 20, 2020

fsck doesn't know about the btree key cache, and non-cached iterators
aren't cache coherent (yet?)
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

6584e84a

bcachefs: Fix an rcu splat · f3020550

Kent Overstreet authored Nov 20, 2020

bch2_bucket_alloc() requires rcu_read_lock() to be held.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

f3020550

bcachefs: Move journal reclaim to a kthread · b7a9bbfc

Kent Overstreet authored Nov 19, 2020

This is to make tracing easier.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

b7a9bbfc

bcachefs: Throttle updates when btree key cache is too dirty · d5425a3b

Kent Overstreet authored Nov 19, 2020

This is needed to ensure we don't deadlock because journal reclaim and
thus memory reclaim isn't making forward progress.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

d5425a3b

bcachefs: Journal reclaim requires memalloc_noreclaim_save() · 9d4582ff

Kent Overstreet authored Nov 19, 2020

Memory reclaim requires journal reclaim to make forward progress - it's
what cleans our caches - thus, while we're in journal reclaim or holding
the journal reclaim lock we can't recurse into memory reclaim.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

9d4582ff

bcachefs: Simplify transaction commit error path · b3c2a06b

Kent Overstreet authored Nov 20, 2020

The transaction restart path traverses all iterators, we don't need to
do it here.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

b3c2a06b