Commits · 53b1c6f44b1a98ea6def11b74c1fde9710f2a0b9 · Kirill Smelkov / linux

22 Oct, 2023 40 commits

bcachefs: Don't use key cache during fsck · 53b1c6f4

Kent Overstreet authored Oct 14, 2022

The btree key cache mainly helps with lock contention, at the cost of
additional memory overhead. During some fsck passes the memory overhead
really matters, but fsck is single threaded so lock contention is an
issue - so skipping the key cache during fsck will help with
performance.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

53b1c6f4

bcachefs: Run check_extents_to_backpointers() in multiple passes · b32f9a57

Kent Overstreet authored Sep 28, 2022

Similer to the previous patch for check_backpointers_to_extents(), if
the alloc + backpointers btrees do not fit in ram we need to run into
multiple passes.

The counting of btree nodes that fit in memory is different here,
because we have to walk the alloc and backpointers btrees at the same
time, since a backpointer could reside in either of them and we don't
know which without checking both.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

b32f9a57

bcachefs: Run bch2_check_backpointers_to_extents() in multiple passes if necessary · 23792a71

Kent Overstreet authored Oct 09, 2022

When the extents + reflink btrees don't fit into memory this fsck pass
becomes _much_ slower, since we're doing random lookups.

This patch changes this pass to check how much of the relevant btrees
will fit into memory, and run in multiple passes if needed.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

23792a71

bcachefs: Don't stop copygc while removing devices · 15949c54

Kent Overstreet authored Oct 09, 2022

With the new backpointer based copygc we don't need an explicit copygc
reserve, we're always evacuating buckets one at a time - so this is no
longer needed, and in fact removing it fixes a deadlock in
bch2_dev_allocator_remove().
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

15949c54

bcachefs: Delete in memory ec backpointers · c9828cea

Kent Overstreet authored Oct 09, 2022

Post btree backpointers, these aren't needed anymore.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

c9828cea

bcachefs: Erasure coding now uses backpointers · dea5647e

Kent Overstreet authored Oct 09, 2022

This is only a start to updating erasure coding for backpointers - it's
still not working yet. The subsequent patch will delete our old in
memory backpointers for copygc, and this fixes a spurious EPERM
bug/error message.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

dea5647e

bcachefs: Copygc now uses backpointers · 8e3f913e

Kent Overstreet authored Mar 18, 2022

Previously, copygc needed to walk the entire extents & reflink btrees to
find extents that needed to be moved.

Now that we have backpointers, this patch implements
bch2_evacuate_bucket() in the move code, which copygc now uses for
evacuating mostly empty buckets.

Also, thanks to the new backpointers code, copygc can now move btree
nodes.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

8e3f913e

bcachefs: New on disk format: Backpointers · a8c752bb

Kent Overstreet authored Mar 17, 2022

This patch adds backpointers: we now have a reverse index from device
and offset on that device (specifically, offset within a bucket) back to
btree nodes and (non cached) data extents.

The first 40 backpointers within a bucket are stored in the alloc key;
after that backpointers spill over to the next backpointers btree. This
is to help avoid performance regressions from additional btree updates
on large streaming workloads.

This patch adds all the code for creating, checking and repairing
backpointers. The next patch in the series is going to use backpointers
for copygc - finally getting rid of the need to scan all extents to do
copygc.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

a8c752bb

bcachefs: Btree write buffer · 920e69bc

Kent Overstreet authored Jan 04, 2023

This adds a new method of doing btree updates - a straight write buffer,
implemented as a flat fixed size array.

This is only useful when we don't need to read from the btree in order
to do the update, and when reading is infrequent - perfect for the LRU
btree.

This will make LRU btree updates fast enough that we'll be able to use
it for persistently indexing buckets by fragmentation, which will be a
massive boost to copygc performance.

Changes:
 - A new btree_insert_type enum, for btree_insert_entries. Specifies
   btree, btree key cache, or btree write buffer.

 - bch2_trans_update_buffered(): updates via the btree write buffer
   don't need a btree path, so we need a new update path.

 - Transaction commit path changes:
   The update to the btree write buffer both mutates global, and can
   fail if there isn't currently room. Therefore we do all write buffer
   updates in the transaction all at once, and also if it fails we have
   to revert filesystem usage counter changes.

   If there isn't room we flush the write buffer in the transaction
   commit error path and retry.

 - A new persistent option, for specifying the number of entries in the
   write buffer.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

920e69bc

bcachefs: Go RW before check_alloc_info() · f2b542ba

Kent Overstreet authored Dec 11, 2022

It's possible to do btree updates before going RW by adding them to the
list of updates for journal replay to do, but this is limited by what
fits in RAM. This patch switches the second alloc info phase to run
after going RW - btree_gc has already ensured the alloc btree itself is
correct - and tweaks the allocation path to deal with the potential
small inconsistencies.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

f2b542ba

bcachefs: Start copygc when first going read-write · 5f5c7466

Kent Overstreet authored Oct 17, 2022

In the distant past, it wasn't possible to start copygc until after
journal replay had finished. Now, the btree iterator code overlays keys
from the journal, so there's no reason not to start it earlier - and it
solves a rare deadlock.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

5f5c7466

bcachefs: Kill trans->flags · 30ca6ece

Kent Overstreet authored Feb 09, 2023

Recursive transaction commits are occasionally necessary - in
particular, for the upcoming btree write buffer's flush path.

This avoids bugs due to trans->flags being accidentally mutated
mid-commit, which can cause c->writes refcount leaks.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

30ca6ece

bcachefs: trans->notrace_relock_fail · 60b55388

Kent Overstreet authored Feb 09, 2023

When we unlock in order to submit IO, the next relock event is likely to
fail if submit_bio() blocked - we shouldn't those events in our _fail
stats, since those are expected events and shouldn't cause test
failures.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

60b55388

bcachefs: Debug mode for c->writes references · d94189ad

Kent Overstreet authored Feb 09, 2023

This adds a debug mode where we split up the c->writes refcount into
distinct refcounts for every codepath that takes a reference, and adds
sysfs code to print the value of each ref.

This will make it easier to debug shutdown hangs due to refcount leaks.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

d94189ad

bcachefs: ec_stripe_delete_work() now takes ref on c->writes · dd81a060
Kent Overstreet authored Feb 09, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
dd81a060

bcachefs: Fix btree_node_write_blocked() not being cleared · 06ab86d5

Kent Overstreet authored Feb 09, 2023

The btree_node_write_blocked bit was a later addition to this code,
it only mirrors the state of the b->write_blocked list (empty or
nonempty) - unfortunately, when it was added it wasn't correctly kept in
sync - oops.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

06ab86d5

bcachefs: Switch a BUG_ON() to a panic() · 434b1c75

Kent Overstreet authored Feb 09, 2023

This assert is popping - rarely - in the CI, this will help us track it
down from the logs.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

434b1c75

bcachefs: Fix btree_path_alloc() · 992fa4e6

Kent Overstreet authored Feb 08, 2023

We need to call bch2_trans_update_max_paths() before marking the new
path as allocated, since we're not initializing it yet.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

992fa4e6

bcachefs: Fix memleak in replicas_table_update() · d7afe651

Brett Holman authored Feb 10, 2023

Signed-off-by: Brett Holman <bholman.devel@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

d7afe651

bcachefs: Use for_each_btree_key_upto() more consistently · c72f687a

Kent Overstreet authored Oct 11, 2022

It's important that in BTREE_ITER_FILTER_SNAPSHOTS mode we always use
peek_upto() and provide an end for the interval we're searching for -
otherwise, when we hit the end of the inode the next inode be in a
different subvolume and not have any keys in the current snapshot, and
we'd iterate over arbitrarily many keys before returning one.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

c72f687a

bcachefs: Don't call bch2_journal_pin_drop() under key cache lock · 5b3008bc

Kent Overstreet authored Mar 02, 2023

This fixes a (harmless) lockdep splat, due to a lock order violation in
the key cache exit path.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

5b3008bc

six locks: Improved optimistic spinning · 91db8066

Kent Overstreet authored Feb 05, 2023

This adds a threshold for the maximum spin time, similar to the rwsem
code, and a flag to the lock itself indicating when we've spun too long
so other threads also refrain from spinning.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

91db8066

bcachefs: Use six_lock_ip() · 94c69faf

Kent Overstreet authored Feb 04, 2023

This uses the new _ip() interface to six locks and hooks it up to
btree_path->ip_allocated, when available.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

94c69faf

six locks: Expose tracepoint IP · f746c62c

Kent Overstreet authored Feb 04, 2023

This adds _ip variations of the various lock functions that allow an IP
to be passed in, which is used by lockstat.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

f746c62c

bcachefs: bch2_trans_in_restart_error() · 12344c7c

Kent Overstreet authored Feb 01, 2023

This replaces various BUG_ON() assertions with panics that tell us where
the restart was done and the restart type.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

12344c7c

bcachefs: Improve btree node read error path · 2e984040

Kent Overstreet authored Feb 01, 2023

This ensures that failure to read a btree node error is treated as a
topology error, and returns the correct error so that the topology
repair pass will be run.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

2e984040

bcachefs: Fix bch2_trans_reset_updates() · 464b4155

Kent Overstreet authored Feb 05, 2023

This should have been resetting trans->fs_usage_deltas as well.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

464b4155

bcachefs: Inline bch2_btree_path_traverse() fastpath · 4e3d1899
Kent Overstreet authored Feb 04, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
4e3d1899

bcachefs: Fix hash_check_key() · 419fc65f

Kent Overstreet authored Feb 01, 2023

On hash collision when we have to check for duplicates or incorrect
hash value, we weren't specifying a snapshot ID to iterate with.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

419fc65f

bcachefs: Don't emit tracepoints for expected events · b8c5b16f
Kent Overstreet authored Jan 25, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
b8c5b16f

bcachefs: Use trylock in bch2_prt_backtrace() · 3e57db65

Kent Overstreet authored Jan 25, 2023

Easy workaround for a lockdep splat - and since bch2_prt_backtrace() is
only used in debug code this is fine.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

3e57db65

bcachefs: bch2_inode_opts_get() · 01ad6737

Kent Overstreet authored Nov 23, 2022

This improves io_opts() and makes it a non-inline function - it's big
enough that it probably shouldn't be.

Also, bch_io_opts no longer needs fields for whether options are
defined, so we can slim it down a bit.

We'd like to stop passing around the full bch_io_opts, but that'll be
tricky because of bch2_rebalance_add_key().
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

01ad6737

bcachefs: Fix bch_alloc_to_text() · f52dd1ae

Kent Overstreet authored Dec 19, 2022

We weren't guarding against the alloc key having an invalid data type.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

f52dd1ae

bcachefs: Better inlining in core write path · 393a1f68

Kent Overstreet authored Nov 24, 2022

Provide inline versions of some allocation functions
 - bch2_alloc_sectors_done_inlined()
 - bch2_alloc_sectors_append_ptrs_inlined()

and use them in the core IO path.

Also, inline bch2_extent_update_i_size_sectors() and
bch2_bkey_append_ptr().

In the core write path, function call overhead matters - every function
call is a jump to a new location and a potential cache miss.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

393a1f68

bcachefs: Better inlining for bch2_alloc_to_v4_mut · 19a614d2

Kent Overstreet authored Jan 30, 2023

This separates out the slowpath into a separate function, and inlines
bch2_alloc_v4_mut into bch2_trans_start_alloc_update(), the main place
it's called.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

19a614d2

bcachefs: Improve btree_reserve_get_fail tracepoint · adf6360b
Kent Overstreet authored Feb 01, 2023
```
Now we include the return code.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
adf6360b

bcachefs: Fix bch2_bucket_alloc_early() · db36c147

Kent Overstreet authored Jan 23, 2023

We were incorrectly retrying after a transaction restart.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

db36c147

bcachefs: Check for lru entries with time=0 · 9fea089a
Kent Overstreet authored Jan 04, 2023
```
These are invalid.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
9fea089a

bcachefs: Fix rereplicate when we already have a cached pointer · d7dd3fb8

Kent Overstreet authored Jan 07, 2023

When we need to add more replicas to an extent, it might be the case
that we already have a replica on every device, but some of them are
cached.

This patch fixes a bug where we'd spin on that extent because the write
path fails to find a device we can allocate from: we allow allocating
from devices that already have cached replicas on them, and change
bch2_data_update_index_update() to drop the cached replica if needed.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

d7dd3fb8

bcachefs: Fix repair path in bch2_mark_reflink_p() · 7c909f65
Kent Overstreet authored Jan 20, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
7c909f65