Commits · 42d237320e9817a94f3a0a2de28156523596b086 · Kirill Smelkov / linux

22 Oct, 2023 40 commits

bcachefs: Snapshot creation, deletion · 42d23732

Kent Overstreet authored Mar 16, 2021

This is the final patch in the patch series implementing snapshots.
This patch implements two new ioctls that work like creation and
deletion of directories, but fancier.

 - BCH_IOCTL_SUBVOLUME_CREATE, for creating new subvolumes and snaphots
 - BCH_IOCTL_SUBVOLUME_DESTROY, for deleting subvolumes and snapshots
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

42d23732

bcachefs: Require snapshot id to be set · a861c722

Kent Overstreet authored Mar 15, 2021

Now that all the existing code has been converted for snapshots, this
patch changes the code for initializing a btree iterator to require a
snapshot to be specified, and also change bkey_invalid() to allow for
non U32_MAX snapshot IDs.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

a861c722

bcachefs: Fix unit & perf tests for snapshots · 6f83cb84

Kent Overstreet authored Dec 15, 2021

This finishes updating the unit & perf tests for snapshots - btrees that
use snapshots now always require the snapshot field of the start
position to be a valid snapshot ID.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

6f83cb84

bcachefs: Update data move path for snapshots · 18443cb9

Kent Overstreet authored Aug 05, 2021

The data move path operates on existing extents, and not within a
subvolume as the regular IO paths do. It needs to change because it may
cause existing extents to be split, and when splitting an existing
extent in an ancestor snapshot we need to make sure the new split has
the same visibility in child snapshots as the existing extent.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

18443cb9

bcachefs: Whiteouts for snapshots · 7a7d17b2

Kent Overstreet authored Feb 02, 2021

This patch adds KEY_TYPE_whiteout, a new type of whiteout for snapshots,
when we're deleting and the key being deleted is in an ancestor
snapshot - and updates the transaction update/commit path to use it.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

7a7d17b2

bcachefs: Convert io paths for snapshots · 8c6d298a

Kent Overstreet authored Mar 12, 2021

This plumbs around the subvolume ID as was done previously for other
filesystem code, but now for the IO paths - the control flow in the IO
paths is trickier so the changes in this patch are more involved.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

8c6d298a

bcachefs: Update fsck for snapshots · ef1669ff

Kent Overstreet authored Apr 20, 2021

This updates the fsck algorithms to handle snapshots - meaning there
will be multiple versions of the same key (extents, inodes, dirents,
xattrs) in different snapshots, and we have to carefully consider which
keys are visible in which snapshot.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

ef1669ff

bcachefs: Plumb through subvolume id · 6fed42bb

Kent Overstreet authored Mar 16, 2021

To implement snapshots, we need every filesystem btree operation (every
btree operation without a subvolume) to start by looking up the
subvolume and getting the current snapshot ID, with
bch2_subvolume_get_snapshot() - then, that snapshot ID is used for doing
btree lookups in BTREE_ITER_FILTER_SNAPSHOTS mode.

This patch adds those bch2_subvolume_get_snapshot() calls, and also
switches to passing around a subvol_inum instead of just an inode
number.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

6fed42bb

bcachefs: BTREE_ITER_FILTER_SNAPSHOTS · c075ff70

Kent Overstreet authored Mar 04, 2021

For snapshots, we need to implement btree lookups that return the first
key that's an ancestor of the snapshot ID the lookup is being done in -
and filter out keys in unrelated snapshots. This patch adds the btree
iterator flag BTREE_ITER_FILTER_SNAPSHOTS which does that filtering.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

c075ff70

bcachefs: Add subvolume to ei_inode_info · 284ae18c

Kent Overstreet authored Mar 16, 2021

Filesystem operations generally operate within a subvolume: at the start
of every btree transaction we'll be looking up (and locking) the
subvolume to get the current snapshot ID, which we then use for our
other btree lookups in BTREE_ITER_FILTER_SNAPSHOTS mode.

But inodes don't record what subvolume they're in - they can't, because
if they did we'd have to update every single inode within a subvolume
when taking a snapshot in order to keep that field up to date. So it
needs to be tracked in memory, based on how we got to that inode.

Hence this patch adds a subvolume field to ei_inode_info, and switches
to iget5() so we can index by it in the inode hash table.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

284ae18c

bcachefs: Per subvolume lost+found · 81ed9ce3

Kent Overstreet authored Apr 19, 2021

On existing filesystems, we have a single global lost+found. Introducing
subvolumes means we need to introduce per subvolume lost+found
directories, because inodes are added to lost+found by their inode
number, and inode numbers are now only unique within a subvolume.

This patch adds support to fsck for per subvolume lost+found.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

81ed9ce3

bcachefs: Add support for dirents that point to subvolumes · b9e1adf5

Kent Overstreet authored Mar 16, 2021

Dirents currently always point to inodes. Subvolumes add a new type of
dirent, with d_type DT_SUBVOL, that instead points to an entry in the
subvolumes btree, and the subvolume has a pointer to the root inode.

This patch adds bch2_dirent_read_target() to get the inode (and
potentially subvolume) a dirent points to, and changes existing code to
use that instead of reading from d_inum directly.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

b9e1adf5

bcachefs: Subvolumes, snapshots · 14b393ee

Kent Overstreet authored Mar 16, 2021

This patch adds subvolume.c - support for the subvolumes and snapshots
btrees and related data types and on disk data structures. The next
patches will start hooking up this new code to existing code.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

14b393ee

bcachefs: Disable quota support · 8948fc8f

Kent Overstreet authored Sep 26, 2021

Existing quota support breaks badly with snapshots. We're not deleting
the code because some of it will be needed when we reimplement quotas
along the lines of btrfs subvolume quotas.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

8948fc8f

Revert "bcachefs: Add more assertions for locking btree iterators out of order" · 3074bc0f

Kent Overstreet authored Sep 15, 2021

Figured out the bug we were chasing, and it had nothing to do with
locking btree iterators/paths out of order.

This reverts commit ff08733dd298c969aec7c7828095458f73fd5374.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

3074bc0f

bcachefs: Improve btree_node_mem_ptr optimization · aae4eea6

Kent Overstreet authored Sep 13, 2021

This patch checks b->hash_val before attempting to lock the node in the
btree, which makes it more equivalent to the "lookup in hash table"
path - and potentially avoids an unnecessary transaction restart if
btree_node_mem_ptr(k) no longer points to the node we want.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

aae4eea6

bcachefs: Add a missing bch2_trans_relock() call · aa76bd33

Kent Overstreet authored Sep 13, 2021

This was causing an assertion to pop in fsck, in one of the repair
paths.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

aa76bd33

bcachefs: Fix some compiler warnings · c79272d1

Kent Overstreet authored Sep 09, 2021

gcc couldn't always deduce that written wasn't used uninitialized
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

c79272d1

bcachefs: Add missing BTREE_ITER_INTENT · 5b5b03e7

Kent Overstreet authored Sep 07, 2021

No reason not to be using it here...
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

5b5b03e7

bcachefs: Better approach to write vs. read lock deadlocks · caaa66aa

Kent Overstreet authored Sep 07, 2021

Instead of unconditionally upgrading read locks to intent locks in
do_bch2_trans_commit(), this patch changes the path that takes write
locks to first trylock, and then if trylock fails check if we have a
conflicting read lock, and restart the transaction if necessary.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

caaa66aa

bcachefs: normalize_read_intent_locks · b301105b

Kent Overstreet authored Sep 07, 2021

This is a new approach to avoiding the self deadlock we'd get if we
tried to take a write lock on a node while holding a read lock - we
simply upgrade the readers to intent.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

b301105b

bcachefs: Consolidate intent lock code in btree_path_up_until_good_node · 8ee0134e

Kent Overstreet authored Sep 07, 2021

We need to take all needed intent locks when relocking an iterator:
bch2_btree_path_traverse() had a special cased, faster version of this,
but it really should be in up_until_good_node() so that set_pos() can
use it too.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

8ee0134e

bcachefs: Optimize btree lookups in write path · db92f2ea

Kent Overstreet authored Sep 07, 2021

This patch significantly reduces the number of btree lookups required in
the extent update path.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

db92f2ea

bcachefs: Add a missing btree_path_make_mut() call · c404f203

Kent Overstreet authored Sep 07, 2021

Also add another small helper, btree_path_clone().
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

c404f203

bcachefs: Enabled shard_inode_numbers by default · 8ffa63cd

Kent Overstreet authored Sep 07, 2021

We'd like performance increasing options to be on by default.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

8ffa63cd

bcachefs: No need to clone iterators for update · cf3c68cd

Kent Overstreet authored Sep 06, 2021

Since btree_path is now internally refcounted, we don't need to clone an
iterator before calling bch2_trans_update() if we'll be mutating it.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

cf3c68cd

bcachefs: Kill retry loop in btree merge path · 22b383ad
Kent Overstreet authored Sep 05, 2021
```
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
```
22b383ad

bcachefs: Drop some fast path tracepoints · f48361b0

Kent Overstreet authored Sep 05, 2021

These haven't turned out to be useful
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

f48361b0

bcachefs: Tighten up btree locking invariants · 1d3ecd7e

Kent Overstreet authored Sep 04, 2021

New rule is: if a btree path holds any locks it should be holding
precisely the locks wanted (accoringing to path->level and
path->locks_want).
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

1d3ecd7e

bcachefs: Extent btree iterators are no longer special · 1ae29c1f

Kent Overstreet authored Sep 04, 2021

Since iter->real_pos was introduced, we no longer have to deal with
extent btree iterators that have skipped past deleted keys - this is a
real performance improvement on btree updates.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

1ae29c1f

bcachefs: Add more assertions for locking btree iterators out of order · 068bcaa5

Kent Overstreet authored Sep 03, 2021

btree_path_traverse_all() traverses btree iterators in sorted order, and
thus shouldn't see transaction restarts due to potential deadlocks - but
sometimes we do. This patch adds some more assertions and tracks some
more state to help track this down.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

068bcaa5

bcachefs: Kill bpos_diff() XXX check for perf regression · 807dda8c

Kent Overstreet authored Aug 30, 2021

This improves the btree iterator lookup code by using
trans_for_each_iter_inorder().
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

807dda8c

bcachefs: btree_path · 67e0dd8f

Kent Overstreet authored Aug 30, 2021

This splits btree_iter into two components: btree_iter is now the
externally visible componont, and it points to a btree_path which is now
reference counted.

This means we no longer have to clone iterators up front if they might
be mutated - btree_path can be shared by multiple iterators, and cloned
if an iterator would mutate a shared btree_path. This will help us use
iterators more efficiently, as well as slimming down the main long lived
state in btree_trans, and significantly cleans up the logic for iterator
lifetimes.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

67e0dd8f

bcachefs: Fix initialization of bch_write_op.nonce · 8f54337d

Kent Overstreet authored Sep 03, 2021

If an extent ends up with a replica that is encrypted an a replica that
isn't encrypted (due the user changing options), and then
copygc/rebalance moves one of the replicas by reading from the
unencrypted replica, we had a bug where we wouldn't correctly initialize
op->nonce - for each crc field in an extent, crc.offset + crc.nonce must
be equal.

This patch fixes that by moving op.nonce initialization to
bch2_migrate_write_init.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

8f54337d

bcachefs: Improve an error message · fbf14104

Kent Overstreet authored Sep 01, 2021

When we detect an invalid key being inserted, we should print what code
was doing the update.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

fbf14104

bcachefs: Add an assertion for removing btree nodes from cache · cab8e233

Kent Overstreet authored Sep 01, 2021

Chasing a bug that has something to do with the btree node cache.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

cab8e233

bcachefs: Kill BTREE_ITER_NODES · f21566f1

Kent Overstreet authored Aug 30, 2021

We really only need to distinguish between btree iterators and btree key
cache iterators - this is more prep work for btree_path.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

f21566f1

bcachefs: Kill BTREE_ITER_NEED_PEEK · deb0e573

Kent Overstreet authored Aug 30, 2021

This was used for an optimization that hasn't existing in quite awhile
- iter->uptodate will probably be going away as well.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

deb0e573

bcachefs: Prefer using btree_insert_entry to btree_iter · 6fba6b83

Kent Overstreet authored Aug 30, 2021

This moves some data dependencies forward, to improve pipelining.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

6fba6b83

bcachefs: More renaming · a0a56879
Kent Overstreet authored Aug 30, 2021
```
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
```
a0a56879