Commits · 73da30e8e0f8ffcc91691934f202ab6e2f985604 · Kirill Smelkov / linux

22 Oct, 2023 40 commits

bcachefs: Fix check_overlapping_extents() · 73da30e8

Kent Overstreet authored May 13, 2023

A error check had a flipped conditional - whoops.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

73da30e8

bcachefs: Replace a BUG_ON() with fatal error · 4a2e5d7b

Kent Overstreet authored May 12, 2023

A user hit this BUG_ON() - it's unclear how it happened, so replace it
with a fatal error that will cause us to go read only, and print out
more information.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

4a2e5d7b

bcachefs: Delete some dead code in bch2_replicas_gc_end() · 92e637ce

Kent Overstreet authored May 08, 2023

bch2_replicas_gc_(start|end) is now only used for journal replicas
entries, which don't have bucket sector counts - so this code is
entirely dead and can be deleted.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

92e637ce

bcachefs: mark journal replicas before journal write submission · a7b29b8d

Brian Foster authored May 04, 2023

The journal write submission path marks the associated replica
entries for journal data in journal_write_done(), which is just
after journal write bio submission. This creates a small window
where journal entries might have been written out, but the
associated replica is not marked such that recovery does not know
that the associated device contains journal data.

Move the replica marking a bit earlier in the write path such that
recovery is guaranteed to recognize that the device contains journal
data in the event of a crash.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

a7b29b8d

bcachefs: Improved comment for bch2_replicas_gc2() · 38e3d93f
Kent Overstreet authored May 02, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
38e3d93f

bcachefs: Fix quotas + snapshots · cb1b479d

Kent Overstreet authored Apr 28, 2023

Now that we can reliably designate and find the master subvolume out of
a tree of snapshots, we can finally make quotas work with snapshots:

That is - quotas will now _ignore_ snapshot subvolumes, and only be in
effect for the master (non snapshot) subvolume.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

cb1b479d

bcachefs: Add otime, parent to bch_subvolume · 653693be

Kent Overstreet authored Mar 29, 2023

Add two new fields to bch_subvolume:
 - otime: creation time
 - parent: For snapshots, this is the id of the subvolume the snapshot
   was created from
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

653693be

bcachefs: BTREE_ID_snapshot_tree · 1c59b483

Kent Overstreet authored Mar 29, 2023

This adds a new btree which gets us a persistent per-snapshot-tree
identifier.

 - BTREE_ID_snapshot_trees
 - KEY_TYPE_snapshot_tree
 - bch_snapshot now has a field that points to a snapshot_tree

This is going to be used to designate one snapshot ID/subvolume out of a
given tree of snapshots as the "main" subvolume, so that we can do quota
accounting in that subvolume and not the rest.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

1c59b483

bcachefs: bch2_bkey_get_empty_slot() · 51e84d3b

Kent Overstreet authored Apr 27, 2023

Add a new helper for allocating a new slot in a btree.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

51e84d3b

bcachefs: bch2_bkey_make_mut() now calls bch2_trans_update() · dbda63bb

Kent Overstreet authored Apr 30, 2023

It's safe to call bch2_trans_update with a k/v pair where the value
hasn't been filled out, as long as the key part has been and the value
is filled out by transaction commit time.

This patch folds the bch2_trans_update() call into bch2_bkey_make_mut(),
eliminating a bit of boilerplate.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

dbda63bb

bcachefs: bch2_bkey_get_mut() now calls bch2_trans_update() · f12a798a

Kent Overstreet authored Apr 30, 2023

It's safe to call bch2_trans_update with a k/v pair where the value
hasn't been filled out, as long as the key part has been and the value
is filled out by transaction commit time.

This patch folds the bch2_trans_update() call into bch2_bkey_get_mut(),
eliminating a bit of boilerplate.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

f12a798a

bcachefs: bch2_bkey_alloc() now calls bch2_trans_update() · f8cb35fd

Kent Overstreet authored Apr 30, 2023

It's safe to call bch2_trans_update with a k/v pair where the value
hasn't been filled out, as long as the key part has been and the value
is filled out by transaction commit time.

This patch folds the bch2_trans_update() call into bch2_bkey_alloc(),
eliminating a bit of boilerplate.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

f8cb35fd

bcachefs: bch2_bkey_get_mut() improvements · 34dfa5db

Kent Overstreet authored Apr 27, 2023

 - bch2_bkey_get_mut() now handles types increasing in size, allocating
   a buffer for the type's current size when necessary
 - bch2_bkey_make_mut_typed()
 - bch2_bkey_get_mut() now initializes the iterator, like
   bch2_bkey_get_iter()

Also, refactor so that most of the code is in functions - now macros are
only used for wrappers.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

34dfa5db

bcachefs: Move bch2_bkey_make_mut() to btree_update.h · d67a16df

Kent Overstreet authored Apr 30, 2023

It's for doing updates - this is where it belongs, and next pathes will
be changing these helpers to use items from btree_update.h.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

d67a16df

bcachefs: bch2_bkey_get_iter() helpers · bcb79a51

Kent Overstreet authored Apr 29, 2023

Introduce new helpers for a common pattern:

  bch2_trans_iter_init();
  bch2_btree_iter_peek_slot();

 - bch2_bkey_get_iter_type() returns -ENOENT if it doesn't find a key of
   the correct type
 - bch2_bkey_get_val_typed() copies the val out of the btree to a
   (typically stack allocated) variable; it handles the case where the
   value in the btree is smaller than the current version of the type,
   zeroing out the remainder.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

bcb79a51

bcachefs: bkey_ops.min_val_size · 174f930b

Kent Overstreet authored Apr 29, 2023

This adds a new field to bkey_ops for the minimum size of the value,
which standardizes that check and also enforces the new rule (previously
done somewhat ad-hoc) that we can extend value types by adding new
fields on to the end.

To make that work we do _not_ initialize min_val_size with sizeof,
instead we initialize it to the size of the first version of those
values.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

174f930b

bcachefs: Converting to typed bkeys is now allowed for err, null ptrs · ab158fce
Kent Overstreet authored Apr 30, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
ab158fce

bcachefs: Btree iterator, update flags no longer conflict · 95b595a5

Kent Overstreet authored Apr 30, 2023

Change btree_update_flags to start after the last btree iterator flag,
so that we can pass both in the same flags argument.

This is needed for the upcoming bch2_bkey_get_mut() helper.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

95b595a5

bcachefs: remove unused key cache coherency flag · 0a23574e

Brian Foster authored May 01, 2023

Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

0a23574e

bcachefs: fix accounting corruption race between reclaim and dev add · 3c434cdf

Brian Foster authored May 01, 2023

When a device is removed from a bcachefs volume, the associated
content is removed from the various btrees. The alloc tree uses the
key cache, so when keys are removed the deletes exist in cache for a
period of time until reclaim comes along and flushes outstanding
updates.

When a device is re-added to the bcachefs volume, the add process
re-adds some of these previously deleted keys. When marking device
superblock locations on device add, the keys will likely refer to
some of the same alloc keys that were just removed. The memory
triggers for these key updates are responsible for further updates,
such as bch2_mark_alloc() calling into bch2_dev_usage_update() to
update per-device usage accounting.

When a new key is added to key cache, the trans update path also
flushes the key to the backing btree for coherency reasons for tree
walks.

With all of this context, if a device is removed and re-added
quickly enough such that some key deletes from the remove are still
pending a key cache flush, the trans update path can view this as
addition of a new key because the old key in the insert entry refers
to a deleted key. However the deleted cached key has not been filled
by absence of a btree key, but rather refers to an explicit deletion
of an existing key that occurred during device removal.

The trans update path adds a new update to flush the key and tags
the original (cached) update to skip running the memory triggers.
This results in running triggers on the non-cached update instead,
which in turn will perform accounting updates based on incoherent
values. For example, bch2_dev_usage_update() subtracts the the old
alloc key dirty sector count in the non-cached btree key from the
newly initialized (i.e. zeroed) per device counters, leading to
underflow and accounting corruption.

There are at least a few ways to avoid this problem, the simplest of
which may be to run triggers against the cached update rather than
the non-cached update. If the key only needs to be flushed when the
key is not present in the tree, however, then this still performs an
unnecessary update. We could potentially use the cached key dirty
state to determine whether the delete is a dirty, cached update vs.
a clean cache fill, but this may require transmitting key cache
dirty state across layers, which adds complexity and seems to be of
limited value. Instead, update flush_new_cached_update() to handle
this by simply checking for the key in the btree and only perform
the flush when a backing key is not present.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

3c434cdf

bcachefs: Mark bch2_copygc() noinline · 958c347b

Kent Overstreet authored Apr 29, 2023

This works around a "stack from too large" error.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

958c347b

bcachefs: Delete obsolete btree ptr check · 3140a3d0

Kent Overstreet authored Apr 27, 2023

This patch deletes a .key_invalid check for btree pointers that only
applies to _very_ old on disk format versions, and potentially
complicates the upgrade process.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

3140a3d0

bcachefs: Always run topology error when CONFIG_BCACHEFS_DEBUG=y · 6b52bcde
Kent Overstreet authored Apr 26, 2023
```
Improved test coverage.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
6b52bcde
bcachefs: Fix a userspace build error · a0668d77
Kent Overstreet authored Apr 26, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
a0668d77

bcachefs: Make sure hash info gets initialized in fsck · c8d5b714

Kent Overstreet authored Apr 25, 2023

We had some bugs with setting/using first_this_inode in the inode
walker in the dirents/xattr code.

This patch changes to not clear first_this_inode until after
initializing the new hash info.

Also, we fix an error message to not print on transaction restart, and
add a comment to related fsck error code.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

c8d5b714

bcachefs: Kill bch2_verify_bucket_evacuated() · 1af5227c

Kent Overstreet authored Apr 21, 2023

With backpointers, it's now impossible for bch2_evacuate_bucket() to be
completely reliable: it can race with an extent being partially
overwritten or split, which needs a new write buffer flush for the
backpointer to be seen.

This shouldn't be a real issue in practice; the previous patch added a
new tracepoint so we'll be able to see more easily if it is.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

1af5227c

bcachefs: Improve move path tracepoints · 5a21764d

Kent Overstreet authored Apr 20, 2023

Move path tracepoints now include the key being moved. Also, add new
tracepoints for the start of move_extent, and evacuate_bucket.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

5a21764d

bcachefs: Drop a redundant error message · 09ebfa61

Kent Overstreet authored Apr 21, 2023

When we're already read-only, we don't need to print out errors from
writing btree nodes.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

09ebfa61

bcachefs: remove bucket_gens btree keys on device removal · 02d51bb9

Brian Foster authored Apr 19, 2023

If a device has keys in the bucket_gens btree associated with its
buckets and is removed from a bcachefs volume, fsck will complain
about the presence of keys associated with an invalid device index.
A repair removes the associated keys and restores correctness.
Update bch2_dev_remove_alloc() to remove device related keys at
device removal time to avoid the problem.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

02d51bb9

bcachefs: fix NULL bch_dev deref when checking bucket_gens keys · 251babb5

Brian Foster authored Apr 18, 2023

fsck removes bucket_gens keys for devices that do not exist in the
volume (i.e., if the device was removed). In 'fsck -n' mode, the
associated fsck_err_on() wrapper returns false to skip the key
removal. This proceeds on to the rest of the function, which
eventually segfaults on a NULL bch_dev because the device does not
exist.

Update bch2_check_bucket_gens_key() to skip out of the rest of the
function when the associated device does not exist, regardless of
running fsck in check or repair mode.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

251babb5

bcachefs: folio pos to bch_folio_sector index helper · bf98ee10

Brian Foster authored Apr 03, 2023

Create a small helper to translate from file offset to the
associated bch_folio_sector index in the underlying bch_folio. The
helper assumes the file offset is covered by the passed folio.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

bf98ee10

bcachefs: Fix a null ptr deref in fsck check_extents() · e3dc75eb

Kent Overstreet authored Apr 16, 2023

It turns out, in rare situations we need to be passing in a disk
reservation, which will be used internally by the transaction commit
path when needed. Pass one in...
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

e3dc75eb

bcachefs: Fix a slab-out-of-bounds · 615fccad

Kent Overstreet authored Apr 16, 2023

In __bch2_alloc_to_v4_mut(), we overrun the buffer we allocate if the
alloc key had backpointers stored in it (which we no longer support).

Fix this with a max() call.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

615fccad

bcachefs: Allow answering y or n to all fsck errors of given type · 853b7393

Kent Overstreet authored Apr 15, 2023

This changes the ask_yn() function used by fsck to accept Y or N,
meaning yes or no for all errors of a given type.

With this, the user can be prompted only for distinct error types -
useful when a filesystem has lots of errors.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

853b7393

bcachefs: use u64 for folio end pos to avoid overflows · 6b9857b2

Brian Foster authored Mar 29, 2023

Some of the folio_end_*() helpers are prone to overflow of signed
64-bit types because the mapping is only limited by the max value of
loff_t and the associated helpers return the start offset of the
next folio. Therefore, a folio_end_pos() of the max allowable folio in a
mapping returns a value that overflows loff_t.

This makes it hard to rely on such values when doing folio
processing across a range of a file, as bcachefs attempts to do with
the recent folio changes. For example, generic/564 causes problems
in the buffered write path when testing writes at max boundary
conditions.

The current understanding is that the pagecache historically limited
the mapping to one less page to avoid this problem and this was
dropped with some of the folio conversions, but may be reinstated to
properly address the problem. In the meantime, update the internal
folio_end_*() helpers in bcachefs to return a u64, and all of the
associated code to use or cast to u64 to avoid overflow problems.
This allows generic/564 to pass and can be reverted back to using
loff_t if at any point the pagecache subsystem can guarantee these
boundary conditions will not overflow.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

6b9857b2

bcachefs: clean up post-eof folios on -ENOSPC · 335f7d4f

Brian Foster authored Mar 29, 2023

The buffered write path batches folio creations in the file mapping
based on the requested size of the write. Under low free space
conditions, it is possible to add a bunch of folios to the mapping
and then return a short write or -ENOSPC due to lack of space. If
this occurs on an extending write, the file size is updated based on
the amount of data successfully written to the file. If folios were
added beyond the final i_size, they may hang around until reclaimed,
truncated or encountered unexpectedly by another operation.

For example, generic/083 reproduces a sequence of events where a
short write leaves around one or more post-EOF folios on an inode, a
subsequent zero range request extends beyond i_size and overlaps
with an aforementioned folio, and __bch2_truncate_folio() happens
across it and complains.

Update __bch2_buffered_write() to keep track of the start offset of
the last folio added to the mapping for a prospective write. After
i_size is updated, check whether this offset starts beyond EOF. If
so, truncate pagecache beyond the latest EOF to clean up any folios
that don't reside at least partially within EOF upon completion of
the write.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

335f7d4f

bcachefs: fix truncate overflow if folio is beyond EOF · 4ad6aa46

Brian Foster authored Mar 29, 2023

generic/083 occasionally reproduces a panic caused by an overflow
when accessing the bch_folio_sector array of the folio being
processed by __bch2_truncate_folio(). The immediate cause of the
overflow is that the folio offset is beyond i_size, and therefore
the sector index calculation underflows on subtraction of the folio
offset.

One cause of this is mainly observed on nocow mounts. When nocow is
enabled, fallocate performs physical block allocation (as opposed to
block reservation in cow mode), which range_has_data() then
interprets as valid data that requires partial zeroing on truncate.
Therefore, if a post-eof zero range request lands across post-eof
preallocated blocks, __bch2_truncate_folio() may actually create a
post-eof folio in order to perform zeroing. To avoid this problem,
update range_has_data() to filter out unwritten blocks from folio
creation and partial zeroing.

Even though we should never create folios beyond EOF like this, the
mere existence of such folios is not necessarily a fatal error. Fix
up the truncate code to warn about this condition and not overflow
the sector array and possibly crash the system. The addition of this
warning without the corresponding unwritten extent fix has shown
that various other fstests are able to reproduce this problem fairly
frequently, but often in ways that doesn't necessarily result in a
kernel panic or a change in user observable behavior, and therefore
the problem goes undetected.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

4ad6aa46

bcachefs: Enable large folios · 550a6a49
Kent Overstreet authored Mar 19, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
550a6a49

bcachefs: Check for folios that don't have bch_folio attached · 34fdcf06

Kent Overstreet authored Mar 27, 2023

With large folios, it's now incidentally possible to end up with a
clean, uptodate folio in the page cache that doesn't have a bch_folio
attached, if a folio has to be split.

This patch fixes __bch2_truncate_folio() to check for this; other code
paths appear to handle it.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

34fdcf06

bcachefs: bch2_readahead() large folio conversion · 9567413c

Kent Overstreet authored Mar 17, 2023

Readahead now uses the new filemap_get_contig_folios_d() helper.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

9567413c