Commits · d79c3f3297f8f2f6bd44a52c95dbae1d06bd3976 · nexedi / MariaDB

11 Dec, 2020 1 commit

MDEV-24353: Adding GROUP BY slows down a query · d79c3f32

Varun Gupta authored Dec 09, 2020

A heuristic in best_access_path says that if for an index
ref access involved key parts which are greater than equal to that
for range access, then range access should not be considered.
The assumption made by this heuristic does not hold when
the range optimizer opted to use the group-by min-max optimization.
So the fix here would be to not consider the heuristic if
the range optimizer picked the usage of group-by min-max
optimization.

d79c3f32

09 Dec, 2020 4 commits

Merge 10.5 into 10.6 · be4d2665
Marko Mäkelä authored Dec 09, 2020

be4d2665

Remove unused DBUG_EXECUTE_IF "ignore_punch_hole" · 0c7c4492

Marko Mäkelä authored Dec 09, 2020

Since commit ea21d630 we
conditionally define a variable that only plays a role on
systems that support hole-punching (explicit creation of sparse files).
However, that broke debug builds on such systems.

It turns out that the debug_dbug label "ignore_punch_hole" is
not at all used in MariaDB server. It would be covered by
the MySQL 5.7 test innodb.table_compress. (Note: MariaDB 10.1
implemented page_compressed tables before something comparable
appeared in MySQL 5.7.)

0c7c4492

Merge 10.5 into 10.6 · ca821692
Marko Mäkelä authored Dec 09, 2020

ca821692

MDEV-12227 Defer writes to the InnoDB temporary tablespace · 5eb53955

Marko Mäkelä authored Dec 09, 2020

The flushing of the InnoDB temporary tablespace is unnecessarily
tied to the write-ahead redo logging and redo log checkpoints,
which must be tied to the page writes of persistent tablespaces.

Let us simply omit any pages of temporary tables from buf_pool.flush_list.
In this way, log checkpoints will never incur any 'collateral damage' of
writing out unmodified changes for temporary tables.

After this change, pages of the temporary tablespace can only be written
out by buf_flush_lists(n_pages,0) as part of LRU eviction. Hopefully,
most of the time, that code will never be executed, and instead, the
temporary pages will be evicted by buf_release_freed_page() without
ever being written back to the temporary tablespace file.

This should improve the efficiency of the checkpoint flushing and
the buf_flush_page_cleaner thread.

Reviewed by: Vladislav Vaintroub

5eb53955

08 Dec, 2020 3 commits

Fix -Wunused-but-set-variable · ea21d630
Marko Mäkelä authored Dec 08, 2020

ea21d630

MDEV-24369 Page cleaner sleeps despite innodb_max_dirty_pages_pct_lwm being exceeded · f0c295e2

Marko Mäkelä authored Dec 08, 2020

MDEV-24278 improved the page cleaner so that it will no longer wake up
once per second on an idle server. However, with innodb_adaptive_flushing
(the default) the function page_cleaner_flush_pages_recommendation()
could initially return 0 even if there is work to do.

af_get_pct_for_dirty(): Remove. Based on a comment here, it appears
that an initial intention of innodb_max_dirty_pages_pct_lwm=0.0
(the default value) was to disable something. That ceased to hold in
MDEV-23855: the value is a pure threshold; the page cleaner will not
perform any work unless the threshold is exceeded.

page_cleaner_flush_pages_recommendation(): Add the parameter dirty_blocks
to ensure that buf_pool.flush_list will eventually be emptied.

f0c295e2

MDEV-24351: S3, same-backend replication: Dropping a table on master... · 6859e80d

Sergei Petrunia authored Dec 08, 2020

..causes error on slave.
Cause: if the master doesn't have the frm file for the table,
DROP TABLE code will call ha_delete_table_force() to drop the table
in all available storage engines.
The issue was that this code path didn't check for
HTON_TABLE_MAY_NOT_EXIST_ON_SLAVE flag for the storage engine,
and so did not add "... IF EXISTS" to the statement that's written
to the binary log.  This can cause error on the slave when it tries to
drop a table that's already gone.

6859e80d

07 Dec, 2020 1 commit
- Simplify clang workarounds. · 3ee24b23
  Vladislav Vaintroub authored Dec 07, 2020
  
  3ee24b23
04 Dec, 2020 2 commits

MDEV-24350 buf_dblwr unnecessarily uses memory-intensive srv_stats counters · 83591a23

Marko Mäkelä authored Dec 04, 2020

The counters in srv_stats use std::atomic and multiple cache lines per
counter. This is an overkill in a case where a critical section already
exists in the code. A regular variable will work just fine, with much
smaller memory bus impact.

83591a23

MDEV-24348 InnoDB shutdown hang with innodb_flush_sync=0 · aa0e3805

Marko Mäkelä authored Dec 04, 2020

This hang was caused by MDEV-23855, and we failed to fix it in
MDEV-24109 (commit 4cbfdeca).

When buf_flush_ahead() is invoked soon before server shutdown
and the non-default setting innodb_flush_sync=OFF is in effect
and the buffer pool contains dirty pages of temporary tables,
the page cleaner thread may remain in an infinite loop
without completing its work, thus causing the shutdown to hang.

buf_flush_page_cleaner(): If the buffer pool contains no
unmodified persistent pages, ensure that buf_flush_sync_lsn= 0
will be assigned, so that shutdown will proceed.

The test case is not deterministic. On my system, it reproduced
the hang with 95% probability when running multiple instances
of the test in parallel, and 4% when running single-threaded.

Thanks to Eugene Kosov for debugging and testing this.

aa0e3805

03 Dec, 2020 13 commits

MDEV-24142: Avoid block_lock alignment loss on 64-bit systems · e9f33b77

Marko Mäkelä authored Dec 03, 2020

sux_lock::recursive: Move right after the 32-bit sux_lock::lock.
This will reduce sizeof(block_lock) from 24 to 16 bytes on
64-bit systems with CMAKE_BUILD_TYPE=RelWithDebInfo. This may be
significant, because there will be one buf_block_t::lock for each
buffer pool page descriptor.

We still have some potential for savings, with sizeof(buf_page_t)==112
and sizeof(buf_block_t)==184 on a GNU/Linux AMD64 system.

Note: On GNU/Linux AMD64, sizeof(index_lock) remains 32 bytes
(16 with PLUGIN_PERFSCHEMA=NO) even tough it would fit in 24 bytes.
This is because sizeof(srw_lock) includes 4 bytes of padding
(to 16 bytes) that index_lock_t::recursive cannot reuse. So,
in total 4+4 bytes will be lost to padding. This is rather
insignificant compared to sizeof(dict_index_t)==400.

e9f33b77

Fixed usage of not initialized memory in LIKE ... ESCAPE · 6033cc85

Monty authored Dec 03, 2020

This was noticed wben running "mtr --valgrind main.precedence"

The problem was that Item_func_like::escape could be left unitialized
when used with views combined with UNIONS like in:

create or replace view v1 as select 2 LIKE 1 ESCAPE 3 IN (SELECT 0 UNION SELECT 1), 2 LIKE 1 ESCAPE (3 IN (SELECT 0 UNION SELECT 1)), (2 LIKE 1 ESCAPE 3) IN (SELECT 0 UNION SELECT 1);

The above query causes in fix_escape_item()
escape_item->const_during_execution() to be true
and
escape_item->const_item() to be false

in which case 'escape' is never calculated.

The fix is to make the main logic of fix_escape_item() out to a
separate function and call that function once in Item.

Other things:
- Reorganized fields in Item_func_like class to make it more compact

6033cc85

MDEV-24142: Remove INFORMATION_SCHEMA.INNODB_MUTEXES · ba2d45dc

Marko Mäkelä authored Dec 03, 2020

Let us remove sux_lock::waits and the associated bookkeeping.
Starting with commit 1669c889
the PERFORMANCE_SCHEMA instrumentation interface is keeping
track of lock waits.

The view INFORMATION_SCHEMA.INNODB_MUTEXES only exported counts
of rw-lock waits.

Also, SHOW ENGINE INNODB MUTEX will no longer export any information
about rw-locks.

ba2d45dc

MDEV-24142: Remove __FILE__,__LINE__ related to buf_block_t::lock · 9702be2c
Marko Mäkelä authored Dec 03, 2020

9702be2c

MDEV-24142: Remove the LatchDebug interface to rw-locks · ac028ec5

Marko Mäkelä authored Dec 03, 2020

The latching order checks for rw-locks have not caught many bugs
in the past few years and they are greatly complicating the code.

Last time the debug checks were useful was in
commit 59caf2c3 (MDEV-13485).

The B-tree hang MDEV-14637 was not caught by LatchDebug,
because the granularity of the checks is not sufficient
to distinguish the levels of non-leaf B-tree pages.

The interface was already made dead code by the grandparent
commit 03ca6495.

ac028ec5

MDEV-24308: Windows improvements · 06efef4b
Marko Mäkelä authored Nov 30, 2020
```
This reverts commit e34e53b5
and defines os_thread_sleep() is a macro on Windows.
```
06efef4b

MDEV-24142: Replace InnoDB rw_lock_t with sux_lock · 03ca6495

Marko Mäkelä authored Dec 03, 2020

InnoDB buffer pool block and index tree latches depend on a
special kind of read-update-write lock that allows reentrant
(recursive) acquisition of the 'update' and 'write' locks
as well as an upgrade from 'update' lock to 'write' lock.
The 'update' lock allows any number of reader locks from
other threads, but no concurrent 'update' or 'write' lock.

If there were no requirement to support an upgrade from 'update'
to 'write', we could compose the lock out of two srw_lock
(implemented as any type of native rw-lock, such as SRWLOCK on
Microsoft Windows). Removing this requirement is very difficult,
so in commit f7e7f487d4b06695f91f6fbeb0396b9d87fc7bbf we
implemented an 'update' mode to our srw_lock.

Re-entrant or recursive locking is mostly needed when writing or
freeing BLOB pages, but also in crash recovery or when merging
buffered changes to an index page. The re-entrancy allows us to
attach a previously acquired page to a sub-mini-transaction that
will be committed before whatever else is holding the page latch.

The SUX lock supports Shared ('read'), Update, and eXclusive ('write')
locking modes. The S latches are not re-entrant, but a single S latch
may be acquired even if the thread already holds an U latch.

The idea of the U latch is to allow a write of something that concurrent
readers do not care about (such as the contents of BTR_SEG_LEAF,
BTR_SEG_TOP and other page allocation metadata structures, or
the MDEV-6076 PAGE_ROOT_AUTO_INC). (The PAGE_ROOT_AUTO_INC field
is only updated when a dict_table_t for the table exists, and only
read when a dict_table_t for the table is being added to dict_sys.)

block_lock::u_lock_try(bool for_io=true) is used in buf_flush_page()
to allow concurrent readers but no concurrent modifications while the
page is being written to the data file. That latch will be released
by buf_page_write_complete() in a different thread. Hence, we use
the special lock owner value FOR_IO.

The index_lock::u_lock() improves concurrency on operations that
involve non-leaf index pages.

The interface has been cleaned up a little. We will use
x_lock_recursive() instead of x_lock() when we know that a
lock is already held by the current thread. Similarly,
a lock upgrade from U to X is only allowed via u_x_upgrade()
or x_lock_upgraded() but not via x_lock().

We will disable the LatchDebug and sync_array interfaces to
InnoDB rw-locks.

The SEMAPHORES section of SHOW ENGINE INNODB STATUS output
will no longer include any information about InnoDB rw-locks,
only TTASEventMutex (cmake -DMUTEXTYPE=event) waits.
This will make a part of the 'innotop' script dead code.

The block_lock buf_block_t::lock will not be covered by any
PERFORMANCE_SCHEMA instrumentation.

SHOW ENGINE INNODB MUTEX and INFORMATION_SCHEMA.INNODB_MUTEXES
will no longer output source code file names or line numbers.
The dict_index_t::lock will be identified by index and table names,
which should be much more useful. PERFORMANCE_SCHEMA is lumping
information about all dict_index_t::lock together as
event_name='wait/synch/sxlock/innodb/index_tree_rw_lock'.

buf_page_free(): Remove the file,line parameters. The sux_lock will
not store such diagnostic information.

buf_block_dbg_add_level(): Define as empty macro, to be removed
in a subsequent commit.

Unless the build was configured with cmake -DPLUGIN_PERFSCHEMA=NO
the index_lock dict_index_t::lock will be instrumented via
PERFORMANCE_SCHEMA. Similar to
commit 1669c889
we will distinguish lock waits by registering shared_lock,exclusive_lock
events instead of try_shared_lock,try_exclusive_lock.
Actual 'try' operations will not be instrumented at all.

rw_lock_list: Remove. After MDEV-24167, this only covered
buf_block_t::lock and dict_index_t::lock. We will output their
information by traversing buf_pool or dict_sys.

03ca6495

MDEV-24142 preparation: Add srw_mutex and srw_lock::u_lock() · d46b4248

Marko Mäkelä authored Dec 03, 2020

The PERFORMANCE_SCHEMA insists on distinguishing read-update-write
locks from read-write locks, so we must add
template<bool support_u_lock> in rd_lock() and wr_lock() operations.

rd_lock::read_trylock(): Add template<bool prioritize_updater=false>
which is used by the srw_lock_low::read_lock() loop. As long as
an UPDATE lock has already been granted to some thread, we will grant
subsequent READ lock requests even if a waiting WRITE lock request
exists. This will be necessary to be compatible with existing usage
pattern of InnoDB rw_lock_t where the holder of SX-latch (which we
will rename to UPDATE latch) may acquire an additional S-latch
on the same object. For normal read-write locks without update operations
this should make no difference at all, because the rw_lock::UPDATER
flag would never be set.

d46b4248

MDEV-24167: Stabilize perfschema.sxlock_func · 3872e585

Marko Mäkelä authored Dec 03, 2020

The extension of the test perfschema.sxlock_func in
commit 1669c889
turned out to be unstable.

Let us filter out purge_sys.latch (trx_purge_latch) from the output,
because it might happen that the purge tasks will not be executed
during the test execution.

3872e585

MDEV-24167 fixup: Improve the PERFORMANCE_SCHEMA instrumentation · 1669c889

Marko Mäkelä authored Dec 03, 2020

Let us try to avoid code bloat for the common case that
performance_schema is disabled at runtime, and use
ATTRIBUTE_NOINLINE member functions for instrumented latch acquisition.

Also, let us distinguish lock waits from non-contended lock requests
by using write_lock,read_lock for the requests that lead to waits,
and try_write_lock,try_read_lock for the wait-free lock acquisitions.
Actual 'try' operations are not being instrumented at all.

1669c889

MDEV-24167 fixup: Avoid hangs in SRW_LOCK_DUMMY · 260161fc

Marko Mäkelä authored Dec 03, 2020

In commit 1fdc161d we introduced
a mutex-and-condition-variable based fallback implementation
for platforms that lack a futex system call. That implementation
is prone to hangs.

Let us use separate condition variables for shared and exclusive requests.

260161fc

Merge 10.5 into 10.6 · a13fac9e
Marko Mäkelä authored Dec 03, 2020

a13fac9e

MDEV-22929 fixup: root_name() clash with clang++ <fstream> · f146969f

Marko Mäkelä authored Dec 03, 2020

The clang++ -stdlib=libc++ header file <fstream> depends on
<filesystem> that defines a member function path::root_name(),
which conflicts with the rather unused #define root_name()
that had been introduced in
commit 7c58e97b.

Because an instrumented -stdlib=libc++ (rather than the default
-stdlib=libstdc++) is easier to build for a working -fsanitize=memory
(cmake -DWITH_MSAN=ON), let us remove the conflicting #define for now.

f146969f

02 Dec, 2020 5 commits
- MDEV-24295: Fix the non-clang build · f3a58ed8
  Marko Mäkelä authored Dec 02, 2020
```
Sorry, only tested commit 4174fc1a
on clang. Other compilers do not define __has_feature().
```
  f3a58ed8
- MDEV-24295: Fix the WITH_MSAN build · 4174fc1a
  Marko Mäkelä authored Dec 02, 2020
```
For some reason, commit 5bb5d4ad
made clang++-11 unhappy about a constexpr declaration.
```
  4174fc1a
- MDEV-20051 fixup: Correct galera.galera_defaults result · 9b725f9a
  Marko Mäkelä authored Dec 02, 2020
```
For some reason, the test was never adjusted for
commit e6a50e41.
```
  9b725f9a
- Merge 10.4 into 10.5 · 6a1e655c
  Marko Mäkelä authored Dec 02, 2020
  
  6a1e655c
- MDEV-15532 after-merge fixes from Monty · 24ec8eaf
  Marko Mäkelä authored Dec 02, 2020
```
The Galera tests were massively failing with debug assertions.
```
  24ec8eaf
01 Dec, 2020 8 commits

Merge 10.3 into 10.4 · 589cf8db
Marko Mäkelä authored Dec 01, 2020

589cf8db
MDEV-22929 MariaBackup option to report and/or continue when corruption is encountered · e30a05f4
Vlad Lesin authored Dec 01, 2020
```
Post-push Windows compilation errors fix.
```
e30a05f4
MDEV-24167 fixup: Improve perfschema.sxlock_func test · e28d9c15
Marko Mäkelä authored Dec 01, 2020

e28d9c15

After merge fixes · 7edfed63

Monty authored Dec 01, 2020

Change thd->mdl_context.release_transactional_locks() to
thd->mdl_release_transactional_locks()

7edfed63

MDEV-24323 Crash on recovery after kill during instant ADD COLUMN · 73f34336
Marko Mäkelä authored Dec 01, 2020
```
row_undo_ins_parse_undo_rec(): Do not try to read non-existing
virtual column information for the metadata record.
```
73f34336
Merge 10.2 into 10.3 · 81ab9ea6
Marko Mäkelä authored Dec 01, 2020

81ab9ea6
MDEV-21962 fixup: Remove buf_pool_contains_zip() · e76e1288
Marko Mäkelä authored Dec 01, 2020
```
The replacement is buf_pool.contains_zip().
```
e76e1288

MDEV-22929 MariaBackup option to report and/or continue when corruption is encountered · e6b3e38d

Vlad Lesin authored Aug 20, 2020

The new option --log-innodb-page-corruption is introduced.

When this option is set, backup is not interrupted if innodb corrupted
page is detected. Instead it logs all found corrupted pages in
innodb_corrupted_pages file in backup directory and finishes with error.

For incremental backup corrupted pages are also copied to .delta file,
because we can't do LSN check for such pages during backup,
innodb_corrupted_pages will also be created in incremental backup
directory.

During --prepare, corrupted pages list is read from the file just after
redo log is applied, and each page from the list is checked if it is allocated
in it's tablespace or not. If it is not allocated, then it is zeroed out,
flushed to the tablespace and removed from the list. If all pages are removed
from the list, then --prepare is finished successfully and
innodb_corrupted_pages file is removed from backup directory. Otherwise
--prepare is finished with error message and innodb_corrupted_pages contains
the list of the pages, which are detected as corrupted during backup, and are
allocated in their tablespaces, what means backup directory contains corrupted
innodb pages, and backup can not be considered as consistent.

For incremental --prepare corrupted pages from .delta files are applied
to the base backup, innodb_corrupted_pages is read from both base in
incremental directories, and the same action is proceded for corrupted
pages list as for full --prepare. innodb_corrupted_pages file is
modified or removed only in base directory.

If DDL happens during backup, it is also processed at the end of backup
to have correct tablespace names in innodb_corrupted_pages.

e6b3e38d

30 Nov, 2020 3 commits

MDEV 15532 Assertion `!log->same_pk' failed in row_log_table_apply_delete · 828471cb

Monty authored Nov 30, 2020

The reason for the failure is that
thd->mdl_context.release_transactional_locks()
was called after commit & rollback even in cases where the current
transaction is still active.

For 10.2, 10.3 and 10.4 the fix is simple:
- Replace all calls to thd->mdl_context.release_transactional_locks() with
  thd->release_transactional_locks(). The thd function will only call
  the mdl_context function if there are no active transactional locks.
  In 10.6 we will better fix where we will change the return value for
  some trans_xxx() functions to indicate if transaction did close the
  transaction or not. This will avoid the need of the indirect call.

Other things:
- trans_xa_commit() and trans_xa_rollback() will automatically
  call release_transactional_locks() if the transaction is closed.
- We can't do that for the other functions as the caller of many of these
  are doing additional work (like close_thread_tables) before calling
  release_transactional_locks().
- Added missing abort_result_set() and missing DBUG_RETURN in
  select_create::send_eof()
- Fixed wrong indentation in injector::transaction::commit()

828471cb

Fixed maria.create test · c5375764
Monty authored Nov 30, 2020

c5375764

MDEV-15532 Assertion `!log->same_pk' failed in row_log_table_apply_delete · a3531775

Monty authored Nov 30, 2020

The real fix for MDEV-15532 will be pushed into 10.2 and 10.6
This is an additional fix for 10.4.

In 10.4 trans_xa_detach was introduced.  However THD::cleanup() assumes
that after trans_xa_detach() is done, there is no registered transactions
anymore. In the 10.2 patch there will be an assert to ensure this, which
will cause 10.4 to fail.

The fix used is to reset the transaction flags in trans_xa_detach().

a3531775