Commits · f9ed4bb6060526772c7e65a0f96950f5b0cb9774 · nexedi / MariaDB

25 Dec, 2022 40 commits

cherry-pick from 10.10: fix failing galera test · f9ed4bb6
Sergei Golubchik authored Dec 20, 2022

f9ed4bb6
ColumnStore still needs mysql* aliases · 75b9aace
Sergei Golubchik authored Dec 24, 2022

75b9aace

Fixed wrong selectivity calculation in table_after_join_selectivity() · 9f95e77a

Monty authored Dec 21, 2022

The old code counted selectivity double in case of queries like:
WHERE key_part1=1 and key_part2 < 100
if the optimizer would decide to use a REF access on key_part1.

The new code in best_access_path() that changes REF access to RANGE
if the RANGE key is longer makes this issue less likely to happen.

I was not able to create a test case for 11.0, however if one ports this
patch to a MariaDB version without the change of REF to RANGE, the
selectivity will be counted double.

9f95e77a

Cache file->index_flags(index, 0, 1) in table->key_info[index].index_flags · e8c3b6a5

Monty authored Dec 20, 2022

The reason for this is that we call file->index_flags(index, 0, 1)
multiple times in best_access_patch()when optimizing a table.
For example, in InnoDB, the calls is not trivial (4 if's and 2 assignments)
Now the function is inlined and is just a memory reference.

Other things:
- handler::is_clustering_key() and pk_is_clustering_key() are now inline.
- Added TABLE::can_use_rowid_filter() to simplify some code.
- Test if we should use a rowid_filter only if can_use_rowid_filter() is
  true.
- Added TABLE::is_clustering_key() to avoid a memory reference.
- Simplify some code using the fact that HA_KEYREAD_ONLY is true implies
  that HA_CLUSTERED_INDEX is false.
- Added DBUG_ASSERT to TABLE::best_range_rowid_filter() to ensure we
  do not call it with a clustering key.
- Reorginized elements in struct st_key to get better memory alignment.
- Updated ha_innobase::index_flags() to not have
  HA_DO_RANGE_FILTER_PUSHDOWN for clustered index

e8c3b6a5

Updated some tests for --valgrind · abe6fb6f

Monty authored Dec 20, 2022

- Increased timeout for binlog_mysqlbinlog_raw_flush.test.
  The old timeout was not enough when running with --valgrind
- Disabled ssl_timeout for --valgrind as it times out
- Disabled binlog_truncate_multi_engine for --valgrind as it does restarts

abe6fb6f

Fixed 'undefined variable' error in mtr · 471098d2
Monty authored Dec 20, 2022
```
This could happen if mtr_grab_file() returned empty (happened to me)
```
471098d2
Make tests work with --view-protocol · d4303e25
Sergei Petrunia authored Dec 16, 2022

d4303e25
Stabilize rocksdb.rocksdb test. · 03571b21
Sergei Petrunia authored Dec 15, 2022

03571b21

MDEV-21095: Make Optimizer Trace support Index Condition Pushdown · bbfa6d06

Sergei Petrunia authored Dec 08, 2022

Fixes over previous patches: do tracing of attached conditions
close to where we generate them.

Fix the tracing code to print the right conditions.

bbfa6d06

MDEV-21092,MDEV-21095,MDEV-29997: Optimizer Trace for index condition... · 346cc0f7

Rex authored Dec 02, 2022

MDEV-21092,MDEV-21095,MDEV-29997: Optimizer Trace for index condition pushdown, partition pruning, exists-to-in

        Add Optimizer Tracing for:
        - Index Condition Pushdown
        - Partition Pruning
        - Exists-to-IN optimization

346cc0f7

Stabilize engines/iuds.type_bit_iuds test · 905ff4d0
Sergei Petrunia authored Dec 14, 2022
```
Make sure the queries use the intended query plan
```
905ff4d0
Remove mysql-test/suite/versioning/r/select,trx_id.rdiff which is empty · 439e406f
Sergei Petrunia authored Dec 14, 2022
```
This seems to confuse windows.
```
439e406f
Update columnstore to include the patch to compile with the new cost model APIs · 16cf2e47
Sergei Petrunia authored Dec 14, 2022

16cf2e47

Removed "<select expression> INTO <destination>" deprication. · 87a409c5

Monty authored Dec 12, 2022

This was done after discussions with Igor, Sanja and Bar.

The main reason for removing the deprication was to ensure that MariaDB
is always backward compatible whenever possible.

Other things:
- Added statistics counters, mainly for the feedback plugin.
  - INTO OUTFILE
  - INTO variable
  - If INTO is using the old syntax (end of query)

87a409c5

Removed diff dates from rdiff files · 28887c15
Monty authored Dec 02, 2022

28887c15

In best_access_path() change record_count to 1.0 if its less than 1.0. · 1ad4a48c

Monty authored Dec 02, 2022

In essence this means that we expect the user query to have at least
one matching row in the end.
This change will not affect the estimated rows for the plan, but will
ensure that the cost for adding a table is not neglected because of
record count being too low.

The reasons for this is that if we have table combination that
together has a very high selectivity then join record_count could
become very low (close to 0)

This would cause costs for all future tables to be so small that they
are irrelevant for the rest of the plan.
This has been shown to be the case in some performance benchmarks and
in a few mtr tests.

There is also still a problem in selectivity calculations as joining two
tables in different order causes a different estimation of total rows.
This can be seen in selectivity_innodb.test, test 'Q20' where joining
nation,supplier is expecting 1.111 rows_out while joining supplier,nation
is expecting 0.04 rows_out.

The reason for 0.04 is that the optimizer estimates 'supplier' to have
10 matching rows, and joining with nation (eq_ref) has 1 row. However
selectivity of n_name = 'UNITED STATES' makes the optimizer things
that there will be only 0.04 matching rows.

This patch avoids this "too low row count" to affect cost
caclulations.

1ad4a48c

Changed some startup warnings to notes · 3e878149

Monty authored Nov 28, 2022

- Changed 'WARNING' of type "You need to use --log-bin to make ... work"
  to 'Note'
- Only print startup Notes if log_warnings >= 4

3e878149

Remove strlen() from Item::cleanup · cacf8044
Monty authored Nov 25, 2022

cacf8044

Do not give warnings about #rocksdb directory information_schema · 0cfd49b6

Monty authored Nov 25, 2022

"select * from information_schema.tables limit 1" was giving the following
warning in the log:

[ERROR] Invalid (old?) table or database name '#rocksdb'

0cfd49b6

MDEV-30032: EXPLAIN FORMAT=JSON output: part #2: print 'loops'. · edabbeb6
Sergei Petrunia authored Nov 21, 2022

edabbeb6
MDEV-30032: EXPLAIN FORMAT=JSON output: print costs · 465fec68
Sergei Petrunia authored Nov 19, 2022
```
Basic printout for join and table execution costs.
```
465fec68
Change BUILD scripts to use wolfss by default · 0c89ebaa
Monty authored Nov 22, 2022

0c89ebaa

Changed a rule to be cost based in test_if_cheaper_ordering · a00b5d3c

Monty authored Nov 22, 2022

- Simplified test by setting read_time=DBL_MAX at start of loop if
  FORCE INDEX is used
- No need to test for 'group by' as the cost compare should handle it.
- Only one test change where index scan was replaced with table scan
 (correct)

a00b5d3c

Simple cleanup of removing QQ comments from sql_select.cc · 018df60c
Monty authored Nov 22, 2022
```
- The comment in test_if_skip_sort_order was removed together with
  a not needed test of 'select'
```
018df60c
Added "override" to ha_heap.h, ha_myisam.h, ha_myisammrg.h and ha_sequence.h · b3081905
Monty authored Nov 21, 2022
```
Added override to a few functions in ha_partition.h
```
b3081905
Change default of histogram_type to JSON_HB · eb021f38
Monty authored Nov 18, 2022

eb021f38

Fixed bug in Aria with aria_log files that are exactly 8K · ad5465bc

Monty authored Nov 14, 2022

In the case one has an old Aria log file that ands with a Aria checkpoint
and the server restarts after next recovery, just after created a
new Aria log file (of 8K), the Aria recovery code would abort.
If one would try to delete all Aria log files after this (but not the
aria_control_file), the server would crash during recovery.

The problem was that translog_get_last_page_addr() would regard a log file
of exactly 8K as illegal and the rest of the code could not handle this
case.

Another issue was that if there was a crash directly after the log file
head was written to the next page, the code in translog_get_next_chunk()
would crash.

This patch fixes most of the issues, but not all. For Sanja to look at!

Things fixed:
- Added code to ignore 8K log files.
- Removed ASSERT in translog_get_next_chunk() that checks if page only
  contains the log page header.

ad5465bc

Small improvements to aria recovery · d4e06f42

Monty authored Nov 09, 2022

I spent 4 hours on work and 12 hours of testing to try to find
the reason for aria crashing in recovery when starting a new test,
in which case the 'data directory' should be a copy of "install.db",
but aria_log.00000001 content was not correct.

The following changes are mostly done to make it a bit easier to find out
more in case of future similar crashes:

- Mark last_checkpoint_lsn volatile (safety).
- Write checkpoint message to aria_recovery.trace
- When compling with DBUG and with HAVE_DBUG_TRANSLOG_SRC,
  use checksum's for Aria log pages. We cannot have it on by default
  for DBUG servers yet as there is bugs when changing CRC between
  restarts.
- Added a message to mtr --verbose when copying the data directory.
- Removed extra linefeed in Aria recovery message (cleanup)

d4e06f42

Added rowid_filter support to Aria · 2e44e613

Monty authored Nov 16, 2022

This includes:
- cleanup and optimization of filtering and pushdown engine code.
- Adjusted costs for rowid filters (based on extensive testing
  and profiling).

This made a small two changes to the handler_rowid_filter_is_active()
API:
- One should not call it with a zero pointer!
- One does not need to call handler_rowid_filter_is_active() for every
  row anymore. It is enough to check if filter is active by calling it
  call it during index_init() or when handler::rowid_filter_changed()
  is called

The changes was to avoid unnecessary function calls and checks if
pushdown conditions and rowid_filter is not used.

Updated costs for rowid_filter_lookup() to be closer to reality.
The old cost was based only on rowid_compare_cost. This is now
changed to take into account the overhead in checking the rowid.

Changed the Range_rowid_filter class to use DYNAMIC_ARRAY directly
instead of Dynamic_array<>. This was done to be able to use the new
append_dynamic() functions which gives a notable speed improvment
compared to the old code.  Removing the abstraction also makes
the code easier to understand.

The cost of filtering is now slightly lower than before, which
is reflected in some test cases that is now using rowid filters.

2e44e613

Set thd->query() for internal (startup) transactions · ef4c23dd

Monty authored Nov 09, 2022

This helps with debugging as 'Query: ' in DBUG traces will show something
useful, for internal transactions, instead of just "".

ef4c23dd

Added MARIADB_NEW_COST_MODEL for ColumnStore to detect new cost model · 2d59e6ed
Sergei Petrunia authored Nov 10, 2022

2d59e6ed
Don't do zerofill of Aria table if it's already zerofilled · 77c4b114
Monty authored Nov 02, 2022
```
This will speed up using tables that are already zerofilled
with aria_chk --zerofill.
```
77c4b114
MDEV-30059: Optimizer Trace: plan_prefix should be a comma-separated-list · 3281970d
Sergei Petrunia authored Nov 21, 2022

3281970d

Added test cases for preceding test · 709c2207

Monty authored Oct 04, 2022

This includes all test changes from
"Changing all cost calculation to be given in milliseconds"
and forwards.

Some of the things that caused changes in the result files:

- As part of fixing tests, I added 'echo' to some comments to be able to
  easier find out where things where wrong.
- MATERIALIZED has now a higher cost compared to X than before. Because
  of this some MATERIALIZED types have changed to DEPENDEND SUBQUERY.
  - Some test cases that required MATERIALIZED to repeat a bug was
    changed by adding more rows to force MATERIALIZED to happen.
- 'Filtered' in SHOW EXPLAIN has in many case changed from 100.00 to
  something smaller. This is because now filtered also takes into
  account the smallest possible ref access and filters, even if they
  where not used. Another reason for 'Filtered' being smaller is that
  we now also take into account implicit filtering done for subqueries
  using FIRSTMATCH.
  (main.subselect_no_exists_to_in)
  This is caluculated in best_access_path() and stored in records_out.
- Table orders has changed because more accurate costs.
- 'index' and 'ALL' for small tables has changed to use 'range' or
   'ref' because of optimizer_scan_setup_cost.
- index can be changed to 'range' as 'range' optimizer assumes we don't
  have to read the blocks from disk that range optimizer has already read.
  This can be confusing in the case where there is no obvious where clause
  but instead there is a hidden 'key_column > NULL' added by the optimizer.
  (main.subselect_no_exists_to_in)
- Scan on primary clustered key does not report 'Using Index' anymore
  (It's a table scan, not an index scan).
- For derived tables, the number of rows is now 100 instead of 2,
  which can be seen in EXPLAIN.
- More tests have "Using index for group by" as the cost of this
  optimization is now more correct (lower).
- A primary key could be preferred for a normal key, even if it would
  access more rows, as it's faster to do 1 lokoup and 3 'index_next' on a
  clustered primary key than one lookup trough a secondary.
  (main.stat_tables_innodb)

Notes:

- There was a 4.7% more calls to best_extension_by_limited_search() in
  the main.greedy_optimizer test.  However examining the test results
  it looked that the plans where slightly better (eq_ref where more
  chained together) so I assume this is ok.
- I have verified a few test cases where there was notable/unexpected
  changes in the plan and in all cases the new optimizer plans where
  faster.  (main.greedy_optimizer and some others)

709c2207

Added range_index to 'range' optimizer_trace output · 80374509
Monty authored Nov 24, 2022
```
Other things:
- Renamed "rowid_filter_key" to "rowid_filter_index" to keep things
  consistent
```
80374509

Fix bug in WITH ties · 8d241abb

Vicențiu Ciorbaru authored Nov 26, 2022

The old code had a bug when the normal sorting code where
where eliminated as part of "Using index for group-by" optimization.
The effect was that the result contained more rows than expected

8d241abb

MDEV-29677 Wrong result with join query and innodb fulltext search · efa070a8

Monty authored Nov 07, 2022

InnoDB FTS scan was used by a subquery. A subquery execution may start
a table read and continue until it finds the first matching record
combination. This can happen before the table read returns EOF.

The next time the subquery is executed, it will start another table read.
InnoDB FTS table read fails to re-initialize its data structures in this
scenario and will try to continue the scan started at the first execution.

Fixed by ha_innobase::ft_init() to stop the FTS scan if there is one.

Author: Sergei Petrunia <sergey@mariadb.com>
Reviewer: Monty

efa070a8

Fixes for 'Filtering' · 39d3e3b5

Monty authored Oct 31, 2022

- table_after_join_selectivity() should use records_init (new bug)
- get_examined_rows() changed to double to get similar results
  as in MariaDB 10.11
- Fixed bug where table_after_join_selectivity() did not correct
  selectivity in the case where a RANGE is used instead of a REF.
  This can happen if the range can use more key_parts than the REF.
  WHERE key_part1=10 and key_part2 < 10

Other things:
- Use JT_RANGE instead of JT_ALL for RANGE access in all parts of the code.
  Before we used JT_ALL for RANGE.
- Force RANGE be used in best_access_path() if the range used more key
  parts than ref. In the original code, this was done much later in
  make_join_select)(). However we need to know in
  table_after_join_selectivity() if we have used RANGE or not.
- Added more information about filtering to optimizer_trace.

39d3e3b5

Updated number of expected rows from 2 to 100 for information_schema tables · 2ae0239d

Monty authored Oct 28, 2022

The reason is that 2 is usually way to low and as information_schema
tables may have implicit locks when accessing rows, it is better that
the optimizer doesn't think that these tables are 'very small and fast'.

This change will affect a very small set of test cases.

2ae0239d

Added optimizer_trace info for index_intersects · 7a63a3d6
Monty authored Oct 28, 2022

7a63a3d6