Skip to content

Mysql string primary key partitioner - #38047

Draft
peterdukelarsen wants to merge 28 commits into
MaterializeInc:mainfrom
peterdukelarsen:pl/mysql-pk-partitioner
Draft

Mysql string primary key partitioner#38047
peterdukelarsen wants to merge 28 commits into
MaterializeInc:mainfrom
peterdukelarsen:pl/mysql-pk-partitioner

Conversation

@peterdukelarsen

@peterdukelarsen peterdukelarsen commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

Motivation

Part of SS-97.

Working on a faster way to identify partition boundaries for string primary keys.

Description

Adds code for partitioning the primary keys of a table with a single-column string primary key.

The core idea for partitioning without needing to understand MySQL's string sort order rules is to walk through string prefixes, in order, estimating the sizes and refining as you go based on table size estimates.

There are a few tweaks from a simple prefix estimate here worth calling out:

  1. Handling short strings, i.e. "a.*" touches most of the keys, but you have an "a" key in the table, how do you step deeper into the prefixes -- if you held onto the "a" prefix by itself it would represent most of the database, so you need an upper bound or something.
  2. Moving to use ranges for representing a sub-partition. Helps with point 1, but also helps if we want to split to the end of the current range... i.e. "a." has 1M records, your target bucket count is 500k, and "aa." has 250k, "ab." has 300k and everything greater than "ab." but still in "a*" is less than your target 500k you can exit early.
  3. Working around observed caps in estimate size -- for the 2.2B row table I've been experimenting with the estimates capped out at exactly 1/2 the approximate row count when they should have covered most/all of the table. This means the algorithm needs to be resilient to pretty severe undercounts on large tables. (TODO: should we skip this?).
  4. Breaking down the buckets smaller to build back up to more even partitions.

Verification

peterdukelarsen and others added 9 commits August 3, 2026 20:39
Add a probe module exposing KeyProber over a string key column of a
table: optimizer row count estimates for half-open key ranges via
EXPLAIN index dives, and first/next key prefix discovery, with all key
ordering done server-side under the column's own collation. Includes a
LIKE-pattern escaping helper so prefixes containing wildcard characters
match literally.

Covered by unit tests plus a live-MySQL test (opt-in via
MZ_TEST_MYSQL_URL) that exercises EXPLAIN estimates, LIKE escaping
against metacharacter keys, and range bounds on both probes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Add the MySql service to the cargo-test composition and hand its URL
to nextest as MZ_TEST_MYSQL_URL, following the pattern POSTGRES_URL and
METADATA_BACKEND_URL already use. Tests gated on the variable skip
locally when it is unset but panic when CI is set, matching the
timestamp oracle's tripwire so a wiring regression cannot silently
retire them. Tests sharing the server run concurrently under nextest,
so each must confine itself to a uniquely named scratch database it
creates itself.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Cover ULID keys (long shared timestamp prefix), UUID keys, keys built
from LIKE metacharacters, case-insensitive vs binary collations, and
stale table statistics. The stale statistics test pins down that range
estimates come from index dives on the real B-tree, so they stay
accurate even while information_schema.tables reports 0 rows. The
metacharacter test asserts the partition property of a prefix walk,
since a key shorter than the prefix length subsumes longer keys
sharing it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A key shorter than the prefix length names an exact key, and the LIKE
anchor skipped every key extending it, leaving such ranges unsplittable
at any depth. next_prefix now steps just past the exact key so its
extensions become prefixes of their own.

Tests: pin charset and collation explicitly everywhere, share table
setup through helpers, and reorganize so behavior-explaining tests lead
and helpers trail. New coverage: basic and case-insensitive traversal,
multibyte keys, EXPLAIN estimate sizing, and sargability asserted via
session handler counters, with a self-check that a non-sargable query
trips them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Move the broader live-MySQL test suite (case sensitivity, wildcards,
multibyte data, ULID/UUID keys, LIKE metacharacters, collations, stale
statistics, sargability) to a stacked follow-up branch so this PR stays
focused on the probing interface itself.
Nothing outside the probe module uses the LIKE escaping, keeping it
private leaves the NO_BACKSLASH_ESCAPES caveat an internal note rather
than a public contract.
@peterdukelarsen
peterdukelarsen force-pushed the pl/mysql-pk-partitioner branch from 6217a19 to 4be8bb5 Compare August 4, 2026 20:57
estimate_range_rows returns the optimizer estimate as an Option
instead of defaulting to 0, and its doc records observed accuracy on a
large static table. The next-prefix LIKE anchor declares an explicit
ESCAPE so the pattern no longer depends on the sql_mode default escape
character, making NO_BACKSLASH_ESCAPES sessions behave identically.
Comments document the charset conversion reasoning and the connection
charset assumption, and range end-bound SQL assembly is deduplicated
into a helper.
Take a concrete mysql_async::Conn instead of a Queryable generic.
Replace the inclusive lower bound with an exclusive one throughout, a
key exactly equal to a bound is skipped as a split point and its
extensions surface through the exclusive bound on re-splits.
Decompose the next-prefix probe into max_key_with_prefix plus
prefix_of_first_key_in_range, drop the client-side character counting
entirely, and rename the probes to say what they return. Privatize
explain_row_estimate, estimates are reached through KeyProber.
The anchor and seek probes read separate snapshots outside a
transaction, and an insert matching the prefix between them makes the
walk see the same prefix again instead of advancing.
Both prefix probes take max_prefix_length, replacing the mismatched
prefix_len and len, and docs say "up to" to match the shorter-key
behavior.
@peterdukelarsen
peterdukelarsen force-pushed the pl/mysql-pk-partitioner branch from 9461398 to 8ac7ad8 Compare August 5, 2026 20:16
upper_bound becomes upper_bound_exclusive to match
lower_bound_exclusive, and estimate_range_rows takes the same names.
Docs state that both bounds are exclusive.
@peterdukelarsen
peterdukelarsen force-pushed the pl/mysql-pk-partitioner branch from 8ac7ad8 to 640b74d Compare August 5, 2026 20:23
first becomes first_key_in_range and next becomes
first_row_not_matching_prefix, the probe method names minus the shared
prefix_of_ stem, short enough that every assertion stays a one-liner.
max_key_with_prefix splices range_filter output after its LIKE
condition instead of using a bespoke append helper. The TRUE fallback
makes that composition uniform, the clause stays a valid predicate
after WHERE or AND even with no bounds.
@peterdukelarsen
peterdukelarsen force-pushed the pl/mysql-pk-partitioner branch from 640b74d to 3acb957 Compare August 5, 2026 20:35
The test wrappers take the full probe method names, trading one-line
assertions for grep-identical naming, and the range_filter doc is
condensed.
@peterdukelarsen
peterdukelarsen force-pushed the pl/mysql-pk-partitioner branch from 3acb957 to 9f60c7a Compare August 5, 2026 20:49
estimate_range_rows accepts an open lower bound like the prefix
probes, and callers pass None for the start of the key space instead
of an empty string sentinel.
@peterdukelarsen
peterdukelarsen force-pushed the pl/mysql-pk-partitioner branch 2 times, most recently from defec50 to e1d8b37 Compare August 5, 2026 22:19
query_string mapped values that fail UTF-8 decoding to None, silently
ending prefix walks early. It now returns MySqlError::NonUtf8KeyValue
so callers can log the condition and fall back explicitly instead of
mistaking it for an exhausted range.
Restores the broader live test suite on top of the probe interface PR:
case-insensitive traversal, LIKE wildcard and metacharacter data,
multibyte and emoji keys, ULID and UUID primary keys, collation
behavior, stale table statistics, and sargability via Handler_read
session counters.
A latin1 table exercises the column-to-connection charset conversion,
including the exact-key step on a key that is one character but two
UTF-8 bytes. A binary key column pins the defensive behavior for
invalid UTF-8: decode failures read as "no next prefix" and end the
walk early instead of erroring. setup_table now derives the charset
from the collation name instead of hardcoding utf8mb4.
Discovers boundaries that split a table string primary key space into
per-worker ranges of roughly equal estimated row counts. Ranges the
optimizer estimates too large are recursively subdivided at each
distinct key prefix one character longer, probing through KeyProber,
then accumulated into per-worker buckets, so discovery costs EXPLAIN
index dives instead of an O(rows) index pass. Inaccurate estimates
skew bucket sizes but never correctness: any ordered boundary list
partitions the key space.

All key ordering happens server-side under the column collation. The
walk guards against non-advancing prefixes and caps children per split
so a misbehaving server cannot hang it. KeyProber steps past exact
keys shorter than the prefix length, so a lone short key among keys
extending it cannot leave a range unsplittable.

Also documents the caller contracts on like_prefix_pattern and
explain_row_estimate.
Drop MAX_DEPTH, MAX_CHILDREN_PER_SPLIT, and the non-advancing prefix
guard. On healthy data the walk terminates because child ranges shrink
and fresh estimates track them. The pathological cases (phantom
estimates, misbehaving servers) will be bounded by the per-table
request budget once it lands, rather than by per-mechanism caps.

Reformulate bucket sizing as a per-worker share divided by
BUCKETS_PER_WORKER, dropping the double-to-8 bucket floor for small
worker counts.
partition() becomes a composition of three stages: bucket_target_rows
(pure sizing math), split_into_ranges (the only stage touching
PartitionDb), and assign_boundaries (pure bucket accumulation). The
pure stages are now directly unit-testable and the signatures document
the data flow.
Rename the sizing knobs to say what they mean
(TARGET_RANGES_PER_WORKER, min_rows_per_worker,
target_max_rows_per_range), trim module and function docs to the
essentials, and stop deduplicating repeated boundary ends. Duplicate
ends only arise from non-advancing servers, and the snapshot layer
validates boundary monotonicity server-side before using boundaries.
MysqlKeyProber is the concrete prober over a live connection, and the
partitioner trait seam takes the KeyProber name. explain_row_estimate
is no longer part of the crate API, callers reach estimates through
MysqlKeyProber.
@peterdukelarsen
peterdukelarsen force-pushed the pl/mysql-pk-partitioner branch from e1d8b37 to 65de2bc Compare August 5, 2026 22:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant