Skip to content

PBS-64: Implement incrementing buffer size strategy in sender_context::get_event() - #200

Merged
kamil-holubicki merged 2 commits into
Percona-Lab:mainfrom
kamil-holubicki:PBS-64
Oct 6, 2026
Merged

kamil-holubicki merged 2 commits into
Percona-Lab:mainfrom
kamil-holubicki:PBS-64

Conversation

@kamil-holubicki

Copy link
Copy Markdown
Contributor

No description provided.

kamil-holubicki and others added 2 commits October 6, 2026 11:37
…:get_event()

https://perconadev.atlassian.net/browse/PBS-64

Problem:
The sender fetches binlog bytes from storage in fixed-size blocks (default
1 MiB). When the first event starting at the block boundary is itself
larger than the block, zero complete events parse out of the fetched
buffer and the sender returns a hard error. From that point on the sender
can never progress past any event larger than one block.

Solution:
Before parsing the fetched buffer, peek the first event's advertised size
from its common header and compare it against the number of bytes actually
received. When the event does not fit, re-issue a single fetch at the
same offset with length set to that exact size; the parse then succeeds
with one complete event.

The common header is always the first 19 bytes of a well-formed event
(matches "default_common_header_length" and the "size_in_bytes" invariant
on "common_header_view_base"). A fetched buffer shorter than that, or an
exact-size retry that still yields no complete event, both indicate a
corrupt or truncated binlog and are reported as errors with no further
retry.

The advertised size is bounded by a new "max_event_size_bytes" constant
set to 1 GiB, matching MySQL's "MAX_MAX_ALLOWED_PACKET" and
"replica_max_allowed_packet" ceiling. A corrupt or malicious header
claiming a multi-GiB event is rejected before any allocation is made.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
https://perconadev.atlassian.net/browse/PBS-64

New MTR test "rs_large_event" exercises the full source -> PBS -> consumer
pipeline with a single row-based event large enough to exceed PBS's fetch
block. A LONGBLOB row of ~1.5 MiB is inserted on the source, PBS fetches
the resulting binlogs into its storage, and mysqlbinlog connects to the
PBS replication-source listener and drains the event back out. The test
then (a) waits for the "sender : block too small for first event" entry
in the PBS log, which only appears when sender_context detected that the
first event in a block did not fit and re-issued the fetch at exact size,
and (b) replays the mysqlbinlog dump into a backup database and verifies
the restored row matches the original by (id, length, MD5). The MD5
comparison is used in place of diff_tables.inc because the default sort
buffer cannot hold a 1.5 MiB blob and would make the test fail for a
reason unrelated to what is being covered.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

@percona-ysorokin percona-ysorokin left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@kamil-holubicki
kamil-holubicki merged commit ee82afa into Percona-Lab:main Oct 6, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants