Skip to content

eventservice: redo scan slices still cause periodic sink stalls under backlog #6027

Description

@asddongmen

Bug Report

Based on #5956 plus the transaction-cursor, redo scan-window cap, and dispatcher prefix-drain follow-ups, a fresh redo changefeed under sustained write backlog still emits sink traffic periodically.

In the exact 30-minute Fix3 validation, the dispatcher prefix drain reduced the original 18.5-minute full-batch stall, but a later interval still held scan_task_count and redo resolved-ts flat for 300 seconds while redo row writes continued. The sink then resumed in large chunks. The scan-window interval remained at the 5-second redo cap throughout.

At this ingestion rate a five-second source-time slice is still much larger than scanner capacity under the 1.6-core limit. Closing that slice delays the durable redo resolved fence, so even prefix draining cannot release newer normal events until the scan boundary completes.

Proposed fix

Use the existing one-second minimum scan interval as the hard cap for redo-enabled changefeeds, while preserving an even smaller syncpoint interval and leaving non-redo behavior unchanged.

Validation

  • targeted tests pass
  • make unit_test_pkg PKG=./pkg/eventservice/...: 184 tests, race/failpoints, 82.8% coverage
  • make fmt, make check, and diff checks pass
  • fresh 30-minute E2E pending on the exact fix commit

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions