e2e: cover scaling out from a single instance - #947
Closed
claudiodekker wants to merge 1 commit into
Closed
claudiodekker wants to merge 1 commit into
claudiodekker wants to merge 1 commit into
Conversation
claudiodekker
force-pushed
the
e2e/scale-up-from-single-instance
branch
from
September 4, 2026 14:01
9238fec to
56320db
Compare
claudiodekker
marked this pull request as draft
September 4, 2026 15:36
The existing scale-out test in replication_test.go grows a cluster from replicas: 3 to 5. Growing a cluster that starts at replicas: 1 is not covered anywhere in e2e, and it is not equivalent: at replicas: 1 the primary has never had semi-sync enabled, because configurePrimary returns early for a single-instance cluster. Scaling out enables it for the first time, at a moment when no replica can acknowledge yet. The test creates a cluster at replicas: 1, writes rows, keeps one commit in flight across the scale-out, scales to 3, and asserts SyncedReplicas == 3 with the rows present on index 2. The concurrent writer is required. With writes quiesced the same scale-out succeeds in about 200s; with one commit in flight it does not converge. On v0.37.0 the primary is left with Rpl_semi_sync_master_status ON, clients 0 and wait_sessions 1, an INSERT stuck in "Waiting for semi-sync ACK from slave", and the recipient's clone fails in PAGE COPY with "ERROR 3862: Clone Donor Error: 1815: Innodb Clone Restart failed". FILE COPY completes first, so data does transfer. Observed on three hosts running k3s v1.36.2 with MOCO v0.37.0 and mysql:8.4.10, and reproduced separately with two mysql:8.4.10 containers. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KgjRthcZN1wqiwM7Acwn1n
claudiodekker
force-pushed
the
e2e/scale-up-from-single-instance
branch
from
September 4, 2026 15:40
56320db to
efea2b2
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this adds
e2e/scaleup_single_test.go, an e2e case for growing a cluster that starts atreplicas: 1. Run it withcd e2e && make start && make test FOCUS="scale up from a single instance".It fails on v0.37.0. It is a reproduction, not a passing test. I figured this would be more helpful compared to creating an issue instead.
Why 1 -> 3 is not the same as 3 -> 5
The existing scale-out test (
e2e/replication_test.go) grows from 3 to 5. Growing from 1 is not covered anywhere in e2e. It is also a different operation: atreplicas: 1the primary has never had semi-sync enabled, becauseconfigurePrimaryreturns early for a single-instance cluster. Scaling out turns it on for the first time.The new test creates a cluster at
replicas: 1, writes rows, keeps one commit in flight across the scale-out, scales to 3, and assertsSyncedReplicas == 3with the rows present on index 2.The failure
Scaling 1 -> 3 never finishes while anything is writing to the primary.
When the cluster reaches
replicas: 3,configurePrimaryturns semi-sync on and requires one replica to acknowledge each commit (wait_for_slave_count = floor(replicas/2),clustering/operations.go:529-538). No replica can do that: the new instances are empty and still have to be cloned from the primary. The wait does not time out either, because the timeout is hardcoded to 24h (pkg/dbop/replication.go:8).Every commit on the primary then blocks. That breaks the clone: the donor cannot finish PAGE COPY while it holds a transaction that never commits, so the replicas never come up to acknowledge anything.
Seen on three hosts running k3s v1.36.2, MOCO v0.37.0 and
mysql:8.4.10, with a writer inserting through the primary Service:SyncedReplicas1, cluster not healthyRpl_semi_sync_master_status ON,Rpl_semi_sync_master_clients 0,Rpl_semi_sync_master_wait_sessions 1INSERT INTO ...in stateWaiting for semi-sync ACK from slaveclone_progress:FILE COPY Completed 73668205 bytes,PAGE COPY FailedERROR 3862: Clone Donor Error: 1815 : Innodb Clone Restart failed, existing clone abortedIf you reproduce this, the logs will show
Clone Apply Restarting State: PAGE COPY. That is the clone plugin restarting its own copy. The mysqld container'srestartCountstays at 0 throughout, and data does transfer: FILE COPY completes before PAGE COPY fails.Reproducing it without Kubernetes
Two
mysql:8.4.10containers, configured the way MOCO configures mysqld, withCLONE INSTANCE FROM 'moco-clone-donor'@'donor':33062. Each run takes about 40 seconds.wait_for_slave_count=1, 0 replicas, idlewait_for_slave_count=1, 0 replicas, one blocked commitThe failure is either the 3862 above, or an indefinite stall with
clone_progresssitting atPAGE COPY, In Progress, 0 bytes.Workaround
Stopping writes for the duration of the scale-out is enough. With writes stopped, the same 1 -> 3 scale-out on the same hosts reached healthy 3/3 in 200s, twice. The freeze has to be complete.
If it has already stalled, it does not clear itself. Stopping the writer at that point does nothing, and neither does killing it: after
KILL CONNECTION,Rpl_semi_sync_master_wait_sessionsstays at 1 and the session is still waiting, because the commit is past the binlog sync and is waiting for an acknowledgement before the engine commit. The row is neither visible nor rolled back.SET GLOBAL rpl_semi_sync_master_enabled=OFFon the primary does clear it. In the container reproduction the blocked commit completed at once and the clone finished within five seconds, ending atclone_status Completedwith the recipient holding the donor's rows, including the write that had been stuck. Two caveats: commits are unacknowledged while semi-sync is off, and MOCO turns it back on at the next reconcile, so writes need to be stopped first or the stall comes back.Two things that do not help:
rpl_semi_sync_master_timeoutthrough the user's own my.cnf (mysqlConfigMapName). It is a hardcoded constant applied withSET GLOBALbydbop.ConfigurePrimary, which runs during the scale-out and overwrites the user's value.floor(2/2)is also 1, so it still asks for an acknowledgement that cannot arrive.Note for anyone fixing this
Ordering the clone before semi-sync is enabled looks like the fix. It is not, and why it fails is not obvious, so I am recording it here.
GatherStatusonly rejects a cluster when the pods are missing (clustering/status.go:173). The StatefulSet creates pods in seconds, but mysqld takes about a minute to initialise, so there is a long window where the pods exist and theirMySQLStatusis still nil. During that window a clone-first ordering finds nothing to clone,configurePrimaryenables semi-sync anyway, and the clone then starts against a donor that is already blocked.I built that fix and it passed an envtest that attached mysqld status to the new pods immediately. On real hardware it failed in exactly the same way as v0.37.0. Making the mock report those instances unreachable for three seconds reproduced the failure locally too.
A fix probably needs to make enabling semi-sync conditional on enough replicas being able to acknowledge, instead of relying on ordering inside one reconcile pass. That decision affects durability, so I have not made it here. It trades against the invariant in
.github/instructions/clustering.instructions.mdand against the 24h timeout, which looks deliberate: MySQL's default is 10s, after which the source falls back to async.🤖 Generated with Claude Code