Skip to content

e2e: cover scaling out from a single instance - #947

Closed
claudiodekker wants to merge 1 commit into
cybozu-go:mainfrom
claudiodekker:e2e/scale-up-from-single-instance
Closed

claudiodekker wants to merge 1 commit into
cybozu-go:mainfrom
claudiodekker:e2e/scale-up-from-single-instance

Conversation

@claudiodekker

@claudiodekker claudiodekker commented Sep 4, 2026 •

Copy link
Copy Markdown

What this adds

e2e/scaleup_single_test.go, an e2e case for growing a cluster that starts at replicas: 1. Run it with cd e2e && make start && make test FOCUS="scale up from a single instance".

It fails on v0.37.0. It is a reproduction, not a passing test. I figured this would be more helpful compared to creating an issue instead.

Why 1 -> 3 is not the same as 3 -> 5

The existing scale-out test (e2e/replication_test.go) grows from 3 to 5. Growing from 1 is not covered anywhere in e2e. It is also a different operation: at replicas: 1 the primary has never had semi-sync enabled, because configurePrimary returns early for a single-instance cluster. Scaling out turns it on for the first time.

The new test creates a cluster at replicas: 1, writes rows, keeps one commit in flight across the scale-out, scales to 3, and asserts SyncedReplicas == 3 with the rows present on index 2.

The failure

Scaling 1 -> 3 never finishes while anything is writing to the primary.

When the cluster reaches replicas: 3, configurePrimary turns semi-sync on and requires one replica to acknowledge each commit (wait_for_slave_count = floor(replicas/2), clustering/operations.go:529-538). No replica can do that: the new instances are empty and still have to be cloned from the primary. The wait does not time out either, because the timeout is hardcoded to 24h (pkg/dbop/replication.go:8).

Every commit on the primary then blocks. That breaks the clone: the donor cannot finish PAGE COPY while it holds a transaction that never commits, so the replicas never come up to acknowledge anything.

Seen on three hosts running k3s v1.36.2, MOCO v0.37.0 and mysql:8.4.10, with a writer inserting through the primary Service:

  • Still unfinished after 431s: SyncedReplicas 1, cluster not healthy
  • New pods: 2/3, never ready
  • Primary: Rpl_semi_sync_master_status ON, Rpl_semi_sync_master_clients 0, Rpl_semi_sync_master_wait_sessions 1
  • Primary processlist: INSERT INTO ... in state Waiting for semi-sync ACK from slave
  • Recipient clone_progress: FILE COPY Completed 73668205 bytes, PAGE COPY Failed
  • Recipient: ERROR 3862: Clone Donor Error: 1815 : Innodb Clone Restart failed, existing clone aborted

If you reproduce this, the logs will show Clone Apply Restarting State: PAGE COPY. That is the clone plugin restarting its own copy. The mysqld container's restartCount stays at 0 throughout, and data does transfer: FILE COPY completes before PAGE COPY fails.

Reproducing it without Kubernetes

Two mysql:8.4.10 containers, configured the way MOCO configures mysqld, with CLONE INSTANCE FROM 'moco-clone-donor'@'donor':33062. Each run takes about 40 seconds.

Donor state Clone
Writable primary, idle succeeds
Semi-sync on, wait_for_slave_count=1, 0 replicas, idle succeeds
Committing sessions, semi-sync off succeeds
Semi-sync on, wait_for_slave_count=1, 0 replicas, one blocked commit fails

The failure is either the 3862 above, or an indefinite stall with clone_progress sitting at PAGE COPY, In Progress, 0 bytes.

Workaround

Stopping writes for the duration of the scale-out is enough. With writes stopped, the same 1 -> 3 scale-out on the same hosts reached healthy 3/3 in 200s, twice. The freeze has to be complete.

If it has already stalled, it does not clear itself. Stopping the writer at that point does nothing, and neither does killing it: after KILL CONNECTION, Rpl_semi_sync_master_wait_sessions stays at 1 and the session is still waiting, because the commit is past the binlog sync and is waiting for an acknowledgement before the engine commit. The row is neither visible nor rolled back.

SET GLOBAL rpl_semi_sync_master_enabled=OFF on the primary does clear it. In the container reproduction the blocked commit completed at once and the clone finished within five seconds, ending at clone_status Completed with the recipient holding the donor's rows, including the write that had been stuck. Two caveats: commits are unacknowledged while semi-sync is off, and MOCO turns it back on at the next reconcile, so writes need to be stopped first or the stall comes back.

Two things that do not help:

  • Lowering rpl_semi_sync_master_timeout through the user's own my.cnf (mysqlConfigMapName). It is a hardcoded constant applied with SET GLOBAL by dbop.ConfigurePrimary, which runs during the scale-out and overwrites the user's value.
  • Scaling 1 -> 2 first. floor(2/2) is also 1, so it still asks for an acknowledgement that cannot arrive.

Note for anyone fixing this

Ordering the clone before semi-sync is enabled looks like the fix. It is not, and why it fails is not obvious, so I am recording it here.

GatherStatus only rejects a cluster when the pods are missing (clustering/status.go:173). The StatefulSet creates pods in seconds, but mysqld takes about a minute to initialise, so there is a long window where the pods exist and their MySQLStatus is still nil. During that window a clone-first ordering finds nothing to clone, configurePrimary enables semi-sync anyway, and the clone then starts against a donor that is already blocked.

I built that fix and it passed an envtest that attached mysqld status to the new pods immediately. On real hardware it failed in exactly the same way as v0.37.0. Making the mock report those instances unreachable for three seconds reproduced the failure locally too.

A fix probably needs to make enabling semi-sync conditional on enough replicas being able to acknowledge, instead of relying on ordering inside one reconcile pass. That decision affects durability, so I have not made it here. It trades against the invariant in .github/instructions/clustering.instructions.md and against the 24h timeout, which looks deliberate: MySQL's default is 10s, after which the source falls back to async.

🤖 Generated with Claude Code

@claudiodekker
claudiodekker force-pushed the e2e/scale-up-from-single-instance branch from 9238fec to 56320db Compare September 4, 2026 14:01
@claudiodekker
claudiodekker marked this pull request as draft September 4, 2026 15:36
The existing scale-out test in replication_test.go grows a cluster from
replicas: 3 to 5. Growing a cluster that starts at replicas: 1 is not covered
anywhere in e2e, and it is not equivalent: at replicas: 1 the primary has never
had semi-sync enabled, because configurePrimary returns early for a
single-instance cluster. Scaling out enables it for the first time, at a moment
when no replica can acknowledge yet.

The test creates a cluster at replicas: 1, writes rows, keeps one commit in
flight across the scale-out, scales to 3, and asserts SyncedReplicas == 3 with
the rows present on index 2.

The concurrent writer is required. With writes quiesced the same scale-out
succeeds in about 200s; with one commit in flight it does not converge. On
v0.37.0 the primary is left with Rpl_semi_sync_master_status ON, clients 0 and
wait_sessions 1, an INSERT stuck in "Waiting for semi-sync ACK from slave", and
the recipient's clone fails in PAGE COPY with "ERROR 3862: Clone Donor Error:
1815: Innodb Clone Restart failed". FILE COPY completes first, so data does
transfer.

Observed on three hosts running k3s v1.36.2 with MOCO v0.37.0 and
mysql:8.4.10, and reproduced separately with two mysql:8.4.10 containers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KgjRthcZN1wqiwM7Acwn1n
@claudiodekker
claudiodekker force-pushed the e2e/scale-up-from-single-instance branch from 56320db to efea2b2 Compare September 4, 2026 15:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant