Skip to content

Add multi-controller assignPod benchmark and fixture - #144

Open
scott-rc wants to merge 1 commit into
sc/fix-cold-start-feedbackfrom
sc/fix-pod-pool-contention
Open

scott-rc wants to merge 1 commit into
sc/fix-cold-start-feedbackfrom
sc/fix-pod-pool-contention

Conversation

@scott-rc

Copy link
Copy Markdown
Contributor

Sets up the test harness for the partitioned pod-assignment work that follows. No behavior change in this PR.

fixture.NewControllerPodAt(name, ip) lets a single fake clientset seed multiple controller pods with distinct PodIPs so their hash rings can agree on N peers. multiControllerFixture (in internal/controller/multicontroller_test.go) builds N Controller instances that share one fake clientset, with peer seeding done before startInformers so the controllers pick the peers up via initial-list without overflowing the fake watcher's small event channel.

BenchmarkAssignPodContention (in internal/controller/pod_bench_test.go) drives the assignment hot path with N=3 controllers calling assignPod concurrently for distinct functions. It reports ns/op, a p99-ms percentile, and a patches/success ratio computed via a patch reactor on the fake clientset. The fake serializes Patch through a tracker mutex and does not implement optimistic concurrency on ResourceVersion, so the wall-clock numbers are not directly comparable to production tail latency; this bench exists to set the forward-looking acceptance bar for the partition work that comes next and to catch regressions in the assignment hot path between commits.

Stacked on #143 (router-side in-flight counter decoupling), which is the prerequisite for this work.

Sets up the test infrastructure for the upcoming pod-pool partition
work. NewControllerPodAt lets a single fake clientset hold multiple
controller pods with distinct PodIPs so their hash rings can agree on
N peers. The multiControllerFixture helper builds N Controllers that
share one fake clientset and starts informers in that pre-seeded
state. BenchmarkAssignPodContention exercises the assignment hot path
under concurrent calls from N controllers and reports ns/op, p99-ms,
and patches/success.

The fake clientset does not simulate optimistic-concurrency on Patch,
so the numbers are not directly comparable to production tail
latency; this baseline exists to catch regressions and to set the
forward-looking acceptance bar for the partitioned implementation.
@scott-rc

Copy link
Copy Markdown
Contributor Author

This change is part of the following stack:

Change managed by git-spice.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit ad5e57d. Configure here.

for range poolPerFunction {
poolObjs = append(poolObjs, benchAvailablePod(fns[c], sharedHost, int32(sharedPort)))
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All functions share same deployment, making pools indistinct

Medium Severity

fixture.NewFunction always sets Deployment to the constant "test". Since getUnassignedPods selects candidates using the key.Deployment.Label selector matching fn.GetDeployment(), all three functions query the same label value "test", making all 60K pods (3 × 20K) a single shared pool. The benchmark claims to drive "distinct functions" but pods aren't actually partitioned per function — every controller contends for the same undifferentiated pool, which misrepresents the contention pattern the benchmark is designed to measure.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit ad5e57d. Configure here.

_, err := ctrl.assignPod(ctx, fn)
elapsed := time.Since(start)
if err != nil {
b.Fatalf("assignPod failed (goroutine %d): %v", gid, err)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

b.Fatalf called from non-test goroutine in RunParallel

Medium Severity

b.Fatalf is called inside the b.RunParallel callback, which runs in goroutines spawned by RunParallel, not the benchmark's main goroutine. The Go testing docs state that FailNow/Fatal/Fatalf "must be called from the goroutine running the test or benchmark function, not from other goroutines created during the test." On Go 1.24+ (this project uses Go 1.26), this can panic or produce unreliable behavior where only the calling goroutine exits while other goroutines continue.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit ad5e57d. Configure here.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant