Skip to content

fix: tag PAC-built image with commit SHA and raise build timeout to 2h - #1166

Merged
verseghy-prow[bot] merged 1 commit into
masterfrom
fix-pac-image-tag-and-timeout
Aug 28, 2026
Merged

verseghy-prow[bot] merged 1 commit into
masterfrom
fix-pac-image-tag-and-timeout

Conversation

@TwoDCube

Copy link
Copy Markdown
Member

Mirrors Verseghy/website_backend2#242, which is verified working end to end (build → SHA tag → automatic website_k8s deploy PR Verseghy/website_k8s#35 within 2 minutes).

Problem 1: image never deploys

The BuildRun pushes the output image untagged, i.e. to :latest. The ImageUpdater CR in website_k8s only considers tags matching ^[0-9a-f]{40}$, so even a successful build would never produce a deploy PR — the deployed frontend is still pinned to a Nov 12, 2025 commit.

Fix: tag the output with {{ revision }} (PAC substitutes the full 40-char commit SHA).

Problem 2: builds time out

Both PAC builds so far (Jun 8 and Jun 26) failed with BuildRunTimeout: failed to finish within 1h0m0s — each died at exactly the 1h wall. Notably the Jun 26 run started the very same second as a backend2 build (which took 47 min that night vs its usual 35), so the two were competing for the node's CPU.

Fix: BuildRun timeout 1h → 2h, with the surrounding timeouts keeping their ordering so a slow build surfaces as a build verdict rather than a pipeline timeout: 2h build < 130m oc wait < 2h30m pipeline.

After merging

The merge push triggers the first build attempt under the new limits. If it succeeds, expect the automatic website_k8s deploy PR ~2 minutes later — note it will jump the deployed frontend from Nov 2025 straight to current master. If it still times out at 2h, the build is likely hung rather than slow and needs investigation of the build log instead of more headroom.

🤖 Generated with Claude Code

Two fixes for the PAC build, mirroring website_backend2#242:

1. The BuildRun pushed the output image untagged (:latest), which
   argocd-image-updater ignores — it only accepts tags matching
   ^[0-9a-f]{40}$. Tag the output with {{ revision }} so successful
   builds actually produce a deploy PR against website_k8s.

2. Both June 2026 builds died at exactly the 1h BuildRun timeout (the
   June 26 one started the same second as a backend2 build and competed
   with it for CPU). Raise the BuildRun timeout to 2h, the oc wait to
   130m and the pipeline timeout to 2h30m, keeping wait > build and
   pipeline > wait so failures surface as verdicts, not timeouts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@verseghy-prow
verseghy-prow Bot requested a review from smrtrfszm August 28, 2026 20:55
@verseghy-prow

verseghy-prow Bot commented Aug 28, 2026

Copy link
Copy Markdown

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: TwoDCube

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@verseghy-prow verseghy-prow Bot added the approved Indicates a PR has been approved by an approver from all required OWNERS files. label Aug 28, 2026
@smrtrfszm

Copy link
Copy Markdown
Member

/lgtm

@verseghy-prow verseghy-prow Bot added the lgtm Indicates that a PR is ready to be merged. label Aug 28, 2026
@verseghy-prow
verseghy-prow Bot merged commit 5ca2e78 into master Aug 28, 2026
2 checks passed
@verseghy-prow
verseghy-prow Bot deleted the fix-pac-image-tag-and-timeout branch August 28, 2026 21:05
verseghy-prow Bot pushed a commit that referenced this pull request Aug 29, 2026
…1167)

The previous timeout fix (#1166) was insufficient — its first build
revealed two deeper limits:

1. Tekton's cluster default TaskRun timeout (60m) killed the pipeline
   task at exactly 1h despite the 2h30m pipeline timeout; raising
   timeouts.pipeline does not lift the per-TaskRun default. Give the
   task an explicit 2h20m timeout (build 2h < wait 130m < task 2h20m
   < pipeline 2h30m).

2. The buildah ClusterBuildStrategy caps build-and-push at 1 CPU / 2Gi.
   The webpack build pinned that CPU and sat at the 2Gi ceiling for 90+
   minutes with no output (GC thrash), then hit the BuildRun's own 2h
   timeout. Override via strategy stepResources to 4 CPU / 6Gi for this
   build only (nodes have 31.5 CPU / 62Gi allocatable).

The BuildRun JSON with the override was validated against the live API
with oc create --dry-run=server; the field survives admission.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

approved Indicates a PR has been approved by an approver from all required OWNERS files. lgtm Indicates that a PR is ready to be merged.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants