Skip to content

fix(ci): wait for ordinary eval agents within job timeout - #330

Merged
mrizzi merged 1 commit into
RHEcosystemAppEng:mainfrom
mrizzi:TC-6783
Oct 7, 2026
Merged

mrizzi merged 1 commit into
RHEcosystemAppEng:mainfrom
mrizzi:TC-6783

Conversation

@mrizzi

@mrizzi mrizzi commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Implements TC-6783, a CI-failure sub-task of TC-6726.

  • Set CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS="0" for ordinary evals so Claude waits for background eval agents rather than stopping them after ten idle minutes.
  • Bound the entire ordinary run-evals job to 90 minutes.
  • Add regression tests that execute the workflow shell and observe the wait setting on both sequential Claude invocations, plus the job timeout contract.

The failing ordinary job explicitly reported terminating all six verify-pr eval agents after ten idle minutes. Anthropic documents the wait setting in its environment variable reference.

Validation

  • Both new regression tests failed against the original workflow: missing wait setting and missing explicit job timeout.
  • Full scripts test suite: 131 passed.
  • Skillsaw: 0 errors, 7 existing warnings.
  • claude plugin validate plugins/sdlc-workflow: passed.
  • git diff --check: passed.

Unit tests validate configuration propagation through the actual workflow shell with a non-inference Claude stub. Actual long-running background-agent completion must be validated in the hosted PR #299 eval after this trusted-main workflow change merges. The 90-minute limit covers setup, all requested skills and publication; a run exceeding it will time out.

Rollout

The runner uses the trusted workflow from main. After merging this PR, synchronize PR #299 with main and trigger a fresh source-bound eval run. Native grading failures from the reported run remain separate work.

Summary by Sourcery

Ensure ordinary CI evaluations wait for background agents while enforcing a finite job runtime.

Bug Fixes:

  • Prevent ordinary evaluation agents from being stopped by the default idle wait ceiling while ensuring the CI job remains bounded.

CI:

  • Set a 90-minute timeout for the ordinary PR evaluation job and configure it to wait for background evaluation agents to complete.
  • Add regression coverage for wait-setting propagation across Claude invocations and the job timeout contract.

Disable the Claude print-mode idle background wait ceiling while bounding the ordinary eval job to 90 minutes. Cover both sequential CLI invocations and the explicit job limit.

Implements TC-6783

Assisted-by: Codex
@sourcery-ai

sourcery-ai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Reviewer's Guide

The workflow now lets ordinary Claude eval invocations wait for background agents indefinitely, while the enclosing run-evals job enforces a 90-minute maximum; tests execute the real workflow shell with a stub to verify both invocations receive the wait and environment-scrubbing settings and assert the timeout configuration.

Sequence diagram for bounded ordinary CI evaluation

sequenceDiagram
    participant Workflow as run-evals workflow
    participant Claude as Claude invocation
    participant Agents as Background eval agents
    participant Job as 90-minute job timeout

    Workflow->>Claude: invoke ordinary evaluation
    Claude->>Agents: start eval agents
    Claude->>Agents: wait with CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS=0
    Agents-->>Claude: complete evaluations
    Claude-->>Workflow: return evaluation result
    Job-->>Workflow: enforce timeout-minutes: 90
    alt 90-minute limit reached
        Job->>Workflow: terminate run-evals job
    end
Loading

Flow diagram for ordinary eval runtime bounds

flowchart LR
    Start["run-evals job starts"] --> Env["Set background wait ceiling to 0"]
    Env --> Invoke["Run Claude evaluation"]
    Invoke --> Wait["Wait for background eval agents"]
    Wait --> Result["Publish evaluation result"]
    Limit["90-minute job timeout"] -.-> Invoke
    Limit -.-> Wait
    Limit -.-> Result
Loading

File-Level Changes

Change Details Files
Ensure ordinary eval agents can complete while keeping the CI job bounded.
  • Set the Claude background-agent wait ceiling to unlimited for ordinary eval invocations.
  • Add a 90-minute timeout to the overall eval job.
  • Add regression coverage for environment propagation across both CLI calls and the workflow timeout contract.
.github/workflows/eval-pr-run.yml
plugins/sdlc-workflow/scripts/test_native_fullsend_eval_ci.py

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

Comment thread .github/workflows/eval-pr-run.yml
@mrizzi

mrizzi commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator Author

Verification Report for TC-6783 (commit e9bcd2b)

Check Result Details
Review Feedback PASS 1 comment thread, classified question (semantics of CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS: "0"); confirmed against docs, no code change requests
Root-Cause Investigation N/A No sub-tasks created
Scope Containment PASS Exactly the 2 task-specified files modified; nothing out-of-scope or missing
Diff Size PASS +30 / -3 across 2 files, proportionate to the fix
Commit Traceability PASS Commit e9bcd2bb references TC-6783 ("Implements TC-6783")
Sensitive Patterns PASS No secrets, keys, or credentials in added lines
CI Status PASS All checks pass (Sourcery neutral, Trigger Eval Dispatch, Plugin Validation, Skill Lint)
Acceptance Criteria PASS 5 of 5 criteria met
Test Quality PASS No repetitive tests; both new tests have docstrings; Eval Quality: N/A
Test Change Classification ADDITIVE +2 test functions, no tests removed or weakened
Verification Commands PASS pytest 131 passed · skillsaw 0 errors / Grade A · plugin validation passed · git diff --check clean

Overall: PASS

The PR correctly disables the Claude print-mode idle background-wait ceiling (CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS: "0" = wait indefinitely, per the official env-var docs) while bounding the run-evals job to timeout-minutes: 90. Credential scrubbing, WIF, sandbox and native-execution settings are preserved. The reviewer's question on the "0" semantics is confirmed resolved (0 = wait indefinitely, not "wait 0 ms"); no code change required.

Note: all four verification commands were run with the sandbox disabled, because the PR branch is checked out in a worktree outside the sandbox's writable allowlist. One unrelated test (test_multiline_credentials_register_individual_nonempty_masks) fails only under that sandbox write restriction — it passes in an isolated checkout and the PR does not touch its masking logic, so it is not a regression.

This report is informational. Merging and Jira transitions remain a human decision.


This comment was AI-generated by sdlc-workflow/verify-pr v0.13.9.

@mrizzi
mrizzi merged commit f3416b4 into RHEcosystemAppEng:main Oct 7, 2026
5 checks passed
@mrizzi
mrizzi deleted the TC-6783 branch October 7, 2026 18:51
mrizzi added a commit that referenced this pull request Oct 7, 2026
Merge main through PR #330 to run the source-bound PR #299 eval with the background wait policy and bounded CI timeout.

Refs TC-6726, TC-6783

Assisted-by: Codex
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant