Skip to content

[DRAFT] verify-pr fullsend CI deployment (TC-6180) - #299

Open
mrizzi wants to merge 190 commits into
mainfrom
verify-pr-fullsend
Open

mrizzi wants to merge 190 commits into
mainfrom
verify-pr-fullsend

Conversation

@mrizzi

@mrizzi mrizzi commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Draft — do not merge. Continuous-CI vehicle for the verify-pr-fullsend feature branch (TC-6180 CI-deployment feature).

Repo CI (validate-plugins, skillsaw, and the path-filtered evals) triggers only on pull_request/push to main. This draft PR gives the feature branch continuous CI against main as sub-tasks land — including the expensive, main-only evals that are intentionally not run on per-sub-task PRs into the feature branch.

Sub-tasks continue to PR into verify-pr-fullsend individually (e.g. #298). This PR is the aggregate gate; it stays draft until the feature is complete and TC-5816 (merge bookend) removes the temporary feature-branch CI entries.

Summary

  • Latest synchronization: e19f2e0a merges main through PR fix(ci): wait for ordinary eval agents within job timeout #330 (f3416b46), importing the ordinary runner wait policy and 90-minute total job timeout. Local validation: 499 passed, 1 skipped; lint/plugin validation passed. A fresh hosted eval is being triggered; previous failures predate this wait fix.

  • TC-6764 fixes TC-6763 in 91698d4d: clarify the native judge’s expected absent/empty stopping boundaries while retaining actual Skill/plugin/tool evidence and rejection of genuine bootstrap failures. Only evals/fullsend/triage-security/judge.md changed; all 21 assertions and product files are unchanged. Saved actual evidence regraded with Opus 4.8: both cases passed in three repeats (6/6); both no-Skill bootstrap controls remained false (2/2 correctly rejected). Original evidence hashes unchanged; no agents rerun. Full local suite: 482 passed, 1 skipped; Skillsaw 0 errors/8 warnings, plugin validation and whitespace checks passed. Hosted activation remains pending: trusted main still pins the older suite source, so this branch push does not activate the new judge in CI. Separate malformed grading inconsistency and hosted valid-case failure are outside this fix.

  • TC-4636: request case 3's eval-review detection evidence in the report; preserve all six verify-pr cases and 68 assertions.

  • TC-6677: retire Fullsend gate cases from the ordinary eval stack, preserving the original triage-security 32 cases and 164 assertions. The native four-case/21-judgment suite is now included here; hosted WIF validation waits for the minimal CI bootstrap in PR ci(fullsend): bootstrap evals from reviewed PR299 source #323.

  • TC-6213: final ordinary integration checks pass: triage-security 164/164, verify-pr 68/68, and 312 local script tests. Hosted run.

  • TC-6726: native suite and normal CI activation delivered in b46b2bf6 and 7b62662c. PR ci(fullsend): bootstrap evals from reviewed PR299 source #323 contains only four CI bootstrap files and pins the reviewed suite commit b46b2bf647be8451b6aedda58dc8e51448f0eab6. After this integration merges, suite execution uses trusted main. Local validation: 423 tests passed, 1 optional resolver skip, including a fresh Python 3.12 environment; lint/plugin/preflight passed. Hosted unit-test collection initially lacked PyYAML; corrected in ed117373 by adding it to the existing unit-test workflow. The obsolete eval run was cancelled before that corrective push; the new hosted unit-test matrix passed on Python 3.11, 3.12, 3.13 and 3.14 (run). Ordinary eval feedback and native WIF validation remain pending. Existing ordinary eval manifests, rendering/sticky comments and TC-4636 are preserved. Merge ci(fullsend): bootstrap evals from reviewed PR299 source #323 first, then rerun this PR’s existing eval flow and require real WIF-backed native21/21 before merging this PR.

  • TC-6726: PR ci(fullsend): bootstrap evals from reviewed PR299 source #323 is merged. Commit fd5c06fb merges trusted main (805068c6) into this branch without rebasing or force-pushing. All ten bootstrap review fixes are retained, together with normal activation, deterministic rendering and sticky comments. The ordinary publisher now rechecks the PR head before replacing its shared report. Native cases/judgments and ordinary manifests are unchanged. Final local validation: 467 passed, 1 optional skip; required script suite 464 passed, 1 optional skip; Skillsaw zero errors/8 existing warnings, plugin validation, native no-inference preflight and resolved-file whitespace checks passed. Main's imported generated baselines retain existing whitespace findings. Hosted WIF-backed native21/21 and ordinary results remain pending; this PR stays draft.

🤖 Generated with Claude Code

Native failure diagnostics — TC-6726

  • Commit c7ca8495f1a83c51d191f55e391b17c5b29e97da adds allowlisted phase/category names and integer exit codes to the safe native result. Raw logs remain private; cases, assertions and credential handling are unchanged. Categories are advisory, not proof of the cause.
  • Local validation: 481 tests passed, 1 skipped; native no-inference preflight, Skillsaw and plugin validation passed. Read-only review found no remaining blockers.
  • PR #326 changes only the trusted main suite pin to this immutable diagnostic source. No native eval files are added to main.
  • After fix(ci): pin reviewed native eval diagnostics #326 is reviewed and merged, rerun the Eval PR trigger 37483925429 to get a fresh consumer using the updated main pin. The original native failure remains unresolved until that diagnostic run is inspected.

Opus 4.8 judge trial after PR #327 merge

  • Synced merged main 27a258d7cdb2cb829a7f2df7ccf300df2dfdc1d3 into this branch as 6a251f7dce1dd921397baf6e7324456c603163c0; both existing and newly merged tests retained.
  • Trusted main wrapper explicitly passes --judge-model claude-opus-4-8. The reviewed native suite pin remains unchanged.
  • Local checks: 482 tests passed, 1 skipped; Skillsaw and plugin validation passed. PR is mergeable.
  • Fresh Eval PR trigger: https://github.com/RHEcosystemAppEng/sdlc-plugins/actions/runs/37500396985 . Hosted model-access hypothesis remains unconfirmed pending the result.

PR328 merged; current main synchronization

  • PR328 merged as ab0c65ae2e3f184b5ddb878efb3afcb50929a9fe; trusted main now selects reviewed judge source 91698d4dca24bc199e763dc6462dc54523e473c0. TC6768 is Closed/Done.
  • Merge commit 43200c82 imports current main without rebasing/force-pushing. Conflict resolutions retain this branch’s post-integration trusted-main activation and the existing fixed-clock 14-day boundary case. Native four-case/21-assertion suite and revised judge remain unchanged.
  • Local validation: 482 tests passed, 1 skipped; Skillsaw0 errors/8 warnings, plugin and whitespace checks passed.
  • Review identified an unresolved mutation-authorized compatibility gap: newly imported Step7.5 release Epic creation and release Task parent linkage cannot be expressed by the existing Fullsend action schema/executor. Global no-external-write/no-confirmation restrictions remain intact. Track this in TC6726 as a release blocker; do not claim mutation-authorized integration complete. Current native cases test report-only behavior and do not cover this gap.
  • Proceed with the existing native/ordinary source-bound hosted eval flow using merged trusted main; integration PR remains draft.

Hosted native validation succeeded

Run37617450108: 21/21 native assertions passed, complete=true, exit_code=0. Head 43200c8265054f13233a4960a2dc2e8563c4e913, tested merge 7fcf9fa7120518fa34f61a2d6fbd2e9345ae8b01, trusted main ab0c65ae2e3f184b5ddb878efb3afcb50929a9fe, reviewed eval/judge source 91698d4dca24bc199e763dc6462dc54523e473c0. Both corrected absent/empty runtime evidence checks pass. Ordinary validation is incomplete despite a green workflow; see the latest result below. Mutation-authorized release-Epic/parent compatibility remains a separate release blocker; this report-only native suite does not cover it.

Ordinary validation incomplete — do not merge

Run37617450108 completed green, but its ordinary review reports No results produced for both skills. The native21/21 result remains valid.

  • triage-security narrated164/164 and reported writing to pr-head/eval-workspace/, rather than the requested /tmp/triage-security-eval-pr. The expected workspace listing is empty; narrated counts are not verified results.
  • verify-pr was stopped by Claude Code’s ten-minute background wait ceiling while eval1 remained unfinished. Some case grades exist, but no complete root summary was produced. No68/68 claim is justified.
  • The trusted ordinary workflow checks process exit but does not fail for missing summaries; its publisher emits a placeholder and the combined status accepts job success. This is a CI result-completeness gap, tracked in TC6726.

Next: fix ordinary workspace/result-completeness enforcement and investigate unfinished verify-pr eval1 before another paid run. Keep existing assertions and credential isolation intact. PR299 remains draft; ordinary validation and the mutation-authorized release-Epic/parent compatibility blocker remain unresolved.

mrizzi and others added 30 commits August 27, 2026 19:15
Confirm fullsend v0.37.0 toolchain, stock claude-runtime image binaries
(all present -> TC-5806 no-op), and root-level harness resolution of the
in-place sdlc-workflow plugin.

Note v0.37.0 constraints for TC-5807: --fullsend-dir must be the repo
root for relative children to resolve, and verify-pr is a valid harness
role but not a valid config.yaml role (enum: fullsend/triage/coder/
review/fix/retro/prioritize/e2e).

Implements TC-5805

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Author harness/verify-pr.yaml at the repo root so its relative children resolve against the repo root and plugins/sdlc-workflow is delivered in place as a whole plugin. Declares every field explicitly with no base composition; omits security: (Go defaults supply it) and any forge block (GitHub is tier-1, Vertex-only provider). Split-trust env: Jira/GitHub tokens on the runner only, read-only context (JIRA_ISSUE_ID, JIRA_BASE_URL) in the sandbox. Also adds the sandbox agent prompt (agents/verify-pr.md) and the Vertex env file (env/gcp-vertex.env).

Referenced children (providers, policy, profile, schema, pre/post scripts) are authored by later epic tasks; this task validates the harness syntactically only.

Implements TC-5807

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Vertex is the sole in-sandbox provider (tier 4). Jira and GitHub are tier-1 (prefetched host-side) and need no in-sandbox provider, so only the Vertex pair is declared locally for the standalone harness.

Copied verbatim from the stock agents Vertex pair: provider type fullsend-vertex-ai matches profile id fullsend-vertex-ai (endpoint *.googleapis.com:443), which the root harness/verify-pr.yaml references via providers:/openshell.profiles:.

Implements TC-5808

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add plugins/sdlc-workflow/policies/verify-pr.yaml for the verify-pr fullsend harness (referenced by harness/verify-pr.yaml). Filesystem is read-only (include_workdir: false); egress is reduced to Anthropic + Vertex AI (*.googleapis.com) inference and telemetry only.

No *.atlassian.net (Jira) and no api.github.com (GitHub) egress: those are tier-1 with runner-only tokens under the split-trust I/O model (prefetch/post-script run host-side). curl and gh are excluded from the binary allowlist to block raw HTTP with injected tokens.

Live fullsend egress verification is deferred to TC-5810/TC-5811, which author the harness's remaining siblings (pre/post scripts, result schema).

Implements TC-5809

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The verify-pr fullsend policy granted Vertex AI (*.googleapis.com) egress but omitted the pi runtime from the binary allowlist. Under fullsend's dual host-AND-binary match rule, the pi runtime — which brokers Vertex inference — would be blocked. Add **/pi alongside **/claude and **/node to match the composed fullsend-vertex-ai profile and every upstream vertex policy.

Implements TC-5880

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…on_loop

Add the split-trust I/O contract for the verify-pr fullsend harness: the
pre_script prefetches everything on the trusted runner (where the Jira and
GitHub tokens live) and the sandbox reads only JSON — no api.github.com or
atlassian egress.

- schemas/verify-pr-input.schema.json — tracker-agnostic pre-script input,
  extended with a `github` tier-1 read bundle (diff, stat, reviews,
  review_comments, issue_comments, commits, headRefName, commit_sha).
- schemas/verify-pr-result.schema.json — agent result contract (ported).
- scripts/pre-verify-pr.sh — validates env, fetches the Jira issue, resolves
  the linked PR, prefetches every GitHub read the skill performs, and writes
  verify-pr-input.json for host_files to mount.
- scripts/pre_verify_pr.py — PR-URL extraction + input transform, extended
  with build_github_bundle and a --github-dir transform mode.
- scripts/{validate-output-schema.sh,strip_extra_properties.py} — kept per the
  Step 5 native-validator decision (below); wired via validation_loop.script.
- scripts/test_pre_verify_pr.py — unit tests incl. the github bundle.
- harness/verify-pr.yaml — set validation_loop.script.

Deviations from the literal task steps, dictated by the verified fullsend
v0.37.0 contract (same class as TC-5805):

- Skip signal (step 4) uses the pre-script output protocol v1 — line-based
  `skipped=true` / `reason=...` appended to $FULLSEND_PRESCRIPT_OUTPUT (guarded,
  exit 0) — not a JSON `{"skipped":true}` document. That env var is the
  skip-signal file, not the input path.
- GitHub bundle is embedded in verify-pr-input.json (single file) rather than
  separate in-sandbox file paths, keeping the harness change scoped to
  validation_loop; input JSON stays at /tmp/fullsend-pre-output (host_files src),
  with PRE_DIR overridable for host-side testing.
- No host-side PR-head checkout (step 3): the prefetched PR diff is sufficient
  for verification, so the sandbox needs no writable checkout.
- Native validator decision (step 5): additionalProperties:false rejects benign
  agent extras, so validate-output-schema.sh + strip_extra_properties.py are
  KEPT and validation_loop.script is set (fullsend also requires a script for
  validation_loop — schema alone is insufficient).

Validated host-side: both schemas are valid JSON; 15 unit tests pass; the
validator strips benign extras and passes, and fails a bad enum (exit 1); the
skip path exits 0 with the correct signal; the happy path (stub gh + jira)
produces a verify-pr-input.json that validates against the input schema with
the full github bundle embedded.

Implements TC-5810

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Assisted-by: Claude Code
gh pr diff has no --stat flag, so the prefetch aborted under
set -euo pipefail before writing verify-pr-input.json. Derive the
per-file stat from the already-fetched patch with git apply --stat,
guarding the empty-diff case so the script still exits 0. Add
regression tests exercising the real stat command and guarding
against reintroducing the unsupported gh flag.

Implements TC-5884

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Assisted-by: Claude Code
The GitHub tier-1 prefetch fetched PR reviews, review comments, and issue
comments without --paginate, so GitHub REST returned only the first ~30
items and the bundle silently truncated on active PRs — degrading the very
review-comment analysis it feeds the verify-pr sub-agents. Add --paginate
--slurp and merge pages with `jq 'add'` into the flat arrays that
build_github_bundle expects. (--slurp is incompatible with gh's built-in
--jq, so a standalone jq performs the merge.)

Also make the sibling git-apply-stat test deterministic by running it from
a non-repo cwd: git apply --stat is CWD-sensitive and reports "0 files
changed" when launched from inside a repo subdirectory.

Implements TC-5885

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
transform_to_input set task.description via fields.get("description", {}),
whose default applies only to an ABSENT key. When Jira returns an explicit
null (an issue with no description), .get returns None, so task.description
became null — violating verify-pr-input.schema.json (which requires an
object) and producing a schema-invalid verify-pr-input.json. Use
`fields.get("description") or {}` to also coerce null to {}, matching the
existing `(fields.get("status") or {})` idiom in the same transform.

Add regression tests: a direct coercion assertion and a full-input
(task + github bundle) validation against verify-pr-input.schema.json that
fails on the pre-fix null description.

Implements TC-5886

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The github.com PR-URL match at pre-verify-pr.sh:100 was start- but not
end-anchored, so a malformed value like `.../pull/42abc` or
`.../pull/42/invalid` matched and BASH_REMATCH truncated the pull number
to 42 — making the pre-script prefetch and embed a different PR than the
Jira custom field identifies. Append `/?$` so the full URL must match: a
trailing non-numeric character or extra path segment is now rejected with
the existing "not a github.com pull request URL" error, while a canonical
URL with or without a trailing slash still parses to owner/repo and number.

Add regression tests that run the script's actual regex (extracted from
source) through bash's [[ =~ ]], asserting well-formed URLs parse to the
correct owner/repo and number and malformed ones are rejected rather than
truncated.

Implements TC-5891

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
gh's `pr view` commits connection is bounded, so `.commits[-1].oid` returns
the last commit of a truncated page rather than the PR head on large PRs — a
silently-wrong commit_sha that still satisfies the result schema's hex pattern.
Read the head ref tip OID directly (`--json headRefOid`), which is correct
regardless of commit count. Also document that the bundled commits list uses the
same bounded connection (best-effort context; head SHA is authoritative).

Adds a behavioral regression test that runs the shipped COMMIT_SHA command
against a gh stub whose headRefOid and commits[-1].oid disagree, plus a source
guard, both of which fail on the old logic and pass on the new.

Implements TC-5924

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…rivations

Two field descriptions in the `github` read bundle of
verify-pr-input.schema.json cited commands that no longer produce their
values (doc-string drift only; the field values are correct):

- `commit_sha` cited `.commits[-1].oid`, but pre-verify-pr.sh derives it
  from `gh pr view --json headRefOid` (TC-5924 — the pr-view commits
  connection is bounded, so `.commits[-1]` can be a truncated-page tip).
- `stat` cited `gh pr diff --stat`, an unsupported flag; it is derived
  from `git apply --stat` on the already-fetched pr.diff (TC-5884).

Documentation-only; the schema remains valid JSON.

Implements TC-5929
Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Port the verify-pr write executor to run host-side after the sandbox is
destroyed. Jira comments and the verification report are posted through the
native `fullsend issues post-comment` sticky-comment CLI (idempotent via a
stable marker), while the irreducible custom Jira writes with no native
primitive yet — sub-tasks, links, and root-cause tasks — call jira-client.py
directly. GitHub writes stay host-side on `gh`.

- post-verify-pr.sh: locate the latest agent-result.json, validate, delegate
- execute-actions.py: resolve {{ref.key}}/{{ref.url}} placeholders, route
  post_comment/post_report Jira side through the native CLI, keep gh for
  GitHub, keep jira-client.py for sub-tasks/links/root-cause
- jira-client.py: create_issue accepts a pre-rendered ADF description and an
  optional parent key (reconciled additively; existing callers unaffected)
- test_execute_actions.py: ref-resolution tests plus native-CLI argv/env,
  stdin body, and post_comment/post_report routing tests

Implements TC-5811

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ira CLI

execute_post_comment read action["body_md"], but the result schema requires
body_adf (an ADF object) and strip_extra_properties.py removes any extra key,
so every schema-valid post_comment raised KeyError before posting. Add
adf_to_markdown to render the ADF body back to the markdown the native fullsend
CLI consumes on stdin, and update the masking test to feed a schema-valid
body_adf action.

TC-5930

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
post-verify-pr.sh only located agent-result.json, but its sibling
validate-output-schema.sh accepts result.json as a fallback. A result
that passed validation via the fallback name was then rejected by the
write path. Prefer agent-result.json per iteration output dir, and fall
back to result.json when absent — matching the validator's precedence.

Implements TC-5931

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
post-verify-pr.sh selected the "most recent" iteration by keeping the last
match in shell glob order, which is lexicographic. Once there are >= 10
iterations, iteration-9 sorts after iteration-20 and a stale iteration is
chosen, so the write path would execute the wrong (older) result.

Iterate the iteration-*/output directories in ascending numeric order via
`sort -V` and keep the highest-numbered one that has a result file. Preserves
the `set -euo pipefail` behavior and the result.json fallback (TC-5931); a
`[[ -d ]]` guard skips the literal glob pattern when no iteration dir exists.

TC-5932

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The ADF block renderer had no taskList/taskItem case, so a schema-valid post_comment task list fell through to the unknown-block fallback and was flattened to plain paragraph text, losing the checklist markers. Add _render_adf_task_list rendering DONE/TODO items as markdown checkboxes, handling both paragraph-wrapped and inline taskItem content. The comment route calls adf_to_markdown directly (not sanitize_adf), so the renderer is the correct fix site.

Implements TC-5936

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… retry

execute_post_report posted the verification report to GitHub via a plain gh pr comment (no sticky marker) before the idempotent Jira post, so a retry after a failed Jira post double-posted the GitHub report comment. Embed a commit-scoped marker in the GitHub report body and, before posting, list the PR comments and PATCH-update an existing same-commit report comment instead of creating a duplicate. Dedup is scoped to the commit SHA so a later commit still gets a fresh comment, preserving the per-run verification history. The Jira side is unchanged and receives the clean body. Updated the module docstring so its idempotency claim matches actual behaviour.

Implements TC-5937

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
_find_report_comment_id listed a PR's issue comments with `gh api ... --paginate`
and parsed the result with a single json.loads. Without `--slurp`, `gh --paginate`
concatenates one JSON array per page (`[...][...]`), which is not valid combined
JSON once the PR has more than one page of comments (>30). json.loads then raised
JSONDecodeError, which the code swallowed by returning None; execute_post_report
read that as "no existing report comment" and created a duplicate GitHub report
comment instead of PATCH-updating the existing one, silently defeating the retry
idempotency TC-5937 introduced.

Add `--slurp` so gh emits a single array-of-pages and flatten the pages into one
comment list, so a same-commit report comment is found even when it lands on a
later page. Stop swallowing JSONDecodeError: with `--slurp` a parse error is a
real failure and now surfaces via sys.exit(1) rather than being misread as "no
existing comment". Added a multi-page test (report comment on a non-first page →
PATCH-update, not duplicate) and a parse-failure test (exits instead of returning
None).

Implements TC-5948

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…teral text

_render_adf_inline emitted a text node's literal value verbatim into the
markdown handed to the native `fullsend issues post-comment` CLI. Because that
CLI re-parses the markdown, literal markdown-active sequences inside ADF text
(`*`, `_`, `[`, `]`, backticks, and backslash) were reinterpreted as formatting
— e.g. a literal `*note*` rendered as emphasized text and `[x](y)` as a link.

Add _escape_markdown, which backslash-escapes those characters (backslash first
so its escapes are not re-escaped), and apply it to a text node's plain value
before the mark wrapping the renderer intentionally adds, so the mark syntax is
not double-escaped. Inline code is exempt: its content is literal to the CLI and
escaping would corrupt inline-code semantics. Link hrefs are also left untouched.
Added tests: literal `* _ [ ]` and backtick are escaped; strong/link marks,
inline-code content, and a link href with an underscore are unaffected.

Implements TC-5949

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
_render_adf_inline previously handled only text and hardBreak nodes; every other inline node fell into the else branch and recursed into a nonexistent content array, so mention/emoji/inlineCard/date/status inline nodes (whose value lives in attrs, not content) rendered as empty strings and were silently dropped -- contradicting the module's own _INLINE_NODE_TYPES set.

Render each from its attrs: mention/emoji/status from attrs.text (escaped as literal text; emoji falls back to attrs.shortName), inlineCard from attrs.url (bare, unescaped), and date from attrs.timestamp via a new _render_adf_date helper (epoch millis to UTC YYYY-MM-DD). The final else is kept for genuinely unknown container-like inline nodes. Works in both paragraph and taskItem inline contexts.

Implements TC-5950

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…result schema

The post_comment action schema accepted any non-empty string for "issue", but the executor post_jira_comment_native requires a hyphenated Jira key and sys.exit(1)s otherwise — a schema-valid {"issue":"12345"} (numeric ID, URL, or other non-key identifier) passed validation and then failed at execution, aborting the whole post_script run.

Constrain the schema "issue" field to the Jira-key pattern ^[A-Z]+-[0-9]+$ (matching jira_issue_id/parent in the same schema) so non-key identifiers are rejected at the producer boundary. The executor rpartition fail-fast guard is kept as defence in depth. Add schema-validation tests asserting the pattern accepts a valid key and rejects non-key issues (numeric ID, URL, lowercase, missing hyphen/number/project).

Implements TC-5953

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Step 9 'Post to GitHub PR' claimed every verification run creates a new PR
comment and never overwrites previous reports, and showed a plain
`gh pr comment` create. That contradicted the shipped executor behavior
(TC-5937): execute_post_report embeds a commit-scoped marker and, via
_find_report_comment_id, PATCH-updates an existing same-commit report comment
instead of duplicating it — a new comment is created only for a later commit.

Rewrite Step 9 to match the code: describe the commit-scoped marker, the
find-then-PATCH-or-create path (gh api --paginate --slurp + PATCH, else
gh pr comment), and reframe the history claim as per-commit (a later commit
gets a new comment) rather than per-run. Keep the (commit <short-sha>) header
convention. Add an eval assertion (evals/verify-pr/evals.json case 3) covering
the commit-scoped report header / per-commit history.

Documentation-only alignment; executor behavior is unchanged.

Implements TC-5954

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…schema

TC-5953 tightened the post_comment 'issue' field to the Jira-key pattern
'^[A-Z]+-[0-9]+$', but execute_post_comment resolves the field through
resolve_refs, which supports the {{<ref>.key}} placeholder so a comment can
target an issue created by an earlier action. Since fullsend's validation_loop
validates the result schema before post_script runs, a placeholder issue was
rejected before the reference could be resolved.

Widen the pattern to '^([A-Z]+-[0-9]+|\\{\\{[a-z0-9-]+\\.key\\}\\})$' so it
accepts either a literal Jira key or a {{<ref>.key}} placeholder, scoped to
.key (not .url) since the field must resolve to a key. Bare numeric IDs, URLs,
and other non-key/non-ref identifiers stay rejected (TC-5953's intent).

Add a schema-validation test that a {{<ref>.key}} post_comment action passes
jsonschema.validate, and extend the negative test with .url placeholders,
uppercase refs, and unanchored/brace-less forms.

Implements TC-5959

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
datetime.fromtimestamp can raise OverflowError/OSError (or ValueError) for an
out-of-range epoch-ms value. The call sat outside the guarded try, so a single
malformed ADF date node in a post_comment/post_report body aborted the entire
execute-actions.py post_script and posted nothing. Move the fromtimestamp call
inside the try and catch OverflowError/OSError alongside the int() parse errors,
falling back to the literal timestamp string so the node degrades gracefully.

Implements TC-5965

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
create_issue declared issue_type: str = "" (reconciled under TC-5811 for the
new description_adf/parent params), so an omitted value serialized to
{"name": ""} and surfaced as an opaque Jira 400 instead of failing fast at
the call site. Restore the prior fail-fast contract: reject an empty (or
whitespace-only) issue_type with a clear stderr message and sys.exit(1)
before any HTTP request, matching the module's existing error style.

Added test_create_issue_fails_fast_on_empty_issue_type (asserts SystemExit
and no request issued); jira-client suite 19 -> 20 passing.

Implements TC-5966

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
post_comment and post_report both posted via post_jira_comment_native with
the single shared STICKY_COMMENT_MARKER, which the native fullsend CLI treats
as the sticky-comment identity. Two comments targeting the same Jira issue in
one run would share that marker, so the second post would overwrite the first
and silently lose a comment.

Thread a marker parameter through post_jira_comment_native (default preserves
the report path's historical marker) and give post_comment a distinct
POST_COMMENT_STICKY_MARKER. Each path keeps its own stable marker, so per-path
re-run idempotency is preserved while cross-purpose clobbering is prevented.
GitHub report dedup (GITHUB_REPORT_MARKER_PREFIX) is untouched.

Added test_post_comment_and_report_use_distinct_sticky_markers; execute_actions
suite 26 -> 27 passing.

Implements TC-5967

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
execute_post_report embedded the raw commit_sha in the GitHub dedup marker and
_find_report_comment_id matched it by substring. The result schema permits
commit_sha to be 7-40 hex chars, so a run recording a full SHA and a retry
recording an abbreviation (or vice-versa) for the same commit produced
different marker strings; the substring lookup missed and the retry posted a
duplicate report comment, defeating the TC-5937 idempotency guarantee.

Normalize the SHA to the schema-minimum 7 chars in the marker (and thus the
lookup, which reuses the same marker). Git abbreviations are always prefixes of
the full SHA, so first-7 is the only fixed length that unifies a full SHA with
any valid abbreviation of the same commit; a longer length (e.g. 12) would
leave a 7-char abbreviation unchanged and still mismatch a full SHA.

Added test_execute_post_report_dedup_marker_invariant_to_sha_length (full-SHA
create then abbreviated-SHA retry -> single PATCH-updated comment);
execute_actions suite 27 -> 28 passing.

Implements TC-5968

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
execute_post_report calls resolve_refs on report_md, which raises an
uncaught KeyError for any {{ref.key}} placeholder whose entity has not yet
been registered. The registry is populated by create_subtask /
create_root_cause_task as main()'s sequential loop runs, so a post_report
ordered before an action it references would abort the whole run with a bare
KeyError. verify-pr emits post_report last today, but that ordering was
undocumented and unenforced.

Defer every post_report until after the actions loop in main() so the
registry is fully populated before any report_md is resolved. post_report is
still a recognized action; only its execution moves to the end. Observed
behavior is unchanged for the always-last emission order.

Added test_post_report_resolves_ref_created_by_later_action, which drives
main() with a temp result JSON that orders post_report before the
create_subtask its report_md interpolates and asserts the resolved key
reaches the GitHub report body; execute_actions suite 28 -> 29 passing.

Implements TC-5969

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
_render_adf_block handled paragraph/heading/rule/codeBlock/bullet+ordered
list/taskList; every other block hit the unknown-block fallback that recurses
into content and flattens structure. So table cells were concatenated with no
grid, blockquote/panel lost their framing, and media/mediaSingle/mediaGroup
(which carry no text content) were dropped entirely. A schema-valid
post_comment.body_adf can carry any of these, so a comment silently lost
content — the same silent-degradation class already closed for taskList
(TC-5936) and inline leaf nodes (TC-5950).

Add explicit renderers: blockquote/panel -> "> "-prefixed lines (panelType
surfaced as a bold label), table -> a GitHub-flavored markdown table (first
tableRow is the header, pipes in cells escaped), and media/mediaSingle/
mediaGroup -> an image link or a non-empty [alt] placeholder, never dropped.
The recursing fallback is kept for genuinely unknown container nodes, and the
adf_to_markdown docstring's node-set contract now lists the covered nodes.

Added rendering tests for table, blockquote+panel, and media, plus one
asserting an unknown block still degrades via the fallback; execute_actions
suite 29 -> 33 passing.

Implements TC-5970

Assisted-by: Claude Code
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Request author, marker, footer, match-count and Test Quality impact evidence from the supplied Reviews list in outputs/report.md. Add focused prompt and fixture regression checks while preserving all six cases and 68 assertion strings.

Implements TC-4636

Assisted-by: Claude Code
mrizzi added 4 commits October 5, 2026 19:23
Deliver all native cases, fixtures, judges, dependency pins and trusted suite execution through the integration branch. Keep ordinary evals and the TC4636 repair unchanged. Include separate host/sandbox credentials, reviewed host validation and lexical symlink rejection.

Implements TC-6677 and TC-6726

Assisted-by: Claude Code
Keep suite execution on trusted main after PR299 lands, with explicit suite source provenance, existing approval and immutable plugin source. Preserve current PR299 ordinary rendering and sticky comment behavior.

Implements TC-6726

Assisted-by: Claude Code
Hosted unit-test collection requires PyYAML for the native suite and workflow contract tests. Preserve the four-file main bootstrap and its reviewed source pin.

Implements TC-6726

Assisted-by: Claude Code
Preserve PR299 activation, rendering and sticky reporting while retaining all merged bootstrap safeguards. Recheck the current PR head before replacing the shared ordinary report.

Implements TC-6726

Assisted-by: Claude Code

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Eval Results

Source head: fd5c06f
Merge: 865f367

Eval Summary: triage-security

Per-Eval Results

Eval Passed Failed Pass Rate
1 11 0 100.0%
2 5 0 100.0%
3 5 0 100.0%
4 5 0 100.0%
5 6 0 100.0%
6 6 0 100.0%
7 5 0 100.0%
8 8 0 100.0%
9 5 0 100.0%
10 5 0 100.0%
11 5 0 100.0%
12 5 0 100.0%
13 5 0 100.0%
14 5 0 100.0%
15 5 0 100.0%
16 7 0 100.0%
17 5 0 100.0%
18 5 0 100.0%
19 5 0 100.0%
20 4 0 100.0%
21 4 0 100.0%
22 4 0 100.0%
23 4 0 100.0%
24 4 0 100.0%
25 4 0 100.0%
26 5 0 100.0%
27 5 0 100.0%
28 5 0 100.0%
29 5 0 100.0%
30 4 0 100.0%
31 4 0 100.0%
32 4 0 100.0%

Aggregate Stats

Metric Mean Stddev
Pass Rate 100.0% 0.0%
Time (seconds) 120.14 48.64
Tokens 54,658 9,060

Summary

  • Total evals: 32
  • Total assertions: 164
  • Total passed: 164
  • Total failed: 0
  • Overall pass rate: 100.0%

No baseline comparison available.


triage-security eval run

verify-pr

No results produced. See workflow logs.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Native Fullsend Eval Results

Head: fd5c06f
Merge: 865f367
Trusted workflow: 805068c
Reviewed eval source: b46b2bf

Native job: failure; complete: false; 0/21 passed.

Implements TC-6726

Assisted-by: Claude Code

@github-actions github-actions Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Native Fullsend Eval Results

Head: c7ca849
Merge: 13ab8ee
Trusted workflow: 710a4d6
Reviewed eval source: c7ca849

No safe native result was produced; native execution/approval failed.

Preserve PR299 trusted-main activation while incorporating human-merged PR326. The resulting tree is identical to the previous PR299 head.

Implements TC-6726

Assisted-by: Claude Code

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Eval Results

Source head: e3b2eae
Merge: 26de0bf

Eval Results: triage-security

Eval Passed Failed Pass Rate
eval-1 11/11 0 100%
eval-2 5/5 0 100%
eval-3 5/5 0 100%
eval-4 5/5 0 100%
eval-5 6/6 0 100%
eval-6 6/6 0 100%
eval-7 5/5 0 100%
eval-8 8/8 0 100%
eval-9 5/5 0 100%
eval-10 5/5 0 100%
eval-11 5/5 0 100%
eval-12 5/5 0 100%
eval-13 5/5 0 100%
eval-14 5/5 0 100%
eval-15 5/5 0 100%
eval-16 7/7 0 100%
eval-17 5/5 0 100%
eval-18 5/5 0 100%
eval-19 5/5 0 100%
eval-20 4/4 0 100%
eval-21 4/4 0 100%
eval-22 4/4 0 100%
eval-23 4/4 0 100%
eval-24 4/4 0 100%
eval-25 4/4 0 100%
eval-26 5/5 0 100%
eval-27 5/5 0 100%
eval-28 5/5 0 100%
eval-29 5/5 0 100%
eval-30 4/4 0 100%
eval-31 4/4 0 100%
eval-32 4/4 0 100%

Pass rate: 100% · Tokens: 59,115 · Duration: 129s


Generated by sdlc-workflow/run-evals v0.13.9

Eval Results

Skill: verify-pr
Date: 2026-10-06
Overall: 68/68 assertions passed (100.0%)

Per-Eval Results

# Name Pass Fail Total Rate Duration Tokens
1 passing 12 0 12 100% 745.5s 76,337
2 failing-criteria 11 0 11 100% 247.4s 69,185
3 review-comments 15 0 15 100% 400.9s 82,194
4 adversarial-injection 10 0 10 100% 265.0s 70,446
5 test-change-classification 10 0 10 100% 417.3s 84,134
6 eval-result-processing 10 0 10 100% 397.6s 83,360

Aggregate Statistics

Metric Mean Stddev Min Max
Duration (s) 412.3 163.4 247.4 745.5
Tokens 77,609 6,063 69,185 84,134

Failing Assertions

None.


Generated by sdlc-workflow/run-evals

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Native Fullsend Eval Results

Head: e3b2eae
Merge: 26de0bf
Trusted workflow: 710a4d6
Reviewed eval source: c7ca849

Native job: failure; complete: false; 0/21 passed.

Preserve both the existing integration tests and the approved judge forwarding regression.

Implements TC-6726

Assisted-by: Claude Code

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Eval Results

Source head: 6a251f7
Merge: 07aaca5

Eval Results: triage-security

Eval Passed Failed Pass Rate
eval-1 11/11 0 100%
eval-2 5/5 0 100%
eval-3 5/5 0 100%
eval-4 5/5 0 100%
eval-5 6/6 0 100%
eval-6 6/6 0 100%
eval-7 5/5 0 100%
eval-8 8/8 0 100%
eval-9 5/5 0 100%
eval-10 5/5 0 100%
eval-11 5/5 0 100%
eval-12 5/5 0 100%
eval-13 5/5 0 100%
eval-14 5/5 0 100%
eval-15 5/5 0 100%
eval-16 7/7 0 100%
eval-17 5/5 0 100%
eval-18 5/5 0 100%
eval-19 5/5 0 100%
eval-20 4/4 0 100%
eval-21 4/4 0 100%
eval-22 4/4 0 100%
eval-23 4/4 0 100%
eval-24 4/4 0 100%
eval-25 4/4 0 100%
eval-26 5/5 0 100%
eval-27 5/5 0 100%
eval-28 5/5 0 100%
eval-29 5/5 0 100%
eval-30 4/4 0 100%
eval-31 4/4 0 100%
eval-32 4/4 0 100%

Pass rate: 100% · Tokens: 56,016 · Duration: 115s


Generated by sdlc-workflow/run-evals v0.13.9

Eval Results: verify-pr

Eval Passed Failed Pass Rate
eval-1 12/12 0 100%
eval-2 11/11 0 100%
eval-3 15/15 0 100%
eval-4 10/10 0 100%
eval-5 10/10 0 100%
eval-6 10/10 0 100%

Pass rate: 100% · Tokens: 100,582 · Duration: 730s


Generated by sdlc-workflow/run-evals v0.13.9

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Native Fullsend Eval Results

Head: 6a251f7
Merge: 07aaca5
Trusted workflow: 27a258d
Reviewed eval source: c7ca849

Native job: failure; complete: true; 19/21 passed.

Accept genuine absent/empty Skill execution up to its required stopping boundary while preserving native evidence and bootstrap failure checks.

Implements TC-6764

Assisted-by: Claude Code
Preserve post-integration trusted-main activation and the existing fixed-clock boundary eval while importing current interactive triage changes.

Implements TC-6726

Assisted-by: Claude Code

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Eval Results

Source head: 43200c8
Merge: 7fcf9fa

triage-security

No results produced. See workflow logs.

verify-pr

No results produced. See workflow logs.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Native Fullsend Eval Results

Head: 43200c8
Merge: 7fcf9fa
Trusted workflow: ab0c65a
Reviewed eval source: 91698d4

Native job: success; complete: true; 21/21 passed.

mrizzi added 2 commits October 7, 2026 17:12
Refs TC-6726

Assisted-by: Claude Code
Refs TC-6726

Assisted-by: Claude Code

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Native Fullsend Eval Results

Head: 89ac2ea
Merge: b91e6e5
Trusted workflow: 9aa885b
Reviewed eval source: 91698d4

Native job: failure; complete: true; 19/21 passed.

Merge main through PR #330 to run the source-bound PR #299 eval with the background wait policy and bounded CI timeout.

Refs TC-6726, TC-6783

Assisted-by: Codex

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Eval Results

Source head: e19f2e0
Merge: e09aa55

Eval Results: triage-security

Eval Passed Failed Pass Rate
eval-1 11/11 0 100%
eval-10 5/5 0 100%
eval-11 5/5 0 100%
eval-12 5/5 0 100%
eval-13 5/5 0 100%
eval-14 5/5 0 100%
eval-15 5/5 0 100%
eval-16 7/7 0 100%
eval-17 5/5 0 100%
eval-18 5/5 0 100%
eval-19 5/5 0 100%
eval-2 5/5 0 100%
eval-20 4/4 0 100%
eval-21 4/4 0 100%
eval-22 4/4 0 100%
eval-23 4/4 0 100%
eval-24 4/4 0 100%
eval-25 4/4 0 100%
eval-26 5/5 0 100%
eval-27 5/5 0 100%
eval-28 5/5 0 100%
eval-29 5/5 0 100%
eval-3 5/5 0 100%
eval-30 4/4 0 100%
eval-31 4/4 0 100%
eval-32 4/4 0 100%
eval-4 5/5 0 100%
eval-5 6/6 0 100%
eval-6 6/6 0 100%
eval-7 5/5 0 100%
eval-8 8/8 0 100%
eval-9 5/5 0 100%

Pass rate: 100% · Tokens: 58,752 · Duration: 114s

Baseline (f1caf724): 96% · 54,746 tokens · 125s


Generated by sdlc-workflow/run-evals v0.13.10

Eval Results: verify-pr

Eval Passed Failed Pass Rate
eval-1 12/12 0 100%
eval-2 11/11 0 100%
eval-3 15/15 0 100%
eval-4 10/10 0 100%
eval-5 10/10 0 100%
eval-6 10/10 0 100%

Pass rate: 100% · Tokens: 96,236 · Duration: 1202s


Generated by sdlc-workflow/run-evals v0.13.10

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Native Fullsend Eval Results

Head: e19f2e0
Merge: e09aa55
Trusted workflow: f3416b4
Reviewed eval source: 91698d4

Native job: failure; complete: true; 20/21 passed.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant