fix(ci): evaluate measured evidence in pipeline quality gates - #710
Merged
proffesor-for-testing merged 1 commit intoSep 27, 2026
Conversation
proffesor-for-testing
approved these changes
Sep 27, 2026
proffesor-for-testing
left a comment
Owner
There was a problem hiding this comment.
Thanks @rudycelekli — great follow-up to #702. The CI quality-gate phase now uses the same evidence loader and evaluator as aqe quality --gate and MCP quality_assess({runGate:true}), and the integration test proves the parity through the real protocol server and CI command on one kernel. Resetting the aggregate artifact first and writing per-gate artifacts is a nice touch. Locally: typecheck/lint clean, 69/69 focused tests. The Performance Gates red is the known RVF timing flake (fixed separately in #718), not this PR. Merging.
Two small follow-ups we'd welcome:
- With the default enforced gate, runs that don't include the gate (
--phase security, a disabled gate phase, or a config without one) now exit 1 instead of 0. Right fail-closed default, but users of.github/actions/aqewith thephase:input can't pass--no-quality-gate— could you add an advisory/quality-gateinput to the action? - A one-line warning when
.aqe-ci.ymlstill setsquality_gate.thresholds(now ignored) would save someone a confusing afternoon.
proffesor-for-testing
merged commit Sep 27, 2026
4da115a
into
proffesor-for-testing:main
23 of 24 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
aqe ci run --phase quality-gatecurrently calls a nonexistent domainevaluate()method, so even complete passing evidence cannot produce a usable gate. It now calls the same timestamped evidence loader and seven-check evaluator asaqe quality --gateand registered MCPquality_assess({ runGate: true }).This is the reachable CLI follow-up invited in the maintainer's #702 review. Gate reporting now separates the actual result from enforcement: omitted/skipped gates cannot become approvals, and
--no-quality-gatekeeps a failed evaluation visible while allowing later phases to continue. Other phase failures remain fatal. An aggregate gate artifact is reset before execution, and individual gates retain separate artifacts.Verification
d47ec76ebefore the product patch, including the passing-evidence CI control.npx vitest run tests/integration/quality-assess-measured-gate.test.ts tests/unit/cli/commands/ci.test.ts tests/unit/cli/quality-evidence.test.ts tests/unit/cli/ci-config.test.ts tests/unit/cli/ci-output.test.ts tests/unit/domains/quality-assessment/quality-evidence.test.ts --maxWorkers=1: 69 passed across 6 files.npm run typecheck,npm run build(TypeScript, ESM, CLI and MCP), and ESLint on the two changed product files passed.HybridMemoryBackendevidence: baseline healthy evidence exits 1 withevaluate is not a function, and baseline filtered gate exits 0 withqualityGatePassed: trueand no executed phases. Patched healthy evidence exits 0 with 7/7 checks; a failing test metric, missing evidence, and 48-hour-old evidence each exit 1; filtering the gate exits 1/not-run; explicit advisory mode exits 0/warning while retaining the failed gate. All 8 baseline/patched cases passed their expected assertions, including JSON/Markdown/artifact agreement.GitHub's full test/coverage execution completed with 23,789 tests passed, 62 skipped, zero failed (988 files passed, 6 skipped).
GitHub currently reports a failure in the unchanged native RVF benchmark (
tests/performance/rvf-pattern-store.test.ts, “ingest 1000 patterns into native HNSW”): 10,125 ms against its existing 10,000 ms bound. The same main-revision check passed at 1,620 ms. This PR changes neither that test, its implementation, dependencies, nor its workflow. The attempted job rerun was denied by GitHub with HTTP 403, so a maintainer rerun is needed to check repeatability. No timing bounds or assertions were changed.Follow-up measurements: the same native gate passed on #711 at 1,779 ms and #712 at 1,644 ms. Main and both passing controls used Ubuntu image
20260907.300.1; the failed run used20260920.314.1, all with Node 24.13.0. That is correlation, not a demonstrated cause. Eight relevant source/config/lock files have identical Git blobs across all four heads.One unchanged exact-head local native run also failed: ingest 10,032 ms, with all 1,000 vectors verified; the subsequent ingest/search case exceeded its whole-test 10-second timeout, although query p95 was 0.31 ms. Native availability and cold-start passed. This is not being dismissed as transient-only. An authorized rerun is the next measurement; a persistent failure needs profiling before any performance repair. No timing limit or assertion was changed. The existing native benchmark exercises both reported limits.
Failure modes
quality-assess-measured-gate.test.tsThe change uses the canonical fixed thresholds; it does not introduce CI threshold customization, alter YAML parsing, or change other phase implementations. Revision/run binding for independently produced metrics remains part of the broader evidence work in #651.
docs/ci-measured-quality-gate.mddescribes the output and exit-code contract, including intentional advisory runs and partial pipelines.