Skip to content

fix(ci): evaluate measured evidence in pipeline quality gates - #710

Merged
proffesor-for-testing merged 1 commit into
proffesor-for-testing:mainfrom
rudycelekli:fix/ci-measured-quality-gate
Sep 27, 2026
Merged

proffesor-for-testing merged 1 commit into
proffesor-for-testing:mainfrom
rudycelekli:fix/ci-measured-quality-gate

Conversation

@rudycelekli

@rudycelekli rudycelekli commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

aqe ci run --phase quality-gate currently calls a nonexistent domain evaluate() method, so even complete passing evidence cannot produce a usable gate. It now calls the same timestamped evidence loader and seven-check evaluator as aqe quality --gate and registered MCP quality_assess({ runGate: true }).

This is the reachable CLI follow-up invited in the maintainer's #702 review. Gate reporting now separates the actual result from enforcement: omitted/skipped gates cannot become approvals, and --no-quality-gate keeps a failed evaluation visible while allowing later phases to continue. Other phase failures remain fatal. An aggregate gate artifact is reset before execution, and individual gates retain separate artifacts.

Verification

  • Regression first: the extended real-kernel/MCP integration suite failed against d47ec76e before the product patch, including the passing-evidence CI control.
  • npx vitest run tests/integration/quality-assess-measured-gate.test.ts tests/unit/cli/commands/ci.test.ts tests/unit/cli/quality-evidence.test.ts tests/unit/cli/ci-config.test.ts tests/unit/cli/ci-output.test.ts tests/unit/domains/quality-assessment/quality-evidence.test.ts --maxWorkers=1: 69 passed across 6 files.
  • npm run typecheck, npm run build (TypeScript, ESM, CLI and MCP), and ESLint on the two changed product files passed.
  • An independent command-level reproducer found a previous approval artifact surviving an omitted gate. The same pass-then-omit sequence passes after adding aggregate invalidation and separate phase artifacts.
  • Built CLI verification in separate subprocesses with actual persisted HybridMemoryBackend evidence: baseline healthy evidence exits 1 with evaluate is not a function, and baseline filtered gate exits 0 with qualityGatePassed: true and no executed phases. Patched healthy evidence exits 0 with 7/7 checks; a failing test metric, missing evidence, and 48-hour-old evidence each exit 1; filtering the gate exits 1/not-run; explicit advisory mode exits 0/warning while retaining the failed gate. All 8 baseline/patched cases passed their expected assertions, including JSON/Markdown/artifact agreement.
  • Runtime checks use fresh project roots and temporary directories. Existing project learning databases are not test fixtures.

GitHub's full test/coverage execution completed with 23,789 tests passed, 62 skipped, zero failed (988 files passed, 6 skipped).

GitHub currently reports a failure in the unchanged native RVF benchmark (tests/performance/rvf-pattern-store.test.ts, “ingest 1000 patterns into native HNSW”): 10,125 ms against its existing 10,000 ms bound. The same main-revision check passed at 1,620 ms. This PR changes neither that test, its implementation, dependencies, nor its workflow. The attempted job rerun was denied by GitHub with HTTP 403, so a maintainer rerun is needed to check repeatability. No timing bounds or assertions were changed.

Follow-up measurements: the same native gate passed on #711 at 1,779 ms and #712 at 1,644 ms. Main and both passing controls used Ubuntu image 20260907.300.1; the failed run used 20260920.314.1, all with Node 24.13.0. That is correlation, not a demonstrated cause. Eight relevant source/config/lock files have identical Git blobs across all four heads.

One unchanged exact-head local native run also failed: ingest 10,032 ms, with all 1,000 vectors verified; the subsequent ingest/search case exceeded its whole-test 10-second timeout, although query p95 was 0.31 ms. Native availability and cold-start passed. This is not being dismissed as transient-only. An authorized rerun is the next measurement; a persistent failure needs profiling before any performance repair. No timing limit or assertion was changed. The existing native benchmark exercises both reported limits.

Failure modes

Case Executed regression
Complete passing evidence crashes instead of evaluating Real kernel-backed CI command and registered MCP positive control in quality-assess-measured-gate.test.ts
One bad metric is lost behind a static/aggregate score All seven per-metric failing controls compare CI and registered MCP checks
Missing, partial, stale, malformed, or future-dated evidence Same integration suite asserts failed CI results and unavailable artifacts
Omitted or disabled gate is reported as passed Integration pass-then-filter case and disabled-gate unit control
Failed advisory gate stops later work or hides its verdict Unit post-gate sentinel executes; failed phase and aggregate verdict remain visible
Advisory mode suppresses an unrelated phase failure Unit non-gate failure remains exit 1
Previous or later gate approval overwrites a failure Same-directory prior-pass/early-stop and failed-first/passed-second controls verify aggregate and individual artifacts
Earlier success hides a selected gate that was never reached Unit early-stop-between-gates control requires every selected gate to execute

The change uses the canonical fixed thresholds; it does not introduce CI threshold customization, alter YAML parsing, or change other phase implementations. Revision/run binding for independently produced metrics remains part of the broader evidence work in #651. docs/ci-measured-quality-gate.md describes the output and exit-code contract, including intentional advisory runs and partial pipelines.

  • Every failure mode mentioned in this PR description has either (a) a test that exercises it, or (b) a linked tracking issue.

@proffesor-for-testing proffesor-for-testing left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @rudycelekli — great follow-up to #702. The CI quality-gate phase now uses the same evidence loader and evaluator as aqe quality --gate and MCP quality_assess({runGate:true}), and the integration test proves the parity through the real protocol server and CI command on one kernel. Resetting the aggregate artifact first and writing per-gate artifacts is a nice touch. Locally: typecheck/lint clean, 69/69 focused tests. The Performance Gates red is the known RVF timing flake (fixed separately in #718), not this PR. Merging.

Two small follow-ups we'd welcome:

  1. With the default enforced gate, runs that don't include the gate (--phase security, a disabled gate phase, or a config without one) now exit 1 instead of 0. Right fail-closed default, but users of .github/actions/aqe with the phase: input can't pass --no-quality-gate — could you add an advisory/quality-gate input to the action?
  2. A one-line warning when .aqe-ci.yml still sets quality_gate.thresholds (now ignored) would save someone a confusing afternoon.

@proffesor-for-testing
proffesor-for-testing merged commit 4da115a into proffesor-for-testing:main Sep 27, 2026
23 of 24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants