@@ -166,11 +166,11 @@ them; they may not replace them.
166166- Verifier. An agent that runs an adversarial fact-check pass over each draft
167167 before human review: resolving citations, checking claims against the sources
168168 they cite, and flagging anything unsupported. It is prompted to find
169- unsupported claims, not to confirm the draft, so it does not inherit the
170- drafter's blind spots. Its report feeds the reviewer's verification (HC-7); it
171- never substitutes for it. The layering exists because a fabrication that
172- reaches a project's maintainers wastes their time and costs the system its
173- credibility.
169+ unsupported claims, not to confirm the draft, making it less likely to inherit
170+ the drafter's blind spots. Its report feeds the reviewer's verification
171+ (HC-7); it never substitutes for it. The layering exists because a fabrication
172+ that reaches a project's maintainers wastes their time and costs the system
173+ its credibility.
174174- Reviewer. Accepts a request to begin work, then reviews and refines each draft
175175 in conversation, verifies findings against source (HC-7), and marks it ready.
176176 Owns the draft's quality, but does not give its final sign-off; that is
@@ -328,12 +328,13 @@ The system is acceptable when, on a pilot assessment:
328328
329329- Quality. The deliverable is scored against an assessment-quality rubric (its
330330 definition is an open item; section 17). Verification is not a token sample:
331- the reviewer checks every rating-bearing finding, or at minimum a set number
332- per criterion, biased toward the highest-risk claims, and records in the
333- deliverable which findings were verified (HC-7). Final sign-off is given by
334- the approver (section 4), who was neither the drafter nor a reviewer of that
335- phase, to avoid signing off on one's own work. The bar is parity with the
336- human baselines (Flatcar, Knative, Helm).
331+ for the pilot, the reviewer checks every rating-bearing finding, biased first
332+ toward the highest-risk claims, and records in the deliverable which findings
333+ were verified (HC-7); any cheaper sampling rule is future work, set with the
334+ rubric (section 17). Final sign-off is given by the approver (section 4), who
335+ was neither the drafter nor a reviewer of that phase, to avoid signing off on
336+ one's own work. The bar is parity with the human baselines (Flatcar, Knative,
337+ Helm).
337338- Safety. Zero writes outside cncf/techdocs, audited from the platform's
338339 activity records, and no unmitigated prompt-injection incident.
339340- Completeness and reproducibility. Every phase's deliverable produced (all
@@ -402,8 +403,9 @@ infrastructure we build and maintain ourselves.
402403 policy. Policy: the GitHub MCP server stays at its read-only default; any
403404 additional MCP server must be read-only with its tools explicitly allowlisted;
404405 no write-capable MCP tool is permitted. MCP configuration lives in repository
405- settings rather than in a versioned file, so the administrator/platform owner
406- records the current configuration in the operational docs whenever it changes.
406+ settings rather than in a versioned file, so its effective state is captured
407+ by the agent-configuration snapshot (section 13), and settings changes appear
408+ in the organization audit log.
407409- Secrets policy. Repository or organization secrets can be made available to
408410 the agent's environment during setup and execution. A secret carrying a
409411 credential (a personal access token, for example) would hand the agent write
@@ -508,6 +510,14 @@ points at the methodology, it never restates it (P-2).
508510- Labels: ` .github/settings.yml ` . An ` assessment ` label, per-phase labels, and
509511 the phase B skip marker (section 12), managed declaratively alongside the
510512 repository's existing label set.
513+ - Agent-configuration snapshot: ` scripts/assessment/ ` . The agent's effective
514+ configuration (MCP servers, firewall state, and custom allowlist) is readable
515+ from a documented endpoint, so a deterministic script captures it alongside
516+ each assessment's data outputs (HC-5): what the agent could reach when a draft
517+ was produced is committed evidence. The organization audit log remains the
518+ authoritative history of who changed a setting and when. The endpoint does not
519+ expose organization-level allowlist entries; the administrator/platform owner
520+ records those by hand when they exist.
511521- Deliverables: ` analyses/<year>/<project>/ ` . The existing convention:
512522 ` analysis.md ` , ` implementation.md ` , and the backlog.
513523- Backlog files: ` analyses/<year>/<project>/issues/ ` . One file per proposed
@@ -570,10 +580,13 @@ The verification profile:
570580
571581- Verifier (every phase). Input: the draft PR's deliverable and the sources it
572582 cites. Output: a report on the PR listing each checked claim as supported,
573- unsupported, or unverifiable, with the failing ones quoted. Its prompt is
574- adversarial (find unsupported claims), not confirmatory (check the draft is
575- fine), so it does not share the drafter's failure mode (G-2). Its tool list is
576- read-only; it edits nothing.
583+ unsupported, or unverifiable, with the failing ones quoted. The report is
584+ delivered whole: a finding tucked behind an all-clear summary line is G-2's
585+ fluent-but-wrong failure mode reappearing in review clothing, and the reviewer
586+ reads the list, not the headline. Its prompt is adversarial (find unsupported
587+ claims), not confirmatory (check the draft is fine), making it less likely to
588+ share the drafter's failure mode (G-2). Its tool list is read-only; it edits
589+ nothing.
577590
578591Model and effort. Profiles state capability requirements, never model names:
579592models change faster than this spec. Drafting warrants the most capable model
@@ -707,12 +720,19 @@ chosen, deliberately, per the caveat in section 10.
707720 agent's branch naming; whether review-comment revisions require write access
708721 (section 12); whether Actions runs on agent PRs wait for approval; how a
709722 verifier run is invoked against an existing PR (section 12); whether
710- delegation reliably opens a draft pull request (section 11); and whether an
711- agent profile can pin a model and effort level, or model choice rides entirely
712- on the per-task picker (section 14).
723+ delegation reliably opens a draft pull request (section 11); whether assigning
724+ an issue to Copilot offers the choice of agent profile, or selection needs the
725+ Agents panel (section 12); whether a full phase A draft fits the session cap
726+ or the drafting needs decomposing (section 11); whether an agent profile can
727+ pin a model and effort level, or model choice rides entirely on the per-task
728+ picker (section 14); and whether the cloud-agent configuration read endpoint,
729+ in public preview as of this writing, returns the fields the snapshot needs
730+ (section 13).
713731- Filing issues into project repos. A separate, opt-in tool to create the
714732 backlog issues in a project's own repository (NG-2). Out of scope for phase
715- one, and it must preserve HC-1.
733+ one. It would be human-run with its own credential and the project's explicit
734+ opt-in, entirely outside the agent's trust boundary; HC-1 continues to bind
735+ the agent unchanged.
716736- Intake relationship. How the request template fits with the existing CNCF
717737 service desk and assistance-program intake (section 7), without a competing
718738 front door. Related: the contribute.cncf.io site needs a page for projects on
0 commit comments