English | 中文
Use the narrowest check that can falsify the change, then widen only when the changed contract crosses more layers. Local checks are fast proof of syntax, schema, generation, and transformation behavior. GPU execution is proof of allocation, serving, workload, and artifact behavior. A smoke run reduces feedback time, but it is not the full-sweep-and-eval evidence required for merge review.
- Sources of truth
- Testing layers
- Local checks
- Smoke, sweep, and eval
- Evidence standard
- Review gates
- Stop conditions
.github/AGENT_OPERATIONS.mddefines sweep labels and modifiers. Its dispatch section defines manual runs and artifact inspection.docs/configuration-procedures.mdis the focused configuration validation procedure..github/workflows/README.mddocuments matrix generation,e2e-tests.yml, PR sweeps, and reuse.run-sweep.ymlis the executable PR/push gate.e2e-tests.ymlis the manually dispatched end-to-end path.docs/PR_REVIEW_CHECKLIST.mdis the merge-review standard. The verifier prompt states how sweep and eval evidence is independently checked.
These sources outrank this guide when behavior changes. Update the English page first, then translate the same structure and evidence into this page's Chinese counterpart.
| Layer | What it can prove | What it cannot prove |
|---|---|---|
| Parse and syntax | Edited YAML loads, and edited Bash parses | Schema validity, runtime routing, or GPU behavior |
| Schema and matrix | A config key validates and emits the intended matrix fields | Runner availability, server startup, or performance |
| Focused Python tests | Changed generator, changelog, result, eval, collection, or reuse contracts behave on covered inputs | Container, accelerator, network, or Slurm behavior |
| Smoke run | One tightly filtered path allocates, starts a server, runs a workload, and emits artifacts | The complete concurrency/search space or merge eligibility |
| Trimmed PR sweep | Each selected single-node group runs its lowest concurrency (sweep-enabled) |
Intermediate concurrency points required by a full sweep |
| Full sweep and eval | The selected untrimmed matrix and eval jobs execute on the reviewed commit | Correctness of evidence that was not inspected, or unrelated configurations |
A green later layer does not erase missing earlier evidence. For example, a green collector can aggregate an empty set, so review must inspect the underlying executed jobs and artifacts.
Run checks from the repository root and replace placeholders with the exact changed path or key.
python3 -c "import yaml; yaml.safe_load(open('configs/<nvidia|amd>-master.yaml')); yaml.safe_load(open('configs/runners.yaml')); yaml.safe_load(open('perf-changelog.yaml'))"
bash -n benchmarks/<path>/<script>.sh
bash -n runners/launch_<cluster>.shParsing is only the first gate. Do not report a YAML parse as matrix validation.
uv run --no-project --with pydantic --with pyyaml --python 3.12 \
utils/matrix_logic/generate_sweep_configs.py test-config \
--config-files configs/<nvidia|amd>-master.yaml \
--runner-config configs/runners.yaml \
--config-keys <exact-key>
uv run --no-project --with pydantic --with pyyaml --python 3.12 \
utils/matrix_logic/generate_sweep_configs.py full-sweep \
--config-files configs/<nvidia|amd>-master.yaml \
--runner-config configs/runners.yaml \
--model-prefix <prefix> \
--framework <framework> \
--precision <precision> \
--runner-type <runner> \
--seq-lens 1k1k 8k1kInspect the emitted values, not only the exit code or row count: config key, model, image, runner, scenario, concurrency, max-model-len, TP/PP/EP/DCP/PCP, prefill/decode workers, hardware, router, KV transfer, eval flags, additional-settings, and spec-decoding. The schema lives in validation.py, and the generator is generate_sweep_configs.py.
| Change | Focused command |
|---|---|
| Matrix schema or generation | python -m pytest utils/matrix_logic/ -v |
| Changelog content or PR gating | python -m pytest utils/test_process_changelog.py utils/changelog_gate_tests/ -v |
| Result processing | python -m pytest utils/test_process_result.py utils/test_aggregate_power.py utils/test_calc_success_rate.py -v |
| Eval dispatch, batching, or patches | python -m pytest utils/evals/ -v |
| Eval collection | python -m pytest utils/test_collect_eval_results.py -v |
| Sweep reuse or reusable artifacts | python -m pytest utils/test_find_reusable_sweep_run.py utils/test_validate_reusable_sweep_artifacts.py -v |
For an edited changelog, also run the same matrix-compatibility validator used by setup, with real base and head refs:
python3 utils/validate_perf_changelog.py \
--changelog-file perf-changelog.yaml \
--base-ref <base-ref> \
--head-ref <head-ref>Its contract is implemented in validate_perf_changelog.py. This check validates the generated matrix and rejects prohibited content changes, but whitespace-only historical deletions can be invisible to its diff reader. Inspect the exact byte diff as a separate evidence gate. Do not rewrite or normalize historical perf-changelog.yaml bytes.
A local matrix cannot prove Slurm allocation or llm-d endpoint discovery. Multi-node recipe changes still require the upstream recipe checker and an execution on the intended fleet, as described in configuration validation.
A smoke run is a manually dispatched e2e-tests.yml run whose generator command is restricted to the changed model/framework/runner/scenario and a minimal supported concurrency. Generate that exact command locally first. Use smoke evidence to answer a narrow question such as “does this image start this server on this runner and produce a result artifact?”
A smoke run is not merge evidence: it intentionally omits configurations and concurrency points. Likewise, agentx-fast is an AgentX preflight, not a canonical AgentX result. The dispatch reference defines the reduced warmup/profile behavior.
sweep-enabledtrims each parallelism group to its lowest concurrency and is the default for most PR feedback.full-sweep-fail-fastis the recommended full-sweep label. It uses the sequential single-node canary and stops each matrix after that matrix's first failure while preserving completed results.- Use a no-canary full-sweep label only when the canary is known to be flaky or unrepresentative. Use
full-sweep-enabledinstead of fail-fast only when every matrix job must continue despite a failure. - Apply exactly one primary sweep label. Modifier-only or conflicting primary labels do not constitute a valid sweep.
The current meanings and eligibility rules are defined in the sweep-label reference and implemented by run-sweep.yml.
Throughput and evals are separate jobs. The default sweep evaluates the selected 8k1k subset. all-evals expands eval selection, and evals-only suppresses throughput. Choose modifiers from the changed scope, but do not substitute an eval-only or preflight run for the required full sweep.
Eval completion is not just a green job. Preserve and inspect meta_env.json, the results*.json files, the score-validation output, the inference image, and the aggregated eval artifact. utils/evals/EVALS.md owns task and artifact behavior. validate_scores.py rejects missing result files, below-threshold scores, and runs with no checked metrics. When expected concurrency metadata is available, it also rejects invalid/incomplete/failed batches. The workflow invokes it without --expected-concs, so reviewers must verify meta_env.json independently for single-concurrency artifacts.
Record enough information for another reviewer to reproduce the claim without guessing:
- Exact commit SHA and whether it is still in the PR.
- Exact local command or workflow generator command, including every filter and modifier.
- Workflow URL, run ID, attempt, job/check name, and conclusion. Preserve failed logs before rerunning.
- Config key plus resolved model, image, runner, framework, precision, scenario, topology, and concurrency range.
- Artifact names and the relevant structured fields or summarized metrics. Do not paste an unbounded raw aggregate.
- For evals, task, expected threshold, observed score, completion metadata, and proof that the evaluated image matches the PR config.
- For a failure, record the first failing layer and the evidence that rules out earlier layers. Use the classification in
troubleshooting.md.
“CI is green,” a screenshot without a run URL, a collector success without executed jobs, or an artifact from a rebased-out commit is not sufficient evidence.
- Before GPU work: parsing, exact-key generation, relevant focused suites, and inspection of emitted fields are green. The exact dispatch generator command has been run locally.
- Before broadening: a smoke or canary proves the changed runtime path. If it fails, diagnose that layer before spending a full sweep.
- Before CODEOWNER sign-off: follow
PR_REVIEW_CHECKLIST.md, including its code-quality, architecture, image provenance, upstream recipe, patch/waiver, chat-template, and AgentX requirements where applicable. - For sweep/eval acceptance: at least one commit currently in the PR has successful, non-skipped executed
single-node */andeval /checks. A successfulcollect-evalsalone is insufficient. Download the corresponding eval artifacts and confirm non-empty, passing accuracy and the same inference image. These are the executable rules in verifier Checks 1 and 2. - For reuse at merge: an authorized
OWNER,MEMBER, orCOLLABORATORposts a whole-line/reuse-sweep-runcommand (optionally with the eligible source run ID) before the supported merge path. The verifier treats a missing or unauthorized command as a failure. See verifier Check 4 and the reuse procedure. - At merge: a CODEOWNER's exact sign-off is independently accepted by
codeowner-signoff-verify.yml. If the PR head changes, reassess and sign the new commit evidence. - After merge: the author confirms the main-branch jobs pass, as required by
CONTRIBUTING.md.
Stop and fix or escalate instead of widening the run or approving when:
- syntax, schema, exact-key generation, changelog validation, or a focused suite fails.
- generated fields differ from the intended config even if the command exits zero.
- the runner/fleet prerequisite cannot be exercised for a multi-node change.
- a smoke or canary fails and the failure layer is still unknown.
- expected jobs are skipped, cancelled, absent, or attached only to a commit no longer in the PR.
- eval artifacts are missing, empty, incomplete, below threshold, or use a different image.
- required evidence, an upstream recipe, a waiver, or a reviewer-owned checklist fact is unknown.
Do not convert an unknown into a pass, rerun until a flake disappears without preserving the original failure, or spend a broader sweep to hide a narrower failure.