Replace the bash complete run driver with tests/complete_run - #876
chengzhuzhang wants to merge 12 commits into
Conversation
Replace tests/main_branch_testing/run_integration_test.bash with a tested Python package, tests/complete_run/, driven by `python -m tests.complete_run.automation`. - Each repository is used through a detached worktree of one resolved commit; a developer's clone is never committed to or switched. - Every stage writes status.json, so a run can be resumed with --start-stage, and always ends with JSON and Markdown reports. - Runs live in a shared, group-writable root; a baseline is a run that `python -m tests.complete_run.promote` points baselines/latest-main at. - Intermediate data (worktrees, zppy's post-processing output) goes to per-user scratch. - Everything that exercises zppy runs in the environment built from the commit under test. - The weekly cfgs, their cases and image-checked tasks are defined once in tests/integration/weekly_cfgs.py. Four legacy cfgs that had become copies of the current ones are removed, and livvkit and legacy 3.1.0 pcmdi_diags plots are now image-checked. - Each run gets an image review summary page across every package. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The image checker's batch script ran `source ~/.bashrc` and then a
per-machine alias (`lcrc_conda`, `compy_conda`, `nersc_conda`). Those aliases
exist only in some people's shell setup, so the job failed immediately for
anyone else ("lcrc_conda: command not found").
Source the conda profile the run is given with --conda-profile, as the
generated task scripts already do, and drop the alias from MachineProfile.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The settings and bash baselines were copied from the output the integration
tests leave in the worktree, but each test deletes that output when it
passes, so a passing run captured none of them ("missing 8 settings
baselines").
Regenerate each one from its dry-run cfg with the zppy under test, pruned
exactly as its test prunes it. On Chrysalis all eight match the current
expected files under the tests' own diff rules.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A run's cfgs source load_latest_e3sm_unified_<machine>.sh, and the manifest listed only checked-out repositories, so nothing machine-readable said which Unified release a baseline was built with, or that any package came from it. Read the version, environment path and key package versions from the package list the run already collects, and record them in status.json and manifest.json. The manifest now lists repositories run from Unified too. `promote show` and the report show the release. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Add --solver (default mamba, falling back to conda when mamba is not installed). It is used only to create environments; activating, running, listing and exporting still use conda, whose output the run parses. On Chrysalis mamba solves zppy's dev.yml in 75 s, where conda 23.3's classic solver had not finished after 15 minutes. MPAS-Analysis's dev-spec.txt cannot name channels, so it was solved against whatever the operator's condarc lists; with only `defaults` it never finishes. Pass conda-forge with strict priority explicitly, as MPAS-Analysis's README directs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
zppy skips a task whose status file begins with OK, WAITING or RUNNING. A cancelled or crashed run leaves WAITING and RUNNING files behind pointing at jobs that no longer exist, so resuming into the same output directory skipped exactly the tasks that never ran. Validation then swept a tree whose every status file read OK and the run could report a pass over work that never happened. Before submitting, drop the status files whose job SLURM no longer has queued. OK and ERROR are left alone: OK means the work is done, and zppy already resubmits ERROR. A job still in the queue keeps its file, since deleting a live one would submit the same task twice. Without the queue there is no way to tell a stale file from a genuinely pending one, so a queue that cannot be read clears nothing. That also keeps the unit tests working where SLURM is absent, since run_command raises on a missing binary even when check is False. Found while resuming a run that died when the filesystem filled: 11 of its global_time_series tasks would have been silently skipped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Running tests/images in CI exposed that test_compare cannot pass anywhere mache is unable to discover a machine: it built the diff directory from the machine's web_portal base_path and $USER, so a GitHub runner failed with "Unable to discover machine from host name". Neither value is under test. The diff images are an artifact of the comparison, so write them to a temporary directory instead. That also stops a unit test leaving diffs behind in the shared web portal. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The image checker names each image relative to the cfg's diff directory, so
names start with the task ("e3sm_diags/..."), but each review page is written
inside that task's directory. Linking the bare name doubled the task
directory, and no image on any review page loaded. write_viewer now links
images relative to the parent directory.
The summary's only links were a trailing "Review" column in a wide table that
scrolls horizontally, so they started out of view. Package names now jump to
their rows among the checks, and the link to each review page is on the task
name.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A baseline whose repositories ran from E3SM-Unified exports an environment only for zppy, so every other repository was skipped as "not compared" and the heading and note, based on zppy alone, said no dependency changed and that image differences were attributable to the code under test. Comparing a dev-solve run against a Unified baseline, that is the opposite of the truth. When the baseline has no export for a repository, compare against the `conda list` its env_descriptions/<task>.txt carries, by version only, with package names matched across pip and conda (a dev environment pip-installs the package under test that Unified has from conda-forge). And never call a run clean while a repository went uncompared, in the note, the viewer heading, or the report. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
E3SM-Unified bundles every tool, while each repository's dev environment holds only what that repository needs, so about 70% of a repository's differences were Unified packages its environment never had, listed as removed. They buried the dependencies that actually differ. Against a Unified baseline, list only packages present in this run's environment, and state how many Unified-only packages were left out. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The CSS escape for the bullet sat in a plain Python string, where "\202" is an octal escape, so every notable package in the environment table was followed by an invisible control character and a literal "2". Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A run's expected image lists are built by walking its own www tree, so a run made with a reduced --cfg or --task selection becomes a baseline holding only those tasks, and every other task then records no comparison against it. The advice narrowed the baseline instead of updating part of it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
@forsyth2 — this is ready for a look. It replaces |
|
Thanks @chengzhuzhang. I won't have time to actually learn this new approach and try it out until after I get the PCMDI PRs (#815, E3SM-Project/zppy-interfaces#49) merged. In the meantime, I've asked Claude to compare the two approaches. I did this by giving it:
I'm pasting its conclusions below. ContextBoth PRs are evolutions of the same weekly zppy integration-test workflow — which started as a fully manual, documented-in-RST checklist, and had already been partially automated into a single Approach A — Extension (bolt automation onto the existing bash script)What it does: keeps
Strengths
Weaknesses
Approach B — Redesign (replace the bash driver with a
|
|
@forsyth2 Thanks for sharing the code analysis. Yeah, no hurry on this, but it should be really straightforward to try out with the instructions in the PR, especially with an agent. |
|
@chengzhuzhang I had a moment to ask Claude a couple more questions. I pasted its comments on Approach A (extension) vs Approach B (redesign) below. A couple thoughts from me:
Functionality found in A but not BWhat A has that B doesn't: A auto-generates the "Step 2: what changed since the baseline" table — it detects, per task, the date the expected results were actually produced (reading B's docs for the equivalent step say, verbatim: "Review each commit log and note commits made since that date... This step is human judgement; the report records what was tested, but deciding which changes matter is yours." B gives you the ingredients — the baseline's commit SHA per repo ( Splitting up this PRA +10k/-4k diff isn't reviewable the way a normal PR is. Nobody can hold that
Bottom-up: complete layers, wire them together at the endEach PR is a fully-built, fully-tested module (or small group of them), in
PRs 1–4 are individually small, low-risk, and fully covered by unit tests PR 5 is also the one that has to delete
Stage-by-stage delegation: the bash script hands off one piece at a timeInstead of building the new package and switching over all at once, the old
By the end of PR 4, the bash script isn't doing any of the actual work anymore Every PR here changes the real weekly test immediately — there's no "built
Bottom line: bottom-up gives you smaller, easier-to-isolate PRs, but the Independent of either path:
|
|
@forsyth2 Thanks for the feedback. This refactor (also borrows feature from e3sm_diags complete run) was done under time pressure, knowing your time share on E3SM will drop, and what I want out of it is a more automated test framework that can be handed off more easily. Good catch on #852 -- it's actually a conflict, not just an overlap. I don't think it's necessary to break the refactor itself into smaller PRs. The first commit is the bulk of it, but the eleven after it are already small and self-contained, and most of the diff is unit tests and deleted legacy cfgs. I agree it should get more testing though. I'll run it on Perlmutter (and setting up team paths for results) as well as Chrysalis. |
This PR is 10,000 lines long.
Having Claude break down the +10,096 added lines by category yields:
Tests are a genuinely large share (37%), and the legacy cfg cleanup is real |
|
Based on the pros and cons discussed here: #876 (comment) I think the new architecture is the better long-term direction. The main concern is how we migrate to it safely, rather than whether we should move in that direction. I think this can be addressed through more testing in practice, for example, by testing with the upcoming E3SM-Unified environment. |
Yes, exactly. The bash script is the result of many incremental changes but a Python-based modular refactor would make things much more maintainable.
Now that Shixuan and I have merged #815 (which unfortunately seems to have introduced merge conflicts here) and its companion PR E3SM-Project/zppy-interfaces#49, I am going to start running tests using both #871 and this PR. It's a bit of duplicated testing, but it's really important to see how they match up before switching over completely. I will try out a complete run using the working NCO (5.3.9) as well as Charlie's fix for #875. I'm using Claude to help setup this test framework to match however I setup #871. A few things Claude has identified as potential differences so far:
I want to actually try it out of course. Also a couple things that don't appear to be addressed here or in #871:
One of the difficulties of testing tests is that there are many ways for code to fail a test. To be fully automated, the test script needs to handle all those things gracefully (e.g., for "don't wait around for jobs that can't ever run" -- you'd never encounter this when testing code that works) and continue as far as it can. So ironically you need failing code (ideally in multiple obscure ways) to improve the tests. Another factor is configurability -- e.g., it needs to work for a cron-based weekly testing of |
Test approach comparison 1: NCO error in
|
@forsyth2 Since they are specifically helping us evaluate the new approach. I think we can post these comparison here. Thank you for taking the time to identify the gaps of this workflow and to run these comparisons! They’ll be really helpful for validating this PR. |
|
Another suggestion for PR 876: enable automatic pinning of dependency versions One issue I'm running into trying to do a main branch test is that because NCO is currently broken (#875), I need to pin an older NCO for these packages myself. This issue applies to either approach, incremental improvement (#871) or refactor (#876). For reference, To address this issue, Claude suggests something like: python -m tests.complete_run.automation \
--machine chrysalis --repo-root ... \
... \
--stop-after prepare \
--run-number 4
# Check status.json's "repos" section for the exact env names it just built, then:
conda install -n test-e3sm_diags-main-20260926_run4 "nco<5.4.0" --yes
conda install -n test-e3sm_to_cmip-main-20260926_run4 "nco<5.4.0" --yes
# ...repeat for each dev-type repo
python -m tests.complete_run.automation \
--machine chrysalis --repo-root ... \
... \
--start-stage generate --tag 20260926_run4which I suppose mimics what could be done for the bash script of #871 as well. It's not very elegant though, and it requires waiting for environments to build, and then doing a manual step before instructing the script to continue on. What would be particularly useful is to be able to specify in the config any dependencies that should be pinned and have the test do this post-creation (Note: I realize I'm suggesting even more features despite expressing concern that the PR is too large already. I think we should at least identify all the issues and possible improvements. Then comes the question of how best to review and integrate the code into |
Test approach comparison 2:
|
| Test name | Total images | Correct images | Identical | Cosmetic only | Missing images | Needs review | Severity |
|---|---|---|---|---|---|---|---|
| bundles_e3sm_diags | 1762 | 1703 | 3 | 1700 | 0 | 59 | 22 moderate, 37 minor, 1700 negligible, 3 identical |
| comprehensive_v2_e3sm_diags | 3806 | 3725 | 3 | 3722 | 1 | 80 | 1 missing, 1 major, 49 moderate, 30 minor, 3722 negligible, 3 identical |
| comprehensive_v3_e3sm_diags | 5369 | 5290 | 3 | 5287 | 0 | 79 | 37 moderate, 42 minor, 5287 negligible, 3 identical |
| legacy_3.1.0_comprehensive_v3_e3sm_diags | 5365 | 5289 | 3 | 5286 | 0 | 76 | 37 moderate, 39 minor, 5286 negligible, 3 identical |
| legacy_3.0.0_comprehensive_v3_e3sm_diags | 5365 | 5289 | 3 | 5286 | 0 | 76 | 37 moderate, 39 minor, 5286 negligible, 3 identical |
mpas_analysis:
| Test name | Total images | Correct images | Identical | Cosmetic only | Missing images | Needs review | Severity |
|---|---|---|---|---|---|---|---|
| comprehensive_v2_mpas_analysis | 856 | 613 | 4 | 609 | 0 | 243 | 63 moderate, 180 minor, 609 negligible, 4 identical |
| comprehensive_v3_mpas_analysis | 1280 | 903 | 6 | 897 | 0 | 377 | 71 moderate, 306 minor, 897 negligible, 6 identical |
| legacy_3.1.0_comprehensive_v3_mpas_analysis | 856 | 580 | 4 | 576 | 0 | 276 | 63 moderate, 213 minor, 576 negligible, 4 identical |
| legacy_3.0.0_comprehensive_v3_mpas_analysis | 856 | 580 | 4 | 576 | 0 | 276 | 63 moderate, 213 minor, 576 negligible, 4 identical |
global_time_series:
| Test name | Total images | Correct images | Identical | Cosmetic only | Missing images | Needs review | Severity |
|---|---|---|---|---|---|---|---|
| bundles_global_time_series | 3 | 3 | 0 | 3 | 0 | 0 | 3 negligible |
| comprehensive_v2_global_time_series | 12 | 12 | 0 | 12 | 0 | 0 | 12 negligible |
| comprehensive_v3_global_time_series | 1404 | 1398 | 0 | 1398 | 0 | 6 | 6 minor, 1398 negligible |
| legacy_3.1.0_comprehensive_v3_global_time_series | 1404 | 1398 | 0 | 1398 | 0 | 6 | 6 minor, 1398 negligible |
| legacy_3.0.0_comprehensive_v3_global_time_series | 90 | 90 | 0 | 90 | 0 | 0 | 90 negligible |
pcmdi_diags:
| Test name | Total images | Correct images | Identical | Cosmetic only | Missing images | Needs review | Severity |
|---|---|---|---|---|---|---|---|
| comprehensive_v3_pcmdi_diags | 647 | 312 | 1 | 311 | 156 | 179 | 156 missing, 4 structural, 36 major, 15 moderate, 124 minor, 311 negligible, 1 identical |
| legacy_3.1.0_comprehensive_v3_pcmdi_diags | 617 | 312 | 1 | 311 | 126 | 179 | 126 missing, 4 structural, 36 major, 15 moderate, 124 minor, 311 negligible, 1 identical |
These results look completely different from the 9/28 test using the incremental improvements. The number of diffs don't match up for e3sm_diags or pcmdi_diags (asides from "Missing images" in the latter), and the 9/28 test didn't show any non-negligible errors for mpas_analysis or global_time_series. This is likely a sign that the expected results baseline is out of sync between the two approaches.
Indeed, looking at the results from python -m tests.complete_run.promote --machine chrysalis show, we can see that the expected results baseline for the refactored test used Unified for all called packages with only zppy itself using a dev environment. Compare that to the expected results baseline for the 9/28 test, where we found the following:
| Test run | Date its results were promoted | Packages whose baseline got updated |
|---|---|---|
| 9/4 | 9/14 | pcmdi_diags |
| 8/28 | 9/4 | e3sm_diags, mpas_analysis |
| 8/12 | 8/14 | global_time_series |
| E3SM Unified 1.13.0 | 5/20 | ilamb, livvkit |
Therefore, this test was not actually apples-to-apples.
Nevertheless, let's continue our analysis of the refactored test's output by returning to the run report. Issues of note here:
- It looks like there might be a unicode issue with some version numbers (alot of
—)
Returning one step back to the results home page. Issues of note here:
- "Compared against baseline 20260918_unified113" implies there can't be varied baselines, which should be addressed per the issues discussed in Enable partial updates of expected results for testing and the testing strategy discussion. We can't be waiting around for every task to pass before updating expected results. The longer ago expected results were generated, the more diffs come up and it becomes harder to know what's a new error.
- It's unintuitive to me that clicking the package link in the "By package" table just brings you to the very row you just clicked. I believe the point is that you can share the link with others, but it makes the link look broken.
- The "By package" section gives us no indicator of the environment used for tasks that don't produce a plot (e.g.,
e3sm_to_cmip), or forzppyitself.
Let's look at some of the "By check" links. For example e3sm_diags weekly_comprehensive_v3. This is pretty nice -- it expand the image checker grid so it's easier to view at a glance. One thing I dislike is that the "Environment" box is open by default -- there are so many dependencies listed it just looks like noise before getting to what we really care about, the image diff grid.
In summary, suggestions for PR 876:
- Check for Unicode issues in the environment listings
- Allow for partial updates of the expected results (maybe even allow the user to set which expected results to use as baseline?)
- Indicate environments used for all tasks, not just the ones generating plots.
- Have the "Environment" box closed by default.
I've also asked Claude for further analysis (copied below):
Claude's further analysis
I went through the linked pages under 20260929_run5 (report, image summary, and index.html) in detail, cross-checked the numbers, and spot-checked a couple of the live conda environments against what the report claims. Summary below, split into things that need work and things that are clear wins over the PR 871 approach.
Needs work
- The Environment diff table can misreport what actually ran. The report says
zppy_interfacesrannco 5.4.0;conda list -n test-zppy_interfaces-main-20260929_run5 ncoshows it's actually5.3.9.e3sm_to_cmipwasn't even included in the comparison (Not compared for: e3sm_to_cmip), but its live env is also correctly5.3.9. So the "This run" column looks like a snapshot taken early (likely right afterconda env create, before the manualnco<5.4.0pins were applied) that never gets refreshed before the final report is written. This needs a fix at the source — capture the environment manifest at or after task execution, not once during setup — ande3sm_to_cmipshould be included in the comparison rather than excluded, since it currently hides exactly this class of bug. (The onencorow that may genuinely be 5.4.0 iszppyitself, which doesn't call NCO, so that particular mismatch has no functional consequence — but it's a coincidence, not evidence the underlying bug is benign elsewhere.) - No support yet for staggered per-package baselines.
promote --machine chrysalis showand the report (Compared against baseline 20260918_unified113) both show one global baseline used for every package except zppy's own dev env. PR 871's config exposesDIAGS_EXPECTED_RESULTS_DATE,MPAS_EXPECTED_RESULTS_DATE, etc. for exactly this reason. Until this is ported, none of the diff counts between the two approaches are apples-to-apples. zppyande3sm_to_cmipare invisible in the "By package" results table. Onlye3sm_diags,ILAMB,LIVVkit,MPAS-Analysis, andzppy-interfacesget a row. There's no result/status indicator at all for the two tasks that don't produce plots.- The "changes since expected results were updated" section from PR 871's report didn't make it into the rewrite. That per-package commit list since the baseline was generated was genuinely useful for triage and should either be ported or explicitly deprioritized.
- The
Environment<details>box on the per-check pages defaults to open (<details class='env' open>), burying the image diff grid under a long dependency list. Should default to closed. - "Needs review" means two different things in two files that sit side by side.
test_images_summary_*.md's per-check "Needs review" column excludes missing images (tracked separately);index.html's "By package" table's "Images needing review" includes them. The underlying numbers are internally consistent, but this will read as a discrepancy to anyone cross-referencing the two files.
Clear improvements over the PR 871 / current method
- The image checker now runs as an automated SLURM job end-to-end — no more manual compute-node step to kick off
pytest tests/integration/test_images.py. - Job-status and integration-test tracking (
test_last_year.py,test_bash_generation.py,test_campaign.py,test_defaults.py,test_bundles.py) is retained and reported cleanly, at parity with PR 871. - The per-check pages (e.g. the
e3sm_diags weekly_comprehensive_v3view) give a much nicer expandable image-diff grid for at-a-glance review than anything in the old report. - The "By package"/"By check" tables give shareable, anchor-linkable views into a single run — a real step up in navigability over a flat markdown report.
- Overall runtime is roughly on par with the current script (~3h in both cases), so the rewrite isn't paying a performance tax for the added structure.
Net: the reporting and automation UX is a real improvement, but before we can use this to judge the refactor against PR 871, we need the environment-snapshot bug fixed (so we can trust any package's reported version) and the staggered-baseline feature ported (so the diff comparison is actually apples-to-apples). Until then, the diff numbers in this run shouldn't be used as evidence either way.
|
A couple more suggestions that come to mind:
|
I’d vote against adding /$USER. It seems unlikely that two testers would run the same test at the same time, and we’d also like to maintain a single shared baseline rather than have baselines scattered across $USER directories. |
@chengzhuzhang These are pretty similar names. All it would take for a name collision is two developers testing different features on the same day. There's no timestamp or any sort of further layer of differentiation.
To clarify, the |
Maybe we could use a timestamp to the test name to provide another layer of differentiation |
| screen -ls # See what screen sessions you have | ||
| tail -f integration_test_runN.log | ||
| Packages that most often move results (``numpy``, ``matplotlib``, ``xarray``, | ||
| ``nco``, ``esmf``, the analysis packages themselves) are listed first and |
There was a problem hiding this comment.
PMP should also be part of "Packages that most often move results"
Replaces
tests/main_branch_testing/run_integration_test.bash(844 lines of bash) withtests/complete_run/, a Python package of 14 modules.How to use
One command runs everything:
Reading the review page
Example, from a run of
mainagainst the20260918_unified113baseline:https://web.lcrc.anl.gov/public/e3sm/zppy_complete_run/runs/20260918_main_run1/
Severity runs IDENTICAL → NEGLIGIBLE (cosmetic) → MINOR → MODERATE → MAJOR → STRUCTURAL → MISSING. Only MINOR and worse need review.
Useful flags:
--cfg,--task--start-stage,--stop-afterprepare → generate → submit → bundles → validate → report--env-type <repo>=dev|unified|baselinedev.yml, use E3SM-Unified, or rebuild what the baseline usedWhen it finishes, open the run's review page.
Promoting a baseline
Expected results are not a copy of a run:
baselines/latest-mainis a symlink pointing at one. Promotion flips that link, so it is atomic, instant, and leaves the previous baseline intact and still promotable.A run never promotes itself; this is always a separate, deliberate command. It refuses a run that has no report, one whose status is not
passed(--allow-failedonce you have decided the new plots are correct), and one built from a feature branch (--allow-non-main).prunenever deletes a run a channel points at.Known gap: partial promotion
A baseline is one whole run, so a task cannot be blessed on its own: if
e3sm_diagschanged legitimately while another package has an unexplained difference in the same run, both are blessed together or neither is.Note that
docs/source/dev_guide/tests/update_expected_results.rstcurrently suggests rerunning the complete test with just the cfg and task you need, then promoting that run. That does not work: a run's expected-image lists are built by walking its ownwww/tree, so a task-limited run becomes a baseline containing only that task, and every other task then records no comparison. The doc fix belongs with whichever design below is chosen, so it is not in this PR.Two designs were considered, neither implemented here:
promoteassembles a new run directory from the current baseline with the blessed task's plot directories symlinked from the newer run, regenerates the image lists from the merged tree, and records per-task provenance in the manifest. A baseline stays one run directory, so nothing downstream changes;env_description.txtis already per task, so each task keeps its own commit andGenerated:date.Major changes vs the current complete test
One resumable command instead of three manual phases. The old driver had
START_PHASE=1|2|3, set by hand in a config file. This runs six stages —prepare → generate → submit → bundles → validate → report— selected with--start-stage/--stop-after. Every stage writesstatus.json, so an interrupted run resumes where it stopped and a crashed one still produces a report.The image check is no longer a manual step. The old README said "test_images.py must be run manually from a compute node." It is now submitted as a SLURM job and its outcome folds into the run's verdict.
Immutable worktrees instead of switching branches. The old script had to be copied out of the repo before use, because it changed branches in the clone. Each repo under test now gets a git worktree, so the clone is never touched and the exact commit tested is recorded.
A report, not just pass/fail. Produces
complete-run-report.md/.jsonplus the HTML review page above, which ranks image differences by severity, so review starts with what actually changed.Provenance and dependency attribution. Records the E3SM-Unified release and every repo's branch, sha and environment, and diffs the run's environments against the baseline's. That distinguishes image differences caused by the code under test from ones caused by a dependency bump. A baseline that ran from E3SM-Unified exports no environment of its own, so the comparison falls back to the
conda listitsenv_descriptions/<task>.txtcarries, and lists only packages the run's own environment has — not the hundreds E3SM-Unified bundles for other tools.Promotion stays explicit. A complete run never promotes its own results;
promote.pyis a separate, deliberate step.The driver itself is tested. 260 unit tests across 13 files (3671 lines). The bash driver had none.
Scheduled runs. A controller script and crontab template for running this unattended.
Also drops the
update_*_expected_files_*.shscripts and the legacy 3.0.0/3.1.0 bundles and comprehensive_v2 cfgs, and updates the testing docs to match.Status
Draft. Exercised end to end on Chrysalis against the
20260918_unified113baseline: all stages ran, 186/186 job status files OK, all five integration tests passed, and the image checker compared 35,912 images across five cfgs.That run also surfaced a resume bug now fixed here — zppy skips tasks whose status file reads
WAITING/RUNNING, so a cancelled run's leftovers made a resume silently skip the tasks that never ran. The harness now drops status files whose SLURM job no longer exists.Reviewing that run's own pages then turned up three viewer bugs, each fixed with a regression test: images were linked one directory too deep and so never loaded; the only links to the per-task pages sat in a column a wide table scrolled out of view; and the environment panel claimed "no dependency changed" when four of five repos had in fact not been compared at all.
🤖 Generated with Claude Code