Skip to content

Restore GitHub Actions CI - #929

Merged
aaronspring merged 8 commits into
mainfrom
claude/cicd-main-issues-89krmb
Sep 21, 2026
Merged

aaronspring merged 8 commits into
mainfrom
claude/cicd-main-issues-89krmb

Conversation

@aaronspring

@aaronspring aaronspring commented Sep 20, 2026 •

Copy link
Copy Markdown
Collaborator

Description

CI on this repository has not run at all since 2026-09-06, and the weekly run was partly inert for a long time before that. This PR is limited to getting it working again.

1. Both workflows were disabled. GitHub had set them to disabled_inactivity, which it does to scheduled workflows after 60 days without a commit (last commit to main was 2026-07-07). That state belongs to the workflow record, which GitHub keys to the file's path — so pushing commits does not clear it, and workflow_dispatch is refused outright. Registering the files at new paths creates fresh workflow records, which start active:

testing.yml         -> ci.yml
upstream-dev-ci.yml -> upstream-dev.yml

Job names, workflow name: values and triggers are unchanged, so branch-protection rules that key off check-run names keep matching. Confirmed working: Upstream Test run #1 under a new workflow id fired on the push that did the rename.

2. The weekly CI never actually ran the tests. detect-ci-trigger is gated on push/pull_request, so on schedule and workflow_dispatch it is skipped — taking the three test-matrix jobs that needs: it down with it. Run 144 (2026-09-06) shows all three as "conclusion": "skipped"; the weekly run had quietly degraded to just the doctest and notebook jobs. Now gated on always() plus an explicit triggered != 'true'.

3. Upstream Test died before running a single test, pinned to Python 3.10 while upstream xclim requires >=3.11:

ERROR: Package 'xclim' requires a different Python: 3.10.21 not in '>=3.11.0'

4. Upstream failure reporting was broken — it used ::set-output, which GitHub removed, so ARTIFACTS_AVAILABLE was never set and the report job never opened its failure issue. Now $GITHUB_OUTPUT. This is plausibly why the weekly failures went unnoticed for so long.

Also fixes the CI badge in README.rst and docs/source/index.rst (it pointed at climpred_testing.yml, a path removed in an earlier rename, so it rendered nothing), and drops a dangling always-true and "lead" clause in _rps that fails the ty pre-commit hook on main.

Type of change

  • Bug fix (non-breaking change which fixes an issue)

How Has This Been Tested?

  • Actions is demonstrably alive again: the rename produced a new workflow record and a run on this PR.

  • Fixes 3 and 4 are confirmed by that run's log — the environment built and pytest ran (instead of failing at install), and the log shows Set output 'artifacts_availability'.

  • pre-commit run --all-files passes, including the workflow-schema hook and ty.

  • Tests added for pytest, if necessary. — not applicable; this is CI configuration.

Checklist (while developing)

  • CHANGELOG is updated with reference to this PR.

Known red checks, both pre-existing on main

  • doctest — the HindcastEnsemble.remove_bias doctest for how="modified_quantile" pins stale values that no longer match current bias_correction output. Reproduces on main; deliberately left for its own PR so this one stays on a single subject.
  • upstream-dev — segfaults in netCDF4's C extension while opening a tutorial dataset in fixture setup (conftest.py:413 → load_dataset → xr.open_dataset). This is the signature of a numpy ABI mismatch: ci/install-upstream-wheels.sh pip-installs nightly numpy/pandas with --no-deps over a conda env whose netCDF4, bottleneck, numba and sklearn were compiled against the conda numpy. It needs its own fix in that script. Note this job is a nightly canary against bleeding-edge upstream mains, not a merge gate — it only runs here because of a [test-upstream] keyword used to validate the Python bump.

🤖 Generated with Claude Code

https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD

The weekly CI has been red and partially inert. Four separate problems:

- testing.yml: `detect-ci-trigger` only runs for push/pull_request, so on
  `schedule` and `workflow_dispatch` it is skipped, which also skipped the
  three dependent test-matrix jobs. The weekly run silently degraded to just
  the doctest and notebook jobs. Gate on `always()` plus an explicit
  `triggered != 'true'` so the matrix runs for every event.

- upstream-dev-ci.yml: pinned to Python 3.10, but upstream xclim now requires
  >=3.11, so `Set up conda environment` failed at install time before any test
  ran. Bumped to 3.13, which testing.yml already exercises.

- upstream-dev-ci.yml: `::set-output` was removed by GitHub, so
  ARTIFACTS_AVAILABLE was never set and the `report` job never opened or
  updated the upstream-failure issue. Use `$GITHUB_OUTPUT`.

- classes.py: the `HindcastEnsemble.remove_bias` doctest for
  `how="modified_quantile"` pinned stale values. Current `bias_correction`
  output reproduces identically under CI's conda/py3.10 env and a pip/py3.11
  env, so the expected values are updated rather than masked.

Also drop a dangling always-true `and "lead"` clause in `_rps` that failed the
`ty` pre-commit hook; the condition is unchanged.

Verified: `pre-commit run --all-files` passes, the `remove_bias` doctest
passes, and 175 rps/crps tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD
The entries added for these CI fixes had no `:pr:` reference, which the
repo's changelog convention and the PR checklist both call for.

The `[test-upstream]` keyword makes upstream-dev-ci.yml's detect-ci-trigger
fire on this PR, so the Python 3.13 bump is exercised here rather than only
after merge on the weekly schedule. It does not affect testing.yml, which
keys off `[skip-ci]`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD
GitHub had set both workflows to `disabled_inactivity` after the repository
went 60 days without a commit (last commit 2026-07-07, last run 2026-09-06).
That state belongs to the workflow record, which GitHub keys to the file's
path, so pushing commits does not clear it and `workflow_dispatch` is refused
outright. Renaming the files registers them as new workflow records, which
start active:

    testing.yml         -> ci.yml
    upstream-dev-ci.yml -> upstream-dev.yml

Job names, workflow `name:` values and triggers are all unchanged, so branch
protection rules that key off check-run names keep matching.

Also fixes the CI badge in README.rst and docs/source/index.rst, which still
pointed at `climpred_testing.yml` from an earlier rename and so rendered no
status at all. Both badges now track the current paths.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD
Seven entries for what amounts to "fix CI" was too granular. The badge fix,
the `::set-output` replacement and the `_rps` lint cleanup are invisible to
users and do not need their own lines; they are part of restoring CI or are
pure internals.

Left with two: one for CI being restored and repaired, one for the doctest
output that users actually see in the docs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD
…tream]

The upstream-dev job now gets far enough to run tests (the 3.13 bump fixed
the install-time failure) but segfaults in netCDF4's C extension while
opening a tutorial dataset in fixture setup:

    netCDF4._netCDF4
    xarray/backends/netCDF4_.py:563 in ds
    src/climpred/tutorial.py:223 in load_dataset
    src/climpred/conftest.py:413 in hindcast_NMME_Nino34
    Segmentation fault (core dumped)

That is the signature of a numpy ABI mismatch: install-upstream-wheels.sh
pip-installs nightly numpy/pandas with --no-deps over a conda env whose
netCDF4, bottleneck, numba and sklearn were built against the conda numpy.

Run 3.12 alongside 3.13 once to tell a 3.13-specific problem apart from
version-independent env fragility, then drop back to one version.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD
Keeps this PR to one subject: getting CI running again. The stale
`how="modified_quantile"` doctest values are a pre-existing failure on main,
not something this PR introduces, and they belong in their own change.

Consequence: the `doctest` job is expected to stay red here for the same
reason it is red on main.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD
@aaronspring aaronspring changed the title Fix CI/CD issues on main Restore GitHub Actions CI Sep 20, 2026
Running 3.12 alongside 3.13 answered the question: both segfault, so the
crash is not caused by choosing 3.13. Back to a single 3.13 entry.

  3.13  failed x2   test_remove_bias_unfair...[basic_quantile-season-same_verifs]
  3.12  failed x2   test_remove_bias_unfair...[additive_mean-month-maximize]

The crash point moves between interpreter versions, which points at memory
corruption rather than a logic bug, consistent with the numpy ABI mismatch
in ci/install-upstream-wheels.sh: nightly numpy/pandas are pip-installed
with --no-deps over a conda env whose netCDF4, bottleneck, numba and sklearn
were compiled against the conda numpy. Fixing that script is its own change.

This commit drops the [test-upstream] keyword. The job stays fully enabled
on its nightly schedule and on any PR that opts in; it is simply no longer
forced onto this CI-restoration PR, where it reports a pre-existing
environment problem this PR does not cause.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD
@aaronspring

Copy link
Copy Markdown
Collaborator Author

upstream-dev segfault — not this PR's defect, needs its own fix

Recording this since it surfaced here and should not be lost.

What fails. upstream-dev crashes hard while opening a tutorial dataset during fixture setup:

netCDF4._netCDF4                                    ← last extension module loaded
xarray/backends/netCDF4_.py:563 in ds
xarray/backends/api.py:619 in open_dataset
src/climpred/tutorial.py:223 in load_dataset
src/climpred/conftest.py:413 in hindcast_NMME_Nino34
Segmentation fault (core dumped)

It is not caused by the Python version bump in this PR. I ran 3.12 alongside 3.13 once to check, and both segfault:

Python Result Crash point
3.13 failed ×2 test_remove_bias_unfair_artificial_skill_over_fair[basic_quantile-season-same_verifs]
3.12 failed ×2 test_remove_bias_unfair_artificial_skill_over_fair[additive_mean-month-maximize]

The crash point moves between interpreter versions, which points at memory corruption rather than a logic bug.

Likely cause. ci/install-upstream-wheels.sh pip-installs nightly numpy/pandas with --no-deps --upgrade on top of a conda environment whose netCDF4, bottleneck, numba and sklearn were compiled against the conda numpy. That is the classic setup for an ABI mismatch, and numpy.ndarray size changed, may indicate binary incompatibility shows up in these environments as the usual precursor. Fixing it properly means reworking how that script layers nightly wheels over conda-built extensions — a separate change from restoring CI, so I have not bundled it here.

Why this PR no longer runs the job. upstream-dev is a nightly canary against bleeding-edge upstream mains, not a merge gate. It only ran on this PR because I added a [test-upstream] keyword to verify that the Python bump fixed the install-time failure. It did — the environment now builds, import climpred succeeds, and the suite runs for several minutes instead of dying in Set up conda environment within 40 seconds. That verification is done, so later commits drop the keyword. The job remains fully enabled on its weekly schedule and on any PR that opts in; nothing is skipped, disabled or quarantined.

One incidental confirmation from the same runs: the ::set-output → $GITHUB_OUTPUT fix works — the logs now show Set output 'artifacts_availability', so the report job will once again open/update the upstream-failure issue. That reporting path had been silently broken, which is plausibly why these weekly failures went unnoticed for so long.


Generated by Claude Code

@codecov

codecov Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 91.27%. Comparing base (9368dca) to head (6452311).
⚠️ Report is 3 commits behind head on main.

Additional details and impacted files
@@           Coverage Diff           @@
##             main     #929   +/-   ##
=======================================
  Coverage   91.27%   91.27%           
=======================================
  Files          59       59           
  Lines        6548     6548           
=======================================
  Hits         5977     5977           
  Misses        571      571           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

…push runs [test-upstream]

Two upstream-CI fixes.

1. install-upstream-wheels.sh resolved numpy and pandas against
   scipy-wheels-nightly, while matplotlib two lines below already used
   scientific-python-nightly-wheels. The scientific Python projects
   consolidated their nightlies on the latter and the former is no longer
   updated, so `--pre --upgrade` picked up a numpy far older than the
   `numpy >=2.0` the conda environment was solved with. Swapping numpy out
   from under conda-built extensions compiled against 2.x is what segfaulted
   the suite in netCDF4 during fixture setup, on 3.12 and 3.13 alike.

2. upstream-dev.yml declared both `on: push` and `on: pull_request`, so every
   branch push started two identical runs. The concurrency group keys on
   `github.ref`, which differs between refs/heads/<branch> and
   refs/pull/<n>/merge, so the pair never deduplicated. Pull requests, the
   weekly schedule and manual dispatch already cover everything `push` did.

The nightly index could not be queried from the sandbox (the egress proxy
blocks pypi.anaconda.org), so fix 1 is verified by this run rather than
locally; [test-upstream] opts the job in for that purpose.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD
@aaronspring
aaronspring merged commit 30c91b9 into main Sep 21, 2026
20 of 22 checks passed
@aaronspring
aaronspring deleted the claude/cicd-main-issues-89krmb branch September 21, 2026 07:48
@aaronspring

Copy link
Copy Markdown
Collaborator Author

Correction: the stale nightly index was not the cause of the upstream-dev segfault

Correcting my comment above, which pointed at ci/install-upstream-wheels.sh resolving numpy/pandas from the unmaintained scipy-wheels-nightly index. That inconsistency is real — matplotlib two lines below already used scientific-python-nightly-wheels — and the switch is worth keeping on its own terms, but it does not fix the segfault.

With the index switched (6452311):

before (5b8a213) after (6452311)
Set up conda environment success success, 64s
Run Tests failed after 5m46s / 7m04s failed after 9m49s
crash point …[basic_quantile-season-same_verifs] (3.13), …[additive_mean-month-maximize] (3.12) …[modified_quantile-season-same_verifs]
pandas C extension pandas._libs.json pandas._libs._ujson

So the change did take effect — a different pandas build was installed and the suite ran ~3 minutes further — but it still ends in Segmentation fault (core dumped), always inside test_remove_bias_unfair_artificial_skill_over_fair, with netCDF4._netCDF4 loaded.

What this points to instead. The crash point moves whenever the installed numpy/pandas change, which is the signature of memory corruption rather than a defect in any one package. The likely root cause is structural: the script force-uninstalls numpy/pandas and reinstalls nightlies with --no-deps, while leaving conda-built extensions (netCDF4, bottleneck, numba, sklearn, pyarrow, scipy) that were compiled against the numpy the environment was originally solved with. Whichever numpy the nightly resolves to is a different C-API build than those extensions expect.

If that is right, the fix is not a URL but a restructuring — either also reinstall the compiled extensions from pip so they build against the nightly numpy, or stop force-uninstalling numpy and let the nightlies come in as a consistent set. That needs iterating in CI against the real environment, so it belongs in its own PR.

Worth restating: upstream-dev is a weekly canary against bleeding-edge upstream mains, not a merge gate, and nothing about it is skipped, disabled or quarantined.


Generated by Claude Code

aaronspring added a commit that referenced this pull request Sep 21, 2026
)

# Description

The `HindcastEnsemble.remove_bias` example for `how="modified_quantile"`
pins values that no longer match what `bias_correction` produces, so the
`Doctests` job fails on `main`:

```diff
-    SST      (lead) float64 80B 0.07628 0.08293 0.08169 ... 0.1577 0.1821 0.2087
+    SST      (lead) float64 80B 0.0766 0.08275 0.08152 ... 0.1573 0.1817 0.2085
```

Split out of #929 so that PR stayed on a single subject. With CI
restored there, the `Doctests` job now actually runs, so this should
take it green.

The new values reproduce **identically** under CI's conda/Python 3.10
environment and under a local pip/Python 3.11 one (numpy 2.4.6, pandas
3.0.6, xarray 2026.7.0) — two quite different environments. That points
to a genuine change in the upstream package rather than environment
noise, which is why the expected output is updated rather than masked
behind an `ELLIPSIS` directive: these numbers are rendered in the
published documentation, so they should be correct rather than merely
non-failing.

## Type of change

-   [x]  Bug fix (non-breaking change which fixes an issue)
-   [x]  Improved Documentation

# How Has This Been Tested?

- `pytest --doctest-modules src/climpred/classes.py -k remove_bias` → 1
passed.
-   `pre-commit run --all-files` passes.

- [ ] Tests added for `pytest`, if necessary. — not applicable; this
corrects an existing doctest's expected output.

## Checklist (while developing)

-   [x]  CHANGELOG is updated with reference to this PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD

---------

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants