Restore GitHub Actions CI - #929
Conversation
The weekly CI has been red and partially inert. Four separate problems: - testing.yml: `detect-ci-trigger` only runs for push/pull_request, so on `schedule` and `workflow_dispatch` it is skipped, which also skipped the three dependent test-matrix jobs. The weekly run silently degraded to just the doctest and notebook jobs. Gate on `always()` plus an explicit `triggered != 'true'` so the matrix runs for every event. - upstream-dev-ci.yml: pinned to Python 3.10, but upstream xclim now requires >=3.11, so `Set up conda environment` failed at install time before any test ran. Bumped to 3.13, which testing.yml already exercises. - upstream-dev-ci.yml: `::set-output` was removed by GitHub, so ARTIFACTS_AVAILABLE was never set and the `report` job never opened or updated the upstream-failure issue. Use `$GITHUB_OUTPUT`. - classes.py: the `HindcastEnsemble.remove_bias` doctest for `how="modified_quantile"` pinned stale values. Current `bias_correction` output reproduces identically under CI's conda/py3.10 env and a pip/py3.11 env, so the expected values are updated rather than masked. Also drop a dangling always-true `and "lead"` clause in `_rps` that failed the `ty` pre-commit hook; the condition is unchanged. Verified: `pre-commit run --all-files` passes, the `remove_bias` doctest passes, and 175 rps/crps tests pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD
The entries added for these CI fixes had no `:pr:` reference, which the repo's changelog convention and the PR checklist both call for. The `[test-upstream]` keyword makes upstream-dev-ci.yml's detect-ci-trigger fire on this PR, so the Python 3.13 bump is exercised here rather than only after merge on the weekly schedule. It does not affect testing.yml, which keys off `[skip-ci]`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD
GitHub had set both workflows to `disabled_inactivity` after the repository
went 60 days without a commit (last commit 2026-07-07, last run 2026-09-06).
That state belongs to the workflow record, which GitHub keys to the file's
path, so pushing commits does not clear it and `workflow_dispatch` is refused
outright. Renaming the files registers them as new workflow records, which
start active:
testing.yml -> ci.yml
upstream-dev-ci.yml -> upstream-dev.yml
Job names, workflow `name:` values and triggers are all unchanged, so branch
protection rules that key off check-run names keep matching.
Also fixes the CI badge in README.rst and docs/source/index.rst, which still
pointed at `climpred_testing.yml` from an earlier rename and so rendered no
status at all. Both badges now track the current paths.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD
Seven entries for what amounts to "fix CI" was too granular. The badge fix, the `::set-output` replacement and the `_rps` lint cleanup are invisible to users and do not need their own lines; they are part of restoring CI or are pure internals. Left with two: one for CI being restored and repaired, one for the doctest output that users actually see in the docs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD
…tream]
The upstream-dev job now gets far enough to run tests (the 3.13 bump fixed
the install-time failure) but segfaults in netCDF4's C extension while
opening a tutorial dataset in fixture setup:
netCDF4._netCDF4
xarray/backends/netCDF4_.py:563 in ds
src/climpred/tutorial.py:223 in load_dataset
src/climpred/conftest.py:413 in hindcast_NMME_Nino34
Segmentation fault (core dumped)
That is the signature of a numpy ABI mismatch: install-upstream-wheels.sh
pip-installs nightly numpy/pandas with --no-deps over a conda env whose
netCDF4, bottleneck, numba and sklearn were built against the conda numpy.
Run 3.12 alongside 3.13 once to tell a 3.13-specific problem apart from
version-independent env fragility, then drop back to one version.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD
Keeps this PR to one subject: getting CI running again. The stale `how="modified_quantile"` doctest values are a pre-existing failure on main, not something this PR introduces, and they belong in their own change. Consequence: the `doctest` job is expected to stay red here for the same reason it is red on main. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD
Running 3.12 alongside 3.13 answered the question: both segfault, so the crash is not caused by choosing 3.13. Back to a single 3.13 entry. 3.13 failed x2 test_remove_bias_unfair...[basic_quantile-season-same_verifs] 3.12 failed x2 test_remove_bias_unfair...[additive_mean-month-maximize] The crash point moves between interpreter versions, which points at memory corruption rather than a logic bug, consistent with the numpy ABI mismatch in ci/install-upstream-wheels.sh: nightly numpy/pandas are pip-installed with --no-deps over a conda env whose netCDF4, bottleneck, numba and sklearn were compiled against the conda numpy. Fixing that script is its own change. This commit drops the [test-upstream] keyword. The job stays fully enabled on its nightly schedule and on any PR that opts in; it is simply no longer forced onto this CI-restoration PR, where it reports a pre-existing environment problem this PR does not cause. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD
|
| Python | Result | Crash point |
|---|---|---|
| 3.13 | failed ×2 | test_remove_bias_unfair_artificial_skill_over_fair[basic_quantile-season-same_verifs] |
| 3.12 | failed ×2 | test_remove_bias_unfair_artificial_skill_over_fair[additive_mean-month-maximize] |
The crash point moves between interpreter versions, which points at memory corruption rather than a logic bug.
Likely cause. ci/install-upstream-wheels.sh pip-installs nightly numpy/pandas with --no-deps --upgrade on top of a conda environment whose netCDF4, bottleneck, numba and sklearn were compiled against the conda numpy. That is the classic setup for an ABI mismatch, and numpy.ndarray size changed, may indicate binary incompatibility shows up in these environments as the usual precursor. Fixing it properly means reworking how that script layers nightly wheels over conda-built extensions — a separate change from restoring CI, so I have not bundled it here.
Why this PR no longer runs the job. upstream-dev is a nightly canary against bleeding-edge upstream mains, not a merge gate. It only ran on this PR because I added a [test-upstream] keyword to verify that the Python bump fixed the install-time failure. It did — the environment now builds, import climpred succeeds, and the suite runs for several minutes instead of dying in Set up conda environment within 40 seconds. That verification is done, so later commits drop the keyword. The job remains fully enabled on its weekly schedule and on any PR that opts in; nothing is skipped, disabled or quarantined.
One incidental confirmation from the same runs: the ::set-output → $GITHUB_OUTPUT fix works — the logs now show Set output 'artifacts_availability', so the report job will once again open/update the upstream-failure issue. That reporting path had been silently broken, which is plausibly why these weekly failures went unnoticed for so long.
Generated by Claude Code
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #929 +/- ##
=======================================
Coverage 91.27% 91.27%
=======================================
Files 59 59
Lines 6548 6548
=======================================
Hits 5977 5977
Misses 571 571 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
…push runs [test-upstream] Two upstream-CI fixes. 1. install-upstream-wheels.sh resolved numpy and pandas against scipy-wheels-nightly, while matplotlib two lines below already used scientific-python-nightly-wheels. The scientific Python projects consolidated their nightlies on the latter and the former is no longer updated, so `--pre --upgrade` picked up a numpy far older than the `numpy >=2.0` the conda environment was solved with. Swapping numpy out from under conda-built extensions compiled against 2.x is what segfaulted the suite in netCDF4 during fixture setup, on 3.12 and 3.13 alike. 2. upstream-dev.yml declared both `on: push` and `on: pull_request`, so every branch push started two identical runs. The concurrency group keys on `github.ref`, which differs between refs/heads/<branch> and refs/pull/<n>/merge, so the pair never deduplicated. Pull requests, the weekly schedule and manual dispatch already cover everything `push` did. The nightly index could not be queried from the sandbox (the egress proxy blocks pypi.anaconda.org), so fix 1 is verified by this run rather than locally; [test-upstream] opts the job in for that purpose. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD
Correction: the stale nightly index was not the cause of the
|
before (5b8a213) |
after (6452311) |
|
|---|---|---|
Set up conda environment |
success | success, 64s |
Run Tests |
failed after 5m46s / 7m04s | failed after 9m49s |
| crash point | …[basic_quantile-season-same_verifs] (3.13), …[additive_mean-month-maximize] (3.12) |
…[modified_quantile-season-same_verifs] |
| pandas C extension | pandas._libs.json |
pandas._libs._ujson |
So the change did take effect — a different pandas build was installed and the suite ran ~3 minutes further — but it still ends in Segmentation fault (core dumped), always inside test_remove_bias_unfair_artificial_skill_over_fair, with netCDF4._netCDF4 loaded.
What this points to instead. The crash point moves whenever the installed numpy/pandas change, which is the signature of memory corruption rather than a defect in any one package. The likely root cause is structural: the script force-uninstalls numpy/pandas and reinstalls nightlies with --no-deps, while leaving conda-built extensions (netCDF4, bottleneck, numba, sklearn, pyarrow, scipy) that were compiled against the numpy the environment was originally solved with. Whichever numpy the nightly resolves to is a different C-API build than those extensions expect.
If that is right, the fix is not a URL but a restructuring — either also reinstall the compiled extensions from pip so they build against the nightly numpy, or stop force-uninstalling numpy and let the nightlies come in as a consistent set. That needs iterating in CI against the real environment, so it belongs in its own PR.
Worth restating: upstream-dev is a weekly canary against bleeding-edge upstream mains, not a merge gate, and nothing about it is skipped, disabled or quarantined.
Generated by Claude Code
) # Description The `HindcastEnsemble.remove_bias` example for `how="modified_quantile"` pins values that no longer match what `bias_correction` produces, so the `Doctests` job fails on `main`: ```diff - SST (lead) float64 80B 0.07628 0.08293 0.08169 ... 0.1577 0.1821 0.2087 + SST (lead) float64 80B 0.0766 0.08275 0.08152 ... 0.1573 0.1817 0.2085 ``` Split out of #929 so that PR stayed on a single subject. With CI restored there, the `Doctests` job now actually runs, so this should take it green. The new values reproduce **identically** under CI's conda/Python 3.10 environment and under a local pip/Python 3.11 one (numpy 2.4.6, pandas 3.0.6, xarray 2026.7.0) — two quite different environments. That points to a genuine change in the upstream package rather than environment noise, which is why the expected output is updated rather than masked behind an `ELLIPSIS` directive: these numbers are rendered in the published documentation, so they should be correct rather than merely non-failing. ## Type of change - [x] Bug fix (non-breaking change which fixes an issue) - [x] Improved Documentation # How Has This Been Tested? - `pytest --doctest-modules src/climpred/classes.py -k remove_bias` → 1 passed. - `pre-commit run --all-files` passes. - [ ] Tests added for `pytest`, if necessary. — not applicable; this corrects an existing doctest's expected output. ## Checklist (while developing) - [x] CHANGELOG is updated with reference to this PR. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD --------- Co-authored-by: Claude <noreply@anthropic.com>
Description
CI on this repository has not run at all since 2026-09-06, and the weekly run was partly inert for a long time before that. This PR is limited to getting it working again.
1. Both workflows were disabled. GitHub had set them to
disabled_inactivity, which it does to scheduled workflows after 60 days without a commit (last commit tomainwas 2026-07-07). That state belongs to the workflow record, which GitHub keys to the file's path — so pushing commits does not clear it, andworkflow_dispatchis refused outright. Registering the files at new paths creates fresh workflow records, which startactive:Job names, workflow
name:values and triggers are unchanged, so branch-protection rules that key off check-run names keep matching. Confirmed working:Upstream Testrun #1 under a new workflow id fired on the push that did the rename.2. The weekly CI never actually ran the tests.
detect-ci-triggeris gated onpush/pull_request, so onscheduleandworkflow_dispatchit is skipped — taking the three test-matrix jobs thatneeds:it down with it. Run 144 (2026-09-06) shows all three as"conclusion": "skipped"; the weekly run had quietly degraded to just the doctest and notebook jobs. Now gated onalways()plus an explicittriggered != 'true'.3. Upstream Test died before running a single test, pinned to Python 3.10 while upstream
xclimrequires>=3.11:4. Upstream failure reporting was broken — it used
::set-output, which GitHub removed, soARTIFACTS_AVAILABLEwas never set and thereportjob never opened its failure issue. Now$GITHUB_OUTPUT. This is plausibly why the weekly failures went unnoticed for so long.Also fixes the CI badge in
README.rstanddocs/source/index.rst(it pointed atclimpred_testing.yml, a path removed in an earlier rename, so it rendered nothing), and drops a dangling always-trueand "lead"clause in_rpsthat fails thetypre-commit hook onmain.Type of change
How Has This Been Tested?
Actions is demonstrably alive again: the rename produced a new workflow record and a run on this PR.
Fixes 3 and 4 are confirmed by that run's log — the environment built and pytest ran (instead of failing at install), and the log shows
Set output 'artifacts_availability'.pre-commit run --all-filespasses, including the workflow-schema hook andty.Tests added for
pytest, if necessary. — not applicable; this is CI configuration.Checklist (while developing)
Known red checks, both pre-existing on
maindoctest— theHindcastEnsemble.remove_biasdoctest forhow="modified_quantile"pins stale values that no longer match currentbias_correctionoutput. Reproduces onmain; deliberately left for its own PR so this one stays on a single subject.upstream-dev— segfaults in netCDF4's C extension while opening a tutorial dataset in fixture setup (conftest.py:413→load_dataset→xr.open_dataset). This is the signature of a numpy ABI mismatch:ci/install-upstream-wheels.shpip-installs nightly numpy/pandas with--no-depsover a conda env whose netCDF4, bottleneck, numba and sklearn were compiled against the conda numpy. It needs its own fix in that script. Note this job is a nightly canary against bleeding-edge upstream mains, not a merge gate — it only runs here because of a[test-upstream]keyword used to validate the Python bump.🤖 Generated with Claude Code
https://claude.ai/code/session_012xXJVZrXR4k3BmmDeTLEaD