Repository navigation
Conversation
samsrabin
force-pushed
the
cirrus-runner-workflows
branch
from
August 6, 2025 21:21
00b95ee to
7114913
Compare
Member
Author
|
The ubuntu runner is currently failing because the |
samsrabin
force-pushed
the
cirrus-runner-workflows
branch
from
July 10, 2026 21:20
d4f93ae to
ee33e2c
Compare
Troubleshooting list-glade-cesm-input: Add `sleep 600`.
list-glade-cesm-input: Delete troubleshooting `sleep`.
…sldev-x86_64-almalinux9-gcc14-mpich:25.08.
Move every stale ARG to what config_machines.xml loads at ccs_config_cesm1.0.88: GCC 14.3.0, netCDF-C 4.9.3, netCDF-Fortran 4.6.2, HDF5 1.14.6, PnetCDF 1.14.1, ESMF 8.9.1, mpi-serial 2.5.3, PIO 2.6.8. MPICH stays 3.4.3, since cray-mpich 8.1.32 is still MPICH-3.4-ABI- derived; only its deviation guard moves. All ten checks now pass. HDF5 needed more than a number. Upstream changed its tag scheme between these releases, from hdf5-1_12_2 to hdf5_1.14.6, so the URL, the tarball name and the extracted directory all change with it. Every other new version's download URL and extracted directory name was checked against upstream and needed nothing. gnu_container.cmake's hardcoded ESMFMKFILE moves with ESMF_VERSION, as the build-time assertion added earlier requires. Not built, not run. The published image still predates this, so the passing check describes the recipe rather than anything on GHCR; see NEXT_STEPS item 6 for why this should not merge ahead of the rebuild. Co-Authored-By: Claude Opus 5 <[email protected]>
build-on-casper.sh located its Dockerfile and build context next to itself, via BASH_SOURCE. That holds when the script is run in place, but PBS executes a copy from /var/spool/pbs/mom_priv/jobs, where nothing of this repo sits, so a qsub submission died before the first Dockerfile instruction with "stat .../jobs/Dockerfile: no such file or directory". Fall back to PBS_O_WORKDIR, accepting either the repo root or this directory, with CTSM_BUILD_CONTEXT as an override and a message naming both places searched when none of them has the file. Co-Authored-By: Claude Opus 5 <[email protected]>
ftp.gnu.org is unreachable from NCAR compute nodes. It times out rather than refusing, so wget exhausts the patient retry budget set in wgetrc and the step exits 4 before anything compiles. DNS resolves, so this is routing, and ftpmirror.gnu.org is no help since it is the same host. Try gcc.gnu.org's own release area first, then a kernel.org mirror, then ftp.gnu.org. Per-attempt --tries and --timeout override the global settings so an unreachable mirror is abandoned in seconds instead of eating the job's walltime, and -O keeps a failed attempt from leaving a partial file for the next one to trip over. The loop exits nonzero when every mirror fails, so tar is never reached with an empty file. Every other download host the Dockerfile uses was checked and is reachable; gnu.org was the only one. Verified by reachability checks and by exercising the loop's exit status in both directions. Not built: the image has not been rebuilt. Co-Authored-By: Claude Opus 5 <[email protected]>
netCDF-C 4.9.3 stopped folding $LIBS into the NC_LIBS it substitutes into nc-config: configure.ac now sets NC_LIBS="$LDFLAGS $NC_LIBS" where 4.9.2 set "$LDFLAGS $NC_LIBS $LIBS". Plain --libs therefore returns a bare -lnetcdf. Against a shared netCDF that is enough, because DT_NEEDED pulls the rest in, but /usr/local/serial is static-only, so the netcdf-fortran configure link test for nc_open failed and the step died with "Could not link to netcdf C library". Pass --libs --static there, which restores -lhdf5_hl -lhdf5 -lm -lz -lxml2 -lcurl. ESMF reaches the same prefix through the same broken idiom: it runs nc-config --libs itself and keeps the -l options. It skips that query when ESMF_NETCDF_LIBS is set in the environment, so the mpiuni ESMF step now sets it. Only the serial ESMF is affected; the two MPI ones link the shared netCDF under /usr/local. Verified by instantiating both versions' real nc-config.in with the values this build logged, and comparing what --libs and --libs --static return. Not built. Co-Authored-By: Claude Opus 5 <[email protected]>
podman's storage here is node-local, and node-local storage is wiped when the allocation ends. The script built, tagged and exited, so a batch build reported success and left nothing behind: the image went with the node. Save to /glade/work/$USER after a successful build, which set -e and pipefail make the only way to reach that point. An existing tarball is never clobbered; the name falls back to one with a timestamp. Skippable with CTSM_BUILD_NO_SAVE=1, relocatable with CTSM_BUILD_SAVEDIR. Co-Authored-By: Claude Opus 5 <[email protected]>
smoke-test.sh asserted six versions written out by hand, so the version bump left it failing a correct image: gcc 14.3.0 against an expected 12.2.0. A second copy of the same numbers was always going to drift from the ARGs, and it did so the first time they moved. Take the expectations from the image's own labels, which the Dockerfile sets from those ARGs. That removes the copy and makes the assertion mean something it did not before: that the toolchain the build produced is the one the ARGs asked for, so a cached or substituted layer is caught rather than assumed away. A missing label fails loudly instead of silently checking nothing. mpichversion's whole output is matched rather than its first line, since nothing guarantees the Version: line comes first. Verified against a stub podman for the label lookup, the <no value> fallback and the missing-label guard, and by checking the real gcc 14.3.0 output against both the new and old expectations. Co-Authored-By: Claude Opus 5 <[email protected]>
Both were labelled on the image but never asserted, so the two libraries most likely to be picked up from the wrong place were the two nothing checked. HDF5 is read from h5dump, not h5cc: a parallel build installs h5pcc instead, so the wrapper's name depends on configuration while the tools are installed either way. The wrappers remain as fallbacks, and an absent version fails rather than passing on an empty string. ESMF is read from ESMF_VERSION_STRING in the esmf.mk that ESMFMKFILE points at, which is the file CTSM consumes, rather than from a directory name that happens to contain digits. Verified against stubs: both sources, the h5pcc fallback, the failure when no tool reports a version, and rejection of a wrong expectation. Co-Authored-By: Claude Opus 5 <[email protected]>
ESMFMKFILE names the MPI debug ESMF. Every mpi-serial case and the whole Fortran unit-test path instead link the mpiuni one, which the baked-in CIME macro selects by a hardcoded path. That path was until now checked only at build time, which says nothing about the image in hand: a mis-tagged image or an older tarball satisfies an assertion that ran in some other build. Read the path back out of the macro in the image and require it to exist, to be an ESMF_COMM=mpiuni build, and to report the labelled version. Verified against stubs for the healthy case and for the three ways it goes wrong: ESMF_VERSION bumped without editing the macro, a macro pointing at an MPI build, and a version that disagrees with the label. Co-Authored-By: Claude Opus 5 <[email protected]>
The image was rebuilt on Casper against derecho's current stack and passes all four validation scripts, including the Fortran unit tests and a single-point system test. Item 6 now covers only publishing to GHCR and repointing cirrus-testing.yml, which is also what still blocks a merge: until then the repo describes a stack nothing in CI runs. Also records that the GCC 12 to 14 jump, flagged as the risky part, produced no compiler diagnostics at all. The two real breakages came from elsewhere, HDF5's tag scheme and nc-config, and are documented at the sites that hit them. The original reasoning is kept, since it remains the right thing to check first on the next compiler jump. Co-Authored-By: Claude Opus 5 <[email protected]>
Whether this image can be built in CI turns on disk more than on time: a hosted runner has roughly 14 GB free, the image is 4.1 GB as a tarball, and the build transiently needs GCC's build tree and three ESMF trees on top of that. That number decides whether item 8 is a straightforward workflow or needs the Dockerfile split into a base image and a thin top layer, so it is worth knowing before any of it is designed. Probe both architectures, since the arm64 runner's availability is part of the same question. Phase A reports what a runner has before and after reclaiming the preinstalled toolchains and takes a couple of minutes; Phase B, opt-in, attempts the real build and reports how far it gets. MAKE_JOBS comes from nproc there, as the Dockerfile's default of 16 would thrash a runner this size. Informational and workflow_dispatch only, following probe-derecho-modules.yml: every step is best-effort and the log is the result. Not run: triggering it needs the branch pushed. Co-Authored-By: Claude Opus 5 <[email protected]>
The build is going to live in a workflow whichever runner it lands on, so the hosted runners' disk limit decides where it runs rather than whether it runs at all. Add gha-runner-ctsm to the spike, since it is the fallback and the same measurements decide it. The reclaim step is skipped there: it is shared, persistent hardware, and deleting its toolchains would not be cleanup. The build step picks docker or podman by what exists, as the NCAR machine may have only podman, and reports podman's graphroot, which is node-local on the Casper machines and may be here too. A job aimed at an offline or busy self-hosted runner queues instead of failing, and timeout-minutes does not bound queue time, so that is called out where someone would otherwise read a pending row as a hung workflow. Co-Authored-By: Claude Opus 5 <[email protected]>
The preceding commit added gha-runner-ctsm to the spike but left item 8 describing Phase A as covering only the two hosted architectures, and said a tight disk result means the base-image split, which is no longer the only alternative. Co-Authored-By: Claude Opus 5 <[email protected]>
The probe pinned ncarenv/23.09 and its gcc, cray-mpich and netcdf-mpi, so once derecho moved to ncarenv/25.10 it reported a module-load failure that said nothing about the runner it exists to test. Load the defaults instead. That cannot go stale, and it is what the Phase 2 cron wants anyway: the question is what derecho has today, not whether it still has what someone wrote down here. The readout now says to compare against config_machines.xml, since a difference there is precisely the drift Phase 2 is for. Also notes that the hdf5 line is now a depends_on rather than a version, HDF5 having become a separate module under ncarenv/25.10. Not run: triggering it needs the branch pushed, and it only runs on the self-hosted runner. Co-Authored-By: Claude Opus 5 <[email protected]>
:20261007 is published, so point cirrus-testing.yml at it and close item 6. Its config digest matches the one in the Casper build log, so what is on GHCR is the artifact that passed validation rather than a re-tag of something else. :20260831 stays on GHCR as a fallback and is still what the September validation record loads; those references are left alone, since they describe runs that happened against that image. Co-Authored-By: Claude Opus 5 <[email protected]>
GHCR associates a package with a repository through org.opencontainers.image.source, which this image never set. The package is therefore absent from the repository-filtered package list, which reads as "the push did not work" even when it did; it is only findable from the organization's unfiltered package page. Setting the label fixes that for images built from here on. It cannot be applied retroactively: :20261007 and earlier carry no source label, so until the next rebuild the package stays linkable only by hand, from the package settings page. Co-Authored-By: Claude Opus 5 <[email protected]>
Loading derecho's stack from anywhere else does not work and cannot be made to. ncarenv and gcc load fine off derecho's tree, but cray-mpich dies -- the path /opt/cray/pe/mpich/8.1.32/ofi/gnu/12.3 does not exist -- because Cray PE is on derecho hardware only, and netcdf-mpi sits below cray-mpich in the hierarchy so it never becomes loadable either. The previous probe would also have loaded the runner's own ncarenv when that tree was on MODULEPATH, reporting a different machine's versions as if they were derecho's. None of it is needed. Every value the check wants is a file under glade: the default version is the "default" symlink beside the modulefiles, the hdf5 pairing is netcdf-mpi's own depends_on line, and netCDF-Fortran comes from derecho's installed nf-config run by absolute path -- nc-config and nf-config are generated shell scripts that echo baked-in strings and load nothing. So the probe now asks whether glade is mounted, resolves derecho's current defaults by reading symlinks, and derives both [snapshot] values without loading a module. Lmod is reported as context. Resolving defaults rather than taking the highest version matters: cray-mpich 9.0.0 is staged in the tree while default still points at 8.1.32. Also correct a claim in the checker: derecho does have a pfunit module now. CTSM still does not use it, so the PFUNIT_PATH reading is unchanged. Probe body run on a casper login node; every version it reports agrees with the Dockerfile ARGs and the recorded snapshot. Not run in Actions. Co-Authored-By: Claude Opus 5 <[email protected]>
Item 5 assumed the drift check had to run as a cron on an NCAR machine, because it assumed the check had to load derecho's modules. It does not: every value it wants is a file under glade, so it can run wherever glade is mounted, and a GitHub workflow opening the drift issue needs only its own GITHUB_TOKEN rather than a PAT parked on a shared machine. Record where each value lives, why loading the stack is not an option off derecho, why defaults must be resolved from the symlinks instead of taken as the highest version, and the four things still undecided -- which runner, the default-branch-only rule for schedules, silent queueing when a self-hosted runner is down, and whether to report success at all. Co-Authored-By: Claude Opus 5 <[email protected]>
Two decisions reshape what is left, both recorded with their evidence in the new DRIFT_CHECK_DESIGN.md, which the items now point to instead of carrying the argument inline. The container's scope becomes gnu + mpi-serial only, so mpich leaves the image along with PnetCDF, the MPI-linked netCDF/HDF5 and two of the three ESMF flavors. The drift check compares built executables rather than version metadata. Metadata comparison is blind to drift that exists now: a container case build links MPIserial_2.5.4 from CTSM's submodule while derecho links its mpi-serial/2.5.3 module, and check-derecho-versions.py calls that a match. Setting MPI_SERIAL_PATH and PIO_LIBDIR in gnu_container.cmake makes the container use derecho's mechanism rather than only matching its numbers. Item 4 is unchecked and rewritten as 4a, the checker, and 4b, the rebuild. Item 5 is rewritten around the above. Unchecked items are now in the order they should be tackled, and each carries its own README updates. The component list in "Where things stand" is corrected to the versions the 2026-10-07 rebuild produced. Also records the podman XDG_RUNTIME_DIR constraint, which bites interactive podman but not build-on-casper.sh. Documentation only; nothing built or run. Co-Authored-By: Claude Opus 5 <[email protected]>
VALIDATION_2026-09.md was reachable only from item 3 of NEXT_STEPS, so a reader starting from the design doc would not find it and might redo the wrapper validation or revisit decisions it already settled. Reference it twice: in the context, as settled work not to reopen, and in the verification section, where phase 4 re-runs its procedure against the rebuilt image. The second reference records why its section ordering matters, since that is the part most likely to be skipped. Documentation only. Co-Authored-By: Claude Opus 5 <[email protected]>
The container's scope is gnu + mpi-serial only, so the check now reads the mpilib="mpi-serial" blocks of config_machines.xml. Gone: the MPI-flavor checks, the cray-mpich guard, the <MPILIBS> assertion and the two-stack twin machinery. cray-libsci is guarded instead; its stand-in carries no version ARG, so deviation mode no longer requires one. MPICH_VERSION and PNETCDF_VERSION stay in the Dockerfile, unchecked, until the image is stripped to serial-only. [snapshot] is re-measured against the serial netcdf module and the hdf5 it pulls in. The numbers did not move, only the provenance; measured on Casper 2026-10-09, with the sources recorded in the ini. New: an existence check resolving every pinned module in derecho's module tree under /glade, because config_machines.xml is CTSM's claim about derecho and nothing else verifies it. That needs /glade, so the workflow now picks its runner per event and the hosted PR gate passes --skip-module-tree. A stdlib unittest file runs as its own CI step. Verified on Casper: the tests pass and the checker passes every check including the live module-tree read. The workflow itself has not been run. Co-Authored-By: Claude Opus 5 <[email protected]>
workflow_dispatch is unavailable until the file reaches the default branch, which leaves the probe unrunnable on the branch that introduced it. pull_request carries no such restriction, and a fork's PR runs in the base repository's context, where gha-runner-ctsm is reachable. The paths filter names the probe itself, so the commit adding the trigger fires it and later PRs do not. Not run; running it is the point. Co-Authored-By: Claude Opus 5 <[email protected]>
The probe exited at its first check, so a negative answer carried no detail. gha-runner-ctsm is a Kubernetes pod, so glade arrives as specific bind mounts rather than as a filesystem; which ones are present decides whether the fix is one more mount or a different host for the check. Not run. Co-Authored-By: Claude Opus 5 <[email protected]>
The probe has now been run. The runner is a Kubernetes pod, so glade arrives as specific bind mounts; the inputdata path is among them and the module tree is not. That decides nothing about the drift check's design, but it rules out the host item 5 assumed and makes the scheduled half of derecho-version-check.yml unrunnable until either a mount is added or the check moves to derecho. Co-Authored-By: Claude Opus 5 <[email protected]>
The second probe run named it: a single read-only NFSv4 export of the campaign cesmdata path, with campaign the only entry under /glade. Splitting the ask in two matters because the modulefile trees are text and the Spack prefix is not, and only netCDF-Fortran's version needs the latter. Co-Authored-By: Claude Opus 5 <[email protected]>
The concretized environment file gives netCDF-Fortran, netCDF-C and HDF5 in one read, and the token in a module's install prefix is a unique prefix of that spec's dag hash, so the versions tie to the module the serial path loads rather than to whatever the environment happens to hold once. It also removes the Spack install tree from what the Cirrus runner would need mounted, leaving three text paths. Verified by resolving the gnu serial netcdf/4.9.3 prefix token against spack.lock; the instrument is not yet wired into any check. Co-Authored-By: Claude Opus 5 <[email protected]>
spack.lock lives under an ncarenv version directory, so naming it in the request would tie the mount to 25.10 and need re-requesting on every bump -- an event this check exists to notice. The parent directories carry no version. Also records the filesystem the request concerns, since the existing mount is the same shape over a different one. Co-Authored-By: Claude Opus 5 <[email protected]>
Splitting it removes the credential entirely: the cron writes a report to the campaign path the Cirrus runner already mounts read-only, and the scheduled workflow reads that and reports with GITHUB_TOKEN. The cron never contacts GitHub, so it needs nothing to authenticate with. Also records where CIRRUS keeps the PAT its runners do require, since that is the posture the split matches rather than fights. Co-Authored-By: Claude Opus 5 <[email protected]>
The /glade/u/apps read-only mount is now requested as CCPP-502. The surrounding paragraph still asserted that the cron fallback needs a long-lived PAT, which the preceding commit established it does not. Co-Authored-By: Claude Opus 5 <[email protected]>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description of changes
This PR is just to test GitHub Workflows on the CIRRUS cloud runners.