Skip to content

[IGNORE] Add cirrus-testing workflow - #3389

Draft
samsrabin wants to merge 159 commits into
ESCOMP:masterfrom
samsrabin:cirrus-runner-workflows
Draft

samsrabin wants to merge 159 commits into
ESCOMP:masterfrom
samsrabin:cirrus-runner-workflows

Conversation

@samsrabin

Copy link
Copy Markdown
Member

Description of changes

This PR is just to test GitHub Workflows on the CIRRUS cloud runners.

@samsrabin samsrabin self-assigned this Aug 6, 2025
@samsrabin
samsrabin force-pushed the cirrus-runner-workflows branch from 00b95ee to 7114913 Compare August 6, 2025 21:21
@samsrabin

Copy link
Copy Markdown
Member Author

The ubuntu runner is currently failing because the CESMDATAROOT variable is undefined. Try giving it some garbage value; it might not actually be needed.

@samsrabin
samsrabin force-pushed the cirrus-runner-workflows branch from d4f93ae to ee33e2c Compare July 10, 2026 21:20
Troubleshooting list-glade-cesm-input: Add `sleep 600`.
list-glade-cesm-input: Delete troubleshooting `sleep`.
samsrabin and others added 30 commits October 6, 2026 13:02
Move every stale ARG to what config_machines.xml loads at
ccs_config_cesm1.0.88: GCC 14.3.0, netCDF-C 4.9.3, netCDF-Fortran 4.6.2,
HDF5 1.14.6, PnetCDF 1.14.1, ESMF 8.9.1, mpi-serial 2.5.3, PIO 2.6.8.
MPICH stays 3.4.3, since cray-mpich 8.1.32 is still MPICH-3.4-ABI-
derived; only its deviation guard moves. All ten checks now pass.

HDF5 needed more than a number. Upstream changed its tag scheme between
these releases, from hdf5-1_12_2 to hdf5_1.14.6, so the URL, the tarball
name and the extracted directory all change with it. Every other new
version's download URL and extracted directory name was checked against
upstream and needed nothing.

gnu_container.cmake's hardcoded ESMFMKFILE moves with ESMF_VERSION, as
the build-time assertion added earlier requires.

Not built, not run. The published image still predates this, so the
passing check describes the recipe rather than anything on GHCR; see
NEXT_STEPS item 6 for why this should not merge ahead of the rebuild.

Co-Authored-By: Claude Opus 5 <[email protected]>
build-on-casper.sh located its Dockerfile and build context next to
itself, via BASH_SOURCE. That holds when the script is run in place, but
PBS executes a copy from /var/spool/pbs/mom_priv/jobs, where nothing of
this repo sits, so a qsub submission died before the first Dockerfile
instruction with "stat .../jobs/Dockerfile: no such file or directory".

Fall back to PBS_O_WORKDIR, accepting either the repo root or this
directory, with CTSM_BUILD_CONTEXT as an override and a message naming
both places searched when none of them has the file.

Co-Authored-By: Claude Opus 5 <[email protected]>
ftp.gnu.org is unreachable from NCAR compute nodes. It times out rather
than refusing, so wget exhausts the patient retry budget set in wgetrc
and the step exits 4 before anything compiles. DNS resolves, so this is
routing, and ftpmirror.gnu.org is no help since it is the same host.

Try gcc.gnu.org's own release area first, then a kernel.org mirror, then
ftp.gnu.org. Per-attempt --tries and --timeout override the global
settings so an unreachable mirror is abandoned in seconds instead of
eating the job's walltime, and -O keeps a failed attempt from leaving a
partial file for the next one to trip over. The loop exits nonzero when
every mirror fails, so tar is never reached with an empty file.

Every other download host the Dockerfile uses was checked and is
reachable; gnu.org was the only one.

Verified by reachability checks and by exercising the loop's exit status
in both directions. Not built: the image has not been rebuilt.

Co-Authored-By: Claude Opus 5 <[email protected]>
netCDF-C 4.9.3 stopped folding $LIBS into the NC_LIBS it substitutes
into nc-config: configure.ac now sets NC_LIBS="$LDFLAGS $NC_LIBS" where
4.9.2 set "$LDFLAGS $NC_LIBS $LIBS". Plain --libs therefore returns a
bare -lnetcdf. Against a shared netCDF that is enough, because DT_NEEDED
pulls the rest in, but /usr/local/serial is static-only, so the
netcdf-fortran configure link test for nc_open failed and the step died
with "Could not link to netcdf C library".

Pass --libs --static there, which restores -lhdf5_hl -lhdf5 -lm -lz
-lxml2 -lcurl.

ESMF reaches the same prefix through the same broken idiom: it runs
nc-config --libs itself and keeps the -l options. It skips that query
when ESMF_NETCDF_LIBS is set in the environment, so the mpiuni ESMF step
now sets it. Only the serial ESMF is affected; the two MPI ones link the
shared netCDF under /usr/local.

Verified by instantiating both versions' real nc-config.in with the
values this build logged, and comparing what --libs and --libs --static
return. Not built.

Co-Authored-By: Claude Opus 5 <[email protected]>
podman's storage here is node-local, and node-local storage is wiped
when the allocation ends. The script built, tagged and exited, so a
batch build reported success and left nothing behind: the image went
with the node.

Save to /glade/work/$USER after a successful build, which set -e and
pipefail make the only way to reach that point. An existing tarball is
never clobbered; the name falls back to one with a timestamp. Skippable
with CTSM_BUILD_NO_SAVE=1, relocatable with CTSM_BUILD_SAVEDIR.

Co-Authored-By: Claude Opus 5 <[email protected]>
smoke-test.sh asserted six versions written out by hand, so the version
bump left it failing a correct image: gcc 14.3.0 against an expected
12.2.0. A second copy of the same numbers was always going to drift from
the ARGs, and it did so the first time they moved.

Take the expectations from the image's own labels, which the Dockerfile
sets from those ARGs. That removes the copy and makes the assertion
mean something it did not before: that the toolchain the build produced
is the one the ARGs asked for, so a cached or substituted layer is
caught rather than assumed away. A missing label fails loudly instead of
silently checking nothing.

mpichversion's whole output is matched rather than its first line, since
nothing guarantees the Version: line comes first.

Verified against a stub podman for the label lookup, the <no value>
fallback and the missing-label guard, and by checking the real
gcc 14.3.0 output against both the new and old expectations.

Co-Authored-By: Claude Opus 5 <[email protected]>
Both were labelled on the image but never asserted, so the two libraries
most likely to be picked up from the wrong place were the two nothing
checked.

HDF5 is read from h5dump, not h5cc: a parallel build installs h5pcc
instead, so the wrapper's name depends on configuration while the tools
are installed either way. The wrappers remain as fallbacks, and an
absent version fails rather than passing on an empty string.

ESMF is read from ESMF_VERSION_STRING in the esmf.mk that ESMFMKFILE
points at, which is the file CTSM consumes, rather than from a directory
name that happens to contain digits.

Verified against stubs: both sources, the h5pcc fallback, the failure
when no tool reports a version, and rejection of a wrong expectation.

Co-Authored-By: Claude Opus 5 <[email protected]>
ESMFMKFILE names the MPI debug ESMF. Every mpi-serial case and the whole
Fortran unit-test path instead link the mpiuni one, which the baked-in
CIME macro selects by a hardcoded path. That path was until now checked
only at build time, which says nothing about the image in hand: a
mis-tagged image or an older tarball satisfies an assertion that ran in
some other build.

Read the path back out of the macro in the image and require it to
exist, to be an ESMF_COMM=mpiuni build, and to report the labelled
version.

Verified against stubs for the healthy case and for the three ways it
goes wrong: ESMF_VERSION bumped without editing the macro, a macro
pointing at an MPI build, and a version that disagrees with the label.

Co-Authored-By: Claude Opus 5 <[email protected]>
The image was rebuilt on Casper against derecho's current stack and
passes all four validation scripts, including the Fortran unit tests and
a single-point system test. Item 6 now covers only publishing to GHCR
and repointing cirrus-testing.yml, which is also what still blocks a
merge: until then the repo describes a stack nothing in CI runs.

Also records that the GCC 12 to 14 jump, flagged as the risky part,
produced no compiler diagnostics at all. The two real breakages came
from elsewhere, HDF5's tag scheme and nc-config, and are documented at
the sites that hit them. The original reasoning is kept, since it
remains the right thing to check first on the next compiler jump.

Co-Authored-By: Claude Opus 5 <[email protected]>
Whether this image can be built in CI turns on disk more than on time:
a hosted runner has roughly 14 GB free, the image is 4.1 GB as a
tarball, and the build transiently needs GCC's build tree and three ESMF
trees on top of that. That number decides whether item 8 is a
straightforward workflow or needs the Dockerfile split into a base image
and a thin top layer, so it is worth knowing before any of it is
designed.

Probe both architectures, since the arm64 runner's availability is part
of the same question. Phase A reports what a runner has before and after
reclaiming the preinstalled toolchains and takes a couple of minutes;
Phase B, opt-in, attempts the real build and reports how far it gets.
MAKE_JOBS comes from nproc there, as the Dockerfile's default of 16
would thrash a runner this size.

Informational and workflow_dispatch only, following
probe-derecho-modules.yml: every step is best-effort and the log is the
result.

Not run: triggering it needs the branch pushed.

Co-Authored-By: Claude Opus 5 <[email protected]>
The build is going to live in a workflow whichever runner it lands on,
so the hosted runners' disk limit decides where it runs rather than
whether it runs at all. Add gha-runner-ctsm to the spike, since it is
the fallback and the same measurements decide it.

The reclaim step is skipped there: it is shared, persistent hardware,
and deleting its toolchains would not be cleanup. The build step picks
docker or podman by what exists, as the NCAR machine may have only
podman, and reports podman's graphroot, which is node-local on the
Casper machines and may be here too.

A job aimed at an offline or busy self-hosted runner queues instead of
failing, and timeout-minutes does not bound queue time, so that is
called out where someone would otherwise read a pending row as a hung
workflow.

Co-Authored-By: Claude Opus 5 <[email protected]>
The preceding commit added gha-runner-ctsm to the spike but left item 8
describing Phase A as covering only the two hosted architectures, and
said a tight disk result means the base-image split, which is no longer
the only alternative.

Co-Authored-By: Claude Opus 5 <[email protected]>
The probe pinned ncarenv/23.09 and its gcc, cray-mpich and netcdf-mpi,
so once derecho moved to ncarenv/25.10 it reported a module-load
failure that said nothing about the runner it exists to test.

Load the defaults instead. That cannot go stale, and it is what the
Phase 2 cron wants anyway: the question is what derecho has today, not
whether it still has what someone wrote down here. The readout now says
to compare against config_machines.xml, since a difference there is
precisely the drift Phase 2 is for.

Also notes that the hdf5 line is now a depends_on rather than a
version, HDF5 having become a separate module under ncarenv/25.10.

Not run: triggering it needs the branch pushed, and it only runs on the
self-hosted runner.

Co-Authored-By: Claude Opus 5 <[email protected]>
:20261007 is published, so point cirrus-testing.yml at it and close
item 6. Its config digest matches the one in the Casper build log, so
what is on GHCR is the artifact that passed validation rather than a
re-tag of something else.

:20260831 stays on GHCR as a fallback and is still what the September
validation record loads; those references are left alone, since they
describe runs that happened against that image.

Co-Authored-By: Claude Opus 5 <[email protected]>
GHCR associates a package with a repository through
org.opencontainers.image.source, which this image never set. The package
is therefore absent from the repository-filtered package list, which
reads as "the push did not work" even when it did; it is only findable
from the organization's unfiltered package page.

Setting the label fixes that for images built from here on. It cannot
be applied retroactively: :20261007 and earlier carry no source label,
so until the next rebuild the package stays linkable only by hand, from
the package settings page.

Co-Authored-By: Claude Opus 5 <[email protected]>
Loading derecho's stack from anywhere else does not work and cannot be
made to. ncarenv and gcc load fine off derecho's tree, but cray-mpich
dies -- the path /opt/cray/pe/mpich/8.1.32/ofi/gnu/12.3 does not exist
-- because Cray PE is on derecho hardware only, and netcdf-mpi sits
below cray-mpich in the hierarchy so it never becomes loadable either.
The previous probe would also have loaded the runner's own ncarenv when
that tree was on MODULEPATH, reporting a different machine's versions
as if they were derecho's.

None of it is needed. Every value the check wants is a file under
glade: the default version is the "default" symlink beside the
modulefiles, the hdf5 pairing is netcdf-mpi's own depends_on line, and
netCDF-Fortran comes from derecho's installed nf-config run by absolute
path -- nc-config and nf-config are generated shell scripts that echo
baked-in strings and load nothing. So the probe now asks whether glade
is mounted, resolves derecho's current defaults by reading symlinks,
and derives both [snapshot] values without loading a module. Lmod is
reported as context.

Resolving defaults rather than taking the highest version matters:
cray-mpich 9.0.0 is staged in the tree while default still points at
8.1.32.

Also correct a claim in the checker: derecho does have a pfunit module
now. CTSM still does not use it, so the PFUNIT_PATH reading is
unchanged.

Probe body run on a casper login node; every version it reports agrees
with the Dockerfile ARGs and the recorded snapshot. Not run in Actions.

Co-Authored-By: Claude Opus 5 <[email protected]>
Item 5 assumed the drift check had to run as a cron on an NCAR machine,
because it assumed the check had to load derecho's modules. It does not:
every value it wants is a file under glade, so it can run wherever glade
is mounted, and a GitHub workflow opening the drift issue needs only its
own GITHUB_TOKEN rather than a PAT parked on a shared machine.

Record where each value lives, why loading the stack is not an option
off derecho, why defaults must be resolved from the symlinks instead of
taken as the highest version, and the four things still undecided --
which runner, the default-branch-only rule for schedules, silent
queueing when a self-hosted runner is down, and whether to report
success at all.

Co-Authored-By: Claude Opus 5 <[email protected]>
Two decisions reshape what is left, both recorded with their evidence
in the new DRIFT_CHECK_DESIGN.md, which the items now point to instead
of carrying the argument inline.

The container's scope becomes gnu + mpi-serial only, so mpich leaves
the image along with PnetCDF, the MPI-linked netCDF/HDF5 and two of the
three ESMF flavors.

The drift check compares built executables rather than version
metadata. Metadata comparison is blind to drift that exists now: a
container case build links MPIserial_2.5.4 from CTSM's submodule while
derecho links its mpi-serial/2.5.3 module, and
check-derecho-versions.py calls that a match. Setting MPI_SERIAL_PATH
and PIO_LIBDIR in gnu_container.cmake makes the container use derecho's
mechanism rather than only matching its numbers.

Item 4 is unchecked and rewritten as 4a, the checker, and 4b, the
rebuild. Item 5 is rewritten around the above. Unchecked items are now
in the order they should be tackled, and each carries its own README
updates. The component list in "Where things stand" is corrected to the
versions the 2026-10-07 rebuild produced.

Also records the podman XDG_RUNTIME_DIR constraint, which bites
interactive podman but not build-on-casper.sh.

Documentation only; nothing built or run.

Co-Authored-By: Claude Opus 5 <[email protected]>
VALIDATION_2026-09.md was reachable only from item 3 of NEXT_STEPS, so a
reader starting from the design doc would not find it and might redo the
wrapper validation or revisit decisions it already settled.

Reference it twice: in the context, as settled work not to reopen, and
in the verification section, where phase 4 re-runs its procedure against
the rebuilt image. The second reference records why its section ordering
matters, since that is the part most likely to be skipped.

Documentation only.

Co-Authored-By: Claude Opus 5 <[email protected]>
The container's scope is gnu + mpi-serial only, so the check now reads
the mpilib="mpi-serial" blocks of config_machines.xml. Gone: the
MPI-flavor checks, the cray-mpich guard, the <MPILIBS> assertion and
the two-stack twin machinery. cray-libsci is guarded instead; its
stand-in carries no version ARG, so deviation mode no longer requires
one. MPICH_VERSION and PNETCDF_VERSION stay in the Dockerfile,
unchecked, until the image is stripped to serial-only.

[snapshot] is re-measured against the serial netcdf module and the
hdf5 it pulls in. The numbers did not move, only the provenance;
measured on Casper 2026-10-09, with the sources recorded in the ini.

New: an existence check resolving every pinned module in derecho's
module tree under /glade, because config_machines.xml is CTSM's claim
about derecho and nothing else verifies it. That needs /glade, so the
workflow now picks its runner per event and the hosted PR gate passes
--skip-module-tree. A stdlib unittest file runs as its own CI step.

Verified on Casper: the tests pass and the checker passes every check
including the live module-tree read. The workflow itself has not been
run.

Co-Authored-By: Claude Opus 5 <[email protected]>
workflow_dispatch is unavailable until the file reaches the default
branch, which leaves the probe unrunnable on the branch that
introduced it. pull_request carries no such restriction, and a fork's
PR runs in the base repository's context, where gha-runner-ctsm is
reachable.

The paths filter names the probe itself, so the commit adding the
trigger fires it and later PRs do not.

Not run; running it is the point.

Co-Authored-By: Claude Opus 5 <[email protected]>
The probe exited at its first check, so a negative answer carried no
detail. gha-runner-ctsm is a Kubernetes pod, so glade arrives as
specific bind mounts rather than as a filesystem; which ones are
present decides whether the fix is one more mount or a different host
for the check.

Not run.

Co-Authored-By: Claude Opus 5 <[email protected]>
The probe has now been run. The runner is a Kubernetes pod, so glade
arrives as specific bind mounts; the inputdata path is among them and
the module tree is not. That decides nothing about the drift check's
design, but it rules out the host item 5 assumed and makes the
scheduled half of derecho-version-check.yml unrunnable until either a
mount is added or the check moves to derecho.

Co-Authored-By: Claude Opus 5 <[email protected]>
The second probe run named it: a single read-only NFSv4 export of
the campaign cesmdata path, with campaign the only entry under
/glade. Splitting the ask in two matters because the modulefile trees
are text and the Spack prefix is not, and only netCDF-Fortran's
version needs the latter.

Co-Authored-By: Claude Opus 5 <[email protected]>
The concretized environment file gives netCDF-Fortran, netCDF-C and
HDF5 in one read, and the token in a module's install prefix is a
unique prefix of that spec's dag hash, so the versions tie to the
module the serial path loads rather than to whatever the environment
happens to hold once. It also removes the Spack install tree from
what the Cirrus runner would need mounted, leaving three text paths.

Verified by resolving the gnu serial netcdf/4.9.3 prefix token
against spack.lock; the instrument is not yet wired into any check.

Co-Authored-By: Claude Opus 5 <[email protected]>
spack.lock lives under an ncarenv version directory, so naming it in
the request would tie the mount to 25.10 and need re-requesting on
every bump -- an event this check exists to notice. The parent
directories carry no version.

Also records the filesystem the request concerns, since the existing
mount is the same shape over a different one.

Co-Authored-By: Claude Opus 5 <[email protected]>
Splitting it removes the credential entirely: the cron writes a report
to the campaign path the Cirrus runner already mounts read-only, and
the scheduled workflow reads that and reports with GITHUB_TOKEN. The
cron never contacts GitHub, so it needs nothing to authenticate with.

Also records where CIRRUS keeps the PAT its runners do require, since
that is the posture the split matches rather than fights.

Co-Authored-By: Claude Opus 5 <[email protected]>
The /glade/u/apps read-only mount is now requested as CCPP-502. The
surrounding paragraph still asserted that the cron fallback needs a
long-lived PAT, which the preceding commit established it does not.

Co-Authored-By: Claude Opus 5 <[email protected]>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant