CI and the gate
For contributors, and for anyone looking at a green check mark on a pull request and wondering what it proves. The short answer is: less than you would assume. The split is deliberate, and it is easy to misread in the dangerous direction.
Where CI runs
CI runs on GitHub Actions, the project’s primary host, from the workflows in
.github/workflows/ on GitHub-hosted Linux x86_64 runners: pr-gate.yml
(per-PR into development), release-gate.yml (release PRs into main and a daily
development-tip run), advisories.yml, docs-pages.yml, the tag-verify / manual-publish
release.yml, and the manual native-arm64 OpenModelica evidence workflows. Actions in the gating
workflows and docs-pages.yml are pinned to full commit SHAs, which check-workflow-gates.sh
enforces for the gating workflows. release.yml and the OpenModelica evidence workflows are
byte-bound by their own approval and evidence manifests, so they keep tag references until their
next reviewed refresh.
The ci.yml in that directory is not the CI: it is the GitHub-era gate that produced the
retained strict-bit evidence, stays byte-identical because it is a bound
source of that evidence, and is disabled in the repository’s Actions settings. (CI briefly ran on
a self-hosted Forgejo instance, now retired.)
Each gating workflow ends in a CI OK job that needs every other job and fails unless each one
succeeded; check-workflow-gates.sh asserts that its needs: lists every other job in the file
and that no gating job carries a job-level if:. Branch protection requires exactly that one
status: the CI OK check from pr-gate.yml on development and from release-gate.yml on
main. Both gates also run on pushes to their branch, so the merged commit is re-checked.
GitHub PR runs test a synthetic merge of the PR head into its base. Branch protection should also require a PR branch to be up to date with its base before merging, so the tested merge is what lands.
Docs-site validation (.github/workflows/docs-pages.yml) is path-filtered to docs/**,
README.md, scripts/docs/**, scripts/authority_claims/**, site/** and its own workflow
file. It reports its own build docs artifact status and is not part of CI OK, so a PR that
does not touch those paths shows no docs check at all. A push to main touching those paths also
deploys the site to GitHub Pages.
One command, one source of truth
.agents/gate.sh is the only place the gate’s command list is written down.
Every other document in this repo — including this page — points at it rather than restating it,
because nine divergent prose copies existed before the script was written and two of them were
materially weaker than CI (see the script’s header). There are two invocations:
bash .agents/gate.sh # light — mirrors the per-PR gate
bash .agents/gate.sh full # full — adds the workspace suite and doctests
CI does not merely mirror that script, it executes it: the gate (light) job at
.github/workflows/pr-gate.yml and gate (full) at
.github/workflows/release-gate.yml.
So every command in the script gates a pull request whether or not pr-gate.yml also runs it as its
own job. Read that as coverage, not as parity, and note that the implication does not run the other way:
gate (light) is bash .agents/gate.sh plus any steps of its own. The Quickstart-executes step
was exactly that for a while — a required check no local run of the script performed — and an
earlier revision of this paragraph used a numeric citation that stopped one line short of it.
Nothing verifies mechanically that the two files still list the same commands. That check was
attempted and withdrawn, and the gate job’s header in the dormant .github/workflows/ci.yml
records why —
every design either compared argv strings that RUSTFLAGS=--cap-lints=allow leaves byte-identical
while neutering clippy, or reimplemented enough of the workflow if:/needs:/matrix semantics to
become its own untested gate.
The script’s steps group into: formatting, file-size and secret hygiene; the repository-invariant
gates (the default build links no database or async runtime; the golden generator cannot bless its
own output as the oracle; package, feature, and publication selection is closed); behavior fixtures
for those gates, because a gate that cannot fail is not a gate; build, clippy and rustdoc under
-D warnings; supply-chain checks; the determinism subset; and two fixture input-hygiene audits. A
failing step never aborts the run, so one round trip reports every problem instead of the first
(see the script’s step function).
The authority index and generated projection add a fast, bounded consistency
check and hostile controls. Native numeric observers run inside the existing oce-api/oce-blocks
subset; package/public/catalog validators retain ownership. The full gate also executes those owners.
This does not check arbitrary Markdown claims or workflow parity, and regeneration is never gated in.
Dev-light, release-heavy
A green PR is not evidence that the change’s own tests pass.
The per-PR gate into development runs the state-determinism subset for oce-api, oce-blocks,
and oce-expr. That is the determinism-matrix job in .github/workflows/pr-gate.yml: it runs
that three-crate subset natively on x86_64 and, cross-compiled, on aarch64 under QEMU user-mode
emulation, each twice — once under debug codegen, once under release codegen. Each architecture
emits populated revision-2 portable and target-bound state vectors. The job compares both across
codegen profiles, requires the portable files to match and the target-bound files to differ across
architectures, then parses and refuses the aarch64 target-bound bytes on x86_64. The emulated leg
excludes one test that re-executes its own binary (an OCE_BLESS truthiness probe, not a
determinism test), which the native leg still runs. Emulation keeps the cross-architecture
comparison on every PR with a single x86_64 runner; it is CI signal, not native aarch64 evidence.
The gate script runs the test commands locally and adds two named
oce-cxf test binaries, which are input hygiene rather than
engine coverage: the port-order audit sweeps 47 CXF documents, of which 46 are Guideline 36 catalog
fixtures and one is a resolver contract; the structural oracle compares the catalog fixtures it can
pair with vendored modelica-json translations
(see the gate script’s fixture input-hygiene section). That oracle compares document structure — instances and undirected
edges — not simulated behavior.
The scoped oce-conformance strict-bit subset also runs per-PR: strict_bits plus the four
affected per-block suite binaries, in Linux x86_64 (native) / aarch64 (emulated) × debug/release,
with two independent captures per cell and a fail-closed cross-cell comparison. The retained evidence
covers exact comparison of 21 pinned Real cases on qualified Linux; unqualified targets retain
the unchanged 1e-12 aligned band. This is not the whole conformance suite or a libm accuracy claim.
The remainder waits for the release/full gate. A change outside the named test subsets can show
a fully green PR having executed none of its own tests.
Before claiming tests pass, run bash .agents/gate.sh full first-hand and read the tail.
Every pull request runs the gate
The per-PR gate runs on every PR, drafts included; there is no draft carve-out, because a gating
job with a job-level if: would read as a failure in CI OK. A PR with no checks still looks a lot
like a PR with no failing checks — confirm CI OK actually reported.
cargo-deny is not skippable, but advisories do not gate a PR
The standalone cargo-deny job in pr-gate.yml runs cargo-deny’s bans, licenses and sources checks
on every PR, and the gate script runs them too (see the script’s cargo-deny step).
advisories is a different story, and the carve-out belongs next to the claim. It is deliberately
excluded from the script — it needs network access and a writable advisory database, neither of
which a sandboxed lane has. It runs daily in advisories.yml and on
release PRs (the release-gate.yml cargo-deny job). advisories.yml has no pull_request trigger at all, so
a PR into development that introduces a dependency with a known RustSec advisory merges green and
is caught by the next scheduled run, not by its own gate.
What the release gate adds
release-gate.yml fires on development → main PRs, on pushes to main, on manual dispatch,
and on a daily cron against the development tip (see its trigger block). It is disjoint from
pr-gate.yml by base
branch, so the two never both fire on one PR. It re-runs the light correctness gates against the
release tip and adds four things:
| Step | What it covers | Where |
|---|---|---|
| workspace nextest | every unit and integration test in the workspace | release-gate.yml, test-suite job, unit + integration step |
| workspace nextest, release codegen | release panic-freedom, debug_assert paths stripped; inherited ci-release runner policy | release-gate.yml, test-suite job, release step |
cargo test --doc | doctests — nextest cannot run them, so this is a separate step | release-gate.yml, test-suite job, doctest step |
two cargo public-api surface gates | exact public API text for oce-api and oce-store | release-gate.yml, test-suite job, per-crate surface steps |
--no-tests=fail is explicit on the nextest steps: a run that discovers zero tests hard-fails
rather than passing, which catches tests that silently stop compiling or being found.
Nextest policy and reports
Local setup and CI pin cargo-nextest 0.9.143; .config/nextest.toml also declares that version as
both required and recommended, so an older local binary exits before testing. The default profile
is fail-fast. Automated debug runs use ci; release-codegen runs use ci-release, which inherits
the same retries, timeout, leak, and reporter policy instead of copying it. The two public-API runs
inherit that policy through separate child profiles because their nested nightly builds need a
longer per-test timeout and separate reports.
Retries are zero and a flaky pass is still a failure. Ordinary tests terminate after 120 seconds;
the public-API surface tests allow 10 minutes for their nested nightly rustdoc builds. A run stops
after 15 minutes, and a child process retaining inherited output handles for more than two seconds
fails as a leak. CI writes Jenkins-compatible JUnit XML to target/nextest/<profile>/junit.xml and
requires every expected report to exist and be non-empty; the reports are no longer uploaded as
artifacts. The emulated aarch64 legs use scripts/ci/nextest-emulated.toml instead of
.config/nextest.toml: the same no-retry, flaky-is-failure policy, with wider time limits because
QEMU runs test binaries several times slower.
Partitioning and build archives are deliberately off: the full test execution takes seconds while compilation dominates, and each determinism leg must execute the complete selected set under its own architecture and codegen mode. Experimental record/replay is also off in CI; enabling a feature that nextest still marks unstable would make the gate depend on a non-stable format. Test groups and thread reservations remain available when measurement identifies a shared resource or heavy test; none is known today.
The public-api baselines are the strongest stability evidence in this repo. They are checked-in
text files — crates/oce-api/tests/public-api.txt (1357 lines) and
crates/oce-store/tests/public-api.txt (1230 lines) — and the tests at
crates/oce-api/tests/public_api.rs and crates/oce-store/tests/public_api.rs diff the crate’s
real surface against them, so any unintended addition, removal or signature change fails the gate
rather than shipping. Two env vars interlock to keep the gate honest: OCE_PUBLIC_API_NIGHTLY arms
it and names the pinned nightly to shell out to, and OCE_REQUIRE_SURFACE_CHECK=1 turns a missing
nightly into a hard panic instead of a silent skip, so disarming the gate turns it red, never green
(see the surface steps’ arming environment). The two crates run as separate steps on purpose: merging the package
selectors would let one surviving crate hide the other’s vanished test.
The exact rows are classified without replacing these signature baselines by the public surface contract and its machine-checked ledger.
What CI cannot observe
- Operating systems other than Linux. Every gating workflow —
pr-gate.yml,release-gate.yml,advisories.yml, anddocs-pages.yml(per-PR ondocs/**,README.md,scripts/docs/**,scripts/authority_claims/**, andsite/**) — runs onubuntu-latest, a Linux x86_64 runner. Cross-architecture is covered for the determinism and strict-bit subsets — x86_64 native and aarch64 emulated, debug and release. macOS and Windows are not built or tested anywhere. - Native aarch64 execution. The per-PR gate runs its aarch64 legs under QEMU user-mode
emulation on x86_64. The retained native qualification in
qualified-linux/is still verified on every PR, but a new native aarch64 capture needs a native arm64 runner (ubuntu-24.04-arm, today used only by the manual OpenModelica evidence workflows). - Anything derived from git history. No workflow sets
fetch-depth, soactions/checkouttakes its default of a single commit. A check that needs history cannot run in CI. The visible consequence: golden provenance records bind to a content digest of the checked-in bytes rather than to the engine revision that produced them (crates/oce-cxf/tests/golden_provenance/mod.rs:3-5). - Line-ending behavior.
.gitattributes:1pins* text=auto eol=lf, but an ubuntu-only CI never performs a CRLF checkout, so that normalization is asserted by git configuration and exercised by no test. Goldens here are compared bit-exactly, which is precisely where a stray\rwould show up.
The script says the rest itself, in its closing report: a green local run
does not prove the cross-arch determinism matrix passes (one machine cannot reproduce it), does
not prove the two cargo public-api surface gates pass (they need the gate-only nightly), does not
prove cargo deny check advisories passes, and does not prove that the script and pr-gate.yml
still agree. (The script’s own closing report still names ubuntu-24.04-arm and ci.yml; the script
is a bound source of the retained strict-bit evidence, so its text changes only with an evidence
refresh.) A light run additionally does not prove the workspace suite or doctests pass, because the
per-PR gate does not run them.
Publishing
release.yml publishes with the crates.io token held in its release GitHub Environment. It is
decoupled from both gates and from
each other’s triggers. Pushing a v* tag runs
verify only — tag/version match, fmt, clippy, a workspace cargo test, and a full
cargo publish --dry-run — with no token and no publish, so a tag can be re-cut safely
(see release.yml, verify job). Publishing is a separate manual workflow_dispatch into the release GitHub
Environment. Cargo’s workspace selection includes the 12 publishable members and skips the five
members with publish = false; the exact split and feature closure are guarded by the
package, feature, and publication policy. No crate is on crates.io
yet, and actual publication remains deferred pending explicit owner authorization.
Related: host-responsibilities.md for what the engine deliberately
leaves to the embedder, and ../TESTING.md for the testing standard a change is
expected to meet.