feat(observability): alert on Forgejo CI job outcomes, not just runner supply #555
No reviewers
Labels
No labels
pull-request
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
webgrip/homelab-cluster!555
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "ryangr0/forgejo-ci-outcome-alerts"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
platform-forgejo-ci watches the runner fleet: scaler errors, queue depth,
pending pods, warm floor. Every one of those is about SUPPLY. None of them can
see a job that fails instantly on a healthy fleet.
That is not hypothetical. ploeg's release job failed on every push for nine
days (2026-07-31 to 08-09) — a composite action pinned at a SHA that had been
amended off its branch, so the runner could not resolve it and the job died
before its first step. Seven runs.
checksgreen on all of them, no queue, nopending pods, nothing to fire on. It was found by hand, and only because
someone went looking for why no rc had been cut.
Adds the missing half:
image, polling the Actions API on a 300s timer and exporting, per trunk
pipeline, the last run's status, the consecutive-failure streak, and the
last-success timestamp. No image to build, publish, scan or renovate, and
nothing to install on a read-only rootfs. The ploeg work-items exporter next
door takes the same stock-image shape.
Three deliberate choices, each forced by what the live data showed:
pull_request runs are excluded. A red PR is already visible to whoever opened
it; the silent class is what lands on trunk. Refs are bounded to the default
branch plus main/development, and dropped entirely for release/schedule events,
because a PR ref is
#36and a release ref is a tag — both unbounded.EXCLUDE_PIPELINES mutes webgrip/workflows' reusable (
workflow_call-only)files. Forgejo schedules them on push and release anyway and they fail
instantly: 74 permanently-red series from one repo, which would have made this
alert storm on its first day. The mute is repo-scoped, not a global workflow
pattern that would also hide a legitimately-named workflow elsewhere, and the
count is exported as forgejo_ci_pipelines_excluded so it stays visible. That
repo's real on_*.yml pipelines still report, and the underlying Forgejo
behaviour is worth fixing at source.
The alert requires the pipeline to have run in the last 7d. Measured against
the live org, four trunk pipelines were red and only ONE had run inside a week;
the rest were branches nobody had pushed in 2-3 weeks, including ploeg's own
main (red since 2026-07-26). Without that guard every abandoned red branch
alerts forever, and an alert that always fires is one nobody reads — which is
how the outage this exists to catch stayed invisible. Anything failing often
enough to matter runs often enough to clear it; ploeg's failed every day or two.
Verified before proposing. The exporter ran against the live instance
(26 repos, 94 series, 7.6s refresh, zero errors) and again as the built
ConfigMap with the Deployment's env, to prove the YAML round-trip. promtool
checks both rules, and a promtool unit test over the logic confirms an
actively-failing pipeline fires, a dormant one does not, a 2-failure blip does
not, and exporter-down fires. Neither this repo's CI nor local tooling
validates PromQL today; wiring
promtool test rulesinto e2e is the follow-upthis makes worth doing.
The exporter runs anonymously and so sees public repos only. FORGEJO_TOKEN is
referenced with optional:true, so provisioning a
forgejo-ci-exporter-tokenSecret (key: token) later starts private repos reporting with no manifest
change.
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
d979b7090a0c28837dd6