feat(observability): alert on Forgejo CI job outcomes, not just runner supply #555

Merged
ryangr0 merged 1 commit from ryangr0/forgejo-ci-outcome-alerts into main 2026-08-09 08:52:43 +00:00 AGit
Owner

platform-forgejo-ci watches the runner fleet: scaler errors, queue depth,
pending pods, warm floor. Every one of those is about SUPPLY. None of them can
see a job that fails instantly on a healthy fleet.

That is not hypothetical. ploeg's release job failed on every push for nine
days (2026-07-31 to 08-09) — a composite action pinned at a SHA that had been
amended off its branch, so the runner could not resolve it and the job died
before its first step. Seven runs. checks green on all of them, no queue, no
pending pods, nothing to fire on. It was found by hand, and only because
someone went looking for why no rc had been cut.

Adds the missing half:

  • forgejo-ci-exporter: stdlib-only Python in a ConfigMap on a stock python
    image, polling the Actions API on a 300s timer and exporting, per trunk
    pipeline, the last run's status, the consecutive-failure streak, and the
    last-success timestamp. No image to build, publish, scan or renovate, and
    nothing to install on a read-only rootfs. The ploeg work-items exporter next
    door takes the same stock-image shape.
  • Two alerts in a platform.forgejo-ci-outcomes group beside the supply ones.

Three deliberate choices, each forced by what the live data showed:

pull_request runs are excluded. A red PR is already visible to whoever opened
it; the silent class is what lands on trunk. Refs are bounded to the default
branch plus main/development, and dropped entirely for release/schedule events,
because a PR ref is #36 and a release ref is a tag — both unbounded.

EXCLUDE_PIPELINES mutes webgrip/workflows' reusable (workflow_call-only)
files. Forgejo schedules them on push and release anyway and they fail
instantly: 74 permanently-red series from one repo, which would have made this
alert storm on its first day. The mute is repo-scoped, not a global workflow
pattern that would also hide a legitimately-named workflow elsewhere, and the
count is exported as forgejo_ci_pipelines_excluded so it stays visible. That
repo's real on_*.yml pipelines still report, and the underlying Forgejo
behaviour is worth fixing at source.

The alert requires the pipeline to have run in the last 7d. Measured against
the live org, four trunk pipelines were red and only ONE had run inside a week;
the rest were branches nobody had pushed in 2-3 weeks, including ploeg's own
main (red since 2026-07-26). Without that guard every abandoned red branch
alerts forever, and an alert that always fires is one nobody reads — which is
how the outage this exists to catch stayed invisible. Anything failing often
enough to matter runs often enough to clear it; ploeg's failed every day or two.

Verified before proposing. The exporter ran against the live instance
(26 repos, 94 series, 7.6s refresh, zero errors) and again as the built
ConfigMap with the Deployment's env, to prove the YAML round-trip. promtool
checks both rules, and a promtool unit test over the logic confirms an
actively-failing pipeline fires, a dormant one does not, a 2-failure blip does
not, and exporter-down fires. Neither this repo's CI nor local tooling
validates PromQL today; wiring promtool test rules into e2e is the follow-up
this makes worth doing.

The exporter runs anonymously and so sees public repos only. FORGEJO_TOKEN is
referenced with optional:true, so provisioning a forgejo-ci-exporter-token
Secret (key: token) later starts private repos reporting with no manifest
change.

Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com

platform-forgejo-ci watches the runner fleet: scaler errors, queue depth, pending pods, warm floor. Every one of those is about SUPPLY. None of them can see a job that fails instantly on a healthy fleet. That is not hypothetical. ploeg's release job failed on every push for nine days (2026-07-31 to 08-09) — a composite action pinned at a SHA that had been amended off its branch, so the runner could not resolve it and the job died before its first step. Seven runs. `checks` green on all of them, no queue, no pending pods, nothing to fire on. It was found by hand, and only because someone went looking for why no rc had been cut. Adds the missing half: - forgejo-ci-exporter: stdlib-only Python in a ConfigMap on a stock python image, polling the Actions API on a 300s timer and exporting, per trunk pipeline, the last run's status, the consecutive-failure streak, and the last-success timestamp. No image to build, publish, scan or renovate, and nothing to install on a read-only rootfs. The ploeg work-items exporter next door takes the same stock-image shape. - Two alerts in a platform.forgejo-ci-outcomes group beside the supply ones. Three deliberate choices, each forced by what the live data showed: pull_request runs are excluded. A red PR is already visible to whoever opened it; the silent class is what lands on trunk. Refs are bounded to the default branch plus main/development, and dropped entirely for release/schedule events, because a PR ref is `#36` and a release ref is a tag — both unbounded. EXCLUDE_PIPELINES mutes webgrip/workflows' reusable (`workflow_call`-only) files. Forgejo schedules them on push and release anyway and they fail instantly: 74 permanently-red series from one repo, which would have made this alert storm on its first day. The mute is repo-scoped, not a global workflow pattern that would also hide a legitimately-named workflow elsewhere, and the count is exported as forgejo_ci_pipelines_excluded so it stays visible. That repo's real on_*.yml pipelines still report, and the underlying Forgejo behaviour is worth fixing at source. The alert requires the pipeline to have run in the last 7d. Measured against the live org, four trunk pipelines were red and only ONE had run inside a week; the rest were branches nobody had pushed in 2-3 weeks, including ploeg's own main (red since 2026-07-26). Without that guard every abandoned red branch alerts forever, and an alert that always fires is one nobody reads — which is how the outage this exists to catch stayed invisible. Anything failing often enough to matter runs often enough to clear it; ploeg's failed every day or two. Verified before proposing. The exporter ran against the live instance (26 repos, 94 series, 7.6s refresh, zero errors) and again as the built ConfigMap with the Deployment's env, to prove the YAML round-trip. promtool checks both rules, and a promtool unit test over the logic confirms an actively-failing pipeline fires, a dormant one does not, a 2-failure blip does not, and exporter-down fires. Neither this repo's CI nor local tooling validates PromQL today; wiring `promtool test rules` into e2e is the follow-up this makes worth doing. The exporter runs anonymously and so sees public repos only. FORGEJO_TOKEN is referenced with optional:true, so provisioning a `forgejo-ci-exporter-token` Secret (key: token) later starts private repos reporting with no manifest change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ryangr0 force-pushed ryangr0/forgejo-ci-outcome-alerts from d979b7090a
Some checks failed
e2e / Lint & static validation (pull_request) Failing after 28s
e2e / Validate Renovate config (pull_request) Successful in 4m12s
e2e / Kyverno Chainsaw (KinD) (pull_request) Successful in 48s
e2e / Flux-local render (pull_request) Has been cancelled
to 0c28837dd6
All checks were successful
e2e / Kyverno Chainsaw (KinD) (pull_request) Successful in 11s
e2e / Flux-local render (pull_request) Successful in 2m58s
e2e / Validate Renovate config (pull_request) Successful in 4m31s
e2e / Lint & static validation (pull_request) Successful in 4m31s
2026-08-09 08:26:04 +00:00
Compare
Sign in to join this conversation.
No reviewers
No labels
pull-request
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
webgrip/homelab-cluster!555
No description provided.