Alert when a Semaphore template or Gitea workflow run fails
Every 15 minutes (on the existing health-check timer) the app reads the latest run of each Semaphore template and each Gitea repo's latest workflow run. A failed one raises one notification, and another when a later run succeeds. It is state-based like the server health alerts, so a job that fails every night alerts on the first failure, not every night. Gitea alerts include the run's link. Toggle: Settings > Notifications. What counts: - Semaphore "error" is a failure, "success" is a pass. A run that is waiting, running, stopped by hand or rejected is neither, so it leaves the previous state alone: a run in progress must not clear a failure it hasn't fixed yet, and a manual stop isn't a failure. - Gitea failure/success likewise; running, waiting, blocked, cancelled and skipped leave things as they were. - A failing template that gets another failing run does not re-alert. Not mistaking "couldn't read" for "fixed": - Semaphore's template listing swallowed per-project errors, so a project that failed to load looked like a project with no templates. A new checkTemplates adapter method reports which projects failed, and their failures are held rather than cleared. - Gitea reports a run it couldn't fetch as null, the same as "no runs"; both leave the repo's state alone. - An unreachable integration holds all of its failures. Nothing is cleared or re-announced while it is down. The first pass only records what is already failing without announcing it, so upgrading (or adding an integration to a fresh install) doesn't produce a wall of alerts about months-old failures. That baseline is not spent while nothing could be read. Maintenance windows on a Semaphore or Gitea integration silence its failure alerts with the same rules as the health alerts: a problem that starts during a window alerts when it ends, and one already announced stays known. The diff logic is reused from the health monitor rather than copied. Maintenance page text updated. Known limit: for Gitea this follows the repo's most recent run on any workflow or branch, matching what the Gitea page shows; a failure in one workflow can be masked by a later success of another. Verified with 53 checks against fake Semaphore and Gitea servers and a webhook receiver: classification, baseline (including not being consumed when nothing is readable), single alert per failure, no repeat, in-progress/ stopped/cancelled runs, recovery and re-failure, unreadable project, unreadable integration, run-fetch errors, maintenance windows (silenced, then announced after), the toggle, disabled integrations and repos without Actions. Real dev database mtime untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
9f1609c4ed
commit
4118062405
11 files changed
+219
-5
No files matched your search
@@ -148,6 +148,19 @@ which can be filtered by one or several tags (the filter is in the URL, so a
|
||||
tag on a server's page links to everything sharing it), and they're searchable
|
||||
from the global search box.
|
||||
|
||||
**Automation failures** — every 15 minutes the app looks at the latest run of
|
||||
each Semaphore template and each Gitea repo's latest workflow run. A failed one
|
||||
raises a single notification (with the project/template or repo, run number, and
|
||||
for Gitea the run's link), and another when a later run succeeds. It's
|
||||
state-based, so a job that fails every night alerts on the first failure rather
|
||||
than every night. A run that's still going, was cancelled, or was stopped by
|
||||
hand leaves things as they were, and anything that couldn't be read (a Semaphore
|
||||
project or Gitea repo that errored, or an integration that's down) is neither
|
||||
cleared nor re-announced. The first check after upgrading only records what's
|
||||
already failing, so old failures aren't announced. Toggle it under Settings →
|
||||
Notifications. For Gitea this follows the repo's most recent run on any
|
||||
workflow or branch, the same as the Gitea page shows.
|
||||
|
||||
**Maintenance mode** silences alerts about one server, integration, or DNS
|
||||
provider while you work on it (server offline / disk, storage and Synology
|
||||
health, Proxmox backup alerts, and "integration down" for that service type).
|
||||
|
||||
Reference in new issue
Block a user