Commit Graph
2 Commits
Author SHA1 Message Date
bobbanandClaude Sonnet 5 4118062405 Alert when a Semaphore template or Gitea workflow run fails
Every 15 minutes (on the existing health-check timer) the app reads the
latest run of each Semaphore template and each Gitea repo's latest
workflow run. A failed one raises one notification, and another when a
later run succeeds. It is state-based like the server health alerts, so a
job that fails every night alerts on the first failure, not every night.
Gitea alerts include the run's link. Toggle: Settings > Notifications.

What counts:
- Semaphore "error" is a failure, "success" is a pass. A run that is
  waiting, running, stopped by hand or rejected is neither, so it leaves
  the previous state alone: a run in progress must not clear a failure it
  hasn't fixed yet, and a manual stop isn't a failure.
- Gitea failure/success likewise; running, waiting, blocked, cancelled and
  skipped leave things as they were.
- A failing template that gets another failing run does not re-alert.

Not mistaking "couldn't read" for "fixed":
- Semaphore's template listing swallowed per-project errors, so a project
  that failed to load looked like a project with no templates. A new
  checkTemplates adapter method reports which projects failed, and their
  failures are held rather than cleared.
- Gitea reports a run it couldn't fetch as null, the same as "no runs";
  both leave the repo's state alone.
- An unreachable integration holds all of its failures. Nothing is cleared
  or re-announced while it is down.

The first pass only records what is already failing without announcing
it, so upgrading (or adding an integration to a fresh install) doesn't
produce a wall of alerts about months-old failures. That baseline is not
spent while nothing could be read.

Maintenance windows on a Semaphore or Gitea integration silence its
failure alerts with the same rules as the health alerts: a problem that
starts during a window alerts when it ends, and one already announced
stays known. The diff logic is reused from the health monitor rather than
copied. Maintenance page text updated.

Known limit: for Gitea this follows the repo's most recent run on any
workflow or branch, matching what the Gitea page shows; a failure in one
workflow can be masked by a later success of another.

Verified with 53 checks against fake Semaphore and Gitea servers and a
webhook receiver: classification, baseline (including not being consumed
when nothing is readable), single alert per failure, no repeat, in-progress/
stopped/cancelled runs, recovery and re-failure, unreadable project,
unreadable integration, run-fetch errors, maintenance windows (silenced,
then announced after), the toggle, disabled integrations and repos without
Actions. Real dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 03:11:59 +02:00
bobbanandClaude Sonnet 5 aae4f0d74f Add maintenance mode to silence alerts while working on a server or integration
Rebooting Proxmox or patching a server triggered failure/offline alerts
you then had to dismiss. A maintenance window silences alerts about one
server, integration, or DNS provider for a chosen time. New Maintenance
page (start with a duration and optional reason, end early, see what's
silenced and what isn't) and a banner in the app shell so every signed-in
user can see what is currently silenced. Starting/ending is operator-only
and audit-logged; starting one on a target that already has a window
restarts its clock instead of stacking.

Silenced for the target: server offline/disk alerts, Proxmox/Synology
storage and health alerts, Proxmox backup alerts, and "integration down"
alerts. Not silenced: expiry and update reminders, DNS change notices.

The design goal is that this cannot hide a real outage:
- Every window has a required end (5 min to 7 days); there is no
  open-ended option, so a forgotten window expires by itself.
- A silenced problem is deliberately NOT recorded as "known". If it is
  still present when the window ends it alerts then, as new. A problem
  that was already alerted before the window stays known, so it isn't
  repeated, and is reported cleared only after the window ends.
- Failure alerts keep counting failures during a window without marking
  themselves alerted, so an outage that outlasts the window alerts on the
  very next failed call.

Known limitation, stated on the page: integration-failure alerts are
tracked per service TYPE (all "proxmox"), not per configured instance, so
a window on one Proxmox integration also silences a failure on a second
Proxmox integration while it's open. Fixing that means threading the
integration id through every adapter and the diagnostic log, which is a
much larger change than this feature.

Also moved the API-error-message helper out of Secrets.tsx into a shared
util now that two pages use it. New table maintenance_windows (migration
0008).

Verified with 44 checks: the condition-key-to-subject mapping (including
server:3 vs server:33), the diff rules with silenced subjects (new problem
not recorded, alerts when the window ends; already-known one carried and
not repeated; clears only after the window), window expiry and
integration/DNS-provider source matching, the failure tracker end to end
against a webhook (silent during a window while an unrelated service still
alerts; outage that outlasts the window alerts on the next failure and
only once; fail-and-recover fully inside a window sends nothing), a full
health pass against a real window, and the real router with a stubbed
session (role rules, duration bounds including the missing-duration case,
extend-not-stack, 404s, deleted targets hidden, audit entries). Real dev
database mtime untouched.

Not done: I haven't clicked through the new page or banner in a browser
(they sit behind the Authentik login); it builds and the API behind it is
tested.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 02:23:06 +02:00