Commit Graph
4 Commits
Author SHA1 Message Date
bobbanandClaude Sonnet 5 df2a5ce42b Add a Proxmox Backup Server integration: datastore/snapshot verification status
Proxmox VE already shows whether the last vzdump push to PBS succeeded, but
has no visibility into PBS's own backup verification, GC/prune health, or
host status. This adds PBS as its own integration (own adapter, page, nav
entry, and Dashboard widget) that reads datastore usage and, for every
stored snapshot, its verification state directly from PBS.

A new daily check (mirroring the existing Proxmox backup-failure check)
notifies when a snapshot has failed verification or a datastore couldn't be
read, with its own toggle in Settings -> Notifications and its own
maintenance-window silencing.

Not verified against a live PBS instance — built from PBS's published API
docs and a scratch test against a mocked PBS server exercising the adapter's
parsing and auth-header format (PBSAPIToken uses a colon separator, unlike
PVE's PVEAPIToken which uses =). See INTEGRATIONS.md for details and the
"not verified" caveat.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-29 20:32:52 +02:00
bobbanandClaude Sonnet 5 bf7f73f6b6 Import maintenance windows from Uptime Kuma
New "Import from Uptime Kuma" action on the Maintenance page (operator):
pick an Uptime Kuma integration and a duration, and it starts (or
extends) a maintenance window here for every server whose address
matches a monitor Uptime Kuma currently reports as being in maintenance,
reusing the existing monitor-to-server matching from the Uptime Kuma
integration itself.

Uptime Kuma's metrics endpoint only exposes a monitor's *current* status,
not its scheduled start/end time (there's no API for that), so this
deliberately doesn't try to mirror Uptime Kuma's own schedule -- it starts
a window for the duration you choose, the same bounded/required-end
window this feature has always used. Running it again while Kuma is
still in maintenance extends the same window rather than stacking a
second one; when Kuma later shows nothing in maintenance, already-active
windows are left alone rather than force-ended, since ending them isn't
something only Uptime Kuma's state should decide. Two Kuma monitors that
match the same server are deduped to one window. Monitors with no
matching server are reported back by name so nothing is silently missed,
and monitors that aren't in maintenance are ignored entirely.

The manual "Start maintenance" endpoint's start-or-extend logic (dedupe,
pruning old rows, the response shape) is now a shared
services/maintenance.ts function instead of living only in that route
handler, so the import path can't drift from how a manual window behaves.
Likewise the server-matching helper gained a small toMatchableServers()
so the existing Uptime Kuma monitors route and this new one build the
same match input the same way instead of each parsing server rows on
their own.

No schema change -- imported windows are ordinary maintenance windows;
their Uptime-Kuma origin is only in the reason text ("Imported from
Uptime Kuma (<integration>): <monitor>"), visible in the Active table and
the audit log like any other window.

Verified with 28 backend checks (the refactored manual start/extend
flow as a regression check; roles; unknown/wrong-type/disabled
integration; validation; nothing-in-maintenance; matching including
same-server dedup and unmatched monitors; re-running extends rather than
duplicating; windows left alone once Kuma exits maintenance; upstream
failure; audit entries) and by driving the real Maintenance page against
the real routers in a browser: import, re-import (extends), the
nothing-in-maintenance state, and the viewer view (no edit card, active
windows still visible). Real dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-29 19:11:35 +02:00
bobbanandClaude Sonnet 5 4118062405 Alert when a Semaphore template or Gitea workflow run fails
Every 15 minutes (on the existing health-check timer) the app reads the
latest run of each Semaphore template and each Gitea repo's latest
workflow run. A failed one raises one notification, and another when a
later run succeeds. It is state-based like the server health alerts, so a
job that fails every night alerts on the first failure, not every night.
Gitea alerts include the run's link. Toggle: Settings > Notifications.

What counts:
- Semaphore "error" is a failure, "success" is a pass. A run that is
  waiting, running, stopped by hand or rejected is neither, so it leaves
  the previous state alone: a run in progress must not clear a failure it
  hasn't fixed yet, and a manual stop isn't a failure.
- Gitea failure/success likewise; running, waiting, blocked, cancelled and
  skipped leave things as they were.
- A failing template that gets another failing run does not re-alert.

Not mistaking "couldn't read" for "fixed":
- Semaphore's template listing swallowed per-project errors, so a project
  that failed to load looked like a project with no templates. A new
  checkTemplates adapter method reports which projects failed, and their
  failures are held rather than cleared.
- Gitea reports a run it couldn't fetch as null, the same as "no runs";
  both leave the repo's state alone.
- An unreachable integration holds all of its failures. Nothing is cleared
  or re-announced while it is down.

The first pass only records what is already failing without announcing
it, so upgrading (or adding an integration to a fresh install) doesn't
produce a wall of alerts about months-old failures. That baseline is not
spent while nothing could be read.

Maintenance windows on a Semaphore or Gitea integration silence its
failure alerts with the same rules as the health alerts: a problem that
starts during a window alerts when it ends, and one already announced
stays known. The diff logic is reused from the health monitor rather than
copied. Maintenance page text updated.

Known limit: for Gitea this follows the repo's most recent run on any
workflow or branch, matching what the Gitea page shows; a failure in one
workflow can be masked by a later success of another.

Verified with 53 checks against fake Semaphore and Gitea servers and a
webhook receiver: classification, baseline (including not being consumed
when nothing is readable), single alert per failure, no repeat, in-progress/
stopped/cancelled runs, recovery and re-failure, unreadable project,
unreadable integration, run-fetch errors, maintenance windows (silenced,
then announced after), the toggle, disabled integrations and repos without
Actions. Real dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 03:11:59 +02:00
bobbanandClaude Sonnet 5 aae4f0d74f Add maintenance mode to silence alerts while working on a server or integration
Rebooting Proxmox or patching a server triggered failure/offline alerts
you then had to dismiss. A maintenance window silences alerts about one
server, integration, or DNS provider for a chosen time. New Maintenance
page (start with a duration and optional reason, end early, see what's
silenced and what isn't) and a banner in the app shell so every signed-in
user can see what is currently silenced. Starting/ending is operator-only
and audit-logged; starting one on a target that already has a window
restarts its clock instead of stacking.

Silenced for the target: server offline/disk alerts, Proxmox/Synology
storage and health alerts, Proxmox backup alerts, and "integration down"
alerts. Not silenced: expiry and update reminders, DNS change notices.

The design goal is that this cannot hide a real outage:
- Every window has a required end (5 min to 7 days); there is no
  open-ended option, so a forgotten window expires by itself.
- A silenced problem is deliberately NOT recorded as "known". If it is
  still present when the window ends it alerts then, as new. A problem
  that was already alerted before the window stays known, so it isn't
  repeated, and is reported cleared only after the window ends.
- Failure alerts keep counting failures during a window without marking
  themselves alerted, so an outage that outlasts the window alerts on the
  very next failed call.

Known limitation, stated on the page: integration-failure alerts are
tracked per service TYPE (all "proxmox"), not per configured instance, so
a window on one Proxmox integration also silences a failure on a second
Proxmox integration while it's open. Fixing that means threading the
integration id through every adapter and the diagnostic log, which is a
much larger change than this feature.

Also moved the API-error-message helper out of Secrets.tsx into a shared
util now that two pages use it. New table maintenance_windows (migration
0008).

Verified with 44 checks: the condition-key-to-subject mapping (including
server:3 vs server:33), the diff rules with silenced subjects (new problem
not recorded, alerts when the window ends; already-known one carried and
not repeated; clears only after the window), window expiry and
integration/DNS-provider source matching, the failure tracker end to end
against a webhook (silent during a window while an unrelated service still
alerts; outage that outlasts the window alerts on the next failure and
only once; fail-and-recover fully inside a window sends nothing), a full
health pass against a real window, and the real router with a stubbed
session (role rules, duration bounds including the missing-duration case,
extend-not-stack, 404s, deleted targets hidden, audit entries). Real dev
database mtime untouched.

Not done: I haven't clicked through the new page or banner in a browser
(they sit behind the Authentik login); it builds and the API behind it is
tested.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 02:23:06 +02:00