Commit Graph
3 Commits
Author SHA1 Message Date
bobbanandClaude Sonnet 5.5 447f33fff6 Add an Alerts page under Operations listing everything that's wrong now
One list of the current problems across servers and integrations, instead
of waiting for a notification or visiting each page: servers that stopped
reporting, full or nearly full disks and volumes (critical from 95%),
Synology volume/disk problems, failed or uncovered Proxmox backups,
failed Proxmox Backup Server verifications, container image updates,
expired or expiring secrets/domains/Tailscale keys, failed Semaphore and
Gitea runs, Uptime Kuma monitors that are down, overdue osTicket tickets,
and integrations whose calls keep failing. Visible to every role, with
severity and kind filters, search, sorting, CSV export and "Check now".

It runs the same checks that send the notifications rather than a second
copy of them: the detection in the health, automation, Proxmox backup,
PBS, Docker update and Tailscale key checks is pulled out into shared
collectors that both the schedulers and the page call, so the two can't
disagree about what counts as a problem. Notification behaviour is
unchanged, including the scheduled backup checks skipping integrations
under a maintenance window. Unlike the notifications the page ignores the
on/off toggles, and keeps problems under a maintenance window, marked
silenced and counted apart.

It reads live, so a result is reused for a minute (and Refresh can't
re-run everything more than once every ten seconds), and every source has
a 20 s limit so one hung integration can't hang the page. Anything it
couldn't read is called out at the top instead of looking like all clear,
and server checks pause for the same 20 minutes after a restart as the
notifications do, with a note saying so.

Also gives the newer integrations (PBS, osTicket, Uptime Kuma, phpIPAM)
proper names in "integration down" notifications instead of their ids.

Verified through the real routes against a scratch database with fake
backends (offline and full-disk servers, secrets and domains, a silenced
server, a fake PBS with failed verification, a hanging integration, a
refused one, a failing-calls streak, caching, the restart grace period,
auth), and by rendering the real page against that data in a browser:
filters, search, silenced toggle, sorting, Check now, dark mode.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 23:07:57 +02:00
bobbanandClaude Sonnet 5 aae4f0d74f Add maintenance mode to silence alerts while working on a server or integration
Rebooting Proxmox or patching a server triggered failure/offline alerts
you then had to dismiss. A maintenance window silences alerts about one
server, integration, or DNS provider for a chosen time. New Maintenance
page (start with a duration and optional reason, end early, see what's
silenced and what isn't) and a banner in the app shell so every signed-in
user can see what is currently silenced. Starting/ending is operator-only
and audit-logged; starting one on a target that already has a window
restarts its clock instead of stacking.

Silenced for the target: server offline/disk alerts, Proxmox/Synology
storage and health alerts, Proxmox backup alerts, and "integration down"
alerts. Not silenced: expiry and update reminders, DNS change notices.

The design goal is that this cannot hide a real outage:
- Every window has a required end (5 min to 7 days); there is no
  open-ended option, so a forgotten window expires by itself.
- A silenced problem is deliberately NOT recorded as "known". If it is
  still present when the window ends it alerts then, as new. A problem
  that was already alerted before the window stays known, so it isn't
  repeated, and is reported cleared only after the window ends.
- Failure alerts keep counting failures during a window without marking
  themselves alerted, so an outage that outlasts the window alerts on the
  very next failed call.

Known limitation, stated on the page: integration-failure alerts are
tracked per service TYPE (all "proxmox"), not per configured instance, so
a window on one Proxmox integration also silences a failure on a second
Proxmox integration while it's open. Fixing that means threading the
integration id through every adapter and the diagnostic log, which is a
much larger change than this feature.

Also moved the API-error-message helper out of Secrets.tsx into a shared
util now that two pages use it. New table maintenance_windows (migration
0008).

Verified with 44 checks: the condition-key-to-subject mapping (including
server:3 vs server:33), the diff rules with silenced subjects (new problem
not recorded, alerts when the window ends; already-known one carried and
not repeated; clears only after the window), window expiry and
integration/DNS-provider source matching, the failure tracker end to end
against a webhook (silent during a window while an unrelated service still
alerts; outage that outlasts the window alerts on the next failure and
only once; fail-and-recover fully inside a window sends nothing), a full
health pass against a real window, and the real router with a stubbed
session (role rules, duration bounds including the missing-duration case,
extend-not-stack, 404s, deleted targets hidden, audit entries). Real dev
database mtime untouched.

Not done: I haven't clicked through the new page or banner in a browser
(they sit behind the Authentik login); it builds and the API behind it is
tested.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 02:23:06 +02:00
bobbanandClaude Sonnet 5 1688de3ea2 Alert when a server goes silent, a disk fills up, or a Synology volume degrades
The data was all being collected (agent last-seen, per-disk usage,
Proxmox storage, Synology volume/disk health) but nothing acted on
it, so a dead server or a full disk was only noticed by opening the
right page.

A new health pass runs every 15 minutes (matching the agent's default
report interval) and raises one notification when a problem starts and
one when it clears: a server's agent silent past a threshold (default
60 min), a server disk / Proxmox storage or root filesystem / Synology
volume at or above a usage threshold (default 90%), and a Synology
volume or disk that isn't "normal", has bad SMART, bad sectors past the
threshold, or life remaining below it. Both thresholds and an on/off
toggle live under Settings -> Notifications.

The parts that make this trustworthy rather than noisy:
- A problem is keyed by identity, so it alerts once and not every run;
  a shared Proxmox storage listed by every node is one problem, not
  one per node.
- Active problems persist across restarts, so a rebuild doesn't
  re-alert everything already known.
- If a source can't be read on a given run (Proxmox/Synology
  unreachable, one node lacking privileges) its existing problems are
  held, not reported "cleared" and then re-alerted when it comes back —
  the integration-failure alert already owns "the integration is down".
- For 20 minutes after startup server-derived problems are held too:
  agents couldn't report while the app was down, so judging them then
  would report every server offline after any restart.
- An offline server's disk figures are stale and are not judged; a
  server that never reported has no agent and raises nothing.
- Tracking continues while the toggle is off (only sending is gated),
  so turning it back on doesn't dump every long-standing problem.

Timestamps without a zone (SQLite's format) are read as UTC; the test
runs on a UTC+2 machine, where reading them as local time gives a
different answer.

Verified with 32 checks: the evaluation rules and the state diff as
pure functions (exact thresholds, the proxmox:1 vs proxmox:10 prefix
trap, the flapping sequence), then a whole pass against a real
Proxmox adapter talking to a fake HTTPS cluster (one node returning
403, the whole API down, a shared storage on two nodes, a node that
recovers), a webhook receiver, the real DB, and the persisted state.

Not exercised end-to-end: the Synology collection path — its rules are
tested on data shaped exactly like the adapter's output types, but I
did not stand up a fake DSM. Real dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 02:15:19 +02:00