Alert when a server goes silent, a disk fills up, or a Synology volume degrades
The data was all being collected (agent last-seen, per-disk usage, Proxmox storage, Synology volume/disk health) but nothing acted on it, so a dead server or a full disk was only noticed by opening the right page. A new health pass runs every 15 minutes (matching the agent's default report interval) and raises one notification when a problem starts and one when it clears: a server's agent silent past a threshold (default 60 min), a server disk / Proxmox storage or root filesystem / Synology volume at or above a usage threshold (default 90%), and a Synology volume or disk that isn't "normal", has bad SMART, bad sectors past the threshold, or life remaining below it. Both thresholds and an on/off toggle live under Settings -> Notifications. The parts that make this trustworthy rather than noisy: - A problem is keyed by identity, so it alerts once and not every run; a shared Proxmox storage listed by every node is one problem, not one per node. - Active problems persist across restarts, so a rebuild doesn't re-alert everything already known. - If a source can't be read on a given run (Proxmox/Synology unreachable, one node lacking privileges) its existing problems are held, not reported "cleared" and then re-alerted when it comes back — the integration-failure alert already owns "the integration is down". - For 20 minutes after startup server-derived problems are held too: agents couldn't report while the app was down, so judging them then would report every server offline after any restart. - An offline server's disk figures are stale and are not judged; a server that never reported has no agent and raises nothing. - Tracking continues while the toggle is off (only sending is gated), so turning it back on doesn't dump every long-standing problem. Timestamps without a zone (SQLite's format) are read as UTC; the test runs on a UTC+2 machine, where reading them as local time gives a different answer. Verified with 32 checks: the evaluation rules and the state diff as pure functions (exact thresholds, the proxmox:1 vs proxmox:10 prefix trap, the flapping sequence), then a whole pass against a real Proxmox adapter talking to a fake HTTPS cluster (one node returning 403, the whole API down, a shared storage on two nodes, a node that recovers), a webhook receiver, the real DB, and the persisted state. Not exercised end-to-end: the Synology collection path — its rules are tested on data shaped exactly like the adapter's output types, but I did not stand up a fake DSM. Real dev database mtime untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
bf03a15337
commit
1688de3ea2
9 files changed
+420
-1
No files matched your search
@@ -135,6 +135,13 @@ real API response (fixed to derive them from `connectedToControl` and
|
||||
`enabledRoutes`). See the git log for the full verification notes per
|
||||
integration.
|
||||
|
||||
Server and storage health is watched every 15 minutes: a server whose agent
|
||||
stops reporting, a server disk / Proxmox storage / Synology volume passing a
|
||||
usage threshold, and a Synology volume or disk that's degraded or failing each
|
||||
raise one notification when the problem starts and one when it clears (both
|
||||
thresholds are set under Settings → Notifications). Active problems are
|
||||
remembered across restarts, so a rebuild doesn't re-alert them.
|
||||
|
||||
The app is installable as a PWA — "Install app" / "Add to Home Screen" from the
|
||||
browser gives it its own icon and a standalone window on phone or desktop. This
|
||||
needs the site to be served over HTTPS (browsers only offer install on secure
|
||||
|
||||
Reference in new issue
Block a user