Alert when a server goes silent, a disk fills up, or a Synology volume degrades
The data was all being collected (agent last-seen, per-disk usage, Proxmox storage, Synology volume/disk health) but nothing acted on it, so a dead server or a full disk was only noticed by opening the right page. A new health pass runs every 15 minutes (matching the agent's default report interval) and raises one notification when a problem starts and one when it clears: a server's agent silent past a threshold (default 60 min), a server disk / Proxmox storage or root filesystem / Synology volume at or above a usage threshold (default 90%), and a Synology volume or disk that isn't "normal", has bad SMART, bad sectors past the threshold, or life remaining below it. Both thresholds and an on/off toggle live under Settings -> Notifications. The parts that make this trustworthy rather than noisy: - A problem is keyed by identity, so it alerts once and not every run; a shared Proxmox storage listed by every node is one problem, not one per node. - Active problems persist across restarts, so a rebuild doesn't re-alert everything already known. - If a source can't be read on a given run (Proxmox/Synology unreachable, one node lacking privileges) its existing problems are held, not reported "cleared" and then re-alerted when it comes back — the integration-failure alert already owns "the integration is down". - For 20 minutes after startup server-derived problems are held too: agents couldn't report while the app was down, so judging them then would report every server offline after any restart. - An offline server's disk figures are stale and are not judged; a server that never reported has no agent and raises nothing. - Tracking continues while the toggle is off (only sending is gated), so turning it back on doesn't dump every long-standing problem. Timestamps without a zone (SQLite's format) are read as UTC; the test runs on a UTC+2 machine, where reading them as local time gives a different answer. Verified with 32 checks: the evaluation rules and the state diff as pure functions (exact thresholds, the proxmox:1 vs proxmox:10 prefix trap, the flapping sequence), then a whole pass against a real Proxmox adapter talking to a fake HTTPS cluster (one node returning 403, the whole API down, a shared storage on two nodes, a node that recovers), a webhook receiver, the real DB, and the persisted state. Not exercised end-to-end: the Synology collection path — its rules are tested on data shaped exactly like the adapter's output types, but I did not stand up a fake DSM. Real dev database mtime untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
bf03a15337
commit
1688de3ea2
9 files changed
+420
-1
No files matched your search
@@ -68,6 +68,7 @@ const updateSchema = z.object({
|
||||
tailscaleKeyCheck: z.boolean(),
|
||||
dockerUpdateCheck: z.boolean(),
|
||||
proxmoxBackupCheck: z.boolean(),
|
||||
healthAlerts: z.boolean(),
|
||||
secretCheckTime: z.string().regex(/^\d{2}:\d{2}$/),
|
||||
timezone: z.string(),
|
||||
integrationFailureAlerts: z.boolean(),
|
||||
@@ -85,6 +86,10 @@ const updateSchema = z.object({
|
||||
.object({ enabled: z.boolean(), retentionDays: z.number().int().min(1).max(3650), intervalHours: z.number().int().min(1).max(720) })
|
||||
.partial()
|
||||
.optional(),
|
||||
healthChecks: z
|
||||
.object({ serverOfflineMinutes: z.number().int().min(15).max(10080), diskUsagePercent: z.number().int().min(50).max(99) })
|
||||
.partial()
|
||||
.optional(),
|
||||
quietHours: z
|
||||
.object({ enabled: z.boolean(), start: z.string().regex(/^\d{2}:\d{2}$/), end: z.string().regex(/^\d{2}:\d{2}$/) })
|
||||
.partial()
|
||||
|
||||
Reference in new issue
Block a user