Alert when a server goes silent, a disk fills up, or a Synology volume degrades

The data was all being collected (agent last-seen, per-disk usage,
Proxmox storage, Synology volume/disk health) but nothing acted on
it, so a dead server or a full disk was only noticed by opening the
right page.

A new health pass runs every 15 minutes (matching the agent's default
report interval) and raises one notification when a problem starts and
one when it clears: a server's agent silent past a threshold (default
60 min), a server disk / Proxmox storage or root filesystem / Synology
volume at or above a usage threshold (default 90%), and a Synology
volume or disk that isn't "normal", has bad SMART, bad sectors past the
threshold, or life remaining below it. Both thresholds and an on/off
toggle live under Settings -> Notifications.

The parts that make this trustworthy rather than noisy:
- A problem is keyed by identity, so it alerts once and not every run;
  a shared Proxmox storage listed by every node is one problem, not
  one per node.
- Active problems persist across restarts, so a rebuild doesn't
  re-alert everything already known.
- If a source can't be read on a given run (Proxmox/Synology
  unreachable, one node lacking privileges) its existing problems are
  held, not reported "cleared" and then re-alerted when it comes back —
  the integration-failure alert already owns "the integration is down".
- For 20 minutes after startup server-derived problems are held too:
  agents couldn't report while the app was down, so judging them then
  would report every server offline after any restart.
- An offline server's disk figures are stale and are not judged; a
  server that never reported has no agent and raises nothing.
- Tracking continues while the toggle is off (only sending is gated),
  so turning it back on doesn't dump every long-standing problem.

Timestamps without a zone (SQLite's format) are read as UTC; the test
runs on a UTC+2 machine, where reading them as local time gives a
different answer.

Verified with 32 checks: the evaluation rules and the state diff as
pure functions (exact thresholds, the proxmox:1 vs proxmox:10 prefix
trap, the flapping sequence), then a whole pass against a real
Proxmox adapter talking to a fake HTTPS cluster (one node returning
403, the whole API down, a shared storage on two nodes, a node that
recovers), a webhook receiver, the real DB, and the persisted state.

Not exercised end-to-end: the Synology collection path — its rules are
tested on data shaped exactly like the adapter's output types, but I
did not stand up a fake DSM. Real dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
bobbanandClaude Sonnet 5 committed 2026-09-26 02:15:19 +02:00
1 parent bf03a15337
commit 1688de3ea2
9 files changed
+420 -1

No files matched your search

+7
View File
@@ -162,6 +162,7 @@ export interface NotificationEvents {
tailscaleKeyCheck: boolean;
dockerUpdateCheck: boolean;
proxmoxBackupCheck: boolean;
healthAlerts: boolean;
secretCheckTime: string;
timezone: string;
integrationFailureAlerts: boolean;
@@ -183,6 +184,11 @@ export interface LogRetentionSettings {
intervalHours: number;
}
export interface HealthCheckSettings {
serverOfflineMinutes: number;
diskUsagePercent: number;
}
export interface QuietHoursSettings {
enabled: boolean;
start: string;
@@ -200,6 +206,7 @@ export interface AppSettings {
display: DisplaySettings;
logRetention: LogRetentionSettings;
quietHours: QuietHoursSettings;
healthChecks: HealthCheckSettings;
}
export type AppSettingsPatch = { [K in keyof AppSettings]?: Partial<AppSettings[K]> };