Alert when a server goes silent, a disk fills up, or a Synology volume degrades

The data was all being collected (agent last-seen, per-disk usage,
Proxmox storage, Synology volume/disk health) but nothing acted on
it, so a dead server or a full disk was only noticed by opening the
right page.

A new health pass runs every 15 minutes (matching the agent's default
report interval) and raises one notification when a problem starts and
one when it clears: a server's agent silent past a threshold (default
60 min), a server disk / Proxmox storage or root filesystem / Synology
volume at or above a usage threshold (default 90%), and a Synology
volume or disk that isn't "normal", has bad SMART, bad sectors past the
threshold, or life remaining below it. Both thresholds and an on/off
toggle live under Settings -> Notifications.

The parts that make this trustworthy rather than noisy:
- A problem is keyed by identity, so it alerts once and not every run;
  a shared Proxmox storage listed by every node is one problem, not
  one per node.
- Active problems persist across restarts, so a rebuild doesn't
  re-alert everything already known.
- If a source can't be read on a given run (Proxmox/Synology
  unreachable, one node lacking privileges) its existing problems are
  held, not reported "cleared" and then re-alerted when it comes back —
  the integration-failure alert already owns "the integration is down".
- For 20 minutes after startup server-derived problems are held too:
  agents couldn't report while the app was down, so judging them then
  would report every server offline after any restart.
- An offline server's disk figures are stale and are not judged; a
  server that never reported has no agent and raises nothing.
- Tracking continues while the toggle is off (only sending is gated),
  so turning it back on doesn't dump every long-standing problem.

Timestamps without a zone (SQLite's format) are read as UTC; the test
runs on a UTC+2 machine, where reading them as local time gives a
different answer.

Verified with 32 checks: the evaluation rules and the state diff as
pure functions (exact thresholds, the proxmox:1 vs proxmox:10 prefix
trap, the flapping sequence), then a whole pass against a real
Proxmox adapter talking to a fake HTTPS cluster (one node returning
403, the whole API down, a shared storage on two nodes, a node that
recovers), a webhook receiver, the real DB, and the persisted state.

Not exercised end-to-end: the Synology collection path — its rules are
tested on data shaped exactly like the adapter's output types, but I
did not stand up a fake DSM. Real dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
bobbanandClaude Sonnet 5 committed 2026-09-26 02:15:19 +02:00
1 parent bf03a15337
commit 1688de3ea2
9 files changed
+420 -1

No files matched your search

+11
View File
@@ -42,6 +42,7 @@ export interface NotificationEvents {
tailscaleKeyCheck: boolean;
dockerUpdateCheck: boolean;
proxmoxBackupCheck: boolean;
healthAlerts: boolean;
secretCheckTime: string; // "HH:MM" — shared by the secret-expiry, Tailscale key-expiry, Docker update, and Proxmox backup checks
timezone: string;
integrationFailureAlerts: boolean;
@@ -78,6 +79,13 @@ export interface QuietHoursSettings {
end: string;
}
export interface HealthCheckSettings {
/** A server whose agent hasn't reported for this long is considered offline (agent default interval is 15 min). */
serverOfflineMinutes: number;
/** Disk / storage / volume usage at or above this percentage is reported. */
diskUsagePercent: number;
}
export interface AppSettings {
gotify: GotifySettings;
ntfy: NtfySettings;
@@ -89,6 +97,7 @@ export interface AppSettings {
display: DisplaySettings;
logRetention: LogRetentionSettings;
quietHours: QuietHoursSettings;
healthChecks: HealthCheckSettings;
}
const DEFAULTS: AppSettings = {
@@ -104,6 +113,7 @@ const DEFAULTS: AppSettings = {
tailscaleKeyCheck: true,
dockerUpdateCheck: true,
proxmoxBackupCheck: true,
healthAlerts: true,
secretCheckTime: "08:00",
timezone: "UTC",
integrationFailureAlerts: true,
@@ -114,6 +124,7 @@ const DEFAULTS: AppSettings = {
display: { dateFormat: "ymd", timeFormat: "24h", pageSize: 20 },
logRetention: { enabled: false, retentionDays: 90, intervalHours: 24 },
quietHours: { enabled: false, start: "22:00", end: "07:00" },
healthChecks: { serverOfflineMinutes: 60, diskUsagePercent: 90 },
};
const KEYS = Object.keys(DEFAULTS) as (keyof AppSettings)[];