Alert when a server goes silent, a disk fills up, or a Synology volume degrades
The data was all being collected (agent last-seen, per-disk usage, Proxmox storage, Synology volume/disk health) but nothing acted on it, so a dead server or a full disk was only noticed by opening the right page. A new health pass runs every 15 minutes (matching the agent's default report interval) and raises one notification when a problem starts and one when it clears: a server's agent silent past a threshold (default 60 min), a server disk / Proxmox storage or root filesystem / Synology volume at or above a usage threshold (default 90%), and a Synology volume or disk that isn't "normal", has bad SMART, bad sectors past the threshold, or life remaining below it. Both thresholds and an on/off toggle live under Settings -> Notifications. The parts that make this trustworthy rather than noisy: - A problem is keyed by identity, so it alerts once and not every run; a shared Proxmox storage listed by every node is one problem, not one per node. - Active problems persist across restarts, so a rebuild doesn't re-alert everything already known. - If a source can't be read on a given run (Proxmox/Synology unreachable, one node lacking privileges) its existing problems are held, not reported "cleared" and then re-alerted when it comes back — the integration-failure alert already owns "the integration is down". - For 20 minutes after startup server-derived problems are held too: agents couldn't report while the app was down, so judging them then would report every server offline after any restart. - An offline server's disk figures are stale and are not judged; a server that never reported has no agent and raises nothing. - Tracking continues while the toggle is off (only sending is gated), so turning it back on doesn't dump every long-standing problem. Timestamps without a zone (SQLite's format) are read as UTC; the test runs on a UTC+2 machine, where reading them as local time gives a different answer. Verified with 32 checks: the evaluation rules and the state diff as pure functions (exact thresholds, the proxmox:1 vs proxmox:10 prefix trap, the flapping sequence), then a whole pass against a real Proxmox adapter talking to a fake HTTPS cluster (one node returning 403, the whole API down, a shared storage on two nodes, a node that recovers), a webhook receiver, the real DB, and the persisted state. Not exercised end-to-end: the Synology collection path — its rules are tested on data shaped exactly like the adapter's output types, but I did not stand up a fake DSM. Real dev database mtime untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
bf03a15337
commit
1688de3ea2
9 files changed
+420
-1
No files matched your search
@@ -42,6 +42,7 @@ export interface NotificationEvents {
|
||||
tailscaleKeyCheck: boolean;
|
||||
dockerUpdateCheck: boolean;
|
||||
proxmoxBackupCheck: boolean;
|
||||
healthAlerts: boolean;
|
||||
secretCheckTime: string; // "HH:MM" — shared by the secret-expiry, Tailscale key-expiry, Docker update, and Proxmox backup checks
|
||||
timezone: string;
|
||||
integrationFailureAlerts: boolean;
|
||||
@@ -78,6 +79,13 @@ export interface QuietHoursSettings {
|
||||
end: string;
|
||||
}
|
||||
|
||||
export interface HealthCheckSettings {
|
||||
/** A server whose agent hasn't reported for this long is considered offline (agent default interval is 15 min). */
|
||||
serverOfflineMinutes: number;
|
||||
/** Disk / storage / volume usage at or above this percentage is reported. */
|
||||
diskUsagePercent: number;
|
||||
}
|
||||
|
||||
export interface AppSettings {
|
||||
gotify: GotifySettings;
|
||||
ntfy: NtfySettings;
|
||||
@@ -89,6 +97,7 @@ export interface AppSettings {
|
||||
display: DisplaySettings;
|
||||
logRetention: LogRetentionSettings;
|
||||
quietHours: QuietHoursSettings;
|
||||
healthChecks: HealthCheckSettings;
|
||||
}
|
||||
|
||||
const DEFAULTS: AppSettings = {
|
||||
@@ -104,6 +113,7 @@ const DEFAULTS: AppSettings = {
|
||||
tailscaleKeyCheck: true,
|
||||
dockerUpdateCheck: true,
|
||||
proxmoxBackupCheck: true,
|
||||
healthAlerts: true,
|
||||
secretCheckTime: "08:00",
|
||||
timezone: "UTC",
|
||||
integrationFailureAlerts: true,
|
||||
@@ -114,6 +124,7 @@ const DEFAULTS: AppSettings = {
|
||||
display: { dateFormat: "ymd", timeFormat: "24h", pageSize: 20 },
|
||||
logRetention: { enabled: false, retentionDays: 90, intervalHours: 24 },
|
||||
quietHours: { enabled: false, start: "22:00", end: "07:00" },
|
||||
healthChecks: { serverOfflineMinutes: 60, diskUsagePercent: 90 },
|
||||
};
|
||||
|
||||
const KEYS = Object.keys(DEFAULTS) as (keyof AppSettings)[];
|
||||
|
||||
Reference in new issue
Block a user