Alert when a server goes silent, a disk fills up, or a Synology volume degrades

The data was all being collected (agent last-seen, per-disk usage,
Proxmox storage, Synology volume/disk health) but nothing acted on
it, so a dead server or a full disk was only noticed by opening the
right page.

A new health pass runs every 15 minutes (matching the agent's default
report interval) and raises one notification when a problem starts and
one when it clears: a server's agent silent past a threshold (default
60 min), a server disk / Proxmox storage or root filesystem / Synology
volume at or above a usage threshold (default 90%), and a Synology
volume or disk that isn't "normal", has bad SMART, bad sectors past the
threshold, or life remaining below it. Both thresholds and an on/off
toggle live under Settings -> Notifications.

The parts that make this trustworthy rather than noisy:
- A problem is keyed by identity, so it alerts once and not every run;
  a shared Proxmox storage listed by every node is one problem, not
  one per node.
- Active problems persist across restarts, so a rebuild doesn't
  re-alert everything already known.
- If a source can't be read on a given run (Proxmox/Synology
  unreachable, one node lacking privileges) its existing problems are
  held, not reported "cleared" and then re-alerted when it comes back —
  the integration-failure alert already owns "the integration is down".
- For 20 minutes after startup server-derived problems are held too:
  agents couldn't report while the app was down, so judging them then
  would report every server offline after any restart.
- An offline server's disk figures are stale and are not judged; a
  server that never reported has no agent and raises nothing.
- Tracking continues while the toggle is off (only sending is gated),
  so turning it back on doesn't dump every long-standing problem.

Timestamps without a zone (SQLite's format) are read as UTC; the test
runs on a UTC+2 machine, where reading them as local time gives a
different answer.

Verified with 32 checks: the evaluation rules and the state diff as
pure functions (exact thresholds, the proxmox:1 vs proxmox:10 prefix
trap, the flapping sequence), then a whole pass against a real
Proxmox adapter talking to a fake HTTPS cluster (one node returning
403, the whole API down, a shared storage on two nodes, a node that
recovers), a webhook receiver, the real DB, and the persisted state.

Not exercised end-to-end: the Synology collection path — its rules are
tested on data shaped exactly like the adapter's output types, but I
did not stand up a fake DSM. Real dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
bobbanandClaude Sonnet 5 committed 2026-09-26 02:15:19 +02:00
1 parent bf03a15337
commit 1688de3ea2
9 files changed
+420 -1

No files matched your search

+7
View File
@@ -162,6 +162,7 @@ export interface NotificationEvents {
tailscaleKeyCheck: boolean;
dockerUpdateCheck: boolean;
proxmoxBackupCheck: boolean;
healthAlerts: boolean;
secretCheckTime: string;
timezone: string;
integrationFailureAlerts: boolean;
@@ -183,6 +184,11 @@ export interface LogRetentionSettings {
intervalHours: number;
}
export interface HealthCheckSettings {
serverOfflineMinutes: number;
diskUsagePercent: number;
}
export interface QuietHoursSettings {
enabled: boolean;
start: string;
@@ -200,6 +206,7 @@ export interface AppSettings {
display: DisplaySettings;
logRetention: LogRetentionSettings;
quietHours: QuietHoursSettings;
healthChecks: HealthCheckSettings;
}
export type AppSettingsPatch = { [K in keyof AppSettings]?: Partial<AppSettings[K]> };
@@ -8,6 +8,7 @@ import {
type WebhookSettings,
type NotificationEvents,
type QuietHoursSettings,
type HealthCheckSettings,
} from "../../api/client";
const TIMEZONES = [
@@ -45,12 +46,14 @@ const DEFAULT_NOTIFICATIONS: NotificationEvents = {
tailscaleKeyCheck: true,
dockerUpdateCheck: true,
proxmoxBackupCheck: true,
healthAlerts: true,
secretCheckTime: "08:00",
timezone: "UTC",
integrationFailureAlerts: true,
integrationFailureThreshold: 3,
};
const DEFAULT_QUIET_HOURS: QuietHoursSettings = { enabled: false, start: "22:00", end: "07:00" };
const DEFAULT_HEALTH_CHECKS: HealthCheckSettings = { serverOfflineMinutes: 60, diskUsagePercent: 90 };
type TestResult = { ok: boolean; message: string } | null;
@@ -76,6 +79,7 @@ export default function NotificationSettings() {
const [webhook, setWebhook] = useState<WebhookSettings>(DEFAULT_WEBHOOK);
const [notifications, setNotifications] = useState<NotificationEvents>(DEFAULT_NOTIFICATIONS);
const [quietHours, setQuietHours] = useState<QuietHoursSettings>(DEFAULT_QUIET_HOURS);
const [healthChecks, setHealthChecks] = useState<HealthCheckSettings>(DEFAULT_HEALTH_CHECKS);
const [queuedCount, setQueuedCount] = useState<number | null>(null);
const [flushing, setFlushing] = useState(false);
const [flushResult, setFlushResult] = useState<string | null>(null);
@@ -99,6 +103,7 @@ export default function NotificationSettings() {
setWebhook(res.settings.webhook);
setNotifications(res.settings.notifications);
setQuietHours(res.settings.quietHours);
setHealthChecks(res.settings.healthChecks);
})
.catch((err) => setLoadError(err instanceof Error ? err.message : String(err)))
.finally(() => setLoading(false));
@@ -113,7 +118,7 @@ export default function NotificationSettings() {
setSaveError(null);
setSaved(false);
try {
await api.settings.update({ gotify, ntfy, smtp, webhook, notifications, quietHours });
await api.settings.update({ gotify, ntfy, smtp, webhook, notifications, quietHours, healthChecks });
setSaved(true);
setTimeout(() => setSaved(false), 3000);
} catch (err) {
@@ -520,6 +525,7 @@ export default function NotificationSettings() {
{ key: "tailscaleKeyCheck" as const, label: "Tailscale key expiry reminder" },
{ key: "dockerUpdateCheck" as const, label: "Docker image update available" },
{ key: "proxmoxBackupCheck" as const, label: "Proxmox backup failed or a guest has no coverage" },
{ key: "healthAlerts" as const, label: "Server offline, disk nearly full, or Synology volume/disk problem" },
{ key: "integrationFailureAlerts" as const, label: "Integration/DNS provider failing repeatedly" },
].map(({ key, label }) => (
<label key={key} className="form-check mb-2">
@@ -552,6 +558,40 @@ export default function NotificationSettings() {
</div>
</div>
</div>
<div className="row g-2 mt-2">
<div className="col-6">
<label className="form-label">Server offline after</label>
<div className="input-group">
<input
type="number"
className="form-control"
min={15}
max={10080}
value={healthChecks.serverOfflineMinutes}
disabled={!notifications.healthAlerts}
onChange={(e) => setHealthChecks((h) => ({ ...h, serverOfflineMinutes: Number(e.target.value) }))}
/>
<span className="input-group-text">min</span>
</div>
<div className="form-hint">Agents report every 15 min by default.</div>
</div>
<div className="col-6">
<label className="form-label">Disk usage alert at</label>
<div className="input-group">
<input
type="number"
className="form-control"
min={50}
max={99}
value={healthChecks.diskUsagePercent}
disabled={!notifications.healthAlerts}
onChange={(e) => setHealthChecks((h) => ({ ...h, diskUsagePercent: Number(e.target.value) }))}
/>
<span className="input-group-text">%</span>
</div>
<div className="form-hint">Servers, Proxmox storage, Synology volumes. Checked every 15 min; a recovery notice follows.</div>
</div>
</div>
{(() => {
const dailyChecksEnabled =
notifications.secretCheck || notifications.tailscaleKeyCheck || notifications.dockerUpdateCheck || notifications.proxmoxBackupCheck;