The data was all being collected (agent last-seen, per-disk usage,
Proxmox storage, Synology volume/disk health) but nothing acted on
it, so a dead server or a full disk was only noticed by opening the
right page.
A new health pass runs every 15 minutes (matching the agent's default
report interval) and raises one notification when a problem starts and
one when it clears: a server's agent silent past a threshold (default
60 min), a server disk / Proxmox storage or root filesystem / Synology
volume at or above a usage threshold (default 90%), and a Synology
volume or disk that isn't "normal", has bad SMART, bad sectors past the
threshold, or life remaining below it. Both thresholds and an on/off
toggle live under Settings -> Notifications.
The parts that make this trustworthy rather than noisy:
- A problem is keyed by identity, so it alerts once and not every run;
a shared Proxmox storage listed by every node is one problem, not
one per node.
- Active problems persist across restarts, so a rebuild doesn't
re-alert everything already known.
- If a source can't be read on a given run (Proxmox/Synology
unreachable, one node lacking privileges) its existing problems are
held, not reported "cleared" and then re-alerted when it comes back —
the integration-failure alert already owns "the integration is down".
- For 20 minutes after startup server-derived problems are held too:
agents couldn't report while the app was down, so judging them then
would report every server offline after any restart.
- An offline server's disk figures are stale and are not judged; a
server that never reported has no agent and raises nothing.
- Tracking continues while the toggle is off (only sending is gated),
so turning it back on doesn't dump every long-standing problem.
Timestamps without a zone (SQLite's format) are read as UTC; the test
runs on a UTC+2 machine, where reading them as local time gives a
different answer.
Verified with 32 checks: the evaluation rules and the state diff as
pure functions (exact thresholds, the proxmox:1 vs proxmox:10 prefix
trap, the flapping sequence), then a whole pass against a real
Proxmox adapter talking to a fake HTTPS cluster (one node returning
403, the whole API down, a shared storage on two nodes, a node that
recovers), a webhook receiver, the real DB, and the persisted state.
Not exercised end-to-end: the Synology collection path — its rules are
tested on data shaped exactly like the adapter's output types, but I
did not stand up a fake DSM. Real dev database mtime untouched.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>