Files
Homelab-manager/NOTIFICATIONS.md
bobbanandClaude Sonnet 5.5 447f33fff6 Add an Alerts page under Operations listing everything that's wrong now
One list of the current problems across servers and integrations, instead
of waiting for a notification or visiting each page: servers that stopped
reporting, full or nearly full disks and volumes (critical from 95%),
Synology volume/disk problems, failed or uncovered Proxmox backups,
failed Proxmox Backup Server verifications, container image updates,
expired or expiring secrets/domains/Tailscale keys, failed Semaphore and
Gitea runs, Uptime Kuma monitors that are down, overdue osTicket tickets,
and integrations whose calls keep failing. Visible to every role, with
severity and kind filters, search, sorting, CSV export and "Check now".

It runs the same checks that send the notifications rather than a second
copy of them: the detection in the health, automation, Proxmox backup,
PBS, Docker update and Tailscale key checks is pulled out into shared
collectors that both the schedulers and the page call, so the two can't
disagree about what counts as a problem. Notification behaviour is
unchanged, including the scheduled backup checks skipping integrations
under a maintenance window. Unlike the notifications the page ignores the
on/off toggles, and keeps problems under a maintenance window, marked
silenced and counted apart.

It reads live, so a result is reused for a minute (and Refresh can't
re-run everything more than once every ten seconds), and every source has
a 20 s limit so one hung integration can't hang the page. Anything it
couldn't read is called out at the top instead of looking like all clear,
and server checks pause for the same 20 minutes after a restart as the
notifications do, with a note saying so.

Also gives the newer integrations (PBS, osTicket, Uptime Kuma, phpIPAM)
proper names in "integration down" notifications instead of their ids.

Verified through the real routes against a scratch database with fake
backends (offline and full-disk servers, secrets and domains, a silenced
server, a fake PBS with failed verification, a hanging integration, a
refused one, a failing-calls streak, caching, the restart grace period,
auth), and by rendering the real page against that data in a browser:
filters, search, silenced toggle, sorting, Check now, dark mode.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 23:07:57 +02:00

127 lines
10 KiB
Markdown

# Notifications
Every notification in Homelab Manager is sent through the same pipe
(`server/src/services/notify.ts`) to whichever channels are enabled under
**Settings → Notifications**: Gotify, ntfy, SMTP (email), and a generic
JSON webhook. All four fire for every notification below — there's no
per-event channel routing, only a per-event on/off toggle (the "Notify
on" list on that same page) and, for the daily ones, one shared
time/timezone.
Two settings apply to *every* notification regardless of what triggered
it:
- **Quiet hours** (Settings → Notifications → Quiet hours): while
enabled and inside the configured window, a notification is held
instead of sent immediately, then delivered as a single digest at the
window's end time. This matters most for the real-time alerts below —
the daily reminders already fire once at a time you pick, usually
outside the window anyway.
- **Maintenance windows** (the Maintenance page): opening one for a
server, integration, or DNS provider silences the alerts that target
it specifically (offline/disk-full for a server; storage, backup, and
automation-failure alerts for an integration; "integration down" for
that whole service type) — see the Maintenance page's own "What gets
silenced" panel for the exact list.
Where a trigger has a configurable threshold, that's noted in its row
below; several are fixed and can't be changed from the UI.
## Daily reminders
These six checks share one schedule: **Settings → Notifications → Daily
reminder time** (default 08:00) and **Timezone** (default UTC). Unlike
the state-based alerts further down, these **re-send every day the
condition is still true** — there's no "only once" de-duplication, so an
expired secret you haven't renewed yet will be mentioned again at the
next day's check, and the day after that.
Each one also runs once at server startup if it hasn't already run
today (e.g. after an upgrade or a period offline), so you're not waiting
until the next scheduled time to catch up.
| Check | "Notify on" toggle | Fires when | Configured at |
|---|---|---|---|
| Secret expiry | *Secret expiry reminder* | Any tracked secret (API token, SSL cert, password, generic) is expired or within its own configured warning window. SSL certificates with a host:port are re-checked live first, so this reflects the real current expiry, not a stale saved date. A certificate that couldn't be re-checked live is mentioned separately, so the expiry shown may be stale. | Per-secret warning threshold, set when adding/editing that secret |
| Domain expiry | *Domain registration expiring or expired (daily reminder)* | Any tracked domain registration is expired or within the warning window; every domain is re-looked-up against its registry first. A domain whose registry lookup has been failing for 3+ days is also mentioned, separately, as "may be stale." Registries that don't publish an expiry (e.g. `.de`, `.eu`) are tracked but never trigger this. | Settings → Notifications → Health checks → **Domain expiry warning** (default 30 days) |
| Tailscale key expiry | *Tailscale key expiry reminder* | Any device's node key (across every enabled Tailscale integration) is within 30 days of expiring, or already expired. Devices with key expiry disabled are skipped. | Fixed at 30 days |
| Docker image updates | *Docker image update available* | Any container (across every enabled Dockhand integration) has an image update available, per Dockhand's own cached update-check results — this doesn't trigger a fresh registry lookup, just reads what the Docker page itself would show. | Not configurable (reflects Dockhand's own check interval) |
| Proxmox backups | *Proxmox backup failed or a guest has no coverage* | Two independent conditions, both under this one toggle: (1) a node's most recent `vzdump` backup task (across every enabled Proxmox integration) didn't succeed; (2) a VM/LXC isn't covered by any enabled backup job at all. | Not configurable |
| Proxmox Backup Server verification | *Proxmox Backup Server snapshot failed verification* | Across every enabled PBS integration: any datastore has at least one stored snapshot whose verification state is "failed", or a datastore couldn't be read at all (e.g. a permissions problem). This is distinct from the Proxmox check above — PVE only knows a backup *ran*, PBS is the only place that knows whether the stored data still verifies. | Not configurable |
## State-based alerts
These two run on a fixed **15-minute interval** (matching the agent's
own default report interval) — not the daily reminder time above, and
not user-configurable. Unlike the daily reminders, these are
**state-based**: a problem is announced once when it first appears, and
once more when it clears — not repeated on every 15-minute pass while it
continues. A condition that can't currently be read (an integration
that's down, a server whose agent hasn't reported) is held exactly as it
was rather than cleared or re-announced, so a temporary read failure
can't fake a recovery. On first-ever run (or right after upgrading to
this feature), whatever's already failing is recorded silently as the
starting baseline rather than announced all at once.
| Check | "Notify on" toggle | Fires when | Configured at |
|---|---|---|---|
| Health — server offline | *Server offline, disk nearly full, or Synology volume/disk problem* | A server's agent hasn't reported for longer than the offline threshold. Held for the first 20 minutes after this app restarts, since agents haven't had a chance to report yet. | Settings → Notifications → Health checks → **Server offline after** (default 60 min) |
| Health — disk usage | *(same toggle)* | A server disk, a Proxmox node's root filesystem or any of its storages, or a Synology volume, is at or above the usage threshold. A server that's currently offline isn't also judged on disk usage (its figures are stale). | Settings → Notifications → Health checks → **Disk usage alert at** (default 90%) |
| Health — Synology status | *(same toggle)* | A Synology volume's status isn't "normal", or a disk's status/SMART result isn't normal, or it's over the bad-sector or under the remaining-life threshold. | Not configurable |
| Automation — failed run | *Semaphore template or Gitea workflow run failed* | A Semaphore template's, or a Gitea repo's, most recent run status is a clear failure. A run that's still going, was cancelled, or was stopped by hand doesn't count as failed or as a recovery — it's left exactly as it was. For Gitea this follows the repo's most recent run on any workflow or branch. | Not configurable |
## Real-time alerts
These fire immediately, as the triggering action happens — not on any
schedule.
| Event | "Notify on" toggle | Fires when |
|---|---|---|
| DNS record added | *DNS record added* | A DNS record is created against any DNS provider through this app (manually, or via any automated action that writes one). |
| DNS record updated | *DNS record updated* | An existing DNS record is edited. |
| DNS record deleted | *DNS record deleted* | A DNS record is deleted. |
| Integration/DNS provider down | *Integration/DNS provider failing repeatedly* | Any integration or DNS provider adapter's calls fail a configurable number of times **in a row**. Tracked per adapter *type* (e.g. "proxmox"), not per individual integration row — with two Proxmox integrations, a streak of failures on either one counts toward the same total. Only fires once per failure streak (not on every failure past the threshold), and is skipped entirely if that type is currently in a maintenance window. | Settings → Notifications → **Alert after** (default 3 consecutive failures) |
| Integration/DNS provider recovered | *(same toggle)* | The next call for that source succeeds, after a "down" alert was already sent for the current streak. |
## Digest
| Notification | Fires when |
|---|---|
| "Notifications from quiet hours" | Once, at quiet hours' configured end time, **only if enabled and only if at least one notification was held** during the window. Bundles every held notification's title and message into one message, then clears the queue. |
## The Alerts page
**Operations → Alerts** shows what these notifications are about — the problems that exist *right now* — as a list you can look at, filter, and
export. It uses the same checks as the notifications above, so a problem appears there for exactly the reason it would be notified, but it differs
in three ways:
- It ignores the "Notify on" toggles. Turning a notification off doesn't hide the problem from the page.
- Problems under a **maintenance window** are kept on the list, marked *silenced* and counted separately, instead of being dropped.
- It also lists things nothing notifies about: Uptime Kuma monitors that are down, osTicket tickets that are overdue, and any integration that's
failing its last few calls (before the threshold that triggers an "integration down" notification).
It runs the checks live when opened (a recent result is reused for a minute), and shows what it couldn't read at the top, so a missing section means
"couldn't check" and not "all clear".
## What does *not* send a notification
Worth calling out explicitly, since it's easy to assume everything in
the app alerts on something:
- **osTicket** — the integration lists open tickets, but there is no
scheduled check or alert wired up for it (e.g. nothing pings you about
a newly-overdue ticket). It's a read-only dashboard/page today.
- **phpIPAM sync**, **Uptime Kuma monitor/maintenance import**, and any
other manual "Sync from X" action — these run when you click the
button and report their result on the page itself, not via a
notification.
- **Test buttons** (Settings → Notifications → each channel's "Send
test") send a one-off message through that one channel only, to
confirm it's wired up correctly — not a real event and not affected by
quiet hours.
- **Consistency reports**, **Diagnostic Log**, and **Audit Log** are all
read-only views you check yourself; none of them push a notification
on their own (the Diagnostic Log's failures are what feed the
integration-down alert above, but reading the log itself never
triggers anything).