Files
Homelab-manager/NOTIFICATIONS.md
T
bobbanandClaude Sonnet 5 4e48377348 Document every notification this app sends in NOTIFICATIONS.md
Covers all 16 notification types (daily reminders, state-based health/
automation alerts, real-time DNS and integration-down alerts, and the
quiet-hours digest), what triggers each, and what does and doesn't
respect quiet hours and maintenance windows. Every trigger condition
and threshold was cross-checked against the actual scheduler/monitor
code, not just the settings labels.

Also fixes the "Daily reminder time" hint text in NotificationSettings,
which had gone stale — it listed only 4 of the 6 checks that actually
share that schedule (missing domain expiry and PBS verification).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-29 23:48:19 +02:00

113 lines
9.1 KiB
Markdown

# Notifications
Every notification in Homelab Manager is sent through the same pipe
(`server/src/services/notify.ts`) to whichever channels are enabled under
**Settings → Notifications**: Gotify, ntfy, SMTP (email), and a generic
JSON webhook. All four fire for every notification below — there's no
per-event channel routing, only a per-event on/off toggle (the "Notify
on" list on that same page) and, for the daily ones, one shared
time/timezone.
Two settings apply to *every* notification regardless of what triggered
it:
- **Quiet hours** (Settings → Notifications → Quiet hours): while
enabled and inside the configured window, a notification is held
instead of sent immediately, then delivered as a single digest at the
window's end time. This matters most for the real-time alerts below —
the daily reminders already fire once at a time you pick, usually
outside the window anyway.
- **Maintenance windows** (the Maintenance page): opening one for a
server, integration, or DNS provider silences the alerts that target
it specifically (offline/disk-full for a server; storage, backup, and
automation-failure alerts for an integration; "integration down" for
that whole service type) — see the Maintenance page's own "What gets
silenced" panel for the exact list.
Where a trigger has a configurable threshold, that's noted in its row
below; several are fixed and can't be changed from the UI.
## Daily reminders
These six checks share one schedule: **Settings → Notifications → Daily
reminder time** (default 08:00) and **Timezone** (default UTC). Unlike
the state-based alerts further down, these **re-send every day the
condition is still true** — there's no "only once" de-duplication, so an
expired secret you haven't renewed yet will be mentioned again at the
next day's check, and the day after that.
Each one also runs once at server startup if it hasn't already run
today (e.g. after an upgrade or a period offline), so you're not waiting
until the next scheduled time to catch up.
| Check | "Notify on" toggle | Fires when | Configured at |
|---|---|---|---|
| Secret expiry | *Secret expiry reminder* | Any tracked secret (API token, SSL cert, password, generic) is expired or within its own configured warning window. SSL certificates with a host:port are re-checked live first, so this reflects the real current expiry, not a stale saved date. A certificate that couldn't be re-checked live is mentioned separately, so the expiry shown may be stale. | Per-secret warning threshold, set when adding/editing that secret |
| Domain expiry | *Domain registration expiring or expired (daily reminder)* | Any tracked domain registration is expired or within the warning window; every domain is re-looked-up against its registry first. A domain whose registry lookup has been failing for 3+ days is also mentioned, separately, as "may be stale." Registries that don't publish an expiry (e.g. `.de`, `.eu`) are tracked but never trigger this. | Settings → Notifications → Health checks → **Domain expiry warning** (default 30 days) |
| Tailscale key expiry | *Tailscale key expiry reminder* | Any device's node key (across every enabled Tailscale integration) is within 30 days of expiring, or already expired. Devices with key expiry disabled are skipped. | Fixed at 30 days |
| Docker image updates | *Docker image update available* | Any container (across every enabled Dockhand integration) has an image update available, per Dockhand's own cached update-check results — this doesn't trigger a fresh registry lookup, just reads what the Docker page itself would show. | Not configurable (reflects Dockhand's own check interval) |
| Proxmox backups | *Proxmox backup failed or a guest has no coverage* | Two independent conditions, both under this one toggle: (1) a node's most recent `vzdump` backup task (across every enabled Proxmox integration) didn't succeed; (2) a VM/LXC isn't covered by any enabled backup job at all. | Not configurable |
| Proxmox Backup Server verification | *Proxmox Backup Server snapshot failed verification* | Across every enabled PBS integration: any datastore has at least one stored snapshot whose verification state is "failed", or a datastore couldn't be read at all (e.g. a permissions problem). This is distinct from the Proxmox check above — PVE only knows a backup *ran*, PBS is the only place that knows whether the stored data still verifies. | Not configurable |
## State-based alerts
These two run on a fixed **15-minute interval** (matching the agent's
own default report interval) — not the daily reminder time above, and
not user-configurable. Unlike the daily reminders, these are
**state-based**: a problem is announced once when it first appears, and
once more when it clears — not repeated on every 15-minute pass while it
continues. A condition that can't currently be read (an integration
that's down, a server whose agent hasn't reported) is held exactly as it
was rather than cleared or re-announced, so a temporary read failure
can't fake a recovery. On first-ever run (or right after upgrading to
this feature), whatever's already failing is recorded silently as the
starting baseline rather than announced all at once.
| Check | "Notify on" toggle | Fires when | Configured at |
|---|---|---|---|
| Health — server offline | *Server offline, disk nearly full, or Synology volume/disk problem* | A server's agent hasn't reported for longer than the offline threshold. Held for the first 20 minutes after this app restarts, since agents haven't had a chance to report yet. | Settings → Notifications → Health checks → **Server offline after** (default 60 min) |
| Health — disk usage | *(same toggle)* | A server disk, a Proxmox node's root filesystem or any of its storages, or a Synology volume, is at or above the usage threshold. A server that's currently offline isn't also judged on disk usage (its figures are stale). | Settings → Notifications → Health checks → **Disk usage alert at** (default 90%) |
| Health — Synology status | *(same toggle)* | A Synology volume's status isn't "normal", or a disk's status/SMART result isn't normal, or it's over the bad-sector or under the remaining-life threshold. | Not configurable |
| Automation — failed run | *Semaphore template or Gitea workflow run failed* | A Semaphore template's, or a Gitea repo's, most recent run status is a clear failure. A run that's still going, was cancelled, or was stopped by hand doesn't count as failed or as a recovery — it's left exactly as it was. For Gitea this follows the repo's most recent run on any workflow or branch. | Not configurable |
## Real-time alerts
These fire immediately, as the triggering action happens — not on any
schedule.
| Event | "Notify on" toggle | Fires when |
|---|---|---|
| DNS record added | *DNS record added* | A DNS record is created against any DNS provider through this app (manually, or via any automated action that writes one). |
| DNS record updated | *DNS record updated* | An existing DNS record is edited. |
| DNS record deleted | *DNS record deleted* | A DNS record is deleted. |
| Integration/DNS provider down | *Integration/DNS provider failing repeatedly* | Any integration or DNS provider adapter's calls fail a configurable number of times **in a row**. Tracked per adapter *type* (e.g. "proxmox"), not per individual integration row — with two Proxmox integrations, a streak of failures on either one counts toward the same total. Only fires once per failure streak (not on every failure past the threshold), and is skipped entirely if that type is currently in a maintenance window. | Settings → Notifications → **Alert after** (default 3 consecutive failures) |
| Integration/DNS provider recovered | *(same toggle)* | The next call for that source succeeds, after a "down" alert was already sent for the current streak. |
## Digest
| Notification | Fires when |
|---|---|
| "Notifications from quiet hours" | Once, at quiet hours' configured end time, **only if enabled and only if at least one notification was held** during the window. Bundles every held notification's title and message into one message, then clears the queue. |
## What does *not* send a notification
Worth calling out explicitly, since it's easy to assume everything in
the app alerts on something:
- **osTicket** — the integration lists open tickets, but there is no
scheduled check or alert wired up for it (e.g. nothing pings you about
a newly-overdue ticket). It's a read-only dashboard/page today.
- **phpIPAM sync**, **Uptime Kuma monitor/maintenance import**, and any
other manual "Sync from X" action — these run when you click the
button and report their result on the page itself, not via a
notification.
- **Test buttons** (Settings → Notifications → each channel's "Send
test") send a one-off message through that one channel only, to
confirm it's wired up correctly — not a real event and not affected by
quiet hours.
- **Consistency reports**, **Diagnostic Log**, and **Audit Log** are all
read-only views you check yourself; none of them push a notification
on their own (the Diagnostic Log's failures are what feed the
integration-down alert above, but reading the log itself never
triggers anything).