fae7089448802372998ed53c5f1867f560348464
13
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
447f33fff6 |
Add an Alerts page under Operations listing everything that's wrong now
One list of the current problems across servers and integrations, instead of waiting for a notification or visiting each page: servers that stopped reporting, full or nearly full disks and volumes (critical from 95%), Synology volume/disk problems, failed or uncovered Proxmox backups, failed Proxmox Backup Server verifications, container image updates, expired or expiring secrets/domains/Tailscale keys, failed Semaphore and Gitea runs, Uptime Kuma monitors that are down, overdue osTicket tickets, and integrations whose calls keep failing. Visible to every role, with severity and kind filters, search, sorting, CSV export and "Check now". It runs the same checks that send the notifications rather than a second copy of them: the detection in the health, automation, Proxmox backup, PBS, Docker update and Tailscale key checks is pulled out into shared collectors that both the schedulers and the page call, so the two can't disagree about what counts as a problem. Notification behaviour is unchanged, including the scheduled backup checks skipping integrations under a maintenance window. Unlike the notifications the page ignores the on/off toggles, and keeps problems under a maintenance window, marked silenced and counted apart. It reads live, so a result is reused for a minute (and Refresh can't re-run everything more than once every ten seconds), and every source has a 20 s limit so one hung integration can't hang the page. Anything it couldn't read is called out at the top instead of looking like all clear, and server checks pause for the same 20 minutes after a restart as the notifications do, with a note saying so. Also gives the newer integrations (PBS, osTicket, Uptime Kuma, phpIPAM) proper names in "integration down" notifications instead of their ids. Verified through the real routes against a scratch database with fake backends (offline and full-disk servers, secrets and domains, a silenced server, a fake PBS with failed verification, a hanging integration, a refused one, a failing-calls streak, caching, the restart grace period, auth), and by rendering the real page against that data in a browser: filters, search, silenced toggle, sorting, Check now, dark mode. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> |
||
|
|
df2a5ce42b |
Add a Proxmox Backup Server integration: datastore/snapshot verification status
Proxmox VE already shows whether the last vzdump push to PBS succeeded, but has no visibility into PBS's own backup verification, GC/prune health, or host status. This adds PBS as its own integration (own adapter, page, nav entry, and Dashboard widget) that reads datastore usage and, for every stored snapshot, its verification state directly from PBS. A new daily check (mirroring the existing Proxmox backup-failure check) notifies when a snapshot has failed verification or a datastore couldn't be read, with its own toggle in Settings -> Notifications and its own maintenance-window silencing. Not verified against a live PBS instance — built from PBS's published API docs and a scratch test against a mocked PBS server exercising the adapter's parsing and auth-header format (PBSAPIToken uses a colon separator, unlike PVE's PVEAPIToken which uses =). See INTEGRATIONS.md for details and the "not verified" caveat. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ca0fa817f8 |
Track domain registration expiry, with daily reminders
New Domains page listing when each domain registration expires, read from the registry. Domains behind the DNS zones already synced are picked up automatically; others can be added by hand. You're reminded daily from N days before expiry (Settings > Notifications, default 30) until it's renewed, and told when an expiry date hasn't been refreshable for several days so a stale date isn't trusted silently. RDAP alone would not have covered this homelab: .se, .nu, .io, .eu and .de are not in IANA's RDAP bootstrap. Lookups therefore try RDAP where the TLD publishes a server and fall back to WHOIS on port 43, found via IANA's own referral, parsing the expiry line out of the free-text answer. Only the expiry date and registrar are read or stored. Verified live against the real registries: .se and .nu via WHOIS, .com/.org/.dev via RDAP. Behaviour worth knowing: - A DNS zone that is a subdomain (lab.example.se) resolves to the registration that actually expires by trying the name and then its parents, so no public-suffix list is needed. Zones already covered by a tracked domain are not looked up again. - "Couldn't ask" is never confused with "not registered": network errors, rate limits and garbled answers are errors, and a transient error at any level stops the walk from concluding the domain doesn't exist. - A failed refresh keeps the last known expiry and records why, rather than blanking a date that's still relied on. - Zones that don't resolve to a real registration (.lan, .local, unregistered names) simply get no row. Zone-derived rows disappear when their zone does; manual rows stay. Zone-derived rows can't be deleted by hand. - Registries that don't publish an expiry (.de, .eu) are tracked with a note instead of a date. - Input like "example.com/path" is refused rather than silently reduced to its host. - Runs on the daily secret-expiry schedule and reminder time, on demand (Check all now / per domain), and once at startup if nothing has been read in a day. Never blocks startup, one lookup at a time with a pause. The warning window lives with the other thresholds in settings. New table domains (migration 0011); two settings fields (toggle and warning days). Verified with 90 checks against fake RDAP/WHOIS backends (name normalization, date formats, WHOIS parsing including rate-limit and no-expiry answers, bootstrap and referral caching, stale-cache fallback, parent walking, add/sync/check/refresh, concurrency guard, alert selection and stale detection, the daily notification and its toggle, role rules) plus a live smoke test against real registries and a browser check of the page against the real router. Real dev database mtime untouched. Not checked: a screenshot of the finished page (the capture timed out); structure, sorting, errors and the viewer view were verified. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
4118062405 |
Alert when a Semaphore template or Gitea workflow run fails
Every 15 minutes (on the existing health-check timer) the app reads the latest run of each Semaphore template and each Gitea repo's latest workflow run. A failed one raises one notification, and another when a later run succeeds. It is state-based like the server health alerts, so a job that fails every night alerts on the first failure, not every night. Gitea alerts include the run's link. Toggle: Settings > Notifications. What counts: - Semaphore "error" is a failure, "success" is a pass. A run that is waiting, running, stopped by hand or rejected is neither, so it leaves the previous state alone: a run in progress must not clear a failure it hasn't fixed yet, and a manual stop isn't a failure. - Gitea failure/success likewise; running, waiting, blocked, cancelled and skipped leave things as they were. - A failing template that gets another failing run does not re-alert. Not mistaking "couldn't read" for "fixed": - Semaphore's template listing swallowed per-project errors, so a project that failed to load looked like a project with no templates. A new checkTemplates adapter method reports which projects failed, and their failures are held rather than cleared. - Gitea reports a run it couldn't fetch as null, the same as "no runs"; both leave the repo's state alone. - An unreachable integration holds all of its failures. Nothing is cleared or re-announced while it is down. The first pass only records what is already failing without announcing it, so upgrading (or adding an integration to a fresh install) doesn't produce a wall of alerts about months-old failures. That baseline is not spent while nothing could be read. Maintenance windows on a Semaphore or Gitea integration silence its failure alerts with the same rules as the health alerts: a problem that starts during a window alerts when it ends, and one already announced stays known. The diff logic is reused from the health monitor rather than copied. Maintenance page text updated. Known limit: for Gitea this follows the repo's most recent run on any workflow or branch, matching what the Gitea page shows; a failure in one workflow can be masked by a later success of another. Verified with 53 checks against fake Semaphore and Gitea servers and a webhook receiver: classification, baseline (including not being consumed when nothing is readable), single alert per failure, no repeat, in-progress/ stopped/cancelled runs, recovery and re-failure, unreadable project, unreadable integration, run-fetch errors, maintenance windows (silenced, then announced after), the toggle, disabled integrations and repos without Actions. Real dev database mtime untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1688de3ea2 |
Alert when a server goes silent, a disk fills up, or a Synology volume degrades
The data was all being collected (agent last-seen, per-disk usage, Proxmox storage, Synology volume/disk health) but nothing acted on it, so a dead server or a full disk was only noticed by opening the right page. A new health pass runs every 15 minutes (matching the agent's default report interval) and raises one notification when a problem starts and one when it clears: a server's agent silent past a threshold (default 60 min), a server disk / Proxmox storage or root filesystem / Synology volume at or above a usage threshold (default 90%), and a Synology volume or disk that isn't "normal", has bad SMART, bad sectors past the threshold, or life remaining below it. Both thresholds and an on/off toggle live under Settings -> Notifications. The parts that make this trustworthy rather than noisy: - A problem is keyed by identity, so it alerts once and not every run; a shared Proxmox storage listed by every node is one problem, not one per node. - Active problems persist across restarts, so a rebuild doesn't re-alert everything already known. - If a source can't be read on a given run (Proxmox/Synology unreachable, one node lacking privileges) its existing problems are held, not reported "cleared" and then re-alerted when it comes back — the integration-failure alert already owns "the integration is down". - For 20 minutes after startup server-derived problems are held too: agents couldn't report while the app was down, so judging them then would report every server offline after any restart. - An offline server's disk figures are stale and are not judged; a server that never reported has no agent and raises nothing. - Tracking continues while the toggle is off (only sending is gated), so turning it back on doesn't dump every long-standing problem. Timestamps without a zone (SQLite's format) are read as UTC; the test runs on a UTC+2 machine, where reading them as local time gives a different answer. Verified with 32 checks: the evaluation rules and the state diff as pure functions (exact thresholds, the proxmox:1 vs proxmox:10 prefix trap, the flapping sequence), then a whole pass against a real Proxmox adapter talking to a fake HTTPS cluster (one node returning 403, the whole API down, a shared storage on two nodes, a node that recovers), a webhook receiver, the real DB, and the persisted state. Not exercised end-to-end: the Synology collection path — its rules are tested on data shaped exactly like the adapter's output types, but I did not stand up a fake DSM. Real dev database mtime untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
7e81306aa7 |
Read SSL certificate expiry from the live server instead of trusting a typed-in date
A certificate secret's expiry was only ever what someone typed in, so a renewed cert (or a wrong date) meant the app's reminders were silently wrong. A certificate secret can now be given a host:port; the app opens a real TLS connection and reads the certificate's actual expiry — on create/edit (if the host changes), daily, and via a per-row "Check now" — and keeps expiryDate in sync. Because the daily refresh runs before the existing expiry check, the reminder is always computed from what's actually being served. Verification is deliberately off for the connection: homelab services routinely serve self-signed/internal-CA certs, and an already-expired one is exactly the case worth reporting, which a verifying connection would refuse before exposing the dates. Failure handling avoids the silent-staleness this is meant to fix: a failed check keeps the last known date, records why on the row (shown as a "Check failed" badge), and is listed in the daily secrets notification. Creating a monitored secret whose host can't be reached and with no manual date is rejected with the reason rather than saved blank. A non-TLS port (the likeliest typo) gets a plain-language error instead of raw OpenSSL output. Server-side connections to a user-supplied host:port need the same operator role that already gates editing secrets (and running Semaphore templates, which is strictly more powerful); the host is validated against a strict character set before any connection is made. New nullable secrets columns (check_host, check_port, last_checked_at, last_check_error) via migration 0007; existing rows are unaffected. Verified against real TLS servers (openssl-generated certs) and the real secrets router with a stubbed session: a live 45-day cert read back as the correct date via both an IP host (no SNI) and a hostname; an already-expired cert reported its past date and shows as expired; refused connections, a server that accepts but never answers (times out), and a plain non-TLS server each produced a descriptive error rather than a hang or crash. Through the router: create with a host and no date reads the date; unreachable host with no date -> 400 with the reason; unreachable with a manual date -> saved with the error recorded; host on a non-certificate type and an invalid host string -> 400; a hand-typed date on a monitored secret is ignored; changing the host re-checks immediately; changing the type away from certificate ends monitoring; a viewer gets 403 on Check now. 23 checks, all passing (a first re-run showed 2 spurious failures that were leftover rows from the previous run's scratch database, confirmed by a clean re-run). Real dev database mtime untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ca61f2a915 |
Surface Proxmox VMs/LXCs with no backup coverage at all
A failing backup run is visible now, but a guest with no backup job covering it in the first place was still a silent gap. Rather than depending on Proxmox's /cluster/backup-info/not-backed-up-guests endpoint (only exists on newer PVE versions), this derives coverage from data already fetched: a guest counts as covered if any enabled job either lists its vmid directly, or backs up "all guests" (scoped to the job's node, if it has one) without excluding it. Adds a warning banner plus a full table to the Proxmox page's Backups card, and extends the existing daily "Proxmox backup failed" notification (relabeled to mention this too) to also list uncovered guests, gated by the same toggle. Verified the coverage logic directly (it's a pure function, so no fake server needed) across 7 cases: no jobs at all, an all-guests job with an exclude list, a specific-vmids job, a node-scoped job that shouldn't cover a guest on a different node, a disabled job providing no real coverage, two jobs whose combined scope covers everything neither would alone, and a realistic mixed scenario — all passed. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
73d649377e |
Add notification quiet hours with digest delivery
Instant notifications (DNS changes, integration failure/recovery alerts) had no way to avoid pinging overnight. Adds a "Quiet hours" window under Settings -> Notifications: notifications that would fire during the window are held in a new notification_queue table instead of sent immediately, then delivered as one combined digest at the end time (in the same timezone already used for the daily checks) via a new scheduled flush job. The scheduled daily checks (secret expiry, Tailscale key, Docker updates, Proxmox backups) already only fire once at a chosen time, so this mainly matters for the instant ones. Includes a live "N queued" indicator with a manual "Flush now" button for visibility, and correctly falls back to sending immediately whenever the feature is disabled (the default). Gating lives at the single choke point every notification already flows through (notify()), so no per-event-type wiring was needed. Verified against a fake webhook receiver on an isolated scratch database: exhaustively checked the midnight-wraparound window math (9 cases including exact-boundary inclusive/exclusive edges) against synthetic "now" values rather than depending on when the test happens to run, then end-to-end through the real notify()/ flushQuietHoursQueue() functions with a window constructed around the actual current time — confirmed a notification during the window queues instead of sending, the flush produces one digest with the original title/message intact and clears the queue, a notification outside the window sends immediately, and disabling the feature entirely sends immediately regardless of the window. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
99db7e1cf0 |
Surface Proxmox backup job status, with a daily failure notification
Proxmox already runs vzdump backups, but nothing in the app said
whether they were actually succeeding — a silent backup failure is
one of the more dangerous blind spots a homelab admin can have. Adds
a "Backups" card to the Proxmox page: configured backup job
schedules (storage target, which guests, enabled/disabled) from
GET /cluster/backup, and recent vzdump task history per node from
GET /nodes/{node}/tasks?typefilter=vzdump, with a banner at the top
if the most recent run didn't succeed.
New "Proxmox backup failed" notification toggle under Settings ->
Notifications, on the same daily schedule as the other checks. The
scheduler checks each node's own most-recent vzdump run independently
(not just the single most recent task overall) so one node's healthy
backup can't mask another node's failing one in a multi-node cluster.
Known limitation, documented in the adapter's own header comment:
Proxmox's task list doesn't reliably expose which specific guest
failed within an "all guests" job — only the task's own log text has
that — so this surfaces job- and task-level status rather than
guessing at per-guest outcomes.
Verified against a fake Proxmox server (real self-signed HTTPS, since
the adapter's node:https usage can't be monkey-patched under ESM)
reproducing the documented /cluster/backup and task-list response
shapes: job parsing (all-guests+exclude vs specific-vmids+disabled)
correct, task OK/failure parsing correct, and the critical multi-node
scenario confirmed — one node's failing latest run flagged, the
other's healthy latest run correctly left alone, with exactly one
notification of the right content. This reproduces Proxmox's
documented API shape rather than a live-verified one; flag if the
real cluster's response differs in some way this didn't anticipate.
Confirmed the real dev database's mtime was untouched throughout.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
1cae35a59e |
Add a daily Docker image-update notification
Dockhand's pending-update counts were only visible if you happened to open the Docker page. New "Docker image update available" toggle under Settings -> Notifications, sharing the same daily time/timezone as the secret and Tailscale key expiry reminders (same node-schedule reschedule-on-settings-change pattern as those two). Reads each enabled Dockhand integration's already-cached update-check results via listContainers() rather than triggering a fresh per-container registry lookup, so it costs nothing extra beyond what the Docker page itself already fetches, and lists every container with an update pending across all environments/integrations in one notification. Verified end-to-end against a fake local Dockhand server (one environment, two containers, one flagged with a pending update) and a fake webhook receiver on an isolated scratch database: the check correctly found only the flagged container (with its newerVersion), sent exactly one notification with the right title/content, excluded the up-to-date container, and sent nothing at all when the setting was toggled off despite still finding the same pending update. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b5a4c6e2d9 |
Alert when an integration or DNS provider fails repeatedly
The Diagnostic Log already records every outbound call's success or failure, but nothing acted on it — you'd only notice an integration was down by happening to open its page. Adds a per-source consecutive- failure counter (in-memory, reset on restart, same durability tier as the diag log's own ring buffer) hooked into recordDiagEntry: crossing the configurable threshold (default 3) sends one "down" notification on every configured channel, and a "recovered" notification fires once it succeeds again — no repeat spam while it stays down. New "Integration/DNS provider failing repeatedly" toggle and threshold field under Settings -> Notifications. Verified end-to-end against an isolated scratch database with a real local HTTP server standing in for the webhook channel: 5 consecutive failures produced exactly one "Down" notification (at the 3rd failure, correctly naming "3 calls"), a subsequent success produced exactly one "Recovered" notification, and two more failures on a fresh streak triggered nothing (below threshold) — confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1ff59afb40 |
Notify on expiring Tailscale device keys
The Tailscale page already showed per-device key expiry; extend the existing daily-reminder infrastructure (currently only for Secrets) to push it out through the configured notification channels too, the same way expiring secrets already are. New tailscaleKeyExpiryScheduler.ts mirrors secretExpiryScheduler.ts: runs once at startup (skipped if already run today) and daily thereafter, checking every enabled Tailscale integration's devices for keys expiring within the warning window and calling notify() with the results. Reuses the exact same daily time/timezone setting as the secret-expiry check (one "Daily reminder time" control, two independent on/off toggles) rather than adding a second schedule for users to configure. Centralized the expiring-soon threshold and check (previously only duplicated in the /synology and /tailscale route summaries) into adapter.ts as `KEY_EXPIRY_WARN_DAYS` / `isKeyExpiringSoon()`, and updated the devices route to use it instead of its own inline copy. New `tailscaleKeyCheck` notification-event toggle (default on) in settings, alongside the existing secret-expiry one. Verified end-to-end against a temp SQLite DB + real migrations: a tailscale integration pointed at a mock Tailscale API (one device expiring in 10 days, one with key-expiry disabled) with the webhook channel enabled and pointed at a mock receiver — confirmed the scheduler's startup check queries the DB correctly, decrypts the integration's credential, calls the adapter, filters out the disabled-expiry device, and delivers a webhook payload naming only the expiring device. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
3255314402 |
Build the Settings module: notification channels, event toggles, DNS badge colors
The Settings page was a "coming soon" placeholder. Port Sloth Manager's settings feature set: Gotify/ntfy/SMTP/webhook notification channels (each with its own test-send button), per-event toggles (DNS record added/updated/deleted, a daily secret-expiry digest with configurable time/timezone), and per-provider DNS badge color customization. Settings persist in the existing `settings` key/value table via a new settingsStore service; a notify service fans a message out to every enabled channel. DNS record add/update/delete now fire notifications, and a node-schedule job re-arms itself whenever the notification settings change. Removed the now-superseded GOTIFY_URL/GOTIFY_TOKEN env vars in favor of in-app configuration. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |