Commit Graph
11 Commits
Author SHA1 Message Date
bobbanandClaude Sonnet 5 1688de3ea2 Alert when a server goes silent, a disk fills up, or a Synology volume degrades
The data was all being collected (agent last-seen, per-disk usage,
Proxmox storage, Synology volume/disk health) but nothing acted on
it, so a dead server or a full disk was only noticed by opening the
right page.

A new health pass runs every 15 minutes (matching the agent's default
report interval) and raises one notification when a problem starts and
one when it clears: a server's agent silent past a threshold (default
60 min), a server disk / Proxmox storage or root filesystem / Synology
volume at or above a usage threshold (default 90%), and a Synology
volume or disk that isn't "normal", has bad SMART, bad sectors past the
threshold, or life remaining below it. Both thresholds and an on/off
toggle live under Settings -> Notifications.

The parts that make this trustworthy rather than noisy:
- A problem is keyed by identity, so it alerts once and not every run;
  a shared Proxmox storage listed by every node is one problem, not
  one per node.
- Active problems persist across restarts, so a rebuild doesn't
  re-alert everything already known.
- If a source can't be read on a given run (Proxmox/Synology
  unreachable, one node lacking privileges) its existing problems are
  held, not reported "cleared" and then re-alerted when it comes back —
  the integration-failure alert already owns "the integration is down".
- For 20 minutes after startup server-derived problems are held too:
  agents couldn't report while the app was down, so judging them then
  would report every server offline after any restart.
- An offline server's disk figures are stale and are not judged; a
  server that never reported has no agent and raises nothing.
- Tracking continues while the toggle is off (only sending is gated),
  so turning it back on doesn't dump every long-standing problem.

Timestamps without a zone (SQLite's format) are read as UTC; the test
runs on a UTC+2 machine, where reading them as local time gives a
different answer.

Verified with 32 checks: the evaluation rules and the state diff as
pure functions (exact thresholds, the proxmox:1 vs proxmox:10 prefix
trap, the flapping sequence), then a whole pass against a real
Proxmox adapter talking to a fake HTTPS cluster (one node returning
403, the whole API down, a shared storage on two nodes, a node that
recovers), a webhook receiver, the real DB, and the persisted state.

Not exercised end-to-end: the Synology collection path — its rules are
tested on data shaped exactly like the adapter's output types, but I
did not stand up a fake DSM. Real dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 02:15:19 +02:00
bobbanandClaude Sonnet 5 73d649377e Add notification quiet hours with digest delivery
Instant notifications (DNS changes, integration failure/recovery
alerts) had no way to avoid pinging overnight. Adds a "Quiet hours"
window under Settings -> Notifications: notifications that would fire
during the window are held in a new notification_queue table instead
of sent immediately, then delivered as one combined digest at the end
time (in the same timezone already used for the daily checks) via a
new scheduled flush job. The scheduled daily checks (secret expiry,
Tailscale key, Docker updates, Proxmox backups) already only fire
once at a chosen time, so this mainly matters for the instant ones.
Includes a live "N queued" indicator with a manual "Flush now" button
for visibility, and correctly falls back to sending immediately
whenever the feature is disabled (the default).

Gating lives at the single choke point every notification already
flows through (notify()), so no per-event-type wiring was needed.

Verified against a fake webhook receiver on an isolated scratch
database: exhaustively checked the midnight-wraparound window math
(9 cases including exact-boundary inclusive/exclusive edges) against
synthetic "now" values rather than depending on when the test
happens to run, then end-to-end through the real notify()/
flushQuietHoursQueue() functions with a window constructed around the
actual current time — confirmed a notification during the window
queues instead of sending, the flush produces one digest with the
original title/message intact and clears the queue, a notification
outside the window sends immediately, and disabling the feature
entirely sends immediately regardless of the window. Confirmed the
real dev database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 20:56:33 +02:00
bobbanandClaude Sonnet 5 99db7e1cf0 Surface Proxmox backup job status, with a daily failure notification
Proxmox already runs vzdump backups, but nothing in the app said
whether they were actually succeeding — a silent backup failure is
one of the more dangerous blind spots a homelab admin can have. Adds
a "Backups" card to the Proxmox page: configured backup job
schedules (storage target, which guests, enabled/disabled) from
GET /cluster/backup, and recent vzdump task history per node from
GET /nodes/{node}/tasks?typefilter=vzdump, with a banner at the top
if the most recent run didn't succeed.

New "Proxmox backup failed" notification toggle under Settings ->
Notifications, on the same daily schedule as the other checks. The
scheduler checks each node's own most-recent vzdump run independently
(not just the single most recent task overall) so one node's healthy
backup can't mask another node's failing one in a multi-node cluster.

Known limitation, documented in the adapter's own header comment:
Proxmox's task list doesn't reliably expose which specific guest
failed within an "all guests" job — only the task's own log text has
that — so this surfaces job- and task-level status rather than
guessing at per-guest outcomes.

Verified against a fake Proxmox server (real self-signed HTTPS, since
the adapter's node:https usage can't be monkey-patched under ESM)
reproducing the documented /cluster/backup and task-list response
shapes: job parsing (all-guests+exclude vs specific-vmids+disabled)
correct, task OK/failure parsing correct, and the critical multi-node
scenario confirmed — one node's failing latest run flagged, the
other's healthy latest run correctly left alone, with exactly one
notification of the right content. This reproduces Proxmox's
documented API shape rather than a live-verified one; flag if the
real cluster's response differs in some way this didn't anticipate.
Confirmed the real dev database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 18:42:22 +02:00
bobbanandClaude Sonnet 5 1cae35a59e Add a daily Docker image-update notification
Dockhand's pending-update counts were only visible if you happened to
open the Docker page. New "Docker image update available" toggle
under Settings -> Notifications, sharing the same daily
time/timezone as the secret and Tailscale key expiry reminders (same
node-schedule reschedule-on-settings-change pattern as those two).
Reads each enabled Dockhand integration's already-cached update-check
results via listContainers() rather than triggering a fresh
per-container registry lookup, so it costs nothing extra beyond what
the Docker page itself already fetches, and lists every container
with an update pending across all environments/integrations in one
notification.

Verified end-to-end against a fake local Dockhand server (one
environment, two containers, one flagged with a pending update) and a
fake webhook receiver on an isolated scratch database: the check
correctly found only the flagged container (with its newerVersion),
sent exactly one notification with the right title/content, excluded
the up-to-date container, and sent nothing at all when the setting
was toggled off despite still finding the same pending update.
Confirmed the real dev database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 23:44:47 +02:00
bobbanandClaude Sonnet 5 b5a4c6e2d9 Alert when an integration or DNS provider fails repeatedly
The Diagnostic Log already records every outbound call's success or
failure, but nothing acted on it — you'd only notice an integration
was down by happening to open its page. Adds a per-source consecutive-
failure counter (in-memory, reset on restart, same durability tier as
the diag log's own ring buffer) hooked into recordDiagEntry: crossing
the configurable threshold (default 3) sends one "down" notification
on every configured channel, and a "recovered" notification fires once
it succeeds again — no repeat spam while it stays down. New
"Integration/DNS provider failing repeatedly" toggle and threshold
field under Settings -> Notifications.

Verified end-to-end against an isolated scratch database with a real
local HTTP server standing in for the webhook channel: 5 consecutive
failures produced exactly one "Down" notification (at the 3rd
failure, correctly naming "3 calls"), a subsequent success produced
exactly one "Recovered" notification, and two more failures on a
fresh streak triggered nothing (below threshold) — confirmed the real
dev database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 14:51:00 +02:00
bobbanandClaude Sonnet 5 10d123b18a Add automatic retention purging for the Diagnostic and Audit logs
The diagnostic log already rings-buffer to 500 rows, but the audit
log had no cap at all and would grow forever. Adds an opt-in
age-based purge under Settings -> Logs: keep entries for N days,
checked on a configurable interval (hourly through monthly), plus a
manual "Purge now" button. Reuses the existing node-schedule-style
reschedule-on-settings-change pattern from the secret/Tailscale
expiry checkers, but as a plain setInterval since "how often" here is
an interval rather than a specific daily time.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 02:22:32 +02:00
bobbanandClaude Sonnet 5 b40234a557 Make table page size a configurable setting, not a hardcoded 20
Pagination just landed hardcoded to 20 rows everywhere; add a way to
change that instead of leaving it fixed for every table in the app.

pageSize joins dateFormat/timeFormat on the existing display settings
object (server-side default 20, 5-500 range enforced by the PUT
schema) rather than becoming its own settings section, since it's the
same kind of thing -- an admin-configured, globally-applied display
preference read by every signed-in role via the already-public GET
/api/settings/display endpoint, same as the date/time format already
works.

Renamed the "Date & Time" settings tab/page/route to "Display" (still
just one component, now covering both date/time format and table
pagination) since its scope no longer matches the old name -- kept
Settings.tsx's usual pattern of one page per concern rather than
adding a second, oddly-scoped tab just for one number field.

New web/src/utils/pageSize.ts mirrors utils/date.ts's existing
module-level "set once at startup, read anywhere without prop-
drilling" pattern; usePagination()'s pageSize parameter now defaults
to getPageSize() instead of a literal 20, evaluated fresh on every
call so it picks up a saved change without touching any of the ten
pages already using the hook.

Verified server-side against a temp SQLite DB: pageSize defaults to
20, a partial update sets it without disturbing dateFormat/timeFormat
and vice versa, and it persists across a fresh settings read. Also
checked the default-parameter mechanics directly (re-evaluates the
global value on every call rather than capturing it once, and an
explicit override still wins).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-18 23:58:14 +02:00
bobbanandClaude Sonnet 5 1ff59afb40 Notify on expiring Tailscale device keys
The Tailscale page already showed per-device key expiry; extend the
existing daily-reminder infrastructure (currently only for Secrets)
to push it out through the configured notification channels too, the
same way expiring secrets already are.

New tailscaleKeyExpiryScheduler.ts mirrors secretExpiryScheduler.ts:
runs once at startup (skipped if already run today) and daily
thereafter, checking every enabled Tailscale integration's devices for
keys expiring within the warning window and calling notify() with the
results. Reuses the exact same daily time/timezone setting as the
secret-expiry check (one "Daily reminder time" control, two
independent on/off toggles) rather than adding a second schedule for
users to configure.

Centralized the expiring-soon threshold and check (previously only
duplicated in the /synology and /tailscale route summaries) into
adapter.ts as `KEY_EXPIRY_WARN_DAYS` / `isKeyExpiringSoon()`, and
updated the devices route to use it instead of its own inline copy.

New `tailscaleKeyCheck` notification-event toggle (default on) in
settings, alongside the existing secret-expiry one.

Verified end-to-end against a temp SQLite DB + real migrations: a
tailscale integration pointed at a mock Tailscale API (one device
expiring in 10 days, one with key-expiry disabled) with the webhook
channel enabled and pointed at a mock receiver — confirmed the
scheduler's startup check queries the DB correctly, decrypts the
integration's credential, calls the adapter, filters out the
disabled-expiry device, and delivers a webhook payload naming only
the expiring device.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 18:59:27 +02:00
bobbanandClaude Sonnet 5 f92f8de96e Add a Date & Time setting and fix inconsistent date formats app-wide
Different tables used different date formats: Servers & Tasks used a
fixed "YYYY-MM-DD HH:mm:ss" (24h), while Audit Log, DNS, Integrations
(Tailscale/Semaphore), and Users called plain .toLocaleString() with
no options, which renders using the browser's own locale — different
per browser/OS, and inconsistent with the other pages' fixed format.

- New Settings -> Date & Time page: pick date order (YYYY-MM-DD,
  DD/MM/YYYY, MM/DD/YYYY) and 12h vs 24h clock, with a live preview.
- utils/date.ts's formatDateTime() now reads these settings instead of
  being hardcoded to sv-SE/24h; the setting is fetched once at app
  startup (alongside /api/me) via a new non-secret GET
  /api/settings/display (any signed-in user, same rationale as the
  badge-color endpoints) and applied immediately on save too, without
  needing a page reload.
- Switched every remaining raw new Date(...).toLocaleString() call
  (Audit Log, DNS zone sync time, Tailscale/Semaphore last-seen,
  Users' last login) over to the shared formatter, so every table now
  renders dates identically.

Verified the formatter's date-order x 12h logic against all six
combinations plus the midnight/noon 12h edge cases, and the settings
endpoints end-to-end (defaults, partial updates, validation
rejection, audit logging) against the real dev server.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-15 22:17:06 +02:00
bobbanandClaude Sonnet 5 ad6783b6a1 Extend badge color settings to cover integration types too
Settings → DNS Badges only let you customize DNS provider badge
colors. Expand it into a general Badges page with a second section
for the six integration types (Tailscale, Proxmox, Synology,
Semaphore, Gitea, Dockhand), applied to the type badge in the Manage
integrations table.

- New integrationColors key on AppSettings, stored/merged the same way
  as providerColors via the existing settingsStore.
- New GET /api/settings/integration-colors — non-secret, any signed-in
  user, mirroring /provider-colors — so the badge color can be read
  without needing admin access to the full settings payload.
- Renamed DnsBadgeSettings.tsx -> BadgeSettings.tsx (route
  /settings/dns-badges -> /settings/badges, sub-nav label "DNS
  Badges" -> "Badges") with both color sections saved together.

Verified end-to-end against the real dev server: PUT persists
integration colors independently of provider colors, the public
integration-colors endpoint reflects updates immediately, and the
change is audit-logged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-15 21:28:26 +02:00
bobbanandClaude Sonnet 5 3255314402 Build the Settings module: notification channels, event toggles, DNS badge colors
The Settings page was a "coming soon" placeholder. Port Sloth Manager's
settings feature set: Gotify/ntfy/SMTP/webhook notification channels
(each with its own test-send button), per-event toggles (DNS record
added/updated/deleted, a daily secret-expiry digest with configurable
time/timezone), and per-provider DNS badge color customization.

Settings persist in the existing `settings` key/value table via a new
settingsStore service; a notify service fans a message out to every
enabled channel. DNS record add/update/delete now fire notifications,
and a node-schedule job re-arms itself whenever the notification
settings change. Removed the now-superseded GOTIFY_URL/GOTIFY_TOKEN
env vars in favor of in-app configuration.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-15 12:14:05 +02:00