bobbanandClaude Sonnet 5 aae4f0d74f Add maintenance mode to silence alerts while working on a server or integration
Rebooting Proxmox or patching a server triggered failure/offline alerts
you then had to dismiss. A maintenance window silences alerts about one
server, integration, or DNS provider for a chosen time. New Maintenance
page (start with a duration and optional reason, end early, see what's
silenced and what isn't) and a banner in the app shell so every signed-in
user can see what is currently silenced. Starting/ending is operator-only
and audit-logged; starting one on a target that already has a window
restarts its clock instead of stacking.

Silenced for the target: server offline/disk alerts, Proxmox/Synology
storage and health alerts, Proxmox backup alerts, and "integration down"
alerts. Not silenced: expiry and update reminders, DNS change notices.

The design goal is that this cannot hide a real outage:
- Every window has a required end (5 min to 7 days); there is no
  open-ended option, so a forgotten window expires by itself.
- A silenced problem is deliberately NOT recorded as "known". If it is
  still present when the window ends it alerts then, as new. A problem
  that was already alerted before the window stays known, so it isn't
  repeated, and is reported cleared only after the window ends.
- Failure alerts keep counting failures during a window without marking
  themselves alerted, so an outage that outlasts the window alerts on the
  very next failed call.

Known limitation, stated on the page: integration-failure alerts are
tracked per service TYPE (all "proxmox"), not per configured instance, so
a window on one Proxmox integration also silences a failure on a second
Proxmox integration while it's open. Fixing that means threading the
integration id through every adapter and the diagnostic log, which is a
much larger change than this feature.

Also moved the API-error-message helper out of Secrets.tsx into a shared
util now that two pages use it. New table maintenance_windows (migration
0008).

Verified with 44 checks: the condition-key-to-subject mapping (including
server:3 vs server:33), the diff rules with silenced subjects (new problem
not recorded, alerts when the window ends; already-known one carried and
not repeated; clears only after the window), window expiry and
integration/DNS-provider source matching, the failure tracker end to end
against a webhook (silent during a window while an unrelated service still
alerts; outage that outlasts the window alerts on the next failure and
only once; fail-and-recover fully inside a window sends nothing), a full
health pass against a real window, and the real router with a stubbed
session (role rules, duration bounds including the missing-duration case,
extend-not-stack, 404s, deleted targets hidden, audit entries). Real dev
database mtime untouched.

Not done: I haven't clicked through the new page or banner in a browser
(they sit behind the Authentik login); it builds and the API behind it is
tested.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 02:23:06 +02:00
2026-09-14 21:57:13 +02:00

Homelab Manager

Repository: git@10.200.5.13:bobban/Homelab-manager.git (gitea.labsconnect.se/bobban/Homelab-manager externally).

A single dashboard for a homelab: Proxmox, Synology DSM, Semaphore, Tailscale, Gitea, and Dockhand/Docker status and basic actions, plus DNS record management, an IP address inventory (IPAM), and a secret-expiry tracker (ported from Sloth Manager) and scheduled-task tracking across Debian/Raspbian hosts (ported from Schedule Task Manager). Looks and feels like a Tabler admin dashboard. Sign-in is delegated to Authentik (OIDC), with local admin/operator/viewer roles.

Status

All modules from the original plan are built:

  • Monorepo scaffold, Tabler-themed app shell/navigation

  • Authentik OIDC login, roles (first user to sign in becomes admin), audit log

  • Dashboard — an overview of every system this app tracks, all sharing one widget-card design (label + status badge, a small stat row, then its own breakdown): DNS (domain/record counts per provider, cached records by type), Secrets (monitored/expiring/expired, by type), and one widget per integration — Tailscale by OS, Proxmox by node (VM/LXC counts too), Dockhand by container state (plus host count), Semaphore and Gitea by last-run status (plus private-repo count), and Synology's CPU/RAM alongside its disk-health breakdown. Breakdowns render as a stacked proportion bar with a legend — no charting library, matching the rest of the app's plain-Tabler-CSS approach.

  • Diagnostic Log (admin-only) — every call this app makes to a DNS provider or integration (Tailscale, Proxmox, Synology, Semaphore, Gitea, Dockhand), success or failure, with latency and the error message if it failed — the last 500 calls, filterable by source/result, for troubleshooting connectivity issues (ported from Sloth Manager's provider-diagnostics log, generalized to cover every integration this app has, not just DNS)

  • Secrets — expiry tracking for API tokens/certs/passwords. An SSL certificate can optionally be given a host:port to watch: the app opens a real TLS connection (daily, and on demand via "Check now"), reads the certificate's actual expiry, and keeps the date current — so a renewed cert is picked up automatically and an unreachable host is flagged instead of silently going stale

  • IP Addresses (IPAM) — inventory of IPs across vendors/locations, with "Sync from Tailscale" and "Sync from Proxmox" actions to pull in tailnet device IPs and VM/LXC IPs (never overwrites a manually-entered IP), and each entry now shows its matching DNS record(s) from the DNS module's cache

  • DNS — zone/record management across Cloudflare, Loopia, Pi-hole, Azure DNS, cPanel, and Technitium; providers are configured in-app (not via env vars) and their credentials are encrypted at rest

  • Servers — cron/systemd tracking across Debian/Raspbian servers via a lightweight push agent (agent/linux/), plus manual entries for things an agent can't see (Docker jobs, backups). The Servers page itself just lists registered servers and (admin-only) adds new ones / issues agent tokens; clicking a server opens its detail page with CPU/RAM/disk status, IP addresses, matching DNS names (looked up from the DNS module's cache), and its scheduled tasks — live hardware from Proxmox for VM/LXC-backed servers, or from the agent's own hardware report for everything else, both showing the same per-disk usage breakdown (an LXC's root filesystem read straight from the host; a QEMU VM's actual mounts via its guest agent, alongside the allocated size Proxmox already knew about without one). Proxmox-linked servers also get start/stop/restart buttons right on the detail page. The detail page also has an Admin Links section (operator/admin to add/edit/remove) for bookmarking that server's own admin UIs — Dockge, Webmin, Cockpit, Portainer, or anything else reachable by URL. Since not every server is a Proxmox VM, an admin can hide the "Proxmox link" card per server ("Not a VM? Hide this" / "+ Show Proxmox link options") — it stays visible regardless once a server actually is linked, so unlinking is always reachable.

  • Tailscale, Proxmox, Synology, Semaphore, Gitea, and Docker each get their own top-level page (backed by the matching integration) instead of living inside a shared Integrations browsing view:

    • Tailscale — device list with online/authorized status, and authorize/deauthorize/remove actions; a live device-count widget.
    • Proxmox — VM/LXC status across every node in the cluster, with start/restart/shutdown/stop actions; a live running/total widget. Each online node also gets its own host-stats card — uptime, CPU usage/cores/ load average, RAM and swap usage, and per-storage usage (local, LVM-thin, ZFS, NFS, etc). Supports self-signed certificates (common in homelab setups).
    • Synology — volume and disk health (read-only by design). Supports self-signed certificates.
    • Semaphore — Ansible run status per template across every project, with a "Run" action to trigger a template; a live template-count widget (with a last-failed warning).
    • Gitea — repo list with each repo's last CI run status, and re-running just the failed jobs in a run; a live repo-count widget (with a failing-build warning).
    • Docker — container status across every Docker host Dockhand manages (one credential covers all of them), with start/stop/restart actions, a host filter, image-update status per container (from Dockhand's own cached update check, plus a button to trigger a fresh one), and a live running/total widget (with an updates-available count).

    The Integrations page itself is now just a list of configured integrations (name/type/status, visible to every role) with an admin-only "Add integration" button and edit/enable/disable/delete actions per row — the six dedicated pages above are where you actually use each one.

  • Every table in the app is click-to-sort on any column (numbers, booleans, and dates/text sort correctly regardless of how the column formats them) and has an "Export CSV" button next to it that exports whatever's currently sorted/filtered. Tables that can realistically grow large (DNS zones/records, IP Addresses, Secrets, Servers, Audit Log, and each integration's device/container/guest/repo/template list) are paginated, 20 rows per page by default — adjustable under Settings → Display — and CSV export still covers every sorted/filtered row, not just the current page.

  • Settings (admin-only) — notification channels (Gotify, ntfy, SMTP, generic webhook) with per-channel test buttons, per-event toggles (DNS record added/updated/deleted, daily secret-expiry reminder and daily Tailscale key-expiry reminder — both sharing one configurable time/timezone), badge-color customization for both DNS providers and integration types, and a Display tab (date order, 12/24-hour clock, and rows-per-page for every paginated table) applied consistently across the app.

All six integrations follow the same config-in-UI + encrypted-credentials pattern, added (and edited — e.g. to rotate an expired API token without recreating the whole integration) through Integrations → Manage integrations. See INTEGRATIONS.md for exactly what credential to create and what access it needs in each target system, for every integration and DNS provider.

Verified for real, end to end: every module above — including all six integrations, both their read-only views and their write actions (start/stop/restart, trigger-a-run, authorize/deauthorize) — has been exercised against the user's actual live homelab, not just built against specs. That pass also found and fixed two real bugs: the Synology adapter assumed HTTPS-only (the NAS is reached over plain HTTP), and the Tailscale adapter read online/isExitNode fields that don't actually exist in the real API response (fixed to derive them from connectedToControl and enabledRoutes). See the git log for the full verification notes per integration.

Server and storage health is watched every 15 minutes: a server whose agent stops reporting, a server disk / Proxmox storage / Synology volume passing a usage threshold, and a Synology volume or disk that's degraded or failing each raise one notification when the problem starts and one when it clears (both thresholds are set under Settings → Notifications). Active problems are remembered across restarts, so a rebuild doesn't re-alert them.

Maintenance mode silences alerts about one server, integration, or DNS provider while you work on it (server offline / disk, storage and Synology health, Proxmox backup alerts, and "integration down" for that service type). Every window has a fixed end (5 minutes to 7 days) and expires on its own, and a problem that began during a window and is still present when it ends alerts then — a forgotten window can't hide an outage. A banner shows what's currently silenced to every signed-in user.

The app is installable as a PWA — "Install app" / "Add to Home Screen" from the browser gives it its own icon and a standalone window on phone or desktop. This needs the site to be served over HTTPS (browsers only offer install on secure origins; localhost also counts). The bundled service worker deliberately caches nothing, so an installed copy always shows the current build.

Requirements

  • Node.js 20+
  • An Authentik instance reachable from wherever this app runs

1. Set up an Authentik application

  1. Create an OAuth2/OpenID Provider:
    • Redirect URI: <APP_BASE_URL>/auth/callback
    • Scopes: openid, email, profile
  2. Create an Application using that provider, and assign the users/groups who should be able to sign in — Authentik controls who can authenticate; the app's own admin/operator/viewer roles control what they can do once in.
  3. Copy the provider's issuer URL, client ID, and client secret into .env.

2. Local development

cp .env.example .env   # fill in AUTHENTIK_*, SESSION_SECRET, CREDENTIALS_ENCRYPTION_KEY
npm install
npm run dev:server   # http://localhost:3000 (API)
npm run dev:web      # http://localhost:5173 (Vite dev server, proxies /api and /auth to :3000)

Visit http://localhost:5173 during development. Database migrations run automatically on server start. SQLite data lands in ./data (gitignored).

Generate SESSION_SECRET and CREDENTIALS_ENCRYPTION_KEY with:

node -e "console.log(require('crypto').randomBytes(32).toString('hex'))"

3. Run with Docker

cp .env.example .env
# edit .env
docker compose -f docker-compose.dev.yml up -d --build   # build locally
# or, once an image is published to your registry:
docker compose up -d

The app listens on HOST_PORT (default 3000); SQLite data persists in ./data on the host.

S
Description
No description provided
Readme
1,009 KiB
Languages
TypeScript 96.5%
PowerShell 1.9%
Shell 1.3%