Files
Homelab-manager/README.md
T
bobbanandClaude Sonnet 5 7e81306aa7 Read SSL certificate expiry from the live server instead of trusting a typed-in date
A certificate secret's expiry was only ever what someone typed in, so
a renewed cert (or a wrong date) meant the app's reminders were
silently wrong. A certificate secret can now be given a host:port; the
app opens a real TLS connection and reads the certificate's actual
expiry — on create/edit (if the host changes), daily, and via a
per-row "Check now" — and keeps expiryDate in sync. Because the daily
refresh runs before the existing expiry check, the reminder is always
computed from what's actually being served.

Verification is deliberately off for the connection: homelab services
routinely serve self-signed/internal-CA certs, and an already-expired
one is exactly the case worth reporting, which a verifying connection
would refuse before exposing the dates.

Failure handling avoids the silent-staleness this is meant to fix: a
failed check keeps the last known date, records why on the row (shown
as a "Check failed" badge), and is listed in the daily secrets
notification. Creating a monitored secret whose host can't be reached
and with no manual date is rejected with the reason rather than saved
blank. A non-TLS port (the likeliest typo) gets a plain-language
error instead of raw OpenSSL output.

Server-side connections to a user-supplied host:port need the same
operator role that already gates editing secrets (and running Semaphore
templates, which is strictly more powerful); the host is validated
against a strict character set before any connection is made.

New nullable secrets columns (check_host, check_port, last_checked_at,
last_check_error) via migration 0007; existing rows are unaffected.

Verified against real TLS servers (openssl-generated certs) and the
real secrets router with a stubbed session: a live 45-day cert read
back as the correct date via both an IP host (no SNI) and a hostname;
an already-expired cert reported its past date and shows as expired;
refused connections, a server that accepts but never answers (times
out), and a plain non-TLS server each produced a descriptive error
rather than a hang or crash. Through the router: create with a host
and no date reads the date; unreachable host with no date -> 400 with
the reason; unreachable with a manual date -> saved with the error
recorded; host on a non-certificate type and an invalid host string
-> 400; a hand-typed date on a monitored secret is ignored; changing
the host re-checks immediately; changing the type away from
certificate ends monitoring; a viewer gets 403 on Check now. 23 checks,
all passing (a first re-run showed 2 spurious failures that were leftover
rows from the previous run's scratch database, confirmed by a clean re-run).
Real dev database mtime untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-25 23:41:29 +02:00

183 lines
9.7 KiB
Markdown

# Homelab Manager
Repository: `git@10.200.5.13:bobban/Homelab-manager.git` ([gitea.labsconnect.se/bobban/Homelab-manager](https://gitea.labsconnect.se/bobban/Homelab-manager) externally).
A single dashboard for a homelab: Proxmox, Synology DSM, Semaphore, Tailscale,
Gitea, and Dockhand/Docker status and basic actions, plus DNS record
management, an IP address inventory (IPAM), and a secret-expiry tracker
(ported from [Sloth Manager](../Sloth%20manager)) and scheduled-task tracking
across Debian/Raspbian hosts (ported from
[Schedule Task Manager](../ScheduleTaskManager)). Looks and feels like a
[Tabler](https://tabler.io) admin dashboard. Sign-in is delegated to
Authentik (OIDC), with local admin/operator/viewer roles.
## Status
All modules from the original plan are built:
- Monorepo scaffold, Tabler-themed app shell/navigation
- Authentik OIDC login, roles (first user to sign in becomes admin), audit log
- **Dashboard** — an overview of every system this app tracks, all sharing
one widget-card design (label + status badge, a small stat row, then its
own breakdown): DNS (domain/record counts per provider, cached records by
type), Secrets (monitored/expiring/expired, by type), and one widget per
integration — Tailscale by OS, Proxmox by node (VM/LXC counts too),
Dockhand by container state (plus host count), Semaphore and Gitea by
last-run status (plus private-repo count), and Synology's CPU/RAM
alongside its disk-health breakdown. Breakdowns render as a stacked
proportion bar with a legend — no charting
library, matching the rest of the app's plain-Tabler-CSS approach.
- **Diagnostic Log** (admin-only) — every call this app makes to a DNS
provider or integration (Tailscale, Proxmox, Synology, Semaphore, Gitea,
Dockhand), success or failure, with latency and the error message if it
failed — the last 500 calls, filterable by source/result, for
troubleshooting connectivity issues (ported from Sloth Manager's
provider-diagnostics log, generalized to cover every integration this app
has, not just DNS)
- **Secrets** — expiry tracking for API tokens/certs/passwords. An SSL
certificate can optionally be given a host:port to watch: the app opens a
real TLS connection (daily, and on demand via "Check now"), reads the
certificate's actual expiry, and keeps the date current — so a renewed cert
is picked up automatically and an unreachable host is flagged instead of
silently going stale
- **IP Addresses (IPAM)** — inventory of IPs across vendors/locations,
with "Sync from Tailscale" and "Sync from Proxmox" actions to pull in
tailnet device IPs and VM/LXC IPs (never overwrites a
manually-entered IP), and each entry now shows its matching DNS
record(s) from the DNS module's cache
- **DNS** — zone/record management across Cloudflare, Loopia, Pi-hole, Azure
DNS, cPanel, and Technitium; providers are configured in-app (not via env
vars) and their credentials are encrypted at rest
- **Servers** — cron/systemd tracking across Debian/Raspbian servers via a
lightweight push agent (`agent/linux/`), plus manual entries for things an
agent can't see (Docker jobs, backups). The Servers page itself just lists
registered servers and (admin-only) adds new ones / issues agent tokens;
clicking a server opens its detail page with CPU/RAM/disk status, IP
addresses, matching DNS names (looked up from the DNS module's cache), and
its scheduled tasks — live hardware from Proxmox for VM/LXC-backed
servers, or from the agent's own hardware report for everything else,
both showing the same per-disk usage breakdown (an LXC's root filesystem
read straight from the host; a QEMU VM's actual mounts via its guest
agent, alongside the allocated size Proxmox already knew about without
one). Proxmox-linked servers also get start/stop/restart buttons right on the
detail page. The detail page also has an **Admin Links** section
(operator/admin to add/edit/remove) for bookmarking that server's own
admin UIs — Dockge, Webmin, Cockpit, Portainer, or anything else reachable
by URL. Since not every server is a Proxmox VM, an admin can hide the
"Proxmox link" card per server ("Not a VM? Hide this" / "+ Show Proxmox
link options") — it stays visible regardless once a server actually is
linked, so unlinking is always reachable.
- **Tailscale**, **Proxmox**, **Synology**, **Semaphore**, **Gitea**,
and **Docker** each get their own top-level page (backed by the
matching integration) instead of living inside a shared Integrations
browsing view:
- **Tailscale** — device list with online/authorized status, and
authorize/deauthorize/remove actions; a live device-count widget.
- **Proxmox** — VM/LXC status across every node in the cluster, with
start/restart/shutdown/stop actions; a live running/total widget. Each
online node also gets its own host-stats card — uptime, CPU usage/cores/
load average, RAM and swap usage, and per-storage usage (local, LVM-thin,
ZFS, NFS, etc). Supports self-signed certificates (common in homelab
setups).
- **Synology** — volume and disk health (read-only by design).
Supports self-signed certificates.
- **Semaphore** — Ansible run status per template across every
project, with a "Run" action to trigger a template; a live
template-count widget (with a last-failed warning).
- **Gitea** — repo list with each repo's last CI run status, and
re-running just the failed jobs in a run; a live repo-count widget
(with a failing-build warning).
- **Docker** — container status across every Docker host Dockhand
manages (one credential covers all of them), with
start/stop/restart actions, a host filter, image-update status per
container (from Dockhand's own cached update check, plus a button
to trigger a fresh one), and a live running/total widget (with an
updates-available count).
The Integrations page itself is now just a list of configured
integrations (name/type/status, visible to every role) with an
admin-only "Add integration" button and edit/enable/disable/delete
actions per row — the six dedicated pages above are where you
actually use each one.
- Every table in the app is click-to-sort on any column (numbers, booleans,
and dates/text sort correctly regardless of how the column formats them)
and has an "Export CSV" button next to it that exports whatever's
currently sorted/filtered. Tables that can realistically grow large
(DNS zones/records, IP Addresses, Secrets, Servers, Audit Log, and each
integration's device/container/guest/repo/template list) are paginated,
20 rows per page by default — adjustable under Settings → Display — and
CSV export still covers every sorted/filtered row, not just the current
page.
- **Settings** (admin-only) — notification channels (Gotify, ntfy, SMTP,
generic webhook) with per-channel test buttons, per-event toggles (DNS
record added/updated/deleted, daily secret-expiry reminder and daily
Tailscale key-expiry reminder — both sharing one configurable
time/timezone), badge-color customization for both DNS
providers and integration types, and a **Display** tab (date order,
12/24-hour clock, and rows-per-page for every paginated table) applied
consistently across the app.
All six integrations follow the same config-in-UI + encrypted-credentials
pattern, added (and edited — e.g. to rotate an expired API token without
recreating the whole integration) through **Integrations → Manage
integrations**. See [INTEGRATIONS.md](INTEGRATIONS.md) for exactly what
credential to create and what access it needs in each target system,
for every integration and DNS provider.
**Verified for real, end to end**: every module above — including all six
integrations, both their read-only views and their write actions
(start/stop/restart, trigger-a-run, authorize/deauthorize) — has been
exercised against the user's actual live homelab, not just built against
specs. That pass also found and fixed two real bugs: the Synology adapter
assumed HTTPS-only (the NAS is reached over plain HTTP), and the Tailscale
adapter read `online`/`isExitNode` fields that don't actually exist in the
real API response (fixed to derive them from `connectedToControl` and
`enabledRoutes`). See the git log for the full verification notes per
integration.
## Requirements
- Node.js 20+
- An Authentik instance reachable from wherever this app runs
## 1. Set up an Authentik application
1. Create an **OAuth2/OpenID Provider**:
- Redirect URI: `<APP_BASE_URL>/auth/callback`
- Scopes: `openid`, `email`, `profile`
2. Create an **Application** using that provider, and assign the users/groups
who should be able to sign in — Authentik controls who can authenticate;
the app's own admin/operator/viewer roles control what they can do once in.
3. Copy the provider's issuer URL, client ID, and client secret into `.env`.
## 2. Local development
```bash
cp .env.example .env # fill in AUTHENTIK_*, SESSION_SECRET, CREDENTIALS_ENCRYPTION_KEY
npm install
npm run dev:server # http://localhost:3000 (API)
npm run dev:web # http://localhost:5173 (Vite dev server, proxies /api and /auth to :3000)
```
Visit `http://localhost:5173` during development. Database migrations run
automatically on server start. SQLite data lands in `./data` (gitignored).
Generate `SESSION_SECRET` and `CREDENTIALS_ENCRYPTION_KEY` with:
```bash
node -e "console.log(require('crypto').randomBytes(32).toString('hex'))"
```
## 3. Run with Docker
```bash
cp .env.example .env
# edit .env
docker compose -f docker-compose.dev.yml up -d --build # build locally
# or, once an image is published to your registry:
docker compose up -d
```
The app listens on `HOST_PORT` (default `3000`); SQLite data persists in
`./data` on the host.