A check of every write path found gaps in what the audit log captured: - Sign-ins and sign-outs are now recorded with the IP they came from (sign-out is recorded first and can't block signing out). - A new account is recorded when it's created on first sign-in, including when the very first user becomes admin - so the log shows who gained access, not only who changed things. - The automatic log purge, which deletes audit entries, now records itself, attributed to "system". recordAudit() takes an optional actor for this. It only records when something was actually deleted. - Settings updates record what changed (before and after) instead of only which sections were touched. The notification channels (Gotify, ntfy, SMTP, webhook) record field names only: they hold credentials, and webhook URLs and public ntfy topics act as secrets, while the audit log is readable by operators and Settings is admin-only. - Integration edits record renames, enabling/disabling, whether credentials were replaced (never the credentials), and which settings fields changed (names only). The Privacy page and README now say sign-ins store an IP in the audit log. Verified through the real routes against a scratch database: user creation, the logout route, the automatic purge, settings and integration edits - including that a secret token and a webhook URL appear nowhere in the stored entries. The sign-in callback itself needs a real identity provider and wasn't run. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
372 lines
22 KiB
Markdown
372 lines
22 KiB
Markdown
# Homelab Manager
|
|
|
|
Repository: `git@10.200.5.13:bobban/Homelab-manager.git` ([gitea.labsconnect.se/bobban/Homelab-manager](https://gitea.labsconnect.se/bobban/Homelab-manager) externally).
|
|
|
|
A single dashboard for a homelab: Proxmox, Synology DSM, Semaphore, Tailscale,
|
|
Gitea, Dockhand/Docker, Uptime Kuma, Proxmox Backup Server, and osTicket status
|
|
and basic actions, plus DNS record
|
|
management, an IP address inventory (IPAM), and a secret-expiry tracker
|
|
(ported from [Sloth Manager](../Sloth%20manager)) and scheduled-task tracking
|
|
across Debian/Raspbian hosts (ported from
|
|
[Schedule Task Manager](../ScheduleTaskManager)). Looks and feels like a
|
|
[Tabler](https://tabler.io) admin dashboard. Sign-in is delegated to
|
|
Authentik (OIDC), with local admin/operator/viewer roles.
|
|
|
|
## Status
|
|
|
|
All modules from the original plan are built:
|
|
|
|
- Monorepo scaffold, Tabler-themed app shell with a grouped sidebar (Infrastructure, Network, Automation, Operations, Administration; groups open on demand, the one holding the current page is always open, and what you leave open is remembered)
|
|
- Authentik OIDC login, roles (first user to sign in becomes admin), audit log
|
|
(changes made in the app, sign-ins and sign-outs with the IP they came from, new accounts, and the log's own
|
|
automatic trimming — attributed to "system"; settings changes show what changed, but never credentials)
|
|
- **Dashboard** — an overview of every system this app tracks, all sharing
|
|
one widget-card design (label + status badge, a small stat row, then its
|
|
own breakdown): DNS (domain/record counts per provider, cached records by
|
|
type), Secrets (monitored/expiring/expired, by type), and one widget per
|
|
integration — Tailscale by OS, Proxmox by node (VM/LXC counts too),
|
|
Dockhand by container state (plus host count), Semaphore and Gitea by
|
|
last-run status (plus private-repo count), and Synology's CPU/RAM
|
|
alongside its disk-health breakdown. Breakdowns render as a stacked
|
|
proportion bar with a legend — no charting
|
|
library, matching the rest of the app's plain-Tabler-CSS approach.
|
|
The Servers widget shows how many servers are online, offline, or have never
|
|
reported, which are offline and for how long, and any disk at or above the
|
|
usage threshold — decided by the same rules as the health alerts, so the
|
|
widget and the notifications always agree, and following the thresholds in
|
|
Settings even for viewers who can't open them.
|
|
The Domains widget shows how many registrations are tracked, which are
|
|
expired or expiring (soonest first), the next one to expire, and how many
|
|
couldn't be refreshed.
|
|
The Uptime Kuma widget shows the monitor count, how many are down, and how
|
|
many are matched to one of your servers.
|
|
The Proxmox Backup Server widget shows the datastore count and how many
|
|
stored snapshots have failed verification or were never verified.
|
|
The osTicket widget shows how many tickets are open, overdue, and awaiting
|
|
a staff reply.
|
|
- **Diagnostic Log** (admin-only) — every call this app makes to a DNS
|
|
provider or integration (Tailscale, Proxmox, Synology, Semaphore, Gitea,
|
|
Dockhand, Uptime Kuma, Proxmox Backup Server, osTicket), success or failure, with latency and the error message if it
|
|
failed — the last 500 calls, filterable by source/result, for
|
|
troubleshooting connectivity issues (ported from Sloth Manager's
|
|
provider-diagnostics log, generalized to cover every integration this app
|
|
has, not just DNS)
|
|
- **Secrets** — expiry tracking for API tokens/certs/passwords. An SSL
|
|
certificate can optionally be given a host:port to watch: the app opens a
|
|
real TLS connection (daily, and on demand via "Check now"), reads the
|
|
certificate's actual expiry, and keeps the date current — so a renewed cert
|
|
is picked up automatically and an unreachable host is flagged instead of
|
|
silently going stale
|
|
- **IP Addresses (IPAM)** — inventory of IPs across vendors/locations,
|
|
with "Sync from Tailscale", "Sync from Proxmox", and "Sync from phpIPAM"
|
|
actions to pull in tailnet device IPs, VM/LXC IPs, and phpIPAM's own
|
|
addresses (never overwrites a manually-entered IP, or one a different sync
|
|
owns), and each entry now shows its matching DNS record(s) from the DNS
|
|
module's cache. Shares the Consistency page's excluded-ranges setting (see
|
|
below) so addresses that aren't interesting to track — a Docker bridge
|
|
network repeating on every host, say — can be hidden here too, with a
|
|
"Hide excluded addresses" toggle and a per-entry "Exclude…" shortcut to add
|
|
a range on the spot; managing ranges from either page updates the other
|
|
- **DNS** — zone/record management across Cloudflare, Loopia, Pi-hole, Azure
|
|
DNS, cPanel, and Technitium; providers are configured in-app (not via env
|
|
vars) and their credentials are encrypted at rest
|
|
- **Servers** — cron/systemd tracking across Debian/Raspbian servers via a
|
|
lightweight push agent (`agent/linux/`), Windows scheduled-task tracking
|
|
through a PowerShell agent (`agent/windows/`, see its README), plus manual
|
|
entries for things an agent can't see (Docker jobs, backups). The Servers page itself just lists
|
|
registered servers and (admin-only) adds new ones / issues agent tokens;
|
|
clicking a server opens its detail page with CPU/RAM/disk status, IP
|
|
addresses, matching DNS names (looked up from the DNS module's cache), and
|
|
its scheduled tasks — live hardware from Proxmox for VM/LXC-backed
|
|
servers, or from the agent's own hardware report for everything else,
|
|
both showing the same per-disk usage breakdown (an LXC's root filesystem
|
|
read straight from the host; a QEMU VM's actual mounts via its guest
|
|
agent, alongside the allocated size Proxmox already knew about without
|
|
one). Proxmox-linked servers also get start/stop/restart buttons right on the
|
|
detail page. The detail page also has an **Admin Links** section
|
|
(operator/admin to add/edit/remove) for bookmarking that server's own
|
|
admin UIs — Dockge, Webmin, Cockpit, Portainer, or anything else reachable
|
|
by URL. **Operations → Admin Links** summarizes every server's admin links
|
|
in one sortable, searchable table (server, label, URL, an "Open" link, and
|
|
the same add/edit/delete as the per-server section — adding one here just
|
|
asks which server it belongs to), so finding or managing one doesn't mean
|
|
visiting each server's own page. Since not every server is a Proxmox VM, an admin can hide the
|
|
"Proxmox link" card per server ("Not a VM? Hide this" / "+ Show Proxmox
|
|
link options") — it stays visible regardless once a server actually is
|
|
linked, so unlinking is always reachable.
|
|
- **Tailscale**, **Proxmox**, **Synology**, **Semaphore**, **Gitea**,
|
|
**Docker**, **Uptime Kuma**, **Proxmox Backup Server**, and **osTicket**
|
|
each get their own top-level page (backed by the matching integration)
|
|
instead of living inside a shared Integrations browsing view:
|
|
- **Tailscale** — device list with online/authorized status, and
|
|
authorize/deauthorize/remove actions; a live device-count widget.
|
|
- **Proxmox** — VM/LXC status across every node in the cluster, with
|
|
start/restart/shutdown/stop actions; a live running/total widget. Each
|
|
online node also gets its own host-stats card — uptime, CPU usage/cores/
|
|
load average, RAM and swap usage, and per-storage usage (local, LVM-thin,
|
|
ZFS, NFS, etc). Supports self-signed certificates (common in homelab
|
|
setups).
|
|
- **Synology** — volume and disk health (read-only by design).
|
|
Supports self-signed certificates.
|
|
- **Semaphore** — Ansible run status per template across every
|
|
project, with a "Run" action to trigger a template; a live
|
|
template-count widget (with a last-failed warning).
|
|
- **Gitea** — repo list with each repo's last CI run status, and
|
|
re-running just the failed jobs in a run; a live repo-count widget
|
|
(with a failing-build warning).
|
|
- **Docker** — container status across every Docker host Dockhand
|
|
manages (one credential covers all of them), with
|
|
start/stop/restart actions, a host filter, image-update status per
|
|
container (from Dockhand's own cached update check, plus a button
|
|
to trigger a fresh one), and a live running/total widget (with an
|
|
updates-available count).
|
|
- **Uptime Kuma** — every monitor's status (up/down/pending/maintenance),
|
|
response time, and 24h/30d uptime, with certificate days remaining where
|
|
applicable. Each monitor is matched to one of your servers when its
|
|
target (an IP, or a TCP/HTTP hostname) lines up with that server's own
|
|
address or hostname, linking straight to it — so you can see what's
|
|
actually being watched on each box, not just a flat monitor list.
|
|
Uptime Kuma has no conventional REST API (the dashboard talks to it over
|
|
Socket.IO), so this reads its Prometheus `/metrics` endpoint instead and
|
|
parses that itself; read-only, no actions.
|
|
- **Proxmox Backup Server** — every configured datastore's usage and
|
|
snapshot count, the PBS host's own CPU/RAM/disk, and each stored
|
|
snapshot's verification status, since Proxmox VE only knows whether a
|
|
backup *ran*, never whether PBS's own verify pass on the stored data
|
|
still passes; a daily check alerts on any snapshot that's failed
|
|
verification or a datastore that couldn't be read. Read-only, no
|
|
actions (nothing here can prune, delete, or trigger a re-verify).
|
|
- **osTicket** — every currently open ticket, with its status, priority,
|
|
department, assigned staff member or team, requester, and two flags
|
|
worth a glance on their own: **overdue** and **awaiting our reply**.
|
|
osTicket's own REST API only supports *creating* tickets, not listing
|
|
them, so this reads osTicket's MySQL/MariaDB database directly with a
|
|
read-only user — the only integration here that isn't a REST API.
|
|
Read-only, no actions.
|
|
|
|
The Integrations page itself is now just a list of configured
|
|
integrations (name/type/status, visible to every role) with an
|
|
admin-only "Add integration" button and edit/enable/disable/delete
|
|
actions per row — the nine dedicated pages above are where you
|
|
actually use each one.
|
|
- Every table in the app is click-to-sort on any column (numbers, booleans,
|
|
and dates/text sort correctly regardless of how the column formats them)
|
|
and has an "Export CSV" button next to it that exports whatever's
|
|
currently sorted/filtered. Tables that can realistically grow large
|
|
(DNS zones/records, IP Addresses, Secrets, Servers, Audit Log, and each
|
|
integration's device/container/guest/repo/template list) are paginated,
|
|
20 rows per page by default — adjustable under Settings → Display — and
|
|
CSV export still covers every sorted/filtered row, not just the current
|
|
page.
|
|
- **Settings** (admin-only) — notification channels (Gotify, ntfy, SMTP,
|
|
generic webhook) with per-channel test buttons, per-event toggles (DNS
|
|
record added/updated/deleted, daily secret-expiry reminder and daily
|
|
Tailscale key-expiry reminder — both sharing one configurable
|
|
time/timezone), badge-color customization for both DNS
|
|
providers and integration types, and a **Display** tab (date order,
|
|
12/24-hour clock, and rows-per-page for every paginated table) applied
|
|
consistently across the app.
|
|
|
|
All six integrations follow the same config-in-UI + encrypted-credentials
|
|
pattern, added (and edited — e.g. to rotate an expired API token without
|
|
recreating the whole integration) through **Integrations → Manage
|
|
integrations**. See [INTEGRATIONS.md](INTEGRATIONS.md) for exactly what
|
|
credential to create and what access it needs in each target system,
|
|
for every integration and DNS provider.
|
|
|
|
**Verified for real, end to end**: every module above — including all six
|
|
integrations, both their read-only views and their write actions
|
|
(start/stop/restart, trigger-a-run, authorize/deauthorize) — has been
|
|
exercised against the user's actual live homelab, not just built against
|
|
specs. That pass also found and fixed two real bugs: the Synology adapter
|
|
assumed HTTPS-only (the NAS is reached over plain HTTP), and the Tailscale
|
|
adapter read `online`/`isExitNode` fields that don't actually exist in the
|
|
real API response (fixed to derive them from `connectedToControl` and
|
|
`enabledRoutes`). See the git log for the full verification notes per
|
|
integration. (Uptime Kuma, Proxmox Backup Server, and osTicket, added
|
|
later, are not part of that "six" — see their own git log entries, and
|
|
[INTEGRATIONS.md](INTEGRATIONS.md), for what was and wasn't verified
|
|
against a real instance.)
|
|
|
|
Server and storage health is watched every 15 minutes: a server whose agent
|
|
stops reporting, a server disk / Proxmox storage / Synology volume passing a
|
|
usage threshold, and a Synology volume or disk that's degraded or failing each
|
|
raise one notification when the problem starts and one when it clears (both
|
|
thresholds are set under Settings → Notifications). Active problems are
|
|
remembered across restarts, so a rebuild doesn't re-alert them.
|
|
|
|
**Tags** — servers can be tagged (prod, media, rack-1, …) from their detail
|
|
page by operators and admins. Tags show as coloured chips on the Servers page,
|
|
which can be filtered by one or several tags (the filter is in the URL, so a
|
|
tag on a server's page links to everything sharing it), and they're searchable
|
|
from the global search box.
|
|
Admins manage tags under Settings → Tags: add tags before any server uses them (they're
|
|
offered as one-click suggestions when tagging), give any tag a colour of your choosing
|
|
(or leave it on the automatic one), rename tags, and delete them. Renaming to a name
|
|
that already exists merges the two, and both rename and delete rewrite every server
|
|
that carries the tag.
|
|
|
|
**Privacy** — a page every signed-in user can open that says what this
|
|
installation stores (accounts, sign-in sessions with their IP and browser, the audit
|
|
and diagnostic logs, server reports, the secrets tracker, credentials), where data
|
|
goes (Authentik, your integrations and DNS providers, the notification channels
|
|
that are switched on, domain registries), what's kept in the browser, who can see
|
|
what, and how to limit or remove data. Retention, integrations and channels are
|
|
read live from the installation; channel addresses are shown to admins only. Each
|
|
user sees their own account and sign-ins there and can download their own data
|
|
(account, sign-ins, audit-log entries) as a JSON file.
|
|
|
|
**Consistency** — a report of where IPAM, DNS and your servers disagree about
|
|
an address: the same address on two servers, a DNS record named after a server
|
|
that points somewhere it isn't, an IPAM entry labelled with a server's name at
|
|
the wrong address, addresses in use that IPAM doesn't list (with an "Add to
|
|
IPAM" button), and server addresses no DNS record points at. It only compares
|
|
data the app already holds — nothing is fetched when you open it — so it says how
|
|
many servers reported addresses and how fresh the synced DNS zones are. Only
|
|
private addresses are compared; ranges you exclude (Docker's, which repeat the same
|
|
subnet on many hosts — 172.16.0.0/12 is excluded by default, remove it if that's a
|
|
real LAN for you) are left out of every source; anything that's fine on purpose
|
|
can be ignored with a reason, and stays ignored. The excluded-ranges list is the
|
|
same one the IP Addresses page manages — edit it from either page and both
|
|
reflect the change.
|
|
|
|
**Domains** — when each domain registration expires, read from the registry
|
|
itself. The domains behind your DNS zones are picked up automatically (a zone
|
|
like `lab.example.se` resolves to the `example.se` registration that actually
|
|
expires); others can be added by hand. Each is looked up daily over RDAP where
|
|
the TLD offers it, and otherwise over WHOIS via IANA's referral — which is what
|
|
makes `.se`, `.nu` and `.io` work, since those aren't in the RDAP bootstrap.
|
|
You're reminded daily from N days before expiry (Settings → Notifications,
|
|
default 30) until it's renewed, and told when an expiry date couldn't be
|
|
refreshed for several days. Registries that don't publish an expiry (`.de`,
|
|
`.eu`) can be tracked but have no date to warn about. Private zones (`.lan`,
|
|
`.local`) are skipped.
|
|
|
|
**Automation failures** — every 15 minutes the app looks at the latest run of
|
|
each Semaphore template and each Gitea repo's latest workflow run. A failed one
|
|
raises a single notification (with the project/template or repo, run number, and
|
|
for Gitea the run's link), and another when a later run succeeds. It's
|
|
state-based, so a job that fails every night alerts on the first failure rather
|
|
than every night. A run that's still going, was cancelled, or was stopped by
|
|
hand leaves things as they were, and anything that couldn't be read (a Semaphore
|
|
project or Gitea repo that errored, or an integration that's down) is neither
|
|
cleared nor re-announced. The first check after upgrading only records what's
|
|
already failing, so old failures aren't announced. Toggle it under Settings →
|
|
Notifications. For Gitea this follows the repo's most recent run on any
|
|
workflow or branch, the same as the Gitea page shows.
|
|
|
|
**Maintenance mode** silences alerts about one server, integration, or DNS
|
|
provider while you work on it (server offline / disk, storage and Synology
|
|
health, Proxmox backup alerts, Proxmox Backup Server verification alerts, and
|
|
"integration down" for that service type).
|
|
Every window has a fixed end (5 minutes to 7 days) and expires on its own, and
|
|
a problem that began during a window and is still present when it ends alerts
|
|
then — a forgotten window can't hide an outage. A banner shows what's currently
|
|
silenced to every signed-in user. If you also run Uptime Kuma, "Import from
|
|
Uptime Kuma" on the Maintenance page reads which of its monitors are
|
|
currently in maintenance and starts (or extends) a window here for whichever
|
|
server each one's target matches, for a duration you choose — Uptime Kuma's
|
|
metrics only say what's in maintenance right now, not for how long, so this
|
|
doesn't try to mirror its schedule, only its current state.
|
|
|
|
**Ports** — each server's detail page has a Ports card for finding free ports
|
|
and remembering what each one is for. "Scan…" runs a TCP scan of a port range
|
|
from the app against the server's address (private addresses only, up to
|
|
20,000 ports at a time) and lists what's open plus the ranges that were
|
|
confirmed free; a free range can be clicked to reserve a port. Any port can
|
|
carry a service name and a comment, and a port with a note counts as taken
|
|
even when nothing is listening. The agent also reports what is bound on the
|
|
host (`ss`), which catches services listening on localhost only — a scan from
|
|
elsewhere can't see those, so they'd otherwise look free. Re-run the agent
|
|
install one-liner on a host to pick that up.
|
|
|
|
**Network → Ports** is the cross-server counterpart: one page listing every
|
|
port every agent currently reports as listening, across all servers at once
|
|
(protocol, address, process, last report time), each linking back to its
|
|
server. Below that, a separate, manually-maintained table is for the ports
|
|
this app can't see on its own — a router's port forward, an edge firewall
|
|
rule, a cloud security group — the same reason people keep a spreadsheet of
|
|
"what did I open and why." Each entry has a label, the external port/protocol,
|
|
an optional link to a tracked server (plus its own internal port, when NAT
|
|
changes it) or a freeform destination, a free-text "source" (which
|
|
router/firewall/service it's actually configured on — this app doesn't talk
|
|
to any firewall, so it can't manage or verify the rule, only record it), and
|
|
a comment. Viewers can see both tables; adding, editing, or deleting a manual
|
|
entry needs operator or admin.
|
|
|
|
The app is installable as a PWA — "Install app" / "Add to Home Screen" from the
|
|
browser gives it its own icon and a standalone window on phone or desktop. This
|
|
needs the site to be served over HTTPS (browsers only offer install on secure
|
|
origins; `localhost` also counts). The bundled service worker deliberately
|
|
caches nothing, so an installed copy always shows the current build.
|
|
|
|
## Requirements
|
|
|
|
- Node.js 20+
|
|
- An Authentik instance reachable from wherever this app runs
|
|
|
|
## 1. Set up an Authentik application
|
|
|
|
1. Create an **OAuth2/OpenID Provider**:
|
|
- Redirect URI: `<APP_BASE_URL>/auth/callback`
|
|
- Scopes: `openid`, `email`, `profile`
|
|
2. Create an **Application** using that provider, and assign the users/groups
|
|
who should be able to sign in — Authentik controls who can authenticate;
|
|
the app's own admin/operator/viewer roles control what they can do once in.
|
|
3. Copy the provider's issuer URL, client ID, and client secret into `.env`.
|
|
|
|
## 2. Local development
|
|
|
|
```bash
|
|
cp .env.example .env # fill in AUTHENTIK_*, SESSION_SECRET, CREDENTIALS_ENCRYPTION_KEY
|
|
npm install
|
|
npm run dev:server # http://localhost:3000 (API)
|
|
npm run dev:web # http://localhost:5173 (Vite dev server, proxies /api and /auth to :3000)
|
|
```
|
|
|
|
Visit `http://localhost:5173` during development. Database migrations run
|
|
automatically on server start. SQLite data lands in `./data` (gitignored).
|
|
|
|
Generate `SESSION_SECRET` and `CREDENTIALS_ENCRYPTION_KEY` with:
|
|
|
|
```bash
|
|
node -e "console.log(require('crypto').randomBytes(32).toString('hex'))"
|
|
```
|
|
|
|
## 3. Run with Docker
|
|
|
|
```bash
|
|
cp .env.example .env
|
|
# edit .env
|
|
docker compose -f docker-compose.dev.yml up -d --build # build locally
|
|
# or, once an image is published to your registry:
|
|
docker compose up -d
|
|
```
|
|
|
|
The app listens on `HOST_PORT` (default `3000`); SQLite data persists in
|
|
`./data` on the host.
|
|
|
|
### arm64 (Raspberry Pi, Apple silicon, ARM servers)
|
|
|
|
Two compose files, one to build the image and one to run the published one:
|
|
|
|
```bash
|
|
# On any machine with Docker: build the arm64 image and push it to the registry
|
|
docker compose -f docker-compose.arm64.build.yml build
|
|
docker compose -f docker-compose.arm64.build.yml push
|
|
|
|
# On the arm64 machine: pull that image and run it (no build there)
|
|
docker compose -f docker-compose.arm64.yml pull
|
|
docker compose -f docker-compose.arm64.yml up -d
|
|
```
|
|
|
|
`docker-compose.arm64.build.yml` can also build and run right on an arm64 machine
|
|
(`up -d --build`). Building on x86 goes through QEMU emulation — slower, and it needs
|
|
a one-time `docker run --privileged --rm tonistiigi/binfmt --install arm64`.
|
|
|
|
Both files use the image `gitea.labsconnect.se/bobban/homelabmanager-homelab-manager:arm64`.
|
|
Set `IMAGE_REPO` and/or `ARM64_TAG` in `.env` to use another registry or to pin a release
|
|
tag (e.g. `ARM64_TAG=1.4.0-arm64`). The arm64 tag is kept separate from the default image,
|
|
so the existing `docker-compose.yml` is unaffected. A `.dockerignore` keeps host
|
|
`node_modules`, `.env` and `data/` out of the build.
|