Reports a Windows machine the way the Linux agent does, replacing the
"planned" stub in agent/windows: scheduled tasks plus hostname, IPv4
addresses, CPU model/cores/current load, memory, every fixed disk, and
TCP/UDP listening ports with the owning process (which feed the Ports
card, localhost-only listeners included).
Scripts (plain ASCII by design -- they are downloaded as text and Windows
PowerShell 5.1 reads BOM-less files as ANSI):
- report-tasks.ps1: collects and POSTs to /api/agent/report. Works in
Windows PowerShell 5.1 and PowerShell 7. -DryRun prints the JSON.
Microsoft's own \Microsoft\ tasks (hundreds) are left out unless
INCLUDE_MICROSOFT_TASKS is set. Triggers are turned into readable text
("Weekly on Mon, Wed at 03:00", "At logon", "..., repeating every 15 min").
Self-signed certificates work via API_INSECURE on both PowerShell
versions (they need different mechanisms).
- install.ps1: elevated only; downloads the agent to ProgramData, writes
agent.json with permissions locked to SYSTEM and Administrators *before*
the token goes in, and registers a SYSTEM scheduled task (every 15 min
plus at startup with a 2 min delay). Reinstalling replaces the task.
- uninstall.ps1: removes the task and only the files the agent installed.
Server: accepts schedule_type "windows_task"; a server can be registered
as Windows (Add a server has an operating system choice); an agent's
reported os_type ("linux"/"windows", anything else ignored) corrects the
stored one. The Servers page shows the right install and uninstall
command for each OS (Windows PowerShell 5.1 one-liners, with a self-signed
variant and a note about PowerShell 7), and Windows tasks are labelled
"Windows scheduled tasks". The Linux commands are unchanged.
Verified on this Windows machine, in both PowerShell 5.1 and 7:
- Real dry runs found and fixed bugs before anything shipped: tasks and
ports came out as one nested item (return , $out wrapped twice), integer
keys in an ordered dictionary index by position (wrong weekday names),
and generic "Trigger" labels.
- End to end against the real agent-report router: HTTP, self-signed HTTPS
refused by default and accepted with API_INSECURE, wrong token gives a
clear one-line error and exit 1, and Swedish letters plus a euro sign
survive JSON -> UTF-8 -> HTTP -> SQLite.
- 35 checks on trigger/action/duration descriptions, 20 on the installer's
building blocks (task parts built but not registered, credentials file
content and ACL, download over HTTP and self-signed HTTPS), 18 on the
server rules, and the generated one-liners run through PowerShell's
parser. The documented one-liners were run through iex and stop at the
administrator check without changing anything.
- Found that PowerShell 7 ignores the ServicePointManager certificate
override, so the installer's own download now uses -SkipCertificateCheck
there.
NOT verified: the elevated install itself. Registering a SYSTEM scheduled
task needs elevation and changes the machine, so it was not run: the task
registration, that the repeating trigger really runs indefinitely, and
the agent running as SYSTEM under Task Scheduler have not been exercised.
Windows 10 / Server 2016 or newer is assumed; older is untested.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
298 lines
17 KiB
Markdown
298 lines
17 KiB
Markdown
# Homelab Manager
|
|
|
|
Repository: `git@10.200.5.13:bobban/Homelab-manager.git` ([gitea.labsconnect.se/bobban/Homelab-manager](https://gitea.labsconnect.se/bobban/Homelab-manager) externally).
|
|
|
|
A single dashboard for a homelab: Proxmox, Synology DSM, Semaphore, Tailscale,
|
|
Gitea, and Dockhand/Docker status and basic actions, plus DNS record
|
|
management, an IP address inventory (IPAM), and a secret-expiry tracker
|
|
(ported from [Sloth Manager](../Sloth%20manager)) and scheduled-task tracking
|
|
across Debian/Raspbian hosts (ported from
|
|
[Schedule Task Manager](../ScheduleTaskManager)). Looks and feels like a
|
|
[Tabler](https://tabler.io) admin dashboard. Sign-in is delegated to
|
|
Authentik (OIDC), with local admin/operator/viewer roles.
|
|
|
|
## Status
|
|
|
|
All modules from the original plan are built:
|
|
|
|
- Monorepo scaffold, Tabler-themed app shell/navigation
|
|
- Authentik OIDC login, roles (first user to sign in becomes admin), audit log
|
|
- **Dashboard** — an overview of every system this app tracks, all sharing
|
|
one widget-card design (label + status badge, a small stat row, then its
|
|
own breakdown): DNS (domain/record counts per provider, cached records by
|
|
type), Secrets (monitored/expiring/expired, by type), and one widget per
|
|
integration — Tailscale by OS, Proxmox by node (VM/LXC counts too),
|
|
Dockhand by container state (plus host count), Semaphore and Gitea by
|
|
last-run status (plus private-repo count), and Synology's CPU/RAM
|
|
alongside its disk-health breakdown. Breakdowns render as a stacked
|
|
proportion bar with a legend — no charting
|
|
library, matching the rest of the app's plain-Tabler-CSS approach.
|
|
- **Diagnostic Log** (admin-only) — every call this app makes to a DNS
|
|
provider or integration (Tailscale, Proxmox, Synology, Semaphore, Gitea,
|
|
Dockhand), success or failure, with latency and the error message if it
|
|
failed — the last 500 calls, filterable by source/result, for
|
|
troubleshooting connectivity issues (ported from Sloth Manager's
|
|
provider-diagnostics log, generalized to cover every integration this app
|
|
has, not just DNS)
|
|
- **Secrets** — expiry tracking for API tokens/certs/passwords. An SSL
|
|
certificate can optionally be given a host:port to watch: the app opens a
|
|
real TLS connection (daily, and on demand via "Check now"), reads the
|
|
certificate's actual expiry, and keeps the date current — so a renewed cert
|
|
is picked up automatically and an unreachable host is flagged instead of
|
|
silently going stale
|
|
- **IP Addresses (IPAM)** — inventory of IPs across vendors/locations,
|
|
with "Sync from Tailscale" and "Sync from Proxmox" actions to pull in
|
|
tailnet device IPs and VM/LXC IPs (never overwrites a
|
|
manually-entered IP), and each entry now shows its matching DNS
|
|
record(s) from the DNS module's cache
|
|
- **DNS** — zone/record management across Cloudflare, Loopia, Pi-hole, Azure
|
|
DNS, cPanel, and Technitium; providers are configured in-app (not via env
|
|
vars) and their credentials are encrypted at rest
|
|
- **Servers** — cron/systemd tracking across Debian/Raspbian servers via a
|
|
lightweight push agent (`agent/linux/`), Windows scheduled-task tracking
|
|
through a PowerShell agent (`agent/windows/`, see its README), plus manual
|
|
entries for things an agent can't see (Docker jobs, backups). The Servers page itself just lists
|
|
registered servers and (admin-only) adds new ones / issues agent tokens;
|
|
clicking a server opens its detail page with CPU/RAM/disk status, IP
|
|
addresses, matching DNS names (looked up from the DNS module's cache), and
|
|
its scheduled tasks — live hardware from Proxmox for VM/LXC-backed
|
|
servers, or from the agent's own hardware report for everything else,
|
|
both showing the same per-disk usage breakdown (an LXC's root filesystem
|
|
read straight from the host; a QEMU VM's actual mounts via its guest
|
|
agent, alongside the allocated size Proxmox already knew about without
|
|
one). Proxmox-linked servers also get start/stop/restart buttons right on the
|
|
detail page. The detail page also has an **Admin Links** section
|
|
(operator/admin to add/edit/remove) for bookmarking that server's own
|
|
admin UIs — Dockge, Webmin, Cockpit, Portainer, or anything else reachable
|
|
by URL. Since not every server is a Proxmox VM, an admin can hide the
|
|
"Proxmox link" card per server ("Not a VM? Hide this" / "+ Show Proxmox
|
|
link options") — it stays visible regardless once a server actually is
|
|
linked, so unlinking is always reachable.
|
|
- **Tailscale**, **Proxmox**, **Synology**, **Semaphore**, **Gitea**,
|
|
and **Docker** each get their own top-level page (backed by the
|
|
matching integration) instead of living inside a shared Integrations
|
|
browsing view:
|
|
- **Tailscale** — device list with online/authorized status, and
|
|
authorize/deauthorize/remove actions; a live device-count widget.
|
|
- **Proxmox** — VM/LXC status across every node in the cluster, with
|
|
start/restart/shutdown/stop actions; a live running/total widget. Each
|
|
online node also gets its own host-stats card — uptime, CPU usage/cores/
|
|
load average, RAM and swap usage, and per-storage usage (local, LVM-thin,
|
|
ZFS, NFS, etc). Supports self-signed certificates (common in homelab
|
|
setups).
|
|
- **Synology** — volume and disk health (read-only by design).
|
|
Supports self-signed certificates.
|
|
- **Semaphore** — Ansible run status per template across every
|
|
project, with a "Run" action to trigger a template; a live
|
|
template-count widget (with a last-failed warning).
|
|
- **Gitea** — repo list with each repo's last CI run status, and
|
|
re-running just the failed jobs in a run; a live repo-count widget
|
|
(with a failing-build warning).
|
|
- **Docker** — container status across every Docker host Dockhand
|
|
manages (one credential covers all of them), with
|
|
start/stop/restart actions, a host filter, image-update status per
|
|
container (from Dockhand's own cached update check, plus a button
|
|
to trigger a fresh one), and a live running/total widget (with an
|
|
updates-available count).
|
|
|
|
The Integrations page itself is now just a list of configured
|
|
integrations (name/type/status, visible to every role) with an
|
|
admin-only "Add integration" button and edit/enable/disable/delete
|
|
actions per row — the six dedicated pages above are where you
|
|
actually use each one.
|
|
- Every table in the app is click-to-sort on any column (numbers, booleans,
|
|
and dates/text sort correctly regardless of how the column formats them)
|
|
and has an "Export CSV" button next to it that exports whatever's
|
|
currently sorted/filtered. Tables that can realistically grow large
|
|
(DNS zones/records, IP Addresses, Secrets, Servers, Audit Log, and each
|
|
integration's device/container/guest/repo/template list) are paginated,
|
|
20 rows per page by default — adjustable under Settings → Display — and
|
|
CSV export still covers every sorted/filtered row, not just the current
|
|
page.
|
|
- **Settings** (admin-only) — notification channels (Gotify, ntfy, SMTP,
|
|
generic webhook) with per-channel test buttons, per-event toggles (DNS
|
|
record added/updated/deleted, daily secret-expiry reminder and daily
|
|
Tailscale key-expiry reminder — both sharing one configurable
|
|
time/timezone), badge-color customization for both DNS
|
|
providers and integration types, and a **Display** tab (date order,
|
|
12/24-hour clock, and rows-per-page for every paginated table) applied
|
|
consistently across the app.
|
|
|
|
All six integrations follow the same config-in-UI + encrypted-credentials
|
|
pattern, added (and edited — e.g. to rotate an expired API token without
|
|
recreating the whole integration) through **Integrations → Manage
|
|
integrations**. See [INTEGRATIONS.md](INTEGRATIONS.md) for exactly what
|
|
credential to create and what access it needs in each target system,
|
|
for every integration and DNS provider.
|
|
|
|
**Verified for real, end to end**: every module above — including all six
|
|
integrations, both their read-only views and their write actions
|
|
(start/stop/restart, trigger-a-run, authorize/deauthorize) — has been
|
|
exercised against the user's actual live homelab, not just built against
|
|
specs. That pass also found and fixed two real bugs: the Synology adapter
|
|
assumed HTTPS-only (the NAS is reached over plain HTTP), and the Tailscale
|
|
adapter read `online`/`isExitNode` fields that don't actually exist in the
|
|
real API response (fixed to derive them from `connectedToControl` and
|
|
`enabledRoutes`). See the git log for the full verification notes per
|
|
integration.
|
|
|
|
Server and storage health is watched every 15 minutes: a server whose agent
|
|
stops reporting, a server disk / Proxmox storage / Synology volume passing a
|
|
usage threshold, and a Synology volume or disk that's degraded or failing each
|
|
raise one notification when the problem starts and one when it clears (both
|
|
thresholds are set under Settings → Notifications). Active problems are
|
|
remembered across restarts, so a rebuild doesn't re-alert them.
|
|
|
|
**Tags** — servers can be tagged (prod, media, rack-1, …) from their detail
|
|
page by operators and admins. Tags show as coloured chips on the Servers page,
|
|
which can be filtered by one or several tags (the filter is in the URL, so a
|
|
tag on a server's page links to everything sharing it), and they're searchable
|
|
from the global search box.
|
|
Admins manage tags under Settings → Tags: add tags before any server uses them (they're
|
|
offered as one-click suggestions when tagging), give any tag a colour of your choosing
|
|
(or leave it on the automatic one), rename tags, and delete them. Renaming to a name
|
|
that already exists merges the two, and both rename and delete rewrite every server
|
|
that carries the tag.
|
|
|
|
**Privacy** — a page every signed-in user can open that says what this
|
|
installation stores (accounts, sign-in sessions with their IP and browser, the audit
|
|
and diagnostic logs, server reports, the secrets tracker, credentials), where data
|
|
goes (Authentik, your integrations and DNS providers, the notification channels
|
|
that are switched on, domain registries), what's kept in the browser, who can see
|
|
what, and how to limit or remove data. Retention, integrations and channels are
|
|
read live from the installation; channel addresses are shown to admins only. Each
|
|
user sees their own account and sign-ins there and can download their own data
|
|
(account, sign-ins, audit-log entries) as a JSON file.
|
|
|
|
**Consistency** — a report of where IPAM, DNS and your servers disagree about
|
|
an address: the same address on two servers, a DNS record named after a server
|
|
that points somewhere it isn't, an IPAM entry labelled with a server's name at
|
|
the wrong address, addresses in use that IPAM doesn't list (with an "Add to
|
|
IPAM" button), and server addresses no DNS record points at. It only compares
|
|
data the app already holds — nothing is fetched when you open it — so it says how
|
|
many servers reported addresses and how fresh the synced DNS zones are. Only
|
|
private addresses are compared; ranges you exclude (Docker's, which repeat the same
|
|
subnet on many hosts — 172.16.0.0/12 is excluded by default, remove it if that's a
|
|
real LAN for you) are left out of every source; anything that's fine on purpose
|
|
can be ignored with a reason, and stays ignored.
|
|
|
|
**Domains** — when each domain registration expires, read from the registry
|
|
itself. The domains behind your DNS zones are picked up automatically (a zone
|
|
like `lab.example.se` resolves to the `example.se` registration that actually
|
|
expires); others can be added by hand. Each is looked up daily over RDAP where
|
|
the TLD offers it, and otherwise over WHOIS via IANA's referral — which is what
|
|
makes `.se`, `.nu` and `.io` work, since those aren't in the RDAP bootstrap.
|
|
You're reminded daily from N days before expiry (Settings → Notifications,
|
|
default 30) until it's renewed, and told when an expiry date couldn't be
|
|
refreshed for several days. Registries that don't publish an expiry (`.de`,
|
|
`.eu`) can be tracked but have no date to warn about. Private zones (`.lan`,
|
|
`.local`) are skipped.
|
|
|
|
**Automation failures** — every 15 minutes the app looks at the latest run of
|
|
each Semaphore template and each Gitea repo's latest workflow run. A failed one
|
|
raises a single notification (with the project/template or repo, run number, and
|
|
for Gitea the run's link), and another when a later run succeeds. It's
|
|
state-based, so a job that fails every night alerts on the first failure rather
|
|
than every night. A run that's still going, was cancelled, or was stopped by
|
|
hand leaves things as they were, and anything that couldn't be read (a Semaphore
|
|
project or Gitea repo that errored, or an integration that's down) is neither
|
|
cleared nor re-announced. The first check after upgrading only records what's
|
|
already failing, so old failures aren't announced. Toggle it under Settings →
|
|
Notifications. For Gitea this follows the repo's most recent run on any
|
|
workflow or branch, the same as the Gitea page shows.
|
|
|
|
**Maintenance mode** silences alerts about one server, integration, or DNS
|
|
provider while you work on it (server offline / disk, storage and Synology
|
|
health, Proxmox backup alerts, and "integration down" for that service type).
|
|
Every window has a fixed end (5 minutes to 7 days) and expires on its own, and
|
|
a problem that began during a window and is still present when it ends alerts
|
|
then — a forgotten window can't hide an outage. A banner shows what's currently
|
|
silenced to every signed-in user.
|
|
|
|
**Ports** — each server's detail page has a Ports card for finding free ports
|
|
and remembering what each one is for. "Scan…" runs a TCP scan of a port range
|
|
from the app against the server's address (private addresses only, up to
|
|
20,000 ports at a time) and lists what's open plus the ranges that were
|
|
confirmed free; a free range can be clicked to reserve a port. Any port can
|
|
carry a service name and a comment, and a port with a note counts as taken
|
|
even when nothing is listening. The agent also reports what is bound on the
|
|
host (`ss`), which catches services listening on localhost only — a scan from
|
|
elsewhere can't see those, so they'd otherwise look free. Re-run the agent
|
|
install one-liner on a host to pick that up.
|
|
|
|
The app is installable as a PWA — "Install app" / "Add to Home Screen" from the
|
|
browser gives it its own icon and a standalone window on phone or desktop. This
|
|
needs the site to be served over HTTPS (browsers only offer install on secure
|
|
origins; `localhost` also counts). The bundled service worker deliberately
|
|
caches nothing, so an installed copy always shows the current build.
|
|
|
|
## Requirements
|
|
|
|
- Node.js 20+
|
|
- An Authentik instance reachable from wherever this app runs
|
|
|
|
## 1. Set up an Authentik application
|
|
|
|
1. Create an **OAuth2/OpenID Provider**:
|
|
- Redirect URI: `<APP_BASE_URL>/auth/callback`
|
|
- Scopes: `openid`, `email`, `profile`
|
|
2. Create an **Application** using that provider, and assign the users/groups
|
|
who should be able to sign in — Authentik controls who can authenticate;
|
|
the app's own admin/operator/viewer roles control what they can do once in.
|
|
3. Copy the provider's issuer URL, client ID, and client secret into `.env`.
|
|
|
|
## 2. Local development
|
|
|
|
```bash
|
|
cp .env.example .env # fill in AUTHENTIK_*, SESSION_SECRET, CREDENTIALS_ENCRYPTION_KEY
|
|
npm install
|
|
npm run dev:server # http://localhost:3000 (API)
|
|
npm run dev:web # http://localhost:5173 (Vite dev server, proxies /api and /auth to :3000)
|
|
```
|
|
|
|
Visit `http://localhost:5173` during development. Database migrations run
|
|
automatically on server start. SQLite data lands in `./data` (gitignored).
|
|
|
|
Generate `SESSION_SECRET` and `CREDENTIALS_ENCRYPTION_KEY` with:
|
|
|
|
```bash
|
|
node -e "console.log(require('crypto').randomBytes(32).toString('hex'))"
|
|
```
|
|
|
|
## 3. Run with Docker
|
|
|
|
```bash
|
|
cp .env.example .env
|
|
# edit .env
|
|
docker compose -f docker-compose.dev.yml up -d --build # build locally
|
|
# or, once an image is published to your registry:
|
|
docker compose up -d
|
|
```
|
|
|
|
The app listens on `HOST_PORT` (default `3000`); SQLite data persists in
|
|
`./data` on the host.
|
|
|
|
### arm64 (Raspberry Pi, Apple silicon, ARM servers)
|
|
|
|
Two compose files, one to build the image and one to run the published one:
|
|
|
|
```bash
|
|
# On any machine with Docker: build the arm64 image and push it to the registry
|
|
docker compose -f docker-compose.arm64.build.yml build
|
|
docker compose -f docker-compose.arm64.build.yml push
|
|
|
|
# On the arm64 machine: pull that image and run it (no build there)
|
|
docker compose -f docker-compose.arm64.yml pull
|
|
docker compose -f docker-compose.arm64.yml up -d
|
|
```
|
|
|
|
`docker-compose.arm64.build.yml` can also build and run right on an arm64 machine
|
|
(`up -d --build`). Building on x86 goes through QEMU emulation — slower, and it needs
|
|
a one-time `docker run --privileged --rm tonistiigi/binfmt --install arm64`.
|
|
|
|
Both files use the image `gitea.labsconnect.se/bobban/homelabmanager-homelab-manager:arm64`.
|
|
Set `IMAGE_REPO` and/or `ARM64_TAG` in `.env` to use another registry or to pin a release
|
|
tag (e.g. `ARM64_TAG=1.4.0-arm64`). The arm64 tag is kept separate from the default image,
|
|
so the existing `docker-compose.yml` is unaffected. A `.dockerignore` keeps host
|
|
`node_modules`, `.env` and `data/` out of the build.
|