Commit Graph
6 Commits
Author SHA1 Message Date
bobbanandClaude Sonnet 5 fea20456e4 Add a Windows agent (PowerShell)
Reports a Windows machine the way the Linux agent does, replacing the
"planned" stub in agent/windows: scheduled tasks plus hostname, IPv4
addresses, CPU model/cores/current load, memory, every fixed disk, and
TCP/UDP listening ports with the owning process (which feed the Ports
card, localhost-only listeners included).

Scripts (plain ASCII by design -- they are downloaded as text and Windows
PowerShell 5.1 reads BOM-less files as ANSI):
- report-tasks.ps1: collects and POSTs to /api/agent/report. Works in
  Windows PowerShell 5.1 and PowerShell 7. -DryRun prints the JSON.
  Microsoft's own \Microsoft\ tasks (hundreds) are left out unless
  INCLUDE_MICROSOFT_TASKS is set. Triggers are turned into readable text
  ("Weekly on Mon, Wed at 03:00", "At logon", "..., repeating every 15 min").
  Self-signed certificates work via API_INSECURE on both PowerShell
  versions (they need different mechanisms).
- install.ps1: elevated only; downloads the agent to ProgramData, writes
  agent.json with permissions locked to SYSTEM and Administrators *before*
  the token goes in, and registers a SYSTEM scheduled task (every 15 min
  plus at startup with a 2 min delay). Reinstalling replaces the task.
- uninstall.ps1: removes the task and only the files the agent installed.

Server: accepts schedule_type "windows_task"; a server can be registered
as Windows (Add a server has an operating system choice); an agent's
reported os_type ("linux"/"windows", anything else ignored) corrects the
stored one. The Servers page shows the right install and uninstall
command for each OS (Windows PowerShell 5.1 one-liners, with a self-signed
variant and a note about PowerShell 7), and Windows tasks are labelled
"Windows scheduled tasks". The Linux commands are unchanged.

Verified on this Windows machine, in both PowerShell 5.1 and 7:
- Real dry runs found and fixed bugs before anything shipped: tasks and
  ports came out as one nested item (return , $out wrapped twice), integer
  keys in an ordered dictionary index by position (wrong weekday names),
  and generic "Trigger" labels.
- End to end against the real agent-report router: HTTP, self-signed HTTPS
  refused by default and accepted with API_INSECURE, wrong token gives a
  clear one-line error and exit 1, and Swedish letters plus a euro sign
  survive JSON -> UTF-8 -> HTTP -> SQLite.
- 35 checks on trigger/action/duration descriptions, 20 on the installer's
  building blocks (task parts built but not registered, credentials file
  content and ACL, download over HTTP and self-signed HTTPS), 18 on the
  server rules, and the generated one-liners run through PowerShell's
  parser. The documented one-liners were run through iex and stop at the
  administrator check without changing anything.
- Found that PowerShell 7 ignores the ServicePointManager certificate
  override, so the installer's own download now uses -SkipCertificateCheck
  there.

NOT verified: the elevated install itself. Registering a SYSTEM scheduled
task needs elevation and changes the machine, so it was not run: the task
registration, that the repeating trigger really runs indefinitely, and
the agent running as SYSTEM under Task Scheduler have not been exercised.
Windows 10 / Server 2016 or newer is assumed; older is untested.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-27 00:09:00 +02:00
bobbanandClaude Sonnet 5 4c11158e98 Add a Ports card to server pages: scan for open ports, find free ones, and keep notes
Each server's detail page now has a Ports card. "Scan…" runs a TCP connect
scan of a chosen range from the app and shows what's open, along with the
ranges that were actually confirmed free; clicking a free range starts a
reservation. Any port can carry a service name and a comment, so the page
also answers "what is this port for". A port with a note counts as taken
even when nothing is listening, which is what makes a reservation work.
Operators can scan and edit; everyone can read. Scans and note changes are
audit-logged.

Details that matter for correctness:
- "Free" means the host actively refused the connection AND nobody has
  claimed the port. A port that never answers (firewall drop, host down)
  is reported as not answering, not as free.
- A scan from elsewhere can't see services bound to localhost only, so the
  agent now also reports what is bound on the host (ss -tulnp) and those
  ports are treated as taken. They show as "local only". Existing agents
  keep working; re-run the install one-liner to add this. The field is
  validated leniently so one odd line can never cost an agent its whole
  report, tasks included.
- If nothing answers at all during a scan, existing results are left
  alone instead of being marked all-closed.
- Scan targets are limited to private addresses (RFC1918, Tailscale
  100.64/10, link-local, IPv6 ULA/link-local); loopback and public
  addresses are refused. Ranges are capped at 20,000 ports, and only one
  scan runs per server at a time.
- Rows exist only while they carry information: an open port, or one with
  a note. A closed port with no note disappears on the next scan; one with
  a note stays as "reserved".

New table server_ports plus two columns on servers (migration 0009).

Verified with 76 backend checks (scanner open/refused/filtered, address
rules, agent report leniency, note/reserve/clear semantics, free-range
calculation including the localhost-only case, roles, concurrency lock,
no-response guard, audit entries, cascade delete) and by driving the real
component against the real router in a browser. Real dev database mtime
untouched.

Not verified: the agent's ss/awk/jq pipeline on a real host — the awk step
was checked against sample ss output and the script passes bash -n, but
jq isn't available here to run the whole thing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 02:39:43 +02:00
bobbanandClaude Sonnet 5 ae41a02864 Fix agent silently failing to report on hosts without a cpuinfo model name
report-tasks.sh's new hardware-collection code ran under set -euo
pipefail, so a single failing command inside it aborted the whole
script before the report was ever sent — with no error message,
since nothing in that path had explicit error handling. On real
hardware this hit immediately: `grep -m1 "model name" /proc/cpuinfo`
exits 1 when there's no match, and many ARM boards (e.g. Raspberry Pi)
have no such line at all. Confirmed via journalctl showing the
systemd service failing every 15 minutes with exit 1 and zero output,
and via a minimal repro of the exact bash control flow.

Wrap the hardware/network collection call so any failure inside it is
non-fatal: task reporting (the actual core function) must never be
taken down by a quirk in the best-effort hardware-gathering code, on
this host or any other. Also fixed a related gap found while testing
the fallback path: the server only accepted the system field being
absent, not explicitly null (what the script now sends if collection
fails outright), which would have turned graceful degradation into a
rejected report.

Verified end-to-end against the real dev server: system:null, system
omitted, and a normal populated report all now return 202 and persist
correctly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-15 19:05:14 +02:00
bobbanandClaude Sonnet 5 b9409d3095 Add server hardware/network detail view with Proxmox sync and agent reporting
Servers & Tasks only tracked scheduled tasks — there was no overview of
the servers themselves and no way to see CPU/RAM/disk/IP info. Add a
clickable server overview (visible to every role, not just admins) that
opens a per-server detail page.

- servers table gains an agent-reported hardware/network snapshot
  (IPs, CPU model/cores/load, memory, disks) and an optional link to a
  Proxmox VM/LXC (integration + node + guest type + vmid).
- Proxmox adapter gains getGuestDetail(): live cores/memory/disk from
  /config, live cpu/mem/uptime from /status/current, and IPs (LXC net
  config directly, QEMU via a best-effort guest-agent call that degrades
  gracefully when the agent isn't installed).
- New GET /api/servers/:id/detail combines whichever hardware source
  applies (live Proxmox vs. last agent report) with a DNS reverse-lookup
  against the DNS module's own record cache, so matching hostnames show
  up next to each IP. New PATCH /api/servers/:id manages the Proxmox
  link.
- agent/linux/report-tasks.sh now also collects and reports IPs,
  CPU/memory/disk info on every check-in (load-average-based CPU number,
  not instantaneous, to keep the agent a cheap oneshot).
- Extracted the per-server task table into a shared ServerTaskTable
  component so the all-servers view and the new detail page render
  tasks identically.

Verified server-side end-to-end against the real dev server (agent
report -> detail endpoint -> DNS match) and the Proxmox adapter against
a mock HTTPS server covering LXC/QEMU config parsing and the
guest-agent-unavailable fallback.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-15 13:02:50 +02:00
bobbanandClaude Sonnet 5 9c2742c30f Add Tailscale integration and harden all routes against crash-on-throw
First of the six planned live integrations (Proxmox, Synology, Semaphore,
Tailscale, Gitea, Dockhand), reusing Sloth Manager's existing Tailscale
adapter logic. Generalizes the "integrations" table already scaffolded in
the foundation pass into a working config-in-UI + encrypted-credentials
flow, following the same pattern as the DNS providers module — an
"Add integration" form only offers types with an implemented adapter
(currently just Tailscale), so the framework is ready for the next five
integrations without further schema/plumbing changes.

- server/src/integrations/tailscale/adapter.ts: ported from Sloth Manager's
  backend/src/adapters/tailscale.js — listDevices/setAuthorized/deleteDevice
  against the Tailscale API, now config-based (tailnet + apiKey) instead of
  reading process.env, and returning a ping() result instead of throwing.
- server/src/routes/integrations.ts: generic integration CRUD (admin) + a
  "test connection" endpoint, plus Tailscale-specific device routes
  (dashboard-and-basic-actions depth per the plan: authorize/deauthorize/
  remove, gated to operator+, audit-logged).
- web: an Integrations page (provider-style manage/browse split, matching
  the DNS page's UX) with a device table, and a live Tailscale widget on
  the Dashboard.

Bug found and fixed while testing: hitting the Tailscale device routes on
a non-Tailscale integration row crashed the ENTIRE server process, not just
that request — the generic adapter registry throws for unimplemented types,
and that throw happened inside an async handler with no surrounding
try/catch, which Express 4 doesn't catch, so it became an unhandled
rejection that (on modern Node) kills the process. Fixed at the source
(check the row's type before ever constructing an adapter) and, since the
same "a helper throws before any local try/catch runs" shape existed
wherever a route calls into loadDnsProviderConfig/loadIntegrationConfig
(both call decryptSecret, which throws if CREDENTIALS_ENCRYPTION_KEY is
ever wrong/missing after data was already encrypted with a different key),
added a small asyncHandler() wrapper and applied it to every route handler
across every router — a single bad request should never be able to take
the whole app down for every user.

Verified: full build passes. Fresh HTTP-layer tests against a running
server (17 checks) cover not-implemented-type rejection, missing-field
validation, a real network call to api.tailscale.com with a bogus key
(clean ok:false, not a crash), role gating at every tier, credential
non-leakage, disabled-integration blocking, and the wrong_type case that
originally crashed the server — confirmed it now returns 400 cleanly and
the server stays up. Re-ran the existing DNS (13 checks) and Servers/Tasks
suites afterward to confirm the asyncHandler sweep didn't regress anything
— all passing. Authorizing/removing a real device still needs a real
Tailscale API key to verify end-to-end.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-14 23:31:19 +02:00
bobbanandClaude Sonnet 5 9712d611a6 Add Servers & Tasks module ported from Schedule Task Manager
Ports cron/systemd task tracking across Debian/Raspbian servers, including
the Linux push agent (install/report/uninstall scripts, rebranded from
"schedule-task-manager-agent" to "homelab-manager-agent") and the
manual-task-entry flow for things an agent can't see (Docker jobs, backups).
The schema (servers/scheduled_tasks tables) was already in place from the
foundation pass, so this is mostly a straight port of the original's
services/routes.

Deliberate change from the original: server/token management (which mints
agent credentials) is now admin-only rather than open to any logged-in user,
and manual task CRUD is gated to operator+ — consistent with how Secrets,
IPAM, and DNS already split "configure credentials" from "everyday edits"
across roles. All mutations are audit-logged.

- server/src/services/tokens.ts, taskSync.ts: ported near-verbatim (agent
  token hashing, the agent-sync-marks-missing-as-stale-not-deleted logic).
- server/src/routes/servers.ts, tasks.ts, agentReport.ts: same contract as
  the original (agent auth is a per-server bearer token, independent of the
  session-based requireAuth used everywhere else).
- web: a single Servers & Tasks page (filter bar, task table grouped by
  server/schedule type, manual task form, and an admin-only server
  management panel with token reveal + copyable install/uninstall commands),
  replacing the original's two separate pages/apps.

Verified: full build passes; a scripted HTTP test against a running server
covers unauthenticated access, role gating at each tier (admin-only server
mgmt, operator+ task mgmt), agent bearer-token auth (valid/invalid/rotated),
manual-vs-agent task edit protection, stale-marking on re-sync, and cascade
delete — 22/22 checks passing. Real agent installation on an actual
Debian/Raspbian host still needs to be tried on the user's network.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-14 23:16:58 +02:00