Commit Graph
12 Commits
Author SHA1 Message Date
bobbanandClaude Sonnet 5 9f1609c4ed Add tags to servers, with a tag filter on the Servers page
Servers can be tagged (prod, media, rack-1, ...) for grouping. Operators
and admins edit tags inline on a server's detail page; the input suggests
tags already used on other servers so the same word ends up spelled the
same way everywhere. Tags show as chips that keep one colour per tag, on
the server cards, in the Manage table (and its CSV export), and on the
detail page, where each chip links to the Servers list filtered by that
tag. The Servers page has a tag bar with counts; picking several tags
narrows to servers that have all of them. The filter lives in the URL, so
it survives a refresh and can be linked to. Tags are also matched by the
global search.

Tags are normalized on the server (trimmed, lowercased, spaces become "-",
duplicates merged, sorted); letters in any language are allowed, plus
digits and - _ . : /, at most 30 characters and 12 per server. Invalid
input is rejected with a message naming the offending tag, and nothing is
saved. Changes are audit-logged with the before and after lists.

Stored as a JSON column on servers (migration 0010). Every server response
now returns tags as an array, and the shared response shaping strips the
token hash in one place instead of five.

Verified with 27 backend checks (normalization edge cases including
Swedish letters, roles, validation, search, audit, PATCH/detail/list
shapes) and by driving the real Servers and detail pages against the real
router in a browser (filtering, editing, invalid tag, viewer view). Real
dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 03:06:29 +02:00
bobbanandClaude Sonnet 5 4c11158e98 Add a Ports card to server pages: scan for open ports, find free ones, and keep notes
Each server's detail page now has a Ports card. "Scan…" runs a TCP connect
scan of a chosen range from the app and shows what's open, along with the
ranges that were actually confirmed free; clicking a free range starts a
reservation. Any port can carry a service name and a comment, so the page
also answers "what is this port for". A port with a note counts as taken
even when nothing is listening, which is what makes a reservation work.
Operators can scan and edit; everyone can read. Scans and note changes are
audit-logged.

Details that matter for correctness:
- "Free" means the host actively refused the connection AND nobody has
  claimed the port. A port that never answers (firewall drop, host down)
  is reported as not answering, not as free.
- A scan from elsewhere can't see services bound to localhost only, so the
  agent now also reports what is bound on the host (ss -tulnp) and those
  ports are treated as taken. They show as "local only". Existing agents
  keep working; re-run the install one-liner to add this. The field is
  validated leniently so one odd line can never cost an agent its whole
  report, tasks included.
- If nothing answers at all during a scan, existing results are left
  alone instead of being marked all-closed.
- Scan targets are limited to private addresses (RFC1918, Tailscale
  100.64/10, link-local, IPv6 ULA/link-local); loopback and public
  addresses are refused. Ranges are capped at 20,000 ports, and only one
  scan runs per server at a time.
- Rows exist only while they carry information: an open port, or one with
  a note. A closed port with no note disappears on the next scan; one with
  a note stays as "reserved".

New table server_ports plus two columns on servers (migration 0009).

Verified with 76 backend checks (scanner open/refused/filtered, address
rules, agent report leniency, note/reserve/clear semantics, free-range
calculation including the localhost-only case, roles, concurrency lock,
no-response guard, audit entries, cascade delete) and by driving the real
component against the real router in a browser. Real dev database mtime
untouched.

Not verified: the agent's ss/awk/jq pipeline on a real host — the awk step
was checked against sample ss output and the script passes bash -n, but
jq isn't available here to run the whole thing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 02:39:43 +02:00
bobbanandClaude Sonnet 5 aae4f0d74f Add maintenance mode to silence alerts while working on a server or integration
Rebooting Proxmox or patching a server triggered failure/offline alerts
you then had to dismiss. A maintenance window silences alerts about one
server, integration, or DNS provider for a chosen time. New Maintenance
page (start with a duration and optional reason, end early, see what's
silenced and what isn't) and a banner in the app shell so every signed-in
user can see what is currently silenced. Starting/ending is operator-only
and audit-logged; starting one on a target that already has a window
restarts its clock instead of stacking.

Silenced for the target: server offline/disk alerts, Proxmox/Synology
storage and health alerts, Proxmox backup alerts, and "integration down"
alerts. Not silenced: expiry and update reminders, DNS change notices.

The design goal is that this cannot hide a real outage:
- Every window has a required end (5 min to 7 days); there is no
  open-ended option, so a forgotten window expires by itself.
- A silenced problem is deliberately NOT recorded as "known". If it is
  still present when the window ends it alerts then, as new. A problem
  that was already alerted before the window stays known, so it isn't
  repeated, and is reported cleared only after the window ends.
- Failure alerts keep counting failures during a window without marking
  themselves alerted, so an outage that outlasts the window alerts on the
  very next failed call.

Known limitation, stated on the page: integration-failure alerts are
tracked per service TYPE (all "proxmox"), not per configured instance, so
a window on one Proxmox integration also silences a failure on a second
Proxmox integration while it's open. Fixing that means threading the
integration id through every adapter and the diagnostic log, which is a
much larger change than this feature.

Also moved the API-error-message helper out of Secrets.tsx into a shared
util now that two pages use it. New table maintenance_windows (migration
0008).

Verified with 44 checks: the condition-key-to-subject mapping (including
server:3 vs server:33), the diff rules with silenced subjects (new problem
not recorded, alerts when the window ends; already-known one carried and
not repeated; clears only after the window), window expiry and
integration/DNS-provider source matching, the failure tracker end to end
against a webhook (silent during a window while an unrelated service still
alerts; outage that outlasts the window alerts on the next failure and
only once; fail-and-recover fully inside a window sends nothing), a full
health pass against a real window, and the real router with a stubbed
session (role rules, duration bounds including the missing-duration case,
extend-not-stack, 404s, deleted targets hidden, audit entries). Real dev
database mtime untouched.

Not done: I haven't clicked through the new page or banner in a browser
(they sit behind the Authentik login); it builds and the API behind it is
tested.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 02:23:06 +02:00
bobbanandClaude Sonnet 5 7e81306aa7 Read SSL certificate expiry from the live server instead of trusting a typed-in date
A certificate secret's expiry was only ever what someone typed in, so
a renewed cert (or a wrong date) meant the app's reminders were
silently wrong. A certificate secret can now be given a host:port; the
app opens a real TLS connection and reads the certificate's actual
expiry — on create/edit (if the host changes), daily, and via a
per-row "Check now" — and keeps expiryDate in sync. Because the daily
refresh runs before the existing expiry check, the reminder is always
computed from what's actually being served.

Verification is deliberately off for the connection: homelab services
routinely serve self-signed/internal-CA certs, and an already-expired
one is exactly the case worth reporting, which a verifying connection
would refuse before exposing the dates.

Failure handling avoids the silent-staleness this is meant to fix: a
failed check keeps the last known date, records why on the row (shown
as a "Check failed" badge), and is listed in the daily secrets
notification. Creating a monitored secret whose host can't be reached
and with no manual date is rejected with the reason rather than saved
blank. A non-TLS port (the likeliest typo) gets a plain-language
error instead of raw OpenSSL output.

Server-side connections to a user-supplied host:port need the same
operator role that already gates editing secrets (and running Semaphore
templates, which is strictly more powerful); the host is validated
against a strict character set before any connection is made.

New nullable secrets columns (check_host, check_port, last_checked_at,
last_check_error) via migration 0007; existing rows are unaffected.

Verified against real TLS servers (openssl-generated certs) and the
real secrets router with a stubbed session: a live 45-day cert read
back as the correct date via both an IP host (no SNI) and a hostname;
an already-expired cert reported its past date and shows as expired;
refused connections, a server that accepts but never answers (times
out), and a plain non-TLS server each produced a descriptive error
rather than a hang or crash. Through the router: create with a host
and no date reads the date; unreachable host with no date -> 400 with
the reason; unreachable with a manual date -> saved with the error
recorded; host on a non-certificate type and an invalid host string
-> 400; a hand-typed date on a monitored secret is ignored; changing
the host re-checks immediately; changing the type away from
certificate ends monitoring; a viewer gets 403 on Check now. 23 checks,
all passing (a first re-run showed 2 spurious failures that were leftover
rows from the previous run's scratch database, confirmed by a clean re-run).
Real dev database mtime untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-25 23:41:29 +02:00
bobbanandClaude Sonnet 5 73d649377e Add notification quiet hours with digest delivery
Instant notifications (DNS changes, integration failure/recovery
alerts) had no way to avoid pinging overnight. Adds a "Quiet hours"
window under Settings -> Notifications: notifications that would fire
during the window are held in a new notification_queue table instead
of sent immediately, then delivered as one combined digest at the end
time (in the same timezone already used for the daily checks) via a
new scheduled flush job. The scheduled daily checks (secret expiry,
Tailscale key, Docker updates, Proxmox backups) already only fire
once at a chosen time, so this mainly matters for the instant ones.
Includes a live "N queued" indicator with a manual "Flush now" button
for visibility, and correctly falls back to sending immediately
whenever the feature is disabled (the default).

Gating lives at the single choke point every notification already
flows through (notify()), so no per-event-type wiring was needed.

Verified against a fake webhook receiver on an isolated scratch
database: exhaustively checked the midnight-wraparound window math
(9 cases including exact-boundary inclusive/exclusive edges) against
synthetic "now" values rather than depending on when the test
happens to run, then end-to-end through the real notify()/
flushQuietHoursQueue() functions with a window constructed around the
actual current time — confirmed a notification during the window
queues instead of sending, the flush produces one digest with the
original title/message intact and clears the queue, a notification
outside the window sends immediately, and disabling the feature
entirely sends immediately regardless of the window. Confirmed the
real dev database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 20:56:33 +02:00
bobbanandClaude Sonnet 5 08a984719f Let admins hide the Proxmox-link card per server
Not every registered server is a Proxmox VM/LXC -- bare-metal boxes
and other hosts had no reason to show a "link to Proxmox" option, but
it appeared unconditionally on every server's detail page.

New hideProxmoxLink column on servers (default false, so existing
behavior is unchanged until someone opts in). A "Not a VM? Hide this"
link in the card's header sets it; once hidden, a small "+ Show
Proxmox link options" link takes its place so it's still reachable,
not buried in a settings form. The card always shows regardless of
this flag once a server IS actually linked, so unlinking never becomes
unreachable by hiding the card out from under an active link.

Verified against a temp SQLite DB with real migrations: a new server
defaults to false, and toggling true/false both persist correctly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-18 21:34:52 +02:00
bobbanandClaude Sonnet 5 655862c94e Add per-server admin-page links (Dockge, Webmin, Cockpit, etc.)
Each server's detail page gets an "Admin Links" card for bookmarking
that host's own web UIs -- container managers, Webmin, Cockpit, or
anything else reachable by URL -- so there's a quick way to jump
there without hunting down the address each time.

New server_links table (serverId FK, label, url, ON DELETE CASCADE so
removing a server cleans up its links automatically) and three new
routes: POST/PATCH/DELETE /api/servers/:id/links, gated to operator+
like the rest of this page's editing actions; the existing GET
/:id/detail now includes the server's links alongside hardware/DNS
info. URLs are validated to start with http:// or https:// server-side
(rejecting e.g. a javascript: URL that would otherwise render as a
clickable link).

Rendered as a row of pill buttons (label opens the URL in a new tab),
with inline edit/remove controls next to each when the viewer can
edit, and an "Add link" form matching the existing task-form style on
the same page.

Verified against a temp SQLite DB with real migrations: create/list/
update/delete all work, and deleting the parent server cascades to
remove its links rather than leaving them orphaned.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 22:51:10 +02:00
bobbanandClaude Sonnet 5 e35da87886 Add a Diagnostic Log, ported and generalized from Sloth Manager
Sloth Manager tracked every API call made to DNS providers for
connectivity troubleshooting. Port that here, generalized to cover
every outbound integration this app makes, not just DNS -- Tailscale,
Proxmox, Synology, Semaphore, Gitea, and Dockhand calls now show up
too, since a broken API token or unreachable host on any of them is
just as worth diagnosing.

New services/diagLog.ts: a generic withDiagLogging(source, adapter)
wraps every async method of any adapter object with timing +
success/failure recording, without touching a single adapter's
request/error-handling internals -- every DNS and integration adapter
interface here is already just a flat set of async methods, so this
one wrapper works for all twelve of them. Applied it at each adapter
factory's own return statement (one line each) rather than at the
route layer, so background jobs that construct adapters directly
(the Tailscale key-expiry scheduler, IPAM sync, agent-driven Proxmox
lookups) get logged too, not just requests through routes/integrations.ts.

New diag_log table (ring-buffered to the last 500 rows, mirroring
Sloth Manager's approach -- this is for live troubleshooting, not a
durable record) and admin-only GET/DELETE /api/diag-log routes, source/
result filters, pagination.

New admin-only Diagnostic Log page: filterable, paginated table with a
Clear button. Also introduces the shared useSortable hook + SortableTh
component used here for the first time -- a follow-up commit applies
the same sorting (and CSV export) to the rest of the app's tables, per
the same request.

Verified end-to-end against a temp SQLite DB with real migrations: a
fake wrapped adapter's successful and failing calls both land correctly
in the log with the right source/operation/latency/error, and the
source/ok filters and clear-log operation all behave correctly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 21:41:06 +02:00
bobbanandClaude Sonnet 5 d714a87754 Add a "Sync from Tailscale" action to IP Addresses (IPAM)
IPAM was entirely manual — Tailscale device IPs never showed up there
even though the Tailscale integration already lists them. Add an
explicit sync action (matching this app's existing pattern of
user-triggered syncs rather than silent background polling).

- New source column on ipam_entries (null = manual, "tailscale" =
  auto-synced) so a re-sync only ever touches rows it created itself —
  a manually-entered IP that happens to collide with a tailnet address
  is left untouched and reported back as skipped, never overwritten.
- POST /api/ipam/sync-tailscale pulls every enabled Tailscale
  integration's device list, upserting by primary IP (label, OS in
  notes, vendor "Tailscale"); one unreachable Tailscale integration
  doesn't block others.
- New "Sync from Tailscale" button on the IP Addresses page, with a
  small "synced" badge marking which rows came from it.

Also fixed a longstanding TODO found in the same file: the "DNS
records" column always showed "-" because matchingDnsRecords was
hardcoded to an empty array from before the DNS module existed. It
now does the same content-based reverse lookup against the DNS
module's record cache used elsewhere in the app.

Verified end-to-end by running the real server with Tailscale's fetch
call intercepted at the process level (its adapter hardcodes
api.tailscale.com with no configurable URL, so it can't be pointed at
a mock server the way Proxmox/Synology can): confirmed add, the
manual-entry skip/never-overwrite behavior, idempotent re-sync
(add -> update), and the DNS-matching fix, all against the real
route and adapter code.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-15 23:16:17 +02:00
bobbanandClaude Sonnet 5 b9409d3095 Add server hardware/network detail view with Proxmox sync and agent reporting
Servers & Tasks only tracked scheduled tasks — there was no overview of
the servers themselves and no way to see CPU/RAM/disk/IP info. Add a
clickable server overview (visible to every role, not just admins) that
opens a per-server detail page.

- servers table gains an agent-reported hardware/network snapshot
  (IPs, CPU model/cores/load, memory, disks) and an optional link to a
  Proxmox VM/LXC (integration + node + guest type + vmid).
- Proxmox adapter gains getGuestDetail(): live cores/memory/disk from
  /config, live cpu/mem/uptime from /status/current, and IPs (LXC net
  config directly, QEMU via a best-effort guest-agent call that degrades
  gracefully when the agent isn't installed).
- New GET /api/servers/:id/detail combines whichever hardware source
  applies (live Proxmox vs. last agent report) with a DNS reverse-lookup
  against the DNS module's own record cache, so matching hostnames show
  up next to each IP. New PATCH /api/servers/:id manages the Proxmox
  link.
- agent/linux/report-tasks.sh now also collects and reports IPs,
  CPU/memory/disk info on every check-in (load-average-based CPU number,
  not instantaneous, to keep the agent a cheap oneshot).
- Extracted the per-server task table into a shared ServerTaskTable
  component so the all-servers view and the new detail page render
  tasks identically.

Verified server-side end-to-end against the real dev server (agent
report -> detail endpoint -> DNS match) and the Proxmox adapter against
a mock HTTPS server covering LXC/QEMU config parsing and the
guest-agent-unavailable fallback.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-15 13:02:50 +02:00
bobbanandClaude Sonnet 5 39b1d1fe2e Add DNS module ported from Sloth Manager
Ports zone/record management across Cloudflare, Loopia, Pi-hole, Azure DNS,
cPanel, and Technitium onto the new stack. Unlike the original (one instance
per provider configured via env vars), providers are now configured through
the UI and support multiple named instances per type, with API
credentials encrypted at rest via the integration_credentials table.

- server/src/dns/adapters/*: each provider ported to a config-based factory
  (no more process.env reads), preserving each provider's original quirks
  (Pi-hole session auth, Azure record-set merging, Loopia XML-RPC, cPanel
  UAPI/API2 fallbacks, Technitium composite record IDs).
- server/src/routes/dns.ts: provider CRUD (admin), a "test connection"
  endpoint, and zone/record browsing+sync+CRUD (operator+), all audit-logged.
- dns_zones_cache/dns_records_cache tables replace Sloth Manager's
  dns-cache.json file, keeping the same "cache is the source of truth for
  display, sync fetches fresh from the provider" behavior.
- web: a DNS page with provider management, zone browsing, and a record
  editor, plus a reusable dynamic provider-config form.

Verified: full build (tsc + vite) passes; a scripted HTTP-layer test against
a running server exercises auth, role gating (403 for viewer), validation
(400 on missing config fields), provider CRUD, credential non-leakage in
list responses, and adapter error propagation (502 against an unreachable
host) — all passing. Real provider connectivity still needs to be checked
against the user's actual DNS accounts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-14 23:00:30 +02:00
bobbanandClaude Sonnet 5 6bd2ed52c1 Scaffold Homelab Manager foundation
Monorepo (Express+TS+Drizzle/libSQL server, React+Vite+Tabler web) matching
the stack used by ScheduleTaskManager and Sloth Manager. Includes Authentik
OIDC login with local admin/operator/viewer roles (first user becomes admin),
a generalized audit log, encrypted-at-rest storage for future integration API
tokens, the DB schema for all planned modules, and the Tabler-styled app
shell/nav. Also ports the Secrets (expiry tracker) and IP Addresses (IPAM)
modules from Sloth Manager onto the new stack.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-14 21:57:13 +02:00