aae4f0d74f505940ce8071200f93afb39403c10e
45
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
aae4f0d74f |
Add maintenance mode to silence alerts while working on a server or integration
Rebooting Proxmox or patching a server triggered failure/offline alerts you then had to dismiss. A maintenance window silences alerts about one server, integration, or DNS provider for a chosen time. New Maintenance page (start with a duration and optional reason, end early, see what's silenced and what isn't) and a banner in the app shell so every signed-in user can see what is currently silenced. Starting/ending is operator-only and audit-logged; starting one on a target that already has a window restarts its clock instead of stacking. Silenced for the target: server offline/disk alerts, Proxmox/Synology storage and health alerts, Proxmox backup alerts, and "integration down" alerts. Not silenced: expiry and update reminders, DNS change notices. The design goal is that this cannot hide a real outage: - Every window has a required end (5 min to 7 days); there is no open-ended option, so a forgotten window expires by itself. - A silenced problem is deliberately NOT recorded as "known". If it is still present when the window ends it alerts then, as new. A problem that was already alerted before the window stays known, so it isn't repeated, and is reported cleared only after the window ends. - Failure alerts keep counting failures during a window without marking themselves alerted, so an outage that outlasts the window alerts on the very next failed call. Known limitation, stated on the page: integration-failure alerts are tracked per service TYPE (all "proxmox"), not per configured instance, so a window on one Proxmox integration also silences a failure on a second Proxmox integration while it's open. Fixing that means threading the integration id through every adapter and the diagnostic log, which is a much larger change than this feature. Also moved the API-error-message helper out of Secrets.tsx into a shared util now that two pages use it. New table maintenance_windows (migration 0008). Verified with 44 checks: the condition-key-to-subject mapping (including server:3 vs server:33), the diff rules with silenced subjects (new problem not recorded, alerts when the window ends; already-known one carried and not repeated; clears only after the window), window expiry and integration/DNS-provider source matching, the failure tracker end to end against a webhook (silent during a window while an unrelated service still alerts; outage that outlasts the window alerts on the next failure and only once; fail-and-recover fully inside a window sends nothing), a full health pass against a real window, and the real router with a stubbed session (role rules, duration bounds including the missing-duration case, extend-not-stack, 404s, deleted targets hidden, audit entries). Real dev database mtime untouched. Not done: I haven't clicked through the new page or banner in a browser (they sit behind the Authentik login); it builds and the API behind it is tested. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1688de3ea2 |
Alert when a server goes silent, a disk fills up, or a Synology volume degrades
The data was all being collected (agent last-seen, per-disk usage, Proxmox storage, Synology volume/disk health) but nothing acted on it, so a dead server or a full disk was only noticed by opening the right page. A new health pass runs every 15 minutes (matching the agent's default report interval) and raises one notification when a problem starts and one when it clears: a server's agent silent past a threshold (default 60 min), a server disk / Proxmox storage or root filesystem / Synology volume at or above a usage threshold (default 90%), and a Synology volume or disk that isn't "normal", has bad SMART, bad sectors past the threshold, or life remaining below it. Both thresholds and an on/off toggle live under Settings -> Notifications. The parts that make this trustworthy rather than noisy: - A problem is keyed by identity, so it alerts once and not every run; a shared Proxmox storage listed by every node is one problem, not one per node. - Active problems persist across restarts, so a rebuild doesn't re-alert everything already known. - If a source can't be read on a given run (Proxmox/Synology unreachable, one node lacking privileges) its existing problems are held, not reported "cleared" and then re-alerted when it comes back — the integration-failure alert already owns "the integration is down". - For 20 minutes after startup server-derived problems are held too: agents couldn't report while the app was down, so judging them then would report every server offline after any restart. - An offline server's disk figures are stale and are not judged; a server that never reported has no agent and raises nothing. - Tracking continues while the toggle is off (only sending is gated), so turning it back on doesn't dump every long-standing problem. Timestamps without a zone (SQLite's format) are read as UTC; the test runs on a UTC+2 machine, where reading them as local time gives a different answer. Verified with 32 checks: the evaluation rules and the state diff as pure functions (exact thresholds, the proxmox:1 vs proxmox:10 prefix trap, the flapping sequence), then a whole pass against a real Proxmox adapter talking to a fake HTTPS cluster (one node returning 403, the whole API down, a shared storage on two nodes, a node that recovers), a webhook receiver, the real DB, and the persisted state. Not exercised end-to-end: the Synology collection path — its rules are tested on data shaped exactly like the adapter's output types, but I did not stand up a fake DSM. Real dev database mtime untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
7e81306aa7 |
Read SSL certificate expiry from the live server instead of trusting a typed-in date
A certificate secret's expiry was only ever what someone typed in, so a renewed cert (or a wrong date) meant the app's reminders were silently wrong. A certificate secret can now be given a host:port; the app opens a real TLS connection and reads the certificate's actual expiry — on create/edit (if the host changes), daily, and via a per-row "Check now" — and keeps expiryDate in sync. Because the daily refresh runs before the existing expiry check, the reminder is always computed from what's actually being served. Verification is deliberately off for the connection: homelab services routinely serve self-signed/internal-CA certs, and an already-expired one is exactly the case worth reporting, which a verifying connection would refuse before exposing the dates. Failure handling avoids the silent-staleness this is meant to fix: a failed check keeps the last known date, records why on the row (shown as a "Check failed" badge), and is listed in the daily secrets notification. Creating a monitored secret whose host can't be reached and with no manual date is rejected with the reason rather than saved blank. A non-TLS port (the likeliest typo) gets a plain-language error instead of raw OpenSSL output. Server-side connections to a user-supplied host:port need the same operator role that already gates editing secrets (and running Semaphore templates, which is strictly more powerful); the host is validated against a strict character set before any connection is made. New nullable secrets columns (check_host, check_port, last_checked_at, last_check_error) via migration 0007; existing rows are unaffected. Verified against real TLS servers (openssl-generated certs) and the real secrets router with a stubbed session: a live 45-day cert read back as the correct date via both an IP host (no SNI) and a hostname; an already-expired cert reported its past date and shows as expired; refused connections, a server that accepts but never answers (times out), and a plain non-TLS server each produced a descriptive error rather than a hang or crash. Through the router: create with a host and no date reads the date; unreachable host with no date -> 400 with the reason; unreachable with a manual date -> saved with the error recorded; host on a non-certificate type and an invalid host string -> 400; a hand-typed date on a monitored secret is ignored; changing the host re-checks immediately; changing the type away from certificate ends monitoring; a viewer gets 403 on Check now. 23 checks, all passing (a first re-run showed 2 spurious failures that were leftover rows from the previous run's scratch database, confirmed by a clean re-run). Real dev database mtime untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ca61f2a915 |
Surface Proxmox VMs/LXCs with no backup coverage at all
A failing backup run is visible now, but a guest with no backup job covering it in the first place was still a silent gap. Rather than depending on Proxmox's /cluster/backup-info/not-backed-up-guests endpoint (only exists on newer PVE versions), this derives coverage from data already fetched: a guest counts as covered if any enabled job either lists its vmid directly, or backs up "all guests" (scoped to the job's node, if it has one) without excluding it. Adds a warning banner plus a full table to the Proxmox page's Backups card, and extends the existing daily "Proxmox backup failed" notification (relabeled to mention this too) to also list uncovered guests, gated by the same toggle. Verified the coverage logic directly (it's a pure function, so no fake server needed) across 7 cases: no jobs at all, an all-guests job with an exclude list, a specific-vmids job, a node-scoped job that shouldn't cover a guest on a different node, a disabled job providing no real coverage, two jobs whose combined scope covers everything neither would alone, and a realistic mixed scenario — all passed. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
73d649377e |
Add notification quiet hours with digest delivery
Instant notifications (DNS changes, integration failure/recovery alerts) had no way to avoid pinging overnight. Adds a "Quiet hours" window under Settings -> Notifications: notifications that would fire during the window are held in a new notification_queue table instead of sent immediately, then delivered as one combined digest at the end time (in the same timezone already used for the daily checks) via a new scheduled flush job. The scheduled daily checks (secret expiry, Tailscale key, Docker updates, Proxmox backups) already only fire once at a chosen time, so this mainly matters for the instant ones. Includes a live "N queued" indicator with a manual "Flush now" button for visibility, and correctly falls back to sending immediately whenever the feature is disabled (the default). Gating lives at the single choke point every notification already flows through (notify()), so no per-event-type wiring was needed. Verified against a fake webhook receiver on an isolated scratch database: exhaustively checked the midnight-wraparound window math (9 cases including exact-boundary inclusive/exclusive edges) against synthetic "now" values rather than depending on when the test happens to run, then end-to-end through the real notify()/ flushQuietHoursQueue() functions with a window constructed around the actual current time — confirmed a notification during the window queues instead of sending, the flush produces one digest with the original title/message intact and clears the queue, a notification outside the window sends immediately, and disabling the feature entirely sends immediately regardless of the window. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
99db7e1cf0 |
Surface Proxmox backup job status, with a daily failure notification
Proxmox already runs vzdump backups, but nothing in the app said
whether they were actually succeeding — a silent backup failure is
one of the more dangerous blind spots a homelab admin can have. Adds
a "Backups" card to the Proxmox page: configured backup job
schedules (storage target, which guests, enabled/disabled) from
GET /cluster/backup, and recent vzdump task history per node from
GET /nodes/{node}/tasks?typefilter=vzdump, with a banner at the top
if the most recent run didn't succeed.
New "Proxmox backup failed" notification toggle under Settings ->
Notifications, on the same daily schedule as the other checks. The
scheduler checks each node's own most-recent vzdump run independently
(not just the single most recent task overall) so one node's healthy
backup can't mask another node's failing one in a multi-node cluster.
Known limitation, documented in the adapter's own header comment:
Proxmox's task list doesn't reliably expose which specific guest
failed within an "all guests" job — only the task's own log text has
that — so this surfaces job- and task-level status rather than
guessing at per-guest outcomes.
Verified against a fake Proxmox server (real self-signed HTTPS, since
the adapter's node:https usage can't be monkey-patched under ESM)
reproducing the documented /cluster/backup and task-list response
shapes: job parsing (all-guests+exclude vs specific-vmids+disabled)
correct, task OK/failure parsing correct, and the critical multi-node
scenario confirmed — one node's failing latest run flagged, the
other's healthy latest run correctly left alone, with exactly one
notification of the right content. This reproduces Proxmox's
documented API shape rather than a live-verified one; flag if the
real cluster's response differs in some way this didn't anticipate.
Confirmed the real dev database's mtime was untouched throughout.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
6d673db9ec |
Add session management: see who's signed in, revoke a session
No visibility existed into who was currently signed in or a way to force a device out. Sessions already live as files via session-file-store, so this reads that store directly rather than adding a new DB table: new Settings-adjacent "Sessions" page (admin-only, alongside Users) lists every live session with the user's name/email/resolved role, IP, a friendly "Browser on OS" summary parsed from the user-agent, last-active time, and expiry, with a Revoke button per row (extra confirmation if you revoke your own current session, since that signs you out immediately). IP and user-agent are now captured into the session at login (auth/router.ts) since express-session doesn't track them itself. session-file-store's own Store type doesn't declare its list() method, so sessionStore.ts adds a narrow local interface for it rather than losing type safety on the rest of the store. Verified against a real session directory seeded through the actual session-file-store APIs (not hand-written JSON): confirmed correct field resolution including a session whose user row was later deleted (role resolves to null instead of crashing), correctly excluded a mid-OIDC-login session with no completed user yet, correctly excluded an already-expired session, and confirmed revoke actually deletes the right session file and only that one. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1cae35a59e |
Add a daily Docker image-update notification
Dockhand's pending-update counts were only visible if you happened to open the Docker page. New "Docker image update available" toggle under Settings -> Notifications, sharing the same daily time/timezone as the secret and Tailscale key expiry reminders (same node-schedule reschedule-on-settings-change pattern as those two). Reads each enabled Dockhand integration's already-cached update-check results via listContainers() rather than triggering a fresh per-container registry lookup, so it costs nothing extra beyond what the Docker page itself already fetches, and lists every container with an update pending across all environments/integrations in one notification. Verified end-to-end against a fake local Dockhand server (one environment, two containers, one flagged with a pending update) and a fake webhook receiver on an isolated scratch database: the check correctly found only the flagged container (with its newerVersion), sent exactly one notification with the right title/content, excluded the up-to-date container, and sent nothing at all when the setting was toggled off despite still finding the same pending update. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
009bb3e027 |
Add a global search / command palette (Ctrl/Cmd+K)
With 6 integrations, DNS, secrets, IPAM, and servers all in one app, there was no single place to type a hostname/IP/name and jump straight to it. New GET /api/search aggregates a LIKE-based search across servers, secrets, IPAM, integrations, DNS providers, DNS zones, and DNS records (joining zone/provider names onto each record result) in one round trip — homelab-scale row counts make a naive LIKE scan plenty fast, no FTS needed. Every underlying resource's own list endpoint already only requires requireAuth (viewer role included), so the aggregate endpoint uses the same single check. Frontend: a self-contained CommandPalette component (Tabler's modal CSS classes driven by React state, since the app doesn't load Bootstrap's JS) opens via a sidebar search button or Ctrl/Cmd+K from anywhere, debounces input, and supports arrow-key navigation. Results link to the right page: servers to their existing /servers/:id detail route; DNS zone/record results deep-link via new ?providerId= &zoneId= query-param handling added to Dns.tsx (auto-selects that provider/zone on load, since Dns.tsx previously held selection only in local state with no URL sync); secrets/IPAM results link with ?q= to prefill each page's existing client-side search box; integrations link to their type's dashboard page (no per-instance route exists yet, so same-type integrations share one link). Verified the query logic (joins, case-insensitive LIKE, correct zone/provider name resolution) against an isolated scratch database seeded with realistic cross-referencing rows — a server name match, a case-mismatched match against both a secret and a DNS provider sharing "cloudflare", an IP address matching both an IPAM entry and the DNS A record pointing at it (confirming the record's joined zone and provider names came through correctly), an integration name match, and a no-match query returning every category empty. Did not re-verify the requireAuth/asyncHandler wiring itself, since it's the same one-line pattern already proven across every other router in this app. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
23eb7f0d70 |
Add passphrase-protected export/import for integrations, DNS providers, and settings
Nothing let you back up or migrate the app's own configuration short of copying the raw SQLite file. Adds Settings -> Backup: export decrypts every integration/DNS provider credential (normally encrypted at rest with this server's CREDENTIALS_ENCRYPTION_KEY) and re-encrypts the whole payload with a passphrase you choose (scrypt- derived key, AES-256-GCM), so the file is portable to a different instance with a different encryption key rather than being tied to this one. Import decrypts with that passphrase and merges settings onto the current ones; integrations/DNS providers are only added when no existing row shares their type+name, so re-running an import never duplicates or overwrites a working credential. Scope is configuration only — no DNS records, secrets, IPAM, servers, or audit/diagnostic log data. Verified end-to-end against two isolated scratch databases with different encryption keys (proving actual cross-instance portability, not just round-tripping through the same key): export -> encrypt -> write file -> decrypt on the other DB -> import -> re-decrypt the newly created integration/provider using the target's own key, confirming the plaintext credentials survived correctly; a wrong passphrase failed loudly (GCM auth failure) as expected; and re-running the same import a second time skipped both rows instead of duplicating them. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b5a4c6e2d9 |
Alert when an integration or DNS provider fails repeatedly
The Diagnostic Log already records every outbound call's success or failure, but nothing acted on it — you'd only notice an integration was down by happening to open its page. Adds a per-source consecutive- failure counter (in-memory, reset on restart, same durability tier as the diag log's own ring buffer) hooked into recordDiagEntry: crossing the configurable threshold (default 3) sends one "down" notification on every configured channel, and a "recovered" notification fires once it succeeds again — no repeat spam while it stays down. New "Integration/DNS provider failing repeatedly" toggle and threshold field under Settings -> Notifications. Verified end-to-end against an isolated scratch database with a real local HTTP server standing in for the webhook channel: 5 consecutive failures produced exactly one "Down" notification (at the 3rd failure, correctly naming "3 calls"), a subsequent success produced exactly one "Recovered" notification, and two more failures on a fresh streak triggered nothing (below threshold) — confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
10d123b18a |
Add automatic retention purging for the Diagnostic and Audit logs
The diagnostic log already rings-buffer to 500 rows, but the audit log had no cap at all and would grow forever. Adds an opt-in age-based purge under Settings -> Logs: keep entries for N days, checked on a configurable interval (hourly through monthly), plus a manual "Purge now" button. Reuses the existing node-schedule-style reschedule-on-settings-change pattern from the secret/Tailscale expiry checkers, but as a plain setInterval since "how often" here is an interval rather than a specific daily time. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b40234a557 |
Make table page size a configurable setting, not a hardcoded 20
Pagination just landed hardcoded to 20 rows everywhere; add a way to change that instead of leaving it fixed for every table in the app. pageSize joins dateFormat/timeFormat on the existing display settings object (server-side default 20, 5-500 range enforced by the PUT schema) rather than becoming its own settings section, since it's the same kind of thing -- an admin-configured, globally-applied display preference read by every signed-in role via the already-public GET /api/settings/display endpoint, same as the date/time format already works. Renamed the "Date & Time" settings tab/page/route to "Display" (still just one component, now covering both date/time format and table pagination) since its scope no longer matches the old name -- kept Settings.tsx's usual pattern of one page per concern rather than adding a second, oddly-scoped tab just for one number field. New web/src/utils/pageSize.ts mirrors utils/date.ts's existing module-level "set once at startup, read anywhere without prop- drilling" pattern; usePagination()'s pageSize parameter now defaults to getPageSize() instead of a literal 20, evaluated fresh on every call so it picks up a saved change without touching any of the ten pages already using the hook. Verified server-side against a temp SQLite DB: pageSize defaults to 20, a partial update sets it without disturbing dateFormat/timeFormat and vice versa, and it persists across a fresh settings read. Also checked the default-parameter mechanics directly (re-evaluates the global value on every call rather than capturing it once, and an explicit override still wins). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
81e55fc792 |
Show per-disk usage for Proxmox-linked servers too, not just a total
The agent path showed a per-mount usage table; a Proxmox-linked server
only ever got a single "Disk (allocated)" figure, since that's all the
VM/LXC config alone can tell you -- it's the attached disk's declared
size, not how full it actually is inside the guest. Fix the actual gap
instead of just matching the display: fetch real usage where Proxmox
can see it.
adapter.getGuestDetail() gained a `disks` field (same {mount,
sizeBytes, usedBytes} shape the agent already reports, so the frontend
renders both identically):
- LXC: the host can read straight into the container's root
filesystem, no agent needed -- status/current's disk/maxdisk fields
are real usage, not just allocation.
- QEMU: the hypervisor can't see inside a virtual disk at all without
help, so this calls the QEMU guest agent's get-fsinfo command (same
"gracefully degrade if the agent's missing/older" tolerance already
used for its IP-address lookup, and independent of it -- one
command failing doesn't take out the other). Pseudo-filesystems
(tmpfs, etc.) are filtered out by checking for a non-empty backing
`disk` array, the common convention for this endpoint.
Extracted the disks-table JSX (previously only in the agent branch)
into a shared DisksTable component and used it in both branches, and
added a note explaining an empty result when a running QEMU VM's
guest agent doesn't support get-fsinfo (an older agent version).
Verified against a mock Proxmox API over real TLS: an LXC's root
usage, a QEMU VM's real fsinfo mounts (with the disk-less tmpfs entry
correctly filtered), and a QEMU VM whose get-fsinfo fails outright --
confirming that degrades to an empty disks list without throwing and
without affecting the separate network-get-interfaces result.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
08a984719f |
Let admins hide the Proxmox-link card per server
Not every registered server is a Proxmox VM/LXC -- bare-metal boxes and other hosts had no reason to show a "link to Proxmox" option, but it appeared unconditionally on every server's detail page. New hideProxmoxLink column on servers (default false, so existing behavior is unchanged until someone opts in). A "Not a VM? Hide this" link in the card's header sets it; once hidden, a small "+ Show Proxmox link options" link takes its place so it's still reachable, not buried in a settings form. The card always shows regardless of this flag once a server IS actually linked, so unlinking never becomes unreachable by hiding the card out from under an active link. Verified against a temp SQLite DB with real migrations: a new server defaults to false, and toggling true/false both persist correctly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
f2a6253ddf |
Show container image-update status from Dockhand
Dockhand already tracks per-container image updates internally (its UI shows this) via a cached check plus an on-demand recheck -- surface that here instead of only showing running/stopped state. adapter.listContainers() now also reads GET /api/containers/pending- updates per environment (a cached read, no registry hit) and merges each container's hasImageUpdate/newerVersion/checkedAt onto it. updateAvailable is a tri-state: true (update pending), false (checked, up to date), or null (never checked) -- distinguishing "no update" from "we don't know yet" matters since a container can sit unchecked indefinitely until someone triggers a check. New adapter.checkForUpdates() triggers a fresh check across every environment (POST /api/containers/check-updates, one registry lookup per container so this can take a while) and new POST /:id/dockhand/ check-updates route (operator+, matching the existing container action's role gating). Docker page: new "Update" column (badge + tooltip with the newer version), an "updates available" count next to the running/total count, and a "Check for updates" button. Dashboard's Dockhand widget also gained an "Updates" mini-stat, swapped in for the less useful "Not running" figure (already inferable from running/total). Field names (hasImageUpdate, newerVersion, checkedAt, the check- updates response shape) came from Dockhand's own published OpenAPI spec, not guessed -- and verified against a local mock Dockhand server covering all three update states (pending, up to date, never checked) plus the check-updates aggregation across environments, since this sandbox can't reach the user's real Dockhand instance. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a5ec099c13 |
Bring Sloth Manager's DNS/Secrets dashboard here, alongside integrations
Sloth Manager's dashboard showed per-system stats for DNS and Secrets (domains/records by type, expiry counts by type) that this app's Dashboard never had -- it only ever showed the six live integration widgets. Port those two sections over, generalized to sit alongside the integrations this app added that Sloth Manager never had. New server/src/dns/stats.ts (getDnsStats(), used by new GET /api/dns/stats): zone counts are fetched live per enabled provider (matching what the DNS page itself shows, since the zone cache table only gets a row once a zone has been synced and would undercount) -- one unreachable provider surfaces its error without blanking the rest. Record-type breakdown comes from the local cache instead, since records are only ever shown from cache elsewhere in this app too (fetching every zone's records live on every dashboard load would be far more expensive for no real accuracy gain). Dashboard.tsx gained three sections under clear headers: DNS (domain/ provider/record stat tiles + a "records by type" breakdown), Secrets (monitored/expiring/expired tiles + a "secrets by type" breakdown, computed client-side from the already-fetched secrets list, same as Sloth Manager did), and the existing Integrations widgets grouped under their own header for visual parity with the two new sections. Breakdowns render as a stacked proportion bar with a color-coded legend rather than a pie chart -- this app has no charting library anywhere yet, and a plain CSS progress bar (an idiom Tabler itself uses) gets the same "see the mix at a glance" value without adding one for a single dashboard. Verified getDnsStats() against a temp SQLite DB with real migrations and an enabled + a disabled DNS provider: enabled/total provider counts, live zone counting, and cached-record-type aggregation (sorted by count) all came back correct, and the disabled provider was correctly excluded from the zone count while still counting toward the total. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
655862c94e |
Add per-server admin-page links (Dockge, Webmin, Cockpit, etc.)
Each server's detail page gets an "Admin Links" card for bookmarking that host's own web UIs -- container managers, Webmin, Cockpit, or anything else reachable by URL -- so there's a quick way to jump there without hunting down the address each time. New server_links table (serverId FK, label, url, ON DELETE CASCADE so removing a server cleans up its links automatically) and three new routes: POST/PATCH/DELETE /api/servers/:id/links, gated to operator+ like the rest of this page's editing actions; the existing GET /:id/detail now includes the server's links alongside hardware/DNS info. URLs are validated to start with http:// or https:// server-side (rejecting e.g. a javascript: URL that would otherwise render as a clickable link). Rendered as a row of pill buttons (label opens the URL in a new tab), with inline edit/remove controls next to each when the viewer can edit, and an "Add link" form matching the existing task-form style on the same page. Verified against a temp SQLite DB with real migrations: create/list/ update/delete all work, and deleting the parent server cascades to remove its links rather than leaving them orphaned. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
e35da87886 |
Add a Diagnostic Log, ported and generalized from Sloth Manager
Sloth Manager tracked every API call made to DNS providers for connectivity troubleshooting. Port that here, generalized to cover every outbound integration this app makes, not just DNS -- Tailscale, Proxmox, Synology, Semaphore, Gitea, and Dockhand calls now show up too, since a broken API token or unreachable host on any of them is just as worth diagnosing. New services/diagLog.ts: a generic withDiagLogging(source, adapter) wraps every async method of any adapter object with timing + success/failure recording, without touching a single adapter's request/error-handling internals -- every DNS and integration adapter interface here is already just a flat set of async methods, so this one wrapper works for all twelve of them. Applied it at each adapter factory's own return statement (one line each) rather than at the route layer, so background jobs that construct adapters directly (the Tailscale key-expiry scheduler, IPAM sync, agent-driven Proxmox lookups) get logged too, not just requests through routes/integrations.ts. New diag_log table (ring-buffered to the last 500 rows, mirroring Sloth Manager's approach -- this is for live troubleshooting, not a durable record) and admin-only GET/DELETE /api/diag-log routes, source/ result filters, pagination. New admin-only Diagnostic Log page: filterable, paginated table with a Clear button. Also introduces the shared useSortable hook + SortableTh component used here for the first time -- a follow-up commit applies the same sorting (and CSV export) to the rest of the app's tables, per the same request. Verified end-to-end against a temp SQLite DB with real migrations: a fake wrapped adapter's successful and failing calls both land correctly in the log with the right source/operation/latency/error, and the source/ok filters and clear-log operation all behave correctly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
5919929dfb |
Fix "No online Proxmox nodes found" masking real host-stats errors
listNodeStats() was silently dropping any node whose /status or /storage call failed and filtering it out of the result — with every node dropped, the page showed the same "no nodes" empty state as a genuinely nodeless tailnet, even though the node was online and the existing VM/LXC table below it worked fine. The likely real cause: Sys.Audit (host status) and Datastore.Audit (storage) are different ACL privileges from the VM/LXC management scope this integration originally needed, so a token created before this feature existed may lack them. getNodeStats() now fetches /status and /storage independently via Promise.allSettled instead of Promise.all, so one endpoint failing doesn't discard data the other successfully returned, and records a per-endpoint error message. listNodeStats() no longer filters failed nodes out at all -- it always returns one entry per online node, with `error` set and the rest of the fields null when nothing could be fetched. The Proxmox page now shows that error inline on the node's card instead of it vanishing. Verified against a mock Proxmox API returning 403 on /storage only (partial data still shows) and on both endpoints (node stays visible with its error surfaced, not silently dropped) — reproducing the reported bug and confirming the fix. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
2d062a9ac4 |
Add host stats to the Proxmox page
The Proxmox page only showed guest (VM/LXC) status; there was no view
of the underlying host(s) themselves -- uptime, CPU/RAM usage, or how
full each storage pool is.
New adapter.listNodeStats() calls GET /nodes/{node}/status and GET
/nodes/{node}/storage for every online node (a node going down
shouldn't blank the whole page, same tolerance as listGuests()) and
returns uptime, CPU usage % + core count + load average, RAM/swap
usage, and per-storage (local, LVM-thin, ZFS, NFS, ...) used/total.
New GET /:id/proxmox/nodes route alongside the existing /guests one.
One card per online node now sits above the VM/LXC table, showing
these stats plus a small storage table with usage bars (reusing the
same green/yellow/red thresholds as the Servers detail page and the
Synology page). Handles a multi-node cluster by rendering one card
per node, and storages with no active/total data (e.g. an offline NFS
mount) render as an "Inactive" badge with blank usage instead of
throwing.
Verified end-to-end against a local mock Proxmox API server over real
TLS (a throwaway self-signed cert, matching how the adapter's
`insecure` option is meant to be used) since this sandbox can't reach
the user's actual Proxmox cluster -- confirmed CPU fraction-to-percent
conversion, load average parsing, byte fields, and that a
null-valued/inactive storage doesn't break parsing.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
1ff59afb40 |
Notify on expiring Tailscale device keys
The Tailscale page already showed per-device key expiry; extend the existing daily-reminder infrastructure (currently only for Secrets) to push it out through the configured notification channels too, the same way expiring secrets already are. New tailscaleKeyExpiryScheduler.ts mirrors secretExpiryScheduler.ts: runs once at startup (skipped if already run today) and daily thereafter, checking every enabled Tailscale integration's devices for keys expiring within the warning window and calling notify() with the results. Reuses the exact same daily time/timezone setting as the secret-expiry check (one "Daily reminder time" control, two independent on/off toggles) rather than adding a second schedule for users to configure. Centralized the expiring-soon threshold and check (previously only duplicated in the /synology and /tailscale route summaries) into adapter.ts as `KEY_EXPIRY_WARN_DAYS` / `isKeyExpiringSoon()`, and updated the devices route to use it instead of its own inline copy. New `tailscaleKeyCheck` notification-event toggle (default on) in settings, alongside the existing secret-expiry one. Verified end-to-end against a temp SQLite DB + real migrations: a tailscale integration pointed at a mock Tailscale API (one device expiring in 10 days, one with key-expiry disabled) with the webhook channel enabled and pointed at a mock receiver — confirmed the scheduler's startup check queries the DB correctly, decrypts the integration's credential, calls the adapter, filters out the disabled-expiry device, and delivers a webhook payload naming only the expiring device. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
c0af546d08 |
Show Tailscale device key expiry
Tailscale node keys expire (180 days by default unless disabled per
device) and an expired key drops the device off the tailnet until
re-authenticated -- worth surfacing before it happens.
adapter.listDevices() now reads `expires` and `keyExpiryDisabled` from
the device list API (`?fields=all`, already being fetched). Go's zero
time ("0001-01-01T00:00:00Z") is what Tailscale returns for "no real
expiry set" and is treated as null rather than shown as a bogus 1AD
date.
New "Key expiry" column on the Tailscale page: "Never" when disabled,
otherwise the date plus a badge (green/yellow/red matching the
Secrets module's ok/expiring/expired convention) using the same
30-day warning window as that module's default. The devices summary
gained `expiringSoon` (<=30 days left, including already-expired),
surfaced in the page header and as a new warning line on the
Dashboard's Tailscale widget, alongside the existing "awaiting
authorization" one.
Verified the parsing (real expiry, disabled/zero-time, already-expired,
far-future) against the compiled adapter with fetch calls to
api.tailscale.com redirected to a local mock server, since this
sandbox can't reach the user's real tailnet.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
9f2d27e2c4 |
Add system/hardware info to the Synology page
Beyond volume/disk health, DSM exposes hostname, CPU, RAM, IP addresses, model/serial/firmware, and uptime -- surface those too as a "System" card on the Synology page. New adapter.getSystemInfo() combines three DSM Web API calls (fields confirmed against the actively-maintained mib1185/py-synologydsm-api client, which documents the same SYNO.Core.System / SYNO.Core.System. Utilization / SYNO.DSM.Network endpoints this adapter already uses the discover-then-call pattern for): - SYNO.Core.System "info" -- model, serial, firmware_ver, cpu_cores, cpu_clock_speed, up_time (a raw uptime-command-style string, not a duration -- parsed client-side into "5d 15h 2m" with a fallback to the raw string if DSM ever returns an unexpected format). - SYNO.Core.System.Utilization "get" -- live CPU load % (user+system+ other) and real memory usage, in KB. - SYNO.DSM.Network "list" -- hostname and interface IP addresses (loopback filtered out). New GET /:id/synology/system route alongside the existing /storage one, read-only like the rest of this integration. Verified end-to-end against a local mock DSM server standing in for the real API (login/session flow, all three endpoint shapes, IP filtering, byte math) since this sandbox can't reach the user's LAN; also unit-checked the client-side uptime regex against DSM's three uptime-command formats (with/without days, minutes-only) plus its unmatched-format fallback. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
e7ac6fea3b |
Add a "Sync from Proxmox" action to IP Addresses (IPAM)
Extends the Tailscale IPAM sync to Proxmox: pulls every VM/LXC's IP address(es) across all enabled Proxmox integrations into the inventory, reusing the same never-overwrite-a-manual-entry semantics. Proxmox guests can have multiple NICs (each IP synced as its own entry) and, for QEMU VMs, IP discovery depends on a responsive guest agent — a VM with none simply contributes zero entries rather than erroring, matching getGuestDetail()'s existing best-effort behavior. listGuests() doesn't include IPs, so this fetches getGuestDetail() per guest; acceptable for an explicit, user-triggered sync rather than a background poll. Extracted the shared "insert, update only if I own this row, else skip and report" upsert logic (previously inline in the Tailscale sync handler) into upsertSyncedEntry(), now used by both sync routes so the core safety rule can't drift between them. Verified end-to-end against a mock Proxmox server covering: an LXC with two NICs (both synced), a QEMU VM with a responsive guest agent, a QEMU VM with no agent (zero entries, no error), a manual-entry collision (skipped and reported, never overwritten), and idempotent re-sync (add -> update on the second run). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
d714a87754 |
Add a "Sync from Tailscale" action to IP Addresses (IPAM)
IPAM was entirely manual — Tailscale device IPs never showed up there even though the Tailscale integration already lists them. Add an explicit sync action (matching this app's existing pattern of user-triggered syncs rather than silent background polling). - New source column on ipam_entries (null = manual, "tailscale" = auto-synced) so a re-sync only ever touches rows it created itself — a manually-entered IP that happens to collide with a tailnet address is left untouched and reported back as skipped, never overwritten. - POST /api/ipam/sync-tailscale pulls every enabled Tailscale integration's device list, upserting by primary IP (label, OS in notes, vendor "Tailscale"); one unreachable Tailscale integration doesn't block others. - New "Sync from Tailscale" button on the IP Addresses page, with a small "synced" badge marking which rows came from it. Also fixed a longstanding TODO found in the same file: the "DNS records" column always showed "-" because matchingDnsRecords was hardcoded to an empty array from before the DNS module existed. It now does the same content-based reverse lookup against the DNS module's record cache used elsewhere in the app. Verified end-to-end by running the real server with Tailscale's fetch call intercepted at the process level (its adapter hardcodes api.tailscale.com with no configurable URL, so it can't be pointed at a mock server the way Proxmox/Synology can): confirmed add, the manual-entry skip/never-overwrite behavior, idempotent re-sync (add -> update), and the DNS-matching fix, all against the real route and adapter code. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
f92f8de96e |
Add a Date & Time setting and fix inconsistent date formats app-wide
Different tables used different date formats: Servers & Tasks used a fixed "YYYY-MM-DD HH:mm:ss" (24h), while Audit Log, DNS, Integrations (Tailscale/Semaphore), and Users called plain .toLocaleString() with no options, which renders using the browser's own locale — different per browser/OS, and inconsistent with the other pages' fixed format. - New Settings -> Date & Time page: pick date order (YYYY-MM-DD, DD/MM/YYYY, MM/DD/YYYY) and 12h vs 24h clock, with a live preview. - utils/date.ts's formatDateTime() now reads these settings instead of being hardcoded to sv-SE/24h; the setting is fetched once at app startup (alongside /api/me) via a new non-secret GET /api/settings/display (any signed-in user, same rationale as the badge-color endpoints) and applied immediately on save too, without needing a page reload. - Switched every remaining raw new Date(...).toLocaleString() call (Audit Log, DNS zone sync time, Tailscale/Semaphore last-seen, Users' last login) over to the shared formatter, so every table now renders dates identically. Verified the formatter's date-order x 12h logic against all six combinations plus the midnight/noon 12h edge cases, and the settings endpoints end-to-end (defaults, partial updates, validation rejection, audit logging) against the real dev server. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ad6783b6a1 |
Extend badge color settings to cover integration types too
Settings → DNS Badges only let you customize DNS provider badge colors. Expand it into a general Badges page with a second section for the six integration types (Tailscale, Proxmox, Synology, Semaphore, Gitea, Dockhand), applied to the type badge in the Manage integrations table. - New integrationColors key on AppSettings, stored/merged the same way as providerColors via the existing settingsStore. - New GET /api/settings/integration-colors — non-secret, any signed-in user, mirroring /provider-colors — so the badge color can be read without needing admin access to the full settings payload. - Renamed DnsBadgeSettings.tsx -> BadgeSettings.tsx (route /settings/dns-badges -> /settings/badges, sub-nav label "DNS Badges" -> "Badges") with both color sections saved together. Verified end-to-end against the real dev server: PUT persists integration colors independently of provider colors, the public integration-colors endpoint reflects updates immediately, and the change is audit-logged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a6ab69316d |
Add a graceful Shutdown action alongside Proxmox's Start/Restart/Stop
Stop maps to Proxmox's hard power-off (/status/stop) — fine for a
crashed guest, but risky for anything with a filesystem that'd rather
flush cleanly first. Proxmox exposes a separate /status/shutdown
endpoint that asks the guest to power itself down (ACPI event for a
VM, SIGTERM-then-wait for a container), so add it as its own action
rather than overloading Stop.
- New shutdownGuest() on the Proxmox adapter, calling /status/shutdown.
- The existing generic action route/loop already dispatches by
${action}Guest, so adding "shutdown" to that list was enough on the
server side — no new route needed.
- New Shutdown button next to Restart/Stop on both the Integrations
page's guest table and a Proxmox-linked server's detail page.
Stop's confirm prompt now explicitly points at Shutdown as the
gentler alternative.
Verified end-to-end against a mock Proxmox server: the shutdown call
hits /status/shutdown (never /status/stop) and is audit-logged as
shutdown_guest.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
ae41a02864 |
Fix agent silently failing to report on hosts without a cpuinfo model name
report-tasks.sh's new hardware-collection code ran under set -euo pipefail, so a single failing command inside it aborted the whole script before the report was ever sent — with no error message, since nothing in that path had explicit error handling. On real hardware this hit immediately: `grep -m1 "model name" /proc/cpuinfo` exits 1 when there's no match, and many ARM boards (e.g. Raspberry Pi) have no such line at all. Confirmed via journalctl showing the systemd service failing every 15 minutes with exit 1 and zero output, and via a minimal repro of the exact bash control flow. Wrap the hardware/network collection call so any failure inside it is non-fatal: task reporting (the actual core function) must never be taken down by a quirk in the best-effort hardware-gathering code, on this host or any other. Also fixed a related gap found while testing the fallback path: the server only accepted the system field being absent, not explicitly null (what the script now sends if collection fails outright), which would have turned graceful degradation into a rejected report. Verified end-to-end against the real dev server: system:null, system omitted, and a normal populated report all now return 202 and persist correctly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
f5d3c25c89 |
Add editing existing integrations, so an expired secret can be rotated
Manage integrations only supported add/toggle/delete — fixing an expired API token meant deleting and recreating the whole integration. - New GET /api/integrations/:id/config returns only the non-secret config fields (never the decrypted secret) so an edit form can pre-fill URL/tailnet/etc. fields. - New POST /api/integrations/:id/test merges the stored, decrypted config with any freshly-typed overrides and pings the real adapter — lets "Test connection" work during an edit without ever sending the current secret back to the browser. - New IntegrationEditForm component: secret fields render blank with a "leave blank to keep the current value" placeholder; submitting only sends the fields that were actually filled in, so a name/URL edit can't accidentally wipe a secret and a secret rotation can't touch anything else. Reuses the existing PATCH /:id route, which already merged partial config updates correctly. Verified end-to-end against the real dev server: confirmed via direct DB decryption that a non-secret-only edit leaves the stored secret byte-for-byte unchanged, and that a secret-only edit rotates it without touching other config; the test route was confirmed to make a real network call (got a genuine "API token invalid" from Tailscale's API against a fake key). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b9409d3095 |
Add server hardware/network detail view with Proxmox sync and agent reporting
Servers & Tasks only tracked scheduled tasks — there was no overview of the servers themselves and no way to see CPU/RAM/disk/IP info. Add a clickable server overview (visible to every role, not just admins) that opens a per-server detail page. - servers table gains an agent-reported hardware/network snapshot (IPs, CPU model/cores/load, memory, disks) and an optional link to a Proxmox VM/LXC (integration + node + guest type + vmid). - Proxmox adapter gains getGuestDetail(): live cores/memory/disk from /config, live cpu/mem/uptime from /status/current, and IPs (LXC net config directly, QEMU via a best-effort guest-agent call that degrades gracefully when the agent isn't installed). - New GET /api/servers/:id/detail combines whichever hardware source applies (live Proxmox vs. last agent report) with a DNS reverse-lookup against the DNS module's own record cache, so matching hostnames show up next to each IP. New PATCH /api/servers/:id manages the Proxmox link. - agent/linux/report-tasks.sh now also collects and reports IPs, CPU/memory/disk info on every check-in (load-average-based CPU number, not instantaneous, to keep the agent a cheap oneshot). - Extracted the per-server task table into a shared ServerTaskTable component so the all-servers view and the new detail page render tasks identically. Verified server-side end-to-end against the real dev server (agent report -> detail endpoint -> DNS match) and the Proxmox adapter against a mock HTTPS server covering LXC/QEMU config parsing and the guest-agent-unavailable fallback. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1f0e6a9d4c |
Split Settings into sub-pages and add DNS cache clearing
Settings was a single long page. Break it into a sub-nav (Notifications, DNS Badges, Cache) with its own route per section, each loading and saving independently. Also add the DNS record-cache clearing action that Sloth Manager had — a new admin-only POST /api/dns/cache/clear truncates the zone/record cache tables so stale data can be wiped and zones re-synced from scratch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
3255314402 |
Build the Settings module: notification channels, event toggles, DNS badge colors
The Settings page was a "coming soon" placeholder. Port Sloth Manager's settings feature set: Gotify/ntfy/SMTP/webhook notification channels (each with its own test-send button), per-event toggles (DNS record added/updated/deleted, a daily secret-expiry digest with configurable time/timezone), and per-provider DNS badge color customization. Settings persist in the existing `settings` key/value table via a new settingsStore service; a notify service fans a message out to every enabled channel. DNS record add/update/delete now fire notifications, and a node-schedule job re-arms itself whenever the notification settings change. Removed the now-superseded GOTIFY_URL/GOTIFY_TOKEN env vars in favor of in-app configuration. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
da34be08d5 |
Fix Tailscale online/exit-node detection against real API data
Found while verifying against the user's real tailnet (25 devices): the Tailscale device object has no "online" or "isExitNode" field at all — this was inherited from Sloth Manager's original adapter, which apparently never had this checked against live data either. With the old mapping, every device showed online:null and isExitNode was always false regardless of reality. Fixed by deriving both from fields that actually exist: - online <- connectedToControl (whether the device currently has an active session with Tailscale's control plane) - isExitNode <- enabledRoutes containing both 0.0.0.0/0 and ::/0 (a device can *advertise* exit-node routes without them being approved; checking enabledRoutes instead of advertisedRoutes reflects whether it's actually acting as one right now) Both fields come back from the list endpoint via `?fields=all`, so this adds no extra requests. Verified against the real tailnet: 25 devices, 24 online / 1 offline (a phone last seen 9 days ago — correct), and the one device actually configured as an exit node identified correctly. This is the last of the five previously-unverified integrations now confirmed against real infrastructure; combined with the earlier Synology (HTTP vs HTTPS) and Proxmox (async task timing) findings, every integration has now had at least one real bug shaken out by testing against the user's actual homelab rather than mocks alone. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
4ed1ac8cad |
Fix Synology adapter to support plain HTTP, not just HTTPS
Found while verifying against the user's real DSM: the adapter hardcoded node:https for every request, but the user's NAS is reached over plain HTTP on port 5000 inside the LAN (HTTPS on 5001 also works, but nothing about the integration should assume one or the other). Now picks http vs https based on the configured URL's own protocol, with the right default port per protocol (5000/5001) and TLS options only applied for https. Verified against the user's real Synology (10.200.5.35): ping, login, and SYNO.Storage.CGI.Storage load_info all confirmed working end-to-end through the actual HTTP route layer — 1 volume (normal, SHR, ~11.5TB/~3.9TB used), 4 disks (all normal, 34-39°C). This is the first of the five previously- unverified integrations confirmed against real hardware. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a83a4b3b11 |
Add Synology integration — completes all six planned live integrations
Sixth and final live integration: volume and disk health across the NAS, read-only per the delivery plan (DSM write actions are riskier and stayed out of scope). Adapter built against the same discover-then-call pattern used by hacf-fr/synologydsm-api (the library behind Home Assistant's Synology integration) since DSM's API paths/versions vary by release and the user's NAS isn't reachable from here to check directly: - GET query.cgi?api=SYNO.API.Info&method=query&query=all once per adapter instance, to learn the real path/version for SYNO.API.Auth and SYNO.Storage.CGI.Storage rather than hardcoding them - SYNO.API.Auth login (account/passwd/format=sid) for a session id, re-used across calls and refreshed on session-related error codes (105/106/119) - SYNO.Storage.CGI.Storage's `load_info` method, which returns disks, volumes, and pools in one call — verified against that library's storage.py field mapping (size.total/used, status, smart_status, temp, exceed_bad_sector_thr, below_remain_life_thr) Uses node:https directly (like Proxmox and the cPanel DNS adapter) for an "allow self-signed certificate" option, since DSM ships one by default. No 2FA support — login surfaces a clear error if the account requires it rather than failing silently. Verified: full build passes. Unreachable from here, so ran an 11-check HTTP test against the live server: confirms all six integration types are now registered, role gating, credential (password) non-leakage, disabled- integration blocking, and the wrong_type crash-safety check — server stays up throughout. Real volume/disk data still needs verification once this app can reach the user's NAS. This completes the integrations phase from the original plan: Tailscale, Gitea, Dockhand, Semaphore, Proxmox, Synology are all built, each following the same config-in-UI + encrypted-credentials pattern. Combined with the foundation, Secrets, IPAM, DNS, and Servers & Tasks modules from earlier, Homelab Manager now covers every feature area from the user's original request. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
bc72b9e152 |
Add Proxmox integration
Fifth live integration: VM/LXC status across every online node in the
cluster, with start/stop/restart actions — matching the "dashboard + basic
actions" depth from the plan. Adapter built against Proxmox VE's own
published API tree (pve.proxmox.com/pve-docs/api-viewer/apidoc.js — parsed
its ~4MB ExtJS tree structure directly since it's not plain JSON) plus
community-verified docs for the token auth header format, since the user's
instance isn't reachable from here.
server/src/integrations/proxmox/adapter.ts: GET /nodes for online nodes,
then GET /nodes/{node}/{qemu,lxc} per node in parallel (Promise.all) and
flattened, tagging each guest with its node and type; POST
/nodes/{node}/{qemu,lxc}/{vmid}/status/{start,stop,reboot} for actions (all
confirmed token-auth-eligible via the spec's "allowtoken" flag). Uses
node:https directly (like the cPanel DNS adapter) rather than fetch, since
Proxmox commonly runs a self-signed certificate in homelab setups — added
an "Allow self-signed certificate" checkbox field for that. Auth is
`Authorization: PVEAPIToken=<tokenId>=<tokenSecret>`, a different shape
from the other three integrations' single bearer token, so the field
schema splits it into two fields (a non-secret token ID like
"root@pam!homelab-manager" and a secret token value).
Verified: full build passes. Unreachable from here, so ran a 16-check HTTP
test against the live server: role gating, credential non-leakage,
disabled-integration blocking, invalid-guest-type and non-numeric-vmid
rejection, checkbox-only config correctly still failing required-field
validation, and the wrong_type crash-safety check for this fifth adapter
type — server stays up throughout. Real node/VM data and the actual
start/stop/restart actions still need verification once this app can reach
the user's Proxmox cluster.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
abb5335a52 |
Add Semaphore integration
Fourth live integration: Ansible run history per template with a
trigger-a-run action — matching the "dashboard + basic actions" depth from
the plan. Adapter built against Semaphore's official published OpenAPI spec
(github.com/semaphoreui/semaphore/blob/develop/api-docs.yml) plus its Go
source directly for the task-status enum, which the spec itself leaves
untyped (db/Task.go, pkg/task_logger/task_logger.go) — instance wasn't
reachable from here (semaphore.int.toolabs.se doesn't resolve outside the
user's LAN), so verified against the real spec/source rather than guessing.
server/src/integrations/semaphore/adapter.ts: GET /api/projects, then GET
/api/project/{id}/templates?sort=name&order=asc per project — each template
in that response already embeds its last_task, so listing every template's
current status across every project needs only one call per project (no
N+1 per-template lookup, unlike Gitea where the run history isn't embedded
in the repo list). POST /api/project/{id}/tasks {template_id} triggers a
run. Follows the same config-in-UI + encrypted-credential pattern as the
other three integrations.
Verified: full build passes. Same as Dockhand — unreachable, so ran a
14-check HTTP test against the live server (real target URL, bogus token):
role gating, credential non-leakage, disabled-integration blocking, and the
wrong_type crash-safety check for this fourth adapter type, confirming the
server stays up throughout. Real template/run data and the trigger-a-run
action still need verification once this app can reach the user's LAN.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
9c45a806c8 |
Add Dockhand integration
Third live integration: container status across every Docker host Dockhand manages, with start/stop/restart actions — matching the "dashboard + basic actions" depth from the plan. Talks to Dockhand's own aggregating REST API (bearer tokens, dh_...) rather than the raw Docker Engine API on each host directly, so one credential covers every host Dockhand is already connected to. Adapter built against Dockhand's real published OpenAPI spec (fetched from https://github.com/strausmann/mcp-dockhand/blob/main/docs/dockhand-openapi.json, 257 documented paths) since the instance itself isn't reachable from here — same "verify against the real shape, don't guess" approach as Gitea, just via the spec instead of a live instance: - GET /api/environments — the Docker hosts Dockhand knows about - GET /api/containers?env=<id>&all=true — containers per host ({id, name, image, state, status}) - POST /api/containers/{id}/{start,stop,restart}?env=<id> server/src/integrations/dockhand/adapter.ts fans the environments call out to one containers call per host (Promise.all) and flattens the result, tagging each container with its environment; one unreachable host returns an empty list for that host rather than failing the whole dashboard view. Follows the same config-in-UI + encrypted-credential pattern as Tailscale and Gitea. Verified: full build passes. Since Dockhand isn't reachable from here, ran a 14-check HTTP test against the live server instead of the real API — role gating, credential non-leakage, disabled-integration blocking, and (critically) re-confirmed the wrong_type crash-safety pattern holds for this third adapter type: hitting the Dockhand routes on a differently-typed integration row returns a clean 400 rather than crashing the process, and the server keeps responding to /health afterward. Real container data and the start/stop/restart actions still need verification once this app runs on the user's LAN where Dockhand is reachable. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
79710aa7a5 |
Add Gitea integration; fix .env never being loaded outside Docker
Second live integration: repo list with last CI run status, and re-running
failed jobs on a workflow run — matching the "dashboard + basic actions"
depth from the plan. Adapter built directly against the real Gitea 1.27
swagger spec (fetched from the user's own instance) rather than guessing at
the API shape: GET /user/repos for the repo list, GET
/repos/{owner}/{repo}/actions/runs?limit=1 for the latest run per repo (only
for repos with Actions enabled), and POST .../rerun-failed-jobs for retrying
just the failed jobs in a run. Follows the same config-in-UI +
encrypted-credential pattern as Tailscale and DNS providers.
Since Gitea collects its own base URL as a config field (unlike Tailscale,
which always talks to a fixed api.tailscale.com), generalized the
"integrations.baseUrl" bookkeeping into resolveBaseUrl() instead of the
one-fixed-URL-per-type map used previously.
Also fixed a real gap found while setting this up: server/src/env.ts reads
process.env directly, but nothing in the app ever loaded .env into
process.env for plain `node dist/index.js` / `tsx src/index.ts` runs — only
Docker's `env_file` config populated it, by injecting vars before Node even
starts. Every local (non-Docker) run silently had every setting at its
insecure default. Added server/src/loadEnv.ts (dotenv, pointed at the
repo-root .env) as the first import in both server/src/index.ts and
server/src/db/migrate.ts's standalone entrypoint.
Verified against the user's real, reachable services — not mocks:
- Authentik (auth.labsconnect.se): full OIDC login completed by the user
through the real UI; confirmed their account landed as admin (first user).
- Gitea (gitea.labsconnect.se): the compiled adapter run directly against a
real API token correctly listed all 11 real repos; a full HTTP-layer test
against the live server (8 checks) additionally covered a real
test-connection ping, credential non-leakage in list responses, and role
gating (403) on the rerun-failed-jobs action even with a valid token
behind it. None of the real repos have any workflow run history yet, so
the success/failure status badge and the rerun action itself are
implemented per the swagger spec but not yet exercised against a real run
— worth checking once one of those repos has actual CI activity.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
9c2742c30f |
Add Tailscale integration and harden all routes against crash-on-throw
First of the six planned live integrations (Proxmox, Synology, Semaphore, Tailscale, Gitea, Dockhand), reusing Sloth Manager's existing Tailscale adapter logic. Generalizes the "integrations" table already scaffolded in the foundation pass into a working config-in-UI + encrypted-credentials flow, following the same pattern as the DNS providers module — an "Add integration" form only offers types with an implemented adapter (currently just Tailscale), so the framework is ready for the next five integrations without further schema/plumbing changes. - server/src/integrations/tailscale/adapter.ts: ported from Sloth Manager's backend/src/adapters/tailscale.js — listDevices/setAuthorized/deleteDevice against the Tailscale API, now config-based (tailnet + apiKey) instead of reading process.env, and returning a ping() result instead of throwing. - server/src/routes/integrations.ts: generic integration CRUD (admin) + a "test connection" endpoint, plus Tailscale-specific device routes (dashboard-and-basic-actions depth per the plan: authorize/deauthorize/ remove, gated to operator+, audit-logged). - web: an Integrations page (provider-style manage/browse split, matching the DNS page's UX) with a device table, and a live Tailscale widget on the Dashboard. Bug found and fixed while testing: hitting the Tailscale device routes on a non-Tailscale integration row crashed the ENTIRE server process, not just that request — the generic adapter registry throws for unimplemented types, and that throw happened inside an async handler with no surrounding try/catch, which Express 4 doesn't catch, so it became an unhandled rejection that (on modern Node) kills the process. Fixed at the source (check the row's type before ever constructing an adapter) and, since the same "a helper throws before any local try/catch runs" shape existed wherever a route calls into loadDnsProviderConfig/loadIntegrationConfig (both call decryptSecret, which throws if CREDENTIALS_ENCRYPTION_KEY is ever wrong/missing after data was already encrypted with a different key), added a small asyncHandler() wrapper and applied it to every route handler across every router — a single bad request should never be able to take the whole app down for every user. Verified: full build passes. Fresh HTTP-layer tests against a running server (17 checks) cover not-implemented-type rejection, missing-field validation, a real network call to api.tailscale.com with a bogus key (clean ok:false, not a crash), role gating at every tier, credential non-leakage, disabled-integration blocking, and the wrong_type case that originally crashed the server — confirmed it now returns 400 cleanly and the server stays up. Re-ran the existing DNS (13 checks) and Servers/Tasks suites afterward to confirm the asyncHandler sweep didn't regress anything — all passing. Authorizing/removing a real device still needs a real Tailscale API key to verify end-to-end. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
9712d611a6 |
Add Servers & Tasks module ported from Schedule Task Manager
Ports cron/systemd task tracking across Debian/Raspbian servers, including the Linux push agent (install/report/uninstall scripts, rebranded from "schedule-task-manager-agent" to "homelab-manager-agent") and the manual-task-entry flow for things an agent can't see (Docker jobs, backups). The schema (servers/scheduled_tasks tables) was already in place from the foundation pass, so this is mostly a straight port of the original's services/routes. Deliberate change from the original: server/token management (which mints agent credentials) is now admin-only rather than open to any logged-in user, and manual task CRUD is gated to operator+ — consistent with how Secrets, IPAM, and DNS already split "configure credentials" from "everyday edits" across roles. All mutations are audit-logged. - server/src/services/tokens.ts, taskSync.ts: ported near-verbatim (agent token hashing, the agent-sync-marks-missing-as-stale-not-deleted logic). - server/src/routes/servers.ts, tasks.ts, agentReport.ts: same contract as the original (agent auth is a per-server bearer token, independent of the session-based requireAuth used everywhere else). - web: a single Servers & Tasks page (filter bar, task table grouped by server/schedule type, manual task form, and an admin-only server management panel with token reveal + copyable install/uninstall commands), replacing the original's two separate pages/apps. Verified: full build passes; a scripted HTTP test against a running server covers unauthenticated access, role gating at each tier (admin-only server mgmt, operator+ task mgmt), agent bearer-token auth (valid/invalid/rotated), manual-vs-agent task edit protection, stale-marking on re-sync, and cascade delete — 22/22 checks passing. Real agent installation on an actual Debian/Raspbian host still needs to be tried on the user's network. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
39b1d1fe2e |
Add DNS module ported from Sloth Manager
Ports zone/record management across Cloudflare, Loopia, Pi-hole, Azure DNS, cPanel, and Technitium onto the new stack. Unlike the original (one instance per provider configured via env vars), providers are now configured through the UI and support multiple named instances per type, with API credentials encrypted at rest via the integration_credentials table. - server/src/dns/adapters/*: each provider ported to a config-based factory (no more process.env reads), preserving each provider's original quirks (Pi-hole session auth, Azure record-set merging, Loopia XML-RPC, cPanel UAPI/API2 fallbacks, Technitium composite record IDs). - server/src/routes/dns.ts: provider CRUD (admin), a "test connection" endpoint, and zone/record browsing+sync+CRUD (operator+), all audit-logged. - dns_zones_cache/dns_records_cache tables replace Sloth Manager's dns-cache.json file, keeping the same "cache is the source of truth for display, sync fetches fresh from the provider" behavior. - web: a DNS page with provider management, zone browsing, and a record editor, plus a reusable dynamic provider-config form. Verified: full build (tsc + vite) passes; a scripted HTTP-layer test against a running server exercises auth, role gating (403 for viewer), validation (400 on missing config fields), provider CRUD, credential non-leakage in list responses, and adapter error propagation (502 against an unreachable host) — all passing. Real provider connectivity still needs to be checked against the user's actual DNS accounts. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
6bd2ed52c1 |
Scaffold Homelab Manager foundation
Monorepo (Express+TS+Drizzle/libSQL server, React+Vite+Tabler web) matching the stack used by ScheduleTaskManager and Sloth Manager. Includes Authentik OIDC login with local admin/operator/viewer roles (first user becomes admin), a generalized audit log, encrypted-at-rest storage for future integration API tokens, the DB schema for all planned modules, and the Tabler-styled app shell/nav. Also ports the Secrets (expiry tracker) and IP Addresses (IPAM) modules from Sloth Manager onto the new stack. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |