7e81306aa7e50f41f31e0142d8bf6b2038ff0b38
14
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
73d649377e |
Add notification quiet hours with digest delivery
Instant notifications (DNS changes, integration failure/recovery alerts) had no way to avoid pinging overnight. Adds a "Quiet hours" window under Settings -> Notifications: notifications that would fire during the window are held in a new notification_queue table instead of sent immediately, then delivered as one combined digest at the end time (in the same timezone already used for the daily checks) via a new scheduled flush job. The scheduled daily checks (secret expiry, Tailscale key, Docker updates, Proxmox backups) already only fire once at a chosen time, so this mainly matters for the instant ones. Includes a live "N queued" indicator with a manual "Flush now" button for visibility, and correctly falls back to sending immediately whenever the feature is disabled (the default). Gating lives at the single choke point every notification already flows through (notify()), so no per-event-type wiring was needed. Verified against a fake webhook receiver on an isolated scratch database: exhaustively checked the midnight-wraparound window math (9 cases including exact-boundary inclusive/exclusive edges) against synthetic "now" values rather than depending on when the test happens to run, then end-to-end through the real notify()/ flushQuietHoursQueue() functions with a window constructed around the actual current time — confirmed a notification during the window queues instead of sending, the flush produces one digest with the original title/message intact and clears the queue, a notification outside the window sends immediately, and disabling the feature entirely sends immediately regardless of the window. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
99db7e1cf0 |
Surface Proxmox backup job status, with a daily failure notification
Proxmox already runs vzdump backups, but nothing in the app said
whether they were actually succeeding — a silent backup failure is
one of the more dangerous blind spots a homelab admin can have. Adds
a "Backups" card to the Proxmox page: configured backup job
schedules (storage target, which guests, enabled/disabled) from
GET /cluster/backup, and recent vzdump task history per node from
GET /nodes/{node}/tasks?typefilter=vzdump, with a banner at the top
if the most recent run didn't succeed.
New "Proxmox backup failed" notification toggle under Settings ->
Notifications, on the same daily schedule as the other checks. The
scheduler checks each node's own most-recent vzdump run independently
(not just the single most recent task overall) so one node's healthy
backup can't mask another node's failing one in a multi-node cluster.
Known limitation, documented in the adapter's own header comment:
Proxmox's task list doesn't reliably expose which specific guest
failed within an "all guests" job — only the task's own log text has
that — so this surfaces job- and task-level status rather than
guessing at per-guest outcomes.
Verified against a fake Proxmox server (real self-signed HTTPS, since
the adapter's node:https usage can't be monkey-patched under ESM)
reproducing the documented /cluster/backup and task-list response
shapes: job parsing (all-guests+exclude vs specific-vmids+disabled)
correct, task OK/failure parsing correct, and the critical multi-node
scenario confirmed — one node's failing latest run flagged, the
other's healthy latest run correctly left alone, with exactly one
notification of the right content. This reproduces Proxmox's
documented API shape rather than a live-verified one; flag if the
real cluster's response differs in some way this didn't anticipate.
Confirmed the real dev database's mtime was untouched throughout.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
6d673db9ec |
Add session management: see who's signed in, revoke a session
No visibility existed into who was currently signed in or a way to force a device out. Sessions already live as files via session-file-store, so this reads that store directly rather than adding a new DB table: new Settings-adjacent "Sessions" page (admin-only, alongside Users) lists every live session with the user's name/email/resolved role, IP, a friendly "Browser on OS" summary parsed from the user-agent, last-active time, and expiry, with a Revoke button per row (extra confirmation if you revoke your own current session, since that signs you out immediately). IP and user-agent are now captured into the session at login (auth/router.ts) since express-session doesn't track them itself. session-file-store's own Store type doesn't declare its list() method, so sessionStore.ts adds a narrow local interface for it rather than losing type safety on the rest of the store. Verified against a real session directory seeded through the actual session-file-store APIs (not hand-written JSON): confirmed correct field resolution including a session whose user row was later deleted (role resolves to null instead of crashing), correctly excluded a mid-OIDC-login session with no completed user yet, correctly excluded an already-expired session, and confirmed revoke actually deletes the right session file and only that one. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1cae35a59e |
Add a daily Docker image-update notification
Dockhand's pending-update counts were only visible if you happened to open the Docker page. New "Docker image update available" toggle under Settings -> Notifications, sharing the same daily time/timezone as the secret and Tailscale key expiry reminders (same node-schedule reschedule-on-settings-change pattern as those two). Reads each enabled Dockhand integration's already-cached update-check results via listContainers() rather than triggering a fresh per-container registry lookup, so it costs nothing extra beyond what the Docker page itself already fetches, and lists every container with an update pending across all environments/integrations in one notification. Verified end-to-end against a fake local Dockhand server (one environment, two containers, one flagged with a pending update) and a fake webhook receiver on an isolated scratch database: the check correctly found only the flagged container (with its newerVersion), sent exactly one notification with the right title/content, excluded the up-to-date container, and sent nothing at all when the setting was toggled off despite still finding the same pending update. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
009bb3e027 |
Add a global search / command palette (Ctrl/Cmd+K)
With 6 integrations, DNS, secrets, IPAM, and servers all in one app, there was no single place to type a hostname/IP/name and jump straight to it. New GET /api/search aggregates a LIKE-based search across servers, secrets, IPAM, integrations, DNS providers, DNS zones, and DNS records (joining zone/provider names onto each record result) in one round trip — homelab-scale row counts make a naive LIKE scan plenty fast, no FTS needed. Every underlying resource's own list endpoint already only requires requireAuth (viewer role included), so the aggregate endpoint uses the same single check. Frontend: a self-contained CommandPalette component (Tabler's modal CSS classes driven by React state, since the app doesn't load Bootstrap's JS) opens via a sidebar search button or Ctrl/Cmd+K from anywhere, debounces input, and supports arrow-key navigation. Results link to the right page: servers to their existing /servers/:id detail route; DNS zone/record results deep-link via new ?providerId= &zoneId= query-param handling added to Dns.tsx (auto-selects that provider/zone on load, since Dns.tsx previously held selection only in local state with no URL sync); secrets/IPAM results link with ?q= to prefill each page's existing client-side search box; integrations link to their type's dashboard page (no per-instance route exists yet, so same-type integrations share one link). Verified the query logic (joins, case-insensitive LIKE, correct zone/provider name resolution) against an isolated scratch database seeded with realistic cross-referencing rows — a server name match, a case-mismatched match against both a secret and a DNS provider sharing "cloudflare", an IP address matching both an IPAM entry and the DNS A record pointing at it (confirming the record's joined zone and provider names came through correctly), an integration name match, and a no-match query returning every category empty. Did not re-verify the requireAuth/asyncHandler wiring itself, since it's the same one-line pattern already proven across every other router in this app. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
10d123b18a |
Add automatic retention purging for the Diagnostic and Audit logs
The diagnostic log already rings-buffer to 500 rows, but the audit log had no cap at all and would grow forever. Adds an opt-in age-based purge under Settings -> Logs: keep entries for N days, checked on a configurable interval (hourly through monthly), plus a manual "Purge now" button. Reuses the existing node-schedule-style reschedule-on-settings-change pattern from the secret/Tailscale expiry checkers, but as a plain setInterval since "how often" here is an interval rather than a specific daily time. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
e35da87886 |
Add a Diagnostic Log, ported and generalized from Sloth Manager
Sloth Manager tracked every API call made to DNS providers for connectivity troubleshooting. Port that here, generalized to cover every outbound integration this app makes, not just DNS -- Tailscale, Proxmox, Synology, Semaphore, Gitea, and Dockhand calls now show up too, since a broken API token or unreachable host on any of them is just as worth diagnosing. New services/diagLog.ts: a generic withDiagLogging(source, adapter) wraps every async method of any adapter object with timing + success/failure recording, without touching a single adapter's request/error-handling internals -- every DNS and integration adapter interface here is already just a flat set of async methods, so this one wrapper works for all twelve of them. Applied it at each adapter factory's own return statement (one line each) rather than at the route layer, so background jobs that construct adapters directly (the Tailscale key-expiry scheduler, IPAM sync, agent-driven Proxmox lookups) get logged too, not just requests through routes/integrations.ts. New diag_log table (ring-buffered to the last 500 rows, mirroring Sloth Manager's approach -- this is for live troubleshooting, not a durable record) and admin-only GET/DELETE /api/diag-log routes, source/ result filters, pagination. New admin-only Diagnostic Log page: filterable, paginated table with a Clear button. Also introduces the shared useSortable hook + SortableTh component used here for the first time -- a follow-up commit applies the same sorting (and CSV export) to the rest of the app's tables, per the same request. Verified end-to-end against a temp SQLite DB with real migrations: a fake wrapped adapter's successful and failing calls both land correctly in the log with the right source/operation/latency/error, and the source/ok filters and clear-log operation all behave correctly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1ff59afb40 |
Notify on expiring Tailscale device keys
The Tailscale page already showed per-device key expiry; extend the existing daily-reminder infrastructure (currently only for Secrets) to push it out through the configured notification channels too, the same way expiring secrets already are. New tailscaleKeyExpiryScheduler.ts mirrors secretExpiryScheduler.ts: runs once at startup (skipped if already run today) and daily thereafter, checking every enabled Tailscale integration's devices for keys expiring within the warning window and calling notify() with the results. Reuses the exact same daily time/timezone setting as the secret-expiry check (one "Daily reminder time" control, two independent on/off toggles) rather than adding a second schedule for users to configure. Centralized the expiring-soon threshold and check (previously only duplicated in the /synology and /tailscale route summaries) into adapter.ts as `KEY_EXPIRY_WARN_DAYS` / `isKeyExpiringSoon()`, and updated the devices route to use it instead of its own inline copy. New `tailscaleKeyCheck` notification-event toggle (default on) in settings, alongside the existing secret-expiry one. Verified end-to-end against a temp SQLite DB + real migrations: a tailscale integration pointed at a mock Tailscale API (one device expiring in 10 days, one with key-expiry disabled) with the webhook channel enabled and pointed at a mock receiver — confirmed the scheduler's startup check queries the DB correctly, decrypts the integration's credential, calls the adapter, filters out the disabled-expiry device, and delivers a webhook payload naming only the expiring device. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
3255314402 |
Build the Settings module: notification channels, event toggles, DNS badge colors
The Settings page was a "coming soon" placeholder. Port Sloth Manager's settings feature set: Gotify/ntfy/SMTP/webhook notification channels (each with its own test-send button), per-event toggles (DNS record added/updated/deleted, a daily secret-expiry digest with configurable time/timezone), and per-provider DNS badge color customization. Settings persist in the existing `settings` key/value table via a new settingsStore service; a notify service fans a message out to every enabled channel. DNS record add/update/delete now fire notifications, and a node-schedule job re-arms itself whenever the notification settings change. Removed the now-superseded GOTIFY_URL/GOTIFY_TOKEN env vars in favor of in-app configuration. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
79710aa7a5 |
Add Gitea integration; fix .env never being loaded outside Docker
Second live integration: repo list with last CI run status, and re-running
failed jobs on a workflow run — matching the "dashboard + basic actions"
depth from the plan. Adapter built directly against the real Gitea 1.27
swagger spec (fetched from the user's own instance) rather than guessing at
the API shape: GET /user/repos for the repo list, GET
/repos/{owner}/{repo}/actions/runs?limit=1 for the latest run per repo (only
for repos with Actions enabled), and POST .../rerun-failed-jobs for retrying
just the failed jobs in a run. Follows the same config-in-UI +
encrypted-credential pattern as Tailscale and DNS providers.
Since Gitea collects its own base URL as a config field (unlike Tailscale,
which always talks to a fixed api.tailscale.com), generalized the
"integrations.baseUrl" bookkeeping into resolveBaseUrl() instead of the
one-fixed-URL-per-type map used previously.
Also fixed a real gap found while setting this up: server/src/env.ts reads
process.env directly, but nothing in the app ever loaded .env into
process.env for plain `node dist/index.js` / `tsx src/index.ts` runs — only
Docker's `env_file` config populated it, by injecting vars before Node even
starts. Every local (non-Docker) run silently had every setting at its
insecure default. Added server/src/loadEnv.ts (dotenv, pointed at the
repo-root .env) as the first import in both server/src/index.ts and
server/src/db/migrate.ts's standalone entrypoint.
Verified against the user's real, reachable services — not mocks:
- Authentik (auth.labsconnect.se): full OIDC login completed by the user
through the real UI; confirmed their account landed as admin (first user).
- Gitea (gitea.labsconnect.se): the compiled adapter run directly against a
real API token correctly listed all 11 real repos; a full HTTP-layer test
against the live server (8 checks) additionally covered a real
test-connection ping, credential non-leakage in list responses, and role
gating (403) on the rerun-failed-jobs action even with a valid token
behind it. None of the real repos have any workflow run history yet, so
the success/failure status badge and the rerun action itself are
implemented per the swagger spec but not yet exercised against a real run
— worth checking once one of those repos has actual CI activity.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
9c2742c30f |
Add Tailscale integration and harden all routes against crash-on-throw
First of the six planned live integrations (Proxmox, Synology, Semaphore, Tailscale, Gitea, Dockhand), reusing Sloth Manager's existing Tailscale adapter logic. Generalizes the "integrations" table already scaffolded in the foundation pass into a working config-in-UI + encrypted-credentials flow, following the same pattern as the DNS providers module — an "Add integration" form only offers types with an implemented adapter (currently just Tailscale), so the framework is ready for the next five integrations without further schema/plumbing changes. - server/src/integrations/tailscale/adapter.ts: ported from Sloth Manager's backend/src/adapters/tailscale.js — listDevices/setAuthorized/deleteDevice against the Tailscale API, now config-based (tailnet + apiKey) instead of reading process.env, and returning a ping() result instead of throwing. - server/src/routes/integrations.ts: generic integration CRUD (admin) + a "test connection" endpoint, plus Tailscale-specific device routes (dashboard-and-basic-actions depth per the plan: authorize/deauthorize/ remove, gated to operator+, audit-logged). - web: an Integrations page (provider-style manage/browse split, matching the DNS page's UX) with a device table, and a live Tailscale widget on the Dashboard. Bug found and fixed while testing: hitting the Tailscale device routes on a non-Tailscale integration row crashed the ENTIRE server process, not just that request — the generic adapter registry throws for unimplemented types, and that throw happened inside an async handler with no surrounding try/catch, which Express 4 doesn't catch, so it became an unhandled rejection that (on modern Node) kills the process. Fixed at the source (check the row's type before ever constructing an adapter) and, since the same "a helper throws before any local try/catch runs" shape existed wherever a route calls into loadDnsProviderConfig/loadIntegrationConfig (both call decryptSecret, which throws if CREDENTIALS_ENCRYPTION_KEY is ever wrong/missing after data was already encrypted with a different key), added a small asyncHandler() wrapper and applied it to every route handler across every router — a single bad request should never be able to take the whole app down for every user. Verified: full build passes. Fresh HTTP-layer tests against a running server (17 checks) cover not-implemented-type rejection, missing-field validation, a real network call to api.tailscale.com with a bogus key (clean ok:false, not a crash), role gating at every tier, credential non-leakage, disabled-integration blocking, and the wrong_type case that originally crashed the server — confirmed it now returns 400 cleanly and the server stays up. Re-ran the existing DNS (13 checks) and Servers/Tasks suites afterward to confirm the asyncHandler sweep didn't regress anything — all passing. Authorizing/removing a real device still needs a real Tailscale API key to verify end-to-end. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
9712d611a6 |
Add Servers & Tasks module ported from Schedule Task Manager
Ports cron/systemd task tracking across Debian/Raspbian servers, including the Linux push agent (install/report/uninstall scripts, rebranded from "schedule-task-manager-agent" to "homelab-manager-agent") and the manual-task-entry flow for things an agent can't see (Docker jobs, backups). The schema (servers/scheduled_tasks tables) was already in place from the foundation pass, so this is mostly a straight port of the original's services/routes. Deliberate change from the original: server/token management (which mints agent credentials) is now admin-only rather than open to any logged-in user, and manual task CRUD is gated to operator+ — consistent with how Secrets, IPAM, and DNS already split "configure credentials" from "everyday edits" across roles. All mutations are audit-logged. - server/src/services/tokens.ts, taskSync.ts: ported near-verbatim (agent token hashing, the agent-sync-marks-missing-as-stale-not-deleted logic). - server/src/routes/servers.ts, tasks.ts, agentReport.ts: same contract as the original (agent auth is a per-server bearer token, independent of the session-based requireAuth used everywhere else). - web: a single Servers & Tasks page (filter bar, task table grouped by server/schedule type, manual task form, and an admin-only server management panel with token reveal + copyable install/uninstall commands), replacing the original's two separate pages/apps. Verified: full build passes; a scripted HTTP test against a running server covers unauthenticated access, role gating at each tier (admin-only server mgmt, operator+ task mgmt), agent bearer-token auth (valid/invalid/rotated), manual-vs-agent task edit protection, stale-marking on re-sync, and cascade delete — 22/22 checks passing. Real agent installation on an actual Debian/Raspbian host still needs to be tried on the user's network. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
39b1d1fe2e |
Add DNS module ported from Sloth Manager
Ports zone/record management across Cloudflare, Loopia, Pi-hole, Azure DNS, cPanel, and Technitium onto the new stack. Unlike the original (one instance per provider configured via env vars), providers are now configured through the UI and support multiple named instances per type, with API credentials encrypted at rest via the integration_credentials table. - server/src/dns/adapters/*: each provider ported to a config-based factory (no more process.env reads), preserving each provider's original quirks (Pi-hole session auth, Azure record-set merging, Loopia XML-RPC, cPanel UAPI/API2 fallbacks, Technitium composite record IDs). - server/src/routes/dns.ts: provider CRUD (admin), a "test connection" endpoint, and zone/record browsing+sync+CRUD (operator+), all audit-logged. - dns_zones_cache/dns_records_cache tables replace Sloth Manager's dns-cache.json file, keeping the same "cache is the source of truth for display, sync fetches fresh from the provider" behavior. - web: a DNS page with provider management, zone browsing, and a record editor, plus a reusable dynamic provider-config form. Verified: full build (tsc + vite) passes; a scripted HTTP-layer test against a running server exercises auth, role gating (403 for viewer), validation (400 on missing config fields), provider CRUD, credential non-leakage in list responses, and adapter error propagation (502 against an unreachable host) — all passing. Real provider connectivity still needs to be checked against the user's actual DNS accounts. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
6bd2ed52c1 |
Scaffold Homelab Manager foundation
Monorepo (Express+TS+Drizzle/libSQL server, React+Vite+Tabler web) matching the stack used by ScheduleTaskManager and Sloth Manager. Includes Authentik OIDC login with local admin/operator/viewer roles (first user becomes admin), a generalized audit log, encrypted-at-rest storage for future integration API tokens, the DB schema for all planned modules, and the Tabler-styled app shell/nav. Also ports the Secrets (expiry tracker) and IP Addresses (IPAM) modules from Sloth Manager onto the new stack. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |