Commit Graph
20 Commits
Author SHA1 Message Date
bobbanandClaude Sonnet 5 ae64cb345c Manage server tags in Settings: pre-add tags, recolour, rename, delete
New Settings > Tags tab (admin) listing every tag -- ones servers use and
ones added ahead of time -- with how many servers carry each:
- Add a tag before anything uses it, optionally with a colour. Such tags
  are offered as one-click "Add:" chips (and datalist suggestions) when
  tagging a server, so the same word gets spelled the same way everywhere.
- Give any tag a colour of your choosing, in use or not, or reset it to the
  automatic one. Changes are drafted with Save/Cancel rather than saved as
  the picker drags. Colours show on the Servers page, its tag filter bar,
  the detail page and the editor, and update everywhere without a reload
  through one shared cached colour map.
- Rename a tag; every server that has it is rewritten. Renaming to a name
  that already exists merges the two after a confirmation naming what will
  happen; the target keeps its own colour unless it had none, and a server
  carrying both ends up with one.
- Delete a tag, which removes it from every server that has it, with a
  confirmation stating how many.

Tags still live on the servers (servers.tags); a new tag_definitions table
holds only what a server can't: existence before use, and a colour. A
defined tag stays listed until an admin deletes it, even with no servers.
Rename and delete change the servers and the catalogue in one transaction
so they can't disagree. Only servers that actually carry the tag are
rewritten and counted -- an earlier draft also counted servers whose tags
merely weren't in sorted order, which the tests caught.

Reading the list and colours is open to everyone signed in (needed to draw
tags anywhere); changing the catalogue is admin-only, while tagging a
server stays an operator action. Names go through the same normalisation
as before, colours must be #rrggbb, and every change is audit-logged.

New table tag_definitions (migration 0013).

Verified with 44 backend checks (list/counts, roles, create/adopt/
duplicate/rejects, colour set/reset, rename incl. defined, undefined and
unused tags, merge colour rules and both-sides servers, delete, audit, and
that normal tagging still works afterwards) and in a browser against the
real routers: add with colour, set/reset a colour and see it change on the
Servers page live, merge with confirmation, delete, and the error path.
Not clicked through: the quick-add chips inside the tag editor on a
server's detail page (typechecked; same colour code as the rest).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 21:17:09 +02:00
bobbanandClaude Sonnet 5 72f8c85406 Add a Privacy page: what is stored, where it goes, and your own data
A page every signed-in user can open. It states what the installation
stores (accounts, sign-in sessions, audit and diagnostic logs, server
reports, the secrets tracker, credentials, inventory) and for how long,
where data goes (Authentik, integrations, DNS providers, notification
channels, domain registries, certificate checks, port scans), what lives
in the browser, who can see what, and how to limit or remove data.

Written from what the code actually does, including the uncomfortable
parts: the session file keeps the user's Authentik ID token plus the IP
and browser from sign-in; audit entries keep a name snapshot after an
account is gone; the app has no delete-account function; cron commands
in agent reports can contain sensitive text. It also says what isn't
there -- no telemetry, update checks, third-party scripts, fonts or
tracking cookies -- which was checked against the web build and the
server's outbound calls before being asserted.

Live values rather than boilerplate: log retention (and whether it's on),
which integration and DNS provider types are enabled, how many domains,
certificate checks and reporting servers, and which notification channels
are on. Channel addresses are shown to admins only, and only the host --
never a path, query string or token -- since a webhook URL can embed a key.

Each user also sees their own account and active sign-ins, and can
"Download my data": their account, their sign-ins and the audit-log entries
made under their account, as JSON. Only their own -- never another user's --
and without session ids or ID tokens. The export is itself audit-logged, so
a later export shows it.

Verified with 23 backend checks (own-vs-others isolation for audit counts,
sessions and export; no session ids, ID tokens or channel secrets in any
response; admin-vs-viewer channel visibility; live counts; audit of the
export; auth) using an isolated session directory so real sessions are
never read, and in a browser against the real router, including the
download. Real dev database and session files untouched.

Not legal text: this is a transparency page for the people using the app,
not a privacy policy or a GDPR compliance document.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 20:28:03 +02:00
bobbanandClaude Sonnet 5 2352689fd3 Add a consistency report across IPAM, DNS and servers
New Consistency page listing where the three places this app records
what lives at an address disagree:
- Address conflicts: the same address reported by more than one server.
- DNS out of date: a record named after a server (its hostname, or its
  short name) that points at an address the server doesn't report.
- IPAM out of date: an entry labelled with a server's name at an
  address the server doesn't report.
- Not in IPAM: addresses a server reports or DNS points at that IPAM
  doesn't list, merged into one finding per address, with a one-click
  "Add to IPAM" that pre-fills a label.
- No DNS record: server LAN addresses no cached A/AAAA record resolves to.

It compares data the app already holds and fetches nothing when opened,
so the page states how many servers had reported addresses and how many
DNS zones are synced (and how old the oldest sync is) -- DNS records are
only cached for zones that have been synced, and a report that silently
treated missing data as "no records" would mislead.

Rules chosen to keep it from crying wolf:
- Only private addresses are compared; public DNS records aren't expected
  to be in IPAM.
- Servers with no reported addresses are never judged.
- Agents report IPv4 only, so records are only compared within an address
  family (an AAAA record isn't "stale" for lacking an IPv6 address).
- Docker bridge networks (172.16/12) are ignored: shared ones aren't
  conflicts, and they aren't listed unless someone put them in DNS.
- Tailscale addresses don't need DNS records (MagicDNS), and IPAM entries
  kept current by the Tailscale/Proxmox syncs aren't second-guessed.

Findings anyone has decided are fine can be ignored (operators) with a
reason. An ignore is keyed on the finding's stable identity so it stays
ignored across runs, its stored text comes from the finding rather than
the request, and it is marked "no longer occurring" once the condition
goes away. Ignore/restore are audit-logged.

New table consistency_ignores (migration 0012). portScan's private-address
helper is now exported and shared.

Verified with 36 checks (each rule and its exclusions, address-family and
case/trailing-dot handling, IPv6 case, ordering, stable keys, the report
route including a garbled agent report, source counts, ignore/unignore
rules and audit entries) and by driving the page against the real routers
in a browser: Add to IPAM actually created the entry, ignore and restore,
severity filter, "show all", the viewer view, and narrow-width layout
(which found and fixed a squeezed badge and clipped buttons). Real dev
database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 04:08:51 +02:00
bobbanandClaude Sonnet 5 ca0fa817f8 Track domain registration expiry, with daily reminders
New Domains page listing when each domain registration expires, read from
the registry. Domains behind the DNS zones already synced are picked up
automatically; others can be added by hand. You're reminded daily from N
days before expiry (Settings > Notifications, default 30) until it's
renewed, and told when an expiry date hasn't been refreshable for several
days so a stale date isn't trusted silently.

RDAP alone would not have covered this homelab: .se, .nu, .io, .eu and .de
are not in IANA's RDAP bootstrap. Lookups therefore try RDAP where the TLD
publishes a server and fall back to WHOIS on port 43, found via IANA's
own referral, parsing the expiry line out of the free-text answer. Only
the expiry date and registrar are read or stored. Verified live against
the real registries: .se and .nu via WHOIS, .com/.org/.dev via RDAP.

Behaviour worth knowing:
- A DNS zone that is a subdomain (lab.example.se) resolves to the
  registration that actually expires by trying the name and then its
  parents, so no public-suffix list is needed. Zones already covered by a
  tracked domain are not looked up again.
- "Couldn't ask" is never confused with "not registered": network errors,
  rate limits and garbled answers are errors, and a transient error at any
  level stops the walk from concluding the domain doesn't exist.
- A failed refresh keeps the last known expiry and records why, rather
  than blanking a date that's still relied on.
- Zones that don't resolve to a real registration (.lan, .local, unregistered
  names) simply get no row. Zone-derived rows disappear when their zone
  does; manual rows stay. Zone-derived rows can't be deleted by hand.
- Registries that don't publish an expiry (.de, .eu) are tracked with a
  note instead of a date.
- Input like "example.com/path" is refused rather than silently reduced
  to its host.
- Runs on the daily secret-expiry schedule and reminder time, on demand
  (Check all now / per domain), and once at startup if nothing has been
  read in a day. Never blocks startup, one lookup at a time with a pause.
  The warning window lives with the other thresholds in settings.

New table domains (migration 0011); two settings fields (toggle and
warning days).

Verified with 90 checks against fake RDAP/WHOIS backends (name
normalization, date formats, WHOIS parsing including rate-limit and
no-expiry answers, bootstrap and referral caching, stale-cache fallback,
parent walking, add/sync/check/refresh, concurrency guard, alert
selection and stale detection, the daily notification and its toggle,
role rules) plus a live smoke test against real registries and a browser
check of the page against the real router. Real dev database mtime
untouched. Not checked: a screenshot of the finished page (the capture
timed out); structure, sorting, errors and the viewer view were verified.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 03:31:02 +02:00
bobbanandClaude Sonnet 5 aae4f0d74f Add maintenance mode to silence alerts while working on a server or integration
Rebooting Proxmox or patching a server triggered failure/offline alerts
you then had to dismiss. A maintenance window silences alerts about one
server, integration, or DNS provider for a chosen time. New Maintenance
page (start with a duration and optional reason, end early, see what's
silenced and what isn't) and a banner in the app shell so every signed-in
user can see what is currently silenced. Starting/ending is operator-only
and audit-logged; starting one on a target that already has a window
restarts its clock instead of stacking.

Silenced for the target: server offline/disk alerts, Proxmox/Synology
storage and health alerts, Proxmox backup alerts, and "integration down"
alerts. Not silenced: expiry and update reminders, DNS change notices.

The design goal is that this cannot hide a real outage:
- Every window has a required end (5 min to 7 days); there is no
  open-ended option, so a forgotten window expires by itself.
- A silenced problem is deliberately NOT recorded as "known". If it is
  still present when the window ends it alerts then, as new. A problem
  that was already alerted before the window stays known, so it isn't
  repeated, and is reported cleared only after the window ends.
- Failure alerts keep counting failures during a window without marking
  themselves alerted, so an outage that outlasts the window alerts on the
  very next failed call.

Known limitation, stated on the page: integration-failure alerts are
tracked per service TYPE (all "proxmox"), not per configured instance, so
a window on one Proxmox integration also silences a failure on a second
Proxmox integration while it's open. Fixing that means threading the
integration id through every adapter and the diagnostic log, which is a
much larger change than this feature.

Also moved the API-error-message helper out of Secrets.tsx into a shared
util now that two pages use it. New table maintenance_windows (migration
0008).

Verified with 44 checks: the condition-key-to-subject mapping (including
server:3 vs server:33), the diff rules with silenced subjects (new problem
not recorded, alerts when the window ends; already-known one carried and
not repeated; clears only after the window), window expiry and
integration/DNS-provider source matching, the failure tracker end to end
against a webhook (silent during a window while an unrelated service still
alerts; outage that outlasts the window alerts on the next failure and
only once; fail-and-recover fully inside a window sends nothing), a full
health pass against a real window, and the real router with a stubbed
session (role rules, duration bounds including the missing-duration case,
extend-not-stack, 404s, deleted targets hidden, audit entries). Real dev
database mtime untouched.

Not done: I haven't clicked through the new page or banner in a browser
(they sit behind the Authentik login); it builds and the API behind it is
tested.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 02:23:06 +02:00
bobbanandClaude Sonnet 5 1688de3ea2 Alert when a server goes silent, a disk fills up, or a Synology volume degrades
The data was all being collected (agent last-seen, per-disk usage,
Proxmox storage, Synology volume/disk health) but nothing acted on
it, so a dead server or a full disk was only noticed by opening the
right page.

A new health pass runs every 15 minutes (matching the agent's default
report interval) and raises one notification when a problem starts and
one when it clears: a server's agent silent past a threshold (default
60 min), a server disk / Proxmox storage or root filesystem / Synology
volume at or above a usage threshold (default 90%), and a Synology
volume or disk that isn't "normal", has bad SMART, bad sectors past the
threshold, or life remaining below it. Both thresholds and an on/off
toggle live under Settings -> Notifications.

The parts that make this trustworthy rather than noisy:
- A problem is keyed by identity, so it alerts once and not every run;
  a shared Proxmox storage listed by every node is one problem, not
  one per node.
- Active problems persist across restarts, so a rebuild doesn't
  re-alert everything already known.
- If a source can't be read on a given run (Proxmox/Synology
  unreachable, one node lacking privileges) its existing problems are
  held, not reported "cleared" and then re-alerted when it comes back —
  the integration-failure alert already owns "the integration is down".
- For 20 minutes after startup server-derived problems are held too:
  agents couldn't report while the app was down, so judging them then
  would report every server offline after any restart.
- An offline server's disk figures are stale and are not judged; a
  server that never reported has no agent and raises nothing.
- Tracking continues while the toggle is off (only sending is gated),
  so turning it back on doesn't dump every long-standing problem.

Timestamps without a zone (SQLite's format) are read as UTC; the test
runs on a UTC+2 machine, where reading them as local time gives a
different answer.

Verified with 32 checks: the evaluation rules and the state diff as
pure functions (exact thresholds, the proxmox:1 vs proxmox:10 prefix
trap, the flapping sequence), then a whole pass against a real
Proxmox adapter talking to a fake HTTPS cluster (one node returning
403, the whole API down, a shared storage on two nodes, a node that
recovers), a webhook receiver, the real DB, and the persisted state.

Not exercised end-to-end: the Synology collection path — its rules are
tested on data shaped exactly like the adapter's output types, but I
did not stand up a fake DSM. Real dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 02:15:19 +02:00
bobbanandClaude Sonnet 5 73d649377e Add notification quiet hours with digest delivery
Instant notifications (DNS changes, integration failure/recovery
alerts) had no way to avoid pinging overnight. Adds a "Quiet hours"
window under Settings -> Notifications: notifications that would fire
during the window are held in a new notification_queue table instead
of sent immediately, then delivered as one combined digest at the end
time (in the same timezone already used for the daily checks) via a
new scheduled flush job. The scheduled daily checks (secret expiry,
Tailscale key, Docker updates, Proxmox backups) already only fire
once at a chosen time, so this mainly matters for the instant ones.
Includes a live "N queued" indicator with a manual "Flush now" button
for visibility, and correctly falls back to sending immediately
whenever the feature is disabled (the default).

Gating lives at the single choke point every notification already
flows through (notify()), so no per-event-type wiring was needed.

Verified against a fake webhook receiver on an isolated scratch
database: exhaustively checked the midnight-wraparound window math
(9 cases including exact-boundary inclusive/exclusive edges) against
synthetic "now" values rather than depending on when the test
happens to run, then end-to-end through the real notify()/
flushQuietHoursQueue() functions with a window constructed around the
actual current time — confirmed a notification during the window
queues instead of sending, the flush produces one digest with the
original title/message intact and clears the queue, a notification
outside the window sends immediately, and disabling the feature
entirely sends immediately regardless of the window. Confirmed the
real dev database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 20:56:33 +02:00
bobbanandClaude Sonnet 5 99db7e1cf0 Surface Proxmox backup job status, with a daily failure notification
Proxmox already runs vzdump backups, but nothing in the app said
whether they were actually succeeding — a silent backup failure is
one of the more dangerous blind spots a homelab admin can have. Adds
a "Backups" card to the Proxmox page: configured backup job
schedules (storage target, which guests, enabled/disabled) from
GET /cluster/backup, and recent vzdump task history per node from
GET /nodes/{node}/tasks?typefilter=vzdump, with a banner at the top
if the most recent run didn't succeed.

New "Proxmox backup failed" notification toggle under Settings ->
Notifications, on the same daily schedule as the other checks. The
scheduler checks each node's own most-recent vzdump run independently
(not just the single most recent task overall) so one node's healthy
backup can't mask another node's failing one in a multi-node cluster.

Known limitation, documented in the adapter's own header comment:
Proxmox's task list doesn't reliably expose which specific guest
failed within an "all guests" job — only the task's own log text has
that — so this surfaces job- and task-level status rather than
guessing at per-guest outcomes.

Verified against a fake Proxmox server (real self-signed HTTPS, since
the adapter's node:https usage can't be monkey-patched under ESM)
reproducing the documented /cluster/backup and task-list response
shapes: job parsing (all-guests+exclude vs specific-vmids+disabled)
correct, task OK/failure parsing correct, and the critical multi-node
scenario confirmed — one node's failing latest run flagged, the
other's healthy latest run correctly left alone, with exactly one
notification of the right content. This reproduces Proxmox's
documented API shape rather than a live-verified one; flag if the
real cluster's response differs in some way this didn't anticipate.
Confirmed the real dev database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 18:42:22 +02:00
bobbanandClaude Sonnet 5 6d673db9ec Add session management: see who's signed in, revoke a session
No visibility existed into who was currently signed in or a way to
force a device out. Sessions already live as files via
session-file-store, so this reads that store directly rather than
adding a new DB table: new Settings-adjacent "Sessions" page
(admin-only, alongside Users) lists every live session with the
user's name/email/resolved role, IP, a friendly "Browser on OS"
summary parsed from the user-agent, last-active time, and expiry, with
a Revoke button per row (extra confirmation if you revoke your own
current session, since that signs you out immediately).

IP and user-agent are now captured into the session at login
(auth/router.ts) since express-session doesn't track them itself.
session-file-store's own Store type doesn't declare its list()
method, so sessionStore.ts adds a narrow local interface for it rather
than losing type safety on the rest of the store.

Verified against a real session directory seeded through the actual
session-file-store APIs (not hand-written JSON): confirmed correct
field resolution including a session whose user row was later deleted
(role resolves to null instead of crashing), correctly excluded a
mid-OIDC-login session with no completed user yet, correctly excluded
an already-expired session, and confirmed revoke actually deletes the
right session file and only that one. Confirmed the real dev
database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-21 20:29:29 +02:00
bobbanandClaude Sonnet 5 1cae35a59e Add a daily Docker image-update notification
Dockhand's pending-update counts were only visible if you happened to
open the Docker page. New "Docker image update available" toggle
under Settings -> Notifications, sharing the same daily
time/timezone as the secret and Tailscale key expiry reminders (same
node-schedule reschedule-on-settings-change pattern as those two).
Reads each enabled Dockhand integration's already-cached update-check
results via listContainers() rather than triggering a fresh
per-container registry lookup, so it costs nothing extra beyond what
the Docker page itself already fetches, and lists every container
with an update pending across all environments/integrations in one
notification.

Verified end-to-end against a fake local Dockhand server (one
environment, two containers, one flagged with a pending update) and a
fake webhook receiver on an isolated scratch database: the check
correctly found only the flagged container (with its newerVersion),
sent exactly one notification with the right title/content, excluded
the up-to-date container, and sent nothing at all when the setting
was toggled off despite still finding the same pending update.
Confirmed the real dev database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 23:44:47 +02:00
bobbanandClaude Sonnet 5 009bb3e027 Add a global search / command palette (Ctrl/Cmd+K)
With 6 integrations, DNS, secrets, IPAM, and servers all in one app,
there was no single place to type a hostname/IP/name and jump
straight to it. New GET /api/search aggregates a LIKE-based search
across servers, secrets, IPAM, integrations, DNS providers, DNS
zones, and DNS records (joining zone/provider names onto each record
result) in one round trip — homelab-scale row counts make a naive
LIKE scan plenty fast, no FTS needed. Every underlying resource's own
list endpoint already only requires requireAuth (viewer role
included), so the aggregate endpoint uses the same single check.

Frontend: a self-contained CommandPalette component (Tabler's modal
CSS classes driven by React state, since the app doesn't load
Bootstrap's JS) opens via a sidebar search button or Ctrl/Cmd+K from
anywhere, debounces input, and supports arrow-key navigation. Results
link to the right page: servers to their existing /servers/:id
detail route; DNS zone/record results deep-link via new ?providerId=
&zoneId= query-param handling added to Dns.tsx (auto-selects that
provider/zone on load, since Dns.tsx previously held selection only
in local state with no URL sync); secrets/IPAM results link with
?q= to prefill each page's existing client-side search box;
integrations link to their type's dashboard page (no per-instance
route exists yet, so same-type integrations share one link).

Verified the query logic (joins, case-insensitive LIKE, correct
zone/provider name resolution) against an isolated scratch database
seeded with realistic cross-referencing rows — a server name match, a
case-mismatched match against both a secret and a DNS provider
sharing "cloudflare", an IP address matching both an IPAM entry and
the DNS A record pointing at it (confirming the record's joined zone
and provider names came through correctly), an integration name
match, and a no-match query returning every category empty. Did not
re-verify the requireAuth/asyncHandler wiring itself, since it's the
same one-line pattern already proven across every other router in
this app. Confirmed the real dev database's mtime was untouched
throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 03:20:05 +02:00
bobbanandClaude Sonnet 5 10d123b18a Add automatic retention purging for the Diagnostic and Audit logs
The diagnostic log already rings-buffer to 500 rows, but the audit
log had no cap at all and would grow forever. Adds an opt-in
age-based purge under Settings -> Logs: keep entries for N days,
checked on a configurable interval (hourly through monthly), plus a
manual "Purge now" button. Reuses the existing node-schedule-style
reschedule-on-settings-change pattern from the secret/Tailscale
expiry checkers, but as a plain setInterval since "how often" here is
an interval rather than a specific daily time.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 02:22:32 +02:00
bobbanandClaude Sonnet 5 e35da87886 Add a Diagnostic Log, ported and generalized from Sloth Manager
Sloth Manager tracked every API call made to DNS providers for
connectivity troubleshooting. Port that here, generalized to cover
every outbound integration this app makes, not just DNS -- Tailscale,
Proxmox, Synology, Semaphore, Gitea, and Dockhand calls now show up
too, since a broken API token or unreachable host on any of them is
just as worth diagnosing.

New services/diagLog.ts: a generic withDiagLogging(source, adapter)
wraps every async method of any adapter object with timing +
success/failure recording, without touching a single adapter's
request/error-handling internals -- every DNS and integration adapter
interface here is already just a flat set of async methods, so this
one wrapper works for all twelve of them. Applied it at each adapter
factory's own return statement (one line each) rather than at the
route layer, so background jobs that construct adapters directly
(the Tailscale key-expiry scheduler, IPAM sync, agent-driven Proxmox
lookups) get logged too, not just requests through routes/integrations.ts.

New diag_log table (ring-buffered to the last 500 rows, mirroring
Sloth Manager's approach -- this is for live troubleshooting, not a
durable record) and admin-only GET/DELETE /api/diag-log routes, source/
result filters, pagination.

New admin-only Diagnostic Log page: filterable, paginated table with a
Clear button. Also introduces the shared useSortable hook + SortableTh
component used here for the first time -- a follow-up commit applies
the same sorting (and CSV export) to the rest of the app's tables, per
the same request.

Verified end-to-end against a temp SQLite DB with real migrations: a
fake wrapped adapter's successful and failing calls both land correctly
in the log with the right source/operation/latency/error, and the
source/ok filters and clear-log operation all behave correctly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 21:41:06 +02:00
bobbanandClaude Sonnet 5 1ff59afb40 Notify on expiring Tailscale device keys
The Tailscale page already showed per-device key expiry; extend the
existing daily-reminder infrastructure (currently only for Secrets)
to push it out through the configured notification channels too, the
same way expiring secrets already are.

New tailscaleKeyExpiryScheduler.ts mirrors secretExpiryScheduler.ts:
runs once at startup (skipped if already run today) and daily
thereafter, checking every enabled Tailscale integration's devices for
keys expiring within the warning window and calling notify() with the
results. Reuses the exact same daily time/timezone setting as the
secret-expiry check (one "Daily reminder time" control, two
independent on/off toggles) rather than adding a second schedule for
users to configure.

Centralized the expiring-soon threshold and check (previously only
duplicated in the /synology and /tailscale route summaries) into
adapter.ts as `KEY_EXPIRY_WARN_DAYS` / `isKeyExpiringSoon()`, and
updated the devices route to use it instead of its own inline copy.

New `tailscaleKeyCheck` notification-event toggle (default on) in
settings, alongside the existing secret-expiry one.

Verified end-to-end against a temp SQLite DB + real migrations: a
tailscale integration pointed at a mock Tailscale API (one device
expiring in 10 days, one with key-expiry disabled) with the webhook
channel enabled and pointed at a mock receiver — confirmed the
scheduler's startup check queries the DB correctly, decrypts the
integration's credential, calls the adapter, filters out the
disabled-expiry device, and delivers a webhook payload naming only
the expiring device.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 18:59:27 +02:00
bobbanandClaude Sonnet 5 3255314402 Build the Settings module: notification channels, event toggles, DNS badge colors
The Settings page was a "coming soon" placeholder. Port Sloth Manager's
settings feature set: Gotify/ntfy/SMTP/webhook notification channels
(each with its own test-send button), per-event toggles (DNS record
added/updated/deleted, a daily secret-expiry digest with configurable
time/timezone), and per-provider DNS badge color customization.

Settings persist in the existing `settings` key/value table via a new
settingsStore service; a notify service fans a message out to every
enabled channel. DNS record add/update/delete now fire notifications,
and a node-schedule job re-arms itself whenever the notification
settings change. Removed the now-superseded GOTIFY_URL/GOTIFY_TOKEN
env vars in favor of in-app configuration.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-15 12:14:05 +02:00
bobbanandClaude Sonnet 5 79710aa7a5 Add Gitea integration; fix .env never being loaded outside Docker
Second live integration: repo list with last CI run status, and re-running
failed jobs on a workflow run — matching the "dashboard + basic actions"
depth from the plan. Adapter built directly against the real Gitea 1.27
swagger spec (fetched from the user's own instance) rather than guessing at
the API shape: GET /user/repos for the repo list, GET
/repos/{owner}/{repo}/actions/runs?limit=1 for the latest run per repo (only
for repos with Actions enabled), and POST .../rerun-failed-jobs for retrying
just the failed jobs in a run. Follows the same config-in-UI +
encrypted-credential pattern as Tailscale and DNS providers.

Since Gitea collects its own base URL as a config field (unlike Tailscale,
which always talks to a fixed api.tailscale.com), generalized the
"integrations.baseUrl" bookkeeping into resolveBaseUrl() instead of the
one-fixed-URL-per-type map used previously.

Also fixed a real gap found while setting this up: server/src/env.ts reads
process.env directly, but nothing in the app ever loaded .env into
process.env for plain `node dist/index.js` / `tsx src/index.ts` runs — only
Docker's `env_file` config populated it, by injecting vars before Node even
starts. Every local (non-Docker) run silently had every setting at its
insecure default. Added server/src/loadEnv.ts (dotenv, pointed at the
repo-root .env) as the first import in both server/src/index.ts and
server/src/db/migrate.ts's standalone entrypoint.

Verified against the user's real, reachable services — not mocks:
- Authentik (auth.labsconnect.se): full OIDC login completed by the user
  through the real UI; confirmed their account landed as admin (first user).
- Gitea (gitea.labsconnect.se): the compiled adapter run directly against a
  real API token correctly listed all 11 real repos; a full HTTP-layer test
  against the live server (8 checks) additionally covered a real
  test-connection ping, credential non-leakage in list responses, and role
  gating (403) on the rerun-failed-jobs action even with a valid token
  behind it. None of the real repos have any workflow run history yet, so
  the success/failure status badge and the rerun action itself are
  implemented per the swagger spec but not yet exercised against a real run
  — worth checking once one of those repos has actual CI activity.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-14 23:55:41 +02:00
bobbanandClaude Sonnet 5 9c2742c30f Add Tailscale integration and harden all routes against crash-on-throw
First of the six planned live integrations (Proxmox, Synology, Semaphore,
Tailscale, Gitea, Dockhand), reusing Sloth Manager's existing Tailscale
adapter logic. Generalizes the "integrations" table already scaffolded in
the foundation pass into a working config-in-UI + encrypted-credentials
flow, following the same pattern as the DNS providers module — an
"Add integration" form only offers types with an implemented adapter
(currently just Tailscale), so the framework is ready for the next five
integrations without further schema/plumbing changes.

- server/src/integrations/tailscale/adapter.ts: ported from Sloth Manager's
  backend/src/adapters/tailscale.js — listDevices/setAuthorized/deleteDevice
  against the Tailscale API, now config-based (tailnet + apiKey) instead of
  reading process.env, and returning a ping() result instead of throwing.
- server/src/routes/integrations.ts: generic integration CRUD (admin) + a
  "test connection" endpoint, plus Tailscale-specific device routes
  (dashboard-and-basic-actions depth per the plan: authorize/deauthorize/
  remove, gated to operator+, audit-logged).
- web: an Integrations page (provider-style manage/browse split, matching
  the DNS page's UX) with a device table, and a live Tailscale widget on
  the Dashboard.

Bug found and fixed while testing: hitting the Tailscale device routes on
a non-Tailscale integration row crashed the ENTIRE server process, not just
that request — the generic adapter registry throws for unimplemented types,
and that throw happened inside an async handler with no surrounding
try/catch, which Express 4 doesn't catch, so it became an unhandled
rejection that (on modern Node) kills the process. Fixed at the source
(check the row's type before ever constructing an adapter) and, since the
same "a helper throws before any local try/catch runs" shape existed
wherever a route calls into loadDnsProviderConfig/loadIntegrationConfig
(both call decryptSecret, which throws if CREDENTIALS_ENCRYPTION_KEY is
ever wrong/missing after data was already encrypted with a different key),
added a small asyncHandler() wrapper and applied it to every route handler
across every router — a single bad request should never be able to take
the whole app down for every user.

Verified: full build passes. Fresh HTTP-layer tests against a running
server (17 checks) cover not-implemented-type rejection, missing-field
validation, a real network call to api.tailscale.com with a bogus key
(clean ok:false, not a crash), role gating at every tier, credential
non-leakage, disabled-integration blocking, and the wrong_type case that
originally crashed the server — confirmed it now returns 400 cleanly and
the server stays up. Re-ran the existing DNS (13 checks) and Servers/Tasks
suites afterward to confirm the asyncHandler sweep didn't regress anything
— all passing. Authorizing/removing a real device still needs a real
Tailscale API key to verify end-to-end.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-14 23:31:19 +02:00
bobbanandClaude Sonnet 5 9712d611a6 Add Servers & Tasks module ported from Schedule Task Manager
Ports cron/systemd task tracking across Debian/Raspbian servers, including
the Linux push agent (install/report/uninstall scripts, rebranded from
"schedule-task-manager-agent" to "homelab-manager-agent") and the
manual-task-entry flow for things an agent can't see (Docker jobs, backups).
The schema (servers/scheduled_tasks tables) was already in place from the
foundation pass, so this is mostly a straight port of the original's
services/routes.

Deliberate change from the original: server/token management (which mints
agent credentials) is now admin-only rather than open to any logged-in user,
and manual task CRUD is gated to operator+ — consistent with how Secrets,
IPAM, and DNS already split "configure credentials" from "everyday edits"
across roles. All mutations are audit-logged.

- server/src/services/tokens.ts, taskSync.ts: ported near-verbatim (agent
  token hashing, the agent-sync-marks-missing-as-stale-not-deleted logic).
- server/src/routes/servers.ts, tasks.ts, agentReport.ts: same contract as
  the original (agent auth is a per-server bearer token, independent of the
  session-based requireAuth used everywhere else).
- web: a single Servers & Tasks page (filter bar, task table grouped by
  server/schedule type, manual task form, and an admin-only server
  management panel with token reveal + copyable install/uninstall commands),
  replacing the original's two separate pages/apps.

Verified: full build passes; a scripted HTTP test against a running server
covers unauthenticated access, role gating at each tier (admin-only server
mgmt, operator+ task mgmt), agent bearer-token auth (valid/invalid/rotated),
manual-vs-agent task edit protection, stale-marking on re-sync, and cascade
delete — 22/22 checks passing. Real agent installation on an actual
Debian/Raspbian host still needs to be tried on the user's network.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-14 23:16:58 +02:00
bobbanandClaude Sonnet 5 39b1d1fe2e Add DNS module ported from Sloth Manager
Ports zone/record management across Cloudflare, Loopia, Pi-hole, Azure DNS,
cPanel, and Technitium onto the new stack. Unlike the original (one instance
per provider configured via env vars), providers are now configured through
the UI and support multiple named instances per type, with API
credentials encrypted at rest via the integration_credentials table.

- server/src/dns/adapters/*: each provider ported to a config-based factory
  (no more process.env reads), preserving each provider's original quirks
  (Pi-hole session auth, Azure record-set merging, Loopia XML-RPC, cPanel
  UAPI/API2 fallbacks, Technitium composite record IDs).
- server/src/routes/dns.ts: provider CRUD (admin), a "test connection"
  endpoint, and zone/record browsing+sync+CRUD (operator+), all audit-logged.
- dns_zones_cache/dns_records_cache tables replace Sloth Manager's
  dns-cache.json file, keeping the same "cache is the source of truth for
  display, sync fetches fresh from the provider" behavior.
- web: a DNS page with provider management, zone browsing, and a record
  editor, plus a reusable dynamic provider-config form.

Verified: full build (tsc + vite) passes; a scripted HTTP-layer test against
a running server exercises auth, role gating (403 for viewer), validation
(400 on missing config fields), provider CRUD, credential non-leakage in
list responses, and adapter error propagation (502 against an unreachable
host) — all passing. Real provider connectivity still needs to be checked
against the user's actual DNS accounts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-14 23:00:30 +02:00
bobbanandClaude Sonnet 5 6bd2ed52c1 Scaffold Homelab Manager foundation
Monorepo (Express+TS+Drizzle/libSQL server, React+Vite+Tabler web) matching
the stack used by ScheduleTaskManager and Sloth Manager. Includes Authentik
OIDC login with local admin/operator/viewer roles (first user becomes admin),
a generalized audit log, encrypted-at-rest storage for future integration API
tokens, the DB schema for all planned modules, and the Tabler-styled app
shell/nav. Also ports the Secrets (expiry tracker) and IP Addresses (IPAM)
modules from Sloth Manager onto the new stack.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-14 21:57:13 +02:00