Commit Graph
110 Commits
Author SHA1 Message Date
bobbanandClaude Sonnet 5 69e9325927 Document role-based menu access in ROLES.md
Adds a reference table of which sidebar items each role (viewer,
operator, admin) can see, plus a summary of which in-page actions on
otherwise-visible pages are held back to operator/admin.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-29 20:41:26 +02:00
bobbanandClaude Sonnet 5 df2a5ce42b Add a Proxmox Backup Server integration: datastore/snapshot verification status
Proxmox VE already shows whether the last vzdump push to PBS succeeded, but
has no visibility into PBS's own backup verification, GC/prune health, or
host status. This adds PBS as its own integration (own adapter, page, nav
entry, and Dashboard widget) that reads datastore usage and, for every
stored snapshot, its verification state directly from PBS.

A new daily check (mirroring the existing Proxmox backup-failure check)
notifies when a snapshot has failed verification or a datastore couldn't be
read, with its own toggle in Settings -> Notifications and its own
maintenance-window silencing.

Not verified against a live PBS instance — built from PBS's published API
docs and a scratch test against a mocked PBS server exercising the adapter's
parsing and auth-header format (PBSAPIToken uses a colon separator, unlike
PVE's PVEAPIToken which uses =). See INTEGRATIONS.md for details and the
"not verified" caveat.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-29 20:32:52 +02:00
bobbanandClaude Sonnet 5 70ba60c7da Add a phpIPAM import to the IP Addresses page
New "Sync from phpIPAM" action on IP Addresses, alongside the existing
"Sync from Tailscale" / "Sync from Proxmox" ones and built the same way:
a new integration type (config in-app, URL + API app ID + app token,
credentials encrypted at rest) that this page pulls from on demand.
Deliberately import-only, not a full integration -- no dedicated page,
dashboard widget, or nav entry, since that's all this was asked for.

Auth is phpIPAM's static "App token" method: create an API app under
Administration -> API with its security set to "SSL with App token", and
its one-time code goes straight in as the `token` header (also sent as
`phpipam-token`, in case a given version expects that name instead) --
no login call, no token to renew. The user/password "User token" method
isn't implemented.

Addresses are read the standard way: GET /subnets/, then GET
/subnets/{id}/addresses/ for each, rather than assuming a single
"all addresses" endpoint exists on every version. phpIPAM wraps every
response as {code, success, data} -- including an empty result: a subnet
with nothing in it answers success:false, message:"No addresses found"
rather than success:true, data:[]. That's read as "nothing here", not a
failure; anything else with success:false throws with phpIPAM's own
message. One subnet failing outright (e.g. the app lacks permission on
it) is skipped with a note rather than aborting the whole sync. Every
address field is read defensively -- optional, independently
type-checked -- so a field phpIPAM renames or drops in some version
leaves that value blank instead of breaking the import.

Imported entries: label from hostname or description, "phpIPAM" as
vendor, the subnet's own description (or its CIDR, if it has none) as
location, and description/note/MAC folded into notes. Existing sync
plumbing (upsertSyncedEntry) gained a location parameter so this and any
future sync can set it; the two existing syncs pass null, unchanged.

Endpoints, the token header, and the address/subnet field names are
cross-checked against phpIPAM's own published API documentation. Not
verified against a live instance -- there wasn't one available while
building this, so if a real sync comes back empty or with the wrong
fields, that's the next thing to check.

Verified with 23 backend checks against a fake phpIPAM server matching
that documented shape (the empty-subnet quirk, a subnet that fails
outright, malformed/missing fields, a non-JSON response) and the real
route (added/updated/skipped counts, a manually-entered IP never
overwritten, roles, no-enabled-integration, upstream failure surfaced
per-integration rather than as a 500, audit entries) plus a browser check
of the real IP Addresses page against the real routers: the sync button,
its result message, re-syncing (updates rather than duplicates), the
manual entry staying untouched, and the viewer view.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-29 19:49:10 +02:00
bobbanandClaude Sonnet 5 bf7f73f6b6 Import maintenance windows from Uptime Kuma
New "Import from Uptime Kuma" action on the Maintenance page (operator):
pick an Uptime Kuma integration and a duration, and it starts (or
extends) a maintenance window here for every server whose address
matches a monitor Uptime Kuma currently reports as being in maintenance,
reusing the existing monitor-to-server matching from the Uptime Kuma
integration itself.

Uptime Kuma's metrics endpoint only exposes a monitor's *current* status,
not its scheduled start/end time (there's no API for that), so this
deliberately doesn't try to mirror Uptime Kuma's own schedule -- it starts
a window for the duration you choose, the same bounded/required-end
window this feature has always used. Running it again while Kuma is
still in maintenance extends the same window rather than stacking a
second one; when Kuma later shows nothing in maintenance, already-active
windows are left alone rather than force-ended, since ending them isn't
something only Uptime Kuma's state should decide. Two Kuma monitors that
match the same server are deduped to one window. Monitors with no
matching server are reported back by name so nothing is silently missed,
and monitors that aren't in maintenance are ignored entirely.

The manual "Start maintenance" endpoint's start-or-extend logic (dedupe,
pruning old rows, the response shape) is now a shared
services/maintenance.ts function instead of living only in that route
handler, so the import path can't drift from how a manual window behaves.
Likewise the server-matching helper gained a small toMatchableServers()
so the existing Uptime Kuma monitors route and this new one build the
same match input the same way instead of each parsing server rows on
their own.

No schema change -- imported windows are ordinary maintenance windows;
their Uptime-Kuma origin is only in the reason text ("Imported from
Uptime Kuma (<integration>): <monitor>"), visible in the Active table and
the audit log like any other window.

Verified with 28 backend checks (the refactored manual start/extend
flow as a regression check; roles; unknown/wrong-type/disabled
integration; validation; nothing-in-maintenance; matching including
same-server dedup and unmatched monitors; re-running extends rather than
duplicating; windows left alone once Kuma exits maintenance; upstream
failure; audit entries) and by driving the real Maintenance page against
the real routers in a browser: import, re-import (extends), the
nothing-in-maintenance state, and the viewer view (no edit card, active
windows still visible). Real dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-29 19:11:35 +02:00
bobbanandClaude Sonnet 5 26de6cb243 Move Uptime Kuma from Infrastructure to Operations in the sidebar
It watches over things rather than being a machine/platform itself,
closer in spirit to Maintenance than to Proxmox or Tailscale. Also
balances the two groups (6/2 -> 5/3).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-29 18:52:41 +02:00
bobbanandClaude Sonnet 5 1a2dd19736 Add an Uptime Kuma integration: monitor status and which server each one watches
New integration, following the existing pattern: config in-app (URL +
API key, credentials encrypted at rest), its own Uptime Kuma page, an
Integrations list entry, a Dashboard widget, and diagnostic-log/
integration-down-alert coverage for free via the shared withDiagLogging
wrapper. Read-only -- no start/stop equivalent exists for a monitor.

Uptime Kuma has no conventional REST API (the dashboard talks to it over
Socket.IO); researched before writing any code, since guessing wrong here
would have cost real time. The one machine-readable, authenticated
endpoint that lists every monitor is its Prometheus exporter at
GET /metrics, gated by HTTP Basic auth -- an API key as the password with
the username left blank on current installs, or the real dashboard
login on installs from before the API-key feature existed. This adapter
authenticates the same way and parses that endpoint's text-exposition
format itself (metrics: monitor_status, monitor_response_time,
monitor_cert_days_remaining, monitor_uptime_ratio; labels: monitor_id,
monitor_name, monitor_type, monitor_url, monitor_hostname, monitor_port),
verified against the documented metric/label set and the actual upstream
source (server/prometheus.js). A malformed line is skipped rather than
failing the whole scrape.

"What server is being monitored for what": each monitor's target (an IP
for TCP checks, or the hostname out of the URL for HTTP/keyword checks)
is matched against your servers' own IPs and hostnames -- reusing the
same kind of match already used in the consistency report -- and linked
to that server's page. Monitors with no single network target (groups,
push monitors, DNS/keyword checks with a complex URL) are left unmatched
rather than guessed at. Uptime Kuma's tags aren't read, since the
Prometheus endpoint doesn't reliably distinguish a tag label from any
other label it might add later.

The username field is the first genuinely optional integration config
field this app has had; IntegrationField gained an `optional` flag
(server validation and both the add/edit web forms honor it) rather than
special-casing Uptime Kuma.

Verified with 48 backend checks (Prometheus text parsing including
escaped quotes, decimals, negative numbers, and malformed lines; TCP vs.
HTTP target/port extraction; every documented status code; server
matching by IP, hostname, and short name, including no-match cases; the
route's real HTTP round trip against a fake Uptime Kuma server, wrong
credentials, upstream failures, roles, wrong/disabled/missing
integration, diagnostic-log entries; the optional-field validation rule)
and by driving the real page and the real Dashboard widget in a browser
against the real routers, including CSV export and column sorting. Real
dev database mtime untouched.

Not verified: a real Uptime Kuma instance. Everything here was checked
against Uptime Kuma's documented metric format, its actual upstream
source, and a fake server built to match both -- not against a live
installation. If your instance's /metrics output differs from what's
documented (older version, unusual monitor types), the parser should
degrade to an empty or partial monitor list rather than error, but that
degradation itself hasn't been observed against the real thing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-29 18:50:27 +02:00
bobbanandClaude Sonnet 5 ada2e648e9 Group the sidebar into collapsible submenus
The flat 22-item menu becomes 7 entries: Dashboard and Secrets as plain
links, and five groups.
- Infrastructure: Servers, Proxmox, Synology, Docker, Tailscale
- Network: DNS, Domains, IP Addresses, Consistency
- Automation: Semaphore, Gitea
- Operations: Maintenance, Generator
- Administration: Integrations, Users, Sessions, Audit Log, Diagnostic
  Log, Settings
Privacy moves to the sidebar footer beside the theme toggle and sign-out,
since it's about the person rather than the homelab.

Behaviour:
- Groups open on demand. The group holding the current page is always
  open (deep routes count: /servers/7 keeps Servers active), and its
  header stays emphasised even if you close it. Which groups you leave
  open is remembered in the browser, and the menu still works if storage
  is blocked.
- Role handling: pages a role can't use are dropped, an emptied group
  disappears, and a group reduced to a single page is drawn as that
  page's own link. A viewer therefore sees a plain Integrations link
  where an admin sees Administration, and an operator sees Administration
  with Integrations and Audit Log.
- Toggling a group on mobile keeps the menu open; choosing a page closes
  it, as before.
- Group headers are real buttons with aria-expanded.

Front end only: one component, plus a small CSS rule tightening submenu
rows so several open groups still fit on one screen. Routes, permissions
and backend are unchanged.

Verified in a browser with the real shell as admin, operator and viewer:
structure per role, open/close and aria state, active highlighting on
direct and deep routes, auto-open of the current page's group, state
surviving a reload, and the mobile toggler behaviour. Not done: the
attention dots on closed groups (deliberately held for a later step).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-27 03:00:22 +02:00
bobbanandClaude Sonnet 5 9229ea2ed6 Add a Domains widget to the dashboard
A Domains card completing the Overview row (DNS, Secrets, Servers,
Domains), shaped like the Secrets card: a status badge (N expired / N
expiring / All OK / No expiry dates / Not configured), tracked, expiring
and expired counts, an OK/expiring/expired/no-date bar, and details --
the expired domains and the soonest-expiring ones (first three, soonest
first), the next one to expire when everything is fine, and a note when
some domains couldn't be refreshed, linking to the Domains page.

Purely front end: the existing /api/domains already carries each domain's
status, days left and last-check error, computed against the warning
window from Settings. A registry that doesn't publish expiry dates (.de,
.eu) shows as "no date" and is not counted as a failed refresh; anything
else that stopped a refresh is.

Verified in a browser against the real domains router across four states:
problems (expired, several expiring, an undated one, a failed refresh),
all healthy, only undated domains, and none tracked. Web build only; no
backend or database changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-27 02:52:05 +02:00
bobbanandClaude Sonnet 5 91796a3c3a Add a Servers widget to the dashboard
A Servers card in the dashboard's Overview row, in the same shape as the
DNS and Secrets cards: a status badge (N offline / N disks nearly full /
All online / No agent data / Not configured), total, online and offline
counts, an online/offline/no-data bar, and the details that matter --
which servers are offline and for how long (the first three, linking to
their pages), which disks are at or above the usage threshold, and a
Linux/Windows split when there are Windows machines.

The numbers come from a new GET /api/servers/summary, computed on the
server with the health monitor's own evaluateHealth. That means "offline"
and "disk full" are decided by exactly the rules the alerts use, so the
widget can't say a server is fine while an alert says it isn't, and it
follows the thresholds set in Settings, which viewers can't read
themselves and so couldn't have applied client-side. Servers under a
maintenance window are marked as such. A server that has never reported
(no agent, or tracked only through Proxmox) counts as "no data" rather
than offline, an offline server's stale disk figure isn't reported, and a
garbled report doesn't blank the widget.

Readable by every signed-in user, like the dashboard itself.

Verified with 23 backend checks (offline/online/no-data classification
including the bare SQLite timestamp format, ordering, colons in Windows
mounts, threshold changes, maintenance flags, agreement with the alert
path's own offline and disk sets, empty and garbled inputs, and the route
as a viewer following Settings) and in a browser against the real router
across five states: problems, disks only, all healthy, no data, and empty.
Real dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-27 02:27:18 +02:00
bobbanandClaude Sonnet 5 fea20456e4 Add a Windows agent (PowerShell)
Reports a Windows machine the way the Linux agent does, replacing the
"planned" stub in agent/windows: scheduled tasks plus hostname, IPv4
addresses, CPU model/cores/current load, memory, every fixed disk, and
TCP/UDP listening ports with the owning process (which feed the Ports
card, localhost-only listeners included).

Scripts (plain ASCII by design -- they are downloaded as text and Windows
PowerShell 5.1 reads BOM-less files as ANSI):
- report-tasks.ps1: collects and POSTs to /api/agent/report. Works in
  Windows PowerShell 5.1 and PowerShell 7. -DryRun prints the JSON.
  Microsoft's own \Microsoft\ tasks (hundreds) are left out unless
  INCLUDE_MICROSOFT_TASKS is set. Triggers are turned into readable text
  ("Weekly on Mon, Wed at 03:00", "At logon", "..., repeating every 15 min").
  Self-signed certificates work via API_INSECURE on both PowerShell
  versions (they need different mechanisms).
- install.ps1: elevated only; downloads the agent to ProgramData, writes
  agent.json with permissions locked to SYSTEM and Administrators *before*
  the token goes in, and registers a SYSTEM scheduled task (every 15 min
  plus at startup with a 2 min delay). Reinstalling replaces the task.
- uninstall.ps1: removes the task and only the files the agent installed.

Server: accepts schedule_type "windows_task"; a server can be registered
as Windows (Add a server has an operating system choice); an agent's
reported os_type ("linux"/"windows", anything else ignored) corrects the
stored one. The Servers page shows the right install and uninstall
command for each OS (Windows PowerShell 5.1 one-liners, with a self-signed
variant and a note about PowerShell 7), and Windows tasks are labelled
"Windows scheduled tasks". The Linux commands are unchanged.

Verified on this Windows machine, in both PowerShell 5.1 and 7:
- Real dry runs found and fixed bugs before anything shipped: tasks and
  ports came out as one nested item (return , $out wrapped twice), integer
  keys in an ordered dictionary index by position (wrong weekday names),
  and generic "Trigger" labels.
- End to end against the real agent-report router: HTTP, self-signed HTTPS
  refused by default and accepted with API_INSECURE, wrong token gives a
  clear one-line error and exit 1, and Swedish letters plus a euro sign
  survive JSON -> UTF-8 -> HTTP -> SQLite.
- 35 checks on trigger/action/duration descriptions, 20 on the installer's
  building blocks (task parts built but not registered, credentials file
  content and ACL, download over HTTP and self-signed HTTPS), 18 on the
  server rules, and the generated one-liners run through PowerShell's
  parser. The documented one-liners were run through iex and stop at the
  administrator check without changing anything.
- Found that PowerShell 7 ignores the ServicePointManager certificate
  override, so the installer's own download now uses -SkipCertificateCheck
  there.

NOT verified: the elevated install itself. Registering a SYSTEM scheduled
task needs elevation and changes the machine, so it was not run: the task
registration, that the repeating trigger really runs indefinitely, and
the agent running as SYSTEM under Task Scheduler have not been exercised.
Windows 10 / Server 2016 or newer is assumed; older is untested.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-27 00:09:00 +02:00
bobbanandClaude Sonnet 5 ae64cb345c Manage server tags in Settings: pre-add tags, recolour, rename, delete
New Settings > Tags tab (admin) listing every tag -- ones servers use and
ones added ahead of time -- with how many servers carry each:
- Add a tag before anything uses it, optionally with a colour. Such tags
  are offered as one-click "Add:" chips (and datalist suggestions) when
  tagging a server, so the same word gets spelled the same way everywhere.
- Give any tag a colour of your choosing, in use or not, or reset it to the
  automatic one. Changes are drafted with Save/Cancel rather than saved as
  the picker drags. Colours show on the Servers page, its tag filter bar,
  the detail page and the editor, and update everywhere without a reload
  through one shared cached colour map.
- Rename a tag; every server that has it is rewritten. Renaming to a name
  that already exists merges the two after a confirmation naming what will
  happen; the target keeps its own colour unless it had none, and a server
  carrying both ends up with one.
- Delete a tag, which removes it from every server that has it, with a
  confirmation stating how many.

Tags still live on the servers (servers.tags); a new tag_definitions table
holds only what a server can't: existence before use, and a colour. A
defined tag stays listed until an admin deletes it, even with no servers.
Rename and delete change the servers and the catalogue in one transaction
so they can't disagree. Only servers that actually carry the tag are
rewritten and counted -- an earlier draft also counted servers whose tags
merely weren't in sorted order, which the tests caught.

Reading the list and colours is open to everyone signed in (needed to draw
tags anywhere); changing the catalogue is admin-only, while tagging a
server stays an operator action. Names go through the same normalisation
as before, colours must be #rrggbb, and every change is audit-logged.

New table tag_definitions (migration 0013).

Verified with 44 backend checks (list/counts, roles, create/adopt/
duplicate/rejects, colour set/reset, rename incl. defined, undefined and
unused tags, merge colour rules and both-sides servers, delete, audit, and
that normal tagging still works afterwards) and in a browser against the
real routers: add with colour, set/reset a colour and see it change on the
Servers page live, merge with confirmation, delete, and the error path.
Not clicked through: the quick-add chips inside the tag editor on a
server's detail page (typechecked; same colour code as the rest).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 21:17:09 +02:00
bobbanandClaude Sonnet 5 72f8c85406 Add a Privacy page: what is stored, where it goes, and your own data
A page every signed-in user can open. It states what the installation
stores (accounts, sign-in sessions, audit and diagnostic logs, server
reports, the secrets tracker, credentials, inventory) and for how long,
where data goes (Authentik, integrations, DNS providers, notification
channels, domain registries, certificate checks, port scans), what lives
in the browser, who can see what, and how to limit or remove data.

Written from what the code actually does, including the uncomfortable
parts: the session file keeps the user's Authentik ID token plus the IP
and browser from sign-in; audit entries keep a name snapshot after an
account is gone; the app has no delete-account function; cron commands
in agent reports can contain sensitive text. It also says what isn't
there -- no telemetry, update checks, third-party scripts, fonts or
tracking cookies -- which was checked against the web build and the
server's outbound calls before being asserted.

Live values rather than boilerplate: log retention (and whether it's on),
which integration and DNS provider types are enabled, how many domains,
certificate checks and reporting servers, and which notification channels
are on. Channel addresses are shown to admins only, and only the host --
never a path, query string or token -- since a webhook URL can embed a key.

Each user also sees their own account and active sign-ins, and can
"Download my data": their account, their sign-ins and the audit-log entries
made under their account, as JSON. Only their own -- never another user's --
and without session ids or ID tokens. The export is itself audit-logged, so
a later export shows it.

Verified with 23 backend checks (own-vs-others isolation for audit counts,
sessions and export; no session ids, ID tokens or channel secrets in any
response; admin-vs-viewer channel visibility; live counts; audit of the
export; auth) using an isolated session directory so real sessions are
never read, and in a browser against the real router, including the
download. Real dev database and session files untouched.

Not legal text: this is a transparency page for the people using the app,
not a privacy policy or a GDPR compliance document.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 20:28:03 +02:00
bobbanandClaude Sonnet 5 07d20bd2b7 Add arm64 compose files: one builds the image, one runs the published one
docker-compose.arm64.build.yml builds the image for linux/arm64 from
source and tags it for the registry (docker compose ... build, then push);
it can also build and run directly on an arm64 machine. Building on x86
works through QEMU, with the one-time binfmt setup noted in the file.

docker-compose.arm64.yml runs that published image on the arm64 machine
with no build there. Both use the same image name and an :arm64 tag, kept
apart from the default image so the existing docker-compose.yml is
untouched. IMAGE_REPO and ARM64_TAG (in .env or the shell) switch the
registry or pin a release tag.

Also adds a .dockerignore, which the repo lacked. The Dockerfile runs
COPY . . after npm ci, so a local build from a developer checkout would
overwrite the container's node_modules with the host's (Windows/x86
binaries) and break the build -- most visibly when cross-building for
arm64. It also keeps .env, data/ and .git out of the build context.

libsql's arm64 musl binary is present in package-lock.json, and the
Dockerfile's node:22-alpine base is multi-arch.

Not built or run: Docker isn't available in this environment, so neither
the emulated arm64 build nor the compose files themselves have been
executed. The YAML was only checked for whitespace and against the
runtime section of docker-compose.yml.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 19:25:52 +02:00
bobbanandClaude Sonnet 5 3f2b5da7be Let the consistency report exclude address ranges
Docker reuses the same subnet on many hosts, and those networks aren't
part of the LAN, so they show up as conflicts and unlisted addresses. The
report only knew about 172.16.0.0/12 through a hardcoded rule; Docker can
just as well pick 192.168.x or 10.x.

The Consistency page now has an Excluded ranges card: CIDR ranges (IPv4
or IPv6) and single addresses, added with a form and removed with one
click, shown to everyone and editable by operators. There's also an
"Exclude range" button on each finding that pre-fills a /24 (or /64) around
its address to edit. Exclusions are applied to servers, IPAM and DNS
before anything is compared, so an excluded address never appears in any
kind of finding, whichever source it came from, and the card says how many
addresses are currently being hidden so it's clear the filter is doing
something.

The old hardcoded rule becomes a visible default (172.16.0.0/12) that can
be removed -- it was silently wrong for anyone using 172.16/12 as a real
LAN. That default is also slightly stronger than before: an address in the
range is now left out even if it is in IPAM or DNS, where the old rule only
skipped it when nothing else mentioned it. Remove or narrow it if that
isn't wanted.

Ranges are validated and normalised on the server (both families, prefix
bounds, no /0, at most 50), a bad one is rejected with a message naming it
and nothing is saved, and changes are audit-logged with before/after.
Matching uses Node's BlockList. Stored as a settings value; managed from
the report rather than admin-only Settings, like ignoring a finding.

Verified with 44 checks (range parsing and rejection, boundary addresses
just inside and outside a range, IPv6, single addresses, exclusion across
all sources and finding kinds, the hidden-address count, route
validation/roles/audit) and in a browser against the real router: add,
invalid, remove the default, exclude from a finding. Real dev database
mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 18:57:28 +02:00
bobbanandClaude Sonnet 5 2352689fd3 Add a consistency report across IPAM, DNS and servers
New Consistency page listing where the three places this app records
what lives at an address disagree:
- Address conflicts: the same address reported by more than one server.
- DNS out of date: a record named after a server (its hostname, or its
  short name) that points at an address the server doesn't report.
- IPAM out of date: an entry labelled with a server's name at an
  address the server doesn't report.
- Not in IPAM: addresses a server reports or DNS points at that IPAM
  doesn't list, merged into one finding per address, with a one-click
  "Add to IPAM" that pre-fills a label.
- No DNS record: server LAN addresses no cached A/AAAA record resolves to.

It compares data the app already holds and fetches nothing when opened,
so the page states how many servers had reported addresses and how many
DNS zones are synced (and how old the oldest sync is) -- DNS records are
only cached for zones that have been synced, and a report that silently
treated missing data as "no records" would mislead.

Rules chosen to keep it from crying wolf:
- Only private addresses are compared; public DNS records aren't expected
  to be in IPAM.
- Servers with no reported addresses are never judged.
- Agents report IPv4 only, so records are only compared within an address
  family (an AAAA record isn't "stale" for lacking an IPv6 address).
- Docker bridge networks (172.16/12) are ignored: shared ones aren't
  conflicts, and they aren't listed unless someone put them in DNS.
- Tailscale addresses don't need DNS records (MagicDNS), and IPAM entries
  kept current by the Tailscale/Proxmox syncs aren't second-guessed.

Findings anyone has decided are fine can be ignored (operators) with a
reason. An ignore is keyed on the finding's stable identity so it stays
ignored across runs, its stored text comes from the finding rather than
the request, and it is marked "no longer occurring" once the condition
goes away. Ignore/restore are audit-logged.

New table consistency_ignores (migration 0012). portScan's private-address
helper is now exported and shared.

Verified with 36 checks (each rule and its exclusions, address-family and
case/trailing-dot handling, IPv6 case, ordering, stable keys, the report
route including a garbled agent report, source counts, ignore/unignore
rules and audit entries) and by driving the page against the real routers
in a browser: Add to IPAM actually created the entry, ignore and restore,
severity filter, "show all", the viewer view, and narrow-width layout
(which found and fixed a squeezed badge and clipped buttons). Real dev
database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 04:08:51 +02:00
bobbanandClaude Sonnet 5 ca0fa817f8 Track domain registration expiry, with daily reminders
New Domains page listing when each domain registration expires, read from
the registry. Domains behind the DNS zones already synced are picked up
automatically; others can be added by hand. You're reminded daily from N
days before expiry (Settings > Notifications, default 30) until it's
renewed, and told when an expiry date hasn't been refreshable for several
days so a stale date isn't trusted silently.

RDAP alone would not have covered this homelab: .se, .nu, .io, .eu and .de
are not in IANA's RDAP bootstrap. Lookups therefore try RDAP where the TLD
publishes a server and fall back to WHOIS on port 43, found via IANA's
own referral, parsing the expiry line out of the free-text answer. Only
the expiry date and registrar are read or stored. Verified live against
the real registries: .se and .nu via WHOIS, .com/.org/.dev via RDAP.

Behaviour worth knowing:
- A DNS zone that is a subdomain (lab.example.se) resolves to the
  registration that actually expires by trying the name and then its
  parents, so no public-suffix list is needed. Zones already covered by a
  tracked domain are not looked up again.
- "Couldn't ask" is never confused with "not registered": network errors,
  rate limits and garbled answers are errors, and a transient error at any
  level stops the walk from concluding the domain doesn't exist.
- A failed refresh keeps the last known expiry and records why, rather
  than blanking a date that's still relied on.
- Zones that don't resolve to a real registration (.lan, .local, unregistered
  names) simply get no row. Zone-derived rows disappear when their zone
  does; manual rows stay. Zone-derived rows can't be deleted by hand.
- Registries that don't publish an expiry (.de, .eu) are tracked with a
  note instead of a date.
- Input like "example.com/path" is refused rather than silently reduced
  to its host.
- Runs on the daily secret-expiry schedule and reminder time, on demand
  (Check all now / per domain), and once at startup if nothing has been
  read in a day. Never blocks startup, one lookup at a time with a pause.
  The warning window lives with the other thresholds in settings.

New table domains (migration 0011); two settings fields (toggle and
warning days).

Verified with 90 checks against fake RDAP/WHOIS backends (name
normalization, date formats, WHOIS parsing including rate-limit and
no-expiry answers, bootstrap and referral caching, stale-cache fallback,
parent walking, add/sync/check/refresh, concurrency guard, alert
selection and stale detection, the daily notification and its toggle,
role rules) plus a live smoke test against real registries and a browser
check of the page against the real router. Real dev database mtime
untouched. Not checked: a screenshot of the finished page (the capture
timed out); structure, sorting, errors and the viewer view were verified.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 03:31:02 +02:00
bobbanandClaude Sonnet 5 017014586f Show when each Gitea repository was last updated
The API already returned Gitea's updated_at for every repo but the page
never displayed it. The Gitea table now has a sortable "Last updated"
column with a relative age ("3 d ago") over the exact date/time (in the
app's configured date format), and the CSV export includes it.

Gitea only exposes a single updated_at per repository (there is no
separate last-push time), so this is Gitea's own notion of "updated".

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 03:20:43 +02:00
bobbanandClaude Sonnet 5 4118062405 Alert when a Semaphore template or Gitea workflow run fails
Every 15 minutes (on the existing health-check timer) the app reads the
latest run of each Semaphore template and each Gitea repo's latest
workflow run. A failed one raises one notification, and another when a
later run succeeds. It is state-based like the server health alerts, so a
job that fails every night alerts on the first failure, not every night.
Gitea alerts include the run's link. Toggle: Settings > Notifications.

What counts:
- Semaphore "error" is a failure, "success" is a pass. A run that is
  waiting, running, stopped by hand or rejected is neither, so it leaves
  the previous state alone: a run in progress must not clear a failure it
  hasn't fixed yet, and a manual stop isn't a failure.
- Gitea failure/success likewise; running, waiting, blocked, cancelled and
  skipped leave things as they were.
- A failing template that gets another failing run does not re-alert.

Not mistaking "couldn't read" for "fixed":
- Semaphore's template listing swallowed per-project errors, so a project
  that failed to load looked like a project with no templates. A new
  checkTemplates adapter method reports which projects failed, and their
  failures are held rather than cleared.
- Gitea reports a run it couldn't fetch as null, the same as "no runs";
  both leave the repo's state alone.
- An unreachable integration holds all of its failures. Nothing is cleared
  or re-announced while it is down.

The first pass only records what is already failing without announcing
it, so upgrading (or adding an integration to a fresh install) doesn't
produce a wall of alerts about months-old failures. That baseline is not
spent while nothing could be read.

Maintenance windows on a Semaphore or Gitea integration silence its
failure alerts with the same rules as the health alerts: a problem that
starts during a window alerts when it ends, and one already announced
stays known. The diff logic is reused from the health monitor rather than
copied. Maintenance page text updated.

Known limit: for Gitea this follows the repo's most recent run on any
workflow or branch, matching what the Gitea page shows; a failure in one
workflow can be masked by a later success of another.

Verified with 53 checks against fake Semaphore and Gitea servers and a
webhook receiver: classification, baseline (including not being consumed
when nothing is readable), single alert per failure, no repeat, in-progress/
stopped/cancelled runs, recovery and re-failure, unreadable project,
unreadable integration, run-fetch errors, maintenance windows (silenced,
then announced after), the toggle, disabled integrations and repos without
Actions. Real dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 03:11:59 +02:00
bobbanandClaude Sonnet 5 9f1609c4ed Add tags to servers, with a tag filter on the Servers page
Servers can be tagged (prod, media, rack-1, ...) for grouping. Operators
and admins edit tags inline on a server's detail page; the input suggests
tags already used on other servers so the same word ends up spelled the
same way everywhere. Tags show as chips that keep one colour per tag, on
the server cards, in the Manage table (and its CSV export), and on the
detail page, where each chip links to the Servers list filtered by that
tag. The Servers page has a tag bar with counts; picking several tags
narrows to servers that have all of them. The filter lives in the URL, so
it survives a refresh and can be linked to. Tags are also matched by the
global search.

Tags are normalized on the server (trimmed, lowercased, spaces become "-",
duplicates merged, sorted); letters in any language are allowed, plus
digits and - _ . : /, at most 30 characters and 12 per server. Invalid
input is rejected with a message naming the offending tag, and nothing is
saved. Changes are audit-logged with the before and after lists.

Stored as a JSON column on servers (migration 0010). Every server response
now returns tags as an array, and the shared response shaping strips the
token hash in one place instead of five.

Verified with 27 backend checks (normalization edge cases including
Swedish letters, roles, validation, search, audit, PATCH/detail/list
shapes) and by driving the real Servers and detail pages against the real
router in a browser (filtering, editing, invalid tag, viewer view). Real
dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 03:06:29 +02:00
bobbanandClaude Sonnet 5 4c11158e98 Add a Ports card to server pages: scan for open ports, find free ones, and keep notes
Each server's detail page now has a Ports card. "Scan…" runs a TCP connect
scan of a chosen range from the app and shows what's open, along with the
ranges that were actually confirmed free; clicking a free range starts a
reservation. Any port can carry a service name and a comment, so the page
also answers "what is this port for". A port with a note counts as taken
even when nothing is listening, which is what makes a reservation work.
Operators can scan and edit; everyone can read. Scans and note changes are
audit-logged.

Details that matter for correctness:
- "Free" means the host actively refused the connection AND nobody has
  claimed the port. A port that never answers (firewall drop, host down)
  is reported as not answering, not as free.
- A scan from elsewhere can't see services bound to localhost only, so the
  agent now also reports what is bound on the host (ss -tulnp) and those
  ports are treated as taken. They show as "local only". Existing agents
  keep working; re-run the install one-liner to add this. The field is
  validated leniently so one odd line can never cost an agent its whole
  report, tasks included.
- If nothing answers at all during a scan, existing results are left
  alone instead of being marked all-closed.
- Scan targets are limited to private addresses (RFC1918, Tailscale
  100.64/10, link-local, IPv6 ULA/link-local); loopback and public
  addresses are refused. Ranges are capped at 20,000 ports, and only one
  scan runs per server at a time.
- Rows exist only while they carry information: an open port, or one with
  a note. A closed port with no note disappears on the next scan; one with
  a note stays as "reserved".

New table server_ports plus two columns on servers (migration 0009).

Verified with 76 backend checks (scanner open/refused/filtered, address
rules, agent report leniency, note/reserve/clear semantics, free-range
calculation including the localhost-only case, roles, concurrency lock,
no-response guard, audit entries, cascade delete) and by driving the real
component against the real router in a browser. Real dev database mtime
untouched.

Not verified: the agent's ss/awk/jq pipeline on a real host — the awk step
was checked against sample ss output and the script passes bash -n, but
jq isn't available here to run the whole thing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 02:39:43 +02:00
bobbanandClaude Sonnet 5 aae4f0d74f Add maintenance mode to silence alerts while working on a server or integration
Rebooting Proxmox or patching a server triggered failure/offline alerts
you then had to dismiss. A maintenance window silences alerts about one
server, integration, or DNS provider for a chosen time. New Maintenance
page (start with a duration and optional reason, end early, see what's
silenced and what isn't) and a banner in the app shell so every signed-in
user can see what is currently silenced. Starting/ending is operator-only
and audit-logged; starting one on a target that already has a window
restarts its clock instead of stacking.

Silenced for the target: server offline/disk alerts, Proxmox/Synology
storage and health alerts, Proxmox backup alerts, and "integration down"
alerts. Not silenced: expiry and update reminders, DNS change notices.

The design goal is that this cannot hide a real outage:
- Every window has a required end (5 min to 7 days); there is no
  open-ended option, so a forgotten window expires by itself.
- A silenced problem is deliberately NOT recorded as "known". If it is
  still present when the window ends it alerts then, as new. A problem
  that was already alerted before the window stays known, so it isn't
  repeated, and is reported cleared only after the window ends.
- Failure alerts keep counting failures during a window without marking
  themselves alerted, so an outage that outlasts the window alerts on the
  very next failed call.

Known limitation, stated on the page: integration-failure alerts are
tracked per service TYPE (all "proxmox"), not per configured instance, so
a window on one Proxmox integration also silences a failure on a second
Proxmox integration while it's open. Fixing that means threading the
integration id through every adapter and the diagnostic log, which is a
much larger change than this feature.

Also moved the API-error-message helper out of Secrets.tsx into a shared
util now that two pages use it. New table maintenance_windows (migration
0008).

Verified with 44 checks: the condition-key-to-subject mapping (including
server:3 vs server:33), the diff rules with silenced subjects (new problem
not recorded, alerts when the window ends; already-known one carried and
not repeated; clears only after the window), window expiry and
integration/DNS-provider source matching, the failure tracker end to end
against a webhook (silent during a window while an unrelated service still
alerts; outage that outlasts the window alerts on the next failure and
only once; fail-and-recover fully inside a window sends nothing), a full
health pass against a real window, and the real router with a stubbed
session (role rules, duration bounds including the missing-duration case,
extend-not-stack, 404s, deleted targets hidden, audit entries). Real dev
database mtime untouched.

Not done: I haven't clicked through the new page or banner in a browser
(they sit behind the Authentik login); it builds and the API behind it is
tested.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 02:23:06 +02:00
bobbanandClaude Sonnet 5 1688de3ea2 Alert when a server goes silent, a disk fills up, or a Synology volume degrades
The data was all being collected (agent last-seen, per-disk usage,
Proxmox storage, Synology volume/disk health) but nothing acted on
it, so a dead server or a full disk was only noticed by opening the
right page.

A new health pass runs every 15 minutes (matching the agent's default
report interval) and raises one notification when a problem starts and
one when it clears: a server's agent silent past a threshold (default
60 min), a server disk / Proxmox storage or root filesystem / Synology
volume at or above a usage threshold (default 90%), and a Synology
volume or disk that isn't "normal", has bad SMART, bad sectors past the
threshold, or life remaining below it. Both thresholds and an on/off
toggle live under Settings -> Notifications.

The parts that make this trustworthy rather than noisy:
- A problem is keyed by identity, so it alerts once and not every run;
  a shared Proxmox storage listed by every node is one problem, not
  one per node.
- Active problems persist across restarts, so a rebuild doesn't
  re-alert everything already known.
- If a source can't be read on a given run (Proxmox/Synology
  unreachable, one node lacking privileges) its existing problems are
  held, not reported "cleared" and then re-alerted when it comes back —
  the integration-failure alert already owns "the integration is down".
- For 20 minutes after startup server-derived problems are held too:
  agents couldn't report while the app was down, so judging them then
  would report every server offline after any restart.
- An offline server's disk figures are stale and are not judged; a
  server that never reported has no agent and raises nothing.
- Tracking continues while the toggle is off (only sending is gated),
  so turning it back on doesn't dump every long-standing problem.

Timestamps without a zone (SQLite's format) are read as UTC; the test
runs on a UTC+2 machine, where reading them as local time gives a
different answer.

Verified with 32 checks: the evaluation rules and the state diff as
pure functions (exact thresholds, the proxmox:1 vs proxmox:10 prefix
trap, the flapping sequence), then a whole pass against a real
Proxmox adapter talking to a fake HTTPS cluster (one node returning
403, the whole API down, a shared storage on two nodes, a node that
recovers), a webhook receiver, the real DB, and the persisted state.

Not exercised end-to-end: the Synology collection path — its rules are
tested on data shaped exactly like the adapter's output types, but I
did not stand up a fake DSM. Real dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 02:15:19 +02:00
bobbanandClaude Sonnet 5 bf03a15337 Make the app installable as a PWA
Adds a web manifest, icons (standard 192/512, a full-bleed maskable
512 with the glyph kept inside the safe zone, and an iOS touch icon),
theme-color/apple meta tags, and a service worker, so "Install app" /
"Add to Home Screen" gives it its own icon and a standalone window.

The service worker is a deliberate no-cache pass-through: it registers
a fetch listener (so browsers treat the app as installable) but never
calls respondWith(), so every request goes to the network exactly as
without a worker. A caching worker would keep serving an old JS bundle
after each rebuild — the same stale-build confusion that already cost
time on the settings layout fix — and this app is a live view of
authenticated data with no useful offline mode. Registered in
production builds only.

Icons are generated procedurally (server-rack glyph on the sidebar's
dark colour) and encoded as real PNGs; visually checked, including
that the maskable variant keeps the glyph inside the safe zone.

Verified in a real Chromium (headless Edge) over the DevTools protocol
against the built bundle served with the same express.static setup as
production: the DevTools installability audit reported no errors, the
manifest parsed with no errors, the worker registered at scope "/",
activated and took control, and a navigation through the worker
returned the app page normally. The manifest is served as
application/manifest+json and sw.js as JavaScript. The in-app preview
pane silently blocks service-worker script fetches, so it could not
be used for this — hence the real browser. Real dev database mtime
untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 02:07:04 +02:00
bobbanandClaude Sonnet 5 7e81306aa7 Read SSL certificate expiry from the live server instead of trusting a typed-in date
A certificate secret's expiry was only ever what someone typed in, so
a renewed cert (or a wrong date) meant the app's reminders were
silently wrong. A certificate secret can now be given a host:port; the
app opens a real TLS connection and reads the certificate's actual
expiry — on create/edit (if the host changes), daily, and via a
per-row "Check now" — and keeps expiryDate in sync. Because the daily
refresh runs before the existing expiry check, the reminder is always
computed from what's actually being served.

Verification is deliberately off for the connection: homelab services
routinely serve self-signed/internal-CA certs, and an already-expired
one is exactly the case worth reporting, which a verifying connection
would refuse before exposing the dates.

Failure handling avoids the silent-staleness this is meant to fix: a
failed check keeps the last known date, records why on the row (shown
as a "Check failed" badge), and is listed in the daily secrets
notification. Creating a monitored secret whose host can't be reached
and with no manual date is rejected with the reason rather than saved
blank. A non-TLS port (the likeliest typo) gets a plain-language
error instead of raw OpenSSL output.

Server-side connections to a user-supplied host:port need the same
operator role that already gates editing secrets (and running Semaphore
templates, which is strictly more powerful); the host is validated
against a strict character set before any connection is made.

New nullable secrets columns (check_host, check_port, last_checked_at,
last_check_error) via migration 0007; existing rows are unaffected.

Verified against real TLS servers (openssl-generated certs) and the
real secrets router with a stubbed session: a live 45-day cert read
back as the correct date via both an IP host (no SNI) and a hostname;
an already-expired cert reported its past date and shows as expired;
refused connections, a server that accepts but never answers (times
out), and a plain non-TLS server each produced a descriptive error
rather than a hang or crash. Through the router: create with a host
and no date reads the date; unreachable host with no date -> 400 with
the reason; unreachable with a manual date -> saved with the error
recorded; host on a non-certificate type and an invalid host string
-> 400; a hand-typed date on a monitored secret is ignored; changing
the host re-checks immediately; changing the type away from
certificate ends monitoring; a viewer gets 403 on Check now. 23 checks,
all passing (a first re-run showed 2 spurious failures that were leftover
rows from the previous run's scratch database, confirmed by a clean re-run).
Real dev database mtime untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-25 23:41:29 +02:00
bobbanandClaude Sonnet 5 0da73711d9 Add a Generator page for server names, usernames, and passwords
New "Generator" nav item, open to every role since it's a pure
client-side utility with no data mutation. Three independent cards:
- Server name: picks from Swedish girl names or Disney characters (or
  both), skipping any name already used by an existing server —
  checked against the real server list, not just avoiding duplicates
  within one session.
- Username: adjective+animal with a configurable separator and an
  optional 2-digit suffix.
- Password: length slider, per-charset toggles, an "exclude ambiguous
  characters" option (0/O, 1/l/I), and a rough entropy/strength
  readout. Uses crypto.getRandomValues with rejection sampling (not
  Math.random or a plain modulo), since a biased password generator
  is a real security footgun.

Nothing generated here is sent to or stored on the server — it's
computed entirely in the browser and only reaches the backend if the
user pastes it into some other form themselves (e.g. Secrets, adding
a server).

Verified generators.ts directly: collision avoidance always returns
the one remaining unused name when every other option in a themed
list is taken, the full-list-exhausted fallback produces a numbered
name that still doesn't collide, matching is case-insensitive,
per-charset password generation only ever produces characters from
the selected sets, excludeAmbiguous holds across 200 40-character
samples, entropy math matches the expected log2 formula, and a
5000-sample single-character distribution came out uniform (469-521
per digit against an expected 500, no modulo bias) confirming the
rejection-sampling RNG is unbiased.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-24 20:56:15 +02:00
bobbanandClaude Sonnet 5 35979cb043 Add bulk actions to Secrets, IP Addresses, and Tailscale devices
One row at a time was tedious for anything beyond a handful of items.
New useSelection hook (id-based Set, so picks made on page 1 survive
moving to page 2 rather than resetting per page) backs a checkbox
column and a bulk-actions bar on each of the three tables:
- Secrets / IP Addresses: bulk delete.
- Tailscale devices: bulk authorize/deauthorize/remove.

No new backend endpoints — each bulk action is a client-side
Promise.allSettled loop over the existing single-item endpoints, so
one failure doesn't block the rest, and the bar reports how many (if
any) failed. Selection clears after a bulk action completes or when
switching Tailscale integrations (device ids aren't comparable across
different tailnets).

Verified the selection hook's Set logic directly (extracted as pure
functions, no React renderer needed): select-all/deselect-all
toggling, individual toggle, and — the case most likely to have a
subtle bug — that selecting items on one page and then selecting
different items on another page keeps both, with toggleAll on either
page only ever touching that page's own ids.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-24 19:13:56 +02:00
bobbanandClaude Sonnet 5 e081eecba4 Stop stretching the settings nav to match content height
Confirmed via real getBoundingClientRect measurements from the user's
own browser that the previous fix's text alignment was already exact
(navItem/heading/saveBtn all centerY=88, pixel-perfect) — so that
wasn't the remaining problem. The actual issue was navCard itself:
1764px tall, because it was stretched (h-100 + row align-items-stretch)
to match the Notifications page's own long column of cards, leaving
~1500px of empty space below the last nav item ("Backup").

Dropped the stretch. The nav card now sits at its own natural, compact
height — still top-aligned with the content column (that part was
never broken) and still using the tightened list-group-item padding
so its text lines up with the content heading, just without forcing
its box to match a content column that can be arbitrarily tall.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 21:46:20 +02:00
bobbanandClaude Sonnet 5 439eccaebb Align settings nav text with the content heading, not just box height
The previous fix made the nav card and content column exactly equal
height, but the two boxes' TOP EDGES were already aligned all along —
what actually looked "off" was that the nav's first item text sat
~12px lower than the content heading/Save button, because Tabler's
default list-group-item padding (1.25rem) is sized for a standalone
list card, not a compact sidebar nav.

Verified with the same pixel-measurement technique as the previous
fix, against the real markup rendered in a browser: before, the nav's
first-item text center sat at y=95.5 while the content heading/button
center sat at y=84 (both card tops already matched at y=64). Tightening
this list-group's own item padding to 0.5rem (scoped via its own
--tblr-list-group-item-padding-y CSS variable, not a global override)
brought the nav text center to y=83.5 — 0.5px off target, imperceptible.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 21:22:45 +02:00
bobbanandClaude Sonnet 5 6ff20d42fd Match the settings sidebar's height to the content column
The nav was a bare list-group with no card wrapper, so it only stood
as tall as its own 6 items while the content column (full of cards)
ran on much longer below it — the two looked like they were floating
at different levels instead of one cohesive layout.

Wrapped the nav in a card (list-group-flush inside it, matching the
visual weight of the content's own cards) and stretched it to the row
's full height. That alone left a stray 16px gap between the two
columns' bottoms — traced to the nav column's mb-3 (meant for spacing
when the columns stack on mobile) counting against the flex-stretch
calculation even at desktop widths, where the columns sit side by
side and don't need it. Changed it to mb-3 mb-md-0, which is what
actually eliminated the gap.

Verified directly in a browser against the compiled markup: measured
the two columns' rendered heights before and after — before, a
16px/only-cosmetic h-100 attempt left a residual gap (traced to the
mb-3 collision above); after fixing that, both columns are exactly
byte-for-byte equal height with matching bottom edges at desktop
width, confirmed via getBoundingClientRect rather than by eye.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 21:08:49 +02:00
bobbanandClaude Sonnet 5 ca61f2a915 Surface Proxmox VMs/LXCs with no backup coverage at all
A failing backup run is visible now, but a guest with no backup job
covering it in the first place was still a silent gap. Rather than
depending on Proxmox's /cluster/backup-info/not-backed-up-guests
endpoint (only exists on newer PVE versions), this derives coverage
from data already fetched: a guest counts as covered if any enabled
job either lists its vmid directly, or backs up "all guests" (scoped
to the job's node, if it has one) without excluding it.

Adds a warning banner plus a full table to the Proxmox page's Backups
card, and extends the existing daily "Proxmox backup failed"
notification (relabeled to mention this too) to also list uncovered
guests, gated by the same toggle.

Verified the coverage logic directly (it's a pure function, so no
fake server needed) across 7 cases: no jobs at all, an all-guests job
with an exclude list, a specific-vmids job, a node-scoped job that
shouldn't cover a guest on a different node, a disabled job providing
no real coverage, two jobs whose combined scope covers everything
neither would alone, and a realistic mixed scenario — all passed.
Confirmed the real dev database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 21:01:48 +02:00
bobbanandClaude Sonnet 5 73d649377e Add notification quiet hours with digest delivery
Instant notifications (DNS changes, integration failure/recovery
alerts) had no way to avoid pinging overnight. Adds a "Quiet hours"
window under Settings -> Notifications: notifications that would fire
during the window are held in a new notification_queue table instead
of sent immediately, then delivered as one combined digest at the end
time (in the same timezone already used for the daily checks) via a
new scheduled flush job. The scheduled daily checks (secret expiry,
Tailscale key, Docker updates, Proxmox backups) already only fire
once at a chosen time, so this mainly matters for the instant ones.
Includes a live "N queued" indicator with a manual "Flush now" button
for visibility, and correctly falls back to sending immediately
whenever the feature is disabled (the default).

Gating lives at the single choke point every notification already
flows through (notify()), so no per-event-type wiring was needed.

Verified against a fake webhook receiver on an isolated scratch
database: exhaustively checked the midnight-wraparound window math
(9 cases including exact-boundary inclusive/exclusive edges) against
synthetic "now" values rather than depending on when the test
happens to run, then end-to-end through the real notify()/
flushQuietHoursQueue() functions with a window constructed around the
actual current time — confirmed a notification during the window
queues instead of sending, the flush produces one digest with the
original title/message intact and clears the queue, a notification
outside the window sends immediately, and disabling the feature
entirely sends immediately regardless of the window. Confirmed the
real dev database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 20:56:33 +02:00
bobbanandClaude Sonnet 5 99db7e1cf0 Surface Proxmox backup job status, with a daily failure notification
Proxmox already runs vzdump backups, but nothing in the app said
whether they were actually succeeding — a silent backup failure is
one of the more dangerous blind spots a homelab admin can have. Adds
a "Backups" card to the Proxmox page: configured backup job
schedules (storage target, which guests, enabled/disabled) from
GET /cluster/backup, and recent vzdump task history per node from
GET /nodes/{node}/tasks?typefilter=vzdump, with a banner at the top
if the most recent run didn't succeed.

New "Proxmox backup failed" notification toggle under Settings ->
Notifications, on the same daily schedule as the other checks. The
scheduler checks each node's own most-recent vzdump run independently
(not just the single most recent task overall) so one node's healthy
backup can't mask another node's failing one in a multi-node cluster.

Known limitation, documented in the adapter's own header comment:
Proxmox's task list doesn't reliably expose which specific guest
failed within an "all guests" job — only the task's own log text has
that — so this surfaces job- and task-level status rather than
guessing at per-guest outcomes.

Verified against a fake Proxmox server (real self-signed HTTPS, since
the adapter's node:https usage can't be monkey-patched under ESM)
reproducing the documented /cluster/backup and task-list response
shapes: job parsing (all-guests+exclude vs specific-vmids+disabled)
correct, task OK/failure parsing correct, and the critical multi-node
scenario confirmed — one node's failing latest run flagged, the
other's healthy latest run correctly left alone, with exactly one
notification of the right content. This reproduces Proxmox's
documented API shape rather than a live-verified one; flag if the
real cluster's response differs in some way this didn't anticipate.
Confirmed the real dev database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 18:42:22 +02:00
bobbanandClaude Sonnet 5 6d673db9ec Add session management: see who's signed in, revoke a session
No visibility existed into who was currently signed in or a way to
force a device out. Sessions already live as files via
session-file-store, so this reads that store directly rather than
adding a new DB table: new Settings-adjacent "Sessions" page
(admin-only, alongside Users) lists every live session with the
user's name/email/resolved role, IP, a friendly "Browser on OS"
summary parsed from the user-agent, last-active time, and expiry, with
a Revoke button per row (extra confirmation if you revoke your own
current session, since that signs you out immediately).

IP and user-agent are now captured into the session at login
(auth/router.ts) since express-session doesn't track them itself.
session-file-store's own Store type doesn't declare its list()
method, so sessionStore.ts adds a narrow local interface for it rather
than losing type safety on the rest of the store.

Verified against a real session directory seeded through the actual
session-file-store APIs (not hand-written JSON): confirmed correct
field resolution including a session whose user row was later deleted
(role resolves to null instead of crashing), correctly excluded a
mid-OIDC-login session with no completed user yet, correctly excluded
an already-expired session, and confirmed revoke actually deletes the
right session file and only that one. Confirmed the real dev
database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-21 20:29:29 +02:00
bobbanandClaude Sonnet 5 1cae35a59e Add a daily Docker image-update notification
Dockhand's pending-update counts were only visible if you happened to
open the Docker page. New "Docker image update available" toggle
under Settings -> Notifications, sharing the same daily
time/timezone as the secret and Tailscale key expiry reminders (same
node-schedule reschedule-on-settings-change pattern as those two).
Reads each enabled Dockhand integration's already-cached update-check
results via listContainers() rather than triggering a fresh
per-container registry lookup, so it costs nothing extra beyond what
the Docker page itself already fetches, and lists every container
with an update pending across all environments/integrations in one
notification.

Verified end-to-end against a fake local Dockhand server (one
environment, two containers, one flagged with a pending update) and a
fake webhook receiver on an isolated scratch database: the check
correctly found only the flagged container (with its newerVersion),
sent exactly one notification with the right title/content, excluded
the up-to-date container, and sent nothing at all when the setting
was toggled off despite still finding the same pending update.
Confirmed the real dev database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 23:44:47 +02:00
bobbanandClaude Sonnet 5 9c07718d2d Fix command palette results being unclickable
The backdrop div was nested inside .modal instead of rendered as its
sibling. Tabler/Bootstrap's backdrop (z-index 1050) is only supposed
to sit below the modal (z-index 1055) because normally they're both
direct children of the same stacking context; nesting it inside
.modal instead made it establish z-index inside .modal's own stacking
context, where it painted on top of .modal-dialog and silently
absorbed every click meant for a result row (confirmed via
elementFromPoint in a real browser: clicking a result hit
.modal-backdrop, not the result div). Moved the backdrop to be a
sibling rendered before .modal, matching the structure elementFromPoint
now confirms resolves to the actual result element.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 23:02:55 +02:00
bobbanandClaude Sonnet 5 5c1376cf1c Make search results open the actual entry, not just its list page
Previously a search result for a secret/IP/integration/DNS record
only filtered or scrolled to the right page — you still had to find
and click it yourself. Now:
- Secrets/IPAM results pass ?editId= and the page auto-opens that
  row's existing edit form on load (viewers still just get the ?q=
  filter, since they can't edit).
- Integration results now link to /integrations?editId= (which
  auto-opens IntegrationEditForm for that row) instead of the type's
  live dashboard page (/proxmox, /tailscale, etc.) — the dashboard
  can't distinguish between two integrations of the same type anyway,
  so it wasn't landing on "the" entry when more than one existed.
- DNS record results now also pass recordType/recordName/recordContent
  so Dns.tsx, after auto-selecting the right zone, finds that exact
  record by (type, name, content) and opens its edit form too. That
  triple was used instead of an id because the records-list endpoint
  returns both the cache table's internal integer id and the
  provider's own string record id under different field names, and
  matching by content sidesteps that ambiguity entirely rather than
  risking picking the wrong one.

Servers and DNS zones/providers already landed on their real entry
(server detail page; zone/provider auto-selected) so those are
unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 19:16:09 +02:00
bobbanandClaude Sonnet 5 009bb3e027 Add a global search / command palette (Ctrl/Cmd+K)
With 6 integrations, DNS, secrets, IPAM, and servers all in one app,
there was no single place to type a hostname/IP/name and jump
straight to it. New GET /api/search aggregates a LIKE-based search
across servers, secrets, IPAM, integrations, DNS providers, DNS
zones, and DNS records (joining zone/provider names onto each record
result) in one round trip — homelab-scale row counts make a naive
LIKE scan plenty fast, no FTS needed. Every underlying resource's own
list endpoint already only requires requireAuth (viewer role
included), so the aggregate endpoint uses the same single check.

Frontend: a self-contained CommandPalette component (Tabler's modal
CSS classes driven by React state, since the app doesn't load
Bootstrap's JS) opens via a sidebar search button or Ctrl/Cmd+K from
anywhere, debounces input, and supports arrow-key navigation. Results
link to the right page: servers to their existing /servers/:id
detail route; DNS zone/record results deep-link via new ?providerId=
&zoneId= query-param handling added to Dns.tsx (auto-selects that
provider/zone on load, since Dns.tsx previously held selection only
in local state with no URL sync); secrets/IPAM results link with
?q= to prefill each page's existing client-side search box;
integrations link to their type's dashboard page (no per-instance
route exists yet, so same-type integrations share one link).

Verified the query logic (joins, case-insensitive LIKE, correct
zone/provider name resolution) against an isolated scratch database
seeded with realistic cross-referencing rows — a server name match, a
case-mismatched match against both a secret and a DNS provider
sharing "cloudflare", an IP address matching both an IPAM entry and
the DNS A record pointing at it (confirming the record's joined zone
and provider names came through correctly), an integration name
match, and a no-match query returning every category empty. Did not
re-verify the requireAuth/asyncHandler wiring itself, since it's the
same one-line pattern already proven across every other router in
this app. Confirmed the real dev database's mtime was untouched
throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 03:20:05 +02:00
bobbanandClaude Sonnet 5 23eb7f0d70 Add passphrase-protected export/import for integrations, DNS providers, and settings
Nothing let you back up or migrate the app's own configuration short
of copying the raw SQLite file. Adds Settings -> Backup: export
decrypts every integration/DNS provider credential (normally
encrypted at rest with this server's CREDENTIALS_ENCRYPTION_KEY) and
re-encrypts the whole payload with a passphrase you choose (scrypt-
derived key, AES-256-GCM), so the file is portable to a different
instance with a different encryption key rather than being tied to
this one. Import decrypts with that passphrase and merges settings
onto the current ones; integrations/DNS providers are only added when
no existing row shares their type+name, so re-running an import never
duplicates or overwrites a working credential.

Scope is configuration only — no DNS records, secrets, IPAM, servers,
or audit/diagnostic log data.

Verified end-to-end against two isolated scratch databases with
different encryption keys (proving actual cross-instance portability,
not just round-tripping through the same key): export -> encrypt ->
write file -> decrypt on the other DB -> import -> re-decrypt the
newly created integration/provider using the target's own key,
confirming the plaintext credentials survived correctly; a wrong
passphrase failed loudly (GCM auth failure) as expected; and
re-running the same import a second time skipped both rows instead of
duplicating them. Confirmed the real dev database's mtime was
untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 21:18:41 +02:00
bobbanandClaude Sonnet 5 b5a4c6e2d9 Alert when an integration or DNS provider fails repeatedly
The Diagnostic Log already records every outbound call's success or
failure, but nothing acted on it — you'd only notice an integration
was down by happening to open its page. Adds a per-source consecutive-
failure counter (in-memory, reset on restart, same durability tier as
the diag log's own ring buffer) hooked into recordDiagEntry: crossing
the configurable threshold (default 3) sends one "down" notification
on every configured channel, and a "recovered" notification fires once
it succeeds again — no repeat spam while it stays down. New
"Integration/DNS provider failing repeatedly" toggle and threshold
field under Settings -> Notifications.

Verified end-to-end against an isolated scratch database with a real
local HTTP server standing in for the webhook channel: 5 consecutive
failures produced exactly one "Down" notification (at the 3rd
failure, correctly naming "3 calls"), a subsequent success produced
exactly one "Recovered" notification, and two more failures on a
fresh streak triggered nothing (below threshold) — confirmed the real
dev database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 14:51:00 +02:00
bobbanandClaude Sonnet 5 10d123b18a Add automatic retention purging for the Diagnostic and Audit logs
The diagnostic log already rings-buffer to 500 rows, but the audit
log had no cap at all and would grow forever. Adds an opt-in
age-based purge under Settings -> Logs: keep entries for N days,
checked on a configurable interval (hourly through monthly), plus a
manual "Purge now" button. Reuses the existing node-schedule-style
reschedule-on-settings-change pattern from the secret/Tailscale
expiry checkers, but as a plain setInterval since "how often" here is
an interval rather than a specific daily time.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 02:22:32 +02:00
bobbanandClaude Sonnet 5 af3f7e77d2 Document required access per integration and DNS provider
Each adapter performs write actions, not just reads (start/stop a
guest, edit a DNS record, rerun a CI job, etc.), so a read-only
credential silently works for the dashboard views but fails the
moment you use an action. Lists the exact endpoints/permissions
needed per target system, drawn from each adapter's own auth code and
header comments (e.g. Proxmox's Sys.Audit/Datastore.Audit split,
Azure's DNS Zone Contributor role, Loopia/Pi-hole/cPanel having no
scoped-credential option at all).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 02:08:40 +02:00
bobban 19f84838cc Revert "Move dark mode toggle from sidebar to Settings > Display"
This reverts commit 4ba86429d5.
2026-09-19 01:58:18 +02:00
bobban 4ba86429d5 Move dark mode toggle from sidebar to Settings > Display 2026-09-19 01:56:08 +02:00
bobbanandClaude Sonnet 5 5c35080e22 Add a dark mode toggle
Personal, per-browser preference (localStorage), not an admin-wide
Display setting like date/time format — every role can pick their own.
Sets data-bs-theme on <html> to hook into Tabler's built-in dark
variant, applied via an inline script in index.html before React
mounts so there's no light-mode flash on load. Toggle button lives in
the sidebar footer next to Sign out.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 01:51:59 +02:00
bobbanandClaude Sonnet 5 82ed25b63f Fix checkbox/label spacing on Notifications and other form checkboxes
Tabler's form-check-single modifier zeroes margin on .form-check-input,
which cancels the -2rem margin-inline-start the base .form-check rule
relies on to offset the checkbox into its label's padding gutter. That
left checkboxes floating almost flush against their label text.

Removed the form-check-single class from all 9 affected checkboxes:
6 in NotificationSettings, 1 each in DnsProviderForm, IntegrationForm,
and IntegrationEditForm.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 01:43:45 +02:00
bobbanandClaude Sonnet 5 08eb0bd87e Make DNS and Secrets dashboard widgets match the integration widgets
DNS and Secrets each rendered as three separate stat-tile cards plus a
fourth full-width breakdown card -- visually a different family from
the single self-contained card each integration widget uses (label +
status badge, a compact stat row, then its own breakdown bar inline).

Extracted the integration widgets' card shell into a shared WidgetCard
component (label + badge header, children below) and rebuilt DNS and
Secrets on top of it instead of StatTile/TypeBreakdown, which are now
unused and removed. Also refactored the Integrations map itself onto
WidgetCard, so all eight dashboard widgets (DNS, Secrets, and the six
integrations) are now literally the same component, not just visually
similar.

DNS and Secrets get a badge like the integrations' Connected/Not
connected -- "Not configured" (gray) when nothing's set up yet, or a
status summary otherwise ("X enabled" for DNS; "All OK" / "X expiring"
/ "X expired" for Secrets, mirroring how a failing integration shows a
warning-colored count instead of a plain badge). They now sit under
one "Overview" heading in a shared row instead of two separate
sections, matching how the integration widgets already share one
row-cards grid.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 00:07:32 +02:00
bobbanandClaude Sonnet 5 b40234a557 Make table page size a configurable setting, not a hardcoded 20
Pagination just landed hardcoded to 20 rows everywhere; add a way to
change that instead of leaving it fixed for every table in the app.

pageSize joins dateFormat/timeFormat on the existing display settings
object (server-side default 20, 5-500 range enforced by the PUT
schema) rather than becoming its own settings section, since it's the
same kind of thing -- an admin-configured, globally-applied display
preference read by every signed-in role via the already-public GET
/api/settings/display endpoint, same as the date/time format already
works.

Renamed the "Date & Time" settings tab/page/route to "Display" (still
just one component, now covering both date/time format and table
pagination) since its scope no longer matches the old name -- kept
Settings.tsx's usual pattern of one page per concern rather than
adding a second, oddly-scoped tab just for one number field.

New web/src/utils/pageSize.ts mirrors utils/date.ts's existing
module-level "set once at startup, read anywhere without prop-
drilling" pattern; usePagination()'s pageSize parameter now defaults
to getPageSize() instead of a literal 20, evaluated fresh on every
call so it picks up a saved change without touching any of the ten
pages already using the hook.

Verified server-side against a temp SQLite DB: pageSize defaults to
20, a partial update sets it without disturbing dateFormat/timeFormat
and vice versa, and it persists across a fresh settings read. Also
checked the default-parameter mechanics directly (re-evaluates the
global value on every call rather than capturing it once, and an
explicit override still wins).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-18 23:58:14 +02:00
bobbanandClaude Sonnet 5 7bf9b03839 Paginate tables that can grow large
Sorting/CSV export were already app-wide; tables with a real chance of
growing into dozens or hundreds of rows (a busy tailnet, a big DNS
zone, a homelab's full IP inventory, a Gitea org with many repos, ...)
had no pagination at all, making them a long unbroken scroll.

New usePagination hook (client-side slicing over an already-sorted/
filtered array, 20 rows per page) and a matching Pagination component
(Prev/Next + "Page X of Y (N total)", hidden entirely when everything
fits on one page). The current page is clamped to the valid range on
every render rather than reset via an effect, so switching to a
smaller data set (a different selected integration, a filter that
narrows the result) can never strand the view on a now-nonexistent
page -- no per-page "reset on change" wiring needed anywhere.

Applied to Audit Log, DNS zones and records, IP Addresses, Secrets,
Servers (manage table), Docker containers, Proxmox guests, Semaphore
templates, Gitea repos, and Tailscale devices. CSV export keeps
exporting the full sorted/filtered array regardless of which page is
currently shown -- pagination only affects what's rendered on screen.
Left the already-small tables (Synology volumes/disks, Users,
Integrations, per-node Proxmox storage) unpaginated, and left the
Diagnostic Log's existing server-driven pagination as-is rather than
bolting a second, different pagination scheme onto it.

Verified the clamping logic directly: a normal page, the trailing
partial page, a requested page beyond the end (clamps to the last
valid page instead of rendering empty), and an empty result set
(clamps to page 0 with a page count of 1 instead of a negative range).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-18 23:51:14 +02:00
bobbanandClaude Sonnet 5 81e55fc792 Show per-disk usage for Proxmox-linked servers too, not just a total
The agent path showed a per-mount usage table; a Proxmox-linked server
only ever got a single "Disk (allocated)" figure, since that's all the
VM/LXC config alone can tell you -- it's the attached disk's declared
size, not how full it actually is inside the guest. Fix the actual gap
instead of just matching the display: fetch real usage where Proxmox
can see it.

adapter.getGuestDetail() gained a `disks` field (same {mount,
sizeBytes, usedBytes} shape the agent already reports, so the frontend
renders both identically):
- LXC: the host can read straight into the container's root
  filesystem, no agent needed -- status/current's disk/maxdisk fields
  are real usage, not just allocation.
- QEMU: the hypervisor can't see inside a virtual disk at all without
  help, so this calls the QEMU guest agent's get-fsinfo command (same
  "gracefully degrade if the agent's missing/older" tolerance already
  used for its IP-address lookup, and independent of it -- one
  command failing doesn't take out the other). Pseudo-filesystems
  (tmpfs, etc.) are filtered out by checking for a non-empty backing
  `disk` array, the common convention for this endpoint.

Extracted the disks-table JSX (previously only in the agent branch)
into a shared DisksTable component and used it in both branches, and
added a note explaining an empty result when a running QEMU VM's
guest agent doesn't support get-fsinfo (an older agent version).

Verified against a mock Proxmox API over real TLS: an LXC's root
usage, a QEMU VM's real fsinfo mounts (with the disk-less tmpfs entry
correctly filtered), and a QEMU VM whose get-fsinfo fails outright --
confirming that degrades to an empty disks list without throwing and
without affecting the separate network-get-interfaces result.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-18 23:32:46 +02:00
bobbanandClaude Sonnet 5 08a984719f Let admins hide the Proxmox-link card per server
Not every registered server is a Proxmox VM/LXC -- bare-metal boxes
and other hosts had no reason to show a "link to Proxmox" option, but
it appeared unconditionally on every server's detail page.

New hideProxmoxLink column on servers (default false, so existing
behavior is unchanged until someone opts in). A "Not a VM? Hide this"
link in the card's header sets it; once hidden, a small "+ Show
Proxmox link options" link takes its place so it's still reachable,
not buried in a settings form. The card always shows regardless of
this flag once a server IS actually linked, so unlinking never becomes
unreachable by hiding the card out from under an active link.

Verified against a temp SQLite DB with real migrations: a new server
defaults to false, and toggling true/false both persist correctly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-18 21:34:52 +02:00