New Consistency page listing where the three places this app records
what lives at an address disagree:
- Address conflicts: the same address reported by more than one server.
- DNS out of date: a record named after a server (its hostname, or its
short name) that points at an address the server doesn't report.
- IPAM out of date: an entry labelled with a server's name at an
address the server doesn't report.
- Not in IPAM: addresses a server reports or DNS points at that IPAM
doesn't list, merged into one finding per address, with a one-click
"Add to IPAM" that pre-fills a label.
- No DNS record: server LAN addresses no cached A/AAAA record resolves to.
It compares data the app already holds and fetches nothing when opened,
so the page states how many servers had reported addresses and how many
DNS zones are synced (and how old the oldest sync is) -- DNS records are
only cached for zones that have been synced, and a report that silently
treated missing data as "no records" would mislead.
Rules chosen to keep it from crying wolf:
- Only private addresses are compared; public DNS records aren't expected
to be in IPAM.
- Servers with no reported addresses are never judged.
- Agents report IPv4 only, so records are only compared within an address
family (an AAAA record isn't "stale" for lacking an IPv6 address).
- Docker bridge networks (172.16/12) are ignored: shared ones aren't
conflicts, and they aren't listed unless someone put them in DNS.
- Tailscale addresses don't need DNS records (MagicDNS), and IPAM entries
kept current by the Tailscale/Proxmox syncs aren't second-guessed.
Findings anyone has decided are fine can be ignored (operators) with a
reason. An ignore is keyed on the finding's stable identity so it stays
ignored across runs, its stored text comes from the finding rather than
the request, and it is marked "no longer occurring" once the condition
goes away. Ignore/restore are audit-logged.
New table consistency_ignores (migration 0012). portScan's private-address
helper is now exported and shared.
Verified with 36 checks (each rule and its exclusions, address-family and
case/trailing-dot handling, IPv6 case, ordering, stable keys, the report
route including a garbled agent report, source counts, ignore/unignore
rules and audit entries) and by driving the page against the real routers
in a browser: Add to IPAM actually created the entry, ignore and restore,
severity filter, "show all", the viewer view, and narrow-width layout
(which found and fixed a squeezed badge and clipped buttons). Real dev
database mtime untouched.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
New Domains page listing when each domain registration expires, read from
the registry. Domains behind the DNS zones already synced are picked up
automatically; others can be added by hand. You're reminded daily from N
days before expiry (Settings > Notifications, default 30) until it's
renewed, and told when an expiry date hasn't been refreshable for several
days so a stale date isn't trusted silently.
RDAP alone would not have covered this homelab: .se, .nu, .io, .eu and .de
are not in IANA's RDAP bootstrap. Lookups therefore try RDAP where the TLD
publishes a server and fall back to WHOIS on port 43, found via IANA's
own referral, parsing the expiry line out of the free-text answer. Only
the expiry date and registrar are read or stored. Verified live against
the real registries: .se and .nu via WHOIS, .com/.org/.dev via RDAP.
Behaviour worth knowing:
- A DNS zone that is a subdomain (lab.example.se) resolves to the
registration that actually expires by trying the name and then its
parents, so no public-suffix list is needed. Zones already covered by a
tracked domain are not looked up again.
- "Couldn't ask" is never confused with "not registered": network errors,
rate limits and garbled answers are errors, and a transient error at any
level stops the walk from concluding the domain doesn't exist.
- A failed refresh keeps the last known expiry and records why, rather
than blanking a date that's still relied on.
- Zones that don't resolve to a real registration (.lan, .local, unregistered
names) simply get no row. Zone-derived rows disappear when their zone
does; manual rows stay. Zone-derived rows can't be deleted by hand.
- Registries that don't publish an expiry (.de, .eu) are tracked with a
note instead of a date.
- Input like "example.com/path" is refused rather than silently reduced
to its host.
- Runs on the daily secret-expiry schedule and reminder time, on demand
(Check all now / per domain), and once at startup if nothing has been
read in a day. Never blocks startup, one lookup at a time with a pause.
The warning window lives with the other thresholds in settings.
New table domains (migration 0011); two settings fields (toggle and
warning days).
Verified with 90 checks against fake RDAP/WHOIS backends (name
normalization, date formats, WHOIS parsing including rate-limit and
no-expiry answers, bootstrap and referral caching, stale-cache fallback,
parent walking, add/sync/check/refresh, concurrency guard, alert
selection and stale detection, the daily notification and its toggle,
role rules) plus a live smoke test against real registries and a browser
check of the page against the real router. Real dev database mtime
untouched. Not checked: a screenshot of the finished page (the capture
timed out); structure, sorting, errors and the viewer view were verified.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Every 15 minutes (on the existing health-check timer) the app reads the
latest run of each Semaphore template and each Gitea repo's latest
workflow run. A failed one raises one notification, and another when a
later run succeeds. It is state-based like the server health alerts, so a
job that fails every night alerts on the first failure, not every night.
Gitea alerts include the run's link. Toggle: Settings > Notifications.
What counts:
- Semaphore "error" is a failure, "success" is a pass. A run that is
waiting, running, stopped by hand or rejected is neither, so it leaves
the previous state alone: a run in progress must not clear a failure it
hasn't fixed yet, and a manual stop isn't a failure.
- Gitea failure/success likewise; running, waiting, blocked, cancelled and
skipped leave things as they were.
- A failing template that gets another failing run does not re-alert.
Not mistaking "couldn't read" for "fixed":
- Semaphore's template listing swallowed per-project errors, so a project
that failed to load looked like a project with no templates. A new
checkTemplates adapter method reports which projects failed, and their
failures are held rather than cleared.
- Gitea reports a run it couldn't fetch as null, the same as "no runs";
both leave the repo's state alone.
- An unreachable integration holds all of its failures. Nothing is cleared
or re-announced while it is down.
The first pass only records what is already failing without announcing
it, so upgrading (or adding an integration to a fresh install) doesn't
produce a wall of alerts about months-old failures. That baseline is not
spent while nothing could be read.
Maintenance windows on a Semaphore or Gitea integration silence its
failure alerts with the same rules as the health alerts: a problem that
starts during a window alerts when it ends, and one already announced
stays known. The diff logic is reused from the health monitor rather than
copied. Maintenance page text updated.
Known limit: for Gitea this follows the repo's most recent run on any
workflow or branch, matching what the Gitea page shows; a failure in one
workflow can be masked by a later success of another.
Verified with 53 checks against fake Semaphore and Gitea servers and a
webhook receiver: classification, baseline (including not being consumed
when nothing is readable), single alert per failure, no repeat, in-progress/
stopped/cancelled runs, recovery and re-failure, unreadable project,
unreadable integration, run-fetch errors, maintenance windows (silenced,
then announced after), the toggle, disabled integrations and repos without
Actions. Real dev database mtime untouched.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Servers can be tagged (prod, media, rack-1, ...) for grouping. Operators
and admins edit tags inline on a server's detail page; the input suggests
tags already used on other servers so the same word ends up spelled the
same way everywhere. Tags show as chips that keep one colour per tag, on
the server cards, in the Manage table (and its CSV export), and on the
detail page, where each chip links to the Servers list filtered by that
tag. The Servers page has a tag bar with counts; picking several tags
narrows to servers that have all of them. The filter lives in the URL, so
it survives a refresh and can be linked to. Tags are also matched by the
global search.
Tags are normalized on the server (trimmed, lowercased, spaces become "-",
duplicates merged, sorted); letters in any language are allowed, plus
digits and - _ . : /, at most 30 characters and 12 per server. Invalid
input is rejected with a message naming the offending tag, and nothing is
saved. Changes are audit-logged with the before and after lists.
Stored as a JSON column on servers (migration 0010). Every server response
now returns tags as an array, and the shared response shaping strips the
token hash in one place instead of five.
Verified with 27 backend checks (normalization edge cases including
Swedish letters, roles, validation, search, audit, PATCH/detail/list
shapes) and by driving the real Servers and detail pages against the real
router in a browser (filtering, editing, invalid tag, viewer view). Real
dev database mtime untouched.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Each server's detail page now has a Ports card. "Scan…" runs a TCP connect
scan of a chosen range from the app and shows what's open, along with the
ranges that were actually confirmed free; clicking a free range starts a
reservation. Any port can carry a service name and a comment, so the page
also answers "what is this port for". A port with a note counts as taken
even when nothing is listening, which is what makes a reservation work.
Operators can scan and edit; everyone can read. Scans and note changes are
audit-logged.
Details that matter for correctness:
- "Free" means the host actively refused the connection AND nobody has
claimed the port. A port that never answers (firewall drop, host down)
is reported as not answering, not as free.
- A scan from elsewhere can't see services bound to localhost only, so the
agent now also reports what is bound on the host (ss -tulnp) and those
ports are treated as taken. They show as "local only". Existing agents
keep working; re-run the install one-liner to add this. The field is
validated leniently so one odd line can never cost an agent its whole
report, tasks included.
- If nothing answers at all during a scan, existing results are left
alone instead of being marked all-closed.
- Scan targets are limited to private addresses (RFC1918, Tailscale
100.64/10, link-local, IPv6 ULA/link-local); loopback and public
addresses are refused. Ranges are capped at 20,000 ports, and only one
scan runs per server at a time.
- Rows exist only while they carry information: an open port, or one with
a note. A closed port with no note disappears on the next scan; one with
a note stays as "reserved".
New table server_ports plus two columns on servers (migration 0009).
Verified with 76 backend checks (scanner open/refused/filtered, address
rules, agent report leniency, note/reserve/clear semantics, free-range
calculation including the localhost-only case, roles, concurrency lock,
no-response guard, audit entries, cascade delete) and by driving the real
component against the real router in a browser. Real dev database mtime
untouched.
Not verified: the agent's ss/awk/jq pipeline on a real host — the awk step
was checked against sample ss output and the script passes bash -n, but
jq isn't available here to run the whole thing.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Rebooting Proxmox or patching a server triggered failure/offline alerts
you then had to dismiss. A maintenance window silences alerts about one
server, integration, or DNS provider for a chosen time. New Maintenance
page (start with a duration and optional reason, end early, see what's
silenced and what isn't) and a banner in the app shell so every signed-in
user can see what is currently silenced. Starting/ending is operator-only
and audit-logged; starting one on a target that already has a window
restarts its clock instead of stacking.
Silenced for the target: server offline/disk alerts, Proxmox/Synology
storage and health alerts, Proxmox backup alerts, and "integration down"
alerts. Not silenced: expiry and update reminders, DNS change notices.
The design goal is that this cannot hide a real outage:
- Every window has a required end (5 min to 7 days); there is no
open-ended option, so a forgotten window expires by itself.
- A silenced problem is deliberately NOT recorded as "known". If it is
still present when the window ends it alerts then, as new. A problem
that was already alerted before the window stays known, so it isn't
repeated, and is reported cleared only after the window ends.
- Failure alerts keep counting failures during a window without marking
themselves alerted, so an outage that outlasts the window alerts on the
very next failed call.
Known limitation, stated on the page: integration-failure alerts are
tracked per service TYPE (all "proxmox"), not per configured instance, so
a window on one Proxmox integration also silences a failure on a second
Proxmox integration while it's open. Fixing that means threading the
integration id through every adapter and the diagnostic log, which is a
much larger change than this feature.
Also moved the API-error-message helper out of Secrets.tsx into a shared
util now that two pages use it. New table maintenance_windows (migration
0008).
Verified with 44 checks: the condition-key-to-subject mapping (including
server:3 vs server:33), the diff rules with silenced subjects (new problem
not recorded, alerts when the window ends; already-known one carried and
not repeated; clears only after the window), window expiry and
integration/DNS-provider source matching, the failure tracker end to end
against a webhook (silent during a window while an unrelated service still
alerts; outage that outlasts the window alerts on the next failure and
only once; fail-and-recover fully inside a window sends nothing), a full
health pass against a real window, and the real router with a stubbed
session (role rules, duration bounds including the missing-duration case,
extend-not-stack, 404s, deleted targets hidden, audit entries). Real dev
database mtime untouched.
Not done: I haven't clicked through the new page or banner in a browser
(they sit behind the Authentik login); it builds and the API behind it is
tested.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The data was all being collected (agent last-seen, per-disk usage,
Proxmox storage, Synology volume/disk health) but nothing acted on
it, so a dead server or a full disk was only noticed by opening the
right page.
A new health pass runs every 15 minutes (matching the agent's default
report interval) and raises one notification when a problem starts and
one when it clears: a server's agent silent past a threshold (default
60 min), a server disk / Proxmox storage or root filesystem / Synology
volume at or above a usage threshold (default 90%), and a Synology
volume or disk that isn't "normal", has bad SMART, bad sectors past the
threshold, or life remaining below it. Both thresholds and an on/off
toggle live under Settings -> Notifications.
The parts that make this trustworthy rather than noisy:
- A problem is keyed by identity, so it alerts once and not every run;
a shared Proxmox storage listed by every node is one problem, not
one per node.
- Active problems persist across restarts, so a rebuild doesn't
re-alert everything already known.
- If a source can't be read on a given run (Proxmox/Synology
unreachable, one node lacking privileges) its existing problems are
held, not reported "cleared" and then re-alerted when it comes back —
the integration-failure alert already owns "the integration is down".
- For 20 minutes after startup server-derived problems are held too:
agents couldn't report while the app was down, so judging them then
would report every server offline after any restart.
- An offline server's disk figures are stale and are not judged; a
server that never reported has no agent and raises nothing.
- Tracking continues while the toggle is off (only sending is gated),
so turning it back on doesn't dump every long-standing problem.
Timestamps without a zone (SQLite's format) are read as UTC; the test
runs on a UTC+2 machine, where reading them as local time gives a
different answer.
Verified with 32 checks: the evaluation rules and the state diff as
pure functions (exact thresholds, the proxmox:1 vs proxmox:10 prefix
trap, the flapping sequence), then a whole pass against a real
Proxmox adapter talking to a fake HTTPS cluster (one node returning
403, the whole API down, a shared storage on two nodes, a node that
recovers), a webhook receiver, the real DB, and the persisted state.
Not exercised end-to-end: the Synology collection path — its rules are
tested on data shaped exactly like the adapter's output types, but I
did not stand up a fake DSM. Real dev database mtime untouched.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds a web manifest, icons (standard 192/512, a full-bleed maskable
512 with the glyph kept inside the safe zone, and an iOS touch icon),
theme-color/apple meta tags, and a service worker, so "Install app" /
"Add to Home Screen" gives it its own icon and a standalone window.
The service worker is a deliberate no-cache pass-through: it registers
a fetch listener (so browsers treat the app as installable) but never
calls respondWith(), so every request goes to the network exactly as
without a worker. A caching worker would keep serving an old JS bundle
after each rebuild — the same stale-build confusion that already cost
time on the settings layout fix — and this app is a live view of
authenticated data with no useful offline mode. Registered in
production builds only.
Icons are generated procedurally (server-rack glyph on the sidebar's
dark colour) and encoded as real PNGs; visually checked, including
that the maskable variant keeps the glyph inside the safe zone.
Verified in a real Chromium (headless Edge) over the DevTools protocol
against the built bundle served with the same express.static setup as
production: the DevTools installability audit reported no errors, the
manifest parsed with no errors, the worker registered at scope "/",
activated and took control, and a navigation through the worker
returned the app page normally. The manifest is served as
application/manifest+json and sw.js as JavaScript. The in-app preview
pane silently blocks service-worker script fetches, so it could not
be used for this — hence the real browser. Real dev database mtime
untouched throughout.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A certificate secret's expiry was only ever what someone typed in, so
a renewed cert (or a wrong date) meant the app's reminders were
silently wrong. A certificate secret can now be given a host:port; the
app opens a real TLS connection and reads the certificate's actual
expiry — on create/edit (if the host changes), daily, and via a
per-row "Check now" — and keeps expiryDate in sync. Because the daily
refresh runs before the existing expiry check, the reminder is always
computed from what's actually being served.
Verification is deliberately off for the connection: homelab services
routinely serve self-signed/internal-CA certs, and an already-expired
one is exactly the case worth reporting, which a verifying connection
would refuse before exposing the dates.
Failure handling avoids the silent-staleness this is meant to fix: a
failed check keeps the last known date, records why on the row (shown
as a "Check failed" badge), and is listed in the daily secrets
notification. Creating a monitored secret whose host can't be reached
and with no manual date is rejected with the reason rather than saved
blank. A non-TLS port (the likeliest typo) gets a plain-language
error instead of raw OpenSSL output.
Server-side connections to a user-supplied host:port need the same
operator role that already gates editing secrets (and running Semaphore
templates, which is strictly more powerful); the host is validated
against a strict character set before any connection is made.
New nullable secrets columns (check_host, check_port, last_checked_at,
last_check_error) via migration 0007; existing rows are unaffected.
Verified against real TLS servers (openssl-generated certs) and the
real secrets router with a stubbed session: a live 45-day cert read
back as the correct date via both an IP host (no SNI) and a hostname;
an already-expired cert reported its past date and shows as expired;
refused connections, a server that accepts but never answers (times
out), and a plain non-TLS server each produced a descriptive error
rather than a hang or crash. Through the router: create with a host
and no date reads the date; unreachable host with no date -> 400 with
the reason; unreachable with a manual date -> saved with the error
recorded; host on a non-certificate type and an invalid host string
-> 400; a hand-typed date on a monitored secret is ignored; changing
the host re-checks immediately; changing the type away from
certificate ends monitoring; a viewer gets 403 on Check now. 23 checks,
all passing (a first re-run showed 2 spurious failures that were leftover
rows from the previous run's scratch database, confirmed by a clean re-run).
Real dev database mtime untouched throughout.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Each adapter performs write actions, not just reads (start/stop a
guest, edit a DNS record, rerun a CI job, etc.), so a read-only
credential silently works for the dashboard views but fails the
moment you use an action. Lists the exact endpoints/permissions
needed per target system, drawn from each adapter's own auth code and
header comments (e.g. Proxmox's Sys.Audit/Datastore.Audit split,
Azure's DNS Zone Contributor role, Loopia/Pi-hole/cPanel having no
scoped-credential option at all).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
DNS and Secrets each rendered as three separate stat-tile cards plus a
fourth full-width breakdown card -- visually a different family from
the single self-contained card each integration widget uses (label +
status badge, a compact stat row, then its own breakdown bar inline).
Extracted the integration widgets' card shell into a shared WidgetCard
component (label + badge header, children below) and rebuilt DNS and
Secrets on top of it instead of StatTile/TypeBreakdown, which are now
unused and removed. Also refactored the Integrations map itself onto
WidgetCard, so all eight dashboard widgets (DNS, Secrets, and the six
integrations) are now literally the same component, not just visually
similar.
DNS and Secrets get a badge like the integrations' Connected/Not
connected -- "Not configured" (gray) when nothing's set up yet, or a
status summary otherwise ("X enabled" for DNS; "All OK" / "X expiring"
/ "X expired" for Secrets, mirroring how a failing integration shows a
warning-colored count instead of a plain badge). They now sit under
one "Overview" heading in a shared row instead of two separate
sections, matching how the integration widgets already share one
row-cards grid.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Pagination just landed hardcoded to 20 rows everywhere; add a way to
change that instead of leaving it fixed for every table in the app.
pageSize joins dateFormat/timeFormat on the existing display settings
object (server-side default 20, 5-500 range enforced by the PUT
schema) rather than becoming its own settings section, since it's the
same kind of thing -- an admin-configured, globally-applied display
preference read by every signed-in role via the already-public GET
/api/settings/display endpoint, same as the date/time format already
works.
Renamed the "Date & Time" settings tab/page/route to "Display" (still
just one component, now covering both date/time format and table
pagination) since its scope no longer matches the old name -- kept
Settings.tsx's usual pattern of one page per concern rather than
adding a second, oddly-scoped tab just for one number field.
New web/src/utils/pageSize.ts mirrors utils/date.ts's existing
module-level "set once at startup, read anywhere without prop-
drilling" pattern; usePagination()'s pageSize parameter now defaults
to getPageSize() instead of a literal 20, evaluated fresh on every
call so it picks up a saved change without touching any of the ten
pages already using the hook.
Verified server-side against a temp SQLite DB: pageSize defaults to
20, a partial update sets it without disturbing dateFormat/timeFormat
and vice versa, and it persists across a fresh settings read. Also
checked the default-parameter mechanics directly (re-evaluates the
global value on every call rather than capturing it once, and an
explicit override still wins).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sorting/CSV export were already app-wide; tables with a real chance of
growing into dozens or hundreds of rows (a busy tailnet, a big DNS
zone, a homelab's full IP inventory, a Gitea org with many repos, ...)
had no pagination at all, making them a long unbroken scroll.
New usePagination hook (client-side slicing over an already-sorted/
filtered array, 20 rows per page) and a matching Pagination component
(Prev/Next + "Page X of Y (N total)", hidden entirely when everything
fits on one page). The current page is clamped to the valid range on
every render rather than reset via an effect, so switching to a
smaller data set (a different selected integration, a filter that
narrows the result) can never strand the view on a now-nonexistent
page -- no per-page "reset on change" wiring needed anywhere.
Applied to Audit Log, DNS zones and records, IP Addresses, Secrets,
Servers (manage table), Docker containers, Proxmox guests, Semaphore
templates, Gitea repos, and Tailscale devices. CSV export keeps
exporting the full sorted/filtered array regardless of which page is
currently shown -- pagination only affects what's rendered on screen.
Left the already-small tables (Synology volumes/disks, Users,
Integrations, per-node Proxmox storage) unpaginated, and left the
Diagnostic Log's existing server-driven pagination as-is rather than
bolting a second, different pagination scheme onto it.
Verified the clamping logic directly: a normal page, the trailing
partial page, a requested page beyond the end (clamps to the last
valid page instead of rendering empty), and an empty result set
(clamps to page 0 with a page count of 1 instead of a negative range).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The agent path showed a per-mount usage table; a Proxmox-linked server
only ever got a single "Disk (allocated)" figure, since that's all the
VM/LXC config alone can tell you -- it's the attached disk's declared
size, not how full it actually is inside the guest. Fix the actual gap
instead of just matching the display: fetch real usage where Proxmox
can see it.
adapter.getGuestDetail() gained a `disks` field (same {mount,
sizeBytes, usedBytes} shape the agent already reports, so the frontend
renders both identically):
- LXC: the host can read straight into the container's root
filesystem, no agent needed -- status/current's disk/maxdisk fields
are real usage, not just allocation.
- QEMU: the hypervisor can't see inside a virtual disk at all without
help, so this calls the QEMU guest agent's get-fsinfo command (same
"gracefully degrade if the agent's missing/older" tolerance already
used for its IP-address lookup, and independent of it -- one
command failing doesn't take out the other). Pseudo-filesystems
(tmpfs, etc.) are filtered out by checking for a non-empty backing
`disk` array, the common convention for this endpoint.
Extracted the disks-table JSX (previously only in the agent branch)
into a shared DisksTable component and used it in both branches, and
added a note explaining an empty result when a running QEMU VM's
guest agent doesn't support get-fsinfo (an older agent version).
Verified against a mock Proxmox API over real TLS: an LXC's root
usage, a QEMU VM's real fsinfo mounts (with the disk-less tmpfs entry
correctly filtered), and a QEMU VM whose get-fsinfo fails outright --
confirming that degrades to an empty disks list without throwing and
without affecting the separate network-get-interfaces result.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Not every registered server is a Proxmox VM/LXC -- bare-metal boxes
and other hosts had no reason to show a "link to Proxmox" option, but
it appeared unconditionally on every server's detail page.
New hideProxmoxLink column on servers (default false, so existing
behavior is unchanged until someone opts in). A "Not a VM? Hide this"
link in the card's header sets it; once hidden, a small "+ Show
Proxmox link options" link takes its place so it's still reachable,
not buried in a settings form. The card always shows regardless of
this flag once a server IS actually linked, so unlinking never becomes
unreachable by hiding the card out from under an active link.
Verified against a temp SQLite DB with real migrations: a new server
defaults to false, and toggling true/false both persist correctly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Dockhand already tracks per-container image updates internally (its
UI shows this) via a cached check plus an on-demand recheck -- surface
that here instead of only showing running/stopped state.
adapter.listContainers() now also reads GET /api/containers/pending-
updates per environment (a cached read, no registry hit) and merges
each container's hasImageUpdate/newerVersion/checkedAt onto it.
updateAvailable is a tri-state: true (update pending), false (checked,
up to date), or null (never checked) -- distinguishing "no update"
from "we don't know yet" matters since a container can sit unchecked
indefinitely until someone triggers a check.
New adapter.checkForUpdates() triggers a fresh check across every
environment (POST /api/containers/check-updates, one registry lookup
per container so this can take a while) and new POST /:id/dockhand/
check-updates route (operator+, matching the existing container
action's role gating).
Docker page: new "Update" column (badge + tooltip with the newer
version), an "updates available" count next to the running/total
count, and a "Check for updates" button. Dashboard's Dockhand widget
also gained an "Updates" mini-stat, swapped in for the less useful
"Not running" figure (already inferable from running/total).
Field names (hasImageUpdate, newerVersion, checkedAt, the check-
updates response shape) came from Dockhand's own published OpenAPI
spec, not guessed -- and verified against a local mock Dockhand server
covering all three update states (pending, up to date, never checked)
plus the check-updates aggregation across environments, since this
sandbox can't reach the user's real Dockhand instance.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Flagged back when Proxmox/Synology got their own pages: once all six
integration types had moved to dedicated pages, the "browsing"
dropdown here only ever showed one of six near-identical "go manage
this on its own page" redirects, and its final fallback branch was
dead code (every IntegrationType was already covered). Now that all
the actual data views have moved out, drop that mode entirely.
The page is now just what "Manage integrations" already was: a
sortable, CSV-exportable table of configured integrations (name/type/
status), visible to every role since none of that is sensitive, with
an admin-only "Add integration" button and per-row edit/enable-disable/
delete actions. No more mode toggle, no provider-picker dropdown, no
per-type redirect cards -- the six dedicated pages (and the sidebar
nav that already points at them) are how you actually use each
integration now.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The six integration cards only ever showed one headline number each
(e.g. "3/5 devices online") -- thin compared to the new DNS/Secrets
sections' stat-row-plus-breakdown layout. Each widget now shows 2-3
mini stats plus a compact breakdown bar, using data these endpoints
already return (so no new API calls except one extra Synology call for
CPU/RAM, matching what the Synology page itself already fetches):
- Tailscale: online/unauthorized/expiring-soon stats + devices by OS
- Proxmox: running/total + VM vs LXC counts + guests by node
- Dockhand: running/total + host count + containers by state
- Semaphore: template/failing counts + templates by last-run status
- Gitea: repo/private/failing counts + repos by last-run status
- Synology: volumes/unhealthy/CPU load + RAM used-of-total + disks by
health status
Extracted the breakdown-bar rendering out of the DNS/Secrets-only
TypeBreakdown into a bare BreakdownBar (no card wrapper, optional
compact sizing) so it drops into each integration card directly, and
factored the per-item counting into a small breakdownFrom() helper
used by all six. Status/state colors reuse the same palette each
integration's own dedicated page already uses for its badges (Docker's
container-state colors, Semaphore/Gitea's run-status colors); open-
ended categories (OS names, Proxmox node names) get a generated
palette instead since there's no fixed enum to hardcode against.
Widget cards moved from a 3-column to a 2-column grid to fit the
extra content without feeling cramped.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sloth Manager's dashboard showed per-system stats for DNS and Secrets
(domains/records by type, expiry counts by type) that this app's
Dashboard never had -- it only ever showed the six live integration
widgets. Port those two sections over, generalized to sit alongside
the integrations this app added that Sloth Manager never had.
New server/src/dns/stats.ts (getDnsStats(), used by new GET
/api/dns/stats): zone counts are fetched live per enabled provider
(matching what the DNS page itself shows, since the zone cache table
only gets a row once a zone has been synced and would undercount) --
one unreachable provider surfaces its error without blanking the rest.
Record-type breakdown comes from the local cache instead, since
records are only ever shown from cache elsewhere in this app too
(fetching every zone's records live on every dashboard load would be
far more expensive for no real accuracy gain).
Dashboard.tsx gained three sections under clear headers: DNS (domain/
provider/record stat tiles + a "records by type" breakdown), Secrets
(monitored/expiring/expired tiles + a "secrets by type" breakdown,
computed client-side from the already-fetched secrets list, same as
Sloth Manager did), and the existing Integrations widgets grouped
under their own header for visual parity with the two new sections.
Breakdowns render as a stacked proportion bar with a color-coded
legend rather than a pie chart -- this app has no charting library
anywhere yet, and a plain CSS progress bar (an idiom Tabler itself
uses) gets the same "see the mix at a glance" value without adding one
for a single dashboard.
Verified getDnsStats() against a temp SQLite DB with real migrations
and an enabled + a disabled DNS provider: enabled/total provider
counts, live zone counting, and cached-record-type aggregation (sorted
by count) all came back correct, and the disabled provider was
correctly excluded from the zone count while still counting toward
the total.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Each server's detail page gets an "Admin Links" card for bookmarking
that host's own web UIs -- container managers, Webmin, Cockpit, or
anything else reachable by URL -- so there's a quick way to jump
there without hunting down the address each time.
New server_links table (serverId FK, label, url, ON DELETE CASCADE so
removing a server cleans up its links automatically) and three new
routes: POST/PATCH/DELETE /api/servers/:id/links, gated to operator+
like the rest of this page's editing actions; the existing GET
/:id/detail now includes the server's links alongside hardware/DNS
info. URLs are validated to start with http:// or https:// server-side
(rejecting e.g. a javascript: URL that would otherwise render as a
clickable link).
Rendered as a row of pill buttons (label opens the URL in a new tab),
with inline edit/remove controls next to each when the viewer can
edit, and an "Add link" form matching the existing task-form style on
the same page.
Verified against a temp SQLite DB with real migrations: create/list/
update/delete all work, and deleting the parent server cascades to
remove its links rather than leaving them orphaned.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Applies the useSortable/SortableTh infrastructure introduced in the
Diagnostic Log commit to the rest of the app's tables: Audit Log,
Users, Servers (manage table), server task tables (grouped by
schedule type, on both the all-servers and per-server views), DNS
providers/zones/records, Docker containers, Gitea repos, Integrations
(manage table), IP Addresses, Proxmox guests, Secrets (already had
CSV, gained sorting), Semaphore templates, Synology volumes/disks, and
Tailscale devices.
Every column header is now click-to-sort (again on the raw field, not
its formatted display -- a byte count sorts numerically even though
the cell shows "1.2 GB", a date sorts chronologically even though the
cell shows "5d 15h 2m"), and every table got an "Export CSV" button
next to its Refresh button, exporting whatever's currently
sorted/filtered via the existing downloadCsv util.
Grouped tables (ServerTaskTable renders one sub-table per schedule
type) can't call the useSortable hook per group without breaking the
Rules of Hooks, so extracted its comparison core as a standalone
sortItems() function, driven by one shared sort-state pair at the
component's top level and applied per group.
Deliberately left two small (1-6 row) tables embedded inside stat
cards unsorted -- Proxmox's per-node storage list and ServerDetail's
agent-reported disk list -- since they're secondary detail inside an
already-scannable card, not primary list content; happy to add if
useful in practice.
Verified the shared sort core (sortItems) directly: numeric-aware
string compare (so "item2" sorts before "item10"), numbers, booleans,
and that null/undefined always sort to the end regardless of
direction.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sloth Manager tracked every API call made to DNS providers for
connectivity troubleshooting. Port that here, generalized to cover
every outbound integration this app makes, not just DNS -- Tailscale,
Proxmox, Synology, Semaphore, Gitea, and Dockhand calls now show up
too, since a broken API token or unreachable host on any of them is
just as worth diagnosing.
New services/diagLog.ts: a generic withDiagLogging(source, adapter)
wraps every async method of any adapter object with timing +
success/failure recording, without touching a single adapter's
request/error-handling internals -- every DNS and integration adapter
interface here is already just a flat set of async methods, so this
one wrapper works for all twelve of them. Applied it at each adapter
factory's own return statement (one line each) rather than at the
route layer, so background jobs that construct adapters directly
(the Tailscale key-expiry scheduler, IPAM sync, agent-driven Proxmox
lookups) get logged too, not just requests through routes/integrations.ts.
New diag_log table (ring-buffered to the last 500 rows, mirroring
Sloth Manager's approach -- this is for live troubleshooting, not a
durable record) and admin-only GET/DELETE /api/diag-log routes, source/
result filters, pagination.
New admin-only Diagnostic Log page: filterable, paginated table with a
Clear button. Also introduces the shared useSortable hook + SortableTh
component used here for the first time -- a follow-up commit applies
the same sorting (and CSV export) to the rest of the app's tables, per
the same request.
Verified end-to-end against a temp SQLite DB with real migrations: a
fake wrapped adapter's successful and failing calls both land correctly
in the log with the right source/operation/latency/error, and the
source/ok filters and clear-log operation all behave correctly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The Proxmox page only showed guest (VM/LXC) status; there was no view
of the underlying host(s) themselves -- uptime, CPU/RAM usage, or how
full each storage pool is.
New adapter.listNodeStats() calls GET /nodes/{node}/status and GET
/nodes/{node}/storage for every online node (a node going down
shouldn't blank the whole page, same tolerance as listGuests()) and
returns uptime, CPU usage % + core count + load average, RAM/swap
usage, and per-storage (local, LVM-thin, ZFS, NFS, ...) used/total.
New GET /:id/proxmox/nodes route alongside the existing /guests one.
One card per online node now sits above the VM/LXC table, showing
these stats plus a small storage table with usage bars (reusing the
same green/yellow/red thresholds as the Servers detail page and the
Synology page). Handles a multi-node cluster by rendering one card
per node, and storages with no active/total data (e.g. an offline NFS
mount) render as an "Inactive" badge with blank usage instead of
throwing.
Verified end-to-end against a local mock Proxmox API server over real
TLS (a throwaway self-signed cert, matching how the adapter's
`insecure` option is meant to be used) since this sandbox can't reach
the user's actual Proxmox cluster -- confirmed CPU fraction-to-percent
conversion, load average parsing, byte fields, and that a
null-valued/inactive storage doesn't break parsing.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The Tailscale page already showed per-device key expiry; extend the
existing daily-reminder infrastructure (currently only for Secrets)
to push it out through the configured notification channels too, the
same way expiring secrets already are.
New tailscaleKeyExpiryScheduler.ts mirrors secretExpiryScheduler.ts:
runs once at startup (skipped if already run today) and daily
thereafter, checking every enabled Tailscale integration's devices for
keys expiring within the warning window and calling notify() with the
results. Reuses the exact same daily time/timezone setting as the
secret-expiry check (one "Daily reminder time" control, two
independent on/off toggles) rather than adding a second schedule for
users to configure.
Centralized the expiring-soon threshold and check (previously only
duplicated in the /synology and /tailscale route summaries) into
adapter.ts as `KEY_EXPIRY_WARN_DAYS` / `isKeyExpiringSoon()`, and
updated the devices route to use it instead of its own inline copy.
New `tailscaleKeyCheck` notification-event toggle (default on) in
settings, alongside the existing secret-expiry one.
Verified end-to-end against a temp SQLite DB + real migrations: a
tailscale integration pointed at a mock Tailscale API (one device
expiring in 10 days, one with key-expiry disabled) with the webhook
channel enabled and pointed at a mock receiver — confirmed the
scheduler's startup check queries the DB correctly, decrypts the
integration's credential, calls the adapter, filters out the
disabled-expiry device, and delivers a webhook payload naming only
the expiring device.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Scheduled-task browsing/editing already lives on each server's detail
page (ServerDetail.tsx), so the overview page's filters, task-group
listing, and add/edit-task form were pure duplication -- the only
things it needs to do are list servers (as clickable cards) and let
an admin register new ones / issue or rotate agent tokens.
Renamed ServersTasks.tsx -> Servers.tsx to match its new, narrower
scope; updated the nav label, route, and ServerDetail's back-link
text from "Servers & Tasks" to "Servers" accordingly. No server-side
or API changes -- this only removes now-redundant frontend code.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Same move as Docker/Tailscale/Semaphore/Gitea: each was buried inside
the generic Integrations browsing view. Give them dedicated pages too
— this was the last pair, so every integration type with a browsing UI
now has its own page.
New web/src/pages/{Proxmox,Synology}.tsx reuse the existing, unchanged
API routes and adapters (no server changes) with their own integration
picker scoped to just that type. Added matching nav items and routes.
Synology stays read-only, matching its original design.
Removed the corresponding state/handlers/tables from Integrations.tsx
(296 lines). It's now down to just the "Manage integrations" CRUD
table plus a browsing dropdown whose six branches are all redirects to
the type's own page — genuinely pointless now that every type has
moved out, worth simplifying separately.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Same move as Docker: each was buried inside the generic Integrations
browsing view, sharing one "pick any integration" dropdown across six
unrelated types. Give them dedicated pages instead.
New web/src/pages/{Tailscale,Semaphore,Gitea}.tsx each reuse the
existing, unchanged API routes and adapters (no server changes) with
their own integration picker scoped to just that type. Added matching
nav items (Tailscale, Semaphore, Gitea, right after Docker) and routes.
Removed the corresponding state/handlers/tables from Integrations.tsx
entirely (374 lines). Adding, editing, enabling/disabling, and
deleting each credential itself still happens under Integrations ->
Manage integrations, same as Proxmox and Synology, which stay on the
Integrations page. Selecting one of these three types in Integrations'
generic browsing dropdown now points to its new page instead of
rendering a table there too.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Container management was buried inside the generic Integrations
browsing view, mixed in with five unrelated integration types behind
a single "pick any integration" dropdown. Give it a dedicated page,
matching how Servers & Tasks and DNS already get their own top-level
spot instead of living inside Integrations.
New web/src/pages/Docker.tsx reuses the existing, unchanged Dockhand
API routes and adapter (no server changes) — it just has its own
integration picker scoped to Dockhand only, rather than sharing
Integrations' any-type dropdown. Added a Docker nav item (right after
Integrations) and route.
Removed the Dockhand-specific state/handlers/table from Integrations
entirely; adding, editing, enabling/disabling, and deleting the
Dockhand credential itself still happens under Integrations -> Manage
integrations like every other integration. Selecting a Dockhand row in
Integrations' generic browsing dropdown now points to the Docker page
instead of rendering a container table there too.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Extends the Tailscale IPAM sync to Proxmox: pulls every VM/LXC's IP
address(es) across all enabled Proxmox integrations into the
inventory, reusing the same never-overwrite-a-manual-entry semantics.
Proxmox guests can have multiple NICs (each IP synced as its own
entry) and, for QEMU VMs, IP discovery depends on a responsive guest
agent — a VM with none simply contributes zero entries rather than
erroring, matching getGuestDetail()'s existing best-effort behavior.
listGuests() doesn't include IPs, so this fetches getGuestDetail() per
guest; acceptable for an explicit, user-triggered sync rather than a
background poll.
Extracted the shared "insert, update only if I own this row, else
skip and report" upsert logic (previously inline in the Tailscale sync
handler) into upsertSyncedEntry(), now used by both sync routes so the
core safety rule can't drift between them.
Verified end-to-end against a mock Proxmox server covering: an LXC
with two NICs (both synced), a QEMU VM with a responsive guest agent,
a QEMU VM with no agent (zero entries, no error), a manual-entry
collision (skipped and reported, never overwritten), and idempotent
re-sync (add -> update on the second run).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
IPAM was entirely manual — Tailscale device IPs never showed up there
even though the Tailscale integration already lists them. Add an
explicit sync action (matching this app's existing pattern of
user-triggered syncs rather than silent background polling).
- New source column on ipam_entries (null = manual, "tailscale" =
auto-synced) so a re-sync only ever touches rows it created itself —
a manually-entered IP that happens to collide with a tailnet address
is left untouched and reported back as skipped, never overwritten.
- POST /api/ipam/sync-tailscale pulls every enabled Tailscale
integration's device list, upserting by primary IP (label, OS in
notes, vendor "Tailscale"); one unreachable Tailscale integration
doesn't block others.
- New "Sync from Tailscale" button on the IP Addresses page, with a
small "synced" badge marking which rows came from it.
Also fixed a longstanding TODO found in the same file: the "DNS
records" column always showed "-" because matchingDnsRecords was
hardcoded to an empty array from before the DNS module existed. It
now does the same content-based reverse lookup against the DNS
module's record cache used elsewhere in the app.
Verified end-to-end by running the real server with Tailscale's fetch
call intercepted at the process level (its adapter hardcodes
api.tailscale.com with no configurable URL, so it can't be pointed at
a mock server the way Proxmox/Synology can): confirmed add, the
manual-entry skip/never-overwrite behavior, idempotent re-sync
(add -> update), and the DNS-matching fix, all against the real
route and adapter code.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Different tables used different date formats: Servers & Tasks used a
fixed "YYYY-MM-DD HH:mm:ss" (24h), while Audit Log, DNS, Integrations
(Tailscale/Semaphore), and Users called plain .toLocaleString() with
no options, which renders using the browser's own locale — different
per browser/OS, and inconsistent with the other pages' fixed format.
- New Settings -> Date & Time page: pick date order (YYYY-MM-DD,
DD/MM/YYYY, MM/DD/YYYY) and 12h vs 24h clock, with a live preview.
- utils/date.ts's formatDateTime() now reads these settings instead of
being hardcoded to sv-SE/24h; the setting is fetched once at app
startup (alongside /api/me) via a new non-secret GET
/api/settings/display (any signed-in user, same rationale as the
badge-color endpoints) and applied immediately on save too, without
needing a page reload.
- Switched every remaining raw new Date(...).toLocaleString() call
(Audit Log, DNS zone sync time, Tailscale/Semaphore last-seen,
Users' last login) over to the shared formatter, so every table now
renders dates identically.
Verified the formatter's date-order x 12h logic against all six
combinations plus the midnight/noon 12h edge cases, and the settings
endpoints end-to-end (defaults, partial updates, validation
rejection, audit logging) against the real dev server.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Settings → DNS Badges only let you customize DNS provider badge
colors. Expand it into a general Badges page with a second section
for the six integration types (Tailscale, Proxmox, Synology,
Semaphore, Gitea, Dockhand), applied to the type badge in the Manage
integrations table.
- New integrationColors key on AppSettings, stored/merged the same way
as providerColors via the existing settingsStore.
- New GET /api/settings/integration-colors — non-secret, any signed-in
user, mirroring /provider-colors — so the badge color can be read
without needing admin access to the full settings payload.
- Renamed DnsBadgeSettings.tsx -> BadgeSettings.tsx (route
/settings/dns-badges -> /settings/badges, sub-nav label "DNS
Badges" -> "Badges") with both color sections saved together.
Verified end-to-end against the real dev server: PUT persists
integration colors independently of provider colors, the public
integration-colors endpoint reflects updates immediately, and the
change is audit-logged.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The Proxmox start/stop/restart routes already existed (used by the
Integrations page's guest table) but weren't reachable from a server's
own detail page, even when that server was linked to a Proxmox guest.
Reuses the existing POST /api/integrations/:id/proxmox/nodes/:node/
:type/:vmid/{start,stop,restart} routes and the operator+ role gate
already enforced there — no server-side changes needed. Buttons only
render for hardware.source === "proxmox", mirroring the Integrations
page's running/stopped button-set logic, with the same confirm-before-
stop prompt for the non-reversible action.
Verified end-to-end against the real dev server with a mock Proxmox
HTTPS server: linked a server to a mock guest, confirmed the detail
endpoint's live status, then triggered restart/stop/start and
confirmed the mock actually received each action and every call was
audit-logged.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Manage integrations only supported add/toggle/delete — fixing an
expired API token meant deleting and recreating the whole integration.
- New GET /api/integrations/:id/config returns only the non-secret
config fields (never the decrypted secret) so an edit form can
pre-fill URL/tailnet/etc. fields.
- New POST /api/integrations/:id/test merges the stored, decrypted
config with any freshly-typed overrides and pings the real adapter —
lets "Test connection" work during an edit without ever sending the
current secret back to the browser.
- New IntegrationEditForm component: secret fields render blank with a
"leave blank to keep the current value" placeholder; submitting only
sends the fields that were actually filled in, so a name/URL edit
can't accidentally wipe a secret and a secret rotation can't touch
anything else. Reuses the existing PATCH /:id route, which already
merged partial config updates correctly.
Verified end-to-end against the real dev server: confirmed via direct
DB decryption that a non-secret-only edit leaves the stored secret
byte-for-byte unchanged, and that a secret-only edit rotates it without
touching other config; the test route was confirmed to make a real
network call (got a genuine "API token invalid" from Tailscale's API
against a fake key).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Servers & Tasks only tracked scheduled tasks — there was no overview of
the servers themselves and no way to see CPU/RAM/disk/IP info. Add a
clickable server overview (visible to every role, not just admins) that
opens a per-server detail page.
- servers table gains an agent-reported hardware/network snapshot
(IPs, CPU model/cores/load, memory, disks) and an optional link to a
Proxmox VM/LXC (integration + node + guest type + vmid).
- Proxmox adapter gains getGuestDetail(): live cores/memory/disk from
/config, live cpu/mem/uptime from /status/current, and IPs (LXC net
config directly, QEMU via a best-effort guest-agent call that degrades
gracefully when the agent isn't installed).
- New GET /api/servers/:id/detail combines whichever hardware source
applies (live Proxmox vs. last agent report) with a DNS reverse-lookup
against the DNS module's own record cache, so matching hostnames show
up next to each IP. New PATCH /api/servers/:id manages the Proxmox
link.
- agent/linux/report-tasks.sh now also collects and reports IPs,
CPU/memory/disk info on every check-in (load-average-based CPU number,
not instantaneous, to keep the agent a cheap oneshot).
- Extracted the per-server task table into a shared ServerTaskTable
component so the all-servers view and the new detail page render
tasks identically.
Verified server-side end-to-end against the real dev server (agent
report -> detail endpoint -> DNS match) and the Proxmox adapter against
a mock HTTPS server covering LXC/QEMU config parsing and the
guest-agent-unavailable fallback.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The Settings page was a "coming soon" placeholder. Port Sloth Manager's
settings feature set: Gotify/ntfy/SMTP/webhook notification channels
(each with its own test-send button), per-event toggles (DNS record
added/updated/deleted, a daily secret-expiry digest with configurable
time/timezone), and per-provider DNS badge color customization.
Settings persist in the existing `settings` key/value table via a new
settingsStore service; a notify service fans a message out to every
enabled channel. DNS record add/update/delete now fire notifications,
and a node-schedule job re-arms itself whenever the notification
settings change. Removed the now-superseded GOTIFY_URL/GOTIFY_TOKEN
env vars in favor of in-app configuration.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>