447f33fff6e6c919d58b1b5bd3049dd2b66f19e8
36
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
447f33fff6 |
Add an Alerts page under Operations listing everything that's wrong now
One list of the current problems across servers and integrations, instead of waiting for a notification or visiting each page: servers that stopped reporting, full or nearly full disks and volumes (critical from 95%), Synology volume/disk problems, failed or uncovered Proxmox backups, failed Proxmox Backup Server verifications, container image updates, expired or expiring secrets/domains/Tailscale keys, failed Semaphore and Gitea runs, Uptime Kuma monitors that are down, overdue osTicket tickets, and integrations whose calls keep failing. Visible to every role, with severity and kind filters, search, sorting, CSV export and "Check now". It runs the same checks that send the notifications rather than a second copy of them: the detection in the health, automation, Proxmox backup, PBS, Docker update and Tailscale key checks is pulled out into shared collectors that both the schedulers and the page call, so the two can't disagree about what counts as a problem. Notification behaviour is unchanged, including the scheduled backup checks skipping integrations under a maintenance window. Unlike the notifications the page ignores the on/off toggles, and keeps problems under a maintenance window, marked silenced and counted apart. It reads live, so a result is reused for a minute (and Refresh can't re-run everything more than once every ten seconds), and every source has a 20 s limit so one hung integration can't hang the page. Anything it couldn't read is called out at the top instead of looking like all clear, and server checks pause for the same 20 minutes after a restart as the notifications do, with a note saying so. Also gives the newer integrations (PBS, osTicket, Uptime Kuma, phpIPAM) proper names in "integration down" notifications instead of their ids. Verified through the real routes against a scratch database with fake backends (offline and full-disk servers, secrets and domains, a silenced server, a fake PBS with failed verification, a hanging integration, a refused one, a failing-calls streak, caching, the restart grace period, auth), and by rendering the real page against that data in a browser: filters, search, silenced toggle, sorting, Check now, dark mode. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> |
||
|
|
88d9c8e097 |
Record sign-ins, new accounts, automatic log trimming and what settings changed
A check of every write path found gaps in what the audit log captured: - Sign-ins and sign-outs are now recorded with the IP they came from (sign-out is recorded first and can't block signing out). - A new account is recorded when it's created on first sign-in, including when the very first user becomes admin - so the log shows who gained access, not only who changed things. - The automatic log purge, which deletes audit entries, now records itself, attributed to "system". recordAudit() takes an optional actor for this. It only records when something was actually deleted. - Settings updates record what changed (before and after) instead of only which sections were touched. The notification channels (Gotify, ntfy, SMTP, webhook) record field names only: they hold credentials, and webhook URLs and public ntfy topics act as secrets, while the audit log is readable by operators and Settings is admin-only. - Integration edits record renames, enabling/disabling, whether credentials were replaced (never the credentials), and which settings fields changed (names only). The Privacy page and README now say sign-ins store an IP in the audit log. Verified through the real routes against a scratch database: user creation, the logout route, the automatic purge, settings and integration edits - including that a secret token and a webhook URL appear nowhere in the stored entries. The sign-in callback itself needs a real identity provider and wasn't run. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> |
||
|
|
236b1da0dc |
Add a Network > Ports page: agent-reported ports plus manual openings
Summarizes every server's agent-reported listening ports in one cross-server table (grouped by protocol+port, addresses merged, loopback-only flagged) - previously this only existed per-server on each server's own detail page. Adds a second table for ports this app has no way to see on its own: manually-recorded openings on a router, edge firewall, or cloud security group, each with a label, external port/protocol, an optional link to a tracked server (with its own internal port when NAT changes it) or a freeform destination, a free-text source, and a comment. Viewer-readable; adding/editing/deleting needs operator or admin. The agent-port grouping logic (dedupe by protocol+port, detect loopback-only sockets) was shared with the existing per-server Ports card via a new agentPorts.ts service instead of duplicating it. Verified with a real HTTP-level test: a genuine Express app with the actual routers, a scratch SQLite DB, and forged admin/viewer sessions, covering grouping correctness, the server-name join, input validation, and role enforcement - the real dev DB was confirmed untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
df2a5ce42b |
Add a Proxmox Backup Server integration: datastore/snapshot verification status
Proxmox VE already shows whether the last vzdump push to PBS succeeded, but has no visibility into PBS's own backup verification, GC/prune health, or host status. This adds PBS as its own integration (own adapter, page, nav entry, and Dashboard widget) that reads datastore usage and, for every stored snapshot, its verification state directly from PBS. A new daily check (mirroring the existing Proxmox backup-failure check) notifies when a snapshot has failed verification or a datastore couldn't be read, with its own toggle in Settings -> Notifications and its own maintenance-window silencing. Not verified against a live PBS instance — built from PBS's published API docs and a scratch test against a mocked PBS server exercising the adapter's parsing and auth-header format (PBSAPIToken uses a colon separator, unlike PVE's PVEAPIToken which uses =). See INTEGRATIONS.md for details and the "not verified" caveat. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
bf7f73f6b6 |
Import maintenance windows from Uptime Kuma
New "Import from Uptime Kuma" action on the Maintenance page (operator):
pick an Uptime Kuma integration and a duration, and it starts (or
extends) a maintenance window here for every server whose address
matches a monitor Uptime Kuma currently reports as being in maintenance,
reusing the existing monitor-to-server matching from the Uptime Kuma
integration itself.
Uptime Kuma's metrics endpoint only exposes a monitor's *current* status,
not its scheduled start/end time (there's no API for that), so this
deliberately doesn't try to mirror Uptime Kuma's own schedule -- it starts
a window for the duration you choose, the same bounded/required-end
window this feature has always used. Running it again while Kuma is
still in maintenance extends the same window rather than stacking a
second one; when Kuma later shows nothing in maintenance, already-active
windows are left alone rather than force-ended, since ending them isn't
something only Uptime Kuma's state should decide. Two Kuma monitors that
match the same server are deduped to one window. Monitors with no
matching server are reported back by name so nothing is silently missed,
and monitors that aren't in maintenance are ignored entirely.
The manual "Start maintenance" endpoint's start-or-extend logic (dedupe,
pruning old rows, the response shape) is now a shared
services/maintenance.ts function instead of living only in that route
handler, so the import path can't drift from how a manual window behaves.
Likewise the server-matching helper gained a small toMatchableServers()
so the existing Uptime Kuma monitors route and this new one build the
same match input the same way instead of each parsing server rows on
their own.
No schema change -- imported windows are ordinary maintenance windows;
their Uptime-Kuma origin is only in the reason text ("Imported from
Uptime Kuma (<integration>): <monitor>"), visible in the Active table and
the audit log like any other window.
Verified with 28 backend checks (the refactored manual start/extend
flow as a regression check; roles; unknown/wrong-type/disabled
integration; validation; nothing-in-maintenance; matching including
same-server dedup and unmatched monitors; re-running extends rather than
duplicating; windows left alone once Kuma exits maintenance; upstream
failure; audit entries) and by driving the real Maintenance page against
the real routers in a browser: import, re-import (extends), the
nothing-in-maintenance state, and the viewer view (no edit card, active
windows still visible). Real dev database mtime untouched.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
1a2dd19736 |
Add an Uptime Kuma integration: monitor status and which server each one watches
New integration, following the existing pattern: config in-app (URL + API key, credentials encrypted at rest), its own Uptime Kuma page, an Integrations list entry, a Dashboard widget, and diagnostic-log/ integration-down-alert coverage for free via the shared withDiagLogging wrapper. Read-only -- no start/stop equivalent exists for a monitor. Uptime Kuma has no conventional REST API (the dashboard talks to it over Socket.IO); researched before writing any code, since guessing wrong here would have cost real time. The one machine-readable, authenticated endpoint that lists every monitor is its Prometheus exporter at GET /metrics, gated by HTTP Basic auth -- an API key as the password with the username left blank on current installs, or the real dashboard login on installs from before the API-key feature existed. This adapter authenticates the same way and parses that endpoint's text-exposition format itself (metrics: monitor_status, monitor_response_time, monitor_cert_days_remaining, monitor_uptime_ratio; labels: monitor_id, monitor_name, monitor_type, monitor_url, monitor_hostname, monitor_port), verified against the documented metric/label set and the actual upstream source (server/prometheus.js). A malformed line is skipped rather than failing the whole scrape. "What server is being monitored for what": each monitor's target (an IP for TCP checks, or the hostname out of the URL for HTTP/keyword checks) is matched against your servers' own IPs and hostnames -- reusing the same kind of match already used in the consistency report -- and linked to that server's page. Monitors with no single network target (groups, push monitors, DNS/keyword checks with a complex URL) are left unmatched rather than guessed at. Uptime Kuma's tags aren't read, since the Prometheus endpoint doesn't reliably distinguish a tag label from any other label it might add later. The username field is the first genuinely optional integration config field this app has had; IntegrationField gained an `optional` flag (server validation and both the add/edit web forms honor it) rather than special-casing Uptime Kuma. Verified with 48 backend checks (Prometheus text parsing including escaped quotes, decimals, negative numbers, and malformed lines; TCP vs. HTTP target/port extraction; every documented status code; server matching by IP, hostname, and short name, including no-match cases; the route's real HTTP round trip against a fake Uptime Kuma server, wrong credentials, upstream failures, roles, wrong/disabled/missing integration, diagnostic-log entries; the optional-field validation rule) and by driving the real page and the real Dashboard widget in a browser against the real routers, including CSV export and column sorting. Real dev database mtime untouched. Not verified: a real Uptime Kuma instance. Everything here was checked against Uptime Kuma's documented metric format, its actual upstream source, and a fake server built to match both -- not against a live installation. If your instance's /metrics output differs from what's documented (older version, unusual monitor types), the parser should degrade to an empty or partial monitor list rather than error, but that degradation itself hasn't been observed against the real thing. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
91796a3c3a |
Add a Servers widget to the dashboard
A Servers card in the dashboard's Overview row, in the same shape as the DNS and Secrets cards: a status badge (N offline / N disks nearly full / All online / No agent data / Not configured), total, online and offline counts, an online/offline/no-data bar, and the details that matter -- which servers are offline and for how long (the first three, linking to their pages), which disks are at or above the usage threshold, and a Linux/Windows split when there are Windows machines. The numbers come from a new GET /api/servers/summary, computed on the server with the health monitor's own evaluateHealth. That means "offline" and "disk full" are decided by exactly the rules the alerts use, so the widget can't say a server is fine while an alert says it isn't, and it follows the thresholds set in Settings, which viewers can't read themselves and so couldn't have applied client-side. Servers under a maintenance window are marked as such. A server that has never reported (no agent, or tracked only through Proxmox) counts as "no data" rather than offline, an offline server's stale disk figure isn't reported, and a garbled report doesn't blank the widget. Readable by every signed-in user, like the dashboard itself. Verified with 23 backend checks (offline/online/no-data classification including the bare SQLite timestamp format, ordering, colons in Windows mounts, threshold changes, maintenance flags, agreement with the alert path's own offline and disk sets, empty and garbled inputs, and the route as a viewer following Settings) and in a browser against the real router across five states: problems, disks only, all healthy, no data, and empty. Real dev database mtime untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
fea20456e4 |
Add a Windows agent (PowerShell)
Reports a Windows machine the way the Linux agent does, replacing the
"planned" stub in agent/windows: scheduled tasks plus hostname, IPv4
addresses, CPU model/cores/current load, memory, every fixed disk, and
TCP/UDP listening ports with the owning process (which feed the Ports
card, localhost-only listeners included).
Scripts (plain ASCII by design -- they are downloaded as text and Windows
PowerShell 5.1 reads BOM-less files as ANSI):
- report-tasks.ps1: collects and POSTs to /api/agent/report. Works in
Windows PowerShell 5.1 and PowerShell 7. -DryRun prints the JSON.
Microsoft's own \Microsoft\ tasks (hundreds) are left out unless
INCLUDE_MICROSOFT_TASKS is set. Triggers are turned into readable text
("Weekly on Mon, Wed at 03:00", "At logon", "..., repeating every 15 min").
Self-signed certificates work via API_INSECURE on both PowerShell
versions (they need different mechanisms).
- install.ps1: elevated only; downloads the agent to ProgramData, writes
agent.json with permissions locked to SYSTEM and Administrators *before*
the token goes in, and registers a SYSTEM scheduled task (every 15 min
plus at startup with a 2 min delay). Reinstalling replaces the task.
- uninstall.ps1: removes the task and only the files the agent installed.
Server: accepts schedule_type "windows_task"; a server can be registered
as Windows (Add a server has an operating system choice); an agent's
reported os_type ("linux"/"windows", anything else ignored) corrects the
stored one. The Servers page shows the right install and uninstall
command for each OS (Windows PowerShell 5.1 one-liners, with a self-signed
variant and a note about PowerShell 7), and Windows tasks are labelled
"Windows scheduled tasks". The Linux commands are unchanged.
Verified on this Windows machine, in both PowerShell 5.1 and 7:
- Real dry runs found and fixed bugs before anything shipped: tasks and
ports came out as one nested item (return , $out wrapped twice), integer
keys in an ordered dictionary index by position (wrong weekday names),
and generic "Trigger" labels.
- End to end against the real agent-report router: HTTP, self-signed HTTPS
refused by default and accepted with API_INSECURE, wrong token gives a
clear one-line error and exit 1, and Swedish letters plus a euro sign
survive JSON -> UTF-8 -> HTTP -> SQLite.
- 35 checks on trigger/action/duration descriptions, 20 on the installer's
building blocks (task parts built but not registered, credentials file
content and ACL, download over HTTP and self-signed HTTPS), 18 on the
server rules, and the generated one-liners run through PowerShell's
parser. The documented one-liners were run through iex and stop at the
administrator check without changing anything.
- Found that PowerShell 7 ignores the ServicePointManager certificate
override, so the installer's own download now uses -SkipCertificateCheck
there.
NOT verified: the elevated install itself. Registering a SYSTEM scheduled
task needs elevation and changes the machine, so it was not run: the task
registration, that the repeating trigger really runs indefinitely, and
the agent running as SYSTEM under Task Scheduler have not been exercised.
Windows 10 / Server 2016 or newer is assumed; older is untested.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
ae64cb345c |
Manage server tags in Settings: pre-add tags, recolour, rename, delete
New Settings > Tags tab (admin) listing every tag -- ones servers use and ones added ahead of time -- with how many servers carry each: - Add a tag before anything uses it, optionally with a colour. Such tags are offered as one-click "Add:" chips (and datalist suggestions) when tagging a server, so the same word gets spelled the same way everywhere. - Give any tag a colour of your choosing, in use or not, or reset it to the automatic one. Changes are drafted with Save/Cancel rather than saved as the picker drags. Colours show on the Servers page, its tag filter bar, the detail page and the editor, and update everywhere without a reload through one shared cached colour map. - Rename a tag; every server that has it is rewritten. Renaming to a name that already exists merges the two after a confirmation naming what will happen; the target keeps its own colour unless it had none, and a server carrying both ends up with one. - Delete a tag, which removes it from every server that has it, with a confirmation stating how many. Tags still live on the servers (servers.tags); a new tag_definitions table holds only what a server can't: existence before use, and a colour. A defined tag stays listed until an admin deletes it, even with no servers. Rename and delete change the servers and the catalogue in one transaction so they can't disagree. Only servers that actually carry the tag are rewritten and counted -- an earlier draft also counted servers whose tags merely weren't in sorted order, which the tests caught. Reading the list and colours is open to everyone signed in (needed to draw tags anywhere); changing the catalogue is admin-only, while tagging a server stays an operator action. Names go through the same normalisation as before, colours must be #rrggbb, and every change is audit-logged. New table tag_definitions (migration 0013). Verified with 44 backend checks (list/counts, roles, create/adopt/ duplicate/rejects, colour set/reset, rename incl. defined, undefined and unused tags, merge colour rules and both-sides servers, delete, audit, and that normal tagging still works afterwards) and in a browser against the real routers: add with colour, set/reset a colour and see it change on the Servers page live, merge with confirmation, delete, and the error path. Not clicked through: the quick-add chips inside the tag editor on a server's detail page (typechecked; same colour code as the rest). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
3f2b5da7be |
Let the consistency report exclude address ranges
Docker reuses the same subnet on many hosts, and those networks aren't part of the LAN, so they show up as conflicts and unlisted addresses. The report only knew about 172.16.0.0/12 through a hardcoded rule; Docker can just as well pick 192.168.x or 10.x. The Consistency page now has an Excluded ranges card: CIDR ranges (IPv4 or IPv6) and single addresses, added with a form and removed with one click, shown to everyone and editable by operators. There's also an "Exclude range" button on each finding that pre-fills a /24 (or /64) around its address to edit. Exclusions are applied to servers, IPAM and DNS before anything is compared, so an excluded address never appears in any kind of finding, whichever source it came from, and the card says how many addresses are currently being hidden so it's clear the filter is doing something. The old hardcoded rule becomes a visible default (172.16.0.0/12) that can be removed -- it was silently wrong for anyone using 172.16/12 as a real LAN. That default is also slightly stronger than before: an address in the range is now left out even if it is in IPAM or DNS, where the old rule only skipped it when nothing else mentioned it. Remove or narrow it if that isn't wanted. Ranges are validated and normalised on the server (both families, prefix bounds, no /0, at most 50), a bad one is rejected with a message naming it and nothing is saved, and changes are audit-logged with before/after. Matching uses Node's BlockList. Stored as a settings value; managed from the report rather than admin-only Settings, like ignoring a finding. Verified with 44 checks (range parsing and rejection, boundary addresses just inside and outside a range, IPv6, single addresses, exclusion across all sources and finding kinds, the hidden-address count, route validation/roles/audit) and in a browser against the real router: add, invalid, remove the default, exclude from a finding. Real dev database mtime untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
2352689fd3 |
Add a consistency report across IPAM, DNS and servers
New Consistency page listing where the three places this app records what lives at an address disagree: - Address conflicts: the same address reported by more than one server. - DNS out of date: a record named after a server (its hostname, or its short name) that points at an address the server doesn't report. - IPAM out of date: an entry labelled with a server's name at an address the server doesn't report. - Not in IPAM: addresses a server reports or DNS points at that IPAM doesn't list, merged into one finding per address, with a one-click "Add to IPAM" that pre-fills a label. - No DNS record: server LAN addresses no cached A/AAAA record resolves to. It compares data the app already holds and fetches nothing when opened, so the page states how many servers had reported addresses and how many DNS zones are synced (and how old the oldest sync is) -- DNS records are only cached for zones that have been synced, and a report that silently treated missing data as "no records" would mislead. Rules chosen to keep it from crying wolf: - Only private addresses are compared; public DNS records aren't expected to be in IPAM. - Servers with no reported addresses are never judged. - Agents report IPv4 only, so records are only compared within an address family (an AAAA record isn't "stale" for lacking an IPv6 address). - Docker bridge networks (172.16/12) are ignored: shared ones aren't conflicts, and they aren't listed unless someone put them in DNS. - Tailscale addresses don't need DNS records (MagicDNS), and IPAM entries kept current by the Tailscale/Proxmox syncs aren't second-guessed. Findings anyone has decided are fine can be ignored (operators) with a reason. An ignore is keyed on the finding's stable identity so it stays ignored across runs, its stored text comes from the finding rather than the request, and it is marked "no longer occurring" once the condition goes away. Ignore/restore are audit-logged. New table consistency_ignores (migration 0012). portScan's private-address helper is now exported and shared. Verified with 36 checks (each rule and its exclusions, address-family and case/trailing-dot handling, IPv6 case, ordering, stable keys, the report route including a garbled agent report, source counts, ignore/unignore rules and audit entries) and by driving the page against the real routers in a browser: Add to IPAM actually created the entry, ignore and restore, severity filter, "show all", the viewer view, and narrow-width layout (which found and fixed a squeezed badge and clipped buttons). Real dev database mtime untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ca0fa817f8 |
Track domain registration expiry, with daily reminders
New Domains page listing when each domain registration expires, read from the registry. Domains behind the DNS zones already synced are picked up automatically; others can be added by hand. You're reminded daily from N days before expiry (Settings > Notifications, default 30) until it's renewed, and told when an expiry date hasn't been refreshable for several days so a stale date isn't trusted silently. RDAP alone would not have covered this homelab: .se, .nu, .io, .eu and .de are not in IANA's RDAP bootstrap. Lookups therefore try RDAP where the TLD publishes a server and fall back to WHOIS on port 43, found via IANA's own referral, parsing the expiry line out of the free-text answer. Only the expiry date and registrar are read or stored. Verified live against the real registries: .se and .nu via WHOIS, .com/.org/.dev via RDAP. Behaviour worth knowing: - A DNS zone that is a subdomain (lab.example.se) resolves to the registration that actually expires by trying the name and then its parents, so no public-suffix list is needed. Zones already covered by a tracked domain are not looked up again. - "Couldn't ask" is never confused with "not registered": network errors, rate limits and garbled answers are errors, and a transient error at any level stops the walk from concluding the domain doesn't exist. - A failed refresh keeps the last known expiry and records why, rather than blanking a date that's still relied on. - Zones that don't resolve to a real registration (.lan, .local, unregistered names) simply get no row. Zone-derived rows disappear when their zone does; manual rows stay. Zone-derived rows can't be deleted by hand. - Registries that don't publish an expiry (.de, .eu) are tracked with a note instead of a date. - Input like "example.com/path" is refused rather than silently reduced to its host. - Runs on the daily secret-expiry schedule and reminder time, on demand (Check all now / per domain), and once at startup if nothing has been read in a day. Never blocks startup, one lookup at a time with a pause. The warning window lives with the other thresholds in settings. New table domains (migration 0011); two settings fields (toggle and warning days). Verified with 90 checks against fake RDAP/WHOIS backends (name normalization, date formats, WHOIS parsing including rate-limit and no-expiry answers, bootstrap and referral caching, stale-cache fallback, parent walking, add/sync/check/refresh, concurrency guard, alert selection and stale detection, the daily notification and its toggle, role rules) plus a live smoke test against real registries and a browser check of the page against the real router. Real dev database mtime untouched. Not checked: a screenshot of the finished page (the capture timed out); structure, sorting, errors and the viewer view were verified. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
4118062405 |
Alert when a Semaphore template or Gitea workflow run fails
Every 15 minutes (on the existing health-check timer) the app reads the latest run of each Semaphore template and each Gitea repo's latest workflow run. A failed one raises one notification, and another when a later run succeeds. It is state-based like the server health alerts, so a job that fails every night alerts on the first failure, not every night. Gitea alerts include the run's link. Toggle: Settings > Notifications. What counts: - Semaphore "error" is a failure, "success" is a pass. A run that is waiting, running, stopped by hand or rejected is neither, so it leaves the previous state alone: a run in progress must not clear a failure it hasn't fixed yet, and a manual stop isn't a failure. - Gitea failure/success likewise; running, waiting, blocked, cancelled and skipped leave things as they were. - A failing template that gets another failing run does not re-alert. Not mistaking "couldn't read" for "fixed": - Semaphore's template listing swallowed per-project errors, so a project that failed to load looked like a project with no templates. A new checkTemplates adapter method reports which projects failed, and their failures are held rather than cleared. - Gitea reports a run it couldn't fetch as null, the same as "no runs"; both leave the repo's state alone. - An unreachable integration holds all of its failures. Nothing is cleared or re-announced while it is down. The first pass only records what is already failing without announcing it, so upgrading (or adding an integration to a fresh install) doesn't produce a wall of alerts about months-old failures. That baseline is not spent while nothing could be read. Maintenance windows on a Semaphore or Gitea integration silence its failure alerts with the same rules as the health alerts: a problem that starts during a window alerts when it ends, and one already announced stays known. The diff logic is reused from the health monitor rather than copied. Maintenance page text updated. Known limit: for Gitea this follows the repo's most recent run on any workflow or branch, matching what the Gitea page shows; a failure in one workflow can be masked by a later success of another. Verified with 53 checks against fake Semaphore and Gitea servers and a webhook receiver: classification, baseline (including not being consumed when nothing is readable), single alert per failure, no repeat, in-progress/ stopped/cancelled runs, recovery and re-failure, unreadable project, unreadable integration, run-fetch errors, maintenance windows (silenced, then announced after), the toggle, disabled integrations and repos without Actions. Real dev database mtime untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
9f1609c4ed |
Add tags to servers, with a tag filter on the Servers page
Servers can be tagged (prod, media, rack-1, ...) for grouping. Operators and admins edit tags inline on a server's detail page; the input suggests tags already used on other servers so the same word ends up spelled the same way everywhere. Tags show as chips that keep one colour per tag, on the server cards, in the Manage table (and its CSV export), and on the detail page, where each chip links to the Servers list filtered by that tag. The Servers page has a tag bar with counts; picking several tags narrows to servers that have all of them. The filter lives in the URL, so it survives a refresh and can be linked to. Tags are also matched by the global search. Tags are normalized on the server (trimmed, lowercased, spaces become "-", duplicates merged, sorted); letters in any language are allowed, plus digits and - _ . : /, at most 30 characters and 12 per server. Invalid input is rejected with a message naming the offending tag, and nothing is saved. Changes are audit-logged with the before and after lists. Stored as a JSON column on servers (migration 0010). Every server response now returns tags as an array, and the shared response shaping strips the token hash in one place instead of five. Verified with 27 backend checks (normalization edge cases including Swedish letters, roles, validation, search, audit, PATCH/detail/list shapes) and by driving the real Servers and detail pages against the real router in a browser (filtering, editing, invalid tag, viewer view). Real dev database mtime untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
4c11158e98 |
Add a Ports card to server pages: scan for open ports, find free ones, and keep notes
Each server's detail page now has a Ports card. "Scan…" runs a TCP connect scan of a chosen range from the app and shows what's open, along with the ranges that were actually confirmed free; clicking a free range starts a reservation. Any port can carry a service name and a comment, so the page also answers "what is this port for". A port with a note counts as taken even when nothing is listening, which is what makes a reservation work. Operators can scan and edit; everyone can read. Scans and note changes are audit-logged. Details that matter for correctness: - "Free" means the host actively refused the connection AND nobody has claimed the port. A port that never answers (firewall drop, host down) is reported as not answering, not as free. - A scan from elsewhere can't see services bound to localhost only, so the agent now also reports what is bound on the host (ss -tulnp) and those ports are treated as taken. They show as "local only". Existing agents keep working; re-run the install one-liner to add this. The field is validated leniently so one odd line can never cost an agent its whole report, tasks included. - If nothing answers at all during a scan, existing results are left alone instead of being marked all-closed. - Scan targets are limited to private addresses (RFC1918, Tailscale 100.64/10, link-local, IPv6 ULA/link-local); loopback and public addresses are refused. Ranges are capped at 20,000 ports, and only one scan runs per server at a time. - Rows exist only while they carry information: an open port, or one with a note. A closed port with no note disappears on the next scan; one with a note stays as "reserved". New table server_ports plus two columns on servers (migration 0009). Verified with 76 backend checks (scanner open/refused/filtered, address rules, agent report leniency, note/reserve/clear semantics, free-range calculation including the localhost-only case, roles, concurrency lock, no-response guard, audit entries, cascade delete) and by driving the real component against the real router in a browser. Real dev database mtime untouched. Not verified: the agent's ss/awk/jq pipeline on a real host — the awk step was checked against sample ss output and the script passes bash -n, but jq isn't available here to run the whole thing. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
aae4f0d74f |
Add maintenance mode to silence alerts while working on a server or integration
Rebooting Proxmox or patching a server triggered failure/offline alerts you then had to dismiss. A maintenance window silences alerts about one server, integration, or DNS provider for a chosen time. New Maintenance page (start with a duration and optional reason, end early, see what's silenced and what isn't) and a banner in the app shell so every signed-in user can see what is currently silenced. Starting/ending is operator-only and audit-logged; starting one on a target that already has a window restarts its clock instead of stacking. Silenced for the target: server offline/disk alerts, Proxmox/Synology storage and health alerts, Proxmox backup alerts, and "integration down" alerts. Not silenced: expiry and update reminders, DNS change notices. The design goal is that this cannot hide a real outage: - Every window has a required end (5 min to 7 days); there is no open-ended option, so a forgotten window expires by itself. - A silenced problem is deliberately NOT recorded as "known". If it is still present when the window ends it alerts then, as new. A problem that was already alerted before the window stays known, so it isn't repeated, and is reported cleared only after the window ends. - Failure alerts keep counting failures during a window without marking themselves alerted, so an outage that outlasts the window alerts on the very next failed call. Known limitation, stated on the page: integration-failure alerts are tracked per service TYPE (all "proxmox"), not per configured instance, so a window on one Proxmox integration also silences a failure on a second Proxmox integration while it's open. Fixing that means threading the integration id through every adapter and the diagnostic log, which is a much larger change than this feature. Also moved the API-error-message helper out of Secrets.tsx into a shared util now that two pages use it. New table maintenance_windows (migration 0008). Verified with 44 checks: the condition-key-to-subject mapping (including server:3 vs server:33), the diff rules with silenced subjects (new problem not recorded, alerts when the window ends; already-known one carried and not repeated; clears only after the window), window expiry and integration/DNS-provider source matching, the failure tracker end to end against a webhook (silent during a window while an unrelated service still alerts; outage that outlasts the window alerts on the next failure and only once; fail-and-recover fully inside a window sends nothing), a full health pass against a real window, and the real router with a stubbed session (role rules, duration bounds including the missing-duration case, extend-not-stack, 404s, deleted targets hidden, audit entries). Real dev database mtime untouched. Not done: I haven't clicked through the new page or banner in a browser (they sit behind the Authentik login); it builds and the API behind it is tested. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1688de3ea2 |
Alert when a server goes silent, a disk fills up, or a Synology volume degrades
The data was all being collected (agent last-seen, per-disk usage, Proxmox storage, Synology volume/disk health) but nothing acted on it, so a dead server or a full disk was only noticed by opening the right page. A new health pass runs every 15 minutes (matching the agent's default report interval) and raises one notification when a problem starts and one when it clears: a server's agent silent past a threshold (default 60 min), a server disk / Proxmox storage or root filesystem / Synology volume at or above a usage threshold (default 90%), and a Synology volume or disk that isn't "normal", has bad SMART, bad sectors past the threshold, or life remaining below it. Both thresholds and an on/off toggle live under Settings -> Notifications. The parts that make this trustworthy rather than noisy: - A problem is keyed by identity, so it alerts once and not every run; a shared Proxmox storage listed by every node is one problem, not one per node. - Active problems persist across restarts, so a rebuild doesn't re-alert everything already known. - If a source can't be read on a given run (Proxmox/Synology unreachable, one node lacking privileges) its existing problems are held, not reported "cleared" and then re-alerted when it comes back — the integration-failure alert already owns "the integration is down". - For 20 minutes after startup server-derived problems are held too: agents couldn't report while the app was down, so judging them then would report every server offline after any restart. - An offline server's disk figures are stale and are not judged; a server that never reported has no agent and raises nothing. - Tracking continues while the toggle is off (only sending is gated), so turning it back on doesn't dump every long-standing problem. Timestamps without a zone (SQLite's format) are read as UTC; the test runs on a UTC+2 machine, where reading them as local time gives a different answer. Verified with 32 checks: the evaluation rules and the state diff as pure functions (exact thresholds, the proxmox:1 vs proxmox:10 prefix trap, the flapping sequence), then a whole pass against a real Proxmox adapter talking to a fake HTTPS cluster (one node returning 403, the whole API down, a shared storage on two nodes, a node that recovers), a webhook receiver, the real DB, and the persisted state. Not exercised end-to-end: the Synology collection path — its rules are tested on data shaped exactly like the adapter's output types, but I did not stand up a fake DSM. Real dev database mtime untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
7e81306aa7 |
Read SSL certificate expiry from the live server instead of trusting a typed-in date
A certificate secret's expiry was only ever what someone typed in, so a renewed cert (or a wrong date) meant the app's reminders were silently wrong. A certificate secret can now be given a host:port; the app opens a real TLS connection and reads the certificate's actual expiry — on create/edit (if the host changes), daily, and via a per-row "Check now" — and keeps expiryDate in sync. Because the daily refresh runs before the existing expiry check, the reminder is always computed from what's actually being served. Verification is deliberately off for the connection: homelab services routinely serve self-signed/internal-CA certs, and an already-expired one is exactly the case worth reporting, which a verifying connection would refuse before exposing the dates. Failure handling avoids the silent-staleness this is meant to fix: a failed check keeps the last known date, records why on the row (shown as a "Check failed" badge), and is listed in the daily secrets notification. Creating a monitored secret whose host can't be reached and with no manual date is rejected with the reason rather than saved blank. A non-TLS port (the likeliest typo) gets a plain-language error instead of raw OpenSSL output. Server-side connections to a user-supplied host:port need the same operator role that already gates editing secrets (and running Semaphore templates, which is strictly more powerful); the host is validated against a strict character set before any connection is made. New nullable secrets columns (check_host, check_port, last_checked_at, last_check_error) via migration 0007; existing rows are unaffected. Verified against real TLS servers (openssl-generated certs) and the real secrets router with a stubbed session: a live 45-day cert read back as the correct date via both an IP host (no SNI) and a hostname; an already-expired cert reported its past date and shows as expired; refused connections, a server that accepts but never answers (times out), and a plain non-TLS server each produced a descriptive error rather than a hang or crash. Through the router: create with a host and no date reads the date; unreachable host with no date -> 400 with the reason; unreachable with a manual date -> saved with the error recorded; host on a non-certificate type and an invalid host string -> 400; a hand-typed date on a monitored secret is ignored; changing the host re-checks immediately; changing the type away from certificate ends monitoring; a viewer gets 403 on Check now. 23 checks, all passing (a first re-run showed 2 spurious failures that were leftover rows from the previous run's scratch database, confirmed by a clean re-run). Real dev database mtime untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ca61f2a915 |
Surface Proxmox VMs/LXCs with no backup coverage at all
A failing backup run is visible now, but a guest with no backup job covering it in the first place was still a silent gap. Rather than depending on Proxmox's /cluster/backup-info/not-backed-up-guests endpoint (only exists on newer PVE versions), this derives coverage from data already fetched: a guest counts as covered if any enabled job either lists its vmid directly, or backs up "all guests" (scoped to the job's node, if it has one) without excluding it. Adds a warning banner plus a full table to the Proxmox page's Backups card, and extends the existing daily "Proxmox backup failed" notification (relabeled to mention this too) to also list uncovered guests, gated by the same toggle. Verified the coverage logic directly (it's a pure function, so no fake server needed) across 7 cases: no jobs at all, an all-guests job with an exclude list, a specific-vmids job, a node-scoped job that shouldn't cover a guest on a different node, a disabled job providing no real coverage, two jobs whose combined scope covers everything neither would alone, and a realistic mixed scenario — all passed. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
73d649377e |
Add notification quiet hours with digest delivery
Instant notifications (DNS changes, integration failure/recovery alerts) had no way to avoid pinging overnight. Adds a "Quiet hours" window under Settings -> Notifications: notifications that would fire during the window are held in a new notification_queue table instead of sent immediately, then delivered as one combined digest at the end time (in the same timezone already used for the daily checks) via a new scheduled flush job. The scheduled daily checks (secret expiry, Tailscale key, Docker updates, Proxmox backups) already only fire once at a chosen time, so this mainly matters for the instant ones. Includes a live "N queued" indicator with a manual "Flush now" button for visibility, and correctly falls back to sending immediately whenever the feature is disabled (the default). Gating lives at the single choke point every notification already flows through (notify()), so no per-event-type wiring was needed. Verified against a fake webhook receiver on an isolated scratch database: exhaustively checked the midnight-wraparound window math (9 cases including exact-boundary inclusive/exclusive edges) against synthetic "now" values rather than depending on when the test happens to run, then end-to-end through the real notify()/ flushQuietHoursQueue() functions with a window constructed around the actual current time — confirmed a notification during the window queues instead of sending, the flush produces one digest with the original title/message intact and clears the queue, a notification outside the window sends immediately, and disabling the feature entirely sends immediately regardless of the window. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
99db7e1cf0 |
Surface Proxmox backup job status, with a daily failure notification
Proxmox already runs vzdump backups, but nothing in the app said
whether they were actually succeeding — a silent backup failure is
one of the more dangerous blind spots a homelab admin can have. Adds
a "Backups" card to the Proxmox page: configured backup job
schedules (storage target, which guests, enabled/disabled) from
GET /cluster/backup, and recent vzdump task history per node from
GET /nodes/{node}/tasks?typefilter=vzdump, with a banner at the top
if the most recent run didn't succeed.
New "Proxmox backup failed" notification toggle under Settings ->
Notifications, on the same daily schedule as the other checks. The
scheduler checks each node's own most-recent vzdump run independently
(not just the single most recent task overall) so one node's healthy
backup can't mask another node's failing one in a multi-node cluster.
Known limitation, documented in the adapter's own header comment:
Proxmox's task list doesn't reliably expose which specific guest
failed within an "all guests" job — only the task's own log text has
that — so this surfaces job- and task-level status rather than
guessing at per-guest outcomes.
Verified against a fake Proxmox server (real self-signed HTTPS, since
the adapter's node:https usage can't be monkey-patched under ESM)
reproducing the documented /cluster/backup and task-list response
shapes: job parsing (all-guests+exclude vs specific-vmids+disabled)
correct, task OK/failure parsing correct, and the critical multi-node
scenario confirmed — one node's failing latest run flagged, the
other's healthy latest run correctly left alone, with exactly one
notification of the right content. This reproduces Proxmox's
documented API shape rather than a live-verified one; flag if the
real cluster's response differs in some way this didn't anticipate.
Confirmed the real dev database's mtime was untouched throughout.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
6d673db9ec |
Add session management: see who's signed in, revoke a session
No visibility existed into who was currently signed in or a way to force a device out. Sessions already live as files via session-file-store, so this reads that store directly rather than adding a new DB table: new Settings-adjacent "Sessions" page (admin-only, alongside Users) lists every live session with the user's name/email/resolved role, IP, a friendly "Browser on OS" summary parsed from the user-agent, last-active time, and expiry, with a Revoke button per row (extra confirmation if you revoke your own current session, since that signs you out immediately). IP and user-agent are now captured into the session at login (auth/router.ts) since express-session doesn't track them itself. session-file-store's own Store type doesn't declare its list() method, so sessionStore.ts adds a narrow local interface for it rather than losing type safety on the rest of the store. Verified against a real session directory seeded through the actual session-file-store APIs (not hand-written JSON): confirmed correct field resolution including a session whose user row was later deleted (role resolves to null instead of crashing), correctly excluded a mid-OIDC-login session with no completed user yet, correctly excluded an already-expired session, and confirmed revoke actually deletes the right session file and only that one. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1cae35a59e |
Add a daily Docker image-update notification
Dockhand's pending-update counts were only visible if you happened to open the Docker page. New "Docker image update available" toggle under Settings -> Notifications, sharing the same daily time/timezone as the secret and Tailscale key expiry reminders (same node-schedule reschedule-on-settings-change pattern as those two). Reads each enabled Dockhand integration's already-cached update-check results via listContainers() rather than triggering a fresh per-container registry lookup, so it costs nothing extra beyond what the Docker page itself already fetches, and lists every container with an update pending across all environments/integrations in one notification. Verified end-to-end against a fake local Dockhand server (one environment, two containers, one flagged with a pending update) and a fake webhook receiver on an isolated scratch database: the check correctly found only the flagged container (with its newerVersion), sent exactly one notification with the right title/content, excluded the up-to-date container, and sent nothing at all when the setting was toggled off despite still finding the same pending update. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
23eb7f0d70 |
Add passphrase-protected export/import for integrations, DNS providers, and settings
Nothing let you back up or migrate the app's own configuration short of copying the raw SQLite file. Adds Settings -> Backup: export decrypts every integration/DNS provider credential (normally encrypted at rest with this server's CREDENTIALS_ENCRYPTION_KEY) and re-encrypts the whole payload with a passphrase you choose (scrypt- derived key, AES-256-GCM), so the file is portable to a different instance with a different encryption key rather than being tied to this one. Import decrypts with that passphrase and merges settings onto the current ones; integrations/DNS providers are only added when no existing row shares their type+name, so re-running an import never duplicates or overwrites a working credential. Scope is configuration only — no DNS records, secrets, IPAM, servers, or audit/diagnostic log data. Verified end-to-end against two isolated scratch databases with different encryption keys (proving actual cross-instance portability, not just round-tripping through the same key): export -> encrypt -> write file -> decrypt on the other DB -> import -> re-decrypt the newly created integration/provider using the target's own key, confirming the plaintext credentials survived correctly; a wrong passphrase failed loudly (GCM auth failure) as expected; and re-running the same import a second time skipped both rows instead of duplicating them. Confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b5a4c6e2d9 |
Alert when an integration or DNS provider fails repeatedly
The Diagnostic Log already records every outbound call's success or failure, but nothing acted on it — you'd only notice an integration was down by happening to open its page. Adds a per-source consecutive- failure counter (in-memory, reset on restart, same durability tier as the diag log's own ring buffer) hooked into recordDiagEntry: crossing the configurable threshold (default 3) sends one "down" notification on every configured channel, and a "recovered" notification fires once it succeeds again — no repeat spam while it stays down. New "Integration/DNS provider failing repeatedly" toggle and threshold field under Settings -> Notifications. Verified end-to-end against an isolated scratch database with a real local HTTP server standing in for the webhook channel: 5 consecutive failures produced exactly one "Down" notification (at the 3rd failure, correctly naming "3 calls"), a subsequent success produced exactly one "Recovered" notification, and two more failures on a fresh streak triggered nothing (below threshold) — confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
10d123b18a |
Add automatic retention purging for the Diagnostic and Audit logs
The diagnostic log already rings-buffer to 500 rows, but the audit log had no cap at all and would grow forever. Adds an opt-in age-based purge under Settings -> Logs: keep entries for N days, checked on a configurable interval (hourly through monthly), plus a manual "Purge now" button. Reuses the existing node-schedule-style reschedule-on-settings-change pattern from the secret/Tailscale expiry checkers, but as a plain setInterval since "how often" here is an interval rather than a specific daily time. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b40234a557 |
Make table page size a configurable setting, not a hardcoded 20
Pagination just landed hardcoded to 20 rows everywhere; add a way to change that instead of leaving it fixed for every table in the app. pageSize joins dateFormat/timeFormat on the existing display settings object (server-side default 20, 5-500 range enforced by the PUT schema) rather than becoming its own settings section, since it's the same kind of thing -- an admin-configured, globally-applied display preference read by every signed-in role via the already-public GET /api/settings/display endpoint, same as the date/time format already works. Renamed the "Date & Time" settings tab/page/route to "Display" (still just one component, now covering both date/time format and table pagination) since its scope no longer matches the old name -- kept Settings.tsx's usual pattern of one page per concern rather than adding a second, oddly-scoped tab just for one number field. New web/src/utils/pageSize.ts mirrors utils/date.ts's existing module-level "set once at startup, read anywhere without prop- drilling" pattern; usePagination()'s pageSize parameter now defaults to getPageSize() instead of a literal 20, evaluated fresh on every call so it picks up a saved change without touching any of the ten pages already using the hook. Verified server-side against a temp SQLite DB: pageSize defaults to 20, a partial update sets it without disturbing dateFormat/timeFormat and vice versa, and it persists across a fresh settings read. Also checked the default-parameter mechanics directly (re-evaluates the global value on every call rather than capturing it once, and an explicit override still wins). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
e35da87886 |
Add a Diagnostic Log, ported and generalized from Sloth Manager
Sloth Manager tracked every API call made to DNS providers for connectivity troubleshooting. Port that here, generalized to cover every outbound integration this app makes, not just DNS -- Tailscale, Proxmox, Synology, Semaphore, Gitea, and Dockhand calls now show up too, since a broken API token or unreachable host on any of them is just as worth diagnosing. New services/diagLog.ts: a generic withDiagLogging(source, adapter) wraps every async method of any adapter object with timing + success/failure recording, without touching a single adapter's request/error-handling internals -- every DNS and integration adapter interface here is already just a flat set of async methods, so this one wrapper works for all twelve of them. Applied it at each adapter factory's own return statement (one line each) rather than at the route layer, so background jobs that construct adapters directly (the Tailscale key-expiry scheduler, IPAM sync, agent-driven Proxmox lookups) get logged too, not just requests through routes/integrations.ts. New diag_log table (ring-buffered to the last 500 rows, mirroring Sloth Manager's approach -- this is for live troubleshooting, not a durable record) and admin-only GET/DELETE /api/diag-log routes, source/ result filters, pagination. New admin-only Diagnostic Log page: filterable, paginated table with a Clear button. Also introduces the shared useSortable hook + SortableTh component used here for the first time -- a follow-up commit applies the same sorting (and CSV export) to the rest of the app's tables, per the same request. Verified end-to-end against a temp SQLite DB with real migrations: a fake wrapped adapter's successful and failing calls both land correctly in the log with the right source/operation/latency/error, and the source/ok filters and clear-log operation all behave correctly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1ff59afb40 |
Notify on expiring Tailscale device keys
The Tailscale page already showed per-device key expiry; extend the existing daily-reminder infrastructure (currently only for Secrets) to push it out through the configured notification channels too, the same way expiring secrets already are. New tailscaleKeyExpiryScheduler.ts mirrors secretExpiryScheduler.ts: runs once at startup (skipped if already run today) and daily thereafter, checking every enabled Tailscale integration's devices for keys expiring within the warning window and calling notify() with the results. Reuses the exact same daily time/timezone setting as the secret-expiry check (one "Daily reminder time" control, two independent on/off toggles) rather than adding a second schedule for users to configure. Centralized the expiring-soon threshold and check (previously only duplicated in the /synology and /tailscale route summaries) into adapter.ts as `KEY_EXPIRY_WARN_DAYS` / `isKeyExpiringSoon()`, and updated the devices route to use it instead of its own inline copy. New `tailscaleKeyCheck` notification-event toggle (default on) in settings, alongside the existing secret-expiry one. Verified end-to-end against a temp SQLite DB + real migrations: a tailscale integration pointed at a mock Tailscale API (one device expiring in 10 days, one with key-expiry disabled) with the webhook channel enabled and pointed at a mock receiver — confirmed the scheduler's startup check queries the DB correctly, decrypts the integration's credential, calls the adapter, filters out the disabled-expiry device, and delivers a webhook payload naming only the expiring device. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
f92f8de96e |
Add a Date & Time setting and fix inconsistent date formats app-wide
Different tables used different date formats: Servers & Tasks used a fixed "YYYY-MM-DD HH:mm:ss" (24h), while Audit Log, DNS, Integrations (Tailscale/Semaphore), and Users called plain .toLocaleString() with no options, which renders using the browser's own locale — different per browser/OS, and inconsistent with the other pages' fixed format. - New Settings -> Date & Time page: pick date order (YYYY-MM-DD, DD/MM/YYYY, MM/DD/YYYY) and 12h vs 24h clock, with a live preview. - utils/date.ts's formatDateTime() now reads these settings instead of being hardcoded to sv-SE/24h; the setting is fetched once at app startup (alongside /api/me) via a new non-secret GET /api/settings/display (any signed-in user, same rationale as the badge-color endpoints) and applied immediately on save too, without needing a page reload. - Switched every remaining raw new Date(...).toLocaleString() call (Audit Log, DNS zone sync time, Tailscale/Semaphore last-seen, Users' last login) over to the shared formatter, so every table now renders dates identically. Verified the formatter's date-order x 12h logic against all six combinations plus the midnight/noon 12h edge cases, and the settings endpoints end-to-end (defaults, partial updates, validation rejection, audit logging) against the real dev server. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ad6783b6a1 |
Extend badge color settings to cover integration types too
Settings → DNS Badges only let you customize DNS provider badge colors. Expand it into a general Badges page with a second section for the six integration types (Tailscale, Proxmox, Synology, Semaphore, Gitea, Dockhand), applied to the type badge in the Manage integrations table. - New integrationColors key on AppSettings, stored/merged the same way as providerColors via the existing settingsStore. - New GET /api/settings/integration-colors — non-secret, any signed-in user, mirroring /provider-colors — so the badge color can be read without needing admin access to the full settings payload. - Renamed DnsBadgeSettings.tsx -> BadgeSettings.tsx (route /settings/dns-badges -> /settings/badges, sub-nav label "DNS Badges" -> "Badges") with both color sections saved together. Verified end-to-end against the real dev server: PUT persists integration colors independently of provider colors, the public integration-colors endpoint reflects updates immediately, and the change is audit-logged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ae41a02864 |
Fix agent silently failing to report on hosts without a cpuinfo model name
report-tasks.sh's new hardware-collection code ran under set -euo pipefail, so a single failing command inside it aborted the whole script before the report was ever sent — with no error message, since nothing in that path had explicit error handling. On real hardware this hit immediately: `grep -m1 "model name" /proc/cpuinfo` exits 1 when there's no match, and many ARM boards (e.g. Raspberry Pi) have no such line at all. Confirmed via journalctl showing the systemd service failing every 15 minutes with exit 1 and zero output, and via a minimal repro of the exact bash control flow. Wrap the hardware/network collection call so any failure inside it is non-fatal: task reporting (the actual core function) must never be taken down by a quirk in the best-effort hardware-gathering code, on this host or any other. Also fixed a related gap found while testing the fallback path: the server only accepted the system field being absent, not explicitly null (what the script now sends if collection fails outright), which would have turned graceful degradation into a rejected report. Verified end-to-end against the real dev server: system:null, system omitted, and a normal populated report all now return 202 and persist correctly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b9409d3095 |
Add server hardware/network detail view with Proxmox sync and agent reporting
Servers & Tasks only tracked scheduled tasks — there was no overview of the servers themselves and no way to see CPU/RAM/disk/IP info. Add a clickable server overview (visible to every role, not just admins) that opens a per-server detail page. - servers table gains an agent-reported hardware/network snapshot (IPs, CPU model/cores/load, memory, disks) and an optional link to a Proxmox VM/LXC (integration + node + guest type + vmid). - Proxmox adapter gains getGuestDetail(): live cores/memory/disk from /config, live cpu/mem/uptime from /status/current, and IPs (LXC net config directly, QEMU via a best-effort guest-agent call that degrades gracefully when the agent isn't installed). - New GET /api/servers/:id/detail combines whichever hardware source applies (live Proxmox vs. last agent report) with a DNS reverse-lookup against the DNS module's own record cache, so matching hostnames show up next to each IP. New PATCH /api/servers/:id manages the Proxmox link. - agent/linux/report-tasks.sh now also collects and reports IPs, CPU/memory/disk info on every check-in (load-average-based CPU number, not instantaneous, to keep the agent a cheap oneshot). - Extracted the per-server task table into a shared ServerTaskTable component so the all-servers view and the new detail page render tasks identically. Verified server-side end-to-end against the real dev server (agent report -> detail endpoint -> DNS match) and the Proxmox adapter against a mock HTTPS server covering LXC/QEMU config parsing and the guest-agent-unavailable fallback. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
3255314402 |
Build the Settings module: notification channels, event toggles, DNS badge colors
The Settings page was a "coming soon" placeholder. Port Sloth Manager's settings feature set: Gotify/ntfy/SMTP/webhook notification channels (each with its own test-send button), per-event toggles (DNS record added/updated/deleted, a daily secret-expiry digest with configurable time/timezone), and per-provider DNS badge color customization. Settings persist in the existing `settings` key/value table via a new settingsStore service; a notify service fans a message out to every enabled channel. DNS record add/update/delete now fire notifications, and a node-schedule job re-arms itself whenever the notification settings change. Removed the now-superseded GOTIFY_URL/GOTIFY_TOKEN env vars in favor of in-app configuration. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
9712d611a6 |
Add Servers & Tasks module ported from Schedule Task Manager
Ports cron/systemd task tracking across Debian/Raspbian servers, including the Linux push agent (install/report/uninstall scripts, rebranded from "schedule-task-manager-agent" to "homelab-manager-agent") and the manual-task-entry flow for things an agent can't see (Docker jobs, backups). The schema (servers/scheduled_tasks tables) was already in place from the foundation pass, so this is mostly a straight port of the original's services/routes. Deliberate change from the original: server/token management (which mints agent credentials) is now admin-only rather than open to any logged-in user, and manual task CRUD is gated to operator+ — consistent with how Secrets, IPAM, and DNS already split "configure credentials" from "everyday edits" across roles. All mutations are audit-logged. - server/src/services/tokens.ts, taskSync.ts: ported near-verbatim (agent token hashing, the agent-sync-marks-missing-as-stale-not-deleted logic). - server/src/routes/servers.ts, tasks.ts, agentReport.ts: same contract as the original (agent auth is a per-server bearer token, independent of the session-based requireAuth used everywhere else). - web: a single Servers & Tasks page (filter bar, task table grouped by server/schedule type, manual task form, and an admin-only server management panel with token reveal + copyable install/uninstall commands), replacing the original's two separate pages/apps. Verified: full build passes; a scripted HTTP test against a running server covers unauthenticated access, role gating at each tier (admin-only server mgmt, operator+ task mgmt), agent bearer-token auth (valid/invalid/rotated), manual-vs-agent task edit protection, stale-marking on re-sync, and cascade delete — 22/22 checks passing. Real agent installation on an actual Debian/Raspbian host still needs to be tried on the user's network. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
6bd2ed52c1 |
Scaffold Homelab Manager foundation
Monorepo (Express+TS+Drizzle/libSQL server, React+Vite+Tabler web) matching the stack used by ScheduleTaskManager and Sloth Manager. Includes Authentik OIDC login with local admin/operator/viewer roles (first user becomes admin), a generalized audit log, encrypted-at-rest storage for future integration API tokens, the DB schema for all planned modules, and the Tabler-styled app shell/nav. Also ports the Secrets (expiry tracker) and IP Addresses (IPAM) modules from Sloth Manager onto the new stack. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |