From 447f33fff6e6c919d58b1b5bd3049dd2b66f19e8 Mon Sep 17 00:00:00 2001 From: Bobban Rydh Date: Fri, 2 Oct 2026 23:07:57 +0200 Subject: [PATCH] Add an Alerts page under Operations listing everything that's wrong now One list of the current problems across servers and integrations, instead of waiting for a notification or visiting each page: servers that stopped reporting, full or nearly full disks and volumes (critical from 95%), Synology volume/disk problems, failed or uncovered Proxmox backups, failed Proxmox Backup Server verifications, container image updates, expired or expiring secrets/domains/Tailscale keys, failed Semaphore and Gitea runs, Uptime Kuma monitors that are down, overdue osTicket tickets, and integrations whose calls keep failing. Visible to every role, with severity and kind filters, search, sorting, CSV export and "Check now". It runs the same checks that send the notifications rather than a second copy of them: the detection in the health, automation, Proxmox backup, PBS, Docker update and Tailscale key checks is pulled out into shared collectors that both the schedulers and the page call, so the two can't disagree about what counts as a problem. Notification behaviour is unchanged, including the scheduled backup checks skipping integrations under a maintenance window. Unlike the notifications the page ignores the on/off toggles, and keeps problems under a maintenance window, marked silenced and counted apart. It reads live, so a result is reused for a minute (and Refresh can't re-run everything more than once every ten seconds), and every source has a 20 s limit so one hung integration can't hang the page. Anything it couldn't read is called out at the top instead of looking like all clear, and server checks pause for the same 20 minutes after a restart as the notifications do, with a note saying so. Also gives the newer integrations (PBS, osTicket, Uptime Kuma, phpIPAM) proper names in "integration down" notifications instead of their ids. Verified through the real routes against a scratch database with fake backends (offline and full-disk servers, secrets and domains, a silenced server, a fake PBS with failed verification, a hanging integration, a refused one, a failing-calls streak, caching, the restart grace period, auth), and by rendering the real page against that data in a browser: filters, search, silenced toggle, sorting, Check now, dark mode. Co-Authored-By: Claude Sonnet 5.5 --- NOTIFICATIONS.md | 14 + README.md | 15 + ROLES.md | 1 + server/src/index.ts | 2 + server/src/routes/alerts.ts | 14 + server/src/services/alertTypes.ts | 28 ++ server/src/services/alerts.ts | 387 ++++++++++++++++++ server/src/services/automationMonitor.ts | 4 +- server/src/services/dockerUpdateScheduler.ts | 23 +- server/src/services/healthMonitor.ts | 20 +- .../src/services/integrationHealthMonitor.ts | 7 + server/src/services/notify.ts | 6 +- .../src/services/pbsVerificationScheduler.ts | 37 +- server/src/services/proxmoxBackupScheduler.ts | 50 ++- .../services/tailscaleKeyExpiryScheduler.ts | 23 +- web/src/App.tsx | 6 +- web/src/api/client.ts | 27 ++ web/src/layout/AppShell.tsx | 2 + web/src/pages/Alerts.tsx | 245 +++++++++++ 19 files changed, 883 insertions(+), 28 deletions(-) create mode 100644 server/src/routes/alerts.ts create mode 100644 server/src/services/alertTypes.ts create mode 100644 server/src/services/alerts.ts create mode 100644 web/src/pages/Alerts.tsx diff --git a/NOTIFICATIONS.md b/NOTIFICATIONS.md index fc1fefb..9b95336 100644 --- a/NOTIFICATIONS.md +++ b/NOTIFICATIONS.md @@ -89,6 +89,20 @@ schedule. |---|---| | "Notifications from quiet hours" | Once, at quiet hours' configured end time, **only if enabled and only if at least one notification was held** during the window. Bundles every held notification's title and message into one message, then clears the queue. | +## The Alerts page + +**Operations → Alerts** shows what these notifications are about — the problems that exist *right now* — as a list you can look at, filter, and +export. It uses the same checks as the notifications above, so a problem appears there for exactly the reason it would be notified, but it differs +in three ways: + +- It ignores the "Notify on" toggles. Turning a notification off doesn't hide the problem from the page. +- Problems under a **maintenance window** are kept on the list, marked *silenced* and counted separately, instead of being dropped. +- It also lists things nothing notifies about: Uptime Kuma monitors that are down, osTicket tickets that are overdue, and any integration that's + failing its last few calls (before the threshold that triggers an "integration down" notification). + +It runs the checks live when opened (a recent result is reused for a minute), and shows what it couldn't read at the top, so a missing section means +"couldn't check" and not "all clear". + ## What does *not* send a notification Worth calling out explicitly, since it's easy to assume everything in diff --git a/README.md b/README.md index f6ddca4..73aa214 100644 --- a/README.md +++ b/README.md @@ -255,6 +255,21 @@ already failing, so old failures aren't announced. Toggle it under Settings → Notifications. For Gitea this follows the repo's most recent run on any workflow or branch, the same as the Gitea page shows. +**Alerts** (Operations → Alerts, visible to every role) lists everything that's +wrong right now in one place, instead of waiting for a notification or visiting +each page: servers that stopped reporting, full or nearly full disks and volumes +(critical from 95%), Synology volume/disk problems, failed or uncovered Proxmox +backups, failed Proxmox Backup Server verifications, container image updates, +secrets, domains and Tailscale keys that are expired or about to be, failed +Semaphore/Gitea runs, Uptime Kuma monitors that are down, overdue osTicket +tickets, and integrations whose calls keep failing. It runs the same checks that +send the notifications — so the two can't disagree — but ignores the on/off +toggles, since it's for looking at rather than being interrupted by. Problems +under a maintenance window stay listed, marked silenced and counted separately. +It checks live (a recent result is reused for a minute; "Check now" forces a +fresh one), and anything it couldn't read is called out at the top rather than +quietly treated as fine. + **Maintenance mode** silences alerts about one server, integration, or DNS provider while you work on it (server offline / disk, storage and Synology health, Proxmox backup alerts, Proxmox Backup Server verification alerts, and diff --git a/ROLES.md b/ROLES.md index 5cce5c5..1bd2051 100644 --- a/ROLES.md +++ b/ROLES.md @@ -37,6 +37,7 @@ page's own top-level link. | Gitea | `/gitea` | ✅ | ✅ | ✅ | | Secrets | `/secrets` | ✅ | ✅ | ✅ | | **Operations** | | | | | +| Alerts | `/alerts` | ✅ | ✅ | ✅ | | Maintenance | `/maintenance` | ✅ | ✅ | ✅ | | Uptime Kuma | `/uptime-kuma` | ✅ | ✅ | ✅ | | osTicket | `/osticket` | ✅ | ✅ | ✅ | diff --git a/server/src/index.ts b/server/src/index.ts index 83e3ded..5c6991f 100644 --- a/server/src/index.ts +++ b/server/src/index.ts @@ -29,6 +29,7 @@ import { consistencyRouter } from "./routes/consistency.js"; import { privacyRouter } from "./routes/privacy.js"; import { tagsRouter } from "./routes/tags.js"; import { portsRouter } from "./routes/ports.js"; +import { alertsRouter } from "./routes/alerts.js"; import { initSecretExpiryScheduler } from "./services/secretExpiryScheduler.js"; import { initTailscaleKeyExpiryScheduler } from "./services/tailscaleKeyExpiryScheduler.js"; import { initLogRetentionScheduler } from "./services/logRetentionScheduler.js"; @@ -104,6 +105,7 @@ app.use("/api/consistency", consistencyRouter); app.use("/api/privacy", privacyRouter); app.use("/api/tags", tagsRouter); app.use("/api/ports", portsRouter); +app.use("/api/alerts", alertsRouter); if (existsSync(webDist)) { app.use(express.static(webDist)); diff --git a/server/src/routes/alerts.ts b/server/src/routes/alerts.ts new file mode 100644 index 0000000..22c9103 --- /dev/null +++ b/server/src/routes/alerts.ts @@ -0,0 +1,14 @@ +import { Router } from "express"; +import { requireAuth } from "../auth/middleware.js"; +import { getAlerts } from "../services/alerts.js"; +import { asyncHandler } from "../utils/asyncHandler.js"; + +export const alertsRouter = Router(); + +alertsRouter.use(requireAuth); + +// Everyone signed in can see it — it's the same server, disk and backup status the other pages already show. +// `?refresh=1` asks for a fresh check instead of a recent one. +alertsRouter.get("/", asyncHandler(async (req, res) => { + res.json(await getAlerts(req.query.refresh === "1")); +})); diff --git a/server/src/services/alertTypes.ts b/server/src/services/alertTypes.ts new file mode 100644 index 0000000..204993b --- /dev/null +++ b/server/src/services/alertTypes.ts @@ -0,0 +1,28 @@ +// Shared by the alerts page's collector (services/alerts.ts) and the scheduled checks it reuses, which can't import +// that file themselves without going round in a circle. + +export type AlertSeverity = "critical" | "warning" | "info"; + +/** What kind of problem — drives the filter on the Alerts page. */ +export type AlertCategory = "offline" | "disk" | "backup" | "updates" | "expiry" | "automation" | "integration" | "monitoring" | "tickets"; + +export interface Alert { + /** Stable, so the page can key rows on it. */ + id: string; + severity: AlertSeverity; + category: AlertCategory; + /** Where it comes from, as a short name: "Server", "Proxmox", "Secrets", ... */ + source: string; + message: string; + /** In-app page that shows more, if there is one. */ + link: string | null; + /** Under an active maintenance window: still a real problem, but notifications for it are held back. */ + silenced: boolean; +} + +/** One integration that couldn't be read while collecting — so the page never mistakes "couldn't check" for "all clear". */ +export interface SourceFailure { + integrationId: number; + integrationName: string; + message: string; +} diff --git a/server/src/services/alerts.ts b/server/src/services/alerts.ts new file mode 100644 index 0000000..2c9b394 --- /dev/null +++ b/server/src/services/alerts.ts @@ -0,0 +1,387 @@ +/** + * Everything that's wrong right now, in one list — the Alerts page. Nothing here decides what counts as a problem: it asks + * the same checks that send the notifications (health, backups, updates, expiry, automation, ...) and turns what they find + * into a flat list, so the page and the notifications can't disagree. Unlike the notifications it ignores the per-event + * on/off toggles (the page is for looking at, not for being interrupted by) and keeps problems that are under a + * maintenance window, flagged as silenced rather than dropped. + * + * It reads live (a handful of API calls per integration), so a result is kept for a short while rather than re-run for + * every viewer, and every source has a time limit so one hung integration can't hang the page. A source that can't be + * read is reported as such — silence from it must never look like "all clear". + */ +import { eq } from "drizzle-orm"; +import { db } from "../db/client.js"; +import { integrations, secrets } from "../db/schema.js"; +import { loadIntegrationConfig } from "../integrations/loadIntegration.js"; +import { createUptimeKumaAdapter } from "../integrations/uptimekuma/adapter.js"; +import { createOsTicketAdapter } from "../integrations/osticket/adapter.js"; +import { activeSubjects, isSourceInMaintenance, subjectOfConditionKey } from "./maintenance.js"; +import { collectSnapshot, evaluateHealth, startupGraceRemainingMs } from "./healthMonitor.js"; +import { collectAutomation, evaluateAutomation } from "./automationMonitor.js"; +import { collectDockerUpdates } from "./dockerUpdateScheduler.js"; +import { collectProxmoxBackupProblems } from "./proxmoxBackupScheduler.js"; +import { collectPbsProblems } from "./pbsVerificationScheduler.js"; +import { collectTailscaleKeyExpiries } from "./tailscaleKeyExpiryScheduler.js"; +import { collectDomainAlerts } from "./domainMonitor.js"; +import { computeSecretStatus } from "./secretStatus.js"; +import { getFailingSources } from "./integrationHealthMonitor.js"; +import { getSettings } from "./settingsStore.js"; +import { sourceLabel } from "./notify.js"; +import type { Alert, AlertCategory, AlertSeverity, SourceFailure } from "./alertTypes.js"; + +export interface AlertsReport { + alerts: Alert[]; + /** Active problems by severity — those under a maintenance window are counted apart, not in these. */ + counts: { critical: number; warning: number; info: number; silenced: number }; + /** Things that couldn't be checked, and which checks that leaves blind. */ + couldntCheck: { name: string; error: string; affects: string[] }[]; + /** Context worth knowing about how complete the picture is. */ + notes: string[]; + generatedAt: string; +} + +/** Where each kind of source lives in the app, and what to call it. */ +const SOURCES: Record = { + server: { label: "Server", link: "/servers" }, + proxmox: { label: "Proxmox", link: "/proxmox" }, + synology: { label: "Synology", link: "/synology" }, + semaphore: { label: "Semaphore", link: "/semaphore" }, + gitea: { label: "Gitea", link: "/gitea" }, + dockhand: { label: "Docker", link: "/docker" }, + tailscale: { label: "Tailscale", link: "/tailscale" }, + pbs: { label: "Proxmox Backup", link: "/pbs" }, + uptimekuma: { label: "Uptime Kuma", link: "/uptime-kuma" }, + osticket: { label: "osTicket", link: "/osticket" }, + secrets: { label: "Secrets", link: "/secrets" }, + domains: { label: "Domains", link: "/domains" }, +}; + +const SOURCE_TIMEOUT_MS = 20_000; +/** How long a result is reused. */ +const CACHE_MS = 60_000; +/** Pressing Refresh over and over shouldn't hammer every integration — a result younger than this is reused even then. */ +const MIN_REFRESH_MS = 10_000; + +const SEVERITY_ORDER: Record = { critical: 0, warning: 1, info: 2 }; + +/** Which kind of source, and which one, a health/automation condition key belongs to. */ +function originOfConditionKey(key: string): { type: string; id: number | null; category: AlertCategory } { + let m = /^(?:offline|disk):server:(\d+)/.exec(key); + if (m) return { type: "server", id: Number(m[1]), category: key.startsWith("offline") ? "offline" : "disk" }; + m = /^disk:proxmox:(\d+):/.exec(key); + if (m) return { type: "proxmox", id: Number(m[1]), category: "disk" }; + m = /^(?:synology-volume|synology-disk|disk:synology):(\d+):/.exec(key); + if (m) return { type: "synology", id: Number(m[1]), category: "disk" }; + m = /^automation:(semaphore|gitea):(\d+):/.exec(key); + if (m) return { type: m[1], id: Number(m[2]), category: "automation" }; + return { type: "server", id: null, category: "disk" }; +} + +function withTimeout(work: Promise): Promise { + let timer: ReturnType; + const limit = new Promise((_, reject) => { + timer = setTimeout(() => reject(new Error(`timed out after ${SOURCE_TIMEOUT_MS / 1000}s`)), SOURCE_TIMEOUT_MS); + }); + return Promise.race([work, limit]).finally(() => clearTimeout(timer)); +} + +const errorText = (err: unknown) => (err instanceof Error ? err.message : String(err)); +const plural = (n: number, one: string, many = `${one}s`) => `${n} ${n === 1 ? one : many}`; + +export async function collectAlerts(): Promise { + const alerts: Alert[] = []; + const notes: string[] = []; + const blind = new Map }>(); + const subjects = await activeSubjects(); + const intRows = await db.select({ id: integrations.id, name: integrations.name, type: integrations.type, enabled: integrations.enabled }).from(integrations); + const intById = new Map(intRows.map((r) => [r.id, r])); + + function push(a: { category: AlertCategory; severity: AlertSeverity; type: string; message: string; key: string; link?: string | null; silenced?: boolean }) { + const meta = SOURCES[a.type]; + alerts.push({ + id: `${a.category}:${a.key}`, + severity: a.severity, + category: a.category, + source: meta?.label ?? sourceLabel(a.type), + message: a.message, + link: a.link === undefined ? (meta?.link ?? null) : a.link, + silenced: !!a.silenced, + }); + } + + function cannotCheck(name: string, error: string, check: string) { + const entry = blind.get(name) ?? { error, affects: new Set() }; + entry.affects.add(check); + blind.set(name, entry); + } + const cannotCheckIntegration = (f: SourceFailure, check: string) => { + const row = intById.get(f.integrationId); + cannotCheck(`${row ? (SOURCES[row.type]?.label ?? row.type) : "Integration"} “${f.integrationName}”`, f.message, check); + }; + const silencedIntegration = (id: number) => subjects.has(`integration:${id}`); + + // One entry per check. A check that blows up, or hangs, only costs its own section of the page. + async function check(name: string, work: () => Promise) { + try { + await withTimeout(work()); + } catch (err) { + cannotCheck(name, errorText(err), name); + } + } + + await Promise.all([ + check("Server and storage health", async () => { + const { healthChecks } = await getSettings(); + const { snapshot, held } = await collectSnapshot(); + const now = Date.now(); + const grace = startupGraceRemainingMs(now); + if (grace > 0) { + notes.push( + `Server-offline and server-disk checks are paused for another ${Math.ceil(grace / 60_000)} min after the app restarted, so agents get a chance to report before any server is judged.`, + ); + } + for (const c of evaluateHealth(snapshot, healthChecks, now, { skipServers: grace > 0 })) { + const origin = originOfConditionKey(c.key); + const subject = subjectOfConditionKey(c.key); + push({ + category: origin.category, + severity: c.severity ?? "warning", + type: origin.type, + message: c.message, + key: c.key, + link: origin.type === "server" && origin.id !== null ? `/servers/${origin.id}` : undefined, + silenced: subject !== null && subjects.has(subject), + }); + } + // Integrations the check couldn't read this time. (A whole-integration entry covers its nodes, so skip those.) + for (const h of held) { + const m = /^(proxmox|synology):(\d+)(?::(.+))?$/.exec(h); + if (!m || (m[3] && held.has(`${m[1]}:${m[2]}`))) continue; + const row = intById.get(Number(m[2])); + cannotCheck(`${SOURCES[m[1]].label} “${row?.name ?? `#${m[2]}`}”${m[3] ? ` (node ${m[3]})` : ""}`, "couldn't be read — see the Diagnostic Log", "Server and storage health"); + } + }), + + check("Automation runs", async () => { + const { items, held } = await collectAutomation(); + for (const c of evaluateAutomation(items).conditions) { + const origin = originOfConditionKey(c.key); + const subject = subjectOfConditionKey(c.key); + push({ category: "automation", severity: "warning", type: origin.type, message: c.message, key: c.key, silenced: subject !== null && subjects.has(subject) }); + } + for (const h of held) { + const m = /^(semaphore|gitea):(\d+)$/.exec(h); + if (!m) continue; + const row = intById.get(Number(m[2])); + cannotCheck(`${SOURCES[m[1]].label} “${row?.name ?? `#${m[2]}`}”`, "couldn't be read — see the Diagnostic Log", "Automation runs"); + } + }), + + check("Proxmox backups", async () => { + const { failures, uncovered, sourceFailures } = await collectProxmoxBackupProblems({ skipSilenced: false }); + for (const f of failures) { + push({ + category: "backup", + severity: "critical", + type: "proxmox", + message: `Latest backup on ${f.node}${f.guestId ? ` (guest ${f.guestId})` : ""} didn't succeed [${f.integrationName}]: ${f.status}`, + key: `proxmox-backup:${f.integrationId}:${f.node}`, + silenced: f.silenced, + }); + } + for (const u of uncovered) { + push({ + category: "backup", + severity: "warning", + type: "proxmox", + message: `${u.guestName} (#${u.vmid}) on ${u.node} isn't covered by any backup job [${u.integrationName}]`, + key: `proxmox-uncovered:${u.integrationId}:${u.vmid}`, + silenced: u.silenced, + }); + } + sourceFailures.forEach((f) => cannotCheckIntegration(f, "Proxmox backups")); + }), + + check("Backup verification", async () => { + const { problems, sourceFailures } = await collectPbsProblems({ skipSilenced: false }); + for (const p of problems) { + push({ + category: "backup", + severity: p.error ? "warning" : "critical", + type: "pbs", + message: p.error + ? `Datastore "${p.datastore}" couldn't be read [${p.integrationName}]: ${p.error}` + : `Datastore "${p.datastore}" has ${plural(p.failedCount, "snapshot")} that failed verification [${p.integrationName}]`, + key: `pbs:${p.integrationId}:${p.datastore}`, + silenced: p.silenced, + }); + } + sourceFailures.forEach((f) => cannotCheckIntegration(f, "Backup verification")); + }), + + check("Image updates", async () => { + const { items, failures } = await collectDockerUpdates(); + for (const u of items) { + push({ + category: "updates", + severity: "info", + type: "dockhand", + message: `${u.containerName} [${u.environmentName}, ${u.integrationName}] has an image update available${u.newerVersion ? ` → ${u.newerVersion}` : ""}`, + key: `docker:${u.integrationId}:${u.environmentName}:${u.containerName}`, + silenced: silencedIntegration(u.integrationId), + }); + } + failures.forEach((f) => cannotCheckIntegration(f, "Image updates")); + }), + + check("Tailscale keys", async () => { + const { items, failures } = await collectTailscaleKeyExpiries(); + for (const k of items) { + push({ + category: "expiry", + severity: k.daysLeft < 0 ? "critical" : "warning", + type: "tailscale", + message: k.daysLeft < 0 ? `Key for ${k.deviceLabel} [${k.integrationName}] has expired` : `Key for ${k.deviceLabel} [${k.integrationName}] expires in ${plural(k.daysLeft, "day")}`, + key: `tailscale-key:${k.integrationId}:${k.deviceLabel}`, + silenced: silencedIntegration(k.integrationId), + }); + } + failures.forEach((f) => cannotCheckIntegration(f, "Tailscale keys")); + }), + + check("Uptime Kuma", async () => { + for (const row of intRows.filter((r) => r.type === "uptimekuma" && r.enabled)) { + try { + const loaded = await loadIntegrationConfig(row.id); + if (!loaded) continue; + for (const m of await createUptimeKumaAdapter(loaded.config as any).listMonitors()) { + if (m.status !== "down") continue; + push({ + category: "monitoring", + severity: "critical", + type: "uptimekuma", + message: `Monitor "${m.name}" is down${m.target ? ` (${m.target}${m.port ? `:${m.port}` : ""})` : ""} [${row.name}]`, + key: `kuma:${row.id}:${m.id}`, + silenced: silencedIntegration(row.id), + }); + } + } catch (err) { + cannotCheckIntegration({ integrationId: row.id, integrationName: row.name, message: errorText(err) }, "Uptime Kuma"); + } + } + }), + + check("osTicket", async () => { + for (const row of intRows.filter((r) => r.type === "osticket" && r.enabled)) { + try { + const loaded = await loadIntegrationConfig(row.id); + if (!loaded) continue; + const overdue = (await createOsTicketAdapter(loaded.config as any).listOpenTickets()).filter((t) => t.isOverdue).length; + if (overdue > 0) { + push({ + category: "tickets", + severity: "warning", + type: "osticket", + message: `${plural(overdue, "open ticket")} ${overdue === 1 ? "is" : "are"} overdue [${row.name}]`, + key: `osticket:${row.id}`, + silenced: silencedIntegration(row.id), + }); + } + } catch (err) { + cannotCheckIntegration({ integrationId: row.id, integrationName: row.name, message: errorText(err) }, "osTicket"); + } + } + }), + + check("Secrets", async () => { + for (const s of await db.select().from(secrets)) { + const status = computeSecretStatus(s.expiryDate, s.warnDays); + if (status.status === "expired") { + push({ category: "expiry", severity: "critical", type: "secrets", message: `Secret "${s.name}" expired ${plural(Math.abs(status.daysLeft), "day")} ago (${s.expiryDate})`, key: `secret:${s.id}` }); + } else if (status.status === "expiring") { + push({ category: "expiry", severity: "warning", type: "secrets", message: `Secret "${s.name}" expires in ${plural(status.daysLeft, "day")} (${s.expiryDate})`, key: `secret:${s.id}` }); + } + if (s.checkHost && s.lastCheckError) { + push({ + category: "expiry", + severity: "warning", + type: "secrets", + message: `Couldn't read the live certificate for "${s.name}" (${s.checkHost}:${s.checkPort ?? 443}) — the expiry shown may be stale: ${s.lastCheckError}`, + key: `secret-check:${s.id}`, + }); + } + } + }), + + check("Domains", async () => { + const { expiring, staleChecks } = await collectDomainAlerts(); + for (const d of expiring) { + push({ + category: "expiry", + severity: d.status === "expired" ? "critical" : "warning", + type: "domains", + message: d.status === "expired" ? `Domain ${d.name} expired on ${d.expiresAt}` : `Domain ${d.name} expires in ${plural(d.daysLeft, "day")} (${d.expiresAt})`, + key: `domain:${d.name}`, + }); + } + for (const s of staleChecks) { + push({ category: "expiry", severity: "info", type: "domains", message: `Couldn't refresh the registration for ${s.name} — the expiry shown may be stale: ${s.error}`, key: `domain-stale:${s.name}` }); + } + }), + + check("Integration failures", async () => { + const { notifications } = await getSettings(); + for (const f of getFailingSources()) { + push({ + category: "integration", + severity: f.alerted ? "critical" : "warning", + type: f.source, + message: `${sourceLabel(f.source)} has failed its last ${plural(f.consecutiveFailures, "call")} in a row${f.alerted ? "" : ` (a notification goes out after ${notifications.integrationFailureThreshold})`} — see the Diagnostic Log`, + key: `failing:${f.source}`, + link: null, + silenced: await isSourceInMaintenance(f.source), + }); + } + }), + ]); + + alerts.sort( + (a, b) => + Number(a.silenced) - Number(b.silenced) || + SEVERITY_ORDER[a.severity] - SEVERITY_ORDER[b.severity] || + a.source.localeCompare(b.source) || + a.message.localeCompare(b.message), + ); + + const active = alerts.filter((a) => !a.silenced); + return { + alerts, + counts: { + critical: active.filter((a) => a.severity === "critical").length, + warning: active.filter((a) => a.severity === "warning").length, + info: active.filter((a) => a.severity === "info").length, + silenced: alerts.length - active.length, + }, + couldntCheck: [...blind.entries()].map(([name, v]) => ({ name, error: v.error, affects: [...v.affects] })), + notes, + generatedAt: new Date().toISOString(), + }; +} + +let cache: { at: number; report: AlertsReport } | null = null; +let inFlight: Promise | null = null; + +/** The current alerts, reusing a recent result unless `force` asks for a fresh one (and even then not more than once every few seconds). */ +export async function getAlerts(force: boolean): Promise { + const age = cache ? Date.now() - cache.at : Infinity; + if (cache && age < (force ? MIN_REFRESH_MS : CACHE_MS)) return { ...cache.report, cached: true }; + inFlight ??= collectAlerts() + .then((report) => { + cache = { at: Date.now(), report }; + return report; + }) + .finally(() => { + inFlight = null; + }); + return { ...(await inFlight), cached: false }; +} diff --git a/server/src/services/automationMonitor.ts b/server/src/services/automationMonitor.ts index a9ee4d7..d1e2ec1 100644 --- a/server/src/services/automationMonitor.ts +++ b/server/src/services/automationMonitor.ts @@ -62,7 +62,7 @@ export function evaluateAutomation(items: AutomationItem[]): { conditions: Healt // ─── Collection (I/O) ─────────────────────────────────────────────────────── -async function collect(): Promise<{ items: AutomationItem[]; held: Set; readable: number }> { +export async function collectAutomation(): Promise<{ items: AutomationItem[]; held: Set; readable: number }> { const items: AutomationItem[] = []; const held = new Set(); let readable = 0; @@ -138,7 +138,7 @@ async function loadState(): Promise { * the first time this feature runs) that would otherwise be a wall of alerts about failures that are months old. */ export async function runAutomationCheck(now: number = Date.now()): Promise<{ added: number; resolved: number; baseline: boolean }> { - const { items, held: readHeld, readable } = await collect(); + const { items, held: readHeld, readable } = await collectAutomation(); const stored = await loadState(); // Nothing could be read at all (or nothing is configured): don't spend the "first pass" on an empty picture. diff --git a/server/src/services/dockerUpdateScheduler.ts b/server/src/services/dockerUpdateScheduler.ts index 91b5bda..29f8a0f 100644 --- a/server/src/services/dockerUpdateScheduler.ts +++ b/server/src/services/dockerUpdateScheduler.ts @@ -5,17 +5,28 @@ import { integrations } from "../db/schema.js"; import { loadIntegrationConfig } from "../integrations/loadIntegration.js"; import { createDockhandAdapter } from "../integrations/dockhand/adapter.js"; import { notifyDockerUpdates } from "./notify.js"; +import type { SourceFailure } from "./alertTypes.js"; import { getSettings, getInternalFlag, setInternalFlag } from "./settingsStore.js"; const LAST_RUN_FLAG = "dockerUpdateCheckLastRunDate"; -async function checkDockerUpdates(): Promise { +export interface DockerUpdate { + integrationId: number; + integrationName: string; + containerName: string; + environmentName: string; + newerVersion: string | null; +} + +/** Every container with an image update waiting, across the enabled Dockhand integrations. Shared with the Alerts page. */ +export async function collectDockerUpdates(): Promise<{ items: DockerUpdate[]; failures: SourceFailure[] }> { const rows = await db .select({ id: integrations.id, name: integrations.name }) .from(integrations) .where(and(eq(integrations.type, "dockhand"), eq(integrations.enabled, true))); - const updatesAvailable: { integrationName: string; containerName: string; environmentName: string; newerVersion: string | null }[] = []; + const updatesAvailable: DockerUpdate[] = []; + const failures: SourceFailure[] = []; for (const row of rows) { try { @@ -28,6 +39,7 @@ async function checkDockerUpdates(): Promise { for (const c of containers) { if (!c.updateAvailable) continue; updatesAvailable.push({ + integrationId: row.id, integrationName: row.name, containerName: c.name, environmentName: c.environmentName, @@ -36,10 +48,15 @@ async function checkDockerUpdates(): Promise { } } catch (err) { console.error(`[dockerUpdate] check failed for integration ${row.id}:`, err); + failures.push({ integrationId: row.id, integrationName: row.name, message: err instanceof Error ? err.message : String(err) }); } } - await notifyDockerUpdates(updatesAvailable); + return { items: updatesAvailable, failures }; +} + +async function checkDockerUpdates(): Promise { + await notifyDockerUpdates((await collectDockerUpdates()).items); } async function checkDockerUpdatesOnce(): Promise { diff --git a/server/src/services/healthMonitor.ts b/server/src/services/healthMonitor.ts index a1294a0..55bec76 100644 --- a/server/src/services/healthMonitor.ts +++ b/server/src/services/healthMonitor.ts @@ -21,6 +21,8 @@ export interface HealthCondition { /** Where the data came from, hierarchically ("server", "proxmox:3", "proxmox:3:pve1"). Used to hold a condition when its source can't be reached. */ source: string; message: string; + /** How bad, for the Alerts page. Notifications don't use it. Left out means "warning". */ + severity?: "critical" | "warning"; } export interface ServerSnapshot { @@ -91,6 +93,8 @@ export function evaluateHealth(snapshot: HealthSnapshot, thresholds: Thresholds, const out: HealthCondition[] = []; const limit = thresholds.diskUsagePercent; const pct = (n: number) => `${Math.round(n)}%`; + /** Over the configured threshold is a warning; practically out of room is critical. */ + const fullness = (p: number): "critical" | "warning" => (p >= 95 ? "critical" : "warning"); if (!opts.skipServers) { for (const s of snapshot.servers) { @@ -103,6 +107,7 @@ export function evaluateHealth(snapshot: HealthSnapshot, thresholds: Thresholds, key: `offline:server:${s.id}`, source: "server", message: `${s.name} hasn't reported for ${formatDuration(age)}`, + severity: "critical", }); } } @@ -115,6 +120,7 @@ export function evaluateHealth(snapshot: HealthSnapshot, thresholds: Thresholds, key: `disk:server:${s.id}:${d.mount}`, source: "server", message: `${s.name}: ${d.mount} is ${pct(p)} full (${formatBytes(d.usedBytes)} of ${formatBytes(d.sizeBytes)})`, + severity: fullness(p), }); } } @@ -132,6 +138,7 @@ export function evaluateHealth(snapshot: HealthSnapshot, thresholds: Thresholds, key: `disk:proxmox:${px.integrationId}:${node.node}:rootfs`, source: `proxmox:${px.integrationId}:${node.node}`, message: `${px.integrationName} / ${node.node}: root filesystem is ${pct(rootP)} full`, + severity: fullness(rootP), }); } for (const st of node.storages) { @@ -146,12 +153,14 @@ export function evaluateHealth(snapshot: HealthSnapshot, thresholds: Thresholds, key: `disk:proxmox:${px.integrationId}:storage:${st.id}`, source: `proxmox:${px.integrationId}:shared`, message: `${px.integrationName}: shared storage "${st.id}" is ${pct(p)} full (${formatBytes(st.usedBytes ?? 0)} of ${formatBytes(st.totalBytes ?? 0)})`, + severity: fullness(p), }); } else { out.push({ key: `disk:proxmox:${px.integrationId}:${node.node}:storage:${st.id}`, source: `proxmox:${px.integrationId}:${node.node}`, message: `${px.integrationName} / ${node.node}: storage "${st.id}" is ${pct(p)} full (${formatBytes(st.usedBytes ?? 0)} of ${formatBytes(st.totalBytes ?? 0)})`, + severity: fullness(p), }); } } @@ -166,6 +175,7 @@ export function evaluateHealth(snapshot: HealthSnapshot, thresholds: Thresholds, key: `synology-volume:${syn.integrationId}:${v.id}`, source: src, message: `${syn.integrationName}: volume ${v.id} status is "${v.status}"`, + severity: "critical", }); } const p = usage(v.sizeUsed, v.sizeTotal); @@ -174,6 +184,7 @@ export function evaluateHealth(snapshot: HealthSnapshot, thresholds: Thresholds, key: `disk:synology:${syn.integrationId}:${v.id}`, source: src, message: `${syn.integrationName}: volume ${v.id} is ${pct(p)} full (${formatBytes(v.sizeUsed ?? 0)} of ${formatBytes(v.sizeTotal ?? 0)})`, + severity: fullness(p), }); } } @@ -188,6 +199,8 @@ export function evaluateHealth(snapshot: HealthSnapshot, thresholds: Thresholds, key: `synology-disk:${syn.integrationId}:${d.id}`, source: src, message: `${syn.integrationName}: disk ${d.name || d.id} — ${problems.join(", ")}`, + // A disk that's no longer "normal" is failing; a SMART warning or a threshold crossing is the early notice. + severity: d.status && d.status.toLowerCase() !== "normal" ? "critical" : "warning", }); } } @@ -247,7 +260,12 @@ export function diffConditions( // ─── Collection (I/O) ─────────────────────────────────────────────────────── -async function collectSnapshot(): Promise<{ snapshot: HealthSnapshot; held: Set }> { +/** How much longer, in ms, servers are exempt from being judged after a restart (their agents get a chance to report first). 0 once it's over. */ +export function startupGraceRemainingMs(now: number = Date.now()): number { + return Math.max(0, STARTUP_GRACE_MS - (now - PROCESS_START)); +} + +export async function collectSnapshot(): Promise<{ snapshot: HealthSnapshot; held: Set }> { const held = new Set(); const snapshot: HealthSnapshot = { servers: [], proxmox: [], synology: [] }; diff --git a/server/src/services/integrationHealthMonitor.ts b/server/src/services/integrationHealthMonitor.ts index 1faf851..d9466e7 100644 --- a/server/src/services/integrationHealthMonitor.ts +++ b/server/src/services/integrationHealthMonitor.ts @@ -17,6 +17,13 @@ interface SourceHealth { */ const health = new Map(); +/** Services whose most recent calls have been failing right now, for the Alerts page. `alerted` means a "down" notification has gone out for the streak. */ +export function getFailingSources(): { source: string; consecutiveFailures: number; alerted: boolean }[] { + return [...health.entries()] + .filter(([, state]) => state.consecutiveFailures > 0) + .map(([source, state]) => ({ source, consecutiveFailures: state.consecutiveFailures, alerted: state.alerted })); +} + /** Called after every diagnostic-log entry is recorded, to track consecutive failures per source and alert on threshold-cross / recovery. */ export async function trackIntegrationHealth(source: string, ok: boolean): Promise { const state = health.get(source) ?? { consecutiveFailures: 0, alerted: false }; diff --git a/server/src/services/notify.ts b/server/src/services/notify.ts index 5a20bc7..c278e0a 100644 --- a/server/src/services/notify.ts +++ b/server/src/services/notify.ts @@ -382,9 +382,13 @@ const SOURCE_LABELS: Record = { semaphore: "Semaphore", gitea: "Gitea", dockhand: "Dockhand", + uptimekuma: "Uptime Kuma", + phpipam: "phpIPAM", + pbs: "Proxmox Backup Server", + osticket: "osTicket", }; -function sourceLabel(source: string): string { +export function sourceLabel(source: string): string { return SOURCE_LABELS[source] ?? source; } diff --git a/server/src/services/pbsVerificationScheduler.ts b/server/src/services/pbsVerificationScheduler.ts index f566b92..87ef729 100644 --- a/server/src/services/pbsVerificationScheduler.ts +++ b/server/src/services/pbsVerificationScheduler.ts @@ -6,21 +6,39 @@ import { loadIntegrationConfig } from "../integrations/loadIntegration.js"; import { createPbsAdapter } from "../integrations/pbs/adapter.js"; import { notifyPbsVerificationFailed } from "./notify.js"; import { getSettings, getInternalFlag, setInternalFlag } from "./settingsStore.js"; -import { isInMaintenance } from "./maintenance.js"; +import { activeSubjects } from "./maintenance.js"; +import type { SourceFailure } from "./alertTypes.js"; const LAST_RUN_FLAG = "pbsVerificationCheckLastRunDate"; -async function checkPbsVerification(): Promise { +export interface PbsProblem { + integrationId: number; + integrationName: string; + datastore: string; + failedCount: number; + error: string | null; + silenced: boolean; +} + +/** + * Datastores with snapshots that failed verification, or that couldn't be read at all, across the enabled PBS + * integrations. The daily check skips integrations under a maintenance window altogether; the Alerts page passes + * skipSilenced: false to see them, flagged. + */ +export async function collectPbsProblems(opts: { skipSilenced: boolean }): Promise<{ problems: PbsProblem[]; sourceFailures: SourceFailure[] }> { const rows = await db .select({ id: integrations.id, name: integrations.name }) .from(integrations) .where(and(eq(integrations.type, "pbs"), eq(integrations.enabled, true))); - const failures: { integrationName: string; datastore: string; failedCount: number; error: string | null }[] = []; + const failures: PbsProblem[] = []; + const sourceFailures: SourceFailure[] = []; + const silencedSubjects = await activeSubjects(); for (const row of rows) { // A PBS host being worked on can't be reached reliably; this check is daily, so tomorrow's pass covers it. - if (await isInMaintenance("integration", row.id)) continue; + const silenced = silencedSubjects.has(`integration:${row.id}`); + if (silenced && opts.skipSilenced) continue; try { const loaded = await loadIntegrationConfig(row.id); if (!loaded) continue; @@ -29,17 +47,22 @@ async function checkPbsVerification(): Promise { for (const d of datastores) { if (d.error) { - failures.push({ integrationName: row.name, datastore: d.name, failedCount: 0, error: d.error }); + failures.push({ integrationId: row.id, integrationName: row.name, datastore: d.name, failedCount: 0, error: d.error, silenced }); } else if (d.failedCount > 0) { - failures.push({ integrationName: row.name, datastore: d.name, failedCount: d.failedCount, error: null }); + failures.push({ integrationId: row.id, integrationName: row.name, datastore: d.name, failedCount: d.failedCount, error: null, silenced }); } } } catch (err) { console.error(`[pbsVerification] check failed for integration ${row.id}:`, err); + sourceFailures.push({ integrationId: row.id, integrationName: row.name, message: err instanceof Error ? err.message : String(err) }); } } - await notifyPbsVerificationFailed(failures); + return { problems: failures, sourceFailures }; +} + +async function checkPbsVerification(): Promise { + await notifyPbsVerificationFailed((await collectPbsProblems({ skipSilenced: true })).problems); } async function checkPbsVerificationOnce(): Promise { diff --git a/server/src/services/proxmoxBackupScheduler.ts b/server/src/services/proxmoxBackupScheduler.ts index 982f323..2ccf657 100644 --- a/server/src/services/proxmoxBackupScheduler.ts +++ b/server/src/services/proxmoxBackupScheduler.ts @@ -6,22 +6,50 @@ import { loadIntegrationConfig } from "../integrations/loadIntegration.js"; import { createProxmoxAdapter, guestsWithoutBackupCoverage, type ProxmoxBackupTask } from "../integrations/proxmox/adapter.js"; import { notifyProxmoxBackupFailure, notifyProxmoxUncoveredGuests } from "./notify.js"; import { getSettings, getInternalFlag, setInternalFlag } from "./settingsStore.js"; -import { isInMaintenance } from "./maintenance.js"; +import { activeSubjects } from "./maintenance.js"; +import type { SourceFailure } from "./alertTypes.js"; const LAST_RUN_FLAG = "proxmoxBackupCheckLastRunDate"; -async function checkProxmoxBackups(): Promise { +export interface BackupFailure { + integrationId: number; + integrationName: string; + node: string; + guestId: string | null; + status: string; + silenced: boolean; +} + +export interface UncoveredGuest { + integrationId: number; + integrationName: string; + guestName: string; + vmid: number; + node: string; + silenced: boolean; +} + +/** + * Failed latest backups and guests no backup job covers, across the enabled Proxmox integrations. The daily check skips + * integrations under a maintenance window altogether (a host being worked on can't run or report backups, and tomorrow's + * pass covers it); the Alerts page passes skipSilenced: false to see them, flagged. + */ +export async function collectProxmoxBackupProblems(opts: { + skipSilenced: boolean; +}): Promise<{ failures: BackupFailure[]; uncovered: UncoveredGuest[]; sourceFailures: SourceFailure[] }> { const rows = await db .select({ id: integrations.id, name: integrations.name }) .from(integrations) .where(and(eq(integrations.type, "proxmox"), eq(integrations.enabled, true))); - const failures: { integrationName: string; node: string; guestId: string | null; status: string }[] = []; - const uncovered: { integrationName: string; guestName: string; vmid: number; node: string }[] = []; + const failures: BackupFailure[] = []; + const uncovered: UncoveredGuest[] = []; + const sourceFailures: SourceFailure[] = []; + const silencedSubjects = await activeSubjects(); for (const row of rows) { - // A host being worked on can't run or report backups; this check is daily, so tomorrow's pass covers it. - if (await isInMaintenance("integration", row.id)) continue; + const silenced = silencedSubjects.has(`integration:${row.id}`); + if (silenced && opts.skipSilenced) continue; try { const loaded = await loadIntegrationConfig(row.id); if (!loaded) continue; @@ -44,18 +72,24 @@ async function checkProxmoxBackups(): Promise { for (const [node, task] of latestByNode) { if (!task.ok && task.status !== "running") { - failures.push({ integrationName: row.name, node, guestId: task.guestId, status: task.status }); + failures.push({ integrationId: row.id, integrationName: row.name, node, guestId: task.guestId, status: task.status, silenced }); } } for (const g of guestsWithoutBackupCoverage(guests, jobs)) { - uncovered.push({ integrationName: row.name, guestName: g.name, vmid: g.vmid, node: g.node }); + uncovered.push({ integrationId: row.id, integrationName: row.name, guestName: g.name, vmid: g.vmid, node: g.node, silenced }); } } catch (err) { console.error(`[proxmoxBackup] check failed for integration ${row.id}:`, err); + sourceFailures.push({ integrationId: row.id, integrationName: row.name, message: err instanceof Error ? err.message : String(err) }); } } + return { failures, uncovered, sourceFailures }; +} + +async function checkProxmoxBackups(): Promise { + const { failures, uncovered } = await collectProxmoxBackupProblems({ skipSilenced: true }); await notifyProxmoxBackupFailure(failures); await notifyProxmoxUncoveredGuests(uncovered); } diff --git a/server/src/services/tailscaleKeyExpiryScheduler.ts b/server/src/services/tailscaleKeyExpiryScheduler.ts index 142fd23..a11fe20 100644 --- a/server/src/services/tailscaleKeyExpiryScheduler.ts +++ b/server/src/services/tailscaleKeyExpiryScheduler.ts @@ -5,17 +5,27 @@ import { integrations } from "../db/schema.js"; import { loadIntegrationConfig } from "../integrations/loadIntegration.js"; import { createTailscaleAdapter, isKeyExpiringSoon } from "../integrations/tailscale/adapter.js"; import { notifyTailscaleKeyExpiry } from "./notify.js"; +import type { SourceFailure } from "./alertTypes.js"; import { getSettings, getInternalFlag, setInternalFlag } from "./settingsStore.js"; const LAST_RUN_FLAG = "tailscaleKeyCheckLastRunDate"; -async function checkTailscaleKeyExpiry(): Promise { +export interface TailscaleKeyExpiry { + integrationId: number; + integrationName: string; + deviceLabel: string; + daysLeft: number; +} + +/** Every device key that's expired or about to, across the enabled Tailscale integrations. Shared with the Alerts page. */ +export async function collectTailscaleKeyExpiries(): Promise<{ items: TailscaleKeyExpiry[]; failures: SourceFailure[] }> { const rows = await db .select({ id: integrations.id, name: integrations.name }) .from(integrations) .where(and(eq(integrations.type, "tailscale"), eq(integrations.enabled, true))); - const expiring: { integrationName: string; deviceLabel: string; daysLeft: number }[] = []; + const expiring: TailscaleKeyExpiry[] = []; + const failures: SourceFailure[] = []; const now = Date.now(); for (const row of rows) { @@ -27,14 +37,19 @@ async function checkTailscaleKeyExpiry(): Promise { for (const d of devices) { if (!isKeyExpiringSoon(d, now)) continue; const daysLeft = Math.floor((new Date(d.keyExpiry!).getTime() - now) / 86_400_000); - expiring.push({ integrationName: row.name, deviceLabel: d.label || d.hostname, daysLeft }); + expiring.push({ integrationId: row.id, integrationName: row.name, deviceLabel: d.label || d.hostname, daysLeft }); } } catch (err) { console.error(`[tailscaleKeyExpiry] check failed for integration ${row.id}:`, err); + failures.push({ integrationId: row.id, integrationName: row.name, message: err instanceof Error ? err.message : String(err) }); } } - await notifyTailscaleKeyExpiry(expiring); + return { items: expiring, failures }; +} + +async function checkTailscaleKeyExpiry(): Promise { + await notifyTailscaleKeyExpiry((await collectTailscaleKeyExpiries()).items); } async function checkTailscaleKeyExpiryOnce(): Promise { diff --git a/web/src/App.tsx b/web/src/App.tsx index 341c744..e1787d9 100644 --- a/web/src/App.tsx +++ b/web/src/App.tsx @@ -24,6 +24,7 @@ import UptimeKuma from "./pages/UptimeKuma"; import PbsBackup from "./pages/PbsBackup"; import OsTicket from "./pages/OsTicket"; import AdminLinks from "./pages/AdminLinks"; +import Alerts from "./pages/Alerts"; import Generator from "./pages/Generator"; import Maintenance from "./pages/Maintenance"; import Domains from "./pages/Domains"; @@ -38,8 +39,8 @@ import CacheSettings from "./pages/settings/CacheSettings"; import LogSettings from "./pages/settings/LogSettings"; import BackupSettings from "./pages/settings/BackupSettings"; import AppShell from "./layout/AppShell"; -import { setDateTimeSettings } from "./utils/date"; import DialogHost from "./components/DialogHost"; +import { setDateTimeSettings } from "./utils/date"; import { setPageSize } from "./utils/pageSize"; const roleRank: Record = { viewer: 0, operator: 1, admin: 2 }; @@ -114,6 +115,7 @@ export default function App() { } /> } /> } /> + } /> } /> + ); - } diff --git a/web/src/api/client.ts b/web/src/api/client.ts index 7937f13..afe70c5 100644 --- a/web/src/api/client.ts +++ b/web/src/api/client.ts @@ -397,6 +397,30 @@ export interface ServerDetail { links: ServerLink[]; } +export type AlertSeverity = "critical" | "warning" | "info"; +export type AlertCategory = "offline" | "disk" | "backup" | "updates" | "expiry" | "automation" | "integration" | "monitoring" | "tickets"; + +export interface AlertItem { + id: string; + severity: AlertSeverity; + category: AlertCategory; + source: string; + message: string; + link: string | null; + /** Under a maintenance window: still a real problem, but its notifications are held back. */ + silenced: boolean; +} + +export interface AlertsReport { + alerts: AlertItem[]; + counts: { critical: number; warning: number; info: number; silenced: number }; + couldntCheck: { name: string; error: string; affects: string[] }[]; + notes: string[]; + generatedAt: string; + /** This is a recent result being reused, not a fresh check. */ + cached: boolean; +} + export interface AdminLink { id: number; serverId: number; @@ -1006,6 +1030,9 @@ export const api = { syncProxmox: () => request("/api/ipam/sync-proxmox", { method: "POST" }), syncPhpIpam: () => request("/api/ipam/sync-phpipam", { method: "POST" }), }, + alerts: { + list: (refresh = false) => request(`/api/alerts${refresh ? "?refresh=1" : ""}`), + }, ports: { agent: () => request("/api/ports/agent"), forwards: { diff --git a/web/src/layout/AppShell.tsx b/web/src/layout/AppShell.tsx index b98507e..b10a929 100644 --- a/web/src/layout/AppShell.tsx +++ b/web/src/layout/AppShell.tsx @@ -32,6 +32,7 @@ import { IconTicket, IconPlug, IconExternalLink, + IconAlertTriangle, } from "@tabler/icons-react"; import { api, type CurrentUser, type MaintenanceWindow } from "../api/client"; import { formatRemaining } from "../utils/duration"; @@ -101,6 +102,7 @@ const NAV: NavEntry[] = [ label: "Operations", icon: , items: [ + { to: "/alerts", label: "Alerts", icon: }, { to: "/maintenance", label: "Maintenance", icon: }, { to: "/uptime-kuma", label: "Uptime Kuma", icon: }, { to: "/osticket", label: "osTicket", icon: }, diff --git a/web/src/pages/Alerts.tsx b/web/src/pages/Alerts.tsx new file mode 100644 index 0000000..7e2b061 --- /dev/null +++ b/web/src/pages/Alerts.tsx @@ -0,0 +1,245 @@ +import { useEffect, useMemo, useState } from "react"; +import { Link } from "react-router-dom"; +import { api, type AlertCategory, type AlertItem, type AlertSeverity, type AlertsReport, type CurrentUser } from "../api/client"; +import { useSortable } from "../hooks/useSortable"; +import SortableTh from "../components/SortableTh"; +import { usePagination } from "../hooks/usePagination"; +import Pagination from "../components/Pagination"; +import { downloadCsv } from "../utils/csv"; +import { formatDateTime } from "../utils/date"; +import { formatAgo } from "../utils/duration"; +import { readableError } from "../utils/errors"; + +const SEVERITY: Record = { + critical: { label: "Critical", badge: "bg-red-lt text-red", rank: 0 }, + warning: { label: "Warning", badge: "bg-yellow-lt text-yellow", rank: 1 }, + info: { label: "Info", badge: "bg-blue-lt text-blue", rank: 2 }, +}; + +const CATEGORY_LABELS: Record = { + offline: "Server down", + disk: "Disk & storage", + backup: "Backups", + updates: "Updates", + expiry: "Expiry", + automation: "Automation", + integration: "Integration failing", + monitoring: "Monitoring", + tickets: "Tickets", +}; + +type Row = AlertItem & { rank: number }; + +export default function Alerts(_props: { user: CurrentUser }) { + const [report, setReport] = useState(null); + const [error, setError] = useState(null); + const [loading, setLoading] = useState(false); + const [severity, setSeverity] = useState<"all" | AlertSeverity>("all"); + const [category, setCategory] = useState<"all" | AlertCategory>("all"); + const [search, setSearch] = useState(""); + const [showSilenced, setShowSilenced] = useState(true); + + function load(refresh = false) { + setLoading(true); + return api.alerts + .list(refresh) + .then((res) => { + setReport(res); + setError(null); + }) + .catch((err) => setError(readableError(err))) + .finally(() => setLoading(false)); + } + + useEffect(() => { + void load(); + }, []); + + const rows: Row[] = useMemo(() => { + if (!report) return []; + const q = search.trim().toLowerCase(); + return report.alerts + .filter((a) => (severity === "all" || a.severity === severity) && (category === "all" || a.category === category) && (showSilenced || !a.silenced)) + .filter((a) => !q || [a.message, a.source, CATEGORY_LABELS[a.category]].some((v) => v.toLowerCase().includes(q))) + .map((a) => ({ ...a, rank: SEVERITY[a.severity].rank })); + }, [report, severity, category, search, showSilenced]); + + // No initial sort: the server already puts active problems first, worst first. + const { sorted, sortKey, sortDir, requestSort } = useSortable(rows); + const { pageItems, page, setPage, pageCount, totalCount } = usePagination(sorted); + + const presentCategories = useMemo(() => [...new Set((report?.alerts ?? []).map((a) => a.category))], [report]); + + function exportCsv() { + downloadCsv( + "alerts.csv", + ["Severity", "Category", "Source", "Alert", "Silenced"], + (sorted ?? []).map((a) => [SEVERITY[a.severity].label, CATEGORY_LABELS[a.category], a.source, a.message, a.silenced ? "yes" : "no"]), + ); + } + + const counts = report?.counts; + const nothingFound = report !== null && report.alerts.length === 0; + const cards: { label: string; value: number | undefined; color: string }[] = [ + { label: "Critical", value: counts?.critical, color: "text-red" }, + { label: "Warnings", value: counts?.warning, color: "text-yellow" }, + { label: "Info", value: counts?.info, color: "text-blue" }, + { label: "Silenced (maintenance)", value: counts?.silenced, color: "text-secondary" }, + ]; + + return ( + <> +
+

Alerts

+
+ + +
+
+
+ What's wrong right now across your servers and integrations — full disks, servers that stopped reporting, failed backups, updates waiting, + things about to expire. It runs the same checks that send notifications, whether or not those notifications are switched on. + {report && ( + <> + {" "} + + Checked {formatAgo(report.generatedAt)} + {report.cached ? " (a recent result, reused)" : ""}. + + + )} +
+ {error &&
{error}
} + + {!report && !error && ( +
+
Checking everything — this can take a few seconds…
+
+ )} + + {report && ( + <> +
+ {cards.map((c) => ( +
+
+
+
{c.label}
+
{c.value ?? "—"}
+
+
+
+ ))} +
+ + {report.couldntCheck.length > 0 && ( +
+
+
Couldn't check everything — what's missing below isn't necessarily fine
+
    + {report.couldntCheck.map((c) => ( +
  • + {c.name} — {c.error} + {c.affects.length > 0 && (affects: {c.affects.join(", ")})} +
  • + ))} +
+
+
+ )} + {report.notes.map((n) => ( +
+ {n} +
+ ))} + + {nothingFound ? ( +
+
+ {report.couldntCheck.length === 0 ? ( + <> +
All clear
+
Nothing needs attention right now.
+ + ) : ( +
No problems found in what could be checked — see the warning above for what couldn't.
+ )} +
+
+ ) : ( +
+
+
+ setSearch(e.target.value)} /> + + + {(counts?.silenced ?? 0) > 0 && ( + + )} +
+
+
+ + + + label="Severity" sortKeyName="rank" activeKey={sortKey} direction={sortDir} onSort={requestSort} /> + label="Kind" sortKeyName="category" activeKey={sortKey} direction={sortDir} onSort={requestSort} /> + label="Source" sortKeyName="source" activeKey={sortKey} direction={sortDir} onSort={requestSort} /> + label="Alert" sortKeyName="message" activeKey={sortKey} direction={sortDir} onSort={requestSort} /> + + + + {(pageItems ?? []).map((a) => ( + + + + + + + ))} + {(sorted ?? []).length === 0 && ( + + + + )} + +
+ {SEVERITY[a.severity].label} + {CATEGORY_LABELS[a.category]}{a.link ? {a.source} : a.source} + {a.message} + {a.silenced && ( + + silenced + + )} +
+ Nothing matches these filters. +
+
+ +
+ )} + + )} + + ); +}